Why can temperature-zero LLM outputs still change when requests are batched?

Instruction: Explain the difference between greedy token selection and reproducible numerical execution.

Context: Investigate batch-dependent numerical behavior before treating every output difference as sampling randomness.

Updated

Example Answer

Temperature zero usually selects the highest-scoring token, but it does not make the computation producing those scores identical. Different batch shapes or kernels can change floating-point rounding. If two token scores are close, that small difference can change the selected token and the rest of the answer. I’d first confirm that the prompt, model, tokenizer, and generation settings really match, then compare the same request alone and in varied batches on a fixed serving stack. If exact reproducibility matters, I’d test a supported batch-invariant configuration and measure its performance cost. A fixed seed alone is not a guarantee across configurations.

Concrete experiment

Run one prompt repeatedly, then alongside short and long requests in different orders. Record token outputs, engine version, hardware, dtype, and relevant kernel settings. Keep this separate from tests that intentionally sample multiple answers.

Limit to explain

vLLM documents batch invariance as beta with compatibility limits. Validate the exact model and backend; do not promise identical results across upgrades.

Follow-up to practice

When does the product need identical text, and when is stable task correctness sufficient?

References

PyTorch: numerical accuracy and batched computation, vLLM: batch invariance

Your preparation path

Work through these questions in order. Read the answer aloud, then explain it in your own words.

Related Questions