Why can temperature-zero LLM outputs still change when requests are batched?
Instruction: Explain the difference between greedy token selection and reproducible numerical execution.
Updated
Example Answer
Temperature zero usually selects the highest-scoring token, but it does not make the computation producing those scores identical. Different batch shapes or kernels can change floating-point rounding. If two token scores are close, that small difference can change the selected token and the rest of the answer. I’d first confirm that the prompt, model, tokenizer, and generation settings really match, then compare the same request alone and in varied batches on a fixed serving stack. If exact reproducibility matters, I’d test a supported batch-invariant configuration and measure its performance cost. A fixed seed alone is not a guarantee across configurations.
Concrete experiment
Run one prompt repeatedly, then alongside short and long requests in different orders. Record token outputs, engine version, hardware, dtype, and relevant kernel settings. Keep this separate from tests that intentionally sample multiple answers.
Limit to explain
vLLM documents batch invariance as beta with compatibility limits. Validate the exact model and backend; do not promise identical results across upgrades.
Follow-up to practice
When does the product need identical text, and when is stable task correctness sufficient?
References
PyTorch: numerical accuracy and batched computation, vLLM: batch invariance
Your preparation path
Work through these questions in order. Read the answer aloud, then explain it in your own words.
1. Start with the foundations
Build the vocabulary and explain the core decisions.
2. Apply it to a real workflow
Practice diagnosis, validation, and everyday tradeoffs.
- What signals matter most when optimizing cost without hiding quality loss? Free sample
- Traffic spikes create queueing even though the model service itself is healthy. How would you investigate? Member answer
- The product is fast on short prompts and unstable on long tool-using requests. What would you inspect? Member answer
3. Prepare for senior discussions
Explain failure boundaries, recovery, and production choices.
Related Questions
-
easy
-
easy
-
easy
-
easy
-
easy
-
easy