What is a KV cache, and why does it matter for long-context concurrency?
Instruction: Explain cached attention state, memory growth, and the limits of prefix reuse.
Updated
Example Answer
A KV cache stores attention keys and values from earlier tokens so generation does not recompute them at every step. It helps decoding, but the stored state grows with context length and active sequences. That can limit concurrency even when the model weights fit in memory. I’d estimate memory using the model’s KV heads, layers, head size, token count, and data type, then measure allocation and eviction under the real workload. Prefix reuse can avoid repeated prefill work for compatible shared prefixes; it is different from returning a previously generated answer.
Small calculation
For one sequence with uniform layers, a simple uncompressed estimate is:
2 × layers × KV heads × head dimension × tokens × bytes per value
With 32 layers, eight KV heads, dimension 128, 4,096 tokens, and two-byte values, the estimate is 536,870,912 bytes, or 0.5 GiB. Allocator overhead, quantization, sharding, and model-specific attention change the real amount. Preserve isolation when deciding whether prefix state can be shared.
Follow-up to practice
How could longer prompts reduce throughput even when output length stays fixed?
Reference
Your preparation path
Work through these questions in order. Read the answer aloud, then explain it in your own words.
1. Start with the foundations
Build the vocabulary and explain the core decisions.
2. Apply it to a real workflow
Practice diagnosis, validation, and everyday tradeoffs.
- What signals matter most when optimizing cost without hiding quality loss? Free sample
- Traffic spikes create queueing even though the model service itself is healthy. How would you investigate? Member answer
- The product is fast on short prompts and unstable on long tool-using requests. What would you inspect? Member answer
3. Prepare for senior discussions
Explain failure boundaries, recovery, and production choices.
Related Questions
-
easy
-
easy
-
easy
-
easy
-
easy
-
easy