What is a KV cache, and why does it matter for long-context concurrency?

Instruction: Explain cached attention state, memory growth, and the limits of prefix reuse.

Context: Distinguish transformer KV state from cached answers and estimate its memory cost.

Updated

Example Answer

A KV cache stores attention keys and values from earlier tokens so generation does not recompute them at every step. It helps decoding, but the stored state grows with context length and active sequences. That can limit concurrency even when the model weights fit in memory. I’d estimate memory using the model’s KV heads, layers, head size, token count, and data type, then measure allocation and eviction under the real workload. Prefix reuse can avoid repeated prefill work for compatible shared prefixes; it is different from returning a previously generated answer.

Small calculation

For one sequence with uniform layers, a simple uncompressed estimate is:

2 × layers × KV heads × head dimension × tokens × bytes per value

With 32 layers, eight KV heads, dimension 128, 4,096 tokens, and two-byte values, the estimate is 536,870,912 bytes, or 0.5 GiB. Allocator overhead, quantization, sharding, and model-specific attention change the real amount. Preserve isolation when deciding whether prefix state can be shared.

Follow-up to practice

How could longer prompts reduce throughput even when output length stays fixed?

Reference

vLLM: prefix caching and KV-state reuse

Your preparation path

Work through these questions in order. Read the answer aloud, then explain it in your own words.

Related Questions