How do you distinguish a prefill bottleneck from a decoding bottleneck?

Instruction: Separate queue delay, prompt processing, and token generation before selecting an optimization.

Context: Diagnose serving latency with workload slices and stage-level measurements.

Updated

Official answer available

Read the opening below, then unlock the full answer and practical guidance.

I’d measure queue time, prefill time, time to first token, decoding time, and output-token latency separately...

Your preparation path

Work through these questions in order. Read the answer aloud, then explain it in your own words.

Related Questions