First separate two clocks. Time to first token includes admission, queue time and processing the prompt. Inter-token latency measures progress after generation starts. Prefill processes many prompt tokens in parallel, while decode repeatedly adds one token per active sequence and reads its KV state. They compete for compute and memory bandwidth in different ways. vLLM's optimization guide says chunked prefill can prioritize decode work and use remaining token budget for prefill. That policy can improve ongoing streams while starving new work if decodes keep consuming the schedule or a long prefill receives too small a residual share.

I would look at queue age by prompt length, allocated prefill tokens per scheduling iteration, prefill chunks until first token, active decode count, KV capacity, and completed requests by class. A mean TTFT across completed requests is dangerous because the worst queued requests may not have completed yet. Include timed-out and cancelled work. Replay a realistic mix of short prompts, long documents and lengthy generations, not a homogeneous benchmark. When can continuous batching improve throughput and make a short request slower? asks about the broad throughput versus latency tradeoff of continuous batching. Here the question is whether a specific scheduler gives a long prefill a bounded route to service while preserving decode SLOs.

If decode reservation takes the entire budget at peak, I would put a measured floor or age-based share on prefill, or isolate long prefill capacity when the workload justifies it. The exact mechanism depends on the engine. Merely raising the maximum batched-token number can change memory and latency pressure, and admitting every waiting long prompt can crush running streams. Bound the number of concurrent long prefills, estimate KV occupancy before admitting, and test against the burst distribution. If prefill and decode are split across pools, the handoff and KV transfer need their own budget and queue metrics.

What if the business says active chats must always win? Then make that policy explicit, reserve capacity for them and publish an honest first-token SLO for long requests or reject them early. A system that silently queues a paid request beyond its deadline has made an admission mistake. The goal is not equal latency for unlike requests. It is a deliberate service policy that accounts for everyone waiting, including requests that never got a first token.