The service moves prefill to one GPU pool and decode to another. Both pools have free compute, and prefill itself gets faster. Yet time to first token gets worse for long prompts. The KV built during prefill now has to reach the decode worker before that worker can produce the next token.

How much data is that? For a conventional full-attention cache, a rough uncompressed estimate is 2 × layers × KV heads × head dimension × cached tokens × bytes per element. The two is for keys and values. Actual bytes depend on cache dtype, attention architecture, which layers retain which positions, transfer granularity and whether some prefix is already local. A million-token advertisement does not mean every request transfers a million tokens, but long prompts can turn this handoff into a material network operation. P/D-Serve includes KV transfer in its disaggregated serving analysis, and NetKV studies how network topology and congestion affect decode placement.

I would break first-token latency into queue time for prefill, prefill execution, KV serialization or registration, transfer wait, actual transfer, decode admission and first decode step. Compare sizes and network paths by request, not fleet-average GPU utilization. A fast prefill worker in one rack paired with a lightly loaded decoder across a congested link may be worse than a busier decoder with local KV. Also check whether the decoder reserves capacity while waiting for KV and whether backpressure allows transfers to pile up. A bandwidth estimate from bytes divided by nominal link rate is only a lower bound when transfers contend.

The scheduler's decision is joint. Choose a prefill worker and decoder based on compute queue, compatible model revision, KV locality, network cost and the request's latency target. For short prompts, local prefill plus decode may beat disaggregation because it avoids the handoff. For long prompts or heavy prefill contention, separation can still win. Cache transfer compression or lower-precision KV can save bytes, but may add compute or change quality, so measure completed-request latency and accuracy together. Do not just optimize transfer microbenchmarks.

If the interviewer asks whether to buy more decode GPUs, I would first inspect the bottleneck. More decode workers cannot fix a saturated shared link and may create more cross-rack traffic. Run a mixed prompt-length load with realistic concurrency, record transfer bytes and p95/p99 per stage, and compare colocated versus disaggregated paths under the same arrival process. Should prefill and decode run on different GPUs? asks whether to separate prefill and decode. This page asks why a separation that looks good on GPU metrics can lose on the KV handoff.