Distributed Reliability · Staff
The user closed the chat. Why are GPUs still finishing their answer?
The question
Interview question
Users close a streaming chat after a few seconds. The edge server stops sending bytes, but the gateway keeps the model request alive and the GPU generates to the maximum output limit. During a traffic spike, abandoned work fills decode slots and raises p99 for users who stayed. Where would you trace and stop it?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The lifetime of the client connection and the lifetime of the upstream computation are separate until the application joins them. The edge can detect a disconnect while a gateway, scheduler or provider call has no cancellation signal, or ignores one. gRPC's cancellation guide describes how a client cancellation is conveyed to a server and why servers need to stop their own and child work. The exact behavior of an HTTP stream and a model provider varies. We must verify every hop rather than assuming socket closure frees the GPU.
I would trace one run ID through client, edge, model gateway, serving engine, scheduler and KV allocator. Measure elapsed time and generated tokens after disconnect. Did the edge cancel its upstream request? Was the gateway still waiting in a queue, already prefilling, or decoding? Did the serving engine remove the request from the batch and reclaim KV blocks? A server may mark the request canceled yet keep buffers until an asynchronous cleanup path executes. Requests time out, but GPU memory stays high. Who still owns the KV blocks? covers KV ownership after a timeout. Here the trigger is the client leaving and the immediate fleet-level wasted work.
Set a clear cancellation policy. For an interactive answer with no durable background purpose, propagate a cancellation signal and deadline through the call tree, stop scheduling new tokens, and reclaim resources once the engine confirms termination. If the answer is intentionally a background job, make that an explicit product state with an owner and retention policy. A client disconnect alone should not silently decide whether an externally consequential action continues. For a read-only model answer it is reasonable to stop. For a payment or email already submitted, cancellation cannot undo it, so track its committed outcome separately. Do not conflate stopping token generation with rolling back tools.
There can be races. The user may reconnect to the same run, a stream may finish while the disconnect event is traveling, or the provider may not support cancellation after submission. Use an idempotent state transition so cancel and complete resolve to one durable outcome. If cancel is best effort, expose that honestly and bound the maximum wasted work with output caps and deadlines. Load-test abandon rates as well as QPS, then measure active decode slots, tokens generated after disconnect, KV release time and cost per completed user answer. A healthy stream latency chart can hide a fleet doing large amounts of work nobody receives.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →