Distributed Reliability · Principal
The client reads tokens slowly. Where does the generated text pile up?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
An inference server produces tokens faster than a mobile client can read its stream. Maybe the radio connection is weak, or the browser tab stopped consuming events. The GPU can keep generating while an API gateway buffers chunks in memory. If that queue has no bound, a few thousand slow clients can exhaust the gateway even when GPU utilization looks normal.
Trace each boundary. The model worker emits token events to the gateway, the gateway writes to an HTTP or RPC stream, and the transport eventually applies flow control. Node.js stream documentation says a writable's write() can return false and a producer should wait for drain. gRPC flow-control guidance similarly explains that stream writers can be held back by receiver capacity. Neither mechanism automatically limits an application's separate queue between a GPU worker and the network handler. That queue must have an explicit policy.
I would cap queued bytes or token events per request and measure lag between generation and delivery. When the cap is reached, pause upstream production if the scheduler and model worker support it. If they cannot pause cleanly without holding scarce KV blocks forever, set a delivery deadline and terminate the request with an honest partial-stream status. Release worker and KV resources after cancellation is acknowledged. Coalescing UI events can reduce per-token overhead, but it cannot silently drop generated content if the API promises the complete answer.
The policy also affects billing and audit. Distinguish tokens generated, tokens queued, tokens delivered and tokens acknowledged if the protocol has acknowledgments. A network write completing only means bytes reached a local buffer, not necessarily the user's screen. Do not claim exactly what the user read without an application-level acknowledgment.
To test it, throttle one client to a few bytes per second while other requests run. Watch gateway resident memory, per-request queue depth, active KV lifetime and cancellation latency. A fast-client load test will miss the failure entirely. The right outcome is bounded memory and explicit lifecycle behavior under slow consumption.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →