Distributed Reliability · Staff
The model emits a token in 200 ms. Why does the browser see nothing for five seconds?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First clarify which clock the 200 ms number stops. It may mean the inference worker produced its first token, not that the API flushed an event and the browser received it. In one common failure, the application emits server-sent events but a reverse proxy buffers the upstream response. It accumulates a chunk and forwards it later. Nginx's proxy module documentation describes the buffering mode and the option to disable it. The mechanism depends on the deployed proxy path, so trace that path rather than asserting Nginx caused every delay.
Put timestamps at model token ready, application stream write, proxy ingress and egress where measurable, and browser event receipt. Send a controlled response with one small event each second. Compare the direct application endpoint with the public route. If direct delivery is immediate and the public route batches events, inspect proxy buffering, compression and any CDN or gateway behavior. A write() call returning does not prove a browser has received bytes.
For the streaming route, configure the relevant intermediaries to flush small events promptly and avoid a compression or transformation layer that coalesces them. Keep heartbeat and idle timeout behavior explicit. Then retest through the real production path on a slow network. Disabling buffering can increase live connections and change backpressure, so reserve capacity and monitor proxy and application memory and connection lifetime. A faster first visible token is not useful if 10,000 slow clients exhaust the service.
If the interviewer asks whether this is the same as a slow reader, no. The client reads tokens slowly. Where does the generated text pile up? starts with a client that cannot consume generated data and asks where it accumulates. Here the browser is ready, but an intermediary withholds data. The two can interact, so a full stream test should cover both. Measure user-visible time to first event, not only worker time to first token.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →