Distributed Reliability · Staff
The gRPC channel is healthy. Why do short tool calls wait behind long streams?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
One channel does not promise unlimited independent RPC capacity. A gRPC channel uses HTTP/2 connections, and each connection can have a limit on concurrent streams. If a service keeps many long-lived model streams open on the same channel used by short authorization or tool calls, new RPCs can wait for stream capacity even while server CPU looks healthy. gRPC's performance guidance calls out per-connection concurrent-stream limits and the possibility of queued RPCs. The exact behavior depends on the client library, server and connection pool.
I would measure time from RPC creation to wire send, then server receive, server processing and return. If a short call waits before the server sees it, changing the tool's handler will not help. Record active and queued streams per connection, channel count, connection limits, stream duration and cancellation. Compare a dedicated channel for short calls against the shared channel under a load test with many long-lived streams. If it works, that points to client-side transport capacity or flow control rather than server application time.
Use separate channel pools for workloads with very different lifetimes when the measured limit warrants it, and cap the number of long-lived streams admitted per pool. Ensure deadlines include queue time, and terminate abandoned streams. Opening arbitrary new channels per request is not the answer because it adds connections, load balancing complexity and resource cost. If the server itself has a global concurrency or quota limit, separate channels cannot manufacture capacity, so measure both sides.
The interviewer might ask whether HTTP/2 multiplexing already solves head-of-line blocking. It lets several streams share one connection, but it still has stream concurrency and flow-control limits. The old statement "one connection means one request" is wrong, and so is "one connection means infinite requests." The client reads tokens slowly. Where does the generated text pile up? concerns a slow client accumulating generated data, and The model emits a token in 200 ms. Why does the browser see nothing for five seconds? concerns a proxy buffering events. Here short calls queue before reaching the server because their shared transport is occupied by long-lived RPC streams.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →