Distributed Reliability · Principal
We gave each agent its own queue. Why did one slow tool still stall everyone?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Separate queues isolate waiting jobs at the front door. They do not isolate a connection pool that every worker uses after dequeue. Imagine investigation jobs and customer-facing actions with separate queues and worker counts, but both call the same tool gateway through a pool of 40 connections. If the investigation tool hangs while holding all 40, a customer action can be admitted, get a worker and then wait at the pool. Queue depth looks fine while end-to-end latency explodes.
Trace a request across the actual scarce resources: worker permit, outbound connection, tool service concurrency, downstream database pool and provider quota. Instrument wait time and occupancy at each boundary, tagged by workload, plus in-flight calls and their timeouts. A stalled connection may keep its slot long after the user's deadline if cancellation is not propagated. AWS's agent resource-isolation guidance calls out starvation in shared pools. The specific pool in this scenario is ours to find.
Give the critical path a reserved amount of capacity at the first shared bottleneck that actually matters. Set concurrency limits before acquiring downstream connections, with a bounded queue and a deadline. Keep an upper bound on slow investigations so they cannot consume the reserve. If the downstream provider itself has one global quota, splitting our local pools will not create provider capacity. Then we need admission control against that quota and a clear priority or degradation policy. An endlessly waiting high-priority request is not protected just because it owns an empty queue.
Would I isolate every tenant and agent in a dedicated deployment? Usually that wastes capacity and can still share the same provider. Start with measured workload classes, reserve for critical work, and load-test a slow tool while ordinary traffic runs. The invariant is that noncritical work cannot occupy all permits required for a critical request, and that the critical request has a bounded failure mode when the dependency itself is down. One agent run opens a hundred tool calls at once deals with one run fanning out too many calls. This one tests whether our supposed bulkhead reaches the real shared bottleneck.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →