If thousands of workflows fail at the same time and use the same deterministic delay, doubling that delay keeps them synchronized. A recovery endpoint sees bursts, not a smooth arrival rate. AWS's retry guidance describes exponential backoff with jitter to spread requests. That helps, but the system also needs a shared admission budget. Fifty thousand individually well-behaved agents can still exceed a provider quota when all are due within a broad window.

I would inspect attempts per original task, wake-up histogram, 429 and 5xx responses, provider retry-after hints, token and request quotas, and in-flight concurrency across regions. Count nested retries. A workflow may retry an activity, the model gateway may retry the HTTP call, and the SDK may retry again. Three layers each allowing three attempts can make one logical step produce many more provider calls than an operator expects. Bound the total attempt and elapsed-time budget at the operation boundary, and make the gateway the owner of provider retry policy where possible.

On recovery, put queued work through a central token or concurrency limiter that reflects the provider's actual capacity and separates urgent interactive requests from background agent steps. Use randomized jitter with a cap, honor retry-after when it is meaningful, and let queue age influence scheduling. If the provider is still unhealthy, fail fast or keep tasks parked rather than spending scarce slots on guaranteed 429s. As capacity returns, increase dispatch gradually and measure successful task completions, not request attempts. Durable jobs need an end state when their deadline has passed. Do not replay a model request if the result has already triggered an external action without checking that action's idempotency boundary.

The interviewer may ask whether a circuit breaker is enough. It can stop calls while the provider is failing, but an unguarded half-open phase is another synchronized release. Give probes a small budget and ramp based on observed success. The queue lease expires while an agent tool call is running is about an individual queue lease expiring during a tool call. This question is about fleet-wide synchronization and retry amplification after an upstream outage. The useful invariant is that recovery load must be controlled by available capacity, not by the number of agents that happened to wake.