Distributed Reliability · Principal
The agent had a 20-second deadline. Why is its tool still running a minute later?
The question
Interview question
An interactive agent gets a 20-second user deadline. It spends 12 seconds in retrieval and another 7 seconds generating a tool call. The tool gateway then starts a fresh 20-second timeout, its SDK starts another on retry, and the database continues work after the user has left. Which clock should control the chain?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
One end-to-end deadline should bound the user-visible operation. Each hop receives the remaining budget and may choose a smaller local timeout. Resetting the clock at each hop converts a 20-second promise into an unbounded chain of individually legal waits. gRPC's deadline guide describes propagation that subtracts elapsed time rather than blindly forwarding an absolute timestamp across unsynchronized clocks. The implementation details depend on the transport and asynchronous workflow, but the principle holds: queueing, inference, retries and tools all spend from the same budget.
I would trace the original deadline and remaining budget at every boundary, including queues. If only one second remains when the tool call is ready, do not start a ten-second action and hope. Return a partial or timeout response, or transition to a separately authorized background task with a new contract. A local database statement timeout and cancellation signal should stop work that no longer helps, but cancellation is not proof an external side effect did not commit. For non-idempotent actions, reconcile the outcome before deciding whether to retry or tell the user it failed.
The scheduler should decide what can finish with the time left. A fast cached answer may be viable while long retrieval is not. Preserve a margin for streaming a coherent response and recording state. For asynchronous agent runs, there can be a durable job deadline separate from the interactive HTTP deadline. Name both and state which operation each governs. Do not hide background continuation behind a user request that was supposedly cancelled.
What if the tool has its own strict minimum runtime? Then either start it earlier, ask for an asynchronous handoff, or decline to promise a 20-second completion for that route. The user closed the chat. Why are GPUs still finishing their answer? covers GPU work continuing after chat disconnect. A model stream fails after the user has seen half an answer covers a stream failing after half an answer. This question is about time budget being recreated at service boundaries, making downstream work exceed the original promise even while every individual component reports a successful timeout policy.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →