Agent Architecture · Principal
An old tool callback arrived after the agent restarted. Which run gets the result?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The callback belongs to the exact operation attempt that launched it, not whichever run currently owns the same business task ID. Imagine an agent requests a document export, times out and is cancelled. The user starts a new run for the same case with different scope. Ten minutes later the export service calls back with the old result. A handler that routes only by case ID can attach the stale export to the new run and let it influence an answer or action that the new user never authorized.
Give each pending operation a durable correlation identity tied to workflow or run identity, step, attempt, tool request ID, and authorization context. Record its expected state before launching the external operation. On callback, verify authenticity, match that exact pending operation, and atomically move it from pending to completed only if the run still accepts it. A duplicate callback becomes an idempotent no-op. A callback for a cancelled or superseded run is recorded as late and cannot be transplanted into the new one. Temporal's activity execution documentation describes asynchronous activity completion as distinct from simply returning from the activity. Its mechanism illustrates why an external completion needs its own tracked execution identity. The application still owns its business authorization and result-use rules.
There is a harder case when the external operation itself had a side effect. Ignoring a late callback does not undo a file export, payment or email. Query the provider by the original operation key, reconcile what happened, and decide whether a compensating action is valid. Keep that reconciliation attached to the old attempt. The new run may be informed of the prior real-world state through an explicit, authorized read, but it should never silently inherit the old callback as its own completed step.
I would race cancel, restart, duplicate callback and callback-before-registration in tests, including two attempts against the same case. The provider fails after a tool result. What can the next model actually continue? asks what context can survive a model-provider failure after a tool result. The user changes their mind while the agent's tool call is in flight. What stops? addresses cancelling an in-flight tool. This case is the arrival path after that boundary: correct routing and fencing of a result when a newer run exists.
Continue reading
Related questions
Read beyond the question
Explore more agent architecture
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →