Agent Architecture · Principal
The workflow replayed after a crash. Why did the agent send the email again?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
In a history-replay workflow engine, workflow code can run again to rebuild its state from recorded events. The workflow should schedule an email activity, not call the email API directly as an unrecorded side effect while calculating the next step. If a direct call sends the email and the worker crashes before recording subsequent workflow progress, replay may run that call again. Some SDKs restrict that kind of I/O in workflow code, but a wrapper or poorly isolated helper can still put the side effect in the wrong place. Temporal's workflow definition requires deterministic replay, and its AI agent reference architecture places nondeterministic external I/O in activities.
Move the provider call into a tracked activity with a stable operation ID. A recorded completed activity is represented by its result when workflow history replays, so replay of the workflow logic does not by itself resend it. Do not infer that activities are exactly once. An activity can also be retried after a timeout or worker crash with an unknown provider outcome. The email or action endpoint needs an idempotency key, a sent-message lookup or another explicit reconciliation path. Use the same stable business action identity across attempts, not a newly generated key on each retry.
I would induce a worker crash at three moments: before scheduling, after the provider accepts the message but before activity completion is recorded, and after completion is in history. Compare provider message IDs, activity attempt IDs and workflow event history. The first should lead to one eventual send, the second may retry and must be deduplicated or reconciled, and the third must not send again on workflow replay. Also test a code deployment against a saved history for deterministic command order.
The refund succeeded but the agent crashed. What happens next? asks how to recover an ambiguous refund after an agent crash. How do you deploy new agent workflow code while 200,000 runs are waiting? deals with deploying workflow changes while many runs wait. This case is a code-placement error: a developer put a real-world side effect in code that is allowed to replay. Knowing exactly which line crosses that boundary is what prevents a durable agent from becoming a duplicate-action machine.
Continue reading
Related questions
Read beyond the question
Explore more agent architecture
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →