Distributed Reliability · Principal
A failed agent job was replayed next week. Does its old approval still count?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
An agent was approved to send a supplier payment on Monday. The job hit an integration error and ended up in a dead-letter queue. On Thursday, the user revokes approval. On Friday, an operator redrives the queue and the payment executes. The retry machinery did its job. The business authorization did not.
A queued message is a record of an intent and past decision, not an evergreen permission. SQS documents dead-letter queues and redrive as transport operations. They do not reauthorize application actions. In this hypothetical system, the consumer must distinguish “this job may be delivered again” from “this side effect is still allowed now.”
I would give the job a durable action ID, actor and tenant, target resource, approved operation and limits, approval reference, decision version, deadline, and original creation time. On every attempt, including redrive, the executor checks live revocation and current resource state with the authority that owns them. The approval may have a defined lifetime and scope. If a policy changed, the outcome depends on the product's explicit versioning rule, but an explicit revocation must not be ignored because an old serialized approved=true flag survived. For high-impact actions, a stale or unavailable authorization source should block execution and request review, with a visible reason.
The interviewer may ask, “What if the payment actually went through on Monday and the acknowledgement was lost?” Rechecking approval alone does not solve duplicate side effects. Query the payment ledger using a stable business operation ID or an idempotency key recognized by the downstream service. If it is settled, reconcile and mark the job complete. If it is definitively absent and authorization is still valid, execute. If status is unknown, quarantine rather than blindly issue another payment. The state machine needs to record attempts, external IDs, authorization checks and terminal outcomes.
Redrive should be a controlled operation with a dry-run inventory of age, action type, revocations and unknown side-effect states. Measure how many replayed jobs were rejected by the fresh gate. Test approval revocation between failure and redrive, lost acknowledgement, and a batch containing both safe and unsafe jobs. The agent's task was approved yesterday. The employee loses access today. Does the scheduled action run? concerns a scheduled action after a principal loses access. This question is the operational replay path, where a dead-letter queue can turn a past intent into a new attempt long after the original failure.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →