Distributed Reliability · Principal
The queue has a deduplication ID. Why did a retry six minutes later create two jobs?
The question
Interview question
A producer submits an agent job to a FIFO queue with a deduplication ID. The send succeeds, but its acknowledgement is lost. Recovery retries the same send after six minutes. The producer thinks it is safe because it reused the same ID. Two jobs run and both try to create a customer credit. Which guarantee ended?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The queue's producer-side deduplication window. Amazon SQS documents a five-minute deduplication interval for FIFO messages and warns that a send retry after that interval can introduce duplicates. That queue behavior says nothing about indefinite, end-to-end exactly-once business execution. In this hypothetical workflow, the same logical operation entered twice because the producer's uncertainty lasted longer than the queue's memory of the first send.
Use one durable business operation ID assigned before enqueue. At the consumer, claim or look up that operation in a transactional state store before doing the work. Both messages can be delivered, but only one logical credit should advance. The claim must be safe under concurrent consumers, not a read-then-write check with a race. If the external credit service supports idempotency keys, pass the same business ID there. If it does not, reconcile using the service's operation ID or a business ledger before retrying an unknown outcome. Queue message ID, deduplication ID and business operation ID serve different scopes.
Now the interviewer asks about an outbox. An outbox can make recording business intent and enqueue work reliable relative to one database transaction, but its dispatcher can still publish twice after losing an acknowledgement. That is fine if consumers and the external side effect are idempotent at the business boundary. Track created, queued, executing, unknown, completed and failure states with evidence, and make duplicate delivery a tested normal path. Do not say “exactly once” without naming which state transition has that property.
Test a send acknowledgement lost at minute zero, a retry inside the queue window, one outside it, and two consumers racing. Then test a credit that succeeded but whose acknowledgement was lost. The refund succeeded but the agent crashed. What happens next? focuses on that consumer-side unknown refund outcome. Here the duplicate begins earlier, at a delayed producer-side resend that outlives transport deduplication.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →