Distributed Reliability · Principal
The index event arrived, but the document transaction rolled back. What should search believe?
The question
Interview question
An ingestion service sends `document updated` to a queue, then commits the metadata row. The queue publish succeeds, but the database transaction rolls back. The indexer fetches the old document, indexes it under the new revision, and acknowledges the event. Why did an apparently reliable queue make the knowledge base less reliable?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
There were two writes to different systems with no shared commit. Publishing first can announce state that never committed. Committing the database first and publishing second can leave a committed update with no event if the process dies between them. A queue can deliver its own messages reliably and still know nothing about whether the source transaction succeeded. AWS's transactional outbox guidance describes writing the business row and an outbox record in one database transaction, then relaying committed outbox records to the broker.
The index event should identify a committed source revision, not claim that the index has already caught up. The consumer fetches that revision or immutable blob, checks that it exists, and applies it conditionally so an older event cannot overwrite a newer indexed revision. If the outbox relay sends the same event twice, the consumer should be idempotent by document ID and revision. If revisions skip, it can fetch current authoritative state or flag a gap rather than fabricating the missing intermediate document. A delete is also a revision with an explicit removal effect.
I would trace transaction IDs, outbox rows, publish attempts, index acknowledgements and source/index revision watermarks. Make a failure-injection test at every boundary: before commit, after commit before relay, after broker acknowledgement before marking the outbox delivered, and after index write before acknowledging consumption. The desired result is eventual convergence without ghost revisions or a permanently missed update. The relay may publish at least once. That is fine if the consumer has a sound version gate.
What if the source is an external SaaS and we cannot use its database outbox? Then the connector needs a different contract, such as a consistent snapshot plus change cursor and periodic reconciliation against authoritative state. Do not pretend the external event and your local metadata are one transaction. Which code actually consumes a changed event? asks which code consumes a change event, and A bulk snapshot races the change stream asks how a bulk snapshot meets a change stream. This question is the producer-side dual-write boundary that can create a false event before search even sees it.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →