Distributed Reliability · Principal
The Kafka transaction aborted. Why did the agent still act on its event?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
An aborted transactional write is still present in the log. Whether the consumer receives it depends on its isolation level. Kafka consumers configured with read_uncommitted can poll records from aborted transactions. read_committed hides them and stops at the last stable offset while earlier transactions remain open. That is the behavior in Kafka's consumer configuration. If the agent bridge used the default read_uncommitted, it may have scheduled a real action before the producer aborted its batch.
Trace the producer transaction ID, record offset, abort marker, consumer isolation.level, and the action's correlation ID. Do not confuse an agent's acknowledgment of a queue job with the producer transaction's commit. If the source uses transactions to ensure an event and its companion records appear together, every consumer that acts on that promise needs read_committed. The change may also increase apparent consumer lag or latency behind a long-running open transaction. Inspect last stable offset separately from log end offset, and alert on transactions that stay open too long.
Now the unpleasant part: changing isolation level does not undo a refund or an email that already happened. Kafka can keep aborted input invisible to a correctly configured reader, but an external side effect has its own commit boundary. Reconcile the action using a stable business ID and decide whether a compensating action is valid. For future actions, verify authorization and relevant state at the external action boundary, then make retries idempotent. There is no single Kafka switch that makes a payment transaction atomic with a Kafka transaction.
An interviewer might say the producer was transactional, so why is the default reader unsafe? Producer guarantees only matter to readers that request committed visibility. Unrelated nontransactional records still pass a read_committed consumer. Kafka says exactly once. Why did the search index apply the document twice? asks why Kafka exactly-once processing does not make an external search index exactly once. This question occurs earlier, when a consumer sees an event from a transaction that never committed at all. The fix starts with read isolation, then deals with the external action already taken.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →