Distributed Reliability · Principal
Kafka acknowledged the event. Why did it vanish after the leader failed?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
With acks=1, the leader can acknowledge a write before a follower has replicated it. If that leader fails first, the acknowledged event can be absent after failover. For an event that starts a training-data ingestion job or records an agent's completed action, I would inspect producer acks, idempotence, retries, topic replication factor, min.insync.replicas, ISR size at the time, leader epoch and whether unclean leader election is permitted. acks=all with a meaningful minimum ISR strengthens the normal single-failure case, but writing to a topic with one in-sync replica cannot magically tolerate losing that only copy. Availability and durability will conflict when too many replicas are unavailable. Choose and document whether writes should fail in that condition.
Reproduce in a test cluster by pausing replication, publishing with acks=1, confirming the producer callback, then stopping the leader before the follower catches up. Compare after failover with the acks=all configuration under its required ISR. Record the partition and offset, not just the producer's generic success log. If the event was lost, reconcile the authoritative source state and republish with a stable event ID. Producer idempotence helps with retry duplicates to Kafka. It does not turn a leader-only acknowledgment into a replicated commit or make downstream side effects exactly once.
A later interviewer push might be, “But the application wrote the row to Postgres, so can we declare success?” The database row and Kafka event need a separate publication protocol, such as an outbox, if both must survive together. Kafka says exactly once. Why did the search index apply the document twice? examines Kafka's exactly-once processing boundary at an external search index. This question is earlier in the path: whether the original producer's acknowledged event survived Kafka leader failover at all.
Continue reading
Related questions
Read beyond the question
Explore more distributed reliability
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →