Data and Knowledge Systems · Principal
The newer customer update was processed first. Why did the old value win?
The question
Interview question
A customer changes their plan from Basic to Pro, then from Pro to Enterprise. The projection processes Enterprise, then later overwrites it with Pro. The system uses Kafka messages keyed by customer ID, and the team assumed that meant all updates for one customer were ordered. What changed?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The topic's partition count grew. The older keyed event was in one partition, while a producer after the expansion routed the same key to a different partition. Each partition is ordered, but there is no total order between them. Kafka's documentation states the per-partition ordering boundary, and its operations guidance warns that increasing partitions can route the same key differently. In this scenario, independent consumers drain the two partitions at different rates and an unconditional upsert allows the older state to win.
I would inspect topic partition counts over time, producer partitioner versions, message keys, source revisions, partition and offset of both events, and the order of successful writes to the projection. A later arrival is not necessarily a later business state. Source timestamps alone can be unreliable if clocks differ or retries reuse payloads. If the source owns a monotonic per-entity revision or a comparable change-log position, carry it with each event and make the sink update conditional: apply revision 12 only if the stored revision is below 12. The exact ordering field must come from the same source authority for both events. Kafka offsets from different partitions are not directly comparable.
What if we cannot change the producer or source? Freeze the partitioning scheme for the ordered stream, introduce an explicit repartitioned topic by stable key with a controlled cutover, or serialize updates through a stateful authority. But a new topic alone cannot retroactively infer the order between in-flight events from old and new partitions. You need a source version, a drain barrier, or a source-of-truth reconciliation. The projection should expose stale-write rejection and lag metrics, with a repair path for entities already overwritten.
Test a partition expansion while one old update is delayed, then deliver the newer update first. Check the final state and source revision, not just the order in one consumer log. One database transaction updated two tables. Why did the AI report see only half of it? asks whether one source transaction is applied partially across tables. This question is about two valid updates to one entity whose order was lost across partitions during a topology change.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →