Data and Knowledge Systems · Principal
We deleted a customer record from Cassandra. Why did repair bring it back?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
A delete in Cassandra writes a tombstone. It is a timestamped piece of data telling replicas that an older value is dead. Imagine replicas A and B receive the tombstone, while C is down and still holds the old value. If the tombstone remains, repair can bring C up to date. If compaction purges the tombstone after its grace period before C receives it, a later repair can treat C's surviving value as live and spread it again. Cassandra's tombstone documentation describes this resurrection risk when a replica is unavailable beyond gc_grace_seconds.
Do not read gc_grace_seconds as “deletion happens after this many seconds.” The tombstone can affect reads immediately, and expiry merely makes it eligible for removal during compaction under Cassandra's rules. Setting grace to zero to reduce tombstone storage can make a missed deletion dangerous. Setting it very high increases read and storage cost if tombstones accumulate. The operational promise is that replicas which could have missed deletes are repaired before tombstones become purgeable, with enough margin for failures and repair duration. The exact safe setting also depends on table and repair strategy.
For the incident, find the deleted partition and the old cell timestamp, inspect whether the replica was down, and establish when tombstone creation, last successful repair and compaction happened. Check TTL expiry too, since TTL also creates tombstones. Do not just delete the row again and close the ticket. That may hide a broken repair schedule.
If the old value was already copied into a search index or training set, those derived copies have their own deletion path and must be checked separately. Source repair alone cannot prove that an assistant no longer quotes the record.
If asked how to prevent this at scale, I would monitor repair coverage and age, node outages relative to grace, tombstone pressure, and resurrection tests in a controlled cluster. A deletion requirement with a strict deadline also needs an authoritative suppression or access check for downstream reads. The core failure is that the negative fact was garbage-collected before every replica learned it.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →