Data and Knowledge Systems · Principal
We reuploaded a deleted file at the same path. Why did search delete the new one?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The indexer may be treating a path as a permanent identity. The source deletes policies/refund.pdf, then uploads a new file under that same key. The new put event reaches the indexer first and creates new chunks. The older delete event arrives afterward. If its handler says “remove everything under this path,” it removes the new generation as well. S3 event notifications are not guaranteed to arrive in event order, and S3 includes a per-key sequencer for relevant put and delete events. Other connectors may have different ordering and version fields, so the protocol needs to be designed for the actual source.
I would give the index a logical key for lookup and a separate source generation for content ownership. A chunk records bucket, key, source version or generation, and the event sequence that installed it. A delete of the old generation may remove only its chunks. It must not clear a newer current generation at the same key. For S3, store and compare the per-key sequencer atomically with the index metadata. AWS says to compare its hexadecimal values for the same key, padding shorter strings before lexicographic comparison. It is not a global timestamp across files. In a versioned bucket, a simple delete creates a delete marker with its own version ID, and an older object version may still physically exist. AWS's delete marker documentation matters when deciding whether the search index should represent the current view or historical versions.
The race is not over when the event handler finishes. Imagine the old put event starts a slow embedding job. The new file is uploaded and indexed, then the old job finishes and writes its chunks late. The final write must compare its expected generation with the source generation currently owned by that key, and refuse to publish if it lost. That condition belongs at index commit, not only at event receipt. For a source that exposes no monotonic sequence, use a stable version plus a source read or reconciliation step before publish. A hash of content by itself cannot tell which identical-looking upload came later.
I would reproduce put A, delete A, put B and deliver notifications in every order, including duplicates and delayed embedding completion. Assert that search points to B, never to a mix of A and B, and that the tombstone does not erase B. A periodic source-to-index reconciliation catches gaps that event ordering logic cannot see. This is a lifecycle problem for a reused name, beyond merely discarding an old value update.
Continue reading
Related questions
Read beyond the question
Explore more data and knowledge systems
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →