Retrieval and RAG · Principal
New documents are indexed. Why did ANN recall fall after a week of updates?
The question
Interview question
The search index accepts inserts and reports a healthy document count. Over a week of updates and deletes, more questions fail to retrieve their answer-bearing passage. A rebuild from the current corpus restores recall. An engineer wants to increase `top_k` from 20 to 100. What would you establish before doing that?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First I need two different reference points. Did the exact nearest neighbors change because the corpus changed, or did the approximate index get worse at finding the same neighbors? For a fixed, authorized query set and a pinned corpus generation, compute exact scores over the eligible vectors, then compare the ANN candidates at the configured search budget. The exact result is a useful retrieval oracle for vector similarity, not a guarantee that the passages answer the question. We need both ANN recall against exact search and answer-bearing recall against human-labeled evidence.
There are several routes to the observed pattern. The query and document embedding versions or normalization recipe may have drifted. Some writes may be acknowledged by one replica but not the one serving queries. An update may leave an old vector live under another ID, or a tombstone filter may remove many candidates after approximate search. Index-specific graph or list maintenance can also change search quality under heavy churn. Faiss documents HNSW search parameters and index operations. In its HNSW implementation, removal is not supported, so a service built on it must handle deletion outside that structure or rebuild. Other engines have different update semantics. Do not infer a particular engine's failure from the word “ANN.”
I would capture query, index generation, replica, embedding version, filter, exact top neighbors, ANN visited or candidate counts where available, and post-filter results. Compare the same query against the live index, a freshly built index over the same source snapshot and an exact scan. If live ANN loses neighbors that fresh ANN finds under the same metric and search settings, inspect incremental construction and maintenance. If both approximate indexes agree but exact neighbors differ from last week, source, embeddings or filters changed. If pre-filter candidates contain the answer but post-filter drops it, debug authorization and tombstones, not graph traversal.
Increasing the search budget or top_k can be a bounded mitigation, but they are different knobs. A larger final result count does not necessarily make the index explore enough of the graph or probe more lists. It can also send many irrelevant passages into the reranker and model. Test actual efSearch, probes or equivalent settings for this engine against latency, candidate recall and authorized result quality. pgvector's HNSW guidance shows such search-depth and maintenance choices for one implementation. A rebuild or generation swap may be warranted if churn has degraded the structure, with a change stream catch-up and rollback plan.
I would keep a small exact-search audit running on a representative query sample, including rare but important documents and recent inserts. Count freshness, tombstone lag and recall by index generation, not just total document count. A green ingest counter proves accepted writes. It does not prove the search path can find the right authorized evidence.
Continue reading
Related questions
Read beyond the question
Explore more retrieval and rag
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →