Retrieval and RAG · Principal
The document did not change. Why did its BM25 score fall after another tenant indexed files?
The question
Interview question
A RAG service rejects hits below BM25 score 7. Tenant A has a useful policy paragraph that scored 8 yesterday. Tenant B bulk indexes thousands of documents, many containing the same query terms. Tenant A's document and question are unchanged, but its score falls below 7 and the assistant says no policy was found. How can another tenant's data change the result of a correctly filtered search?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
BM25 is not an absolute confidence measure. Among its ingredients are term frequency in a document, document length and collection statistics such as how rare a term is. If a shared index computes those statistics across documents outside tenant A, adding tenant B's corpus can change the score assigned to A's hit. Shard placement and replica statistics can add differences too. Elastic's BM25 explanation describes the role of term rarity, and its consistent-scoring guidance discusses index and shard statistics. The particular cross-tenant threshold failure is our hypothetical application design.
I would replay the identical query and tenant filter against index snapshots before and after the bulk load. Compare analyzed query tokens, the returned hit IDs, per-term document frequency, field length statistics, shard selection and the score explanation. First establish whether the ranking changed or only the arbitrary cutoff did. A document can remain the best permitted hit and still cross a fixed numeric threshold. A lexical score of 7 from one query is not necessarily comparable to a 7 from another query either.
Fix the answerability decision rather than tuning 7 to 6.9. Retrieve a bounded candidate set under the permission filter, then test whether the candidate actually supports the requested claim. If a cheap relevance gate is needed, train or calibrate it on labeled query-document pairs from the intended tenant and query mix, with versioned features. Measure abstention errors and citation support, not just a ranking metric. Depending on scale, per-tenant indices or routing may isolate statistics, but that costs shards and operations. Search types that gather more global statistics can reduce shard score inconsistency, not make a raw BM25 score a timeless probability.
Suppose the interviewer asks whether this is a data-leak incident. No private document was returned in this scenario. Still, shared statistics can create cross-tenant coupling in ranking, and permission filtering must happen before any content reaches the model. Investigate both properties separately. A federated search merges incomparable scores covers merging scores from different search systems. Here one lexical system changes a hit's score as its own underlying corpus changes.
Continue reading
Related questions
Read beyond the question
Explore more retrieval and rag
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →