Retrieval and RAG · Staff
The exact error code ranks first in keyword search. Why did hybrid search bury it?
The question
Interview question
A support assistant searches for `E180`. BM25 returns the incident note containing that exact code at rank one. Dense retrieval finds broadly similar troubleshooting articles. The hybrid service adds the raw scores and the exact note drops below the context cutoff. Does a combined score necessarily combine evidence?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
No. BM25 and vector similarity are different measurements. Their numeric ranges can change with query, corpus and scoring implementation. Adding BM25 + cosine without calibration lets whichever side has larger numbers dominate. Even a fixed weight does not make an uncalibrated sum meaningful. Elasticsearch's approximate kNN guidance calls out the unrelated score scales and explains that reciprocal rank fusion combines positions instead of raw scores. RRF is a useful baseline, not a guarantee that every exact match wins.
I would inspect both candidate lists before fusion, their scores, candidate limits, filters and deduplication. Is the exact document present in both lists, only the lexical list, or absent because the hybrid query applied different tenant or date filters? Then replay with a small fixture of codes that differ by one character. A dense model can treat E180 and E108 as similar strings while the distinction is the entire task. Protect exact identifiers with an explicit lexical signal or typed lookup, and use semantic search for surrounding explanation. Evaluate answer correctness and citation support, not only whether a document appeared somewhere in top 100.
For rank fusion, choose candidate depth and the RRF rank constant on held-out queries. If you need weighted score fusion, normalize or calibrate on representative queries and monitor shifts when the corpus changes. A reranker can help among candidates, but it cannot rescue the exact note if first-stage retrieval discarded it. Do not relax permissions to improve recall.
The interviewer might say BM25 already solved this case, so why keep hybrid search? Because the next user may describe the symptom without the code. The product needs both intents. Why did hybrid search fill every slot with one old document? covers one old document filling every hybrid slot because of chunk duplication. This question is about incompatible raw score scales and identifier fidelity, even with distinct candidate documents.
Continue reading
Related questions
Read beyond the question
Explore more retrieval and rag
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →