Ask how the evidence entered the evaluation. If the benchmark hands the generator a gold passage selected by a human, it measures what the generator does given that passage. Production first has to retrieve, permission-check, rank, pack and fit the evidence into context. A 95% oracle-context score is a useful ceiling for a particular question set. It is not an end-to-end answer rate. The original RAG paper explicitly combines a retriever with generation, and the BRIGHT benchmark's RAG case study illustrates a gold relevant example helping where ordinary retrieval did not.

Run paired conditions on the same eligible questions. Give the reader gold evidence in one condition, production-retrieved evidence in another, and optionally no evidence as a baseline. Record whether the answer-bearing passage was in the indexed corpus, retrievable under the user's permissions, present in the candidate set, retained after reranking, and actually included in the final context. Then inspect whether the generator used it correctly. This locates the gap. Do not label a question as a retrieval failure if the authoritative source was never in the index, or call a generator wrong when the necessary passage was removed by an access filter.

Be careful with the oracle condition too. A gold passage may include the answer in a more direct form than a real source would, or come from a document the user cannot access. It can even leak an answer written after the query's as-of date. The oracle set should obey source and permission rules when it is used to diagnose a deployable system. Otherwise call it a capability probe, not a production ceiling.

An interviewer might ask whether the best next investment is a larger reader. I would compare the gap by failure stage first. If the reader fails with valid gold evidence, work on interpretation, reasoning or abstention. If it succeeds with gold but the live pipeline rarely presents that evidence, fix ingestion and retrieval. The new RAG stack wins. Was it the model, the retriever, or their combination? compares combinations of retriever and model changes. This question catches the evaluation harness quietly doing the hardest retrieval work for the system.