We cannot decide from that number. The labels were collected from one candidate distribution. A missing judgment means nobody assessed the document, not that the document is irrelevant. If the new system retrieves more unjudged items, scoring them as zero favors the old system whose results filled the pool. TREC has long used pooled judgments because assessing every document is infeasible, and its overview discusses the incompleteness assumption behind evaluating unpooled results. A sufficiently diverse, deep pool can still be useful. We need to measure whether this particular pool covers the new system, rather than dismiss all offline metrics.

First quantify judged coverage at each rank for both retrievers and by query class. Count new documents, old documents, restricted results and answer-bearing passages separately. Inspect the top unjudged hits from the new system without telling reviewers which model returned them. Add a pooled sample from both systems and relevant lexical or human-found baselines, with stable document revisions. Have qualified assessors label whether each passage supports the query, including the version and permission context. If answer quality is the product goal, also evaluate final answers with the evidence each retriever supplies. NDCG for passage relevance and grounded task success are different measures.

Do not fix the metric by marking all unjudged documents relevant either. Some fresh results will be noise, or a page can mention the query without containing the answer. Use a judged union at a defined depth, plus a stratified sample from outside that union to estimate what remains unobserved. Report the uncertainty when coverage is thin. A system that retrieves an unjudged policy update may be genuinely better. It might also be surfacing a copied or unauthorized version. Label the actual content the user could see at evaluation time, not only a document ID that now points somewhere else.

I would preserve a separate held-out query set and label budget. If we repeatedly select queries where the new retriever looks good, relabel, tune and repeat, the refreshed set becomes a development set. Use it to understand the mechanism, then confirm on untouched queries and an online test where permitted. For an online comparison, measure successful evidence use and user outcomes, with guardrails for latency and restricted material. A few human anecdotes can identify a flaw in the old qrels, but they do not establish a fleet-wide win.

If the interviewer insists the new NDCG is down by ten points, I would show the score on commonly judged results alongside judged coverage, then the score after balanced adjudication. The first comparison may still reveal a real loss. The second can expose pooling bias. Neither should be hidden. The decision is about retrieval quality on the current authorized corpus, not loyalty to the candidates last year's system happened to show.