Evaluation and Quality · Principal
Reviewers agree more after calibration. Why do they all miss the same failure?
The question
Interview question
Reviewers used to disagree about whether an answer was grounded. After rubric calibration, agreement rises and an LLM judge now matches them. An external auditor finds that they all accept a cited revenue number that is absent from the cited table. Is the new agreement evidence that quality improved?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It is evidence that the reviewers apply a rule more consistently. It is not evidence that the rule is right. They may have learned to treat a relevant-looking citation as proof that every number in a sentence is supported. The judge can learn the same shortcut from examples or rubric text. If the audit has found a shared miss, three agreeing reviewers and a matching model are not four independent checks.
I would inspect the actual source and the claim at the smallest useful unit. Which number was asserted, which table cell or calculation would support it, and what does the citation actually contain? Then sample other answers blindly from before and after calibration, with source material available to adjudicators who did not train on this rubric. Include paired cases where only the unsupported number changes while the document, citation and surrounding prose stay the same. A grader who still approves both is testing citation presence or topical relevance, not numeric support. RubricBench is a research example of why nuanced and misleading rubric cases matter. It does not tell us this system's failure rate.
The measurement has two pieces. A targeted set of known traps checks whether the corrected rubric catches the failure. An independently drawn probability sample estimates how often the failure occurs in the actual traffic mix. Keep answer type, source format, model version and reviewer cohort in the strata so a change in traffic does not masquerade as a quality change. Record both the old and corrected grades on the same examples. Do not extrapolate prevalence from the auditor's handpicked bad cases.
Rewrite the rule so a numerical claim passes only when the cited material contains the value or a reproducible calculation from cited values, with the relevant scope and units. Make reviewers point to the cell or inputs. If the table is ambiguous or inaccessible, allow an uncertain grade. Run a fresh blind calibration on cases the reviewers have not memorized, then regrade a representative slice of history under a versioned rubric. Historical dashboards should say which rubric produced each number. Agreement can still be useful, but report it beside independently adjudicated accuracy and the shared-error rate.
What if experts disagree too? I would ask whether they are disagreeing about arithmetic, the meaning of a table heading, or what level of evidence the product requires. Write that boundary into the claim rule and retain unresolved cases as uncertainty. Forcing one gold label where the source itself is unclear makes the evaluation look cleaner than the product. Two reviewers disagree whether the question is answerable is about a disagreement that the process has to resolve. This is more uncomfortable: calibration succeeded at making people agree on a bad shortcut.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or return to the full Interview Prep index.
Browse this area →