Evaluation and Quality · Staff
The citation is relevant, but it does not support the number. Is the answer correct?
The question
Interview question
An answer states a numeric value and cites a relevant passage, but that passage does not support the exact value. How should the eval score it? The value happens to be correct according to another source.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would score two different things. Is the number true under the right date, unit, and population? And did the cited evidence actually support the claim the user saw? The first may be yes while the second is no. If the product promises grounded answers, a lucky correct number with a decorative citation has still failed the evidence contract.
Take an example. The assistant says, “The service met 99.95 percent availability in August,” and cites an incident page that only says August had one regional outage. The incident page is relevant. It does not contain uptime, denominator, region coverage, or a calculation from which 99.95 follows. The number might appear in a verified monthly report elsewhere. Then factual correctness can get credit after checking that report, but citation support should fail. The retrieval or citation assembly path may have missed the actual source.
I would split an answer into material claims and ask for the smallest supporting span or valid derivation for each. For a number, check the quantity, unit, time window, entity, qualifiers, and whether it is direct or calculated. If the source gives 999,500 successful requests out of 1,000,000 eligible requests, a deterministic calculation can support 99.95 percent if those counts refer to the same service and month. Cite the input counts and the calculation, not a neighboring paragraph that merely mentions availability. If the definition excludes scheduled maintenance, that exclusion has to survive into the answer too.
This is why exact string match is too weak. The number could be rounded, converted, or derived. And a judge that sees “99.95” in both the answer and a nearby page can approve a claim with the wrong year or denominator. I would have the eval label the required source identity and fields, then use code for arithmetic where possible and human review for disputed scope. A model judge can triage claims, but it should be given the cited span and the rubric. It should not use its own memory to fill a missing premise. OpenAI's evaluation best practices recommend task-specific criteria and human calibration of automated graders. Here the criterion is entailment of the numeric claim by permitted evidence.
The final number happens to be correct. That changes the factual-correctness label, not the citation-support label. I would record the failure as “correct claim, unsupported citation” rather than collapse everything into a single wrong-answer bit. That distinction tells the team where to repair the system. Perhaps the good report was retrieved and dropped from context, or perhaps the model supplied the number from prior knowledge. Neither makes the incident page a valid citation.
What if a different retrieved source supports the number but was not cited? The answer may be factually supported in the internal trace, but the user still cannot verify it through the displayed citation. Fix the citation mapping and test the visible response. If the different source is restricted to a principal who cannot read it, do not simply swap in that citation. Re-run the answer under the actual principal's allowed evidence. Correctness in a privileged eval run does not license disclosure in the product.
For release, I would report claim-level support separately from overall answer correctness, and break it out for numeric, temporal, and policy claims. A fluent answer with five supported claims and one unsupported decisive number should fail the relevant case. Averaging six claims into a high score would miss the harm. The citation is part of the claim, not a footnote we score for topical similarity.
Continue practicing
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or return to the full Interview Prep index.
Browse this area →