Model and Inference Engineering · Principal
The reward says correct. Why did the model's final answer give another number?
The question
Interview question
A math rollout mentions the correct number in a rejected candidate, then gives a different final answer. A substring verifier awards full reward. How should the answer be parsed and graded?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Look at what the verifier extracts, not just the ground-truth value. Imagine a math rollout says, “One candidate is 42, but that was a mistake. Final answer: 41.” The expected answer is 42. A reward function that searches for 42 anywhere in the entire text gives a success signal even though the response delivered to the user is wrong. During RL, the policy can learn to include likely answers in a list or quote the problem's number, then finish with something else.
The evaluator must have a defined answer contract: a parseable final-answer field or delimiter, one unambiguous value, and clear handling for malformed or missing finals. Parse only the intended final region, then compare a normalized typed value under task-specific rules. For code, the equivalent is executing the submitted artifact, not awarding credit because a passing snippet appeared in a comment. OpenAI's grader documentation distinguishes grader types and supports custom code grading, while OpenAI Evals shows that matching behavior is a design choice of the eval, not a property of the model.
Before training, attack the reward function with adversarial completions: multiple numbers, a correct number in a quote, a correct intermediate result followed by a wrong final, an empty final, Unicode digits, units and a refusal. Write down which should score. Then run a small sample of actual rollouts through both the automatic verifier and human review. Track disagreement and the fraction the parser cannot classify. Do not silently count parse errors as correct.
If we tighten the parser after a run, old reward curves are not directly comparable to new ones. Regrade a fixed rollout sample with both versions to measure how much of the apparent gain came from the grading loophole. Keep verifier version with every reward record.
An interviewer might suggest a stronger LLM judge. It can help on open answers, but a narrow verifiable task benefits from a deterministic, well-tested final-answer contract. The important question is whether the reward matches the user's observable answer, not whether a correct token appeared somewhere in the model's work.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →