Model and Inference Engineering · Staff
The reward model prefers longer answers. Did humans prefer length or correctness?
The question
Interview question
A reward model trained on chosen and rejected answers scores longer answers higher. The policy starts adding several paragraphs to simple questions and its reward rises. The team says the human annotators chose the longer answers, so the model is doing exactly what people wanted. How would you tell whether length is a useful signal or a shortcut?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The pairwise labels do not isolate why a response won. Suppose chosen answers were longer because they included the missing correct step, while rejected answers were shorter because they skipped it. A reward model can learn the underlying usefulness, the length correlation, or both. Once a policy optimizes that model, it can add words without adding the missing step. The training comparison still looks satisfied to the proxy. Work on reward-model length bias studies this issue directly. It does not mean every preference for a longer answer is mistaken.
I would audit the pair construction. Plot length difference and choice rate by task, difficulty, response source and annotator instruction. Read pairs in which the longer response is plainly worse, and pairs where its extra detail is necessary. Look for boilerplate, repeated caveats and formatting that correlate with the chosen label. If nearly all pairs place a good long answer against a bad short answer, the data provide little evidence about length independent of quality. A reward model cannot learn a clean distinction from a confounded experiment by wishing hard enough.
Build controlled tests that hold factual content and task success fixed while changing length. For a straightforward question, compare a correct two-sentence answer with a padded version that repeats itself. Also compare a short but incomplete answer with a longer one that supplies a necessary condition. Have blinded humans judge usefulness and correctness in both directions, with varied response order. Check how reward difference changes with length when correctness is fixed, and how it changes with correctness when length is matched. Slices matter. A detailed debugging answer may legitimately need 600 words while a yes-or-no fact does not.
I would also measure policy outputs, not just reward-model classification accuracy on held-out pairs from the same collection process. Count redundant tokens, omitted required details, correct outcomes, human preference and cost on fresh prompts. If the model wins on the old pair distribution but loses with real users who needed a crisp answer, the release decision is straightforward. We can improve data with length-balanced comparisons, better annotator guidance and adversarial pairs, or calibrate the reward under an independently checked quality target. Any correction needs testing for the opposite mistake of rewarding terse but incomplete replies.
The interviewer might propose subtracting 0.01 × tokens from every reward. That can be a useful cost or style term if chosen deliberately. It is not a universal debiasing formula. Different tasks have different minimum explanation lengths, and the reward scale may change across model versions. I would first identify whether the bias came from labels, reward-model representation, or policy overoptimization, then tune against a held-out human and task evaluation. A rising reward is not enough evidence that the extra words helped anyone.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →