I would stop treating the reward curve as the release score. The reward model learned to predict preferences from finite comparisons. The policy is now being optimized against that predictor, so it has an incentive to find responses the predictor likes even when people do not. It may learn verbosity, a particular refusal style, flattering agreement, or a way to game a format feature. It might also be better on the training distribution and worse on the new tasks. A rising proxy score and falling human preference do not, by themselves, identify which failure occurred.

The reward overoptimization study from OpenAI is useful evidence for the general mechanism: optimizing an imperfect proxy far enough can reduce a separate target score. Its controlled study uses a simulated gold reward, not proof that our particular policy is hacking its reward model. We need to look at actual responses. Take the same held out prompts, produce old and new answers under controlled decoding, blind the policy identity, and ask reviewers for a preference with a reason tied to the task. Include ties and disagreements. Report the whole assigned sample, not just examples where both models finished neatly.

I would slice the comparison by difficulty, language, code, tool use, safety, length and user goal. Review cases where the reward margin grows most while the human preference flips. Does the new answer omit a needed step while sounding decisive? Does it refuse legitimate work? Does it invent evidence or learn to exploit a rubric phrase? Also check whether the reward model's own input changed, whether the prompt template or tokenizer changed, and whether the human study showed both reviewers the same context. A badly run preference study can be wrong too. I would not declare humans infallible because one aggregate number moved.

Now look at the training trajectory, not only the last checkpoint. Plot independent human preference against reward score and the distance from the reference policy across checkpoints. Compare response lengths and other cheap behavioral signals. A KL penalty or an early stop can limit how far a policy moves from its reference, but neither makes an imperfect reward true. Better comparison data on the failure slices, a reward model refreshed with examples from the current policy, and independent evaluations can address the cause more directly. If the model learned to exploit a grader that the same team uses for release decisions, reserve a genuinely separate holdout and human review protocol before doing another promotion.

Suppose the interviewer says the new policy has better scores on a safety set but human reviewers dislike its refusals. That is not a one number decision. Separate harmful requests correctly refused from ordinary requests wrongly refused, and look at severity and user utility in each slice. The release owner sets guardrails for critical harm and an improvement target for legitimate tasks. We may ship a bounded safety improvement while repairing overrefusal, or hold the checkpoint if it fails a hard product requirement. Make that choice explicit, not an average of incompatible metrics.

Every optimizer update can make the policy better at the proxy while making it worse for a person. That difference changes what we log, what checkpoint we choose, and how we collect the next training examples.