The labels say which of two answers a reviewer preferred. DPO also asks how the current policy changes the relative odds of those answers compared with a reference policy. That comparison is in the loss. It is not just a training artifact that can be replaced after the fact.

For one prompt, let Δpolicy be the policy log probability of the chosen completion minus that of the rejected completion. Let Δref be the same difference under the frozen reference. The common sigmoid DPO loss for that pair is −log σ(β × (Δpolicy − Δref)). The original DPO paper derives the method, and the TRL trainer documentation writes the loss with both policy and reference log probabilities. If the policy currently gives the chosen answer twice the odds of the rejected answer, but R1 gave it four times the odds, the policy is behind that reference on this pair. If R2 gave the two equal odds, the very same policy is ahead. The chosen label did not move, but the margin and gradient did.

This does not mean the reference decides what humans want. It sets the baseline for the policy's relative movement, and β scales the comparison in this loss. The effect across a full dataset depends on how the reference assigns probabilities to each completion and on which examples dominate training. A new reference may be stronger overall yet make some preferred completions less likely. Looking only at its average benchmark score misses the per pair change the optimizer sees.

I would verify that the experiment really changed only the reference. Compare the exact frozen checkpoint hash, tokenizer, chat template, special tokens, prompt and completion boundaries, truncation, masking and sequence log probability computation. A subtle bug can score the prompt as part of one response, use stale precomputed reference log probabilities from R1, or activate an adapter while claiming the reference is frozen. Compute the two reference margins on a few hand checked rows, then compare their distribution over the whole set by task and response length. Hold the initial policy fixed while measuring the first batch's loss and gradients under R1 and R2. That makes the mechanism visible before a long run adds optimizer and data order effects.

The next pushback is usually, “Could we change the reference halfway through training because R2 is better?” You can design an algorithm that updates its reference, but silently swapping it changes the objective at that step. I would make it a named new training stage with a checkpoint, logged reference version and a fresh evaluation plan. Do not compare its loss curve directly with the old run as though the scale and target remained fixed. Evaluate actual generated behavior on independent tasks, not just DPO loss or the percentage of pairs with a positive margin. A low loss can coexist with regression outside the preference set.

The reference is part of what the optimizer was asked to improve relative to, even when every preference label is identical.