Model and Inference Engineering · Staff
The preference pairs stayed the same. Why did a new DPO reference model change the policy?
The question
Interview question
A team trains with direct preference optimization. Each row still has the same prompt, chosen answer and rejected answer. They switch the frozen reference from checkpoint R1 to R2 and keep the policy initialization and all other training settings fixed. The learned policy changes. An engineer says the labels are unchanged, so the reference should not matter. What would you show them?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The labels say which of two answers a reviewer preferred. DPO also asks how the current policy changes the relative odds of those answers compared with a reference policy. That comparison is in the loss. It is not just a training artifact that can be replaced after the fact.
For one prompt, let Δpolicy be the policy log probability of the chosen completion minus that of the rejected completion. Let Δref be the same difference under the frozen reference. The common sigmoid DPO loss for that pair is −log σ(β × (Δpolicy − Δref)). The original DPO paper derives the method, and the TRL trainer documentation writes the loss with both policy and reference log probabilities. If the policy currently gives the chosen answer twice the odds of the rejected answer, but R1 gave it four times the odds, the policy is behind that reference on this pair. If R2 gave the two equal odds, the very same policy is ahead. The chosen label did not move, but the margin and gradient did.
This does not mean the reference decides what humans want. It sets the baseline for the policy's relative movement, and β scales the comparison in this loss. The effect across a full dataset depends on how the reference assigns probabilities to each completion and on which examples dominate training. A new reference may be stronger overall yet make some preferred completions less likely. Looking only at its average benchmark score misses the per pair change the optimizer sees.
I would verify that the experiment really changed only the reference. Compare the exact frozen checkpoint hash, tokenizer, chat template, special tokens, prompt and completion boundaries, truncation, masking and sequence log probability computation. A subtle bug can score the prompt as part of one response, use stale precomputed reference log probabilities from R1, or activate an adapter while claiming the reference is frozen. Compute the two reference margins on a few hand checked rows, then compare their distribution over the whole set by task and response length. Hold the initial policy fixed while measuring the first batch's loss and gradients under R1 and R2. That makes the mechanism visible before a long run adds optimizer and data order effects.
The next pushback is usually, “Could we change the reference halfway through training because R2 is better?” You can design an algorithm that updates its reference, but silently swapping it changes the objective at that step. I would make it a named new training stage with a checkpoint, logged reference version and a fresh evaluation plan. Do not compare its loss curve directly with the old run as though the scale and target remained fixed. Evaluate actual generated behavior on independent tasks, not just DPO loss or the percentage of pairs with a positive margin. A low loss can coexist with regression outside the preference set.
The reference is part of what the optimizer was asked to improve relative to, even when every preference label is identical.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →