Model and Inference Engineering · Principal
The DPO reference and policy score different prompt tokens. What does the margin mean?
The question
Interview question
A custom DPO trainer renders a conversation with a new chat template for the trainable policy. It reads cached reference log probabilities computed months ago with the old template. The chosen and rejected text is unchanged. Loss decreases normally. Is the preference objective still comparing the two models on the same event?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Not necessarily. DPO compares how the policy and reference score the chosen versus rejected completion under a prompt. Schematically, its margin contains the policy's log-probability difference for chosen and rejected, minus the corresponding reference difference. The original DPO paper defines that comparison, and TRL's DPO trainer documentation describes rendering conversational data with the model's template. If the cached reference used another prompt serialization, its log probabilities are conditioned on different tokens. The loss is still a number, but it is no longer the intended policy-versus-reference contrast under the same context.
This is more than cosmetic role punctuation. Templates can change system instructions, tool-result markers, assistant-start tokens and whether an end-of-turn token belongs to the prompt or completion. Even if the visible words are the same, token boundaries and conditional probabilities can differ. Check the exact token IDs and completion masks for policy and reference, not only the raw JSON row. The cached reference log probabilities should be keyed by model revision, tokenizer and special tokens, chat template, serialized prompt and completion, truncation rule and scoring mask. A hash of the text alone will miss a template change.
I would take a small fixed pair and recompute reference log probabilities using the current policy's intended serialization, without updating either model. Compare to the cached values token by token. Find the first divergence and whether the decisive completion tokens were truncated. If the two checkpoints require incompatible tokenizers or vocabularies, we cannot simply feed one set of IDs to both and call the result a DPO reference. Align the scoring contract or choose a compatible reference and re-evaluate the recipe. Do not paper over it by tuning beta until the loss looks familiar.
One pushback is that DPO can use a different reference checkpoint. Yes. The preference pairs stayed the same. Why did a new DPO reference model change the policy? covers how that legitimate choice changes the policy target. This case is a broken comparison even when the checkpoint choice is intentional, because one side scores a different serialized event. The DPO pair has a chosen and rejected answer. Did both answer the same prompt? checks whether chosen and rejected saw the same real-world prompt. Here those candidates can be perfectly matched while policy and reference disagree about what the prompt bytes are. Both contracts must hold.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →