Model and Inference Engineering · Staff
The DPO pair has a chosen and rejected answer. Did both answer the same prompt?
The question
Interview question
A preference dataset joins feedback by conversation ID. The chosen answer came after a tool returned the latest price. The rejected answer came from an earlier retry before the tool returned. The training row contains one prompt field and both completions, so the DPO trainer runs without error. What signal is it learning?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The preference comparison assumes the candidates are responses to the same conditioning context. The original DPO paper defines a preference between two responses for a prompt, and TRL's DPO trainer expects a prompt, chosen completion and rejected completion. If one candidate saw a price that the other never saw, the preference may be about information availability, not answer quality. Reformatting them under a shared prompt can also create an impossible example, where a chosen answer refers to a fact absent from the supplied context. The loss still has numbers. It cannot detect that the label asks the model to know the future.
I would inspect the raw trace, not only the flattened pair. Record message and tool-result IDs, their order, timestamps, tool versions and the exact context each candidate saw. Hash a canonical prompt boundary after rendering the chat template, including system and developer instructions, relevant tool messages, images if present and truncation settings. A conversation ID is too broad. Even two candidates at the same turn may have seen different retrieved documents or different versions of a tool output. Pair only candidates with a compatible context. When they are not comparable, exclude them from pairwise preference training or relabel under a deliberately shared replay.
There is a second failure in preprocessing. If chosen and rejected prompts are the same in raw data but are truncated differently because their completions have different lengths, the effective model inputs can diverge. Check the actual tokenized sequences and the completion mask for each side. Keep the decisive evidence in both, or drop the pair when that cannot fit. A perfectly valid optimization run against mismatched contexts can reward hallucinated facts and penalize reasonable uncertainty.
What if the rater genuinely preferred the newer answer? That observation is useful for the product. It says getting fresh price data helped. It does not by itself say the model should prefer that response without the tool result. Build a tool-use or retrieval evaluation for the workflow, and use a matched-context preference pair to teach response quality. The preference pairs stayed the same. Why did a new DPO reference model change the policy? asks how changing the reference model changes DPO. This question asks whether the supposed preference pair represented a valid comparison before the loss was computed.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →