Model and Inference Engineering · Principal
The preference data was valid. Why did training reward the rejected answer?
The question
Interview question
A human reviewer chose response A over B. The stored record is right. During an export, one job emits `(prompt, A, B)` and another emits `(prompt, B, A)`. Both are valid strings, both fit the tensor shapes, and the DPO loss decreases. Yet a safety evaluation gets worse. Where would you look?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First, the objective only knows which side the data calls chosen. The DPO paper defines a pairwise objective that increases the preferred response's relative likelihood against a reference policy. Swap the two responses and the training signal reverses for that pair. A falling loss means the optimizer is fitting the labels it received. It cannot certify that the labels still reflect the reviewer's decision.
I would trace one bad example by immutable review ID from UI decision through extraction, joins, serialization, tokenization and trainer batch. Preserve named fields rather than positional tuples at the boundaries. Hash the prompt and both response texts, retain the review choice as an ID, and assert that chosen_response_id still points to the same text after every transformation. A join on prompt alone can accidentally combine candidates from different review rounds. Tokenization should not silently collapse the two sides into an identical or truncated response either.
If the interviewer says this affects only 2% of pairs, I would ask which 2%. A small global rate can concentrate in refusal, tool use or a newly imported source and dominate that slice's behavior. Compare signed preference margins on a fixed, human-inspected canary set before and after training, with an independent parser that reads the original review record. Break down flip rate by source and transformation version. Pair reversal in training can be exposed with synthetic examples whose direction is unambiguous, but that alone will not validate a production join. Put both checks in the data contract and block training when sampled lineage disagrees.
Rollback uses the last clean checkpoint and a rebuilt dataset, not a cosmetic label fix in the evaluation dashboard. Quarantine affected shards, estimate exactly which runs consumed them, and retrain or revert according to impact. The DPO pair has a chosen and rejected answer. Did both answer the same prompt? asks whether the preferred and rejected answers were compared under the same prompt. Here they were, but their ordering changed on the route from human judgment to objective.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →