Model and Inference Engineering · Staff
The DPO pair is valid in the dataset. Did truncation remove the preference?
The question
Interview question
Two answers share a long opening and differ only near the end. The chosen one refuses an unsafe action. The rejected one gives the action. In the raw row, the preference label is perfectly clear. But the trainer has a maximum token length and keeps the beginning of each prompt plus completion. If it cuts both answers before the decisive sentence, what is it training on?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The loss compares sequence log probabilities of the chosen and rejected completions under the policy and reference model. It cannot learn a distinction that was removed from the actual token sequences. At the extreme, if the two tokenized completions become identical after truncation, the pair contributes no useful preference between those endings. More often they retain a little unrelated difference and the optimizer learns to favor that instead. TRL's DPO configuration documents a maximum tokenized sequence length and a truncation mode, which are part of the effective dataset.
I would inspect tokenized rows after the exact chat template, tokenizer, prompt split and trainer truncation. Compute how often chosen and rejected sequences still differ, where the first differing token falls, how many completion tokens survive, and whether the human reason for the label is still present. Sample the worst cases and decode them. A length histogram alone will miss pairs that technically fit a cap but lose their important tail under a separate prompt limit.
Fixing it is not just setting a huge context window. Longer pairs cost memory and time, and truncating prompts from the other end can remove the facts needed to judge the answer. Curate or shorten prompts while preserving the decision context, raise the length limit where affordable, or filter pairs whose preference cannot survive preprocessing. Keep the same preprocessing in evaluation.
If the interviewer says the raw dataset passed QA, that is exactly why I would look here. The model never trains on a spreadsheet row. It trains on token IDs and masks. The preference needs to remain visible there.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →