They are optimizing different margins. Standard DPO compares the policy's chosen-versus-rejected log-probability difference with the same difference under a reference model. A completion probability is a product of conditional token probabilities, so its log is a sum. Dividing each completion by its length changes the quantity into an average per-token score. For responses of different lengths, it can change the sign and magnitude of the pairwise margin even if every token probability remains the same. It also changes how the reference baseline cancels. Calling both runs “same DPO data” misses an objective change.

I would inspect a few actual pairs with token IDs, completion masks, policy and reference log-probability sums, lengths and resulting margins. Make sure the prompt tokens are excluded consistently, as The DPO reference and policy score different prompt tokens. What does the margin mean? discusses. Padding and EOS handling matter as well. If the team meant to control verbosity, measure whether humans prefer the longer answer because it is more useful or because annotation and judge processes reward verbosity. Research on length in DPO shows that length can affect preference optimization, but it does not make arbitrary length normalization an equivalent correction.

Run the two objectives as separate experiments, holding data, reference, beta, tokenizer and evaluation fixed. Slice by pair length ratio, task type, factual support, refusal and tool use. Report output-length distribution and human preference with length-controlled comparisons. A short model can look more concise while dropping a necessary step. A longer model can look more helpful while padding an answer. Neither conclusion follows from offline loss alone.

An interviewer may ask whether per-token averaging is always wrong. No. It can be a deliberate alternative scoring or training choice. It just needs its own derivation and evaluation. The reward model prefers longer answers. Did humans prefer length or correctness? asks whether a reward model learned length instead of quality. This question is about the mathematical effect of changing the sequence log-probability used in a preference objective after the pairs have already been chosen.