Because the loss asks a relative question. For one prompt, DPO takes the policy's log probability of the chosen completion minus that of the rejected completion. It then subtracts the same difference under a fixed reference policy, multiplies by beta and applies a logistic loss. It is teaching the policy to prefer chosen over rejected relative to the reference. It is not maximum likelihood training on chosen answers. That distinction is in the original DPO objective, and the fall in absolute chosen likelihood has also been studied in DPO-Shift.

Say the chosen completion's sequence log probability moves from -10 to -11. The rejected completion moves from -12 to -15. The chosen answer got less likely, but the policy's chosen-minus-rejected margin moved from 2 to 4. With the same fixed reference and beta, the DPO loss for this pair improves. These are full completion log probabilities conditioned on the same prompt, with prompt tokens excluded. If you compare a length-normalized diagnostic against an unnormalized training loss, you can manufacture a second apparent contradiction.

First I would plot policy log probability for chosen and rejected separately, their margin, the reference margin, and the resulting DPO logit on a fixed evaluation set. Check masking, tokenization, reference checkpoint, beta and whether an implementation silently changed the length reduction. A falling training loss on changing minibatches cannot by itself tell us whether the same examples improved. Then sample from the new policy. Does it give useful answers, or has it merely learned to suppress a particular kind of rejected answer? Look at held-out preferences, capability regressions and answer length alongside the training curve.

If the chosen likelihood keeps falling and generation quality suffers, I would inspect pair quality and distribution shift before tuning a coefficient. A supervised likelihood term or a different preference objective may be reasonable if preserving a chosen-answer anchor is an actual goal. But adding it changes the objective, so test that change. A low DPO loss is evidence about pair discrimination under this loss. It is not a certificate that the chosen answer became common, or that users now prefer sampled outputs.