Look at the margin that enters the sigmoid. For one prompt, compare the policy's log probability of chosen versus rejected against the same difference under the reference. Call that difference of differences m. The usual DPO loss is -log sigmoid(beta × m). The DPO paper derives this objective, and TRL exposes beta as a configuration parameter.

Raising beta does not simply “train harder.” At m = 0, the derivative magnitude with respect to m grows with beta. For pairs that already have a positive margin, a large beta × m pushes the sigmoid near one and their gradient toward zero. For strongly negative margins, it can amplify the pressure in the other direction. Which pairs still drive the update changes as training progresses. That is why a run can move aggressively at first and then stop learning from many pairs even though the dataset did not change.

The reference model and the way sequence log probabilities are summed matter too. A margin of 2 means nothing by itself if one run uses different token masks, lengths or a different reference. Compare beta experiments using the unscaled margin distribution, loss per slice, fraction of saturated positive and negative pairs, policy-reference KL or another drift measure, and held-out human preference. A flat training loss can be easy pairs saturated while difficult pairs still fail.

Would I immediately lower beta? I would first verify the exact objective and preprocessing. A wrong chosen/rejected order, masked completion, or prompt included in one side can mimic a beta problem. Once those are ruled out, sweep beta with a fixed reference and comparable token budget. Watch for both under-learning and excessive movement away from the reference. There is no single beta that transfers automatically across model, data and length distribution.

The useful interview move is to reason through beta × m and its gradient, not to memorize that a bigger beta is “more conservative.” It changes which examples have leverage at each point in training.