Frozen weights do not mean a deterministic forward pass. If the reference stays in training mode and has active dropout, two passes over the same prompt and completion can produce different log probabilities. The chosen-minus-rejected reference margin in DPO then moves even though neither the preference data nor the weights changed. The TRL DPO trainer settings expose disable_dropout, defaulting to true in the documented version, and precompute_ref_log_probs. A custom loop can violate that intended scoring setup.

Before changing the learning rate, I would score a hand-checked pair repeatedly with the reference alone. Hold token IDs and masks fixed. Log the two completion log probabilities, their difference, model mode and dropout configuration. Then use evaluation mode for the frozen reference and check whether the repeated scores become stable within the numerical behavior of the execution path. Also make sure a training loop has not accidentally called train() on a parent module that contains the reference. A zero gradient and a frozen parameter flag do not switch off dropout.

Precomputing the reference scores can make this easier to spot, but only if the cache is built with the intended deterministic model and keyed by checkpoint, tokenizer, template, truncation and scoring mask. Caching one accidental dropout draw per example makes that draw permanent. If both policy and reference are stochastic, the noise can appear on both sides of the DPO margin. Establish the intended objective explicitly and measure margin variance before claiming an optimizer instability.

What if inference still has small logit differences after dropout is disabled? Distributed kernels and precision can vary, so I would compare against a tolerance and isolate whether changes in chosen-minus-rejected margin are large enough to affect updates. This is different from The preference pairs stayed the same. Why did a new DPO reference model change the policy?, where the reference checkpoint intentionally changes, and The DPO reference and policy score different prompt tokens. What does the margin mean?, where the two models score different prompt tokens. Here a supposedly fixed reference can change its score between two evaluations of the exact same token sequence.