Model and Inference Engineering · Staff
The fine-tune loss falls. Why is the model learning the token after next?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
For a causal language model, the logits at position i should predict token i+1. The input at i can see itself and earlier tokens, not the future. With tokens A B C D, the training pairs are A -> B, B -> C, C -> D. A falling loss only says the model is getting better at the targets the loss actually compares. It does not say those are the targets generation needs.
One nasty integration bug is shifting labels in the data pipeline and then passing them to a model that shifts again inside its loss. The pipeline turns labels into B C D ignore. The model compares its logit at A with the next label, which is now C. It has been taught A -> C. The Hugging Face causal language modeling guide shows a language-model collator that prepares labels from inputs, and a documented model implementation explicitly pairs logits[..., :-1, :] with labels[..., 1:]. Do not assume every custom model or trainer follows that contract. Read the loss path you actually run.
I would print one tiny tokenized batch after collation, including token IDs, labels, attention mask and ignored positions. Then check which target is paired with each logit at the point cross-entropy is called. Use a four-token synthetic example with known IDs and manually compute the loss. Compare the framework loss with a reference loss from unshifted labels. Also check prompt/completion masking: when the first answer token is supervised, the predictor is the logit at the preceding prompt token. Shifting a mask can silently remove that boundary token even when the rest of the answer is aligned.
If training loss is already low, ask where it was evaluated. A validation job that uses the same wrong collator measures the same two-step task and can look reassuring. Recompute held-out next-token loss with a separately checked alignment, then sample continuations and inspect the first answer token. Fix the duplicate shift and restart the affected fine-tune from a clean checkpoint. Adjusting the learning rate cannot turn a two-step target into a one-step target.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →