Model and Inference Engineering · Staff
Training loss is excellent. Why does generation fail when the attention mask sees the future?
The question
Interview question
A decoder-only model is trained to predict the next token. After an attention-kernel rewrite, training loss falls sharply. During generation, the model rambles and cannot complete simple prompts. The training mask accidentally permits position `i` to attend to position `i + 1`, whose token is the label at `i`. Why is the low loss a warning rather than a success?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
With teacher forcing, the full training sequence is already present. The model should predict the next token using only current and previous positions. If its hidden state at i can inspect the embedding of the token at i + 1, it can read the answer it is being graded on. It can learn a shortcut that does not exist during autoregressive generation, where future tokens have not been supplied. A tiny mask error can therefore produce excellent next-token loss and poor standalone generation. The Transformer paper describes masking in the decoder so a prediction cannot depend on later positions. The diagram below shows the single forbidden cell that exposes the next-token label.
I would inspect the actual attention mask handed to the kernel after batching, packing and any fused conversion, not just the intended mask in config. For a four-token sequence, capture the permitted key positions for each query and verify that query 0 sees key 0 but never key 1. Some models shift labels inside the loss, so trace token IDs, label indices and query/key indices together. An off-by-one elsewhere could produce a similar symptom. Compare reference and optimized kernels on a tiny tensor where changing a future token must not change earlier-position logits. That is a stronger invariant than checking tensor shapes or an average perplexity number.
Then rerun loss under a correct mask and test free generation. If the model was trained for a long time with leakage, a serving-only mask fix may remove the shortcut and make quality look worse before retraining. We need a clean checkpoint or a validated recovery recipe, not a claim that the training curve was real. Test packed boundaries too. Sequence packing doubled training throughput. Why can one document attend to another? covers one document attending into another during packing. Here even within one document, a position sees the label it should predict. Both can coexist, so mask tests should cover each boundary separately.
What if the model still had a good held-out validation loss? If validation used the same leaky mask and full teacher-forced sequences, it measured the same shortcut. Use both correct-mask next-token evaluation and autoregressive tasks with unseen prompts. The general principle is simple: training and inference must offer the model the same information at each prediction step.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →