Model and Inference Engineering · Staff
The attention map points to the right token. Why does the model still answer wrongly?
The question
Interview question
A team inspects one head while debugging a model that answers the wrong version of a policy. The query token assigns its largest attention weight to the correct version number. They say retrieval inside the model worked, so the answer must be a decoding bug. What does that attention weight actually prove?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It proves a particular query and key produced a relatively high score in that head at that layer. It does not prove the information in the version token was copied into the output, preserved by later layers, or used to choose the answer.
The basic calculation helps. For one head, attention weights come from softmax(QKᵀ / √d), with masking where required. The head output is those weights multiplied by V. A high weight on position 42 means that position contributes its value vector, not its visible text directly. The value projection may encode something other than the literal version number. Other positions also contribute. Heads are concatenated, projected, added through residual paths and processed by later attention and MLP blocks before the final logits. The original Transformer paper gives this multi-head computation. Looking at one colorful heatmap skips most of the path to the answer.
I would first check that the displayed map belongs to the actual request, layer, head and query position. Is the model seeing the correct version number in its tokenized input, with the expected document boundaries? Does the head attend to that token when the two policy versions are swapped but the rest of the prompt is held fixed? Then trace the resulting value contribution and the residual stream around the layer. Compare the correct-answer logit or an appropriate answer score before and after replacing that source span. A counterfactual edit to the evidence is more informative than declaring one attention weight an explanation. It still needs care, because editing tokens can shift positions and change many activations at once.
There are several genuine failure points. The head may attend to “v4” because it is a useful section marker while a later MLP uses a memorized rule from v3. A different head may read a stale exception. The correct source may be present but the prompt asks the model to combine it with a condition it missed. Or decoding could be involved, but we would need the logits and sampling trace to show that the correct answer had a high probability and was lost only at selection. Inspecting the final text alone cannot distinguish these.
If the interviewer says every head attends to the correct span, I would still ask what values they carry and what later layers do. Attention patterns can be diagnostic, but they are not a certificate of causal use. Research on attention as explanation makes that caution explicit. I would grade the model on a set of paired policy versions and decisive exceptions, then use interventions and logit comparisons to locate where the version information stops influencing the answer. The fix might be clearer source construction, better training examples, or a model change. The heatmap by itself does not choose one.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →