Model and Inference Engineering · Staff
The attention scores use model width instead of head width. What changes?
The question
Interview question
A multi-head attention refactor divides query-key dot products by `√d_model` instead of `√d_head`. The model still trains, shapes are correct, and the loss does not become NaN. An engineer calls the scale a harmless constant. Would you agree, and how would you find the effect?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
No. Softmax is sensitive to the relative differences in its input scores. Multiplying all scores by a smaller factor changes how concentrated the attention distribution is. If d_model = 4096 and each of 32 heads has d_head = 128, the usual divisor is about 11.3. Dividing by √4096 = 64 instead makes logits about 5.7 times smaller, assuming the underlying query and key vectors are unchanged. More positions then receive comparable weight. That can look like the model is attending broadly rather than selecting a needed token. It is not a shape error, so a unit test that checks only tensor dimensions will pass.
Why use the head dimension? If query and key components are roughly independent with unit variance at initialization, their dot product sums d_head terms, so its variance grows with d_head and its standard deviation with √d_head. The scale controls logit magnitude before softmax. The Transformer paper describes this reason and notes that large unscaled dot products can push softmax into regions with small gradients. Real trained activations may not satisfy those simple assumptions, but the head dimension remains the natural reference for the dot product each head actually computes.
I would take a fixed batch with representative lengths and compare the old and refactored kernel at the point immediately before softmax. Confirm the head shape and scale in the exact production path, including fused attention. Log score distribution, attention entropy and gradients by head and layer, then compare output vectors and task behavior on examples that require choosing a narrow piece of evidence. A lower entropy is not automatically better, and a higher entropy is not automatically wrong. We need to see whether the change harms the intended computation. Include short and long contexts, since a hotter softmax spreads mass over more tokens as the candidate set grows.
If this was a refactor of a trained checkpoint, I would restore the original scale and compare logits on the same inputs. Calling it a harmless constant would require proof that another learned parameter or calibration exactly compensated, which a code change after training cannot assume. If the model was trained from scratch with the new scale, it may adapt through Q and K projections, but it is a different recipe and needs its own stability and quality review. A silently changed temperature in every head can shift gradients, head specialization and downstream behavior even when loss initially looks ordinary.
The interviewer may ask whether 1/√d_head is the only valid formula. It is not. Architectures can deliberately use other attention scaling, normalization or learned temperatures. The answer is to name the intended architecture and test it. Here the engineer claims equivalence while changing the softmax temperature by a large factor. That claim does not survive the first equation.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →