Model and Inference Engineering · Staff
The tiled attention kernel returns NaNs. What did it forget when the row maximum changed?
The question
Interview question
An engineer writes a tiled attention kernel so the full score matrix never lands in high-bandwidth memory. Each tile computes exponentials, adds them to a running denominator, and accumulates weighted values. Short random tests pass. A longer sequence with a very large score in a later tile produces NaNs or a wildly different output. What invariant did the running calculation miss?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Softmax for one query row weights score s_j by exp(s_j - m) where m is the maximum over all visible keys. Subtracting the maximum does not change the probability because the same factor cancels from numerator and denominator. It prevents unnecessarily large exponentials. With tiles, the maximum may rise later. If an earlier tile used old maximum m_old, its running denominator and value sum have to be rescaled by exp(m_old - m_new) before adding terms from the new tile. Without that rescale, the two sums are expressed on different scales. A naive direct exp(s) may also overflow before the later maximum is known.
The FlashAttention paper uses an exact tiled attention computation with online softmax statistics. “Exact” here means the same mathematical attention operation, subject to normal floating-point differences in execution order. It is not a license to concatenate independently normalized tile outputs. Each query row must carry its running maximum, normalized exponential sum, and weighted value accumulator across all permitted tiles. Masked keys contribute zero, and a fully masked row needs a defined convention rather than an accidental 0/0.
I would compare the kernel with a high-precision reference on one deliberately constructed row. Put moderate scores in the first tile and the highest score in the last tile. Then reverse them, vary tile size, insert masked entries, and test score ranges near the dtype's limits. Capture the running maximum and denominator after each tile. This shows whether the bug is overflow, missing rescaling, wrong mask semantics or a separate QK scaling mistake. The attention scores use model width instead of head width. What changes? covers using model width instead of head width in the QK scale. Here the scale can be correct and the online normalization still wrong.
What if subtracting the local maximum from every tile makes all exponentials finite? That only makes each tile internally stable. It does not make their local denominators comparable. The answer should explain what gets multiplied when a new maximum arrives, and why both the denominator and accumulated weighted values must move to the new scale.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →