Softmax for one query row weights score s_j by exp(s_j - m) where m is the maximum over all visible keys. Subtracting the maximum does not change the probability because the same factor cancels from numerator and denominator. It prevents unnecessarily large exponentials. With tiles, the maximum may rise later. If an earlier tile used old maximum m_old, its running denominator and value sum have to be rescaled by exp(m_old - m_new) before adding terms from the new tile. Without that rescale, the two sums are expressed on different scales. A naive direct exp(s) may also overflow before the later maximum is known.

The FlashAttention paper uses an exact tiled attention computation with online softmax statistics. “Exact” here means the same mathematical attention operation, subject to normal floating-point differences in execution order. It is not a license to concatenate independently normalized tile outputs. Each query row must carry its running maximum, normalized exponential sum, and weighted value accumulator across all permitted tiles. Masked keys contribute zero, and a fully masked row needs a defined convention rather than an accidental 0/0.

I would compare the kernel with a high-precision reference on one deliberately constructed row. Put moderate scores in the first tile and the highest score in the last tile. Then reverse them, vary tile size, insert masked entries, and test score ranges near the dtype's limits. Capture the running maximum and denominator after each tile. This shows whether the bug is overflow, missing rescaling, wrong mask semantics or a separate QK scaling mistake. The attention scores use model width instead of head width. What changes? covers using model width instead of head width in the QK scale. Here the scale can be correct and the online normalization still wrong.

What if subtracting the local maximum from every tile makes all exponentials finite? That only makes each tile internally stable. It does not make their local denominators comparable. The answer should explain what gets multiplied when a new maximum arrives, and why both the denominator and accumulated weighted values must move to the new scale.