"A trillion tokens" tells me how many token positions the model trained on. It does not tell me how those positions were arranged. Attention compares positions within a sequence. With dense causal attention, one sequence of length L has roughly L squared query-key work, even though there are only L tokens. Holding the total token count fixed, moving from many 4K sequences to fewer 32K sequences can increase the attention arithmetic per token by roughly 8 times. That is only the attention part, not an 8 times prediction for the whole training job.

There is another trap in the usual training estimate of about 6 times parameters times tokens. That is a useful rough estimate when the model and sequence regime make parameter matmuls dominant. It hides the length-dependent attention term. FlashAttention avoids materializing the whole attention matrix in high bandwidth memory, but it still computes exact attention. Better memory traffic does not turn dense attention into linear arithmetic.

I would ask for actual sequence-length histograms after packing and truncation, attention implementation, tokens per optimizer step, measured forward and backward time, device utilization and achieved versus estimated FLOPs. Two runs can have the same nominal maximum context but very different length distributions. Padding, attention masks, checkpointing and the shape of the kernel also affect wall time. Measure a short and a long batch on the same hardware before attributing the entire bill to the attention formula.

If the interviewer asks whether we can simply pack shorter documents into 32K blocks, I would ask which attention boundary we want. A block-diagonal mask prevents one document from reading another and changes which token pairs are computed by an implementation that can skip masked regions. A plain causal mask over the full packed block permits cross-document attention. That is the separate correctness problem in Sequence packing doubled training throughput. Why can one document attend to another?. For this run, compare quality at the long-context tasks we actually need, and report both tokens and training FLOPs or device hours. Equal token budgets are not equal experiment budgets.