Not from those observations. In full causal self-attention, each token can attend to earlier tokens. The number of query-key interactions grows roughly with the square of sequence length. FlashAttention tiles the work and computes the softmax without writing the whole score matrix to high-bandwidth memory. That can save a great deal of memory traffic and allow longer sequences to fit. It does not make full attention's pairwise computation disappear. The original FlashAttention paper calls it exact attention with IO-aware tiling. “Exact” matters here. We have changed how the work is scheduled and stored, not replaced full attention with a sparse approximation.

A useful back-of-the-envelope check is to double prompt length while holding batch, model, hardware and kernel path fixed. The pairwise attention work for a full causal layer grows by about four times, while projections and MLP work grow roughly with token count, about two times. Whole-model prefill is a mix of those terms, so its latency need not rise by exactly four. At 64K, the attention term may become dominant, or some other stage may be worse. A context window number alone does not tell us which.

I would profile the request in pieces. Measure tokenization and multimodal preprocessing if present, queue time, data transfer, projection and MLP kernels, attention kernels, and KV writes. Check the actual kernel dispatch and dtype for the 64K shape. A mask type, head dimension, architecture or dynamic shape can select a different implementation. Inspect occupancy and HBM traffic rather than assuming every slow prefill is arithmetic. Compare a short and a long prompt on the same path and benchmark the attention kernel independently with matching batch, heads, head dimension, sequence length and causal mask.

If the full-attention work is genuinely the limit, there is no free switch that preserves every possible dependency while making it linear. We can use a model trained and evaluated for local or sparse attention, compress or retrieve evidence before inference, prefill a shared prefix once when valid, or change the product to process a document in stages. Each can lose information or add complexity. A chunked prefill schedule can improve responsiveness and coexistence with decode traffic, but it does not remove total attention work for the same full-context model. Test the actual answer tasks, especially distant evidence and exceptions, before treating a faster prompt path as equivalent.

The interviewer may point to an almost flat GPU memory graph as evidence that the algorithm made attention cheap. It shows the quadratic intermediate was not materialized, which is good. It says little about how many multiply-adds and softmax operations were performed or how long other requests waited. I would ship on measured quality, time to first token and fleet goodput for real prompt lengths, with the chosen context semantics clearly stated.