The name can mislead us. In Megatron-style sequence parallelism, certain activations such as LayerNorm and dropout are split along the sequence dimension across tensor-parallel ranks. It does not mean each rank owns only its fraction of the entire attention computation and all its KV activations. Megatron Core's context-parallel documentation makes this distinction and introduces context parallelism for partitioning the network inputs and activations across the sequence dimension.

At 64K, I would look at the actual peak, not infer it from total allocated memory after the OOM. Is the peak attention workspace, QKV activations retained for backward, an unoptimized attention matrix, communication buffers, or a large microbatch? FlashAttention-style kernels avoid materializing a full sequence-by-sequence attention matrix, but they do not remove the need to compute and backpropagate over long contexts. Check whether the chosen attention kernel is actually used for this shape and mask. A silent fallback can make memory jump even though the flags say sequence parallelism is on.

If attention states are the issue, context parallelism can distribute sequence work more broadly, but it has to exchange remote KV information or otherwise compute attention across sequence shards. That trades memory for communication and implementation complexity. Activation recomputation trades memory for repeated compute. Smaller microbatches, a shorter sequence, or a different attention pattern are other choices, each with a different effect on the training objective and throughput. I would test one representative 64K step with peak memory broken down by rank and timed collectives before choosing.

Can we just add more sequence-parallel ranks? That can shrink the subset of activations SP actually shards, but it cannot solve a separate attention peak by arithmetic magic. More tensor-parallel ranks may also raise collective cost. Name the specific memory owner before selecting the parallelism axis.