The causal mask explains the observation. It stops a token from seeing later positions. A token in B is later than every token in A, so the ordinary lower triangular mask lets it read A. An end of sequence token tells the model about a boundary. It does not remove the earlier keys from attention. Packing filled what used to be padding, but the team also changed which information is available to each prediction.

The first token of example B sees A under a causal mask, but not when segment boundaries are enforced
The end marker identifies a boundary, but only segment-aware attention prevents B from reading A.

I would ask what objective was intended before calling this a bad run. Some language model pretraining deliberately concatenates documents and trains next token prediction across an end of document marker. In that recipe, B seeing A is part of the data construction. Here the team said these are independent examples. If they are separate instruction and answer pairs, unrelated customers, or evaluation examples whose answer must not depend on a preceding sample, a causal mask alone violates that contract. This is an objective and data boundary decision, not a rule that every packed run needs isolation.

For independent examples, I want three boundaries checked separately. First, attention: queries in B should only reach keys in B. A block diagonal causal mask can express that, or a variable length attention kernel can use cumulative sequence lengths to avoid computing cross segment attention at all. NVIDIA's packed sequence guide describes that cu_seqlens route. Second, labels: the boundary position, often the end marker after A, must not be trained to predict the first token of B unless that transition is explicitly intended. A masked attention row does not fix an incorrectly shifted target. Third, positions and loss: check position IDs, prompt token loss masking, and any padding or truncation at the join. Resetting positions alone does not block attention, and a correct attention mask does not prove the target or loss mask is correct.

I would inspect one packed batch before debating a full run. Log the input token IDs, example IDs per position, boundaries, shifted labels, loss mask and the exact attention metadata sent to the chosen kernel. On a tiny two example fixture, compare logits and gradients for B with A changed to an unrelated document. If the recipe promises isolation, B's output should not change because A changed. Test the actual fused kernel path used in training, including checkpointed backward, rather than only a CPU mask constructed by the data loader. The Megatron Core dataset configuration exposes independent controls for attention, positions and end of document loss, which is a useful reminder that one flag does not establish all three properties.

If the team says validation loss improved, that does not settle it. Validation may use the same packing and reward the leaked context. Run an isolated validation set and compare the suspected run with the previous checkpoint on tasks that reflect the intended objective. Also check source family and permission boundaries. An example from another tenant in the same packed sequence makes the error much more serious than a minor modeling choice.

My default for this scenario is to stop promotion, fix the packer and kernel contract, and retrain from the last checkpoint before contaminated batches entered the optimizer. Whether the earlier compute is usable depends on the objective, exposure and measured behavior. I would not wave it through because the token throughput chart looks good. The token IDs can be perfectly stable while the model is trained with the wrong visibility and target boundaries.