I would ask whether "equal weight" means equal example weight or equal token weight. Those are different objectives. Suppose we average token loss inside each answer, then average the two answer losses. Every token in the 20-token answer contributes one twentieth of half the batch loss. Every token in the 2,000-token answer contributes one two-thousandth of half. A short-answer token has 100 times the weight of a long-answer token. That may be intentional if each user task should count equally, but it is not the same as treating every supervised token equally.

If instead we sum losses over all nonignored target positions and divide by the total 2,020 positions, each token contributes equally. The long answer then has 100 times the total weight of the short one. PyTorch's cross-entropy documentation defines how reduction and ignored targets affect the denominator. A custom loop can add a second mean across examples after already computing per-example means, so inspect the actual tensors and masks rather than relying on a config label.

For an instruction fine-tune, neither objective wins automatically. Per-example weighting prevents one very long response from dominating a batch of short requests. Per-token weighting reflects the number of supervised positions and may better match next-token likelihood. The choice also interacts with how much of a long answer is truncated, how many prompt and tool tokens are masked, and whether a batch packs multiple examples together. I would specify the desired population objective, then compare the effective contribution by task type and length bucket.

A useful check is a two-example microbatch with known token losses. Compute the expected scalar and gradient contribution by hand, then compare with the trainer. Record total valid targets, number of examples and the reduction applied at every stage. Ranks process different token counts. Why does averaging their losses change the objective? concerns uneven token counts across distributed ranks after local averaging. Here one rank can be perfectly implemented and still give short examples much greater per-token influence because of the chosen per-example objective.