Only if the supervised positions match the objective. A causal LM can compute next-token loss over every token in a transcript. That may be a deliberate language-modeling objective, but for assistant SFT we often want user and system turns to be conditioning context with ignored labels, and assistant response tokens as targets. TRL's SFT documentation describes assistant-only loss and notes that, when explicit labels or masks are absent, labels can default to input IDs. The chat-template documentation explains the generation markers used to construct assistant masks. Check the actual library and template version in the run, not just a configuration flag.

I would render three raw conversations through the complete path to the batch that reaches the model. Print token IDs, decoded tokens, role boundaries, labels after the causal shift and the loss mask. Count supervised tokens by role, by dataset source and by sequence position. Watch for a template upgrade that removed {% generation %} boundaries, a custom collator that discarded assistant_masks, or packing that put a boundary in the wrong place. A metric averaged over all labeled tokens can improve simply because user prompts are repetitive or easy to predict. It is not an assistant-quality metric.

Fix the intended objective explicitly. If we want to supervise assistant tool calls, include those assistant tokens. If tool responses are environment observations, mask them unless a separate objective justifies predicting them. Decide how to treat assistant analysis, final text and turn terminators according to the actual serving format. Preserve role masks through truncation and packing, and assert that only the intended tokens have nonignored labels after all preprocessing. Make a small canary batch and verify gradients come from the expected spans before launching the full run.

Someone may argue that predicting user text teaches the model to understand users. It might have a role in a separate continued-pretraining recipe. That is a choice to evaluate, not what this SFT run claimed to train. Compare fixed-token-budget runs with the intended masks on held-out instruction tasks, tool use and role adherence, not just train loss. The tool fine tune is fluent, but the model invents tool results. Which tokens were training targets? concerns tool outputs that became SFT targets and encouraged invented observations. Here the broader failure is that the user and system turns were also part of the supervised target.