I would inspect one example as token IDs, roles and labels. In next-token training, a rendered transcript is just a sequence unless the chat template and label mask tell the objective which parts the assistant is meant to produce. If the trainer puts loss on a tool's response body, it is rewarding the model for predicting the tool's bytes. That can make the model good at imitating a returned JSON object and bad at waiting for the real tool. A low training loss may partly reflect easy, repetitive tool payloads.

Take a small exchange. The user asks for account status. The assistant emits a structured get_account call. The tool returns {"status":"suspended"}. The assistant explains the result. At inference, the assistant can learn to choose the call and write the final explanation, but the external tool result is supplied by the runtime. I would set target labels only on the assistant spans intended for generation, including the tool call syntax if that is generated by the assistant. User, system and tool response tokens should be context with ignored loss for that SFT objective. The exact boundaries depend on the model's chat template and trainer. TRL's SFT documentation documents assistant_only_loss and its need for generation markers in the chat template. A flag name does not prove your custom template marks the right spans.

Print a decoded batch with each token's role, loss mask and shifted target. I would check the first token after every boundary, especially where assistant tool call ends and tool content begins. Confirm that packing did not join two unrelated examples in a way that trains a fake transition. Verify that tool calls have the same schema and special tokens at training and serving time. If the tool output is an assistant role in the stored data, fixing the mask alone may not be enough. The data itself claims the model spoke those bytes.

There is another possible cause. Even with a correct mask, the final assistant reply might contain the answer from the earlier tool result. A model can learn a shortcut from highly predictable tasks and skip the call. So after fixing the data contract, evaluate actual rollouts with the tool available, unavailable and returning a surprising value. Count calls, fabricated tool messages, answers that disagree with the observed result, and behavior when the result is delayed. Feed a different status for the same user question in a controlled fixture. The final answer should track the result received, not the common value from training.

I would compare the old and corrected fine tunes on this fixture before a broad run, while keeping a small sample of raw token traces for human review. The fix is not merely to make the model “less hallucinated” through another instruction. First establish which bytes the optimizer was paid to generate and which bytes only the runtime can provide.