Model and Inference Engineering · Staff
The tool fine tune is fluent, but the model invents tool results. Which tokens were training targets?
The question
Interview question
A supervised fine tune uses transcripts that contain user turns, assistant tool calls, JSON returned by tools and final assistant answers. After training, the model often writes a plausible tool response as its own message without making a call. The trainer's loss fell, and the transcripts look sensible when printed as text. Where would you look first?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would inspect one example as token IDs, roles and labels. In next-token training, a rendered transcript is just a sequence unless the chat template and label mask tell the objective which parts the assistant is meant to produce. If the trainer puts loss on a tool's response body, it is rewarding the model for predicting the tool's bytes. That can make the model good at imitating a returned JSON object and bad at waiting for the real tool. A low training loss may partly reflect easy, repetitive tool payloads.
Take a small exchange. The user asks for account status. The assistant emits a structured get_account call. The tool returns {"status":"suspended"}. The assistant explains the result. At inference, the assistant can learn to choose the call and write the final explanation, but the external tool result is supplied by the runtime. I would set target labels only on the assistant spans intended for generation, including the tool call syntax if that is generated by the assistant. User, system and tool response tokens should be context with ignored loss for that SFT objective. The exact boundaries depend on the model's chat template and trainer. TRL's SFT documentation documents assistant_only_loss and its need for generation markers in the chat template. A flag name does not prove your custom template marks the right spans.
Print a decoded batch with each token's role, loss mask and shifted target. I would check the first token after every boundary, especially where assistant tool call ends and tool content begins. Confirm that packing did not join two unrelated examples in a way that trains a fake transition. Verify that tool calls have the same schema and special tokens at training and serving time. If the tool output is an assistant role in the stored data, fixing the mask alone may not be enough. The data itself claims the model spoke those bytes.
There is another possible cause. Even with a correct mask, the final assistant reply might contain the answer from the earlier tool result. A model can learn a shortcut from highly predictable tasks and skip the call. So after fixing the data contract, evaluate actual rollouts with the tool available, unavailable and returning a surprising value. Count calls, fabricated tool messages, answers that disagree with the observed result, and behavior when the result is delayed. Feed a different status for the same user question in a controlled fixture. The final answer should track the result received, not the common value from training.
I would compare the old and corrected fine tunes on this fixture before a broad run, while keeping a small sample of raw token traces for human review. The fix is not merely to make the model “less hallucinated” through another instruction. First establish which bytes the optimizer was paid to generate and which bytes only the runtime can provide.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →