Model and Inference Engineering · Principal
The tool returned JSON. Why is the RL policy getting credit for those tokens?
The question
Interview question
An agent policy chooses a tool call, receives JSON from the tool, and then answers the user. During RL training, the trajectory is serialized as one long token sequence. The trainer assigns policy loss across every token after the prompt, including the JSON observation. Success reward rises on training tasks, but the policy begins emitting fake tool-result-looking text. Where is the action boundary?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The policy controls its own generated messages and tool-call arguments. It does not sample the environment's tool response. Policy-gradient logic applies log probability to actions drawn from the policy, conditioned on observations. If the loss treats observed JSON as policy action, the update rewards the model for predicting or emitting bytes it did not choose in the rollout. That can teach an unwanted imitation of tool output and distort the ratio between current and rollout policy probabilities. The policy-gradient formulation makes the action log-probability boundary explicit. A recent agent-trajectory study also distinguishes agent-authored action tokens from environment observation tokens in supervised training. It explores observation supervision as a deliberate SFT choice, which is a different objective from incorrectly counting observations as sampled policy actions during RL.
I would inspect the trajectory as structured events, not as a string. Mark assistant reasoning or message tokens, tool-call arguments, tool response, user continuation and terminal reward. Compare the action mask used for log probabilities and KL with the actual sampler provenance. The tool response should be context for later decisions, with its bytes and tool version recorded. If it is inserted into the same model input, that does not make it a policy action. Recalculate a small batch's loss with a hand-checked mask and measure how much of the current gradient came from observation spans.
The tool may be stochastic or fail. That makes attribution more important. Two identical policy actions can produce different observations and eventual rewards. The policy should learn how to choose and respond under that environment distribution, not learn to pretend that one favorable observation was under its control. Preserve environment outcomes for replay and use a consistent definition of action and state for any importance ratios. If the model is intentionally trained in a separate auxiliary task to predict observations, label and weight that objective explicitly and evaluate whether it helps.
The tool fine tune is fluent, but the model invents tool results. Which tokens were training targets? covers SFT labels that teach a model to invent tool results. This question concerns an RL rollout where the environment supplied a real result, but the training implementation put those observation tokens inside the policy loss. Fixing the mask is a correctness change to the objective, so rerun evaluation rather than simply comparing old and new reward curves.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →