The policy controls its own generated messages and tool-call arguments. It does not sample the environment's tool response. Policy-gradient logic applies log probability to actions drawn from the policy, conditioned on observations. If the loss treats observed JSON as policy action, the update rewards the model for predicting or emitting bytes it did not choose in the rollout. That can teach an unwanted imitation of tool output and distort the ratio between current and rollout policy probabilities. The policy-gradient formulation makes the action log-probability boundary explicit. A recent agent-trajectory study also distinguishes agent-authored action tokens from environment observation tokens in supervised training. It explores observation supervision as a deliberate SFT choice, which is a different objective from incorrectly counting observations as sampled policy actions during RL.

I would inspect the trajectory as structured events, not as a string. Mark assistant reasoning or message tokens, tool-call arguments, tool response, user continuation and terminal reward. Compare the action mask used for log probabilities and KL with the actual sampler provenance. The tool response should be context for later decisions, with its bytes and tool version recorded. If it is inserted into the same model input, that does not make it a policy action. Recalculate a small batch's loss with a hand-checked mask and measure how much of the current gradient came from observation spans.

The tool may be stochastic or fail. That makes attribution more important. Two identical policy actions can produce different observations and eventual rewards. The policy should learn how to choose and respond under that environment distribution, not learn to pretend that one favorable observation was under its control. Preserve environment outcomes for replay and use a consistent definition of action and state for any importance ratios. If the model is intentionally trained in a separate auxiliary task to predict observations, label and weight that objective explicitly and evaluate whether it helps.

The tool fine tune is fluent, but the model invents tool results. Which tokens were training targets? covers SFT labels that teach a model to invent tool results. This question concerns an RL rollout where the environment supplied a real result, but the training implementation put those observation tokens inside the policy loss. Fixing the mask is a correctness change to the objective, so rerun evaluation rather than simply comparing old and new reward curves.