Model and Inference Engineering · Staff
PPO clipped almost no tokens. Did we recompute the old log probabilities after updating?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
That is the first thing I would inspect. PPO's ratio for a sampled action is exp(new_logp - old_logp). The denominator is the probability assigned by the policy that generated the rollout. If we calculate both terms with the current weights after an optimizer update, the ratio is close to one by construction. The clipping statistics look calm even when the policy has moved a long way from the behavior policy. OpenAI's PPO explanation defines the old and new policies, and its implementation stores old log probabilities with the rollout and computes the new probabilities during training.
Suppose an agent selected a tool with probability 0.1 under the rollout policy. After training, the current policy gives it 0.5. The ratio should be 5. Replacing the recorded old probability with a fresh 0.5 produces a ratio of 1. The precise threshold and sign of the advantage decide how PPO clips that sample, but the bug has already removed the intended trust-region signal. A low reported clip fraction is not proof of a conservative update.
Store action tokens, the exact context visible when each action was sampled, the behavior-policy log probability, policy version and mask for tokens actually chosen by the policy. Do not mix in tool-output or prompt tokens. Keep the behavior log probabilities fixed through the optimization epochs for that rollout. At the first update, replay a sample before any optimizer step and check its ratio is approximately one. After a controlled parameter change, check that the stored denominator remains fixed while the numerator moves. Compare ratio distribution, approximate KL, clip fraction and held-out behavior, not a single metric.
If rollout workers are stale, this is a separate concern. The recorded denominator must still describe their actual sampling policy, and the learner needs a deliberate rule for admitting or discarding stale data. RL workers are three checkpoints behind. Are their rollouts still usable? covers that policy-lag decision. This question catches a more basic accounting error that can make even a fresh rollout look safe while optimization changes the model.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →