Model and Inference Engineering · Principal
RL workers are three checkpoints behind. Are their rollouts still usable?
The question
Interview question
An asynchronous post-training system has rollout workers generating agent trajectories while a learner updates the policy. A queue contains trajectories from checkpoints three updates old. The reward is valid for the action that happened. The training team wants to mix them into the next update without tagging their origin because the prompts are still relevant. What would you require before accepting those gradients?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The reward can remain a valid observation of what happened in an environment. The gradient estimator is a different question. A rollout was sampled from a particular behavior policy, with particular sampling settings, tool schema and environment state. If the learner treats those actions as if its current policy sampled them, it has changed the distribution in the objective without accounting for it. The size of the error depends on how far those policies differ, not on a fixed count of three checkpoints.
Each trajectory needs a policy version and the log probabilities of the sampled actions under that behavior policy, with prompt, masks, temperature and tokenization pinned. For an agent, also record tool and environment versions and the actual observations after actions. Compute the current-to-behavior probability ratio on the same eligible tokens and look at its distribution, not just its mean. In PPO-like training, importance ratios and clipping limit certain update sizes, but clipping cannot magically make arbitrarily stale, low-support data equivalent to fresh on-policy samples. TRL's asynchronous GRPO documentation explicitly provides a way to tell workers the live policy version so they can tag or discard stale samples. Exact controls depend on the trainer.
I would replay a small set of queued trajectories through the current policy to measure how much probability moved. If the actions now have tiny likelihood, ratios can be unstable and the old behavior may tell us little about what the new policy will do. If the environment has changed, an old reward might not even answer today's task. Put bounds on age, ratio tails and environment compatibility, and quarantine data that fail them. The limit should come from pilot quality and variance, not “three steps is always safe.” Refresh workers or slow the learner if queue depth keeps exceeding the bound. Asynchronous throughput is useful only when the samples still support the chosen algorithm.
There is a second trap in multi-turn agent rollouts. The current policy's probability of later actions is conditioned on earlier observations that followed the old policy's actions. You cannot take a cached outcome from a different path and call it a counterfactual for the current policy. For a verifier reward, a stored trajectory may still be useful for an explicitly off-policy objective or for supervised distillation, with a separate evaluation. Name that algorithm honestly and validate it. Do not quietly feed old traces to a code path written for fresh rollouts.
For release evidence I would compare a tightly synchronized pilot with the asynchronous system at matched compute and environment tasks. Track policy lag in steps and probability space, effective sample contribution after clipping, reward and actual task success on fresh held-out rollouts. If the older data raise a training reward but fresh policy behavior gets worse, the queue has become a way of optimizing yesterday's policy.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →