Not by treating them as fresh evidence of success under the new tool. A trajectory is a sequence of observations, chosen actions, tool outcomes and rewards under a particular environment. If issue_credit(amount) used to take dollars and now takes cents, the same action text can have a different effect. If a tool field became required, yesterday's apparently successful call might fail today. OpenAI's RL introduction defines a trajectory in terms of actions and environment transitions. NVIDIA's agent training harness concepts likewise identify tools and APIs as part of the environment. The version binding is an engineering conclusion from those definitions, not a guarantee that every training framework enforces it.

Keep the tool schema, implementation version, observation builder, policy or prompt version, reward or verifier version, and any fixture state with the rollout. Then classify the change. A field rename with provably identical semantics might permit an explicit migration of old supervised examples. A change in units, permission checks or side effects needs fresh execution or strong evidence of equivalent transitions. Do not feed a recorded success reward to a new action that was never executed under the new rules.

The algorithm matters. An on-policy PPO update cannot simply treat a large archive of old trajectories as current rollout data. An off-policy or supervised method may intentionally use old data, but it still needs valid labels and a supported action vocabulary. Compare a sample of migrated trajectories against a versioned test fixture, and run fresh rollouts on new-schema tasks before trusting the gain. Track invalid-call rate, authorized action success, outcome reward and regressions on old valid cases.

RL workers are three checkpoints behind. Are their rollouts still usable? asks whether policy weights on rollout workers are too stale for the learner. An agent resumes after approval. Why does its saved tool call now fail? asks what a paused production run does when its tool schema changed. This page is about the training dataset. A successful historical trace is evidence about the old environment, not proof that the new agent can repeat it.