No. An average over a sampled prompt distribution and generated tokens cannot guarantee every slice behaves similarly. Most training prompts may be ordinary prose, so large drift on a small tool-calling slice can contribute little to the overall metric. Even within a response, a few decision-critical tokens can flip a tool name, argument value, or refusal while the majority stay close to reference. InstructGPT's paper uses a KL penalty to discourage excessive drift during policy optimization. That is a regularizer in a particular training setup, not a proof that all downstream contracts remain intact.

I would ask what “KL” in the dashboard actually measures. Is it a sampled log-probability ratio to the reference, an exact next-token distribution divergence on selected positions, or a training loss term? Over which prompts, policy version, reference version, sequence lengths and tokens is it averaged? Does the reported mean hide a heavy tail? Then inspect per-slice behavior for tool-heavy tasks, rare schemas, multi-turn handoffs and different languages. Compare parse validity, authorization, tool argument semantics and successful task completion before and after training. A passing grammar rate alone can still hide wrong values.

If the important failures occur on prompts almost absent from the rollout distribution, raising the global KL penalty is a blunt fix. It may constrain useful improvements everywhere while missing the real data gap. Add representative prompts to the training and held-out evaluations, protect tool formats with interface-level checks, and consider targeted constraints or a mixture strategy. Before changing the algorithm, replay examples against the reference and new policy with the same tools and sampling settings. Some apparent policy regressions are actually parser, template or serving changes. Prove where the first divergence appears.

The pushback is that per-slice KL itself may be noisy for rare prompts. True. Report sample sizes and uncertainty. A low mean does not justify a blanket safety claim, and a high slice metric alone does not prove a user-visible failure. Pair distribution diagnostics with concrete behavioral tests and incident severity. The reward model score rises while people prefer the old assistant. What did optimization learn? asks why optimizing reward can exploit a proxy. This question asks why a global drift budget fails to bound a small, important interface slice, even if reward optimization is otherwise working as intended.