Model and Inference Engineering · Principal
The RL policy stays within its KL budget. Why does it break tool calls on rare prompts?
The question
Interview question
A post-training run limits its average divergence from a reference policy. The reported KL stays close to target and reward improves. Yet rare prompts requiring a structured tool call start producing invalid arguments. The team argues that a small mean KL should rule out a large behavior change. Does it?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
No. An average over a sampled prompt distribution and generated tokens cannot guarantee every slice behaves similarly. Most training prompts may be ordinary prose, so large drift on a small tool-calling slice can contribute little to the overall metric. Even within a response, a few decision-critical tokens can flip a tool name, argument value, or refusal while the majority stay close to reference. InstructGPT's paper uses a KL penalty to discourage excessive drift during policy optimization. That is a regularizer in a particular training setup, not a proof that all downstream contracts remain intact.
I would ask what “KL” in the dashboard actually measures. Is it a sampled log-probability ratio to the reference, an exact next-token distribution divergence on selected positions, or a training loss term? Over which prompts, policy version, reference version, sequence lengths and tokens is it averaged? Does the reported mean hide a heavy tail? Then inspect per-slice behavior for tool-heavy tasks, rare schemas, multi-turn handoffs and different languages. Compare parse validity, authorization, tool argument semantics and successful task completion before and after training. A passing grammar rate alone can still hide wrong values.
If the important failures occur on prompts almost absent from the rollout distribution, raising the global KL penalty is a blunt fix. It may constrain useful improvements everywhere while missing the real data gap. Add representative prompts to the training and held-out evaluations, protect tool formats with interface-level checks, and consider targeted constraints or a mixture strategy. Before changing the algorithm, replay examples against the reference and new policy with the same tools and sampling settings. Some apparent policy regressions are actually parser, template or serving changes. Prove where the first divergence appears.
The pushback is that per-slice KL itself may be noisy for rare prompts. True. Report sample sizes and uncertainty. A low mean does not justify a blanket safety claim, and a high slice metric alone does not prove a user-visible failure. Pair distribution diagnostics with concrete behavioral tests and incident severity. The reward model score rises while people prefer the old assistant. What did optimization learn? asks why optimizing reward can exploit a proxy. This question asks why a global drift budget fails to bound a small, important interface slice, even if reward optimization is otherwise working as intended.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →