Model and Inference Engineering · Principal
The RL run rewards short proofs. Did the model finish, or did generation stop it?
The question
Interview question
In a reasoning fine-tune, a verifier reads each generated proof and gives a score. Scores rise, but the model increasingly emits a promising first half and never closes the argument. Most rollouts hit the configured generation limit. What did the reward actually label?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Reaching max_completion_length is an external cutoff. It is not the model choosing an EOS token. If the grader awards partial credit for plausible steps, the optimizer may learn to spend its limited tokens on a persuasive opening, while never learning how to finish within the product budget. If the grader assigns zero to every cutoff, it may punish good but longer attempts as if the model chose to stop badly. Both policies can be defensible for particular objectives, but neither should be an accidental consequence of conflating cutoff and completion.
Keep the termination reason with each rollout: EOS, other stop condition, context limit, wall-clock timeout or cancellation. Score mathematical validity on genuinely complete proofs, and set an explicit policy for incomplete ones. A product with a hard 512-token limit may reasonably penalize an answer that cannot finish within it. For a training experiment trying to discover whether the model can solve the problem, use a larger cap or analyze capped samples separately. TRL's GRPO documentation exposes truncated-completion measurements and an option to mask them from the loss. That option avoids one noisy target, but masking every long attempt can itself bias which tasks and trajectories drive learning. TRL's PPO guidance describes an explicit missing-EOS penalty. Neither switch decides the product objective for you.
I would plot reward and solved-rate against termination reason, generated length and problem difficulty. Then manually inspect high-reward cutoffs and run the same checkpoint with a higher evaluation cap, keeping the training and serving caps visible. If solutions now complete, we have a budget or incentive issue to resolve. If they still trail off, the model has a deeper capability problem. The code solution passes, but its verifier timed out. What reward should it get? asks what to do when an external code verifier times out. This question is about the model generation cutoff and the false impression of a completed, rewarded proof.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →