Reaching max_completion_length is an external cutoff. It is not the model choosing an EOS token. If the grader awards partial credit for plausible steps, the optimizer may learn to spend its limited tokens on a persuasive opening, while never learning how to finish within the product budget. If the grader assigns zero to every cutoff, it may punish good but longer attempts as if the model chose to stop badly. Both policies can be defensible for particular objectives, but neither should be an accidental consequence of conflating cutoff and completion.

Keep the termination reason with each rollout: EOS, other stop condition, context limit, wall-clock timeout or cancellation. Score mathematical validity on genuinely complete proofs, and set an explicit policy for incomplete ones. A product with a hard 512-token limit may reasonably penalize an answer that cannot finish within it. For a training experiment trying to discover whether the model can solve the problem, use a larger cap or analyze capped samples separately. TRL's GRPO documentation exposes truncated-completion measurements and an option to mask them from the loss. That option avoids one noisy target, but masking every long attempt can itself bias which tasks and trajectories drive learning. TRL's PPO guidance describes an explicit missing-EOS penalty. Neither switch decides the product objective for you.

I would plot reward and solved-rate against termination reason, generated length and problem difficulty. Then manually inspect high-reward cutoffs and run the same checkpoint with a higher evaluation cap, keeping the training and serving caps visible. If solutions now complete, we have a budget or incentive issue to resolve. If they still trail off, the model has a deeper capability problem. The code solution passes, but its verifier timed out. What reward should it get? asks what to do when an external code verifier times out. This question is about the model generation cutoff and the false impression of a completed, rewarded proof.