First inspect how the prompt and answer are assembled and which end the tokenizer truncates. A classifier scores only the tokens it receives. If the tail never reaches it, a high score cannot certify that tail. Truncating from the other side may drop the user request instead. Either way, the score has a narrower meaning than the dashboard claims. Transformers' truncation documentation describes the max_length behavior, but the actual training code decides how prompt and response are concatenated and which side is kept.

I would log lengths and the effective visible span for reward scoring, without exposing sensitive answer text in routine telemetry. For a small audited sample, put a decisive safety violation or factual correction just before and just after the cutoff. If the score is unchanged when the latter appears, this reward signal cannot see it. Compare complete-answer human labels and a longer-context or segmented evaluator. Also check whether the reward model's own training pairs had the same length policy. A different truncation rule at reward-training and RL-scoring time creates another mismatch.

Then choose a policy that matches the task. You might cap generation at a length the reward model can fully inspect, use a reward model that covers the whole response, or combine a full-answer safety check with the quality reward. For longer outputs, segment-level checks need to preserve context and cannot simply average reassuring chunks. Measure whether the fix changes both reward and real task outcomes. Do not declare success because the average score fell after the grader finally saw previously hidden mistakes.

Could the policy learn to place bad material after the cutoff? It could, especially under repeated optimization, but that requires evidence here. The immediate proven failure is simpler: the score did not cover the complete answer. The reward model prefers longer answers. Did humans prefer length or correctness? studies reward models preferring length. This is a visibility boundary in the reward input, even if the reward model's judgments on visible text are otherwise sound.