Model and Inference Engineering · Principal
The reward model praises a safe answer. Did it read the last 2,000 tokens?
The question
Interview question
An RL policy starts producing long answers. The reward dashboard says they are increasingly safe and helpful. Human reviewers find an unsafe instruction near the end of several highly scored answers. The reward model's input is capped at 1,024 tokens. What exactly was scored?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First inspect how the prompt and answer are assembled and which end the tokenizer truncates. A classifier scores only the tokens it receives. If the tail never reaches it, a high score cannot certify that tail. Truncating from the other side may drop the user request instead. Either way, the score has a narrower meaning than the dashboard claims. Transformers' truncation documentation describes the max_length behavior, but the actual training code decides how prompt and response are concatenated and which side is kept.
I would log lengths and the effective visible span for reward scoring, without exposing sensitive answer text in routine telemetry. For a small audited sample, put a decisive safety violation or factual correction just before and just after the cutoff. If the score is unchanged when the latter appears, this reward signal cannot see it. Compare complete-answer human labels and a longer-context or segmented evaluator. Also check whether the reward model's own training pairs had the same length policy. A different truncation rule at reward-training and RL-scoring time creates another mismatch.
Then choose a policy that matches the task. You might cap generation at a length the reward model can fully inspect, use a reward model that covers the whole response, or combine a full-answer safety check with the quality reward. For longer outputs, segment-level checks need to preserve context and cannot simply average reassuring chunks. Measure whether the fix changes both reward and real task outcomes. Do not declare success because the average score fell after the grader finally saw previously hidden mistakes.
Could the policy learn to place bad material after the cutoff? It could, especially under repeated optimization, but that requires evidence here. The immediate proven failure is simpler: the score did not cover the complete answer. The reward model prefers longer answers. Did humans prefer length or correctness? studies reward models preferring length. This is a visibility boundary in the reward input, even if the reward model's judgments on visible text are otherwise sound.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →