Splitting individual preference pairs at random can put one pair for a prompt in training and another pair for the same prompt in evaluation. The second pair has new row IDs, perhaps new response text, yet the model has already learned the task, likely answer pattern or preferred refusal for that prompt. Held-out pair accuracy then tests a narrow ability to rank more candidates on familiar prompts. It is not evidence of generalization to new user problems. RewardBench 2 emphasizes new human prompts when constructing a more rigorous reward-model evaluation.

I would reconstruct the data lineage before changing the split. Canonicalize and group by prompt plus relevant system instruction, tool context and task fixture. Look for paraphrased prompts, repeated coding tasks with changed variable names, and the same source answer under different IDs. Split by those groups, then check how many evaluation prompts have near neighbors in training. Also hold out domains or customers if that is the deployment claim. A strict split may lower the reported score. That is useful information about the previous estimate, not a regression of the model itself.

The two tests answer different questions. A same-prompt, new-candidate set is valuable for testing whether the reward model can rank additional outputs for tasks it already knows. A genuinely new-prompt set tests transfer. Report both with their names and do not market the first as the second. Inspect downstream effects too: the reward model can win pairwise accuracy yet guide the policy toward verbosity, brittle refusals or reward hacking when optimized hard. The reward model prefers longer answers. Did humans prefer length or correctness? covers length preference as one such confound.

If the data comes from an interactive agent, the “prompt” is not just the last user sentence. The available tools, retrieved observations and environment version define the decision context. A split keyed only on text can still leak the same underlying task through different trajectories. Train and eval have different document IDs. Why did the model see the test answers? concerns train/eval leakage from duplicate documents despite different IDs. This page applies the isolation boundary to preference pairs and the claim a reward model makes about unseen decisions.