Model and Inference Engineering · Principal
Reviewers prefer A to B, B to C, and C to A. What should the reward model learn?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First I would make sure these are three comparisons for the same prompt and rubric. If each pair answered a different user request, the cycle says little about a model that scores responses conditional on their prompts. If it is the same context and the judgments are reliable, no single scalar score can satisfy all three strict inequalities. A score with A greater than B and B greater than C must put A above C. A common Bradley-Terry style pairwise reward model assumes that scalar ordering. Research on preference models beyond Bradley-Terry studies cases where a scalar ranking misses intransitive or stochastic judgments.
Do not call the reviewers wrong as soon as a cycle appears. With one vote per edge, noise alone can create it. Repeat the comparisons with randomized order and independent reviewers. Record the rubric they used. Perhaps A is concise, B is thorough and C is cautious, and people weigh those differently. Check whether the cycle persists within each rubric and reviewer group or only after aggregating unlike populations. If the evaluations disagree on safety constraints, do not average that away into a neat single ranking.
For training, label uncertainty is part of the data. A scalar reward model can fit the majority pattern approximately but cannot make all three preferences true at once. Measure held-out pairwise error by prompt type and reviewer population. If the product genuinely needs multiple objectives, keep separate criteria or a conditional preference model, and decide explicitly which trade-off governs a given task. A human adjudication or policy rule may be needed for nonnegotiable safety boundaries.
The interviewer may ask whether a larger reward model will fix it. More capacity helps only if the inconsistency came from missing context that we can actually supply. It cannot satisfy a mathematically inconsistent strict ordering for one fixed context with a single scalar. Annotators split 50/50 on refusal. What should the preference label teach the model? deals with disagreement on one refusal example. Here each pair may have a clear winner while the three winners together expose a limit of the assumed reward representation.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →