Model and Inference Engineering · Principal
One answer scores 9 and another scores 6. Can you compare them across prompts?
The question
Interview question
A reward model was trained on human comparisons between two answers to the *same* prompt. A serving router now sends answer A for customer question X and answer B for unrelated question Y through that reward model. A scores 9, B scores 6. The team concludes A is better and uses the raw score to decide which customer should receive an expensive second pass. Is that conclusion justified?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Not from pairwise training alone. In a common Bradley-Terry formulation, the probability of preferring A to B for prompt X depends on sigmoid(r(X,A) - r(X,B)). Add any prompt-dependent constant to both rewards for X and that probability stays exactly the same. Comparisons within X cannot identify that offset. The shared neural network may happen to choose a scale and offset, but the training objective does not guarantee that raw 9 for X and raw 6 for Y mean comparable quality. The InstructGPT paper uses human rankings of outputs for a given prompt to train a reward model. Cross-prompt routing is an additional use that needs its own validation.
First I would pin down the actual decision. If we choose between two candidates for one request, score both under the same full prompt and compare their margin. If we are deciding which different requests need escalation, we need a calibrated risk or expected-improvement estimate, not an unexamined reward value. Difficulty, prompt format, language, answer length and the model's uncertainty can all shift the score distribution. A score trained for relative preference is not automatically a probability of correctness, safety, or customer harm.
Build an independently labeled set of requests spanning those slices. Measure whether raw score predicts the outcome we actually care about across requests, using reliability plots or risk by score bucket, then check calibration on held-out customers and changed prompt distributions. A separate error predictor or a rubric with absolute labels may suit escalation better. If you calibrate the reward model empirically, name the scope and refresh it as the data changes. Calibration cannot create information about a rare safety failure absent from its labels.
An interviewer may ask, “But can't RL use raw rewards across prompts?” It can, with choices about normalization and baseline that affect optimization. That fact does not make a particular raw score a calibrated cross-prompt quality measure for serving. The reward model ranks answers the same. Why did multiplying scores change the RL policy? asks why rescaling reward scores changes the policy even when within-prompt ranking is preserved. Here the proposed mistake happens before optimization, when incomparable scores are treated as an absolute measure across different tasks.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →