Not from pairwise training alone. In a common Bradley-Terry formulation, the probability of preferring A to B for prompt X depends on sigmoid(r(X,A) - r(X,B)). Add any prompt-dependent constant to both rewards for X and that probability stays exactly the same. Comparisons within X cannot identify that offset. The shared neural network may happen to choose a scale and offset, but the training objective does not guarantee that raw 9 for X and raw 6 for Y mean comparable quality. The InstructGPT paper uses human rankings of outputs for a given prompt to train a reward model. Cross-prompt routing is an additional use that needs its own validation.

First I would pin down the actual decision. If we choose between two candidates for one request, score both under the same full prompt and compare their margin. If we are deciding which different requests need escalation, we need a calibrated risk or expected-improvement estimate, not an unexamined reward value. Difficulty, prompt format, language, answer length and the model's uncertainty can all shift the score distribution. A score trained for relative preference is not automatically a probability of correctness, safety, or customer harm.

Build an independently labeled set of requests spanning those slices. Measure whether raw score predicts the outcome we actually care about across requests, using reliability plots or risk by score bucket, then check calibration on held-out customers and changed prompt distributions. A separate error predictor or a rubric with absolute labels may suit escalation better. If you calibrate the reward model empirically, name the scope and refresh it as the data changes. Calibration cannot create information about a rare safety failure absent from its labels.

An interviewer may ask, “But can't RL use raw rewards across prompts?” It can, with choices about normalization and baseline that affect optimization. That fact does not make a particular raw score a calibrated cross-prompt quality measure for serving. The reward model ranks answers the same. Why did multiplying scores change the RL policy? asks why rescaling reward scores changes the policy even when within-prompt ranking is preserved. Here the proposed mistake happens before optimization, when incomparable scores are treated as an absolute measure across different tasks.