No. A pairwise ranking metric is invariant to a positive rescaling, but the RL objective generally uses reward magnitude. In a simple reward-minus-KL objective, scaling reward by ten while keeping the KL coefficient fixed makes improvement in reward ten times more valuable relative to staying near the reference policy. The optimizer may move further to exploit the reward model. The InstructGPT paper describes reward with a per-token KL penalty and a coefficient governing that tradeoff. The same ordering does not imply the same optimum under an unchanged regularizer.

I would first identify the actual training pipeline. Are scores centered, standardized, clipped, normalized per batch, or transformed into advantages? Does a value model learn the new scale? Is the KL coefficient adapted automatically? Some transformations can reduce the effect of simple rescaling, and clipping can introduce a different nonlinearity. We should not assert a tenfold policy change just because raw scores grew tenfold. But a silent calibration change is not harmless until we show the full objective and optimizer inputs are effectively invariant. Log raw and processed rewards, KL, advantage distributions, value errors, entropy, rollout lengths and behavior by task slice.

To make the effect concrete, compare a candidate policy change that improves expected reward by 1 and incurs a KL cost weighted at 2. It is unattractive under the original scale. After multiplying reward by 10 with the penalty unchanged, that same change looks attractive under this simplified objective. Real PPO or another RL method has clipping, estimates and finite optimization, so do not treat the toy comparison as an exact policy prediction. It shows why rank preservation is insufficient.

I would version the reward model and its calibration in the training recipe. Replay a fixed rollout batch through both score paths, verify any normalization, and run a controlled small training comparison at matched effective reward-to-KL pressure. Independent human or task outcomes decide whether the new policy is better. The reward model score rises while people prefer the old assistant. What did optimization learn? asks why optimizing reward can exploit a proxy even when the reported score improves. This asks a narrower first-principles question: why a score transformation that preserves every preference pair can still change the optimization target.