Model and Inference Engineering · Principal
The reward differences are tiny. Why can GRPO give them a large update?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Look at how the advantage is made before looking at the policy optimizer. With group-relative normalization, completions for one prompt are scored together. Each reward loses the group mean, then is divided by the standard deviation of that same group's rewards. TRL's GRPO documentation describes this as its scale_rewards="group" behavior.
Take two groups of two completions. One has rewards [0, 1]. Another has [0.49, 0.51]. After centering and dividing by each group's own standard deviation, the better completion in each group gets the same normalized sign and magnitude under the same standard-deviation convention. The second raw gap is fifty times smaller. The normalization has deliberately removed that scale. In a real training step, sequence lengths, clipping and other terms affect the final parameter update, so identical advantages do not imply identical gradient vectors.
There is a more subtle effect with sparse binary rewards. A group with one success among several failures can give the success a larger normalized advantage than a group with a more balanced mix. This changes which prompt difficulties carry weight. A group where every answer gets the same reward has no within-group ranking signal at all. A numerical epsilon may prevent division by zero, but it cannot invent a preference.
I would log raw reward spreads, group standard deviations, fraction of all-equal groups, advantage distributions and update contribution by task slice. If tiny differences come from noisy reward judgments, normalizing them can give that noise a loud voice. Check reward reliability before tuning the optimizer.
There are choices. TRL also exposes scale_rewards="batch", which keeps per-prompt centering but uses a batch-wide standard deviation, and scale_rewards="none", which skips standard-deviation scaling. Neither setting fixes a bad reward function. A batch-wide scale makes task weight depend on batch composition, while no scaling makes raw reward units matter. The interview answer is to name the implicit weighting policy, test it on the task mix, and report it with the training result.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →