Model and Inference Engineering · Staff
GRPO rewards the easy answer and punishes the best hard answer. Were the groups mixed?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
GRPO compares several sampled completions for the same prompt. It centers and scales their rewards within that group to form relative advantages. If a batching bug groups completions from different prompts, the baseline loses its meaning. A correct answer to a hard proof with reward 0.6 can be assigned a negative advantage next to easy arithmetic answers scoring 1.0, even if 0.6 was the best proof attempt. The DeepSeekMath paper introduces GRPO and its group-relative reward construction.
Consider two prompts. The hard prompt has rewards 0.2, 0.6. The easy prompt has 0.9, 1.0. Within the hard prompt, 0.6 is the better sample. Within the easy prompt, 0.9 is worse than 1.0. If all four are normalized as one group, the 0.6 sample falls below the combined mean of 0.675 and gets a negative centered advantage, while 0.9 gets a positive one. The optimizer is now being told the opposite of the intended within-prompt comparison for both of those samples. A standard deviation term changes magnitude, not those signs.
I would inspect the group key through sampling, async reward collection, shuffling and learner minibatches. The stable key must identify the exact prompt and relevant context, not just task category or row position. A timeout that removes one sample must not cause the next prompt's sample to fill the empty slot. Log prompt IDs, reward vector, group mean and standard deviation, per-token action mask and resulting advantages for a tiny deterministic batch. Test mixed difficulty and out-of-order reward responses.
Could we make reward scores globally calibrated and compare prompts? That would be a different objective and would need its own justification. It does not make an incorrectly assembled GRPO group correct. Every sampled solution gets zero reward. Why is the GRPO training run doing almost no learning? covers groups where all rewards are zero, so there is almost no within-group signal. One answer scores 9 and another scores 6. Can you compare them across prompts? warns that reward scores across prompts may not be comparable. This case shows a concrete systems bug that turns that comparability problem into the wrong gradient sign.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →