GRPO compares several sampled completions for the same prompt. It centers and scales their rewards within that group to form relative advantages. If a batching bug groups completions from different prompts, the baseline loses its meaning. A correct answer to a hard proof with reward 0.6 can be assigned a negative advantage next to easy arithmetic answers scoring 1.0, even if 0.6 was the best proof attempt. The DeepSeekMath paper introduces GRPO and its group-relative reward construction.

Consider two prompts. The hard prompt has rewards 0.2, 0.6. The easy prompt has 0.9, 1.0. Within the hard prompt, 0.6 is the better sample. Within the easy prompt, 0.9 is worse than 1.0. If all four are normalized as one group, the 0.6 sample falls below the combined mean of 0.675 and gets a negative centered advantage, while 0.9 gets a positive one. The optimizer is now being told the opposite of the intended within-prompt comparison for both of those samples. A standard deviation term changes magnitude, not those signs.

I would inspect the group key through sampling, async reward collection, shuffling and learner minibatches. The stable key must identify the exact prompt and relevant context, not just task category or row position. A timeout that removes one sample must not cause the next prompt's sample to fill the empty slot. Log prompt IDs, reward vector, group mean and standard deviation, per-token action mask and resulting advantages for a tiny deterministic batch. Test mixed difficulty and out-of-order reward responses.

Could we make reward scores globally calibrated and compare prompts? That would be a different objective and would need its own justification. It does not make an incorrectly assembled GRPO group correct. Every sampled solution gets zero reward. Why is the GRPO training run doing almost no learning? covers groups where all rewards are zero, so there is almost no within-group signal. One answer scores 9 and another scores 6. Can you compare them across prompts? warns that reward scores across prompts may not be comparable. This case shows a concrete systems bug that turns that comparability problem into the wrong gradient sign.