For the usual group-centered advantage, each sample is compared with the mean reward of the samples for the same prompt. Four zeros have mean zero. The centered task advantage for every sample is zero. Four ones have the same problem. Dividing by a standard deviation with an epsilon does not create a preference signal where none exists. DeepSeekMath introduced group-relative policy optimization, and later research on advantage collapse analyzes homogeneous groups. An optimizer may still show movement from a KL term, regularization or implementation-specific losses, so “no task learning signal from these groups” is more precise than “all gradients are zero.”

Before changing the algorithm, I would measure the fraction of all-zero, all-one and mixed-reward groups by problem family and policy checkpoint. Then inspect the verifier. An infrastructure failure encoded as zero is not a difficult problem, as The code solution passes, but its verifier timed out. What reward should it get? explains. If the verifier is sound, a prompt may be too hard for this policy to produce even one correct candidate at the current sampling temperature and group size. In the independent-sample approximation, if per-sample success probability is p and group size is G, the chance of seeing both successes and failures is 1 - p^G - (1 - p)^G. For p near zero, four samples rarely produce a useful contrast. Real samples are correlated, so estimate this from actual grouped rollouts rather than trusting the formula as a forecast.

I would adjust the data curriculum or sampling to create informative groups, with a held-out measure of performance on the original hard distribution. Raising group size costs rollouts and still may fail when p is extremely low. Mix in solvable tasks, improve exploration, or use a validated reward signal with more information than pass/fail. Algorithm variants can use other baselines or anchors, but they change the objective. Do not label an invalid or timed-out answer as correct just to create reward variance.

What if all four are correct? We also lack a relative correctness preference. That can be a sign the task is now too easy for this stage. Track pass rate and effective informative groups together. A training chart that counts four rollouts per prompt is counting work. It has not shown that those rollouts taught the policy anything about the target task.