Model and Inference Engineering · Principal
One GRPO candidate timed out. Why did the other candidates' advantages change?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
GRPO compares several sampled completions for the same prompt. A common formulation centers each candidate's reward by its group's mean and scales by the group's standard deviation. So a candidate's advantage is not just a property of its own reward. DeepSeekMath introduced group-relative policy optimization, and TRL's GRPO documentation describes group rewards and handling of truncated completions in a concrete trainer. A request timeout is not automatically the same thing as a completion that hit a token limit. Check the implementation's group assembly and reward rules.
Say two completed candidates score 0 and 1. Centering and scaling those observed rewards gives roughly -1 and +1, using population standard deviation. If a third candidate had completed with reward 2, the rewards 0, 1 and 2 would center on 1, making the middle candidate's advantage zero. We do not know the timed-out candidate's counterfactual reward. The example shows why silently dropping it changes the relative target for the other candidates. Assigning it zero is also a choice that treats infrastructure failure as a low-quality policy answer.
I would log prompt ID, intended group size, generated candidate IDs, terminal cause, reward availability and which candidates entered normalization. Then inspect whether hard, long, tool-heavy completions time out more often. If so, the observed groups are a selected sample. Dropping failures can reward short easy completions, even when the task needed more work. For a recoverable infrastructure timeout, retry with the same policy snapshot and generation settings if the algorithm's sampling contract permits it. Otherwise mark the group incomplete and decide explicitly whether to resample the whole group, skip its update, or use a well-defined censored-outcome objective. Each choice has cost and bias.
The interviewer may ask if a timed-out candidate deserves zero reward because it failed the user's deadline. If meeting that deadline is part of the task, yes, a defined timeout penalty can be the right reward. But a crashed rollout worker is not a model failure. Keep service deadline failure separate from missing telemetry and use the same rule in training and evaluation. GRPO rewards the easy answer and punishes the best hard answer. Were the groups mixed? concerns completions from different prompts being mixed into one group. This case keeps prompt grouping correct and tests what missing members do to the group's statistics.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →