The verifier observed a timeout, not a failed test. Whether that timeout is evidence about the candidate depends on where time was spent. A hard per-solution CPU limit may legitimately be part of the task specification. A wait for an overloaded worker says almost nothing about solution correctness. If the pipeline collapses both into zero, it feeds systematic false negatives to the policy. Research on RL with imperfect verifiers analyzes how false positives and false negatives change training, and this incident gives a concrete operational cause. The policy can learn to favor fast-to-grade outputs or short programs even when the intended reward is functional correctness.

I would split outcome states: tests passed, tests failed with evidence, solution exceeded a defined resource limit, verifier infrastructure failed, and unknown after timeout. Record queue wait, setup time, compile time, test execution time, resource usage, image version and test-suite revision. A timeout budget should begin at the intended boundary. Replay a sample of timeouts on healthy workers with a larger diagnostic limit, but keep the production task limit unchanged for the final label. The estimated false-negative rate may differ by problem difficulty and solution style, so a global correction factor is unlikely to be enough.

For training, do not silently assign wrong to unknown. Quarantine or retry infrastructure failures under a bounded budget. If an answer still exceeds the task's specified execution limit when run fairly, that is a valid failure for that task. If the verifier is intermittently flaky, use repeated grading or a trusted reference implementation to audit the label. Track the fraction of rollouts with unknown outcomes and who bears their cost, since simply dropping all timeouts can also bias the training distribution. The batch-selection policy should be explicit and reviewed on hard problems.

The interviewer may ask whether rerunning gives the policy another chance to game tests. The candidate code and test inputs must be immutable during verification, and retries must be independent of candidate edits. A coding policy earns full reward by changing the tests. What was the verifier allowed to trust? covers a policy earning reward by changing its tests. Here it may get no reward despite a correct immutable solution because the verifier never completed its work. A reliable reward channel needs a third state when evidence is absent.