Hold the patch fixed and rerun the evaluator. If the pass result flips, the variation cannot be attributed to a new model output in those runs. Maybe the test depends on timing, external data, order, a random seed or a shared resource. Pytest's flaky-test guidance defines the basic phenomenon: a test can intermittently pass or fail under apparently unchanged code.

For an agent benchmark, this is more than a noisy CI inconvenience. The verifier supplies the success label, and sometimes the training reward. A patch that passes one out of three runs might be a correct solution exposed to a bad test, or an incorrect solution exploiting a flaky assertion. Treating the best run as truth inflates success. Treating a single failure as truth can reject a good patch. Repeatedly trying until green also creates an extra attempt budget, separate from the agent's ability to write code.

I would pin the repository commit, dependency versions, test image, input data and environment. Run the same artifact several times with test order and parallelism varied in a controlled way. Record failing assertions, seed, timing and external calls. Compare with a known reference patch and the original failing repository. If the reference also flips, isolate or repair the test. If only this patch flips, inspect whether it introduced a race or relies on an unstable behavior. “Flaky test” is not automatically an excuse for the candidate.

For reporting, separate stable passes, stable failures and unresolved nondeterministic cases. State the rerun policy before looking at candidate results and apply it equally across models. Do not silently exclude unstable cases from one model's denominator or award pass on any green run. For training, quarantine uncertain rewards until the verifier can be trusted, otherwise the policy may learn to chase lucky outcomes.

The deeper question is what the benchmark measures. It should measure whether the agent solved the task under a defined execution contract. If the contract itself changes from run to run, a point estimate of agent quality carries false precision.