pass@10 means at least one of ten candidates succeeds under the evaluator. pass@1 is the success rate for a single candidate from a defined policy. The HumanEval paper makes the distinction explicit and shows why sampling more candidates can raise the first measure. The “at least one” metric assumes an oracle that can recognize success for reporting. If production has no reliable selector, a correct candidate sitting among nine wrong ones may not help its user.

Do not back-calculate pass@1 as if every retry were independent with one shared success probability. Tasks have very different difficulty, attempts can be correlated, and an agent retry may learn from failed tests or change its tools and context. Even the mathematical relation under independent identical trials would describe a specific sampling model, not an observed production policy.

I would rerun the eval with the production contract: one initial attempt if that is the actual limit, the same tools and timeouts, and the same selection or verification available to users. Report first-attempt success, success after each allowed retry, wall time, tokens, dollars, and any human intervention. Keep the task denominator fixed. If users can genuinely run ten attempts with automatic tests, measure the selected answer after ten under that exact budget, not just whether an oracle saw one good draft.

An interviewer may ask whether retries are “cheating.” No. They are a product capability with a cost and a latency. The mistake is advertising an oracle-assisted ten-try result as the probability that one unattended run succeeds. Show the curve as budget grows and let the deployment choice determine which point matters.