Evaluation and Quality · Principal
The agent solves 90% with ten tries. What will one production run solve?
The question
Interview question
The honest answer is that the 90% alone does not tell us. Did each task get ten independent sampled attempts, one agent with a retry loop, or ten candidates followed by a test-based selector? Was success “at least one of ten passes the hidden test,” or “the deployed selector picked a passing candidate”? Those are different systems with different budgets.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
pass@10 means at least one of ten candidates succeeds under the evaluator. pass@1 is the success rate for a single candidate from a defined policy. The HumanEval paper makes the distinction explicit and shows why sampling more candidates can raise the first measure. The “at least one” metric assumes an oracle that can recognize success for reporting. If production has no reliable selector, a correct candidate sitting among nine wrong ones may not help its user.
Do not back-calculate pass@1 as if every retry were independent with one shared success probability. Tasks have very different difficulty, attempts can be correlated, and an agent retry may learn from failed tests or change its tools and context. Even the mathematical relation under independent identical trials would describe a specific sampling model, not an observed production policy.
I would rerun the eval with the production contract: one initial attempt if that is the actual limit, the same tools and timeouts, and the same selection or verification available to users. Report first-attempt success, success after each allowed retry, wall time, tokens, dollars, and any human intervention. Keep the task denominator fixed. If users can genuinely run ten attempts with automatic tests, measure the selected answer after ten under that exact budget, not just whether an oracle saw one good draft.
An interviewer may ask whether retries are “cheating.” No. They are a product capability with a cost and a latency. The mistake is advertising an oracle-assisted ten-try result as the probability that one unattended run succeeds. Show the curve as budget grows and let the deployment choice determine which point matters.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →