Evaluation and Quality · Principal
The new model wins on completed eval cases. Where did its timed-out cases go?
The question
Interview question
A candidate agent is more accurate than production on the evaluation dashboard. It also times out more often on difficult repository tasks. The scoring job compares only rows where both agents returned a final answer. Can we call it a win?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
First ask what the deployment promises. If a user needs a working answer before a deadline, a timeout is an outcome, not a missing row. Google's SRE guidance treats a request past a committed deadline as an error when that is the service promise. Let the assigned test set contain 1,000 tasks. Suppose the old system finishes 950 with 760 correct answers and the new one finishes 800 with 720 correct. Conditional accuracy rises from 80% to 90%. Successful answers per assigned task fall from 76% to 72%. Both numbers are true. Only one reflects the user's chance of getting a correct answer within this budget.
The same problem shows up in a paired comparison. If the candidate's hardest tasks disappear, the set of scored pairs is selected by the candidate's behavior. You can report quality conditional on completion as a diagnostic, but keep the original task IDs and define a primary end-to-end metric on the assigned population. Report completion, deadline misses, quality among completed tasks, and cost separately. Where a failure is an infrastructure outage rather than the model's behavior, apply a predeclared rerun rule symmetrically and preserve both the first attempt and the rerun outcome. Do not improvise exclusions after seeing which model won.
I would stratify by task difficulty and length, then inspect the missing cases. Maybe the candidate is much better when it finishes but needs a larger time budget. That could be a valid product choice if the product can afford it. Measure a common deadline curve, rather than moving only the candidate's deadline until its dashboard looks good. Confidence intervals should resample the independent assigned tasks, with paired outcomes on the same task whenever both systems were assigned it. The model looks accurate on labeled cases. What happened to cases still waiting for an outcome? covers labels that arrive late from the world. Here the ground truth exists. The candidate's own timeout removes its hard cases from the scoring denominator.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →