Evaluation and Quality · Principal
Your eval has 10,000 cases from 50 customers. How many independent wins did you measure?
The question
Interview question
An assistant improves from 82% to 84% accuracy on 10,000 questions. The team reports a very narrow confidence interval by treating every question as an independent draw. But most questions are variants of the same customer documents, and only 50 customers contributed. Would you trust the interval?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The number of rows is not the number of independent pieces of evidence. Cases from one customer can share document structure, product vocabulary, permissions, an incident, or the same root failure. If a new parser fixes one customer's PDF, it may improve 200 near-duplicate questions together. Pretending those are 200 unrelated wins makes the uncertainty look artificially small for a claim about new customers. The cluster bootstrap analysis by Indeed's engineering team demonstrates why repeated responses per input require resampling at the independent input level. Here the likely cluster is customer, possibly with nested documents or tasks.
I would first ask what population the claim targets. Is it the next question from these same 50 customers, the next customer, or the entire market? The sampling unit follows that choice. For a new-customer claim, hold out customers before prompt tuning and resample customers, keeping related questions together within each bootstrap sample. Because both models answered the same cases, compare their paired differences rather than two unrelated rates. Report the spread across customers, not only the aggregate. If one large customer supplies half the rows, say whether the target metric weights traffic or customers equally. Both are legitimate questions with different answers.
Fifty clusters still cannot give a precise guarantee for every industry or rare failure. Show the per-customer improvement distribution, how many actually improved or regressed, and the range of document and query types. A grouped train/test split is essential if variants of the same source appear on both sides. Do not bootstrap questions after leaking one document's variants into tuning and expect statistics to repair the leakage. An eval can also be large and still unrepresentative, so audit source selection and label quality before arguing about confidence intervals.
If the interviewer says production traffic comes mostly from those existing customers, I would use a second analysis weighted for that traffic mix. It estimates a different outcome and may have more information within a customer, though temporal drift and correlated failures still matter. For launch, I would combine a customer-held-out result with a production-weighted result and inspect severe regressions. “10,000 examples” sounds strong. The useful statement is what was sampled, what can vary together, and to whom the two-point gain should generalize.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →