The number of rows is not the number of independent pieces of evidence. Cases from one customer can share document structure, product vocabulary, permissions, an incident, or the same root failure. If a new parser fixes one customer's PDF, it may improve 200 near-duplicate questions together. Pretending those are 200 unrelated wins makes the uncertainty look artificially small for a claim about new customers. The cluster bootstrap analysis by Indeed's engineering team demonstrates why repeated responses per input require resampling at the independent input level. Here the likely cluster is customer, possibly with nested documents or tasks.

I would first ask what population the claim targets. Is it the next question from these same 50 customers, the next customer, or the entire market? The sampling unit follows that choice. For a new-customer claim, hold out customers before prompt tuning and resample customers, keeping related questions together within each bootstrap sample. Because both models answered the same cases, compare their paired differences rather than two unrelated rates. Report the spread across customers, not only the aggregate. If one large customer supplies half the rows, say whether the target metric weights traffic or customers equally. Both are legitimate questions with different answers.

Fifty clusters still cannot give a precise guarantee for every industry or rare failure. Show the per-customer improvement distribution, how many actually improved or regressed, and the range of document and query types. A grouped train/test split is essential if variants of the same source appear on both sides. Do not bootstrap questions after leaking one document's variants into tuning and expect statistics to repair the leakage. An eval can also be large and still unrepresentative, so audit source selection and label quality before arguing about confidence intervals.

If the interviewer says production traffic comes mostly from those existing customers, I would use a second analysis weighted for that traffic mix. It estimates a different outcome and may have more information within a customer, though temporal drift and correlated failures still matter. For launch, I would combine a customer-held-out result with a production-weighted result and inspect severe regressions. “10,000 examples” sounds strong. The useful statement is what was sampled, what can vary together, and to whom the two-point gain should generalize.