An aggregate score weights cases by how many examples each slice contributes. Suppose ordinary FAQ cases make up 95% of the eval and improve from 90% to 94%, while payment-dispute cases make up 5% and fall from 80% to 50%. Overall accuracy moves from 89.5% to 91.8%. That is a genuine aggregate gain and a serious regression for the small slice. The arithmetic is simple enough to calculate directly. Google's guidance on slicing data warns that mix shifts and aggregation can hide or reverse patterns.

Before deciding whether to ship, define slices by failure consequence, not only by traffic volume. Payment disputes, security-sensitive actions and unsupported answers may need separate release gates even if they are rare. Report numerator, denominator and uncertainty for each slice. A five-case segment is a signal to collect more evidence, not a precise estimate of a 30-point loss.

I would rerun old and new systems on the same fixed cases with the same grader and source snapshot. Then look at current production mix separately. If traffic composition changed, the aggregate can move even when within-slice behavior did not. In the example, both effects are possible, so do not attribute the gain to the model until the weighting and paired outcomes are clear. Review actual failed cases to see whether the root cause is retrieval, model behavior, policy or tool execution.

The release decision can be a targeted rollback, routing rule or gated rollout while the high-risk slice is fixed. A rule that sends those users to an older model has its own operational and fairness costs, so monitor its routing and expiry. Keep a critical-slice gate in the eval rather than hoping the overall average will protect it.

If the interviewer says “But the new model wins for most users,” agree with the arithmetic and ask what promise the product makes to the remaining users. Staff and Principal judgment is choosing which regression is acceptable, with evidence and ownership, not treating a single mean as the full product.