Evaluation and Quality · Principal
Safety reviewers found 11% violations. Why isn't the production violation rate 11%?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Ask how the review queue was sampled. If it deliberately oversamples suspected high-risk requests, its plain average describes the review queue, not all production requests. That is often the right way to find rare failures efficiently. It is the wrong denominator for a population-rate claim. CDC's sample-weighting tutorial explains why oversampled groups need weights to make population estimates representative. The same sampling logic applies here, though the product's actual strata and inclusion probabilities must be known.
Say one percent of production requests are in a predefined high-risk stratum and 99 percent are in a lower-risk stratum. Auditors review 500 from each. They find violations in 20 percent of high-risk cases and 2 percent of lower-risk cases. The unweighted review rate is 11 percent. The estimated production request rate is 0.01 × 0.20 + 0.99 × 0.02 = 0.0218, or 2.18 percent, provided those within-stratum samples are random enough and the stratum proportions are correct. Neither number says how many people were harmed. That requires a separate outcome definition and unit of analysis.
I would publish both the weighted overall estimate and the stratum-specific rates, with uncertainty. Preserve inclusion probabilities, account for repeated requests from the same user and distinguish review nonresponse from true negatives. If the risky stratum is defined by a fallible classifier, report its coverage and look for unsafe requests outside it. Unreviewed or timed-out items cannot be silently counted as safe. As the traffic mix changes, update the weights rather than carrying last month's 1 percent forward.
A Principal-level decision might still block a release because the high-risk stratum is at 20 percent even if the weighted overall rate looks small. The purpose of weighting is accurate communication about the population, not permission to average away a serious cohort. The safety filter catches 99% of harmful requests. Why is its review queue mostly harmless? covers base rates in a safety-filter review queue. This page is about estimating the production failure rate from an intentionally unequal audit sample.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →