Evaluation and Quality · Staff
The safety filter catches 99% of harmful requests. Why is its review queue mostly harmless?
The question
Interview question
A classifier flags requests for human review. On a labeled benchmark it catches 99% of harmful requests and falsely flags 1% of harmless ones. In production only one in a thousand requests is genuinely harmful. The review queue is full of harmless cases. The team thinks one of the metrics must be wrong. Does it?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Take 100,000 production requests at that prevalence. About 100 are harmful and 99,900 are harmless. At 99% recall, the filter flags 99 harmful requests and misses one. At a 1% false-positive rate, it flags about 999 harmless requests. The queue has roughly 1,098 items, of which only 99 are harmful. Precision is about 9%. All three numbers can be correct. The false-positive rate is small relative to harmless traffic, but harmless traffic is enormous relative to the harmful class. Scikit-learn's metric definitions express precision as true positives divided by all predicted positives. This is why recall alone cannot size a review team.
I would verify that the one-in-a-thousand prevalence was measured for the same unit and period. Is it a request, user, conversation, or actually harmful model output? Are adversarial attempts grouped? How were the unflagged cases audited? If only flagged items receive human labels, the estimated production prevalence and recall can both be biased. Sample unflagged traffic with appropriate privacy controls and expert review. Report uncertainty, especially for the one missed harmful item in this toy calculation. A benchmark with curated positive examples may not represent current production attacks.
Then decide what the review queue is for. If reviewers must see every flag before any response, 1,098 items per 100,000 requests may be operationally impossible at scale. If low-risk false positives can receive a safe response automatically, a second stage can reserve human time for uncertain or high-severity cases. Thresholds can differ by action and harm severity, with a clear escalation path. Measure harmful cases missed, benign users blocked or delayed, review capacity, time to decision and downstream outcomes. A higher threshold can improve precision while lowering recall, and that may be unacceptable for the most severe class. We need a cost and risk decision, not a single universal F1 score.
The interviewer may offer a model that catches 95% with a 0.1% false-positive rate. Under the same toy prevalence it catches 95 harmful requests and flags about 100 harmless ones, giving roughly 49% precision. That is a far smaller queue but four more missed harmful requests per 100,000. Whether it is better depends on the harm, available reviewers and what happens to missed cases. We should compare calibrated operating points on fresh traffic and with confidence intervals, not pick the prettier percentage.
Finally, attackers may adapt after launch and change prevalence or the positive distribution. Keep sampled audits and incident feedback rather than freezing the benchmark. The arithmetic explains why the review queue can be mostly harmless today. It does not promise that tomorrow's 99% recall is real.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →