Take 100,000 production requests at that prevalence. About 100 are harmful and 99,900 are harmless. At 99% recall, the filter flags 99 harmful requests and misses one. At a 1% false-positive rate, it flags about 999 harmless requests. The queue has roughly 1,098 items, of which only 99 are harmful. Precision is about 9%. All three numbers can be correct. The false-positive rate is small relative to harmless traffic, but harmless traffic is enormous relative to the harmful class. Scikit-learn's metric definitions express precision as true positives divided by all predicted positives. This is why recall alone cannot size a review team.

I would verify that the one-in-a-thousand prevalence was measured for the same unit and period. Is it a request, user, conversation, or actually harmful model output? Are adversarial attempts grouped? How were the unflagged cases audited? If only flagged items receive human labels, the estimated production prevalence and recall can both be biased. Sample unflagged traffic with appropriate privacy controls and expert review. Report uncertainty, especially for the one missed harmful item in this toy calculation. A benchmark with curated positive examples may not represent current production attacks.

Then decide what the review queue is for. If reviewers must see every flag before any response, 1,098 items per 100,000 requests may be operationally impossible at scale. If low-risk false positives can receive a safe response automatically, a second stage can reserve human time for uncertain or high-severity cases. Thresholds can differ by action and harm severity, with a clear escalation path. Measure harmful cases missed, benign users blocked or delayed, review capacity, time to decision and downstream outcomes. A higher threshold can improve precision while lowering recall, and that may be unacceptable for the most severe class. We need a cost and risk decision, not a single universal F1 score.

The interviewer may offer a model that catches 95% with a 0.1% false-positive rate. Under the same toy prevalence it catches 95 harmful requests and flags about 100 harmless ones, giving roughly 49% precision. That is a far smaller queue but four more missed harmful requests per 100,000. Whether it is better depends on the harm, available reviewers and what happens to missed cases. We should compare calibrated operating points on fresh traffic and with confidence intervals, not pick the prettier percentage.

Finally, attackers may adapt after launch and change prevalence or the positive distribution. Keep sampled audits and incident feedback rather than freezing the benchmark. The arithmetic explains why the review queue can be mostly harmless today. It does not promise that tomorrow's 99% recall is real.