Accuracy and confidence calibration are different claims. If an assistant's score means "probability this answer is correct," then among answered cases assigned about 95 percent confidence, roughly 95 percent should be correct in a large representative sample. An assistant can be 90 percent accurate overall and still be dangerously overconfident on the difficult cases it does answer. Research on calibrated selective classification explicitly distinguishes selective accuracy from the reliability of confidence on accepted predictions.

I would pin down how the score is produced. Is it a trained correctness predictor, a heuristic from model token probabilities, or the model writing "95%" in prose? Those are not interchangeable. Build reliability bins on a held-out set of actual answered requests, with outcome labels and uncertainty intervals. Keep the denominator of all assigned requests as well, because a system that abstains on hard questions can make answered-only accuracy rise while useful coverage falls. Report risk against coverage, calibration among answered cases, and the rate of high-confidence errors on the harm-sensitive slices.

Suppose 80 percent of requests get an answer and 90 percent of those are correct. That means 72 percent of assigned requests receive a correct answer, assuming the remaining 20 percent are abstentions. The calculation says nothing about whether the 95-percent bucket is 95-percent correct. If that bucket is only 75-percent correct on policy exceptions, a UI that highlights its confidence can make the error more harmful. We need labels from those exceptions, not just the easy FAQ traffic that dominates the average.

Could we fix this by lowering every displayed confidence? That may reduce overstatement but does not make the ranking useful for routing or abstention. Calibrate a score on representative data, verify that calibration survives new topics and changed retrieval, and use a separate decision threshold chosen for the cost of wrong answers versus abstentions. Choose an answer-or-abstain threshold without gaming accuracy covers choosing that answer-or-abstain threshold. This page tests whether the number presented as confidence means what users and downstream systems think it means after selection.