Accuracy among answered requests is conditional on the model deciding to answer. It can reach 100 percent by answering only one easy question. Put coverage beside risk: coverage is the fraction of eligible questions answered, and selective risk is the error or unsupported-claim rate among those answered. Measure false abstention on questions with sufficient authorized evidence too. The selective prediction literature treats abstention as a decision with its own accuracy and coverage trade-off. For this product I would add “ask for a version” as a third action, because a useful clarification is different from a flat refusal.

Build a held-out set with the question, user and version context, available authorized sources, answerability label, and a defensible answer or acceptable clarification. Separate cases with no source, conflicting sources, missing version, and genuinely answerable but hard reasoning. Score whether the policy chose the right action and, when it answered, whether every important claim was supported. A confidence number emitted by the same generator is not automatically a calibrated probability. Validate the actual score against outcomes, then tune thresholds on a calibration set and report performance on a separate evaluation set. Reusing the same set for many threshold changes invites overfitting.

I would choose a point on the risk-coverage curve from the product's cost of mistakes and missed help, not from a round number like 0.8. For a billing rule, an unsupported confident answer may cost much more than asking for a version. For a low-stakes navigation question, unnecessary abstention may cost more. Even if the service has one global policy, show curves by tenant, document age, source authority, language, and query type. The customer with old manuals might have a high aggregate abstention rate because the system lacks the requested versions, not because its users ask worse questions. Fix the coverage of the knowledge source if that is the cause. A threshold cannot create missing evidence.

If the team says “overall unsafe-answer rate fell,” check the denominator. Unsafe answers divided by all questions can fall simply because the system answers fewer. Report both unconditional harm per eligible request and harm among answered requests, plus useful resolutions after clarification. If a customer's slice is small, show uncertainty rather than declaring its threshold safe from three examples. Monitor shifts after release, because a newly updated manual changes what is answerable. The purpose of abstention is to make the right decision with the evidence available, not to hide difficult cases from the score.