No. A click is an observation made under the old ranking and interface. The overview page was probably seen more often, and users may click it to orient themselves even when the exception page contains the answer. A document below the fold cannot be judged irrelevant because it was not clicked. Training on those clicks can repeat the old exposure pattern and then make it stronger. Joachims and colleagues' work on learning to rank from biased feedback formalizes why position bias matters. It does not tell us that every click in this product has a particular meaning.

I would replay the affected queries with the old and new candidate lists and inspect where the exception is lost. Was it never retrieved, reranked out, or dropped by the six-slot context builder? Did a general-rule page receive a high label because it was clicked, despite lacking the customer's qualifying condition? Check queries by customer type, policy version, rare exceptions, and whether a user asked for the exception explicitly. NDCG on clicked labels may be rising while the measure we care about, answer-bearing authorized evidence at the context boundary, is falling.

For the high-impact slice, label the actual evidence task. A reviewer sees the question, the governing policy and exception, the source revisions, and a candidate span. They mark whether that span supports the requested decision, including its conditions. The ranking gate can use recall of required evidence within the final context budget and answer support, alongside ordinary search engagement. A clicked overview can still be useful, but it should not displace a necessary exception on a question that depends on it.

If the team wants to learn from clicks, log what was shown, at which position, under which ranker, and what the user did after the answer. Use a small, safe exploration or interleaving design where appropriate to learn about results that were otherwise never exposed. Propensity methods need known or credibly estimated exposure probabilities and overlap. They cannot recover a label for a page that had no chance to be shown in the logged setting. For a sensitive policy slice, explicit adjudication may be cheaper than pretending clicks settle the semantics.

Suppose the exception page is old and the current overview incorporates it. Then the rare page should not automatically be promoted. Verify the current policy and its effective time. The target is sufficient, authoritative evidence, not a quota for exception documents. Conversely, if the exception is still governing but never clicked because it was never visible, a better click metric is telling us almost nothing about that customer's answer.

I would hold the rollout on the affected slice, fix the stage that drops required evidence, and compare old and new systems on the same source snapshot. The follow-up online test should watch useful resolution and wrong-policy outcomes, with a human review path for rare costly failures. A reranker can be more engaging and less correct at the same time.