First show both reviewers the exact world the assistant had: customer scope, effective policy revision, authorized passages, and the wording of the question. A disagreement may be a missing fact in the annotation packet, a vague rubric, or a genuine ambiguity in the user request. More votes can reduce random label noise. They cannot tell us what plan the customer meant when it was never specified. Research on reasons for human disagreement in language inference finds distinct causes including uncertain meaning and task artifacts. Applying that lesson here is a design inference, not a claim that every disagreement in support QA is irreducible.

Ask reviewers to record a short reason and the decisive evidence, not just “answerable” or “not answerable.” If the policy gives one rule for every plan, the missing plan might not matter. If caps differ, the safe response may be “Which plan?” plus what can be said generally. If the source has an authoritative default, record that rule. If two interpretations are both legitimate, the gold artifact should list acceptable actions and the condition for each. Do not force an answer that chooses one unstated plan simply because three out of five reviewers assumed it.

The adjudication process should repair what can be repaired. Clarify the question, fix the rubric, add missing source context, and rerun blind labels.

For unresolved ambiguity, mark the case as such and score whether the assistant identified the missing variable. Keep the raw labels and reasons so a later policy change or reviewer training does not erase the fact that the task was ambiguous. Agreement statistics are useful to locate troubled slices, but a high agreement obtained by excluding all hard cases gives a clean dashboard and a weak eval.

The pushback is “we need one number for the release meeting.” Give a primary metric on adjudicated cases with clear evidence and report the ambiguous slice separately: clarification rate, unsafe forced-answer rate, and reviewer agreement by cause. The manager can still make a decision. They should see that a model which always answers may look accurate under a forced single label yet fail the real user by guessing a plan. Gold data is a model of a task, not a vote that creates facts missing from the task.