I would treat the disagreement as a release blocker for the affected workflow, not as proof that humans or the judge are wrong. The prompt under test has changed repeatedly in response to this very score. That is enough to overfit a fixed evaluation, even when the grader and cases never change. The assistant may learn to write a style that sounds well supported to the judge, cite more often, or put uncertainty in a sentence the judge rewards while still making an unsupported claim.

First inspect a small stratified set of changed outcomes. Pair each generated claim with its cited span, source revision, and the user's actual permission. Did the judge see the source text, or only the answer and citation marker? If it saw a whole document, did it verify the sentence cited rather than a nearby paragraph? Did the prompt change produce more correct answers, more abstentions, or merely more judge friendly phrasing? Show reviewers the old and new answer blind to version and score, then adjudicate disagreements. This is a diagnosis of the measurement, not a contest of who sounds more confident.

Keep separate measures for factual correctness, visible citation support, answerability, and usefulness. A true fact cited to the wrong paragraph fails support. A safe abstention on an answerable question has a different cost from a confident unsupported number. An LLM judge can help triage these cases, but I would calibrate it against expert labeled claims and report errors by question type, source version, and difficulty. OpenAI's evaluation guidance recommends calibrating automated scores with human feedback. Its reinforcement fine-tuning guidance explicitly notes that a model can learn to reward hack a grader. Here we may be overfitting an application prompt rather than training weights, but the measurement failure is similar.

“But the judge and cases never changed.” That actually makes the problem sharper. A static metric becomes less independent with every selection made against it. I would freeze an untouched holdout sampled from real, recent tasks, keep access restricted, and use it only for promotion decisions. Rotate in newly adjudicated production failures and retire cases whose sources have changed. A second judge with a different prompt or model can be a useful check, but agreement between two judges is not proof when both miss the same citation problem. For high impact slices, blinded human review remains part of the release evidence.

Then compare the prompt versions on the same corpus and a genuinely fresh holdout. Record source and reranker versions so retrieval drift does not masquerade as a prompt effect. The deployment gate is a noninferiority floor on unsupported claims and critical violations, plus an improvement in useful outcomes. If the new prompt improves fluency but degrades support, I would roll it back on the affected slice even if the composite judge number is higher.

If the judge sees citations but human reviewers do not, the human study is poorly specified too. Give both the same evidence and ask reviewers to score explicit claims, or measure end user trust as a separate outcome. I need to know what evidence each evaluator saw and what event its score represents.