The order swap is a diagnostic intervention. If “B is better” survives only when B appears first, the reported 58 percent mixes answer quality with presentation. A systematic study of LLM judge position bias finds that such effects vary across judges and tasks. I would rerun each item in both orders, with candidate identity hidden, and tabulate four outcomes: B wins both, A wins both, flips, and ties or invalid grades. Report per-item pairing, not just two independent aggregate percentages. The 58 and 55 alone do not tell us the overlap or uncertainty, so they cannot establish a winner.

Then look at what the judge was asked to judge. If this is a documentation assistant, give both answers the same question and allowed source spans, and use separate criteria for factual support, completeness, usefulness, and needless verbosity. A long answer can score well on a vague “which is better?” prompt while adding unsupported claims. Blind the model name, position, formatting cues, and cost where possible. Preserve any context needed to decide correctness. If a citation supports a number only in one answer, a human or evidence-based checker should be able to explain the decision at that claim, not simply say the prose sounds stronger.

Would averaging the two orders fix it? It may reduce an order effect in the aggregate, but it does not make the judge valid. If many items flip, set them aside for human review or mark them unresolved. Check how often a judge agrees with domain reviewers on a blinded sample, including ties and hard policy exceptions. Use the same rubrics and source packet for that review. The judge may be consistently biased toward length in both orders. Test matched pairs where one answer adds harmless words and matched pairs where it adds one false claim.

Finally, define a release rule that handles uncertainty. If B really wins on supported, task-relevant answers after the blind paired review, show the magnitude by slice and its cost. If quality is indistinguishable, a faster or cheaper model may win a product decision, but call it a cost decision. The point of a judge is to help inspect many cases. The order-swap result is telling us this judge needs calibration before its percentage becomes evidence of progress.