We cannot infer that from a raw win rate. Research on self-preference bias in LLM judges reports that familiarity of text can affect judgments relative to human evaluations. Another controlled study finds that apparent self-preference can shrink after quality and evaluator effects are controlled. So I would treat style affinity as a hypothesis to test, not declare every same-family judge invalid. A judge can also fail for simpler reasons, such as missing source documents, length bias, position bias or an ambiguous rubric.

Start with a small set where independent human reviewers can inspect the exact question, evidence and both answers. Blind source identity, randomize answer order, and give a rubric that separates factual support, completeness, tool correctness and helpfulness. Ask the judge for claim-level evidence when support is the criterion. Then create controlled pairs: hold facts fixed while editing length and style, and hold style roughly fixed while introducing a factual error. If the judge flips for style and misses the error, its score is not a reliable release gate for that task. Compare judges from different families, but do not assume a panel vote removes correlated bias.

The interviewer may ask for one number to drive weekly releases. I would first define the costly failure we care about, then give it a separate metric and audit. A judge may be useful for high-volume triage after calibration on a current, human-labeled set. Release decisions can combine the judge's estimate with targeted human review, deterministic checks for tool effects, and minimum performance on critical slices. Keep the judge prompt, model version and evidence bundle fixed during comparisons, and report uncertainty. Recheck after the policy or answer style changes, since a tuned candidate can learn to please the judge without helping users.

What if humans also like the polished answer? That is legitimate product value if it survives the task rubric and does not conceal unsupported claims. We should not strip style automatically. The decision is whether the measured win represents the outcome promised to users. The judge changes its winner when answer order swaps examines answer-order bias. The judge score rises while human reviewers find less support finds that judge scores rise while human support ratings fall. This question isolates a plausible mechanism for that disagreement and gives an experiment that could falsify it.