I would trust neither number without its denominator and sampling path. The vote rate is positive votes among people who decided to vote. It is not positive experiences among all exposed users. Maybe the new UI makes happy users vote more. Maybe users with a bad answer abandon before they can rate it. Maybe the release changes which tasks reach the voting screen. The random audit can also be wrong if its rubric ignores what users needed, but its random selection is a much better starting point for estimating the rate of a specified failure over all eligible requests.

Start with a request ledger: eligible exposures, completed responses, failures before an answer, vote prompt shown, vote actually submitted, task and tenant slice, version assignment, and later user action. Compare the distribution of the 4 percent who voted with the full exposed population. If response quality was measured only after a response completed, include errors and abandonments in the product denominator. Google researchers have studied response-rate bias in satisfaction data. That work is not a calibration for this assistant, but it supports treating voluntary feedback as a selected observation rather than population truth.

Keep thumbs as a useful diagnostic. They can surface complaints and examples to inspect. For a release decision, sample a known fraction of traffic for blind review and label concrete criteria: factual support, task completion, unsafe action, and user usefulness. If an invited survey has nonresponse, track invitation and response separately. If we know or can estimate inclusion probabilities, weighted analysis may help, but no weight repairs an unobserved group whose response mechanism and quality are both unknown. I would also inspect a cohort of sessions with no votes, especially users who left at a tool error or did not see an answer.

What if product argues that auditors dislike a concise answer users love? Then compare the two on the same sampled sessions, with a user outcome measure and the evidence available to the assistant. A response may be pleasant yet unsupported. That is a quality trade-off to discuss openly, not a reason to relabel a factual error as success. A release can improve satisfaction for answered easy cases while raising unsafe answers on hard cases. Report both, with counts and uncertainty by slice. The 80 percent thumbs rate is one signal from people who spoke, not a referendum from everyone the system served.