Evaluation and Quality · Principal
Human-reviewed accuracy rose. Did the assistant improve, or did review routing change?
The question
Interview question
The assistant's quality dashboard uses cases sent to human review. A new router sends high-risk cases to a separate specialist queue that is not included in this dashboard. Reviewed accuracy jumps from 86% to 94%. Product traffic and model weights did not change. What does the new number measure?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It measures accuracy on the cases that still entered that review sample. It does not establish improvement across all customer requests. The routing rule selected which labels are observed, and that selection depends on risk, which may correlate with the model's error rate. Research on selective labels explains the broader problem of evaluating predictions when observed outcomes depend on decisions made by the system. Its application here is an inference from the review design, not a claim that the paper studied this particular assistant.
I would compare the total request population before and after, the probability that each type reached each review queue, and label coverage by risk, tenant, language, task and model version. Bring the specialist labels back into an evaluation view while keeping their distinct severity and rubric. Establish a small stratified random audit sample independent of the router so every important cohort has a chance of review. Weighting reviewed cases back to the population can help only if inclusion probabilities are known and nonzero in the slices of interest. It cannot recover quality for a group the system never labels without new audit data.
Keep separate metrics for routine requests and high-risk tasks. If the specialist queue sees harder cases, its raw accuracy may be lower without implying its reviewers or the model are worse. Report both the conditional rates and an overall estimate with its sampling design and uncertainty. Check whether the router itself changed user outcomes through human intervention. The case can be safe after a specialist fixes it while the assistant's original answer was wrong. Define what event the quality metric credits.
The interviewer may ask if this is just a reporting bug. The dashboard query is part of it, but the deeper issue is that the evaluation population shifted when the product policy changed. The model looks accurate on labeled cases. What happened to cases still waiting for an outcome? covers missing labels because outcomes have not arrived yet. The agent succeeds after humans rescue its hardest cases. Whose success are we measuring? covers human rescue counted as agent success. Here labels exist in another queue, but the decision to send hard cases there made a selected sample look better.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →