Evaluation and Quality · Principal
The grader says the agent improved, but customers still call back. Which result wins?
The question
Interview question
An offline grader says a new support agent resolves 87% of cases, up from 79%. In a randomized rollout, repeat contacts within seven days rise. The agent also closes tickets faster. Product says the repeat contact metric is noisy because customers can call about something else. The evaluation team says their grader was calibrated against human labels. Decide what evidence should control the release. Then you learn only successful conversations were sent to the grader.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Neither number wins by being called the “real metric.” They answer different questions, and the last detail means the 87% is not even measuring the rollout population. First define the customer outcome. For a disputed charge, it could be a correct adjustment or a truthful explanation plus a completed next step. A polite close is not that outcome. A repeat contact may be a different problem, but a customer who gave up may never contact again. We need a case-level adjudication method and an outcome window appropriate to the task.
I would reconstruct the experiment from assignment, not from finished transcripts. Count every eligible case assigned to old or new agent, including timeouts, handoffs, refusals, abandoned chats, and cases without a final message. Check stable randomization by customer or issue so a follow-up does not switch arms, and examine assignment and exposure counts for imbalance. Microsoft's experiment guidance on sample ratio mismatch explains why unexpected arm proportions are a data quality warning. The exact unit here must be chosen for the support workflow, not copied from a game experiment.
Now audit a stratified sample of both arms using the customer goal, account state, tool receipts, follow-up records, and the actual message shown. Reviewers should be blinded to model version where feasible. Ask whether the underlying issue was resolved, not whether the answer sounds helpful. Include cases that the grader never received. If missingness is related to failures, scoring only successes inflates the reported result even when the grader agrees perfectly with humans on the transcripts it saw. OpenAI's grader guidance supports comparing grader outputs to expert judgments. That calibration does not repair a biased input sample or an outcome the transcript cannot observe.
For repeat contact, link by issue as well as account and time. Separate a second message about the same disputed charge from a new shipping problem. Check whether the new agent encourages another contact as part of a legitimate pending process, and whether the old agent prematurely closes without telling the customer to return. The seven-day window may miss slow provider outcomes. Show a sensitivity analysis at different windows, with a fixed plan before inspecting the favorable one. Randomization helps with causal comparison if assignment and measurement held up, but it does not make a flawed proxy equal to resolution.
Suppose the audit finds the agent gives correct explanations but misses a provider confirmation, so customers call back to ask whether their credit posted. Then the grader may be measuring answer quality correctly and the product experience still got worse. Fix the workflow and status message, keep cases pending until confirmed, and rerun the test. If the rise is driven by a changed contact routing rule applied only in one arm, fix the experiment and do not claim a treatment effect yet.
I would pause expansion while the discrepancy is unresolved. Severe wrong financial actions would stop the affected path immediately. For a low-risk repeat-contact increase, keep a bounded canary long enough to investigate, with a named release owner and customer effort guardrail. Report the audited resolution estimate with uncertainty, the complete assigned population, repeat contact by issue and cohort, and the grader's coverage and errors. The support agent closes more tickets. Did it solve more problems? asks whether closing a ticket solved the problem. Here the tougher question is what to believe when the evaluation pipeline and randomized product outcome disagree.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or return to the full Interview Prep index.
Browse this area →