Evaluation and Quality · Principal
The support agent closes more tickets. Did it solve more problems?
The question
Interview question
After a support agent rollout, automatic closure rises from 35 to 55 percent and human escalations fall. A random review finds some closed cases where the customer's underlying problem remains. The business team calls this a productivity win. Decide what to measure, how to investigate the failure, and whether to continue the rollout. Then the interviewer says many customers never reopen a bad ticket, and human review cannot inspect every case.
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Closure is a system event. Resolution is a claim about the customer's problem. They coincide only if the closure rule is tied to evidence of the requested outcome. An agent can improve a closure metric by writing a convincing final message, changing the ticket state, or avoiding an escalation. None of those actions proves that an API key works, an invoice was corrected, or an incident stopped recurring.
I would start with the case ledger, not the agent's closing sentence. Record the original customer goal, relevant account state before and after, the exact actions taken, provider receipts, the final customer message, closure reason, and any later contact about the same issue. Define different success rules by task. For a password reset, a verified reset event may be enough to claim the action completed, but only the customer can tell us whether access was restored. For a billing adjustment, a posted ledger entry and customer-visible amount matter. An information request can be resolved by a supported answer. A case waiting on a third party should remain pending, even if the agent has done all it can today.
Then compare the rollout and control on the population assigned to them, including cases the agent failed to answer, routed away, or timed out. Report closure, independently verified resolution, escalation, time to real resolution, repeat contact over an appropriate window, unsafe action, and customer effort. Keep task mix and risk slices visible. A falling escalation rate may mean the agent handles routine cases well, or that it suppresses a necessary human handoff. The investigation samples both closed and unclosed cases, with extra sampling of high-risk actions and low-confidence closure paths. It also traces the exact point where “waiting for provider” became “resolved.” OpenAI's agent evaluation guide describes traces across calls, tools and handoffs. That is useful for locating a workflow mistake, though a trace alone cannot observe whether the customer's problem stayed fixed. Amazon Connect's metric definitions separately track closures, reopens and resolution outcomes in a contact center, illustrating why the labels are not interchangeable.
Suppose the random audit finds 8 unresolved cases among 100 closed ones in the new arm, versus 2 among 100 in control. These are illustrative small samples, so I would report uncertainty and investigate severity before pretending the difference is a precise population estimate. If the failures include money, access or service continuity, stop autonomous closure on those paths now and route cases for review while the broader experiment continues only where safe. Preserve the ability to reopen and notify affected customers. A polished final answer from the agent does not override a missing provider receipt.
But customers rarely reopen a bad ticket. Reopen rate is an incomplete proxy and can even fall when the customer gives up or starts a new case. Join contacts by issue and account within a defensible window, sample customers for follow-up, look at product telemetry where it directly tests the goal, and audit a stratified set of silent closures. Be careful with privacy and with selection: people who answer a survey are not all customers. For cases whose outcome cannot be observed, label them unknown, not success. Human review cannot cover everything, so use random audits to estimate a failure rate with uncertainty, targeted reviews to find severe patterns, and deterministic checks for things the runtime can verify, such as a posted adjustment before billing closure.
If the team argues that the agent deserves credit for closing a ticket the customer never replies to, I would separate “agent completed its permitted work” from “customer outcome verified.” Both can be useful operational metrics. The release gate depends on the promised product. If we market resolved issues, we need evidence for resolved issues. If the product only drafts next steps for a human, assess the draft and handoff instead. This distinction changes the metric, the UI wording, and which actions the agent may take without review.
Continue reading
Related questions
Read beyond the question
Explore more evaluation and quality
Follow another question in this area, or return to the full Interview Prep index.
Browse this area →