It tells me the combined workflow often reaches a resolved state, if the resolution label is actually verified. It does not tell me the agent can handle 94 percent of cases alone. The human is part of the system that produced the number. Removing that person, or letting the agent take riskier actions without them, is a different intervention.

I would reconstruct the funnel from assignment, not from the cases that completed in the agent UI. For every eligible case, record whether the agent finished unaided, asked a clarifying question, needed a routine required approval, was rescued after a bad step, handed off cleanly, timed out, or failed without resolution. Record what the human changed and what evidence establishes the final customer outcome. “Human touched the case” is too crude. A policy required approval of a correct plan is not the same as a person discovering an invented refund or repairing a tool action the agent got wrong.

Then publish three numbers on the same assigned population. How many cases reached a verified outcome with no substantive rescue? How many reached one with human assistance and what work did that require? How many never reached the outcome, including silent abandonments and unresolved handoffs? Show these by case type and risk. If the agent succeeds on 90 percent of easy cases but humans take almost every high value failure, an overall 94 percent can be useful as a service metric while being dangerously misleading as an autonomy claim. Do not improve the rate by declaring the hard cases “ineligible” after seeing that they needed rescue. Set eligibility rules before assignment and report excluded cases separately.

There is also a question about value, not just attribution. The agent may save a human twenty minutes by collecting evidence and then handing over a clean case. That is a good product result even if the human makes the final decision. Measure human minutes, queue delay, correction burden, quality and harm as well as completed cases. The human can also spend more time undoing a plausible but wrong plan than they would have spent starting fresh. OpenAI's agent evaluation guide describes traces of model calls, tools and handoffs. Those traces help locate the intervention, but final resolution and labor still need their own evidence.

If product asks whether deploying the agent improves the service, compare the agent plus human workflow with the relevant human workflow under a credible assignment design, and account for shared human queue capacity. If product asks how often the agent can operate autonomously, use the assigned denominator and a predeclared definition of unaided completion. These are different estimands. A randomized rollout can estimate the effect of the workflow, but cannot magically tell us what would have happened without rescuing the dangerous cases. We should not withhold required human help merely to make an autonomy metric clean. Use bounded, lower risk cases for any autonomous expansion and keep the review path available.

Suppose the interviewer says every rescued case ends happily, so there is no customer problem. Good, then the assisted workflow may work. I would still reject the proposal to expand autonomous permissions on the 94 percent figure. Show the unaided rate and its high value slice, inspect severe near misses, define what requires a person, and test a guarded change in that boundary. The system being proposed after that change is not the one that produced the 94 percent.