Model and Inference Engineering · Principal
The voice agent heard 'approve the transfer.' Did it miss the next two words?
The question
Interview question
A caller says, “Approve the transfer... no, wait.” There is a pause before “no, wait.” A voice-activity timeout closes the turn after the first phrase. The agent treats that partial utterance as a final instruction and starts a money-movement workflow. The transcription of the first phrase was accurate. What failed?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Turn segmentation is part of the decision, not just a speech-processing detail. A VAD detects speech and silence, not the end of a person's intent. Google Cloud's voice activity documentation describes activity events and configurable timeouts. The hypothetical system chose to interpret a timeout as authorization. That is an application error, even if the ASR accurately transcribed every audio chunk it received.
I would separate partial transcripts, final segments, end-of-turn inference and authorization to act. Keep the audio and timestamps around the cutoff, including the words after it. Inspect whether a new speech segment arrived while the action was still pending, whether the agent was allowed to interrupt or cancel, and whether the external transfer had actually committed. Measure the distribution of pause lengths for this task, including people who self-correct. A longer timeout reduces early commits but adds response latency. The right answer depends on the cost of waiting versus the cost of a false action.
For an irreversible or expensive action, a VAD timeout should never be sufficient consent. The agent can summarize the intent and ask for an explicit confirmation that names the payee and amount, then bind that confirmation to an action preview with a version or expiry. If the user says “no, wait” while approval is pending, cancel the pending attempt and state whether any external effect already happened. A spoken “yes” still needs an identity and policy check appropriate to the operation. UI and voice flows may use different confirmation friction, but the action boundary must be clear.
Would waiting for the entire call fix it? It avoids this early commit and makes the experience unusable for normal turns. Use streaming for conversation, but treat high-impact tool calls differently. The transcript changes ‘do’ to ‘don't’ after the assistant starts speaking. What was committed? covers partial ASR hypotheses being mistaken for stable text. Here the first words really were spoken and accurately recognized. The missing information is the continuation and correction after a pause, which changes the user's intended command.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →