Model and Inference Engineering · Staff
The voice agent interrupted itself. Who did its microphone actually hear?
The question
Interview question
A voice agent speaks a confirmation through the user's device speaker. The microphone picks up part of that synthesized speech. Voice-activity detection fires and ASR transcribes “Yes, confirm,” which the agent treats as the user's barge-in. The actual user was silent. What evidence would separate a person speaking from the system hearing its own output?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Full-duplex audio has two streams: playback and captured microphone. Acoustic echo can carry playback into capture with delay and distortion. WebRTC's media-device guidance describes an echo-cancellation constraint, but requesting a browser feature is not a guarantee it worked on every device or acoustic path. A VAD only detects speech-like activity. ASR can transcribe the echo perfectly and still attribute it to the wrong actor. If the agent treats any transcript as user intent, it can act on its own words.
I would align the playback reference waveform, microphone capture, VAD events and ASR spans on one clock. Did the suspicious speech occur shortly after matching TTS audio? Compare the decoded words and audio correlation, then test with headphones, speakerphone, quiet rooms, noisy rooms and different device volume. Check whether echo cancellation was enabled and actually active, whether the app routed the TTS reference into the canceller, and whether audio mixing or resampling introduced delay. Preserve a sample of raw capture for diagnosis under privacy controls, since a transcript alone cannot prove speaker identity.
The product needs a barge-in policy. Detect real user speech during playback, stop TTS promptly when confidence is sufficient, and suppress playback echo rather than simply muting the mic for the entire answer. Muting would miss genuine interruptions. For consequential commands, do not accept a phrase identical to the system's own TTS as consent without another confirmation boundary. Distinguish user audio, assistant playback and uncertain overlap in the event stream. Measure false interruptions and missed interruptions, not only ASR word error rate.
An interviewer might suggest ignoring all recognized speech for 500 milliseconds after each TTS chunk. That can suppress the echo and also miss a real “stop.” A reference-aware canceller and a policy for uncertain overlap are safer than one fixed delay across devices. The voice agent heard 'approve the transfer.' Did it miss the next two words? concerns a user's correction after a pause being missed. Here the user said nothing, and the system invented a user turn from its own output.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →