Full-duplex audio has two streams: playback and captured microphone. Acoustic echo can carry playback into capture with delay and distortion. WebRTC's media-device guidance describes an echo-cancellation constraint, but requesting a browser feature is not a guarantee it worked on every device or acoustic path. A VAD only detects speech-like activity. ASR can transcribe the echo perfectly and still attribute it to the wrong actor. If the agent treats any transcript as user intent, it can act on its own words.

I would align the playback reference waveform, microphone capture, VAD events and ASR spans on one clock. Did the suspicious speech occur shortly after matching TTS audio? Compare the decoded words and audio correlation, then test with headphones, speakerphone, quiet rooms, noisy rooms and different device volume. Check whether echo cancellation was enabled and actually active, whether the app routed the TTS reference into the canceller, and whether audio mixing or resampling introduced delay. Preserve a sample of raw capture for diagnosis under privacy controls, since a transcript alone cannot prove speaker identity.

The product needs a barge-in policy. Detect real user speech during playback, stop TTS promptly when confidence is sufficient, and suppress playback echo rather than simply muting the mic for the entire answer. Muting would miss genuine interruptions. For consequential commands, do not accept a phrase identical to the system's own TTS as consent without another confirmation boundary. Distinguish user audio, assistant playback and uncertain overlap in the event stream. Measure false interruptions and missed interruptions, not only ASR word error rate.

An interviewer might suggest ignoring all recognized speech for 500 milliseconds after each TTS chunk. That can suppress the echo and also miss a real “stop.” A reference-aware canceller and a policy for uncertain overlap are safer than one fixed delay across devices. The voice agent heard 'approve the transfer.' Did it miss the next two words? concerns a user's correction after a pause being missed. Here the user said nothing, and the system invented a user turn from its own output.