Model and Inference Engineering · Staff
The transcript changes ‘do’ to ‘don't’ after the assistant starts speaking. What was committed?
The question
Interview question
A voice assistant streams speech recognition and starts answering while the caller is still speaking. An early transcript says “do send the summary.” The final turn says “don't send the summary.” The assistant has already played half a sentence promising to send it. No external send has happened yet. Design the turn boundary and the repair behavior. Would you handle “what time is the meeting” the same way?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The first question is what we mean by committed. The ASR system can emit provisional text to reduce latency. The conversation state can include a finalized user turn. The assistant can generate audio, the client can play some of it, and an external action can be submitted. Those are four different events. Treating the first partial transcript as an irrevocable user instruction lets a single late negation change the meaning after the assistant has already committed to a response.
For this caller, I would let provisional text drive low-risk preparation, such as finding the candidate summary, but keep the response and send decision conditional until the relevant utterance is stable enough. A send action needs final or explicitly confirmed intent, a named recipient and content. The scenario says no send happened, so stop further playback, mark the already heard promise as incorrect, and say something plain: “I heard the correction. I have not sent it.” Do not write a tidy transcript that makes it look as though the assistant never spoke. The user heard half a sentence.
An implementation should key every partial and completion to a turn identity. OpenAI's realtime transcription documentation distinguishes incremental deltas from a completed transcript and warns that completion events for different turns may arrive out of order. A product should not assume every ASR engine revises partials in the same form, but it should model partial hypotheses as provisional and associate all updates with the right utterance. Keep an explicit state for generated, queued and played audio. A text transcript of generated audio cannot tell you exactly what reached the speaker if playback was buffered.
For the information question, we can speculate sooner. If the caller says “what time is the meeting,” begin looking it up and perhaps begin a short response once the intent is clear. If later audio changes the meeting name or adds “tomorrow,” stop or revise before stating an unverified time. The threshold depends on the cost of being wrong and whether a response is reversible. It is reasonable to trade a few hundred milliseconds for a settled negation in an action request. It is wasteful to make every low-risk spoken question wait for an unusually long silence if the system can interrupt itself cleanly.
I would test recorded audio with corrections, accents, background speech and late negations. Measure time to first useful speech, words heard before correction, incorrect external actions, and whether the repair statement matches the action ledger. A low word error rate alone hides the costly case where one changed word reverses intent. The voice experience can be quick without pretending that early transcript bytes have the authority of a confirmed instruction.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →