The first question is what we mean by committed. The ASR system can emit provisional text to reduce latency. The conversation state can include a finalized user turn. The assistant can generate audio, the client can play some of it, and an external action can be submitted. Those are four different events. Treating the first partial transcript as an irrevocable user instruction lets a single late negation change the meaning after the assistant has already committed to a response.

For this caller, I would let provisional text drive low-risk preparation, such as finding the candidate summary, but keep the response and send decision conditional until the relevant utterance is stable enough. A send action needs final or explicitly confirmed intent, a named recipient and content. The scenario says no send happened, so stop further playback, mark the already heard promise as incorrect, and say something plain: “I heard the correction. I have not sent it.” Do not write a tidy transcript that makes it look as though the assistant never spoke. The user heard half a sentence.

An implementation should key every partial and completion to a turn identity. OpenAI's realtime transcription documentation distinguishes incremental deltas from a completed transcript and warns that completion events for different turns may arrive out of order. A product should not assume every ASR engine revises partials in the same form, but it should model partial hypotheses as provisional and associate all updates with the right utterance. Keep an explicit state for generated, queued and played audio. A text transcript of generated audio cannot tell you exactly what reached the speaker if playback was buffered.

For the information question, we can speculate sooner. If the caller says “what time is the meeting,” begin looking it up and perhaps begin a short response once the intent is clear. If later audio changes the meeting name or adds “tomorrow,” stop or revise before stating an unverified time. The threshold depends on the cost of being wrong and whether a response is reversible. It is reasonable to trade a few hundred milliseconds for a settled negation in an action request. It is wasteful to make every low-risk spoken question wait for an unusually long silence if the system can interrupt itself cleanly.

I would test recorded audio with corrections, accents, background speech and late negations. Measure time to first useful speech, words heard before correction, incorrect external actions, and whether the repair statement matches the action ledger. A low word error rate alone hides the costly case where one changed word reverses intent. The voice experience can be quick without pretending that early transcript bytes have the authority of a confirmed instruction.