The recording test and the live path may feed different audio to the recognizer. A voice activity detector decides when speech starts. If it fires 150 ms after a quiet consonant or a soft first syllable, and the pipeline sends audio only from that detection point, the speech model never receives the missing sound. It can give an internally plausible transcript of the shortened clip. Training a better language model will not recover a syllable that was never in its input.

I would preserve a short rolling audio buffer before the detected start and send some of that prefix with the turn. OpenAI's Realtime VAD guide describes prefix_padding_ms as audio included before detected speech. This establishes the mechanism, not a universal number of milliseconds that works for every microphone and noise environment. Too little prefix clips onsets. Excessive prefix can bring in a prior speaker or noise, which matters when an agent acts on a named customer.

The useful debug artifact is a time-aligned comparison of the raw microphone capture, the segment passed to ASR, the VAD start event and the final transcript. Listen to the first 300 ms around the boundary. Does the syllable exist in the raw capture but disappear in the segment? If yes, change segmentation and measure. If it is already absent in the raw capture, inspect microphone gating, echo cancellation and network packet loss. If it reaches ASR intact, investigate recognition or vocabulary instead. Do not use transcript accuracy alone to choose among these causes.

I would evaluate onset recall on quiet speech, names, acronyms and code-switching, along with end-to-end latency and false-start rate. For a tool action, repeat or confirm an uncertain account name before submission. The call lasted ten seconds. Why did the speech pipeline hear only three? concerns a sample-rate mismatch that changes how much audio is heard. Here duration can be correct while the start of each utterance is cut off.