At the audio it actually received. Combining channels can hide a quiet speaker under a louder one, create destructive interference if signals have unfortunate phase relationships, or make overlapping words harder to separate. A transcript from the mono mix cannot restore information the preprocessing discarded. The original channel labels may also provide strong speaker attribution without requiring a diarization model to guess identity from one-word segments. Amazon Transcribe's multi-channel guidance describes transcribing channels separately, and Google Cloud Speech-to-Text's guidance returns channel tags when separate recognition is configured. Those are examples of preserving the source structure, not guarantees that either service always recognizes the reply.

I would inspect the original left and right waveforms and listen to each channel separately around the missing reply. Then compare the mono mix, sample rate, channel map, clipping, gain and any voice-activity detector that cut quiet intervals. Did one channel get dropped rather than mixed? Was the right channel delayed relative to the left? Did preprocessing treat a very short utterance as silence? Run ASR on each original channel without the suspect transform and compare word timestamps. The first stage where the reply disappears determines whether the defect is capture, conversion, VAD or recognition.

For a call where each participant normally has a dedicated channel, keep those channels through transcription and attach channel IDs to words and citations. A participant may leak into the other channel, transfer the call, or use a conference bridge, so channel is not a permanent person identity. Check that assignment at each segment. If only a mixed recording exists, diarization and source separation may help, but they are estimates. A high-stakes action cannot treat a guessed “yes” as verified consent just because a transcript contains the word.

To release a fix, build an evaluation set with overlaps, short acknowledgments, quiet callers, channel swaps and real call transfers. Score recall of decision-critical words and correct channel attribution, not only overall word error rate. The transcript says the chair approved it. Did the right speaker say yes? asks whether an overlapping “yes” was attributed to the chair. Here the customer reply was discarded before diarization or language reasoning could even see it. Preserve the separate source signal first.