Look at the samples, not just the channel labels. A simple downmix is roughly M(t) = [L(t) + R(t)] / 2. If the caller's waveform on the right is nearly the negative of the left, R(t) ≈ -L(t), adding them cancels the caller. This can happen after a polarity inversion in capture or a previous processing stage. It does not mean every stereo downmix loses speech, and it does not mean the ASR model is defective. FFmpeg's audio filter documentation shows how channel mixes are constructed and how to select channels explicitly.

For one affected clip, keep the original and measure energy and speech intelligibility on left, right and mono over the same timestamps. Inspect channel correlation or compare aligned waveforms around the missing phrase. A near-negative correlation where speech disappears is strong evidence for phase cancellation. It may also be a channel-routing bug: perhaps one channel contains a different speaker, or the configured layout is wrong. Test those hypotheses before applying a global fix.

I would preserve the two channels through ingest, with channel layout and provenance, and choose a downmix only after checking representative recordings from each source. If one channel is consistently clean, transcribe it directly. If both have useful speakers, transcribe per channel and reconcile timestamps and speaker labels. If a polarity inversion is confirmed, correct it at the appropriate boundary, then rerun a small evaluation set. Blindly flipping every right channel can break recordings that were already correct.

The call recording has two channels. Why did transcription lose the customer's reply? asks how channel handling lost the customer's reply when speakers were separated. This page is subtler: both channels visibly contain the same speech, yet their sum removes it. A short waveform comparison usually explains the failure faster than another ASR prompt.