Two people can occupy the same time interval. A pipeline that assigns exactly one speaker and one text span to every slice of audio has no place to represent that overlap. A standard ASR pass on the mixed signal may favor the louder speaker or produce a blend that neither person said. Diarization that labels who spoke when does not itself reconstruct words the recognizer missed. Overlap-aware diarization research allows two speakers on overlapping frames, and work combining diarization and speech separation studies recognition on separated sources. Those approaches help, but they do not promise perfect recovery on a particular noisy call.

I would play the original interval and inspect the channel layout. If speakers occupy separate channels, process those channels before trying to separate a mono mixture. Otherwise mark the overlap interval, run a suitable separation or overlap-capable recognition path, and compare against a human transcript for the disputed words. Keep both candidate speaker tracks with their own timestamps and uncertainty. Do not force one into a sequential conversation by moving its words later. The fact that someone objected during the approval is material.

The answer system should not infer consensus from a transcript known to have untranscribed overlap at the decision point. It can quote the clear approval and flag the overlapping objection as unresolved until reviewed. Evaluate the full pipeline on overlap-specific word error and speaker attribution, not only average word error on single-speaker audio. One missed two-word objection can matter more than many harmless filler errors.

Fixing the speaker label on the surviving sentence does not recover the missing objection. Listen to the disputed interval before treating a clean-looking transcript as a complete record of the decision.