Model and Inference Engineering · Principal
The transcript says the chair approved it. Did the right speaker say yes?
The question
Interview question
A meeting assistant produces action items and says the chair approved a budget increase. The decisive line is a short “yes” during overlapping speech. The words are transcribed correctly, but the diarization system assigns the line to the chair's speaker label. Someone wants to ship approval extraction because the word error rate is low. What is missing?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
The content of a word and the identity of its speaker are separate predictions. Low word error rate can coexist with wrong attribution, especially for a one-word interruption where there is little voice evidence and two people speak at once. Diarization often produces anonymous local labels such as SPEAKER_00. Linking one of those labels to a named chair is another step with its own evidence. The pyannote diarization toolkit explicitly deals with speaker changes and overlapping speech. NIST's meeting recognition evaluations score diarization and speaker-attributed transcription as distinct tasks. Neither source makes this meeting's “yes” a valid approval.
I would play the audio around that line, not just inspect the diarized text. Check the overlapping interval, microphones or channels, speaker embeddings on the short segment, and whether the prior and next turns support the same speaker. Inspect how SPEAKER_00 was mapped to the chair. Did someone identify the chair by a clean introduction, by seat location, or simply by being the first voice? If there is video, lip movement may help, but camera cuts and timing must be checked. A confidence score from a model is a hint, not a substitute for a source that establishes the decision.
Then ask what “approved” means in this organization. A chair saying yes to “can we discuss it next week?” is not a budget approval. The system needs the proposal, amount, decision procedure and any later correction. A signed minute, recorded vote or explicit confirmation may be authoritative. If that evidence is absent, the answer should be “the transcript contains a yes during overlapping speech, but I cannot verify who said it or whether it approved the increase.” It can queue the item for the meeting owner rather than writing it as a committed decision.
At the pipeline level, keep word timestamps, overlap indicators, anonymous diarization labels, identity mapping evidence and source audio references separate. Do not flatten them into “Chair: yes” and then ask a language model to reason from that fabricated certainty. Test short acknowledgments, interruptions, similar voices, remote participants, crosstalk and role changes across meetings. Evaluate speaker-attributed decision accuracy, not only WER and overall diarization error. A model can do well over hours of clear speech and still fail on the one second that matters.
Suppose the interviewer says the chair confirmed later in an email. Then the action item can be recorded as approved with the email as its authority, and the meeting transcript remains an ambiguous supporting source. That is an evidence upgrade. It is not a reason to rewrite the audio record as though the original attribution was correct.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →