Model and Inference Engineering · Staff
The speaker says fifteen, but the slide shows fifty. What should the assistant report?
The question
Interview question
A meeting assistant summarizes a recorded budget review. The transcript says “fifteen million” at 09:42. A slide visible at that moment reads “50 million.” A person asks for the approved budget, and the assistant produces 50 million with a citation to the slide. No decision record is attached. Did it answer the question?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It found two conflicting claims. It has not found which number was approved. The slide may be a draft, the speaker may have corrected a typo verbally, the ASR may have misheard “fifty,” or the visible slide might belong to a different discussion because audio and video clocks are offset. Choosing the sharper visual text is not a proof of approval, and choosing the spoken phrase is not either.
I would first preserve the evidence with its actual provenance. What audio interval generated “fifteen,” with surrounding words and a confidence or uncertainty signal if the ASR provides one? Which frame or slide revision contains “50,” and was it displayed at the corresponding presentation time? Was there a later correction, a vote, a sign-off, or minutes that say “approved”? FFprobe's timestamp fields are one way to inspect the source streams when synchronization is in doubt. OpenAI's realtime transcription guide distinguishes provisional deltas from completed transcripts. Neither source declares one modality authoritative for a budget decision.
If all I have is this recording, I would answer in the language of the evidence: “The slide shows 50 million, while the speaker appears to say fifteen million at 09:42. I cannot tell from this segment which amount was approved.” Link each claim to the correct source location. If the assistant can ask the user, ask for the approved minutes or owner. If it must produce a working draft, mark the figure as disputed, not a single settled amount. That is still useful. It tells the learner exactly what remains unresolved rather than burying the conflict behind a polished sentence.
For the system design, store audio segments, visual spans and decisions as separate evidence objects with timestamps, source revisions and extraction methods. Join them on a timeline with tolerances, not by assuming frame 100 and transcript row 100 describe the same moment. A model may propose that two statements refer to the same budget, but a downstream verifier should surface incompatible numeric claims and require a resolving source for an “approved” answer. Test 15 versus 50, 1.5 versus 15, negations, revised slides and later verbal corrections. Evaluate whether the final answer represents uncertainty, not only whether one number appears somewhere in the input.
If the interviewer reveals that the signed minutes say 15 million, the answer can become 15 million and explain the discrepancy with the draft slide. If the minutes say 50, reverse it. The source that establishes the decision changes the conclusion. The assistant should not pretend that a clear OCR reading was a decision record.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →