The audio preprocessor changed the time axis. If ten seconds of silence before the statement were removed, timestamp 03:00 in the compressed stream maps to 03:10 in the original. The two clocks are no longer interchangeable. Word accuracy can remain high while temporal alignment is wrong. OpenAI Whisper's transcription implementation explicitly adds a time_offset when mapping timestamp positions within a processed segment. A pipeline that concatenates speech ranges needs its own mapping from compressed offsets back to original media time. Whisper does not automatically know about a separate silence-removal step invented upstream.

I would inspect one bad citation end to end: original media timestamp, audio extraction offset, detected speech ranges, concatenated audio offset, ASR segment and word times, and video frame presentation timestamp. Build a piecewise map for each retained range, including its start and duration in original media and its start in compressed audio. Apply that map before merging with video or creating citations. Store source asset ID and timestamp units. Mixing milliseconds and seconds or wall-clock time with media presentation time can cause a similar symptom, so verify those separately. For variable-frame-rate video, seek using presentation timestamps, not frame_number / nominal_fps as though frames were uniform.

The ten-second example is easy if silence was cut only once. With several cuts, the offset accumulates and changes by interval. A single global +10 correction will fail later. If an utterance straddles a cut, segmentation and word alignment need closer inspection. A robust pipeline can keep silence in the ASR input when cost permits, or maintain a reversible edit decision list and validate a sample of mapped word boundaries against the original audio. Do not use the model's textual summary to infer timing after the fact.

If the interviewer says the user only wants meeting notes, citation timing may seem minor. It becomes central when a reviewer clicks the cited frame, when audio is reconciled with a slide, or when an action depends on who said something at a particular moment. The video summary reverses two events. Which timestamps reached the model? covers reversing the order of events already in a video summary. Here the ordering may be fine, but the audio and video are indexed on different clocks. The source timestamp should remain the authority throughout.