A video question-answering service finds the frame where a valve begins leaking. It cites 00:40, but the player shows the leak around 00:56. The frame classifier is right. The service calculated frame_index / 30 and assumed a constant 30 frames per second. This recording has variable frame rate and a gap after the camera stopped sending frames.

Frame ordinal is a position in a decoded sequence. Presentation time is when that frame belongs on the media timeline. In variable-frame-rate material, dividing the ordinal by a nominal or average FPS does not recover presentation time. Encoded streams can also have decode order different from display order, so a packet's decode timestamp is not automatically the citation timestamp. FFprobe exposes frame PTS and duration, and FFmpeg's frame output includes pts_time and best_effort_timestamp_time. Those fields need interpretation with the source stream and any transformations applied downstream.

I would keep a stable mapping from sampled model input back to source asset ID, stream, original presentation timestamp and transformation version. Sampling “every fifth frame” is fine for cost control, but a selected sample must retain its source time. If we clip, transcode or normalize timestamps, record both source and derived timeline positions. If the player uses a clip-relative zero while the answer uses source-relative time, even correct PTS can look wrong. Seek to the cited time in the actual user player during validation.

The harder push is whether one frame can support “the valve starts leaking at 00:56.” It may only show that leakage is visible by then. To claim onset, inspect adjacent frames at a temporal resolution fine enough for the question, and report an interval if the onset lies between samples or across a recording gap. A single timestamp with millisecond precision can imply certainty the evidence does not have. For temporal order, sort by presentation time after handling discontinuities rather than assuming decode index is wall-clock order.

I would build a fixture with variable frame durations, a pause, a clip starting at a nonzero PTS and a reordered video stream. Check the cited seek lands on the supporting visual evidence. The spoken words are right. Why does the video citation point ten seconds early? covers audio and video clocks after silence removal. This is the video frame-to-player time mapping even when audio is absent.