Model and Inference Engineering · Staff
The video summary reverses two events. Which timestamps reached the model?
The question
Interview question
A warehouse video summary says a worker closed a gate before a cart passed it. In the original clip, the cart passed first. The vision model identifies both events in still frames. The ingestion job samples “one frame every two seconds” by frame index and stores images in object storage. Audio from a separate track has its own timestamps. What would you inspect before changing the prompt?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
I would ask to see the exact frame sequence that the model received, with a time attached to every image. If the source video has variable frame rate, frame index / nominal fps is not necessarily the time the frame was presented. If the pipeline decoded packets and retained decode order rather than presentation order, reordered frames can make the sequence misleading. If workers completed image uploads out of order and the manifest later sorted by object key, the labels “before” and “after” may be wrong even if extraction was correct. The model cannot reconstruct time that the loader discarded.
Use presentation timestamps, interpreted with the stream time base and any start-time offset, to sample near requested elapsed times. FFprobe documents showing packet PTS and duration fields, which is a useful way to inspect the source rather than trusting a nominal rate. Decode and order the selected frames according to their presentation times. Keep the source clip ID, stream ID, timestamp, extraction rule and frame hash in the manifest passed to the model. If audio is involved, convert its timestamps to the same timeline. A clip cut from a longer source may have its own offset, and audio and video may not begin at exactly the same timestamp.
There is an information limit too. One frame every two seconds can miss the gate moving between samples. If a frame at 10 seconds shows the cart approaching and one at 12 shows it past the gate, that does not by itself establish precisely when the gate shut. Better timestamp handling cannot recover an event that was never sampled. For this task I would add frames around the gate and cart transitions, or use motion and event detection to choose denser windows, then give the model the ordered evidence. The summary should say “the sampled frames do not establish the order” when that remains true.
To debug, align three views on the same timeline: the source player, the extracted frames with actual PTS, and the model's evidence references. Hand-check a small set where the event order is obvious in full video, then vary variable frame rate, edited clips, dropped frames and out-of-order uploads. Compare the original manifest to one with corrected timestamps without changing the model. If the answer flips, the ingestion contract was the likely cause. If it still reverses events while receiving correct, sufficiently dense frames, investigate temporal reasoning and prompting.
I would grade the order claim separately from object recognition. “Gate” and “cart” can both be right while the causal narrative is false. That is a more useful metric than one score for the whole summary, and it points us to either the timeline or the model rather than sending us straight to prompt tuning.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →