Model and Inference Engineering · Principal
The video model says nothing happened. What if the event lasted one frame?
The question
Interview question
A warehouse camera records at 30 frames per second. An assistant samples one frame every two seconds and reports that no safety gate was opened. An operator points to a brief gate opening between sampled frames. The vision model correctly described every frame it received. What should the system have said?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
It should not turn “not observed in the sampled frames” into “did not happen in the video.” Sampling discards temporal evidence before the model sees it. A single visible event lasting about 33 milliseconds can fall entirely between samples separated by two seconds. Even if we sampled faster, motion blur, occlusion and camera position may limit what is visible. The FPS-Bench paper studies rapid events that low-frame-rate video inputs miss. This is a coverage problem as much as a model-reasoning problem.
I would first preserve the original clip and its timeline, not just the sampled frames. Inspect sampling timestamps, variable-frame-rate decoding, dropped frames, clock alignment and any preprocessor that resized or skipped segments. For the reported interval, review the source at native or sufficient frame rate. If the system must answer many long videos cheaply, use a staged design: cheap coarse scan, motion or domain-specific triggers, then dense frames around candidate events. But triggers can miss quiet or subtle gate changes, so measure their recall against annotated clips. For a high-consequence negative claim, either cover the full relevant interval at sufficient temporal resolution or state the coverage limitation.
The temporal requirement comes from the event. A person standing in a room for minutes may be captured at sparse intervals. A gate flicker, handoff or brief signal needs different sampling and perhaps a dedicated detector. Count what was actually decoded, not only what the API requested. Timestamped evidence should let a reviewer replay the exact region. If several cameras are involved, synchronize their clocks before combining observations. A second camera may prove the event, but it cannot repair a false claim that the first camera had complete coverage.
Suppose the interviewer offers “just sample all 30 fps.” That can be too costly over days of video, and even then an event between exposure windows or behind an obstruction may be invisible. Set a detectable-event duration and confidence target, then design acquisition, storage and processing for that target. The video summary reverses two events. Which timestamps reached the model? covers a model reversing two events it was shown. Here the missing event never entered its visual context. The safe answer distinguishes absent evidence from evidence of absence.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →