Model and Inference Engineering · Staff
The video found the right badge. Why did it say the wrong person wore it?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
Finding an object in a frame is not the same as tracking a person across frames. Two people walk past each other, one is briefly occluded, and a tracker swaps their IDs when they reappear. A later frame clearly shows a badge on one person. If the assistant joins that badge observation to the earlier identity using the switched track ID, every individual observation can be correct while the sentence about who wore it is false. ByteTrack is one primary example of how tracking associates detections across frames. Google's video person-detection documentation distinguishes per-person tracks and time-stamped boxes.
I would ask for the exact claim and the frame evidence supporting each part. Identify the frame that establishes the person's identity, the frame where the badge is visible, and the association linking the two. Inspect boxes and track IDs around the crossing and occlusion, not only the polished summary. A track ID is a model-produced association, not a legal identity or a guarantee that the same human appears on both sides of an occlusion.
For a consequential answer, require a continuous enough visible path or corroborating cues, such as clothing and scene position, and keep confidence about the association separate from confidence about detecting the badge. When association is uncertain, answer that a badge is visible at a timestamp but the wearer cannot be identified reliably from this clip. More sparse frame samples can make the association worse even if each sampled frame is sharp. Review the dense window around the crossing before escalating to a larger video model.
An interviewer might say the system has perfect object detection. Good, then object detection is not the failure. Replay with both people crossing and check identity-switch rate and claim-level person attribution, not just badge-detection accuracy. The video summary reverses two events. Which timestamps reached the model? concerns event order and The model found the right moment in a video. Why is its timestamp wrong? timestamp mapping. This is a temporal identity link between two otherwise correct observations.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →