Model and Inference Engineering · Staff
Audio and video align at the start. Why is the live agent seconds off after twenty minutes?
Take a few minutes to form your approach. Then open a worked answer and compare the decisions.
Reveal a worked answer
A one-time offset is not enough if the streams have different clocks or the pipeline advances them using different rules. Audio may carry timestamps in sample-clock units while video timestamps use another rate. If the ingest code assigns video time from frame count at a nominal 30 frames per second and audio time from actual captured samples, dropped frames or variable frame rate can make the two timelines diverge. Separate device clocks can drift too. RFC 3550 uses RTP timestamps and RTCP sender reports to relate media clocks to a common time for inter-media synchronization. Do not compare raw RTP timestamp numbers across audio and video as though they share units.
I would plot the difference between audio presentation time and video presentation time at the beginning, middle and end. A constant gap points to an origin or buffering offset. A gap that grows with time points to rate conversion, timestamp synthesis or clock drift. A sudden jump points to dropped packets, a discontinuity or segment stitching. Trace capture timestamps, packet timestamps, decoder output presentation timestamps and the timestamp finally sent to the model. A correct incoming stream can still be ruined when a transcoder resets a segment's time base without preserving its offset.
For a concrete failure, suppose the pipeline receives 29 frames during one second of a nominal 30 fps stream and then labels each received frame with the next frame-count timestamp. It advances video time by about 0.967 seconds for that second while audio advances by one second. Repeated loss accumulates a mismatch. The right treatment is to preserve timestamps and represent missing frames as missing, not compress time to make frame indices contiguous. A playback system may choose to drop, duplicate or resample for user experience. The evidence pipeline still needs a stable mapping back to source time.
Test with a visible clap and audible clap at several points in a long recording, then compare the model's audiovisual claim to the mapped timestamps. Include a reconnect and a transcoder restart. A single clap at the start can pass while the timeline is several seconds wrong by the end.
Continue reading
Related questions
Read beyond the question
Explore more model and inference engineering
Follow another question in this area, or search the complete Question Library.
Browse this area →Browse Question Library →