A lot of video footage is surprisingly static, and running a full interference sweep on unchanging frames is a waste of precious GPUs. We're optimizing the pipeline by using codec motion vectors, used for video codec compression, to crudely detect motion for free. This is a simple heuristic that gates a full YOLO run on the entire frame, saving up to 70% for low motion scenes, but could easily be extended to specific regions of the video and more expensive models. Everything uses my open source real-time media transport (moq.dev) allowing the AI model to only run on demand, saving even more money for rarely monitored security footage.
YOLO is running on my desktop PC at home so I can process individual frames. It's connected to my global CDN (moq.pro) which it uses to both fetch the camera footage and publish the AI detection results (as a track). The web player subscribes to tracks on demand based on the camera selected and options. We emulate live footage using VAST's sample videos slowly tricked over the network (ffmpeg -re) and everything is real-time. Object detections and motion vectors (extracted by the desktop) trail the footage by at least 2 frames (plus RTT) because we intentionally don't synchronize the sources. Real-time latency is critical for some use-cases, like drones, so we felt it would make a more honest demo.