One laptop camera watches the class. On the laptop, YOLO11-pose, YOLOE and YAMNet track the student and spot loud things by name (a blender, a drill) before they make a sound, plus crowding, someone rushing in or someone staying too close. As risk rises, calming sound fades into the student's AirPods early and the teacher's phone gets an alert with the right IEP accommodation. Covering ears or rocking triggers an overload alert.
NVIDIA Cosmos3-Reason on CoreWeave GPUs reads the last few seconds of video and says what is about to get overwhelming. Canary-1B transcribes the lecture, so a student with ADHD can ask "what did I miss?" about the minutes they looked away; Llama 3.3 70B on W&B Inference answers. Weave traces every model call.
On VAST, we scored all 2,352 clips in our archive for sensory load from the Cosmos captions and YOLO11 counts (771 high, 566 calm) and built a sensory map: the SF street cameras are the most overwhelming (76% high load), the neighborhood camera the calmest (97% calm). Our VM app adds a Cosmos risk timeline over archive video, VAST search with clip streaming, re-ingest with our sensory prompt, and W&B summaries of the calmest places and times.