Inspiration: Companies record how their best people work, and almost nobody watches it. New hires still learn by shadowing, and mistakes get caught late. We wanted that unwatched video to coach the next person live.
How each sponsor is used:
- VAST Data (VAST AI OS): stores the 2,352-clip archive and our own recorded takes in S3. We segmented, captioned, embedded and inserted our takes into the team VastDB ourselves (2,409 rows). VAST semantic search powers "Ask the archive".
- NVIDIA Cosmos Reason: live scene reading about once a second during coaching, plus clip captions and task recognition.
- NVIDIA Cosmos Embed: vectors for the archive map, task recognition and search.
- CoreWeave: GPUs serving all the Cosmos and YOLO models.
- Weights & Biases: serverless inference. Llama 3.3 70B writes the score feedback and names the map clusters, DeepSeek V4 Flash is the fallback judge, and Weave traces every live check.
- Cursor / SpaceXAI: Used for creating the project