tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Warehouse best buddy
Hui Xuabout 2 hours agoContributorJudging locked: Event build

Warehouse best buddy

A real-time video agent for warehouse safety and operations. It watches every camera through the VSS stack on VAST, catches near misses between workers and forklifts, finds idle labor and idle machines, prices the idle time per shift, and answers questions with cited clips.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Visit project websiteWatch demo videoProject gallery
Demo video
Watch demo video
Project description
For warehouse safety managers and operations leads: the cameras already record every shift, but nobody has time to watch them. Real time: every 30 s the agent checks what VSS has indexed, analyzes new footage once it is fully indexed, and the app updates within seconds. Safety you can trust: motion tracks from the YOLO detections propose candidate near misses, NVIDIA Cosmos Reason checks that exact time window and image region, and the camera views vote. Asked directly, Cosmos called 168 of 180 clips a near miss, so the agent never relies on that question. With verification, 0 of 12 random control moments were flagged and all 3 drill near misses were caught. Operations: idle labor, idle machines, congestion and underused zones, with recommendations tied to the evidence and an idle-cost estimate with editable assumptions (about $859 recoverable per shift, about $215k a year at the defaults). Ask and report: plain-English questions answered by NVIDIA Nemotron-3-Ultra on W&B Inference, grounded in the analyzer's facts and VSS search, citing alerts and clips. Every number in the shift report is checked against the data. Results go back to VastDB (alerts, flags, per-segment metrics), keyed by each segment's S3 URI, so they join with the VSS captions and embeddings. Built with NVIDIA Cosmos Reason and Nemotron, VAST S3, DataEngine and VastDB, CoreWeave GPUs, W&B Inference and Cursor, on 60 videos from 33 camera views.
Project links
  • GitHub repository
  • Project website
  • Demo video
Tools used
  • VAST Data logoVAST Data
  • SpaceXAI (Cursor) logoSpaceXAI (Cursor)
  • CoreWeave (Weights & Biases) logoCoreWeave (Weights & Biases)
  • NVIDIA logoNVIDIA
Tools used
  • VAST Data logoVAST Data
  • SpaceXAI (Cursor) logoSpaceXAI (Cursor)
  • CoreWeave (Weights & Biases) logoCoreWeave (Weights & Biases)
  • NVIDIA logoNVIDIA