tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Auraafit
Tarun Theegelaabout 2 hours agoContributorJudging locked: Event build

Auraafit

AURAAFIT turns street-camera video into outfit analytics a merchandiser can act on: what people wear and carry, with a VAST clip and a W&B Weave trace behind every label, and without guessing who anyone is.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Visit project website
Demo video
Watch demo video
Project description
Stores count foot traffic, but they can't see what the street is wearing. AURAAFIT reads outfits from video. It starts from event footage indexed in VAST VSS (612 clips, 13 cameras). NVIDIA YOLO11 finds and tracks each person, and W&B Inference vision (Gemma 4) labels what they wear and carry: top, bottom, outer layer, bags. Every labelling call is traced in W&B Weave. Requiring a real garment type in our prompt took typed labels from 49 to 213 out of 214. NVIDIA Cosmos3-Reason adds a scene caption for each 5-second piece. The app at auraafit.tech replays recorded street cameras from San Francisco and New York with the labels on the boxes. Buyers can search the labels ("backpack"), and each match links to its clip and timestamp in VAST. A dashboard rolls up 214 person-sightings into colour mix, style mix and carried items, with CSV export. A merchandising agent reads only the totals and drafts store actions, such as putting carry accessories near the entrance because 32% of sightings carry a backpack. Privacy by design: a guard filter keeps clothing and carried items only. Age, gender, race and faces are never inferred, and nothing is tracked across cameras. Counts are sightings, not unique people. Built for retailers, landlords and business districts deciding what to stock or lease on a block. Next step: a consenting storefront camera, over repeated days.
Tools used
  • VAST Data logoVAST Data
  • ElevenLabs Conversational AI QA - Delete Safe logoElevenLabs Conversational AI QA - Delete Safe
  • Claude Code Router logoClaude Code Router
  • NVIDIA logoNVIDIA
  • SpaceXAI (Cursor) logoSpaceXAI (Cursor)
Watch demo video
Project gallery
Project links
  • GitHub repository
  • Project website
  • Demo video
Tools used
  • VAST Data logoVAST Data
  • ElevenLabs Conversational AI QA - Delete Safe logoElevenLabs Conversational AI QA - Delete Safe
  • Claude Code Router logoClaude Code Router
  • NVIDIA logoNVIDIA
  • SpaceXAI (Cursor) logoSpaceXAI (Cursor)
  • CoreWeave (Weights & Biases) logoCoreWeave (Weights & Biases)
  • CoreWeave (Weights & Biases) logoCoreWeave (Weights & Biases)