tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Grabbit Mining Co.
Jacobabout 2 hours agoContributorJudging locked: Event build

Grabbit Mining Co.

Describe a dangerous moment in one sentence; Grabbit digs through pre-indexed warehouse and road footage and hands you a verified, exportable test set of edge cases for robotics/AV policy evaluation. AI proposes, a human decides.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Visit project websiteWatch demo videoProject gallery
Demo video

Video demos are proof context. Repo, stack, and build notes stay attached so visitors can inspect what was actually built.

Project description
Robotics and AV teams own thousands of hours of footage nobody watched. Grabbit Mining Co. turns one sentence ("forklift within 2 m of a worker in an aisle") into a human-verified test set of edge cases. How it works (every step runs on the sponsor stack): 01 Expand — W&B Inference rewrites the sentence into 3 query variants (traced in Weave). 02 Search — Cosmos Embed vectors in VastDB over 2,352 segments indexed by VAST DataEngine; all cameras, deduped, top 12. 03 Evidence — Cosmos Reason captions + YOLO11 detections from the VSS pipeline on CoreWeave GPUs. We re-ingested 8 warehouse chunks with a custom RISK prompt (distance to forklift, walkway, running) because the default captions lacked those facts. 04 Severity — W&B Inference returns strict JSON {is_match, severity 1-5, tags, rationale}; median 402 ms per clip. 05 Review — a human presses Confirm or Reject on the page. Grabbit can't reach its own button; that's the safety feature. 06 Act — export test set (JSONL) for policy eval. Key finding: the only severity-5 clip for the forklift query sat at search rank 9, and at rank 7 for person-vs-vehicle. Similarity search alone would have buried both. You can't just search; you have to mine. Today's run: 3 queries, 36 hits, 17 flagged, 15 human-reviewed, 60% confirmed, 9 in the test set. Built in Cursor with its VAST skills. The page is a static replay of today's run so nothing breaks on stage; the pipeline is live on the VM (mine.py).
Project links
  • GitHub repository
  • Project website
  • Demo video
Tools used
  • VDVAST Data
  • C(CoreWeave (Weights & Biases)
  • NNVIDIA
  • S(SpaceXAI (Cursor)
Tools used
  • VDVAST Data
  • C(CoreWeave (Weights & Biases)
  • NNVIDIA
  • S(SpaceXAI (Cursor)