Robotics and AV teams own thousands of hours of footage nobody watched. Grabbit Mining Co. turns one sentence ("forklift within 2 m of a worker in an aisle") into a human-verified test set of edge cases.
How it works (every step runs on the sponsor stack):
01 Expand — W&B Inference rewrites the sentence into 3 query variants (traced in Weave).
02 Search — Cosmos Embed vectors in VastDB over 2,352 segments indexed by VAST DataEngine; all cameras, deduped, top 12.
03 Evidence — Cosmos Reason captions + YOLO11 detections from the VSS pipeline on CoreWeave GPUs. We re-ingested 8 warehouse chunks with a custom RISK prompt (distance to forklift, walkway, running) because the default captions lacked those facts.
04 Severity — W&B Inference returns strict JSON {is_match, severity 1-5, tags, rationale}; median 402 ms per clip.
05 Review — a human presses Confirm or Reject on the page. Grabbit can't reach its own button; that's the safety feature.
06 Act — export test set (JSONL) for policy eval.
Key finding: the only severity-5 clip for the forklift query sat at search rank 9, and at rank 7 for person-vs-vehicle. Similarity search alone would have buried both. You can't just search; you have to mine.
Today's run: 3 queries, 36 hits, 17 flagged, 15 human-reviewed, 60% confirmed, 9 in the test set.
Built in Cursor with its VAST skills. The page is a static replay of today's run so nothing breaks on stage; the pipeline is live on the VM (mine.py).