tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. CrossWise
SAIRAM KUMAR REDDYabout 2 hours agoContributorJudging locked: Event build

CrossWise

CrossWise is a crossing copilot for sidewalk delivery robots that watches street video and decides GO, WAIT or NO-GO every 0.2 seconds, so robots stop getting stuck at crosswalks.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Watch demo video
Demo video
Watch demo video
Project description
Problem. Delivery robots get stuck at crosswalks: they spin in place, miss the walk phase, or start into traffic. What it does. CrossWise watches street video and decides GO, WAIT or NO-GO every 0.2 seconds, with the evidence drawn on the video. When a robot would wait too long, it escalates to a reroute or a remote operator. A scorecard ranks 7 crossings by difficulty for route planning. How it works. Vehicles are tracked from the YOLO11 boxes VAST stores for every frame, with camera-shake removal and 1.5-second path prediction. The pedestrian signal is read from pixels. A vision model on W&B Inference finds cones, barriers and the free path. Unclear evidence forces WAIT. NVIDIA Cosmos Reason captions and Cosmos Embed search, on CoreWeave GPUs, found the crossing footage in the VAST archive. Decisions are stored in VastDB and model calls are traced in W&B Weave. What we learned
Tools used
  • VAST Data logoVAST Data
  • SpaceXAI (Cursor) logoSpaceXAI (Cursor)
  • CoreWeave (Weights & Biases) logoCoreWeave (Weights & Biases)
  • NVIDIA logoNVIDIA
Project gallery
. Our first version had an LLM judge 5-second captions. It confused the car light with the pedestrian signal and parked cars with traffic. We rebuilt the core as frame-level perception and measured both on 5 clips: starts against DON'T WALK fell from 3 to 0, crossings rose from 6 to 8, and time from WALK to rolling dropped from 3.4 s to 1.0 s. Limits. Decision support on recorded footage; it does not control a robot. Built end to end with Cursor.
Project links
  • GitHub repository
  • Demo video
Tools used
  • VAST Data logoVAST Data
  • SpaceXAI (Cursor) logoSpaceXAI (Cursor)
  • CoreWeave (Weights & Biases) logoCoreWeave (Weights & Biases)
  • NVIDIA logoNVIDIA