tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. CityPulse
Anonymous builder36 minutes agoJudging locked: Event build

CityPulse

CityPulse lets you ask a city a natural-language question and receive a grounded 30-second narrated video answer assembled from the most relevant moments in its indexed footage.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Visit project websiteWatch demo videoProject gallery
Demo video

Video demos are proof context. Repo, stack, and build notes stay attached so visitors can inspect what was actually built.

Project description
CityPulse turns urban video archives into an interactive visual storytelling agent. A user asks a question such as “How do pedestrians and vehicles interact at Toronto intersections?” CityPulse searches the indexed Toronto driving footage, surfaces the strongest evidence, selects six complementary shots, writes a grounded narration, and automatically edits them into a 30-second narrated film. The interface makes the agent’s work visible: source routes illuminate as relevant moments are found, evidence clips appear during retrieval, selected shots are ordered into the final cut, and the resulting film answers the original question. CityPulse builds directly on the VAST Video Search & Summary stack. VAST DataEngine stores and retrieves the indexed video evidence. NVIDIA Cosmos Embed powers semantic video retrieval, while Cosmos Reason provides scene understanding and descriptions. YOLO11 detections provide object-level evidence for vehicles, pedestrians, cyclists and other road users. Weights & Biases serverless inference on CoreWeave powers the higher-level AI director that selects shots and writes the narration. The final response is not generated footage: every shot comes from the licensed source archive and remains grounded in retrieved evidence.
Project links
  • GitHub repository
  • Project website
  • Demo video
Tools used
  • VDVAST Data
  • S(SpaceXAI (Cursor)
  • C(CoreWeave (Weights & Biases)
  • NNVIDIA
Tools used
  • VDVAST Data
  • S(SpaceXAI (Cursor)
  • C(CoreWeave (Weights & Biases)
  • NNVIDIA