Where should the bench go? Designers answer that by watching a plaza for weeks or scrubbing hours of footage, like William Whyte did. Most streets never get that attention, so seating, lighting and bus stop decisions get guessed. Linger is video agent that does this observation at scale. It scans ever camera in the VAST video library, ranks the places where people stop, wait and gather (linger), and lets a planner explore one: where people stop by day and night, what tends to happen next (a bus arrives, then a group forms where it stopped), and three plausible futures, like seating or lighting, each citing the real clips behind it. We re-ingested the footage with a behavior-only prompt, so NVIDIA Cosmos3-Reason never records faces, identity, or clothing. YOLO11 detects people and vehicles on CoreWeave GPUs. The agent reasons with W&B Inference, traces every step in Weave, and a Weave evaluation checks it never cites a missing clip or claims cause.
Built for planners, transit agencies and campuses: weeks of observation in an afternoon, backed by evidence anyone can watch.