WORLDSTATE turns video from passive footage into operational memory.
Instead of only detecting objects or describing individual frames, WORLDSTATE watches repeated executions of a physical process and learns the normal sequence of states and transitions without manually labeling each step.
In our robot pick-and-place demo, WORLDSTATE learns the normal process: approach, grasp, lift, carry, descend, and place. When an object unexpectedly shifts sideways, the system detects the exact moment reality diverges from what it learned — at 1.904 seconds — and flags it as a novel failure.
NVIDIA Cosmos then explains what physically happened in natural language, while YOLO tracks the objects involved.
The operator can click Remember This Failure, adding the event to WORLDSTATE's memory. When a different run later fails in the same way, it is immediately recognized as a known failure instead of an unknown anomaly.
WORLDSTATE also includes searchable event memory and a 3D digital twin for visualizing failures and simulated recovery paths.
We additionally integrated our team's real VAST warehouse corpus: 40 clips with Cosmos Reason captions and YOLO11 detections that can be searched directly from the WORLDSTATE interface.
The goal is simple: give robots, factories, and warehouses a video agent that learns what should happen — and remembers when reality breaks.