As somebody who loves to travel around Asia, it's always very difficult to buy local. We use the video understanding models to convert video footage of markets into real high-quality e-commerce stores so that local people can reach travelers and expats, investing more dollars into local economies.
Built today on the event stack: walkthrough videos go into our team's VAST S3 bucket; on the workshop VM each video is cut into 5-second pieces, NVIDIA Cosmos Reason on CoreWeave GPUs lists every item for sale, Llama 3.3 70B on Weights & Biases Inference turns that into clean catalogue entries, Cosmos Embed1 indexes every item and moment, and it all lives in VastDB. The shop reads that catalogue: 97 pieces from markets in New York, Bali and Chiang Mai, each linked to the exact 5 seconds it appears on camera. A 60-second walkthrough takes about 90 to 120 s of compute on one worker and scales out linearly.