1. Perception: MediaPipe Hand Landmarker reads 21 hand landmarks from each OpenCV webcam frame. We turn them into a 65-number feature vector. 2. Classification: A scikit-learn logistic regression, trained on public HaGRIDv2 landmarks plus synthetic and calibration data, labels each frame. A hold timer weighted by confidence keeps random hand movement from firing commands. 3. Action: PyObjC calls macOS directly (Quartz events, the Accessibility API, overlays). It sends keystrokes, moves windows with its own snapping, and draws on-screen feedback. 4. Self-improvement loop (the hackathon part): Every gesture event is logged. The intent pipeline goes segments → video clips → Cosmos Reason → JEV check → tagger → review → export: - Cosmos Reason (the shared NIM, cosmos3-nano-reasoner) watches each clip and judges what the user meant to do: a real command, a misfire, or a missed gesture. - A JEV pass (over OpenRouter) cross-checks Cosmos's verdict. - An LLM tagger with escalation to a stronger model settles disagreements and writes corrected labels. - The corrected labels go back into retraining.