Ward runs a security engineer's loop on an AI app, triggered by a GitHub push webhook. No human presses run.
1. Discover: scans code with our own Semgrep AI-security ruleset and probes the live AI agent with prompt injection. On our reference helpdesk bot it finds a hardcoded secret and a SQL injection, rejects an eval() decoy, and gets the bot to email out an internal document.
2. Remediate: fixes code and agent, then re-attacks until zero exploits reproduce. Verdicts are deterministic code (tests, replayed probes, scanner), never the agent grading itself.
3. Detective: rebuilds the attack from the app's logs alone. It recovers 2 of 3 steps (67%); the exfiltration never shows in the logs, so Ward flags the blind spot and writes a detection rule. Penetrable is half the risk; blind is the other half.
4. Output: a real GitHub PR for human approval, an email, and a case file mapped to OWASP-LLM, CWE and MITRE ATLAS/ATT&CK. A run takes ~2-3 min and ~$0.05.
Sponsors Stack: Semgrep (custom ruleset), ClickHouse (attack telemetry: 1M simulated events, 50-209 ms queries), Guild.ai (published ward-triage governance agent: APPROVE_PR / NEEDS_HUMAN_REVIEW / BLOCK).
Stack: TypeScript (Node), OpenAI (gpt-4o agent, gpt-4o-mini target), MongoDB Atlas, GitHub webhooks + PRs, Resend.
Scope: runs on our reference app today; a config adapter for any HTTP AI endpoint is next.