I built an autonomous attack-fix-verify agent for AI-generated APIs, it attacks a running API like a real adversary over HTTP (IDOR, SQL injection, sensitive-data exposure, missing auth), confirms each bug by comparing the actual responses, then uses an LLM (Gemini) to propose a minimal patch for just the vulnerable handler and crucially, it doesn't trust the AI's fix, it trusts the re-attack: deterministic code re-runs the exact exploit and only marks a fix "verified" if the attack now fails, the legitimate good path still works, and the app still starts (else it retries up to 3× or flags for human review), after which it opens a GitHub pull request with the patch, auto-generated regression tests, and evidence. Semgrep is layered in as the static-analysis second opinion, it pre-scans the code for suspects (file/line/CWE), feeds those into the fix prompt, re-scans after patching, and most importantly cross-checks static vs. dynamic findings. On my app Semgrep caught only the SQL injection and missed the IDOR, data-exposure, and missing-auth logic bugs, which is exactly the point static analysis shows where code looks wrong, while the active attacker proves what's actually exploitable, so the two together are far stronger than either alone.