I hid "Ignore all previous instructions" inside a normal-looking security blog, telling AI readers to report a fake breach. My swarm read it. The fake breach never shipped.
Hive is a swarm of cheap open-weight agents that reads untrusted web pages all day, plus a ClickHouse monitoring plane that watches the agents themselves. The scraper is the attack surface; watching the swarm is the product.
In the demo, a live heartbeat runs on real security news and SEC filings about CrowdStrike, Palo Alto Networks, Zscaler and more:
- The dashboard lights up with every agent call by model and role, streamed into ClickHouse.
- The poisoned blog gets caught by a prompt canary, quarantined, and its source's trust zeroed.
- A news site says $45M, the 8-K says $450M. A judge on a different model sides with the filing.
- The writer builds a competitor profile only from claims approved in Senso, every line cited. Ask it something it can't verify and it refuses and files the question to Senso's gap report.
By the numbers:
- 0 planted lies reached a brief, in every run
- 100% injection catch on fixtures, 0 false alarms on 88 live docs
- 96 docs, 745 agent calls, 285 verified claims for 13 cents on AkashML
Built solo in one day, orchestrating four parallel Claude Code sessions. Semgrep plus my own review caught an SSRF in my own scraper along the way.