Agent Honeypot is canarytokens.org for AI agents: detection infrastructure that reveals which agents in a fleet have already been compromised by prompt injection.
Prevention tools try to stop injection. None tell you which of your agents already fell for it. Agent Honeypot answers that by planting inert bait and logging who takes it. The bait comes in two forms: a decoy tool named get_admin_credentials that no legitimate task should ever call, and canary documents in markdown, HTML, and PDF, each carrying a unique tracking token. When a compromised agent calls the decoy or reaches the hidden callback in a document, a Flask listener records the hit: which token, the timestamp, source IP, user agent, and which agent tripped it. A live dashboard shows every hit as it happens.
The core idea is turning this into a susceptibility benchmark. We send identical bait through two agents: an undefended model and Fable 5.1, a defended model. The undefended agent trips every canary while Fable 5.1 stays dark. That contrast measures exploitability instead of just flagging one incident, so a security team can rank which agents in their fleet are actually safe to deploy.
The project targets threat discovery and attack intelligence, and it emphasizes detection over prevention, because the hardest problem in agent security is knowing you have already been breached.
Tools: Python, Flask, SQLite, a small undefended victim model, and Fable 5.1 as the defended baseline.