tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Agent Honeypot
Om Joshiabout 2 hours agoContributorJudging locked: Event build

Agent Honeypot

It is a detection infrastructure that is supposed to reveal which AI agents have already been compromised by prompt injection, by logging every time one takes planted bait.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Visit project website
Demo video
Watch demo video
Project description
Agent Honeypot is canarytokens.org for AI agents: detection infrastructure that reveals which agents in a fleet have already been compromised by prompt injection. Prevention tools try to stop injection. None tell you which of your agents already fell for it. Agent Honeypot answers that by planting inert bait and logging who takes it. The bait comes in two forms: a decoy tool named get_admin_credentials that no legitimate task should ever call, and canary documents in markdown, HTML, and PDF, each carrying a unique tracking token. When a compromised agent calls the decoy or reaches the hidden callback in a document, a Flask listener records the hit: which token, the timestamp, source IP, user agent, and which agent tripped it. A live dashboard shows every hit as it happens. The core idea is turning this into a susceptibility benchmark. We send identical bait through two agents: an undefended model and Fable 5.1, a defended model. The undefended agent trips every canary while Fable 5.1 stays dark. That contrast measures exploitability instead of just flagging one incident, so a security team can rank which agents in their fleet are actually safe to deploy. The project targets threat discovery and attack intelligence, and it emphasizes detection over prevention, because the hardest problem in agent security is knowing you have already been breached. Tools: Python, Flask, SQLite, a small undefended victim model, and Fable 5.1 as the defended baseline.
Tools used
  • SSemgrep
  • ClickHouse logoClickHouse
  • PPi
Watch demo video
Project gallery
Project links
  • GitHub repository
  • Project website
  • Demo video
Tools used
  • SSemgrep
  • ClickHouse logoClickHouse
  • PPi