tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Honeypot generator for RL environments
Soham Pardeshiabout 2 hours agoContributorJudging locked: Event build

Honeypot generator for RL environments

Generates honeypot RL environments for detecting misaligned AI behavior

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Demo video
Watch demo video
Project description
The recent OpenAI + HuggingFace incident reveals the dangers of reward hacking during post-training. Models with misaligned behavior may collude and hack into services to complete their tasks. As more SMB's are beginning to finetune their own models using open-weights models, it is critical that folks have confidence in their agent. But how can we detect models with misaligned behavior? One option is to generate honeypots. These are RL environments that cannot be solved by an LLM without performing misaligned behaviors. The only way that a model can complete the task is by performing an exploit. Models that succeed in a honeypot environment are necessarily dangerous.  Our tool takes existing RL environments on HuggingFace and mutates them to generate honeypot tasks. We use SemGrep to detect vulnerabilities to (1) find candidate RL environments and (2) confirm that our honeypots are exploitable. We use AkashAI
Tools used
  • Guild.ai logoGuild.ai
  • Claude Code Router logoClaude Code Router
  • AAkash
  • SSemgrep
  • ClickHouse logo
Watch demo videoProject gallery
for open weights models and
Anthropic
SDK for closed weights models to run our evaluations. We use
ClickHouse
to store task data, exploit data, and results. We use
Guild AI
agents to download candidate RL environments from HuggingFace and run our prechecks and validations on them.
Project links
  • GitHub repository
  • Demo video
Tools used
  • Guild.ai logoGuild.ai
  • Claude Code Router logoClaude Code Router
  • AAkash
  • SSemgrep
  • ClickHouse logoClickHouse
  • HFHugging Face
ClickHouse
  • HFHugging Face