tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Differential Swarm - Self Improving Security Evals
Hardik Goyalabout 2 hours agoContributorJudging locked: Event build

Differential Swarm - Self Improving Security Evals

Differential Swarm is a set of dynamic security evals that runs on every PR: a swarm of red-team agents attacks the old and new builds in paired sandboxes and flags any regression in the agent’s safety.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Watch demo videoProject gallery
Demo video
Watch demo video
Project description
Differential Swarm tests how an AI agent’s behavior changes with every pull request. It maps the change’s blast radius, generates targeted attacks, and sends a swarm of red-team agents (scam callers, manipulative sellers, poisoned listings) against the old and new builds in isolated paired sandboxes. Comparing each pair shows regressions, fixes and pre-existing failures. Our demo uses a shopping agent facing fraudulent voice calls, manipulative sellers and malicious listings. A live dashboard shows the swarm’s progress and lets reviewers inspect conversations, tool calls and payment evidence side by side. Each finding becomes a GitHub remediation issue and a verified Semgrep rule, which turns behavioral discoveries into reusable static checks. That’s how the evals improve themselves.
Project links
  • GitHub repository
  • Demo video
Tools used
  • ElevenLabs Conversational AI QA - Delete Safe logoElevenLabs Conversational AI QA - Delete Safe
  • OpenAI Codex CLI logoOpenAI Codex CLI
  • SSemgrep
  • ClickHouse logoClickHouse
  • PPi
Tools used
  • ElevenLabs Conversational AI QA - Delete Safe logoElevenLabs Conversational AI QA - Delete Safe
  • OpenAI Codex CLI logoOpenAI Codex CLI
  • SSemgrep
  • ClickHouse logoClickHouse
  • PPi