tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. Ward
Pratham Agarwal 2104about 2 hours agoContributorJudging locked: Event build

Ward

Ward is an autonomous security engineer for AI apps: on every git push it attacks your code and your AI agent, fixes what it broke into, re-attacks to prove the fix, then rebuilds the attack from your logs alone to find what your monitoring missed, and opens a PR for a human to approve.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Watch demo video
Demo video
Watch demo video
Project description
Ward runs a security engineer's loop on an AI app, triggered by a GitHub push webhook. No human presses run. 1. Discover: scans code with our own Semgrep AI-security ruleset and probes the live AI agent with prompt injection. On our reference helpdesk bot it finds a hardcoded secret and a SQL injection, rejects an eval() decoy, and gets the bot to email out an internal document. 2. Remediate: fixes code and agent, then re-attacks until zero exploits reproduce. Verdicts are deterministic code (tests, replayed probes, scanner), never the agent grading itself. 3. Detective: rebuilds the attack from the app's logs alone. It recovers 2 of 3 steps (67%); the exfiltration never shows in the logs, so Ward flags the blind spot and writes a detection rule. Penetrable is half the risk; blind is the other half. 4. Output: a real GitHub PR for human approval, an email, and a case file mapped to OWASP-LLM, CWE and MITRE ATLAS/ATT&CK. A run takes ~2-3 min and ~$0.05. Sponsors Stack: Semgrep (custom ruleset), ClickHouse (attack telemetry: 1M simulated events, 50-209 ms queries), Guild.ai (published ward-triage governance agent: APPROVE_PR / NEEDS_HUMAN_REVIEW / BLOCK). Stack: TypeScript (Node), OpenAI (gpt-4o agent, gpt-4o-mini target), MongoDB Atlas, GitHub webhooks + PRs, Resend. Scope: runs on our reference app today; a config adapter for any HTTP AI endpoint is next.
Tools used
  • Guild.ai logoGuild.ai
  • SSemgrep
  • ClickHouse logoClickHouse
Project gallery
Project links
  • GitHub repository
  • Demo video
Tools used
  • Guild.ai logoGuild.ai
  • SSemgrep
  • ClickHouse logoClickHouse