For developers and security teams who let coding agents run commands on their machines. Agents act with the developer's permissions, and nobody reviews what they did across hundreds of sessions.
We collected one developer's real logs: 345 Claude Code and 199 Codex sessions from a Mac and a linux-vm, plus the proxy and cost ledger. Read from those logs:
- 292 destructive commands (forced deletes, hard resets, forced pushes)
- 24 tool results containing a key or token pattern
- 166 of 345 sessions ran with permission checks off at some point
- 25 tool calls blocked by a hook
- 123+ prompts where the developer had to correct the agent; the agent asked 162 questions while 137 of its calls were rejected
The agent turns this into action: it watches new sessions as they end, writes a rule for each piece of context the agent lacked, publishes a weekly counts-only report, and proposes fixes to CLAUDE.md, skills, hooks or settings as pull requests the developer approves. Every number cites the sessions and records behind it.
Built today: log collection from both machines, the measurements above, and the PRD and technical spec (linked).