Anonymous builderabout 2 months agoJudging locked: Self-Evolving Agents Hackathon

JudgeAndJury

An LLM Judge that improves over usage that improves user coding intent by coding agents.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Project description
I have an LLM Judge that tracks coding agent turns and user input and judges the coding agent output and if it's satisfactory for the user. The goal is to reduce LLM coding friction and increase coding agent understanding of user intention. The platform also supports trace-tracking for when errors occur, or when critical infrastructure problems arise. The traces are logged, and using deterministic checks, it would not occur again. So, in total, it tracks errors, mistakes, coding agent turns, user turns, and improves the LLM judge over time through these traces. Users are also able to fine-tune their own model through these traces and specifically create a model just for themselves.
Tools used
  • Actian
  • Pioneer
  • Replay.io
  • Senso