LevelForge is a testbed for long-horizon agent memory — an LLM repairs a 2D level action-by-action while a deterministic validator, not the model, judges whether it actually solved it.
Review the project
Start with the source code, then open the demo or video if available.