Tag
TRUSS is an evidence-guided framework for generating functionally effective and safety-reliable Agent Skills, improving task effectiveness from 17.11% to 52.94% and security rate to 100%.
Proposes Anchored Self-Play (ASP), a method for scaling code repair supervision via generator–fixer self-play with an embedding-similarity reward and reference bug mixing, achieving +24% relative improvement in fix rates over standard self-play on a new benchmark BugSourceBench.
This paper empirically analyzes the cost-effectiveness of code execution in LLM-based program repair agents, finding that execution is used heavily but often indiscriminately, and that restricting execution can save significant cost with minimal impact on repair success.