program-repair

Tag

Cards List
#program-repair

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

arXiv cs.AI · 6d ago Cached

TRUSS is an evidence-guided framework for generating functionally effective and safety-reliable Agent Skills, improving task effectiveness from 17.11% to 52.94% and security rate to 100%.

0 favorites 0 likes
#program-repair

@carolineschoi: I’ll be at the poster today in Hall A, 5–7:45pm KST. Come say hi if you want to talk about synthetic data, self-play, o…

X AI KOLs Following · 2026-07-08 Cached

Proposes Anchored Self-Play (ASP), a method for scaling code repair supervision via generator–fixer self-play with an embedding-similarity reward and reference bug mixing, achieving +24% relative improvement in fix rates over standard self-play on a new benchmark BugSourceBench.

0 favorites 0 likes
#program-repair

To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair

Hugging Face Daily Papers · 2026-06-25 Cached

This paper empirically analyzes the cost-effectiveness of code execution in LLM-based program repair agents, finding that execution is used heavily but often indiscriminately, and that restricting execution can save significant cost with minimal impact on repair success.

0 favorites 0 likes
← Back to home

Submit Feedback