Tag
The paper introduces WideSWE, a benchmark for evaluating coding agents on cross-repository tasks, based on 120 real-world tasks, showing that agent success rates vary and joint execution can improve performance.