SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Summary
Introduces SWE-Touch, a benchmark framework that injects conflicting user edits during agent coding trajectories, showing that current coding agents significantly degrade in collaborative settings despite strong standalone benchmark performance.
View Cached Full Text
Cached at: 08/04/26, 05:37 AM
Paper page - SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Source: https://huggingface.co/papers/2608.02499 SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
Current coding agent benchmarks (SWE-Bench, etc.) evaluate agents working alone on a static codebase. But in real development, users actively inspect and modify code while the agent is working — our analysis of SWE-chat data shows 59% of sessions contain user-authored repository changes.
**What we do:**We introduce SWE-Touch, a framework that stress-tests coding agents in a shared workspace by injecting validated, task-conflicting user edits during an ongoing repair trajectory. The edits are small, plausible code changes that conflict with task completion, delivered with contextual user messages when the agent reaches the relevant code.
Key findings across 9 frontier models on SWE-bench Verified:
- Counter-Edit lowers average resolve rate by 7.7 points and reshuffles model rankings
- 63.3% of failed runs simply retain the user’s conflicting code untouched
- Only Claude Opus 4.8 and GPT 5.5 show strong resilience; open-source models that score competitively on autonomous benchmarks degrade substantially (up to 16.5 points)
- The degradation persists on longer-horizon tasks (SWE-Bench Pro & DeepSWE)
- Ablations show the code edit itself — not the accompanying message — drives the performance drop
The results suggest that optimization focused on static leaderboard performance does not ensure robustness in collaborative settings. Agents need to detect workspace changes, reconcile conflicts, and re-validate affected behavior — capabilities that current models largely lack.
Similar Articles
SWE-Together: Evaluating Coding Agents in Interactive User Sessions
SWE-Together is a multi-turn coding benchmark created from real user-agent interactions, featuring a reactive LLM simulator to evaluate agents based on both final correctness and interaction efficiency.
SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
SWE-Interact is a new testbed that evaluates coding agents in realistic multi-turn, user-driven software engineering tasks, revealing that strong single-turn benchmark performance does not reliably transfer to interactive, iterative workflows where agents must discover user intent and adapt to evolving requirements.
SWE-chat: Coding Agent Interactions From Real Users in the Wild
SWE-chat introduces a 6,000-session dataset of real-world coding agent interactions, revealing that only 44% of agent-generated code survives in commits and highlighting inefficiencies and security issues in current AI-assisted development.
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
SWE-bench Science introduces a repository-level benchmark for evaluating coding agents on scientific software repair tasks, revealing failure mechanisms and mixed effects of scientific guidance.
SWE Context Bench just proved something I think a lot of coding agent users already feel
A new benchmark paper 'SWE Context Bench' tests whether coding agents can reuse knowledge across tasks, highlighting a gap in existing benchmarks that only evaluate isolated problem-solving. The author discusses solutions like external memory and mentions tools such as langmem, mem0, supermemory, and Greplica.