SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Hugging Face Daily Papers 06/29/26, 12:00 AM Papers

coding-agents benchmark multi-turn interaction-simulator user-simulator llm repository-level-tasks

Summary

SWE-Together is a multi-turn coding benchmark created from real user-agent interactions, featuring a reactive LLM simulator to evaluate agents based on both final correctness and interaction efficiency.

Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals, adding constraints, and correcting mistakes over multiple turns. We introduce SWE-Together, a multi-turn benchmark reconstructed from real user-agent coding sessions. To make real interactions verifiable, we curate 109 repository-level tasks from 11,260 recorded sessions, selecting sessions with recoverable repository states, clear user goals, and observable outcomes. To replay these interactions across agents, we build a reactive LLM-based user simulator that preserves the original users' intents and provides feedback when the coding agent's progress requires it. To evaluate agents as collaborators, we measure both final repository correctness and the number of corrective feedback turns required during the interaction. Experiments with frontier coding agents show that stronger agents generally achieve higher final success rates while requiring fewer interventions, suggesting an improved user experience.

Original Article

View Cached Full Text

Cached at: 06/30/26, 07:37 PM

Paper page - SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Source: https://huggingface.co/papers/2606.29957 Authors:

Abstract

Mostcoding-agent benchmarksare static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals, adding constraints, and correcting mistakes over multiple turns. We introduce SWE-Together, amulti-turn benchmarkreconstructed from realuser-agent coding sessions. To make real interactions verifiable, we curate 109repository-level tasksfrom 11,260 recorded sessions, selecting sessions with recoverable repository states, clear user goals, and observable outcomes. To replay these interactions across agents, we build a reactive LLM-baseduser simulatorthat preserves the original users’ intents and provides feedback when the coding agent’s progress requires it. To evaluate agents as collaborators, we measure bothfinal repository correctnessand the number ofcorrective feedback turnsrequired during the interaction. Experiments with frontier coding agents show that stronger agents generally achieve higher final success rates while requiring fewer interventions, suggesting an improved user experience.

View arXiv page View PDF Project page GitHub2 Add to collection

Get this paper in your agent:

hf papers read 2606\.29957

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2606.29957 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2606.29957 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2606.29957 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Paper page - SWE-Together: Evaluating Coding Agents in Interactive User Sessions

Abstract

Models citing this paper0

Datasets citing this paper0

Spaces citing this paper0

Collections including this paper0

Similar Articles

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

SWE-chat: Coding Agent Interactions From Real Users in the Wild

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents

Submit Feedback

Similar Articles

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

SWE-chat: Coding Agent Interactions From Real Users in the Wild

Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills

EvoCode-Bench: Evaluating Coding Agents in Multi-Turn Iterative Interactions

Dialogue SWE-Bench: A Benchmark for Dialogue-Driven Coding Agents