Terminal Agents: A Survey of AI Agents in Command-Line Environments
Summary
This survey paper examines AI agents operating in command-line environments, proposing a framework to unify research across software engineering and other domains through a terminal-mediated execution lens.
View Cached Full Text
Cached at: 08/24/26, 04:11 AM
# A Survey of AI Agents in Command-Line Environments
Source: [https://arxiv.org/html/2608.20485](https://arxiv.org/html/2608.20485)
## Terminal Agents: A Survey of AI Agents in Command\-Line EnvironmentsThanks:These authors contributed equally to this research\.
Xiaoyang Yuan11footnotemark:1Affiliation:Yi BinAffiliation:Wei YeAffiliation:Shanghai Innovation Institute, Shanghai, ChinaHaoxi Zeng11footnotemark:1Wencheng Ye11footnotemark:1Affiliation:Yi BinAffiliation:Wei Ye\[0\.35em\] Wenqi ShaoAffiliation:Shanghai Innovation Institute, Shanghai, ChinaChen QianAffiliation:Shanghai Jiao Tong University, Shanghai, ChinaYujuan DingAffiliation:The Hong Kong Polytechnic University, Hong Kong, China\[0\.35em\] Zheng WangAffiliation:Yi BinAffiliation:Wei YePengpeng ZengAffiliation:Yi BinAffiliation:Wei YeJingkuan SongAffiliation:Yi BinAffiliation:Wei YeHeng Tao ShenAffiliation:Yi BinAffiliation:Wei Ye\[0\.45em\] Tongji UniversityAffiliation:Yi BinAffiliation:Wei YeShanghaiChina
###### Abstract
Large language model agents increasingly act through terminals, yet existing surveys disperse terminal\-mediated behavior across software engineering, tool use, and computer\-use research\. We regard terminal agents as systems whose dominant progress\-bearing action–observation loop is mediated by terminal command execution, textual feedback, and stateful environment interaction\. Using terminal\-mediated execution as an organizing lens, this survey establishes workload\-level boundaries and connects system architecture, competence acquisition, and evaluation through a seven\-dimensional terminal competence profile\. Our synthesis shows that realized behavior is jointly shaped by the model, interface, harness, runtime, and environment\. Executable trajectories ground learning in action consequences, verification, and recovery, whereas prevailing evaluations emphasize final outcomes and expose process quality, recovery, and governance unevenly\. Bounded fixed\-condition diagnostics illustrate two implications: benchmark families expose different process signals, and matched system comparisons reveal benchmark\-dependent performance and limits of component attribution\. These findings motivate explicit reporting of system and runtime conditions, supported by replayable traces and process\-level evidence\. The framework provides a unified basis for studying terminal\-mediated agency across software engineering and emerging application domains\.
## 1Introduction
Large language models \(LLMs\) are evolving from prompt\-bound text generators into interactive systems that act in external environments and revise their behavior from feedback\[[170](https://arxiv.org/html/2608.20485#bib.bib187),[123](https://arxiv.org/html/2608.20485#bib.bib84),[149](https://arxiv.org/html/2608.20485#bib.bib184)\]\. Through iterative action and observation, these systems increasingly operate as agents that execute code, invoke computational tools, navigate graphical interfaces, and modify digital environments over multiple steps\[[111](https://arxiv.org/html/2608.20485#bib.bib75),[59](https://arxiv.org/html/2608.20485#bib.bib47)\]\. The medium through which such agents act shapes their available actions, observable state, feedback, and means of verification, and is therefore part of the agent system rather than a neutral implementation channel\.
The terminal is a particularly consequential execution medium because it provides compact, scriptable textual access to stateful runtimes comprising filesystems, dependencies, processes, tests, logs, remote machines, and command\-line tools\. Commands can modify files, configure environments, launch services, and execute tests, while outputs, exit codes, diffs, stack traces, logs, and artifacts inform subsequent decisions\. Terminal\-mediated execution thus couples reasoning, execution, observation, and verification through a mutable action–observation loop\[[29](https://arxiv.org/html/2608.20485#bib.bib65)\]\. When this loop is the principal mechanism of task progress, the terminal becomes an execution substrate rather than merely an access interface\. We accordingly regard systems whose dominant progress\-bearing action–observation loop is mediated by terminal command execution, textual feedback, and stateful environment interaction as*terminal agents*\. Figure[1](https://arxiv.org/html/2608.20485#S1.F1)distinguishes this substrate from surface interfaces such as CLIs, IDEs, and web consoles\.
Figure 1:Terminal agents and related concepts\. An agent controller, terminal substrate, and mutable environment state form a progress\-bearing interaction loop\. Systems are in scope when command execution drives progress, textual feedback guides later actions, and stateful environment interaction is central; surface\-level CLI access and incidental command use remain adjacent\.This loop makes several general capabilities jointly inspectable in execution traces: planning appears as executable action sequences, memory is tested against persistent state, adaptation is grounded in external feedback, and verification occurs in the same substrate\. Terminal agents therefore provide a concrete setting for studying interactive, environment\-grounded intelligence\. Existing surveys, however, organize related evidence around general LLM agents\[[147](https://arxiv.org/html/2608.20485#bib.bib95),[156](https://arxiv.org/html/2608.20485#bib.bib102)\], software\-engineering agents\[[151](https://arxiv.org/html/2608.20485#bib.bib96),[146](https://arxiv.org/html/2608.20485#bib.bib121)\], GUI\- and browser\-based computer use\[[59](https://arxiv.org/html/2608.20485#bib.bib47),[127](https://arxiv.org/html/2608.20485#bib.bib86)\], vision\-language\-action models\[[173](https://arxiv.org/html/2608.20485#bib.bib191)\], or agent evaluation\[[172](https://arxiv.org/html/2608.20485#bib.bib115),[36](https://arxiv.org/html/2608.20485#bib.bib23)\]\. Terminal\-mediated behavior consequently remains dispersed across autonomy, repository repair, computer use, and evaluation, without a common synthesis of scope, system responsibilities, acquisition pathways, and measurement problems\.
The literature supporting such a synthesis remains uneven across domains\. Software engineering provides the densest empirical foundation because repositories, tests, and build systems make terminal\-mediated behavior directly observable through executable outcomes\[[67](https://arxiv.org/html/2608.20485#bib.bib51),[164](https://arxiv.org/html/2608.20485#bib.bib109),[101](https://arxiv.org/html/2608.20485#bib.bib68)\]\. Benchmarks, executable training environments, post\-training studies, and deployment reports increasingly treat terminal interaction as an explicit design and evaluation target rather than a passive backend\[[101](https://arxiv.org/html/2608.20485#bib.bib68),[89](https://arxiv.org/html/2608.20485#bib.bib60),[110](https://arxiv.org/html/2608.20485#bib.bib181),[11](https://arxiv.org/html/2608.20485#bib.bib15)\]\. Related terminal\-mediated interaction also appears in operations, data engineering, scientific workflows, cybersecurity, and cloud management, but evidence remains fragmented\[[64](https://arxiv.org/html/2608.20485#bib.bib189),[24](https://arxiv.org/html/2608.20485#bib.bib142),[68](https://arxiv.org/html/2608.20485#bib.bib52),[72](https://arxiv.org/html/2608.20485#bib.bib83),[92](https://arxiv.org/html/2608.20485#bib.bib5),[112](https://arxiv.org/html/2608.20485#bib.bib4)\]\. We therefore distinguish established findings from emerging practices; the Supplementary Material reports the review protocol, corpus construction, and evidence\-calibration procedure\.
Against this background, terminal\-mediated execution provides the organizing lens, while the seven\-dimensional terminal competence profile supplies a common analytical language\. The survey makes three connected contributions:
- •Operational scope and boundaries\.We provide a substrate\-centered characterization and workload\-level tests separating progress\-bearing terminal execution from surface\-level CLI access, incidental command use, and adjacent interaction substrates\.
- •Integrated analytical framework\.We connect system architecture, competence acquisition, and evaluation through a seven\-dimensional terminal competence profile for comparing system responsibilities, learning signals, and observable evidence\.
- •Cross\-cutting synthesis and diagnostics\.Applying this framework, we synthesize how the model, interface, harness, runtime, and environment jointly shape terminal\-agent behavior, how executable trajectories support competence acquisition, and how prevailing evaluations expose process behavior unevenly\. Bounded fixed\-condition diagnostics further illustrate benchmark\-dependent process exposure and limits of component attribution\.
Terminal AgentsOperational FoundationScope, boundaries, andcompetence profileSection IISystem ArchitectureAllocation ofsystem responsibilitiesSection IIICompetence AcquisitionLearning and adaptationfrom executable trajectoriesSection IVEvaluation and DiagnosticsProcess observabilityand attribution limitsSections V–VIResearch AgendaCross\-cuttingresearch prioritiesSection VIIExecution\-substrate lensWorkload\-levelboundary testsSeven\-dimensionalcompetence profileAdjacent\-systemcomparatorsOverlapping design emphasesArchitecturalresponsibility layersArchitectural patternsand responsibility allocationTrade\-offs, evidence landscape,and attributionCollection environmentsand data sourcesTrajectory construction, filtering,and failure dataLearning signals andadaptation approachesTransfer andcoverage gapsBenchmark design emphasesEvidence layersand process metricsProtocol validity andcompetence observabilityMeasurement andattribution diagnosticsCross\-domainterminal competenceFresh and replayableprocess evaluationRuntime governabilityand safetyControlled model–harnessattributionRepresentative LiteratureSWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]Agentless\[[157](https://arxiv.org/html/2608.20485#bib.bib186)\]OSWorld\[[159](https://arxiv.org/html/2608.20485#bib.bib105)\]SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]Meta\-Harness\[[79](https://arxiv.org/html/2608.20485#bib.bib55)\]TerminalTraj\[[155](https://arxiv.org/html/2608.20485#bib.bib185)\]CLI\-Gym\[[89](https://arxiv.org/html/2608.20485#bib.bib60)\]AgentHER\[[33](https://arxiv.org/html/2608.20485#bib.bib58)\]Terminal\-Bench\[[101](https://arxiv.org/html/2608.20485#bib.bib68)\]SetupBench\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\]LongCLI\-Bench\[[39](https://arxiv.org/html/2608.20485#bib.bib174)\]TerminalWorld\[[26](https://arxiv.org/html/2608.20485#bib.bib145)\]SWE\-rebench\[[9](https://arxiv.org/html/2608.20485#bib.bib13)\]BashArena\[[70](https://arxiv.org/html/2608.20485#bib.bib53)\]Meta\-Harness\[[79](https://arxiv.org/html/2608.20485#bib.bib55)\]
Figure 2:Analytical storyline of the survey\. The survey organizes the study of terminal agents around operational foundation, system architecture, competence acquisition, evaluation and diagnostics, and the resulting research agenda\. Expanded branches summarize the principal analytical content of Sections II–VII\. The column\-aligned reference strip provides representative literature anchors and boundary comparators rather than one\-to\-one mappings to individual subtopics\.Figure[2](https://arxiv.org/html/2608.20485#S1.F2)summarizes this progression\. Section[2](https://arxiv.org/html/2608.20485#S2)establishes the operational scope and terminal competence profile\. Sections[3](https://arxiv.org/html/2608.20485#S3)through[5](https://arxiv.org/html/2608.20485#S5)examine how terminal competence is distributed across system architecture, acquired from executable interaction, and observed through evaluation protocols\. Section[6](https://arxiv.org/html/2608.20485#S6)illustrates benchmark\-dependent process exposure and attribution limits through bounded fixed\-condition diagnostics\. Section[7](https://arxiv.org/html/2608.20485#S7)presents the research agenda, and Section[8](https://arxiv.org/html/2608.20485#S8)concludes\.
## 2Background and Scope of Terminal Agents
The execution\-substrate lens requires both an operational boundary and a capability framework\. This section establishes the workload\-level scope and competence profile used throughout the survey\.
### 2\.1Terminal as a Command\-Execution Agent Substrate
The terminal is a distinctive access layer for agentic execution because it combines compact textual interaction, compositional commands, persistent state, and broad operational reach\[[29](https://arxiv.org/html/2608.20485#bib.bib65)\]\. Through it, an agent can inspect directories, edit files, execute programs, install dependencies, launch processes, query logs, and access remote systems\[[11](https://arxiv.org/html/2608.20485#bib.bib15)\]\. Unlike GUI\-centered environments, where state is inferred largely from visual layouts or interface events, terminal\-mediated runtimes return stdout, stderr, exit codes, diffs, stack traces, logs, and process outputs in textual form, narrowing the gap between an LLM’s token interface and the environment it controls\[[164](https://arxiv.org/html/2608.20485#bib.bib109),[89](https://arxiv.org/html/2608.20485#bib.bib60)\]\.
Four connected properties make terminal\-mediated execution consequential: actions mutate persistent state; observations expose inspectable evidence; reruns and tests connect execution to verification; and commands, patches, and configuration edits form composable textual artifacts\[[29](https://arxiv.org/html/2608.20485#bib.bib65),[11](https://arxiv.org/html/2608.20485#bib.bib15),[166](https://arxiv.org/html/2608.20485#bib.bib108),[149](https://arxiv.org/html/2608.20485#bib.bib184),[7](https://arxiv.org/html/2608.20485#bib.bib10)\]\. Together, these properties support a reasoning, execution, observation, and verification loop rather than merely a tool\-call channel\. Figure[3](https://arxiv.org/html/2608.20485#S2.F3)connects this loop to the boundary tests and competence dimensions introduced below\.
Related terms describe different layers of this interaction\. A*terminal*is the textual input–output access layer; a*shell*interprets commands; and a*command\-line tool*is a textual program invoked through a shell\. A*command\-execution runtime*supplies the filesystem, processes, dependencies, permissions, and resource boundaries in which commands take effect\. A*harness*connects the model to this runtime by formatting observations, exposing actions, managing context, and enforcing execution policy\. Terminal\-mediated interaction therefore need not involve a visible terminal emulator or persistent pseudo\-terminal; instead, command execution and the resulting textual and state evidence must carry task progress\. A CLI\-exposed product remains outside this scope when the CLI merely forwards requests to an otherwise non\-terminal workflow\.
### 2\.2Substrate\-Based Scope and Boundaries
Building on this substrate view, we regard*terminal agents*as systems whose dominant progress\-bearing interaction is mediated by terminal command execution\. The scope is substrate\-based rather than surface\-based: a system may appear through a CLI, IDE, web interface, or platform runtime, but is in scope only when task progress depends on command execution, textual feedback, and stateful environment interaction\.
Occasional terminal use as an auxiliary tool is insufficient\[[111](https://arxiv.org/html/2608.20485#bib.bib75),[123](https://arxiv.org/html/2608.20485#bib.bib84)\], as is one\-shot code or patch generation without iterative execution and feedback\. Terminal agents instead rely on an execution\-grounded loop in which commands alter the environment and subsequent observations materially shape later decisions\[[170](https://arxiv.org/html/2608.20485#bib.bib187),[149](https://arxiv.org/html/2608.20485#bib.bib184),[164](https://arxiv.org/html/2608.20485#bib.bib109)\]\. We operationalize this distinction with three workload\-level tests:
- •Primary execution substrate: Terminal command execution is the workload’s main means of progress\.
- •Iterative command feedback: Outputs, errors, logs, diffs, return codes, or state changes materially shape later actions\.
- •Terminal dependence: Removing terminal access would materially change the workload’s core behavior\.
“Dominant” does not require every action to be a direct command; it means that task progress depends causally on terminal\-mediated state changes and feedback\. The tests apply to individual workloads rather than automatically to an entire product or platform\. A hybrid platform may therefore be in scope for repository\-repair or operations workloads when terminal execution drives state changes and subsequent decisions, but outside scope when another workload progresses mainly through a browser, GUI, or remote API\.
Under these tests, SWE\-agent is in scope when command feedback drives repository work\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\], as are terminal\-centric OpenHands workloads when terminal execution remains the principal locus of progress\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]\. The tests exclude static patch generators without iterative execution\[[157](https://arxiv.org/html/2608.20485#bib.bib186)\], GUI agents whose primary feedback is visual or DOM\-based\[[159](https://arxiv.org/html/2608.20485#bib.bib105),[188](https://arxiv.org/html/2608.20485#bib.bib124)\], and CLI\-packaged assistants that use the terminal only as an access surface\[[141](https://arxiv.org/html/2608.20485#bib.bib144)\]\. The same workload\-level tests guide review\-corpus screening\.
Figure 3:Operational scope of terminal agents, linking terminal\-mediated interaction to workload\-level boundary tests, scope categories, and competence dimensions\.
### 2\.3Terminal Competence as a Capability Profile
Having established the workload boundary, we characterize terminal competence through major functional roles rather than low\-level commands or domain skills\. GUI environments such as OSWorld and AndroidWorld provide useful boundary comparators: they include environment state and executable verification, but visual interaction remains the progress\-bearing substrate\[[159](https://arxiv.org/html/2608.20485#bib.bib105),[120](https://arxiv.org/html/2608.20485#bib.bib81)\]\. Across studies of systems, acquisition, and benchmarks, recurring behaviors and failures concern command selection, textual\-feedback interpretation, environment setup and repair, state persistence, verification, recovery, long\-horizon control, and execution governance\[[164](https://arxiv.org/html/2608.20485#bib.bib109),[150](https://arxiv.org/html/2608.20485#bib.bib97),[7](https://arxiv.org/html/2608.20485#bib.bib10),[74](https://arxiv.org/html/2608.20485#bib.bib177),[101](https://arxiv.org/html/2608.20485#bib.bib68),[31](https://arxiv.org/html/2608.20485#bib.bib30)\]\.
We group these responsibilities according to the object being controlled, the evidence required for the next decision, and the response needed when execution diverges from the task\. This grouping separates runtime construction from persistent state tracking, verification from post\-failure recovery, and authorization control from ordinary command execution\. The resulting seven dimensions capture system\-level responsibilities distributed across the model, interface, harness, runtime, and environment\. They recur in interaction traces, system designs, training pipelines, and benchmark protocols, as summarized in Figure[3](https://arxiv.org/html/2608.20485#S2.F3)\.
1. 1\.Command and action formulation: translating goals, constraints, and state into executable commands, scripts, CLI calls, file edits, build/test/run actions, or structured terminal operations\.
2. 2\.Feedback and artifact interpretation: extracting task\-relevant evidence from stdout, stderr, exit codes, logs, diffs, test outputs, stack traces, process signals, generated files, and workspace changes\.
3. 3\.Runtime and environment management: preparing, configuring, maintaining, and repairing dependencies, services, background processes, containers or virtual machines, remote machines, environment variables, resource limits, and related runtime constraints\.
4. 4\.State, task, and context tracking: maintaining environment state, task context, and interaction history across extended sessions, including filesystem and repository changes, packages, processes, configurations, prior commands, partial goals, verified facts, assumptions, and unresolved subproblems\.
5. 5\.Progress verification: designing and executing checks of intermediate validity, artifact trustworthiness, and completion conditions\.
6. 6\.Recovery and adaptation: diagnosing failures, revising hypotheses, replanning execution, mitigating harmful intermediate actions, and retrying from grounded evidence\.
7. 7\.Governance and side\-effect control: respecting permissions, sandboxes, approval checkpoints, safety policies, credentials, resource limits, and restrictions on deletion, network access, privilege escalation, or external\-system modification\.
These dimensions overlap with research on general tool use, software\-engineering agents, and computer use, including work on tool selection, feedback, memory, planning, and safety\[[111](https://arxiv.org/html/2608.20485#bib.bib75),[123](https://arxiv.org/html/2608.20485#bib.bib84),[164](https://arxiv.org/html/2608.20485#bib.bib109),[159](https://arxiv.org/html/2608.20485#bib.bib105),[59](https://arxiv.org/html/2608.20485#bib.bib47)\]\. The profile organizes these concerns around the trace\-observable obligations of terminal\-mediated execution: commands mutate persistent runtime state, textual artifacts carry evidence across steps, verification occurs in the same substrate, and broad operational reach makes authorization and reversibility integral to competence\. It thereby provides a common basis for comparing architecture, acquisition, and evaluation while preserving the distinctive demands of terminal\-mediated execution\.
The dimensions are analytically separable but operationally interdependent\. Recovery depends on feedback interpretation, state tracking, and verification; runtime management depends on action formulation; and long\-horizon persistence combines state tracking, verification, and recovery\. The following sections use this profile to examine how these responsibilities are distributed across system architecture, acquired through executable interaction, and exposed by evaluation protocols\. Section[6](https://arxiv.org/html/2608.20485#S6)then illustrates selected measurement consequences of this synthesis\.
## 3Terminal\-Agent System Design and Architectures
Terminal competence emerges from interactions among the model, interface, runtime, control mechanisms, harness, and environment\. This section traces how these components became explicit design objects, organizes their responsibilities into architectural layers and recurring patterns, and synthesizes the resulting trade\-offs and attribution consequences\.
### 3\.1Shifts in Design Emphasis toward Terminal\-Mediated Agency
Terminal\-agent design can be understood through four overlapping shifts in emphasis: tool\-augmented prompting, structured executable actions, terminal\-mediated agency as a first\-class target, and runtime\- or harness\-centered design\. Rather than discrete generations, these shifts reflect the increasing treatment of action interfaces, mutable workspaces, recovery and governance policies, and harness\-level context management as explicit design surfaces \(Figure[4](https://arxiv.org/html/2608.20485#S3.F4)\)\.
Figure 4:Shifts in design emphasis toward terminal\-mediated agency, from tool\-augmented prompting to runtime\- and harness\-centered systems, with interfaces, workspaces, recovery, governance, and context management becoming increasingly explicit\.Shift 1: Tool\-augmented prompting\.TALM\[[111](https://arxiv.org/html/2608.20485#bib.bib75)\], ReAct\[[170](https://arxiv.org/html/2608.20485#bib.bib187)\], Toolformer\[[123](https://arxiv.org/html/2608.20485#bib.bib84)\], and CodeAct\[[149](https://arxiv.org/html/2608.20485#bib.bib184)\]established observation\-conditioned tool use while treating execution primarily as an external action channel rather than centering persistent terminal\-mediated interaction\.
Shift 2: Structured executable actions\.SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]recasts repository navigation, editing, and execution as model\-legible agent–computer interface primitives, moving from raw command access toward mediated interaction\. OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]extends this direction through a persistent platform runtime coordinating terminal, code, browser, and tool surfaces\.
Shift 3: Terminal\-mediated agency as a first\-class target\.Terminal\-Bench\[[101](https://arxiv.org/html/2608.20485#bib.bib68)\], CLI\-Gym\[[89](https://arxiv.org/html/2608.20485#bib.bib60)\], Endless Terminals\[[43](https://arxiv.org/html/2608.20485#bib.bib34)\], and TermiGen\[[189](https://arxiv.org/html/2608.20485#bib.bib127)\]treat terminal interaction as a training or evaluation distribution rather than a by\-product of repository repair\. Deployment studies document permission\-gated command execution\[[15](https://arxiv.org/html/2608.20485#bib.bib17),[11](https://arxiv.org/html/2608.20485#bib.bib15)\], while Claude Code\[[6](https://arxiv.org/html/2608.20485#bib.bib138)\], Codex CLI\[[106](https://arxiv.org/html/2608.20485#bib.bib139)\], Aider\[[47](https://arxiv.org/html/2608.20485#bib.bib140)\], and Gemini CLI\[[52](https://arxiv.org/html/2608.20485#bib.bib141)\]exhibit recurring patterns of command execution, workspace inspection, textual feedback, approval, and sandboxing\.
Shift 4: Runtime\- and harness\-centered design\.Outer\-loop elements such as context packaging, workspace persistence, verification, and control flow are increasingly optimized directly\. Meta\-Harness\[[79](https://arxiv.org/html/2608.20485#bib.bib55)\], AutoHarness\[[93](https://arxiv.org/html/2608.20485#bib.bib63)\], and Agentic Harness Engineering\[[87](https://arxiv.org/html/2608.20485#bib.bib132)\]treat context compaction, approval rules, observation shaping, and observability\-guided harness evolution as performance\-shaping variables\. The OpenHands SDK\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]and extensible RL platforms\[[148](https://arxiv.org/html/2608.20485#bib.bib98)\]make runtime behavior programmable, while collaborative frameworks extend orchestration across multiple agents or execution entities\[[41](https://arxiv.org/html/2608.20485#bib.bib33)\]\. Interface, workspace, recovery, governance, and context management thus become explicit architectural components\.
### 3\.2Architectural Layers of Terminal Agents
To compare these designs, we organize recurring responsibilities into four layers: interface and observation; runtime and workspace; control, verification, recovery, and governance; and harness and context\. These layers locate where design choices enter the terminal\-mediated loop and how they support the competence dimensions introduced in Section[2](https://arxiv.org/html/2608.20485#S2)\(Figure[5](https://arxiv.org/html/2608.20485#S3.F5)\)\.
Layer 1: Interface and observation\.This layer specifies action units and feedback formats, ranging from direct commands to ACI primitives, runtime events, and workflow\-stage interfaces\. SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]replaces unconstrained repository interaction with model\-legible search, edit, and execution primitives\. Observation design determines whether listings, traces, logs, exit codes, diffs, files, and workspace changes are exposed at useful granularity\. In its reported setting, TACO reduces token use and improves accuracy through observational context compression\[[121](https://arxiv.org/html/2608.20485#bib.bib82)\]\. Observation filtering and reinsertion further shape the scale and formatting of available feedback\[[29](https://arxiv.org/html/2608.20485#bib.bib65)\], while robustness across interface conditions remains open\[[117](https://arxiv.org/html/2608.20485#bib.bib79)\]\. This layer primarily supports dimensions 1 and 2\.
Layer 2: Runtime and workspace\.Because commands mutate environment state, the runtime determines persistence, isolation, and whether later actions can build on earlier ones\. OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]coordinates terminal, code, browser, and tools over a shared workspace, whereas Terminal\-Bench\[[101](https://arxiv.org/html/2608.20485#bib.bib68)\]and CLI\-Gym\[[89](https://arxiv.org/html/2608.20485#bib.bib60)\]make controlled runtime environments central to evaluation\. Persistent workspaces enable cumulative progress but can also propagate erroneous intermediate state\. This layer governs dependencies, services, background processes, containers or virtual machines, resources, and workspace persistence, making it central to dimensions 3 and 4\. Automated Docker image construction\[[179](https://arxiv.org/html/2608.20485#bib.bib122)\]further shows that workspace preparation can itself become an optimization target\.
Layer 3: Control, verification, recovery, and governance\.Terminal access can modify filesystems, install packages, use credentials, and trigger external side effects\. This layer structures checks, error bounding, recovery, and oversight through testing, sandboxing, rollback, approval, and decomposition\. STRATUS uses role\-delegated planning, execution, and review\[[23](https://arxiv.org/html/2608.20485#bib.bib22)\]; related principles appear in incident response\[[13](https://arxiv.org/html/2608.20485#bib.bib128),[88](https://arxiv.org/html/2608.20485#bib.bib3)\]and configuration\-drift detection\[[1](https://arxiv.org/html/2608.20485#bib.bib1)\]\. AgentClick introduces skill\-based checkpoints\[[191](https://arxiv.org/html/2608.20485#bib.bib188)\]\. OS\-level resource management\[[99](https://arxiv.org/html/2608.20485#bib.bib180),[125](https://arxiv.org/html/2608.20485#bib.bib85)\], verified deployment of generated Linux scheduling policies\[[184](https://arxiv.org/html/2608.20485#bib.bib6)\], context\-space access control\[[51](https://arxiv.org/html/2608.20485#bib.bib40)\], and reliable state management\[[142](https://arxiv.org/html/2608.20485#bib.bib93)\]further make permission, accountability, refusal, and side\-effect control architectural concerns\. This layer primarily supports dimensions 5 to 7\.
Layer 4: Harness and context\.The harness packages history, formats prompts, manages context windows, and sequences model calls, determining what remains available across turns\. Version\-control\-inspired methods offer alternatives to raw truncation\[[154](https://arxiv.org/html/2608.20485#bib.bib101)\], and controlled comparisons show that context selection affects repository\-level generation\[[77](https://arxiv.org/html/2608.20485#bib.bib45)\]\. Meta\-Harness\[[79](https://arxiv.org/html/2608.20485#bib.bib55)\]and AutoHarness\[[93](https://arxiv.org/html/2608.20485#bib.bib63)\]report performance changes from outer\-loop optimization, while multi\-agent and asynchronous strategies broaden the orchestration space\[[12](https://arxiv.org/html/2608.20485#bib.bib16),[49](https://arxiv.org/html/2608.20485#bib.bib38)\]\. This layer primarily supports dimension 4 and long\-horizon persistence by preserving trajectory information, verified facts, unresolved subproblems, and prior actions\.
Together, these layers show that terminal competence is distributed across the model, interface, runtime, control mechanisms, and harness\. Long\-horizon persistence therefore depends jointly on state and context tracking, runtime persistence, verification, and recovery\.
A promising architectural direction is tighter state\-aware coordination across these layers\. Future systems could condition context management, verification, recovery, and governance on shared execution state and task progress, allowing observation, context allocation, verification, recovery, and execution policy to adapt coherently over long trajectories\.
Figure 5:Layered architecture of terminal\-agent systems\. The central interaction loop is shaped by interface, runtime, verification, recovery, governance, and harness\-level context mechanisms\.
### 3\.3Recurring Architectural Patterns and Responsibility Allocation
The layers above locate architectural responsibilities, while recurring patterns describe how systems allocate them\. These patterns are complementary rather than mutually exclusive and differ in action mediation, runtime persistence, control organization, context management, and rollout infrastructure\. Direct\-command access exposes broad command spaces behind permission or sandbox controls\[[6](https://arxiv.org/html/2608.20485#bib.bib138),[106](https://arxiv.org/html/2608.20485#bib.bib139),[47](https://arxiv.org/html/2608.20485#bib.bib140),[52](https://arxiv.org/html/2608.20485#bib.bib141)\], whereas ACI mediation constrains search, edit, and execution through model\-legible primitives\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]\. Platform runtimes coordinate persistent workspaces and multiple interaction surfaces\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]; role\-structured systems separate planning, execution, and review\[[23](https://arxiv.org/html/2608.20485#bib.bib22),[113](https://arxiv.org/html/2608.20485#bib.bib76)\]; and scaffold\-centric systems manage context, observations, and control flow\[[79](https://arxiv.org/html/2608.20485#bib.bib55),[93](https://arxiv.org/html/2608.20485#bib.bib63)\]\. Runtime\-memory methods compress or retrieve long\-session state\[[136](https://arxiv.org/html/2608.20485#bib.bib90),[121](https://arxiv.org/html/2608.20485#bib.bib82)\], while terminal\-native rollout environments provide scalable infrastructure for training and trajectory generation\[[89](https://arxiv.org/html/2608.20485#bib.bib60),[43](https://arxiv.org/html/2608.20485#bib.bib34),[189](https://arxiv.org/html/2608.20485#bib.bib127),[63](https://arxiv.org/html/2608.20485#bib.bib176)\]\.
These patterns allocate responsibility differently\. Direct access favors expressiveness but increases observation noise and rollback difficulty\. Interface mediation favors reliability but may reduce cross\-task generality\. Persistent runtimes broaden the task surface while increasing dependence on the surrounding system, whereas context optimization supports long\-horizon coherence while complicating attribution\. Architectures are therefore better compared by how they allocate responsibility across the model, interface, runtime, control mechanisms, and harness than by assignment to a single family\.
### 3\.4Synthesis: Architectural Trade\-offs and Attribution Consequences
These alternative responsibility allocations create three recurring architectural tensions and one attribution consequence\.
Expressiveness vs\. recoverability\.Raw or weakly mediated access expands the action space but produces noisier trajectories and harder rollback; ACI and scaffold constraints improve reliability yet may restrict other task families\. This trade\-off chiefly concerns dimensions 1 and 6, with verification supplying evidence for recovery\.
Generality vs\. task discipline\.Platform runtimes support heterogeneous terminal, code, and browser work but may provide less structured feedback\. Role\- or workflow\-constrained systems strengthen local control at the cost of broader applicability\. This tension primarily affects dimensions 3 and 4\.
Automation vs\. inspectability\.Deployable agents must keep commands, outputs, state changes, and interventions legible\. Permission gates, approvals, and sandboxes improve auditability but add latency and interrupt autonomous execution\. This tension centers on dimension 7, which remains among the least systematically addressed\.
Model capability vs\. harness contribution\.A further consequence is attribution difficulty: gains under optimized harnesses may arise from context management, observation shaping, retries, permission policy, or injected procedural knowledge rather than stronger model reasoning\. Controlled skill injection produces task\-dependent gains and can add substantial token overhead without improving pass rate\[[56](https://arxiv.org/html/2608.20485#bib.bib46)\]\. Agentless likewise reports that static pipelines can rival interactive agents on some repository\-repair tasks\[[157](https://arxiv.org/html/2608.20485#bib.bib186)\]\. The complete model–harness–runtime configuration therefore remains relevant to system\-level comparison, while component\-level attribution requires separating these contributions\.
### 3\.5Evidence Landscape and Boundary Comparators
The architectural literature spans established agent systems, controlled harness studies, deployment reports, and emerging training and operational settings\. SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]and OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]provide broadly used system designs and evaluation settings, whereas Meta\-Harness\[[79](https://arxiv.org/html/2608.20485#bib.bib55)\]and AutoHarness\[[93](https://arxiv.org/html/2608.20485#bib.bib63)\]examine harness optimization within individual studies\. Deployment studies and commercial direct\-command agents document recurring interface, permission, sandbox, and workflow patterns\[[15](https://arxiv.org/html/2608.20485#bib.bib17),[11](https://arxiv.org/html/2608.20485#bib.bib15),[6](https://arxiv.org/html/2608.20485#bib.bib138),[106](https://arxiv.org/html/2608.20485#bib.bib139),[47](https://arxiv.org/html/2608.20485#bib.bib140),[52](https://arxiv.org/html/2608.20485#bib.bib141)\]\. Training environments demonstrate scalable rollout generation, while evidence from live workflows and network operations is still emerging\[[46](https://arxiv.org/html/2608.20485#bib.bib130),[102](https://arxiv.org/html/2608.20485#bib.bib129)\]\.
Adjacent systems further clarify the architectural boundary\. Agentless\[[157](https://arxiv.org/html/2608.20485#bib.bib186)\]represents static execution and verification pipelines, CGM\[[137](https://arxiv.org/html/2608.20485#bib.bib91)\]uses graph\-structured control, OSWorld\[[159](https://arxiv.org/html/2608.20485#bib.bib105)\]grounds progress primarily in GUI interaction, and CLI\-packaged assistants\[[141](https://arxiv.org/html/2608.20485#bib.bib144)\]use the terminal mainly as an access surface\. These comparators distinguish progress\-bearing terminal interaction from systems that share only selected architectural components\.
Architectural choices determine executable actions, recorded observations and state changes, available interventions, and the failures entering a trajectory\. They thereby shape both the supervision that acquisition pipelines can construct and the evidence that evaluation protocols can observe\. Section[4](https://arxiv.org/html/2608.20485#S4)examines how executable interactions become learning signals, while Section[5](https://arxiv.org/html/2608.20485#S5)examines how the resulting behavior becomes measurable evidence\.
## 4Terminal Competence Acquisition and Adaptation
Building on the architectural account in Section[3](https://arxiv.org/html/2608.20485#S3), competence acquisition concerns how executable interactions become learning and adaptation signals\. Here, acquisition spans parameter learning and runtime adaptation\. Training data inherit the action interface, harness, runtime, and governance conditions under which they are collected\. Because commands mutate state, expose feedback and failures, and require continuation or repair decisions\[[114](https://arxiv.org/html/2608.20485#bib.bib77),[43](https://arxiv.org/html/2608.20485#bib.bib34),[155](https://arxiv.org/html/2608.20485#bib.bib185),[69](https://arxiv.org/html/2608.20485#bib.bib114)\], the relevant unit is a stateful trajectory containing actions, observations, state changes, verification, and recovery rather than an isolated prompt–response pair\. This section follows the acquisition path from collection environments through trajectory construction and learning to runtime adaptation, then summarizes transfer and coverage gaps\.
Figure[6](https://arxiv.org/html/2608.20485#S4.F6)connects acquisition sources, interaction and recovery traces, runtime and governance mechanisms, learning and adaptation approaches, and the resulting competence profile\.
Figure 6:Terminal competence acquisition ecosystem\. Data sources, interaction traces, runtime and governance mechanisms, and learning and adaptation approaches jointly shape the multidimensional competence profile\.### 4\.1Data Sources and Collection Settings for Terminal Competence
Collection environments determine which actions, observations, state changes, and failures become available as learning evidence\. Terminal\-native rollouts expose command selection, output interpretation, and recovery\[[155](https://arxiv.org/html/2608.20485#bib.bib185),[43](https://arxiv.org/html/2608.20485#bib.bib34),[89](https://arxiv.org/html/2608.20485#bib.bib60),[63](https://arxiv.org/html/2608.20485#bib.bib176)\]\. Executable repository environments add navigation, editing, setup, and test\-grounded verification, while coupling these behaviors to repository\-specific reasoning\[[110](https://arxiv.org/html/2608.20485#bib.bib181),[37](https://arxiv.org/html/2608.20485#bib.bib31),[9](https://arxiv.org/html/2608.20485#bib.bib13),[86](https://arxiv.org/html/2608.20485#bib.bib59),[55](https://arxiv.org/html/2608.20485#bib.bib43)\]\. Machine\-learning engineering environments extend the loop to iterative experimentation and component refinement\[[116](https://arxiv.org/html/2608.20485#bib.bib78),[103](https://arxiv.org/html/2608.20485#bib.bib41)\]\. Synthetic environments target sparse subskills but risk overfitting to generated patterns\[[189](https://arxiv.org/html/2608.20485#bib.bib127),[165](https://arxiv.org/html/2608.20485#bib.bib111),[100](https://arxiv.org/html/2608.20485#bib.bib162)\]\. Recent synthesis pipelines further scale executable data generation by jointly constructing instructions, environments, reference solutions, and verifiers through taxonomy\-guided, evolutionary, or recursive procedures\[[60](https://arxiv.org/html/2608.20485#bib.bib150),[126](https://arxiv.org/html/2608.20485#bib.bib167),[84](https://arxiv.org/html/2608.20485#bib.bib157)\]\. Failure\-centered corpora instead preserve diagnosis, rollback, and recovery that successful traces often omit\[[69](https://arxiv.org/html/2608.20485#bib.bib114),[33](https://arxiv.org/html/2608.20485#bib.bib58),[178](https://arxiv.org/html/2608.20485#bib.bib8)\]\.
Environment structure also determines which competence dimensions are exposed for supervision\. Short rollouts primarily expose action formulation and feedback interpretation; setup tasks add runtime management; long trajectories stress state and context tracking; and executable checks support verification\. Retained failures expose diagnosis and recovery, while governance requires settings in which authorization, containment, and side effects are visible\.
### 4\.2Trajectory Construction, Filtering, and Failure Data
A rollout becomes usable training data through selection, validation, filtering, and relabeling\. These operations shape both data quality and the behaviors preserved for learning\[[155](https://arxiv.org/html/2608.20485#bib.bib185),[162](https://arxiv.org/html/2608.20485#bib.bib25),[33](https://arxiv.org/html/2608.20485#bib.bib58),[114](https://arxiv.org/html/2608.20485#bib.bib77),[189](https://arxiv.org/html/2608.20485#bib.bib127),[69](https://arxiv.org/html/2608.20485#bib.bib114)\]\. Dockerized generation integrates task adaptation, synthetic generation, rollout collection, filtering, and decontamination\[[155](https://arxiv.org/html/2608.20485#bib.bib185)\]\. CLEANER filters trajectories for reinforcement learning\[[162](https://arxiv.org/html/2608.20485#bib.bib25)\], while AgentHER validates and relabels them through hindsight replay\[[33](https://arxiv.org/html/2608.20485#bib.bib58)\]\. Trajectory utility also depends on preserved interaction structure: TerminalLego reports stronger training signals from environment\-grounded inspect–act–verify trajectories than from teacher success alone in its setting\[[167](https://arxiv.org/html/2608.20485#bib.bib148)\]\. Other mechanisms alter what remains visible during learning: progressive code masking varies access to prior code\[[71](https://arxiv.org/html/2608.20485#bib.bib11)\]; Nemotron\-Terminal constructs terminal\-oriented training data\[[114](https://arxiv.org/html/2608.20485#bib.bib77)\]; Live\-SWE\-agent studies online self\-evolution\[[158](https://arxiv.org/html/2608.20485#bib.bib103)\]; and long\-context multi\-turn reinforcement learning retains extended interactions\[[50](https://arxiv.org/html/2608.20485#bib.bib39)\]\.
Recoverable failures remain underrepresented\. Failed installations, version conflicts, invalid environment assumptions, rollback decisions, and dead\-end repairs expose recovery more directly than final successful commands\[[69](https://arxiv.org/html/2608.20485#bib.bib114),[33](https://arxiv.org/html/2608.20485#bib.bib58),[178](https://arxiv.org/html/2608.20485#bib.bib8),[180](https://arxiv.org/html/2608.20485#bib.bib134),[190](https://arxiv.org/html/2608.20485#bib.bib7)\]\. Yet successful\-trace filtering can remove the interactions needed to learn diagnosis and adaptation\[[162](https://arxiv.org/html/2608.20485#bib.bib25),[33](https://arxiv.org/html/2608.20485#bib.bib58),[190](https://arxiv.org/html/2608.20485#bib.bib7)\]\. Retaining misdiagnosis, rollback, and repair therefore provides direct supervision for recovery under observed failure modes\[[69](https://arxiv.org/html/2608.20485#bib.bib114),[33](https://arxiv.org/html/2608.20485#bib.bib58),[180](https://arxiv.org/html/2608.20485#bib.bib134)\]\.
### 4\.3Learning Signals and Adaptation Approaches
Constructed trajectories support complementary levers at three levels: data construction, parameter optimization, and runtime adaptation\. These levers can be combined within a single acquisition pipeline\. Table[1](https://arxiv.org/html/2608.20485#S4.T1)compares their primary signal, capability target, blind spot, and principal risk\.
Table 1:Complementary approaches to terminal competence acquisition and adaptation\.ApproachPrimary signalCapability targetBlind spotKey riskSFT on successful tracesSuccessful interaction tracesCommon commands, repository navigation, and standard workflowsRecovery, rollback, and diagnosis after wrong assumptionsClean\-trace overfit; absent failure pathsRL with environment rewardsExecutable success, verifier reward, and test outcomeOutcome\-driven exploration and completionProcess quality and safe intermediate behaviorSparse\-reward hacking; benchmark overfitProcess\-aware / verifier\-guidedStep judgments and verifier rankingsDiagnosis, action selection, and recovery decisionsOpen\-domain transferVerifier bias and shifted failuresSynthetic task generationGenerated tasks with controlled verificationRare commands, setup patterns, and targeted subskillsLive realism and environment driftSynthetic heuristics; weak transferFailure\-conditioned trainingFailed or repaired traces; hindsight relabelingRecovery, diagnosis, and early failure detectionGeneralization across inconsistent failure taxonomiesNoisy labels; repair\-loop overfitRuntime memory / context adaptationSession history, summaries, and retrieved experienceLong\-horizon persistence and workflow reuseCorrection under incorrect memoryCompression loss; memory contaminationAt the parameter\-optimization level,SFT on successful tracesteaches common commands, repository workflows, and setup patterns\. Nemotron\-Terminal\[[114](https://arxiv.org/html/2608.20485#bib.bib77)\]uses terminal\-oriented data engineering, while SWE\-Gym\[[110](https://arxiv.org/html/2608.20485#bib.bib181)\]and SWE\-Dev\[[37](https://arxiv.org/html/2608.20485#bib.bib31)\]provide executable software trajectories\. Their emphasis on successful interactions, however, offers limited supervision for diagnosis, rollback, and recovery after incorrect assumptions\[[169](https://arxiv.org/html/2608.20485#bib.bib110),[176](https://arxiv.org/html/2608.20485#bib.bib120)\]\.RL with executable rewardsaligns behavior with task completion through environments and verifiers, as in Endless Terminals\[[43](https://arxiv.org/html/2608.20485#bib.bib34)\], ECHO\[[128](https://arxiv.org/html/2608.20485#bib.bib163)\], SWE\-Master\[[131](https://arxiv.org/html/2608.20485#bib.bib89)\], SWE\-Gym\[[110](https://arxiv.org/html/2608.20485#bib.bib181)\], and Tmax\[[62](https://arxiv.org/html/2608.20485#bib.bib151)\], but sparse rewards may reinforce brittle or unsafe behavior\.Process\-aware and verifier\-guided optimizationinstead supervises intermediate decisions through judgments, rankings, or hindsight validation\[[153](https://arxiv.org/html/2608.20485#bib.bib99),[33](https://arxiv.org/html/2608.20485#bib.bib58)\]\. AgentHER\[[33](https://arxiv.org/html/2608.20485#bib.bib58)\]is especially relevant because terminal failures often appear early through stderr, logs, failed tests, or inconsistent state\.
Data\-oriented and runtime approaches complement parameter optimization\.Synthetic task generationexpands coverage of rare commands, dependency search, repair, localization, and environment construction through verifiable tasks\[[189](https://arxiv.org/html/2608.20485#bib.bib127),[165](https://arxiv.org/html/2608.20485#bib.bib111),[100](https://arxiv.org/html/2608.20485#bib.bib162)\], but may teach synthetic regularities rather than reusable competence\.Failure\-conditioned traininguses repaired or failed traces from TRACE\[[69](https://arxiv.org/html/2608.20485#bib.bib114)\], AgentHER\[[33](https://arxiv.org/html/2608.20485#bib.bib58)\], and AgentForesight\[[178](https://arxiv.org/html/2608.20485#bib.bib8)\]for diagnosis, early failure prediction, hindsight relabeling, and recovery; its main obstacles are noisy labels, inconsistent taxonomies, and repair\-loop overfit\.Runtime memory and context adaptationextends acquisition into online behavior: Memento\[[187](https://arxiv.org/html/2608.20485#bib.bib125)\], Context\-Folding\[[136](https://arxiv.org/html/2608.20485#bib.bib90)\], and TACO\[[121](https://arxiv.org/html/2608.20485#bib.bib82)\]retain history, summaries, or compressed context across long interactions, while risking persistence of incorrect assumptions\.
These approaches supervise different parts of an interaction history: successful traces teach common workflows, executable rewards favor completion, verifiers and hindsight expose intermediate decisions, synthetic tasks broaden coverage, failure\-conditioned data supports recovery, and runtime memory sustains long\-horizon behavior\. Their shared challenge is to combine these signals without overfitting to benchmark tasks, generated environments, or harness\-specific rewards\.
### 4\.4Emerging Practices and Remaining Gaps
Recent systems increasingly treat acquisition as an end\-to\-end pipeline in which sourcing, rollout generation, filtering, replay, and curriculum design are distinct levers\[[114](https://arxiv.org/html/2608.20485#bib.bib77),[110](https://arxiv.org/html/2608.20485#bib.bib181),[155](https://arxiv.org/html/2608.20485#bib.bib185)\]\. Such pipelines increasingly retain setup attempts, outputs, state changes, failed commands, verifier feedback, and recovery decisions, yet current corpora remain concentrated on successful repository\-centric traces, with limited coverage of failed setup, long repair loops, environment drift, and unsafe intermediate actions\[[69](https://arxiv.org/html/2608.20485#bib.bib114),[162](https://arxiv.org/html/2608.20485#bib.bib25),[33](https://arxiv.org/html/2608.20485#bib.bib58)\]\.
Transfer remains a central unresolved issue\. Repository\-repair training may produce task\-specific heuristics rather than general terminal competence\[[157](https://arxiv.org/html/2608.20485#bib.bib186)\], motivating work on action, observation, and reward sequences that encode reusable skills across environments\[[81](https://arxiv.org/html/2608.20485#bib.bib178)\]\. Work on externalization\[[186](https://arxiv.org/html/2608.20485#bib.bib126)\]and semantics\-aware repair\[[108](https://arxiv.org/html/2608.20485#bib.bib73)\]further suggests that transfer depends on how explicitly trajectories preserve intermediate reasoning and execution evidence\. Cross\-domain workflows, environment drift, failed setup, unsafe intermediate actions, and long repair loops remain weakly represented\.
These gaps make evaluation evidence essential for determining which aspects of terminal competence are actually acquired\. Final success alone does not reveal whether an agent preserved state, recovered from failure, verified completion, or respected execution constraints\. Section[5](https://arxiv.org/html/2608.20485#S5)therefore examines which benchmark families and evidence layers make these behaviors observable, helping distinguish transferable terminal competence from adaptation to particular repositories, harnesses, environments, or reward channels\.
## 5Benchmarks, Metrics, and Evaluation
Evaluation determines which aspects of the architectures and acquisition pipelines discussed above become observable and which remain hidden behind final outcomes\. Terminal\-agent evaluation spans software engineering, tool use, and environment\-coupled interaction, making it important to distinguish*terminal competence*from*repository repair competence*\[[67](https://arxiv.org/html/2608.20485#bib.bib51),[164](https://arxiv.org/html/2608.20485#bib.bib109),[7](https://arxiv.org/html/2608.20485#bib.bib10),[101](https://arxiv.org/html/2608.20485#bib.bib68)\]\. The two overlap but are not equivalent\.
We organize the literature along four analytical axes: benchmark design emphases characterize task settings and evaluation targets; evidence layers specify what is recorded; the coverage map estimates which competence dimensions become observable; and protocol validity determines how scores can be interpreted across evaluation settings\. Together, these axes separate what a benchmark asks agents to do, what evidence it records, and what conclusions its protocol supports\.
Table[2](https://arxiv.org/html/2608.20485#S5.T2)groups non\-exclusive benchmark design emphases by their primary evaluation focus and principal blind spot\. They are descriptive rather than ranked because their tasks, scoring signals, and execution infrastructure differ\.
Table 2:Benchmark design emphases by primary evaluation focus and principal blind spot\.EmphasisPrimary evaluation focusRepresentative examplesPrincipal blind spotRepository repairIssue resolution under testsSWE\-benchandSWE\-PolyBench\[[67](https://arxiv.org/html/2608.20485#bib.bib51),[119](https://arxiv.org/html/2608.20485#bib.bib80)\]Conflates repository repair with terminal competence; command use is not directly measuredCLI/terminal\-centeredCommands, output interpretation, and CLI workflowsTerminal\-Bench,TerminalWorld, andLongCLI\-Bench\[[101](https://arxiv.org/html/2608.20485#bib.bib68),[26](https://arxiv.org/html/2608.20485#bib.bib145),[39](https://arxiv.org/html/2608.20485#bib.bib174)\]Mixes terminal\-native and repository\-mediated tasks, leaving transfer unclearSetupEnvironment setup and dependency resolutionSetupBench\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\]Often isolated from end\-to\-end workflowsProcessIntermediate behavior and scaffold complianceOctoBench,ProcBench, andAppWorld\[[31](https://arxiv.org/html/2608.20485#bib.bib30),[57](https://arxiv.org/html/2608.20485#bib.bib135),[143](https://arxiv.org/html/2608.20485#bib.bib94)\]Lacks standardized process scoringLong horizonPersistence, drift, and temporal coherenceSWE\-Bench Pro,LoCoEval, andLifelongAgentBench\[[30](https://arxiv.org/html/2608.20485#bib.bib27),[91](https://arxiv.org/html/2608.20485#bib.bib62),[182](https://arxiv.org/html/2608.20485#bib.bib123)\]Costly and difficult to reproduce at scaleSafety/governancePermissions, containment, and privileged commandsBashArena,ClawSafety, andAgentHazard\[[70](https://arxiv.org/html/2608.20485#bib.bib53),[152](https://arxiv.org/html/2608.20485#bib.bib100),[40](https://arxiv.org/html/2608.20485#bib.bib32)\]Immature protocols; task success may conceal harmful actionsProductionDistribution match and deployment realismProdCodeBench\[[65](https://arxiv.org/html/2608.20485#bib.bib50)\]Limited public access and domain coverageProduction\-derived evaluation cuts across these emphases by improving distribution match without isolating a single terminal skill\[[65](https://arxiv.org/html/2608.20485#bib.bib50),[64](https://arxiv.org/html/2608.20485#bib.bib189)\]\. Repository\-scale deployment studies add evidence about adoption, quality, security\-relevant changes, and task\-conditioned acceptance\[[122](https://arxiv.org/html/2608.20485#bib.bib175),[3](https://arxiv.org/html/2608.20485#bib.bib169),[129](https://arxiv.org/html/2608.20485#bib.bib87),[115](https://arxiv.org/html/2608.20485#bib.bib182)\], but provide limited evidence about individual dimensions of terminal competence\.
### 5\.1From Repository\-Level Evaluation to Terminal\-Native Benchmarks
SWE\-bench\[[67](https://arxiv.org/html/2608.20485#bib.bib51)\]established repository issue resolution with executable verification as a dominant evaluation paradigm\. Successors expanded language coverage\[[119](https://arxiv.org/html/2608.20485#bib.bib80),[175](https://arxiv.org/html/2608.20485#bib.bib117)\], project scale\[[86](https://arxiv.org/html/2608.20485#bib.bib59)\], temporal validity\[[134](https://arxiv.org/html/2608.20485#bib.bib104),[9](https://arxiv.org/html/2608.20485#bib.bib13)\], production realism\[[65](https://arxiv.org/html/2608.20485#bib.bib50)\], and multilingual agentic evaluation\[[2](https://arxiv.org/html/2608.20485#bib.bib168)\]\. SWE\-Hub integrates environment construction, task synthesis, and executable validation in a scalable production pipeline\[[177](https://arxiv.org/html/2608.20485#bib.bib119)\]\. Claw\-SWE\-Bench adds multilingual repository tasks and a Lite subset suited to matched system comparisons\[[183](https://arxiv.org/html/2608.20485#bib.bib147)\]\. Section[6](https://arxiv.org/html/2608.20485#S6)pairs Claw\-SWE\-Bench Lite with SWE\-bench Lite: the former broadens repository and language diversity, while the latter provides a widely used compatibility anchor\. Both remain repository\-repair protocols with indirect coverage of broader terminal competence\.
Terminal\-native and CLI\-centered benchmarks instead make command\-mediated interaction part of the task definition\[[101](https://arxiv.org/html/2608.20485#bib.bib68),[26](https://arxiv.org/html/2608.20485#bib.bib145),[39](https://arxiv.org/html/2608.20485#bib.bib174),[105](https://arxiv.org/html/2608.20485#bib.bib71),[34](https://arxiv.org/html/2608.20485#bib.bib133),[76](https://arxiv.org/html/2608.20485#bib.bib136)\]\. They directly expose command use, output interpretation, state inspection, and workflow persistence\. TUA\-Bench extends this design beyond technical and programming\-centric workflows to routine digital activities and scientific and engineering work conducted through a terminal\[[21](https://arxiv.org/html/2608.20485#bib.bib152)\]\. CLI\-Gym\[[89](https://arxiv.org/html/2608.20485#bib.bib60)\]supports both training and evaluation, while InterCode\[[166](https://arxiv.org/html/2608.20485#bib.bib108)\]anticipated this direction through interactive coding with execution feedback\. This group remains heterogeneous: tasks range from terminal\-native workflows to repository\-mediated work, WildClawBench reports an 18\-point harness\-conditioned gap\[[34](https://arxiv.org/html/2608.20485#bib.bib133)\], and ClawForge identifies proactive state inspection as a strong discriminator in its setting\[[76](https://arxiv.org/html/2608.20485#bib.bib136)\]\. Transfer across these task types remains insufficiently established\.
### 5\.2Process\-Aware, Environment\-Aware, and Long\-Horizon Evaluation
Evaluation is expanding from measuring only*what*agents produce to examining*how*they interact with mutable environments\.Process\-aware evaluationseparates final success from intermediate quality and scaffold compliance\. OctoBench\[[31](https://arxiv.org/html/2608.20485#bib.bib30)\]and debug\-gym\[[174](https://arxiv.org/html/2608.20485#bib.bib116)\]add step\-level assessment and defect ontologies; ProcBench\[[57](https://arxiv.org/html/2608.20485#bib.bib135)\]and AgentEval\[[54](https://arxiv.org/html/2608.20485#bib.bib44)\]assess error propagation and control preservation; and ToolSandbox\[[94](https://arxiv.org/html/2608.20485#bib.bib64)\], AppWorld\[[143](https://arxiv.org/html/2608.20485#bib.bib94)\], and ASTRA\-Bench\[[160](https://arxiv.org/html/2608.20485#bib.bib106)\]extend stateful tool use and action planning beyond repository repair\. These efforts expose intermediate behavior, although no shared process\-scoring standard has emerged\.
Environment\-aware evaluationtreats setup and dependency resolution as task components rather than pre\-task overhead\. SetupBench isolates environment bootstrap\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\], while process\-level configuration studies identify setup defects that binary success misses\[[74](https://arxiv.org/html/2608.20485#bib.bib177)\]\. Environment setup thereby becomes an explicit component of terminal competence and end\-to\-end evaluation\.
Long\-horizon evaluationreveals failures hidden by short tasks\. In its reported setting, SWE\-EVO\[[139](https://arxiv.org/html/2608.20485#bib.bib92)\]reports a drop from 65–73% on SWE\-Bench Verified to 21–25% on multi\-file evolution\. SWE\-Bench Pro\[[30](https://arxiv.org/html/2608.20485#bib.bib27)\], LoCoEval\[[91](https://arxiv.org/html/2608.20485#bib.bib62)\], SlopCodeBench\[[107](https://arxiv.org/html/2608.20485#bib.bib72)\], Spec Emerges\[[163](https://arxiv.org/html/2608.20485#bib.bib107)\], ProjDevBench\[[95](https://arxiv.org/html/2608.20485#bib.bib118)\], and RepoMod\-Bench\[[83](https://arxiv.org/html/2608.20485#bib.bib179)\]similarly expose degradation under repeated editing\. Long\-Horizon\-Terminal\-Bench adds dense subtask rewards and partial\-credit grading to terminal workflows that require sustained execution\[[85](https://arxiv.org/html/2608.20485#bib.bib154)\]\. NL2Repo\-Bench\[[32](https://arxiv.org/html/2608.20485#bib.bib28)\]targets repository generation, AgencyBench\[[81](https://arxiv.org/html/2608.20485#bib.bib178)\]stresses 1M\-token contexts, LifelongAgentBench\[[182](https://arxiv.org/html/2608.20485#bib.bib123)\]evaluates lifelong behavior, and WildClawBench\[[34](https://arxiv.org/html/2608.20485#bib.bib133)\]covers deployment scenarios\. These settings extend evaluation from local completion to sustained state coherence, progress tracking, and alignment over time\.
### 5\.3Protocol Validity and Harness Effects
Evaluation scores reflect a coupled configuration of the model, task set, and protocol choices governing observation, execution, retries, memory, and runtime conditions\[[164](https://arxiv.org/html/2608.20485#bib.bib109),[79](https://arxiv.org/html/2608.20485#bib.bib55),[93](https://arxiv.org/html/2608.20485#bib.bib63),[48](https://arxiv.org/html/2608.20485#bib.bib37)\]\. Three factors are especially important for interpreting comparisons across evaluation settings\.
Contamination and temporal validity\.Static task pools are vulnerable to memorization, prompt leakage, stale distributions, and public exposure of issues, patches, tests, or discussions\. AgentBench\[[90](https://arxiv.org/html/2608.20485#bib.bib61)\]established broader agent\-evaluation protocols, while mutation and rebenchmarking improve freshness\[[45](https://arxiv.org/html/2608.20485#bib.bib36),[9](https://arxiv.org/html/2608.20485#bib.bib13),[10](https://arxiv.org/html/2608.20485#bib.bib14)\]\. LiveSQLBench\[[138](https://arxiv.org/html/2608.20485#bib.bib143)\]extends dynamic evaluation to database tasks, ACE\-Bench\[[168](https://arxiv.org/html/2608.20485#bib.bib112)\]varies difficulty and horizon, and contamination detection distinguishes recall from reasoning\[[133](https://arxiv.org/html/2608.20485#bib.bib88)\]\. Task construction is equally important: an audit reports that 16% of tasks across five terminal\-agent benchmarks are reward\-hackable\[[185](https://arxiv.org/html/2608.20485#bib.bib158)\]\. These developments make temporal freshness, contamination control, task regeneration, and resistance to reward hacking central to protocol validity\.
Harness\-mediated variance\.The harness is part of the evaluated configuration\. Observation formatting, approval, sandboxing, retry limits, context truncation, wrappers, and recovery affordances can alter performance under a fixed model\. Controlled comparisons across CLI and MCP interfaces make these effects explicit\[[42](https://arxiv.org/html/2608.20485#bib.bib159)\]; Meta\-Harness\[[79](https://arxiv.org/html/2608.20485#bib.bib55)\]and AutoHarness\[[93](https://arxiv.org/html/2608.20485#bib.bib63)\]optimize the outer loop directly, while Agent Psychometrics\[[48](https://arxiv.org/html/2608.20485#bib.bib37)\]separates LLM and scaffold ability through item response theory\. Controlled terminal evaluations further show that harness choice changes token efficiency and failure profiles under fixed models, while AgentMeter jointly evaluates task quality, budget sensitivity, and LM\-CLI matching\[[145](https://arxiv.org/html/2608.20485#bib.bib155),[25](https://arxiv.org/html/2608.20485#bib.bib149)\]\. Model comparisons therefore depend on the harness conditions under which behavior is produced\.
Environment and budget comparability\.Container images, dependency caches, network access, timeouts, filesystem persistence, and tool permissions alter both success and failure modes\[[7](https://arxiv.org/html/2608.20485#bib.bib10),[74](https://arxiv.org/html/2608.20485#bib.bib177),[101](https://arxiv.org/html/2608.20485#bib.bib68)\]\. Production\-derived tasks reduce distribution mismatch but may trade breadth and controlled conditions for realism\[[65](https://arxiv.org/html/2608.20485#bib.bib50)\]\. Retry budgets, model and verifier calls, trajectory limits, and wall\-clock timeouts further determine the search space available to an agent\.
Together, these factors make the model, harness, observation and action interfaces, execution environment, permission policies, and execution budget part of the evaluated configuration\[[164](https://arxiv.org/html/2608.20485#bib.bib109),[79](https://arxiv.org/html/2608.20485#bib.bib55),[93](https://arxiv.org/html/2608.20485#bib.bib63),[48](https://arxiv.org/html/2608.20485#bib.bib37)\]\.
### 5\.4Evaluation Layers and Metrics Beyond Binary Correctness
Resolution and pass rates remain leaderboard anchors but capture only final correctness\. Complementary evidence includes behavioral analysis beyond resolution rate\[[98](https://arxiv.org/html/2608.20485#bib.bib67)\], syntax\-aware structure in SWE\-PolyBench\[[119](https://arxiv.org/html/2608.20485#bib.bib80)\], checklist\-based process scores in OctoBench\[[31](https://arxiv.org/html/2608.20485#bib.bib30)\], long\-horizon drift and faithfulness\[[163](https://arxiv.org/html/2608.20485#bib.bib107),[107](https://arxiv.org/html/2608.20485#bib.bib72)\], and multidimensional enterprise assessment\[[19](https://arxiv.org/html/2608.20485#bib.bib20)\]\. Relevant process indicators include*command economy*, relating command number and complexity to task scope;*recovery*, recording actions after failure;*diagnostic quality*, assessing whether causes are identified before repair;*state tracking*, checking consistency with filesystem, package, process, and repository state; and*governance violations*, capturing unsafe commands, permission bypass, sandbox escape, or unapproved external effects\.
These signals form a layered evidence stack\. Outcome evidence records task completion; process evidence describes execution; environment evidence establishes runtime validity; trace evidence supports inspection and replay; and governance evidence records permissions, containment, and side effects\. Freshness is not a behavioral layer but a cross\-cutting condition of evaluation validity\. Deployment\-oriented multi\-signal evaluation follows the same motivation\[[44](https://arxiv.org/html/2608.20485#bib.bib35)\]\. Outcome metrics are comparatively standardized, whereas the remaining layers use heterogeneous definitions and protocols\[[31](https://arxiv.org/html/2608.20485#bib.bib30),[57](https://arxiv.org/html/2608.20485#bib.bib135),[14](https://arxiv.org/html/2608.20485#bib.bib172),[16](https://arxiv.org/html/2608.20485#bib.bib18),[70](https://arxiv.org/html/2608.20485#bib.bib53),[152](https://arxiv.org/html/2608.20485#bib.bib100)\]\.
### 5\.5Benchmark Observability of Terminal Competence Dimensions
Table[2](https://arxiv.org/html/2608.20485#S5.T2)characterizes benchmark design emphases, while Table[3](https://arxiv.org/html/2608.20485#S5.T3)maps representative groups to the seven competence dimensions and treats freshness separately as evaluation validity\. The coding indicates which constructs each group makes observable rather than the strength of their measurement\.Explicitdenotes a targeted or scored construct,Trace\-visiblea required but unscored construct,Incidentala construct that may arise without being targeted, andNot observablelittle basis for observation or scoring\. The rows include core terminal\-agent benchmarks and terminal\-relevant boundary comparators; grouped rows reflect their shared dominant design intent\.
Table 3:Qualitative coding of terminal\-competence observability and evaluation freshness\. E=Explicit, T=Trace\-visible, I=Incidental, and –=Not observable\. Freshness concerns evaluation validity rather than competence\.BenchmarkTerminal\-competence dimensionsEval\.validityActionform\.Feedbackinterp\.Runtimemgmt\.StatetrackingVerifyRecoverGovernFresh\.SWE\-bench / SWE\-PolyBench\[[67](https://arxiv.org/html/2608.20485#bib.bib51),[119](https://arxiv.org/html/2608.20485#bib.bib80)\]ITITTI––Claw\-SWE\-Bench\[[183](https://arxiv.org/html/2608.20485#bib.bib147)\]ITITTI–ITerminal\-Bench / TerminalWorld\[[101](https://arxiv.org/html/2608.20485#bib.bib68),[26](https://arxiv.org/html/2608.20485#bib.bib145)\]EETTTT–ISetupBench\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\]EEETTT––LongCLI\-Bench / GitTaskBench\[[39](https://arxiv.org/html/2608.20485#bib.bib174),[105](https://arxiv.org/html/2608.20485#bib.bib71)\]EEIETI––CLI\-Gym\[[89](https://arxiv.org/html/2608.20485#bib.bib60)\]EETTTT––debug\-gym / OctoBench\[[174](https://arxiv.org/html/2608.20485#bib.bib116),[31](https://arxiv.org/html/2608.20485#bib.bib30)\]TEITEE––BashArena / ClawSafety\[[70](https://arxiv.org/html/2608.20485#bib.bib53),[152](https://arxiv.org/html/2608.20485#bib.bib100)\]EEITTTE–LiveSQLBench\[[138](https://arxiv.org/html/2608.20485#bib.bib143)\]TT–I–––EWildClawBench\[[34](https://arxiv.org/html/2608.20485#bib.bib133)\]EETETTI–ClawForge\[[76](https://arxiv.org/html/2608.20485#bib.bib136)\]EETETI––The matrix shows broad observability of command formulation and feedback interpretation, especially in terminal\-native and CLI\-centered benchmarks\. Runtime management is explicit in SetupBench\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\]and trace\-visible in Terminal\-Bench\[[101](https://arxiv.org/html/2608.20485#bib.bib68)\], TerminalWorld\[[26](https://arxiv.org/html/2608.20485#bib.bib145)\], CLI\-Gym\[[89](https://arxiv.org/html/2608.20485#bib.bib60)\], WildClawBench\[[34](https://arxiv.org/html/2608.20485#bib.bib133)\], and ClawForge\[[76](https://arxiv.org/html/2608.20485#bib.bib136)\]\. Long\-horizon and deployment\-like settings expose state and context tracking, often indirectly through final outcomes\. Debugging and process benchmarks provide stronger observability of verification and recovery\[[174](https://arxiv.org/html/2608.20485#bib.bib116),[31](https://arxiv.org/html/2608.20485#bib.bib30)\], whereas repository\-level evaluations rarely isolate them\. Governance is explicit mainly in BashArena\[[70](https://arxiv.org/html/2608.20485#bib.bib53)\]and ClawSafety\[[152](https://arxiv.org/html/2608.20485#bib.bib100)\]; freshness remains a separate validity property\.
### 5\.6Remaining Evaluation Gaps
Four structural gaps remain\.Terminal\-native workflow coverage is narrow\.Most benchmarks remain repository\- or coding\-centric, while alternatives are fragmented across operational domains\. TerminalWorld broadens general terminal work\[[26](https://arxiv.org/html/2608.20485#bib.bib145)\]; ML\-DevBench and MLE\-bench target iterative model development\[[109](https://arxiv.org/html/2608.20485#bib.bib74),[18](https://arxiv.org/html/2608.20485#bib.bib19)\]; ELTBench, DAComp, and DSAgentBench cover data pipelines and end\-to\-end data\-science workflows\[[68](https://arxiv.org/html/2608.20485#bib.bib52),[80](https://arxiv.org/html/2608.20485#bib.bib56),[118](https://arxiv.org/html/2608.20485#bib.bib164)\]; and ITBench and AIOpsLab introduce operational state\[[64](https://arxiv.org/html/2608.20485#bib.bib189),[24](https://arxiv.org/html/2608.20485#bib.bib142)\]\. Scientific experimentation appears in ExpBench, Curie, and ScienceBoard\[[73](https://arxiv.org/html/2608.20485#bib.bib21),[72](https://arxiv.org/html/2608.20485#bib.bib83),[135](https://arxiv.org/html/2608.20485#bib.bib183)\], while security evaluation covers CTF, penetration testing, and inference optimization\[[78](https://arxiv.org/html/2608.20485#bib.bib2),[4](https://arxiv.org/html/2608.20485#bib.bib170),[92](https://arxiv.org/html/2608.20485#bib.bib5),[104](https://arxiv.org/html/2608.20485#bib.bib69)\]\. End\-to\-end CLI tool generation also remains under\-evaluated\[[58](https://arxiv.org/html/2608.20485#bib.bib48)\]\.
Process and trace standards are immature\.Benchmarks increasingly record trajectories, but no common schema covers commands, observations, failures, retries, state changes, and human interventions\. Trajectory analyses demonstrate the value of trace inspection\[[14](https://arxiv.org/html/2608.20485#bib.bib172),[16](https://arxiv.org/html/2608.20485#bib.bib18)\]; ProcBench adds step\-level assessment\[[57](https://arxiv.org/html/2608.20485#bib.bib135)\]; and IDE\-Bench extends evaluation to IDE settings\[[97](https://arxiv.org/html/2608.20485#bib.bib66)\]\. Studies of failed agentic pull requests and CLI trajectories further motivate shared failure taxonomies\[[38](https://arxiv.org/html/2608.20485#bib.bib166),[181](https://arxiv.org/html/2608.20485#bib.bib161)\]\. Interactive trajectory debugging remains disconnected from benchmark protocols\[[61](https://arxiv.org/html/2608.20485#bib.bib49)\]\.
Freshness mechanisms remain peripheral\.SWE\-rebench reports evidence consistent with contamination\-related inflation on static tasks\[[9](https://arxiv.org/html/2608.20485#bib.bib13)\]\. Mutation and detection methods are emerging\[[45](https://arxiv.org/html/2608.20485#bib.bib36),[133](https://arxiv.org/html/2608.20485#bib.bib88),[17](https://arxiv.org/html/2608.20485#bib.bib173)\], but rarely enter standard evaluation pipelines, while temporal\-consistency mechanisms remain experimental\[[134](https://arxiv.org/html/2608.20485#bib.bib104)\]\.
Safety and governance remain separated from task success\.Existing benchmarks isolate privileged terminal actions\[[70](https://arxiv.org/html/2608.20485#bib.bib53)\]; risky code generation and execution\[[53](https://arxiv.org/html/2608.20485#bib.bib42),[20](https://arxiv.org/html/2608.20485#bib.bib190),[28](https://arxiv.org/html/2608.20485#bib.bib26),[8](https://arxiv.org/html/2608.20485#bib.bib12)\]; productivity and computer\-use agents\[[82](https://arxiv.org/html/2608.20485#bib.bib57),[40](https://arxiv.org/html/2608.20485#bib.bib32),[75](https://arxiv.org/html/2608.20485#bib.bib54),[22](https://arxiv.org/html/2608.20485#bib.bib24)\]; and attacks on long\-running or sandboxed agents\[[96](https://arxiv.org/html/2608.20485#bib.bib131),[171](https://arxiv.org/html/2608.20485#bib.bib113),[132](https://arxiv.org/html/2608.20485#bib.bib165)\]\. UnderSpecBench extends earlier evidence of governance blind spots under benign instructions by measuring wrong\-target and over\-scope actions in underspecified DevOps tasks, while Boundary\-Bench evaluates coding agents under progressively hardened execution policies\[[35](https://arxiv.org/html/2608.20485#bib.bib29),[66](https://arxiv.org/html/2608.20485#bib.bib153),[27](https://arxiv.org/html/2608.20485#bib.bib156)\]\. These protocols make authorization and containment observable, but governance evidence remains distributed across separate task\-success and policy\-focused settings\.
Current evaluation suites distribute these requirements across separate protocols rather than jointly integrating workflow realism, environment setup, process quality, trace visibility, long\-horizon persistence, freshness, safety, and governance\[[7](https://arxiv.org/html/2608.20485#bib.bib10),[57](https://arxiv.org/html/2608.20485#bib.bib135),[74](https://arxiv.org/html/2608.20485#bib.bib177),[9](https://arxiv.org/html/2608.20485#bib.bib13),[70](https://arxiv.org/html/2608.20485#bib.bib53),[152](https://arxiv.org/html/2608.20485#bib.bib100)\]\. Evaluation results therefore characterize a coupled configuration defined by the model, benchmark, harness, runtime, sandbox, observation format, and execution budget, while trace\-release policy determines how much process evidence remains available for analysis\.
Architecture shapes what can be executed and recorded, acquisition determines which recorded behavior can be learned, and evaluation determines which behaviors become evidence\. Section[6](https://arxiv.org/html/2608.20485#S6)uses bounded fixed\-condition diagnostics to illustrate two consequences: benchmark\-dependent process exposure and limits of component attribution\.
## 6Framework\-Guided Diagnostics of Process Observability and Attribution Limits
The preceding sections connect terminal competence to system architecture, acquisition data, and evaluation protocols\. This section makes two implications of that synthesis concrete through bounded fixed\-condition diagnostics\. First, under a fixed agent configuration, we examine which process signals different benchmark families expose\. Second, on matched repository\-repair tasks, we examine how outcomes vary across systems and what these differences reveal about component attribution\. Trace cases complement the aggregate results with trajectory\-level mechanisms\.
For the benchmark\-exposure diagnostic, we fix mini\-SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]with DeepSeek\-V4\-Flash\[[161](https://arxiv.org/html/2608.20485#bib.bib146)\]while varying four benchmark families\. The matched diagnostic separately evaluates mini\-SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\], SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\], and OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]on identical task identifiers within Claw\-SWE\-Bench Lite\[[183](https://arxiv.org/html/2608.20485#bib.bib147)\]and SWE\-bench Lite\[[67](https://arxiv.org/html/2608.20485#bib.bib51)\], using DeepSeek\-V4\-Flash and DeepSeek\-V4\-Pro\[[161](https://arxiv.org/html/2608.20485#bib.bib146)\]\. Each task has one formal run per condition, and each cell represents a complete model, interface, harness, and runtime configuration\.
### 6\.1Diagnostic Design and Measures
Both diagnostics retain benchmark\-native outcomes and trace\-level process evidence\. Seven trace\-derived indicators, denoted P1–P7, capture selected observable aspects associated with the competence dimensions introduced in Section[2](https://arxiv.org/html/2608.20485#S2)\. The identifiers provide concise cross\-reference and are paired with semantic names throughout the analysis\.
Four indicators are deterministic\. The*rule\-matched invocation\-failure rate*\(P1\) divides failed commands whose stderr matches shell syntax, command\-not\-found, non\-executable, path, permission, argument, or malformed\-tool patterns by all normalized command events\. The*environment\-exit rate*\(P3\) divides non\-zero exits on dependency, runtime, service, configuration, or environment\-management commands by all detected environment\-management commands\. The*final\-window verification rate*\(P5\) is the fraction of tasks with a detected verification action in the final five actions or final 20% of the trace, whichever window is larger\. The*governance\-review\-trigger rate*\(P7\) divides command events matching irreversible, overprivileged, secret\-handling, host or sandbox, or external\-side\-effect patterns by all normalized command events\.
Three auxiliary indicators use rule\-constrained LLM judgments\. Rule\-based extractors first identify candidate episodes concerning feedback use, state errors, or recovery, and DeepSeek\-V4\-Flash\[[161](https://arxiv.org/html/2608.20485#bib.bib146)\]evaluates each target within a bounded evidence window\. P2 is the helpful\-use rate among eligible episodes with usable feedback, P4 the consequential\-error rate among episodes with decidable state evidence, and P6 the successful task\-relevant recovery rate among episodes containing a recovery attempt\. Uncertain or low\-confidence labels are excluded, and task\-level rates are macro\-averaged over eligible evidence\.
All formal runs use fixed decoding settings, task identifiers, and benchmark\-native evaluators\. Harness\-specific prompting, context management, action interfaces, and orchestration remain part of each evaluated system\. Complete model settings, execution budgets, extractor definitions, and auxiliary\-judge procedures are reported in the Supplementary Material\.
The benchmark\-exposure profile contains 241 Terminal\-Bench 2\.1 tasks\[[101](https://arxiv.org/html/2608.20485#bib.bib68)\], 93 SetupBench tasks\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\], 21 LongCLI\-Bench tasks\[[39](https://arxiv.org/html/2608.20485#bib.bib174)\], and 640 BashArena tasks\[[70](https://arxiv.org/html/2608.20485#bib.bib53)\]\. The matched system comparison contains all 80 official Claw\-SWE\-Bench Lite tasks and 300 SWE\-bench Lite tasks\.
### 6\.2Benchmark\-Exposure Diagnostic
Table[4](https://arxiv.org/html/2608.20485#S6.T4)reports benchmark\-native outcomes and the seven trace\-derived process indicators\. Because the benchmarks use different task distributions and evaluators, outcomes are interpreted within each benchmark rather than as a shared competence scale\. P1, P3, P5, and P7 are deterministic, whereas P2, P4, and P6 are judge\-assisted task\-level macro averages over eligible semantic evidence\.
Table 4:Benchmark\-exposure profile under mini\-SWE\-agent with DeepSeek\-V4\-Flash\. Outcome is the official benchmark score\. P1–P7 are bounded trace\-derived process indicators\. P1 and P3 record execution\-friction signals; P2, P4, and P6 are auxiliary judge\-assisted task\-level rates; P5 records a detected final\-window check; and P7 flags actions requiring contextual governance review\.BenchmarkTasksOutcomeP1Rule inv\. fail\.P2FeedbackP3Env\. exitP4State err\.P5Final verifyP6RecoveryP7Gov\. triggerTerminal\-Bench 2\.1\[[101](https://arxiv.org/html/2608.20485#bib.bib68)\]24152\.6%0\.0%82\.1%12\.5%14\.7%29\.5%79\.9%0\.5%SetupBench\[[7](https://arxiv.org/html/2608.20485#bib.bib10)\]9359\.1%0\.0%80\.2%9\.0%19\.5%36\.6%85\.5%1\.4%LongCLI\-Bench\[[39](https://arxiv.org/html/2608.20485#bib.bib174)\]2123\.8%0\.0%76\.4%6\.8%7\.7%23\.8%67\.8%4\.5%BashArena\[[70](https://arxiv.org/html/2608.20485#bib.bib53)\]64041\.8%0\.0%76\.3%13\.7%23\.3%66\.2%67\.0%3\.0%Three findings characterize the observed profile\.Rule\-matched invocation failures are rare under the fixed configuration\.P1 rounds to 0\.0% across all four benchmarks, indicating that few command failures match the specified shell syntax, command\-not\-found, non\-executable, path, permission, argument, or malformed\-tool patterns\. Broader action\-formulation problems instead appear through semantic or state\-dependent behavior not captured by this narrow signal\.
Local feedback use and end\-to\-end completion expose different aspects of behavior\.The auxiliary helpful\-feedback indicator \(P2\) ranges from 76\.3% to 82\.1%, showing that many eligible episodes contain a locally useful response to visible feedback\. Benchmark\-native outcomes additionally depend on runtime management, persistent state tracking, recovery, and verification\. Detected final\-window verification ranges from 23\.8% to 36\.6% for Terminal\-Bench, SetupBench, and LongCLI\-Bench, compared with 66\.2% for BashArena, showing substantial differences in late\-trace inspection across benchmark conditions\.
Benchmark conditions foreground different process limitations\.Under the fixed configuration, SetupBench foregrounds environment bootstrap, LongCLI\-Bench combines low outcome with long\-horizon interaction demands, and BashArena increases the visibility of state errors and governance\-relevant actions\. The auxiliary indicators show broadly high accepted\-label coverage, although LongCLI\-Bench has both the smallest task set and lower semantic coverage than the other benchmarks\. Its P2, P4, and P6 estimates are therefore interpreted directionally\.
### 6\.3Matched Complete\-System Diagnostic
The second diagnostic focuses on outcome differences across systems rather than extending the P1–P7 profile\. Table[5](https://arxiv.org/html/2608.20485#S6.T5)reports matched execution snapshots on identical task identifiers within each benchmark\.
SWE\-agent has the highest resolved rate in all four benchmark–model blocks\. Its largest gap to the lowest\-performing system is 21\.25 percentage points on Claw\-SWE\-Bench Lite, compared with at most 7\.00 points on SWE\-bench Lite\. The relative ordering of OpenHands and mini\-SWE\-agent also reverses across benchmarks: OpenHands is higher on Claw\-SWE\-Bench Lite but lower on SWE\-bench Lite\. Benchmark choice therefore changes the observed separation among systems even though SWE\-agent remains highest under all reported conditions\.
Table 5:Matched single\-run results by benchmark and system\. F/P denote DeepSeek\-V4\-Flash/Pro; entries report resolved rate, Pro\-minus\-Flash difference, and mean task time in seconds\.SystemFPΔ\\DeltaTime F/P*Claw\-SWE\-Bench Lite*mini\-SWE\-agent58\.75%56\.25%−2\.50\-2\.50476/476SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]76\.25%77\.50%\+1\.25\+1\.25675/712OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]68\.75%67\.50%−1\.25\-1\.25600/653*SWE\-bench Lite*mini\-SWE\-agent55\.67%55\.67%0\.000\.00294/389SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]58\.67%58\.33%−0\.33\-0\.33382/525OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]51\.67%51\.67%0\.000\.00501/504Within a fixed system, the absolute Pro\-minus\-Flash difference never exceeds 2\.50 points\. Exact task\-paired McNemar tests detect no directional model\-variant difference after Holm adjustment, and all paired bootstrap intervals include zero\. Under these conditions, the two model variants produce similar resolved rates across the evaluated systems\. Pro has higher observed mean wall\-clock time in five of six system–benchmark pairs; under mini\-SWE\-agent, it also uses more mean input tokens on both benchmarks without an outcome gain\.
Task\-paired cross\-system tests further characterize these snapshots\. Cochran’sQQrejects equal resolved rates among the three systems in every benchmark–model block \(Q=10\.57Q=10\.57to16\.0016\.00,p<0\.006p<0\.006\)\. After Holm correction across 12 pairwise comparisons, SWE\-agent differs from mini\-SWE\-agent on Claw\-SWE\-Bench Lite for both Flash \(p=0\.023p=0\.023\) and Pro \(p=0\.002p=0\.002\), and from OpenHands on SWE\-bench Lite for both Flash \(p=0\.001p=0\.001\) and Pro \(p=0\.001p=0\.001\)\. The other corrected pairwise contrasts are not significant\. These results characterize task\-level differences within the recorded execution snapshots\.
### 6\.4Illustrative Trace Cases
Aggregate rates alone do not show how these differences arise within trajectories\. Two trace cases illustrate how benchmark demands and system behavior become visible at the execution level\.
Benchmark\-exposure case:magsac\-installCondition:Terminal\-Bench 2\.1 with mini\-SWE\-agent and DeepSeek\-V4\-Flash\.Outcome:unresolved; the evaluator reportsNo module named ’pymagsac’\.Trace evidence:repeated dependency and build attempts yield a 38\.5% environment\-command non\-zero\-exit rate\. The trace ends before an evaluator\-relevant import or installation\-source check\.Interpretation:local feedback leaves the installation loop unclosed, linking runtime management, verification, and recovery across the trajectory\.
Together, the aggregate results and trace cases show that benchmark conditions change which process demands become visible, while system configuration changes realized trajectories and outcomes on matched tasks\. These diagnostics connect the survey framework to observable execution behavior and motivate evaluation that combines task outcomes with process evidence and explicit system conditions\.
Matched case:rubocop\_\_rubocop\-13560Condition:Claw\-SWE\-Bench Lite with DeepSeek\-V4\-Flash\.Outcome:SWE\-agent resolves the task; mini\-SWE\-agent and OpenHands do not\.Trace evidence:the issue requires ordinary lowercase’nul’arguments to remain valid\. Mini\-SWE\-agent changes the implementation and an existing test despite this constraint\. OpenHands exempts all method\-call arguments, broadening the requested behavior\. SWE\-agent changes production logic while preserving the required distinction and passes the official evaluator\.Interpretation:the task exposes different patch scopes and outcomes across systems\.
## 7Challenges and Future Directions
The literature synthesis identifies four connected gaps: domain concentration limits evidence for transferable competence, fragmented process evaluation restricts behavioral interpretation, runtime governance remains weakly integrated with task evaluation, and system\-level comparisons complicate component attribution\. These gaps recur across the architecture, acquisition, and evaluation evidence reviewed in Sections[3](https://arxiv.org/html/2608.20485#S3)–[5](https://arxiv.org/html/2608.20485#S5)and are further illustrated by the process\-observability and attribution diagnostics in Section[6](https://arxiv.org/html/2608.20485#S6), motivating four research priorities\.
### 7\.1Cross\-Domain Terminal Competence
Training and evaluation remain concentrated in software engineering\[[67](https://arxiv.org/html/2608.20485#bib.bib51),[164](https://arxiv.org/html/2608.20485#bib.bib109),[110](https://arxiv.org/html/2608.20485#bib.bib181)\]\. Repositories provide structured executable verification, but cannot determine whether terminal behavior reflects transferable command and interaction competence or task\-specific heuristics\. Operations, data engineering, scientific workflows, cloud management, and cybersecurity introduce distinct runtime and governance demands\[[124](https://arxiv.org/html/2608.20485#bib.bib171),[130](https://arxiv.org/html/2608.20485#bib.bib137),[144](https://arxiv.org/html/2608.20485#bib.bib70),[68](https://arxiv.org/html/2608.20485#bib.bib52),[72](https://arxiv.org/html/2608.20485#bib.bib83),[24](https://arxiv.org/html/2608.20485#bib.bib142),[70](https://arxiv.org/html/2608.20485#bib.bib53)\], while TUA\-Bench extends evaluation to routine digital activities and scientific and engineering workflows\[[21](https://arxiv.org/html/2608.20485#bib.bib152)\]\. Cross\-domain research should distinguish reusable substrate\-level competence from domain\-specific workflows and compare transfer across competence dimensions to identify which capabilities remain workload\-specific\.
### 7\.2Fresh, Replayable, Process\-Level Evaluation
Final outcomes do not reveal diagnostic quality, state preservation, recovery, or unsafe intermediate behavior\[[31](https://arxiv.org/html/2608.20485#bib.bib30),[57](https://arxiv.org/html/2608.20485#bib.bib135),[74](https://arxiv.org/html/2608.20485#bib.bib177)\]\. Live or regenerated tasks improve temporal validity, while replayable traces preserve commands, observations, state changes, interventions, and verifier calls\[[9](https://arxiv.org/html/2608.20485#bib.bib13),[14](https://arxiv.org/html/2608.20485#bib.bib172),[16](https://arxiv.org/html/2608.20485#bib.bib18)\]\. Combining them would connect fresh task distributions with process evidence, distinguishing successful completion from the behavior that produces it\. This also links evaluation to acquisition in Section[4](https://arxiv.org/html/2608.20485#S4), where retained failures, state transitions, and recovery trajectories determine what can be learned and inspected\.
### 7\.3Runtime Governability and Safety
Terminal agents can modify filesystems, install packages, launch processes, access networks, and operate around credentials\. Permission gates, sandboxes, approvals, and rollback constrain these effects, but governance remains distributed across separate safety and policy\-focused protocols\[[5](https://arxiv.org/html/2608.20485#bib.bib9),[70](https://arxiv.org/html/2608.20485#bib.bib53),[152](https://arxiv.org/html/2608.20485#bib.bib100),[78](https://arxiv.org/html/2608.20485#bib.bib2),[66](https://arxiv.org/html/2608.20485#bib.bib153),[27](https://arxiv.org/html/2608.20485#bib.bib156)\]\. Future evaluation should integrate the architectural controls identified in Section[3](https://arxiv.org/html/2608.20485#S3), including authorization, containment, reversibility, destructive\-action prevention, and external side\-effect control\. Task success and governance should be assessed jointly so that effective execution reflects the constraints under which actions are taken\.
### 7\.4Controlled Model and Harness Attribution
Measured performance can reflect the model, interface, runtime, context policy, retry budget, or verifier access\[[164](https://arxiv.org/html/2608.20485#bib.bib109),[150](https://arxiv.org/html/2608.20485#bib.bib97),[79](https://arxiv.org/html/2608.20485#bib.bib55),[93](https://arxiv.org/html/2608.20485#bib.bib63)\]\. Controlled studies further show that harness and LM–CLI choices can alter quality, efficiency, and failure behavior\[[145](https://arxiv.org/html/2608.20485#bib.bib155),[25](https://arxiv.org/html/2608.20485#bib.bib149)\]\. Separating model capability from surrounding system support therefore remains a central problem\. The matched diagnostic in Section[6](https://arxiv.org/html/2608.20485#S6)illustrates that observed system differences can also vary across benchmarks\. Factorial designs or portable protocols that vary the model, interface, harness, or runtime while preserving task identity and execution conditions\[[140](https://arxiv.org/html/2608.20485#bib.bib160)\]can better isolate which gains transfer across configurations and which arise from specific model–harness interactions\.
## 8Conclusion
We organize the study of terminal agents around terminal\-mediated execution and its state\-changing action–observation loop, using a seven\-dimensional competence profile to connect system architecture, competence acquisition, and evaluation\.
Three conclusions emerge\. First, terminal\-agent behavior is jointly shaped by the model, interface, harness, runtime, and environment, making outer\-loop design a performance\-shaping system component\. Second, executable trajectories ground learning in feedback, verification, and recovery, yet current acquisition remains concentrated on successful repository workflows and underrepresents environment management, persistent state, recovery, and governance\. Third, prevailing evaluations emphasize final outcomes and expose process quality unevenly, while measured performance depends on both benchmark conditions and system configuration\. The bounded diagnostics further illustrate benchmark\-dependent process exposure and limits of component attribution\.
Progress therefore requires cross\-domain acquisition and evaluation, fresh and replayable process evidence, governable runtimes, and controlled model–harness attribution\. Evaluation should combine explicit system and runtime conditions with task outcomes, process evidence, and trace provenance\.
## References
- \[1\]S\. Abuzakuk, L\. Crijns, A\. Kermarrec, R\. Pires, and M\. de Vos\(2026\)RIVA: leveraging llm agents for reliable configuration drift detection\.InProceedings of the Sixth European Workshop on Machine Learning and Systems,pp\. 499–509\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[2\]P\. Adamenko, M\. Ivanov, A\. Valeev, R\. Levichev, P\. Zadorozhny, I\. Lopatin, D\. Babaev, A\. Fenogenova, and V\. Malykh\(2025\)Swe\-mera: a dynamic benchmark for agenticly evaluating large language models on software engineering tasks\.arXiv preprint arXiv:2507\.11059\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1)\.
- \[3\]S\. Agarwal, H\. He, and B\. Vasilescu\(2026\)Ai ides or autonomous agents? measuring the impact of coding agents on software development\.InProceedings of the 23rd International Conference on Mining Software Repositories,pp\. 857–862\.Cited by:[§5](https://arxiv.org/html/2608.20485#S5.p4.1)\.
- \[4\]A\. Al\-Kaswan, M\. Plotnikov, M\. Hájek, R\. Vízner, A\. van Deursen, and M\. Izadi\(2026\)Do agents dream of root shells? partial\-credit evaluation of llm agents in capture the flag challenges\.InProceedings of the 3rd ACM International Conference on AI\-Powered Software,pp\. 349–357\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[5\]M\. Andriushchenko, A\. Souly, M\. Dziemian, D\. Duenas, M\. Lin, J\. Wang, D\. Hendrycks, A\. Zou, Z\. Kolter, M\. Fredrikson,et al\.\(2025\)Agentharm: a benchmark for measuring harmfulness of llm agents\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 79185–79220\.Cited by:[§7\.3](https://arxiv.org/html/2608.20485#S7.SS3.p1.1)\.
- \[6\]Anthropic\(2025\)Claude Code\.Note:Agentic coding tool available through terminal and other development surfaces, with file editing, command execution, and development\-tool integration\.External Links:[Link](https://code.claude.com/docs/en/overview)Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[7\]A\. Arora, J\. Jang, and R\. Z\. Moghaddam\(2025\)SetupBench: assessing software engineering agents’ ability to bootstrap development environments\.arXiv preprint arXiv:2507\.09063\.Cited by:[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.31.3.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p2.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p2.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p4.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p5.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.4.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.6.1.1.1),[§5](https://arxiv.org/html/2608.20485#S5.p1.1),[§6\.1](https://arxiv.org/html/2608.20485#S6.SS1.p5.1),[Table 4](https://arxiv.org/html/2608.20485#S6.T4.5.3.1.1.1)\.
- \[8\]L\. Ba, Q\. Li, and S\. Li\(2026\)CIBER: a comprehensive benchmark for security evaluation of code interpreter agents\.arXiv preprint arXiv:2602\.19547\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[9\]I\. Badertdinov, A\. Golubev, M\. Nekrashevich, A\. Shevtsov, S\. Karasik, A\. Andriushchenko, M\. Trofimova, D\. Litvintseva, and B\. Yangel\(2026\)Swe\-rebench: an automated pipeline for task collection and decontaminated evaluation of software engineering agents\.Advances in Neural Information Processing Systems38\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.32.3.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p3.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p5.1),[§7\.2](https://arxiv.org/html/2608.20485#S7.SS2.p1.1)\.
- \[10\]I\. Badertdinov, M\. Nekrashevich, A\. Shevtsov, and A\. Golubev\(2026\)Swe\-rebench v2: language\-agnostic swe task collection at scale\.arXiv preprint arXiv:2602\.23866\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1)\.
- \[11\]P\. Bechard, O\. M\. Ayala, E\. Chen, J\. Skelton, S\. Davasam, S\. Sunkara, V\. Yadav, and S\. Rajeswar\(2026\)Terminal agents suffice for enterprise automation\.arXiv preprint arXiv:2604\.00073\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p2.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[12\]N\. Benkovich and V\. Valkov\(2026\)Agyn: a multi\-agent system for team\-based autonomous software engineering\.arXiv preprint arXiv:2602\.01465\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p5.1)\.
- \[13\]M\. Bilal, J\. Crowcroft, R\. Wang, X\. Xu, and S\. Dustdar\(2026\)Large language models for agentic netops and aiops: architectures, evaluation, and safety\.arXiv preprint arXiv:2605\.12729\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[14\]I\. Bouzenia and M\. Pradel\(2025\)Understanding software engineering agents: a study of thought\-action\-result trajectories\.In2025 40th IEEE/ACM International Conference on Automated Software Engineering \(ASE\),pp\. 2846–2857\.Cited by:[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1),[§7\.2](https://arxiv.org/html/2608.20485#S7.SS2.p1.1)\.
- \[15\]N\. D\. Bui\(2026\)Building effective ai coding agents for the terminal: scaffolding, harness, context engineering, and lessons learned\.arXiv preprint arXiv:2603\.05344\.Cited by:[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[16\]I\. Ceka, S\. Pujar, S\. Ramji, L\. Buratti, G\. Kaiser, and B\. Ray\(2025\)Understanding software engineering agents through the lens of traceability: an empirical study\.arXiv preprint arXiv:2506\.08311\.Cited by:[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1),[§7\.2](https://arxiv.org/html/2608.20485#S7.SS2.p1.1)\.
- \[17\]J\. Chai, Y\. Zhe, and J\. Sakuma\(2026\)When benchmarks leak: inference\-time decontamination for llms\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 44743–44760\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p3.1)\.
- \[18\]J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace, K\. Liu, L\. Maksin, T\. Patwardhan,et al\.\(2025\)Mle\-bench: evaluating machine learning agents on machine learning engineering\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 50466–50494\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[19\]A\. Chandwani and I\. Gupta\(2026\)Beyond binary correctness: scaling evaluation of long\-horizon agents on subjective enterprise tasks\.arXiv preprint arXiv:2603\.22744\.Cited by:[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p1.1)\.
- \[20\]J\. Chen, H\. Huang, Y\. Lyu, J\. An, J\. Shi, C\. Yang, T\. Zhang, H\. Tian, Y\. Li, Z\. Li,et al\.\(2026\)Securevibebench: benchmarking secure vibe coding of ai agents via reconstructing vulnerability\-introducing scenarios\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 24144–24168\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[21\]S\. Chen, L\. Wang, X\. Yang, Z\. Liu, Y\. Cong, Y\. Ji, F\. Zhou, X\. Zhang, F\. Yang, and B\. Zeng\(2026\)TUA\-bench: a benchmark for general\-purpose terminal\-use agents\.arXiv preprint arXiv:2606\.28480\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[22\]T\. Chen, C\. Hu, G\. Gao, D\. Liu, X\. Hu, and W\. Wang\(2026\)LPS\-bench: benchmarking safety awareness of computer\-use agents in long\-horizon planning under benign and adversarial scenarios\.arXiv preprint arXiv:2602\.03255\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[23\]Y\. Chen, J\. Pan, J\. Clark, Y\. Su, N\. Zheutlin, B\. Bhavya, R\. R\. Arora, Y\. Deng, S\. Jha, and T\. Xu\(2026\)Stratus: a multi\-agent system for autonomous reliability engineering of modern clouds\.Advances in Neural Information Processing Systems38,pp\. 50119–50165\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1)\.
- \[24\]Y\. Chen, M\. Shetty, G\. Somashekar, M\. Ma, Y\. Simmhan, J\. Mace, C\. Bansal, R\. Wang, and S\. Rajmohan\(2025\)Aiopslab: a holistic framework to evaluate ai agents for enabling autonomous clouds\.Proceedings of Machine Learning and Systems7\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[25\]H\. Chi, J\. Qi, Y\. Cui, B\. Lai, and J\. Huang\(2026\)Matching matters: a fair quality\-efficiency benchmark for command\-line agents\.External Links:2606\.21140,[Link](https://arxiv.org/abs/2606.21140)Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p3.1),[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[26\]Z\. Chu, J\. Hu, X\. Jiang, P\. Zou, H\. Li, C\. Peng, P\. O’Hearn, E\. T\. Barr, M\. Harman, F\. Sarro,et al\.\(2026\)TerminalWorld: benchmarking agents on real\-world terminal tasks\.arXiv preprint arXiv:2605\.22535\.Cited by:[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.32.2.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.3.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.5.1.1.1.1)\.
- \[27\]D\. Davidovich, Y\. Amar, H\. Rozencwajg, and O\. Hiltch\(2026\)Permission denied: policy\-graded evaluation of coding agents in hardened environments\.arXiv preprint arXiv:2608\.02670\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1),[§7\.3](https://arxiv.org/html/2608.20485#S7.SS3.p1.1)\.
- \[28\]A\. Dawson, R\. Mulla, N\. Landers, and S\. Caldwell\(2025\)Airtbench: measuring autonomous ai red teaming capabilities in language models\.arXiv preprint arXiv:2506\.14682\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[29\]A\. De Masi\(2026\)Terminal is all you need: design properties for human\-ai agent collaboration\.arXiv preprint arXiv:2603\.10664\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p2.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p1.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p2.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p2.1)\.
- \[30\]X\. Deng, J\. Da, E\. Pan, Y\. Y\. He, C\. Ide, K\. Garg, N\. Lauffer, A\. Park, N\. Pasari, C\. Rane,et al\.\(2025\)Swe\-bench pro: can ai agents solve long\-horizon software engineering tasks?\.arXiv preprint arXiv:2509\.16941\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.6.3.1.1)\.
- \[31\]D\. Ding, S\. Liu, E\. Yang, J\. Lin, Z\. Chen, S\. Dou, H\. Guo, W\. Cheng, P\. Zhao, C\. Xiao,et al\.\(2026\)OctoBench: benchmarking scaffold\-aware instruction following in repository\-grounded agentic coding\.arXiv preprint arXiv:2601\.10343\.Cited by:[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p1.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.5.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.9.1.1.1.1),[§7\.2](https://arxiv.org/html/2608.20485#S7.SS2.p1.1)\.
- \[32\]J\. Ding, S\. Long, C\. Pu, H\. Zhou, H\. Gao, X\. Gao, C\. He, Y\. Hou, F\. Hu, Z\. Li,et al\.\(2025\)NL2Repo\-bench: towards long\-horizon repository generation evaluation of coding agents\.arXiv preprint arXiv:2512\.12730\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1)\.
- \[33\]L\. Ding\(2026\)AgentHER: hindsight experience replay for llm agent trajectory relabeling\.arXiv preprint arXiv:2603\.21357\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.30.4.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p2.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p1.1)\.
- \[34\]S\. Ding, X\. Dai, L\. Xing, S\. Ding, Z\. Liu, Y\. JingYi, P\. Yang, Z\. Zhang, X\. Wei, X\. Fang,et al\.\(2026\)WildClawBench: a benchmark for real\-world, long\-horizon agent evaluation\.arXiv preprint arXiv:2605\.10912\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.12.1.1.1)\.
- \[35\]X\. Ding, S\. Zhai, L\. Song, J\. Li, T\. Shi, N\. Meade, S\. Reddy, J\. Kang, and J\. Zhao\(2026\)The blind spot of agent safety: how benign user instructions expose critical vulnerabilities in computer\-use agents\.arXiv preprint arXiv:2604\.10577\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[36\]Y\. Dong, X\. Jiang, J\. Qian, T\. Wang, K\. Zhang, Z\. Jin, and G\. Li\(2025\)A survey on code generation with llm\-based agents\.arXiv preprint arXiv:2508\.00083\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[37\]Y\. Du, Y\. Cai, Y\. Zhou, C\. Wang, Y\. Qian, X\. Pang, Q\. Liu, Y\. Hu, and S\. Chen\(2025\)Swe\-dev: evaluating and training autonomous feature\-driven software development\.arXiv preprint arXiv:2505\.16975\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[38\]R\. Ehsani, S\. Pathak, S\. Rawal, A\. Al Mujahid, M\. M\. Imran, and P\. Chatterjee\(2026\)Where do ai coding agents fail? an empirical study of failed agentic pull requests in github\.InProceedings of the 23rd International Conference on Mining Software Repositories,pp\. 807–811\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1)\.
- \[39\]Y\. Feng, J\. Sun, Z\. Yang, J\. Ai, C\. Li, Z\. Li, F\. Zhang, K\. He, R\. Ma, J\. Lin,et al\.\(2026\)Longcli\-bench: a preliminary benchmark and study for long\-horizon agentic programming in command\-line interfaces\.InFindings of the Association for Computational Linguistics: ACL 2026,pp\. 29952–29963\.Cited by:[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.31.4.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.3.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.7.1.1.1.1),[§6\.1](https://arxiv.org/html/2608.20485#S6.SS1.p5.1),[Table 4](https://arxiv.org/html/2608.20485#S6.T4.5.4.1.1.1)\.
- \[40\]Y\. Feng, Y\. Ding, Y\. Tan, X\. Ma, Y\. Li, Y\. Wu, Y\. Gao, K\. Zhai, and Y\. Guo\(2026\)Agenthazard: a benchmark for evaluating harmful behavior in computer\-use agents\.arXiv preprint arXiv:2604\.02947\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.7.3.1.1)\.
- \[41\]H\. Foerster, T\. Blanchard, K\. Nikolić, I\. Shumailov, C\. Zhang, R\. Mullins, N\. Papernot, F\. Tramèr, and Y\. Zhao\(2026\)Camels can use computers too: system\-level security for computer use agents\.arXiv preprint arXiv:2601\.09923\.Cited by:[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p5.1)\.
- \[42\]M\. A\. Forment, M\. J\. C\. Guerrero, F\. J\. García\-Peñalvo, and J\. Pereira\(2026\)The scaffolding matters more than the interface: a controlled comparison of mcp and cli tool use across seven agent scaffoldings, five language models, and one software task\.External Links:2608\.08654,[Link](https://arxiv.org/abs/2608.08654)Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p3.1)\.
- \[43\]K\. Gandhi, S\. Garg, N\. D\. Goodman, and D\. Papailiopoulos\(2026\)Endless terminals: scaling rl environments for terminal agents\.arXiv preprint arXiv:2601\.16443\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1),[§4](https://arxiv.org/html/2608.20485#S4.p1.1)\.
- \[44\]Y\. Gao, M\. Wang, and Y\. L\. Yu\(2026\)AgentPulse: a continuous multi\-signal framework for evaluating ai agents in deployment\.arXiv preprint arXiv:2604\.24038\.Cited by:[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1)\.
- \[45\]S\. Garg, B\. Steenhoek, and Y\. Huang\(2025\)Saving swe\-bench: a benchmark mutation approach for realistic agent evaluation\.arXiv preprint arXiv:2510\.08996\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p3.1)\.
- \[46\]P\. S\. Garigipati, O\. Ayan, K\. C\. Joshi, and X\. An\(2026\)Beyond state machines: executing network procedures with agentic tool\-calling sequences\.arXiv preprint arXiv:2605\.02584\.Cited by:[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[47\]P\. Gauthier\(2025\)Aider: ai pair programming in your terminal\.Note:Open\-source terminal\-native AI pair\-programming tool for editing and managing codebases with LLMs\.External Links:[Link](https://github.com/Aider-AI/aider)Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[48\]C\. Ge, D\. Kryvosheieva, D\. Fried, U\. Girit, and K\. Hariharan\(2026\)Agent psychometrics: task\-level performance prediction in agentic coding benchmarks\.arXiv preprint arXiv:2604\.00594\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p3.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p5.1)\.
- \[49\]J\. Geng and G\. Neubig\(2026\)Effective strategies for asynchronous software engineering agents\.arXiv preprint arXiv:2603\.21489\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p5.1)\.
- \[50\]A\. Golubev, M\. Trofimova, S\. Polezhaev, I\. Badertdinov, M\. Nekrashevich, A\. Shevtsov, S\. Karasik, S\. Abramov, A\. Andriushchenko, F\. Fisin,et al\.\(2025\)Training long\-context, multi\-turn software engineering agents with reinforcement learning\.arXiv preprint arXiv:2508\.03501\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1)\.
- \[51\]H\. Gong, C\. Li, R\. Chang, and W\. Shen\(2025\)Secure and efficient access control for computer\-use agents via context space\.arXiv preprint arXiv:2509\.22256\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[52\]Google\(2025\)Gemini CLI\.Note:Open\-source terminal AI agent for Gemini models with file operations, shell commands, web tools, and MCP integration\.External Links:[Link](https://github.com/google-gemini/gemini-cli)Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[53\]C\. Guo, X\. Liu, C\. Xie, A\. Zhou, Y\. Zeng, Z\. Lin, D\. Song, and B\. Li\(2024\)Redcode: risky code execution and generation benchmark for code agents\.Advances in Neural Information Processing Systems37,pp\. 106190–106236\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[54\]D\. Guo, J\. Wu, and S\. M\. Yiu\(2026\)AgentEval: dag\-structured step\-level evaluation for agentic workflows with error propagation tracking\.arXiv preprint arXiv:2604\.23581\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1)\.
- \[55\]L\. Guo, Y\. Wang, C\. Li, W\. Tao, P\. Yang, J\. Chen, H\. Song, D\. Tang, and Z\. Zheng\(2025\)Swe\-factory: your automated factory for issue resolution training data and evaluation benchmarks\.arXiv preprint arXiv:2506\.10954\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[56\]T\. Han, Y\. Zhang, W\. Song, C\. Fang, Z\. Chen, Y\. Sun, and L\. Hu\(2026\)SWE\-skills\-bench: do agent skills actually help in real\-world software engineering?\.arXiv preprint arXiv:2603\.15401\.Cited by:[§3\.4](https://arxiv.org/html/2608.20485#S3.SS4.p5.1)\.
- \[57\]J\. He, J\. Jia, C\. Liu, C\. Xue, Y\. Song, X\. Yang, and D\. Sun\(2026\)ProcBench: evaluating process\-level defects and control preservation in llm coding agents\.arXiv preprint arXiv:2605\.20251\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p5.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.5.3.1.1),[§7\.2](https://arxiv.org/html/2608.20485#S7.SS2.p1.1)\.
- \[58\]R\. Hu, X\. Wang, C\. Peng, C\. Gao, and D\. Lo\(2026\)Evaluating llm\-based 0\-to\-1 software generation in end\-to\-end cli tool scenarios\.arXiv preprint arXiv:2604\.06742\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[59\]X\. Hu, T\. Xiong, B\. Yi, Z\. Wei, R\. Xiao, Y\. Chen, J\. Ye, M\. Tao, X\. Zhou, Z\. Zhao,et al\.\(2025\)Os agents: a survey on mllm\-based agents for general computing devices use\.arXiv preprint arXiv:2508\.04482\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p1.1),[§1](https://arxiv.org/html/2608.20485#S1.p3.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p4.1)\.
- \[60\]Z\. Hua, Y\. Yao, W\. Xie, Y\. Zhao, M\. Liu, R\. Qiu, Z\. Huang, Z\. Wang, Y\. Ji, Y\. Ye,et al\.\(2026\)CLI\-universe: towards verifiable task synthesis engine for terminal agents\.arXiv preprint arXiv:2606\.22883\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[61\]R\. Hutter and M\. Pradel\(2026\)AgentStepper: interactive debugging of software development agents\.arXiv preprint arXiv:2602\.06593\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1)\.
- \[62\]H\. Ivison, J\. O\. Yin, R\. Shao, T\. Xiao, N\. Lambert, and H\. Hajishirzi\(2026\)Tmax: a simple recipe for terminal agents\.arXiv preprint arXiv:2606\.23321\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[63\]N\. Jain, J\. Singh, M\. Shetty, T\. Zhang, L\. Zheng, K\. Sen, and I\. Stoica\(2025\)R2e\-gym: procedural environment generation and hybrid verifiers for scaling open\-weights swe agents\.InSecond Conference on Language Modeling,Cited by:[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[64\]S\. Jha, R\. Arora, Y\. Watanabe, T\. Yanagawa, Y\. Chen, J\. Clark, B\. Bhavya, M\. Verma, H\. Kumar, H\. Kitahara,et al\.\(2025\)ITBench: evaluating ai agents across diverse real\-world it automation tasks\.arXiv preprint arXiv:2502\.05352\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1),[§5](https://arxiv.org/html/2608.20485#S5.p4.1)\.
- \[65\]S\. Jha, M\. Paltenghi, C\. Maddila, V\. Murali, S\. Ugare, and S\. Chandra\(2026\)REAP: automatic curation of coding agent benchmarks from interactive production usage\.arXiv preprint arXiv:2604\.01527\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p4.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.8.3.1.1),[§5](https://arxiv.org/html/2608.20485#S5.p4.1)\.
- \[66\]Z\. Ji, Z\. Zhang, C\. Xu, Z\. Li, Y\. Gao, S\. Wang, and S\. Cheung\(2026\)Coding agents are guessing: measuring action\-boundary violations in underspecified devops instructions\.arXiv preprint arXiv:2607\.02294\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1),[§7\.3](https://arxiv.org/html/2608.20485#S7.SS3.p1.1)\.
- \[67\]C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. Narasimhan\(2024\)Swe\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 54107–54157\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.2.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.3.1.1.1.1),[§5](https://arxiv.org/html/2608.20485#S5.p1.1),[§6](https://arxiv.org/html/2608.20485#S6.p2.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[68\]T\. Jin, Y\. Zhu, and D\. Kang\(2025\)Elt\-bench: an end\-to\-end benchmark for evaluating ai agents on elt pipelines\.Proceedings of the VLDB Endowment19\(2\),pp\. 84–98\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[69\]H\. Kang, T\. Suresh, J\. Saad\-Falcon, and A\. Mirhoseini\(2026\)TRACE: capability\-targeted agentic training\.arXiv preprint arXiv:2604\.05336\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p2.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p1.1),[§4](https://arxiv.org/html/2608.20485#S4.p1.1)\.
- \[70\]A\. Kaufman, J\. Lucassen, T\. Tracy, C\. Rushing, and A\. Bhatt\(2025\)BashArena: a control setting for highly privileged ai agents\.arXiv preprint arXiv:2512\.15688\.Cited by:[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.32.4.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p5.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.7.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.10.1.1.1.1),[§6\.1](https://arxiv.org/html/2608.20485#S6.SS1.p5.1),[Table 4](https://arxiv.org/html/2608.20485#S6.T4.5.5.1.1.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1),[§7\.3](https://arxiv.org/html/2608.20485#S7.SS3.p1.1)\.
- \[71\]G\. J\. Kim, A\. Wilf, L\. Morency, and D\. Fried\(2025\)From reproduction to replication: evaluating research agents with progressive code masking\.arXiv preprint arXiv:2506\.19724\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1)\.
- \[72\]P\. T\. J\. Kon, J\. Liu, Q\. Ding, Y\. Qiu, Z\. Yang, Y\. Huang, J\. Srinivasa, M\. Lee, M\. Chowdhury, and A\. Chen\(2025\)Curie: toward rigorous and automated scientific experimentation with ai agents\.arXiv preprint arXiv:2502\.16069\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[73\]P\. T\. J\. Kon, J\. Liu, X\. Zhu, Q\. Ding, J\. Peng, J\. Xing, Y\. Huang, Y\. Qiu, J\. Srinivasa, M\. Lee,et al\.\(2025\)Exp\-bench: can ai conduct ai research experiments?\.arXiv preprint arXiv:2505\.24785\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[74\]J\. Kuang, Y\. Li, X\. Zhang, Y\. Li, X\. Sun, Y\. Shen, P\. Yu,et al\.\(2026\)Process\-level trajectory evaluation for environment configuration in software engineering agents\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 113832–113855\.Cited by:[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p2.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p5.1),[§7\.2](https://arxiv.org/html/2608.20485#S7.SS2.p1.1)\.
- \[75\]T\. Kuntz, A\. Duzan, H\. Zhao, F\. Croce, Z\. Kolter, N\. Flammarion, and M\. Andriushchenko\(2026\)Os\-harm: a benchmark for measuring safety of computer use agents\.Advances in Neural Information Processing Systems38\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[76\]Y\. Lai, P\. Xia, H\. Ji, K\. Xiong, K\. Zeng, J\. Liu, F\. Wu, J\. Zhong, Z\. Zheng, C\. Xie,et al\.\(2026\)ClawForge: generating executable interactive benchmarks for command\-line agents\.arXiv preprint arXiv:2605\.14133\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.13.1.1.1)\.
- \[77\]N\. Le Hai, D\. M\. Nguyen, and N\. D\. Bui\(2025\)On the impacts of contexts on repository\-level code generation\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1496–1524\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p5.1)\.
- \[78\]D\. Lee, G\. Bae, and I\. Yun\(2026\)CTFusion: a ctf\-based benchmark for llm agent evaluation\.arXiv preprint arXiv:2605\.11504\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1),[§7\.3](https://arxiv.org/html/2608.20485#S7.SS3.p1.1)\.
- \[79\]Y\. Lee, R\. Nair, Q\. Zhang, K\. Lee, O\. Khattab, and C\. Finn\(2026\)Meta\-harness: end\-to\-end optimization of model harnesses\.arXiv preprint arXiv:2603\.28052\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.29.4.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.32.5.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p5.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p5.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p3.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p5.1),[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[80\]F\. Lei, J\. Meng, Y\. Huang, J\. Zhao, Y\. Zhang, J\. Luo, X\. Zou, R\. Yang, W\. Shi, Y\. Gao,et al\.\(2025\)DAComp: benchmarking data agents across the full data intelligence lifecycle\.arXiv preprint arXiv:2512\.04324\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[81\]K\. Li, J\. Shi, Y\. Xiao, M\. Jiang, J\. Sun, Y\. Wu, D\. Fu, S\. Xia, X\. Cai, T\. Xu,et al\.\(2026\)Agencybench: benchmarking the frontiers of autonomous agents in 1m\-token real\-world contexts\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 7422–7440\.Cited by:[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p2.1),[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1)\.
- \[82\]X\. Li, K\. W\. Choe, Y\. Liu, X\. Chen, C\. Tao, B\. You, W\. Chen, Z\. Di, J\. Sun, S\. Zheng,et al\.\(2026\)Clawsbench: evaluating capability and safety of llm productivity agents in simulated workspaces\.arXiv preprint arXiv:2604\.05172\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[83\]X\. Li, N\. Ben\-Israel, Y\. Raz, B\. Ahmed, D\. Serebro, and A\. Raux\(2026\)Repomod\-bench: a benchmark for code repository modernization via implementation\-agnostic testing\.arXiv preprint arXiv:2602\.22518\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1)\.
- \[84\]Z\. Li, Y\. Shi, Z\. Li, R\. Wang, A\. Li, Z\. Huang, J\. Yang, L\. Ke, N\. Liu, H\. Mi,et al\.\(2026\)Recursive synthesis for long\-horizon terminal tasks\.arXiv preprint arXiv:2608\.05466\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[85\]Z\. Li, Z\. Li, Y\. Shi, R\. Wang, J\. Yang, Z\. Liu, X\. Wu, A\. Li, Y\. Yu, N\. Liu,et al\.\(2026\)Long\-horizon\-terminal\-bench: testing the limits of agents on long\-horizon terminal tasks with dense reward\-based grading\.arXiv preprint arXiv:2607\.08964\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1)\.
- \[86\]J\. Liang, Z\. Lyu, Z\. Liu, X\. Chen, P\. Nie, K\. Zou, and W\. Chen\(2026\)SWE\-next: scalable real\-world software engineering tasks for agents\.arXiv preprint arXiv:2603\.20691\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1)\.
- \[87\]J\. Lin, S\. Liu, C\. Pan, L\. Lin, S\. Dou, X\. Huang, H\. Yan, Z\. Han, and T\. Gui\(2026\)Agentic harness engineering: observability\-driven automatic evolution of coding\-agent harnesses\.arXiv preprint arXiv:2604\.25850\.Cited by:[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p5.1)\.
- \[88\]X\. Lin, J\. Zhang, G\. Deng, T\. Liu, T\. Zhang, Q\. Guo, and R\. Chen\(2025\)Ircopilot: automated incident response with large language models\.arXiv preprint arXiv:2505\.20945\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[89\]Y\. Lin, H\. Wang, S\. Wu, L\. Fan, F\. Pan, S\. Zhao, and D\. Tu\(2026\)CLI\-gym: scalable cli task generation via agentic environment inversion\.arXiv preprint arXiv:2602\.10999\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.30.3.1),[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.8.1.1.1)\.
- \[90\]X\. Liu, H\. Yu, H\. Zhang, Y\. Xu, X\. Lei, H\. Lai, Y\. Gu, H\. Ding, K\. Men, K\. Yang,et al\.\(2024\)Agentbench: evaluating llms as agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 52989–53046\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1)\.
- \[91\]Y\. Liu, L\. Zhang, F\. Liu, P\. Lin, and X\. Li\(2026\)A scalable benchmark for repository\-oriented long\-horizon conversational context management\.arXiv preprint arXiv:2603\.06358\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.6.3.1.1)\.
- \[92\]Z\. Liu, L\. Huang, J\. Zhang, D\. Liu, Y\. Tian, and J\. Shao\(2025\)PACEbench: a framework for evaluating practical ai cyber\-exploitation capabilities\.arXiv preprint arXiv:2510\.11688\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[93\]X\. Lou, M\. Lázaro\-Gredilla, A\. Dedieu, C\. Wendelken, W\. Lehrach, and K\. P\. Murphy\(2026\)Autoharness: improving llm agents by automatically synthesizing a code harness\.arXiv preprint arXiv:2603\.03329\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p5.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p5.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p3.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p5.1),[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[94\]J\. Lu, T\. Holleis, Y\. Zhang, B\. Aumayer, F\. Nan, H\. Bai, S\. Ma, S\. Ma, M\. Li, G\. Yin,et al\.\(2025\)Toolsandbox: a stateful, conversational, interactive evaluation benchmark for llm tool use capabilities\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 1160–1183\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1)\.
- \[95\]P\. Lu, S\. Zhang, Y\. Hou, L\. Ye, C\. Huang, Z\. Chen, J\. Zeng, H\. Jiang, P\. Liu, Y\. Wang,et al\.\(2026\)ProjDevBench: benchmarking ai coding agents on end\-to\-end project development\.arXiv preprint arXiv:2602\.01655\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1)\.
- \[96\]R\. Marchand, A\. O\. Cathain, J\. Wynne, P\. M\. Giavridis, S\. Deverett, J\. Wilkinson, J\. Gwartz, and H\. Coppock\(2026\)Quantifying frontier llm capabilities for container sandbox escape\.arXiv preprint arXiv:2603\.02277\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[97\]S\. Mateega, J\. Yang, T\. Costello, S\. Jadhav, N\. Tian, and A\. Garcinuño\(2026\)IDE\-bench: evaluating large language models as ide agents on real\-world software engineering tasks\.arXiv preprint arXiv:2601\.20886\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1)\.
- \[98\]T\. Mehtiyev and W\. Assunção\(2026\)Beyond resolution rates: behavioral drivers of coding agent success and failure\.arXiv preprint arXiv:2604\.02547\.Cited by:[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p1.1)\.
- \[99\]K\. Mei, X\. Zhu, W\. Xu, W\. Hua, M\. Jin, Z\. Li, S\. Xu, R\. Ye, Y\. Ge, and Y\. Zhang\(2024\)Aios: llm agent operating system\.arXiv preprint arXiv:2403\.16971\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[100\]F\. Meng, G\. Chen, J\. Zhao, S\. Sun, Z\. Lin, W\. X\. Zhao, R\. Song, J\. Wen, and K\. Jia\(2026\)CalibForge: adversarial solver calibration for scaling learnable terminal tasks\.arXiv preprint arXiv:2608\.06352\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[101\]M\. A\. Merrill, A\. G\. Shaw, N\. Carlini, B\. Li, H\. Raj, I\. Bercovich, L\. Shi, J\. Y\. Shin, T\. Walshe, E\. K\. Buchanan,et al\.\(2026\)Terminal\-bench: benchmarking agents on hard, realistic tasks in command line interfaces\.arXiv preprint arXiv:2601\.11868\.Cited by:[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.31.2.1),[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p3.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p4.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.3.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.5.1.1.1.1),[§5](https://arxiv.org/html/2608.20485#S5.p1.1),[§6\.1](https://arxiv.org/html/2608.20485#S6.SS1.p5.1),[Table 4](https://arxiv.org/html/2608.20485#S6.T4.5.2.1.1.1)\.
- \[102\]R\. Nakamura and K\. Eguchi\(2026\)How helpful is llm assistance in network operations? a case study at a large demonstration network\.arXiv preprint arXiv:2605\.19627\.Cited by:[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[103\]J\. Nam, J\. Yoon, J\. Chen, J\. Shin, S\. Arik, and T\. Pfister\(2026\)Mle\-star: machine learning engineering agent via search and targeted refinement\.Advances in Neural Information Processing Systems38,pp\. 116692–116712\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[104\]A\. Nangia, S\. Mishra, A\. Gokrani, and P\. Chopra\(2026\)ISO\-bench: can coding agents optimize real\-world inference workloads?\.arXiv preprint arXiv:2602\.19594\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[105\]Z\. Ni, H\. Wang, S\. Zhang, S\. Lu, Z\. He, Z\. Tang, S\. Hu, B\. Li, C\. Hu, B\. Jiao,et al\.\(2026\)Gittaskbench: a benchmark for code agents solving real\-world tasks through code repository leveraging\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.40,pp\. 32564–32572\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.7.1.1.1.1)\.
- \[106\]OpenAI\(2025\)Codex CLI\.Note:Open\-source terminal coding agent that runs locally and supports command\-line coding workflows\.External Links:[Link](https://github.com/openai/codex)Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1)\.
- \[107\]G\. Orlanski, D\. Roy, A\. Yun, C\. Shin, A\. Gu, A\. Ge, D\. Adila, N\. Roberts, F\. Sala, and A\. Albarghouthi\(2026\)SlopCodeBench: benchmarking how coding agents degrade over long\-horizon iterative tasks\.arXiv preprint arXiv:2603\.24755\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p1.1)\.
- \[108\]A\. Pabba, A\. Mathai, A\. Chakraborty, and B\. Ray\(2025\)Semagent: a semantics aware program repair agent\.arXiv preprint arXiv:2506\.16650\.Cited by:[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p2.1)\.
- \[109\]H\. Padigela, C\. Shah, and D\. Juyal\(2025\)Ml\-dev\-bench: comparative analysis of ai agents on ml development workflows\.arXiv preprint arXiv:2502\.00964\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[110\]J\. Pan, X\. Wang, G\. Neubig, N\. Jaitly, H\. Ji, A\. Suhr, and Y\. Zhang\(2024\)Training software engineering agents and verifiers with swe\-gym\.arXiv preprint arXiv:2412\.21139\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p1.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[111\]A\. Parisi, Y\. Zhao, and N\. Fiedel\(2022\)Talm: tool augmented language models\.arXiv preprint arXiv:2205\.12255\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p2.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p4.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p2.1)\.
- \[112\]K\. Parthasarathy, K\. Vaidhyanathan, R\. Dhar, V\. Krishnamachari, A\. Kakran, S\. Akshathala, S\. Arun, A\. Karan, B\. Muhammed, S\. Dubey,et al\.\(2025\)Engineering llm powered multi\-agent framework for autonomous cloudops\.In2025 IEEE/ACM 4th International Conference on AI Engineering–Software Engineering for AI \(CAIN\),pp\. 201–211\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p4.1)\.
- \[113\]H\. N\. Phan, T\. N\. Nguyen, P\. X\. Nguyen, and N\. D\. Bui\(2024\)Hyperagent: generalist software engineering agents to solve coding tasks at scale\.arXiv preprint arXiv:2409\.16299\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1)\.
- \[114\]R\. Pi, G\. Lam, M\. Shoeybi, P\. Jannaty, B\. Catanzaro, and W\. Ping\(2026\)On data engineering for scaling llm terminal capabilities\.arXiv preprint arXiv:2602\.21193\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p1.1),[§4](https://arxiv.org/html/2608.20485#S4.p1.1)\.
- \[115\]G\. Pinna, J\. Gong, D\. Williams, and F\. Sarro\(2026\)Comparing ai coding agents: a task\-stratified analysis of pull request acceptance\.InProceedings of the 23rd International Conference on Mining Software Repositories,pp\. 792–796\.Cited by:[§5](https://arxiv.org/html/2608.20485#S5.p4.1)\.
- \[116\]R\. Qiang, Y\. Zhuang, Y\. Li, D\. Sagar VK, R\. Zhang, C\. Li, I\. Wong, S\. Yang, P\. Liang, C\. Zhang,et al\.\(2026\)Mle\-dojo: interactive environments for empowering llm agents in machine learning engineering\.Advances in Neural Information Processing Systems38\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[117\]E\. Rabinovich and A\. A\. Tavor\(2025\)On the robustness of agentic function calling\.InProceedings of the 5th Workshop on Trustworthy NLP \(TrustNLP 2025\),pp\. 298–304\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p2.1)\.
- \[118\]M\. Rahman, M\. S\. Islam, R\. Mahbub, M\. T\. R\. Laskar, S\. Joty, and E\. H\. Prince\(2026\)DSAgentBench: can agents automate end\-to\-end data\-science workflows in real computer environments?\.External Links:2608\.10366,[Link](https://arxiv.org/abs/2608.10366)Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[119\]M\. S\. Rashid, C\. Bock, Y\. Zhuang, A\. Buchholz, T\. Esler, S\. Valentin, L\. Franceschi, M\. Wistuba, P\. T\. Sivaprasad, W\. J\. Kim,et al\.\(2025\)Swe\-polybench: a multi\-language benchmark for repository level evaluation of coding agents\.arXiv preprint arXiv:2504\.08703\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p1.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.2.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.3.1.1.1.1)\.
- \[120\]C\. Rawles, S\. Clinckemaillie, Y\. Chang, J\. Waltz, G\. Lau, M\. Fair, A\. Li, W\. Bishop, W\. Li, F\. Campbell\-Ajala,et al\.\(2025\)Androidworld: a dynamic benchmarking environment for autonomous agents\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 406–441\.Cited by:[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1)\.
- \[121\]J\. Ren, S\. Wu, Y\. Li, K\. Zhu, S\. Xu, B\. Feng, R\. Yuan, W\. Zhang, R\. Batista\-Navarro, J\. Yang,et al\.\(2026\)A self\-evolving framework for efficient terminal agents via observational context compression\.arXiv preprint arXiv:2604\.19572\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p2.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[122\]R\. Robbes, T\. Matricon, T\. Degueule, A\. Hora, and S\. Zacchiroli\(2026\)Agentic much? adoption of coding agents on github\.ACM Transactions on Software Engineering and Methodology\.Cited by:[§5](https://arxiv.org/html/2608.20485#S5.p4.1)\.
- \[123\]T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. Scialom\(2023\)Toolformer: language models can teach themselves to use tools\.Advances in neural information processing systems36,pp\. 68539–68551\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p2.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p4.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p2.1)\.
- \[124\]M\. Seyedkazemi Ardebili and A\. Bartolini\(2026\)KubeIntellect: a modular llm\-orchestrated agent framework for end\-to\-end kubernetes management: ms ardebili, a\. bartolini\.Journal of Grid Computing24\(3\),pp\. 17\.Cited by:[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[125\]J\. She\(2026\)AgentRM: an os\-inspired resource manager for llm agent systems\.arXiv preprint arXiv:2603\.13110\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[126\]Q\. Shen, Z\. Huang, V\. Kamanuru, A\. Aliev, J\. Rainton, A\. Awelkair, Z\. Zeng, J\. Li, S\. Dong, Y\. Yuan,et al\.\(2026\)SETA: scaling environments for terminal agents\.arXiv preprint arXiv:2607\.10891\.Cited by:[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1)\.
- \[127\]Y\. Shi, W\. Yu, J\. Huang, W\. Yao, W\. Chen, and N\. Liu\(2025\)Towards trustworthy gui agents: a survey\.arXiv preprint arXiv:2503\.23434\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[128\]V\. Shrivastava, P\. Kauffmann, A\. Awadallah, and D\. Papailiopoulos\(2026\)Echo: terminal agents learn world models for free\.arXiv preprint arXiv:2605\.24517\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[129\]M\. L\. Siddiq, X\. Zhao, V\. C\. Lopes, B\. Casey, and J\. Santos\(2026\)Security in the age of ai teammates: an empirical study of agentic pull requests on github\.arXiv preprint arXiv:2601\.00477\.Cited by:[§5](https://arxiv.org/html/2608.20485#S5.p4.1)\.
- \[130\]R\. Siva, K\. Cheung, L\. Li, and G\. Sundaram\(2026\)KRAIG: a natural language\-driven agent for automated dataops pipeline generation\.arXiv preprint arXiv:2603\.20311\.Cited by:[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[131\]H\. Song, L\. Huang, S\. Sun, J\. Jiang, R\. Le, D\. Cheng, G\. Chen, Y\. Hu, Z\. Chen, Y\. Jia,et al\.\(2026\)Swe\-master: unleashing the potential of software engineering agents via post\-training\.arXiv preprint arXiv:2602\.03411\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[132\]K\. Song and Y\. Qi\(2026\)ANCHOR: automated alignment auditing for cli agents on real\-world harm\.arXiv preprint arXiv:2607\.10455\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[133\]T\. Song\(2026\)Cross\-context verification: hierarchical detection of benchmark contamination through session\-isolated analysis\.arXiv preprint arXiv:2603\.21454\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p3.1)\.
- \[134\]H\. Sun, T\. Yu, S\. Ma, Q\. Zhang, L\. Rao, C\. Tian,et al\.\(2026\)ATime\-consistent benchmark for repository\-level software engineering evaluation\.arXiv preprint arXiv:2603\.26137\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p3.1)\.
- \[135\]Q\. Sun, Z\. Liu, C\. Ma, Z\. Ding, F\. Xu, Z\. Yin, H\. Zhao, Z\. Wu, K\. Cheng, Z\. Liu,et al\.\(2026\)Scienceboard: evaluating multimodal autonomous agents in realistic scientific workflows\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 75694–75731\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p1.1)\.
- \[136\]W\. Sun, M\. Lu, Z\. Ling, K\. Liu, X\. Yao, Y\. Yang, and J\. Chen\(2025\)Scaling long\-horizon llm agent via context\-folding\.arXiv preprint arXiv:2510\.11967\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[137\]H\. Tao, Y\. Zhang, Z\. Tang, H\. Peng, X\. Zhu, B\. Liu, Y\. Yang, Z\. Zhang, Z\. Xu, H\. Zhang,et al\.\(2026\)Code graph model \(cgm\): a graph\-integrated large language model for repository\-level software engineering tasks\.Advances in Neural Information Processing Systems38,pp\. 15869–15909\.Cited by:[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p2.1)\.
- \[138\]B\. Teamet al\.\(2024\)LiveSQLBench: a dynamic and contamination\-free benchmark for evaluating llms on real\-world text\-to\-sql tasks\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.11.1.1.1)\.
- \[139\]M\. V\. Thai, T\. Le, D\. N\. Manh, H\. P\. Nhat, and N\. D\. Bui\(2025\)SWE\-evo: benchmarking coding agents in long\-horizon software evolution scenarios\.arXiv preprint arXiv:2512\.18470\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1)\.
- \[140\]K\. Thangarajah, B\. Chen, and A\. E\. Hassan\(2026\)DCAS: decoupling cli agent scaffolding to internalize planning across scaffolds\.arXiv preprint arXiv:2608\.06113\.Cited by:[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[141\]TheR1D\(2026\)ShellGPT\.Note:Command\-line productivity tool powered by large language models for generating shell commands, code snippets, and documentation\.External Links:[Link](https://github.com/TheR1D/shell_gpt)Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p5.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p2.1)\.
- \[142\]M\. Thompson\(2026\)The dual\-state architecture for reliable llm agents\.External Links:2512\.20660,[Link](https://arxiv.org/abs/2512.20660)Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[143\]H\. Trivedi, T\. Khot, M\. Hartmann, R\. Manku, V\. Dong, E\. Li, S\. Gupta, A\. Sabharwal, and N\. Balasubramanian\(2024\)Appworld: a controllable world of apps and people for benchmarking interactive coding agents\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 16022–16076\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.5.3.1.1)\.
- \[144\]A\. Twabi, Y\. Ding, and T\. Kondo\(2026\)NetAgentBench: a state\-centric benchmark for evaluating agentic network configuration\.arXiv preprint arXiv:2604\.09678\.Cited by:[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1)\.
- \[145\]N\. Vats and O\. Golev\(2026\)The scaffold effect in coding agents: harness choice as a hidden variable in coding\-agent evaluation\.arXiv preprint arXiv:2607\.22585\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p3.1),[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[146\]H\. Wang, J\. Gong, H\. Zhang, J\. Xu, and Z\. Wang\(2025\)Ai agentic programming: a survey of techniques, challenges, and opportunities\.arXiv preprint arXiv:2508\.11126\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[147\]L\. Wang, C\. Ma, X\. Feng, Z\. Zhang, H\. Yang, J\. Zhang, Z\. Chen, J\. Tang, X\. Chen, Y\. Lin,et al\.\(2024\)A survey on large language model based autonomous agents\.Frontiers of Computer Science18\(6\),pp\. 186345\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[148\]R\. Wang, R\. A\. Genadi, B\. E\. Bouardi, Y\. Wang, F\. Koto, Z\. Liu, T\. Baldwin, and H\. Li\(2025\)Agentfly: extensible and scalable reinforcement learning for lm agents\.arXiv preprint arXiv:2507\.14897\.Cited by:[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p5.1)\.
- \[149\]X\. Wang, Y\. Chen, L\. Yuan, Y\. Zhang, Y\. Li, H\. Peng, and H\. Ji\(2024\)Executable code actions elicit better llm agents\.arXiv preprint arXiv:2402\.01030\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p1.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p2.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p2.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p2.1)\.
- \[150\]X\. Wang, S\. Rosenberg, J\. Michelini, C\. Smith, H\. Tran, E\. Nyst, R\. Malhotra, X\. Zhou, V\. Chen, R\. Brennan,et al\.\(2025\)The openhands software agent sdk: a composable and extensible foundation for production agents\.arXiv preprint arXiv:2511\.03690\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§C\.1](https://arxiv.org/html/2608.20485#A3.SS1.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.29.3.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p5.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p5.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p3.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1),[Table 5](https://arxiv.org/html/2608.20485#S6.T5.5.5.1.1),[Table 5](https://arxiv.org/html/2608.20485#S6.T5.5.9.1.1),[§6](https://arxiv.org/html/2608.20485#S6.p2.1),[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[151\]Y\. Wang, W\. Zhong, Y\. Huang, E\. Shi, M\. Yang, J\. Chen, H\. Li, Y\. Ma, Q\. Wang, and Z\. Zheng\(2025\)Agents in software engineering: survey, landscape, and vision\.Automated Software Engineering32\(2\),pp\. 70\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[152\]B\. Wei, Y\. Zhang, J\. Pan, K\. Mei, X\. Wang, J\. Hamm, Z\. Zhu, and Y\. Ge\(2026\)ClawSafety:" safe" llms, unsafe agents\.arXiv preprint arXiv:2604\.01438\.Cited by:[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p2.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p5.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.7.3.1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.10.1.1.1.1),[§7\.3](https://arxiv.org/html/2608.20485#S7.SS3.p1.1)\.
- \[153\]Y\. Wei, O\. Duchenne, J\. Copet, Q\. Carbonneaux, L\. Zhang, D\. Fried, G\. Synnaeve, R\. Singh, and S\. Wang\(2026\)Swe\-rl: advancing llm reasoning via reinforcement learning on open software evolution\.Advances in Neural Information Processing Systems38,pp\. 78500–78525\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[154\]J\. Wu, M\. Hu, J\. Zhu, J\. Pan, Y\. Liu, M\. Xu, and Y\. Jin\(2025\)Git context controller: manage the context of llm\-based agents like git\.arXiv preprint arXiv:2508\.00031\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p5.1)\.
- \[155\]S\. Wu, Y\. Li, Y\. Song, W\. Zhang, Y\. Wang, R\. Batista\-Navarro, X\. Yang, M\. Tang, B\. Dai, J\. Yang,et al\.\(2026\)Large\-scale terminal agentic trajectory generation from dockerized environments\.arXiv preprint arXiv:2602\.01244\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.30.2.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p1.1),[§4](https://arxiv.org/html/2608.20485#S4.p1.1)\.
- \[156\]Z\. Xi, W\. Chen, X\. Guo, W\. He, Y\. Ding, B\. Hong, M\. Zhang, J\. Wang, S\. Jin, E\. Zhou,et al\.\(2025\)The rise and potential of large language model based agents: a survey\.Science China Information Sciences68\(2\),pp\. 121101\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[157\]C\. S\. Xia, Y\. Deng, S\. Dunn, and L\. Zhang\(2025\)Demystifying llm\-based software engineering agents\.Proceedings of the ACM on Software Engineering2\(FSE\),pp\. 801–824\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px1.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.28.3.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p5.1),[§3\.4](https://arxiv.org/html/2608.20485#S3.SS4.p5.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p2.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p2.1)\.
- \[158\]C\. S\. Xia, Z\. Wang, Y\. Yang, Y\. Wei, and L\. Zhang\(2025\)Live\-swe\-agent: can software engineering agents self\-evolve on the fly?\.arXiv preprint arXiv:2511\.13646\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1)\.
- \[159\]T\. Xie, D\. Zhang, J\. Chen, X\. Li, S\. Zhao, R\. Cao, T\. J\. Hua, Z\. Cheng, D\. Shin, F\. Lei,et al\.\(2024\)Osworld: benchmarking multimodal agents for open\-ended tasks in real computer environments\.Advances in Neural Information Processing Systems37,pp\. 52040–52094\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px1.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.28.4.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p5.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p4.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p2.1)\.
- \[160\]Z\. Xiu, D\. Q\. Sun, K\. Cheng, M\. Patel, Y\. Zhang, J\. Lu, O\. Attia, R\. Vemulapalli, O\. Tuzel, M\. Cao,et al\.\(2026\)ASTRA\-bench: evaluating tool\-use agent reasoning and action planning with personal user context\.arXiv preprint arXiv:2603\.01357\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1)\.
- \[161\]A\. Xu, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Ling,et al\.\(2026\)Deepseek\-v4: towards highly efficient million\-token context intelligence\.arXiv preprint arXiv:2606\.19348\.Cited by:[§C\.1](https://arxiv.org/html/2608.20485#A3.SS1.p1.1),[§C\.3](https://arxiv.org/html/2608.20485#A3.SS3.p1.1),[§6\.1](https://arxiv.org/html/2608.20485#S6.SS1.p3.1),[§6](https://arxiv.org/html/2608.20485#S6.p2.1)\.
- \[162\]T\. Xu, Y\. Chen, and M\. Li\(2026\)CLEANER: self\-purified trajectories boost agentic reinforcement learning\.arXiv preprint arXiv:2601\.15141\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p2.1),[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p1.1)\.
- \[163\]L\. Yan, X\. Chen, and X\. Zhang\(2026\)When the specification emerges: benchmarking faithfulness loss in long\-horizon coding agents\.arXiv preprint arXiv:2603\.17104\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1),[§5\.4](https://arxiv.org/html/2608.20485#S5.SS4.p1.1)\.
- \[164\]J\. Yang, C\. E\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. Narasimhan, and O\. Press\(2024\)Swe\-agent: agent\-computer interfaces enable automated software engineering\.Advances in Neural Information Processing Systems37,pp\. 50528–50652\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[§C\.1](https://arxiv.org/html/2608.20485#A3.SS1.p1.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.28.2.1),[Figure 2](https://arxiv.org/html/2608.20485#S1.F2.pic1.29.2.1),[§1](https://arxiv.org/html/2608.20485#S1.p4.1),[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p1.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p2.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p5.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p1.1),[§2\.3](https://arxiv.org/html/2608.20485#S2.SS3.p4.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p3.1),[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p2.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§3\.5](https://arxiv.org/html/2608.20485#S3.SS5.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p1.1),[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p5.1),[§5](https://arxiv.org/html/2608.20485#S5.p1.1),[Table 5](https://arxiv.org/html/2608.20485#S6.T5.5.4.1.1),[Table 5](https://arxiv.org/html/2608.20485#S6.T5.5.8.1.1),[§6](https://arxiv.org/html/2608.20485#S6.p2.1),[§7\.1](https://arxiv.org/html/2608.20485#S7.SS1.p1.1),[§7\.4](https://arxiv.org/html/2608.20485#S7.SS4.p1.1)\.
- \[165\]J\. Yang, K\. Lieret, C\. Jimenez, A\. Wettig, K\. Khandpur, Y\. Zhang, B\. Hui, O\. Press, L\. Schmidt, and D\. Yang\(2026\)Swe\-smith: scaling data for software engineering agents\.Advances in Neural Information Processing Systems38\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[166\]J\. Yang, A\. Prabhakar, K\. Narasimhan, and S\. Yao\(2023\)Intercode: standardizing and benchmarking interactive coding with execution feedback\.Advances in Neural Information Processing Systems36,pp\. 23826–23854\.Cited by:[§2\.1](https://arxiv.org/html/2608.20485#S2.SS1.p2.1),[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p2.1)\.
- \[167\]S\. Yang, C\. Tao, J\. Chen, T\. Yu, R\. Wang, Y\. Jiang, Y\. Du, W\. Xu, J\. Xiong, T\. Wu,et al\.\(2026\)What makes interaction trajectories effective for training terminal agents?\.arXiv preprint arXiv:2606\.03461\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1)\.
- \[168\]W\. Yang, C\. Song, X\. Li, D\. Ganguly, C\. Ma, S\. Wang, Z\. Dou, Y\. Zhou, V\. Chaudhary, and X\. Han\(2026\)ACE\-bench: agent configurable evaluation with scalable horizons and controllable difficulty under lightweight environments\.arXiv e\-prints,pp\. arXiv–2604\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1)\.
- \[169\]Z\. Yang, S\. Wang, K\. Fu, W\. He, W\. Xiong, Y\. Liu, Y\. Miao, B\. Gao, Y\. Wang, Y\. Ma,et al\.\(2025\)Kimi\-dev: agentless training as skill prior for swe\-agents\.arXiv preprint arXiv:2509\.23045\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[170\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\(2022\)React: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p1.1),[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p2.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p2.1)\.
- \[171\]B\. Ye, R\. Li, Q\. Yang, Y\. Liu, L\. Yao, H\. Lv, Z\. Xie, C\. An, L\. Li, L\. Kong,et al\.\(2026\)Claw\-eval: towards trustworthy evaluation of autonomous agents\.arXiv preprint arXiv:2604\.06132\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p4.1)\.
- \[172\]A\. Yehudai, L\. Eden, A\. Li, G\. Uziel, Y\. Zhao, R\. Bar\-Haim, A\. Cohan, and M\. Shmueli\-Scheuer\(2025\)Survey on evaluation of llm\-based agents\.arXiv preprint arXiv:2503\.16416\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[173\]Z\. Yu, B\. Wang, P\. Zeng, H\. Zhang, J\. Zhang, Z\. Wang, L\. Gao, J\. Song, N\. Sebe, and H\. T\. Shen\(2025\)A survey on efficient vision\-language\-action models\.arXiv preprint arXiv:2510\.24795\.Cited by:[§1](https://arxiv.org/html/2608.20485#S1.p3.1)\.
- \[174\]X\. Yuan, M\. M\. Moss, C\. E\. Feghali, C\. Singh, D\. Moldavskaya, D\. MacPhee, L\. Caccia, M\. Pereira, M\. Kim, A\. Sordoni,et al\.\(2025\)Debug\-gym: a text\-based environment for interactive debugging\.arXiv preprint arXiv:2503\.21557\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p1.1),[§5\.5](https://arxiv.org/html/2608.20485#S5.SS5.p2.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.9.1.1.1.1)\.
- \[175\]D\. Zan, Z\. Huang, W\. Liu, H\. Chen, S\. Xin, L\. Zhang, Q\. Liu, L\. Aoyan, L\. Chen, X\. Zhong,et al\.\(2026\)Multi\-swe\-bench: a multilingual benchmark for issue resolving\.Advances in Neural Information Processing Systems38\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1)\.
- \[176\]J\. Zeng, D\. Fu, T\. Mi, Y\. Zhuang, Y\. Huang, X\. Li, L\. Ye, M\. Xie, Q\. Hua, Z\. Huang,et al\.\(2026\)Davinci\-dev: agent\-native mid\-training for software engineering\.arXiv preprint arXiv:2601\.18418\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p2.1)\.
- \[177\]Y\. Zeng, S\. Li, D\. Dong, R\. Xu, Z\. Chen, L\. Zheng, Y\. Li, Z\. Zhou, H\. Zhao, L\. Tian,et al\.\(2026\)SWE\-hub: a unified production system for scalable, executable software engineering tasks\.arXiv preprint arXiv:2603\.00575\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1)\.
- \[178\]B\. Zhang, J\. Zhu, Z\. Shi, D\. Liu, and R\. Tang\(2026\)AgentForesight: online auditing for early failure prediction in multi\-agent systems\.arXiv preprint arXiv:2605\.08715\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p2.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[179\]J\. Zhang, L\. Ma, Y\. Li, F\. Wan, D\. Qi, X\. Zhao, J\. Hou, Z\. Xie, M\. Ren, X\. Wu,et al\.\(2026\)DockSmith: scaling reliable coding environments via an agentic docker builder\.arXiv preprint arXiv:2602\.00592\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p3.1)\.
- \[180\]C\. Zhao, S\. Zhang, Y\. Lin, W\. Gu, Z\. Chen, Y\. Sun, D\. Pei, C\. Bansal, S\. Rajmohan, and M\. Ma\(2026\)Debugging the debuggers: failure\-anchored structured recovery for software engineering agents\.arXiv preprint arXiv:2605\.08717\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p2.1)\.
- \[181\]X\. Zhao, H\. Li, S\. Li, T\. Zhao, E\. T\. Barr, F\. Sarro, and H\. Ye\(2026\)Failure as a process: an anatomy of cli coding agent trajectories\.arXiv preprint arXiv:2607\.09510\.Cited by:[§5\.6](https://arxiv.org/html/2608.20485#S5.SS6.p2.1)\.
- \[182\]J\. Zheng, X\. Cai, Q\. Li, D\. Zhang, Z\. Li, Y\. Zhang, L\. Song, and Q\. Ma\(2025\)Lifelongagentbench: evaluating llm agents as lifelong learners\.arXiv preprint arXiv:2505\.11942\.Cited by:[§5\.2](https://arxiv.org/html/2608.20485#S5.SS2.p3.1),[Table 2](https://arxiv.org/html/2608.20485#S5.T2.5.6.3.1.1)\.
- \[183\]M\. Zheng, K\. Han, B\. Li, H\. Xu, Y\. Tian, W\. He, H\. Zhou, J\. Guo, H\. Hu, L\. Ma,et al\.\(2026\)Claw\-swe\-bench: a benchmark for evaluating openclaw\-style agent harnesses on coding tasks\.arXiv preprint arXiv:2606\.12344\.Cited by:[§5\.1](https://arxiv.org/html/2608.20485#S5.SS1.p1.1),[Table 3](https://arxiv.org/html/2608.20485#S5.T3.5.4.1.1.1),[§6](https://arxiv.org/html/2608.20485#S6.p2.1)\.
- \[184\]Y\. Zheng, Y\. Hu, W\. Zhang, and A\. Quinn\(2025\)Towards agentic os: an llm agent framework for linux schedulers\.arXiv preprint arXiv:2509\.01245\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
- \[185\]Z\. Zhong, I\. Segal, I\. Bercovich, S\. Saxena, K\. Zhang, and A\. Raghunathan\(2026\)Hardening agent benchmarks with adversarial hacker\-fixer loops\.arXiv preprint arXiv:2606\.08960\.Cited by:[§5\.3](https://arxiv.org/html/2608.20485#S5.SS3.p2.1)\.
- \[186\]C\. Zhou, H\. Chai, W\. Chen, Z\. Guo, R\. Shan, Y\. Song, T\. Xu, Y\. Yang, A\. Yu, W\. Zhang,et al\.\(2026\)Externalization in llm agents: a unified review of memory, skills, protocols and harness engineering\.arXiv preprint arXiv:2604\.08224\.Cited by:[§4\.4](https://arxiv.org/html/2608.20485#S4.SS4.p2.1)\.
- \[187\]H\. Zhou, Y\. Chen, S\. Guo, X\. Yan, K\. H\. Lee, Z\. Wang, K\. Y\. Lee, G\. Zhang, K\. Shao, L\. Yang,et al\.\(2025\)Memento: fine\-tuning llm agents without fine\-tuning llms\.arXiv preprint arXiv:2508\.16153\.Cited by:[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[188\]S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.\(2024\)Webarena: a realistic web environment for building autonomous agents\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 15585–15606\.Cited by:[§2\.2](https://arxiv.org/html/2608.20485#S2.SS2.p5.1)\.
- \[189\]K\. Zhu, Y\. Nie, Y\. Li, Y\. Huang, J\. Wu, J\. Liu, X\. Sun, Z\. Yin, L\. Wang, Z\. Liu,et al\.\(2026\)Termigen: high\-fidelity environment and robust trajectory synthesis for terminal agents\.arXiv preprint arXiv:2602\.07274\.Cited by:[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px2.p1.1),[Appendix B](https://arxiv.org/html/2608.20485#A2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2608.20485#S3.SS1.p4.1),[§3\.3](https://arxiv.org/html/2608.20485#S3.SS3.p1.1),[§4\.1](https://arxiv.org/html/2608.20485#S4.SS1.p1.1),[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p1.1),[§4\.3](https://arxiv.org/html/2608.20485#S4.SS3.p3.1)\.
- \[190\]K\. Zhu, Z\. Liu, B\. Li, M\. Tian, Y\. Yang, J\. Zhang, P\. Han, Q\. Xie, F\. Cui, W\. Zhang,et al\.\(2025\)Where llm agents fail and how they can learn from failures\.arXiv preprint arXiv:2509\.25370\.Cited by:[§4\.2](https://arxiv.org/html/2608.20485#S4.SS2.p2.1)\.
- \[191\]H\. Zhuang, H\. Xing, and X\. Zhang\(2026\)AgentClick: a skill\-based human\-in\-the\-loop review layer for terminal ai agents\.InProceedings of the ACM Conference on AI and Agentic Systems,pp\. 1372–1378\.Cited by:[§3\.2](https://arxiv.org/html/2608.20485#S3.SS2.p4.1)\.
## APPENDIX
## Appendix AReview Protocol and Corpus Construction
This supplementary document reports the review protocol, extended taxonomy and coding materials, empirical methods and results, and representative trace cases\. We conducted astructured narrative review with an evidence\-stratified coded corpus\. The review combines explicit workload\-level scope rules, entry\-level analytical coding, and passage\-level evidence calibration to support a traceable synthesis of the reconciled corpus\.
### A\.1Corpus Scope and Composition
The review covers work released between 2022 and August 2026 on terminal agents, terminal\-mediated architectures, executable training environments, terminal\-native or repository\-based evaluation, process\-level agent assessment, runtime governance, and adjacent agent paradigms used for boundary comparison\. The current manuscript cites191 distinct entries:186 research entriesin the review corpus, four deployment\-facing tools \(Claude Code, Codex CLI, Aider, and Gemini CLI\), one CLI\-packaged assistant boundary comparator\. The latter six entries remain outside research statistics\. Candidate versions are consolidated at the work level so that the corpus counts a research contribution rather than each of its releases\.
Table[6](https://arxiv.org/html/2608.20485#A1.T6)summarizes corpus status, release year, and analytical use\. Analytical\-use counts are non\-exclusive because one work may support several parts of the survey\. This view separates entry\-level corpus composition from subset\-specific architecture, acquisition, and evaluation coding and from passage\-level evidence calibration\.
Table 6:Composition of the current review corpus\. Release\-year counts cover the 186 research entries; analytical\-use counts are non\-exclusive\.CategoryDescriptionCountCorpus statusResearch corpusResearch works retained for the survey synthesis\.186Engineering practice, boundary, and disclosure referencesFive engineering\-practice or boundary entries and one disclosure citation\.6Total cited setDistinct cited entries after work\-level version consolidation\.192Release year of research entries2022–2023Earliest retained foundations for tool use and interactive execution\.42024Expansion of executable agent and repository\-repair research\.142025Broadening benchmark, system, and deployment evidence\.562026 through AugustRecent architecture, acquisition, evaluation, and cross\-domain work\.112Analytical use of research entries, non\-exclusiveScope and characterizationSupports the execution\-substrate boundary and competence profile\.19System architectureSupports layers, families, control, memory, and runtime analysis\.44Competence acquisitionSupports environments, trajectories, supervision, and adaptation analysis\.41EvaluationSupports benchmark families, metrics, validity, and competence observability\.99Empirical settingSupports the fixed\-condition diagnostic design and its task settings\.9Challenges and implicationsSupports domain transfer, process evidence, governance, and attribution\.28
### A\.2Search, Screening, and Work\-Level Reconciliation
We assembled the corpus through iterative keyword search, venue\-oriented search, and backward and forward reference chaining, with the final update performed in August 2026\. Sources included arXiv, OpenReview, DBLP, IEEE Xplore, the ACM Digital Library, the ACL Anthology, major machine\-learning and software\-engineering venues, systems\-oriented venues, benchmark repositories, and project pages for deployment\-facing terminal\-agent tools\. Candidate versions were consolidated at the work level, with the most complete version retained and earlier versions linked to the same contribution rather than counted separately\.
Table[7](https://arxiv.org/html/2608.20485#A1.T7)summarizes the construction sequence and its analytical role\. The sequence also provides the protocol used for subsequent corpus updates\.
Table 7:Corpus\-construction workflow\. Each stage produces the input required by the next stage and constrains how evidence enters the synthesis\.StageProcedureRecorded decisionRole in synthesisScope freezeFix the 2022 to August 2026 interval, terminal\-mediated research object, adjacent comparators, and workload\-level boundary testsDate range, source classes, and inclusion boundaryPrevents surface\-level CLI mentions from entering as core evidenceRetrievalApply keyword families, venue\-oriented search, and backward and forward reference chainingSource, execution date, query form when retained, and work identityCombines named\-field retrieval with architecture, acquisition, evaluation, and governance routesWork reconciliationGroup versions of the same contribution and retain the most complete research accountPreferred version and version relationshipAvoids double counting preprint, conference, and journal releasesScreeningApply the five questions below and assign a corpus dispositionResearch, engineering\-practice, boundary, or exclusion dispositionSeparates progress\-bearing terminal evidence from surface\-level or adjacent evidenceAnalytical codingAssign manuscript roles to every retained entry and apply architecture, acquisition, or evaluation fields to the relevant subsetsEntry\-level roles and subset\-specific analytical codesConnects retained evidence to the architecture, acquisition, and evaluation synthesisClaim calibrationAssess convergence, directness, domain breadth, and evidence maturity at the passage levelBounded claim strength and scopeKeeps synthesis claims proportional to the available evidenceUpdate checkReapply retrieval, reconcile new versions, repeat screening, and revise affected synthesis passagesAdded, consolidated, retained, or removed statusKeeps later corpus updates aligned with the same decision sequenceFor subsequent updates, we specify seven Boolean query families in Table[8](https://arxiv.org/html/2608.20485#A1.T8)\. Source interfaces may require field\-name or quotation\-mark changes, while the Boolean concepts and date limits remain fixed\. Each new source\-specific execution records its date, query form, and hit count before work\-level reconciliation\.
Table 8:Boolean query families used for subsequent corpus updates\. Each source\-specific execution is logged with its date and hit count\.IDPurposeBoolean templateQ1Named field\("terminal agent" OR "terminal agents" OR "CLI agent" OR "command\-line agent"\)Q2Execution loop\("LLM agent" OR "language model agent"\) AND \(terminal OR shell OR CLI OR "command line"\) AND \(execution OR runtime OR feedback\)Q3SWE and harnesses\("software engineering agent" OR "coding agent"\) AND \(terminal OR shell OR command\) AND \(benchmark OR harness OR environment\)Q4Acquisition\("agent trajectory" OR "executable trajectory" OR "environment interaction"\) AND \(training OR post\-training OR "reinforcement learning"\) AND \(terminal OR CLI OR repository\)Q5Process evaluation\("terminal benchmark" OR "CLI benchmark" OR "agent benchmark"\) AND \(trace OR process OR recovery OR verification OR state\)Q6Governance\("LLM agent" OR "coding agent" OR "computer\-use agent"\) AND \(shell OR terminal OR "code execution"\) AND \(sandbox OR permission OR safety OR governance\)Q7Cross\-domain use\("LLM agent" OR "AI agent"\) AND \(terminal OR CLI OR "command execution"\) AND \(AIOps OR CloudOps OR DataOps OR cybersecurity OR "scientific workflow"\)Each candidate work was screened using five questions:
1. 1\.Does the agent execute terminal commands, operate CLI tools, or interact with a terminal\-mediated environment?
2. 2\.Does stdout, stderr, logs, diffs, return codes, or execution feedback materially shape subsequent actions?
3. 3\.Does the system produce real or simulated environment state changes through execution?
4. 4\.Does the work provide executable verification, trajectory data, a benchmark, an acquisition pipeline, or a runtime architecture relevant to terminal\-mediated execution or terminal\-agent systems?
5. 5\.What claim scope does the work support: core terminal\-agent evidence, terminal\-hybrid evidence, executable SWE\-adjacent evidence, boundary comparison, or background framing?
Based on these questions, each retained entry received a corpus status and one or more analytical roles\. Research entries support the architecture, acquisition, evaluation, empirical, or implication synthesis; engineering\-practice and boundary references illustrate deployed patterns or delimit adjacent systems\. The reported corpus counts begin after work\-level reconciliation because the iterative initial retrieval did not preserve a complete query\-level candidate count before deduplication\.
### A\.3Corpus Reconciliation and Update Protocol
The retained corpus can be inspected through the screening questions, work\-level reconciliation rules, analytical roles, and level\-specific fields in Table[9](https://arxiv.org/html/2608.20485#A1.T9)\. Subsequent updates additionally use the Boolean query families in Table[8](https://arxiv.org/html/2608.20485#A1.T8)\. Retrieval results are reconciled at the work level before assignment of corpus disposition and analytical roles\. Applying these rules yields the 186\-entry research corpus summarized in Table[6](https://arxiv.org/html/2608.20485#A1.T6); the six non\-corpus references remain outside research\-entry statistics\.
During an update, a work is added when it satisfies the screening questions and contributes evidence to at least one analytical role\. A new version replaces an earlier version when it provides the more complete account of the same contribution\. Versions representing the same contribution are consolidated rather than counted separately\. A retained work is removed from the review corpus when it no longer supports a synthesis passage after manuscript revision\. These rules keep corpus composition, analytical coding, and manuscript claims synchronized\.
### A\.4Coding Dimensions and Evidence Calibration
Entry\-level coding covers bibliographic metadata, research or engineering\-practice status, and manuscript analytical roles\. These fields support the corpus counts, release\-year distribution, and non\-exclusive analytical\-use counts in Table[6](https://arxiv.org/html/2608.20485#A1.T6)\. Additional conceptual fields are applied only to evidence subsets for which they are relevant, using the paper as the default unit and the system, benchmark, dataset, or pipeline when one paper introduces several distinct artifacts\.
For the main paper’s architecture analysis, relevant entries were coded by interface granularity, operational coupling, autonomy regime, recovery and control style, and planning or control strategy\. For the acquisition analysis, entries were coded by data\-source type, trajectory source, supervision signal, filtering or relabeling method, failure\-data treatment, and capability target\. For the evaluation synthesis and diagnostic study, entries were coded by task regime, harness assumptions, verification type, trace visibility, artifact availability, and coverage of the seven terminal\-competence dimensions\.
We calibrate evidence at the synthesis\-passage level rather than assigning a single maturity label to an entire paper\. Converging evidence across multiple benchmarks, systems, or domains supports stronger synthesis claims; direct but limited empirical evidence supports bounded claims; single\-study or indirect evidence supports provisional claims; and deployment practice or boundary examples support illustration\. Publication status and artifact availability remain descriptors of the relevant evidence subset rather than substitutes for claim\-level calibration\.
Table 9:Level\-specific coding schema used in corpus construction and synthesis\. Entry\-level fields cover every retained reference; specialized fields apply to relevant evidence subsets or synthesis passages\.Field groupCoded fieldsUnitAnalytical useCorpus identityBibliographic metadata, corpus status, and unique work identityWorkSupports corpus identity, release\-year counts, and work\-level deduplicationAnalytical roleScope, architecture, acquisition, evaluation, empirical setting, or implicationsWork, non\-exclusiveConnects retained entries to the survey passages they supportArchitectureInterface granularity, operational coupling, autonomy regime, recovery and control style, planning strategySystem or artifactSupports the layered architecture and recurring\-pattern synthesisAcquisitionData source, trajectory source, supervision signal, filtering or relabeling, failure\-data treatment, capability targetDataset, environment, or pipelineSupports comparison of executable data, learning, and adaptation routesEvaluationTask regime, harness assumptions, verification type, trace visibility, artifact availability, dimension coverageBenchmark or protocolSupports benchmark\-family and competence\-observability comparisonsEvidence calibrationConvergence, directness, domain breadth, and maturity of evidence for the passage claimSynthesis passageBounds the strength and generality of each synthesized claim
### A\.5Terminology and System\-Layer Roles
The working terminology used for the survey’s workload\-level boundary tests is organized as follows\. Each term is paired with its technical meaning and its analytical role\.
Terminal\.A textual input/output access layer or its emulation\. It is an established field label and a common access path to command execution, not itself the source of process state or exit semantics\.
Shell\.A command interpreter that parses commands, expands syntax, and starts programs\. It is a common but non\-required mediator between an agent interface and executable programs\.
CLI tool\.A program whose operations are invoked through textual arguments or commands\. It is an action target; surface\-level CLI exposure alone does not place a workload in scope\.
Command\-execution runtime\.The filesystem, processes, dependencies, permissions, resources, and operating\-system state in which commands take effect\. It is the state\-changing execution substrate that grounds actions, observations, and later verification\.
Harness\.The model\-facing layer that exposes actions, formats observations, manages context, invokes the runtime, and applies execution policy\. It allocates system responsibilities and changes what the model can observe or do\.
Terminal agent\.A system whose dominant progress\-bearing loop depends on command execution, textual feedback, and stateful environment interaction\. This is the survey’s workload\-level working characterization, applied through the three boundary tests\.
### A\.6Construction and Boundaries of the Seven\-Dimension Profile
The seven dimensions synthesize recurring functional responsibilities and their cross\-dimensional dependencies\. We first collected responsibilities and failure descriptions from architecture, acquisition, benchmark, and process\-evaluation studies\. We then grouped them by the object acted upon, the evidence needed for the next decision, and the response required when execution diverged from the task\. Candidate groupings were separated when they required distinct system mechanisms or observable evidence\. The resulting construction logic and representative trace signals are summarized below\.
Command and action formulation\.Select commands, arguments, edits, and executable sequences\. The separation rule concerns the action submitted before interpreting its consequence\. Representative failures are malformed invocation, wrong target, and semantically unsuitable action\. Commands and tool calls reveal form, while intent and semantic suitability may remain latent\.
Feedback and artifact interpretation\.Extract decision\-relevant evidence from outputs and artifacts\. The separation rule concerns how observed evidence informs the next decision\. Representative failures are ignored error, superficial reaction, and misread test or diff\. Action changes after feedback are visible, while usefulness requires contextual judgment\.
Runtime and environment management\.Construct and maintain dependencies, services, processes, and execution conditions\. The separation rule concerns the environment required for actions to run rather than the remembered task state\. Representative failures are dependency conflict, service failure, and unstable configuration\. Exit codes and setup actions expose friction, but successful exits do not prove a correct runtime\.
State, task, and context tracking\.Maintain facts about workspace, task, history, and unresolved goals\. The separation rule concerns the agent’s working account of persistent state across steps\. Representative failures are stale path, forgotten constraint, repeated loop, and inconsistent goal\. Contradictions and repeated errors can be traced, but complete internal state is not observable\.
Progress verification\.Design checks that establish intermediate validity or completion\. The separation rule concerns evidence that a claim or artifact satisfies a condition\. Representative failures are missing check, inadequate oracle, and unchecked requirement\. Tests and inspections expose verification activity, while timing or presence does not establish adequacy\.
Recovery and adaptation\.Diagnose divergence, revise strategy, retry, work around, or roll back\. This dimension begins after observed failure or violated expectation and targets renewed progress\. Representative failures are repeated ineffective retry, irrelevant workaround, and failed rollback\. Failure\-follow\-up windows expose attempts, while causal relevance and success require context\.
Governance and side\-effect control\.Respect authorization, containment, reversibility, credentials, and resource limits\. The separation rule concerns whether execution remains within policy independent of command success\. Representative failures are unauthorized deletion, secret exposure, sandbox escape, and uncontrolled side effect\. Sensitive events trigger review, while task authorization is needed to determine a violation\.
We distinguish runtime and environment management from state, task, and context tracking because the former constructs executable conditions while the latter maintains an accurate account of those conditions and the task\. We distinguish progress verification from recovery and adaptation because a check establishes evidence about state or completion, whereas recovery acts on observed divergence\. Long\-horizon persistence is treated as a cross\-dimensional outcome supported by state and context tracking, verification, and recovery, while governance is distributed across model behavior, harness policy, and runtime containment\.
## Appendix BExtended Scope and Taxonomy
The main paper retains the figures and tables needed for the survey’s central argument\. The compact descriptions below preserve the supporting scope examples and mappings without introducing additional full\-width floats\.
#### Workload\-level scope examples\.
SWE\-agent\-style repository work\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]is in scope because command feedback drives progress and removing terminal access changes the workload’s behavior\.OpenHands terminal\-centric workloads\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]are conditional: inclusion depends on whether terminal execution remains the dominant locus of progress\.Static patch generation / Agentless\[[157](https://arxiv.org/html/2608.20485#bib.bib186)\]is out of scope because patch generation does not require iterative terminal execution and feedback\.Browser or desktop agents\[[159](https://arxiv.org/html/2608.20485#bib.bib105)\]are out of scope when visual, GUI, or DOM feedback is the primary execution substrate\.CLI\-packaged API assistants\[[141](https://arxiv.org/html/2608.20485#bib.bib144)\]remain boundary cases when the terminal is only an access surface rather than the progress\-bearing execution substrate\.
#### Architecture and runtime\-infrastructure patterns\.
Direct\-command accessprovides broad command access with approval, permission, or sandbox controls, as in Claude Code, Codex CLI, Aider, and Gemini CLI\[[6](https://arxiv.org/html/2608.20485#bib.bib138),[106](https://arxiv.org/html/2608.20485#bib.bib139),[47](https://arxiv.org/html/2608.20485#bib.bib140),[52](https://arxiv.org/html/2608.20485#bib.bib141)\]; its trade\-off is expressiveness versus noise, safety burden, and rollback difficulty\.ACI mediationexposes structured search, edit, and execution primitives, as in SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\], trading reliability against cross\-task generality\.Platform runtimesprovide persistent workspaces with coordinated interaction surfaces, as in OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\], trading breadth against dependence on the surrounding runtime\.Role\-structured controlseparates planning, execution, and review, as in STRATUS and HyperAgent\[[23](https://arxiv.org/html/2608.20485#bib.bib22),[113](https://arxiv.org/html/2608.20485#bib.bib76)\], improving diagnosability at the cost of coordination overhead\.Scaffold\-centric controloptimizes context, observation, and control flow, as in Meta\-Harness and AutoHarness\[[79](https://arxiv.org/html/2608.20485#bib.bib55),[93](https://arxiv.org/html/2608.20485#bib.bib63)\], which creates an attribution trade\-off\.Runtime memoryuses compression, retrieval, or context adaptation, as in Context\-Folding and TACO\[[136](https://arxiv.org/html/2608.20485#bib.bib90),[121](https://arxiv.org/html/2608.20485#bib.bib82)\], trading persistence against information loss\.Rollout environmentsprovide terminal\-native worlds for training and trajectory generation, as in CLI\-Gym, Endless Terminals, and TermiGen\[[89](https://arxiv.org/html/2608.20485#bib.bib60),[43](https://arxiv.org/html/2608.20485#bib.bib34),[189](https://arxiv.org/html/2608.20485#bib.bib127)\], trading scale against transfer uncertainty\.
#### Sources for terminal competence acquisition and adaptation\.
Terminal\-native rolloutsuse Dockerized generation, Endless Terminals, or CLI\-Gym\[[155](https://arxiv.org/html/2608.20485#bib.bib185),[43](https://arxiv.org/html/2608.20485#bib.bib34),[89](https://arxiv.org/html/2608.20485#bib.bib60)\]to expose command selection, feedback interpretation, and recovery, while transfer beyond generated tasks remains uncertain\.Executable repositoriessuch as SWE\-Gym, SWE\-Dev, and SWE\-rebench\[[110](https://arxiv.org/html/2608.20485#bib.bib181),[37](https://arxiv.org/html/2608.20485#bib.bib31),[9](https://arxiv.org/html/2608.20485#bib.bib13)\]expose navigation, editing, setup, and test\-grounded verification, although terminal behavior remains entangled with repository repair\.Synthetic subskillsfrom TermiGen, SWE\-smith, and CalibForge\[[189](https://arxiv.org/html/2608.20485#bib.bib127),[165](https://arxiv.org/html/2608.20485#bib.bib111),[100](https://arxiv.org/html/2608.20485#bib.bib162)\]target localization, dependency search, repair, and environment construction, with synthetic\-pattern overfit as a limitation\.Failure\-centered tracessuch as TRACE, AgentHER, and AgentForesight\[[69](https://arxiv.org/html/2608.20485#bib.bib114),[33](https://arxiv.org/html/2608.20485#bib.bib58),[178](https://arxiv.org/html/2608.20485#bib.bib8)\]expose diagnosis, rollback, and recovery, but remain limited by sparse traces and inconsistent recovery labels\.
#### Evaluation evidence layers\.
Outcomeasks whether the task was completed and uses pass/fail or executable verification, but hides process quality\.Processasks how execution proceeded and uses step scores, recovery, command economy, and defect evidence, while lacking a shared metric standard\.Environmentasks whether execution was valid and realistic and examines runtime state and dependency resolution, which are often detached from end\-to\-end workflows\.Traceasks whether behavior is inspectable and replayable through commands, observations, state changes, and interventions, although reporting schemas differ\.Governanceasks whether execution was contained and authorized through permissions, sandboxing, reversibility, and side effects, while protocols remain immature\.Freshnessasks whether tasks are temporally valid through mutation, live tasks, and contamination checks; it is a cross\-cutting validity condition, and static pools still dominate\.
## Appendix CEmpirical Diagnostic Study: Protocol and Extended Results
### C\.1Experimental Conditions and Evidence Chain
The empirical study provides two fixed\-condition diagnostic views\. The benchmark\-exposure diagnostic fixes mini\-SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\]with DeepSeek\-V4\-Flash\[[161](https://arxiv.org/html/2608.20485#bib.bib146)\]while varying the benchmark family\. It contains 241 Terminal\-Bench 2\.1 tasks, 93 SetupBench tasks, 21 LongCLI\-Bench tasks, and 640 BashArena tasks\. The matched\-system diagnostic compares mini\-SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\], SWE\-agent\[[164](https://arxiv.org/html/2608.20485#bib.bib109)\], and OpenHands\[[150](https://arxiv.org/html/2608.20485#bib.bib97)\]on the same official task identifiers within each benchmark, using DeepSeek\-V4\-Flash and DeepSeek\-V4\-Pro\[[161](https://arxiv.org/html/2608.20485#bib.bib146)\]\. It contains all 80 Claw\-SWE\-Bench Lite tasks and 300 SWE\-bench Lite tasks\. Each benchmark, system, and model cell therefore represents a complete model, interface, harness, and runtime configuration\.
For each formal task, the evidence chain connects the official evaluator outcome to the complete execution trace, normalized events, task\-level indicators, and benchmark\-level summary\. The model\-call policy fixes temperature to 0\.0, top\-ppto 1\.0, one sample per step, a 32,768\-token maximum output budget, and non\-streaming responses where supported\. Thinking mode is enabled, and high reasoning effort is requested through the harness\-specific adapter where exposed by the provider and harness\. The benchmark\-exposure runs use an 80\-step task ceiling and a 3,600\-second wall\-clock ceiling\. The matched\-system runs use a 200\-step or 200\-call task ceiling and a 7,200\-second wall\-clock ceiling\. Seed 42 controls task ordering and local random sources and is passed to the provider when supported\. Provider\-side context limits, unsupported controls, and harness\-specific context truncation remain part of the evaluated configuration\.
Official benchmark outcomes form a separate result channel from the trace\-derived process indicators\. P1, P3, P5, and P7 use deterministic trace rules\. P2, P4, and P6 use rule\-constrained LLM judgments for semantic distinctions that deterministic succession cannot resolve\. Rule\-based extractors first identify candidate episodes; target\-specific prompts then inspect bounded evidence windows and return schema\-valid decisions with visible evidence\-step identifiers, an evidence statement, and a confidence value\. Technical API failures and invalid responses are retried and must be resolved before aggregation\. Uncertain or low\-confidence semantic labels do not enter semantic numerators or denominators\. A task without eligible semantic evidence for P2, P4, or P6 remains missing for that indicator\. Because different harnesses expose different action schemas, cross\-system process indicators are treated as interface\-sensitive trace evidence rather than directly interchangeable action units\.
The final matched\-system snapshot passed checksum and semantic validation before aggregation\. For the SWE\-agent and OpenHands records in this snapshot, strict semantic checks accepted all 1,520 task records, and stale local records were excluded before outcomes and normalized summaries were computed\.
### C\.2Operational Definitions of P1–P7
The seven trace\-derived process indicators are defined compactly below\. Numerator and denominator rules are applied within each task before task\-level macro averaging where applicable\.
P1\.Rule\-matched invocation failure: failed command events whose stderr matches shell syntax, command\-not\-found, non\-executable, path, permission, argument, or malformed\-tool patterns, divided by all normalized command events\. This is a narrow execution\-friction signal\.
P2\.Helpful feedback utilization: high\-confidence episodes in which visible feedback is used and judged helpful, divided by high\-confidence episodes with usable feedback\. This is a judge\-assisted positive rate\.
P3\.Environment\-command non\-zero\-exit rate: non\-zero exits on dependency, runtime, service, configuration, or environment\-management commands, divided by all detected environment\-management commands\. This is an execution\-friction signal\.
P4\.Consequential state\-tracking error: high\-confidence state or context errors with blocking or inefficient task impact, divided by high\-confidence episodes with decidable state evidence\. This is a judge\-assisted error rate\.
P5\.Final\-window verification rate: the fraction of tasks for which a rule\-matched verification action occurs within the final five actions or final 20% of the trace, whichever window is larger\. This is an activity\-presence and timing signal\.
P6\.Successful task\-relevant recovery: high\-confidence attempted recoveries that succeed and are directly or indirectly relevant to the preceding failure, divided by high\-confidence episodes containing a recovery attempt\. This is a judge\-assisted positive rate\.
P7\.Governance\-review\-trigger rate: normalized command events whose command or associated output matches irreversible, overprivileged, secret\-handling, host or sandbox, or external\-side\-effect patterns, divided by all normalized command events\. This is a contextual governance\-review signal\.
For P2, P4, and P6, letni,pn\_\{i,p\}andmi,pm\_\{i,p\}denote the semantic numerator and denominator for taskiiand indicatorpp\. The task\-level rate is
ri,p=ni,pmi,p,mi,p\>0\.r\_\{i,p\}=\\frac\{n\_\{i,p\}\}\{m\_\{i,p\}\},\\qquad m\_\{i,p\}\>0\.\(1\)The reported benchmark value is the macro\-average over tasks with a defined semantic denominator,
Rp=1\|Ip\|∑i∈Ipri,p,Ip=\{i:mi,p\>0\}\.R\_\{p\}=\\frac\{1\}\{\|I\_\{p\}\|\}\\sum\_\{i\\in I\_\{p\}\}r\_\{i,p\},\\qquad I\_\{p\}=\\\{i:m\_\{i,p\}\>0\\\}\.\(2\)P5 is a binary task indicator and is averaged over all tasks\. For P1, P3, and P7, task\-level values are computed when the corresponding deterministic denominator is defined\. Missing evidence is not assigned a zero\.
For P2, eligible episodes have high\-confidence, non\-uncertain labels andfeedback\_present=true; the numerator additionally requiresfeedback\_used=trueandusefulness=helpful\. For P4, eligible episodes have a decidablestate\_error; the numerator requires an error withtask\_impactequal toblockingorinefficient\. For P6, eligible episodes haverecovery\_attempted=true; the numerator requiresrecovery\_successful=trueand direct or indirect task relevance\. The confidence threshold of 0\.7 is an engineering filter over the judge’s reported confidence rather than a calibrated probability\.
### C\.3Rule\-Constrained LLM\-as\-a\-Judge Configuration and Targeted Human Audit
All formal P2, P4, and P6 labels were generated byDeepSeek\-V4\-Flash\[[161](https://arxiv.org/html/2608.20485#bib.bib146)\]\. Separate prompts ask whether visible terminal feedback is used helpfully, whether a state or context error has consequential task impact, and whether a recovery addresses the original failure\. The judge sees one rule\-extracted episode window at a time, the task context supplied to that run, and only the requested semantic decision\. Table[10](https://arxiv.org/html/2608.20485#A3.T10)records the request and filtering configuration\.
Table 10:Configuration of the semantic episode judge used for formal P2, P4, and P6 labels\.FieldValueFieldValueJudge modelDeepSeek\-V4\-FlashPrompt designSeparate prompts for feedback use, state tracking, and recoveryTemperature0\.0Top\-pp1\.0Initial maximum output4,096 tokensLocal seed42, passed when supportedResponse modeOne non\-streaming JSON objectRetry policyAt most three retries; output cap doubles on reasoning\-only truncation, up to 32,768 tokensEligibility filterConfidence≥0\.7\\geq 0\.7and not uncertainAggregationTask\-level semantic rates over eligible evidenceThe three target\-specific prompt templates share the following instruction: use only the visible episode window; do not use a final benchmark result unless it is visible inside that window; return exactly one JSON object; use JSON booleans or null for unknown Boolean fields; cite non\-empty visible step identifiers; and provide a short evidence summary and a one\- or two\-sentence rationale without hidden reasoning\. The variable task context and episode JSON are inserted after these instructions\. The target\-specific decision rules are reproduced below\.
P2 prompt: feedback utilizationDecide whether the agent used terminal feedback in this episode\.feedback\_presentis true only when visible stdout, stderr, logs, test output, or a command result gives actionable information\.feedback\_usedis true only when a later visible action is grounded in that feedback; a merely different next command is insufficient\.usefulnessis one ofhelpful,superficial,irrelevant, oruncertain\. The JSON fields arefeedback\_present,feedback\_used,usefulness,confidence,evidence\_step\_ids,evidence, andrationale\.
P4 prompt: state, task, and context trackingDecide whether the episode shows the agent acting on stale, contradicted, forgotten, or inconsistent state\. Repetition is not automatically an error because it may be verification or a retry after state change\. A path failure is an error only when the visible prior context supplied enough information to avoid it\.state\_error\_typeis one ofduplicate\_loop,stale\_state,forgotten\_fact,wrong\_path\_state,inconsistent\_goal, oruncertain;task\_impactisblocking,inefficient,minor,none, oruncertain\. The remaining fields arestate\_error,confidence,evidence\_step\_ids,evidence, andrationale\.
P2 uses two preceding and four following steps, P4 uses eight preceding and two following steps, and P6 uses two preceding and eight following steps; the nearest later verification action is added when present\. Command output is clipped with an explicit truncation marker\. The official evaluator outcome is not supplied unless it is already visible inside the episode\.
P6 prompt: failure recoveryDecide whether a later visible action diagnoses, fixes, works around, or otherwise addresses the original visible failure\.recovery\_successfulis true only when the follow\-up resolves that failure or clearly advances the task past it\. An unrelated successful command is not recovery\.recovery\_typeis one ofparameter\_fix,dependency\_fix,path\_fix,code\_fix,strategy\_shift,rollback,workaround, oruncertain;recovery\_relevanceisdirect,indirect,irrelevant, oruncertain\. The remaining fields arerecovery\_attempted,recovery\_successful,confidence,evidence\_step\_ids,evidence, andrationale\.
#### Calibration and targeted human audit
We calibrated the prompts on sampled trajectories and manually inspected calibration episodes spanning low\-confidence or uncertain outputs and retained labels for P2, P4, and P6 across the available calibration benchmarks\. Each episode was reviewed using the task context, complete visible episode window, and indicator codebook\. We recorded the operative target fields and an evidence\-grounded note, treating subordinate fields as inapplicable whenfeedback\_present,state\_error, orrecovery\_attemptedis false\.
This targeted review provides a qualitative check on label behavior, and P2, P4, and P6 are treated as auxiliary process indicators\. Formal aggregation retains the prespecified confidence and uncertainty filters; manual adjudications are kept separate from the reported benchmark\-level rates\.
### C\.4Semantic\-Label Coverage for the Benchmark\-Exposure Diagnostic
The P2, P4, and P6 coverage summary is compactly reported in prose\. Terminal\-Bench 2\.1 contains 241 tasks with 167, 145, and 162 tasks eligible for P2, P4, and P6, respectively; episode\-level coverage and uncertainty are 88\.0% and 9\.6%\. SetupBench contains 93 tasks with 76, 50, and 71 eligible tasks, with 86\.7% coverage and 11\.2% uncertainty\. LongCLI\-Bench contains 21 tasks with 21, 16, and 21 eligible tasks, with 68\.7% coverage and 28\.2% uncertainty\. BashArena contains 640 tasks with 471, 444, and 468 eligible tasks, with 88\.7% coverage and 8\.5% uncertainty\. Here, eligibility is task\-level, whereas coverage and uncertainty are episode\-level rates over extracted semantic candidates\. LongCLI\-Bench has the smallest task set and the lowest accepted semantic\-label coverage among the four benchmark conditions, so its auxiliary semantic indicators are interpreted directionally\. P5 records the presence and timing of a rule\-matched late\-trace verification action, while P7 records command events that trigger contextual governance review\.
### C\.5Matched\-System Outcome and Efficiency Matrix
Figure 7:Matched\-system outcomes for all 12 benchmark–system–model cells\. Points show resolved\-task rates and horizontal intervals show 95% Wilson confidence intervals; marker shape identifies the model variant and color identifies the system\. The plotted cells use aligned task identifiers and official evaluator outcomes\.#### Efficiency values\.
Action counts follow harness\-specific schemas and are not cross\-system units\. For Claw Lite, mini\-SWE\-agent uses 59\.69 actions and 476\.34 s on average with Flash, with median 305\.2 s \[216\.8, 591\.6\], and 55\.25 actions and 475\.68 s with Pro, with median 406\.4 s \[245\.5, 609\.5\]\. SWE\-agent uses 96\.18 actions and 674\.50 s with Flash, with median 519\.3 s \[394\.8, 709\.8\], and 85\.26 actions and 711\.94 s with Pro, with median 601\.4 s \[416\.0, 839\.8\]\. OpenHands uses 55\.69 actions and 600\.16 s with Flash, with median 445\.1 s \[340\.3, 675\.6\], and 51\.50 actions and 653\.08 s with Pro, with median 491\.3 s \[382\.6, 853\.5\]\. For SWE\-bench Lite, mini\-SWE\-agent uses 48\.78 actions and 293\.62 s with Flash, with median 215\.8 s \[147\.4, 350\.4\], and 48\.21 actions and 388\.92 s with Pro, with median 294\.7 s \[198\.1, 466\.4\]\. SWE\-agent uses 63\.99 actions and 382\.09 s with Flash, with median 301\.8 s \[212\.8, 465\.3\], and 56\.95 actions and 525\.13 s with Pro, with median 414\.5 s \[270\.7, 674\.6\]\. OpenHands uses 52\.44 actions and 501\.41 s with Flash, with median 382\.7 s \[291\.1, 599\.9\], and 46\.15 actions and 504\.40 s with Pro, with median 412\.9 s \[312\.5, 572\.3\]\.
For mini\-SWE\-agent, the same provider accounting fields were available across both benchmarks and model variants\. Per\-task token use, reported as input millions \(M\) and output thousands \(k\), is 2\.153/20\.811 for Claw Lite with Flash, 2\.223/18\.116 for Claw Lite with Pro, 1\.285/16\.427 for SWE\-bench Lite with Flash, and 1\.391/15\.298 for SWE\-bench Lite with Pro\. These token totals characterize realized API traffic under mini\-SWE\-agent, including repeated context and step count, rather than an intrinsic cross\-system efficiency measure\.
### C\.6Paired Model\-Variant Analysis
Table 11:Task\-paired outcome transitions between Flash and Pro\. “Pro only” counts tasks passed only by Pro, and “Flash only” counts tasks passed only by Flash\. Exact McNemar tests use the discordant pairs; Holm adjustment covers all six comparisons\. Difference intervals use 20,000 task\-paired bootstrap resamples with seed 42\.BenchmarkSystemBoth failPro onlyFlash onlyBoth passP−\-F \(pp\)95% CI \(pp\)ExactppHolmppClaw Litemini\-SWE\-agent285740−2\.50\-2\.50\[−11\.25\-11\.25, 6\.25\]0\.7741\.000Claw LiteSWE\-agent163259\+1\.25\+1\.25\[−3\.75\-3\.75, 6\.25\]1\.0001\.000Claw LiteOpenHands223451−1\.25\-1\.25\[−7\.50\-7\.50, 5\.00\]1\.0001\.000SWE\-bench Litemini\-SWE\-agent12013131540\.000\.00\[−3\.33\-3\.33, 3\.33\]1\.0001\.000SWE\-bench LiteSWE\-agent1121213163−0\.33\-0\.33\[−3\.67\-3\.67, 3\.00\]1\.0001\.000SWE\-bench LiteOpenHands13312121430\.000\.00\[−3\.00\-3\.00, 3\.33\]1\.0001\.000None of the six paired model\-variant contrasts is statistically significant after Holm adjustment, and all paired bootstrap intervals include zero\. Under the reported conditions, the two model variants therefore show no directional resolved\-rate difference across the evaluated system–benchmark pairs\. The task pairing controls task identity, while each cell remains a single formal run per task\.
### C\.7Paired Cross\-System Analysis
Cochran’sQQtest, using the asymptoticχ2\\chi^\{2\}reference distribution with two degrees of freedom, rejects equal resolved rates among the three matched systems in all four benchmark–model blocks: Claw Lite with Flash,Q=10\.57Q=10\.57,p=0\.0051p=0\.0051; Claw Lite with Pro,Q=15\.50Q=15\.50,p=0\.0004p=0\.0004; SWE\-bench Lite with Flash,Q=15\.49Q=15\.49,p=0\.0004p=0\.0004; and SWE\-bench Lite with Pro,Q=16\.00Q=16\.00,p=0\.0003p=0\.0003\. Table[12](https://arxiv.org/html/2608.20485#A3.T12)localizes these omnibus differences through task\-paired transitions\.
Table 12:Task\-paired cross\-system contrasts\. “A only” and “B only” count discordant tasks resolved by one system\. Exact McNemar tests use the discordant pairs; Holm adjustment covers all 12 system contrasts\. The difference is A minus B in percentage points\.BenchmarkModelSystem ASystem BA onlyB onlyA−\-B \(pp\)ExactppHolmppClaw LiteFlashmini\-SWE\-agentSWE\-agent317−17\.50\-17\.500\.00260\.0232Claw LiteFlashmini\-SWE\-agentOpenHands816−10\.00\-10\.000\.15160\.4033Claw LiteFlashSWE\-agentOpenHands93\+7\.50\+7\.500\.14600\.4033Claw LitePromini\-SWE\-agentSWE\-agent219−21\.25\-21\.250\.00020\.0022Claw LitePromini\-SWE\-agentOpenHands514−11\.25\-11\.250\.06360\.4033Claw LiteProSWE\-agentOpenHands124\+10\.00\+10\.000\.07680\.4033SWE\-bench LiteFlashmini\-SWE\-agentSWE\-agent716−3\.00\-3\.000\.09310\.4033SWE\-bench LiteFlashmini\-SWE\-agentOpenHands2311\+4\.00\+4\.000\.05760\.4033SWE\-bench LiteFlashSWE\-agentOpenHands254\+7\.00\+7\.000\.00010\.0011SWE\-bench LitePromini\-SWE\-agentSWE\-agent513−2\.67\-2\.670\.09630\.4033SWE\-bench LitePromini\-SWE\-agentOpenHands2210\+4\.00\+4\.000\.05010\.4008SWE\-bench LiteProSWE\-agentOpenHands233\+6\.67\+6\.670\.00010\.0011After Holm correction, significant contrasts are benchmark\-specific: SWE\-agent differs from mini\-SWE\-agent on Claw Lite for both model variants and from OpenHands on SWE\-bench Lite for both variants\. The remaining pairwise contrasts do not cross the corrected threshold\. These results characterize task\-level differences within the recorded matched execution snapshots rather than a benchmark\-independent system ordering\.
## Appendix DExtended Empirical Case Notes
The main paper reports two representative traces\. The additional cases below illustrate how environment management, state tracking, verification, recovery, and governance signals arise across the four benchmark\-exposure conditions and the matched\-system comparison\.
Each note follows a common reading sequence\. The task condition and evaluator outcome establish what occurred; process indicators or trace evidence locate the relevant episode behavior; and the mechanism\-level interpretation relates that behavior to runtime state and task requirements\. This sequence separates outcome evidence, trace\-derived signals, and qualitative interpretation while preserving their connections\.
The cases cover environment construction, dependency intervention, state\-tracking and closure, execution\-substrate damage, governance triggers, and system\-dependent patch scope\. Their ordering moves from benchmark\-exposure trajectories to the matched\-system comparison, providing concrete counterparts to both the indicator profiles and paired outcome analysis\.
Across these examples, identical outcome scores arise from different process paths, while similar trace signals depend on task authorization and runtime context\. Reading outcome, indicators, and mechanism together clarifies where progress stalls, how recovery proceeds, and which system layer shapes the observed behavior\.
Terminal\-Benchmagsac\-install: environment construction without final verification
Outcome:failed; official score 0\.0\.Key indicators:P3 environment\-command non\-zero\-exit rate 38\.5%, P5 final\-window verification rate 0\.0%, P6 successful task\-relevant recovery rate 41\.4%\.Mechanism:the task requires building OpenCV from source and installingpymagsac\. The trajectory contains repeated dependency installation, submodule updates, OpenCV builds, background build attempts, and import checks, but the final evaluator still reportsNo module named ’pymagsac’\. The case shows environment friction together with incomplete late\-stage verification and recovery: the agent reacts to build feedback but does not converge to an installation state satisfying the evaluator\.
SetupBenchprometheus\-bd5b2ea: dependency\-resolution loop
Outcome:failed; official score 0\.0\.Key indicators:P3 environment\-command non\-zero\-exit rate 41\.9%, P5 final\-window verification rate 0\.0%, P6 successful task\-relevant recovery rate 18\.8%\.Mechanism:the trajectory alternates among Go installation,go build,go mod download, proxy settings, checksum settings, timeout wrappers, and cache\-lock cleanup\. Many attempts end in timeout or dependency\-download failure\. The case illustrates that environment management requires not only individual setup actions but also a stable runtime strategy and evidence that configuration changes alter the build state\.
LongCLI\-Benchcmu15\_445\_p0: local progress without long\-horizon closure
Outcome:failed; official score 0\.0\.Key indicators:P2 helpful feedback utilization 58\.3%, P4 consequential state\-tracking error rate 18\.2%, P5 final\-window verification rate 0\.0%, P6 successful task\-relevant recovery rate 33\.3%\.Mechanism:the agent explores the BusTub project, reads source files, reacts to missing test paths, edits code, and later encounterssqllogictestand shell failures\. Some missing\-path feedback is used productively, but the trajectory does not maintain a stable account of which defects remain unresolved or which tests must pass\. The case illustrates why long\-horizon terminal tasks require state tracking and final verification in addition to local feedback use\.
BashArena12249547: state mutation breaks the execution substrate
Outcome:failed; official score 0\.0\.Key indicators:P2 helpful feedback utilization 12\.5%, P3 environment\-command non\-zero\-exit rate 56\.2%, P4 consequential state\-tracking error rate 73\.1%, P5 final\-window verification rate 0\.0%, P6 successful task\-relevant recovery rate 5\.8%, P7 governance\-review\-trigger rate 2\.5%\.Mechanism:the trajectory searches architecture\-specific library paths and then attempts library recovery\. Subsequent commands repeatedly fail withbash: error while loading shared libraries: libc\.so\.6, and the official evaluator cannot run because the shell substrate itself is damaged\. The case illustrates recovery in a mutable runtime: after a destructive state change, the agent may need to repair the same execution substrate required to perform the repair\.
Terminal\-Benchsanitize\-git\-repo: governance signal under task authorization
Outcome:passed; official score 1\.0\.Key indicators:P7 governance\-review\-trigger rate 25\.9%, P2 helpful feedback utilization 66\.7%, P6 successful task\-relevant recovery rate 40\.0%\.Mechanism:the task explicitly asks the agent to locate and remove API keys from a repository\. Broad secret search and replacement checks therefore produce a high P7 signal, but the behavior is within the task’s authorization scope and the official tests pass\. The case illustrates why P7 is a contextual governance\-review trigger rather than a direct unauthorized\-action rate\.
Claw\-SWE\-Bench Literubocop\_\_rubocop\-13560: task\-matched system divergence
Condition:DeepSeek\-V4\-Flash on the same task identifier\.Outcomes:mini\-SWE\-agent unresolved, SWE\-agent resolved, OpenHands unresolved\.Trace evidence:the issue asks that ordinary lower\-case’nul’data remain valid\. Mini\-SWE\-agent also edits an existing test despite the instruction not to change tests; OpenHands exempts all method\-call arguments, which is broader than requested; SWE\-agent changes production logic only and passes the official evaluator\.Interpretation:the matched task exposes different patch scopes and outcomes across system configurations\.
## Appendix EExtended Research Roadmap
The main paper identifies four research priorities\. The compact roadmap below expands them into study\-design, reporting, and infrastructure considerations while preserving the research questions developed in the main paper\.
Cross\-domain competence\.*Study design:*evaluate matched systems across multiple operational domains and compare transfer by competence dimension\.*Minimum reporting:*domain distribution, environment diversity, training overlap, outcomes, and process profiles\.*Infrastructure:*shared task and trace formats\.
Fresh evaluation\.*Study design:*combine live or regenerated tasks with replayable execution traces\.*Minimum reporting:*environment version, event schema, missingness, recovery, interventions, and verifier calls\.*Infrastructure:*versioned runtimes and trace schemas\.
Runtime governability\.*Study design:*evaluate authorization, containment, reversibility, and side effects together with task success\.*Minimum reporting:*permissions, sandbox policy, approvals, destructive\-action prevention, and external effects\.*Infrastructure:*threat taxonomies and recoverable sandboxes\.
Model–harness attribution\.*Study design:*use factorial or portable protocols with matched tasks and controlled system changes\.*Minimum reporting:*model version, action interface, context policy, budgets, runtime, and controlled factors\.*Infrastructure:*portable harness descriptions and intervention\-ready runtimes\.Similar Articles
TMax: A Simple Recipe for Terminal Agents
TMax presents a straightforward method for building AI agents that operate in terminal environments, combining practical design principles for effective command-line automation.
Designing an agent-friendly CLI, what am I missing?
A discussion on designing command-line interfaces that are optimized for use by AI agents, seeking input on potential missing considerations.
agents-cli
Agents-cli is a command-line interface tool that enables coding agents to ship AI agents.
AgentOS
AgentOS provides a unified control layer for managing AI agents, tasks, and workspaces.
Terminal-Bench-Science: Evaluating AI agents on scientific research workflows
Terminal-Bench-Science is a benchmark developed by Stanford University researchers to evaluate AI agents on real scientific research workflows, aiming to drive AI capabilities in science.