Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
Summary
提出一种自进化 agent harness 框架:同一冻结模型先作为 solver 解题、再作为 proposer 直接编辑自己的 harness 代码,在多任务上进化后于分布外基准上显著超越 Codex(提升 12.64 分)并达成匹配表现。
View Cached Full Text
Cached at: 10/01/26, 09:40 AM
# Self-Evolving Harness on Multiple Tasks with the Agent as Its Own Optimizer
Source: [https://arxiv.org/html/2609.38372](https://arxiv.org/html/2609.38372)
###### Abstract
A harness is the code around a language\-model agent that organizes prompts, calls tools, manages context, and controls execution\. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self\-evolving harnesses\. In most existing methods, a separate proposer running on a human\-designed harness modifies the solver’s harness, and a separate harness is evolved for each benchmark\. Real\-world tasks come from many domains, so both the evolution and the evaluation of a harness should cover a diverse range of tasks\. We propose a framework close to recursive self\-improvement: the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it\. Each evolution batch draws tasks from five benchmarks in different domains\. To measure generalization, training and held\-out tasks are strictly separated, and we additionally evaluate on five out\-of\-distribution benchmarks never used during evolution\. We frame the evolution process as deep\-learning training with two stages, multi\-task pretraining and continual training\. Starting from a 49\-line seed harness, the harness obtained at the end of the first stage improves the average score by 4\.48 points on the in\-distribution benchmarks and by 12\.64 points on the out\-of\-distribution benchmarks, surpassing Codex on the former and matching it on the latter\. In the second stage, continued evolution on Claw\-Eval, one of the out\-of\-distribution benchmarks, further raises the score on that benchmark from 66\.17 to 68\.06, exceeding Codex\. We also provide an in\-depth analysis of the mechanisms that emerged during evolution, including output truncation, history compaction, and independent review\.
11footnotetext:Nanjing University\. Correspondence to: Qiankai Xu <qiankaixu6@gmail\.com\>\.Figure 1:Overview of the framework\. The solver and the proposer are the same agent, run in separate sessions with the same frozen modelMMand the same harness, which consists of the codeCtC\_\{t\}and the evolution notesNtN\_\{t\}\. In Iterations 1–5, each batch mixes new tasks from five benchmarks; in Iterations 6–7, it combines new Claw\-Eval tasks with success anchors from the first stage\. The solver runs the batch and produces the run recordsEtE\_\{t\}: trajectories, scores, and grader outputs\. The proposer reads these records and the notes, modifies the harness code within the budget of new linesata\_\{t\}, and updates the notes\. Once the new version passes the engineering checks, its codeCt\+1C\_\{t\+1\}and notesNt\+1N\_\{t\+1\}are used in the next iteration\.## 1Introduction
The performance of a language\-model agent depends both on the model itself and on the control code around it\. This code decides what information the model sees at each step, which tools it can call, what to keep when the conversation grows long, and when a task counts as finished; it is usually called the harness\. Work such as ReAct, CodeAct, and SWE\-agent shows that the design of action formats and tool interfaces substantially changes what an agent can do\[[1](https://arxiv.org/html/2609.38372#bib.bib1),[2](https://arxiv.org/html/2609.38372#bib.bib2),[3](https://arxiv.org/html/2609.38372#bib.bib3)\], and with the model held fixed, changing only the harness can also produce large performance differences\[[4](https://arxiv.org/html/2609.38372#bib.bib4),[5](https://arxiv.org/html/2609.38372#bib.bib5)\]\. Most harnesses are still written by engineers, and developers of frontier models build dedicated harnesses such as Codex and Claude Code for their own models\.
Self\-evolving harnesses aim to hand this engineering work to the agent\. Recent work mostly follows the same basic framework: a solver works on the tasks of a benchmark with the current harness and leaves trajectories; a proposer reads these trajectories and modifies the solver’s harness; and the loop repeats\. The proposer is usually driven by a strong model and runs on a fixed, human\-designed harness\. For example, the proposer of Meta\-Harness is Claude Code driven by Claude Opus 4\.6\[[4](https://arxiv.org/html/2609.38372#bib.bib4)\], and HarnessBank uses Claude Opus 4\.8 to improve the harness of Qwen3\.6\-27B\[[6](https://arxiv.org/html/2609.38372#bib.bib6)\]\. The outer loop often adds further human\-designed steps, such as a dedicated agent that organizes trajectories\[[5](https://arxiv.org/html/2609.38372#bib.bib5),[7](https://arxiv.org/html/2609.38372#bib.bib7)\]\. We revisit this framework from three angles: what an ideal and practical self\-improvement framework for frontier models should look like \(Section[1\.1](https://arxiv.org/html/2609.38372#S1.SS1)\); how to make the improvements generalize to unseen tasks \(Section[1\.2](https://arxiv.org/html/2609.38372#S1.SS2)\); and whether the evolution process can be viewed as model training \(Section[1\.3](https://arxiv.org/html/2609.38372#S1.SS3)\)\.
### 1\.1Self\-Improvement Framework with Minimal Human Priors
If the goal is to improve the harnesses used by frontier agents, this framework falls short of practical needs in four respects:
- •Models of the proposer and solver\.Harness improvements often help weaker models more\[[4](https://arxiv.org/html/2609.38372#bib.bib4),[7](https://arxiv.org/html/2609.38372#bib.bib7),[8](https://arxiv.org/html/2609.38372#bib.bib8)\], so it is unsurprising that a strong model can build an effective harness for a weaker one\. The more interesting question is whether a frontier model can improve its own harness beyond human\-designed ones\.
- •Harnesses of the proposer and solver\.The proposer runs on a fixed, human\-designed harness\. However much the solver’s harness improves, the proposer keeps working the same way, so the gains never compound in the improver itself\.
- •Scope of edits\.Methods such as AHE, HarnessX, and Self\-Harness fix in advance the components, edit types, or configuration points that may change\[[5](https://arxiv.org/html/2609.38372#bib.bib5),[7](https://arxiv.org/html/2609.38372#bib.bib7),[9](https://arxiv.org/html/2609.38372#bib.bib9)\], which leaves no room for mechanisms outside the predefined categories\.
- •Human\-designed outer pipelines\.Trajectory\-organizing agents and preset component structures mainly compensate for the weaknesses of current models and may become unnecessary as models improve\[[10](https://arxiv.org/html/2609.38372#bib.bib10),[11](https://arxiv.org/html/2609.38372#bib.bib11)\]\. mini\-SWE\-agent, for instance, keeps only a bash tool and a linear conversation history111[https://github\.com/SWE\-agent/mini\-swe\-agent](https://github.com/SWE-agent/mini-swe-agent)\.
To address these four issues, we adopt a setting close to recursive self\-improvement \(RSI\)\[[12](https://arxiv.org/html/2609.38372#bib.bib12),[13](https://arxiv.org/html/2609.38372#bib.bib13),[14](https://arxiv.org/html/2609.38372#bib.bib14)\]: a frontier model starts from a minimal initial harness and modifies the harness that runs it\. The same agent plays two roles\. As the solver it solves tasks, and as the proposer it reads the complete run records of the previous batch and edits the harness directly \(Figure[1](https://arxiv.org/html/2609.38372#S0.F1)\)\. Both roles use the same frozen model and the same version of the harness\. The proposer itself runs on the harness it is modifying, so an improvement to the solver is also an improvement to the proposer\. The setting follows two principles:first, minimize human\-designed priors, so that the mechanisms in the harness emerge from the model’s own experience;second, give the model maximal freedom to modify the harness, with no restriction on what can be changed or how\. Section[3\.2](https://arxiv.org/html/2609.38372#S3.SS2)describes the concrete design\.
This setting reflects our long\-term vision\. Once models are capable enough, harnesses will no longer need human design: a model can start from scratch and keep iterating on incoming tasks, its own attempts, and the feedback it receives, evolving a harness that fits both the model and its task stream while remaining general and capable\. Our experimental design follows from this vision\.
### 1\.2Generalization of Harness Updates
A harness is ultimately used on tasks it has not seen, so the improvements found by evolution must generalize beyond the training tasks\. Prior studies find that evolution tends to memorize training tasks or overfit to a particular benchmark, and that the gains shrink markedly, or even vanish, on new tasks or out\-of\-distribution benchmarks\[[15](https://arxiv.org/html/2609.38372#bib.bib15),[16](https://arxiv.org/html/2609.38372#bib.bib16),[17](https://arxiv.org/html/2609.38372#bib.bib17),[18](https://arxiv.org/html/2609.38372#bib.bib18),[19](https://arxiv.org/html/2609.38372#bib.bib19)\]\. Rethinking the Evaluation of Harness Evolution and HarnessDev both find smaller improvements once evaluation tasks are separated from evolution tasks\[[20](https://arxiv.org/html/2609.38372#bib.bib20),[21](https://arxiv.org/html/2609.38372#bib.bib21)\]\. To assess and improve the generalization of harness evolution, we take five measures:
- •Train/held\-out separation\.The tasks of each benchmark are split into disjoint training and held\-out sets, and the effect of evolution is measured only on held\-out tasks\. Some prior work reports scores directly on the tasks used for evolution\[[4](https://arxiv.org/html/2609.38372#bib.bib4),[5](https://arxiv.org/html/2609.38372#bib.bib5)\]\.
- •Out\-of\-distribution evaluation\.We also evaluate on five benchmarks never used in the first stage\.
- •Multi\-benchmark mixing\.Most prior work evolves a separate harness for each benchmark \(Table[2](https://arxiv.org/html/2609.38372#S2.T2)\)\. In practice, an agent faces tasks from many domains and a harness cannot be tailored to each of them, so one harness has to work across many benchmarks\. Each of our training batches therefore draws tasks from five benchmarks in different domains, and an edit has to work for all task types at once\[[22](https://arxiv.org/html/2609.38372#bib.bib22)\]\. Evaluation likewise covers multiple benchmarks and uses the average score to measure overall performance\.
- •Decaying edit budget\.Early iterations allow larger changes to establish broadly useful mechanisms; later iterations tighten the budget, leaving less room to memorize individual tasks\.
- •Generalization\-oriented prompt\.The prompt given to the proposer stresses that improvements must generalize to unseen tasks and that memorizing specific tasks is worthless\.
### 1\.3Harness Evolution as Training
Figure 2:Harness evolution viewed as two\-stage training \(schematic\)\. In the first stage \(Iterations 1–5\), each batch takes 12 new tasks from each of five benchmarks, 60 tasks in total\. In the second stage \(Iterations 6–7\), each batch combines 40 tasks from a new benchmark with 20 success anchors from the first stage, rerun with the current harness\. The shaded surfaces sketch the loss of the two objectives over the space of harness code\. After Iteration 5 the objective changes while the code stays atC5C\_\{5\}\(dashed line\), and Iterations 6 and 7 continue from there\. Seed, Iteration 5, and Iteration 7, the three versions evaluated on held\-out tasks, have 49, 430, and 524 lines of Python\. Bottom left: the budget of new linesata\_\{t\}shrinks during the first stage, is raised again when the second stage begins, and then shrinks again, like a learning\-rate schedule\. Bottom right: each update turns the run recordsEtE\_\{t\}and the notesNtN\_\{t\}into a code edit and revised notesNt\+1N\_\{t\+1\}\. Surface shapes, step lengths, and bar heights are illustrative\.Table 1:Correspondence between deep\-learning training and harness evolution in this paper\.
We formulate harness evolution as the optimization of a state external to the frozen model, and we train the harness the way model parameters are trained \(Figure[2](https://arxiv.org/html/2609.38372#S1.F2)\)\. SkillOpt trains skill documents in this way\[[23](https://arxiv.org/html/2609.38372#bib.bib23)\]; we extend the optimization target to the entire harness source code\. The source code plays the role of the parameters, solving and grading a batch of tasks play the roles of the forward pass and the supervision signal, and the proposer’s rewrite of the source is a parameter update\. Table[1](https://arxiv.org/html/2609.38372#S1.T1)gives the full correspondence\. Our experimental design broadly mirrors this optimization process; we highlight three aspects:
- •Data mixture\.Pretraining of large models mixes data from many sources to acquire broad foundational knowledge, and more diverse data tends to yield better generalization\[[24](https://arxiv.org/html/2609.38372#bib.bib24)\]\. Likewise, every first\-stage batch mixes five benchmarks, so that the harness first develops mechanisms that are useful across task types\.
- •Learning\-rate annealing\.Training lowers the learning rate gradually: large early steps make fast progress, and small later steps keep updates stable\[[25](https://arxiv.org/html/2609.38372#bib.bib25)\]\. The budget of new lines likewise tightens across iterations\.
- •Continual training with replay\.Continued training on a new domain tends to erode existing abilities\[[26](https://arxiv.org/html/2609.38372#bib.bib26)\], and a common remedy is to mix old data into the new data as replay\[[27](https://arxiv.org/html/2609.38372#bib.bib27),[28](https://arxiv.org/html/2609.38372#bib.bib28)\]\. Our second stage switches to tasks from a new domain and mixes in old tasks solved in the first stage as replay\.
In our experiments, the harness grows from 49 lines to 430 in the first stage and 524 in the second, gaining tool\-output truncation, history compaction, and independent review along the way, as well as a visual inspection budget in the second stage\. On the held\-out tasks of the five benchmarks used for evolution, the evolved harness scores about 4\.5 points higher on average than the seed harness and also exceeds Codex\. On five out\-of\-distribution benchmarks never used in evolution, its average score is about 12\.6 points above the seed harness and on par with Codex\.
Our main contributions are as follows:
1. 1\.Self\-evolving framework\.A self\-evolving harness framework with minimal human priors: the same model solves tasks on the same harness and modifies the harness that runs it\. Both evolution and evaluation mix benchmarks from multiple domains, and generalization is measured on held\-out tasks and out\-of\-distribution benchmarks\.
2. 2\.Training analogy\.We map harness evolution onto deep\-learning training \(Table[1](https://arxiv.org/html/2609.38372#S1.T1)\), use the budget of new lines as the learning rate, and organize the experiments into two stages, multi\-task pretraining and continual training\.
3. 3\.Experiments and analysis\.We compare the seed harness, the evolved harness, and Codex on ten benchmarks\. Drawing on the code changes and trajectories, we analyze how each mechanism emerged and changed and how it affected task performance, and we discuss the costs of the framework\.
## 2Related Work
Table 2:Comparison of recent self\-evolving harness work with ours\.*Same model*: the proposer that modifies the harness and the solver that solves tasks use the same base model\.*Same harness*: the proposer itself runs on the harness being evolved\.*Evolution tasks*: the tasks from which one harness receives feedback during evolution;*single benchmark*means one benchmark, and*same\-type benchmarks*means several benchmarks or datasets of the same task type\. Many methods are evaluated on several benchmarks but evolve a separate harness for each, which still counts as a single benchmark\.*Acceptance*: how a modified version is admitted to the next round\.†The Self\-Harness paper states that its proposer is invoked by the same model under the current harness, so we mark it with ✓\. According to the paper and its public code repository \([https://github\.com/qzzqzzb/Self\-Harness](https://github.com/qzzqzzb/Self-Harness)\), the proposer receives a piece of text and outputs a structured edit plan, which a separate program parses and applies to the harness; our proposer runs as a coding agent on the same harness as the solver and edits files directly\.
#### Self\-improving agents and harness evolution\.
For language models, Gödel Agent lets an agent modify its own logic at run time, and DGM and SICA let coding agents rewrite their own code bases\[[29](https://arxiv.org/html/2609.38372#bib.bib29),[30](https://arxiv.org/html/2609.38372#bib.bib30),[31](https://arxiv.org/html/2609.38372#bib.bib31)\]\. Since 2026 the focus has shifted to harnesses\[[4](https://arxiv.org/html/2609.38372#bib.bib4),[5](https://arxiv.org/html/2609.38372#bib.bib5),[7](https://arxiv.org/html/2609.38372#bib.bib7)\], and later work has introduced designs such as candidate banks, population merging, and module\-wise evolution\[[6](https://arxiv.org/html/2609.38372#bib.bib6),[32](https://arxiv.org/html/2609.38372#bib.bib32),[16](https://arxiv.org/html/2609.38372#bib.bib16)\]\. Table[2](https://arxiv.org/html/2609.38372#S2.T2)compares representative methods\.
#### Harness gains and model capability\.
The gains from harness improvements depend on model capability\. On Terminal\-Bench 2\.0, Meta\-Harness improves Claude Haiku 4\.5 more than Claude Opus 4\.6\[[4](https://arxiv.org/html/2609.38372#bib.bib4)\], and HarnessX likewise observes that the weakest task model benefits most\[[7](https://arxiv.org/html/2609.38372#bib.bib7)\]\. Harness Updating Is Not Harness Benefit finds that the strongest tier of models gains less from updates than mid\-tier models\[[8](https://arxiv.org/html/2609.38372#bib.bib8)\]\. OEO finds that a sufficiently strong optimizer model performs better when it organizes the optimization process itself than when it follows a prescribed pipeline\[[10](https://arxiv.org/html/2609.38372#bib.bib10)\]\.
#### Evolution tasks and evaluation\.
SICA, Meta\-Harness, and HarnessDev let one harness receive feedback from multiple benchmarks or datasets, all of the same task type\[[31](https://arxiv.org/html/2609.38372#bib.bib31),[4](https://arxiv.org/html/2609.38372#bib.bib4),[21](https://arxiv.org/html/2609.38372#bib.bib21)\]\. SEAGym and Evo\-Bench find that both the diversity of experience sources and the task domain affect the outcome of evolution\[[33](https://arxiv.org/html/2609.38372#bib.bib33),[34](https://arxiv.org/html/2609.38372#bib.bib34)\]\. Simple Baselines and Rethinking the Evaluation of Harness Evolution stress that evolution methods should be compared with simple methods under the same budget\[[35](https://arxiv.org/html/2609.38372#bib.bib35),[20](https://arxiv.org/html/2609.38372#bib.bib20)\]\.
#### Optimization\-style evolution\.
TextGrad and GEPA update prompts with textual feedback and trajectory reflection, respectively\[[36](https://arxiv.org/html/2609.38372#bib.bib36),[37](https://arxiv.org/html/2609.38372#bib.bib37)\]\. SkillOpt uses bounded edits and a textual learning\-rate budget when training skill documents, which directly inspired our training analogy\[[23](https://arxiv.org/html/2609.38372#bib.bib23)\]; the concurrent RRSI also anneals the number of edits a candidate may contain across iterations\[[15](https://arxiv.org/html/2609.38372#bib.bib15)\]\. Adaptive Auto\-Harness and Continual Harness study continual evolution, targeting streams of heterogeneous incoming tasks and embodied tasks without environment resets, respectively\[[38](https://arxiv.org/html/2609.38372#bib.bib38),[39](https://arxiv.org/html/2609.38372#bib.bib39)\]\.
## 3Method
### 3\.1Problem Setup
LetMMdenote the frozen base model\. At the end of iterationtt, the harness state isHt=\(Ct,Nt\)H\_\{t\}=\(C\_\{t\},N\_\{t\}\), whereCtC\_\{t\}is the entire source code that runs the agent, including the system prompt, tool definitions and implementations, context handling, and execution control, andNtN\_\{t\}denotes the evolution notes stored alongside the code \(Section[3\.5](https://arxiv.org/html/2609.38372#S3.SS5)\)\.H0H\_\{0\}is the seed harness\.
The harness runs on a fixed execution interface𝒦\\mathcal\{K\}, which calls the model, executes commands, and records the entire run\. Given a taskxx, one run of the agent is
τ=𝒜𝒦\(M,Ht,x\),\\tau=\\mathcal\{A\}\_\{\\mathcal\{K\}\}\(M,H\_\{t\},x\),\(1\)whereτ\\taucontains the model replies, tool calls, command outputs, and final artifacts\. The grader of benchmarkkkreturns
\(r,v\)=Vk\(x,τ\),\(r,v\)=V\_\{k\}\(x,\\tau\),\(2\)whererris the task score andvvis the test log or grading explanation\.
### 3\.2Solver, Proposer, and Seed Harness
Following the two principles of Section[1\.1](https://arxiv.org/html/2609.38372#S1.SS1), minimizing human\-designed priors and giving the model maximal freedom to modify the harness, the framework makes three design choices\.
Thesolverreceives benchmark tasks and solves them, and theproposerreceives an update taskItI\_\{t\}that asks it to improve the harness\. Both roles are played by the same agent with the sameMMand the sameHtH\_\{t\}, and each runs in a fresh session\.
The outer program only dispatches tasks and packages run records; there is no step that organizes trajectories or summarizes errors\. The update task provides three kinds of material: the source code and evolution notes of the current harness; the complete run records of the previous batch, including trajectories, task scores, and grader outputs, with both successful and failed tasks provided as is, without summarization or selection; and, from the second iteration of each stage on, the run record of the previous update session\. The update prompt is short \(Appendix[B](https://arxiv.org/html/2609.38372#A2)gives the full text\)\. It tells the proposer that the modified harness will be evaluated on unseen tasks and on entirely different benchmarks, so it should look for improvements that generalize; that it may add or delete files and introduce any component; and that the notes travel with the harness to the next round\. It also states the input limit per request and the budget of new lines for the current iteration\. Beyond this, the prompt does not prescribe which records to read first, how to identify problems, where to start, or which part to change\.
The seedH0H\_\{0\}is a single 49\-line Python file \(Appendix[A](https://arxiv.org/html/2609.38372#A1)\) containing a system prompt of a few sentences, a bash tool, and a minimal loop: call the model, execute the tool calls in its reply and append the results to the history, and stop once the model replies without tool calls\. The seed has no output truncation, command timeouts, context compaction, or error recovery; whether these are needed, and how to implement them, is left to evolution\. The framework also neither partitions the harness into components nor restricts the scope of edits: apart from the execution interface𝒦\\mathcal\{K\}, any part of any file in the harness can be modified, and files can be created or deleted\.
### 3\.3One Iteration: Rollout, Update, and Check
Iterationt\+1t\+1starts fromHtH\_\{t\}and consists of three steps: rollout, update, and check \(Figure[1](https://arxiv.org/html/2609.38372#S0.F1)and Algorithm[1](https://arxiv.org/html/2609.38372#alg1)\)\. In the rollout, the solver runs the task batchBtB\_\{t\}of the iteration withHtH\_\{t\}, once per task, producing the records
Et=\{\(x,τx,rx,vx\):x∈Bt\}\.E\_\{t\}=\\\{\(x,\\tau\_\{x\},r\_\{x\},v\_\{x\}\):x\\in B\_\{t\}\\\}\.\(3\)In the update, the outer program packagesEtE\_\{t\}together with the other materials described in Section[3\.2](https://arxiv.org/html/2609.38372#S3.SS2)into the update taskItI\_\{t\}and passes it to the same agent:
H~t\+1=Extract\(𝒜𝒦\(M,Ht,It\)\),\\widetilde\{H\}\_\{t\+1\}=\\operatorname\{Extract\}\\\!\\left\(\\mathcal\{A\}\_\{\\mathcal\{K\}\}\(M,H\_\{t\},I\_\{t\}\)\\right\),\(4\)whereExtract\\operatorname\{Extract\}retrieves the modified code and notes at the end of the session\. In the check, the new version must pass a few engineering checksGtG\_\{t\}: the edits touch only the harness itself, there is at least one change besides the notes, the number of new lines does not exceedata\_\{t\}, and the harness completes a full task run with a mock model:
Ht\+1=\{H~t\+1,Gt\(H~t\+1,Ht\)=1,Ht,otherwise\.H\_\{t\+1\}=\\begin\{cases\}\\widetilde\{H\}\_\{t\+1\},&G\_\{t\}\(\\widetilde\{H\}\_\{t\+1\},H\_\{t\}\)=1,\\\\ H\_\{t\},&\\text\{otherwise\}\.\\end\{cases\}\(5\)If the check fails, the outer program starts a repair session\.
Algorithm 1The harness self\-evolution loop shared by both stages1:frozen model
MMand execution interface
𝒦\\mathcal\{K\}, seed
H0H\_\{0\}, task batches
\{Bt\}t=06\\\{B\_\{t\}\\\}\_\{t=0\}^\{6\}and budgets
\{at\}t=06\\\{a\_\{t\}\\\}\_\{t=0\}^\{6\}of the iterations
2:harness states
H1,…,H7H\_\{1\},\\ldots,H\_\{7\}and all run records
3:for
t=0,…,6t=0,\\ldots,6do
4:
Et←∅E\_\{t\}\\leftarrow\\varnothing
5:for
x∈Btx\\in B\_\{t\}do⊳\\trianglerightsolver, parallelizable
6:
τx←𝒜𝒦\(M,Ht,x\)\\tau\_\{x\}\\leftarrow\\mathcal\{A\}\_\{\\mathcal\{K\}\}\(M,H\_\{t\},x\)
7:
\(rx,vx\)←Vk\(x\)\(x,τx\)\(r\_\{x\},v\_\{x\}\)\\leftarrow V\_\{k\(x\)\}\(x,\\tau\_\{x\}\)
8:
Et←Et∪\{\(x,τx,rx,vx\)\}E\_\{t\}\\leftarrow E\_\{t\}\\cup\\\{\(x,\\tau\_\{x\},r\_\{x\},v\_\{x\}\)\\\}
9:endfor
10:build the update task
ItI\_\{t\}from
EtE\_\{t\}, the record of the previous update session,
HtH\_\{t\}, and
ata\_\{t\}
11:
H~t\+1←Extract\(𝒜𝒦\(M,Ht,It\)\)\\widetilde\{H\}\_\{t\+1\}\\leftarrow\\operatorname\{Extract\}\(\\mathcal\{A\}\_\{\\mathcal\{K\}\}\(M,H\_\{t\},I\_\{t\}\)\)⊳\\trianglerightproposer
12:check with Eq\. \([5](https://arxiv.org/html/2609.38372#S3.E5)\) to obtain
Ht\+1H\_\{t\+1\}
13:endfor
14:freeze
H5H\_\{5\}and
H7H\_\{7\}and evaluate them on the corresponding held\-out tasks
### 3\.4Two Stages: Multi\-Task Pretraining and Continual Training
Stage 1: multi\-task evolution \(Iterations 1–5\)\.There areK=5K=5benchmarks, each split in advance into a training setDktrD\_\{k\}^\{\\mathrm\{tr\}\}and a disjoint held\-out setDkhoD\_\{k\}^\{\\mathrm\{ho\}\}\. Each iteration takesb=12b=12previously unused tasks from each training set to form a batch of 60 tasks:
Bt=⋃k=1KBt,k,Bt,k⊂Dktr,\|Bt,k\|=b,Bt,k∩Bu,k=∅\(t≠u\)\.B\_\{t\}=\\bigcup\_\{k=1\}^\{K\}B\_\{t,k\},\\qquad B\_\{t,k\}\\subset D\_\{k\}^\{\\mathrm\{tr\}\},\\qquad\|B\_\{t,k\}\|=b,\\qquad B\_\{t,k\}\\cap B\_\{u,k\}=\\varnothing\\ \(t\\neq u\)\.\(6\)Together, the five iterations use each of the 60 training tasks of every benchmark exactly once, and the held\-out sets are never run during evolution\.
Stage 2: continued evolution on a new domain \(Iterations 6–7\)\.Starting from Iteration 5, the training tasks switch to tasksBtnewB\_\{t\}^\{\\mathrm\{new\}\}from a new domain, mixed with a small number of old tasksAtA\_\{t\}:
Bt=Btnew∪At,t∈\{5,6\}\.B\_\{t\}=B\_\{t\}^\{\\mathrm\{new\}\}\\cup A\_\{t\},\\qquad t\\in\\\{5,6\\\}\.\(7\)The two iterations use different new tasks\.AtA\_\{t\}is drawn from the training tasks solved in the first stage, spread evenly across the five benchmarks; we call these tasks success anchors\. They are rerun with the current harness, and their results are given to the proposer together with the records of the new tasks, so that while adapting the harness to the new domain the proposer also sees how its edits affect the original tasks\.
### 3\.5Optimizer State and Learning\-Rate Schedule
Table[1](https://arxiv.org/html/2609.38372#S1.T1)lists the full correspondence; here we give the implementation details of the optimizer state and the learning\-rate schedule\.
#### Evolution notes as optimizer state\.
The first moment in Adam accumulates past gradients with exponential weighting, so that each update reflects both the current gradient and the direction of earlier updates\[[40](https://arxiv.org/html/2609.38372#bib.bib40),[41](https://arxiv.org/html/2609.38372#bib.bib41)\]\. The evolution notes play a similar role: in each iteration, the proposer first reads the existing notes and then, in light of the new batch of run records, writes down which edits worked as expected, which earlier judgments need correction, and what to try next\. Unlike numerical momentum, the notes are free text that can be rewritten entirely, and the proposer can also overturn earlier conclusions; Section[5](https://arxiv.org/html/2609.38372#S5)gives concrete examples\.
#### Edit budget as learning rate\.
The budgetata\_\{t\}of each iteration limits the number of lines added between two versions of the source code, i\.e\.,A\(Ct,C~t\+1\)≤atA\(C\_\{t\},\\widetilde\{C\}\_\{t\+1\}\)\\leq a\_\{t\}; deleted lines and edits to the notes do not count toward the budget\. The first\-stage budget follows a cosine learning\-rate schedule\[[25](https://arxiv.org/html/2609.38372#bib.bib25)\]and decreases fromamax=500a\_\{\\max\}=500toamin=100a\_\{\\min\}=100:
at=round\[amin\+amax−amin2\(1\+cosπtTpre−1\)\],t=0,…,Tpre−1,a\_\{t\}=\\operatorname\{round\}\\\!\\left\[a\_\{\\min\}\+\\frac\{a\_\{\\max\}\-a\_\{\\min\}\}\{2\}\\left\(1\+\\cos\\frac\{\\pi t\}\{T\_\{\\mathrm\{pre\}\}\-1\}\\right\)\\right\],\\qquad t=0,\\ldots,T\_\{\\mathrm\{pre\}\}\-1,\(8\)whereTpre=5T\_\{\\mathrm\{pre\}\}=5, giving budgets of 500, 441, 300, 159, and 100 lines for the five iterations\. When the second stage enters the new domain, the budget is raised back to 200 lines and then lowered to 100 lines, mirroring the practice in continual pretraining of first re\-warming and then decaying the learning rate\[[28](https://arxiv.org/html/2609.38372#bib.bib28)\]\.
## 4Experiments
### 4\.1Experimental Setup
#### Benchmarks\.
The first stage uses five benchmarks that differ in interaction style, task length, and grading \(Table[3](https://arxiv.org/html/2609.38372#S4.T3)\)\. Terminal\-Bench 2\.1 consists of long\-horizon engineering tasks, such as software builds and system configuration, carried out in a terminal\[[42](https://arxiv.org/html/2609.38372#bib.bib42)\]\. GAIA2 involves handling asynchronous requests in a simulated mobile\-app environment that pushes new messages over time\[[43](https://arxiv.org/html/2609.38372#bib.bib43)\]\. OfficeQA Pro requires retrieving evidence from nearly a century of U\.S\. Treasury bulletins and performing numerical reasoning\[[44](https://arxiv.org/html/2609.38372#bib.bib44)\]\.τ3\\tau^\{3\}\-Bench banking involves multi\-turn conversations with a simulated user to handle banking requests under business rules\[[45](https://arxiv.org/html/2609.38372#bib.bib45)\]\. SWE\-Bench Pro asks the agent to fix real issues in large code repositories\[[46](https://arxiv.org/html/2609.38372#bib.bib46)\]\.
To test the generalization of the evolved harness, we also select five out\-of\-distribution \(OOD\) benchmarks never used in the first stage: SWE\-bench Verified, issue resolution in Python repositories\[[47](https://arxiv.org/html/2609.38372#bib.bib47),[48](https://arxiv.org/html/2609.38372#bib.bib48)\]; DeepSWE, feature development in repositories written in multiple programming languages\[[49](https://arxiv.org/html/2609.38372#bib.bib49)\]; BrowseComp\-Plus, question answering by retrieval over a fixed corpus of about 100,000 web pages\[[50](https://arxiv.org/html/2609.38372#bib.bib50)\], which is placed in the task environment as files; APEX\-Agents, professional work in investment banking, consulting, and law\[[51](https://arxiv.org/html/2609.38372#bib.bib51)\]; and Claw\-Eval, which mixes service orchestration, multimodal understanding, and multi\-turn dialogue\[[52](https://arxiv.org/html/2609.38372#bib.bib52)\]\. Time limits and grading follow the official settings of each benchmark\.
Table 3:Benchmarks and task counts\. Benchmarks with few held\-out tasks are evaluated repeatedly and averaged: each of the 29 Terminal\-Bench 2\.1 tasks is run three times, written as 29×\\times3; for OfficeQA Pro andτ3\\tau^\{3\}\-Bench banking we take two complete runs, written as 73×\\times2 and 37×\\times2\. The out\-of\-distribution benchmarks have no first\-stage training tasks, and the Claw\-Eval training tasks are used only in the second stage\.BenchmarkTask typeTotalTrainHeld\-outIn\-distribution benchmarksTerminal\-Bench 2\.1Long\-horizon engineering tasks in a terminal896029×\\times3GAIA2Asynchronous interaction in dynamic app environments80060100OfficeQA ProTreasury document retrieval and numerical reasoning1336073×\\times2τ3\\tau^\{3\}\-Bench bankingMulti\-turn service dialogue with a simulated user976037×\\times2SWE\-Bench ProIssue resolution in code repositories73160100Out\-of\-distribution benchmarksSWE\-bench VerifiedIssue resolution in Python repositories500–100DeepSWEFeature development in multi\-language repositories113–53BrowseComp\-PlusDeep\-research question answering over a fixed corpus830–100APEX\-AgentsProfessional work in banking, consulting, and law480–100Claw\-EvalService orchestration, multimodal understanding, and multi\-turn dialogue30080100
#### Model and baselines\.
All experiments use GPT\-5\.6 Sol with high reasoning effort, an input limit of 258,400 tokens, and an output limit of 65,536 tokens per request; the solver and the proposer have identical configurations\. We compare three harnesses: the seed harness \(Seed\), Iteration 5 at the end of the first stage, and the off\-the\-shelf coding agent Codex222[https://github\.com/openai/codex](https://github.com/openai/codex)\(Codex CLI 0\.146\.0\), which uses the same model and keeps its own tools and context\-compaction mechanism\. Iteration 7 is compared with Iteration 5 only on the new domain of the second stage\.
#### Preventing reward hacking\.
To prevent reward hacking, such as bypassing the task to obtain answers directly or exploiting loopholes in grading\[[53](https://arxiv.org/html/2609.38372#bib.bib53),[54](https://arxiv.org/html/2609.38372#bib.bib54)\], we apply two safeguards in all evaluations, a preventive constraint and a post\-hoc audit, identically for all three harnesses\. Beforehand, every task description ends with a fixed requirement that forbids the agent from obtaining answers from published solutions, upstream patches, or evaluation files, and reference answers and grading tests are placed in the environment only after the agent has finished\. Afterward, we audit all evaluation trajectories\. A script first scans the commands and outputs of every trajectory and flags suspicious behavior such as downloading upstream fixes, searching for published answers or dataset copies, reading grading materials, or modifying test files; for SWE\-Bench Pro, it also compares the content returned from the network with the reference patch, record by record\. Independent agents then check the flagged records one by one and decide whether the answers came from these sources\. We also inspect the trajectories of records with unusually high or low scores\. A record confirmed as cheating scores 0 and is not rerun\. No cheating was confirmed in any evaluation record reported in this paper\.
### 4\.2Main Results
Figure 3:Held\-out scores of the three harnesses with GPT\-5\.6 Sol \(high\)\. \(a\) The five in\-distribution benchmarks; \(b\) the five out\-of\-distribution benchmarks\. Scores on APEX\-Agents and Claw\-Eval are the mean official score multiplied by 100; on the other benchmarks they are success rates \(%\)\. Averages weight the five benchmarks equally\.#### In\-distribution results\.
Figure[3](https://arxiv.org/html/2609.38372#S4.F3)\(a\) shows the scores of the three harnesses on the held\-out tasks of the five in\-distribution benchmarks\. Iteration 5 averages 60\.17 across the five benchmarks, 4\.48 points above Seed and 2\.99 above Codex\. The gains are concentrated on Terminal\-Bench 2\.1 and SWE\-Bench Pro, which rise from 66\.67 and 51\.00 to 83\.91 and 61\.00, clearly above Codex as well\. Both benchmarks consist mainly of executing commands and editing files, and their tasks often take dozens of tool calls; the output truncation, command time limits, history compaction, and independent review added during evolution target exactly such long\-horizon tasks \(Section[5](https://arxiv.org/html/2609.38372#S5)\)\. GAIA2 improves slightly, and Seed scores highest on OfficeQA Pro andτ3\\tau^\{3\}\-Bench banking\.
#### Out\-of\-distribution results\.
Figure[3](https://arxiv.org/html/2609.38372#S4.F3)\(b\) shows the results on the five out\-of\-distribution benchmarks\. Iteration 5 averages 70\.45, 12\.64 points above Seed and within one point of Codex \(70\.57\)\. The largest change is on BrowseComp\-Plus, where Iteration 5 rises from Seed’s 34\.00 to 90\.00, compared with 93\.00 for Codex, mainly owing to output truncation \(Section[5\.2](https://arxiv.org/html/2609.38372#S5.SS2)\)\. Iteration 5 scores highest on DeepSWE, and the three harnesses are close on SWE\-bench Verified, APEX\-Agents, and Claw\-Eval\. Even excluding BrowseComp\-Plus, the average of Iteration 5 over the other four benchmarks \(65\.56\) exceeds those of Seed \(63\.76\) and Codex \(64\.97\)\. These benchmarks differ considerably from the tasks used in evolution, yet the mechanisms developed during evolution remain effective on them\.
#### Where Codex trails the seed\.
On OfficeQA Pro andτ3\\tau^\{3\}\-Bench banking, Codex scores below Seed\. A task\-by\-task comparison of the trajectories of the three harnesses shows that neither gap is related to harness mechanisms: on OfficeQA Pro, the questions answered incorrectly are themselves ambiguous; onτ3\\tau^\{3\}\-Bench banking, the failures of Codex are spread over different tasks and stem from business\-judgment and calculation errors on individual tasks\. Section[5\.6](https://arxiv.org/html/2609.38372#S5.SS6)explains why Iteration 5 falls below Seed on these two benchmarks\.
### 4\.3Continual Training on Claw\-Eval
Figure 4:Results on the Claw\-Eval held\-out tasks; the four harnesses use the same 100 tasks and the same model\. \(a\) Mean official score over the 100 tasks, multiplied by 100\. \(b\) Number of tasks with a score of at least 0\.75 in each category; parentheses give the number of tasks in the category\.The second stage starts from Iteration 5 and runs two consecutive iterations on Claw\-Eval, yielding Iteration 7\. Each iteration mixes 40 Claw\-Eval training tasks, different for the two iterations, with 20 success anchors from the five first\-stage benchmarks\.
Figure[4](https://arxiv.org/html/2609.38372#S4.F4)compares Iteration 5 with Iteration 7, with Seed and Codex as references\. Iteration 7 raises the overall score from 66\.17 to 68\.06 and the number of passed tasks from 52 to 57, both the highest among the four harnesses\. The gain comes mainly from multimodal tasks, whose average score rises from 39\.3 to 47\.8, consistent with the visual inspection budget added in these two iterations \(Section[5\.5](https://arxiv.org/html/2609.38372#S5.SS5)\)\.
## 5Analysis
### 5\.1How the Harness Evolved
Table[4](https://arxiv.org/html/2609.38372#S5.T4)lists the code changes over the seven iterations\. The edits fall into five broad groups of components: the tool layer that executes commands \(output truncation, command time limits, image viewing\), context management \(history compaction and a progress file\), end\-of\-task checks \(from a completion check to independent review in a fresh context\), the interaction protocol \(deciding when a conversation ends\), and task\-specific rules \(e\.g\., set operations, deadlines, and stopping conditions for search\)\. The second stage added a visual inspection budget and delivery deadlines\.
Table 4:Total Python lines of the harness at each iteration \(including blank lines and comments\) and the main code changes\.The first iteration made the largest changes\. The notes of Iteration 1 record several problems exposed by the seed: one command printed about 16 MB of logs and pushed the next request over the context limit; some tasks spent their time budget on a hung command; and in image tasks the agent, unable to view images, resorted to parsing pixels one by one\. Based on these observations, the proposer added output truncation, command time limits, history compaction, and an image\-viewing tool in a single update\.
Later iterations mostly revised mechanisms added in earlier rounds\. End\-of\-conversation detection first matched tool names with a regular expression \(Iteration 3\), switched to word\-level matching after theendinsidesendwas found to trigger it falsely \(Iteration 4\), and was later extended to read tool descriptions \(Iteration 7\)\. The image tool first gained a size cap \(Iteration 2\), then format validation by byte signature \(Iteration 4\), and then limits on the number of images attached per reply and per task \(Iteration 6\)\. The notes also assess the previous round’s edits and correct earlier judgments\. For example, the notes of Iteration 4 state that “the previous note’s claim that the image ceiling fully prevented visual\-request crashes was too broad\.”
The next four subsections analyze four mechanisms that affect multiple benchmarks: output truncation, history compaction, independent review, and the visual inspection budget added in the second stage\. Appendix[C](https://arxiv.org/html/2609.38372#A3)gives the trigger rates of the first three on the ten benchmarks\. Section[5\.6](https://arxiv.org/html/2609.38372#S5.SS6)analyzes the benchmarks on which Iteration 5 falls below Seed, and Section[5\.7](https://arxiv.org/html/2609.38372#S5.SS7)discusses the costs of the two design principles\.
### 5\.2Output Truncation
Output truncation was added in Iteration 1, and later iterations never changed its thresholds\. When the output of a single command exceeds 48,000 characters, the harness keeps only its beginning and end and replaces the middle with a one\-line note that suggests a more precise command or writing the output to a file; command outputs older than the six most recent ones are further cut to at most 6,000 characters\. The seed puts command outputs into the context verbatim, so a single overlong output pushes the next request over the input limit and the task aborts\.
Such aborts occurred with the seed on nine of the ten benchmarks and never with Iteration 5\. BrowseComp\-Plus was affected most: its corpus is large, a broadrgsearch can print tens of millions of characters, and the seed aborted on 65 tasks for this reason\. Receiving truncated output, Iteration 5 narrows its searches and reads candidate files one at a time, and it answers 57 of these 65 tasks correctly\. On SWE\-Bench Pro, the seed aborted on 12 tasks because of long test or build output, and Iteration 5 solves 7 of them\. Truncation was added based on training records from terminal and software\-engineering tasks, yet its largest gain comes on a retrieval benchmark never seen during evolution\.
### 5\.3History Compaction
History compaction was also added in Iteration 1 and has likewise remained unchanged\. When the total length of the messages exceeds 360,000 characters, the harness keeps only the first two messages and about 170,000 characters of the most recent messages, and it prompts the model to re\-check the files and the progress file and not to assume that the deleted commands all succeeded\. With output truncation in place, the context of coding tasks rarely grows to this threshold; compaction occurs mainly in long tasks that search or view images repeatedly \(Appendix[C](https://arxiv.org/html/2609.38372#A3)\)\.
On BrowseComp\-Plus, the seed aborted from context overflow on three questions that Iteration 5 answered correctly after compaction; q0285 is one of them\. On this question, after three model calls by the seed, one search command printed about 3\.19 million characters, and the fourth request exceeded the input limit\. Iteration 5 made 64 model calls on the same question and compacted its history three times along the way\. After each compaction, it first checked the progress file and the answer file, then resumed searching and verifying, and eventually answered correctly\.
### 5\.4Independent Review
Review was introduced in Iteration 2 and revised three times afterwards\. Iteration 2 appends a completion\-check request the first time the model replies without calling a tool\. The notes of Iteration 3 argue that when the same session re\-checks its own result with the whole implementation trajectory in context, the second pass tends to repeat the reasoning of the first\. For tasks that need no interaction, the harness therefore clears the conversation, keeps only the original task and the file system, and lets the model review the result independently in a fresh context; interactive tasks keep their history to avoid repeating actions already taken\. Iteration 5 requires the review to read the progress file first and to prioritize requirements not yet checked\. In the second stage, Iteration 6 further caps the review at 150 seconds or 8 calls\. Iteration 5 enters review on nearly every non\-interactive task \(Appendix[C](https://arxiv.org/html/2609.38372#A3)\)\.
During review, the model can gather evidence and recompute, and it revises the answer when it finds a problem\. In one run of OfficeQA Pro task uid0029, Iteration 5 first took 1960 as the first year of a table and computed 0\.77667; during review it realigned the year column, counted from 1959, obtained 0\.88525, and passed grading\.
### 5\.5Visual Inspection Budget
The visual inspection budget was added on Claw\-Eval in the second stage\. Multimodal tasks require viewing images or video frames repeatedly, yet an answer or a media file must be delivered within 600 seconds; Iteration 5 failed to deliver before the deadline on 22 of the 34 multimodal held\-out tasks\. Iteration 6 limits the number of images viewed: at most 4 per reply and 24 per task, and images already used are not sent again\. The notes of Iteration 7 observe that after the image quota ran out, the model kept making revisions through the command line, so Iteration 7 added delivery deadlines\. Once at least 12 images have been viewed, the model is told at 360 seconds to stop collecting evidence and start building the deliverable, and at 500 seconds the tools are withdrawn and the current result must be submitted; after the image quota is used up, at most 8 more model calls are allowed\. Iteration 7 no longer timed out on these 34 tasks, and the scores of multimodal tasks rose accordingly \(Section[4\.3](https://arxiv.org/html/2609.38372#S4.SS3)\)\.
On the video speed\-change task M097, Iteration 5 requested images 41 times, timed out at 600 seconds, and scored 0\.2\. Iteration 7 made 24 requests, received the notice that the image quota was used up at about 304 seconds, then generated and checked the speed\-changed video and the time\-interval file, and scored 1\.0\.
### 5\.6Regressions Relative to the Seed
Iteration 5 scores below Seed onτ3\\tau^\{3\}\-Bench banking and OfficeQA Pro\. Onτ3\\tau^\{3\}\-Bench banking it passes four fewer runs than Seed, and the gap comes from two mechanisms\. The first is history compaction \(Section[5\.3](https://arxiv.org/html/2609.38372#S5.SS3)\)\. In multi\-turn dialogue the customer’s request exists only in the conversation history, and all 11 runs in which Iteration 5 triggered compaction deleted the customer’s original request at the first compaction; only one of them passed\. The second is the interface to task tools\. From Iteration 1 on, the harness exposes all tools provided by the environment directly to the model, includingconfigure\_run, which sets run parameters\. In four runs, Iteration 5 used it to lower the conversation step limit from 200, and two of these runs were cut off before the conversation finished\. Neither Seed nor Codex ever called this tool\.
The gap on OfficeQA Pro is unrelated to harness mechanisms\. The questions answered incorrectly are themselves ambiguous, for example leaving open whether to compute by fiscal or calendar year or how to take the numerator and denominator of a ratio, and all three harnesses gave the same wrong answers on these questions\.
### 5\.7Costs of Simplicity and Freedom
Our framework is deliberately simple: the outer loop has no analysis or validation step, and the scope of edits is unrestricted\. This design comes with three costs, some of which should diminish as models become more capable\.
#### Local fixes with side effects\.
Each round, the proposer reads only one batch of run records and focuses its edits on the most common failures in that batch, so side effects on other tasks may be noticed only after they recur in later batches\. Both regressions in Section[5\.6](https://arxiv.org/html/2609.38372#S5.SS6)come from mechanisms added for other tasks\. Mixing task types makes such conflicts more likely: an edit has to serve all tasks at once, while the proposer can attend to only some of them each time\.
#### Acceptance without validation\.
A new version is accepted as long as it runs; the framework has no validation set and does not compare the scores of the old and new versions, so whether an edit helps rests almost entirely on the model’s own judgment\. This is where harness evolution differs most from deep\-learning training\. A sufficiently small step along the negative gradient is guaranteed to lower the loss on the current batch, whereas here the update direction is decided by the model, may be wrong, and is not checked by any gate\. The weaker the model, the greater this risk\[[10](https://arxiv.org/html/2609.38372#bib.bib10),[8](https://arxiv.org/html/2609.38372#bib.bib8)\]\. Validating each update at an acceptable cost is a missing piece of the framework\.
#### Freedom without diversity\.
Although the scope of edits is unrestricted, the actual edits narrow to a few mechanisms, and components mentioned in the update prompt, such as skills, memory, and sub\-agents, never appeared\. Rethinking the Evaluation of Harness Evolution also finds that evolution edits usually land on the system prompt first and then move to middleware and the tool layer\[[20](https://arxiv.org/html/2609.38372#bib.bib20)\]\. Methods that partition components in advance at least make each component type an explicit target of modification\[[5](https://arxiv.org/html/2609.38372#bib.bib5)\]; once the scope is opened up, the variety of edits may even shrink\. One possible remedy is to keep multiple versions, as DGM does, so that evolution proceeds along several paths at once\[[30](https://arxiv.org/html/2609.38372#bib.bib30)\]\.
## 6Discussion and Conclusion
#### Limitations\.
We report results for only one strong model; what kind of outer pipeline weaker models would need remains to be studied experimentally \(Section[5\.7](https://arxiv.org/html/2609.38372#S5.SS7)\)\. Evolution is itself stochastic: rerunning the same setup may produce different mechanisms, and estimating this variation would require multiple independent evolution chains\. Some benchmarks have few held\-out tasks\. We evaluate them repeatedly and average the results to reduce variance, but as the analyses in Sections[4\.2](https://arxiv.org/html/2609.38372#S4.SS2)and[5\.6](https://arxiv.org/html/2609.38372#S5.SS6)show, the run\-to\-run randomness of the model still has a sizable effect on individual benchmarks, and small differences should be interpreted with caution\. The second stage uses only one new domain, and whether the conclusions extend to other domains remains to be tested\.
#### Conclusion\.
We propose a new self\-evolving harness framework in which the same frontier model solves tasks on the same harness and directly modifies the harness that runs it, with no human\-designed analysis step in the outer loop and no restriction on the scope of edits\. Evolution tasks mix benchmarks from multiple domains, and the process is organized as the two stages of deep\-learning training, multi\-task pretraining and continual training\. Experiments show that the approach is effective\. Starting from a 49\-line seed, the evolved harness surpasses both the seed harness and Codex in average score on the in\-distribution benchmarks; on the out\-of\-distribution benchmarks never used in evolution, it scores 12\.64 points above the seed harness and on par with Codex; and continual training on a new domain further improves the score in that domain\. The analysis also identifies where the framework can improve: local edits have side effects, updates lack validation, and edit directions tend to narrow\. Validating each update at an acceptable cost and letting evolution proceed along several paths at once are promising ways to further improve harness self\-evolution\.
## References
- \[1\]Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao\.ReAct: Synergizing Reasoning and Acting in Language Models, 2022\.URL[https://arxiv\.org/abs/2210\.03629](https://arxiv.org/abs/2210.03629)\.
- \[2\]Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, and Heng Ji\.Executable Code Actions Elicit Better LLM Agents, 2024\.URL[https://arxiv\.org/abs/2402\.01030](https://arxiv.org/abs/2402.01030)\.
- \[3\]John Yang, Carlos E\. Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press\.SWE\-agent: Agent\-Computer Interfaces Enable Automated Software Engineering, 2024\.URL[https://arxiv\.org/abs/2405\.15793](https://arxiv.org/abs/2405.15793)\.
- \[4\]Yoonho Lee, Roshen Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn\.Meta\-Harness: End\-to\-End Optimization of Model Harnesses, 2026\.URL[https://arxiv\.org/abs/2603\.28052](https://arxiv.org/abs/2603.28052)\.
- \[5\]Jiahang Lin, Shichun Liu, Chengjun Pan, Lizhi Lin, Shihan Dou, Zhiheng Xi, et al\.Agentic Harness Engineering: Observability\-Driven Automatic Evolution of Coding\-Agent Harnesses, 2026a\.URL[https://arxiv\.org/abs/2604\.25850](https://arxiv.org/abs/2604.25850)\.
- \[6\]Xiaotian Luo, Dizhan Xue, Fengxingyu Wang, Chuanrui Hu, and Yafeng Deng\.HarnessBank: Semantic Gene\-Bank Search with Gated Verification for Agent\-Harness Self\-Evolution, 2026\.URL[https://arxiv\.org/abs/2607\.13683](https://arxiv.org/abs/2607.13683)\.
- \[7\]Tingyang Chen, Shuo Lu, Kang Zhao, Weicheng Meng, Hanlin Teng, Tianhao Li, et al\.HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry, 2026\.URL[https://arxiv\.org/abs/2606\.14249](https://arxiv.org/abs/2606.14249)\.
- \[8\]Minhua Lin, Juncheng Wu, Zijun Wang, Zhan Shi, Yisi Sang, Bing He, et al\.Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self\-Evolving LLM Agents, 2026b\.URL[https://arxiv\.org/abs/2605\.30621](https://arxiv.org/abs/2605.30621)\.
- \[9\]Hangfan Zhang, Shao Zhang, Kangcong Li, Chen Zhang, Yang Chen, Yiqun Zhang, Lei Bai, and Shuyue Hu\.Self\-Harness: Harnesses That Improve Themselves, 2026a\.URL[https://arxiv\.org/abs/2606\.09498](https://arxiv.org/abs/2606.09498)\.
- \[10\]Hui Xue and Fan Yang\.Rethinking Self\-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?, 2026\.URL[https://arxiv\.org/abs/2608\.09629](https://arxiv.org/abs/2608.09629)\.
- \[11\]Richard S\. Sutton\.The Bitter Lesson\.Essay, March 13, 2019\.URL[http://www\.incompleteideas\.net/IncIdeas/BitterLesson\.html](http://www.incompleteideas.net/IncIdeas/BitterLesson.html)\.March 13, 2019\.
- \[12\]Irving John Good\.Speculations Concerning the First Ultraintelligent Machine\.In*Advances in Computers*, volume 6, pages 31–88\. Academic Press, 1965\.
- \[13\]Jürgen Schmidhuber\.Gödel Machines: Self\-Referential Universal Problem Solvers Making Provably Optimal Self\-Improvements, 2003\.URL[https://arxiv\.org/abs/cs/0309048](https://arxiv.org/abs/cs/0309048)\.
- \[14\]Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai\.Self\-Taught Optimizer \(STOP\): Recursively Self\-Improving Code Generation, 2023\.URL[https://arxiv\.org/abs/2310\.02304](https://arxiv.org/abs/2310.02304)\.COLM 2024\.
- \[15\]Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhuang, Yoonho Lee, et al\.RRSI: Regularized Recursive Self\-Improvement of Agent Harnesses, 2026\.URL[https://arxiv\.org/abs/2609\.24972](https://arxiv.org/abs/2609.24972)\.
- \[16\]Siwei Wu, Jincheng Ren, Yizhi Li, Haau\-Sing Li, Chengran Yang, Yuxuan Zhang, et al\.ModularRSI: Modular and Generalizable Recursive Harness Self\-Improvement, 2026a\.URL[https://arxiv\.org/abs/2609\.14857](https://arxiv.org/abs/2609.14857)\.
- \[17\]Luan Zhang, Ruochen Zhou, Dandan Song, Zhengyu Chen, Yuhang Tian, Jun Yang, et al\.HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses, 2026b\.URL[https://arxiv\.org/abs/2608\.01918](https://arxiv.org/abs/2608.01918)\.
- \[18\]Wen Jiang, Mingmin Chu, Yimeng Tian, Qianxin Zhang, Haofei Yang, Rui Yang, et al\.HarnessEvolve: Learning from Reference Trajectories for Reliable Agent Self\-Evolution, 2026\.URL[https://arxiv\.org/abs/2609\.00829](https://arxiv.org/abs/2609.00829)\.
- \[19\]Hao Zhou, Haichuan Hu, Tianyu Luo, Ye Shang, Chunrong Fang, Zhenyu Chen, et al\.Self\-Evolving Coding Agents, 2026\.URL[https://arxiv\.org/abs/2608\.03392](https://arxiv.org/abs/2608.03392)\.
- \[20\]Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, et al\.Rethinking the Evaluation of Harness Evolution for Agents, 2026\.URL[https://arxiv\.org/abs/2607\.12227](https://arxiv.org/abs/2607.12227)\.
- \[21\]Yuhao Wu, Jingyuan Zhang, Jiajun Shi, Xinping Lei, Qingshui Gu, Yuxuan Zhang, et al\.HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?, 2026b\.URL[https://arxiv\.org/abs/2609\.01437](https://arxiv.org/abs/2609.01437)\.
- \[22\]Rich Caruana\.Multitask Learning\.*Machine Learning*, 28\(1\):41–75, 1997\.doi:10\.1023/A:1007379606734\.
- \[23\]Yifan Yang, Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, Zisu Huang, et al\.SkillOpt: Executive Strategy for Self\-Evolving Agent Skills, 2026\.URL[https://arxiv\.org/abs/2605\.23904](https://arxiv.org/abs/2605.23904)\.
- \[24\]Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy\.The Pile: An 800GB Dataset of Diverse Text for Language Modeling, 2020\.URL[https://arxiv\.org/abs/2101\.00027](https://arxiv.org/abs/2101.00027)\.
- \[25\]Ilya Loshchilov and Frank Hutter\.SGDR: Stochastic Gradient Descent with Warm Restarts\.In*International Conference on Learning Representations*, 2017\.doi:10\.48550/arXiv\.1608\.03983\.URL[https://arxiv\.org/abs/1608\.03983](https://arxiv.org/abs/1608.03983)\.
- \[26\]Michael McCloskey and Neal J\. Cohen\.Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem\.In*Psychology of Learning and Motivation*, volume 24, pages 109–165\. Academic Press, 1989\.
- \[27\]David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy P\. Lillicrap, and Greg Wayne\.Experience Replay for Continual Learning\.In*Advances in Neural Information Processing Systems*, 2019\.URL[https://arxiv\.org/abs/1811\.11682](https://arxiv.org/abs/1811.11682)\.
- \[28\]Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L\. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish\.Simple and Scalable Strategies to Continually Pre\-train Large Language Models, 2024\.URL[https://arxiv\.org/abs/2403\.08763](https://arxiv.org/abs/2403.08763)\.
- \[29\]Xunjian Yin, Xinyi Wang, Liangming Pan, Li Lin, Xiaojun Wan, and William Yang Wang\.Gödel Agent: A Self\-Referential Agent Framework for Recursive Self\-Improvement, 2024\.URL[https://arxiv\.org/abs/2410\.04444](https://arxiv.org/abs/2410.04444)\.ACL 2025\.
- \[30\]Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune\.Darwin Godel Machine: Open\-Ended Evolution of Self\-Improving Agents, 2025\.URL[https://arxiv\.org/abs/2505\.22954](https://arxiv.org/abs/2505.22954)\.
- \[31\]Maxime Robeyns, Martin Szummer, and Laurence Aitchison\.A Self\-Improving Coding Agent, 2025\.URL[https://arxiv\.org/abs/2504\.15228](https://arxiv.org/abs/2504.15228)\.
- \[32\]Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, et al\.DarwinX: Evolving Agent Harnesses Through Natural Selection, 2026c\.URL[https://arxiv\.org/abs/2608\.07545](https://arxiv.org/abs/2608.07545)\.
- \[33\]Congjie Zheng, Chuanyi Xue, Bin Liang, Jun Yang, and Changshui Zhang\.SEAGym: An Evaluation Environment for Self\-Evolving LLM Agents, 2026\.URL[https://arxiv\.org/abs/2606\.17546](https://arxiv.org/abs/2606.17546)\.
- \[34\]Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, et al\.Evo\-Bench: Can Language Models Improve Agent Harness?, 2026\.URL[https://arxiv\.org/abs/2608\.09096](https://arxiv.org/abs/2608.09096)\.
- \[35\]Yonatan Gideoni, Sebastian Risi, and Yarin Gal\.Simple Baselines are Competitive with Code Evolution, 2026\.URL[https://arxiv\.org/abs/2602\.16805](https://arxiv.org/abs/2602.16805)\.
- \[36\]Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Zhi Huang, Carlos Guestrin, and James Zou\.TextGrad: Automatic "Differentiation" via Text, 2024\.URL[https://arxiv\.org/abs/2406\.07496](https://arxiv.org/abs/2406.07496)\.
- \[37\]Lakshya A Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl\-Ong, et al\.GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning, 2025\.URL[https://arxiv\.org/abs/2507\.19457](https://arxiv.org/abs/2507.19457)\.
- \[38\]Zewen Liu, Zhan Shi, Yisi Sang, Bing He, Minhua Lin, Tianxin Wei, et al\.Adaptive Auto\-Harness: Sustained Self\-Improvement for Agentic System Deployment on Open\-Ended Task Streams, 2026\.URL[https://arxiv\.org/abs/2606\.01770](https://arxiv.org/abs/2606.01770)\.
- \[39\]Seth Karten, Joel Zhang, Tersoo Upaa, Jr\., Ruirong Feng, Wenzhe Li, Chengshuai Shi, et al\.Continual Harness: Online Adaptation for Self\-Improving Foundation Agents, 2026\.URL[https://arxiv\.org/abs/2605\.09998](https://arxiv.org/abs/2605.09998)\.
- \[40\]Diederik P\. Kingma and Jimmy Ba\.Adam: A Method for Stochastic Optimization\.In*International Conference on Learning Representations*, 2015\.URL[https://arxiv\.org/abs/1412\.6980](https://arxiv.org/abs/1412.6980)\.
- \[41\]Ilya Loshchilov and Frank Hutter\.Decoupled Weight Decay Regularization\.In*International Conference on Learning Representations*, 2019\.URL[https://arxiv\.org/abs/1711\.05101](https://arxiv.org/abs/1711.05101)\.
- \[42\]Mike A\. Merrill, Alexander G\. Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, et al\.Terminal\-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces, 2026\.URL[https://arxiv\.org/abs/2601\.11868](https://arxiv.org/abs/2601.11868)\.
- \[43\]Romain Froger, Pierre Andrews, Matteo Bettini, Amar Budhiraja, Ricardo Silveira Cabral, Virginie Do, et al\.Gaia2: Benchmarking LLM Agents on Dynamic and Asynchronous Environments, 2026\.URL[https://arxiv\.org/abs/2602\.11964](https://arxiv.org/abs/2602.11964)\.
- \[44\]Krista Opsahl\-Ong, Arnav Singhvi, Jasmine Collins, Ivan Zhou, Cindy Wang, Ashutosh Baheti, et al\.OfficeQA Pro: An Enterprise Benchmark for End\-to\-End Grounded Reasoning, 2026\.URL[https://arxiv\.org/abs/2603\.08655](https://arxiv.org/abs/2603.08655)\.
- \[45\]Quan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan, and Victor Barres\.τ\\tau\-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge, 2026\.URL[https://arxiv\.org/abs/2603\.04370](https://arxiv.org/abs/2603.04370)\.
- \[46\]Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, et al\.SWE\-Bench Pro: Can AI Agents Solve Long\-Horizon Software Engineering Tasks?, 2025\.URL[https://arxiv\.org/abs/2509\.16941](https://arxiv.org/abs/2509.16941)\.
- \[47\]Carlos E\. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan\.SWE\-bench: Can Language Models Resolve Real\-World GitHub Issues?In*International Conference on Learning Representations*, 2024\.URL[https://arxiv\.org/abs/2310\.06770](https://arxiv.org/abs/2310.06770)\.
- \[48\]OpenAI\.Introducing SWE\-bench Verified\.[https://openai\.com/index/introducing\-swe\-bench\-verified/](https://openai.com/index/introducing-swe-bench-verified/), 2024\.
- \[49\]Datacurve AI\.DeepSWE: Measuring Frontier Coding Agents on Original, Long\-Horizon Engineering Tasks\.[https://github\.com/datacurve\-ai/deep\-swe](https://github.com/datacurve-ai/deep-swe), 2026\.Version 1\.1\.
- \[50\]Zijian Chen, Xueguang Ma, Shengyao Zhuang, Ping Nie, Kai Zou, Andrew Liu, et al\.BrowseComp\-Plus: A More Fair and Transparent Evaluation Benchmark of Deep\-Research Agent, 2025\.URL[https://arxiv\.org/abs/2508\.06600](https://arxiv.org/abs/2508.06600)\.
- \[51\]Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, et al\.APEX\-Agents, 2026\.URL[https://arxiv\.org/abs/2601\.14242](https://arxiv.org/abs/2601.14242)\.
- \[52\]Bowen Ye, Rang Li, Qibin Yang, Yuanxin Liu, Linli Yao, Hanglong Lv, Zhihui Xie, Chenxin An, Lei Li, Lingpeng Kong, Qi Liu, Zhifang Sui, and Tong Yang\.Claw\-Eval: Towards Trustworthy Evaluation of Autonomous Agents, 2026\.URL[https://arxiv\.org/abs/2604\.06132](https://arxiv.org/abs/2604.06132)\.
- \[53\]Sydney Von Arx, Lawrence Chan, and Beth Barnes\.Recent Frontier Models Are Reward Hacking\.METR blog,[https://metr\.org/blog/2025\-06\-05\-recent\-reward\-hacking/](https://metr.org/blog/2025-06-05-recent-reward-hacking/), 2025\.June 5, 2025\.
- \[54\]Sayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir, Zachary S\. Siegel, Boyi Wei, et al\.Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation, 2025\.URL[https://arxiv\.org/abs/2510\.11977](https://arxiv.org/abs/2510.11977)\.
## Appendix ASeed Harness Code
The complete code of the seed harness \(Section[3\.2](https://arxiv.org/html/2609.38372#S3.SS2)\) is shown below\.
Listing 1:Complete code of the seed harness \(49 lines\)\.1SYSTEM\_PROMPT=\(
2"YouareanautonomousagentworkinginsideaLinuxcontainer\.\\n"
3"Usethebashtooltoinspecttheenvironmentandcompletethetask\.\\n"
4"Nevercheat:donotobtaintheanswerbysearchingtheinternetorany"
5"similarshortcut;solvethetaskyourself\.\\n"
6"Whenthetaskisdone,replywithafinalmessageandnotoolcalls\."
7\)
8
9TOOLS=\[
10\{
11"type":"function",
12"function":\{
13"name":"bash",
14"description":"Runabashcommandinthetaskcontainerandreturnitsoutput\.",
15"parameters":\{
16"type":"object",
17"properties":\{
18"command":\{"type":"string","description":"Thecommandtorun\."\}
19\},
20"required":\["command"\],
21\},
22\},
23\}
24\]
25
26
27asyncdefrun\(instruction,model,world\):
28messages=\[
29\{"role":"system","content":SYSTEM\_PROMPT\},
30\{"role":"user","content":instruction\},
31\]
32whileTrue:
33reply=awaitmodel\.chat\(messages,tools=TOOLS\)
34messages\.append\(reply\.message\)
35ifnotreply\.tool\_calls:
36return
37forcallinreply\.tool\_calls:
38ifcall\.name=="bash":
39r=awaitworld\.exec\(str\(call\.arguments\.get\("command",""\)\)\)
40parts=\[r\.stdout\]
41ifr\.stderr:
42parts\.append\("\[stderr\]\\n"\+r\.stderr\)
43parts\.append\(f"\[exit\{r\.return\_code\}\]"\)
44output="\\n"\.join\(pforpinpartsifp\)
45else:
46output=f"unknowntool:\{call\.name\}"
47messages\.append\(
48\{"role":"tool","tool\_call\_id":call\.id,"content":output\}
49\)
## Appendix BUpdate Task Prompt
The update prompt that the proposer receives at each iteration is shown below\. It is given as an ordinary task to the agent running on the current harness, whose system prompt is the current harness’s own system prompt; in the first iteration, this is the system prompt in Listing[1](https://arxiv.org/html/2609.38372#LST1)\. The last paragraph is the fixed requirement appended to every task description \(Section[4\.1](https://arxiv.org/html/2609.38372#S4.SS1)\)\.<a\_t\>is the budget of new lines for the iteration \(Section[3\.5](https://arxiv.org/html/2609.38372#S3.SS5)\); apart from it, the prompt is identical in both stages\.
Listing 2:Full text of the update prompt\.Youareimprovingtheagentharnessthatrunsyou\.Thegoalisaharnessthatgeneralizes:itwillbeevaluatedontasksithasnotseen,somefromthesesamebenchmarksandsomefromentirelydifferentones\.Memorizingspecifictasksortheirsolutionsisthereforeworthless,andsoishard\-codingwordingsorrulesthatrestateasingletrial,suchasafixedclarificationquestionoradictatedreportphrasing\.Lookforthegeneral,high\-levelimprovementstheevidencepointsto;guidanceshouldteachamethodandleavethewordingtothemodel\.
‘/app/harness\_workspace‘holdtheharness\.‘/app/harness\_kernel‘isthefixedkernelthatloadstheworkspace;readitforcontext,butitisnoteditableinthistask\.‘/app/eval\_trials/‘holdsthetrialrecordsofthisharness’spastruns\.Editscountonlyunder‘/app/harness\_workspace‘\.
Youarefreetoreshapetheharnesshoweveryouseefit\-\-add,split,ordeletefiles,andintroducewhatevercomponentsyoujudgeitneeds:skills,tools,prompts,memory,subagents,contextmanagement,planning,oranythingofyourowndesign\.Overmanyroundsthisshouldgrowintoacomplete,well\-roundedharnessthathandlesallkindsoftasks,soexploreboldlyratherthankeepreworkingonespot\.
Thisisoneroundofaloop\.Theharnessyouleavehereistheonethatrunsthenextbatchoftasksandtheonethatrunsthiseditingtaskagainonthosetrials\-\-roundafterround\.
Theworkspacecarriesapermanentnotebook,‘harness\_workspace/evolve\_notes\.md‘,thattravelswiththeharnessfromroundtoround;bringituptodatebeforeyoufinish\.Itcanrecordwhetherearliereditsdeliveredwhatwasexpected,whyorwhynot,andwhethertheirgainsgeneralizedtothisround’stasks;yourreadingofthecurrentevidence;whatyouchangedandwhatyouexpectittoimprove;whichideashaveprovenineffective;andwhateverelseyoujudgeworthrecording\.Feelfreetorewriteordeletenotesthathavebecomeoutdatedordisproven\.Donottrytorecordthespecificsofanyonetaskinthenotebook\-\-thesametaskwillnotappearagain\.
Fileswrittenunder‘/logs/agent‘inanytaskcontainer,thiseditingoneincluded,arecollectedintothattrial’s‘agent/‘directory,andthosetrialsarewhatthenexteditingroundsees\.Thatishowthemodel,nowandduringthecomingtasks,chooseswhattokeepforfuturerounds:atrajectoryinitsownformat,notesonwhatwentwrong,whateveritjudgesworthkeeping\.
ThemodelAPIacceptsatmost258,400inputtokensperrequest;overthat,therequestfailswith‘context\_length\_exceeded‘\.
Acceptancechecksrunafteryoufinish:changesmustbeconfinedto‘harness\_workspace/‘andmustnotbeemptyasidefrom‘evolve\_notes\.md‘,and‘cd/app&&python\-mharness\_kernel\.smoke/app/harness\_workspace‘mustprint‘SMOKEOK‘\.Thisroundyouareallowedtoaddatmost<a\_t\>newlines;editsto‘evolve\_notes\.md‘donotcount,andneitherdodeletedlines\.Thesmokecommandcountsandchecksthisforyou\.Leaveyourchangesuncommitted\.
Nevercheat:donotobtaintheanswerfromanysourcethatcarriesitratherthansolvingthetask\-\-theupstreamfixoritspullrequest,apublishedsolutionortestfile,adatasetcopyoftheanswer,oranoracle/evaluationartifactleftintheenvironment\.Readingordinarydocumentationandlibrarysourceisfine\.Solvethetaskyourself\.’
## Appendix CMechanism Trigger Rates
Figure[5](https://arxiv.org/html/2609.38372#A3.F5)shows, for each benchmark, how often Seed aborted from context overflow and how often the three mechanisms of Iteration 5 were triggered\. Overflow in Seed concentrates on BrowseComp\-Plus\. Truncation is most common in retrieval and document tasks, compaction is rare overall, and independent review runs on nearly every non\-interactive task\. GAIA2 andτ3\\tau^\{3\}\-Bench banking never trigger truncation or review: their tools return short outputs, and both are handled as interactive tasks\.
Figure 5:Share of tasks \(%\) on each benchmark in which Seed aborted because of context overflow, and in which each of three mechanisms of Iteration 5 was triggered, computed over all Seed and Iteration 5 evaluation records used in Figure[3](https://arxiv.org/html/2609.38372#S4.F3)\. Truncation: the model received a clipped tool output\. Compaction: earlier turns of the conversation were removed\. Independent review: at the end of the task, the result was checked again in a fresh context\.Similar Articles
@NFTCPS: HarnessX is pretty interesting: an agent architecture that can modify itself. Previously, architectural changes relied entirely on manual tuning. When a new model came out, Anthropic removed the planning steps from Claude Code, and Manus refactored its agents five times in six months, each time simplifying. What to change and when to change it — all decided by humans.
HarnessX introduces a framework for self-evolving AI agent harnesses that treats the runtime harness as a first-class object, enabling automatic adaptation via trace-driven reinforcement learning. It achieves average gains of +14.5% across five benchmarks, with larger improvements for weaker models.
@geekbb: Auto-optimization tool for Agent harness. It takes over the heavy lifting of harness optimization: you provide a benchmark command and a target repository, and it automatically generates proposals, runs evaluations, records results, keeps the best, discards the rest, and automatically improves the agent's prompts, configurations, and source code. https…
autoharness is an automated agent harness optimization tool that automatically generates proposals and runs evaluations based on benchmark commands to improve an agent's prompts, configurations, and source code. It supports Codex and Claude.
Self-Harness: Harnesses That Improve Themselves
Self-Harness introduces a new paradigm where LLM-based agents iteratively improve their own operating harness by mining model-specific weaknesses, proposing harness modifications, and validating them through regression testing, achieving substantial performance gains on Terminal-Bench-2.0 across multiple base models.
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
HarnessX is a foundry for composable, adaptive, and evolvable AI agent harnesses that uses compositional primitives and trace-driven evolution to improve agent performance. Across five benchmarks, it achieves an average gain of +14.5% (up to +44.0%), demonstrating that runtime interface evolution is a complementary lever to model scaling.
@akshay_pachaar: self-evolving harnesses are here. (100% open-source) today you pick a fixed harness, and every task runs through it. a …
JIT-Agent is an open-source 27B model that dynamically generates task-specific harnesses for AI agents, outperforming hand-built systems with improved token efficiency.