VERSE:面向 Agent Harness 的可验证自进化优化器
摘要
论文提出了 VERSE,一个可验证的自进化优化器,利用基于执行的验证机制,同时改进执行器 agent 的 harness 以及自身所用的提示词、工具和流程。在留出的 SWE-rebench 任务上,VERSE 达到了 42.3% 和 37.7% 的准确率,而最强基线仅为 39.2% 和 29.3%。
查看缓存全文
缓存时间: 2026/10/05 09:58
# VERSE: Verified Self-Evolving Optimizer for Agent Harnesses
Source: [https://arxiv.org/html/2610.02616](https://arxiv.org/html/2610.02616)
Zekai Wang , Yingqiang Ge, Zekun Wang††footnotemark:, Hai Wang, Yuhui Xu,††thanks:Work done during an internship at Amazon\. Correspondence to: Zekai Wang <zekai@mit\.edu\>, Yingqiang Ge <gyq@amazon\.com\>\.Joshua Frandsen, Shancong Fu, Ashia C\. Wilson, Chandan K\. ReddyAffiliation:MITAffiliation:Amazon
###### Abstract
Harness evolution improves an LLM agent’s prompts, tools, and workflow, while the optimizer’s own tools and procedures often remain fixed\. We study whether an optimizer can improve another agent more effectively by also improving how it diagnoses failures, develops edits, and tests their effects\. Two observations guide our design\. In a controlled study, optimizer self\-evolution fails to improve performance without execution\-based verification, but achieves the best result of that study when verification is available\. Across five executors, self\-evolving optimizers build their own tools for failure analysis, verification, training audits, and workflow control\. Motivated by these findings, we introduce VERSE, aVerifiedSelf\-Evolving optimizer for agent harnesses\. VERSE lets the optimizer test draft edits, replay failures, and perturb suspected steps before submission, while tracking fixes and regressions across rounds\. Using this feedback, the optimizer revises both the executor harness and its own prompts, skills, tools, hooks, and notes, while the weights of the optimizer and executor models stay fixed\. Under a shared protocol with disjoint training, validation, and test tasks, VERSE improves all four evaluated harness optimizers on held\-out SWE\-rebench tasks and newer out\-of\-distribution tasks in five languages\. Its best validation\-selected harness reaches42\.3%42\.3\\%and37\.7%37\.7\\%accuracy, respectively, against39\.2%39\.2\\%and29\.3%29\.3\\%for the strongest baselines\. Code is available at[https://github\.com/wzekai/VERSE](https://github.com/wzekai/VERSE)\.
## 1Introduction
An LLM agent is a model wrapped in a harness: the prompts, tools, memory, and control flow that turn a language model into a system that acts\. On agentic benchmarks such as SWE\-bench\([Jimenez et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib17)\)and Terminal\-Bench\([Merrill et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib27)\), the harness often matters as much as the model\. In*harness evolution*, an LLM optimizer reads the failed trajectories of a frozen executor agent and edits the executor’s harness, over several rounds\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21);[Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23);[Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55);[Chen et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib9);[Nie et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib29)\)\. The optimizer’s own harness, however, stays fixed: it is a hand\-built pipeline or an off\-the\-shelf coding agent\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21);[Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23);[Chen et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib9);[Park et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib31)\), or, when one model plays both roles, a fixed loop that proposes and accepts edits\([Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55)\)\. Lessons from earlier rounds shape the executor’s next harness, but not how the optimizer works\. Can the optimizer evolve its own harness as well?
Simply allowing it did not help \(Section[2](https://arxiv.org/html/2610.02616#S2)\)\. When we added self\-evolution to Meta\-Harness\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21)\), the optimizer wrote untested lessons into its own harness, and no later round beat the initial harness on validation\. The optimizer learns only from trajectories and scores, and it sees the effect of an edit one round later\. Reading trajectories can mislead: optimizers can report failures that never occurred\([Wang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib41)\), and evolved harnesses can gain little on new tasks\([Wang et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib42)\)\. Once the same optimizer could run a draft edit on a training task before submitting it, self\-evolution gave the best result among the settings we compared\. Execution checks give a self\-evolving optimizer evidence about which of its lessons hold\.
A second study shows what a self\-evolving optimizer builds for itself\. Across five executors, it tends to build tools that find the causes of failures, checks that a draft works, notes on which fixes held, and rules for its own workflow\. Some recent self\-improving systems also revise their own improvement procedure\([Wang et al\., 2026d](https://arxiv.org/html/2610.02616#bib.bib44);[Zhang et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib57);[Zhou, 2026](https://arxiv.org/html/2610.02616#bib.bib59)\)\(Appendix[A](https://arxiv.org/html/2610.02616#A1)\)\. Together, the two studies suggest a design: give the optimizer reliable execution checks from the first round, and let it evolve how it uses them\.
These observations motivate our method, VERSE, which combines execution checks with optimizer self\-evolution\. VERSE provides from the first round what the self\-evolving optimizer built for itself\. Trace minimization cuts a failed trajectory down to the few steps that still reproduce the failure\. We design three verification tools, and the optimizer decides when and how to call them: they run a draft on the tasks it targets, replay a recorded failure to check that it reproduces, and change a single step to test whether that step causes the failure\. A training audit records which fixes held in later rounds and which broke\. Between rounds, the optimizer revises its own prompts, skills, tools, hooks, and notes\. This makes harness evolution a bilevel process: the inner loop edits the executor’s harness, and the outer loop edits the optimizer’s own harness \(Section[3\.1](https://arxiv.org/html/2610.02616#S3.SS1)\)\.
Figure 1:Fixed versus self\-evolving optimizers\.Left: in prior harness optimizers, the optimizer is fixed: it reads trajectories and scores, edits the executor harnesshhthat wraps a frozen base model, and keeps its own harness unchanged\. Right: a self\-evolving optimizer adds an outer loop: after each round, it updates its own prompts, skills, tools, hooks, and notes from the round’s results\.We study this design in theory and in experiments\. Our contributions are as follows\.
- •Two studies: optimizer self\-evolution helped only with execution\-based verification, and the optimizer builds its own tools for failure analysis and verification\. They motivate VERSE, which provides such tools and lets the optimizer evolve its own harness \(Sections[2](https://arxiv.org/html/2610.02616#S2)–[3](https://arxiv.org/html/2610.02616#S3)\)\.
- •A theoretical framework that views harness evolution as bilevel learning to generalize to new tasks\. We prove that when the recorded evidence is ambiguous, reading trajectories alone leaves a minimum rate of wrong repairs, which executed checks can lower, and that wrong repairs limit what self\-evolution can gain \(Section[3\.2](https://arxiv.org/html/2610.02616#S3.SS2)\)\.
- •Experiments on SWE\-rebench, where VERSE improves all four published harness optimizers it is added to, both on the test set and on newer out\-of\-distribution tasks\. In the evolution logs, few draft edits fix their targets, and self\-evolving optimizers add regression checks \(Section[4](https://arxiv.org/html/2610.02616#S4)\)\.
## 2Setting and Motivating Observations
Figure 2:Verification during harness evolution\.Prior optimizers review training trajectories differently \(Appendix[E](https://arxiv.org/html/2610.02616#A5)\), then edit the executor harness and rerun training tasks to collect new trajectories\. Simple verification lets the optimizer execute a draft while editing and send its result and trajectory back for review, before submission and within budget\. The validation summary has no task content, to limit overfitting \(Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2)\)\. For the self\-evolving optimizer \(ours\), an outer loop updates the optimizer’s own harness between rounds \(Figure[1](https://arxiv.org/html/2610.02616#S1.F1)\)\.We first describe the setup and the task splits used throughout the paper\. We then present two studies that motivate VERSE\. The first evaluates the role of verification in optimizer self\-evolution\. The second examines the tools, instructions, and notes that the optimizer builds for itself\.
### 2\.1Setup
A self\-evolving optimizer can rewrite five parts of its own harness\.*Prompts*: the optimizer’s system prompts\.*Skills*: instruction files added to the optimizer’s prompt every round\.*Tools*: functions it writes and can call\.*Hooks*: code that runs automatically at fixed points of the optimizer’s own loop, for example to check a tool call before it runs or to add a reminder after each turn\.*Notes*: its memory of earlier rounds\. Self\-written code is loaded only after it passes a safety check and a test run\. Everything else is fixed: how harnesses are evaluated and selected, and the turn and execution budgets of each round\. Keeping these fixed makes methods comparable under the same protocol\.
For Observation 1, the optimizer is Qwen3\.8\-Flash\-Next \(176B total parameters, 6B active\) and the executor is Qwen3\.8\-27B\. Both stay frozen through six evolution rounds on SWE\-rebench\([Badertdinov et al\., 2025](https://arxiv.org/html/2610.02616#bib.bib3)\)\. We start from Meta\-Harness\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21)\)and add simple verification \(Figure[2](https://arxiv.org/html/2610.02616#S2.F2)\), self\-evolution \(Figure[1](https://arxiv.org/html/2610.02616#S1.F1)\), or both \(Table[1](https://arxiv.org/html/2610.02616#S2.T1)\)\. For Observation 2, we run the self\-evolving optimizer with simple verification with five executors and audit what it builds for itself\.
### 2\.2Training, validation, and test splits
Many harness optimizers search and report results on the same benchmark\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21);[Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23);[Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55);[Chen et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib9);[Zhang et al\., 2026d](https://arxiv.org/html/2610.02616#bib.bib58)\)\.[Wang et al\. \(2026b\)](https://arxiv.org/html/2610.02616#bib.bib42)point out that such a search resembles test\-time scaling, and find that its gains can come from overfitting to task\-specific patterns rather than from a better harness design\. We study how to evolve harnesses that generalize to unseen tasks\. We therefore use three disjoint task sets, as is standard in machine learning\([Hastie et al\., 2009](https://arxiv.org/html/2610.02616#bib.bib12)\), and measure final performance only on tasks never used during evolution\.
In the main text, we use SWE\-rebench\([Badertdinov et al\., 2025](https://arxiv.org/html/2610.02616#bib.bib3)\), a continuously refreshed pool of GitHub issue tasks in the SWE\-bench format\([Jimenez et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib17)\)\. We draw 110 training and 50 validation tasks from a pool of 593 tasks in its monthly releases \(January 2025 to February 2026\)\. The test set is the 108 gradable tasks of the March 2026 release\. All three sets consist of Python repositories, so we treat this test set as in\-distribution\. For an out\-of\-distribution \(OOD\) test set, we use the July 2026 release: its 107 gradable tasks are newer than every task in the three sets and span Go, Java, Python, Rust, and TypeScript, with only 20 in Python\.
Table 1:Self\-evolution with and without a verification tool\(SWE\-rebench avg@3 accuracy, %,±\\pmSEM over three test evaluations; one evolution run per row\)\.*Val\-selected*: the harness selected by validation accuracy, from roundr⋆r^\{\\star\};*Last*: the final\-round harness\. Atr⋆=0r^\{\\star\}=0, validation selectsh0h\_\{0\}itself, and the entry is an independent re\-evaluation of it\.We select harnesses by validation accuracy, the percentage of validation tasks a harness resolves\. This harness selection mirrors model selection in machine learning, where the checkpoint with the best validation performance is kept rather than the last one\([Prechelt, 1998](https://arxiv.org/html/2610.02616#bib.bib32)\)\. We report the validation\-selected harness as our main result, together with the final\-round harness\. On the in\-distribution test sets of Tables[1](https://arxiv.org/html/2610.02616#S2.T1)–[5](https://arxiv.org/html/2610.02616#S4.F5), the validation\-selected harness matches or beats the final\-round harness in 15 of 18 runs, by up to 6\.8 points, and trails it by at most 1\.6 points\. These results support validation\-based selection over always reporting the final round\.
Validation accuracy can also guide hyperparameter and architecture search\([Snoek et al\., 2012](https://arxiv.org/html/2610.02616#bib.bib36);[Zoph & Le, 2017](https://arxiv.org/html/2610.02616#bib.bib60)\); both uses reduce the validation set to a single number\. Our optimizer is a language model and, like other LLM\-based optimizers\([Yuksekgonul et al\., 2025](https://arxiv.org/html/2610.02616#bib.bib52);[Agrawal et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib1)\), can also learn from text feedback\. After each round, we therefore give it a short text summary of its validation results: its validation accuracy in every round, the identifiers of the validation tasks its latest edit fixed or broke and whether each change held on a rerun, and how often its harness code raised errors\. The optimizer never sees the content of validation tasks or their trajectories\. This summary shows how well its edits generalize, which training accuracy cannot, because the optimizer has read every training trajectory\. Validation therefore both guides later edits and selects the final harness, so validation accuracy can overstate test accuracy; Appendix[C\.3\.2](https://arxiv.org/html/2610.02616#A3.SS3.SSS2)bounds this gap by how much the feedback reveals about the validation set\([Dwork et al\., 2015](https://arxiv.org/html/2610.02616#bib.bib10);[Blum & Hardt, 2015](https://arxiv.org/html/2610.02616#bib.bib6)\)\.
### 2\.3Two Observations on Self\-Evolving Optimizers
Observation 1: self\-evolution helped with verification\.Here, verification means letting the optimizer try an executor harness edit on a task*before submitting it*\. The baseline reads past trajectories, submits a new harness, and learns how it performs from the next evaluation\. With simple verification, the optimizer can first run one training task with its edited harness, see whether the tests pass, and read the executor’s trajectory\. It can then change the harness and try again within the remaining call budget before submitting \(Figure[2](https://arxiv.org/html/2610.02616#S2.F2)\)\.
Without verification, adding self\-evolution to Meta\-Harness\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21)\)made it worse: no round beat round 0 on validation, so validation selected the initial harness over every evolved one \(r⋆=0r^\{\\star\}=0in Table[1](https://arxiv.org/html/2610.02616#S2.T1)\)\. The optimizer could not test its edits before submitting them, so it could not tell which ones were wrong \(Appendix[J\.2](https://arxiv.org/html/2610.02616#A10.SS2)shows the failed edits and why they failed\)\. This suggests that, without verification, a self\-evolving optimizer can be misled by noisy validation results and write wrong lessons into its own harness\. With the verification tool, the same optimizer could run a draft on a training task before submission and reached41\.05%41\.05\\%test accuracy, the highest in Table[1](https://arxiv.org/html/2610.02616#S2.T1)\.
Observation 2: what the optimizer builds for itself\.We run the self\-evolving optimizer with simple verification \(Figure[2](https://arxiv.org/html/2610.02616#S2.F2)\) on SWE\-rebench with five different executor models, and audit how the optimizer evolves its own harness to improve the executor’s performance \(Figure[3](https://arxiv.org/html/2610.02616#S3.F3)\)\. Every run starts from an empty workspace\. Appendix[K](https://arxiv.org/html/2610.02616#A11)lists and classifies every artifact\.
The five runs evolve 51 artifacts in total, each a prompt, skill, tool, hook, or notes file\. Regardless of the executor, we find that the optimizer tends to build four kinds of artifacts\.*Attribution*: tools and skills that find and rank the causes of failures\.*Verification*: checks that a draft harness runs correctly and actually improves the executor\.*Training audit*: notes that track across rounds which tasks were fixed, which regressed, and which fixes were later undone\.*Workflow*: rules and hooks that change how the optimizer itself works, such as budget reminders and instructions for its parallel candidates\.
## 3VERSE
Section[2](https://arxiv.org/html/2610.02616#S2)shows that self\-evolution helped when the optimizer could verify its edits by execution, and that the optimizer repeatedly built its own tools for attribution, verification, and training audit\. VERSE therefore provides these tools from the start and lets the optimizer adapt how it uses them\.
### 3\.1Overview of VERSE
VERSE is added to an existing harness optimizer, which we call the host, such as Meta\-Harness\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21)\)\. The hosts differ mainly in how they analyze failed trajectories and report the causes to the optimizer \(Appendix[E](https://arxiv.org/html/2610.02616#A5)\)\. VERSE keeps this report and adds four components on top: attribution, verification, a training audit, and self\-evolution of the optimizer harnessgg\(Figure[4](https://arxiv.org/html/2610.02616#S3.F4); pseudocode in Appendix[B](https://arxiv.org/html/2610.02616#A2)\)\.
Attribution\.Attribution proposes which steps of a failed training trajectory caused the failure\. Failed trajectories are long, which makes the cause hard to find, so VERSE adds trace minimization: it re\-runs subsets of its steps and keeps the few steps that still produce the same error\. In our replay study, the median failed trajectory shrinks from 129 steps to 8 \(Appendix[I](https://arxiv.org/html/2610.02616#A9)\)\.
Figure 3:What the optimizer builds for itself\.The self\-evolving optimizer with simple verification runs six rounds on SWE\-rebench with five open\-weight executors \(Appendix[K](https://arxiv.org/html/2610.02616#A11)gives the setup of each run\)\. Across these runs the optimizer evolved 51 artifacts for its own harness, which we audit: \(a\) artifacts by executor, colored by function; \(b\) the same artifacts by type\.Verification\.Simple verification \(Section[2](https://arxiv.org/html/2610.02616#S2)\) runs a draft on one training task and leaves the result for the optimizer to interpret\. VERSE adds three tools, each answering one question before submission\. Theverificationtool asks whether a draft fixes the tasks it targets: it runs the draft on up to three of them and reports the failing tests and any errors in harness code\. A task counts as fixed only if it passes twice, since a single pass can be luck\. Thereplaytool asks whether a recorded failure is reproducible: it re\-runs the recorded trajectory\. Theperturbationtool asks whether the failure depends on a given step: it removes or replaces that step in a reproducible trajectory and replays the rest\. The first tool checks a fix; the other two check a diagnosis\.
Training audit\.In Observation 2, every run kept round\-by\-round notes of its outcomes\. VERSE keeps this record from the start\. For each failure mode found by attribution, it tracks which training tasks still fail with it after every round\. The optimizer can thus see which failures earlier edits removed and which returned\.
Figure 4:Overview of VERSE\.Each round, attribution and the training audit report on the latest training runs, and the optimizer calls the verification tools while it edits the executor harness\. Between rounds, the optimizer revises its own harness; the verification tools stay fixed, and it can build new tools on top of them\.Optimizer self\-evolution\.At the end of each round, the optimizer spends up to 30 turns on a self\-update\. It reviews the round’s results, including validation accuracy, whether the tasks it expected to fix were fixed, and its tool usage and errors\. It then edits its own harnessgg\(prompts, skills, tools, hooks, and notes; Section[2\.1](https://arxiv.org/html/2610.02616#S2.SS1)\)\. It also keeps simple verification, which shares the budget of the verification tools\. These tools stay fixed, but its own code can call them\.
### 3\.2Theoretical Analysis
Improving a harness requires deciding what to fix and how to fix it\. We analyze how verification helps identify the right repair, and how self\-evolution can improve the search for a harness that performs well on new tasks\. Formal statements and proofs appear in Appendix[C](https://arxiv.org/html/2610.02616#A3)\.
Framework\.A taskxxincludes a problem, an environment, and a success criterion\. The executor has fixed model weightsθ\\thetaand a harnesshhspecifying its prompts, tools, and workflow\. Its actions and the environment responses form a trajectoryy∼πθ\(⋅∣x;h\)y\\sim\\pi\_\{\\theta\}\(\\cdot\\mid x;h\)\. The lossℓ\(x,y\)\\ell\(x,y\)is00for success and11for failure\. The population lossLθ,D\(h\)L\_\{\\theta,D\}\(h\)averages this loss over new tasksx∼Dx\\sim Dand their execution outcomes\. The optimizer has fixed weightsω\\omegaand a separate harnessggconsisting of its prompts, skills, tools, hooks, and notes \(Section[2](https://arxiv.org/html/2610.02616#S2)\); let𝒯\(g\)\\mathcal\{T\}\(g\)denote its tools\. The initial optimizer harnessg0g\_\{0\}has tools𝒯0=𝒯\(g0\)=𝒯basic∪𝒯verify\\mathcal\{T\}\_\{0\}=\\mathcal\{T\}\(g\_\{0\}\)=\\mathcal\{T\}\_\{\\mathrm\{basic\}\}\\cup\\mathcal\{T\}\_\{\\mathrm\{verify\}\}, where
𝒯basic\\displaystyle\\mathcal\{T\}\_\{\\mathrm\{basic\}\}=\{read,search,bash,write\_notes,submit\_proposal\},\\displaystyle=\\\{\\texttt\{read\},\\ \\texttt\{search\},\\ \\texttt\{bash\},\\ \\texttt\{write\\\_notes\},\\ \\texttt\{submit\\\_proposal\}\\\},\(1\)𝒯verify\\displaystyle\\mathcal\{T\}\_\{\\mathrm\{verify\}\}=\{verification,replay,perturbation\}\.\\displaystyle=\\\{\\texttt\{verification\},\\ \\texttt\{replay\},\\ \\texttt\{perturbation\}\\\}\.All methods share𝒯basic\\mathcal\{T\}\_\{\\mathrm\{basic\}\}\(Appendix[C](https://arxiv.org/html/2610.02616#A3)\), and Section[3\.1](https://arxiv.org/html/2610.02616#S3.SS1)describes𝒯verify\\mathcal\{T\}\_\{\\mathrm\{verify\}\}\. For a self\-evolving optimizer,𝒯verify\\mathcal\{T\}\_\{\\mathrm\{verify\}\}also includes simple verification\. Later optimizer harnesses keep all tools in𝒯0\\mathcal\{T\}\_\{0\}, so𝒯0⊆𝒯\(g\)\\mathcal\{T\}\_\{0\}\\subseteq\\mathcal\{T\}\(g\), and can add new tools built from them\.
We separate tasks into training, validation, and test setsStrain,Sval,StestS\_\{\\mathrm\{train\}\},S\_\{\\mathrm\{val\}\},S\_\{\\mathrm\{test\}\}, with empirical lossesL^train,L^val,L^test\\widehat\{L\}\_\{\\mathrm\{train\}\},\\widehat\{L\}\_\{\\mathrm\{val\}\},\\widehat\{L\}\_\{\\mathrm\{test\}\}on these sets\. In the inner loop, the optimizer edits and checkshth\_\{t\}on training tasks, generating candidates𝒞t\(g\)\\mathcal\{C\}\_\{t\}\(g\)within budgetBtB\_\{t\}\. Validation losses and outcome summaries from earlier rounds are available in both loops, for every method\. In the outer loop, they and the round audits guide updates togg\. The objectives are
inner:\\displaystyle\\text\{inner:\}ht\+1\(g\)∈argminh∈𝒞t\(g\)L^train\(h\),\\displaystyle h\_\{t\+1\}\(g\)\\in\\argmin\_\{h\\in\\mathcal\{C\}\_\{t\}\(g\)\}\\widehat\{L\}\_\{\\mathrm\{train\}\}\(h\),\(2\)outer:\\displaystyle\\text\{outer:\}ming𝔼\[L^val\(ht\+1\(g\)\)\]\.\\displaystyle\\min\_\{g\}\\ \\mathbb\{E\}\[\\widehat\{L\}\_\{\\mathrm\{val\}\}\(h\_\{t\+1\}\(g\)\)\]\.The expectation is taken over candidate generation and evaluation, with the task sets fixed\. The inner loop optimizes the executor harnesshhfor a givengg; the outer loop updates the optimizer harnessggto improve how it generates and checks candidates\([Zakerinia et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib53);[Sucker et al\., 2025](https://arxiv.org/html/2610.02616#bib.bib37)\)\. The final harness is selected onSvalS\_\{\\mathrm\{val\}\}, andStestS\_\{\\mathrm\{test\}\}is used only to evaluate it\.
Verification: choosing the right repair\.A tool call may fail because of a timeout or incorrect arguments, requiring different repairs\. We model this as a binary hypothesis test\([Nielsen, 2014](https://arxiv.org/html/2610.02616#bib.bib30)\)with two possible causesW∈\{1,2\}W\\in\\\{1,2\\\}, prior probabilitiesλ\\lambdaand1−λ1\-\\lambda, and repairsh1,h2h\_\{1\},h\_\{2\};hWh\_\{W\}is the repair for the actual cause\. The optimizer does not knowWW: it examines the available evidence, runs checks, and selects a repairh^\\widehat\{h\}\. The overlap of a piece of evidence is the Bhattacharyya coefficient\([Bhattacharyya, 1943](https://arxiv.org/html/2610.02616#bib.bib5)\)between its distributions under the two causes:11if it is equally likely under both,00if it reveals the cause\. Letβ0\\beta\_\{0\}be the overlap of all evidence available before the checks, such as the recorded traces, earlier feedback, and notes\. Letρj\\rho\_\{j\}bound the overlap of checkjj, which may depend on earlier results \(Appendix[C\.1](https://arxiv.org/html/2610.02616#A3.SS1)\)\.
###### Theorem 1\(Repair choice\)\.
Letpn∗p\_\{n\}^\{\*\}be the smallest probability of choosing the wrong repair that any decision rule achieves from the evidence afternnchecks, and letε≥0\\varepsilon\\geq 0be the optimizer’s excess error over this best rule\. Then
ℙ\(h^≠hW\)=pn∗\+ε≤12β0∏j=1nρj\+ε\.\\mathbb\{P\}\(\\widehat\{h\}\\neq h\_\{W\}\)=p\_\{n\}^\{\*\}\+\\varepsilon\\leq\\tfrac\{1\}\{2\}\\,\\beta\_\{0\}\\prod\_\{j=1\}^\{n\}\\rho\_\{j\}\+\\varepsilon\.\(3\)If the checks add no information about the cause beyond the existing evidence, as when they only re\-analyze it, thenpn∗≥λ\(1−λ\)β02p\_\{n\}^\{\*\}\\geq\\lambda\(1\-\\lambda\)\\,\\beta\_\{0\}^\{2\}for everynn\.
Self\-evolution: finding and selecting a good harness\.Theorem[1](https://arxiv.org/html/2610.02616#Thmtheorem1)concerns how to repair a given failure\. The optimizer must also decide which failure to repair, and the training and validation losses must then select a good candidate\. Fix a target lossℓ0∈\[0,1\]\\ell\_\{0\}\\in\[0,1\]and call a harness*good*if its loss on new tasks is at mostℓ0\\ell\_\{0\}\. A run makesK=TMK=TMproposals,MMcandidates in each ofTTrounds\. Let𝒞\\mathcal\{C\}containh0h\_\{0\}and allKKcandidates, and letεsel=𝔼\[Lθ,D\(h^\)−minh∈𝒞Lθ,D\(h\)\]\\varepsilon\_\{\\mathrm\{sel\}\}=\\mathbb\{E\}\[L\_\{\\theta,D\}\(\\widehat\{h\}\)\-\\min\_\{h\\in\\mathcal\{C\}\}L\_\{\\theta,D\}\(h\)\]be the selection error, the expected loss gap between the selected harnessh^\\widehat\{h\}and the best harness in𝒞\\mathcal\{C\}\.
###### Theorem 2\(Finding and selecting a good harness\)\.
Suppose that, as long as no good harness has been found, proposaliitargets a failure whose correct repair yields a good harness with probability at leastrir\_\{i\}\. Given such a target, suppose it submits a wrong repair with probability at mosteie\_\{i\}\. Then
𝔼Lθ,D\(h^\)≤ℓ0\+\(1−ℓ0\)∏i=1K\[1−ri\(1−ei\)\]⏟no good harness found\+εsel⏟selection error\.\\mathbb\{E\}L\_\{\\theta,D\}\(\\widehat\{h\}\)\\leq\\ell\_\{0\}\+\\underbrace\{\(1\-\\ell\_\{0\}\)\\prod\_\{i=1\}^\{K\}\\bigl\[1\-r\_\{i\}\(1\-e\_\{i\}\)\\bigr\]\}\_\{\\text\{no good harness found\}\}\+\\underbrace\{\\varepsilon\_\{\\mathrm\{sel\}\}\}\_\{\\text\{selection error\}\}\.\(4\)
## 4Experiments
### 4\.1Setup
We use the SWE\-rebench splits of Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2): all training, validation, and in\-distribution test tasks are in Python, while 87 of the 107 OOD test tasks are in Go, Java, Rust, or TypeScript\. We evaluate every harness on both test sets\. As in Observation 1, the optimizer is Qwen3\.8\-Flash\-Next, the executor is Qwen3\.8\-27B, and each run has six rounds\. Every method generates three candidate executor harnesses per round with the same proposal\-turn limit; the verification tools account for 2 to 7% of all executor runs during evolution, and total compute differs across methods \(Appendix[D\.3](https://arxiv.org/html/2610.02616#A4.SS3)\)\. We also run experiments on Terminal\-Bench\([Merrill et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib27)\); Appendix[D](https://arxiv.org/html/2610.02616#A4)gives its splits and setup, and Appendix[F](https://arxiv.org/html/2610.02616#A6)its results\.
We reimplement Meta\-Harness\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21)\), AHE\([Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23)\), Self\-Harness\([Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55)\), and HarnessX\([Chen et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib9)\)under this protocol, keeping each method’s evidence interface \(Appendix[E](https://arxiv.org/html/2610.02616#A5)\), and add VERSE to each as its host\. We report avg@3 accuracy: the percentage of test tasks resolved, averaged over three evaluations of a harness, with the standard error \(SEM\)\. Blank, the initial harnessh0h\_\{0\}, is evaluated the same way\. Every run starts fromh0h\_\{0\}and a new optimizer workspace\.
### 4\.2Main results
Table 2:Test accuracy of VERSE on SWE\-rebench\(avg@3 accuracy, %,±\\pmSEM over three test evaluations; one evolution run per row\)\. In\-distribution: the sealed test set \(108 tasks\); out\-of\-distribution \(OOD\): 107 newer tasks in five languages, tested with the same harnesses\.r⋆r^\{\\star\}: the round selected on validation;*Val\-selected*: the harness of that round;*Last*: the final\-round harness\. Atr⋆=0r^\{\\star\}=0the entry re\-evaluatesh0h\_\{0\}\.Bold: higher mean than the baseline of its block\.VERSE improves every baseline on both test sets\.With self\-evolution, VERSE scores above its unequipped baseline in all 16 reported comparisons of Table[2](https://arxiv.org/html/2610.02616#S4.T2), on both test sets and for both the selected and the final harness, by0\.60\.6to10\.310\.3points\. Its best harness, on the Meta\-Harness host, is3\.13\.1points above the strongest baseline in distribution\. The four hosts analyze failed trajectories in different ways \(Appendix[E](https://arxiv.org/html/2610.02616#A5)\), and VERSE improves each of them, so it is not tied to one particular form of failure analysis\.
The gains of VERSE generalize out of distribution\.On the OOD test set, the validation\-selected baselines stay within two points of blank, including Meta\-Harness, whose gain over blank falls from 5\.2 points in distribution to 1\.2\. With self\-evolution, VERSE is 2\.8 to 10\.3 points above blank on every host\. What VERSE learns on Python tasks thus carries over to Go, Java, Rust, and TypeScript: across the five languages, it raises its host in 16 of the 20 pairs of host and language \(Appendix[G\.1](https://arxiv.org/html/2610.02616#A7.SS1)\)\.
Self\-evolution keeps VERSE above the baseline in every comparison\.The VERSE tools already help without self\-evolution: in distribution, they select a better harness than the baseline on three of the four hosts, by4\.04\.0to10\.510\.5points\. Without self\-evolution, however, they fall below the baseline in 8 of the 16 comparisons of Table[2](https://arxiv.org/html/2610.02616#S4.T2), while VERSE with self\-evolution stays above it in all 16\. Self\-evolution does not raise every score: compared with VERSE without self\-evolution, it raises the selected accuracy on Meta\-Harness on both test sets and on HarnessX out of distribution, and lowers it in the other five cases\. Its selected and final harnesses stay within1\.61\.6points in distribution, but differ by5\.05\.0points on HarnessX out of distribution\. Test accuracy is not available when one chooses a host or a round, so staying above the baseline in every comparison matters in practice\.
### 4\.3Ablations
Table 3:Ablating VERSE on the AHE host\(SWE\-rebench in\-distribution test set; avg@3 accuracy, %,±\\pmSEM over three test evaluations; one evolution run per row\)\. Every configuration runs without self\-evolution\. Each indented row removes components from the full configuration; the last row is AHE alone\.r⋆r^\{\\star\},*Val\-selected*, and*Last*are as in Table[2](https://arxiv.org/html/2610.02616#S4.T2)\.Figure 5:What verification does during evolution, over the 18 runs of Tables[1](https://arxiv.org/html/2610.02616#S2.T1)–[5](https://arxiv.org/html/2610.02616#S4.F5)\. Top: draft evaluations in the ten runs with theverificationtool \(light blue: only regression checks passed; gray: inconclusive\)\. Middle: failing training tasks that submitted edits were predicted to fix, checked in the next training run\. Bottom: validation outcome changes, rerun once\.The diagnosis reports make the checks informative\.In Table[5](https://arxiv.org/html/2610.02616#S4.F5), removing attribution and minimization, which also removes the training audit, costs about half of the gain that VERSE adds to AHE \(2\.22\.2of4\.04\.0points\); removing the training audit alone costs1\.21\.2\. Without the reports, the optimizer calls the verification tools more than three times as often, yet selects a weaker harness\. Execution checks likely help most when they test well\-chosen hypotheses, which the reports supply\.
Without self\-evolution, the optimizer uses the verification tools less\.Removing them barely changes the selected\-round accuracy: on this host, the fixed optimizer’s first nine calls all refuted its drafts, and it stopped calling after the second round\. Across the four hosts of Table[2](https://arxiv.org/html/2610.02616#S4.T2), self\-evolving optimizers call the tools more than twice as often, and more than three times as often in the last three rounds\. On Meta\-Harness, VERSE without self\-evolution scores below both the baseline and simple verification alone \(Tables[1](https://arxiv.org/html/2610.02616#S2.T1)and[2](https://arxiv.org/html/2610.02616#S4.T2)\)\. Its optimizer kept calling theverificationtool, but only on tasks that the current harness fails, so none of these calls could show whether an edit breaks a task that already passes\. In 16 of the 23 drafts that the self\-evolving optimizer ran on this host, the targets included such a task\. In Appendix[H](https://arxiv.org/html/2610.02616#A8), random or fixed experiment schedules and a restricted workspace also cost accuracy\.
### 4\.4What verification does during evolution
Few draft edits fix their targets, and execution shows it before submission\.Figure[5](https://arxiv.org/html/2610.02616#S4.F5)aggregates the logs of the 18 runs behind Tables[1](https://arxiv.org/html/2610.02616#S2.T1)–[5](https://arxiv.org/html/2610.02616#S4.F5)\. With each submitted edit, the optimizer names the failing training tasks it expects the edit to fix; only one in ten of these tasks passed in the next training run\. Theverificationtool reveals such failures before submission: only 17 of its 170 draft evaluations, counting repeated evaluations of the same draft, found that the draft fixed a failing task it targeted \(Appendix[D\.4](https://arxiv.org/html/2610.02616#A4.SS4)\)\. In another 30, the draft passed only on tasks that the current harness already solves, which the optimizer had added to check that the draft does not break them\. Such regression checks matter because an edit can fix its targets and still break other tasks: across the 18 runs, submitted edits fixed 148 validation tasks and broke 134, counting only changes that reproduced on a rerun\. Half of all validation outcome changes did not reproduce\. Because single runs are this noisy, the tool counts a task as fixed only after two passes\.
Self\-evolution changes how the optimizer verifies\.Self\-evolving optimizers learned to add regression checks before submission \(Appendix[L](https://arxiv.org/html/2610.02616#A12)\): 34 of their 81 draft evaluations included a training task that the current harness already solves, against 2 of 89 for fixed optimizers\. Self\-evolving optimizers also respond when a check fails: when the tool showed that a draft fixed none of its targets, they revised the draft before submitting in 30 of 38 cases \(79%79\\%\), against 42 of 74 \(57%57\\%\) without self\-evolution\. Evaluated drafts fixed a failing target at similar rates in both groups \(9 of 81 against 8 of 89\); the clearer difference lies in what the optimizer tests and how it responds when a test fails\.
## 5Conclusion
VERSE combines execution\-based verification with self\-evolution of the harness optimizer, while keeping both models’ weights fixed\. The optimizer can test draft edits, replay failures, and perturb suspected steps before submission, then use the resulting evidence to revise both the executor harness and its own prompts, skills, tools, hooks, and notes\. Our analysis identifies conditions under which informative checks can reduce repair\-choice error and relates downstream performance to the quality of target selection, repair, and harness selection\. In our experiments, adding VERSE to four reimplemented harness optimizers improves all four on held\-out Python tasks and on newer out\-of\-distribution tasks spanning five programming languages\. The evolution logs reveal how the optimization process changes: self\-evolving optimizers more often include regression checks and revise unsuccessful drafts before submission\. Although the incremental benefit of self\-evolution varies across hosts, these findings support a broader principle: harness evolution should improve not only the agent’s configuration, but also the procedures that propose, test, and revise it\.
### AI use statement
In this work, we used generative AI tools to draft and edit the manuscript, search and check related work and citations, implement experiment infrastructure, help analyze experiment logs, and help write and check the theoretical results and proofs\. Benchmark tasks come from existing datasets; the trajectories, diagnoses, edits, and optimizer artifacts that we analyze are outputs of the models under study\. The research ideas, the ideas and structure of the proofs, the experimental design, and the interpretation of the analyses are the authors’ own\. All AI\-assisted text, code, analyses, and proofs were reviewed, verified, and tested by the authors\. We take responsibility for the final content of this work, including text, claims, and artifacts produced with the aid of generative AI\.
### Ethics statement
This work studies methods for improving LLM agent harnesses on software engineering benchmarks in sandboxed environments\. It does not involve human subjects or sensitive data\. Self\-modifying optimizers raise oversight questions\. In our system, the internals of the verification tools are fixed, all optimizer actions are logged and auditable, self\-written code is loaded only after a safety check and a test run, the executor acts only inside task containers without network access, and the executor model is frozen, which bounds the scope of self\-modification\.
### Reproducibility statement
Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2)and Appendix[D](https://arxiv.org/html/2610.02616#A4)specify the task pools, splits, and budgets; Section[4\.1](https://arxiv.org/html/2610.02616#S4.SS1)specifies the models and the evaluation protocol; Section[3\.1](https://arxiv.org/html/2610.02616#S3.SS1)specifies the verification tools and optimizer self\-evolution; Appendix[B](https://arxiv.org/html/2610.02616#A2)gives pseudocode, including the selection rules; Appendix[E](https://arxiv.org/html/2610.02616#A5)describes the baseline reimplementations; and Appendix[C](https://arxiv.org/html/2610.02616#A3)states the theoretical assumptions and proofs\. Our repository at[https://github\.com/wzekai/VERSE](https://github.com/wzekai/VERSE)contains the code \(the evolution engine, the verification tools, the baseline reimplementations, and the evaluation scripts\), the configuration of every reported evolution run, and the task IDs of every split; for the main results table, it also contains the selected executor and optimizer harnesses and their per\-task test results\.
## References
- Agrawal et al\. \(2026\)Lakshya A\. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems, Rishi Khare, Krista Opsahl\-Ong, Arnav Singhvi, Herumb Shandilya, Michael J\. Ryan, Meng Jiang, Christopher Potts, Koushik Sen, Alex Dimakis, Ion Stoica, Dan Klein, Matei Zaharia, and Omar Khattab\.GEPA: Reflective prompt evolution can outperform reinforcement learning\.In*International Conference on Learning Representations \(ICLR\)*, 2026\.
- Alzubi et al\. \(2026\)Salaheddin Alzubi, Noah Provenzano, Jaydon Bingham, Christian Alexander Calvo, Weiyuan Chen, and Tu Vu\.EvoSkill: Automated skill discovery for multi\-agent systems\.In*Conference on Language Modeling \(COLM\)*, 2026\.
- Badertdinov et al\. \(2025\)Ibragim Badertdinov, Alexander Golubev, Maksim Nekrashevich, Anton Shevtsov, Simon Karasik, Andrei Andriushchenko, Maria Trofimova, Daria Litvintseva, and Boris Yangel\.SWE\-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents\.In*Advances in Neural Information Processing Systems \(NeurIPS\), Datasets and Benchmarks Track*, 2025\.
- Bertran et al\. \(2026\)Martin Andres Bertran, Aaron Roth, and Zhiwei Steven Wu\.What fits \(into few tokens\) doesn’t overfit: Compression and generalization in ML research agents\.*arXiv preprint arXiv:2606\.11045*, 2026\.
- Bhattacharyya \(1943\)A\. Bhattacharyya\.On a measure of divergence between two statistical populations defined by their probability distributions\.*Bulletin of the Calcutta Mathematical Society*, 35:99–109, 1943\.
- Blum & Hardt \(2015\)Avrim Blum and Moritz Hardt\.The ladder: A reliable leaderboard for machine learning competitions\.In*International Conference on Machine Learning \(ICML\)*, 2015\.
- Chen et al\. \(2026a\)Guhong Chen, Yingcheng Shi, Yongbin Li, Binhua Li, Xander Xu, Hu Wei, Shiwen Ni, Min Yang, and Jieping Ye\.EvoTrainer: Co\-evolving LLM policies and training harnesses for autonomous agentic reinforcement learning\.*arXiv preprint arXiv:2606\.03108*, 2026a\.
- Chen et al\. \(2026b\)Mengzhuo Chen, Junjie Wang, Zhe Liu, Yawen Wang, Haiming Zheng, and Qing Wang\.From failed trajectories to reliable LLM agents: Diagnosing and repairing harness flaws\.*arXiv preprint arXiv:2606\.06324*, 2026b\.
- Chen et al\. \(2026c\)Tingyang Chen, Shuo Lu, Kang Zhao, Weicheng Meng, Hanlin Teng, Tianhao Li, Chao Li, Xule Liu, Jian Liang, Zhizhong Zhang, Yuan Xie, Heng Qu, Kun Shao, and Jian Luan\.HarnessX: A composable, adaptive, and evolvable agent harness foundry\.*arXiv preprint arXiv:2606\.14249*, 2026c\.
- Dwork et al\. \(2015\)Cynthia Dwork, Vitaly Feldman, Moritz Hardt, Toni Pitassi, Omer Reingold, and Aaron Roth\.Generalization in adaptive data analysis and holdout reuse\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2015\.
- Fernando et al\. \(2024\)Chrisantha Fernando, Dylan Sunil Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel\.Promptbreeder: Self\-referential self\-improvement via prompt evolution\.In*International Conference on Machine Learning \(ICML\)*, 2024\.
- Hastie et al\. \(2009\)Trevor Hastie, Robert Tibshirani, and Jerome Friedman\.*The Elements of Statistical Learning: Data Mining, Inference, and Prediction*\.Springer, 2nd edition, 2009\.
- Hebbar et al\. \(2026\)Prannay Hebbar, Yogendra Manawat, Samuel Verboomen, Alesia Ivanova, Selvam Palanimalai, Kunal Bhatia, and Vignesh Baskaran\.SIA: Self improving AI with harness & weight updates\.*arXiv preprint arXiv:2605\.27276*, 2026\.
- Hoeffding \(1963\)Wassily Hoeffding\.Probability inequalities for sums of bounded random variables\.*Journal of the American Statistical Association*, 58\(301\):13–30, 1963\.
- Huang et al\. \(2026\)Chengsong Huang, Zifeng Wang, Rujun Han, Jun Yan, Yanfei Chen, Zoey CuiZhu, Ke Jiang, Peng Xia, Han Yu, Yufan Zhuang, Yifei Ming, Jiaqi Pan, Bhavana Dalvi Mishra, Jiaxin Huang, Burak Gokturk, Tomas Pfister, and Chen\-Yu Lee\.EnvHarness: Awakening static worlds for agent learning\.*arXiv preprint arXiv:2608\.19880*, 2026\.
- Jiang et al\. \(2026\)Wen Jiang, Mingmin Chu, Yimeng Tian, Qianxin Zhang, Haofei Yang, Rui Yang, Yang Liu, Tao Lv, and Fangming Li\.HarnessEvolve: Learning from reference trajectories for reliable agent self\-evolution\.*arXiv preprint arXiv:2609\.00829*, 2026\.
- Jimenez et al\. \(2024\)Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R Narasimhan\.SWE\-bench: Can language models resolve real\-world GitHub issues?In*International Conference on Learning Representations \(ICLR\)*, 2024\.
- Kailath \(1967\)Thomas Kailath\.The divergence and Bhattacharyya distance measures in signal selection\.*IEEE Transactions on Communication Technology*, 15\(1\):52–60, 1967\.
- Khattab et al\. \(2024\)Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T\. Joshi, Hanna Moazam, Heather Miller, Matei Zaharia, and Christopher Potts\.DSPy: Compiling declarative language model calls into state\-of\-the\-art pipelines\.In*International Conference on Learning Representations \(ICLR\)*, 2024\.
- Lee et al\. \(2026a\)Hyunin Lee, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Matei Zaharia, and Yujin Tang\.Recursive Harness Self\-Improvement\.*arXiv preprint arXiv:2607\.15524*, 2026a\.
- Lee et al\. \(2026b\)Yoonho Lee, Roshen Sanjay Nair, Qizheng Zhang, Kangwook Lee, Omar Khattab, and Chelsea Finn\.Meta\-Harness: End\-to\-end optimization of model harnesses\.In*Conference on Language Modeling \(COLM\)*, 2026b\.
- Li et al\. \(2026\)Han Li, Yifan Yao, Letian Zhu, Rili Feng, Hongyi Ye, Jiaming Wang, Yancheng He, Pengyu Zou, Lehan Zhang, Xinping Lei, Haoyang Huang, Ken Deng, Ming Sun, Zhaoxiang Zhang, He Ye, and Jiaheng Liu\.CodeTracer: Towards traceable agent states\.*arXiv preprint arXiv:2604\.11641*, 2026\.
- Lin et al\. \(2026\)Jiahang Lin, Shichun Liu, Chengjun Pan, Lizhi Lin, Shihan Dou, Zhiheng Xi, Xuanjing Huang, Hang Yan, Zhenhua Han, Tao Gui, and Yu\-Gang Jiang\.Agentic Harness Engineering: Observability\-driven automatic evolution of coding\-agent harnesses\.*arXiv preprint arXiv:2604\.25850*, 2026\.
- Lu et al\. \(2026\)Mengyin Lu, Cong Feng, Huimin Han, Guangming Lu, Yu Sun, Xiaonan Ding, Shihui Long, Fengyi Li, and Tanvi Motwani\.SPEAR: Code\-augmented agentic prompt optimization\.*arXiv preprint arXiv:2605\.26275*, 2026\.
- Luo et al\. \(2026a\)Haochen Luo, Yi Huang, Sichun Luo, Fengyuan Liu, Lei Li, Zefa Hu, Junlan Feng, and Qi Liu\.Harness\-Aware Self\-Evolving: Co\-evolving model weights, harness, and task solutions\.*arXiv preprint arXiv:2607\.03935*, 2026a\.
- Luo et al\. \(2026b\)Xiaotian Luo, Dizhan Xue, Fengxingyu Wang, Chuanrui Hu, and Yafeng Deng\.HarnessBank: Semantic gene\-bank search with gated verification for agent\-harness self\-evolution\.*arXiv preprint arXiv:2607\.13683*, 2026b\.
- Merrill et al\. \(2026\)Mike A Merrill, Alexander Glenn Shaw, Nicholas Carlini, Boxuan Li, Harsh Raj, Ivan Bercovich, Lin Shi, Jeong Yeon Shin, Thomas Walshe, E\. Kelly Buchanan, Junhong Shen, Guanghao Ye, Haowei Lin, Jason Poulos, Maoyu Wang, Marianna Nezhurina, Di Lu, Orfeas Menis Mastromichalakis, Zhiwei Xu, Zizhao Chen, Yue Liu, Robert Zhang, Leon Liangyu Chen, Anurag Kashyap, Jan\-Lucas Uslu, Jeffrey Li, Jianbo Wu, Minghao Yan, Song Bian, Vedang Sharma, Ke Sun, Steven Dillmann, Akshay Anand, Andrew Lanpouthakoun, Bardia Koopah, Changran Hu, Etash Kumar Guha, Gabriel H\. S\. Dreiman, Jiacheng Zhu, Karl Krauth, Li Zhong, Niklas Muennighoff, Robert Kwesi Amanfu, Shangyin Tan, Shreyas Pimpalgaonkar, Tushar Aggarwal, Xiangning Lin, Xin Lan, Xuandong Zhao, Yiqing Liang, Yuanli Wang, Zilong Wang, Changzhi Zhou, David Heineman, Hange Liu, Harsh Trivedi, John Yang, Junhong Lin, Manish Shetty, Michael Yang, Nabil Omi, Negin Raoof, Shanda Li, Terry Yue Zhuo, Wuwei Lin, Yiwei Dai, Yuxin Wang, Wenhao Chai, Shang Zhou, Dariush Wahdany, Ziyu She, Jiaming Hu, Zhikang Dong, Yuxuan Zhu, Sasha Cui, Ahson Saiyed, Arinbjörn Kolbeinsson, Christopher Michael Rytting, Ryan Marten, Yixin Wang, Jenia Jitsev, Alex Dimakis, Andy Konwinski, and Ludwig Schmidt\.Terminal\-Bench: Benchmarking agents on hard, realistic tasks in command line interfaces\.In*International Conference on Learning Representations \(ICLR\)*, 2026\.
- Ni et al\. \(2026\)Jingwei Ni, Yihao Liu, Xinpeng Liu, Yutao Sun, Mengyu Zhou, Pengyu Cheng, Dexin Wang, Erchao Zhao, Xiaoxi Jiang, and Guanjun Jiang\.Trace2Skill: Distill trajectory\-local lessons into transferable agent skills\.*arXiv preprint arXiv:2603\.25158*, 2026\.
- Nie et al\. \(2026\)Jun Nie, Yonggang Zhang, Jun Song, Qianshu Cai, Dahai Yu, Yike Guo, Xinmei Tian, and Bo Han\.TTHE: Test\-time harness evolution\.*arXiv preprint arXiv:2607\.08124*, 2026\.
- Nielsen \(2014\)Frank Nielsen\.Generalized Bhattacharyya and Chernoff upper bounds on Bayes error using quasi\-arithmetic means\.*Pattern Recognition Letters*, 42:25–34, 2014\.
- Park et al\. \(2026\)Sungho Park, Wonjoong Kim, Rongyuan Tan, Jue Zhang, Wook\-Shin Han, Pengfei Gao, Chanyoung Park, Yongqiang Yao, Rao Fu, Elsie Nallipogu, Qingwei Lin, Saravan Rajmohan, and Dongmei Zhang\.AutoSaddler: Automatic harness optimization with durable updates from agent execution traces\.*arXiv preprint arXiv:2608\.23041*, 2026\.
- Prechelt \(1998\)Lutz Prechelt\.Early stopping – but when?In Genevieve B\. Orr and Klaus\-Robert Müller \(eds\.\),*Neural Networks: Tricks of the Trade*, pp\. 55–69\. Springer, 1998\.
- Russo & Zou \(2016\)Daniel Russo and James Zou\.Controlling bias in adaptive data analysis using information theory\.In*International Conference on Artificial Intelligence and Statistics \(AISTATS\)*, 2016\.
- Seong et al\. \(2026\)Haebin Seong, Li Yin, Haoran Zhang, and Zhan Shi\.The last harness you’ll ever build\.*arXiv preprint arXiv:2604\.21003*, 2026\.
- Shinn et al\. \(2023\)Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\.Reflexion: Language agents with verbal reinforcement learning\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2023\.
- Snoek et al\. \(2012\)Jasper Snoek, Hugo Larochelle, and Ryan P\. Adams\.Practical Bayesian optimization of machine learning algorithms\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2012\.
- Sucker et al\. \(2025\)Michael Sucker, Jalal Fadili, and Peter Ochs\.Learning\-to\-optimize with PAC\-Bayesian guarantees: Theoretical considerations and practical implementation\.*Journal of Machine Learning Research*, 26\(211\):1–53, 2025\.
- Tan et al\. \(2026\)Zihan Tan, Leixin Sun, Zitong Shi, Yitao Liu, Jiajun Wu, Nathaniel Brooks, Jiaru Qian, Xiaoran Shang, Suyuan Huang, Yi Ding, Yangxu Liao, Mukai Li, Qiushi Sun, Shudong Liu, Xuankun Rong, Xiaohang Yu, Zhuo Chen, Hejia Geng, Chenxin Li, Aozhou Wang, Zengji Tu, Robert Tang, Yuxin Zhan, Eric Jiang, Yuxin Wu, Jianqing Zhang, Xiao Liang, Fang Wu, Haochi Zhang, Alexander Marlow, and Guancheng Wan\.MetaRSI / RSI2: A meta\-recursive self\-improving system for recursive self\-improving systems themselves\.*arXiv preprint arXiv:2609\.06396*, 2026\.
- Tao et al\. \(2026\)Wangcheng Tao, Han Wu, and Weng\-Fai Wong\.SePO: Self\-evolving prompt agent for system prompt optimization\.*arXiv preprint arXiv:2606\.04465*, 2026\.
- Ursekar et al\. \(2026\)Varun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue, and Samuel Marc Denton\.VeRO: A harness for agents to optimize agents\.In*International Conference on Machine Learning \(ICML\)*, 2026\.
- Wang et al\. \(2026a\)Su Wang, Pin Qian, Yifan Lin, Jingzhou Xu, Yihang Chen, Xiaochong Jiang, Lifei Liu, and Haoran Yu\.Phantom Guardrails: When self\-improving agent harnesses fix failures that never happened\.In*KDD Workshop on Evaluation and Trustworthiness of Agentic AI*, 2026a\.
- Wang et al\. \(2026b\)Yike Wang, Huaisheng Zhu, Zhengyu Hu, Yige Yuan, Zhengyu Chen, Shakti Senthil, Hannaneh Hajishirzi, Yulia Tsvetkov, Pradeep Dasigi, and Teng Xiao\.Rethinking the evaluation of harness evolution for agents\.In*COLM 2026 The 2nd Workshop on Lifelong Agents: Learning, Aligning, and Evolving*, 2026b\.
- Wang et al\. \(2026c\)Zefeng Wang, Minxi Yan, Jinhe Bi, Sikuan Yan, Volker Tresp, and Yunpu Ma\.MetaSkill\-Evolve: Recursive self\-improvement of LLM agents via two\-timescale meta\-skill evolution\.*arXiv preprint arXiv:2607\.05297*, 2026c\.
- Wang et al\. \(2026d\)Zekun Wang, Anant Gupta, Zihan Dong, and Christopher J\. MacLellan\.Self\-consolidating language models: Continual knowledge incorporation from context\.*arXiv preprint arXiv:2605\.07076*, 2026d\.
- Wu et al\. \(2026\)Siwei Wu, Jincheng Ren, Yizhi Li, Haau\-Sing Li, Chengran Yang, Yuxuan Zhang, Weicheng Gu, Jian Yang, Riza Batista\-Navarro, Chuanyi Zhang, Xianglong Liu, Ming Zhou, Bryan Dai, and Chenghua Lin\.ModularRSI: Modular and generalizable recursive harness self\-improvement\.*arXiv preprint arXiv:2609\.14857*, 2026\.
- Xia et al\. \(2026\)Peng Xia, Rujun Han, Zifeng Wang, Yanfei Chen, Yufan Zhuang, Yoonho Lee, Chengsong Huang, Han Yu, Zhongying CuiZhu, Yifei Ming, Huaxiu Yao, Burak Gokturk, Tomas Pfister, and Chen\-Yu Lee\.RRSI: Regularized recursive self\-improvement of agent harnesses\.*arXiv preprint arXiv:2609\.24972*, 2026\.
- Xu & Raginsky \(2017\)Aolin Xu and Maxim Raginsky\.Information\-theoretic analysis of generalization capability of learning algorithms\.In*Advances in Neural Information Processing Systems \(NeurIPS\)*, 2017\.
- Yang et al\. \(2024\)Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V\. Le, Denny Zhou, and Xinyun Chen\.Large language models as optimizers\.In*International Conference on Learning Representations \(ICLR\)*, 2024\.
- Ye et al\. \(2026\)Haoran Ye, Xuning He, Vincent Arak, Haonan Dong, and Guojie Song\.Meta Context Engineering via agentic skill evolution\.In*International Conference on Machine Learning \(ICML\)*, 2026\.
- Yin et al\. \(2025\)Xunjian Yin, Xinyi Wang, Liangming Pan, Li Lin, Xiaojun Wan, and William Yang Wang\.Gödel agent: A self\-referential agent framework for recursively self\-improvement\.In*Annual Meeting of the Association for Computational Linguistics \(ACL\)*, 2025\.
- Yu et al\. \(2026\)Simon Yu, Derek Chong, Ananjan Nandi, Dilara Soylu, Jiuding Sun, Christopher D Manning, and Weiyan Shi\.Shepherd: Enabling programmable meta\-agents via reversible agentic execution traces\.*arXiv preprint arXiv:2605\.10913*, 2026\.
- Yuksekgonul et al\. \(2025\)Mert Yuksekgonul, Federico Bianchi, Joseph Boen, Sheng Liu, Pan Lu, Zhi Huang, Carlos Guestrin, and James Zou\.Optimizing generative AI by backpropagating language model feedback\.*Nature*, 639\(8055\):609–616, 2025\.
- Zakerinia et al\. \(2024\)Hossein Zakerinia, Amin Behjati, and Christoph H\. Lampert\.More flexible PAC\-Bayesian meta\-learning by learning learning algorithms\.In*International Conference on Machine Learning \(ICML\)*, 2024\.
- Zelikman et al\. \(2024\)Eric Zelikman, Eliana Lorch, Lester Mackey, and Adam Tauman Kalai\.Self\-taught optimizer \(STOP\): Recursively self\-improving code generation\.In*Conference on Language Modeling \(COLM\)*, 2024\.
- Zhang et al\. \(2026a\)Hangfan Zhang, Shao Zhang, Kangcong Li, Chen Zhang, Yang Chen, Yiqun Zhang, Lei Bai, and Shuyue Hu\.Self\-Harness: Harnesses that improve themselves\.*arXiv preprint arXiv:2606\.09498*, 2026a\.
- Zhang et al\. \(2026b\)Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, and Jeff Clune\.Darwin Gödel Machine: Open\-ended evolution of self\-improving agents\.In*International Conference on Learning Representations \(ICLR\)*, 2026b\.
- Zhang et al\. \(2026c\)Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang, Sam Devlin, and Tatiana Shavrina\.Hyperagents\.*arXiv preprint arXiv:2603\.19461*, 2026c\.
- Zhang et al\. \(2026d\)Jiayi Zhang, Yongfeng Gu, Jianhao Ruan, Maojia Song, Yiran Peng, Zhiguang Han, Jinyu Xiang, Zhitao Wang, Caiyin Yang, Yixi Ouyang, Bang Liu, Chenglin Wu, and Yuyu Luo\.Harnessing agentic evolution\.*arXiv preprint arXiv:2605\.13821*, 2026d\.
- Zhou \(2026\)Tailin Zhou\.Hierarchical self\-improvement: A framework for task\-specific evolvable agent harnesses\.*arXiv preprint arXiv:2608\.08466*, 2026\.
- Zoph & Le \(2017\)Barret Zoph and Quoc V\. Le\.Neural architecture search with reinforcement learning\.In*International Conference on Learning Representations \(ICLR\)*, 2017\.
## Appendix contents
## Appendix ARelated work
Harness evolution\.Recent methods improve an agent’s harness with an LLM proposer that reads execution traces and scores; our four baselines differ in how the proposer reads this evidence\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21);[Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23);[Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55);[Chen et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib9)\)\(Appendix[E](https://arxiv.org/html/2610.02616#A5)\)\. Related work adapts harnesses at test time\([Nie et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib29)\), refines prompt\-level harnesses from pairwise comparisons\([Lee et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib20)\), optimizes prompts and programs with LLM feedback\([Shinn et al\., 2023](https://arxiv.org/html/2610.02616#bib.bib35);[Yang et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib48);[Khattab et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib19)\), updates model weights\([Hebbar et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib13)\), or distills skills from trajectories\([Ni et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib28);[Alzubi et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib2)\)\.[Wang et al\. \(2026b\)](https://arxiv.org/html/2610.02616#bib.bib42)note that many results search and report on the same tasks; ModularRSI evolves on tasks disjoint from its benchmarks\([Wu et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib45)\), and RRSI regularizes proposal and selection and reports out\-of\-distribution gains\([Xia et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib46)\)\.
Verifying harness edits\.Execution is also used to screen candidate harnesses with gates\([Luo et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib26)\), repair attributed failures\([Chen et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib8)\), and revert and replay agent runs\([Yu et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib51)\)\. SPEAR gives a prompt optimizer an evaluate tool that it calls itself\([Lu et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib24)\), and VeRO finds that agent optimizers mostly edit prompts\([Ursekar et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib40)\)\. Among contemporaneous systems, AutoSaddler keeps a patch only if it improves a mini\-batch, then evaluates it on a development set\([Park et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib31)\), HarnessEvolve gates updates and locates errors against reference trajectories\([Jiang et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib16)\), and EnvHarness accepts environment components by success rates over fresh rollouts\([Huang et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib15)\)\.
Evolving the optimizer\.In most harness optimizers, the optimizer’s own procedure stays fixed during a run; when the optimizer and the executor share a base model, the optimizer runs under its own prompts or scaffold\([Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23);[Nie et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib29)\)\. In Self\-Harness and HASE, the model proposes edits under the harness it is editing, so one harness serves both roles\([Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55);[Luo et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib25)\)\. In self\-referential agents, one program both solves tasks and improves itself: STOP runs a program improver on its own code\([Zelikman et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib54)\), Gödel Agent modifies its own logic\([Yin et al\., 2025](https://arxiv.org/html/2610.02616#bib.bib50)\), the Darwin Gödel Machine lets coding agents modify their own code\([Zhang et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib56)\), and Hyperagents place a task agent and a meta agent in one editable program, so the meta agent can rewrite itself\([Zhang et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib57)\)\.
Other systems keep the optimizer separate and revise its procedure\. Some add a meta\-level above it: a meta\-evolution loop over the evolution blueprint\([Seong et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib34)\), a meta\-agent that edits the procedure controlling future evolution\([Zhang et al\., 2026d](https://arxiv.org/html/2610.02616#bib.bib58)\)or a base agent’s context\-engineering skills\([Ye et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib49)\), a meta\-evolver that improves a harness evolver’s strategy\([Zhou, 2026](https://arxiv.org/html/2610.02616#bib.bib59)\), and a meta\-level policy that schedules improvement operators\([Tan et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib38)\)\. In others, the optimizer revises its own procedure\. Promptbreeder evolves the mutation prompts that improve its task prompts\([Fernando et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib11)\), SePO optimizes a prompt agent’s own system prompt\([Tao et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib39)\), MetaSkill\-Evolve applies its improvement pipeline to its own meta\-skill\([Wang et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib43)\), and EvoTrainer’s trainer revises its own harness across training runs\([Chen et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib7)\)\.
Position of VERSE\.Like the last group, the VERSE optimizer revises its own procedure, but its harnessggis a full agent harness kept separate from the executor harnesshh: it rewrites its own prompts, skills, notes, hooks, and executable tools\. While editinghh, it calls verification tools that test the causal claim behind an edit before submission; AutoSaddler, HarnessEvolve, and EnvHarness also verify edits, with fixed optimization procedures\. The executor model stays frozen and is evaluated on disjoint task splits\. In our controlled study, self\-evolution lowered accuracy without verification and gave the best result with it \(Section[2\.3](https://arxiv.org/html/2610.02616#S2.SS3)\)\.
## Appendix BPseudocode for the bilevel learning protocol
A harness is a complete executable configuration, including prompts, tools, control flow, dependencies, and persistent state\. The executor weightsθ\\thetaand optimizer weightsω\\omegaremain fixed;hhcontrols task execution andggcontrols candidate generation and checking\. All tasks, trajectories, harnesses, and evidence records have finite encodings\. Their spaces are countable, with all subsets measurable;ℋ\\mathcal\{H\}denotes the space of executor harnesses\. Learningggtreats the optimizer as a learned algorithm\([Zakerinia et al\., 2024](https://arxiv.org/html/2610.02616#bib.bib53)\), using execution feedback for program revision\([Agrawal et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib1)\)\.
The inner loop compares the nonempty candidate set𝒞t\(g\)\\mathcal\{C\}\_\{t\}\(g\)on the same training tasks; if a prescribed training subset is used, the comparison uses its mean loss\. The outer objective conditions on the current state and task sets and averages over candidate generation and evaluation\. Candidates may run in parallel, and their results need not be independent\.
Algorithm 1Verified self\-evolving harness learning1:Fixed executor weights
θ\\thetaand optimizer weights
ω\\omega; training
StrainS\_\{\\mathrm\{train\}\},
2:validation
SvalS\_\{\\mathrm\{val\}\}, and test
StestS\_\{\\mathrm\{test\}\}; initial
h0,g0h\_\{0\},g\_\{0\};
3:round count
T≥1T\\geq 1, candidates per round
M≥1M\\geq 1, budgets
BtB\_\{t\}
4:Frozen executor harness
h^\\widehat\{h\}and its test evaluation
5:Evaluate
h0h\_\{0\}on
SvalS\_\{\\mathrm\{val\}\}and archive the harness and result
6:for
t=0,…,T−1t=0,\\ldots,T\-1do⊳\\trianglerightOuter loop: improve the optimizer
7:Fix
gtg\_\{t\};
𝒞t←∅\\mathcal\{C\}\_\{t\}\\leftarrow\\varnothing; training record
Et←∅E\_\{t\}\\leftarrow\\varnothing
8:for
c=1,…,Mc=1,\\ldots,Mdo⊳\\trianglerightInner loop: search for executor harnesses
9:
h~←ht\\widetilde\{h\}\\leftarrow h\_\{t\}
10:whilethe candidate’s turn limit is not reacheddo
11:Read training traces, earlier validation feedback
F<tF\_\{<t\}, and prior evidence under
gtg\_\{t\}
12:Propose an edit to
h~\\widetilde\{h\}or a diagnostic check on
StrainS\_\{\\mathrm\{train\}\}
13:ifa check is requested and the round’s shared budget remainsthen
14:Execute the check; append raw outcomes and costs to
EtE\_\{t\}
15:endif
16:Revise
h~\\widetilde\{h\}using the available evidence
17:endwhile
18:
𝒞t←𝒞t∪\{h~\}\\mathcal\{C\}\_\{t\}\\leftarrow\\mathcal\{C\}\_\{t\}\\cup\\\{\\widetilde\{h\}\\\}
19:endfor
20:Evaluate
𝒞t\\mathcal\{C\}\_\{t\}on the training comparison tasks
21:Choose
ht\+1∈argminh∈𝒞tL^train\(h\)h\_\{t\+1\}\\in\\argmin\_\{h\\in\\mathcal\{C\}\_\{t\}\}\\widehat\{L\}\_\{\\mathrm\{train\}\}\(h\), or keep
hth\_\{t\}if
𝒞t=∅\\mathcal\{C\}\_\{t\}=\\varnothing
22:Smoke\-check
ht\+1h\_\{t\+1\}and let the optimizer repair it if needed; record results in
EtE\_\{t\}
23:Evaluate
ht\+1h\_\{t\+1\}on
SvalS\_\{\\mathrm\{val\}\}to obtain feedback
FtF\_\{t\}; archive both
24:Use the optimizer to revise
gtg\_\{t\}from
Et,FtE\_\{t\},F\_\{t\}, obtaining
gt\+1g\_\{t\+1\}
25:endfor
26:Re\-evaluate the two archived harnesses with the lowest validation losses; keep the better as
h^\\widehat\{h\}
27:Freeze
h^\\widehat\{h\}and evaluate it on
StestS\_\{\\mathrm\{test\}\}
28:return
h^\\widehat\{h\}and its test evaluation
Evidence and feedback\.EtE\_\{t\}contains the training checks and comparison results of roundtt\.FtF\_\{t\}is the validation feedback returned to the optimizer, the same for every method: the validation loss of each round, the identifiers of up to 30 validation tasks whose outcome changed and whether each change held on a rerun, and the rate and two most frequent messages of harness\-code errors on validation runs\. It never contains task statements, tests, or trajectories\. Test tasks enter neither loop\.
Checks\.All candidates and the optimizer self\-update draw checks from one execution budget per round, and every check runs on training tasks\. A check that cannot run, because the round’s budget or time is spent or the draft cannot be applied, returns an inconclusive result; a draft that cannot be applied costs no budget\. No check result blocks a submission: the optimizer decides whether to revise, check again, or submit\.
Choosing the harness of a round\.The candidates are compared by the net number of training tasks they fix on a shared subset of up to 12 training tasks\. A tie goes to the candidate whose submitted edit has more tasks that theverificationtool confirmed as passing, and then to a fixed order\. The chosen candidate must pass a smoke check that includes one run on a training task\. If it fails, the optimizer repairs it in at most three attempts, and the repaired version is validated and archived asht\+1h\_\{t\+1\}without being screened again\. If no attempt passes, or if no candidate was submitted, the round keepshth\_\{t\}and skips validation; the self\-update still runs\.
Validation and final selection\.A validation task whose outcome differs from that of the parent harness is rerun once, and a change that does not reproduce is reverted to the parent’s outcome\. At the end, the two archived harnesses with the lowest validation losses are each evaluated once more, and the one with the lower mean loss is selected; ties favor the later round\. Appendix[C\.3](https://arxiv.org/html/2610.02616#A3.SS3)analyzes this rule and the reliability of validation losses under feedback\.
Self\-updates\.A self\-update of the optimizer harness takes effect only if its code passes a safety check and a test run; otherwise the previous version stays in place\.
## Appendix CAdditional theoretical analysis
This appendix gives the formal versions of Theorems[1](https://arxiv.org/html/2610.02616#Thmtheorem1)and[2](https://arxiv.org/html/2610.02616#Thmtheorem2)and the analysis of selection\. Tasks, trajectories, harnesses, and evidence have finite encodings, so all spaces are countable and all subsets measurable\. Logarithms are natural\.
Tools of the optimizer\.All methods in our experiments share the basic tools𝒯basic\\mathcal\{T\}\_\{\\mathrm\{basic\}\}of equation[1](https://arxiv.org/html/2610.02616#S3.E1)\. They read and search trajectories, run shell commands for analysis, write notes that carry over to later rounds, and submit proposed edits to the executor harness\. Attribution and the training audit are not in𝒯0\\mathcal\{T\}\_\{0\}; they run automatically each round and add their reports to the optimizer’s evidence\. Self\-evolving optimizers also keep simple verification \(Section[2\.3](https://arxiv.org/html/2610.02616#S2.SS3)\), which theverificationtool generalizes to up to three tasks; their self\-written tools can call both\. Later optimizer harnesses keep all tools in𝒯0\\mathcal\{T\}\_\{0\}, can add new tools built from them, and can revise the other parts ofgg\.
### C\.1Choosing a repair from verification evidence
#### C\.1\.1Checks as an adaptive hypothesis test
Fix two distinct repairsh1,h2h\_\{1\},h\_\{2\}for two possible causesW∈\{1,2\}W\\in\\\{1,2\\\}with prior probabilitiesλ,1−λ\\lambda,1\-\\lambda,λ∈\[0,1\]\\lambda\\in\[0,1\]\. The optimizer starts from the evidenceℐ0\\mathcal\{I\}\_\{0\}gathered before any check, such as trajectories, attribution reports, earlier feedback, and notes, and may make up tonnchecks\. Checkjjconsists of a queryQjQ\_\{j\}and feedbackYjY\_\{j\}:
Qj∼Πj\(⋅∣ℐj−1\),Yj∼κW\(⋅∣ℐj−1,Qj\),ℐj=\(ℐj−1,Qj,Yj\)\.Q\_\{j\}\\sim\\Pi\_\{j\}\(\\cdot\\mid\\mathcal\{I\}\_\{j\-1\}\),\\qquad Y\_\{j\}\\sim\\kappa\_\{W\}\(\\cdot\\mid\\mathcal\{I\}\_\{j\-1\},Q\_\{j\}\),\\qquad\\mathcal\{I\}\_\{j\}=\(\\mathcal\{I\}\_\{j\-1\},Q\_\{j\},Y\_\{j\}\)\.\(5\)The query policyΠj\\Pi\_\{j\}belongs to the optimizer\. It may depend on the whole history, including revisions of the optimizer harness, but not onWW, which the optimizer cannot observe\. The feedback lawκw\\kappa\_\{w\}belongs to the sandbox and may differ between causes\. A run that stops early makes null checks with constant feedback untilnn; they carry no information\.
For probability massesP,P′P,P^\{\\prime\}on a countable set, the Bhattacharyya coefficient isAff\(P,P′\)=∑yP\(y\)P′\(y\)\\operatorname\{Aff\}\(P,P^\{\\prime\}\)=\\sum\_\{y\}\\sqrt\{P\(y\)P^\{\\prime\}\(y\)\}\([Bhattacharyya, 1943](https://arxiv.org/html/2610.02616#bib.bib5)\)\. It lies in\[0,1\]\[0,1\], equals11for identical distributions, and equals00for disjoint supports\. The coefficient of checkjjis
ρj=supℐ∑qΠj\(q∣ℐ\)Aff\(κ1\(⋅∣ℐ,q\),κ2\(⋅∣ℐ,q\)\),\\rho\_\{j\}=\\sup\_\{\\mathcal\{I\}\}\\sum\_\{q\}\\Pi\_\{j\}\(q\\mid\\mathcal\{I\}\)\\,\\operatorname\{Aff\}\\\!\\left\(\\kappa\_\{1\}\(\\cdot\\mid\\mathcal\{I\},q\),\\kappa\_\{2\}\(\\cdot\\mid\\mathcal\{I\},q\)\\right\),\(6\)where the supremum runs over histories at stepj−1j\-1with positive probability under both causes, andρj=0\\rho\_\{j\}=0if there are none\. It averages the ambiguity of the feedback over the checks the optimizer chooses at each history\. The sandbox determinesκ1,κ2\\kappa\_\{1\},\\kappa\_\{2\}and the optimizer determinesΠj\\Pi\_\{j\}; either can makeρj\\rho\_\{j\}small\.
#### C\.1\.2Proof of Theorem[1](https://arxiv.org/html/2610.02616#Thmtheorem1)
###### Lemma C\.1\.1\(Bayes error and the Bhattacharyya coefficient\)\.
LetP1,P2P\_\{1\},P\_\{2\}be the laws of the evidence under the two causes andβ=Aff\(P1,P2\)\\beta=\\operatorname\{Aff\}\(P\_\{1\},P\_\{2\}\)\. The smallest error probability of any decision rule,p∗=∑ℐmin\{λP1\(ℐ\),\(1−λ\)P2\(ℐ\)\}p^\{\*\}=\\sum\_\{\\mathcal\{I\}\}\\min\\\{\\lambda P\_\{1\}\(\\mathcal\{I\}\),\(1\-\\lambda\)P\_\{2\}\(\\mathcal\{I\}\)\\\}, satisfies
λ\(1−λ\)β2≤p∗≤λ\(1−λ\)β≤12β\.\\lambda\(1\-\\lambda\)\\beta^\{2\}\\leq p^\{\*\}\\leq\\sqrt\{\\lambda\(1\-\\lambda\)\}\\,\\beta\\leq\\tfrac\{1\}\{2\}\\beta\.
The upper bound is the Bhattacharyya bound on the Bayes error and the lower bound is due to[Kailath \(1967\)](https://arxiv.org/html/2610.02616#bib.bib18);[Nielsen \(2014\)](https://arxiv.org/html/2610.02616#bib.bib30)presents both\.
###### Proof\.
The rule that picks the cause with the larger posterior mass at eachℐ\\mathcal\{I\}leaves the smaller mass as its error, and no rule, randomized or not, does better pointwise\. Writea=λP1\(ℐ\)a=\\lambda P\_\{1\}\(\\mathcal\{I\}\)andb=\(1−λ\)P2\(ℐ\)b=\(1\-\\lambda\)P\_\{2\}\(\\mathcal\{I\}\), so that∑min\(a,b\)=p∗\\sum\\min\(a,b\)=p^\{\*\},∑max\(a,b\)=1−p∗\\sum\\max\(a,b\)=1\-p^\{\*\}, and∑ab=λ\(1−λ\)β\\sum\\sqrt\{ab\}=\\sqrt\{\\lambda\(1\-\\lambda\)\}\\,\\beta\. The upper bounds follow frommin\(a,b\)≤ab\\min\(a,b\)\\leq\\sqrt\{ab\}andλ\(1−λ\)≤14\\lambda\(1\-\\lambda\)\\leq\\tfrac\{1\}\{4\}\. For the lower bound,ab=min\(a,b\)max\(a,b\)\\sqrt\{ab\}=\\sqrt\{\\min\(a,b\)\\max\(a,b\)\}and the Cauchy–Schwarz inequality giveλ\(1−λ\)β2≤p∗\(1−p∗\)≤p∗\\lambda\(1\-\\lambda\)\\beta^\{2\}\\leq p^\{\*\}\(1\-p^\{\*\}\)\\leq p^\{\*\}\. ∎
###### Lemma C\.1\.2\(Each check multiplies the coefficient\)\.
LetPw\(j\)P\_\{w\}^\{\(j\)\}be the law ofℐj\\mathcal\{I\}\_\{j\}under causewwandβj=Aff\(P1\(j\),P2\(j\)\)\\beta\_\{j\}=\\operatorname\{Aff\}\(P\_\{1\}^\{\(j\)\},P\_\{2\}^\{\(j\)\}\)\. Thenβj≤ρjβj−1\\beta\_\{j\}\\leq\\rho\_\{j\}\\beta\_\{j\-1\}, andβj=βj−1\\beta\_\{j\}=\\beta\_\{j\-1\}whenκ1=κ2\\kappa\_\{1\}=\\kappa\_\{2\}at checkjj\.
###### Proof\.
Under equation[5](https://arxiv.org/html/2610.02616#A3.E5),Pw\(j\)\(ℐ,q,y\)=Pw\(j−1\)\(ℐ\)Πj\(q∣ℐ\)κw\(y∣ℐ,q\)P\_\{w\}^\{\(j\)\}\(\\mathcal\{I\},q,y\)=P\_\{w\}^\{\(j\-1\)\}\(\\mathcal\{I\}\)\\,\\Pi\_\{j\}\(q\\mid\\mathcal\{I\}\)\\,\\kappa\_\{w\}\(y\\mid\\mathcal\{I\},q\), and the factorΠj\\Pi\_\{j\}is the same under both causes\. Hence
βj\\displaystyle\\beta\_\{j\}=∑ℐP1\(j−1\)\(ℐ\)P2\(j−1\)\(ℐ\)∑qΠj\(q∣ℐ\)∑yκ1\(y∣ℐ,q\)κ2\(y∣ℐ,q\)≤ρjβj−1,\\displaystyle=\\sum\_\{\\mathcal\{I\}\}\\sqrt\{P\_\{1\}^\{\(j\-1\)\}\(\\mathcal\{I\}\)P\_\{2\}^\{\(j\-1\)\}\(\\mathcal\{I\}\)\}\\sum\_\{q\}\\Pi\_\{j\}\(q\\mid\\mathcal\{I\}\)\\sum\_\{y\}\\sqrt\{\\kappa\_\{1\}\(y\\mid\\mathcal\{I\},q\)\\kappa\_\{2\}\(y\\mid\\mathcal\{I\},q\)\}\\leq\\rho\_\{j\}\\beta\_\{j\-1\},\(7\)because the inner double sum is at mostρj\\rho\_\{j\}at every history with positive mass under both causes, and a history with zero mass under either cause contributes nothing\. Ifκ1=κ2\\kappa\_\{1\}=\\kappa\_\{2\}, the innermost sum equals11and the expression equalsβj−1\\beta\_\{j\-1\}\. ∎
Theorem[1](https://arxiv.org/html/2610.02616#Thmtheorem1)\(formal\)\.Letpn∗=infψℙ\{ψ\(ℐn\)≠hW\}p\_\{n\}^\{\*\}=\\inf\_\{\\psi\}\\mathbb\{P\}\\\{\\psi\(\\mathcal\{I\}\_\{n\}\)\\neq h\_\{W\}\\\}over decision rulesψ\\psion complete histories\. Then
pn∗≤λ\(1−λ\)β0∏j=1nρj≤12β0∏j=1nρj,p\_\{n\}^\{\*\}\\leq\\sqrt\{\\lambda\(1\-\\lambda\)\}\\,\\beta\_\{0\}\\prod\_\{j=1\}^\{n\}\\rho\_\{j\}\\leq\\tfrac\{1\}\{2\}\\,\\beta\_\{0\}\\prod\_\{j=1\}^\{n\}\\rho\_\{j\},\(8\)and ifκ1=κ2\\kappa\_\{1\}=\\kappa\_\{2\}at every check, thenpn∗≥λ\(1−λ\)β02p\_\{n\}^\{\*\}\\geq\\lambda\(1\-\\lambda\)\\beta\_\{0\}^\{2\}\. The optimizer’s choiceh^\\widehat\{h\}, which may use the history and private randomness, errs with probabilitypn∗\+εp\_\{n\}^\{\*\}\+\\varepsilonfor someε≥0\\varepsilon\\geq 0\.
###### Proof\.
Lemma[C\.1\.2](https://arxiv.org/html/2610.02616#A3.SS1.Thmlemma2)givesβn≤β0∏jρj\\beta\_\{n\}\\leq\\beta\_\{0\}\\prod\_\{j\}\\rho\_\{j\}, withβn=β0\\beta\_\{n\}=\\beta\_\{0\}when no check depends on the cause\. Lemma[C\.1\.1](https://arxiv.org/html/2610.02616#A3.SS1.Thmlemma1)applied to the laws ofℐn\\mathcal\{I\}\_\{n\}gives both bounds, andε≥0\\varepsilon\\geq 0becausepn∗p\_\{n\}^\{\*\}is the minimum over all rules\. ∎
The lower bound holds for every query policy, hence for every optimizer harness\. A check that only re\-analyzes the recorded evidence, such as re\-reading traces or running static analysis over them, hasκ1=κ2\\kappa\_\{1\}=\\kappa\_\{2\}and leaves the coefficient unchanged\. A check can lower it only if its feedback distinguishes the causes beyond the existing evidence; executing a draft repair provides such feedback\. For example,perturbationcan replace a command whose arguments may be wrong and replay the remaining steps\.
The confirmation rerun of theverificationtool lowers the coefficient further\. Consider a check that runs the draft repairh1h\_\{1\}on a target training task, and let it pass with probabilityq1q\_\{1\}when cause11is present andq2q\_\{2\}otherwise\. The coefficient of this check isρ=q1q2\+\(1−q1\)\(1−q2\)\\rho=\\sqrt\{q\_\{1\}q\_\{2\}\}\+\\sqrt\{\(1\-q\_\{1\}\)\(1\-q\_\{2\}\)\}\. Theverificationtool reruns the task after a pass\. With independent runs the outcomes are a fail, a pass followed by a fail, and two passes, and the coefficient becomes
\(1−q1\)\(1−q2\)\+q1q2ρ=ρ−q1q2\(1−ρ\),\\sqrt\{\(1\-q\_\{1\}\)\(1\-q\_\{2\}\)\}\+\\sqrt\{q\_\{1\}q\_\{2\}\}\\,\\rho=\\rho\-\\sqrt\{q\_\{1\}q\_\{2\}\}\\,\(1\-\\rho\),which is belowρ\\rhowheneverq1q2\>0q\_\{1\}q\_\{2\}\>0andρ<1\\rho<1\. A wrong repair that passes by chance with probabilityq2q\_\{2\}passes twice with probabilityq22q\_\{2\}^\{2\}\.
### C\.2Finding a good harness
#### C\.2\.1Conditions and proof of Theorem[2](https://arxiv.org/html/2610.02616#Thmtheorem2)
WriteL\(h\)=Lθ,D\(h\)L\(h\)=L\_\{\\theta,D\}\(h\)for a fixed executor and task distribution; expectations run over the whole run\. The run makesKKproposals in a fixed order;K=TMK=TMin Algorithm[1](https://arxiv.org/html/2610.02616#alg1), where the candidates of one round may run in parallel and their outcomes may be dependent\. Leth\(i\)h^\{\(i\)\}be the harness submitted by proposalii, repaired if it fails the smoke check \(Appendix[B](https://arxiv.org/html/2610.02616#A2)\), and let𝒞=\{h0,h\(1\),…,h\(K\)\}\\mathcal\{C\}=\\\{h\_\{0\},h^\{\(1\)\},\\ldots,h^\{\(K\)\}\\\}\. If proposaliiproduces no candidate,h\(i\)h^\{\(i\)\}is the current harness, so such failures only lowerrir\_\{i\}\. The selected harnessh^\\widehat\{h\}lies in𝒞\\mathcal\{C\}\.
Fixℓ0∈\[0,1\]\\ell\_\{0\}\\in\[0,1\]\. LetAiA\_\{i\}be the event thath0h\_\{0\}andh\(1\),…,h\(i\)h^\{\(1\)\},\\ldots,h^\{\(i\)\}all have loss aboveℓ0\\ell\_\{0\}, withA0=\{L\(h0\)\>ℓ0\}A\_\{0\}=\\\{L\(h\_\{0\}\)\>\\ell\_\{0\}\\\}\. LetOiO\_\{i\}be the event that proposaliitargets a failure whose correct repair has loss at mostℓ0\\ell\_\{0\}, andJiJ\_\{i\}the event that it submits that repair, so thatOi∩Ji⊆\{L\(h\(i\)\)≤ℓ0\}O\_\{i\}\\cap J\_\{i\}\\subseteq\\\{L\(h^\{\(i\)\}\)\\leq\\ell\_\{0\}\\\}\. The conditions of Theorem[2](https://arxiv.org/html/2610.02616#Thmtheorem2)are, for constantsri,ei∈\[0,1\]r\_\{i\},e\_\{i\}\\in\[0,1\],
ℙ\(Oi∣Ai−1\)≥ri,ℙ\(Jic∣Oi,Ai−1\)≤ei,\\mathbb\{P\}\(O\_\{i\}\\mid A\_\{i\-1\}\)\\geq r\_\{i\},\\qquad\\mathbb\{P\}\(J\_\{i\}^\{c\}\\mid O\_\{i\},A\_\{i\-1\}\)\\leq e\_\{i\},\(9\)whenever the conditioning events have positive probability; ifℙ\(Oi∣Ai−1\)=0\\mathbb\{P\}\(O\_\{i\}\\mid A\_\{i\-1\}\)=0we setri=0r\_\{i\}=0\. These conditions allow proposals to depend on earlier outcomes and on revisions of the optimizer harness; independence between proposals is not required\. The repair must lower the population loss, not only fix the training task it was found on\.
Theorem[2](https://arxiv.org/html/2610.02616#Thmtheorem2)\(formal\)\.Under equation[9](https://arxiv.org/html/2610.02616#A3.E9), withεsel=𝔼\[L\(h^\)−minh∈𝒞L\(h\)\]\\varepsilon\_\{\\mathrm\{sel\}\}=\\mathbb\{E\}\[L\(\\widehat\{h\}\)\-\\min\_\{h\\in\\mathcal\{C\}\}L\(h\)\],
𝔼L\(h^\)≤ℓ0\+\(1−ℓ0\)∏i=1K\[1−ri\(1−ei\)\]\+εsel\.\\mathbb\{E\}L\(\\widehat\{h\}\)\\leq\\ell\_\{0\}\+\(1\-\\ell\_\{0\}\)\\prod\_\{i=1\}^\{K\}\[1\-r\_\{i\}\(1\-e\_\{i\}\)\]\+\\varepsilon\_\{\\mathrm\{sel\}\}\.\(10\)Ifri≥rr\_\{i\}\\geq randei≤ee\_\{i\}\\leq efor allii, the middle term is at most\(1−ℓ0\)\[1−r\(1−e\)\]K\(1\-\\ell\_\{0\}\)\[1\-r\(1\-e\)\]^\{K\}\.
###### Proof\.
Whenℙ\(Ai−1\)\>0\\mathbb\{P\}\(A\_\{i\-1\}\)\>0andℙ\(Oi∣Ai−1\)\>0\\mathbb\{P\}\(O\_\{i\}\\mid A\_\{i\-1\}\)\>0,
ℙ\{L\(h\(i\)\)≤ℓ0∣Ai−1\}≥ℙ\(Oi∣Ai−1\)ℙ\(Ji∣Oi,Ai−1\)≥ri\(1−ei\),\\mathbb\{P\}\\\{L\(h^\{\(i\)\}\)\\leq\\ell\_\{0\}\\mid A\_\{i\-1\}\\\}\\geq\\mathbb\{P\}\(O\_\{i\}\\mid A\_\{i\-1\}\)\\,\\mathbb\{P\}\(J\_\{i\}\\mid O\_\{i\},A\_\{i\-1\}\)\\geq r\_\{i\}\(1\-e\_\{i\}\),and the bound also holds when either probability is zero\. SinceAi=Ai−1∩\{L\(h\(i\)\)\>ℓ0\}A\_\{i\}=A\_\{i\-1\}\\cap\\\{L\(h^\{\(i\)\}\)\>\\ell\_\{0\}\\\}, this givesℙ\(Ai\)≤\[1−ri\(1−ei\)\]ℙ\(Ai−1\)\\mathbb\{P\}\(A\_\{i\}\)\\leq\[1\-r\_\{i\}\(1\-e\_\{i\}\)\]\\mathbb\{P\}\(A\_\{i\-1\}\), and iterating fromℙ\(A0\)≤1\\mathbb\{P\}\(A\_\{0\}\)\\leq 1,
ℙ\{minh∈𝒞L\(h\)\>ℓ0\}=ℙ\(AK\)≤∏i=1K\[1−ri\(1−ei\)\]\.\\mathbb\{P\}\\Bigl\\\{\\min\_\{h\\in\\mathcal\{C\}\}L\(h\)\>\\ell\_\{0\}\\Bigr\\\}=\\mathbb\{P\}\(A\_\{K\}\)\\leq\\prod\_\{i=1\}^\{K\}\[1\-r\_\{i\}\(1\-e\_\{i\}\)\]\.The minimum loss is at mostℓ0\\ell\_\{0\}offAKA\_\{K\}and at most11on it, so𝔼min𝒞L≤ℓ0\+\(1−ℓ0\)ℙ\(AK\)\\mathbb\{E\}\\min\_\{\\mathcal\{C\}\}L\\leq\\ell\_\{0\}\+\(1\-\\ell\_\{0\}\)\\mathbb\{P\}\(A\_\{K\}\)\. Addingεsel\\varepsilon\_\{\\mathrm\{sel\}\}proves the claim\. ∎
#### C\.2\.2Connecting verification to the repair error
Theorem[1](https://arxiv.org/html/2610.02616#Thmtheorem1)boundseie\_\{i\}when the repair decision of proposaliiis a test between two causes\. Condition onOiO\_\{i\},Ai−1A\_\{i\-1\}, and the proposal’s context, and suppose that within this context the optimizer runs the adaptive test of Appendix[C\.1](https://arxiv.org/html/2610.02616#A3.SS1)with at mostnin\_\{i\}checks and submits the repair it selects\. Letρi,j\\rho\_\{i,j\}bound the coefficient of checkjjandεi\\varepsilon\_\{i\}the optimizer’s excess error, uniformly over contexts\. Applying Theorem[1](https://arxiv.org/html/2610.02616#Thmtheorem1)in each context and averaging gives
ℙ\(Jic∣Oi,Ai−1\)≤ei:=min\{1,12∏j=1niρi,j\+εi\},\\mathbb\{P\}\(J\_\{i\}^\{c\}\\mid O\_\{i\},A\_\{i\-1\}\)\\leq e\_\{i\}:=\\min\\\!\\left\\\{1,\\frac\{1\}\{2\}\\prod\_\{j=1\}^\{n\_\{i\}\}\\rho\_\{i,j\}\+\\varepsilon\_\{i\}\\right\\\},\(11\)which can be substituted into equation[10](https://arxiv.org/html/2610.02616#A3.E10)\.
If every proposal makesnnchecks with coefficient at mostρ<1\\rho<1and excess error at mostε\\varepsilon, thenei≤min\{1,12ρn\+ε\}e\_\{i\}\\leq\\min\\\{1,\\tfrac\{1\}\{2\}\\rho^\{n\}\+\\varepsilon\\\}for everyii\. The first term decays geometrically in the number of checks, so a few checks bring the bound close toε\\varepsilon; further gains come from lowerε\\varepsilonand higherrir\_\{i\}, which self\-evolution can act on\.
The two factors ofri\(1−ei\)r\_\{i\}\(1\-e\_\{i\}\)multiply\. Increasingrir\_\{i\}byΔ\\Deltaraises the lower bound on the success probability of a proposal by onlyΔ\(1−ei\)\\Delta\(1\-e\_\{i\}\), so better target selection helps little while repairs often fail\. Within each benchmark, all methods use the sameKK\(Section[4\.1](https://arxiv.org/html/2610.02616#S4.SS1)\), so in equation[4](https://arxiv.org/html/2610.02616#S3.E4)they differ only throughrir\_\{i\},eie\_\{i\}, andεsel\\varepsilon\_\{\\mathrm\{sel\}\}\.
### C\.3Selection error
#### C\.3\.1Two stages
Theorem[2](https://arxiv.org/html/2610.02616#Thmtheorem2)leaves the selection errorεsel\\varepsilon\_\{\\mathrm\{sel\}\}: even when a good harness has been proposed, the run must recognize it from measured losses\. Selection has two stages \(Algorithm[1](https://arxiv.org/html/2610.02616#alg1)\)\. Within a round, the candidate with the lowest training loss is kept, and these candidates together withh0h\_\{0\}form the archive𝒱\\mathcal\{V\}; at the end, the archived harness with the lowest validation loss is selected\. The first stage uses tasks whose trajectories the optimizer has read; the second uses tasks of which it has seen only validation losses and outcome summaries\. The lemma charges each stage for its estimation error\.
###### Lemma C\.3\.1\(Selection in two stages\)\.
Suppose that with probability at least1−δ1\-\\delta, every training loss is withinatraina\_\{\\mathrm\{train\}\}of its candidate’s population loss and every validation loss on𝒱\\mathcal\{V\}is withinavala\_\{\\mathrm\{val\}\}\. Then, withc=min\{1,2atrain\+2aval\}c=\\min\\\{1,2a\_\{\\mathrm\{train\}\}\+2a\_\{\\mathrm\{val\}\}\\\}, we haveεsel≤c\+\(1−c\)δ\\varepsilon\_\{\\mathrm\{sel\}\}\\leq c\+\(1\-c\)\\delta\.
###### Proof\.
Leth∗h\_\{\*\}minimizeLLon𝒞\\mathcal\{C\}\. On the stated event, ifh∗h\_\{\*\}was proposed in a round whose kept candidate ishscrh^\{\\mathrm\{scr\}\}, then
L\(hscr\)≤L^train\(hscr\)\+atrain≤L^train\(h∗\)\+atrain≤L\(h∗\)\+2atrain,L\(h^\{\\mathrm\{scr\}\}\)\\leq\\widehat\{L\}\_\{\\mathrm\{train\}\}\(h^\{\\mathrm\{scr\}\}\)\+a\_\{\\mathrm\{train\}\}\\leq\\widehat\{L\}\_\{\\mathrm\{train\}\}\(h\_\{\*\}\)\+a\_\{\\mathrm\{train\}\}\\leq L\(h\_\{\*\}\)\+2a\_\{\\mathrm\{train\}\},and ifh∗=h0h\_\{\*\}=h\_\{0\}it is in𝒱\\mathcal\{V\}already; somin𝒱L≤min𝒞L\+2atrain\\min\_\{\\mathcal\{V\}\}L\\leq\\min\_\{\\mathcal\{C\}\}L\+2a\_\{\\mathrm\{train\}\}\. The same comparison for validation selection on𝒱\\mathcal\{V\}givesL\(h^\)≤min𝒱L\+2avalL\(\\widehat\{h\}\)\\leq\\min\_\{\\mathcal\{V\}\}L\+2a\_\{\\mathrm\{val\}\}\. The gapL\(h^\)−min𝒞LL\(\\widehat\{h\}\)\-\\min\_\{\\mathcal\{C\}\}Lis therefore at mostccon the event and at most11off it; take expectations\. ∎
This is the usual comparison between empirical and population minimizers\([Hastie et al\., 2009](https://arxiv.org/html/2610.02616#bib.bib12)\), applied twice\. It also covers the final step of the protocol, which re\-evaluates the two archived harnesses with the lowest validation losses and keeps the one with the lower mean: the minimizer ofLLon𝒱\\mathcal\{V\}is either among the two, and then the winner’s mean loss is no larger than its own, or it is not, and then both finalists have first losses no larger than its own\. In both casesL\(h^\)≤min𝒱L\+2avalL\(\\widehat\{h\}\)\\leq\\min\_\{\\mathcal\{V\}\}L\+2a\_\{\\mathrm\{val\}\}as long as first losses and mean losses are withinavala\_\{\\mathrm\{val\}\}of the population loss\. The optimizer builds its candidates from the training trajectories, soatraina\_\{\\mathrm\{train\}\}is a condition on the training comparison rather than a consequence of sampling\. In our implementation, a candidate that fails the smoke check is repaired before validation without being screened again \(Appendix[B](https://arxiv.org/html/2610.02616#A2)\); such repairs address loading and runtime errors\. The lemma covers this case with the repaired version as the candidate in𝒞\\mathcal\{C\}and its training loss measured before the repair\. The condition onatraina\_\{\\mathrm\{train\}\}then compares this measured loss with the population loss of the repaired version\. The next subsection boundsavala\_\{\\mathrm\{val\}\}\.
#### C\.3\.2Validation losses under feedback
The optimizer never executes validation tasks, but their losses and outcome summaries are returned to it and shape later candidates\. Each round therefore chooses a harness after looking at the validation set through this feedback\. This is adaptive data analysis: how much the validation loss can flatter the chosen harness is governed by how much the feedback reveals, and can be bounded by the number of possible values of the feedback\([Dwork et al\., 2015](https://arxiv.org/html/2610.02616#bib.bib10);[Blum & Hardt, 2015](https://arxiv.org/html/2610.02616#bib.bib6)\)or by its entropy\([Russo & Zou, 2016](https://arxiv.org/html/2610.02616#bib.bib33);[Xu & Raginsky, 2017](https://arxiv.org/html/2610.02616#bib.bib47)\)\.[Bertran et al\. \(2026\)](https://arxiv.org/html/2610.02616#bib.bib4)apply the same argument to research agents that receive one bit of validation feedback per submission\. We give both forms\.
LetX1,…,Xm∼DX\_\{1\},\\ldots,X\_\{m\}\\sim Dbe the validation tasks and, for each harnesshhand run indexrr, letYi\(r\)\(h\)∼πθ\(⋅∣Xi;h\)Y\_\{i\}^\{\(r\)\}\(h\)\\sim\\pi\_\{\\theta\}\(\\cdot\\mid X\_\{i\};h\)be a complete trajectory; given the task, reruns are independent\. Collect everything about taskiiinZi=\(Xi,\(Yi\(r\)\(h\)\)h,r\)Z\_\{i\}=\(X\_\{i\},\(Y\_\{i\}^\{\(r\)\}\(h\)\)\_\{h,r\}\)\. TheZiZ\_\{i\}are i\.i\.d\. and independent ofUU, the training data and the random seeds that drive the optimizer and every training\-side execution\. A validation estimate is an averageϕ^=1m∑iϕ\(Zi\)\\widehat\{\\phi\}=\\frac\{1\}\{m\}\\sum\_\{i\}\\phi\(Z\_\{i\}\)of a per\-task statisticϕ\(Zi\)∈\[0,1\]\\phi\(Z\_\{i\}\)\\in\[0,1\]with meanϕ¯\\bar\{\\phi\}\. The validation loss ofhhuses one run per task,L^val\(h\)=1m∑iℓ\(Xi,Yi\(1\)\(h\)\)\\widehat\{L\}\_\{\\mathrm\{val\}\}\(h\)=\\frac\{1\}\{m\}\\sum\_\{i\}\\ell\(X\_\{i\},Y\_\{i\}^\{\(1\)\}\(h\)\), and has meanL\(h\)L\(h\)\.
LetF=\(F0,…,FT−1\)F=\(F\_\{0\},\\ldots,F\_\{T\-1\}\)be the feedback returned across rounds\. We assume thatFFtakes values in a finite setℱ\\mathcal\{F\}, as it does when every summary has bounded length\. Our summaries list at most 30 task identifiers and the two most frequent error messages\. The number of possible summaries can still be large, and the bounds below are informative only whenlog\|ℱ\|\\log\|\\mathcal\{F\}\|is small relative tomm\. GivenUUandF0,…,Ft−1F\_\{0\},\\ldots,F\_\{t\-1\}, the run is deterministic up to the validation of roundtt: the validated harness, the statistics that evaluate it, and the optimizer harnessgtg\_\{t\}are all functions of the feedback so far\.
###### Proposition C\.3\.2\(Uniform deviation of validation estimates\)\.
Let each validated harness receive at mostRRestimates\. For0<δ<10<\\delta<1, with probability at least1−δ1\-\\deltaevery estimate the run computes satisfies
\|ϕ^−ϕ¯\|≤εm=log\|ℱ\|\+log\(2R\(T\+1\)/δ\)2m,\|\\widehat\{\\phi\}\-\\bar\{\\phi\}\|\\leq\\varepsilon\_\{m\}=\\sqrt\{\\frac\{\\log\|\\mathcal\{F\}\|\+\\log\(2R\(T\+1\)/\\delta\)\}\{2m\}\},and so does every average of such estimates\.
###### Proof\.
Condition onU=uU=u; theZiZ\_\{i\}keep their law\. The harness of roundt∈\{0,…,T\}t\\in\\\{0,\\ldots,T\\\}, where round00validatesh0h\_\{0\}, and each of itsRRstatistics are fixed onceF0,…,Ft−1F\_\{0\},\\ldots,F\_\{t\-1\}are, and this prefix takes at most\|ℱ\|\|\\mathcal\{F\}\|values\. So at mostR\(T\+1\)\|ℱ\|R\(T\+1\)\|\\mathcal\{F\}\|statistics, each fixed before theZiZ\_\{i\}are drawn, can ever be computed\. Hoeffding’s inequality\([Hoeffding, 1963](https://arxiv.org/html/2610.02616#bib.bib14)\)givesℙ\(\|ϕ^−ϕ¯\|\>ε\)≤2e−2mε2\\mathbb\{P\}\(\|\\widehat\{\\phi\}\-\\bar\{\\phi\}\|\>\\varepsilon\)\\leq 2e^\{\-2m\\varepsilon^\{2\}\}for each, and a union bound atε=εm\\varepsilon=\\varepsilon\_\{m\}leaves failure probabilityδ\\delta\. Averages inherit the bound by the triangle inequality; integrate overuu\. ∎
###### Proposition C\.3\.3\(Expected optimism\)\.
Leth^\\widehat\{h\}be an archived harness chosen by any rule\. Then, with entropies in nats,
𝔼\[L\(h^\)−L^val\(h^\)\]≤H\(F∣U\)\+log\(T\+1\)2m≤log\|ℱ\|\+log\(T\+1\)2m\.\\mathbb\{E\}\\bigl\[L\(\\widehat\{h\}\)\-\\widehat\{L\}\_\{\\mathrm\{val\}\}\(\\widehat\{h\}\)\\bigr\]\\leq\\sqrt\{\\frac\{H\(F\\mid U\)\+\\log\(T\+1\)\}\{2m\}\}\\leq\\sqrt\{\\frac\{\\log\|\\mathcal\{F\}\|\+\\log\(T\+1\)\}\{2m\}\}\.
###### Proof\.
Condition onU=uU=u\. The archive is a function ofFFandh^\\widehat\{h\}is its entry at some indexJ∈\{0,…,T\}J\\in\\\{0,\\ldots,T\\\}, soH\(h^∣U=u\)≤H\(F∣U=u\)\+log\(T\+1\)H\(\\widehat\{h\}\\mid U=u\)\\leq H\(F\\mid U=u\)\+\\log\(T\+1\)\. The lossℓ\(Xi,Yi\(1\)\(h\)\)∈\[0,1\]\\ell\(X\_\{i\},Y\_\{i\}^\{\(1\)\}\(h\)\)\\in\[0,1\]is12\\tfrac\{1\}\{2\}\-sub\-Gaussian for everyhh, andh^\\widehat\{h\}is a function ofZ1:mZ\_\{1:m\}\. Theorem 1 of[Xu & Raginsky \(2017\)](https://arxiv.org/html/2610.02616#bib.bib47)withσ=12\\sigma=\\tfrac\{1\}\{2\}bounds the expected gap byI\(Z1:m;h^∣U=u\)/\(2m\)\\sqrt\{I\(Z\_\{1:m\};\\widehat\{h\}\\mid U=u\)/\(2m\)\}, andI\(Z1:m;h^∣U=u\)≤H\(h^∣U=u\)I\(Z\_\{1:m\};\\widehat\{h\}\\mid U=u\)\\leq H\(\\widehat\{h\}\\mid U=u\)\. Average overuuwith Jensen’s inequality, and useH\(F∣U\)≤log\|ℱ\|H\(F\\mid U\)\\leq\\log\|\\mathcal\{F\}\|\. ∎
Both bounds depend on the feedback, not on the harness\. The size and content of what the optimizer writes do not enter them, and the optimizer harnessgtg\_\{t\}is itself a function ofUUand the feedback so far, so both loops are covered by the same feedback\. What matters islog\|ℱ\|\\log\|\\mathcal\{F\}\|\. With no feedback it is00and only the choice amongT\+1T\+1archived harnesses remains; a validation loss per round contributes at mostlog\(m\+1\)\\log\(m\+1\); an improved\-or\-not bit contributeslog2\\log 2; a list of which tasks changed outcome contributes up tomlog3m\\log 3; and returning validation trajectories would make\|ℱ\|\|\\mathcal\{F\}\|far larger\. This is the trade\-off behind the protocol of Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2)\. Feedback can raiserir\_\{i\}and lowereie\_\{i\}in Theorem[2](https://arxiv.org/html/2610.02616#Thmtheorem2)by showing which edits held up on unseen tasks, and it enlargesavala\_\{\\mathrm\{val\}\}throughlog\|ℱ\|\\log\|\\mathcal\{F\}\|; restricting its possible values lowers this cost\.
The protocol of Appendix[B](https://arxiv.org/html/2610.02616#A2)also reruns a validation task whose outcome differs from the parent’s and keeps the change only if it reproduces\. Such a confirmed estimate is an average of per\-task statistics and is covered by Proposition[C\.3\.2](https://arxiv.org/html/2610.02616#A3.SS3.Thmlemma2), but its mean is not exactlyL\(h\)L\(h\)\. On a task wherehhfails with probabilitypp, letaabe its first outcome,bbthe parent’s, andrrthe rerun; the confirmed outcome equalsaawhena=ba=bandrrotherwise, so𝔼\[c−a∣x\]=pℙ\(a=0,b=1\)−\(1−p\)ℙ\(a=1,b=0\)\\mathbb\{E\}\[c\-a\\mid x\]=p\\,\\mathbb\{P\}\(a=0,b=1\)\-\(1\-p\)\\,\\mathbb\{P\}\(a=1,b=0\), and both terms lie in\[0,p\(1−p\)\]\[0,p\(1\-p\)\]\. The mean of a confirmed estimate is therefore withinv\(h\)=𝔼x∼D\[ph\(x\)\(1−ph\(x\)\)\]≤14v\(h\)=\\mathbb\{E\}\_\{x\\sim D\}\[p\_\{h\}\(x\)\(1\-p\_\{h\}\(x\)\)\]\\leq\\tfrac\{1\}\{4\}ofL\(h\)L\(h\), withv\(h\)=0v\(h\)=0when outcomes are deterministic given the task\. In Lemma[C\.3\.1](https://arxiv.org/html/2610.02616#A3.SS3.Thmlemma1)this term adds toavala\_\{\\mathrm\{val\}\}: with the two estimates per harness that the final step uses,aval=εma\_\{\\mathrm\{val\}\}=\\varepsilon\_\{m\}atR=2R=2plusmaxhv\(h\)\\max\_\{h\}v\(h\)\.
## Appendix DAdditional experiment details
### D\.1Data and grading
SWE\-rebench splits\.The training and validation pools draw from the monthly leaderboard pool of the benchmark\. The test split was frozen separately, before any method development, and an n\-gram overlap scan between the training or validation task texts and the test tasks found no contamination\. The sealed test pool holds 110 tasks\. Two of them fail their own official reference solution under our grading pipeline and are excluded, so in\-distribution SWE\-rebench accuracies use 108 tasks\.
Out\-of\-distribution test set\.The OOD test set of Table[2](https://arxiv.org/html/2610.02616#S4.T2)is the July 2026 release of the SWE\-rebench leaderboard, which covers tasks from May 15 to July 1, 2026\. It holds 111 tasks in Go \(21\), Java \(21\), Python \(20\), Rust \(25\), and TypeScript \(24\)\. Four Java tasks fail their reference solution in our environment and are excluded, so OOD accuracies use 107 tasks: Go 21, Java 17, Python 20, Rust 25, and TypeScript 24\. On the 87 tasks outside Python, blank scores25\.3%25\.3\\%, and VERSE with self\-evolution scores35\.6%35\.6\\%on the Meta\-Harness host and37\.9%37\.9\\%on the HarnessX host, the two hosts with the largest OOD gains in Table[2](https://arxiv.org/html/2610.02616#S4.T2)\. Three of the 20 Python tasks come from a repository that also appears in the training and validation sets; the in\-distribution test set shares no repository with them\.
Terminal\-Bench splits\.The Terminal\-Bench\([Merrill et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib27)\)pools follow the suite generations, which share no tasks\. Training uses 104 Terminal\-Bench 1\.x tasks, and the test split is the 89 tasks of Terminal\-Bench 2\.x at a pinned commit, so training never sees a test environment\. The validation pool holds 98 tasks \(20 from Terminal\-Bench 1\.x and 78 from Terminal\-Bench\-Lite\), and each run samples 50 of them with a fixed seed\.
Grading\.Each task is graded by its own test script, and task environments are frozen as image snapshots, so the checker version cannot drift during a run\. Every validation task passed a screening check that its official reference solution passes our grading pipeline\. Grader logs are available on training tasks only\.
### D\.2Models and budgets
Models\.Sections[2](https://arxiv.org/html/2610.02616#S2)and[4](https://arxiv.org/html/2610.02616#S4)use Qwen3\.8\-Flash\-Next as the optimizer and Qwen3\.8\-27B as the executor, served through vLLM; the executor runs with low reasoning effort and greedy decoding\. The appendix also reports results with closed\-weight models: the small tier of an anonymized model family as the executor and its mid\-size tier as the optimizer\. We refer to them as the small model and the mid\-size model\. These models satisfy the two conditions[Wang et al\. \(2026b\)](https://arxiv.org/html/2610.02616#bib.bib42)identify for informative harness evolution studies: the executor has headroom, and performance depends strongly on the harness\. Observation 2 \(Figures[3](https://arxiv.org/html/2610.02616#S3.F3)and[13](https://arxiv.org/html/2610.02616#A11.F13)\) also uses four further open\-weight executors: Qwen3\-Coder\-480B, DeepSeek V3\.2, Mistral Large 3, and GLM\-5\. The first two also serve as target executors in the generalization study of Appendix[G](https://arxiv.org/html/2610.02616#A7)\.
Budget\.Every method runsT=6T=6evolution rounds on SWE\-rebench andT=3T=3on Terminal\-Bench, with three candidates generated per round, the same proposal\-turn limit, the same sandbox parallelism, and the same harness edit space \(executable hooks plus prompts\)\. The only axis that varies between methods is what the optimizer is given: the evidence interface, the verification tools, and the right to edit its own harness\.
Simple verification budget\.In Table[1](https://arxiv.org/html/2610.02616#S2.T1), simple verification executes one training task under an optional draft edit in an isolated copy of the executor harness\. We call each such execution, and each call to a VERSE verification tool, a*probe*\. Each call returns the test verdict, the trajectory path, and the remaining execution budget\. The configured allowance is eight calls per candidate, scaled by the three candidates to a shared pool of 24 calls per round\. The three candidates and the subsequent optimizer self\-update draw from this same pool\. A passing probe does not automatically accept an edit; the optimizer decides whether to revise, probe again, or submit, after which the shared screening and validation procedure applies\.
Candidate search and verification costs\.Each candidate has at most 60 optimizer turns\. The candidates are compared on a shared subset of up to 12 training tasks within that round; the same comparison procedure and validation\-selection rule apply to baselines and VERSE\. A failed generation can leave fewer than three candidates\.
For variants with verification, the configured allowance is eight experiment units per candidate, pooled into a shared limit of 24 units per round\. This is one common pool, not three independently enforced limits: calls to the three verification tools and, for self\-evolving optimizers, to simple verification, including calls during the optimizer self\-update, draw from it\. In the four self\-evolving runs of Table[2](https://arxiv.org/html/2610.02616#S4.T2), simple verification was called 21 times, against 140 calls to the three verification tools\. Averificationcall runs at most three target tasks and costs one unit per target task; areplayorperturbationcall costs one unit\. Units count experiments, not physical executions: a successfulverificationcall includes a confirmation rerun, and aperturbationcall can require both a baseline and a modified replay\. All these calls useStrainS\_\{\\mathrm\{train\}\}only\.
Baselines without verification receive no probe allowance, and these extra executions are not offset by reducing other runs\. Automatic attribution has a separate time budget; optimizer self\-updates also require additional model calls\. Thus the comparisons match candidate count, proposal\-turn limits, and the evaluation protocol, but not total execution or token cost; Appendix[D\.3](https://arxiv.org/html/2610.02616#A4.SS3)reports the cost of each run\. Fixed and self\-evolving variants with verification share the same probe allowance\.
### D\.3Cost of evolution
Table 4:Cost of each evolution run of Table[2](https://arxiv.org/html/2610.02616#S4.T2), counted from the run logs\.*Agent runs*: runs of the executor on one task, with model calls, excluding test evaluation;*Verification*: the agent runs of the verification tools and simple verification, including their confirmation reruns;*Units*: the verification budget spent;*Replays*: re\-executions of recorded commands without model calls;*Tokens*: executor tokens in the training, screening, and validation runs \(millions\);*Opt\. calls*: optimizer model turns while generating candidates and during self\-updates;*Hours*: wall\-clock time from the start of round 1 to the end of harness selection\.Table[4](https://arxiv.org/html/2610.02616#A4.T4)reports what each run executed\. An agent run is one run of the executor on one task\. Agent runs comprise the training, screening, and validation runs of every round, the re\-evaluation of the two finalists, the confirmation reruns of validation changes, the smoke checks, and the agent runs of the verification tools\. Replays re\-execute recorded commands in a fresh container without calling a model; trace minimization and thereplaytool use one replay, and aperturbationcall uses two, one with the original step and one with the changed step\.
Units count verification experiments rather than executions: averificationcall costs one unit per target task, but each target that passes also gets a confirmation run\. Optimizer tokens were not logged, so the table counts optimizer model turns instead\. Test evaluation is the same for every method and adds three runs of the 108 in\-distribution and 107 out\-of\-distribution tasks for each evaluated harness\.
Training, screening, and validation runs account for 91 to 98% of the agent runs of every evolution run\. The verification tools add 27 to 104 agent runs, 2 to 7% of a run’s agent runs, and at most 37 replays\. The larger gaps in agent runs between a baseline and its VERSE versions come from rounds in which none of the baseline’s candidates could be applied, so the round ran no screening or validation \(AHE and HarnessX\)\. The self\-evolving optimizer on the AHE host ends each candidate after a median of 26 turns instead of 58, so it makes fewer optimizer calls than the baseline despite its self\-updates\.
VERSE runs take 1\.3 to 2\.5 times as long as their baselines in wall\-clock time\. Waiting for verification runs while candidates are generated accounts for 1\.0 to 4\.7 hours of this\. Most of the rest is screening and validation, which take longer when an evolved harness makes each executor run longer and when a baseline skipped rounds; the load on the shared executor servers also differed between runs\.
### D\.4Counting draft evaluations
The top panel of Figure[5](https://arxiv.org/html/2610.02616#S4.F5)and the draft counts of Section[4\.4](https://arxiv.org/html/2610.02616#S4.SS4)follow these rules\. In the evolution runs of Tables[1](https://arxiv.org/html/2610.02616#S2.T1)–[5](https://arxiv.org/html/2610.02616#S4.F5), theverificationtool made 170 draft evaluations\. Each run of the tool counts as one evaluation, so a draft tested twice counts twice\. The tool reports a tested task as passed only after it passes twice\. For each passed task, we check whether the harness being edited passed it in its own training run\. This sorts the 170 evaluations into four groups:
- •in 17, the draft fixed a task that was failing;
- •in 30, it passed only on tasks that already passed, added by the optimizer as regression checks;
- •in 112, it passed none of its tasks;
- •11 were inconclusive: 10 inconclusive verdicts and one pass whose task is missing from the log\.
Some drafts in the first two groups became the round’s harness unchanged\. Across these drafts, all 5 tasks they fixed and 11 of the 12 tasks they used as regression checks passed again in the next training run\. When we report how often optimizers revised a draft that passed none of its tasks, a draft counts as revised if the final submission of its candidate differs from it\.
Table 5:The four baselines as implemented in their original papers, from the respective publications\. In our experiments \(Section[4\.1](https://arxiv.org/html/2610.02616#S4.SS1)\) all methods are reimplemented under one protocol: same executor, same edit space, same candidate count and proposal\-turn limit, and the same optimizer model for every method; only the evidence interface of each method is preserved\.
## Appendix EBaseline reimplementation details
Table[5](https://arxiv.org/html/2610.02616#A4.T5)compares how the four published harness optimizers are implemented in their original papers\. All four keep the optimizer fixed while it evolves the executor harness; they differ in how the optimizer is built, which model powers it, and how an edit is checked\. VERSE differs on the two axes this paper studies: the optimizer harness is itself optimized, and the execution check is a tool the optimizer calls before submitting an edit, not an outer\-loop gate applied afterwards\.
Evidence interfaces\.The four methods differ in the evidence interface through which the optimizer digests failures: Meta\-Harness reads raw traces directly, AHE distills them into per\-task root\-cause reports, Self\-Harness mines model\-specific weakness patterns, and HarnessX summarizes trajectories on demand\.
Evaluation protocols of the original papers\.The original papers use different evaluation protocols\. Meta\-Harness uses separate test sets in several domains and reports its Terminal\-Bench 2 results on the tasks used for search\([Lee et al\., 2026b](https://arxiv.org/html/2610.02616#bib.bib21)\); AHE and HarnessX report results on the task set used for evolution\([Lin et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib23);[Chen et al\., 2026c](https://arxiv.org/html/2610.02616#bib.bib9)\); Self\-Harness hides held\-out traces from the proposer and uses held\-out scores to accept edits\([Zhang et al\., 2026a](https://arxiv.org/html/2610.02616#bib.bib55)\)\. We therefore evaluate all methods under the protocol of Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2)\.
Our reimplementations preserve each method’s evidence interface, the column that distinguishes the four systems in Table[5](https://arxiv.org/html/2610.02616#A4.T5)\. Everything else is unified for the controlled comparison: all four run inside our engine with the same executor and harness edit space \(executable hooks plus prompts\), the same optimizer model, the same candidate count and proposal\-turn limit, the same round structure, the same harness selection rule on the validation split, and the same evaluation protocol \(Section[4\.1](https://arxiv.org/html/2610.02616#S4.SS1)\)\. The original systems use different executors, edit spaces, optimizer models, acceptance rules, and evaluation protocols\.
Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2)explains how our shared data splits separate optimization feedback from final test performance\. Test\-time scaling methods are not comparable baselines under this protocol, because they produce per\-instance adaptations rather than a reusable harness\.
## Appendix FAdditional main results
This section repeats the comparison of Section[4\.2](https://arxiv.org/html/2610.02616#S4.SS2)with the small model as the executor and the mid\-size model as the optimizer \(Appendix[D\.2](https://arxiv.org/html/2610.02616#A4.SS2)\)\. It covers SWE\-rebench and Terminal\-Bench, and repeats the evolution three times for the four baselines and for VERSE on the HarnessX host\.
Table 6:Each baseline equipped with VERSE, small and mid\-size models\(avg@3 accuracy, %,±\\pmSEM over three test evaluations; one evolution run per entry\)\. “\+\+VERSE \(w/o self\-evolve\)” adds the verification tools to the baseline above; “w/ self\-evolve” also lets the optimizer edit its own harness\.r⋆r^\{\\star\}: the round selected on validation;*Val\-selected*: the harness of that round;*Last*: the final\-round harness\.Bold: higher mean than the baseline of its block\.Underline: w/ self\-evolve better than w/o\.Table[6](https://arxiv.org/html/2610.02616#A6.T6)reports one evolution run per entry, Table[7](https://arxiv.org/html/2610.02616#A6.T7)the mean of three runs per method, and Table[8](https://arxiv.org/html/2610.02616#A6.T8)every run\. Every run starts fromh0h\_\{0\}; atr⋆=0r^\{\\star\}=0an entry re\-evaluatesh0h\_\{0\}\.
The verification tools help every baseline\.On SWE\-rebench, adding them improves all four baselines at the selected round, and every equipped baseline reaches or exceeds the strongest unequipped one\. On Terminal\-Bench, the equipped baselines match or improve on their unequipped versions, with margins close to the evaluation noise\.
VERSE is the strongest configuration on both benchmarks\.Its three\-run mean is the highest in both views on both benchmarks, and at the selected round each of its runs scores above the mean of every baseline\. The three\-run study uses the HarnessX host; the AHE\-hosted run of Table[6](https://arxiv.org/html/2610.02616#A6.T6)scores higher on SWE\-rebench \(34\.8834\.88\) but is a single run\. Terminal\-Bench is less sensitive to the harness in this setting: most methods stay within about two points of blank, in line with[Wang et al\. \(2026b\)](https://arxiv.org/html/2610.02616#bib.bib42)\.
The final round is less stable without verification tools\.On SWE\-rebench, three of the four baselines lose accuracy from the selected to the final round, and a single run can lose several points\. With verification tools, the two views usually stay within two points of each other, although some runs lose more\.
Table 7:Three\-run means with the small and mid\-size models\(avg@3 accuracy, %,±\\pmSEM across three evolution runs, except for Blank, where it is over three test evaluations ofh0h\_\{0\}; every run is listed in Table[8](https://arxiv.org/html/2610.02616#A6.T8)\)\.*Val\-selected*: the harness of the round selected on validation;*Last*: the final\-round harness\. VERSE \(bold\) uses the HarnessX host\.Theunderlinedrun of each method in Table[8](https://arxiv.org/html/2610.02616#A6.T8)is the run shown in Table[6](https://arxiv.org/html/2610.02616#A6.T6); the generalization study of Appendix[G](https://arxiv.org/html/2610.02616#A7)starts from its harness\.
Figure[6](https://arxiv.org/html/2610.02616#A6.F6)shows the three runs of each method in Table[7](https://arxiv.org/html/2610.02616#A6.T7)\. On Terminal\-Bench, the lowest VERSE run scores above every baseline run\.
Figure 6:The evolution runs behind Table[7](https://arxiv.org/html/2610.02616#A6.T7)\.Validation\-selected avg@3 test accuracy \(%\) of each of the three evolution runs per method \(dots; values in Table[8](https://arxiv.org/html/2610.02616#A6.T8)\), with the three\-run mean \(bar\) and blank \(dashed line\)\.Table 8:Every evolution run behind Table[7](https://arxiv.org/html/2610.02616#A6.T7)\(avg@3 accuracy, %,±\\pmSEM over three test evaluations\)\. Each method is evolved three times from scratch, one run per row;*avg*rows give the mean and the SEM across runs\.r⋆r^\{\\star\}: the round selected on validation;*Val\-selected*: the harness of that round;*Last*: the final\-round harness\. Atr⋆=0r^\{\\star\}=0the entry re\-evaluatesh0h\_\{0\}\.Underline: the run shown in Table[6](https://arxiv.org/html/2610.02616#A6.T6), from which the generalization study starts\.
## Appendix GGeneralization study
This appendix asks what an evolved harness carries beyond the setting in which it was evolved\. We study three shifts, and they separate what transfers from what does not\. With the same executor and the same kind of task, the gains of VERSE carry over from Python to newer tasks in four other programming languages \(Section[G\.1](https://arxiv.org/html/2610.02616#A7.SS1)\)\. When the task domain changes, the executor harness stays specific to its source domain, and what carries over is the optimizer harness learned by self\-evolution \(Section[G\.2](https://arxiv.org/html/2610.02616#A7.SS2)\)\. When the executor model changes, executor harnesses tuned to one model mostly stop helping \(Section[G\.3](https://arxiv.org/html/2610.02616#A7.SS3)\)\. Section[G\.1](https://arxiv.org/html/2610.02616#A7.SS1)uses the harnesses of Table[2](https://arxiv.org/html/2610.02616#S4.T2); the other two sections start from the harnesses of Table[6](https://arxiv.org/html/2610.02616#A6.T6), evolved with the small and mid\-size models\. Throughout, every score is compared to blank, the initial harnessh0h\_\{0\}, measured on the same test set with the same executor\.
### G\.1Transfer to new programming languages
Figure 7:OOD accuracy by programming language\.The validation\-selected harnesses of Table[2](https://arxiv.org/html/2610.02616#S4.T2)on the 107 OOD tasks, split by language \(mean over three test evaluations\)\. Each host has one color: the hollow marker is the baseline, the filled marker is the baseline with VERSE and self\-evolution, and the line joins the two\. All harnesses were evolved on Python tasks\.All training and validation tasks of Section[2\.2](https://arxiv.org/html/2610.02616#S2.SS2)are in Python, while 87 of the 107 OOD test tasks of Table[2](https://arxiv.org/html/2610.02616#S4.T2)are in Go, Java, Rust, or TypeScript\. Each OOD task runs in its own released container image, which provides the toolchain of its language, and is graded by its own test script with the SWE\-rebench log parser for that language\. Figure[7](https://arxiv.org/html/2610.02616#A7.F7)splits the OOD accuracy of Table[2](https://arxiv.org/html/2610.02616#S4.T2)by language\. Outside Python, the four baselines range from 1\.1 points below blank to 4\.2 points above it, while VERSE with self\-evolution is 3\.4 to 12\.6 points above blank\. VERSE raises its host in 16 of the 20 pairs of host and language, and in all five languages on the Meta\-Harness and HarnessX hosts\. Averaged over the four hosts, it raises accuracy in every language: by 10\.3 points in Go, 10\.3 in Java, 2\.3 in Rust, 4\.9 in TypeScript, and 4\.2 in Python\. Self\-evolution matters most on these two hosts: without it, the evolved harnesses stay near their baselines outside Python \(28\.7% and 26\.8%, against 29\.5% and 26\.1%\); with it, they reach 35\.6% and 37\.9%\. On the AHE and Self\-Harness hosts, the tools alone raise accuracy outside Python \(to 35\.2% and 30\.3%, from 24\.1% and 29\.1%\)\.
### G\.2Transfer to a new task domain
Table[9](https://arxiv.org/html/2610.02616#A7.T9)moves the evolved harnesses between SWE\-rebench and Terminal\-Bench and allows two rounds of evolution in the target domain\.
Table 9:Cross\-domain transfer \(SWE↔\\leftrightarrowTerminal\-Bench\)\.Each harness of Table[6](https://arxiv.org/html/2610.02616#A6.T6)is moved to the other domain, evaluated frozen \(zero\-shot\), and evolved for two more rounds there\. An evolved optimizer inherits its harness from the final round of the source run; a cold start begins ath0h\_\{0\}, so its zero\-shot entry evaluatesh0h\_\{0\}and is not bolded\. The VERSE rows start from the HarnessX\-hosted VERSE harness\. Avg@3 accuracy \(%,±\\pmSEM over three test evaluations\)\.Bold: higher mean than the reference, which is the baseline above for “\+\+VERSE” rows and the best baseline in the column for self\-evolving rows\.Unlike a change of language, a change of domain changes the task itself: a Terminal\-Bench task is graded by a program that checks the final state of its container, not by hidden tests applied to a repository diff\. Zero\-shot, a transplanted executor harness can fall below the target\-domain blank: it encodes its source domain, from tool conventions to the failures it fixes\. One or two rounds of adaptation bring it back to about the blank level\. Gains beyond blank come from the inherited optimizer harness\. From SWE\-rebench to Terminal\-Bench, VERSE with its inherited optimizer is above every frozen\-optimizer entry in all three columns, and with the executor harness restarted from blank, one self\-evolving configuration matches the best Terminal\-Bench score of Table[6](https://arxiv.org/html/2610.02616#A6.T6)within two rounds\. The reverse direction moves less, consistent with how flat Terminal\-Bench is\. What crosses domains appears to be the optimizer’s methodology, its skills, notes, and tools for diagnosing and verifying failures, while the executor harness stays specific to its source domain\.
### G\.3Transfer to a new executor model
We hand the evolved harnesses to two open\-weight executors: Qwen3\-Coder\-480B and DeepSeek V3\.2\. Each harness from Table[6](https://arxiv.org/html/2610.02616#A6.T6)\(evolved with the small model\) is re\-evaluated frozen \(zero\-shot\) and then adapted for two rounds with the new executor under the main protocol\.
Table 10:Transfer to a new executor on SWE\-rebench\.The harnesses of Table[6](https://arxiv.org/html/2610.02616#A6.T6), evolved with the small model, are handed to Qwen3\-Coder\-480B and DeepSeek V3\.2\. Each is evaluated frozen with the new executor \(zero\-shot\) and then evolved for two more rounds with it \(Rounds 1 and 2\)\. Evolved optimizers inherit their harness from the final round of the source run; a cold start restarts the executor harness ath0h\_\{0\}, so its zero\-shot entry evaluatesh0h\_\{0\}and is not bolded\. The VERSE rows start from the HarnessX\-hosted VERSE harness\. Avg@3 accuracy \(%,±\\pmSEM over three test evaluations\)\.Bold: higher mean than the reference, which is the baseline above for “\+\+VERSE” rows and the best baseline in the column for self\-evolving rows\.Zero\-shot transfer to a new executor model\.The zero\-shot columns of Table[10](https://arxiv.org/html/2610.02616#A7.T10)freeze the ten SWE harnesses of Table[6](https://arxiv.org/html/2610.02616#A6.T6)and swap the executor\. A harness evolved for one executor mostly stops helping when the executor changes: averaged over the ten harnesses, zero\-shot transfer is close to zero\. Two patterns remain\. The equipped version keeps its advantage over its baseline in three of four pairs on each executor, and transfer is larger for the executor that handles basic tool use worse under blank \(Qwen\)\. The clear exception is the HarnessX family on Qwen: both variants sit above blank, and HarnessX\+\+VERSE gives the largest zero\-shot gain in the table\.
To see why the HarnessX\+\+VERSE harness transfers, we compared all 108 Qwen test trajectories under the blank harness with the same tasks under this harness\. That harness is a minimal prompt plus an 18KB hook file with three components: astr\_replace\_editortool whose edit content travels as a tool call argument rather than through the shell, reminders that escalate at fixed turn counts if the agent has not yet edited a source file or run the tests, and fixes for two common environment errors\.
Under blank, Qwen never uses an editor tool and writes every patch through shell heredocs; in all 8 tasks that the transplant flips, the blank trajectory dies on heredoc quoting before the patch ever lands\. Under the transplant, all 108 trajectories use the editor, all 8 flipped tasks land their edits, and trajectories become shorter\. VERSE’s own evolved harness, which encodes behavioral fixes specific to the small model, does not transfer\. The contrast suggests a mechanism:*harness components that fix tool use errors common across models, such as shell quoting, transfer; components tuned to one model’s behavioral quirks do not*\.
Table 11:Transfer to a new executor on Terminal\-Bench\.The Terminal\-Bench harnesses of Table[6](https://arxiv.org/html/2610.02616#A6.T6), evolved with the small model, are handed to Qwen3\-Coder\-480B and DeepSeek V3\.2\. Each is evaluated frozen with the new executor \(zero\-shot\) and then evolved for two more rounds with it \(Rounds 1 and 2\)\. Evolved optimizers inherit their harness from the final round of the source run; a cold start restarts the executor harness ath0h\_\{0\}, so its zero\-shot entry evaluatesh0h\_\{0\}and is not bolded\. The VERSE rows start from the HarnessX\-hosted VERSE harness\. Avg@3 accuracy \(%,±\\pmSEM over three test evaluations\)\.Bold: higher mean than the reference, which is the baseline above for “\+\+VERSE” rows and the best baseline in the column for self\-evolving rows\.The same transfer on Terminal\-Bench\.Table[11](https://arxiv.org/html/2610.02616#A7.T11)repeats the executor swap for the Terminal\-Bench harnesses of Table[6](https://arxiv.org/html/2610.02616#A6.T6)\. Zero\-shot, most transplanted harnesses stay near the blank, and adaptation recovers the blank level but rarely exceeds it\. There is no Terminal\-Bench analogue of the HarnessX jump: the Terminal\-Bench executor issues one shell command per turn and never passes edits through shell quoting, so the friction that the transferred editor tool removes on SWE\-rebench does not arise\. The one consistent signal again favors the inherited optimizer harness: on Qwen, VERSE is the only method above the best baseline in all three columns, and its cold\-started row also exceeds the best baseline in both adaptation rounds\.
Warm\-started re\-evolution on a new executor\.The Round 1 and Round 2 columns of Table[10](https://arxiv.org/html/2610.02616#A7.T10)ask whether one evolution run can be reused across executor models\. In our runs, the best warm\-started configuration outscores the best cold\-started one on each executor after two rounds\. Adaptation helps most when the starting harness is weak on the new executor and little when it is already strong, as for HarnessX\+\+VERSE on Qwen\. This suggests that evolving once and then adapting for a round or two can be a practical way to move to a new executor\.
## Appendix HAdditional ablations
This appendix uses the small and mid\-size models of Appendix[F](https://arxiv.org/html/2610.02616#A6)on SWE\-rebench\. Table[12](https://arxiv.org/html/2610.02616#A8.T12)collects three ablations\. The first repeats the component ablation of Section[4\.3](https://arxiv.org/html/2610.02616#S4.SS3)with these models, again on AHE without self\-evolution, the host that gains most from the tools in Table[6](https://arxiv.org/html/2610.02616#A6.T6)\. The second keeps this configuration and its probe budget, but the optimizer no longer chooses the probes\. The third uses the setting of Observation 1, the self\-evolving optimizer with simple verification on Meta\-Harness, and restricts what the optimizer may write for itself\. Figure[8](https://arxiv.org/html/2610.02616#A8.F8)varies the probe budget\. Each configuration in this appendix is a single evolution run, and Table[8](https://arxiv.org/html/2610.02616#A6.T8)shows that runs of one method can differ by several points, so we read these ablations for their trends\.
Table 12:Ablations with the small and mid\-size models\(SWE\-rebench; avg@3 accuracy, %,±\\pmSEM over three test evaluations; one evolution run per row\)\. Top and middle: AHE\+\+VERSE without self\-evolution, with components removed \(top\) or with the probes scheduled by a rule instead of the optimizer \(middle\); both change the full configuration in the first row\. Bottom: the self\-evolving optimizer with simple verification on Meta\-Harness, as in Observation 1, restricted in what it may write for itself; the first row is an independent run without the restriction\.r⋆r^\{\\star\}: the round selected on validation;*Val\-selected*: the harness of that round;*Last*: the final\-round harness\.Figure 8:The probe budget ablation\.Avg@3 test accuracy \(%,±\\pmSEM over three test evaluations; one evolution run per budget\) of AHE\+\+VERSE \(w/o self\-evolve\) as the verification budget per candidate varies; the three candidates of a round share a pool of three times this number of units, and the full configuration allows 8 per candidate \(24 per round\)\. The dashed line is AHE without any verification tools\.The training audit matters most\.Removing attribution and minimization, which also removes the training audit, costs about half of VERSE’s margin over plain AHE\. Removing only the training audit, with attribution and minimization still running, costs more, and its final round is the lowest in the table\. Attributing individual failures without tracking them across rounds may pull the optimizer toward causes it would otherwise deprioritize\.
The optimizer choosing its own probes is part of the gain\.At the same probe budget, a random schedule costs accuracy, and a fixed hand\-written schedule falls below plain AHE at the selected round: probes the optimizer did not ask for can be worse than none\. This suggests that the value of verification lies in the optimizer’s own choice of experiments, not in the extra computation alone\.
No single kind of self\-written equipment replaces the full workspace\.Restricting the self\-evolving optimizer to only its prompts, only its notes and skills, or only its tools lowers selected\-round accuracy by about three points in each case\.
The gain needs probe budget\.With a budget of one or two units per candidate, most of the gain disappears; from four units on, most of it returns \(Figure[8](https://arxiv.org/html/2610.02616#A8.F8)\)\. This is consistent with Appendix[C\.2\.2](https://arxiv.org/html/2610.02616#A3.SS2.SSS2), where the bound on the repair error falls geometrically with the number of checks, then levels off\.
## Appendix IReplay study details
Each verification tool of Section[3\.1](https://arxiv.org/html/2610.02616#S3.SS1)rests on an empirical premise: attribution and minimization assume that failures compress to a few load\-bearing steps, theperturbationtool assumes that single steps carry the outcome, and theverificationtool assumes that repairing the diagnosed step can flip it\. We test each premise by re\-executing recorded traces \(Figure[9](https://arxiv.org/html/2610.02616#A9.F9)\)\. The measurements below support these premises, and a fourth finding shapes the design: reading a trace identifies the kind of failure far more reliably than the failing step, so localization needs execution\.
### I\.1Corpora and replay
The study uses two corpora\. The main corpus holds roughly 300 SWE\-rebench task instances\([Badertdinov et al\., 2025](https://arxiv.org/html/2610.02616#bib.bib3)\), each with a successful and a failed trajectory of an OpenHands agent driven by Qwen3\-Coder\-480B under one fixed harness\. A step is one message of a trajectory: the task prompt, a model turn, or a tool result\. Minimization and repair use 299 successful and 298 failed traces; the perturbation experiment injects faults into 297 successful traces\. A closed\-weight LLM supervisor, a different model from the executor, produces every diagnosis, and every claim is checked by re\-executing the trace in the task’s Docker environment against the task’s own tests\.
The second corpus is CodeTraceBench\([Li et al\., 2026](https://arxiv.org/html/2610.02616#bib.bib22)\), a public collection of executed trajectories from four agent frameworks \(OpenHands, mini\-SWE\-agent, SWE\-agent, and Terminus 2\) whose failed runs carry human step\-level annotations of the steps on the causal chain to the failure\. We use its verified split; the 208 failed trajectories for which a human label and a supervisor diagnosis both exist enter the attribution comparison below\.
How replay works\.Each replay starts from a fresh container of the task at its original commit\. The recorded actions of the kept steps, that is, the agent’s commands and file edits, are executed again in order, so every tool output comes from the environment rather than from the recording\. After the last kept action, the agent continues on its own until it stops, and its final patch is graded with the task’s own tests\. The replays inside VERSE stop after the recorded commands and call no model\.
A reduced subset is accepted only if its replay fails the same tests as the full trace, or, for a successful trace, still resolves the task; 230 of the 235 accepted subsets of failed traces keep at least one of the agent’s own edits, and in the other 5 the failure appears without any agent edit\. A repair counts only if a second replay also resolves the task\. Before a fault is injected into a solved trace, the unmodified trace is replayed once and must still resolve the task\. The supervisor that proposes subsets, diagnoses, and repairs reads the task, the recorded trajectory, and the failing tests, but never the reference patch\.
### I\.2Minimization and repair
Trace minimization searches for a reduced subset of steps that still reproduces the original outcome under replay \(Figure[9](https://arxiv.org/html/2610.02616#A9.F9), left\)\. The median failed trace reproduces its exact failure, with the same failing tests, using 8 of its 129 recorded steps \(median per\-trace compression14\.4×14\.4\\times; 78\.9% of 298 traces accept a reduced subset\), and the median successful trace still solves the task with 4 of its 127 recorded steps \(median per\-trace compression28\.2×28\.2\\times; 87\.6% of 299\); most subsets are accepted at the supervisor’s first proposal \(68\.8% and 80\.3%\)\.
Figure 9:The replay study: minimization, repair, and calibration\.Left: median steps of the original trace against its reduced reproducing subset, confirmed by replay against the task’s own tests \(299 successful and 298 failed traces\); the compression factor is the median of the per\-trace ratios, not the ratio of the two medians\. Center: fraction of the 298 failed traces that flip to success when the supervisor may apply up tokkminimal edits at the diagnosed steps\. Right: attribution confidence against actual accuracy on the injected\-fault corpus \(889 scored diagnoses\); expected calibration error 0\.162\.A worked example\.Figure[10](https://arxiv.org/html/2610.02616#A9.F10)shows one failed trace from the main corpus\. The full trajectory has 103 steps and fails 14 of the task’s tests\. The supervisor’s first candidate subset kept 8 steps and a second kept 12, but both replays reproduced only 5 of the 14 failing tests and were rejected: a reduced trace must reproduce the*same*failure, not just some failure\. The accepted subset also keeps 12 steps, including ten mid\-trajectory actions of which six are writes, and its replay fails exactly the same 14 tests, an8\.6×8\.6\\timescompression\.
Figure 10:A worked minimization example\.Each cell is one step of a 103\-step failed trace from the main corpus of the replay study; the 12 blue steps are the reduced subset whose replay fails exactly the same 14 tests as the full trace\.In the repair test \(Figure[9](https://arxiv.org/html/2610.02616#A9.F9), center\), a single\-step fix at the diagnosed step flips 50\.3% of 298 failures to success, and allowing up to five edits raises this to 67\.1%, with only 1\.65 edits applied on average\.
### I\.3The perturbation spectrum
We injected fifteen kinds of single\-step faults into 297 solved traces and replayed each perturbed trace, for 3,868 valid injections \(Figure[11](https://arxiv.org/html/2610.02616#A9.F11)\)\. Overall, 35\.7% flip the run to failure, and the spectrum is sharply structured: faults in tool calls, logic, and edits flip 44 to 50% of runs, faults in context flip 6\.7%, and individual operators range from 99\.2% \(duplicating a retry loop\) down to 0\.4% \(truncating context\)\. Two faults injected at once flip 81\.0% of runs, more than either fault alone\. Step\-level interventions measure a real, structured quantity, which theperturbationtool exploits\.
Figure 11:The perturbation spectrum\.Left: for eleven of the fifteen fault kinds, the fraction of solved runs that flip to failure when a single fault of that kind is injected and the trace is replayed \(3,868 valid injections of all fifteen kinds into 297 solved traces\), colored by the part of the harness the fault touches\. Right: three fault pairs, each fault alone against both injected together \(780 paired injections\): two faults flip more runs than either alone\.
### I\.4Attribution against ground truth
Table[13](https://arxiv.org/html/2610.02616#A9.T13)scores attribution on 1,354 traces where we injected a known fault, so the ground truth holds by construction: the supervisor lists the correct fault class among its candidates 98\.7% of the time, and its first choice is correct 62\.0% of the time, but it names the broken step exactly only 36\.9% of the time and within one step 57\.3% of the time\. These 1,354 traces are the injections of Figure[11](https://arxiv.org/html/2610.02616#A9.F11)that flipped a solved run to failure and received a diagnosis \(1,354 of 1,382\); the calibration panel of Figure[9](https://arxiv.org/html/2610.02616#A9.F9)uses a separate scoring run of 1,001 diagnoses, of which the 889 with a single injected fault enter the plot\.
On the 208 human\-annotated failed traces of CodeTraceBench, drawn from four different agent frameworks, its top step choice lands on a human\-flagged step 22\.6% of the time and within three steps of one 39\.9% of the time; a uniformly random step choice would land on a flagged step 9\.7% of the time\. Reading identifies what kind of failure occurred far more reliably than where, which is why VERSE gives the optimizer execution tools to test diagnoses and draft edits before submission; whether and how it uses them is part of the optimizer’s policy\. Attribution confidence is well ordered but overconfident in the middle bins \(Figure[9](https://arxiv.org/html/2610.02616#A9.F9), right\): the expected calibration error is 0\.162, and above confidence 0\.85 accuracy reaches 0\.95 at 45% coverage\.
Table 13:Attribution against constructed ground truth\(1,354 traces, each with one injected fault\)\.*Fault class found*: the harness layer of the injected fault is among the layers the supervisor lists \(2\.9 on average; its first choice is correct in 62\.0% of traces\)\.*Exact step*/*±1\\pm 1step*: the predicted step matches the injected one exactly or within one; a step is one message of the trace\. Rows list the six most frequent faults; the overall row covers all injections\. When a fault removes content \(premature finish, deleted edit\), the supervisor names the message right after the injection point \(in all 169 and in 75 of 90 traces\), which only the±1\\pm 1column counts; a duplicated retry spreads the diagnosis over the retry loop\.
## Appendix JAdditional results for Observation 1
### J\.1Observation 1 on Terminal\-Bench
Table[14](https://arxiv.org/html/2610.02616#A10.T14)repeats the comparison of Table[1](https://arxiv.org/html/2610.02616#S2.T1)on Terminal\-Bench with the small model as the executor and the mid\-size model as the optimizer \(Appendix[D\.2](https://arxiv.org/html/2610.02616#A4.SS2)\), over three evolution rounds\. The combination of self\-evolution and simple verification is again the strongest configuration in both views, and simple verification alone improves on the frozen optimizer\. Self\-evolution alone gives no gain: its validation\-selected harness scores the same as Meta\-Harness \(37\.08%37\.08\\%\), and its final harness is within one standard error of it\.
Table 14:Observation 1 on Terminal\-Bench with the small and mid\-size models\(avg@3 accuracy, %,±\\pmSEM over three test evaluations; one evolution run per row\)\.*Self\-evolving*: the optimizer edits its own harness;*Verified*: it can run a draft edit on a training task before submitting it\.r⋆r^\{\\star\}: the validation\-selected round;*Val\-selected*: the harness of that round;*Last*: the final\-round harness\.
### J\.2Self\-evolution with and without verification, round by round
Figure[12](https://arxiv.org/html/2610.02616#A10.F12)follows the two self\-evolving runs of Table[1](https://arxiv.org/html/2610.02616#S2.T1)round by round, reconstructed from the optimizer’s tool calls, the round logs, and the prompts, notes, and tools the optimizer wrote for itself\. The two runs use the same host, models, number of candidates, turn limit, and training tasks, and both read the trajectories and scores of the training tasks\. The only difference is that one run has simple verification, with a budget of eight calls per candidate \(Appendix[D\.2](https://arxiv.org/html/2610.02616#A4.SS2)\)\.
Without the tool\.In round 1, the three candidates each spent about 58 shell commands reading the 110 training trajectories\. The winner found three failure causes: empty patches, repeated commands, and fixes that the repository’s own tests never checked\. It submitted a 40 KB harness in eight files: hook code, a system prompt, a workflow, three skills, tool notes, and a memory file\.
One hook, meant to rescue runs that stop early, returned output that the hook runtime rejects\. The error appeared in 13 of the 50 validation runs and in 40 of the 110 training runs of the next round\. Validation accuracy fell from 20 to 19 of 50 tasks, and a rerun confirmed one broken task\. The optimizer found the error only in round 2, from the training logs\. In rounds 3 and 4, it added a cache that answered repeated commands with stale output and then widened the rescue hook; it ran neither on a task before submitting it, and reruns confirmed three more broken tasks\.
Before round 4, it wrote its own tool for checking drafts, which loads the hook code and feeds it made\-up model replies\. This tool could show that the code ran, but not whether the executor did better on real tasks\. From round 3 on, its notes told it to “stop paying for code layers”\. In round 5, it discarded its edits and returned to the initial harness plus one small file\-editing tool; in round 6, it added one hook\. After the last round, it wrote in its own prompt that harnesses differing by 9 to 30 KB of code “score 3 apart for reasons I cannot see”\. Validation accuracy never rose above its round\-0 value, so validation selected the initial harness itself \(r⋆=0r^\{\\star\}=0in Table[1](https://arxiv.org/html/2610.02616#S2.T1)\)\. The final\-round harness scored31\.79%31\.79\\%\.
With the tool\.In round 1, the candidates read the trajectories of the same 110 training tasks with about 61 shell commands each\. The winner then ran one training task under the unchanged harness and read which tests failed\. It drafted its hook code and ran another task with the draft: the hooks fired without error\. The submitted 24 KB harness raised no errors in the 50 validation runs\.
In later rounds, the optimizer wrote its own tools: one that classifies failed training runs, one that checks draft code, one that edits files, and a probe that runs the current and the edited harness on the same training tasks\. The three candidates called these tools 11 to 23 times per round\. The optimizer’s prompt came to state that reading transcripts gives only a hypothesis, and that only running the executor verifies it\. Validation accuracy rose from 17 to 20 of 50 tasks, and the final harness scored41\.05%41\.05\\%, the highest in Table[1](https://arxiv.org/html/2610.02616#S2.T1)\.
The difference\.Both optimizers read trajectories of the same training tasks and found similar failure causes in round 1\. What differed was whether they could test an edit before submitting it\. Without the tool, errors surfaced only after a round had been spent, and the optimizer could not tell which of its edits changed validation accuracy\. Its self\-evolution turned toward doing less, until it discarded its edits\. With the tool, the optimizer checked its drafts on real tasks, and its self\-evolution went into better ways of testing them\.
Without verification \(Table[1](https://arxiv.org/html/2610.02616#S2.T1), row 3\)Round 1 reads110 training trajectories, about 58 shell commands; finds empty patches, repeated commands, and untested fixesSubmitsa 40 KB harness in eight files: hook code plus prompt, workflow, skills, notes, and memoryA rescue hook is broken: the runtime rejects its output in 13 of 50 validation runs and 40 of 110 training runsValidation 20→\\to19; the error is found only in round 2, from the training logsRounds 3–4: a stale\-output cache and a wider rescue hook, neither run on a task; three more broken tasks confirmed by rerunRound 5: discards its edits, back toh0h\_\{0\}plus one small tool\. After the last round, its prompt: scores differ “for reasons I cannot see”\.r⋆=0r^\{\\star\}=0; final round31\.7931\.79With simple verification \(row 5\)Round 1 readstrajectories of the same 110 tasks, about 61 shell commands, thenruns one task unchangedand reads which tests failDraftsthe hook code andruns a second task with the draft applied: the hooks fire, no errorSubmitsa 24 KB harness; no errors in 50 validation runsRounds 2–6: writes tools to classify failures, check drafts, edit files, andrun the current and edited harness on the same tasks; 11–23 calls per roundIts prompt: reading transcripts gives only a hypothesis; only running the executor verifies itValidation 17→\\to20; test41\.0541\.05, the highest in Table[1](https://arxiv.org/html/2610.02616#S2.T1)Figure 12:The two self\-evolving runs of Table[1](https://arxiv.org/html/2610.02616#S2.T1), without and with simple verification, reconstructed from the optimizer’s tool calls and round logs\. Red: edits submitted without being run on a task, and their consequences; green: steps based on running the executor\.
## Appendix KAdditional results for Observation 2
The audit covers the five runs of Figure[3](https://arxiv.org/html/2610.02616#S3.F3), one per executor: Qwen3\.8\-27B \(the last row of Table[1](https://arxiv.org/html/2610.02616#S2.T1)\), Qwen3\-Coder\-480B, DeepSeek V3\.2, GLM\-5, and Mistral Large 3\. Each run evolves six rounds with the self\-evolving optimizer and simple verification\. The optimizer workspace has five editable slots: a prompt file added to the optimizer’s own instructions, a candidate\-guidance file whose blocks are handed to its parallel candidates, a notes file, a skills directory, and one Python file that may declare callable tools and automatic hooks on the optimizer’s own agent loop\. All five runs start with the three files blank, the skills directory empty, and no Python file\. We read the workspace after the final round\.
An artifact is one self\-written unit: a prompt file, a skill file, a callable tool declared in the Python file, an automatic hook, or the notes file\. An automatic hook is one check or reminder with its own trigger condition; a series of reminders keyed only to the turn counter counts once\. Each artifact receives one primary function: attribution \(locating and ranking failure causes\), verification \(checking a draft before submission\), training audit \(recording outcomes across rounds\), or workflow \(how the optimizer organizes its own process, including budget, turn allocation, and the division of work among its candidates\)\. Superseded versions, temporary scripts, executor harness files, and the pipeline’s own reports are not counted\. Table[15](https://arxiv.org/html/2610.02616#A11.T15)lists all 51 artifacts\.
Three of the five runs cover all four functions; the GLM\-5 run wrote no attribution artifact and the Mistral Large 3 run no verification artifact\. Every run kept a training audit and workflow artifacts, but the implementations differ\. Two runs \(DeepSeek V3\.2 and GLM\-5\) wrote no callable tool and placed their verification logic in submission\-time hooks instead; the Qwen3\.8\-27B run wrote the largest toolkit, including a probe tool that batches head\-versus\-edit comparisons over simple verification calls\. Every run rewrote both prompt files and kept a notes file, and the notes are where confirmed fixes and regressions, prediction outcomes, and process lessons accumulate\.
Presence does not establish correctness or benefit\. The Qwen3\.8\-27B run’s own notes record a triage tool that raised on every call in one round and a probe tool that misread transcripts until a later revision, and the Python file is loaded only if it passes the same safety check as the executor’s hooks and a test run, so a rejected version leaves the previous one in place\.
Figure[13](https://arxiv.org/html/2610.02616#A11.F13)plots the validation accuracy of these five runs across evolution rounds\. The best validation round exceeds round 0 on every executor, and the final round stays above round 0 on four of the five\.
Figure 13:Validation accuracy per evolution roundof the self\-evolving optimizer with simple verification, run from scratch on SWE\-rebench with the five executors of Figure[3](https://arxiv.org/html/2610.02616#S3.F3)\. Validation pool: 50 tasks per round; round 0 is the initial harness\.Table 15:Inventory of self\-written optimizer artifactsbehind Figure[3](https://arxiv.org/html/2610.02616#S3.F3)\. One row is one artifact; the identifier is the file in the optimizer workspace, the tool name declared inextra\_tools, or the hook point that implements an automatic check\. Each artifact carries one primary function\. Descriptions summarize the retained content and do not assert that it worked\.ExecutorTypeIdentifierFunctionContentQwen3\-Coder\-480BPromptprompt\.mdWorkflowEleven standing rules: read notes first, triage in bulk, validate before submitting, keep edits additive\.Promptsearch\.mdWorkflowThree focus blocks that assign a different target to each parallel candidate\.Skillskills/triage\_playbook\.mdAttributionCount failure signatures before opening transcripts; cross\-check notes for categories already fixed\.Skillskills/validate\_and\_ship\.mdVerificationRun 2–3 failing and 1–2 passing tasks with the draft before writing the submission\.Skillskills/hooks\_audit\_constraints\.mdVerificationThe safety\-check rules a hook file must pass; lint before submitting\.Skillskills/destructive\_mechanism\_caution\.mdVerificationHigher validation bar for mechanisms that mutate or revert the executor’s files\.Tooltriage\_scanAttributionCounts failure\-signature substrings across a trajectory directory in one call\.Toollint\_hooksVerificationRegex check of a draft hook file against the executor’s safety\-check rules\.Hookon\_turn\_endWorkflowTurn 8: read notes and triage; turn 20: spend probes; turn 30: lint before submitting\.Notesnotes\.mdTraining auditPer\-round outcomes, ranked failure categories, process lessons, and an open task queue\.DeepSeek V3\.2Promptprompt\.mdWorkflowStructural facts about the loop, the current baseline state, and standing rules for each round\.Promptsearch\.mdWorkflowThree focus blocks: repair last round’s regression, audit the full hook stack, target turn exhaustion\.Skillskills/investigation\.mdAttributionGrep recipes for failure classes; separates root causes from symptoms; checks cross\-round oscillation\.Skillskills/validation\.mdVerificationProbe protocol: failing targets, passing regression tasks, one unrelated repository, then calibrate predictions\.Hookbefore\_tool\(submit\)VerificationReturns the submission once if predicted fixes were never probed with the final edits\.Hookbefore\_tool\(submit\)VerificationReturns the submission once if probes covered fewer than three repositories\.Hookon\_turn\_endWorkflowReminds after six probe calls that the pool is shared across candidates\.Notesnotes\.mdTraining auditPer\-round fixed and regressed tasks, harness changes, and two structural findings about selection\.GLM\-5Promptprompt\.mdWorkflowTest a new regex\-based hook on synthetic strings before spending probes\.Promptsearch\.mdWorkflowSafe\-fallback, diversity, and alternate\-mechanism slots for the three candidates\.Skillskills/regression\_safety\.mdVerificationRisk tiers for edits, false\-positive probes that must reach the guarded branch, prediction rules\.Hookbefore\_tool\(submit\)VerificationFlags benchmark task identifiers inside submitted file content\.Hookbefore\_tool\(submit\)VerificationRequires probes with edits on at least two distinct tasks before the first submission attempt\.Hookbefore\_tool\(submit\)VerificationFlags predicted fixes that the optimizer’s own probes with these edits returned as failing\.Hookbefore\_tool\(submit\)VerificationAsks for evidence that a probe reached a new blocking hook before it ships\.Hookbefore\_tool\(submit\)VerificationSuggests shipping a revert alone rather than bundled with a new mechanism\.Hookon\_turn\_endWorkflowReserves remaining probes for passing\-task regression checks after six calls\.Notesnotes\.mdTraining auditPost\-mortems of rounds 5 and 6: unexercised branches, entangled reverts, screening ties\.Mistral Large 3Promptprompt\.mdWorkflowEight standing instructions: read memory, triage in bulk, budget probes, prefer narrow edits\.Promptsearch\.mdWorkflowCandidate A revives a validated but unshipped mechanism; B finds a new signature; C is free\.Skillskills/triage\.mdAttributionFind the highest\-leverage mechanical cause in the first turns; weight aggregate evidence\.Toolsurvey\_failuresAttributionBatch\-scans trajectory files for a list of regex signatures and reports counts with examples\.Hookbefore\_toolWorkflowParses a stringified edits argument into a list, or blocks the call with an explanation\.Hookon\_turn\_endWorkflowEvery four probe calls: commit to one confirmed hypothesis and reserve regression checks\.Hookon\_turn\_endWorkflowAfter 25 shell calls: switch to the batch scanning tool\.Notesnotes\.mdTraining auditPer\-round root causes and outcomes, process lessons, and a checklist for the next round\.Qwen3\.8\-27BPromptprompt\.mdWorkflowWhat is scored, how selection behaves, evidence rules, edit\-design rules, and a turn schedule\.Promptsearch\.mdWorkflowA shared floor for all candidates plus three focus blocks: shrink the current harness, two failure groups\.Skillskills/corpus\-mine\.mdAttributionTwo\-call corpus digest: bucket runs by verdict and audit which current\-harness checks fire on passes\.Skillskills/loss\-mode\-map\.mdAttributionMaps each failure bucket to the harness levers that can and cannot move it\.Skillskills/verifier\-verdicts\.mdAttributionReads the grader’s verdict taxonomy; tests that never ran and tests that failed need opposite fixes\.Skillskills/evidence\-protocol\.mdVerificationA/B before proposing, measure task noise with repeats, keep a passing regression task, falsify the edit\.Skillskills/harness\-anatomy\.mdWorkflowVerified facts about the executor loop, hook surface, knobs, audit rules, and scoring\.TooltriageAttributionFailure map by grader verdict with cheap probe targets; grep and bucket views on request\.ToolprobeVerificationRuns the current and the edited harness on the same tasks with repeats and reports per\-run telemetry\.ToollintVerificationRuns the pipeline’s real audit function over a draft hook file\.TooleditWorkflowAnchored patching of a draft with all\-or\-nothing anchors; stages the file for probing\.Hookon\_turn\_endWorkflowScheduled notes at turns 3, 8, 15, 21, 26, and 30\+: triage, probe, draft, A/B, assemble, submit\.Hookbefore\_tool / after\_toolWorkflowAnswers a third identical shell command from cache instead of rerunning it\.Hookon\_turn\_endWorkflowReports failures of the self\-written tools so their silence is not mistaken for evidence\.Notesnotes\.mdTraining auditRound outcomes with reproduced flips, selection behaviour, the harness’s own failure record\.
## Appendix LExamples of optimizer\-written artifacts
This appendix shows what the optimizer wrote for itself in our best run: VERSE with self\-evolution on the Meta\-Harness host, with Qwen3\.8\-Flash\-Next as the optimizer and Qwen3\.8\-27B as the executor \(Table[2](https://arxiv.org/html/2610.02616#S4.T2);42\.28%42\.28\\%in distribution and37\.69%37\.69\\%out of distribution\)\. Validation selected round 2\. We show the optimizer harness that produced this round, which the optimizer wrote in its self\-update after round 1\. Its notes, skills, and code were empty at the start; VERSE seeds only its prompt and candidate guidance with instructions for the verification tools, and Appendix[L\.3](https://arxiv.org/html/2610.02616#A12.SS3)shows only the lines the optimizer added to them\. Excerpts are verbatim, except that internal names are replaced by the terms used in this paper; omissions are marked\[\.\.\.\]\.
The logs of this run show three lessons behind its result\. First, the optimizer measured before it edited\. In round 1, it classified the 79 failed training runs: 41 ran tests with a Python interpreter that lacked the test framework, 28 tried to install packages offline, and 26 tried to reach the network\. Its round\-1 edit to the executor harness addressed these failure classes as a whole rather than individual tasks\. When the new harness ran on the training tasks, the three counts fell to 2, 0, and 0, and training accuracy rose from 31 to 37 of 110 tasks\. In its self\-update, the optimizer turned this analysis into a tool it could call in one step \(Appendix[L\.1](https://arxiv.org/html/2610.02616#A12.SS1)\)\.
Second, the optimizer learned from a regression it caused\. The round\-1 edit fixed three validation tasks but broke two that had passed\. In its self\-update, the optimizer traced this to its verification calls, which had tested only tasks it expected the edit to fix\. It added a rule to its prompt, skill, and notes: check every draft on two tasks it should fix and one task the current harness already passes \(Appendices[L\.3](https://arxiv.org/html/2610.02616#A12.SS3)and[L\.4](https://arxiv.org/html/2610.02616#A12.SS4)\)\. It also wrote a hook that blocks, once, a submission made without executed evidence \(Appendix[L\.2](https://arxiv.org/html/2610.02616#A12.SS2)\)\.
In round 2, all three verification calls included a passing task, compared with none of the 3 calls in round 1 and none of the 12 calls in the run with the same host and tools but without self\-evolution\. The round\-2 edit fixed both earlier regressions and four other validation tasks and broke one, raising validation accuracy from 17 to 22 of 50 tasks\. It also added exceptions to the round\-1 interpreter rewrite for projects that manage their own environments, a change the notes had listed as a lead\.
Third, once a failure class was gone, the optimizer moved to the next one\. After round 1, the largest remaining class was runs whose visible tests passed but whose hidden tests failed \(30 of 73 failures\)\. The notes list this as the next lead \(Appendix[L\.4](https://arxiv.org/html/2610.02616#A12.SS4)\), and the round\-2 edit made the executor check its patch against the requirements stated in the issue and run the wider test directory before submitting\.
### L\.1A tool for measuring failures
The optimizer wrotecorpus\_triagein its first self\-update, after spending 56 to 62 shell calls per candidate on this analysis in round 1\. The tool classifies all failed training runs in one call and lists tasks that the current harness passes, which the optimizer then uses as regression checks\.
defextra\_tools\(self\):
return\[\{
”name”:”corpus\_triage”,
”description”:\(”Runthestandardfailure\-modetriageoverthelatestround’s”
”trajectoriesinONEcall:pass/failcounts,failure\-classflags”
”\(never\_ran\_tests,no\_edit\_seen,base\_python\_no\_pytest,”
”pip\_try\_offline,green\_but\_graded\_fail,mangled\_nodeids,”
”net\_attempt\)withexampletaskids,andthepassing\-tasklistto”
”pickregressioncanariesfrom\.Replacesdozensofexploratory”
”bashcalls\.Optionalarg:dir\(pathtoatrajectoriesdir\)\.”\),
”input\_schema”:\{”type”:”object”,
”properties”:\{”dir”:\{”type”:”string”,
”description”:”trajectoriesdir\(default:latestround\)”\}\}\},
\}\]
defrun\_tool\(self,name,args,env,state\):
ifname\!=”corpus\_triage”:
return”unknowntool%s”%name
try:
d=\(\(argsor\{\}\)\.get\(”dir”\)or””\)\.strip\(\)\.replace\(”’”,””\)
argv=”\-%s”%difdelse”\-”
returnenv\.bash\(”python3%s<<’TRIAGEPY’\\n%s\\nTRIAGEPY\\n”%\(argv,\_TRIAGE\)\)
exceptExceptionase:
return”corpus\_triagefailed:%r\(fallbacktomanualbashtriage\)”%\(e,\)
ran=any\(re\.search\(r”pytest\|\-munittest”,c\)forcincmds\)
green=bool\(re\.search\(r”\\d\+passed”,t\)\)andnotre\.search\(r”\\d\+failed”,t\)
base=bool\(re\.search\(r”Nomodulenamed\.?pytest”,t\)\)
pip=any\(re\.search\(r”\\b\(pip\|uv\|conda\)\\b\.\*install”,c\)forcincmds\)
edit=any\(re\.search\(r”sed\-i\|write\_text\|gitapply\|patch\-p\|\\\.write\\\(”,c\)forcincmds\)
rows\.append\(\(tid,ok\)\)
ifnotok:
fork,vin\[\(”never\_ran\_tests”,notran\),\(”no\_edit\_seen”,notedit\),
\(”base\_python\_no\_pytest”,base\),\(”pip\_try\_offline”,pip\),
\(”green\_but\_graded\_fail”,green\),\(”mangled\_nodeids”,”ERROR:notfound”int\),
\(”net\_attempt”,”NewConnectionError”int\)\]:
ifv:flag\[k\]\.append\(tid\)
Running the same script on the training runs that enter round 2 gives the following output\.
triageovertrajectories:110tasks,37pass,73fail
green\_but\_graded\_fail30e\.g\.\[’CQCL\_\_guppylang\-1259’,’Spenhouet\_\_confluence\-markdown\-exporter\-92’,\[…\]\]
mangled\_nodeids15e\.g\.\[’All\-Hands\-AI\_\_OpenHands\-10628’,’davep\_\_braindrop\-96’,\[…\]\]
no\_edit\_seen4e\.g\.\[’avaiga\_\_taipy\-2797’,’bvanelli\_\_actualpy\-117’,\[…\]\]
never\_ran\_tests3e\.g\.\[’avaiga\_\_taipy\-2797’,’bvanelli\_\_actualpy\-117’,\[…\]\]
base\_python\_no\_pytest2e\.g\.\[’Spenhouet\_\_confluence\-markdown\-exporter\-92’,’sissbruecker\_\_linkding\-1291’,\[…\]\]
passingset\(37\):pickregressioncanariesfromthese
ASPP\_\_pelita\-863,DiamondLightSource\_\_mx\-bluesky\-1519,PlasmaControl\_\_DESC\-1901,PyPSA\_\_linopy\-438,\[…\]
The optimizer also summarized what went wrong in round 1 at the top of its code\. Its count of about 110 shell calls per session is its own: round 1 was interrupted and restarted, and the log it read held both attempts, each with 53 to 62 shell calls per candidate\.
”””GovernorhooksforMYoptimizationloop\(authoredafterround1\)\.
Measuredfailurepatternsofmyownoptimizationsessions\(round1evidence\):
\*~110bashcorpus\-miningcallspercandidatesessionvs4\-6verificationcalls:the
scarceresources\(turns,sharedexecutionpool\)werespentonexploration
thatasinglereusabletriagescriptdoesinonecall\.
\*Theacceptedchange\-set’stransforminghooksfixed3tasksbutREGRESSED2
previously\-passingones:probesonlytestedpredicted\-fixdirections,never
a”staysgreen”canaryfromthecurrently\-passingset\.
\*4of9observedflipsdidnotreproduce:single\-runflipevidenceis~coin
flip;conclusionsshouldweightconfirmed\(twice\-run\)flips\.
\*Everysessionstillsubmitted\(good\)\-keeptheanti\-deadlockreleases\.
Design:everyhookisfail\-open\(identityonanyexception\),blocksarecapped
andself\-releasing,andtheonlyadditivetoolrunsapure\-computationtriage
scriptovertherundirectory\.
”””
### L\.2Hooks on the optimizer’s own loop
Before each tool call, one hook caps the verification calls of each candidate session so that parallel sessions share the pool, caps exploratory shell calls, and blocks a submission once if no draft has been verified yet\. Two further hooks inject reminders at fixed turns and append rules to the optimizer’s system prompt, each citing the round\-1 measurement behind it\.
BASH\_SOFT=40\#bashminingcallsallowedbeforeguidancekicksin
BASH\_HARD=70\#absolutecaponcedrafting/probingisunderway
POOL\_CAP=4\#verification\+simple\_verificationcallspersession\(poolsharedacrosssessions\)
defbefore\_tool\(self,name,args,state\):
try:
ifnamein\(”verification”,”simple\_verification”\):
used=state\.get\(”pool”,0\)
ifused\>=POOL\_CAP:
return\(”executionpoolreserved”,
”BLOCKED:thissessionalreadyusedits%dshared\-executionslots”
”\(otherparallelsessionsneedthem\)\.Reviseyourdraftusing”
”evidenceyoualreadyhave,thensubmit\.”%POOL\_CAP\)
state\[”pool”\]=used\+1
returnargs
ifname==”bash”:
\[…\]\#capscorpus\-miningbashcalls\(40soft/70hard\),anti\-deadlockrelease
ifname==”submit\_proposal”:
edits=\[\]
try:
edits=\(argsor\{\}\)\.get\(”edits”\)or\[\]
exceptException:
returnargs
ifeditsandnotstate\.get\(”pool”\)andnotstate\.get\(”warned”\):
state\[”warned”\]=True
return\(”probethedraftfirst”,
”BLOCKED\(once\):zeroexecutedevidenceyetthissession\.Draft”
”thesesamebytesin/scratchandverify2predicted\-fixtargets”
”\+1passingcanary;aprobereportingyourhookcallcountsbeats”
”anyamountofreading\.Thenresubmit\(thisblockfiresonlyonce\)\.”\)
returnargs
\[…\]
defon\_turn\_end\(self,turn,state\):
\[…\]
try:
ift==20andnotstate\.get\(”draft”\):
return\(”Milestonet20:stopbroadmining\.Writeyourtop\-2rootcausesand”
”drafteditfilesto/scratchnow\(corpus\_triagecoversthestats\)\.”\)
ift==32andnotstate\.get\(”pool”\):
return\(”Milestonet32:noexecutedevidenceyet\.Verifyyourcurrentdraft”
”on2fix\-targets\+1canaryNOW;poolleftthissession:%d\.”
%max\(0,POOL\_CAP\-state\.get\(”pool”,0\)\)\)
ift==45:
return\(”Milestonet45:finalprobe\(samebytesassubmission\)thenSUBMIT”
”withinafewturns\.Unsubmitted=0forthewholesession\.”\)
ift==52:
return\(”SUBMITNOWwithyourbestcurrentchange\-set\-budgetisalmostout\.”
”Amediocresubmittedchange\-setbeatsaperfectunsubmittedone\.”\)
defsystem\_prompt\(self,assembled\):
try:
return\(assembledor””\)\+\(
”\\n\\nBUDGETDISCIPLINE\(measuredround1:~110miningbashcallspersession”
”vs4probes\-backwards\)\.\(1\)Usethecorpus\_triagetool\(ortheonetriage”
”script\)FIRST;thentargetedgrepsonly\.Committoyourtop2rootcausesby”
”turn~25anddrafteditsin/scratch\.\(2\)Everyprobebatchmustincludeat”
”leastoneREGRESSIONCANARY:atasktheCURRENTharnessalreadypasses\-”
”round\-1transforminghooksfixed3andregressed2becausenocanarywas”
”everprobed\.\(3\)Asingle\-runflipis~50%noise:countonlytwice\-confirmed”
”flips;preferkillingawholefailureclassoverchasingtaskluck\.”
”\(4\)Reservethelast~8turns:finalprobeoftheexactbytesyousubmit,”
”thensubmit\.Unsubmittedsession=zero\.”\)
### L\.3A skill and prompt rules
The optimizer wrote one skill in its first self\-update, on writing executor hooks safely\. Rule 5 is the regression check it derived from round 1: test each draft on failing tasks and on one task that already passes\. It also added the following lines to its seeded prompt\.
\#Skill:writingexecutorhooksthatdon’tblowuptherun
Theexecutor’sharness\_code/hooks\.pyrunsonEVERYrunofEVERYtask\.Onebad
transformisarun\-wideregression;oneraisepercallisreportedinyouraudit\.
\#\#Contractfacts\(verifyinround\_NN/executor\_source/\{executor\.py,hooks\_runtime\.py\}\)
\-Arunendswhenareplyhasnotool\_useblockandstop\_reason=end\_turn;thescore
istheaccumulated‘gitdiff‘\.after\_llmcanappendatool\_useblocktokeepthe
runalive–appendONLY‘text‘/‘tool\_use‘typedblocks\(a‘thinking‘blockin
thereturnedlistmadetheAPIvalidatorrejectthereplyinarealincident\)\.
\-before\_toolreturninga\(reason,synthetic\_result\)tupleblocksthecommandand
injectsthesyntheticobservation:zerocontainercost,zerowastedturn\.
\-after\_toolseestherawobservationbeforethemodel;themodel’sowncommandsare
visibleinbefore\_toolargs\.Statedictpersistsperrun\.
\#\#Safetyrules
1\.Failopen:wrapeverymethodbodyintry/exceptreturningtheinputunchanged;
ahookthatraisesdegradestoidentityANDprintsanerrorinyourreportcard\.
2\.Actontight,anchoredpatterns,neverbroadkeywords;bailout\(returninput\)
themomentanythingunexpectedappears\(quotedheredocs,pipes,envvars,
explicitalternativeinterpreterslike‘uv‘/‘poetry‘/venvpaths–an
interpreter\-rewritethatignoresthesesankpreviously\-passingtasksonce\)\.
3\.Preferannotatinganobservationtorewritingacommand;preferblockinga
never\-can\-workaction\(offlinefetch\)totransformingamaybe\-validone\.
4\.One\-shotmessages:gatethegate\.Asubmit\-gatemustfireonce,notloopthe
run;akeep\-alivemustcountitsrescues\(cap2\-3\)\.
5\.Proveitwithverificationon2failingtargets\+1passingcanary\.Readtheprobe’s
per\-hookcall/changedcountsandtracebacks–called112x/changed0xmeansyour
conditionnevermatched\(deadcode\),notsuccess\.
6\.Perturbbeforeyoucodify:onlydistillabehaviorintoahook/skillwhen
perturbationshowsitisload\-bearinginpassingtrajectories\.
0\.Yournotes\.mdalreadysummarizesthecurrentharnessandlastround’smeasuredfailureclasses–startthere,thenrunthe‘corpus\_triage‘tool\(onecall:pass/failcounts,failure\-classflags,passing\-tasklistforcanaries\)\.Dotargetedgrepsonlyafterthat\.Committoyourtop\-2rootcausesbyturn~25andstartdraftingeditfilesin/scratch;afullre\-derivationofthestatswastesthehalfofthesessionwhereprobeslive\.
1\.\[…\]BEFOREsubmitting,draftyoureditsandverifythemon2tasksyoupredicttheyfixPLUS1REGRESSIONCANARY–ataskthecurrentharnessalreadyPASSES\.Historicalpredicted\-fixhitrateswithouttestingare0\-16%;andaround\-1proposalthatprobedonlyfix\-directionsfixed3tasksandREGRESSED2previously\-passingones\.\[…\]
2\.Flipevidenceisnoisy:~halfofsingle\-runflipsfailtoreproduceonrerun\(r1:5/9\)\.Onetwice\-confirmedflipbeatsthreesingle\-runflips;prefermechanismsthatkillaWHOLEfailureclass\(theysurvivethenoise\)overclevernessaimedatonetask\.
\#\#Sessionpacing\(self\-enforced;thegovernorhookwillnagyou\)
\-t<=20:triage\+targetedgreps;writeroot\-causehypotheses\+intendededitsto/scratch\.
\-t20\-35:draftalleditfiles;probe\#1\(2fixtargets\+1canary\)\.
\-t35\-45:revisefromprobeoutput\(readthehookcallcounts/tracebacks,notjustPASS/FAIL\);probe\#2withfinalbytes\.
\-t45\+:submitbyte\-identicaltotheprobedbytes;ifnothingisready,submityourbestunprobedchange\-setanyway\(betterthanzero\)\.
\-Sharedpool:atmost4verification/simple\_verificationcallspersession–parallelsessionsandthefinalistsessionneedtherest\.
### L\.4Notes
After round 1, the notes record the optimizer’s scores, its confirmed fixes and regressions, the failure classes it measured on the training runs, the rules it derived, and leads for the next round\. Two sections, which describe the current harness and the executor’s loop, are omitted\. The notes are the optimizer’s own account, and two of their statements differ from the logs\. The training accuracy listed for the round\-1 harness, 31 of 110, is that of the initial harness; the round\-1 harness solved 37 of 110 when it ran in round 2\. The notes also mention a second session for the winning candidate; the protocol runs none, and the only other session of a round is the self\-update\.
\#Optimizermemory\(cross\-round\)–condensed,keepupdatedeveryself\-update
\#\#Scoreboard
\-r0harness:val16/50,train31/110\.
\-r1\(kept,MIXED\):env\-normalising\+patch\-protectinghooks\.py\+promptset\-\>val17/50,train31/110\.Net\+3fixed/\-2regressed;only5/9flipsreproducedonrerun\(confirmedfixed:2oftheclaimed;thelitellm\-styleREGRESSIONconfirmedreal\)\.LESSON:single\-runflips~coinflips;transforminghookscarryrealregressionrisk\.
\-r1screeningmechanics:3candidatesessionswiththe3candidate\-guidancelenses;each~55\-60turns;winnergetsasecond\(final\)session;executionpool\(verification\+simple\_verification\)sharedround\-wide\(~24slots;11used\)\.Winnerr1=lens1\(OMISSION\)\.
\[…\]
\#\#Corpusfacts\(r1triage,110tasks,usecorpus\_triagetoolforfreshnumbers\)
\-31pass/79fail\.Flags:base\_python\_no\_pytest~41\-50,pip\_try\_offline28,net\_attempt26,green\_but\_graded\_fail18,never\_ran\_tests17,mangled\_nodeids16,no\_edit\_seen14\.
\-Classes:FAILED\-with\-a\-patch~40\(CORRECTNESSgap–biggestuntouched,nextfrontier\),NOT\_RUN~22\(severalUNWINNABLE:theverifier’sownpytestdies–offlineuv,git\-applyconflict,parametrizedidtruncatedinevalscript;don’tchase\),NO\_PATCH13,REGRESSION3\.
\-16failureshadfullygreenlastpytestrun\-\>gradedbyhiddentestsmatchingissuewording\(names/messages/edges\)\.”Existingtestspass”isnotthebar\.
\-Passingruns\(mean~29turns\):grep\-locate\-\>sed/pythonedit\-\>repo\-envpytestontouchedfile\-\>wholedir\-\>gitdiff\-\>submit\.0/31skippedtests;17/79failsneverrantests\.
\[…\]
\#\#Methodrules\(hard\-won–obey\)
1\.VerifytheDRAFTBYTESbeforesubmit:2predicted\-fixtargets\+\>=1REGRESSIONCANARY\(taskthecurrentharnesspasses\)\.Submitbyte\-identicaltotheprobedset\.r1’sbiggestmiss:neverprobedacanary\-\>\-2regressions\.
2\.Arefutedprobedowngradesspecifictaskclaims,notthemechanism\(empty\-diff\-\>real\-patch=mechanismverifiedevenat0flips\)\.
3\.Max2probesperhypothesis,thenmoveon;neverburnthesharedpool\(\>4/sessionstarvesthefinalsession\)\.
4\.Neverendwithoutsubmit\_proposal\.Anunsubmittedsession=0\.
5\.~50%ofsingle\-runflipsarenoise;thereportcardrerunsflips–buildclaimsthatsurvivetwice\.
6\.\[…\]minimize/perturbonFAILINGtracesrarelyflips–perturbPASSINGtracestofindload\-bearingbehaviorsanddistillthemintoskills\.
7\.Don’traiseobs\_cap/max\_turnsforfree–checkcontextpressurefirst\(r1leftloopmostlydefault\)\.
\#\#Openleadsfornextround
\-CORRECTNESS\(40FAILED\-with\-patch\):helptheexecutorholdthediffagainsttheissue’sassertions–e\.g\.anextratoolthatdiffsrequirementsvspatch\(meteredself\-reviewviaenv\.llm\),orarun\_teststoolthatalsoprintsthetouchedmodule’sfulltestresultscompactly\.Verifywiththeverificationtool\(correctnessflipsarenoisy–probetwiceifpoolallows\)\.
\-Regression\-proofingthecurrentharness:guardtransforms\(skiprewritewhenthecommandmentionsanexplicitdifferentinterpreter/venv/uv/poetry;skipgateifapriorsubmitwasblockedalready\)\.
\-NO\_PATCHsurvivors:keep\-aliveexists;checkwhethertherescuecommandsactuallyrun\(after\_llmchanged0/112callsinprobes–rescueFIRESbuttheappendedcommandmaynotbeexecuted?investigatebeforeaddingmore\)\.
\-Verifier\-NOT\_RUNtasksaredeadbudget:identifyfromverifieroutputandblacklistmentally\(theyrotate;re\-identifyeachroundfromfreshoutput,notfromstaleids\)\.相似文章
VeriHarness:扩展长时程任务的智能体验证能力
VeriHarness 将 LLM 生成器转变为智能体式验证器——为其配备工作区、证据工具、可复用的验证技能、分歧消解器和共识挑战器——从而在没有参考答案的情况下,为长时程任务筛选可靠输出。在五个基准测试中,它取得了最高的选择得分,使用 Gemini 3.5 Flash 和 Claude Opus 4.8 分别比单次采样提升 6.2 至 6.4 个点,并发布了约 26,000 次采样结果,成本超过 100,000 美元。
SBCO:面向规划智能体的自监督、基于验证器的框架优化
介绍了SBCO,一种面向规划智能体的自监督、基于验证器的框架优化器,它通过近似块坐标上升法改进智能体输出,以远少于自修改基线的计算量达到或超越其性能。
基于多任务的自进化 harness:agent 作为自身的优化器
提出一种自进化 agent harness 框架:同一个冻结模型先作为 solver 求解任务,再作为 proposer 直接编辑自身 harness 的代码;在多任务上完成进化后,于分布外基准上显著超越 Codex(提升 12.64 分)并达到相当的表现水平。
来自 Meta 的一篇关于智能体框架(agent harness)优化的优秀论文。Meta-Harness 式的搜索方法只使用一个开发集和一个提议……
# Meta(联合 Duke 与 UC Davis)提出用于智能体 harness 优化的「自改进分支混合」框架 Meta 联合杜克大学(Duke)与加州大学戴维斯分校(UC Davis)提出了一种 **Mixture of Self-Improving Branches(自改进分支混合)** 框架,用于优化智能体 harness(运行框架)。该方法将原本单轨迹的 **Meta-Harness** 搜索拆分为多个**自适应分支**:每个分支拥有不断演进的开发子集(development subsets)和提案策略(proposal policies),并引入一个**路由器(router)**,针对每个输入动态选择表现最佳的分支。 该框架在多个基准上取得了显著提升: - **Olympiad 级别数学任务**:相对提升最高可达 **+34.8%** - **Terminal-Bench 2.0**:提升 **+11.6%** - **SWE-bench Lite**:提升 **+3.8%**
@dair_ai: 关于自进化代理框架的精彩论文。自进化代理框架有两个实际问题:1. 搜索很慢…
论文提出了 Ecdysis,一个用于训练LLM代理运行时框架的框架,它通过识别重复的失败模式来提高效率和准确性,实现了1.84倍更快的训练和18.56%更高的推理准确性。