Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Summary
This paper introduces the Mendel Gödel Machine, a recursive self-improving framework that applies comparative evolution to iteratively improve coding agents.
View Cached Full Text
Cached at: 08/11/26, 08:02 AM
# Recursive Self-Improving Coding Agents via Comparative Evolution
Source: [https://arxiv.org/html/2608.07645](https://arxiv.org/html/2608.07645)
Changzhi Liu∗,§,\\ast,\\S,\\resizebox\{\}\{6\.83331pt\}\{\\par\\par\\par \\hbox to1580\.81pt\{\\vbox to1185\.6pt\{\\pgfpicture\\makeatletter\\hbox\{\\thinspace\\lower\-1492\.00569pt\\hbox to0\.0pt\{\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\definecolor\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@stroke\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@color@rgb@fill\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@setlinewidth\{\\the\\pgflinewidth\}\\pgfsys@invoke\{ \}\\nullfont\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\pgfsys@invoke\{ \}\\pgfsys@endscope\\hbox to0\.0pt\{\\pgfsys@beginscope\\pgfsys@invoke\{ \}\{\}\{\}\{\{\}\}\{\{\}\} \\pgfsys@beginscope\\pgfsys@invoke\{ \}\{\{\}\}\\pgfsys@eorulefalse\\pgfsys@invoke\{ \} \{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\definecolor\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@stroke\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\definecolor\{pgffillcolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@fill\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@setlinewidth\{\\the\\pgflinewidth\}\\pgfsys@invoke\{ \}\{\}\\pgfsys@moveto\{1796\.80685pt\}\{\-1491\.20569pt\}\\pgfsys@lineto\{217\.60083pt\}\{\-1491\.20569pt\}\\pgfsys@lineto\{217\.60083pt\}\{\-307\.20117pt\}\\pgfsys@lineto\{1796\.80685pt\}\{\-307\.20117pt\}\\pgfsys@lineto\{1796\.80685pt\}\{\-1491\.20569pt\}\\pgfsys@closepath\\pgfsys@moveto\{1756\.8067pt\}\{\-1424\.00543pt\}\\pgfsys@lineto\{1756\.8067pt\}\{\-374\.40143pt\}\\pgfsys@lineto\{1232\.8047pt\}\{\-899\.20343pt\}\\pgfsys@lineto\{1756\.8067pt\}\{\-1424\.00543pt\}\\pgfsys@closepath\\pgfsys@moveto\{1729\.6066pt\}\{\-346\.40132pt\}\\pgfsys@lineto\{284\.80109pt\}\{\-346\.40132pt\}\\pgfsys@lineto\{979\.20374pt\}\{\-1040\.80397pt\}\\pgfsys@curveto\{986\.6704pt\}\{\-1048\.27065pt\}\{996\.0038pt\}\{\-1052\.00401pt\}\{1007\.20384pt\}\{\-1052\.00401pt\}\\pgfsys@curveto\{1018\.40388pt\}\{\-1052\.00401pt\}\{1027\.73727pt\}\{\-1048\.27065pt\}\{1035\.20395pt\}\{\-1040\.80397pt\}\\pgfsys@lineto\{1729\.6066pt\}\{\-346\.40132pt\}\\pgfsys@closepath\\pgfsys@moveto\{1728\.8066pt\}\{\-1452\.00554pt\}\\pgfsys@lineto\{1204\.8046pt\}\{\-927\.20354pt\}\\pgfsys@lineto\{1063\.20406pt\}\{\-1068\.00407pt\}\\pgfsys@curveto\{1047\.204pt\}\{\-1084\.00414pt\}\{1028\.53728pt\}\{\-1092\.00417pt\}\{1007\.20384pt\}\{\-1092\.00417pt\}\\pgfsys@curveto\{985\.33711pt\}\{\-1092\.00417pt\}\{966\.40369pt\}\{\-1084\.00414pt\}\{950\.40363pt\}\{\-1068\.00407pt\}\\pgfsys@lineto\{809\.60309pt\}\{\-927\.20354pt\}\\pgfsys@lineto\{284\.80109pt\}\{\-1452\.00554pt\}\\pgfsys@lineto\{1728\.8066pt\}\{\-1452\.00554pt\}\\pgfsys@closepath\\pgfsys@moveto\{781\.60298pt\}\{\-899\.20343pt\}\\pgfsys@lineto\{256\.80098pt\}\{\-374\.40143pt\}\\pgfsys@lineto\{256\.80098pt\}\{\-1424\.00543pt\}\\pgfsys@lineto\{781\.60298pt\}\{\-899\.20343pt\}\\pgfsys@closepath\\pgfsys@fillstroke\\pgfsys@invoke\{ \} \\pgfsys@invoke\{ \}\\pgfsys@endscope \\pgfsys@invoke\{ \}\\pgfsys@endscope \\par \\pgfsys@invoke\{ \}\\pgfsys@endscope\{\}\{\}\{\}\\hss\}\\pgfsys@discardpath\\pgfsys@invoke\{ \}\\pgfsys@endscope\\hss\}\}\\endpgfpicture\}\} \\par\}Yilun Liu†,‡,§,\\dagger,\\ddagger,\\S,\\resizebox\{\}\{6\.83331pt\}\{\\par\\par\\par \\hbox to1580\.81pt\{\\vbox to1185\.6pt\{\\pgfpicture\\makeatletter\\hbox\{\\thinspace\\lower\-1492\.00569pt\\hbox to0\.0pt\{\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\definecolor\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@stroke\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@color@rgb@fill\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@setlinewidth\{\\the\\pgflinewidth\}\\pgfsys@invoke\{ \}\\nullfont\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\pgfsys@invoke\{ \}\\pgfsys@endscope\\hbox to0\.0pt\{\\pgfsys@beginscope\\pgfsys@invoke\{ \}\{\}\{\}\{\{\}\}\{\{\}\} \\pgfsys@beginscope\\pgfsys@invoke\{ \}\{\{\}\}\\pgfsys@eorulefalse\\pgfsys@invoke\{ \} \{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\definecolor\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@stroke\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\definecolor\{pgffillcolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@fill\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@setlinewidth\{\\the\\pgflinewidth\}\\pgfsys@invoke\{ \}\{\}\\pgfsys@moveto\{1796\.80685pt\}\{\-1491\.20569pt\}\\pgfsys@lineto\{217\.60083pt\}\{\-1491\.20569pt\}\\pgfsys@lineto\{217\.60083pt\}\{\-307\.20117pt\}\\pgfsys@lineto\{1796\.80685pt\}\{\-307\.20117pt\}\\pgfsys@lineto\{1796\.80685pt\}\{\-1491\.20569pt\}\\pgfsys@closepath\\pgfsys@moveto\{1756\.8067pt\}\{\-1424\.00543pt\}\\pgfsys@lineto\{1756\.8067pt\}\{\-374\.40143pt\}\\pgfsys@lineto\{1232\.8047pt\}\{\-899\.20343pt\}\\pgfsys@lineto\{1756\.8067pt\}\{\-1424\.00543pt\}\\pgfsys@closepath\\pgfsys@moveto\{1729\.6066pt\}\{\-346\.40132pt\}\\pgfsys@lineto\{284\.80109pt\}\{\-346\.40132pt\}\\pgfsys@lineto\{979\.20374pt\}\{\-1040\.80397pt\}\\pgfsys@curveto\{986\.6704pt\}\{\-1048\.27065pt\}\{996\.0038pt\}\{\-1052\.00401pt\}\{1007\.20384pt\}\{\-1052\.00401pt\}\\pgfsys@curveto\{1018\.40388pt\}\{\-1052\.00401pt\}\{1027\.73727pt\}\{\-1048\.27065pt\}\{1035\.20395pt\}\{\-1040\.80397pt\}\\pgfsys@lineto\{1729\.6066pt\}\{\-346\.40132pt\}\\pgfsys@closepath\\pgfsys@moveto\{1728\.8066pt\}\{\-1452\.00554pt\}\\pgfsys@lineto\{1204\.8046pt\}\{\-927\.20354pt\}\\pgfsys@lineto\{1063\.20406pt\}\{\-1068\.00407pt\}\\pgfsys@curveto\{1047\.204pt\}\{\-1084\.00414pt\}\{1028\.53728pt\}\{\-1092\.00417pt\}\{1007\.20384pt\}\{\-1092\.00417pt\}\\pgfsys@curveto\{985\.33711pt\}\{\-1092\.00417pt\}\{966\.40369pt\}\{\-1084\.00414pt\}\{950\.40363pt\}\{\-1068\.00407pt\}\\pgfsys@lineto\{809\.60309pt\}\{\-927\.20354pt\}\\pgfsys@lineto\{284\.80109pt\}\{\-1452\.00554pt\}\\pgfsys@lineto\{1728\.8066pt\}\{\-1452\.00554pt\}\\pgfsys@closepath\\pgfsys@moveto\{781\.60298pt\}\{\-899\.20343pt\}\\pgfsys@lineto\{256\.80098pt\}\{\-374\.40143pt\}\\pgfsys@lineto\{256\.80098pt\}\{\-1424\.00543pt\}\\pgfsys@lineto\{781\.60298pt\}\{\-899\.20343pt\}\\pgfsys@closepath\\pgfsys@fillstroke\\pgfsys@invoke\{ \} \\pgfsys@invoke\{ \}\\pgfsys@endscope \\pgfsys@invoke\{ \}\\pgfsys@endscope \\par \\pgfsys@invoke\{ \}\\pgfsys@endscope\{\}\{\}\{\}\\hss\}\\pgfsys@discardpath\\pgfsys@invoke\{ \}\\pgfsys@endscope\\hss\}\}\\endpgfpicture\}\} \\par\}Sikuan Yan†,‡\\dagger,\\ddaggerVolker Tresp†,‡\\dagger,\\ddaggerYunpu Ma†,‡,\\dagger,\\ddagger,\\resizebox\{\}\{6\.83331pt\}\{\\par\\par\\par \\hbox to1580\.81pt\{\\vbox to1185\.6pt\{\\pgfpicture\\makeatletter\\hbox\{\\thinspace\\lower\-1492\.00569pt\\hbox to0\.0pt\{\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\definecolor\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@stroke\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@color@rgb@fill\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@setlinewidth\{\\the\\pgflinewidth\}\\pgfsys@invoke\{ \}\\nullfont\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\pgfsys@invoke\{ \}\\pgfsys@endscope\\hbox to0\.0pt\{\\pgfsys@beginscope\\pgfsys@invoke\{ \}\{\}\{\}\{\{\}\}\{\{\}\} \\pgfsys@beginscope\\pgfsys@invoke\{ \}\{\{\}\}\\pgfsys@eorulefalse\\pgfsys@invoke\{ \} \{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\}\{\{\}\}\{\}\{\{\}\}\{\}\{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\{\}\{\{\}\}\{\} \{\}\{\} \{\}\{\} \{\}\{\} \{\}\\pgfsys@beginscope\\pgfsys@invoke\{ \}\\definecolor\{pgfstrokecolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@stroke\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\definecolor\{pgffillcolor\}\{rgb\}\{0,0,0\}\\pgfsys@color@rgb@fill\{0\}\{0\}\{0\}\\pgfsys@invoke\{ \}\\pgfsys@setlinewidth\{\\the\\pgflinewidth\}\\pgfsys@invoke\{ \}\{\}\\pgfsys@moveto\{1796\.80685pt\}\{\-1491\.20569pt\}\\pgfsys@lineto\{217\.60083pt\}\{\-1491\.20569pt\}\\pgfsys@lineto\{217\.60083pt\}\{\-307\.20117pt\}\\pgfsys@lineto\{1796\.80685pt\}\{\-307\.20117pt\}\\pgfsys@lineto\{1796\.80685pt\}\{\-1491\.20569pt\}\\pgfsys@closepath\\pgfsys@moveto\{1756\.8067pt\}\{\-1424\.00543pt\}\\pgfsys@lineto\{1756\.8067pt\}\{\-374\.40143pt\}\\pgfsys@lineto\{1232\.8047pt\}\{\-899\.20343pt\}\\pgfsys@lineto\{1756\.8067pt\}\{\-1424\.00543pt\}\\pgfsys@closepath\\pgfsys@moveto\{1729\.6066pt\}\{\-346\.40132pt\}\\pgfsys@lineto\{284\.80109pt\}\{\-346\.40132pt\}\\pgfsys@lineto\{979\.20374pt\}\{\-1040\.80397pt\}\\pgfsys@curveto\{986\.6704pt\}\{\-1048\.27065pt\}\{996\.0038pt\}\{\-1052\.00401pt\}\{1007\.20384pt\}\{\-1052\.00401pt\}\\pgfsys@curveto\{1018\.40388pt\}\{\-1052\.00401pt\}\{1027\.73727pt\}\{\-1048\.27065pt\}\{1035\.20395pt\}\{\-1040\.80397pt\}\\pgfsys@lineto\{1729\.6066pt\}\{\-346\.40132pt\}\\pgfsys@closepath\\pgfsys@moveto\{1728\.8066pt\}\{\-1452\.00554pt\}\\pgfsys@lineto\{1204\.8046pt\}\{\-927\.20354pt\}\\pgfsys@lineto\{1063\.20406pt\}\{\-1068\.00407pt\}\\pgfsys@curveto\{1047\.204pt\}\{\-1084\.00414pt\}\{1028\.53728pt\}\{\-1092\.00417pt\}\{1007\.20384pt\}\{\-1092\.00417pt\}\\pgfsys@curveto\{985\.33711pt\}\{\-1092\.00417pt\}\{966\.40369pt\}\{\-1084\.00414pt\}\{950\.40363pt\}\{\-1068\.00407pt\}\\pgfsys@lineto\{809\.60309pt\}\{\-927\.20354pt\}\\pgfsys@lineto\{284\.80109pt\}\{\-1452\.00554pt\}\\pgfsys@lineto\{1728\.8066pt\}\{\-1452\.00554pt\}\\pgfsys@closepath\\pgfsys@moveto\{781\.60298pt\}\{\-899\.20343pt\}\\pgfsys@lineto\{256\.80098pt\}\{\-374\.40143pt\}\\pgfsys@lineto\{256\.80098pt\}\{\-1424\.00543pt\}\\pgfsys@lineto\{781\.60298pt\}\{\-899\.20343pt\}\\pgfsys@closepath\\pgfsys@fillstroke\\pgfsys@invoke\{ \} \\pgfsys@invoke\{ \}\\pgfsys@endscope \\pgfsys@invoke\{ \}\\pgfsys@endscope \\par \\pgfsys@invoke\{ \}\\pgfsys@endscope\{\}\{\}\{\}\\hss\}\\pgfsys@discardpath\\pgfsys@invoke\{ \}\\pgfsys@endscope\\hss\}\}\\endpgfpicture\}\} \\par\} ∗\\astUniversity of Electronic Science and Technology of China †\\daggerLudwig Maximilian University of Munich‡\\ddaggerMunich Center for Machine Learning changzhiliu1@gmail\.comyilun\.liu@tum\.decognitive\.yunpu@gmail\.com §Equal contribution\.
###### Abstract
Self\-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks\. However, existing solutions generally derive self\-modification from a single failure trajectory at a time, overlooking rich comparative signals available in the agent’s expanding archive of past attempts\. According to Mendelian principles of controlled inheritance, we introduce Mendel Gödel Machine \(MGM\)\. In addition to the general single\-trajectory*clonal mutation*, MGM includes two new types of self\-modification that better utilizes evidences accumulated: the*reaction\-norm mutation*edits an agent based on its trajectories on multiple tasks simultaneously, and the*cross\-lineage hybridization*edits an agent using the trajectory of a reference agent from another lineage on the same task\. Under an additive fitness landscape model, we prove theoretically and demonstrate via controlled surrogate simulation that the new strategies facilitate a faster and better convergence over single\-trajectory baselines\. Experiments on SWE\-bench and Polyglot confirm MGM’s consistent improvement in performance, efficiency, and generalizability\.
Project Page:[https://reallcz\.github\.io/MGM/](https://reallcz.github.io/MGM/)
Code:[https://github\.com/RealLcz/MGM](https://github.com/RealLcz/MGM)
## 1Introduction
The vision of artificial intelligence recursively rewriting itself to become better traces back several decades\(Schmidhuber,[1987](https://arxiv.org/html/2608.07645#bib.bib28)\)\. Gödel Machine\(Schmidhuber,[2003](https://arxiv.org/html/2608.07645#bib.bib29);[2007](https://arxiv.org/html/2608.07645#bib.bib30)\)conceives a mathematically rigorous, self\-referential modification process when a provable \(measurable\) benefit can be derived\. Recent advances in large language models \(LLMs\) and coding agents have started to empirically realize such an idea\(Yang et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib41); Hu et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib13); Gao et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib10)\)\.Robeyns et al\. \([2025](https://arxiv.org/html/2608.07645#bib.bib26)\)show that an agent equipped with basic file\-editing tools can autonomously refactor its own codebase and lift its performance\.Zhang et al\. \([2026a](https://arxiv.org/html/2608.07645#bib.bib44)\)reframe this loop as open\-ended evolution, maintaining an expanding archive of agent variants, from which at each iteration one is sampled to seed the next self\-modification\.Wang et al\. \([2026](https://arxiv.org/html/2608.07645#bib.bib36)\)improve the sampling policy by using the aggregated performance of agents’ descendants as signals guiding evaluation and expansion operations\.
However, progress so far has focused on improving the archive that stores generated agents and evaluated traces, and on the process that samples which agent to evaluate or edit next\. The*self\-modification*process, during which the agent edits its own source code, has remained essentially underexplored: each self\-modification step is conditioned only on one agent’s single trajectory \(typically a recent failure\) on one task\. The expanding archive, which records all agent variants ever generated together with their behavior on all evaluated tasks, is used only as a leaderboard for sampling, overlooking rich comparative evidence that can facilitate better\-informed edits\.
Figure 1:Mendel Gödel Machine\. MGM organizes self\-modification via controlled inheritance based on evidence across tasks and lineages\. The archive maintains a lineage tree of agent variants; each stores its source code as genotype and evaluation outcomes as phenotype\. Each iteration appliesπ\\pi\-sampling that selects an operation to perform with the necessary resources from the archive\.φ\\varphi\-evaluation executes the agent on untested tasks and records its trajectory and results\. Clonal mutationΦCM\\varPhi\_\{\\mathrm\{CM\}\}edits the agent based on a single failure trajectory on target taskτt\\tau\_\{\\mathrm\{t\}\}\. Reaction\-norm mutationΦRM\\varPhi\_\{\\mathrm\{RM\}\}uses the agent’s trajectories across reference tasksτr\\tau\_\{\\mathrm\{r\}\}\. Cross\-lineage hybridizationΦCH\\varPhi\_\{\\mathrm\{CH\}\}uses a reference agentara\_\{\\mathrm\{r\}\}from a different lineage that attempted the same task\.We identify two types of such comparative signals\. First, when an agent is evaluated across multiple tasks, the pattern of its successes and failures forms a reaction norm\(Woltereck,[1909](https://arxiv.org/html/2608.07645#bib.bib38); Pigliucci,[2001](https://arxiv.org/html/2608.07645#bib.bib22)\): a stable, genotype\-specific profile of how performance varies across environments\. Recurring failure modes can more plausibly distinct genotype\-level defects from task\-specific accidents\. Second, when multiple agents have attempted the same task, their trajectories reveal transferable behavioral traits\. Conditioning edits on these contrastive evidence enables targeted ability transfer across lineages and reduce redundant exploration\. Building on these observations, we introduceMendel Gödel Machine \(MGM\), which isolates heritable effects through controlled comparisons as in Mendelian genetics\. As shown in Figure[1](https://arxiv.org/html/2608.07645#S1.F1), we structure self\-modification into three operators:*clonal mutation*denotes standard single\-agent, single\-trajectory self\-modification;*reaction\-norm mutation*edits an agent conditioned on its trajectories across multiple tasks; and*cross\-lineage hybridization*edits an agent using a reference agent’s trajectory on the same task\. All strategies directly operate on trajectories accumulated during routine archive evaluation and therefore incur no extra task evaluations\.
To isolate the contribution of each dimension rigorously, we develop a formal additive fitness landscape model in which each agent is represented as a binary vector in genotype space, and self\-improvement progress is measured as the reduction in Hamming distance to an oracle genotype\. We demonstrate that MGM yields a strictly faster expected convergence than single\-trajectory baselines, and we validate this in controlled Monte Carlo surrogate simulations\.
Experiments on Polyglot\(Gauthier,[2024](https://arxiv.org/html/2608.07645#bib.bib11)\)and SWE\-bench\(Jimenez et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib14)\)confirm MGM’s consistent gains over baselines in both performance and efficiency\. Purely through evolving the agent scaffold, MGM advances Qwen3\.6\-35B\-A3B on Polyglot from 50\.8% to 93\.3%, surpassing closed\-source GPT\-5\(Singh et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib33)\)with∼\\sim117×\\timesfewer parameters, as Figure[1](https://arxiv.org/html/2608.07645#footnote1)shows\. MGM also exhibits stronger generalizability across both unseen benchmarks and backbone LLMs\. Notably, transferring the Qwen\-evolved scaffold to DeepSeek\-V4\-Pro yields 96\.9% on Polyglot\. These findings demonstrate that MGM’s richer comparative conditioning discovers reusable, workflow\-level improvements that support strong coding\-agent performance, indicating its potential as a scalable approach for future self\-improving agent development\.
## 2Preliminaries
We denote an executable coding agenta∈𝒜a\\in\\mathcal\{A\}, together with its auxiliary scaffolding, as the genotype that self\-improvement seeks to evolve\. Runningaaon a taskτ∼𝒟\\tau\\sim\\mathcal\{D\}yields an evaluation trajectoryφ\(a,τ\)\\varphi\(a,\\tau\)and a binary outcomer\(a,τ\)∈\{0,1\}r\(a,\\tau\)\\in\\\{0,1\\\}, which is viewed as its phenotype\. The expected utility ofaais
U\(a\)=𝔼τ∼𝒟\[r\(a,τ\)\]\.\\displaystyle U\(a\)=\\mathbb\{E\}\_\{\\tau\\sim\\mathcal\{D\}\}\[r\(a,\\tau\)\]\.\(1\)A self\-modification processΦ\\varPhilets the parent agent edit its own code using diagnostic evidenceEE, a finite set of stored trajectory–outcome pairs:
a′←Φ\(a,E\),E⊆\{\(φ\(a,τ\),r\(a,τ\)\)\}\.\\displaystyle a^\{\\prime\}\\leftarrow\\varPhi\(a,E\),\\qquad E\\subseteq\\\{\(\\varphi\(a,\\tau\),r\(a,\\tau\)\)\\\}\.\(2\)
DGM\(Zhang et al\.,[2026a](https://arxiv.org/html/2608.07645#bib.bib44)\)introduces an archive\-based self\-improvement framework by maintaining an expanding tree of generated agents and their evaluation trajectories\. HGM\(Wang et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib36)\)further formulates it as a fixed\-budget tree\-search problem\. Let𝒢t\\mathcal\{G\}\_\{t\}be the archive tree at stepttand𝒱t\\mathcal\{V\}\_\{t\}its node set\. For nodeai∈𝒱ta\_\{i\}\\in\\mathcal\{V\}\_\{t\}, letSiS\_\{i\}be its evaluated tasks,Fi⊆SiF\_\{i\}\\subseteq S\_\{i\}its failed tasks, and
ns\(ai\)=\|Si\|−\|Fi\|,nf\(ai\)=\|Fi\|\.\\displaystyle n\_\{\\mathrm\{s\}\}\(a\_\{i\}\)=\|S\_\{i\}\|\-\|F\_\{i\}\|,\\qquad n\_\{\\mathrm\{f\}\}\(a\_\{i\}\)=\|F\_\{i\}\|\.\(3\)At each steptt, HGM chooses to allocate either an evaluation or an expansion to its archive\. An expansion happens when
Ntα≥\|𝒱t\|,Nt=∑a∈𝒱t\(ns\(a\)\+nf\(a\)\),\\displaystyle N\_\{t\}^\{\\alpha\}\\geq\|\\mathcal\{V\}\_\{t\}\|,\\qquad N\_\{t\}=\\sum\_\{a\\in\\mathcal\{V\}\_\{t\}\}\(n\_\{\\mathrm\{s\}\}\(a\)\+n\_\{\\mathrm\{f\}\}\(a\)\),\(4\)whereα∈\[0,1\]\\alpha\\in\[0,1\]is a widening parameter, and otherwise another task evaluation is performed\.
The evaluation policy samplesaausing Thompson sampling from the node\-level posterior
πa∼Beta\(κ\(1\+ns\(a\)\),κ\(1\+nf\(a\)\)\),\\displaystyle\\pi\_\{a\}\\sim\\mathrm\{Beta\}\(\\kappa\(1\+n\_\{\\mathrm\{s\}\}\(a\)\),\\kappa\(1\+n\_\{\\mathrm\{f\}\}\(a\)\)\),\(5\)whereκ\>0\\kappa\>0is the concentration parameter controlling the exploration–exploitation trade\-off\. The expansion policy samplesaato expand using clade\-level evidence\. LetCt\(a\)C\_\{t\}\(a\)denote the subtree rooted ataa, and define
nsC\(a\)=∑a′∈Ct\(a\)ns\(a′\),nfC\(a\)=∑a′∈Ct\(a\)nf\(a′\)\.\\displaystyle n\_\{\\mathrm\{s\}\}^\{C\}\(a\)=\\sum\_\{a^\{\\prime\}\\in C\_\{t\}\(a\)\}n\_\{\\mathrm\{s\}\}\(a^\{\\prime\}\),\\qquad n\_\{\\mathrm\{f\}\}^\{C\}\(a\)=\\sum\_\{a^\{\\prime\}\\in C\_\{t\}\(a\)\}n\_\{\\mathrm\{f\}\}\(a^\{\\prime\}\)\.\(6\)HGM samples expansion candidates from
πaC∼Beta\(κ\(1\+nsC\(a\)\),κ\(1\+nfC\(a\)\)\)\.\\displaystyle\\pi\_\{a\}^\{C\}\\sim\\mathrm\{Beta\}\(\\kappa\(1\+n\_\{\\mathrm\{s\}\}^\{C\}\(a\)\),\\kappa\(1\+n\_\{\\mathrm\{f\}\}^\{C\}\(a\)\)\)\.\(7\)
## 3Mendel Gödel Machine
MGM builds on the aforementioned tree\-search framework of selection, evaluation, and expansion policies\. Instead of relying on single\-trajectory self\-modification, MGM partitions the expansion operatorΦ\\varPhiinto three specialized sub\-operators based on the type of diagnostic evidenceEEavailable in the archive: clonal mutationΦCM\\varPhi\_\{\\rm CM\}, reaction\-norm mutationΦRM\\varPhi\_\{\\rm RM\}, and cross\-lineage hybridizationΦCH\\varPhi\_\{\\rm CH\}, as Figure[1](https://arxiv.org/html/2608.07645#S1.F1)shows\.
### 3\.1Mendelian Self\-Modification Operators
Analogous to Mendelian genetics, which seeks to isolate heritable effects through controlled comparisons, we design self\-modification operators as diagnostic processes that ask the selected agent to make general improvements to its genotype based on different phenotype evidence\.
#### Clonal Mutation\.
Clonal mutation is the standard single\-agent, single\-trajectory self\-improvement operator\. It is used when MGM has only one informative failure or cannot construct a reliable comparison\. Given a selected nodeiiand a failed taskτ∈Fi\\tau\\in F\_\{i\}, the evidence is
ECM\(i,τ\)=\{\(φ\(ai,τ\),r\(ai,τ\)\)\}\.E\_\{\\mathrm\{CM\}\}\(i,\\tau\)=\\\{\(\\varphi\(a\_\{i\},\\tau\),r\(a\_\{i\},\\tau\)\)\\\}\.\(8\)The editor diagnoses the failure and modifiesaia\_\{i\}to avoid similar failures in future tasks:
a′←ΦCM\(ai,ECM\)\.a^\{\\prime\}\\leftarrow\\varPhi\_\{\\mathrm\{CM\}\}\(a\_\{i\},E\_\{\\mathrm\{CM\}\}\)\.\(9\)This operator preserves the behavior of HGM\-style self\-modification and ensures that the search can proceed even when the archive is still small\.
#### Reaction\-norm Mutation\.
A reaction norm describes how one genotype expresses different phenotypes under different environments\(Woltereck,[1909](https://arxiv.org/html/2608.07645#bib.bib38); Pigliucci,[2001](https://arxiv.org/html/2608.07645#bib.bib22)\)\. MGM incorporates this idea and designs reaction\-norm mutationΦRM\\varPhi\_\{\\mathrm\{RM\}\}by comparing multiple phenotypes of the same genotype across different tasks\. If the same agent fails, or behaves inconsistently, across multiple tasks, the resulting pattern can provide richer diagnostic evidence of the agent’s genuine weakness, rather than accidental task\-specific errors\.
Formally,ΦRM\\varPhi\_\{\\mathrm\{RM\}\}becomes available foraia\_\{i\}when it has accumulated enoughφ\(ai,τ\)\\varphi\(a\_\{i\},\\tau\)trajectories
\|Si\|≥mRM,\|S\_\{i\}\|\\geq m\_\{\\mathrm\{RM\}\},\(10\)and there exist at least two trajectories, with a failed one as the targetτt∈Si\\tau\_\{\\mathrm\{t\}\}\\in S\_\{i\}\. The referenceτr\\tau\_\{\\mathrm\{r\}\}may be any other trajectory inSiS\_\{i\}, preferably a failed one if available\. The evidence
ERM\(ai,τt,τr\)=\{\\displaystyle E\_\{\\mathrm\{RM\}\}\(a\_\{i\},\\tau\_\{\\mathrm\{t\}\},\\tau\_\{\\mathrm\{r\}\}\)=\\\{\(φ\(ai,τt\),r\(ai,τt\)\),\(φ\(ai,τr\),r\(ai,τr\)\)\}\.\\displaystyle\\big\(\\varphi\(a\_\{i\},\\tau\_\{\\mathrm\{t\}\}\),r\(a\_\{i\},\\tau\_\{\\mathrm\{t\}\}\)\\big\),\\big\(\\varphi\(a\_\{i\},\\tau\_\{\\mathrm\{r\}\}\),r\(a\_\{i\},\\tau\_\{\\mathrm\{r\}\}\)\\big\)\\\}\.\(11\)is given to the self\-modification process, where the agent is asked to identify a recurring or contrastive behavioral pattern shared by the provided trajectories and to implement a general improvement
a′←ΦRM\(ai,ERM\)\.a^\{\\prime\}\\leftarrow\\varPhi\_\{\\mathrm\{RM\}\}\(a\_\{i\},E\_\{\\mathrm\{RM\}\}\)\.\(12\)
#### Cross\-lineage Hybridization\.
Cross\-lineage hybridization compares different genotypes under the same task environment\. It is available when two nodes have attempted at least one common target task, and the task is not already solved by both:
∃j≠i,∃τt∈Si∩Sj\\displaystyle\\exists j\\neq i,\\ \\exists\\tau\_\{\\mathrm\{t\}\}\\in S\_\{i\}\\cap S\_\{j\}s\.t\.¬\(r\(ai,τt\)=1∧r\(aj,τt\)=1\)\.\\displaystyle\\quad\\text\{s\.t\.\}\\quad\\neg\\big\(r\(a\_\{i\},\\tau\_\{\\mathrm\{t\}\}\)=1\\land r\(a\_\{j\},\\tau\_\{\\mathrm\{t\}\}\)=1\\big\)\.\(13\)MGM designates one failed agent as the targetata\_\{\\mathrm\{t\}\}to improve\. If the reference agentara\_\{\\mathrm\{r\}\}fails likewise, MGM uses the comparison to identify complementary failure modes; ifara\_\{\\mathrm\{r\}\}solvedτt\\tau\_\{\\mathrm\{t\}\}, differences in their genotypes are used as corrective signals to guideata\_\{\\mathrm\{t\}\}’s self\-modification\. The evidence is
ECH\(at,ar,τt\)=\{\\displaystyle E\_\{\\mathrm\{CH\}\}\(a\_\{\\mathrm\{t\}\},a\_\{\\mathrm\{r\}\},\\tau\_\{\\mathrm\{t\}\}\)=\\\{\(φ\(at,τt\),r\(at,τt\)\),\(φ\(ar,τt\),r\(ar,τt\)\)\}\.\\displaystyle\(\\varphi\(a\_\{\\mathrm\{t\}\},\\tau\_\{\\mathrm\{t\}\}\),r\(a\_\{\\mathrm\{t\}\},\\tau\_\{\\mathrm\{t\}\}\)\),\(\\varphi\(a\_\{\\mathrm\{r\}\},\\tau\_\{\\mathrm\{t\}\}\),r\(a\_\{\\mathrm\{r\}\},\\tau\_\{\\mathrm\{t\}\}\)\)\\\}\.\(14\)whereata\_\{\\mathrm\{t\}\}is the target agent andara\_\{\\mathrm\{r\}\}is the reference agent\. The child is attached to the primary lineage:
a′←ΦCH\(at,ECH\)\.a^\{\\prime\}\\leftarrow\\varPhi\_\{\\mathrm\{CH\}\}\(a\_\{\\mathrm\{t\}\},E\_\{\\mathrm\{CH\}\}\)\.\(15\)
This hybridization operation is a diagnostic process\. MGM does not splice source files from one agent into another\. Instead, it asks the target agent itself to extract a transferable behavioral trait from the reference trajectory and adapt that trait to the target agent’s own codebase\. This encourages the failing lineage to derive and inherit a genuine improvement rather than task\-specific behaviors\.
### 3\.2Sampling Tasks and Operators
MGM additionally maintains a global pool
𝒫t=⋃i∈𝒱tFi,\\mathcal\{P\}\_\{t\}=\\bigcup\_\{i\\in\\mathcal\{V\}\_\{t\}\}F\_\{i\},\(16\)which stores tasks that have exposed failures in any previously evaluated agent\. This pool is not a separate evaluation benchmark and does not introduce extraφ\\varphi\-evaluations\. It is maintained to control how future evaluation tasks are sampled\. When selecting a new task for an agent, MGM samples from tasks not yet attempted by that agent, but assigns a predetermined weight to tasks in𝒫t\\mathcal\{P\}\_\{t\}:
wi\(τ\)=\{βfail,τ∈𝒫t,1,τ∉𝒫t,τ∉Si,w\_\{i\}\(\\tau\)=\\begin\{cases\}\\beta\_\{\\mathrm\{fail\}\},&\\tau\\in\\mathcal\{P\}\_\{t\},\\\\ 1,&\\tau\\notin\\mathcal\{P\}\_\{t\},\\end\{cases\}\\qquad\\tau\\notin S\_\{i\},\(17\)whereβfail\\beta\_\{\\mathrm\{fail\}\}is the failed\-pool boost\. This task sampling design concentrates evaluation on tasks that are known to reveal weaknesses in at least one lineage, increasing the diagnostic value of eachφ\\varphi\-evaluation\. It also deliberately creates overlap across lineages\. SinceΦCH\\varPhi\_\{\\textrm\{CH\}\}requires agents to have attempted a shared task, the pool increases the availability of controlled cross\-lineage comparisons without spending additional evaluation budget, making transferable behavioral traits easier to discover\.
For strategy selection, MGM inherits HGM’s Thompson\-sampling policyπ\\pi, which at each step chooses between initiating aφ\\varphi\-evaluation for an existing node or aΦ\\varPhi\-expansion of the evolution tree, using the clade\-level and node\-level Beta posteriors defined in Section[2](https://arxiv.org/html/2608.07645#S2)\. The difference lies in how an expansion is executed\. MGM partitions the singleΦ\\varPhioperator that HGM and DGM use \(ΦCM\\varPhi\_\{\\mathrm\{CM\}\}\) into three sub\-operators\{ΦCM,ΦRM,ΦCH\}\\\{\\varPhi\_\{\\mathrm\{CM\}\},\\varPhi\_\{\\mathrm\{RM\}\},\\varPhi\_\{\\mathrm\{CH\}\}\\\}\. Whenπ\\piselects aΦ\\varPhi\-expansion for parentaia\_\{i\}, MGM first constructs the set of eligible operatorsΩi⊆\{ΦCM,ΦRM,ΦCH\}\\Omega\_\{i\}\\subseteq\\\{\\varPhi\_\{\\mathrm\{CM\}\},\\varPhi\_\{\\mathrm\{RM\}\},\\varPhi\_\{\\mathrm\{CH\}\}\\\}from the archive\.ΦCM\\varPhi\_\{\\mathrm\{CM\}\}is eligible whenever the selected agent has at least one failed task, i\.e\.,Fi≠∅F\_\{i\}\\neq\\varnothing\.ΦRM\\varPhi\_\{\\mathrm\{RM\}\}becomes eligible if the agent has been evaluated on at leastmRMm\_\{\\mathrm\{RM\}\}distinct tasks \(Equation[10](https://arxiv.org/html/2608.07645#S3.E10)\), and there exist trajectories for two distinct tasksτt,τr∈Si\\tau\_\{\\mathrm\{t\}\},\\tau\_\{\\mathrm\{r\}\}\\in S\_\{i\}with at least one failure\.ΦCH\\varPhi\_\{\\mathrm\{CH\}\}becomes eligible if agents from two lineages have attempted a shared taskτt\\tau\_\{\\mathrm\{t\}\}that is not solved by both \(Equation[13](https://arxiv.org/html/2608.07645#S3.E13)\)\.
MGM then samples among eligible operators with configurable weightsλCM,λRM,λCH\\lambda\_\{\\mathrm\{CM\}\},\\lambda\_\{\\mathrm\{RM\}\},\\lambda\_\{\\mathrm\{CH\}\}:
Pr\(σ∣i\)=λσ∑σ′∈Ωiλσ′,σ∈Ωi\.\\Pr\(\\sigma\\mid i\)=\\frac\{\\lambda\_\{\\sigma\}\}\{\\sum\_\{\\sigma^\{\\prime\}\\in\\Omega\_\{i\}\}\\lambda\_\{\\sigma^\{\\prime\}\}\},\\qquad\\sigma\\in\\Omega\_\{i\}\.\(18\)The selected operator determines the evidenceEσE\_\{\\sigma\}, and the child is produced by
a′←Φσ\(at,Eσ\)\.a^\{\\prime\}\\leftarrow\\varPhi\_\{\\sigma\}\(a\_\{\\mathrm\{t\}\},E\_\{\\sigma\}\)\.\(19\)ForΦCM\\varPhi\_\{\\mathrm\{CM\}\}andΦRM\\varPhi\_\{\\mathrm\{RM\}\}, the target agent is the selected parentat=aia\_\{\\mathrm\{t\}\}=a\_\{i\}\. ForΦCH\\varPhi\_\{\\mathrm\{CH\}\}, if exactly one agent solves the shared task, the failing agent is edited using the successful agent as reference; if both fail, the higher\-utility lineage serves as the primary target\. IfΩi=∅\\Omega\_\{i\}=\\varnothing, MGM skips the expansion andπ\\piallocates anotherφ\\varphi\-evaluation instead\.
The complete pseudocode of MGM is provided in Appendix[A\.1](https://arxiv.org/html/2608.07645#A1.SS1)\.
## 4Simulations
Figure 3:Additive fitness landscape model\.Each agentaacarries a binary genotype with loci that are either correct or mismatched relative to an oraclea∞a\_\{\\infty\}\. The genotype is not directly observable and can only be examined throughφ\\varphi\-evaluation, where each taskτi\\tau\_\{i\}requires a fixed subset of loci to be all correct for a successful phenotype\.Φ\\varPhi\-expansion reads phenotype records and modifies the genotype by flipping the examined loci, which may correct mismatched ones or corrupt correct ones under certain probabilities\.To validate MGM’s design decisions in isolation of implementation details and benchmark choices, we instantiate controlled surrogate models for methods mentioned in Sections[2](https://arxiv.org/html/2608.07645#S2)and[3](https://arxiv.org/html/2608.07645#S3)\. We demonstrate that comparative evidence can improve the effective fix probability of self\-modification by reducing diagnostic uncertainty, and study how such diagnostic advantage affects performance and efficiency\.
### 4\.1Additive Fitness Landscape
Each agent’s genotype is modeled as a binary vector𝐠∈\{0,1\}L\\mathbf\{g\}\\in\\\{0,1\\\}^\{L\}withLLloci, each corresponding to a minimum scaffold\-level capability or implementation choice\. A fixed oracle genotype𝐠∗∈\{0,1\}L\\mathbf\{g\}^\{\*\}\\in\\\{0,1\\\}^\{L\}represents the optimal program; without loss of generality, we set𝐠∗=𝟏\\mathbf\{g\}^\{\*\}=\\mathbf\{1\}\. The genotype is not directly exposed; instead, its Hamming distance to the oracle is
d\(𝐠\)=∑ℓ=1L𝟏\[gℓ≠gℓ∗\],d\(\\mathbf\{g\}\)\\;=\\;\\sum\_\{\\ell=1\}^\{L\}\\mathbf\{1\}\\\!\\left\[g\_\{\\ell\}\\neq g^\{\*\}\_\{\\ell\}\\right\],\(20\)whered\(𝐠\)=0d\(\\mathbf\{g\}\)=0if and only if𝐠=𝐠∗\\mathbf\{g\}=\\mathbf\{g\}^\{\*\}, serving as a zero\-error lower bound\. Each run starts from an initial genotype withd0d\_\{0\}mismatched loci, which self\-improvement must correct to reach the oracle\.
The task pool containsNNtasks\. Each taskτ\\tauexamines a subset of lociRτ⊆\[L\]R\_\{\\tau\}\\subseteq\[L\]with\|Rτ\|=k\|R\_\{\\tau\}\|=k\. The agent solves the task only when all required loci are correct:
r\(a,τ\)=1⟺Rτ∩M\(a\)=∅,\\displaystyle r\(a,\\tau\)=1\\quad\\Longleftrightarrow\\quad R\_\{\\tau\}\\cap M\(a\)=\\varnothing,\(21\)M\(a\)=\{ℓ∈\[L\]:gℓ\(a\)≠gℓ∗\}\\displaystyle M\(a\)=\\\{\\ell\\in\[L\]:g\_\{\\ell\}\(a\)\\neq g^\{\*\}\_\{\\ell\}\\\}\(22\)whereM\(a\)M\(a\)denotes the set of incorrect loci of agentaa\. Under independent locus sampling, the resulting per\-task success probability for an agent at edit distanceddis
P\(r=1∣d\)=\(L−dL\)k\.P\(r\{=\}1\\mid d\)=\\Bigl\(\\tfrac\{L\-d\}\{L\}\\Bigr\)^\{k\}\.\(23\)At the start of each run, this probability equals\(\(L−d0\)/L\)k\(\(L\-d\_\{0\}\)/L\)^\{k\}\. This models coding tasks as requiring multiple scaffold\-level capabilities to be simultaneously correct, such as localization, reasoning, editing, and validation\. All parameter values are listed in Table[5](https://arxiv.org/html/2608.07645#A2.T5)\.
### 4\.2Comparative Evidence as Diagnostic Compression
We demonstrate how the comparative operators in MGM can induce a higher effective fix probability than single\-trajectory mutation\. A self\-modification operator does not observe the hidden incorrect loci directly\. Instead, given diagnostic evidenceEE, it transiently constructs an implicit candidate setCσ\(E\)⊆\[L\]C\_\{\\sigma\}\(E\)\\subseteq\[L\]of loci that may explain the observed failure, whereσ∈\{CM,RM,CH\}\\sigma\\in\\\{\\mathrm\{CM\},\\mathrm\{RM\},\\mathrm\{CH\}\\\}\. Suppose that once the editor targets an actually incorrect locus, it repairs it with probabilitys∈\(0,1\]s\\in\(0,1\]\. Then the effective fix probability of operatorσ\\sigmais
pfσ=s⋅Prℓ∼Cσ\(E\)\[ℓ∈M\(a\)\]\.p\_\{f\}^\{\\sigma\}=s\\cdot\\Pr\_\{\\ell\\sim C\_\{\\sigma\}\(E\)\}\[\\ell\\in M\(a\)\]\.\(24\)Therefore, the fix probability increases when the evidence yields a candidate set with a higher density of truly incorrect loci\.
ΦCM\\varPhi\_\{\\textrm\{CM\}\}observes a single failed taskτt\\tau\_\{t\}\. Since this evidence only impliesRτt∩M\(a\)≠∅R\_\{\\tau\_\{t\}\}\\cap M\(a\)\\neq\\varnothing, the natural candidate set isCCM=RτtC\_\{\\mathrm\{CM\}\}=R\_\{\\tau\_\{t\}\}\. By contrast,ΦRM\\varPhi\_\{\\textrm\{RM\}\}compares multiple trajectories of the same genotype\. When two failures share a recurring scaffold\-level defect, the common explanatory region is compressed toCRM=Rτt∩RτrC\_\{\\mathrm\{RM\}\}=R\_\{\\tau\_\{t\}\}\\cap R\_\{\\tau\_\{r\}\}, which is smaller thanRτtR\_\{\\tau\_\{t\}\}in expectation\.ΦCH\\varPhi\_\{\\textrm\{CH\}\}compares different genotypes on the same task\. When a reference agent succeeds on the task while the target agent fails, the reference trajectory acts as a contrastive control and filters out non\-causal task\-relevant loci fromRτR\_\{\\tau\}\. This gives the following proposition\.
#### Proposition 1\.
Under the aforementioned model and sound comparative evidence,ΦRM\\varPhi\_\{\\textrm\{RM\}\}andΦCH\\varPhi\_\{\\textrm\{CH\}\}have strictly higher effective fix probability thanΦCM\\varPhi\_\{\\textrm\{CM\}\}:
pfRM\>pfCM,pfCH\>pfCM\.p\_\{f\}^\{\\mathrm\{RM\}\}\>p\_\{f\}^\{\\mathrm\{CM\}\},\\qquad p\_\{f\}^\{\\mathrm\{CH\}\}\>p\_\{f\}^\{\\mathrm\{CM\}\}\.\(25\)The full derivation is provided in Appendix[B\.1](https://arxiv.org/html/2608.07645#A2.SS1)\. Intuitively, comparative evidence improves self\-modification not by making the editor intrinsically stronger, but by reducing diagnostic uncertainty: the editor searches over a smaller and cleaner set of candidate defects\.
### 4\.3Monte Carlo Simulation
Guided by Proposition 1, we instantiate the diagnostic advantage of comparative evidence through a controllable fix\-probability ratio
ρ=pfRMpfCM=pfCHpfCM\.\\rho=\\frac\{p\_\{f\}^\{\\mathrm\{RM\}\}\}\{p\_\{f\}^\{\\mathrm\{CM\}\}\}=\\frac\{p\_\{f\}^\{\\mathrm\{CH\}\}\}\{p\_\{f\}^\{\\mathrm\{CM\}\}\}\.\(26\)The null settingρ=1\\rho=1corresponds to the case where comparative evidence provides no additional diagnostic benefit, whileρ\>1\\rho\>1models increasing levels of diagnostic compression\.
The simulation isolates axes on which the three methods differ\. For budget allocation, DGM assigns a fixednevaln\_\{\\mathrm\{eval\}\}evaluations to every population member per generation, distributing budget uniformly regardless of node quality\. HGM uses adaptive allocation, concentrating evaluations on promising nodes and triggering edits according to its tree\-search policy\. DGM and HGM use only clonal mutationΦCM\\varPhi\_\{\\mathrm\{CM\}\}\. MGM follows the same evaluation\-expansion framework but additionally employsΦRM\\varPhi\_\{\\mathrm\{RM\}\}andΦCH\\varPhi\_\{\\mathrm\{CH\}\}when their availability conditions hold\.
To ensure that the comparison isolates the value of diagnostic evidence rather than unequal compute, all edit operators incur the same cost:
cCM=cRM=cCH=cφ\.c\_\{\\mathrm\{CM\}\}=c\_\{\\mathrm\{RM\}\}=c\_\{\\mathrm\{CH\}\}=c\_\{\\varphi\}\.\(27\)Each method runs for a fixed total budget ofBBresource units withnseedsn\_\{\\mathrm\{seeds\}\}independent Monte Carlo seeds\. At each seed, a fresh task pool and initial genotype are regenerated\. The minimum edit distance across all active nodes is recorded at evenly spaced budget checkpoints; final performance is reported as mean±\\pm95 % CI across seeds\. A separate parameter sweep varies the initial edit distanced0d\_\{0\}and the diagnostic\-advantage ratioρ\\rhoto test robustness of the ordering\. Detailed configurations for our simulation can be found in Table[5](https://arxiv.org/html/2608.07645#A2.T5)\.
The parameter sensitivity results in Figure[5](https://arxiv.org/html/2608.07645#S4.F5)and Figure[5](https://arxiv.org/html/2608.07645#S4.F5)test the robustness of MGM across a broad range of initial difficultiesd0d\_\{0\}and operator\-quality settingsρ\\rho\. MGM consistently outperforms other baselines in both final performance and convergence speed across alld0d\_\{0\}andρ\>1\.0\\rho\>1\.0settings, demonstrating reliable gains across both easier and harder initial conditions\. The null caseρ=1\\rho=1suggests that when comparative operators provide no fix\-quality advantage, MGM collapses back to HGM\-like behavior\. This confirms that the simulated gain is caused directly by the diagnostic\-quality advantage formalized in Proposition 1 in Section[4\.2](https://arxiv.org/html/2608.07645#S4.SS2.SSS0.Px1)and Appendix[B\.1](https://arxiv.org/html/2608.07645#A2.SS1)\. Asρ\\rhoincreases, MGM’s advantage grows steadily, and this trend is especially pronounced at smallerd0d\_\{0\}, where fewer but more precise self\-improvements are required and higher\-quality diagnostic signals translate into clearer performance gains\. Figure[5](https://arxiv.org/html/2608.07645#S4.F5)additionally shows the distribution of final performance across seeds, where MGM achieves the lowest mean and tightest spread, while the broader distributions of other baselines reflect less stable budget allocation\.
d0=80d\_\{0\}=80
d0=40d\_\{0\}=40
d0=20d\_\{0\}=20
d0=10d\_\{0\}=10
ρ=2\.0\\rho=2\.0ρ=1\.5\\rho=1\.5ρ=1\.2\\rho=1\.2ρ=1\.0\\rho=1\.0Figure 4:Simulated performance evolution over cumulative budget spent\.Results compared across task difficulty and edit effectiveness\. Rows vary the initial edit distanced0d\_\{0\}; columns vary the fix\-probability advantage ratioρ\\rho\. Each panel shows the edit distance results \(lower is better\) averaged across random seeds with 95 % CIs\. Dashed lines mark oracle optima\.
d0=80d\_\{0\}=80
d0=40d\_\{0\}=40
d0=20d\_\{0\}=20
d0=10d\_\{0\}=10
ρ=2\.0\\rho=2\.0ρ=1\.5\\rho=1\.5ρ=1\.2\\rho=1\.2ρ=1\.0\\rho=1\.0Figure 5:Simulated final performance distribution\.Results compared across task difficulty and edit effectiveness\. Rows vary the initial edit distanced0d\_\{0\}; columns vary the fix\-probability advantage ratioρ\\rho\. Each panel shows the distribution of final edit distance results \(lower is better\) across all random seeds, with dashed lines marking per\-method means\.
## 5Experiments
We evaluate MGM on challenging software\-engineering benchmarks to answer three research questions: Does MGM exhibit better performance and efficiency? Can agents evolved by MGM generalize to other challenging tasks or models? What is the contribution of each key component in MGM?
Our experiments span several challenging and representative benchmarks, including SWE\-bench Verified\(Jimenez et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib14)\), SWE\-bench Pro\(Deng et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib8)\), SWE\-bench Multilingual\(Khandpur & the SWE\-bench Team,[2025](https://arxiv.org/html/2608.07645#bib.bib16); Yang et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib42)\), and Polyglot\(Gauthier,[2024](https://arxiv.org/html/2608.07645#bib.bib11)\)\. In all experiments, agents are not given access to private test cases or test results during evolution\. To ensure fair comparisons and control computational cost, unless otherwise specified, all results reported in this section are evaluated on the same two 60\-task subsets of SWE\-bench Verified and Polyglot asWang et al\. \([2026](https://arxiv.org/html/2608.07645#bib.bib36)\)andZhang et al\. \([2026a](https://arxiv.org/html/2608.07645#bib.bib44)\)\.
### 5\.1Performance
To assess whether MGM exhibits greater self\-improvement potential than HGM, following the setting ofZhang et al\. \([2026a](https://arxiv.org/html/2608.07645#bib.bib44)\); Wang et al\. \([2026](https://arxiv.org/html/2608.07645#bib.bib36)\), we evaluate both methods on SWE\-bench Verified and Polyglot under an identical computational budget of 200 evaluations\. We adopt Qwen3\.6\-35B\-A3B\(Qwen Team,[2026](https://arxiv.org/html/2608.07645#bib.bib25)\)as the backbone LLM\. To ensure a fair comparison, both methods start from the same ancestor agent\. We report the accuracy of the best\-belief agent evolved by each method\.
Table 1:Performance of coding agents evolved on SWE\-bench Verified and Polyglot\.Results evolved using Qwen3\.6\-35B\-A3B after 200φ\\varphi\-evaluations and 24Φ\\varPhi\-expansions\. For each benchmark, HGM and MGM start from the same initial scaffold\. Superscripts in accuracy denote absolute percentage\-point improvements over corresponding initial agents\. Time reported as CPU wall\-clock time with 8×\\timesNVIDIA H100 GPUs\.![[Uncaptioned image]](https://arxiv.org/html/2608.07645v1/x34.png)Figure 2:Polyglot performance\.Result marked with asterisk is from Polyglot\-60; all other scores are from the complete Polyglot‑225111Sourced from Aider\-Polyglot Leaderboard[https://llm\-stats\.com/benchmarks/aider\-polyglot](https://llm-stats.com/benchmarks/aider-polyglot)\.\. Closed\-source model sizes follows estimates reported byLi \([2026](https://arxiv.org/html/2608.07645#bib.bib17)\)\.
As shown in Table[1](https://arxiv.org/html/2608.07645#S5.T1), MGM consistently outperforms HGM\. On SWE\-bench Verified, both methods start from the same initial accuracy of 68\.3%, which HGM improves to 73\.3%, and MGM improves to 78\.3%\. On Polyglot, starting from the same 50\.8% agent, HGM improves to 77\.9%, while MGM reaches 93\.2%\. These result shows that MGM’s comparative self\-modification operators are effective on both standalone coding tasks as in Polyglot, and also repository\-level software\-engineering tasks, where failures often involve incompetency in localization, environment understanding, and repair workflows\. Since both methods use exactly the same number of evaluation and expansion operations, this performance gap cannot be attributed to a larger search budget\. Instead, it supports our central hypothesis that reusing archived trajectories through reaction\-norm mutation and cross\-lineage hybridization provides more informative diagnostic evidence than conditioning each self\-modification step on a single failure trajectory\. Evaluation on the full 225\-task Polyglot benchmark \(Figure[1](https://arxiv.org/html/2608.07645#footnote1)and Appendix[E\.1](https://arxiv.org/html/2608.07645#A5.SS1)\) further confirms the same results\.
We further analyze the token consumption of each step in the experiments\. Figure[5\.2](https://arxiv.org/html/2608.07645#S5.SS2.SSS0.Px2)shows that all evolutionary strategies used by HGM and MGM have comparable average token costs, indicating that the observed performance gains are not achieved simply by increased token expenditure\.
### 5\.2Generalization
A key question for self\-improving coding agents is whether the evolved improvements remain useful beyond the exact setting in which evolution is performed\(Zhang et al\.,[2026a](https://arxiv.org/html/2608.07645#bib.bib44); Wang et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib36)\)\. Since MGM modifies the agent scaffold rather than the parameters of the underlying LLM, a successful evolution method should ideally discover reusable workflow\-level improvements that transfer across both benchmarks and backbone models\(Hu et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib13)\)\. We therefore evaluate generalization from two complementary perspectives: cross\-benchmark generalization and cross\-model scaffold transfer\.
#### Cross\-benchmark generalization\.
We first evaluate whether scaffolds evolved on Polyglot\-60 can transfer to different software\-engineering settings without further self\-improvement\. After evolution, we freeze the evolved scaffolds and evaluate them on subsets of SWE\-bench Pro and SWE\-bench Multilingual\(Khandpur & the SWE\-bench Team,[2025](https://arxiv.org/html/2608.07645#bib.bib16); Yang et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib42)\)\. We report the tasks used in Appendix[D](https://arxiv.org/html/2608.07645#A4)\. We choose these two benchmarks because they test complementary forms of out\-of\-distribution generalization\. SWE\-bench Pro provides a more challenging repository\-level software\-engineering setting, where solving a task often requires understanding larger codebases, localizing the relevant files, and producing robust patches\. SWE\-bench Multilingual further stresses the language generality of the evolved scaffold by covering repositories implemented in diverse programming languages\.
Table 2:Cross\-benchmark generalization from Polyglot to SWE\-bench Pro and SWE\-bench Multilingual\.All agents evolved on Polyglot are evaluated zero\-shot on held\-out SWE\-bench variants, using Qwen3\.6\-35B\-A3B\. Superscripts in accuracy denote absolute percentage\-point improvements over corresponding initial agents\.As shown in Table[2](https://arxiv.org/html/2608.07645#S5.T2), the initial scaffold obtains 16\.7% accuracy on SWE\-bench Pro and 41\.7% on SWE\-bench Multilingual\. The HGM\-evolved scaffold shows limited transfer: it decreases to 13\.3% on SWE\-bench Pro, corresponding to a−3\.4\-3\.4percentage\-point change, and slightly improves to 43\.3% on SWE\-bench Multilingual, corresponding to a\+1\.6\+1\.6percentage\-point gain\. In contrast, MGM achieves positive transfer on both target benchmarks\. It reaches 26\.7% on SWE\-bench Pro, giving a\+10\.0\+10\.0percentage\-point improvement over the initial scaffold, and 55\.0% on SWE\-bench Multilingual, giving a\+13\.3\+13\.3percentage\-point improvement\.
These results suggest that MGM’s comparative self\-modification operators help discover scaffold changes that remain useful beyond the source benchmark\. The gains on both SWE\-bench Pro and SWE\-bench Multilingual indicate that MGM does not overfit to Polyglot\-specific failure patterns and is more likely to acquire reusable software\-engineering workflows that transfer across benchmark distributions and programming\-language settings\.
#### Cross\-model generalization\.
We further evaluate whether the evolved scaffolds remain effective when paired with different backbone LLMs\. In this experiment, the scaffolds are evolved on SWE\-bench Verified\-60 using Qwen3\.6\-35B\-A3B\. We then freeze the evolved scaffold and replace only the inference backbone with DeepSeek\-V4\-Flash and DeepSeek\-V4\-Pro\(DeepSeek\-AI et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib7)\)\. This protocol isolates whether the improvement is encoded in the scaffold itself, rather than being specific to the original Qwen3\.6 backbone\.
Table 3:Cross\-model transfer on SWE\-bench Verified\-60\.The Qwen3\.6\-35B\-A3B block reports the original evolved agents, while the DeepSeek blocks report transferred performance by using the scaffolds evolved on Qwen3\.6\-35B\-A3B and evaluating with LLM backbone replaced\. Superscripts in accuracy denote absolute percentage\-point improvements over corresponding initial agents\.Table[3](https://arxiv.org/html/2608.07645#S5.T3)shows that both HGM and MGM produce scaffolds that transfer across models, but MGM transfers more strongly and consistently\. Under the original Qwen3\.6\-35B\-A3B backbone, HGM improves the initial scaffold from 68\.3% to 73\.3%, whereas MGM improves it to 78\.3%, giving MGM a \+5\.0 percentage\-point advantage over HGM\. After transferring the same scaffolds to DeepSeek\-V4\-Flash, HGM reaches 60\.0%, while MGM further improves to 66\.7%\. On DeepSeek\-V4\-Pro, HGM reaches 70\.0%, whereas MGM reaches 75\.0%\. Averaged over the two transferred DeepSeek backbones, MGM achieves 70\.8% accuracy, compared with 65\.0% for HGM and 47\.5% for the initial scaffold\. The result confirms that MGM does not simply tune the scaffold to the quirks of a single foundation model\. Instead, its evolved changes remain beneficial when the same scaffold is executed by substantially different inference backbones\.
Beyond the subset\-level transfer results above, we further evaluate whether the scaffold evolved on Polyglot with Qwen3\.6\-35B\-A3B remains effective when paired with a stronger inference backbone\. Specifically, we freeze the MGM\-evolved Polyglot scaffold, replace only the backbone model with DeepSeek\-V4\-Pro, and evaluate it on the complete 225\-task Polyglot benchmark\. As Figure[1](https://arxiv.org/html/2608.07645#footnote1)shows, this transferred configuration achieves96\.89%accuracy, further advancing the frontier result achieved with Qwen above\. This result provides additional evidence that MGM truely discovers reusable workflow\-level improvements that can be further amplified by stronger foundation models\.
Together, the cross\-benchmark and cross\-model results demonstrate that MGM helps improve not only in\-domain performance, but also the generalizability of the evolved agent scaffold across benchmarks and models\. This points to a potential path for scalable self\-improving agent development, where scaffolds evolved on smaller datasets and cheaper backbones can be reused on stronger models\. Appendices[F\.1](https://arxiv.org/html/2608.07645#A6.SS1)and[F\.3](https://arxiv.org/html/2608.07645#A6.SS3)further provide qualitative analysis demonstrating how these gains actually correspond to reusable workflow\-level skills\.
![[Uncaptioned image]](https://arxiv.org/html/2608.07645v1/x35.png)Figure 6:Token costs of each operators for HGM and MGM evolved on Polyglot\.The distribution of token costs for all evaluation and expansion operations lies in similar order of magnitude\.
### 5\.3Ablation Study
To understand the contribution of each component in MGM, we conduct an ablation study on Polyglot\-60\. All variants are evaluated under the same computational budget of 200 evaluations, with two parallel workers enabled during evolution\. We use Qwen3\.6\-35B\-A3B as the backbone LLM for all settings\. Each variant starts from the same initial ancestor agent, which achieves 50\.8% accuracy on Polyglot\-60\. We compare the full MGM with two ablated variants: MGM withoutΦRM\\varPhi\_\{\\mathrm\{RM\}\}and withoutΦCH\\varPhi\_\{\\mathrm\{CH\}\}\. Appendix[E\.2](https://arxiv.org/html/2608.07645#A5.SS2)visualizes the corresponding evolution trees for the full model, HGM, and ablations\. As shown in TableLABEL:tab:mgm\_ablation, the full MGM achieves the best performance, and removingΦRM\\varPhi\_\{\\mathrm\{RM\}\}andΦCH\\varPhi\_\{\\mathrm\{CH\}\}leads to a clear performance degradation\. This suggests thatΦRM\\varPhi\_\{\\mathrm\{RM\}\}plays an important role in guiding the self\-improvement process toward more promising agents, and thatΦCH\\varPhi\_\{\\mathrm\{CH\}\}is even more critical for effective self\-improvement, likely because it helps preserve and reuse useful evolutionary information across iterations\. The ablated variants have comparable training hours, which also confirms that the gains of MGM come from the combined effect of its key components rather than from differences in computational cost\.
## 6Related Work
LLM agent systems and automated agent\-design methods study how prompts, tools, and workflows can be optimized as executable scaffolds\(Yao et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib43); Schick et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib27); Wu et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib39); Yang et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib41); Zhang et al\.,[2026b](https://arxiv.org/html/2608.07645#bib.bib45); Gao et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib10)\)\.Hu et al\. \([2025](https://arxiv.org/html/2608.07645#bib.bib13)\)andHong et al\. \([2024](https://arxiv.org/html/2608.07645#bib.bib12)\)show that meta\-agents can iteratively propose new agent designs from an archive of prior discoveries, yielding agents that transfer across tasks and models\. This line of work establishes agent scaffolding as a searchable program space, with a focus on discovering new agent designs\.
Self\-improving coding agents instantiate this idea in an inherited self\-modification setting\(Madaan et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib18); Shinn et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib32); Xia et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib40)\)\.Robeyns et al\. \([2025](https://arxiv.org/html/2608.07645#bib.bib26)\)show that a coding agent can evaluate itself, edit its own codebase, and improve over iterations\.Zhang et al\. \([2026a](https://arxiv.org/html/2608.07645#bib.bib44)\)extend this process into open\-ended evolution by maintaining an archive of self\-modified agents, whileWang et al\. \([2026](https://arxiv.org/html/2608.07645#bib.bib36)\)improve archive expansion by estimating the future self\-improvement potential of clades\. MGM follows this archive\-based self\-improvement setting, but targets a complementary bottleneck: improving the evidence used for each expansion\. By reusing existing trajectories across tasks and lineages, MGM improves each self\-modification step without requiring additional evaluations\.
Software\-engineering benchmarks such as SWE\-bench\(Jimenez et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib14)\), SWE\-bench Pro\(Deng et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib8)\), SWE\-bench Multilingual\(Khandpur & the SWE\-bench Team,[2025](https://arxiv.org/html/2608.07645#bib.bib16); Yang et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib42)\), and Polyglot\(Gauthier,[2024](https://arxiv.org/html/2608.07645#bib.bib11)\)evaluate repository\-level repair, long\-horizon reasoning, and cross\-language coding ability, testing whether coding agents can make workflow\-level improvements\.
A detailed discussion of additional related work on self\-evolving agents, runtime adaptation methods, and datasets is provided in Appendix[G](https://arxiv.org/html/2608.07645#A7)\.
## 7Conclusion
We introduce the Mendel Gödel Machine \(MGM\), an archive\-based self\-improving coding\-agent framework that conditions self\-modification on comparative evidence already present in evaluation trajectories\. Beyond clonal mutation, MGM adds reaction\-norm mutation and cross\-lineage hybridization, which reuse archived phenotypes across tasks and lineages without additional task evaluations\. Under an additive fitness landscape, theory and controlled simulations show that richer diagnostic evidence can raise effective fix probability and accelerate convergence relative to single\-trajectory baselines\. Experiments on SWE\-bench and Polyglot confirm MGM’s consistent gains over HGM in performance and efficiency under a matched budget\. Ablations show that both comparative operators contribute to these gains\. Held\-out evaluations further suggest that the evolved scaffolds transfer across benchmarks and backbone LLMs, indicating that richer comparative conditioning can discover reusable workflow\-level improvements rather than narrow task\-specific patches\. Within the constraints of scaffold‑level evolution under sandboxed, budgeted archive search, MGM potentially points to a compute‑efficient path for scalable self\-improving agent development, where scaffolds evolved on smaller and cheaper backbones can later be reused on stronger models\.
## Limitations
Self\-improving coding\-agent experiments remain expensive\. Even though MGM reuses existing trajectories for its comparative operators, evolution and evaluation on repository\-level tasks still consume substantial wall\-clock and GPU resources\. This constrains the number of independent evolution seeds and the breadth of hyperparameter sweeps we can report\.
MGM is history\-dependent\. Its advantage comes from comparative evidence in the archive\. Reaction\-norm mutation needs multiple trajectories from the same agent, and cross\-lineage hybridization needs overlapping tasks across lineages\. When the archive is small, task overlap is sparse, or informative failures have not yet appeared, MGM has little comparative evidence and may behave like single\-trajectory baselines\. The failed\-task pool partially mitigates this by increasing diagnostic overlap, but it cannot create informative contrasts before failures accumulate\.
MGM improves the evidence given to self\-modification, but it does not guarantee that the resulting edit is correct, general, or maintainable\. The Mendelian operators expose useful behavioral contrasts, yet the actual scaffold change is still produced by an LLM\-based editor\. High\-quality evidence therefore need not yield a high\-quality modification, and failed edits can waste budget\. Relatedly, comparative evidence helps only when the backbone can diagnose failure mechanisms from trajectories\. A stronger coding specialist with weaker general reasoning may still produce weaker self\-improvement under the same operators\.
Our formal analysis and Monte Carlo study use an additive fitness surrogate\. They isolate diagnostic compression under controlled assumptions and do not capture the full complexity of editable agent scaffolds\. Empirically, primary evolution is reported on fixed 60\-task subsets under a single matched budget, with broader checks on full Polyglot and held\-out SWE\-bench variants\. Subset selection and limited seed diversity leave residual uncertainty about variance across random restarts and alternative task samples\. Finally, our claims concern coding\-agent scaffolds evaluated on public software\-engineering benchmarks\. They do not establish that the same operators would transfer unchanged to non\-coding agents or open\-ended real\-world software maintenance\.
## Ethical Considerations
MGM advances self\-improving coding agents that can modify their own source code\. Such systems have a different risk profile from static models because an erroneous or adversarial self\-edit can persist in later descendants\. In our experiments, self\-modification and evaluation run inside isolated containers with no network access and with read\-only mounts of the host file system, following the safety protocol of prior archive\-based self\-improving agents\. Any deployment beyond this sandboxed coding setting requires a fresh review of the execution boundary\. Giving a self\-modifying agent access to production systems, external APIs, or writable storage outside controlled containers could allow unintended changes to spread in ways that are hard to audit or reverse\. Practitioners should retain the isolated execution model and subject any broader action space to explicit safety review before deployment\.
The same capability also has dual\-use implications\. Scaffold\-level self\-improvement can amplify both beneficial coding assistance and harmful automation, including generation of malicious software or exploitation of vulnerable systems, if isolation is removed\. Our public release is intended for research on sandboxed self\-evolution\. We discourage unconstrained deployment of evolved scaffolds outside controlled environments\.
The benchmarks and backbone models used here may also inherit harmful or biased content\. SWE\-bench, SWE\-bench Pro, and Polyglot draw on public open\-source repositories, and the backbone LLMs are trained on large web corpora, all of which may contain insecure code patterns and may overrepresent particular languages, ecosystems, or problem domains\. We did not introduce dedicated bias or safety audits of the evolved agents, so inherited fairness and security issues may persist\. In addition, public benchmark material may overlap with pretraining data, which limits how strongly offline scores should be read as evidence of robust real\-world competence\. This work focuses on self\-evolution of coding agents only, and our evolved agents are not equipped to act outside software\-engineering related/ repositories\. Task\-specific safety evaluation remains necessary before any use beyond general\-purpose coding assistance\.
## Acknowledgments
The authors gratefully acknowledge the scientific support and HPC resources provided by the hessian\.AI Service Center \(funded by the Federal Ministry of Research, Technology and Space, BMFTR, grant no\. 16IS22091\), the hessian\.AI Innovation Lab \(funded by the Hessian Ministry for Digital Strategy and Innovation, grant no\. S\-DIW04/0013/003\), and the Karlsruhe Institute of Technology National High Performance Computing Center \(NHR@KIT\) under the NHR projects 22560, and 24767\. The HoreKa supercomputer at NHR@KIT is funded by the Ministry of Science, Research and the Arts Baden\-Württemberg and by the Federal Ministry of Education and Research of Germany\. This work also receives support from the Munich Center for Machine Learning \(MCML\) and the Program of China Scholarship Council \(Grant No\.202508080292\)\. The funding bodies had no role in the design of methodology and experiments, the analysis and interpretation of results, or the writing of manuscript\.
## References
- Austin et al\. \(2021\)Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton\.Program synthesis with large language models, 2021\.URL[https://arxiv\.org/abs/2108\.07732](https://arxiv.org/abs/2108.07732)\.
- Cai et al\. \(2024\)Tianle Cai, Xuezhi Wang, Tengyu Ma, Xinyun Chen, and Denny Zhou\.Large language models as tool makers\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=qV83K9d5WB](https://openreview.net/forum?id=qV83K9d5WB)\.
- Chen et al\. \(2021\)Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert\-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N\. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba\.Evaluating large language models trained on code, 2021\.URL[https://arxiv\.org/abs/2107\.03374](https://arxiv.org/abs/2107.03374)\.
- Clune \(2020\)Jeff Clune\.Ai\-gas: Ai\-generating algorithms, an alternate paradigm for producing general artificial intelligence, 2020\.URL[https://arxiv\.org/abs/1905\.10985](https://arxiv.org/abs/1905.10985)\.
- Cully \(2021\)Antoine Cully\.Multi\-emitter map\-elites: improving quality, diversity and data efficiency with heterogeneous sets of emitters\.In*Proceedings of the Genetic and Evolutionary Computation Conference*, GECCO ’21, pp\. 84–92\. ACM, 2021\.doi:10\.1145/3449639\.3459326\.URL[http://dx\.doi\.org/10\.1145/3449639\.3459326](http://dx.doi.org/10.1145/3449639.3459326)\.
- DeepSeek\-AI et al\. \(2025\)DeepSeek\-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H\. Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoran Wei, Haowei Zhang, Haowen Luo, Haozhe Ji, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, Jialiang Huang, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jingchang Chen, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexin Huang, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Liang Zhao, Liangsheng Yin, Lihua Guo, Lingxiao Luo, Linwang Ma, Litong Wang, Liyue Zhang, M\. S\. Di, M\. Y Xu, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Panpan Huang, Peixin Cong, Peiyi Wang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, S\. H\. Liu, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Songyang Zhou, Tao Ni, Tao Yun, Tian Pei, Tian Ye, Tianyuan Yue, Wangding Zeng, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjun Gao, Wentao Zhang, Xi Gao, Xiangwen Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyuan Li, Xu Chen, Xuecheng Su, Xuehai Pan, Xuheng Lin, Xuwei Fu, Y\. Q\. Wang, Yang Zhang, Yanhong Xu, Yanru Ma, Yao Li, Yao Li, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yiliang Xiong, Ying He, Ying Zhou, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuduan Wang, Yue Gong, Yuhan Wu, Yuheng Zou, Yukun Li, Yunfan Xiong, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuyang Zhou, Z\. F\. Wu, Z\. Z\. Ren, Zehua Zhao, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhiyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Zizheng Pan, Zongqing Yao, Bei Feng, Hui Li, J\. L\. Cai, Jiaqi Ni, Lei Xu, Meng Li, Ning Tian, R\. J\. Chen, R\. L\. Jin, S\. S\. Li, Shuang Zhou, Tianyu Sun, X\. Q\. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xinnan Song, Xinyi Zhou, Y\. X\. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, Dongjie Ji, Jian Liang, Jianzhong Guo, Jin Chen, Leyi Xia, Miaojun Wang, Mingming Li, Peng Zhang, Ruyi Chen, Shangmian Sun, Shaoqing Wu, Shengfeng Ye, T\. Wang, W\. L\. Xiao, Wei An, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Ying Tang, Yukun Zha, Zekai Zhang, Zhe Ju, Zhen Zhang, and Zihua Qu\.Deepseek\-v3\.2: Pushing the frontier of open large language models, 2025\.URL[https://arxiv\.org/abs/2512\.02556](https://arxiv.org/abs/2512.02556)\.
- DeepSeek\-AI et al\. \(2026\)DeepSeek\-AI, Anyi Xu, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chenchen Ling, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chengyu Hou, Chenhao Xu, Chenze Shao, Chong Ruan, Conner Sun, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Donghao Li, Dongjie Ji, Erhang Li, Fang Wei, Fangyun Lin, Fangzhou Yuan, Feiyu Xia, Fucong Dai, Guangbo Hao, Guanting Chen, Guoai Cao, Guolai Meng, Guowei Li, Han Yu, Han Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoling Zhang, Haoming Luo, Haoran Wei, Haotian Yuan, Haowei Zhang, Haowen Luo, Haoyu Chen, Haozhe Ji, Hengqing Zhang, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, J Yang, JQ Zhu, Jia Luo, Jia Song, Jia Yu, Jialiang Huang, Jialu Cai, Jian Liang, Jiangting Zhou, Jiasheng Ye, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jieyu Yang, Jin Chen, Jin Yan, Jingchang Chen, Jingli Zhou, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jingzi Zhou, Jinhua Zhu, Jiping Yu, Joseph Sun, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junmin Zheng, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Leyi Xia, Li Zhang, Liang Zhao, Lihua Guo, Lingxiao Luo, Linwang Ma, Linyan Zhu, Litong Wang, Liyu Cai, Liyue Zhang, Longhao Chen, MS Di, MY Xu, Max Mei, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Mingxu Zhou, Minmin Han, Ning Wang, Panpan Huang, Panpan Wang, Peixin Cong, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Qiwei Jiang, Rui Tian, Ruifan Xu, Ruijie Lu, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqian Chen, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, Ruyi Chen, SH Liu, Shanghao Lu, Shangmian Sun, Shangyan Zhou, Shanhuang Chen, Shaofei Cai, Shaoheng Nie, Shaoqing Wu, Shaoyuan Chen, Shengding Hu, Shengyu Liu, Shiqiang Hu, Shirong Ma, Shiyu Wang, Shuiping Yu, Shunfeng Zhou, Shuting Pan, Shuying Yu, Songyang Zhou, Tao Ni, Tao Yun, Tian Jin, Tian Pei, Tian Ye, Tianle Lin, Tianran Ji, Tianyi Cui, Tianyuan Yue, Tingting Yu, Tun Wang, W Zhang, WL Xiao, Wangding Zeng, Wei An, Weilin Zhao, Wen Liu, Wenfeng Liang, Wenjie Pang, Wenjing Luo, Wenjing Yao, Wenjun Gao, Wenkai Yang, Wenlve Huang, Wenqing Hou, Wentao Zhang, Wenting Ma, Xi Gao, Xiang He, Xiangwen Wang, Xianzu Wang, Xiao Bi, Xiaodong Liu, Xiaohan Wang, Xiaokang Chen, Xiaokang Zhang, Xiaotao Nie, Xiaowen Sun, Xiaoxiang Wang, Xin Cheng, Xin Liu, Xin Xie, Xingchao Liu, Xingchen Liu, Xingkai Yu, Xingyou Li, Xinyu Yang, Xinyu Zhang, Xu Chen, Xuanyu Wang, Xuecheng Su, Xueyin Chen, Xuheng Lin, Xuwei Fu, YC Yan, YQ Wang, YW Ma, Yanfeng Luo, Yang Zhang, Yanhong Xu, Yanru Ma, Yanwen Huang, Yao Li, Yao Li, Yao Xu, Yao Zhao, Yaofeng Sun, Yaohui Wang, Yi Qian, Yi Shao, Yi Yu, Yichao Zhang, Yifan Ding, Yifan Shi, Yijia Wu, Yiliang Xiong, Yiling Ma, Ying He, Ying Tang, Ying Zhou, Yingjia Luo, Yinmin Zhong, Yishi Piao, Yisong Wang, Yixiang Zhang, Yixiao Chen, Yixuan Tan, Yixuan Wei, Yiyang Ma, Yiyuan Liu, Yonglun Yang, Yongqiang Guo, Yongtong Wu, Yu Wu, YuKun Li, Yuan Cheng, Yuan Ou, Yuanfan Xu, Yuanhao Li, Yuduan Wang, Yuehan Yang, Yuer Xu, Yuhan Wu, Yuhao Meng, Yuheng Zou, Yukun Zha, Yunfan Xiong, Yupeng Chen, Yuping Lin, Yuqian Cao, Yuqian Wang, Yushun Zhang, Yuting Yan, Yutong Lin, Yuxian Gu, Yuxiang Luo, Yuxiang You, Yuxuan Liu, Yuxuan Zhou, Yuyang Zhou, Yuzhen Huang, ZF Wu, Zehao Wang, Zehua Zhao, Zehui Ren, Zekai Zhang, Zhangli Sha, Zhe Fu, Zhe Ju, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zheren Gao, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhixuan Chen, Zhiyu Wu, Zhizhou Ren, Zhongyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihua Qu, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Ziyi Wan, Zizheng Pan, and Zongqing Yao\.Deepseek\-v4: Towards highly efficient million\-token context intelligence, 2026\.URL[https://arxiv\.org/abs/2606\.19348](https://arxiv.org/abs/2606.19348)\.
- Deng et al\. \(2025\)Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Nitin Pasari, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy, Sean Hendryx, Zifan Wang, Vijay Bharadwaj, Jeff Holm, Raja Aluri, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler\.Swe\-bench pro: Can ai agents solve long\-horizon software engineering tasks?, 2025\.URL[https://arxiv\.org/abs/2509\.16941](https://arxiv.org/abs/2509.16941)\.
- Ellenberg et al\. \(2025\)Jordan S\. Ellenberg, Cristofero S\. Fraser\-Taliente, Thomas R\. Harvey, Karan Srivastava, and Andrew V\. Sutherland\.Generative modeling for mathematical discovery, 2025\.URL[https://arxiv\.org/abs/2503\.11061](https://arxiv.org/abs/2503.11061)\.
- Gao et al\. \(2026\)Huan\-ang Gao, Jiayi Geng, Wenyue Hua, Mengkang Hu, Xinzhe Juan, Hongzhang Liu, Shilong Liu, Jiahao Qiu, Xuan Qi, Qihan Ren, Yiran Wu, Hongru WANG, Han Xiao, Yuhang Zhou, Shaokun Zhang, Jiayi Zhang, Jinyu Xiang, Yixiong Fang, Qiwen Zhao, Dongrui Liu, Cheng Qian, Zhenhailong Wang, Minda Hu, Huazheng Wang, Qingyun Wu, Heng Ji, and Mengdi Wang\.A survey of self\-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence\.*Transactions on Machine Learning Research*, 2026\.ISSN 2835\-8856\.URL[https://openreview\.net/forum?id=CTr3bovS5F](https://openreview.net/forum?id=CTr3bovS5F)\.Survey Certification\.
- Gauthier \(2024\)Paul Gauthier\.o1 tops aider’s new polyglot leaderboard\.[https://aider\.chat/2024/12/21/polyglot\.html](https://aider.chat/2024/12/21/polyglot.html), December 2024\.Accessed: 2026\-05\-22\.
- Hong et al\. \(2024\)Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber\.MetaGPT: Meta programming for a multi\-agent collaborative framework\.In*The Twelfth International Conference on Learning Representations*, 2024\.URL[https://openreview\.net/forum?id=VtmBAGCN7o](https://openreview.net/forum?id=VtmBAGCN7o)\.
- Hu et al\. \(2025\)Shengran Hu, Cong Lu, and Jeff Clune\.Automated design of agentic systems\.In Y\. Yue, A\. Garg, N\. Peng, F\. Sha, and R\. Yu \(eds\.\),*International Conference on Learning Representations*, volume 2025, pp\. 21344–21377, 2025\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2025/file/36b7acf6f6010652b3f2a433774a66fe\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2025/file/36b7acf6f6010652b3f2a433774a66fe-Paper-Conference.pdf)\.
- Jimenez et al\. \(2024\)Carlos E Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan\.Swe\-bench: Can language models resolve real\-world github issues?In B\. Kim, Y\. Yue, S\. Chaudhuri, K\. Fragkiadaki, M\. Khan, and Y\. Sun \(eds\.\),*International Conference on Learning Representations*, volume 2024, pp\. 54107–54157, 2024\.URL[https://proceedings\.iclr\.cc/paper\_files/paper/2024/file/edac78c3e300629acfe6cbe9ca88fb84\-Paper\-Conference\.pdf](https://proceedings.iclr.cc/paper_files/paper/2024/file/edac78c3e300629acfe6cbe9ca88fb84-Paper-Conference.pdf)\.
- Karpas et al\. \(2022\)Ehud Karpas, Omri Abend, Yonatan Belinkov, Barak Lenz, Opher Lieber, Nir Ratner, Yoav Shoham, Hofit Bata, Yoav Levine, Kevin Leyton\-Brown, Dor Muhlgay, Noam Rozen, Erez Schwartz, Gal Shachaf, Shai Shalev\-Shwartz, Amnon Shashua, and Moshe Tenenholtz\.Mrkl systems: A modular, neuro\-symbolic architecture that combines large language models, external knowledge sources and discrete reasoning, 2022\.URL[https://arxiv\.org/abs/2205\.00445](https://arxiv.org/abs/2205.00445)\.
- Khandpur & the SWE\-bench Team \(2025\)Kabir Khandpur and the SWE\-bench Team\.SWE\-bench Multilingual\.[https://www\.swebench\.com/multilingual\.html](https://www.swebench.com/multilingual.html), 2025\.300 tasks across 9 languages and 42 repositories\. Official citation guidance points toyang2025swesmith\.
- Li \(2026\)Bojie Li\.Incompressible knowledge probes: Estimating black\-box llm parameter counts via factual capacity, 2026\.URL[https://arxiv\.org/abs/2604\.24827](https://arxiv.org/abs/2604.24827)\.
- Madaan et al\. \(2023\)Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark\.Self\-refine: Iterative refinement with self\-feedback\.In A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(eds\.\),*Advances in Neural Information Processing Systems*, volume 36, pp\. 46534–46594\. Curran Associates, Inc\., 2023\.doi:10\.52202/075280\-2019\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/91edff07232fb1b55a505a9e9f6c0ff3\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/91edff07232fb1b55a505a9e9f6c0ff3-Paper-Conference.pdf)\.
- MiniMax et al\. \(2025\)MiniMax, :, Aili Chen, Aonian Li, Bangwei Gong, Binyang Jiang, Bo Fei, Bo Yang, Boji Shan, Changqing Yu, Chao Wang, Cheng Zhu, Chengjun Xiao, Chengyu Du, Chi Zhang, Chu Qiao, Chunhao Zhang, Chunhui Du, Congchao Guo, Da Chen, Deming Ding, Dianjun Sun, Dong Li, Enwei Jiao, Haigang Zhou, Haimo Zhang, Han Ding, Haohai Sun, Haoyu Feng, Huaiguang Cai, Haichao Zhu, Jian Sun, Jiaqi Zhuang, Jiaren Cai, Jiayuan Song, Jin Zhu, Jingyang Li, Jinhao Tian, Jinli Liu, Junhao Xu, Junjie Yan, Junteng Liu, Junxian He, Kaiyi Feng, Ke Yang, Kecheng Xiao, Le Han, Leyang Wang, Lianfei Yu, Liheng Feng, Lin Li, Lin Zheng, Linge Du, Lingyu Yang, Lunbin Zeng, Minghui Yu, Mingliang Tao, Mingyuan Chi, Mozhi Zhang, Mujie Lin, Nan Hu, Nongyu Di, Peng Gao, Pengfei Li, Pengyu Zhao, Qibing Ren, Qidi Xu, Qile Li, Qin Wang, Rong Tian, Ruitao Leng, Shaoxiang Chen, Shaoyu Chen, Shengmin Shi, Shitong Weng, Shuchang Guan, Shuqi Yu, Sichen Li, Songquan Zhu, Tengfei Li, Tianchi Cai, Tianrun Liang, Weiyu Cheng, Weize Kong, Wenkai Li, Xiancai Chen, Xiangjun Song, Xiao Luo, Xiao Su, Xiaobo Li, Xiaodong Han, Xinzhu Hou, Xuan Lu, Xun Zou, Xuyang Shen, Yan Gong, Yan Ma, Yang Wang, Yiqi Shi, Yiran Zhong, Yonghong Duan, Yongxiang Fu, Yongyi Hu, Yu Gao, Yuanxiang Fan, Yufeng Yang, Yuhao Li, Yulin Hu, Yunan Huang, Yunji Li, Yunzhi Xu, Yuxin Mao, Yuxuan Shi, Yuze Wenren, Zehan Li, Zelin Li, Zhanxu Tian, Zhengmao Zhu, Zhenhua Fan, Zhenzhen Wu, Zhichao Xu, Zhihang Yu, Zhiheng Lyu, Zhuo Jiang, Zibo Gao, Zijia Wu, Zijian Song, and Zijun Sun\.Minimax\-m1: Scaling test\-time compute efficiently with lightning attention, 2025\.URL[https://arxiv\.org/abs/2506\.13585](https://arxiv.org/abs/2506.13585)\.
- Mouret & Clune \(2015\)Jean\-Baptiste Mouret and Jeff Clune\.Illuminating search spaces by mapping elites, 2015\.URL[https://arxiv\.org/abs/1504\.04909](https://arxiv.org/abs/1504.04909)\.
- Novikov et al\. \(2025\)Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po\-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J\. R\. Ruiz, Abbas Mehrabian, M\. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, and Matej Balog\.Alphaevolve: A coding agent for scientific and algorithmic discovery, 2025\.URL[https://arxiv\.org/abs/2506\.13131](https://arxiv.org/abs/2506.13131)\.
- Pigliucci \(2001\)Massimo Pigliucci\.*Phenotypic Plasticity: Beyond Nature and Nurture*\.Johns Hopkins University Press, 2001\.
- Qian et al\. \(2023\)Cheng Qian, Chi Han, Yi Fung, Yujia Qin, Zhiyuan Liu, and Heng Ji\.CREATOR: Tool creation for disentangling abstract and concrete reasoning of large language models\.In Houda Bouamor, Juan Pino, and Kalika Bali \(eds\.\),*Findings of the Association for Computational Linguistics: EMNLP 2023*, pp\. 6922–6939, Singapore, December 2023\. Association for Computational Linguistics\.doi:10\.18653/v1/2023\.findings\-emnlp\.462\.URL[https://aclanthology\.org/2023\.findings\-emnlp\.462/](https://aclanthology.org/2023.findings-emnlp.462/)\.
- Qiu et al\. \(2025\)Jiahao Qiu, Xuan Qi, Tongcheng Zhang, Xinzhe Juan, Jiacheng Guo, Yifu Lu, Yimin Wang, Zixin Yao, Qihan Ren, Xun Jiang, Xing Zhou, Dongrui Liu, Ling Yang, Yue Wu, Kaixuan Huang, Shilong Liu, Hongru Wang, and Mengdi Wang\.Alita: Generalist agent enabling scalable agentic reasoning with minimal predefinition and maximal self\-evolution, 2025\.URL[https://arxiv\.org/abs/2505\.20286](https://arxiv.org/abs/2505.20286)\.
- Qwen Team \(2026\)Qwen Team\.Qwen3\.6\-35B\-A3B: Agentic coding power, now open to all, April 2026\.URL[https://qwen\.ai/blog?id=qwen3\.6\-35b\-a3b](https://qwen.ai/blog?id=qwen3.6-35b-a3b)\.
- Robeyns et al\. \(2025\)Maxime Robeyns, Martin Szummer, and Laurence Aitchison\.A self\-improving coding agent\.In*Scaling Self\-Improving Foundation Models without Human Supervision*, 2025\.URL[https://openreview\.net/forum?id=rShJCyLsOr](https://openreview.net/forum?id=rShJCyLsOr)\.
- Schick et al\. \(2023\)Timo Schick, Jane Dwivedi\-Yu, Roberto Dessi, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom\.Toolformer: Language models can teach themselves to use tools\.In*Thirty\-seventh Conference on Neural Information Processing Systems*, 2023\.URL[https://openreview\.net/forum?id=Yacmpz84TH](https://openreview.net/forum?id=Yacmpz84TH)\.
- Schmidhuber \(1987\)Jurgen Schmidhuber\.Evolutionary principles in self\-referential learning\. on learning now to learn: The meta\-meta\-meta…\-hook\.Diploma thesis, Technische Universitat Munchen, Germany, 1987\.URL[https://mediatum\.ub\.tum\.de/?id=813180](https://mediatum.ub.tum.de/?id=813180)\.
- Schmidhuber \(2003\)Jürgen Schmidhuber\.Goedel machines: Self\-referential universal problem solvers making provably optimal self\-improvements\.*CoRR*, cs\.LO/0309048, 2003\.URL[http://arxiv\.org/abs/cs/0309048](http://arxiv.org/abs/cs/0309048)\.
- Schmidhuber \(2007\)Jürgen Schmidhuber\.*Gödel Machines: Fully Self\-referential Optimal Universal Self\-improvers*, pp\. 199–226\.Springer Berlin Heidelberg, Berlin, Heidelberg, 2007\.ISBN 978\-3\-540\-68677\-4\.doi:10\.1007/978\-3\-540\-68677\-4\_7\.URL[https://doi\.org/10\.1007/978\-3\-540\-68677\-4\_7](https://doi.org/10.1007/978-3-540-68677-4_7)\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y\. K\. Li, Y\. Wu, and Daya Guo\.Deepseekmath: Pushing the limits of mathematical reasoning in open language models, 2024\.URL[https://arxiv\.org/abs/2402\.03300](https://arxiv.org/abs/2402.03300)\.
- Shinn et al\. \(2023\)Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\.Reflexion: language agents with verbal reinforcement learning\.In A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(eds\.\),*Advances in Neural Information Processing Systems*, volume 36, pp\. 8634–8652\. Curran Associates, Inc\., 2023\.doi:10\.52202/075280\-0377\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2023/file/1b44b878bb782e6954cd888628510e90\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2023/file/1b44b878bb782e6954cd888628510e90-Paper-Conference.pdf)\.
- Singh et al\. \(2026\)Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El\-Kishky, Aidan McLaughlin, Aiden Low, AJ Ostrow, Akhila Ananthram, Akshay Nathan, Alan Luo, Alec Helyar, Aleksander Madry, Aleksandr Efremov, Aleksandra Spyra, Alex Baker\-Whitcomb, Alex Beutel, Alex Karpenko, Alex Makelov, Alex Neitz, Alex Wei, Alexandra Barr, Alexandre Kirchmeyer, Alexey Ivanov, Alexi Christakis, Alistair Gillespie, Allison Tam, Ally Bennett, Alvin Wan, Alyssa Huang, Amy McDonald Sandjideh, Amy Yang, Ananya Kumar, Andre Saraiva, Andrea Vallone, Andrei Gheorghe, Andres Garcia Garcia, Andrew Braunstein, Andrew Liu, Andrew Schmidt, Andrey Mereskin, Andrey Mishchenko, Andy Applebaum, Andy Rogerson, Ann Rajan, Annie Wei, Anoop Kotha, Anubha Srivastava, Anushree Agrawal, Arun Vijayvergiya, Ashley Tyra, Ashvin Nair, Avi Nayak, Ben Eggers, Bessie Ji, Beth Hoover, Bill Chen, Blair Chen, Boaz Barak, Borys Minaiev, Botao Hao, Bowen Baker, Brad Lightcap, Brandon McKinzie, Brandon Wang, Brendan Quinn, Brian Fioca, Brian Hsu, Brian Yang, Brian Yu, Brian Zhang, Brittany Brenner, Callie Riggins Zetino, Cameron Raymond, Camillo Lugaresi, Carolina Paz, Cary Hudson, Cedric Whitney, Chak Li, Charles Chen, Charlotte Cole, Chelsea Voss, Chen Ding, Chen Shen, Chengdu Huang, Chris Colby, Chris Hallacy, Chris Koch, Chris Lu, Christina Kaplan, Christina Kim, CJ Minott\-Henriques, Cliff Frey, Cody Yu, Coley Czarnecki, Colin Reid, Colin Wei, Cory Decareaux, Cristina Scheau, Cyril Zhang, Cyrus Forbes, Da Tang, Dakota Goldberg, Dan Roberts, Dana Palmie, Daniel Kappler, Daniel Levine, Daniel Wright, Dave Leo, David Lin, David Robinson, Declan Grabb, Derek Chen, Derek Lim, Derek Salama, Dibya Bhattacharjee, Dimitris Tsipras, Dinghua Li, Dingli Yu, DJ Strouse, Drew Williams, Dylan Hunn, Ed Bayes, Edwin Arbus, Ekin Akyurek, Elaine Ya Le, Elana Widmann, Eli Yani, Elizabeth Proehl, Enis Sert, Enoch Cheung, Eri Schwartz, Eric Han, Eric Jiang, Eric Mitchell, Eric Sigler, Eric Wallace, Erik Ritter, Erin Kavanaugh, Evan Mays, Evgenii Nikishin, Fangyuan Li, Felipe Petroski Such, Filipe de Avila Belbute Peres, Filippo Raso, Florent Bekerman, Foivos Tsimpourlas, Fotis Chantzis, Francis Song, Francis Zhang, Gaby Raila, Garrett McGrath, Gary Briggs, Gary Yang, Giambattista Parascandolo, Gildas Chabot, Grace Kim, Grace Zhao, Gregory Valiant, Guillaume Leclerc, Hadi Salman, Hanson Wang, Hao Sheng, Haoming Jiang, Haoyu Wang, Haozhun Jin, Harshit Sikchi, Heather Schmidt, Henry Aspegren, Honglin Chen, Huida Qiu, Hunter Lightman, Ian Covert, Ian Kivlichan, Ian Silber, Ian Sohl, Ibrahim Hammoud, Ignasi Clavera, Ikai Lan, Ilge Akkaya, Ilya Kostrikov, Irina Kofman, Isak Etinger, Ishaan Singal, Jackie Hehir, Jacob Huh, Jacqueline Pan, Jake Wilczynski, Jakub Pachocki, James Lee, James Quinn, Jamie Kiros, Janvi Kalra, Jasmyn Samaroo, Jason Wang, Jason Wolfe, Jay Chen, Jay Wang, Jean Harb, Jeffrey Han, Jeffrey Wang, Jennifer Zhao, Jeremy Chen, Jerene Yang, Jerry Tworek, Jesse Chand, Jessica Landon, Jessica Liang, Ji Lin, Jiancheng Liu, Jianfeng Wang, Jie Tang, Jihan Yin, Joanne Jang, Joel Morris, Joey Flynn, Johannes Ferstad, Johannes Heidecke, John Fishbein, John Hallman, Jonah Grant, Jonathan Chien, Jonathan Gordon, Jongsoo Park, Jordan Liss, Jos Kraaijeveld, Joseph Guay, Joseph Mo, Josh Lawson, Josh McGrath, Joshua Vendrow, Joy Jiao, Julian Lee, Julie Steele, Julie Wang, Junhua Mao, Kai Chen, Kai Hayashi, Kai Xiao, Kamyar Salahi, Kan Wu, Karan Sekhri, Karan Sharma, Karan Singhal, Karen Li, Kenny Nguyen, Keren Gu\-Lemberg, Kevin King, Kevin Liu, Kevin Stone, Kevin Yu, Kristen Ying, Kristian Georgiev, Kristie Lim, Kushal Tirumala, Kyle Miller, Lama Ahmad, Larry Lv, Laura Clare, Laurance Fauconnet, Lauren Itow, Lauren Yang, Laurentia Romaniuk, Leah Anise, Lee Byron, Leher Pathak, Leon Maksin, Leyan Lo, Leyton Ho, Li Jing, Liang Wu, Liang Xiong, Lien Mamitsuka, Lin Yang, Lindsay McCallum, Lindsey Held, Liz Bourgeois, Logan Engstrom, Lorenz Kuhn, Louis Feuvrier, Lu Zhang, Lucas Switzer, Lukas Kondraciuk, Lukasz Kaiser, Manas Joglekar, Mandeep Singh, Mandip Shah, Manuka Stratta, Marcus Williams, Mark Chen, Mark Sun, Marselus Cayton, Martin Li, Marvin Zhang, Marwan Aljubeh, Matt Nichols, Matthew Haines, Max Schwarzer, Mayank Gupta, Meghan Shah, Melody Y\. Guan, Melody Huang, Meng Dong, Mengqing Wang, Mia Glaese, Micah Carroll, Michael Lampe, Michael Malek, Michael Sharman, Michael Zhang, Michele Wang, Michelle Pokrass, Mihai Florian, Mikhail Pavlov, Miles Wang, Ming Chen, Mingxuan Wang, Minnia Feng, Mo Bavarian, Molly Lin, Moose Abdool, Mostafa Rohaninejad, Nacho Soto, Natalie Staudacher, Natan LaFontaine, Nathan Marwell, Nelson Liu, Nick Preston, Nick Turley, Nicklas Ansman, Nicole Blades, Nikil Pancha, Nikita Mikhaylin, Niko Felix, Nikunj Handa, Nishant Rai, Nitish Keskar, Noam Brown, Ofir Nachum, Oleg Boiko, Oleg Murk, Olivia Watkins, Oona Gleeson, Pamela Mishkin, Patryk Lesiewicz, Paul Baltescu, Pavel Belov, Peter Zhokhov, Philip Pronin, Phillip Guo, Phoebe Thacker, Qi Liu, Qiming Yuan, Qinghua Liu, Rachel Dias, Rachel Puckett, Rahul Arora, Ravi Teja Mullapudi, Raz Gaon, Reah Miyara, Rennie Song, Rishabh Aggarwal, RJ Marsan, Robel Yemiru, Robert Xiong, Rohan Kshirsagar, Rohan Nuttall, Roman Tsiupa, Ronen Eldan, Rose Wang, Roshan James, Roy Ziv, Rui Shu, Ruslan Nigmatullin, Saachi Jain, Saam Talaie, Sam Altman, Sam Arnesen, Sam Toizer, Sam Toyer, Samuel Miserendino, Sandhini Agarwal, Sarah Yoo, Savannah Heon, Scott Ethersmith, Sean Grove, Sean Taylor, Sebastien Bubeck, Sever Banesiu, Shaokyi Amdo, Shengjia Zhao, Sherwin Wu, Shibani Santurkar, Shiyu Zhao, Shraman Ray Chaudhuri, Shreyas Krishnaswamy, Shuaiqi, Xia, Shuyang Cheng, Shyamal Anadkat, Simón Posada Fishman, Simon Tobin, Siyuan Fu, Somay Jain, Song Mei, Sonya Egoian, Spencer Kim, Spug Golden, SQ Mah, Steph Lin, Stephen Imm, Steve Sharpe, Steve Yadlowsky, Sulman Choudhry, Sungwon Eum, Suvansh Sanjeev, Tabarak Khan, Tal Stramer, Tao Wang, Tao Xin, Tarun Gogineni, Taya Christianson, Ted Sanders, Tejal Patwardhan, Thomas Degry, Thomas Shadwell, Tianfu Fu, Tianshi Gao, Timur Garipov, Tina Sriskandarajah, Toki Sherbakov, Tomek Korbak, Tomer Kaftan, Tomo Hiratsuka, Tongzhou Wang, Tony Song, Tony Zhao, Troy Peterson, Val Kharitonov, Victoria Chernova, Vineet Kosaraju, Vishal Kuo, Vitchyr Pong, Vivek Verma, Vlad Petrov, Wanning Jiang, Weixing Zhang, Wenda Zhou, Wenlei Xie, Wenting Zhan, Wes McCabe, Will DePue, Will Ellsworth, Wulfie Bain, Wyatt Thompson, Xiangning Chen, Xiangyu Qi, Xin Xiang, Xinwei Shi, Yann Dubois, Yaodong Yu, Yara Khakbaz, Yifan Wu, Yilei Qian, Yin Tat Lee, Yinbo Chen, Yizhen Zhang, Yizhong Xiong, Yonglong Tian, Young Cha, Yu Bai, Yu Yang, Yuan Yuan, Yuanzhi Li, Yufeng Zhang, Yuguang Yang, Yujia Jin, Yun Jiang, Yunyun Wang, Yushi Wang, Yutian Liu, Zach Stubenvoll, Zehao Dou, Zheng Wu, and Zhigang Wang\.Openai gpt\-5 system card, 2026\.URL[https://arxiv\.org/abs/2601\.03267](https://arxiv.org/abs/2601.03267)\.
- Team et al\. \(2026a\)Kimi Team, Tongtong Bai, Yifan Bai, Yiping Bao, S\. H\. Cai, Yuan Cao, Y\. Charles, H\. S\. Che, Cheng Chen, Guanduo Chen, Huarong Chen, Jia Chen, Jiahao Chen, Jianlong Chen, Jun Chen, Kefan Chen, Liang Chen, Ruijue Chen, Xinhao Chen, Yanru Chen, Yanxu Chen, Yicun Chen, Yimin Chen, Yingjiang Chen, Yuankun Chen, Yujie Chen, Yutian Chen, Zhirong Chen, Ziwei Chen, Dazhi Cheng, Minghan Chu, Jialei Cui, Jiaqi Deng, Muxi Diao, Hao Ding, Mengfan Dong, Mengnan Dong, Yuxin Dong, Yuhao Dong, Angang Du, Chenzhuang Du, Dikang Du, Lingxiao Du, Yulun Du, Yu Fan, Shengjun Fang, Qiulin Feng, Yichen Feng, Garimugai Fu, Kelin Fu, Hongcheng Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Chengyang Gong, Xiaochen Gong, Zhuoma Gongque, Qizheng Gu, Xinran Gu, Yicheng Gu, Longyu Guan, Yuanying Guo, Xiaoru Hao, Weiran He, Wenyang He, Yunjia He, Chao Hong, Hao Hu, Jiaxi Hu, Yangyang Hu, Zhenxing Hu, Ke Huang, Ruiyuan Huang, Weixiao Huang, Zhiqi Huang, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yu Jing, Guokun Lai, Aidi Li, C\. Li, Cheng Li, Fang Li, Guanghe Li, Guanyu Li, Haitao Li, Haoyang Li, Jia Li, Jingwei Li, Junxiong Li, Lincan Li, Mo Li, Weihong Li, Wentao Li, Xinhang Li, Xinhao Li, Yang Li, Yanhao Li, Yiwei Li, Yuxiao Li, Zhaowei Li, Zheming Li, Weilong Liao, Jiawei Lin, Xiaohan Lin, Zhishan Lin, Zichao Lin, Cheng Liu, Chenyu Liu, Hongzhang Liu, Liang Liu, Shaowei Liu, Shudong Liu, Shuran Liu, Tianwei Liu, Tianyu Liu, Weizhou Liu, Xiangyan Liu, Yangyang Liu, Yanming Liu, Yibo Liu, Yuanxin Liu, Yue Liu, Zhengying Liu, Zhongnuo Liu, Enzhe Lu, Haoyu Lu, Zhiyuan Lu, Junyu Luo, Tongxu Luo, Yashuo Luo, Long Ma, Yingwei Ma, Shaoguang Mao, Yuan Mei, Xin Men, Fanqing Meng, Zhiyong Meng, Yibo Miao, Minqing Ni, Kun Ouyang, Siyuan Pan, Bo Pang, Yuchao Qian, Ruoyu Qin, Zeyu Qin, Jiezhong Qiu, Bowen Qu, Zeyu Shang, Youbo Shao, Tianxiao Shen, Zhennan Shen, Juanfeng Shi, Lidong Shi, Shengyuan Shi, Feifan Song, Pengwei Song, Tianhui Song, Xiaoxi Song, Hongjin Su, Jianlin Su, Zhaochen Su, Lin Sui, Jinsong Sun, Junyao Sun, Tongyu Sun, Flood Sung, Yunpeng Tai, Chuning Tang, Heyi Tang, Xiaojuan Tang, Zhengyang Tang, Jiawen Tao, Shiyuan Teng, Chaoran Tian, Pengfei Tian, Ao Wang, Bowen Wang, Chensi Wang, Chuang Wang, Congcong Wang, Dingkun Wang, Dinglu Wang, Dongliang Wang, Feng Wang, Hailong Wang, Haiming Wang, Hengzhi Wang, Huaqing Wang, Hui Wang, Jiahao Wang, Jinhong Wang, Jiuzheng Wang, Kaixin Wang, Linian Wang, Qibin Wang, Shengjie Wang, Shuyi Wang, Si Wang, Wei Wang, Xiaochen Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yipu Wang, Yiqin Wang, Yucheng Wang, Yuzhi Wang, Zhaoji Wang, Zhaowei Wang, Zhengtao Wang, Zhexu Wang, Zihan Wang, Zizhe Wang, Chu Wei, Ming Wei, Chuan Wen, Zichen Wen, Chengjie Wu, Haoning Wu, Junyan Wu, Rucong Wu, Wenhao Wu, Yuefeng Wu, Yuhao Wu, Yuxin Wu, Zijian Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Yuchong Xie, Yifei Xin, Bowei Xing, Boyu Xu, Jianfan Xu, Jing Xu, Jinjing Xu, L\. H\. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinbo Xu, Xinran Xu, Yangchuan Xu, Yichang Xu, Yuemeng Xu, Zelai Xu, Ziyao Xu, Junjie Yan, Yuzi Yan, Guangyao Yang, Hao Yang, Junwei Yang, Kai Yang, Ningyuan Yang, Ruihan Yang, Xiaofei Yang, Xinlong Yang, Ying Yang, Yi Yang, Yi Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Dan Ye, Wenjie Ye, Zhuorui Ye, Bohong Yin, Chengzhen Yu, Longhui Yu, Tao Yu, Tianxiang Yu, Enming Yuan, Mengjie Yuan, Xiaokun Yuan, Yang Yue, Weihao Zeng, Dunyuan Zha, Haobing Zhan, Dehao Zhang, Hao Zhang, Jin Zhang, Puqi Zhang, Qiao Zhang, Rui Zhang, Xiaobin Zhang, Y\. Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yushun Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Chenguang Zhao, Feifan Zhao, Jinxiang Zhao, Shuai Zhao, Xiangyu Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Ruihan Zheng, Shaojie Zheng, Tengyang Zheng, Junfeng Zhong, Longguang Zhong, Weiming Zhong, M\. Zhou, Runjie Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Liya Zhu, Xinhao Zhu, Yuxuan Zhu, Zhen Zhu, Jingze Zhuang, Weiyu Zhuang, Ying Zou, and Xinxing Zu\.Kimi k2\.5: Visual agentic intelligence, 2026a\.URL[https://arxiv\.org/abs/2602\.02276](https://arxiv.org/abs/2602.02276)\.
- Team et al\. \(2026b\)Kimi Team, Yifan Bai, Yiping Bao, Y\. Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, Zhuofu Chen, Jialei Cui, Hao Ding, Mengnan Dong, Angang Du, Chenzhuang Du, Dikang Du, Yulun Du, Yu Fan, Yichen Feng, Kelin Fu, Bofei Gao, Chenxiao Gao, Hongcheng Gao, Peizhong Gao, Tong Gao, Yuyao Ge, Shangyi Geng, Qizheng Gu, Xinran Gu, Longyu Guan, Haiqing Guo, Jianhang Guo, Xiaoru Hao, Tianhong He, Weiran He, Wenyang He, Yunjia He, Chao Hong, Hao Hu, Yangyang Hu, Zhenxing Hu, Weixiao Huang, Zhiqi Huang, Zihao Huang, Tao Jiang, Zhejun Jiang, Xinyi Jin, Yongsheng Kang, Guokun Lai, Cheng Li, Fang Li, Haoyang Li, Ming Li, Wentao Li, Yang Li, Yanhao Li, Yiwei Li, Zhaowei Li, Zheming Li, Hongzhan Lin, Xiaohan Lin, Zongyu Lin, Chengyin Liu, Chenyu Liu, Hongzhang Liu, Jingyuan Liu, Junqi Liu, Liang Liu, Shaowei Liu, T\. Y\. Liu, Tianwei Liu, Weizhou Liu, Yangyang Liu, Yibo Liu, Yiping Liu, Yue Liu, Zhengying Liu, Enzhe Lu, Haoyu Lu, Lijun Lu, Yashuo Luo, Shengling Ma, Xinyu Ma, Yingwei Ma, Shaoguang Mao, Jie Mei, Xin Men, Yibo Miao, Siyuan Pan, Yebo Peng, Ruoyu Qin, Zeyu Qin, Bowen Qu, Zeyu Shang, Lidong Shi, Shengyuan Shi, Feifan Song, Jianlin Su, Zhengyuan Su, Lin Sui, Xinjie Sun, Flood Sung, Yunpeng Tai, Heyi Tang, Jiawen Tao, Qifeng Teng, Chaoran Tian, Chensi Wang, Dinglu Wang, Feng Wang, Hailong Wang, Haiming Wang, Jianzhou Wang, Jiaxing Wang, Jinhong Wang, Shengjie Wang, Shuyi Wang, Si Wang, Xinyuan Wang, Yao Wang, Yejie Wang, Yiqin Wang, Yuxin Wang, Yuzhi Wang, Zhaoji Wang, Zhengtao Wang, Zhengtao Wang, Zhexu Wang, Chu Wei, Qianqian Wei, Haoning Wu, Wenhao Wu, Xingzhe Wu, Yuxin Wu, Chenjun Xiao, Jin Xie, Xiaotong Xie, Weimin Xiong, Boyu Xu, Jinjing Xu, L\. H\. Xu, Lin Xu, Suting Xu, Weixin Xu, Xinran Xu, Yangchuan Xu, Ziyao Xu, Jing Xu, Jing Xu, Junjie Yan, Yuzi Yan, Hao Yang, Xiaofei Yang, Yi Yang, Ying Yang, Zhen Yang, Zhilin Yang, Zonghan Yang, Haotian Yao, Xingcheng Yao, Wenjie Ye, Zhuorui Ye, Bohong Yin, Longhui Yu, Enming Yuan, Hongbang Yuan, Mengjie Yuan, Siyu Yuan, Haobing Zhan, Dehao Zhang, Hao Zhang, Wanlu Zhang, Xiaobin Zhang, Yadong Zhang, Yangkun Zhang, Yichi Zhang, Yizhi Zhang, Yongting Zhang, Yu Zhang, Yutao Zhang, Yutong Zhang, Zheng Zhang, Haotian Zhao, Yikai Zhao, Zijia Zhao, Huabin Zheng, Shaojie Zheng, Longguang Zhong, Jianren Zhou, Xinyu Zhou, Zaida Zhou, Jinguo Zhu, Zhen Zhu, Weiyu Zhuang, and Xinxing Zu\.Kimi k2: Open agentic intelligence, 2026b\.URL[https://arxiv\.org/abs/2507\.20534](https://arxiv.org/abs/2507.20534)\.
- Wang et al\. \(2026\)Wenyi Wang, Piotr Piękos, Li Nanbo, Firas Laakom, Yimeng Chen, Mateusz Ostaszewski, Mingchen Zhuge, and Jürgen Schmidhuber\.Huxley\-g\\”odel machine: Human\-level coding agent development by an approximation of the optimal self\-improving machine\.In*The Fourteenth International Conference on Learning Representations*, 2026\.URL[https://openreview\.net/forum?id=T0EiEuhOOL](https://openreview.net/forum?id=T0EiEuhOOL)\.
- Wei et al\. \(2022\)Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al\.Chain\-of\-Thought Prompting Elicits Reasoning in Large Language Models\.*Advances in Neural Information Processing Systems*, 35:24824–24837, 2022\.URL[https://proceedings\.neurips\.cc/paper%5Ffiles/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4\-Abstract\-Conference\.html](https://proceedings.neurips.cc/paper%5Ffiles/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html)\.
- Woltereck \(1909\)Richard Woltereck\.Weitere experimentelle Untersuchungen über Artveränderung, speziell über das Wesen quantitativer Artunterschiede bei Daphniden\.*Verhandlungen der deutschen zoologischen Gesellschaft*, 19:110–173, 1909\.URL[https://cir\.nii\.ac\.jp/crid/1573668924237108864](https://cir.nii.ac.jp/crid/1573668924237108864)\.
- Wu et al\. \(2023\)Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang\.Autogen: Enabling next\-gen llm applications via multi\-agent conversation, 2023\.URL[https://arxiv\.org/abs/2308\.08155](https://arxiv.org/abs/2308.08155)\.
- Xia et al\. \(2025\)Chunqiu Steven Xia, Zhe Wang, Yan Yang, Yuxiang Wei, and Lingming Zhang\.Live\-swe\-agent: Can software engineering agents self\-evolve on the fly?, 2025\.URL[https://arxiv\.org/abs/2511\.13646](https://arxiv.org/abs/2511.13646)\.
- Yang et al\. \(2024\)John Yang, Carlos Jimenez, Alexander Wettig, Kilian Lieret, Shunyu Yao, Karthik Narasimhan, and Ofir Press\.Swe\-agent: Agent\-computer interfaces enable automated software engineering\.In A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(eds\.\),*Advances in Neural Information Processing Systems*, volume 37, pp\. 50528–50652\. Curran Associates, Inc\., 2024\.doi:10\.52202/079017\-1601\.URL[https://proceedings\.neurips\.cc/paper\_files/paper/2024/file/5a7c947568c1b1328ccc5230172e1e7c\-Paper\-Conference\.pdf](https://proceedings.neurips.cc/paper_files/paper/2024/file/5a7c947568c1b1328ccc5230172e1e7c-Paper-Conference.pdf)\.
- Yang et al\. \(2025\)John Yang, Kilian Lieret, Carlos E\. Jimenez, Alexander Wettig, Kabir Khandpur, Yanzhe Zhang, Binyuan Hui, Ofir Press, Ludwig Schmidt, and Diyi Yang\.Swe\-smith: Scaling data for software engineering agents, 2025\.URL[https://arxiv\.org/abs/2504\.21798](https://arxiv.org/abs/2504.21798)\.
- Yao et al\. \(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao\.React: Synergizing reasoning and acting in language models\.In*The Eleventh International Conference on Learning Representations*, 2023\.URL[https://openreview\.net/forum?id=WE\_vluYUL\-X](https://openreview.net/forum?id=WE_vluYUL-X)\.
- Zhang et al\. \(2026a\)Jenny Zhang, Shengran Hu, Cong Lu, Robert Tjarko Lange, and Jeff Clune\.Darwin gödel machine: Open\-ended evolution of self\-improving agents\.In*The Fourteenth International Conference on Learning Representations*, 2026a\.URL[https://openreview\.net/forum?id=pUpzQZTvGY](https://openreview.net/forum?id=pUpzQZTvGY)\.
- Zhang et al\. \(2026b\)Jenny Zhang, Bingchen Zhao, Wannan Yang, Jakob Foerster, Jeff Clune, Minqi Jiang, Sam Devlin, and Tatiana Shavrina\.Hyperagents, 2026b\.URL[https://arxiv\.org/abs/2603\.19461](https://arxiv.org/abs/2603.19461)\.
## Appendix Contents
## Appendix AMethod Details
### A\.1Pseudocode of Mendel Gödel Machine
Here, we present pseudocode of MGM, which helps understand the full evolution pipeline\.
1
Input:Initial coding agent
a0a\_\{0\}, benchmark suite
ℬ\\mathcal\{B\}, maximum iterations
TT, strategy weights
λA,λB,λC\\lambda\_\{A\},\\lambda\_\{B\},\\lambda\_\{C\}, failed task pool weights
βfail\\beta\_\{\\mathrm\{fail\}\}
Output:Archive of agents
𝒜\\mathcal\{A\}
2
r\(a0,ℬ\)←evaluate\(a0,ℬ\)r\(a\_\{0\},\\mathcal\{B\}\)\\leftarrow\\mathrm\{evaluate\}\(a\_\{0\},\\mathcal\{B\}\)
//Evaluate the base agent
3
initialize
𝒜←\{\(a0,r\(a0,ℬ\)\}\\mathcal\{A\}\\leftarrow\\\{\(a\_\{0\},r\(a\_\{0\},\\mathcal\{B\}\)\\\}
//Start with the base agent
4
initialize failed\-task pool
𝒫←∅\\mathcal\{P\}\\leftarrow\\emptyset
//Store informative failed tasks
5
6for*t←1t\\leftarrow 1toTT*do
7
decide action
u∈\{expand,evaluate\}u\\in\\\{\\mathrm\{expand\},\\mathrm\{evaluate\}\\\}
//Use Selection Policy
8
9if*u=expandu=\\mathrm\{expand\}*then
10
a′←SelectParent\(𝒜\)a^\{\\prime\}\\leftarrow\\mathrm\{SelectParent\}\(\\mathcal\{A\}\)
//Select a promising lineage
11
Ωa′←EligibleStrategies\(a′,𝒜\)\\Omega\_\{a^\{\\prime\}\}\\leftarrow\\mathrm\{EligibleStrategies\}\(a^\{\\prime\},\\mathcal\{A\}\)
//Find available mutation operators
12
σ←SampleStrategy\(Ωa′;λCM,λRM,λCH\)\\sigma\\leftarrow\\mathrm\{SampleStrategy\}\(\\Omega\_\{a^\{\\prime\}\};\\lambda\_\{\\mathrm\{CM\}\},\\lambda\_\{\\mathrm\{RM\}\},\\lambda\_\{\\mathrm\{CH\}\}\)
//Choose CM, RM, or CH
13
Eσ←BuildEvidence\(a′,σ,𝒜\)E\_\{\\sigma\}\\leftarrow\\mathrm\{BuildEvidence\}\(a^\{\\prime\},\\sigma,\\mathcal\{A\}\)
//Construct diagnostic evidence
14
c←Editσ\(a′,Eσ\)c\\leftarrow\\mathrm\{Edit\}\_\{\\sigma\}\(a^\{\\prime\},E\_\{\\sigma\}\)
//Self\-modification
15
16if*c\.is\_valid\(\)c\.\\mathrm\{is\\\_valid\}\(\)*then
𝒜←𝒜∪\{\(c,∅\)\}\\mathcal\{A\}\\leftarrow\\mathcal\{A\}\\cup\\\{\(c,\\emptyset\)\\\}
//Keep valid child agent
17
18
19else
20
a←SelectAgent\(𝒜\)a\\leftarrow\\mathrm\{SelectAgent\}\(\\mathcal\{A\}\)
//Choose agent to test
21
τ←SelectTask\(a,ℬ,𝒫,βfail\)\\tau\\leftarrow\\mathrm\{SelectTask\}\(a,\\mathcal\{B\},\\mathcal\{P\},\\beta\_\{\\mathrm\{fail\}\}\)
//Consider informative failed tasks
22
r\(a,τ\)←evaluate\(a,τ\)r\(a,\\tau\)\\leftarrow\\mathrm\{evaluate\}\(a,\\tau\)
//Evaluate on one task
23
update
𝒜\\mathcal\{A\}with
\(a,τ,r\(a,τ\)\)\(a,\\tau,r\(a,\\tau\)\)
//Store trajectory and result
24
25if*r\(a,τ\)=0r\(a,\\tau\)=0*then
𝒫←𝒫∪\{τ\}\\mathcal\{P\}\\leftarrow\\mathcal\{P\}\\cup\\\{\\tau\\\}
//Update failed\-task pool
26
27
28
29
30return
𝒜\\mathcal\{A\}
Algorithm 1Mendel Gödel MachineAlgorithm[1](https://arxiv.org/html/2608.07645#algorithm1)gives the complete MGM loop\. The procedure follows the standard archive\-based self\-improvement structure: each budgeted step either evaluates an existing agent on a task or expands the archive by self\-modifying an agent\. Evaluations update both the agent archive and the failed\-task pool\. During expansion, MGM constructs the set of eligible operators from already collected trajectories and samples one of clonal mutation, reaction\-norm mutation, or cross\-lineage hybridization\. The selected operator determines how evidence is assembled for the editor: from one failed trajectory, multiple trajectories of the same agent, or shared\-task trajectories across lineages\. Therefore, MGM improves the informativeness of each self\-modification step without adding extra task evaluations\.
### A\.2Primary\-Lineage Selection in Cross\-Lineage Hybridization
In the main algorithm, the policyπ\\pifirst selects an archive node as an anchor for expansion\. For Clonal Mutation and Reaction\-norm Mutation, this anchor is also the primary agent whose codebase is edited\. Cross\-lineage Hybridization differs because the purpose of the operator is to transfer a useful behavior from one lineage to another\. Therefore, the anchor node is used to retrieve a comparable peer, but the final primary lineage is determined by the shared\-task outcomes\.
Concretely, suppose the anchor agentaia\_\{i\}and a peer agentaja\_\{j\}have both attempted a shared taskτ⋆\\tau^\{\\star\}, and the task is not solved by both agents\. MGM constructs the comparison evidence
EC\(i,j,τ⋆\)=\{\(φ\(ai,τ⋆\),r\(ai,τ⋆\)\),\\displaystyle E\_\{C\}\(i,j,\\tau^\{\\star\}\)=\\\{\(\\varphi\(a\_\{i\},\\tau^\{\\star\}\),r\(a\_\{i\},\\tau^\{\\star\}\)\),\(28\)\(φ\(aj,τ⋆\),r\(aj,τ⋆\)\)\}\.\\displaystyle\(\\varphi\(a\_\{j\},\\tau^\{\\star\}\),r\(a\_\{j\},\\tau^\{\\star\}\)\)\\\}\.If exactly one of the two agents solvesτ⋆\\tau^\{\\star\}, the failing agent is selected as the primary agentapa\_\{p\}, and the successful agent is used as the context donoraqa\_\{q\}\. The resulting child is attached to the failing lineage:
a′←ΦC\(ap,EC\),a′∈children\(ap\)\.a^\{\\prime\}\\leftarrow\\varPhi\_\{C\}\(a\_\{p\},E\_\{C\}\),\\ a^\{\\prime\}\\in\\mathrm\{children\}\(a\_\{p\}\)\.\(29\)This direction is intentional: the goal is not to further edit the already successful lineage, but to let the failing lineage inherit a transferable behavior observed in the successful trajectory\.
If both agents fail onτ⋆\\tau^\{\\star\}, neither trajectory provides a direct success demonstration\. In this case, MGM uses the higher\-utility lineage as the primary agent and the other lineage as contrastive context\. This fallback preserves the archive policy’s preference for promising lineages while still allowing the editor to diagnose complementary failure modes from the same\-task comparison\.
Thus, the node selected byπ\\pishould be interpreted as an anchor for constructing a controlled comparison, not necessarily as the final parent edited by CH\. The final child is always attached to the primary lineage determined by the CH comparison\. This implementation matches the diagnostic role of hybridization: successful trajectories act as donors when available, and failed–failed comparisons are used to identify general weaknesses rather than task\-specific patches\.
## Appendix BTheory and Simulation Details
### B\.1Theoretical Justification of Comparative Fix Probability
This section provides the full derivation for Proposition 1 in Section[4\.2](https://arxiv.org/html/2608.07645#S4.SS2.SSS0.Px1)\. The goal is to justify why the comparative operators used by MGM can induce a higher effective fix probability than single\-trajectory clonal mutation\.
#### Diagnostic model\.
Let an agentaabe represented by a binary genotypeg\(a\)∈\{0,1\}Lg\(a\)\\in\\\{0,1\\\}^\{L\}, and let the oracle genotype beg⋆=𝟏g^\{\\star\}=\\mathbf\{1\}\. Define the set of incorrect loci as
M\(a\)=\{ℓ∈\[L\]:gℓ\(a\)≠gℓ⋆\}\.M\(a\)=\\\{\\ell\\in\[L\]:g\_\{\\ell\}\(a\)\\neq g^\{\\star\}\_\{\\ell\}\\\}\.\(30\)Each taskτ\\taurequires a subset of lociRτ⊆\[L\]R\_\{\\tau\}\\subseteq\[L\], with\|Rτ\|=k\|R\_\{\\tau\}\|=k\. The task is solved if and only if all required loci are correct:
r\(a,τ\)=1⟺Rτ∩M\(a\)=∅\.r\(a,\\tau\)=1\\quad\\Longleftrightarrow\\quad R\_\{\\tau\}\\cap M\(a\)=\\varnothing\.\(31\)Therefore, a failed task only reveals that at least one of its required loci is incorrect:
r\(a,τ\)=0⟹Rτ∩M\(a\)≠∅\.r\(a,\\tau\)=0\\quad\\Longrightarrow\\quad R\_\{\\tau\}\\cap M\(a\)\\neq\\varnothing\.\(32\)
We view a self\-modification operatorΦσ\\varPhi\_\{\\sigma\}as a diagnostic procedure\. Given evidenceEE, it constructs an implicit candidate setCσ\(E\)⊆\[L\]C\_\{\\sigma\}\(E\)\\subseteq\[L\]of loci that may explain the observed failure\. The editor then attempts to modify one locus inCσ\(E\)C\_\{\\sigma\}\(E\)\. Suppose that if the selected locus is truly incorrect, the editor repairs it with probabilitys∈\(0,1\]s\\in\(0,1\]\. Then the effective fix probability of operatorσ\\sigmais
pfσ=s⋅Prℓ∼Cσ\(E\)\[ℓ∈M\(a\)\]\.p\_\{f\}^\{\\sigma\}=s\\cdot\\Pr\_\{\\ell\\sim C\_\{\\sigma\}\(E\)\}\[\\ell\\in M\(a\)\]\.\(33\)Thus, an operator has higher fix probability when its evidence produces a candidate set with higher posterior density of truly incorrect loci\.
#### Clonal mutation\.
Clonal mutation uses a single failed trajectory\(φ\(a,τt\),0\)\(\\varphi\(a,\\tau\_\{t\}\),0\)\. Since this evidence only implies that at least one locus inRτtR\_\{\\tau\_\{t\}\}is incorrect, the natural candidate set is
CCM=Rτt\.C\_\{\\mathrm\{CM\}\}=R\_\{\\tau\_\{t\}\}\.\(34\)Let the number of truly incorrect loci among thekkloci required by the failed task be denoted as
c=\|Rτt∩M\(a\)\|\.c=\|R\_\{\\tau\_\{t\}\}\\cap M\(a\)\|\.\(35\)Then we have
pfCM=s⋅ck\.p\_\{f\}^\{\\mathrm\{CM\}\}=s\\cdot\\frac\{c\}\{k\}\.\(36\)In the sparse\-defect case where a failure is caused by one dominant missing capability,c=1c=1, and therefore
pfCM=sk\.p\_\{f\}^\{\\mathrm\{CM\}\}=\\frac\{s\}\{k\}\.\(37\)
#### Reaction\-norm mutation\.
Reaction\-norm mutation uses multiple trajectories from the same genotype\. Consider two failed tasksτt\\tau\_\{t\}andτr\\tau\_\{r\}of the same agent\. We assume that these two failures share a recurring causal defectb∈M\(a\)b\\in M\(a\), so that
b∈Rτt∩Rτr\.b\\in R\_\{\\tau\_\{t\}\}\\cap R\_\{\\tau\_\{r\}\}\.\(38\)This models the case where the same agent repeatedly fails because of the same scaffold\-level weakness rather than unrelated task\-specific accidents\.
Under this sound\-comparison assumption, the common explanatory region is
CRM=Rτt∩Rτr\.C\_\{\\mathrm\{RM\}\}=R\_\{\\tau\_\{t\}\}\\cap R\_\{\\tau\_\{r\}\}\.\(39\)Assume the remainingk−1k\-1required loci of each task are sampled independently from\[L\]∖\{b\}\[L\]\\setminus\\\{b\\\}\. Then the expected size of the intersection is
𝔼\[\|CRM\|\]=1\+\(k−1\)2L−1\.\\mathbb\{E\}\[\|C\_\{\\mathrm\{RM\}\}\|\]=1\+\\frac\{\(k\-1\)^\{2\}\}\{L\-1\}\.\(40\)SinceL\>kL\>k, we have
1\+\(k−1\)2L−1<k\.1\+\\frac\{\(k\-1\)^\{2\}\}\{L\-1\}<k\.\(41\)Because the recurring causal defectbbis contained inCRMC\_\{\\mathrm\{RM\}\}, the probability of targeting a truly incorrect locus is larger than in the single\-trajectory candidate set\. In the sparse\-defect case, this gives
pfRM≥s⋅11\+\(k−1\)2L−1\>sk=pfCM\.p\_\{f\}^\{\\mathrm\{RM\}\}\\geq s\\cdot\\frac\{1\}\{1\+\\frac\{\(k\-1\)^\{2\}\}\{L\-1\}\}\>\\frac\{s\}\{k\}=p\_\{f\}^\{\\mathrm\{CM\}\}\.\(42\)Therefore, reaction\-norm mutation improves fix probability by compressing the candidate set from the full task\-relevant regionRτtR\_\{\\tau\_\{t\}\}to the intersection of multiple failures that share the same genotype\-level defect\.
#### Cross\-lineage hybridization\.
Cross\-lineage hybridization compares different genotypes on the same task\. Consider a target agentata\_\{t\}that fails on taskτ\\tau, and a reference agentara\_\{r\}from another lineage that succeeds:
r\(at,τ\)=0,r\(ar,τ\)=1\.r\(a\_\{t\},\\tau\)=0,\\qquad r\(a\_\{r\},\\tau\)=1\.\(43\)The target failure implies
Rτ∩M\(at\)≠∅,R\_\{\\tau\}\\cap M\(a\_\{t\}\)\\neq\\varnothing,\(44\)whereas the reference success implies
Rτ∩M\(ar\)=∅\.R\_\{\\tau\}\\cap M\(a\_\{r\}\)=\\varnothing\.\(45\)Thus, the reference agent provides a contrastive control: the same task\-relevant loci are sufficient for success in the reference lineage, but at least one of them is defective in the target lineage\.
Let the number of causal target defects within the task\-relevant region be denoted as
c=\|Rτ∩M\(at\)\|\.c=\|R\_\{\\tau\}\\cap M\(a\_\{t\}\)\|\.\(46\)Clonal mutation must search over the wholeRτR\_\{\\tau\}, so
pfCM=s⋅ck\.p\_\{f\}^\{\\mathrm\{CM\}\}=s\\cdot\\frac\{c\}\{k\}\.\(47\)By comparing the failed target trajectory with the successful reference trajectory, cross\-lineage hybridization can filter out loci that are task\-relevant but unlikely to explain the target\-specific failure\. Let the resulting contrastive candidate set be
CCH=\(Rτ∩M\(at\)\)∪N,C\_\{\\mathrm\{CH\}\}=\(R\_\{\\tau\}\\cap M\(a\_\{t\}\)\)\\cup N,\(48\)whereNNdenotes non\-causal differences that remain after comparison\. Leth=\|N\|h=\|N\|\. Then
pfCH=s⋅cc\+h\.p\_\{f\}^\{\\mathrm\{CH\}\}=s\\cdot\\frac\{c\}\{c\+h\}\.\(49\)Whenever the reference comparison removes at least one irrelevant candidate, we have
Therefore,
pfCH=s⋅cc\+h\>s⋅ck=pfCM\.p\_\{f\}^\{\\mathrm\{CH\}\}=s\\cdot\\frac\{c\}\{c\+h\}\>s\\cdot\\frac\{c\}\{k\}=p\_\{f\}^\{\\mathrm\{CM\}\}\.\(51\)Thus, cross\-lineage hybridization improves fix probability by using a successful or contrastive reference lineage to remove non\-causal explanations from the candidate set\.
#### Discussion of assumptions\.
The above derivation is not intended to claim that comparative operators are unconditionally superior\. It relies on four assumptions\. First, failures are caused by relatively sparse causal defects, so that narrowing the candidate set meaningfully increases the density of true defects\. Second, reaction\-norm mutation is most beneficial when the compared trajectories share a recurring genotype\-level weakness\. Third, cross\-lineage hybridization requires an informative reference trajectory that serves as a useful contrastive control\. Fourth, the editor must be capable of exploiting the comparative evidence\. When these assumptions fail, the comparative operators may provide little or no fix\-quality advantage; this case is explicitly represented in our simulation by the null settingρ=1\\rho=1\.
### B\.2Simulation Parameters
Table 5:Parameter settings for simulation\.All edit costs are equal so that the HGM–MGM comparison isolates diagnostic quality alone\. The fix\-probability advantage ratio for the main simulation isρ=pfRM/pfCM=2\.0\\rho=p\_\{f\}^\{\\mathrm\{RM\}\}/p\_\{f\}^\{\\mathrm\{CM\}\}=2\.0; the sweep grid variesρ\\rhoandd0d\_\{0\}to test robustness\.
## Appendix CExperiment Details
### C\.1Compute and Software Environment
All experiments were run on cluster nodes with eight NVIDIA H100 GPUs \(80 GB GPU memory each\) and approximately 2 TB of host memory\. The Qwen models used in our experiments222[https://huggingface\.co/Qwen/Qwen3\.6\-35B\-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B)and[https://huggingface\.co/Qwen/Qwen3\-Coder\-Next](https://huggingface.co/Qwen/Qwen3-Coder-Next)was served with vLLM through an OpenAI\-compatible API using tensor parallelism across the visible GPUs\. We set the maximum context length to 262,144 tokens and the GPU memory utilization to 92%\. Benchmark tasks were executed inside Apptainer containers built from a Python 3\.10 base image, with the repository bind\-mounted into the container\. By default, host networking was disabled, and the container addressed the vLLM endpoint through the host node IP\. For DeepSeek\-V4 models, we directly accessed their preview release via the official DeepSeek API333[https://api\-docs\.deepseek\.com/news/news260424/](https://api-docs.deepseek.com/news/news260424/)\. For generation, the Qwen models were run with thinking enabled, a temperature of 1\.0, top\-ppof 0\.95, and top\-kkof 20, following the sampling configuration loaded by vLLM from the models’ generation configuration\. No explicit reasoning\-effort level or thinking\-token budget was imposed\. The DeepSeek\-V4 models were evaluated in non\-thinking mode, with the API\-default temperature and top\-ppvalues of 1\.0\. Parallel tool calls were disabled for both model families\.
### C\.2Hyper\-parameter Settings
In this section, to improve transparency and reproducibility, we present the hyper\-parameter settings used in our code implementations, as shown in Table[6](https://arxiv.org/html/2608.07645#A3.T6)\.
Table 6:Hyper\-parameter settings used in all experiments\.
## Appendix DBenchmark Details
To control evaluation cost, instead of evaluating MGM on the full SWE\-bench Pro and SWE\-bench Multilingual, we follow the procedures byZhang et al\. \([2026a](https://arxiv.org/html/2608.07645#bib.bib44)\)and utilize ChatGPT to randomly and separately choose 60 representative tasks for each benchmark\.
For SWE\-bench Pro, the selective 60\-task subset is required to be representative and involve all the six Programming Languages,JavaScript, Python, Java, C\+\+, TypeScript, Go\. Here we report the 60 tasks we used for our SWE\-bench Pro evaluation:
- •ansible\_\_ansible\-0ea40e09d1b35bcb69ff4d9cecf3d0defa4b36e8
- •ansible\_\_ansible\-189fcb37f973f0b1d52b555728208eeb9a6fce83
- •ansible\_\_ansible\-3889ddeb4b780ab4bac9ca2e75f8c1991bcabe83
- •ansible\_\_ansible\-5260527c4a71bfed99d803e687dd19619423b134
- •ansible\_\_ansible\-a20a52701402a12f91396549df04ac55809f68e9
- •internetarchive\_\_openlibrary\-308a35d6999427c02b1dbf5211c033ad3b352556
- •internetarchive\_\_openlibrary\-30bc73a1395fba2300087c7f307e54bb5372b60a
- •internetarchive\_\_openlibrary\-4b7ea2977be2747496ba792a678940baa985f7ea
- •internetarchive\_\_openlibrary\-7edd1ef09d91fe0b435707633c5cc9af41dedddf
- •internetarchive\_\_openlibrary\-9bdfd29fac883e77dcbc4208cab28c06fd963ab2
- •qutebrowser\_\_qutebrowser\-66cfa15c372fa9e613ea5a82d3b03e4609399fb6
- •qutebrowser\_\_qutebrowser\-6b320dc18662580e1313d2548fdd6231d2a97e6d
- •qutebrowser\_\_qutebrowser\-8cd06741bb56cdca49f5cdc0542da97681154315
- •qutebrowser\_\_qutebrowser\-8f46ba3f6dc7b18375f7aa63c48a1fe461190430
- •qutebrowser\_\_qutebrowser\-99029144b5109bb1b2a53964a7c129e009980cd9
- •flipt\-io\_\_flipt\-02e21636c58e86c51119b63e0fb5ca7b813b07b1
- •flipt\-io\_\_flipt\-5aef5a14890aa145c22d864a834694bae3a6f112
- •flipt\-io\_\_flipt\-9d25c18b79bc7829a6fb08ec9e8793d5d17e2868
- •flipt\-io\_\_flipt\-b68b8960b8a08540d5198d78c665a7eb0bea4008
- •flipt\-io\_\_flipt\-e2bd19dafa7166c96b082fb2a59eb54b4be0d778
- •future\-architect\_\_vuls\-78b52d6a7f480bd610b692de9bf0c86f57332f23
- •future\-architect\_\_vuls\-86b60e1478e44d28b1aff6b9ac7e95ceb05bc5fc
- •future\-architect\_\_vuls\-e049df50fa1eecdccc5348e27845b5c783ed7c76
- •gravitational\_\_teleport\-3fa6904377c006497169945428e8197158667910
- •gravitational\_\_teleport\-3ff75e29fb2153a2637fe7f83e49dc04b1c99c9f
- •gravitational\_\_teleport\-73cc189b0e9636d418c4470ecce0d9af5dae2f02
- •gravitational\_\_teleport\-ba6c4a135412c4296dd5551bd94042f0dc024504
- •navidrome\_\_navidrome\-66b74c81f115c78cb69910b0472eeb376750efc4
- •navidrome\_\_navidrome\-812dc2090f20ac4f8ac271b6ed95be5889d1a3ca
- •navidrome\_\_navidrome\-c90468b895f6171e33e937ff20dc915c995274f0
- •element\-hq\_\_element\-web\-1077729a19c0ce902e713cf6fab42c91fb7907f1
- •element\-hq\_\_element\-web\-41dfec20bfe9b62cddbbbf621bef2e9aa9685157
- •element\-hq\_\_element\-web\-53a9b6447bd7e6110ee4a63e2ec0322c250f08d1
- •element\-hq\_\_element\-web\-9a31cd0fa849da810b4fac6c6c015145e850b282
- •element\-hq\_\_element\-web\-b007ea81b2ccd001b00f332bee65070aa7fc00f9
- •NodeBB\_\_NodeBB\-397835a05a8e2897324e566b41c5e616e172b4af
- •NodeBB\_\_NodeBB\-51d8f3b195bddb13a13ddc0de110722774d9bb1b
- •NodeBB\_\_NodeBB\-97c8569a798075c50e93e585ac741ab55cb7c28b
- •NodeBB\_\_NodeBB\-be43cd25974681c9743d424238b7536c357dc8d3
- •NodeBB\_\_NodeBB\-f48ed3658aab7be0f1165d4c1f89af48d7865189
- •protonmail\_\_webclients\-01b519cd49e6a24d9a05d2eb97f54e420740072e
- •protonmail\_\_webclients\-08bb09914d0d37b0cd6376d4cab5b77728a43e7b
- •protonmail\_\_webclients\-51742625834d3bd0d10fe0c7e76b8739a59c6b9f
- •protonmail\_\_webclients\-6f8916fbadf1d1f4a26640f53b5cf7f55e8bedb7
- •protonmail\_\_webclients\-8142704f447df6e108d53cab25451c8a94976b92
- •tutao\_\_tutanota\-09c2776c0fce3db5c6e18da92b5a45dce9f013aa
- •tutao\_\_tutanota\-12a6cbaa4f8b43c2f85caca0787ab55501539955
- •tutao\_\_tutanota\-1e516e989b3c0221f4af6b297d9c0e4c43e4adc3
- •tutao\_\_tutanota\-1ff82aa365763cee2d609c9d19360ad87fdf2ec7
- •tutao\_\_tutanota\-219bc8f05d7b980e038bc1524cb021bf56397a1b
- •tutao\_\_tutanota\-40e94dee2bcec2b63f362da283123e9df1874cc1
- •tutao\_\_tutanota\-4b4e45949096bb288f2b522f657610e480efa3e8
- •tutao\_\_tutanota\-51818218c6ae33de00cbea3a4d30daac8c34142e
- •tutao\_\_tutanota\-8513a9e8114a8b42e64f4348335e0f23efa054c4
- •tutao\_\_tutanota\-b4934a0f3c34d9d7649e944b183137e8fad3e859
- •tutao\_\_tutanota\-d1aa0ecec288bfc800cfb9133b087c4f81ad8b38
- •tutao\_\_tutanota\-db90ac26ab78addf72a8efaff3c7acc0fbd6d000
- •tutao\_\_tutanota\-de49d486feef842101506adf040a0f00ded59519
- •tutao\_\_tutanota\-fb32e5f9d9fc152a00144d56dd0af01760a2d4dc
- •tutao\_\_tutanota\-fe240cbf7f0fdd6744ef7bef8cb61676bcdbb621
SWE\-bench Multilingual contains 300 tasks from 42 repositories and spans nine programming languages:C,C\+\+,Go,Java,JavaScript,TypeScript,PHP,Ruby, andRust\. We construct the subset to cover these language groups while keeping the evaluation cost manageable\. Here we report the exact tasks used in our SWE\-bench Multilingual evaluation below\.
- •apache\_\_druid\-14092
- •apache\_\_lucene\-13494
- •apache\_\_lucene\-13704
- •astral\-sh\_\_ruff\-15356
- •astral\-sh\_\_ruff\-15443
- •axios\_\_axios\-4731
- •babel\_\_babel\-16130
- •briannesbitt\_\_carbon\-3005
- •briannesbitt\_\_carbon\-3041
- •briannesbitt\_\_carbon\-3103
- •caddyserver\_\_caddy\-4943
- •caddyserver\_\_caddy\-5404
- •caddyserver\_\_caddy\-5870
- •caddyserver\_\_caddy\-5995
- •facebook\_\_docusaurus\-9183
- •fastlane\_\_fastlane\-19207
- •fastlane\_\_fastlane\-20642
- •fastlane\_\_fastlane\-20975
- •fluent\_\_fluentd\-3917
- •fmtlib\_\_fmt\-2317
- •fmtlib\_\_fmt\-2457
- •gohugoio\_\_hugo\-12579
- •google\_\_gson\-1014
- •immutable\-js\_\_immutable\-js\-2006
- •jekyll\_\_jekyll\-8771
- •jqlang\_\_jq\-2235
- •jqlang\_\_jq\-2658
- •jqlang\_\_jq\-2839
- •jqlang\_\_jq\-2919
- •laravel\_\_framework\-51195
- •laravel\_\_framework\-53914
- •laravel\_\_framework\-53949
- •nushell\_\_nushell\-13605
- •php\-cs\-fixer\_\_php\-cs\-fixer\-7635
- •phpoffice\_\_phpspreadsheet\-3570
- •phpoffice\_\_phpspreadsheet\-4114
- •preactjs\_\_preact\-2757
- •preactjs\_\_preact\-3454
- •preactjs\_\_preact\-3562
- •preactjs\_\_preact\-4436
- •projectlombok\_\_lombok\-3009
- •projectlombok\_\_lombok\-3350
- •projectlombok\_\_lombok\-3422
- •projectlombok\_\_lombok\-3479
- •projectlombok\_\_lombok\-3594
- •prometheus\_\_prometheus\-10633
- •prometheus\_\_prometheus\-12874
- •prometheus\_\_prometheus\-14861
- •redis\_\_redis\-11734
- •redis\_\_redis\-13115
- •rubocop\_\_rubocop\-13375
- •rubocop\_\_rubocop\-13479
- •rubocop\_\_rubocop\-13627
- •sharkdp\_\_bat\-2650
- •tokio\-rs\_\_axum\-1730
- •tokio\-rs\_\_tokio\-6752
- •tokio\-rs\_\_tokio\-6838
- •uutils\_\_coreutils\-6575
- •uutils\_\_coreutils\-6682
- •vuejs\_\_core\-11739
## Appendix EAdditional Results
### E\.1Full Evaluation on Polyglot
To verify that the Polyglot improvement reported in the main experiments is not an artifact of the 60\-task subset, we additionally evaluate the best MGM\-discovered agent on the full Polyglot benchmark\. This evaluation is performed only after evolution is complete: the agent is not further updated, and the full benchmark is used only for post\-evolution evaluation and is not used to further update the agent\.
As shown in Figure[7](https://arxiv.org/html/2608.07645#A5.F7), the MGM\-evolved agent solves 210 out of 225 tasks, achieving an overall accuracy of93\.3%\. This closely matches the 93\.2% accuracy observed on Polyglot\-60 in the main experiment, suggesting that the improvement is stable when moving from the subset evaluation to the full benchmark\. The gains are also broadly distributed across languages: the agent solves 25/26 C\+\+ tasks, 36/39 Go tasks, 43/47 Java tasks, 46/49 JavaScript tasks, 33/34 Python tasks, and 27/30 Rust tasks\.
These results support the claim that MGM learns reusable, language\-agnostic workflow improvements rather than narrow heuristics tied to a particular language or subset\. In particular, the consistently high resolution rate across C\+\+, Go, Java, JavaScript, Python, and Rust is consistent with the role of reaction\-norm mutation and cross\-lineage hybridization: the former identifies recurring agent\-level weaknesses across tasks, while the latter transfers useful behavioral traits across lineages\. The remaining failures are not concentrated in a single language, indicating that future improvements should likely target harder residual failure modes rather than language\-specific specialization\.
Figure 7:Polyglot\-225 performance of the MGM\-discovered agent by language\. The agent is evaluated after evolution without further self\-modification and solves 210 out of 225 tasks overall\. The area of pie slices report the proportion of resolved and unresolved tasks within each language, with labels showing the ratio of resolved to total tasks\. MGM maintains high accuracy across C\+\+, Go, Java, JavaScript, Python, and Rust, indicating that the evolved improvements transfer across languages rather than specializing to a single language subset\.
### E\.2Visualizations of Evolution Trees
\(a\)MGM
\(b\)HGM
\(c\)w/oΦRM\\varPhi\_\{\\textrm\{RM\}\}\(ΦCM\\varPhi\_\{\\textrm\{CM\}\}andΦCH\\varPhi\_\{\\textrm\{CH\}\}only\)
\(d\)w/oΦCH\\varPhi\_\{\\textrm\{CH\}\}\(ΦCM\\varPhi\_\{\\textrm\{CM\}\}andΦRM\\varPhi\_\{\\textrm\{RM\}\}only\)
Figure 8:Evolution trees of MGM, HGM, and MGM’s ablated variants\.Results after 200φ\\varphi\-evaluations and 24Φ\\varPhi\-expansions on Polyglot\. Nodes are independently colored by utility estimates aggregated over each tree’sφ\\varphi\-evaluation results\.All four figures use the same visual encoding: node fill color denotes evaluation accuracy, while edge color denotes the child node’s self\-improvement operator \(ΦCM\\varPhi\_\{\\mathrm\{CM\}\}/ΦRM\\varPhi\_\{\\mathrm\{RM\}\}/ΦCH\\varPhi\_\{\\mathrm\{CH\}\}\)\. Starred nodes mark the representative agent chosen for each experimental setting\.
Figure[8\(a\)](https://arxiv.org/html/2608.07645#A5.F8.sf1)shows the evolution tree produced by the full MGM system under a budget of 200 task evaluations, yielding 24 nodes\. All three operators are active: Clonal Mutation, Reaction\-norm Mutation, and Cross\-lineage Hybridization\. Node fill color encodes evaluation accuracy\. The node \#20 and \#23 has the highest utility 1\.00 but only with 15 and 4 evaluations, respectively\. The node \#16 has utility 0\.91 after 35 evaluations and is selected as the final result\.
Figure[8\(b\)](https://arxiv.org/html/2608.07645#A5.F8.sf2)shows the evolution tree of the HGM baseline, which uses Clonal Mutation as its only self\-modification operator\. Under the same 200\-evaluation budget, HGM also produces 24 nodes, but all child edges correspond to single\-trajectory clonal edits\. Compared with full MGM, HGM still explores multiple branches, but it lacks the additional comparative evidence channels represented by edges with different colors\. Although node \#24 has the highest utility 1\.00, it has only been evaluated 5 times, so its color reflects a high\-variance estimate rather than a reliable final selection\. The node \#18 is selected as result, with utility 0\.71 after 14 evaluations\.
Figure[8\(c\)](https://arxiv.org/html/2608.07645#A5.F8.sf3)shows the ablation that removes Reaction\-norm Mutation\. Clonal Mutation and Cross\-lineage Hybridization remain enabled\. The node \#23 with highest utility 1\.00 got only 2 evaluations, and the node \#14 with utility 0\.84 after 50 evaluations is selected as result\.
Figure[8\(d\)](https://arxiv.org/html/2608.07645#A5.F8.sf4)shows the ablation that removes Cross\-lineage Hybridization\. Clonal Mutation and Reaction\-norm Mutation remain active\. All improvement is confined to within\-lineage evolution\. The node \#19 has utility 1\.00 after only 2 evaluations, and the selected result is node \#6 with utility 0\.88, after 16 evaluations\.
Figure 9:Final per\-node utilities for MGM, HGM, and the ablations\.Each point is one of the 24 evolved nodes\. Vertical position and fill color encode utility\. Marker area is proportional to the number of evaluations\. The marked node in each column is the selected final result\.In Figure[9](https://arxiv.org/html/2608.07645#A5.F9)we additionally report the final utilities of all 24 nodes under each configuration\. Point size reflects evaluation count, making high\-utility nodes with few measurements visually distinct from better\-supported estimates\. The raw maximum is often attained by a sparsely evaluated node, whereas the selected result favors a high utility backed by more evaluations\.
## Appendix FAdditional Discussion
### F\.1When Does Hybridization Help? A Case Study
Figure 10:Illustration of Cross\-lineage Hybridization\. Nodes are colored by their outcome on the shared diagnostic taskjavascript\_\_queen\-attack: red nodes fail and green nodes solve it\.Just as Figure[10](https://arxiv.org/html/2608.07645#A6.F10)illustrates,Cross\-lineage Hybridizationenables the evolutionary process to transfer skills discovered in one archive to another, thereby improving evolutionary efficiency\. In this example, both node \#1 and node \#2 initially fail on the same diagnostic task,javascript\_\_queen\-attack\. Along the donor lineage, node \#2 later produces node \#6, which is the first descendant in this branch to solve the task\. The evolved behavior can be summarized as a test\-contract skill\. Before implementing, the agent learns to read tests as strict behavioral contracts, preserving exact public APIs, expected values, and error messages\. This capability is then preserved in node \#8, which also solvesjavascript\_\_queen\-attack\. Crucially, node \#8 does not merely improve its own lineage\. Through Cross\-lineage Hybridization, it serves as a successful context donor for a separate failing lineage rooted at node \#3\. The resulting hybrid child, node \#15, also solves the previously failed task\. This case shows that hybridization is not simply another mutation operator, instead it acts as a cross\-archive mechanism for capability transfer, allowing a reusable skill evolved in one branch to accelerate progress in another branch that had not discovered it independently\.
### F\.2Bigger Model Does Not Always Lead to Better Results
Since recursive self\-improvement ultimately edits the agent’s own codebase, it is tempting to view self\-evolution as another coding task\. Under this view, a larger or more coding\-specialized foundation model should naturally lead to stronger self\-improvement: if a model is better at coding benchmarks, it should also be better at modifying the agent scaffold\. However, our experiments here show that this intuition is incomplete\. Self\-improvement is indeed implemented through code editing, but the quality of the edit depends critically on preceding diagnosis steps, where the model must infer why the current scaffold failed and what general modification should be made\.
Table 7:Performance of coding agents evolved on SWE\-bench Verified with different models\.Results evolved using 200φ\\varphi\-evaluations and 24Φ\\varPhi\-expansions\. For each benchmark, HGM and MGM start from the same initial scaffold\. Superscripts in accuracy denote absolute percentage\-point improvements over corresponding initial agents\. Time reported as CPU wall\-clock time with 8×\\timesNVIDIA H100 GPUs\.To examine this, we conduct an additional experiment on SWE\-bench Verified using Qwen3\-Coder\-Next\-80B\-A3B as the backbone model\. Qwen3\-Coder\-Next is larger and explicitly optimized for coding with comparable coding ability but poorer reasoning ability, so one might expect it to produce stronger self\-improving agents\. Surprisingly, this is not what we observe\. As shown in Table[7](https://arxiv.org/html/2608.07645#A6.T7), the initial scaffold obtains 33\.3% accuracy\. Under the same budget of 200 task evaluations and 24 self\-modification expansions, HGM improves the scaffold to 40\.0%, while MGM improves it to 41\.7%\. MGM still outperforms HGM under this backbone, but the final evolved performance is substantially lower than the corresponding Qwen3\.6\-35B\-A3B setting, where MGM reaches 61\.7% on SWE\-bench Verified\-60\.
This suggests that self\-improvement performance cannot be predicted solely from model size or coding specialization\. The model must inspect the trajectory, diagnose the underlying failure, and transform that diagnosis into a robust scaffold\-level edit\. This diagnosis step is not equivalent to ordinary code completion or local bug fixing\. It requires the model to reason about the interaction between prompts, tools, control flow, repository exploration, execution feedback, and prior agent behavior\. If the diagnosis is shallow or incorrect, the resulting code edit may still be syntactically valid, but it can target the wrong mechanism, overfit to superficial symptoms, or introduce brittle workflow changes\.
Table 8:Performance comparison between models\.Results are compared across general knowledge, science, and mathematics reasoning benchmarks\.To better understand this result, as shown in Table[8](https://arxiv.org/html/2608.07645#A6.T8), we compare the broader reasoning abilities of Qwen3\-Coder\-Next\-80B\-A3B and Qwen3\.6\-35B\-A3B\. Although Qwen3\-Coder\-Next is larger and coding\-oriented, Qwen3\.6\-35B\-A3B performs substantially better on a wide range of general and reasoning\-intensive benchmarks\. Qwen3\.6 scores 93\.3 on MMLU\-Redux, compared with 91\.18 for Qwen3\-Coder\-Next, and 85\.2 on MMLU\-Pro, compared with 80\.52\. The gaps become larger on more difficult reasoning tasks: Qwen3\.6 outperforms Qwen3\-Coder\-Next by 11\.51 points on GPQA, 7\.25 points on SuperGPQA, 20\.49 points on HMMT Feb 2025, 13\.53 points on HMMT Nov 2025, and 21\.47 points on LiveCodeBench v6\. These results indicate that Qwen3\.6 has a stronger general reasoning profile, despite having fewer total parameters\.
This reasoning advantage helps explain why Qwen3\.6 leads to stronger self\-improvement\. In a self\-evolving coding agent, the model is not only asked to write code; it is asked to reason about why the current agent failed and how the scaffold should be changed to prevent similar failures in future tasks\. The edit itself is a coding operation, but deciding what to edit is a diagnosis and abstraction problem\. A coding\-specialized model may be strong at implementing local changes, yet still produce weaker self\-improvement if it cannot reliably infer the correct failure mechanism from trajectories\. Conversely, a model with stronger reasoning ability can generate more accurate diagnoses and therefore produce more reusable scaffold\-level modifications\.
This observation is particularly important for MGM\. MGM introduces reaction\-norm mutation and cross\-lineage hybridization, both of which increase the amount of comparative evidence available to the self\-modification step\. However, richer evidence is only useful if the backbone model can reason over it\. Reaction\-norm mutation requires the model to compare multiple trajectories of the same agent and identify recurring behavioral patterns\. Cross\-lineage hybridization requires the model to compare different agents on the same task and extract transferable scaffold\-level traits\. These operations make the diagnosis stage more informative, but also more reasoning\-intensive\.
Overall, our results do not suggest that larger or more coding\-specialized models are ineffective for self\-improvement\. Qwen3\-Coder\-Next still enables both HGM and MGM to improve over the initial scaffold, showing that strong coding backbones can support meaningful self\-evolution\. However, the improvement is not necessarily monotonic with model size or coding specialization\. Although self\-evolution ultimately edits code, the quality of the edit depends on whether the model can diagnose the failure mechanism before modifying the scaffold\. A larger coding model may therefore produce valid self\-edits, but not necessarily better self\-edits, if its trajectory\-level diagnosis and abstraction ability are weaker\. This suggests that backbone selection for self\-improving agents should consider a broader capability profile, especially general reasoning and failure\-diagnosis ability, rather than relying only on parameter count or coding benchmark performance\.
### F\.3Are skills evolved from MGM really more general?
The visualization[11](https://arxiv.org/html/2608.07645#A6.F11)suggests that MGM produces changes that concentrate around a reusable workflow\-level capability\. Most MGM points lie in a compact semantic region centered on exact test\-contract extraction, API\-contract adherence, and pre\-implementation verification\. These are not tied to a particular programming language, benchmark instance, or repository\-specific bug\. Instead, they jointly describe a general procedure incorporating reading tests and existing interfaces, extracting the expected contract, implementing against that contract, and verifying exact agreement before finalizing the patch\. This kind of skills can transfer across many programming tasks because almost all repair problems involve some form of latent contract between tests, existing code, and the expected implementation\.
Figure 11:Semantic visualization of evolved skills\.A content\-level view of the evolved “To Implement” descriptions from MGM and HGM\. Each point corresponds to one evolved skill or workflow change, embedded by the semantics of its To Implement text\. The marker shape indicates the evolution method, while color indicates the main semantic theme of the proposed change\.In contrast, HGM points are more widely dispersed in the semantic map\. HGM also discovers useful ideas, such as iterative testing, API alignment, context discovery, and auxiliary tool utilities, but these ideas appear less consolidated into a single reusable default workflow\. The broader spread of HGM points suggests a more heterogeneous search over possible interventions\. Some HGM changes add or wire in tools, while others propose isolated prompt phases or iterative loops\. These can be beneficial, but they are less consistently organized around one general mechanism that would apply by default across unseen tasks\.
This distinction is important because generalizability should not be measured only by whether a change uses broad language such as general capability\. A more operational criterion is whether the evolved skill abstracts away from the original training instance and becomes a reusable procedure for a broad class of future tasks\. Under this criterion, MGM appears more general at the workflow level\. It repeatedly evolves skills that convert task\-specific feedback into a language\-agnostic contract\-extraction and verification protocol\. HGM, by comparison, evolves a wider variety of interventions, but they are more fragmented and less clearly integrated into a stable general workflow\.
## Appendix GAdditional Related Work
As foundation models continue to advance, includingDeepSeek\-AI et al\. \([2025](https://arxiv.org/html/2608.07645#bib.bib6);[2026](https://arxiv.org/html/2608.07645#bib.bib7)\); Team et al\. \([2026b](https://arxiv.org/html/2608.07645#bib.bib35);[a](https://arxiv.org/html/2608.07645#bib.bib34)\); Singh et al\. \([2026](https://arxiv.org/html/2608.07645#bib.bib33)\); MiniMax et al\. \([2025](https://arxiv.org/html/2608.07645#bib.bib19)\), the attainable performance ceiling of coding agents is also steadily increasing\. The development of agentic systems has progressed from human\-designed scaffolds\(Cai et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib2); Qian et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib23)\), to automated agent design\(Hu et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib13)\), and more recently to self\-improving agents\(Zhang et al\.,[2026a](https://arxiv.org/html/2608.07645#bib.bib44); Wang et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib36); Qiu et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib24); Gao et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib10)\)\.
### G\.1From Agent Design to Inherited Self\-Modification
Modern LLM agents build on a broad line of work on tool use, reasoning\-action interleaving, and multi\-agent orchestration, including modular tool\-augmented systems\(Karpas et al\.,[2022](https://arxiv.org/html/2608.07645#bib.bib15)\), ReAct\-style reasoning\-and\-acting\(Yao et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib43)\), self\-supervised tool use\(Schick et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib27)\), and multi\-agent frameworks such as AutoGen and MetaGPT\(Wu et al\.,[2023](https://arxiv.org/html/2608.07645#bib.bib39); Hong et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib12)\)\. These general agentic paradigms become especially important in software engineering, where an agent must inspect repositories, edit files, execute commands, and validate patches\. Systems such as SWE\-agent emphasize the importance of the agent\-computer interface: a model’s ability to inspect files, edit code, execute commands, and run tests depends heavily on the tools and interaction protocol exposed to it\(Yang et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib41)\)\. HyperAgent instead explores a multi\-agent decomposition of software engineering work, assigning specialized roles such as planning, navigation, editing, and execution to different agents\(Zhang et al\.,[2026b](https://arxiv.org/html/2608.07645#bib.bib45)\)\. These systems show that scaffold design is central to coding\-agent performance, but the scaffold is still largely engineered by humans\.
ADAS shifts scaffold design from manual engineering to automated search\. Meta Agent Search represents agents as code and uses a meta\-agent to invent new agent programs from an archive of prior discoveries\(Hu et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib13)\)\. This view is important because it treats prompts, tools, and control flow as searchable program components rather than fixed infrastructure\. However, ADAS\-style methods usually preserve a separation between the designer and the designed agent\. The meta\-agent is responsible for generating new target agents, whereas the target agent does not necessarily improve itself through its own execution history\.
Self\-improving coding agents reduce this separation\. SICA shows that a coding agent can use basic file\-editing tools to modify its own codebase and improve on benchmarks\(Robeyns et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib26)\)\. DGM extends this into open\-ended evolution by maintaining an archive of self\-modified agents and sampling from it to create new descendants\(Zhang et al\.,[2026a](https://arxiv.org/html/2608.07645#bib.bib44)\)\. HGM further observes that an agent’s current benchmark performance is not always aligned with its future self\-improvement potential, and therefore estimates descendant\-based metaproductivity to guide archive expansion\(Wang et al\.,[2026](https://arxiv.org/html/2608.07645#bib.bib36)\)\. These works establish the XGM setting: an agent population evolves through persistent, heritable edits to executable scaffolds\.
MGM focuses on a different bottleneck inside this setting\. Prior Gödel Machine\-style methods mainly decide which node in the archive should be expanded\. MGM instead asks what evidence should be given to the editor once an expansion is triggered\. A single failed trajectory can be noisy: it may reflect a task\-specific accident, a bad local choice, or a general weakness in the agent’s design\. MGM reduces this ambiguity by constructing controlled comparisons from trajectories that are already present in the archive\. Reaction\-norm mutation compares the same genotype across multiple environments, making recurring failures more likely to reveal a stable design weakness\. Cross\-lineage hybridization compares different genotypes on the same task, making behavioral differences easier to interpret as transferable skills\. Therefore, MGM improves the diagnostic quality of self\-modification without requiring additional task evaluations\.
### G\.2Runtime Adaptation versus Training\-Time Self\-Evolution
Live\-SWE\-agent is the most relevant concurrent work to distinguish from MGM\. It starts from a minimal agent scaffold and lets the agent expand or revise its own capabilities during the process of solving a real\-world software issue\(Xia et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib40)\)\. This is a powerful runtime adaptation mechanism: the agent can create helper tools, refine its own execution procedure, and specialize its workflow to the current repository\. Its central advantage is immediacy\. It does not require an offline evolution loop before deployment, and the agent can adapt to the concrete structure of the issue it is currently solving\.
MGM addresses a different question\. Rather than asking how an agent should adapt inside one episode, MGM asks how a population of agents should improve across many episodes so that future agents inherit better scaffolds\. The modifications produced by MGM are persistent changes to the agent program\. They are evaluated over a task distribution and stored as descendants in an archive\. This makes MGM a training\-time self\-evolution method, even though the training signal is not gradient\-based\. The objective is to discover general capabilities—better planning, validation, context management, debugging, or tool usage—that remain useful beyond the tasks that exposed them\.
This distinction is similar to the difference between chain\-of\-thought prompting and policy optimization\. Chain\-of\-thought prompting allocates more inference\-time computation to a single response, improving the current trajectory without changing the model parameters\(Wei et al\.,[2022](https://arxiv.org/html/2608.07645#bib.bib37)\)\. GRPO, by contrast, is a reinforcement learning algorithm that updates the policy so that future trajectories improve\(Shao et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib31)\)\. Live\-SWE\-agent plays a role analogous to test\-time reasoning or skill construction: it improves the current problem\-solving process\. MGM plays a role analogous to policy learning: it changes the inherited scaffold that future agents use\. These paradigms are complementary\. A strong practical system could first use MGM to evolve robust base agents offline and then use live runtime adaptation to specialize the evolved agent to the current issue\.
### G\.3Open\-Ended Search and Algorithm Discovery
MGM is also related to open\-ended and quality\-diversity search\. In these paradigms, progress does not come only from optimizing a single incumbent solution, but from maintaining an archive of diverse candidates that can serve as stepping stones for future discovery\. Quality\-diversity methods such as MAP\-Elites maintain structured archives of high\-performing yet behaviorally diverse solutions, improving both search coverage and downstream adaptability\(Mouret & Clune,[2015](https://arxiv.org/html/2608.07645#bib.bib20); Cully,[2021](https://arxiv.org/html/2608.07645#bib.bib5)\)\. Open\-ended evolution further emphasizes that preserving diverse lineages can expose stepping stones that would be missed by purely exploitative optimization\(Clune,[2020](https://arxiv.org/html/2608.07645#bib.bib4)\)\.
Recent foundation\-model\-based discovery systems instantiate a related idea in code and scientific domains\. AlphaEvolve uses a coding agent to iteratively propose, evaluate, and improve programs for scientific and algorithmic discovery\(Novikov et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib21)\)\. Generative modeling has also been used to search for mathematical objects and conjecture\-relevant structures, suggesting that learned generative models can act as proposal mechanisms for discovery under external evaluators\(Ellenberg et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib9)\)\. MGM differs from these systems in that its search space is not a standalone algorithm or mathematical object, but the executable scaffold of a coding agent itself\. Rather than only preserving diverse candidates as archive entries, MGM reuses the archive as a source of comparative diagnostic evidence for inherited self\-modification\.
### G\.4Benchmark Coverage and Evaluation Motivation
SWE\-bench is the canonical benchmark for repository\-level issue resolution\. Unlike function\-level benchmarks such as HumanEval\(Chen et al\.,[2021](https://arxiv.org/html/2608.07645#bib.bib3)\)or MBPP\(Austin et al\.,[2021](https://arxiv.org/html/2608.07645#bib.bib1)\), SWE\-bench gives the model a real repository and a natural\-language GitHub issue, and evaluates whether the generated patch passes tests derived from the corresponding pull request\(Jimenez et al\.,[2024](https://arxiv.org/html/2608.07645#bib.bib14)\)\. This setting stresses repository navigation, fault localization, patch construction, and validation\. SWE\-bench Verified further improves reliability by using a human\-filtered subset, which is why it has become a standard benchmark for comparing software\-engineering agents\.
However, SWE\-bench Verified alone is not sufficient for evaluating self\-evolving agents\. First, it is mostly Python\-centric, so an evolved agent may overfit to Python idioms, test frameworks, or repository layouts\. Second, many tasks are relatively short compared with professional software engineering work\. SWE\-bench Pro addresses the second limitation by introducing harder long\-horizon tasks from a broader range of actively maintained repositories, often requiring deeper investigation and multi\-file changes\(Deng et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib8)\)\. Polyglot addresses the first limitation by testing coding across several languages, including C\+\+, Go, Java, JavaScript, Python, and Rust\(Gauthier,[2024](https://arxiv.org/html/2608.07645#bib.bib11)\)\. Multilingual repository\-level variants of SWE\-bench provide another related direction by extending issue resolution beyond Python repositories\(Khandpur & the SWE\-bench Team,[2025](https://arxiv.org/html/2608.07645#bib.bib16); Yang et al\.,[2025](https://arxiv.org/html/2608.07645#bib.bib42)\)\.
These benchmark choices are aligned with MGM’s objective\. Reaction\-norm mutation is intended to identify general weaknesses that recur across tasks, so it should be evaluated on settings where task diversity matters\. Cross\-lineage hybridization is intended to transfer useful behaviors between agents, so it should be tested on benchmarks where different strategies may solve different subsets of tasks\. SWE\-bench Verified, SWE\-bench Pro, and Polyglot therefore provide complementary evidence: standard real\-world issue resolution, long\-horizon robustness, and cross\-language generality\.
## Appendix HBest Discovered Agents
### H\.1MGM on Polyglot
Compared with the initial agent, whoseforward\(\)routine performs a single\-turn code generation call without structured repository analysis or test feedback, agent \(node \#16\) accumulates three successive modifications along the lineageinitial→\\rightarrow\#2→\\rightarrow\#9→\\rightarrow\#16\. The first patch \(\#2, introduced via cross\-lineage hybridization\) adds a read\-first prompt directive that requires the agent to inspect stubs, tests, and class hierarchies with theeditortool and to implement only the interfaces already defined in the repository\. The second patch \(\#9, a reaction\-norm mutation\) replaces the single\-turn workflow with a two\-phase pipeline: in Phase 1, the agent extracts a structuredcontract\_planJSON containing exact class signatures, constructors, method definitions, and verbatim error messages from test and stub files\. In Phase 2, it implements the solution under strict adherence to that plan and performs an explicit checklist\-based self\-audit before submission\. The third patch \(\#16, also a reaction\-norm mutation\) extends this design with a test\-driven verification loop: after implementation, the agent repeatedly executes the task’s test suite \(via a newrun\_tests\(\)method and supporting utilities inutils/test\_utils\.py\), feeds raw test output back to the model together with the remaining contract, and iteratively repairs the solution for up to five attempts until tests pass or the budget is exhausted\. Cumulatively, these changes transform the initial agent from a prompt\-only code generator into a test\-contract\-guided, self\-auditing, and empirically self\-correcting coding agent capable of reducing API hallucination, enforcing exact test constraints, and recovering from runtime failures through language\-agnostic test feedback\.
### H\.2MGM on SWE\-bench Verified
Listing below shows an example self\-modification discovered by MGM on SWE\-bench Verified\. Unlike a task\-specific repository patch, this modification changes the agent’s general debugging workflow\. The evolved scaffold turns the original single\-pass repair process into a more trace\-aware and patch\-constrained procedure\.
Concretely, the agent first extracts test function names from the problem statement and test description, and then builds an explicit code\-path trace before editing\. When relevant failing tests are available, the scaffold generates acode\_path\_tracethat links the failing test to the likely implementation path and provides a targeted fix direction\. After the initial repair attempt, the agent further inspects the message history to identify remaining failed tests and can perform additional revision rounds conditioned on the traced failure path\. This makes the repair process less dependent on a vague natural\-language issue description and more directly grounded in executable regression signals\.
The modification also introduces lightweight provenance tracking for repository exploration\. By registering callbacks around the editor tool, the agent records which files were actually viewed during the debugging process\. This information is then used by a new diff\-minimality filter, which removes patch blocks that are weakly related to the problem statement or to the files inspected by the agent\. As a result, the evolved scaffold encourages localized fixes and discourages broad, accidental, or speculative edits\.
The concrete utility added in this example focuses on tracing docstring and autodoc\-style failures, but the underlying scaffold\-level skill is more general: MGM discovers a workflow that first grounds the repair in failing tests, then traces the relevant code path, and finally constrains the submitted diff to files supported by the agent’s own investigation\. This provides qualitative evidence that MGM can evolve reusable repository\-level debugging habits on SWE\-bench Verified, rather than merely memorizing a solution to a single benchmark instance\.
Listing 1:Figure 17: Code modification on Polyglotdiff
index39cb47e\.\.d824c3d100644
\-\-\-a/coding\_agent\.py
\+\+\+b/coding\_agent\.py
@@\-172,7\+172,21@@classAgenticSystem:
Yourtaskistomakechangestothefilesinthe\{self\.git\_dir\}directorytoaddressthe<problem\_description\>\.Ihavealreadytakencareoftherequireddependencies\.
"""
\-instruction=f"\{task\}\\n\\nPleaseanalyzetheproblemdescriptioncarefully\.Thenmakeeditstothecodefilestocompletetheinstruction\."
\+inspection\_directive="""
\+\*\*CRITICALPRE\-IMPLEMENTATIONREQUIREMENT:\*\*
\+
\+Beforewritinganycode,carefullyinspecttheexistingrepositorystructure\.Usethe‘editor‘toolto‘view‘allrelevantstubfiles,testfiles,andrelatedclasses\.Extractandstrictlyadheretotheirexactmethodsignatures,classhierarchies,andconstructorrequirements\.Donotassumeorinventinterfaces;implementexactlywhattheexistingcodeexpects\.
\+
\+\*\*SpecificInstructions:\*\*
\+1\.First,exploretherepositorystructureusing‘view‘commandstounderstandthefilelayout
\+2\.Readallstubfiles,interfacedefinitions,andtestfilestounderstandtheexpectedAPI
\+3\.Summarizetherequiredmethodsignatures,constructorparameters,andclasshierarchies
\+4\.Onlyafterthoroughlyunderstandingtheexistingcodestructure,proceedwithyourimplementation
\+5\.Implementexactlywhatthestubsandtestsexpect\-neverinventnewinterfaces
\+
\+\*\*Reminder:\*\*ThecodingagentneedstodealwithdifferentlanguagesincludingC\+\+,Go,Java,JavaScript,Python,andRust\.Adheretotheexactmethodsignatures,classhierarchies,andconstructorrequirementsdefinedintheexistingcode,regardlessofprogramminglanguage\.
\+"""
\+instruction=f"\{task\}\\n\\n\{inspection\_directive\}\\nPleaseanalyzetheproblemdescriptioncarefully\.Thenmakeeditstothecodefilestocompletetheinstruction\."
chat\_history,n\_llm\_calls\_used=chat\_with\_agent\(
instruction,
model=self\.code\_model,
diff\-\-gita/coding\_agent\.pyb/coding\_agent\.py
indexd824c3d\.\.25f5411100644
\-\-\-a/coding\_agent\.py
\+\+\+b/coding\_agent\.py
@@\-1,6\+1,7@@
\#Thisfileisadaptedfromhttps://github\.com/jennyzzt/dgm\.
importargparse
\+importjson
importlogging
importos
importsubprocess
@@\-160,41\+161,242@@classAgenticSystem:
returnnew\_msg\_history
\+defextract\_json\_from\_response\(self,response\_content\):
\+"""
\+ExtractJSONcontentfromLLMresponsewhichmaycontainmarkdownorothertext\.
\+"""
\+\#TrytofindJSONincodeblocks
\+importre
\+
\+json\_pattern=r"‘‘‘\(?:json\)?\\s\*\(\\\{\.\*?\\\}\)\\s\*‘‘‘"
\+match=re\.search\(json\_pattern,response\_content,re\.DOTALL\)
\+ifmatch:
\+try:
\+returnjson\.loads\(match\.group\(1\)\)
\+exceptjson\.JSONDecodeError:
\+pass
\+
\+\#TrytofindaJSONobjectdirectly
\+try:
\+\#Findthefirstopeningbraceandlastclosingbrace
\+start=response\_content\.find\("\{"\)
\+end=response\_content\.rfind\("\}"\)\+1
\+ifstart\>=0andend\>start:
\+json\_str=response\_content\[start:end\]
\+returnjson\.loads\(json\_str\)
\+except\(json\.JSONDecodeError,ValueError\):
\+pass
\+
\+\#Ifallelsefails,returntherawcontent
\+returnresponse\_content
\+
defforward\(self,timeout\):
"""
TheforwardfunctionfortheAgenticSystem\.
\+Implementsatwo\-phaseworkflow:
\+1\.ContractExtractionPhase:Agentreadstest/stubfilesandoutputsastructuredcontract\_planJSON
\+2\.Implementation&VerificationPhase:Agentimplementssolutionadheringtotheplanwithself\-audit
"""
\-task=f"""Ihaveuploadedacoderepositoryinthedirectory\{self\.git\_dir\}\.Helpsolvethefollowingproblem\.
\+\#Phase1:ContractExtraction
\+contract\_plan\_phase\_instruction=f"""Ihaveuploadedacoderepositoryinthedirectory\{self\.git\_dir\}\.Helpsolvethefollowingproblem\.
<problem\_description\>
\{self\.problem\_statement\}
</problem\_description\>
\-Yourtaskistomakechangestothefilesinthe\{self\.git\_dir\}directorytoaddressthe<problem\_description\>\.Ihavealreadytakencareoftherequireddependencies\.
\+\#PHASE1:CONTRACTEXTRACTION
\+
\+YourFIRSTtaskistoextracttheexactAPIcontractfromtherepository\.DoNOTwriteanyimplementationcodeinthisphase\.
\+
\+\*\*InstructionsforPhase1:\*\*
\+
\+1\.\*\*Exploretherepositorystructure\*\*usingthe‘editor‘tooltounderstandthefilelayout\.
\+
\+2\.\*\*ReadALLrelevantfiles\*\*:
\+\-Testfiles\(filescontainingteststhatdefineexpectedbehavior\)
\+\-Stub/Interfacefiles\(filescontainingclassdefinitions,methodsignatures,typehints\)
\+\-Anyspecificationordocumentationfiles
\+
\+3\.\*\*Extracttheexactcontract\*\*whichincludes:
\+\-\*\*Classnamesandtheirexacthierarchy\*\*\(parentclasses,interfacesimplemented\)
\+\-\*\*Constructorsignatures\*\*\(exactparameternames,types,andorder\)
\+\-\*\*Methodsignatures\*\*\(exactparameternames,types,returntypes\)
\+\-\*\*Errormessages\*\*\(verbatimerrorstringsthatmustberaised\)
\+\-\*\*Module\-levelfunctions\*\*\(exactsignaturesifapplicable\)
\+\-\*\*Filepaths\*\*thatneedmodification
\+
\+4\.\*\*OutputyourfindingsasaJSONobject\*\*withthefollowingstructure:
\+
\+‘‘‘json
\+\{\{
\+"classes":\[
\+\{\{
\+"name":"ClassName",
\+"parent":"ParentClassNameornull",
\+"fields":\[
\+\{\{"name":"field\_name","type":"Type"\}\}
\+\],
\+"methods":\[
\+\{\{
\+"name":"method\_name",
\+"parameters":\[
\+\{\{"name":"param","type":"Type","required":true\}\}
\+\],
\+"return\_type":"ReturnType",
\+"is\_constructor":false
\+\}\}
\+\],
\+"file\_path":"path/to/file"
\+\}\}
\+\],
\+"functions":\[
\+\{\{
\+"name":"function\_name",
\+"parameters":\[
\+\{\{"name":"param","type":"Type","required":true\}\}
\+\],
\+"return\_type":"ReturnType",
\+"file\_path":"path/to/file"
\+\}\}
\+\],
\+"error\_messages":\[
\+\{\{"message":"exacterrorstringfromtests","context":"whenthiserrorisraised"\}\}
\+\],
\+"files\_to\_modify":\["path/to/file1","path/to/file2"\],
\+"test\_files":\["path/to/test\_file1","path/to/test\_file2"\],
\+"key\_observations":\[
\+"Observation1:e\.g\.,Constructortakesnoarguments",
\+"Observation2:e\.g\.,Error’NotImplementedError’mustberaisedwithmessage’featurenotimplemented’"
\+\]
\+\}\}
\+‘‘‘
\+
\+\*\*CRITICALREQUIREMENTSFORPHASE1:\*\*
\+\-ExtractsignaturesEXACTLYasdefinedinthetest/stubfiles
\+\-DoNOTinferorguessmissinginformation\-ifsomethingisunclear,noteitinkey\_observations
\+\-Preserveexacterrormessagestrings\(case\-sensitive,includingpunctuation\)
\+\-NotetheEXACTfilepathsthatcontainthecontracts
\+\-ThisphaseisforANALYSISONLY\-donotwriteanyimplementationcode
\+
\+OutputonlytheJSONobjectandabriefsummaryofyourfindings\.Noimplementationcodeshouldbegeneratedinthisphase\.
"""
\-inspection\_directive="""
\-\*\*CRITICALPRE\-IMPLEMENTATIONREQUIREMENT:\*\*
\-Beforewritinganycode,carefullyinspecttheexistingrepositorystructure\.Usethe‘editor‘toolto‘view‘allrelevantstubfiles,testfiles,andrelatedclasses\.Extractandstrictlyadheretotheirexactmethodsignatures,classhierarchies,andconstructorrequirements\.Donotassumeorinventinterfaces;implementexactlywhattheexistingcodeexpects\.
\+safe\_log\("="\*50\)
\+safe\_log\("PHASE1:CONTRACTEXTRACTION"\)
\+safe\_log\("="\*50\)
\+
\+\#Phase1:Extractcontractfromtest/stubfiles
\+contract\_chat\_history,contract\_calls=chat\_with\_agent\(
\+contract\_plan\_phase\_instruction,
\+model=self\.code\_model,
\+msg\_history=\[\],
\+logging=safe\_log,
\+timeout=timeout,
\+\)
\+
\+\#ExtracttheJSONcontractplanfromtheresponse
\+contract\_response=""
\+ifcontract\_chat\_history:
\+\#Getthelastassistantmessage
\+formsginreversed\(contract\_chat\_history\):
\+ifisinstance\(msg,dict\):
\+ifmsg\.get\("role"\)=="assistant":
\+content=msg\.get\("content",""\)
\+ifcontent:
\+contract\_response=content
\+break
\+else:
\+\#Handlenon\-dictresponseobjects
\+contract\_response=str\(msg\)
\+break
\+
\+contract\_plan=self\.extract\_json\_from\_response\(contract\_response\)
\+
\+ifnotisinstance\(contract\_plan,dict\):
\+safe\_log\(f"Warning:CouldnotextractvalidJSONcontractplan\.Rawresponse:\{contract\_response\[:500\]\}"\)
\+contract\_plan=\{
\+"classes":\[\],
\+"functions":\[\],
\+"error\_messages":\[\],
\+"files\_to\_modify":\[\],
\+"test\_files":\[\],
\+"key\_observations":\[
\+"Couldnotparsestructuredcontract\.Agentresponseshouldbereviewed\."
\+\],
\+\}
\+
\+safe\_log\(f"Extractedcontractplanwith\{len\(contract\_plan\.get\(’classes’,\[\]\)\)\}classesand\{len\(contract\_plan\.get\(’functions’,\[\]\)\)\}functions"\)
\+safe\_log\(f"Filestomodify:\{contract\_plan\.get\(’files\_to\_modify’,\[\]\)\}"\)
\+safe\_log\(f"Keyobservations:\{contract\_plan\.get\(’key\_observations’,\[\]\)\}"\)
\+
\+\#Phase2:Implementation&Verification
\+phase2\_instruction=f"""Ihaveuploadedacoderepositoryinthedirectory\{self\.git\_dir\}\.Helpsolvethefollowingproblem\.
\+
\+<problem\_description\>
\+\{self\.problem\_statement\}
\+</problem\_description\>
\+
\+\#PHASE2:IMPLEMENTATION&VERIFICATION
\-\*\*SpecificInstructions:\*\*
\-1\.First,exploretherepositorystructureusing‘view‘commandstounderstandthefilelayout
\-2\.Readallstubfiles,interfacedefinitions,andtestfilestounderstandtheexpectedAPI
\-3\.Summarizetherequiredmethodsignatures,constructorparameters,andclasshierarchies
\-4\.Onlyafterthoroughlyunderstandingtheexistingcodestructure,proceedwithyourimplementation
\-5\.Implementexactlywhatthestubsandtestsexpect\-neverinventnewinterfaces
\+\#\#CONTRACTPLAN\(fromPhase1\)
\-\*\*Reminder:\*\*ThecodingagentneedstodealwithdifferentlanguagesincludingC\+\+,Go,Java,JavaScript,Python,andRust\.Adheretotheexactmethodsignatures,classhierarchies,andconstructorrequirementsdefinedintheexistingcode,regardlessofprogramminglanguage\.
\+Basedonanalysisoftestfilesandstubfiles,hereistheexactAPIcontractyouMUSTadhereto:
\+
\+‘‘‘json
\+\{json\.dumps\(contract\_plan,indent=2\)\}
\+‘‘‘
\+
\+\#\#YOURTASK
\+
\+Nowyoumustimplementthesolution\.YouareSTRICTLYBOUNDbythecontractabove\.
\+
\+\#\#\#Step1:ReviewtheContract
\+\-Carefullyreadthroughthecontract\_plan
\+\-Identifyallclasses,methods,constructors,anderrormessagesthatneedtobeimplemented
\+\-Notetheexactfilepathswheremodificationsareneeded
\+
\+\#\#\#Step2:ImplementtheSolution
\+\-Usethe‘editor‘tooltomakechangestothefileslistedin‘files\_to\_modify‘
\+\-Followtheexactsignaturesfromthecontract\-doNOTinventnewmethodsorchangeexistingsignatures
\+\-ImplementerrorhandlingthatraisestheEXACTerrormessagesspecifiedin‘error\_messages‘
\+\-Ensureallconstructorsmatchtheparametersexactlyasspecified
\+
\+\#\#\#Step3:Self\-AuditChecklist
\+Beforeoutputtingyourpatch,youMUSTperformaself\-audit\.Reviewyourimplementationagainstthischecklist:
\+
\+1\.\[\]\*\*ClassNames\*\*:Doallclassnamesmatchexactly\(includinginheritance\)?
\+2\.\[\]\*\*ConstructorSignatures\*\*:Doallconstructorsignaturesmatchthecontractparametersexactly?
\+3\.\[\]\*\*MethodSignatures\*\*:Doallmethodsignatures\(name,parameters,returntypes\)match?
\+4\.\[\]\*\*ErrorMessages\*\*:DoallerrorraisesusetheEXACTerrorstringsspecified?
\+5\.\[\]\*\*FilePaths\*\*:Areallmodificationsmadetothecorrectfiles?
\+6\.\[\]\*\*NoExtraCode\*\*:Haveyouavoidedaddingmethodsorfeaturesnotinthecontract?
\+
\+Foreachitem,explicitlystatePASSorFAILandprovideevidencefromyourimplementation\.
\+
\+\#\#\#Step4:OutputtheFinalPatch
\+Aftertheself\-audit\(allitemsmustPASS\),outputyourfinalimplementation\.Makesure:
\+\-Youusethe‘editor‘tooltoeditthefiles
\+\-Theimplementationstrictlyfollowsthecontract
\+\-Allmethodsareproperlyimplementedtopassthetests
\+
\+\*\*IMPORTANT\*\*:DoNOTwriteanycodethatdeviatesfromthecontract\.Everymethodsignatureanderrormessagemustmatchexactly\.
"""
\-instruction=f"\{task\}\\n\\n\{inspection\_directive\}\\nPleaseanalyzetheproblemdescriptioncarefully\.Thenmakeeditstothecodefilestocompletetheinstruction\."
\-chat\_history,n\_llm\_calls\_used=chat\_with\_agent\(
\-instruction,
\+
\+safe\_log\("="\*50\)
\+safe\_log\("PHASE2:IMPLEMENTATION&VERIFICATION"\)
\+safe\_log\("="\*50\)
\+
\+\#Phase2:Implementandverify
\+implementation\_chat\_history,impl\_calls=chat\_with\_agent\(
\+phase2\_instruction,
model=self\.code\_model,
\-msg\_history=\[\],
\+msg\_history=contract\_chat\_history,\#Continuefromphase1
logging=safe\_log,
timeout=timeout,
\)
\-chat\_history\_str=str\(chat\_history\)
\+
\+\#Returnthefullchathistoryforlogging
\+returnimplementation\_chat\_history
defmain\(\):
diff\-\-gita/coding\_agent\.pyb/coding\_agent\.py
index25f5411\.\.e98de16100644
\-\-\-a/coding\_agent\.py
\+\+\+b/coding\_agent\.py
@@\-193,9\+193,10@@classAgenticSystem:
defforward\(self,timeout\):
"""
TheforwardfunctionfortheAgenticSystem\.
\-Implementsatwo\-phaseworkflow:
\+Implementsamulti\-phaseworkflow:
1\.ContractExtractionPhase:Agentreadstest/stubfilesandoutputsastructuredcontract\_planJSON
\-2\.Implementation&VerificationPhase:Agentimplementssolutionadheringtotheplanwithself\-audit
\+2\.ImplementationPhase:Agentimplementssolutionadheringtotheplan
\+3\.VerificationLoop:Agentrunstests,analyzesresults,anditeratesuntilpassingormaxattempts
"""
\#Phase1:ContractExtraction
contract\_plan\_phase\_instruction=f"""Ihaveuploadedacoderepositoryinthedirectory\{self\.git\_dir\}\.Helpsolvethefollowingproblem\.
@@\-329,14\+330,14@@OutputonlytheJSONobjectandabriefsummaryofyourfindings\.Noimplementat
safe\_log\(f"Filestomodify:\{contract\_plan\.get\(’files\_to\_modify’,\[\]\)\}"\)
safe\_log\(f"Keyobservations:\{contract\_plan\.get\(’key\_observations’,\[\]\)\}"\)
\-\#Phase2:Implementation&Verification
\+\#Phase2:Implementation
phase2\_instruction=f"""Ihaveuploadedacoderepositoryinthedirectory\{self\.git\_dir\}\.Helpsolvethefollowingproblem\.
<problem\_description\>
\{self\.problem\_statement\}
</problem\_description\>
\-\#PHASE2:IMPLEMENTATION&VERIFICATION
\+\#PHASE2:IMPLEMENTATION
\#\#CONTRACTPLAN\(fromPhase1\)
@@\-373,20\+374,17@@Beforeoutputtingyourpatch,youMUSTperformaself\-audit\.Reviewyourimpleme
Foreachitem,explicitlystatePASSorFAILandprovideevidencefromyourimplementation\.
\-\#\#\#Step4:OutputtheFinalPatch
\-Aftertheself\-audit\(allitemsmustPASS\),outputyourfinalimplementation\.Makesure:
\-\-Youusethe‘editor‘tooltoeditthefiles
\-\-Theimplementationstrictlyfollowsthecontract
\-\-Allmethodsareproperlyimplementedtopassthetests
\+\#\#\#Step4:OutputtheFinalImplementation
\+Aftertheself\-audit\(allitemsmustPASS\),makeyourfinalimplementationusingthe‘editor‘tool\.
\*\*IMPORTANT\*\*:DoNOTwriteanycodethatdeviatesfromthecontract\.Everymethodsignatureanderrormessagemustmatchexactly\.
"""
safe\_log\("="\*50\)
\-safe\_log\("PHASE2:IMPLEMENTATION&VERIFICATION"\)
\+safe\_log\("PHASE2:IMPLEMENTATION"\)
safe\_log\("="\*50\)
\-\#Phase2:Implementandverify
\+\#Phase2:Implement
implementation\_chat\_history,impl\_calls=chat\_with\_agent\(
phase2\_instruction,
model=self\.code\_model,
@@\-395,8\+393,129@@Aftertheself\-audit\(allitemsmustPASS\),outputyourfinalimplementation\.Ma
timeout=timeout,
\)
\+\#Phase3:VerificationLoop
\+current\_chat\_history=implementation\_chat\_history
\+max\_iterations=5
\+
\+\#Determineifweshouldruntestsbasedonlanguage
\+should\_run\_tests=True
\+
\+ifshould\_run\_tests:
\+\#Preparethecurrenteditsfortheagenttosee
\+current\_edits\_msg=self\.get\_current\_edits\(\)
\+
\+foriterationinrange\(1,max\_iterations\+1\):
\+safe\_log\("="\*50\)
\+safe\_log\(f"PHASE3:VERIFICATION\-Iteration\{iteration\}/\{max\_iterations\}"\)
\+safe\_log\("="\*50\)
\+
\+\#Runtests
\+test\_result=self\.run\_tests\(\)
\+test\_output=test\_result\.get\("output",""\)
\+test\_exit\_code=test\_result\.get\("exit\_code",\-1\)
\+
\+safe\_log\(f"Testexitcode:\{test\_exit\_code\}"\)
\+
\+iftest\_exit\_code==0:
\+\#Testspass\!
\+safe\_log\("Alltestspassed\!"\)
\+break
\+
\+\#Testsfailed\-providerawoutputtoagentforanalysis
\+verification\_instruction=f"""\#\#TESTRESULTS\(Iteration\{iteration\}\)
\+
\+Testsfailedwithexitcode\{test\_exit\_code\}\.Hereistherawtestoutput\-analyzeittounderstandwhatwentwrong:
\+
\+\{test\_output\}
\+
\+\#\#REMAININGCONTRACT
\+
\+‘‘‘json
\+\{json\.dumps\(contract\_plan,indent=2\)\}
\+‘‘‘
\+
\+\#\#YOURTASK
\+
\+1\.\*\*Analyzethetestoutputabove\*\*tounderstandwhat’sfailing
\+2\.\*\*Identifywhichspecifictestsorfunctionalityarefailing\*\*
\+3\.\*\*Usethe‘editor‘tooltofixtheimplementation\*\*
\+4\.\*\*Focusonfixingtheissuesrevealedbythetestoutput\*\*
\+
\+Maketargetedfixestopassthefailingtests\.Thenwewillruntestsagain\.
\+"""
\+
\+\#Addcurrenteditscontext
\+ifcurrent\_edits\_msgandcurrent\_edits\_msg\!=\[\]:
\+verification\_instruction\+="\\n\\n\#\#CURRENTCHANGES\\n"
\+ifisinstance\(current\_edits\_msg,list\)andlen\(current\_edits\_msg\)\>0:
\+edit\_msg=current\_edits\_msg\[\-1\]ifisinstance\(current\_edits\_msg\[\-1\],dict\)elsecurrent\_edits\_msg
\+ifisinstance\(edit\_msg,dict\)andedit\_msg\.get\("content"\):
\+content=edit\_msg\["content"\]
\+ifisinstance\(content,list\):
\+foritemincontent:
\+ifisinstance\(item,dict\)anditem\.get\("type"\)=="input\_text":
\+verification\_instruction\+=item\.get\("text",""\)
\+elifisinstance\(item,str\):
\+verification\_instruction\+=item
\+else:
\+verification\_instruction\+=str\(content\)
\+
\+\#Continuewithagenttofixissues
\+current\_chat\_history,\_=chat\_with\_agent\(
\+verification\_instruction,
\+model=self\.code\_model,
\+msg\_history=current\_chat\_history,
\+logging=safe\_log,
\+timeout=timeout,
\+\)
\+
\+safe\_log\(f"Agentmadechangesiniteration\{iteration\}\.Testingagain\.\.\."\)
\+
\#Returnthefullchathistoryforlogging
\-returnimplementation\_chat\_history
\+returncurrent\_chat\_history
\+
\+defrun\_tests\(self\):
\+"""
\+Runtestsforthecurrentrepository\.
\+Returnsadictwith’output’\(str\)and’exit\_code’\(int\)\.
\+"""
\+try:
\+\#Importthetestutility
\+fromutils\.test\_utilsimportrun\_test\_command
\+
\+\#Getthetestcommandforthelanguage
\+test\_cmd=TEST\_COMMANDS\.get\(self\.language,\[\["pytest","\-rA","\-\-tb=short"\]\]\)
\+
\+\#Buildthecommandasastringforbash
\+ifisinstance\(test\_cmd,list\):
\+iflen\(test\_cmd\)==1andisinstance\(test\_cmd\[0\],list\):
\+cmd\_parts=test\_cmd\[0\]
\+else:
\+cmd\_parts=test\_cmd
\+else:
\+cmd\_parts=test\_cmd
\+
\+\#Converttobashcommandstring
\+ifisinstance\(cmd\_parts\[0\],list\):
\+bash\_cmd="&&"\.join\(\[’’\.join\(cmd\)forcmdincmd\_parts\]\)
\+else:
\+bash\_cmd="&&"\.join\(cmd\_parts\)
\+
\+\#Runthetestcommandusingtheutility
\+result=run\_test\_command\(bash\_cmd,cwd=self\.git\_dir\)
\+
\+return\{
\+"output":result\["output"\],
\+"exit\_code":result\["exit\_code"\]
\+\}
\+exceptExceptionase:
\+safe\_log\(f"Errorrunningtests:\{str\(e\)\}"\)
\+return\{
\+"output":f"Errorrunningtests:\{str\(e\)\}",
\+"exit\_code":\-1
\+\}
defmain\(\):
diff\-\-gita/tests/test\_test\_utils\.pyb/tests/test\_test\_utils\.py
newfilemode100644
index0000000\.\.8ac038f
\-\-\-/dev/null
\+\+\+b/tests/test\_test\_utils\.py
@@\-0,0\+1,119@@
\+\#Testsfortest\_utils\.py
\+
\+importpytest
\+fromutils\.test\_utilsimportparse\_test\_output
\+
\+
\+classTestParseTestOutput:
\+"""Testsforparse\_test\_outputfunction\."""
\+
\+deftest\_all\_tests\_passed\(self\):
\+"""Testparsingoutputwhenalltestspass\."""
\+output="""
\+=============================testsessionstarts==============================
\+platformlinux\-\-Python3\.10\.20,pytest\-9\.0\.3,pluggy\-1\.6\.0
\+rootdir:/test
\+collected5items
\+
\+tests/test\_a\.py\.\.\[40%\]
\+tests/test\_b\.py\.\.\.\[100%\]
\+
\+==============================5passedin0\.1s==============================
\+"""
\+result=parse\_test\_output\(output,0\)
\+assertresult\["passed"\]==5
\+assertresult\["failed"\]==0
\+assertresult\["error"\]==0
\+assertresult\["success"\]==True
\+assertresult\["exit\_code"\]==0
\+
\+deftest\_some\_tests\_failed\(self\):
\+"""Testparsingoutputwhensometestsfail\."""
\+output="""
\+=============================testsessionstarts==============================
\+platformlinux\-\-Python3\.10\.20,pytest\-9\.0\.3,pluggy\-1\.6\.0
\+rootdir:/test
\+collected10items
\+
\+tests/test\_a\.py\.\.F\.\.\.\[80%\]
\+tests/test\_b\.py\.\.\.\[100%\]
\+
\+===================================FAILURES===================================
\+\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_TestClass\.test\_failed\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_
\+
\+deftest\_failed\(self\):
\+\>assertFalse
\+EassertFalse
\+
\+tests/test\_a\.py:6:AssertionError
\+===========================1failed,9passedin0\.2s===========================
\+"""
\+result=parse\_test\_output\(output,1\)
\+assertresult\["passed"\]==9
\+assertresult\["failed"\]==1
\+assertresult\["success"\]==False
\+assertresult\["exit\_code"\]==1
\+
\+deftest\_tests\_with\_errors\(self\):
\+"""Testparsingoutputwhenthereareerrors\."""
\+output="""
\+=============================testsessionstarts==============================
\+platformlinux\-\-Python3\.10\.20,pytest\-9\.0\.3,pluggy\-1\.6\.0
\+rootdir:/test
\+collected5items
\+
\+tests/test\_a\.pyE\.\.\.\[80%\]
\+tests/test\_b\.py\.\.\[100%\]
\+
\+====================================ERRORS====================================
\+\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_ERRORatsetupoftest\_error\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_
\+FileNotFoundError:\[Errno2\]Nosuchfileordirectory
\+===========================1error,4passedin0\.1s===========================
\+"""
\+result=parse\_test\_output\(output,1\)
\+assertresult\["passed"\]==4
\+assertresult\["error"\]==1
\+assertresult\["success"\]==False
\+
\+deftest\_tests\_with\_skipped\(self\):
\+"""Testparsingoutputwhenthereareskippedtests\."""
\+output="""
\+=============================testsessionstarts==============================
\+platformlinux\-\-Python3\.10\.20,pytest\-9\.0\.3,pluggy\-1\.6\.0
\+rootdir:/test
\+collected5items
\+
\+tests/test\_a\.py\.\.s\.\.\[100%\]
\+
\+===========================4passed,1skippedin0\.1s===========================
\+"""
\+result=parse\_test\_output\(output,0\)
\+assertresult\["passed"\]==4
\+assertresult\["skipped"\]==1
\+assertresult\["success"\]==True
\+
\+deftest\_empty\_output\(self\):
\+"""Testparsingemptyoutput\."""
\+result=parse\_test\_output\("",0\)
\+assertresult\["passed"\]==0
\+assertresult\["failed"\]==0
\+assertresult\["success"\]==True
\+
\+deftest\_exit\_code\_parameter\(self\):
\+"""Testexitcodeparameterisused\."""
\+output="testcompleted"
\+result=parse\_test\_output\(output,5\)
\+assertresult\["exit\_code"\]==5
\+assertresult\["success"\]==False
\+
\+deftest\_success\_summary\(self\):
\+"""Testsummaryforpassingtests\."""
\+output="5passedin0\.1s"
\+result=parse\_test\_output\(output,0\)
\+assertresult\["summary"\]=="Testspassed"
\+
\+deftest\_failure\_summary\(self\):
\+"""Testsummaryforfailingtests\."""
\+output="2failed,8passedin0\.2s"
\+result=parse\_test\_output\(output,1\)
\+assert"2test\(s\)failed"inresult\["summary"\]
diff\-\-gita/utils/test\_utils\.pyb/utils/test\_utils\.py
newfilemode100644
index0000000\.\.3284dda
\-\-\-/dev/null
\+\+\+b/utils/test\_utils\.py
@@\-0,0\+1,112@@
\+\#Utilityfunctionsforrunningtestsandanalyzingresults\.
\+
\+importre
\+fromtypingimportDict,Any
\+
\+
\+defparse\_test\_output\(output:str,exit\_code:int=0\)\-\>Dict\[str,Any\]:
\+"""
\+Parsetestoutputtoextractrelevantinformation\.
\+Returnsadictwith’passed’,’failed’,’error\_count’,’summary’\.
\+
\+Args:
\+output:Rawtestoutputstring
\+exit\_code:Exitcodefromtestexecution
\+
\+Returns:
\+Dictwithparsedtestresults
\+"""
\+result=\{
\+"passed":0,
\+"failed":0,
\+"error":0,
\+"skipped":0,
\+"exit\_code":exit\_code,
\+"success":exit\_code==0,
\+"summary":"",
\+"raw\_output":output
\+\}
\+
\+\#Countpatterns
\+passed\_match=re\.search\(r"\(\\d\+\)\\s\+passed",output\)
\+ifpassed\_match:
\+result\["passed"\]=int\(passed\_match\.group\(1\)\)
\+
\+failed\_match=re\.search\(r"\(\\d\+\)\\s\+failed",output\)
\+iffailed\_match:
\+result\["failed"\]=int\(failed\_match\.group\(1\)\)
\+
\+error\_match=re\.search\(r"\(\\d\+\)\\s\+error",output,re\.IGNORECASE\)
\+iferror\_match:
\+result\["error"\]=int\(error\_match\.group\(1\)\)
\+
\+skipped\_match=re\.search\(r"\(\\d\+\)\\s\+skipped",output\)
\+ifskipped\_match:
\+result\["skipped"\]=int\(skipped\_match\.group\(1\)\)
\+
\+\#Determinesuccessfromexitcodeandcontent
\+ifexit\_code==0:
\+result\["success"\]=True
\+result\["summary"\]="Testspassed"
\+else:
\+result\["success"\]=False
\+ifresult\["failed"\]\>0:
\+result\["summary"\]=f"\{result\[’failed’\]\}test\(s\)failed"
\+elifresult\["error"\]\>0:
\+result\["summary"\]=f"\{result\[’error’\]\}error\(s\)occurred"
\+else:
\+result\["summary"\]=f"Testsfailedwithexitcode\{exit\_code\}"
\+
\+returnresult
\+
\+
\+defrun\_test\_command\(command:str,cwd:str=None\)\-\>Dict\[str,Any\]:
\+"""
\+Runatestcommandandreturntheresults\.
\+
\+Args:
\+command:Testcommandtorun
\+cwd:Workingdirectory
\+
\+Returns:
\+Dictwith’output’,’exit\_code’,andparsedresults
\+"""
\+try:
\+fromtools\.bashimporttool\_function
\+
\+\#Executethecommand
\+ifcwd:
\+full\_command=f"cd\{cwd\}&&\{command\}"
\+else:
\+full\_command=command
\+
\+output=tool\_function\(full\_command\)
\+
\+\#Trytoextractexitcodefromoutput
\+exit\_code=0
\+exit\_match=re\.search\(r"exitcode\[:\\s\]\+\(\\d\+\)",output,re\.IGNORECASE\)
\+ifexit\_match:
\+exit\_code=int\(exit\_match\.group\(1\)\)
\+
\+\#Parsetheoutput
\+test\_result=parse\_test\_output\(output,exit\_code\)
\+
\+return\{
\+"output":output,
\+"exit\_code":test\_result\["exit\_code"\],
\+"success":test\_result\["success"\],
\+"summary":test\_result\["summary"\],
\+"passed":test\_result\["passed"\],
\+"failed":test\_result\["failed"\],
\+"error":test\_result\["error"\],
\+\}
\+exceptExceptionase:
\+return\{
\+"output":f"Errorrunningtests:\{str\(e\)\}",
\+"exit\_code":\-1,
\+"success":False,
\+"summary":f"Testexecutionfailed:\{str\(e\)\}",
\+"passed":0,
\+"failed":0,
\+"error":0,
\+\}
Listing 2:Figure 18: Code modification on SWE\-benchdiff
indexd19efd6\.\.98f364a100644
\-\-\-a/coding\_agent\.py
\+\+\+b/coding\_agent\.py
@@\-3,16\+3,19@@
importargparse
importlogging
importos
\+importre
importsubprocess
importthreading
fromlogging\.handlersimportRotatingFileHandler
fromtimeimporttime
fromllm\_withtoolsimport\(CLAUDE\_MODEL,OPENAI\_MODEL,chat\_with\_agent,
\-convert\_msg\_history\)
\-fromutils\.eval\_utilsimport\(get\_report\_score,msg\_history\_to\_report,
\-score\_tie\_breaker\)
\+convert\_msg\_history,set\_explored\_files\_callback,
\+clear\_explored\_files\_callback\)
\+fromutils\.eval\_utilsimport\(enforce\_diff\_minimality,get\_report\_score,
\+msg\_history\_to\_report,score\_tie\_breaker\)
fromutils\.git\_utilsimportapply\_patch,diff\_versus\_commit,reset\_to\_commit
\+fromutils\.trace\_utilsimportextract\_test\_function\_names,generate\_trace\_instruction
\#Thread\-localstorageforloggerinstances
thread\_local=threading\.local\(\)
@@\-96,6\+99,9@@classAgenticSystem:
self\.instance\_id=instance\_idifnotself\_improveelse"hgm"
self\.code\_model=model
\+\#Trackallfilesexploredbytheagentviatheeditortool’sviewcommand
\+self\.explored\_files=\[\]
\+
\#Initializeloggerandstoreitinthread\-localstorage
self\.logger=setup\_logger\(chat\_history\_file\)
@@\-172,12\+178,88@@Yourtaskistoruntheregressiontestsinthe\{self\.git\_tempdir\}directoryto
\)
returntest\_report
\+def\_extract\_test\_failures\_from\_history\(self,msg\_history\):
\+failed\_tests=\[\]
\+formsginmsg\_history:
\+ifnotisinstance\(msg,dict\):
\+continue
\+content=msg\.get\("content"\)ormsg\.get\("output"\)or""
\+ifnotisinstance\(content,str\):
\+content=str\(content\)
\+forpatternin\(
\+r"FAILED\\s\+\\S\+?\\\.py::\(test\_\\w\+\)",
\+r"\\b\(test\_\\w\+\)\\b\.\*FAILED",
\+\):
\+formatchinre\.findall\(pattern,content,re\.IGNORECASE\):
\+name=match\[0\]ifisinstance\(match,tuple\)elsematch
\+ifname\.startswith\("test\_"\)andnamenotinfailed\_tests:
\+failed\_tests\.append\(name\)
\+returnfailed\_tests
\+
\+def\_build\_trace\_context\(self,test\_names\):
\+ifnottest\_names:
\+return""
\+blocks=\[\]
\+fortest\_nameintest\_names\[:2\]:
\+blocks\.append\(
\+generate\_trace\_instruction\(
\+instance\_id=self\.instance\_id,
\+test\_function\_name=test\_name,
\+target\_object=test\_name\.replace\("test\_",""\),
\+git\_tempdir=self\.git\_tempdir,
\+\)
\+\)
\+return"\\n"\.join\(blocks\)
\+
\+def\_apply\_diff\_minimality\(self,diff\_str\):
\+cleaned\_diff=enforce\_diff\_minimality\(
\+problem\_statement=self\.problem\_statement,
\+explored\_files=self\.explored\_files,
\+diff\_str=diff\_str,
\+git\_tempdir=self\.git\_tempdir,
\+\)
\+ifcleaned\_diff\!=diff\_str:
\+reset\_to\_commit\(self\.git\_tempdir,self\.base\_commit\)
\+result=subprocess\.run\(
\+\["git","\-C",self\.git\_tempdir,"apply","\-"\],
\+input=cleaned\_diff,
\+text=True,
\+capture\_output=True,
\+\)
\+ifresult\.returncode\!=0:
\+print\(f"Warning:Failedtoapplyminimizeddiff:\{result\.stderr\}"\)
\+else:
\+print\("Appliedminimizeddiffsuccessfully"\)
\+safe\_log\(f"Diffminimalitycomplete\.Exploredfiles:\{self\.explored\_files\}"\)
\+
defforward\(self,timeout=3600\):
timeout\-=60
start\_time=time\(\)
\-"""
\-TheforwardfunctionfortheAgenticSystem\.
\-"""
\+set\_explored\_files\_callback\(self\.explored\_files\.append\)
\+
\+hinted\_tests=extract\_test\_function\_names\(self\.test\_descriptionor""\)
\+hinted\_tests\+=extract\_test\_function\_names\(self\.problem\_statementor""\)
\+deduped=\[\]
\+seen=set\(\)
\+fortest\_nameinhinted\_tests:
\+iftest\_namenotinseen:
\+seen\.add\(test\_name\)
\+deduped\.append\(test\_name\)
\+hinted\_tests=deduped
\+
\+pre\_trace\_context=self\.\_build\_trace\_context\(hinted\_tests\)
\+trace\_preamble=""
\+ifpre\_trace\_context:
\+trace\_preamble=\(
\+f"\\n<code\_path\_trace\>\\n\{pre\_trace\_context\}\\n</code\_path\_trace\>\\n"
\+"Usethistracetolocatethecorrectfixpathbeforeediting\.\\n"
\+\)
\+else:
\+trace\_preamble=\(
\+"\\nBeforeediting,runthemostrelevantfailingtest\(s\)forthisissue"
\+f"in\{self\.git\_tempdir\}soyoucanobserveactualvsexpectedoutput\.\\n"
\+\)
\+
instruction=f"""IhaveuploadedaPythoncoderepositoryinthedirectory\{self\.git\_tempdir\}\.Helpsolvethefollowingproblem\.
<problem\_description\>
@@\-187,17\+269,47@@Yourtaskistoruntheregressiontestsinthe\{self\.git\_tempdir\}directoryto
<test\_description\>
\{self\.test\_description\}
</test\_description\>
\-
\-Yourtaskistomakechangestothefilesinthe\{self\.git\_tempdir\}directorytoaddressthe<problem\_description\>\.Ihavealreadytakencareoftherequireddependencies\.
\+\{trace\_preamble\}
\+Makechangesin\{self\.git\_tempdir\}toaddresstheproblem\.
"""
\-chat\_history,n\_llm\_calls\_used=chat\_with\_agent\(
\+msg\_history,\_=chat\_with\_agent\(
instruction,
model=self\.code\_model,
msg\_history=\[\],
logging=safe\_log,
timeout=timeout\-\(time\(\)\-start\_time\),
\)
\-chat\_history\_str=str\(chat\_history\)
\+
\+generic\_history=convert\_msg\_history\(msg\_history,self\.code\_model\)
\+failed\_tests=self\.\_extract\_test\_failures\_from\_history\(generic\_history\)
\+ifnotfailed\_testsandhinted\_tests:
\+failed\_tests=hinted\_tests\[:1\]
\+
\+forattemptinrange\(2\):
\+remaining=timeout\-\(time\(\)\-start\_time\)
\+ifnotfailed\_testsorremaining<=120:
\+break
\+trace\_instruction=self\.\_build\_trace\_context\(failed\_tests\[:1\]\)
\+follow\_up=f"""Thefollowingtest\(s\)failed:\{’,’\.join\(failed\_tests\)\}\.
\+
\+\{trace\_instruction\}
\+
\+Reviseyourfixin\{self\.git\_tempdir\}sothepatchaddressesthetracedcodepathandmakesthefailingtestpass\.
\+"""
\+msg\_history,\_=chat\_with\_agent\(
\+follow\_up,
\+model=self\.code\_model,
\+msg\_history=msg\_history,
\+logging=safe\_log,
\+timeout=remaining,
\+\)
\+generic\_history=convert\_msg\_history\(msg\_history,self\.code\_model\)
\+failed\_tests=self\.\_extract\_test\_failures\_from\_history\(generic\_history\)
\+ifnotfailed\_tests:
\+break
\+
\+clear\_explored\_files\_callback\(\)
\+self\.\_apply\_diff\_minimality\(self\.get\_current\_edits\(\)\)
defmain\(\):
diff\-\-gita/llm\_withtools\.pyb/llm\_withtools\.py
indexba7ea87\.\.b35bae3100644
\-\-\-a/llm\_withtools\.py
\+\+\+b/llm\_withtools\.py
@@\-13,6\+13,28@@importopenai
fromllmimportcreate\_client
fromtoolsimportload\_all\_tools
\+\#Module\-levelcallbackfortrackingexploredfiles
\+\_EXPLORED\_FILES\_CALLBACK=None
\+
\+
\+defset\_explored\_files\_callback\(callback\):
\+global\_EXPLORED\_FILES\_CALLBACK
\+\_EXPLORED\_FILES\_CALLBACK=callback
\+
\+
\+defclear\_explored\_files\_callback\(\):
\+global\_EXPLORED\_FILES\_CALLBACK
\+\_EXPLORED\_FILES\_CALLBACK=None
\+
\+
\+def\_track\_editor\_view\(tool\_name,tool\_input\):
\+"""Helpertotrackwheneditorviewiscalled\."""
\+if\_EXPLORED\_FILES\_CALLBACKandtool\_name=="editor":
\+command=tool\_input\.get\("command",""\)
\+path=tool\_input\.get\("path",""\)
\+ifcommand=="view"andpath:
\+\_EXPLORED\_FILES\_CALLBACK\(path\)
\+
CLAUDE\_MODEL="anthropic/claude\-sonnet\-4"
OPENAI\_MODEL="gpt\-5"
MAX\_XML\_TOOL\_FORMAT\_RETRIES=2
@@\-80,6\+102,9@@def\_assistant\_message\_with\_tool\_call\(message,tool\_use\):
defprocess\_tool\_call\(tools\_dict,tool\_name,tool\_input\):
\+\#Trackeditortoolusage
\+\_track\_editor\_view\(tool\_name,tool\_input\)
\+
try:
iftool\_nameintools\_dict:
returntools\_dict\[tool\_name\]\["function"\]\(\*\*tool\_input\)
diff\-\-gita/utils/eval\_utils\.pyb/utils/eval\_utils\.py
index1c6e117\.\.15a9d83100644
\-\-\-a/utils/eval\_utils\.py
\+\+\+b/utils/eval\_utils\.py
@@\-125,3\+125,167@@Yourresponsewillbeautomaticallyparsed,soensurethatthestringresponsei
exceptExceptionase:
logging\(f"Errorinscore\_tie\_breaker:\{e\}"\)
returnbest\_score\_index
\+
\+
\+def\_extract\_file\_from\_diff\_header\(line:str\)\-\>str:
\+"""
\+Extractthefilepathfromagitdiffheaderlinelike:
\+’diff\-\-gita/path/to/file\.pyb/path/to/file\.py’
\+Returnsthepathafter’a/’\(whichshouldmatchthepathafter’b/’\)\.
\+"""
\+\#Matchthepattern:diff\-\-gita/\.\.\.b/\.\.\.
\+importre
\+match=re\.search\(r’diff\-\-gita/\(\.\+?\)b/\(\.\+\)$’,line\)
\+ifmatch:
\+\#Returnthepathfromeitherside\(theyshouldbethesame\)
\+returnmatch\.group\(1\)\.split\(’/’\)\[\-1\]ormatch\.group\(2\)\.split\(’/’\)\[\-1\]
\+returnNone
\+
\+
\+def\_get\_relevant\_test\_files\(explored\_files:list,all\_files:set\)\-\>set:
\+"""
\+Identifytestfilesthatarelikelyrelatedtotheexploredfiles\.
\+Atestfileisrelevantif:
\+\-It’sina’test\_’prefixor’\_test’suffixpattern
\+\-It’sina’tests/’or’test/’directory
\+\-Thebasemodulenamematches\(e\.g\.,inspect\.py\-\>test\_inspect\.py\)
\+"""
\+relevant=set\(\)
\+forexploredinexplored\_files:
\+\#Getthefilenamewithoutextension
\+importos
\+base=os\.path\.splitext\(os\.path\.basename\(explored\)\)\[0\]
\+\#Checkifthisisatest\-relatedfile\(e\.g\.,test\_inspect\.pyforinspect\.py\)
\+\#Lookforfilesthathavetest\_prefixor\_testsuffixwithmatchingbasename
\+forfinall\_files:
\+f\_base=os\.path\.splitext\(os\.path\.basename\(f\)\)\[0\]
\+f\_dir=os\.path\.dirname\(f\)
\+\#Checktestdirectorypattern
\+ifos\.path\.basename\(f\_dir\)\.lower\(\)in\(’test’,’tests’\):
\+relevant\.add\(f\)
\+continue
\+\#Checknamingpattern:test\_<module\>\.pyor<module\>\_test\.py
\+if\(f\_base==f’test\_\{base\}’or
\+f\_base==f’\{base\}\_test’or
\+f==f’test\_\{base\}\.py’or
\+f==f’\{base\}\_test\.py’\):
\+relevant\.add\(f\)
\+returnrelevant
\+
\+
\+defenforce\_diff\_minimality\(
\+problem\_statement:str,
\+explored\_files:list,
\+diff\_str:str,
\+git\_tempdir:str=None,
\+\)\-\>str:
\+"""
\+Enforcesdiffminimalitybyfilteringoutchangestofilesthatarenot
\+relevanttotheproblemstatementorthefilesexploredbytheagent\.
\+
\+Thisfunction:
\+1\.Parsesthedifftoextractmodifiedfilepaths
\+2\.Computesarelevancescoreforeachfilebasedon:
\+\-Whetherthefileappearsintheproblemstatement\(exactmatchorkeywordmatch\)
\+\-Whetherthefilewasexploredbytheagent
\+\-Whetherit’satestfilerelatedtoanexploredfile
\+3\.Returnsafiltereddiffcontainingonlyhigh\-relevancefiles
\+
\+Args:
\+problem\_statement:Theoriginalproblemstatement/issuedescription
\+explored\_files:Listoffilepathsthattheagentviewedduringexecution
\+diff\_str:Thecompletediffstringtofilter
\+git\_tempdir:Optionalpathtothegitrepository\(usedtofindallfiles\)
\+
\+Returns:
\+Afiltereddiffstringcontainingonlychangestorelevantfiles
\+"""
\+importos
\+ifnotdiff\_str:
\+returndiff\_str
\+
\+\#Normalizeproblemstatementformatching
\+problem\_lower=problem\_statement\.lower\(\)
\+
\+\#Extractfilepathsfromthediff
\+importre
\+diff\_lines=diff\_str\.split\(’\\n’\)
\+diff\_blocks=\[\]
\+current\_block=\[\]
\+
\+forlineindiff\_lines:
\+ifline\.startswith\(’diff\-\-git’\):
\+ifcurrent\_block:
\+diff\_blocks\.append\(’\\n’\.join\(current\_block\)\)
\+current\_block=\[line\]
\+else:
\+current\_block\.append\(line\)
\+ifcurrent\_block:
\+diff\_blocks\.append\(’\\n’\.join\(current\_block\)\)
\+
\+\#Analyzeeachdiffblock
\+relevant\_blocks=\[\]
\+forblockindiff\_blocks:
\+ifnotblock\.strip\(\):
\+continue
\+
\+\#Getthefilepathfromthediffheader
\+header\_line=block\.split\(’\\n’\)\[0\]
\+filename=\_extract\_file\_from\_diff\_header\(header\_line\)
\+iffilenameisNone:
\+\#Ifwecan’tparsetheheader,keeptheblocktobesafe
\+relevant\_blocks\.append\(block\)
\+continue
\+
\+\#Computerelevancescore
\+score=0
\+
\+\#Checkifthefilewasexploredbytheagent
\+forexploredinexplored\_files:
\+explored\_name=os\.path\.basename\(explored\)ifexploredelseNone
\+ifexplored\_name==filename:
\+score\+=10\#Highrelevance\-wasexplored
\+break
\+\#Alsocheckifthepathcontainsthefilename
\+iffilenameinexplored:
\+score\+=5
\+
\+\#Checkiffilenameappearsinproblemstatement
\+iffilename\.lower\(\)inproblem\_lower:
\+score\+=10
\+
\+\#Checkifanyparentdirectoriesappearinproblemstatement
\+forpartinfilename\.split\(’/’\):
\+ifpart\.lower\(\)inproblem\_lower:
\+score\+=5
\+break
\+
\+\#Checkifit’satestfilerelatedtoexploredmodules
\+ifgit\_tempdir:
\+test\_files=\_get\_relevant\_test\_files\(explored\_files,set\(\)\)
\+ifany\(filenameinfforfintest\_files\):
\+score\+=8
\+
\+\#Keepblockswithsufficientrelevancescore
\+\#Thresholdof5ensuresweonlykeepclearlyrelevantchanges
\+ifscore\>=5:
\+relevant\_blocks\.append\(block\)
\+else:
\+\#Logfordebugging\(optional\)
\+print\(f"\[diff\_minimality\]Filteringout\{filename\}\(score:\{score\}\)"\)
\+
\+\#Reconstructthediffwithonlyrelevantblocks
\+result=’\\n\\n’\.join\(relevant\_blocks\)
\+
\+\#Ensuretrailingnewline
\+ifresultandnotresult\.endswith\(’\\n’\):
\+result\+=’\\n’
\+
\+returnresult
\+
\+
\+def\_ensure\_patch\_trailing\_newline\(patch\_str\):
\+"""Ensurepatchstringendswithnewline\."""
\+ifpatch\_strandnotpatch\_str\.endswith\(’\\n’\):
\+returnpatch\_str\+’\\n’
\+returnpatch\_str
diff\-\-gita/utils/trace\_utils\.pyb/utils/trace\_utils\.py
newfilemode100644
index0000000\.\.b030135
\-\-\-/dev/null
\+\+\+b/utils/trace\_utils\.py
@@\-0,0\+1,475@@
\+\#Utilityfunctionsfortracingdocstringcodepathsinautodoc\-relatedissues\.
\+
\+importos
\+importre
\+importsubprocess
\+frompathlibimportPath
\+fromtypingimportDict,List,Optional,Tuple
\+
\+
\+defrun\_grep\(pattern:str,directory:str,file\_pattern:str="\*\.py",ignore\_case:bool=True,extended:bool=False\)\-\>List\[str\]:
\+"""
\+Rungrepinadirectoryforapatterninfilesmatchingfile\_pattern\.
\+Returnsalistofmatchinglines\.
\+
\+Args:
\+pattern:Theregexpatterntosearchfor
\+directory:Thedirectorytosearchin
\+file\_pattern:Thefilepatterntomatch\(e\.g\.,"\*\.py"\)
\+ignore\_case:Whethertousecase\-insensitivematching
\+extended:Whethertouseextendedregexmode\(\-Eflag\)
\+"""
\+cmd=\["grep","\-r","\-\-include="\+file\_pattern\]
\+ifignore\_case:
\+cmd\.append\("\-i"\)
\+ifextended:
\+cmd\.append\("\-E"\)
\+cmd\.extend\(\["\-n","\-\-color=never",pattern,directory\]\)
\+try:
\+result=subprocess\.run\(
\+cmd,
\+capture\_output=True,
\+text=True,
\+timeout=30,
\+\)
\+ifresult\.returncodein\(0,1\):\#0=matchfound,1=nomatch
\+returnresult\.stdout\.strip\(\)\.split\("\\n"\)ifresult\.stdout\.strip\(\)else\[\]
\+return\[\]
\+exceptsubprocess\.TimeoutExpired:
\+return\[\]
\+exceptException:
\+return\[\]
\+
\+
\+defrun\_agrep\(pattern:str,directory:str,file\_pattern:str="\*\.py"\)\-\>List\[str\]:
\+"""
\+Runag\(thesilversearcher\)forfastergrep\-likesearching\.
\+Fallsbacktogrepifagisnotavailable\.
\+"""
\+cmd=\["ag","\-\-python","\-\-literal","\-\-color=never",pattern,directory\]
\+try:
\+result=subprocess\.run\(
\+cmd,
\+capture\_output=True,
\+text=True,
\+timeout=30,
\+\)
\+ifresult\.returncodein\(0,1\):
\+returnresult\.stdout\.strip\(\)\.split\("\\n"\)ifresult\.stdout\.strip\(\)else\[\]
\+return\[\]
\+except\(subprocess\.TimeoutExpired,FileNotFoundError\):
\+\#Fallbacktogrep
\+returnrun\_grep\(pattern,directory,file\_pattern\)
\+
\+
\+deffind\_test\_file\(git\_tempdir:str,test\_function\_name:str\)\-\>Optional\[str\]:
\+"""
\+Findthetestfilecontainingthegiventestfunctionname\.
\+Returnsthefilepathrelativetogit\_tempdir,orNoneifnotfound\.
\+"""
\+\#Usegrepwithextendedregextofindthefunctiondefinition
\+\#Pattern:deftest\_function\_name\(
\+results=run\_grep\(
\+rf’def\\s\+\{re\.escape\(test\_function\_name\)\}\\s\*\\\(’,
\+git\_tempdir,
\+file\_pattern="\*\.py",
\+extended=True
\+\)
\+ifresults:
\+forlineinresults:
\+if":"inlineandline\.strip\(\):
\+\#Extractfilepath\(formatisusually"filename:line\_number:content"\)
\+file\_path=line\.split\(":"\)\[0\]
\+returnfile\_path
\+returnNone
\+
\+
\+defextract\_do\_autodoc\_call\(test\_content:str\)\-\>Optional\[Dict\[str,str\]\]:
\+"""
\+Extractthedo\_autodoc\(app,’<type\>’,’<object\>’\)callfromtestcontent\.
\+Returnsadictwith’type’and’object’keys,orNoneifnotfound\.
\+"""
\+\#Matchpatternslike:do\_autodoc\(app,’class’,’SomeClass’\)
\+\#or:do\_autodoc\(app,"class","SomeClass"\)
\+pattern=r"do\_autodoc\\s\*\\\(\\s\*app\\s\*,\\s\*\[’\\"\]\(\[^’\\"\]\+\)\[’\\"\]\\s\*,\\s\*\[’\\"\]\(\[^’\\"\]\+\)\[’\\"\]\\s\*\\\)"
\+match=re\.search\(pattern,test\_content\)
\+ifmatch:
\+return\{
\+"type":match\.group\(1\),
\+"object":match\.group\(2\),
\+\}
\+returnNone
\+
\+
\+defextract\_expected\_output\(test\_content:str\)\-\>Optional\[str\]:
\+"""
\+Extracttheexpectedoutputfromanassertioninthetest\.
\+Looksforpatternslike:
\+\-assert’expectedtext’inresult
\+\-assertresult==’expectedtext’
\+\-assert"expectedtext"inresult
\+"""
\+\#Patternforassertresult==’expected’orassertresult=="expected"
\+pattern=r"assert\\s\+\.\*?==\\s\*\[’\\"\]\(\.\+?\)\[’\\"\]"
\+matches=re\.findall\(pattern,test\_content,re\.DOTALL\)
\+ifmatches:
\+returnmatches\[\-1\]\#Returnthelastmatch\(usuallythemostrelevant\)
\+
\+\#Patternforassert’expected’inresultorassert"expected"inresult
\+pattern=r"assert\\s\+\[’\\"\]\(\.\+?\)\[’\\"\]\\s\+in\\s\+\.\*?"
\+matches=re\.findall\(pattern,test\_content,re\.DOTALL\)
\+ifmatches:
\+returnmatches\[\-1\]
\+
\+returnNone
\+
\+
\+deffind\_documenter\_class\(git\_tempdir:str,obj\_type:str\)\-\>Optional\[str\]:
\+"""
\+Findthedocumenterclassthathandlesthegivenobjecttype\.
\+Forexample,’class’\-\>’DataDocumenter’,’module’\-\>’ModuleDocumenter’,etc\.
\+"""
\+\#Mapcommonautodoctypestodocumenternamingpatterns
\+type\_to\_pattern=\{
\+"class":r"\(?:^\|\\s\)\(DataDocumenter\|ClassDocumenter\)\\b",
\+"function":r"\(?:^\|\\s\)\(FunctionDocumenter\|MethodDocumenter\)\\b",
\+"method":r"\(?:^\|\\s\)\(MethodDocumenter\)\\b",
\+"attribute":r"\(?:^\|\\s\)\(AttributeDocumenter\|DataDocumenter\)\\b",
\+"module":r"\(?:^\|\\s\)\(ModuleDocumenter\)\\b",
\+"exception":r"\(?:^\|\\s\)\(ExceptionDocumenter\|DataDocumenter\)\\b",
\+"data":r"\(?:^\|\\s\)\(DataDocumenter\)\\b",
\+\}
\+
\+pattern=type\_to\_pattern\.get\(obj\_type\.lower\(\),r"\(?:^\|\\s\)\(\\w\+Documenter\)\\b"\)
\+
\+\#Searchfordocumenterclasseswithextendedregex
\+results=run\_grep\(pattern,git\_tempdir,file\_pattern="\*\.py",extended=True\)
\+forlineinresults:
\+match=re\.search\(pattern,line\)
\+ifmatch:
\+returnmatch\.group\(1\)
\+
\+\#Ifnospecificmatch,lookforDocumenterbaseclassreferences
\+results=run\_grep\(r"class\\s\+\\w\+Documenter\\s\*\\\(",git\_tempdir,file\_pattern="\*\.py",extended=True\)
\+forlineinresults:
\+if"Documenter"inline:
\+match=re\.search\(r"class\\s\+\(\\w\+Documenter\)",line\)
\+ifmatch:
\+returnmatch\.group\(1\)
\+
\+returnNone
\+
\+
\+deftrace\_get\_object\_doc\(git\_tempdir:str,documenter\_class:str\)\-\>Optional\[Dict\]:
\+"""
\+Tracetheget\_object\_docmethodinthegivendocumenterclass\.
\+Returnsinformationabouthowdocstringsareread\.
\+"""
\+\#Useextendedregexforbetterpatternmatching
\+results=run\_grep\(
\+r’def\\s\+get\_object\_doc\\s\*\\\(’,
\+git\_tempdir,
\+file\_pattern="\*\.py",
\+extended=True
\+\)
\+
\+forlineinresults:
\+ifdocumenter\_class\.lower\(\)inline\.lower\(\)or"Documenter"inline:
\+\#Extractfilepath
\+parts=line\.split\(":"\)
\+iflen\(parts\)\>=2:
\+file\_path=parts\[0\]
\+\#Readthefunctiontocheckfor\_\_doc\_\_access
\+try:
\+content=Path\(file\_path\)\.read\_text\(\)
\+\#Lookforself\.object\.\_\_doc\_\_orself\.object\.\_\_doc\_\_
\+if"self\.object\.\_\_doc\_\_"incontentor"self\.object\.\_\_doc\_\_"incontent:
\+return\{
\+"file":file\_path,
\+"method":"get\_object\_doc",
\+"doc\_source":"\_\_doc\_\_",
\+"uses\_self\_object\_doc":True,
\+\}
\+\#Checkfor\.\_\_doc\_\_attributeaccessmoregenerally
\+doc\_access\_pattern=r"self\\\.object\\s\*\\\.\\s\*\_\_doc\_\_"
\+ifre\.search\(doc\_access\_pattern,content\):
\+return\{
\+"file":file\_path,
\+"method":"get\_object\_doc",
\+"doc\_source":"\_\_doc\_\_",
\+"uses\_self\_object\_doc":True,
\+\}
\+exceptException:
\+pass
\+returnNone
\+
\+
\+defcheck\_module\_analyzer\(git\_tempdir:str\)\-\>dict:
\+"""
\+CheckifModuleAnalyzerisimportedandusedinthecodebase\.
\+Returnsadictwithimportstatusandusageinformation\.
\+"""
\+\#CheckforModuleAnalyzerimportsusingextendedregex
\+\#Use\\S\+insteadof\\w\+tohandledottedmodulepaths
\+import\_patterns=\[
\+r’from\\s\+\\S\+\\s\+import\\s\+\.\*ModuleAnalyzer’,
\+r’import\\s\+\\S\+\\\.ModuleAnalyzer’,
\+\]
\+
\+imported=False
\+import\_lines=\[\]
\+
\+forpatterninimport\_patterns:
\+results=run\_grep\(pattern,git\_tempdir,file\_pattern="\*\.py",extended=True\)
\+import\_lines\.extend\(results\)
\+ifresults:
\+imported=True
\+
\+\#Checkforget\_commentsusage\(ModuleAnalyzer’smethodforgetting\#:comments\)
\+get\_comments\_results=run\_grep\(
\+r’get\_comments’,
\+git\_tempdir,
\+file\_pattern="\*\.py",
\+extended=True
\+\)
\+uses\_get\_comments=len\(get\_comments\_results\)\>0
\+
\+\#Checkforcomment\-baseddocstringhandling
\+comment\_patterns=\[
\+r’\#:\\s\*\\w’,
\+r’comment\.\*docstring’,
\+r’docstring\.\*comment’,
\+\]
\+
\+has\_comment\_handling=False
\+forpatternincomment\_patterns:
\+ifrun\_grep\(pattern,git\_tempdir,file\_pattern="\*\.py",extended=True\):
\+has\_comment\_handling=True
\+break
\+
\+return\{
\+"is\_imported":imported,
\+"import\_lines":import\_lines\[:5\],\#Limittofirst5lines
\+"uses\_get\_comments":uses\_get\_comments,
\+"has\_comment\_handling":has\_comment\_handling,
\+\}
\+
\+
\+deffind\_fix\_location\(git\_tempdir:str,documenter\_class:str,doc\_info:dict\)\-\>str:
\+"""
\+Determinethelikelyfixlocationbasedontheanalysis\.
\+Returnsahuman\-readabledescriptionofwheretofix\.
\+"""
\+ifnotdoc\_info:
\+return"Couldnotdeterminefixlocationfromcodepathanalysis\."
\+
\+file\_path=doc\_info\.get\("file","unknown"\)
\+doc\_source=doc\_info\.get\("doc\_source","unknown"\)
\+
\+module\_analyzer\_info=check\_module\_analyzer\(git\_tempdir\)
\+
\+ifdoc\_source=="\_\_doc\_\_"andnotmodule\_analyzer\_info\.get\("is\_imported",False\):
\+fix\_suggestion=\(
\+f"Fix:Modify‘\{documenter\_class\}\.get\_object\_doc\(\)‘tofallbackto"
\+f"‘ModuleAnalyzer\.get\_comments\(\)‘when‘\_\_doc\_\_‘isempty\."
\+f"File:‘\{file\_path\}‘"
\+\)
\+elifdoc\_source=="\_\_doc\_\_"andmodule\_analyzer\_info\.get\("is\_imported",False\):
\+fix\_suggestion=\(
\+f"Fix:Modify‘\{documenter\_class\}\.get\_object\_doc\(\)‘in‘\{file\_path\}‘tocheck"
\+f"‘ModuleAnalyzer\.get\_comments\(\)‘asafallbackwhen‘self\.object\.\_\_doc\_\_‘isNone/empty\."
\+f"Thecurrentcodeonlyreadsfrom‘\_\_doc\_\_‘andignorescomment\-baseddocstrings\."
\+\)
\+else:
\+fix\_suggestion=\(
\+f"Examine‘\{documenter\_class\}\.get\_object\_doc\(\)‘in‘\{file\_path\}‘tounderstand"
\+f"howdocstringsarecurrentlyretrieved\.Consideraddingfallbacklogic\."
\+\)
\+
\+returnfix\_suggestion
\+
\+
\+deftrace\_docstring\_path\(
\+instance\_id:str,
\+test\_function\_name:str,
\+target\_object:str,
\+git\_tempdir:str,
\+\)\-\>str:
\+"""
\+Tracethedocstringcodepathforautodoc\-relatedissues\.
\+
\+Thisfunction:
\+1\.Parsesthetestfiletoextractthedo\_autodoc\(app,’<type\>’,’<object\>’\)call
\+2\.Findsthedocumenterclassthathandlesthegiventype
\+3\.Tracesthroughget\_object\_doc\(\)toidentifywhereself\.object\.\_\_doc\_\_isread
\+4\.ChecksifModuleAnalyzerisimportedandused
\+5\.Outputsasummaryofthecodepathwithfixsuggestions
\+
\+Args:
\+instance\_id:TheinstanceIDfortheissue
\+test\_function\_name:Thenameofthefailingtestfunction
\+target\_object:Thetargetobjectbeingdocumented
\+git\_tempdir:Pathtothegitrepositorydirectory
\+
\+Returns:
\+Aformattedsummarystringdescribingthecodepathandfixlocation\.
\+"""
\+\#Step1:Findandparsethetestfile
\+test\_file=find\_test\_file\(git\_tempdir,test\_function\_name\)
\+
\+ifnottest\_file:
\+return\(
\+f"\[TraceSummary\]Couldnotfindtestfilecontaining’\{test\_function\_name\}’"
\+f"in\{git\_tempdir\}\.Manualinvestigationrequired\."
\+\)
\+
\+\#Readthetestfilecontent
\+try:
\+test\_path=os\.path\.join\(git\_tempdir,test\_file\)
\+test\_content=Path\(test\_path\)\.read\_text\(\)
\+exceptExceptionase:
\+returnf"\[TraceSummary\]Couldnotreadtestfile\{test\_file\}:\{e\}"
\+
\+\#Step2:Extractdo\_autodoccallinformation
\+autodoc\_info=extract\_do\_autodoc\_call\(test\_content\)
\+
\+ifautodoc\_info:
\+obj\_type=autodoc\_info\["type"\]
\+obj\_name=autodoc\_info\["object"\]
\+else:
\+\#Fallback:usetarget\_objectifautodoccallnotfound
\+obj\_type="class"\#Defaultassumption
\+obj\_name=target\_object
\+
\+expected\_output=extract\_expected\_output\(test\_content\)
\+
\+\#Step3:Findthedocumenterclass
\+documenter\_class=find\_documenter\_class\(git\_tempdir,obj\_type\)
\+
\+ifnotdocumenter\_class:
\+return\(
\+f"\[TraceSummary\]Couldnotidentifythedocumenterclasshandling’\{obj\_type\}’type\."
\+f"Testfunction:\{test\_function\_name\},Target:\{target\_object\}"
\+\)
\+
\+\#Step4:Traceget\_object\_doc
\+doc\_info=trace\_get\_object\_doc\(git\_tempdir,documenter\_class\)
\+
\+\#Step5:CheckModuleAnalyzer
\+module\_analyzer\_info=check\_module\_analyzer\(git\_tempdir\)
\+
\+\#Step6:Findfixlocation
\+fix\_suggestion=find\_fix\_location\(git\_tempdir,documenter\_class,doc\_info\)
\+
\+\#Buildthesummary
\+summary\_lines=\[
\+"="\*60,
\+"DOCSTRINGCODEPATHTRACESUMMARY",
\+"="\*60,
\+"",
\+f"Instance:\{instance\_id\}",
\+f"TestFunction:\{test\_function\_name\}",
\+f"TargetObject:\{target\_object\}",
\+f"TestFile:\{test\_file\}",
\+"",
\+"\-\-\-AnalysisDetails\-\-\-",
\+"",
\+f"1\.autodoc\(\)calldetected:do\_autodoc\(app,’\{obj\_type\}’,’\{obj\_name\}’\)",
\+\]
\+
\+ifexpected\_output:
\+summary\_lines\.append\(f"Expectedoutputcontains:’\{expected\_output\[:100\]\}\.\.\.’"\)
\+
\+summary\_lines\.extend\(\[
\+"",
\+f"2\.Documenterclass:\{documenter\_class\}\(handles’\{obj\_type\}’objects\)",
\+\]\)
\+
\+ifdoc\_info:
\+summary\_lines\.append\(
\+f"3\.Docstringsource:\{documenter\_class\}\.get\_object\_doc\(\)"
\+f"in\{doc\_info\.get\(’file’,’unknown’\)\}readsfrom\{doc\_info\.get\(’doc\_source’,’unknown’\)\}"
\+\)
\+
\+summary\_lines\.extend\(\[
\+"",
\+"4\.ModuleAnalyzerstatus:",
\+f"\-Imported:\{’Yes’ifmodule\_analyzer\_info\[’is\_imported’\]else’No’\}",
\+f"\-Usesget\_comments\(\):\{’Yes’ifmodule\_analyzer\_info\[’uses\_get\_comments’\]else’No’\}",
\+f"\-Hascomment\-basedhandling:\{’Yes’ifmodule\_analyzer\_info\[’has\_comment\_handling’\]else’No’\}",
\+"",
\+"\-\-\-RecommendedFix\-\-\-",
\+"",
\+fix\_suggestion,
\+"",
\+"="\*60,
\+\]\)
\+
\+return"\\n"\.join\(summary\_lines\)
\+
\+
\+defextract\_test\_function\_names\(text:str\)\-\>List\[str\]:
\+"""Extractpytest\-styletestfunctionnamesmentionedinfreetext\."""
\+ifnottext:
\+return\[\]
\+names=re\.findall\(r"\\b\(test\_\[A\-Za\-z0\-9\_\]\+\)\\b",text\)
\+seen=set\(\)
\+out=\[\]
\+fornameinnames:
\+ifnamenotinseen:
\+seen\.add\(name\)
\+out\.append\(name\)
\+returnout
\+
\+
\+defgenerate\_trace\_instruction\(instance\_id:str,test\_function\_name:str,target\_object:str,git\_tempdir:str\)\-\>str:
\+"""
\+Generateaformattedinstructionfortheagenttotracethedocstringpath\.
\+Thisisawrapperthatformatstheoutputforconsumptionbytheagent\.
\+"""
\+trace\_result=trace\_docstring\_path\(instance\_id,test\_function\_name,target\_object,git\_tempdir\)
\+
\+instruction=f"""
\+<code\_path\_trace\>
\+\{trace\_result\}
\+</code\_path\_trace\>
\+
\+Pleaseexaminethecodepathidentifiedabove\.Thetraceshowshowdocstringsarebeingreadandwherethefixshouldbeapplied\.Focusontherecommendedfixlocationandensureyourchangesaddresstherootcause\.
\+"""
\+returninstruction
\+
\+
\+deftest\_trace\_docstring\_path\(\):
\+"""
\+Testfunctiontoverifytrace\_docstring\_pathworkscorrectly\.
\+Thisisaself\-testfunctionthatdoesn’trequirearealrepository\.
\+"""
\+\#Thisisabasicsanitycheck\-actualtestingrequiresarealrepository
\+try:
\+result=trace\_docstring\_path\(
\+instance\_id="test\_instance",
\+test\_function\_name="test\_placeholder",
\+target\_object="TestClass",
\+git\_tempdir="/tmp"
\+\)
\+assert"TraceSummary"inresultor"Couldnot"inresult
\+returnTrue
\+exceptExceptionase:
\+print\(f"Testerror:\{e\}"\)
\+returnFalse
\+
\+
\+if\_\_name\_\_=="\_\_main\_\_":
\+importsys
\+iflen\(sys\.argv\)\>=4:
\+instance\_id=sys\.argv\[1\]
\+test\_function\_name=sys\.argv\[2\]
\+target\_object=sys\.argv\[3\]
\+git\_tempdir=sys\.argv\[4\]iflen\(sys\.argv\)\>4else"/tmp"
\+
\+result=trace\_docstring\_path\(instance\_id,test\_function\_name,target\_object,git\_tempdir\)
\+print\(result\)
\+else:
\+\#Runself\-test
\+success=test\_trace\_docstring\_path\(\)
\+print\(f"Self\-test\{’passed’ifsuccesselse’failed’\}"\)Similar Articles
Self-Evolving Coding Agents
This paper surveys self-evolving coding agents, which improve their future behavior by updating frameworks, memory, skills, tools, or models from prior coding interactions, and presents a taxonomy of what evolves, when, and what software-specific evidence drives it.
The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators
This paper introduces the Red Queen Gödel Machine (RQGM), an evolutionary framework for recursive self-improvement under non-stationary utilities, where agents and evaluators co-evolve, improving performance on coding tasks, scientific writing, and Olympiad-level proof grading.
@niclane7: Just in time for ICML week, we are sharing our take on a key question for recursive self-improving AI. How can AI keep …
The Red Queen Gödel Machine enables recursive self-improvement in AI by co-evolving the agent and evaluator, achieving better coding performance with fewer tokens.
The first experimental evidence of recursive self-improvement (3 minute read)
Researchers present AIDE², a system with recursive auto-research loops that improved its own code over 100 iterations, discovering seven improvements and beating a hand-tuned agent on held-out benchmarks.
@AlphaSignalAI: https://x.com/AlphaSignalAI/status/2054201045346287766
The article discusses new research from Sakana AI and Meta on self-improving AI agents, specifically the Darwin-Gödel Machine and Hyperagents, which autonomously rewrite their own code and infrastructure to enhance performance without human intervention.