ELMER: Evolutionary Language Model that Explores and Refines

arXiv cs.LG Papers

Summary

Introduces ELMER, an evolutionary language model that searches over natural-language policy descriptions and compiles them into executable programs, using fine-tuned Qwen3-8B with Direct Preference Optimization to control mutation strength and improve search efficiency.

arXiv:2608.10196v1 Announce Type: new Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a larger rewrite can preserve the same execution trace. We introduce an Evolutionary Language Model that searches over natural-language policy descriptions and compiles typed programs for execution. A fully fine-tuned Qwen3-8B model learns three task-conditioned operations: conditional semantic mutation, natural language to domain-specific language (GPTL) compilation, and GPTL to natural language translation. The model is fine-tuned with conditional input on the mutation strength (low, medium, high) using Direct Preference Optimization (oDPO). Across 252 fixed-budget evolutionary searches, oDPO improves both behavioral calibration and finite-budget search efficiency. Natural-language attains the highest observed held-out fitness. Our analysis shows that the condition input (mutation strength) systematically changes semantic edit composition and that language mutations preserve more parent fitness at matched small-to-moderate behavioral displacement. These results show that language can serve as a steerable, execution-grounded search representation over executable program space.
Original Article
View Cached Full Text

Cached at: 08/12/26, 08:28 AM

# ELMER: Evolutionary Language Model that Explores and Refines
Source: [https://arxiv.org/html/2608.10196](https://arxiv.org/html/2608.10196)
Matthew Siper Nof1, New York Universitysiper\.matthew@gmail\.comAhmed Khalifa Nof1, University of Maltaahmed\.khalifa@um\.edu\.mtJulian Togelius Nof1, New York Universityjulian\.togelius@gmail\.com

###### Abstract

Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space\. Syntactic edit size is an unreliable proxy: a small code change can alter nearly every action, while a larger rewrite can preserve the same execution trace\. We introduce an Evolutionary Language Model that searches over natural\-language policy descriptions and compiles typed programs for execution\. A fully fine\-tuned Qwen3\-8B model learns three task\-conditioned operations: conditional semantic mutation, natural language to domain\-specific language \(GPTL\) compilation, and GPTL to natural language translation\. The model is fine\-tuned with conditional input on the mutation strength \(low, medium, high\) using Direct Preference Optimization \(oDPO\)\. Across 252 fixed\-budget evolutionary searches, oDPO improves both behavioral calibration and finite\-budget search efficiency\. Natural\-language attains the highest observed held\-out fitness\. Our analysis shows that the condition input \(mutation strength\) systematically changes semantic edit composition and that language mutations preserve more parent fitness at matched small\-to\-moderate behavioral displacement\. These results show that language can serve as a steerable, execution\-grounded search representation over executable program space\.

## 1Introduction

Program evolution can measure whether a mutation improves fitness, but it often cannot control how far that mutation moves in behavior space\. Continuous optimizers expose an explicit update scale in that program search typically relies on syntactic proxies such as token, tree, or subtree edits\. These proxies can be poorly aligned with execution in that a small comparator change may alter nearly every action a policy takes, while a larger rewrite may preserve the same action sequence\. The missing object is a mutation scale whose requested magnitude predicts realized behavior\.

We ask whether a language model can learn this scale from execution\. Our system separates the representation used for variation from the representation used for evaluation\. Evolution mutates standardized natural\-language \(NL\) descriptions of trading policies, while a learned compiler maps each child into typed Genetic Programming Trading Language \(GPTL\) code for deterministic execution\. Parent and child programs are run on the same historical market data, and disagreement between their action sequences defines realized behavioral displacement\. A low, medium, or high request targets an ordered, overlapping behavioral regime\.

A single fine\-tuned Qwen3\-8B large language model \(LLM\) learns three task\-conditioned operations: NL\-to\-GPTL compilation, GPTL\-to\-NL translation, and conditional NL mutation\. We construct mutation supervision from grammar\-valid GPTL degradation chains whose endpoints can be executed and assigned deterministic behavior distances, then translate those transitions into language\. Behavior\-grounded supervised fine\-tuning \(SFT\) learns partial regime control\. Common\-parent offset Direct Preference Optimization \(oDPO\) sharpens that control by preferring sibling children whose executed displacement better matches the requested regime\. While fitness remains the outer\-loop selection signal, behavioral displacement trains the variation operator\.

We use financial trading as a testbed for stochastic, state\-dependent programmatic control\. Fixed historical trajectories permit deterministic, highly parallel replay and strict temporal holdouts\. Across 252 fixed\-budget searches, oDPO significantly improves mutation\-strength calibration and validation\-trajectory area under the curve \(AUC\) relative to behavior\-grounded SFT, native abstract syntax tree \(AST\) mutation, and a matched code\-output oDPO model\. The matched model holds the backbone, training transitions, preference data, conditioning, and objective fixed while changing the mutation representation from language to GPTL code\. Natural\-language oDPO also discovers the highest observed held\-out policy, while held\-out mean and upper\-tail results are treated descriptively\. Mechanistic analyses show that requested strength changes semantic edit composition and that language mutations preserve more parent fitness at matched small\-to\-moderate behavioral displacement\.

Our contributions are:

- •an execution\-grounded definition of mutation displacement and a conditional variation operator over ordered behavioral regimes;
- •a multitask language model trained for semantic mutation, compilation, and translation using behaviorally calibrated executable transitions;
- •controlled ablations that isolate domain fine\-tuning, behavior\-grounded conditioning, oDPO, and the natural\-language representation; and
- •mechanistic analyses of calibration, semantic edit composition, syntax–behavior locality, and mutation quality at matched displacement\.

The rest of the paper is organized as follows\. We first review related work on program evolution, LLM\-guided search, preference optimization\. We then introduce the behavior\-space mutation operator, describe the multitask training data and oDPO objective, and define the mutation systems used in evolution\. Next, we present the experimental design and report the calibration, search\-efficiency, representation, and mechanistic results\. Finally, we discuss the implications and limitations of the approach before concluding\.

## 2Related Work

### Program Evolution, Semantics, and Locality

Genetic programming \(GP\) searches executable structures, so representation and variation jointly determine the neighborhood exposed to selection\(Koza[1992](https://arxiv.org/html/2608.10196#bib.bib1); Rothlauf[2006](https://arxiv.org/html/2608.10196#bib.bib2)\)\. Syntactically small edits can produce large output changes, motivating semantic operators\(Moraglioet al\.[2012](https://arxiv.org/html/2608.10196#bib.bib14)\)\. MAP\-Elites preserves elites across behavioral niches\(Mouret and Clune[2015](https://arxiv.org/html/2608.10196#bib.bib3)\)\. Transformer Semantic GP learns semantically related proposals\(Antheset al\.[2025](https://arxiv.org/html/2608.10196#bib.bib15)\)\. Continuous Program Search learns a behavior\-aware continuous space for typed trading programs\(Siperet al\.[2026b](https://arxiv.org/html/2608.10196#bib.bib12)\)\. We instead learn categorical control directly in the proposal distribution from executed parent–child displacement\.

### LLM\-Guided Evolution and Automated Discovery

Evolution through LLMs established language models as learned variation operators\(Lehmanet al\.[2023](https://arxiv.org/html/2608.10196#bib.bib16)\)\. Evaluator\-guided systems have been employed to evolve functions, model code, heuristics, metaheuristics, and scientific programs\(Romera\-Paredeset al\.[2024](https://arxiv.org/html/2608.10196#bib.bib4); Morriset al\.[2024](https://arxiv.org/html/2608.10196#bib.bib11); Liuet al\.[2024](https://arxiv.org/html/2608.10196#bib.bib17); van Stein and Bäck[2025](https://arxiv.org/html/2608.10196#bib.bib18); Novikovet al\.[2025](https://arxiv.org/html/2608.10196#bib.bib19); Adaption Research Staff[2026](https://arxiv.org/html/2608.10196#bib.bib13)\)\. ProFiT evolves executable trading programs from historical market feedback\(Siperet al\.[2026a](https://arxiv.org/html/2608.10196#bib.bib8)\)\. These systems evaluate proposals through fitness or execution\. We additionally train how far the proposal operator moves in behavior space\.

### Preference Optimization for Behavior\-Grounding

Scalar fitness and validity describehow wella program performs, but they fail to capturehow fara mutation has moved from its parent in behavior space\. To resolve this ambiguity, we adapt the reverse\-degradation principle from Path of Destruction\(Siperet al\.[2022](https://arxiv.org/html/2608.10196#bib.bib7)\)\. By reversing synthetic degradation chains, we construct a diverse dataset of directionally improving program transitions across multiple discrete behavioral displacements\. We then optimize our mutation operator using a margin\-aware preference framework\. While standard Direct Preference Optimization \(DPO\) optimizes for a binary quality gradient\(Rafailovet al\.[2023](https://arxiv.org/html/2608.10196#bib.bib5)\), oDPO incorporates pair\-dependent margins to scale the preference update\(Aminiet al\.[2024](https://arxiv.org/html/2608.10196#bib.bib6)\)\. By computing these margins directly from the observed trajectory displacement of common\-parent sibling pairs, we repurpose preference optimization\. Rather than simply ranking output quality, the model learns a steerable, ordinal geometry of the policy mutation space\.

## 3A Behavior\-Space Mutation Operator

### Behavior\-Space Step Size

Let𝒢\\mathcal\{G\}be the set of valid executable programs,ℛNL\\mathcal\{R\}\_\{\\mathrm\{NL\}\}the set of language descriptions,𝒮\\mathcal\{S\}the market\-state space, andℬ=\{0,1\}4\\mathcal\{B\}=\\\{0,1\\\}^\{4\}the raw signal space for long entry, short entry, long exit, and short exit\. A policyπ:𝒮→ℬ\\pi:\\mathcal\{S\}\\rightarrow\\mathcal\{B\}maps a market state to a four\-signal vector\. For a shared evaluation trajectory𝐬1:N=\(s1,…,sN\)∈𝒮N\\mathbf\{s\}\_\{1:N\}=\(s\_\{1\},\\ldots,s\_\{N\}\)\\in\\mathcal\{S\}^\{N\}, behavioral displacement is equivalent to the raw\-signal disagreement:

dbeh​\(πa,πb\)=1N​∑t=1N𝕀​\[πa​\(st\)≠πb​\(st\)\]\.d\_\{\\mathrm\{beh\}\}\(\\pi\_\{a\},\\pi\_\{b\}\)=\\frac\{1\}\{N\}\\sum\_\{t=1\}^\{N\}\\mathbb\{I\}\\\!\\left\[\\pi\_\{a\}\(s\_\{t\}\)\\neq\\pi\_\{b\}\(s\_\{t\}\)\\right\]\.\(1\)A learned compilerC:ℛNL→𝒢C:\\mathcal\{R\}\_\{\\mathrm\{NL\}\}\\rightarrow\\mathcal\{G\}maps child descriptionxcx\_\{c\}to a program inducingπc\\pi\_\{c\}\. Samplingxc∼pθ\(⋅∣xp,s\)x\_\{c\}\\sim p\_\{\\theta\}\(\\cdot\\mid x\_\{p\},s\)and compiling inducesqθ​\(πc∣πp,s\)q\_\{\\theta\}\(\\pi\_\{c\}\\mid\\pi\_\{p\},s\)fors∈\{low,medium,high\}s\\in\\\{\\mathrm\{low\},\\mathrm\{medium\},\\mathrm\{high\}\\\}\. Letμs\\mu\_\{s\}be the empirical center of regimess\. An idealized target is

minθ⁡𝔼πp,s​𝔼πc∼qθ\(⋅∣πp,s\)​\[\|dbeh​\(πp,πc\)−μs\|\]\.\\min\_\{\\theta\}\\;\\mathbb\{E\}\_\{\\pi\_\{p\},s\}\\,\\mathbb\{E\}\_\{\\pi\_\{c\}\\sim q\_\{\\theta\}\(\\cdot\\mid\\pi\_\{p\},s\)\}\\left\[\\left\|d\_\{\\mathrm\{beh\}\}\(\\pi\_\{p\},\\pi\_\{c\}\)\-\\mu\_\{s\}\\right\|\\right\]\.\(2\)Because execution is nondifferentiable, common\-parent oDPO instead ranks children by proximity toμs\\mu\_\{s\}, subject to

Dlow≺Dmedium≺Dhigh\.D\_\{\\mathrm\{low\}\}\\prec D\_\{\\mathrm\{medium\}\}\\prec D\_\{\\mathrm\{high\}\}\.\(3\)These categories are ordered, overlapping regimes rather than exact numerical radii\. Each representation\-operator pair induces its own practical reachability distribution\.

![Refer to caption](https://arxiv.org/html/2608.10196v1/x1.png)Figure 1:Learning a behavior\-space mutation operator\. A parent policy and requested strength produce a language mutation that is compiled and executed\. Displacement trains the variation operator, while validation fitness drives selection\. The lower panels show common\-parent oDPO preferences and the experimental decomposition\.
### Behavior\-Grounded Multitask Mutation Model

We fine\-tune a single Qwen3\-8B checkpoint\(Yanget al\.[2025](https://arxiv.org/html/2608.10196#bib.bib10)\)on three task\-conditioned operations\.Compilationmaps a natural\-language policy description to a typed GPTL program that can be checked, executed, and evaluated through backtesting\. Program\-to\-languagetranslationmaps a GPTL program into the standardized natural\-language representation consumed by the semantic mutator\.Mutationreceives a parent description, its backtest feedback, and a requested strengths∈\{low,medium,high\}s\\in\\\{\\mathrm\{low\},\\mathrm\{medium\},\\mathrm\{high\}\\\}, then produces a child description whose executed behavioral displacement should fall within the corresponding regime\. Together with deterministic backtesting and fitness\-based selection, these operations form a closed evolutionary loop\. Policies are mutated in language, compiled into GPTL code, executed to obtain fitness, and retained or rejected by the outer optimization algorithm\. Translation allows a hand\-designed GPTL seed, a directly mutated program, or another code\-space discovery to enter the language representation used by the mutator\. During ordinary natural\-language evolution, the child description is inherited directly and its compiled GPTL program is stored as the evaluated artifact, so translation is not repeatedly applied along the lineage\.

The same loop provides a controlled way to construct behavior\-grounded mutation supervision\. Starting from high\-fitness GPTL programs, defined as training Sharpe greater than 0\.5, we apply grammar\-valid AST edits and retain edges that directionally reduce fitness, producing degradation chains\. Reversing one\-, two\-, and four\-link segments yields directionally improving parent\-child transitions at several mutation scales\. Because both endpoints are executable programs, we run them on the same training\-only market windows and compute their behavioral displacement using Eq\.[1](https://arxiv.org/html/2608.10196#S3.E1)\. GLM 5\.2, employed as an annotation model with reasoning off and temperature equal to 0\.1, then renders both endpoints in the standardized natural\-language format\. The resulting examples pair a natural\-language semantic edit with a displacement measured deterministically from the original executable programs\. The model therefore learns to mutate in language while developing an implicit association between semantic edit patterns and their realized behavioral regimes\.

The final supervised corpus contains 45,639 examples: 13,551 conditional mutations, 16,044 compilations, and 16,044 program\-to\-language translations\. These collections are split into 44,727 training and 912 validation examples\. The mutation data are derived from 2,811 degradation chains and 44,920 graded reverse transitions, of which 44,917 execute successfully\. Strength labels are assigned using equal\-frequency behavioral\-distance bins with cut points 0\.165 and 0\.466 as determined by the data distribution\. We remove null transitions below 0\.005 and transitions near the category boundaries, then balance the retained examples across strength and asset\. Distances are measured on six training\-only windows of 6,000 hourly bars\. Mutation prompts provide the available domain primitives, the parent policy \(both its GPTL and NL representations\), its fitness, and seven backtest statistics\. Conditioned prompts add the requested mutation strength\.

### Common\-Parent Preferences and oDPO

For each eligible parent, we select one representative child per strength category\. Given requestss, we designate as preferredy\+y^\{\+\}the child nearestμs\\mu\_\{s\}and as rejectedy−y^\{\-\}a child outside the requested category\. We filter out pairs whose behavioral separation falls below0\.150\.15times the training interquartile range \(IQR\) to ensure a distinct preference gap\. Each retained pair receives

m=min⁡\(1,\|d\+−d−\|Q0\.95​\(\|d\+−d−\|\)\)\.m=\\min\\\!\\left\(1,\\frac\{\|d^\{\+\}\-d^\{\-\}\|\}\{Q\_\{0\.95\}\(\|d^\{\+\}\-d^\{\-\}\|\)\}\\right\)\.\(4\)HereQ0\.95Q\_\{0\.95\}, the 95th percentile of training separations, robustly scales typical pairs\. Clipping at one prevents extreme outliers, from dominating the oDPO offset and gradient\. The procedure yields 14,735 training and 775 validation pairs from 6,369 parents\.

Letrθ​\(y∣x\)=log⁡πθ​\(y∣x\)−log⁡πref​\(y∣x\)r\_\{\\theta\}\(y\\mid x\)=\\log\\pi\_\{\\theta\}\(y\\mid x\)\-\\log\\pi\_\{\\mathrm\{ref\}\}\(y\\mid x\), whereπref\\pi\_\{\\mathrm\{ref\}\}is the frozen behavior\-grounded SFT model\. We optimize

ℒoDPO=−𝔼​\[log⁡σ​\(β​Δ​rθ−λ​m\)\],\\mathcal\{L\}\_\{\\mathrm\{oDPO\}\}=\-\\mathbb\{E\}\\\!\\left\[\\log\\sigma\\\!\\left\(\\beta\\Delta r\_\{\\theta\}\-\\lambda m\\right\)\\right\],\(5\)withΔ​rθ=rθ​\(y\+∣x\)−rθ​\(y−∣x\)\\Delta r\_\{\\theta\}=r\_\{\\theta\}\(y^\{\+\}\\mid x\)\-r\_\{\\theta\}\(y^\{\-\}\\mid x\),β=0\.3\\beta=0\.3, andλ=0\.2\\lambda=0\.2\. SFT uses one epoch at2×10−52\\times 10^\{\-5\}; oDPO uses one epoch at2×10−62\\times 10^\{\-6\}\. Checkpoints must exceed 95% on compilation validity, translation signal preservation, and output format, and pass a held\-out monotonicity gate with Spearmanρ≥0\.4\\rho\\geq 0\.4\.

### Mutation Systems and Evolution

The ablation ladder separates four effects\. Stock Qwen3\-8B measures zero\-shot mutation; no\-label SFT isolates domain and mutation\-task fine\-tuning; prompt\-only labels test whether strength words alone induce control; behavior\-grounded SFT tests supervised conditioning on measured displacement; and oDPO isolates the additional effect of common\-parent preference optimization after conditional SFT\. Two code controls isolate representation and operator effects\. The native baseline compares learned language mutation with a hand\-designed weighted AST operator spanning parameter, indicator, subtree, comparator, logical, clause, crossover, and signal\-leg edits\. This same operator system was used to generate the degradation chains that ultimately produced the multi\-task training dataset\. The matched code\-output model holds the backbone, training transitions, preference pairs, strength labels, and oDPO objective fixed while changing the output substrate from natural language to GPTL\. It therefore provides the cleanest test of natural language as a search representation\.

## 4Experimental Design

Experiments use hourly continuous futures from January 2008 through October 2025: the E\-mini S&P 500 contract \(ES; 106,685 bars\), Silver \(SI; 107,085\), and U\.S\. Treasury Bond \(US; 105,124\)\. The backtester starts with $10,000, applies 0\.005% commission and 0\.01% slippage, and uses unit long, short, or flat positions\. Each asset has three non\-overlapping temporal folds, split chronologically into 70% training, 15% validation, and 15% test, with a ten\-day embargo\. During search, the mutation model receives training\-fold feedback; validation scores drive evolutionary selection, and the best\-validation policy is evaluated once on the corresponding held\-out test fold after the 1,000\-evaluation budget ends\. The learned\-operator matrix contains six operators, three algorithms, three assets, and four seeds \(216 searches\)\. A separate 36\-search AST matrix supplies the native baseline, and the 36 NL oDPO runs are shared across analyses\. We report validation Sharpe, held\-out test Sharpe, top\-quartile mean test Sharpe, maximum test Sharpe, validity, and area under the best\-so\-far validation curve on the common 1,000\-evaluation grid \(AUC\)\.

Lineage analysis measures normalized AST token distance, Eq\.[1](https://arxiv.org/html/2608.10196#S3.E1), edit taxonomy, entropy, and fitness change in shared behavioral\-distance bins\. A mutation is neutral whendbeh=0d\_\{\\mathrm\{beh\}\}=0and catastrophic when an execution\-valid child’s validation Sharpe falls by more than one; failures are tracked separately through validity\. Runs are paired by algorithm, asset, and seed, yielding 36 matched observations per contrast\. Six confirmatory contrasts use two\-sided Wilcoxon signed\-rank tests with Holm correction, 10,000\-resample paired\-bootstrap intervals, and rank\-biserial effects\. Held\-out test fitness is the endpoint for the three capability and conditioning contrasts; AUC is the endpoint for the three oDPO and representation contrasts\. Representation\-level mean, top\-quartile, and maximum test fitness are descriptive\. Calibration uses run\-level Spearman correlations followed by paired Wilcoxon tests with Holm correction; pooled low–high Cliff’sδ\\deltais descriptive\.

## 5Results

### A Requested Strength Becomes a Behavioral Control

Figure[2](https://arxiv.org/html/2608.10196#S5.F2)shows the pooled ordinal response\. In the run\-level analysis used for inference, mean Spearman correlation between requested strength and realized displacement is0\.8240\.824for oDPO,0\.3100\.310for behavior\-grounded SFT, and0\.0180\.018for prompt\-only conditioning\. Every one of the 36 paired runs favors oDPO in both comparisons \(rrb=1\.00r\_\{\\mathrm\{rb\}\}=1\.00,pHolm<0\.0001p\_\{\\mathrm\{Holm\}\}<0\.0001\)\. Descriptive pooled low–high Cliff’s deltas are0\.980\.98,0\.360\.36, and0\.020\.02\. oDPO therefore turns partial supervised ordering into consistent regime control; the broad medium distribution and high\-strength saturation still preclude exact numerical calibration\.

![Refer to caption](https://arxiv.org/html/2608.10196v1/x2.png)Figure 2:Requested strength versus realized action\-sequence displacement\. Behavior\-grounded SFT learns partial ordering, oDPO sharpens it, and prompt\-only labels produce no measurable response\. Bars mark medians\.
### Where the Mutation Capability Comes From

Table[1](https://arxiv.org/html/2608.10196#S5.T1)and Figure[3](https://arxiv.org/html/2608.10196#S5.F3)A decompose the operator\. The untrained model yields negative validation, test, and AUC\. No\-label SFT becomes functional\. Prompt\-only labels lower every average metric relative to no\-label SFT, showing that grounding gives the conditioning interface its operational meaning\.

A\. Descriptive aggregates \(36 searches per system\)Mutation systemMeanValidationMeanTestTop\-QuartileTestMaximumTestMeanAUC*Natural\-language training ablation*Untrained base\-0\.1844\-0\.21630\.27450\.4027\-109\.0No\-label SFT0\.91220\.15830\.80211\.1715491\.1Label\-only checkpoint0\.7836\-0\.05750\.50591\.3383436\.9Behavior\-grounded SFT1\.12320\.48051\.53861\.9841527\.9Behavior\-grounded NL oDPO1\.43640\.44781\.50632\.3445751\.2*Representation controls*Direct code \(AST\)1\.19490\.35260\.99591\.2598569\.9Code\-output oDPO1\.29380\.30241\.19671\.6319693\.0B\. Confirmatory paired contrasts \(36 matched runs\)ContrastEndpointΔ\\Delta95% CIrrbr\_\{\\mathrm\{rb\}\}pHolmp\_\{\\mathrm\{Holm\}\}No\-label SFT−\-BaseTest\+0\.375\+0\.375\[0\.189,0\.559\]\[0\.189,\\ 0\.559\]0\.640\.640\.0024Label\-only−\-No\-labelTest−0\.216\-0\.216\[−0\.417,0\.003\]\[\-0\.417,\\ 0\.003\]−0\.51\-0\.510\.0251BG\-SFT−\-Label\-onlyTest\+0\.538\+0\.538\[0\.220,0\.869\]\[0\.220,\\ 0\.869\]0\.560\.560\.0111NL oDPO−\-BG\-SFTAUC\+223\.3\+223\.3\[134\.8,320\.0\]\[134\.8,\\ 320\.0\]0\.820\.820\.0001NL oDPO−\-Direct ASTAUC\+181\.3\+181\.3\[85\.0,295\.1\]\[85\.0,\\ 295\.1\]0\.580\.580\.0078NL oDPO−\-Code oDPOAUC\+58\.2\+58\.2\[2\.3,123\.9\]\[2\.3,\\ 123\.9\]0\.390\.390\.0425Table 1:Training, representation, and confirmatory results\. Panel A reports descriptive aggregates\. Panel B reports paired mean differences and bootstrap intervals;pHolmp\_\{\\mathrm\{Holm\}\}comes from two\-sided Wilcoxon signed\-rank tests with Holm correction across the six contrasts\. Test denotes held\-out fitness and AUC denotes evaluation\-index search\-curve area\. Representation\-level held\-out summaries and maxima remain descriptive\.Panel B shows that domain SFT creates a functional mutator, prompt\-only labels significantly degrade performance, and behavior\-grounded SFT restores a useful conditional interface\. All three oDPO AUC contrasts remain significant after Holm correction, establishing its primary gain in finite\-budget search efficiency\. Panel A provides descriptive held\-out context: behavior\-grounded SFT has the highest mean and top\-quartile test fitness, while NL oDPO attains the highest observed test fitness \(2\.34452\.3445\)\. The label\-only bootstrap interval targets the paired mean, whereas its Wilcoxon test targets a paired rank shift, so the interval can narrowly cross zero while the corrected rank test remains significant\.

![Refer to caption](https://arxiv.org/html/2608.10196v1/x3.png)Figure 3:Best\-so\-far validation fitness\. Panel A decomposes training\. Panel B compares NL with native AST mutation and a matched code\-output oDPO model\. Curves show mean±\\pmstandard error \(SE\) over 36 runs\.![Refer to caption](https://arxiv.org/html/2608.10196v1/x4.png)Figure 4:Held\-out test fitness across representation–operator systems\. Points are searches; marker shape denotes the outer algorithm\. Lines show mean, top\-quartile mean, and maximum\.![Refer to caption](https://arxiv.org/html/2608.10196v1/x5.png)Figure 5:Realized edit composition by requested strength\. The behavior\-grounded operator changes edit types and entropy, while the prompt\-only checkpoint remains nearly invariant\.![Refer to caption](https://arxiv.org/html/2608.10196v1/x6.png)Figure 6:Why language is a productive mutation substrate\. Panels A–B relate AST distance to executed behavior for NL and native code mutations\. Panels C–D compare fitness change at matched displacement fordbeh≤0\.25d\_\{\\mathrm\{beh\}\}\\leq 0\.25\.
### Language Improves Finite\-Budget Search Efficiency

Figure[3](https://arxiv.org/html/2608.10196#S5.F3)B and Panel B of Table[1](https://arxiv.org/html/2608.10196#S5.T1)compare NL oDPO with native AST mutation and the matched code\-output control\. Both AUC contrasts remain significant after Holm correction\. Because the matched model holds the backbone, transitions, preferences, labels, and objective fixed, its contrast isolates the representation used to express variation\. The code\-output effect is modest and near the corrected threshold, but the two comparisons support a finite\-budget search\-efficiency advantage for language\. Figure[4](https://arxiv.org/html/2608.10196#S5.F4)provides descriptive held\-out context\. NL oDPO has higher mean test fitness than native AST \(0\.44780\.4478versus0\.35260\.3526\) and matched code output \(0\.44780\.4478versus0\.30240\.3024\), as well as the highest top\-quartile mean \(1\.50631\.5063\) and observed maximum \(2\.34452\.3445\)\. These values indicate a stronger observed upper tail but are not confirmatory representation endpoints\. All systems include negative outcomes, and the outer\-loop pattern is heterogeneous: MAP\-Elites and ProFiT favor NL on mean test, whereas\(μ,λ\)\(\\mu,\\lambda\)\-ES favors matched code output\. Archive memory may help preserve NL’s structured exploration, but the experiment does not isolate this interaction\.

### Strength Selects Different Mutation Regimes

Figure[5](https://arxiv.org/html/2608.10196#S5.F5)shows that parameter changes constitute 86\.4% of low\-strength edits, 35\.5% at medium, and 0\.46% at high\. High strength reallocates probability toward indicator substitutions \(39\.5%\), comparator swaps \(19\.7%\), structural rewrites \(12\.4%\), directional flips \(11\.4%\), and clause operations\. Entropy rises from0\.820\.82to2\.232\.23to2\.402\.40bits, while prompt\-only remains near1\.861\.86bits\. Requested strength therefore selects distinct semantic mutation regimes rather than merely changing surface wording\.

### Why Language Produces More Useful Moves

Compiled NL mutations have a stronger syntax–behavior relationship \(ρ=0\.88\\rho=0\.88,r=0\.66r=0\.66\) than native AST mutation \(ρ=0\.44\\rho=0\.44,r=0\.31r=0\.31\) and a lower catastrophic rate \(9% versus 14%\)\. Catastrophic denotes an execution\-valid child losing more than one validation Sharpe; failures are counted separately\. Both mappings remain noisy\. Within the common\-support rangedbeh≤0\.25d\_\{\\mathrm\{beh\}\}\\leq 0\.25, NL preserves more parent fitness than native AST mutation and generally more than matched code output, supporting more coherent moves at comparable executed distance\.

## 6Discussion

Low, medium, and high mutation\-strengths acquire operational meaning through executed parent–child displacement\. Prompt\-only labels leave behavior nearly unchanged, while behavior\-grounded SFT establishes partial ordering, and oDPO makes that ordering consistent across matched runs\. The stages play complementary roles\. SFT creates a functional semantic mutator and yields the strongest descriptive mean and top\-quartile held\-out performance, while oDPO primarily improves calibration and finite\-budget search efficiency\. The strongest representation\-level evidence is validation\-trajectory AUC, where NL oDPO significantly exceeds native AST mutation and matched code\-output oDPO, while its higher held\-out mean, upper tail, and held\-out maximum remain descriptive\. Because the matched control fixes the backbone, training transitions, preference pairs, labels, and objective, this contrast most directly tests the representation used for variation\. Mechanistic results indicate that language produced more behaviorally ordered and less damaging small\-to\-moderate moves\. Variation across outer loops and uniform strength sampling motivate adaptive mutation scheduling as a future research focus\.

## 7Conclusion

Program evolution has long been able to reward a mutation after the fact but has been far less able to control how far that mutation moves before evaluation\. We address this gap by separating semantic variation from deterministic execution: a multitask language model edits policy intent in natural language, compiles each proposal into typed GPTL, and learns mutation strength from the action trajectories produced by execution\. Behavior\-grounded SFT creates the conditional mutation capability, while common\-parent oDPO turns low, medium, and high requests into reliable ordinal behavioral regimes and significantly improves finite\-budget search efficiency over behavior\-grounded SFT, native AST mutation, and a matched code\-output model\. The resulting operator also discovers the highest observed held\-out policy, while mechanistic analyses show that requested strength reshapes semantic edit composition and preserves more parent fitness at matched small\-to\-moderate behavioral displacement\. These results move program evolution beyond blind syntactic perturbation toward deliberate, execution\-grounded movement through program space with models that learn not only what to change, but how far to move\.

## References

- AutoScientist: automating the science of model training\.Note:https://adaptionlabs\.ai/blog/autoscientistAccessed: 2026\-07\-19Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- A\. Amini, T\. Vieira, and R\. Cotterell \(2024\)Direct preference optimization with an offset\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 9954–9972\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.592)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx3.p1.1)\.
- P\. Anthes, D\. Sobania, and F\. Rothlauf \(2025\)Transformer semantic genetic programming for symbolic regression\.InProceedings of the Genetic and Evolutionary Computation Conference,New York, NY, USA,pp\. 952–960\.External Links:[Document](https://dx.doi.org/10.1145/3712256.3726412)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx1.p1.1)\.
- J\. R\. Koza \(1992\)Genetic programming: on the programming of computers by means of natural selection\.MIT Press,Cambridge, MA\.Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx1.p1.1)\.
- J\. Lehman, J\. Gordon, S\. Jain, K\. Ndousse, C\. Yeh, and K\. O\. Stanley \(2023\)Evolution through large models\.InHandbook of Evolutionary Machine Learning,W\. Banzhaf, P\. Machado, and M\. Zhang \(Eds\.\),pp\. 331–366\.External Links:[Document](https://dx.doi.org/10.1007/978-981-99-3814-8%5F11)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- F\. Liu, X\. Tong, M\. Yuan, X\. Lin, F\. Luo, Z\. Wang, Z\. Lu, and Q\. Zhang \(2024\)Evolution of heuristics: towards efficient automatic algorithm design using large language model\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 32201–32223\.Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- A\. Moraglio, K\. Krawiec, and C\. G\. Johnson \(2012\)Geometric semantic genetic programming\.InParallel Problem Solving from Nature—PPSN XII,Lecture Notes in Computer Science, Vol\.7491,Berlin, Heidelberg,pp\. 21–31\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-32937-1%5F3)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx1.p1.1)\.
- C\. Morris, M\. Jurado, and J\. Zutty \(2024\)LLM guided evolution—the automation of models advancing models\.InProceedings of the Genetic and Evolutionary Computation Conference,New York, NY, USA,pp\. 377–384\.External Links:[Document](https://dx.doi.org/10.1145/3638529.3654178)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- J\. Mouret and J\. Clune \(2015\)Illuminating search spaces by mapping elites\.External Links:1504\.04909,[Document](https://dx.doi.org/10.48550/arXiv.1504.04909)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx1.p1.1)\.
- A\. Novikov, N\. Vũ, M\. Eisenberger, E\. Dupont, P\. Huang, A\. Z\. Wagner, S\. Shirobokov, B\. Kozlovskii, F\. J\. R\. Ruiz, A\. Mehrabian, M\. P\. Kumar, A\. See, S\. Chaudhuri, G\. Holland, A\. Davies, S\. Nowozin, P\. Kohli, and M\. Balog \(2025\)AlphaEvolve: a coding agent for scientific and algorithmic discovery\.External Links:2506\.13131,[Document](https://dx.doi.org/10.48550/arXiv.2506.13131)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 53728–53741\.External Links:[Document](https://dx.doi.org/10.52202/075280-2338)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx3.p1.1)\.
- B\. Romera\-Paredes, M\. Barekatain, A\. Novikov, M\. Balog, M\. P\. Kumar, E\. Dupont, F\. J\. R\. Ruiz, J\. S\. Ellenberg, P\. Wang, O\. Fawzi, P\. Kohli, and A\. Fawzi \(2024\)Mathematical discoveries from program search with large language models\.Nature625\(7995\),pp\. 468–475\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06924-6)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- F\. Rothlauf \(2006\)Representations for genetic and evolutionary algorithms\.2nd edition,Springer,Berlin, Heidelberg\.External Links:[Document](https://dx.doi.org/10.1007/3-540-32444-5)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx1.p1.1)\.
- M\. Siper, A\. Khalifa, L\. B\. Soros, M\. U\. Nasir, J\. Azhang, and J\. Togelius \(2026a\)ProFiT: program search for financial trading\.InProceedings of the Genetic and Evolutionary Computation Conference,New York, NY, USA,pp\. 727–734\.External Links:[Document](https://dx.doi.org/10.1145/3795095.3805175)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- M\. Siper, A\. Khalifa, and J\. Togelius \(2022\)Path of destruction: learning an iterative level generator using a small dataset\.In2022 IEEE Symposium Series on Computational Intelligence \(SSCI\),pp\. 337–343\.External Links:[Document](https://dx.doi.org/10.1109/SSCI51031.2022.10022073)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx3.p1.1)\.
- M\. Siper, M\. U\. Nasir, A\. Khalifa, L\. Soros, J\. Azhang, and J\. Togelius \(2026b\)Continuous program search\.External Links:2602\.07659,[Document](https://dx.doi.org/10.48550/arXiv.2602.07659)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx1.p1.1)\.
- N\. van Stein and T\. Bäck \(2025\)LLaMEA: a large language model evolutionary algorithm for automatically generating metaheuristics\.IEEE Transactions on Evolutionary Computation29\(2\),pp\. 331–345\.External Links:[Document](https://dx.doi.org/10.1109/TEVC.2024.3497793)Cited by:[§2](https://arxiv.org/html/2608.10196#S2.SSx2.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025\)Qwen3 technical report\.External Links:2505\.09388,[Document](https://dx.doi.org/10.48550/arXiv.2505.09388)Cited by:[§3](https://arxiv.org/html/2608.10196#S3.SSx2.p1.1)\.

Similar Articles

Evolution through large models

OpenAI Blog

This paper demonstrates that large language models trained on code can significantly enhance genetic programming mutation operators, enabling the generation of hundreds of thousands of functional Python programs for robot design in the Sodarace domain without prior training data. The approach, called Evolution through Large Models (ELM), combines LLMs with MAP-Elites to bootstrap new conditional models for context-specific artifact generation.

Rethinking Experience Utilization in Self-Evolving Language Model Agents

arXiv cs.CL

This paper introduces ExpWeaver, a framework that optimizes how self-evolving language model agents utilize past experiences during runtime decision-making. It demonstrates that selectively invoking experience based on reasoning uncertainty improves performance across various environments and models.