Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
Summary
This paper critiques single-point evaluations for LLM evolutionary search, showing that optimal performance depends on varying seeds and iterations, and proposes a budget-grid evaluation protocol.
View Cached Full Text
Cached at: 09/18/26, 09:01 AM
# Evolution or Illusion? Rethinking Evaluation in LLM Evolutionary Search
Source: [https://arxiv.org/html/2609.19799](https://arxiv.org/html/2609.19799)
Roi PonyOshri NaparstekUdi BarzelayAffiliation:IBM Research
###### Abstract
LLM\-driven evolutionary search finds programs by launching seeds and iterating each one\. Papers report a single budget setting, usually one seed run for a fixed number of iterations, and rank methods from that one point\. We show this is not enough\. We evaluate three evolutionary search strategies on five optimization tasks, commonly used by papers in the genre to report results\. We run the analysis over a full grid of seeds and iterations\. Our findings suggest that the best way to split a fixed budget between more seeds \(width\) and more iterations \(depth\) changes with the strategy, the task, and the total budget\. Furthermore, we observe that the ranking of strategies also changes with the budget\. On one task the strategy that looks worst at one seed is best at forty seeds\. On another the best number of iterations is well below the value common in practice, so extra depth wastes budget that more seeds would turn into score\. We provide a measurement protocol that reports the seeds\-by\-iterations frontier and practical guidance for using it\.
Figure 1:An illustration of the need for rigorous evaluation methodology in LLM\-based evolutionary search\.## 1Introduction
Recent years have seen the emergence of a new paradigm for algorithm discovery, in which large language models generate and iteratively refine candidate algorithms through evolutionary search\([Ma et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib4);[Fernando et al\., 2023](https://arxiv.org/html/2609.19799#bib.bib22);[Guo et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib5);[Madaan et al\., 2023](https://arxiv.org/html/2609.19799#bib.bib21)\)\. FunSearch found new mathematical constructions this way\([Romera\-Paredes et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib1)\), AlphaEvolve extended the idea to algorithms and hardware kernels\([Novikov et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib2)\), and open engines such as OpenEvolve reproduce the loop for public use\([Sharma, 2025](https://arxiv.org/html/2609.19799#bib.bib6)\)\. The execution of these methods is determined by two dimensions, width and depth\. Width is the number of independent seeds launched, and the best result over seeds is kept\. Depth is the number of iterations each seed runs\. Every seed and every iteration spends LLM calls, so a study has a proposal budgetB=kB=kseeds×t\\times\\,titerations\. A method is usually run for one seed, sometimes up to three, for a fixed number of iterations, and methods are ranked from that single point\. The budget is chosen by habit\. This paper shows that a single point is not enough to evaluate or compare these methods\([Henderson et al\., 2018](https://arxiv.org/html/2609.19799#bib.bib23)\)\.
In this work, we address the methodological gap by evaluating three evolutionary search strategies on five optimization tasks over a grid of forty seeds by two hundred iterations\. We find that the optimal width\-versus\-depth split depends on the strategy, task, and budget; that the strategy ranking changes with the budget, so on one task the strategy worst at one seed is best at forty; and that the best number of iterations is often well below common practice, wasting budget more seeds would turn into score\.
The message is practical\. Report the seeds\-by\-iterations frontier, not one point\. Run several seeds\. Tune depth per strategy and task\.
### Contributions\.
- •A budget\-grid evaluation protocol and practical reporting guidance: replay logged trajectories and compute the expected best score per split\(k,t\)\(k,t\)with exact order statistics \(§[2](https://arxiv.org/html/2609.19799#S2), §[5](https://arxiv.org/html/2609.19799#S5)\)\.
- •Evidence across three strategies and five tasks that the optimal width\-versus\-depth split depends on strategy, task, and budget \(§[4](https://arxiv.org/html/2609.19799#S4)\)\.
- •A ranking inversion: single\-seed evaluation reorders the strategies relative to a multi\-seed budget, so the common protocol can name the wrong winner \(§[4](https://arxiv.org/html/2609.19799#S4)\)\.
### Related work\.
LLM evolutionary search pairs an LLM proposer with an evaluator in an evolutionary loop: FunSearch\([Romera\-Paredes et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib1)\), AlphaEvolve\([Novikov et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib2)\), OpenEvolve\([Sharma, 2025](https://arxiv.org/html/2609.19799#bib.bib6)\), GEPA\([Agrawal et al\., 2026](https://arxiv.org/html/2609.19799#bib.bib17)\), heuristic\-design systems\([Liu et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib13);[Ye et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib14)\), sample\-efficient variants\([Lange et al\., 2026](https://arxiv.org/html/2609.19799#bib.bib15)\), and recent work on harness engineering\([Ishibashi et al\., 2026](https://arxiv.org/html/2609.19799#bib.bib18)\)\. ADRS\-Bench collects systems\-optimization tasks for this setting\([Cheng et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib7)\)\. Across these methods the budget is fixed in advance and runs are reported singly or in a handful\. A parallel line studies how to spend a fixed sampling or search budget\([Gideoni et al\., 2026](https://arxiv.org/html/2609.19799#bib.bib9);[Ellis and Castro, 2026](https://arxiv.org/html/2609.19799#bib.bib8);[Brown et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib11);[Schaeffer et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib19);[Kazdan et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib20);[Snell et al\., 2024](https://arxiv.org/html/2609.19799#bib.bib12);[Inoue et al\., 2026](https://arxiv.org/html/2609.19799#bib.bib16)\), none comparing evolutionary strategies across a seeds\-by\-iterations budget as we do\. Finally, because individual runs are noisy, a best\-of\-kkscore inherently scales withkk\([Smith and Winkler, 2006](https://arxiv.org/html/2609.19799#bib.bib24)\)\. We connect these threads: we measure the width\-versus\-depth split for each strategy and task, and show the reported winner depends on it\.
## 2Method
### Setup\.
A strategy runs independent seeds\. Seediiproduces a trajectory of candidate scoresf\(xi,1\),…,f\(xi,t\)f\(x\_\{i,1\}\),\\dots,f\(x\_\{i,t\}\), whereffis the task objective \(combined\_score, higher is better\)\. Its value at depthttis the running best
Mi\(t\)=max1≤s≤tf\(xi,s\)\.M\_\{i\}\(t\)=\\max\_\{1\\leq s\\leq t\}f\(x\_\{i,s\}\)\.\(1\)Launchingkkseeds and keeping the best gives
ℬk\(t\)=max1≤i≤kMi\(t\),\\mathcal\{B\}\_\{k\}\(t\)=\\max\_\{1\\leq i\\leq k\}M\_\{i\}\(t\),\(2\)and the budget isB=ktB=k\\,t\.
### Budget grid\.
From the logged trajectories of then=40n\{=\}40seeds we compute, for every split\(k,t\)\(k,t\), the expected best score𝔼\[ℬk\(t\)\]\\mathbb\{E\}\\\!\\left\[\\mathcal\{B\}\_\{k\}\(t\)\\right\]over random size\-kksubsets of the4040observed seeds \(a finite\-sample expectation\)\. We use exact order statistics on the per\-seed values\{Mi\(t\)\}\\\{M\_\{i\}\(t\)\\\}, so there is no resampling noise\. This yields a score surface over seeds and iterations for each strategy and task \(Fig\.[2](https://arxiv.org/html/2609.19799#S2.F2)\)\. We examine three axes:
- •Width marginal:𝔼\[ℬk\(T\)\]\\mathbb\{E\}\[\\mathcal\{B\}\_\{k\}\(T\)\]againstkkat full depthT=200T\{=\}200, the value of more seeds\.
- •Depth marginal:𝔼\[ℬ1\(t\)\]\\mathbb\{E\}\[\\mathcal\{B\}\_\{1\}\(t\)\]againstttfor one seed, the value of more iterations\.
- •Frontier: for each total budgetBB, the best expected score over all splits withkt≤Bk\\,t\\leq B, and the split\(k,t\)\(k,t\)that reaches it\. That split is the recommended allocation at budgetBB\.
Replaying logged trajectories ensures exact comparisons without requiring redundant runs\.
Figure 2:Expected bestcombined\_scoreover the seeds \(kk\) by iterations \(tt\) grid \(EvoX top, OpenEvolve middle, AdaEvolve bottom\)\. Log\-scaled axes make the white iso\-budget lines \(k⋅tk\\cdot t\) straight\. Stars mark the score\-maximizing cells\. The star sits in a different place in each panel: the best split depends on the strategy and the task, and the best depth is often well short of the full two hundred iterations\.
### Ranking\.
At a budget point we rank strategies by expected best score\. The single\-seed protocol common in practice is the pointk=1,t=Tk\{=\}1,\\,t\{=\}T\. We compare its ranking to the multi\-seed budgetk=40,t=Tk\{=\}40,\\,t\{=\}T\.
### Ranking reversion\.
We also ask how often a partial budget names the wrong order\. We bootstrap the seeds: for each of two thousand resamples we rank the strategies by best\-of\-seeds score at a budget point and at the full budget, and record whether the orders differ\. The reversion probability is the share of resamples that differ, zero at the full budget by construction\. We read it over iterations at full width, seeds at full depth, and total budget at its best split\.
## 3Experimental Setup
### Engine and strategies\.
All runs use the ADRS engine\([Cheng et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib7)\)withgpt\-5\-mini, keeping the model, objective, and evaluator identical so only the strategy varies\. The three strategies areOpenEvolve\(islands plus archive, population 40;[Sharma, 2025](https://arxiv.org/html/2609.19799#bib.bib6)\),EvoX\(co\-evolving its selection rule,[Liu et al\., 2026a](https://arxiv.org/html/2609.19799#bib.bib3)\), andAdaEvolve\(adaptive variant,[Cemri et al\., 2026](https://arxiv.org/html/2609.19799#bib.bib10)\)\. Following prior work, we budget by proposals rather than raw LLM calls, using an 8,000\-proposal limit \(40×20040\\times 200\) to ensure an equal comparison\.
### Tasks and objective\.
We use three ADRS\-Bench system tasks\([Cheng et al\., 2025](https://arxiv.org/html/2609.19799#bib.bib7)\):PRISM\(LLM\-serving scheduler, measured by goodput\);Cloudcast\(broadcast planner, transfer cost\); andtransaction scheduling\(throughput\)\. We optimize their nativecombined\_scores \(PRISMreciprocal\(max\_kvpr\)\+success\_rate\\text\{reciprocal\}\(\\text\{max\\\_kvpr\}\)\+\\text\{success\\\_rate\}, transaction scheduling106/\(1\+makespan\)10^\{6\}/\(1\+\\text\{makespan\}\), Cloudcast cost\-based\), gating out invalid or reward\-hacking candidates via validity checks for comparability\. But the transaction\-scheduling evaluator is incomplete: it does not verify that a schedule covers the whole workload, so a partial\-schedule exploit passes \(Appendix[D](https://arxiv.org/html/2609.19799#A4)\); we keep whatever it accepts, as prior work does\. PRISM and Cloudcast are deterministic; transaction scheduling draws one random workload per evaluation, so we re\-evaluate each seed’s best program to remove single\-draw noise\. We also analyze two math tasks \(Appendix[C](https://arxiv.org/html/2609.19799#A3)\): Circle Packing \(Square\), and Heilbronn \(triangle\)\.
## 4Results
### The optimal split depends on strategy, task, and budget\.
Figure[2](https://arxiv.org/html/2609.19799#S2.F2)shows the seeds\-by\-iterations surface for each strategy and task, with iso\-budget lines and the best cell marked\. The best cell sits in different places across panels\. Table[1](https://arxiv.org/html/2609.19799#S4.T1)reports the optimal split at full and 10% budgets, highlighting two patterns\. First, allocations diverge by strategy: on PRISM at 10% budget, EvoX is best deep and narrow \(four seeds,196196iterations\) while OpenEvolve is best wide and shallow \(3636,2222\)\. Second, depth saturates early\. AdaEvolve on transaction scheduling peaks at4040iterations and EvoX on PRISM at6060; further depth wastes proposals better spent on seeds\. Consequently, the optimal allocation axis is highly strategy\- and task\-specific\.
Figure 3:Probability that the strategy order declared at a partial budget differs from the full\-budget order, one curve per task\. Left: iterations at full width\. Middle: seeds at full depth\. Right: total budget at its best split\. Every curve falls to zero as the budget approaches full\. The order settles late with depth but steadily with more seeds\. Shaded bands are bootstrap95%95\\%confidence intervals\.Table 1:Optimal budget split\(k,t\)\(k,t\)withkt≤Bkt\\leq Bat the full and≈\\approx10% budgets\. Allocations vary by strategy and budget\. On PRISM at 10%, prescriptions are opposite \(EvoX deep/narrow, OpenEvolve wide/shallow\)\. AdaEvolve on transaction scheduling peaks at4040iterations; further depth is wasted\.Table 2:Expected bestcombined\_scoreat one and forty seeds \(full depth\), with rank in parentheses \(higher is better\)\. The order changes with the seed count on every task\.
### The ranking inverts with the budget\.
Table[2](https://arxiv.org/html/2609.19799#S4.T2)gives each strategy at one seed and at forty seeds, both at full depth\. On Cloudcast the order reverses: EvoX is last at one seed and first at forty seeds, and OpenEvolve moves from second to last\. On transaction scheduling EvoX and OpenEvolve swap the second and third places\. On PRISM the single\-seed scores are within0\.030\.03\(a near tie\) while at forty seeds AdaEvolve leads by0\.530\.53\. These crossings arise because strategies use budgets differently: added seeds help AdaEvolve far more than EvoX on PRISM, while EvoX gains on both axes on Cloudcast, so single\-point measurements observe only one corner of this surface\. Figure[3](https://arxiv.org/html/2609.19799#S4.F3)turns this into a probability: small budgets yield unreliable rankings; while more depth corrects the order late, more seeds fix it steadily\. On every task a large part of the budget is needed before the order is reliable\. More seeds also surfaced an AdaEvolve program scoring over4×4\\timesany prior result, a validator exploit, not a better scheduler \(Appendix[D](https://arxiv.org/html/2609.19799#A4)\); that more seeds expose such flaws is part of our argument\.
## 5Conclusion
Evaluating LLM evolutionary search at one budget point is not enough\. The best split between seeds and iterations, and even which method wins, depend on the strategy, the task, and the total budget\. We recommend three practices\. Report the seeds\-by\-iterations frontier and the optimal split at each budget, not a single score\. Run several seeds, since single\-seed rankings are unreliable\. Tune depth per strategy and task rather than fixing iterations by habit\. All three follow from replaying logged trajectories, so they cost no extra runs\.
## Limitations
While this study introduces a rigorous assessment methodology necessary for validating improvements in evolutionary processes, we acknowledge that this protocol imposes a higher computational and financial barrier\. Evaluating a full grid of seeds and iterations requires significantly more LLM API calls per assessment than traditional single\-point evaluations\. Additionally, our budget\-grid protocol is an offline evaluation tool designed to accurately rank methods post\-hoc\. While it reveals that the optimal width\-versus\-depth split is task\-dependent, a prediction of this optimal split, without first exploring the grid, is still an open question\.
## References
- Agrawalet al\.\(2026\)L\. A\. Agrawal, S\. Tan, D\. Soylu, N\. Ziems, R\. Khare, K\. Opsahl\-Ong, A\. Singhvi, H\. Shandilya, M\. J\. Ryan, M\. Jiang,et al\.Gepa: reflective prompt evolution can outperform reinforcement learning\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 8479–8565\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Brownet al\.\(2024\)B\. Brown, J\. Juravsky, R\. Ehrlich, R\. Clark, Q\. V\. Le, C\. Ré, and A\. MirhoseiniLarge language monkeys: scaling inference compute with repeated sampling\.arXiv preprint arXiv:2407\.21787\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Cemriet al\.\(2026\)M\. Cemri, S\. Agrawal, A\. Gupta, S\. Liu, A\. Cheng, Q\. Mang, A\. Naren, L\. E\. Erdogan, K\. Sen, M\. Zaharia,et al\.Adaevolve: adaptive llm driven zeroth\-order optimization\.arXiv preprint arXiv:2602\.20133\.Cited by:[§3](https://arxiv.org/html/2609.19799#S3.SS0.SSS0.Px1.p1.1)\.
- Chenget al\.\(2025\)A\. Cheng, L\. Liu, S\. Wang, I\. Stoica,et al\.Barbarians at the gate: how ai is upending systems research\.arXiv preprint arXiv:2510\.06189\.External Links:[Link](https://arxiv.org/abs/2510.06189)Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.19799#S3.SS0.SSS0.Px1.p1.1),[§3](https://arxiv.org/html/2609.19799#S3.SS0.SSS0.Px2.p1.1)\.
- Ellis and Castro \(2026\)M\. Ellis and P\. CastroDon’t gamble, gamble: an analytical framework for ai\-driven research systems\.arXiv preprint arXiv:2606\.02863\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Fernandoet al\.\(2023\)C\. Fernando, D\. Banarse, H\. Michalewski, S\. Osindero, and T\. RocktäschelPromptbreeder: self\-referential self\-improvement via prompt evolution\.arXiv preprint arXiv:2309\.16797\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Gideoniet al\.\(2026\)Y\. Gideoni, S\. Risi, and Y\. GalSimple baselines are competitive with code evolution\.arXiv preprint arXiv:2602\.16805\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Guoet al\.\(2024\)Q\. Guo, R\. Wang, J\. Guo, B\. Li, K\. Song, X\. Tan, G\. Liu, J\. Bian, and Y\. YangConnecting large language models with evolutionary algorithms yields powerful prompt optimizers\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Hendersonet al\.\(2018\)P\. Henderson, R\. Islam, P\. Bachman, J\. Pineau, D\. Precup, and D\. MegerDeep reinforcement learning that matters\.InProceedings of the AAAI conference on artificial intelligence,Vol\.32\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Inoueet al\.\(2026\)Y\. Inoue, K\. Misaki, Y\. Imajuku, S\. Kuroki, T\. Nakamura, and T\. AkibaWider or deeper? scaling llm inference\-time compute with adaptive branching tree search\.Advances in Neural Information Processing Systems38,pp\. 35448–35484\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Ishibashiet al\.\(2026\)Y\. Ishibashi, T\. Yano, and M\. OyamadaEffective harness engineering for algorithm discovery with coding agents\.arXiv preprint arXiv:2605\.15221\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Kazdanet al\.\(2025\)J\. Kazdan, R\. Schaeffer, Y\. Allouah, C\. Sullivan, K\. Yu, N\. Levi, and S\. KoyejoEfficient prediction of pass@ k scaling in large language models\.arXiv preprint arXiv:2510\.05197\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Langeet al\.\(2026\)R\. Lange, Y\. Imajuku, and E\. CetinShinkaevolve: towards open\-ended and sample\-efficient program evolution\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 74026–74078\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2024\)F\. Liu, T\. Xialiang, M\. Yuan, X\. Lin, F\. Luo, Z\. Wang, Z\. Lu, and Q\. ZhangEvolution of heuristics: towards efficient automatic algorithm design using large language models\.InInternational Conference on Machine Learning \(ICML\),Note:arXiv:2401\.02051Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2026a\)S\. Liu, S\. Agarwal, M\. Maheswaran, M\. Cemri, Z\. Li, Q\. Mang, A\. Naren, E\. Boneh, A\. Cheng, M\. Z\. Pan,et al\.Evox: meta\-evolution for automated discovery\.arXiv preprint arXiv:2602\.23413\.Cited by:[§3](https://arxiv.org/html/2609.19799#S3.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2026b\)S\. Liu, M\. Cemri, S\. Agarwal, A\. Krentsel, A\. Naren, Q\. Mang, Z\. Li, A\. Gupta, M\. Maheswaran, A\. Cheng,et al\.Skydiscover: a flexible, adaptive framework for ai\-driven scientific and algorithmic discovery\.InProceedings of the ACM Conference on AI and Agentic Systems,pp\. 1223–1227\.Cited by:[Appendix A](https://arxiv.org/html/2609.19799#A1.p1.1)\.
- Maet al\.\(2024\)Y\. J\. Ma, W\. Liang, G\. Wang, D\. Huang, O\. Bastani, D\. Jayaraman, Y\. Zhu, L\. Fan, and A\. AnandkumarEureka: human\-level reward design via coding large language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Madaanet al\.\(2023\)A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang,et al\.Self\-refine: iterative refinement with self\-feedback\.Advances in neural information processing systems36,pp\. 46534–46594\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Novikovet al\.\(2025\)A\. Novikov, N\. Vũ, M\. Eisenberger, E\. Dupont, P\. Huang, A\. Z\. Wagner, S\. Shirobokov, B\. Kozlovskii, F\. J\. R\. Ruiz, A\. Mehrabian, M\. P\. Kumar, A\. See, S\. Chaudhuri, G\. Holland, A\. Davies, S\. Nowozin, P\. Kohli, and M\. BalogAlphaEvolve: a coding agent for scientific and algorithmic discovery\.Technical reportGoogle DeepMind\.External Links:[Link](https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/AlphaEvolve.pdf)Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Romera\-Paredeset al\.\(2024\)B\. Romera\-Paredes, M\. Barekatain, A\. Novikov, M\. Balog, M\. P\. Kumar, E\. Dupont, F\. J\. R\. Ruiz, J\. S\. Ellenberg, P\. Wang, O\. Fawzi, P\. Kohli, and A\. FawziMathematical discoveries from program search with large language models\.Nature625\(7995\),pp\. 468–475\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06924-6)Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19799#S1.p1.1)\.
- Schaefferet al\.\(2025\)R\. Schaeffer, J\. Kazdan, J\. Hughes, J\. Juravsky, S\. Price, A\. Lynch, E\. Jones, R\. Kirk, A\. Mirhoseini, and S\. KoyejoHow do large language monkeys get their power \(laws\)?\.arXiv preprint arXiv:2502\.17578\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Sharma \(2025\)A\. SharmaOpenEvolve: an open\-source implementation of alphaevolve\.Note:[https://github\.com/codelion/openevolve](https://github.com/codelion/openevolve)Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.19799#S1.p1.1),[§3](https://arxiv.org/html/2609.19799#S3.SS0.SSS0.Px1.p1.1)\.
- Smith and Winkler \(2006\)J\. E\. Smith and R\. L\. WinklerThe optimizer’s curse: skepticism and postdecision surprise in decision analysis\.Management Science52\(3\),pp\. 311–322\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Snellet al\.\(2024\)C\. Snell, J\. Lee, K\. Xu, and A\. KumarScaling llm test\-time compute optimally can be more effective than scaling model parameters\.External Links:2408\.03314Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
- Yeet al\.\(2024\)H\. Ye, J\. Wang, Z\. Cao, F\. Berto, C\. Hua, H\. Kim, J\. Park, and G\. SongReevo: large language models as hyper\-heuristics with reflective evolution\.Advances in neural information processing systems37,pp\. 43571–43608\.Cited by:[§1](https://arxiv.org/html/2609.19799#S1.SS0.SSS0.Px2.p1.1)\.
## Appendix AImplementation details
All experiments run on SkyDiscover\([Liu et al\., 2026b](https://arxiv.org/html/2609.19799#bib.bib25)\), an open framework for AI\-driven algorithmic discovery, using its strategy implementations, task evaluators, and logging\. Building on it keeps our pipeline reproducible\.
## Appendix BBudget frontier
Figure[4](https://arxiv.org/html/2609.19799#A2.F4)plots the best expected score against the total budget, using the optimal split at each budget\. The gaps between strategies open and close as the budget grows, and the leader changes on some tasks, the effect Figure[3](https://arxiv.org/html/2609.19799#S4.F3)quantifies\.
Figure 4:Best expectedcombined\_scoreagainst the LLM\-call budget, one column per task\. Top: value of more seeds at full depth\. Middle: value of more iterations for one seed\. Bottom: the frontier, the best score at each total budget using the optimal split\. The gaps between strategies open and close as the budget grows, so the strategy to prefer depends on how many calls are available\. The no\-evolve best\-of\-N line is a reference, not part of the strategy comparison\. Shaded bands span the 25th–75th percentile of the best\-of\-kkdistribution\.Figure[4](https://arxiv.org/html/2609.19799#A2.F4)keeps every seed, so a few lucky seeds set the top of the best\-of\-k and can mask the typical trend\. Figure[5](https://arxiv.org/html/2609.19799#A2.F5)repeats it with each technique’s best seed dropped \(top two on transaction scheduling, top one on the other tasks, and the best draw of the no\-evolve baseline\)\. Read it for the trend, not for absolute scores\.
Figure 5:Same as Figure[4](https://arxiv.org/html/2609.19799#A2.F4)with the best seed\(s\) per technique removed\. Trend\-reading only\.
## Appendix CMath tasks
The two math tasks show the same budget effects as the systems tasks \(Figure[3](https://arxiv.org/html/2609.19799#S4.F3)\)\. Figure[6](https://arxiv.org/html/2609.19799#A3.F6)plots the seeds\-by\-iterations surface: the best cell again lands in different places across strategies, and depth saturates before the full budget\. Figure[7](https://arxiv.org/html/2609.19799#A3.F7)shows the frontier rising with the budget, with the gaps between strategies opening and closing as before\. Table[3](https://arxiv.org/html/2609.19799#A3.T3)confirms the ranking inversion: on Heilbronn, AdaEvolve leads at one seed but EvoX overtakes it at forty, exactly the crossing a single\-seed protocol would miss; circle packing demonstrates a near tie between EvoX and AdaEvolve at both budgets\. Both tasks support the same conclusion as the systems tasks\.
Figure 6:Seeds\-by\-iterations surface for the two math tasks; the best cell is marked\. As in the systems tasks, it moves across strategies and depth saturates early\.Figure 7:Best expected score against the budget for the math tasks\. Top: seeds at full depth\. Middle: iterations for one seed\. Bottom: the frontier at the optimal split\.Table 3:Expected best score at one and forty seeds \(full depth\), rank in parentheses \(higher is better\)\. On Heilbronn, AdaEvolve leads at one seed but EvoX at forty; circle packing is a near tie\.
## Appendix DThe discovered transaction scheduler
The transaction scheduling column of Table[2](https://arxiv.org/html/2609.19799#S4.T2)looks like our strongest result, but part of it is an artifact of the evaluator\. The validity check,validate\_schedule, only tests that the returned sequence is a permutation of the indices it contains\. It never checks that every transaction was scheduled\. Each workload holds100100transactions, so a program can drop some of them, lower its makespan, and still pass withvalidity=1\.0=1\.0\.
Running more seeds exposed this gap, and different strategies use it to different degrees \(Table[4](https://arxiv.org/html/2609.19799#A4.T4)\)\. AdaEvolve is the extreme case: its best program scorescombined\_score2127721277\(makespan4646\), about4×4\\timesany prior result, by scheduling a single transaction per workload and skipping the other9999\. Two of its forty seeds do this\. OpenEvolve’s best program \(49734973\) is subtler: it drops about4545of the100100transactions in one workload and completes the other two, which is why its makespan \(200200\) sits below EvoX’s \(227227\)\. Only EvoX schedules every transaction\.
We score under the official evaluator to stay comparable with prior work, so we keep every program in the tables and figures\. But it means the leaderboard on this task partly reflects how much of the workload a program quietly drops, not how well it schedules\. A sound evaluator would need a coverage check, which this benchmark lacks\.
We think this makes our point stronger\. A one\- or three\-seed run would rarely reach these programs, while forty seeds surface them, change the ranking, and stress\-test the benchmark itself, exposing an evaluator flaw a small budget would miss\.
Table 4:The best transaction\-scheduling program found by each strategy at forty seeds, under the official evaluator\. Each of the three workloads holds100100transactions, for300300in total\. AdaEvolve and OpenEvolve both pass the validity check while leaving transactions unscheduled, which lowers their makespan and inflatescombined\_score; only EvoX schedules the full workload\. The evaluator accepts all three becausevalidate\_schedulenever checks coverage\.Similar Articles
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
Large-scale study of 15 LLMs across 8 tasks reveals that optimization success hinges on maintaining localized search trajectories rather than initial problem-solving ability or solution novelty.
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
This paper proposes using Evolution Strategies (ES) instead of Reinforcement Learning for post-training LLMs, showing that ES improves solution coverage (pass@k) and achieves better results on math benchmarks.
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
This paper investigates Evolution Strategies (ES) as a post-training paradigm for LLM reasoning, showing that ES provides broader reasoning coverage and better Pass@K performance than GRPO through sparse functional updates and population diversity.
Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
This paper proposes a method for co-evolving evaluation metrics and skills in self-improving LLM agent systems, demonstrating that metrics can be evolved and that a co-evolution approach recovers most of the performance of a ground-truth-driven oracle across code generation, text-to-SQL, and report generation tasks.
Beyond Static Evaluation: Co-Evolutionary Mechanisms for LLM-Driven Strategy Evolution in Adversarial Games
This paper proposes three co-evolutionary mechanisms (evaluator co-evolution, hierarchical deep evaluation, and weakness pressure) for LLM-driven code evolution in adversarial multi-agent games, achieving state-of-the-art results on the MCTF 2026 maritime capture-the-flag task.