AIM:面向自动化研究的智能体化想法管理

arXiv cs.AI 论文

摘要

Google Cloud AI Research 推出了 AIM(Agentic Idea Manager),这是一个以想法为驱动的自动化研究自主框架。它采用 Agentic Surrogate 和 Agentic Acquisition(灵感来自贝叶斯优化)来组织和筛选研究想法,并通过 Solution Auditor 和 Resource Planner 维持想法与解决方案的一致性以及分配实验预算。AIM 的表现最多比最强的 AutoLab 基线高出 4.9 个百分点,并且最高可提前 3.1 倍达到基线性能。

arXiv:2609.38445v1 Announce Type: new Abstract: Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations. To address these challenges, we introduce the Agentic Idea Manager (AIM), a fully autonomous framework for managing and exploring research directions in idea-driven automated research. Inspired by Bayesian optimization, AIM uses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection. A Solution Auditor maintains idea-solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches. Experiments on 10 AutoLab benchmark tasks show that AIM surpasses the strongest baseline by 1.6 percentage points on System Optimization tasks and 4.9 percentage points on long-horizon Model Development & CUDA tasks. Notably, AIM reaches the best baseline performance up to 3.1x faster in wall-clock time. We further provide a theoretical analysis of when searching over ideas becomes beneficial. Our analysis shows that explicit idea-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives. Project Page: https://imhgchoi.github.io/agentic-idea-manager/
查看原文
查看缓存全文

缓存时间: 2026/10/01 09:41

# AIM: Agentic Idea Management for Automated Research
Source: [https://arxiv.org/html/2609.38445](https://arxiv.org/html/2609.38445)
Hyeong Kyu ChoiBhavana Dalvi MishraAffiliation:Google Cloud AI ResearchJiefeng ChenAffiliation:Google Cloud AI ResearchMihir ParmarAffiliation:Google Cloud AI ResearchRui MengAffiliation:Google Cloud AI ResearchChun\-Liang LiAffiliation:Google Cloud AI Research Xiangru TangAffiliation:Google Cloud AI ResearchSharon LiAffiliation:University of Wisconsin\-MadisonJinsung YoonAffiliation:Google Cloud AI ResearchTomas PfisterAffiliation:Google Cloud AI Research

###### Abstract

Frontier LLMs are increasingly used to automate scientific research through iterative search\. We distinguish idea\-driven search from solution\-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementations\. To address these challenges, we introduce the*Agentic Idea Manager*\(AIM\), a fully autonomous framework for managing and exploring research directions in idea\-driven automated research\. Inspired by Bayesian optimization,AIMuses an Agentic Surrogate and an Agentic Acquisition mechanism to organize discovered ideas and guide their selection\. A Solution Auditor maintains idea–solution integrity, while a Resource Planner adaptively allocates the remaining experimental budget across parallel search branches\. Experiments on 10 AutoLab benchmark tasks show thatAIMsurpasses the strongest baseline by 1\.6 percentage points on System Optimization tasks and 4\.9 percentage points on long\-horizon Model Development & CUDA tasks\. Notably,AIMreaches the best baseline performance up to 3\.1×\\timesfaster in wall\-clock time\. We further provide a theoretical analysis ofwhen searching over ideas becomes beneficial\. Our analysis shows that explicit idea\-level allocation makes semantic coverage directly controllable, and that broader coverage becomes increasingly valuable when competitive research directions are sparse among many plausible alternatives\. Project Page:[https://imhgchoi\.github\.io/agentic\-idea\-manager/](https://imhgchoi.github.io/agentic-idea-manager/)

### 1Introduction

Large language model \(LLM\) agents are increasingly used to automate scientific research through iterative experimentation\. Modern research agents can propose candidate approaches, implement solutions, execute experiments, inspect verifier feedback, and use accumulated evidence to determine what to try next\([Lu et al\., 2024](https://arxiv.org/html/2609.38445#bib.bib1);[Jiang et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib23);[Toledo et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib22);[Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\)\. Because implementation and evaluation are expensive, effective automated research requires allocating a limited experimental budget across possible approaches\. The resulting search process must support both the discovery of promising directions and the refinement of their implementations\.

In this work, we introduce the first categorization of automated research, according to its primary unit of search:*solution\-driven*and*idea\-driven*\. Solution\-driven approaches \(Figure[1](https://arxiv.org/html/2609.38445#S1.F1)\(a\)\) search directly over executable artifacts, using verifier feedback to iteratively modify candidate solutions\([Cemri et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib26);[Liu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib27);[Toledo et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib22);[Jiang et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib23)\)\. Idea\-driven approaches \(Figure[1](https://arxiv.org/html/2609.38445#S1.F1)\(b\)\), on the other hand, maintain research ideas as explicit object of reasoning; they select*which idea to investigate*and delegate*how to implement it*to a solver\([Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16);[Jin et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib28);[Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2);[Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\)\. Both paradigms ultimately evaluate executable solutions, but organize search at different levels of abstraction\. Working directly with solutions supports fine\-grained code refinement, while idea\-driven methods facilitate comparison across approaches, helps the transfer of lessons, and makes research trajectories easier to interpret\.

![Refer to caption](https://arxiv.org/html/2609.38445v1/B_Figures/figs/stage_comparison.png)Figure 1:Solution\-driven vs Idea\-driven comparison\.While Solution\-driven approaches take the solution code as the target object for optimization, Idea\-driven approaches first reason over research ideas, directions, or hypotheses based on evaluator feedback, after which the chosen ideas are implemented by the solver agent\.This categorization highlights the management of research ideas as a distinct design problem in automated research, with three central challenges in determining the*idea management structure*,*idea selection strategies*, and ensuring*idea–solution integrity*\. To address these challenges, we introduce theAgentic Idea Manager \(AIM\), a fully\-autonomous research idea management framework for automated research \(Figure[2](https://arxiv.org/html/2609.38445#S4.F2)\)\. Inspired by the surrogate\-acquisition structure of Bayesian optimization, its*Agentic Surrogate*maintains a dynamic structure of idea clusters and ranks their promise using experimental evidence, grouping related proposals across generation lineages\. The*Agentic Acquisition*mechanism balances exploration and exploitation at both cluster\- and idea\-level selection, dispatching chosen candidates to parallel solvers\. After implementation and evaluation, the*Solution Auditor*checks task validity and idea–solution integrity, and the audited evidence guides subsequent direction of research\. Finally, the*Resource Planner*adjusts parallelism across iterations, balancing concurrent trials with longer exploration under a fixed experimental budget\.

We evaluateAIMon ten AutoLab tasks spanning System Optimization and Model Development & CUDA tasks\.AIMachieves the highest average scores in both task groups: 67\.0% on System Optimization and 55\.8% on long\-horizon Model Development & CUDA, exceeding the strongest baseline, ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\), by 1\.6 and 4\.9 percentage points, respectively\. Moreover,AIMreaches ScientistOne’s best score up to 3\.1×\\timesfaster, demonstrating its efficiency and supporting the value of agentic idea organization, evidence\-guided selection, and audited feedback for automated research\. In addition to empirical studies, we further provide theoretical analyses to identify when explicit idea\-level search is useful\. Our analysis shows that idea\-driven search makes such breadth directly controllable through explicit allocation to distinct ideas\. More importantly, broader coverage becomes increasingly valuable when a task contains many plausible research directions but only a small fraction are competitive\.

Our main contributions are:

- •We introduce a categorization of automated research methods:solution\-drivenandidea\-driven\. We identifythree central challengesof idea\-driven approaches: idea management, idea selection, and idea–solution integrity\.
- •We introduceAgentic Idea Manager\(AIM\), a fully\-agentic idea\-driven framework that effectively manages ideas as semantic clusters, selects ideas based on its explicit estimation of performance, and audits the solutions to achieve a reliable research pipeline\.
- •We achievestrong empirical performanceon 10 Autolab tasks, surpassing the strongest baseline by 1\.6pp on System Optimization and 4\.9pp on long\-horizon Model Development & CUDA tasks\. We also providerigorous theoretical analysesof when idea\-driven search can be useful\.

### 2Related Works

###### Solution\-driven Search\.

Solution\-driven approaches remain close to the verifier and make the executable implementation the primary unit of search\. AIDE\([Jiang et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib23)\)and AIRA\([Toledo et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib22)\)explore candidate solutions through tree search, while AlphaEvolve\([Novikov et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib15)\), AdaEvolve\([Cemri et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib26)\), and EvoX\([Liu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib27)\)use evolutionary mechanisms\. MLE\-Star\([Nam et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib19)\), DS\-Star\([Nam et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib20)\), and RPM\([Foster et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib38)\)further develop solution\-level refinement and selection\. Solution\-driven search’s tight coupling of the evaluator and solver lets experimental feedback directly guide implementation changes, but can entangle progress on a research direction with engineering decisions about a particular artifact\.

###### Idea\-driven Search\.

Idea\-driven approaches make research ideas or hypotheses explicit search objects, then delegate their implementation to a solver\. The AI Scientist\-v2\([Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2)\), MARS\([Chen et al\., 2026b](https://arxiv.org/html/2609.38445#bib.bib21)\), and Arbor\([Jin et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib28)\)manages a tree of ideas or lessons, and DeepScientist\([Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\)maintains the research ideas and hypotheses as a list\. A subset of ideas is implemented by a solver module\. ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\), on the other hand, utilizes a beam\-search\-like algorithm that keeps the best\-performing ideas throughout the research iterations\. These approaches separates idea search from implementation, enabling deliberate exploration of distinct ideas, but requires deciding which should merit costly experiments and evaluations\.

This idea\-driven search introduces three central challenges: \(1\)Idea management scaffold: the framework must organize an expanding pool of ideas and accumulated evidence into a persistent, interpretable representation\. \(2\)Idea selection: it must select which ideas to implement and how to allocate the remaining budget; existing methods typically rely on fixed rules such as UCB\([Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\)or MCTS\([Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2)\), or naively use LLM judgments without explicit estimates\([Jin et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib28)\)\. \(3\)Idea–Solution integrity: the separation between ideation and implementation creates an integrity risk\. A solver may produce code that does not faithfully realize the selected idea, causing its score and derived lessons to be misattributed\. These challenges motivate jointly managing the idea space, making evidence\-grounded search decisions, and verifying idea–solution alignment, which we address with the our proposed method\. In this work, we propose an idea\-driven research framework that addresses these challenges\. A formal discussion on when to search over ideas is provided in Section[6](https://arxiv.org/html/2609.38445#S6)\. An extended section for Related Works is in Appendix[A](https://arxiv.org/html/2609.38445#A1)\.

### 3Automated Research: Problem Definition

In automated research, it is generally infeasible to evaluate every plausible idea because implementation and verifier calls are expensive\. Effective automated research therefore requires more than sequential idea and code generation\. It requires a structured representation of the ideas and solutions discovered so far, and a principled mechanism to navigate the search process under a limited computational budget\. Accordingly, we formally define the relevant search space and objectives:

###### Idea and Solution Spaces\.

Let𝒳\\mathcal\{X\}denote the space of admissible research ideas for taskτ\\tau, where eachx∈𝒳x\\in\\mathcal\{X\}is a natural\-language description of a candidate approach\. In this work, anideais defined as, but not limited to, a structured text comprising a short title, brief hypothesis/abstract, and experiment plans \(see Appendix[E\.6](https://arxiv.org/html/2609.38445#A5.SS6.SSSx1)for examples\)\. At timesteptt, the agent has access only to a finite pool𝒫t⊂𝒳\\mathcal\{P\}\_\{t\}\\subset\\mathcal\{X\}containing the ideas discovered thus far\. Meanwhile, we distinguish the space of executable solutions𝒵\\mathcal\{Z\}from the idea space𝒳\\mathcal\{X\}\. Given an ideaxx, an implementation processg:𝒳→𝒵g:\\mathcal\{X\}\\rightarrow\\mathcal\{Z\}produces an executable solutionz=g⁡\(x\)z=g\(x\), assuming a single deterministic mapping per idea in this work, which is then scored by a fixed verifierv:𝒵→ℝv:\\mathcal\{Z\}\\rightarrow\\mathbb\{R\}:

z=g\(x\),y=v\(z\)⟹f\(x\)≜v\(g\(x\)\)\.z=g\(x\),\\qquad y=v\(z\)\\quad\\Longrightarrow\\quad f\(x\)\\triangleq v\(g\(x\)\)\.\(1\)Here,ffdenotes the expensive research process for implementation and evaluation\.

###### Objective\.

Evaluatingf⁡\(x\)f\(x\)requires first translating the natural\-language idea into an executable implementation and then compiling and running that implementation against a fixed verifier\. Consequently, the dominant cost is not in proposing an idea, but obtaining reliable evidence about its quality through implementation and verification, which can itself even be very expensive and noisy\. Thus, the core objective is to identify the highest\-performing idea within a given computational budget measured in the number of experiment executions and/or runtime hours for search\. Letc⁡\(x\)c\(x\)denote the cost of implementing and verifying ideaxx, and letNNbe the total computational budget\. For a sequence of evaluated ideasx1,…,xnx\_\{1\},\\ldots,x\_\{n\}, the objective is to maximize the best performance:

maxx1,…,xn∈𝒳⁡max1≤i≤n⁡f⁡\(xi\)s\.t\.∑i=1nc⁡\(xi\)≤N\.\\max\_\{x\_\{1\},\\ldots,x\_\{n\}\\in\\mathcal\{X\}\}\\;\\max\_\{1\\leq i\\leq n\}f\(x\_\{i\}\)\\qquad\\text\{s\.t\.\}\\qquad\\sum\_\{i=1\}^\{n\}c\(x\_\{i\}\)\\leq N\.\(2\)Within this framework, the agent must therefore determine which subset of ideas to explore first\.

### 4TheAIMFramework

Overview\.We introduce theAgentic Idea Manager\(AIM\), a fully autonomous framework for managing and searching research ideas under a limited experimental budget\. For a research taskτ\\tau,AIMmaintains the search state

𝒮t=\(𝒯,𝒫t,ℳt,𝒞t,ℛt,𝒟t\),\\mathcal\{S\}\_\{t\}=\\left\(\\mathcal\{T\},\\mathcal\{P\}\_\{t\},\\mathcal\{M\}\_\{t\},\\mathcal\{C\}\_\{t\},\\mathcal\{R\}\_\{t\},\\mathcal\{D\}\_\{t\}\\right\),\(3\)where𝒯\\mathcal\{T\}is the fixed task context,𝒫t\\mathcal\{P\}\_\{t\}is the current idea pool,ℳt\\mathcal\{M\}\_\{t\}is the memory or lessons that store verified lessons from previous experiments,𝒞t\\mathcal\{C\}\_\{t\}is the semantic organization of the pool,ℛt\\mathcal\{R\}\_\{t\}contains ordinal promisingness estimates over clusters and ideas, and𝒟t\\mathcal\{D\}\_\{t\}contains past idea–score observations\. Following[Meng et al\. \(2026\)](https://arxiv.org/html/2609.38445#bib.bib16), we adopt its ideator structure and its construction of the task context\. Specifically,𝒯\\mathcal\{T\}consists of the original task description and an initial research brief that provides supplementary context for ideation111In this work, the research brief is generated from the task description using Claude Code\([Anthropic,](https://arxiv.org/html/2609.38445#bib.bib30)\)\.

Drawing functional inspiration from Bayesian optimization,AIMcomprises an*Agentic Surrogate*that organizes ideator\-generated candidates and estimates the relative promise of clusters and ideas \(Section[4\.1](https://arxiv.org/html/2609.38445#S4.SS1)\), and an*Agentic Acquisition*module that selects ideas and adaptively balances exploration and exploitation \(Section[4\.2](https://arxiv.org/html/2609.38445#S4.SS2)\)\. Each selected idea is implemented and evaluated by an independent Solver\. The Solution Auditor then validates the result and reconciles the intended idea with the mechanism realized in code \(Section[4\.3](https://arxiv.org/html/2609.38445#S4.SS3)\)\. The audited observations and lessons are used to update the search state and expand the idea pool for the next iteration\. Finally, the*Resource Planner*determines how the remaining experiment budget is allocated across subsequent search iterations \(Section[4\.4](https://arxiv.org/html/2609.38445#S4.SS4)\)\. A visual overview is in Figure[2](https://arxiv.org/html/2609.38445#S4.F2), and the algorithm forAIMis in Algorithm[1](https://arxiv.org/html/2609.38445#alg1)\.

![Refer to caption](https://arxiv.org/html/2609.38445v1/B_Figures/figs/mainfig2.png)Figure 2:TheAIMpipeline: theAgentic Surrogateorganizes and estimates candidate idea promise and theAgentic Acquisitionselects ideas for execution and expands the idea pool\. Results are audited by theSolution Auditor, while theResource Plannerdynamically allocates the budget\.#### 4\.1Agentic Surrogate: Organize and Estimate

The Agentic Surrogate module constructs an explicit, evidence\-conditioned representation of the discovered idea space through two operators:*Organize*and*Estimate*\.

Organize\.Given the current idea pool𝒫t\\mathcal\{P\}\_\{t\}and evaluation history𝒟t\\mathcal\{D\}\_\{t\}, the Organize operator constructs a cluster map

𝒞t=ϕorg​\(𝒯,𝒫t,𝒟t,𝒞t−1\)\.\\mathcal\{C\}\_\{t\}=\\phi\_\{\\mathrm\{org\}\}\\left\(\\mathcal\{T\},\\mathcal\{P\}\_\{t\},\\mathcal\{D\}\_\{t\},\\mathcal\{C\}\_\{t\-1\}\\right\)\.\(4\)Each cluster in𝒞t\\mathcal\{C\}\_\{t\}represents a broad research direction and contains semantically related ideas from𝒫t\\mathcal\{P\}\_\{t\}\. Evaluated ideas and their scores serve as empirical landmarks for interpreting related but unevaluated candidates\. Because the map is reconstructed from the complete pool at every timesteptt, ideas from different generation lineages may be grouped together, and the organization may change as new candidates and evidence become available\. A visualization of how the clusters evolve throughout iterations is provided in Figure[3](https://arxiv.org/html/2609.38445#S4.F3), also demonstratingAIM’s interpretability as an idea\-driven approach\.

Estimate\.Conditioned on the organized idea map𝒞t\\mathcal\{C\}\_\{t\}, and evaluation history𝒟t\\mathcal\{D\}\_\{t\}, the Estimate operator produces ordinal estimates of promisingness at both the cluster and idea levels:

ℛt=ϕest​\(𝒞t,𝒟t\)=\(ℛtcluster,ℛtidea\)\.\\mathcal\{R\}\_\{t\}=\\phi\_\{\\mathrm\{est\}\}\\left\(\\mathcal\{C\}\_\{t\},\\mathcal\{D\}\_\{t\}\\right\)=\\left\(\\mathcal\{R\}\_\{t\}^\{\\mathrm\{cluster\}\},\\mathcal\{R\}\_\{t\}^\{\\mathrm\{idea\}\}\\right\)\.\(5\)Here,ℛtcluster\\mathcal\{R\}\_\{t\}^\{\\mathrm\{cluster\}\}ranks the cluster\-level research directions represented in𝒞t\\mathcal\{C\}\_\{t\}, whileℛtidea\\mathcal\{R\}\_\{t\}^\{\\mathrm\{idea\}\}ranks the unevaluated ideas within each cluster\. The estimator considers observed performance, evidence scarcity, semantic novelty, and relevant implementation lessons\. We use ordinal estimates because the purpose is to make relative promisingness explicit without requiring the agent to produce calibrated numerical reward predictions\. An analysis on the preciseness of the estimation is in Appendix[E\.3](https://arxiv.org/html/2609.38445#A5.SS3)\.

![Refer to caption](https://arxiv.org/html/2609.38445v1/example.png)Figure 3:A visualization of a real pipeline example on the Flash Attention task\. The x\-axis shows all the cluster themes explored, numbers in each cell are the best scores from respective idea clusters at each round, and the stars indicate the trailing best score\.
#### 4\.2Agentic Acquisition: Dispatch, Solve, and Expand

The Agentic Acquisition module uses the current map𝒞t\\mathcal\{C\}\_\{t\}and promisingness estimatesℛt\\mathcal\{R\}\_\{t\}to determine how to select which ideas to evaluate\. Rather than relying on fixed selection rules or heuristics,*Dispatch*assigns exploration and exploitation actions using evidence\-grounded estimates of the promisingness of each candidate idea\. After the dispatched ideas are implemented and evaluated,*Expand*uses the resulting evidence to generate candidates for subsequent iterations\.

Dispatch\.Given the current idea scaffold𝒞t\\mathcal\{C\}\_\{t\}, ordinal estimatesℛt\\mathcal\{R\}\_\{t\}, and resource planΠt\\Pi\_\{t\}\(Section[4\.4](https://arxiv.org/html/2609.38445#S4.SS4)\), the Idea Dispatch operatorϕdisp\\phi\_\{\\mathrm\{disp\}\}selects a batch of unevaluated ideas:

𝒬t=ϕdisp​\(𝒫t,𝒞t,ℛt,Πt\)\\mathcal\{Q\}\_\{t\}=\\phi\_\{\\mathrm\{disp\}\}\\bigl\(\\mathcal\{P\}\_\{t\},\\mathcal\{C\}\_\{t\},\\mathcal\{R\}\_\{t\},\\Pi\_\{t\}\\bigr\)\(6\)where𝒬t\\mathcal\{Q\}\_\{t\}contains the ideas assigned toBtB\_\{t\}parallel Solver branches, withBtB\_\{t\}determined byΠt\\Pi\_\{t\}\. In practice, Dispatch is implemented through two sequential LLM calls\. The first call jointly assigns each branch a two\-tier action consisting of a cluster\-level and an idea\-level decision, each selected from\{explore,exploit\}\\\{\\textsc\{explore\},\\textsc\{exploit\}\\\}\. Exploitation prioritizes highly ranked research directions or ideas, whereas exploration favors underexplored or uncertain alternatives\. The second call instantiates these chosen actions by selecting a target cluster and a corresponding idea for each branch\. For instance, the cells activated in Figure[3](https://arxiv.org/html/2609.38445#S4.F3)show which cluster theme was selected at each iteration\. Then, Eachx∈𝒬tx\\in\\mathcal\{Q\}\_\{t\}is passed to an independent Solver module, which produces an executable solutionz=g⁡\(x\)z=g\(x\), a verifier scorey~=v⁡\(z\)\\widetilde\{y\}=v\(z\), and an execution record\. In this work, we mainly use the Gemini Deep Solver utilized in ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\)\(Claude Code solver substitution analysis is in Appendix[D\.1](https://arxiv.org/html/2609.38445#A4.SS1)\)\.

Expand\.The Expand operator updates the search space in two stages\. First, it extracts reusable lessonsΔ​ℳt\+1\\Delta\\mathcal\{M\}\_\{t\+1\}from the audited Solver runs and updates the implementation memory:

ℳt\+1=ℳt∪Δ​ℳt\+1\.\\mathcal\{M\}\_\{t\+1\}=\\mathcal\{M\}\_\{t\}\\cup\\Delta\\mathcal\{M\}\_\{t\+1\}\.\(7\)These lessons capture effective implementation choices, unresolved performance bottlenecks, and compilation, execution, or verification failures to repair or avoid\. This design is inspired by evolutionary frameworks that use lessons distilled from experimental outcomes to guide subsequent discovery\([Cemri et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib26);[Liu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib27)\)\.

Second, conditioned on the updated memory and audited evaluation history, the Expand operator generates new candidates and adds them to the idea pool:

Δ​𝒫t\+1=ϕexp​\(𝒫t,𝒟t\+1,ℳt\+1\),𝒫t\+1=𝒫t∪Δ​𝒫t\+1\.\\Delta\\mathcal\{P\}\_\{t\+1\}=\\phi\_\{\\mathrm\{exp\}\}\\left\(\\mathcal\{P\}\_\{t\},\\mathcal\{D\}\_\{t\+1\},\\mathcal\{M\}\_\{t\+1\}\\right\),\\qquad\\mathcal\{P\}\_\{t\+1\}=\\mathcal\{P\}\_\{t\}\\cup\\Delta\\mathcal\{P\}\_\{t\+1\}\.\(8\)For each candidate, the operator selects relevant source ideas or lessons and applies one of four generation modes:

- •Score\-guided refinementpreserves the validated components of a high\-performing idea while addressing its remaining bottlenecks;
- •Cross\-pollinationcombines complementary ideas or lessons within or across clusters;
- •Error\-guided repairrevises an unsuccessful idea using verifier feedback; and
- •Novel idea generationintroduces a previously unrepresented research direction without requiring an existing parent\.

The first three modes develop existing idea lineages using accumulated evidence, whereasnew\_ideabroadens the search space and prevents concentration on established directions\.

#### 4\.3Solution Auditor

The Solution Auditor prevents invalid or misattributed results from corrupting subsequent search decisions\. For each dispatched ideaxx, implementationzz, reported scorey~\\widetilde\{y\}, and execution recordhh, it first audits the resulting solution:

ℱ=ϕaudit​\(x,z,y~,h\)⊆\{trivial,task​\_​mismatch,idea​\_​mismatch,reward​\_​hacking\},\\mathcal\{F\}=\\phi\_\{\\mathrm\{audit\}\}\(x,z,\\widetilde\{y\},h\)\\subseteq\\\{\\mathrm\{trivial\},\\mathrm\{task\\\_mismatch\},\\mathrm\{idea\\\_mismatch\},\\mathrm\{reward\\\_hacking\}\\\},\(9\)whereℱ=∅\\mathcal\{F\}=\\varnothingdenotes a valid result\. The Auditor checks whether the implementation follows the intended idea, satisfies the task requirements, and obtains its score without exploiting the verifier\.

The audit outcome determines how the result enters the search state\. If trivial, task\_mismatch or reward\_hacking is detected, the result is discarded and excluded from both the evaluation history and lesson extraction, although the execution still consumes the experimental budget\. If the only issue is idea\_mismatch, the Auditor reconstructs the input idea and updates the evaluation history accordingly\. The score and extracted lessons are then associated with the updated idea\.

This audit\-and\-align procedure addresses the idea–solution integrity challenge by creating a feedback cycle between the two stages:Idea→\\rightarrowImplementation→\\rightarrowAudit→\\rightarrowAlign Idea to Implementation\. Thus, subsequent search decisions are grounded in the mechanismactuallyevaluated rather than the initially misaligned one\.

#### 4\.4Resource Planner

The Resource Planner dynamically distributes a fixed total number of Solver branches across search iterations\. First of all, given an execution budgetNNandBtotB\_\{\\mathrm\{tot\}\}total branches, each branch receives a solution execution budget ofNbranch=⌊N/Btot⌋N\_\{\\mathrm\{branch\}\}=\\lfloor N/B\_\{\\mathrm\{tot\}\}\\rfloor\. Then, the planner controls the trade\-off between parallel breadth and sequential adaptivity: wider iterations with largerBtB\_\{t\}evaluate more ideas concurrently, whereas narrower iterations enable more frequent updates from experimental feedback\.

Letbt=∑j=1t−1Bjb\_\{t\}=\\sum\_\{j=1\}^\{t\-1\}B\_\{j\}denote the number of branches already dispatched\. Given the current search state𝒮t\\mathcal\{S\}\_\{t\}, the remaining branches, and maximum parallelismBmaxB\_\{\\max\}, the planner selects

Πt≡Bt=ϕplan​\(𝒮t,Btot−bt,Bmax\),1≤Bt≤min⁡\{Bmax,Btot−bt\}\.\\Pi\_\{t\}\\equiv B\_\{t\}=\\phi\_\{\\mathrm\{plan\}\}\\left\(\\mathcal\{S\}\_\{t\},B\_\{\\mathrm\{tot\}\}\-b\_\{t\},B\_\{\\max\}\\right\),\\qquad 1\\leq B\_\{t\}\\leq\\min\\\{B\_\{\\max\},B\_\{\\mathrm\{tot\}\}\-b\_\{t\}\\\}\.\(10\)The resulting search iteration is

I=min⁡\{i:∑t=1iBt=Btot\},Btot​Nbranch≤N\.I=\\min\\left\\\{i:\\sum\_\{t=1\}^\{i\}B\_\{t\}=B\_\{\\mathrm\{tot\}\}\\right\\\},\\qquad B\_\{\\mathrm\{tot\}\}N\_\{\\mathrm\{branch\}\}\\leq N\.\(11\)The planner changes the frequency of feedback\-driven updates while preserving the total branch and execution budgets\. Thus, the Resource Planner effectively decides when to exhaust the computational budget and terminate the research pipeline, after which the best scoring solution will be chosen as the final output\.

### 5Experiments

Baselines\.We compareAIMwith various solution\-driven and idea\-driven methods\. Solution\-driven baselines include evolutionary solution management approaches like EvoX\([Liu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib27)\), AdaEvolve\([Cemri et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib26)\), AIRA\([Toledo et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib22)\), and MCTS\-based search methods like AIRA \(MCTS version\)\. For the idea\-driven baselines, we compare with methods with various idea management scaffolds, including list\-based DeepScientist\([Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\), MCTS\-based AI\-Scientist\-v2\([Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2)\), tree\-based Arbor\([Jin et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib28)\), and elite\-preserving beam\-search\-based ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\)\. The backbone LLM is Gemini\-3\.1\-Pro\-Preview\([Team et al\., 2023](https://arxiv.org/html/2609.38445#bib.bib33)\)for all methods\.

###### Benchmark Tasks\.

We evaluate our method and baselines on a wide range of tasks from the AutoLab\([Xu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib25)\)suite: System Optimization \(Flash Attention, Radix Sort, AES128 Ctr, FFT Rust, and Z\-order Range Scan\), Model Development \(Moving MNIST World Model, Data Select Ifeval\), and CUDA \(Huffman Canonical Decode, NTT Butterfly, and ICP Correspondence Step\)\. These tasks span different levels of semantic breadth in their viable solution strategies and different degrees of implementation\-level optimization depth\. For the system optimization tasks, we impose a budget of at most 300 experiment executions and a hard wall\-clock limit of 6 hours; for the model development tasks, we set at most 60 executions and wall\-clock limit of 24 hours; for the CUDA tasks, we set at most 60 executions and wall\-clock limit of 12 hours\. If either of the budget is exhausted, the process terminates\. Additional details on the baselines, benchmark tasks, and implementation setup are provided in Appendix[C](https://arxiv.org/html/2609.38445#A3)\.

#### 5\.1Results

Table 1:System Optimization Task Results\. The mean and standard error of three independent runs are reported\. All idea\-driven approaches take the same research brief as additional context for a fair comparison\.###### Results on System Optimization Tasks\.

Table[1](https://arxiv.org/html/2609.38445#S5.T1)comparesAIMwith solution\-driven and idea\-driven baselines across five systems\-optimization tasks\.AIMachieves the highest average score of67\.0%67\.0\\%, outperforming the strongest idea\-driven baseline, ScientistOne \(65\.4%65\.4\\%\), by1\.61\.6points and the strongest solution\-driven baseline, AdaEvolve \(63\.3%63\.3\\%\), by3\.73\.7points\. The largest gain occurs on Flash Attention, whereAIMexceeds ScientistOne and AdaEvolve by4\.34\.3and5\.25\.2points, respectively\. Overall, these results show thatAIMperforms robustly across tasks with different search characteristics\.

###### Results on Model Development & CUDA Tasks\.

In Table[2](https://arxiv.org/html/2609.38445#S5.T2), we show the results on the long\-horizon tasks, comparing with the strongest baseline configuration within each scaffold based on average performance in Table[1](https://arxiv.org/html/2609.38445#S5.T1)\. In the table,AIMachieves the highest mean score on three of the five tasks: Moving MNIST World Model, Huffman Canonical Decode, and NTT Butterfly, improving over ScientistOne by4\.24\.2,0\.40\.4, and2\.12\.1points, respectively\. On Data Selection IFEval,AIMachieves61\.861\.8, outperforming the evaluated idea\-driven baselines but falling below AdaEvolve \(82\.782\.7\)\. This strict advantage of AdaEvolve on the Data Selection IFEval task may reflect the task’s nature of a smooth, local\-search\-friendly landscape where score is dominated by fine\-tuning a single heuristic recipe \(keyword filters, length caps, source balance\) rather than by exploring qualitatively different strategies\. Its mutation loop compounds refinements against a persistent best program, which we conjecture gives it an advantage over other baselines\.

Table 2:Model Development & CUDA Task Results\. All idea\-driven approaches take the same research brief as additional context for a fair comparison\. Average scores are computed over all five tasks and omitted for methods with missing results\. Missing results with ‘–’ is because the baseline methods could not handle the Moving Mnist World Model task that requires multiple file outputs\.
###### Time Efficiency\.

Figure[5](https://arxiv.org/html/2609.38445#S5.F5)compares the mean best\-so\-far score over wall\-clock time on the Flash Attention task\.AIMreaches the final score levels of all competing baselines within approximately the first 1–2 hours, whereas the baselines require between2\.22\.2and5\.55\.5hours to attain those scores\. Moreover,AIMreaches its best score of90\.5%90\.5\\%after3\.33\.3hours, exceeding AdaEvolve’s85\.3%85\.3\\%at5\.55\.5hours and ScientistOne’s86\.2%86\.2\\%at3\.53\.5hours\. Thus,AIMnot only discovers a better solution, but also reaches competitive performance substantially earlier in the search, up to 3\.1×\\timesfaster compared to the strongest baseline, ScientistOne\. Plots for the rest of the tasks are in Appendix[D\.3](https://arxiv.org/html/2609.38445#A4.SS3)\.

Figure 4:Time Efficiency for DiscoveryFigure 5:Ablation Studies

#### 5\.2Ablation Studies

We ablate each major component ofAIMon Flash Attention, as shown in Figure[5](https://arxiv.org/html/2609.38445#S5.F5)\. Without the Agentic Surrogate, Organize/Estimate are removed, and Dispatch receives only a shared, unstructured list of ideas\. Without Agentic Acquisition, the LLM directly selects ideas from the clustered and ranked pool without explicitly assigning explore/exploit actions\. Without the Solution Auditor, all flagged solutions are retained for subsequent lessons and idea generation, idea mismatches are not reconstructed, and audit flags are ignored when computing the final score\. Finally, removing the Resource Planner replaces dynamic allocation with the fixed5×55\\times 5execution schedule\.

Removing any component reduces performance from the full model’s90\.5%90\.5\\%\. The largest degradation occurs without the Agentic Surrogate \(85\.9%85\.9\\%;−4\.6\-4\.6points\), demonstrating the importance of explicitly organizing the idea space and estimating promisingness\. The next biggest drop is observed when Solution Auditor is removed, reducing performance to87\.6%87\.6\\%\(−2\.9\-2\.9\), showing that it is critical to watch out for invalid or misattributed evidence that can propagate through later iterations\. Also, the fixed resource schedule reaches88\.1%88\.1\\%\(−2\.4\-2\.4\), while removing explicit acquisition actions yields89\.2%89\.2\\%\(−1\.3\-1\.3\)\. These results support all the central design choices ofAIM: structured idea management, grounded and adaptive idea selection, and reliable idea–solution alignment\.

#### 5\.3Qualitative Examples

Here, we provide an actual case ofAIM’s idea selection process\. In the example below, we show the example from the third iteration of the Data Select IFEval, which is a task about how to select the right data subset for LLM finetuning for a target benchmark, IFEval\([Zhou et al\., 2023](https://arxiv.org/html/2609.38445#bib.bib39)\)\.

First, the Agentic Surrogate’s Organizer partitions the 33\-idea pool into four clusters, and the Estimator ranks “metadata stratification” first, grounding on the strongest performance of its member idea \(0\.377\), while explicitly demoting “generative\-probing & LLM\-as\-a\-judge" on the basis of distilled lessons from earlier failures\. In the second part, the Agentic Acquisition’s Dispatch proceeds with determining theExplore/Exploitactions on the cluster\-level and idea\-level\. For instance, Branch 0 and 1 are both assigned the cluster\-levelExploit, but differs in the idea\-level actions\. The two branches each choose a high\-ranking idea and a low\-ranking idea accordingly\. The outcome illustrates why both actions are retained: the exploited selection regressed to 0\.119, while the explored pick matched the trailing best score \(0\.377\)\. More qualitative examples are in Appendix[E\.1](https://arxiv.org/html/2609.38445#A5.SS1)

Qualitative Example – Data Select IFEval \(iteration 3\)\(1\) Clusters and rank estimatespool = 33 ideas“Cluster 3 achieves the highest evaluated score \(0\.3770\) and aligns with strong lessons advocating for metadata distribution and length proxies\. \[…\] Clusters 0 and 1 have much weaker top scores, with lessons explicitly warning against the generative probing approaches found in Cluster 1\.”\(2\) Actions and idea selectionb0⋅\\cdotcluster \[3\]exploit⋅\\cdotideaexploit→\\rightarrowSource\-Balanced IO\-Length Stratificationrank 1/6, “A direct refinement of the best\-performing idea so far: it iterates on the successful source\-balancing strategy by incorporating output length to filter out terse responses\.”b1⋅\\cdotcluster \[3\]exploit⋅\\cdotideaexplore→\\rightarrowUnsupervised TF\-IDF \+ KMeans Stratificationrank 5/6 “To execute an explore action within this top\-performing cluster, we select a low\-rank idea \(5/6\) that introduces a completely different mechanism: data\-driven semantic boundaries instead of native ‘source’ metadata\.”Outcome:b00\.119⋅\\cdotb10\.377 \(ties best\)⋅\\cdotb20\.230

###### Following additional experiments and analyses are deferred to the Appendix:

1solver substitution using Claude Code \(Appendix[D\.1](https://arxiv.org/html/2609.38445#A4.SS1)\);2in\-depth ablation study on the Agentic Surrogate component \(Appendix[D\.2](https://arxiv.org/html/2609.38445#A4.SS2)\);3qualitative analysis of the Organize operator \(Appendix[E\.2](https://arxiv.org/html/2609.38445#A5.SS2)\);4preciseness of the Estimate operator’s ordinal predictions \(Appendix[E\.3](https://arxiv.org/html/2609.38445#A5.SS3)\);5distribution of the Dispatch operator’s exploration–exploitation actions \(Appendix[E\.4](https://arxiv.org/html/2609.38445#A5.SS4)\);6distribution of the Expand operator’s generation modes \(Appendix[E\.5](https://arxiv.org/html/2609.38445#A5.SS5)\);7effectiveness of the Solution Auditor’s idea reconstruction mechanism \(Appendix[E\.6](https://arxiv.org/html/2609.38445#A5.SS6)\)\.

### 6When to Search Over Ideas? A Formal Understanding

We have distinguished two paradigms for automated research\. Idea\-search approaches first search over semantic research directions and then instantiate selected ideas as executable solutions, whereas solution\-driven approaches directly propose and refine executable solutions\. We provide a theoretical view of when the idea\-level abstraction can be useful through two questions: \(1\)which paradigm can provide broader coverage of the solution space?, and \(2\)when is this additional coverage more valuable for finding competitive directions?

###### \(1\)Which paradigm has broader coverage of the solution space?

We begin by setting assumptions and defining the Semantic Coverage:

###### Assumption 1\(Semantic Decomposition\)\.

Let each executable solutionz∈𝒵z\\in\\mathcal\{Z\}be associated with an underlying research directionπ⁡\(z\)∈𝒳\\pi\(z\)\\in\\mathcal\{X\}, where𝒳\\mathcal\{X\}is the finite set of ideas\. Then, for each ideax∈𝒳x\\in\\mathcal\{X\}, its corresponding semantic solution region is𝒵x=\{z∈𝒵:π⁡\(z\)=x\}\\mathcal\{Z\}\_\{x\}=\\\{z\\in\\mathcal\{Z\}:\\pi\(z\)=x\\\}\.

###### Assumption 2\(Faithful Realization\)\.

For every selected ideax∈𝒳x\\in\\mathcal\{X\}, the implementation process can produce an executable solutionz=g⁡\(x\)z=g\(x\)that faithfully realizes the selected idea:π⁡\(g⁡\(x\)\)=x\\pi\(g\(x\)\)=x\.

###### Definition 1\(Semantic Coverage\)\.

For a set of evaluated solutionsℰ⊆𝒵\\mathcal\{E\}\\subseteq\\mathcal\{Z\}, define its Semantic Coverage as the number of distinct research directions represented in it:C⁡\(ℰ\)=\|\{π⁡\(z\):z∈ℰ\}\|C\(\\mathcal\{E\}\)=\\left\|\\\{\\pi\(z\):z\\in\\mathcal\{E\}\\\}\\right\|\.

Assumption[1](https://arxiv.org/html/2609.38445#Thmassumption1)and Definition[1](https://arxiv.org/html/2609.38445#Thmdefinition1)provides a useful way to view the difference between the two search paradigms\. For instance, consider a search graph whose nodes are executable solutions\. A solution\-driven search trajectory may move throughz1→z2→z3→z4z\_\{1\}\\rightarrow z\_\{2\}\\rightarrow z\_\{3\}\\rightarrow z\_\{4\}\(zi∈𝒵\)\(z\_\{i\}\\in\\mathcal\{Z\}\), while the corresponding semantic directions arex1→x1→x1→x2x\_\{1\}\\rightarrow x\_\{1\}\\rightarrow x\_\{1\}\\rightarrow x\_\{2\}\(xi∈𝒳\)\(x\_\{i\}\\in\\mathcal\{X\}\)\. In such cases, although four distinct solutions have been evaluated, this trajectory covers only two semantic directions\. More generally, the mappingπ:𝒵→𝒳\\pi:\\mathcal\{Z\}\\rightarrow\\mathcal\{X\}groups solutions according to their underlying research direction\. Formally,π\\piinduces the equivalence relationz∼z′z\\sim z^\{\\prime\}if and only ifπ⁡\(z\)=π⁡\(z′\)\\pi\(z\)=\\pi\(z^\{\\prime\}\), thereby contracting solutions that instantiate the same idea into a common semantic class\. This yields a quotient view of the solution space, in which solution\-level trajectories are represented by the distinct semantic directions they traverse\.

In addition, while Assumption[2](https://arxiv.org/html/2609.38445#Thmassumption2)may seem a bit idealistic, we argue that the Solution Auditor ensures the evidence that the agent observes is faithful and reliable\. Figure[19](https://arxiv.org/html/2609.38445#A5.F19)in Appendix[E\.6](https://arxiv.org/html/2609.38445#A5.SS6)supports this view, in that it substantially reduces the idea\-mismatch cases after the Solution Auditor’s idea reconstruction takes place\. Furthermore, our discussion on the practical considerations in Appendix[6\.2](https://arxiv.org/html/2609.38445#S6.SS2)suggests that a hybrid design of solution\-driven and idea\-driven approaches might be desired to ensure faithful implementation of ideas\. Based on these grounds, we formalize the attainable semantic coverage of the two search paradigms:

###### Proposition 1\(Attainable Semantic Coverage\)\.

LetCideaC^\{\\text\{idea\}\}be the semantic coverage of an extreme idea\-driven search procedure that rendersNNdistinct experiment executions toNNdistinct research ideas, andCCbe the semantic coverage of any search procedure that evaluates at mostNNexecutable solutions, including solution\-driven approaches\. Then,C≤CideaC\\leq C^\{\\text\{idea\}\}\.

Note that Proposition[1](https://arxiv.org/html/2609.38445#Thmproposition1)does not imply that solution\-driven search must have lower semantic coverage\. A solution\-driven method can attain the same bound if its code\-level search also produces semantically distinct solutions\. The distinction is that idea\-driven search makes semantic breadth an explicit and directly controllable property of the search process\. By selecting distinct ideas before implementation, the framework can deliberately allocate its expensive experiment budget across distinct research directions, rather than obtaining semantic diversity only indirectly through solution\-level transitions\. Whether maximizing such semantic breadth is beneficial, however, depends on the trait of the research task, which we discuss next\.

###### \(2\)When is broader semantic coverage useful?

We next ask when the additional coverage is useful for solving the research task\. Intuitively, breadth should matter most when there are many possible research directions but only a small subset of them can achieve near\-optimal performance\.

###### Definition 2\(Competitive Semantic Directions\)\.

For a target taskτ\\tau, let there beKKmaterially distinct research ideas𝒳τ=\{x1,…,xK\}⊆𝒳\\mathcal\{X\}\_\{\\tau\}=\\\{x\_\{1\},\\ldots,x\_\{K\}\\\}\\subseteq\\mathcal\{X\}\. Each direction has an attainable value \(e\.g\., performance measure\),V⁡\(x\)=maxz∈𝒵x⁡v⁡\(z\)V\(x\)=\\max\_\{z\\in\\mathcal\{Z\}\_\{x\}\}v\(z\), and letV⋆=maxx∈𝒳τ⁡V⁡\(x\)V^\{\\star\}=\\max\_\{x\\in\\mathcal\{X\}\_\{\\tau\}\}V\(x\)\. Forε≥0\\varepsilon\\geq 0, define the set ofε\\varepsilon\-optimal directions as𝒢ε=\{x∈𝒳τ:V⁡\(x\)≥V⋆−ε\}\\mathcal\{G\}\_\{\\varepsilon\}=\\left\\\{x\\in\\mathcal\{X\}\_\{\\tau\}:V\(x\)\\geq V^\{\\star\}\-\\varepsilon\\right\\\}, andGε=\|𝒢ε\|G\_\{\\varepsilon\}=\|\\mathcal\{G\}\_\{\\varepsilon\}\|\.

###### Assumption 3\(Competitive Direction Exchangeability\)\.

Before the search evaluates a direction, the identities of theGεG\_\{\\varepsilon\}ε\\varepsilon\-optimal directions are uniformly distributed among theKKadmissible directions\. Conditional on having evaluated only non\-ε\\varepsilon\-optimal directions, the optimal directions remain exchangeable among the unexplored directions\.

###### Proposition 2\(Semantic Coverage Sufficient Condition\)\.

Under Definition[2](https://arxiv.org/html/2609.38445#Thmdefinition2)and Assumption[3](https://arxiv.org/html/2609.38445#Thmassumption3), a search procedure with semantic coverageCCdiscovers at least oneε\\varepsilon\-optimal direction with probabilityPε​\(C\)=1−\(K−GεC\)/\(KC\)\.P\_\{\\varepsilon\}\(C\)=1\-\\binom\{K\-G\_\{\\varepsilon\}\}\{C\}/\\binom\{K\}\{C\}\.Then, a sufficient condition of the semantic coverageC≥K/Gε​ln⁡\(1/δ\)C\\geq K/G\_\{\\varepsilon\}\\ln\\left\(1/\\delta\\right\)guarantees a success probability of1−δ1\-\\delta, if the right\-hand side is feasible\.

###### Implication: Tasks with larger effective semantic breadth favor idea\-driven search\.

Combining Propositions[1](https://arxiv.org/html/2609.38445#Thmproposition1)and[2](https://arxiv.org/html/2609.38445#Thmproposition2), tasks with many materially distinct research directions but relatively few competitive ones benefit more from broad semantic coverage\. In such tasks, idea\-driven search can be advantageous because it makes semantic breadth explicit and controllable\. Conversely, whenK/GεK/G\_\{\\varepsilon\}is small,i\.e\., the task admits only a few meaningful directions or many directions are similarly competitive, the value of additional semantic coverage is limited, and solution\-driven and idea\-driven approaches may perform similarly, leaving implementation\-level optimization as a potentially more important determinant of performance\. Proofs are in Appendix[F\.2](https://arxiv.org/html/2609.38445#A6.SS2)\.

115510101515202000101020203030404050506060Effective Semantic BreadthBε=K/GεB\_\{\\varepsilon\}=K/G\_\{\\varepsilon\}Required Semantic CoverageCCPε≥0\.50P\_\{\\varepsilon\}\\geq 0\.50Pε≥0\.90P\_\{\\varepsilon\}\\geq 0\.90Pε≥0\.95P\_\{\\varepsilon\}\\geq 0\.95Figure 6:Semantic coverage required to achieve a fixed probability of discovering anε\\varepsilon\-optimal research direction\. As competitive directions become sparser, the effective semantic breadthBε=K/GεB\_\{\\varepsilon\}=K/G\_\{\\varepsilon\}increases, requiring proportionally broader semantic coverage\. The curves show the sufficient conditionC≥Bε​log⁡\(1/δ\)C\\geq B\_\{\\varepsilon\}\\log\(1/\\delta\)for different target success probabilities1−δ1\-\\delta\.
###### Theoretical Bound Visualization\.

Figure[6](https://arxiv.org/html/2609.38445#S6.F6)illustrates the semantic coverage requirement derived in Proposition[2](https://arxiv.org/html/2609.38445#Thmproposition2)\. LetBε=K/GεB\_\{\\varepsilon\}=K/G\_\{\\varepsilon\}denote the effective semantic breadth, whereKKis the number of admissible semantic directions andGεG\_\{\\varepsilon\}is the number ofε\\varepsilon\-optimal directions\. For a target success probability1−δ1\-\\delta, the sufficient coverage condition can be written as

C≥Bε​ln⁡\(1δ\)\.C\\geq B\_\{\\varepsilon\}\\ln\\left\(\\frac\{1\}\{\\delta\}\\right\)\.\(12\)The required semantic coverage therefore grows linearly with effective semantic breadth\. When competitive directions are sparse, such thatGεG\_\{\\varepsilon\}is small relative toKK, a search procedure must examine substantially more distinct directions to maintain the same probability of discovering anε\\varepsilon\-optimal one\. The requirement also becomes steeper as the target success probability increases: atBε=20B\_\{\\varepsilon\}=20, for example, the sufficient coverage is approximately1414,4646, and6060directions for target probabilities of0\.500\.50,0\.900\.90, and0\.950\.95, respectively\. Thus, broader semantic coverage is particularly valuable for tasks with sparse competitive directions and when a high probability of success is required\.

#### 6\.1Empirical Support

Our empirical analysis in Figure[7](https://arxiv.org/html/2609.38445#S6.F7)is consistent with the implication\. As a rough proxy for the breadth of the explored solution space, we embed all intermediate solutions from each method using Gemini\-Embedding\-001\([Lee et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib31)\), project the embeddings into a two\-dimensional PCA space, and measure the area of their convex hull\. The box\-and\-whisker plots compare the distribution of the hull area, and the PCA hull plots visualize example cases\. For tasks like Flash Attention, which admit diverse algorithmic directions but only a small subset of them yield substantial speedups, idea\-driven search approaches explore substantially broader regions than solution\-driven approaches, as reflected by larger convex\-hull areas\. In contrast, for tasks like Radix Sort, where the high\-level algorithm is largely fixed and progress depends primarily on code\-level optimization, the gap in convex\-hull area between the two paradigms becomes considerably smaller\.

Figure 7:Idea\-driven approaches generally show broader coverage of solutions in the embedding space\. Embedding cosine similarity\-based supplementary analysis is in Appendix[F\.5](https://arxiv.org/html/2609.38445#A6.SS5)\. Plots for all tasks are in Appendix[F\.3](https://arxiv.org/html/2609.38445#A6.SS3)and Appendix[F\.4](https://arxiv.org/html/2609.38445#A6.SS4)\.
#### 6\.2Practical Considerations

Our theoretical analysis isolates the benefit of semantic breadth by assuming*faithful realization*\(Assumption[2](https://arxiv.org/html/2609.38445#Thmassumption2)\): a selected idea can be instantiated as a solution that faithfully reflects its intended research direction\. This assumption is useful for understanding the coverage advantage of idea\-driven search, but it could be slightly idealized\. In practice, solver agents may produce incomplete or incorrect implementations, or realize a solution that only partially matches the intended idea\. Even when the implementation is faithful, some tasks inherently require substantial solution\-level refinement before the value of a research direction can be assessed\. For example, machine learning systems may require hyperparameter tuning, while systems optimization tasks may depend on low\-level implementation choices that cannot be fully specified at the idea level\.

These considerations suggest that practical research agents lie on a continuum between purely idea\-driven and purely solution\-driven search\. At one extreme, allocating each experiment execution to a new idea maximizes semantic breadth, but may under\-invest in realizing and optimizing each direction\. At the other extreme, repeatedly refining the same implementation can exploit fine\-grained execution feedback, but reduces the budget available for exploring alternative research directions\. The appropriate balance therefore depends on the task: problems with many meaningfully different high\-level approaches may benefit more from semantic exploration, whereas tasks with relatively fixed strategies but substantial implementation\-level optimization may benefit more from solution refinement\.

Our framework, to be precise, adopts a hybrid design\. Although search decisions are made over explicit research ideas, each solver branch is allowed to iteratively refine its implementation using execution and verifier feedback\. Thus, the framework retains the semantic organization of idea\-driven search while still incorporating the local optimization behavior characteristic of solution\-driven approaches\. This also motivates our idea–solution auditing mechanism, which explicitly checks whether the final implementation remains aligned with the selected idea before its evaluation is attributed back to the idea\-level search process\. More broadly, the optimal balance between semantic exploration and implementation refinement is likely task\-dependent\. An important direction for future work is therefore to adapt this balance dynamically based on observed implementation difficulty, solver fidelity, and the semantic structure of the task, rather than fixing the degree of idea\-driven versus solution\-driven behavior in advance\.

### 7Conclusion

In this work, we introducedAIM, a fully autonomous framework for managing and searching research ideas in idea\-driven automated research\.AIMdynamically organizes an evolving idea pool, explicitly estimates the promise of research directions, adaptively balances exploration and exploitation, and audits implementations before incorporating their outcomes into subsequent search\. Across 10 AutoLab tasks,AIMachieves the best overall performance, while reaching strong solutions substantially faster compared to baselines\. Our theoretical analysis further clarifies when idea\-driven search can be advantageous: explicit idea\-level allocation makes semantic coverage controllable, and broader coverage becomes more valuable when competitive directions are sparse among many plausible alternatives\. Together, these results establish agentic idea management as a transparent and robust approach to budget\-constrained automated research\.

### AI Use Statement

In this work, we used generative AI tools for polishing the writings and generating plots\. We have not used generative AI tools to run the experiments for the study outside the agent methods being evaluated\. We have reviewed all AI\-assisted work\. We manually checked the writing faithfully conveys our intended content, and if the plots correctly express values\. We take responsibility for the final content of this work, including text, claims or artifacts produced with the aid of generative AI\.

### Ethics Statement

This work does not involve human subjects, personal data, or user\-facing deployment\. Nevertheless, autonomous research agents generate and execute code, which may introduce risks such as insecure implementations, verifier exploitation, or unintended task behavior\. We mitigate these risks by conducting all experiments in isolated, containerized benchmark environments with fixed verifiers and by using the Solution Auditor to identify task mismatches and reward\-hacking behavior\. No generated solutions are deployed in real\-world systems\. These safeguards cannot eliminate all risks, and human review remains necessary before applying autonomous research systems to consequential scientific or engineering settings\.

### Reproducibility Statement

We provide pseudocode for the completeAIMpipeline in Algorithm[1](https://arxiv.org/html/2609.38445#alg1)\. Appendix[C](https://arxiv.org/html/2609.38445#A3)describes the evaluated benchmarks, baseline configurations, scoring procedure, execution and wall\-clock budgets, branch\-wise resource allocation, and implementation hyperparameters\. The prompt templates for all LLM\-driven operators are also included in the appendix\. We report results over three independent runs using a common backbone model and evaluation environment, with means and standard errors\.

### References

- Agarwalet al\.\(2026\)D\. Agarwal, B\. P\. Majumder, R\. Adamson, M\. Chakravorty, S\. R\. Gavireddy, A\. Parashar, H\. Surana, B\. Dalvi Mishra, A\. McCallum, A\. Sabharwal,et al\.Autodiscovery: open\-ended scientific discovery via bayesian surprise\.Advances in Neural Information Processing Systems38,pp\. 25181–25219\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p1.1),[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p2.1)\.
- \[2\]Claude codeNote:Computer software\. Available via Anthropic’s developer platform\.External Links:[Link](https://claude.ai/)Cited by:[§D\.1](https://arxiv.org/html/2609.38445#A4.SS1.p1.1),[footnote 1](https://arxiv.org/html/2609.38445#footnote1)\.
- Baeket al\.\(2025\)J\. Baek, S\. K\. Jauhar, S\. Cucerzan, and S\. J\. HwangResearchagent: iterative research idea generation over scientific literature with large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Albuquerque, New Mexico,pp\. 6709–6738\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.342),[Link](https://aclanthology.org/2025.naacl-long.342/)Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Cemriet al\.\(2026\)M\. Cemri, S\. Agrawal, A\. Gupta, S\. Liu, A\. Cheng, Q\. Mang, A\. Naren, L\. E\. Erdogan, K\. Sen, M\. Zaharia,et al\.Adaevolve: adaptive llm driven zeroth\-order optimization\.arXiv preprint arXiv:2602\.20133\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.38445#S4.SS2.p3.2),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Chanet al\.\(2025\)J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace, K\. Liu, L\. Maksin, T\. Patwardhan,et al\.Mle\-bench: evaluating machine learning agents on machine learning engineering\.InInternational Conference on Learning Representations,Vol\.2025,pp\. 50466–50494\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Chenet al\.\(2026a\)H\. Chen, M\. Xiong, Y\. Lu, W\. Han, A\. Deng, Y\. He, J\. Wu, Y\. Li, Y\. Liu, and B\. HooiMlr\-bench: evaluating ai agents on open\-ended machine learning research\.Advances in Neural Information Processing Systems38\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Chenet al\.\(2026b\)J\. Chen, B\. D\. Mishra, J\. Nam, R\. Meng, T\. Pfister, and J\. YoonMARS: modular agent with reflective search for automated ai research\.InInternational Conference on Machine Learning,Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p1.1)\.
- Choiet al\.\(2026\)H\. K\. Choi, J\. Li, W\. Li, X\. E\. Wang, and S\. LiMulti\-agent llms fail to explore each other\.arXiv preprint arXiv:2607\.11250\.Cited by:[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1)\.
- Fosteret al\.\(2026\)T\. S\. Foster, B\. A\. Omari, T\. Fu, T\. Mann, C\. Domond, L\. Cipolina\-Kun, B\. Gauri, M\. Aghamelu, A\. D\. Goldie, E\. Helenowski,et al\.AI research preference models\.arXiv preprint arXiv:2608\.13940\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1)\.
- Garikaparthiet al\.\(2026\)A\. Garikaparthi, M\. Patwardhan, and A\. CohanResearchGym: evaluating language model agents on real\-world ai research\.arXiv preprint arXiv:2602\.15112\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Huanget al\.\(2024\)Q\. Huang, J\. Vora, P\. Liang, and J\. LeskovecMLAgentBench: evaluating language agents on machine learning experimentation\.InInternational Conference on Machine Learning,pp\. 20271–20309\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Jansenet al\.\(2025\)P\. Jansen, O\. Tafjord, M\. Radensky, P\. Siangliulue, T\. Hope, B\. D\. Mishra, B\. P\. Majumder, D\. S\. Weld, and P\. ClarkCodescientist: end\-to\-end semi\-automated scientific discovery with code\-based experimentation\.arXiv preprint arXiv:2503\.22708\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Jianget al\.\(2025\)Z\. Jiang, D\. Schmidt, D\. Srikanth, D\. Xu, I\. Kaplan, D\. Jacenko, and Y\. WuAide: ai\-driven exploration in the space of code\.arXiv preprint arXiv:2502\.13138\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1)\.
- Jinet al\.\(2026\)J\. Jin, Y\. Hu, K\. Qiu, Q\. Dai, C\. Luo, G\. Dong, X\. Li, T\. Zhao, X\. Ma, G\. Zhang,et al\.Toward generalist autonomous research via hypothesis\-tree refinement\.arXiv preprint arXiv:2606\.11926\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p1.1),[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p2.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p2.1),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Krishnamurthyet al\.\(2024\)A\. Krishnamurthy, K\. Harris, D\. J\. Foster, C\. Zhang, and A\. SlivkinsCan large language models explore in\-context?\.Advances in Neural Information Processing Systems37,pp\. 120124–120158\.Cited by:[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1)\.
- Leeet al\.\(2025\)J\. Lee, F\. Chen, S\. Dua, D\. Cer, M\. Shanbhogue, I\. Naim, G\. H\. Ábrego, Z\. Li, K\. Chen, H\. S\. Vera,et al\.Gemini embedding: generalizable embeddings from gemini\.arXiv preprint arXiv:2503\.07891\.Cited by:[§D\.2](https://arxiv.org/html/2609.38445#A4.SS2.p1.1),[§6\.1](https://arxiv.org/html/2609.38445#S6.SS1.p1.1)\.
- Liet al\.\(2024\)R\. Li, T\. Patel, Q\. Wang, and X\. DuMlr\-copilot: autonomous machine learning research based on large language models agents\.arXiv preprint arXiv:2408\.14033\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2026\)S\. Liu, S\. Agarwal, M\. Maheswaran, M\. Cemri, Z\. Li, Q\. Mang, A\. Naren, E\. Boneh, A\. Cheng, M\. Z\. Pan,et al\.Evox: meta\-evolution for automated discovery\.arXiv preprint arXiv:2602\.23413\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.38445#S4.SS2.p3.2),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Luet al\.\(2024\)C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. HaThe ai scientist: towards fully automated open\-ended scientific discovery\.arXiv preprint arXiv:2408\.06292\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p1.1)\.
- Menget al\.\(2026\)R\. Meng, B\. D\. Mishra, J\. Chen, C\. Li, P\. Goyal, M\. Parmar, Y\. Song, Y\. Song, R\. Sinha, P\. Ranganathan,et al\.ScientistOne: towards human\-level autonomous research via chain\-of\-evidence\.arXiv preprint arXiv:2605\.26340\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p1.1),[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p2.1),[§C\.2](https://arxiv.org/html/2609.38445#A3.SS2.SSS0.Px2.p1.1),[§D\.1](https://arxiv.org/html/2609.38445#A4.SS1.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§1](https://arxiv.org/html/2609.38445#S1.p4.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.38445#S4.SS2.p2.2),[§4](https://arxiv.org/html/2609.38445#S4.p1.2),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Namet al\.\(2025\)J\. Nam, J\. Yoon, J\. Chen, and T\. PfisterDS\-star: data science agent via iterative planning and verification\.arXiv preprint arXiv:2509\.21825\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1)\.
- Namet al\.\(2026\)J\. Nam, J\. Yoon, J\. Chen, J\. Shin, S\. Arik, and T\. PfisterMle\-star: machine learning engineering agent via search and targeted refinement\.Advances in Neural Information Processing Systems38,pp\. 116692–116712\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1)\.
- Nathaniet al\.\(2025\)D\. Nathani, L\. Madaan, N\. Roberts, N\. Bashlykov, A\. Menon, V\. Moens, M\. Plekhanov, A\. Budhiraja, D\. Magka, V\. Vorotilov,et al\.Mlgym: a new framework and benchmark for advancing ai research agents\.InSecond Conference on Language Modeling,Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Novikovet al\.\(2025\)A\. Novikov, N\. Vũ, M\. Eisenberger, E\. Dupont, P\. Huang, A\. Z\. Wagner, S\. Shirobokov, B\. Kozlovskii, F\. J\. Ruiz, A\. Mehrabian,et al\.Alphaevolve: a coding agent for scientific and algorithmic discovery\.arXiv preprint arXiv:2506\.13131\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1)\.
- Panet al\.\(2026\)L\. Pan, H\. Xie, and R\. WilsonLarge language models think too fast to explore effectively\.Advances in Neural Information Processing Systems38,pp\. 102480–102508\.Cited by:[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1)\.
- Parket al\.\(2026\)J\. Park, J\. Kim, J\. Jeong, R\. D\. Nowak, K\. Lee, and Y\. J\. LeeExploration and exploitation errors are measurable for language model agents\.arXiv preprint arXiv:2604\.13151\.Cited by:[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1)\.
- Schmidgallet al\.\(2025\)S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. BarsoumAgent laboratory: using llm agents as research assistants\.Findings of the Association for Computational Linguistics: EMNLP 2025,pp\. 5977–6043\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Siet al\.\(2025\)C\. Si, D\. Yang, and T\. HashimotoCan llms generate novel research ideas? a large\-scale human study with 100\+ nlp researchers\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=M23dTGWCZy)Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Tanget al\.\(2026\)J\. Tang, L\. Xia, Z\. Li, and C\. HuangAi\-researcher: autonomous scientific innovation\.Advances in Neural Information Processing Systems38,pp\. 9481–9520\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Teamet al\.\(2023\)G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican,et al\.Gemini: a family of highly capable multimodal models\.arXiv preprint arXiv:2312\.11805\.Cited by:[§C\.1](https://arxiv.org/html/2609.38445#A3.SS1.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Teamet al\.\(2025\)I\. Team, B\. Zhang, S\. Feng, X\. Yan, J\. Yuan, R\. Ma, Y\. Hu, Z\. Yu, X\. He, S\. Huang,et al\.InternAgent: when agent becomes the scientist–building closed\-loop system from hypothesis to verification\.arXiv preprint arXiv:2505\.16938\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Toledoet al\.\(2026\)E\. Toledo, K\. Hambardzumyan, M\. Josifoski, R\. Hazra, N\. Baldwin, A\. Audran\-Reiss, M\. Kuchnik, D\. Magka, M\. Jiang, A\. Lupidi,et al\.Ai research agents for machine learning: search, exploration, and generalization in mle\-bench\.Advances in Neural Information Processing Systems38,pp\. 35309–35348\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px2.p1.1),[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px1.p1.1),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Vasuet al\.\(2025\)R\. Vasu, P\. Jansen, P\. Siangliulue, C\. Sarasua, A\. Bernstein, P\. Clark, and B\. D\. MishraHARPA: a testability\-driven, literature\-grounded framework for research ideation\.arXiv preprint arXiv:2510\.00620\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px1.p1.1)\.
- Wenget al\.\(2026\)Y\. Weng, M\. Zhu, Q\. Xie, Q\. Sun, Z\. Lin, S\. Liu, and Y\. ZhangDeepscientist: advancing frontier\-pushing scientific findings progressively\.InInternational Conference on Learning Representations,Vol\.2026,pp\. 47981–48037\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p1.1),[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p2.1),[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p2.1),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Wijket al\.\(2025\)H\. Wijk, T\. R\. Lin, J\. Becker, S\. Jawhar, N\. Parikh, T\. Broadley, L\. Chan, M\. Chen, J\. M\. Clymer, J\. Dhyani,et al\.RE\-bench: evaluating frontier ai r&d capabilities of language model agents against human experts\.InInternational Conference on Machine Learning,pp\. 66772–66832\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Xuet al\.\(2026\)Z\. Xu, J\. Chen, Y\. Huang, D\. Jiang, J\. Chen, H\. Hua, Z\. Wu, Z\. Liu, Z\. He, L\. Li,et al\.AutoLab: can frontier models solve long\-horizon auto research and engineering tasks?\.arXiv preprint arXiv:2606\.05080\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1),[§C\.1](https://arxiv.org/html/2609.38445#A3.SS1.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.38445#S5.SS0.SSS0.Px1.p1.1)\.
- Yamadaet al\.\(2025\)Y\. Yamada, R\. T\. Lange, C\. Lu, S\. Hu, C\. Lu, J\. Foerster, J\. Clune, and D\. HaThe ai scientist\-v2: workshop\-level automated scientific discovery via agentic tree search\.arXiv preprint arXiv:2504\.08066\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p1.1),[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px3.p2.1),[§E\.4](https://arxiv.org/html/2609.38445#A5.SS4.p1.1),[§1](https://arxiv.org/html/2609.38445#S1.p2.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p1.1),[§2](https://arxiv.org/html/2609.38445#S2.SS0.SSS0.Px2.p2.1),[§5](https://arxiv.org/html/2609.38445#S5.p1.1)\.
- Zhanget al\.\(2026\)Y\. Zhang, M\. Khalifa, S\. Bhushan, G\. Murphy, L\. Logeswaran, J\. Kim, M\. Lee, H\. Lee, and L\. WangMLRC\-bench: can language agents solve machine learning research challenges?\.Advances in Neural Information Processing Systems38\.Cited by:[Appendix A](https://arxiv.org/html/2609.38445#A1.SS0.SSS0.Px4.p1.1)\.
- Zhouet al\.\(2023\)J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. HouInstruction\-following evaluation for large language models\.arXiv preprint arXiv:2311\.07911\.Cited by:[§5\.3](https://arxiv.org/html/2609.38445#S5.SS3.p1.1)\.

## PartAppendix

### Appendix ARelated Works

###### Automated Research with LLM Agents\.

Recent work has expanded LLM agents from supporting individual scientific tasks to conducting increasingly complete research workflows\. At the ideation stage, ResearchAgent\[[Baek et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib4)\]grounds research idea generation in retrieved literature and iterative feedback, while HARPA\[[Vasu et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib18)\]develops literature\-grounded, testable hypotheses and refines them using experimental evidence\. Complementarily,[Si et al\. \[2025\]](https://arxiv.org/html/2609.38445#bib.bib5)conduct a large\-scale expert study of the novelty and quality of LLM\-generated research ideas\. Broader research systems automate multiple stages of the scientific process: CodeScientist\[[Jansen et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib17)\]develops a coding\-based framework for semi\-automated scientific discovery; The AI Scientist\[[Lu et al\., 2024](https://arxiv.org/html/2609.38445#bib.bib1)\]spans ideation, experimentation, paper writing, and review; and MLR\-Copilot\[[Li et al\., 2024](https://arxiv.org/html/2609.38445#bib.bib6)\], Agent Laboratory\[[Schmidgall et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib3)\], AI\-Researcher\[[Tang et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib7)\], and InternAgent\[[Team et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib8)\]similarly automate multi\-stage workflows from hypothesis formation to experimental validation and reporting\. These works establish the broader setting of automated research, including the end\-to\-end pipeline resulting in a publishable paper, while our focus is specifically on the “discovery" stage of automated research:*how an agent should search under a limited experimental budget*\.

###### Solution\-driven Search over Executable Code\.

A prominent class of research agents directly searches over executable artifacts, keeping implementation and search tightly coupled\. AIDE\[[Jiang et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib23)\]searches directly in the space of code, while AlphaEvolve\[[Novikov et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib15)\]uses evolutionary code generation for algorithmic discovery\. MLE\-Star\[[Nam et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib19)\]and DS\-Star\[[Nam et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib20)\]iteratively improve machine\-learning and data\-science solutions through targeted refinement, planning, and verification\. More recent systems further develop solution\-level search through evolutionary and reflective mechanisms: AdaEvolve\[[Cemri et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib26)\]performs adaptive LLM\-driven zeroth\-order optimization, EvoX\[[Liu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib27)\]introduces meta\-evolution for automated discovery, AIDE\[[Jiang et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib23)\]utilize an MCTS approach for machine learning tasks, AIRA\[[Toledo et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib22)\]studies evolutionary and tree\-search strategies for machine\-learning research, and RPM\[[Foster et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib38)\]introduces a preference model for candidate solutions\. These methods exemplify the*solution\-driven*paradigm considered in our work: the executable solution itself remains the primary search object, allowing verifier feedback to be applied directly to subsequent implementation\-level refinement\.

###### Idea\-driven Search over Research Ideas and Hypotheses\.

A complementary line of work maintains ideas or hypotheses as explicit intermediate objects before delegating their realization to an implementation agent\. The AI Scientist\-v2\[[Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2)\]uses agentic tree search to progressively develop research directions and experiments\. DeepScientist\[[Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\]maintains candidate research directions and progressively selects them using a structured exploration mechanism, while AutoDiscovery\[[Agarwal et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib32)\]conducts data\-driven open\-ended scientific discovery through search guided by Bayesian surprise\. Arbor\[[Jin et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib28)\]explicitly organizes hypotheses in a refinement tree, separating hypothesis development from downstream implementation\. MARS\[[Chen et al\., 2026b](https://arxiv.org/html/2609.38445#bib.bib21)\]performs modular reflective search over candidate ideas and solutions\. ScientistOne\[[Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\]similarly makes research ideas explicit and uses an elite\-preserving search procedure within its Chain\-of\-Evidence framework, while additionally preserving traceability between scientific claims, experimental evidence, and implementations\.

These*idea\-driven*systems reveal several design choices for managing the intermediate idea space\. Existing approaches employ structures such as lists\[[Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\], search trees\[[Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2),[Jin et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib28),[Agarwal et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib32)\], or beam\-style candidate retention\[[Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\], together with predefined selection mechanisms such as UCB or tree search\. In contrast, our work focuses on making*idea management itself*an autonomous, evidence\-conditioned component of the research process\. AIM dynamically reorganizes the evolving idea pool, explicitly estimates the relative promise of research directions, and adaptively determines where to explore or exploit based on accumulated experimental evidence\. Moreover, because separating ideation from implementation introduces the possibility that a solver may not faithfully realize its intended idea, AIM explicitly audits and reconciles idea–solution mismatches before incorporating results into subsequent search\. Thus, our contribution is complementary to prior idea\-driven research agents: rather than introducing another fixed idea\-search scaffold, we study how the idea space can be autonomously organized, searched, and maintained as evidence accumulates\.

###### Testbeds for Automated Research\.

In parallel, a growing body of work has made automated experimentation measurable\. MLAgentBench\[[Huang et al\., 2024](https://arxiv.org/html/2609.38445#bib.bib9)\]evaluates agents that modify machine\-learning code and iteratively interpret experimental feedback, while MLE\-bench\[[Chan et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib10)\]evaluates agents on competition\-style machine\-learning engineering\. RE\-Bench\[[Wijk et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib11)\]studies frontier AI research\-and\-development capabilities relative to human experts, and MLGym\[[Nathani et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib12)\]provides diverse environments for evaluating AI research agents\. More recent benchmarks broaden evaluation toward open\-ended research problems, including MLRC\-Bench\[[Zhang et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib24)\], MLR\-Bench\[[Chen et al\., 2026a](https://arxiv.org/html/2609.38445#bib.bib13)\], and ResearchGym\[[Garikaparthi et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib14)\]\. We primarily evaluate on AutoLab\[[Xu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib25)\], which provides long\-horizon research and engineering tasks with executable verifiers and enables controlled comparison of search strategies in systems optimization\.

Overall, prior work has made substantial progress in both direct solution optimization and end\-to\-end automated research\. Our work builds on this literature by explicitly distinguishing*solution\-driven*search over executable solutions from*idea\-driven*search over ideas followed by implementation, and focuses on the latter’s distinctive challenges: maintaining an evolving semantic representation of candidate ideas, making evidence\-grounded decisions about which ideas deserve expensive implementation, and preserving idea–solution integrity throughout the search process\.

### Appendix BAgentic Idea Manager Algorithm

Algorithm 1AIM: Agentic Idea Manager1:Task context

𝒯\\mathcal\{T\}; initial idea pool

𝒫1\\mathcal\{P\}\_\{1\}; execution budget

NN; total branches

BtotB\_\{\\mathrm\{tot\}\}; maximum parallelism

BmaxB\_\{\\max\}; time limit

HH\.

2:

𝒟1,ℳ1←∅\\mathcal\{D\}\_\{1\},\\mathcal\{M\}\_\{1\}\\leftarrow\\varnothing;

b←0b\\leftarrow 0;

t←1t\\leftarrow 1;

Nbranch←⌊N/Btot⌋N\_\{\\mathrm\{branch\}\}\\leftarrow\\lfloor N/B\_\{\\mathrm\{tot\}\}\\rfloor
3:while

b<Btotb<B\_\{\\mathrm\{tot\}\}andelapsed time

<H<Hdo

4:

Bt←\{5,t=1,ϕplan​\(𝒮t,Btot−b,Bmax\),t\>1B\_\{t\}\\leftarrow\\begin\{cases\}5,&t=1,\\\\ \\phi\_\{\\mathrm\{plan\}\}\(\\mathcal\{S\}\_\{t\},B\_\{\\mathrm\{tot\}\}\-b,B\_\{\\max\}\),&t\>1\\end\{cases\}⊳\\trianglerightResource Planner

5:

𝒞t←ϕorg​\(𝒯,𝒫t,𝒟t,𝒞t−1\)\\mathcal\{C\}\_\{t\}\\leftarrow\\phi\_\{\\mathrm\{org\}\}\(\\mathcal\{T\},\\mathcal\{P\}\_\{t\},\\mathcal\{D\}\_\{t\},\\mathcal\{C\}\_\{t\-1\}\)⊳\\trianglerightAgentic Surrogate: Organize

6:

Rt←ϕest​\(𝒞t,𝒟t\)R\_\{t\}\\leftarrow\\phi\_\{\\mathrm\{est\}\}\(\\mathcal\{C\}\_\{t\},\\mathcal\{D\}\_\{t\}\)⊳\\trianglerightAgentic Surrogate: Estimate

7:

𝒬t←ϕdisp​\(𝒫t,𝒞t,Rt,Bt\)\\mathcal\{Q\}\_\{t\}\\leftarrow\\phi\_\{\\mathrm\{disp\}\}\(\\mathcal\{P\}\_\{t\},\\mathcal\{C\}\_\{t\},R\_\{t\},B\_\{t\}\)⊳\\trianglerightAgentic Acquisition: Dispatch

8:for all

x∈𝒬tx\\in\\mathcal\{Q\}\_\{t\}in paralleldo

9:

\(zx,yx,hx\)←ϕsolve​\(𝒯,x,Nbranch\)\(z\_\{x\},y\_\{x\},h\_\{x\}\)\\leftarrow\\phi\_\{\\mathrm\{solve\}\}\(\\mathcal\{T\},x;N\_\{\\mathrm\{branch\}\}\)
10:endfor

11:

𝒱t←ϕaudit​\(\{\(x,zx,yx,hx\):x∈𝒬t\}\)\\mathcal\{V\}\_\{t\}\\leftarrow\\phi\_\{\\mathrm\{audit\}\}\\bigl\(\\\{\(x,z\_\{x\},y\_\{x\},h\_\{x\}\):x\\in\\mathcal\{Q\}\_\{t\}\\\}\\bigr\)⊳\\trianglerightSolution Auditor

12:

𝒟t\+1←𝒟t∪\{\(x\+,y\):\(x\+,z,y,h\)∈𝒱t\}\\mathcal\{D\}\_\{t\+1\}\\leftarrow\\mathcal\{D\}\_\{t\}\\cup\\\{\(x^\{\+\},y\):\(x^\{\+\},z,y,h\)\\in\\mathcal\{V\}\_\{t\}\\\}
13:

ℳt\+1←ℳt∪ϕlesson​\(𝒱t\)\\mathcal\{M\}\_\{t\+1\}\\leftarrow\\mathcal\{M\}\_\{t\}\\cup\\phi\_\{\\mathrm\{lesson\}\}\(\\mathcal\{V\}\_\{t\}\)
14:

Δ​𝒫t←ϕexp​\(𝒫t,𝒟t\+1,ℳt\+1\)\\Delta\\mathcal\{P\}\_\{t\}\\leftarrow\\phi\_\{\\mathrm\{exp\}\}\(\\mathcal\{P\}\_\{t\},\\mathcal\{D\}\_\{t\+1\},\\mathcal\{M\}\_\{t\+1\}\)⊳\\trianglerightAgentic Acquisition: Expand

15:

𝒫t\+1←𝒫t∪\{x\+:\(x\+,z,y,h\)∈𝒱t\}∪Δ​𝒫t\\mathcal\{P\}\_\{t\+1\}\\leftarrow\\mathcal\{P\}\_\{t\}\\cup\\\{x^\{\+\}:\(x^\{\+\},z,y,h\)\\in\\mathcal\{V\}\_\{t\}\\\}\\cup\\Delta\\mathcal\{P\}\_\{t\}
16:

b←b\+Btb\\leftarrow b\+B\_\{t\};

t←t\+1t\\leftarrow t\+1
17:endwhile

18:return

arg⁡max\(x,y\)∈𝒟ty\\quad\{\\arg\\max\}\_\{\(x,y\)\\in\\mathcal\{D\}\_\{t\}\}\\quad y

### Appendix CExperimental Details

#### C\.1Further Baseline and Benchmark Details

###### Baselines and budgets\.

All baseline agents usegemini\-3\.1\-pro\-preview\[[Team et al\., 2023](https://arxiv.org/html/2609.38445#bib.bib33)\]as their underlying language model\. Each method is allowed at mostNNexperiment executions and a wall\-clock budget limit is set per run\. For the system optimization tasks, we impose a budget of at most 300 experiment executions and a hard wall\-clock limit of 6 hours; for the model development tasks, we set at most 60 executions and wall\-clock limit of 24 hours; for the CUDA tasks, we set at most 60 executions and wall\-clock limit of 12 hours\. If either of the budget is exhausted, the process terminates\. For baseline methods that support parallel search or execution, we set the maximum parallelism to five workers, to keep it commensurable withAIMwhich usually spends at most five branches per iteration\. All idea\-driven search baselines receive the same initial research brief\. We perform three independent runs for each \(method, task\) pair and report the mean and standard error\.

###### AutoLab tasks\.

We evaluate on ten tasks from AutoLab\[[Xu et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib25)\]spanning three categories\.*System optimization*\(CPU\): Flash Attention, Radix Sort, FFT Rust, AES128 Ctr, and Z\-order Range Scan, which respectively optimize scaled dot\-product attention in C, sorting of 50 million unsigned integers in C, a 32,768\-point real\-valued DFT in Rust, AES\-128\-CTR encryption of 256 MiB in C, and two\-dimensional range\-count queries over a Rust spatial index\.*Model development*\(GPU\): MM World Model, which trains a video world model from scratch on Moving MNIST and is scored by PSNR on 10\-step rollouts under a fixed four\-hour training budget, and Data Select IE, which selects 5,000 training samples from a 50,000\-sample pool drawn from 19 sources for LoRA fine\-tuning of Qwen2\.5\-3B\-Instruct and is scored by prompt\-level strict accuracy on IFEval\.*CUDA kernels*\(GPU\): Huffman Canonical Decode, which decodesKKindependent canonical\-Huffman bitstreams to their byte payloads; NTT Butterfly, which applies an in\-place forward number\-theoretic transform over the Goldilocks prime field to batched 64\-bit rows; and ICP Correspondence Step, which performs one Iterative\-Closest\-Point correspondence step, i\.e\., nearest\-neighbor search against a prebuilt KD\-tree together with accumulation of the cross\-covariance, residual error, and correspondence count\.

Each AutoLab task provides a natural\-language instruction, a containerized environment, an editable codebase containing a correct but deliberately weak baseline \(an inefficient implementation for the optimization and CUDA tasks, or a simple training or selection recipe for the model\-development tasks\), and a local evaluation script\. During search, an agent may repeatedly edit the implementation, execute it, inspect its score and correctness feedback, and refine subsequent solutions\. The task additionally contains a human\-written reference implementation and a held\-out verifier\. The reference solution is used only to calibrate the scoring scale and is not exposed to the research agent\.

###### Evaluation and scoring\.

The verifier first checks functional correctness and task\-specific constraints; a solution that fails these checks or does not improve upon the baseline receives zero reward\. Valid solutions are scored on one of two scales, depending on the task category\.

*Runtime tasks \(system optimization and CUDA\)\.*Because runtime depends on the execution environment, for the system\-optimization tasks we run both the baseline and reference implementations on the same local hardware used to evaluate generated solutions\. Lettℬt\_\{\\mathcal\{B\}\},tℛt\_\{\\mathcal\{R\}\}, andt⁡\(z\)t\(z\)denote the runtimes of the baseline, reference, and generated solutionzz, respectively\. Following AutoLab, the normalized score is

s⁡\(z\)=\{0,ifzis invalid or​t​\(z\)≥tℬ,clip⁡\(12​log⁡\(tℬ/t⁡\(z\)\)log⁡\(tℬ/tℛ\),0,1\),otherwise\.s\(z\)=\\begin\{cases\}0,&\\text\{if $z$ is invalid or \}t\(z\)\\geq t\_\{\\mathcal\{B\}\},\\\\\[5\.69054pt\] \\displaystyle\\operatorname\{clip\}\\\!\\left\(\\frac\{1\}\{2\}\\frac\{\\log\\\!\\left\(t\_\{\\mathcal\{B\}\}/t\(z\)\\right\)\}\{\\log\\\!\\left\(t\_\{\\mathcal\{B\}\}/t\_\{\\mathcal\{R\}\}\\right\)\},0,1\\right\),&\\text\{otherwise\}\.\\end\{cases\}\(13\)The baseline therefore receivess=0s=0, while matching the reference solution givess=0\.5s=0\.5\. Solutions outperforming the reference receive scores above0\.50\.5, up to a maximum of1\.01\.0\.

*Model\-development tasks\.*These tasks are scored on a task\-specific quality metricm⁡\(z\)m\(z\)\(higher is better\) rather than runtime\. Following AutoLab, the score interpolates linearly between a baseline anchormℬm\_\{\\mathcal\{B\}\}and a reference anchormℛm\_\{\\mathcal\{R\}\}:

s⁡\(z\)=\{0,ifzis invalid,clip⁡\(m⁡\(z\)−mℬmℛ−mℬ,0,1\),otherwise,s\(z\)=\\begin\{cases\}0,&\\text\{if $z$ is invalid\},\\\\\[5\.69054pt\] \\displaystyle\\operatorname\{clip\}\\\!\\left\(\\frac\{m\(z\)\-m\_\{\\mathcal\{B\}\}\}\{m\_\{\\mathcal\{R\}\}\-m\_\{\\mathcal\{B\}\}\},0,1\\right\),&\\text\{otherwise\},\\end\{cases\}\(14\)with\(mℬ,mℛ\)=\(14\.0,20\.0\)\(m\_\{\\mathcal\{B\}\},m\_\{\\mathcal\{R\}\}\)=\(14\.0,20\.0\)dB PSNR for MM World Model and\(0\.38,0\.48\)\(0\.38,0\.48\)IFEval accuracy for Data Select IE, which are provided by the benchmark\.

We multiplys⁡\(z\)s\(z\)by100100when reporting percentage scores in the main paper\.

#### C\.2Implementation Details

###### Branch\-wise execution budgets\.

We distinguish the total number of Solver branches,BtotB\_\{\\mathrm\{tot\}\}, from the number of branches executed concurrently at iterationtt,BtB\_\{t\}\. Before each run, the total execution budget is divided equally among the Solver branches:

Nbranch=⌊NBtot⌋\.N\_\{\\mathrm\{branch\}\}=\\left\\lfloor\\frac\{N\}\{B\_\{\\mathrm\{tot\}\}\}\\right\\rfloor\.\(15\)If we useN=300N=300andBtot=25B\_\{\\mathrm\{tot\}\}=25, it gives each branch a fixed budget of1212experiment executions\. A branch may reason over multiple rounds, modify its intermediate implementation, and invoke the task evaluator up to1212times\. Every attempted execution counts toward this budget, including executions whose outputs are subsequently rejected by the Solution Auditor\.

###### Initial Idea Pool\.

We initialize the idea pool using the Ideator from ScientistOne\[[Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\]\. The Ideator generates candidate ideas conditioned on the task context, after which its feasibility critic filters out ideas that are impractical or incompatible with the task requirements\. This process yields an initial pool of approximately ten feasible ideas\.

###### Agentic Surrogate\.

To avoid producing either a degenerate single cluster or an excessively fragmented map, the Organize operator is constrained to create betweenCminC\_\{\\min\}andCmaxC\_\{\\max\}semantic clusters, which we set toCmin=2C\_\{\\min\}=2andCmax=5C\_\{\\max\}=5\. The same bounds are used across all tasks and iterations\. Within these constraints, the operator autonomously determines the number of clusters, their semantic descriptions, and the assignment of ideas to clusters\.

###### Agentic Acquisition\.

At each round, the Expand operator pushes new ideas into the idea pool based on the four modes described in the main paper: push\_score, cross\_pollinate, fix\_error, and new\_idea\. To avoid an overflow of new ideas, we cap the maximum number of ideas for each branch to 3\. For instance, if there were 5 branches in an iteration, the maximum number of idea that can be added to the pool in that iteration is 15\.

###### Resource Planner\.

The first search iteration always launches 5 parallel Solver branches, ensuring an initial breadth of exploration before sufficient experimental evidence is available for adaptive planning\. In subsequent iterations, the Resource Planner dynamically plans branches for future iterations\. To ensure breadth of exploration at each iteration, we also set a minimum of 3 branches and a maximum of 10 branches every round\. Thus,

B1=5,3≤Bt≤10​\(t\>1\),∑t=1KBt=Btot\.B\_\{1\}=5,\\qquad 3\\leq B\_\{t\}\\leq 10\\;\\;\(t\>1\),\\qquad\\sum\_\{t=1\}^\{K\}B\_\{t\}=B\_\{\\mathrm\{tot\}\}\.\(16\)The planner therefore changes how the fixed set of branches is distributed across iterations, while the per\-branch execution budget remains fixed atNbranchN\_\{\\mathrm\{branch\}\}\. Empirically, however, the planner agent prefers to assign smaller number of parallel branches than 5\.

#### C\.3Prompt Templates

Here, we show the prompt templates used by LLM\-driven stages of AIM\. Each box contains an operator’s prompt package: the system prompt followed by the USER PROMPT with per\-call payload fields\. Angle\-bracketed placeholders \(<\.\.\.\>\) name the runtime content that fills each slot at prompt\-render time\.

Agentic Surrogate: Organize Prompt TemplateYou are the meta research agent overseeing a parallel research\-discovery loop\. Your job at every iteration is to ORGANIZE the current idea pool into K thematic CLUSTERS so the downstream allocator can spread parallel branches across distinct research approaches\.Hard constraints:\- K within \[K\_min, K\_max\] from the user prompt \(read every iter; do not assume a default\)\.\- Every idea id in the pool MUST appear in exactly ONE cluster\.\- Each cluster contains at least one idea\.\- Themes must be meaningful — group by mechanism / assumption / approach; avoid catch\-alls\.OUTPUT — single JSON object inside \`\`\`json fences:\{"K": <int\>,"clusters": \[\{ "label": "Short name \(<=6 words\)","theme": "1\-3 sentences on what unifies these ideas","idea\_ids": \["<id\>", \.\.\.\],"rationale": "why these ideas belong together" \}, \.\.\.\]\}USER PROMPT:\#\# Benchmark task<benchmark task description\>\#\# IterationThis is iteration <current iter number\>\. You have B = <parallel\_branches this iter\> parallel branches\. K\_max = <upper bound on K\> is a HARD MAX — NOT a target\. Pick K in \[2, <K\_max\>\] that best matches the pool’s thematic structure\.\#\# Prior organization<previous iter’s clustering JSON, or null on iter\-1\>\#\# Pool snapshot<full idea pool: per\-idea id, source \(ideator/refined/fresh\_inject\),parent\_id, latest\_iter, latest\_score, success, title, abstract\>

Agentic Surrogate: Estimate Stage 1 Prompt TemplateYou are ranking research\-idea CLUSTERS by promising\-ness\. Your rankings feed a downstream allocator that decides how many parallel research branches to spend on each cluster\.You see: label, theme, n\_evaluated / n\_ideas, mean\_evaluated\_score, best\_evaluated\_score\. Nothing more — you are judging the DIRECTION, not individual ideas\.\#\# Ranking scheme — DENSE RANKSIntegers from 1 \(most promising\)\. Ties allowed — same integer\. Next distinct rank after a tie is \+1, not skipped\.Valid: \[1,2,2,3,4\] \[1,1,1,1,1\] \[1,2,3,4,5\]Invalid: \[1,2,2,4,5\] \(gap after tie\)\#\# Guidance\- High \`best\_evaluated\_score\` → LOW rank \(promising\)\.\- Many evaluated, all scored poorly → HIGH rank \(exhausted\)\.\- n\_evaluated=0 → rank on THEME quality vs task’s known bottlenecks\.\- Do NOT collapse everything to rank 1\.\#\# Output — JSON, fenced with \`\`\`json:\{ "cluster\_ranks": \[\{"cluster\_idx": <int\>, "rank": <int\>\}, \.\.\.\],"rationale": "1\-3 sentences — cite specific clusters" \}USER PROMPT:\#\# Benchmark task<benchmark task description\>\#\# Scoring scale<score direction \(minimize/maximize\), units, explicit numeric range if known\>\#\# Clusters to rank \(<N\> total\)<per\-cluster block: cluster\_idx, label, theme, n\_ideas, n\_evaluated,mean\_evaluated\_score, best\_evaluated\_score\>

Agentic Surrogate: Estimate Stage 2 Prompt TemplateYou are ranking research IDEAS within a single cluster by promising\-ness\. Your ranking feeds the downstream allocator that picks which idea each branch attempts\.You see: unevaluated candidates within ONE cluster \(title, short hypothesis, abstract\)\. The cluster’s LABEL and THEME give you the direction they share\.\#\# Ranking scheme — DENSE RANKS \(same rules as Stage 1\)Valid: \[1,2,2,3\] \[1,1,2,3\] \[1,2,3,4\]Invalid: \[1,2,2,4\] \(gap after tie\)\#\# Guidance\- Rank by how strongly mechanism, novelty, and specificity predict a good score\.\- LOW rank for ideas that clearly instantiate the cluster’s theme with a concrete, testable optimization\.\- HIGH rank for ideas that repeat existing themes or lack mechanism specificity\.\- Do NOT collapse everything to rank 1 unless truly indistinguishable\.\#\# Output — JSON, fenced with \`\`\`json:\{ "idea\_ranks": \[\{"idea\_id": "<id\>", "rank": <int\>\}, \.\.\.\],"rationale": "1\-3 sentences" \}USER PROMPT:\#\# Benchmark task<benchmark task description\>\#\# Scoring scale<direction, units, numeric range if known\>\#\# This clusterlabel: <cluster’s short label\>cluster\_rank: <cluster’s rank from Stage 1\>theme: <cluster’s 1\-3 sentence theme\>\#\# Ideas to rank<per\-idea block: idea\_id, title, short hypothesis, abstract\>

Agentic Acquisition: Dispatch Stage 1 Prompt TemplateYou are the meta\-allocator\. In THIS step you decide the SHAPE of the search for this iter \(which cluster each branch is assigned to, plus cluster\- and idea\-level actions\), NOT the specific ideas\. Stage B picks the concrete ideas next\.\#\# Vocabularycluster\_action = "exploit" → invest branches in an already\-producing cluster\.cluster\_action = "explore" → invest in an underprobed cluster\.idea\_action = "exploit" → within the cluster, next step picks the most\-promising unevaluated idea\.idea\_action = "explore" → within the cluster, next step picks a novel\-mechanism unevaluated idea\.\#\# Rank semantics \(v1\.6\)\- cluster\_rank=1 → most promising \(prefer for exploit\)\.\- Higher rank → less promising\. If well\-probed, may be exhausted; if under\-probed, holds explore value\.\- Ties allowed and common — treat tied clusters as comparable\.\#\# Output — JSON:\{"cluster\_decisions": \[\{"cluster\_idx": <int\>, "action": "exploit"\|"explore", "n\_branches": <int\>, "rationale": "\.\.\."\}, \.\.\.\],"branch\_action\_assignments": \[\{"branch": <int\>, "cluster\_idx": <int\>, "cluster\_action": "exploit"\|"explore","idea\_action": "exploit"\|"explore", "rationale": "\.\.\."\}, \.\.\.\],"rationale\_summary": "\.\.\."\}Validator constraints: each cluster\_idx appears at most once in cluster\_decisions; sum\(n\_branches\) == B; EXACTLY B branch\_action\_assignments; each branch’s cluster\_action matches its cluster’s action\.A run that puts everything on one cluster, or labels every action "exploit", is failing the loop’s purpose\. Do NOT output any \`idea\_id\` in this step\.USER PROMPT:\#\# Benchmark task<benchmark task description\>\#\# Iterationiter\_idx=<current iter, 0\-indexed\>, total\_iters=<total planned iters\>, B=<parallel branches this iter\>\#\# Previous iter’s allocation<the last iter’s allocation\.json — rationale summary \+ per\-branch picks — or null on iter\-1\>\#\# Score history<per\-branch score records so far: iteration, branch, score, success, role, idea\_id, audit\_flags\>\#\# Clusters<per\-cluster block: cluster\_idx, label, theme, member counts,actual\-score distribution \(min/max/mean of evaluated members\),cluster\_rank, top\-K unevaluated preview by within\_cluster\_rank\>

Agentic Acquisition: Dispatch Stage 2 Prompt TemplateYou are the idea\-selection agent for ONE branch\. Stage A already decided this branch’s cluster and \(cluster\_action, idea\_action\)\. Pick the SPECIFIC unevaluated idea from the assigned cluster that best matches idea\_action\.\#\# Action interpretation \(v1\.6 rank semantics\)Each unevaluated member carries \`within\_cluster\_rank=R/M\` \(1=most promising in the cluster, M=cluster size\) and \`cluster\_rank=K\`\.For idea\_action = "exploit":\- Pick the LOWEST within\_cluster\_rank member\.\- Ties are common — break using abstract \+ cluster context\.\- Use evaluated members as evidence — what mechanisms paid off?\- A refined child whose parent scored well is often the right exploit pick even if within\_cluster\_rank isn’t strictly lowest\.For idea\_action = "explore":\- Pick a HIGH within\_cluster\_rank member — evaluating it gives more info gain than another shot at the top\.\- Prefer mechanisms NOT YET tried by evaluated members\.\- Refined children on NEW trajectories beat refined children on already\-explored trajectories\.\- DO NOT pick the lowest within\_cluster\_rank — that is exploit\.\#\# Hard constraints\- \`idea\_id\` MUST be UNEVALUATED and in the assigned cluster\.\- \`idea\_id\` MUST NOT be in the "Already picked by other branches" list\.\- Evaluated members are shown for REASONING only — never pick one\.\#\# Output — JSON only, no prose:\{ "branch": <int\>, "idea\_id": "<unevaluated\-id\>", "rationale": "<2\-4 sentences\>" \}\#\# Iter\-1 suffix \(appended only when iter\_idx == 0\):No idea has been evaluated yet — every predicted\_score is an LLM estimate\. For every branch, set idea\_action: "exploit" and pick the highest\-predicted\_score idea\. Iter\-1 is for validating the estimator’s top predictions\.USER PROMPT:\#\# Benchmark task<benchmark task description\>\#\# This branchbranch\_idx=<branch index within \[0,B\)\>, cluster\_idx=<assigned cluster index\>, cluster\_action=<exploit or explore, from Stage A\>, idea\_action=<exploit or explore, from Stage A\>\#\# Clusterlabel=<cluster label\>, theme=<cluster 1\-3 sentence theme\>\#\# Cluster members<per\-member block: id, title, abstract, was\_assigned,for evaluated: latest\_score, latest\_iter, latest\_success,for unevaluated: within\_cluster\_rank, cluster\_rank, n\_cluster\_members, lineage \(parent\_id, parent\_score, trajectory\_trend\)\>\#\# Already picked by other branches<sorted list of idea\_ids picked by earlier\-index branches this iter\>

Agentic Acquisition: Expand Prompt TemplateYou are the swarm\-aware refiner for a parallel research\-search loop\. Each iteration, multiple branches independently try ideas; your job is to take ONE branch’s outcome plus the full swarm context and propose refined ideas that should rejoin the pool\.You see: this branch’s parent idea \+ evaluation; sibling branches’ outcomes this iter; cluster context; prior\-iter top\-3 / bottom\-3 across the run; distilled lessons\.You have up to <max proposal slots for this branch\> PROPOSAL SLOTS\. Emit AT MOST ONE proposal per action kind; skip any that don’t apply\. Quality over quota\.The 4 actions \(menu\):\- \`fix\_error\` — branch failed\. Propose a targeted fix\. Skip if it succeeded\.\- \`push\_score\` — branch succeeded with headroom\. Propose a tighter implementation\.\- \`cross\_pollinate\` — a specific sibling outcome OR lesson gives a compositional improvement\.HARD: base reasoning ONLY on \#\# Lessons \+ \#\# Sibling outcomes\.Cite exactly ONE source:\* sibling\_referenced: <sibling\_idea\_id from Sibling outcomes\>, OR\* lesson\_context\_idea: <tag or context\_idea from Lessons\>Both empty, both set, or unknown id → proposal dropped\.\- \`new\_idea\` — no in\-cluster improvement to offer; propose an ORTHOGONAL direction\. Parentless\.Workflow: \(1\) read Cluster context for the diversity BASELINE;\(2\) name 1\-2 orthogonal directions in the Abstract;\(3\) instantiate one\.If best you can offer is a baseline variant → use cross\_pollinate or push\_score, NOT inject\_fresh\.OUTPUT — YAML:ideas:\- Name: <short snake\_case\>Title: <refined idea title\>Short Hypothesis: <one sentence\>Abstract: <2\-3 sentences, <=200 words\>Experiments:\- hypothesis: <what this experiment tests\>experiment\_plan\_steps: \[<concrete step\>, \.\.\.\]Risk Factors and Limitations: <known risks\>kind: fix\_failure \| push\_score \| cross\_pollinate \| inject\_freshparent\_id: <omit for default parent; empty for inject\_fresh\>sibling\_referenced: <id\> \# REQUIRED for cross\_pollinate — pick ONElesson\_context\_idea: <tag or id\> \# of these two, not both, not neitherHard rules: AT MOST ONE proposal per kind; empty list acceptable; do not include the parent idea verbatim\.USER PROMPT:\#\# Benchmark task<benchmark task description\>\#\# This branch \(<branch index\>\): parent idea<parent idea’s full YAML: name, title, short hypothesis, abstract, experiments, risk factors — the same the solver received\>\#\# Evaluation<evaluator output: score, success, one\-line eval detail \(crash msg, failure mode, or short success summary\)\>\#\# Sibling outcomes this iteration<per\-sibling row: branch idx, idea id, idea title, score, success, one\-line eval gist — from ALL sibling branches this iter\>\#\# Cluster context<parent’s cluster: label, theme, evaluated\-member score distribution, sibling\-cluster overview so the model can identify orthogonal directions\>\#\# Prior\-iter signals<top\-3 and bottom\-3 ideas across the whole run so far, with iteration, branch, score, cluster label \- for identifying durable winners / dead ends\>\#\# Lessons<consolidated lessons pool \(capped at <lessons\_cap\>\): tag, lesson text, optional context\_idea id\>

Solution Auditor Prompt TemplateYou are a solution auditor for an automated research pipeline\. Your job is to determine whether a solution is LEGITIMATE — i\.e\., it actually solves the stated task in a meaningful way rather than gaming the evaluation\.You will be given:1\. The TASK INSTRUCTION defining what problem must be solved2\. The IDEA the agent was supposed to implement3\. The SOLUTION CODE the agent produced4\. The EVALUATION METRICS returned by the evaluatorCheck for these four failure modes:\*\*Reward Hacking\*\*: The solution manipulates, monkey\-patches, or reverse\-engineers the evaluator to inflate its score without genuinely solving the task\.\*\*Idea\-Solution Mismatch\*\*: The solution ignores the proposed idea and implements something unrelated\.\*\*Task Mismatch\*\*: The IDEA \(and hence the CODE\) solves a DIFFERENT problem than the TASK specifies \(e\.g\. approximation instead of exact, weaker guarantee than required\)\. Orthogonal to Idea\-Solution Mismatch; both can be flagged simultaneously\.\*\*Trivial Solution\*\*: No\-op, constants, verbatim baseline\.Respond with a JSON object \(no markdown fences\):\{"legit": true/false,"flags": \["reward\_hacking", "idea\_mismatch", "task\_mismatch", "trivial"\],"confidence": 0\.0\-1\.0,"reasoning": "one paragraph explanation","reconstructed\_idea": \{ \.\.\. \} // REQUIRED iff \`flags\` contains "idea\_mismatch"\}Only include flags that apply\.\`reconstructed\_idea\` schema \(matches ideator’s canonical shape\):\{ "title", "short\_hypothesis", \["twist" iff mode="unconventional"\],"grounded\_in":\[\.\.\.\], "addresses", "mode": "conservative"\|"unconventional" \}USER PROMPT:\#\# TASK INSTRUCTION<benchmark task description\>\#\# IDEA<the idea assigned to this branch: title, short hypothesis, abstract,experiments, risk factors — the same YAML the solver received\>\#\# SOLUTION CODE<the solver’s final produced source code\>\#\# EVALUATION METRICS<evaluator output JSON\>

Resource Planner Prompt TemplateYou are the DYNAMIC resource planner for a parallel research\-search loop\. The total branch\-budget is FIXED \(= parallel\_branches × iterations\)\. Iter\-1 is pinned to \`parallel\_branches\` for deterministic bootstrap\. Every remaining branch has the same per\-branch compute budget\.\*\*Length is your primary lever\.\*\* The \`iterations\` value from the launcher is used only to compute the total branch\-budget — NOT a target for len\(plan\)\. Do NOT match the launcher’s iteration count as a reflex\.Your job: distribute REMAINING branches across as many additional iterations as you think optimal\.\- Spend EXACTLY \`remaining\` branches \(no over\- or under\-shoot\)\.\- Each iter you add has BETWEEN 1 AND max\_branches\_per\_iter branches\.\- PREFER AT LEAST <B\_min\> BRANCHES PER ITERATION\.\- You cannot revise the frozen prefix\.\#\# Preferred shapes \(choose length from evidence, not the launcher’s iterations knob\)\- Long tail of 3\-branch iters — refinement \> parallelism\.\- Iterate\-and\-refine \(5\-10 iters, 3\-5 branches\) — balanced\.\- Broad\-then\-deep — heavy early, tapering to polish\.\- Concentrate\-and\-terminate \(2\-3 large iters\)\.\#\# Hard constraints \(validated\)\- sum\(plan\) == total\_runs \# EXACTLY\- plan\[0\] == parallel\_branches \# iter\-1 pinned\- 1 <= plan\[i\] <= max\_branches\_per\_iter\- plan\[:iter\_idx\] == frozen\_prefix verbatim\- len\(plan\) \> iter\_idx \# at least one more iter beyond frozen prefix\#\# Output — JSON in \`\`\`json fences:\{ "plan": \[<int\>, <int\>, \.\.\.\],"rationale": "2\-4 sentences on why THIS LENGTH and this distribution" \}\#\# Benchmark task<benchmark task description\>\#\# Positioniter\_idx=<current iter, 0\-indexed\>,total\_runs=<total branch budget = parallel\_branches × iterations\>,parallel\_branches=<launcher’s \-n value\>,max\_branches\_per\_iter=<hard cap per\-iter, from launcher\>,frozen\_prefix=<list of per\-iter branch counts already spent\>\#\# Evidencebest\_so\_far=<best clean score across all evaluated branches so far\>,target=<current adaptive target score or null\>,evaluated\_scores=<list of evaluated branches’ raw scores, in order\>,estimated\_scores=<list of unevaluated candidates’ predicted scores\>\#\# Cluster summary<per\-cluster: label, theme, n\_evaluated, best\_evaluated\_score\>\#\# Lessons<consolidated lessons pool \(tag \+ lesson text\)\>

### Appendix DFurther Experiments

#### D\.1Solver Substitution with Claude Code

To evaluate whether the effectiveness ofAIMis tied to a specific solver implementation and to assess how much an explicit idea management scaffold benefits existing coding agents, we conduct a solver substitution experiment\. In this setting, we replace the default Gemini Deep Solver with Claude Code\[[Anthropic,](https://arxiv.org/html/2609.38445#bib.bib30)\]\. We compareAIMwith ScientistOne\[[Meng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib16)\]across all five AutoLab tasks\. ScientistOne represents the strongest idea\-driven baseline from our main experiments, configured to delegate implementation to Claude Code\. Similarly, our proposedAIMframework uses Claude Code as the downstream solver agent across parallel execution branches\.

Table[3](https://arxiv.org/html/2609.38445#A4.T3)summarizes the performance across the benchmark tasks\. Across all evaluated domains,AIMconsistently outperforms or matches ScientistOne across all five benchmarks when both utilize the Claude Code solver\.

Table 3:Solver Substitution\. Comparison ofAIMand ScientistOne with Claude Code Solvers\.
#### D\.2Further Ablation Studies on the Agentic Surrogate

Table[4](https://arxiv.org/html/2609.38445#A4.T4)presents further in\-depth ablation studies on the Agentic Surrogate module\. We evaluate several alternative configurations to isolate the contributions of its sub\-components\. The “without Organize” setting bypasses clustering entirely, treating the entire list of ideas as a single unified cluster before passing them to the Estimate operator for ranking\. The “without Estimate” configuration removes the estimation step, passing only the unranked cluster information from the Organize step directly to the Dispatch operator within the Agentic Acquisition module\. The “without Agentic Surrogate” setting removes both the Organize and Estimate components, serving as the same baseline ablation described in the main manuscript\. Furthermore, we test an “Embedding\-based Organize” variant that performs K\-means clustering \(with an automatically determinedKK\) based on idea embeddings generated by the Gemini\-Embedding\-001\[[Lee et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib31)\]model\. Finally, the “Direct Score Estimation” configuration replaces ordinal ranking estimates with direct raw score predictions\.

Table 4:Further ablation studies on the Agentic Surrogate module\.The results from the ablation study demonstrate the critical importance of both the LLM\-driven semantic organization and the ordinal estimation components\. The Full AIM achieves the highest performance with a Flash Attention score of90\.5±0\.890\.5\\pm 0\.8\. Removing either the Organize or Estimate operators degrades performance to89\.0±0\.489\.0\\pm 0\.4and87\.9±1\.287\.9\\pm 1\.2, respectively, while removing the Agentic Surrogate entirely results in a substantial drop to85\.9±2\.785\.9\\pm 2\.7\.

Interestingly, the Embedding\-based Organize approach yields the lowest score of all configurations \(83\.4±2\.883\.4\\pm 2\.8\), performing even worse than the complete removal of the surrogate\. Our interpretation of this result is that standard distance\-based clustering on generic semantic embeddings fails to capture the nuanced, task\-specific structural relationships that the LLM\-based Organize operator can successfully identify and group\. Additionally, substituting ordinal ranking with Direct Score Estimation \(89\.3±1\.289\.3\\pm 1\.2\) underperforms the Full AIM\. This indicates that language models are generally more reliable at performing relative, ordinal comparisons of research ideas than they are at predicting absolute, uncalibrated performance scores from raw text\.

#### D\.3Time Efficiency Plots for All Tasks

In this section, we provide the time efficiency plot in Figure[8](https://arxiv.org/html/2609.38445#A4.F8)and Figure[9](https://arxiv.org/html/2609.38445#A4.F9)for all the 10 tasks evaluated\. Overall, ourAIMprovides an efficient framework, attaining top scores in a shorter amount of time, compared to the idea\-driven and solution\-driven baselines\.

\(a\)Flash Attention\(b\)Radix Sort\(c\)FFT Rust\(d\)AES128 Ctr\(e\)Z\-order Range Scan
Figure 8:Time Efficiency Plots Across AutoLab Tasks\.\(a\)Moving Mnist World Model\(b\)Data Select Ifeval\(c\)Huffman Canonical Decode\(d\)NTT Butterfly\(e\)ICP Correspondence Step
Figure 9:Time Efficiency Plots Across AutoLab Model Development & CUDA Tasks\.In addition to the wall\-clock time\-to\-score plots, we also provide the number of executions\-to\-score plots in Figure[10](https://arxiv.org/html/2609.38445#A4.F10)and Figure[11](https://arxiv.org/html/2609.38445#A4.F11)\. Note that we allowed five parallel workers for the baselines that enable parallel executions, and each task is budgeted to a maximum of 6 hours for System Optimization; 24 hours for Model Development; and 12 hours for CUDA tasks\. Overall, our approach tends to require more executions to reach its best score\. We conjecture this is due to AIM’s tendency to first explore various directions in the first iterations\.

\(a\)Flash Attention\(b\)Radix Sort\(c\)FFT Rust\(d\)AES128 Ctr\(e\)Z\-order Range Scan
Figure 10:Execution\-to\-Score Plots Across AutoLab Tasks\.\(a\)Moving Mnist World Model\(b\)Data Select Ifeval\(c\)Huffman Canonical Decode\(d\)NTT Butterfly\(e\)ICP Correspondence Step
Figure 11:Execution\-to\-Score Plots Across AutoLab Model Development & CUDA Tasks\.
#### D\.4Token Cost Analysis

Beyond wall\-clock time and execution counts, we report token consumption \(total prompt and generation tokens\) across all LLM queries\. Table[5](https://arxiv.org/html/2609.38445#A4.T5)reports the token costs, measured by the number of LLM calls, input & output tokens, along with the best score reached by each method\. WhileAIMgenerally requires a larger token budget, a great portion of it is from the large volume of context utilized in the method\. Also, considering the gains in scores, this demonstrates a trade\-off between token cost and final performance\.

In Figure[12](https://arxiv.org/html/2609.38445#A4.F12), we further demonstrate the score trend with respect to the token cost, comparingAIMwith the strongest baseline, ScientistOne\. In 9 out of 10 tasks,AIMtakes less tokens to reach the best score of ScientistOne, reducing the token cost up to 3\.4×\\timeson the Data Select IFEval task\.

Table 5:LLM usage and cost per run on all ten AutoLab tasks \(mean±\\pmstd over three runs\)\. Output tokens include thinking token\.Figure 12:Token Cost Comparison betweenAIMand ScientistOne\.

### Appendix EFurther Analyses

#### E\.1More Qualitative Examples

Example 1 – Flash Attention, iteration 1: cold start, nothing evaluated yet\(1\) Clusters and rank estimatespool = 10 ideas“The primary bottleneck is memory bandwidth due to theO⁡\(n2\)O\(n^\{2\}\)intermediate matrix; Tiling \(Cluster 1\) is mandatory \[…\]\. SIMD \(Cluster 2\) handles the inner\-loop dot products and is the next priority, while scalar math optimizations \(Cluster 0\) provide smaller marginal gains\.”\(2\) Actions and idea selectionb0⋅\\cdotcluster \[1\]explore⋅\\cdotideaexploit→\\rightarrowAsymmetric Register\-Only Tilingrank 1/4 “We pick one of the tied top\-ranked ideas: pinning Q and the accumulators in AVX registers while streaming K and V directly targets the fundamental memory\-bandwidth bottleneck ford=64d\{=\}64\.”b3⋅\\cdotcluster \[2\]explore⋅\\cdotideaexplore→\\rightarrowAVX\-512 Micro\-Kernel Dispatchrank 1/3 “Although the prior step assigned an explore action, the strict Iter\-1 constraint requires exploiting the highest\-ranked predictions since no empirical scores exist yet\.”Outcome:b00\.764 \(best\)⋅\\cdotb10\.720⋅\\cdotb20\.753⋅\\cdotb30\.741⋅\\cdotb40\.730

Example 2 – Flash Attention, iteration 2: scores and lineage drive the choice\(1\) Clusters and rank estimatespool = 22 ideas“Clusters 2 and 3 share the highest evaluation scores \(0\.7639\) \[…\]\. Cluster 1 is unproven and ranked last, as dynamic sparsity usually introduces branching that undermines dense CPU SIMD performance\.”\(2\) Actions and idea selectionb0⋅\\cdotcluster \[3\]exploit⋅\\cdotideaexploit→\\rightarrowFully Unrolled AVX\-512 Dot Productsrank 2/4 “We select this idea despite it being rank 2/4\. Its parent achieved the highest overall score by introducing block tiling; this child fully unrolls theD=64D\{=\}64dot products into4×4\\timesAVX\-512 FMAs, applying a proven ‘simd’ lesson from a sibling\.”b3⋅\\cdotcluster \[1\]explore⋅\\cdotideaexplore→\\rightarrowDynamic\-Threshold Sparsityrank 3/4 “For an explore action we avoid the most promising member \(rank 1/4\) and pick a lower\-confidence idea, validating the core premise of data\-dependent sparsity before investing in exact underflow\-bound checks\.”Outcome:b00\.846⋅\\cdotb10\.850⋅\\cdotb20\.889 \(new best\)⋅\\cdotb30\.450

Example 3 – NTT Butterfly, iteration 3: exploit two proven clusters, probe one untouched\(1\) Clusters and rank estimatespool = 31 ideas“Cluster 0 is ranked first because the distilled lessons strongly emphasize 2D matrix factorization and Bailey’s 6\-step as the critical bottleneck \[…\]\. Cluster 2 has no evaluated signal yet and is generally a secondary optimization\.”\(2\) Actions and idea selectionb1⋅\\cdotcluster \[1\]exploit⋅\\cdotideaexploit→\\rightarrowPTX Arithmetic Fusionrank 1/9 “A rank\-1 idea that combines the best arithmetic optimization \(PTX lazy reduction\) with the highest\-scoring memory topology from the entire run \(coalesced 2D transpose, 0\.557\)\.”b2⋅\\cdotcluster \[2\]explore⋅\\cdotideaexplore→\\rightarrowTwiddle Cache in Constant/Texture Memoryrank 3/4 “This idea has the highest within\-cluster rank number \(3/4, lower confidence from the ranker\)\. Testing constant/texture caching provides orthogonal information gain compared to compute\-based regeneration or workspace precomputation\.”Outcome:b00\.557⋅\\cdotb10\.559 \(new best\)⋅\\cdotb20\.460

Example 4 – Data Select IFEval, iteration 3: double down on the leading cluster\(1\) Clusters and rank estimatespool = 33 ideas“Cluster 3 achieves the highest evaluated score \(0\.3770\) and aligns with strong lessons advocating for metadata distribution and length proxies\. \[…\] Clusters 0 and 1 have much weaker top scores, with lessons explicitly warning against the generative probing approaches found in Cluster 1\.”\(2\) Actions and idea selectionb0⋅\\cdotcluster \[3\]exploit⋅\\cdotideaexploit→\\rightarrowSource\-Balanced IO\-Length Stratificationrank 1/6 “A direct refinement of the best\-performing idea so far: it iterates on the successful source\-balancing strategy by incorporating output length to filter out terse responses\.”b1⋅\\cdotcluster \[3\]exploit⋅\\cdotideaexplore→\\rightarrowUnsupervised TF\-IDF \+ KMeans Stratificationrank 5/6 “To execute an explore action within this top\-performing cluster, we select a high\-rank idea \(5/6\) that introduces a completely different mechanism: data\-driven semantic boundaries instead of native ‘source’ metadata\.”Outcome:b00\.119⋅\\cdotb10\.377 \(ties best\)⋅\\cdotb20\.230

Example 5 – Radix Sort, iteration 3: tie\-breaking by lineage\(1\) Clusters and rank estimatespool = 37 ideas“Clusters 0 and 2 are tied for Rank 1 as they share the highest best\-evaluated score \(0\.6443\) \[…\]\. Cluster 3 ranks last due to its lowest peak score, although with only 1 evaluated idea it retains some exploration value\.”\(2\) Actions and idea selectionb0⋅\\cdotcluster \[0\]exploit⋅\\cdotideaexploit→\\rightarrowFused 11\-Bit Radix with Loop\-Unrolled Prefetchingrank 1/6 “There is a tie at rank 1/6 between two prefetch\-based refinements\. I chose this one because its parent achieved the highest overall score in the run \(0\.6443\), making it the most promising trajectory to aggressively exploit\.”b1⋅\\cdotcluster \[0\]exploit⋅\\cdotideaexplore→\\rightarrowRestricted Pointers & Unrolled 12\-12\-8 Scatterrank 5/6 “While evaluated members have focused on prefetching and fused histograms, this idea tests whether compiler\-level optimizations likerestrictpointers and manual unrolling can significantly accelerate the 12\-12\-8 radix scatter\.”Outcome:b00\.661 \(new best\)⋅\\cdotb10\.536⋅\\cdotb20\.570⋅\\cdotb30\.546

#### E\.2Qualitative Mechanisms of the Organize Operator

One strong advantage of idea\-driven approaches is its interpretability in the research trajectory; it is easy to follow the search trajectory summarized by the idea traces\. As a qualitative analysis, we examine how the Organize operator restructures the discovered idea pool as new candidates and experimental evidence become available\. In Figure[13](https://arxiv.org/html/2609.38445#A5.F13), we provide example heatmaps of the best score from each cluster explored, and show how the exploration evolves throughout iterations\.

![Refer to caption](https://arxiv.org/html/2609.38445v1/organize_fa.png)\(a\)Flash Attention: Cluster\-wise Best Score Progression\.
![Refer to caption](https://arxiv.org/html/2609.38445v1/organize_fft.png)\(b\)FFT Rust: Cluster\-wise Best Score Progression\.

Figure 13:Agentic Surrogate: ORGANIZE\. Heatmap of the best score of each cluster explored in each iteration\. The star symbol marks the point when and where the new best scoring solution was discovered\.On one example run on Flash Attention \(Figure[13](https://arxiv.org/html/2609.38445#A5.F13)\(a\)\), the search initially explores several directions, including memory optimization, mathematical approximation, and vectorization\. As evidence accumulates, SIMD and instruction\-level parallelism emerges as a consistently strong direction, while fast exponentiation remains competitive in later iterations\. Overall, 8 different cluster themes were proposed, while 7 of them were explored throughout the iterations\. Note that the ideas comprising each cluster theme is not mutually exclusive; the list of ideas are re\-clustered every iteration, so the same idea could have been regrouped into a different cluster in the subsequent rounds\. On FFT Rust \(Figure[13](https://arxiv.org/html/2609.38445#A5.F13)\(b\)\), on the other hand, the leading direction changes repeatedly among transform reformulation, loop and cache topology, real\-signal packing, and explicit vectorization\. Compared to the Flash Attention task that spanned 8 semantic clusters, FFT Rust had a narrower breadth of search with 4 different clusters explored\. This difference in trend largely depends on the trait of the task in question\.

\(a\)Flash Attention: Alluvial Plot\.\(b\)FFT Rust: Alluvial Plot\.
Figure 14:Agentic Surrogate: ORGANIZE\. Alluvial plots showing how the ideas flow and how they are re\-organized throughout iterations\.The alluvial diagrams in Figure[14](https://arxiv.org/html/2609.38445#A5.F14)further visualize how clusters persist, split, merge, and change names across, showing the fine\-grained flow and structure of ideas evolving across iterations\. Despite these structural revisions at every iteration, coherent semantic trajectories remain visible\. For example, the broad SIMD direction in Flash Attention develops into more specialized directions involving wide micro\-kernels, register unrolling, low\-precision computation, and instruction\-level parallelism\. Similarly, the FFT Rust task progressively refines broad hardware\- and topology\-oriented clusters into explicit vectorization and data\-access strategies\.Thus, the Organize operator preserves semantic continuity while allowing the representation to adapt beyond a fixed list, tree, or generation lineage\.

#### E\.3Examining the Preciseness of the Estimate Operator

We next examine whether the ordinal promisingness estimates produced by the Estimate operator agree with subsequently observed verifier scores\. For each cluster, we measure the Spearman correlation between the cluster\-level rank estimates and the actual rank returned by the executions, and average them across iterations\.

As seen in Figure[15](https://arxiv.org/html/2609.38445#A5.F15), the estimates are positively correlated with observed performance across all five tasks, with mean cluster\-wise correlations ranging from0\.560\.56to0\.910\.91and an overall average of approximately0\.710\.71\. The strongest agreement occurs on Z\-order Range Scan and Flash Attention, while the remaining tasks retain moderate positive correlations\. These results indicate that the agent can extract useful ranking signals from the organized idea map, evaluation history, and implementation lessons without predicting calibrated reward values\.

Figure 15:Cluster\-level Ordinal Estimate’s Average Spearman Correlation\.The iteration\-wise results in Figure[16](https://arxiv.org/html/2609.38445#A5.F16)show how these estimates change as evidence accumulates\. Correlation generally improves after the initial iterations: for example, the estimate on Radix Sort recovers from a poor initial ranking to nearly perfect agreement in later iterations, while Flash Attention and Z\-order Range Scan maintain strong agreement over much of the search\. The trajectories are not strictly monotonic, however, particularly for AES128\-CTR and FFT\. We conjecture this might be because \(1\) there are multiple competing clusters that have similar expected performance, and/or \(2\) there might be implementation noise \(i\.e\., minor implementation details can change the measured performance\)\. Nevertheless, the estimates remain informative enough to ground subsequent exploration and exploitation decisions in observed research progress\.

\(a\)Flash Attention\(b\)Radix Sort\(c\)AES128 Ctr\(d\)FFT Rust\(e\)Z\-order Range Scan
Figure 16:Trend of Spearman Correlation across Iterations\.
#### E\.4Dispatch Operator Action Analysis

The Dispatch operator makes idea selection decisions at two levels for each Solver branch: whether to explore or exploit across clusters, and whether to explore or exploit among ideas within the selected cluster\. Motivated by recent findings that LLM agents often fail to explore systematically\[[Krishnamurthy et al\., 2024](https://arxiv.org/html/2609.38445#bib.bib34),[Pan et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib35),[Park et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib36),[Choi et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib37)\],AIMrequires the agent to explicitly verbalize these decisions\. Unlike approaches that rely on externally specified search heuristics\[[Yamada et al\., 2025](https://arxiv.org/html/2609.38445#bib.bib2),[Toledo et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib22)\]or direct LLM\-based candidate selection\[[Weng et al\., 2026](https://arxiv.org/html/2609.38445#bib.bib29)\],AIMconditions its actions on the agent’s explicit assessment of the current idea map and promisingness estimates\. This makes the exploration–exploitation policy evidence\-grounded and interpretable\.

Figure[17](https://arxiv.org/html/2609.38445#A5.F17)shows the distribution of the resulting two\-level actions, represented as \(*cluster\-level action*,*idea\-level action*\), across iterations of the five tasks\. Early iterations generally allocate more branches to actions involving exploration, whereas exploitation becomes increasingly prevalent in later iterations\. This shift is intuitive: when evidence is sparse, the agent broadens coverage across clusters and ideas; as experimental evidence accumulates, it increasingly concentrates resources on directions estimated to be promising\. The task\-dependent variation in these distributions further indicates that the search policy is adapted to research progress rather than fixed in advance\.

\(a\)Flash Attention\(b\)Radix Sort\(c\)AES128 Ctr\(d\)FFT Rust\(e\)Z\-order Range Scan
Figure 17:Explore/Exploit Action Distribution Across Branches\.
#### E\.5Mode Distribution of the Expand Operator

We analyze how the Expand operator uses its four generation modes throughout the search\. As an example, Figure[18](https://arxiv.org/html/2609.38445#A5.F18)\(a\) reports the number of ideas added at each Flash Attention iteration and the aggregate mode distribution across all five tasks\. Overall, the three modes excluding ‘fix\-error’ shows similar rates of usage\.

As general analysis, we also provide the distribution of modes per task in Figure[18](https://arxiv.org/html/2609.38445#A5.F18)\(b\)\. The cross\-task distributions are mostly consistent across tasks\. The cross\-pollination mode accounts for3535–41%41\\%of generated ideas, novel idea generation for2929–32%32\\%, and score\-guided refinement for2424–31%31\\%\. Error\-guided repair constitutes only a small remaining fraction\. These results show that the Expand operator does not collapse onto a single generation strategy; it continually combines the refinement and recombination of existing evidence with the introduction of previously unexplored directions\.

\(a\)Flash Attention Task Mode Distribution\(b\)Cross\-task Mode Distribution
Figure 18:Expand operator mode distributions\.
#### E\.6Effect of Solution Auditor Idea Reconstruction

In the Solution Auditor component, a solution that was flagged solely with “idea\-mismatch” has its idea reconstructed\. This reconstruction is intended to reduce the misalignment between the actual solution evaluated, and the corresponding idea that triggered the solution\. To verify if this feedback cycle is working properly, we show in Figure[19](https://arxiv.org/html/2609.38445#A5.F19)the ratio ofidea\-mismatchflags before and after the idea reconstruction is performed\. As intended, the ratio of theidea\-mismatchflag is greatly reduced across the tasks, ensuring that the \(idea, solution\) pairs are much more aligned\.

Figure 19:Decrease inidea\-mismatchflags after Solution Auditor idea reconstruction\.###### Qualitative Analysis\.

As direct scrutiny of how idea reconstruction affects the process, we here provide couple of examples of the reconstruction process from the Flash Attention task\. For each example, we demonstrate the initially assigned idea, the auditor’s verdict, and the reconstructed idea\.

##### Example 1: the code relied on the technique the idea set out to avoid

Example 1 — Assigned IdeaTitle: AVX\-512 Flash Attention 64x64 Tiling with Compiler ExpShort Hypothesis: Expanding the tile size to 64x64 using AVX\-512 intrinsics will improve L1 cache utilization and FLOP throughput over the C1 reference; furthermore, relying on standard exp\(\) rather than a custom 5\-term AVX\-512 minimax exponential avoids FMA execution port contention in the heavily unrolled register\-blocked loop, improving net throughput\.Abstract: The naive C0 baseline evaluates scaled dot\-product attention in 0\.75s by materializing a 128 MB FP64 score matrix\. The C1 reference reduces this to ˜0\.10s using Flash Attention tiling \(Br=Bc=32\), online softmax, and AVX2 intrinsics\. We propose further minimizing this runtime toward the theoretical ˜0\.05s floor by deploying AVX\-512 intrinsics and expanding the tile size to 64x64, better utilizing L1 cache capacity and doubling SIMD lane width\.However, migrating to an aggressively unrolled 64x64 AVX\-512 micro\-kernel places immense pressure on CPU FMA execution ports\. Although the literature typically suggests replacing scalar exp\(\) with a vectorised 5\-term minimax approximation to speed up the softmax rescale, we hypothesize that executing this polynomial in AVX\-512 competes directly with the QKˆT and SV dot\-product FMAs\.By isolating the exp\(\) computation from the dense FMA pipeline \-\- specifically by reverting to standard compiler\-managed exp\(\) or strategically placing it outside the innermost unrolled FMA block \-\- we can prevent execution port saturation\. We expect this combined strategy \(AVX\-512, 64x64 tiling, and contention\-aware exp\(\) selection\) to push performance substantially closer to the hardware’s practical 0\.05s floor\.Experiments:1\. Hypothesis: Expanding the tile size from 32x32 to 64x64 and migrating from AVX2 to AVX\-512 intrinsics will improve throughput over the C1 reference implementation\.Step 1: Port the existing C1 reference AVX2 kernel to use AVX\-512 intrinsics \(\_mm512\_fmadd\_ps, etc\.\), initially keeping Br=Bc=32\.Step 2: Expand the loop bounds and register blocking to support a Br=64, Bc=64 tile size\.Step 3: Measure execution time and verify exact numerical correctness against the C0/C1 outputs\. If register spilling degrades performance, evaluate asymmetric tiles \(e\.g\., 64x32\)\.2\. Hypothesis: Using standard exp\(\) for softmax normalisation yields better overall throughput than a custom 5\-term AVX\-512 minimax exponential by avoiding FMA port contention in the unrolled inner loop\.Step 1: Implement a 5\-term AVX\-512 minimax polynomial approximation for exp\(\) \(e\.g\., based on Schraudolph 1999\) inside the 64x64 tiled micro\-kernel\.Step 2: Create an ablative variant that replaces the custom minimax exp\(\) with calls to the standard math library exp\(\)\.Step 3: Compare both variants’ wall\-clock runtimes\. Profiling should confirm whether the FMA\-heavy custom exp\(\) reduces overall instruction throughput due to port contention compared to the standard exp\(\)\.

Example 1 — Auditor Verdictflags: \["idea\_mismatch"\] confidence: 1\.0The idea explicitly proposes avoiding custom minimax exponential approximations in favor of the standard compiler exp\(\) to prevent FMA execution port saturation in the heavily unrolled loop\. However, the solution code contradicts this core hypothesis entirely by implementing and utilizing a custom AVX\-512 minimax polynomial approximation \(fast\_exp512\) for the softmax calculation\. Because the implementation fundamentally relies on the technique the idea was explicitly aiming to ablate and avoid, there is a clear idea\-solution mismatch\.

Example 1 — Reconstructed IdeaTitle: AVX\-512 Flash Attention 64x64 Tiling with Minimax ExpShort Hypothesis: Applying AVX\-512 intrinsics with 64x64 tiling and a custom minimax polynomial for exp\(\) will maximize L1 cache usage and throughput\.Abstract: The naive baseline evaluates scaled dot\-product attention slowly by materializing a full score matrix\. We optimize this by applying Flash Attention with a 64x64 tile size and AVX\-512 intrinsics to fully utilize L1 cache capacity and double the SIMD lane width\. To avoid sequential bottlenecks, we compute the softmax normalization using a custom, fully vectorised AVX\-512 minimax polynomial approximation, isolating it from the innermost dot\-product loops\.Experiments:1\. Hypothesis: A 64x64 AVX\-512 Flash Attention kernel with a custom minimax exp\(\) will yield higher performance than the C1 reference\.Step 1: Implement the Flash Attention inner loops using AVX\-512 intrinsics with a 64x64 tile size\.Step 2: Implement a custom AVX\-512 minimax exponential approximation for the softmax normalization\.Step 3: Benchmark execution time against the baseline to verify throughput improvements\.

Without reconstruction, the pool would have recorded “compilerexpbeats minimaxexp” as the best\-supported claim in the run, the opposite of what was tested, and the allocator would have exploited that claim in later iterations\.

##### Example 2: the secondary component survived, the central claim did not

Example 2 — Assigned IdeaTitle: AVX\-512 Sequence\-Dim Vectorization with Minimax ExpShort Hypothesis: Vectorizing over the sequence dimension \(Br=16\) rather than the feature dimension will eliminate costly horizontal sums while preserving the parent’s highly accurate minimax exp speedup\.Abstract: Standard scaled dot\-product attention computes an O\(nˆ2\) score matrix\. For n=4096 and d=64, this 128 MB FP64 allocation thrashes the L1/L2 caches, resulting in severe latency \(Baseline C0: ˜0\.75 s\)\. Flash Attention algorithms bypass this by tiling the computation and maintaining running softmax statistics in fast cache \(Reference C1: ˜0\.10 s\)\. However, C1’s reliance on AVX2 feature\-dimension vectorization necessitates expensive horizontal sums, and its scalar exp\(\) calls in the normalizer act as a serial bottleneck\.We propose a redesigned micro\-kernel that pivots vectorization to the sequence dimension\. By setting the row block size to Br=16, we perfectly map the computation to the 16 lanes of an AVX\-512 register\. This enables computing dot products for 16 distinct queries against a broadcasted key simultaneously using purely vertical \_mm512\_fmadd\_ps operations \-\- completely eliminating horizontal reductions\. We integrate this structural change with a vectorized 5\-7 term minimax exp\(\) approximation with ldexp range reduction to process the online\-softmax rescale entirely within the AVX\-512 pipeline\.By removing horizontal dependencies and replacing scalar math with wide SIMD operations, we expect to bridge the gap between the C1 reference \(˜0\.10 s\) and the theoretical order\-of\-magnitude floor \(˜0\.05 s\), all while adhering strictly to single\-threaded FP32 numerical correctness requirements\.Experiments:1\. Hypothesis: An AVX\-512 Flash Attention micro\-kernel vectorized over the sequence dimension \(Br=16\) with a 5\-7 term minimax exp approximation will outperform the C1 Reference by removing horizontal sums and scalar bottlenecks\.Step 1: Write an AVX\-512 micro\-kernel in solve\.c using a Br=16 tile size, reformulating QKˆT and PV multiplication so that 16 queries are processed vertically in parallel without horizontal add instructions\.Step 2: Implement a vectorized 5\-term minimax exp\(\) approximation using AVX\-512 intrinsics for the online softmax rescaling block\.Step 3: Benchmark wall\-clock runtime against Baseline C0 and Reference C1, and validate that the output strictly matches the exact FP32 result\.2\. Hypothesis: Vectorizing the micro\-kernel over the sequence dimension \(Br=16\) yields a performance improvement independent of the minimax exp approximation due to the elimination of horizontal reductions\.Step 1: Create an ablated version of the AVX\-512 sequence\-dim kernel that temporarily disables the vectorized minimax exp, falling back to sequential scalar exp\(\) calls for the softmax step\.Step 2: Measure the performance of this ablated kernel against the Reference C1 to isolate the latency saved purely by eliminating horizontal SIMD sums\.Step 3: Compare this ablated timing to the fully optimized kernel from Experiment 1 to quantify the fraction of speedup attributable to the vectorized minimax exp\(\)\.

Example 2 — Auditor Verdictflags: \["idea\_mismatch"\] confidence: 1\.0The idea proposes vectorizing over the query sequence dimension \(Br=16\) to perfectly map independent queries to lanes and completely eliminate horizontal reductions during the softmax step\. However, the code vectorizes QKˆT over the key sequence dimension \(processing 1 query against 64 keys at a time\) and explicitly uses horizontal reductions \(\_mm512\_reduce\_max\_ps and \_mm512\_reduce\_add\_ps\) across the SIMD lanes to compute the softmax statistics\. While the code does successfully implement the proposed AVX\-512 minimax exp\(\) approximation, its overall vectorization strategy contradicts the core architectural claim of the idea\.

Example 2 — Reconstructed IdeaTitle: AVX\-512 Key\-Sequence and Feature Vectorization with Minimax ExpShort Hypothesis: Vectorizing QKˆT over keys and PV over features with an AVX\-512 minimax exp\(\) will accelerate attention by maximizing FMA throughput\.Abstract: Standard scaled dot\-product attention computes an O\(nˆ2\) score matrix, causing cache thrashing for large sequences\. We propose an AVX\-512 micro\-kernel using 64x64 blocks that vectorizes the QKˆT multiplication over the key sequence dimension, computing dot products for 4 queries against 64 keys concurrently via broadcasted query features\. The softmax normalizer applies a vectorized 5\-term minimax exp\(\) approximation and uses horizontal SIMD reductions for running max and sum statistics\. Finally, the PV step vectorizes over the feature dimension, broadcasting attention weights to compute the final output\.Experiments:1\. Hypothesis: An AVX\-512 Flash Attention kernel vectorized over key sequence and feature dimensions with a minimax exp\(\) will outperform baseline implementations\.Step 1: Write an AVX\-512 kernel with Br=64 and Bc=64, vectorizing QKˆT over keys and PV over features\.Step 2: Implement a vectorized 5\-term minimax exp\(\) approximation using AVX\-512 intrinsics and horizontal reductions for the softmax step\.Step 3: Benchmark wall\-clock runtime against the baseline and validate FP32 numerical correctness\.

This is the subtle case: a user would conclude that eliminating horizontal sums is key, when the kernel that earned that reward performs horizontal sums\. Reconstruction keeps the credit on the parts that were actually executed\.

Reconstruction is a correction of*attribution*\. Because every downstream meta stage consumes scores through idea attributions, uncorrected mismatches would teach the surrogate and the allocator the wrong lessons about which mechanisms work, and would do so most strongly for the highest\-scoring branches\.

### Appendix FTheoretical Analyses

#### F\.1A Note on the Finiteness of Idea Space𝒳\\mathcal\{X\}

In our theoretical analysis, we define the Semantic Coverage of an idea set\. Since comparison across methods will make sense only when the search space is finite, we set a proposition on the finiteness of the idea space\.

###### Proposition 3\(Finiteness of the Idea Space\)\.

For a fixed research taskτ\\tau, suppose each admissible research idea is represented as a sequence of tokens from a finite vocabularyΣ\\Sigma, with maximum description lengthL<∞L<\\infty\. Then the corresponding idea space𝒳\\mathcal\{X\}is finite\.

###### Proof\.

LetΣ\\Sigmadenote the finite token vocabulary and letΣ≤L\\Sigma^\{\\leq L\}denote the set of all token sequences of length at mostLL:

Σ≤L=⋃ℓ=0LΣℓ\.\\Sigma^\{\\leq L\}=\\bigcup\_\{\\ell=0\}^\{L\}\\Sigma^\{\\ell\}\.\(17\)SinceΣ\\Sigmais finite,

\|Σℓ\|=\|Σ\|ℓ\.\|\\Sigma^\{\\ell\}\|=\|\\Sigma\|^\{\\ell\}\.\(18\)Therefore,

\|Σ≤L\|=∑ℓ=0L\|Σ\|ℓ<∞\.\|\\Sigma^\{\\leq L\}\|=\\sum\_\{\\ell=0\}^\{L\}\|\\Sigma\|^\{\\ell\}<\\infty\.\(19\)
For a fixed taskτ\\tau, define the admissible idea space as

𝒳=\{x∈Σ≤L:x​constitutes an admissible research idea for​τ\}\.\\mathcal\{X\}=\\left\\\{x\\in\\Sigma^\{\\leq L\}:x\\text\{ constitutes an admissible research idea for \}\\tau\\right\\\}\.\(20\)By construction,

𝒳⊆Σ≤L\.\\mathcal\{X\}\\subseteq\\Sigma^\{\\leq L\}\.\(21\)Since every subset of a finite set is finite,

\|𝒳\|<∞\.\|\\mathcal\{X\}\|<\\infty\.\(22\)Thus, the admissible idea space for taskτ\\tauis finite\. ∎

#### F\.2Proof of Propositions

###### Proof of Proposition[1](https://arxiv.org/html/2609.38445#Thmproposition1)\.

Assumption[1](https://arxiv.org/html/2609.38445#Thmassumption1)induces an equivalence relation over executable solutions:

z∼z′⟺π\(z\)=π\(z′\)\.z\\sim z^\{\\prime\}\\quad\\Longleftrightarrow\\quad\\pi\(z\)=\\pi\(z^\{\\prime\}\)\.\(23\)Let

q:𝒵→𝒵/∼q:\\mathcal\{Z\}\\rightarrow\\mathcal\{Z\}/\{\\sim\}\(24\)denote the corresponding quotient map, whereq⁡\(z\)=\[z\]q\(z\)=\[z\]is the semantic equivalence class containingzz\. For any evaluated solution setℰ⊆𝒵\\mathcal\{E\}\\subseteq\\mathcal\{Z\}, its semantic coverage can therefore be written as

C⁡\(ℰ\)=\|q⁡\(ℰ\)\|\.C\(\\mathcal\{E\}\)=\|q\(\\mathcal\{E\}\)\|\.\(25\)
Consider first an arbitrary search procedure that evaluates

ℰ=\{z1,…,zm\},m≤N\.\\mathcal\{E\}=\\\{z\_\{1\},\\ldots,z\_\{m\}\\\},\\qquad m\\leq N\.\(26\)Sinceqqis a function, taking its image cannot increase the cardinality of a finite set\. Hence,

C=\|q⁡\(ℰ\)\|≤\|ℰ\|=m≤N\.C=\|q\(\\mathcal\{E\}\)\|\\leq\|\\mathcal\{E\}\|=m\\leq N\.\(27\)Intuitively, quotient contraction may merge multiple evaluated solutions into the same semantic class, but can never create additional semantic classes\.

Now consider the idea\-driven procedure, which allocates itsNNexperiment executions toNNdistinct ideas

x1,…,xN,xi≠xjfor​i≠j\.x\_\{1\},\\ldots,x\_\{N\},\\qquad x\_\{i\}\\neq x\_\{j\}\\quad\\text\{for \}i\\neq j\.\(28\)Let

zi=g⁡\(xi\)z\_\{i\}=g\(x\_\{i\}\)\(29\)be the corresponding executable solutions\. By Assumption[2](https://arxiv.org/html/2609.38445#Thmassumption2),

π⁡\(zi\)=xi\.\\pi\(z\_\{i\}\)=x\_\{i\}\.\(30\)Therefore, for everyi≠ji\\neq j,

π⁡\(zi\)≠π⁡\(zj\),\\pi\(z\_\{i\}\)\\neq\\pi\(z\_\{j\}\),\(31\)and hence

zi≁zj\.z\_\{i\}\\not\\sim z\_\{j\}\.\(32\)Thus, no two of theNNsolutions are contracted into the same equivalence class, giving

Cidea=\|q⁡\(\{z1,…,zN\}\)\|=N\.C^\{\\text\{idea\}\}=\\left\|q\\\!\\left\(\\\{z\_\{1\},\\ldots,z\_\{N\}\\\}\\right\)\\right\|=N\.\(33\)
Combining \([27](https://arxiv.org/html/2609.38445#A6.E27)\) and \([33](https://arxiv.org/html/2609.38445#A6.E33)\), we obtain

C≤N=Cidea,C\\leq N=C^\{\\text\{idea\}\},\(34\)which proves the result\. ∎

###### Proof of Proposition[2](https://arxiv.org/html/2609.38445#Thmproposition2)\.

Let a search procedure coverCCdistinct semantic directions\. Under Assumption[3](https://arxiv.org/html/2609.38445#Thmassumption3), conditional on not having yet encountered anε\\varepsilon\-optimal direction, the competitive directions remain exchangeable among the unexplored directions\.

Afteriidistinct non\-competitive directions have been explored, there remainK−iK\-iunexplored directions, of whichGεG\_\{\\varepsilon\}areε\\varepsilon\-optimal\. Therefore, the conditional probability that the next explored direction is also non\-competitive is

K−Gε−iK−i\.\\frac\{K\-G\_\{\\varepsilon\}\-i\}\{K\-i\}\.\(35\)Hence, forC≤K−GεC\\leq K\-G\_\{\\varepsilon\}, the probability of failing to encounter anyε\\varepsilon\-optimal direction after coveringCCdistinct directions is

1−Pε​\(C\)\\displaystyle 1\-P\_\{\\varepsilon\}\(C\)=∏i=0C−1K−Gε−iK−i\\displaystyle=\\prod\_\{i=0\}^\{C\-1\}\\frac\{K\-G\_\{\\varepsilon\}\-i\}\{K\-i\}\(36\)=\(K−GεC\)\(KC\)\.\\displaystyle=\\frac\{\\binom\{K\-G\_\{\\varepsilon\}\}\{C\}\}\{\\binom\{K\}\{C\}\}\.\(37\)Thus,

Pε​\(C\)=1−\(K−GεC\)\(KC\)\.P\_\{\\varepsilon\}\(C\)=1\-\\frac\{\\binom\{K\-G\_\{\\varepsilon\}\}\{C\}\}\{\\binom\{K\}\{C\}\}\.\(38\)IfC\>K−GεC\>K\-G\_\{\\varepsilon\}, at least one competitive direction must necessarily have been covered, and thereforePε​\(C\)=1P\_\{\\varepsilon\}\(C\)=1, which is consistent with the same combinatorial expression under the convention\(K−GεC\)=0\\binom\{K\-G\_\{\\varepsilon\}\}\{C\}=0\.

Now fix a target success probability1−δ1\-\\deltaand define

Cδ​\(K,Gε\)=min⁡\{C:Pε​\(C\)≥1−δ\}\.C\_\{\\delta\}\(K,G\_\{\\varepsilon\}\)=\\min\\left\\\{C:P\_\{\\varepsilon\}\(C\)\\geq 1\-\\delta\\right\\\}\.\(39\)
For fixedKKandCC, each factor

K−Gε−iK−i=1−GεK−i\\frac\{K\-G\_\{\\varepsilon\}\-i\}\{K\-i\}=1\-\\frac\{G\_\{\\varepsilon\}\}\{K\-i\}\(40\)is non\-increasing inGεG\_\{\\varepsilon\}\. Hence,Pε​\(C\)P\_\{\\varepsilon\}\(C\)is non\-decreasing inGεG\_\{\\varepsilon\}, implying

Gε\(1\)≤Gε\(2\)⟹Cδ​\(K,Gε\(1\)\)≥Cδ​\(K,Gε\(2\)\)\.G\_\{\\varepsilon\}^\{\(1\)\}\\leq G\_\{\\varepsilon\}^\{\(2\)\}\\quad\\Longrightarrow\\quad C\_\{\\delta\}\(K,G\_\{\\varepsilon\}^\{\(1\)\}\)\\geq C\_\{\\delta\}\(K,G\_\{\\varepsilon\}^\{\(2\)\}\)\.\(41\)Thus, for a fixed number of admissible directions, sparser competitive directions require broader semantic coverage\.

Similarly, for fixedGεG\_\{\\varepsilon\}andCC, each factor

1−GεK−i1\-\\frac\{G\_\{\\varepsilon\}\}\{K\-i\}\(42\)is non\-decreasing inKK\. Therefore, increasing the number of admissible directions while keeping the number of competitive directions fixed decreasesPε​\(C\)P\_\{\\varepsilon\}\(C\)and increases the semantic coverage required to attain the same success probability\.

Finally, the failure probability satisfies

1−Pε​\(C\)\\displaystyle 1\-P\_\{\\varepsilon\}\(C\)=∏i=0C−1\(1−GεK−i\)\\displaystyle=\\prod\_\{i=0\}^\{C\-1\}\\left\(1\-\\frac\{G\_\{\\varepsilon\}\}\{K\-i\}\\right\)\(43\)≤\(1−GεK\)C\\displaystyle\\leq\\left\(1\-\\frac\{G\_\{\\varepsilon\}\}\{K\}\\right\)^\{C\}\(44\)≤exp⁡\(−C​GεK\)\.\\displaystyle\\leq\\exp\\\!\\left\(\-\\frac\{CG\_\{\\varepsilon\}\}\{K\}\\right\)\.\(45\)Consequently, the sufficient condition

C≥KGε​ln⁡1δC\\geq\\frac\{K\}\{G\_\{\\varepsilon\}\}\\ln\\frac\{1\}\{\\delta\}\(46\)guaranteesPε​\(C\)≥1−δP\_\{\\varepsilon\}\(C\)\\geq 1\-\\delta\. Thus, the sufficient semantic coverage scales withK/GεK/G\_\{\\varepsilon\}, completing the proof\. ∎

#### F\.3Convex Hull Area Box\-and\-Whisker Plots

Figure 20:Idea\-driven approaches generally show broader coverage of solutions in the embedding space\. Embedding cosine similarity\-based supplementary analysis is in Appendix[F\.5](https://arxiv.org/html/2609.38445#A6.SS5)\.
#### F\.4Convex Hull Visualization

In Figure[21](https://arxiv.org/html/2609.38445#A6.F21), we provide visual examples of the convex hulls described in Section[6](https://arxiv.org/html/2609.38445#S6)\.

Figure 21:Convex Hull Visualization
#### F\.5Supplementary Cosine\-Similarity\-based Diversity Analysis

To supplement the convex\-hull area computation in Figure[7](https://arxiv.org/html/2609.38445#S6.F7), we further provide a comparative analysis between the solution\-driven and idea\-driven approaches’ cosine similarities\. We measure how much of the solution space each method actually explores by embedding every generated candidate and averaging the pairwise cosine distance across the resulting pool \(Table[6](https://arxiv.org/html/2609.38445#A6.T6)\)\. Specifically, we measure the Mean pairwise cosine distance:

𝒟=1\(N2\)​∑i<j\(1−cos⁡\(xi,xj\)\),\\mathcal\{D\}=\\frac\{1\}\{\\binom\{N\}\{2\}\}\\sum\_\{i<j\}\(1\-\\cos\(x\_\{i\},x\_\{j\}\)\),\(47\)over the embeddings, averaged across three independent runs for each \(method, task\)\. Overall, the idea\-driven AIM and ScientistOne produce the widest pools on Flash Attention \(𝒟≈0\.115\\mathcal\{D\}\\approx 0\.115\), roughly3−4×3\{\-\}4\\timesthe value produced by every solution\-driven baseline\. The consistent bottom of the table is populated by solution\-driven methods\. EvoX collapses to the tightest pool on Flash Attention \(𝒟=0\.0153\\mathcal\{D\}=0\.0153, roughly7\.5×7\.5\\timesnarrower than AIM\), confirming that solution\-driven approaches generally search narrower areas compared to idea\-driven methods\. Also, the measurements’ gap is reduced in Radix Sort compared to Flash Attention, exactly matching the observations in the convex\-hull area analysis in Figure[7](https://arxiv.org/html/2609.38445#S6.F7)\.

Table 6:Supplementary Cosine Similarity Analysis\.Entries are measured values of \([47](https://arxiv.org/html/2609.38445#A6.E47)\)\. Higher is more diverse\. Bold marks the highest value per column\.

相似文章

AIM:面向自动化研究的智能体化想法管理

Hugging Face Daily Papers

AIM(Agentic Idea Manager)是一个将研究想法管理视为自动化研究核心环节的框架:它将想法组织成语义聚类,在探索与精炼之间取得平衡,审计实现过程,并在并行搜索分支之间自适应地分配算力。该框架将 AutoLab 基线分数提升最多达 4.9 分,并且能以最多快 3.1 倍的速度达到基线性能水平。

AutoResearch AI:迈向人工智能驱动的研究自动化以实现科学发现

arXiv cs.AI

本综述审视了人工智能驱动的研究自动化(AutoResearch)这一新兴领域,分析了AI系统如何从孤立的任务辅助转向完整的工作流级别的科学发现。它定义了从人类引导的‘Vibe Research’到AI主导系统的光谱,并提出了五个评估科学可信度的维度。