ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI

arXiv cs.AI Papers

Summary

ScientistTwo is a fully autonomous multi-agent framework that conducts end-to-end scientific research, generating expert-level papers and codebases that outperform human state-of-the-art models and meet acceptance standards at top-tier AI conferences like ICLR and NeurIPS.

arXiv:2609.19644v1 Announce Type: new Abstract: Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory. The ultimate vision for AI in science is problem-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge. In this paper, we introduce ScientistTwo, a fully autonomous multi-agent framework designed to realize this vision. Specifically, ScientistTwo takes an initial problem as input, establishes state-of-the-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end-to-end discovery cycle without human intervention. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed-loop simulated peer-review rebuttal engine. To evaluate ScientistTwo's capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top-tier conferences such as ICLR, ICML, and NeurIPS. As a result, ScientistTwo autonomously generates expert-level, publishable papers and fully verified, executable codebases. Its solutions consistently outperform human state-of-the-art models, and achieve higher average review ratings than human-authored papers under automated AI review agents. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery. Project website: https://scientist-two.github.io/
Original Article
View Cached Full Text

Cached at: 09/18/26, 09:22 AM

# ScientistTwo: Pioneering the Human Knowledge Frontier with Autonomous AI
Source: [https://arxiv.org/html/2609.19644](https://arxiv.org/html/2609.19644)
\\uselogo\\reportnumber

Jinsung YoonAffiliation:Google Cloud AI ResearchYanzhou PanAffiliation:Google Cloud AI ResearchYubo WangAffiliation:University of WaterlooRui MengAffiliation:Google Cloud AI ResearchParthasarathy RanganathanAffiliation:Google Cloud AI ResearchTomas PfisterAffiliation:Google Cloud AI Research

###### Abstract

Scientific discovery is defined by the ability to identify the boundaries of existing knowledge and venture into unexplored territory\. The ultimate vision for AI in science is problem\-driven autonomous research: given a fundamental challenge by a human expert, the AI independently navigates the scientific landscape, uncovers theoretical and empirical bottlenecks, and systematically expands the frontier of knowledge\. In this paper, we introduce*ScientistTwo*, a fully autonomous multi\-agent framework designed to realize this vision\. Specifically, ScientistTwo takes an initial problem as input, establishes state\-of\-the\-art baselines, formulates novel hypotheses, and coordinates specialized agents to orchestrate an end\-to\-end discovery cycle without human intervention\. Moreover, the framework rigorously conducts experiments using diverse datasets and metrics, refines methodologies through automated ablation studies, and validates research findings via a closed\-loop simulated peer\-review rebuttal engine\. To evaluate ScientistTwo’s capabilities against the highest standards of human scientific achievement, we benchmark it across papers accepted at top\-tier conferences such as ICLR, ICML, and NeurIPS\. As a result, ScientistTwo autonomously generates expert\-level, publishable papers and fully verified, executable codebases\. Its solutions consistently outperform human state\-of\-the\-art models, and achieve higher average review ratings than human\-authored papers under automated AI review agents\. These results show that ScientistTwo is not merely an assistive tool but an autonomous scientific pioneer capable of pushing the frontiers of human discovery\. Project website:[https://scientist\-two\.github\.io/](https://scientist-two.github.io/)

![Refer to caption](https://arxiv.org/html/2609.19644v1/sc2_new_fig2.png)Figure 1:Teaser\.ScientistTwo pushes the frontier of human knowledge across diverse research domains \(e\.g\., LLMs, robotics, neuroscience, speech, robustness, reinforcement learning, game theory, privacy, optimization, and time series\) by generating publication\-quality papers and fully verified codebases, with the resulting methodologies consistently outperforming human state\-of\-the\-art baselines\. Specifically, ScientistTwo improves 86 out of 107 papers \(an 80\.4% success rate\) with an average relative improvement of 25\.2% over human state\-of\-the\-art methods\. Furthermore, papers generated by ScientistTwo surpass the average scores of accepted papers at ICLR 2026 and NeurIPS 2025 under the Stanford Agentic Reviewer, demonstrating its capability to produce manuscripts that reach the empirical acceptance standards of top\-tier AI venues\.## 1Introduction

![Refer to caption](https://arxiv.org/html/2609.19644v1/page3.png)![Refer to caption](https://arxiv.org/html/2609.19644v1/page7.png)![Refer to caption](https://arxiv.org/html/2609.19644v1/page8.png)![Refer to caption](https://arxiv.org/html/2609.19644v1/page9.png)

Figure 2:Case study\.ScientistTwo independently establishes a novel approach that achieves a 10\.9% relative improvement over the human\-designed state\-of\-the\-art baseline\([Jiang and Gong, 2026](https://arxiv.org/html/2609.19644#bib.bib50)\)\. Moreover, evaluated under the ICLR peer\-review procedures, this AI\-authored paper is accepted, receiving impressive scores of 8\.0 from ScholarPeer and 6\.5 from the Stanford Agentic Reviewer\.Scientific discovery has long been the hallmark of human ingenuity, defined by the ability to identify the boundaries of current knowledge and venture into the unknown\. With the rapid advancement of foundation models\([Team et al\., 2023](https://arxiv.org/html/2609.19644#bib.bib107);[Liu et al\., 2024](https://arxiv.org/html/2609.19644#bib.bib66);[Singh et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib99)\), artificial intelligence \(AI\) is transitioning from a passive conversational assistant to an active participant in the scientific process\([Xu and Peng, 2025](https://arxiv.org/html/2609.19644#bib.bib122);[Meng et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib77);[Jansen et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib48)\)\. The ultimate ambition of this paradigm is purely*problem\-driven autonomous discovery*: a human researcher simply specifies a scientific challenge, and an autonomous AI independently navigates the landscape of human knowledge, diagnoses theoretical and empirical bottlenecks, formulates novel hypotheses, and executes the end\-to\-end research lifecycle on its own to pioneer the human knowledge frontier\.

Despite recent progress in the field of autonomous research agents\([Tang et al\., 2026a](https://arxiv.org/html/2609.19644#bib.bib104);[Weng et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib117);[Yamada et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib123)\), a substantial gap remains between automated assistant systems and rigorous empirical scientific standards\. Existing systems primarily focus on optimizing single\-scalar metrics on isolated benchmarks, lacking the multi\-dimensional reasoning capabilities required to tackle complex scientific problems\([Chen et al\., 2026c](https://arxiv.org/html/2609.19644#bib.bib12);[Jin et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib52)\)\. More importantly, current agents lack the closed\-loop empirical rigor of human scientists who continuously iterate based on evidence\. They cannot systematically conduct ablation studies to isolate causal mechanisms for improvement, nor can they engage in dynamic peer\-review processes essential for validating ideas and addressing methodological critiques through targeted supplementary experiments\.

To overcome these fundamental challenges, we introduce*ScientistTwo*, an expert\-level autonomous multi\-agent framework designed to pioneer the frontier of human knowledge \(see Figure[1](https://arxiv.org/html/2609.19644#S0.F1)\)\. Specifically, ScientistTwo operates on a foundational division of labor: the human researcher defines the target scientific challenge, while the AI autonomously orchestrates the path to advance it\. To rigorously demonstrate this capability against the highest standard of scientific rigor, we subject ScientistTwo to the real\-world testing ground: given competitive peer\-reviewed human research from top\-tier AI venues \(e\.g\., ICLR, ICML, NeurIPS\), ScientistTwo must independently discover meaningful advancements beyond the established state\-of\-the\-art, while simultaneously demonstrating that its autonomous scientific discovery process is measurable, reproducible, and transparent\.

![Refer to caption](https://arxiv.org/html/2609.19644v1/scientisttwo_overview.png)Figure 3:Overview\.ScientistTwo leverages previous experiments, ablation studies, and reviewer to generate and validate novel ideas that advance human state\-of\-the\-art baselines, verifying each idea across multiple datasets and evaluation metrics\. Additionally, it simulates the peer\-review process to iteratively refine drafts into publication\-ready papers\.To achieve expert\-level rigor across diverse disciplines demanded by top\-tier science, ScientistTwo coordinates a collaborative ecosystem of specialized agents \(see Figure[3](https://arxiv.org/html/2609.19644#S1.F3)\):

- •*Holistic Benchmark Reasoning & Efficient Screening*: Rather than overfitting to a single metric, ScientistTwo evaluates proposed ideas across comprehensive, multi\-dataset benchmarks\. To balance computation efficiency, it employs a subset\-first evaluation strategy, rapidly filtering ideas on representative benchmark slices before allocating compute to full\-scale experiments\.
- •*Ablation\-Driven Hypothesis Refinement*: Emulating the empirical rigor of expert researchers, ScientistTwo autonomously designs and executes ablation studies to isolate individual component contributions, dynamically pruning ineffective components and refining its core hypothesis\.
- •*Dynamic Peer\-Review & Rebuttal Loops*: To ensure publication\-grade validity, generated drafts are critiqued by a simulated Peer\-Review Agent\. Rather than treating review comments as passive editorial guidance, a dedicated Rebuttal Agent actively conceives, codes, and executes targeted supplementary experiments to address critique\. A Meta\-Review Agent oversees this cycle, triggering recursive refinement loops until rigorous acceptance criteria are satisfied\.

We evaluate ScientistTwo across 107 diverse research challenges encompassing diverse areas such as optimization, reinforcement learning, large language models \(LLMs\), and theory \(see Figure[1](https://arxiv.org/html/2609.19644#S0.F1)\)\. Specifically, ScientistTwo successfully advances 80\.4% of the target problems, delivering an average relative performance gain of 25\.2% over the original human state\-of\-the\-art baselines\. Furthermore, the autonomously generated manuscripts and executable codebases consistently achieve high acceptance rates across independent automated peer\-review evaluations with high scores\. These results demonstrate that ScientistTwo moves beyond passive empirical problem\-solving, serving as an autonomous pioneer capable of systematically expanding the frontiers of scientific knowledge\.

## 2Related Work

AI Agents\.The rapid advancement of LLMs has sparked significant interest in autonomous AI agents capable of planning and using tools\. Early general\-purpose agents such as ReAct\([Yao et al\., 2022](https://arxiv.org/html/2609.19644#bib.bib126)\), HuggingGPT\([Shen et al\., 2023](https://arxiv.org/html/2609.19644#bib.bib98)\), and AutoGen\([Wu et al\., 2024](https://arxiv.org/html/2609.19644#bib.bib120)\)leverage external tools to break down and execute complex, open\-ended tasks\. In the field of software development, systems such as Voyager\([Wang et al\., 2023](https://arxiv.org/html/2609.19644#bib.bib110)\), AlphaCode\([Li et al\., 2022](https://arxiv.org/html/2609.19644#bib.bib64)\), SWE\-agent\([Yang et al\., 2024](https://arxiv.org/html/2609.19644#bib.bib124)\), and Claude Code\([Liu et al\., 2026b](https://arxiv.org/html/2609.19644#bib.bib68)\)use interfaces that recognize execution feedback and the environment to iteratively debug and solve coding problems\. Recently, this paradigm has expanded into the fields of automated machine learning \(ML\) engineering and data science\. Frameworks such as MLAgentBench\([Huang et al\., 2023](https://arxiv.org/html/2609.19644#bib.bib45)\), OpenHands\([Wang et al\., 2025c](https://arxiv.org/html/2609.19644#bib.bib115)\), AIDE\([Jiang et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib51)\), MLE\-STAR\([Nam et al\., 2026a](https://arxiv.org/html/2609.19644#bib.bib80)\), and MARS\([Chen et al\., 2026c](https://arxiv.org/html/2609.19644#bib.bib12)\)automate end\-to\-end ML workflows by executing and improving modeling pipelines, while data science agents such as DA\-Agent\([Huang et al\., 2024](https://arxiv.org/html/2609.19644#bib.bib46)\), Data Interpreter\([Hong et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib43)\), and DS\-STAR\([Nam et al\., 2026b](https://arxiv.org/html/2609.19644#bib.bib81)\)tackle heterogeneous data challenges through structured planning and recursive verification\. In this paper, we introduce ScientistTwo, a specialized AI agent for autonomous research, from idea generation and implementation to the writing of high\-quality publishable academic papers\.

Autonomous Research Agents\.Autonomous research agents have rapidly evolved from ML problems into sophisticated, multi\-stage pipelines that coordinate the entire scientific workflow—from literature\-based analysis and hypothesis generation to execution and drafting of research papers\. Early end\-to\-end pipelines, such as AI Scientist\([Lu et al\., 2024](https://arxiv.org/html/2609.19644#bib.bib73)\), established this automation paradigm but suffered from execution instability and issues with hallucinated writing; subsequent versions, such as AI Scientist\-v2\([Yamada et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib123)\), mitigated these problems through a best\-first\-tree search for experimental branches and review\-based reporting\. At the same time, modular systems have focused on resolving specific bottlenecks; for example, PaperOrchestra\([Song et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib101)\)compiles unconstrained experiment logs and ideas into LaTeX manuscripts, while ScholarPeer\([Goyal et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib36)\)deploys a multi\-agent system that acts as peer\-reviewers\. Other frameworks introduce human\-in\-the\-loop gates\([Schmidgall et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib96)\), apply evolutionary optimization to algorithmic discovery\([Novikov et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib84);[Lyu et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib74)\), or focus on single\-metric optimization for a target codebase\([Li et al\., 2026d](https://arxiv.org/html/2609.19644#bib.bib65);[Liu et al\., 2026a](https://arxiv.org/html/2609.19644#bib.bib67);[Tang et al\., 2026a](https://arxiv.org/html/2609.19644#bib.bib104);[Weng et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib117)\); however, they are generally limited to optimizing a single scalar metric on a single dataset\. Crucially, even verifiability\-centric systems like ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib77)\)lack the ability to perform an active analytical expansion\. Our ScientistTwo addresses these fundamental gaps by conducting systematic experiments on various benchmark datasets and incorporating a dynamic peer\-review loop, automating the iterative process of performing complementary ablation experiments, refining ideas, and revising manuscripts, thus producing expert\-level papers that meet the rigorous standards of top academic venues\.

Figure 4:Python\-style pseudocode for each pipeline stage\.At each stage, ScientistTwo evaluates candidate artifacts using a critic agent\. If accepted, the artifact is returned; if rejected, it is discarded; otherwise, it is iteratively refined based on the critic’s feedback up to a maximum number of rounds\.defstage\(candidate,critic,refine,max\_rounds\):

for\_inrange\(max\_rounds\):

verdict,feedback=critic\(candidate\)

ifverdict=="accept":

returncandidate

elifverdict=="refine":

candidate=refine\(candidate,feedback\)

else:

returnNone

returnNone

Table 1:ScientistTwo overview\.At each stage, ScientistTwo generates acandidateusing a specialized AI agent and employs a correspondingcriticagent to decide whether to accept the output or invoke arefinement agent to improve and supplement it\.stagecandidatecriticrefineGenerating Novel Seed Ideas \(§\\lx@sectionsign[3\.1](https://arxiv.org/html/2609.19644#S3.SS1)\)Finding limitationsSet of limitationsCan guide novel improvement?Add missing limitationsSeed idea generationSeed ideasIs it novel?Add more novel ideasEvaluating Ideas \(§\\lx@sectionsign[3\.2](https://arxiv.org/html/2609.19644#S3.SS2)\)Reproduce baseline on subsetSOTA baseline result\-\-Idea experiment on subsetIdea, Code, ResultsIs it better than the reproduced baseline?Refine idea through engineeringIdea experiment on full\-setIdea, Code, ResultsIs it better than the original SOTA result?Refine idea through engineeringRefining Ideas \(§\\lx@sectionsign[3\.3](https://arxiv.org/html/2609.19644#S3.SS3)\)Idea evolutionEvolved ideaIdea experimentEvolve idea from tracesSelect best ideaBest ideaWhat is the best idea from traces?\-Ablation Studies \(§\\lx@sectionsign[3\.4](https://arxiv.org/html/2609.19644#S3.SS4)\)Ablation studyIdea, Ablation resultsIs the component breakdown clean?Refine the methodDrafting & Meta\-Reviewing \(§\\lx@sectionsign[3\.5](https://arxiv.org/html/2609.19644#S3.SS5),§\\lx@sectionsign[3\.6](https://arxiv.org/html/2609.19644#S3.SS6)\)Initial draftingManuscript\-\-Peer\-ReviewManuscriptIs review score good enough?Run rebuttal experimentsMeta\-ReviewManuscript, ReviewDoes it meet the venue bar?Refine idea and analyze again

## 3ScientistTwo: An Expert\-Level Autonomous Research Agent

ScientistTwo is an expert\-level autonomous research agent that effectively orchestrates specialized AI agents throughout the scientific discovery pipeline\. First, ScientistTwo identifies the limitations of the human state\-of\-the\-art method regarding a given scientific problem and generates ideas to address them, filtering for those with high novelty \(§\\lx@sectionsign[3\.1](https://arxiv.org/html/2609.19644#S3.SS1)\)\. Next, ScientistTwo implements the selected ideas; to optimize computational efficiency, it validates the ideas on a benchmark subset before scaling to the full dataset \(§\\lx@sectionsign[3\.2](https://arxiv.org/html/2609.19644#S3.SS2)\)\. The agent then refines the ideas based on previous experimental results \(§\\lx@sectionsign[3\.3](https://arxiv.org/html/2609.19644#S3.SS3)\) and conducts ablation studies to analyze the sources of performance gain \(§\\lx@sectionsign[3\.4](https://arxiv.org/html/2609.19644#S3.SS4)\)\. During manuscript drafting \(§\\lx@sectionsign[3\.5](https://arxiv.org/html/2609.19644#S3.SS5)\), ScientistTwo dynamically enhances paper quality through an iterative review process, using a Review Agent to provide critical feedback and a Rebuttal Agent to conduct supplementary experiments\. Finally, in the meta\-review phase \(§\\lx@sectionsign[3\.6](https://arxiv.org/html/2609.19644#S3.SS6)\), a Meta\-Review Agent evaluates the manuscript’s overall quality and determines whether it meets the submission standards\. If further improvements are required, ScientistTwo incorporates feedback, updates idea and ablation analyses, and iterates through the pipeline until the manuscript is approved for submission\. For clarity, we abstract all stages \(see overview in Table[1](https://arxiv.org/html/2609.19644#S2.T1)\) using the Python\-style pseudocode in Listing[4](https://arxiv.org/html/2609.19644#S2.F4)\.

Problem Setup: Problem\-Driven Autonomous Discovery\.Scientific advancement in AI is inherently cumulative: modern progress relies on identifying the limitations of state\-of\-the\-art work and developing methodologies to surpass them\. To mirror this realistic research workflow, our goal is to build an autonomous research agent𝒜\\mathcal\{A\}capable of executing continuous research iterations\. Given a scientific problem𝒢\\mathcal\{G\}, the framework generates a paper𝒫\+\\mathcal\{P\}^\{\+\}paired with a reproducible codebase𝒞\+\\mathcal\{C\}^\{\+\}, which represents new frontier knowledge on the given problem\. Formally, we define this transformation as:

\(𝒫\+,𝒞\+\)=𝒜⁡\(𝒢\),where𝒜=\{𝒜1,𝒜2,…,𝒜Na\}\(\\mathcal\{P\}^\{\+\},\\mathcal\{C\}^\{\+\}\)=\\mathcal\{A\}\(\\mathcal\{G\}\),\\quad\\text\{where\}\\quad\\mathcal\{A\}=\\\{\\mathcal\{A\}\_\{1\},\\mathcal\{A\}\_\{2\},\\dots,\\mathcal\{A\}\_\{N\_\{\\mathrm\{a\}\}\}\\\}\(1\)where𝒜\\mathcal\{A\}comprisesNaN\_\{\\mathrm\{a\}\}specialized AI agents, each assigned to distinct phases of the research life cycle \(e\.g\., limitation extraction, idea generation, code modification, empirical validation, etc\.\)\. In addition,𝒫\+\\mathcal\{P\}^\{\+\}should identify and resolve key methodological or empirical bottlenecks present in𝒢\\mathcal\{G\}and𝒞\+\\mathcal\{C\}^\{\+\}should correctly implement the proposed idea while maintaining execution reproducibility and demonstrating measurable performance gains\.

### 3\.1Generating Novel Seed Ideas: Targeting Limitation Resolution

Figure 5:Generating Novel Seed Ideas\.ScientistTwo begins its research by generating novel ideas that can address the limitations of the human state\-of\-the\-art \(see§\\lx@sectionsign[3\.1](https://arxiv.org/html/2609.19644#S3.SS1)\)\.Finding Limitations\.ScientistTwo initiates the research pipeline by identifying the core limitations of the human state\-of\-the\-art method with respect to a given scientific problem𝒢\\mathcal\{G\}\. Specifically, initially, the Limitation Extractor extracts a set of limitations from𝒢\\mathcal\{G\}\. Then, the Limitation Verifier verifies whether this set is sufficient to guide novel improvements to𝒢\\mathcal\{G\}\. If the Limitation Verifier considers the set insufficient, the Limitation Extractor again identifies missing limitations or weaknesses and expands the collection\. This verification loop repeats until the Limitation Verifier confirms that all actionable limitations have been thoroughly extracted or maximum number of iterations is reached\.

Generating Novel Seed Ideas\.ScientistTwo first generates an initial ideah0h\_\{0\}specifically designed to address the identified limitations, and computes its novelty score using the Novelty Checker\. Then, starting with the initial set of candidatesℋ0=\{h0\}\\mathcal\{H\}\_\{0\}=\\\{h\_\{0\}\\\}, ScientistTwo leverages the Idea Generator Agent to iteratively expand the candidate poolℋ0\\mathcal\{H\}\_\{0\}with distinct, higher\-novelty ideas\. This process continues until a collection ofNseedN\_\{\\mathrm\{seed\}\}candidate ideas is gathered\. Finally, the seed ideas inℋ0=\{hi\}i=0Nseed−1\\mathcal\{H\}\_\{0\}=\\\{h\_\{i\}\\\}\_\{i=0\}^\{N\_\{\\mathrm\{seed\}\}\-1\}are sorted in descending order according to their novelty scores\{si\}i=0Nseed−1\\\{s\_\{i\}\\\}\_\{i=0\}^\{N\_\{\\mathrm\{seed\}\}\-1\}, ensuring thatsi≥sjs\_\{i\}\\geq s\_\{j\}wheneveri<ji<j\. This allows ScientistTwo to prioritize implementing the ideas with the highest originality\.

### 3\.2Evaluating Ideas: Experimenting from Subset to Full\-Set

Generating Baseline Results on Subset\.To establish a reliable reference, ScientistTwo first employs a Baseline Coding Agent to reproduce the primary experiments of𝒢\\mathcal\{G\}on the benchmark subset, producing baseline experimental resultsℰbase\\mathcal\{E\}\_\{\\mathrm\{base\}\}alongside a reproducible subset codebase𝒞base\\mathcal\{C\}\_\{\\mathrm\{base\}\}\.

Evaluating an Idea on the Subset\.For a candidate ideahh\(here, we select only the top\-N0N\_\{0\}highest\-ranked ideas based on\{si\}i=0N0−1\\\{s\_\{i\}\\\}\_\{i=0\}^\{N\_\{0\}\-1\}fromℋ0\\mathcal\{H\}\_\{0\}\), ScientistTwo leverages a Subset Coding Agent to implementhhby modifying𝒞base\\mathcal\{C\}\_\{\\mathrm\{base\}\}, producing the resulting logsℰsubh\\mathcal\{E\}\_\{\\mathrm\{sub\}\}^\{h\}alongside a reproducible codebase𝒞subh\\mathcal\{C\}\_\{\\mathrm\{sub\}\}^\{h\}\. Then, to evaluate performance, a Subset Critic Agent comparesℰsubh\\mathcal\{E\}\_\{\\mathrm\{sub\}\}^\{h\}againstℰbase\\mathcal\{E\}\_\{\\mathrm\{base\}\}and emits a categorical decisiondhd^\{h\}with feedbackrhr^\{h\}:

1. 1\.Bad\\mathrm\{Bad\}: If performance is substantially inferior to the baseline,hhis discarded \(dh=Badd^\{h\}=\\mathrm\{Bad\}\)\.
2. 2\.Good\\mathrm\{Good\}: Ifhhconsistently outperforms the baseline, it is approved for scale\-up \(dh=Goodd^\{h\}=\\mathrm\{Good\}\)\.
3. 3\.Engineer\\mathrm\{Engineer\}: Ifhhshows potential but requires hyperparameter tuning or code adjustments, a Subset Engineering Agent refineshhand𝒞subh\\mathcal\{C\}\_\{\\mathrm\{sub\}\}^\{h\}guided by strategyrhr^\{h\}, generating corresponding resultsℰsubh\\mathcal\{E\}\_\{\\mathrm\{sub\}\}^\{h\}\.

This refinement loop repeats untildh∈\{Good,Bad\}d^\{h\}\\in\\\{\\mathrm\{Good\},\\mathrm\{Bad\}\\\}or a budget ofNengN\_\{\\mathrm\{eng\}\}iterations is exhausted\. IfNengN\_\{\\mathrm\{eng\}\}is reached without achievingd=Goodd=\\mathrm\{Good\},hhis designated asBad\\mathrm\{Bad\}and pruned—mirroring practical research settings where unpromising avenues are abandoned after bounded optimization\.

Figure 6:Evaluating Ideas\.ScientistTwo utilizes𝒜Coder\\mathcal\{A\}\_\{\\mathrm\{Coder\}\}to implement the generated idea through subset testing, critic evaluation, engineering refinement loops, and full\-set scaling \(see§\\lx@sectionsign[3\.2](https://arxiv.org/html/2609.19644#S3.SS2)\)\.Scaling Up to the Full\-Set\.For ideas validated asGood\\mathrm\{Good\}, ScientistTwo scales the evaluation to the full benchmark\. A Full\-Set Coding Agent adapts𝒞subh\\mathcal\{C\}\_\{\\mathrm\{sub\}\}^\{h\}to run across the entire benchmark suite\. A Full\-Set Critic Agent and Full\-Set Engineer perform final validation and engineering against the full benchmark, producing final execution outputsℰfullh\\mathcal\{E\}\_\{\\mathrm\{full\}\}^\{h\}, updated codebase𝒞fullh\\mathcal\{C\}\_\{\\mathrm\{full\}\}^\{h\}, and terminal decisiondhd^\{h\}\.

Unified Coder Interface\.To streamline subsequent rounds of idea improvement, we abstract this entire idea experiment pipeline into a single high\-level Idea Implementer Agent𝒜Coder\\mathcal\{A\}\_\{\\mathrm\{Coder\}\}\. Formally,

h,ℰh,𝒞h,dh,rh=𝒜Coder​\(𝒢,h\)\.h,\\mathcal\{E\}^\{h\},\\mathcal\{C\}^\{h\},d^\{h\},r^\{h\}=\\mathcal\{A\}\_\{\\mathrm\{Coder\}\}\(\\mathcal\{G\},h\)\.\(2\)

### 3\.3Refining Ideas: Improving Ideas Using Experimental Results

In the initial evaluation round \(k=0k=0\), ScientistTwo executes the top\-N0N\_\{0\}seed ideas fromℋ0\\mathcal\{H\}\_\{0\}using𝒜Coder\\mathcal\{A\}\_\{\\mathrm\{Coder\}\}, yielding a set of execution tracesℛ0=\{\(h,ℰh,𝒞h,dh,rh\)∣h∈ℋ0\}\\mathcal\{R\}\_\{0\}=\\\{\(h,\\mathcal\{E\}^\{h\},\\mathcal\{C\}^\{h\},d^\{h\},r^\{h\}\)\\mid h\\in\\mathcal\{H\}\_\{0\}\\\}\. To continuously enhance idea quality, we propose an evolution strategy that uses these execution traces as feedback\.

Idea Evolution\.In refinement roundk≥1k\\geq 1, ScientistTwo aggregates all historic execution traces𝐑<k=⋃i=0k−1ℛi\\mathbf\{R\}\_\{<k\}=\\bigcup\_\{i=0\}^\{k\-1\}\\mathcal\{R\}\_\{i\}to generate a setℐk\\mathcal\{I\}\_\{k\}ofNkN\_\{k\}evolved ideas\. An Idea Evolver Agent𝒜Evolve\\mathcal\{A\}\_\{\\mathrm\{Evolve\}\}analyzes both successful results \(dh=Goodd^\{h\}=\\mathrm\{Good\}\) and diagnostic failure logs \(dh=Badd^\{h\}=\\mathrm\{Bad\}\) to propose refined hypotheses\.

Exploration and Exploitation\.Relying exclusively on𝒜Evolve\\mathcal\{A\}\_\{\\mathrm\{Evolve\}\}risks trapping the optimization process in local optima centered around early seed ideas\. To ensure broad coverage of the solution space, ScientistTwo complements the evolved setℐk\\mathcal\{I\}\_\{k\}withNeN\_\{e\}previously unevaluated seed ideas drawn fromℋ0\\mathcal\{H\}\_\{0\}in descending order of novelty score\. Formally, the total candidate poolℋk\\mathcal\{H\}\_\{k\}for roundkkisℐk∪ℋ0\(k\)\\mathcal\{I\}\_\{k\}\\cup\\mathcal\{H\}\_\{0\}^\{\(k\)\},whereℋ0\(k\)⊂ℋ0\\mathcal\{H\}\_\{0\}^\{\(k\)\}\\subset\\mathcal\{H\}\_\{0\}contains the nextNeN\_\{e\}highest\-ranked unevaluated seed ideas\.

Figure 7:Refining Ideas\.ScientistTwo improves ideas based on prior experimental results, ultimately selecting the best hypothesis for ablation studies and paper drafting \(see§\\lx@sectionsign[3\.3](https://arxiv.org/html/2609.19644#S3.SS3)\)\.Experimenting Ideas\.For each candidateh∈ℋkh\\in\\mathcal\{H\}\_\{k\}, ScientistTwo invokes𝒜Coder\\mathcal\{A\}\_\{\\mathrm\{Coder\}\}to evaluate the idea:

ℛk=\{\(h,ℰh,𝒞h,dh,rh\)\|h,ℰh,𝒞h,dh,rh=𝒜Coder\(𝒢,h\),h∈ℋk\}\.\\mathcal\{R\}\_\{k\}=\\left\\\{\\left\(h,\\mathcal\{E\}^\{h\},\\mathcal\{C\}^\{h\},d^\{h\},r^\{h\}\\right\)\\;\\middle\|\\;h,\\mathcal\{E\}^\{h\},\\mathcal\{C\}^\{h\},d^\{h\},r^\{h\}=\\mathcal\{A\}\_\{\\mathrm\{Coder\}\}\(\\mathcal\{G\},h\),\\;h\\in\\mathcal\{H\}\_\{k\}\\right\\\}\.\(3\)This evolutionary loop iterates untilSSsuccessful ideas \(dh=Goodd^\{h\}=\\mathrm\{Good\}\) are collected, or the maximum refinement limitKKis reached\. Formally, idea refinement terminates successfully when∑i=0k∑h∈ℋi𝕀⁡\(dh=Good\)≥S\\sum\_\{i=0\}^\{k\}\\sum\_\{h\\in\\mathcal\{H\}\_\{i\}\}\\mathbb\{I\}\(d^\{h\}=\\mathrm\{Good\}\)\\geq S, where𝕀⁡\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\. If roundKKis reached with zero successful ideas \(∑i=0K∑h∈ℋi𝕀⁡\(dh=Good\)=0\\sum\_\{i=0\}^\{K\}\\sum\_\{h\\in\\mathcal\{H\}\_\{i\}\}\\mathbb\{I\}\(d^\{h\}=\\mathrm\{Good\}\)=0\), ScientistTwo terminates the entire process\.

Selecting the Best Idea\.Upon discovering at least one successful idea, ScientistTwo selects the optimal candidatehbesth\_\{\\mathrm\{best\}\}for downstream ablation analysis\. A Selector Agent𝒜Selector\\mathcal\{A\}\_\{\\mathrm\{Selector\}\}compares performance metrics and execution logs across all validated ideas\{h∣dh=Good\}\\\{h\\mid d^\{h\}=\\mathrm\{Good\}\\\}evaluated on the full benchmark:

hbest,ℰbest,𝒞best=𝒜Selector​\(𝒢,\{\(h,ℰh,𝒞h\)\|dh=Good\}\)\.h\_\{\\mathrm\{best\}\},\\mathcal\{E\}\_\{\\mathrm\{best\}\},\\mathcal\{C\}\_\{\\mathrm\{best\}\}=\\mathcal\{A\}\_\{\\mathrm\{Selector\}\}\\left\(\\mathcal\{G\},\\left\\\{\(h,\\mathcal\{E\}^\{h\},\\mathcal\{C\}^\{h\}\)\\;\\middle\|\\;d^\{h\}=\\mathrm\{Good\}\\right\\\}\\right\)\.\(4\)

### 3\.4Ablation Studies: Analyzing Source of Gain and Refining Ideas

Once the optimal candidate ideahbesth\_\{\\mathrm\{best\}\}is selected, ScientistTwo conducts systematic component\-level ablation studies\. This phase isolates the explicit sources of empirical gain for scientific interpretation, and leverages fine\-grained ablation feedback to perform an additional round of idea refinement\.

Ablation Planning and Execution\.To evaluate individual components, an Ablation Planner Agent automatically formulates a set ofNpN\_\{p\}executable ablation plans\{p1,…,pNp\}\\\{p\_\{1\},\\dots,p\_\{N\_\{p\}\}\\\}tailored tohbesth\_\{\\mathrm\{best\}\}\. For each ablation planpip\_\{i\}, an Ablation Coding Agent modifies the validated codebase𝒞best\\mathcal\{C\}\_\{\\mathrm\{best\}\}to execute, yielding an ablation outcomecic\_\{i\}\. The aggregated ablation results are collected asℰabl=\{c1,…,cNp\}\\mathcal\{E\}\_\{\\mathrm\{abl\}\}=\\\{c\_\{1\},\\dots,c\_\{N\_\{p\}\}\\\}\.

Refining the Idea via Ablation Insights\.Emulating expert human research practices, where removing redundant or counterproductive components often yields a better method, ScientistTwo usesℰabl\\mathcal\{E\}\_\{\\mathrm\{abl\}\}to further refinehbesth\_\{\\mathrm\{best\}\}\. An Ablation Critic Agent𝒜AblCritic\\mathcal\{A\}\_\{\\mathrm\{AblCritic\}\}inspects the component breakdown to determine whetherhbesth\_\{\\mathrm\{best\}\}is optimal or requires additional modification, generatingdabl,rabld\_\{\\mathrm\{abl\}\},r\_\{\\mathrm\{abl\}\}wheredabl∈\{Good,Refine\}d\_\{\\mathrm\{abl\}\}\\in\\\{\\mathrm\{Good\},\\mathrm\{Refine\}\\\}\. Ifdabl=Goodd\_\{\\mathrm\{abl\}\}=\\mathrm\{Good\},hbesth\_\{\\mathrm\{best\}\}is finalized and passed to the paper drafting stage\. Ifdabl=Refined\_\{\\mathrm\{abl\}\}=\\mathrm\{Refine\}, ScientistTwo re\-engages the Full\-Set Engineering Agent𝒜FullEng\\mathcal\{A\}\_\{\\mathrm\{FullEng\}\}to produce a refined hypothesishnewh\_\{\\mathrm\{new\}\}, updated resultsℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}, and revised codebase𝒞new\\mathcal\{C\}\_\{\\mathrm\{new\}\}guided by critiquerablr\_\{\\mathrm\{abl\}\}\.

Robust Verification and Loop\.Because structural refinements do not guarantee improved performance, ScientistTwo verifies whetherℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}strictly outperformsℰbest\\mathcal\{E\}\_\{\\mathrm\{best\}\}\. The baseline variables are updated—\(hbest,ℰbest,𝒞best\)←\(hnew,ℰnew,𝒞new\)\(h\_\{\\mathrm\{best\}\},\\mathcal\{E\}\_\{\\mathrm\{best\}\},\\mathcal\{C\}\_\{\\mathrm\{best\}\}\)\\leftarrow\(h\_\{\\mathrm\{new\}\},\\mathcal\{E\}\_\{\\mathrm\{new\}\},\\mathcal\{C\}\_\{\\mathrm\{new\}\}\)—if and only ifℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}is preferred thanℰbest\\mathcal\{E\}\_\{\\mathrm\{best\}\}by the Result Comparison Agent\. When an update occurs, ScientistTwo re\-executes the ablation planning phase on the updated candidate\. This refinement loop repeats for a maximum ofNablN\_\{\\mathrm\{abl\}\}iterations or until𝒜AblCrit\\mathcal\{A\}\_\{\\mathrm\{AblCrit\}\}emitsdabl=Goodd\_\{\\mathrm\{abl\}\}=\\mathrm\{Good\}, ensuring a fully optimized hypothesis prior to manuscript generation\.

Figure 8:From Ablation Study to Drafting\.ScientistTwo conducts ablation studies and simulates the peer\-review and rebuttal processes\. Beyond this, ScientistTwo dynamically refines the idea and enhances paper quality using feedback from each stage \(see§\\lx@sectionsign[3\.4](https://arxiv.org/html/2609.19644#S3.SS4),§\\lx@sectionsign[3\.5](https://arxiv.org/html/2609.19644#S3.SS5),§\\lx@sectionsign[3\.6](https://arxiv.org/html/2609.19644#S3.SS6)\)\.
### 3\.5Manuscript Drafting: Simulating the Peer\-Review Process

Initial Drafting\.First, an Initial Drafter Agent𝒜Draft\\mathcal\{A\}\_\{\\mathrm\{Draft\}\}\(incorporating PaperOrchestra;[Song et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib101)\) synthesizes the selected ideahbesth\_\{\\mathrm\{best\}\}, main benchmark resultsℰbest\\mathcal\{E\}\_\{\\mathrm\{best\}\}, and component ablation studiesℰabl\\mathcal\{E\}\_\{\\mathrm\{abl\}\}into a full, conference\-formatted manuscript𝒫new\\mathcal\{P\}\_\{\\mathrm\{new\}\}\.

Enhancing the Draft via Simulated Review\-Rebuttal\.To rigorously elevate manuscript quality, ScientistTwo simulates an interactive peer\-review and rebuttal process\. A Peer\-Reviewer Agent𝒜Reviewer\\mathcal\{A\}\_\{\\mathrm\{Reviewer\}\}\(i\.e\., ScholarPeer;[Goyal et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib36)\) critically evaluates𝒫new\\mathcal\{P\}\_\{\\mathrm\{new\}\}, generating a detailed evaluationℛnew\\mathcal\{R\}\_\{\\mathrm\{new\}\}containing identified strengths, weaknesses, targeted questions, and an overall numerical scoresreview∈\[1,10\]s\_\{\\mathrm\{review\}\}\\in\[1,10\]based on the standard ICLR grading scale\. If thesreviews\_\{\\mathrm\{review\}\}is below the acceptance threshold \(e\.g\., 8\), ScientistTwo initiates an automated rebuttal stage to address reviewer concerns\. A Rebuttal Planner Agent𝒜RebPlan\\mathcal\{A\}\_\{\\mathrm\{RebPlan\}\}analyzesℛnew\\mathcal\{R\}\_\{\\mathrm\{new\}\}to formulate a set ofNtN\_\{t\}supplementary experimental tasks\{t1,…,tNt\}\\\{t\_\{1\},\\dots,t\_\{N\_\{t\}\}\\\}designed to resolve reviewer queries\. Next, a Rebuttal Coding Agent𝒜RebCoder\\mathcal\{A\}\_\{\\mathrm\{RebCoder\}\}implements and executes each planned tasktit\_\{i\}using codebase𝒞best\\mathcal\{C\}\_\{\\mathrm\{best\}\}, generating supplementary resultseie\_\{i\}, yielding the aggregated supplementary experiment resultsℰreb=\{e1,…,eNt\}\\mathcal\{E\}\_\{\\mathrm\{reb\}\}=\\\{e\_\{1\},\\dots,e\_\{N\_\{t\}\}\\\}\.

Iterative Enhancement\.A Paper Enhancer Agent𝒜Enhancer\\mathcal\{A\}\_\{\\mathrm\{Enhancer\}\}integrates theℛnew\\mathcal\{R\}\_\{\\mathrm\{new\}\}and supplementary findingsℰreb\\mathcal\{E\}\_\{\\mathrm\{reb\}\}into𝒫new\\mathcal\{P\}\_\{\\mathrm\{new\}\}, revising narrative claims and updating empirical tables and figures\. The updated manuscript is re\-evaluated by𝒜Reviewer\\mathcal\{A\}\_\{\\mathrm\{Reviewer\}\}, updatingℛnew\\mathcal\{R\}\_\{\\mathrm\{new\}\}and scoresreviews\_\{\\mathrm\{review\}\}\. This review\-rebuttal cycle repeats untilsnew≥8s\_\{\\mathrm\{new\}\}\\geq 8or the maximum budget ofNpeerN\_\{\\mathrm\{peer\}\}review iterations is reached, producing a polished, thoroughly validated final manuscript𝒫new\\mathcal\{P\}\_\{\\mathrm\{new\}\}\.

### 3\.6Meta\-Reviewing: Final Assessment and Review\-Driven Refinement

To mirror the complete lifecycle of academic publishing, ScientistTwo integrates a Meta\-Review Agent𝒜Meta\\mathcal\{A\}\_\{\\mathrm\{Meta\}\}\. Specifically,𝒜Meta\\mathcal\{A\}\_\{\\mathrm\{Meta\}\}evaluates the revised manuscript𝒫new\\mathcal\{P\}\_\{\\mathrm\{new\}\}alongside the generated peer reviewℛnew\\mathcal\{R\}\_\{\\mathrm\{new\}\}to make a final publication assessment and generate actionable strategic feedback\.

Meta\-Review Decision\.Formally,𝒜Meta\\mathcal\{A\}\_\{\\mathrm\{Meta\}\}takes the paper draft and reviewer feedback as inputs to produce a decisiondmetad\_\{\\mathrm\{meta\}\}and meta\-critiquermetar\_\{\\mathrm\{meta\}\}, wheredmeta∈\{Accept,Refine\}d\_\{\\mathrm\{meta\}\}\\in\\\{\\mathrm\{Accept\},\\mathrm\{Refine\}\\\}\. Ifdmeta=Acceptd\_\{\\mathrm\{meta\}\}=\\mathrm\{Accept\},𝒫new\\mathcal\{P\}\_\{\\mathrm\{new\}\}is judged to meet top\-tier conference standards\. ScientistTwo finalizes the process and exports the final improved paper and its codebase as𝒫\+←𝒫new\\mathcal\{P\}^\{\+\}\\leftarrow\\mathcal\{P\}\_\{\\mathrm\{new\}\}and𝒞\+←𝒞best\\mathcal\{C\}^\{\+\}\\leftarrow\\mathcal\{C\}\_\{\\mathrm\{best\}\}\.

Review\-Driven Idea Refinement\.Whendmeta=Refined\_\{\\mathrm\{meta\}\}=\\mathrm\{Refine\}, indicating that the meta\-reviewer identified a critical algorithmic or empirical weakness, ScientistTwo initiates a deep idea refinement phase\. Using meta\-critiquermetar\_\{\\mathrm\{meta\}\}as guidance, the Full\-Set Engineering Agent𝒜FullEng\\mathcal\{A\}\_\{\\mathrm\{FullEng\}\}updateshbesth\_\{\\mathrm\{best\}\}to construct a revised hypothesishnewh\_\{\\mathrm\{new\}\}, updated resultsℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}, and modified codebase𝒞new\\mathcal\{C\}\_\{\\mathrm\{new\}\}\.

Verification and Re\-Drafting\.To ensure that the meta\-review modification yields true scientific progression, ScientistTwo re\-evaluatesℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}againstℰbest\\mathcal\{E\}\_\{\\mathrm\{best\}\}using the Result Comparison Agent\. Ifℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}is verified as strictly superior , ScientistTwo updates the core state\(hbest,ℰbest,𝒞best\)←\(hnew,ℰnew,𝒞new\)\(h\_\{\\mathrm\{best\}\},\\mathcal\{E\}\_\{\\mathrm\{best\}\},\\mathcal\{C\}\_\{\\mathrm\{best\}\}\)\\leftarrow\(h\_\{\\mathrm\{new\}\},\\mathcal\{E\}\_\{\\mathrm\{new\}\},\\mathcal\{C\}\_\{\\mathrm\{new\}\}\)\. Because the fundamental idea has changed, ScientistTwo re\-executes downstream ablation planning, ablation execution, manuscript re\-drafting, and simulated peer\-review cycles\. Conversely, ifℰnew\\mathcal\{E\}\_\{\\mathrm\{new\}\}fails to outperform the baseline, the refinement is discarded, and ScientistTwo terminates the process using the previous best outputs \(𝒫\+←𝒫new,𝒞\+←𝒞best\\mathcal\{P\}^\{\+\}\\leftarrow\\mathcal\{P\}\_\{\\mathrm\{new\}\},\\mathcal\{C\}^\{\+\}\\leftarrow\\mathcal\{C\}\_\{\\mathrm\{best\}\}\)\. This meta\-refinement loop repeats for a maximum ofNmetaN\_\{\\mathrm\{meta\}\}iterations or untildmeta=Acceptd\_\{\\mathrm\{meta\}\}=\\mathrm\{Accept\}, yielding a rigorously validated final contribution\.

## 4Experiments

In this section, we empirically validate the effectiveness of ScientistTwo\. Specifically, in§\\lx@sectionsign[4\.1](https://arxiv.org/html/2609.19644#S4.SS1), we present quantitative and qualitative comparisons against existing autonomous research agents\. In§\\lx@sectionsign[4\.2](https://arxiv.org/html/2609.19644#S4.SS2), we conduct component\-wise ablation studies to evaluate each constituent part of ScientistTwo\. In§\\lx@sectionsign[4\.3](https://arxiv.org/html/2609.19644#S4.SS3), we provide a discussion including cost analysis and case studies\.

Common Setup\.We evaluate the effectiveness of ScientistTwo on 107 scientific problems drawn from top machine learning venues, including ICLR, ICML, and NeurIPS \(see Appendix[A\.1](https://arxiv.org/html/2609.19644#A1.SS1)for details\)\. All experiments are conducted using Gemini 3\.6 Flash and Claude Opus 4\.8, unless otherwise specified\. For evaluation, we primarily report review scores from automated AI review agents: ScholarPeer\([Goyal et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib36)\)and Stanford Agentic Reviewer111[https://paperreview\.ai/](https://paperreview.ai/)\. ScholarPeer serves as an in\-distribution evaluation, as it is also used to refine the draft quality generated by ScientistTwo\. Conversely, Stanford Agentic Reviewer serves as a held\-out evaluator that was unseen during development by both the baselines and our method\. Full implementation details and configurations are provided in Appendix[A\.2](https://arxiv.org/html/2609.19644#A1.SS2)\.

### 4\.1Main Results

Table 2:Comparison with autonomous research agents\.We report the average review ratings, standard deviations \(1–10 scale\), and acceptance rates \(%\) given by ScholarPeer and Stanford Agentic Reviewer\. The column “\# Papers” represents the number of publicly released AI\-generated papers by each autonomous research agent used to aggregate performance\. Bold indicates the best performance\.Quantitative Comparison\.As shown in Table[2](https://arxiv.org/html/2609.19644#S4.T2), ScientistTwo achieves a 91\.9% acceptance rate under ScholarPeer \(nearly doubling the review score from ScientistOne’s 3\.8 to 7\.5\)\. More importantly, while all baselines fail to achieve acceptance from the Stanford Agentic Reviewer, ScientistTwo is the only agent that produces research where 72\.1% of generated papers meet high acceptance standards\. This result shows that while current autonomous research agents cannot produce expert\-level research results, ScientistTwo possesses such capabilities\.

![Refer to caption](https://arxiv.org/html/2609.19644v1/scientistone_8.png)![Refer to caption](https://arxiv.org/html/2609.19644v1/scientistone_9.png)\(a\) Experiment section generated by ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib77)\)![Refer to caption](https://arxiv.org/html/2609.19644v1/ours_7.png)![Refer to caption](https://arxiv.org/html/2609.19644v1/ours_8.png)\(b\) Experiment section generated byScientistTwo \(Ours\)

Figure 9:Qualitative comparison\.Side\-by\-side comparison of the experiment section generated by ScientistTwo against the highest\-reviewed paper from ScientistOne\([Meng et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib77)\)\.Qualitative Comparison\.To assess the depth and completeness of the generated manuscripts, we compare ScientistTwo against the strongest baseline, ScientistOne\. As illustrated in Figure[9](https://arxiv.org/html/2609.19644#S4.F9), ScientistTwo demonstrates substantially broader experimental coverage and analytical rigor\. While ScientistOne generates only 4 figures and 2 tables—evaluating a single metric on an isolated benchmark with limited baselines—ScientistTwo produces 9 figures and 12 tables \(including the Appendix\)\. Furthermore, ScientistTwo reports multi\-metric evaluations across diverse datasets and incorporates an extensive set of baseline comparisons, reflecting a publication\-ready experimental design\.

Table 3:Comparison with human\-author papers and AI\-generated papers\.We report the average review ratings, standard deviations \(1–10 scale\), and acceptance rates \(%\) evaluated by ScholarPeer\([Goyal et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib36)\)and Stanford Agentic Reviewer\. The top section evaluates papers accepted at each venue, which serve as inputs for ScientistTwo\. The bottom section reports the evaluation of papers generated by ScientistTwo using the accepted papers from the corresponding venue\. The column “\# Papers” indicates the number of accepted reference papers evaluated or the number of papers successfully generated by ScientistTwo\.†\\daggerdenotes AI\-generated papers\.ScholarPeerStanford Agentic ReviewerVenue\# PapersAvg\. RatingAccept RateAvg\. RatingAccept RateAgent4Science 2025 Accepted†43\.0±\\pm0\.00\.03\.8±\\pm0\.40\.0ICLR 2026 Accepted56\.8±\\pm1\.660\.05\.2±\\pm0\.760\.0NeurIPS 2025 Accepted386\.2±\\pm1\.965\.85\.5±\\pm0\.776\.3ICML 2026 Spotlight646\.9±\\pm1\.579\.76\.1±\\pm0\.596\.9ScientistTwo \(Ours\)ICLR 2026†4/57\.0±\\pm1\.2100\.05\.4±\\pm0\.275\.0NeurIPS 2025†33/387\.3±\\pm1\.787\.95\.6±\\pm0\.775\.8ICML 2026 Spotlight†49/647\.6±\\pm1\.093\.95\.7±\\pm0\.669\.4Overall†86/1077\.5±\\pm1\.391\.95\.7±\\pm0\.672\.1

Comparison with Human Researchers\.We evaluate the research quality produced by ScientistTwo against accepted publications\. Specifically, we use ScholarPeer and the Stanford Agentic Reviewer to evaluate papers accepted at NeurIPS 2025, ICLR 2026, and ICML 2026 \(including spotlight presentations\), whose problem specifications and codebases serve as benchmark tasks for ScientistTwo\. As shown in Table[3](https://arxiv.org/html/2609.19644#S4.T3), ScientistTwo successfully executes research on 86 out of 107 problems, achieving an 80\.4% success rate\. Furthermore, papers generated by ScientistTwo surpass the average scores of accepted papers at ICLR 2026 and NeurIPS 2025 under both ScholarPeer and Stanford Agentic Reviewer\. While ScientistTwo does not yet achieve spotlight\-level quality, these results still demonstrate that it functions as an expert\-level research agent capable of producing manuscripts that meet the acceptance threshold of top\-tier AI venues\. Finally, we include comparisons with AI\-generated papers accepted at Agent4Science 2025—the first venue dedicated to AI\-generated research—to demonstrate that prior systems were incapable of meeting top conference standards\.

Table 4:Comparison with AutoSOTA\.We report the number of tasks successfully done by each framework, and performance gain \(%\) across successful cases\. For NeurIPS 2025 and ICLR 2026, evaluations use the same set of 33 and 4 input papers, respectively\. Bold indicates the best score\.Comparison with AutoSOTA\.While AutoSOTA\([Li et al\., 2026d](https://arxiv.org/html/2609.19644#bib.bib65)\)operates in a distinct setting, i\.e\., modifying existing codebases to optimize a single scalar metric, ScientistTwo is designed for end\-to\-end scientific paper generation\. For completeness, we provide a comparison based on the performance gains reported over human state\-of\-the\-art baselines\. To extract these gains systematically, we parse the main tables for 10 times using Gemini 3\.6 Flash and averaged them\. As shown in Table[4](https://arxiv.org/html/2609.19644#S4.T4), ScientistTwo shows superior performance to AutoSOTA in terms of average and median improvement across average of all tasks and NeurIPS 2025 tasks, but performance drops slightly in the ICLR 2026 tasks\. We hypothesize that this advantage stems from ScientistTwo proposing novel methodological innovations to overcome baseline limitations, whereas AutoSOTA relies primarily on searching within the existing hyperparameter and execution space to improve the given method\. Appendix[B](https://arxiv.org/html/2609.19644#A2)examines the five ICLR 2026 papers one by one and supports this account: none of AutoSOTA’s five changes introduces a new algorithmic component, and each is a configuration\-level edit of at most a few lines\.

### 4\.2Ablation Studies

Here, we evaluate on the 49 target problems sourced from ICML 2026 Spotlight papers\.

Figure 10:Ablation on idea improvement\.\(a\)Relative performance gains are largest during the initial improvement rounds and steadily diminish in later stages\.\(b\)Top\-performing ideas emerge early in the refinement process\. Across iterations, evolved ideas are selected for the majority of tasks\.Effectiveness of Idea Evolution\.As shown in Figure[10](https://arxiv.org/html/2609.19644#S4.F10)\(a\), the relative gain over the human state\-of\-the\-art baseline steadily increases as ScientistTwo proceeds through iterative idea refinement\. In particular, the magnitude of improvement is most pronounced during the early refinement stages\.

Balancing Exploration and Exploitation\.As shown in Figure[10](https://arxiv.org/html/2609.19644#S4.F10)\(b\), the majority of the top\-performing ideas chosen by the Selector Agent are identified in the early stages\. This indicates that the initial seed ideas generated by ScientistTwo to overcome the limitations of human state\-of\-the\-art methods are already well\-formed and competitive\. Conversely, when seed ideas yield marginal gains or fail, the Idea Evolver Agent leverages these experimental traces to evolve them into stronger hypotheses that overcome baseline shortcomings\. Consequently, as refinement progresses, newly selected best ideas are substantially more likely to be drawn from evolved candidates rather than original seed ideas \(as evidenced by the dominant share of the blue bars after the initial round\)\.

Table 5:Effectiveness of peer\-review simulation\.Average review ratings \(1–10 scale\) and acceptance rates \(%\) evaluated by ScholarPeer and Stanford Agentic Reviewer\. Best score is bolded\.Effectiveness of the Rebuttal Agent\.During the simulated dynamic peer\-review phase, ScientistTwo leverages the Rebuttal Agent to conduct supplementary experiments and uses the findings to enhance the draft\. As shown in Table[5](https://arxiv.org/html/2609.19644#S4.T5), this rebuttal procedure effectively addresses the weaknesses flagged by ScholarPeer\. While it is not surprising that ScientistTwo gradually achieves good scores on ScholarPeer, as we utilize ScholarPeer reviews, we found that our peer\-review simulation framework is generalized across Review Agent\. Specifically, incorporating ScholarPeer reviews also improves acceptance rate from Stanford Agentic Reviewer, with this effect being particularly pronounced in the first review round\. Furthermore, the first and second rows of Table[5](https://arxiv.org/html/2609.19644#S4.T5)demonstrate that even without the Rebuttal Agent, ScientistTwo produces initial drafts of significantly higher quality than the prior state\-of\-the\-art baseline, ScientistOne\. This indicates that the preceding autonomous research cycle—specifically the holistic benchmark reasoning, idea refinement, and ablation\-driven hypothesis refinement—is highly effective for drafting a robust manuscript, achieving a nearly 50% acceptance rate when evaluated by both ScholarPeer and Stanford Agentic Reviewer\.

Table 6:Ablation study on review\-driven idea refinement\.We report the main benchmark results generated by ScientistTwo using[Kwon et al\. \(2026\)](https://arxiv.org/html/2609.19644#bib.bib57)as input\. Specifically, the unlearning performance on TOFU \(𝚏𝚘𝚛𝚐𝚎𝚝𝟷𝟶\\mathtt\{forget10\}split, Llama\-3\.2\-1B\-Instruct\)\. Bold text indicates the best performance\.Effectiveness of Review\-Driven Idea Refinement\.In the final stage, ScientistTwo employs the Meta\-Review Agent to evaluate whether the generated paper meets the acceptance standards of a top\-tier venue\. If the agent indicates that further improvement is required, ScientistTwo refines the underlying idea by incorporating the review as feedback\. As shown in Table[6](https://arxiv.org/html/2609.19644#S4.T6), this refinement process enables ScientistTwo to generate more effective ideas that advance beyond the current knowledge frontier\. For instance, prior to review\-driven refinement, ScientistTwo already established a new state\-of\-the\-art method, LFR\-Engram, outperforming the human baseline Engram\([Kwon et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib57)\)\. However, the Meta\-Review Agent deemed LFR\-Engram insufficient to meet expert standards\. By leveraging the reviewer feedback for refinement, ScientistTwo subsequently synthesized FCD\-Engram, which consistently outperforms LFR\-Engram\. This case study demonstrates that review\-driven refinement is essential for producing more novel, high\-impact research ideas\.

Table 7:Integrity Audit results\.Evaluation of paper and code integrity following[Meng et al\. \(2026\)](https://arxiv.org/html/2609.19644#bib.bib77)\.Score Verif\.\(↑\\uparrow\) indicates that every claimed results are reproducible\.Spec\. Violat\.\(↓\\downarrow\) indicates code specification violations or reward hacking issues\.Ref\. Verif\.\(↓\\downarrow\) checks for hallucinated references\.Method\-Code\(↑\\uparrow\) measures alignment between paper methodology and code implementation\.CoE Integrity Audit\.The CoE Integrity Audit\([Meng et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib77)\)is a post\-hoc evaluation framework that verifies whether claims in a generated paper are supported by its artifacts: code, empirical outputs, and bibliography\. It comprises four integrity checks: \(1\)Score verification, which evaluates codebase reproducibility by comparing reported scores against those obtained from re\-executing the repository; \(2\)Specification compliance, which ensures that solution code adheres strictly to task rules without reward hacking; \(3\)Reference verification, which guarantees that the bibliography contains no hallucinated citations; and \(4\)Method\-code alignment, which ensures that the paper faithfully and accurately describes the codebase implementation\. Following[Meng et al\. \(2026\)](https://arxiv.org/html/2609.19644#bib.bib77), AI\-generated papers and their corresponding codebases must be rigorously verifiable across all four dimensions\.

To guarantee these properties, we introduce three dedicated refinement agents alongside careful agent prompt design\. For score verification, the Coding Agent is prompted during the experimentation phase to output self\-contained, reproducible scripts and execution instructions, ensuring full reproducibility without requiring additional post\-hoc refinement\. For specification compliance, we introduce a validation filter that uses the Coding Agent to detect and discard rule\-violating solutions immediately after experimentation\. For reference verification, a search\-augmented LLM identifies hallucinated citations, enabling the Writer Agent to ground and correct the bibliography using live search results\. Finally, for method\-code alignment, the Coding Agent audits the repository against the manuscript to produce an audit report, which the Writer Agent then uses to rectify any discrepancies in the method section\. As shown in Table[7](https://arxiv.org/html/2609.19644#S4.T7), ScientistTwo faithfully passes all four audits; removing these refinement agents leads to sporadic audit failures, highlighting ScientistTwo’s capability to generating reliable and verifiable scientific contributions\.222ScientistTwo successfully completes an additional task without a refinement agent for specification compliance \(see the first row of Table[7](https://arxiv.org/html/2609.19644#S4.T7)\)\. However, since the corresponding codebase contains reward hacking, it must be filtered\.

Table 8:Coding Agent Generalizability\.We report success rate \(SR; %\) on 5 tasks sourced from ICLR 2026 accepted papers, along with average performance gain \(%\), review ratings \(1–10 scale\), and acceptance rates \(%\) evaluated by ScholarPeer and Stanford Agentic Reviewer across successful cases\. Claude Code and Antigravity are powered by Opus 4\.8 and Gemini 3\.8 Flash, respectively\.Generalizability Across Coding Agents\.Throughout our primary experiments, we leverage Claude Code with Opus 4\.8 whenever coding capabilities are required\. To evaluate generalizability across agent backends—a core component for testing generated ideas across multiple benchmarks—we replace Claude Code with Antigravity powered by Gemini 3\.8 Flash, and run ScientistTwo on 5 tasks sourced from ICLR 2026 accepted papers\. As shown in Table[8](https://arxiv.org/html/2609.19644#S4.T8), ScientistTwo remains capable of generating state\-of\-the\-art methods across different coding agents\. Furthermore, the codebases generated by Antigravity remain fully reproducible, exhibit no specification violations, and faithfully align with the designs proposed by ScientistTwo\.

### 4\.3Discussion

Figure 11:Computational cost of ScientistTwo\.\(a\)ScientistTwo requires 2\.5 days on average to improve a single paper\.\(b\)Idea refinement accounts for the majority of overall time and computational cost, primarily driven by iterative idea evolution and experimental execution\.Cost Analysis\.For cost analysis, we analyze on the 33 target problems sourced from NeurIPS 2025 papers\. As shown in Figure[11](https://arxiv.org/html/2609.19644#S4.F11)\(a\), ScientistTwo requires an average of 2–3 days to complete the entire research cycle, demonstrating significantly faster execution compared to human researchers and substantially accelerating the exploration of new research directions\. Figure[11](https://arxiv.org/html/2609.19644#S4.F11)\(b\) highlights that the majority of execution time is concentrated in the Idea Refinement, Dynamic Peer\-Review, and Meta\-Review stages\. This overhead stems from iterative benchmark evaluation and code execution—traditionally the most time\-consuming phase in empirical research\. Across the full pipeline, ScientistTwo incurs an average cost of $3765, including token usage costs and virtual machine costs \(Figure[11](https://arxiv.org/html/2609.19644#S4.F11)\(b\), right\)\. Cost expenditure increases proportionally to execution time, which is primarily due to long\-term experimentation and the large amount of context and generation requirements\.

Table 9:Iterative Frontier Expansion\.Sequential discovery results where ScientistTwo uses its own newly discovered state\-of\-the\-art method as context to drive further algorithmic improvements\.Iterative Frontier Expansion\.Having demonstrated that ScientistTwo surpasses human\-designed baselines, a natural follow\-up question is whether ScientistTwo can iteratively compound these gains—namely, can it use its own newly discovered state\-of\-the\-art solutions as priors to discover even better methods? As a positive answer, as summarized in Table[9](https://arxiv.org/html/2609.19644#S4.T9), ScientistTwo initially discovers VD\-STrans, achieving a 10\.9% improvement over the existing state of the art for incremental Byte Pair Encoding \(BPE\) tokenization\([Jiang and Gong, 2026](https://arxiv.org/html/2609.19644#bib.bib50)\)\. When VD\-STrans is subsequently provided as context in the next discovery cycle, ScientistTwo generates BXT\-Transducer, yielding an additional 9\.6% relative improvement over VD\-STrans\. Moreover, in the third iteration, ScientistTwo proposes SBR\-Transducer, again yielding an additional 8\.2% improvement over BXT\-Transducer\. These results demonstrate that ScientistTwo is capable of compounding discovery, sequentially pushing algorithmic performance beyond both human baselines and its own solutions\.

Table 10:Human Reviewers Evaluation\.Standalone scores represent mean absolute ratings for ScientistTwo on a 1–5 Likert scale \(\> 3\.0 indicates positive endorsement\)\. Comparative scores represent relative preference between human\-written and ScientistTwo\-generated papers on a 1–5 scale \(3\.0 = Parity, \> 3\.0 favors ScientistTwo\)\.Human Evaluation\.To evaluate the research quality of manuscripts produced by ScientistTwo, we conducted a human expert evaluation across 33 papers generated from NeurIPS\-derived problems, evaluated by 9 experienced human reviewers\. Reviewers first scored each ScientistTwo\-generated paper standalone on a 1–5 Likert scale across six research dimensions\. As shown in Table[10](https://arxiv.org/html/2609.19644#S4.T10), ScientistTwo consistently received positive endorsements across all evaluated criteria, achieving notable strengths in ablation design\. Furthermore, in pairwise comparative assessments against accepted human\-authored papers, ScientistTwo achieved overall parity and was favored in experimental execution—specifically in benchmark breadth\. While human researchers retained a slight advantage in methodological rigor, the results demonstrate that ScientistTwo is capable of producing publication\-grade manuscripts competitive with human\-authored papers at top\-tier venues\.

![Refer to caption](https://arxiv.org/html/2609.19644v1/figures/fig_framework_architecture.jpg)Figure 12:Case Study: Main Figure of DynaSpec\-RAG developed by ScientistTwo\.Overall architecture of the DynaSpec\-RAG framework for zero\-shot time series forecasting\.Table 11:Case Study: Main Results of DynaSpec\-RAG developed by ScientistTwo\.Zero\-shot long\-term forecasting performance \(T=512,L=64T=512,L=64\) across seven benchmarks, reported as MSE / MAE \(lower is better\)\. Best results are inbold, second\-best areunderlined\. “—” indicates datasets present in a model’s pretraining corpus for which zero\-shot numbers are omitted\.Case Study\.Prior retrieval augmented forecasting methods\([Ning et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib83)\)attempt to improve future predictions by fetching similar historical trajectories from external databases, but they treat these retrieved curves as single, indivisible blocks\. Such approach creates three critical bottlenecks: it introduces unnatural jumps right at the boundary where current observations end and the forecast begins, it tangles steady long\-term trends with erratic short\-term ripples, and it frequently degrades performance when the retrieved data is noisy or irrelevant\. To overcome these issues, ScientistTwo developed DynaSpec\-RAG, a lightweight framework that refines forecasts dynamically\. DynaSpec\-RAG resolves boundary jumps by smoothly anchoring retrieved curves to the final known observation, uses real\-FFT Fourier decomposition to split trajectories into clean macro\-trends and seasonal details, deploys a fine\-grained gating network that evaluates how much to trust each frequency band at every individual time step, and introduces a validation safety switch that dials retrieval influence down to zero whenever it fails to add value \(see Figure[12](https://arxiv.org/html/2609.19644#S4.F12)\)\.

The novelty of DynaSpec\-RAG lies in replacing crude copy\-pasting with a surgical, frequency\-aware filter that actively anticipates and guards against bad data\. Rather than assuming all retrieved history is helpful, it independently regulates trust across different temporal scales and incorporates an explicit*do\-no\-harm*safety fallback\. This design provides high practical contribution: because it operates strictly on the output space of frozen time\-series foundation models, it requires no costly backbone retraining and introduces only 0\.27M trainable parameters\. Despite this tiny footprint, it consistently surpasses strong baselines across standard benchmarks \(see Table[11](https://arxiv.org/html/2609.19644#S4.T11)\) and transfers zero\-shot to completely unseen multi\-domain datasets without retraining\. For our evaluation of ScientistTwo, this case demonstrates that the agent does not simply run shallow trial\-and\-error experiments; it autonomously identifies root failure modes in existing literature and invests mathematically sound, parameter\-efficient solutions that earn acceptance from competitive peer review\. In Appendices[C](https://arxiv.org/html/2609.19644#A3)and[D](https://arxiv.org/html/2609.19644#A4), we provide extensive qualitative artifacts, including generated ideas and limitations, evaluation and ablation reports, Critic Agent feedback, CoE audit reports, and a complete generated paper\.

## 5Conclusion

In this paper, we introduce ScientistTwo, an autonomous multi\-agent framework designed to advance the frontier of scientific discovery without human intervention\. By coupling holistic benchmark evaluation and ablation\-driven hypothesis refinement with a closed\-loop peer\-review and rebuttal engine, ScientistTwo emulates the empirical rigor of expert human researchers\. Across an extensive evaluation on 107 competitive research challenges from top\-tier venues \(ICLR, ICML, NeurIPS\), ScientistTwo successfully advanced 80\.4% of target problems, achieving an average relative improvement of 25\.2% over human state\-of\-the\-art baselines\. Furthermore, the resulting manuscripts and codebases consistently met top\-tier conference acceptance thresholds while faithfully passing rigorous multi\-dimensional integrity audits\. These results demonstrate that ScientistTwo moves beyond narrow metric optimization to generate verifiable, publication\-grade scientific contributions\.

Limitations\.While ScientistTwo consistently exceeds the acceptance threshold for standard conference publications \(surpassing average scores from ICLR 2026 and NeurIPS 2025\), it does not yet consistently achieve the caliber of spotlight or oral presentations that introduce paradigm\-shifting conceptual breakthroughs\. Future work will focus on expanding multi\-agent exploration beyond local algorithmic refinements toward discovering fundamentally new theoretical formulations\.

It should also be noted that ScientistTwo costs approximately $3,800 to execute a single task\. This may be a limiting factor in the widespread use of ScientistTwo by academic laboratories or independent researchers\. On the other hand, improving the system’s cost\-efficiency, such as by replacing proprietary models with open\-source models, is an interesting direction for future research\.

## References

- Abe et al\. \(2026\)K\. Abe, M\. Sakamoto, K\. Ariu, and A\. Iwasaki\.Asymmetric perturbation in solving bilinear saddle\-point optimization\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Arora et al\. \(2026\)A\. Arora, Z\. Wu, J\. Steinhardt, and S\. Schwettmann\.Language model circuits are sparse in the neuron basis\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Balazadeh et al\. \(2025\)V\. Balazadeh, H\. Kamkari, V\. Thomas, J\. Ma, B\. Li, J\. C\. Cresswell, and R\. Krishnan\.CausalPFN: Amortized causal effect estimation via in\-context learning\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Behnam and Wang \(2025\)A\. Behnam and B\. Wang\.Measure\-theoretic anti\-causal representation learning\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Benard \(2025\)C\. Benard\.Tree ensemble explainability through the hoeffding functional decomposition and treeHFD algorithm\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Bhalla et al\. \(2026\)U\. Bhalla, A\. Oesterling, C\. M\. Verdun, H\. Lakkaraju, and F\. Calmon\.Temporal sparse autoencoders: Leveraging the sequential nature of language for interpretability\.In*The Fourteenth International Conference on Learning Representations*, 2026\.
- Butler et al\. \(2025\)L\. Butler, A\. Agarwal, J\. S\. Kang, Y\. E\. Erginbas, B\. Yu, and K\. Ramchandran\.ProxySPEX: Inference\-efficient interpretability via sparse feature interactions in LLMs\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Candogan and Foussoul \(2026\)O\. Candogan and A\. Foussoul\.Deep flow networks\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Chen et al\. \(2025a\)B\. Chen, Z\. Zhou, L\. Peng, and Z\. Wang\.Balanced active inference\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025a\.
- Chen et al\. \(2026a\)D\. Chen, A\. Manolache, M\. Niepert, and K\. Borgwardt\.Protein fold classification at scale: Benchmarking and pretraining\.In*Forty\-third International Conference on Machine Learning*, 2026a\.
- Chen et al\. \(2026b\)J\. Chen, Y\. Luo, and L\. Pan\.Mechanistic data attribution: Tracing the training origins of interpretable LLM units\.In*Forty\-third International Conference on Machine Learning*, 2026b\.
- Chen et al\. \(2026c\)J\. Chen, B\. D\. Mishra, J\. Nam, R\. Meng, T\. Pfister, and J\. Yoon\.Mars: Modular agent with reflective search for automated ai research\.*arXiv preprint arXiv:2602\.02660*, 2026c\.
- Chen et al\. \(2025b\)M\. Chen, Z\. Cui, X\. Liu, J\. Xiang, C\. Zheng, J\. Li, and E\. Shlizerman\.SAVVY: Spatial awareness via audio\-visual LLMs through seeing and hearing\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025b\.
- Chen et al\. \(2026d\)M\. Chen, T\. Berrett, T\. Damoulas, and M\. Caprio\.Bulk\-calibrated credal ambiguity sets: Fast, tractable decision making under out\-of\-sample contamination\.In*Forty\-third International Conference on Machine Learning*, 2026d\.
- Chen et al\. \(2026e\)P\. Chen, H\. Zhao, X\. Tang, Y\. Wang, and S\. Deng\.Towards optimal robustness in learning\-augmented paging\.In*Forty\-third International Conference on Machine Learning*, 2026e\.
- Chi et al\. \(2026\)H\. Chi, Q\. Wu, Z\. Zhou, J\. Light, E\. Dodwell, and Y\. Ma\.Unifying and optimizing data values for selection via sequential decision\-making\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Choi et al\. \(2026\)S\. Choi, S\. Mittal, V\. Elvira, J\. Park, and E\. S\. Whitammer\.Reinforced sequential monte carlo for amortised sampling\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Davoodi et al\. \(2026\)A\. G\. Davoodi, N\. Rezazadeh, S\. P\. M\. Davoudi, and P\. Pezeshkpour\.Geometry\-aware decoding with wasserstein\-regularized truncation and mass penalties for large language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Dorman et al\. \(2026\)J\. M\. Dorman, E\. Gillman, D\. C\. Rose, J\. F\. Mair, and J\. P\. Garrahan\.Rare event analysis of large language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Du et al\. \(2026\)Z\. Du, J\. Zhao, and B\. Li\.On the difficulty of learning a meta\-network for training data selection\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Durasinovic et al\. \(2026\)S\. Durasinovic, J\. B\. Lasserre, and V\. Magron\.Mixtures closest to a given measure: A semidefinite programming approach\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Eldesokey et al\. \(2025\)A\. Eldesokey, A\. Cvejić, B\. Ghanem, and P\. Wonka\.Mind\-the\-glitch: Visual correspondence for detecting inconsistencies in subject\-driven generation\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Erata et al\. \(2026\)F\. Erata, O\. Paradise, T\. Typaldos, T\. Antonopoulos, T\. Nguyen, S\. Goldwasser, and R\. Piskac\.Learning randomized reductions\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Fay et al\. \(2025\)Y\. L\. Fay, N\. Chopin, and S\. Barthelmé\.Least squares variational inference\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Feng et al\. \(2025\)Y\. Feng, J\. Li, J\. Hu, Y\. Zhang, L\. Tan, and J\. Ji\.MDReID: Modality\-decoupled learning for any\-to\-any multi\-modal object re\-identification\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Ferrere et al\. \(2026\)B\. Ferrere, N\. Bousquet, F\. Gamboa, J\.\-M\. Loubes, and J\. Muré\.Exact functional ANOVA decomposition for categorical inputs models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Flores et al\. \(2025\)G\. Flores, A\. H\. Smith, J\. Fukuyama, and A\. C\. Wilson\.Aligning evaluation with clinical priorities: Calibration, label shift, and error costs\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Forstenhäusler et al\. \(2025\)M\. Forstenhäusler, D\. Külzer, C\. Anagnostopoulos, S\. A\. P\. Parambath, and N\. Weber\.STaRFormer: Semi\-supervised task\-informed representation learning via dynamic attention\-based regional masking for sequential data\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Fu et al\. \(2026\)J\. Fu, Y\. Jiang, P\. WU, C\. Liu, J\. T\. Zhou, and X\. Yang\.Rethinking LLM ensembling from the perspective of mixture models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Fu et al\. \(2025\)Y\. Fu, F\. Wang, Z\. Shao, B\. Diao, L\. Wu, Z\. An, C\. Yu, Y\. Li, and Y\. Xu\.On the integration of spatial\-temporal knowledge: A lightweight approach to atmospheric time series forecasting\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Gan and Isola \(2026\)Y\. Gan and P\. Isola\.Neural thickets: Diverse task experts are dense around pretrained weights\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Gao et al\. \(2026\)L\. Gao, Z\. Jia, Z\. Xing, W\. Sun, H\. Duan, G\. Zhai, and X\. Min\.EEmo\-logic: A unified dataset and multi\-stage framework for comprehensive image\-evoked emotion assessment\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Gao et al\. \(2025\)Y\. Gao, Q\. Yan, Y\. Leng, and R\. Liao\.Neural MJD: Neural non\-stationary merton jump diffusion for time series prediction\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Geiping et al\. \(2026\)J\. Geiping, X\. Yang, and G\. Su\.Efficient parallel samplers for recurrent\-depth models and their connection to diffusion language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Geng et al\. \(2025\)Z\. Geng, M\. Deng, X\. Bai, J\. Z\. Kolter, and K\. He\.Mean flows for one\-step generative modeling\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Goyal et al\. \(2026\)P\. Goyal, M\. Parmar, Y\. Song, H\. Palangi, T\. Pfister, and J\. Yoon\.Scholarpeer: A context\-aware multi\-agent framework for automated peer review\.*arXiv preprint arXiv:2601\.22638*, 2026\.
- Grebe et al\. \(2026\)J\. H\. Grebe, T\. Braun, A\. Rohrbach, and M\. Rohrbach\.GEM: Geometric erasure by contrastive velocity matching in rectified flows\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Grontas et al\. \(2026\)P\. D\. Grontas, A\. Terpin, E\. C\. Balta, R\. D’Andrea, and J\. Lygeros\.Pinet: Optimizing hard\-constrained neural networks with orthogonal projection layers\.In*The Fourteenth International Conference on Learning Representations*, 2026\.
- Hashemi et al\. \(2025\)B\. Hashemi, K\. Pasque, C\. Teska, and R\. Yoshida\.Tropical attention: Neural algorithmic reasoning for combinatorial algorithms\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- He et al\. \(2025\)H\. He, K\. Yi, Y\. Ma, Q\. Zhang, Z\. Niu, and G\. Pang\.SEMPO: Lightweight foundation models for time series forecasting\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Henaff et al\. \(2026\)M\. Henaff, S\. Fujimoto, M\. Matthews, and M\. Rabbat\.Scalable option learning in high\-throughput environments\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Heyman and Vandeputte \(2026\)G\. Heyman and F\. Vandeputte\.Steer like the LLM: Activation steering that mimics prompting\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Hong et al\. \(2025\)S\. Hong, Y\. Lin, B\. Liu, B\. Liu, B\. Wu, C\. Zhang, D\. Li, J\. Chen, J\. Zhang, J\. Wang, et al\.Data interpreter: An llm agent for data science\.In*Findings of the Association for Computational Linguistics: ACL 2025*, 2025\.
- Howard et al\. \(2026\)S\. Howard, N\. Nüsken, and J\. Pidstrigach\.Control consistency losses for diffusion bridges\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Huang et al\. \(2023\)Q\. Huang, J\. Vora, P\. Liang, and J\. Leskovec\.Mlagentbench: Evaluating language agents on machine learning experimentation\.*arXiv preprint arXiv:2310\.03302*, 2023\.
- Huang et al\. \(2024\)Y\. Huang, J\. Luo, Y\. Yu, Y\. Zhang, F\. Lei, Y\. Wei, S\. He, L\. Huang, X\. Liu, J\. Zhao, et al\.Da\-code: Agent data science code generation benchmark for large language models\.In*Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing*, 2024\.
- Huang et al\. \(2026\)Y\. Huang, W\. He, and Z\.\-X\. Cui\.Thinking in flow: A dissipative stabilization operator for robust autoregressive reasoning\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Jansen et al\. \(2025\)P\. Jansen, O\. Tafjord, M\. Radensky, P\. Siangliulue, T\. Hope, B\. Dalvi, B\. P\. Majumder, D\. S\. Weld, and P\. Clark\.Codescientist: End\-to\-end semi\-automated scientific discovery with code\-based experimentation\.In*Findings of the Association for Computational Linguistics: ACL 2025*, 2025\.
- Jeon et al\. \(2025\)K\. Jeon, M\. Muehlebach, and M\. Tao\.Fast non\-log\-concave sampling under nonconvex equality and inequality constraints with landing\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Jiang and Gong \(2026\)S\. Jiang and R\. Gong\.Incremental BPE tokenization\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Jiang et al\. \(2025\)Z\. Jiang, D\. Schmidt, D\. Srikanth, D\. Xu, I\. Kaplan, D\. Jacenko, and Y\. Wu\.Aide: Ai\-driven exploration in the space of code\.*arXiv preprint arXiv:2502\.13138*, 2025\.
- Jin et al\. \(2026\)J\. Jin, Y\. Hu, K\. Qiu, Q\. Dai, C\. Luo, G\. Dong, X\. Li, T\. Zhao, X\. Ma, G\. Zhang, et al\.Toward generalist autonomous research via hypothesis\-tree refinement\.*arXiv preprint arXiv:2606\.11926*, 2026\.
- Kechris et al\. \(2026\)C\. Kechris, J\. Dan, and D\. Atienza\.Time series saliency maps: Explaining models across multiple domains\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Kim et al\. \(2026\)W\. Kim, S\. Hyeon, J\. Oh, and J\. Do\.VALUEFLOW: Toward pluralistic and steerable value\-based alignment in large language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Kiyani et al\. \(2026\)S\. Kiyani, S\. Noorani, G\. J\. Pappas, and H\. Hassani\.When to trust the cheap check: Weak and strong verification for reasoning\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Kreitner et al\. \(2026\)L\. Kreitner, P\. Hager, J\. Mengedoht, G\. Kaissis, D\. Rueckert, and M\. J\. Menten\.Efficient numeracy in language models through single\-token number embeddings\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Kwon et al\. \(2026\)J\. Kwon, D\.\-K\. Kim, J\. Kim, Y\. Kim, W\. Kook, and M\. Cha\.AI engram: In search of memory traces in artificial intelligence\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Lev et al\. \(2026\)O\. Lev, M\. Shenfeld, V\. Srinivasan, K\. Ligett, and A\. C\. Wilson\.Near\-optimal private linear regression via iterative hessian mixing\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Li et al\. \(2026a\)C\. Li, Y\. Wang, Y\. Wang, W\. Li, D\. Jaeger, and A\. Wu\.A factorized low\-rank RNN framework for uncovering independent neural latent dynamics and connectivity\.In*Forty\-third International Conference on Machine Learning*, 2026a\.
- Li et al\. \(2026b\)L\. Li, Y\. Wang, J\. Yan, W\. Zhang, J\. Deng, H\. Sun, Z\. Han, and Y\. Gong\.From text to forecasts: Bridging modality gap with temporal evolution semantic space\.In*Forty\-third International Conference on Machine Learning*, 2026b\.
- Li et al\. \(2025a\)T\. Li, Y\. Huang, L\. Jiang, C\. Liu, Q\. Xie, W\. Du, L\. Wang, and K\. Wu\.FedWMSAM: Fast and flat federated learning via weighted momentum and sharpness\-aware minimization\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025a\.
- Li et al\. \(2025b\)X\. Li, Y\. Luo, H\. Wang, H\. Li, L\. Peng, F\. Liu, Y\. Guo, K\. Zhang, and M\. Gong\.Towards accurate time series forecasting via implicit decoding\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025b\.
- Li et al\. \(2026c\)X\. Li, D\. Liu, and K\. Kawaguchi\.Initialization is half the battle: Generating diverse images from a guidance potential posterior\.In*Forty\-third International Conference on Machine Learning*, 2026c\.
- Li et al\. \(2022\)Y\. Li, D\. Choi, J\. Chung, N\. Kushman, J\. Schrittwieser, R\. Leblond, T\. Eccles, J\. Keeling, F\. Gimeno, A\. Dal Lago, et al\.Competition\-level code generation with alphacode\.*Science*, 2022\.
- Li et al\. \(2026d\)Y\. Li, C\. Shao, X\. Liu, R\. Zhao, P\. Liu, H\. Su, Z\. Chen, Q\. Yang, A\. Xu, Y\. Fang, et al\.Autosota: An end\-to\-end automated research system for state\-of\-the\-art ai model discovery\.*arXiv preprint arXiv:2604\.05550*, 2026d\.
- Liu et al\. \(2024\)A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang, C\. Ruan, et al\.Deepseek\-v3 technical report\.*arXiv preprint arXiv:2412\.19437*, 2024\.
- Liu et al\. \(2026a\)J\. Liu, S\. Qiu, M\. Li, B\. Li, H\. Ji, S\. Han, X\. Ye, P\. Xia, Z\. Dong, M\. Chen, et al\.Autoresearchclaw: Self\-reinforcing autonomous research with human\-ai collaboration\.*arXiv preprint arXiv:2605\.20025*, 2026a\.
- Liu et al\. \(2026b\)J\. Liu, X\. Zhao, X\. Shang, and Z\. Shen\.Dive into claude code: The design space of today’s and future ai agent systems\.*arXiv preprint arXiv:2604\.14228*, 2026b\.
- Liu and Ye \(2026\)S\.\-Y\. Liu and H\.\-J\. Ye\.Tabswift: An efficient tabular foundation model with row\-wise attention\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Liu et al\. \(2026c\)T\. Liu, E\. Dobriban, and F\. Orabona\.Online conformal prediction via universal portfolio algorithms\.In*Forty\-third International Conference on Machine Learning*, 2026c\.
- Liu et al\. \(2026d\)Y\. Liu, Y\. Zhao, Z\. Xie, Q\. Ye, J\. Jiao, Y\. Hu, S\. Cao, and Y\. Liu\.Balancing understanding and generation in discrete diffusion models\.In*Forty\-third International Conference on Machine Learning*, 2026d\.
- Liu et al\. \(2025\)Z\. Liu, M\. Cheng, G\. Zhao, J\. Yang, Q\. Liu, and E\. Chen\.Improving time series forecasting via instance\-aware post\-hoc revision\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Lu et al\. \(2024\)C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha\.The ai scientist: Towards fully automated open\-ended scientific discovery\.*arXiv preprint arXiv:2408\.06292*, 2024\.
- Lyu et al\. \(2026\)Y\. Lyu, X\. Zhang, X\. Yi, Y\. Zhao, S\. Guo, W\. Hu, J\. Piotrowski, J\. Kaliski, J\. Urbani, Z\. Meng, et al\.Evoscientist: Towards multi\-agent evolving ai scientists for end\-to\-end scientific discovery\.*arXiv preprint arXiv:2603\.08127*, 2026\.
- Ma et al\. \(2025\)Y\. Ma, H\. Wu, H\. Zhou, H\. Weng, J\. Wang, and M\. Long\.Physense: Sensor placement optimization for accurate physics sensing\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Martens et al\. \(2026\)T\. Martens, L\. Devos, L\. Cascioli, W\. Meert, H\. Blockeel, and J\. Davis\.OC\-space: a unifying perspective on verification of tree ensembles\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Meng et al\. \(2026\)R\. Meng, B\. D\. Mishra, J\. Chen, C\.\-L\. Li, P\. Goyal, M\. Parmar, Y\. Song, Y\. Song, R\. Sinha, P\. Ranganathan, et al\.Scientistone: Towards human\-level autonomous research via chain\-of\-evidence\.*arXiv preprint arXiv:2605\.26340*, 2026\.
- Monod et al\. \(2025\)M\. Monod, A\. Micheli, and S\. Bhatt\.Neuralsurv: Deep survival analysis with bayesian uncertainty quantification\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Muni et al\. \(2026\)A\. Muni, V\. Taboga, E\. Derman, P\.\-L\. Bacon, and E\. Delage\.Reward redistribution for CVar MDPs using a bellman operator on l\-infinity\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Nam et al\. \(2026a\)J\. Nam, J\. Yoon, J\. Chen, J\. Shin, S\. Arik, and T\. Pfister\.Mle\-star: Machine learning engineering agent via search and targeted refinement\.*Advances in Neural Information Processing Systems*, 2026a\.
- Nam et al\. \(2026b\)J\. Nam, J\. Yoon, J\. Chen, R\. Sinha, J\. Shin, and T\. Pfister\.Ds\-star: Data science agent for solving diverse tasks across heterogeneous formats and open\-ended queries\.*arXiv preprint arXiv:2509\.21825*, 2026b\.
- Ni et al\. \(2026\)Z\. Ni, S\. Wang, Y\. Yue, T\. Yu, W\. Zhao, Y\. Hua, T\. Chen, J\. Song, C\. Yu, B\. Zheng, and G\. Huang\.The flexibility trap: Rethinking the value of arbitrary order in diffusion language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Ning et al\. \(2025\)K\. Ning, Z\. Pan, Y\. Liu, Y\. Jiang, J\. Y\. Zhang, K\. Rasul, A\. Schneider, L\. Ma, Y\. Nevmyvaka, and D\. Song\.TS\-RAG: Retrieval\-augmented generation based time series foundation models are stronger zero\-shot forecaster\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Novikov et al\. \(2025\)A\. Novikov, N\. Vũ, M\. Eisenberger, E\. Dupont, P\.\-S\. Huang, A\. Z\. Wagner, S\. Shirobokov, B\. Kozlovskii, F\. J\. Ruiz, A\. Mehrabian, et al\.Alphaevolve: A coding agent for scientific and algorithmic discovery\.*arXiv preprint arXiv:2506\.13131*, 2025\.
- Ohnemus et al\. \(2026\)J\. Ohnemus, M\. Fochesato, R\. Zuliani, and J\. Lygeros\.Loss\-aware distributionally robust optimization via trainable optimal transport ambiguity sets\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Ortiz et al\. \(2026\)J\. J\. G\. Ortiz, A\. Gupta, C\. Rinard, and D\. Blalock\.Flashoptim: Optimizers for memory\-efficient training\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Pan et al\. \(2026\)W\. Pan, Z\. Liu, X\. Wang, Y\. Haining, and X\. Jia\.Towards long\-horizon interpretability: Efficient and faithful multi\-token attribution for reasoning LLMs\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Pan et al\. \(2025\)Y\. Pan, Z\. Cao, C\. GU, L\. Liu, P\. Zhao, Y\. Chen, and F\. Lin\.Multi\-task vehicle routing solver via mixture of specialized experts under state\-decomposable MDP\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Park et al\. \(2025\)J\. Park, Y\. Choi, and J\. Lee\.Multi\-class support vector machine with differential privacy\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Pfaff et al\. \(2026\)N\. Pfaff, T\. Cohn, S\. Zakharov, R\. Cory, and R\. Tedrake\.Scenesmith: Agentic generation of simulation\-ready indoor scenes\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Prinster et al\. \(2026\)D\. Prinster, C\. Fannjiang, J\. W\. Park, K\. Cho, A\. Liu, S\. Saria, and S\. D\. Stanton\.Conformal policy control\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Qiu et al\. \(2026\)X\. Qiu, S\. Gu, P\. Wu, J\. Hu, Y\. Wen, Y\. Pan, X\. Luo, B\. XU, and G\. Li\.SVL: Empowering spiking neural networks for efficient 3d open\-world understanding\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Rodionov et al\. \(2025\)G\. Rodionov, R\. Garipov, A\. Shutova, G\. Yakushev, E\. Schultheis, V\. Egiazarian, A\. Sinitsin, D\. Kuznedelev, and D\. Alistarh\.Hogwild\! inference: Parallel LLM generation via concurrent attention\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Sadi et al\. \(2026\)B\. Sadi, E\. Saig, and N\. Rosenfeld\.Welfare\-optimal classification with accuracy auctions\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Santis et al\. \(2026\)F\. D\. Santis, G\. Ciravegna, G\. D\. Felice, A\. Casanova, F\. Giannini, M\. Diligenti, J\. Schneider, D\. Giordano, M\. E\. Zarlenga, and P\. Barbiero\.Mixture of concept bottleneck experts\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Schmidgall et al\. \(2025\)S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum\.Agent laboratory: Using llm agents as research assistants\.*Findings of the Association for Computational Linguistics: EMNLP 2025*, 2025\.
- Schur et al\. \(2026\)F\. Schur, N\. Pfister, P\. Ding, S\. Mukherjee, and J\. Peters\.Many experiments, few repetitions, unpaired data, and sparse effects: Is causal inference possible?In*Forty\-third International Conference on Machine Learning*, 2026\.
- Shen et al\. \(2023\)Y\. Shen, K\. Song, X\. Tan, D\. Li, W\. Lu, and Y\. Zhuang\.Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face\.*Advances in Neural Information Processing Systems*, 2023\.
- Singh et al\. \(2025\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram, et al\.Openai gpt\-5 system card\.*arXiv preprint arXiv:2601\.03267*, 2025\.
- Słupiński and Lipinski \(2026\)M\. Słupiński and P\. Lipinski\.RED\-HDP\-HMM: Observation\-dependent durations for bayesian nonparametric sequential models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Song et al\. \(2026\)Y\. Song, Y\. Song, T\. Pfister, and J\. Yoon\.Paperorchestra: A multi\-agent framework for automated ai research paper writing\.*arXiv preprint arXiv:2604\.05018*, 2026\.
- Su et al\. \(2026\)L\. Su, M\. Zhang, Y\. Xiong, T\. LIU, S\. Zhang, X\. Chen, and L\. Sun\.TG\-RAG: A retrieval\-augmented framework for reasoning guidance in specialized domains\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Tajwar et al\. \(2026\)F\. Tajwar, G\. Zeng, Y\. Zhou, Y\. Song, D\. Arora, Y\. Jiang, J\. Schneider, R\. Salakhutdinov, H\. Feng, and A\. Zanette\.Maximum likelihood reinforcement learning\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Tang et al\. \(2026a\)J\. Tang, L\. Xia, Z\. Li, and C\. Huang\.Ai\-researcher: Autonomous scientific innovation\.*Advances in Neural Information Processing Systems*, 2026a\.
- Tang et al\. \(2026b\)L\. Tang, Y\. Meng, J\. Costa, Y\. Zhang, M\. Ye, and Z\. Xi\.The value of variance: Mitigating debate collapse in multi\-agent systems via uncertainty\-driven policy optimization\.In*Forty\-third International Conference on Machine Learning*, 2026b\.
- Tang et al\. \(2025\)Z\. Tang, B\. Wang, C\. Wen, and J\. Teng\.Accelerating feature conformal prediction via taylor approximation\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Team et al\. \(2023\)G\. Team, R\. Anil, S\. Borgeaud, J\.\-B\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican, et al\.Gemini: a family of highly capable multimodal models\.*arXiv preprint arXiv:2312\.11805*, 2023\.
- Tjanaka et al\. \(2026\)B\. Tjanaka, H\. Chen, M\. C\. Fontaine, and S\. Nikolaidis\.Discount model search for quality diversity optimization in high\-dimensional measure spaces\.In*The Fourteenth International Conference on Learning Representations*, 2026\.
- Tran et al\. \(2025\)V\.\-H\. Tran, T\. Tran, T\. Chu, T\. Le, and T\. M\. Nguyen\.Tree\-sliced entropy partial transport\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Wang et al\. \(2023\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. Anandkumar\.Voyager: An open\-ended embodied agent with large language models\.*arXiv preprint arXiv:2305\.16291*, 2023\.
- Wang et al\. \(2025a\)H\. Wang, zhengnan li, Z\. Chen, X\. Chen, S\. He, G\. Liu, H\. Li, and Z\. Lin\.Iterative missing data imputation with model form adaptation and non\-missing feature supervision\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025a\.
- Wang et al\. \(2025b\)J\. Wang, W\. Tu, and J\. Cheng\.Hierarchical shortest\-path graph kernel network\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025b\.
- Wang et al\. \(2026\)S\. Wang, X\. Ouyang, T\. Xu, Y\. Hu, J\. Liu, G\. Chen, T\. Zhang, J\. Zheng, K\. Yang, X\. Ren, D\. Liu, and L\. Zhang\.OPUS: Towards efficient and principled data selection in large language model pre\-training in every iteration\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Wang and Dobriban \(2026\)T\. Wang and E\. Dobriban\.Optimal decision\-making based on prediction sets\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Wang et al\. \(2025c\)X\. Wang, B\. Li, Y\. Song, F\. F\. Xu, X\. Tang, M\. Zhuge, J\. Pan, Y\. Song, B\. Li, J\. Singh, et al\.Openhands: An open platform for ai software developers as generalist agents\.In*International Conference on Learning Representations*, 2025c\.
- Wei et al\. \(2025\)T\. Wei, B\.\-L\. Wang, J\.\-X\. Shi, Y\.\-F\. Li, and M\.\-L\. Zhang\.X\-mahalanobis: Transformer feature mixing for reliable OOD detection\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Weng et al\. \(2025\)Y\. Weng, M\. Zhu, Q\. Xie, Q\. Sun, Z\. Lin, S\. Liu, and Y\. Zhang\.Deepscientist: Advancing frontier\-pushing scientific findings progressively\.*arXiv preprint arXiv:2509\.26603*, 2025\.
- Wróbel et al\. \(2026\)A\. Wróbel, S\. Gairola, J\. Tabor, B\. Schiele, B\. M\. Zieliński, and D\. D\. Rymarczyk\.DAVE: Distribution\-aware attribution via vit gradient decomposition\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Wu and Silwal \(2025\)F\. Wu and S\. Silwal\.Efficient training\-free online routing for high\-volume multi\-LLM serving\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Wu et al\. \(2024\)Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, et al\.Autogen: Enabling next\-gen llm applications via multi\-agent conversations\.In*First conference on language modeling*, 2024\.
- Xie et al\. \(2026\)T\. Xie, H\. Luo, H\. Tang, H\. Yiwen, J\. K\. Liu, Q\. Ren, Y\. Wang, X\. Zhao, R\. Yan, B\. Su, C\. Luo, and B\. Guo\.Controlled LLM training on spectral sphere\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Xu and Peng \(2025\)R\. Xu and J\. Peng\.A comprehensive survey of deep research: Systems, methodologies, and applications\.*arXiv preprint arXiv:2506\.12594*, 2025\.
- Yamada et al\. \(2025\)Y\. Yamada, R\. T\. Lange, C\. Lu, S\. Hu, C\. Lu, J\. Foerster, J\. Clune, and D\. Ha\.The ai scientist\-v2: Workshop\-level automated scientific discovery via agentic tree search\.*arXiv preprint arXiv:2504\.08066*, 2025\.
- Yang et al\. \(2024\)J\. Yang, C\. E\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. R\. Narasimhan, and O\. Press\.Swe\-agent: Agent\-computer interfaces enable automated software engineering\.In*The Thirty\-eighth Annual Conference on Neural Information Processing Systems*, 2024\.
- Yang et al\. \(2025\)Y\. Yang, D\. Zhang, Y\. Liang, H\. Lu, G\. Chen, and H\. Li\.Not all data are good labels: On the self\-supervised labeling for time series forecasting\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Yao et al\. \(2022\)S\. Yao, J\. Zhao, D\. Yu, I\. Shafran, K\. R\. Narasimhan, and Y\. Cao\.React: Synergizing reasoning and acting in language models\.In*NeurIPS 2022 Foundation Models for Decision Making Workshop*, 2022\.
- Ye et al\. \(2026\)F\. X\.\-F\. Ye, X\. Li, A\. Yu, M\.\-C\. Chang, L\. CHU, and D\. Wertheimer\.Flashsinkhorn: IO\-aware entropic optimal transport on GPU\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Yu et al\. \(2026\)G\. Yu, J\. Wang, C\. Yang, J\. Qin, A\. I\. Aviles\-Rivero, and S\. Wang\.Decentralized attention fails centralized signals: Rethinking transformers for medical time series\.In*The Fourteenth International Conference on Learning Representations*, 2026\.
- Yuksekgonul et al\. \(2026\)M\. Yuksekgonul, D\. Koceja, X\. Li, F\. Bianchi, J\. McCaleb, X\. Wang, J\. Kautz, Y\. Choi, J\. Zou, C\. Guestrin, and Y\. Sun\.Learning to discover at test time\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Zeng et al\. \(2026\)Y\. Zeng, Y\. Shi, T\. Tan, X\. Li, Y\. Qin, Z\. Lu, W\. Yang, J\.\-H\. Xue, and Q\. Liao\.Egotactile: Learning grasp pressure for everyday objects from egocentric video\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Zhan et al\. \(2026\)S\. Zhan, Y\. Lai, Z\. Liu, L\. Hai, S\. Li, X\. Cai, Z\. Lin, W\. Huang, and H\.\-T\. Zheng\.3viewsense: Spatial and mental perspective reasoning from orthographic views in vision\-language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Zhang et al\. \(2026\)H\. Zhang, Y\. Li, Z\. Wang, Z\. Wang, S\. Zhang, X\. Qu, and Y\. Cheng\.Characterizing, evaluating, and optimizing complex reasoning\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Zhang et al\. \(2025a\)K\. Zhang, S\. Zhang, D\. Zhou, and Y\. Zhou\.Wasserstein transfer learning\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025a\.
- Zhang et al\. \(2025b\)X\. Zhang, Z\. He, C\. Fu, and C\. Xie\.IA\-GGAD: Zero\-shot generalist graph anomaly detection via invariant and affinity learning\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025b\.
- Zhao et al\. \(2026a\)S\. Zhao, X\. Zhang, W\. Li, J\. Li, L\. zhang, T\. Xue, and J\. Zhang\.Reasoning as representation: Rethinking visual reinforcement learning in image quality assessment\.In*The Fourteenth International Conference on Learning Representations*, 2026a\.
- Zhao et al\. \(2026b\)Z\. Zhao, K\.\-C\. Mo, S\.\-H\. Ho, B\. Amos, and K\. Wang\.A fully first\-order layer for differentiable optimization\.In*Forty\-third International Conference on Machine Learning*, 2026b\.
- Zhou et al\. \(2026\)Q\. Zhou, E\. Aleshina, A\. Lovyagin, O\. Somov, M\. Seleznyov, A\. Panchenko, I\. Oseledets, E\. Tutubalina, and I\. Y\. Tyukin\.Harnessing non\-adversarial robustness in large language models\.In*Forty\-third International Conference on Machine Learning*, 2026\.
- Zhou et al\. \(2025\)Y\. Zhou, J\. Wu, Z\. Ren, Z\. Yao, W\. Lu, K\. Peng, Q\. Zheng, C\. Song, W\. Ouyang, and C\. Gou\.CSBrain: A cross\-scale spatiotemporal brain foundation model for EEG decoding\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025\.
- Zhu et al\. \(2025a\)W\. Zhu, J\. Wang, B\. Gao, Y\. Jia, H\. Tan, Y\.\-Q\. Zhang, W\.\-Y\. Ma, and Y\. Lan\.AANet: Virtual screening under structural uncertainty via alignment and aggregation\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025a\.
- Zhu et al\. \(2025b\)Z\. Zhu, Y\. QI, H\. Ma, W\. Lu, and J\. Feng\.Stochastic forward\-forward learning through representational dimensionality compression\.In*The Thirty\-ninth Annual Conference on Neural Information Processing Systems*, 2025b\.

## Appendix AExperiment Details

### A\.1Benchmark

NeurIPS 2025 Papers\.We utilize 38 papers accepted at NeurIPS 2025\. Specifically, these papers are drawn from benchmarks used in AutoSOTA\([Li et al\., 2026d](https://arxiv.org/html/2609.19644#bib.bib65)\)\. We provide the full list of papers below\.

Table 12:NeurIPS 2025 Papers\.We utilize 38 accepted papers\.ICLR 2026 Papers\.We utilize 5 papers accepted at ICLR 2026\. Specifically, these papers are drawn from benchmakrs used in AutoSOTA\([Li et al\., 2026d](https://arxiv.org/html/2609.19644#bib.bib65)\)\. We provide the full list of papers below\.

Table 13:ICLR 2026 Papers\.We utilize 5 accepted papers\.ICML 2026 Spotlight Papers\.We utilize 64 papers accepted as spotlight presentations at ICML 2026\. Specifically, these 64 papers were selected from among all spotlight papers by strictly adhering to AutoSOTA’s filtering process \(e\.g\., verifying reproducibility\)\. We provide the full list of papers below\.

Table 14:ICML 2026 Spotlight Papers\.We utilize 64 papers accepted as spotlight presentations\.
### A\.2Configuration

Unless otherwise specified, we employ Gemini 3\.6 Flash for all agents, except for the Idea Experiment Coding Agent, the Ablation Study Agent, the Rebuttal Agent, and the Draft Enhancer, which use Claude Code with Opus 4\.8\. To evaluate novelty, ScientistTwo retrieves two reference papers via Google Search\. We extract limitations for a maximum of 16 rounds\. In each idea experimentation round, we evaluate two candidates: one selected from the seed ideas and the other an evolved idea\. We run this experimentation loop for up to four rounds, terminating early once four successful ideas are obtained\. If the Idea Critic Agent flags an idea for engineering refinement, we apply engineering techniques for at most two rounds\. Similarly, when the Ablation Critic Agent recommends refinement based on ablation results, we refine the idea at most once\. Finally, the peer\-review simulation runs for at most two rounds and terminates early if the ScholarPeer review score reaches 8, with review\-based refinement conducted at most once\. We use the ICLR 2025 format for drafting, following the PaperOrchestra\.

## Appendix BDetailed Comparison with AutoSOTA

Table 15:Aggregate differencesover the five papers\. The two systems apply different acceptance rules: AutoSOTA halts as soon as a pre\-registered numerical target is cleared \(e\.g\. after a single iteration on DMSQD\([Tjanaka et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib108)\)\), whereas ScientistTwo has no target metric and additionally requires the gain to be attributable to the proposed mechanism in ablation\.We run ScientistTwo on the same five ICLR 2026 submissions \(Table[13](https://arxiv.org/html/2609.19644#A1.T13)\) that AutoSOTA reports\. The two systems are built for different objectives\. AutoSOTA is an end\-to\-end, metric\-driven optimizer: starting from a raw paper, it locates the repository, reconstructs a runnable baseline, and distills the paper’s headline result into a single numerical target, then proposes and benchmarks edits until that target is cleared\. It does construct a multi\-dimensional rubric, but the optimization loop is driven by that one result\-match number\. In contrast, ScientistTwo operates without predefined numerical targets—it must independently diagnose research limitations, formulate a novel methodology, implement and ablate the proposed approach, and produce a complete scientific manuscript\.

Table[15](https://arxiv.org/html/2609.19644#A2.T15)summarizes the quantitative differences, while Table[16](https://arxiv.org/html/2609.19644#A2.T16)provides a qualitative comparison of the modifications introduced by each system\. Because each system measures against its own reproduced baseline on different hardware, the twoΔ\\Deltacolumns are not a head\-to\-head on a common metric; they characterize the*nature*and*scope*of each change\.

What gets optimized\.For these ICLR2026 papers, every AutoSOTA improvement is a configuration change inside an existing code path: solver iterations and float precision\([Grontas et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib38)\), emitter count\([Tjanaka et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib108)\), model width and optimizer\([Yu et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib128)\), a threshold in the evaluation harness\([Bhalla et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib6)\), and the mixing weights of three already\-implemented pooling paths\([Zhao et al\., 2026a](https://arxiv.org/html/2609.19644#bib.bib135)\)\. No new algorithmic component appears in any of the five, and the median change is under ten lines\. ScientistTwo’s accepted solutions instead introduce transferable mechanisms: a nullspace re\-parameterization that makes the equality constraints of Pinet\([Grontas et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib38)\)exact by construction, a sheaf\-Laplacian/Koopman training objective for T\-SAE\([Bhalla et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib6)\), adversarial content–quality disentanglement with region\-level quality segmentation for RALI\([Zhao et al\., 2026a](https://arxiv.org/html/2609.19644#bib.bib135)\), and a dimension\-adaptive smoothness penalty on the discount model for DMSQD\([Tjanaka et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib108)\)\.

What counts as success\.The two systems disagree about what a success*is*, and this drives most of the divergence\. On TeCh\([Yu et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib128)\), AutoSOTA reports\+4\.45%\+4\.45\\%accuracy from widening the backbone and adding label smoothing\. ScientistTwo found the same lever: itsDMC\-TeChvariant beat the baseline on 5 of 6 metrics, but the ablation critic rejected it because “the gains were primarily driven by general training controls \(EMA and label smoothing\) rather than the multi\-core architectural innovation itself,” so it was not accepted as a contribution\. A metric\-movement criterion and an attribution criterion score generic capacity tuning in opposite directions\.

Robustness and scope\.Optimizing a single registered number exposes two classic failure modes\.*Unmeasured trade\-offs*: on Pinet\([Grontas et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib38)\)the dominant edit removes exactly the solver work that controls constraint violation, which is not in the registered metric, so the−16\.7%\-16\.7\\%latency is reported with no feasibility number\.*Optimizing the evaluator*: on T\-SAE\([Bhalla et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib6)\)the winning change alters how activations are displayed to the LLM judge, leaving the SAE untouched, and AutoSOTA’s own re\-evaluation returns0\.75570\.7557, below the0\.75860\.7586baseline\. AutoSOTA does declare red lines against exactly these failures \(R1–R2 forbid altering the evaluation script or metric parameters, and R4 constrains cross\-metric trade\-offs\), but they fail empirically on both cases here\. ScientistTwo instead blocks both structurally, via a reproduction re\-run \(I1\), a protocol\-immutability audit \(I2\), and a method–code alignment audit \(I4\) rather than a prompt\-level list of prohibitions, and evaluates on the paper’s full benchmark grid rather than the single registered split\.

Complementarity\.The two systems are not strictly ordered, and DMSQD\([Tjanaka et al\., 2026](https://arxiv.org/html/2609.19644#bib.bib108)\)shows why: both improve QD score, but along orthogonal axes\. AutoSOTA raises the emitter count→2015\\\!\\to\\\!20, buying its\+33%\+33\\%evaluations per iteration—a pure compute\-scaling knob that our idea generator deliberately does not propose\. ScientistTwo instead leaves the compute budget fixed and improves the discount model itself, adding a dimension\-adaptive contact\-form smoothness penalty \(LC\-FTT\) that lifts DMS in high\-dimensional measure spaces \(\+422\+422QD/\+1\.16/\\,\{\+\}1\.16pp coverage averaged, reproduced bit\-exact\)\. The two changes touch disjoint parts of the pipeline and compose\. We therefore see the systems as complementary stages: ScientistTwo produces a method, and a tuner such as AutoSOTA can then optimize its deployment constants—and where a headline metric is near\-monotone in compute, a short tuning loop is the cheaper tool for that last mile\.

Table 16:Paper\-by\-paper comparisonon the five submissions of Table[13](https://arxiv.org/html/2609.19644#A1.T13)\.Δ\\Deltais self\-reported by each system against its own reproduced baseline; the two columns optimize different metrics in three of five cases and are not a head\-to\-head\.
## Appendix CQualitative Results

In this section, we present qualitative artifacts illustrating the discovery of Procrustes\-DS, an out\-of\-distribution \(OOD\) detection method designed by ScientistTwo that outperforms X\-Mahalanobis\([Wei et al\., 2025](https://arxiv.org/html/2609.19644#bib.bib116)\)\. We provide representative excerpts generated by ScientistTwo across its research pipeline below, selected from the broader trajectory of hypotheses, explorations, and ablation studies conducted by the framework:

- •Limitations of X\-Mahalanobis \(pp\. 29–30\)
- •Ideation and Method Proposal \(pp\. 31–34\)
- •Experimental Evaluation Report \(pp\. 35–37\)
- •Ablation Study Report \(pp\. 38–40\)
- •Critic Agent Feedback \(p\. 41\)
- •Reproducibility Audit Report \(p\. 42\)
- •Specification Verification and Method\-Code Alignment Audit \(pp\. 43–45\)

For an example of a rebuttal report obtained through a simulated peer\-review process, since Procrustes\-DS did not undergo this process, the example of TABHARMONY is presented \(pp\. 46–50\)\.

## Appendix DCase Study: DynaSpec\-RAG

We provide the final draft of DynaSpec\-RAG, which is developed by our ScientistTwo \(pp\. 51–65\)\.

![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/limitation_1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/limitation_2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/idea_1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/idea_2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/result_1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/result_2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/ablation_2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/audit1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/audit_1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/audit_2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/rebuttal_1.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/rebuttal_2.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/rebuttal_4.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/rebuttal_5.png)![[Uncaptioned image]](https://arxiv.org/html/2609.19644v1/ts_rag4.png)

Similar Articles

Towards End-to-End Automation of AI Research

arXiv cs.AI

A paper presenting The AI Scientist, a system that automates the entire research lifecycle from idea generation to peer review, demonstrating AI's growing capacity for scientific contribution.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

arXiv cs.AI

ScientistOne introduces Chain-of-Evidence, a verifiability framework for autonomous research agents that ensures every claim is traceable to evidence, achieving zero hallucinated references, perfect score verification, and the highest method-code alignment across 75 papers while matching or exceeding human expert performance on frontier research tasks.

EvoScientist: Towards Multi-Agent Evolving AI Scientists for End-to-End Scientific Discovery

Papers with Code Trending

EvoScientist is an adaptive multi-agent framework for end-to-end scientific discovery that continuously improves through persistent memory modules, comprising three specialized agents for idea generation, experiment execution, and knowledge distillation. It outperforms 7 state-of-the-art systems in scientific idea generation and improves code execution success rates through multi-agent evolution.