How Far Are We From True Auto-Research?
Summary
This paper introduces ResearchArena, a scaffold for evaluating auto-research agents, and finds that while agent-generated papers appear competitive under manuscript-only review, artifact-aware review reveals severe failures in experimental rigor, with no paper meeting top-tier acceptance standards.
View Cached Full Text
Cached at: 05/20/26, 08:28 AM
# How Far Are We From True Auto-Research?
Source: [https://arxiv.org/html/2605.19156](https://arxiv.org/html/2605.19156)
Zhengxin ZhangNing Wang11footnotemark:1Sainyam GalhotraClaire Cardie Cornell University \{zz865, nw366\}@cornell\.edu
###### Abstract
Recent auto\-research systems*can*produce complete papers, but feasibility is not the same as quality, and the field still lacks a systematic study of how good agent\-generated papers actually are\. We introduceResearchArena, a minimal scaffold that lets off\-the\-shelf agents \(Claude Code using Opus 4\.6, Codex using GPT\-5\.4, and Kimi Code using K2\.5\) carry out the full research loop themselves \(ideation, experimentation, paper writing, self\-refinement\) under only lightweight guidance\. Across 13 computer science seeds and 3 trials per agent\-domain pair, ResearchArena yields 117 agent\-generated papers, each evaluated under three complementary lenses: a manuscript\-only reviewer \(SAR\), an artifact\-aware peer review \(PR\) in which agents inspect the workspace alongside the manuscript, and an human conducted meta\-review\. Under SAR alone the picture is optimistic: Claude Code obtains the highest score, outperforms Analemma’s FARS, and matches the weighted\-average human ICLR 2025 submission, suggesting that minimally scaffolded agents can produce papers that look competitive on manuscript\-only review\. Manual inspection, however, reveals this picture is overstated: SAR scores are poorly aligned with its actual acceptance decisions and reward plausible framing without verifying experimental substance\. Under artifact\-aware PR scores drop sharply, and manual auditing identifies experimental rigor as the major bottleneck, decomposing into three failure modes \(*fabricated results*,*underpowered experiments*, and*plan/execution mismatch*\) that are highly agent\-dependent: Codex 5%/8% paper\-vs\-artifact mismatch / fabricated references versus Kimi Code 77%/72%, a∼\\sim15×\\timesspread that tracks distinct research*personas*the agents develop\. None of the 117 agent\-generated papers reaches the acceptance bar of a top\-tier venue\. This suggests that we are still gaped from the true auto\-research\.
## 1Introduction
Large language models \(LLMs\) have rapidly evolved from passive text generators into autonomous agents capable of interleaving reasoning with actionsYaoet al\.\([2022](https://arxiv.org/html/2605.19156#bib.bib11)\), invoking external toolsSchicket al\.\([2023](https://arxiv.org/html/2605.19156#bib.bib12)\), browsing the webNakanoet al\.\([2021](https://arxiv.org/html/2605.19156#bib.bib13)\); Zhouet al\.\([2023](https://arxiv.org/html/2605.19156#bib.bib14)\), writing and executing code in real software environmentsYanget al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib16)\); Jimenezet al\.\([2023](https://arxiv.org/html/2605.19156#bib.bib17)\), and operating over long horizons in open\-ended settingsWanget al\.\([2023](https://arxiv.org/html/2605.19156#bib.bib15),[2024](https://arxiv.org/html/2605.19156#bib.bib18)\)\. Recent studiesLuet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib2)\); Yamadaet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib19)\); Analemma Intelligence \([2026](https://arxiv.org/html/2605.19156#bib.bib3)\); Baeket al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib21)\); Schmidgallet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib20)\); Schmidgall and Moor \([2025](https://arxiv.org/html/2605.19156#bib.bib22)\)have begun chaining these capabilities into end\-to\-end scientific research pipelines that take a seed topic and produce a complete research artifact\. For example, the AI ScientistLuet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib2)\); Yamadaet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib19)\)brainstorms research ideas, writes and runs code, summarizes results, and drafts full manuscripts, with its second version producing the first peer\-review\-accepted workshop paper authored entirely by an AI system\. Analemma’s Fully Automated Research System \(FARS\)Analemma Intelligence \([2026](https://arxiv.org/html/2605.19156#bib.bib3)\)pursues a similar full\-pipeline objective at substantially larger compute scale\. These works demonstrate that agentic systems*can*produce complete papers\. However, they primarily establish feasibility rather than quality: we still lack a systematic study of how good agent\-generated papers actually are\.
To study this question, we build a minimal scaffold for off\-the\-shelf agents, calledResearchArena, that lets general\-purpose agents carry out the full research loop themselves: ideation, experimentation, paper writing, and self\-refinement, with only lightweight guidance\. Whereas prior benchmarks such as MLR\-BenchChenet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib27)\)evaluate open\-ended ML research through modular scaffolds with stage\-wise and end\-to\-end evaluation, ResearchArena studies a complementary setting: a single off\-the\-shelf agent operates autonomously across broader computer science domains, rather than a pipeline of stage\-specific components\. We evaluate three frontier agents: Claude Code with Opus 4\.6Anthropic \([2026](https://arxiv.org/html/2605.19156#bib.bib4)\), Codex with GPT\-5\.4OpenAI \([2026](https://arxiv.org/html/2605.19156#bib.bib6)\), and Kimi Code with K2\.5Moonshot AI \([2026](https://arxiv.org/html/2605.19156#bib.bib8)\)\. To test these systems across diverse research settings, we select 13 computer science domains, including 5 CPU\-only and 8 GPU\-intensive fields, and run 3 trials for each agent\-domain pair\. This yields 117 agent\-generated papers, together with their accompanying with experimental artifacts\. Every paper is then evaluated under three complementary lenses: an manuscript\-only agentic reviewer \(SAR\)Stanford ML Group \([2025](https://arxiv.org/html/2605.19156#bib.bib10)\), our artifact\-aware peer review \(PR\) in which agents inspect the workspace alongside the manuscript, and human inspection\.
SAR\-only evaluation paints an optimistic picture and serves as a useful manuscript\-only calibration lens: Claude Code obtains the highest average score among the three agents, outperforms Analemma’s FARS system, and reaches a score comparable to the weighted\-average human\-authored ICLR 2025 submission\. This suggests that minimally scaffolded agents can produce papers that look competitive under manuscript\-only review\. Manual inspection, however, reveals that this picture is overstated: SAR scores are poorly aligned with actual ICLR acceptance decisions, and SAR rewards plausible\-but\-non\-workable ideas, polished framing, and honest\-looking negative results without verifying experimental substance\. Our artifact\-aware PR and human inspection tell a different story\. Under PR, where reviewers see the code and logs alongside the manuscript, scores drop sharply and almost all papers fall below the acceptance threshold\. Manual inspection identifies*experimental rigor*as the major bottleneck across all agents, decomposing into three distinct failure modes:*fabricated results*\(numbers reported in the paper do not match the underlying outputs\),*underpowered experiments*\(narrow scope on a single small dataset and a single model\), and*plan/execution mismatch*\(the experiment does not include all the components from ideation\)\. These modes are agent\-dependent: Codex shows mostly underpowered experiments and the fewest integrity issues \(results\-vs\-artifact mismatches and fabricated references in only 5% / 8% of papers\), Kimi Code combines fabrication and plan/execution mismatch \(77% / 72%\), and Claude Code falls in between \(31% / 36%\); this∼\\sim15×\\timesspread tracks the distinct research*personas*the agents develop: Codex as careful empirical scientist, Kimi Code as ambitious system builder, and Claude Code as full\-stack researcher \(§[4](https://arxiv.org/html/2605.19156#S4)\)\.
Taken together, our findings show that despite producing papers that look polished on manuscript\-only review, the actual quality of agent\-generated research, measured by artifact\-aware peer review and manual auditing, remains far below human\-authored work, andnone of the 117 agent\-generated papers reaches the acceptance bar of a top\-tier venue\. To support the community in tracking progress as models advance, we release the full corpus: 117 papers with their code and logs, 351 PR reviews, 117 SAR scores, human inspection results, and the configurable harness\.
## 2Related Work
Auto\-research systems\.A growing body of workKarpathy \([2026](https://arxiv.org/html/2605.19156#bib.bib1)\); Analemma Intelligence \([2026](https://arxiv.org/html/2605.19156#bib.bib3)\); Luet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib2)\); Yamadaet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib19)\); Baeket al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib21)\); Schmidgallet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib20)\); Schmidgall and Moor \([2025](https://arxiv.org/html/2605.19156#bib.bib22)\)has demonstrated end\-to\-end agent\-driven research\. For example, the AI ScientistLuet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib2)\)pioneered the full loop of ideation, experiments, writing, and automated review with a linear multi\-agent pipeline, evaluated on three ML subfields: diffusion modeling, transformer\-based language modeling, and learning dynamics\. Its successorYamadaet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib19)\)replaces the linear loop with agentic tree search, adds VLM feedback for figures and parallel experiment execution, and produced the first peer\-review\-accepted workshop paper\. Analemma’s Fully Automated Research System \(FARS\)Analemma Intelligence \([2026](https://arxiv.org/html/2605.19156#bib.bib3)\), in contrast, is a closed multi\-agent pipeline reportedly run at substantial compute scale \($104,000 reported\) and produced over 100 agent\-generated papers\. Karpathy’s Auto\-ResearchKarpathy \([2026](https://arxiv.org/html/2605.19156#bib.bib1)\)is, in contrast, a minimal single\-agent demonstration that iteratively edits a fixedtrain\.pyto explore architecture and hyperparameter choices against a held\-out validation metric, automating only the coding\-and\-experimentation stages\. ResearchAgentBaeket al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib21)\)targets only the early stages \(problem definition, method proposal, and experiment design\), iteratively refined by multiple LLM\-based reviewing agents calibrated to human criteria and grounded in an academic citation graph plus a cross\-paper concept store, with no code, experiment execution, or paper writing\. Agent LaboratorySchmidgallet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib20)\)structures the full process as three sequential phases driven by specialized LLM agents: literature review, anmle\-solvermodule for experimentation, and apaper\-solvermodule for report writing, with an optional human\-in\-the\-loop co\-pilot mode\. AgentRxivSchmidgall and Moor \([2025](https://arxiv.org/html/2605.19156#bib.bib22)\)adds a shared preprint server through which multiple agent laboratories upload and retrieve each other’s reports across runs, allowing successive runs to build on prior research rather than operating in isolation\.
Benchmarks for LLM research agents\.Existing benchmarksHuanget al\.\([2023](https://arxiv.org/html/2605.19156#bib.bib23)\); Chanet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib25)\); Zhanget al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib30)\); Wijket al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib28)\); Chenet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib24),[2025](https://arxiv.org/html/2605.19156#bib.bib27)\); Staraceet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib26)\); Siegelet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib29)\)evaluate language agents on partial slices of the research process, and fall into three groups\.*Fixed\-task ML engineering benchmarks*score agents on predefined tasks against objective leaderboard metrics: MLAgentBenchHuanget al\.\([2023](https://arxiv.org/html/2605.19156#bib.bib23)\)\(13 tasks from CIFAR\-10 to BabyLM and Kaggle challenges\), MLE\-benchChanet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib25)\)\(75 Kaggle competitions\), MLRC\-BenchZhanget al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib30)\)\(7 ML research\-competition tasks targeting novel\-methodology proposal and implementation\), and RE\-BenchWijket al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib28)\)\(7 open\-ended ML R&D environments pitting agents against human experts under matched time budgets\)\.*Open\-ended research\-task benchmarks*draw tasks from peer\-reviewed publications and require self\-contained research artifacts: ScienceAgentBenchChenet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib24)\)extracts 102 data\-driven discovery problems from 44 papers across four disciplines, and MLR\-BenchChenet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib27)\)contains 201 open\-ended ML research tasks taken from NeurIPS / ICLR / ICML workshops\.*Replication benchmarks*ask agents to reach a known target: PaperBenchStaraceet al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib26)\)evaluates replicating 20 ICML 2024 spotlights from scratch via hierarchically decomposed rubrics, while CORE\-BenchSiegelet al\.\([2024](https://arxiv.org/html/2605.19156#bib.bib29)\)measures reproduction of computational results from already\-published papers \(an adjacent, complementary line of work\)\.
## 3ResearchArena
In this secton, we first give an overview of ResearchArena §[3\.1](https://arxiv.org/html/2605.19156#S3.SS1), describe the setup in §[3\.2](https://arxiv.org/html/2605.19156#S3.SS2)and describe the three complementary evaluation lenses: the Stanford Agentic Reviewer \(SAR\), our artifacts\-aware peer review \(PR\), and a human inspection in §[3\.3](https://arxiv.org/html/2605.19156#S3.SS3)–[3\.5](https://arxiv.org/html/2605.19156#S3.SS5)\.
### 3\.1Overview
As shown in Figure[1](https://arxiv.org/html/2605.19156#S3.F1), each agent receives a CS\-domain seed and runs a four\-stage research loop: ideation, experiments, paper writing, and review\. Stages 1–3 each include a self\-refinement loop\. At each of these three stages, the agent is paired with a concise domain\-specific guideline that fixes the deliverable but not the research itself, distilled from established research practice \(e\.g\., Schulman’s ML research notesSchulman \([2020](https://arxiv.org/html/2605.19156#bib.bib31)\), the ResearchAgent methodologyBaeket al\.\([2025](https://arxiv.org/html/2605.19156#bib.bib21)\), Peyton Jones’s writing advicePeyton Jones \([2017](https://arxiv.org/html/2605.19156#bib.bib33)\), and the submission and reviewer instructions\)\. The guidelines are intentionally kept short so they act as minimal scaffolding rather than as a step\-by\-step recipe\. We provide example guidelines in Appendix[B](https://arxiv.org/html/2605.19156#A2)\. Stage 4 evaluates the resulting paper through three complementary lenses: the Stanford Agentic Reviewer \(SAR, manuscript\-only\), our artifacts\-aware peer review \(PR, in which three agents inspect the workspace alongside the manuscript\), and human inspection\.
Figure 1:The ResearchArena pipeline\.
### 3\.2Setup
ResearchArena spans 13 research seeds across two compute platforms\. The 5 CPU seeds \(causal learning, compiler optimization, data integration & cleaning, operating system design, probabilistic methods\) target systems / databases / programming\-language venues\. The 8 GPU seeds \(AI for biology, computer vision, datasets & benchmarks, generative models, interpretability, NLP, privacy in ML, supervised representation learning\) target ML venues\. Hardware: 1×\\timesNVIDIA RTX A6000 \(48 GB\) with 4 CPUs and 60 GB RAM for the main experiments\. We re\-run all GPU seeds on 1×\\timesH100 \(80 GB\) to test compute scaling \(§[5](https://arxiv.org/html/2605.19156#S5)\)\.
### 3\.3Stanford Agentic Reviewer \(SAR\)
SARStanford ML Group \([2025](https://arxiv.org/html/2605.19156#bib.bib10)\)is an automatic agentic paper reviewer that is calibrated to the ICLR scale \(0–10\) and returns an overall score together with strengthes and weaknesses for any submitted manuscripts\. We use SAR for three purposes: \(i\) to score all 117 agent\-generated papers from the manuscript–only perspective; \(ii\) to anchor these scores against human\-authored papers by additionally scoring 200 ICLR 2025 papers \(100 accepted, 100 rejected\); and \(iii\) to compare against an existing automated research system by scoring 102 FARS\-generated papers\.
### 3\.4Artifacts\-Aware Peer Review \(PR\)
All three agents review every paper \(351 reviews=117papers×3=117\\text\{ papers\}\\times 3\)\. We distill a domain\-specific reviewer guideline \(Appendix[B](https://arxiv.org/html/2605.19156#A2.SS0.SSS0.Px5)\), standardize all domains on the ICLR 0–10 scoring scale, and break each review down into nine dimensions:*novelty*,*soundness*,*significance*,*clarity*,*reproducibility*,*experimental rigor*,*references*,*reference integrity*, and*results integrity*\. Reviewers check results integrity against experimental artifacts and reference integrity by online lookups against arXiv, Semantic Scholar, and CrossRef\. Each reviewer is given*read\-only*access to the workspace; the read\-only restriction prevents a reviewer agent from silently modifying the artifacts under review, so the paper\-vs\-artifact comparison reflects what the authoring agent actually produced\.
### 3\.5Human inspection
The authors serve as*meta\-reviewers*\. For every paper, two authors jointly assess both the manuscript, experimental artifacts, SAR review, and PR reviews\. The meta\-review deliberately focuses on integrity rather than novelty\. First, integrity is*objectively verifiable*against the artifacts: a reported number either matchesresults\.jsonor it does not, and a citation either resolves to a real bibliographic entry or it does not\. Novelty, by contrast, are inherently subjective and remain the responsibility of the SAR and PR scores\. Second, the official reviewer instructions of top\-tier ML conferences all caution reviewers against using “lack of novelty” as a sole rejection criterion and ask them to remain open\-minded about new ideas; treating novelty as the discriminator for paper quality would therefore run against the field’s own reviewing norms\.
## 4Capabilities
In this section, we first compare the three agents against both an automated research system \(FARS\) and human ICLR papers in §[4\.1](https://arxiv.org/html/2605.19156#S4.SS1), identify three research personas across the agents in §[4\.2](https://arxiv.org/html/2605.19156#S4.SS2), and analyze the progamming language and time usage in §[4\.3](https://arxiv.org/html/2605.19156#S4.SS3)\.
### 4\.1Comparison against automated systems and human baselines
Figure 2:SAR score distributions\.Table 1:SAR scores\.
Figure[2](https://arxiv.org/html/2605.19156#S4.F2)shows the SAR score distributions for the three agents and Analemma’s FARSAnalemma Intelligence \([2026](https://arxiv.org/html/2605.19156#bib.bib3)\), and Table[1](https://arxiv.org/html/2605.19156#S4.T1)reports the per\-system means and standard deviations\. Mean scores rank asClaude Code \(5\.45\)\>\>FARS \(5\.06\)\>\>Codex \(4\.93\)\>\>Kimi Code \(4\.24\): Claude Code outperforms FARS by 0\.39 SAR points and Codex achieves similar performance to FARS \(4\.93 vs\. 5\.06\), all while our entire three\-agent run cost∼\\sim$1,000 \(≈\\approx$9 per paper across the 117 papers\), versus FARS’s reported $104,000 \(∼\\sim$1,040 per paper\), roughly 100×\\timescheaper per paper\. Kimi Code lags the other automated systems\. The picture above the ICLR acceptance threshold \(SAR≥6\\geq 6\) is even more lopsided: Claude Code produces 21% \(8/39\) of papers, versus 10% for Codex, only 1% \(1/102\) for FARS, and 0% for Kimi Code\. Together, these results validate the effectiveness of ResearchArena: a minimal scaffold around an off\-the\-shelf agent matches or surpasses a heavily engineered, closed\-source auto\-research system\. Against the 200 ICLR 2025 baselines in Table[1](https://arxiv.org/html/2605.19156#S4.T1), Claude Code \(5\.45\) sits between rejected \(5\.34\) and accepted \(5\.59\) human submissions and*exceeds*the weighted\-average human submission \(5\.42\), where the weighted average mixes the accepted and rejected means in proportion to ICLR’s∼\\sim32% acceptance rate\.
### 4\.2Three research personas
Figure 3:Word cloud of the most frequent content words in each agent’s paper titles\.During our human inspection of all 117 papers \(§[3\.5](https://arxiv.org/html/2605.19156#S3.SS5)\), we find that the three agents have developed fundamentally different research personas\. To make this concrete, we further run research\-type analysis on every paper along with the title and the paper structure breakdown, summarized in Table[2](https://arxiv.org/html/2605.19156#S4.T2)and complemented by the per\-agent title word cloud in Figure[3](https://arxiv.org/html/2605.19156#S4.F3)\.
Table 2:Per\-agent persona signals on research type, title, and paper structure\.Claude Code: the full\-stack researcher\.Claude Code produces the most balanced portfolio: 46% method papers, 46% empirical studies, and 8% benchmark papers\. It writes the longest papers \(4,023 words on average\) with the most figures \(4\.8\) and tables \(6\.0\), and includes complexity analysis in 77% of papers\. Title style favors essayistic “*The X of Y*” framing, e\.g\.*“The Algebra of Compiler Passes: An Empirical Study of Idempotency,”**“The Bandwidth Knapsack: Optimal Migration Scheduling,”*and*“The Functional Anatomy of Sparse Features in Language Models\.”*Title vocabulary leans analytical and mechanistic \(learning,when,causal,adaptive,pipelines,contrastive\)\. The full\-stack persona is the most ambitious of the three; when Claude Code does fail, the failure mode is narrow\-but\-occasionally\-fabricated experiments rather than wholesale fabrication or method/implementation mismatch \(§[5\.2](https://arxiv.org/html/2605.19156#S5.SS2)\)\.
Codex: the empirical scientist\.Codex is overwhelmingly empirical \(87% of papers\), while producing only 13% method papers and*zero*benchmark papers\. Its papers are mid\-length \(3,421 words\), with the fewest equations \(2\.3 vs\. 3\.8 / 4\.0 for Claude Code / Kimi Code\),*zero*algorithm blocks, and*zero*theorems, consistent with an empiricist style that defers from formal claims\. Codex has the highest question\-title rate at 28% \(vs\. Claude Code 10% and Kimi Code 0%\), framed as controlled studies:*“Do Shared Decoders Improve Prototype\-Edit Reusability?”*,*“When Does Clarification Supervision Transfer to Formal Reasoning?”*,*“How Much Signal Is in Early Training Trajectories?”*\. Title vocabulary clusters around controlled\-study and pilot\-study terms \(study,benchmark,negative,matched,controlled,pilot\)\. The empiricist persona buys high integrity \(Codex has the fewest fabricated references; §[5\.2](https://arxiv.org/html/2605.19156#S5.SS2)\) but at the cost of empirical breadth: many Codex papers are explicitly scoped as pilot or feasibility studies that are underpowered\.
Kimi Code: the system builder\.Kimi Code reframes 79% of its papers as methods, the highest method\-paper rate of any agent\. Titles are acronym\-heavy named frameworks \(51% acronym rate, 85% “Name: Subtitle” colon structure\) and*never*questions: e\.g\.*“CAGER: Causal Geometric Explanation Recovery,”**“DU\-VPT: Decomposed Uncertainty\-Guided Visual Prompt Tuning,”*and*“VAST: Velocity\-Adaptive Spatially\-varying Timesteps\.”*Title vocabulary leans toward method\-name modifiers \(adaptive,aware,guided,dynamic,gradient\)\. Despite the system\-builder framing, Kimi Code writes the shortest papers \(2,461 words\) with by far the fewest figures \(0\.8 on average, sometimes none, vs\. Claude Code’s 4\.8\) and substitutes formal cues \(the most equations at 4\.0 and the most theorems at 0\.4\) for visual evidence\.
### 4\.3Programming language and time usage
Programming\-language usage\.We further conduct analysis on the programming language of the experiments, shown in Figure[4](https://arxiv.org/html/2605.19156#S4.F4)\(left\)\. We find that all three agents overwhelmingly default to Python regardless of the research domain with the remainder all shell scripts\. Notably, we find zero C/C\+\+/Rust/Go files in any agent’s output, even on CPU\-only seeds where those languages would be more idiomatic \(e\.g\., C/C\+\+ for operating system design\)\.
Wall\-clock time per pipeline stage\.We analyze wall\-clock time by pipeline stage for each agent in Figure[4](https://arxiv.org/html/2605.19156#S4.F4)\(right\)\. All three agents spend the majority of their time on experiments, where Claude Code \(13\.0h total\) is roughly 3×\\timesslower than Kimi Code \(4\.1h\) and 2×\\timesslower than Codex \(6\.8h\)\. This is consistent with Kimi Code’s higher fabrication rate: it does not fully use its compute budget for conducting experiments\. Claude Code’s longer experimentation time aligns with its lowest underpowered and plan/execution\-mismatch rates in §[5\.2](https://arxiv.org/html/2605.19156#S5.SS2)\. For ideation, Codex spends the most time; for paper writing, Claude Code takes the longest\. Notably, self\-refinement takes only a very small share of the total time, almost negligible compared with the other stages\.

\\phantomsubcaption

\\phantomsubcaption
Figure 4:Programming\-language usage \(left\) and wall\-clock time per stage \(right\)\.
## 5Limitations
In this section, we first show that SAR alone cannot be trusted as a reliable reviewer \(§[5\.1](https://arxiv.org/html/2605.19156#S5.SS1)\)\. Artifacts\-aware peer reviews and human inspections deliver the three failure modes \(§[5\.2](https://arxiv.org/html/2605.19156#S5.SS2)\)\. We then break the scores down by research domain and compute platform \(§[5\.3](https://arxiv.org/html/2605.19156#S5.SS3)\) and rule out compute as the bottleneck \(§[5\.4](https://arxiv.org/html/2605.19156#S5.SS4)\)\. A self\-refinement and reviewer\-severity\-drift analysis is in Appendix[G](https://arxiv.org/html/2605.19156#A7)\.
### 5\.1SAR cannot be trusted in isolation
Table 3:SAR vs\. human review\.We find in Table[3](https://arxiv.org/html/2605.19156#S5.T3)thatSAR is a weaker discriminator than human reviewers\. Comparing SAR scores to the average human review score for each of the same 200 ICLR papers, the human accept\-vs\-reject score gap is 1\.52 points \(6\.54 vs\. 5\.02\), but SAR compresses that gap to only 0\.25 points \(5\.59 vs\. 5\.34\)\. Because SAR does not provide an accept/reject decision, we manually inspect every SAR review and label each paper \(Appendix[H](https://arxiv.org/html/2605.19156#A8)\)\. The resulting acceptance rates make the same point: SAR accepts 76% of human\-accepted ICLR papers and 52% of human\-rejected ones, but only 41% of Claude Code’s, 22% of FARS’s, 13% of Codex’s, and 5% of Kimi Code’s\. Mean scores overstate how close agents are to top\-tier acceptance; the underlying acceptance gap is much larger, and SAR cannot be the sole evaluator of agent\-generated papers\.
### 5\.2Artifacts\-aware peer review surfaces three failure modes

Figure 5:PR breakdown scores\.
Figure 6:Fabricated results\.
Figure 7:SAR vs\. PR\.Under artifacts\-aware PR review, every agent’s score drops below its SAR score \(Figure[7](https://arxiv.org/html/2605.19156#S5.F7)\): Claude Code−0\.85\-0\.85, Codex−0\.42\-0\.42, Kimi Code−0\.86\-0\.86\. Through per\-dimension PR scores \(Figure[5](https://arxiv.org/html/2605.19156#S5.F5)\) localise the drop: Codex leads on every reliability\-leaning dimension \(reproducibility, references, reference and results integrity\), Claude Code leads on creative dimensions \(novelty, significance\), Kimi Code lags on every dimension simultaneously, andexperimental rigor is the lowest dimension across all agents\. To further investigate the experiment rigor problems\. We manually verify three failure modes \(*fabricated results*,*underpowered experiments*, and*plan/execution mismatch*\)\.
Fabricated results\.We classify fabricated results into 4 categories:*results mismatch only*\(numbers reported in the paper do not matchresults\.json, logs, or experiment outputs\),*setting mismatch only*\(the paper claims components not implemented in the code, or hyperparameters in the text differ from the config\),*both*\(the paper exhibits both results and setting mismatches\), and*fake reference*\(citations that do not exist, have fabricated authors, or have incorrect bibliographic metadata\)\. As shown in Figure[6](https://arxiv.org/html/2605.19156#S5.F6), Kimi Code shows by far the highest rates \(77% paper\-vs\-artifact mismatch, 72% fake references\): it invents experimental results directly \([Case 4](https://arxiv.org/html/2605.19156#A9.SS4)\) or reports baselines that were never run \([Case 5](https://arxiv.org/html/2605.19156#A9.SS5)\)\. Claude Code follows at 31%/36% \(occasional fabrication when experiments fail;[Case 2](https://arxiv.org/html/2605.19156#A9.SS2)\); Codex stays clean at 5%/8%\.
Underpowered experiments\.A paper is flagged as*underpowered*by having limited experiments\. For example, a single small dataset where multiple are expected, one model size where a ladder is expected, or one random seed for what should be a stochastic comparison, or when the paper is explicitly framed as a pilot or feasibility study with limited evidence\. As shown in Table[4](https://arxiv.org/html/2605.19156#S5.T4), Kimi Code 82\.1%\>\>Codex 41\.0%\>\>Claude Code 25\.6%; for every agent the rate is higher on GPU than on CPU \(Claude Code 33% vs\. 13%, Codex 42% vs\. 40%, Kimi Code 92% vs\. 67%\), consistent with GPU work being broader in scope than the CPU\-only seeds\. Codex’s papers are often explicitly framed as pilot/feasibility studies \([Case 3](https://arxiv.org/html/2605.19156#A9.SS3)\), which reduces fabrication but limits the empirical evidence\.
Plan/execution mismatch\.During ideation, each agent is aware of the resources \(i\.e\., hardware and time budget\) to write an experimental plan based on the proposed ideas\. We define*plan/execution mismatch*as cases where the executed artifacts diverge from that plan or from the manuscript that follows it: a baseline named in the plan but missing from the code, an ablation specified but never run\. As shown in Table[4](https://arxiv.org/html/2605.19156#S5.T4), Kimi Code 33\.3%\>\>Codex 20\.5%\>\>Claude Code 17\.9%; the CPU\-vs\-GPU pattern differs by agent \(Claude Code 13% / 21%, Codex 33% / 13%, Kimi Code 33% / 33%\): Claude Code’s mismatch concentrates on GPU work, Codex’s on CPU, while Kimi Code splits evenly\. Kimi Code plans the most experiments \(13\.2/trial\), and its overambitious planning exceeds what it can execute, producing the highest plan/execution mismatch \(33\.3%\) and underpowered rate \(82\.1%\)\. Codex plans the most conservatively \(5\.7/trial\), which keeps mismatch low \(20\.5%\) but inflates the underpowered rate \(41\.0%\) via frequent pilot/feasibility framing\. Claude Code plans moderately \(10\.6/trial\) and has the lowest rates on both axes \(17\.9% mismatch, 25\.6% underpowered\)\. A shared failure across all agents is the tendency to compare against*older*baselines rather than recent ones, even when newer baselines are mentioned in the related\-work section\.
Besides the three failure modes, we observe that PR review and SAR review both credit the honesty towards negative results, where human reviewers would not credit this as a strength\.
Table 4:Per\-agent breakdown of*underpowered*and*plan/execution mismatch*ratios\.
### 5\.3Per\-domain analysis
Per\-domain breakdown\.Figure[8](https://arxiv.org/html/2605.19156#S5.F8)shows mean PR scores per agent across the 13 research domains \(the parallel SAR breakdown is in Appendix[D](https://arxiv.org/html/2605.19156#A4)due to the space limit\)\. Patterns vary by agent\. Claude Code peaks on Probabilistic Methods \(5\.32\) and Computer Vision \(5\.10\) but dips to∼\\sim3\.78 on Privacy in ML and Supervised Repr\. Learning\. Codex stays in a tighter band \(4\.20–4\.89\), with its highest mark on Supervised Repr\. Learning \(4\.89\)\. Kimi Code is consistently the lowest, with its weakest scores on Generative Models \(2\.44\) and Privacy in ML \(2\.66\)\.
Figure 8:Per\-domain mean PR scores by agent across the 13 research domains\.CPU vs\. GPU: opposite trends in SAR and PR\.PR and SAR move in opposite directions across the CPU/GPU split\(Table[9](https://arxiv.org/html/2605.19156#A4.T9)in Appendix[D](https://arxiv.org/html/2605.19156#A4)\)\. Under*PR*, all three agents score higher on CPU than on GPU \(Claude Code\+0\.26\+0\.26, Codex\+0\.03\+0\.03, Kimi Code\+0\.50\+0\.50\), with Kimi Code showing the largest gap\. Under*SAR*, Codex and Kimi Code score*higher*on GPU \(Codex−0\.61\-0\.61, Kimi Code−0\.20\-0\.20\), while only Claude Code is roughly platform\-invariant\. GPU domains \(vision, NLP, generative models\) are well\-established fields where agents can produce better\-looking papers \(more polished prose, more figures, familiar baselines\), but GPU experiments are also harder to execute correctly: CUDA issues, memory limits, and training instabilities lead to more incomplete runs and mismatched results when reviewers verify the code\. CPU tasks are simpler to run and verify, yielding more reliable experiments\. This divergence further illustrates that SAR alone is insufficient: it rewards presentation quality over experimental substance\.
### 5\.4Compute is not the bottleneck
We re\-run all 8 GPU seeds with Codex on 8×\\timesNVIDIA H100 \(80 GB\) for 3 trials each, with budget matched to the A6000 runs\. The result is no consistent improvement: Codex PR drops from 4\.51 \(A6000\) to 4\.26 \(H100\), confirming that the limiting factor is not compute but the agent’s experiment design capabilities\. The per\-domain H100 vs\. A6000 breakdown is in Appendix[F](https://arxiv.org/html/2605.19156#A6)\.
## 6Future Directions
Can we trust agentic reviewers for agent\-generated papers?SAR and PR both over\-credit agent\-generated papers relative to human reviewers \(§[5\.1](https://arxiv.org/html/2605.19156#S5.SS1)\)\. Future automated reviewers should combine with principled calibration against human review\.
Faithfulness over complex tasks\.Frontier model providers increasingly advertise faithfulness as a core capability of their agents, yet under our open\-ended end\-to\-end research setting we still observe substantial fabrication \(§[5\.2](https://arxiv.org/html/2605.19156#S5.SS2)\)\. The claim of faithful behaviour does not yet survive contact with sufficiently complex tasks\. Future work should focus on training agents to be faithful end\-to\-end rather than only on individual reasoning traces\.
Better experiment\-planning agents and scaffolds\.The major challenge for auto\-research is experimental rigor \(§[5\.2](https://arxiv.org/html/2605.19156#S5.SS2)\)\. Closing this gap will require improvements: stronger agent capabilities and scaffolds that harness agents for designing and executing rigorous experiments end\-to\-end\.
## 7Conclusion
In this paper, we systematically investigate the auto\-research capabilities and limitations of three frontier agents across 13 CS domains using a minimal scaffold, ResearchArena\. We find that experimental rigor is the number\-one weakness: agents routinely fail to plan, execute, and faithfully report experiments, limiting both the scope and significance of their papers\. Fabricated results and a manuscript\-only reviewer that systematically favours honest but narrow framings further raise faithfulness concerns for today’s frontier models\. In terms of paper quality, all current agents still fall well short of the threshold for top\-tier venues\. There is still a long way to go for true auto\-research\.
## References
- \[1\]Analemma Intelligence\(2026\)Introducing fars: fully automated research system\.Note:[https://analemma\.ai/blog/introducing\-fars/](https://analemma.ai/blog/introducing-fars/)Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1),[§2](https://arxiv.org/html/2605.19156#S2.p1.1),[Table 1](https://arxiv.org/html/2605.19156#S4.F2.2.1.6.5.1),[§4\.1](https://arxiv.org/html/2605.19156#S4.SS1.p1.9)\.
- \[2\]Anthropic\(2026\)Claude opus 4\.6\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-6](https://www.anthropic.com/news/claude-opus-4-6)Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p2.1)\.
- \[3\]J\. Baek, S\. K\. Jauhar, S\. Cucerzan, and S\. J\. Hwang\(2025\)Researchagent: iterative research idea generation over scientific literature with large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 6709–6738\.Cited by:[Appendix B](https://arxiv.org/html/2605.19156#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2605.19156#S1.p1.1),[§2](https://arxiv.org/html/2605.19156#S2.p1.1),[§3\.1](https://arxiv.org/html/2605.19156#S3.SS1.p1.1)\.
- \[4\]J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace, K\. Liu, L\. Maksin, T\. Patwardhan,et al\.\(2024\)Mle\-bench: evaluating machine learning agents on machine learning engineering\.arXiv preprint arXiv:2410\.07095\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[5\]H\. Chen, M\. Xiong, Y\. Lu, W\. Han, A\. Deng, Y\. He, J\. Wu, Y\. Li, Y\. Liu, and B\. Hooi\(2025\)Mlr\-bench: evaluating ai agents on open\-ended machine learning research\.arXiv preprint arXiv:2505\.19955\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p2.1),[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[6\]Z\. Chen, S\. Chen, Y\. Ning, Q\. Zhang, B\. Wang, B\. Yu, Y\. Li, Z\. Liao, C\. Wei, Z\. Lu,et al\.\(2024\)Scienceagentbench: toward rigorous assessment of language agents for data\-driven scientific discovery\.arXiv preprint arXiv:2410\.05080\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[7\]Q\. Huang, J\. Vora, P\. Liang, and J\. Leskovec\(2023\)Mlagentbench: evaluating language agents on machine learning experimentation\.arXiv preprint arXiv:2310\.03302\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[8\]C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. R\. Narasimhan\(2023\)Swe\-bench: can language models resolve real\-world github issues?\.InThe twelfth international conference on learning representations,Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[9\]A\. Karpathy\(2026\)Auto research\.Note:[https://github\.com/karpathy/autoresearch](https://github.com/karpathy/autoresearch)Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p1.1)\.
- \[10\]C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. Ha\(2024\)The ai scientist: towards fully automated open\-ended scientific discovery\.arXiv preprint arXiv:2408\.06292\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1),[§2](https://arxiv.org/html/2605.19156#S2.p1.1)\.
- \[11\]Moonshot AI\(2026\)Kimi k2\.5\.Note:[https://www\.kimi\.com/ai\-models/kimi\-k2\-5](https://www.kimi.com/ai-models/kimi-k2-5)Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p2.1)\.
- \[12\]R\. Nakano, J\. Hilton, S\. Balaji, J\. Wu, L\. Ouyang, C\. Kim, C\. Hesse, S\. Jain, V\. Kosaraju, W\. Saunders,et al\.\(2021\)Webgpt: browser\-assisted question\-answering with human feedback\.arXiv preprint arXiv:2112\.09332\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[13\]OpenAI\(2026\)Introducing gpt\-5\.4\.Note:[https://openai\.com/index/introducing\-gpt\-5\-4/](https://openai.com/index/introducing-gpt-5-4/)Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p2.1)\.
- \[14\]S\. Peyton Jones\(2017\)How to write a great research paper\.Note:[https://www\.microsoft\.com/en\-us/research/academic\-program/write\-great\-research\-paper/](https://www.microsoft.com/en-us/research/academic-program/write-great-research-paper/)Cited by:[Appendix B](https://arxiv.org/html/2605.19156#A2.SS0.SSS0.Px4.p1.1),[§3\.1](https://arxiv.org/html/2605.19156#S3.SS1.p1.1)\.
- \[15\]T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. Scialom\(2023\)Toolformer: language models can teach themselves to use tools\.Advances in neural information processing systems36,pp\. 68539–68551\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[16\]S\. Schmidgall and M\. Moor\(2025\)Agentrxiv: towards collaborative autonomous research\.arXiv preprint arXiv:2503\.18102\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1),[§2](https://arxiv.org/html/2605.19156#S2.p1.1)\.
- \[17\]S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum\(2025\)Agent laboratory: using llm agents as research assistants\.Findings of the Association for Computational Linguistics: EMNLP 2025,pp\. 5977–6043\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1),[§2](https://arxiv.org/html/2605.19156#S2.p1.1)\.
- \[18\]J\. Schulman\(2020\)An opinionated guide to ML research\.Note:[http://joschu\.net/blog/opinionated\-guide\-ml\-research\.html](http://joschu.net/blog/opinionated-guide-ml-research.html)Cited by:[Appendix B](https://arxiv.org/html/2605.19156#A2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2605.19156#S3.SS1.p1.1)\.
- \[19\]C\. Si, D\. Yang, and T\. Hashimoto\(2024\)Can llms generate novel research ideas? a large\-scale human study with 100\+ nlp researchers\.arXiv preprint arXiv:2409\.04109\.Cited by:[Appendix B](https://arxiv.org/html/2605.19156#A2.SS0.SSS0.Px1.p1.1)\.
- \[20\]Z\. S\. Siegel, S\. Kapoor, N\. Nagdir, B\. Stroebl, and A\. Narayanan\(2024\)Core\-bench: fostering the credibility of published research through a computational reproducibility agent benchmark\.arXiv preprint arXiv:2409\.11363\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[21\]Stanford ML Group\(2025\)Stanford agentic reviewer\.Note:[https://paperreview\.ai/](https://paperreview.ai/)Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p2.1),[§3\.3](https://arxiv.org/html/2605.19156#S3.SS3.p1.1)\.
- \[22\]G\. Starace, O\. Jaffe, D\. Sherburn, J\. Aung, J\. S\. Chan, L\. Maksin, R\. Dias, E\. Mays, B\. Kinsella, W\. Thompson,et al\.\(2025\)PaperBench: evaluating ai’s ability to replicate ai research\.arXiv preprint arXiv:2504\.01848\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[23\]G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. Anandkumar\(2023\)Voyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[24\]L\. Wang, C\. Ma, X\. Feng, Z\. Zhang, H\. Yang, J\. Zhang, Z\. Chen, J\. Tang, X\. Chen, Y\. Lin,et al\.\(2024\)A survey on large language model based autonomous agents\.Frontiers of Computer Science18\(6\),pp\. 186345\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[25\]H\. Wijk, T\. Lin, J\. Becker, S\. Jawhar, N\. Parikh, T\. Broadley, L\. Chan, M\. Chen, J\. Clymer, J\. Dhyani,et al\.\(2024\)Re\-bench: evaluating frontier ai r&d capabilities of language model agents against human experts\.arXiv preprint arXiv:2411\.15114\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[26\]Y\. Yamada, R\. T\. Lange, C\. Lu, S\. Hu, C\. Lu, J\. Foerster, J\. Clune, and D\. Ha\(2025\)The ai scientist\-v2: workshop\-level automated scientific discovery via agentic tree search\.arXiv preprint arXiv:2504\.08066\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1),[§2](https://arxiv.org/html/2605.19156#S2.p1.1)\.
- \[27\]J\. Yang, C\. E\. Jimenez, A\. Wettig, K\. Lieret, S\. Yao, K\. Narasimhan, and O\. Press\(2024\)Swe\-agent: agent\-computer interfaces enable automated software engineering\.Advances in Neural Information Processing Systems37,pp\. 50528–50652\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[28\]S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao\(2022\)React: synergizing reasoning and acting in language models\.arXiv preprint arXiv:2210\.03629\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
- \[29\]Y\. Zhang, M\. Khalifa, S\. Bhushan, G\. D\. Murphy, L\. Logeswaran, J\. Kim, M\. Lee, H\. Lee, and L\. Wang\(2025\)MLRC\-bench: can language agents solve machine learning research challenges?\.arXiv preprint arXiv:2504\.09702\.Cited by:[§2](https://arxiv.org/html/2605.19156#S2.p2.1)\.
- \[30\]S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried,et al\.\(2023\)Webarena: a realistic web environment for building autonomous agents\.arXiv preprint arXiv:2307\.13854\.Cited by:[§1](https://arxiv.org/html/2605.19156#S1.p1.1)\.
## Appendix ALimitations
Due to budget constraints, our study evaluates only three agents and therefore does not cover the full space of available agentic coding systems\. We also focus on computer\-science research domains, leaving evaluation in other scientific and engineering fields to future work\. Finally, our end\-to\-end analysis is conducted under our specific ResearchArena setting; although this setting reflects many emerging auto\-research workflows, the findings may not fully generalize to all possible auto\-research systems, scaffolds, or evaluation protocols\.
## Appendix BPer\-stage guidelines
Each pipeline stage is anchored by a domain\-aware guideline document\. We reproduce the ML\-domain guidelines below verbatim as the canonical reference; the other five domain families \(systems, databases, PL, theory, security\) follow the same structure with domain\-specific phrasing\.
#### Ideation guidelines \(stage 1\)\.
The agent proposes a research idea framed as a hypothesis with a falsifiable prediction; output isidea\.json\. Distilled from Schulman’s research advice\[[18](https://arxiv.org/html/2605.19156#bib.bib31)\], the human study of LLM ideation by Si et al\.\[[19](https://arxiv.org/html/2605.19156#bib.bib32)\], and the ResearchAgent methodology\[[3](https://arxiv.org/html/2605.19156#bib.bib21)\]\.
\#IdeaGenerationGuidelines
Howtogofromaseedfieldtoanovel,feasibleresearchidea\.
DistilledfromJohnSchulman's"OpinionatedGuidetoMLResearch",the
ResearchAgentmethodology,"CanLLMsGenerateNovelResearchIdeas?"\(Sietal\.\),
andstandardacademicresearchpractices\.
\#\#Step1:Explorethefield
Startbyunderstandingwhatalreadyexists\.DONOTskipthisstep\.
\#\#\#Searchforexistingwork\(newesttooldest\)
\-SearcharXiv\(arxiv\.org\),SemanticScholar\(semanticscholar\.org\),and
GoogleScholar\(scholar\.google\.com\)forpapersinyourseedfield
\-\*\*Startwiththenewestpapersfirst\*\*âsortbydate,readthemost
recentworkbeforegoingtoolderfoundationalpapers\.Thisensures
youknowthecurrentfrontierbeforeproposingsomethingnew\.
\-Recommendedsearchorder:
1\.Last6monthsâwhat'shappeningrightnow?
2\.Last1\-2yearsâwhatarethecurrentstate\-of\-the\-artmethods?
3\.Foundationalpapersâwhataretheclassicapproaches?
\-Lookfor:
\-Surveypapersâtheysummarizethelandscapeandlistopenproblems
\-Highly\-citedrecentpapersâtheydefinethecurrentstateoftheart
\-Workshoppapersâtheyoftencontainearly\-stageideasandemergingtrends
\#\#\#Buildamentalmap
\-Whatarethemainapproachesinthisarea?
\-Whataretheestablishedbenchmarksandmetrics?
\-Whataretheknownlimitationsofcurrentmethods?
\-Whatproblemsareconsidered"open"or"unsolved"?
\-WhatrecenttechniquesfromOTHERfieldscouldapplyhere?
\#\#\#Findthegaps
\-Readthe"Limitations"and"FutureWork"sectionsofrecentpapers
\-Lookforrecurringcomplaintsinreviews\(onOpenReview,ifavailable\)
\-Identifyassumptionsthatcurrentmethodsmakeâcanyourelaxthem?
\-Lookforproblemswheresimplebaselinesstillperformsurprisinglywell
\(thissignalsthecommunityhasn'tcrackedityet\)
\#\#Step2:Generatecandidateideas
\#\#\#Twoapproaches\(chooseoneorcombine\)
\*\*Goal\-driven\*\*\(recommended\):Startwithaproblemyouwanttosolve\.
\-"CurrentmethodsforXfailwhenYhappens\.Howcanwefixthat?"
\-"TaskZrequirestoomuchlabeleddata\.Canwedoitwithless?"
\-Thegoalconstrainsyoursearchandmakesthecontributionclear\.
\*\*Idea\-driven\*\*:Startwithatechniqueandfindwhereitapplies\.
\-"TechniqueAworkswellforB\.CoulditalsoworkforC?"
\-Riskierâyoumayfindtheideaalreadyexistsordoesn'twork\.
\#\#\#Whatmakesagoodresearchidea
\-\*\*Novel\*\*:Notalreadydone\.YouMUSTverifythis\(Step3\)\.
\-\*\*Feasible\*\*:Canbeimplementedandtestedwithinyourresourceconstraints\.
\-\*\*Clear\*\*:Thecontributioniseasytoexplaininonesentence\.
\-\*\*Testable\*\*:There'saconcretewaytoevaluatewhetheritworks\.
\-\*\*Significant\*\*:Ifitworks,thecommunitywouldcare\.
\#\#\#WhatmakesaBADresearchidea
\-Toobroad\("improveNLP"\)âneedsaspecificproblemandapproach
\-Tooincremental\("changehyperparameterXfrom0\.1to0\.01"\)
\-Notverifiable\(nowaytotestifitworked\)
\-Requiresresourcesyoudon'thave\(100GPUs,proprietarydata\)
\-Alreadyexists\(youdidn'tchecktheliterature\)
\#\#Step3:Verifynovelty\(CRITICALâdonotskip\)
Beforecommittingtoanidea,verifyithasn'tbeendone:
\#\#\#Searchspecificallyforyouridea
\-SearchSemanticScholarandarXivwithkeywordsfromyourproposedmethod
\-SearchforthePROBLEMyou'resolving,notjustyourapproach
\-Checkifyourideaisaspecialcaseofsomethingmoregeneralthatexists
\-Lookatthe"RelatedWork"sectionsofpapersclosesttoyouridea
\#\#\#Commonnoveltytraps
\-Yourideaexistsbutunderadifferentname\(jargonvariesacrosssubfields\)
\-Yourideawastriedanddidn'twork\(checkfornegativeresultstoo\)
\-Yourideaisaminorvariationofanexistingapproach
\-Aconcurrentpaper\(postedinthelastfewmonths\)doesthesamething
\#\#\#Ifyourideaalreadyexists
\-DON'Tgiveupimmediately\.Ask:canyouimproveonit?Applyittoa
newdomain?Combineitwithsomethingelse?Scaleitup?
\-Ifittrulyexistswithnoroomforimprovement,gobacktoStep2
\#\#Step4:Produceoutputs
YoumustproduceTHREEoutputsinthisstep:
\#\#\#4\.1proposal\.mdâResearchProposal
Athoroughdocumentwiththesesections:
\-\*\*Introduction\*\*:Context,problemstatement,keyinsight,hypothesis
\-\*\*ProposedApproach\*\*:Overview,methoddetails,keyinnovations
\-\*\*RelatedWork\*\*:Keypapers,howyourideadiffers,positioning
\-\*\*Experiments\*\*:Plannedsetup,benchmarks,metrics,expectedresults
\-\*\*SuccessCriteria\*\*:Whatwouldconfirmorrefuteyourhypothesis
\-\*\*References\*\*:Fullcitationlist\(allmustbereal,verifiablepapers\)
\#\#\#4\.2idea\.jsonâStructuredSummary
AJSONobjectwithatleastthesefields:
\-\*\*description\*\*:1\-3sentencesexplainingwhatyou'reproposing
\-\*\*title\*\*:papertitle
\-\*\*motivation\*\*:whythisproblemmatters,whatgapyou'refilling
\-\*\*proposed\_approach\*\*:yourhigh\-levelmethodandwhyitshouldwork
\-\*\*related\_work\*\*:keyexistingpapersandhowyourideadiffers
\(useREALpapersyoufoundinSteps1and3âincludetitlesandauthors\)
\-\*\*hypothesis\*\*:testablehypothesis
\-\*\*success\_criteria\*\*:whatwouldconfirm/refutethehypothesis
\#\#\#4\.3references/âParsedReferencePapers
Createadirectorywithkeyreferencepapers\.Foreachpaper,createa
subdirectorycontainingthepaper'scontentparsedintosections:
\`\`\`
references/
âââPaper\-Title\-One/
ââââmeta/
âââââmeta\_info\.txt\#title,authors,venue,year,URL
âââââbibtex\.txt\#BibTeXentry
ââââsections/
ââââabstract\.md
ââââ1Introduction\.md
ââââ2RelatedWork\.md
ââââ\.\.\.
âââPaper\-Title\-Two/
ââââ\.\.\.
\`\`\`
Thisgroundsyourproposalinrealliteratureandensuresreferencesare
verifiablebyreviewers\.
\#\#\#Sanitychecksbeforemovingon
\-Canyouexplaintheideainonesentencetoanon\-expert?
\-Isthereaclearexperimentthatwouldtesttheidea?
\-Doyouhavetheresources\(data,compute,time\)todoit?
\-Istheexpectedcontributionlargeenoughforapaper?
\-Areallreferencesreal,verifiablepublications?
\#\#Generalprinciples
\#\#\#FromJohnSchulman
\-Yourabilitytochoosetherightproblemismoreimportantthanrawskill
\-Watchwhichideasprosperandwhichareforgottenâthisdevelopstaste
\-Goal\-drivenresearchhaslowerscoopingriskthanidea\-drivenresearch
\-There'snoshameinworkingonideassuggestedbyothersorbytheliterature
\#\#\#From"CanLLMsGenerateNovelResearchIdeas?"\(Sietal\.\)
\-AI\-generatedideastendtobenovelbutlackfeasibilityâgroundyours
inpracticalconstraints
\-Vagueimplementationdetailsarethe\#1weaknessâbespecificabouthow
yourmethodactuallyworks
\-Missingbaselinesandunrealisticassumptionsarecommonfailures
\-Verifyyourideaagainstexistingworkâ80%ofreviewerrejections
citeexistingpapersthattheauthorsmissed
\#\#\#FromResearchAgent\(Baeketal\.\)
\-Connectideasacrosspapers,notjustwithinonepaper
\-Lookforsharedconceptsacrossdifferentsubfields
\-Iterativerefinementimprovesideaqualityâbutdiminishingreturns
after2\-3rounds
\-BothcitationrelationshipsANDunderlyingconceptsmatterfornovelty
#### Plan guidelines \(stage 2, planning\)\.
The agent converts the idea into a concrete experiment plan covering datasets, baselines, metrics, expected outcomes under the hypothesis, ablations, and statistical tests; output isplan\.json\.
\#ExperimentPlanGuidelines
Howtodesignarigorousexperimentplanbasedonyourresearchproposal\.
Readproposal\.mdandidea\.jsonfirstâyourplanmusttesttheclaims
andhypothesisdescribedthere\.
\#\#PlanFormat
Saveyourplanasplan\.jsonâaJSONarrayofexperimentsteps:
\`\`\`json
\[
\{
"category":"<category\>",
"title":"shortdescriptivetitle",
"description":"whatthisstepdoesandwhy",
"steps":\{
"step1":"detailedinstructionwithspecifics",
"step2":"\.\.\.",
\.\.\.
\}
\},
\.\.\.
\]
\`\`\`
Suggestedcategories\(addyourownasneeded\):
\-\*\*EnvironmentConfiguration\*\*âdependencies,setup
\-\*\*DataPreparation\*\*âdownload,preprocess,splits,statistics
\-\*\*BaselineExperiment\*\*âexistingmethodstocompareagainst
\-\*\*MainExperiment\*\*âyourproposedmethod
\-\*\*AnalysisExperiment\*\*âablations,robustness,sensitivity
\-\*\*EffectivenessEvaluation\*\*âsuccesscriteria,statisticaltests
\-\*\*Visualization\*\*âfigures,tables,plotsforthepaper
\#\#ExperimentDesignPrinciples
\#\#\#Formulatetestableclaims
\-Stateyourhypothesisasatestableclaim
\-Designexperimentsthatcouldfailâiftheycan'tproduceanegative
result,they'renotinformative
\-DefinewhatwouldDISPROVEyourclaim
\#\#\#Choosetherightexperimenttype
\|Claimtype\|Experimenttype\|Whattomeasure\|
\|\-\-\-\|\-\-\-\|\-\-\-\|
\|"OurmethodoutperformsX"\|Empiricalcomparison\|Metricsonsharedbenchmarks\|
\|"ComponentAiscritical"\|Ablationstudy\|Performancewith/withoutA\|
\|"Thisscalesbetter"\|Scalingexperiment\|Performancevs\.data/compute/params\|
\|"OurtheorypredictsX"\|Theoreticalvalidation\|Syntheticsetupwithknowngroundtruth\|
\|"Thispropertyholds"\|Analysis/probing\|Measurementsonexistingmodels/data\|
\|"Thisisfaster/cheaper"\|Systemsexperiment\|Latency,throughput,memory,FLOPs\|
\|"Thisbenchmarkisbetter"\|Benchmarkevaluation\|Existingmethodsonnewbenchmark\|
\#\#\#Selectmetricscarefully
\-Usestandardmetricsforyourtask
\-ReportALLstandardmetrics,notjusttheonewhereyouwin
\-ConsiderbothperformanceANDcost\(FLOPs,latency,memory\)
\#\#\#Choosedatasetsthattestyourclaim
\-Usestandardbenchmarkswhenpossible
\-Ifclaimingrobustness,testondistribution\-shifteddata
\-Ifclaiminggeneralization,testonmultipledatasets
\-Document:datasource,size,splits,preprocessing
\#\#\#Selectbaselinesfairly
\-Atleast2meaningfulbaselines\(onesimple,onestrong/recent\)
\-Runallbaselineswithequivalenteffort\(samecompute,sametuning\)
\-Nevercompareagainstintentionallyweakbaselines
\#\#\#Planablationstudies
\-Foreachnovelcomponent,plantoremoveitandmeasureimpact
\-PlanablationsBEFORErunningexperiments,notafterseeingresults
\#\#\#Thinkaboutconfounders
\-Couldtheimprovementcomefrommoreparameters/data/compute?
\-Arecomparisonsfair\(samepreprocessing,splits,computebudget\)?
\-Ifusingpublishedbaselines,arethesetupstrulycomparable?
\#\#\#Rigorousevaluation
\-Useafixedrandomseedforreproducibility
\-UsetheSAMEseedformethodandallbaselines\(faircomparison\)
\-Avoiddataleakage\(preprocessingstatsfromtrainonly,etc\.\)
\-TestsetusedONCEforfinalevaluation,notformodelselection
\#\#\#Commonpitfallstoavoid
\-Don'ttunehyperparametersonthetestset
\-Don'tcompareagainstbaselineswithdifferentpreprocessingorsplits
\-Don'treportonlythemetricwhereyouwin
\-Don'tclaimSOTAwithoutcomparingagainstactualSOTAmethods
\-Don'tignorenegativeresultsâreportthemhonestly
\-Don'tuseasingletrain/testsplit
\-Don'tassumedeeplearningisalwaysbetterâtestsimpleralternatives
\#\#FeasibilityandRuntimeEstimation
Beforefinalizingtheplan,estimatewhetheritfitstheresourcebudget:
1\.\*\*Estimateper\-experimentruntime\*\*:Howlongdoesonetrainingrun/
evaluation/APIcalltake?Usepublishedbenchmarksorroughestimates
basedonmodelsizeanddatasetsize\.
2\.\*\*Countindependentexperiments\*\*:Howmanybaselines,seeds,ablations,
anddatasets?Identifywhichcanruninparallel\.
3\.\*\*Dividebyavailableparallelism\*\*:IfyouhaveKGPUsandNindependent
GPUexperiments,parallelruntimeâN/KÃper\-experimenttime\.ForCPU
tasks,dividebyavailablecores\.
4\.\*\*Comparetobudget\*\*:Ifestimatedparallelruntimeexceedsthetime
budget,simplifytheplan\(fewerseeds,smallermodels,fewerdatasets\)
BEFOREfinalizingânotduringexecution\.
Example:
\-5baselines\+1method\+3ablations=9experiments
\-Eachrun:~30minon1GPU
\-With8GPUs:9/8â2batchesÃ30min=~1hour
\-Budgetis8hoursâplentyofroomforvisualization\+analysis
\#\#PlanQualityChecklist
Beforefinalizingplan\.json,verify:
\-\[\]Eachstephasaclearcategory,title,description,andsub\-steps
\-\[\]Sub\-stepsaredetailedenoughtofollowwithoutambiguity
\-\[\]Specificdatasets,metrics,andhyperparametersarenamed
\-\[\]Atleast2meaningfulbaselinesareincluded
\-\[\]Ablationstudiesplannedforeachnovelcomponent
\-\[\]Fixedrandomseedforreproducibility
\-\[\]Successcriteriaclearlydefined
\-\[\]Runtimeestimatefitstheresourcebudgetwithparallelexecution
\-\[\]PlanaccountsforALLavailableGPUsandCPUs
\-\[\]Visualizationstepincludedforpaperfigures
#### Experiment guidelines \(stage 2, execution\)\.
The agent writes self\-contained Python underexp/that prepares data, trains baselines and the proposed method, runs ablations, and writesresults\.jsoncontaining all reported metrics; per\-command logs are kept\.
\#ExperimentExecutionGuidelines
Howtoexecuteyourexperimentplanefficientlyandrigorously\.
Experimentdesignprinciplesareinplan\_guidelines\.mdâyoushould
havealreadyappliedthemwhencreatingplan\.json\.
\#\#Phase1:MaximizeResourceUsage
YourgoalistouseALLavailableresourcesefficiently\.Checkwhatyou
havebeforestarting:
\`\`\`bash
nvidia\-smi\#GPUs:count,memory,currentusage
nproc\#CPUcores
free\-h\#RAM
\`\`\`
\#\#\#Parallelexecutionstrategy
Identifyindependentexperimentsinyourplan\(differentseeds,baselines,
ablations,datasets\)andrunthemsimultaneously:
\`\`\`bash
\#Example:8GPUs,6independentexperiments
CUDA\_VISIBLE\_DEVICES=0pythonexp/baseline1/run\.py&
CUDA\_VISIBLE\_DEVICES=1pythonexp/baseline2/run\.py&
CUDA\_VISIBLE\_DEVICES=2pythonexp/method/run\.py&
CUDA\_VISIBLE\_DEVICES=3pythonexp/ablation1/run\.py&
CUDA\_VISIBLE\_DEVICES=4pythonexp/ablation2/run\.py&
CUDA\_VISIBLE\_DEVICES=5pythonexp/dataset2/run\.py&
wait\#waitforalltofinish
\`\`\`
ForCPU\-boundwork\(datapreprocessing,evaluation,APIcalls\):
\`\`\`bash
\#RunmultipleCPUtasksinparallel
pythonexp/preprocess\_dataset1\.py&
pythonexp/preprocess\_dataset2\.py&
pythonexp/preprocess\_dataset3\.py&
wait
\`\`\`
\#\#\#GPUutilization
\-\*\*PinexperimentstoGPUs\*\*with\`CUDA\_VISIBLE\_DEVICES=N\`
\-IfamodelusesonlypartofGPUmemory,runmultipleexperiments
perGPU\(e\.g\.,2smallmodelsonone48GBGPU\)
\-IncreasebatchsizetofillGPUmemoryâlargerbatches=fastertraining
\-Forinference\-onlyexperiments\(embedding,evaluation\),consider
runningseveralonthesameGPU
\#\#\#CPUutilization
\-Use\`multiprocessing\`or\`concurrent\.futures\.ProcessPoolExecutor\`for
CPU\-bounddataprocessing
\-Parallelizedataloadingwith\`num\_workers\`inPyTorchDataLoaders
\-ForAPI\-basedexperiments\(LLMscoring\),use\`asyncio\`orthreadpools
tomakeconcurrentAPIcalls
\#\#\#Followtheplanâdonotscopedown
Yourplan\.jsonwasdesignedwiththeavailableresourcesinmind\.
ExecuteALLsteps\.Ifasteptrulycannotrun\(dependencyfailure,
out\-of\-memory\),documentitinthatstep's\`SKIPPED\.md\`andmoveon\.
DoNOTdropexperimentsjustbecauseasingle\-GPUpilotlooksslowâ
useparallelexecutionacrossallGPUs\.
\#\#\#Prioritizeexecutionorder
Runexperimentsindependencyorder:
1\.Datapreparation\(mustfinishfirst\)
2\.Baselines\+method\(canruninparallel\)
3\.Ablations\(canruninparallelaftermethodworks\)
4\.Analysis\+visualization\(afterresultsarein\)
\#\#Phase2:WorkspaceStructure
Organizeexperimentssothateachstephasitsownfolderwithcode,results,
andlogs\.Thismakesiteasytoverifywhichcodeproducedwhichresults\.
\`\`\`
exp/
âââ<experiment\_name\>/\#onefolderperexperiment/condition
âââârun\.py\#experimentscript
ââââconfig\.yaml\(or\.json\)\#hyperparameters,settings
ââââresults\.json\#per\-experimentresults
ââââlogs/\#training/evallogs,stdout
â
âââ<baseline\_name\>/
âââârun\.py
ââââresults\.json
ââââlogs/
â
âââ<ablation\_name\>/
âââârun\.py
ââââresults\.json
ââââlogs/
â
âââshared/\#sharedutilitiesacrossexperiments
âââdata\_loader\.py\#dataloading,preprocessing
âââmetrics\.py\#evaluationmetrics
âââmodels\.py\#modeldefinitions
âââutils\.py\#commonhelpers
data/\#downloaded/processeddatasets
figures/\#generatedfiguresforthepaper
\`\`\`
\#\#\#Per\-experimentresults
Each\`exp/<name\>/results\.json\`shouldcapturethatexperiment'soutput:
\`\`\`json
\{
"experiment":"<name\>",
"metrics":\{"metric1":\{"mean":0\.87,"std":0\.002\},\.\.\.\},
"config":\{"lr":0\.001,"epochs":50,"seed":42,\.\.\.\},
"runtime\_minutes":45
\}
\`\`\`
\#\#\#Figures
Savepublication\-readyfiguresto\`figures/\`:
\-Comparisonplots\(yourmethodvsbaselines\)
\-Ablationcharts\(impactofeachcomponent\)
\-Trainingcurves\(loss/metricoverepochs\)
\-Analysisvisualizations\(distributions,embeddings,etc\.\)
Eachfigureshouldbeself\-containedwithaxislabels,legends,andtitles\.
\#\#Phase4:PlanCompliance
Followplan\.jsonstepbystep:
\-Executeeverystepinorder
\-Createasubfolderunder\`exp/\`foreachplanstep
\-Ifastepisinfeasible,documentwhyinthatstep'sfolder\(createa
\`SKIPPED\.md\`withthereason\)andmoveon
\-Afterallsteps,verifythattheplan'ssuccesscriteriaaremet
\-Ifresultscontradictthehypothesis,reportthishonestlyânegative
resultswithgoodanalysisarevaluable
\#\#ReproducibilityChecklist
Beforefinishing,verify:
\-\[\]Fixedrandomseedsusedthroughout
\-\[\]Atleast2meaningfulbaselinescomparedfairly
\-\[\]Fixedrandomseedusedforreproducibility
\-\[\]Ablationstudyforeachnovelcomponent
\-\[\]Nodataleakage\(verified\)
\-\[\]Allconfigurationdocumentedinper\-experimentresults
\-\[\]Eachexperimenthasitsownfolderunderexp/withcodeandresults
\-\[\]Figuressavedforkeyresults
\-\[\]Negativeresultsreportedhonestly\(ifany\)
#### Paper\-writing guidelines \(stage 3\)\.
The agent produces a NeurIPS\-stylepaper\.texthat cites the items inreferences\.bib; reported numbers must matchresults\.jsonverbatim and the guideline explicitly forbids fabricating or extrapolating results\. Distilled from Peyton Jones’s research\-writing advice\[[14](https://arxiv.org/html/2605.19156#bib.bib33)\]and the formatting requirements of NeurIPS, ICML, and ICLR\.
\#PaperWritingGuidelines
DistilledfromSimonPeytonJones,NeurIPS/ICML/ICLRformattingrequirements,
andtechnicalwritingbestpractices\.
\#\#CorePrinciple
Yourpapertellsastory:problemâwhyitmattersâyourapproachâevidence
itworksâwhatitmeans\.Everysectionservesthisnarrative\.
\#\#Startfromproposal\.md
Youalreadywrotearesearchproposal\(proposal\.md\)withintroduction,
approach,relatedwork,andreferences\.\*\*Useitasyourfoundation\*\*:
\-\*\*Introduction\*\*:Adaptfromproposal\.md'sIntroductionsection\.Add
concreteresultsnowthatexperimentsaredone\.
\-\*\*RelatedWork\*\*:Expandfromproposal\.md'sRelatedWorksection\.Use
theBibTeXentriesinreferences/foryourbibliography\.
\-\*\*Method\*\*:Expandfromproposal\.md'sProposedApproachsection\.Add
fulltechnicaldetails,notation,andalgorithmdescriptions\.
\-\*\*References\*\*:Startfromthecitationsinproposal\.mdandreferences/\.
Addanynewpapersdiscoveredduringexperiments\.
DoNOTrewritefromscratchârefineandexpandwhatyoualreadyhave\.
\#\#Structure
Writeinthisorder\(nottheordertheyappearinthepaper\):
1\.MethodsâExperimentsâContributionslistâConclusion
2\.ThenIntroduction\(nowyouknowwhattointroduce\)
3\.ThenRelatedWork
4\.AbstractLAST\(summarizethecompletedpaper\)
Finalpaperorder:
\`\`\`
1\.Title
2\.Abstract\(150\-250words,oneparagraph\)
3\.Introduction\(problem,gap,contributionslist,paperroadmap\)
4\.RelatedWork\(funnel:broadânarrow,endwithyourpositioning\)
5\.Method\(complete,reproducibledescription\)
6\.Experiments\(setup,resultstables,ablations,analysis\)
7\.Discussion/Limitations
8\.Conclusion
9\.References
\`\`\`
\#\#Abstract
\-ONEparagraph,150\-250words
\-Structure:contextâproblemâmethodâkeyresultâimplication
\-Mustbeself\-containedâreadablewithouttherestofthepaper
\-Nocitationsintheabstract
\-Includeoneconcretequantitativeresultifpossible
\#\#Introduction
\-Startwithwhatisknown\(context\)
\-Identifythegap\(what'smissingorbroken\)
\-Stateyourapproach\(onesentence\)
\-Listcontributionsexplicitly:
\`\`\`latex
Ourcontributionsare:
\\begin\{itemize\}
\\itemWeproposeX,whichaddressesY\.
\\itemWeshowthatZthroughexperimentsonAandB\.
\\itemWereleaseourcodeanddataat\[URL\]\.
\\end\{itemize\}
\`\`\`
\-Endwitharoadmap:"Section2reviews\.\.\.,Section3describes\.\.\.,Section4presents\.\.\."
\#\#RelatedWork
\-Organizebyapproach/concept,NOTchronologically
\-Funnelstructure:broadfieldâspecificsubproblemâdirectlycompetingmethods
\-Foreachgroupofrelatedpapers,explain:
1\.Whattheydo
2\.Howyourworkdiffers
\-Endwith:"Unlike\[priorwork\],ourapproach\.\.\."
\-EverycitedpapermustbeREALandverifiable\.SearchSemanticScholar
\(semanticscholar\.org\)toconfirmpapersexistbeforecitingthem\.
\#\#Method
\-Completeenoughthatanexpertcanreimplementfromthepaperalone
\-Stateallassumptionsexplicitly
\-Include:modelarchitecture,lossfunction,trainingalgorithm
\-Useclearnotation,defineeverysymbolonfirstuse
\-Includeamethodoverviewfigureiftheapproachhasmultiplecomponents
\#\#Experiments
\-Structure:SetupâMainresultsâAblationsâAnalysis
\#\#\#Setupsubsection
\-Datasets:name,size,splits,preprocessing
\-Baselines:whattheyare,whychosen,howtrained\(faircomparison\)
\-Metrics:whichones,whyappropriate
\-Implementation:optimizer,lr,epochs,batchsize,hardware,trainingtime
\-Seeds:howmany,whichvalues
\#\#\#Resultssubsection
\-Maincomparisontablewithyourmethodvsallbaselines
\-Boldthebestvalueineachcolumn
\-Includeâorâtoindicateifhigher/lowerisbetter
\-Everynumberinthepapermustmatchexperimentresultsinexp/exactly
\#\#\#Ablationsubsection
\-Onetableshowing:fullmethod,thenremoveeachcomponent
\-Proveseverycomponentcontributes
\#\#\#Analysissubsection\(optionalbutstrengthenspaper\)
\-Failurecases:wheredoesyourmethodfailandwhy?
\-Qualitativeexamples:showwhatthemodelactuallyproduces
\-Trainingcurves:showconvergencebehavior
\#\#Tables
\`\`\`latex
\\begin\{table\}\[t\]
\\caption\{Comparisonon\[Dataset\]\.Bestresultsin\\textbf\{bold\}\.âmeanshigherisbetter\.\}
\\label\{tab:main\}
\\centering
\\begin\{tabular\}\{lccc\}
\\toprule
Method&Accuracyâ&F1â&Latency\(ms\)â\\\\
\\midrule
BaselineA&82\.1±0\.3&79\.4±0\.5&12\.3\\\\
BaselineB&84\.7±0\.2&81\.2±0\.4&15\.7\\\\
\\textbf\{Ours\}&\\textbf\{87\.3±0\.2\}&\\textbf\{84\.1±0\.3\}&14\.1\\\\
\\bottomrule
\\end\{tabular\}
\\end\{table\}
\`\`\`
Rules:
\-Usebooktabs\(\\toprule,\\midrule,\\bottomrule\)ânoverticallines
\-CaptiongoesABOVEthetable
\-Self\-containedcaption:readablewithoutmaintext
\-Referenceeverytableinthetext:"AsshowninTable~\\ref\{tab:main\}\.\.\."
\-Numbersmustmatchexperimentresultsexactly
\#\#Figures
\`\`\`latex
\\begin\{figure\}\[t\]
\\centering
\\includegraphics\[width=0\.8\\linewidth\]\{figures/training\_curve\.pdf\}
\\caption\{Traininglossoverepochs\.Ourmethod\(blue\)convergesfaster
thanBaselineA\(orange\)andBaselineB\(green\)\.\}
\\label\{fig:training\}
\\end\{figure\}
\`\`\`
Rules:
\-CaptiongoesBELOWthefigure
\-Self\-containedcaption
\-UsePDForvectorformatwhenpossible\(notlow\-resPNG\)
\-Readableatprintsize\(fontâ¥8ptinthefigure\)
\-Referenceeveryfigure:"Figure~\\ref\{fig:training\}shows\.\.\."
\-Useconsistentcolorsacrossallfigures
\#\#Discussion/Limitations
\-Discusswhattheresultsmean,notjustwhattheyare
\-Honestlyacknowledgelimitations:
\-"OurmethodassumesX,whichmaynotholdinYscenarios"
\-"WeevaluatedonZdatasets;generalizationtootherdomainsisuntested"
\-Reviewersrewardhonestyâhidinglimitationsgetspapersrejected
\#\#Conclusion
\-Restatetheproblemandyourapproach\(onesentenceeach\)
\-Summarizekeyfindingswithconcretenumbers
\-Statebroaderimplications
\-Suggestfuturework
\-DONOTintroducenewresultsorclaimshere
\-0\.5\-1page
\#\#References\(CRITICAL\)
\-EVERYreferencemustbeaREAL,VERIFIABLEpublication
\-SearchSemanticScholar\(semanticscholar\.org\)tofindandverifypapers
\-Fakeorhallucinatedcitationsunderminescientificintegrity
\-Usecorrectformat:authors,title,venue,year
\-Preferpublishedconference/journalpapersoverarXivpreprints
\-Include15\-30referencesforatypicalMLpaper
\-Use\\citep\{\}forparenthetical:"\(Smithetal\.,2023\)"
\-Use\\citet\{\}fortextual:"Smithetal\.\(2023\)showed\.\.\."
\#\#LaTeXBestPractices
\-Usethevenue'sofficialstylefile\(neurips\_2025\.sty,etc\.\)
\-Usebooktabsfortables\(noverticallines\)
\-Use\\usepackage\{hyperref\}forclickablereferences
\-Definenotationwith\\newcommandforconsistency
\-Use~fornon\-breakingspacesbeforereferences:Table~\\ref\{tab:main\}
\-Compileatleasttwicetoresolvereferences
\-8\-10pagesformaincontent\(excludingreferencesandappendix\)
\#\#CommonMistakesThatGetPapersRejected
\-Noexplicitcontributionslistintheintroduction
\-Claimsnotsupportedbyevidenceintheexperiments
\-Numbersintextdon'tmatchexperimentresults
\-Noablationstudy
\-Unfairbaselinecomparisons
\-Fabricatedreferences
\-Nolimitationsdiscussion
\-Poorwritingquality/unclearmaincontribution
#### Reviewer guideline \(stage 4\)\.
Used by all three reviewer agents \(§[3\.4](https://arxiv.org/html/2605.19156#S3.SS4)\)\. The ML guideline shown below is one of six per\-domain reviewer guidelines; the others share the same nine 0–10 dimensions and ICLR scale, so reviewer scores are cross\-domain comparable\.
\#ReviewerGuidelines
DistilledfromofficialreviewerinstructionsofNeurIPS,ICML,ICLR,ACL,andTMLR\.
\#\#YourRole
Youarereviewingaresearchpaper\.Yourprimaryjobistoevaluatethe
scientificcontributionâthenovelty,soundness,significance,andclarity
ofthework\.Berigorousbutfair\.Bespecific,notvague\.
Youalsohaveaccesstotheexperimentworkspace\(code,logs,results\)for
asanitycheckonresultsintegrity\.
\#\#OverallScore\(ICLRscale:0\-10,evennumbersonly\)
\|Score\|Meaning\|
\|\-\-\-\|\-\-\-\|
\|10\|Top5%ofacceptedpapers,seminalpaper\|
\|8\|Clearaccept,strongcontribution\|
\|6\|Marginal,needsrevision\|
\|4\|Belowthreshold,reject\|
\|2\|Strongrejection,significantflaws\|
\|0\|Trivial,wrong,orfabricated\|
UseONLYthesevalues:0,2,4,6,8,10\.
Acceptancethresholdis8\.Score6triggersarevisionloop\.Score<6isrejected\.
\#\#Per\-DimensionScores\(each1\-10\)
\#\#\#1\.Novelty\(mostimportant\)
\-Doesthepaperpresentgenuinelynewideas,methods,orinsights?
\-\*\*YouMUSTperformatleast5distinctonlinesearches\*\*beforeassessing
novelty\.DoNOTaccepttheauthors'noveltyclaimsatfacevalue\.
Requiredsearchstrategies\(doALLofthem\):
a\)SearchtheexactpapertitleonGoogleScholarandSemanticScholar
b\)Searchthecoretechniquename\+thedomain\(e\.g\.,"adaptivemarginmetriclearning"\)
c\)Searchforeachkeybaseline/relatedworkcitedtofindpapersTHEYcite
d\)Searchforthemethod'skeycomponentscombined\(e\.g\.,"CLIPtextencodermarginloss"\)
e\)Searchrecentproceedings\(last3years\)ofthetargetvenueforsimilarideas
\-Ifyoufindapaperthatproposesasubstantiallysimilarmethod,scorenoveltyâ¤4
\-NovelcombinationsofexistingtechniquescountIFclearlyreasoned
andthecombinationitselfprovidesnewinsight
\-Incrementalimprovementsneedstrongjustificationforwhythe
incrementmatters
\-Lackofstate\-of\-the\-artresultsalonedoesNOTjustifyrejection
\#\#\#2\.Soundness
\-Areclaimswell\-supportedbytheoryorexperiments?
\-Isthemethodologyappropriatefortheproblem?
\-Areproofscorrect?Isexperimentaldesignvalid?
\-Areassumptionsstatedandreasonable?
\-Dotheresultsactuallysupporttheclaimsmade?
\#\#\#3\.Significance
\-Doesthisworkmattertothecommunity?
\-Wouldpractitionersorresearchersbenefitfromknowingthesefindings?
\-Doesitopennewresearchdirectionsorsolvearealproblem?
\-NegativeresultswithhonestanalysisCANbesignificant
\#\#\#4\.Clarity
\-Isthepaperwell\-writtenandorganized?
\-Arecontributionsexplicitlystatedintheintroduction?
\-Arefigures/tablesself\-containedwithdescriptivecaptions?
\-Couldanexpertreproducetheworkfromthepaperalone?
\-Isthenotationconsistentandwell\-defined?
\#\#\#5\.Reproducibility
\-Areallhyperparameters,architectures,andtrainingdetailsspecified?
\-Isthedatadescribed\(splits,sizes,preprocessing\)?
\-Iscomputespecified\(hardware,runtime\)?
\-Areenoughdetailsprovidedforanindependentreimplementation?
\#\#\#6\.ExperimentalRigor
\-Arethereatleast2meaningfulbaselines?
\-Isthereanablationstudyshowingeachcomponent'scontribution?
\-Areerrorbars/confidenceintervalsreported?
\-Areresultsfrommultipleruns\(differentseeds\)?
\-Arestatisticalsignificancetestsusedwhenclaimingsuperiority?
\-Arecomparisonsfair\(samedata,samecomputebudget\)?
\#\#\#7\.References
\-Arekeyrelatedworkscitedandproperlydiscussed?
\-Isthepaperwell\-positionedrelativetopriorwork?
\#\#\#8\.ReferenceIntegrity
\-Areallreferencesreal,verifiablepublications?
\-\*\*SearchSemanticScholarorGoogleScholar\*\*toverifythatcited
papersactuallyexistwiththestatedtitles,authors,andvenues
\-Dothecitedtitles,authors,andvenuesmatchtheactualpublications?
\-Arethereanyhallucinatedorfabricatedcitations?
\#\#\#9\.ResultsIntegrity\(sanitycheckâbutviolationsmeanreject\)
Youhaveaccesstotheexperimentworkspace\(code,logs,results\.json\)\.
YouMUSTverifyALLofthefollowing:
\-Readresults\.jsonandcompareEVERYnumberinthepaper'stablesagainstit
\-Checkthatexperimentsourcecode\(\.pyfiles\)existsintheworkspace\.
IfNOsourcecodeispresent,thisisamajorintegrityconcern\(scoreâ¤4\)
\-Readexperimentlogsandverifytheyshowactualtrainingruns\(epochs,losses,etc\.\)
\-Checkthatthecodeimplementswhatthepaperdescribes\(notadifferentmethod\)
\-Verifyfiguresaregeneratedfromtheactualresults,notfabricated
Theprimaryevaluationisthescientificcontribution\.However,anyof
thefollowingaregroundsfor\*\*automaticrejection\*\*:
\-Referencesthatdon'texist\(fakecitations\)
\-Experimentcodethatcannotrunordoesn'tproducetheclaimedresults
\-Logsthatshowdifferentnumbersthanwhatthepaperreports
\-Numbersinthepaperthatdon'tmatchresults\.json
\-Missingexperimentsourcecodewithnoexplanation
Thesearenotminorissuesâtheyindicatetheresearchisnottrustworthy\.
\#\#DecisionGuidelines
Youroverall\_scoredeterminesthedecision:
\|Score\|Decision\|
\|\-\-\-\|\-\-\-\|
\|10\|accept\|
\|8\|accept\|
\|6\|revision\(marginal,needsrevision\)\|
\|4\|reject\|
\|2\|reject\(strong\)\|
\|0\|reject\(fabricated/trivial\)\|
\#\#ReviewStructure
Yourreviewmustinclude:
1\.\*\*Summary\*\*:2\-3sentenceoverviewofwhatthepaperdoes\(nocritiquehere\)
2\.\*\*Noveltyassessment\*\*:Whatyoufoundwhensearchingforexistingwork
3\.\*\*Strengths\*\*:Specificpositiveswithevidencefromthepaper
4\.\*\*Weaknesses\*\*:Specificissuesâbeconstructiveandactionable
5\.\*\*Detailedfeedback\*\*:Howtoimprovethepaper
6\.\*\*Questions\*\*:Pointsthatcouldchangeyourassessment
7\.\*\*Integritycheck\*\*:Briefnoteonwhetherresultsappeargenuine
\#\#CommonReviewMistakestoAvoid
\-Don'tdismissresultsas"obviousinretrospect"
\-Don'trequireSOTAresultswhenthepaperdoesn'tclaimSOTA
\-Don'trejectforacknowledgedlimitations
\-Don'tdemandexperimentsbeyondthepaper'sstatedscope
\-Don'tusevaguecriticism\("thepaperisunclear"\)âbespecific
\-Don'tletpersonalmethodologypreferencesbiasyourreview
\-Evaluateeachcontributionindependently,notasabundle
\-Don'tconflate"Idon'tfindthisinteresting"with"thisisnotnovel"
## Appendix CPer\-seed paper titles and scores
TablesLABEL:tab:per\_seed\_claude–LABEL:tab:per\_seed\_kimilist all 117 generated papers, split per agent, with the seed, trial index, full title, SAR overall score, and mean PR score \(over 3 reviewers\)\.
Table 5:Per\-seed paper titles and scores forClaude Code\.SeedTrialTitleSARPRAI for Biologyt1When Does Coarse\-to\-Fine Training Help? Ablation Insights from Curriculum Contrastive Learning for Enzyme Function Prediction4\.504\.67t2EpiGNN: Multi\-Mutation Protein Fitness Prediction via Message Passing on Language Model\-Derived Residue Coupling Graphs5\.204\.67t3Supervised Learning on PLM Embeddings for Multi\-Mutant Protein Fitness Prediction: When Do Structural Priors Help?5\.204\.00Causal Learningt1E\-Valued Causal Discovery: Constraint\-Based Structure Learning with Anytime\-Valid FDR Control5\.705\.33t2Know Your Assumptions: Assumption\-Adaptive Edge Orientation for Robust Causal Discovery via Data\-Driven Diagnostics5\.603\.33t3When Do Causal Discovery Algorithms Disagree? Diagnosing Assumption Violations via Per\-Edge Profiling6\.304\.67Compiler Optimizationt1The Algebra of Compiler Passes: An Empirical Study of Idempotency, Commutativity, and Convergence in LLVM Optimization Pipelines5\.605\.33t2ShapleyPass: Game\-Theoretic Attribution and Interaction Analysis of Compiler Optimization Passes5\.605\.33t3ShapleyPass: Quantifying Higher\-Order Interactions Among Compiler Optimization Passes via Shapley Interaction Indices4\.504\.00Computer Visiont1Entropy\-Guided Adaptive Token Merging for Robust and Efficient Vision Transformers5\.806\.00t2Attention Entropy Profiling for Training\-Free Out\-of\-Distribution Detection in Vision Transformers4\.404\.00t3Spectral Token Gating for Vision Transformer Robustness: A Negative Result with Insights on Frequency\-Domain Corruption Detection in Embedding Space5\.805\.33Data Integration & Cleaningt1Structural Sparsity in Constraint Interactions: An Empirical Study of Multi\-Constraint Data Repair Decomposition5\.203\.33t2Error Amplification in Entity Resolution Pipelines: A Formal Analysis of Stage\-Wise Quality Propagation3\.905\.33t3Characterizing Operator Interaction Effects in Data Cleaning Pipelines5\.605\.33Datasets & Benchmarkst1FlipBench: Measuring Directional Reasoning Asymmetry in Large Language Models5\.604\.67t2consistbench\{6\.305\.33t3SkillStack: A Procedurally Generated Benchmark for Measuring Compositional Cognitive Skill Gaps in Large Language Models5\.804\.67Generative Modelst1Spectral Consistency Distillation: Frequency\-Adaptive Teacher Supervision for Few\-Step Flow Matching3\.104\.67t2Prediction Coherence is Not a Quality Signal: A Negative Result on Verifier\-Free Inference\-Time Scaling for Diffusion Models6\.305\.33t3Conditioning\-Space Guidance for Diffusion Transformers: When Does Single\-Pass Classifier\-Free Guidance Work?6\.304\.67Interpretability of Learned Repr\.t1The Convergent Core: Connecting Seed Stability, Cross\-Model Universality, and Causal Importance of Sparse Autoencoder Features5\.604\.00t2The Functional Anatomy of Sparse Features in Language Models5\.805\.33t3Faithful by Consensus: Identifying Causally Important Features Through Multi\-Seed Sparse Autoencoder Agreement5\.804\.00Natural Language Processingt1Context\-Contrastive Uncertainty Decomposition for Reliable Retrieval\-Augmented Generation6\.104\.00t2SpecCheck: Testing Confidence Monotonicity Across Specificity Levels for LLM Hallucination Detection—A Negative Result5\.804\.67t3Know When to Look: Parametric\-Retrieval Agreement as a Calibration Signal for Retrieval\-Augmented Generation5\.805\.33Operating System Designt1MarkovTier: Anticipatory Page Migration via Markov Phase Models for Tiered Memory Systems6\.104\.67t2The Bandwidth Knapsack: Optimal Migration Scheduling for Tiered Memory Systems5\.204\.67t3Invisible Cycles: Characterizing and Quantifying CPU Time Displacement from Asynchronous Kernel Execution in Modern Linux6\.104\.00Privacy in MLt1MemPrune: Investigating Gradient Dispersion as a Memorization\-Aware Neural Network Pruning Criterion4\.903\.33t2Difficulty\-Calibrated Unlearning Auditing: Exposing Per\-Sample Privacy Gaps in Machine Unlearning5\.203\.33t3The Compounding Cost: How Differential Privacy and Model Compression Jointly Amplify Fairness Degradation5\.604\.67Probabilistic Methodst1Confidence Sequences for Markov Chain Monte Carlo: Anytime\-Valid Estimation with Sequential Guarantees5\.205\.33t2Optimal Error Budgeting for Heterogeneous Sketch Pipelines in Approximate Stream Processing6\.105\.33t3Sublinear\-Memory Confidence Sequences for Streaming Quantiles5\.705\.33Supervised Repr\. Learningt1Confusion\-Geometric Supervised Contrastive Learning: Shaping Embedding Geometry from Training Dynamics4\.404\.00t2The Neural Collapse–Calibration Connection is Dataset\-Dependent: An Empirical Investigation via Controlled Within\-Class Geometry5\.804\.00t3Confusion\-Calibrated Supervised Contrastive Learning: Adaptive Class\-Pair Reweighting from Training Dynamics5\.003\.33Table 6:Per\-seed paper titles and scores forCodex\.SeedTrialTitleSARPRAI for Biologyt1Does Systema\-Style Perturbed\-Reference Residualization Help as a Training Target for Unseen Single\-Cell Perturbation Prediction? A Pre\-Registered Benc5\.004\.00t2Masked\-Child Surrogate Calibration for Safe EC Prefix Decisions: A Leakage\-Controlled Benchmark for Future\-Child Emergence5\.004\.67t3SPARE\-Gain: A Low\-Compute Benchmark of Baseline\-Relative Routing for Unseen\-Perturbation Pseudobulk Prediction4\.504\.67Causal Learningt1Benchmarking Subset Aggregation for Classical Causal Discovery Under Marginalization Error5\.004\.67t2PACER\-Cert as a Benchmark for Stopping\-Certificate Calibration in CPU\-Only Active Causal Discovery5\.004\.67t3When Do Path\-Dependent Setup Costs Matter in Sequential Causal Design?3\.204\.67Compiler Optimizationt1DebtAware Beyond LastRunTracking? A Proxy Feasibility Boundary Study of Typed Rerun Suppression in LLVM2\.704\.67t2Typed Skip\-Versus\-Global Scheduling for LLVM Cleanup Pipelines: A Negative Proxy Study5\.004\.00t3Optimization Remarks as a Feasibility Signal for Low\-Budget LLVM Micro\-Search A Pilot Against Random and ProbeDelta3\.604\.00Computer Visiont1Do Corruption\-Family Text Residuals Help Zero\-Shot CLIP? A Controlled Negative Result on CIFAR\-C5\.004\.00t2FOCUS: Object\-Centric Evidence as a Causal Update Gate for Realistic Online Vision\-Language Adaptation5\.804\.67t3Object Units or Pixels? A Negative Proxy Feasibility Study for Online Segmentation Adaptation3\.904\.67Data Integration & Cleaningt1A Preliminary Artifact\-Backed Cautionary Study of Benchmark\-Conditional Admissibility for Robustness Evaluation in Schema and Entity Matching3\.904\.67t2CanopyER: A Short Systems Note on Budgeted Rewrite\-vs\-Match Scheduling for Progressive Entity Matching3\.804\.67t3StressAudit\-SM: A Compact Robustness Audit for Schema Matching Under Metadata Stress5\.004\.00Datasets & Benchmarkst1DriftAnswer\-Py: A 12\-Item Executable Pilot of Accepted Python Stack Overflow Answers Under Documented Library Drift6\.304\.67t2RevisionBench: A Reproducible Failed\-Construction Case Study for Abstract\-Local Scientific Claim Updates6\.204\.00t3TwinBench: A Synthetic Procedural\-Core Pilot for Coupled Invariance and Boundary Sensitivity in QA6\.304\.67Generative Modelst1When Assignment Matters: A Pilot Study of DAAM\-Assisted Compositional Text\-to\-Image Reranking4\.404\.00t2Decomposed Early Ranking Targets Under a Local Surrogate Evaluator for Compositional Text\-to\-Image Generation4\.504\.00t3bf ParaDG: An Exploratory Negative Study of Disagreement\-Gated Paraphrase Blending for Text\-to\-Image Diffusion5\.804\.67Interpretability of Learned Repr\.t1Do Shared Decoders Improve Prototype\-Edit Reusability on Frozen CLIP Features?5\.004\.00t2Pair\-Supervised Regularization for Selective Counterfactual Edits in Frozen Vision SAEs5\.004\.67t3Benchmarking Weakly Supervised Factor\-Localized Sparse Autoencoders on Frozen Vision Features5\.804\.67Natural Language Processingt1When Does Clarification Supervision Transfer to Formal Reasoning? A Controlled Pilot of Validator\-Clean vs\. Matched Noisy Missing\-Fact Tuning4\.904\.00t2LIMS\-RAG: Localized Minimal\-Support Perturbations Are a Weak but Measurable Feature Family for Sentence\-Level RAG Verification5\.004\.67t3LateBind: Controlled Additive\-Value Tests for Timing\-Aware Shortcut Mitigation in Text Classification6\.304\.67Operating System Designt1Replay\-Scoped Evidence for Bias\-Corrected Counterfactual Policy Ranking in One Shared Linux Page Cache5\.804\.67t2ShareArb: Evictor\-Side Responsibility for Shared Linux Page\-Cache Arbitration5\.804\.67t3ShadowCache: A Simulator\-Backed Trace Study of How Much Observable State Policy Ranking Needs5\.004\.67Privacy in MLt1Who Was in the Recent Window? A Rigorous Audit of Online Test\-Time Adaptation Privacy5\.205\.33t2Matched\-Budget Evaluation of Weak\-View Residuals for One\-Run Differential Privacy Auditing4\.704\.00t3Are Early Artifact Forecasts Actionable for Membership Privacy? A Budget\-Matched Study of Selective Intervention3\.904\.67Probabilistic Methodst1CoSBC: Dependence\-Specialized Enriched SBC with Symmetric Pooled Ranking4\.504\.67t2Hierarchical Diagonal\-GMM Posteriors for Localized Conformal Prediction: A Scoped Negative Result5\.004\.67t3How Far Does Probe\-Only Recalibration Transfer in Bayesian Quadrature?5\.004\.67Supervised Repr\. Learningt1STRIDE: Reliability\-Gated Class\-Relation Smoothing for Self\-Supervised Transfer4\.504\.67t2Adaptive Prototype Granularity in Frozen\-Feature Contrastive Adaptation5\.004\.67t3How Much Signal Is in Early Training Trajectories? A Matched\-Budget Study of Pseudo\-Group Inference5\.805\.33Table 7:Per\-seed paper titles and scores forKimi Code\.SeedTrialTitleSARPRAI for Biologyt1contextstab: Context\-Aware Protein Stability Prediction Using Real Single\-Cell Transcriptomics4\.404\.00t2Tri\-Con: Tri\-Hierarchy Contrastive Learning for Cell Ontology\-Guided Single\-Cell Analysis4\.403\.33t3CellStratCP: Cell\-Type\-Stratified Adaptive Conformal Prediction for Calibrated Uncertainty in Single\-Cell RNA\-seq Imputation5\.205\.33Causal Learningt1AIT\-LCD: Adaptive Information\-Theoretic Local Causal Discovery with Explicit Conditioning Set Awareness4\.102\.67t2SPICED: Structural Prior Integration for Constrained Estimation of Directed Information5\.204\.00t3On the Challenges of Multi\-Fidelity Conditional Independence Testing for Causal Discovery: An Empirical Study2\.004\.00Compiler Optimizationt1From Branches to Bytes: A Negative Result in Extending Learned Static Prediction to Data Layout Optimization4\.404\.00t2LEOPARD: Lightweight Learned Guidance for Equality Saturation in Compiler Optimization4\.402\.00t3Joint Compute and Layout Optimization via Hierarchical E\-Graphs4\.402\.67Computer Visiont1DU\-VPT: Decomposed Uncertainty\-Guided Visual Prompt Tuning for Test\-Time Adaptation4\.402\.67t2Adaptive Prototype\-Aware Consistency with Learnable Augmentation Policies for Single\-Image Test\-Time Adaptation3\.603\.33t3CASS\-ViM: Content\-Adaptive Selective Scanning for Vision State Space Models4\.404\.00Data Integration & Cleaningt1CESF: A Controllable Error Synthesis Framework for Reproducible Data Cleaning Evaluation3\.903\.33t2Towards LLM\-as\-Compiler for Data Cleaning: A Feasibility Study4\.104\.00t3CleanBP: Making Belief Propagation Practical for Holistic Data Repair via FD\-Specific Sparsification4\.104\.00Datasets & Benchmarkst1textbf\{CompViz: A Dynamic Benchmark for Compositional Visual Reasoning with Sub\-100ms Generation5\.203\.33t2IntrospectBench: A Cross\-Domain Benchmark for Evaluating Step\-Level Reasoning Introspection in Large Language Models5\.202\.67t3DynaScale: Dynamic Difficulty Scaling for Maintaining Discriminative Power in AI Benchmarks5\.204\.00Generative Modelst1Flow\-Guided Token Routing: Adaptive Computation Allocation for Efficient Flow Matching3\.702\.00t2Distance\-Aware Flow Matching for LiDAR Point Cloud Generation4\.903\.33t3VAST: Velocity\-Adaptive Spatially\-varying Timesteps for Training\-Free Acceleration of Diffusion Models4\.102\.00Interpretability of Learned Repr\.t1CAGER: Causal Geometric Explanation Recovery A Framework for Grounding Interpretability in Causal Subspace Geometry4\.404\.00t2Intervention Fidelity Scoring: Characterizing the Precision\-Magnitude Trade\-off in SAE\-Based Steering3\.703\.33t3PhaseMine: Detecting Feature Emergence Phase Transitions via Dynamic Sparse Probing4\.404\.00Natural Language Processingt1SAE\-GUIDE: Sparse Autoencoder\-Guided Uncertainty\-aware Information Detection and Enhancement for Multi\-Hop Retrieval4\.102\.67t2Confidence\-Dynamic Heterogeneous Reasoning: protect A Study of Challenges in Adaptive Strategy Selection3\.302\.67t3Entropy\-Guided Stepwise Revision: In\-Chain Self\-Correction for Efficient Reasoning3\.603\.33Operating System Designt1UniSched: A Critical Analysis of Simulation\-Based Evaluation for CXL\-Aware CPU Scheduling4\.403\.33t2WattSched: Adaptive Workload\-Aware Energy Scheduling for Heterogeneous Multi\-Core Systems using sched\_ext4\.404\.00t3KAPHE: Kernel\-Aware Performance Heuristic Extraction vspace\{0\.2em3\.604\.67Privacy in MLt1Post\-Hoc Compression\-Aware Differential Privacy: Optimizing DP Training for Deployed Compressed Models5\.300\.00t2textbf\{G3P: Gradient\-Guided Privacy\-Preserving Pruning via Train\-Test Gradient Saliency4\.404\.00t3On the Limitations of Gradient\-Based Verification for Machine Unlearning5\.204\.00Probabilistic Methodst1Decaying HyperLogLog: Continuous\-Time Cardinality Estimation with Exponential Aging2\.404\.00t2Streaming Multi\-Scale Adaptive Kernel Conformal Prediction5\.204\.00t3Comparative Analysis of Adaptation Criteria for Gradient\-Based Discrete MCMC: When Acceptance\-Rate Trumps Jump\-Distance5\.204\.67Supervised Repr\. Learningt1Feature\-Diversity\-Aware Supervised Contrastive Learning: Mitigating Feature Suppression through Adaptive Pair Weighting5\.202\.00t2Gradient\-Confusion Aware Supervised Contrastive Learning2\.804\.00t3ETF\-SCL: Equiangular Tight Frame Guided Supervised Contrastive Learning for Long\-Tail Recognition2\.502\.67
## Appendix DPer\-domain SAR and PR breakdown
Table[8](https://arxiv.org/html/2605.19156#A4.T8)reports SAR and PR mean scores for each of the 13 research domains, averaged across the three agents \(n=9n=9papers per domain\)\. Domains are ordered by SAR within each platform\.
Table 8:Per\-domain SAR and PR mean scores, averaged across the three agents\.PlatformDomainSARPRSAR−\-PRCPUOperating System Design5\.164\.370\.79Probabilistic Methods4\.924\.740\.18Causal Learning4\.684\.220\.46Compiler Optimization4\.474\.000\.47Data Integration & Cleaning4\.394\.300\.09GPUDatasets & Benchmarks5\.794\.221\.57Interpretability of Reps5\.064\.220\.84NLP4\.994\.000\.99Privacy in ML4\.933\.701\.23AI for Biology4\.824\.370\.45Computer Vision4\.794\.300\.49Generative Models4\.793\.850\.94Supervised Repr\. Learning4\.563\.850\.71Figure[8](https://arxiv.org/html/2605.19156#S5.F8)\(in §[5\.3](https://arxiv.org/html/2605.19156#S5.SS3)\) visualises the per\-domain PR scores; the parallel SAR view is Figure[9](https://arxiv.org/html/2605.19156#A4.F9)below\. Table[9](https://arxiv.org/html/2605.19156#A4.T9)reports the same data collapsed to the CPU/GPU split per agent\. Figure[10](https://arxiv.org/html/2605.19156#A4.F10)gives the per\-\(seed, trial\) score grids for both SAR and PR alongside each agent\.
Table 9:SAR vs\. PR by agent and compute platform\.Figure 9:Per\-domain mean SAR scores by agent \(companion to Figure[8](https://arxiv.org/html/2605.19156#S5.F8)\)\.Figure 10:Per\-\(seed, trial\) SAR \(top\) and PR \(bottom\) score heatmaps per agent\. CPU seeds above the divider, GPU below; columns are trial indices \(t1t\_\{1\},t2t\_\{2\},t3t\_\{3\}\)\.
## Appendix ETime analysis
#### Per\-stage time\.
Figure[11](https://arxiv.org/html/2605.19156#A5.F11)reports per\-agent wall\-clock time per pipeline stage in two views: grouped means per stage \(left\) and a stacked breakdown of total minutes per average run \(right\)\. Experiments dominate \(67–83% of total\) and self\-refinement is only 3–8%\. The same totals \(Claude Code 13\.0 h, Codex 6\.8 h, Kimi Code 4\.1 h\) appear as the right panel of Figure[4](https://arxiv.org/html/2605.19156#S4.F4)in §[4\.3](https://arxiv.org/html/2605.19156#S4.SS3)\.
Figure 11:Mean wall\-clock per pipeline stage \(left, in minutes\) and total time per average run with stages stacked \(right, in minutes\), for each agent\.
#### Per\-paper distribution\.
Figure[12](https://arxiv.org/html/2605.19156#A5.F12)shows the per\-paper total wall\-clock distribution from the releasedtracker\.jsonlogs\. Claude Code’s distribution has the longest tail \(\>\>40 h on some runs, driven by experiment\-execution self\-refinement loops\); Codex and Kimi Code are tighter and shorter\.
Figure 12:Per\-paper total wall\-clock distribution per agent, fromtracker\.jsonlogs\.
## Appendix FH100 scaling experiment \(per\-domain breakdown\)
We re\-ran all 8 GPU seeds with Codex on 8×\\timesNVIDIA H100 \(80 GB\) for 3 trials each, with budget matched to the A6000 runs \(§[5\.4](https://arxiv.org/html/2605.19156#S5.SS4)\)\. Table[10](https://arxiv.org/html/2605.19156#A6.T10)reports per\-domain Codex PR scores on H100 vs\. A6000\. The aggregate change is small \(−0\.24\-0\.24on average\) and agent\-level SAR also drops, indicating that compute is not the binding constraint\.
Table 10:Per\-domain Codex PR mean scores on 8×\\timesNVIDIA H100 vs\. 1×\\timesA6000, budget\-matched\.
## Appendix GReviewer analysis
#### Self\-refinement effectiveness\.
At each of the three authoring stages \(ideation, experiment, paper\), the agent self\-reviews and revises if the score falls below threshold \(up to 3 rounds\)\. Table[11](https://arxiv.org/html/2605.19156#A7.T11)reports, for each \(agent, stage\), the share of revision rounds in which the score*improved*, stayed the*same*, or*dened*, plus the mean delta\. Self\-refinement is effective for ideation and experiment\-execution \(avg\.\+2\.0\+2\.0to\+3\.0\+3\.0for Claude Code / Kimi Code at those gates\), but limited for paper writing, where revising tends to leave scores unchanged or lower them\.
Table 11:Self\-refinement effectiveness per agent and gate \(across all revision rounds\)\.
#### Peer\-review bias\.
Each paper is scored by all three agents acting as reviewers\. Reviewer severity differs sharply \(Table[12](https://arxiv.org/html/2605.19156#A7.T12), 2\.8\-point mean spread between strictest and most lenient\), and the reviewer\-by\-reviewee matrix in Table[13](https://arxiv.org/html/2605.19156#A7.T13)and Figure[13](https://arxiv.org/html/2605.19156#A7.F13)shows that the bias is mostly an across\-the\-board reviewer effect rather than an agent\-specific self\-favouring effect: Codex gives every reviewee its lowest scores, Kimi Code gives every reviewee its highest scores\. Notably, Kimi Code does*not*score its own papers highest \(5\.0\), but is still substantially more lenient overall, including on agents that produce stronger papers\. Single\-reviewer self\-evaluation therefore drifts toward the reviewer’s own bias, motivating the triple\-reviewer protocol used throughout the paper\.
Figure 13:Score distribution by \(reviewer, reviewee\) across all 117 papers\. Codex is the strictest reviewer; Kimi Code is the most lenient\. Reviewer effects dominate over self\-favouring\.Table 12:Per\-reviewer severity: distribution of overall scores each agent assigns across all 117 papers it reviews\.Table 13:Mean PR score by \(reviewer, reviewee\)\. Rows are reviewers; columns are paper authors\.
## Appendix HManually annotated SAR final\-decision acceptance rates
SAR’s continuous 0–10 score does not map directly to accept/reject\. We manually inspected every SAR review across the 117 agent\-generated papers, the 200 ICLR 2025 papers, and the 102 FARS papers, and assigned a binary decision from the verbal recommendation \(treating “borderline accept”, “accept with revision”, and “conditional accept” as accepts\)\. The ICLR\-Weighted row in Table[14](https://arxiv.org/html/2605.19156#A8.T14)mixes accepted and rejected rates in proportion to ICLR’s∼\\sim32% acceptance rate\.
Table 14:Manually annotated SAR final\-decision acceptance rates per system\.
## Appendix ICase studies
We illustrate the failure modes documented in §[5](https://arxiv.org/html/2605.19156#S5)with six representative papers\.
### I\.1Case 1: When Do Causal Discovery Algorithms Disagree? Diagnosing Assumption Violations via Per\-Edge Profiling
Claude Code can occasionally produce a paper with a real insight, but weak evidence still prevents it from being convincing\.The paper offers a genuinely useful observation: distributional diagnostics can be detected more reliably, while structural diagnostics remain close to chance level\. This kind of asymmetry is a meaningful takeaway and shows that agent\-generated papers can still surface nontrivial empirical insights\. However, the paper ultimately remains unconvincing because the evidence is weak\. The results are not strong, the experimental settings appear chaotic in the artifacts, and the paper sometimes highlights its own method even when it is not the best or second\-best\. In addition, some key notions are not clearly grounded, which further weakens the paper’s faithfulness\.
Figure 14:Page\-level thumbnail of the Case 1 paper\.
### I\.2Case 2: The Algebra of Compiler Passes: An Empirical Study of Idempotency, Commutativity, and Convergence in LLVM Optimization Pipelines
Claude Code may overclaim or present unsupported results when experiments are weak\.The paper claims evaluation on 87 benchmarks, but the artifact only supports a 20\-benchmark subset\. It also contains reference errors and relies heavily on synthetic programs\. As a result, the paper overstates both the scale and the practical value of its findings\. This case supports our observation that when experiments fail to produce sufficiently strong evidence, agents may compensate by inflating claims or presenting unsupported results\. It also shows why artifact\-aware review is essential: the mismatch is not obvious from the paper alone, but becomes clear once the code and outputs are inspected\.
Figure 15:Page\-level thumbnail of the Case 2 paper\.
### I\.3Case 3: Do Corruption\-Family Text Residuals Help Zero\-Shot P? A Controlled Baseline Study
Codex reduces fabrication partly by running much narrower experiments\.Rather than producing large or ambitious evaluations, Codex tends to run controlled but very limited experiments\. Here the study uses only a single frozen P backbone on CIFAR\-10, making the empirical scope too narrow to support broad conclusions\. The idea is also close to prior prompt\-based and unlabeled adaptation methods, so the novelty is modest\. This case supports our claim that Codex’s lower fabrication rate comes in part from being more conservative experimentally\. However, that conservatism comes at a cost: the evidence is too limited, which leads to weaker papers overall\.
Figure 16:Page\-level thumbnail of the Case 3 paper\.
### I\.4Case 4: DU\-VPT: Decomposed Uncertainty\-Guided Visual Prompt Tuning for Test\-Time Adaptation
Kimi Code fabricates experimental results directly rather than actually running the experiments\.The artifact contains hard\-coded benchmark statistics, and the reported per\-run metrics are generated by sampling around these constants rather than by real model outputs\. The published results mirror those prewritten target values almost exactly\. Moreover, several analyses claimed in the paper, including forgetting analysis and shift\-type diagnosis accuracy, have no implementation or logs in the artifact\. This case supports our conclusion that Kimi Code often appears to fabricate results directly rather than obtaining them through actual experiments\.
Figure 17:Page\-level thumbnail of the Case 4 paper\.
### I\.5Case 5: UniSched: A Critical Analysis of Simulation\-Based Evaluation for CXL\-Aware CPU Scheduling
Even when Kimi Code produces code, the method and implementation often do not match\.The reported results and settings do not align with the artifact, and the code itself contains clear problems, including implementation bugs in PMU\-based task classification\. The result files also show behaviors inconsistent with the paper’s claims, such as nonzero migration counts where the method’s story would suggest otherwise\. Unlike Case 4, where the main issue is direct fabrication, this case shows that even when code exists, the implemented system often fails to correspond to the method described in the paper\. It therefore supports our claim that Kimi Code’s failures are not limited to fake numbers, but also include deeper mismatches between method, code, and evaluation\.
Figure 18:Page\-level thumbnail of the Case 5 paper\.
### I\.6Case 6: Characterizing Operator Interaction Effects in Data Cleaning Pipelines
A common failure is missing relevant baselines even when the paper substantially overlaps with prior work\.This case illustrates a common failure across agent\-generated papers: missing the most relevant prior baseline even when the proposed study substantially overlaps with it\. Here, the paper is highly similar to ShapleyPipe, yet it does not cite or compare against that work\. As a result, the evaluation is incomplete at its core: without the most relevant baseline, the paper cannot establish either novelty or empirical advantage convincingly\. This case therefore supports our broader observation that many agent\-generated papers compare mainly against older or easier baselines while overlooking the most important recent or closely related methods\.
Figure 19:Page\-level thumbnail of the Case 6 paper\.Similar Articles
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
This paper introduces AutoResearchEval, an evaluation framework for AI agents in automated scientific research, revealing a critical lack of metacognitive abilities as a recurring failure pattern across models.
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
This survey examines the emerging field of AI-powered research automation (AutoResearch), analyzing how AI systems are moving from isolated task assistance to full workflow-level scientific discovery. It defines a spectrum from human-steered 'Vibe Research' to AI-led systems, and proposes five evaluation dimensions for scientific credibility.
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
A survey paper examining the transition of AI from task-specific assistants to workflow-level research automators, defining AutoResearch as the spectrum of AI-powered scientific workflow automation and analyzing challenges in autonomy, reproducibility, and accountability.
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
This paper introduces ARAC-Bench, a benchmark for evaluating the alignment, logical coherence, and completeness of Auto-Research systems' research processes against human methodology. Experiments on 11 state-of-the-art frameworks show a best alignment score of only 67.9/100, highlighting a significant gap in simulating rigorous human research behavior.
@rohanpaul_ai: New Meta, Stanford, Google and many other top labs paper proposes AutoResearchClaw. Shows that automated research impro…
A new paper from Meta, Stanford, and Google introduces AutoResearchClaw, which improves automated research by integrating failure recovery, debate, and selective human input. It outperforms AI Scientist v2 by 54.7% on ARC-Bench and reveals that autonomy is enhanced when constrained by process rather than given unlimited freedom.