Personalized Auto-Research: Towards a True AI Co-Scientist
Summary
The paper introduces a framework for personalized auto-research systems that condition every stage of the research process on individual scientist representations, arguing that personalization is essential for AI to serve as true co-scientists rather than generic instruments.
View Cached Full Text
Cached at: 08/18/26, 10:12 AM
# Towards a True AI Co-Scientist Source: [https://arxiv.org/html/2608.14881](https://arxiv.org/html/2608.14881) ## Personalized Auto\-Research: Towards a True AI Co\-ScientistConference:ACM AI Leadership Summit; 2026;ACM AI Leadership Summit 2026CCS:Computing methodologies Artificial intelligenceCCS:Computing methodologies Machine learningCCS:Information systems Data mining Bo Ni,Franck Dernoncourtemail:[dernonco@adobe\.com](mailto:[email protected])Affiliation:Adobe Research,San Jose,California,USA,Hongjie Chenemail:[hojiechen@gmail\.com](mailto:[email protected])Affiliation:Dolby Laboratories,San Francisco,California,USA,Yu Wangemail:[yu\.wang6@uga\.edu](mailto:[email protected])Affiliation:University of Georgia,Athens,Georgia,USA,Nesreen K\. Ahmedemail:[n\.kamel@gmail\.com](mailto:[email protected])Affiliation:Cisco AI Research,San Jose,California,USA,Zhengzhong Tuemail:[tzz@tamu\.edu](mailto:[email protected])Affiliation:Texas A&M University,College Station,Texas,USA,Tyler Derremail:[tyler\.derr@vanderbilt\.edu](mailto:[email protected])Affiliation:Vanderbilt University,Nashville,Tennessee,USAandRyan A\. Rossiemail:[ryarossi@gmail\.com](mailto:[email protected])Affiliation:Adobe Research,San Jose,California,USA 2026© , 2026; ###### Abstract\. AI co\-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out\. Despite this rapid progress, state\-of\-the\-art systems remain*researcher\-agnostic*: given a research goal, they optimize novelty, validity, or reviewer score while ignoring the individual scientist who will use the output\. This overlooks a fundamental fact about research, namely, that what counts as novel, valuable, or feasible depends on the researcher, including their prior work, methodological repertoire, and the collaborators and communities in which they are embedded\. In this work, we introduce the problem of*personalized auto\-research*, which conditions every stage of the research process on a representation of the individual researcher\. We argue that personalization is not a convenience layer, but rather the fundamental property that allows an AI system to serve as a genuine co\-scientist rather than a generic instrument\. To address this problem, we propose a general and flexible framework that threads a graph\-grounded researcher context through retrieval, hypothesis search, experimentation, writing, and review\. The framework consists of three fundamental components: \(i\) graph\-grounded researcher representations, \(ii\) personalization across the full research pipeline, and \(iii\) evaluation grounded in the individual\. Notably, we highlight a one\-size\-fits\-all failure mode where distinct researchers issuing the same goal receive essentially the same research, erasing the tacit knowledge through which novel ideas arise\. Finally, we discuss fundamental open problems and challenges\. ###### Keywords: Auto\-research, AI co\-scientists, personalization goalggresearcher\-agnosticco\-scientistfi\(g,o<i\)f\_\{i\}\(g,o\_\{<i\}\)𝒬\\mathcal\{Q\}same for*every*useruu\(a\)Researcher\-agnosticu1u\_\{1\}fi\(g,o<i∣cu1\)f\_\{i\}\(g,o\_\{<i\}\\mid c\_\{u\_\{1\}\}\)𝒬u1\\mathcal\{Q\}\_\{u\_\{1\}\}𝐳u1→cu1\\mathbf\{z\}\_\{u\_\{1\}\}\\\!\\rightarrow\\\!c\_\{u\_\{1\}\}\(b\)Personalizedu2u\_\{2\}fi\(g,o<i∣cu2\)f\_\{i\}\(g,o\_\{<i\}\\mid c\_\{u\_\{2\}\}\)𝒬u2\\mathcal\{Q\}\_\{u\_\{2\}\}𝐳u2→cu2\\mathbf\{z\}\_\{u\_\{2\}\}\\\!\\rightarrow\\\!c\_\{u\_\{2\}\}same goalgg≠\\neqdifferent for*each*useruuFigure 1\.\(a\)A researcher\-agnostic co\-scientist maps the same abstract goalggto an identical research package𝒬\\mathcal\{Q\}for every researcher\.\(b\)Personalized auto\-research conditions every stage on a graph\-derived contextcuc\_\{u\}\. For example, researcheru1u\_\{1\}lies in a dense network cluster, whileu2u\_\{2\}bridges a structural hole\. Consequently, the same goal naturally yields distinct evidence, search trajectories, and output packages \(𝒬u1≠𝒬u2\\mathcal\{Q\}\_\{u\_\{1\}\}\\neq\\mathcal\{Q\}\_\{u\_\{2\}\}\)\. Each package is thereby optimized for feasibility, alignment, and novelty relative to the individual researcher’s domain\.## 1\.Introduction Scientific publication has grown exponentially for decades\([27](https://arxiv.org/html/2608.14881#bib.bib13)\), and no individual researcher can fully navigate the literature of their own field, let alone integrate insights from adjacent disciplines\. This fundamental problem motivates the recent line of work on AI co\-scientists and auto\-research systems, that is, language\-model agents that generate hypotheses, retrieve and synthesize related work, design and execute experiments, and draft manuscripts\([12](https://arxiv.org/html/2608.14881#bib.bib3);[4](https://arxiv.org/html/2608.14881#bib.bib1);[33](https://arxiv.org/html/2608.14881#bib.bib16);[23](https://arxiv.org/html/2608.14881#bib.bib4);[25](https://arxiv.org/html/2608.14881#bib.bib19)\)\. Since the 2024 AI Scientist prototype, the field has moved rapidly to progressive agentic tree search, experiment\-manager control, shared preprint\-style archives, cross\-run memory, human\-in\-the\-loop co\-research, and explicit provenance and safety mechanisms\([22](https://arxiv.org/html/2608.14881#bib.bib20);[36](https://arxiv.org/html/2608.14881#bib.bib21);[26](https://arxiv.org/html/2608.14881#bib.bib22);[11](https://arxiv.org/html/2608.14881#bib.bib23);[28](https://arxiv.org/html/2608.14881#bib.bib24);[6](https://arxiv.org/html/2608.14881#bib.bib25)\)\. Despite their importance and rapid progress, these systems share a fundamental structural limitation: they are*researcher\-agnostic*\. Given the same goal, they produce the same distribution of outputs whether the requester is a first\-year doctoral student or a senior professor, a graph\-mining researcher or a computational biologist \(Figure[1](https://arxiv.org/html/2608.14881#S0.F1)\)\. The output may be competent, but it is also interchangeable, and this interchangeability contradicts how scientific ideas are often discovered\. Intuitively, novel directions frequently arise from a scientist’s idiosyncratic combination of prior failures, methodological taste, and hard\-won experimental intuition, and a homogeneous system erases precisely the heterogeneity that makes research creative\. The stakes span every level: \(i\)*epistemic*, since most scientific capability is tacit and absent from the literature, and is thus unreachable by any literature\-conditioned system; \(ii\)*systemic*, since many researchers querying one generic engine drives the field toward a scientific monoculture, racing the same ideas while counterfactually valuable directions go unexplored; and \(iii\)*practical*, since researchers can only verify, and will only adopt, directions matched to their expertise and resources\. In this work, we introduce*personalized auto\-research*, the problem of conditioning every stage of the research process on a representation of the individual researcher\. Notably, this problem is fundamentally different from both existing AI co\-scientist systems and prior work on personalized language models\([21](https://arxiv.org/html/2608.14881#bib.bib12);[18](https://arxiv.org/html/2608.14881#bib.bib10)\): existing co\-scientists automate research but largely ignore the researcher, whereas existing personalization work models user\-specific outputs but not the sequence of scientific decisions that constitute research\. In personalized auto\-research, the object being personalized is an end\-to\-end research trajectory, not merely a response\. Summary of contributions\.The key contributions of this work are as follows:\(1\)We formalize the problem of personalized auto\-research \(§[3](https://arxiv.org/html/2608.14881#S3)\)\.\(2\)We propose a general and flexible framework \(Algorithm[1](https://arxiv.org/html/2608.14881#alg1)\) for this problem that personalizes every step such as retrieval, hypothesis search, experimentation, writing, citation, and review, etc \(§[4](https://arxiv.org/html/2608.14881#S4)\)\.\(3\)We present a vision for personalized auto\-research organized around three fundamental components: researcher representation, personalization across the full research pipeline, and evaluation grounded in the individual \(§[5](https://arxiv.org/html/2608.14881#S5)\)\.\(4\)Finally, we discuss open problems and challenges\. \(§[6](https://arxiv.org/html/2608.14881#S6)\)\. ## 2\.Background AI Co\-Scientists and Auto\-Research\.Long\-horizon reasoning and tool use\([29](https://arxiv.org/html/2608.14881#bib.bib7);[35](https://arxiv.org/html/2608.14881#bib.bib8)\)have enabled agents that carry out research end\-to\-end\. The AI Scientist established the autonomous loop of proposing, implementing, writing up, and reviewing experiments\([12](https://arxiv.org/html/2608.14881#bib.bib3);[13](https://arxiv.org/html/2608.14881#bib.bib17)\), and AI Scientist\-v2 replaced its template\-dependent loop with progressive agentic tree search, an experiment manager, parallel execution, and vision\-language figure feedback\([33](https://arxiv.org/html/2608.14881#bib.bib16)\)\. A growing set of systems extends this across multi\-agent hypothesis generation, staged human\-feedback workflows, automated data science, and long\-horizon discovery\([4](https://arxiv.org/html/2608.14881#bib.bib1);[5](https://arxiv.org/html/2608.14881#bib.bib18);[23](https://arxiv.org/html/2608.14881#bib.bib4);[25](https://arxiv.org/html/2608.14881#bib.bib19);[34](https://arxiv.org/html/2608.14881#bib.bib5);[31](https://arxiv.org/html/2608.14881#bib.bib26);[17](https://arxiv.org/html/2608.14881#bib.bib27);[19](https://arxiv.org/html/2608.14881#bib.bib28)\)\. A parallel line builds the surrounding infrastructure, namely shared archives, cross\-run memory, and population\-level search over code states\([22](https://arxiv.org/html/2608.14881#bib.bib20);[36](https://arxiv.org/html/2608.14881#bib.bib21);[28](https://arxiv.org/html/2608.14881#bib.bib24);[11](https://arxiv.org/html/2608.14881#bib.bib23);[6](https://arxiv.org/html/2608.14881#bib.bib25);[26](https://arxiv.org/html/2608.14881#bib.bib22)\), and a third studies evaluation and risk through discovery benchmarks, manuscript verification, safety, and critiques of implementation and evaluation bottlenecks\([2](https://arxiv.org/html/2608.14881#bib.bib29);[14](https://arxiv.org/html/2608.14881#bib.bib30);[24](https://arxiv.org/html/2608.14881#bib.bib31);[39](https://arxiv.org/html/2608.14881#bib.bib32);[40](https://arxiv.org/html/2608.14881#bib.bib33);[15](https://arxiv.org/html/2608.14881#bib.bib34)\); recent surveys organize the space by research stage and autonomy level\([16](https://arxiv.org/html/2608.14881#bib.bib35);[3](https://arxiv.org/html/2608.14881#bib.bib36);[37](https://arxiv.org/html/2608.14881#bib.bib37);[20](https://arxiv.org/html/2608.14881#bib.bib38);[30](https://arxiv.org/html/2608.14881#bib.bib39);[38](https://arxiv.org/html/2608.14881#bib.bib40)\)\. Auto\-research is thus no longer speculative, spanning agentic search, automated implementation, full\-paper generation, and safety controls, with domain systems such as AlphaFold showing how AI complements expert judgment in high\-stakes science\([7](https://arxiv.org/html/2608.14881#bib.bib6)\)\. Yet the dominant objective is task\-, benchmark\-, or field\-conditioned: quality is optimized with respect to the goal, literature, evidence, or reviewer model, never with respect to the individual researcher who will adopt the output\. Even scientist\-in\-the\-loop systems\([4](https://arxiv.org/html/2608.14881#bib.bib1);[23](https://arxiv.org/html/2608.14881#bib.bib4)\)condition only on manual, session\-level guidance, so interaction is not personalization, which requires a learned, persistent representation of what the scientist cannot articulate in a prompt\. Auto\-research and co\-scientist are therefore not synonyms: the former names a capability \(the automation of research\) while the latter names a relationship with a specific researcher\. Table[1](https://arxiv.org/html/2608.14881#S2.T1)makes this explicit: personalization is orthogonal to the autonomy axis along which the field has advanced, and the bottom\-right quadrant, where interaction updates the personalization itself, is the true AI co\-scientist we pose and the gap this work addresses\. Table 1\.Personalization is orthogonal to autonomy\. Existing systems advance along the autonomy axis while remaining researcher\-agnostic, whereas our vision ofPersonalized Auto\-Research\(Alg\.[1](https://arxiv.org/html/2608.14881#alg1)\) and ourTrue Personalized AI Co\-Scientist\(Alg\.[1](https://arxiv.org/html/2608.14881#alg1)with human\-in\-the\-loop\)\.Personalization\.Recommender systems learn latent user representations from interactions\([8](https://arxiv.org/html/2608.14881#bib.bib9)\), while preference alignment\([18](https://arxiv.org/html/2608.14881#bib.bib10)\), retrieval\-augmented generation\([10](https://arxiv.org/html/2608.14881#bib.bib11)\), and the LongLaMP benchmark\([9](https://arxiv.org/html/2608.14881#bib.bib2)\)establish that user\-conditioned language modeling is tractable and beneficial\. However, personalizing a recommendation or a single output is far narrower than personalizing science\. In personalized auto\-research, the object is a sequence of decisions: what to retrieve, what to hypothesize, which experiments to run, how to frame the write\-up, how to revise the resulting artifact, etc\. ## 3\.Problem Formulation More formally, let𝒰\\mathcal\{U\}denote the population of researchers and let𝒢=\(𝒱,ℰ𝒢,τV,τE\)\\mathcal\{G\}=\(\\mathcal\{V\},\\mathcal\{E\}\_\{\\mathcal\{G\}\},\\tau\_\{V\},\\tau\_\{E\}\)denote a heterogeneous graph over the research landscape, where𝒱\\mathcal\{V\}contains researcher, paper, venue, institution, method, dataset, and topic nodes;ℰ𝒢\\mathcal\{E\}\_\{\\mathcal\{G\}\}contains co\-authorship, citation, publication, affiliation, usage, and topic\-assignment edges; andτV,τE\\tau\_\{V\},\\tau\_\{E\}assign node and edge types \(ℰ𝒢\\mathcal\{E\}\_\{\\mathcal\{G\}\}is distinct from the literature indexℐ\\mathcal\{I\}used below\)\. For a researcheru∈𝒰u\\in\\mathcal\{U\}, let𝒮u\\mathcal\{S\}\_\{u\}denote observed signals \(papers, citations, code, review history, venue preferences, explicit constraints\), let𝐳u∈ℝd\\mathbf\{z\}\_\{u\}\\in\\mathbb\{R\}^\{d\}be a graph\-derived representation ofuu, and letcuc\_\{u\}be the operational context given to the auto\-research system: \(1\)𝐳u=Enc𝒢\(u,𝒢\),cu=Φ\(𝒮u,𝐳u\)\.\\mathbf\{z\}\_\{u\}=\\mathrm\{Enc\}\_\{\\mathcal\{G\}\}\(u;\\mathcal\{G\}\),\\qquad c\_\{u\}=\\Phi\(\\mathcal\{S\}\_\{u\},\\mathbf\{z\}\_\{u\}\)\. ###### Definition 3\.1 \(Personalized Auto\-Research\)\. Letggbe a research goal and𝒫=\(p1,…,pK\)\\mathcal\{P\}=\(p\_\{1\},\\ldots,p\_\{K\}\)the stages of the research process \(literature retrieval, hypothesis generation, experiment design, code execution, writing, citation, refinement, review\)\. Whereas a researcher\-agnostic co\-scientist parameterizes each stage by the goal alone,oi=fi\(g,o<i\)o\_\{i\}=f\_\{i\}\(g,o\_\{<i\}\), personalized auto\-research learns stage models conditioned on the researcher, \(2\)oi=fi\(g,o<i∣cu\),o\_\{i\}=f\_\{i\}\(g,o\_\{<i\}\\mid c\_\{u\}\),whereo<io\_\{<i\}denotes outputs of earlier stages\. The desiderata are threefold: the output should be \(i\)*feasible*, respecting the researcher’s capabilities, resources, and constraints; \(ii\)*aligned*, compatible with the researcher’s scientific identity, community, and style; and \(iii\)*novel*, new to the field and distinct from the researcher’s prior work\. Feasibility and alignment are personalized properties, whereas novelty is partly field\-level and partly user\-relative\. This distinction is fundamentally important, since without it, personalization collapses into a recommender that predicts more of the same\. ## 4\.Personalized Auto\-Research Algorithm[1](https://arxiv.org/html/2608.14881#alg1)gives the proposed personalized auto\-research procedure\. The key idea is to formulate personalization as a modification of the SOTA agentic auto\-research loop, that is, agentic tree search over partial research states with an experiment manager, code execution, figure refinement, manuscript writing, automated review, and provenance logging\([33](https://arxiv.org/html/2608.14881#bib.bib16);[11](https://arxiv.org/html/2608.14881#bib.bib23);[28](https://arxiv.org/html/2608.14881#bib.bib24);[6](https://arxiv.org/html/2608.14881#bib.bib25)\), rather than the linear 2024 template loop\. Personalization adds a useruu, a contextcuc\_\{u\}, a personalized evidence setℛu\(g\)\\mathcal\{R\}\_\{u\}\(g\), and a user\-conditioned utility, where retrieval scores documentsddin corpus𝒟\\mathcal\{D\}jointly by goal and researcher andTopk\\TopOp\_\{k\}returns thekkhighest\-scoring items: \(3\)su\(d∣g\)\\displaystyle s\_\{u\}\(d\\mid g\)=⟨η\(d\),ρ\(g,cu\)⟩,ℛu\(g\)=Topk\(𝒟;su\(⋅∣g\)\),\\displaystyle=\\big\\langle\\eta\(d\),\\,\\rho\(g,c\_\{u\}\)\\big\\rangle,\\quad\\mathcal\{R\}\_\{u\}\(g\)=\\TopOp\_\{k\}\\\!\\big\(\\mathcal\{D\};s\_\{u\}\(\\cdot\\mid g\)\\big\),\(4\)U\(h∣g,u\)\\displaystyle U\(h\\mid g,u\)=αNov\(h,ℐ,cu\)\+βRel\(h,g,cu\)\+γFeas\(h,𝒲0,cu\)\.\\displaystyle=\\alpha\\Nov\(h,\\mathcal\{I\},c\_\{u\}\)\+\\beta\\Rel\(h,g,c\_\{u\}\)\+\\gamma\\Feas\(h,\\mathcal\{W\}\_\{0\},c\_\{u\}\)\.Notably, Eq\. \([4](https://arxiv.org/html/2608.14881#S4.E4)\) is one key distinction: hypotheses are*ranked*by personalized novelty, relevance, and feasibility, rather than filtered by a binary, global novelty test\. Intuitively, the same idea can be infeasible for one researcher, obvious to another, and transformative for a third whose graph position makes a new bridge credible\. Algorithm 1Personalized Auto\-Research1:language\-model agents π\\pi, experiment manager μ\\mu, vision\-language reviewer ω\\omega, research goal gg, initial workspace 𝒲0\\mathcal\{W\}\_\{0\}, seed archive 𝒥\\mathcal\{J\}, literature index ℐ\\mathcal\{I\}over corpus 𝒟\\mathcal\{D\}, research graph 𝒢\\mathcal\{G\}, user uuwith signals 𝒮u\\mathcal\{S\}\_\{u\}, context encoder Φ\\Phi, tree\-search budget NN, branch factor bb, package count mm, exp\. budget BB 2:personalized reprod\. research packages 𝒜u\\mathcal\{A\}\_\{u\} 3:// Personalized Context and Literature Grounding 4: 𝐳u←Enc𝒢\(u,𝒢\)\\mathbf\{z\}\_\{u\}\\leftarrow\\mathrm\{Enc\}\_\{\\mathcal\{G\}\}\(u;\\mathcal\{G\}\) 5: cu←Φ\(𝒮u,𝐳u\)c\_\{u\}\\leftarrow\\Phi\(\\mathcal\{S\}\_\{u\},\\mathbf\{z\}\_\{u\}\) 6: ℛu\(g\)←Topk\(𝒟;su\(⋅∣g\)\)\\mathcal\{R\}\_\{u\}\(g\)\\leftarrow\\TopOp\_\{k\}\(\\mathcal\{D\};s\_\{u\}\(\\cdot\\mid g\)\) 7: ℬu←\{\(𝒲0,∅,∅,∅,0\)\}\\mathcal\{B\}\_\{u\}\\leftarrow\\\{\(\\mathcal\{W\}\_\{0\},\\emptyset,\\emptyset,\\emptyset,0\)\\\} 8:// Personalized Hypothesis Search and Experimentation 9:for r=1r=1to NNdo 10: x←μ\(select,ℬu,g,ℛu\(g\),cu\)x\\leftarrow\\mu\(\\texttt\{select\},\\mathcal\{B\}\_\{u\},g,\\mathcal\{R\}\_\{u\}\(g\),c\_\{u\}\) 11: 𝒳←π\(⋅∣expand,x,g,𝒥,ℛu\(g\),cu,b\)\\mathcal\{X\}\\leftarrow\\pi\(\\cdot\\mid\\texttt\{expand\},x,g,\\mathcal\{J\},\\mathcal\{R\}\_\{u\}\(g\),c\_\{u\},b\) 12:foreach child state x′∈𝒳x^\{\\prime\}\\in\\mathcal\{X\}do 13: h∼π\(⋅∣hypothesize,x′,g,𝒥,ℛu\(g\),cu\)h\\sim\\pi\(\\cdot\\mid\\texttt\{hypothesize\},x^\{\\prime\},g,\\mathcal\{J\},\\mathcal\{R\}\_\{u\}\(g\),c\_\{u\}\) 14: U\(h∣g,u\)←αNov\(h,ℐ,cu\)\+βRel\(h,g,cu\)\+γFeas\(h,𝒲0,cu\)U\(h\\mid g,u\)\\leftarrow\\alpha\\Nov\(h,\\mathcal\{I\},c\_\{u\}\)\+\\beta\\Rel\(h,g,c\_\{u\}\)\+\\gamma\\Feas\(h,\\mathcal\{W\}\_\{0\},c\_\{u\}\) 15: 𝒞∼π\(⋅∣implement,x′,h,𝒲0,cu\)\\mathcal\{C\}\\sim\\pi\(\\cdot\\mid\\texttt\{implement\},x^\{\\prime\},h,\\mathcal\{W\}\_\{0\},c\_\{u\}\) 16: \(𝒪,ℱ,Λu\)←Execute\(𝒞,h,B,cu\)\(\\mathcal\{O\},\\mathcal\{F\},\\Lambda\_\{u\}\)\\leftarrow\\Execute\(\\mathcal\{C\},h,B,c\_\{u\}\) 17: ℱ∼ω\(⋅∣figure\-review,ℱ,𝒪,g,cu\)\\mathcal\{F\}\\sim\\omega\(\\cdot\\mid\\texttt\{figure\-review\},\\mathcal\{F\},\\mathcal\{O\},g,c\_\{u\}\) 18: qu←μ\(score,U\(h∣g,u\),h,𝒞,𝒪,ℱ,g,cu\)q\_\{u\}\\leftarrow\\mu\(\\texttt\{score\},U\(h\\mid g,u\),h,\\mathcal\{C\},\\mathcal\{O\},\\mathcal\{F\},g,c\_\{u\}\) 19: ℬu←ℬu∪\{\(h,𝒞,𝒪,ℱ,qu,Λu\)\}\\mathcal\{B\}\_\{u\}\\leftarrow\\mathcal\{B\}\_\{u\}\\cup\\\{\(h,\\mathcal\{C\},\\mathcal\{O\},\\mathcal\{F\},q\_\{u\},\\Lambda\_\{u\}\)\\\} 20:endfor 21:optionally: er←u\(feedback,Top1\(ℬu;qu\)\)e\_\{r\}\\leftarrow u\\big\(\\texttt\{feedback\},\\TopOp\_\{1\}\(\\mathcal\{B\}\_\{u\};q\_\{u\}\)\\big\), 𝒮u←𝒮u∪\{er\}\\mathcal\{S\}\_\{u\}\\leftarrow\\mathcal\{S\}\_\{u\}\\cup\\\{e\_\{r\}\\\}, cu←Φ\(𝒮u,𝐳u\)c\_\{u\}\\leftarrow\\Phi\(\\mathcal\{S\}\_\{u\},\\mathbf\{z\}\_\{u\}\) 22:endfor 23:// Personalized Research Package Synthesis 24: 𝒜u←∅\\mathcal\{A\}\_\{u\}\\leftarrow\\emptyset 25:foreach state \(h,𝒞,𝒪,ℱ,qu,Λu\)∈Topm\(ℬu;qu\)\(h,\\mathcal\{C\},\\mathcal\{O\},\\mathcal\{F\},q\_\{u\},\\Lambda\_\{u\}\)\\in\\TopOp\_\{m\}\(\\mathcal\{B\}\_\{u\};q\_\{u\}\)do 26: y∼π\(⋅∣write,g,h,𝒞,𝒪,ℱ,ℛu\(g\),cu\)y\\sim\\pi\(\\cdot\\mid\\texttt\{write\},g,h,\\mathcal\{C\},\\mathcal\{O\},\\mathcal\{F\},\\mathcal\{R\}\_\{u\}\(g\),c\_\{u\}\) 27: y∼π\(⋅∣cite\-refine,y,ℛu\(g\),𝒪,ℱ,cu\)y\\sim\\pi\(\\cdot\\mid\\texttt\{cite\-refine\},y,\\mathcal\{R\}\_\{u\}\(g\),\\mathcal\{O\},\\mathcal\{F\},c\_\{u\}\) 28: v∼π\(⋅∣review,y,g,𝒪,ℱ,cu\)v\\sim\\pi\(\\cdot\\mid\\texttt\{review\},y,g,\\mathcal\{O\},\\mathcal\{F\},c\_\{u\}\) 29: 𝒬u←\(h,U\(h∣g,u\),𝒞,𝒪,ℱ,y,v,Λu,cu,ℛu\(g\)\)\\mathcal\{Q\}\_\{u\}\\leftarrow\(h,U\(h\\mid g,u\),\\mathcal\{C\},\\mathcal\{O\},\\mathcal\{F\},y,v,\\Lambda\_\{u\},c\_\{u\},\\mathcal\{R\}\_\{u\}\(g\)\) 30: 𝒜u←𝒜u∪\{𝒬u\}\\mathcal\{A\}\_\{u\}\\leftarrow\\mathcal\{A\}\_\{u\}\\cup\\\{\\mathcal\{Q\}\_\{u\}\\\} 31:endfor 32:return 𝒜u\\mathcal\{A\}\_\{u\} ### Personalized context and evidence Lines[4](https://arxiv.org/html/2608.14881#alg1.l4)–[6](https://arxiv.org/html/2608.14881#alg1.l6)construct the user\-specific state: the graph encoder produces𝐳u\\mathbf\{z\}\_\{u\}, the context encoder producescuc\_\{u\}, and retrieval producesℛu\(g\)\\mathcal\{R\}\_\{u\}\(g\)\. Notably, this evidence set is not simply the topically closest literature togg\. It is the one that is relevant to the goal*and*useful for this researcher, given their prior work, collaborators, methods, resources, and position in𝒢\\mathcal\{G\}\. ### Personalized hypothesis search Lines[9](https://arxiv.org/html/2608.14881#alg1.l9)–[19](https://arxiv.org/html/2608.14881#alg1.l19)replace a universal agentic tree with a user\-conditioned treeℬu\\mathcal\{B\}\_\{u\}: the manager selects partial states, the language model expands them, and every hypothesis is scored byU\(h∣g,u\)U\(h\\mid g,u\), all undercuc\_\{u\}\. Two researchers with the same goal therefore induce*different search trees*, since their contexts change the selection policy, the expansion distribution, and the utility\. Furthermore, candidates are implemented and executed under the user’s constraints \(Lines[15](https://arxiv.org/html/2608.14881#alg1.l15)–[16](https://arxiv.org/html/2608.14881#alg1.l16)\), and even vision\-language figure feedback may depend oncuc\_\{u\}\(Line[17](https://arxiv.org/html/2608.14881#alg1.l17)\): a graph\-mining researcher and a biomedical collaborator need different visual encodings and terminology for the same quantitative result\. The researcher also remains in the loop: at any iteration,uumay optionally critique the current best state \(Line[21](https://arxiv.org/html/2608.14881#alg1.l21)\)\. Notably, this feedback does more than steer the session, as in existing human\-in\-the\-loop systems\([4](https://arxiv.org/html/2608.14881#bib.bib1);[23](https://arxiv.org/html/2608.14881#bib.bib4);[11](https://arxiv.org/html/2608.14881#bib.bib23)\), where guidance evaporates when the session ends\. Here it is folded into𝒮u\\mathcal\{S\}\_\{u\}and re\-encoded intocuc\_\{u\}, and thus the representation of the relationship itself is updated, both within a run and across runs\. ### Personalized package synthesis Lines[25](https://arxiv.org/html/2608.14881#alg1.l25)–[32](https://arxiv.org/html/2608.14881#alg1.l32)write, cite, refine, and review the top\-mmstates under the same context\. Each terminal artifact is a reproducible research package𝒬u=\(h,U\(h∣g,u\),𝒞,𝒪,ℱ,y,v,Λu,cu,ℛu\(g\)\)\\mathcal\{Q\}\_\{u\}=\(h,U\(h\\mid g,u\),\\mathcal\{C\},\\mathcal\{O\},\\mathcal\{F\},y,v,\\Lambda\_\{u\},c\_\{u\},\\mathcal\{R\}\_\{u\}\(g\)\)consisting of the hypothesis, its personalized score, code, outputs, figures, paper, automated review, provenance log, and the context and evidence that conditioned the run, that is, the metadata needed to audit*why this package was produced for this researcher*\. ## 5\.Personalized Auto\-Research Vision The distinction between an instrument and a collaborator is not one of capability but of relationship: a capable instrument returns high\-quality outputs, while a collaborator returns outputs tailored to the person it works with\. Personalization is therefore not a marginal improvement to AI co\-scientists, but rather what makes the metaphor accurate\. We organize the resulting research agenda around three fundamental components\. ### Researcher Representation\. A researcher’s publications alone are a thin description of their scientific identity\. Intuitively, two researchers with similar publication lists can have fundamentally different research identities if one lies in a dense cluster of theorists while the other bridges that cluster to computational biology\. This leads us to propose learning𝐳u\\mathbf\{z\}\_\{u\}from the researcher’s position in𝒢\\mathcal\{G\}, aggregating multi\-hop neighborhoods that capture who their collaborators work with, where their extended network publishes, and which topics are adjacent to their community\. Notably, this grounds personalization in structural properties of science: the most valuable direction for a researcher is often not an extension of their work but a*structural hole*, that is, a bridge between a region they inhabit and a nearby region not yet connected\([1](https://arxiv.org/html/2608.14881#bib.bib15)\)\. ### Pipeline Personalization\. Personalizing only hypothesis generation while leaving retrieval, experiment design, writing, citation, and review researcher\-agnostic is internally inconsistent: a hypothesis can be aligned yet require resources the researcher lacks and cite a literature that omits their community\. The design principle is to separate*exploitation*, which extends the researcher’s trajectory, from*exploration*, which uses the profile to judge feasibility and framing while keeping the candidate space broad\. The goal is to personalize the*how*without narrowing the*what*\. Personalization also need not stop at the stages: the components Algorithm[1](https://arxiv.org/html/2608.14881#alg1)treats as fixed can themselves be learned per researcher\. The weights in Eq\. \([4](https://arxiv.org/html/2608.14881#S4.E4)\) and theNov\\mathrm\{Nov\},Rel\\mathrm\{Rel\}, andFeas\\mathrm\{Feas\}functions can be fit from research traces such as revisions, reviews, and abandoned versus published projects; feasibility can be grounded in the researcher’s actual repositories, compute, and datasets rather than a self\-reported profile, and review calibrated to what they can rigorously verify, keeping autonomous science inside their expertise and auditable\. ### Evaluation Grounded in the Individual\. Predicting a researcher’s next papers is a flawed proxy: held\-out papers record what they happened to work on under path\-dependent incentives, not what they*should*have, and the proxy penalizes a co\-scientist that recommends something better than anything they pursued\. Evaluation should instead combine \(i\)*feasibility alignment*, whether directions are executable given documented capabilities; \(ii\)*expert\-assessed quality*, whether blinded experts judge ideas novel, significant, and appropriate for the profile; and \(iii\)*longitudinal impact*, whether co\-scientist use yields more influential work over multi\-year horizons\. Held\-out papers nonetheless serve as a scalable*necessary condition*: holding out a paper at timettand buildingcuc\_\{u\}from the record prior tott, one conditions the system on the general concept and measures*fidelity*, whether it recovers a hypothesis and experimental path close to what the researcher pursued, and*contrast*, whether the researcher\-agnostic system stays generic while different contexts diverge on the same concept\. High fidelity with high contrast showscuc\_\{u\}carries signal; failure falsifies the representation\. Fully automatable from bibliographic data, this protocol validates the personalization signal rather than the quality ceiling, and building benchmark infrastructure for all of these criteria is itself a major missing contribution\. ## 6\.Open Challenges We now discuss key open problems and challenges, which are deep tensions in the problem itself rather than engineering obstacles\. ### Creativity Collapse\. A researcher\-agnostic system applies one map from goal to output distribution, so the epistemic loss compounds at population scale: when many researchers query the same engine, the field’s portfolio of explored hypotheses collapses toward a monoculture, and globally “best” ideas are raced redundantly while directions of high counterfactual value, those only a particular researcher is positioned to pursue, go unexplored\. Personalization is therefore a decorrelation mechanism for collective discovery, not merely a convenience for individuals\. The challenge is to exploit a researcher’s experience without collapsing into biographical mimicry, treating it as a signal for experiments a generic system would not propose yet this researcher can develop\. This calls for rewarding*counterfactual complementarity*: ideas unlikely under both a universal model and the user’s past work alone, yet plausible and valuable given their accumulated experience\. One realization trains a mimicry model of the researcher and optimizes for high utilityU\(h∣g,u\)U\(h\\mid g,u\)at low likelihood under both the mimicry and universal models\. ### Lifecycle Dependence and Cold Start\. For a senior researcher with a dense graph, value lies in adjacent unexplored territory and bridging structural holes, whereas for an early\-career researcher with sparse history, the system must help*establish*an identity rather than extend one\. Notably, this exceeds the classical cold\-start problem, since the objective function itself changes with career stage\. ### Team\-Personalized Auto\-Research\. Most impactful research is carried out by teams rather than individuals\([32](https://arxiv.org/html/2608.14881#bib.bib14)\), yet the formulation above personalizes for a single researcher\. The framework extends naturally by replacinguuwith a teamT⊆𝒰T\\subseteq\\mathcal\{U\}and definingcT=Ψ\(\{\(𝒮u,𝐳u\)\}u∈T\)c\_\{T\}=\\Psi\(\\\{\(\\mathcal\{S\}\_\{u\},\\mathbf\{z\}\_\{u\}\)\\\}\_\{u\\in T\}\), where a team is simply a subgraph of𝒢\\mathcal\{G\}\. Notably, the desiderata aggregate asymmetrically: feasibility is a union, since the team can execute what any member can execute; novelty is relative to the union of prior work; and alignment is closer to an intersection, since framing must be evaluable by every community the team spans\. This asymmetry is precisely why scientists collaborate, and it enables fundamentally new capabilities, including routing experiments to the member best positioned to execute them, writing and citing for the union of the team’s communities, and recommending the collaborator whose addition closes the structural hole that makes a hypothesis credible\([1](https://arxiv.org/html/2608.14881#bib.bib15)\)\. Notably, Algorithm[1](https://arxiv.org/html/2608.14881#alg1)supports this setting with only local modifications: Lines[4](https://arxiv.org/html/2608.14881#alg1.l4)–[5](https://arxiv.org/html/2608.14881#alg1.l5)encode each member and aggregate viaΨ\\Psito obtaincTc\_\{T\}; retrieval \(Line[6](https://arxiv.org/html/2608.14881#alg1.l6)\) scores documents against the team context; the utility \(Line[14](https://arxiv.org/html/2608.14881#alg1.l14)\) becomesU\(h∣g,T\)U\(h\\mid g,T\)withFeas\(h,𝒲0,cT\)=maxu∈TFeas\(h,𝒲0,cu\)\\Feas\(h,\\mathcal\{W\}\_\{0\},c\_\{T\}\)=\\max\_\{u\\in T\}\\Feas\(h,\\mathcal\{W\}\_\{0\},c\_\{u\}\)and alignment taken as a minimum over members; implementation and execution \(Lines[15](https://arxiv.org/html/2608.14881#alg1.l15)–[16](https://arxiv.org/html/2608.14881#alg1.l16)\) assign each candidate toargmaxu∈TFeas\(h,𝒲0,cu\)\\arg\\max\_\{u\\in T\}\\Feas\(h,\\mathcal\{W\}\_\{0\},c\_\{u\}\); and the feedback step \(Line[21](https://arxiv.org/html/2608.14881#alg1.l21)\) collects critiques from multiple members, updating each𝒮u\\mathcal\{S\}\_\{u\}and re\-aggregatingcTc\_\{T\}\. However, aggregating conflicting member preferences into a singlecTc\_\{T\}is a social\-choice problem, the max\-based feasibility assumes frictionless handoffs between members, and multi\-party privacy becomes harder when the user is itself a group\. ### Evaluation Without Ground Truth\. The fundamental difficulty is that the quantity we need to evaluate, namely the value of a recommended direction for a specific researcher, is never observed\. History records only the single trajectory each researcher actually followed, and that trajectory was shaped by funding, advisors, reviewing, and chance rather than by an oracle over alternatives\. Scoring a system by similarity to this record therefore rewards mimicry and penalizes better recommendations \(§[5](https://arxiv.org/html/2608.14881#S5.SS0.SSS0.Px3)\)\. The held\-out protocol of §[5](https://arxiv.org/html/2608.14881#S5.SS0.SSS0.Px3)confirms only a necessary condition, namely thatcuc\_\{u\}carries researcher\-specific signal; it cannot certify that a recommended direction is*good*, since the value of an unpursued alternative is never recorded\. Resolving that distinction requires human judgment \(expert panels\), time \(longitudinal studies\), or interventions \(counterfactual designs\), each of which necessitates further research\. ## 7\.Conclusion We introduced*personalized auto\-research*, the problem of conditioning the full research process on a representation of the individual researcher\. The core claim is simple: a system cannot be a true co\-scientist if it does not know whom it is collaborating with\. The goal is complementarity rather than similarity, just as scientists choose collaborators for the expertise they lack, and so a stronger universal engine does not resolve the problem; it still returns the same high\-scoring package to everyone, which makes personalization orthogonal to, and required on top of, the current SOTA\. We therefore pose personalized auto\-research as a grand challenge for the AI ecosystem, spanning graph learning for representation, agentic systems for execution, HCI for consent and control, and community benchmarks for individual\-grounded evaluation\. No single group can deliver it alone, and the tensions it raises around novelty, equity, privacy, lifecycle, and evaluation define the field as much as its algorithms do\. ## References - Burt \(2004\)R\. S\. BurtStructural holes and good ideas\.American Journal of Sociology110\(2\),pp\. 349–399\.External Links:[Document](https://dx.doi.org/10.1086/421787)Cited by:[§5](https://arxiv.org/html/2608.14881#S5.SS0.SSS0.Px1.p1.1),[§6](https://arxiv.org/html/2608.14881#S6.SS0.SSS0.Px3.p1.1)\. - Chenet al\.\(2025\)T\. Chen, S\. Anumasa, B\. Lin, V\. Shah, A\. Goyal, and D\. LiuAuto\-Bench: an automated benchmark for scientific discovery in LLMs\.External Links:2502\.15224,[Document](https://dx.doi.org/10.48550/arXiv.2502.15224),[Link](https://arxiv.org/abs/2502.15224)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Egeret al\.\(2025\)S\. Eger, Y\. Cao, J\. D’Souza, A\. Geiger, C\. Greisinger, S\. Gross, Y\. Hou, B\. Krenn, A\. Lauscher, Y\. Li, C\. Lin, N\. S\. Moosavi, W\. Zhao, and T\. MillerTransforming science with large language models: a survey on AI\-assisted scientific discovery, experimentation, content generation, and evaluation\.External Links:2502\.05151,[Document](https://dx.doi.org/10.48550/arXiv.2502.05151),[Link](https://arxiv.org/abs/2502.05151)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Gottweiset al\.\(2025\)J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, A\. Palepu, P\. Sirkovic, A\. Myaskovsky, F\. Weissenberger, K\. Rong, R\. Tanno, K\. Saab, D\. Popovici, J\. Blum, F\. Zhang, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, P\. Kohli, Y\. Matias, A\. Carroll, K\. Kulkarni, N\. Tomasev, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penadés, G\. Peltz, Y\. Xu, A\. Pawlosky, A\. Karthikesalingam, and V\. NatarajanTowards an AI co\-scientist\.External Links:2502\.18864,[Document](https://dx.doi.org/10.48550/arXiv.2502.18864),[Link](https://arxiv.org/abs/2502.18864)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[Table 1](https://arxiv.org/html/2608.14881#S2.T1.6.3.2.1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p2.1),[§4](https://arxiv.org/html/2608.14881#S4.SS0.SSS0.Px2.p1.1)\. - Gottweiset al\.\(2026\)J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, P\. Sirkovic, A\. Myaskovsky, G\. Glowaty, F\. Weissenberger, A\. Orlandi, D\. Popovici, A\. Palepu, K\. Rong, R\. Tanno, K\. Saab, F\. Zhang, J\. Blum, A\. Carroll, K\. Kulkarni, N\. Tomašev, D\. Zverinski, I\. Rendulic, E\. Vedadi, F\. Hasler, L\. Rimanic, M\. Boia, I\. Budiselic, B\. Feinstein, M\. Bellaiche, T\. Sheffer, J\. Freyberg, J\. Ratcliff, O\. Bertolli, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penadés, G\. Peltz, Y\. Matias, J\. Manyika, D\. Hassabis, Y\. Xu, P\. Kohli, A\. Pawlosky, A\. Karthikesalingam, and V\. NatarajanAccelerating scientific discovery with co\-scientist\.Nature\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10644-y),[Link](https://www.nature.com/articles/s41586-026-10644-y)Cited by:[Table 1](https://arxiv.org/html/2608.14881#S2.T1.6.3.2.1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Jeddiet al\.\(2026\)A\. Jeddi, M\. N\. Le, H\. C\. Karaimer, K\. G\. Derpanis, and B\. TaatiGEAR: genetic AutoResearch for agentic code evolution\.External Links:2605\.13874,[Document](https://dx.doi.org/10.48550/arXiv.2605.13874),[Link](https://arxiv.org/abs/2605.13874)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1),[§4](https://arxiv.org/html/2608.14881#S4.p1.1)\. - Jumperet al\.\(2021\)J\. Jumper, R\. Evans, A\. Pritzel, T\. Green, M\. Figurnov, O\. Ronneberger, K\. Tunyasuvunakool, R\. Bates, A\. Žídek, A\. Potapenko, A\. Bridgland, C\. Meyer, S\. A\. A\. Kohl, A\. J\. Ballard, A\. Cowie, B\. Romera\-Paredes, S\. Nikolov, R\. Jain, J\. Adler, T\. Back, S\. Petersen, D\. Reiman, E\. Clancy, M\. Zielinski, M\. Steinegger, M\. Pacholska, T\. Berghammer, S\. Bodenstein, D\. Silver, O\. Vinyals, A\. W\. Senior, K\. Kavukcuoglu, P\. Kohli, and D\. HassabisHighly accurate protein structure prediction with AlphaFold\.Nature596\(7873\),pp\. 583–589\.External Links:[Document](https://dx.doi.org/10.1038/s41586-021-03819-2)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p2.1)\. - Korenet al\.\(2009\)Y\. Koren, R\. Bell, and C\. VolinskyMatrix factorization techniques for recommender systems\.Computer42\(8\),pp\. 30–37\.External Links:[Document](https://dx.doi.org/10.1109/MC.2009.263)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p3.1)\. - Kumaret al\.\(2024\)I\. Kumar, S\. Viswanathan, S\. Yerra, A\. Salemi, R\. A\. Rossi, F\. Dernoncourt, H\. Deilamsalehy, X\. Chen, R\. Zhang, S\. Agarwal,et al\.Longlamp: a benchmark for personalized long\-form text generation\.arXiv preprint arXiv:2407\.11016\.Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p3.1)\. - Lewiset al\.\(2020\)P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. KielaRetrieval\-augmented generation for knowledge\-intensive NLP tasks\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 9459–9474\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p3.1)\. - Liuet al\.\(2026\)J\. Liu, S\. Qiu, M\. Li, B\. Li, H\. Ji, S\. Han, X\. Ye, P\. Xia, Z\. Dong, M\. Chen, C\. Zhang, L\. Zhang, G\. Chen, H\. Tu, X\. Yang, L\. Feng, X\. Zhao, H\. Chen, J\. Zhou, X\. Wang, W\. Zhang, H\. Zhu, Y\. Li, J\. Mei, H\. Fei, J\. Zhang, L\. Li, L\. Zhang, Y\. Zhou, S\. Wang, C\. Xiong, J\. Zou, Z\. Zheng, C\. Xie, M\. Ding, and H\. YaoAutoResearchClaw: self\-reinforcing autonomous research with human\-AI collaboration\.External Links:2605\.20025,[Document](https://dx.doi.org/10.48550/arXiv.2605.20025),[Link](https://arxiv.org/abs/2605.20025)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[Table 1](https://arxiv.org/html/2608.14881#S2.T1.6.3.2.1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1),[§4](https://arxiv.org/html/2608.14881#S4.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2608.14881#S4.p1.1)\. - Luet al\.\(2024\)C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. HaThe AI scientist: towards fully automated open\-ended scientific discovery\.External Links:2408\.06292,[Document](https://dx.doi.org/10.48550/arXiv.2408.06292),[Link](https://arxiv.org/abs/2408.06292)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[Table 1](https://arxiv.org/html/2608.14881#S2.T1.6.2.2.1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Luet al\.\(2026\)C\. Lu, C\. Lu, R\. T\. Lange, Y\. Yamada, S\. Hu, J\. Foerster, D\. Ha, and J\. CluneTowards end\-to\-end automation of AI research\.Nature651,pp\. 914–919\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10265-5),[Link](https://www.nature.com/articles/s41586-026-10265-5)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Luoet al\.\(2025a\)E\. Luo, J\. Jia, Y\. Xiong, X\. Li, X\. Guo, B\. Yu, M\. Hao, L\. Wei, and X\. ZhangBenchmarking AI scientists for omics data driven biological discovery\.External Links:2505\.08341,[Document](https://dx.doi.org/10.48550/arXiv.2505.08341),[Link](https://arxiv.org/abs/2505.08341)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Luoet al\.\(2025b\)Z\. Luo, A\. Kasirzadeh, and N\. B\. ShahThe more you automate, the less you see: hidden pitfalls of AI scientist systems\.External Links:2509\.08713,[Document](https://dx.doi.org/10.48550/arXiv.2509.08713),[Link](https://arxiv.org/abs/2509.08713)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Luoet al\.\(2025c\)Z\. Luo, Z\. Yang, Z\. Xu, W\. Yang, and X\. DuLLM4SR: a survey on large language models for scientific research\.External Links:2501\.04306,[Document](https://dx.doi.org/10.48550/arXiv.2501.04306),[Link](https://arxiv.org/abs/2501.04306)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Mitcheneret al\.\(2025\)L\. Mitchener, A\. Yiu, B\. Chang, M\. Bourdenx, T\. Nadolski, A\. Sulovari, E\. C\. Landsness, D\. L\. Barabasi, S\. Narayanan, N\. Evans, S\. Reddy, M\. Foiani, A\. Kamal, L\. P\. Shriver, F\. Cao, A\. T\. Wassie, J\. M\. Laurent, E\. Melville\-Green, M\. Caldas, A\. Bou, K\. F\. Roberts, S\. Zagorac, T\. C\. Orr, M\. E\. Orr, K\. J\. Zwezdaryk, A\. E\. Ghareeb, L\. McCoy, B\. Gomes, E\. A\. Ashley, K\. E\. Duff, T\. Buonassisi, T\. Rainforth, R\. J\. Bateman, M\. Skarlinski, S\. G\. Rodriques, M\. M\. Hinks, and A\. D\. WhiteKosmos: an AI scientist for autonomous discovery\.External Links:2511\.02824,[Document](https://dx.doi.org/10.48550/arXiv.2511.02824),[Link](https://arxiv.org/abs/2511.02824)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. L\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray, J\. Schulman, J\. Hilton, F\. Kelton, L\. Miller, M\. Simens, A\. Askell, P\. Welinder, P\. Christiano, J\. Leike, and R\. LoweTraining language models to follow instructions with human feedback\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 27730–27744\.External Links:[Link](https://arxiv.org/abs/2203.02155)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p3.1),[§2](https://arxiv.org/html/2608.14881#S2.p3.1)\. - Puet al\.\(2025\)Y\. Pu, T\. Lin, and H\. ChenPiFlow: principle\-aware scientific discovery with multi\-agent collaboration\.External Links:2505\.15047,[Document](https://dx.doi.org/10.48550/arXiv.2505.15047),[Link](https://arxiv.org/abs/2505.15047)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Renet al\.\(2025\)S\. Ren, C\. Xie, P\. Jian, Z\. Ren, C\. Leng, and J\. ZhangTowards scientific intelligence: a survey of LLM\-based scientific agents\.External Links:2503\.24047,[Document](https://dx.doi.org/10.48550/arXiv.2503.24047),[Link](https://arxiv.org/abs/2503.24047)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Salemiet al\.\(2024\)A\. Salemi, S\. Mysore, M\. Bendersky, and H\. ZamaniLaMP: when large language models meet personalization\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics,pp\. 7370–7392\.External Links:[Link](https://aclanthology.org/2024.acl-long.399/)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p3.1)\. - Schmidgall and Moor \(2025\)S\. Schmidgall and M\. MoorAgentRxiv: towards collaborative autonomous research\.External Links:2503\.18102,[Document](https://dx.doi.org/10.48550/arXiv.2503.18102),[Link](https://arxiv.org/abs/2503.18102)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Schmidgallet al\.\(2025\)S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. BarsoumAgent laboratory: using LLM agents as research assistants\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 5977–6043\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.320),[Link](https://aclanthology.org/2025.findings-emnlp.320/)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p2.1),[§4](https://arxiv.org/html/2608.14881#S4.SS0.SSS0.Px2.p1.1)\. - Sonet al\.\(2025\)G\. Son, J\. Hong, H\. Fan, H\. Nam, H\. Ko, S\. Lim, J\. Song, J\. Choi, G\. Paulo, Y\. Yu, and S\. BidermanWhen AI co\-scientists fail: SPOT\-a benchmark for automated verification of scientific research\.External Links:2505\.11855,[Document](https://dx.doi.org/10.48550/arXiv.2505.11855),[Link](https://arxiv.org/abs/2505.11855)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Tanget al\.\(2025\)J\. Tang, L\. Xia, Z\. Li, and C\. HuangAI\-researcher: autonomous scientific innovation\.External Links:2505\.18705,[Document](https://dx.doi.org/10.48550/arXiv.2505.18705),[Link](https://arxiv.org/abs/2505.18705)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Tieet al\.\(2026\)G\. Tie, J\. Shi, D\. Song, Y\. Huang, Z\. Sheng, X\. Zhou, D\. Liu, P\. Zhou, Y\. Chen, R\. Xu, L\. He, Q\. Wen, M\. Li, C\. Lu, S\. Li, P\. Xie, Y\. Yuan, R\. Meng, L\. Xing, L\. Sun, C\. Xiong, P\. S\. Yu, and J\. GaoAutoResearch AI: towards AI\-powered research automation for scientific discovery\.External Links:2605\.23204,[Document](https://dx.doi.org/10.48550/arXiv.2605.23204),[Link](https://arxiv.org/abs/2605.23204)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Wang and Barabási \(2021\)D\. Wang and A\. BarabásiThe science of science\.Cambridge University Press\.External Links:[Document](https://dx.doi.org/10.1017/9781108610834)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1)\. - Wang and Luan \(2026\)Y\. Wang and Z\. LuanPARNESS: a paper harness for end\-to\-end automated scientific research with dynamic workflows, full\-text indexing, and cross\-run knowledge accumulation\.External Links:2605\.05258,[Document](https://dx.doi.org/10.48550/arXiv.2605.05258),[Link](https://arxiv.org/abs/2605.05258)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1),[§4](https://arxiv.org/html/2608.14881#S4.p1.1)\. - Weiet al\.\(2022\)J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. Chi, Q\. V\. Le, and D\. ZhouChain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 24824–24837\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Weiet al\.\(2025\)J\. Wei, Y\. Yang, X\. Zhang, Y\. Chen, X\. Zhuang, Z\. Gao, D\. Zhou, G\. Wang, Z\. Gao, J\. Cao, Z\. Qiu, M\. Hu, C\. Ma, S\. Tang, J\. He, C\. Song, X\. He, Q\. Zhang, C\. You, S\. Zheng, N\. Ding, W\. Ouyang, N\. Dong, Y\. Cheng, S\. Sun, L\. Bai, and B\. ZhouFrom AI for science to agentic science: a survey on autonomous scientific discovery\.External Links:2508\.14111,[Document](https://dx.doi.org/10.48550/arXiv.2508.14111),[Link](https://arxiv.org/abs/2508.14111)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Wenget al\.\(2025\)Y\. Weng, M\. Zhu, Q\. Xie, Q\. Sun, Z\. Lin, S\. Liu, and Y\. ZhangDeepScientist: advancing frontier\-pushing scientific findings progressively\.External Links:2509\.26603,[Document](https://dx.doi.org/10.48550/arXiv.2509.26603),[Link](https://arxiv.org/abs/2509.26603)Cited by:[Table 1](https://arxiv.org/html/2608.14881#S2.T1.6.2.2.1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Wuchtyet al\.\(2007\)S\. Wuchty, B\. F\. Jones, and B\. UzziThe increasing dominance of teams in production of knowledge\.Science316\(5827\),pp\. 1036–1039\.External Links:[Document](https://dx.doi.org/10.1126/science.1136099)Cited by:[§6](https://arxiv.org/html/2608.14881#S6.SS0.SSS0.Px3.p1.1)\. - Yamadaet al\.\(2025\)Y\. Yamada, R\. T\. Lange, C\. Lu, S\. Hu, C\. Lu, J\. Foerster, J\. Clune, and D\. HaThe AI scientist\-v2: workshop\-level automated scientific discovery via agentic tree search\.External Links:2504\.08066,[Document](https://dx.doi.org/10.48550/arXiv.2504.08066),[Link](https://arxiv.org/abs/2504.08066)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[Table 1](https://arxiv.org/html/2608.14881#S2.T1.6.2.2.1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1),[§4](https://arxiv.org/html/2608.14881#S4.p1.1)\. - Yanget al\.\(2025\)X\. Yang, X\. Yang, S\. Fang, Y\. Zhang, J\. Wang, B\. Xian, Q\. Li, J\. Li, M\. Xu, Y\. Li, H\. Pan, Y\. Zhang, W\. Liu, Y\. Shen, W\. Chen, and J\. BianR&D\-Agent: an LLM\-agent framework towards autonomous data science\.External Links:2505\.14738,[Document](https://dx.doi.org/10.48550/arXiv.2505.14738),[Link](https://arxiv.org/abs/2505.14738)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. R\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=WE_vluYUL-X)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Zhanget al\.\(2025\)P\. Zhang, X\. Hu, G\. Huang, Y\. Qi, H\. Zhang, X\. Li, J\. Song, J\. Luo, Y\. Li, S\. Yin, C\. Dai, E\. H\. Jiang, X\. Zhou, Z\. Yin, B\. Yuan, J\. Dong, G\. Su, G\. Qiao, H\. Tang, A\. Du, L\. Pan, Z\. Lan, and X\. LiuaiXiv: a next\-generation open access ecosystem for scientific discovery generated by AI scientists\.External Links:2508\.15126,[Document](https://dx.doi.org/10.48550/arXiv.2508.15126),[Link](https://arxiv.org/abs/2508.15126)Cited by:[§1](https://arxiv.org/html/2608.14881#S1.p1.1),[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Zhenget al\.\(2025\)T\. Zheng, Z\. Deng, H\. T\. Tsang, W\. Wang, J\. Bai, Z\. Wang, and Y\. SongFrom automation to autonomy: a survey on large language models in scientific discovery\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,Suzhou, China,pp\. 17733–17750\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.895),[Link](https://aclanthology.org/2025.emnlp-main.895/)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Zhouet al\.\(2025\)L\. Zhou, H\. Ling, C\. Fu, Y\. Huang, M\. Sun, W\. Yu, X\. Wang, X\. Li, X\. Su, J\. Zhang, X\. Chen, C\. Liang, X\. Qian, H\. Ji, W\. Wang, M\. Zitnik, and S\. JiAutonomous agents for scientific discovery: orchestrating scientists, language, code, and physics\.External Links:2510\.09901,[Document](https://dx.doi.org/10.48550/arXiv.2510.09901),[Link](https://arxiv.org/abs/2510.09901)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Zhuet al\.\(2025a\)K\. Zhu, J\. Zhang, Z\. Qi, N\. Shang, Z\. Liu, P\. Han, Y\. Su, H\. Yu, and J\. YouSafeScientist: toward risk\-aware scientific discoveries by LLM agents\.External Links:2505\.23559,[Document](https://dx.doi.org/10.48550/arXiv.2505.23559),[Link](https://arxiv.org/abs/2505.23559)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\. - Zhuet al\.\(2025b\)M\. Zhu, Q\. Xie, Y\. Weng, J\. Wu, Z\. Lin, L\. Yang, and Y\. ZhangAI scientists fail without strong implementation capability\.External Links:2506\.01372,[Document](https://dx.doi.org/10.48550/arXiv.2506.01372),[Link](https://arxiv.org/abs/2506.01372)Cited by:[§2](https://arxiv.org/html/2608.14881#S2.p1.1)\.
Similar Articles
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
A survey paper examining the transition of AI from task-specific assistants to workflow-level research automators, defining AutoResearch as the spectrum of AI-powered scientific workflow automation and analyzing challenges in autonomy, reproducibility, and accountability.
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
This survey examines the emerging field of AI-powered research automation (AutoResearch), analyzing how AI systems are moving from isolated task assistance to full workflow-level scientific discovery. It defines a spectrum from human-steered 'Vibe Research' to AI-led systems, and proposes five evaluation dimensions for scientific credibility.
AI for Auto-Research: Roadmap & User Guide
This paper surveys the capabilities and limitations of AI across the full research lifecycle, from idea generation to dissemination, identifying a sharp boundary between reliable assistance and unreliable autonomy. It provides a taxonomy, benchmark suite, tool inventory, and design principles for human-governed AI collaboration in research.
Towards End-to-End Automation of AI Research
A paper presenting The AI Scientist, a system that automates the entire research lifecycle from idea generation to peer review, demonstrating AI's growing capacity for scientific contribution.
Some insights on Personal Research Work and interview preparation
This article discusses the current limitations of AI in research-level work, arguing that while AI excels at using existing packages and engineering solutions, it still struggles with the deep hypothesis-driven iteration required for genuine research. The author also warns against extreme views on AI's capabilities and uses AlphaFold as an example to illustrate that structuring the problem is the hardest part, not the optimization.