Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis
Summary
Brain Researcher is an agentic platform that enhances AI-driven neuroimaging data analysis by enforcing analytic rigor, improving tool selection accuracy, and ensuring reproducibility. The study demonstrates substantial performance gains and integrates methodological judgment into the workflow.
View Cached Full Text
Cached at: 08/21/26, 10:10 AM
# Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis
Source: [https://arxiv.org/html/2608.19902](https://arxiv.org/html/2608.19902)
\\unnumbered
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the claim is limited to what the evidence supports\. Agents may reproduce failures including selective analysis, premature declarations of success and optimization of imperfect criteria\. We present Brain Researcher, an agentic research harness operating in a neuroimaging researcher’s computational environment under rules for admissible analyses, required checks and claim scope\. In benchmarks, Brain Researcher increased first\-choice tool\-selection accuracy across seven models by 70\.2 percentage points \(23\.3% without it versus 93\.6% with it\) and verifiable grounding from 4\.6% to 22\.0%\. In collaborator\-led and self\-evolving studies, multiverse analyses exposed analytic\-choice sensitivity, and scientific review classified claims as accepted, qualified, revised, blocked, rejected or deferred\. By linking decisions to evidence and provenance, Brain Researcher embeds methodological judgment within the workflow, not after it\.
Nicholas LuAffiliation:Stanford University, Stanford, CA, USAXinhui LiAffiliation:Tri\-institutional Center for Translational Research in Neuroimaging and Data Science \(TReNDS\), Georgia State University, Georgia Institute of Technology, Emory University, Atlanta, GA, USAJocelyn A\. RicardAffiliation:Stanford University, Stanford, CA, USACe JuAffiliation:Inria, CEA, Université Paris\-Saclay, Palaiseau, FranceHuan H\. WangAffiliation:Stanford University, Stanford, CA, USAChristian KindermannAffiliation:Stanford University, Stanford, CA, USAJeanette A\. MumfordAffiliation:Stanford University, Stanford, CA, USASteven DillmannAffiliation:Stanford University, Stanford, CA, USAJames KentAffiliation:The University of Texas at Austin, Austin, TX, USAAlejandro de la VegaAffiliation:The University of Texas at Austin, Austin, TX, USASanmi KoyejoAffiliation:Stanford University, Stanford, CA, USAVince D\. CalhounAffiliation:Tri\-institutional Center for Translational Research in Neuroimaging and Data Science \(TReNDS\), Georgia State University, Georgia Institute of Technology, Emory University, Atlanta, GA, USAJoshua W\. BuckholtzAffiliation:Stanford University, Stanford, CA, USAJuan Helen ZhouAffiliation:National University of Singapore, SingaporeSteffen BollmannAffiliation:Stanford University, Stanford, CA, USAAffiliation:The University of Queensland, Brisbane, QLD, AustraliaRussell A\. PoldrackAffiliation:Stanford University, Stanford, CA, USA
###### keywords
neuroimaging; scientific agents; reproducibility; research infrastructure; analytic flexibility
Science often advances by absorbing its former frontiers into infrastructure: what once marked the edge of scientific practice, such as sequencing a genome or preprocessing a brain scan, becomes a routine step in a larger workflow\. AI agents may represent the next phase of this progression\. They increasingly interpret goals, call external tools, observe intermediate results, and choose subsequent actions\([5](https://arxiv.org/html/2608.19902#bib.bib2);[52](https://arxiv.org/html/2608.19902#bib.bib26);[27](https://arxiv.org/html/2608.19902#bib.bib45);[30](https://arxiv.org/html/2608.19902#bib.bib15);[41](https://arxiv.org/html/2608.19902#bib.bib20)\)\. Yet scientific research is not only a sequence of procedures, and executing an analysis is not the same as establishing a claim\. This distinction is especially consequential in neuroimaging, where large, heterogeneous datasets enter long analysis pipelines with multiple defensible choices\. Decisions about preprocessing, parcellation, confound adjustment, and model specification can materially reshape results\([45](https://arxiv.org/html/2608.19902#bib.bib21);[36](https://arxiv.org/html/2608.19902#bib.bib35);[20](https://arxiv.org/html/2608.19902#bib.bib11);[9](https://arxiv.org/html/2608.19902#bib.bib6)\); when seventy teams analyzed the same neuroimaging dataset, no two used the same workflow and their conclusions differed substantially\([6](https://arxiv.org/html/2608.19902#bib.bib3)\)\.
Neuroimaging has also developed one of the most mature open\-science ecosystems for computational automation\. BIDS and OpenNeuro standardize data organization and sharing\([46](https://arxiv.org/html/2608.19902#bib.bib22);[39](https://arxiv.org/html/2608.19902#bib.bib19)\); fMRIPrep automates functional MRI preprocessing\([21](https://arxiv.org/html/2608.19902#bib.bib12)\); and Nipype integrates software packages into reproducible workflows\([26](https://arxiv.org/html/2608.19902#bib.bib13)\)\. BIDS Statistical Models and FitLins extend this standardization to machine\-readable statistical models and their execution\([12](https://arxiv.org/html/2608.19902#bib.bib1);[40](https://arxiv.org/html/2608.19902#bib.bib36)\), while guided multiverse analysis makes alternative workflows more navigable\([13](https://arxiv.org/html/2608.19902#bib.bib7)\)\. Together, these tools make procedures, and increasingly their alternatives, reusable and comparable\. They do not, however, bind a selected route to the evidence and methodological conditions required for the claim it is used to support\. The challenge is therefore not simply to automate neuroimaging analysis, but to ensure that automation does not obscure the methodological conditions under which an analysis may support a claim\.
Tool\-using agents create both an opportunity and a risk\. They can connect fragmented stages of an analysis and adapt after observing intermediate outcomes, but successful command execution does not guarantee a scientifically valid analysis\. An agent may select an inappropriate tool, overlook a data or model incompatibility, stop prematurely, or optimize a criterion that is misaligned with the intended claim\. These failures parallel questionable research practices in human\-executed science\([48](https://arxiv.org/html/2608.19902#bib.bib49);[32](https://arxiv.org/html/2608.19902#bib.bib50)\), and are compounded by premature declarations of task completion\([29](https://arxiv.org/html/2608.19902#bib.bib52);[33](https://arxiv.org/html/2608.19902#bib.bib51)\)and reward hacking\([50](https://arxiv.org/html/2608.19902#bib.bib53);[22](https://arxiv.org/html/2608.19902#bib.bib54)\)\. Scientific agents therefore need more than access to tools: they need reliable routing, evidence linked to its methodological conditions, explicit exposure of alternative specifications, and a visible boundary between formalizable checks and judgments that remain with the researcher\.
Here we present Brain Researcher, a researcher\-governed, domain\-specific agentic harness that operates inside a neuroimaging researcher’s existing computational environment\. Brain Researcher is designed to preserve rather than replace scientific judgment: it operationalizes the parts that researchers can state in advance and records the rest for inspection\. For prospectively governed analyses, the researcher specifies the question, admissible analyses, required checks, and scope of any resulting claim\. A tool registry and the Brain Researcher Knowledge Graph connect analysis routes to evidence and method conditions; a Model Context Protocol server mediates model actions; and execution and review layers return an audit bundle linking the committed plan, tool calls, artifacts, evidence, and claim verdicts \(Fig\.[1](https://arxiv.org/html/2608.19902#S1.F1); Supplementary Methods S1\)\. Multiverse analyses expose sensitivity to defensible specifications\([51](https://arxiv.org/html/2608.19902#bib.bib25)\), while commitment and claim cards preserve what was decided before and after results were observed\. In self\-evolving research, intermediate evidence can redirect the trajectory only within a researcher\-defined action space specifying the admissible analyses, datasets, and evaluation budget; each successor analysis is frozen before execution\. Judgments that resist formalization remain with the researcher\.
We evaluated Brain Researcher in three settings\. First, a paired seven\-model tool\-calling benchmark and a separate evidence\-citation benchmark tested whether the harness improves upstream tool selection and evidence citation\. Second, three collaborator\-led studies tested whether multiverse analyses and explicit constraints make claim sensitivity and status visible\. Third, two self\-evolving research episodes tested whether evidence from one stage could be carried into a frozen successor analysis\. Together, these evaluations ask whether AI assistance becomes more capable and more defensible when methodological commitments, evidence, and claim scope are made explicit and auditable\.
## 1Results
Figure 1:Brain Researcher: workspace\-centric infrastructure for auditable neuroimaging research\.Brain Researcher runs inside the researcher’s existing computational environment and exposes neuroimaging analyses as structured,*auditable*operations: every choice, check, and input is recorded as it happens, so that a frozen record of the completed analysis can be read and audited by someone other than the person who ran it, without re\-executing it\. \(a\) The tool ecosystem: established neuroimaging software for preprocessing, modeling, meta\-analysis, machine learning, quality control, and reporting, each represented by a machine\-readable specification and executed in a version\-pinned container\. \(b\) Each specification declares its inputs, outputs, parameters, version, evidence anchors, and validation rules; the rule checker tests every proposed call against these clauses before it runs, and records which clauses passed or failed\. \(c\) The episode workflow: the researcher frames a question and approves the plan at the commitment gate, valid actions are dispatched to version\-pinned executors, and the resulting audit bundle, containing the committed plan, tool versions, evidence consulted, artifacts, logs, provenance, and the checks each claim passed, feeds the review layer, which writes condition\-tagged claims back to memory\. This audit bundle is what makes an analysis auditable: a completed run is a fully inspectable research object rather than a one\-off result, exported as a compact claim card a reviewer can reopen field by field\. Methodological judgment remains the researcher’s; the system makes it visible at each stage\.Figure 2:BR\-KG: provenance\-linked semantic integration for grounded, auditable neuroimaging reasoning\.BR\-KG integrates existing ontologies, repositories, data resources, and literature into a single graph \(745,949 nodes, 2,461,469 edges; 2026\-07\-07 release snapshot\) aligned to the OpenNeuro Vocabulary \(ONVOC\)\([44](https://arxiv.org/html/2608.19902#bib.bib27)\), which normalizes heterogeneous terms to shared identifiers so a query resolves consistently across sources\. The central graph links neuroimaging concepts \(tasks, contrasts, cognitive constructs\), neural representations \(brain regions, statistical maps\), and research resources \(datasets, tools\) through typed relationships, with literature evidence attached\. Crucially, source\-backed facts carry explicit provenance \(their source, and where available a verbatim supporting quote and grounding label\), so a retrieved claim can be traced back to the study and passage that support it rather than taken on trust; coverage is partial and tracked, and this is what makes retrieval here*auditable*\. Downstream panels show the payoff: grounded query answering and multi\-hop reasoning over concept\-task\-map paths, each hop inspectable down to its underlying nodes, edges, and cited evidence, which lets the review layer attach a recommendation’s method\-condition checks \(cohort, paradigm, preprocessing, statistical model\) before it is accepted\.### 1\.1Brain Researcher converts research questions into auditable claim records
At the center of Brain Researcher is a complete, auditable record of a single run: a persistent*episode*linking the question to its assembled evidence, the analyses the researcher deemed admissible, the committed plan, the execution, the review, and a condition\-tagged conclusion \(Supplementary Methods S2\)\. Two dated records can anchor it\. A*commitment card*, written before any analysis runs, fixes the question, the allowed alternatives, and the success and failure criteria, and is sealed with a content hash so that any later change to the plan is detectable\. A*claim card*, written afterward by the review layer, records the resulting claim, its assigned state, scope, and the checks it passed and failed; the supplement includes a specific worked example, built on public Neurosynth data, that a reader can open and inspect field by field \(Supplementary Methods S8\.5\.1; Appendix G\)\. A reviewer can then reopen and audit the record without re\-executing the analysis\. Across the collaborator\-led and self\-evolving evaluations reported below, every evaluated claim is assigned one of six states \(accepted, qualified, revised, blocked, rejected, or deferred\), each defined by an explicit adjudication rule\. For example, when a researcher asks whether two groups differ on a functional\-connectivity measure, the commitment card records the cohort, the subject groups, and the estimator \(either entered by the researcher or resolved from the dataset\) and fixes the checks that must pass before anything runs \(aligned subject groups, a full\-rank design matrix, no feature leakage across folds\)\. If the difference then holds in only some admissible specifications, the claim is recorded as qualified, with its conditions attached\. These methodological decisions are the researcher’s; Brain Researcher records each one and enforces the checks it implies\.
### 1\.2Brain Researcher improves tool calling and evidence citation
We first isolated the infrastructure’s effect on two decisions that precede scientific review: choosing the correct analysis tool and citing checkable evidence\. The tool\-calling benchmark pairs each request with a hidden scoring target and reports Correct route/tool@k \(all\-or\-nothing match of the analysis family and required capabilities\), Capability@k \(graded capability coverage\), and Handoff score@k \(whether the proposed route is specified completely enough to execute, with the required inputs present, so it can be handed to a downstream executor and run without gaps\) \(Supplementary Methods S11\.1\.1\); three condition\-blind LLM judges credit any response that reaches the required capabilities, whether through a Brain Researcher call or an equivalent executable route, so the measured gain reflects reaching those capabilities rather than credit for naming Brain Researcher’s specific tools\. Both conditions used the same seven frontier models\([4](https://arxiv.org/html/2608.19902#bib.bib38);[43](https://arxiv.org/html/2608.19902#bib.bib39);[25](https://arxiv.org/html/2608.19902#bib.bib40);[54](https://arxiv.org/html/2608.19902#bib.bib41);[14](https://arxiv.org/html/2608.19902#bib.bib42);[42](https://arxiv.org/html/2608.19902#bib.bib43);[1](https://arxiv.org/html/2608.19902#bib.bib44)\)and general\-purpose tools, the without\-BR condition lacking only Brain Researcher’s registry \(Supplementary Methods S5\), knowledge graph, and constraint layer\. Across 60 tool\-calling tasks and seven models, scores without versus with Brain Researcher were 23\.3% versus 93\.6% for first\-action correct route/tool selection \(with\-BR 95% CI 88\.8–97\.1, task\-clustered\), 49\.8% versus 94\.5% for mean Capability@1, and 47\.4% versus 76\.1% for handoff sufficiency\. The reference route for each task was fixed before either condition ran and curated with a co\-author who does not develop the Brain Researcher system; equivalent non\-BR routes were set by two model reviewers and one human \(Supplementary Methods S11\.1\.2\)\. All seven models improved on all three tool\-calling metrics \(7/7 positive paired differences; exact two\-sided Wilcoxon signed\-rankp=0\.016p=0\.016for each\), with mean gains of 70\.2 percentage points for Correct route/tool@1, 44\.7 points for Capability@1, and 28\.7 points for Handoff score@1\. In a routing ablation across 60 tasks and seven models, Brain Researcher without direct KG calls selected an acceptable exact top\-1 route in 362 of 420 episodes \(86\.2%; model range, 81\.7–90\.0%; details in Supplementary Methods S11\.1\.5\)\. A separate 50\-question benchmark counted a claim as grounded only when its cited evidence could be located and judged supportive; under a three\-judge majority vote, the descriptive question\-level verified\-groundedness rate rose from 4\.6% to 22\.0% \(95% CI 16\.8–27\.2, question\-clustered\), a 4\.8\-fold increase, though most evidence rows still failed, so grounding improved substantially without being solved\. Among the 444 non\-verified with\-BR rows present in all three judge outputs, 65% received an exact\-label majority of real but off\-topic and 28% of partial support; the remaining 7% lacked an exact\-label majority or could not be judged, and none had a fabricated or malformed majority \(Supplementary Methods S11\.1\.1\)\. Inter\-judge reliability and its dependence on judge strictness are reported in Supplementary Methods S11\.1\.2\. Secondary single\-judge safeguards confirmed that the gain was not accompanied by more unrelated citations or lower answer correctness \(Supplementary Methods S11\.1\.1; Appendix J; Fig\.[3](https://arxiv.org/html/2608.19902#S1.F3)\)\. Because the without\-BR condition removes Brain Researcher’s registry, knowledge graph, and constraint layer together, this contrast measures the harness as a whole rather than isolating any single component; and because the reference routes were curated with a co\-author, target construction may share vocabulary with the registry\.
Figure 3:Summary of Brain Researcher effects across quantitative benchmark tasks\.Without\-BR \(gray\) and with\-BR \(blue\) benchmark performance\. Capability@k is mean coverage of required task capabilities after the first k non\-neutral actions\. The left column reports Capability@1 and @3 across the seven model variants \(Claude Opus 4\.8, Codex GPT\-5\.5, Gemini 3\.1 Pro, GLM\-5\.1, DeepSeek\-V4\-Pro, Kimi K2\.5, Qwen3\.6\-Plus\)\. Upper\-right panels break Capability@1 down by task domain\. Lower panels report Handoff score@1 and @3 \(whether the first route carries enough information for another agent to continue\) and a Gemini 2\.5 Flash single\-judge safeguard: precision among claims marked grounded \(fraction whose cited evidence was both locatable and judged supportive\)\. Correct route/tool@1, the first\-action selection accuracy reported in the text \(23\.3% to 93\.6%\), is detailed in Supplementary Methods S11\.1\.1\. Metrics are interpreted within panel, as denominators and scoring rules differ across benchmarks\.
### 1\.3Brain Researcher runs multiverse analyses to expose claim sensitivity
We next evaluated Brain Researcher on three active neuroimaging research questions from collaborating scientists \(schizophrenia NeuroMark connectivity, cocaine\-use\-disorder connectivity, and cross\-cultural social\-cognition meta\-analysis\), chosen for heterogeneity in evidence structure without regard to the specific outcomes \(Fig\.[4](https://arxiv.org/html/2608.19902#S1.F4)\)\. Every reported analysis case was run by a coding agent on the local system, which called Brain Researcher for grounding, logging, and review \(Supplementary Methods S11\.2\)\. The NeuroMark case starts from a single, well\-established pipeline that its developers use as their standard, giving the audit one clearly defined baseline to build the multiverse around; the other two cases have no such established single pipeline, and test whether the workflow extends to that more difficult setting\. The collaborators’ hypotheses were pre\-specified in their own protocols rather than sealed as commitment cards, so the NeuroMark record is a post\-hoc audit of the completed multiverse\.
A collaborator studying schizophrenia functional network connectivity using the NeuroMark framework\([16](https://arxiv.org/html/2608.19902#bib.bib30);[31](https://arxiv.org/html/2608.19902#bib.bib31)\)brought three pre\-specified hypotheses for robustness audit: latent connectivity factors outperform individual edges for patient\-versus\-control classification \(NM\-H1\); between\-domain connections show larger group differences than within\-domain ones \(NM\-H2\); and latent factors concentrate loading mass on between\-domain edges \(NM\-H3\)\. We evaluated these hypotheses in the FBIRN cohort\([34](https://arxiv.org/html/2608.19902#bib.bib32)\)\(N=363N=363; 181 controls, 182 patients\), parcellated through NeuroMark 2\.2 template\-based independent component analysis into 5,460 edges per subject\. In the collaborator’s workspace, Brain Researcher expanded the analysis into a 480\-specification multiverse spanning connectivity, confound, dimensionality\-reduction, classifier, and domain\-granularity choices, and recorded and reviewed the resulting runs\.
None of the three hypotheses was supported uniformly across specifications; all were recorded as qualified, but for different patterns of conditional support\. Under the corrected sign\-aware criterion \(p<0\.05p<0\.05andΔmean\|d\|\>0\\Delta\\,\\mathrm\{mean\}\|d\|\>0\), 12 of 24 unique connectivity–confound–domain contrasts favored NM\-H2\. This pooled fraction obscured a complete estimator split: 100% of contrasts were favorable under Pearson and Spearman and 0% under partial correlation and mutual information\. NM\-H2 is therefore an estimator\-regime–dependent finding rather than a generally robust effect\. Because partial correlation and mutual information alter the dependence measure in non\-equivalent ways, distinguishing shared covariance from estimator scale, power, or nonlinearity requires targeted follow\-up\. NM\-H1 and NM\-H3 were also weak: edges outperformed latent factors in aggregate \(medianΔ\\DeltaAUC=−0\.032=\-0\.032; only 18\.8% of specifications favored latent features\), and only 26\.0% favored between\-domain loading mass\. These claims were qualified rather than rejected because support persisted within identifiable analytic subfamilies \(Supplementary Methods S11\.2\.1; Fig\.[4](https://arxiv.org/html/2608.19902#S1.F4)A–C\); a claim with no supporting subfamily is rejected instead \(Supplementary Methods S8\.4\)\.
NM\-H2 also supplied the audit’s governance lesson: automated review missed an error that a human caught\. After a server\-side fault triggered fallback to a general\-purpose coding agent, the agent scored any specification with permutationp<0\.05p<0\.05as favorable regardless of sign, inflating apparent support for a directional hypothesis to near\-universal levels\. The review layer did not flag the error; a human reviewer detected it by inspecting the code, outputs, and specification curve, leading to the corrected rescoring above\. Two checks were then added to the Brain Researcher skillset: a directionality test requiring the statistic and acceptance rule to match the hypothesized sign, and a warning whenever execution falls back to a general\-purpose agent\. The review missed this error\. The record nevertheless provided value by binding each claim to its hypothesis, statistic, and conditions, thereby turning a one\-off correction into an enforced check\.
Figure 4:Multiverse sensitivity and claim\-review outcomes across three collaborator episodes\.\(A–C\)Schizophrenia NeuroMark audit: \(A\) group\-mean functional connectivity for controls \(HC,N=181N=181\), patients \(SZ,N=182N=182\), and their difference across four estimators; \(B\) NM\-H2 \(between\- versus within\-domain\) specification curve over the 480\-specification multiverse; after sign\-aware rescoring, its estimand comprises 24 unique connectivity–confound–domain contrasts, with favorable support at 100% for Pearson and Spearman and 0% for partial correlation and mutual information\. This complete estimator partition, rather than the pooled 12\-of\-24 fraction, is the informative result: NM\-H2 is measure\-dependent, and the mechanism underlying the partition remains unresolved\. \(C\) Marginal influence of each analytic choice on NM\-H2\.\(D, E\)Cocaine\-use\-disorder episode: \(D\) multiverse stability of systemic\-segregation associations across 36 specifications with SDMA\-GLS consensus; \(E\) single\-specification versus multiverse SDMA\-GLS maps for five network–outcome pairs\.\(F\)Cross\-cultural social cognition: culture\-stratified ALE maps contrasting Euro\-American trust networks with East Asian social\-cognition networks\.In the other two episodes, prespecified checks in the scientific review layer determined whether a result could receive confirmatory status\. The SUDMEX CONN \(OpenNeuro ds003346;N=138N=138\)\([2](https://arxiv.org/html/2608.19902#bib.bib28);[23](https://arxiv.org/html/2608.19902#bib.bib29)\)example assessed associations between brain connectivity and behavior; a 36\-specification multiverse rejected all five pre\-specified connectivity–behavior associations under same\-dataset meta\-analysis \(SDMA\-GLS\)\([35](https://arxiv.org/html/2608.19902#bib.bib16)\)\(allZ<1\.24Z<1\.24, false discovery rate \[FDR\]q\>0\.58q\>0\.58\), and an exploratory screen over 70 combinations surfaced no FDR\-surviving effect, so the system blocked it from confirmatory promotion and converted the null into a replication plan \(Fig\.[4](https://arxiv.org/html/2608.19902#S1.F4)D,E\)\. In another test case that applied coordinate\-based neuroimaging meta\-analysis to a small cross\-cultural neuroscience literature, subgroup activation\-likelihood estimation \(ALE\) on 21 studies\([18](https://arxiv.org/html/2608.19902#bib.bib9);[19](https://arxiv.org/html/2608.19902#bib.bib10);[47](https://arxiv.org/html/2608.19902#bib.bib23);[15](https://arxiv.org/html/2608.19902#bib.bib8)\)produced a medial prefrontal cortex \(mPFC\)\-topology interpretation, but the system blocked it as exploratory: the subgroups held onlyk=6k=6–88entries \(below the recommendedk≥17k\\geq 17\), paradigm composition was imbalanced, and centroid shifts alone cannot establish non\-overlapping distributions\. The case ended in a paradigm\-matched follow\-up with no settled claim \(Fig\.[4](https://arxiv.org/html/2608.19902#S1.F4)F\)\. Full statistics are in Supplementary Methods S11\.2\.1 and Appendix J\.
Across the three episodes, the multiverse exposed which findings were sensitive to analytic choices, while scientific review determined what each result could support\. Prespecified review checks withheld confirmatory status from the SUDMEX exploratory screen and the underpowered cross\-cultural ALE; in the post\-hoc NeuroMark audit, a human reviewer identified the sign\-blind scoring error\.
### 1\.4Brain Researcher converts adaptive searches into frozen successor analyses in two self\-evolving episodes
We next asked whether Brain Researcher could transform open\-ended exploration into frozen, auditable successor analyses\. We examined two extended research episodes that differed in what was searched\. Using the Human connectome Project \(HCP\) data, Brain Researcher searched over candidate analysis workflows for a fixed question about connectivity\-based prediction of behavioural variation\. Using the TRIBE foundation model, it searched over candidate scientific questions about a model’s internal representations and then converted one question into a frozen test on newly sampled stimuli \(Supplementary Methods S11\.3\)\.
In the HCP episode, we began with a published study with openly available code and shared analysis materials\([37](https://arxiv.org/html/2608.19902#bib.bib17)\)\. Brain Researcher allocated 116 candidate prediction\-pipeline evaluations for Cognition, of which 104 returned scored results in the parent runs\. Following a selector audit, the researcher designated a frozen selected workflow\. In 10 repeated same\-cohort nested\-cross\-validation splits, the frozen selected workflow achieved a higher pooled out\-of\-fold correlation than a matched local reconstruction of the published procedure \(medianΔr=\.098\\Delta r=\.098; conditional one\-sidedp=\.006p=\.006\)\. When the frozen selected workflow was refit to four additional behavioural outcomes, it again produced higher correlations in 37 of 40 comparisons, giving the same direction in 47 of 50 comparisons across all five outcomes\. Median out\-of\-sampleR2R^\{2\}was positive only for two variables \(Cognition and Tobacco Use\), and multiplicity\-aware transfer inference remained inconclusive \(Fig\.[5](https://arxiv.org/html/2608.19902#S1.F5); Supplementary Methods S11\.3\.1\)\.
Figure 5:Brain Researcher searches 116 HCP prediction pipelines and identifies a workflow that consistently exceeds a matched reference\.A\.Brain Researcher first evaluated 20 candidate pipelines for Cognition prediction, reaching a best discovery score ofr=\.373r=\.373\. Brain researcher then launched a 96\-candidate expansion; 84 candidates returned scores and 12 ended in transport failure\. Within the expanded episode, Brain Researcher adapted its proposals to the accumulating results: 27 candidates exceeded the initial search maximum, and the highest discovery score wasr=\.487r=\.487, obtained with whole\-band coherence and ridge regression\. Following a selector audit, the researcher froze a related coherence\-based workflow for matched evaluation\. Across 10 repeated family\-grouped5×35\\times 3nested\-cross\-validation runs, this workflow achieved medianr=\.332r=\.332, compared with\.235\.235for the matched reference \(medianΔr=\.098\\Delta r=\.098; conditional one\-sidedp=\.006p=\.006\), and was higher in all 10 runs\.B\.The same frozen selected workflow was then refit for each of four additional behavioural outcomes without target\-specific retuning\. It produced a higher median correlation for every outcome and exceeded the matched reference in 37 of 40 repeat\-level comparisons, giving 47 of 50 directional wins across all five outcomes\.TRIBE v2\([17](https://arxiv.org/html/2608.19902#bib.bib37)\)is a tri\-modal foundation model that predicts human fMRI responses from video, audio, and language inputs\. We asked how natural\-sound category geometry changes across its internal audio layers\. Brain Researcher screened category contrasts without choosing one in advance and ranked them by changes in source\-held\-out discrimination \(Fig\.[6](https://arxiv.org/html/2608.19902#S1.F6)A\)\. Tools–voice showed the largest change but varied across sound collections\. Brain Researcher instead proposed speech–tools for follow\-up because early layers strongly separated the categories, whereas later layers brought them closer while largely preserving the same representational direction\. This contrast could also be tested prospectively using new recordings sampled from multiple collections and matched on seven prespecified acoustic measurements\. The researcher approved this direction and froze the hypothesis and analysis before the new stimuli were evaluated\.
Brain Researcher then evaluated three successive, non\-overlapping 48\-item panels\. The normalized speech–tools separation became smaller in later layers in 11 of 12 collection\-by\-panel comparisons, and all three panels met the prespecified directional criterion\. In most collections, later TRIBE layers preserved the representational direction separating speech from tools while bringing the categories closer together\. The result was a direction\-preserving contraction of speech–tools geometry\. After the pattern recurred across all three panels, Brain Researcher extended the test to four previously unused sound collections\. Three showed the same geometry, although the corrected collection\-level test remained inconclusive \(Holm\-adjustedp=\.396p=\.396; Fig\.[6](https://arxiv.org/html/2608.19902#S1.F6)B,C; Supplementary Methods S11\.3\.2\)\.
Figure 6:Brain Researcher turns an open question about TRIBE into successive tests with new sounds and collections\.A\.Brain Researcher began by asking how TRIBE changes natural\-sound representations from early to late layers\. It screened category contrasts by their change in held\-out distinguishability \(AUC\), without choosing a target in advance\. Tools–voice changed most, but the pattern varied across collections\. Rather than simply following the top\-ranked result, Brain Researcher identified speech–tools as a clearer lead: the categories moved closer in later layers while usually keeping the same representational direction\. The contrast could also be retested with new, acoustically matched sounds from several collections\. The researcher approved this direction and froze the prediction and analysis\.B\.Brain Researcher then evaluated three non\-overlapping 48\-item panels\. All three showed a smaller speech–tools separation on average in later layers\. In 11 of 12 collection\-by\-panel comparisons, the categories became less separated while retaining the prespecified direction\.C\.After the pattern recurred across all three panels, Brain Researcher extended the test to four previously unused sound collections\. Three of four showed the same geometry, and the late\-layer separation was again smaller on average \(ΔS=−0\.198\\Delta S=\-0\.198\)\. In all geometry plots, horizontal position shows the late\-minus\-early change in normalized separation \(ΔS\\Delta S\), and vertical position shows late directional alignment \(CC\); the upper\-left quadrant therefore marks smaller separation with retained direction\. A uses fold\-specific references, whereas B and C use the frozen speech–tools reference\.Both episodes converted adaptive searches into frozen follow\-up analyses\. In separately initiated sessions without Brain Researcher, the same coding agent completed substantial analyses, but neither session generated and froze a follow\-up study\. Because these sessions were not matched controls, this contrast is descriptive and does not establish that Brain Researcher caused the transition from an initial result to a frozen follow\-up \(Supplementary Methods S11\.3\.3\)\.
## 2Discussion
Our central contribution is to treat the unit of AI\-assisted research as a governed research trajectory rather than a model output\. The paired benchmarks tested whether models could reach relevant tools and evidence\. The collaborator studies showed how multiverse analysis and explicit constraints change what can be claimed from a completed analysis\. The HCP and TRIBE episodes went one step further: a result from one stage became an input to the next\. Taken together, these evaluations support a view of scientific agents not as systems that generate a final answer, but as infrastructure that augments human judgments by keeping questions, decisions, evidence, and claim states connected as a project evolves\.
This is the sense in which the research episodes were self\-evolving\. In HCP, Brain Researcher searched over candidate workflows for a fixed question; following a selector audit, the researcher designated the frozen selected workflow and carried it into a matched comparison and four additional behavioural outcomes\. In TRIBE, Brain Researcher did not simply promote the contrast with the largest change\. It set aside an unstable lead, proposed a more coherent speech–tools question and, after the researcher froze it, carried that question into newly sampled panels\. In both cases, intermediate evidence changed the next analysis without rewriting the analysis already under test\. The research trajectory evolved, but the evidentiary standard did not\.
This trajectory\-level view also changes the role of verification\. Formal criteria can operate prospectively when they are specified in advance; multiverse analysis can show how a result depends on the enumerated defensible choices\([51](https://arxiv.org/html/2608.19902#bib.bib25);[49](https://arxiv.org/html/2608.19902#bib.bib24);[13](https://arxiv.org/html/2608.19902#bib.bib7);[36](https://arxiv.org/html/2608.19902#bib.bib35);[7](https://arxiv.org/html/2608.19902#bib.bib4)\); and judgments that resist formalization remain with the researcher \(Supplementary Methods S6\.1\)\. Relative to systems evaluated primarily for task execution or output correctness\([52](https://arxiv.org/html/2608.19902#bib.bib26);[30](https://arxiv.org/html/2608.19902#bib.bib15);[5](https://arxiv.org/html/2608.19902#bib.bib2);[27](https://arxiv.org/html/2608.19902#bib.bib45);[24](https://arxiv.org/html/2608.19902#bib.bib46);[53](https://arxiv.org/html/2608.19902#bib.bib47);[10](https://arxiv.org/html/2608.19902#bib.bib33);[11](https://arxiv.org/html/2608.19902#bib.bib34);[3](https://arxiv.org/html/2608.19902#bib.bib48)\), Brain Researcher makes the relationship among the estimator, comparison, evidence base, and claim part of the persistent research record\. The NeuroMark example illustrates the limit of this formal layer: automated review missed sign\-blind scoring, and a human reviewer found the error\. Brain Researcher therefore makes formalizable conditions visible and auditable; it does not make expert inspection unnecessary\.
Once decisions and claim states persist, qualified, negative, and failed results need not be terminal outputs\. They can narrow the next question, retire an unproductive branch, or define a frozen successor analysis\. This is the broader infrastructure implication of self\-evolving research: progress can accumulate across successive episodes instead of restarting from an unstructured prompt each time\. The mechanism is not unconstrained model autonomy, but the combination of an adaptive trajectory with researcher\-defined scope, explicit evidence standards, and durable records of what was tried and why it changed\. Researchers retain authority to define the action space, select and freeze successor questions, interpret the evidence, and decide whether the trajectory should continue\.
This work has several important limitations \(Supplementary Methods S12\)\. Brain Researcher produces auditable evidence but leaves interpretation and writing to the researcher\. Its foundation\-model and retrieval priors reflect the literature, datasets, and instrumented tools, which may favor well\-represented, operationalized questions over negative results, low\-resource populations, and unusual paradigms\([28](https://arxiv.org/html/2608.19902#bib.bib14)\)\. Several episodes rely primarily on same\-dataset multiverse or internal validation; these improve auditability but do not replace independent replication or constitute external confirmation\([8](https://arxiv.org/html/2608.19902#bib.bib5);[38](https://arxiv.org/html/2608.19902#bib.bib18)\)\. Several of these datasets are public \(HCP, OpenNeuro\), so a frontier model may have encountered the associated published findings during training; we cannot rule out memorization, which is a further reason not to treat these results as novel detections\. Evidence grounding was scored by condition\-blind LLM judges \(three frontier models that are also among the seven evaluated, a potential source of self\-preference\)\. A reproducible human audit of 20% of scored results \(272 of roughly 1,360 items\) agreed with 96% of verdicts \(Cohen’sκ=0\.94\\kappa=0\.94\); discrepancies were one\-step severity differences, never reversals between supported and unsupported, and the judges erred strictly \(Supplementary Methods S11\.1\.3\)\. Runtime and researcher effort were not measured\. Review\-layer error was estimated against a 60\-case calibration library \(16 invalid, 5 valid controls, 39 warn\), which produced no false\-accepts \(0 of 16; rule\-of\-three 95% upper bound 19%\) and no false\-blocks \(0 of 5, a loose bound\); this library was assembled after the sign\-direction check identified through the NeuroMark case and is not an independent, field\-scale estimate \(Appendix G; Supplementary Methods S11\.3\.4\)\. The calibration therefore measures internal consistency on canonical scenarios, not how often flawed claims escape review in deployed research workflows\. Independent replication and field\-scale adjudication remain separate tests of scientific validity, which will require labeled real analyses\. Finally, claim records are exportable files, but their value as shared, contestable infrastructure across laboratories remains a future objective\. AI assistance should make the conditions under which results become reproducible knowledge easier to see, test, and share\.
## 3Online Methods
Detailed methods, including the runtime stack, the BR\-KG substrate and sources, the operation registry, execution backends, benchmark scoring contracts, multiverse and validation\-gated search protocols, and all per\-case statistics, are provided in the Supplementary Information \(Supplementary Methods S1–S12; Appendices A–K, with the per\-case episode reports and the item\-level benchmark audit sheet released as extended\-data Appendices L and M; Supplementary Figures\)\.
#### Supplementary information
Supplementary Information accompanies this manuscript\.
#### Funding
JAR is supported by Stanford University Knight\-Hennessy Scholars Program, National Academies of Sciences, Engineering, and Medicine’s Ford Foundation Predoctoral Fellowship, Institute of International Education Quad Fellowship, the National Science Foundation’s Graduate Research Fellowship Program, the Center for Mind, Brain, Computation and Technology, and the Wu Tsai Neurosciences Institute\. V\.D\.C\. received support from NSF 2112455 and NIH R01MH123610\. A\.d\.l\.V\., R\.P\. and J\.K\. were supported by the National Institute of Mental Health under award R01MH096906\. Z\.C\. and R\.P\. received cloud\-computing credits through the 2025 HAI\-Google Cloud Credits Grant Program to support Brain Researcher API development and computation\.
#### Competing interests
S\.K\. reports part\-time employment with Meta, which began recently and after most of the work reported here\. The other authors declare no competing interests\.
#### Ethics, consent and materials availability
Not applicable: this work analyzed only previously collected, publicly available or collaborator\-provided de\-identified neuroimaging data under their original ethics approvals and consents, and generated no new human\- or animal\-subjects data or materials\.
#### Data availability
BR\-KG is archived at Zenodo \([https://doi\.org/10\.5281/zenodo\.21966011](https://doi.org/10.5281/zenodo.21966011)\) and linked from the public project site \([https://brain\-researcher\.com/](https://brain-researcher.com/)\)\. The release includes graph snapshots, node and edge schemas, provenance fields, registry links, benchmark manifests, scoring tables, aggregate outputs, figure source data, run\-bundle schemas, a worked auditable claim\-record example \(an exported claim card with its evidence verdicts, on public Neurosynth data\), and deployment notes\. Users can access the public MCP interface and released Brain Researcher skills from the project site, which describes how users can suggest additions or corrections to BR\-KG\. Source neuroimaging datasets remain under their original terms: public resources are cited and linked in Supplementary Methods S4 and Appendix C, and controlled\-access, collaborator\-provided, or license\-restricted human\-subject data are not redistributed\. Artifact and provenance records are described in Supplementary Methods S7\.4 and Appendix F; benchmark records in Supplementary Methods S11\.1 and Appendix J\. To keep the Supplementary Information self\-contained, the full audit ledgers it condenses are released in the same archival repository as an extended\-data package: the complete BR\-KG, evidence\-bundle, dataset, tool\-registry, constraint, execution–provenance, and memory data cards \(Appendices A–F and H\), the full per\-rule review registry \(Appendix G9\.1–G9\.4 and G9\.6\), the automatically generated per\-case episode reports \(Appendix L\), the current HCP and TRIBE research\-line reports and their supporting run bundles, and the item\-level benchmark human\-audit sheet \(Appendix M\); a crosswalk maps each condensed Supplementary section to its archived file\.
#### Code availability
The Brain Researcher system \(Python package, CLI, agent runtime, MCP server with versioned tool contracts, orchestrator, web UI, and deployment recipes\) is available under the MIT license at[https://github\.com/brain\-researcher/brain\-researcher\-public](https://github.com/brain-researcher/brain-researcher-public), with the companion agent layer \(skills, agent templates, MCP adapters, and AutoResearch evaluation rubrics\) at[https://github\.com/brain\-researcher/brain\-researcher\-agent\-kit](https://github.com/brain-researcher/brain-researcher-agent-kit); both are linked from the project site \([https://brain\-researcher\.com/](https://brain-researcher.com/)\), and the v0\.3\.0 release is archived at Zenodo \([https://doi\.org/10\.5281/zenodo\.21966011](https://doi.org/10.5281/zenodo.21966011)\)\. Analysis code and per\-specification outputs for the NeuroMark collaborator case are available at[https://github\.com/XinhuiLi/BR\-NeuroMark](https://github.com/XinhuiLi/BR-NeuroMark)\.
#### Author contributions
Z\.C\. and R\.P\. initiated and conceived the project\. Z\.C\. designed and implemented the Brain Researcher system, ran the experiments and analyses, generated the main results, and drafted the manuscript\. R\.P\. supervised the project and contributed to conceptual framing, study design, hands\-on system testing, evaluation feedback, interpretation, and manuscript revision\. J\.H\.Z\. provided early supervision and initial computational resources for the project\. N\.L\. gathered background information, including dataset lists and literature\-review materials, helped design the benchmark questions, evaluated system outputs, and tested performance for the quantitative benchmark section\. X\.L\. and V\.D\.C\. designed and ran the NeuroMark schizophrenia functional\-network\-connectivity case\. J\.R\. and R\.P\. designed and ran the cocaine\-use\-disorder connectivity case\. H\.W\. designed and ran the cross\-cultural social cognition case\. C\.K\., J\.K\., and A\.d\.l\.V\. provided feedback on knowledge\-graph design and contributed data and design requirements for the knowledge\-graph and source\-integration components\. S\.B\. provided feedback on agent design, MCP infrastructure, backend integration, and execution design\. J\.M\. provided feedback on the scientific\-review layer\. S\.D\. contributed suggestions and ideas on the agent harness and validation\-gated research design, including the bounded\-validation framing; S\.K\. provided feedback on the agent harness\. C\.J\. contributed to system testing\. J\.W\.B\. provided feedback on the manuscript\. All authors reviewed and approved the manuscript\.
## References
- Alibaba Cloud \(2026\)Alibaba CloudQwen3\.6\-Plus: towards real world agents\.Note:Accessed 21 May 2026External Links:[Link](https://www.alibabacloud.com/blog/603005)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.
- Angeles\-Valdezet al\.\(2022\)D\. Angeles\-Valdez, J\. Rasgado\-Toledo, V\. Issa\-Garcia, T\. Balducci, V\. Villicaña, A\. Valencia, J\. J\. Gonzalez\-Olvera, E\. Reyes\-Zamorano, E\. A\. Garza\-Villarreal,et al\.The Mexican magnetic resonance imaging dataset of patients with cocaine use disorder: SUDMEX CONN\.Scientific Data9\(1\),pp\. 133\.External Links:[Document](https://dx.doi.org/10.1038/s41597-022-01251-3)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Anthropic \(2026a\)AnthropicClaude Science, an AI workbench for scientists, is now available\.Note:[https://www\.anthropic\.com/news/claude\-science\-ai\-workbench](https://www.anthropic.com/news/claude-science-ai-workbench)Anthropic news announcement, 30 June 2026Anthropic \(2026\)\. Claude Science, an AI workbench for scientists, is now available\. https://www\.anthropic\.com/news/claude\-science\-ai\-workbenchCited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Anthropic \(2026b\)AnthropicIntroducing Claude Opus 4\.8\.Note:Accessed 5 June 2026External Links:[Link](https://www.anthropic.com/news/claude-opus-4-8)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.
- Boikoet al\.\(2023\)D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. GomesAutonomous chemical research with large language models\.Nature624,pp\. 570–578\.Note:Boiko, D\. A\., MacKnight, R\., Kline, B\., & Gomes, G\. \(2023\)\. Autonomous chemical research with large language models\. Nature, 624, 570–578\. https://doi\.org/10\.1038/s41586\-023\-06792\-0External Links:[Link](https://doi.org/10.1038/s41586-023-06792-0),[Document](https://dx.doi.org/10.1038/s41586-023-06792-0)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Botvinik\-Nezeret al\.\(2020\)R\. Botvinik\-Nezer, F\. Holzmeister, C\. F\. Camerer,et al\.Variability in the analysis of a single neuroimaging dataset by many teams\.Nature582,pp\. 84–88\.Note:Botvinik\-Nezer, R\., Holzmeister, F\., Camerer, C\. F\., et al\. \(2020\)\. Variability in the analysis of a single neuroimaging dataset by many teams\. Nature, 582, 84–88\. https://doi\.org/10\.1038/s41586\-020\-2314\-9External Links:[Link](https://doi.org/10.1038/s41586-020-2314-9),[Document](https://dx.doi.org/10.1038/s41586-020-2314-9)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Burkhardt and Giessing \(2026\)M\. Burkhardt and C\. GiessingThe Comet Toolbox: Improving robustness in network neuroscience through multiverse analysis\.Imaging Neuroscience4,pp\. IMAG\.a\.1122\.Note:Burkhardt, M\., & Giessing, C\. \(2026\)\. The Comet Toolbox: Improving robustness in network neuroscience through multiverse analysis\. Imaging Neuroscience, 4, IMAG\.a\.1122\. https://doi\.org/10\.1162/IMAG\.a\.1122External Links:[Link](https://doi.org/10.1162/IMAG.a.1122),[Document](https://dx.doi.org/10.1162/IMAG.a.1122)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Buttonet al\.\(2013\)K\. S\. Button, J\. P\. A\. Ioannidis, C\. Mokrysz, B\. A\. Nosek, J\. Flint, E\. S\. J\. Robinson, and M\. R\. MunafoPower failure: Why small sample size undermines the reliability of neuroscience\.Nature Reviews Neuroscience14,pp\. 365–376\.Note:Button, K\. S\., Ioannidis, J\. P\. A\., Mokrysz, C\., Nosek, B\. A\., Flint, J\., Robinson, E\. S\. J\., & Munafo, M\. R\. \(2013\)\. Power failure: Why small sample size undermines the reliability of neuroscience\. Nature Reviews Neuroscience, 14, 365–376\. https://doi\.org/10\.1038/nrn3475External Links:[Link](https://doi.org/10.1038/nrn3475),[Document](https://dx.doi.org/10.1038/nrn3475)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p5.1)\.
- Carp \(2012\)J\. CarpThe secret lives of experiments: Methods reporting in the fMRI literature\.NeuroImage63\(1\),pp\. 289–300\.Note:Carp, J\. \(2012\)\. The secret lives of experiments: Methods reporting in the fMRI literature\. NeuroImage, 63\(1\), 289–300\. https://doi\.org/10\.1016/j\.neuroimage\.2012\.07\.004External Links:[Link](https://doi.org/10.1016/j.neuroimage.2012.07.004),[Document](https://dx.doi.org/10.1016/j.neuroimage.2012.07.004)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Chanet al\.\(2024\)J\. S\. Chan, N\. Chowdhury, O\. Jaffe, J\. Aung, D\. Sherburn, E\. Mays, G\. Starace,et al\.MLE\-bench: Evaluating machine learning agents on machine learning engineering\.Note:Publication Title: arXivExternal Links:[Link](https://doi.org/10.48550/arXiv.2410.07095),[Document](https://dx.doi.org/10.48550/arXiv.2410.07095)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Chenet al\.\(2025\)Z\. Chen, S\. Chen, Y\. Ning, Q\. Zhang, B\. Wang, B\. Yu, Y\. Li,et al\.ScienceAgentBench: Toward rigorous assessment of language agents for data\-driven scientific discovery\.Note:ICLR 2025; Publication Title: arXivExternal Links:[Link](https://doi.org/10.48550/arXiv.2410.05080),[Document](https://dx.doi.org/10.48550/arXiv.2410.05080)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Community \(2024\)B\. CommunityBIDS Stats Models Specification\.Note:BIDS Community\. \(2024\)\. BIDS Stats Models Specification\. https://bids\-standard\.github\.io/stats\-models/External Links:[Link](https://bids-standard.github.io/stats-models/)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- Dafflonet al\.\(2022\)J\. Dafflon, P\. F\. Costa, F\. Vasa,et al\.A guided multiverse study of neuroimaging analyses\.Nature Communications13,pp\. 3758\.Note:Dafflon, J\., da Costa, P\. F\., Vasa, F\., et al\. \(2022\)\. A guided multiverse study of neuroimaging analyses\. Nature Communications, 13, 3758\. https://doi\.org/10\.1038/s41467\-022\-31347\-8External Links:[Link](https://doi.org/10.1038/s41467-022-31347-8),[Document](https://dx.doi.org/10.1038/s41467-022-31347-8)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- DeepSeek \(2026\)DeepSeekDeepSeek V4 preview release\.Note:Accessed 21 May 2026External Links:[Link](https://api-docs.deepseek.com/news/news260424)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.
- Dockeset al\.\(2020\)J\. Dockes, R\. A\. Poldrack, R\. Primet, H\. Gozukan, T\. Yarkoni, F\. Suchanek, B\. Thirion, and G\. VaroquauxNeuroQuery, comprehensive meta\-analysis of human brain mapping\.eLife9,pp\. e53385\.Note:Dockes, J\., Poldrack, R\. A\., Primet, R\., Gozukan, H\., Yarkoni, T\., Suchanek, F\., Thirion, B\., & Varoquaux, G\. \(2020\)\. NeuroQuery, comprehensive meta\-analysis of human brain mapping\. eLife, 9, e53385\. https://doi\.org/10\.7554/eLife\.53385External Links:[Link](https://doi.org/10.7554/eLife.53385),[Document](https://dx.doi.org/10.7554/eLife.53385)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Duet al\.\(2020\)Y\. Du, Z\. Fu, J\. Sui, S\. Gao, Y\. Xing, D\. Lin, M\. Salman, A\. Abrol, M\. A\. Rahaman, J\. Chen, L\. E\. Hong, P\. Kochunov, E\. A\. Osuch, and V\. D\. CalhounNeuroMark: An automated and adaptive ICA\-based pipeline to identify reproducible fMRI markers of brain disorders\.NeuroImage: Clinical28,pp\. 102375\.External Links:[Document](https://dx.doi.org/10.1016/j.nicl.2020.102375)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p2.1)\.
- d’Ascoliet al\.\(2026\)S\. d’Ascoli, J\. Rapin, Y\. Benchetrit, T\. Brooks, K\. Begany, J\. Raugel, H\. Banville, and J\. KingA foundation model of vision, audition, and language for in\-silico neuroscience\.Note:Publication Title: arXivExternal Links:[Link](https://doi.org/10.48550/arXiv.2605.04326),[Document](https://dx.doi.org/10.48550/arXiv.2605.04326)Cited by:[§1\.4](https://arxiv.org/html/2608.19902#S1.SS4.p3.1)\.
- Eickhoffet al\.\(2009\)S\. B\. Eickhoff, A\. R\. Laird, C\. Grefkes, L\. E\. Wang, K\. Zilles, and P\. T\. FoxCoordinate\-based activation likelihood estimation meta\-analysis of neuroimaging data: A random\-effects approach based on empirical estimates of spatial uncertainty\.Human Brain Mapping30\(9\),pp\. 2907–2926\.Note:Eickhoff, S\. B\., Laird, A\. R\., Grefkes, C\., Wang, L\. E\., Zilles, K\., & Fox, P\. T\. \(2009\)\. Coordinate\-based activation likelihood estimation meta\-analysis of neuroimaging data: A random\-effects approach based on empirical estimates of spatial uncertainty\. Human Brain Mapping, 30\(9\), 2907–2926\. https://doi\.org/10\.1002/hbm\.20718External Links:[Link](https://doi.org/10.1002/hbm.20718),[Document](https://dx.doi.org/10.1002/hbm.20718)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Eickhoffet al\.\(2016\)S\. B\. Eickhoff, T\. E\. Nichols, A\. R\. Laird,et al\.Behavior, sensitivity, and power of activation likelihood estimation characterized by massive empirical simulation\.NeuroImage137,pp\. 70–85\.Note:Eickhoff, S\. B\., Nichols, T\. E\., Laird, A\. R\., et al\. \(2016\)\. Behavior, sensitivity, and power of activation likelihood estimation characterized by massive empirical simulation\. NeuroImage, 137, 70–85\. https://doi\.org/10\.1016/j\.neuroimage\.2016\.04\.072External Links:[Link](https://doi.org/10.1016/j.neuroimage.2016.04.072),[Document](https://dx.doi.org/10.1016/j.neuroimage.2016.04.072)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Eklundet al\.\(2016\)A\. Eklund, T\. E\. Nichols, and H\. KnutssonCluster failure: Why fMRI inferences for spatial extent have inflated false\-positive rates\.Proceedings of the National Academy of Sciences113\(28\),pp\. 7900–7905\.Note:Eklund, A\., Nichols, T\. E\., & Knutsson, H\. \(2016\)\. Cluster failure: Why fMRI inferences for spatial extent have inflated false\-positive rates\. Proceedings of the National Academy of Sciences, 113\(28\), 7900–7905\. https://doi\.org/10\.1073/pnas\.1602413113External Links:[Link](https://doi.org/10.1073/pnas.1602413113),[Document](https://dx.doi.org/10.1073/pnas.1602413113)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Estebanet al\.\(2019\)O\. Esteban, C\. J\. Markiewicz, R\. W\. Blair,et al\.fMRIPrep: A robust preprocessing pipeline for functional MRI\.Nature Methods16,pp\. 111–116\.Note:Esteban, O\., Markiewicz, C\. J\., Blair, R\. W\., et al\. \(2019\)\. fMRIPrep: A robust preprocessing pipeline for functional MRI\. Nature Methods, 16, 111–116\. https://doi\.org/10\.1038/s41592\-018\-0235\-4External Links:[Link](https://doi.org/10.1038/s41592-018-0235-4),[Document](https://dx.doi.org/10.1038/s41592-018-0235-4)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- Gaoet al\.\(2023\)L\. Gao, J\. Schulman, and J\. HiltonScaling laws for reward model overoptimization\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 10835–10866\.External Links:[Link](https://proceedings.mlr.press/v202/gao23h.html)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p4.1)\.
- Garza\-Villarrealet al\.\(2026\)E\. A\. Garza\-Villarreal, J\. J\. Gonzalez Olvera, T\. Balducci, D\. Angeles Valdez, A\. Valencia, and J\. RasgadoSUDMEX\_CONN: The Mexican dataset of cocaine use disorder patients\.OpenNeuro\.Note:OpenNeuro datasetExternal Links:[Document](https://dx.doi.org/10.18112/openneuro.ds003346.v1.1.3),[Link](https://doi.org/10.18112/openneuro.ds003346.v1.1.3)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Ghareebet al\.\(2026\)A\. E\. Ghareeb, B\. Chang, L\. Mitchener, A\. Yiu, C\. J\. Szostkiewicz, D\. Shved, G\. J\. Gyimesi, J\. M\. Laurent, S\. M\. Wright, M\. T\. Razzak, A\. D\. White, S\. C\. Finnemann, M\. M\. Hinks, and S\. G\. RodriquesA multi\-agent system for automating scientific discovery\.Nature\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10652-y),[Link](https://www.nature.com/articles/s41586-026-10652-y)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Google DeepMind \(2026\)Google DeepMindGemini 3\.1 Pro: model card\.Note:Accessed 21 May 2026External Links:[Link](https://deepmind.google/models/model-cards/gemini-3-1-pro/)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.
- Gorgolewskiet al\.\(2011\)K\. Gorgolewski, C\. D\. Burns, C\. Madison, D\. Clark, Y\. O\. Halchenko, M\. L\. Waskom, and S\. S\. GhoshNipype: A flexible, lightweight and extensible neuroimaging data processing framework in Python\.Frontiers in Neuroinformatics5,pp\. 13\.Note:Gorgolewski, K\., Burns, C\. D\., Madison, C\., Clark, D\., Halchenko, Y\. O\., Waskom, M\. L\., & Ghosh, S\. S\. \(2011\)\. Nipype: A flexible, lightweight and extensible neuroimaging data processing framework in Python\. Frontiers in Neuroinformatics, 5, 13\. https://doi\.org/10\.3389/fninf\.2011\.00013External Links:[Link](https://doi.org/10.3389/fninf.2011.00013),[Document](https://dx.doi.org/10.3389/fninf.2011.00013)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- Gottweiset al\.\(2026\)J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, P\. Sirkovic, A\. Myaskovsky, G\. Glowaty, F\. Weissenberger, A\. Orlandi, D\. Popovici, A\. Palepu, K\. Rong, R\. Tanno, K\. Saab, F\. Zhang, J\. Blum, A\. Carroll, K\. Kulkarni, N\. Tomašev, D\. Zverinski, I\. Rendulic, E\. Vedadi, F\. Hasler, L\. Rimanic, M\. Boia, I\. Budiselic, B\. Feinstein, M\. Bellaiche, T\. Sheffer, J\. Freyberg, J\. Ratcliff, O\. Bertolli, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penadés, G\. Peltz, Y\. Matias, J\. Manyika, D\. Hassabis, Y\. Xu, P\. Kohli, A\. Pawlosky, A\. Karthikesalingam, and V\. NatarajanAccelerating scientific discovery with Co\-Scientist\.Nature\.External Links:[Document](https://dx.doi.org/10.1038/s41586-026-10644-y),[Link](https://www.nature.com/articles/s41586-026-10644-y)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Haoet al\.\(2026\)Q\. Hao, F\. Xu, Y\. Li,et al\.Artificial intelligence tools expand scientists’ impact but contract science’s focus\.Nature649,pp\. 1237–1243\.Note:Hao, Q\., Xu, F\., Li, Y\., et al\. \(2026\)\. Artificial intelligence tools expand scientists’ impact but contract science’s focus\. Nature, 649, 1237–1243\. https://doi\.org/10\.1038/s41586\-025\-09922\-yExternal Links:[Link](https://doi.org/10.1038/s41586-025-09922-y),[Document](https://dx.doi.org/10.1038/s41586-025-09922-y)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p5.1)\.
- Hasan and Biswas \(2026\)A\. A\. Hasan and S\. BiswasWhat breaks when LLMs code? characterizing operational safety failures of agentic code assistants\.External Links:2605\.30777,[Document](https://dx.doi.org/10.48550/arXiv.2605.30777),[Link](https://arxiv.org/abs/2605.30777)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p4.1)\.
- Huanget al\.\(2025\)K\. Huang, S\. Zhang, H\. Wang, Y\. Qu, Y\. Lu, Y\. Roohani,et al\.Biomni: A general\-purpose biomedical AI agent\.Note:Publication Title: bioRxivHuang, K\., Zhang, S\., Wang, H\., Qu, Y\., Lu, Y\., Roohani, Y\., et al\. \(2025\)\. Biomni: A general\-purpose biomedical AI agent\. bioRxiv\. https://doi\.org/10\.1101/2025\.05\.30\.656746External Links:[Link](https://doi.org/10.1101/2025.05.30.656746),[Document](https://dx.doi.org/10.1101/2025.05.30.656746)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Irajiet al\.\(2023\)A\. Iraji, Z\. Fu, A\. Faghiri, M\. Duda, J\. Chen, S\. Rachakonda, T\. DeRamus, P\. Kochunov, B\. M\. Adhikari, A\. Belger, J\. M\. Ford, D\. H\. Mathalon, G\. D\. Pearlson, S\. G\. Potkin, A\. Preda, J\. A\. Turner, T\. G\. M\. van Erp, J\. R\. Bustillo, K\. Yang, K\. Ishizuka, A\. Faria, A\. Sawa, K\. Hutchison, E\. A\. Osuch, J\. Theberge, C\. Abbott, B\. A\. Mueller, D\. Zhi, C\. Zhuo, S\. Liu, Y\. Xu, M\. Salman, J\. Liu, Y\. Du, J\. Sui, T\. Adali, and V\. D\. CalhounIdentifying canonical and replicable multi\-scale intrinsic connectivity networks in 100k\+ resting\-state fmri datasets\.Human Brain Mapping44\(17\),pp\. 5729–5748\.External Links:[Document](https://dx.doi.org/https%3A//doi.org/10.1002/hbm.26472),[Link](https://onlinelibrary.wiley.com/doi/abs/10.1002/hbm.26472),https://onlinelibrary\.wiley\.com/doi/pdf/10\.1002/hbm\.26472Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p2.1)\.
- Johnet al\.\(2012\)L\. K\. John, G\. Loewenstein, and D\. PrelecMeasuring the prevalence of questionable research practices with incentives for truth telling\.Psychological Science23\(5\),pp\. 524–532\.External Links:[Document](https://dx.doi.org/10.1177/0956797611430953)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p4.1)\.
- Kaddouret al\.\(2026\)J\. Kaddour, S\. Patel, G\. Dovonon, L\. Richter, P\. Minervini, and M\. J\. KusnerAgentic uncertainty reveals agentic overconfidence\.External Links:2602\.06948,[Document](https://dx.doi.org/10.48550/arXiv.2602.06948),[Link](https://arxiv.org/abs/2602.06948)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p4.1)\.
- Keatoret al\.\(2016\)D\. B\. Keator, T\. G\.M\. van Erp, J\. A\. Turner, G\. H\. Glover, B\. A\. Mueller, T\. T\. Liu, J\. T\. Voyvodic, J\. Rasmussen, V\. D\. Calhoun, H\. J\. Lee, A\. W\. Toga, S\. McEwen, J\. M\. Ford, D\. H\. Mathalon, M\. Diaz, D\. S\. O’Leary, H\. Jeremy Bockholt, S\. Gadde, A\. Preda, C\. G\. Wible, H\. S\. Stern, A\. Belger, G\. McCarthy, B\. Ozyurt, and S\. G\. PotkinThe function biomedical informatics research network data repository\.NeuroImage124,pp\. 1074–1079\.External Links:ISSN 1053\-8119,[Document](https://dx.doi.org/10.1016/j.neuroimage.2015.09.003),[Link](https://www.sciencedirect.com/science/article/pii/S1053811915007995)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p2.1)\.
- Lefort\-Besnardet al\.\(2025\)J\. Lefort\-Besnard, T\. E\. Nichols, and C\. MaumetStatistical inference for neuroimaging multiverse analyses with the same\-data meta\-analysis\.Imaging Neuroscience\.Note:Lefort\-Besnard, J\., Nichols, T\. E\., & Maumet, C\. \(2025\)\. Statistical inference for neuroimaging multiverse analyses with the same\-data meta\-analysis\. Imaging Neuroscience\. https://doi\.org/10\.1162/imag\_a\_00513External Links:[Link](https://doi.org/10.1162/imag_a_00513),[Document](https://dx.doi.org/10.1162/imag%5Fa%5F00513)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Liet al\.\(2024\)X\. Li, N\. Bianchini Esper, L\. Ai, S\. Giavasis, H\. Jin, E\. Feczko, T\. Xu,et al\.Moving beyond processing\- and analysis\-related variation in resting\-state functional brain imaging\.Nature Human Behaviour8,pp\. 2003–2017\.External Links:[Link](https://doi.org/10.1038/s41562-024-01942-4),[Document](https://dx.doi.org/10.1038/s41562-024-01942-4)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Liuet al\.\(2025\)Z\. Liu, A\. I\. Luppi, J\. Y\. Hansen, Y\. E\. Tian, A\. Zalesky, B\. T\. T\. Yeo, B\. D\. Fulcher, and B\. MisicBenchmarking methods for mapping functional connectivity in the brain\.Nature Methods22\(7\),pp\. 1593–1602\.External Links:[Link](https://doi.org/10.1038/s41592-025-02704-4),[Document](https://dx.doi.org/10.1038/s41592-025-02704-4)Cited by:[§1\.4](https://arxiv.org/html/2608.19902#S1.SS4.p2.1)\.
- Mareket al\.\(2022\)S\. Marek, B\. Tervo\-Clemmens, F\. J\. Calabro,et al\.Reproducible brain\-wide association studies require thousands of individuals\.Nature603,pp\. 654–660\.Note:Marek, S\., Tervo\-Clemmens, B\., Calabro, F\. J\., et al\. \(2022\)\. Reproducible brain\-wide association studies require thousands of individuals\. Nature, 603, 654–660\. https://doi\.org/10\.1038/s41586\-022\-04492\-9External Links:[Link](https://doi.org/10.1038/s41586-022-04492-9),[Document](https://dx.doi.org/10.1038/s41586-022-04492-9)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p5.1)\.
- Markiewiczet al\.\(2021\)C\. J\. Markiewicz, K\. J\. Gorgolewski, F\. Feingold,et al\.The OpenNeuro resource for sharing of neuroscience data\.eLife10,pp\. e71774\.Note:Markiewicz, C\. J\., Gorgolewski, K\. J\., Feingold, F\., et al\. \(2021\)\. The OpenNeuro resource for sharing of neuroscience data\. eLife, 10, e71774\. https://doi\.org/10\.7554/eLife\.71774External Links:[Link](https://doi.org/10.7554/eLife.71774),[Document](https://dx.doi.org/10.7554/eLife.71774)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- Markiewiczet al\.\(2022\)C\. J\. Markiewicz, A\. De La Vega, A\. Wagner, Y\. O\. Halchenko, K\. Finc, R\. Ciric, M\. Goncalves, D\. M\. Nielson, J\. D\. Kent, J\. A\. Lee, S\. Bansal, R\. A\. Poldrack, and K\. J\. GorgolewskiPoldracklab/fitlins: 0\.11\.0\.Note:ZenodoVersion 0\.11\.0External Links:[Link](https://doi.org/10.5281/zenodo.7217447),[Document](https://dx.doi.org/10.5281/zenodo.7217447)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- Mitcheneret al\.\(2025\)L\. Mitchener, A\. Yiu, B\. Chang,et al\.Kosmos: An AI Scientist for Autonomous Discovery\.Note:Publication Title: arXivMitchener, L\., Yiu, A\., Chang, B\., et al\. \(2025\)\. Kosmos: An AI Scientist for Autonomous Discovery\. arXiv\. https://doi\.org/10\.48550/arXiv\.2511\.02824External Links:[Link](https://doi.org/10.48550/arXiv.2511.02824),[Document](https://dx.doi.org/10.48550/arXiv.2511.02824)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Moonshot AI \(2026\)Moonshot AIKimi K2\.5\.Note:Accessed 21 May 2026External Links:[Link](https://platform.kimi.ai/docs/guide/kimi-k2-5-quickstart)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.
- OpenAI \(2026\)OpenAIIntroducing GPT\-5\.5\.Note:Accessed 21 May 2026External Links:[Link](https://openai.com/index/introducing-gpt-5-5/)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.
- OpenNeuro \(2026\)OpenNeuroOpenNeuro Vocabulary \(ONVOC\)\.Note:BioPortal, National Center for Biomedical OntologyAccessed 2026External Links:[Link](https://bioportal.bioontology.org/ontologies/ONVOC)Cited by:[Figure 2](https://arxiv.org/html/2608.19902#S1.F2)\.
- Poldracket al\.\(2017\)R\. A\. Poldrack, C\. I\. Baker, J\. Durnez,et al\.Scanning the horizon: Towards transparent and reproducible neuroimaging research\.Nature Reviews Neuroscience18,pp\. 115–126\.Note:Poldrack, R\. A\., Baker, C\. I\., Durnez, J\., et al\. \(2017\)\. Scanning the horizon: Towards transparent and reproducible neuroimaging research\. Nature Reviews Neuroscience, 18, 115–126\. https://doi\.org/10\.1038/nrn\.2016\.167External Links:[Link](https://doi.org/10.1038/nrn.2016.167),[Document](https://dx.doi.org/10.1038/nrn.2016.167)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Poldracket al\.\(2024\)R\. A\. Poldrack, C\. J\. Markiewicz, S\. Appelhoff,et al\.The past, present, and future of the Brain Imaging Data Structure \(BIDS\)\.Imaging Neuroscience2,pp\. 1–19\.Note:Poldrack, R\. A\., Markiewicz, C\. J\., Appelhoff, S\., et al\. \(2024\)\. The past, present, and future of the Brain Imaging Data Structure \(BIDS\)\. Imaging Neuroscience, 2, 1–19\. https://doi\.org/10\.1162/imag\_a\_00103External Links:[Link](https://doi.org/10.1162/imag_a_00103),[Document](https://dx.doi.org/10.1162/imag%5Fa%5F00103)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p3.1)\.
- Saloet al\.\(2023\)T\. Salo, T\. Yarkoni, T\. E\. Nichols, J\.\-B\. Poline, M\. Bilgel, K\. L\. Bottenhorn,et al\.NiMARE: Neuroimaging Meta\-Analysis Research Environment\.Aperture Neuro3,pp\. 1–32\.Note:Salo, T\., Yarkoni, T\., Nichols, T\. E\., Poline, J\.\-B\., Bilgel, M\., Bottenhorn, K\. L\., et al\. \(2023\)\. NiMARE: Neuroimaging Meta\-Analysis Research Environment\. Aperture Neuro, 3, 1–32\. https://doi\.org/10\.52294/001c\.87681External Links:[Link](https://doi.org/10.52294/001c.87681),[Document](https://dx.doi.org/10.52294/001c.87681)Cited by:[§1\.3](https://arxiv.org/html/2608.19902#S1.SS3.p5.1)\.
- Simmonset al\.\(2011\)J\. P\. Simmons, L\. D\. Nelson, and U\. SimonsohnFalse\-positive psychology: undisclosed flexibility in data collection and analysis allows presenting anything as significant\.Psychological Science22\(11\),pp\. 1359–1366\.External Links:[Document](https://dx.doi.org/10.1177/0956797611417632)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p4.1)\.
- Simonsohnet al\.\(2020\)U\. Simonsohn, J\. P\. Simmons, and L\. D\. NelsonSpecification curve analysis\.Nature Human Behaviour4,pp\. 1208–1214\.Note:Simonsohn, U\., Simmons, J\. P\., & Nelson, L\. D\. \(2020\)\. Specification curve analysis\. Nature Human Behaviour, 4, 1208–1214\. https://doi\.org/10\.1038/s41562\-020\-0912\-zExternal Links:[Link](https://doi.org/10.1038/s41562-020-0912-z),[Document](https://dx.doi.org/10.1038/s41562-020-0912-z)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Skalseet al\.\(2022\)J\. Skalse, N\. Howe, D\. Krasheninnikov, and D\. KruegerDefining and characterizing reward gaming\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 9460–9471\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2022/hash/3d719fee332caa23d5038b8a90e81796-Abstract-Conference.html)Cited by:[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p4.1)\.
- Steegenet al\.\(2016\)S\. Steegen, F\. Tuerlinckx, A\. Gelman, and W\. VanpaemelIncreasing transparency through a multiverse analysis\.Perspectives on Psychological Science11\(5\),pp\. 702–712\.Note:Steegen, S\., Tuerlinckx, F\., Gelman, A\., & Vanpaemel, W\. \(2016\)\. Increasing transparency through a multiverse analysis\. Perspectives on Psychological Science, 11\(5\), 702–712\. https://doi\.org/10\.1177/1745691616658637External Links:[Link](https://doi.org/10.1177/1745691616658637),[Document](https://dx.doi.org/10.1177/1745691616658637)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p5.1)\.
- Swansonet al\.\(2025\)K\. Swanson, W\. Wu, N\. L\. Bulaong, J\. E\. Pak, and J\. ZouThe Virtual Lab of AI agents designs new SARS\-CoV\-2 nanobodies\.Nature646,pp\. 716–723\.Note:Swanson, K\., Wu, W\., Bulaong, N\. L\., Pak, J\. E\., & Zou, J\. \(2025\)\. The Virtual Lab of AI agents designs new SARS\-CoV\-2 nanobodies\. Nature, 646, 716–723\. https://doi\.org/10\.1038/s41586\-025\-09442\-9External Links:[Link](https://doi.org/10.1038/s41586-025-09442-9),[Document](https://dx.doi.org/10.1038/s41586-025-09442-9)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1),[Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis](https://arxiv.org/html/2608.19902#p2.1)\.
- Wanget al\.\(2026\)C\. Wang, Z\. He, Z\. Peng, S\. Liu, Y\. Hu, C\. Yang, L\. He, L\. Sun, X\. Li, and Y\. YuanNeuroClaw Technical Report\.Note:Publication Title: arXiv \(2604\.24696\)Wang, C\., He, Z\., Peng, Z\., Liu, S\., Hu, Y\., Yang, C\., He, L\., Sun, L\., Li, X\., & Yuan, Y\. \(2026\)\. NeuroClaw Technical Report\. arXiv:2604\.24696\. Closed\-loop agentic AI for executable and reproducible neuroimaging research\.External Links:[Link](https://arxiv.org/abs/2604.24696),[Document](https://dx.doi.org/10.48550/arXiv.2604.24696)Cited by:[§2](https://arxiv.org/html/2608.19902#S2.p3.1)\.
- Z\.AI \(2026\)Z\.AIGLM\-5\.1\.Note:Accessed 21 May 2026External Links:[Link](https://docs.z.ai/guides/llm/glm-5.1)Cited by:[§1\.2](https://arxiv.org/html/2608.19902#S1.SS2.p1.1)\.Similar Articles
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse
This paper benchmarks agentic AI systems on the task of loading, understanding, and reformatting fragmented neuroscience data, finding that while agents perform well on subtasks, they rarely achieve fully error-free end-to-end solutions and human oversight remains necessary.
Ressearch AI
Ressearch AI is an AI-powered workspace aimed at enabling reproducible scientific research.
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
This survey examines the emerging field of AI-powered research automation (AutoResearch), analyzing how AI systems are moving from isolated task assistance to full workflow-level scientific discovery. It defines a spectrum from human-steered 'Vibe Research' to AI-led systems, and proposes five evaluation dimensions for scientific credibility.
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
A survey paper examining the transition of AI from task-specific assistants to workflow-level research automators, defining AutoResearch as the spectrum of AI-powered scientific workflow automation and analyzing challenges in autonomy, reproducibility, and accountability.
Most “agentic AI” conversations feel too abstract. Here is how my agentic research system looks like
The author shares a practical breakdown of an agentic research system they built to identify and evaluate AI use cases within companies. The system uses six agents for discovery, evaluation, and context extraction, emphasizing human-in-the-loop decision-making over full autonomy.