Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods

arXiv cs.CL Papers

Summary

This study examines how LLMs suggest research methods (datasets, models, metrics) when prompted only with a research question, finding that LLMs exhibit a strong provider bias and propose a much narrower range of methods compared to actual papers, potentially narrowing researchers' methodological search space.

arXiv:2606.26130v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear. Here, we prompt GPT-5.1, Gemini 3 Pro, and DeepSeek-V3.2 with an LLM-extracted research question from each of 1,000 recent arXiv computer-science papers and compare the resulting methodology suggestions against a paper-derived experimental inventory. Since we provide only the research question, the differences we measure reflect initial suggestions and not how optimal those suggestions are. We extract structured method features from both sources, map them into a shared taxonomy, and quantify divergence across multiple taxonomy dimensions including model provider, dataset task type, and evaluation metric type. The strongest imbalance appears in provider choice, with Jensen-Shannon divergence about 3-5x larger than any other taxonomy dimension. Other/Academic single-occurrence models are underrepresented by 23-24 percentage points, while reused academic/community models are slightly overrepresented (4-6pp). LLMs also suggest a much narrower range of methods overall: the effective number of model entities contracts from 1,232 to 59-96, and inter-LLM rank correlations (0.55-0.68) generally exceed LLM-to-paper correlations (0.33-0.56), so the distortions are largely shared across models. Popularity baselines, BM25 retrieval calibration, and paper-level similarity tests confirm that the outputs are query-specific responses, but filtered through a narrower set of options. Researchers who rely on LLM suggestions without cross-checking therefore risk narrowing their methodological search space toward a more concentrated default.
Original Article
View Cached Full Text

Cached at: 06/26/26, 05:14 AM

# Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods
Source: [https://arxiv.org/html/2606.26130](https://arxiv.org/html/2606.26130)
Francesca Carlon1, 2,![[Uncaptioned image]](https://arxiv.org/html/2606.26130v1/panda2.png) [0009\-0004\-2152\-2745](https://orcid.org/0009-0004-2152-2745)&Brecht Verbeken1,2 [0000\-0002\-7506\-3298](https://orcid.org/0000-0002-7506-3298) &Vincent Ginis1,2,3 [0000\-0003\-0063\-9608](https://orcid.org/0000-0003-0063-9608)&Andres Algaba1,2 [0000\-0002\-0532\-3066](https://orcid.org/0000-0002-0532-3066) 1Data Analytics Lab, Vrije Universiteit Brussel, Pleinlaan 5, 1050 Brussels, Belgium 2imec\-SMIT, Vrije Universiteit Brussel, Pleinlaan 9, 1050 Brussels, Belgium 3School of Engineering and Applied Sciences, Harvard University, Cambridge, Massachusetts 02138, USA

###### Abstract

Large Language Models \(LLMs\) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal prompting remain unclear\. Here, we prompt GPT\-5\.1, Gemini 3 Pro, and DeepSeek\-V3\.2 with an LLM\-extracted research question from each of 1,000 recent arXiv computer\-science papers and compare the resulting methodology suggestions against a paper\-derived experimental inventory\. Since we provide only the research question, the differences we measure reflect initial suggestions and not how optimal those suggestions are\. We extract structured method features from both sources, map them into a shared taxonomy, and quantify divergence across multiple taxonomy dimensions including model provider, dataset task type, and evaluation metric type\. The strongest imbalance appears in provider choice, with Jensen–Shannon divergence about 3–5×\\timeslarger than any other taxonomy dimension\. Other/Academic single\-occurrence models are underrepresented by 23–24 percentage points, while reused academic/community models are slightly overrepresented \(4–6 pp\)\. LLMs also suggest a much narrower range of methods overall: the effective number of model entities contracts from 1,232 to 59–96, and inter\-LLM rank correlations \(0\.550\.55–0\.680\.68\) generally exceed LLM\-to\-paper correlations \(0\.330\.33–0\.560\.56\), so the distortions are largely shared across models\. Popularity baselines, BM25 retrieval calibration, and paper\-level similarity tests confirm that the outputs are query\-specific responses, but filtered through a narrower set of options\. Researchers who rely on LLM suggestions without cross\-checking therefore risk narrowing their methodological search space toward a more concentrated default\.

††footnotetext:![[Uncaptioned image]](https://arxiv.org/html/2606.26130v1/panda2.png)Corresponding author:[francesca\.carlon@vub\.be](https://arxiv.org/html/2606.26130v1/mailto:[email protected])
*K*eywordsAI scientist⋅\\cdotlarge language models⋅\\cdotresearch methodology⋅\\cdotscience of science⋅\\cdotscientometrics

## 1Introduction

![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/new_figure_overview.png)Figure 1:Overview of the study design\.We compare methodology suggestions from three LLMs, each prompted only with a research question, against paper\-derived reference inventories for 1,000 recent arXiv computer\-science papers on LLM\-related topics\. For each paper, GPT\-5\.1 extracts the datasets, models, and evaluation metrics actually used \(the paper\-derived reference inventory\) and generates a single research question from the title, abstract, and full text\. Each of GPT\-5\.1, Gemini 3 Pro, and DeepSeek\-V3\.2 then proposes datasets, models, metrics, and a short pipeline from that question alone\. Normalised entity\-name analyses use the full 1,000\-paper outputs per source, while taxonomy\- and matched\-paper analyses use smaller classified or shared\-paper subsets reported in the relevant captions and tables\. This design isolates first\-pass recommendation behaviour under deliberately sparse, research\-question\-only input\.Large Language Models \(LLMs\) are increasingly integrated into scientific workflows, not only as writing assistants\(Lundet al\.,[2023](https://arxiv.org/html/2606.26130#bib.bib9); Jain and Jain,[2024](https://arxiv.org/html/2606.26130#bib.bib4)\)but as tools that may reshape scientific practice more broadly\(Musslicket al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib36)\)\. Researchers now consult LLMs for paper feedback\(Lianget al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib24)\), research\-idea and hypothesis generation\(Siet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib7); Baeket al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib13); Qiet al\.,[2023](https://arxiv.org/html/2606.26130#bib.bib12)\), knowledge synthesis and literature\-guided research support\(Liaoet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib3); Skarlinskiet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib25)\), and experimental design across natural language processing, computer vision, chemistry, materials science, social science, and others\(Liaoet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib3); Jain and Jain,[2024](https://arxiv.org/html/2606.26130#bib.bib4); Lianget al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib24); Siet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib7); Baeket al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib13); Qiet al\.,[2023](https://arxiv.org/html/2606.26130#bib.bib12); Boikoet al\.,[2023](https://arxiv.org/html/2606.26130#bib.bib26); Skarlinskiet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib25); Hewittet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib45); Manninget al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib46); Kusumegiet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib28); Fortunatoet al\.,[2018](https://arxiv.org/html/2606.26130#bib.bib35)\)\. Recent evidence suggests that scientific production itself is already shifting in response to LLM adoption\(Kusumegiet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib28)\)\. Several systems now automate larger parts of the research cycle, from autonomous idea generation and scientific discovery\(Luet al\.,[2026](https://arxiv.org/html/2606.26130#bib.bib1); Yamadaet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib33); Mitcheneret al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib15); Gottweiset al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib47); Elbadawiet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib18)\)to multi\-agent research workflows\(Liet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib2); Villaescusa\-Navarroet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib10); Schmidgallet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib5); Liet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib8); Wanget al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib14); Gridachet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib6); Liuet al\.,[2025a](https://arxiv.org/html/2606.26130#bib.bib16)\), literature screening\(Delgado\-Chaveset al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib32)\), and research\-code or methodology generation\(Gandhiet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib17); Novikovet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib34)\)\.

The science\-of\-science question these developments raise is not whether LLMs answer correctly, but which methods they make salient by default\. When a researcher asks a frontier model for first\-pass guidance on datasets, models, or evaluation metrics, the resulting suggestions define an initial menu of methodological options that may anchor downstream decisions\(Musslicket al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib36); Fortunatoet al\.,[2018](https://arxiv.org/html/2606.26130#bib.bib35)\)\. If that menu is systematically compressed or skewed, it could reshape the distribution of experimental designs across a field before deeper literature review even begins\.

Prior work already suggests that reinforcement learning from human feedback can reduce output diversity\(Kirket al\.,[2023](https://arxiv.org/html/2606.26130#bib.bib21); Luoet al\.,[2026](https://arxiv.org/html/2606.26130#bib.bib22)\), that LLM\-generated scholarship can reproduce unequal scientific recognition and citation patterns\(Liuet al\.,[2025b](https://arxiv.org/html/2606.26130#bib.bib29); Algabaet al\.,[2025b](https://arxiv.org/html/2606.26130#bib.bib31),[a](https://arxiv.org/html/2606.26130#bib.bib11); Mobiniet al\.,[2026](https://arxiv.org/html/2606.26130#bib.bib30)\), and that LLM use can contribute to benchmark saturation and homogenization in downstream evaluation and creative tasks\(Ballonet al\.,[2026](https://arxiv.org/html/2606.26130#bib.bib27); Bommasaniet al\.,[2022](https://arxiv.org/html/2606.26130#bib.bib40); Doshi and Hauser,[2024](https://arxiv.org/html/2606.26130#bib.bib41); Andersonet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib42)\)\. These concerns matter because science already operates under well\-documented pressures around citation inequality and reproducibility\(Nielsen and Andersen,[2021](https://arxiv.org/html/2606.26130#bib.bib23); Ioannidis,[2005](https://arxiv.org/html/2606.26130#bib.bib19); Open Science Collaboration,[2015](https://arxiv.org/html/2606.26130#bib.bib20); Camereret al\.,[2016](https://arxiv.org/html/2606.26130#bib.bib44)\)\. However, it remains unclear how LLMs reshape methodological attention in a contemporary research field, and which specific providers, model families, benchmarks, and evaluation criteria become more or less salient when frontier models serve as first\-pass intermediaries\.

Here, we study one such case: recent computer\-science papers on LLM\-related topics, where researchers are especially likely to use LLMs to choose benchmarks, model families, and evaluation criteria\. We present a large\-scale empirical comparison of methodology suggestions from GPT\-5\.1 \(documented in the GPT\-5 system card; exact model ID reported in[Section˜A\.3](https://arxiv.org/html/2606.26130#A1.SS3)\)\(Singhet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib57)\), Gemini 3 Pro\(Google DeepMind,[2025](https://arxiv.org/html/2606.26130#bib.bib79)\), and DeepSeek\-V3\.2\(DeepSeek\-AIet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib74)\)against paper\-derived methodology inventories extracted from 1,000 recent arXiv papers\. For each paper, we extract the datasets, models, and evaluation metrics used by the authors \(which we refer to as the paper\-derived reference inventory\), generate a single research question from the paper’s title, abstract, and full text, and ask each LLM to propose datasets, models, metrics, and a short pipeline from that question alone\. We compare suggestions against the paper\-derived inventories at three levels: individual entity frequencies, broad category distributions, and patterns of which methods appear together \([Figure˜1](https://arxiv.org/html/2606.26130#S1.F1)\)\. Full methodological details are provided in[Appendix˜A](https://arxiv.org/html/2606.26130#A1)\. Ablations and robustness analyses are reported in[Appendices˜B](https://arxiv.org/html/2606.26130#A2)and[C](https://arxiv.org/html/2606.26130#A3)\.

We find two regularities\. First, divergence concentrates in model provider\. All three LLMs overweight a small set of major commercial providers, and under a broader regrouping the main deficit lies in the singleton\-defined long tail rather than in reused academic or community models\. Second, LLMs generate question\-sensitive but sharply compressed method menus\. Effective model diversity contracts by 13–21×\\times, inter\-LLM rank correlations exceed LLM\-to\-paper correlations, and many exact\-name misses collapse to family\- or provider\-level matches\. We interpret divergence as redistribution of methodological attention relative to a paper\-derived reference corpus, not deviation from a normative optimum, and we study these effects in a deliberately sparse regime where each model sees only a research question\. Popularity baselines, BM25 calibration, and shuffled\-paper tests rule out generic templating\. Frontier LLMs respond to the question, but through a narrower and more provider\-concentrated vocabulary\.

## 2Results

We classify analyses by whether the taxonomy classifier sees only entity names \[EL\] or also the generated pipeline \[WP\]\. Section[2\.2](https://arxiv.org/html/2606.26130#S2.SS2)uses the \[EL\] baseline, whereas Section[2\.3](https://arxiv.org/html/2606.26130#S2.SS3)and the taxonomy\-classified analyses in Sections[C\.7](https://arxiv.org/html/2606.26130#A3.SS7)–[C\.8](https://arxiv.org/html/2606.26130#A3.SS8)use \[WP\]\. Entity\-name recall \(Section[C\.7](https://arxiv.org/html/2606.26130#A3.SS7)\) and normalisation\-threshold sensitivity \(Section[C\.10](https://arxiv.org/html/2606.26130#A3.SS10)\) are branch\-independent\. Section[C\.11](https://arxiv.org/html/2606.26130#A3.SS11)contains branch\-independent audits of extraction, normalisation, and introducedness, plus a \[WP\] audit of taxonomy labels; provider reliability is taken to apply to both \[EL\] and \[WP\] because provider labels barely move across branches \([Section˜B\.1](https://arxiv.org/html/2606.26130#A2.SS1)\)\. Figures based on normalised entity names \([Figures˜2](https://arxiv.org/html/2606.26130#S2.F2)and[3](https://arxiv.org/html/2606.26130#S2.F3)\) precede classification and carry no branch tag\. Provider remains the most divergent dimension in both settings \([Figure˜4](https://arxiv.org/html/2606.26130#S2.F4)a,[Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)\)\. Co\-occurrence matrices are unavailable in the \[EL\] branch because it stores only aggregate category counts\.

### 2\.1Question\-conditioned but compressed method menus

Table 1:LLMs operate with a substantially reduced model and metric vocabulary while comparatively preserving more dataset diversity\.Computed from deterministically normalised entity inventories after default fuzzy clustering \(T=90T=90\) for 1,000 papers per source\. Unique entity names, total mentions, and the ratio of unique names to total mentions for the paper\-derived reference inventory and each LLM\. Total mentions exceed the number of papers because each paper may contain multiple entities of each type\.We first compare individual dataset, model, and evaluation metric names between the paper\-derived reference inventory and the LLM suggestions\. Unless otherwise noted, entity\-name analyses use deterministic normalisation followed by fuzzy clustering at the default thresholdT=90T=90\([Section˜A\.6](https://arxiv.org/html/2606.26130#A1.SS6)\), and[Table˜1](https://arxiv.org/html/2606.26130#S2.T1)summarises vocabulary size and total mentions by entity type\. The clearest pattern is vocabulary compression\. LLMs operate with an order\-of\-magnitude smaller model vocabulary than the reference inventory \([Tables˜1](https://arxiv.org/html/2606.26130#S2.T1)and[5](https://arxiv.org/html/2606.26130#A3.T5)\), and coverage is correspondingly sparse\. Roughly four out of five reference datasets and metrics, and nine out of ten reference models, are not suggested by any LLM \([Figure˜2](https://arxiv.org/html/2606.26130#S2.F2)a–c\)\. These vocabulary totals are pipeline\-specific estimates rather than exact counts, because the blinded audit finds only moderate agreement on extraction \(inter\-model ICC\(2,1\)=0\.469\(2,1\)=0\.469\) and the perfect precision/recall/F1 in[Table˜E11](https://arxiv.org/html/2606.26130#A5.T11)applies only to the consensus subset\. LLMs also introduce many corpus\-novel entities, so they simultaneously omit entities present in full paper inventories and substitute alternatives not used in the target literature\.

![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_diversity_coverage_combined.png)Figure 2:LLMs cover only a small fraction of the entities in paper\-derived reference inventories and do so through markedly more concentrated distributions\.Panels use normalised entity inventories \([Section˜A\.6](https://arxiv.org/html/2606.26130#A1.SS6)\) from 1,000 papers per source\. \(a–c\) Share of reference\-inventory entities covered by 0, 1, 2, or all 3 LLMs\. \(d–f\) Recall of reference\-inventory entities by frequency decile \(decile 1 = most frequent\) for datasets, models, and metrics\. \(g\) Effective number of entities \(exp⁡\(H\)\\exp\(H\)\) for each source across entity types\. \(h\) Gini coefficient measuring frequency inequality\. The coverage gap is heavily concentrated in the long tail, with around 90% of reference models receiving no LLM coverage\.Information\-theoretic diversity measures show that this gap reflects concentration as well as missing coverage\. The effective number of model entities contracts by1313–21×21\\timesbetween the reference inventory and LLM suggestions, a reduction that is stable across fuzzy\-clustering thresholds \([Figure˜2](https://arxiv.org/html/2606.26130#S2.F2)g,h;[Table˜5](https://arxiv.org/html/2606.26130#A3.T5)\), and Gini coefficients rise in the same direction, indicating heavy reliance on a small set of names\. Coverage is also strongly frequency\-dependent\. The union of all three LLMs recovers most top\-decile datasets and metrics but only a small fraction of bottom\-decile entities \([Figure˜2](https://arxiv.org/html/2606.26130#S2.F2)d–f\)\. Among shared entities, log\-log regressions of the LLM\-to\-reference frequency ratio on reference rank yield significantly positive slopes \(allp<0\.001p<0\.001\), indicating that LLMs amplify the rarer entities they do cover\. Representative entity\-level examples are shown in[Figure˜3](https://arxiv.org/html/2606.26130#S2.F3)\.

![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure1_entity_preferences.png)Figure 3:The top\-ranked datasets, models, and metrics in paper\-derived reference inventories are overrepresented in LLM suggestions, while the long tail is truncated\.Computed from normalised entity inventories \([Section˜A\.6](https://arxiv.org/html/2606.26130#A1.SS6)\) aggregated over 1,000 papers per source\. Normalised frequency \(%\) of the top\-20 datasets \(a\), models \(b\), and top\-15 evaluation metrics \(c\), sorted by reference\-inventory frequency\. Bars show the reference inventory \(blue\), GPT\-5\.1 \(orange\), Gemini 3 Pro \(green\), and DeepSeek\-V3\.2 \(red\)\. For models, exact\-name comparisons should be interpreted with caution because the paper corpus \(June–December 2025\) partly postdates the training\-data cutoffs of the evaluated LLMs\. The figure makes visible how LLM suggestions cluster on well\-known benchmarks and major commercial models while under\-covering the long tail present in full paper\-level inventories\.Distortions are correlated across the three systems\. Inter\-LLM Spearman correlations reach0\.550\.55–0\.680\.68, generally exceeding reference\-to\-LLM correlations of0\.330\.33–0\.560\.56\([Table˜2](https://arxiv.org/html/2606.26130#S2.T2)\), so the three LLMs compress the methodological vocabulary in similar ways rather than making independent errors\. Because GPT\-5\.1 is also used for research\-question generation, paper\-side entity extraction, and taxonomy classification \([Appendix˜A](https://arxiv.org/html/2606.26130#A1)\), reference–GPT\-5\.1 overlap comparisons are not independent and may overstate GPT\-5\.1’s apparent closeness to the reference corpus relative to Gemini 3 Pro and DeepSeek\-V3\.2\. The inter\-LLM comparisons do not share this confound\.

Compression does not, however, imply fixed output\. Paper\-level similarity tests \([Section˜C\.6](https://arxiv.org/html/2606.26130#A3.SS6)\) show that same\-paper taxonomy similarity between the reference inventory and LLM suggestions significantly exceeds shuffled baselines \(p<0\.001p<0\.001across all nine entity type×\\timesLLM combinations\), with effect sizes spanning0\.0750\.075–0\.2660\.266from models to metrics \([Table˜E7](https://arxiv.org/html/2606.26130#A5.T7)\)\. Suggestions are conditioned on individual research questions, but each question is answered through a narrower vocabulary than the paper ecosystem provides\.

Table 2:Inter\-LLM rank correlations generally exceed paper\-to\-LLM correlations, indicating shared distortions rather than independent errors\.Computed from normalised entity\-count series \([Section˜A\.6](https://arxiv.org/html/2606.26130#A1.SS6)\) aggregated over 1,000 papers per source\. Spearman rank correlation\(Spearman,[1904](https://arxiv.org/html/2606.26130#bib.bib68)\)\(ρ\\rho\) is computed on shared support \(the intersection of entities with non\-zero counts in both sources\) and Jaccard similarity\(Jaccard,[1912](https://arxiv.org/html/2606.26130#bib.bib69)\)atK=20K\{=\}20for each pair of sources across datasets, models, and evaluation metrics\. Reference–GPT\-5\.1 pairs are not fully symmetric because GPT\-5\.1 also generates the research question and the paper\-side inventories \([Appendix˜A](https://arxiv.org/html/2606.26130#A1)\); these values may overstate GPT\-5\.1’s closeness to the paper\-derived reference inventory relative to Gemini 3 Pro and DeepSeek\-V3\.2, whereas inter\-LLM pairs are unaffected\.We turn next to whether these shared distortions reflect systematic concentration in specific method categories \(model provider, dataset modality, evaluation type\) that entity\-level analysis alone cannot reveal\.

### 2\.2Provider concentration dominates taxonomy\-level divergence

We classify all entities into structured category schemes \([Section˜A\.8](https://arxiv.org/html/2606.26130#A1.SS8)\) and compare the resulting distributions across sources\. Unless otherwise noted, the summary category results in this subsection use labels assigned from the entity lists alone and are based on the classified outputs available at analysis time \(n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, 998 DeepSeek\-V3\.2 papers; the lower counts for Gemini 3 Pro and DeepSeek\-V3\.2 arise from classification failures during processing, and entries with broken or invalid classifications were excluded\)\. We keep that entity\-list\-only baseline separate from the richer setting in which the classifier also sees the suggested pipeline\. Appendix B measures how much labels move when pipeline or web search are added, and Appendix C recomputes robustness statistics under the richer labelling setting\. Across all 15 category dimensions, provider is the dominant divergence\. In the entity\-list\-only baseline \([Figure˜4](https://arxiv.org/html/2606.26130#S2.F4)a\) and in the with\-pipeline robustness setting alike, provider carries the largest JSD, about33–5×5\\timesthe next\-largest, and also the largest Cramér’sVV, though the Cramér’sVVmargin is smaller \([Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)\)\. By contrast, low\-JSD dimensions such as openness and data quality show much closer aggregate distributions, but only openness reaches the strong audit tier\. Data quality remains exploratory because it falls below theκ\\kappathreshold\.

![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_provider_composite.png)Figure 4:Model\-provider choice is the single most divergent taxonomy dimension, with the deficit concentrated in the singleton\-defined long tail rather than in reused academic/community models\.Composite figure combining one baseline branch with two robustness summaries; a dashed separator marks the branch switch\. \(a–b\) \[EL, mixed strong/moderate/tentative/exploratory dimensions\]: \(a\) Jensen–Shannon divergence \(JSD, base\-2\) between the reference inventory and each LLM for all 15 taxonomy dimensions, grouped by entity type; computed in the entity\-list\-only baseline using classified outputs from n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, and 998 DeepSeek\-V3\.2 papers\. \(b\) Excess share \(percentage\-point difference from reference inventory\) for each model provider×\\timesLLM\. Black borders mark “own provider” cells; all three LLMs overrepresent Meta AI and OpenAI while underrepresenting the aggregate Other/Academic category\. \(c–d\) \[WP, strong: provider\]: \(c\) Provider distribution under a broader four\-way provider regrouping \([Section˜C\.8](https://arxiv.org/html/2606.26130#A3.SS8)\), computed on the all\-three\-LLM shared model\-paper subset \(n=915=915\)\. \(d\) Mean per\-paper model recall under exact\-name, family\-level, and provider\-level matching \(provider recall deduplicated per paper;[Section˜C\.9](https://arxiv.org/html/2606.26130#A3.SS9)\), on the same n=915=915shared\-paper subset\. The substantial rise from exact to family and provider matching shows that many LLM misses are errors of granularity, not complete failures to recover the relevant model family\.As detailed in the blinded audit \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\), the 15 taxonomy dimensions fall into four validation tiers\. Strong dimensions \(provider, openness\) anchor our main claims\. Moderate dimensions \(modality, evaluation type, linguistic scope\) meet theκ≥0\.5\\kappa\\geq 0\.5and accuracy≥75%\\geq 75\\%thresholds with narrower margins, and tentative size is adequate overall but drops on the model\-side LLM stratum, so size\-dependent claims carry more uncertainty\. All remaining dimensions are exploratory and provide descriptive context rather than definitive structural claims\. Full tier statistics,κ\\kappavalues, and accuracies are reported in[Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)andLABEL:tab:annotation\_classification\.

All three LLMs share nearly the same provider\-concentration profile \([Figure˜4](https://arxiv.org/html/2606.26130#S2.F4)b\)\. They overrepresent Meta AI and OpenAI while underrepresenting the aggregate Other/Academic category and Alibaba/Qwen\. In the paper\-derived reference inventory, academic and independent models together account for roughly two\-fifths of model mentions, and all three LLMs roughly halve that share while inflating Meta AI \(panel b gives per\-provider percentage\-point deltas,[Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)givesχ2\\chi^\{2\}and effect sizes\)\. The fine\-grained Other/Academic category aggregates reused academic/community models with singletons that appear in only one paper\. Under a broader four\-way regrouping that separates these two groups, the deficit falls in the singleton\-defined long tail, and reused academic/community models are modestly overrepresented \([Figure˜4](https://arxiv.org/html/2606.26130#S2.F4)c,[Section˜C\.8](https://arxiv.org/html/2606.26130#A3.SS8)\)\. The deficit is therefore major\-provider concentration plus long\-tail suppression, not a blanket suppression of academic or community models\. Only GPT\-5\.1 shows clear own\-provider self\-preferencing, so the shared concentration toward a small set of commercial providers is difficult to explain as simple self\-promotion\.

Most other dimensions are closer to the reference inventory, but the remaining deviations are directionally consistent\. Dataset and model size rank next by JSD, and evaluation type shifts toward accuracy\-like metrics while user\-experience and efficiency metrics decline \([Figure˜D1](https://arxiv.org/html/2606.26130#A4.F1),[Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)\)\. LLMs consistently overweight the top model\-size bucket, though size labels are coarse analysis buckets rather than exact parameter\-count bands, especially for closed models whose sizes are estimated from public information\. In the with\-pipeline robustness setting \([Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)\), provider remains dominant and size forms the next tier\. Model architecture appears elevated in some summaries but stays exploratory and label\-sensitive under the audit and ablation checks \([Sections˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)and[B\.2](https://arxiv.org/html/2606.26130#A2.SS2)\)\.

The remaining dataset dimensions \(domain, annotation type, linguistic scope, cognitive/affective properties, granularity, data quality\) all exhibit low aggregate divergence \(mean JSD==0\.002–0\.006\), with small shifts toward self\-supervised annotation, multilingual datasets, and away from education\-domain content \([Figure˜D1](https://arxiv.org/html/2606.26130#A4.F1)caption\)\. These deviations are an order of magnitude smaller than the provider distortions documented above\.

The observed distributional divergences are substantially smaller than what a deterministic popularity baseline produces, and larger than what a popularity\-proportional sampler produces\. For provider, model size, and metric evaluation type alike, each LLM sits between a stochastic sampled baseline \(near\-zero JSD\) and a deterministic top\-kkbaseline \(one to two orders of magnitude higher JSD\), always closer to the sampled floor than to the top\-kkceiling \([Table˜E4](https://arxiv.org/html/2606.26130#A5.T4)\)\. LLM suggestions therefore track paper\-level content without collapsing to a fixed popularity ranking, even though they exhibit distributional biases that a popularity\-proportional sampler would not\. A complementary calibration using leave\-one\-out BM25 retrieval over generated research questions improves substantially over global popularity but remains more divergent than the LLMs on all three main dimensions \([Table˜E5](https://arxiv.org/html/2606.26130#A5.T5)\)\. Equal\-count comparisons that match entities per paper between reference and LLM preserve the hierarchy, with provider the dominant model dimension \([Section˜C\.4](https://arxiv.org/html/2606.26130#A3.SS4)\)\. These results, together with the paper\-level question\-sensitivity tests in[Section˜2\.1](https://arxiv.org/html/2606.26130#S2.SS1), suggest that provider concentration is a systematic distributional property of LLM suggestions, not an artefact of fixed template outputs\.

![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure4_cooccurrence_evtype_provider.png)Figure 5:LLM suggestions shift evaluation\-type×\\timesprovider co\-occurrences toward Meta AI and OpenAI and away from the aggregated Other/Academic category \[WP\]\.Both axes are validated dimensions \(provider: strong, 89\.7–94\.1% accuracy,κ=0\.923\\kappa=0\.923–1\.0001\.000; evaluation type: moderate, 89\.7% LLM accuracy,κ=0\.588\\kappa=0\.588\), though evaluation\-type reliability is asymmetric and rests on the LLM stratum \(reference\-stratumκ=0\.061\\kappa=0\.061onn=17n=17rows\)\. Panels show the percentage of papers \(within each source\) in which a given evaluation\-type row co\-occurs with a given provider column, computed from the with\-pipeline classification branch; source files contain n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, and 998 DeepSeek\-V3\.2 papers, with each panel restricted to papers with non\-empty labels for both entity types\. Heatmaps show the top 7 row and column categories\. A complementary exploratory figure for dataset task type×\\timesmodel provider \(task\-type labels have only 10–32% audit accuracy in the primary strata\) is in[Figure˜D5](https://arxiv.org/html/2606.26130#A4.F5)\. Among the reported co\-occurrence analyses, this evaluation type×\\timesprovider pair is the most interpretable because provider is strongly validated and evaluation type is moderately validated in the LLM stratum; reference\-stratum reliability for evaluation type is the only weak point \(κ=0\.061\\kappa=0\.061,n=17n=17\)\.
### 2\.3LLMs narrow provider\-centred co\-occurrence patterns

Beyond frequency, we test whether LLMs reproduce method combinations\. The analyses in this subsection use the with\-pipeline classification branch \([Section˜A\.9](https://arxiv.org/html/2606.26130#A1.SS9)\) and serve as a structural robustness check, not a direct numeric continuation of the entity\-list\-only category distributions in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)\.[Figure˜5](https://arxiv.org/html/2606.26130#S2.F5)shows the evaluation type×\\timesmodel provider co\-occurrence across sources\. Evaluation\-type patterns are broadly distributed across providers in the reference inventory, whereas in LLM suggestions they concentrate on Meta AI and OpenAI, and the aggregated Other/Academic category declines\. Much of that decline reflects the singleton\-defined long tail identified in[Section˜C\.8](https://arxiv.org/html/2606.26130#A3.SS8)\. Among the reported co\-occurrence analyses, this pair is the most interpretable because provider is strongly validated and evaluation type is moderately validated in the LLM stratum; reference\-stratum reliability for evaluation type is the only weak point \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\. A complementary exploratory analysis of dataset task type×\\timesmodel provider is reported in[Figure˜D5](https://arxiv.org/html/2606.26130#A4.F5); task\-type labels have only 10–32% audit accuracy in the primary strata and provide descriptive context rather than precise semantic claims\.

Across all 12 taxonomy pairs, provider\-based combinations show the largest structural divergence \([Figure˜D2](https://arxiv.org/html/2606.26130#A4.F2)\)\. Some architecture\-based pairs also show non\-trivial divergence, but architecture labels are exploratory and sensitive to classification context \([Sections˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)and[B\.2](https://arxiv.org/html/2606.26130#A2.SS2)\), so those patterns are descriptive only\. Openness\-based pairs are the most faithfully preserved\. Averaged across all 12 pairs, Gemini 3 Pro is closest to the reference inventory and DeepSeek\-V3\.2 farthest, though the ordering varies by pair\.

Residual analysis that holds row and column totals fixed shows that LLMs do not invent an entirely new interaction structure, though they do sharpen it\. Residual correlations for the focal provider pairs are moderate, and sign flips are rare \([Table˜E6](https://arxiv.org/html/2606.26130#A5.T6)\)\. LLMs therefore usually preserve the direction of associations while amplifying a narrower subset of provider\-centred combinations\. This narrowing is not uniform across entity types\. The inter\-paper Jaccard decomposition \([Table˜E7](https://arxiv.org/html/2606.26130#A5.T7)\) shows that most of the increased inter\-paper overlap for datasets is accounted for by marginal frequency shifts, that the pattern for metrics is mixed, and that excess homogenisation beyond vocabulary compression is concentrated in model suggestions, where all three LLMs show significantly positiveΔexcess\\Delta\_\{\\mathrm\{excess\}\}\([Section˜C\.6](https://arxiv.org/html/2606.26130#A3.SS6)\)\. The co\-occurrence narrowing is therefore mainly a model and provider effect, not an equally strong flattening across all entity types, with the practical consequence that fewer provider\-centred combinations remain salient on the suggested menu\.

### 2\.4Robustness: granularity, long tails, and label reliability

Three complementary analyses sharpen our interpretation of the compression and concentration patterns documented above \([Sections˜C\.7](https://arxiv.org/html/2606.26130#A3.SS7),[E8](https://arxiv.org/html/2606.26130#A5.T8),[C\.8](https://arxiv.org/html/2606.26130#A3.SS8),[E9](https://arxiv.org/html/2606.26130#A5.T9),[C\.9](https://arxiv.org/html/2606.26130#A3.SS9)and[E10](https://arxiv.org/html/2606.26130#A5.T10),[Figure˜D8](https://arxiv.org/html/2606.26130#A4.F8)\)\. They clarify where the measured divergence reflects granularity mismatches, where it reflects long\-tail underweighting, and where the evidence is strongest\.

Exact\-name model recall is low \(4–7% per paper\), but relaxed matching substantially improves it\. Mean per\-paper recall rises to 20–28% at family level and 45–53% at provider level \(deduplicated per paper,[Section˜C\.9](https://arxiv.org/html/2606.26130#A3.SS9)\)\. These are per\-paper robustness scores on the shared\-paper model subset, not direct re\-estimates of the raw 1,000\-paper entity\-frequency branch\. They show that many apparent misses are errors of granularity, in which an LLM names a different variant from the same model family, not complete failures to recover the relevant model family\.

Under the broader four\-way provider regrouping \([Section˜C\.8](https://arxiv.org/html/2606.26130#A3.SS8)\), the dominant deficit falls in the singleton\-defined long tail \(roughly−\-24 pp across LLMs\), and reused academic/community models are modestly overrepresented \(\+\+4–6 pp\)\. The pattern is therefore concentration on major commercial providers alongside long\-tail suppression, not academic erasure\. Removing singleton entities from the reference inventory improves recall but also removes many established methods\. The singleton filter has only 7\.5% precision as a proxy for paper\-specificity \([Tables˜E15](https://arxiv.org/html/2606.26130#A5.T15)and[C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\), so singleton exclusion is a long\-tail sensitivity analysis rather than a principled fairness correction, and the compression remains large even after these adjustments\.

Finally, the blinded cross\-model audit \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\) classifies the 15 taxonomy dimensions into four validation tiers\. Provider and openness are strong\. Modality, evaluation type, and linguistic scope are moderate, while size is tentative \(adequate overall but weaker on the model\-side LLM stratum\)\. All remaining dimensions are exploratory and provide descriptive context rather than definitive structural claims\. Full tier statistics are reported inLABEL:tab:annotation\_classificationand[6](https://arxiv.org/html/2606.26130#A3.T6)\.

## 3Discussion

Two findings emerge from these analyses about how LLMs redistribute methodological attention under research\-question\-only prompting\.

First, LLM suggestions are commercially concentrated and suppress Other/Academic singleton models\. Provider divergence is about 3–5×\\timeslarger than the next\-largest taxonomy dimension, and the same pattern survives broader provider regrouping, popularity\-baseline comparisons, BM25 retrieval calibration, and cross\-model robustness checks in which Claude Opus 4\.6 re\-extracts entities from a 94\-paper subset and re\-classifies 200 papers’ entities in place of GPT\-5\.1 \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\. When providers are grouped into four broad categories, the shortfall concentrates in singleton Other/Academic models \(those appearing in exactly one paper in the reference corpus,−\-23\.5 to−\-24\.3 pp\), while established academic and community models are modestly overrepresented \(\+\+3\.6 to\+\+5\.8 pp\)\. LLMs therefore favour well\-known providers rather than suppressing non\-commercial work across the board\. The concentration also propagates to provider\-centred method combinations\. Evaluation type×\\timesprovider is the most interpretable co\-occurrence pair among those reported because provider is strongly validated and evaluation type is moderately validated in the LLM stratum \(LLMκ=0\.588\\kappa=0\.588, 89\.7% accuracy\); the weak point is reference\-stratum reliability for evaluation type \(κ=0\.061\\kappa=0\.061onn=17n=17pairwise rows\)\. Excess homogenisation beyond vocabulary compression is significantly positive for models \([Section˜2\.3](https://arxiv.org/html/2606.26130#S2.SS3)\)\.

Second, LLMs narrow first\-pass method menus more broadly\. Model diversity contracts from an effective 1,232 entities in the paper\-derived inventory to 59–96 in LLM suggestions \(these effective numbers inherit extraction and normalisation uncertainty,[Sections˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)and[C\.10](https://arxiv.org/html/2606.26130#A3.SS10)\)\. At the corpus level, 78–90% of paper\-side entities receive no coverage from any model, though many exact\-name misses are errors of granularity\. Per\-paper recall rises to 20–28% at family level and 45–53% at provider level \(per\-paper means on then=915n=915shared\-paper subset, not directly comparable to the corpus\-level exact\-name statistics above\)\. Inter\-LLM rank correlations generally exceed LLM\-to\-paper correlations, so the compressions are shared rather than idiosyncratic, with the three models converging on similar subsets of the methodology space\.

First\-pass suggestions can shape the initial menu of options researchers consider before deeper literature review, so these patterns matter for downstream design\. If multiple frontier models share similar provider\-centred and popularity\-weighted tendencies, over\-reliance on them may narrow the subset of experimental designs that receive early consideration\(Bommasaniet al\.,[2022](https://arxiv.org/html/2606.26130#bib.bib40); Doshi and Hauser,[2024](https://arxiv.org/html/2606.26130#bib.bib41); Andersonet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib42); Fortunatoet al\.,[2018](https://arxiv.org/html/2606.26130#bib.bib35); Musslicket al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib36)\)\. These results do not show that LLMs ignore the question\. Popularity baselines, BM25 retrieval calibration, equal\-count comparisons, and paper\-by\-paper similarity tests all show that suggestions still vary with the input problem \([Section˜2\.1](https://arxiv.org/html/2606.26130#S2.SS1)\)\. LLMs respond to the task, but through a compressed and commercially concentrated vocabulary\. This conclusion rests primarily on the strong provider dimension, and claims involving moderate or lower\-tier dimensions should be weighted accordingly \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\.

#### Limitations\.

This study isolates the recommendation behaviour of LLMs under research\-question\-only prompting in one field and evaluates it against a paper\-derived reference corpus rather than a normative optimum\. It therefore identifies output\-level redistribution of methodological attention, not the mechanism producing it, and it does not by itself establish generality beyond recent arXiv computer\-science papers on LLM\-related topics\. Because research\-question generation, paper\-side entity extraction, and taxonomy classification all use GPT\-5\.1, reference\-inventory–GPT\-5\.1 comparisons are not independent\. Our main claims rest on patterns shared across all three models rather than on GPT\-5\.1\-specific closeness, and a model\-swap robustness check using Claude Opus 4\.6 for both extraction and classification confirms the same provider\-concentration pattern \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\. Exact\-name model misses are also partly inflated by temporal mismatch, because the paper corpus \(June–December 2025\) postdates the known training\-data cutoffs of GPT\-5\.1 and Gemini 3 Pro\. The suggestion prompt explicitly asks for “one or more” specific names and a “structured and short” pipeline \([Appendix˜G](https://arxiv.org/html/2606.26130#A7)\), so part of the observed compression reflects this concise, single\-shot elicitation rather than a property intrinsic to LLM\-assisted methodology search more broadly\. Adding richer context \(abstracts, full papers, iterative follow\-up, or retrieval augmentation\) might reduce the observed compression, so the results bound a lower\-context baseline rather than the full range of LLM\-assisted research workflows\(Baeket al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib13); Liet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib2); Skarlinskiet al\.,[2024](https://arxiv.org/html/2606.26130#bib.bib25)\)\.

Audit confidence varies across dimensions\. Provider and openness are strong \(κ≥0\.923\\kappa\\geq 0\.923\)\. Modality, evaluation type, and linguistic scope are moderate \(κ=0\.588\\kappa=0\.588–0\.7560\.756\)\. Size is tentative \(model\-side LLM accuracy drops to 47\.1%\), and all others are exploratory \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\. These confidence tiers apply to all consensus\-row diagnostics, which themselves are best\-case checks covering 59–95% of audited rows rather than corpus\-level accuracy estimates\. These constraints bound the interpretation but do not alter the central descriptive result, which rests on the strongly validated provider dimension\.

The architecture schema itself is a source of overlap\. It combines backbone types \(Transformer, CNN, RNN/LSTM, GNN\), functional roles \(Generative\), and training paradigms \(Reinforcement Learning\), and this mixing contributes to the label sensitivity of the architecture category and helps explain why results shift sharply under the web\-search ablation\. The schema notes already acknowledge leakage between task types and adjacent semantic categories\. The resulting low reliability reflects a structural limitation of the taxonomy, which conflates distinct dimensions, rather than a pure annotation issue\. Making the limitation explicit clarifies where and why the measurement becomes unreliable\.

More broadly, LLM\-based research support should be evaluated not only for relevance or correctness, but also for how it redistributes methodological attention across providers, families, and long\-tail alternatives\. The vocabulary compression and provider concentration documented here are output\-level distributional regularities\. Our study does not distinguish whether they arise from training\-data exposure, benchmark popularity, or internal model preferences, so disentangling these mechanisms is an important direction for future work\.

#### Future directions\.

The most direct extension is to test whether richer suggestion\-time context, iterative prompting, or retrieval reduces the convergence patterns we observe here\. Replicating the analysis in other scientific domains would show whether the same distortions generalise beyond computer science, and longitudinal analyses could test whether they weaken or reinforce as training data changes\. Causal studies of how LLM advice changes researchers’ actual methodological choices would help clarify whether the descriptive concentration patterns documented here translate into measurable shifts in experimental practice\.

## Acknowledgements

This research was supported by funding from the Flemish Government under the “Onderzoeksprogramma Artificiële Intelligentie \(AI\) Vlaanderen” program\. Andres Algaba acknowledges support from the Francqui Foundation \(Belgium\) through a Francqui Start\-Up Grant and a fellowship from the Research Foundation Flanders \(FWO\) under Grant No\. 1286924N\. Vincent Ginis acknowledges support from Research Foundation Flanders \(FWO\) under Grant Nos\. G032822N and G0K9322N\. The resources and services used in this work were provided by the VSC \(Flemish Supercomputer Center\), funded by Research Foundation Flanders \(FWO\) and the Flemish Government\.

## Data and code availability statement

The code for the full pipeline, including data collection, arXiv license\-metadata collection, paper\-side entity extraction, LLM suggestion generation, entity normalisation, taxonomy classification, and figure generation, is available at[https://github\.com/francescacarlon/Thinking\-Like\-a\-Scientist](https://github.com/francescacarlon/Thinking-Like-a-Scientist)\. The processed analysis outputs, including derived entity inventories, frequency counts, ratio tables, taxonomy classifications, co\-occurrence matrices, and per\-paper arXiv license metadata, are included in the repository\. We do not redistribute arXiv PDFs, extracted full texts, substantial expressive excerpts from papers, or API batch/input/output files containing paper full text\. Raw arXiv papers can be retrieved from arXiv using the collection script, subject to arXiv’s API terms and the license associated with each paper\. Local raw PDFs and extracted full\-text files, and provider\-side batch/input/output file objects under our control, were deleted after the structured annotations had been generated and verified\.

## References

- How deep do large language models internalize scientific literature and citation practices?\.arXiv preprint arXiv:2504\.02767\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- A\. Algaba, C\. Mazijn, V\. Holst, F\. Tori, S\. Wenmackers, and V\. Ginis \(2025b\)Large language models reflect human citation patterns with a heightened citation bias\.InFindings of the Association for Computational Linguistics: NAACL 2025,pp\. 6844–6879\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-naacl.381),[Link](https://doi.org/10.18653/v1/2025.findings-naacl.381)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- B\. R\. Anderson, J\. H\. Shah, and M\. Kreminski \(2024\)Homogenization effects of large language models on human creative ideation\.InProceedings of the 16th Conference on Creativity and Cognition,pp\. 413–425\.External Links:[Document](https://dx.doi.org/10.1145/3635636.3656204)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1),[§3](https://arxiv.org/html/2606.26130#S3.p4.1)\.
- Anthropic \(2026a\)Anthropic commercial terms of service\.Note:[https://www\.anthropic\.com/legal/commercial\-terms](https://www.anthropic.com/legal/commercial-terms)Effective 2025\-06\-17; accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1),[Table 4](https://arxiv.org/html/2606.26130#A1.T4.3.1.1.1.5.4.3.1.1)\.
- Anthropic \(2026b\)How long do you store my data?\.Note:[https://privacy\.claude\.com/en/articles/10023548\-how\-long\-do\-you\-store\-my\-data](https://privacy.claude.com/en/articles/10023548-how-long-do-you-store-my-data)Article dated 2026\-03\-16; accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1),[Table 4](https://arxiv.org/html/2606.26130#A1.T4.3.1.1.1.5.4.3.1.1)\.
- arXiv \(2026a\)ArXiv license information\.Note:[https://info\.arxiv\.org/help/license/index\.html](https://info.arxiv.org/help/license/index.html)Accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p1.1)\.
- arXiv \(2026b\)Terms of use for arxiv apis\.Note:[https://info\.arxiv\.org/help/api/tou\.html](https://info.arxiv.org/help/api/tou.html)Accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p1.1)\.
- J\. Baek, S\. K\. Jauhar, S\. Cucerzan, and S\. J\. Hwang \(2025\)ResearchAgent: iterative research idea generation over scientific literature with large language models\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 6709–6738\.External Links:[Link](https://doi.org/10.18653/v1/2025.naacl-long.342),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.342)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1),[§3](https://arxiv.org/html/2606.26130#S3.SS0.SSS0.Px1.p1.1)\.
- M\. Ballon, A\. Algaba, B\. Verbeken, and V\. Ginis \(2026\)Benchmarks saturate when the model gets smarter than the judge\.arXiv preprint arXiv:2601\.19532\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- Belgian Federal Public Service Economy \(2022\)European directive on copyright and related rights in the digital single market – transposition in belgian law\.Note:[https://economie\.fgov\.be/en/themes/intellectual\-property/intellectual\-property\-rights/copyright\-and\-related\-rights/copyright/european\-directive\-copyright](https://economie.fgov.be/en/themes/intellectual-property/intellectual-property-rights/copyright-and-related-rights/copyright/european-directive-copyright)Accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1)\.
- Y\. Benjamini and Y\. Hochberg \(1995\)Controlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal Statistical Society: Series B \(Methodological\)57\(1\),pp\. 289–300\.External Links:[Document](https://dx.doi.org/10.1111/j.2517-6161.1995.tb02031.x)Cited by:[§C\.5](https://arxiv.org/html/2606.26130#A3.SS5.p1.7)\.
- D\. A\. Boiko, R\. MacKnight, B\. Kline, and G\. Gomes \(2023\)Autonomous chemical research with large language models\.Nature624\(7992\),pp\. 570–578\.External Links:[Document](https://dx.doi.org/10.1038/s41586-023-06792-0),ISSN 1476\-4687,[Link](https://doi.org/10.1038/s41586-023-06792-0)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- R\. Bommasani, K\. A\. Creel, A\. Kumar, D\. Jurafsky, and P\. Liang \(2022\)Picking on the same person: does algorithmic monoculture lead to outcome homogenization?\.InAdvances in Neural Information Processing Systems,Vol\.35,pp\. 3663–3678\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1),[§3](https://arxiv.org/html/2606.26130#S3.p4.1)\.
- C\. F\. Camerer, A\. Dreber, E\. Forsell, T\. Ho, J\. Huber, M\. Johannesson, M\. Kirchler, J\. Almenberg, A\. Altmejd, T\. Chan, E\. Heikensten, F\. Holzmeister, T\. Imai, S\. Isaksson, G\. Nave, T\. Pfeiffer, M\. Razen, and H\. Wu \(2016\)Evaluating replicability of laboratory experiments in economics\.Science351\(6280\),pp\. 1433–1436\.External Links:[Document](https://dx.doi.org/10.1126/science.aaf0918),ISSN 1095\-9203,[Link](https://doi.org/10.1126/science.aaf0918)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- J\. Cohen \(1960\)A coefficient of agreement for nominal scales\.Educational and Psychological Measurement20\(1\),pp\. 37–46\.Cited by:[§C\.11](https://arxiv.org/html/2606.26130#A3.SS11.p1.5)\.
- J\. Cohen \(1988\)Statistical power analysis for the behavioral sciences\.2nd edition,Lawrence Erlbaum Associates,Hillsdale, NJ\.Cited by:[§C\.1](https://arxiv.org/html/2606.26130#A3.SS1.p1.11)\.
- H\. Cramér \(1946\)Mathematical methods of statistics\.Princeton University Press,Princeton, NJ\.Cited by:[§C\.1](https://arxiv.org/html/2606.26130#A3.SS1.p1.6)\.
- DeepSeek\-AI, A\. Liu, A\. Mei, B\. Lin, B\. Xue, B\. Wang, B\. Xu, B\. Wu, B\. Zhang, C\. Lin, C\. Dong, C\. Lu, C\. Zhao, C\. Deng, C\. Xu,et al\.\(2025\)DeepSeek\-V3\.2: pushing the frontier of open large language models\.arXiv preprint arXiv:2512\.02556\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.02556),[Link](https://arxiv.org/abs/2512.02556)Cited by:[3rd item](https://arxiv.org/html/2606.26130#A1.I2.i3.p1.1),[§1](https://arxiv.org/html/2606.26130#S1.p4.1)\.
- DeepSeek \(2025\)DeepSeek\-v3\.2 release\.Note:[https://api\-docs\.deepseek\.com/news/news251201](https://api-docs.deepseek.com/news/news251201)Published 2025\-12\-01; accessed 2026\-06\-01Cited by:[3rd item](https://arxiv.org/html/2606.26130#A1.I2.i3.p1.1)\.
- DeepSeek \(2026\)DeepSeek terms of use\.Note:[https://cdn\.deepseek\.com/policies/en\-US/deepseek\-terms\-of\-use\.html](https://cdn.deepseek.com/policies/en-US/deepseek-terms-of-use.html)Last updated 2026\-03\-27; accessed 2026\-06\-01Cited by:[Table 4](https://arxiv.org/html/2606.26130#A1.T4.3.1.1.1.4.3.3.1.1)\.
- F\. M\. Delgado\-Chaves, M\. J\. Jennings, A\. Atalaia, J\. Wolff, R\. Horvath, Z\. M\. Mamdouh, J\. Baumbach, and L\. Baumbach \(2025\)Transforming literature screening: the emerging role of large language models in systematic reviews\.Proceedings of the National Academy of Sciences122\(2\),pp\. e2411962122\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2411962122),ISSN 1091\-6490,[Link](https://doi.org/10.1073/pnas.2411962122)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- A\. R\. Doshi and O\. P\. Hauser \(2024\)Generative ai enhances individual creativity but reduces the collective diversity of novel content\.Science Advances10\(28\),pp\. eadn5290\.External Links:[Document](https://dx.doi.org/10.1126/sciadv.adn5290)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1),[§3](https://arxiv.org/html/2606.26130#S3.p4.1)\.
- M\. Elbadawi, H\. Li, A\. W\. Basit, and S\. Gaisford \(2024\)The role of artificial intelligence in generating original scientific research\.International Journal of Pharmaceutics652,pp\. 123741\.Note:Epub 2024 Jan 3External Links:[Document](https://dx.doi.org/10.1016/j.ijpharm.2023.123741)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- European Parliament and Council of the European Union \(2019\)Directive \(eu\) 2019/790 of the european parliament and of the council on copyright and related rights in the digital single market\.Note:[https://eur\-lex\.europa\.eu/eli/dir/2019/790/oj/eng](https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng)Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1)\.
- S\. Fortunato, C\. T\. Bergstrom, K\. Börner, J\. A\. Evans, D\. Helbing, S\. Milojević, A\. M\. Petersen, F\. Radicchi, R\. Sinatra, B\. Uzzi, A\. Vespignani, L\. Waltman, D\. Wang, and A\. Barabási \(2018\)Science of science\.Science359\(6379\),pp\. eaao0185\.External Links:[Document](https://dx.doi.org/10.1126/science.aao0185),ISSN 1095\-9203,[Link](https://doi.org/10.1126/science.aao0185)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1),[§1](https://arxiv.org/html/2606.26130#S1.p2.1),[§3](https://arxiv.org/html/2606.26130#S3.p4.1)\.
- S\. Gandhi, D\. Shah, M\. Patwardhan, L\. Vig, and G\. Shroff \(2025\)ResearchCodeAgent: an llm multi\-agent system for automated codification of research methodologies\.InAI for Research and Scalable, Efficient Systems,pp\. 3–37\.External Links:[Document](https://dx.doi.org/10.1007/978-981-96-8912-5%5F1),[Link](https://doi.org/10.1007/978-981-96-8912-5_1),ISSN 1865\-0937,ISBN 9789819689125Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- C\. Gini \(1912\)Variabilità e mutabilità: contributo allo studio delle distribuzioni e delle relazioni statistiche\.Tipografia di Paolo Cuppini,Bologna\.External Links:[Link](https://www.byterfly.eu/islandora/object/librib%3A680892)Cited by:[§A\.7](https://arxiv.org/html/2606.26130#A1.SS7.p2.5)\.
- Google DeepMind \(2025\)Gemini 3 pro model card\.Technical reportGoogle DeepMind\.External Links:[Link](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-Pro-Model-Card.pdf)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p4.1)\.
- Google \(2026\)Gemini api additional terms of service\.Note:[https://ai\.google\.dev/gemini\-api/terms](https://ai.google.dev/gemini-api/terms)Effective 2026\-03\-23; accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1),[Table 4](https://arxiv.org/html/2606.26130#A1.T4.3.1.1.1.3.2.3.1.1)\.
- J\. Gottweis, W\. Weng, A\. Daryin, T\. Tu, A\. Palepu, P\. Sirkovic, A\. Myaskovsky, F\. Weissenberger, K\. Rong, R\. Tanno, K\. Saab, D\. Popovici, J\. Blum, F\. Zhang, K\. Chou, A\. Hassidim, B\. Gokturk, A\. Vahdat, P\. Kohli, Y\. Matias, A\. Carroll, K\. Kulkarni, N\. Tomasev, Y\. Guan, V\. Dhillon, E\. D\. Vaishnav, B\. Lee, T\. R\. D\. Costa, J\. R\. Penadés, G\. Peltz, Y\. Xu, A\. Pawlosky, A\. Karthikesalingam, and V\. Natarajan \(2025\)Towards an ai co\-scientist\.arXiv preprint arXiv:2502\.18864\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- M\. Gridach, J\. Nanavati, K\. Z\. E\. Abidine, L\. Mendes, and C\. Mack \(2025\)Agentic ai for scientific discovery: a survey of progress, challenges, and future directions\.arXiv preprint arXiv:2503\.08979\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- L\. Hewitt, A\. Ashokkumar, I\. Ghezae, and R\. Willer \(2024\)Predicting results of social science experiments using large language models\.Note:Working paperExternal Links:[Link](https://ai4pb.stanford.edu/projects/predicting-results-of-social-science-experiments-using-large-language-models)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- J\. P\. Ioannidis \(2005\)Why most published research findings are false\.PLoS medicine2\(8\),pp\. e124\.External Links:[Document](https://dx.doi.org/10.1371/journal.pmed.0020124),ISSN 1549\-1676,[Link](https://doi.org/10.1371/journal.pmed.0020124)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- P\. Jaccard \(1912\)The distribution of the flora in the alpine zone\.New Phytologist11\(2\),pp\. 37–50\.External Links:[Document](https://dx.doi.org/10.1111/j.1469-8137.1912.tb05611.x)Cited by:[§A\.7](https://arxiv.org/html/2606.26130#A1.SS7.p4.6),[§C\.6](https://arxiv.org/html/2606.26130#A3.SS6.SSS0.Px2.p1.8),[Table 2](https://arxiv.org/html/2606.26130#S2.T2)\.
- R\. Jain and A\. Jain \(2024\)Generative ai in writing research papers: a new type of algorithmic bias and uncertainty in scholarly work\.InIntelligent Systems and Applications,K\. Arai \(Ed\.\),Cham,pp\. 656–669\.External Links:ISBN 9783031663291,[Document](https://dx.doi.org/10.1007/978-3-031-66329-1%5F42),ISSN 2367\-3389,[Link](https://doi.org/10.1007/978-3-031-66329-1_42)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- L\. Jost \(2006\)Entropy and diversity\.Oikos113\(2\),pp\. 363–375\.External Links:[Document](https://dx.doi.org/10.1111/j.2006.0030-1299.14714.x),ISSN 1600\-0706,[Link](https://doi.org/10.1111/j.2006.0030-1299.14714.x)Cited by:[§A\.7](https://arxiv.org/html/2606.26130#A1.SS7.p2.5)\.
- R\. Kirk, I\. Mediratta, C\. Nalmpantis, J\. Luketina, E\. Hambro, E\. Grefenstette, and R\. Raileanu \(2023\)Understanding the effects of RLHF on LLM generalisation and diversity\.arXiv preprint arXiv:2310\.06452\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- K\. Kusumegi, X\. Yang, P\. Ginsparg, M\. de Vaan, T\. Stuart, and Y\. Yin \(2025\)Scientific production in the era of large language models\.Science390\(6779\),pp\. 1240–1243\.External Links:[Document](https://dx.doi.org/10.1126/science.adw3000),ISSN 1095\-9203,[Link](https://doi.org/10.1126/science.adw3000)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- V\. I\. Levenshtein \(1966\)Binary codes capable of correcting deletions, insertions and reversals\.Soviet Physics Doklady10\(8\),pp\. 707–710\.Note:English translation of the 1965 Russian originalCited by:[§A\.6](https://arxiv.org/html/2606.26130#A1.SS6.SSS0.Px2.p1.5)\.
- L\. Li, W\. Xu, J\. Guo, R\. Zhao, X\. Li, Y\. Yuan, B\. Zhang, Y\. Jiang, Y\. Xin, R\. Dang, Y\. Rong, D\. Zhao, T\. Feng, and L\. Bing \(2025\)Chain of ideas: revolutionizing research via novel idea development with LLM agents\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 8971–9004\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.477/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.477),ISBN 979\-8\-89176\-335\-7Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- R\. Li, T\. Patel, Q\. Wang, and X\. Du \(2024\)Mlr\-copilot: autonomous machine learning research based on large language models agents\.arXiv preprint arXiv:2408\.14033\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1),[§3](https://arxiv.org/html/2606.26130#S3.SS0.SSS0.Px1.p1.1)\.
- W\. Liang, Y\. Zhang, H\. Cao, B\. Wang, D\. Y\. Ding, X\. Yang, K\. Vodrahalli, S\. He, D\. S\. Smith, Y\. Yin, D\. A\. McFarland, and J\. Zou \(2024\)Can large language models provide useful feedback on research papers? A large\-scale empirical analysis\.NEJM AI1\(8\),pp\. AIoa2400196\.External Links:[Document](https://dx.doi.org/10.1056/AIoa2400196),ISSN 2836\-9386,[Link](https://doi.org/10.1056/AIoa2400196)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- Z\. Liao, M\. Antoniak, I\. Cheong, E\. Y\. Cheng, A\. Lee, K\. Lo, J\. C\. Chang, and A\. X\. Zhang \(2024\)Llms as research tools: a large scale survey of researchers’ usage and perceptions\.arXiv preprint arXiv:2411\.05025\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- J\. Lin \(1991\)Divergence measures based on the Shannon entropy\.IEEE Transactions on Information Theory37\(1\),pp\. 145–151\.External Links:[Document](https://dx.doi.org/10.1109/18.61115)Cited by:[§C\.1](https://arxiv.org/html/2606.26130#A3.SS1.p1.2)\.
- C\. Liu, C\. Wang, J\. Cao, J\. Ge, K\. Wang, L\. Zhang, M\. Cheng, P\. Zhao, T\. Li, X\. Jia,et al\.\(2025a\)A vision for auto research with llm agents\.arXiv preprint arXiv:2504\.18765\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- Y\. Liu, Á\. Elekes, J\. Lu, R\. Dorantes\-Gilardi, and A\. Barabási \(2025b\)Unequal scientific recognition in the age of LLMs\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 23558–23568\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1279),[Link](https://doi.org/10.18653/v1/2025.findings-emnlp.1279)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- C\. Lu, C\. Lu, R\. T\. Lange, Y\. Yamada, S\. Hu, J\. Foerster, D\. Ha, and J\. Clune \(2026\)Towards end\-to\-end automation of ai research\.Nature651\(8107\),pp\. 914–919\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- B\. D\. Lund, T\. Wang, N\. R\. Mannuru, B\. Nie, S\. Shimray, and Z\. Wang \(2023\)ChatGPT and a new academic reality: artificial intelligence\-written research papers and the ethics of large language models in scholarly publishing\.Journal of the Association for Information Science and Technology74\(5\),pp\. 570–581\.External Links:ISSN 2330\-1643,[Link](http://dx.doi.org/10.1002/asi.24750),[Document](https://dx.doi.org/10.1002/asi.24750)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- Q\. Luo, G\. King, M\. Puett, and M\. D\. Smith \(2026\)Inducing sustained creativity and diversity in large language models\.arXiv preprint arXiv:2603\.19519\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- B\. S\. Manning, K\. Zhu, and J\. J\. Horton \(2024\)Automated social science: language models as scientist and subjects\.Working PaperTechnical Report32381,National Bureau of Economic Research\.External Links:[Document](https://dx.doi.org/10.3386/w32381),[Link](https://www.nber.org/papers/w32381)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- L\. Mitchener, A\. Yiu, B\. Chang, M\. Bourdenx, T\. Nadolski, A\. Sulovari, E\. C\. Landsness, D\. L\. Barabasi, S\. Narayanan, N\. Evans,et al\.\(2025\)Kosmos: an ai scientist for autonomous discovery\.arXiv preprint arXiv:2511\.02824\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- M\. Mobini, V\. Holst, F\. Tori, A\. Algaba, and V\. Ginis \(2026\)Structurally human, semantically biased: detecting LLM\-generated references with embeddings and GNNs\.arXiv preprint arXiv:2601\.20704\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- S\. Musslick, L\. K\. Bartlett, S\. H\. Chandramouli, M\. Dubova, F\. Gobet, T\. L\. Griffiths, J\. Hullman, R\. D\. King, J\. N\. Kutz, C\. G\. Lucas, S\. Mahesh, F\. Pestilli, S\. J\. Sloman, and W\. R\. Holmes \(2025\)Automating the practice of science: opportunities, challenges, and implications\.Proceedings of the National Academy of Sciences122\(5\),pp\. e2401238121\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2401238121),ISSN 1091\-6490,[Link](https://doi.org/10.1073/pnas.2401238121)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1),[§1](https://arxiv.org/html/2606.26130#S1.p2.1),[§3](https://arxiv.org/html/2606.26130#S3.p4.1)\.
- M\. W\. Nielsen and J\. P\. Andersen \(2021\)Global citation inequality is on the rise\.Proceedings of the National Academy of Sciences118\(7\),pp\. e2012208118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2012208118),ISSN 1091\-6490,[Link](https://doi.org/10.1073/pnas.2012208118)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- A\. Novikov, N\. Vũ, M\. Eisenberger, E\. Dupont, P\. Huang, A\. Z\. Wagner, S\. Shirobokov, B\. Kozlovskii, F\. J\. R\. Ruiz, A\. Mehrabian, M\. P\. Kumar, A\. See, S\. Chaudhuri, G\. Holland, A\. Davies, S\. Nowozin, P\. Kohli, and M\. Balog \(2025\)AlphaEvolve: a coding agent for scientific and algorithmic discovery\.arXiv preprint arXiv:2506\.13131\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- Open Science Collaboration \(2015\)Estimating the reproducibility of psychological science\.Science349\(6251\),pp\. aac4716\.External Links:[Document](https://dx.doi.org/10.1126/science.aac4716),ISSN 1095\-9203,[Link](https://doi.org/10.1126/science.aac4716)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p3.1)\.
- OpenAI \(2026a\)Data controls in the openai platform\.Note:[https://platform\.openai\.com/docs/guides/your\-data](https://platform.openai.com/docs/guides/your-data)Accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1),[Table 4](https://arxiv.org/html/2606.26130#A1.T4.3.1.1.1.2.1.3.1.1)\.
- OpenAI \(2026b\)OpenAI services agreement\.Note:[https://openai\.com/policies/services\-agreement/](https://openai.com/policies/services-agreement/)Effective 2026\-01\-01; accessed 2026\-06\-01Cited by:[§A\.2](https://arxiv.org/html/2606.26130#A1.SS2.p2.1),[Table 4](https://arxiv.org/html/2606.26130#A1.T4.3.1.1.1.2.1.3.1.1)\.
- K\. Pearson \(1900\)On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling\.Philosophical Magazine50\(302\),pp\. 157–175\.External Links:[Document](https://dx.doi.org/10.1080/14786440009463897)Cited by:[§C\.5](https://arxiv.org/html/2606.26130#A3.SS5.p1.3)\.
- B\. Qi, K\. Zhang, H\. Li, K\. Tian, S\. Zeng, Z\. Chen, and B\. Zhou \(2023\)Large language models are zero shot hypothesis proposers\.arXiv preprint arXiv:2311\.05965\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- S\. Schmidgall, Y\. Su, Z\. Wang, X\. Sun, J\. Wu, X\. Yu, J\. Liu, M\. Moor, Z\. Liu, and E\. Barsoum \(2025\)Agent laboratory: using LLM agents as research assistants\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 5977–6043\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.320/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.320),ISBN 979\-8\-89176\-335\-7Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- SeatGeek \(2023\)Thefuzz: fuzzy string matching in Python\.Note:[https://github\.com/seatgeek/thefuzz](https://github.com/seatgeek/thefuzz)Software libraryCited by:[§A\.6](https://arxiv.org/html/2606.26130#A1.SS6.SSS0.Px2.p1.5)\.
- C\. E\. Shannon \(1948\)A mathematical theory of communication\.The Bell system technical journal27\(3\),pp\. 379–423\.Cited by:[§A\.7](https://arxiv.org/html/2606.26130#A1.SS7.p2.5)\.
- P\. E\. Shrout and J\. L\. Fleiss \(1979\)Intraclass correlations: uses in assessing rater reliability\.Psychological Bulletin86\(2\),pp\. 420–428\.Cited by:[§C\.11](https://arxiv.org/html/2606.26130#A3.SS11.p1.5)\.
- C\. Si, D\. Yang, and T\. Hashimoto \(2025\)Can llms generate novel research ideas? a large\-scale human study with 100\+ nlp researchers\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://proceedings.iclr.cc/paper_files/paper/2025/hash/ea94957d81b1c1caf87ef5319fa6b467-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov,et al\.\(2025\)OpenAI GPT\-5 system card\.arXiv preprint arXiv:2601\.03267\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.03267),[Link](https://arxiv.org/abs/2601.03267)Cited by:[1st item](https://arxiv.org/html/2606.26130#A1.I2.i1.p1.1),[§1](https://arxiv.org/html/2606.26130#S1.p4.1)\.
- M\. D\. Skarlinski, S\. Cox, J\. M\. Laurent, J\. D\. Braza, M\. Hinks, M\. J\. Hammerling, M\. Ponnapati, S\. G\. Rodriques, and A\. D\. White \(2024\)Language agents achieve superhuman synthesis of scientific knowledge\.arXiv preprint arXiv:2409\.13740\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1),[§3](https://arxiv.org/html/2606.26130#S3.SS0.SSS0.Px1.p1.1)\.
- C\. Spearman \(1904\)The proof and measurement of association between two things\.American Journal of Psychology15\(1\),pp\. 72–101\.External Links:[Document](https://dx.doi.org/10.2307/1412159)Cited by:[Table 2](https://arxiv.org/html/2606.26130#S2.T2)\.
- R\. J\. Tibshirani and B\. Efron \(1993\)An introduction to the bootstrap\.Monographs on statistics and applied probability57\(1\),pp\. 1–436\.Cited by:[§C\.1](https://arxiv.org/html/2606.26130#A3.SS1.p2.16),[§C\.6](https://arxiv.org/html/2606.26130#A3.SS6.SSS0.Px2.p1.8)\.
- F\. Villaescusa\-Navarro, B\. Bolliet, P\. Villanueva\-Domingo, A\. E\. Bayer, A\. Acquah, C\. Amancharla, A\. Barzilay\-Siegal, P\. Bermejo, C\. Bilodeau, P\. C\. Ramírez,et al\.\(2025\)The denario project: deep knowledge ai agents for scientific discovery\.arXiv preprint arXiv:2510\.26887\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- Q\. Wang, D\. Downey, H\. Ji, and T\. Hope \(2024\)SciMON: scientific inspiration machines optimized for novelty\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 279–299\.External Links:[Link](https://aclanthology.org/2024.acl-long.18/),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.18)Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.
- Y\. Yamada, R\. T\. Lange, C\. Lu, S\. Hu, C\. Lu, J\. Foerster, J\. Clune, and D\. Ha \(2025\)The AI scientist\-v2: workshop\-level automated scientific discovery via agentic tree search\.arXiv preprint arXiv:2504\.08066\.Cited by:[§1](https://arxiv.org/html/2606.26130#S1.p1.1)\.

## Appendix AMethods

### A\.1Data collection

We retrieved 1,000 papers from arXiv by running an automated collection script on December 4, 2025\. The script queried arXiv for Computer Science \(CS\) papers with “LLM” or “Large Language Model” in the title, appended a runtimesubmittedDatefilter covering the preceding 182 days \(approximately June–December 2025\), sorted results by submission date, filtered the returned records to papers published in 2025, and then randomly sampled 1,000 papers from that pool\. We chose this time frame so that the papers would post\-date the known training\-data cutoffs of GPT\-5\.1 \(September 30, 2024\) and Gemini 3 Pro \(January 2025\), yielding suggestions that are unlikely to rely on knowledge of the papers themselves\. No official cutoff has been published for DeepSeek\-V3\.2, so prior exposure cannot be ruled out for that model\. We filtered for papers in the Computer Science \(CS\) category with “LLM” or “Large Language Model” in the title, using the core query:

```
(ti:"LLM*" OR ti:"Large Language Model*") AND cat:cs.*
```

with the date\-window constraint added programmatically at runtime\. The subcategories \(e\.g\., Computation and Language, Artificial Intelligence, Computer Vision\) were left unrestricted within the CS main category\. For each paper, we downloaded the arXiv\-hosted PDF and extracted full text using PyMuPDF\. We also recorded the license URI declared in the arXiv metadata for each paper\. The per\-paper license metadata are included in the released derived\-data table, and the aggregate license distribution is reported in[Table˜3](https://arxiv.org/html/2606.26130#A1.T3)\.

### A\.2Copyright, licensing, and text\-and\-data mining

We did not treat availability on arXiv, nor the arXiv default perpetual non\-exclusive distribution license, as a general public license to redistribute or republish full text\. arXiv states that e\-prints remain subject to copyright protection, that redistribution requires permission unless a permissive license applies, and that its default non\-exclusive license gives arXiv limited distribution rights while limiting reuse by other entities or individuals\(arXiv,[2026a](https://arxiv.org/html/2606.26130#bib.bib80),[b](https://arxiv.org/html/2606.26130#bib.bib81)\)\. For this reason, we restricted our handling of the PDFs to non\-public computational processing\. This distinction is also consistent with arXiv’s API terms, which permit users to retrieve, store, and use arXiv e\-print content for research purposes, while prohibiting users from storing and serving PDFs, source files, or other e\-print content from their own servers unless authorised by the copyright holder or permitted by the e\-print’s license\(arXiv,[2026b](https://arxiv.org/html/2606.26130#bib.bib81)\)\.

The full\-text collection and processing were conducted at Vrije Universiteit Brussel, a Belgian research organisation, for scientific research purposes, using works lawfully accessible through arXiv\. We relied on the scientific text\-and\-data\-mining exception in Article 3 of Directive \(EU\) 2019/790, as transposed into Belgian law, specifically Article XI\.191/1, § 1, 7∘of the Belgian Code of Economic Law, which permits reproductions and extractions by research organisations for scientific TDM over works to which they have lawful access and requires retained copies to be stored with an appropriate level of security\(European Parliament and Council of the European Union,[2019](https://arxiv.org/html/2606.26130#bib.bib82); Belgian Federal Public Service Economy,[2022](https://arxiv.org/html/2606.26130#bib.bib83)\)\. The Directive defines research organisations to include universities, defines TDM as automated analysis of digital text/data to generate information, treats freely available online content as lawful access, and recognises that research organisations may rely on private partners or their technological tools to carry out TDM\(European Parliament and Council of the European Union,[2019](https://arxiv.org/html/2606.26130#bib.bib82)\)\. Article 4’s reservation mechanism for general TDM does not affect Article 3 scientific TDM, and contractual provisions contrary to the Article 3 exception are unenforceable under Article 7\(European Parliament and Council of the European Union,[2019](https://arxiv.org/html/2606.26130#bib.bib82)\)\. Third\-party API processing followed the controls summarised in[Section˜A\.5](https://arxiv.org/html/2606.26130#A1.SS5): OpenAI API submissions were not used for model training or improvement absent opt\-in; Anthropic API/commercial use was not used for model training absent opt\-in, and model\-improvement was disabled where applicable for Claude Code; and Gemini calls were made from VUB/Belgium/EEA, where Google’s Gemini API terms apply paid\-services data\-use protections to all Gemini API/AI Studio services, including unpaid quota\(OpenAI,[2026a](https://arxiv.org/html/2606.26130#bib.bib84),[b](https://arxiv.org/html/2606.26130#bib.bib85); Google,[2026](https://arxiv.org/html/2606.26130#bib.bib86); Anthropic,[2026a](https://arxiv.org/html/2606.26130#bib.bib88),[b](https://arxiv.org/html/2606.26130#bib.bib89)\)\. The TDM outputs analysed and released in this study are derived annotations and aggregate statistics\. We do not redistribute arXiv PDFs, extracted full texts, substantial expressive excerpts from the papers, or API batch/input/output files containing paper full text\.

Table 3:arXiv license distribution in the 1,000\-paper corpus\.License metadata were retrieved from the arXiv metadata record for each paper\. The table is reported for transparency and to delimit downstream reuse\. Inclusion in the computational analysis did not rely on permissive Creative Commons licensing, but on the scientific TDM basis described in[Section˜A\.2](https://arxiv.org/html/2606.26130#A1.SS2)\.
### A\.3Paper\-side entity extraction and research question generation

For each of these 1,000 papers, we prompted GPT\-5\.1 with the title, abstract, and full text to generate a main research question \(RQ\) and to extract the paper\-side reference entities, under the copyright/TDM and API data\-handling procedures described in[Sections˜A\.2](https://arxiv.org/html/2606.26130#A1.SS2)and[A\.5](https://arxiv.org/html/2606.26130#A1.SS5)\. This was an inference\-only use of the OpenAI Batch API\. The submitted texts were not used for model training or fine\-tuning, and the OpenAI batch input/output file objects under our control were deleted after processing\. The resulting RQ was then used as the input prompt for the suggestion step described below\.

#### Paper\-side entity extraction\.

We prompted GPT\-5\.1 via the OpenAI Batch API, providing it with the title, abstract, and full text, to extract:

1. 1\.Datasetsused in the paper’s experiments \(excluding those only mentioned in related work\)\.
2. 2\.Modelsand architectures used experimentally\.
3. 3\.Evaluation metricsreported in experiments\.

We requested JSON with aresearch\_questionfield andGroundTruth\.\.\.entity lists, with each entity name limited to 1–3 words for consistency\. In the saved raw batch outputs, some responses instead used genericdatasets/models/metricskeys; downstream parsing harmonized both key variants into the standardized paper\-side schema used for analysis\. The full extraction prompt is provided in[Appendix˜F](https://arxiv.org/html/2606.26130#A6)\.

#### LLM suggestion generation\.

Using the generated RQ, we prompted three LLMs to suggest methodology components given only the research question\. Gemini 3 Pro and DeepSeek\-V3\.2 did not receive the arXiv PDFs, abstracts, or extracted full texts:

- •GPT\-5\.1\(OpenAI\), exact experimental model IDgpt\-5\.1\-2025\-11\-13, documented in the GPT\-5 system card\(Singhet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib57)\), knowledge cutoff September 30, 2024, accessed via the OpenAI Batch API \(/v1/responsesendpoint, reasoning effort set tomedium\)\.
- •Gemini 3 Pro\(Google\), model IDgemini\-3\-pro\-preview, knowledge cutoff January 2025, accessed via the Google GenAI Batch API\.
- •DeepSeek\-V3\.2\(DeepSeek\)\(DeepSeek\-AIet al\.,[2025](https://arxiv.org/html/2606.26130#bib.bib74); DeepSeek,[2025](https://arxiv.org/html/2606.26130#bib.bib75)\), accessed via the DeepSeek API aliasmodel="deepseek\-reasoner"on the OpenAI\-compatible/v1/chat/completionsendpoint\. DeepSeek announced DeepSeek\-V3\.2 as live on its App, Web, and API on December 1, 2025, and documented V3\.2 support for thinking/tool\-use mode\. Our API calls ran between December 2025 and January 2026\. Because DeepSeek compatibility aliases are time\-dependent and may later be remapped on the live Models & Pricing page, we report the version used at execution time according to our run logs and API configuration\.

Each model was asked to suggest suitable datasets, models, evaluation metrics, and a structured experimental pipeline describing how these components would be used together to address the RQ\. Responses were collected in JSON format: GPT\-5\.1 via the OpenAI Batch API, Gemini 3 Pro via the Google GenAI Batch API, and DeepSeek\-V3\.2 via asynchronous calls to the standard DeepSeek/v1/chat/completionsendpoint\. All API calls were executed between December 2025 and January 2026\. The suggestion prompt, identical across all three models except for LLM\-specific JSON key prefixes, is provided in[Appendix˜G](https://arxiv.org/html/2606.26130#A7)\.

### A\.4Comparison scope and analysis denominators

The comparison is intentionally asymmetric: the paper\-derived reference inventory is an automated extraction of the datasets, models, and metrics used experimentally in each paper \(with extraction reliability quantified by ICC\(2,1\)=0\.469\(2,1\)=0\.469;[Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\), whereas the suggestion prompt asks for a short set of suitable first\-pass choices and a concise “straightforward” pipeline from the research question alone\. The analysis therefore targets compression and reweighting of first\-pass methodological menus rather than exact reconstruction of full paper\-level experimental inventories\. Throughout the narrative and in visible table labels we use “paper\-derived reference inventory” \(or “reference inventory”, “Reference”\); the tag “GT” or “Ground Truth” persists only in pipeline artefacts \(JSON schema keys such asGroundTruthDatasets, CSV column headers, filenames, and a small number of pre\-generated figure panels\) that would be fragile to rename\. This reference corpus is itself an automated extraction with only moderate inter\-model agreement on counts \(ICC\(2,1\)=0\.469\(2,1\)=0\.469;[Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\. The measured object is therefore the compound pipelinepaper→\\rightarrowGPT\-5\.1\-generated research question→\\rightarrowmodel suggestion→\\rightarrowGPT\-5\.1\-derived ontology, not an independent model\-only benchmark\. Our prompt explicitly asks for “one or more” specific names and a “structured and short” straightforward pipeline \([Appendix˜G](https://arxiv.org/html/2606.26130#A7)\), so the observed compression likely reflects both model\-side methodological priors and a concise\-assistant elicitation frame\. Unless otherwise noted, normalised entity\-name analyses use the full 1,000\-paper outputs for each source\. Taxonomy\-based analyses use the classified files available at analysis time \(n=1,000=1\{,\}000reference, 1,000 GPT\-5\.1, 981 Gemini 3 Pro, 998 DeepSeek\-V3\.2\), and analyses requiring non\-empty entities in both sources use smaller intersections reported in the corresponding captions and tables\.

### A\.5Third\-party API data handling

All third\-party model calls were inference\-only calls\. We did not fine\-tune, train, or ask any provider to train models on the arXiv corpus\. Provider\-side file objects under our control were deleted after processing, while provider security, abuse\-monitoring, or safety logs may be retained according to the applicable provider terms\. Table[4](https://arxiv.org/html/2606.26130#A1.T4)summarises the data submitted to each provider and the relevant data\-handling controls\. No provider was granted any independent right to redistribute the corpus, to expose full texts to users, or to use the paper full texts for model training\.

Table 4:Third\-party API/model data handling\.The table distinguishes full\-text processing from downstream suggestion and validation calls\.
### A\.6Entity normalisation

Raw entity names from different sources exhibit surface\-level variation \(e\.g\., “GPT\-4o” vs\. “GPT4o”, “F1\-score” vs\. “F1”\)\. We applied a two\-stage normalisation:

#### Stage 1: Deterministic normalisation\.

We applied a sequence of regex\-based transformations: removing bracketed content, replacing separators \(/,;,\|\) with commas, converting underscores and hyphens to spaces, stripping remaining punctuation, collapsing whitespace, and lowercasing\. For evaluation metrics, manual inspection of the raw outputs revealed systematic sub\-variant proliferation, so we applied additional hardcoded aggregation rules: all ROUGE variants \(rouge\-1, rouge\-2, rouge\-L\) map to “rouge”, all BLEU variants to “bleu”, Pearson and Spearman correlation variants \(including common misspellings\) to their canonical forms, and all F1 score variants to “f1”\. Leading numeric prefixes \(e\.g\., “1accuracy”\) were stripped\.

#### Stage 2: Fuzzy clustering\.

We clustered the deterministically normalised names using a greedy representative\-matching procedure\. Names are sorted by length \(longest first\) so that more specific names tend to become early cluster representatives\. For each name, we compute the token\-sort ratio against all existing cluster representatives\. The token\-sort ratio tokenises both strings, sorts the tokens alphabetically, and computes the normalised Levenshtein edit distance\(Levenshtein,[1966](https://arxiv.org/html/2606.26130#bib.bib73)\)on a\[0,100\]\[0,100\]scale \(implemented inthefuzz\(SeatGeek,[2023](https://arxiv.org/html/2606.26130#bib.bib76)\)\)\. If any representative scores≥90\\geq 90, the name joins the best\-scoring cluster; otherwise it seeds a new cluster\. A prefix\-safety check zeroes the score whenever one string is a strict prefix of the other \(e\.g\., “gpt4”≠\\neq“gpt4o”\), preventing over\-merging of distinct model variants\. Manual inspection identified one additional hardcoded correction: the concatenated string “gpt4o3mini” was split into its constituent models\. A sensitivity analysis over thresholdsT∈\{80,85,90,95,100\}T\\in\\\{80,85,90,95,100\\\}confirms that the headline vocabulary\-compression and coverage metrics are stable across this range \([Section˜C\.10](https://arxiv.org/html/2606.26130#A3.SS10)\); unless otherwise noted, main\-text entity\-name analyses use the defaultT=90T=90clustering\.

### A\.7Frequency analysis and comparison

For each entity type and each source \(reference inventory, GPT\-5\.1, Gemini 3 Pro, DeepSeek\-V3\.2\), we counted how often each entity appeared\. The reference inventory contains significantly more unique models and evaluation metrics than the LLMs suggest, although the number of unique datasets is more comparable \([Table˜1](https://arxiv.org/html/2606.26130#S2.T1)\)\. To obtain a fair comparison, we normalised the LLM suggestion counts into relative percentages against the reference totals\. We report within\-source percentages \(the entity’s count divided by the total count for that source\) and normalised\-to\-reference percentages \(the entity’s count divided by the total reference count\)\.

Vocabulary compression is quantified using two complementary measures\. The Shannon entropy\(Shannon,[1948](https://arxiv.org/html/2606.26130#bib.bib63)\)of the frequency distribution isH=−∑i=1Spi​ln⁡piH=\-\\sum\_\{i=1\}^\{S\}p\_\{i\}\\ln p\_\{i\}wherepip\_\{i\}is the relative frequency of entityiiamongSStypes \(computed on mention counts, natural logarithm\)\. The effective number of entitiesexp⁡\(H\)\\exp\(H\)\(Jost,[2006](https://arxiv.org/html/2606.26130#bib.bib39)\)represents the number of equally frequent entities producing the same entropy\. The Gini coefficient\(Gini,[1912](https://arxiv.org/html/2606.26130#bib.bib70)\)

G=2​∑i=1ni​x\(i\)n​∑i=1nx\(i\)−n\+1nG=\\frac\{2\\sum\_\{i=1\}^\{n\}i\\,x\_\{\(i\)\}\}\{n\\sum\_\{i=1\}^\{n\}x\_\{\(i\)\}\}\-\\frac\{n\+1\}\{n\}\(1\)for frequency counts sorted in ascending orderx\(1\)≤⋯≤x\(n\)x\_\{\(1\)\}\\leq\\cdots\\leq x\_\{\(n\)\}, measures concentration \(G=0G=0is perfect equality,G=1G=1is maximal concentration\)\.

To quantify rank\-dependent amplification among shared entities, we fitlog10⁡\(riLLM/riref\)=α\+β​log10⁡\(rankiref\)\\log\_\{10\}\(r\_\{i\}^\{\\mathrm\{LLM\}\}/r\_\{i\}^\{\\mathrm\{ref\}\}\)=\\alpha\+\\beta\\,\\log\_\{10\}\(\\mathrm\{rank\}\_\{i\}^\{\\mathrm\{ref\}\}\)by ordinary least squares, whererir\_\{i\}denotes the relative frequency of entityiiand the superscript “ref” refers to the paper\-derived reference inventory\. A positive slopeβ^\\hat\{\\beta\}indicates that LLMs amplify rarer shared entities proportionally more than frequent ones\.

Inter\-source agreement on top\-ranked entities is measured by Jaccard@KK\(Jaccard,[1912](https://arxiv.org/html/2606.26130#bib.bib69)\):J​@​K=\|TK\(a\)∩TK\(b\)\|/\|TK\(a\)∪TK\(b\)\|J@K=\|T\_\{K\}^\{\(a\)\}\\cap T\_\{K\}^\{\(b\)\}\|\\,/\\,\|T\_\{K\}^\{\(a\)\}\\cup T\_\{K\}^\{\(b\)\}\|, whereTK\(s\)T\_\{K\}^\{\(s\)\}is the set ofKKmost frequent entities in sourcess\(ties broken alphabetically\)\. We reportK=20K=20throughout\.

### A\.8Taxonomy classification

We prompted GPT\-5\.1 via the Batch API to classify all reference\-inventory and LLM\-suggested entities according to structured category dimensions:

Datasetswere classified along the following dimensions: modality \(text, image, audio, video, multimodal, …\), task type \(classification, QA, generation, reasoning, …\), domain \(general, scientific, healthcare, legal, …\), annotation type \(supervised, semi\-supervised, crowdsourced, …\), size \(small:<<10K samples, medium: 10K–100K, large:\>\>100K\), granularity \(document, sentence, token, …\), linguistic scope \(monolingual, multilingual, cross\-lingual\), cognitive/affective properties \(reasoning, emotion, decision making, …\), and data quality \(noisy, curated\)\.

Modelswere classified by architecture \(Transformer, CNN, RNN/LSTM, GNN, …\), training paradigm \(supervised, self\-supervised, few\-shot, fine\-tuning, RAG, …\), provider \(OpenAI, Meta AI, Google DeepMind, Anthropic, …\), openness \(open or closed\), and size \(small:<<1B, medium: 1–10B, large: 10–100B, extra\-large:\>\>100B parameters\)\. The classifier itself uses this four\-way model\-size schema\. In the downstream comparison helpers used for the summary size analyses, however, model size is normalised to three analysis bins by folding Extra\-large into Large; the size percentages reported in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)should therefore be read as a combined top\-end bucket rather than as the 10–100B band alone\.

Metricswere classified by evaluation type \(accuracy, ranking, regression, fairness, safety, efficiency, robustness, …\)\. The full classification prompt with all allowed values for each dimension is provided in[Appendix˜H](https://arxiv.org/html/2606.26130#A8)\. All classifier settings use this same schema\. Unless otherwise noted, the main category results in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)use the entity\-list\-only setting, in which the classifier sees only entity lists\.[Appendix˜B](https://arxiv.org/html/2606.26130#A2)quantifies how labels change when the generated experimental pipeline or web search are added, whereas[Appendix˜C](https://arxiv.org/html/2606.26130#A3)recomputes the summary robustness tables when the classifier also sees pipeline context\. The co\-occurrence analyses use exports derived from that same richer setting\.

### A\.9Co\-occurrence analysis

We computed pairwise co\-occurrence frequencies between taxonomy dimensions to assess whether LLMs reproduce the combinatorial structure of real research methodology\. These co\-occurrence tables come from the setting where the classifier also sees the generated pipeline, so they are not a direct reuse of the entity\-list\-only category distributions in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)\. We analyse all 12 cross\-entity co\-occurrence pairs: 9 dataset–model pairs \(task type, size, and cognitive/affective×\\timesarchitecture, provider, and openness\) and 3 metric–model pairs \(evaluation type×\\timesarchitecture, provider, and openness\)\. For each pair of categories, we computed the percentage of papers \(within each source\) in which both categories co\-occur\. In∼\{\\sim\}0\.04% of classifier outputs, a comma\-delimited multi\-label string \(e\.g\., “Retrieval, Reasoning”\) was returned as a single field value rather than as separate items; these conjunction labels are carried through as atomic categories in the co\-occurrence matrices and residual heatmaps, but their low frequency does not materially affect the results\. The main text focuses on the two pairs with the clearest interpretive value, but the evaluation type×\\timesprovider pair is the more strongly validated of the two because both axes are comparatively reliable in the blinded audit \(though evaluation type’s moderateκ\\kapparests on the LLM stratum; the reference\-inventory stratum hasκ=0\.061\\kappa=0\.061onn=17n=17rows\); task type×\\timesprovider remains useful descriptive context\. The displayed heatmaps are restricted to the top 7 row categories and top 7 provider columns selected by the largest row\-wise and column\-wise cell maxima after summing the aligned source matrices for each focal pair, whereas[Figure˜D2](https://arxiv.org/html/2606.26130#A4.F2)and the residual analyses use the full aligned matrices subject to their stated support filters\. For each co\-occurrence pair, the row\-wise JSD is computed as follows: for each row categoryrrshared between both the reference\-inventory and LLM matrices, column counts are normalised to probability vectors and the JSD \([Equation˜2](https://arxiv.org/html/2606.26130#A3.E2), base\-2\) is computed; the reported mean is the average over all shared row categories\. Residual analysis \([Section˜C\.5](https://arxiv.org/html/2606.26130#A3.SS5)\) complements this by separating structural co\-occurrence shifts from simple overall frequency changes\.

## Appendix BAblation studies

This appendix is organized to separate checks about the labeling step from checks about whether the main conclusions survive alternative assumptions\.[Section˜B\.1](https://arxiv.org/html/2606.26130#A2.SS1)tests whether giving the taxonomy classifier the suggested pipeline changes the resulting overall category distributions, whereas[Section˜B\.2](https://arxiv.org/html/2606.26130#A2.SS2)tests whether web search mainly affects how GPT\-5\.1’s suggestions are labeled\. These analyses therefore tell us how sensitive the labeling step is; they do not mean the LLMs changed which entities they suggested\.

### B\.1Pipeline context ablation

In the suggestion step, each LLM produces not only entity lists but also a free\-form experimental pipeline\. We tested whether giving that extra context to the taxonomy classifier changes the resulting distributions by comparing two conditions: \(1\) classifier sees the suggested entities and the experimental pipeline; and \(2\) classifier sees only the entity lists\. The entity\-list\-only condition is the baseline used for the main category figures in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2); this appendix quantifies how far the distributions move under the richer labeling setting\.

[Figure˜D3](https://arxiv.org/html/2606.26130#A4.F3)shows the percentage\-point change across dataset category dimensions for all three LLMs\. The most consistent pattern across models is an increase in “generation” task\-type classification when pipeline context is provided:\+3\.6\+3\.6percentage points for GPT\-5\.1,\+3\.1\+3\.1for Gemini 3 Pro, and\+3\.5\+3\.5for DeepSeek\-V3\.2\. Document\-level granularity also increases consistently \(\+3\.2\+3\.2for GPT\-5\.1,\+2\.6\+2\.6for Gemini 3 Pro,\+2\.6\+2\.6for DeepSeek\-V3\.2; not shown in[Figure˜D3](https://arxiv.org/html/2606.26130#A4.F3)\)\.

For model taxonomy, the largest shift is in architecture classification: with pipeline context, GPT\-5\.1’s “Generative” category increases from 18\.4% to 23\.8% \(\+5\.3\+5\.3percentage points\) while “Transformer \(all\)” decreases from 74\.6% to 69\.6% \(−5\.0\-5\.0pp\)\. This suggests that for GPT\-5\.1, pipeline context helps the classifier disambiguate between the broader Transformer family and more specific generative model types\. Gemini 3 Pro shows smaller but directionally consistent shifts \(\+\+1\.8 pp Generative,−\-1\.7 pp Transformer\), whereas DeepSeek\-V3\.2 shows small shifts in the opposite direction \(−\-1\.4 pp Generative,\+\+1\.6 pp Transformer\), indicating that pipeline context does not uniformly disambiguate architecture labels across models\.[Table˜E1](https://arxiv.org/html/2606.26130#A5.T1)reports all model and metric subcategories where at least one LLM exceeds\|Δ​pp\|≥2\.0\|\\Delta\\text\{pp\}\|\\geq 2\.0; only the Generative/Transformer pair crosses this threshold, confirming that pipeline context has minimal effect on model and metric taxonomy distributions\. For provider specifically, pipeline context reclassifies a small number of models from Other/Academic to specific commercial providers \(GPT\-5\.1: 19/350, Gemini 3 Pro: 17/414, DeepSeek\-V3\.2: 8/268 models reclassified\), but the overall impact on provider distributions is below 0\.5 pp per category\.

[Table˜E2](https://arxiv.org/html/2606.26130#A5.T2)summarises the ablation conditions across models\.

### B\.2Web search ablation

For GPT\-5\.1 only, we compared taxonomy classification with and without web search enabled\. The most substantial effect is on model architecture classification: the “Generative” category increases from 18\.4% to 35\.4% with web search \(\+17\.0\+17\.0pp\), while “Transformer \(all\)” drops from 74\.6% to 58\.4% \(−16\.2\-16\.2pp\)\. This shift indicates that architecture labels are sensitive to label operationalisation and retrieval context, in the sense that the classifier’s access to web information changes how it categorises the same set of suggested models, rather than reflecting a change in the LLMs’ underlying suggestion behaviour\. Architecture\-based findings throughout the paper should therefore be interpreted cautiously, as partly reflecting classifier sensitivity rather than stable properties of the suggestions themselves\.

Training paradigm distributions also shift: few\-shot and zero\-shot learning both decrease \(from∼14\.5%\{\\sim\}14\.5\\%each to∼7\.9%\{\\sim\}7\.9\\%\), while multi\-task learning increases from 7\.4% to 11\.8%\. Provider distributions remain largely stable \(<<2 percentage point changes\), and model size distributions are minimally affected\. These results suggest that web search primarily helps the classifier make finer\-grained architecture distinctions rather than changing the fundamental distribution of suggested entities\.[Figure˜D4](https://arxiv.org/html/2606.26130#A4.F4)visualises the effect across dataset, model, and metric dimensions\.

## Appendix CStatistical robustness

The preceding ablation studies \([Appendix˜B](https://arxiv.org/html/2606.26130#A2)\) test the robustness of the taxonomy classification step\. Here we report complementary robustness checks on the comparison methodology itself: effect sizes, popularity baselines, comparisons that equalize entity counts per paper, co\-occurrence checks that separate structure from overall frequency shifts, paper\-by\-paper similarity tests, filtering of likely paper\-specific entities \([Section˜C\.7](https://arxiv.org/html/2606.26130#A3.SS7)\), broader provider groupings \([Section˜C\.8](https://arxiv.org/html/2606.26130#A3.SS8)\), looser matching rules for model families \([Section˜C\.9](https://arxiv.org/html/2606.26130#A3.SS9)\), normalisation\-threshold sensitivity \([Section˜C\.10](https://arxiv.org/html/2606.26130#A3.SS10)\), and a blinded cross\-model audit of the extraction, classification, normalisation, and the step that asks whether an entity is pre\-existing or paper\-specific \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\)\. Unless stated otherwise, the tables in this appendix are recomputed in the setting where the classifier also sees the generated pipeline and therefore function as robustness checks for the conclusions of[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2), not as exact numeric copies of the main\-text figures\.

The section is organized by claim\.[Sections˜C\.1](https://arxiv.org/html/2606.26130#A3.SS1),[C\.2](https://arxiv.org/html/2606.26130#A3.SS2),[C\.3](https://arxiv.org/html/2606.26130#A3.SS3)and[C\.4](https://arxiv.org/html/2606.26130#A3.SS4)quantify the main taxonomy results, especially the strength and interpretation of provider divergence\.[Sections˜C\.5](https://arxiv.org/html/2606.26130#A3.SS5)and[C\.6](https://arxiv.org/html/2606.26130#A3.SS6)address the structural and question\-specificity claims that support the co\-occurrence results\.[Sections˜C\.7](https://arxiv.org/html/2606.26130#A3.SS7),[C\.8](https://arxiv.org/html/2606.26130#A3.SS8),[C\.9](https://arxiv.org/html/2606.26130#A3.SS9)and[C\.10](https://arxiv.org/html/2606.26130#A3.SS10)test whether vocabulary compression and provider concentration survive stricter or looser matching assumptions and changes to the fuzzy\-clustering threshold\. Finally,[Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)audits the extraction, classification, normalisation, and established\-versus\-paper\-specific labeling steps that underlie the automated pipeline\.

### C\.1Effect sizes and bootstrapped confidence intervals

The Jensen–Shannon divergence\(Lin,[1991](https://arxiv.org/html/2606.26130#bib.bib64)\)between discrete distributionsPPandQQis

JSD​\(P∥Q\)=12​DKL​\(P∥M\)\+12​DKL​\(Q∥M\),M=12​\(P\+Q\)\\mathrm\{JSD\}\(P\\\|Q\)=\\tfrac\{1\}\{2\}\\,D\_\{\\mathrm\{KL\}\}\(P\\\|M\)\+\\tfrac\{1\}\{2\}\\,D\_\{\\mathrm\{KL\}\}\(Q\\\|M\),\\quad M=\\tfrac\{1\}\{2\}\(P\+Q\)\(2\)whereDKLD\_\{\\mathrm\{KL\}\}is the Kullback–Leibler divergence computed with base\-2 logarithms, yielding JSD in bits bounded by\[0,1\]\[0,1\]\. Cramér’sVV\(Cramér,[1946](https://arxiv.org/html/2606.26130#bib.bib66)\)quantifies the strength of association in ar×cr\\times ccontingency table:

V=χ2n⋅\(min⁡\(r,c\)−1\)V=\\sqrt\{\\frac\{\\chi^\{2\}\}\{n\\cdot\(\\min\(r,c\)\-1\)\}\}\(3\)wherennis the sample size andχ2\\chi^\{2\}the chi\-square statistic\. As a rough guide,V<0\.10V<0\.10is conventionally considered small andV\>0\.30V\>0\.30large\(Cohen,[1988](https://arxiv.org/html/2606.26130#bib.bib67)\), though these thresholds are approximate whenmin⁡\(r,c\)\>2\\min\(r,c\)\>2\.

[Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)reports JSD with bootstrapped 95% confidence intervals\(Tibshirani and Efron,[1993](https://arxiv.org/html/2606.26130#bib.bib71)\)\(2,000 paper\-level resamples, percentile method, seed==42\) and Cramér’sVVfor all 15 taxonomy dimensions×\\times3 LLMs in the same richer labeling setting where the classifier also sees the generated pipeline\. Provider exhibits the largest effect size \(V=0\.33V=0\.33–0\.350\.35, medium\-to\-large\) and the largest JSD \(0\.1010\.101–0\.1380\.138\), with narrow confidence intervals confirming precise estimation\. Most other dimensions haveV<0\.15V<0\.15\(small effect\)\. Dataset size \(V=0\.16V=0\.16–0\.210\.21\) and model architecture \(V=0\.13V=0\.13–0\.170\.17\) are, by Cramér’sVV, the only other dimensions approaching medium effect sizes\. Note that the ranking byVVdiffers slightly from the no\-pipeline JSD ranking in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2), where dataset size and model size rank second and third; this reflects the sensitivity of each measure to sample size and number of categories\. Because the underlying contingency tables tabulate label instances that are multi\-label and nested within papers and entities, theχ2\\chi^\{2\}p\-values reported in[Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)are descriptive rather than fully inferential; statistical inference rests on the JSD bootstrap CIs \(paper\-level resampling\), Cramér’sVV, and the robustness analyses in[Sections˜C\.2](https://arxiv.org/html/2606.26130#A3.SS2),[C\.3](https://arxiv.org/html/2606.26130#A3.SS3),[C\.4](https://arxiv.org/html/2606.26130#A3.SS4),[C\.7](https://arxiv.org/html/2606.26130#A3.SS7),[C\.8](https://arxiv.org/html/2606.26130#A3.SS8),[C\.9](https://arxiv.org/html/2606.26130#A3.SS9)and[C\.10](https://arxiv.org/html/2606.26130#A3.SS10)\.

### C\.2Popularity baseline comparison

Using the same richer labeling setting, we test whether LLM category distributions can be explained by simply recommending popular entities by comparing each LLM’s divergence from the reference inventory against two popularity baselines \([Table˜E4](https://arxiv.org/html/2606.26130#A5.T4)\)\. For each paper, we construct a leave\-one\-out popularity distribution from all other reference papers \(excluding the target paper to prevent corpus leakage\) and “suggest”kkentities, wherekkmatches the number of entities that paper’s LLM suggested\. The deterministic top\-kkbaseline always assigns thekkmost frequent entities; the stochastic sampled baseline drawskkentities without replacement with probability proportional to reference\-corpus frequency \(seed==42\)\.[Table˜E4](https://arxiv.org/html/2606.26130#A5.T4)reports the three dimensions emphasized in the main text: model provider, model size, and metric evaluation type\. In all three cases, LLMs produce substantially lower JSD than the deterministic top\-kkbaseline, demonstrating that their suggestions are conditioned on the research question rather than defaulting to a fixed popularity ranking\. The stochastic sampled baseline achieves near\-zero JSD because drawing entities proportionally to reference\-corpus frequency naturally reconstructs the aggregate taxonomy distribution; the LLMs’ higher divergence relative to this floor indicates systematic distributional biases beyond what popularity sampling would produce\.

### C\.3Local retrieval calibration baseline

As a complementary calibration, we replace global popularity with a simple content\-conditioned non\-generative adviser\. For each target paper, we build a leave\-one\-out BM25 neighborhood over the generated research questions and retrieve the topN=25N=25most similar reference papers within the same reference–LLM shared subset\. We then aggregate dataset, model, or metric names across that local neighborhood and emit the deterministic topkkentities, wherekkmatches the number of entities suggested by the LLM for that paper\.[Table˜E5](https://arxiv.org/html/2606.26130#A5.T5)reports the same three dimensions emphasized in the main text\. This local retrieval baseline is far closer than global top\-kkpopularity on all three dimensions, showing that question\-conditioned lexical retrieval already captures a large share of the signal available under the sparse research\-question\-only input\. However, it remains more divergent than the LLMs on model provider \(BM25 JSD==0\.245–0\.290 versus 0\.101–0\.138 for the LLMs\), model size \(0\.071–0\.089 versus 0\.014–0\.036\), and metric evaluation type \(0\.059–0\.084 versus 0\.011–0\.021\)\. The calibration therefore sharpens, rather than replaces, the popularity result: the main findings are not reducible either to global popularity or to a simple question\-conditioned lexical retriever\.

### C\.4Equal\-count comparisons

Again in the same richer labeling setting, LLMs and the reference inventory may differ in the number of entities per paper, which could inflate divergence estimates\. These matched\-count comparisons use the reference–LLM shared\-paper subsets for the relevant entity type \(datasets n=904/891/910=904/891/910, models n=942/928/948=942/928/948, metrics n=937/924/943=937/924/943for GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2\)\. To control for this, we subsample the larger entity set to match the smaller per paper \(1,000 iterations, seed==42\) and recompute JSD on the matched distributions\. Matched JSD values are slightly lower than full estimates but preserve the same divergence hierarchy: provider remains the dominant model dimension under matched counts \(0\.088–0\.112 across LLMs\), followed by model size \(0\.011–0\.028\)\.

### C\.5Co\-occurrence structure after accounting for overall frequencies

To disentangle structural co\-occurrence changes from marginal frequency shifts, we compute Pearson residuals for each cell of the co\-occurrence matrix\. We first exclude rows and columns with marginal totals below 5 to ensure adequate expected cell counts\. For the remaining cells, the expected count under independence is

Ei​j=Ri​Cj/NE\_\{ij\}=R\_\{i\}\\,C\_\{j\}\\,/\\,N\(4\)whereRiR\_\{i\}andCjC\_\{j\}are the row and column marginal totals andNNis the grand total\. The Pearson residual\(Pearson,[1900](https://arxiv.org/html/2606.26130#bib.bib72)\)is

ei​j=\(Oi​j−Ei​j\)/Ei​je\_\{ij\}=\(O\_\{ij\}\-E\_\{ij\}\)\\,/\\,\\sqrt\{E\_\{ij\}\}\(5\)which is approximately standard\-normal under the null hypothesis of independence\. Cell\-levelpp\-values are obtained from the two\-sided normal distribution and corrected for multiple comparisons using the Benjamini–Hochberg procedure\(Benjamini and Hochberg,[1995](https://arxiv.org/html/2606.26130#bib.bib65)\)atα=0\.05\\alpha=0\.05, applied independently within each co\-occurrence matrix\. Note that marginal filtering on totals≥5\\geq 5does not guaranteeEi​j≥5E\_\{ij\}\\geq 5in every cell, so the normal approximation remains approximate for sparse cells\.

[Table˜E6](https://arxiv.org/html/2606.26130#A5.T6)reports the Pearson correlation between reference\-inventory and LLM residual matrices \(on shared rows and columns\), the number of significant cells in each source, and the number of sign flips among cells significant in at least one source\.[Figures˜D6](https://arxiv.org/html/2606.26130#A4.F6)and[D7](https://arxiv.org/html/2606.26130#A4.F7)show the full residual heatmaps\.

### C\.6Paper\-by\-paper specificity tests

To assess whether LLM suggestions track individual paper content rather than defaulting to generic recommendations, we conduct two complementary tests on pairwise reference–LLM shared\-paper subsets\.

#### Same\-paper vs\. shuffled\-paper similarity\.

For each shared paper, we build a taxonomy profile \(a count vector over all taxonomy categories for that entity type\) from both the reference inventory and the LLM, then compute cosine similaritycos⁡\(𝐯ref,𝐯LLM\)=𝐯ref⋅𝐯LLM/\(‖𝐯ref‖​‖𝐯LLM‖\)\\cos\(\\mathbf\{v\}^\{\\mathrm\{ref\}\},\\mathbf\{v\}^\{\\mathrm\{LLM\}\}\)=\\mathbf\{v\}^\{\\mathrm\{ref\}\}\\cdot\\mathbf\{v\}^\{\\mathrm\{LLM\}\}\\,/\\,\(\\\|\\mathbf\{v\}^\{\\mathrm\{ref\}\}\\\|\\,\\\|\\mathbf\{v\}^\{\\mathrm\{LLM\}\}\\\|\), set to0when either vector has zero norm\. We compare the mean same\-paper similaritySsameS\_\{\\mathrm\{same\}\}against a shuffled baselineSshuffledS\_\{\\mathrm\{shuffled\}\}obtained by randomly permuting the mapping between reference\-inventory and LLM paper indices \(1,000 permutations, seed==42\)\. The one\-sidedpp\-value is the fraction of permutations whereSshuffled≥SsameS\_\{\\mathrm\{shuffled\}\}\\geq S\_\{\\mathrm\{same\}\}\. Across all nine entity type×\\timesLLM combinations,SsameS\_\{\\mathrm\{same\}\}significantly exceedsSshuffledS\_\{\\mathrm\{shuffled\}\}\(p<0\.001p<0\.001in all cases;[Table˜E7](https://arxiv.org/html/2606.26130#A5.T7)\), with deltas ranging from 0\.075 \(models\) to 0\.266 \(metrics\)\. The pairwise shared\-paper counts are n=904/891/910=904/891/910for datasets, n=942/928/948=942/928/948for models, and n=937/924/943=937/924/943for metrics in the reference\-versus\-GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2 comparisons\. This provides strong evidence that LLM suggestions are conditioned on individual research questions, not generated from a fixed generic template\.

#### Inter\-paper entity Jaccard decomposition\.

We decompose the increase in inter\-paper entity overlap into a vocabulary\-compression component and an LLM\-specific excess\. For each source, we sample 10,000 unordered paper pairs and compute the mean pairwise Jaccard index\(Jaccard,[1912](https://arxiv.org/html/2606.26130#bib.bib69)\)J​\(a,b\)=\|Ea∩Eb\|/\|Ea∪Eb\|J\(a,b\)=\|E\_\{a\}\\cap E\_\{b\}\|\\,/\\,\|E\_\{a\}\\cup E\_\{b\}\|on entity name setsEa,EbE\_\{a\},E\_\{b\}\(within\-source, lowercased\)\. A random\-draw baseline controls for vocabulary compression: for each paper, we drawkik\_\{i\}entities \(matching the paper’s actual cardinality\) without replacement from the source’s corpus\-wide frequency distribution \(100 draws per pair\)\. We defineJexcess​\(s\)=Jactual​\(s\)−Jrandom​\(s\)J\_\{\\mathrm\{excess\}\}\(s\)=J\_\{\\mathrm\{actual\}\}\(s\)\-J\_\{\\mathrm\{random\}\}\(s\)andΔexcess=Jexcess​\(LLM\)−Jexcess​\(ref\)\\Delta\_\{\\mathrm\{excess\}\}=J\_\{\\mathrm\{excess\}\}\(\\mathrm\{LLM\}\)\-J\_\{\\mathrm\{excess\}\}\(\\mathrm\{ref\}\), with paired bootstrap 95% confidence intervals\(Tibshirani and Efron,[1993](https://arxiv.org/html/2606.26130#bib.bib71)\)\(2,000 resamples of paper pairs, percentile method\)\. For datasets,Δexcess\\Delta\_\{\\mathrm\{excess\}\}is near zero overall: it is indistinguishable from zero for Gemini 3 Pro and DeepSeek\-V3\.2 and small but positive for GPT\-5\.1, indicating that most of the increased inter\-paper overlap is explained by vocabulary compression rather than a large additional homogenisation term\. For models,Δexcess\>0\\Delta\_\{\\mathrm\{excess\}\}\>0\(0\.004–0\.012\), reflecting genuine excess homogenisation consistent with provider concentration\. For metrics, the pattern is mixed: GPT\-5\.1 and Gemini 3 Pro show negativeΔexcess\\Delta\_\{\\mathrm\{excess\}\}\(more content\-specific than the reference inventory relative to their vocabulary\), while DeepSeek\-V3\.2 shows modest positive excess\.

### C\.7Long\-tail sensitivity: robustness after removing low\-frequency entities

Many reference\-inventory entities appear in only one paper, some because they are genuinely paper\-specific and others because they are established but niche\. To test whether the main findings survive once low\-frequency entities are removed from the reference set, we apply simple rule\-based filters to the reference entity set while leaving LLM suggestions unfiltered\. These are long\-tail sensitivity analyses, not principled fairness corrections: the singleton filter has only 7\.5% precision as a proxy for paper\-specificity \([Table˜E15](https://arxiv.org/html/2606.26130#A5.T15)\), and the introducedness audit on which it is calibrated has weak inter\-model agreement \(κ=0\.108\\kappa=0\.108\)\.

We define three rule\-of\-thumb proxies for paper\-introduced entities: \(1\) singleton filter: entities appearing in exactly one paper in the reference corpus; \(2\) title\-match filter: entities whose name appears in the focal paper’s title; \(3\) combined filter: entities satisfying both criteria\. These are rarity proxies, not perfect tests of whether an entity is truly new to a paper, so they may remove genuinely established but niche entities\. We evaluate all three filters, but the manuscript table focuses on singleton exclusion because the title\-match and combined filters remove little and change the results negligibly\. Because these robustness analyses compare the reference inventory against all three LLMs under common inclusion rules, they use the all\-three\-LLM shared\-paper subsets: n=878=878for datasets, n=915=915for models, and n=911=911for metrics\. We calibrate these rules with a blinded cross\-model audit over a stratified sample of 300 entity–paper pairs using Claude Opus 4\.6 and GPT\-5\.4 \([Sections˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11),[E14](https://arxiv.org/html/2606.26130#A5.T14)and[E15](https://arxiv.org/html/2606.26130#A5.T15)\)\.

#### Entity\-name recall\.

After excluding singletons \(the most aggressive filter\), recall improves substantially: dataset recall rises from 15\.9% to 42\.7% \(GPT\-5\.1\), 16\.2% to 43\.2% \(Gemini 3 Pro\), and 12\.7% to 35\.9% \(DeepSeek\-V3\.2\)\. Model recall improves from 6\.8% to 20\.7% \(GPT\-5\.1\), 7\.9% to 24\.3% \(Gemini 3 Pro\), and 5\.1% to 16\.9% \(DeepSeek\-V3\.2\)\. Metric recall improves from 14\.1% to 42\.6% \(GPT\-5\.1\), 13\.1% to 42\.2% \(Gemini 3 Pro\), and 12\.2% to 40\.2% \(DeepSeek\-V3\.2\)\. This indicates that much of the apparent coverage gap is driven by singleton or otherwise rare entities\. However, singleton exclusion does not improve paper\-level model coverage: model zero\-coverage rises slightly from 65\.1/65\.2/83\.5% to 66\.6/67\.1/85\.2% for GPT\-5\.1, Gemini 3 Pro, and DeepSeek\-V3\.2, respectively\. The singleton filter therefore mainly shrinks the target universe rather than making more papers recoverable\. The title\-match and combined filters have minimal impact \(removing 0\.3–3\.3% of entities depending on entity type\), consistent with most paper\-specific entities not appearing in titles\.

#### Provider divergence persists\.

Provider JSD changes only modestly after removing singletons: it decreases for GPT\-5\.1 \(0\.111→0\.0920\.111\\to 0\.092\) and Gemini 3 Pro \(0\.102→0\.0900\.102\\to 0\.090\) but edges up for DeepSeek\-V3\.2 \(0\.137→0\.1400\.137\\to 0\.140\), confirming that provider concentration persists even under this aggressive filter for paper\-specific items\. The residual divergence remains substantial and follows the same pattern of Major commercial overrepresentation documented in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)\.[Table˜E8](https://arxiv.org/html/2606.26130#A5.T8)reports full results for all entity types and filter levels;[Figure˜D8](https://arxiv.org/html/2606.26130#A4.F8)a provides a visual comparison\.

### C\.8Broader provider grouping

The original provider taxonomy uses fine\-grained categories for major commercial providers but aggregates all academic and independent models into a single “Other/Academic” category\. This asymmetry may amplify the apparent concentration of LLM suggestions around commercial providers\. To address this, we introduce a broader four\-way grouping: \(1\) Major commercial \(OpenAI, Meta AI, Google DeepMind, Anthropic, Alibaba/Qwen, DeepSeek, Mistral AI\); \(2\) Other commercial \(Cohere, Hugging Face, Stability AI, Microsoft Research, NVIDIA/NeMo, Databricks/MosaicML\); \(3\) Academic/Community reused \(Other/Academic entities appearing in 2\+ papers\); \(4\) Other/Academic singleton \(Other/Academic entities appearing in exactly one paper after the same deterministic normalisation and fuzzy clustering used in the main entity\-level analysis\)\. This frequency\-based split therefore preserves the paper’s canonical model\-identity rule and requires no additional audit labels\.

Under this broader grouping, the reference inventory is 54\.9% Major commercial, 27\.7% Other/Academic singleton, 14\.5% Academic/Community reused, and 2\.9% Other commercial\. LLM suggestions overrepresent Major commercial \(\+\+17 to\+\+19 pp excess\) at the expense of Other/Academic singletons \(−\-23\.5 to−\-24\.3 pp\)\. Academic/Community reused is moderately overrepresented \(\+\+3\.6 to\+\+5\.8 pp\), suggesting that LLMs slightly favour established academic models over Other/Academic singletons\. The JSD under this broader grouping \(0\.082–0\.092\) is lower than under the original provider taxonomy \(0\.102–0\.137\), indicating that some of the original divergence reflects the asymmetric granularity of commercial versus academic provider categories; however, the remaining divergence remains substantial\. The cleaner interpretation is therefore commercial concentration combined with long\-tail suppression, not a wholesale disappearance of reused academic/community models\. Full distributions and excess shares are reported in[Table˜E9](https://arxiv.org/html/2606.26130#A5.T9);[Figure˜D8](https://arxiv.org/html/2606.26130#A4.F8)b visualises the broader grouping\.

### C\.9Credit for naming the right model family or provider

Exact\-name matching is the harshest possible evaluation: if the reference inventory uses “Llama\-3\-8B” and the LLM suggests “Llama\-3\-70B”, exact matching treats this as a complete miss\. We relax matching to three granularity levels: \(1\) exact match: entity names must match after normalisation; \(2\) family match: model names are mapped to families via regex patterns \(e\.g\., all Llama variants→\\to“llama”, all GPT\-4 variants→\\to“gpt4”\); \(3\) provider match: only the provider label must match\.

Mean per\-paper model recall improves substantially with relaxed matching\. At the exact level, recall is 6\.2% \(GPT\-5\.1\), 6\.5% \(Gemini 3 Pro\), and 3\.7% \(DeepSeek\-V3\.2\)\. At the family level, recall rises to 28\.2%, 27\.9%, and 20\.0%\. At the provider level \(deduplicated per paper\), recall reaches 53\.1%, 53\.3%, and 44\.9%\. This demonstrates that LLMs often suggest the correct model family or provider even when they miss the exact version, size, or fine\-tuned variant\. The gap between family\-level and provider\-level recall \(∼25\{\\sim\}25pp\) quantifies the extent to which LLMs suggest models from the correct provider but from the wrong family within that provider’s offerings\. Note that these per\-paper recall figures are not directly comparable to the corpus\-level recall in[Section˜C\.7](https://arxiv.org/html/2606.26130#A3.SS7), which counts unique entities across the entire corpus rather than averaging per\-paper overlap\.[Table˜E10](https://arxiv.org/html/2606.26130#A5.T10)reports mean and median recall at each level;[Figure˜D8](https://arxiv.org/html/2606.26130#A4.F8)c summarises the pattern\.

### C\.10Normalisation\-threshold sensitivity

The fuzzy\-clustering threshold \(token\-sort ratio≥T\\geq T\) governs how aggressively near\-duplicate entity names are merged \([Section˜A\.6](https://arxiv.org/html/2606.26130#A1.SS6)\)\. At the defaultT=90T=90, the blinded audit finds 76% merge precision \([Table˜E13](https://arxiv.org/html/2606.26130#A5.T13)\)\. To verify that headline metrics are not an artefact of this particular threshold,[Table˜5](https://arxiv.org/html/2606.26130#A3.T5)reports vocabulary size, effective number, Gini coefficient, and zero\-LLM\-coverage for model entities acrossT∈\{80,85,90,95,100\}T\\in\\\{80,85,90,95,100\\\}\. AsTTincreases \(stricter merging\), vocabulary sizes grow monotonically because fewer names are merged, but the relative compression ratio \(reference vocabulary to LLM vocabulary\) and the ordering of all information\-theoretic summaries remain stable across the full range\. Dataset and metric results follow the same pattern \(full data in the supplementary CSV\)\.

Table 5:The vocabulary\-compression hierarchy is stable across fuzzy\-clustering thresholds, ruling out normalisation sensitivity as an artefact\.Vocabulary size, effective numberexp⁡\(H\)\\exp\(H\), Gini coefficient, and zero\-LLM\-coverage for model entities across fuzzy\-clustering thresholdsT∈\{80,85,90,95,100\}T\\in\\\{80,85,90,95,100\\\}\. LLM columns show the range across GPT\-5\.1, Gemini 3 Pro, and DeepSeek\-V3\.2\. Bold row: default threshold used throughout the manuscript\. Dataset and metric results follow the same pattern \(full data in supplementary CSV\)\.
### C\.11Blinded two\-model audit of the pipeline

The pipeline uses GPT\-5\.1 for two automated steps, entity extraction and taxonomy classification, together with deterministic rules and fuzzy string clustering \(threshold 90\) for entity normalisation\. We audit these steps with a blinded two\-model protocol using Claude Opus 4\.6 \(high\) in Claude Code and GPT\-5\.4 \(xhigh\) in OpenAI Codex as independent model raters, under the API data\-handling procedure described in[Section˜A\.5](https://arxiv.org/html/2606.26130#A1.SS5)\. Both systems annotate stratified samples across four tasks: \(1\) extraction validation \(90 rows: 30 papers×\\times3 entity types\), verifying whether pipeline\-extracted entities genuinely appear in each paper’s experiments; \(2\) classification validation \(180 sampled entities, stratified by source: reference inventory and LLM outputs; each entity is audited on all taxonomy dimensions applicable to its entity type, so the audit unit is the entity×\\timesapplicable\-dimension judgment, and the per\-dimensionncons\.n\_\{\\mathrm\{cons\.\}\}andnpair\.n\_\{\\mathrm\{pair\.\}\}counts inLABEL:tab:annotation\_classificationsum across these entity–dimension pairs rather than across the 180 sampled entities\), verifying taxonomy label assignments; \(3\) normalisation validation \(100 fuzzy\-merge decisions, balanced between merged pairs and near\-miss non\-merged pairs with similarity scores 80–89\), verifying merge correctness; and \(4\) introducedness validation \(300 entity–paper pairs: 120 model, 100 dataset, 80 metric, quota\-sampled across frequency bands\), classifying whether entities are pre\-existing, newly introduced by the paper, or paper\-specific derivatives of existing items\. Inter\-model reliability is measured with Cohen’sκ\\kappa\(Cohen,[1960](https://arxiv.org/html/2606.26130#bib.bib77)\)\(classification, normalisation, introducedness\) and intraclass correlation\(Shrout and Fleiss,[1979](https://arxiv.org/html/2606.26130#bib.bib78)\)\(extraction counts\)\. We do not adjudicate disagreements; instead, final pipeline\-result metrics are computed only on rows where both systems agree after multi\-label answers are put into a consistent order, while agreement metrics use all rows where both systems provided usable labels\.

#### Pipeline validation\.

Extraction validation reports mean per\-paper precision \(correct / extracted\), recall \(correct / \(correct\+\+missed\)\), F1, and hallucination rate on rows where both systems agree \([Table˜E11](https://arxiv.org/html/2606.26130#A5.T11)\)\. The audited taxonomy labels are drawn from the with\-pipeline \[WP\] classifier outputs used throughout[Appendix˜C](https://arxiv.org/html/2606.26130#A3); for dimensions whose labels are sensitive to pipeline context \(notably architecture and task type,[Section˜B\.2](https://arxiv.org/html/2606.26130#A2.SS2)\), the reliability tiers below pertain to the \[WP\] branch, whereas provider labels barely move across branches \([Section˜B\.1](https://arxiv.org/html/2606.26130#A2.SS1)\), so provider tier assignments apply to both \[EL\] and \[WP\]\. Classification validation reports exact\-match accuracy against the original pipeline labels on those agreed rows, withκ\\kappavalues by source stratum and dimension computed on all rows where both systems gave usable labels \(LABEL:tab:annotation\_classification\)\. Normalisation validation reports merge precision on merged pairs and error rate on the near\-miss band on agreed rows, with overallκ\\kappaon the full sample \([Table˜E13](https://arxiv.org/html/2606.26130#A5.T13)\)\. Agreement is highest for normalisation \(κ=0\.898\\kappa=0\.898\), moderate across classification dimensions \(medianκ=0\.660\\kappa=0\.660\), lower for extraction counts \(overall ICC\(2,1\)=0\.469\(2,1\)=0\.469\), and weakest for introducedness \(overallκ=0\.108\\kappa=0\.108\)\.[Table˜6](https://arxiv.org/html/2606.26130#A3.T6)provides a compact reliability overview; conditional\-on\-consensus diagnostics \(precision, recall, accuracy\) are reported in the task\-specific tables\. These diagnostics describe only the subset of rows on which both auditors converged and serve as best\-case cleanliness checks for agreed cases, not as corpus\-level accuracy estimates\. Consensus rates vary substantially: normalisation reaches 95%, classification 77%, and introducedness 79%, but extraction consensus is only 59% \(53 of 90 rows\), so the perfect precision/recall/F1 reported there applies to a consensus subset, not the full audit\.

#### Four\-tier validation of taxonomy dimensions\.

The classification audit supports a four\-tier interpretation of the 15 taxonomy dimensions based on consensus\-row accuracy, inter\-modelκ\\kappa, and sample size\. Strong dimensions achieve the highest agreement: provider \(κ=0\.923\\kappa=0\.923–1\.0001\.000, accuracy 89\.7–94\.1%\) and openness \(κ=1\.000\\kappa=1\.000, accuracy 94\.1–96\.8%\)\. Moderate dimensions meet theκ≥0\.5\\kappa\\geq 0\.5and accuracy≥\\geq75% thresholds in at least one source stratum but with narrower margins: modality \(Reference:κ=0\.756\\kappa=0\.756, accuracy 85\.7%,npair\.=17n\_\{\\mathrm\{pair\.\}\}=17; LLM:κ=0\.654\\kappa=0\.654, accuracy 85\.7%,npair\.=44n\_\{\\mathrm\{pair\.\}\}=44\), evaluation type \(LLM:κ=0\.588\\kappa=0\.588, accuracy 89\.7%; Reference:κ=0\.061\\kappa=0\.061on onlyn=17n=17pairwise rows, a small\-sample artifact\), and linguistic scope \(LLM:κ=0\.660\\kappa=0\.660, accuracy 85\.0%,npair\.=23n\_\{\\mathrm\{pair\.\}\}=23\)\. Size is tentative: it meets accuracy andκ\\kappathresholds in the reference strata \(dataset Reference: 90\.0%,κ=0\.642\\kappa=0\.642; model Reference: 87\.5%,κ=1\.000\\kappa=1\.000\) but the dataset LLM stratum drops toκ=0\.406\\kappa=0\.406, individual strata narrowly miss the sample\-size floor \(npair\.=8n\_\{\\mathrm\{pair\.\}\}=8–1313in the stronger strata\), and model\-side LLM size accuracy drops to 47\.1%; size\-dependent claims therefore carry more uncertainty than provider or evaluation\-type claims\. The remaining dimensions are exploratory for distinct reasons: task type, domain, annotation, architecture, and training paradigm have consensus\-row accuracy below 50% in at least one stratum; granularity and data quality achieve high accuracy but fall below theκ\\kappathreshold \(κ=0\.25\\kappa=0\.25–0\.370\.37andκ=0\.00\\kappa=0\.00–0\.500\.50, respectively\); and cognitive/affective has insufficient sample size \(npair\.≤6n\_\{\\mathrm\{pair\.\}\}\\leq 6\)\. Analyses resting on exploratory dimensions \(including task type×\\timesprovider co\-occurrence\) provide descriptive context rather than definitive structural claims\. This four\-tier distinction is flagged explicitly in figure annotations and throughout the results narrative\.

#### How well the singleton and title rules capture paper\-specific entities\.

The singleton and title\-match rules from[Section˜C\.7](https://arxiv.org/html/2606.26130#A3.SS7)are proxies for paper\-introduced entities\. To calibrate them, each of the 300 entity–paper pairs receives an audit\-confirmed label saying whether the entity is pre\-existing reusable, paper\-introduced, paper\-specific derivative, or unclear\.[Table˜E14](https://arxiv.org/html/2606.26130#A5.T14)reports the agreed label distribution;[Table˜E15](https://arxiv.org/html/2606.26130#A5.T15)reports the precision, recall, and specificity of those rules against the agreed labels\. No agreed row remained labelled “unclear,” so the conservative \(unclear→\\topre\-existing\) and liberal \(unclear→\\topaper\-introduced\) analyses coincide numerically\. The sample is quota\-based across frequency bands and entity types; raw percentages are sample\-level estimates, not corpus\-weighted distributions\.

#### Mismatch type characterisation\.

For each entity–paper pair, both systems also annotate the mismatch type that best characterises why the entity might be missed by LLM suggestions: alias/variant, same family, paper\-specific, or established but absent\. We treat this field as auxiliary context rather than a quantitative calibration target, because overall inter\-model agreement is low \(κ=0\.040\\kappa=0\.040\)\.

#### Caveats\.

This is a blinded model\-assisted audit, not human manual validation\. Claude Opus 4\.6 relied more heavily ondomain\-knowledgetags in dataset and model classification, whereas GPT\-5\.4 more often marked fieldsunresolved\-after\-revieworpaper\-reviewed; to avoid adjudicating across these evidence\-use profiles, final pipeline metrics are reported only on consensus rows\. The validation samples are designed to test pipeline decisions under stricter conditions, not to estimate corpus\-level prevalence\. Classification accuracy should be interpreted per source stratum\. Extraction metrics are computed per paper and averaged over consensus rows\. The introducedness sample is quota\-based, so extrapolation to corpus\-level distributions requires caution\.

#### Model\-swap robustness\.

We further test whether the headline provider\-concentration finding is an artefact of using GPT\-5\.1 for extraction or classification through three supplementary robustness checks that are orthogonal to the blinded audit above\. First, a deterministic regex\-based mapping \(60\+ hand\-crafted rules, requiring no LLM\) agrees with 92–98% of GPT\-5\.1 provider labels among regex\-classifiable model entities \(56–78% coverage per source\) across all four corpus sources \(n=1,000n=1\{,\}000papers each for the reference inventory and GPT\-5\.1;n=998n=998for DeepSeek\-V3\.2;n=981n=981for Gemini 3 Pro\)\. Second, we re\-extract entities from 94 papers using Claude Opus 4\.6 with reconstructed full text, under the same non\-public scientific TDM and API data\-handling procedure described in[Sections˜A\.2](https://arxiv.org/html/2606.26130#A1.SS2)and[A\.5](https://arxiv.org/html/2606.26130#A1.SS5); the resulting provider distribution is near\-identical to the GPT\-5\.1 extraction \(Spearmanρ=0\.991\\rho=0\.991,p<0\.0001p<0\.0001; mean Jaccard similarity on model names=0\.62=0\.62\)\. Third, we re\-classify the GPT\-5\.1\-extracted entities from 200 papers with Claude Opus 4\.6 using the identical taxonomy prompt\. Cross\-classifier agreement is almost perfect for provider \(κ=0\.961\\kappa=0\.961,n=1,498n=1\{,\}498model entities\) and openness \(κ=0\.912\\kappa=0\.912\), substantial for evaluation type \(κ=0\.711\\kappa=0\.711,n=996n=996metric entities\) and modality \(κ=0\.602\\kappa=0\.602,n=659n=659dataset entities\), fair for architecture \(κ=0\.339\\kappa=0\.339\) and the tentative size dimension \(κ=0\.280\\kappa=0\.280\), and slight for training paradigm \(κ=0\.200\\kappa=0\.200\)\. This gradient is broadly consistent with the reliability hierarchy established by the blinded audit: the strong dimensions remain strong under classifier substitution, while the exploratory and tentative dimensions show the same instability regardless of which model performs the classification\.

Table 6:Blinded cross\-model audit: inter\-model reliability\.Claude Opus 4\.6 \(high\) in Claude Code and GPT\-5\.4 \(xhigh\) in OpenAI Codex independently audited stratified validation samples for the four pipeline stages\. Inter\-model agreement is computed on all pairwise\-complete rows\. The consensus rate shows the fraction of rows on which both auditors converged after label canonicalisation; conditional\-on\-consensus diagnostics \(precision, recall, accuracy\) are reported in the task\-specific tables \([Tables˜E11](https://arxiv.org/html/2606.26130#A5.T11),LABEL:tab:annotation\_classification,[E13](https://arxiv.org/html/2606.26130#A5.T13)and[E14](https://arxiv.org/html/2606.26130#A5.T14)\)\. For classification,n=180n=180refers to sampled entities; each entity contributes one judgment per applicable taxonomy dimension, so the per\-dimension counts inLABEL:tab:annotation\_classificationsum across entity×\\timesdimension pairs\.

## Appendix DSupplementary figures

![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure2_taxonomy_divergence.png)Figure D1:Across the five most\-referenced taxonomy dimensions, the most pronounced LLM\-versus\-reference divergence is concentrated in model provider distributions \[EL\]\.Computed in the baseline setting where the classifier sees only entity lists, not the generated pipeline, using classified outputs from n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, and 998 DeepSeek\-V3\.2 papers\. Grouped horizontal bar charts for five taxonomy dimensions: dataset modality \(a\), dataset task type \(b\), model provider \(c\), model openness \(d\), and evaluation metric type \(e\)\. For visual clarity, only GPT\-5\.1 and Gemini 3 Pro are shown; DeepSeek\-V3\.2 exhibits similar patterns\. In this baseline, LLMs consistently overweight the top combined model\-size bucket \(GPT\-5\.1 59\.1%, Gemini 3 Pro 58\.3%, DeepSeek\-V3\.2 63\.2%, vs\. 43\.4% in the reference inventory\), and metric evaluation type shifts toward accuracy\-like metrics \(reference 41\.6%; LLMs 44\.8–55\.9%\) while user\-experience and efficiency metrics decline\. Within\-dataset category shifts are modest: self\-supervised annotation rises from 4\.7% to 7\.6–8\.1%, multilingual datasets from 4\.9% to 5\.6–8\.2%, and the education domain falls from 11\.2% to∼\\sim8\.3%\. The take\-away is that LLMs overweight major commercial providers relative to the paper corpus, with a Jensen–Shannon divergence about 3–5×\\timeslarger than the next\-largest taxonomy dimension \([Table˜E3](https://arxiv.org/html/2606.26130#A5.T3)\)\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_cooccurrence_jsd.png)Figure D2:Provider\-based co\-occurrence pairs show the largest structural divergence between LLMs and the reference inventory, while openness\-based pairs are the most faithfully preserved \[WP\]\.Computed from the with\-pipeline classification branch\. Source files contain n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, and 998 DeepSeek\-V3\.2 papers, but each co\-occurrence pair uses the subset of papers with non\-empty labels for both participating entity types, so support varies by pair and source\. Mean row\-wise Jensen–Shannon divergence between the reference\-inventory and LLM co\-occurrence matrices is shown for all 12 cross\-entity taxonomy combinations: 9 dataset–model pairs \(task type, size, and cognitive/affective×\\timesarchitecture, provider, and openness\) and 3 metric–model pairs \(evaluation type×\\timesarchitecture, provider, and openness\)\. For the two focal provider pairs, evaluation type×\\timesprovider has mean row\-wise JSD of 0\.124/0\.105/0\.153 \(GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2\) and task type×\\timesprovider of 0\.185/0\.100/0\.201; other provider\-based pairs show comparable divergence\. Architecture\- and task\-type\-based pairs should be read descriptively because those dimensions are label\-sensitive or exploratory\. Averaged across all 12 pairs, Gemini 3 Pro sits closest to the reference\-inventory co\-occurrence structure and DeepSeek\-V3\.2 farthest, but the ordering varies by pair\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_ablation_pipeline.png)Figure D3:Pipeline context mainly relabels generation and document\-level granularity categories, leaving most dataset dimensions nearly unchanged\.Computed on the classified suggestion outputs available for each model \(n=1,000=1\{,\}000GPT\-5\.1 papers, 981 Gemini 3 Pro papers, 998 DeepSeek\-V3\.2 papers\)\. Percentage\-point change in dataset category labels when the classifier sees the generated pipeline in addition to the entity lists, relative to the baseline where it sees only the entity lists, across three LLMs \(GPT\-5\.1, Gemini 3 Pro, DeepSeek\-V3\.2\); positive values indicate increased classification frequency with pipeline context\. The shift is small in aggregate, indicating that the entity\-list\-only baseline in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)is not artificially dominated by labelling sensitivity\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_web_search_ablation.png)Figure D4:Enabling web search at classification time mainly relabels architecture categories without shifting provider, size, or evaluation\-type distributions\.Computed on the 1,000 GPT\-5\.1 suggestion rows classified with and without web search\. Percentage\-point change in category labels when web search is enabled versus disabled, for datasets \(a\), models \(b\), and metrics \(c\); only subcategories with\|Δ\|≥1\|\\Delta\|\\geq 1pp are shown\. This confirms that architecture findings are label\-sensitive and exploratory, whereas the provider\-level claims in[Section˜2\.2](https://arxiv.org/html/2606.26130#S2.SS2)are stable under this ablation\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_appendix_cooccurrence_tasktype_provider.png)Figure D5:Across major task types, Meta AI and OpenAI dominate in LLM suggestions whereas Other/Academic models dominate in the paper\-derived reference inventory \[WP, exploratory\]\.Task\-type labels are exploratory \(10–32% audit accuracy in the primary strata\), so these panels provide descriptive context rather than definitive structural claims\. Layout and conventions as in[Figure˜5](https://arxiv.org/html/2606.26130#S2.F5)but for task type×\\timesprovider\. In the reference inventory \(a\), Other/Academic models dominate classification, reasoning, and question answering; in LLM suggestions \(b–d\), this is inverted, with DeepSeek\-V3\.2 showing the strongest Other/Academic suppression\. The fine\-grained Other/Academic category aggregates reused academic/community models with the singleton\-defined long tail \([Section˜C\.8](https://arxiv.org/html/2606.26130#A3.SS8)\), so the co\-occurrence shift partly reflects long\-tail suppression rather than displacement of established academic models\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_b3_TaskTypes_x_Provider.png)Figure D6:LLMs produce sharper task\-type×\\timesprovider residual patterns than the paper\-derived reference inventory, indicating a narrower combinatorial landscape \[WP\]\.Residual matrices are built from pair\-specific with\-pipeline source subsets in the classified files \(n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, 998 DeepSeek\-V3\.2 papers\); residual correlations and significance counts after marginal filtering are reported in[Table˜E6](https://arxiv.org/html/2606.26130#A5.T6)\. Cells show how much each pairing appears above or below what would be expected from the overall row and column frequencies, for the reference inventory \(a\), GPT\-5\.1 \(b\), Gemini 3 Pro \(c\), and DeepSeek\-V3\.2 \(d\); asterisks mark FDR\-significant cells \(q<0\.05q<0\.05\), positive \(red\) residuals indicate co\-occurrence above expectation, and negative \(blue\) below\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_b3_EvType_x_Provider.png)Figure D7:Evaluation\-type×\\timesprovider residuals are sparser and more concentrated in LLM suggestions than in the paper\-derived reference inventory \[WP\]\.Layout as in[Figure˜D6](https://arxiv.org/html/2606.26130#A4.F6), with the same source\-count and residual\-analysis conventions\. The pattern of more significant residual cells in LLMs than in the reference inventory reinforces that LLMs narrow the combinatorial landscape around provider\-centred pairings\.![Refer to caption](https://arxiv.org/html/2606.26130v1/figures/figure_robustness_dashboard.png)Figure D8:Provider concentration and long\-tail suppression survive singleton exclusion, broader provider regrouping, and relaxed family/provider matching\.Panelauses the all\-three\-LLM shared\-paper subsets \(datasets n=878=878, models n=915=915, metrics n=911=911\); it compares entity names only and is therefore branch\-independent\. Panelsbandcuse the 915 shared model papers with taxonomy labels from the with\-pipeline setting \[WP\]\. \(a\) Corpus\-level recall before and after excluding singleton entities from the reference inventory; hatched bars show recall after singleton exclusion\. \(b\) Provider distribution under the broader four\-way provider grouping \[WP\]; LLMs overrepresent major commercial providers and underrepresent Other/Academic singletons\. \(c\) Mean per\-paper model recall under exact\-name, family\-level, and provider\-level matching \(provider recall deduplicated per paper\) \[WP\]\. The sharp jump from exact to family/provider matching shows that many apparent misses are granularity errors rather than complete neighbourhood failures\.
## Appendix ESupplementary tables

The tables below provide the detailed numerical counterparts to the appendix claims summarized above: first the label\-sensitivity checks, then the effect\-size tables, paper\-by\-paper similarity estimates, robustness filters, looser matching rules, normalisation\-threshold sensitivity, and audit outputs\.

Table E1:Largest shifts in model and metric labels when the classifier sees the generated pipeline\.Computed on the classified suggestion outputs available for each model \(n=1,000=1\{,\}000GPT\-5\.1 papers, 981 Gemini 3 Pro papers, 998 DeepSeek\-V3\.2 papers\)\. Each value represents the change in the share of a subcategory when the classifier sees the generated pipeline as well as the entity lists\. Only subcategories where at least one LLM exhibits\|Δ​pp\|≥2\.0\|\\Delta\\text\{pp\}\|\\geq 2\.0are shown\. Bold values indicate\|Δ​pp\|≥2\.0\|\\Delta\\text\{pp\}\|\\geq 2\.0\.Table E2:Ablation conditions across models\.All three LLMs are evaluated with and without pipeline context\. Web search is tested only for GPT\-5\.1\.Table E3:Provider exhibits the largest Jensen–Shannon divergence and the largest Cramér’sVVamong the 15 taxonomy dimensions\. The provider JSD is about 3–5×\\timesthe next\-largest JSD; the Cramér’sVVmargin is smaller\.Computed in the setting where the classifier also sees the generated pipeline\. Pairwise reference–LLM shared\-paper counts are n=904/891/910=904/891/910for datasets, n=942/928/948=942/928/948for models, and n=937/924/943=937/924/943for metrics \(GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2\)\. Jensen–Shannon divergence \(JSD, base\-2\) with bootstrapped 95% confidence intervals, Cramér’sVVeffect sizes, andχ2\\chi^\{2\}test statistics are reported for all category dimensions\.nndenotes the total number of label instances in the contingency table\.χ2\\chi^\{2\}p\-values are descriptive given the multi\-label, paper\-nested structure of the contingency tables; inferential weight rests on JSD with paper\-level bootstrap CIs and the robustness analyses in this appendix\.Table E4:LLM suggestions fall between a popularity\-proportional sampler and a deterministic top\-kkranker, confirming that they respond to the research question while still exhibiting distributional biases\.Computed on the with\-pipeline classification branch using pairwise reference–LLM shared\-paper subsets\. Model\-provider and model\-size comparisons use n=942/928/948=942/928/948shared papers for GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2; metric evaluation type uses n=937/924/943=937/924/943\. Jensen–Shannon divergence between the paper\-derived reference inventory and each LLM \(JSDLLM\), a deterministic top\-kkbaseline \(JSDtop​\-​k\{\}\_\{\\mathrm\{top\\text\{\-\}k\}\}\), and a stochastic popularity\-sampled baseline \(JSDsampled\) is reported\. Lower values indicate closer alignment with the reference inventory\.Table E5:Local BM25 retrieval sharpens, rather than replaces, the popularity calibration: LLMs remain closer to the reference inventory than a content\-conditioned retriever\.Computed on the with\-pipeline classification branch using pairwise reference–LLM shared\-paper subsets and leave\-one\-out BM25 retrieval over generated research questions with top\-NNneighborhoods \(N=25\)\. Model\-provider and model\-size comparisons use n=942/928/948=942/928/948shared papers for GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2; metric evaluation type uses n=937/924/943=937/924/943\. Jensen–Shannon divergence between the paper\-derived reference inventory and each LLM \(JSDLLM\) and the deterministic local BM25 baseline \(JSDlocal​\-​BM25\{\}\_\{\\mathrm\{local\\text\{\-\}BM25\}\}\) is reported\. Lower values indicate closer alignment with the reference inventory\.Table E6:LLMs preserve the sign of co\-occurrence associations but sharpen them into a narrower pattern\.Residual correlations are computed on shared rows and columns after marginal filtering\. The underlying co\-occurrence matrices come from pair\-specific with\-pipeline source subsets in the classified files \(n=1,000=1\{,\}000reference\-inventory papers, 1,000 GPT\-5\.1 papers, 981 Gemini 3 Pro papers, 998 DeepSeek\-V3\.2 papers\)\. Pearson correlation of residual matrices between the paper\-derived reference inventory and each LLM, number of FDR\-significant cells, and sign flips among significant cells\.Table E7:LLM suggestions track individual paper content, while excess homogenisation beyond vocabulary compression concentrates in model suggestions\.Pairwise shared\-paper counts are n=904/891/910=904/891/910for datasets, n=942/928/948=942/928/948for models, and n=937/924/943=937/924/943for metrics in the reference\-versus\-GPT\-5\.1/Gemini 3 Pro/DeepSeek\-V3\.2 comparisons\.Top:Same\-paper vs\. shuffled taxonomy cosine similarity;Δ=Ssame−Sshuffled\\Delta=S\_\{\\mathrm\{same\}\}\-S\_\{\\mathrm\{shuffled\}\}\.Bottom:Inter\-paper entity Jaccard decomposition;Δexcess\\Delta\_\{\\mathrm\{excess\}\}isolates homogenisation beyond vocabulary compression, with bootstrapped 95% CIs\. NegativeΔexcess\\Delta\_\{\\mathrm\{excess\}\}indicates LLMs are more content\-specific than the paper\-derived reference inventory relative to their vocabulary size\.Test 1: Same\-paper vs\. shuffled taxonomy similarityEntity typeLLMSsameS\_\{\\mathrm\{same\}\}SshuffledS\_\{\\mathrm\{shuffled\}\}Δ\\DeltappDatasetsGPT\-5\.10\.7210\.4960\.225<<0\.001Gemini 3 Pro0\.7270\.5080\.219<<0\.001DeepSeek\-V3\.20\.7160\.4930\.223<<0\.001ModelsGPT\-5\.10\.7900\.7100\.081<<0\.001Gemini 3 Pro0\.7900\.7040\.086<<0\.001DeepSeek\-V3\.20\.7640\.6890\.075<<0\.001MetricsGPT\-5\.10\.7110\.4510\.261<<0\.001Gemini 3 Pro0\.6930\.4270\.266<<0\.001DeepSeek\-V3\.20\.6990\.4640\.234<<0\.001Test 2: Inter\-paper entity Jaccard decompositionEntity typeSourceJactualJ\_\{\\mathrm\{actual\}\}JrandomJ\_\{\\mathrm\{random\}\}JexcessJ\_\{\\mathrm\{excess\}\}\[95% CI\]DatasetsReference0\.00140\.0022\-0\.0008GPT\-5\.10\.00610\.00590\.0002Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(GPT\-5\.1\)––0\.0010 \[0\.0001, 0\.0019\]Reference0\.00160\.0022\-0\.0006Gemini 3 Pro0\.00540\.0061\-0\.0007Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(Gemini 3 Pro\)––\-0\.0000 \[\-0\.0008, 0\.0008\]Reference0\.00140\.0022\-0\.0008DeepSeek\-V3\.20\.00490\.0056\-0\.0007Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(DeepSeek\-V3\.2\)––0\.0001 \[\-0\.0007, 0\.0010\]ModelsReference0\.00790\.00590\.0019GPT\-5\.10\.04540\.03990\.0055Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(GPT\-5\.1\)––0\.0036 \[0\.0018, 0\.0055\]Reference0\.00800\.00600\.0020Gemini 3 Pro0\.06860\.05420\.0144Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(Gemini 3 Pro\)––0\.0124 \[0\.0097, 0\.0150\]Reference0\.00800\.00600\.0020DeepSeek\-V3\.20\.07830\.06410\.0142Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(DeepSeek\-V3\.2\)––0\.0122 \[0\.0092, 0\.0152\]MetricsReference0\.02930\.01630\.0130GPT\-5\.10\.03670\.03240\.0043Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(GPT\-5\.1\)––\-0\.0087 \[\-0\.0107, \-0\.0067\]Reference0\.02740\.01600\.0114Gemini 3 Pro0\.01960\.01880\.0008Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(Gemini 3 Pro\)––\-0\.0106 \[\-0\.0123, \-0\.0089\]Reference0\.02790\.01590\.0120DeepSeek\-V3\.20\.07250\.05200\.0205Δexcess\\Delta\_\{\\mathrm\{excess\}\}\(DeepSeek\-V3\.2\)––0\.0085 \[0\.0058, 0\.0113\]

Table E8:Provider concentration persists after removing singleton entities, confirming it is not an artefact of rare paper\-specific references\.These analyses use the all\-three\-LLM shared\-paper subsets: datasets n=878=878, models n=915=915, metrics n=911=911\. “All” retains the full reference\-inventory entity set; “Excl\. singletons” removes entities appearing in exactly one paper\. Provider JSD is reported only for models \(the only entity type with a provider taxonomy dimension\)\. Zero\-coverage rate denominator = papers with≥\\geq1 reference\-inventory entity after filtering\. Title\-match and combined filters are omitted as they remove 0\.3–3\.3% of entities and produce negligible changes\.nrefn\_\{\\text\{ref\}\}denotes the number of unique normalised reference\-inventory entity names in the shared\-paper subset retained under each filter\.Table E9:Under a broader provider regrouping, the deficit concentrates in singleton\-defined long\-tail models while reused academic/community models are modestly overrepresented\.Computed on the all\-three\-LLM shared model\-paper intersection \(n=915=915\)\. Share \(%\) shows the proportion of model mentions falling in each broad provider category; excess is the percentage\-point difference \(LLM−\-Reference\)\. JSD \(original\) uses the fine\-grained provider taxonomy; JSD \(broader grouping\) uses the four\-category taxonomy\.Table E10:Exact\-name model recall is low, but family\- and provider\-level matching recovers most of the apparent miss\.Mean per\-paper recall at three matching granularities\. Exact: normalised entity names must match\. Family: model names mapped to families via regex \(e\.g\., all Llama variants→\\to“llama”\)\. Provider \(dedup\): only the provider label must match, deduplicated per paper\. These per\-paper figures are not comparable to the corpus\-level recall in[Table˜E8](https://arxiv.org/html/2606.26130#A5.T8), which counts unique entities across the entire corpus\.n=915n=915shared papers \(papers where the paper\-derived reference inventory and all three LLMs contain model entities after normalisation\)\.Table E11:Conditional\-on\-consensus extraction diagnostics; not corpus\-level or full\-audit accuracy\.Precision = correctly extracted / total extracted; recall = correctly extracted / \(correctly extracted\+\+missed\); hallucination rate = hallucinated / total extracted\. Precision, recall, F1, and hallucination rate are mean per\-paper values over consensus rows only \(53 of 90 audited rows; overall extraction consensus 58\.9%\)\. The 100% precision/recall/F1 therefore describe the majority\-selected subsample, not the full audit\. ICC\(2,1\) is computed on the full pairwise\-complete count series between Claude Opus 4\.6 and GPT\-5\.4\.Table E12:Blinded cross\-model audit: classification validation\.Exact\-match accuracy requires set equality for multi\-label dimensions after alphabetical canonicalisation\. Accuracy is computed on consensus rows only;κ\\kappais computed on all pairwise\-complete rows\. Source separates paper\-derived reference\-inventory entities \(Reference\) from LLM\-suggested entities pooled across GPT\-5\.1, Gemini 3 Pro, and DeepSeek\-V3\.2 \(LLM\)\. The four\-tier validation assignments \([Section˜C\.11](https://arxiv.org/html/2606.26130#A3.SS11)\) are based on the Reference and pooled\-LLM \(LLM\) strata, which provide the largest sample sizes and the cleanest source separation; the per\-model strata \(DeepSeek\-V3\.2, GPT\-5\.1, Gemini 3 Pro\) are supplementary descriptive checks with smaller samples and should not be used to override tier assignments\.Entity typeSourceDimensionncons\.n\_\{\\mathrm\{cons\.\}\}Accuracy \(%\)κ\\kappanpair\.n\_\{\\mathrm\{pair\.\}\}DatasetReferenceModality1485\.70\.75617Task type1010\.00\.51818Domain714\.30\.32918Annotation80\.00\.71010Size1090\.00\.64213Granularity977\.80\.36716Linguistic scope988\.91\.0009Cognitive/affective450\.00\.6006Data quality16100\.00\.00018LLMModality3585\.70\.65444Task type2532\.00\.52044Domain2951\.70\.59245Annotation1816\.70\.74721Size1687\.50\.40624Granularity2475\.00\.25437Linguistic scope2085\.00\.66023Cognitive/affective1100\.00\.2503Data quality3997\.40\.49544DeepSeek\-V3\.2Modality1190\.90\.57716Task type850\.00\.43416Domain1353\.80\.76816Annotation714\.31\.0007Size475\.00\.6675Granularity771\.4\-0\.10011Linguistic scope666\.71\.0006Cognitive/affective1100\.01\.0001Data quality1593\.30\.63616GPT\-5\.1Modality1586\.70\.70917Task type1330\.80\.69118Domain955\.60\.41718Annotation714\.30\.5619Size785\.70\.42111Granularity1172\.70\.36516Linguistic scope785\.70\.5919Data quality15100\.00\.00018Gemini 3 ProModality977\.80\.66711Task type40\.00\.31810Domain742\.90\.56911Annotation425\.00\.6675Size5100\.00\.0008Granularity683\.30\.31010Linguistic scope7100\.00\.7048Data quality9100\.00\.73710ModelReferenceArchitecture1625\.00\.81317Training paradigm1010\.00\.33917Provider1794\.11\.00017Openness1794\.11\.00017Size887\.51\.0008LLMArchitecture3240\.61\.00032Training paradigm2010\.00\.32932Provider2989\.70\.92331Openness3196\.81\.00031Size1747\.10\.88518DeepSeek\-V3\.2Architecture1020\.01\.00010Training paradigm714\.30\.51610Provider966\.71\.0009Openness9100\.01\.0009Size560\.01\.0005GPT\-5\.1Architecture1258\.31\.00012Training paradigm70\.00\.26812Provider11100\.00\.89312Openness1291\.71\.00012Size862\.51\.0008Gemini 3 ProArchitecture1040\.01\.00010Training paradigm616\.70\.13010Provider9100\.00\.86510Openness10100\.01\.00010Size40\.00\.0005MetricReferenceEvaluation type785\.70\.06117LLMEvaluation type2989\.70\.58842DeepSeek\-V3\.2Evaluation type10100\.00\.76212GPT\-5\.1Evaluation type771\.40\.19818Gemini 3 ProEvaluation type1291\.71\.00012Table E12:Blinded cross\-model audit: classification validation\.\(continued\)Table E13:Blinded cross\-model audit: normalisation validation\.Merge precision is the fraction of merged pairs that both systems judge should merge; near\-miss error rate is the fraction of 80–89 similarity pairs that both systems judge should merge\. The overallκ\\kappais computed on all 100 pairwise\-complete rows; value percentages use consensus rows only\.Table E14:Blinded cross\-model audit: introducedness label distribution\.Each entity–paper pair is classified as pre\-existing reusable, paper\-introduced, paper\-specific derivative, or unclear by Claude Opus 4\.6 and GPT\-5\.4\. Percentages are computed on consensus rows only;κ\\kappauses all pairwise\-complete rows within each entity type\. The sample is quota\-based across entity types and frequency bands, so percentages are sample\-level estimates rather than corpus\-weighted distributions\.Table E15:Heuristic calibration against audit\-confirmed introducedness labels\.Precision is the fraction of heuristic\-flagged entities that are audit\-confirmed paper\-specific \(paper\-introduced or paper\-specific derivative\); recall is the fraction of audit\-confirmed paper\-specific entities that the heuristic flags; specificity is the fraction of audit\-confirmed pre\-existing entities correctly not flagged\. Effective evaluation counts use consensus introducedness rows only\. Because no consensus introducedness row remained labelled unclear, the conservative and liberal analyses coincide numerically\.
## Appendix FPaper\-side entity extraction prompt

The following system prompt was sent to GPT\-5\.1 via the OpenAI Batch API for each of the 1,000 papers\. The model received the paper’s title, abstract, and full text as user input under the non\-public scientific TDM and API data\-handling procedure described in[Sections˜A\.2](https://arxiv.org/html/2606.26130#A1.SS2)and[A\.5](https://arxiv.org/html/2606.26130#A1.SS5); the batch input/output file objects under our control were deleted after processing\.

System promptYou are an academic assistant\. Given the title, abstract, and full text of a paper:1\. Generate a single concise research question\.2\. Extract:\- datasets used\- models used\- evaluation metrics usedRespond in JSON with:\{"research\_question": "\.\.\.","GroundTruthDatasets": \["\.\.\."\],"GroundTruthModels": \["\.\.\."\],"GroundTruthMetrics": \["\.\.\."\]\}Only report the datasets, models, and metrics used in the experiments and not from the literature review or related work sections\.Each dataset, model and evaluation metric name must be composed from one to three words tops\.

This is the requested schema\. In the saved raw batch outputs, some responses instead returned genericdatasets/models/metricskeys; before analysis, the parsing step harmonized both forms into the standardizedGroundTruth\.\.\.JSON fields used by the analysis pipeline \(CSV headers and internal schema\)\. The narrative and visible table labels use “paper\-derived reference inventory”/“Reference” throughout\.

## Appendix GLLM suggestion prompt

The following system prompt was sent identically to GPT\-5\.1, Gemini 3 Pro, and DeepSeek\-V3\.2, with only the JSON key prefixes varying per model \(e\.g\.,GPT51\_suggested\_dataset,gemini\_suggested\_dataset,deepseek\_suggested\_dataset\)\. Each model received only the research question extracted from the paper; Gemini 3 Pro and DeepSeek\-V3\.2 did not receive the paper PDF, abstract, or extracted full text\.

System promptYou are an expert AI research assistant\.Given the following research question:"\{research\_question\}"Please suggest one or more:1\. suitable datasets to address the question\.2\. appropriate machine learning models, architectures, or Large Language Models to use\.3\. relevant evaluation metrics for measuring the model’s performance\.4\. A straightforward pipeline or methodology for solving it\.Only respond with one or more specific dataset names, one or more specific model names, one or more specific evaluation metric names and keep the pipeline structured and short\.Each dataset, model and evaluation metric name must be composed from one to three words tops\.In the pipeline, explain how you want to run the experiment to solve the research question in bullet points\.Respond in valid JSON with keys:\{"\[LLM\]\_suggested\_dataset": \["\.\.\."\],"\[LLM\]\_suggested\_model": \["\.\.\."\],"\[LLM\]\_suggested\_evaluation\_metric": \["\.\.\."\],"\[LLM\]\_suggested\_pipeline": "\.\.\."\}

## Appendix HTaxonomy classification schema

All paper\-derived reference\-inventory and LLM\-suggested entities were classified by GPT\-5\.1 via the OpenAI Batch API using the following taxonomy schema\. This classification step operated on extracted entity names and, where applicable, generated pipeline text, rather than on the original PDFs or full extracted paper texts\. For each entity, the classifier was instructed to assign one or more values from each applicable dimension and return the result in structured JSON\. The schema below lists the intended allowed values for each dimension\.

1. 1\.Dataset dimensions\.For each dataset, classify under: - •Modalities:Text, Audio, Image, Video, Time series, Graph, Spatial, Multimodal\. - •Task types:Classification, Regression, Sequence labeling, Generation, Summarization, Translation, Question answering, Reasoning, Dialogue, Object detection, Forecasting, Retrieval, Alignment, Multimodal integration, Clustering, Reinforcement learning\.111In practice, the classifier occasionally assigned labels from adjacent dimensions \(e\.g\.,*Decision Making*and*Problem Solving*from the cognitive/affective schema;*Safety*,*Ranking*, and*Evaluation*from metric evaluation types;*Segmentation*adjacent to object detection\) to the task\-type field in≤\\leq0\.4% of assignments per source\. A single off\-schema modality label \(*Vision*, 0\.04% of reference\-inventory modality assignments\) also appears\. These were retained in the analysis rather than discarded, which is why some appendix figures show categories beyond the lists above\. - •Domains:General, Media, Scientific/academic, Healthcare, Legal, Economics, Social, Geospatial, Robotics, Vision, Entertainment, Education, Infrastructure, Ontology, Biology, Chemistry, Environmental\. - •Annotation:Fully supervised, Weakly supervised, Self\-supervised, Semi\-supervised, Reinforcement feedback, Crowdsourced, Expert annotations\. - •Size:Small \(<<10K items\), Medium \(10K–100K items\), Large \(\>\>100K items\)\. - •Granularity:Document\-level, Sentence\-level, Token\-level, Frame\-level, Pixel\-level, Object\-level\. - •Linguistic scope:Monolingual \(with language specification from: English, Chinese, Spanish, French, German, Russian, Portuguese, Italian, Dutch, Arabic, Japanese, Korean, Turkish, Polish, Vietnamese, Indonesian, Hebrew, Swedish, Czech, Hungarian, Other\), Multilingual, Cross\-lingual\. - •Cognitive/affective:Attention, Memory, Problem solving, Reasoning, Decision making, Perception, Learning, Cognitive load, Emotion, Empathy, Theory of mind, Social reasoning, Moral cognition, Personality\. - •Data quality:Noisy, Curated\.
2. 2\.Model dimensions\.For each model, classify under: - •Architecture:Transformer \(Encoder/Decoder/Encoder–Decoder\), Generative, CNN, RNN/LSTM, GNN, Tree\-based, Linear, Kernel models, Probabilistic, Reinforcement learning\.222In figures and text, “Transformer \(all\)” aggregates the three Transformer subtypes \(Encoder, Decoder, Encoder–Decoder\) into a single category\. - •Training paradigm:Supervised learning, Self\-supervised learning, Unsupervised learning, Reinforcement learning, Multi\-task learning, Few\-shot learning, Zero\-shot learning, Fine\-tuning, RAG\. - •Provider:OpenAI, Anthropic, Meta AI, Google DeepMind, Mistral AI, Alibaba/Qwen, Cohere, Hugging Face, Stability AI, Microsoft Research, NVIDIA/NeMo, Databricks/MosaicML, DeepSeek, Other/Academic\. - •Openness:Closed, Open\. - •Size:Small \(<<1B\), Medium \(1–10B\), Large \(10–100B\), Extra\-large \(\>\>100B parameters\)\.
3. 3\.Metric dimensions\.For each evaluation metric, classify under: - •Evaluation type:Accuracy, Ranking, Regression, Continuous prediction, Probability, Uncertainty, Fairness, Safety, Efficiency/latency, Explainability, Robustness, User experience\.

Similar Articles

Measuring the Gap Between Human and LLM Research Ideas

Hugging Face Daily Papers

This research paper introduces a framework to measure the distributional gap between human-generated and LLM-generated research ideas, finding that LLM ideas are concentrated around specific opportunity patterns and synthesis methods, while human ideas are more diverse.

Evaluating LLMs as Human Surrogates in Controlled Experiments

arXiv cs.CL

This paper evaluates whether off-the-shelf LLMs can reliably simulate human responses in controlled behavioral experiments by comparing LLM-generated data with human survey responses on accuracy perception. The findings show that while LLMs capture directional effects and aggregate belief-updating patterns, they do not consistently match human-scale effect magnitudes, clarifying when synthetic LLM data can serve as behavioral proxies.