Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
Summary
This study evaluates how AI chatbots like ChatGPT, Claude, and Gemini retrieve clinical studies for medical questions, finding significant performance differences by model and user role, with a bias toward larger sample sizes.
View Cached Full Text
Cached at: 08/17/26, 10:05 AM
# Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
Source: [https://arxiv.org/html/2608.13786](https://arxiv.org/html/2608.13786)
Qingfang LiuAddress:National Institute on Drug Abuse Intramural Research Program, National Institutes of Health, Baltimore, MD, USA E\-mail: qingfang\.liu@nih\.govQiao JinAddress:National Library of Medicine, National Institutes of Health, Bethesda, MD, USA E\-mail: qiao\.jin@nih\.govJoe D\. MenkeAddress:School of Information Sciences, University of Illinois Urbana\-Champaign, Champaign, IL, USA E\-mail: jmenke2@illinois\.eduThorsten KahntAddress:National Institute on Drug Abuse Intramural Research Program, National Institutes of Health, Baltimore, MD, USA E\-mail: thorsten\.kahnt@nih\.govZhiyong Lu†Address:National Library of Medicine, National Institutes of Health, Bethesda, MD, USA †E\-mail: zhiyong\.lu@nih\.gov
###### Abstract
Large language model \(LLM\) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies\. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of the retrieved studies and the factors driving their selection, particularly for newer models with stronger reasoning capabilities\. In this study, we evaluated three recent, general\-purpose LLM chatbots: Claude Sonnet 5, Gemini 3\.1 Pro, and ChatGPT GPT\-5\.5\. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence\-synthesis researcher user roles\. For each review question, we queried each of the three chatbots under each of the three user roles, with four independent repetitions per chatbot–role combination, yielding 720 responses in total \(33chatbots×\\times33user roles×\\times44repetitions×\\times2020review questions\)\. Each chatbot was asked to support its answers with primary clinical citations, which we then benchmarked against the included and excluded study sets of the corresponding Cochrane reviews\. On average, a single chatbot response retrieved 39\.2%±\\pm29\.8% \(mean±\\pmSD across all 720 responses\) of Cochrane included studies, while citing 5\.0%±\\pm9\.4% of Cochrane excluded studies\. Recall of Cochrane included studies varied significantly by model and user role\. ChatGPT achieved higher recall than Claude or Gemini \(63\.1%±\\pm29\.5% vs\. 37\.0%±\\pm23\.8% vs\. 17\.3%±\\pm13\.1%; blocked permutation test,p=2\.0×10−5p=2\.0\\times 10^\{\-5\}\)\. The researcher role yielded higher recall than the clinician or patient roles \(42\.8%±\\pm30\.8% vs\. 38\.6%±\\pm28\.9% vs\. 36\.1%±\\pm29\.3%;p=2\.0×10−5p=2\.0\\times 10^\{\-5\}\)\. Controlling for publication year, citations per year, and open\-access status, sample size was the only independently significant predictor of retrieval \(odds ratio 1\.80 per 1\-unit increase in log sample size, 95% CI 1\.37–2\.36,p=2\.34×10−5p=2\.34\\times 10^\{\-5\}\)\. These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes\. Collected LLM responses and analysis code are available at https://github\.com/QingfangLiu/llm\-evidence\-retrieval\-bias\.
###### keywords
Large Language Models; Evidence Retrieval; Retrieval Bias; Systematic Reviews; Cochrane Reviews; Evidence Synthesis; Medical Question Answering
\\copyrightinfo
© 2024 The Authors\. Open Access chapter published by World Scientific Publishing Company and distributed under the terms of the Creative Commons Attribution Non\-Commercial \(CC BY\-NC\) 4\.0 License\.
## 1Introduction
An increasing number of individuals are consulting chatbots to address medical questions, including requests to retrieve evidence from existing literature to better understand medical conditions\[costagomes2026healthqueries,shpiner2026chatbotfirst\], encouraged in part by strong chatbot performance on medical knowledge benchmarks\[singhal2023clinicalknowledge,Jin2026LLMMedicalResearch\]\. Yet the ability of general\-purpose chatbots to identify and cite relevant clinical studies remains less well characterized\[somer2026studyidentification,chelli2024hallucination,gwon2024scientificsearches,sarker2024natural\], especially in light of prior findings that large language models can fabricate or hallucinate citations\[walters2023fabrication,wu2025sourcecheckup,wang\-etal\-2025\-medcite,jin2026med,topaz2026fabricated\]\. At the same time, systematic reviews and meta\-analyses synthesize primary studies selected by experts to produce higher\-quality conclusions for medical questions\[lefebvre2025searching\]\. This raises an important question: how closely do the studies cited by chatbots align with those included in systematic reviews?
Some prior work has begun to evaluate how well general\-purpose chatbots retrieve primary studies, generally by comparing chatbot\-cited studies against the included\-study lists of published systematic reviews\. These evaluations generally report low recall, whereas precision and citation fabrication vary across systems and evaluation methods\. Somer et al\.\[somer2026studyidentification\]evaluated six publicly accessible systems against the 14 studies in an obstetric meta\-analysis; the best\-performing system, Claude 3\.7, identified five studies\. Chelli et al\.\[chelli2024hallucination\]tested three LLMs across 11 systematic reviews of rotator\-cuff interventions, reporting precision of 0–13\.4% and hallucination rates of 28\.6–91\.4%\. Gwon et al\.\[gwon2024scientificsearches\]found that ChatGPT and Bing AI identified only one and two benchmark randomized trials, respectively, out of 24 randomized trials\. Sidhu et al\.\[sidhu2026trust\]compared the authenticity, quality, and geographic provenance of over 1,200 references across nine contemporary chatbot configurations\. Low et al\.\[low2025clinicalquestions\]found that general\-purpose LLMs produced very few relevant, evidence\-based answers \(2–10% of questions\)\. However, many of these studies relied on models predating advanced reasoning and agentic systems, leaving it unclear how modern, more capable models behave in these contexts\. Additionally, much prior work has focused on a single medical domain\[somer2026studyidentification,chelli2024hallucination,gwon2024scientificsearches\], limiting the generalizability of their conclusions\.
Outside biomedicine, retrieval\-benchmark work such as LitSearch\[ajith2024litsearch\]showed that retrieval performance varies substantially with query specificity and query construction\. PaperAsk evaluated GPT\-4o, GPT\-5, and Gemini\-2\.5\-Flash across a range of scientific fields and found that citation retrieval failed in 48\-98% of multi\-reference queries\[wu2025paperaskbenchmarkreliabilityevaluation\]\. Gao et al\. evaluated five LLMs for 40 randomly selected original articles and found that the models failed to retrieve correct reference data 47\.8% of the time\[gao2026errors\]\. Related work has probed adjacent failure modes rather than recall itself: whether generated citations actually support their associated claims\[wu2025sourcecheckup\], how often citations are outright fabricated or contain metadata errors\[walters2023fabrication\], and how reference\-generation performance varies by model, discipline, and publication recency across large literature\-review corpora\[tang2025literaturereview\]\. Other work asked chatbots to generate search strategies for bibliographic databases such as PubMed and Embase rather than to select studies directly, but this setup may not reflect how general users interact with chatbots\[tam2026database\]\.
A further consideration is that different types of users \(e\.g\., patients, clinicians, evidence synthesis researchers\) may seek access to primary clinical studies for medical question answering\[easterlin2020child\], and whether such variation affects evidence retrieval remains an open question\. Task\-aligned role\-play has been shown to help with reasoning benchmarks\[kong2024roleplay\], while assigning the model a persona does not reliably improve accuracy and can sometimes reduce it\[zheng2024personas\]\. Persona information more broadly shifts model predictions in ways that do not always track genuine human variation\[hu2024persona\]\. Prompt architecture and framing can induce systematic, non\-neutral bias even when no persona is involved\[brucks2025promptarchitecture\], and patient question framing alone has been shown to change a model’s medical conclusions even when the underlying evidence is held fixed\[yun2026framing\]\. However, no prior study has tested whether a user’s self\-identified role changes which primary studies a chatbot retrieves\.
In this study, we evaluated three contemporary general\-purpose chatbot systems \(GPT\-5\.5, Claude Sonnet 5, and Gemini 3\.1 Pro, all equipped with the ability to perform web searches\) on medical questions adapted from 20 Cochrane review topics\. For each topic, prompts were framed from the perspective of a patient, clinician, or evidence\-synthesis researcher, and each role–topic combination was repeated four times\. GPT\-5\.5 achieved the highest recall of Cochrane included\-study sets, and evidence\-synthesis researcher framing yielded higher recall than clinician or patient framing\. After controlling for publication year, citation rate, and open\-access status, larger sample size remained the only significant predictor of study retrieval\. These findings indicate that chatbot\-based evidence retrieval is incomplete, context\-dependent, and biased toward larger clinical trials, with implications for patients who use chatbots for medical information, as well as for clinicians, LLM developers, medical AI researchers, and policymakers\[meyer2023chatgpt,weissenbacher2026enhancing,sahoo2024large\]\.
## 2Methods
### 2\.1Review selection
We used Cochrane intervention reviews published in Issues 6 and 7 of the 2026Cochrane Database of Systematic Reviews\. These were the two most recent complete issues available at the time of the experiment and were published after the knowledge cutoff dates of all three chatbots \(see below\)\. The 20 eligible intervention reviews covered diverse clinical areas, evaluating pharmacological or biologic interventions \(n = 7\), procedural, surgical, or laboratory techniques \(n = 4\), bedside\-care or clinical\-management strategies \(n = 3\), nutritional supplementation \(n = 2\), rehabilitative or conservative physical treatments \(n = 2\), behavioral interventions \(n = 1\), and telehealth or service\-delivery interventions \(n = 1\)\. We excluded protocols and reviews addressing diagnostic, prognostic, or other non\-intervention questions\.
### 2\.2Experiment
We evaluated three widely used, general\-purpose chatbots using the reasoning\-enabled model configurations listed in Table[2\.2](https://arxiv.org/html/2608.13786#S2.SS2)\. These systems were selected to represent prominent consumer chatbot ecosystems\. To provide external context for model capability, we additionally reference Humanity’s Last Exam \(HLE\), a multimodal, frontier\-level academic benchmark comprising 3,000 expert\-authored questions across dozens of subjects and designed to probe advanced knowledge and reasoning beyond standard benchmark suites\. Each response was generated independently in a fresh anonymous web session with no prior conversational context\. Data were collected from the U\.S\. East Coast between mid\- and late July 2026\.
\\tblChatbot systems, configurations, and benchmark performance\.\\topruleModelKnowledge cutoffRelease dateReasoning configurationHLE \(no tools\)HLE \(with tools\)\\colruleAnthropic Claude Sonnet 5January 2026June 30, 2026Medium effort; thinking enabled43\.2%57\.4%Google Gemini 3\.1 ProJanuary 31, 2025February 19, 2026Extended thinking44\.4%51\.4%OpenAI GPT\-5\.5December 1, 2025April 23, 2026High reasoning41\.4%52\.2%\\botrule\\tabnoteNote: Knowledge cutoff indicates the latest date through which a model’s documented training knowledge extends\. HLE \(Humanity’s Last Exam\) is a multimodal benchmark comprising 3,000 expert\-authored questions\.
For each review question, we developed three prompts representing different user roles\. Except for the user role stated at the beginning of each prompt, the wording was kept as consistent as possible\. We also explicitly instructed the chatbots to rely on primary evidence \(e\.g\., original research studies and clinical trials\) rather than secondary evidence \(e\.g\., systematic reviews, meta analyses, narrative reviews, committee opinions, practice guidelines, editorials, or commentaries\) when answering\. To account for variability in chatbot responses, each prompt was run four times\. An example is provided in Table\.
\\tblExample of the three user\-role prompt variants for one review question\.\\topruleUser roleRole\-specific opening\\colrulePatientI have high blood pressure, and I am considering taking medicine to help me lose weight\.ClinicianI am a clinician caring for people with high blood pressure who are considering taking medicine to help them lose weight\.Evidence\-synthesis researcherI am an evidence\-synthesis researcher studying weight\-loss medicines in people with high blood pressure\.Similar Articles
Study: AI-powered chatbots respond to everyday health-related questions from general users with nearly 76% accuracy
A Penn State study found that AI chatbots like ChatGPT respond to everyday health queries with nearly 76% accuracy, raising concerns about trustworthiness in real-world healthcare applications. The research highlights that AI tools may be best used by physicians rather than patients.
Navigating health questions with ChatGPT
OpenAI publishes guidance on using ChatGPT to navigate health-related questions, addressing how users can leverage the model while understanding its limitations in medical contexts.
ChatGPT, Gemini, Claude, Grok Fail Accuracy Test on Election Topics: Forum AI
A study by Forum AI found that major chatbots like ChatGPT, Gemini, Claude, and Grok fail to provide accurate and unbiased election information, with 90% of responses containing errors or bias.
An A.I. Aggregator?
A user shares their experience using ChatGPT for complex medical caregiving and proposes the idea of aggregating multiple AI models to improve reliability by seeking consensus among different LLMs.
Improving health intelligence in ChatGPT
OpenAI announces significant improvements in health-related responses within ChatGPT using GPT-5.5 Instant, achieving accuracy comparable to frontier models and reducing factuality issues by 71% through physician-led evaluations.