When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG
Summary
A large-scale study across 5 models (7B–72B), 10 biomedical QA datasets, 4 retrieval methods, and 4 corpora finds that RAG yields only small and inconsistent gains (1–2 points) over no-retrieval baselines in biomedical question answering. The study concludes that the main bottleneck is not retrieval quality but models' limited ability to effectively use retrieved evidence.
View Cached Full Text
Cached at: 06/05/26, 02:12 AM
# When Retrieval Doesn’t Help: A Large-Scale Study of Biomedical RAG
Source: [https://arxiv.org/html/2606.04127](https://arxiv.org/html/2606.04127)
Anthony Rios The University of Texas at San Antonio \{Erfan\.Nourbakhsh, Rocky\.Slavin, Ke\.Yang, Anthony\.Rios\}@utsa\.edu
###### Abstract
Medical question answering is a high\-stakes setting where factual errors can have serious consequences\. Retrieval\-augmented generation \(RAG\) is widely viewed as a promising solution, and prior work has reported substantial gains for large medical QA models\. We revisit this assumption across a broad range of open\-weight instruction\-tuned models spanning 7B to 72B parameters\. Across five models, ten biomedical QA datasets, four retrieval methods, and four retrieval corpora, we find that retrieval yields only small and inconsistent improvements over a no\-retrieval baseline, typically within 1–2 points\. In contrast, the choice of backbone model has a much larger effect than the choice of retriever or corpus, and expert and layman retrieval sources perform similarly in most settings\. These results suggest that the main bottleneck is not retrieval quality alone, but the model’s limited ability to use retrieved evidence effectively\. Code is available here:[https://github\.com/erfan\-nourbakhsh/BioMedicalRAG](https://github.com/erfan-nourbakhsh/BioMedicalRAG)
When Retrieval Doesn’t Help: A Large\-Scale Study of Biomedical RAG
Erfan Nourbakhsh, Rocky Slavin, Ke Yang, and Anthony RiosThe University of Texas at San Antonio\{Erfan\.Nourbakhsh, Rocky\.Slavin, Ke\.Yang, Anthony\.Rios\}@utsa\.edu
## 1Introduction
Accurate and reliable medical question answering is a high\-stakes problem, where errors can have direct consequences for patient safety\. Large language models \(LLMs\) have recently shown strong performance on a range of biomedical question answering tasks[Singhalet al\.](https://arxiv.org/html/2606.04127#bib.bib24); Hendryckset al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib10)\); Jinet al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib6)\)\. However, they remain prone to hallucination, producing fluent but factually incorrect responsesJiet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib25)\), and to knowledge staleness due to their reliance on fixed training corpora\. In the medical domain, these limitations are especially problematic because even small factual errors can lead to harmful downstream decisions\.
Retrieval\-augmented generation \(RAG\)Lewiset al\.\([2020](https://arxiv.org/html/2606.04127#bib.bib23)\)has become a leading approach for addressing these limitations by grounding model outputs in retrieved external evidence\. By incorporating supporting documents at inference time, RAG offers a mechanism for improving factuality, transparency, and access to more current knowledge\. As a result, RAG has been adopted widely in biomedical NLP, where recent work has reported substantial gains from retrieval\-based methods\. For example,Xionget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib13)\)showed that MedRAG improves biomedical QA accuracy by as much as 18% over chain\-of\-thought prompting, whileTanget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib15)\)found that multi\-agent LLM systems can further improve medical reasoning performance\. These findings have led to growing interest in retrieval\-centered biomedical QA systems, with increasing attention to the choice of corpora, retrieval methods, and model backbones\.
Figure 1:Overview of our motivation and main finding: across models from 7B to 72B, retrieval yields only small gains, suggesting that the main bottleneck is evidence use rather than retrieval quality\.However, an important gap remains\. Prior systematic studies of medical RAGXionget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib13)\)have largely focused on large proprietary or 70B\-scale models \(GPT\-3\.5, GPT\-4, Mixtral\-8x7B, Llama2\-70B\) under zero\-shot multiple\-choice evaluation, leaving unclear whether their gains carry over to 7B–8B models that are far more practical under real hardware constraints\. Existing evaluations have also focused primarily on expert\-level biomedical questions, with little attention to consumer\-health queries or community\-generated retrieval sources\.
In this paper, we revisit biomedical RAG under a substantially different and more comprehensive setting\. We evaluate five open\-weight instruction\-tuned models spanning 7B to 72B parameters: Qwen2\.5\-7B\-InstructYanget al\.\([2025](https://arxiv.org/html/2606.04127#bib.bib20)\), Llama\-3\.1\-8B\-InstructGrattafioriet al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib21)\), Mistral\-7B\-Instruct\-v0\.3Jianget al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib22)\), LLaMA\-3\.1\-70B\-InstructGrattafioriet al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib21)\), and Qwen2\.5\-72B\-InstructYanget al\.\([2025](https://arxiv.org/html/2606.04127#bib.bib20)\), across ten biomedical QA datasets spanning both lay and expert questions and covering both open\-ended and multiple\-choice formats\. We compare four retrieval methods, BM25, TF\-IDF, MedCPT, and Hybrid RRF, across four retrieval corpora, including both expert biomedical resources and consumer\-facing health sources: PubMed abstracts, medical textbooks, Yahoo Answers, and HealthCareMagic\. We also evaluate against a no\-retrieval baseline in order to isolate the contribution of retrieval itself\.
Our results challenge the prevailing picture from prior studies\. Across all five models, retrieval yields only small and inconsistent gains: the gap between the best retrieval configuration and the no\-retrieval baseline is usually within 1–2 points \(e\.g\., BERTScore 62\.88 vs\. 61\.72 for Llama\-8B, 63\.23 vs\. 61\.28 for Qwen\-7B\), and differences across retrieval corpora are similarly modest even for the larger 70B models\. By contrast, backbone model choice has a much larger effect than retriever or corpus selection, and expert versus lay retrieval sources differ by less than 2 points in most settings\. Figure[1](https://arxiv.org/html/2606.04127#S1.F1)illustrates the key implication: the limiting factor is not retrieval quality but the generator’s capacity to incorporate retrieved evidence\.
Our contributions are: \(1\) A large\-scale evaluation of biomedical RAG covering 5 models from 7B to 72B parameters, 10 QA datasets, 4 retrieval methods, and 4 retrieval corpora\. \(2\) We show that retrieval yields only small and inconsistent improvements across all model scales \(typically within 2 points\), challenging the gains reported in prior large\-model studies\. \(3\) We show that backbone model choice matters more than retriever or corpus choice, and provide evidence that the main bottleneck is the model’s weak use of retrieved evidence\.
Figure 2:Experimental pipeline overview\.
## 2Related Work
Retrieval\-Augmented Generation\.RAG was introduced byLewiset al\.\([2020](https://arxiv.org/html/2606.04127#bib.bib23)\)as a method to enhance language models on knowledge\-intensive tasks by conditioning generation on documents retrieved from a non\-parametric memory\. The approach combines a parametric sequence\-to\-sequence model with a dense passage retrieval componentKarpukhinet al\.\([2020](https://arxiv.org/html/2606.04127#bib.bib26)\)and has been extended in numerous directions, including iterative retrievalTrivediet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib27)\); Shaoet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib28)\), self\-reflective retrievalAsaiet al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib29)\), and query rewritingMaet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib30)\); for a broad survey of RAG paradigms and architectures, seeGaoet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib31)\)\. In the biomedical domain, RAG has been applied to clinical decision supportXionget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib13)\), scientific literature searchJinet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib18)\), and consumer health QALiet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib9)\)\. However, most prior biomedical RAG studies either lack systematic comparison across retrieval configurations or are limited in dataset coverage\.
Benchmarking Medical RAG\.The most directly related work to ours is the MIRAGE benchmark and MedRAG toolkit byXionget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib13)\), which evaluates 41 combinations of corpora, retrievers, and backbone LLMs on five medical QA datasets restricted to multiple\-choice questions\. MIRAGE shows that RAG can improve LLM accuracy by up to 18% and identifies PubMed combined with BM25 or MedCPT as strong retrieval configurations\. However, MIRAGE exclusively uses zero\-shot prompting and evaluates primarily large models \(GPT\-4, GPT\-3\.5, Mixtral\-8×7B, Llama2\-70B\), leaving open the question of whether these gains hold for smaller, more widely deployable models\. Concurrent large\-model evaluations, such asNoriet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib32)\), who find that GPT\-4 surpasses the USMLE passing threshold by over 20 points even without retrieval augmentation, further underscore that model scale is a critical confound in existing medical benchmarks\.Tanget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib15)\)propose a zero\-shot multi\-agent framework achieving competitive GPT\-4 performance on MMLU Medical, yet neither this nor MIRAGE examines retrieval for open\-ended or consumer\-health queries at the 7–8B scale\.Shiet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib33)\)show that irrelevant retrieved passages can mislead LLMs, a concern especially acute for smaller models, whileOvadiaet al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib34)\)find retrieval augmentation outperforms knowledge fine\-tuning primarily for large models, further motivating our cross\-scale evaluation\.
Biomedical Question Answering Datasets\.Biomedical QA has long served as a testbed for evaluating NLP systems in medicineKritharaet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib11)\); Jinet al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib6)\); Palet al\.\([2022](https://arxiv.org/html/2606.04127#bib.bib8)\); Hendryckset al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib10)\)\. Expert\-oriented benchmarks such as BioASQNentidiset al\.\([2025](https://arxiv.org/html/2606.04127#bib.bib2)\), MedQA\-USMLEJinet al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib6)\), and MedMCQAPalet al\.\([2022](https://arxiv.org/html/2606.04127#bib.bib8)\)test clinical and examination\-level knowledge, while consumer\-health datasets such as MeQSumBen Abacha and Demner\-Fushman \([2019b](https://arxiv.org/html/2606.04127#bib.bib1)\), MedRedQANguyenet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib3)\), MedicationQAAbachaet al\.\([2019](https://arxiv.org/html/2606.04127#bib.bib5)\), MASH\-QAZhuet al\.\([2020](https://arxiv.org/html/2606.04127#bib.bib7)\), and ChatDoctor\-iCliniqLiet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib9)\)reflect more informal, everyday health information needs\. MedQuADBen Abacha and Demner\-Fushman \([2019a](https://arxiv.org/html/2606.04127#bib.bib4)\)and MMLU MedicalHendryckset al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib10)\)bridge the two groups by covering structured NIH\-sourced QA and standardised medical knowledge\. Despite this rich landscape, most RAG studies focus on MCQ\-format expert benchmarks and omit the open\-ended and layman query types that constitute the bulk of real\-world health information needs, a gap we directly address\.
Retrieval Methods in Biomedical NLP\.Sparse retrieval methods have been dominant in biomedical information retrieval\. BM25Robertson and Zaragoza \([2009](https://arxiv.org/html/2606.04127#bib.bib16)\), a probabilistic bag\-of\-words ranking function, remains a strong baseline and is adopted as the primary retriever in MedRAGXionget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib13)\)\. TF\-IDFSPARCK JONES \([1972](https://arxiv.org/html/2606.04127#bib.bib17)\), a simpler precursor that models term specificity without BM25’s saturation and length normalisation, provides a useful lower bound for sparse retrieval\.
Dense retrieval with domain\-adapted encoders has gained traction\. MedCPTJinet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib18)\)was trained contrastively on large\-scale PubMed search logs and demonstrates strong zero\-shot biomedical retrieval, outperforming general\-domain encoders on medical tasks\. Fusion methods such as Reciprocal Rank Fusion \(RRF\)Cormacket al\.\([2009](https://arxiv.org/html/2606.04127#bib.bib19)\)combine sparse and dense ranked lists and have been shown to outperform individual retrievers without requiring additional training\. While MedRAG includes RRF as a configuration, it does not systematically isolate the contribution of each component retriever across diverse query types and corpora, which we do in this study\.
## 3Experiments
Our experiments systematically compare sparse, dense, and hybrid retrieval strategies across four corpora, ten QA datasets spanning expert and layman health queries, and five open\-weight instruction\-tuned models ranging from 7B to 72B parameters\. Figure[2](https://arxiv.org/html/2606.04127#S1.F2)provides a visual overview of the full experimental pipeline\.
Evaluation Dataset and Knowledge Base\.
Evaluation Datasets\.We evaluate across ten biomedical and consumer\-health question answering datasets grouped into two user types:laymandatasets reflecting everyday consumer\-health language, andexpertdatasets targeting biomedical professionals or medical students\. Dataset statistics and split sizes are summarised in Table[7](https://arxiv.org/html/2606.04127#A1.T7)in Appendix[A](https://arxiv.org/html/2606.04127#A1)\. For all datasets, examples lacking a question or a reference answer are discarded before any split is finalised, and whenever random sampling is needed it is performed with a fixed seed of 42\.
Layman datasets\.MeQSumBen Abacha and Demner\-Fushman \([2019b](https://arxiv.org/html/2606.04127#bib.bib1)\)contains 1,000 consumer health questions from the U\.S\. National Library of Medicine\. FollowingZhanget al\.\([2022](https://arxiv.org/html/2606.04127#bib.bib14)\), we reserve 500 examples for evaluation and use the remaining 500 as the few\-shot query pool\.MedRedQANguyenet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib3)\)provides over 51,000 consumer question–physician answer pairs from Reddit’s/r/AskDocs; we sample 1,000 evaluation examples from the official test split \(5,099 examples\) and combine the training \(40,792\) and validation \(5,100\) splits into the query pool\.MedicationQAAbachaet al\.\([2019](https://arxiv.org/html/2606.04127#bib.bib5)\)contains 690 real consumer medication questions; we randomly sample 500 for evaluation and retain the remaining 189 as the query pool\.MASH\-QAZhuet al\.\([2020](https://arxiv.org/html/2606.04127#bib.bib7)\)offers over 34,000 WebMD\-derived healthcare Q&A pairs; we randomly sample 1,000 examples from the official test file \(2,614 entries\) and use the full training set \(19,989 examples\) as the query pool\.ChatDoctor\-iCliniqLiet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib9)\)comprises 7,321 real patient–physician conversations from iCliniq\.com; we randomly sample 1,000 for evaluation and retain the remaining 6,321 as the query pool\.
Expert datasets\.BioASQ Task BNentidiset al\.\([2025](https://arxiv.org/html/2606.04127#bib.bib2)\)provides expert biomedical questions grounded in PubMed literature; following the official benchmark protocol, we use the Task 13B golden test set \(restricted to summary\-type questions, 80 examples\) and the Task 13B training set \(1,283 examples\) as the query pool\.MedQuADBen Abacha and Demner\-Fushman \([2019a](https://arxiv.org/html/2606.04127#bib.bib4)\)contains 47,457 medical Q&A pairs from 12 NIH websites, of which 16,407 are publicly available; we randomly sample 1,000 for evaluation and use the remaining 15,407 as the query pool\.MedQA\-USMLEJinet al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib6)\)provides USMLE clinical vignette MCQs; we use the official test split \(1,273 examples\) for evaluation and the official training split \(10,178 examples\) as the query pool\.MedMCQAPalet al\.\([2022](https://arxiv.org/html/2606.04127#bib.bib8)\)contains 194k\+ MCQs from AIIMS and NEET PG medical entrance exams; as the official test split is unlabelled, we randomly sample 1,000 from the validation set \(6,150 examples\) for evaluation and use the full training set \(182,822 examples\) as the query pool\.MMLU MedicalHendryckset al\.\([2021](https://arxiv.org/html/2606.04127#bib.bib10)\): followingTanget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib15)\), we restrict to six medical sub\-tasks,anatomy,clinical\_knowledge,college\_biology,college\_medicine,medical\_genetics, andprofessional\_medicine, totalling 1,242 examples\. Roughly 100 examples per sub\-task \(600 in total\) are used for evaluation; the remaining 642 form the query pool\.
Knowledge Bases\.We build four retrieval corpora covering both expert biomedical and layman health domains, as summarised in Table[8](https://arxiv.org/html/2606.04127#A1.T8)in Appendix[A](https://arxiv.org/html/2606.04127#A1)\. All corpora are indexed as whole records without further chunking\. For Q&A\-style corpora \(Yahoo Answers and HealthCareMagic\), each document concatenates the question or title with the corresponding answer body\.
BioASQ / PubMedKritharaet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib11)\)consists of 16\.2 million PubMed abstracts with human\-assigned MeSH annotations and serves as the primary expert\-domain knowledge base\.
Medical TextbooksXionget al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib13)\)provides 125,847 retrieval\-friendly chunks \(≤\\leq1,000 characters each\) drawn from 18 authoritative biomedical textbooks spanning anatomy, physiology, pharmacology, pathology, and clinical medicine\.
Yahoo AnswersYahoo\! Research \([2009](https://arxiv.org/html/2606.04127#bib.bib12)\)is an open\-domain community Q&A corpus; from the original 1\.4 million records we retain 1,238,506 after quality filtering, discarding entries whose answer body contains fewer than five words or whose combined question–answer text falls below ten words\.
HealthCareMagicLiet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib9)\)contains 112,165 real\-world patient symptom queries paired with detailed physician responses across more than ten clinical specialties\.
Table 1:ROUGE\-L by model and retrieval corpus \(open\-ended datasets\)\.Retrieval Approaches\.We compare four retrieval strategies that differ in their document and query representations\.
BM25Robertson and Zaragoza \([2009](https://arxiv.org/html/2606.04127#bib.bib16)\)is a classic sparse probabilistic retrieval model that scores documents by the weighted overlap of query terms, applying a term\-frequency saturation function and a document\-length normalisation penalty\. BM25 has long served as a strong baseline for ad\-hoc retrieval and remains competitive with many neural approaches\. We adopt BM25 parametersk1=0\.9k\_\{1\}\{=\}0\.9andb=0\.4b\{=\}0\.4, and apply a title\-boost factor of 2 by repeating title tokens at indexing time to approximate field\-weighted BM25F scoring\.
TF\-IDFSPARCK JONES \([1972](https://arxiv.org/html/2606.04127#bib.bib17)\)represents both documents and queries as bag\-of\-words vectors weighted by term frequency–inverse document frequency, and ranks candidates by cosine similarity\. Unlike BM25, TF\-IDF applies no term\-frequency saturation or document\-length penalty, making it a simpler baseline for sparse lexical matching\. We build a TF\-IDF index with a vocabulary capped at 50,000 features and standard English stop\-word removal\.
MedCPTJinet al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib18)\)is a biomedical dense retrieval model consisting of aquery encoderand anarticle encodertrained contrastively on large\-scale PubMed user search logs\. Documents are encoded offline by the article encoder and stored as L2\-normalised embeddings; at query time, the query encoder produces a query embedding and retrieval proceeds by maximum inner\-product search\. By capturing semantic similarity beyond exact term overlap, MedCPT is particularly well\-suited to the biomedical domain\.
Hybrid BM25 \+ MedCPT via RRFCormacket al\.\([2009](https://arxiv.org/html/2606.04127#bib.bib19)\)combines the BM25 and MedCPT ranked lists using Reciprocal Rank Fusion \(RRF\)\. Each documentddranked at positionrrin a ranked list receives a score1k\+r\\frac\{1\}\{k\{\+\}r\}; withk=60k\{=\}60, and the scores are summed across both lists\. The final ranking is by descending combined RRF score\. RRF is parameter\-light and has been shown to consistently outperform individual rankers as well as more complex score\-fusion methodsCormacket al\.\([2009](https://arxiv.org/html/2606.04127#bib.bib19)\)\.
For all retrieval conditions, we retrieve the topk=5k\{=\}5documents and concatenate them as the retrieved context prepended to the generator prompt\.
Implementation Details\.All generation experiments are conducted with five open\-source instruction\-tuned models spanning two scales\. The 7–8B models areQwen2\.5\-7B\-InstructYanget al\.\([2025](https://arxiv.org/html/2606.04127#bib.bib20)\),Llama\-3\.1\-8B\-InstructGrattafioriet al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib21)\), andMistral\-7B\-Instruct\-v0\.3Jianget al\.\([2023](https://arxiv.org/html/2606.04127#bib.bib22)\), each servable on a single GPU\. The 70B\-scale models areLLaMA\-3\.1\-70B\-InstructGrattafioriet al\.\([2024](https://arxiv.org/html/2606.04127#bib.bib21)\)andQwen2\.5\-72B\-InstructYanget al\.\([2025](https://arxiv.org/html/2606.04127#bib.bib20)\), which serve as large\-scale reference points to contextualise the small\-model results\.
All models are run in half\-precision \(FP16\) with greedy decoding and a maximum of 300 newly generated tokens per response\.
Experimental Setting\.Each experimental condition is defined by a triple \(retriever,corpus,query dataset\)\. The retriever dimension covers five options:No retrieval\(baseline\), BM25, TF\-IDF, MedCPT, and Hybrid \(BM25 \+ MedCPT via RRF\)\. The corpus dimension covers four knowledge bases: BioASQ/PubMed and Medical Textbooks as expert corpora, and Yahoo Answers and HealthCareMagic as layman corpora \(for the baseline condition both retriever and corpus are set to none\)\. The query dimension covers the ten datasets described in Section[3](https://arxiv.org/html/2606.04127#S3), split evenly between layman \(MeQSum, MedRedQA, MedicationQA, MASH\-QA, ChatDoctor\-iCliniq\) and expert \(BioASQ Task B, MedQuAD, MedQA\-USMLE, MedMCQA, MMLU Medical\) user types\.
For thew/o RAGcondition the model receives only the question in its prompt, with no retrieved context\. For retrieval\-augmented conditions, the top\-kkretrieved passages are prepended to the question in a fixed prompt template\. Each condition is run independently for every generator model, and all per\-dataset query pools described in Section[3](https://arxiv.org/html/2606.04127#S3)are also available for few\-shot prompting ablations\. The complete set of conditions spans every combination of retriever, corpus, and query dataset, yielding a large\-scale cross\-model, cross\-retriever, cross\-dataset evaluation\.
## 4Results
We present results separately for open\-ended QA, evaluated with ROUGE\-L as the primary metric \(ROUGE\-1, ROUGE\-2, METEOR, BLEU, and BERTScore in Appendix[C](https://arxiv.org/html/2606.04127#A3)\), and for multiple\-choice QA, evaluated with accuracy\.
Open\-ended QA\.Table[1](https://arxiv.org/html/2606.04127#S3.T1)reports ROUGE\-L across seven open\-ended datasets \(five layman and two expert\), averaged over all retrieval conditions per corpus\. Across all models, retrieval yields small and inconsistent improvements over the no\-retrieval baseline\. The largest gains appear on the BioASQ open\-ended task, where the BioASQ/PubMed corpus consistently provides the strongest lift: for example, LLaMA\-3\.1\-8B improves from 21\.65 to 27\.43 ROUGE\-L\. However, for the remaining six datasets, changes from the baseline are typically under 1 ROUGE\-L point and often negative\. Averaged across all seven datasets, the maximum retrieval benefit over no\-retrieval is 1\.18 points \(LLaMA\-3\.1\-8B: 13\.06 baseline vs\. 14\.24 with BioASQ\); for all other models the gain is smaller still\. ROUGE\-1, ROUGE\-2, METEOR, BLEU, and BERTScore results \(Appendix[C](https://arxiv.org/html/2606.04127#A3)\) show the same pattern\.
A consistent observation is that backbone model choice matters far more than retrieval configuration\. Mistral\-7B lags behind both LLaMA\-3\.1\-8B and Qwen2\.5\-7B regardless of the retrieval setup, and the 70B\-scale models \(LLaMA\-3\.1\-70B, Qwen2\.5\-72B\) are consistently stronger than all 7–8B variants\. The gap between any two retrieval conditions for the same model is almost always smaller than the gap between two different backbone models using the same conditions\. Expert and layman retrieval corpora produce similar results in most open\-ended settings, differing by less than 1 ROUGE\-L point on average\.
Table 2:Accuracy by model and retrieval corpus\. MCQA denotes MedMCQA, MQA denotes MedQA, and HCM denotes HealthCareMagic\.Multiple\-choice QA\.Table[2](https://arxiv.org/html/2606.04127#S4.T2)reports accuracy grouped by dataset subset \(MCQA: MedQA \+ MedMCQA; QA: open\-ended expert; MMLU: six MMLU medical subjects\)\. For smaller models \(LLaMA\-3\.1\-8B, Mistral\-7B, Qwen2\.5\-7B\), retrieval frequently*hurts*accuracy relative to the no\-retrieval baseline\. Mistral\-7B drops from 75\.7 to 68\.6–72\.3 across all retrieval corpora\. The larger models \(LLaMA\-3\.1\-70B, Qwen2\.5\-72B\) are more robust, maintaining accuracy within 1–2 points of the baseline across all conditions, but still show no consistent gain\. As with the open\-ended setting, backbone choice dominates: Qwen2\.5\-72B’s no\-retrieval accuracy of 85\.6 exceeds the best retrieval configuration of any 7B model by over 2 points\.
Effect of Retrieval Method\.Table[3](https://arxiv.org/html/2606.04127#S4.T3)reports accuracy averaged across the three close\-ended datasets from Table[2](https://arxiv.org/html/2606.04127#S4.T2)\(MedMCQA, MedQA\-USMLE, and MMLU Medical\), broken down by retrieval method rather than corpus\. Table[4](https://arxiv.org/html/2606.04127#S5.T4)reports ROUGE\-L averaged across the seven open\-ended datasets from Table[1](https://arxiv.org/html/2606.04127#S3.T1)\(BioASQ, ChatDoctor\-iCliniq, MashQA, MedicationQA, MedQuAD, MedRedQA, and MeQSum\), again broken down by retrieval method\. Together, these two tables allow direct comparison of BM25, Hybrid \(RRF\), MedCPT, and TF\-IDF across question types and retrieval corpora\. Differences among methods are within 1–2 points for any model–corpus combination\. The Hybrid retriever shows marginal advantages in several configurations, but no method consistently dominates\. MedCPT, despite domain\-specific training, does not systematically outperform lexical BM25\. Full per\-metric breakdowns by retriever type \(ROUGE\-1, ROUGE\-2, METEOR, BLEU, BERTScore\) are in Appendix[C](https://arxiv.org/html/2606.04127#A3)\.
Table 3:Accuracy by retrieval method \(close\-ended datasets\)\.
## 5Ablation Study
We conduct two ablations to understand how retrieval depth and few\-shot context affect performance\. Both use a stratified subset of the test queries with BM25 retrieval from BioASQ\. Additional open\-ended metric trends are visualised in Appendix[D](https://arxiv.org/html/2606.04127#A4)\(Figures[7](https://arxiv.org/html/2606.04127#A4.F7)and[8](https://arxiv.org/html/2606.04127#A4.F8)\)\.
Table 4:ROUGE\-L by retrieval method \(open\-ended datasets\)\.Figure 3:Close\-ended accuracy across shot counts \(1, 3, 5, 10\)\.Number of Retrieved Documents \(Top\-kk\)\.Figures[5](https://arxiv.org/html/2606.04127#S5.F5)and[6](https://arxiv.org/html/2606.04127#S5.F6)show accuracy and ROUGE\-L askkvaries over\{1,3,5,10,25,50\}\\\{1,3,5,10,25,50\\\}\. For open\-ended metrics, performance reaches a plateau byk=5k\{=\}5: ROUGE\-L changes by less than 0\.2 points betweenk=5k\{=\}5andk=50k\{=\}50for all models, indicating that additional retrieved documents add no useful signal once the context budget is satisfied\. For close\-ended accuracy the picture is less uniform: LLaMA\-3\.1\-8B peaks atk=5k\{=\}5\(72\.83%\) before declining, while Qwen2\.5\-7B and LLaMA\-3\.1\-70B reach their best performance atk≥25k\{\\geq\}25\. Mistral\-7B declines steadily afterk=3k\{=\}3, reaching 51\.22% atk≥25k\{\\geq\}25\. These results confirm thatk=5k\{=\}5is a reasonable default: it matches or closely approaches the optimum for most models while keeping context length manageable\. Additional open\-ended metric trends across allkkvalues are shown in Figure[8](https://arxiv.org/html/2606.04127#A4.F8)in the appendix\.
Figure 4:Open\-ended ROUGE\-L across shot counts \(1, 3, 5, 10\)\.Few\-shot Prompting\.Figures[3](https://arxiv.org/html/2606.04127#S5.F3)and[4](https://arxiv.org/html/2606.04127#S5.F4)show accuracy and ROUGE\-L as the number of in\-context examples varies over\{1,3,5,10\}\\\{1,3,5,10\\\}\. Larger models \(LLaMA\-3\.1\-70B, Qwen2\.5\-72B\) are essentially unaffected by shot count across all metrics, suggesting they can extract the task pattern from a single example or from zero\-shot prompting equally well\. In contrast, smaller 7–8B models show sharp degradation at 5 and 10 shots: LLaMA\-3\.1\-8B accuracy collapses from 82\.89% \(1\-shot\) to 10\.06% \(10\-shot\), and ROUGE\-L drops from 14\.29 to 8\.38, as the long few\-shot context overwhelms the model’s ability to locate the target instruction\. Mistral\-7B and Qwen2\.5\-7B follow the same pattern\. Notably, 3\-shot prompting is the sweet spot for open\-ended ROUGE\-L: LLaMA\-3\.1\-8B reaches 17\.19 at 3 shots \(vs\. 14\.29 at 1\-shot\), and Qwen2\.5\-7B reaches 17\.22, before degrading at higher shot counts\. For MCQ accuracy, even 3 shots already reduces performance for most small models, pointing to the inherent tension between providing helpful demonstrations and staying within the model’s effective context capacity\. Additional open\-ended metric trends across all shot counts are shown in Figure[7](https://arxiv.org/html/2606.04127#A4.F7)in the appendix\.
Figure 5:Close\-ended accuracy across top\-kk\(1, 3, 5, 10, 25, 50\)\.Figure 6:Open\-ended ROUGE\-L across top\-kkvalues\.Table 5:Accuracy of LLMs across retrieval methods in the oracle retrieval setting, where all retrieved documents are relevant \(clean context\)\.Quality of the Retrieval Analysis\.Tables[5](https://arxiv.org/html/2606.04127#S5.T5)and[6](https://arxiv.org/html/2606.04127#S5.T6)show two important problems for retrieval\-augmented generation in the biomedical domain\. For this analysis, we use the BioASQ corpus as the retrieval source and evaluate on PubMedQAJinet al\.\([2019](https://arxiv.org/html/2606.04127#bib.bib35)\), a benchmark of expert\-annotated yes/no/maybe biomedical research questions derived from PubMed abstracts\. Since both the retrieval corpus and evaluation dataset come from PubMed, this provides a controlled setting for studying whether retrieved biomedical papers help models answer research questions\. To evaluate retrieval quality, we use an LLM\-as\-a\-judge framework to determine whether the retrieved context contains enough information to answer the question correctly\. We then select 100 questions where all retrieval methods retrieved context judged to be relevant\. The questions are the same across all retrieval methods, but the retrieved documents can differ depending on the retriever\.
Table[5](https://arxiv.org/html/2606.04127#S5.T5)shows that even when all retrieved contexts contain the correct information, retrieval only leads to limited and inconsistent improvements\. For example, LLaMA3\.1\-70B improves substantially with BM25 retrieval \(0\.410→\\rightarrow0\.660\), while Qwen2\.5\-72B shows almost no improvement across retrieval methods\. In several cases, simple sparse retrieval methods such as BM25 and TF\-IDF perform better than MedCPT\. These results suggest that retrieving relevant evidence alone is not enough to guarantee better performance\. Instead, many models still struggle to correctly use and reason over the retrieved information\. Table[6](https://arxiv.org/html/2606.04127#S5.T6)further shows that current models are highly sensitive to irrelevant context\. When we add 20 unrelated documents to the retrieved evidence, performance drops substantially across nearly all models and retrieval methods\. For example, LLaMA3\.1\-70B decreases from 0\.660 to 0\.260 under BM25 retrieval, while Mistral\-7B drops from 0\.530 to 0\.340 under TF\-IDF retrieval\. In many cases, performance becomes worse than using no retrieval at all\. Overall, these results show that current biomedical RAG systems remain brittle\. Even when relevant evidence is retrieved successfully, small amounts of distracting context can strongly reduce answer accuracy\.
Table 6:Accuracy of LLMs across retrieval methods in the noisy retrieval setting, where 20 unrelated documents are mixed with retrieved results \(distracted context\)\.Implications\.Our results suggest a more cautious view of biomedical RAG\. Retrieval can help, but only when the system retrieves information that is actually relevant to the question\. This is not guaranteed, especially when the answer is absent from the corpus or when the retrieved passages are only loosely related\. In these cases, retrieval may add little useful information and can introduce misleading context\.
Even when relevant evidence is retrieved, the model still has to understand and use it correctly\. Our clean retrieval analysis shows that relevant context does not always improve performance, suggesting that evidence use is a major bottleneck\. The noisy retrieval results make this concern stronger: adding unrelated documents to useful evidence often hurts performance, sometimes making RAG worse than no retrieval at all\. Future biomedical RAG systems, therefore, need better evidence filtering, reranking, and generation methods that can identify useful passages while ignoring distractors\.
## 6Conclusion
We presented a large\-scale evaluation of retrieval\-augmented generation for biomedical question answering using five open\-weight, instruction\-tuned models ranging from 7B to 72B parameters\. Across all five models, ten datasets, four retrieval methods, and four retrieval corpora, retrieval yields only small and inconsistent improvements over a no\-retrieval baseline, typically within 1–2 points on any metric\. In contrast, backbone model choice has a substantially larger effect: the gap between a 7B model and its 70B counterpart often exceeds the gain from any retrieval configuration\. Expert and layman retrieval corpora also perform similarly in most settings, and differences across retrieval methods \(BM25, TF\-IDF, MedCPT, Hybrid\) remain minor throughout\.
Our ablation studies further support this overall pattern\. Increasing the number of retrieved documents beyondk=5k\{=\}5provides little additional benefit for open\-ended settings, and few\-shot prompting yields a modest gain at 3 shots for smaller models but degrades sharply at higher counts, with small models struggling under longer few\-shot contexts\. Larger models are comparatively stable across shot counts, but they also show limited benefit from retrieval augmentation\.
Taken together, these findings suggest that improving retrieval quality alone may not be sufficient to substantially improve biomedical QA performance in these settings\. One possible explanation is that current models, especially smaller ones, do not consistently make effective use of retrieved evidence, though our experiments do not directly measure evidence utilization or grounding\. This points to several directions for future work, including training or fine\-tuning methods that better support evidence integration, post\-retrieval re\-ranking or filtering to reduce context noise, and evaluation frameworks that more directly assess faithfulness and grounding rather than relying only on reference\-based metrics\. More broadly, an important open question is when retrieval is actually necessary, and whether we can better identify cases where the required knowledge is already contained within the model\.
## Acknowledgments
This material is based upon work supported by the National Science Foundation \(NSF\) under Grant No\. 2145357\.
## Limitations
Our study has several limitations\. First, we evaluate retrieval\-augmented generation using only reference\-based downstream metrics such as ROUGE\-L, BLEU, METEOR, BERTScore, and accuracy, rather than direct measures of faithfulness or evidence grounding\. For example, a model may produce a correct answer from its parametric knowledge without actually using the retrieved documents, or it may copy surface details from retrieval without truly improving medical correctness\. This limitation is not major for our study because our main goal is a comparative evaluation across models, retrievers, corpora, and no\-retrieval baselines within a single, consistent framework, and these metrics are sufficient to support our central finding that retrieval provides only small and inconsistent gains\.
Second, our experiments are limited to five open\-weight instruction\-tuned models and do not include proprietary frontier systems such as GPT\-4\-class medical assistants\. It is possible that stronger closed models use retrieved evidence more effectively, especially in cases requiring multi\-step reasoning over documents\. This limitation is not major for our study because our paper is specifically motivated by practical, deployable biomedical QA settings, where open\-weight 7B–72B models are realistic choices, and we also include both small and large open models to test whether the observed pattern holds across scales\.
Third, our retrieval setup uses a fixed top\-kkpipeline with four retrievers and four corpora, but does not explore more complex retrieval strategies such as adaptive retrieval, document re\-ranking, iterative retrieval, or task\-specific chunking\. For instance, some questions may require retrieving fewer but more precise passages, while others may benefit from multi\-hop retrieval or filtering noisy evidence before generation\. This limitation is not major for our study because we intentionally focus on strong, standard retrieval baselines widely used in prior biomedical RAG work, which makes the comparison clean and allows us to show that, even with several commonly used retrieval choices, the gains remain modest\.
Finally, our evaluation mixes expert and layman biomedical QA datasets, but the study does not separately analyze all possible sources of variation across question type, answer length, or knowledge intensity\. For example, retrieval may be more useful for highly specialized factoid questions than for common consumer\-health questions that models may already answer from pretraining alone\. This limitation is not major for our study because the breadth of datasets is a strength of the paper overall: the consistency of the pattern across ten datasets suggests that the weak benefit of retrieval is not tied to a single benchmark or user population\.
## References
- A\. B\. Abacha, Y\. Mrabet, M\. Sharp, T\. R\. Goodwin, S\. E\. Shooshan, and D\. Demner\-Fushman \(2019\)Bridging the gap between consumers’ medication questions and trusted answers\.Stud Health Technol Inform264,pp\. 25–29\.Note:1879\-8365 Abacha, Asma Ben Mrabet, Yassine Sharp, Mark Goodwin, Travis R Shooshan, Sonya E Demner\-Fushman, Dina Journal Article Netherlands 2019/08/24 Stud Health Technol Inform\. 2019 Aug 21;264:25\-29\. doi: 10\.3233/SHTI190176\.External Links:ISSN 0926\-9630,[Document](https://dx.doi.org/10.3233/shti190176)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.4.3.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p4.1)\.
- A\. Asai, Z\. Wu, Y\. Wang, A\. Sil, and H\. Hajishirzi \(2024\)Self\-RAG: learning to retrieve, generate, and critique through self\-reflection\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=hSyW5go0v8)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- A\. Ben Abacha and D\. Demner\-Fushman \(2019a\)A question\-entailment approach to question answering\.BMC Bioinformatics20\(1\),pp\. 511\.External Links:ISSN 1471\-2105,[Document](https://dx.doi.org/10.1186/s12859-019-3119-4),[Link](https://doi.org/10.1186/s12859-019-3119-4)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.8.7.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p5.1)\.
- A\. Ben Abacha and D\. Demner\-Fushman \(2019b\)On the summarization of consumer health questions\.InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics,A\. Korhonen, D\. Traum, and L\. Màrquez \(Eds\.\),Florence, Italy,pp\. 2228–2234\.External Links:[Link](https://aclanthology.org/P19-1215/),[Document](https://dx.doi.org/10.18653/v1/P19-1215)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.2.1.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p4.1)\.
- G\. V\. Cormack, C\. L\. A\. Clarke, and S\. Buettcher \(2009\)Reciprocal rank fusion outperforms condorcet and individual rank learning methods\.InProceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’09,New York, NY, USA,pp\. 758–759\.External Links:ISBN 9781605584836,[Link](https://doi.org/10.1145/1571941.1572114),[Document](https://dx.doi.org/10.1145/1571941.1572114)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p5.1),[§3](https://arxiv.org/html/2606.04127#S3.p15.4),[§3](https://arxiv.org/html/2606.04127#S3.p15.4.1)\.
- Y\. Gao, Y\. Xiong, X\. Gao, K\. Jia, J\. Pan, Y\. Bi, Y\. Dai, J\. Sun, H\. Wang, and H\. Wang \(2023\)Retrieval\-augmented generation for large language models: a survey\.2\(1\),pp\. 32\.External Links:[Link](https://arxiv.org/abs/2312.10997)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. Canton Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. Arrieta Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. Vasuden Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. Singh Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. Silveira Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, and T\. Speckbacher \(2024\)The Llama 3 Herd of Models\.pp\. arXiv:2407\.21783\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2407.21783),2407\.21783Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p4.1),[§3](https://arxiv.org/html/2606.04127#S3.p17.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.11.10.2.1.1),[§1](https://arxiv.org/html/2606.04127#S1.p1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p5.1)\.
- Z\. Ji, N\. Lee, R\. Frieske, T\. Yu, D\. Su, Y\. Xu, E\. Ishii, Y\. J\. Bang, A\. Madotto, and P\. Fung \(2023\)Survey of hallucination in natural language generation\.55\(12\)\.External Links:ISSN 0360\-0300,[Link](https://doi.org/10.1145/3571730),[Document](https://dx.doi.org/10.1145/3571730)Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p1.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier,et al\.\(2023\)Mistral 7b\.External Links:[Link](https://arxiv.org/abs/2310.06825)Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p4.1),[§3](https://arxiv.org/html/2606.04127#S3.p17.1)\.
- D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. Szolovits \(2021\)What disease does this patient have? a large\-scale open domain question answering dataset from medical exams\.Applied SciencesCureusSci DataFound\. Trends Inf\. Retr\.Journal of DocumentationBioinformaticsarXiv preprint arXiv:2505\.09388arXiv e\-printsarXiv preprint arXiv:2310\.06825NatureACM Comput\. Surv\.arXiv preprint arXiv:2312\.10997arXiv preprint arXiv:2303\.1337511\(14\)\.External Links:[Link](https://www.mdpi.com/2076-3417/11/14/6421),ISSN 2076\-3417Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.9.8.2.1.1),[§1](https://arxiv.org/html/2606.04127#S1.p1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p5.1)\.
- Q\. Jin, B\. Dhingra, Z\. Liu, W\. Cohen, and X\. Lu \(2019\)PubMedQA: a dataset for biomedical research question answering\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),K\. Inui, J\. Jiang, V\. Ng, and X\. Wan \(Eds\.\),Hong Kong, China,pp\. 2567–2577\.External Links:[Link](https://aclanthology.org/D19-1259/),[Document](https://dx.doi.org/10.18653/v1/D19-1259)Cited by:[§5](https://arxiv.org/html/2606.04127#S5.p4.1)\.
- Q\. Jin, W\. Kim, Q\. Chen, D\. C\. Comeau, L\. Yeganova, W\. J\. Wilbur, and Z\. Lu \(2023\)MedCPT: contrastive pre\-trained transformers with large\-scale pubmed search logs for zero\-shot biomedical information retrieval\.39\(11\),pp\. btad651\.External Links:ISSN 1367\-4811,[Document](https://dx.doi.org/10.1093/bioinformatics/btad651),[Link](https://doi.org/10.1093/bioinformatics/btad651),https://academic\.oup\.com/bioinformatics/article\-pdf/39/11/btad651/52799559/btad651\.pdfCited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1),[§2](https://arxiv.org/html/2606.04127#S2.p5.1),[§3](https://arxiv.org/html/2606.04127#S3.p14.1.1)\.
- V\. Karpukhin, B\. Oguz, S\. Min, P\. Lewis, L\. Wu, S\. Edunov, D\. Chen, and W\. Yih \(2020\)Dense passage retrieval for open\-domain question answering\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),B\. Webber, T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 6769–6781\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.550/),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.550)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- A\. Krithara, A\. Nentidis, K\. Bougiatiotis, and G\. Paliouras \(2023\)BioASQ\-qa: a manually curated corpus for biomedical question answering\.10\(1\),pp\. 170\.Note:2052\-4463 Krithara, Anastasia Orcid: 0000\-0003\-0491\-4507 Nentidis, Anastasios Bougiatiotis, Konstantinos Paliouras, Georgios Dataset Journal Article England 2023/03/28 Sci Data\. 2023 Mar 27;10\(1\):170\. doi: 10\.1038/s41597\-023\-02068\-4\.External Links:ISSN 2052\-4463,[Document](https://dx.doi.org/10.1038/s41597-023-02068-4)Cited by:[Table 8](https://arxiv.org/html/2606.04127#A1.T8.1.3.1.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p7.1.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. Kiela \(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InAdvances in Neural Information Processing Systems,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 9459–9474\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/6b493230205f780e1bc26945df7481e5-Paper.pdf)Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p2.1),[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- Y\. Li, Z\. Li, K\. Zhang, R\. Dan, S\. Jiang, and Y\. Zhang \(2023\)Chatdoctor: a medical chat model fine\-tuned on a large language model meta\-ai \(llama\) using medical domain knowledge\.15\(6\)\.External Links:[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC10364849/)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.6.5.2.1.1),[Table 8](https://arxiv.org/html/2606.04127#A1.T8.1.5.3.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p10.1.1),[§3](https://arxiv.org/html/2606.04127#S3.p4.1)\.
- X\. Ma, Y\. Gong, P\. He, H\. Zhao, and N\. Duan \(2023\)Query rewriting in retrieval\-augmented large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 5303–5315\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.322/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.322)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- A\. Nentidis, G\. Katsimpras, A\. Krithara, M\. Krallinger, M\. Rodríguez\-Ortega, E\. Rodriguez\-López, N\. Loukachevitch, A\. Sakhovskiy, E\. Tutubalina, D\. Dimitriadis, G\. Tsoumakas, G\. Giannakoulas, A\. Bekiaridou, A\. Samaras, G\. M\. D\. Nunzio, N\. Ferro, S\. Marchesin, M\. Martinelli, G\. Silvello, and G\. Paliouras \(2025\)Overview of bioasq 2025: the thirteenth bioasq challenge on large\-scale biomedical semantic indexing and question answering\.InExperimental IR Meets Multilinguality, Multimodality, and Interaction,Cham\.External Links:ISBN 978\-3\-032\-04354\-2,[Document](https://dx.doi.org/10.1007/978-3-032-04354-2%5F12),[Link](https://link.springer.com/10.1007/978-3-032-04354-2_12)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.7.6.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p5.1)\.
- V\. Nguyen, S\. Karimi, M\. Rybinski, and Z\. Xing \(2023\)MedRedQA for medical consumer question answering: dataset, tasks, and neural baselines\.InProceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics \(Volume 1: Long Papers\),J\. C\. Park, Y\. Arase, B\. Hu, W\. Lu, D\. Wijaya, A\. Purwarianti, and A\. A\. Krisnadhi \(Eds\.\),Nusa Dua, Bali,pp\. 629–648\.External Links:[Link](https://aclanthology.org/2023.ijcnlp-main.42/),[Document](https://dx.doi.org/10.18653/v1/2023.ijcnlp-main.42)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.3.2.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p4.1)\.
- H\. Nori, N\. King, S\. M\. McKinney, D\. Carignan, and E\. Horvitz \(2023\)Capabilities of gpt\-4 on medical challenge problems\.External Links:[Link](https://arxiv.org/abs/2303.13375)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p2.1)\.
- O\. Ovadia, M\. Brief, M\. Mishaeli, and O\. Elisha \(2024\)Fine\-tuning or retrieval? comparing knowledge injection in LLMs\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 237–250\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.15/),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.15)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p2.1)\.
- A\. Pal, L\. K\. Umapathi, and M\. Sankarasubbu \(2022\)MedMCQA: a large\-scale multi\-subject multi\-choice dataset for medical domain question answering\.InProceedings of the Conference on Health, Inference, and Learning,G\. Flores, G\. H\. Chen, T\. Pollard, J\. C\. Ho, and T\. Naumann \(Eds\.\),Proceedings of Machine Learning Research, Vol\.174,pp\. 248–260\.External Links:[Link](https://proceedings.mlr.press/v174/pal22a.html)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.10.9.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p5.1)\.
- S\. Robertson and H\. Zaragoza \(2009\)The probabilistic relevance framework: bm25 and beyond\.3\(4\),pp\. 333–389\.External Links:ISSN 1554\-0669,[Link](https://doi.org/10.1561/1500000019),[Document](https://dx.doi.org/10.1561/1500000019)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p4.1),[§3](https://arxiv.org/html/2606.04127#S3.p12.2.1)\.
- Z\. Shao, Y\. Gong, Y\. Shen, M\. Huang, N\. Duan, and W\. Chen \(2023\)Enhancing retrieval\-augmented large language models with iterative retrieval\-generation synergy\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 9248–9274\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.620/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.620)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- F\. Shi, X\. Chen, K\. Misra, N\. Scales, D\. Dohan, E\. H\. Chi, N\. Schärli, and D\. Zhou \(2023\)Large language models can be easily distracted by irrelevant context\.InProceedings of the 40th International Conference on Machine Learning,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 31210–31227\.External Links:[Link](https://proceedings.mlr.press/v202/shi23a.html)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p2.1)\.
- \[27\]K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl, P\. Payne, M\. Seneviratne, P\. Gamble, C\. Kelly, A\. Babiker, N\. Schärli, A\. Chowdhery, P\. Mansfield, D\. Demner\-Fushman, B\. Agüera y Arcas, D\. Webster, G\. S\. Corrado, Y\. Matias, K\. Chou, J\. Gottweis, N\. Tomasev, Y\. Liu, A\. Rajkomar, J\. Barral, C\. Semturs, A\. Karthikesalingam, and V\. NatarajanLarge language models encode clinical knowledge\.620\(7972\),pp\. 172–180\.Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p1.1)\.
- K\. SPARCK JONES \(1972\)A statistical interpretation of term specificity and its application in retrieval\.28\(1\),pp\. 11–21\.External Links:ISSN 0022\-0418,[Document](https://dx.doi.org/10.1108/eb026526),[Link](https://doi.org/10.1108/eb026526)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p4.1),[§3](https://arxiv.org/html/2606.04127#S3.p13.1.1)\.
- X\. Tang, A\. Zou, Z\. Zhang, Z\. Li, Y\. Zhao, X\. Zhang, A\. Cohan, and M\. Gerstein \(2024\)MedAgents: large language models as collaborators for zero\-shot medical reasoning\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 599–621\.External Links:[Link](https://aclanthology.org/2024.findings-acl.33/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.33)Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p2.1),[§2](https://arxiv.org/html/2606.04127#S2.p2.1),[§3](https://arxiv.org/html/2606.04127#S3.p5.1)\.
- H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. Sabharwal \(2023\)Interleaving retrieval with chain\-of\-thought reasoning for knowledge\-intensive multi\-step questions\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),A\. Rogers, J\. Boyd\-Graber, and N\. Okazaki \(Eds\.\),Toronto, Canada,pp\. 10014–10037\.External Links:[Link](https://aclanthology.org/2023.acl-long.557/),[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.557)Cited by:[§2](https://arxiv.org/html/2606.04127#S2.p1.1)\.
- G\. Xiong, Q\. Jin, Z\. Lu, and A\. Zhang \(2024\)Benchmarking retrieval\-augmented generation for medicine\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 6233–6251\.External Links:[Link](https://aclanthology.org/2024.findings-acl.372/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.372)Cited by:[Table 8](https://arxiv.org/html/2606.04127#A1.T8.1.1.3.1.1),[§1](https://arxiv.org/html/2606.04127#S1.p2.1),[§1](https://arxiv.org/html/2606.04127#S1.p3.1),[§2](https://arxiv.org/html/2606.04127#S2.p1.1),[§2](https://arxiv.org/html/2606.04127#S2.p2.1),[§2](https://arxiv.org/html/2606.04127#S2.p4.1),[§3](https://arxiv.org/html/2606.04127#S3.p8.1.1)\.
- Yahoo\! Research \(2009\)Yahoo\! Webscope Datasets Catalog\.Technical reportYahoo\! Inc\.\.Note:19 Datasets Available\. Accessed via Stanford InfoLabExternal Links:[Link](http://infolab.stanford.edu/%CB%9Cullman/mining/2009/YahooData.pdf)Cited by:[Table 8](https://arxiv.org/html/2606.04127#A1.T8.1.4.2.2.1.1),[§3](https://arxiv.org/html/2606.04127#S3.p9.1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.External Links:[Link](https://arxiv.org/abs/2505.09388)Cited by:[§1](https://arxiv.org/html/2606.04127#S1.p4.1),[§3](https://arxiv.org/html/2606.04127#S3.p17.1)\.
- M\. Zhang, S\. Dou, Z\. Wang, and Y\. Wu \(2022\)Focus\-driven contrastive learning for medical question summarization\.InProceedings of the 29th International Conference on Computational Linguistics,N\. Calzolari, C\. Huang, H\. Kim, J\. Pustejovsky, L\. Wanner, K\. Choi, P\. Ryu, H\. Chen, L\. Donatelli, H\. Ji, S\. Kurohashi, P\. Paggio, N\. Xue, S\. Kim, Y\. Hahm, Z\. He, T\. K\. Lee, E\. Santus, F\. Bond, and S\. Na \(Eds\.\),Gyeongju, Republic of Korea,pp\. 6176–6186\.External Links:[Link](https://aclanthology.org/2022.coling-1.539/)Cited by:[§3](https://arxiv.org/html/2606.04127#S3.p4.1)\.
- M\. Zhu, A\. Ahuja, D\. Juan, W\. Wei, and C\. K\. Reddy \(2020\)Question answering with long multiple\-span answers\.InFindings of the Association for Computational Linguistics: EMNLP 2020,T\. Cohn, Y\. He, and Y\. Liu \(Eds\.\),Online,pp\. 3840–3849\.External Links:[Link](https://aclanthology.org/2020.findings-emnlp.342/),[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.342)Cited by:[Table 7](https://arxiv.org/html/2606.04127#A1.T7.1.5.4.2.1.1),[§2](https://arxiv.org/html/2606.04127#S2.p3.1),[§3](https://arxiv.org/html/2606.04127#S3.p4.1)\.
## Appendix ADatasets
Our experiments use ten biomedical and consumer\-health query datasets and four retrieval corpora\. Table[7](https://arxiv.org/html/2606.04127#A1.T7)summarises each evaluation dataset, its source size, and the number of examples selected for evaluation; Table[8](https://arxiv.org/html/2606.04127#A1.T8)describes the knowledge bases indexed for retrieval\. Full details on dataset splits and the few\-shot query pools are provided in Section[3](https://arxiv.org/html/2606.04127#S3)\.
Table 7:Evaluation query datasets grouped by user type\.Layman: consumer health queries in everyday language;Expert: biomedical/clinical questions\. “Query Set”: few\-shot pool size; “Test Set”: evaluation examples used\.Datasets are grouped by intended user type:laymandatasets reflect consumer health inquiries in everyday language, whileexpertdatasets target biomedical professionals or medical students\. TheQuery Setcolumn reports the eligible pool size used for few\-shot prompting; theTest Setcolumn reports the number of examples used as evaluation queries\. All corpora in Table[8](https://arxiv.org/html/2606.04127#A1.T8)are indexed as whole records without further chunking; for Q&A\-format corpora \(Yahoo Answers and HealthCareMagic\) each document concatenates the question or title with the answer body\.
Table 8:Retrieval corpora grouped by user type\.Expert: technical biomedical sources;Layman: community health and general Q&A\. “Documents Retained”: records after quality filtering\.
## Appendix BPrompt Templates
All generation experiments use two system prompts and two user message templates, combined via each model’s native chat template\.
### Layman system prompt\.
Used when the query dataset belongs to thelaymanuser type \(MeQSum, MedRedQA, MedicationQA, MASH\-QA, ChatDoctor\-iCliniq\):
Layman System PromptYou are a helpful health assistant answering questions from members of the general public\. Use simple, everyday language that a non\-medical person can easily understand\. Avoid medical jargon\. Be clear, friendly, and concise\.
### Expert system prompt\.
Used forexpertdatasets \(BioASQ Task B, MedQuAD, MedQA\-USMLE, MedMCQA, MMLU Medical\):
Expert System PromptYou are a clinical decision support assistant\. Answer questions from healthcare professionals using precise medical terminology\. Provide evidence\-based, clinically detailed responses with relevant diagnostic and therapeutic considerations\.
User message \(without retrieval\)
User MessageQUESTION: \{query\} ANSWER:
User message \(with retrieval\)
User MessageUse the following retrieved passages to help answer the question\. RETRIEVED CONTEXT: \{context\} QUESTION: \{query\} ANSWER:
where\{context\}is a concatenation of the top\-kkretrieved passages, each formatted as\[Passage N \(source: \{source\}\)\]: \{text\}\. The final prompt is produced by wrapping these system and user messages in each model’s chat template\.
## Appendix CFull Results by Metric
This section provides per\-dataset performance tables for all evaluated metrics beyond ROUGE\-L \(reported in the main paper\)\. Tables[9](https://arxiv.org/html/2606.04127#A3.T9)–[13](https://arxiv.org/html/2606.04127#A3.T13)report ROUGE\-2, ROUGE\-1, BERTScore, METEOR, and BLEU respectively, broken down by model and retrieval corpus across the seven open\-ended datasets\. Tables[14](https://arxiv.org/html/2606.04127#A3.T14)–[18](https://arxiv.org/html/2606.04127#A3.T18)further break down ROUGE\-1, ROUGE\-2, BLEU, METEOR, and BERTScore by retrieval method \(BM25, Hybrid, MedCPT, TF\-IDF\) across all four corpora\.
### ROUGE\-2 \(Table[9](https://arxiv.org/html/2606.04127#A3.T9)\)\.
The pattern mirrors ROUGE\-L: the BioASQ/PubMed corpus produces the largest gains, and only on the BioASQ expert open\-ended task\. For example, LLaMA\-3\.1\-8B improves from 14\.03 to 18\.34, and LLaMA\-3\.1\-70B from 14\.43 to 19\.25, while all other datasets see gains under 1 point or negative effects\. The average improvement over baseline is at most 0\.53 points \(LLaMA\-3\.1\-8B: 5\.82→\\to6\.35\), and for Qwen2\.5\-72B Yahoo Answers produces the best average \(6\.39\), marginally ahead of BioASQ \(6\.37\), illustrating how small these corpus\-level differences are\.
### ROUGE\-1 \(Table[10](https://arxiv.org/html/2606.04127#A3.T10)\)\.
Again, the BioASQ corpus helps on the BioASQ dataset \(gains of 3–5 points for all models\) while effects on lay datasets are within±\\pm1 point\. Mistral\-7B shows an above\-average improvement with BioASQ corpus on the BioASQ open\-ended task \(37\.56→\\to40\.46 with BioASQ corpus; 40\.53 with Qwen2\.5\-72B\), confirming domain\-matched retrieval has local benefit\. Averaged over all datasets the maximum gain is 0\.65 points \(LLaMA\-3\.1\-8B baseline 22\.49→\\tobest 22\.88\)\.
### BERTScore \(Table[11](https://arxiv.org/html/2606.04127#A3.T11)\)\.
BERTScore is notably more stable than any ROUGE metric: the gap between the no\-retrieval baseline and the best retrieval condition is under 0\.7 points for all models\. For instance, LLaMA\-3\.1\-8B moves from 52\.47 \(baseline\) to at best 52\.85 \(BioASQ corpus\), a gain of just 0\.38 points\. This suggests that while retrieved context can slightly shift surface n\-gram overlap, the overall semantic content of model outputs barely changes, consistent with the view that 7–8B models are not effectively incorporating the retrieved evidence\.
### METEOR \(Table[12](https://arxiv.org/html/2606.04127#A3.T12)\)\.
METEOR shows small, mixed effects: the BioASQ corpus provides a modest boost on the BioASQ dataset \(e\.g\., LLaMA\-3\.1\-8B: 29\.84→\\to30\.92; Qwen2\.5\-72B: 31\.11→\\to34\.05\), but on lay datasets retrieval often slightly lowers METEOR, particularly for the 70B models where the baseline exceeds all retrieval conditions on several tasks \(e\.g\., LLaMA\-3\.1\-70B average: 18\.92 baseline vs\. 17\.03–17\.12 across all corpora\)\.
### BLEU \(Table[13](https://arxiv.org/html/2606.04127#A3.T13)\)\.
BLEU scores are generally very low for lay datasets \(<2<2across all conditions\), underscoring that n\-gram precision is a weak signal for open\-ended health QA\. The BioASQ corpus produces notable gains on the expert BioASQ dataset \(LLaMA\-3\.1\-8B: 12\.92→\\to19\.08; LLaMA\-3\.1\-70B: 13\.32→\\to19\.76\), but all other datasets improve by less than 0\.3 BLEU points, and many worsen\. Average BLEU across all datasets improves by 0\.82 points at most\.
### Retrieval method breakdown \(Tables[14](https://arxiv.org/html/2606.04127#A3.T14)–[18](https://arxiv.org/html/2606.04127#A3.T18)\)\.
Across all five metrics, differences among BM25, Hybrid \(RRF\), MedCPT, and TF\-IDF are consistently within 0\.5 metric points for any model–corpus combination\. The Hybrid retriever shows a slight edge in several configurations \(particularly ROUGE\-L and ROUGE\-1\), while TF\-IDF is competitive with BM25 despite its greater simplicity\. No single retrieval method dominates across all metrics and models, reinforcing the conclusion that retrieval architecture choice is secondary to corpus and model selection\.
Table 9:ROUGE\-2 by model and retrieval corpus \(open\-ended datasets\)\.Table 10:ROUGE\-1 by model and retrieval corpus \(open\-ended datasets\)\.Table 11:BERTScore by model and retrieval corpus \(open\-ended datasets\)\.Table 12:METEOR by model and retrieval corpus \(open\-ended datasets\)\.Table 13:BLEU by model and retrieval corpus \(open\-ended datasets\)\.Table 14:ROUGE\-1 by retrieval method\.Table 15:ROUGE\-2 by retrieval method\.Table 16:BLEU by retrieval method\.Table 17:METEOR by retrieval method\.Table 18:BERTScore by retrieval method\.
## Appendix DAblation Study: Additional Figures
This section provides additional figures for the two ablation studies described in Section[5](https://arxiv.org/html/2606.04127#S5)\. Figure[7](https://arxiv.org/html/2606.04127#A4.F7)shows BERTScore, METEOR, BLEU, ROUGE\-2, and ROUGE\-1 trends under few\-shot prompting across all five models on open\-ended questions\. Figure[8](https://arxiv.org/html/2606.04127#A4.F8)shows the same metrics across top\-kkvalues\.
### Few\-shot: additional metrics \(Figure[7](https://arxiv.org/html/2606.04127#A4.F7)\)\.
All six open\-ended metrics tell a consistent story\. For the larger models \(LLaMA\-3\.1\-70B and Qwen2\.5\-72B\), all metrics are flat across all shot counts: for example, METEOR stays at 17\.44–17\.45 for LLaMA\-3\.1\-70B and BERTScore stays at 52\.69 for Qwen2\.5\-72B regardless of shot count\. For smaller models, the 3\-shot sweet spot and subsequent collapse are visible in every metric\. Specifically, ROUGE\-1 peaks at 3 shots for LLaMA\-3\.1\-8B \(26\.29\) and Qwen2\.5\-7B \(26\.61\) before collapsing to 12\.19 and 13\.52 at 10 shots\. METEOR follows the same pattern: LLaMA\-3\.1\-8B peaks at 19\.99 \(3 shots\) vs\. 9\.12 \(10 shots\), and Qwen2\.5\-7B at 20\.54 \(3 shots\) vs\. 10\.43 \(10 shots\)\. BERTScore shows a more severe drop for Mistral\-7B: from 54\.32 at 1 shot to 40\.96 at 10 shots, a 13\-point collapse\. BLEU, while numerically small, also collapses dramatically, LLaMA\-3\.1\-8B drops from 5\.15 \(3 shots\) to 1\.15 \(10 shots\), and Mistral\-7B from 4\.66 \(3 shots\) to 0\.88 \(10 shots\), confirming that higher shot counts severely degrade output quality at this scale\.
### Top\-kk: additional metrics \(Figure[8](https://arxiv.org/html/2606.04127#A4.F8)\)\.
The plateau behavior seen in ROUGE\-L \(main paper\) extends to all six additional metrics\. ROUGE\-1 changes by at most 0\.13 points fromk=5k\{=\}5tok=50k\{=\}50across all models\. ROUGE\-2 is similarly stable: for example, LLaMA\-3\.1\-70B moves from 7\.44 \(k=5k\{=\}5\) to 7\.45 \(k=50k\{=\}50\), a negligible change\. METEOR plateaus byk=5k\{=\}5for most models, with variations under 0\.1 betweenk=5k\{=\}5andk=50k\{=\}50\. BERTScore is the most stable metric of all: for LLaMA\-3\.1\-8B, it ranges only from 51\.33 \(k=1k\{=\}1\) to 51\.53 \(k=10k\{=\}10\), a 0\.20\-point spread across all sixkkvalues\. The transition fromk=1k\{=\}1tok=5k\{=\}5accounts for nearly all the variation, and additional passages beyondk=5k\{=\}5provide no measurable benefit in any metric for any model\.
Figure 7:Additional open\-ended metrics across shot counts \(few\-shot ablation\)\.Figure 8:Additional open\-ended metrics across top\-kkvalues\.Similar Articles
Where Does Retrieval Fail? Evaluating RAG Architectures for Agricultural Advisory
This paper evaluates retrieval quality in RAG systems for Bengali agricultural advisory, finding performance varies by query type and language conditions, and introduces a benchmark dataset to highlight the need for disaggregated evaluation in low-resource settings.
RAG-Safety-Bench: Reliable Evaluation of Retrieval-Augmented LLM Safety
This paper introduces RAG-Safety-Bench, a benchmark for evaluating the safety of retrieval-augmented large language models, demonstrating that RAG can lead to safety degradation even with benign retrieved documents.
Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms
This paper presents a controlled scaling study comparing lexical, dense, graph-based, and agentic RAG paradigms across corpus sizes from 1,000 to 512,000 documents, finding that BM25 provides the best accuracy-cost tradeoff, while graph-based RAG faces high construction costs that limit scalability.
"Most RAG benchmarks lie about real-world corpora." Test data from 3 production websites.
This article argues that most RAG benchmarks are misleading because they assume uniform corpus quality, while real-world corpora vary significantly in content density. Using data from three production websites, it shows that a tiered approach and a 'yield score' can better predict retrieval effectiveness.
When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering
Introduces OGCaReBench, a free-form retrieval benchmark for evaluating LLMs on clinical questions that require reasoning beyond standard guidelines. Experiments show that even the best model achieves only 56% accuracy, but retrieval augmentation boosts performance to 82%.