REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement

arXiv cs.AI Papers

Summary

REPAIR is a self-evolving data augmentation framework for scientific dense retrievers that resolves long-tail confusion via fact-verified iterative refinement, demonstrating significant performance improvements on materials science and biomedical benchmarks.

arXiv:2609.18262v1 Announce Type: new Abstract: Precise retrieval of scientific information is fundamentally constrained by long-tailed concepts and high fact-sensitivity of scientific corpora. These challenges often limit the effectiveness of dense retrievers and hallucination-prone LLM augmentation. To address this, we present REPAIR, a self-evolving data augmentation framework for scientific dense retrievers. REPAIR iteratively synthesizes training data to address knowledge gaps by cycling through diagnosis of long-tail concepts, API-guided evidence expansion, and differentiation via hard negative mining. This process effectively grounds retrieval in factual reality to resolve fine-grained distinctions. Extensive experiments demonstrate that REPAIR significantly outperforms 19 strong baselines on nine materials science and biomedical benchmarks. Our work highlights that diagnosing and factually augmenting data to long-tail deficits is essential for robust scientific retrieval.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:35 AM

# Resolving Long-Tail Confusion in Scientific Retrieversvia Fact-Verified Iterative Refinement
Source: [https://arxiv.org/html/2609.18262](https://arxiv.org/html/2609.18262)
###### Abstract

Precise retrieval of scientific information is fundamentally constrained bylong\-tailed conceptsandhigh fact\-sensitivityof scientific corpora\. These challenges often limit the effectiveness of dense retrievers and hallucination\-prone LLM augmentation\. To address this, we presentREPAIR, a self\-evolving data augmentation framework for scientific dense retrievers\.REPAIRiteratively synthesizes training data to address knowledge gaps by cycling through diagnosis of long\-tail concepts, API\-guided evidence expansion, and differentiation via hard negative mining\. This process effectively grounds retrieval in factual reality to resolve fine\-grained distinctions\. Extensive experiments demonstrate thatREPAIRsignificantly outperforms 19 strong baselines on nine materials science and biomedical benchmarks\. Our work highlights that diagnosing and factually augmenting data to long\-tail deficits is essential for robust scientific retrieval\.111Our code is available at[https://github\.com/yerimoh/REPAIR](https://github.com/yerimoh/REPAIR)

## 1Introduction

In highly specialized fields such as materials science and biomedicine, the continuous influx of new literature makes efficient knowledge discovery a critical challenge[Sharma et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib48);[Choudhary et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib15);[Kononova et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib16)\. To address this, retrieval\-augmented generation \(RAG\) has emerged as a promising methodology to dynamically incorporate up\-to\-date domain knowledge\. By grounding generation in precise information retrieval \(IR\), RAG enables reliable downstream applications, including question answering[Sohn et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib50);[Zhang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib49), knowledge discovery[Ocana et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib10);[Pei et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib9), and scientific decision\-making[Chiang et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib51);[Ong et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib11)\. The success of these applications fundamentally depends on the accuracy of the underlying IR models\.

However, while state\-of\-the\-art dense retrievers excel on general\-domain text, they suffer significant performance degradation when applied to scientific corpora\. This lexical and semantic gap is widely recognized as the domain shift issue[Kamalloo et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib54)\. To mitigate this, recent studies have adapted retrievers by aggregating domain\-specific datasets or utilizing synthetic data augmentation[Jin et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib12);[Zhang et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib52);[Singh et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib70);[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)\. Despite yielding empirical improvements, these approaches largely treat scientific documents as standard text, overlooking the intrinsic characteristics that distinguish scientific literature from general domains\.

\(a\) Comparison of augmented pairs

Domain\-level Long\-tail

Concept\-level Long\-tail

\(b\) Long\-tail distributions

Figure 1:Motivation of REPAIR\. \(a\) While existing data augmentation methods generate structuralhallucinations\(e\.g\., lexically similarBCL1andBCL2critically corrupt scientific facts\), REPAIR accurately groundscondition\-sensitive scientific facts\. \(b\) Log\-frequency analysis on scientific vs\. general corpora reveals that the severe long\-tail in scientific domains \(left\) is predominantly driven byscientific concepts\(right\)\.Specifically, scientific retrieval is governed by two structural properties:long\-tailed concept distributionandhigh fact\-sensitivity\. Scientific corpora exhibit extreme long\-tail distributions composed of irreplaceable entities such as chemical formulas, rare molecular structures, and specific gene or protein families[Oh et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib55)\. Unlike general\-domain terms, these entities lack semantic substitutes, making naive data augmentation ineffective\. As shown in Figure[1](https://arxiv.org/html/2609.18262#S1.F1), specific proteins likeBCL2represent such entities that cannot be loosely generalized\. A multi\-corpus statistical characterization of this long tail is given in Appendix[A\.3](https://arxiv.org/html/2609.18262#A1.SS3)\. Moreover, scientific outcomes are hypersensitive to precise terminology and experimental conditions\. A single\-character hallucination fromBCL2toBCL1invalidates the generated query\. Such plausible but incorrect LLM\-generated data are harmful as it trains retrievers with scientific falsehoods[Pal et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib75)\.

To address these limitations, we proposeREPAIR\(Retriever viaEpistemicAPI\-GuidedIterativeRefinement\), an iterative, epistemic self\-evolving procedure of data augmentation for training of an LLM\-based scientific retriever\. REPAIR first initializes a seed retriever with compact scientific corpora, then iteratively refines it through data synthesis of three stages: \(1\)Diagnosisidentifies long\-tail concepts that the retriever tends to confuse, \(2\)Expansiongrounds these concepts in fact\-verified documents mined from scientific APIs, and \(3\)Differentiationresolves fine\-grained factual distinctions using fact\-contrastive hard negatives to trains the model\. Through iterative refinement, REPAIR progressively corrects the retriever’s long\-tail confusions using externally verified evidence\.

We conduct comprehensive experiments across nine diverse scientific retrieval benchmarks in both materials science and biomedical domains\. The results show thatREPAIRenhances retrieval precision on specialized scientific tasks while exhibiting robust generalization capabilities\. By iteratively correcting long\-tail confusion with externally verified evidence,REPAIRoutperforms 19 strong baselines, demonstrating the effectiveness of our iterative self\-evolving strategy\. In summary, this work presents the following contributions:

- •We proposeREPAIR, a self\-evolving data augmentation framework that curtails structural hallucinations of scientific retrievers by resolving long\-tailed concept confusion and high fact\-sensitivity, which have largely been overlooked in prior work\.
- •Scaling retriever parameters from 500M to 7B,REPAIRachieves new state\-of\-the\-art results across nine materials science and biomedical benchmarks, outperforming 19 strong baselines while using less training data\.
- •Our extensive experiments show that the three\-stage data augmentation pipeline of diagnosis, expansion, and differentiation outperforms naive synthetic data scaling in correcting long\-tail retrieval errors\.

## 2Related Work

A broad range of studies has investigated representation learning for text retrieval, progressing from latent semantic models to neural embedding\-based approaches[Blei et al\. \(2003\)](https://arxiv.org/html/2609.18262#bib.bib19);[Hofmann \(1999\)](https://arxiv.org/html/2609.18262#bib.bib56);[Deerwester et al\. \(1990\)](https://arxiv.org/html/2609.18262#bib.bib18)\. In recent years, dense retrieval with transformer encoders has become the dominant paradigm\. Further gains have been achieved by scaling retrievers or applying instruction tuning with LLMs[Izacard et al\. \(2021a\)](https://arxiv.org/html/2609.18262#bib.bib21);[Yu et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib24);[Chen et al\. \(2024a\)](https://arxiv.org/html/2609.18262#bib.bib20);[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib58);[Ni et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib57);[Neelakantan et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib13)\. However, such improvements are largely attained in general\-domain settings characterized by abundant supervised data\.

As LLMs are increasingly applied to scientific reasoning, accurate retrieval of domain\-specific knowledge has become critical for reliability[Zhang et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib28);[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib27);[Pilania \(2021\)](https://arxiv.org/html/2609.18262#bib.bib2);[Olivetti et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib3)\. Although retrievers have been adapted to scientific domains via domain\-specific pretraining and task\-oriented training[Jin et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib12);[Zhang et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib52);[Singh et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib70), scientific retrieval remains fundamentally challenged by distributional shifts[Kamalloo et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib54), particularly due to severe data scarcity and long\-tailed entity distributions\.

To address data scarcity, recent studies have adopted LLM\-based data augmentation[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)\. However, such generative methods risk hallucinations, undermining the factual reliability essential for science[Pal et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib75)\. Furthermore, while prior work on long\-tailed distributions has focused on model\-centric adaptations, such as specialized tokenization[Oh et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib55)or domain\-adaptive learning[Kim et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib1), we argue that the fundamental bottleneck lies in the data themselves\. Unlike previous model\-centric strategies that attempt to adapt parameters to noisy or scarce distributions, our approach directly targets the quality and factual grounding of the retrieval data\. We introduce a data\-centric framework designed to mitigate the risks of hallucination and effectively cover long\-tailed scientific concepts\.

![Refer to caption](https://arxiv.org/html/2609.18262v1/main+fin.png)Figure 2:Overview of theREPAIRframework\. The model is initialized with scientific seed corpora and iteratively refined through self\-diagnosed factual expansion, including diagnosis, expansion, and differentiation\.
## 3Methodology: REPAIR

We focus on improving the reliability of dense retrievers for scientific domains \(§[3\.1](https://arxiv.org/html/2609.18262#S3.SS1)\) by explicitly addressing long\-tail concept confusion and high factual sensitivity\. Starting from a retriever initialized with compact scientific seed corpora \(§[3\.2](https://arxiv.org/html/2609.18262#S3.SS2)\), we refine it through three stages:*Diagnosis*of long\-tail concept uncertainty \(§[3\.3](https://arxiv.org/html/2609.18262#S3.SS3)\),*Expansion*with externally verified scientific evidence \(§[3\.4](https://arxiv.org/html/2609.18262#S3.SS4)\), and*Differentiation*via fact\-contrastive hard negatives \(§[3\.5](https://arxiv.org/html/2609.18262#S3.SS5)\)\. This refinement is iteratively optimized using contrastive learning \(§[3\.6](https://arxiv.org/html/2609.18262#S3.SS6)\)\. The overall procedure is illustrated in Figure[2](https://arxiv.org/html/2609.18262#S2.F2), and implementation details are provided in Appendix[A](https://arxiv.org/html/2609.18262#A1)\.

### 3\.1Retriever Formulation

Let𝒬\\mathcal\{Q\}be a set of queries and𝒟\\mathcal\{D\}a document corpus\. The dense retriever represents queries and documents as dense embeddings using a shared decoder\-only language modelℳθ\\mathcal\{M\}\_\{\\theta\}\(e\.g\., Qwen\-2\.5[Yang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib37)\) with parameter scales of 500M, 1\.5B, and 7B\. For a query–document pair\(q,d\)\(q,d\), we compute the embeddings as𝐞q∈ℝD\\mathbf\{e\}\_\{q\}\\in\\mathbb\{R\}^\{D\}and𝐞d∈ℝD\\mathbf\{e\}\_\{d\}\\in\\mathbb\{R\}^\{D\}; we append an end\-of\-sequence token to them and compute their dense representations by EOS pooling over the final\-layer hidden states:

𝐞q/d=PoolEOS​\(ℳθ​\(q/d⊕\[EOS\]\)\)\.\\mathbf\{e\}\_\{q/d\}=\\mathrm\{Pool\}\_\{\\texttt\{EOS\}\}\\left\(\\mathcal\{M\}\_\{\\theta\}\(q/d\\oplus\\texttt\{\[EOS\]\}\)\\right\)\.\(1\)We score relevance by the dot product:sθ​\(q,d\)=𝐞q⊤​𝐞ds\_\{\\theta\}\(q,d\)=\\mathbf\{e\}\_\{q\}^\{\\top\}\\mathbf\{e\}\_\{d\}\. For each queryqq, the retriever returns a ranked listTopKθ​\(q\)⊂𝒟\\mathrm\{TopK\}\_\{\\theta\}\(q\)\\subset\\mathcal\{D\}undersθs\_\{\\theta\}\.

### 3\.2Initialization from Scientific Seed Corpora

REPAIRstarts from a seed training set𝒯0=\{\(qi,di\+\)\}i=1N0\\mathcal\{T\}\_\{0\}=\\\{\(q\_\{i\},d\_\{i\}^\{\+\}\)\\\}\_\{i=1\}^\{N\_\{0\}\}constructed from well\-recognized, domain\-curated public scientific corpora\. The full list is shown in Table[4](https://arxiv.org/html/2609.18262#A0.T4)with further details in Appendix[A\.1](https://arxiv.org/html/2609.18262#A1.SS1)\. This seed set provides minimal in\-domain alignment, but may be insufficient to cover long\-tailed entities and condition\-sensitive relations\. Thus,REPAIRrefines the retriever by iteratively incorporating verified supervision\.

### 3\.3Stage I: Diagnosis

In the first stage, the retriever self\-diagnoses its weaknesses by identifying unreliable low\-margin queries and extracts their long\-tail concepts from its own scoring behavior\.

##### Low\-Margin Query Selection\.

For each training queryqqwith a seed positive documentd\+​\(q\)d^\{\+\}\(q\), we define the positive–negative separation margin:

Δθ​\(q\)=sθ​\(q,d\+​\(q\)\)−maxd∈TopKθ​\(q\)∖\{d\+​\(q\)\}⁡sθ​\(q,d\)\.\\Delta\_\{\\theta\}\(q\)=s\_\{\\theta\}\(q,d^\{\+\}\(q\)\)\-\\hskip\-6\.0pt\\max\_\{d\\in\\mathrm\{TopK\}\_\{\\theta\}\(q\)\\setminus\\\{d^\{\+\}\(q\)\\\}\}s\_\{\\theta\}\(q,d\)\.\(2\)A smallΔθ​\(q\)\\Delta\_\{\\theta\}\(q\)means that the retriever assigns nearly indistinguishable scores tod\+​\(q\)d^\{\+\}\(q\)and top\-ranked negatives, indicating local unreliability\. We form the confusion query set by selecting the lowest\-margin queries:

𝒬conf=Bottom​\-​p%​\(\{Δθ​\(q\)\}q∈𝒬\),\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}=\\mathrm\{Bottom\}\\text\{\-\}p\\%\\big\(\\\{\\Delta\_\{\\theta\}\(q\)\\\}\_\{q\\in\\mathcal\{Q\}\}\\big\),\(3\)where we setp=40%p=40\\%, whose empirical analysis is presented in §[4\.3](https://arxiv.org/html/2609.18262#S4.SS3.SSS0.Px1)\.

#### Long\-tail Confusing Concept Mining

While low margins reveal*where*retrieval fails, REPAIR explains*why*by identifying long\-tail distractors driving model confusion\. We first extract candidate concepts from𝒬conf\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}, and then isolate the actual distractor concepts via two complementary intra\- and inter\-query distractor mining\.

##### Extraction of Candidate Concepts\.

For subsequent analysis, we extract candidate conceptse∈ℰe\\in\\mathcal\{E\}such as scientific concepts and chemical formulas, by applyingMatDetector[Oh et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib55)andChemDataExtractor[Swain and Cole \(2016\)](https://arxiv.org/html/2609.18262#bib.bib8)to the confusing queriesq∈𝒬confq\\in\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}and their retrieved documents\.

##### Intra\-Query Distractor Mining\.

To identify distractors specific to each queryq∈𝒬confq\\in\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}, we contrast the positive documentd\+​\(q\)d^\{\+\}\(q\)against a highly scored negative set𝒟−​\(q\)\\mathcal\{D\}^\{\-\}\(q\)as the top\-kkretrieved documents\. Aggregating over the confusion set,

𝒟\+=\{d\+\(q\)\}q∈𝒬conf,𝒟−=∪q∈𝒬conf𝒟−\(q\),\\mathcal\{D\}^\{\+\}=\\\{d^\{\+\}\(q\)\\\}\_\{q\\in\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\},\\,\\mathcal\{D\}^\{\-\}=\\cup\_\{q\\in\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\}\\mathcal\{D\}^\{\-\}\(q\),\(4\)we score each extracted concepteeusing the Confusing Concept Score \(CCS\):

CCS⁡\(e\)=df⁡\(e,𝒟−\)df⁡\(e,𝒟\+\)\+ϵ,\\mathrm\{CCS\}\(e\)=\\frac\{\\mathrm\{df\}\(e;\\mathcal\{D\}^\{\-\}\)\}\{\\mathrm\{df\}\(e;\\mathcal\{D\}^\{\+\}\)\+\\epsilon\},\(5\)wheredf⁡\(e,⋅\)\\mathrm\{df\}\(e;\\cdot\)is the document frequency ofee, andϵ\>0\\epsilon\>0prevents division by zero\. Thus, a high CCS explicitly identifies distractor concepts that frequently occur in highly scored negative documents \(𝒟−\\mathcal\{D\}^\{\-\}\) but remain rare in the positive documents \(𝒟\+\\mathcal\{D\}^\{\+\}\)\. Finally, we construct𝒞intra\\mathcal\{C\}\_\{\\mathrm\{intra\}\}by selecting the highest\-CCS concept per query\.

##### Inter\-Query Distractor Mining\.

To complement the intra\-query analysis, we identify systemic distractors by clustering queries within𝒬conf\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}that exhibit shared confusion patterns, using the FINCH algorithm[Sarfraz et al\. \(2019\)](https://arxiv.org/html/2609.18262#bib.bib45)\. Rather than relying on complex adjacency matrices, FINCH directly captures the mutual dependency between queries by grouping those that share first nearest neighbors\. This parameter\-free approach is critical for discovering confusion clusters without requiring a predefined cluster number\. From each resulting cluster, we aggregate the previously extracted candidate concepts and select the most frequent one\. This process transforms the clustered query groups into the inter\-query concept set𝒞inter\\mathcal\{C\}\_\{\\mathrm\{inter\}\}\.

##### The Distractor Concept Set\.

Finally, the distractor concept set is defined by𝒞conf=𝒞intra∪𝒞inter\\mathcal\{C\}\_\{\\mathrm\{conf\}\}=\\mathcal\{C\}\_\{\\mathrm\{intra\}\}\\cup\\mathcal\{C\}\_\{\\mathrm\{inter\}\}, which identifies long\-tail scientific concepts responsible for confusion\. Then𝒞conf\\mathcal\{C\}\_\{\\mathrm\{conf\}\}is used in Stage II for the expansion of verified evidence\.

### 3\.4Stage II: Expansion

This stage expands the training data by grounding the distractor concept set𝒞conf\\mathcal\{C\}\_\{\\mathrm\{conf\}\}into verifiable evidence\. This yields rigorously validated training tuples\(qnew,d\+,𝒟cand−\)\(q\_\{\\mathrm\{new\}\},d^\{\+\},\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}\), consisting of a newly augmented queryqnewq\_\{\\mathrm\{new\}\}, its positive documentd\+d^\{\+\}, and its negative set𝒟cand−\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}\.

##### Concept Grounding via External Metadata\.

We ground each concept in𝒞conf\\mathcal\{C\}\_\{\\mathrm\{conf\}\}using external databases such asPubChem[Kim et al\. \(2019\)](https://arxiv.org/html/2609.18262#bib.bib35)andMatProj[Jain et al\. \(2013\)](https://arxiv.org/html/2609.18262#bib.bib7), from which we extract diverse chemical and physical attributes of each concept \(e\.g\., synonyms, molecular weight\)\.

##### Confusing Query Generation\.

From the grounded concepts, we randomly sample 1\-\-3 concepts to prompt the model222Note that we use an LLM\-based retriever \(§[3\.1](https://arxiv.org/html/2609.18262#S3.SS1)\)\., which is instructed to generate a candidate query \(qmodelq\_\{\\mathrm\{model\}\}\) that it finds inherently ambiguous or difficult to resolve\.

##### Verification and Hard Negative Mining\.

To filter out hallucinatedqmodelq\_\{\\mathrm\{model\}\}, we query external APIs \(Semantic Scholar,PubChem, andMatProj\) and discard it if no results are returned\. For valid searches, the top\-matching document defines the ground truth: its title becomes the updated queryqnewq\_\{\\mathrm\{new\}\}, and its content serves as the positive documentd\+d^\{\+\}\. The remaining highly similar documents form𝒟cand−\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}as hard negative candidates\. In §[4\.3](https://arxiv.org/html/2609.18262#S4.SS3.SSS0.Px1), we experiment with the effect of its size\|𝒟cand−\|\|\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}\|on performance\. This generation and verification process continues until the number of valid tuples\(qnew,d\+,𝒟cand−\)\(q\_\{\\mathrm\{new\}\},d^\{\+\},\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}\)reaches twice the size of the initial training data\.

### 3\.5Stage III: Differentiation

Instead of random in\-batch negatives, we construct training triplets\(qnew,d\+,d−\)\(q\_\{\\mathrm\{new\}\},d^\{\+\},d^\{\-\}\)by selecting a single hard negatived−d^\{\-\}from the API\-verified candidates𝒟cand−\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}\. Once mining noise is filtered out, we choosed−d^\{\-\}to confuse the current retriever the most;d−d^\{\-\}is both factually plausible \(API\-ranked\) and empirically challenging \(model\-scored\)\.

##### Consistency Filtering\.

To prevent mining noise[Wang et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib36), we define a localized candidate pool as𝒟pool=\{d\+\}∪𝒟cand−\\mathcal\{D\}\_\{\\mathrm\{pool\}\}=\\\{d^\{\+\}\\\}\\cup\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}, and retain a tuple only if the current retriever ranks the positive documentd\+d^\{\+\}within the topκ=2\\kappa=2of this pool:

𝕀keep\(qnew\)=\[rank\(d\+∣qnew;𝒟pool\)≤κ\]\.\\mathbb\{I\}\_\{\\mathrm\{keep\}\}\(q\_\{\\mathrm\{new\}\}\)=\\mathbb\{1\}\\\!\\left\[\\mathrm\{rank\}\\\!\\left\(d^\{\+\}\\mid q\_\{\\mathrm\{new\}\};\\mathcal\{D\}\_\{\\mathrm\{pool\}\}\\right\)\\leq\\kappa\\right\]\.\(6\)This ensures the query is answerable, keeping the subsequent hard negative mining informative\.

##### Single Hard Negative Selection\.

We extract the hardest negatived−d^\{\-\}from the candidate set𝒟cand−\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}by maximizing the current retriever’s similarity scoresθs\_\{\\theta\}:

d−=argmaxd∈𝒟cand−sθ​\(qnew,d\)\.d^\{\-\}=\\operatorname\*\{argmax\}\_\{d\\in\\mathcal\{D\}^\{\-\}\_\{\\mathrm\{cand\}\}\}s\_\{\\theta\}\(q\_\{\\mathrm\{new\}\},d\)\.\(7\)
Using this single negative, we expect a more semantically meaningful decision boundary than when using multiple easy negatives\. This stage completes a set of verified training triplets\(qnew,d\+,d−\)\(q\_\{\\mathrm\{new\}\},d^\{\+\},d^\{\-\}\)\.

### 3\.6Iterative Contrastive Optimization

The retriever parametersθ\\thetaare updated via contrastive learning; we minimize an InfoNCE objective over the verified triplets\(qnew,d\+,d−\)\(q\_\{\\mathrm\{new\}\},d^\{\+\},d^\{\-\}\):

ℒ⁡\(qnew\)=−log⁡esθ​\(qnew,d\+\)/τ∑d∈\{d\+\}∪𝒩esθ​\(qnew,d\)/τ,\\mathcal\{L\}\(q\_\{\\mathrm\{new\}\}\)=\-\\log\\frac\{e^\{s\_\{\\theta\}\(q\_\{\\mathrm\{new\}\},d^\{\+\}\)/\\tau\}\}\{\\sum\_\{d\\in\\\{d^\{\+\}\\\}\\cup\\mathcal\{N\}\}e^\{s\_\{\\theta\}\(q\_\{\\mathrm\{new\}\},d\)/\\tau\}\},\(8\)whereτ\\tauis a temperature\. As each update shifts the margin landscape\{Δθ​\(q\)\}\\\{\\Delta\_\{\\theta\}\(q\)\\\},REPAIRrepeats the generation\-verification pipeline \(Stages I–III\) for two iterations\. We empirically study how performance varies with the number of iterations in §[4\.3](https://arxiv.org/html/2609.18262#S4.SS3)\. This iterative refinement progressively reshapes the embedding space toward reliable scientific discrimination while avoiding hallucinated supervision or overfitting\.

## 4Experiments

### 4\.1Experiment Setups

##### Tasks and Datasets\.

To assess the model’s robustness, we use an extensive collection of benchmarks that cover a broad range ofscientificdisciplines, from general inquiries tomaterialandbiomedical\-specific challenges\. They evaluate varied retrieval\-oriented tasks, including four IR datasets \(NFCorpus[Boteva et al\. \(2016\)](https://arxiv.org/html/2609.18262#bib.bib67), SciFact[Wadden et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib68), SciDocs[Cohan et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib53), and TREC\-COVID[Voorhees et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib69)\), three QA datasets \(iCliniq[Chen et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib40), and the materials and biomedical subsets of ChemLit\-QA[Wellawatte et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib41)\), one entity linking \(MeSH[Lipscomb \(2000\)](https://arxiv.org/html/2609.18262#bib.bib25)\), one paper recommendation \(RELISH[Singh et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib70);[Brown et al\. \(2019\)](https://arxiv.org/html/2609.18262#bib.bib23)\), and one sentence similarity dataset \(BIOSSES[Soğancıoğlu et al\. \(2017\)](https://arxiv.org/html/2609.18262#bib.bib39)\)\. Full details about datasets are provided in Appendix[B\.2](https://arxiv.org/html/2609.18262#A2.SS2)\.

##### Baselines\.

We compare our method with an extensive set of 19 baselines\. They include one sparse retriever such as BM25[Robertson and Zaragoza \(2009\)](https://arxiv.org/html/2609.18262#bib.bib76)and 14 dense retrievers across various model scales, such as Contriever[Izacard et al\. \(2021b\)](https://arxiv.org/html/2609.18262#bib.bib43), Dragon[Lin et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib74), InstructOR\-L/XL[Su et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib46), E5\-Large\-v2[Wang et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib36), BGE\-Large[Chen et al\. \(2024b\)](https://arxiv.org/html/2609.18262#bib.bib60), DRAMA\-L/1B[Ma et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib62), GTR\-XL/XXL[Ni et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib57), SGPT\-1\.3B/2\.7B[Muennighoff \(2022\)](https://arxiv.org/html/2609.18262#bib.bib31)and Llama2Vec[Li et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib66), RepLLaMA[Ma et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib65), LLM2Vec[BehnamGhader et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib38), E5\-Mistral[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib58), CPT\-text\-XL[Neelakantan et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib13), and Promptriever[Weller et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib64)\. We also include four models specialized for scientific domains: SciMult[Zhang et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib52), SPECTER 2\.0[Singh et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib70), MedCPT[Jin et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib12), andBMRetrieverseries \(410M/2B/7B\)[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)\. Details about the baselines are provided in the Appendix[B\.1](https://arxiv.org/html/2609.18262#A2.SS1)\.

##### Training\.

We train Qwen2\.5\-0\.5B/1\.5B/7B with scientific seed data with a particular focus on materials science[Tshitoyan et al\. \(2019\)](https://arxiv.org/html/2609.18262#bib.bib14);[Gupta et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib5);[Trewartha et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib4)and biomedical domains[Bajaj et al\. \(2016\)](https://arxiv.org/html/2609.18262#bib.bib34);[Wang et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib59);[Xiong et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib33);[Chen et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib29)\. More training details are provided in Appendix[A\.2](https://arxiv.org/html/2609.18262#A1.SS2)\.

TaskScale\# PairsDataAug\.Standard IRAVG\.Sent\. Sim\.AVG\.ModelNFCorpusSciFactSciDocsCOVIDBIOSSESBM25\-\-✓0\.3250\.6650\.1580\.6560\.451\-Contriever110M1\.5B✓0\.3280\.6770\.1650\.5960\.4420\.8330\.520Dragon110M28\.5M✓0\.3390\.6790\.1590\.7590\.4840\.8190\.540SPECTER 2\.0110M3\.3M0\.2280\.671\-0\.584\-\-\-SciMult110M5\.5M0\.3080\.707\-0\.712\-\-\-MedCPT220M255M✓0\.3400\.7240\.1230\.6970\.4710\.8370\.532InstructOR\-L335M1\.24M✓0\.3410\.6430\.1860\.5810\.4380\.8440\.505E5\-Large\-v2†660M271M✓0\.3710\.7260\.2010\.6650\.4910\.8360\.548BGE\-Large∗‡895M2\.8B✓0\.3450\.7230\.2220\.7530\.5110\.8040\.560BMRetriever\-410M410M11\.4M✓0\.3210\.7110\.1670\.8310\.5080\.8400\.563DRAMA\-L300M127M✓0\.3240\.6510\.1380\.5000\.4030\.7250\.442REPAIR\-500M\(ours\)500M4M✓0\.3760\.6800\.1960\.8120\.5160\.8530\.583InstructOR\-XL1\.5B1\.24M✓0\.3600\.6460\.1740\.7130\.4730\.8420\.547GTR\-XL1\.2B2\.7B✓0\.3430\.6350\.1590\.5840\.4300\.7890\.502GTR\-XXL4\.8B2\.7B✓0\.3420\.6620\.1610\.5010\.4170\.8190\.497SGPT\-1\.3B1\.3Bunknown✓0\.3200\.6820\.1620\.7300\.4730\.8300\.545SGPT\-2\.7B2\.7Bunknown✓0\.3390\.7010\.1660\.7520\.4890\.8480\.561BMRetriever\-2B2B10M✓0\.3510\.7600\.1990\.8630\.5430\.8280\.600DRAMA\-1B1B127M✓0\.1580\.7070\.1450\.4120\.3550\.7650\.419REPAIR\-1\.5B\(ours\)1\.5B4M✓0\.3760\.7570\.2010\.8530\.5460\.8490\.607Llama2Vec7B21\.5M✓0\.3720\.7570\.1720\.8530\.539\-\-RepLLaMA7B500K✓0\.3780\.7560\.1810\.8470\.541\-\-LLM2Vec7B2\.7M✓0\.3930\.7880\.2250\.7760\.5450\.8520\.606E5\-Mistral7B1\.8M✓0\.3860\.7640\.1620\.8720\.5460\.8550\.608CPT\-text\-XL175Bunknown0\.4070\.754\-0\.649\-\-\-BMRetriever\-7B7B11\.4M✓0\.3640\.7780\.2010\.8610\.5510\.8470\.610Promptriever7B1M✓0\.3760\.7600\.1760\.8350\.5370\.8610\.602REPAIR\-7B\(ours\)7B4M✓0\.4130\.7890\.2270\.8420\.5680\.8460\.623

Table 1:Experiments on scientific text representation tasks across various model scales\. All scores are reported in nDCG@10\.†\\daggerand‡\\ddaggerdenote the use of reranker distillation and hybrid retrieval, respectively\. We highlight thescientificdomain\-specific retrieval models\. "\#Pairs" and "Sent\. Sim\." stand for the total number of query\-document pairs used for training and Sentence Similarity, respectively\. The best\-performing results are highlighted inboldface, whileunderlinedrepresent the second\-highest scores\.Table 2:Experiments on retrieval\-oriented material and biomedical NLP applications across materials science and biomedical domains\. Here,ChemLit\-QAmat\\text\{ChemLit\-QA\}\_\{\\text\{mat\}\}andChemLit\-QAbiomed\\text\{ChemLit\-QA\}\_\{\\text\{biomed\}\}denote the materials and biomedical categories of the ChemLit\-QA task, respectively\. nDCG refers to nDCG@20, except for the paper recommendation task\. The best\-performing results are highlighted inboldface, whileunderlinerepresent the second\-highest scores\.
##### Evaluation\.

To ensure rigorous evaluation, we follow all experiment setups of BMRetriever\([Xu et al\., 2024](https://arxiv.org/html/2609.18262#bib.bib47)\), including dataset curation, task formulation, baseline selection, and evaluation metrics\. Following this framework, we categorize our evaluation into two distinct areas: fundamental text representation tasks \(Table[1](https://arxiv.org/html/2609.18262#S4.T1)\) and retrieval\-oriented material and biomedical applications \(Table[2](https://arxiv.org/html/2609.18262#S4.T2)\)\. Standard information retrieval is evaluated with nDCG@10, and sentence similarity with Spearman’s rank correlation over cosine similarity\. For material and biomedical applications, we report Recall@\{5, 20\} and nDCG@20 for question answering, mean reciprocal rank \(MRR\)@5 and Recall@\{1, 5\} for entity linking, and mean average precision \(MAP\) and nDCG for paper recommendation\([Singh et al\., 2023](https://arxiv.org/html/2609.18262#bib.bib70)\)\.

### 4\.2Main Results

##### Results on Text Representation Tasks\.

Table[1](https://arxiv.org/html/2609.18262#S4.T1)presents a comprehensive evaluation of embedding quality across four science IR and one sentence similarity benchmarks\. Across different scales,REPAIRconsistently outperforms baseline methods\. While some strong baselines heavily rely on computationally expensive reranker distillation \(e\.g\., E5\-Large\-v2†[Wang et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib36)\) or complex hybrid systems requiring sparse inverted indices \(e\.g\., BGE\-Large‡[Chen et al\. \(2024b\)](https://arxiv.org/html/2609.18262#bib.bib60)\),

REPAIRdemonstrates exceptional parameter and data efficiency\. First, in terms of parameter efficiency, it exhibits competitive performance against substantially larger baselines\. Specifically,REPAIR\-500M successfully surpasses both the SGPT\-2\.7B[Muennighoff \(2022\)](https://arxiv.org/html/2609.18262#bib.bib31)and the GTR\-XXL[Ni et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib57)with 4\.8B parameters\. Furthermore,REPAIR\-1\.5B outperforms massive 7B LLM\-based retrievers, such as LLM2Vec[BehnamGhader et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib38)and Promptriever[Weller et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib64)\. Second, from a data efficiency perspective,REPAIRuses only 4M fact\-verified instances\. In stark contrast, it significantly exceeds the performance ofBMRetriever\-2B, which consumes 11\.4M synthetic pairs, as well as BGE\-Large, a model trained on an extensive corpus of 2\.8B text pairs\.

##### Results on Retrieval\-Oriented Material and Biomedical Applications\.

Figure 3:Performance variations ofREPAIRmodels according to the number of iterations\. Iteration 0 indicates the seed\-only training baseline\. The percentages above the bars are the relative nDCG@10 improvement of Iteration 2 over Iteration 0\.Figure 4:Effect of fact\-verified data across model capacities\. The evaluation is based on the nDCG@10 metric using three different model sizes \(0\.5B, 1\.5B, and 7B\)\.Table 3:A case study ofREPAIRgenerating fact\-verified triplets \(qnewq\_\{\\text\{new\}\},d\+d^\{\+\},d−d^\{\-\}\) to resolvelong\-tail conceptconfusion in the materials science and biomedical domains\.Table[2](https://arxiv.org/html/2609.18262#S4.T2)highlights the robust generalization ofREPAIRacross specialized material and biomedical downstream tasks\. With mid\-sized parameters,BMRetriever\-2B[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)exhibits slightly higher overall performance in the biomedical domains, these marginal gaps are primarily attributable to its larger scale and an exhaustive multi\-task instruction fine\-tuning\. That is,BMRetriever\-2Bis explicitly aligned with downstream tasks by aggregating human\-annotated datasets and synthesizing task\-specific scenarios to adapt to various input formats\.

In contrast,REPAIRachieves exceptional generalization without this task\-specific engineering\. Not only doesREPAIR\-1\.5Bdirectly outperform the largerBMRetriever\-2Bon several specific tasks, but our 500M and 7B variants consistently achieves the best performance across all evaluated tasks\. By simply utilizing a unified query format,REPAIReliminates the overhead of curating diverse query\-passage pairs\.REPAIRcan seamlessly adapts to diverse and complex scenarios, including question answering to entity linking, demonstrating remarkable parameter efficiency\. Our fact\-verified grounding approach can establish a universally adaptable semantic space, rather than memorizing downstream task instructions\.

### 4\.3Ablation Studies and Analyses

We perform ablation studies to isolate the contributions of two key design choices inREPAIR, iterative refinement and fact\-verified data augmentation\. We also conduct an empirical analysis to validate the single\-positive assumption underlying our diagnosis stage\. Detailed quantitative results are provided in Appendix[C](https://arxiv.org/html/2609.18262#A3)\.

##### Effect of Iterative Refinement\.

Figure[3](https://arxiv.org/html/2609.18262#S4.F3)evaluates iterative self\-diagnosis acrossREPAIRmodels of different sizes \(500M, 1\.5B, 7B\)\. By recomputing the margin landscapeΔθ​\(q\)\{\\Delta\_\{\\theta\}\(q\)\}at each step, our approach dynamically resolves long\-tail failure modes, yielding consistent performance improvements across all model capacities as iterations progress\. Two iterations raise the average nDCG@10 by10\.0%10\.0\\%to11\.4%11\.4\\%over the seed\-only baseline, as annotated above the bars\. While performance continues to rise with additional iterations, the marginal gains progressively diminish, whereas the per\-iteration cost stays constant \(Appendix[C\.5](https://arxiv.org/html/2609.18262#A3.SS5)\)\. Given that the most substantial improvements occur within the first two rounds, we set the default number of iterations to two for all of our experiments\.

##### Robustness of the Selection Parameters\.

To avoid tuning the pipeline for each model, we use one setting for every model size and every iteration:p=40%p=40\\%andk=30k=30\. Both values come from separate measurements\. Raisingppbeyond40%40\\%finds few new concepts \(Tables[9](https://arxiv.org/html/2609.18262#A3.T9)and[10](https://arxiv.org/html/2609.18262#A3.T10)\), and six measures of negative quality all point tok=30k=30\(Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\)\. With this one setting, the average nDCG@10 improves at every iteration for all three model sizes \(0\.530→0\.547→0\.5830\.530\\to 0\.547\\to 0\.583for 500M,0\.546→0\.586→0\.6070\.546\\to 0\.586\\to 0\.607for 1\.5B, and0\.559→0\.589→0\.6230\.559\\to 0\.589\\to 0\.623for 7B\), and it keeps improving up to the fourth iteration \(Table[8](https://arxiv.org/html/2609.18262#A2.T8)\)\. One setting is therefore enough across model sizes and iterations, with no re\-tuning\.

##### Effect of Fact\-Verified Data Beyond Model Capacity\.

To verify thatREPAIR’s improvements stem from our data refinement rather than the Qwen2\.5’s inherent capacity, we isolate the effect of the augmented training data\. As shown in Figure[4](https://arxiv.org/html/2609.18262#S4.F4), we train the Qwen2\.5 models entirely on the augmented dataset used in a strong baseline,BMRetriever\. Across all parameter scales, these models yield lower retrieval performance compared toREPAIR\. This confirms that our core approach, resolving long\-tail concept confusion through API\-guided, fact\-verified iterative refinement, is the fundamental driver of enhanced scientific retrieval, proving that the quality of well\-curated data outweighs the backbone capacity\. We reach the same conclusion when we replace the backbone instead of the data, applying our pipeline to four backbones from different model families \(Appendix[C\.4](https://arxiv.org/html/2609.18262#A3.SS4)\)\.

##### Validity of the Single\-Positive Assumption in Diagnosis\.

To efficiently isolate long\-tail confusions, our diagnosis stage extracts distractor concepts by treating the retrieved documents as negatives against a single positive\. To ensure false negatives do not skew this diagnosis, we analyze the direct citation relationships between the anchor positives and the retrieved negatives\. In scientific literature, direct citation relationships are established as a rigorous proxy for true semantic equivalence[Cohan et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib53)\. Our analysis reveals an overwhelmingly low citation overlap of just<0\.00001%<0\.00001\\%\. For comparison, we ran the same check on SciDocs pairs that are known to cite each other, and only5\.80%5\.80\\%of them showed a citation link\. This is the highest rate our lookup can detect, and our hard negatives fall far below it, at the same level as randomly paired documents \(Appendix[C\.7](https://arxiv.org/html/2609.18262#A3.SS7)\)\. Since scientific text exhibits extreme fact\-sensitivity, lexically similar documents without citation links are overwhelmingly true hard negatives rather than false negatives\. Furthermore, as our concept mining statistically aggregates signals across a large query set, this infinitesimally small noise is heavily diluted\. This confirms that our approach robustly captures genuine diagnostic signals without contamination\.

Finally, beyond the nine benchmarks of Tables[1](https://arxiv.org/html/2609.18262#S4.T1)and[2](https://arxiv.org/html/2609.18262#S4.T2),REPAIRretains its advantage on three held\-out scientific benchmarks spanning multi\-aspect scientific IR, physics community QA, and broad\-coverage science QA \(Appendix[C\.6](https://arxiv.org/html/2609.18262#A3.SS6)\)\.

### 4\.4Case study

Table[3](https://arxiv.org/html/2609.18262#S4.T3)demonstrates howREPAIRresolves the confusion of long\-tail concepts using augmented fact\-verified triplets \(qnewq\_\{\\text\{new\}\},d\+d^\{\+\},d−d^\{\-\}\)\. By grounding identified concepts \(𝒞conf\\mathcal\{C\}\_\{\\text\{conf\}\}\), the framework generates targeted queries that probe precise yet underrepresented distinctions\. For example, grounding*Pb\-based perovskite*constructs a query that pairs a positive document \(d\+d^\{\+\}\) detailing*dynamic symmetry breaking*with a fact\-contrastive hard negative \(d−d^\{\-\}\) addressing*static Rashba splitting*\. Similarly, grounding*poly\(dl\-lactic acid\)*yields a query that retrieves a positive document \(d\+d^\{\+\}\) detailing its*size\-dependent hydrolytic degradation*, while isolating a fact\-contrastive negative \(d−d^\{\-\}\) that discusses its chemical formula*C3H6O3*in the unrelated context of*zinc\-ion batteries \(AZIBs\)*\. This concept\-driven, evidence\-based expansion enables the retriever to resolve fine\-grained factual distinctions\.

## 5Conclusion

We presentedREPAIR, a self\-evolving framework that has effectively addressed the persistent challenges of long\-tailed entities and high fact\-sensitivity in scientific retrieval\. By grounding iterative refinement in API\-guided evidence, we demonstrated that diagnosing specific knowledge gaps outperforms indiscriminate data augmentation\. While we focused on materials science and biomedicine, movingREPAIRto a new domain is straightforward\. Stages I and III depend only on the retriever and its training data, so they transfer unchanged, and only the evidence source in Stage II has to be replaced\. Within science this means plugging in resources such as ChEMBL, the NIST WebBook, or NASA ADS\. Beyond it, the same recipe applies to any field that has an authoritative database, such as USPTO for patents or SEC EDGAR for finance\. We leave a full study of physics, engineering, and the social sciences to future work\.

Ultimately, our work established a new paradigm, proving that integrating external verification into the training loop is essential for trustworthy knowledge discovery, and encouraging future research to prioritize rigorous factual verification\.

## Limitations

REPAIRimproves retrieval through an iterative loop, and each iteration carries an additional cost\. In practice this cost is bounded, since the loop saturates at the second iteration across all model scales; we report the per\-stage breakdown in Appendix[C\.5](https://arxiv.org/html/2609.18262#A3.SS5)\. The other side of that saturation is a limitation: deeper iterations buy little, with average nDCG@10 improving by at most\+0\.005\+0\.005beyond the second iteration\. Simply extending the loop is therefore not a route to further gains, and widening the evidence expansion within each iteration is a more promising direction we leave to future work\.

A second limitation is thatREPAIRis bounded by its external verifiers\. Concepts the scientific APIs cannot resolve are discarded rather than approximated \(§[3\.4](https://arxiv.org/html/2609.18262#S3.SS4)\), which keeps supervision factual but leaves those regions of the long tail untouched\. Coverage thus extends only as far as the available scientific resources do, and reaching domains beyond materials science and biomedicine requires plugging in an appropriate API for that field\.

## Acknowledgements

This work was supported by Institute of Information & communications Technology Planning & Evaluation \(IITP\) grant funded by the Korea government \(MSIT\) \(No\. RS\-2022\-II220156, Fundamental research on continual meta\-learning for quality enhancement of casual videos and their 3D metaverse transformation\), Institute of Information & communications Technology Planning & Evaluation \(IITP\) grant funded by the Korea government\(MSIT\) \(No\.RS\-2026\-25524173, Ultro\-Long\-Term Hierarchical Memory and Reasoning Architecture for Next\-Generation Omnimodal Agents\), Basic Science Research Program through the National Research Foundation of Korea\(NRF\) funded by the Ministry of Education\(RS\-2023\-00274280\), the Institute of Information & Communications Technology Planning & Evaluation\(IITP\) grant funded by the Korea government\(MSIT\) \(RS\-2025\-25442338, AI star Fellowship Support Program\(Seoul National Univ\.\)\), Institute of Information & communications Technology Planning & Evaluation \(IITP\) grant funded by the Korea government \(MSIT\) \(No\. RS\-2021\-II211343, Artificial Intelligence Graduate School Program \(Seoul National University\)\) , and the AI Seoul Tech Research Support Program of the Seoul Future Foundation\. Gunhee Kim is the corresponding author\.

## References

- Andonianet al\.\(2023\)A\. Andonian, S\. Biderman, S\. Black, P\. Gali, L\. Gao, E\. Hallahan, J\. Levy\-Kramer, C\. Leahy, L\. Nestler, K\. Parker,et al\.GPT\-neox: large scale autoregressive language modeling in pytorch\.Zenodo\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.15.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.16.2)\.
- Bajajet al\.\(2016\)P\. Bajaj, D\. Campos, N\. Craswell, L\. Deng, J\. Gao, X\. Liu, R\. Majumder, A\. McNamara, B\. Mitra, T\. Nguyen,et al\.Ms marco: a human generated machine reading comprehension dataset\.arXiv preprint arXiv:1611\.09268\.Cited by:[§A\.3](https://arxiv.org/html/2609.18262#A1.SS3.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- BehnamGhaderet al\.\(2024\)P\. BehnamGhader, V\. Adlakha, M\. Mosbach, D\. Bahdanau, N\. Chapados, and S\. ReddyLlm2vec: large language models are secretly powerful text encoders\.arXiv preprint arXiv:2404\.05961\.Cited by:[3rd item](https://arxiv.org/html/2609.18262#A2.I4.i3.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.21.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px1.p2.1)\.
- Beltagyet al\.\(2019\)I\. Beltagy, K\. Lo, and A\. CohanSciBERT: a pretrained language model for scientific text\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing,pp\. 3615–3620\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.5.2)\.
- Bidermanet al\.\(2023\)S\. Biderman, H\. Schoelkopf, Q\. G\. Anthony, H\. Bradley, K\. O’Brien, E\. Hallahan, M\. A\. Khan, S\. Purohit, U\. S\. Prashanth, E\. Raff,et al\.Pythia: a suite for analyzing large language models across training and scaling\.InInternational conference on machine learning,pp\. 2397–2430\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.11.2)\.
- Bleiet al\.\(2003\)D\. M\. Blei, A\. Y\. Ng, and M\. I\. JordanLatent dirichlet allocation\.Journal of machine Learning research3\(Jan\),pp\. 993–1022\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p1.1)\.
- Botevaet al\.\(2016\)V\. Boteva, D\. Gholipour, A\. Sokolov, and S\. RiezlerA full\-text learning to rank dataset for medical information retrieval\.InEuropean Conference on Information Retrieval,pp\. 716–722\.Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px1.p1.1),[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Brownet al\.\(2019\)P\. Brown, R\. Consortium, and Y\. ZhouLarge expert\-curated database for benchmarking document similarity detection in biomedical literature search\.Database J\. Biol\. Databases Curation2019,pp\. baz085\.External Links:[Link](https://doi.org/10.1093/database/baz085),[Document](https://dx.doi.org/10.1093/DATABASE/BAZ085)Cited by:[§B\.2\.5](https://arxiv.org/html/2609.18262#A2.SS2.SSS5.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.23.2)\.
- Chenet al\.\(2024a\)J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. LiuBge m3\-embedding: multi\-lingual, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.arXiv preprint arXiv:2402\.032164\(5\)\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p1.1)\.
- Chenet al\.\(2024b\)J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. LiuM3\-embedding: multi\-linguality, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.InFindings of the association for computational linguistics: ACL 2024,pp\. 2318–2335\.Cited by:[8th item](https://arxiv.org/html/2609.18262#A2.I2.i8.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.10.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px1.p1.1)\.
- Chenet al\.\(2021\)Q\. Chen, A\. Allot, and Z\. LuLitCovid: an open database of covid\-19 literature\.Nucleic acids research49\(D1\),pp\. D1534–D1540\.Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.8.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- Chenet al\.\(2020\)S\. Chen, Z\. Ju, X\. Dong, H\. Fang, S\. Wang, Y\. Yang, J\. Zeng, R\. Zhang, R\. Zhang, M\. Zhou, P\. Zhu, and P\. XieMedDialog: A large\-scale medical dialogue dataset\.CoRRabs/2004\.03329\.External Links:[Link](https://arxiv.org/abs/2004.03329),2004\.03329Cited by:[§B\.2\.3](https://arxiv.org/html/2609.18262#A2.SS2.SSS3.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Chianget al\.\(2025\)Y\. Chiang, E\. Hsieh, C\. Chou, and J\. RiebesellLLaMP: large language model made powerful for high\-fidelity materials knowledge retrieval\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 25200–25232\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Choudharyet al\.\(2022\)K\. Choudhary, B\. DeCost, C\. Chen, A\. Jain, F\. Tavazza, R\. Cohn, C\. W\. Park, A\. Choudhary, A\. Agrawal, S\. J\. Billinge,et al\.Recent advances and applications of deep learning methods in materials science\.npj Computational Materials8\(1\),pp\. 59\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Cohanet al\.\(2020\)A\. Cohan, S\. Feldman, I\. Beltagy, D\. Downey, and D\. S\. WeldSpecter: document\-level representation learning using citation\-informed transformers\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 2270–2282\.Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px3.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1),[§4\.3](https://arxiv.org/html/2609.18262#S4.SS3.SSS0.Px4.p1.1)\.
- Conneauet al\.\(2020\)A\. Conneau, K\. Khandelwal, N\. Goyal, V\. Chaudhary, G\. Wenzek, F\. Guzmán, E\. Grave, M\. Ott, L\. Zettlemoyer, and V\. StoyanovUnsupervised cross\-lingual representation learning at scale\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 8440–8451\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.10.2)\.
- Deerwesteret al\.\(1990\)S\. Deerwester, S\. T\. Dumais, G\. W\. Furnas, T\. K\. Landauer, and R\. HarshmanIndexing by latent semantic analysis\.Journal of the American society for information science41\(6\),pp\. 391–407\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p1.1)\.
- Guet al\.\(2021\)Y\. Gu, R\. Tinn, H\. Cheng, M\. Lucas, N\. Usuyama, X\. Liu, T\. Naumann, J\. Gao, and H\. PoonDomain\-specific language model pretraining for biomedical natural language processing\.ACM Transactions on Computing for Healthcare3,pp\. 1–23\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.6.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.7.2)\.
- Guptaet al\.\(2022\)T\. Gupta, M\. Zaki, N\. A\. Krishnan, and MausamMatSciBERT: a materials domain language model for text mining and information extraction\.npj Computational Materials8,pp\. 102\.Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.3.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- Hofmann \(1999\)T\. HofmannProbabilistic latent semantic indexing\.InProceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval,pp\. 50–57\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p1.1)\.
- Hoogeveenet al\.\(2015\)D\. Hoogeveen, K\. M\. Verspoor, and T\. BaldwinCQADupStack: a benchmark data set for community question\-answering research\.InProceedings of the 20th Australasian Document Computing Symposium \(ADCS\),pp\. 1–8\.Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px6.p1.1),[§C\.6](https://arxiv.org/html/2609.18262#A3.SS6.p1.1)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InThe Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25\-29, 2022,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[§A\.2](https://arxiv.org/html/2609.18262#A1.SS2.p1.1)\.
- Izacardet al\.\(2021a\)G\. Izacard, M\. Caron, L\. Hosseini, S\. Riedel, P\. Bojanowski, A\. Joulin, and E\. GraveUnsupervised dense information retrieval with contrastive learning\.arXiv preprint arXiv:2112\.09118\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p1.1)\.
- Izacardet al\.\(2021b\)G\. Izacard, M\. Caron, L\. Hosseini, S\. Riedel, P\. Bojanowski, A\. Joulin, and E\. GraveUnsupervised dense information retrieval with contrastive learning\.Transactions on Machine Learning Research\.Cited by:[1st item](https://arxiv.org/html/2609.18262#A2.I2.i1.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.3.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Jainet al\.\(2013\)A\. Jain, S\. P\. Ong, G\. Hautier, W\. Chen, W\. D\. Richards, S\. Dacek, S\. Cholia, D\. Gunter, D\. Skinner, G\. Ceder,et al\.Commentary: The Materials Project: a materials genome approach to accelerating materials innovation\.APL Materials1\(1\),pp\. 011002\.Cited by:[§3\.4](https://arxiv.org/html/2609.18262#S3.SS4.SSS0.Px1.p1.1)\.
- Jianget al\.\(2025\)X\. Jiang, W\. Wang, S\. Tian, H\. Wang, T\. Lookman, and Y\. SuApplications of natural language processing and large language models in materials discovery\.npj Computational Materials11\(1\),pp\. 79\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p2.1)\.
- Jinet al\.\(2023\)Q\. Jin, W\. Kim, Q\. Chen, D\. C\. Comeau, L\. Yeganova, W\. J\. Wilbur, and Z\. LuMedcpt: contrastive pre\-trained transformers with large\-scale pubmed search logs for zero\-shot biomedical information retrieval\.Bioinformatics39\(11\),pp\. btad651\.Cited by:[5th item](https://arxiv.org/html/2609.18262#A2.I2.i5.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.7.1),[§1](https://arxiv.org/html/2609.18262#S1.p2.1),[§2](https://arxiv.org/html/2609.18262#S2.p2.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Kamallooet al\.\(2024\)E\. Kamalloo, N\. Thakur, C\. Lassance, X\. Ma, J\. Yang, and J\. LinResources for brewing beir: reproducible reference models and statistical analyses\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’24,New York, NY, USA,pp\. 1431–1440\.External Links:ISBN 9798400704314,[Link](https://doi.org/10.1145/3626772.3657862),[Document](https://dx.doi.org/10.1145/3626772.3657862)Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p2.1),[§2](https://arxiv.org/html/2609.18262#S2.p2.1)\.
- Kimet al\.\(2024\)J\. Kim, Y\. Kim, J\. Park, Y\. Oh, S\. Kim, and S\. LeeMELT: materials\-aware continued pre\-training for language model adaptation to materials science\.arXiv preprint arXiv:2410\.15126\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p3.1)\.
- Kimet al\.\(2019\)S\. Kim, J\. Chen, T\. Cheng, A\. Gindulyte, J\. He, S\. He, Q\. Li, B\. A\. Shoemaker, P\. A\. Thiessen, B\. Yu,et al\.PubChem 2019 update: improved access to chemical data\.Nucleic acids research47\(D1\),pp\. D1102–D1109\.Cited by:[§3\.4](https://arxiv.org/html/2609.18262#S3.SS4.SSS0.Px1.p1.1)\.
- Kononovaet al\.\(2021\)O\. Kononova, T\. He, H\. Huo, A\. Trewartha, E\. A\. Olivetti, and G\. CederOpportunities and challenges of text mining in materials research\.Iscience24\(3\)\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Labraket al\.\(2024\)Y\. Labrak, A\. Bazoge, E\. Morin, P\. Gourraud, M\. Rouvier, and R\. DufourBiomistral: a collection of open\-source pretrained large language models for medical domains\.InFindings of the association for computational linguistics: acl 2024,pp\. 5848–5864\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.24.2)\.
- Liet al\.\(2024\)C\. Li, Z\. Liu, S\. Xiao, Y\. Shao, and D\. LianLlama2vec: unsupervised adaptation of large language models for dense retrieval\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3490–3500\.Cited by:[1st item](https://arxiv.org/html/2609.18262#A2.I4.i1.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.19.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Linet al\.\(2023\)S\. Lin, A\. Asai, M\. Li, B\. Oguz, J\. Lin, Y\. Mehdad, W\. Yih, and X\. ChenHow to train your dragon: diverse augmentation towards generalizable dense retrieval\.InFindings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6\-10, 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Findings of ACL, Vol\.EMNLP 2023,pp\. 6385–6400\.External Links:[Link](https://doi.org/10.18653/v1/2023.findings-emnlp.423),[Document](https://dx.doi.org/10.18653/V1/2023.FINDINGS-EMNLP.423)Cited by:[2nd item](https://arxiv.org/html/2609.18262#A2.I2.i2.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.4.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Lipscomb \(2000\)C\. E\. LipscombMedical subject headings \(mesh\)\.Bulletin of the Medical Library Association88\(3\),pp\. 265\.Cited by:[§B\.2\.4](https://arxiv.org/html/2609.18262#A2.SS2.SSS4.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Loet al\.\(2020\)K\. Lo, L\. L\. Wang, M\. Neumann, R\. Kinney, and D\. S\. WeldS2ORC: the semantic scholar open research corpus\.InProceedings of the 58th annual meeting of the association for computational linguistics,pp\. 4969–4983\.Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.5.2)\.
- Loshchilov and Hutter \(2019\)I\. Loshchilov and F\. HutterDecoupled weight decay regularization\.In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6\-9, 2019,External Links:[Link](https://openreview.net/forum?id=Bkg6RiCqY7)Cited by:[§A\.2](https://arxiv.org/html/2609.18262#A1.SS2.p1.1)\.
- Maet al\.\(2025\)X\. Ma, X\. V\. Lin, B\. Oguz, J\. Lin, W\. Yih, and X\. ChenDRAMA: diverse augmentation from large language models to smaller dense retrievers\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 30170–30186\.Cited by:[10th item](https://arxiv.org/html/2609.18262#A2.I2.i10.p1.1),[5th item](https://arxiv.org/html/2609.18262#A2.I3.i5.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.18.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Maet al\.\(2024\)X\. Ma, L\. Wang, N\. Yang, F\. Wei, and J\. LinFine\-tuning llama for multi\-stage text retrieval\.InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 2421–2425\.Cited by:[2nd item](https://arxiv.org/html/2609.18262#A2.I4.i2.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.20.1),[§C\.5](https://arxiv.org/html/2609.18262#A3.SS5.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Menget al\.\(2024\)R\. Meng, Y\. Liu, S\. R\. Joty, C\. Xiong, Y\. Zhou, and S\. YavuzSFR\-Embedding\-Mistral: enhance text retrieval with transfer learning\.Note:Salesforce AI Research BlogExternal Links:[Link](https://www.salesforce.com/blog/sfr-embedding/)Cited by:[§C\.5](https://arxiv.org/html/2609.18262#A3.SS5.SSS0.Px3.p1.1)\.
- Merityet al\.\(2017\)S\. Merity, C\. Xiong, J\. Bradbury, and R\. SocherPointer sentinel mixture models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§A\.3](https://arxiv.org/html/2609.18262#A1.SS3.p1.1)\.
- Muennighoff \(2022\)N\. MuennighoffSgpt: gpt sentence embeddings for semantic search\.arXiv preprint arXiv:2202\.08904\.Cited by:[3rd item](https://arxiv.org/html/2609.18262#A2.I3.i3.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.15.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.16.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px1.p2.1)\.
- Neelakantanet al\.\(2022\)A\. Neelakantan, T\. Xu, R\. Puri, A\. Radford, J\. M\. Han, J\. Tworek, Q\. Yuan, N\. Tezak, J\. W\. Kim, C\. Hallacy,et al\.Text and code embeddings by contrastive pre\-training\.arXiv preprint arXiv:2201\.10005\.Cited by:[5th item](https://arxiv.org/html/2609.18262#A2.I4.i5.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.23.1),[§2](https://arxiv.org/html/2609.18262#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Niet al\.\(2022\)J\. Ni, C\. Qu, J\. Lu, Z\. Dai, G\. H\. Abrego, J\. Ma, V\. Zhao, Y\. Luan, K\. Hall, M\. Chang,et al\.Large dual encoders are generalizable retrievers\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 9844–9855\.Cited by:[2nd item](https://arxiv.org/html/2609.18262#A2.I3.i2.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.12.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.13.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.14.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.3.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.4.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.8.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.9.2),[§2](https://arxiv.org/html/2609.18262#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px1.p2.1)\.
- Ocanaet al\.\(2025\)A\. Ocana, A\. Pandiella, C\. Privat, I\. Bravo, M\. Luengo\-Oroz, E\. Amir, and B\. GyorffyIntegrating artificial intelligence in drug discovery and early drug development: a transformative approach\.Biomarker Research13,pp\. 45\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Ohet al\.\(2025\)Y\. Oh, J\. Park, J\. Kim, S\. Kim, and S\. LeeIncorporating domain knowledge into materials tokenization\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 9623–9644\.External Links:[Link](https://aclanthology.org/2025.acl-long.474/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.474),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p3.1),[§2](https://arxiv.org/html/2609.18262#S2.p3.1),[§3\.3](https://arxiv.org/html/2609.18262#S3.SS3.SSSx1.Px1.p1.1)\.
- Olivettiet al\.\(2020\)E\. A\. Olivetti, J\. M\. Cole, E\. Kim, O\. Kononova, G\. Ceder, T\. Y\. Han, and A\. M\. HiszpanskiData\-driven materials research enabled by natural language processing and information extraction\.Applied Physics Reviews7\(4\),pp\. 21106\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p2.1)\.
- Onget al\.\(2025\)J\. C\. L\. Ong, L\. Jin, K\. Elangovan, G\. Y\. San Lim, D\. Y\. Z\. Lim, G\. G\. R\. Sng, Y\. H\. Ke, J\. Y\. M\. Tung, R\. J\. Zhong, C\. M\. Y\. Koh,et al\.Large language model as clinical decision support system augments medication safety in 16 clinical specialties\.Cell Reports Medicine6\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Palet al\.\(2023\)A\. Pal, L\. K\. Umapathi, and M\. SankarasubbuMed\-HALT: medical domain hallucination test for large language models\.InProceedings of the 27th Conference on Computational Natural Language Learning \(CoNLL\),J\. Jiang, D\. Reitter, and S\. Deng \(Eds\.\),Singapore,pp\. 314–334\.External Links:[Link](https://aclanthology.org/2023.conll-1.21/),[Document](https://dx.doi.org/10.18653/v1/2023.conll-1.21)Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p3.1),[§2](https://arxiv.org/html/2609.18262#S2.p3.1)\.
- Peiet al\.\(2025\)Z\. Pei, J\. Yin, and J\. ZhangLanguage models for materials discovery and sustainability: progress, challenges, and opportunities\.Progress in Materials Science,pp\. 101495\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Pilania \(2021\)G\. PilaniaMachine learning in materials science: from explainable predictions to autonomous design\.Computational Materials Science193,pp\. 110360\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p2.1)\.
- Robertson and Zaragoza \(2009\)S\. Robertson and H\. ZaragozaThe probabilistic relevance framework: bm25 and beyond\.Vol\.4,Now Publishers Inc\.Cited by:[1st item](https://arxiv.org/html/2609.18262#A2.I1.i1.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Sarfrazet al\.\(2019\)S\. Sarfraz, V\. Sharma, and R\. StiefelhagenEfficient parameter\-free clustering using first neighbor relations\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 8934–8943\.Cited by:[§3\.3](https://arxiv.org/html/2609.18262#S3.SS3.SSSx1.Px3.p1.1)\.
- Sharmaet al\.\(2025\)K\. Sharma, P\. Kumar, and Y\. LiOG\-rag: ontology\-grounded retrieval\-augmented generation for large language models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 32950–32969\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Singhet al\.\(2023\)A\. Singh, M\. D’Arcy, A\. Cohan, D\. Downey, and S\. FeldmanScirepeval: a multi\-format benchmark for scientific document representations\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 5548–5566\.Cited by:[3rd item](https://arxiv.org/html/2609.18262#A2.I2.i3.p1.1),[§B\.2\.5](https://arxiv.org/html/2609.18262#A2.SS2.SSS5.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.5.1),[§1](https://arxiv.org/html/2609.18262#S1.p2.1),[§2](https://arxiv.org/html/2609.18262#S2.p2.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px4.p1.1)\.
- Soğancıoğluet al\.\(2017\)G\. Soğancıoğlu, H\. Öztürk, and A\. ÖzgürBIOSSES: a semantic sentence similarity estimation system for the biomedical domain\.Bioinformatics33\(14\),pp\. i49–i58\.Cited by:[§B\.2\.2](https://arxiv.org/html/2609.18262#A2.SS2.SSS2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Sohnet al\.\(2025\)J\. Sohn, Y\. Park, C\. Yoon, S\. Park, H\. Hwang, M\. Sung, H\. Kim, and J\. KangRationale\-guided retrieval augmented generation for medical question answering\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 12739–12753\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Suet al\.\(2023\)H\. Su, W\. Shi, J\. Kasai, Y\. Wang, Y\. Hu, M\. Ostendorf, W\. Yih, N\. A\. Smith, L\. Zettlemoyer, and T\. YuOne embedder, any task: instruction\-finetuned text embeddings\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 1102–1121\.Cited by:[6th item](https://arxiv.org/html/2609.18262#A2.I2.i6.p1.1),[1st item](https://arxiv.org/html/2609.18262#A2.I3.i1.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.12.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.8.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Swain and Cole \(2016\)M\. C\. Swain and J\. M\. ColeChemDataExtractor: a toolkit for automated extraction of chemical information from the scientific literature\.Journal of Chemical Information and Modeling56,pp\. 1894–1904\.Cited by:[§3\.3](https://arxiv.org/html/2609.18262#S3.SS3.SSSx1.Px1.p1.1)\.
- Teamet al\.\(2024\)G\. Team, T\. Mesnard, C\. Hardin, R\. Dadashi, S\. Bhupatiraju, S\. Pathak, L\. Sifre, M\. Rivière, M\. S\. Kale, J\. Love,et al\.Gemma: open models based on gemini research and technology\.arXiv preprint arXiv:2403\.08295\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.17.2)\.
- Thakuret al\.\(2021\)N\. Thakur, N\. Reimers, A\. Rücklé, A\. Srivastava, and I\. GurevychBEIR: a heterogeneous benchmark for zero\-shot evaluation of information retrieval models\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px6.p1.1),[§C\.6](https://arxiv.org/html/2609.18262#A3.SS6.p1.1)\.
- Touvronet al\.\(2023\)H\. Touvron, L\. Martin, K\. Stone, P\. Albert, A\. Almahairi, Y\. Babaei, N\. Bashlykov, S\. Batra, P\. Bhargava, S\. Bhosale,et al\.Llama 2: open foundation and fine\-tuned chat models\.arXiv preprint arXiv:2307\.09288\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.25.2)\.
- Trewarthaet al\.\(2022\)A\. Trewartha, N\. Walker, H\. Huo, S\. Lee, K\. Cruse, J\. Dagdelen, A\. Dunn, K\. A\. Persson, G\. Ceder, and A\. JainQuantifying the advantage of domain\-specific pre\-training on named entity recognition tasks in materials science\.Patterns3\(4\),pp\. 100488\.Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.4.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- Tshitoyanet al\.\(2019\)V\. Tshitoyan, J\. Dagdelen, L\. Weston, A\. Dunn, Z\. Rong, O\. Kononova, K\. A\. Persson, G\. Ceder, and A\. JainUnsupervised word embeddings capture latent knowledge from materials science literature\.Nature571\(7763\),pp\. 95–98\.Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.2.2),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. PolosukhinAttention is all you need\.Advances in neural information processing systems30\.Cited by:[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.13.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.14.2)\.
- Voorheeset al\.\(2021\)E\. Voorhees, T\. Alam, S\. Bedrick, D\. Demner\-Fushman, W\. R\. Hersh, K\. Lo, K\. Roberts, I\. Soboroff, and L\. L\. WangTREC\-covid: constructing a pandemic information retrieval test collection\.ACM SIGIR Forum54\(1\),pp\. 1–12\.Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px4.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Waddenet al\.\(2020\)D\. Wadden, S\. Lin, K\. Lo, L\. L\. Wang, M\. van Zuylen, A\. Cohan, and H\. HajishirziFact or fiction: verifying scientific claims\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 7534–7550\.Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Wanget al\.\(2023\)J\. Wang, K\. Wang, X\. Wang, P\. Naidu, L\. Bergen, and R\. PaturiDORIS\-MAE: scientific document retrieval using multi\-level aspect\-based queries\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px5.p1.1),[§C\.6](https://arxiv.org/html/2609.18262#A3.SS6.p1.1)\.
- Wanget al\.\(2022\)L\. Wang, N\. Yang, X\. Huang, B\. Jiao, L\. Yang, D\. Jiang, R\. Majumder, and F\. WeiText embeddings by weakly\-supervised contrastive pre\-training\.CoRRabs/2212\.03533\.Cited by:[7th item](https://arxiv.org/html/2609.18262#A2.I2.i7.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.9.1),[§3\.5](https://arxiv.org/html/2609.18262#S3.SS5.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px1.p1.1)\.
- Wanget al\.\(2024\)L\. Wang, N\. Yang, X\. Huang, L\. Yang, R\. Majumder, and F\. WeiImproving text embeddings with large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 11897–11916\.Cited by:[4th item](https://arxiv.org/html/2609.18262#A2.I4.i4.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.22.1),[§C\.5](https://arxiv.org/html/2609.18262#A3.SS5.SSS0.Px2.p1.1),[§C\.5](https://arxiv.org/html/2609.18262#A3.SS5.SSS0.Px3.p1.1),[§E\.1](https://arxiv.org/html/2609.18262#A5.SS1.p1.1),[§2](https://arxiv.org/html/2609.18262#S2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.
- Wanget al\.\(2020\)L\. L\. Wang, K\. Lo, Y\. Chandrasekhar, R\. Reas, J\. Yang, D\. Burdick, D\. Eide, K\. Funk, Y\. Katsis, R\. M\. Kinney, Y\. Li, Z\. Liu, W\. Merrill, P\. Mooney, D\. A\. Murdick, D\. Rishi, J\. Sheehan, Z\. Shen, B\. Stilson, A\. D\. Wade, K\. Wang, N\. X\. R\. Wang, C\. Wilhelm, B\. Xie, D\. M\. Raymond, D\. S\. Weld, O\. Etzioni, and S\. KohlmeierCORD\-19: the covid\-19 open research dataset\.InProceedings of the 1st Workshop on NLP for COVID\-19 at ACL 2020,Online\.External Links:[Link](https://www.aclweb.org/anthology/2020.nlpcovid19-acl.1)Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.6.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- Welblet al\.\(2017\)J\. Welbl, N\. F\. Liu, and M\. GardnerCrowdsourcing multiple choice science questions\.InProceedings of the 3rd Workshop on Noisy User\-generated Text \(W\-NUT\),pp\. 94–106\.Cited by:[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.Px7.p1.1),[§C\.6](https://arxiv.org/html/2609.18262#A3.SS6.p1.1)\.
- Wellawatteet al\.\(2025\)G\. P\. Wellawatte, H\. Guo, M\. Lederbauer, A\. S\. Borisova, M\. Hart, M\. Brucka, and P\. SchwallerChemLit\-qa: a human evaluated dataset for chemistry RAG tasks\.Mach\. Learn\. Sci\. Technol\.6\(2\),pp\. 20601\.External Links:[Link](https://doi.org/10.1088/2632-2153/adc2d6),[Document](https://dx.doi.org/10.1088/2632-2153/ADC2D6)Cited by:[§B\.2\.3](https://arxiv.org/html/2609.18262#A2.SS2.SSS3.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px1.p1.1)\.
- Welleret al\.\(2025\)O\. Weller, B\. V\. Durme, D\. J\. Lawrie, A\. Paranjape, Y\. Zhang, and J\. HesselPromptriever: instruction\-trained retrievers can be prompted like language models\.InThe Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24\-28, 2025,External Links:[Link](https://openreview.net/forum?id=odvSjn416y)Cited by:[7th item](https://arxiv.org/html/2609.18262#A2.I4.i7.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.25.1),[§C\.5](https://arxiv.org/html/2609.18262#A3.SS5.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px1.p2.1)\.
- Xionget al\.\(2024\)G\. Xiong, Q\. Jin, Z\. Lu, and A\. ZhangBenchmarking retrieval\-augmented generation for medicine\.arXiv preprint arXiv:2402\.13178\.Cited by:[Table 4](https://arxiv.org/html/2609.18262#A0.T4.2.1.7.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px3.p1.1)\.
- Xuet al\.\(2024\)R\. Xu, W\. Shi, Y\. Yu, Y\. Zhuang, Y\. Zhu, M\. D\. Wang, J\. C\. Ho, C\. Zhang, and C\. YangBmretriever: tuning large language models as better biomedical text retrievers\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 22234–22254\.Cited by:[9th item](https://arxiv.org/html/2609.18262#A2.I2.i9.p1.1),[4th item](https://arxiv.org/html/2609.18262#A2.I3.i4.p1.1),[6th item](https://arxiv.org/html/2609.18262#A2.I4.i6.p1.1),[§B\.2\.1](https://arxiv.org/html/2609.18262#A2.SS2.SSS1.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.11.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.17.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.24.1),[§C\.5](https://arxiv.org/html/2609.18262#A3.SS5.SSS0.Px2.p1.1),[Appendix D](https://arxiv.org/html/2609.18262#A4.p1.1),[§E\.1](https://arxiv.org/html/2609.18262#A5.SS1.p1.1),[Figure 1](https://arxiv.org/html/2609.18262#S1.F1.2.2.1.1.3),[§1](https://arxiv.org/html/2609.18262#S1.p2.1),[§2](https://arxiv.org/html/2609.18262#S2.p3.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px4.p1.1),[§4\.2](https://arxiv.org/html/2609.18262#S4.SS2.SSS0.Px2.p1.1)\.
- Yanget al\.\(2024\)A\. Yang, B\. Yang, B\. Hui, B\. Zheng, B\. Yu, C\. Zhou, C\. Li, C\. Li, D\. Liu, F\. Huang, G\. Dong, H\. Wei, H\. Lin, J\. Tang, J\. Wang, J\. Yang, J\. Tu, J\. Zhang, J\. Ma, J\. Xu, J\. Zhou, J\. Bai, J\. He, J\. Lin, K\. Dang, K\. Lu, K\. Chen, K\. Yang, M\. Li, M\. Xue, N\. Ni, P\. Zhang, P\. Wang, R\. Peng, R\. Men, R\. Gao, R\. Lin, S\. Wang, S\. Bai, S\. Tan, T\. Zhu, T\. Li, T\. Liu, W\. Ge, X\. Deng, X\. Zhou, X\. Ren, X\. Zhang, X\. Wei, X\. Ren, Y\. Fan, Y\. Yao, Y\. Zhang, Y\. Wan, Y\. Chu, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. FanQwen2 technical report\.arXiv preprint arXiv:2407\.10671\.Cited by:[§A\.2](https://arxiv.org/html/2609.18262#A1.SS2.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.26.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.27.2),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.28.2),[§3\.1](https://arxiv.org/html/2609.18262#S3.SS1.p1.1)\.
- Yuet al\.\(2022\)Y\. Yu, C\. Xiong, S\. Sun, C\. Zhang, and A\. OverwijkCoco\-dr: combating distribution shifts in zero\-shot dense retrieval with contrastive and distributionally robust learning\.arXiv preprint arXiv:2210\.15212\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p1.1)\.
- Zhanget al\.\(2024\)H\. Zhang, Y\. Song, Z\. Hou, S\. Miret, and B\. LiuHoneyComb: a flexible llm\-based agent system for materials science\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 3369–3382\.Cited by:[§1](https://arxiv.org/html/2609.18262#S1.p1.1)\.
- Zhanget al\.\(2025\)Y\. Zhang, S\. A\. Khan, A\. Mahmud, H\. Yang, A\. Lavin, M\. Levin, J\. Frey, J\. Dunnmon, J\. Evans, A\. Bundy,et al\.Exploring the role of large language models in the scientific method: from hypothesis to discovery\.npj Artificial Intelligence1\(1\),pp\. 14\.Cited by:[§2](https://arxiv.org/html/2609.18262#S2.p2.1)\.
- Zhanget al\.\(2023\)Y\. Zhang, H\. Cheng, Z\. Shen, X\. Liu, Y\. Wang, and J\. GaoPre\-training multi\-task contrastive learning models for scientific literature understanding\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 12259–12275\.Cited by:[4th item](https://arxiv.org/html/2609.18262#A2.I2.i4.p1.1),[Table 7](https://arxiv.org/html/2609.18262#A2.T7.2.1.6.1),[§1](https://arxiv.org/html/2609.18262#S1.p2.1),[§2](https://arxiv.org/html/2609.18262#S2.p2.1),[§4\.1](https://arxiv.org/html/2609.18262#S4.SS1.SSS0.Px2.p1.1)\.

DomainDatasetSizeLineMaterialMat2Vec[Tshitoyan et al\. \(2019\)](https://arxiv.org/html/2609.18262#bib.bib14)1\.5 Mhttps://github\.com/materialsintelligence/mat2vec/MatSciBERT[Gupta et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib5)0\.1 Mhttps://github\.com/M3RG\-IITD/MatSciBERTMatBERT[Trewartha et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib4)2 Mhttps://github\.com/lbnlp/MatBERTBioMedicalS2ORC[Lo et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib22)600Khttps://github\.com/allenai/s2orcMeadow[Wang et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib59)460khttps://huggingface\.co/datasets/medalpaca/medical\_meadow\_cord19Textbooks[Xiong et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib33)50Khttps://huggingface\.co/datasets/MedRAG/textbooksLitCovid[Chen et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib29)70Khttps://huggingface\.co/datasets/KushT/LitCovid\_BioCreative

Table 4:Statistics of the public scientific corpora used for model initialization, categorized by domain\.## Appendix ADetails of Implementation and Setup

### A\.1Initial Corpus Construction

We prioritize data quality and domain breadth over sheer scale\. Unlike standard baselines that rely on massive, noisy web\-crawled corpora, we constructed a compact yet highly diverse dataset spanning a wider range of scientific disciplines, specifically integrating large\-scale biomedical benchmarks with our newly constructed materials science data \(see Table[4](https://arxiv.org/html/2609.18262#A0.T4)\)\. Although smaller in total volume compared to general\-domain pre\-training corpora, this curated mixture undergoes rigorous cleaning to ensure superior density of scientific information\.

For the materials science domain, which specifically lacks unified public resources, we crawled documents via DOIs and addressed the substantial inconsistency in notation \(e\.g\.,α\\alpha\-Fe2O3vs\.alpha\-Fe2O3\)\. We applied a materials\-aware normalization pipeline adapted from the MatSciBERT framework, including NFKC normalization, HTML entity mapping, and chemical formula hyphen reconnection\. Crucially, we deliberately excluded standard normalization steps that would destroy materials\-specific semantics, such as replacing numbers with placeholders or normalizing stoichiometric formulas \(e\.g\.,Ni0\.5Fe0\.5→\\toFeNi\)\.

Finally, we maximized data efficiency through strict quality filtering and consistent instruction formatting\. We removed entries with missing metadata, as well as those exceeding context limits or lacking sufficient information\. To leverage the instruction\-following capabilities of the base model, we format every queryqqwith a specific task instruction:

‘‘Given a query, retrieve passages that are relevant to the query\.\\nQuery: \{text\} \{eos\}’’

This results in a refined corpus that is surface\-consistent yet semantically precise, enabling the model to learn robust scientific representations from a smaller but more potent dataset\.

### A\.2Details of Implementation

All models are trained using PyTorch with Distributed Data Parallel \(DDP\) on two NVIDIA H200 GPUs\. We adopt Qwen2\.5\-0\.5B, Qwen2\.5\-1\.5B, and Qwen2\.5\-7B[Yang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib37)as backbone encoders, initialized from publicly released checkpoints\. Training is performed inbfloat16precision with gradient checkpointing enabled to reduce memory consumption\. We apply parameter\-efficient fine\-tuning with LoRA[Hu et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib73), using rankr=16r=16, scaling factorα=32\\alpha=32, and dropout rate 0\.05, and update only the LoRA parameters while keeping the backbone frozen\. Optimization is carried out using AdamW[Loshchilov and Hutter \(2019\)](https://arxiv.org/html/2609.18262#bib.bib71), with a learning rate of2×10−52\\times 10^\{\-5\}for the 7B model, and training proceeds for two epochs with a global batch size of 256 across GPUs\. Input queries and passages are tokenized with a maximum sequence length of 512 and encoded using an EOS\-based last\-token pooling strategy to obtain fixed\-dimensional representations\. The retriever is trained with an InfoNCE contrastive objective, leveraging in\-batch negatives as well as cross\-device negatives enabled by DDP synchronization\. Model checkpoints are saved periodically during training, and all hyperparameters are fixed across runs unless otherwise specified\.

#### A\.2\.1Verified Query Generation

To generate challenging queries in*Expansion*\(§[3\.4](https://arxiv.org/html/2609.18262#S3.SS4)\), from the detected long\-tail scientific concepts, we employ a prompt\-based query generation strategy\. Given a target concept identified during self\-diagnosis, we instruct a large language model to produce a single, specific research\-oriented query grounded in materials science\. The prompt template used for query generation is shown below\.

Your task is to generate a single research query that is inherently ambiguous or difficult to resolve for standard retrieval models\. The query should sound like a real paper title or research question — using indirect, context\-dependent, or metaphorical language — so that a retrieval model cannot easily find the answer without deep understanding\. Concept 1:\{Concept\_1\} \(Attributes: \{Attributes\_1\}\) Concept 2:\{Concept\_2\} \(Attributes: \{Attributes\_2\}\) Concept 3:\{Concept\_3\} \(Attributes: \{Attributes\_3\}\)Query:

### A\.3Statistical Characterization of the Scientific Long Tail

Figure[1](https://arxiv.org/html/2609.18262#S1.F1)\(b\) contrasts a scientific corpus with MS MARCO\. To check that this contrast does not depend on a single reference corpus, we compare five frequency distributions: scientific concepts, the science corpus as a whole, the general words inside that corpus, and two independent general\-domain corpora, MS MARCO[Bajaj et al\. \(2016\)](https://arxiv.org/html/2609.18262#bib.bib34)and WikiText\-103[Merity et al\. \(2017\)](https://arxiv.org/html/2609.18262#bib.bib82)\. The same regular\-expression tokenizer is applied to all five so that the counts are comparable\.

Table[5](https://arxiv.org/html/2609.18262#A1.T5)reports the rank\-frequency statistics\. The tail of the scientific concepts is more than twice as flat as either general\-domain corpus, with a log\-log Zipf slope of0\.820\.82against1\.821\.82for MS MARCO and1\.761\.76for WikiText\-103\. The gap is even clearer in how rare the terms are:62\.4%62\.4\\%of scientific concepts appear exactly once and93\.9%93\.9\\%appear five times or fewer in a 36\.2M\-token corpus, against3737–40%40\\%and6767–69%69\\%for the two general\-domain corpora\. In other words, there is almost no dense head from which a retriever could learn these concepts, which is why resampling the training data internally cannot fix the problem and why the Expansion stage draws evidence from outside the corpus\.

Table[6](https://arxiv.org/html/2609.18262#A1.T6)tests the separation directly\. Against both general\-domain corpora the two\-sample Kolmogorov\-Smirnov distance is0\.290\.29–0\.310\.31with a p\-value below10−30010^\{\-300\}\. The control comparison, scientific concepts against the science corpus they are drawn from, is far smaller atD=0\.06D=0\.06\. The separation is therefore between the scientific and general domains, not between two samples of the same corpus\.

DistributionZipf slopeZipfα\\alphaHapaxfreq≤5\\leq 5freq≤10\\leq 10ss\(R2\)\(MLE\)\(%\)\(%\)\(%\)Science Concepts0\.823 \(0\.904\)1\.86262\.4393\.8796\.52Science Corpus \(overall\)1\.097 \(0\.922\)1\.75156\.5189\.7293\.49General Words \(in\-science\)1\.446 \(0\.962\)1\.60645\.1582\.1287\.90MS MARCO1\.818 \(0\.981\)1\.45537\.0166\.7975\.13WikiText\-1031\.761 \(0\.983\)1\.47639\.5568\.6277\.26

Table 5:Rank\-frequency statistics of five distributions under an identical tokenizer\. A smaller Zipf slopessmeans a heavier tail, and Hapax is the share of types that occur exactly once\.Table 6:Distributional distance from Science Concepts, measured inlog10\\log\_\{10\}frequency space\. The last row compares the scientific concepts with the corpus they are drawn from and serves as a within\-domain control\.

## Appendix BDetails of Evaluation

### B\.1Baselines for Retrieval Tasks

In this section, we provide detailed descriptions of the baseline models used in our experiments\. A comprehensive summary of their architectural characteristics and methodological components, alongside our proposed REPAIR framework, is provided in Table[7](https://arxiv.org/html/2609.18262#A2.T7)\.

##### Sparse Retrieval Models\.

Sparse retrieval approaches estimate relevance by matching keywords between queries and documents\.

- •BM25[Robertson and Zaragoza \(2009\)](https://arxiv.org/html/2609.18262#bib.bib76)serves as the standard probabilistic baseline for lexical retrieval\. It utilizes a term\-frequency inverse\-document\-frequency \(TF\-IDF\) based scoring function to compute similarity scores between high\-dimensional sparse vectors, effectively weighting term importance\.

##### Dense Retrieval Models\.

Dense retrieval models encode queries and documents into continuous vector spaces to capture semantic relationships\. We evaluate models across three distinct scales:

- •Contriever[Izacard et al\. \(2021b\)](https://arxiv.org/html/2609.18262#bib.bib43)is a dual\-encoder model \(110M\) trained via unsupervised contrastive learning\. It leverages a massive corpus comprising data from Wikipedia and CC\-Net to learn robust representations without labeled supervision\.
- •Dragon[Lin et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib74)is a BERT\-base scale model \(110M\) that adopts a progressive training strategy\. It utilizes diverse supervision signals and data augmentation techniques to enhance general retrieval capabilities\.
- •SPECTER 2\.0[Singh et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib70)is specifically tailored for scientific document representation \(110M\)\. It employs a multi\-task training objective that covers various scientific tasks, allowing the model to generate embeddings adaptable to different formats and downstream applications\.
- •SciMult[Zhang et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib52)is a domain\-specialized retriever \(110M\) for scientific literature\. It integrates instruction tuning within a multi\-task contrastive learning framework to better align representations with scientific query intents\.
- •MedCPT[Jin et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib12)focuses on biomedical information retrieval \(220M\)\. Its representations are learned from a large\-scale dataset of 255 million user search logs from PubMed, effectively capturing the semantics of medical queries and documents\.
- •InstructOR\-L[Su et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib46)is an instruction\-finetuned model \(335M\) capable of generating task\-specific embeddings\. By conditioning on natural language instructions, it adapts to diverse domains without further fine\-tuning\.
- •E5\-Large\-v2[Wang et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib36)employs a two\-stage training pipeline \(335M\): initial contrastive pre\-training on weakly labeled text pairs followed by supervised fine\-tuning on high\-quality datasets with mined hard negatives\.
- •BGE\-Large[Chen et al\. \(2024b\)](https://arxiv.org/html/2609.18262#bib.bib60)is a strong baseline \(335M\) trained with a multi\-stage approach similar to E5 but enhanced by improved negative sampling and a diverse training mixture\.
- •BMRetriever\-410M[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)is the compact variant of a retrieval family tailored for biology and medicine\. It is pre\-trained on extensive domain\-specific corpora and subsequently fine\-tuned using augmented data synthesized by Large Language Models\.
- •DRAMA\-L[Ma et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib62)represents a lightweight baseline designed for efficient retrieval, balancing performance with computational constraints\.

MethodBackboneScaleDomainBM25\-\-General✗✗✗✗✗✗Contriever \([2021b](https://arxiv.org/html/2609.18262#bib.bib43)\)BERT\-base \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)110MGeneral✓✓✓✗✗✗Dragon \([2023](https://arxiv.org/html/2609.18262#bib.bib74)\)BERT\-base \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)110MGeneral✗✓✓✗✗✗SPECTER 2\.0 \([2023](https://arxiv.org/html/2609.18262#bib.bib70)\)SciBERT \([2019](https://arxiv.org/html/2609.18262#bib.bib44)\)110MScientific✓✗✓✗✗✗SciMult \([2023](https://arxiv.org/html/2609.18262#bib.bib52)\)PubMedBERT \([2021](https://arxiv.org/html/2609.18262#bib.bib6)\)110MScientific✓✗✓✗✗✗MedCPT \([2023](https://arxiv.org/html/2609.18262#bib.bib12)\)PubMedBERT \([2021](https://arxiv.org/html/2609.18262#bib.bib6)\)220MBiomedical✗✓✓✗✗✗InstructOR\-L \([2023](https://arxiv.org/html/2609.18262#bib.bib46)\)GTR\-Large \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)335MGeneral✗✓✓✗✗✗E5\-Large\-v2† \([2022](https://arxiv.org/html/2609.18262#bib.bib36)\)BERT\-large \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)660MGeneral✓✓✓✗✗✗BGE\-Large‡ \([2024b](https://arxiv.org/html/2609.18262#bib.bib60)\)RoBERTa\-large \([2020](https://arxiv.org/html/2609.18262#bib.bib63)\)895MGeneral✓✓✓✗✗✗BMRetriever\-410M \([2024](https://arxiv.org/html/2609.18262#bib.bib47)\)Pythia\-410M \([2023](https://arxiv.org/html/2609.18262#bib.bib61)\)410MBiomedical✓✓✓✗✗✗InstructOR\-XL \([2023](https://arxiv.org/html/2609.18262#bib.bib46)\)GTR\-XL \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)1\.5BGeneral✗✓✓✗✗✗GTR\-XL \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)T5\-XL \([2017](https://arxiv.org/html/2609.18262#bib.bib30)\)1\.2BGeneral✓✓✓✗✗✗GTR\-XXL \([2022](https://arxiv.org/html/2609.18262#bib.bib57)\)T5\-XXL \([2017](https://arxiv.org/html/2609.18262#bib.bib30)\)4\.8BGeneral✓✓✓✗✗✗SGPT\-1\.3B \([2022](https://arxiv.org/html/2609.18262#bib.bib31)\)GPT\-Neo \([2023](https://arxiv.org/html/2609.18262#bib.bib32)\)1\.3BGeneralunk✓✗✗✗✗SGPT\-2\.7B \([2022](https://arxiv.org/html/2609.18262#bib.bib31)\)GPT\-Neo \([2023](https://arxiv.org/html/2609.18262#bib.bib32)\)2\.7BGeneralunk✓✗✗✗✗BMRetriever\-2B \([2024](https://arxiv.org/html/2609.18262#bib.bib47)\)Gemma \([2024](https://arxiv.org/html/2609.18262#bib.bib17)\)2BBiomedical✓✓✓✗✗✗DRAMA\-1B \([2025](https://arxiv.org/html/2609.18262#bib.bib62)\)LLaMA\-3\.2\-1B1BGeneral✗✓✓✗✗✗Llama2Vec \([2024](https://arxiv.org/html/2609.18262#bib.bib66)\)LLaMA\-2\-7B7BGeneral✓✓✓✗✗✗RepLLaMA \([2024](https://arxiv.org/html/2609.18262#bib.bib65)\)LLaMA\-2\-7B7BGeneral✗✓✓✗✗✗LLM2Vec \([2024](https://arxiv.org/html/2609.18262#bib.bib38)\)Mistral\-7B7BGeneral✓✓✓✗✗✗E5\-Mistral \([2024](https://arxiv.org/html/2609.18262#bib.bib58)\)Mistral\-7B7BGeneral✗✓✓✗✗✗CPT\-text\-XL \([2022](https://arxiv.org/html/2609.18262#bib.bib13)\)GPT \([2020](https://arxiv.org/html/2609.18262#bib.bib26)\)175BGeneralunkunk✗✗✗✗BMRetriever\-7B \([2024](https://arxiv.org/html/2609.18262#bib.bib47)\)BioMistral \([2024](https://arxiv.org/html/2609.18262#bib.bib72)\)7BBiomedical✓✓✓✗✗✗Promptriever \([2025](https://arxiv.org/html/2609.18262#bib.bib64)\)llama2\-7b \([2023](https://arxiv.org/html/2609.18262#bib.bib42)\)7BGeneral✓✓✓✗✗✗REPAIR\-500M \(ours\)Qwen2\.5\-0\.5B \([2024](https://arxiv.org/html/2609.18262#bib.bib37)\)500MScientific✓✓✓✓✓✓REPAIR\-1\.5B \(ours\)Qwen2\.5\-1\.5B \([2024](https://arxiv.org/html/2609.18262#bib.bib37)\)1\.5BScientific✓✓✓✓✓✓REPAIR\-7B \(ours\)Qwen2\.5\-7B \([2024](https://arxiv.org/html/2609.18262#bib.bib37)\)7BScientific✓✓✓✓✓✓

Table 7:Comprehensive comparison of baseline retrieval models and the proposed REPAIR framework\. The table delineates backbone architectures, model scales, target domains, and specific training methodologies\. Methodological components are abbreviated as follows: Contra Pretrain\. \(Contrastive Pretraining\), Data Aug\. \(Data Augmentation\), Hard Neg\. \(Hard Negative Mining\), Iter Refine\. \(Iterative Refinement\), Confus Diag\. \(Confusion Diagnosis\), and Fact Verif\. \(Factual Verification\)\.- •InstructOR\-XL[Su et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib46)scales the instruction\-based training methodology to 1\.5B parameters, offering improved generalization and instruction\-following capabilities compared to its smaller counterpart\.
- •GTR\-XL / GTR\-XXL[Ni et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib57)are Generalizable T5\-based Retrievers\. Initialized from T5, they undergo pre\-training on community QA pairs followed by fine\-tuning on NQ and MS MARCO\. We report results for the 1\.2B and 4\.8B variants\.
- •SGPT\-1\.3B / SGPT\-2\.7B[Muennighoff \(2022\)](https://arxiv.org/html/2609.18262#bib.bib31)adapt decoder\-only GPT architectures for symmetric search\. By freezing the backbone and fine\-tuning only the bias tensors and position\-weighted pooling layers, they transform generative models into effective retrievers\.
- •BMRetriever\-2B[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)scales the biomedical\-focused architecture to 2 billion parameters, allowing for deeper semantic understanding of scientific texts\.
- •DRAMA\-1B[Ma et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib62)is the billion\-scale iteration of the DRAMA series, providing a middle\-ground baseline between efficiency and capacity\.

- •Llama2Vec[Li et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib66)converts LLaMA\-7B into a retriever using two novel pre\-training tasks: Embedding\-Based Auto\-Encoding \(EBAE\) and Embedding\-Based Next Sentence Prediction \(EBNSP\)\.
- •RepLLaMA[Ma et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib65)performs full fine\-tuning of the LLaMA\-7B model on MS MARCO, directly optimizing the generative backbone for passage retrieval tasks\.
- •LLM2Vec[BehnamGhader et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib38)enables bidirectional attention in causal LLMs through masked next\-token prediction\. This unsupervised approach transforms standard LLMs into powerful text encoders\.
- •E5\-Mistral[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib58)initializes from Mistral\-7B and is trained with a wide variety of synthetic data generated by LLMs, achieving state\-of\-the\-art performance on the MTEB benchmark\.
- •CPT\-text\-XL[Neelakantan et al\. \(2022\)](https://arxiv.org/html/2609.18262#bib.bib13)is a web\-scale contrastive model \(175B\)\. We include it as a reference point for performance achievable with massive\-scale pre\-training, rather than a direct comparison due to its size\.
- •BMRetriever\-7B[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)is the largest model in its series, leveraging 7 billion parameters to maximize retrieval accuracy in specialized scientific domains\.
- •Promptriever[Weller et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib64)is a bi\-encoder retrieval model initialized from an LLM backbone\. Unlike standard retrievers, it is fine\-tuned on a massive dataset of MS MARCO pairs augmented with instance\-level instructions and “instruction negatives, enabling it to follow complex, per\-instance natural language prompts to dynamically adjust relevance criteria without further training\.

### B\.2Evaluation Task and Dataset

In this section, we provide detailed descriptions of the datasets employed in our experiments\. We categorize these benchmarks into five primary retrieval\-oriented groups: Information Retrieval \(IR\), Sentence Similarity, Question Answering \(QA\), Entity Linking, and Paper Recommendation\.

#### B\.2\.1Information Retrieval

Following prior work[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47), we evaluate passage retrieval performance in scientific and biomedical domains using four datasets from the BEIR benchmark[Boteva et al\. \(2016\)](https://arxiv.org/html/2609.18262#bib.bib67)\. These benchmarks require retrieving relevant passages from corpora containing complex, terminology\-intensive documents\.

##### NFCorpus

[Boteva et al\. \(2016\)](https://arxiv.org/html/2609.18262#bib.bib67): A biomedical information retrieval dataset consisting of 323 natural\-language queries related to nutrition facts, evaluated over a corpus of approximately 3\.6K PubMed documents\. The task is formulated as document retrieval, where models are given a question and are required to retrieve documents that best answer the query\.

##### SciFact

[Wadden et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib68): A scientific fact\-verification dataset comprising 300 queries, where the task is to retrieve abstracts that provide supporting or refuting evidence for a given scientific claim\. The corpus consists of approximately 5K scientific papers\.

##### SciDocs

[Cohan et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib53): A citation\-oriented retrieval dataset consisting of 1K queries derived from scientific paper titles, evaluated over a corpus of 25K scientific articles\. The task requires retrieving abstracts of papers that are cited by the given paper\.

##### TREC\-COVID

[Voorhees et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib69): A biomedical information retrieval dataset focused on COVID\-19\-related literature, comprising 50 queries evaluated over a corpus of approximately 171K documents\. Each query is associated with a dense set of relevant documents, averaging 493\.5 per query, and the task requires retrieving documents that answer the given COVID\-19 query\.

We additionally evaluate on three scientific benchmarks that lie outside the nine used in the main experiments, in order to probe generalization to unseen task formats and to scientific subareas beyond materials science and biomedicine \(Appendix[C\.6](https://arxiv.org/html/2609.18262#A3.SS6)\)\. None of the three is used at any point during training\.

##### DORIS\-MAE

[Wang et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib78): A multidisciplinary scientific document retrieval dataset built from computer science literature, in which each of the 100 queries is a multi\-sentence research summary decomposed into several aspects\. Relevance is graded over a corpus of 8,591 abstracts, and the multi\-aspect query format differs markedly from the single\-intent queries of the four benchmarks above\.

##### CQA\-physics

[Hoogeveen et al\. \(2015\)](https://arxiv.org/html/2609.18262#bib.bib79);[Thakur et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib80): The physics subforum of CQADupStack as distributed in BEIR, consisting of 1,039 community question\-answering queries over a corpus of 38,316 posts\. The task requires retrieving duplicate or answer\-bearing posts, and its informal, user\-written style contrasts with the formal scientific prose of the other benchmarks\.

##### SciQ

[Welbl et al\. \(2017\)](https://arxiv.org/html/2609.18262#bib.bib81): A broad\-coverage science QA dataset spanning physics, chemistry, biology, and earth science\. We cast it as retrieval by pairing each of the 884 test questions with the supporting passage that contains its answer, over a corpus of 12,241 deduplicated support passages\.

Table 8:Comparison of retrieval performance across iterations and model scales\. The highlighted row marks our default setting \(Iter 2\)\. Beyond it, the average nDCG@10 improves by at most\+0\.005\+0\.005at any scale\.

#### B\.2\.2Sentence Similarity\.

For sentence\-level retrieval, we employBIOSSES[Soğancıoğlu et al\. \(2017\)](https://arxiv.org/html/2609.18262#bib.bib39), a biomedical sentence similarity dataset consisting of 100 sentence pairs extracted from PubMed articles\. Each pair is annotated by human experts with a similarity score on a 5\-point scale, ranging from 0 \(no semantic relation\) to 4 \(semantically equivalent\)\. The task is formulated as sentence retrieval, where models are given a sentence and are required to retrieve sentences with the same meaning\.

#### B\.2\.3Question Answering\.

We extend our evaluation to retrieval\-augmented downstream tasks using three QA datasets:

##### iCliniq

[Chen et al\. \(2020\)](https://arxiv.org/html/2609.18262#bib.bib40): A biomedical conversational question answering dataset constructed from patient–clinician interactions collected from a public health forum, comprising approximately 7\.3K questions and 7\.3K responses\. The task is formulated as retrieval\-based QA, where models are given a question with conversational context and are required to retrieve responses that best answer the query\.

##### ChemLit\-QA

[Wellawatte et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib41): A literature\-based scientific QA and Retrieval\-Augmented Generation \(RAG\) benchmark\. This dataset evaluates the model’s ability to generate faithful and precise answers based on chemical literature contexts\. It was rigorously validated by four experts with backgrounds in chemistry and chemical engineering\. For our experiments, we specifically utilized the subsets categorized under biomedical and material science domains to align with our target tasks\.

#### B\.2\.4Entity Linking\.

To assess the model’s capability in identifying and linking domain\-specific concepts, we useMeSH[Lipscomb \(2000\)](https://arxiv.org/html/2609.18262#bib.bib25), a biomedical entity linking benchmark designed to evaluate the identification and normalization of domain\-specific concepts\. The dataset comprises approximately 29\.6K biomedical concepts and corresponding textual entries from the Medical Subject Headings \(MeSH\) thesaurus\. The task is formulated as retrieval\-based entity linking, where models are given a biomedical concept mention and are required to retrieve passages that define or correspond to the correct MeSH concept\.

#### B\.2\.5Paper Recommendation\.

We evaluate retrieval performance on a paper recommendation task using theRELISHdataset[Singh et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib70);[Brown et al\. \(2019\)](https://arxiv.org/html/2609.18262#bib.bib23)\. The benchmark consists of approximately 3\.2K query articles and a corpus of 191\.2K PubMed abstracts\. The task requires retrieving literature relevant to a given article, with relevance annotated using graded similarity scores ranging from 0 \(not similar\) to 2 \(highly similar\)\.

## Appendix CDetails of Ablation Studies and Analyses

This section provides comprehensive experimental details and extended results for ablation studies introduced in §[4\.3](https://arxiv.org/html/2609.18262#S4.SS3)\. Specifically, we further investigate the individual contributions of key design choices inREPAIRby presenting detailed analyses on the iterative refinement process \(§[C\.1](https://arxiv.org/html/2609.18262#A3.SS1)\), the low\-margin query selection ratio \(§[C\.2](https://arxiv.org/html/2609.18262#A3.SS2)\), and the impact of the number of analyzed negativeskk\(§[C\.3](https://arxiv.org/html/2609.18262#A3.SS3)\)\.

### C\.1Detailed Analysis of Iterative Refinement

Table[8](https://arxiv.org/html/2609.18262#A2.T8)illustrates the performance trajectory across iterations\. The primary driver of these gains is the resolution of long\-tail concept confusion rather than inherent model capacity\. To isolate this effect, we compare the same Qwen2\.5 backbones trained on our refined data versus a strong baseline \(BMRetriever\)\. Across all parameter scales, models trained withREPAIRconsistently outperform those trained on baseline datasets, proving that fact\-verified data quality outweighs backbone size\.

The efficacy of this refinement is further evidenced by the representational margin shift\. For the 173 "persistent queries" that remained in the confusion set after Iteration 1, the average margin shifted from−2\.7×10−3\-2\.7\\times 10^\{\-3\}to\+3\.5×10−3\+3\.5\\times 10^\{\-3\}in Iteration 2\. This positive shift indicates that the iterative process successfully expands the model’s embedding space to distinguish fine\-grained scientific concepts that were previously collapsed\. Consequently, the refinement process ensures that the model’s improvements are grounded in factual differentiation rather than biased stagnation\.

Table 9:Number of unique long\-tail concepts extracted across different query selection margins \(pp\)\. The Total Unique \(A∪BA\\cup B\) shows the footprint of epistemic uncertainty captured by the diagnosis stage\. The marginal increase \(Δ\\Delta\) significantly drops afterp=40%p=40\\%, indicating diminishing returns\. Selecting beyond this threshold primarily introduces well\-resolved concepts that act as noise during the data expansion stage\.Table 10:Detailed retrieval performance \(nDCG@10\) across five target datasets at varying low\-margin query selection ratios \(pp\)\. The highlighted row \(p=40%p=40\\%\) indicates the optimal threshold that provides a strong balance between performance and concept efficiency\.
### C\.2Detail Analysis of Low\-Margin Query Selection Ratio

To evaluate the effectiveness of our margin\-based selection in concentrating diagnostic signals for long\-tail errors, we tests the selection ratioppbased onΔθ​\(q\)\\Delta\_\{\\theta\}\(q\)\. In conjunction with this visual summary, Table[9](https://arxiv.org/html/2609.18262#A3.T9)and Table[10](https://arxiv.org/html/2609.18262#A3.T10)provide the complete empirical results supporting our choice to fixp=40%p=40\\%\.

##### Concept Extraction Scale and Diminishing Returns\.

Table[9](https://arxiv.org/html/2609.18262#A3.T9)details the number of unique long\-tail concepts extracted via Path A and Path B as the selection ratioppincreases from5%5\\%to100%100\\%\. The total number of unique concepts \(A∪BA\\cup B\) demonstrates rapid initial growth\. However, the marginal increase \(Δ\\Delta\) column clearly illustrates a point of diminishing returns\. Up top=40%p=40\\%, the diagnosis stage efficiently extracts422,183422,183unique concepts\. Beyond this threshold, increasing the ratio requires processing a significantly larger volume of queries, but the marginal discovery of novel concepts drops\. This indicates that queries above the 40th percentile of the retrieval marginΔθ​\(q\)\\Delta\_\{\\theta\}\(q\)are largely well\-resolved by the base retriever and contribute little to the footprint of epistemic uncertainty\.

##### Downstream Retrieval Performance\.

Table[10](https://arxiv.org/html/2609.18262#A3.T10)reports the exact nDCG@10 scores across the five individual target datasets \(NFCorpus, SciFact, SciDocs, COVID, BIOSSES\)\. The average nDCG@10 score rises steadily from0\.5300\.530atp=10%p=10\\%to0\.5470\.547atp=40%p=40\\%\. Beyondp=40%p=40\\%, the performance exhibits a clear saturation effect\. While processing100%100\\%of the queries yields the absolute maximum average of0\.5540\.554, the gain fromp=40%p=40\\%is minimal \(\+0\.007\+0\.007\)\. Given the substantial computational cost of the data expansion stage, introducing the remaining60%60\\%of queries primarily acts as noise\. Therefore,p=40%p=40\\%provides an optimal balance, maximizing diagnostic value while maintaining high retrieval accuracy\.

### C\.3Detailed Analysis of the Number of Analyzed Negatives \(kk\)

##### Setup and Motivation\.

The core strength of the REPAIR framework lies in its ability to precisely isolate long\-tail confusions without being polluted by semantic noise or irrelevant distractors\. During the diagnosis stage, identifying the optimal number of analyzed top\-ranked negatives \(kk\) is critical: inspecting too few negatives might fail to capture systemic error patterns, while inspecting too many risks introducing semantic drift that degrades the factual fidelity of the extracted concepts\.

To systematically justify the optimal boundary ofk=30k=30, we evaluate the neighborhood stability and negative hardness employing the 0\.5B retriever at the initial iteration\. For each confused queryq∈𝒬confq\\in\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}, we construct a ranked list of negatives𝒩k​\(q\)=\{d^1,…,d^k\}\\mathcal\{N\}\_\{k\}\(q\)=\\\{\\hat\{d\}\_\{1\},\\ldots,\\hat\{d\}\_\{k\}\\\}using the current encoder while strictly excluding the paired positive documentd\+d^\{\+\}\. By fixing the confusion selection ratio atp=40%p=40\\%, we obtain confused queries and subsequently sweep the parameterkkacross the set\{5,10,15,…,100\}\\\{5,10,15,\\dots,100\\\}\. All measurements utilize cosine similarities between L2\-normalized embeddings derived via end\-of\-sequence last\-token pooling, which are efficiently computed through a cached top\-100 retrieval matrix\.

#### C\.3\.1Quantitative Analysis

The core objective of the REPAIR framework is to accurately diagnose the model’s vulnerabilities by exposing it to genuine hard negatives, that is, documents that are highly confusable with the true positive\. However, determining the optimal number of analyzed negatives \(kk\) presents a critical trade\-off\. Inspecting too few candidates provides an insufficient signal to capture the model’s precise confusion boundary\. Conversely, expanding the pool too broadly risks diluting the diagnostic process with easily distinguishable, out\-of\-domain noise that distorts the semantic focus\.

To systematically justify our selection ofk=30k=30, we analyze the empirical results across six distinct metrics \(visualized in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\)\. Rather than relying on arbitrary thresholds, this analysis demonstrates howk=30k=30provides an effective balance between maximizing diagnostic yield and mitigating semantic drift\. For baseline comparisons, we define𝒩30​\(q\)\\mathcal\{N\}\_\{30\}\(q\)as the reference negative set\.

##### Average Negative Similarity\.

To quantify how effectively the extracted distractors capture genuine confusion, we measure the average negative similarity\. This metric reflects the overall difficulty of the negative pool, a higher value indicates that the retrieved documents remain highly competitive and structurally close to the query\.

Aavg​\(k\)=1\|𝒬conf\|​∑q∈𝒬conf1k​∑d∈𝒩k​\(q\)sθ​\(q,d\)A\_\{\\mathrm\{avg\}\}\(k\)=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\\in\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\}\\frac\{1\}\{k\}\\sum\_\{d\\in\\mathcal\{N\}\_\{k\}\(q\)\}s\_\{\\theta\}\(q,d\)\(9\)
As illustrated in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\(a\), the pool maintains a high level of hardness up tok=30k=30\. Beyond this threshold, the similarity sharply declines, demonstrating that larger pools dilute the diagnostic quality with easily distinguishable documents\. This validatesk=30k=30as the optimal boundary for preserving concentrated hardness\.

##### Marginal Negative Similarity\.

While the overall average similarity demonstrates general pool hardness, it can mask the diminishing quality of documents added at lower ranks\. To isolate the exact diagnostic value of incrementally expanding the negative pool, we measure the marginal negative similarity\. This metric specifically tracks the average similarity of the newly added documents between consecutive boundskprev<kk\_\{\\mathrm\{prev\}\}<k:

μk=1\|𝒬conf\|​∑q1k−kprev​∑j=kprev\+1ksθ​\(q,d^j\)\\mu\_\{k\}=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\}\\frac\{1\}\{k\-k\_\{\\mathrm\{prev\}\}\}\\sum\_\{j=k\_\{\\mathrm\{prev\}\}\+1\}^\{k\}s\_\{\\theta\}\(q,\\hat\{d\}\_\{j\}\)\(10\)
By comparingμk\\mu\_\{k\}to the mean positive similarity \(s¯\+=1\|𝒬conf\|​∑qsθ​\(q,dq\+\)\\bar\{s\}^\{\+\}=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\}s\_\{\\theta\}\(q,d\_\{q\}^\{\+\}\)\), we intuitively determine whether the freshly incorporated negatives are actually harder than the true positive\. As illustrated in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\(b\), oncekkexceeds 30,μk\\mu\_\{k\}drops significantly below the positive baseline\. This confirms that documents ranked beyond 30 are, on average, easier for the model to distinguish than the true positive itself\. Because they offer no meaningful diagnostic value, restricting the expansion tok=30k=30is strictly justified\.

##### Concept Drift\.

To determine whether expanding the negative pool inadvertently introduces semantic noise, we quantify concept drift using the Jaccard distance relative to thek=30k=30reference set\. This metric intuitively evaluates neighborhood stability, a value approaching 0 signifies strong alignment with the target semantic neighborhood, whereas higher values indicate substantial deviation\.

Bdrift​\(k\)=1\|𝒬conf\|​∑q\(1−\|𝒩k​\(q\)∩𝒩30​\(q\)\|\|𝒩k​\(q\)∪𝒩30​\(q\)\|\)B\_\{\\mathrm\{drift\}\}\(k\)=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\}\\left\(1\-\\frac\{\|\\mathcal\{N\}\_\{k\}\(q\)\\cap\\mathcal\{N\}\_\{30\}\(q\)\|\}\{\|\\mathcal\{N\}\_\{k\}\(q\)\\cup\\mathcal\{N\}\_\{30\}\(q\)\|\}\\right\)\(11\)
As demonstrated in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\(c\), the concept drift remains remarkably constrained up tok=30k=30but escalates rapidly thereafter\. This sharp increase indicates that enlargingkkbeyond 30 progressively pulls negatives from entirely different semantic neighborhoods, which compromises the precision of the diagnostic pool\. Consequently, these results establishk=30k=30as the critical limit for maintaining semantic stability\.

##### Hard\-Negative Ratio\.

To assess the concentration of high\-quality distractors within the pool, we calculate the hard\-negative ratio\. By defining a strict hardness thresholdτh=Aavg​\(5\)\\tau\_\{h\}=A\_\{\\mathrm\{avg\}\}\(5\), this metric intuitively quantifies pool dilution; a lowerρk\\rho\_\{k\}implies that the retrieval space is saturated with easily distinguishable, non\-informative documents\.

ρk=1\|𝒬conf\|∑q1k∑d∈𝒩k​\(q\)𝕀\[sθ\(q,d\)≥τh\]\\rho\_\{k\}=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\}\\frac\{1\}\{k\}\\sum\_\{d\\in\\mathcal\{N\}\_\{k\}\(q\)\}\\mathbb\{I\}\[s\_\{\\theta\}\(q,d\)\\geq\\tau\_\{h\}\]\(12\)
As depicted in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\(d\), maintainingk=30k=30preserves a dense fraction of effective distractors\. Expanding the pool beyond this boundary results in severe dilution, establishingk=30k=30as the strict limit for maintaining the diagnostic quality of the negative set\.

##### Similarity Spread\.

To observe the heterogeneity of the analyzed documents, we measure the similarity spread by computing the within\-query standard deviation \(σk\\sigma\_\{k\}\) of the negative similarities\. An increasing trend visually indicates a mixed pool of hard and easy documents rather than a dense, confusable cluster\. Lets¯k​\(q\)\\bar\{s\}\_\{k\}\(q\)be the average negative similarity for queryqqwithin the top\-kkpool, defined as1k​∑d′∈𝒩k​\(q\)sθ​\(q,d′\)\\frac\{1\}\{k\}\\sum\_\{d^\{\\prime\}\\in\\mathcal\{N\}\_\{k\}\(q\)\}s\_\{\\theta\}\(q,d^\{\\prime\}\)\. We measure the similarity spreadσk\\sigma\_\{k\}as the within\-query standard deviation:

σk=1\|𝒬conf\|​∑q1k​∑d∈𝒩k​\(q\)\(sθ​\(q,d\)−s¯k​\(q\)\)2\\sigma\_\{k\}=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\}\\sqrt\{\\frac\{1\}\{k\}\\sum\_\{d\\in\\mathcal\{N\}\_\{k\}\(q\)\}\\left\(s\_\{\\theta\}\(q,d\)\-\\bar\{s\}\_\{k\}\(q\)\\right\)^\{2\}\}\(13\)
As shown in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\(e\), the spread remains constrained up tok=30k=30\. Keeping the boundary here ensures the diagnosis mechanism focuses exclusively on a tightly packed cluster of errors\.

##### Positive\-Negative Gap\.

To directly quantify the degree of model confusion, we calculate the positive\-negative gap, measuring the absolute difference between the true positive score and the average negative score\. A value ofγk<0\\gamma\_\{k\}<0highlights genuine confusion where negatives are scored higher than the positive\.

γk=1\|𝒬conf\|​∑q\(sθ​\(q,dq\+\)−1k​∑d∈𝒩k​\(q\)sθ​\(q,d\)\)\\gamma\_\{k\}=\\frac\{1\}\{\|\\mathcal\{Q\}\_\{\\mathrm\{conf\}\}\|\}\\sum\_\{q\}\\left\(s\_\{\\theta\}\(q,d\_\{q\}^\{\+\}\)\-\\frac\{1\}\{k\}\\sum\_\{d\\in\\mathcal\{N\}\_\{k\}\(q\)\}s\_\{\\theta\}\(q,d\)\\right\)\(14\)
As shown in Figure[5](https://arxiv.org/html/2609.18262#A3.F5)\(f\),γk\\gamma\_\{k\}becomes increasingly positive askkgrows past 30, meaning the average negative document becomes drastically easier than the positive document\. This solidifiesk=30k=30as the tipping point where true confusion is lost to general retrieval noise\.

Figure 5:Fine\-grainedkkablation \(k≤100k\\leq 100\) for the 0\.5B Iter\-0 retriever\. The vertical dash\-dot line marks the selected settingk=30k\{=\}30\.

### C\.4Effect of Fact\-Verified Data Beyond Backbone Choice

To further verify that the gain stems from our data rather than a particular backbone, we replace the backbone instead of the data\. We apply theREPAIRpipeline to the four backbones used byBMRetriever\(Pythia\-410M, Pythia\-1B, Gemma\-2B, and BioMistral\-7B\), and compare each model against the releasedBMRetrievermodel built on the same backbone\. Both models of a pair follow the setup of §[A\.2](https://arxiv.org/html/2609.18262#A1.SS2)and are evaluated on the five benchmarks of Table[1](https://arxiv.org/html/2609.18262#S4.T1)\. As shown in Table[11](https://arxiv.org/html/2609.18262#A3.T11),REPAIRimproves the average nDCG@10 on all four backbones, by\+0\.007\+0\.007to\+0\.019\+0\.019, and wins 18 of the 20 per\-benchmark comparisons\. The two exceptions both occur on Pythia\-1B, where SciFact ties and BIOSSES favorsBMRetriever\. This confirms that our fact\-verified data refinement drives the improvement across heterogeneous backbone families, rather than benefiting from the specific capacity of Qwen2\.5\.

Table 11:Experiments on the effect of fact\-verified data across the four backbones used byBMRetriever\. Each pair trains the same backbone onBMRetriever’s data and on ours\. All scores are reported in nDCG@10\. The best\-performing results within each backbone are highlighted inboldface\.
### C\.5Extended Iterations and Computational Cost

##### Saturation Beyond Two Iterations\.

Table[8](https://arxiv.org/html/2609.18262#A2.T8)extends the refinement loop to four iterations at every model scale under the same\(p,k\)\(p,k\)setting\. Moving from the second to the third iteration raises the average nDCG@10 by\+0\.005\+0\.005\(500M\),\+0\.003\+0\.003\(1\.5B\) and\+0\.003\+0\.003\(7B\), and a fourth iteration adds a further\+0\.001\+0\.001in all three cases, while each additional round consumes another 0\.5M training pairs\. The saturation point is thus the same across scales and is reached without re\-tuningpporkk, which is why we fix the number of iterations to two throughout the paper\.

##### Cost Structure\.

The cost profile ofREPAIRdiffers structurally from that of LLM\-based augmentation\. The Stage II API calls \(Semantic Scholar,PubChem,MatProj\) are issued offline in batch and are fully decoupled from the contrastive training loop, so they consume no GPU time\. Approaches that synthesize training data with a generative model instead pay an inference cost at every augmentation step, together with the downstream cost of filtering the hallucinations this introduces[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47);[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib58)\.REPAIRsecures fact\-verified evidence without incurring either\.

##### Per\-Stage Wall\-Clock Cost\.

Table[12](https://arxiv.org/html/2609.18262#A3.T12)itemizes the wall\-clock cost of a single iteration for each model scale, measured on two NVIDIA H200 GPUs under the training configuration of §[A\.2](https://arxiv.org/html/2609.18262#A1.SS2)\. Even for the 7B model, one iteration takes∼13\{\\sim\}13h, so the two iterations used throughout the paper amount to∼52\{\\sim\}52GPU\-hours\. For reference, RepLLaMA\-7B reports four days on16×16\{\\times\}V100 \(∼1,500\{\\sim\}1\{,\}500GPU\-hours\)[Ma et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib65), and Promptriever\-7B follows the same training recipe[Weller et al\. \(2025\)](https://arxiv.org/html/2609.18262#bib.bib64); fine\-tuning E5\-Mistral for SFR\-Embedding alone requires 120 GPU\-hours \(15h on8×8\{\\times\}A100\)[Meng et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib77), excluding its weakly\-supervised pre\-training stage[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib58)\. The total compute ofREPAIRis thus one to two orders of magnitude below that of comparable 7B\-scale baselines\.

Table 12:Per\-iteration wall\-clock cost of eachREPAIRstage, measured on2×2\{\\times\}H200 GPUs\. Stages I and III are GPU\-bound, while Stage II is bound by API latency and runs offline in batch without GPU cost\. All values are approximate\.

### C\.6Generalization to Additional Scientific Benchmarks

The nine benchmarks of §[4\.3](https://arxiv.org/html/2609.18262#S4.SS3)already span five task families, but they are drawn from materials science and biomedicine\. To test whether the long\-tail resolution mechanism ofREPAIRcarries to task formats and subareas it was never tuned for, we evaluate on three further scientific benchmarks that appear nowhere in training: DORIS\-MAE[Wang et al\. \(2023\)](https://arxiv.org/html/2609.18262#bib.bib78), whose queries are multi\-aspect research summaries; CQA\-physics[Hoogeveen et al\. \(2015\)](https://arxiv.org/html/2609.18262#bib.bib79);[Thakur et al\. \(2021\)](https://arxiv.org/html/2609.18262#bib.bib80), whose queries are informal community posts from a physics forum; and SciQ[Welbl et al\. \(2017\)](https://arxiv.org/html/2609.18262#bib.bib81), which spans physics, chemistry, biology, and earth science\. Dataset statistics are given in §[B\.2](https://arxiv.org/html/2609.18262#A2.SS2), and all scores are nDCG@10\.

Table[13](https://arxiv.org/html/2609.18262#A3.T13)groups the results by parameter tier\. Within every tierREPAIRoutperforms the correspondingBMRetrievermodel, by\+0\.040\+0\.040at 500M,\+0\.005\+0\.005at 1\.5B and\+0\.022\+0\.022at 7B, andREPAIR\-7B attains the highest average overall \(0\.65690\.6569\) despite training on 4M pairs\. The gains are largest on DORIS\-MAE, whereREPAIRleads at all three tiers, indicating that resolving long\-tail confusion transfers to the multi\-aspect query format the model never saw\. Two comparisons are closer\.REPAIR\-500M is on par with BGE\-Large \(0\.61430\.6143vs\.0\.61490\.6149\), which is trained on 2\.8B pairs, roughly700×700\\timesour data; andREPAIR\-1\.5B trailsBMRetriever\-2B on SciQ alone \(0\.79440\.7944vs\.0\.80130\.8013\) while remaining ahead on average\. Overall, performance holds up outside the domains and task formats the framework was developed on\.

Table 13:Experiments on three additional scientific benchmarks that are held out from training, grouped by parameter scale\. All scores are reported in nDCG@10 and given to four decimals, since several comparisons differ only in the fourth\. The best\-performing results within each scale are highlighted inboldface\.
### C\.7Citation Adjacency of Mined Hard Negatives

Table[14](https://arxiv.org/html/2609.18262#A3.T14)reports the citation check behind the claim in §[4\.3](https://arxiv.org/html/2609.18262#S4.SS3)\. For each source we sample 500 anchor documents, look up every pair inSemantic Scholar, and count a pair as adjacent if either document cites the other\. Two reference points frame the result\. Randomly paired documents give0\.000%0\.000\\%, the floor of the measurement\. SciDocs positive pairs, which are built from citation links and should therefore give100%100\\%, give only5\.80%5\.80\\%: the lookup finds a citation for just 29 of 500 pairs, becauseSemantic Scholarindexes few references for older papers\. This5\.80%5\.80\\%is thus the highest rate the check can return, not the true rate\. The hard negatives mined byREPAIRsit below0\.00001%0\.00001\\%, far closer to the random floor than to this ceiling, so the documents our diagnosis treats as negatives are almost never overlooked positives\.

Table 14:Citation rates of three sources of document pairs, measured with the sameSemantic Scholarlookup\. SciDocs positive pairs are already linked by citation, so their5\.80%5\.80\\%is the highest rate the lookup can detect rather than a true rate\.Table 15:Examples of LLM\-generated dataset errors\. Red text indicates hallucinated entities, misattributed scientific facts, or context stripped from negative documents\.

## Appendix DAnalysis of LLM\-Generated Dataset Errors

Table[15](https://arxiv.org/html/2609.18262#A3.T15)presents representative failure cases identified from a qualitative inspection of a subset of the synthetic dataset333[https://huggingface\.co/datasets/BMRetriever/biomed\_retrieval\_dataset](https://huggingface.co/datasets/BMRetriever/biomed_retrieval_dataset)generated by existing LLM\-based augmentation methods[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)\. Although our analysis is confined to a limited sample, the severity and fundamental nature of the uncovered errors suggest a risk that such structural hallucinations may be present throughout the corpus\. A more comprehensive investigation is warranted to determine the full spectrum of these critical flaws\. While naive prompt\-based generation has shown empirical success in general\-domain retrieval, our findings reveal that current LLMs fundamentally struggle with thelong\-tailed concept distribution \(P1\)andhigh fact\-sensitivity \(P2\)of scientific texts\. This limitation inevitably leads to the generation of harmful, hallucinatory data that degrades retriever performance\. Based on our manual review, we categorize the observed vulnerabilities into four primary failure modes, explicitly highlighting why our proposed methodology is strictly necessary to overcome these bottlenecks\.

### Failure Mode 1: Entity Number and Sub\-variant Swap \(P1 & P2\)

LLMs frequently treat structurally similar but biologically distinct entities as interchangeable tokens, especially within long\-tailed biomedical concepts\.

- •Analysis of Case 1 & 5:In Case 1, the LLM confusesBCL1withBCL2under the exact same context of "prosurvival myeloma proteins\." Similarly, in Case 5,CDK6is swapped withCDK4\. To a general\-domain LLM, a single\-digit difference represents a negligible semantic shift\. However, in the biomedical domain, this minor perturbation completely invalidates the scientific fact\.
- •Why our method is required:Naive generative models cannot self\-correct these single\-token factual violations\. Our methodology specifically addresses this by enforcing strict entity\-grounding constraints, ensuring that long\-tailed numerical variants are perfectly aligned between the query and the positive document\.

### Failure Mode 2: Fact Direction Reversal \(P2\)

Medical literature is highly sensitive to the directionality of outcomes \(e\.g\., increase vs\. decrease, inhibit vs\. promote\)\. LLMs often hallucinate these directional markers because opposite terms frequently co\-occur in similar training contexts\.

- •Analysis of Case 2:The generated query asks about adecreasedcardiovascular risk associated with Microalbuminuria \(MAU\), whereas the positive document explicitly states anenhancedrisk\. The LLM successfully grasped the topic \(MAU and cardiovascular risk\) but completely inverted the medical conclusion\.
- •Why our method is required:This demonstrates that semantic similarity alone is insufficient for scientific retrieval\. Our approach directly addresses High Fact\-Sensitivity \(P2\) by verifying the causal and directional consistency of the generated triplets, preventing the model from learning biologically fatal contradictions\.

### Failure Mode 3: Disease Entity Confusion via Context Stripping

LLMs often suffer from attention leakage when processing multiple documents, mistakenly integrating concepts from negative documents into the query intended for the positive document\.

- •Analysis of Case 3:The positive document describes a device to treatobesity\. However, the LLM insertsdiabetesinto the query\. This hallucination occurs because the surrounding negative documents \(or the LLM’s internal prior\) strongly associate obesity treatments with diabetes, causing a cross\-contamination of concepts\.
- •Why our method is required:This proves that providing LLMs with negative documents as prompt context often degrades query quality rather than improving it\. Our pipeline introduces a robust isolation mechanism that prevents negative context bleeding, maintaining the exact conceptual boundaries of the target document\.

### Failure Mode 4: Chemical Substitution

Similar to numerical swaps, LLMs fail to distinguish between fundamental chemical compounds that share functional or structural categories\.

- •Analysis of Case 4:The LLM replacessucrosewithfructose\. While both are sugars, the specific vacuole accumulation process described in the document is exclusive to sucrose in this experimental context\.
- •Why our method is required:Our proposed filtering and generation strategy explicitly penalizes out\-of\-context chemical substitutions\. By leveraging domain\-specific hard\-negative mining, we force the retriever to learn the precise distinctions between such granular entities, a capability entirely absent in datasets generated by baseline LLM approaches\.

##### Conclusion on Novelty

The examples delineated in Table[15](https://arxiv.org/html/2609.18262#A3.T15)are not mere edge cases; they are systemic failures stemming from the inherent architectural limitations of unconstrained LLMs\. Generating training data with these undetected hallucinations forces retrieval models to learn scientifically false representations\. The novelty of our proposed methodology lies in its structural capability to categorically eliminate these failure modes, specifically addressing long\-tailed entity swaps and fact\-direction reversals, thereby producing a high\-fidelity, factually rigorous dataset that significantly elevates biomedical retrieval performance\.

DatasetUser QueryModelTop\-1 Retrieved Snippet \(Truncated\)MatchSciFact
\(Biology/Fact\)Less than 10% of the gabonese children with SFM had a plasma lactate of more than 5mmol/L\.REPAIR\[Correct\]…measured body compartment volumes inGabonese childrenwith malaria…OBMR\[Irrelevant\]Compound heterozygous ZMPSTE24 mutations reduce prelamin A processing…XE5M\[Lexical Trap\]Lacticacidosis in patients withdiabetestreated with metformin…XNFCorpus
\(Nutrition\)red teaREPAIR\[Correct\]…elucidate health benefit ofherbal teas…green tea,black tea…OBMR\[Lexical Trap\]Colorredreduces snack food soft drink intake…XE5M\[Partial\]…antimutagenic activitywhite teacomparison green tea…XChemLit
\(Chemistry Proc\.\)What is the step before heating the solution in the process?REPAIR\[Correct\]…TheTeflon screw top was closedon the J\-young NMR tube, and thesolution was heated…OBMR\[Irrelevant\]Treatment of \[TpMo\(CO\)3\] with 1 equiv of gray Se in THF\-d8… failed to produce…XE5M\[Partial\]…solution was heated at 50 °C on a hot plate… turned from yellow to dark red…△\\triangleTable 16:Comparative Case Study of Retrieval Performance Across Diverse Domains

## Appendix EExtended Case Study Results

### E\.1Qualitative Analysis of Retrieval Capabilities

To explicitly demonstrate the superiority of the REPAIR framework over existing strong dense retrieval baselines,BMRetriever[Xu et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib47)and E5\-Mistral[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18262#bib.bib58), we present an in\-depth qualitative comparison\. We specifically targeted three highly specialized domains that challenge distinct retrieval capabilities: biomedical fact\-verification \(SciFact\), nutritional literature \(NFCorpus\), and chemical procedural reasoning \(ChemLit\)\. As illustrated in Table[16](https://arxiv.org/html/2609.18262#A4.T16), conventional models frequently fall into the trap of superficial lexical overlap or fail to capture complex relational logic\. In contrast, REPAIR successfully isolates deep semantic structures, factual nuances, and procedural causality\. This robustness directly stems from our self\-evolving methodology, which trains the model to comprehend holistic context rather than relying on token\-level matching\.

##### Case 1: Resolving Complex Factual Constraints \(SciFact\)\.

In the SciFact example, the user query demands the precise intersection of demographic data \("Gabonese children"\) and clinical measurements \("plasma lactate"\)\. While E5\-Mistral is completely derailed by the keyword "lactic" and retrieves an irrelevant document about lactic acidosis in diabetes \(a classic lexical trap\), REPAIR accurately localizes the specific demographic and clinical context\. This highlights REPAIR’s novelty in maintaining multi\-hop factual integrity without being distracted by high\-frequency medical jargon\.

##### Case 2: Ontological Understanding over Lexical Matching \(NFCorpus\)\.

The "red tea" query exposes the limitations of traditional semantic models in handling ambiguous, real\-world terms\.BMRetrievererroneously focuses on the exact color "red" in an entirely unrelated context \(snack food packaging\)\. Conversely, REPAIR exhibits a sophisticated understanding of ontological categories, successfully retrieving documents conceptually mapped to "herbal teas," "green tea," and "black tea\." This demonstrates REPAIR’s capability to map queries to broader semantic clusters, proving its effectiveness in domains where exact keyword overlaps are sparse\.

##### Case 3: Procedural and Temporal Reasoning \(ChemLit\)\.

Perhaps the most striking evidence of REPAIR’s novelty lies in the ChemLit domain, which strictly requires sequential reasoning\. The query explicitly asks for the stepbeforea specific action \("heating the solution"\)\. While E5\-Mistral retrieves a snippet that simply describes the heating process \(a partial match that entirely misses the temporal prerequisite\), REPAIR accurately identifies the chronological predecessor \("The Teflon screw top was closed"\)\. This proves that REPAIR goes beyond static semantic matching to comprehend dynamic, procedural causality, a significant and novel advancement over current baseline models\.

## Appendix FRobustness to Concept Extraction Noise

A fundamental strength of the proposed REPAIR framework is its capacity for continuous epistemic renewal\. Rather than stagnating in a self\-reinforcing feedback loop of existing model biases, the iterative refinement process dynamically resolves prior confusions while continuously uncovering novel epistemic boundaries\. This structural advantage is guaranteed by the Expansion stage, which anchors newly diagnosed concepts in externally verified knowledge bases \(e\.g\., Semantic Scholar,PubChem,MatProj\) rather than relying solely on internal model\-generated distributions\.

To empirically validate this dynamic self\-correction and demonstrate that the model does not merely reinforce its own bias, we analyze the evolution of the diagnosed confused concept set \(𝒞conf\\mathcal\{C\}\_\{\\text\{conf\}\}\) and the confusion query set \(𝒬conf\\mathcal\{Q\}\_\{\\text\{conf\}\}\) across consecutive iterations\. We track the transition from the seed model on the initial corpus𝒯0\\mathcal\{T\}\_\{0\}\(Iteration 1\) to the refined model on the augmented corpus𝒯1\\mathcal\{T\}\_\{1\}\(Iteration 2\) using the REPAIR\-500M setup\. Both iterations employ a selection ratio ofp=40%p=40\\%andk=30k=30negatives\. We utilize three key metrics to capture the nature of this representational shift:

##### Concept\-Level Set Overlap \(Jaccard Similarity\)\.

We first investigate whether the model is simply trapped in a cycle of repeating its past mistakes\. To quantify this, we calculate the Jaccard similarity between the confused concept set from Iteration 1 \(𝒞conf\(1\)\\mathcal\{C\}\_\{\\text\{conf\}\}^\{\(1\)\}\) and Iteration 2 \(𝒞conf\(2\)\\mathcal\{C\}\_\{\\text\{conf\}\}^\{\(2\)\}\)\.

Intuitively, if the refinement process were merely reinforcing existing biases, we would observe a high overlap; this would indicate that the model continually struggles with the exact same concepts \(epistemic stagnation\)\. Conversely, a low overlap demonstrates that the model successfully resolves past confusions and progresses to discover new, uncharted boundaries\.

As shown in Table[17](https://arxiv.org/html/2609.18262#A6.T17), the Jaccard similarity is remarkably low at0\.1460\.146\. This low overall overlap is driven by two highly positive outcomes: first, nearly half \(48\.6%48\.6\\%\) of the concepts that confused the Iteration\-1 model are completely resolved after just one refinement step\. Second, the vast majority \(83\.1%83\.1\\%\) of the concepts diagnosed in Iteration 2 are entirely novel\. Together, these statistics provide clear evidence that the model is actively expanding its knowledge rather than stagnating in a feedback loop\.

Table 17:Concept set overlap between𝒞conf\(1\)\\mathcal\{C\}\_\{\\text\{conf\}\}^\{\(1\)\}and𝒞conf\(2\)\\mathcal\{C\}\_\{\\text\{conf\}\}^\{\(2\)\}, REPAIR\-500M\.
##### Top\-KKSeverity Persistence and Rank Correlation\.

Beyond general set overlap, it is critical to determine whether the most severe confusions persist\. If a bias feedback loop were active, the highest\-ranked confusion targets \(measured by CCS score\) would remain anchored at the top of the distribution\. Table[18](https://arxiv.org/html/2609.18262#A6.T18)demonstrates that the Jaccard similarity for the top\-100 highest\-CCS concepts is strictly zero\. Extending this observation to the top\-1,000 yields a near\-zero similarity of0\.0030\.003\. Furthermore, among the fractional subset of concepts that do persist across both iterations, their severity ordering is fundamentally disrupted; the Spearman rank correlation \(ρ\\rho\) of their CCS scores is merely0\.1110\.111\(Table[19](https://arxiv.org/html/2609.18262#A6.T19)\)\. This confirms that the refinement process decisively dismantles the most severe representational bottlenecks\.

Table 18:Top\-KKconcept Jaccard by CCS rank, REPAIR\-500M\.MetricIter\-1Iter\-2Δ\\DeltaCCS Median2\.093\.89\+1\.80\+1\.80CCS Mean6\.885\.59−1\.28\-1\.28Spearmanρ\\rho\(CCS rank\)0\.111\\mathbf\{0\.111\}Table 19:CCS statistics for persistent concepts \(𝒞\(1\)∩𝒞\(2\)\\mathcal\{C\}^\{\(1\)\}\\cap\\mathcal\{C\}^\{\(2\)\}\), REPAIR\-500M\.
##### Query\-Level Margin Shift\.

Finally, we track the evolutionary trajectory at the query level\. The query Jaccard similarity stands at0\.00010\.0001\(Table[20](https://arxiv.org/html/2609.18262#A6.T20)\), indicating that the augmented corpus𝒯1\\mathcal\{T\}\_\{1\}successfully provides the necessary supervision to resolve nearly all queries that confused the Iteration\-1 model\. Crucially, we isolate the behavior of the 173 persistent queries that remain in the confused set during Iteration 2\. For this specific subset, we observe a positive margin shift from−2\.7×10−3\-2\.7\\times 10^\{\-3\}to\+3\.5×10−3\+3\.5\\times 10^\{\-3\}\. This metric directly illustrates that even when a query necessitates multiple refinement rounds, the model’s representational margins are actively expanding and separating, firmly countering any hypothesis of biased stagnation\.

Table 20:𝒬conf\\mathcal\{Q\}\_\{\\text\{conf\}\}overlap and average margin shift across iterations for REPAIR\-500M \(p=40%p=40\\%\)\.

Similar Articles

REVES: REvision and VErification--Augmented Training for Test-Time Scaling

Hugging Face Daily Papers

Proposes REVES, a two-stage iterative framework that alternates between data augmentation and policy optimization to improve LLM reasoning by leveraging intermediate correction steps, achieving superior performance on coding benchmarks and constraint satisfaction problems.

When Retrieval Doesn't Help: A Large-Scale Study of Biomedical RAG

arXiv cs.CL

A large-scale study across 5 models (7B–72B), 10 biomedical QA datasets, 4 retrieval methods, and 4 corpora finds that RAG yields only small and inconsistent gains (1–2 points) over no-retrieval baselines in biomedical question answering. The study concludes that the main bottleneck is not retrieval quality but models' limited ability to effectively use retrieved evidence.