BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
Summary
A controlled scaling study of retrieval-augmented generation paradigms finds that BM25 lexical retrieval outperforms agentic and graph-based retrieval at scale, while agentic search only leads on small corpora.
View Cached Full Text
Cached at: 07/31/26, 05:53 AM
Paper page - BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms
Source: https://huggingface.co/papers/2607.26497
Abstract
Retrieval-augmentedgeneration(RAG)spanslexicalanddenseretrieval,graph-basedindexing,andagenticsearch,buttheseparadigmsareusuallyevaluatedondifferentbenchmarksatonecorpussize,leavingtheiraccuracy-costscalingunclear.Tobridgethisgap,wepresentacontrolledstudythatvariescorpussizealong28strictlynestedtiersspanningroughly450-fold,whileholdingquestionsandafixedbedrockofrelevantandadversarialdocumentsunchanged.Underonereadermodelandonejudgingprotocol,wemeasureofficialaccuracy,constructionandquerytokens,andlatency.Theresultsrevealascale-dependentcrossoverratherthananunconditionalwinner.File-SystemAgentleadsatthesmallestsharedtiers,butitssequentialexplorationcosts39timesmorequerytokensatthebedrockandbecomeslesseffectiveasthesearchspacegrows.Around10millioncorpustokens,BM25overtakesitandleadsateverylargersharedtier,withamarginapproaching20pointsatfullscale.BM25alsoanchorsthelow-costendoftheParetofrontierwithoutLLM-basedconstruction.Denseretrievalremainsefficientbutlessaccurate,whereasgraph-basedRAGencountersconstructionwallsbeforedeploymentscaleanditsscalablevariantsremainbelowBM25atsharedtiers.Overall,corpusgrowthincreasinglyfavorsglobalcandidateranking:lexicalretrievalisthestrongestscalabledefault,whileagenticreasoningworksbestafterrankeddiscoveryratherthaninplaceofit.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.26497
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.26497 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.26497 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.26497 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms
This paper presents a controlled scaling study comparing lexical, dense, graph-based, and agentic RAG paradigms across corpus sizes from 1,000 to 512,000 documents, finding that BM25 provides the best accuracy-cost tradeoff, while graph-based RAG faces high construction costs that limit scalability.
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
Introduces MKG-RAG-Bench, a cross-domain benchmark for evaluating retrieval in multimodal knowledge graph-augmented generation, demonstrating that effective multimodal retrieval remains challenging and critical for downstream generation quality.
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
AACL-IJCNLP 2026 Findings paper showing that classic lexical retrieval (BM25 plus a small cross-encoder) matches or beats agentic skill retrieval on agent skill selection across 3 of 4 settings, at roughly half the cost, arguing agentic retrieval should be the fallback rather than the default.
Retrieval-augmented generation solves a problem most teams don't actually have
The article argues that retrieval-augmented generation (RAG) is often misapplied in AI systems, where the real issue is context curation rather than retrieval. It suggests that RAG is only truly beneficial for large, frequently changing corpora.
MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A
MM-BizRAG is a multimodal retrieval-augmented generation system for enterprise Q&A that uses document structure-aware splitting and layout-aware parsing to outperform vision-centric baselines by up to 32% on heterogeneous enterprise documents. The paper also introduces FastRAGEval, a cost-efficient LLM-based evaluation metric with stronger human alignment than RAGChecker.