Follow the Entities: A Corpus Map for Agentic Search

Hugging Face Daily Papers Papers

Summary

CorpusMap is an entity-based navigation layer for agentic search that organizes large document collections around recurring entities to improve evidence discovery and answer quality while reducing token usage.

Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as a project's approval recorded in one, its requirements in another, and its latest status in a third. Recent LLM agents approach this by iteratively searching the full corpus rather than reading only a fixed set of top-ranked documents. However, when the corpus is exposed only as a flat collection of files, a relevant document gives no indication of how it relates to others, so the agent must rediscover these relationships for every query, often missing complementary evidence while simultaneously consuming substantial additional tokens. To address this, we introduce CorpusMap, a navigation layer that organizes the corpus around its recurring entities, which are identifiable from the documents themselves and can link a single document to many others across sources. Specifically, CorpusMap represents each recurring entity as an Entity Page that aggregates information about it and links to every document that refers to it, forming a graph between entities and documents that the agent can traverse to gather otherwise disconnected evidence. Moreover, since CorpusMap is constructed offline by resolving mentions of the same entity across documents, its links are shared across queries rather than rediscovered repeatedly at inference time. Using 7 different models with 3 benchmark datasets, we show that CorpusMap improves both evidence discovery and answer quality over raw-corpus agentic search while using fewer tokens on average, and further outperforms 4 alternative navigation layers, suggesting that entities serve as effective anchors for navigating large document collections.
Original Article
View Cached Full Text

Cached at: 09/30/26, 08:20 AM

Paper page - Follow the Entities: A Corpus Map for Agentic Search

Source: https://huggingface.co/papers/2609.37226

Abstract

Answeringquestionsandcompletingtasksoverlargedocumentcollectionsoftenrequiresconnectingevidencespreadacrossmultipledocuments,suchasaproject’sapprovalrecordedinone,itsrequirementsinanother,anditslateststatusinathird.RecentLLMagentsapproachthisbyiterativelysearchingthefullcorpusratherthanreadingonlyafixedsetoftop-rankeddocuments.However,whenthecorpusisexposedonlyasaflatcollectionoffiles,arelevantdocumentgivesnoindicationofhowitrelatestoothers,sotheagentmustrediscovertheserelationshipsforeveryquery,oftenmissingcomplementaryevidencewhilesimultaneouslyconsumingsubstantialadditionaltokens.Toaddressthis,weintroduceCorpusMap,anavigationlayerthatorganizesthecorpusarounditsrecurringentities,whichareidentifiablefromthedocumentsthemselvesandcanlinkasingledocumenttomanyothersacrosssources.Specifically,CorpusMaprepresentseachrecurringentityasanEntityPagethataggregatesinformationaboutitandlinkstoeverydocumentthatreferstoit,formingagraphbetweenentitiesanddocumentsthattheagentcantraversetogatherotherwisedisconnectedevidence.Moreover,sinceCorpusMapisconstructedofflinebyresolvingmentionsofthesameentityacrossdocuments,itslinksaresharedacrossqueriesratherthanrediscoveredrepeatedlyatinferencetime.Using7differentmodelswith3benchmarkdatasets,weshowthatCorpusMapimprovesbothevidencediscoveryandanswerqualityoverraw-corpusagenticsearchwhileusingfewertokensonaverage,andfurtheroutperforms4alternativenavigationlayers,suggestingthatentitiesserveaseffectiveanchorsfornavigatinglargedocumentcollections.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.37226

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.37226 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.37226 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.37226 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Beyond Semantic Similarity: Rethinking Retrieval for Agentic Search via Direct Corpus Interaction

Hugging Face Daily Papers

The paper introduces Direct Corpus Interaction (DCI), a novel approach allowing AI agents to query raw text directly using standard terminal tools instead of traditional embedding-based retrieval. By bypassing fixed similarity interfaces and offline indexing, DCI significantly outperforms conventional sparse, dense, and reranking baselines across multiple IR and agentic search benchmarks.