MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG

arXiv cs.AI Papers

Summary

Mosaic is a training-free framework for query-aware exploration in GraphRAG that adapts retrieval policies based on query requirements, showing significant improvements in answer correctness and efficiency on standard benchmarks.

arXiv:2609.11065v1 Announce Type: new Abstract: Graph Retrieval-Augmented Generation (GraphRAG) can connect evidence distributed across a corpus graph, but most systems use largely shared exploration procedures across queries. This creates a structural mismatch: direct facts may need compact local neighborhoods, comparisons need balanced coverage of multiple targets, and mediated questions may require deeper paths through weakly related connectors. We present Mosaic, a training-free framework that formulates GraphRAG retrieval as a per-query control problem. An LLM analyzer converts query-specific evidence requirements into a bounded policy over seed selection, graph traversal, stopping, and evidence selection, while the corpus graph, indexes, scoring functions, grounding procedure, and answer generator remain shared. On GraphRAG-Bench, Mosaic achieves query-weighted Answer Correctness of 76.97 on Medical and 64.33 on Novel, improving over the strongest previously reported overall results by 5.13 and 4.43 points. On Medical, it reaches 95.1 Evidence Recall and 86.1 Context Relevancy. Controlled comparisons on an identical graph and generator show that no fixed narrow, medium, or wide policy is consistently optimal; Mosaic improves by 9.96 points over the strongest canonical fixed policy. Relative to Fixed Wide, it evaluates 81.9% fewer paths and retains 47.2% fewer evidence items. Transfer experiments on HotpotQA, MuSiQue, and 2WikiMultiHopQA further show that the policy interface can be applied without benchmark-specific retriever training.
Original Article
View Cached Full Text

Cached at: 09/12/26, 08:22 AM

# Query-Aware Exploration Policy Adaptation for GraphRAG Technical Report
Source: [https://arxiv.org/html/2609.11065](https://arxiv.org/html/2609.11065)
Kyeong\-Jin OhJinwon KimHye Woo LeeAffiliation:Minsang Song, Hyeongjun Jang, Junyoung YounAffiliation:KT Corporation

September 2026

###### Abstract

Graph Retrieval\-Augmented Generation \(GraphRAG\) can connect evidence distributed across a corpus graph, but most systems execute a largely shared exploration procedure for every query\. This creates a structural mismatch: a direct fact may require a compact local neighborhood, a comparison requires balanced coverage of multiple targets, and a mediated question may require a deeper path through a weakly query\-related connector\. We presentMosaic, a training\-free framework that formulates GraphRAG retrieval as a per\-query control problem\. An LLM analyzer translates the evidence requirements implied by a query into a bounded, executable policy over seed selection, graph traversal, cumulative\-gain stopping, and evidence selection\. The corpus graph, indexes, scoring functions, grounding procedure, and answer generator remain shared\. The current analyzer uses six representative considerations—Mediatedness, Output Cardinality, Seed Coverage, Anchoring Need, Identity–Relation Dependence, and Completeness—as composable reasoning signals rather than mutually exclusive query classes or a closed taxonomy\.

On GraphRAG\-Bench,Mosaicachieves query\-weighted Answer Correctness of 76\.97 on Medical and 64\.33 on Novel, improving over the strongest previously reported overall results by 5\.13 and 4\.43 points, respectively\. On Medical, it reaches 95\.1 Evidence Recall and 86\.1 Context Relevancy\. Controlled comparisons on an identical graph and generator show that no fixed narrow, medium, or wide policy is consistently optimal;Mosaicimproves by 9\.96 points over the strongest canonical fixed policy\. Relative to Fixed Wide, it evaluates 81\.9% fewer paths and retains 47\.2% fewer evidence items, although its separate analyzer call increases end\-to\-end latency\. Transfer experiments on HotpotQA, MuSiQue, and 2WikiMultiHopQA further show that the policy interface can be applied without benchmark\-specific retriever training\. These results support analyzer\-driven, query\-specific exploration as a general and extensible design principle for GraphRAG\.

Technical\-report scopeThis report consolidates the method, implementation details, controlled analyses, transfer experiments, and qualitative cases in one document\.Mosaicis*not*a six\-rule retrieval system: it is an analyzer\-driven framework for instantiating query\-specific retrieval policies\. The six considerations in the current implementation are representative and extensible design signals derived from observed retrieval failures\.

## 1Introduction

Retrieval\-Augmented Generation \(RAG\) grounds language\-model outputs in external evidence\[[9](https://arxiv.org/html/2609.11065#bib.bib9)\]\. GraphRAG extends this idea by representing entities, relations, passages, or communities as a graph and retrieving connected evidence for questions that require relation following, multi\-hop reasoning, or synthesis across sources\[[3](https://arxiv.org/html/2609.11065#bib.bib3),[4](https://arxiv.org/html/2609.11065#bib.bib4)\]\. The graph provides useful structure, but it also creates a control problem: retrieval must decide where to start, how far and broadly to traverse, when to stop, and which paths to preserve\.

These decisions are usually configured globally\. The query changes similarity scores or initial nodes, yet the main operating limits—seed count, depth, width, stopping behavior, and final evidence budget—remain largely shared\. A global policy is attractive because it is simple, but it assumes that heterogeneous questions require comparable evidence structures\. That assumption is often false\. A narrowly stated fact can be diluted by broad exploration; a comparison can fail if seeds cover only one target; and a mediated question can be unreachable under a shallow beam even when the answer\-bearing region exists in the graph\.

Mosaicreplaces the search for one globally optimal configuration with explicit per\-query policy construction\. Before retrieval, an analyzer interprets the question and emits bounded controls for the shared pipeline\. During retrieval, cumulative structural gain determines whether the current evidence has saturated, and role\-aware selection can preserve complementary paths\. Answer generation is held fixed: the analyzer changes what evidence is retrieved, not how the final answer is freely generated\.

The contribution is therefore not the observation that queries influence retrieval; many prior systems are query\-conditioned, and recent systems adapt routes, edges, constraints, or operators\. The technical distinction is the representation and target of adaptation:

query⟶explicit multi\-stage retrieval policy⟶shared corpus\-graph pipeline\.\\text\{query\}\\;\\longrightarrow\\;\\text\{explicit multi\-stage retrieval policy\}\\;\\longrightarrow\\;\\text\{shared corpus\-graph pipeline\}\.
The policy jointly controls seed breadth and formulation, traversal depth and width, anchor and connector behavior, stopping sensitivity, and evidence retention\. This interface is training\-free, validated before execution, and extensible to new evidence requirements\.

Our main findings are:

- •No fixed exploration scope is uniformly effective\. On Medical, Fixed Medium reaches 67\.01% ACC, while both narrower \(65\.45%\) and wider \(66\.35%\) exploration are worse\.
- •Query\-specific control substantially improves retrieval and generation\.Mosaicobtains 76\.97% overall ACC on Medical and 64\.33% on Novel, with Evidence Recall of 95\.1% and 90\.2%\.
- •The gain is not explained by exhaustive search\. Relative to Fixed Wide,Mosaicevaluates 81\.9% fewer paths and retains 47\.2% fewer evidence items\.
- •The analyzer emits diverse and requirement\-aligned policies\. For example, connector traversal is activated for 65\.6% of multi\-hop queries but 21\.6% of other queries\.
- •The framework transfers without benchmark\-specific retriever training to three standard multi\-hop QA benchmarks, while leaving room for stronger answer\-format calibration\.

## 2Related Work

RAG combines parametric generation with non\-parametric retrieval\[[9](https://arxiv.org/html/2609.11065#bib.bib9),[8](https://arxiv.org/html/2609.11065#bib.bib8)\]\. GraphRAG systems retrieve entities, relations, paths, passages, or graph communities: Microsoft GraphRAG supports local and global community\-oriented search\[[3](https://arxiv.org/html/2609.11065#bib.bib3)\]; LightRAG combines entity\- and relation\-level retrieval\[[4](https://arxiv.org/html/2609.11065#bib.bib4)\]; HippoRAG propagates query relevance through personalized PageRank\[[5](https://arxiv.org/html/2609.11065#bib.bib5)\]; and PathRAG prunes relational paths using flow\-based signals\[[2](https://arxiv.org/html/2609.11065#bib.bib2)\]\. Learned graph retrievers such as GFM\-RAG and G\-Reasoner encode textual and structural relevance with pretrained graph models\[[10](https://arxiv.org/html/2609.11065#bib.bib10),[11](https://arxiv.org/html/2609.11065#bib.bib11)\]\. Adaptive textual RAG methods decide whether or when to retrieve\[[6](https://arxiv.org/html/2609.11065#bib.bib6),[1](https://arxiv.org/html/2609.11065#bib.bib1),[7](https://arxiv.org/html/2609.11065#bib.bib7)\]\. Closer graph systems adapt the retrieval paradigm, query\-side evidence graph, path constraints, or operator composition\. EA\-GraphRAG routes between dense and graph retrieval\[[12](https://arxiv.org/html/2609.11065#bib.bib12)\]; Relink constructs a query\-driven evidence graph on the fly\[[13](https://arxiv.org/html/2609.11065#bib.bib13)\]; DOTRAG generates query\-conditioned constraints for path exploration\[[14](https://arxiv.org/html/2609.11065#bib.bib14)\]; and PAGE\-RAG composes heterogeneous evidence operators under bounded budgets\[[15](https://arxiv.org/html/2609.11065#bib.bib15)\]\.Mosaicdoes not claim that query\-aware GraphRAG is itself new\. It differs by constructing an explicit stage\-wise operating policy inside one shared corpus\-graph retriever, jointly controlling seeds, traversal, stopping, and retained evidence without training an additional router or graph retriever\.

Table 1:High\-level location of query adaptation in representative systems\. “Shared” means globally configured at inference time\.
## 3Problem Formulation

LetG=\(V,E\)G=\(V,E\)be a corpus graph constructed once from a document collection, and letqqbe a user query\. A conventional graph retriever applies a globally selected configurationπ¯\\bar\{\\pi\}:

y^q=ℳans​\(q,ℛ⁡\(G,q,π¯\)\)\.\\hat\{y\}\_\{q\}=\\mathcal\{M\}\_\{\\mathrm\{ans\}\}\\\!\\left\(q,\\mathcal\{R\}\(G,q;\\bar\{\\pi\}\)\\right\)\.The objective ofMosaicis to replaceπ¯\\bar\{\\pi\}with a query\-specific but bounded policyπq\\pi\_\{q\}while keepingGG, the retrieval implementationℛ\\mathcal\{R\}, and the answer modelℳans\\mathcal\{M\}\_\{\\mathrm\{ans\}\}shared:

πq=𝒜⁡\(q\)=\(πqseed,πqtrav,πqfilter\)\.\\pi\_\{q\}=\\mathcal\{A\}\(q\)=\\left\(\\pi\_\{q\}^\{\\mathrm\{seed\}\},\\pi\_\{q\}^\{\\mathrm\{trav\}\},\\pi\_\{q\}^\{\\mathrm\{filter\}\}\\right\)\.\(1\)The policy is generated only from the question\. Gold answers, gold evidence, benchmark category labels, and evaluation annotations are not available to the analyzer\. All categorical outputs are checked against an allowlist, numerical values are clamped to valid ranges, and invalid or missing fields fall back to conservative defaults\.

This formulation separates three concepts that are easily conflated:

1. 1\.Query conditioning: the query changes similarity or relevance scores\.
2. 2\.Query classification: the query is assigned to one of several predefined types, each linked to a fixed pipeline\.
3. 3\.Query\-specific policy construction: the query is translated into composable stage\-level controls executed by a common pipeline\.

Mosaicimplements the third\. Diagnostic labels may be logged for analysis or defensive fallback, but they do not select separate retrievers\.

## 4From Retrieval Failures to Policy Signals

We examined retrieval traces in which relevant evidence was present or reachable but the shared policy produced an incorrect answer\. The recurring failures did not define six mutually exclusive question types\. Instead, they exposed six questions the analyzer should ask when configuring retrieval\.[Table2](https://arxiv.org/html/2609.11065#S4.T2)summarizes the current instantiation\.

Table 2:Representative evidence\-requirement signals and their operational effects\. Multiple rows may apply to one query\.The six signals are deliberately not a closed taxonomy\. Another corpus, graph schema, domain, or pipeline may expose a new failure pattern\. The framework can incorporate it by adding a reasoning cue and mapping it to an existing or newly introduced bounded control\. The core contribution is the analyzer\-to\-policy interface, not the number or names of the present cues\.

### 4\.1Why controls must be composed

Individual controls are not independent treatments\. A comparison may simultaneously require per\-target seeds, connector traversal, and a larger evidence budget\. A mediated question may require greater depth but weaker anchoring so that an apparently generic connector is not pruned\. Strong anchoring can reduce drift, yet the same anchor can block a necessary mediator\. Similarly, stopping quality depends on the evidence produced by the seed and traversal stages\. Therefore, control\-specific counterfactuals are diagnostic rather than an additive decomposition of total performance\.

## 5MOSAIC

### 5\.1End\-to\-end architecture

QueryqqLLM Analyzer𝒜⁡\(q\)\\mathcal\{A\}\(q\)Query\-Specific Policyπq\\pi\_\{q\}Seed Policymode⋅\\cdotbudget⋅\\cdotscoringTraversal & Stop Policydepth⋅\\cdotwidth⋅\\cdotanchorconnector⋅\\cdotlow\-gain stopEvidence Policybudget⋅\\cdotfloor⋅\\cdotreservationSeed SelectionPath Traversaladaptive stopping during explorationEvidence SelectionFixed AnswerGenerationnot policy\-controlled

Figure 1:Mosaicmaps a query to one validated policy with stage\-specific controls\. Seed, traversal and stopping, and evidence controls configure a shared GraphRAG pipeline, while answer generation remains fixed\.Given a policy, the complete computation is

Sq\\displaystyle S\_\{q\}=𝒮⁡\(G,q,πqseed\),\\displaystyle=\\mathcal\{S\}\(G,q;\\pi\_\{q\}^\{\\mathrm\{seed\}\}\),\(2\)Pq\\displaystyle P\_\{q\}=𝒯⁡\(G,Sq,πqtrav\),\\displaystyle=\\mathcal\{T\}\(G,S\_\{q\};\\pi\_\{q\}^\{\\mathrm\{trav\}\}\),\(3\)E~q\\displaystyle\\widetilde\{E\}\_\{q\}=ℱ⁡\(q,Pq,πqfilter\),\\displaystyle=\\mathcal\{F\}\(q,P\_\{q\};\\pi\_\{q\}^\{\\mathrm\{filter\}\}\),\(4\)y^q\\displaystyle\\widehat\{y\}\_\{q\}=ℳans​\(q,Ground⁡\(E~q\)\)\.\\displaystyle=\\mathcal\{M\}\_\{\\mathrm\{ans\}\}\(q,\\operatorname\{Ground\}\(\\widetilde\{E\}\_\{q\}\)\)\.\(5\)

### 5\.2Policy space and validation

The analyzer starts from conservative defaults and overrides a control only when the question provides a clear signal\.[Table3](https://arxiv.org/html/2609.11065#S5.T3)reports the paper\-run interface\.

Table 3:Bounded policy controls\.The analyzer also returns target entities, target facets, a short retrieval\-risk description, and a rationale for diagnostics\. These fields do not bypass policy validation\. Unsupported categorical values are rejected, numeric fields are clamped, and missing values revert to defaults\. This ensures that the LLM controls a finite interface rather than directly executing arbitrary code or inventing new retrieval operators\.

### 5\.3Seed selection

The selected query mode searches fixed entity and relation indexes with the complete question, a compact query, or one query per extracted target\. Letsent​\(v,qe\)s\_\{\\mathrm\{ent\}\}\(v,q\_\{e\}\)andsrel​\(v,qe\)s\_\{\\mathrm\{rel\}\}\(v,q\_\{e\}\)denote entity\- and relation\-index scores for nodevv\. Relation\-only selection uses

sseed​\(v,q\)=srel​\(v,qe\),s\_\{\\mathrm\{seed\}\}\(v,q\)=s\_\{\\mathrm\{rel\}\}\(v,q\_\{e\}\),\(6\)while hybrid selection uses

sseed​\(v,q\)=0\.3​sent​\(v,qe\)\+0\.7​srel​\(v,qe\)\.s\_\{\\mathrm\{seed\}\}\(v,q\)=0\.3s\_\{\\mathrm\{ent\}\}\(v,q\_\{e\}\)\+0\.7s\_\{\\mathrm\{rel\}\}\(v,q\_\{e\}\)\.\(7\)The seed set isSq=TopKKq⁡\{v:sseed​\(v,q\)\}S\_\{q\}=\\operatorname\{TopK\}\_\{K\_\{q\}\}\\\{v:s\_\{\\mathrm\{seed\}\}\(v,q\)\\\}\. In per\-target mode, the two best candidates for each target are preserved before node\-level deduplication and final truncation\. This prevents a comparison from allocating all seeds to its most salient entity\.

### 5\.4Policy\-conditioned traversal

Letp=\(v0,e1,v1,…,eh,vh\)p=\(v\_\{0\},e\_\{1\},v\_\{1\},\\ldots,e\_\{h\},v\_\{h\}\)be a path withhhedges\. The structure\-aware score combines relation similarityr¯\\bar\{r\}, target relevancet¯\\bar\{t\}, specificitys¯\\bar\{s\}, anchor consistencya¯\\bar\{a\}, transition alignmentx¯\\bar\{x\}, path coherencec¯\\bar\{c\}, hub exposureh¯\\bar\{h\}, repetitionρrep\\rho\_\{\\mathrm\{rep\}\}, and question\-like labelsρql\\rho\_\{\\mathrm\{ql\}\}:

spath​\(p,q\)=wr​r¯\+wt​t¯\+ws​s¯\+wa​a¯\+wx​x¯\+wc​c¯−wh​h¯−wrep​ρreph−wql​ρql\.s\_\{\\mathrm\{path\}\}\(p,q\)=w\_\{r\}\\bar\{r\}\+w\_\{t\}\\bar\{t\}\+w\_\{s\}\\bar\{s\}\+w\_\{a\}\\bar\{a\}\+w\_\{x\}\\bar\{x\}\+w\_\{c\}\\bar\{c\}\-w\_\{h\}\\bar\{h\}\-w\_\{\\mathrm\{rep\}\}\\frac\{\\rho\_\{\\mathrm\{rep\}\}\}\{h\}\-w\_\{\\mathrm\{ql\}\}\\rho\_\{\\mathrm\{ql\}\}\.\(8\)Default weights are\(0\.45,0\.15,0\.15,0\.10,0\.15,0\.05,0\.20,1\.0,1\.0\)\(0\.45,0\.15,0\.15,0\.10,0\.15,0\.05,0\.20,1\.0,1\.0\)\. Whenaq≥0\.4a\_\{q\}\\geq 0\.4, target, anchor, transition, and hub terms increase withaqa\_\{q\}, changing a general relation\-oriented scorer into a target\- and transition\-aware scorer\. At depthdd, all retained paths are expanded, connector\-view candidates are added when enabled, and the bestWqW\_\{q\}paths are kept\.DqD\_\{q\}remains the hard cap\.

Connector exploration is a second view, not a separate pipeline\. Standard beam expansion grows outward from each seed; the connector view searches for paths that join regions initialized by distinct targets\. Both views enter the same scoring and pruning pool\.

### 5\.5Cumulative\-gain stopping

LetAd−1NA^\{N\}\_\{d\-1\}andAd−1EA^\{E\}\_\{d\-1\}be nodes and edges accumulated before depthdd, andNd,EdN\_\{d\},E\_\{d\}the structure reached at the current depth\. Novelty gain is

gd=\|Nd∖Ad−1N\|\+\|Ed∖Ad−1E\|max⁡\(1,\|Ad−1N\|\+\|Ad−1E\|\)\.g\_\{d\}=\\frac\{\|N\_\{d\}\\setminus A^\{N\}\_\{d\-1\}\|\+\|E\_\{d\}\\setminus A^\{E\}\_\{d\-1\}\|\}\{\\max\(1,\|A^\{N\}\_\{d\-1\}\|\+\|A^\{E\}\_\{d\-1\}\|\)\}\.\(9\)After minimum depthdmin=2d\_\{\\min\}=2, a step is eligible for early stopping whengd≤ϵrg\_\{d\}\\leq\\epsilon\_\{r\}and cumulative evidence completenesscd≥0\.35c\_\{d\}\\geq 0\.35\. With patience one,

Stop\(q,d\)=𝕀\[d≥Dq\]∨𝕀\[d≥2∧gd≤ϵrq∧cd≥0\.35\]\.\\operatorname\{Stop\}\(q,d\)=\\mathbb\{I\}\[d\\geq D\_\{q\}\]\\lor\\mathbb\{I\}\[d\\geq 2\\land g\_\{d\}\\leq\\epsilon\_\{r\_\{q\}\}\\land c\_\{d\}\\geq 0\.35\]\.\(10\)The paper\-run thresholds are 5\.00, 0\.35, and 0\.24 for shallow, medium, and deep regimes\. The large shallow threshold intentionally stops at the first completeness\-eligible depth; it does not mean novelty is numerically small in an absolute sense\.

### 5\.6Evidence selection and source grounding

After traversal, paths are deduplicated and ranked\. A shared path filter operates on at most ten candidates\. When adaptive reservation is disabled, selection uses the global top paths\. When enabled, it first preserves up to four core paths, then reserves up to four additional positions for complementary roles: a connector path, an uncovered target, an uncovered facet, or a previously unseen first\-relation label\. Remaining positions are filled by path score\.

The ranked paths are flattened into deduplicated evidence units\.BqB\_\{q\}caps retained evidence, whileLqL\_\{q\}restores high\-ranked excluded units when selection falls below the preservation floor\. Importantly, the evaluatedMosaicconfiguration does not add a separate LLM call for post\-traversal noise filtering\. Noise is controlled through policy\-conditioned traversal, path scoring, cumulative\-gain stopping, role\-aware reservation, and local source\-window reranking\.

Graph evidence is mapped back to source chunks\. Candidate chunks are divided into 1,600\-character windows with a stride of 800\. A local BAAI/bge\-large\-en\-v1\.5 bi\-encoder ranks the windows by cosine similarity to the question, and the top 15 windows are passed to GPT\-4o\-mini at temperature zero\. This preserves graph structure during exploration while returning full textual detail for answer generation\.

Algorithm 1: MOSAIC query\-specific graph retrievalInput:queryqq, corpus graphGG, entity and relation indexes\.Output:answery^q\\hat\{y\}\_\{q\}, grounded windowsWqW\_\{q\}\.1\.zq←AnalyzeRequirements​\(q\)z\_\{q\}\\leftarrow\\textsc\{AnalyzeRequirements\}\(q\);πq←ValidateAndInstantiate​\(zq\)\\pi\_\{q\}\\leftarrow\\textsc\{ValidateAndInstantiate\}\(z\_\{q\}\)\.2\.Sq←RetrieveSeeds​\(q,πqseed\)S\_\{q\}\\leftarrow\\textsc\{RetrieveSeeds\}\(q,\\pi\_\{q\}^\{\\mathrm\{seed\}\}\); initializeℬ0←Sq\\mathcal\{B\}\_\{0\}\\leftarrow S\_\{q\}andPq←∅P\_\{q\}\\leftarrow\\varnothing\.3\.Ford=1,…,πq\.Dmaxd=1,\\ldots,\\pi\_\{q\}\.D\_\{\\max\}: expandℬd−1\\mathcal\{B\}\_\{d\-1\}; rank and prune withπqtrav\\pi\_\{q\}^\{\\mathrm\{trav\}\}; merge retained paths intoPqP\_\{q\}; stop whend≥2d\\geq 2and the policy\-conditioned stopping test succeeds\.4\.Eq←SelectEvidence​\(Pq,q,πqfilter\)E\_\{q\}\\leftarrow\\textsc\{SelectEvidence\}\(P\_\{q\},q,\\pi\_\{q\}^\{\\mathrm\{filter\}\}\)\. If\|Eq\|<πq\.L\|E\_\{q\}\|<\\pi\_\{q\}\.L, restore the highest\-ranked excluded evidence up to the floor\.5\.Wq←GroundToSource​\(Eq,q\)W\_\{q\}\\leftarrow\\textsc\{GroundToSource\}\(E\_\{q\},q\);y^q←GenerateAnswer​\(q,Wq\)\\hat\{y\}\_\{q\}\\leftarrow\\textsc\{GenerateAnswer\}\(q,W\_\{q\}\); return\(y^q,Wq\)\(\\hat\{y\}\_\{q\},W\_\{q\}\)\.

## 6Experimental Setup

#### Primary benchmarks\.

GraphRAG\-Bench Medical contains 2,062 questions over 2,406 medical\-guideline documents; Novel contains 2,010 questions over 461 literary documents\. Both contain Fact Retrieval \(FR\), Complex Reasoning \(CR\), Contextual Summarization \(CS\), and Creative Generation \(CG\)\. We use every question and the original corpora\.

#### Metrics\.

Answer Correctness \(ACC\) is the primary end\-to\-end metric because it is defined for all four categories\. Overall ACC is query\-count weighted, not the unweighted mean of category scores\. Retrieval is evaluated with Evidence Recall \(ER\) and Context Relevancy \(CR\)\. Task\-specific ROUGE\-L, coverage, and faithfulness are reported as secondary metrics\.

#### Implementation\.

Each corpus graph is constructed once with a LightRAG\-style entity–relation extraction pipeline\. OpenAI text\-embedding\-3\-large supplies semantic representations\. GPT\-4o\-mini at temperature 0 is used for the analyzer and answer generation\. The graph, indexes, retrieval code, and generator are shared byMosaicand fixed\-policy controls\.Mosaicrequires no additional training or fine\-tuning\.

#### Baselines\.

We compare with reported GraphRAG\-Bench systems and with AutoPrunedRetriever and G\-Reasoner\. Because published baselines were not rerun in our implementation, claims involving them are comparisons to externally reported values\. To isolate policy adaptation, we additionally implement Fixed Narrow, Medium, Wide, and a Budget\-Matched Fixed policy on the identical graph and generator\.

## 7Results

### 7\.1Answer correctness

Table 4:Answer Correctness \(%\) on GraphRAG\-Bench\. Overall is query\-count weighted\. Best values are bold\.Mosaicobtains 76\.97 on Medical and 64\.33 on Novel, improving over the strongest previously reported overall results by 5\.13 and 4\.43 points\. On Medical it leads FR, CR, and CS; it does not lead CG\. On Novel, it leads overall and FR but trails specialized systems on CR, CS, and CG\. We therefore interpret the result as consistent overall accuracy across heterogeneous evidence requirements, not uniform dominance on every task category\.

### 7\.2Retrieval quality

Table 5:Retrieval performance \(%\)\. ER: Evidence Recall; CR: Context Relevancy\.Mosaicachieves the highest Evidence Recall on both domains, exceeding G\-Reasoner by 1\.3 points on Medical and 2\.5 points on Novel\. It also achieves the highest reported Medical Context Relevancy\. Novel reveals a useful trade\-off: HippoRAG2 has higher Context Relevancy, but substantially lower Evidence Recall\.

### 7\.3Why one fixed policy is insufficient

Table 6:Controlled fixed\-policy comparison on Medical\. All systems share graph and generator\.PolicySeedsDepthWidthEvidenceFRCRCSCGOverallNarrow3241–3\.620\.674\.733\.686\.6545Medium5485–8\.636\.683\.758\.703\.6701Wide1051612–15\.630\.672\.765\.683\.6635Mosaicadaptive per query\.761\.769\.853\.688\.7697The relationship between scope and accuracy is non\-monotonic\. Medium improves on Narrow, but Wide declines despite more seeds, deeper search, a larger beam, and more evidence\. Fixed Wide helps CS but hurts FR, CR, and CG relative to Medium\. Benchmark category also does not uniquely determine the correct policy: some fact questions require mediated evidence, while some summaries are supported by a compact neighborhood\.Mosaicexceeds the strongest canonical fixed policy by 9\.96 points, showing that the key benefit is query\-level allocation rather than choosing a single larger operating point\.

### 7\.4Policy diversity and requirement alignment

Across 2,062 Medical queries, the analyzer produces 19 unique policy combinations\. Assigned depth is distributed across 2, 3, and 4 hops for 26%, 43%, and 31% of queries\. Seed count ranges from 5 to 12; evidence budgets range from 8 to 16\. Coupled connector/two\-view exploration is activated for 32%, adaptive evidence selection for 35%, hybrid identity–relation scoring for 19%, and per\-target seeding for 19%\. This rules out collapse to a single dominant configuration\.

\(a\)Policy diversity\.\(b\)Requirement–behavior alignment\.
Figure 2:Analyzer outputs on 2,062 Medical questions\. Error bars in \(b\) are 95% Wilson intervals\.Independent post\-hoc proxies further show that variation is purposeful\. Two\-view traversal is activated for 90\.5% of multi\-target queries versus 30\.7% of others; a large evidence budget for 58\.4% of multi\-answer queries versus 12\.1%; connector traversal for 65\.6% of multi\-hop queries versus 21\.6%; shallow traversal for 48\.6% of local/direct queries versus 0\.8%; and early\-stop behavior for 57\.2% of closed\-form queries versus 10\.8%\. The proxies use deterministic surface, gold\-answer\-structure, and benchmark\-label rules only for analysis; they are never provided to the analyzer\.

### 7\.5Diagnostic counterfactuals

Table 7:Queries improved by the complete policy relative to disabling the corresponding execution control\. Counts are diagnostic and non\-additive\.These counterfactuals measure the operational effect of a control while retaining the full analyzer\. They should not be summed: the same query can improve through several interacting controls\. The strongest single count belongs to output cardinality, consistent with the tendency of high\-scoring paths to crowd out sibling evidence needed for list\-like answers\.

### 7\.6Latency and graph\-search effort

Table 8:Accuracy and average per\-query latency on Medical\. ACC uses all 2,062 queries; latency uses the same 197\-query subset\.Table 9:Graph\-search effort relative to Fixed Wide\.Mosaicis not faster end\-to\-end in the current implementation\. Its separate policy\-construction call adds 2\.82 seconds on average, producing 9\.68 seconds total versus 5\.29 for Fixed Medium\. The relevant efficiency result is different: additional LLM computation is used to allocate much less graph search more effectively\. The analyzer and seed\-keyword extraction are currently separate and unbatched; caching, prompt compression, call fusion, or distillation could reduce control latency without changing the retrieval interface\.

## 8Transfer to Multi\-Hop QA

We evaluate the same policy space on 1,000\-question subsets of HotpotQA, MuSiQue, and 2WikiMultiHopQA\. No benchmark training split or task\-specific retriever fine\-tuning is used\. One or two short answer\-format instructions are added for each benchmark\.

Table 10:Transfer results \(%\)\. G\-Reasoner values are externally reported and use target\-benchmark training;Mosaicis not trained on these benchmarks\.Mosaicis close to the trained reference on HotpotQA and exceeds it by 1\.8 EM on MuSiQue, while a substantial gap remains on 2WikiMultiHopQA\. A separate LLM\-based semantic correctness evaluation yields 77\.9, 51\.3, and 75\.1, suggesting that lexical EM/F1 penalize some semantically correct but differently formatted answers\. Because the systems were not rerun under a common pipeline and G\-Reasoner uses benchmark\-specific supervision, these results support transfer but do not establish a state\-of\-the\-art claim\.

## 9Qualitative Analysis

Table 11:Representative controlled cases from Medical\. Percentages denote ACC\.The successful cases show why controls must be combined\. More depth alone does not recover the kidney\-tumor answer because the larger context can still be dominated by a nearby but incorrect staging interpretation\. The failure case is equally important: an expressive controller can overreact\. Guardrails constrain the output range, but they cannot guarantee that the evidence requirement itself is inferred correctly\.

## 10Discussion

### 10\.1What the results establish

The controlled experiments establish that query\-specific policy adaptation improves over several globally fixed operating points on a shared graph and generator\. Policy logs show that the controller neither collapses to one configuration nor varies arbitrarily: activations align with independent requirement proxies\. The graph\-search measurements show that gains are not produced by uniformly retrieving more evidence\.

### 10\.2What the results do not establish

The experiments do not isolate a causal contribution for every prompt phrase or every low\-level scoring term\. Controls interact, and the counterfactual counts are not additive\. Comparisons to published GraphRAG systems are not component\-matched reproductions\. Transfer results compare against external G\-Reasoner numbers under asymmetric training conditions\. Finally, the current latency reflects an unoptimized separate analyzer call\.

### 10\.3Extensibility and deployment

The bounded policy schema offers a practical separation of concerns\. Retrieval engineers define safe controls and defaults; the analyzer maps natural\-language evidence requirements to those controls; monitoring uses policy logs and retrieval traces\. New failure patterns can be addressed without introducing a new end\-to\-end pipeline for each query type\. In production, policies can be cached for recurring query structures, distilled into a smaller controller, or constrained further by domain rules\. Because the answer model consumes grounded source windows and remains outside the controller’s free\-form action space, the framework also supports clearer auditing of how a query changed retrieval\.

## 11Limitations

First, the analyzer is an LLM and can misclassify evidence structure, as the SLNB case demonstrates\. Second, the current signals and valid ranges were developed from observed failures in the evaluated setting; transfer to different graph schemas may require additional calibration\. Third, graph construction quality bounds retrieval: missing or incorrectly merged entities and relations cannot be repaired solely by a better policy\. Fourth, answer quality still depends on source grounding and generation, particularly for open\-ended CG questions whereMosaicdoes not uniformly improve\. Fifth, latency and monetary cost are higher than fixed retrieval in the current implementation\. Finally, some comparisons rely on externally reported baselines, and a fully component\-matched reproduction across all methods remains future work\.

## 12Conclusion

Mosaictreats GraphRAG retrieval as a per\-query control problem\. An analyzer converts evidence requirements into a bounded policy spanning seed selection, traversal, stopping, and evidence retention, while the corpus graph and answer generator remain shared\. The approach improves overall answer correctness and evidence recall on GraphRAG\-Bench and uses substantially less graph search than a uniformly wide policy\. Its six current considerations are composable implementation signals, not a fixed taxonomy\. The broader result is that effective GraphRAG should adapt not merely relevance scores, but the operating policy of graph exploration itself\.

## Artifact and Reproducibility Statement

The evaluated implementation contains proprietary components\. For result verification, the authors intend to provide benchmark configurations, dependencies, graph construction and retrieval code required for reproduction, inference scripts, generated outputs, and official evaluation commands through a controlled\-access repository, subject to organizational approval and applicable benchmark licenses\.

## References

- \[1\]A\. Asai et al\. Self\-RAG: Learning to Retrieve, Generate, and Critique through Self\-Reflection\. ICLR, 2024\.
- \[2\]B\. Chen et al\. PathRAG: Pruning Graph\-Based Retrieval Augmented Generation with Relational Paths\. AAAI, 2026\.
- \[3\]D\. Edge et al\. From Local to Global: A Graph RAG Approach to Query\-Focused Summarization\. arXiv:2404\.16130, 2024\.
- \[4\]Z\. Guo et al\. LightRAG: Simple and Fast Retrieval\-Augmented Generation\. arXiv:2410\.05779, 2024\.
- \[5\]B\. Gutiérrez et al\. HippoRAG: Neurobiologically Inspired Long\-Term Memory for Large Language Models\. NeurIPS, 2024\.
- \[6\]Z\. Jiang et al\. Active Retrieval Augmented Generation\. EMNLP, 2023\.
- \[7\]S\. Jeong et al\. Adaptive\-RAG: Learning to Adapt Retrieval\-Augmented Large Language Models through Question Complexity\. NAACL, 2024\.
- \[8\]V\. Karpukhin et al\. Dense Passage Retrieval for Open\-Domain Question Answering\. EMNLP, 2020\.
- \[9\]P\. Lewis et al\. Retrieval\-Augmented Generation for Knowledge\-Intensive NLP Tasks\. NeurIPS, 2020\.
- \[10\]L\. Luo et al\. GFM\-RAG: Graph Foundation Model for Retrieval Augmented Generation\. NeurIPS, 2025\.
- \[11\]L\. Luo et al\. G\-Reasoner: Foundation Models for Unified Reasoning over Graph\-Structured Knowledge\. arXiv preprint, 2025\.
- \[12\]S\. Dong, Q\. Zhang, Y\. Xiao, S\. Chen, C\. Zhou, and X\. Huang\. Use Graph When It Needs: Efficiently and Adaptively Integrating Retrieval\-Augmented Generation with Graphs\. arXiv:2602\.03578, 2026\.
- \[13\]M\. Huang, C\. Bu, Y\. He, X\. Zhuo, and X\. Wu\. Relink: Constructing Query\-Driven Evidence Graph On\-the\-Fly for GraphRAG\. arXiv:2601\.07192, 2026\.
- \[14\]L\. Moore, N\. Deng, R\. Mihalcea, and F\. Jahanbakhsh\. DOTRAG: Retrieval\-Time Reasoning Along Paths\. arXiv:2605\.18760, 2026\.
- \[15\]X\. Chen, J\. An, J\. Guo, and L\. Wang\. PAGE\-RAG: Evidence\-Grounded Adaptive Graph Retrieval for Long\-Document Question Answering\. arXiv:2607\.19301, 2026\.
- \[16\]Y\. Xiao, J\. Dong, C\. Zhou, S\. Dong, Q\. Zhang, D\. Yin, X\. Sun, and X\. Huang\. GraphRAG\-Bench: Challenging Domain\-Specific Reasoning for Evaluating Graph Retrieval\-Augmented Generation\. arXiv:2506\.02404, 2025\.

Similar Articles

Accurate and Efficient Long-Term Memory for LLM Agents

arXiv cs.AI

MOSAIC is a structured, conflict-aware long-term memory framework for LLM agents that uses entity-typed graph storage, hash-accelerated retrieval, and active conflict detection to achieve high accuracy and efficiency on long-conversation QA and factual conflict detection tasks.

GRASP: GRanularity-Aware Search Policy for Agentic RAG

Hugging Face Daily Papers

Introduces GRASP, a reinforcement learning framework that trains agents to adaptively coordinate semantic search, keyword search, and paragraph reading during multi-step reasoning, improving retrieval recall and question answering performance on multi-hop benchmarks.