A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation
Summary
This paper proposes HyGRAG, a hierarchical graph RAG framework that integrates contextual and relational information for multi-hop reasoning, achieving a 9.7% average accuracy improvement over existing methods.
View Cached Full Text
Cached at: 06/17/26, 05:41 AM
# A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation
Source: [https://arxiv.org/html/2606.18075](https://arxiv.org/html/2606.18075)
\(2026\)
###### Abstract\.
Retrieval\-Augmented Generation \(RAG\) has emerged as a paradigm for enhancing large language models \(LLMs\) with external knowledge, yet existing graph\-based methods face a fundamental limitation: entity\-centric and chunk\-centric approaches operate on representations anchored to original text without true knowledge fusion\. While entity\-centric methods connect logically related content and chunk\-centric methods preserve context, both retrieve information separately through similarity search, missing emergent understanding from their synthesis\. In this paper, we proposeHyGRAG, a hierarchical graph RAG framework that transcends source documents by addressing three core challenges: constructing summaries that genuinely integrate contextual and relational information, leveraging these synthesized representations to access emergent knowledge during retrieval, and efficiently updating hierarchical structures for dynamic corpora\. Specifically, we design hierarchical index structures over hybrid graphs with both chunk and entity nodes, then iteratively cluster them and generate LLM\-based summaries\. Then, we design context and relation\-aware retrieval that searches across all abstraction levels while expanding through community membership\. Moreover, we enable dynamic knowledge update through attachment\-based algorithms with only local re\-summarization\. Experimental results show thatHyGRAGimproves the average accuracy of multi\-hop reasoning tasks by 9\.7%, while maintaining reasonable efficiency\.111Our codes are available at[https://github\.com/zjunet/HyGRAG](https://github.com/zjunet/HyGRAG)\.
Graph RAG; context\-aware; relation\-aware\.
††journalyear:2026††copyright:cc††conference:Proceedings of the ACM Web Conference 2026; April 13–17, 2026; Dubai, United Arab Emirates††booktitle:Proceedings of the ACM Web Conference 2026 \(WWW ’26\), April 13–17, 2026, Dubai, United Arab Emirates††doi:10\.1145/3774904\.3792720††isbn:979\-8\-4007\-2307\-0/2026/04††ccs:Networks Online social networks††ccs:Computing methodologies Knowledge representation and reasoning## 1\.Introduction
Figure 1\.Illustration of method performance using Qwen3\-8B across relation\-aware \(MuSiQue\) and context\-aware \(MultiHop\-RAG\) datasets, including representative cases\.RAG systems have been developed to enhance LLMs by integrating external knowledge sources\(Gaoet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib122); Lewiset al\.,[2021](https://arxiv.org/html/2606.18075#bib.bib128)\)\. This integration allows LLMs to access up\-to\-date information and domain\-specific knowledge, addressing the inherent limitations of parametric memory\(Fanet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib123)\)\. Recent advances in RAG have moved beyond simple document retrieval to incorporate graph structures that explicitly model relationships between concepts, entities, and documents\(Huet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib125); Xianget al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib126)\)\. These graph RAG systems leverage the rich structural information in knowledge graphs to improve both retrieval accuracy and reasoning capabilities, making them particularly effective for complex queries that require understanding interconnected information\(Maet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib127); Gutiérrezet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib106)\)\.
Current Graph RAG methods have explored two distinct directions to harness graph structures for retrieval augmentation\. Entity\-centric approaches such as GraphRAG\(Edgeet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib105)\), HippoRAG\(Gutiérrezet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib106)\), and HiRAG\(Huanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib108)\)build knowledge graphs where nodes represent entities extracted from text and edges encode their semantic relationships\. These methods excel at multi\-hop reasoning by enabling traversal along relational paths—for instance, answering ”Which company acquired the developer of GPT\-3?” by following edges from GPT\-3 to OpenAI to Microsoft\. In contrast, chunk\-centric methods like RAPTOR\(Sarthiet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib103)\)and EraRAG\(Zhanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib104)\)organize text chunks into hierarchical structures based on semantic similarity\. They create tree\-like indexes where leaf nodes contain original text chunks and parent nodes store increasingly abstract summaries\. This design preserves full contextual information while enabling retrieval at different levels of granularity\.
However, both approaches exhibit fundamental limitations that restrict their effectiveness\. Entity\-centric methods suffer from information loss during the entity extraction process\. Named entity recognition and relation extraction models introduce errors that compound throughout the system, and the abstraction from text to entities discards valuable contextual information needed for factual question answering\. Our analysis shows that these methods often underperform simple dense retrieval on factual QA benchmarks\. Chunk\-centric methods face the opposite problem: while they maintain high accuracy on factual queries, they cannot capture explicit relationships between entities scattered across different chunks\. This makes them ineffective for queries requiring logical reasoning, such as finding all products affected by a specific supply chain disruption\. The fundamental issue is that existing methods optimize for either contextual completeness or relational structure, but real\-world queries often require both\.
Our key insight is that simply combining chunk and entity representations in a hybrid graph does not solve the fundamental problem: both representations remain anchored to their original textual sources without true knowledge fusion\. While knowledge graphs can connect logically related content that is textually distant, they cannot integrate the logical information from relations with the background context from chunks\. During retrieval, these two aspects are still accessed separately through similarity search, missing the opportunity for synergistic understanding\. Consider a query about ”the impact of renewable energy adoption on manufacturing industries”—entity\-based retrieval might find relationships between solar panels and factories, while chunk\-based retrieval might return documents about energy costs, but neither captures the emergent understanding that comes from synthesizing these perspectives\. This motivates our approach: we cluster entities and chunks together and generate summaries that simultaneously encode both relational logic and contextual background, creating new knowledge representations that transcend the original text\.
Implementing this knowledge fusion approach presents three core technical challenges\. First, how to construct summaries that genuinely integrate context and relations rather than simply concatenating them\. Second, how to leverage these hybrid summaries during retrieval to access knowledge beyond the original corpus\. The summaries represent emergent understanding that doesn’t exist in any single source document—the retrieval mechanism must be able to match queries to these synthesized insights while still providing access to supporting details\. Third, how to efficiently update such complex hierarchical structures when the corpus changes\. Real\-world knowledge bases receive continuous updates, but rebuilding the entire hierarchy for each new document would be computationally prohibitive\.
To address these challenges, we proposeHyGRAG, a unified framework with three key innovations\. For summary construction, we design a two\-stage approach: first clustering nodes using graph structure\-aware embeddings that capture both textual similarity and relational connectivity, then prompting LLMs with structured templates that explicitly request synthesis of contextual and relational information within each community\. This produces summaries that represent genuine knowledge fusion rather than simple aggregation\. For retrieval, we implement a bi\-level mechanism that operates across both context and relation dimensions\. Context\-aware retrieval searches across all hierarchy levels—from specific chunks to abstract community summaries—enabling access to emergent knowledge\. Relation\-aware retrieval then expands results by collecting entities from retrieved communities and filtering their associated triplets, ensuring logical completeness\. For dynamic updates, we develop an attachment\-based algorithm where new content is matched to the most similar community at the appropriate abstraction level, with local re\-summarization propagating only along affected paths\. This preserves the stability of unrelated portions while incorporating new knowledge\. Experimental results show thatHyGRAGimproves the average accuracy of multi\-hop reasoning tasks by 9\.7%, while maintaining reasonable efficiency\.
In summary, we make three contributions:
- •We propose hierarchical indexing for Graph RAG that unifies context and relational information through multi\-level abstraction, enabling retrieval across different granularities\.
- •We develop a retrieval mechanism that combines context\-aware and relation\-aware strategies, leveraging community structures to find information that flat approaches miss\.
- •We achieve significant performance gains across a range of tasks, including a 6\.2% improvement in Factual Accuracy and a 9\.7% increase in Multi\-Hop Reasoning, with improvements reaching up to 12\.2% on the HotpotQA dataset\.
## 2\.Related Work
### 2\.1\.Graph RAG
RAG\(Gaoet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib122)\)mitigates hallucinations\(Huanget al\.,[2025b](https://arxiv.org/html/2606.18075#bib.bib121)\)in LLMs by integrating external knowledge during generation\. RAG involves two coupled steps: retrieval and generation\. Given a query, the system retrieves the top\-k relevant text chunks from a preprocessed external corpus\. The LLM then generates an answer conditioned on both the retrieved chunks and the query, producing more factual and contextually appropriate responses\.
Vanilla RAG\(Lewiset al\.,[2020](https://arxiv.org/html/2606.18075#bib.bib102)\)improves factual grounding but treats the corpus as unstructured text, limiting inter\-document reasoning and multi\-hop inference\. To overcome this, recent work incorporates graph structures into retrieval\(Penget al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib124)\), giving rise to graph RAG methods\. These can be categorized as context\-aware, focusing on text chunks, and relation\-aware, explicitly leveraging graph structures\.
Context\-awaremethods enhance generation by retrieving segmented text chunks as context\. Vanilla RAG retrieves top\-k chunks based on embedding similarity\. Advanced approaches cluster and summarize chunks to capture high\-order information\. RAPTOR\(Sarthiet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib103)\)recursively clusters chunks into hierarchical summaries, while EraRAG\(Zhanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib104)\)organizes the corpus into multi\-layered summaries using LSH, handling high\-order information in dynamic corpora\.
Relation\-awaremethods construct textual knowledge graphs to encode structured relationships\. LightRAG\(Guoet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib109)\)integrates graph structures into the text indexing ,and retrieve by high/low\-level keywords\. GraphRAG\(Edgeet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib105)\)cluster communities and summarize them to provide holistic context\. HiRAG\(Huanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib108)\)and ArchRAG\(Wanget al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib110)\)further leverage hierarchical graphs to bridge local entity details and global insights by applying attributed entities\. HippoRAG\(Gutiérrezet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib106)\)and HippoRAG 2\(gutiérrez2025ragmemorynonparametriccontinual\)enhance multi\-hop reasoning by navigating graph connectivity via Personalized PageRank, improving performance on complex QA tasks\.
### 2\.2\.Dynamic Retrieval
Most graph RAG systems assume a static corpus, making updates costly\. LightRAG\(Guoet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib109)\)uses a modular retriever to add new documents dynamically\. HippoRAG\(Gutiérrezet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib106)\)also supports incremental updates, as it leverages the KG only for retrieval\. EraRAG\(Zhanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib104)\)introduces LSH\-based clustering, enabling localized updates by re\-segmenting and re\-summarizing only affected parts\. However, entity\-based community clustering methods still require full reconstruction\.
Figure 2\.Overall architecture ofHyGRAG\.
## 3\.Method
Overview\.Our proposed framework,HyGRAG, addresses the fundamental trade\-off between context\-aware and relation\-aware retrieval through a unified architecture\. The system consists of four key components: \(1\)Hierarchical Index Structure Constructionconstructs a hierarchical indexing structure with hybrid graph that captures both chunk\-based contextual information and entity\-based relational information; \(2\)Context and Relation\-Aware Retrievalperforms retrieval across micro and macro granularities; \(3\)Retrieval\-Augmented Efficient Generationensemble heirarchical context; \(4\)Dynamic Knowledge Updatesupports incremental knowledge integration without full reconstruction\.
### 3\.1\.Hierarchical Index Structure Construction
This module operates entirely in the offline phase, ensuring no impact on online retrieval efficiency\. We construct a hybrid knowledge graph from raw text that simultaneously captures chunk\-based contextual information and entity\-based relational structures\.
Text Chunking and Chunk\-Level Graph Construction\.Given a corpus𝒟\\mathcal\{D\}, we first segment it into overlapping text chunks𝒞=\{c1,c2,…,cn\}\\mathcal\{C\}=\\\{c\_\{1\},c\_\{2\},\\ldots,c\_\{n\}\\\}\. Each chunkcic\_\{i\}preserves the contextual information of the original text, enabling inference LLMs to better understand its semantic content\. For each chunk, we compute its embedding representation:
\(1\)𝐞ic=LMEmbedding\(ci\),\\mathbf\{e\}\_\{i\}^\{c\}=\\text\{LM\}\_\{\\text\{Embedding\}\}\(c\_\{i\}\),where we employ BGE\-M3 as the embedding model to encode text chunks into dense vector representations\. Based on embedding similarity, we construct a chunk\-based graph𝒢c=\(𝒱c,ℰc\)\\mathcal\{G\}^\{c\}=\(\\mathcal\{V\}^\{c\},\\mathcal\{E\}^\{c\}\), where𝒱c=𝒞\\mathcal\{V\}^\{c\}=\\mathcal\{C\}\. We create edges between chunks based on shared entities rather than shallow embedding similarity\. Specifically, if two chunkscic\_\{i\}andcjc\_\{j\}share more thanllcommon entities, we create an edge\(ci,cj\)∈ℰc\(c\_\{i\},c\_\{j\}\)\\in\\mathcal\{E\}^\{c\}:
\(2\)\(ci,cj\)∈ℰc⇔\|ℰi∩ℰj\|\>l,\(c\_\{i\},c\_\{j\}\)\\in\\mathcal\{E\}^\{c\}\\Leftrightarrow\|\\mathcal\{E\}\_\{i\}\\cap\\mathcal\{E\}\_\{j\}\|\>l,whereℰi\\mathcal\{E\}\_\{i\}andℰj\\mathcal\{E\}\_\{j\}are the entity sets extracted from chunkscic\_\{i\}andcjc\_\{j\}respectively, andllis set to 3 in our implementation\. This approach establishes meaningful connections between chunks through shared entities, providing stronger semantic relationships than surface\-level embedding similarity\.
Entity\-Level Graph Construction\.Then we employ LLMs to extract knowledge graph triplets from the raw text\. Specifically, for each text chunkcic\_\{i\}, the LLM identifies entities and their relationships, generating a set of triplets:
\(3\)𝒯i=\{\(h,r,t\)∣h,t∈ℰi,r∈ℛi\},\\mathcal\{T\}\_\{i\}=\\\{\(h,r,t\)\\mid h,t\\in\\mathcal\{E\}\_\{i\},r\\in\\mathcal\{R\}\_\{i\}\\\},whereℰi\\mathcal\{E\}\_\{i\}andℛi\\mathcal\{R\}\_\{i\}are the entity and relation sets extracted fromcic\_\{i\}, respectively\. By merging triplets from all chunks, we obtain the global entity set𝒱e=⋃i=1nℰi\\mathcal\{V\}^\{e\}=\\bigcup\_\{i=1\}^\{n\}\\mathcal\{E\}\_\{i\}and relation setℰe=⋃i=1n𝒯i\\mathcal\{E\}^\{e\}=\\bigcup\_\{i=1\}^\{n\}\\mathcal\{T\}\_\{i\}, forming the entity\-based knowledge graph𝒢e=\(𝒱e,ℰe\)\\mathcal\{G\}^\{e\}=\(\\mathcal\{V\}^\{e\},\\mathcal\{E\}^\{e\}\)\.
Hybrid Graph Construction\.To bridge the gap between chunk\-based context and entity\-based relations, we connect the two graphs through membership relationships\. For each entitye∈𝒱ee\\in\\mathcal\{V\}^\{e\}, we identify the set of text chunks containing that entity:
\(4\)𝒞\(e\)=\{ci∈𝒞∣eis mentioned inci\}\.\\mathcal\{C\}\(e\)=\\\{c\_\{i\}\\in\\mathcal\{C\}\\mid e\\text\{ is mentioned in \}c\_\{i\}\\\}\.We create cross\-layer edgesℰce\\mathcal\{E\}^\{ce\}connecting entities to their containing chunks:
\(5\)ℰce=\{\(e,c\)∣e∈𝒱e,c∈𝒞\(e\)\}\.\\mathcal\{E\}^\{ce\}=\\\{\(e,c\)\\mid e\\in\\mathcal\{V\}^\{e\},c\\in\\mathcal\{C\}\(e\)\\\}\.The final hybrid knowledge graph is defined as:
\(6\)𝒢hybrid=\(𝒱c∪𝒱e,ℰc∪ℰe∪ℰce\)\.\\mathcal\{G\}\_\{\\text\{hybrid\}\}=\(\\mathcal\{V\}^\{c\}\\cup\\mathcal\{V\}^\{e\},\\mathcal\{E\}^\{c\}\\cup\\mathcal\{E\}^\{e\}\\cup\\mathcal\{E\}^\{ce\}\)\.This graph contains two types of nodes \(chunk nodes and entity nodes\) and three types of edges \(chunk\-to\-chunk edges, entity\-to\-entity edges, and cross\-layer edges\)\.
Hierarchical Indexing\.While the hybrid graph enriches information representation, it also increases the complexity of online retrieval\. To efficiently retrieve both relevant contextual and relational information simultaneously, we propose to construct a hierarchical tree\-structured index for the hybrid graph\.
First, we use Cleora to generate structure\-aware embeddings for all nodesv∈𝒱c∪𝒱ev\\in\\mathcal\{V\}^\{c\}\\cup\\mathcal\{V\}^\{e\}in the hybrid graph:
\(7\)𝐳v=Cleora\(𝒢hybrid,v\)\.\\mathbf\{z\}\_\{v\}=\\text\{Cleora\}\(\\mathcal\{G\}\_\{\\text\{hybrid\}\},v\)\.
Then, we employ hyperplane\-based Locality\-Sensitive Hashing \(LSH\) for efficient clustering\. For each node embedding𝐳v∈ℝd\\mathbf\{z\}\_\{v\}\\in\\mathbb\{R\}^\{d\}, we randomly samplekkhyperplanes\{𝐡1,…,𝐡k\}\\\{\\mathbf\{h\}\_\{1\},\\ldots,\\mathbf\{h\}\_\{k\}\\\}and compute its hash code:
\(8\)hash\(𝐳v\)=\[sign\(𝐳v⋅𝐡1\),…,sign\(𝐳v⋅𝐡k\)\]\\text\{hash\}\(\\mathbf\{z\}\_\{v\}\)=\[\\text\{sign\}\(\\mathbf\{z\}\_\{v\}\\cdot\\mathbf\{h\}\_\{1\}\),\\ldots,\\text\{sign\}\(\\mathbf\{z\}\_\{v\}\\cdot\\mathbf\{h\}\_\{k\}\)\]Nodes with similar hash codes \(small Hamming distance\) are assigned to the same bucket\. For each bucketBB, if its size satisfiesSmin≤\|B\|≤SmaxS\_\{\\min\}\\leq\|B\|\\leq S\_\{\\max\}, it is retained; otherwise, split or merge operations are performed:
\(9\)Badjusted=\{Split\(B\)if\|B\|\>SmaxMerge\(B,Bneighbor\)if\|B\|<SminBotherwiseB\_\{\\text\{adjusted\}\}=\\begin\{cases\}\\text\{Split\}\(B\)&\\text\{if \}\|B\|\>S\_\{\\max\}\\\\ \\text\{Merge\}\(B,B\_\{\\text\{neighbor\}\}\)&\\text\{if \}\|B\|<S\_\{\\min\}\\\\ B&\\text\{otherwise\}\\end\{cases\}
For each adjusted bucket \(forming a community\)CC, we generate summary representations through a two\-step process:
\(10\)tC=LLMsummarize\(\{v∣v∈C\}\),t\_\{C\}=\\text\{LLM\}\_\{\\text\{summarize\}\}\(\\\{v\\mid v\\in C\\\}\),
\(11\)sC=LMEmbedding\(tC\),s\_\{C\}=\\text\{LM\}\_\{\\text\{Embedding\}\}\(t\_\{C\}\),
where we utilize Llama3\.1\-8B\-Instruct to generate concise textual summariestCt\_\{C\}that capture the semantic essence of nodes within each community, and then employ BGE\-M3 to encode these summaries into dense vector representationssCs\_\{C\}\.
We treat summary nodes as nodes in the next layer and recursively apply the above hashing, bucketing, and summarization process up to a predefined number of layers\. Letℒℓ\\mathcal\{L\}\_\{\\ell\}denote the node set at layerℓ\\ell, whereℒ0=𝒱c∪𝒱e\\mathcal\{L\}0=\\mathcal\{V\}^\{c\}\\cup\\mathcal\{V\}^\{e\}represents the leaf nodes\. The hierarchical construction continues until the layer indexℓ\\ellreaches the preset maximum depthLL, resulting in anLL\-layer hierarchical index tree:
\(12\)𝒯index=ℒ0,ℒ1,…,ℒL\.\\mathcal\{T\}\{\\text\{index\}\}=\{\\mathcal\{L\}\_\{0\},\\mathcal\{L\}\_\{1\},\\ldots,\\mathcal\{L\}\_\{L\}\}\.This hierarchical tree serves as an efficient retrieval database during the dual\-stage retrieval phase, where leaf nodes correspond to original nodes from the hybrid graph and upper\-layer nodes contain semantic summaries at coarser granularities, supporting multi\-scale context and relation\-aware retrieval\.
### 3\.2\.Context and Relation\-Aware Retrieval
During the online retrieval phase, we employ a bi\-level approach that captures information at both context and relation levels to provide comprehensive and contextually rich results\.
Query Encoding\.Given a user queryqq, we first encode it using the same embedding model employed during indexing:
\(13\)𝐪=LMEmbedding\(q\),\\mathbf\{q\}=\\text\{LM\}\_\{\\text\{Embedding\}\}\(q\),where we utilize BGE\-M3 to generate the query embedding𝐪∈ℝd\\mathbf\{q\}\\in\\mathbb\{R\}^\{d\}that serves as the basis for similarity computation across all retrieval levels\.
Context\-Aware Retrieval\.From the hierarchical indexing structure𝒯index\\mathcal\{T\}\_\{\\text\{index\}\}containing community, chunk, and entity nodes, we first perform context\-aware retrieval by computing similarity scores across all node types\. We retrieve the top\-kkmost similar nodes for each type:
\(14\)𝒫retrieved=TopK\(\{C∈⋃ℓ=1Lℒℓ∣sim\(𝐪,sC\)\},k\),\\mathcal\{P\}\_\{\\text\{retrieved\}\}=\\text\{TopK\}\(\\\{C\\in\\bigcup\_\{\\ell=1\}^\{L\}\\mathcal\{L\}\_\{\\ell\}\\mid\\text\{sim\}\(\\mathbf\{q\},s\_\{C\}\)\\\},k\),
\(15\)𝒞retrieved=TopK\(\{c∈𝒱c∣sim\(𝐪,𝐞cc\)\},k\),\\mathcal\{C\}\_\{\\text\{retrieved\}\}=\\text\{TopK\}\(\\\{c\\in\\mathcal\{V\}^\{c\}\\mid\\text\{sim\}\(\\mathbf\{q\},\\mathbf\{e\}\_\{c\}^\{c\}\)\\\},k\),
\(16\)ℰretrievedcontext=TopK\(\{e∈𝒱e∣sim\(𝐪,𝐳e\)\},k\)\.\\mathcal\{E\}\_\{\\text\{retrieved\}\}^\{\\text\{context\}\}=\\text\{TopK\}\(\\\{e\\in\\mathcal\{V\}^\{e\}\\mid\\text\{sim\}\(\\mathbf\{q\},\\mathbf\{z\}\_\{e\}\)\\\},k\)\.
These retrievals capture content most relevant to the query from a contextual perspective across different granularities\.
Relation\-Aware Retrieval\.To capture logical relationships, we construct a comprehensive entity set by combining entities from retrieved communities and context\-based entity retrieval:
\(17\)ℰall=ℰretrievedcontext∪⋃C∈𝒫retrieved\{e∣e∈𝒱e,\(e,C\)∈ℰce\}\.\\mathcal\{E\}\_\{\\text\{all\}\}=\\mathcal\{E\}\_\{\\text\{retrieved\}\}^\{\\text\{context\}\}\\cup\\bigcup\_\{C\\in\\mathcal\{P\}\_\{\\text\{retrieved\}\}\}\\\{e\\mid e\\in\\mathcal\{V\}^\{e\},\(e,C\)\\in\\mathcal\{E\}^\{ce\}\\\}\.
We then extract all relations connected to these entities from leaf nodes:
\(18\)ℛall=\{\(h,r,t\)∈ℰe∣h∈ℰall∨t∈ℰall\}\.\\mathcal\{R\}\_\{\\text\{all\}\}=\\\{\(h,r,t\)\\in\\mathcal\{E\}^\{e\}\\mid h\\in\\mathcal\{E\}\_\{\\text\{all\}\}\\vee t\\in\\mathcal\{E\}\_\{\\text\{all\}\}\\\}\.
Since\|ℛall\|\|\\mathcal\{R\}\_\{\\text\{all\}\}\|can be large, potentially introducing noise and LLM overhead, we filter by computing triplet embeddings and selecting the top\-kkmost relevant:
\(19\)ℛretrieved=TopK\(\{\(h,r,t\)∈ℛall∣𝐳\(h,r,t\)=LMEmbedding\(h⊕r⊕t\)\},k\),\\mathcal\{R\}\_\{\\text\{retrieved\}\}=\\text\{TopK\}\(\\\{\(h,r,t\)\\in\\mathcal\{R\}\_\{\\text\{all\}\}\\mid\\mathbf\{z\}\_\{\(h,r,t\)\}=\\text\{LM\}\_\{\\text\{Embedding\}\}\(h\\oplus r\\oplus t\)\\\},k\),where⊕\\oplusdenotes concatenation\.
Integrated Retrieval Strategy\.The final retrieval combines context\-aware and relation\-aware perspectives, yielding four sets ofkkelements each:
\(20\)ℛfinal=\{𝒫retrieved,𝒞retrieved,ℰretrievedcontext,ℛretrieved\},\\mathcal\{R\}\_\{\\text\{final\}\}=\\\{\\mathcal\{P\}\_\{\\text\{retrieved\}\},\\mathcal\{C\}\_\{\\text\{retrieved\}\},\\mathcal\{E\}\_\{\\text\{retrieved\}\}^\{\\text\{context\}\},\\mathcal\{R\}\_\{\\text\{retrieved\}\}\\\},providing community summaries for high\-level understanding, chunks for detailed context, entities for key concepts, and relations for logical connections\.
Computational Complexity Analysis\.Similar to other retrieval\-augmented approaches,HyGRAGemploys a FAISS vector store with HNSW indexing to efficiently retrieve the top\-kkmost similar items\. LetNNdenote the total number of indexed nodes andddthe embedding dimension\. The HNSW\-based search achieves logarithmic complexity for each index:𝒪\(logNi⋅d\),\\mathcal\{O\}\(\\log N\_\{i\}\\cdot d\),and retrieving from entity\-, chunk\-, and community\-level indexes results in:
\(21\)𝒪\(\(logNe\+logNc\+logNp\)⋅d\)=𝒪\(logN⋅d\)\.\\mathcal\{O\}\(\(\\log N\_\{e\}\+\\log N\_\{c\}\+\\log N\_\{p\}\)\\cdot d\)=\\mathcal\{O\}\(\\log N\\cdot d\)\.For each retrieved entity, relation extraction and relevance scoring introduce an additional cost of𝒪\(ke⋅d¯⋅d\)\\mathcal\{O\}\(k\_\{e\}\\cdot\\bar\{d\}\\cdot d\), whered¯\\bar\{d\}is the average node degree\. The overall retrieval complexity is thus:
\(22\)𝒪\(logN⋅d\+ke⋅d¯⋅d\),\\mathcal\{O\}\(\\log N\\cdot d\+k\_\{e\}\\cdot\\bar\{d\}\\cdot d\),which maintains sub\-linear scaling with respect to graph size while enabling multi\-granularity and relation\-aware retrieval\.
Compared with traditional chunk\-aware RAG systems that only perform a single chunk\-level retrieval𝒪\(logNc⋅d\)\\mathcal\{O\}\(\\log N\_\{c\}\\cdot d\), our method introduces additional but lightweight costs from entity and community retrievals𝒪\(\(logNe\+logNp\)⋅d\+ke⋅d¯⋅d\)\\mathcal\{O\}\(\(\\log N\_\{e\}\+\\log N\_\{p\}\)\\cdot d\+k\_\{e\}\\cdot\\bar\{d\}\\cdot d\)\.
### 3\.3\.Retrieval\-Augmented Efficient Generation
Given the retrieved nodes from bi\-level retrieval, we extract and organize information for LLM generation\. For community nodesC∈𝒫retrievedC\\in\\mathcal\{P\}\_\{\\text\{retrieved\}\}, we obtain their summariestCt\_\{C\}\. For entity nodese∈ℰretrievedcontexte\\in\\mathcal\{E\}\_\{\\text\{retrieved\}\}^\{\\text\{context\}\}, we extract their textual representations\. For relation triplets\(h,r,t\)∈ℛretrieved\(h,r,t\)\\in\\mathcal\{R\}\_\{\\text\{retrieved\}\}, we preserve their structured format\. For chunk nodesc∈𝒞retrievedc\\in\\mathcal\{C\}\_\{\\text\{retrieved\}\}, we directly use their textual content\.
We combine these four information types using a structured prompt template:
\(23\)Context=𝒫\(\{tC\}C∈𝒫retrieved,ℰretrievedcontext,ℛretrieved,𝒞retrieved\),\\text\{Context\}=\\mathcal\{P\}\(\\\{t\_\{C\}\\\}\_\{C\\in\\mathcal\{P\}\_\{\\text\{retrieved\}\}\},\\mathcal\{E\}\_\{\\text\{retrieved\}\}^\{\\text\{context\}\},\\mathcal\{R\}\_\{\\text\{retrieved\}\},\\mathcal\{C\}\_\{\\text\{retrieved\}\}\),where𝒫\\mathcal\{P\}organizes community summaries, entities, relation triplets, and chunk contexts hierarchically\.
The final response is generated by:
\(24\)y=LLMgenerate\(q,Context\),y=\\text\{LLM\}\_\{\\text\{generate\}\}\(q,\\text\{Context\}\),leveraging community summaries for high\-level understanding, entities for key concepts, relation triplets for logical reasoning, and chunk contexts for detailed background information\.
### 3\.4\.Dynamic Knowledge Update
In real\-world scenarios, knowledge corpora evolve continuously, requiring efficient update mechanisms\. Our clustering\-based hierarchical structure enables fast integration of new content without full graph reconstruction\.
Update Processing\.Given new text contentdnewd\_\{\\text\{new\}\}, we first segment it into chunks𝒞new=\{cnew1,…,cnewm\}\\mathcal\{C\}\_\{\\text\{new\}\}=\\\{c\_\{\\text\{new\}\}^\{1\},\\ldots,c\_\{\\text\{new\}\}^\{m\}\\\}and extract knowledge graph triplets:
\(25\)𝒯new=⋃i=1m\{\(h,r,t\)∣h,t∈ℰinew,r∈ℛinew\}\.\\mathcal\{T\}\_\{\\text\{new\}\}=\\bigcup\_\{i=1\}^\{m\}\\\{\(h,r,t\)\\mid h,t\\in\\mathcal\{E\}\_\{i\}^\{\\text\{new\}\},r\\in\\mathcal\{R\}\_\{i\}^\{\\text\{new\}\}\\\}\.
We generate a summary representation for the new content:
\(26\)tnew=LLMsummarize\(𝒞new∪𝒯new\),t\_\{\\text\{new\}\}=\\text\{LLM\}\_\{\\text\{summarize\}\}\(\\mathcal\{C\}\_\{\\text\{new\}\}\\cup\\mathcal\{T\}\_\{\\text\{new\}\}\),\(27\)𝐬new=LMEmbedding\(tnew\)\.\\mathbf\{s\}\_\{\\text\{new\}\}=\\text\{LM\}\_\{\\text\{Embedding\}\}\(t\_\{\\text\{new\}\}\)\.
Hierarchical Attachment\.We traverse the hierarchical index from bottom to top\. Starting at layerℒ1\\mathcal\{L\}\_\{1\}, we find the most similar community:
\(28\)C∗=argmaxC∈ℒ1sim\(𝐬new,sC\)\.C^\{\*\}=\\arg\\max\_\{C\\in\\mathcal\{L\}\_\{1\}\}\\text\{sim\}\(\\mathbf\{s\}\_\{\\text\{new\}\},s\_\{C\}\)\.
Ifsim\(𝐬new,sC∗\)\>τattach\\text\{sim\}\(\\mathbf\{s\}\_\{\\text\{new\}\},s\_\{C^\{\*\}\}\)\>\\tau\_\{\\text\{attach\}\}, we attach the new content toC∗C^\{\*\}\. Otherwise, we proceed to the next layer:
\(29\)ℓ∗=min\{ℓ∣∃C∈ℒℓ:sim\(𝐬new,sC\)\>τattach\}\.\\ell^\{\*\}=\\min\\\{\\ell\\mid\\exists C\\in\\mathcal\{L\}\_\{\\ell\}:\\text\{sim\}\(\\mathbf\{s\}\_\{\\text\{new\}\},s\_\{C\}\)\>\\tau\_\{\\text\{attach\}\}\\\}\.
Upon attachment at layerℓ∗\\ell^\{\*\}, we update all ancestor community summaries along the path to the root:
\(30\)∀j∈\{ℓ∗,…,L\},tCj←LLMsummarize\(Children\(Cj\)\),\\forall j\\in\\\{\\ell^\{\*\},\\ldots,L\\\},\\quad t\_\{C\_\{j\}\}\\leftarrow\\text\{LLM\}\_\{\\text\{summarize\}\}\(\\text\{Children\}\(C\_\{j\}\)\),\(31\)sCj←LMEmbedding\(tCj\)\.s\_\{C\_\{j\}\}\\leftarrow\\text\{LM\}\_\{\\text\{Embedding\}\}\(t\_\{C\_\{j\}\}\)\.
Additionally, we establish connections between new nodes and existing leaf nodes following the original hybrid graph construction rules, creating edges when entities are shared or semantic similarity exceeds thresholds\. Since all updates occur offline, online retrieval efficiency remains unaffected\.
## 4\.Experiments
In this section, we answer the following questions to validate the effectiveness of our method:
- •RQ1\.How doesHyGRAGperform compared with existing baselines on different type of static QA tasks?
- •RQ2\.How efficient is theHyGRAGapproach?
- •RQ3\.How robust isHyGRAGin corpus expansion?
- •RQ4\.What is the quality of our communities and relations, and how does it impact the performance of RAG?
### 4\.1\.Experimental Setup
#### 4\.1\.1\.Datasets and Metrics
For static QA, we adopt five datasets:PopQA\(Mallenet al\.,[2023](https://arxiv.org/html/2606.18075#bib.bib116)\)\(factual accuracy\),MuSiQue\(Trivediet al\.,[2022](https://arxiv.org/html/2606.18075#bib.bib119)\), andHotpotQA\(Yanget al\.,[2018](https://arxiv.org/html/2606.18075#bib.bib120)\)\(multi\-hop reasoning\), andMultiHop\-RAG\(Tang and Yang,[2024](https://arxiv.org/html/2606.18075#bib.bib118)\),QuALITY\(Panget al\.,[2022](https://arxiv.org/html/2606.18075#bib.bib117)\)\(reading comprehension\)\. For evaluation, static datasets reportAccuracyandRecall\(QuALITY uses Accuracy only\)\.
Table 1\.Overall QA results\(Accuracy and Recall\) on static RAG query datasets using Llama\-3\.1\-8B\-Instruct\.
#### 4\.1\.2\.Baseline
We compare our approach with three categories of RAG baselines: \(1\) LLM\-only:ZeroShot,Chain\-of\-Thought \(CoT\)\(Kojimaet al\.,[2022](https://arxiv.org/html/2606.18075#bib.bib101)\); \(2\) Context\-aware:VanillaRAG\(Lewiset al\.,[2020](https://arxiv.org/html/2606.18075#bib.bib102)\),RAPTOR\(Sarthiet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib103)\)\(GMM/K\-means clustering algorithm\), andEraRAG\(Zhanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib104)\); \(3\) Relation\-aware:L\-GraphRAG\(Edgeet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib105)\)\(local mode of Microsoft GraphRAG\),HippoRAG\(Gutiérrezet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib106)\),HippoRAG2\(gutiérrez2025ragmemorynonparametriccontinual\),HiRAG\(Huanget al\.,[2025a](https://arxiv.org/html/2606.18075#bib.bib108)\),ArchRAG\(Wanget al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib110)\)andLightRAG\(Guoet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib109)\)\. Variants of LightRAG, namely Local, Global, and Hybrid, are denoted asL/G/H\-LightRAGfor brevity\.
#### 4\.1\.3\.Implementation Details
All static\-query experiments use Llama\-3\.1\-8B\-Instruct\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib114)\)as backbone under the vLLM\(Kwonet al\.,[2023](https://arxiv.org/html/2606.18075#bib.bib113)\)engine with greedy decoding and top\-k=5 retrieval\. For a fair comparison, we maintain retrieval sizes across all node types, i\.e\.,entity nodes=community \+ chunk nodes=relations=top\-k\\text\{entity nodes\}=\\text\{community \+ chunk nodes\}=\\text\{relations\}=\\text\{top\-\}k\. Embeddings are generated by BGE\-M3\(Chenet al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib111)\), a SOTA embedding model that supports both multilingual and multi\-granularity retrieval, splitting text into 1,200\-token chunks with 100\-token overlap\. Baselines follow the framework\(Zhouet al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib112)\)or official code with default hyperparameters\. Runs exceeding two days are marked asOOT\.
### 4\.2\.Static QA Performace
To answer the question RQ1, we present the experimental results of question answering \(QA\) on various types of static queries, including factual accuracy, multi\-hop reasoning, and reading comprehension tasks\. Our method is systematically compared with a range of baseline models to assess its effectiveness across different reasoning scenarios\. Furthermore, the statistics of different datasets are summarized in the Appendix[C](https://arxiv.org/html/2606.18075#A3)\.
#### 4\.2\.1\.QA results
As shown in Table[1](https://arxiv.org/html/2606.18075#S4.T1), different categories of baseline methods exhibit substantial variations in static QA performance\. LLM inference\-only approaches \(Zero\-shot and CoT\) perform poorly across all datasets, with accuracies significantly lower than retrieval\-based methods\. This indicates that relying solely on the parametric knowledge of LLMs is insufficient for factual and reasoning\-intensive tasks\.
Among context\-level methods, VanillaRAG, RAPTOR, and EraRAG achieve comparable performance on factual accuracy and multi\-hop reasoning datasets\. This suggests that the advantage of high\-order community structures is not always benificial, and the effectiveness largely depends on whether multi\-hop content is aggregated into the same community during the offline indexing stage\. On the QuALITY dataset, however, RAPTOR and EraRAG exhibit clear advantages, with EraRAG achieving the best overall performance\. We attribute this to EraRAG’s retrieval strategy, which enforces the selection of one non\-leaf node and four leaf nodes, thereby reducing irrelevant information and facilitating semantic understanding by the LLM\. Overall, RAPTOR slightly outperforms RAPTOR\-K and EraRAG, mainly due to its use of a superior but expensive clustering method, though the improvements are marginal\.
For relation\-level methods, the LightRAG variants demonstrate divergent strengths\. L\-LightRAG achieves the highest recall \(39\.97%\) on PopQA, showing that relation\-aware retrieval can effectively support LLMs in factual accuracy tasks\. In contrast, G\-LightRAG and H\-LightRAG perform relatively poorly, likely because global higher\-level concepts are less suited for highly specific queries\. ArchRAG and HiRAG, as improved versions of L\-GraphRAG, refines reasoning path generation by leveraging semantically similar summary entities\. HiRAG achieves better performance than L\-GraphRAG in multi\-hop reasoning, though still leaving room for improvement\. HippoRAG and its improved version, HippoRAG2, improving retrieval performance through relations, exhibit competitive results on tasks such as MultiHop\-RAG and PopQA but consistently fall short of our method in terms of overall accuracy\.
Overall,HyGRAGachieves the best or near\-best performance across all tasks\. On PopQA, it reaches 72\.34% accuracy and 43\.51% recall, outperforming other methods by 6\.2% and 8\.9%, respectively\. On the QuALITY dataset, it achieves the second\-best result\. In multi\-hop reasoning tasks,HyGRAGattains 65\.41% accuracy on MultiHop\-RAG, a 6\.0% improvement over the strongest baseline\(the metric in parentheses denotes results without explicitly prompting the model to state “Insufficient Information” when possible\), and shows stable advantages on MuSiQue and HotpotQA, with an average gain of approximately 11\.1%\. With the addition of inference augmentation \(\+Inference\), which add a rule in prompt to encourage the model to leverage our domain knowledge for reasoning, MuSiQue and HotpotQA further improve by 2\.0% and 0\.4%, respectively\. These results demonstrate thatHyGRAGconsistently delivers robustness and superior effectiveness across factual QA, reading comprehension, and multi\-hop reasoning tasks through a synergistic integration of relation and context\. In addition, we record in the Appendix[D](https://arxiv.org/html/2606.18075#A4)the results of further analyses using different dense retrievers and replacing the base LLM with Qwen3\-8B\(Yanget al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib115)\)\. The detailed results are shown in Table[10](https://arxiv.org/html/2606.18075#A3.T10)and Table[9](https://arxiv.org/html/2606.18075#A2.T9)\. We observe that across various embedding models,HyGRAGconsistently maintains stable and competitive performance, indicating that its advantages are not limited to a specific dense retriever\.
When replacing the base LLM with the newer model Qwen3\-8B\(Yanget al\.,[2025](https://arxiv.org/html/2606.18075#bib.bib115)\)\(using a consistentno\-thinkingmode\), the overall trend remains consistent with previous analyses\. Notably, we observe an unexpected performance drop of context\-aware methods on the MuSiQue and HotpotQA datasets after replacing the base LLM\. In contrast,HyGRAGmaintains competitive results, achieving up to 10% higher accuracy than other relation\-aware approaches\. This demonstrates its robustness and adaptability across different model architectures and reasoning paradigms\.
Figure 3\.Comparison of query efficiency\.
#### 4\.2\.2\.Efficiency ofHyGRAG
To answer the question RQ2, we compare the time cost and token usage ofHyGRAGwith those of other baseline methods\. As shown in Figure[3](https://arxiv.org/html/2606.18075#S4.F3),HyGRAGdemonstrates significant time and cost efficiency especially among relation\-aware methods for online queries\. However, compared with context\-aware methods, the cost ofHyGRAGis higher\.
Figure 4\.Corpus expansion performance\.
#### 4\.2\.3\.Robustness to Corpus Expansion
As RAG systems are increasingly deployed in real\-world applications, they must become more adaptable to scenarios of continuous learning where the retrieval corpus keeps expanding\. To answer the question RQ3, in order to evaluate the ability ofHyGRAGto handle incremental corpus insertions, we designed an experiment that divides the initial corpus into different proportions and then incrementally inserts the remaining corpus as updates\. We then tested the QA performance under varying proportions of incremental corpus, and the results are shown in Figure[4](https://arxiv.org/html/2606.18075#S4.F4)\.
Figure 5\.Corpus expansion indexing cost\.From the results, we observe that the larger the initial corpus proportion, the better the final QA performance\. This is mainly because, compared to static construction, incremental corpus insertion leads to a slight decrease in community quality\. However, this decrease is minor, around 1–2%, which shows the capability ofHyGRAGto handle continual learning scenarios effectively\. In addition, we measured the reconstruction time and token consumption required by various baseline methods when 20% of the corpus is inserted, as shown in Figure[5](https://arxiv.org/html/2606.18075#S4.F5)\. The results show that our proposedHyGRAGachieves favorable efficiency and competitive performance, indicating its strong practicality in dynamic corpus\.
### 4\.3\.Detailed Analysis
To answerRQ4, we conduct detailed analyses to investigate how each component ofHyGRAGcooperatively affects the overall performance of the system\.
#### 4\.3\.1\.Ablation Study
Table[2](https://arxiv.org/html/2606.18075#S4.T2)presents the results of our ablation studies conducted on multiple datasets\. To assess the individual contributions of each retrieval strategy, we progressively removed specific components and examined the resulting performance changes\.
As shown in Table[2](https://arxiv.org/html/2606.18075#S4.T2), removing the chunk\-level structure leads to the most pronounced degradation across all datasets, highlighting its crucial role in providing fundamental contextual grounding and factual support\. Similarly, eliminating entity and relation information causes a decline in accuracy, confirming the importance of relational reasoning for both factual and multi\-hop question answering tasks\. Interestingly, on the quality dataset, this removal yields a slight performance improvement\. We attribute this to the fact that, in reading comprehension tasks—particularly for smaller LLMs—explicit relational cues may sometimes mislead the model, whereas context\-aware semantic signals are more essential when the answer requires inference rather than direct textual matching\. In contrast, removing the community\-level representation results in a relatively mild yet stable decrease, indicating that community structures primarily contribute to aggregating query\-relevant higher\-order knowledge, providing semantic summarization, and optimizing relation retrieval\. Overall, these findings confirm thatHyGRAGbenefits substantially from the joint design of context\-aware and relation\-aware mechanisms\.
Table 2\.Results of ablation study\.Table 3\.Comparative performance and token cost against combining HiRAG and RAPTOR\.
#### 4\.3\.2\.Comparative Analysis
To further examine the complementary effects of thecontext\-awareandrelation\-awarecomponents, we conducted comparative experiments\. Specifically, we combined the retrieved chunks and communities from RAPTOR \(the best\-performing context\-aware method\) with the entities and relations from HiRAG \(the modified method of MS GraphRAG in multi\-hop reasoning path generation\)\. We then compared the results with and without applying the same system prompt used inHyGRAG\.
As shown in Table[3](https://arxiv.org/html/2606.18075#S4.T3), combining HiRAG and RAPTOR yields moderate improvements after applying theHyGRAG\-style prompt but still underperforms our model in both accuracy and recall\. Notably,HyGRAGachieves superior or comparable performance with substantially fewer tokens, demonstrating its efficiency in balancing reasoning depth and retrieval precision\. These results demonstrate that the integrated design ofHyGRAGis not a mere concatenation of relation and context representations, but rather an organic and unified framework that jointly models both aspects, enabling LLMs to more effectively leverage contextual and relational knowledge during reasoning\.
#### 4\.3\.3\.Case Study
To further illustrate the qualitative advantages of our framework, we present representative examples from the PopQA and MuSiQue datasets\. For each dataset, we compareHyGRAGwith the best\-performing baseline method\. Detailed case studies of model responses are provided in the Appendix[A](https://arxiv.org/html/2606.18075#A1)\. The results demonstrate thatHyGRAGeffectively integrates relevant entities, relations, contextual information, and community\-level knowledge to assist the model in constructing coherent reasoning chains and producing factually consistent answers\. For instance, in the MuSiQue multi\-hop reasoning task, baseline models such as LLightRAG and VanillaRAG fail to retrieve the correct passages through semantic search or to establish the connection between Jan Klapáč’s birthplace and the target entity Prague Castle\. In contrast,HyGRAGsuccessfully leverages relational and contextual information from both the chunk and community levels to capture the latent associations and infer the correct answer\. Similarly, in a factual query from PopQA, under a setting involving multiple person entities, the model accurately identifies the target entity Paul Walker and retrieves the supporting evidence from key relation that he was born in Kilwinning\. These cases collectively illustrate thatHyGRAGnot only achieves unified modeling of relational and contextual knowledge but also enhances entity disambiguation and relational consistency during reasoning, thereby improving both factual accuracy and interpretability\.
## 5\.Conclusion
This work introduces a unified framework with context and relation\-aware retrieval\.HyGRAGaddresses the problem that entity\-centric and chunk\-centric methods operate on representations anchored to original text\. By clustering hybrid nodes and generating communities integrating contextual and relational information, we create summary representations beyond source documents\. Our bi\-level retrieval enables access to insights neither approach could achieve independently\. Furthermore,HyGRAG’s attachment\-based update capability ensures efficient incorporation of new information\. Overall,HyGRAGshows 9\.7% average improvements in multi\-hop reasoning while maintaining efficiency for dynamic corpora\.
## Acknowledgments
This work is supported by NSFC \(No\. 62322606, No\. 62441605\)\.
## References
- @articlesun2025mlnc, author = Yifei Sun and Zemin Liu and Bryan Hooi and Yang Yang and Rizal Fathony and Jia Chen and Bingsheng He, … \(2025\)Multi\-label node classification with label influence propagation\.InInternational Conference on Learning Representations,Cited by:[Appendix C](https://arxiv.org/html/2606.18075#A3.p1.1)\.
- J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. Liu \(2024\)BGE m3\-embedding: multi\-lingual, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.External Links:2402\.03216,[Link](https://arxiv.org/abs/2402.03216)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.18075#S4.SS1.SSS3.p1.1)\.
- D\. Edge, H\. Trinh, N\. Cheng, J\. Bradley, A\. Chao, A\. Mody, S\. Truitt, D\. Metropolitansky, R\. O\. Ness, and J\. Larson \(2025\)From local to global: a graph rag approach to query\-focused summarization\.External Links:2404\.16130,[Link](https://arxiv.org/abs/2404.16130)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p4.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.6)\.
- W\. Fan, Y\. Ding, L\. Ning, S\. Wang, H\. Li, D\. Yin, T\. Chua, and Q\. Li \(2024\)A survey on rag meeting llms: towards retrieval\-augmented large language models\.External Links:2405\.06211,[Link](https://arxiv.org/abs/2405.06211)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1)\.
- Y\. Gao, Y\. Xiong, X\. Gao, K\. Jia, J\. Pan, Y\. Bi, Y\. Dai, J\. Sun, M\. Wang, and H\. Wang \(2024\)Retrieval\-augmented generation for large language models: a survey\.External Links:2312\.10997,[Link](https://arxiv.org/abs/2312.10997)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1),[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian,et al\.\(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.18075#S4.SS1.SSS3.p1.1)\.
- Z\. Guo, L\. Xia, Y\. Yu, T\. Ao, and C\. Huang \(2025\)LightRAG: simple and fast retrieval\-augmented generation\.External Links:2410\.05779,[Link](https://arxiv.org/abs/2410.05779)Cited by:[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p4.1),[§2\.2](https://arxiv.org/html/2606.18075#S2.SS2.p1.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.11)\.
- B\. J\. Gutiérrez, Y\. Shu, Y\. Gu, M\. Yasunaga, and Y\. Su \(2025\)HippoRAG: neurobiologically inspired long\-term memory for large language models\.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA\.External Links:ISBN 9798331314385Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1),[§1](https://arxiv.org/html/2606.18075#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p4.1),[§2\.2](https://arxiv.org/html/2606.18075#S2.SS2.p1.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.7)\.
- Y\. Hu, Z\. Lei, Z\. Zhang, B\. Pan, C\. Ling, and L\. Zhao \(2025\)GRAG: graph retrieval\-augmented generation\.External Links:2405\.16506,[Link](https://arxiv.org/abs/2405.16506)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1)\.
- H\. Huang, Y\. Huang, J\. Yang, Z\. Pan, Y\. Chen, K\. Ma, H\. Chen, and J\. Cheng \(2025a\)Retrieval\-augmented generation with hierarchical knowledge\.External Links:2503\.10150,[Link](https://arxiv.org/abs/2503.10150)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p4.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.9)\.
- L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. Liu \(2025b\)A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions\.ACM Transactions on Information Systems43\(2\),pp\. 1–55\.External Links:ISSN 1558\-2868,[Link](http://dx.doi.org/10.1145/3703155),[Document](https://dx.doi.org/10.1145/3703155)Cited by:[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p1.1)\.
- T\. N\. Kipf and M\. Welling \(2016\)Semi\-supervised classification with graph convolutional networks\.CoRRabs/1609\.02907\.External Links:[Link](http://arxiv.org/abs/1609.02907),1609\.02907Cited by:[Appendix C](https://arxiv.org/html/2606.18075#A3.p1.1)\.
- T\. Kojima, S\. S\. Gu, M\. Reid, Y\. Matsuo, and Y\. Iwasawa \(2022\)Large language models are zero\-shot reasoners\.InProceedings of the 36th International Conference on Neural Information Processing Systems,NIPS ’22,Red Hook, NY, USA\.External Links:ISBN 9781713871088Cited by:[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.2)\.
- W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. Stoica \(2023\)Efficient memory management for large language model serving with pagedattention\.InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles,Cited by:[§4\.1\.3](https://arxiv.org/html/2606.18075#S4.SS1.SSS3.p1.1)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. Kiela \(2020\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.InProceedings of the 34th International Conference on Neural Information Processing Systems,NIPS ’20,Red Hook, NY, USA\.External Links:ISBN 9781713829546Cited by:[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p2.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.3)\.
- P\. Lewis, E\. Perez, A\. Piktus, F\. Petroni, V\. Karpukhin, N\. Goyal, H\. Küttler, M\. Lewis, W\. Yih, T\. Rocktäschel, S\. Riedel, and D\. Kiela \(2021\)Retrieval\-augmented generation for knowledge\-intensive nlp tasks\.External Links:2005\.11401,[Link](https://arxiv.org/abs/2005.11401)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1)\.
- S\. Ma, C\. Xu, X\. Jiang, M\. Li, H\. Qu, C\. Yang, J\. Mao, and J\. Guo \(2025\)Think\-on\-graph 2\.0: deep and faithful large language model reasoning with knowledge\-guided retrieval augmented generation\.External Links:2407\.10805,[Link](https://arxiv.org/abs/2407.10805)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1)\.
- A\. Mallen, A\. Asai, V\. Zhong, R\. Das, D\. Khashabi, and H\. Hajishirzi \(2023\)When not to trust language models: investigating effectiveness of parametric and non\-parametric memories\.External Links:2212\.10511,[Link](https://arxiv.org/abs/2212.10511)Cited by:[§4\.1\.1](https://arxiv.org/html/2606.18075#S4.SS1.SSS1.p1.1)\.
- R\. Y\. Pang, A\. Parrish, N\. Joshi, N\. Nangia, J\. Phang, A\. Chen, V\. Padmakumar, J\. Ma, J\. Thompson, H\. He, and S\. R\. Bowman \(2022\)QuALITY: question answering with long input texts, yes\!\.External Links:2112\.08608,[Link](https://arxiv.org/abs/2112.08608)Cited by:[§4\.1\.1](https://arxiv.org/html/2606.18075#S4.SS1.SSS1.p1.1)\.
- B\. Peng, Y\. Zhu, Y\. Liu, X\. Bo, H\. Shi, C\. Hong, Y\. Zhang, and S\. Tang \(2024\)Graph retrieval\-augmented generation: a survey\.External Links:2408\.08921,[Link](https://arxiv.org/abs/2408.08921)Cited by:[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p2.1)\.
- P\. Sarthi, S\. Abdullah, A\. Tuli, S\. Khanna, A\. Goldie, and C\. D\. Manning \(2024\)RAPTOR: recursive abstractive processing for tree\-organized retrieval\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p3.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.4)\.
- Y\. Sun, H\. Deng, Y\. Yang, C\. Wang, J\. Xu, R\. Huang, L\. Cao, Y\. Wang, and L\. Chen \(2022\)Beyond homophily: structure\-aware path aggregation graph neural network\.InProceedings of the Thirty\-First International Joint Conference on Artificial Intelligence, IJCAI\-22,L\. D\. Raedt \(Ed\.\),pp\. 2233–2240\.Note:Main TrackExternal Links:[Document](https://dx.doi.org/10.24963/ijcai.2022/310),[Link](https://doi.org/10.24963/ijcai.2022/310)Cited by:[Appendix C](https://arxiv.org/html/2606.18075#A3.p1.1)\.
- Y\. Sun, Y\. Yang, X\. Feng, Z\. Wang, H\. Zhong, C\. Wang, and L\. Chen \(2025\)Handling feature heterogeneity with learnable graph patches\.InKnowledge Discovery and Data Mining,External Links:[Link](https://dl.acm.org/doi/10.1145/3690624.3709242)Cited by:[Appendix C](https://arxiv.org/html/2606.18075#A3.p1.1)\.
- Y\. Sun, Q\. Zhu, Y\. Yang, C\. Wang, T\. Fan, J\. Zhu, and L\. Chen \(2024\)Fine\-tuning graph neural networks by preserving graph generative patterns\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.38,pp\. 9053–9061\.Cited by:[Appendix C](https://arxiv.org/html/2606.18075#A3.p1.1)\.
- Y\. Tang and Y\. Yang \(2024\)MultiHop\-rag: benchmarking retrieval\-augmented generation for multi\-hop queries\.External Links:2401\.15391,[Link](https://arxiv.org/abs/2401.15391)Cited by:[§4\.1\.1](https://arxiv.org/html/2606.18075#S4.SS1.SSS1.p1.1)\.
- H\. Trivedi, N\. Balasubramanian, T\. Khot, and A\. Sabharwal \(2022\)MuSiQue: multihop questions via single\-hop question composition\.External Links:2108\.00573,[Link](https://arxiv.org/abs/2108.00573)Cited by:[§4\.1\.1](https://arxiv.org/html/2606.18075#S4.SS1.SSS1.p1.1)\.
- L\. Wang, N\. Yang, X\. Huang, L\. Yang, R\. Majumder, and F\. Wei \(2024\)Multilingual e5 text embeddings: a technical report\.External Links:2402\.05672,[Link](https://arxiv.org/abs/2402.05672)Cited by:[§D\.1](https://arxiv.org/html/2606.18075#A4.SS1.p1.1)\.
- S\. Wang, Y\. Fang, Y\. Zhou, X\. Liu, and Y\. Ma \(2025\)ArchRAG: attributed community\-based hierarchical retrieval\-augmented generation\.External Links:2502\.09891,[Link](https://arxiv.org/abs/2502.09891)Cited by:[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p4.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.10)\.
- Z\. Xiang, C\. Wu, Q\. Zhang, S\. Chen, Z\. Hong, X\. Huang, and J\. Su \(2025\)When to use graphs in rag: a comprehensive analysis for graph retrieval\-augmented generation\.External Links:2506\.05690,[Link](https://arxiv.org/abs/2506.05690)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui,et al\.\(2025\)Qwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§4\.2\.1](https://arxiv.org/html/2606.18075#S4.SS2.SSS1.p4.1),[§4\.2\.1](https://arxiv.org/html/2606.18075#S4.SS2.SSS1.p5.1)\.
- Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. W\. Cohen, R\. Salakhutdinov, and C\. D\. Manning \(2018\)HotpotQA: a dataset for diverse, explainable multi\-hop question answering\.External Links:1809\.09600,[Link](https://arxiv.org/abs/1809.09600)Cited by:[§4\.1\.1](https://arxiv.org/html/2606.18075#S4.SS1.SSS1.p1.1)\.
- F\. Zhang, Z\. Huang, Y\. Zhou, Q\. Guo, Z\. Li, W\. Luo, D\. Jiang, Y\. Fang, and X\. Zhou \(2025a\)EraRAG: efficient and incremental retrieval augmented generation for growing corpora\.External Links:2506\.20963,[Link](https://arxiv.org/abs/2506.20963)Cited by:[§1](https://arxiv.org/html/2606.18075#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.18075#S2.SS1.p3.1),[§2\.2](https://arxiv.org/html/2606.18075#S2.SS2.p1.1),[§4\.1\.2](https://arxiv.org/html/2606.18075#S4.SS1.SSS2.p1.1.5)\.
- Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin, F\. Huang, and J\. Zhou \(2025b\)Qwen3 embedding: advancing text embedding and reranking through foundation models\.External Links:2506\.05176,[Link](https://arxiv.org/abs/2506.05176)Cited by:[§D\.1](https://arxiv.org/html/2606.18075#A4.SS1.p1.1)\.
- Y\. Zhou, Y\. Su, Y\. Sun, S\. Wang, T\. Wang, R\. He, Y\. Zhang, S\. Liang, X\. Liu, Y\. Ma, and Y\. Fang \(2025\)In\-depth analysis of graph\-based rag in a unified framework\.External Links:2503\.04338,[Link](https://arxiv.org/abs/2503.04338)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.18075#S4.SS1.SSS3.p1.1)\.
Table 4\.Case Study on MuSiQue\.## Appendix ACase Study
Table 5\.Case Study on PopQA\.As shown in Table[4](https://arxiv.org/html/2606.18075#A0.T4), the baselines fail to establish the correct reasoning chain required to answer the question\. LLightRAG retrieves partial contextual evidence \(i\.e\., Jan Klapáč’s birthplace\) but lacks the relational linkage to the target entity, while VanillaRAG produces an entirely irrelevant response due to the absence of structured retrieval\. In contrast,HyGRAGsuccessfully integrates entity\-level, relational, and community\-level information\. It first identifies the entity Jan Klapáč and his birthplace \(Prague\), then connects this with the relation located in and the community summary mentioning Prague Castle\. This multi\-level reasoning enables the model to recover the implicit connection between the birthplace and the target entity, ultimately leading to the correct answer \(Prague Castle\)\. This case illustrates thatHyGRAGeffectively captures and organizes cross\-entity relations and contextual dependencies, allowing the LLM to perform more accurate and interpretable multi\-hop reasoning\.
Table[5](https://arxiv.org/html/2606.18075#A1.T5)presents a case from the PopQA dataset illustratingHyGRAG’s superiority in handling entity ambiguity\. The query “In what city was Paul Walker born?” involves multiple homonymous entities\. Baseline models like LLightRAG and RAPTOR confuse different individuals, producing inconsistent or irrelevant answers \(e\.g\., Balwyn, Colac, or Glendale\)\. In contrast,HyGRAGcombines entity–relation reasoning with chunk\-level to filter out unrelated entities and consolidate consistent evidence\. It correctly identifies Kilwinning and North Ayrshire—matching the gold answer—demonstrating that its relational and hierarchical design substantially improves factual accuracy and interpretability\.
Table 6\.Statistic ofHyGRAGusing Llama\-3\.1\-8B\-Instruct\.Table 7\.Dataset StatisticsTable 8\.Hyperparameter settings used in our experiments\.Algorithm 1HyGRAG Hierarchical Index Construction1:Corpus
DD, Max layers
LL, Split thresholds
Smin,SmaxS\_\{min\},S\_\{max\}
2:Hierarchical Index Tree
TindexT\_\{index\}
3:Phase 1: Hybrid Graph Construction
4:
C←Chunking\(D\)C\\leftarrow\\text\{Chunking\}\(D\)⊳\\trianglerightSplit documents into chunks
5:
Vc←CV^\{c\}\\leftarrow C;
Ec←∅E^\{c\}\\leftarrow\\emptyset
6:forpair
\(ci,cj\)\(c\_\{i\},c\_\{j\}\)in
CCdo
7:if
\|Entities\(ci\)∩Entities\(cj\)\|\>ϵ\|\\text\{Entities\}\(c\_\{i\}\)\\cap\\text\{Entities\}\(c\_\{j\}\)\|\>\\epsilonthen
8:
Ec←Ec∪\{\(ci,cj\)\}E^\{c\}\\leftarrow E^\{c\}\\cup\\\{\(c\_\{i\},c\_\{j\}\)\\\}⊳\\trianglerightEq\. 2
9:endif
10:endfor
11:
Ge←ExtractTriplets\(C\)G^\{e\}\\leftarrow\\text\{ExtractTriplets\}\(C\)⊳\\trianglerightLLM extracts\(h,r,t\)\(h,r,t\)
12:
Ece←\{\(e,c\)∣e∈Ve,c∈C\(e\)\}E^\{ce\}\\leftarrow\\\{\(e,c\)\\mid e\\in V^\{e\},c\\in C\(e\)\\\}⊳\\trianglerightEq\. 6
13:
Ghybrid←\(Vc∪Ve,Ec∪Ee∪Ece\)G\_\{hybrid\}\\leftarrow\(V^\{c\}\\cup V^\{e\},E^\{c\}\\cup E^\{e\}\\cup E^\{ce\}\)
14:Phase 2: Hierarchical Clustering
15:
Z←Cleora\(Ghybrid\)Z\\leftarrow\\text\{Cleora\}\(G\_\{hybrid\}\)⊳\\trianglerightStructure\-aware embeddings
16:
ℒ0←V\(Ghybrid\)\\mathcal\{L\}\_\{0\}\\leftarrow V\(G\_\{hybrid\}\)
17:for
l=1l=1to
LLdo
18:
𝒞l←LSH\_Clustering\(ℒl−1,Z\)\\mathcal\{C\}\_\{l\}\\leftarrow\\text\{LSH\\\_Clustering\}\(\\mathcal\{L\}\_\{l\-1\},Z\)⊳\\trianglerightEq\. 8
19:
ℒl←∅\\mathcal\{L\}\_\{l\}\\leftarrow\\emptyset
20:for
bucket∈𝒞lbucket\\in\\mathcal\{C\}\_\{l\}do
21:
B←AdjustSize\(bucket,Smin,Smax\)B\\leftarrow\\text\{AdjustSize\}\(bucket,S\_\{min\},S\_\{max\}\)
22:
tB←LLMsumm\(\{v∣v∈B\}\)t\_\{B\}\\leftarrow\\text\{LLM\}\_\{summ\}\(\\\{v\\mid v\\in B\\\}\)⊳\\trianglerightEq\. 10
23:
sB←LMemb\(tB\)s\_\{B\}\\leftarrow\\text\{LM\}\_\{emb\}\(t\_\{B\}\)⊳\\trianglerightEq\. 11
24:
Z\(sB\)←sBZ\(s\_\{B\}\)\\leftarrow s\_\{B\}⊳\\trianglerightUpdate embeddings for next layer
25:
ℒl←ℒl∪\{sB\}\\mathcal\{L\}\_\{l\}\\leftarrow\\mathcal\{L\}\_\{l\}\\cup\\\{s\_\{B\}\\\}
26:
Parent\(B\)←sB\\text\{Parent\}\(B\)\\leftarrow s\_\{B\}
27:endfor
28:endfor
29:return
Tindex=\{ℒ0,…,ℒL\}T\_\{index\}=\\\{\\mathcal\{L\}\_\{0\},\\dots,\\mathcal\{L\}\_\{L\}\\\}
Algorithm 2Context and Relation\-Aware Retrieval1:Query
qq, Index
TindexT\_\{index\}, Top\-
kkparam
kk
2:Generated Answer
AA
3:
qemb←LMemb\(q\)q\_\{emb\}\\leftarrow\\text\{LM\}\_\{emb\}\(q\)⊳\\trianglerightEncode query
4:Step 1: Context\-Aware Retrieval
5:
𝒮pool←\(⋃l=1Lℒl\)∪Vc\\mathcal\{S\}\_\{pool\}\\leftarrow\(\\bigcup\_\{l=1\}^\{L\}\\mathcal\{L\}\_\{l\}\)\\cup V^\{c\}
6:
𝒩ctx←TopK\(𝒮pool,qemb,k\)\\mathcal\{N\}\_\{ctx\}\\leftarrow\\text\{TopK\}\(\\mathcal\{S\}\_\{pool\},q\_\{emb\},k\)
7:
𝒫ret←\{n∈𝒩ctx∣nis Community\}\\mathcal\{P\}\_\{ret\}\\leftarrow\\\{n\\in\\mathcal\{N\}\_\{ctx\}\\mid n\\text\{ is Community\}\\\}⊳\\trianglerightSeparate for prompt
8:
Cret←\{n∈𝒩ctx∣nis Chunk\}C\_\{ret\}\\leftarrow\\\{n\\in\\mathcal\{N\}\_\{ctx\}\\mid n\\text\{ is Chunk\}\\\}
9:
Eretctx←TopK\(Ve,qemb,k\)E^\{ctx\}\_\{ret\}\\leftarrow\\text\{TopK\}\(V^\{e\},q\_\{emb\},k\)⊳\\trianglerightEntities retrieved separately
10:Step 2: Relation\-Aware Retrieval
11:
Eexpand←⋃C∈𝒫ret\{e∣\(e,C\)∈Ece\}E\_\{expand\}\\leftarrow\\bigcup\_\{C\\in\\mathcal\{P\}\_\{ret\}\}\\\{e\\mid\(e,C\)\\in E^\{ce\}\\\}
12:
Eall←Eretctx∪EexpandE\_\{all\}\\leftarrow E^\{ctx\}\_\{ret\}\\cup E\_\{expand\}⊳\\trianglerightIntegrate entities
13:
Rall←\{\(h,r,t\)∣h,t∈Eall\}R\_\{all\}\\leftarrow\\\{\(h,r,t\)\\mid h,t\\in E\_\{all\}\\\}⊳\\trianglerightExtract connected triplets
14:
Rret←TopK\(Rall,qemb,k\)R\_\{ret\}\\leftarrow\\text\{TopK\}\(R\_\{all\},q\_\{emb\},k\)⊳\\trianglerightFilter by similarity
15:Step 3: Generation
16:
Ctx←Prompt\(𝒫ret,Cret,Eretctx,Rret\)Ctx\\leftarrow\\text\{Prompt\}\(\\mathcal\{P\}\_\{ret\},C\_\{ret\},E^\{ctx\}\_\{ret\},R\_\{ret\}\)⊳\\trianglerightStructured Prompt
17:
A←LLMgen\(q,Ctx\)A\\leftarrow\\text\{LLM\}\_\{gen\}\(q,Ctx\)
18:return
AA
## Appendix BAlgorithm Details
The index construction process is illustrated in Algorithm[1](https://arxiv.org/html/2606.18075#alg1)\. Subsequently, the logic of the retrieval phase is detailed in Algorithm[2](https://arxiv.org/html/2606.18075#alg2)\.HyGRAGseparates its operational logic into two distinct phases: hierarchical index construction and dual\-aware retrieval\. Algorithm[1](https://arxiv.org/html/2606.18075#alg1)delineates the creation of a hybrid graph that bridges unstructured chunks with structured entity\-relations, followed by a bottom\-up community abstraction process\. Specifically, it employs LSH\-based clustering on Cleora embeddings and generates LLM summaries creating emergent knowledge\. Algorithm[2](https://arxiv.org/html/2606.18075#alg2)details the retrieval strategy, which synergizes context with relational triplets to provide a comprehensive context for the final generation\. Specifically, it performs bi\-level retrieval across hierarchy layers while expanding entity coverage through communities for comprehensive context\.
Table 9\.Overall QA results\(Accuracy and Recall\) on static RAG query datasets using Qwen3\-8B\. “OOT” denotes results that could not be obtained within two days\.
## Appendix CExperiment Details
We show the hybrid graph statistics, dataset statistics and Hyperparameter settings in Table[6](https://arxiv.org/html/2606.18075#A1.T6), Table[7](https://arxiv.org/html/2606.18075#A1.T7)and Table[8](https://arxiv.org/html/2606.18075#A1.T8)\. Exploring GNN\(Kipf and Welling,[2016](https://arxiv.org/html/2606.18075#bib.bib136); Sunet al\.,[2022](https://arxiv.org/html/2606.18075#bib.bib135),[2024](https://arxiv.org/html/2606.18075#bib.bib134),[2025](https://arxiv.org/html/2606.18075#bib.bib133); @articlesun2025mlnc, author = Yifei Sun and Zemin Liu and Bryan Hooi and Yang Yang and Rizal Fathony and Jia Chen and Bingsheng He, …,[2025](https://arxiv.org/html/2606.18075#bib.bib132)\)to learn more expressive representations over heterogeneous structures is a future work\. Regarding efficiency \(on Musique\),HyGRAG’s offline stage \(23,042s\) is 21% faster than the efficient GraphRAG representative \(LightRAG\)\. Online stage \(10\.7s\) is 7% faster than all graph\-based baselines\. During the process, it uses 9\.97GB memory and 2\.50GB GPU \(for embedding\-model only\)\. Clustering takes 31s \(GMM\-clustering in RAPTOR took 10,965s\)\. Summarization stage requires 1 hour, about 15% of overall indexing time\. Results showHyGRAGis efficient in both offline and online phases\. Regarding hallucinations in LLM\-generated community summaries, due to inherent limitations of LLM paradigm, using LLMs inevitably introduces certain probability of hallucinations \(errors\)\. However this issue can be mitigated as LLMs become more powerful\. Our framework allows replacing the LLM component, and even lightweight models still improve performance\. We manually evaluate 100 sampled summaries on Musique and found 17% hallucination rate, mostly due to added external knowledge\. After switching to a stronger model \(gpt\-4o\-mini\) the rate dropped to 7%, indicating improved summary quality\.
Table 10\.Performance comparison of different embedding models under Vanilla andHyGRAGsettings\.
## Appendix DExtended Experiments
### D\.1\.Robustness in Dense Retrievers
We experimented with a range of SOTA embedding models\. As shown in Table[10](https://arxiv.org/html/2606.18075#A3.T10),HyGRAGconsistently outperforms the corresponding vanilla dense retrieval baseline across all embedding models\. Moreover, the performance ofHyGRAGimproves progressively as the quality of the underlying embedding model increases, demonstrating its strong scalability and sensitivity to embedding fidelity\. In our experiments, we specifically evaluatedHyGRAGusing Qwen3\-Embedding\-0\.6B\(Zhanget al\.,[2025b](https://arxiv.org/html/2606.18075#bib.bib129)\), text\-embedding\-v3\(Zhanget al\.,[2025b](https://arxiv.org/html/2606.18075#bib.bib129)\), and multilingual\-e5\-large\-instruct\(Wanget al\.,[2024](https://arxiv.org/html/2606.18075#bib.bib130)\)as the underlying dense retrievers\. This demonstrates thatHyGRAG’s retrieval framework effectively amplifies embedding quality, creating synergistic gains beyond vanilla dense retrieval approaches\.
### D\.2\.Robustness in Base LLM
We present the QA performance of various methods using the base LLM Qwen3\-8B, as summarized in Table[9](https://arxiv.org/html/2606.18075#A2.T9)\. When employing Qwen3\-8B for both indexing and question answering,HyGRAGconsistently achieves competitive accuracy and recall across most datasets, maintaining leading performance on MuSiQue and HotpotQA, with accuracy gains up to 21\.36% over other methods\.
For context\-level methods, VanillaRAG, RAPTOR, and EraRAG show slight improvements compared to their performance with the previous base LLM\. However, they exhibit noticeable drops on the MuSiQue and HotpotQA datasets, indicating that their performance is sensitive to changes in the base LLM\. Similarly, the relation\-level method HippoRAG2 also shows performance declines on these datasets, and its results on QuALITY remain suboptimal\.
In contrast,HyGRAGmaintains stable and superior performance across all tasks, demonstrating robustness and adaptability to different LLM architectures\. This suggests thatHyGRAG’s synergistic integration of relation\- and context\-level information provides resilience against base model changes, while other approaches are more sensitive to variations in the underlying LLM\.
## Appendix EPrompts Used forHyGRAG
The Answer Generation Prompt is designed to synthesize a comprehensive and evidence\-grounded answer to a user’s query from a structured, multi\-source context\. The process begins by presenting the model with a rich contextual payload retrieved from a hierarchical knowledge system, which includes: hierarchical community summaries with similarity scores, a ranked list of the most relevant entities, key entity relationships structured as triplets, and the content of the most relevant source documents\. Subsequently, the prompt provides a set of explicit operational rules to constrain the generation process\. The framework mandates that the model must report any informational gaps by stating ”Insufficient information” and is forbidden from fabricating information not present in the context, ensuring a high degree of fidelity to the source material\.
The Community Summarization Prompt outlines a framework for distilling a collection of content, referred to as a knowledge community, into a structured and semantically dense summary\. The prompt’s primary objective is to generate a summary specifically tailored for the downstream task of creating high\-quality semantic embeddings\. It guides the model by defining a clear, four\-part structure for the output, requiring the summary to comprehensively cover: 1\) key themes and topics, 2\) important entities and their roles, 3\) relationships and connections between these entities, and 4\) the overall context and significance of the information\. By specifying both a word limit and the explicit need for rich semantic content, the prompt ensures the generation of a concise yet information\-rich text that captures the core essence of the knowledge community\.Similar Articles
ContextRAG: Extraction-Free Hierarchical Graph Construction for Retrieval-Augmented Generation
ContextRAG introduces an extraction-free method for constructing hierarchical graph indices for retrieval-augmented generation, using Residual-Quantization K-Means and Formal Concept Analysis to reduce LLM calls and tokens by orders of magnitude while maintaining competitive F1 scores on multi-hop questions.
HyCE-RAG: Hypergraph Chain-of-Evidence Retrieval-Augmented Generation for Explainable Multi-hop Question Answering
HyCE-RAG is a novel hypergraph-based retrieval-augmented generation framework for multi-hop question answering that constructs explicit evidence chains via confidence-aware heuristic search, outperforming standard RAG and graph-based RAG methods in accuracy, relevance, and faithfulness.
LightRAG: Simple and Fast Retrieval-Augmented Generation
The article introduces LightRAG, an open-source framework that enhances Retrieval-Augmented Generation by integrating graph structures for improved contextual awareness and efficient information retrieval.
Text-Graph Synergy: A Bidirectional Verification and Completion Framework for RAG
This paper introduces TGS-RAG, a bidirectional verification and completion framework that synergizes text-based and graph-based Retrieval-Augmented Generation to improve multi-hop reasoning accuracy.
FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
A novel framework called FAIR GraphRAG integrates FAIR Digital Objects with graph-based retrieval to enhance retrieval-augmented generation for semantic data analysis, improving question answering accuracy and adherence to FAIR principles, demonstrated on a biomedical dataset.