Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering
Summary
This paper benchmarks Knowledge Graph-based augmentation against Retrieval Augmented Generation for culturally specific question answering, demonstrating significant error reduction on Latin American culture data using the LatamQA dataset.
View Cached Full Text
Cached at: 09/17/26, 09:16 AM
# Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering Source: [https://arxiv.org/html/2609.18317](https://arxiv.org/html/2609.18317) Pablo PoulenardAffiliation:DCC, Universidad de Chile, Santiago, ChileAffiliation:École Polytechnique, Palaiseau, FranceYannis KarmimAffiliation:DCC, Universidad de Chile, Santiago, ChileAffiliation:Inria, Almanach, Paris, FranceAffiliation:Inria Chile, Santiago, ChileValentin BarrièreAffiliation:DCC, Universidad de Chile, Santiago, ChileAffiliation:CENIA, Macul, ChileCorrespondence:[pablo\.poulenard@polytechnique\.edu](mailto:[email protected]) ###### Abstract Large language models \(LLMs\) suffer from a long\-tail deficit: culturally specific facts, particularly those concerning underrepresented regions such as Latin America, appear too rarely in pretraining corpora to be reliably memorized\. Retrieval\-Augmented Generation \(RAG\) addresses this by grounding generation in external text, but structured alternatives such as Knowledge Graphs \(KGs\) offer tighter control over what enters the context, along with potential gains in explainability and updatability\. We benchmark Graph\-RAG against standard RAG on LatamQA, a culturally grounded multiple\-choice dataset spanning eight thematic categories\. The graphs are built end\-to\-end from Wikipedia articles with KGGen, a recent open\-domain extractor, without manual curation in our main setting\. G\-Retriever is competitive with RAG and reduces the error of the base LLM by 72% with a standard KG and 78% with a benchmark\-aware variant, the gap to RAG narrowing further as the graph is oriented toward task\-relevant content\. The trained projection transfers zero\-shot to Portuguese without target\-language fine\-tuning, indicating multilingual reach\. Our code is available[here](https://github.com/Payblito/KG-vs-RAG-QA)\. ## 1Introduction and Related Work LLMs acquire factual knowledge as a by\-product of next\-token prediction, so recall reliability scales with pretraining frequency[Mallen et al\. \(2023\)](https://arxiv.org/html/2609.18317#bib.bib18)\. Culturally specific facts, and especially those pertaining to underrepresented regions, appear too rarely to be memorised reliably, and the gap manifests as confabulation rather than abstention\. Latin American culture is a canonical instance: despite Spanish being a high\-resource language, state\-of\-the\-art LLMs answer questions about Iberian Spanish culture substantially more accurately than equivalent questions about Latin American culture[Karmim et al\. \(2026\)](https://arxiv.org/html/2609.18317#bib.bib3)\. Parametric adaptation does not close this gap: continued pretraining on a new distribution induces catastrophic forgetting[Yang et al\. \(2026\)](https://arxiv.org/html/2609.18317#bib.bib21), and supervised fine\-tuning is superseded by retrieval methods at scale[Ovadia et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib12)\. Retrieval\-Augmented Generation \(RAG\) addresses these limitations by grounding generation in an external text store without modifying model weights, and is highly effective in low\-frequency factual regimes[Mallen et al\. \(2023\)](https://arxiv.org/html/2609.18317#bib.bib18)\. Its main drawback is token cost: retrieved passages fill the context window, and the practitioner has little control over what enters the prompt\. A complementary line of work targets not what is retrieved but how it is reasoned over:[Ranaldi et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib1)improve multilingual RAG by having the model compare and reconcile heterogeneous retrieved passages through dialectic argumentation\. Knowledge Graphs \(KGs\) offer a structured alternative: explicit relational triples are compact, easy to inspect, update, and trace back to sources\. Automatic KG construction at scale has recently become tractable with KGGen[Mo et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib8), which converts raw text into triples via structured LLM prompting and reduces graph sparsity through entity and relation resolution\. G\-Retriever[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15)integrates a GNN\-based soft prompt with PCST subgraph retrieval and is the leading graph\-augmented LLM pipeline\. Its evaluations use ExplaGraphs[Saha et al\. \(2021\)](https://arxiv.org/html/2609.18317#bib.bib7), SceneGraphs[Hudson and Manning \(2019\)](https://arxiv.org/html/2609.18317#bib.bib6), and WebQSP[Yih et al\. \(2016\)](https://arxiv.org/html/2609.18317#bib.bib4): datasets that pair each question with a small, clean, dedicated graph and require multi\-hop reasoning, conditions structurally favorable to graph\-based methods\. We evaluate on a deliberately harder regime: a single large, noisy KG per thematic category with single\-hop questions, so that RAG is naturally advantaged and the true cost of converting text into triples can be measured\. Figure 1:Overview of our method\. From 5,848 Spanish Wikipedia articles, KGGen[Mo et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib8)extracts one schema\-free KG per thematic domain\. For a LatamQA[Karmim et al\. \(2026\)](https://arxiv.org/html/2609.18317#bib.bib3)question, G\-Retriever[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15)selects a subgraph via PCST and prepends its encoding<G\>to a frozen LLM; the same triples serve as text \(DD\) for our RAG and Top\-kkbaselines\.We benchmark Graph\-RAG against standard RAG on LatamQA[Karmim et al\. \(2026\)](https://arxiv.org/html/2609.18317#bib.bib3), a culturally grounded MCQ dataset spanning eight Latin American thematic categories\. Our contributions are:\(i\)a large\-scale application of the recent open\-domain extractor KGGen[Mo et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib8)to 5,848 Wikipedia articles, evaluated on downstream QA rather than intrinsic extraction metrics;\(ii\)a systematic comparison of RAG, Top\-kktriples, and G\-Retriever[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15);\(iii\)ablation studies isolating the contributions of graph structure and extraction quality; and\(iv\)zero\-shot multilingual transfer of the trained projection to Portuguese\.[Figure1](https://arxiv.org/html/2609.18317#S1.F1)present an overview of our proposed pipeline\. ## 2Proposed experimental pipeline Our setup combines a culturally grounded QA benchmark, KGs extracted from the articles it is built from, and a small LM to answer the questions\. #### Dataset We build upon LatamQA[Karmim et al\. \(2026\)](https://arxiv.org/html/2609.18317#bib.bib3), a multiple\-choice benchmark of culturally grounded factual knowledge extracted from Wikipedia articles on the cultures of Latin American countries\. We define eight thematic categories from the Wikipedia ontology \(Musica,Literatura,Cinema, …, see[Table1](https://arxiv.org/html/2609.18317#S2.T1)\), scrape each country\-theme mother category \(e\.g\.Gastronomía de Chile\) recursively up to depth two, and intersect the result with LatamQA, yielding 5,848 articles\. Each article supports exactly one question, with one correct answer and three distractors grounded in its content, so no question requires composition across articles\. #### Graph construction We rely on KGGen[Mo et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib8), a recent extractor that predicts entities and relations from raw text via structured LLM prompting, then resolves duplicates by iterative clustering \(prompts in Appendix[B](https://arxiv.org/html/2609.18317#A2)\)\. It has so far been evaluated only on English corpora of at most 5M tokens; we apply it to 50M characters of Spanish Wikipedia, a regime in which entity resolution dominates runtime\. Articles are chunked into 5,000\-character segments and per\-article graphs are aggregated by category with entity and edge resolution, yieldingone unified KG per category, linked with a central node in one global KG\. Each triple is traced through the pipeline, providing provenance used as retrieval ground truth\. We additionally construct abenchmark\-awarevariant per category by injecting entity and relation hints derived from LatamQA questions into the extraction prompt, steering the KG toward task\-relevant content\. Mistral Small 3\.2111Mistral\-Small\-3\.2\-24B\-Instruct\-2506is the KGGen backbone\. Table 1:Statistics of the eight thematic KGs, after entity and edge resolution #### Language Model Our main experiments useQwen2\.5\-3B\-Instruct[Qwen et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib9)\. Measuring the contribution of external knowledge requires a model that has not already memorized the target facts, since strong zero\-shot performance would leave little headroom to attribute retrieval gains\. At 60\.17 zero\-shot accuracy, parametric knowledge does not saturate the benchmark\. #### Query encoding for retrieval All four systems share the same encoder and query format\.jinaai/jina\-embeddings\-v3[Sturua et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib11), a multilingual encoder with an 8,192\-token context window, is used throughout, with its task\-specific LoRA adapters for queries and passages \(details in Appendix[A\.1](https://arxiv.org/html/2609.18317#A1.SS1)\)\. The query concatenates the question with its four options: this symmetrically enriches the lexical and semantic signal available to the retriever while preserving the integrity of the task, since no option is privileged at retrieval time\.[Section3](https://arxiv.org/html/2609.18317#S3)describes how each system uses this signal to select context\. ## 3Compared retrieval methods All systems below use the same language model and embedding model described in[Section2](https://arxiv.org/html/2609.18317#S2), and differ only in the context they retrieve\. RAG operates on raw text and serves as our reference point; the three graph\-based systems consume the KG with increasing use of its structure, from an unordered set of triples to a trained subgraph encoder\. #### RAG Each article is segmented into chunks of at most 512 tokens with an overlap of 64 tokens, encoded with the same model as the queries\. Thek=5k=5chunks with the highest similarity to the query are passed to the reader\. Retrieval is restricted to the articles of the corresponding category, matching the scope of the KG used by the graph\-based methods\. These values were selected by grid search; the full sweep is reported in Appendix[A\.2](https://arxiv.org/html/2609.18317#A1.SS2)\. #### Top\-kktriples This training\-free baseline scores every triple independently: each\(s,r,o\)\(s,r,o\)is encoded, ranked by cosine similarity to the query, and the top\-kktriples are verbalized into the prompt\. This follows the similarity\-based filtering of KAPING[Baek et al\. \(2023\)](https://arxiv.org/html/2609.18317#bib.bib5)but omits its entity\-linking stage, which restricts candidates to the one\-hop neighborhood of the question entities: since our graph comes from open extraction rather than a canonical knowledge base, no exact correspondence between question and graph entities is guaranteed\. The graph is treated as an unordered set of triples, making this a natural lower bound for the structure\-aware methods below\. #### G\-Retriever G\-Retriever[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15)first extracts a subgraph by solving a Prize\-Collecting Steiner Tree over the KG, using query similarity to edges and vertices as node prizes and edge costs\. A graph encoder222We use a Graph Transformer[Yun et al\. \(2019\)](https://arxiv.org/html/2609.18317#bib.bib14)rather than the Graph Attention Network[Veličković et al\. \(2018\)](https://arxiv.org/html/2609.18317#bib.bib13)of the original work\.maps this subgraph to a single vector, which an MLP projects into the LM embedding space as a soft prompt\. Both modules are trained end\-to\-end on a training split to produce the correct answer\. In[Section4\.2](https://arxiv.org/html/2609.18317#S4.SS2)we also evaluate a variant replacing the graph encoder by mean pooling over node embeddings, where topology determines which nodes are retrieved but is never encoded\. ## 4Results and Analysis We first compare all four systems on LatamQA[Karmim et al\. \(2026\)](https://arxiv.org/html/2609.18317#bib.bib3)\([Section4\.1](https://arxiv.org/html/2609.18317#S4.SS1.SSS0.Px1)\), then isolate the contribution of each component of G\-Retriever\([Section4\.2](https://arxiv.org/html/2609.18317#S4.SS2)\), and finally test whether the trained projection transfers to an unseen language \([Section4\.3](https://arxiv.org/html/2609.18317#S4.SS3)\)\. ### 4\.1Main comparison #### RAG vs Top\-kktriples vs G\-Retriever Table[2](https://arxiv.org/html/2609.18317#S4.T2)compares all methods against zero\-shot \(60\.17\)\. All retrieval methods substantially outperform the unaugmented model\. The training\-free Top\-kktriples baseline \(\+15\.4 pp\) confirms that even unordered triples carry discriminative signal, yet it remains well behind G\-Retriever and RAG\. RAG is the strongest overall \(93\.38 vs\. 89\.71\), a gap attributable to information loss at triple extraction\. Per\-category results reveal substantial heterogeneity; notably, G\-Retriever overtakes RAG on Gastronomía \(93\.04 vs\. 87\.37\), where relational abstraction filters out lexically crowded dense\-retrieval noise\. Table 2:Accuracy \(%\) per category on LatamQA ### 4\.2Ablation Studies Table[3](https://arxiv.org/html/2609.18317#S4.T3)reports ablation results\. #### Graph structure Removing the trained projection \(PCST as plain text, 73\.10\) falls below Top\-kktriples, replicating the corresponding ablation of[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15): the trained continuous prefix is the essential component\. Interestingly, a linear projection matches or exceeds the graph encoder \(88\.90 vs\. 86\.13 without the textualized graph, 89\.50 vs\. 89\.71 with it\)\. We attribute this to the single\-hop nature of the task: what the trained module supplies is a task\-adapted continuous summary of the retrieved node and edge embeddings, functionally a form of prompt tuning[Lester et al\. \(2021\)](https://arxiv.org/html/2609.18317#bib.bib17);[Li and Liang \(2021\)](https://arxiv.org/html/2609.18317#bib.bib16)conditioned on retrieved content, and message passing is an unnecessarily expressive way of producing it here\. #### Benchmark\-aware extraction Steering extraction with question\-derived entity hints improves both Top\-kktriples \(\+2\.3 pp\) and G\-Retriever \(\+1\.7 pp, reaching 91\.41\), confirming that standard extraction discards task\-relevant content\. The persistent gap with RAG \(91\.41 vs\. 93\.38\) shows that extraction, not retrieval, is the main bottleneck\. Table 3:Accuracy under different KGGen selection methods, graph encoders, and benchmark settings\. GraphEnc denotes the trained graph encoder of G\-Retriever, Linear denotes mean\-pooled node embeddings with a single projection\. Bench KGGen refers to a KG constructed with LatamQA\-specific relations/entities\. ### 4\.3Multilingual transfer The five checkpoints from the Spanish cross\-validation folds are applied without further training to a Portuguese KG built from 397 Literatura articles, with questions, triples and generation all in Portuguese\. Since the projection aligns a pooled subgraph representation with the reader’s input space rather than modeling any particular language, and since the encoder is multilingual[Sturua et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib11), the mapping should be largely language\-agnostic\. G\-Retriever reaches 91\.18, above its 89\.44 on Spanish Literatura, while the zero\-shot baseline drops from 60\.06 to 53\.40 \(Table[4](https://arxiv.org/html/2609.18317#S4.T4)\)\. A single trained projection can therefore serve several languages, provided the retrieval backbone is itself multilingual\. Table 4:Zero\-shot cross\-lingual transfer on the Portuguese Literatura graph \(397 articles\)\. ### 4\.4Inference efficiency Graph\-RAG is also lighter at inference: subgraph retrieval returns a compact set of triples rather than full passages, reducing the average context from28142814to875875tokens \(Table[5](https://arxiv.org/html/2609.18317#S4.T5)\)\. The cost moves offline: KG extraction cost 72\.26 EUR in API calls and training the graph encoder took∼\\sim9,942 s per fold \(Appendix[E](https://arxiv.org/html/2609.18317#A5)\)\. Table 5:Inference\-time cost comparison\. ## 5Conclusion We benchmarked Graph\-RAG against RAG on a culturally grounded, single\-hop MCQ dataset spanning eight Latin American thematic categories\. G\-Retriever reduces base LLM error by 74% on a 3\.2×\\timesshorter context, and a benchmark\-aware KG narrows the residual gap with RAG to 2 pp, confirming extraction quality rather than retrieval as the bottleneck; the projection also transfers zero\-shot to Portuguese\. Both LatamQA and our KGs derive from Spanish Wikipedia, whose coverage skews toward documented over orally transmitted knowledge, and addressing this bias will require participatory sources[Zhou et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib20)\. ## 6Limitations #### Evaluation format The four\-way multiple\-choice format admits a 25% chance floor and lets a system succeed by elimination rather than recall: a partially relevant triple may suffice to discard distractors without stating the answer itself[Balepur et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib2)\. Extraction\-induced information loss is therefore penalised only when it removes discriminative content, so the gap we report between flat retrieval and graph\-based augmentation is likely a lower bound on the true cost of converting text into triples\. #### Single reader All results use one small reader, Qwen2\.5\-3B\-Instruct, chosen to leave headroom for retrieval to matter\. The G\-Retriever projection is trained to align with this specific frozen reader, and we do not test whether the learned mapping transfers across backbones\. #### Narrow cross\-lingual evidence Our multilingual transfer experiment covers a single category \(Literatura\) and a language typologically close to Spanish, on a graph an order of magnitude smaller than its Spanish counterpart, which likely eases retrieval\. Broader claims would require matched\-size graphs and typologically distant, lower\-resource languages[Hu et al\. \(2020\)](https://arxiv.org/html/2609.18317#bib.bib19)\. #### Coverage bias Both LatamQA and our graphs are derived from Spanish Wikipedia, whose category distribution is highly uneven \(1,306 articles for Música against 140 for Artesanía\) and reflects editorial attention rather than cultural salience\. Domains transmitted orally or through artisanal practice are underrepresented, so our benchmark measures Latin American culture as Wikipedia records it\. ## References - J\. Baek, A\. F\. Aji, and A\. SaffariKnowledge\-augmented language model prompting for zero\-shot knowledge graph question answering\.External Links:2306\.04136,[Link](https://arxiv.org/abs/2306.04136)Cited by:[§3](https://arxiv.org/html/2609.18317#S3.SS0.SSS0.Px2.p1.1)\. - Balepuret al\.\(2024\)N\. Balepur, S\. Palta, and R\. RudingerIt’s not easy being wrong: large language models struggle with process of elimination reasoning\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 10143–10166\.External Links:[Link](https://aclanthology.org/2024.findings-acl.604/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.604)Cited by:[§6](https://arxiv.org/html/2609.18317#S6.SS0.SSS0.Px1.p1.1)\. - Heet al\.\(2024\)B\. He, H\. Li, Y\. K\. Jang, M\. Jia, X\. Cao, A\. Shah, A\. Shrivastava, and S\. LimMA\-LMM: Memory\-Augmented Large Multimodal Model for Long\-Term Video Understanding\.InCVPR,pp\. 13504–13514\.External Links:[Link](http://arxiv.org/abs/2404.05726)Cited by:[1st item](https://arxiv.org/html/2609.18317#A1.I1.i1.p1.1),[2nd item](https://arxiv.org/html/2609.18317#A1.I1.i2.p1.1),[3rd item](https://arxiv.org/html/2609.18317#A1.I1.i3.p1.1),[Figure 1](https://arxiv.org/html/2609.18317#S1.F1),[§1](https://arxiv.org/html/2609.18317#S1.p3.1),[§1](https://arxiv.org/html/2609.18317#S1.p4.1),[§3](https://arxiv.org/html/2609.18317#S3.SS0.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2609.18317#S4.SS2.SSS0.Px1.p1.1)\. - Huet al\.\(2020\)J\. Hu, S\. Ruder, A\. Siddhant, G\. Neubig, O\. Firat, and M\. JohnsonXTREME: A Massively Multilingual Multi\-task Benchmark for Evaluating Cross\-lingual Generalization\.External Links:[Link](http://arxiv.org/abs/2003.11080)Cited by:[§6](https://arxiv.org/html/2609.18317#S6.SS0.SSS0.Px3.p1.1)\. - Hudson and Manning \(2019\)D\. A\. Hudson and C\. D\. ManningGQA: a new dataset for real\-world visual reasoning and compositional question answering\.External Links:1902\.09506,[Link](https://arxiv.org/abs/1902.09506)Cited by:[§1](https://arxiv.org/html/2609.18317#S1.p3.1)\. - Karmimet al\.\(2026\)Y\. Karmim, R\. Pino, H\. Contreras, H\. Lira, S\. Cifuentes, S\. Escoffier, L\. Martí, D\. Seddah, and V\. BarrièreLeveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America\.arXiv\.External Links:2603\.10001,[Document](https://dx.doi.org/10.48550/arXiv.2603.10001)Cited by:[Figure 1](https://arxiv.org/html/2609.18317#S1.F1),[§1](https://arxiv.org/html/2609.18317#S1.p1.1),[§1](https://arxiv.org/html/2609.18317#S1.p4.1),[§2](https://arxiv.org/html/2609.18317#S2.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.18317#S4.p1.1)\. - Lesteret al\.\(2021\)B\. Lester, R\. Al\-Rfou, and N\. ConstantThe Power of Scale for Parameter\-Efficient Prompt Tuning\.External Links:[Link](http://arxiv.org/abs/2104.08691),[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.243)Cited by:[§4\.2](https://arxiv.org/html/2609.18317#S4.SS2.SSS0.Px1.p1.1)\. - Li and Liang \(2021\)X\. L\. Li and P\. LiangPrefix\-tuning: Optimizing continuous prompts for generation\.InACL\-IJCNLP 2021 \- 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Proceedings of the Conference,pp\. 4582–4597\.External Links:ISBN 9781954085527,[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.353)Cited by:[§4\.2](https://arxiv.org/html/2609.18317#S4.SS2.SSS0.Px1.p1.1)\. - Mallenet al\.\(2023\)A\. Mallen, A\. Asai, V\. Zhong, R\. Das, D\. Khashabi, and H\. HajishirziWhen Not to Trust Language Models: Investigating Effectiveness of Parametric and Non\-Parametric Memories\.Proceedings of the Annual Meeting of the Association for Computational Linguistics1\(Section 6\),pp\. 9802–9822\.External Links:ISBN 9781959429722,[Document](https://dx.doi.org/10.18653/v1/2023.acl-long.546),ISSN 0736587XCited by:[§1](https://arxiv.org/html/2609.18317#S1.p1.1),[§1](https://arxiv.org/html/2609.18317#S1.p2.1)\. - Moet al\.\(2025\)B\. Mo, K\. Yu, J\. Kazdan, J\. Cabezas, P\. Mpala, L\. Yu, C\. Cundy, C\. Kanatsoulis, and S\. KoyejoKGGen: Extracting Knowledge Graphs from Plain Text with Language Models\.External Links:2502\.09956,[Document](https://dx.doi.org/10.48550/arXiv.2502.09956)Cited by:[§B\.1](https://arxiv.org/html/2609.18317#A2.SS1.p1.1),[Appendix E](https://arxiv.org/html/2609.18317#A5.p2.1),[Figure 1](https://arxiv.org/html/2609.18317#S1.F1),[§1](https://arxiv.org/html/2609.18317#S1.p2.1),[§1](https://arxiv.org/html/2609.18317#S1.p4.1),[§2](https://arxiv.org/html/2609.18317#S2.SS0.SSS0.Px2.p1.1)\. - Ovadiaet al\.\(2024\)O\. Ovadia, M\. Brief, M\. Mishaeli, and O\. ElishaFine\-Tuning or Retrieval? Comparing Knowledge Injection in LLMs\.EMNLP 2024 \- 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference,pp\. 237–250\.External Links:ISBN 9798891761643,[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.15)Cited by:[§1](https://arxiv.org/html/2609.18317#S1.p1.1)\. - Qwenet al\.\(2025\)Qwen, A\. Yang, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Li, D\. Liu, F\. Huang, H\. Wei, H\. Lin, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Lin, K\. Dang, K\. Lu, K\. Bao, K\. Yang, L\. Yu, M\. Li, M\. Xue, P\. Zhang, Q\. Zhu, R\. Men, R\. Lin, T\. Li, T\. Tang, T\. Xia, X\. Ren, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Cui, Z\. Zhang, and Z\. QiuQwen2\.5 Technical Report\.arXiv\.External Links:2412\.15115,[Document](https://dx.doi.org/10.48550/arXiv.2412.15115)Cited by:[§2](https://arxiv.org/html/2609.18317#S2.SS0.SSS0.Px3.p1.1)\. - Ranaldiet al\.\(2025\)L\. Ranaldi, F\. Ranaldi, F\. M\. Zanzotto, B\. Haddow, and A\. BirchImproving multilingual retrieval\-augmented language models through dialectic reasoning argumentations\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 9064–9085\.External Links:[Link](https://aclanthology.org/2025.emnlp-main.461/),[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-main.461),ISBN 979\-8\-89176\-332\-6Cited by:[§1](https://arxiv.org/html/2609.18317#S1.p2.1)\. - Sahaet al\.\(2021\)S\. Saha, P\. Yadav, L\. Bauer, and M\. BansalExplaGraphs: an explanation graph generation task for structured commonsense reasoning\.External Links:2104\.07644,[Link](https://arxiv.org/abs/2104.07644)Cited by:[§1](https://arxiv.org/html/2609.18317#S1.p3.1)\. - Sturuaet al\.\(2024\)S\. Sturua, I\. Mohr, M\. K\. Akram, M\. Günther, B\. Wang, M\. Krimmel, F\. Wang, G\. Mastrapas, A\. Koukounas, N\. Wang, and H\. XiaoJina\-embeddings\-v3: Multilingual Embeddings With Task LoRA\.arXiv\.External Links:2409\.10173,[Document](https://dx.doi.org/10.48550/arXiv.2409.10173)Cited by:[§A\.1](https://arxiv.org/html/2609.18317#A1.SS1.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.18317#S2.SS0.SSS0.Px4.p1.1),[§4\.3](https://arxiv.org/html/2609.18317#S4.SS3.p1.1)\. - Veličkovićet al\.\(2018\)P\. Veličković, A\. Casanova, P\. Liò, G\. Cucurull, A\. Romero, and Y\. BengioGraph attention networks\.6th International Conference on Learning Representations, ICLR 2018 \- Conference Track Proceedings,pp\. 1–12\.External Links:[Document](https://dx.doi.org/10.1007/978-3-031-01587-8%5F7)Cited by:[footnote 2](https://arxiv.org/html/2609.18317#footnote2)\. - Wanget al\.\(2024\)L\. Wang, N\. Yang, X\. Huang, L\. Yang, R\. Majumder, and F\. WeiMultilingual E5 Text Embeddings: A Technical Report\.arXiv\.External Links:2402\.05672,[Document](https://dx.doi.org/10.48550/arXiv.2402.05672)Cited by:[§A\.1](https://arxiv.org/html/2609.18317#A1.SS1.SSS0.Px1.p1.1)\. - Yanget al\.\(2026\)Z\. Yang, Y\. Song, I\. Ahmed, and I\. HarrisFine\-Tuning vs\. RAG for Multi\-Hop Question Answering with Novel Knowledge\.InGEM 2026 @ ACL,pp\. 384–392\.External Links:[Link](http://arxiv.org/abs/2601.07054)Cited by:[§1](https://arxiv.org/html/2609.18317#S1.p1.1)\. - Yihet al\.\(2016\)W\. Yih, M\. Richardson, C\. Meek, M\. Chang, and J\. SuhThe Value of Semantic Parse Labeling for Knowledge Base Question Answering\.InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),K\. Erk and N\. A\. Smith \(Eds\.\),Berlin, Germany,pp\. 201–206\.External Links:[Document](https://dx.doi.org/10.18653/v1/P16-2033)Cited by:[§1](https://arxiv.org/html/2609.18317#S1.p3.1)\. - Yunet al\.\(2019\)S\. Yun, M\. Jeong, R\. Kim, J\. Kang, and H\. J\. KimGraph transformer networks\.Advances in Neural Information Processing Systems32\(NeurIPS\)\.External Links:ISSN 10495258Cited by:[2nd item](https://arxiv.org/html/2609.18317#A1.I1.i2.p1.1),[footnote 2](https://arxiv.org/html/2609.18317#footnote2)\. - Zhouet al\.\(2025\)R\. Zhou, G\. Wan, S\. Gabriel, S\. Li, A\. J\. Gates, M\. Sap, and T\. HartvigsenDisparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English\.External Links:[Link](http://arxiv.org/abs/2503.04099)Cited by:[§5](https://arxiv.org/html/2609.18317#S5.p1.1)\. ## Appendix AImplementation Details ### A\.1Embedding and Query #### Embedding model All retrieval components rely onjinaai/jina\-embeddings\-v3[Sturua et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib11), a state\-of\-the\-art multilingual encoder matching the Spanish of both the corpus and the questions\. Its 8,192\-token context window — against 512 for comparable encoders such asmultilingual\-e5\-large\-instruct[Wang et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib10)— is a deliberate design choice: it allows coarse retrieval units, up to entire articles, and thus lets us vary retrieval granularity while holding the encoder fixed\. Queries and Passages were encoded using the task\-specific LoRA adapters of the model\. #### Query construction For all retrieval methods, the query is the concatenation of the question and its four answer options, without any indication of which option is correct\. In a multiple\-choice setting, the question alone is often an underspecified retrieval cue: it may lack the named entities and surface forms that anchor the relevant subgraph, whereas these frequently appear in the options themselves\. Including all four options symmetrically enriches the lexical and semantic signal available to the retriever while preserving the integrity of the task, since no option is privileged at retrieval time\. This design also aligns the evaluation with the inference\-time setting of the reader, which observes the question and all options jointly\. ### A\.2Retrieval Method Hyperparameters Optimization We ran an exhaustive grid search for the three retrieval modes—TOP\-K TRIPLES, RAG and GRAPH\-RAG \(PCST\)—on a held\-out subset of 500 questions sampled uniformly at random from the benchmark\. For RAG we varied the chunk size \(32–512 tokens\) and the number of retrieved chunksk∈\{5,10,25\}k\\in\\\{5,10,25\\\}; for top\-kktriples, the number of retrieved triplesk∈\{3,…,100\}k\\in\\\{3,\\dots,100\\\}; for Graph\-RAG, the node and edge budgets and the PCST edge costce∈\{0\.01,0\.1,0\.5\}c\_\{e\}\\in\\\{0\.01,0\.1,0\.5\\\}\. For the Graph\-RAG grid we searched over the PCST retriever alone, without the trained GNN\-based soft\-prompting module, in the same configuration as the ablation of Section[4\.2](https://arxiv.org/html/2609.18317#S4.SS2)\. This choice is dictated by cost: a full grid would have required retraining the GNN for every cell, at roughly10410^\{4\}seconds per fold \(Table[5](https://arxiv.org/html/2609.18317#S4.T5)\)\. It rests on the assumption that the retrieval component and the trained projection module are approximately separable, i\.e\., that the ranking of subgraph budgets induced by retrieval quality is preserved when the projection module is added downstream\. ### A\.3G\-Retriever - •Subgraph Retrieval via PCST\.Following[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15), we first encode node and edge textual attributes with a pretrained language model and compute their cosine similarity to the query embedding, which serves as node prizes and edge costs\. We then solve the Prize\-Collecting Steiner Tree \(PCST\) problem to retrieve a connected subgraphS∗=\(V∗,E∗\)S^\{\*\}=\(V^\{\*\},E^\{\*\}\)maximizing query relevance while penalizing edge costs: S∗=argmax∑v∈VS⊆Gprize\(v\)−∑e∈Ecoste\.S^\{\*\}=\\arg\\max\_\{S\\subseteq G\}\\sum\_\{v\\in V\}\\text\{prize\}\(v\)\-\\sum\_\{e\\in E\}\\text\{cost\}\_\{e\}\.\(1\)We settop\_knodes=15top\\\_k\_\{nodes\}=15,top\_kedges=20top\\\_k\_\{edges\}=20andcoste=0\.5\\text\{cost\}\_\{e\}=0\.5, hyperparameters optimized via grid search\. - •Graph Encoder\.To model the structure of the retrieved subgraphS∗S^\{\*\}, we depart from the Graph Attention Network used in[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15)and instead employ a Graph Transformer[Yun et al\. \(2019\)](https://arxiv.org/html/2609.18317#bib.bib14), implemented via multi\-head Transformer convolution layers that incorporate edge features into the attention computation\. Our encoder stacksL=4L=4layers with 8 attention heads, residual connections, layer normalization and dropout \(p=0\.1p=0\.1\), operating on node and edge embeddings of dimension 1024\. Node representations are then aggregated into a single graph token via mean pooling:hg=POOL\(GraphTransformerϕ1\(S∗\)\)∈ℝdgh\_\{g\}=\\text\{POOL\}\(\\text\{GraphTransformer\}\_\{\\phi\_\{1\}\}\(S^\{\*\}\)\)\\in\\mathbb\{R\}^\{d\_\{g\}\}, withdg=1024d\_\{g\}=1024\. Each layer updates node representations through multi\-head attention over neighboring nodes, where attention coefficients are conditioned on both node and edge embeddings: xi′=W1xi\+∑j∈𝒩\(i\)αijW2xjx\_\{i\}^\{\\prime\}=W\_\{1\}x\_\{i\}\+\\sum\_\{j\\in\\mathcal\{N\}\(i\)\}\\alpha\_\{ij\}W\_\{2\}x\_\{j\}αij=softmaxj\(\(W3xi\)⊤\(W4xj\+W5eij\)d\)\.\\alpha\_\{ij\}=\\text\{softmax\}\_\{j\}\\left\(\\frac\{\(W\_\{3\}x\_\{i\}\)^\{\\top\}\(W\_\{4\}x\_\{j\}\+W\_\{5\}e\_\{ij\}\)\}\{\\sqrt\{d\}\}\\right\)\. - •Projection Layer and Prompt Tuning\.A multilayer perceptron aligns the graph token with the hidden space of the frozen LLM \(dl=1024d\_\{l\}=1024\):h^g=MLPϕ2\(hg\)∈ℝdl\\hat\{h\}\_\{g\}=\\text\{MLP\}\_\{\\phi\_\{2\}\}\(h\_\{g\}\)\\in\\mathbb\{R\}^\{d\_\{l\}\}\. This graph token acts as a soft prompt, prepended to the embeddings of the textualized subgraph and the query; while the LLM parametersθ\\thetaremain frozen, gradients flow throughh^g\\hat\{h\}\_\{g\}, enabling the optimization ofϕ1\\phi\_\{1\}andϕ2\\phi\_\{2\}by standard backpropagation[He et al\. \(2024\)](https://arxiv.org/html/2609.18317#bib.bib15)\. ### A\.4Cross\-Validation All the experiments were run on the full dataset using akk\-fold train\-val\-test cross\-validation\. ## Appendix BKG Generation ### B\.1Base pipeline KGGen[Mo et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib8)extracts entities and relations from raw text via structured LLM prompting\. Each Wikipedia article is segmented into chunks of 5,000 characters and processed independently, yielding a per\-article graph\. Article\-level graphs are then aggregated by category, and entity and edge resolution is applied at category scale so that each thematic category is represented by a single unified KG\. Extraction is carried out in Spanish using Mistral Small 3\.2 \(Mistral\-Small\-3\.2\-24B\-Instruct\-2506\) accessed via API\. All LLM calls are routed through LiteLLM\. Each triple is traced throughout the resolution process, so that for any given article we can recover the full set of triples originating from it in the final graph; this provenance information provides the ground truth against which retrieval is evaluated\. ### B\.2Benchmark\-aware knowledge graph For each category, we additionally construct abenchmark\-awareaugmented KG\. Given the questions and answer options of LatamQA — including distractors — we extract the entities and relations required to discriminate between the candidate answers using Mistral Small 3\.2\. The resulting entity and relation hints are injected into the extraction prompt, steering the KG towards elements relevant to the downstream task rather than arbitrary factual content\. Extraction otherwise follows the base pipeline, and entity and edge resolution is applied unchanged\. The procedure yields a second graph per category aligned with the question distribution while relying on the same source articles\. This graph is used as an upper\-bound diagnostic: it is not a deployable configuration, as it presupposes access to the evaluation questions at construction time\. ## Appendix CQA Prompts ### C\.1Vanilla Prompt The Vanilla prompt to answer the LatamQA MCQ is shown in Figure[2](https://arxiv.org/html/2609.18317#A3.F2)\. Answer the following multiple\-choice question based on your general knowledge\. Question:\{question\}Options: A\)\{options\[’A’\]\} B\)\{options\[’B’\]\} C\)\{options\[’C’\]\} D\)\{options\[’D’\]\}Answer \(single letter\):Figure 2:Prompt used for the LatamQA benchmark\. ### C\.2Triplet\-enhanced Prompt The prompt using context extracted from the graph is shown in Figure[3](https://arxiv.org/html/2609.18317#A3.F3)\. Here is some context that may help answer the following multiple\-choice question\. If the context is irrelevant or unclear, rely on your general knowledge\.Context: \{context\_block\}Question:\{question\}Options: A\)\{options\[’A’\]\} B\)\{options\[’B’\]\} C\)\{options\[’C’\]\} D\)\{options\[’D’\]\}Answer \(single letter\):Figure 3:Prompt used for the LatamQA benchmark\.context\_blockcontains the extracted triplets using the Top\-kktriplet or PCST\. ## Appendix DRetrieval Oracle Analysis To separate retrieval error from reading error, we restrict the candidate pool to the gold article for each question \(oracle condition\)\. Table[6](https://arxiv.org/html/2609.18317#A4.T6)shows that RAG benefits most \(\+3\.3%, 93\.38 to 96\.72\), and Top\-kktriples gains 2\.2%, whereas G\-Retriever gains only 0\.6% — within its cross\-fold dispersion\. G\-Retriever’s residual error thus stems from information lost during triple extraction, not from retrieval failure\. The gap with RAG widens under the oracle condition \(from 3\.7 to 6\.5%\), confirming that the ceiling of the graph representation is set by extraction, not retrieval\. Table 6:Full\-graph vs\. per\-article \(oracle\) retrieval, accuracy \(%\)\. ## Appendix EInference Cost Details The two pipelines distribute their cost very differently across the system lifecycle\. The text\-based RAG baseline concentrates its expense at inference time: retrieved passages are injected verbatim into the prompt, yielding an average context of2814\.32814\.3tokens, roughly3\.2×3\.2\\timeslonger than the874\.8874\.8tokens required by Graph\-RAG\. This reduction is a direct consequence of retrieval granularity: subgraph retrieval returns a compact set of triples rather than full passages, discarding surrounding prose that contributes tokens without contributing evidence\. These figures depend on retrieval hyperparameters and are not intrinsic to either paradigm\. Conversely, Graph\-RAG front\-loads its cost into an offline construction phase that RAG does not incur\. Extracting the knowledge graphs with KGGen[Mo et al\. \(2025\)](https://arxiv.org/html/2609.18317#bib.bib8)on a 51M\-character dataset required 72\.26 EUR in API calls, and training the Graph encoder took 9,942 s on average per fold\. Table 7:Full cost comparison between RAG and Graph\-RAG\. Inference context is mean±\\pmstd over the evaluation set; GraphEnc training is averaged overkk\-fold runs\.
Similar Articles
Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering
This paper extends PubHealthBench into a retrieval-augmented setting and evaluates retrieval and generation choices for public health QA, showing hybrid retrieval improves accuracy and that smaller models with retrieval can match larger ones.
SelfGraphRAG: Bridging the Supervision Gap in Graph-Based RAG with Synthetic QA Generation
SelfGraphRAG introduces a framework that generates synthetic question-answer pairs from knowledge graphs to address the supervision gap in graph-based retrieval-augmented generation, enhancing retrieval precision and reasoning performance.
VisKG-LM: Compiling Knowledge Graphs into Visual Memory for Multiple-Choice Question Answering
This paper proposes VisKG-LM, a method that compiles knowledge graphs into visual memory for efficient multiple-choice question answering, achieving performance gains over baselines by decoupling graph encoding from language reasoning.
Repair Before Reinforce: Context-Augmented Knowledge Graph Reasoning for Multi-Hop Question Answering
This paper proposes a context-augmented training framework for multi-hop question-answering, showing that combining context graphs with knowledge graphs and using reinforcement learning improves performance in biomedical domains.
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
Introduces MKG-RAG-Bench, a cross-domain benchmark for evaluating retrieval in multimodal knowledge graph-augmented generation, demonstrating that effective multimodal retrieval remains challenging and critical for downstream generation quality.