LLMs是否理解上下文?基于知识图谱的评估框架

arXiv cs.AI 论文

摘要

本文提出一种基于知识图谱的评估框架,用于评估大型语言模型在问答任务中的上下文理解能力,并引入一种新的相似度度量S3KG,该度量在性能上优于现有的基线方法。

arXiv:2609.30484v1 Announce Type: new Abstract: While large language models (LLMs) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses. However, traditional methods such as BiLingual Evaluation Understudy (BLEU) and perplexity simply measure surface-level performance. This reveals a critical gap in question answering (QA), where responses must be contextually grounded rather than simply being memorized associations. To fill this void, we propose a novel knowledge graph (KG) based evaluation framework for LLM contextual understanding in QA. Central to this is Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure combining structural and semantic signals into a single score. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine-grained analysis of model failures. Together, across nine benchmarks, S3KG achieves F1 gains of up to $+7.6$ points over the strongest baseline and AUROC up to $0.973$.
查看原文
查看缓存全文

缓存时间: 2026/09/28 09:37

# Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework
Source: [https://arxiv.org/html/2609.30484](https://arxiv.org/html/2609.30484)
## Do LLMs Understand Context? A Knowledge Graph\-Based Evaluation FrameworkThanks:\*Equal contribution\.

Mamta NallaretnamKithuni WickramasingheChamath GunapalaPragatheeswaran VipulanandanAffiliation:Department of Electrical and Computer Engineering, University of Miami, USAKamal PremaratneAffiliation:Department of Electrical and Computer Engineering, University of Miami, USAUthayasanker ThayasivamEmail:[\{ subavarshanaa\.21, nallaretnam\.21, kithuni\.21, chamathg\.21\}@cse\.mrt\.ac\.lk](mailto:)Affiliation:Department of Electrical and Computer Engineering, University of Miami, USAAffiliation:Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka

###### Abstract

While large language models \(LLMs\) have achieved remarkable linguistic capabilities, a profound question lingers at their core: do these models truly comprehend context or simply excel at pattern matching on an unprecedented scale? Contextual understanding in LLMs refers to the ability to correctly extract relevant information from a given context, integrate it into a coherent internal representation, and reason over it to produce factually consistent and contextually grounded responses\. However, traditional methods such as BiLingual Evaluation Understudy \(BLEU\) and perplexity simply measure surface\-level performance\. This reveals a critical gap in question answering \(QA\), where responses must be contextually grounded rather than simply being memorized associations\. To fill this void, we propose a novel knowledge graph \(KG\) based evaluation framework for LLM contextual understanding in QA\. Central to this is Semantic Structural Similarity for KGs \(S3KG\), a hybrid similarity measure combining structural and semantic signals into a single score\. In addition, a diagnostic analysis framework is developed to identify and categorize reasoning errors at the triplet level, enabling fine\-grained analysis of model failures\. Together, across nine benchmarks, S3KG achieves F1 gains of up to\+7\.6\+7\.6points over the strongest baseline and AUROC up to0\.9730\.973\.

## 1Introduction

LLMs, such as Bidirectional Encoder Representations from Transformers \(BERT\)\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.30484#bib.bib23)\), the Generative Pre\-trained Transformer \(GPT\) models\([Radford et al\., 2018](https://arxiv.org/html/2609.30484#bib.bib3)\), and their advanced variants, have fundamentally transformed natural language processing \(NLP\)\. These models exhibit extraordinary proficiency in generating coherent, human\-like text, answering complex questions, and executing a broad spectrum of tasks\([Brown et al\., 2020](https://arxiv.org/html/2609.30484#bib.bib1)\)\. However, a profound question lingers at the core of these models: do they genuinely understand the content they process or do they merely produce plausible outputs without true contextual comprehension\([Zhu et al\., 2024](https://arxiv.org/html/2609.30484#bib.bib6)\)\. Hallucination detection literature has largely bypassed this dimension, focusing instead on output\-level signals such as semantic entropy and token sequence probabilities\([Vipulanandan et al\., 2026](https://arxiv.org/html/2609.30484#bib.bib7)\)\.

Traditional evaluation metrics—such as perplexity, BLEU\([Papineni et al\., 2002](https://arxiv.org/html/2609.30484#bib.bib9)\), and other token\-level matching methods on standardized benchmarks—primarily measure syntactic\- or surface\-level performance and fail to capture the depth of semantic comprehension or contextual understanding\([Reiter, 2018](https://arxiv.org/html/2609.30484#bib.bib29)\)\. This limitation raises concerns about the reliability of LLMs in high\-stakes scenarios such as QA systems in medicine and healthcare, defense, and legal analysis\.

## 2Related Work

Evaluating how LLMs understand and utilize contextual information remains a key challenge in QA systems, as fluent and plausible responses do not necessarily reflect faithful use of the provided context\([Wallat et al\., 2024](https://arxiv.org/html/2609.30484#bib.bib2)\)\.[Zhu et al\. \(2024\)](https://arxiv.org/html/2609.30484#bib.bib6)evaluate LLMs across four core tasks—coreference resolution, discourse relation classification, dialogue state tracking, and query rewriting—showing that while they capture general contextual patterns, they fail to reveal which parts of the context are misunderstood, and their analysis does not extend to the QA domain\. Complementing this,[Yan et al\. \(2024\)](https://arxiv.org/html/2609.30484#bib.bib8)probe LLM reasoning by manipulating in\-context examples, including logical modifications such as swapping “AND”/“OR”, finding that models do not consistently obey formal reasoning rules nor exhibit clearly identifiable error patterns\. Together, these works establish that current LLMs demonstrate strong surface\-level comprehension yet still struggle with fine\-grained contextual interpretation and logical reasoning\. Evaluating these limitations in long\-form LLM answers is particularly challenging, making the use of KGs which involve accurate extraction of relational triplets from text a promising evaluation approach\.

KG construction has progressed from fixed\-schema supervised pipelines to joint extraction architectures enabled by pretrained language models\([Shang et al\., 2022](https://arxiv.org/html/2609.30484#bib.bib24)\)\. Embedding models such as TransE\([Bordes et al\., 2013](https://arxiv.org/html/2609.30484#bib.bib19)\), which represents relations as vector translations, and RotatE\([Sun et al\., 2019](https://arxiv.org/html/2609.30484#bib.bib20)\), which models relations as complex\-space rotations to capture symmetry and composition, makes them well\-suited for link prediction on static KGs but ill\-suited for cross\-graph similarity\. The Weisfeiler\-Lehman \(WL\) kernel\([Shervashidze et al\., 2011](https://arxiv.org/html/2609.30484#bib.bib21)\)compares graphs by iteratively aggregating neighbourhood labels into histograms, while the Wasserstein WL \(WWL\) kernel\([Togninalli et al\., 2019](https://arxiv.org/html/2609.30484#bib.bib22)\)replaces histogram comparison with Wasserstein distance for better handling of continuous attributes\. Both treat node labels as opaque symbols, so semantically equivalent but syntactically distinct labels receive zero credit—a key limitation\.[Haskins and Adams \(2025\)](https://arxiv.org/html/2609.30484#bib.bib4)partially address this through SBERT\-based\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.30484#bib.bib23)\)semantic clustering with few\-shot instruction tuning, yet the alignment remains lossy and falls short of a principled similarity measure for heterogeneous, independently constructed KGs\.

## 3Our Contributions

Our work makes three main contributions\.

- •Semantic Structural Similarity for KGs \(S3KG\)is a hybrid semantic\-structural similarity metric that converts LLM responses and reference answers into KG triplets and produces a single interpretable evaluation score\.
- •Contextual Understanding Score \(CUS\)is a model\-level aggregate of two complementary dimensions, factual accuracy \(GoldSim\) and contextual faithfulness \(CtxSim\), enabling cross\-model comparison across benchmarks\.
- •Triplet Analyzing Unit \(TAU\)is a diagnostic component that identifies and classifies reasoning failures at the triplet level for fine\-grained behavioral analysis of model outputs\.

## 4Methodology

Our framework evaluates LLM contextual understanding using a KG\-based pipeline \(see Figure[1](https://arxiv.org/html/2609.30484#S4.F1)\)\. Given a question and its supporting context, the LLM generates a response, from which KGs are constructed alongside those derived from the gold \(or reference or ground truth\) answer and context\. These KGs are then compared to measure similarity\. Low scoring pairs are further analyzed using a triplet analyzing unit to identify reasoning errors\.

![Refer to caption](https://arxiv.org/html/2609.30484v1/figures/flowchart_new.png)Figure 1:Methodology pipeline for LLM comparison and evaluation\.### 4\.1LLM Answer Collection

Each model is prompted with a question paired with its associated supporting context, and the generated response is recorded alongside the gold answer to form the inputs for downstream evaluation\. To establish a reproducible baseline, responses are first generated deterministically at temperature zero, yielding outputs that closely adhere to the provided context; subsequent runs are at progressively higher temperature settings allowing for us to examine how increasing generation diversity affects the model’s ability to retain and utilize contextual information\. These collected responses—together with the gold answers and supporting contexts—feed directly into the KG construction stage\.

### 4\.2KG Construction

For each QA instance, knowledge graphs are constructed from three sources: the gold answer, the model\-generated response, and the supporting context\. To ensure that the resulting KGs are comparable, we adopt the single few\-shot prompting strategy with instruction tuning used in[Sansford et al\. \(2024\)](https://arxiv.org/html/2609.30484#bib.bib5)and[Haskins and Adams \(2025\)](https://arxiv.org/html/2609.30484#bib.bib4), applying a shared extraction prompt uniformly across all three sources\. This ensures that entities and relations are extracted under the same schema, encouraging a consistent entity and relation label space across all 3 KGs\. Full details of the extraction prompt appear in Appendix[B](https://arxiv.org/html/2609.30484#A2)\.

Following extraction, an additional NLP normalization step is applied uniformly to every entity and relation label across all 3 KGs\. This includes lowercasing, lemmatization, and whitespace normalization, ensuring that any residual syntactic variation introduced during extraction is not retained in the final KG representations\. Together, the consistency\-aware prompting and post\-extraction normalization ensure that the 3 KGs are structurally compatible and ready for meaningful comparison using S3KG\.

### 4\.3S3KG: Semantic Structural Similarity for KGs

We denote a KG as𝒢=𝒢⁡\(𝒯\)\\mathcal\{G\}=\\mathcal\{G\}\(\\mathcal\{T\}\), where𝒯=\{\(hi,ri,ti\),i∈ℐ\}\\mathcal\{T\}=\\\{\(h\_\{i\},r\_\{i\},t\_\{i\}\),\\;i\\in\\mathcal\{I\}\\\}is a collection of triplets enumerated via a finite index setℐ\\mathcal\{I\}\. In theii\-th triplet\(hi,ri,ti\)\(h\_\{i\},r\_\{i\},t\_\{i\}\),hih\_\{i\},rir\_\{i\}, andtit\_\{i\}are the head entity, relation, and tail entity, respectively\.

S3KG computes the similarity between two KGs𝒢1​\(𝒯1\)\\mathcal\{G\}\_\{1\}\(\\mathcal\{T\}\_\{1\}\)and𝒢2​\(𝒯2\)\\mathcal\{G\}\_\{2\}\(\\mathcal\{T\}\_\{2\}\)\. Here, fork=1,2k=1,2,𝒯k=\{\(hk,i,rk,i,tk,i\),i∈ℐk\}\\mathcal\{T\}\_\{k\}=\\\{\(h\_\{k,i\},r\_\{k,i\},t\_\{k,i\}\),\\;i\\in\\mathcal\{I\}\_\{k\}\\\}\. Rather than comparing KGs by a single criterion, S3KG operates at two levels: \(1\) astructural scoreSWL\\text\{S\}\_\{\\text\{WL\}\}computed node\- and edge\-wise using the WL kernel over soft\-label aligned KGs, and \(2\) asemantic scoreSSBERT\\text\{S\}\_\{\\text\{SBERT\}\}computed triplet\-wise from mean\-pooled SBERT embeddings\([Reimers and Gurevych, 2019](https://arxiv.org/html/2609.30484#bib.bib23)\)\. These are blended via a mixing coefficientα\\alphato get the final combined similarity score \(see \([2](https://arxiv.org/html/2609.30484#S4.E2)\)\)\.

Triplet\-Level Matching\.Triplets are serialised as natural language \(NL\) strings\. SBERT embeddings are computed for all triplets in both sets𝒯1\\mathcal\{T\}\_\{1\}and𝒯2\\mathcal\{T\}\_\{2\}usingparaphrase\-MPNet\-base\-v2\. For each triplet in𝒯1\\mathcal\{T\}\_\{1\}, the most semantically similar triplet in𝒯2\\mathcal\{T\}\_\{2\}is selected by cosine similarity, producing a filtered set𝒯^2⊆𝒯2\\widehat\{\\mathcal\{T\}\}\_\{2\}\\subseteq\\mathcal\{T\}\_\{2\}that anchors the comparison to semantically relevant content\. This unidirectional matching strategy is adopted deliberately\. By treating𝒯1\\mathcal\{T\}\_\{1\}as the reference set, the method primarily evaluates how well the content of𝒯1\\mathcal\{T\}\_\{1\}is covered by𝒯2\\mathcal\{T\}\_\{2\}, thereby emphasising recall while not penalising additional triplets present in𝒯2\\mathcal\{T\}\_\{2\}\. In contrast, a symmetric bidirectional formulation—computed by averaging matches from both𝒯^2\\widehat\{\\mathcal\{T\}\}\_\{2\}and𝒯^1\\widehat\{\\mathcal\{T\}\}\_\{1\}, similar to the F1 formulation of BERTScore\([Zhang et al\., 2020](https://arxiv.org/html/2609.30484#bib.bib16)\)—would account for both recall and precision by also penalising unmatched surplus triplets\. Investigating the empirical differences between the unidirectional and bidirectional variants is left for future work\.

Soft Label Alignment\.The Standard WL kernel compares graphs by matching node labels exactly: two labels contribute to the similarity score only if they are syntactically identical strings\. This means that semantically equivalent but syntactically different entity or relation labels \(e\.g\.,“found”and“discovered”\) are treated as entirely distinct, and no credit is awarded for any semantic equivalence\. S3KG resolves this through a soft label alignment step applied before kernel computation\. Each node in𝒢\\mathcal\{G\}carries an entity label, and each edge carries a relation label\. These are aligned independently to prevent cross\-type collisions: for each entity labelℓ\\ellin𝒢1\\mathcal\{G\}\_\{1\}, if its maximum cosine similarity, computed via SBERT embeddingsϕ⁡\(⋅\)\\phi\(\\cdot\)to any entity label in𝒢2\\mathcal\{G\}\_\{2\}exceeds a thresholdτ=0\.65\\tau=0\.65, it is mapped to the canonical identifier of the best\-matching label in𝒢2\\mathcal\{G\}\_\{2\}\(e\.g\.,node\_0,node\_1\); otherwise it is left unchanged\. The same procedure is applied independently to relation labels \(e\.g\.,rel\_0,rel\_1\)\. After alignment, syntactically different but semantically equivalent labels share the same canonical identifier, allowing the WL kernel to recognise them as matching\.

WL Kernel Structural Similarity\.Following soft label alignment, the two KGs are compared using the WL graph kernel\([Shervashidze et al\., 2011](https://arxiv.org/html/2609.30484#bib.bib21)\)\. The WL kernel operates iteratively: at iteration00, each node is characterised by its initial \(soft aligned\) entity label\. At each subsequent iterationkk, every node aggregates its current label with the multiset of its neighbours’ labels and the connecting relation labels, producing a new refined label that encodes the node’skk\-hop neighbourhood structure\. We useK=5K=5iterations, so each node’s final label summarises structural patterns up to 5 hops away\. The kernel score is the normalised inner product between the resulting label\-count histograms of the two graphs, yielding a structural similarity scoreSWL∈\[0,1\]\\text\{S\}\_\{\\text\{WL\}\}\\in\[0,1\]regardless of KG size\.

SBERT Mean\-Pool Semantic Similarity\.Each triplet\(h,r,t\)∈𝒯\(h,r,t\)\\in\\mathcal\{T\}is encoded by SBERT into an embeddingϕ⁡\(h,r,t\)\\phi\(h,r,t\); the graph\-level representation is the mean𝐞¯𝒯=1\|𝒯\|​∑𝒯ϕ⁡\(h,r,t\)\\bar\{\\mathbf\{e\}\}\_\{\\mathcal\{T\}\}=\\dfrac\{1\}\{\|\\mathcal\{T\}\|\}\\sum\_\{\\mathcal\{T\}\}\\phi\(h,r,t\)\. Semantic similarity is the cosine between the mean\-pooled representations of𝒯1\\mathcal\{T\}\_\{1\}and𝒯^2\\widehat\{\\mathcal\{T\}\}\_\{2\}\(the semantically filtered reference triplets from Step 1\), clipped to\[0,1\]\[0,1\]:

SSBERT​\(𝒯1,𝒯^2\)=max⁡\(0,𝐞¯𝒯1⋅𝐞¯𝒯^2‖𝐞¯𝒯1‖⋅‖𝐞¯𝒯^2‖\)\.\\text\{S\}\_\{\\text\{SBERT\}\}\(\\mathcal\{T\}\_\{1\},\\,\\widehat\{\\mathcal\{T\}\}\_\{2\}\)=\\max\\left\(0,\\frac\{\\bar\{\\mathbf\{e\}\}\_\{\\mathcal\{T\}\_\{1\}\}\\cdot\\bar\{\\mathbf\{e\}\}\_\{\\widehat\{\\mathcal\{T\}\}\_\{2\}\}\}\{\\\|\\bar\{\\mathbf\{e\}\}\_\{\\mathcal\{T\}\_\{1\}\}\\\|\\cdot\\\|\\bar\{\\mathbf\{e\}\}\_\{\\widehat\{\\mathcal\{T\}\}\_\{2\}\}\\\|\}\\right\)\.\(1\)This captures sentence\-level meaning that discrete WL label refinement cannot\.

Combined Score\.The structural and semantic scores are combined via mixing coefficientα\\alphaas

SS3KG=\(1−α\)​SWL\+α​SSBERT\.α∈\[0,1\]\.\\text\{S\}\_\{\\text\{S3KG\}\}=\(1\-\\alpha\)\\,\\text\{S\}\_\{\\text\{WL\}\}\+\\alpha\\,\\text\{S\}\_\{\\text\{SBERT\}\}\.\\;\\alpha\\in\[0,1\]\.\(2\)We useα=0\.5\\alpha=0\.5so that the score equally weights structural fidelity from WL neighbourhood aggregation over aligned labels and semantic coherence from SBERT mean\-pooling, thus encoding both local relational patterns and global meaning within a single score\.

### 4\.4Contextual Understanding Score \(CUS\)

With the 3 KGs—𝐾𝐺LLM\\mathit\{KG\}\_\{\\text\{LLM\}\}associated with the model\-generated response,𝐾𝐺gold\\mathit\{KG\}\_\{\\text\{gold\}\}associated with the gold answer, and𝐾𝐺ctx\\mathit\{KG\}\_\{\\text\{ctx\}\}associated with the supporting context—in hand, for each QA instanceqq, we apply S3KG to get \(1\)GoldSim⁡\(q\)=SS3KG​\(𝐾𝐺LLM\(q\),𝐾𝐺gold\(q\)\)\\mathrm\{GoldSim\}\(q\)=\\text\{S\}\_\{\\text\{S3KG\}\}\(\\mathit\{KG\}\_\{\\text\{LLM\}\}^\{\(q\)\},\\mathit\{KG\}\_\{\\text\{gold\}\}^\{\(q\)\}\)which measures factual accuracy by comparing the LLM response against the gold answer; and \(2\)CtxSim⁡\(q\)=SS3KG​\(𝐾𝐺LLM\(q\),𝐾𝐺ctx\(q\)\)\\mathrm\{CtxSim\}\(q\)=\\text\{S\}\_\{\\text\{S3KG\}\}\(\\mathit\{KG\}\_\{\\text\{LLM\}\}^\{\(q\)\},\\mathit\{KG\}\_\{\\text\{ctx\}\}^\{\(q\)\}\)which measures contextual faithfulness by comparing the LLM response against the supporting context\. Since neither dimension alone reflects true understanding, we employ the harmonic mean to generate a*Contextual Understanding Score \(CUS\)*as

CUS⁡\(q\)=2⋅GoldSim⁡\(q\)⋅CtxSim⁡\(q\)GoldSim⁡\(q\)\+CtxSim⁡\(q\)\.\\mathrm\{CUS\}\(q\)=\\frac\{2\\cdot\\mathrm\{GoldSim\}\(q\)\\cdot\\mathrm\{CtxSim\}\(q\)\}\{\\mathrm\{GoldSim\}\(q\)\+\\mathrm\{CtxSim\}\(q\)\}\.\(3\)This harmonic mean penalises imbalanced profiles, ranking a model having one strong and one weak score below one having a pair of moderate scores\. The dataset\-level CUS is the mean ofCUS⁡\(q\)\\mathrm\{CUS\}\(q\)over allNNsamples\.

### 4\.5Triplet Analysis Unit \(TAU\)

For the 5% of lowest\-scoring QA pairs, we apply a triplet analysis unit \(TAU\) to identify where and how the LLM generated KG diverges from the gold KG\. Each triplet is converted into an NL sentence and encoded using a sentence transformer model\. Cosine similarity is computed between gold and LLM triplet embeddings; aligned triplets are identified by thresholding the cosine similarity between triplet sentence embeddings\.

After removing aligned triplets, residual pairs are categorized into interpretable error classes based on component\-wise cosine similarities for head, relation, and tail: \(1\)Relation mismatch: entities match, but the relation differs\. \(2\)Entity mismatch: the relation aligns, but the entity pair is inconsistent\. \(3\)Extra triplets: hallucinated or additional triplets generated by the LLM\. \(4\)Missing triplets: relevant triplets that were not extracted by the LLM\. Full TAU evaluation details appear in Appendix[C](https://arxiv.org/html/2609.30484#A3)\.

## 5Experiments

### 5\.1Datasets

Two context\-rich QA datasets containing long\-form answers are used for evaluation purposes\. \(1\)PubMedQA\([Jin et al\., 2019](https://arxiv.org/html/2609.30484#bib.bib26)\)contains 273,518 biomedical QA pairs drawn from research articles, with answers typically exceeding 100 words, providing a testbed for domain\-specific detailed response evaluation\. \(2\)MesaQA\([Wang et al\., 2025](https://arxiv.org/html/2609.30484#bib.bib25)\)comprises approximately 6,100 QA pairs from consumer healthcare documents, featuring abstractive answers averaging 70 words that require multi\-span evidence integration\. Together, these two datasets provide evaluation coverage across academic biomedical reasoning and practical healthcare knowledge synthesis\.

### 5\.2LLM Answer Collection

We evaluate 4 instruction\-tuned open\-source language models with 7\-billion parameters:Llama\-2\-7b\-chat\-hf,Gemma\-7b\-it,Mistral\-7B\-Instruct\-v0\.2, andFalcon\-7B\-Instruct\. These specific models are selected because they operate at a similar scale, thus allowing for a fair comparison without the influence of model size\. Each model is prompted with a question and its associated supporting context, and the model\-generated response is recorded alongside the gold answer for evaluation\. To establish a baseline, responses are first generated deterministically at00temperature, yielding outputs that closely adhere to the provided context\. Additional responses are then generated at higher temperature settings \(we use0\.30\.3,0\.70\.7, and1\.01\.0\) to examine how increasing generation diversity influences the model’s ability to retain and utilize contextual information\. Full temperature analysis appears in Appendix[A](https://arxiv.org/html/2609.30484#A1)\.

### 5\.3Benchmarking Dataset Collection

To assess generalisation across diverse text types, we use 10 datasets spanning three structural categories\. Each dataset is cast as a binary classification task: given a text pair\(s1,s2\)\(s\_\{1\},s\_\{2\}\), predict whether the pair is semantically equivalent \(label=1=1\) or not \(label=0=0\)\. All datasets are balanced atNNpositive andNNnegative pairs \(we useN=400N=400\)\. Performance is measured by maximum F1 score \(obtained via threshold sweep\) and AUROC\.

Short\-Text with Human\-Annotated\.We use 3 datasets containing sentence pairs averaging 10–22 words with crowd\-sourced or expert equivalence labels:MRPC\([Dolan and Brockett, 2005](https://arxiv.org/html/2609.30484#bib.bib10)\), consisting of news sentence pairs with paraphrase labels;PAWS\-Wiki\([Zhang et al\., 2019](https://arxiv.org/html/2609.30484#bib.bib11)\), adversarially constructed paraphrase pairs from Wikipedia where lexical overlap is deliberately an unreliable signal; andSTS12\([Agirre et al\., 2012](https://arxiv.org/html/2609.30484#bib.bib12)\), sentence similarity pairs drawn from multiple NLP tasks\.

KG\-Perturbed Paragraphs\.We use 6 datasets containing paragraphs averaging 69–126 words, constructed by perturbing entity relationships in KG\-derived paragraph representations\. The 5 evaluated datasets—SK\-Codex 400,SK\-Combined,SK\-FindKG,SK\-GloBI, andSK\-Oregano—differ in their underlying KG source, covering general encyclopaedic \(Codex\([Safavi and Koutra, 2022](https://arxiv.org/html/2609.30484#bib.bib13)\)\), financial/economic \(FindKG\([Li and Sanna Passino, 2024](https://arxiv.org/html/2609.30484#bib.bib28)\)\), biological interaction \(GloBI\([Poelen et al\., 2014](https://arxiv.org/html/2609.30484#bib.bib14)\)\), and food ontology \(Oregano[Boudin et al\. \(2023\)](https://arxiv.org/html/2609.30484#bib.bib27)\) domains\.

NLP\-Perturbed Paragraphs\.NLP\-Perturbed Paragraphs\.TheWikipedia Entity\-Swapdataset \(399 pairs\) replaces named entities in Wikipedia passages using four NLP\-based perturbations—node replacement, node deletion, edge deletion, and edge replacement via WordNet antonyms[Miller \(1992\)](https://arxiv.org/html/2609.30484#bib.bib32)—with no KG involvement at any stage[Rico et al\. \(2016\)](https://arxiv.org/html/2609.30484#bib.bib30);[Wei and Zou \(2019\)](https://arxiv.org/html/2609.30484#bib.bib31)\. It serves as an anti\-circularity probe: were our gains an artifact of circular evaluation, performance here should collapse, but it does not\.

### 5\.4Methods Evaluated

Table 1:Similarity scores on two PAWS\-Wiki pairs \(threshold=0\.5=0\.5;✓ = correct,×\\times= incorrect\)\.The positive pair differs only in word order; the negative pair swaps the subject and object of the winning relation\. All seven baselines assign near\-identical high scores to both pairs, failing on the negative example\. For thepositive pair, S3KG achieves the optimal score of1\.001\.00, whereas surface\-form methods such as BLEU \(0\.580\.58\) and ROUGE\-L \(0\.850\.85\) underperform by penalising inconsequential word\-order variation\.For thenegative pair, S3KG scores0\.480\.48—the only sub\-threshold result—correctly predictingNot Similar\. The score is not zero because the sentences share substantial content; only the relational direction differs\. The KG component isolates this reversal via the directed triplet, reducing the score from the near\-1\.01\.0surface baseline to just below the decision threshold, whileα=0\.5\\alpha=0\.5balances surface and structural similarity\.S3KGis evaluated with a mixing coefficientα\\alpha\(see \([2](https://arxiv.org/html/2609.30484#S4.E2)\)\) which controls the blend between KG structural signal and sentence\-transformer signal:α=0\.0\\alpha=0\.0recovers a pure KG structural embedding ;α=1\.0\\alpha=1\.0recovers a pure dense sentence\-transformer representation\. Pure KG structural embeddings capture relational and ontological structure between concepts but lack linguistic flexibility and contextual expressiveness, whereas sentence transformers excel at contextual and semantic similarity yet remain blind to the underlying KG topology\. S3KG bridges this gap by interpolating between both signals, enabling richer matching that is sensitive to both conceptual structure and NL meaning\. For each dataset, we report the best\-performing variant selected by maximum F1 across the sweepα∈\{0\.0,0\.1,…,1\.0\}\\alpha\\in\\\{0\.0,0\.1,\\ldots,1\.0\\\}\. Results for a more complete per\-datasetα\\alphasweep appear in Appendix[E](https://arxiv.org/html/2609.30484#A5)\. We compare against 7 standard baselines: ROUGE\-1, ROUGE\-2, ROUGE\-L\([Lin, 2004](https://arxiv.org/html/2609.30484#bib.bib15)\), BLEU\([Papineni et al\., 2002](https://arxiv.org/html/2609.30484#bib.bib9)\), BERTScore\([Zhang et al\., 2020](https://arxiv.org/html/2609.30484#bib.bib16)\), MiniLM\([Wang et al\., 2020](https://arxiv.org/html/2609.30484#bib.bib17)\), and sentence\-T5\-base\([Ni et al\., 2022](https://arxiv.org/html/2609.30484#bib.bib18)\)\. Table[1](https://arxiv.org/html/2609.30484#S5.T1)provides a concrete worked example demonstrating how S3KG detects relational reversals that all 7 baselines fail to distinguish\.

## 6Results

We evaluate S3KG against seven baselines on 9 benchmarks spanning short\-text paraphrase detection, KG\-perturbed paragraphs, and an anti\-circularity entity\-swap control\. Table[2](https://arxiv.org/html/2609.30484#S6.T2)summarises F1 and AUROC across every dataset×\\timesmethod cell; the per\-dataset bestα\\alphatogether with the headline scores appear in Table[3](https://arxiv.org/html/2609.30484#S6.T3)\. Full per\-dataset performance tables \(short\-text, KG\-perturbed paragraph, Wikipedia entity\-swap\) appear in Appendix[D](https://arxiv.org/html/2609.30484#A4); the completeα\\alphasweep appears in Appendix[E](https://arxiv.org/html/2609.30484#A5)\.

Table 2:F1 Score and ROC\-AUC of S3KG and seven baselines across all nine benchmark datasets\. The best\-performing method per dataset isbolded\. S3KG is shown using the bestα\\alphavariant per dataset, selected by maximum F1 overα\\alphaswept overα∈\{0\.0,0\.1,…,1\.0\}\\alpha\\in\\\{0\.0,0\.1,\\ldots,1\.0\\\}\(full sweep in Appendix[E](https://arxiv.org/html/2609.30484#A5)\)\. S3KG attains the highest score on 6 of 9 datasets and the highest*meaningful*score on the Wikipedia Entity\-Swap anti\-circularity control, with gains of up to\+7\.6\+7\.6F1 over the strongest baseline on KG\-rich paragraph datasets\.F1 Score

AUROC

Table 3:Best S3KG variant per dataset, selected by maximum F1 via grid search overα∈\{0\.0,0\.1,…,1\.0\}\\alpha\\in\\\{0\.0,0\.1,\\ldots,1\.0\\\}\. KG\-rich paragraph datasets gain up to\+7\.6\+7\.6F1 over the strongest baseline; sparse or noisy KGs yield reduced margins\.Text Richness Drives KG Performance\.S3KG performance scales with text length and relational density\. On short texts \(MRPC, STS12; 10–22 words average\), the structural KG signal is sparse because only a few well\-formed triplets can be extracted, and S3KG is competitive but trails sentence\-T5\-base by 6–7 F1 points\. On paragraph\-level datasets \(69–126 words\), S3KG reaches top\-1 performance on 4 of 5 KG\-perturbed benchmarks—SK\-Codex 400 \(F1=0\.932=0\.932,\+5\.7\+5\.7over MiniLM\), SK\-Combined \(0\.8340\.834,\+6\.4\+6\.4over sentence\-T5\-base\), SK\-GloBI \(0\.8920\.892,\+7\.6\+7\.6over BERTScore\), and SK\-Oregano \(0\.8120\.812,\+1\.2\+1\.2over sentence\-T5\-base\)—confirming that sufficient relational content is required for the structural channel to pay off\. On the adversarial PAWS\-Wiki paraphrase benchmark, S3KG still leads \(F1=0\.766=0\.766\) while ROUGE\-1 collapses to near\-random \(AUC=0\.490=0\.490\), reflecting the well\-known failure ofnn\-gram overlap on surface\-form\-preserving rephrasings\.

KG Extraction Quality is a Bottleneck\.On SK\-FindKG, where financial and economic vocabulary degrades triplet extraction quality, sentence\-T5\-base leads \(F1=0\.848=0\.848\) and S3KG drops to0\.7670\.767\. The bestα\\alphafor this dataset is0\.00\.0\(pure KG\), but the absolute score remains capped by noisy triplets, evidence that S3KG’s gains depend on the underlying graph being faithfully recoverable from text\.

Anti\-Circularity Validation\.On the Wikipedia Entity\-Swap control, which is constructed independently of any KG used during S3KG development, S3KG attains the best meaningful score \(F1=0\.872=0\.872, AUC=0\.890=0\.890\), exceeding MiniLM \(F1=0\.821=0\.821\) and sentence\-T5\-base \(F1=0\.762=0\.762\)\. ROUGE\-1 is excluded from the meaningful comparison because entity\-swapped pairs share nearly all surrounding tokens, making unigram overlap trivially near\-perfect, a dataset artifact rather than a real signal\. The result rules out circularity as an explanation for S3KG’s KG\-perturbed gains\.

Optimalα\\alphais Dataset\-Dependent\.Lowerα\\alphavalues favour datasets where perturbations are primarily structural \(SK\-FindKG:α=0\.0\\alpha=0\.0; STS12, Wiki Swap:α=0\.1\\alpha=0\.1\)\. Higher values are preferred when KG and dense signals are complementary \(SK\-Codex 400 and SK\-Combined:α=0\.5\\alpha=0\.5; SK\-GloBI:α=0\.6\\alpha=0\.6\)\. The full sweep appears in Appendix[E](https://arxiv.org/html/2609.30484#A5)\.

Baseline Behaviour\.Token\-overlap baselines \(ROUGE, BLEU\) are strong on KG\-perturbed datasets where perturbations alter surface form, but unreliable on adversarial datasets \(PAWS\-Wiki: ROUGE\-1 AUC=0\.490=0\.490; Wiki Swap: ROUGE\-L AUC=0\.311=0\.311\)\. BERTScore is more stable but consistently underperforms S3KG on KG\-perturbed data\. MiniLM and sentence\-T5\-base are the strongest baselines overall but require full fine\-tuned transformer inference, whereas S3KG’s KG component is comparatively lightweight at inference time\.

### 6\.1LLM Comparison on QA Datasets

Having validated S3KG as a reliable KG similarity measure, we apply it as an evaluation instrument to address the following question:given a question and its supporting context, to what extent does an LLM capture the relational knowledge of the reference answer, and how faithfully does its response reflect the provided context?Traditional metrics such as BLEU or exact match are insufficient for this purpose, as they assess syntactic\-level token overlap rather than the relational knowledge structure of a response\.

Using the CUS evaluation pipeline in Section[4\.4](https://arxiv.org/html/2609.30484#S4.SS4), Table[4](https://arxiv.org/html/2609.30484#S6.T4)reports mean GoldSim, CtxSim, and CUS for four 7B\-parameter models on PubMedQA and MesaQA \(N=400N=400,α=0\.5\\alpha=0\.5\)\.

Table 4:LLM Evaluation Results \(mean overN=400N\{=\}400samples,,α=0\.5\\alpha\{=\}0\.5\)\. GoldSim measures factual alignment with the reference answer; CtxSim measures faithfulness to the supporting context; CUS \(Equation[3](https://arxiv.org/html/2609.30484#S4.E3)\) is their harmonic mean, penalising imbalanced profiles\. Mistral\-7B achieves the best CUS on both datasets, reflecting consistently balanced factual and contextual understanding, while Falcon\-7B underperforms across all metrics and both domains\.MesaQA Results\.Gemma\-7B, Llama\-2\-7B, and Mistral\-7B achieve similar CUS scores \(0\.6750\.675–0\.6780\.678\), while Falcon\-7B scores notably lower \(0\.6060\.606\)\. Gemma\-7B leads on GoldSim \(0\.6890\.689\), while Mistral\-7B leads on CtxSim \(0\.7360\.736\) and achieves the best overall CUS \(0\.6780\.678\) by balancing both dimensions\. CtxSim consistently exceeds GoldSim across all models, indicating that models draw effectively from context but add content not present in the reference answer\.

PubMedQA Results\.Performance drops substantially, with CUS ranging from0\.4800\.480\(Falcon\-7B\) to0\.5920\.592\(Mistral\-7B\), reflecting the difficulty of matching precise biomedical reference answers\. Mistral\-7B again leads in CUS, supported by the highest CtxSim \(0\.7330\.733\)\. Gemma\-7B scores highest on GoldSim \(0\.5240\.524\) but lowest on CtxSim \(0\.6400\.640\), indicating closer alignment with reference content but weaker use of biomedical context\. Falcon\-7B is weakest across all metrics on both datasets\.

Cross\-Dataset Observations\.CtxSim exceeds GoldSim in all 8 model–dataset combinations, consistent with instruction\-tuned models elaborating on context rather than producing concise reference\-style responses\. The∼\\sim10\-point CUS gap between MesaQA and PubMedQA across all models points to domain complexity as the primary factor, with biomedical vocabulary and reasoning posing challenges irrespective of model architecture\.

Triplet Analysis of the Least Similar KGs\.Across the 5% lowest\-similarity cases in both datasets, the four models show clear differences, and we focus on this tail subset because these challenging instances make model failures more diagnostic and reveal systematic weaknesses that can be masked by strong average scores\. Mistral is consistently the most reliable on the hard examples, preserving more reference facts and producing more aligned triplets than the other models\. Llama generally falls in the middle, while Gemma shows the weakest performance, especially on the general health QA set where it often fails to recover many gold\-aligned facts\. A second consistent pattern is that, in these difficult cases, the generated KGs tend to align more closely with the contextual KG than with the reference KG\. This suggests that many failures are not simply random errors, but cases where the model either drifts toward a different interpretation, omits key reference facts, or introduces unsupported additions\. Finally, the PubMed setting is noticeably harder for all models: aligned\-triplet recovery drops sharply, indicating that technical biomedical terminology, abbreviations, and entity variability dominate the failure modes in the hardest cases\.

## 7Discussion

The proposed KG\-based framework provides a structured approach to evaluate LLM understanding, moving beyond syntactic\-level metrics to assess relational and structural fidelity\. S3KG’s type\-separated, one\-to\-one label alignment addresses fundamental limitations of clustering\-based approaches while maintaining the WL kernel’s ability to capture multi\-hop neighborhood similarity\.

The consistent performance gap between PubMedQA and MesaQA scores across all 3 models reveals a measurable difference in LLM capability for domain\-specific versus general health knowledge\. The two\-dimensional S3KG evaluation further distinguishes factual accuracy \(gold similarity\) from contextual faithfulness \(context similarity\), exposing model\-specific trade\-offs that aggregate metrics cannot capture\. The TAU additionally enables pinpointing specific reasoning failures, whether due to incorrect entity substitution, relation errors, or broader inconsistencies\.

## 8Conclusion

We presented a novel KG\-based evaluation framework for assessing LLM contextual understanding in QA\. By constructing canonicalized KGs from LLM outputs, gold answers, and context, and comparing them using S3KG, we move beyond syntactic\-level accuracy toward verifiable, graph\-theoretic comprehension measurement\. S3KG achieves best\-per\-dataset F1 of0\.7660\.766–0\.9320\.932and AUROC up to0\.9730\.973, consistently outperforming lexical and neural baselines on KG\-rich datasets while remaining competitive on short\-text settings\. The TAU provides interpretable, fine\-grained diagnostics of reasoning failures\. Evaluations on PubMedQA and MesaQA demonstrate consistent model\-specific strengths and weaknesses, establishing a reproducible pipeline that can be extended to other datasets and tasks to support trustworthy AI development\. In this sense, whether LLMs understand context becomes empirically testable by measuring how well their generated responses preserve the relational knowledge expressed in the reference answer and supporting context\.

## Limitations

Key limitations of this work include: \(1\) sensitivity of similarity scores to KG extraction quality, as noisy or incomplete triplet extraction directly degrades S3KG performance; \(2\) computational cost of SBERT inference at scale, which may be prohibitive for very large evaluation sets without GPU acceleration; and \(3\) evaluation is currently restricted to open\-source 7B\-parameter models; extending to larger proprietary models such as GPT\-4 remains future work\. Additionally, the current alignment scheme does not handle directional semantic equivalence \(e\.g\.,daughter\_ofvs\.mother\_of\), which would require attention\-based mechanisms\.

## Acknowledgments

The work of Kamal Premaratne \(KP\) was supported by the French/US joint project LUCAS between the Agence Nationale de la Recherche \(ANR\) \(grant ANR\-25\-CE23\-2189\) and the U\.S\. National Science Foundation \(NSF\) \(grant numbers 2530255 and 2530256\)\.

## References

- Agirreet al\.\(2012\)E\. Agirre, D\. Cer, M\. Diab, and A\. Gonzalez\-AgirreSemEval\-2012 task 6: a pilot on semantic textual similarity\.InProceedings of the First Joint Conference on Lexical and Computational Semantics \(\*SEM\),Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p2.1)\.
- Bordeset al\.\(2013\)A\. Bordes, N\. Usunier, A\. García\-Durán, J\. Weston, and O\. YakhnenkoTranslating embeddings for modeling multi\-relational data\.InAdvances in Neural Information Processing Systems,Vol\.26,pp\. 2787–2795\.Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p2.1)\.
- Boudinet al\.\(2023\)M\. Boudin, G\. Diallo, M\. Drancé, and F\. MouginThe OREGANO knowledge graph for computational drug repurposing\.Scientific Data10,pp\. 871\.Note:Food ontology and natural compound knowledge graph for drug repurposingExternal Links:[Document](https://dx.doi.org/10.1038/s41597-023-02757-0),[Link](https://www.nature.com/articles/s41597-023-02757-0)Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p3.1)\.
- Brownet al\.\(2020\)T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.Language models are few\-shot learners\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 1877–1901\.Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p1.1)\.
- Dolan and Brockett \(2005\)W\. B\. Dolan and C\. BrockettAutomatically constructing a corpus of sentential paraphrases\.InProceedings of the Third International Workshop on Paraphrasing \(IWP2005\),Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p2.1)\.
- Haskins and Adams \(2025\)R\. Haskins and B\. AdamsKea explain: explanations of hallucinations using graph kernel analysis\.Note:arXiv preprint arXiv:2507\.03847Cited by:[Appendix B](https://arxiv.org/html/2609.30484#A2.p1.1),[§2](https://arxiv.org/html/2609.30484#S2.p2.1),[§4\.2](https://arxiv.org/html/2609.30484#S4.SS2.p1.1)\.
- Jinet al\.\(2019\)Q\. Jin, B\. Dhingra, Z\. Liu, W\. Cohen, and X\. LuPubMedQA: a dataset for biomedical research question answering\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 2567–2577\.Cited by:[§5\.1](https://arxiv.org/html/2609.30484#S5.SS1.p1.1)\.
- Li and Sanna Passino \(2024\)X\. V\. Li and F\. Sanna PassinoFinDKG: dynamic knowledge graphs with large language models for detecting global trends in financial markets\.InProceedings of the 5th ACM International Conference on AI in Finance \(ICAIF 2024\),pp\. 573–581\.Note:Financial knowledge graph extracted from news articles using LLMsExternal Links:[Document](https://dx.doi.org/10.1145/3677052.3698603),[Link](https://arxiv.org/abs/2407.10909),2407\.10909Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p3.1)\.
- Lin \(2004\)C\.\-Y\. LinROUGE: a package for automatic evaluation of summaries\.InText Summarization Branches Out: Proceedings of the ACL\-04 Workshop,pp\. 74–81\.Cited by:[§5\.4](https://arxiv.org/html/2609.30484#S5.SS4.p1.1)\.
- Miller \(1992\)G\. A\. MillerWordNet: a lexical database for English\.InSpeech and Natural Language: Proceedings of a Workshop Held at Harriman, New York, February 23\-26, 1992,External Links:[Link](https://aclanthology.org/H92-1116/)Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p4.1)\.
- Niet al\.\(2022\)J\. Ni, N\. C\. Gustavo Hernandez Abrego, K\. H\. Ji Ma, D\. Cer, and Y\. YangSentence\-t5: scalable sentence encoders from pre\-trained text\-to\-text models\.InFindings of the Association for Computational Linguistics: ACL 2022,pp\. 1864–1874\.External Links:[Document](https://dx.doi.org/10.18653/v1/2022.findings-acl.146)Cited by:[§5\.4](https://arxiv.org/html/2609.30484#S5.SS4.p1.1)\.
- Papineniet al\.\(2002\)K\. Papineni, S\. Roukos, T\. Ward, and W\.\-J\. ZhuBLEU: a method for automatic evaluation of machine translation\.InProceedings of the 40th Annual Meeting of the Association for Computational Linguistics,Philadelphia, PA, USA,pp\. 311–318\.Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p2.1),[§5\.4](https://arxiv.org/html/2609.30484#S5.SS4.p1.1)\.
- Poelenet al\.\(2014\)J\. H\. Poelen, J\. D\. Simons, and C\. J\. MungallGloBI: global biotic interactions\.Note:\[Online\]\. Available:[https://www\.globalbioticinteractions\.org](https://www.globalbioticinteractions.org/)Accessed: Jan\. 15, 2025Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p3.1)\.
- Radfordet al\.\(2018\)A\. Radford, K\. Narasimhan, T\. Salimans, and I\. SutskeverImproving language understanding by generative pre\-training\.Technical reportOpenAI\.External Links:[Link](https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf)Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p1.1)\.
- Reimers and Gurevych \(2019\)N\. Reimers and I\. GurevychSentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),Hong Kong, China,pp\. 3982–3992\.External Links:[Document](https://dx.doi.org/10.18653/v1/D19-1410)Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p1.1),[§2](https://arxiv.org/html/2609.30484#S2.p2.1),[§4\.3](https://arxiv.org/html/2609.30484#S4.SS3.p2.1)\.
- Reiter \(2018\)E\. ReiterA structured review of the validity of BLEU\.Computational Linguistics44\(3\),pp\. 393–401\.External Links:[Document](https://dx.doi.org/10.1162/coli%5Fa%5F00322)Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p2.1)\.
- Ricoet al\.\(2016\)S\. Rico, H\. Barry, and B\. AlexandraNeural machine translation of rare words with subword units\.InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Berlin, Germany,pp\. 1715–1725\.External Links:[Link](https://aclanthology.org/P16-1162/),[Document](https://dx.doi.org/10.18653/v1/P16-1162)Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p4.1)\.
- Safavi and Koutra \(2022\)T\. Safavi and D\. KoutraCoDEx: a comprehensive knowledge graph completion benchmark\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing\(EMNLP\),pp\. 8328–8350\.Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p3.1)\.
- Sansfordet al\.\(2024\)H\. Sansford, N\. Richardson, H\. P\. Maretic, and J\. N\. SaadaGrapheval: a knowledge\-graph based llm hallucination evaluation framework\.Note:arXiv preprint arXiv:2407\.10793Cited by:[Appendix B](https://arxiv.org/html/2609.30484#A2.p1.1),[§4\.2](https://arxiv.org/html/2609.30484#S4.SS2.p1.1)\.
- Shanget al\.\(2022\)Y\. Shang, H\. Huang, and X\. MaoOneRel: joint entity and relation extraction with one module in one step\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.36,pp\. 11285–11293\.Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p2.1)\.
- Shervashidzeet al\.\(2011\)N\. Shervashidze, P\. Schweitzer, E\. J\. van Leeuwen, K\. Mehlhorn, and K\. M\. BorgwardtWeisfeiler\-Lehman graph kernels\.Journal of Machine Learning Research12,pp\. 2539–2561\.Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p2.1),[§4\.3](https://arxiv.org/html/2609.30484#S4.SS3.p5.1)\.
- Sunet al\.\(2019\)Z\. Sun, Z\. Deng, J\. Nie, and J\. TangRotatE: knowledge graph embedding by relational rotation in complex space\.InProceedings of the 7th International Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p2.1)\.
- Togninalliet al\.\(2019\)M\. Togninalli, E\. Ghisu, F\. Llinares\-López, B\. Rieck, and K\. BorgwardtWasserstein Weisfeiler–Lehman graph kernels\.InAdvances in Neural Information Processing Systems,Vol\.32,pp\. 6439–6449\.Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p2.1)\.
- Vipulanandanet al\.\(2026\)P\. Vipulanandan, K\. Premaratne, and D\. SarkarSemantic uncertainty quantification of hallucinations in llms: a quantum tensor network based method\.arXiv preprint arXiv:2601\.20026\.Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p1.1)\.
- Wallatet al\.\(2024\)J\. Wallat, M\. Heuss, M\. de Rijke, and A\. AnandCorrectness is not faithfulness in rag attributions\.arXiv preprint arXiv:2412\.18004\.Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p1.1)\.
- Wanget al\.\(2025\)J\. Wang, H\. Huang, and H\. ChenMESAQA: a dataset for multi\-span contextual and evidence\-grounded question answering\.InProceedings of the 31st International Conference on Computational Linguistics,Abu Dhabi, UAE,pp\. 10891–10901\.External Links:[Link](https://aclanthology.org/2025.coling-main.724/)Cited by:[§5\.1](https://arxiv.org/html/2609.30484#S5.SS1.p1.1)\.
- Wanget al\.\(2020\)W\. Wang, F\. Wei, L\. Dong, H\. Bao, N\. Yang, and M\. ZhouMiniLM: deep self\-attention distillation for task\-agnostic compression of pre\-trained transformers\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 5776–5788\.Cited by:[§5\.4](https://arxiv.org/html/2609.30484#S5.SS4.p1.1)\.
- Wei and Zou \(2019\)J\. Wei and K\. ZouEDA: easy data augmentation techniques for boosting performance on text classification tasks\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),Hong Kong, China,pp\. 6382–6388\.External Links:[Link](https://aclanthology.org/D19-1670/),[Document](https://dx.doi.org/10.18653/v1/D19-1670)Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p4.1)\.
- Yanet al\.\(2024\)J\. Yan, C\. Wang, J\. Huang, and W\. ZhangDo large language models understand logic or just mimick context?\.Note:arXiv preprint arXiv:2402\.12091Cited by:[§2](https://arxiv.org/html/2609.30484#S2.p1.1)\.
- Zhanget al\.\(2020\)T\. Zhang, V\. Kishore, F\. Wu, K\. Q\. Weinberger, and Y\. ArtziBERTScore: evaluating text generation with BERT\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§4\.3](https://arxiv.org/html/2609.30484#S4.SS3.p3.1),[§5\.4](https://arxiv.org/html/2609.30484#S5.SS4.p1.1)\.
- Zhanget al\.\(2019\)Y\. Zhang, J\. Baldridge, and L\. HePAWS: paraphrase adversaries from word scrambling\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 1298–1308\.External Links:[Document](https://dx.doi.org/10.18653/v1/N19-1131)Cited by:[§5\.3](https://arxiv.org/html/2609.30484#S5.SS3.p2.1)\.
- Zhuet al\.\(2024\)Y\. Zhu, J\. R\. A\. Moniz, S\. Bhargava, J\. Lu, D\. Piraviperumal, S\. Li, Y\. Zhang, H\. Yu, and B\. TsengCan large language models understand context?\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 2004–2018\.Cited by:[§1](https://arxiv.org/html/2609.30484#S1.p1.1),[§2](https://arxiv.org/html/2609.30484#S2.p1.1)\.

## Appendix AEffect of Temperature on Contextual Understanding

Tables[5](https://arxiv.org/html/2609.30484#A1.T5)and[6](https://arxiv.org/html/2609.30484#A1.T6)report mean GoldSim and ContextSim respectively for each model across temperature settingsT∈\{0\.0,0\.3,0\.7,1\.0\}T\\in\\\{0\.0,0\.3,0\.7,1\.0\\\}on both datasets\. Most models show little sensitivity to temperature, with score variations within±0\.01\\pm 0\.01–0\.020\.02across all settings\. The exception is Falcon\-7B on MesaQA, where GoldSim drops substantially from0\.60360\.6036atT=0\.0T=0\.0to0\.46620\.4662atT=1\.0T=1\.0, indicating that higher sampling randomness significantly degrades factual alignment for this model\. Scores tend to peak mildly atT=0\.3T=0\.3for most models before declining atT=1\.0T=1\.0, suggesting that a small degree of randomness can marginally improve contextual grounding without sacrificing factual accuracy\. PubMedQA scores are notably more stable across temperatures than MesaQA, likely due to the constrained nature of biomedical answers\.

Table 5:Mean GoldSim per model across temperatures\. CUS remains stable across temperature settings, withT=0\.0T=0\.0serving as a reliable default for controlled evaluation\.Table 6:Mean ContextSim per model across temperatures\. CtxSim is more sensitive to temperature than GoldSim, yet remains stable for most models, withT=0\.3T=0\.3yielding peak contextual faithfulness across both datasets before declining at higher temperatures\.
## Appendix BKG Construction Prompt

KGs are extracted using a structured chat\-style prompt inspired by[Sansford et al\. \(2024\)](https://arxiv.org/html/2609.30484#bib.bib5)and[Haskins and Adams \(2025\)](https://arxiv.org/html/2609.30484#bib.bib4)\. The prompt instructs the model to perform four sequential steps across all three input texts:

1. 1\.Entity detection:Extract all named entities, concepts, attributes, quantities, dates, locations, and roles comprehensively\.
2. 2\.Coreference resolution:Replace all pronouns with their referent entity names, using consistent labels across all three texts\.
3. 3\.Relation extraction:Identify semantic relationships as simple, concise phrases, decomposing compound sentences into one triplet per fact\.
4. 4\.Knowledge graph refinement:Where the same entity or relation appears across multiple graphs, use the same label consistently without merging distinct facts\.

The model is instructed to return a JSON object with exactly 3 keys \(knowledge\_graph1,knowledge\_graph2, andknowledge\_graph3\), each containing a list of\[subject, relation, object\]triples\. Few\-shot examples are included in the system prompt to ground the expected output format and label consistency behaviour\. The user turn specifies the three input sources explicitly:

- •TEXT1— reference answer\.
- •TEXT2— model\-generated response\.
- •TEXT3— supporting context\.

## Appendix CTAU Diagnosing Context and Answer Deviations

### C\.1Dataset and Annotation

We evaluate the TAU on a manually annotated subset of theMessaQAdataset, which contains general health\-related questions with long form answers\. For each QA instance, two KGs are constructed: \(i\) a gold KG extracted from the reference answer, and \(ii\) an LLM generated \(Llama and Gamma was used\) KG extracted from the model output\. The task of the triplet analysis unit is to identify semantically aligned triplet pairs between these two graphs\. To obtain reliable evaluation labels, we created a gold set of aligned triplet pairs \(GT↔\\leftrightarrowLLM\)\. Alignment was independently annotated by three medical students following a fixed guideline defining semantic equivalence at the triplet level\. Disagreements were resolved through adjudication\.

### C\.2Annotation Quality

Inter annotator agreement was measured over candidate aligned triplet pairs\. The results indicate strong consistency, with percent agreement of 0\.9433, pairwise F1 scores of 0\.9708, 0\.9825, and 0\.9882, and pairwise Jaccard scores of 0\.9433, 0\.9657, and 0\.9766, confirming the reliability of the annotations\.

### C\.3Evaluation Protocol

Performance is evaluated against the annotated alignments using precision, recall, and F1 score, reported using both micro averaged metrics \(aggregated over all triplets\) and macro averaged metrics \(computed per QA instance and averaged\), capturing both overall performance and consistency across samples\.

### C\.4Results

As shown in Table[7](https://arxiv.org/html/2609.30484#A3.T7), the KEA baseline achieves higher precision due to its conservative component wise matching, but exhibits low recall, missing many valid alignments\. In contrast, our sentence\-level alignment approach significantly improves recall \(\+34\.8%\), resulting in higher Micro and Macro F1 scores\. This improvement arises from robustness to lexical variation in relations\. For example, semantically equivalent relations such as“treats”and“used for”are often not aligned by KEA, whereas our method captures such equivalence through sentence level semantic similarity\. The lower precision reflects the expected trade\-off when moving from strict lexical matching to semantic matching, while the overall F1 improvement indicates a better balance between sensitivity and specificity\. Additionally, our method enables residual error analysis by categorizing mismatches into relation wrong and entity wrong types\.

Table 7:Comparison of TAU alignment performance between our method and the KEA baseline, evaluated on manually annotated MesaQA and PubMed subsets\.

## Appendix DPer\-Dataset Performance Tables

This appendix decomposes the headline results of Table[2](https://arxiv.org/html/2609.30484#S6.T2)into the 3 benchmark families, with F1 and AUROC shown in adjacent columns so that threshold\-dependent and ranking behaviour can be inspected side by side\. All datasets share the same protocol:N=400N=400balanced pairs, per\-method threshold sweep for F1, and AUROC on the raw scores\. For S3KG, the bestα\\alphaidentified in Appendix[E](https://arxiv.org/html/2609.30484#A5)is used per dataset, and the best score per column isbolded\. The Wikipedia Entity\-Swap table \(Table[9](https://arxiv.org/html/2609.30484#A4.T9)\) additionally reports precision and recall to expose ROUGE\-1’s near\-perfect score as a token\-overlap artifact rather than a meaningful signal\.

Table 8:Performance on Short\-Text Datasets\. S3KG leads on PAWS\-Wiki \(F1 0\.766, AUROC 0\.795\) but trails sentence\-T5\-base on MRPC and STS12, where dense semantic representations have a natural advantage over structural signals\.ROUGE\-1 achieves a near\-perfect score on this dataset due to a known artifact: entity\-swapped pairs differ only in the swapped entity tokens while sharing nearly identical surrounding surface\-form tokens, making unigram overlap trivially high\. ROUGE\-1 is therefore excluded from the meaningful comparison\.

Table 9:Results on Wikipedia Entity\-Swap \(N=400N=400\)\. S3KG attains the highest F1 and AUROC across all retained baselines; ROUGE\-1 is excluded as its near\-perfect score is a token\-overlap artifact rather than a meaningful signal\.Table 10:Performance on KG\-Perturbed Paragraph Datasets\. C400: SK\-Codex 400; Comb\.: SK\-Combined; Find: SK\-FindKG; Oreg\.: SK\-Oregano\. S3KG leads on four of five datasets; sentence\-T5\-base outperforms on SK\-FindKG, the only dataset where pure dense embeddings dominate \(α=0\.0\\alpha=0\.0is optimal\)\.
## Appendix EHyper\-parameter Selection

The mixing coefficientα\\alphain \([2](https://arxiv.org/html/2609.30484#S4.E2)\) is the sole hyper\-parameter of S3KG, controlling the trade\-off between the structural WL kernel signal \(α=0\.0\\alpha=0\.0\) and the SBERT semantic signal \(α=1\.0\\alpha=1\.0\), with intermediate values blending both\. We tuneα\\alphaper dataset via grid search over\{0\.0,0\.1,…,1\.0\}\\\{0\.0,0\.1,\\ldots,1\.0\\\}, selecting the value maximising binary\-classification F1; AUROC is reported alongside to confirm the result is not an artefact of threshold sensitivity\.

Table[11](https://arxiv.org/html/2609.30484#A5.T11)consolidates the full sweep across all nine benchmark datasets\. The bestα\\alphaper dataset \(selected by maximum F1\) isboldedtogether with its F1 / AUROC entries\. Three patterns emerge: \(i\) datasets dominated by perturbations \(SK\-FindKG, Wiki Swap, STS12\) favour lowα∈\{0\.0,0\.1\}\\alpha\\in\\\{0\.0,0\.1\\\}, where the WL signal carries most of the discriminative power; \(ii\) datasets where both relational and lexical paraphrase signals are simultaneously informative \(SK\-Codex 400, SK\-Combined, PAWS\-Wiki\) peak in the balanced rangeα∈\[0\.4,0\.6\]\\alpha\\in\[0\.4,0\.6\]; \(iii\) the AUROC surface is consistently flatter than the F1 surface, indicating thatα\\alphaprimarily reshapes the score distribution near the decision boundary\.

Table 11:Full S3KGα\\alphasweep across short\-text and combined datasets \(F1 / AUROC\)\. The bestα\\alphaper dataset \(by maximum F1\) isbolded\. Balanced datasets such as PAWS\-Wiki and SK\-Combined peak atα∈\[0\.4,0\.6\]\\alpha\\in\[0\.4,0\.6\], while STS12 favours the KG\-heavy end \(α=0\.1\\alpha=0\.1\)\.Table 12:Full S3KGα\\alphasweep across KG\-perturbed and entity\-swap datasets \(F1 / AUROC\)\. The bestα\\alphaper dataset \(by maximum F1\) isbolded\. Structure\-dominated datasets \(SK\-FindKG, Wiki Swap\) favour lowα∈\{0\.0,0\.1\}\\alpha\\in\\\{0\.0,0\.1\\\}, confirming the WL kernel’s advantage on relational perturbations\.

相似文章

超越准确率:大型语言模型统计推理的多维评估

arXiv cs.CL

本文提出了一个用于评估大型语言模型统计推理能力的多维评估框架,结合响应准确率、响应行为、结构主题建模和词汇相似性分析,应用于15个LLM和90道考试题。研究发现,仅靠准确率不足以刻画LLM的统计推理能力,并且存在厂商特有的风格差异。

基于外部子图生成的大语言模型逐步推理增强

arXiv cs.CL

本文提出了SGR框架,通过查询相关的子图生成将外部知识图谱与大语言模型相结合,融合基于Cypher的推理与协同推理集成,从而增强大语言模型的逐步推理能力。在CWQ、WebQSP、GrailQA和KQA Pro上的实验表明,该框架相比标准提示方法和知识增强基线具有更高的推理准确性。