Tag
SWORD introduces a Wikidata-based benchmark to evaluate cross-lingual inconsistencies in LLMs' rejection of factual errors, revealing that models rely on distributional familiarity and show performance degradation in Asian languages.
This paper introduces Graph Alignment Topology as an inductive bias for grounding detection, using a graph neural network to model alignment structure between reference information and LLM outputs. The method achieves state-of-the-art results on multiple hallucination and question-answering datasets, outperforming GPT-4o.