XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
Summary
This paper introduces XAI-Arena, an LLM-as-a-judge framework for scalable and reproducible evaluation of explainable AI explanation quality, showing strong correlation with human judgments.
View Cached Full Text
Cached at: 09/11/26, 08:37 AM
# XAI-Arena: Can LLMs Assess the Quality of XAI Explanations? Source: [https://arxiv.org/abs/2609.09428](https://arxiv.org/abs/2609.09428) [View PDF](https://arxiv.org/pdf/2609.09428) > Abstract:Evaluating the quality of explanations produced by explainable AI \(XAI\) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies\. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations\. We introduce XAI\-Arena, an LLM\-as\-a\-judge framework for scalable, reproducible, multidimensional, and stakeholder\-sensitive evaluation of XAI explanation quality\. XAI\-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability\. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas\. Human validation shows a strong positive association between LLM\-generated and human ratings \(Spearman's rho=\.693, p<\.001\)\. Together, LLM\-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations\. ## Submission history From: Stefan Feuerriegel \[[view email](https://arxiv.org/show-email/a7a15288/2609.09428)\] **\[v1\]**Tue, 8 Sep 2026 20:25:46 UTC \(2,119 KB\)
Similar Articles
Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids
This paper investigates whether high-quality Natural Language Explanations (NLEs) generated by LLMs from XAI outputs actually improve task performance, finding they do not aid accuracy but inflate confidence, revealing a quality-usefulness gap.
Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
A systematic review of XAI research in the context of the EU Right to Explanation, analyzing gaps between legal requirements and technical implementations across GDPR and the AI Act.
Explainable artificial intelligence (XAI): From inherent explainability to large language models
This paper examines the progression from inherent explainability in artificial intelligence to the development and application of explainable methods for large language models.
Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap
This survey identifies three critical gaps in explainable AI for Arabic NLP—method, task, and linguistic—and proposes a taxonomy and research agenda for linguistically grounded explanations.
Evaluating Explainability in Safety-Critical ATR Systems: Limitations of Post-Hoc Methods and Paths Toward Robust XAI
This paper evaluates explainability methods in safety-critical Automatic Target Recognition (ATR) systems, highlighting the limitations of post-hoc techniques like saliency and attention maps. It proposes a taxonomy and assessment framework to address issues such as spurious explanations and instability, advocating for more robust, causally grounded XAI approaches.