XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?

arXiv cs.AI Papers

Summary

This paper introduces XAI-Arena, an LLM-as-a-judge framework for scalable and reproducible evaluation of explainable AI explanation quality, showing strong correlation with human judgments.

arXiv:2609.09428v1 Announce Type: new Abstract: Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.
Original Article
View Cached Full Text

Cached at: 09/11/26, 08:37 AM

# XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?
Source: [https://arxiv.org/abs/2609.09428](https://arxiv.org/abs/2609.09428)
[View PDF](https://arxiv.org/pdf/2609.09428)

> Abstract:Evaluating the quality of explanations produced by explainable AI \(XAI\) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies\. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations\. We introduce XAI\-Arena, an LLM\-as\-a\-judge framework for scalable, reproducible, multidimensional, and stakeholder\-sensitive evaluation of XAI explanation quality\. XAI\-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability\. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas\. Human validation shows a strong positive association between LLM\-generated and human ratings \(Spearman's rho=\.693, p<\.001\)\. Together, LLM\-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations\.

## Submission history

From: Stefan Feuerriegel \[[view email](https://arxiv.org/show-email/a7a15288/2609.09428)\] **\[v1\]**Tue, 8 Sep 2026 20:25:46 UTC \(2,119 KB\)

Similar Articles