Can Conversational XAI Improve User Performance? An Experimental Study

arXiv cs.LG Papers

Summary

This paper presents an experimental study investigating whether conversational XAI assistants improve user performance in terms of prediction accuracy, model understanding, and error identification compared to Q&A-based assistance, with preliminary results showing no significant performance differences.

arXiv:2605.20439v1 Announce Type: new Abstract: Explainable AI (XAI) techniques aim to provide insights into predictive models and enhance user performance, yet they often fall short of these expectations. Conversational XAI assistants promise to overcome such limitations, but empirical evidence on their impact on objective performance measures remains limited. We propose an experimental design for evaluating explanation assistance through prediction accuracy, model understanding, and error identification. Using an explainable-by-design prediction model, we create conditions where users can outperform the model by identifying and compensating for systematic errors. We compare conversational assistance against Q&A-based assistance to assess which better supports users in working with model explanations. Preliminary results from testing our experimental design show that participants (N=42) in both treatments significantly outperformed the model but reveal no performance differences between assistance types and modest engagement overall. These findings inform refinements for our planned full study, including enhanced engagement interventions and investigation of the mechanisms driving improved predictions.
Original Article
View Cached Full Text

Cached at: 05/21/26, 06:26 AM

# Can Conversational XAI Improve User Performance? An Experimental Study
Source: [https://arxiv.org/abs/2605.20439](https://arxiv.org/abs/2605.20439)
[View PDF](https://arxiv.org/pdf/2605.20439)

> Abstract:Explainable AI \(XAI\) techniques aim to provide insights into predictive models and enhance user performance, yet they often fall short of these expectations\. Conversational XAI assistants promise to overcome such limitations, but empirical evidence on their impact on objective performance measures remains limited\. We propose an experimental design for evaluating explanation assistance through prediction accuracy, model understanding, and error identification\. Using an explainable\-by\-design prediction model, we create conditions where users can outperform the model by identifying and compensating for systematic errors\. We compare conversational assistance against Q&A\-based assistance to assess which better supports users in working with model explanations\. Preliminary results from testing our experimental design show that participants \(N=42\) in both treatments significantly outperformed the model but reveal no performance differences between assistance types and modest engagement overall\. These findings inform refinements for our planned full study, including enhanced engagement interventions and investigation of the mechanisms driving improved predictions\.

## Submission history

From: Julian Rosenberger \[[view email](https://arxiv.org/show-email/80e6267b/2605.20439)\] **\[v1\]**Tue, 19 May 2026 19:47:17 UTC \(1,218 KB\)

Similar Articles

Researchers gave 1,222 people AI assistants, then took them away after 10 minutes. Performance crashed below the control group and people stopped trying. UCLA, MIT, Oxford, and Carnegie Mellon call it the "boiling frog" effect.

Reddit r/artificial

A multi-institutional study of 1,222 participants found that brief AI assistant use (10 minutes) led to measurable cognitive decline and reduced effort on subsequent tasks compared to control groups, termed the 'boiling frog' effect. The research provides causal evidence that even short-term AI reliance may impair independent problem-solving performance.