Can Conversational XAI Improve User Performance? An Experimental Study
Summary
This paper presents an experimental study investigating whether conversational XAI assistants improve user performance in terms of prediction accuracy, model understanding, and error identification compared to Q&A-based assistance, with preliminary results showing no significant performance differences.
View Cached Full Text
Cached at: 05/21/26, 06:26 AM
# Can Conversational XAI Improve User Performance? An Experimental Study Source: [https://arxiv.org/abs/2605.20439](https://arxiv.org/abs/2605.20439) [View PDF](https://arxiv.org/pdf/2605.20439) > Abstract:Explainable AI \(XAI\) techniques aim to provide insights into predictive models and enhance user performance, yet they often fall short of these expectations\. Conversational XAI assistants promise to overcome such limitations, but empirical evidence on their impact on objective performance measures remains limited\. We propose an experimental design for evaluating explanation assistance through prediction accuracy, model understanding, and error identification\. Using an explainable\-by\-design prediction model, we create conditions where users can outperform the model by identifying and compensating for systematic errors\. We compare conversational assistance against Q&A\-based assistance to assess which better supports users in working with model explanations\. Preliminary results from testing our experimental design show that participants \(N=42\) in both treatments significantly outperformed the model but reveal no performance differences between assistance types and modest engagement overall\. These findings inform refinements for our planned full study, including enhanced engagement interventions and investigation of the mechanisms driving improved predictions\. ## Submission history From: Julian Rosenberger \[[view email](https://arxiv.org/show-email/80e6267b/2605.20439)\] **\[v1\]**Tue, 19 May 2026 19:47:17 UTC \(1,218 KB\)
Similar Articles
How AI Agents Reshape Knowledge Work: Autonomy, Efficiency, and Scope
This study uses production data from Perplexity to compare AI agents versus conversational assistants, finding that agents reduce completion time by 87% and costs by 94% while expanding the scope and quality of knowledge work.
Is AI actually getting better at understanding context in long conversations, or does it still fall apart?
This article discusses the limitations of AI models in maintaining context over long conversations, highlighting recency bias and the distinction between context window size and actual comprehension. It suggests practical workarounds like restating constraints and using running context documents.
Researchers gave 1,222 people AI assistants, then took them away after 10 minutes. Performance crashed below the control group and people stopped trying. UCLA, MIT, Oxford, and Carnegie Mellon call it the "boiling frog" effect.
A multi-institutional study of 1,222 participants found that brief AI assistant use (10 minutes) led to measurable cognitive decline and reduced effort on subsequent tasks compared to control groups, termed the 'boiling frog' effect. The research provides causal evidence that even short-term AI reliance may impair independent problem-solving performance.
Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
This paper introduces a conversational voice agent system that uses a lightweight on-device 'Talker' model to start responding immediately, then incorporates knowledge from a frontier LLM 'Reasoner' as it becomes available, achieving 7-19x faster time-to-first-response while approaching frontier-level performance on a laptop.
Do you find yourself genuinely building skills with AI assistance, or do you notice your baseline abilities getting softer over time because you reach for the tool first?
Reflection on whether using AI tools like ChatGPT and Claude genuinely builds skills or erodes baseline abilities due to over-reliance on shortcuts, comparing the phenomenon to calculators and search engines.