The complexities of patient-centred conversational artificial intelligence

arXiv cs.AI Papers

Summary

This paper analyzes 2,053 real patient-chatbot conversations to show that communication styles vary widely and can significantly alter triage outcomes, finding that patient simulators that model emotional state and conversational strategy produce conversations nearly indistinguishable from real ones in a Turing test.

arXiv:2607.08625v1 Announce Type: new Abstract: Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients. We analysed 2,053 real patient-chatbot conversations and found that communication patterns and expression of emotions vary widely across users. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style. In a Turing-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%. We used five distinct patient personae, across 1,164 clinician-graded cases, to evaluate the performance of four LLMs in urgency assessment. We found that communication style can significantly alter triage outcomes. Patient-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world.
Original Article
View Cached Full Text

Cached at: 07/10/26, 06:09 AM

# The complexities of patient-centred conversational artificial intelligence
Source: [https://arxiv.org/abs/2607.08625](https://arxiv.org/abs/2607.08625)
Authors:[João Matos](https://arxiv.org/search/cs?searchtype=author&query=Matos,+J),[Olivia Buege](https://arxiv.org/search/cs?searchtype=author&query=Buege,+O),[Donny Cheung](https://arxiv.org/search/cs?searchtype=author&query=Cheung,+D),[Gary S\. Collins](https://arxiv.org/search/cs?searchtype=author&query=Collins,+G+S),[Paula Dhiman](https://arxiv.org/search/cs?searchtype=author&query=Dhiman,+P),[Nan Li](https://arxiv.org/search/cs?searchtype=author&query=Li,+N),[Bingyu Mao](https://arxiv.org/search/cs?searchtype=author&query=Mao,+B),[Benjamin W\. Nelson](https://arxiv.org/search/cs?searchtype=author&query=Nelson,+B+W),[Michail Ouroutzoglou](https://arxiv.org/search/cs?searchtype=author&query=Ouroutzoglou,+M),[Paul Varghese](https://arxiv.org/search/cs?searchtype=author&query=Varghese,+P),[Jonathan Amar](https://arxiv.org/search/cs?searchtype=author&query=Amar,+J)

[View PDF](https://arxiv.org/pdf/2607.08625)

> Abstract:Consumer\-facing health chatbots powered by large language models \(LLMs\) are increasingly used for symptom assessment\. However, chatbot development and evaluation often rely on cooperative, articulate, simulated patients\. We analysed 2,053 real patient\-chatbot conversations and found that communication patterns and expression of emotions vary widely across users\. We developed a patient simulator that separately models clinical content, emotional state, conversational strategy, and communication style\. In a Turing\-inspired evaluation of realism with 15 human graders, simulated conversations were nearly indistinguishable from real ones, with human graders achieving an accuracy of 55%\. We used five distinct patient personae, across 1,164 clinician\-graded cases, to evaluate the performance of four LLMs in urgency assessment\. We found that communication style can significantly alter triage outcomes\. Patient\-centred conversational artificial intelligence must accommodate communication diversity: systems designed for idealised, rather than realistic, interactions risk underperforming and amplifying health disparities when deployed in the real world\.

## Submission history

From: João Matos \[[view email](https://arxiv.org/show-email/acdd0efe/2607.08625)\] **\[v1\]**Thu, 9 Jul 2026 15:56:55 UTC \(2,716 KB\)

Similar Articles

Training AI chatbots to be warm and empathetic makes them less factually accurate

Reddit r/artificial

New research shows that training AI chatbots to be warmer and more empathetic significantly reduces their factual accuracy, leading to higher error rates in medical advice and increased agreement with user misconceptions. The findings challenge the common assumption that conversational style can be adjusted without compromising factual correctness.

Imperfectly Cooperative Human-AI Interactions: Comparing the Impacts of Human and AI Attributes in Simulated and User Studies

arXiv cs.CL

This research paper investigates how human personality traits and AI design characteristics jointly impact human-AI interactions in imperfectly cooperative scenarios using both simulated datasets (2,000 simulations) and human subjects experiments (290 participants). The study finds significant divergences between simulation and real-world interactions, with AI transparency emerging as a critical factor in actual human-AI encounters.

The other half of AI safety

Hacker News Top

The article critiques the AI safety field's focus on catastrophic risks while neglecting everyday mental health harms from chatbots like ChatGPT, citing OpenAI's own data on millions of users showing signs of psychosis, mania, or suicidal ideation yet receiving only redirects instead of hard gating.