ConsultMind:Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning
Summary
This paper presents ConsultMind and AutoDisym, frameworks that leverage Bayesian networks and uncertainty-aware reasoning to automate diagnostic consultation, demonstrating improved diagnostic accuracy and explanation quality across medical domains.
View Cached Full Text
Cached at: 09/28/26, 09:47 AM
# ConsultMind: Towards Automated Diagnostic Consultation via Uncertainty-Aware Reasoning
Source: [https://arxiv.org/html/2609.30796](https://arxiv.org/html/2609.30796)
Yuming YangYun ChenJiang ZhongJunnan ZhuXinyi JiangHaoyang ZengRuirui ChenYining WangXinyu ZhouRong TangKaiwen Wei
###### Abstract
Diagnostic consultation is an online sequential decision\-making process in which clinicians gather evidence through patient interaction until a diagnosis is sufficiently supported\. Automating this process requires adaptive inquiry and interpretable decisions\. Bayesian networks offer a natural foundation by updating diagnostic posteriors as evidence accumulates, but their use in open\-ended consultation raises two challenges: linking diagnostic hypotheses to potential inquiries and translating evolving posteriors into consultation decisions\. We introduceAutoDisym, an automated pipeline that integrates diagnostic knowledge with heterogeneous diagnosis\-labeled clinical narratives to construct a Disorder–Symptom Bayesian Network \(DSBN\)\. Building on the DSBN, we proposeConsultMind, an uncertainty\-aware framework that updates disorder posteriors after each response and uses posterior uncertainty to guide inquiry and diagnosis\. We evaluate both methods across psychiatry, respiratory medicine, fever clinics, and three public datasets\. The results show that AutoDisym can automatically construct high\-quality DSBNs and that ConsultMind consistently improves diagnostic performance and explanation soundness\. For example, AutoDisym achieves macro\-averaged F1 scores of 81\.37 for canonical symptoms and 72\.19 for manifestations using GPT\-5\.6\-Sol\. ConsultMind improves Top\-1 and Top\-3 diagnostic accuracy by up to 22\.15 and 37\.89 percentage points, respectively\. Physician evaluation further shows that ConsultMind improves the quality of ranking explanations, differential diagnoses, and diagnosis rationales across LLMs of different scales\. This work offers a promising approach to automatic diagnostic consultation\.
1College of Computer Science, Chongqing University
2College of Computer Science, Hunan University3AI Labs, Unisound
4Institute of Automation, Chinese Academy of Sciences5Psycharity, Chongqing Medical University
6School of Computer Science and Engineering, University of New South Wales
7Institute of Advanced Intelligence and Computing \(IAIC\),
Agency for Science, Technology and Research \(A\*STAR\), Singapore
8Technological Information, Chongqing Center for Disease Control and Prevention
sunx@stu\.cqu\.edu\.cn, \{zhongjiang, weikaiwen\}@cqu\.edu\.cn
$\\dagger$$\\dagger$footnotetext:Co\-corresponding authors\.## Introduction
Figure 1:Comparison of automated diagnostic consultation methods\.Protocol\- and LLM\-driven systems suffer from rigid inquiry and opaque decisions, respectively, whereasConsultMindenables adaptive and interpretable consultation through uncertainty\-aware reasoning\.Diagnostic consultation often begins with patients providing partial or imprecise descriptions of their conditions, requiring clinically relevant information to be recovered through multi\-turn interaction\([Zeng et al\. 2020](https://arxiv.org/html/2609.30796#bib.bib33);[Shi et al\. 2023](https://arxiv.org/html/2609.30796#bib.bib34)\)\. Clinicians respond with targeted questions about symptoms and medical history, progressively narrowing plausible diagnoses\([Tang et al\. 2016](https://arxiv.org/html/2609.30796#bib.bib37);[Wei et al\. 2018](https://arxiv.org/html/2609.30796#bib.bib38);[Yuan and Yu 2024](https://arxiv.org/html/2609.30796#bib.bib27)\)\. Consultation is therefore an iterative decision\-making process that gathers evidence until it is sufficient to support a diagnosis\([Li et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib22);[Werthaim et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib24)\)\.
Recent advances in large language models \(LLMs\) have accelerated research on automatic diagnostic consultation\([Tu et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib30);[Saab et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib29)\), in which a system interacts with a patient through successive inquiries until it can formulate a diagnostic hypothesis\([Li et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib22);[Werthaim et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib24)\)\. Existing systems generally follow either a protocol\-driven\([Wei et al\. 2018](https://arxiv.org/html/2609.30796#bib.bib38)\)or an LLM\-driven paradigm\([Ren et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib31)\)\. Protocol\-driven systems organize consultations around structured symptom spaces, symptom checklists, or diagnostic pathways, offering standardized and controllable procedures\([Yuan and Yu 2024](https://arxiv.org/html/2609.30796#bib.bib27)\)\. LLM\-driven systems generate questions and diagnostic outputs from the dialogue context, allowing greater adaptation to patient\-specific information\([Sanghvi et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib26)\)\.
Despite these strengths, recent evaluations have identified persistent weaknesses in the inquiry quality, clinical reasoning, and decision reliability of LLM\-based consultation systems\([Johri et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib35);[Johri et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib23)\)\. Existing systems have two key limitations\. \(1\)Rigid inquiry\. Protocol\-driven systems confine inquiries to predefined sequences or transition rules\([Xia et al\. 2020](https://arxiv.org/html/2609.30796#bib.bib39)\), potentially overlooking diagnostically informative symptoms that fall outside the prescribed paths\. \(2\)Opaque decisions\. LLM\-driven systems derive consultation decisions directly from the dialogue context, leaving the evolving diagnostic state and the rationale for each decision opaque\([Gong et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib25);[Qiao et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib36)\)\. Diagnostic consultation, however, is an online sequential decision\-making process in which clinicians update their diagnostic assessment after each patient response and use it to guide the next decision\([Sun et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib28)\)\. Existing systems therefore lack a mechanism that tracks the evolving diagnostic state, adapts inquiry accordingly, and makes each decision interpretable\.
As shown in Figure[1](https://arxiv.org/html/2609.30796#Sx1.F1), within a structured variable space, Bayesian inference updates disorder posteriors as evidence accumulates, while explicit probabilistic dependencies make the basis of each update inspectable\. Diagnostic consultation, however, lacks a predefined variable space because patient descriptions and potential inquiries are open\-ended\. Applying Bayesian networks to this setting therefore raises two unresolved questions:\(1\) how to link diagnostic hypotheses to potential inquiriesand\(2\) how to translate evolving posteriors into consultation decisions\.
To address the first question, we introduceAutoDisym, an automated pipeline that integrates diagnostic knowledge with diagnosis\-labeled clinical narratives to construct a Disorder–Symptom Bayesian Network \(DSBN\)\. The DSBN represents diagnostic hypotheses as candidate disorders and potential inquiries as clinically grounded symptom variables\. Through collaboration among specialized agents, AutoDisym derives a canonical symptom schema from diagnostic knowledge and maps symptom observations from heterogeneous clinical narratives, including consultation dialogues, electronic medical records, and clinical case, onto this schema\. Finally, AutoDisym refines the schema through corpus\-level feedback from recurrent unmatched observations and uses the grounded observations to parameterize the DSBN\.
Building on this, we further proposeConsultMind, an uncertainty\-aware reasoning framework that explicitly grounds each consultation decision in the evolving disorder posterior\. ConsultMind comprises two mechanisms: \(1\)Uncertainty Awareness, which updates this posterior after each patient response and quantifies the remaining diagnostic uncertainty; and \(2\)Adaptive Reasoning, which supports three inquiry strategies and dynamically selects among them based on the evolving posterior state\.Explorationseeks potentially relevant symptoms;Differentiationtargets symptoms that distinguish competing hypotheses; andConsolidationgathers further evidence for the leading hypothesis\. Together, these mechanisms enable ConsultMind to guide an LLM through successive consultation decisions and provide a diagnosis once sufficient evidence has accumulated\.
We evaluate AutoDisym and ConsultMind across psychiatry, respiratory medicine, fever clinics, and three public datasets\. The results show that AutoDisym improves DSBN construction quality and that ConsultMind consistently improves diagnostic performance and explanation soundness\. For example,AutoDisymachieves macro\-averaged F1 scores of 81\.37 for canonical symptoms and 72\.19 for manifestations using GPT\-5\.6\-Sol\. ConsultMind improves Top\-1 and Top\-3 diagnostic accuracy by up to 22\.15 and 37\.89 percentage points, respectively\. Physician evaluation further shows that ConsultMind improves ranking explanations, differential diagnoses, and diagnosis rationales across LLMs, resulting in more clinically sound diagnostic explanations\. Our contributions are as follows:
- •We introduceAutoDisym, an automated pipeline that constructs Disorder–Symptom Bayesian Networks from diagnostic knowledge and clinical narratives, linking diagnostic hypotheses to potential inquiries\.
- •Building on these networks, we proposeConsultMind, an uncertainty\-aware reasoning framework that translates evolving disorder posteriors into adaptive and interpretable consultation decisions\.
- •Across three clinical settings and external datasets, evaluations show that AutoDisym improves DSBN construction quality and that ConsultMind consistently improves diagnostic performance and explanation soundness\.
## Related Work
Medical knowledge has been incorporated into LLM\-based diagnostic systems through several complementary paradigms\([Singhal et al\. 2023](https://arxiv.org/html/2609.30796#bib.bib32)\)\. In\-context approaches embed clinical guidelines\([Kresevic et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib1);[Wang et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib2)\)or diagnostic criteria directly into prompts\([Li et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib3);[Savage et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib4)\), whereas retrieval\-augmented generation retrieves case\-relevant evidence from external medical corpora\([Xiong et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib5);[Gaber et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib6)\)\. Knowledge\-graph\-based retrieval further represents diseases\([Jia et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib7)\), symptoms\([Song et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib8)\), examinations\([Gao et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib9)\), and treatments through structured relations\([Alber et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib10);[Zhou et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib11)\), enabling relation\-aware grounding\([Jiang et al\. 2023](https://arxiv.org/html/2609.30796#bib.bib12)\)\. However, existing systems exhibit rigid and insufficiently adaptive inquiry\.
Automatic diagnostic consultation has evolved from structured control\([Sohn et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib13)\)to learned and generative decision\-making\([Xu et al\. 2023](https://arxiv.org/html/2609.30796#bib.bib14)\)\. Protocol\-driven systems organize interactions\([Andreadis et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib15)\)around clinical scales\([Sittig et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib16)\), symptom checklists\([Ben\-Shabat et al\. 2022](https://arxiv.org/html/2609.30796#bib.bib17)\), decision trees\([You et al\. 2023](https://arxiv.org/html/2609.30796#bib.bib18)\), and medical flowcharts\([Liu et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib19)\)\. Reinforcement\-learning and task\-oriented dialogue methods\([Hou et al\. 2023](https://arxiv.org/html/2609.30796#bib.bib20)\)learn symptom\-inquiry policies within predefined state and action spaces, whereas LLM\-based systems generate follow\-up questions and diagnoses directly from the dialogue context\([Tu et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib21)\)\. However, LLM\-based consultation decisions often remain opaque and difficult to interpret\.
## AutoDisym
AutoDisymuses specialized agents to construct a Disorder–Symptom Bayesian Network \(DSBN\) from diagnostic knowledge and labeled clinical narratives\.
#### Data Preparation\.
AutoDisymtakes two inputs: a curated diagnostic knowledge repository and a corpus of diagnosis\-labeled clinical narratives\. The repository contains diagnostic criteria from the International Classification of Diseases, 11th Revision \(ICD\-11\), clinical guidelines, and selected consensus statements, with𝒦ver\\mathcal\{K\}\_\{\\mathrm\{ver\}\}denoting the guideline and consensus subset\. The corpus includes consultation dialogues, electronic medical records, and case reports\. Let𝒟=\{dk\}k=1K\\mathcal\{D\}=\\\{d\_\{k\}\\\}\_\{k=1\}^\{K\}be the set of candidate disorders\. We represent the corpus as𝒞=\{\(ni,yi\)\}i=1N\\mathcal\{C\}=\\\{\(n\_\{i\},y\_\{i\}\)\\\}\_\{i=1\}^\{N\}, wherenin\_\{i\}is a clinical narrative andyi∈𝒟y\_\{i\}\\in\\mathcal\{D\}is its primary diagnosis\. From these inputs,AutoDisymconstructs the DSBN in four stages\. Details are provided in Appendix\.
Figure 2:Overview ofAutoDisym\. An automated pipeline that constructs Disorder–Symptom Bayesian Networks \(DSBN\) from diagnostic knowledge and clinical narratives\.
#### Stage 1: Knowledge\-Guided Schema Construction\.
For each disorderdk∈𝒟d\_\{k\}\\in\\mathcal\{D\}, the Symptom Schema Agent \(SSA\) processes its diagnostic criteria in two steps:
- •Initialization:SSA invokesExtract\(dk\)\\operatorname\{Extract\}\(d\_\{k\}\)and dynamically orchestrates subagents to identify symptoms and their initial manifestations\.
- •Schema refinement:SSA clusters similar candidate expressions and assigns each cluster to an adjudication subagent, whichmerges,retains, orremoves\.
SSA consolidates these inventories into the canonical schema𝒮=\{Sj\}j=1M\\mathcal\{S\}=\\\{S\_\{j\}\\\}\_\{j=1\}^\{M\}, where eachSjS\_\{j\}has a manifestation state space𝒱j\\mathcal\{V\}\_\{j\}\(details see Appendix\)\.
#### Stage 2: Observation Extraction and Grounding\.
The Clinical Observation Agent \(COA\) segments each narrativenin\_\{i\}and extracts symptom mentions, manifestations, and supporting spans\. It retrieves related canonical symptoms and grounds each observation to a symptom–state pair, yielding three outcomes:
- •Match:Maps the observation to\(Sj,v\)\(S\_\{j\},v\), wherev∈𝒱jv\\in\\mathcal\{V\}\_\{j\}\.
- •Unobserved:Leaves an unmentioned canonical symptom unassigned\.
- •Abstention:Defers an observation with no reliable mapping to the current schema\.
COA retains each abstention and its supporting span as either a manifestation candidate\(Sj,v~\)\(S\_\{j\},\\tilde\{v\}\)or a symptom candidate\(S~,v~\)\(\\tilde\{S\},\\tilde\{v\}\)for schema verification\.
#### Stage 3: Retrieval\-Augmented Schema Verification\.
The Schema Verification Agent \(SVA\) verifies COA’s manifestation and symptom candidates against𝒦ver\\mathcal\{K\}\_\{\\mathrm\{ver\}\}\. For retrieval,𝒦ver\\mathcal\{K\}\_\{\\mathrm\{ver\}\}is divided into fixed\-size chunksℬver\\mathcal\{B\}\_\{\\mathrm\{ver\}\}and indexed by the BGE\-M3\([Chen et al\. 2024a](https://arxiv.org/html/2609.30796#bib.bib51)\)encoderf\(⋅\)f\(\\cdot\)\. For each candidatecc, SVA builds a queryqcq\_\{c\}from the candidate and its supporting span, then retrieves theLLmost similar chunks:
ℛc=Top\-Lp∈ℬversim\(f\(qc\),f\(p\)\)\.\\mathcal\{R\}\_\{c\}=\\underset\{p\\in\\mathcal\{B\}\_\{\\mathrm\{ver\}\}\}\{\\operatorname\{Top\}\\text\{\-\}L\}\\operatorname\{sim\}\\bigl\(f\(q\_\{c\}\),f\(p\)\\bigr\)\.\(1\)
Evidence\-based schema expansion\.Usingℛc\\mathcal\{R\}\_\{c\}, SVA accepts candidates supported by clinical evidence and rejects the others\. This process captures guideline\- or consensus\-supported symptoms and manifestations absent from the diagnostic criteria\. SVA sends accepted candidates to SSA, which addsv~\\tilde\{v\}to𝒱j\\mathcal\{V\}\_\{j\}forc=\(Sj,v~\)c=\(S\_\{j\},\\tilde\{v\}\)or creates a symptom–state entry forc=\(S~,v~\)c=\(\\tilde\{S\},\\tilde\{v\}\)\. COA then re\-grounds the affected observations and records each candidate’s frequency\. Details are provided in Appendix\.
#### Stage 4: Network Construction and Parameterization\.
Finally, the canonical schema, grounded observations, and diagnosis labels define the DSBN\. LetDDdenote the disorder variable and𝐒=\(S1,…,SM\)\\mathbf\{S\}=\(S\_\{1\},\\ldots,S\_\{M\}\)the symptom variables, with edgesD→SjD\\rightarrow S\_\{j\}\. The joint distribution factorizes as
P\(D,𝐒\)=P\(D\)∏j=1MP\(Sj∣D\)\.P\(D,\\mathbf\{S\}\)=P\(D\)\\prod\_\{j=1\}^\{M\}P\(S\_\{j\}\\mid D\)\.\(2\)
The prior and conditional distributions are estimated from diagnosis labels and grounded observations using add\-one smoothing \(details see Appendix\)\.
## ConsultMind
Figure 3:Automated Diagnostic Consultation viaConsultMind\. It grounds patient responses into canonical symptom states, updates diagnostic posteriors through the DSBN, and uses posterior uncertainty to guide inquiries until reaching a diagnosis\.ConsultMindconducts sequential consultations using the DSBN constructed by AutoDisym\. At each turn,Uncertainty Awarenessgrounds the patient response and updates the diagnostic state, whichAdaptive Reasoninguses to continue the inquiry or make a diagnosis\.
### Uncertainty Awareness
Given a responsertr\_\{t\}to a query about symptomSjtS\_\{j\_\{t\}\}, an LLM\-based grounding module mapsrtr\_\{t\}to a canonical statev∈𝒱jtv\\in\\mathcal\{V\}\_\{j\_\{t\}\}\. If no reliable mapping is supported, the module abstains without changing the evidence\. Explicit negations map toabsent, whereas unmentioned symptoms remain unobserved\. Let𝒪t⊆\{1,…,M\}\\mathcal\{O\}\_\{t\}\\subseteq\\\{1,\\ldots,M\\\}denote the symptoms observed by turntt, with grounded evidence
𝐞t=\{Sj=et,j∣j∈𝒪t\}\.\\mathbf\{e\}\_\{t\}=\\\{S\_\{j\}=e\_\{t,j\}\\mid j\\in\\mathcal\{O\}\_\{t\}\\\}\.
For each binary disorder nodeDkD\_\{k\}, DSBN inference computesqt,k=P\(Dk=1∣𝐞t\)q\_\{t,k\}=P\(D\_\{k\}=1\\mid\\mathbf\{e\}\_\{t\}\)\. Normalizing these values gives the diagnostic posterior:
pt\(dk\)=qt,k∑ℓ=1Kqt,ℓ\.p\_\{t\}\(d\_\{k\}\)=\\frac\{q\_\{t,k\}\}\{\\sum\_\{\\ell=1\}^\{K\}q\_\{t,\\ell\}\}\.\(3\)
Only grounded observations update the posterior; without evidence, it reduces to the normalized disorder priors\. The remaining uncertainty is measured by posterior entropy:
Ht=−∑k=1Kpt\(dk\)log2pt\(dk\)\.H\_\{t\}=\-\\sum\_\{k=1\}^\{K\}p\_\{t\}\(d\_\{k\}\)\\log\_\{2\}p\_\{t\}\(d\_\{k\}\)\.\(4\)here,HtH\_\{t\}measures diagnostic uncertainty and informs subsequent reasoning\.
### Adaptive Reasoning
ConsultMindselects a strategy\-specific inquiry target from the evolving diagnostic state\. It encodes the posterior and accumulated evidence into a six\-dimensional state vector𝐳t\\mathbf\{z\}\_\{t\}, then compares it with prototypes for three strategies \(details see Appendix\):𝒜=\{exp,diff,con\}\\mathcal\{A\}=\\\{\\mathrm\{exp\},\\mathrm\{diff\},\\mathrm\{con\}\\\}, corresponding toExploration,Differentiation, andConsolidation\. It computes weighted cosine similarities with additional emphasis on evidence density, normalizes them with softmax, and selects the highest\-probability strategyat∈𝒜a\_\{t\}\\in\\mathcal\{A\}\. For symptom selection, let𝒯t\\mathcal\{T\}\_\{t\}denote the top\-NNdisorders,𝒰t=\{1,…,M\}∖𝒪t\\mathcal\{U\}\_\{t\}=\\\{1,\\ldots,M\\\}\\setminus\\mathcal\{O\}\_\{t\}the unobserved symptoms, anddt⋆=argmaxd∈𝒟pt\(d\)d\_\{t\}^\{\\star\}=\\arg\\max\_\{d\\in\\mathcal\{D\}\}p\_\{t\}\(d\)the leading disorder\.
#### Exploration\.
Exploration prioritizes symptoms characteristic of at least one leading disorder relative to the others:
sexp\(j\)=maxd∈𝒯tP\(Sj≠abs∣d\)P\(Sj≠abs∣¬d\)\.s\_\{\\mathrm\{exp\}\}\(j\)=\\max\_\{d\\in\\mathcal\{T\}\_\{t\}\}\\frac\{P\(S\_\{j\}\\neq\\mathrm\{abs\}\\mid d\)\}\{P\(S\_\{j\}\\neq\\mathrm\{abs\}\\mid\\neg d\)\}\.\(5\)
#### Differentiation\.
Differentiation prioritizes symptoms with the greatest distributional differences among competing:
sdiff\(j\)=meand≠d′∈𝒯tJS\(CPDj\(d\),CPDj\(d′\)\)\.s\_\{\\mathrm\{diff\}\}\(j\)=\\operatorname\*\{mean\}\_\{d\\neq d^\{\\prime\}\\in\\mathcal\{T\}\_\{t\}\}\\mathrm\{JS\}\\\!\\left\(\\mathrm\{CPD\}\_\{j\}\(d\),\\mathrm\{CPD\}\_\{j\}\(d^\{\\prime\}\)\\right\)\.\(6\)
#### Consolidation\.
Consolidation seeks further evidence for the leading disorder by prioritizing its likely symptoms:
scon\(j\)=P\(Sj≠abs∣dt⋆\)\.s\_\{\\mathrm\{con\}\}\(j\)=P\(S\_\{j\}\\neq\\mathrm\{abs\}\\mid d\_\{t\}^\{\\star\}\)\.\(7\)here,abs\\mathrm\{abs\}is the absent state, andCPDj\(d\)\\mathrm\{CPD\}\_\{j\}\(d\)is the categorical distribution ofSjS\_\{j\}under disorderdd\. Given the selected strategyata\_\{t\}, the next inquiry target is
jt\+1=argmaxj∈𝒰tsat\(j\)\.j\_\{t\+1\}=\\arg\\max\_\{j\\in\\mathcal\{U\}\_\{t\}\}s\_\{a\_\{t\}\}\(j\)\.\(8\)
An LLM convertsSjt\+1S\_\{j\_\{t\+1\}\}into a patient\-facing question, whose response begins the next turn\.
#### Termination\.
After each posterior update, ConsultMind decides whether to continue based on the current uncertainty and its reduction over a rolling window\. The consultation ends when uncertainty is sufficiently low or no longer decreases; symptom exhaustion and the turn limit serve as fallback conditions \(details see Appendix\)\. It then returns the disorders ranked by posterior probability:
𝝆t=argsortd∈𝒟↓pt\(d\)\.\\boldsymbol\{\\rho\}\_\{t\}=\\operatorname\{argsort\}^\{\\downarrow\}\_\{d\\in\\mathcal\{D\}\}p\_\{t\}\(d\)\.\(9\)
## Experiments
PsychiatryRespiratory MedicineFever ClinicConfigurationTop\-1↑\\uparrowTop\-3↑\\uparrowMRR↑\\uparrowTop\-1↑\\uparrowTop\-3↑\\uparrowMRR↑\\uparrowTop\-1↑\\uparrowTop\-3↑\\uparrowMRR↑\\uparrowQwen3\-8BDirect21\.05%35\.79%0\.27917\.42%31\.64%0\.25514\.26%26\.91%0\.222w/ RAG28\.31%\+7\.2651\.11%\+15\.320\.383\+0\.10422\.27%\+4\.8542\.56%\+10\.920\.328\+0\.07319\.01%\+4\.7536\.38%\+9\.470\.284\+0\.062w/ConsultMind41\.80%\+20\.7569\.84%\+34\.050\.538\+0\.25931\.28%\+13\.8655\.91%\+24\.270\.438\+0\.18327\.84%\+13\.5847\.95%\+21\.040\.376\+0\.154Llama\-3\.1\-8BDirect20\.53%40\.53%0\.30016\.08%35\.87%0\.27413\.68%30\.42%0\.238w/ RAG28\.28%\+7\.7551\.24%\+10\.710\.388\+0\.08821\.88%\+5\.8043\.77%\+7\.900\.334\+0\.06019\.03%\+5\.3537\.37%\+6\.950\.289\+0\.051w/ConsultMind42\.68%\+22\.1564\.33%\+23\.800\.521\+0\.22132\.65%\+16\.5753\.42%\+17\.550\.423\+0\.14928\.96%\+15\.2845\.86%\+15\.440\.365\+0\.127Qwen3\-32BDirect33\.33%40\.32%0\.36731\.64%36\.91%0\.35126\.84%33\.28%0\.319w/ RAG39\.75%\+6\.4256\.30%\+15\.980\.470\+0\.10336\.22%\+4\.5848\.91%\+12\.000\.420\+0\.06931\.07%\+4\.2343\.00%\+9\.720\.376\+0\.057w/ConsultMind51\.67%\+18\.3475\.83%\+35\.510\.625\+0\.25844\.72%\+13\.0863\.58%\+26\.670\.524\+0\.17338\.92%\+12\.0854\.87%\+21\.590\.461\+0\.142Gemini\-3\.1\-ProDirect37\.37%40\.00%0\.38534\.18%37\.26%0\.36727\.18%33\.54%0\.321w/ RAG42\.34%\+4\.9757\.05%\+17\.050\.483\+0\.09838\.29%\+4\.1149\.48%\+12\.220\.437\+0\.07031\.19%\+4\.0142\.67%\+9\.130\.374\+0\.053w/ConsultMind51\.58%\+14\.2177\.89%\+37\.890\.629\+0\.24445\.91%\+11\.7364\.42%\+27\.160\.543\+0\.17638\.64%\+11\.4653\.82%\+20\.280\.454\+0\.133DeepSeek\-V4\-ProDirect35\.48%41\.40%0\.38232\.03%37\.84%0\.35927\.92%34\.16%0\.329w/ RAG40\.01%\+4\.5356\.40%\+15\.000\.469\+0\.08735\.65%\+3\.6248\.38%\+10\.540\.416\+0\.05731\.40%\+3\.4842\.61%\+8\.450\.375\+0\.046w/ConsultMind48\.42%\+12\.9474\.74%\+33\.340\.600\+0\.21842\.36%\+10\.3361\.27%\+23\.430\.501\+0\.14237\.86%\+9\.9452\.94%\+18\.780\.443\+0\.114GPT\-5\.6\-SolDirect33\.68%38\.42%0\.36032\.48%38\.76%0\.36929\.78%36\.42%0\.356w/ RAG39\.58%\+5\.9053\.81%\+15\.390\.456\+0\.09637\.86%\+5\.3850\.65%\+11\.890\.441\+0\.07234\.85%\+5\.0747\.11%\+10\.690\.424\+0\.068w/ConsultMind50\.53%\+16\.8572\.63%\+34\.210\.599\+0\.23947\.86%\+15\.3865\.18%\+26\.420\.549\+0\.18044\.26%\+14\.4860\.18%\+23\.760\.526\+0\.170Claude\-Sonnet\-5Direct34\.74%43\.68%0\.39033\.27%39\.18%0\.37429\.13%35\.68%0\.344w/ RAG41\.19%\+6\.4560\.02%\+16\.340\.496\+0\.10638\.54%\+5\.2750\.74%\+11\.560\.445\+0\.07133\.77%\+4\.6445\.47%\+9\.790\.405\+0\.061w/ConsultMind53\.16%\+18\.4280\.00%\+36\.320\.654\+0\.26548\.34%\+15\.0764\.87%\+25\.690\.552\+0\.17842\.38%\+13\.2557\.44%\+21\.760\.497\+0\.153
Table 1:Comparison of LLMs under direct diagnosis, diagnostic\-criteria RAG, andConsultMindacross clinical specialties\. Top\-kkdenotes diagnostic accuracy; MRR is the mean reciprocal rank of the correct diagnosis\.### Data Collection
We collected diagnosis\-labeled cases from two tertiary hospitals and one center for disease control and prevention across different regions\. The dataset comprises psychiatric consultation dialogues, respiratory electronic health records \(EHRs\), and clinical cases from fever clinics\. Primary labels were derived from EHR discharge diagnoses\.
#### Network Construction Corpus\.
We constructed a separate DSBN for each clinical setting from 14,541 cases: 4,102 respiratory, 4,644 psychiatric, and 5,795 fever\-clinic cases\. The psychiatric cases pair dialogue and record narratives, allowing cross\-format consistency evaluation of AutoDisym diagnostic posteriors\.
#### Evaluation Set\.
The disjoint evaluation set contains 990 cases spanning 78 disorders: 287 respiratory cases from 27 disorders, 312 psychiatric cases from 27 disorders, and 391 fever\-clinic cases from 24 disorders\.
#### Virtual Standard Patient Setup\.
We construct each clinical narrative\-based virtual standardized patient \(VSP\) from a structured clinical record\. At each turn, the VSP generates a patient response and cites its supporting record span\. An independent auditor checks record consistency, evidence support, and unsupported clinical claims\. For undocumented information, the VSP must express uncertainty rather than infer symptom absence\. Responses that fail the audit are revised using its feedback until accepted\. Details are provided in Appendix\. Separate GPT\-5\.5 instances serve as the patient and auditor at temperature 0\.
#### Ethics, Privacy, and Availability\.
Data processing followed an approved ethics protocol and theDeclaration of Helsinki\. Evaluation records were institutionally de\-identified and locally rewritten with Qwen3\-32B to protect privacy\. Clinicians verified that diagnostically relevant information was preserved\. Further data and reproducibility details appear in the appendix\. Institution names and ethics protocol identifiers are omitted for double\-blind review\. Upon acceptance, we will disclose them and release the evaluation set and learned DSBN parameters, while the source records used to construct the DSBN will remain restricted\.
#### LLM Baselines\.
We evaluate 10 LLMs in 3 groups: \(1\) Small LLMs: Qwen2\.5\-7B, Qwen3\-8B and 32B\([Team 2025](https://arxiv.org/html/2609.30796#bib.bib47)\), Llama\-3\.1\-8B\([Grattafiori et al\. 2024](https://arxiv.org/html/2609.30796#bib.bib40)\); \(2\) Large LLMs: GPT\-5\.6\-Sol\([OpenAI 2026](https://arxiv.org/html/2609.30796#bib.bib41)\), Gemini\-3\.1\-Pro\([Google DeepMind 2026](https://arxiv.org/html/2609.30796#bib.bib43)\), DeepSeek\-V4\-Pro\([DeepSeek\-AI 2026](https://arxiv.org/html/2609.30796#bib.bib44)\), Claude\-Sonnet\-5\([Anthropic 2026](https://arxiv.org/html/2609.30796#bib.bib42)\); and \(3\) Medical LLMs: ClinicalGPT\-R1\([Lan et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib45)\), HuatuoGPT\-O1\-7B\([Chen et al\. 2024b](https://arxiv.org/html/2609.30796#bib.bib46)\)\. Model configurations and inference details see Appendix
### Experiment Results
Table 2:Projected VSP reliability with and without auditing\.#### RQ1: Are Virtual Standardized Patients Reliable?
We assess VSP reliability using three patient and auditor backbones on 60 cases with seven LLMs, yielding 1,260 dialogues\. Fourteen crowdworkers double\-annotated the sampled dialogues, with agreement exceedingκ=0\.85\\kappa=0\.85\. As shown in Table[2](https://arxiv.org/html/2609.30796#Sx5.T2), audited VSPs achieve over 88\.9% coverage and 93\.2% precision, with hallucination below 3\.1% across all backbone sizes\. Auditing improves coverage by up to 16\.51 points and reduces hallucination by up to 4\.72 points\. These results demonstrate the reliability of VSPs across model capacities and the central role of the auditor\.
#### RQ2: DoesConsultMindImprove Diagnostic Accuracy across LLMs and Specialties?
We compare seven LLMs using direct prompting, RAG\-enhanced prompting, andConsultMindacross psychiatry, respiratory medicine, and fever clinics\. The RAG baseline follows a Naive RAG setup with 256\-token chunks and Top\-5 retrieval from the clinical materials used for DSBN construction, including diagnostic criteria, guidelines, and expert consensus \(details see Appendix\)\. As shown in Table[1](https://arxiv.org/html/2609.30796#Sx5.T1), RAG improves over direct prompting, whileConsultMindfurther improves all 21 model–specialty combinations\. Compared with direct prompting, the gains reach 22\.15 percentage points in Top\-1 accuracy, 37\.89 points in Top\-3 accuracy, and 0\.265 in MRR\.Claude\-Sonnet\-5performs best in psychiatry and is closely matched byGPT\-5\.6\-Solin respiratory medicine\.GPT\-5\.6\-Solperforms best in the fever clinic, reaching 44\.26% Top\-1 accuracy, 60\.18% Top\-3 accuracy, and 0\.526 MRR\. These results show thatConsultMindconsistently outperforms both direct prompting and RAG\-enhanced diagnosis across LLMs and specialties\.
Figure 4:Comparison ofConsultMindwith medical LLMs\.
#### RQ3: How Does ConsultMind Compare with Medical LLMs?
We compare ConsultMind with medically specialized LLMs built on the same Qwen2\.5\-7B base model\. As shown in Figure[4](https://arxiv.org/html/2609.30796#Sx5.F4), ConsultMind outperforms both medical LLMs across all diagnostic metrics\. It achieves 40\.42% Top\-1 and 66\.35% Top\-3 accuracy, surpassing Clinical\-R1 by 7\.56 and 12\.06 points, respectively\. ConsultMind also achieves the highest MRR while requiring fewer dialogue turns than either medical LLM\. These results show that ConsultMind improves the shared base model beyond specialized SFT\.
#### RQ4: Does ConsultMind Generalize across Datasets?
We evaluate ConsultMind on three external datasets: MentalHospital\([Yang et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib48)\), MedSP1000\([Liang et al\. 2026](https://arxiv.org/html/2609.30796#bib.bib49)\), and AIHospital\([Fan et al\. 2025](https://arxiv.org/html/2609.30796#bib.bib50)\)\. The evaluation covers 310 psychiatric virtual patients from MentalHospital, 40 psychiatric and 48 respiratory cases from MedSP1000, and 14 fever\-clinic cases from AIHospital\. As shown in Table[3](https://arxiv.org/html/2609.30796#Sx5.T3), ConsultMind improves all nine model and dataset combinations, with gains of up to 16\.78 points in Top\-1 accuracy, 24\.51 points in Top\-3 accuracy, and 0\.187 in MRR\. These consistent gains demonstrate robust generalization across external datasets and clinical settings\.
ModelTop\-1↑\\uparrowTop\-3↑\\uparrowMRR↑\\uparrowMentalHospitalPsychiatryLlama\-3\.1\-8B16\.77%34\.19%0\.249w/ConsultMind33\.55%\+16\.7851\.61%\+17\.420\.430\+0\.181Qwen3\-32B27\.74%35\.81%0\.325w/ConsultMind40\.65%\+12\.9160\.32%\+24\.510\.512\+0\.187GPT\-5\.6\-Sol29\.03%34\.19%0\.329w/ConsultMind39\.68%\+10\.6556\.13%\+21\.940\.487\+0\.158MedSP1000Psychiatry, Respiratory medicineLlama\-3\.1\-8B14\.77%32\.95%0\.230w/ConsultMind27\.27%\+12\.5047\.73%\+14\.780\.375\+0\.145Qwen3\-32B25\.00%34\.09%0\.301w/ConsultMind35\.23%\+10\.2353\.41%\+19\.320\.450\+0\.149GPT\-5\.6\-Sol27\.27%32\.95%0\.312w/ConsultMind37\.50%\+10\.2351\.14%\+18\.190\.452\+0\.140AIHospitalFever ClinicLlama\-3\.1\-8B21\.43%35\.71%0\.280w/ConsultMind35\.71%\+14\.2850\.00%\+14\.290\.433\+0\.153Qwen3\-32B28\.57%35\.71%0\.327w/ConsultMind42\.86%\+14\.2957\.14%\+21\.430\.510\+0\.183GPT\-5\.6\-Sol28\.57%35\.71%0\.330w/ConsultMind42\.86%\+14\.2957\.14%\+21\.430\.507\+0\.177
Table 3:Diagnostic performance of LLMs with and withoutConsultMindacross datasets\.
#### RQ5: Does ConsultMind Improve the Clinical Soundness of Diagnostic Explanations?
Figure 5:Physician ratings of diagnostic explanation quality\.Figure 6:Consult comparison with/withoutConsultMind\.We compare paired explanations from three representative LLMs on 50 correctly diagnosed cases per model, using their base and ConsultMind\-enhanced versions\. Each explanation follows a three\-part template: Ranking Explanation \(RE\), Differential Diagnosis \(DD\), and Diagnosis Rationale \(DR\)\. Specialists from the corresponding departments rate the anonymized and shuffled explanations on a five\-point scale \(details see Appendix\)\. As shown in Figure[5](https://arxiv.org/html/2609.30796#Sx5.F5), ConsultMind improves models, with the largest overall gain for Llama\-3\.1\-8B from 2\.30 to 3\.74\. Its RE score shows the largest component gain, increasing from 2\.04 to 3\.58\. These results demonstrate that ConsultMind produces more clinically sound diagnostic explanations\.
#### RQ6: How Reliable Are AutoDisym\-Generated Symptom Schemas?
We compare AutoDisym\-generated schemas with expert gold schemas for six disorders across three departments:Influenza,Dengue fever,Emphysema,Acute bronchiolitis,Bipolar type II disorder, andRecurrent depressive disorder\. Two experts independently annotate canonical symptoms, normalized manifestations, and their pairs from the same diagnostic materials, while a third resolves disagreements \(details see Appendix\)\. Using the same backbones, we compare Direct LLM, RAG\-enhanced LLM, and AutoDisym\. Direct LLM uses parametric knowledge, while the other methods share a Naive RAG setting with 256\-token chunks and Top\-5 retrieval\. After synonym normalization, one\-to\-one matches are correct only when both the symptom and manifestation match the gold schema above 0\.85 similarity\. Table[4](https://arxiv.org/html/2609.30796#Sx5.T4)reports disorder\-level macro precision, recall, and F1\. AutoDisym performs best with both backbones, reaching 81\.37 symptom F1 and 72\.19 manifestation F1 with GPT\-5\.6\-Sol\. It also reduces the manifestation\-F1 gap between backbones from 9\.68 to 5\.96, demonstrating more reliable schemas and less dependence on backbone capability\.
Table 4:Schema agreement across configurations\.
### Case Study
Figure[6](https://arxiv.org/html/2609.30796#Sx5.F6)compares GPT\-5\.6\-Sol with and without ConsultMind on a patient reporting increased energy, reduced sleep, and persecutory beliefs\. Without ConsultMind, the model follows the medication\-related narrative and predicts bipolar I disorder\. With ConsultMind, it probes grandiosity, fixed false beliefs, and suspiciousness, yielding stronger evidence for schizoaffective disorder\. This case shows that ConsultMind reduces narrative drift by directing inquiry toward discriminative symptoms in overlapping presentations\.
## Conclusion
We introducedAutoDisym, an automated pipeline for constructing Disorder–Symptom Bayesian Networks \(DSBNs\), andConsultMind, an uncertainty\-aware framework for sequential diagnostic consultation\. AutoDisym integrates diagnostic knowledge with diagnosis\-labeled clinical narratives, while ConsultMind updates disorder posteriors and uses uncertainty to guide inquiry and diagnosis\. Evaluations across three clinical settings and three public datasets show that AutoDisym constructs high\-quality DSBNs and ConsultMind consistently improves diagnostic performance and explanation soundness\. ConsultMind increases Top\-1 and Top\-3 accuracy by up to 22\.15 and 37\.89 percentage points, respectively\. Physician evaluation further confirms improvements in ranking explanation, differential diagnosis, and diagnosis rationale across LLMs of different scales\. This work offers a promising approach to automatic diagnostic consultation\.
## References
- Alberet al\.\(2025\)D\. A\. Alber, Z\. Yang, A\. Alyakin, E\. Yang, S\. Rai, A\. A\. Valliani, J\. Zhang, G\. R\. Rosenbaum, A\. K\. Amend\-Thomas, D\. B\. Kurland,et al\.Medical large language models are vulnerable to data\-poisoning attacks\.Nature Medicine31\(2\),pp\. 618–626\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Andreadiset al\.\(2024\)K\. Andreadis, D\. R\. Newman, C\. Twan, A\. Shunk, D\. M\. Mann, and E\. R\. StevensMixed methods assessment of the influence of demographics on medical advice of chatgpt\.Journal of the American Medical Informatics Association31\(9\),pp\. 2002–2009\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Anthropic \(2026\)AnthropicClaude Sonnet 5 System Card\.Note:AnthropicSystem cardExternal Links:[Link](https://www.anthropic.com/claude-sonnet-5-system-card)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Ben\-Shabatet al\.\(2022\)N\. Ben\-Shabat, G\. Sharvit, B\. Meimis, D\. B\. Joya, A\. Sloma, D\. Kiderman, A\. Shabat, A\. M\. Tsur, A\. Watad, and H\. AmitalAssessing data gathering of chatbot based symptom checkers\-a clinical vignettes study\.International Journal of Medical Informatics168,pp\. 104897\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Chenet al\.\(2024a\)J\. Chen, S\. Xiao, P\. Zhang, K\. Luo, D\. Lian, and Z\. LiuBGE m3\-embedding: multi\-lingual, multi\-functionality, multi\-granularity text embeddings through self\-knowledge distillation\.External Links:2402\.03216Cited by:[Stage 3: Retrieval\-Augmented Schema Verification\.](https://arxiv.org/html/2609.30796#Sx3.SS0.SSS0.Px4.p1.1)\.
- Chenet al\.\(2024b\)J\. Chen, Z\. Cai, K\. Ji, X\. Wang, W\. Liu, R\. Wang, J\. Hou, and B\. WangHuatuoGPT\-o1, towards medical complex reasoning with llms\.External Links:2412\.18925,[Link](https://arxiv.org/abs/2412.18925)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- DeepSeek\-AI \(2026\)DeepSeek\-AIDeepSeek\-V4: Towards Highly Efficient Million\-Token Context Intelligence\.Note:Technical reportExternal Links:[Link](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Fanet al\.\(2025\)Z\. Fan, L\. Wei, J\. Tang, W\. Chen, W\. Siyuan, Z\. Wei, and F\. HuangAi hospital: benchmarking large language models in a multi\-agent medical interaction simulator\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 10183–10213\.Cited by:[RQ4: Does ConsultMind Generalize across Datasets?](https://arxiv.org/html/2609.30796#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Gaberet al\.\(2025\)F\. Gaber, M\. Shaik, F\. Allega, A\. J\. Bilecz, F\. Busch, K\. Goon, V\. Franke, and A\. AkalinEvaluating large language model workflows in clinical decision support for triage and referral and diagnosis\.npj Digital Medicine8\(1\),pp\. 263\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Gaoet al\.\(2025\)S\. Gao, K\. Yu, Y\. Yang, S\. Yu, C\. Shi, X\. Wang, N\. Tang, and H\. ZhuLarge language model powered knowledge graph construction for mental health exploration\.Nature Communications16\(1\),pp\. 7526\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Gonget al\.\(2025\)L\. Gong, A\. Wang, Y\. Lai, W\. Ma, and Y\. LiuThe dialogue that heals: a comprehensive evaluation of doctor agents’ inquiry capability\.External Links:2509\.24958,[Link](https://arxiv.org/abs/2509.24958)Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p3.1)\.
- Google DeepMind \(2026\)Google DeepMindGemini 3\.1 Pro Model Card\.Note:Google DeepMindModel cardExternal Links:[Link](https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Houet al\.\(2023\)Z\. Hou, Y\. Cen, Z\. Liu, D\. Wu, B\. Wang, X\. Li, L\. Hong, and J\. TangMtdiag: an effective multi\-task framework for automatic diagnosis\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.37,pp\. 14241–14248\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Jiaet al\.\(2025\)M\. Jia, J\. Duan, Y\. Song, and J\. WangMedikal: integrating knowledge graphs as assistants of llms for enhanced clinical diagnosis on emrs\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 9278–9298\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Jianget al\.\(2023\)P\. Jiang, C\. Xiao, A\. R\. Cross, and J\. SunGraphcare: enhancing healthcare predictions with personalized knowledge graphs\.InThe Twelfth International Conference on Learning Representations,Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Johriet al\.\(2025\)S\. Johri, J\. Jeong, B\. A\. Tran, D\. I\. Schlessinger, S\. Wongvibulsin, L\. A\. Barnes, H\. Zhou, Z\. R\. Cai, E\. M\. Van Allen, D\. Kim,et al\.An evaluation framework for clinical use of large language models in patient interaction tasks\.Nature medicine31\(1\),pp\. 77–86\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p3.1)\.
- Johriet al\.\(2024\)S\. Johri, J\. Jeong, B\. A\. Tran, D\. I\. Schlessinger, S\. Wongvibulsin, Z\. R\. Cai, R\. Daneshjou, and P\. RajpurkarCRAFT\-md: a conversational evaluation framework for comprehensive assessment of clinical llms\.InAAAI 2024 Spring Symposium on Clinical Foundation Models,Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p3.1)\.
- Kresevicet al\.\(2024\)S\. Kresevic, M\. Giuffrè, M\. Ajcevic, A\. Accardo, L\. S\. Crocè, and D\. L\. ShungOptimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation\-based framework\.NPJ digital medicine7\(1\),pp\. 102\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Lanet al\.\(2025\)W\. Lan, W\. Wang, C\. Ji, G\. Yang, Y\. Zhang, X\. Liu, S\. Wu, and G\. WangClinicalGPT\-r1: pushing reasoning capability of generalist disease diagnosis with large language model\.External Links:2504\.09421,[Link](https://arxiv.org/abs/2504.09421)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Liet al\.\(2024\)S\. S\. Li, V\. Balachandran, S\. Feng, J\. S\. Ilgen, E\. Pierson, P\. W\. Koh, and Y\. TsvetkovMediq: question\-asking llms and a benchmark for reliable interactive clinical reasoning\.Advances in Neural Information Processing Systems37,pp\. 28858–28888\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Liet al\.\(2026\)W\. Li, Y\. Zhang, C\. Wang, Y\. Li, X\. He, A\. L\. Wang, M\. Xu, F\. Zhang, H\. Sun, K\. Wang,et al\.CARE: a clinical agentic reasoning engine to enhance real\-world diagnostic accuracy via structured medical reasoning\.Expert Systems with Applications,pp\. 131476\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Lianget al\.\(2026\)C\. Liang, P\. Qiu, Y\. Zhang, Y\. Wang, C\. Wu, and W\. XieEvaluating large language models in dynamic clinical decision\-making with standardized patient cases\.External Links:2606\.05112,[Link](https://arxiv.org/abs/2606.05112)Cited by:[RQ4: Does ConsultMind Generalize across Datasets?](https://arxiv.org/html/2609.30796#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Liuet al\.\(2026\)Y\. Liu, S\. Yu, H\. Jin, J\. Wen, A\. Qian, T\. Lee, M\. Ramsis, G\. W\. Choi, L\. Qin, X\. Liu,et al\.A multi\-agent framework combining large language models with medical flowcharts for self\-triage\.Nature Health,pp\. 1–10\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- OpenAI \(2026\)OpenAIGPT\-5\.6 System Card\.Note:OpenAI Deployment Safety HubSystem cardExternal Links:[Link](https://deploymentsafety.openai.com/gpt-5-6)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Qiaoet al\.\(2026\)C\. Qiao, J\. Huang, D\. Zhao, Z\. Liu, Y\. Shen, B\. Cheng, W\. Lin, and K\. WuMedConsultBench: a full\-cycle, fine\-grained, process\-aware benchmark for medical consultation agents\.External Links:2601\.12661,[Link](https://arxiv.org/abs/2601.12661)Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p3.1)\.
- Renet al\.\(2025\)W\. Ren, T\. Zhao, L\. Wang, T\. Wang, and V\. G\. HonavarDiaLLMs: ehr\-enhanced clinical conversational system for clinical test recommendation and diagnosis prediction\.InFindings of the Association for Computational Linguistics: ACL 2025,pp\. 25622–25635\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Saabet al\.\(2026\)K\. Saab, C\. Park, T\. Strother, J\. Freyberg, D\. G\. Barrett, Y\. Cheng, W\. Weng, D\. Stutz, N\. Tomasev, A\. Palepu,et al\.Advancing conversational diagnostic ai with multimodal reasoning\.Nature Medicine,pp\. 1–11\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Sanghviet al\.\(2026\)A\. Sanghvi, N\. Akash, R\. Imam, A\. Sharma, and M\. JainMeDxAgent: multi\-agent consultation for interactive medical diagnosis\.External Links:2606\.03416,[Link](https://arxiv.org/abs/2606.03416)Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Savageet al\.\(2024\)T\. Savage, A\. Nayak, R\. Gallo, E\. Rangan, and J\. H\. ChenDiagnostic reasoning prompts reveal the potential for large language model interpretability in medicine\.NPJ Digital Medicine7\(1\),pp\. 20\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Shiet al\.\(2023\)X\. Shi, Z\. Liu, C\. Wang, H\. Leng, K\. Xue, X\. Zhang, and S\. ZhangMidMed: towards mixed\-type dialogues for medical consultation\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8145–8157\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1)\.
- Singhalet al\.\(2023\)K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl,et al\.Large language models encode clinical knowledge\.Nature620\(7972\),pp\. 172–180\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Sittiget al\.\(2024\)D\. F\. Sittig, A\. Boxwala, A\. Wright, C\. Zott, N\. A\. Gauthreaux, J\. Swiger, E\. A\. Lomotan, and P\. DullabhPatient\-centered clinical decision support challenges and opportunities identified from workflow execution models\.Journal of the American Medical Informatics Association31\(8\),pp\. 1682–1692\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Sohnet al\.\(2026\)J\. Sohn, B\. Ha, S\. Park, J\. Kim, E\. Lee, H\. Oh, S\. Lee, and E\. KimSystematic review and meta analysis of chatbots in the management of depressive and anxiety symptoms\.NPJ Digital Medicine\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Songet al\.\(2025\)J\. Song, Z\. Xu, M\. He, J\. Feng, and B\. ShenGraph retrieval augmented large language models for facial phenotype associated rare genetic disease\.NPJ digital medicine8\(1\),pp\. 543\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Sunet al\.\(2026\)X\. Sun, Y\. Yang, X\. Jiang, Y\. Tian, J\. Zhu, J\. Zhong, Q\. Lei, J\. Huang, H\. Zeng, X\. Zhou,et al\.MentalSeek\-dx: towards progressive hypothetico\-deductive reasoning for real\-world psychiatric diagnosis\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 26600–26636\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p3.1)\.
- Tanget al\.\(2016\)K\. Tang, H\. Kao, C\. Chou, and E\. Y\. ChangInquire and diagnose: neural symptom checking ensemble using deep reinforcement learning\.InNIPS workshop on deep reinforcement learning,Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1)\.
- Team \(2025\)Q\. TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[LLM Baselines\.](https://arxiv.org/html/2609.30796#Sx5.SSx1.SSS0.Px5.p1.1)\.
- Tuet al\.\(2024\)T\. Tu, A\. Palepu, M\. Schaekermann, K\. Saab, J\. Freyberg, R\. Tanno, A\. Wang, B\. Li, M\. Amin, N\. Tomasev, S\. Azizi, K\. Singhal, Y\. Cheng, L\. Hou, A\. Webson, K\. Kulkarni, S\. S\. Mahdavi, C\. Semturs, J\. Gottweis, J\. Barral, K\. Chou, G\. S\. Corrado, Y\. Matias, A\. Karthikesalingam, and V\. NatarajanTowards conversational diagnostic ai\.External Links:2401\.05654,[Link](https://arxiv.org/abs/2401.05654)Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Tuet al\.\(2025\)T\. Tu, M\. Schaekermann, A\. Palepu, K\. Saab, J\. Freyberg, R\. Tanno, A\. Wang, B\. Li, M\. Amin, Y\. Cheng,et al\.Towards conversational diagnostic artificial intelligence\.Nature642\(8067\),pp\. 442–450\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Wanget al\.\(2024\)L\. Wang, X\. Chen, X\. Deng, H\. Wen, M\. You, W\. Liu, Q\. Li, and J\. LiPrompt engineering in consistency and reliability with the evidence\-based guideline for llms\.NPJ digital medicine7\(1\),pp\. 41\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Weiet al\.\(2018\)Z\. Wei, Q\. Liu, B\. Peng, H\. Tou, T\. Chen, X\. Huang, K\. Wong, and X\. DaiTask\-oriented dialogue system for automatic diagnosis\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 201–207\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Werthaimet al\.\(2026\)M\. Werthaim, M\. Kimhi, A\. Apartsin, and Y\. ApersteinA benchmark for evaluating diagnostic questioning efficiency of llms in patient conversations\.Scientific Reports\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Xiaet al\.\(2020\)Y\. Xia, J\. Zhou, Z\. Shi, C\. Lu, and H\. HuangGenerative adversarial regularized mutual information policy gradient framework for automatic diagnosis\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 1062–1069\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p3.1)\.
- Xionget al\.\(2024\)G\. Xiong, Q\. Jin, Z\. Lu, and A\. ZhangBenchmarking retrieval\-augmented generation for medicine\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 6233–6251\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.
- Xuet al\.\(2023\)K\. Xu, W\. Hou, Y\. Cheng, J\. Wang, and W\. LiMedical dialogue generation via dual flow modeling\.InFindings of the Association for Computational Linguistics: ACL 2023,pp\. 6771–6784\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Yanget al\.\(2026\)Y\. Yang, X\. Sun, Y\. Zou, Z\. Wu, Y\. Chen, J\. Zhong, H\. Zeng, J\. Huang, and K\. WeiMentalHospital: a virtual environment for evaluating psychiatric clinical encounters\.External Links:2607\.08257,[Link](https://arxiv.org/abs/2607.08257)Cited by:[RQ4: Does ConsultMind Generalize across Datasets?](https://arxiv.org/html/2609.30796#Sx5.SSx2.SSS0.Px4.p1.1)\.
- Youet al\.\(2023\)Y\. You, C\. Tsai, Y\. Li, F\. Ma, C\. Heron, and X\. GuiBeyond self\-diagnosis: how a chatbot\-based symptom checker should respond\.ACM Transactions on Computer\-Human Interaction30\(4\),pp\. 1–44\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p2.1)\.
- Yuan and Yu \(2024\)H\. Yuan and S\. YuEfficient symptom inquiring and diagnosis via adaptive alignment of reinforcement learning and classification\.Artificial Intelligence in Medicine148,pp\. 102748\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1),[Introduction](https://arxiv.org/html/2609.30796#Sx1.p2.1)\.
- Zenget al\.\(2020\)G\. Zeng, W\. Yang, Z\. Ju, Y\. Yang, S\. Wang, R\. Zhang, M\. Zhou, J\. Zeng, X\. Dong, R\. Zhang,et al\.MedDialog: large\-scale medical dialogue datasets\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 9241–9250\.Cited by:[Introduction](https://arxiv.org/html/2609.30796#Sx1.p1.1)\.
- Zhouet al\.\(2026\)H\. Zhou, F\. Liu, J\. Wu, W\. Zhang, G\. Huang, L\. Clifton, D\. Eyre, H\. Luo, F\. Liu, K\. Branson,et al\.A collaborative large language model for drug analysis\.Nature Biomedical Engineering10\(5\),pp\. 870–881\.Cited by:[Related Work](https://arxiv.org/html/2609.30796#Sx2.p1.1)\.Similar Articles
WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis
WiseMind is a knowledge-guided multi-agent framework that uses LLMs for psychiatric diagnosis by combining a "Reasonable Mind" agent for evidence-based logic with an "Emotional Mind" agent for empathetic communication, achieving 85.6% diagnostic accuracy on simulated and real patient interactions. The framework leverages DSM-5 structured knowledge graphs to reduce hallucinations and outperforms single-agent baselines by 15-54 percentage points while maintaining clinical soundness and psychological support.
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
MedDDC-Eval introduces a diagnosis-decoupled evaluation testbed for multi-turn medical consultation agents, isolating the policy-elicited conversation history from diagnosis generation to enable cleaner measurement of evidence acquisition and diagnostic usefulness.
COTCAgent: Preventive Consultation via Probabilistic Chain-of-Thought Completion
COTCAgent is a hierarchical reasoning framework for longitudinal electronic health records that uses a probabilistic chain-of-thought completion approach, achieving 90.47% Top-1 accuracy on a self-built dataset and outperforming existing medical agents.
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
This paper presents an evaluation of multi-turn multimodal diagnostic reasoning using challenging real-world clinical cases, aiming to assess AI models' ability to handle complex medical scenarios.
Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning
This paper examines how human interventions at fault points in multi-agent medical systems affect diagnostic accuracy, showing improvements with correct interventions and degradation with incorrect ones.