From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators

arXiv cs.CL Papers

Summary

This paper introduces DischargeBench, a persona-grounded simulation to evaluate LLMs as discharge educators, measuring patient understanding through multi-turn dialogues with virtual patients.

arXiv:2609.20827v1 Announce Type: new Abstract: Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.
Original Article
View Cached Full Text

Cached at: 09/21/26, 08:56 AM

# From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators
Source: [https://arxiv.org/html/2609.20827](https://arxiv.org/html/2609.20827)
Won Seok Jang1,Zonghai Yao1,Hong Yu1, 1University of Massachusetts Lowell 2University of Massachusetts Amherst Correspondence:[WonSeok\_Jang@student\.uml\.edu](mailto:[email protected])

###### Abstract

Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient’s literacy, recall, and personality\. Existing LLM evaluations target static or artifact\-generation tasks and do not measure patient understanding under open\-ended dialogue\. We introduceDischargeBench, a persona\-grounded simulation in which a candidate LLM educator conducts a multi\-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal\. We curateMIMIC\-IV\-Ext\-DischargeBench, 477 cases over 24 ICD chapters with persona axes \(personality, education level, health literacy, past\-medical\-history recall\) for stratified analysis\. Each simulation is scored on four axes — Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency — by an LLM\-as\-a\-Judge aligned against physician annotations\. Across closed\- and open\-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source\-answer agreement\. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone\.

From Discharge Notes to Patient Understanding: Persona\-Grounded, Open\-Ended Simulation of LLMs as Discharge Educators

Won Seok Jang1, Zonghai Yao1, Hong Yu1,1University of Massachusetts Lowell2University of Massachusetts AmherstCorrespondence:[WonSeok\_Jang@student\.uml\.edu](mailto:[email protected])

## 1Introduction

![Refer to caption](https://arxiv.org/html/2609.20827v1/x1.png)Figure 1:Overview of DischargeBench\. Two stages: \(1\) An Educator \(system under test\) conducts a multi\-turn session with a Virtual Patient, supervised by an Education Monitor Agent that regulates patient realism and signals session completion \(477 cases\)\. \(2\) An LLM\-as\-a\-Judge scores each dialogue along Conversation Quality, Topic Checklist Score \(TCS\), post\-education Comprehension Score, and Factual Consistency\.Patients leave hospitals with prescriptions, follow\-up appointments, and ideally, education about their discharge instructions\. Patient education represents a core dimension of the clinician–patient relationship, providing the knowledge and behavioral guidance that supports safe recovery[M\. V\. Williams, C\. White\-Williams, and J\. Li \(2026\)](https://arxiv.org/html/2609.20827#bib.bib67);[D\. C\. Gonçalves\-Bradley, N\. A\. Lannin, L\. Clemson, I\. D\. Cameron, and S\. Shepperd \(2022\)](https://arxiv.org/html/2609.20827#bib.bib45);[37](https://arxiv.org/html/2609.20827#bib.bib64)\. At the moment of discharge in particular, the quality of this education has direct and measurable consequences: better\-informed patients are less likely to be readmittedOhet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib33)\); Beckeret al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib34)\); Yumena \([2025](https://arxiv.org/html/2609.20827#bib.bib35)\); Rasmussenet al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib36)\), experience fewer postoperative complications and recover more fullyGillespieet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib57)\); Kanget al\.\([2022](https://arxiv.org/html/2609.20827#bib.bib38)\), and report higher satisfaction with their careDeSaiet al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib11)\); Zandifaret al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib58)\)\. Yet in practice, discharge education is routinely neglected: clinicians acknowledge its importance but often lack the time to deliver itTrivediet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib18)\)\.

Researchers are exploring ways to apply Large Language Models \(LLMs\) to clinical communication: generating discharge education materialsWillet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib41)\), simplifying discharge notes into lay languageChuaet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib50)\); Liet al\.\([2026](https://arxiv.org/html/2609.20827#bib.bib49)\); Hainset al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib60)\), training task\-specific discharge chatbotsJanget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib10)\), and evaluating LLMs as discharge educators in controlled simulationsYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\)\. Nevertheless, these settings do not test the full task\. First, education\-material generation, summarization, lay\-language simplification, and clinical QA benchmarksSinghalet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib62)\); Zhanget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib70)\); Jianget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib68)\); Kweonet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib59)\)evaluate artifacts or accuracy on text rather than whether a patient understands\. Second, model\-development studies fix a specific system \(e\.g\., NoteAid\-ChatbotJanget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib10)\)\), making it difficult to compare arbitrary LLMs under matched conditions\. Third, the closest dialogue benchmark, DischargeSimYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\), runs 49 cases through a stage\-structured \(3 personas; MCQ comprehension\) rather than open\-ended evaluation\. None of these jointly measure whether an LLM can teach an open\-ended, heterogeneous patient to understand from the source discharge note\.

To this end, we introduce DischargeBench, an open\-ended, persona\-grounded simulation framework for evaluating AI models as discharge educators\. Each session pairs a candidate Educator with a Virtual Patient grounded in real MIMIC\-IV\-Note discharge informationJohnsonet al\.\([2023a](https://arxiv.org/html/2609.20827#bib.bib7),[b](https://arxiv.org/html/2609.20827#bib.bib37)\), regulated by an Education Monitor Agent, and is scored along four clinically grounded axes — Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency\. We summarize our contributions as follows:

\(1\) We formulate hospital discharge education as aninteractive teaching taskfor LLMs, in which the target outcome is post\-educationpatient understandinggrounded in the source discharge note, rather than text\-quality, simplification, or question\-answering metrics\. \(2\) We curateMIMIC\-IV\-Ext\-DischargeBench, 477 cases derived from MIMIC\-IV and MIMIC\-IV\-Note spanning 24 ICD chapters and annotated with virtual\-patient attribute axes \(personality, education level, health literacy, past\-medical\-history recall\) that enable stratified evaluation across heterogeneous patients\. \(3\) We design a multi\-agent simulation in which a Virtual Patient is regulated by an Education Monitor Agent that intervenes only on the patient side, preserving the integrity of the evaluation signal for the model under test\. \(4\) Benchmarking closed\- and open\-source LLMs, we find that aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source\-answer agreement\.

## 2DischargeBench

DischargeBench consists of two stages: a multi\-agent simulation of discharge conversations between an educator LLM and a Virtual Patient, and an LLM\-as\-judge evaluation along four clinically grounded axes \(Figure[1](https://arxiv.org/html/2609.20827#S1.F1)\)\. We describe the dataset, simulation, evaluation, and failure analysis\.

### 2\.1MIMIC\-IV\-Ext\-DischargeBench

#### 2\.1\.1Dataset Overview

The dataset was derived from the MIMIC\-IV \(v3\.1\) databaseJohnsonet al\.\([2023a](https://arxiv.org/html/2609.20827#bib.bib7)\)and MIMIC\-IV Note \(v2\.2\)Johnsonet al\.\([2023b](https://arxiv.org/html/2609.20827#bib.bib37)\), collectively forming MIMIC\-IV\-Ext\-DischargeBench\. A total of 477 cases were sampled to construct the dataset\. To ensure clinical relevance and dataset quality, a systematic patient selection pipeline was employed, incorporating both ICD\-9 and ICD\-10 diagnosis code filtering and manual auditing\. The resulting dataset encompasses patients spanning 24 distinct ICD chapters of primary diagnosis\. Detailed descriptions of the curation and validation procedures are provided in Appendix[A\.1](https://arxiv.org/html/2609.20827#A1.SS1)\.

#### 2\.1\.2Dataset Access

Because MIMIC\-IV\-Ext\-DischargeBench is derived from MIMIC\-IV, we will release it through the PhysioNet platformGoldbergeret al\.\([2000](https://arxiv.org/html/2609.20827#bib.bib54)\)under the same access regime as the source data\. To use our data, one must hold credentialed PhysioNet access, sign the PhysioNet Credentialed Health Data Use Agreement \(v1\.5\.0\), and complete the CITI Data or Specimens Only Research training\.

### 2\.2Simulating Discharge Education

#### 2\.2\.1Virtual Patient

Among the primary contributions of this work is a simulation framework centered on a Virtual Patient \(VP\)\. The VP is driven by a structured prompt applied to the open\-source Llama\-3\.3\-70B\-Instruct modelGrattafioriet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib65)\), encoding rich behavioral rules that condition the model on persona attributes and constrain its output to in\-character patient utterances\. Each VP persona is parameterized across five axes: \(1\)Medical Profile, integrating demographics, main diagnoses, medications, allergies, medical history, reason for admission, and family history extracted from MIMIC\-IV noteJohnsonet al\.\([2023b](https://arxiv.org/html/2609.20827#bib.bib37)\); \(2\)Education Level\(elementary, high school, or college\), which shapes vocabulary and instruction\-following depth; \(3\)Health Literacy\(low or high\), which governs the patient’s ability to interpret medical information and translate plan into self\-management; \(4\)Personality, drawn from five distinct profiles \(neutral, anxious, distrustful, high\-conscientiousness, and minimiser\) grounded in the Five\-Factor ModelMcCrae and John \([1992](https://arxiv.org/html/2609.20827#bib.bib24)\)and prior work on persona\-driven patient simulationKyunget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib25)\); Redelmeieret al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib26)\); and \(5\)Past Medical History Recall\(poor, partial, or accurate\), which determines how reliably the patient can volunteer prior diagnoses, surgeries, and medications during the encounter, and whether they fill memory gaps with approximations or admit uncertainty\. This trait was also motivated by realistic simulation studies ofKyunget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib25)\); Zhonget al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib43)\)\. A description of the VP trait descriptions and the prompt is provided in Appendix[A\.2](https://arxiv.org/html/2609.20827#A1.SS2)\.

#### 2\.2\.2Education Monitor Agent

Motivated bySchmidgallet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib31)\); Yuet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib32)\), which demonstrate the strength of multi\-agent orchestration for stable and realistic clinical simulation, we developed an Education Monitor Agent \(EMA\) that operates as a silent quality controller running in parallel with the simulation\. Qwen3\.5\-9BYanget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib5)\), was used as the EMA backend\. EMA serves two functions: \(1\)turn\-level oversightof the ongoing dialogue, and \(2\)session\-termination control\.

##### Turn\-level oversight

On every conversational turn, the EMA receives the agent’s system prompt alongside its output and returns a structured verdict \(PASS / WARN / FAIL\) tagging any of seven failure categories, together with a severity \(minor / moderate / severe\) and a recommended action \(Appendix[A\.3](https://arxiv.org/html/2609.20827#A1.SS3)\)\. This oversight design follows prior work in which a third\-party LLM monitors and critiques agent behaviorWuet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib52)\); Tuet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib16)\); Vedadiet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib53)\); Schmidgallet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib31)\)\. The VP is held to an active intervention policy: WARN and FAIL verdicts at minor or moderate severity trigger a soft correction appended to the VP’s subsequent system prompt, severe failures \(e\.g\. unrecoverable repetition loops or sustained character drift\) cause the offending turn to be discarded and regenerated, and the session is hard\-stopped only when the per\-turn retry budget \(max\_retries=2\\text\{max\\\_retries\}\{=\}2by default\) is exhausted on a persistent severe failure\. The Educator is held to a passive observation policy to preserve an authentic evaluation signal\.

##### Session\-termination control

The EMA is also responsible for ending each session, whether by intervention or by natural conclusion\. A session is marked complete only when the educator has covered all required discharge domains and both parties have exchanged explicit closing acknowledgments, with a minimum\-turn guard \(≥10\\geq 10\) preventing premature termination after only an opening exchange\. Detailed specifications of the EMA prompt, intervention policy, and session\-completion logic are provided in Appendix[A\.3](https://arxiv.org/html/2609.20827#A1.SS3)\.

### 2\.3Evaluation Methodology

#### 2\.3\.1Automated Evaluation

##### LLM\-as\-a\-Judge

We employed the LLM\-as\-a\-Judge evaluation framework, following approaches established in prior conversational diagnostic agent developmentSaabet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib47)\); Tuet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib16)\)and benchmark studiesYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\)\. As the judge model, we utilized Gemma\-4\-31B\-it[42](https://arxiv.org/html/2609.20827#bib.bib27)\.

##### Conversation Quality

This metric evaluates the overall quality of the educator\-patient dialogue\. The LLM Judge scores each conversation along four criteria on a 1\-to\-5 Likert scale, with an additional automatic readability measure:

Naturalness \(NA\): whether the conversation flows naturally, without repetition, and exhibits a clear opening and closing\.Responsiveness \(RE\): how effectively the educator addresses the patient’s concerns, questions, and emotions\.Clarity \(CL\): whether the educator’s responses are clear and easy to understand, avoiding unexplained medical jargon and addressing one topic per turn\.Clinical Relevance \(CR\): whether the educator’s responses are clinically valid and aligned with established medical practice\.Readability \(RD\)111This score does not involve the LLM Judge; we use a Python library to compute the Flesch\-Kincaid grade level\.: the Flesch\-Kincaid grade levelKincaidet al\.\([1975](https://arxiv.org/html/2609.20827#bib.bib66)\)of the educator’s turns\.

These criteria are adapted from prior work on medical\-domain conversational agentsGudipati \([2025](https://arxiv.org/html/2609.20827#bib.bib15)\); Tuet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib16)\); Wanget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib17)\)\. The LLM Judge provides a written justification for each Likert\-scale score \(prompt in Appendix[8](https://arxiv.org/html/2609.20827#A1.F8)\)\.

##### Topic Checklist Score

This metric measures how thoroughly the dialogue covers the discharge topics\. FollowingDeSaiet al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib11)\); Trivediet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib18)\); Janget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib10)\), we curated a set of yes/no questions covering six clinical areas: Discharge Diagnosis, New Medication, Treatment During Stay, Post\-Discharge Treatment, Indications for Return to Hospital, and Follow\-Up Appointment \(Appendix Table[3](https://arxiv.org/html/2609.20827#A1.T3)\)\. Each parent question carries a weight of11and each sub\-question a weight of0\.50\.5; questions deemed not applicable to a given patient note are excluded from both numerator and denominator\. LetQQdenote the full Topic Checklist question set\. The Topic Checklist Score \(TCS\) for sessioniiis the total weight of correctly addressed questions normalised by the total weight of applicable questionsQi⊆QQ\_\{i\}\\subseteq Q, and the reported TCS is the mean over all successful sessions \(nn\) \(Eq\.[1](https://arxiv.org/html/2609.20827#S2.E1)\)\.QQand the per\-case applicable subsetQiQ\_\{i\}are constructed during dataset curation \(Appendix[A\.1](https://arxiv.org/html/2609.20827#A1.SS1)\)\.

T​C​S=1n​∑i=1n∑q∈Qisq⋅𝕀​\[y^i,q=1\]∑q∈QisqTCS=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\frac\{\\sum\_\{q\\in Q\_\{i\}\}s\_\{q\}\\cdot\\mathbb\{I\}\[\\hat\{y\}\_\{i,q\}=1\]\}\{\\sum\_\{q\\in Q\_\{i\}\}s\_\{q\}\}\(1\)

##### Comprehension Score

This evaluates the VP’s comprehension using six clinically motivated open\-ended questions adapted fromTrivediet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib18)\); DeSaiet al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib11)\); Janget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib10)\)\(Appendix Table[4](https://arxiv.org/html/2609.20827#A1.T4)\) using the same clinical areas from TCS\. For each discharge note, a reference answer is first extracted via the LLM Judge\. The VP is then queried per question — post\-education, with the full conversation history — and each answer is scored against the reference on a three\-level rubric: correct \(1\.0\), partially correct \(0\.5\), or incorrect \(0\.0\)\. The final Comprehension Score is the mean across the six questions, computed post\-education; the score reflects how effectively the educator conveyed discharge information given the patient’s persona, education level, and health literacy\. Unlike the TCS, which measures the educator’s topic coverage directly, the Comprehension Score evaluates the VP’s understanding conditioned on the dialogue history, providing an indirect measure of how well the educator delivered medical information\.

##### Factual Consistency

Inspired by QA\-based factuality evaluationScialomet al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib19)\); Fabbriet al\.\([2022](https://arxiv.org/html/2609.20827#bib.bib20)\), this metric estimates how faithfully the completed dialogue preserves the content of the source discharge note\. The LLM Judge is queried with a tailored variant of the six discharge\-focused questions used for the Comprehension Score \(Appendix Figure[12](https://arxiv.org/html/2609.20827#A1.F12)\), once conditioned on the discharge note and once on the dialogue record\. We then compute ROUGE\-LLin \([2004](https://arxiv.org/html/2609.20827#bib.bib21)\)and cosine similarity \(similarity\) between the two sets of answers, with cosine similarity computed over SentenceTransformer embeddingsReimers and Gurevych \([2019](https://arxiv.org/html/2609.20827#bib.bib28)\)\.

#### 2\.3\.2Human Evaluation

##### Virtual Patient Evaluation

We further evaluated the quality of the simulated VP by asking two licensed physicians \(specialized in emergency\) to interact with 48 randomly selected VP, as the realism and naturalness of its interactions are critical for assessing the model’s ability to engage with patients effectively\. The VP was evaluated across the following criteria:Personality \(PE\): Whether the VP adequately represents the persona it is intended to convey\.Education Level \(EL\): Whether the VP’s use of language reflects the education level it is role\-playing\.Health Literacy \(HL\): Whether the VP’s use of language reflects the health literacy level it has been assigned to role\-play\.Recall Level \(RL\): Whether the VP’s ability to recall medical and personal information is consistent with its assigned recall level\.Medical Coherency \(MC\): Whether the VP’s portrayal is coherent with its assigned medical scenario\. Here, the PE, EL, HL and RL criteria were motivated byKyunget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib25)\), which evaluated extensively on developing patient simulator\. The physicians rated each criterion with a 4\-point likert scale \(1: strongly disagree, 4: strongly agree\)\.

##### LLM\-as\-a\-Judge Validation

Since the automated evaluation in DischargeBench is delivered by an LLM Judge, we validated its outputs against physician annotations on the three judge\-driven tracks: Conversation Quality, TCS, and Comprehension Score\. Factual Consistency was excluded from this validation because it is computed deterministically from ROUGE\-L and cosine similarity and therefore involves no judge\-side scoring that could diverge from a clinician’s assessment\. We sampled 70 simulated cases stratified by educator model and asked two licensed physicians to annotate them following the same rubric given to the LLM Judge\. The set was partitioned into 25 cases assigned exclusively to each physician and 20 cases shared between the two, yielding two complementary measurements: agreement between each physician and the judge — quantified with Cohen’sκ\\kappaMcHugh \([2012](https://arxiv.org/html/2609.20827#bib.bib55)\)and Spearman’sρ\\rhoSchoberet al\.\([2018](https://arxiv.org/html/2609.20827#bib.bib56)\)— and inter\-annotator agreement \(IAA\) between the two physicians on the 20 shared cases, quantified with Cohen’sκ\\kappa\.

### 2\.4Simulation Failure Analysis

We conducted a failure analysis of the simulated conversations along three axes\. First, for each educator model we report the counts of two outcome classes: a simulation isclean\-completewhen the EMA markssession\_complete=trueafter the required minimum of 10 turns, anderroneouswhen the EMA stops the session via its early\-termination flag or when the conversation reaches the max turn \(i=100i=100\) ceiling without natural closure\. Second, we stratified the erroneous\-simulation rate across five case\-level axes — the four virtual\-patient trait dimensions \(personality, health literacy, education level, past\-medical\-history recall\) and ICD chapter222Due to space constraints, we defer the ICD\-chapter stratification to Appendix[A\.7](https://arxiv.org/html/2609.20827#A1.SS7)\.— computing the proportion of erroneous cases within each criteria to test whether specific patient traits or clinical domains disproportionately induce failures\. Third, to verify that the clean\-complete classification reflects substantively legitimate conversational closure \(rather than premature judge agreement\), we auditedsession\_completesessions where full results are in Appendix[A\.7](https://arxiv.org/html/2609.20827#A1.SS7)\.

Table 1:Evaluation of dialogues across different axes\.nn: cases with a valid evaluation output \(clean\-complete plus erroneous\-but evaluable; system\-level failures excluded\)\. Conversation Quality evaluates the Naturalness \(NA\), Responsiveness \(RE\), Clarity \(CL\), Clinical Relevance \(CR\) and Readability \(RD\)\. Topic Checklist Score \(TCS\) represents the topic coverage score based on the medical guidelines\. Comprehension Score measures the VP’s understanding of discharge instructions conditioned on the dialogue history\. Factual Consistency measures the dialogue’s factuality score based on ROUGE\-L and cosine similarity, using the discharge note as the reference\.

## 3Experiments

### 3\.1Benchmark Models

We evaluated both open\- and closed\-source large language models on DischargeBench\. Open\-source models included Llama\-3\.3Grattafioriet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib65)\), Qwen3Yanget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib5)\), and MedGemmaSellergrenet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib22)\)\. The closed\-source model evaluated were GPT\-5 variantsSinghet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib23)\)from OpenAI\.

### 3\.2Evaluation Results

#### 3\.2\.1Conversation Quality

As shown in Table[1](https://arxiv.org/html/2609.20827#S2.T1), the GPT\-5 family consistently led on Conversation Quality: GPT\-5\.4\-nano and GPT\-5\.5 tied for the top Naturalness score \(4\.616\), GPT\-5\.4\-nano led on Responsiveness \(4\.956\), GPT\-5\.5 led on Clarity \(4\.994\), and all three GPT\-5 variants saturated Clinical Relevance at 5\.000\. MedGemma\-27b\-text\-it was the strongest open\-source model, scoring within 0\.1 of the GPT\-5 family on every axis\. For Readability, GPT\-5\.4\-mini and GPT\-5\.4\-nano produced the most verbose output \(Flesch–Kincaid Grade Level≈\\approx10\.4\) — markedly harder to read than the Qwen3 and MedGemma\-27b\-text\-it outputs \(6\.5–6\.9\) — whereas GPT\-5\.5 dropped to a Grade Level of 8\.6, closing much of the gap with the open\-source models\. The smaller models lagged behind: Qwen3\-4B had the lowest Naturalness \(2\.499\), while MedGemma\-4b\-it scored lowest on Responsiveness, Clarity, and Clinical Relevance \(4\.258, 3\.945, and 3\.803, respectively\)\.

#### 3\.2\.2Topic Checklist Score

GPT\-5\.5 achieved the highest TCS at 0\.708, ahead of the other GPT\-5 variants \(GPT\-5\.4\-nano: 0\.666; GPT\-5\.4\-mini: 0\.663\) and every open\-source model \(Table[1](https://arxiv.org/html/2609.20827#S2.T1)\)\. Because the TCS serves as a proxy for topic coverage, this indicates that the GPT\-5 family — and GPT\-5\.5 in particular — was the most consistent at delivering the discharge topics required by the medical guideline\. Among the open\-source models, Qwen3\-32B was the strongest \(0\.658\), narrowly ahead of Qwen3\-4B \(0\.627\) and MedGemma\-27b\-text\-it \(0\.624\); Llama\-3\.3\-70B\-Instruct \(0\.569\) and MedGemma\-4b\-it \(0\.566\) trailed the rest, suggesting that simply scaling open\-source capacity does not guarantee topic coverage\.

#### 3\.2\.3Comprehension Score

The Comprehension Score followed the same ordering as the TCS, with the GPT\-5 family leading \(Table[1](https://arxiv.org/html/2609.20827#S2.T1)\)\. GPT\-5\.5 was again the strongest at 0\.599, followed by GPT\-5\.4\-mini \(0\.566\) and GPT\-5\.4\-nano \(0\.536\)\. Among the open\-source models, Qwen3\-32B was the best at 0\.479, with Qwen3\-4B \(0\.450\) and MedGemma\-27b\-text\-it \(0\.447\) close behind\. MedGemma\-4b\-it was the lowest at 0\.384, and Llama\-3\.3\-70B\-Instruct trailed at 0\.396 despite being the largest open\-source model evaluated — a notable inversion of the size\-vs\-capability trend, suggesting that raw parameter count does not by itself guarantee that the patient learns the educator’s content\. Because the Comprehension Score is conditioned on the dialogue history rather than measuring topic coverage directly, this ordering also corroborates the TCS results: the models that cover more of the required topics also tend to leave the patient with more accurate post\-education understanding\.

#### 3\.2\.4Factual Consistency

The GPT\-5 family was the most factually consistent overall \(Table[1](https://arxiv.org/html/2609.20827#S2.T1)\)\. GPT\-5\.4\-mini achieved the highest ROUGE\-L \(0\.373\), narrowly ahead of GPT\-5\.5 \(0\.369\) and GPT\-5\.4\-nano \(0\.359\), while GPT\-5\.5 edged out GPT\-5\.4\-mini on cosine similarity \(0\.681 vs\. 0\.681 at the displayed precision; 0\.6809 vs\. 0\.6807 at four decimal places\)\. The gap to the open\-source models was substantial: the best open\-source ROUGE\-L was 0\.275 \(Qwen3\-4B\) and the best open\-source similarity was 0\.624 \(MedGemma\-27b\-text\-it\), roughly 0\.1 and 0\.06 below the GPT\-5 leaders on the two metrics, respectively\. Collectively, these results indicate that the GPT\-5 models reproduce the content of the source discharge note more faithfully — both lexically and semantically — than any of the evaluated open\-source models\.

![Refer to caption](https://arxiv.org/html/2609.20827v1/x2.png)Figure 2:Per\-model performance stratified across the 24 ICD chapters along all evaluation axes:Conversation Quality\(Naturalness, Responsiveness, Clarity, Clinical Relevance, Readability Grade Level\),TCS,Comprehension Score, andFactual Consistency\(ROUGE\-L and cosine similarity to the source discharge note\)\. The abbreviations can be found at Appendix[A\.7](https://arxiv.org/html/2609.20827#A1.SS7)\.
#### 3\.2\.5Stratification by ICD Chapter

Stratifying performance by ICD chapter reveals substantial cross\-chapter heterogeneity \(Figure[2](https://arxiv.org/html/2609.20827#S3.F2)\)\. On Conversation Quality, the GPT\-5 family and MedGemma\-27b\-text\-it received consistently high ratings across chapters\. MedGemma\-4b\-it was the clearest exception: its Clinical Relevance score dropped sharply on several chapters, bottoming out at 3\.10 in Neoplasms \(NEO\)\. MedGemma\-27b\-text\-it produced lower Readability Grade Level scores than the GPT\-5 family in every ICD chapter, indicating that its responses were more accessible to lay readers\. TCS, by contrast, varied markedly across chapters even for the strongest models: GPT\-5\.5 scored highest in NEO \(0\.78\) and lowest in Diseases of the Blood and Blood\-Forming Organs \(BLD; 0\.61\), a gap of 0\.17 points\. For weaker models the spread was wider still — MedGemma\-4b\-it ranged from 0\.68 in Congenital Anomalies \(CON\) down to 0\.45 in Complications of Pregnancy, Childbirth, and the Puerperium \(CPC\), a gap of 0\.23 points\. Comprehension Score likewise varied across chapters; notably, scores in Factors Influencing Health Status and Contact with Health Services \(FHS\) were particularly low across all models, including GPT\-5\.5\. Factual Consistency also varied across chapters for both ROUGE\-L and cosine similarity\. The strongest performance came from GPT\-5\.4\-mini, whose scores spanned 0\.30–0\.43 on ROUGE\-L and 0\.62–0\.73 on cosine similarity, while Llama\-3\.3\-70B\-Instruct produced the lowest scores on both metrics\.

#### 3\.2\.6Stratification by Personality

![Refer to caption](https://arxiv.org/html/2609.20827v1/x3.png)Figure 3:Stratification by Personality: Neutral \(NEU\), High Conscientiousness \(HCO\), Minimiser \(MIN\), Anxious \(ANX\), Distrustful \(DIS\)\. Small models \(Qwen3\-4B, MedGemma\-4b\-it\) exhibit a noticeable drop in performance on the more challenging Anxious and Distrustful personalities, whereas larger models remain comparatively stable across personality types\.Stratifying by patient personality showed that Minimiser \(MIN\), Anxious \(ANX\), and Distrustful \(DIS\) patients were systematically harder for the educator agents, mirroring difficulties documented in real clinical encounters \(Figure[3](https://arxiv.org/html/2609.20827#S3.F3)\)\. On Naturalness, every model — including the GPT\-5 family and MedGemma\-27b\-text\-it — scored lower on MIN/ANX/DIS than on Neutral \(NEU\) and High\-Conscientiousness \(HCO\) personas\. Readability Grade Levels rose for HCO, ANX, and DIS patients, indicating that the educators produced longer, more verbose explanations for these personas\. TCS was also personality\-dependent: for GPT\-5\.5, NEU, HCO, and MIN patients all received≥0\.72\\geq 0\.72, whereas ANX and DIS dropped to 0\.70 and 0\.67 respectively\. The GPT\-5 family remained comparatively stable across personalities, while Llama\-3\.3\-70B\-Instruct showed the widest gap — its lowest TCS came on DIS \(0\.51\) and its highest on HCO \(0\.62\), a 0\.11\-point swing driven entirely by patient temperament\. Comprehension scores followed the same pattern: MIN, ANX, and DIS personas elicited lower scores than NEU patients across all models, with MedGemma\-4b\-it and Llama\-3\.3\-70B\-Instruct showing the steepest drops relative to their NEU baseline\. ROUGE\-L and cosine similarity also degraded on the difficult personas, and notably a smaller but consistent degradation was visible even within the GPT\-5 family, indicating that conversations with MIN/ANX/DIS patients are systematically more prone to factually deviated dialogue regardless of model scale\.

### 3\.3Physician Evaluation

##### Virtual Patient Evaluation

Physicians rated the VP as*Agree*or*Strongly Agree*on at least 92% of annotations across all five criteria \(Appendix Figure[13](https://arxiv.org/html/2609.20827#A1.F13)\), and no case was ever marked*Strongly Disagree*\. Personality drew the highest disagreement rate \(8%\), identifying persona portrayal as the most challenging dimension for the VP while still leaving the overall representation well within a range that clinicians judged reliable\.

##### Validating LLM\-as\-a\-Judge

For Conversation Quality, the physician–LLM Judge mean Spearmanρ\\rhowas 0\.276 and mean weightedκ\\kappawas 0\.238, showing weak positive correlation and agreement\. For TCS, pooledκ\\kappawas 0\.249, and on Comprehension Score it was 0\.171, indicating slight agreement\. These results show that the LLM Judge’s calibration must be interpreted with caution\. Inter\-physician agreement on the 20 shared cases provides the corresponding ceiling: only fair agreement for Conversation Quality \(mean Spearmanρ=0\.213\\rho=0\.213, mean weightedκ=0\.223\\kappa=0\.223\), but moderate agreement for TCS \(pooledκ=0\.572\\kappa=0\.572\) and Comprehension Score \(weightedκ=0\.562\\kappa=0\.562\)\. Full results are reported in Appendix[A\.6\.4](https://arxiv.org/html/2609.20827#A1.SS6.SSS4.Px2)\.

### 3\.4Failure Analysis Results

![Refer to caption](https://arxiv.org/html/2609.20827v1/x4.png)Figure 4:Failure Analysis\. \(A\) Distribution of clean\-complete and erroneous simulation cases\. \(B\) Failed cases stratified by the VP’s settings \(Personality — NEU: Neutral, HCO: High Conscientiousness, MIN: Minimiser, ANX: Anxious, DIS: Distrustful; Health Literacy — HI: High, LO: Low; Education Level — EL: Elementary, HS: High School, CO: College; Past Medical History Recall Level — ACC: Accurate, PAR: Partial, POO: Poor\)\.Completion rates varied substantially by educator backbone \(Figure[4](https://arxiv.org/html/2609.20827#S3.F4)A\)\. GPT\-5\.4\-nano achieved the highestclean\-completionrate at 96\.9%, marginally above Llama\-3\.3\-70B\-Instruct \(96\.6%\)\. MedGemma\-4b\-it had the most failures \(72 cases, 15\.1%\), followed by Qwen3\-4B \(38 cases, 8\.0%\)\. Overall, every model completed at least 84% of simulations cleanly\. A small number of system\-level failures — context\-length overruns and JSON\-decoding errors \(6 of 3,816 simulations\) — occurred but did not materially affect the evaluation\.

In the stratified analysis \(Figure[4](https://arxiv.org/html/2609.20827#S3.F4)B\), patient personality drove the largest share of failures: Anxious \(ANX\) and Distrustful \(DIS\) personas accounted for most errors across models\. In Distrustful patients in particular, every model exhibited its highest per\-personality failure rate, ranging from 7% for Llama\-3\.3\-70B\-Instruct \(most robust\) to 26% for MedGemma\-4b\-it \(least\)\. Conversely, all models performed comparatively well on Neutral \(NEU\) and High\-Conscientiousness \(HCO\) patients relative to the harder Minimiser, Anxious, and Distrustful types\. For health literacy, patterns were model\-dependent: GPT\-5\.4\-nano, Qwen3\-32B, and Qwen3\-4B failed more often on low\-health\-literacy VPs, while the remaining models showed the opposite trend\. For education level, every model had its lowest error rate on elementary\-school VPs, with high\-school and college rates varying inconsistently\. For past\-medical\-history recall, behavior split by family — Qwen models showed monotonically rising error rates as recall declined, MedGemma models failed least onpartial\-recall patients, and the GPT\-5 family was largely insensitive to recall level\.

## 4Related Work

##### Discharge Education with LLMs

Recently, studies have started utilizing LLMs for patient educationYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\); Janget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib10)\); Willet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib41)\); Zhouet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib42)\)\. The most similar study from benchmarked open\- and closed\-source models on discharge patient education using 49 real\-world cases, but the scopes are limitedYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\)\.Janget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib10)\)used a lightweight LLM trained with Proximal Policy Optimization \(PPO\) on a patient discharge education scenario, but only tested it on a handful of cases\.Willet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib41)\); Zhouet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib42)\)used LLMs to generate patient\-education materials, which were found helpful for patients\. However, these studies focus on generating educational materials\.

##### Patient Simulation

Kyunget al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib25)\)proposed a realistic simulation grounded in clinical literature\. They showed that, using fine\-grained factors and persona descriptions, it is possible to role\-play a virtual patient in a diagnostic conversation\.Yaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\); Cooket al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib48)\)also simulated patients using ChatGPT with diverse input parameters and by prompt engineering\.Louieet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib29)\)likewise implemented patient simulation by prompting LLMs; their system dynamically incorporates human experts’ perspectives, adjusting and refining the model’s initial output\. However, the framework constantly requires human intervention\.

## 5Conclusion

DischargeBench evaluates LLMs as discharge educators through persona\-grounded, open\-ended simulation against the source discharge note\. We find that aggregate scores conceal clinically relevant variations across ICD chapters and patient personas, with difficult personas exposing coverage failures, comprehension gaps, and reduced source\-answer agreement\. Evaluation of LLMs for discharge education should center post\-education patient understanding, not text quality or answer accuracy in isolation\.

## 6Limitations

DischargeBench has several limitations\. First, the benchmark currently considers only English\-speaking scenarios and assumes that both the educator and the patient are fluent in English\. Second, the underlying data is drawn from the MIMIC\-IV databaseJohnsonet al\.\([2023a](https://arxiv.org/html/2609.20827#bib.bib7)\), which does not capture the full diversity of real patient populations; in addition, our curation explicitly excludes patients with psychiatric or cognitive conditions \(Appendix[A\.1](https://arxiv.org/html/2609.20827#A1.SS1)\) and therefore assumes a baseline ability to engage in dialogue — an assumption that will not hold for every clinical scenario\. Together, these scoping choices mean our results reflect only a partial view of model capabilities\. Third, we evaluate a limited set of closed\- and open\-source models, which may not span the full range of contemporary systems\. Fourth, DischargeBench has not yet been validated against real patient\-education encounters\. Fifth, results are reported from a single run per case; we do not provide variance estimates from repeated trials\. Sixth, Factual Consistency captures source\-answer agreement, not clinical safety\. ROUGE\-L and embedding similarity cannot flag unsafe advice that is plausibly worded aligned with the source\.

## 7Ethical Considerations

Our experiments use data derived from MIMIC\-IV \(v3\.1\)Johnsonet al\.\([2023a](https://arxiv.org/html/2609.20827#bib.bib7)\)and MIMIC\-IV\-Note \(v2\.2\)Johnsonet al\.\([2023b](https://arxiv.org/html/2609.20827#bib.bib37)\)\. We accessed both databases under the PhysioNet credentialed\-access Data Use Agreement \(DUA\) \(v1\.5\.0\)\. MIMIC\-IV is itself a de\-identified resource, and MIMIC\-IV\-Ext\-DischargeBench introduces no additional identifiable information; the benchmark inherits MIMIC\-IV’s de\-identification guarantees and will be released through PhysioNet under the same credentialed\-access regime\. All OpenAI API calls were issued through an organization endpoint with zero data retention enabled, consistent with the PhysioNet credentialed\-access DUA\. And all open source models were run locally\.

A central ethical concern is the potential impact of LLM\-generated medical advice on patient outcomes\. DischargeBench does not rigorously evaluate models on clinical safety, and we therefore recommend that any model validated on it undergo additional safety review — including expert auditing of model outputs for unsafe behavior — before deployment in a real\-world clinical setting\. DischargeBench is intended for evaluation only: the released artifact contains no training, development, or test split, and is not designed to be used as a training corpus\. Strong performance on DischargeBench likewise does not guarantee that a model will perform comparably in practice\. Although our framework is designed to approximate a realistic discharge\-education environment, we have not yet tested how DischargeBench scores translate to clinician\-rated performance in genuine encounters, and we leave this simulation\-to\-practice validation as future work\.

## 8Acknowledgment

We used Claude \(Anthropic\) throughout the manuscript for language polishing and for drafting Appendix section\. All AI\-generated text was reviewed and edited by the authors, who verified its accuracy and take full responsibility for the content, claims, and analyses\. All of the experiments were aided by Claude Code, where the codes and results were manually reviewed by the authors\.

## References

- C\. Becker, S\. Zumbrunn, K\. Beck, A\. Vincent, N\. Loretz, J\. Müller, S\. A\. Amacher, R\. Schaefert, and S\. Hunziker \(2021\)Interventions to Improve Communication at Hospital Discharge and Rates of Readmission: A Systematic Review and Meta\-analysis\.JAMA Network Open4\(8\),pp\. e2119346\.External Links:ISSN 2574\-3805,[Link](https://doi.org/10.1001/jamanetworkopen.2021.19346),[Document](https://dx.doi.org/10.1001/jamanetworkopen.2021.19346)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- C\. E\. Chua, N\. L\. Y\. Clara, M\. S\. Furqan, J\. L\. W\. Kit, A\. Makmur, Y\. C\. Tham, A\. Santosa, and K\. Y\. Ngiam \(2024\)Integration of customised LLM for discharge summary generation in real\-world clinical settings: a pilot study on RUSSELL GPT\.The Lancet Regional Health – Western Pacific51\(English\)\.External Links:ISSN 2666\-6065,[Link](https://www.thelancet.com/journals/lanwpc/article/PIIS2666-6065(24)00205-0/fulltext),[Document](https://dx.doi.org/10.1016/j.lanwpc.2024.101211)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- Virtual Patients Using Large Language Models: Scalable, Contextualized Simulation of Clinician\-Patient Dialogue With Feedback\.Journal of Medical Internet Research27,pp\. e68486\(en\)\.External Links:ISSN 1438\-8871,[Link](https://www.jmir.org/2025/1/e68486),[Document](https://dx.doi.org/10.2196/68486)Cited by:[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px2.p1.1)\.
- C\. DeSai, K\. Janowiak, B\. Secheli, E\. Phelps, S\. McDonald, G\. Reed, and A\. Blomkalns \(2021\)Empowering patients: simplifying discharge instructions\.BMJ Open Quality10\(3\) \(en\)\.External Links:ISSN 2399\-6641,[Link](https://bmjopenquality.bmj.com/content/10/3/e001419),[Document](https://dx.doi.org/10.1136/bmjoq-2021-001419)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px3.p1.8),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px4.p1.1)\.
- A\. Fabbri, C\. Wu, W\. Liu, and C\. Xiong \(2022\)QAFactEval: Improved QA\-Based Factual Consistency Evaluation for Summarization\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,M\. Carpuat, M\. de Marneffe, and I\. V\. Meza Ruiz \(Eds\.\),Seattle, United States,pp\. 2587–2601\.External Links:[Link](https://aclanthology.org/2022.naacl-main.187/),[Document](https://dx.doi.org/10.18653/v1/2022.naacl-main.187)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px5.p1.1)\.
- B\. M\. Gillespie, L\. Thalib, E\. Harbeck, G\. Tobiano, E\. Kang, S\. Tobiano, M\. Tong, J\. Clark, B\. Patel, and W\. Chaboyer \(2023\)Effectiveness of discharge education for patients undergoing general surgery: A systematic review and meta\-analysis\.International Journal of Nursing Studies140,pp\. 104471\(en\)\.External Links:ISSN 00207489,[Link](https://linkinghub.elsevier.com/retrieve/pii/S0020748923000366),[Document](https://dx.doi.org/10.1016/j.ijnurstu.2023.104471)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- A\. L\. Goldberger, L\. A\. N\. Amaral, L\. Glass, J\. M\. Hausdorff, P\. Ch\. Ivanov, R\. G\. Mark, J\. E\. Mietus, G\. B\. Moody, C\.\-K\. Peng, and H\. E\. Stanley \(2000\)PhysioBank, PhysioToolkit, and PhysioNet: components of a new research resource for complex physiologic signals\.Circulation101\(23\),pp\. e215–e220\.Note:Circulation Electronic Pages: http://circ\.ahajournals\.org/content/101/23/e215\.full PMID:1085218; doi: 10\.1161/01\.CIR\.101\.23\.e215Cited by:[§2\.1\.2](https://arxiv.org/html/2609.20827#S2.SS1.SSS2.p1.1)\.
- D\. C\. Gonçalves\-Bradley, N\. A\. Lannin, L\. Clemson, I\. D\. Cameron, and S\. Shepperd \(2022\)Discharge planning from hospital\.Cochrane Database of Systematic Reviews\(2\) \(en\)\.External Links:ISSN 1465\-1858,[Link](https://www.cochranelibrary.com/cdsr/doi/10.1002/14651858.CD000313.pub6/full?highlightAbstract=planning%7Chospit%7Cdischarg%7Cdischarge%7Cfrom%7Cto%7Chospital%7Cplan%7Chome),[Document](https://dx.doi.org/10.1002/14651858.CD000313.pub6)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. v\. d\. Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. v\. d\. Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. d\. Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The Llama 3 Herd of Models\.arXiv\.Note:arXiv:2407\.21783 \[cs\.AI\]External Links:[Link](http://arxiv.org/abs/2407.21783),[Document](https://dx.doi.org/10.48550/arXiv.2407.21783)Cited by:[§2\.2\.1](https://arxiv.org/html/2609.20827#S2.SS2.SSS1.p1.1),[§3\.1](https://arxiv.org/html/2609.20827#S3.SS1.p1.1)\.
- S\. K\. Gudipati \(2025\)Chatbot Evaluation Frameworks: From BLEU and F1 to Multi\-Dimensional Real\-World Benchmarks\.InProceedings of the 2025 International Conference on Management Science and Computer Engineering,pp\. 228–235\.External Links:ISBN 979\-8\-4007\-1596\-9,[Link](https://dl.acm.org/doi/10.1145/3760023.3760060)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px2.p3.1)\.
- L\. Hains, O\. Kleinig, A\. Murugappa, S\. Gluck, J\. Marks, T\. Gilbert, and S\. Bacchi \(2025\)Large language model discharge summary preparation using real‐world electronic medical record data shows promise\.Internal Medicine Journal55\(7\),pp\. 1188–1192\.External Links:ISSN 1444\-0903,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC12240013/),[Document](https://dx.doi.org/10.1111/imj.70073)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- \[12\]\(2015\-05\)Health Literacy\.\(EN\)\.External Links:[Link](https://www.nih.gov/institutes-nih/nih-office-director/office-communications-public-liaison/clear-communication/health-literacy)Cited by:[§A\.2\.3](https://arxiv.org/html/2609.20827#A1.SS2.SSS3.p1.1)\.
- W\. S\. Jang, H\. Tran, M\. Mistry, S\. Gandluri, Y\. Zhang, S\. Sultana, S\. Kown, Y\. Zhang, Z\. Yao, and H\. Yu \(2025\)Chatbot To Help Patients Understand Their Health\.arXiv\.Note:arXiv:2509\.05818 \[cs\]External Links:[Link](http://arxiv.org/abs/2509.05818),[Document](https://dx.doi.org/10.48550/arXiv.2509.05818)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px3.p1.8),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px4.p1.1),[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px1.p1.1)\.
- Y\. Jiang, K\. C\. Black, G\. Geng, D\. Park, J\. Zou, A\. Y\. Ng, and J\. H\. Chen \(2025\)MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents\.NEJM AI2\(9\),pp\. AIdbp2500144\.External Links:[Link](https://ai.nejm.org/doi/full/10.1056/AIdbp2500144),[Document](https://dx.doi.org/10.1056/AIdbp2500144)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- A\. E\. W\. Johnson, L\. Bulgarelli, L\. Shen, A\. Gayles, A\. Shammout, S\. Horng, T\. J\. Pollard, S\. Hao, B\. Moody, B\. Gow, L\. H\. Lehman, L\. A\. Celi, and R\. G\. Mark \(2023a\)MIMIC\-IV, a freely accessible electronic health record dataset\.Scientific Data10\(1\),pp\. 1\(en\)\.Note:Number: 1External Links:ISSN 2052\-4463,[Link](https://www.nature.com/articles/s41597-022-01899-x),[Document](https://dx.doi.org/10.1038/s41597-022-01899-x)Cited by:[§A\.1](https://arxiv.org/html/2609.20827#A1.SS1.SSS0.Px1.p1.1),[§A\.2\.1](https://arxiv.org/html/2609.20827#A1.SS2.SSS1.p1.1),[§1](https://arxiv.org/html/2609.20827#S1.p3.1),[§2\.1\.1](https://arxiv.org/html/2609.20827#S2.SS1.SSS1.p1.1),[§6](https://arxiv.org/html/2609.20827#S6.p1.1),[§7](https://arxiv.org/html/2609.20827#S7.p1.1)\.
- A\. Johnson, T\. Pollard, S\. Horng, L\. A\. Celi, and R\. Mark \(2023b\)MIMIC\-IV\-Note: Deidentified free\-text clinical notes\.PhysioNet\.Note:Version 2\.2External Links:[Document](https://dx.doi.org/10.13026/1n74-ne17),[Link](https://doi.org/10.13026/1n74-ne17)Cited by:[§A\.1](https://arxiv.org/html/2609.20827#A1.SS1.SSS0.Px1.p1.1),[§A\.2\.1](https://arxiv.org/html/2609.20827#A1.SS2.SSS1.p1.1),[§1](https://arxiv.org/html/2609.20827#S1.p3.1),[§2\.1\.1](https://arxiv.org/html/2609.20827#S2.SS1.SSS1.p1.1),[§2\.2\.1](https://arxiv.org/html/2609.20827#S2.SS2.SSS1.p1.1),[§7](https://arxiv.org/html/2609.20827#S7.p1.1)\.
- E\. Kang, B\. M\. Gillespie, G\. Tobiano, and W\. Chaboyer \(2022\)Development of a web\-based discharge education intervention to improve the postdischarge recovery of general surgical patients\.Journal of Nursing Scholarship54\(2\),pp\. 143–151\(영어\)\.Note:Num Pages: 143\-151External Links:[Link](https://www.proquest.com/docview/2649764562/abstract/2F82071B48ED4F5APQ/1),[Document](https://dx.doi.org/10.1111/jnu.12717)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- J\. P\. Kincaid, Jr\. Fishburne, R\. Robert P\., C\. Richard L\., and Brad S\. \(1975\)Derivation of New Readability Formulas \(Automated Readability Index, Fog Count and Flesch Reading Ease Formula\) for Navy Enlisted Personnel:\.Technical reportDefense Technical Information Center,Fort Belvoir, VA\(en\)\.External Links:[Link](https://apps.dtic.mil/sti/citations/tr/ADA006655),[Document](https://dx.doi.org/10.21236/ADA006655)Cited by:[§A\.2\.2](https://arxiv.org/html/2609.20827#A1.SS2.SSS2.p1.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px2.p2.1)\.
- S\. Kweon, J\. Kim, H\. Kwak, D\. Cha, H\. Yoon, K\. Kim, J\. Yang, S\. Won, and E\. Choi \(2024\)EHRNoteQA: An LLM Benchmark for Real\-World Clinical Practice Using Discharge Summaries\.\(en\)\.External Links:[Link](https://arxiv.org/abs/2402.16040v5)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- D\. Kyung, H\. Chung, S\. Bae, J\. Kim, J\. H\. Sohn, T\. Kim, S\. K\. Kim, and E\. Choi \(2025\)PatientSim: A Persona\-Driven Simulator for Realistic Doctor\-Patient Interactions\.arXiv\.Note:arXiv:2505\.17818 \[cs\]External Links:[Link](http://arxiv.org/abs/2505.17818),[Document](https://dx.doi.org/10.48550/arXiv.2505.17818)Cited by:[§2\.2\.1](https://arxiv.org/html/2609.20827#S2.SS2.SSS1.p1.1),[§2\.3\.2](https://arxiv.org/html/2609.20827#S2.SS3.SSS2.Px1.p1.1),[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px2.p1.1)\.
- W\. Li, H\. Feng, C\. Hu, M\. Xu, and L\. Cheng \(2026\)Accurate discharge summary generation using fine tuned large language models with self evaluation\.Scientific Reports16\(1\),pp\. 5607\(en\)\.External Links:ISSN 2045\-2322,[Link](https://www.nature.com/articles/s41598-026-35552-z),[Document](https://dx.doi.org/10.1038/s41598-026-35552-z)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- C\. Lin \(2004\)Rouge: A package for automatic evaluation of summaries\.InText summarization branches out,pp\. 74–81\.External Links:[Link](https://aclanthology.org/W04-1013.pdf)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px5.p1.1)\.
- R\. Louie, A\. Nandi, W\. Fang, C\. Chang, E\. Brunskill, and D\. Yang \(2024\)Roleplay\-doh: Enabling Domain\-Experts to Create LLM\-simulated Patients via Eliciting and Adhering to Principles\.arXiv\.Note:arXiv:2407\.00870 \[cs\]External Links:[Link](http://arxiv.org/abs/2407.00870),[Document](https://dx.doi.org/10.48550/arXiv.2407.00870)Cited by:[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px2.p1.1)\.
- R\. R\. McCrae and O\. P\. John \(1992\)An Introduction to the Five\-Factor Model and its Applications\.\.Journal of Personality60\(2\),pp\. 175–215\(eng\)\.External Links:ISSN 0022\-3506,[Link](https://research.ebsco.com/linkprocessor/plink?id=8391d694-037f-3c06-b999-ddf88fe9d5df),[Document](https://dx.doi.org/10.1111/j.1467-6494.1992.tb00970.x)Cited by:[§A\.2\.5](https://arxiv.org/html/2609.20827#A1.SS2.SSS5.p1.1),[§2\.2\.1](https://arxiv.org/html/2609.20827#S2.SS2.SSS1.p1.1)\.
- M\. L\. McHugh \(2012\)Interrater reliability: the kappa statistic\.Biochemia Medica22\(3\),pp\. 276–282\.External Links:ISSN 1330\-0962,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC3900052/)Cited by:[§2\.3\.2](https://arxiv.org/html/2609.20827#S2.SS3.SSS2.Px2.p1.3)\.
- S\. Oh, H\. Choi, E\. G\. Oh, and J\. Y\. Lee \(2023\)Effectiveness of discharge education using teach\-back method on readmission among heart failure patients: A systematic review and meta\-analysis\.Patient Education and Counseling107,pp\. 107559\.External Links:ISSN 0738\-3991,[Link](https://www.sciencedirect.com/science/article/pii/S0738399122008254),[Document](https://dx.doi.org/10.1016/j.pec.2022.11.001)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- L\. F\. Rasmussen, L\. B\. Grode, J\. Lange, I\. Barat, and M\. Gregersen \(2021\)Impact of transitional care interventions on hospital readmissions in older medical patients: a systematic review\.BMJ Open11\(1\),pp\. e040057\(en\)\.External Links:ISSN 2044\-6055, 2044\-6055,[Link](https://bmjopen.bmj.com/content/11/1/e040057),[Document](https://dx.doi.org/10.1136/bmjopen-2020-040057)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- D\. A\. Redelmeier, U\. Najeeb, and E\. E\. Etchells \(2021\)Understanding Patient Personality in Medical Care: Five\-Factor Model\.Journal of General Internal Medicine36\(7\),pp\. 2111–2114\.External Links:ISSN 0884\-8734,[Link](https://pmc.ncbi.nlm.nih.gov/articles/PMC7840072/),[Document](https://dx.doi.org/10.1007/s11606-021-06598-8)Cited by:[§A\.2\.5](https://arxiv.org/html/2609.20827#A1.SS2.SSS5.p1.1),[§2\.2\.1](https://arxiv.org/html/2609.20827#S2.SS2.SSS1.p1.1)\.
- N\. Reimers and I\. Gurevych \(2019\)Sentence\-BERT: Sentence Embeddings using Siamese BERT\-Networks\.\(en\)\.External Links:[Link](https://arxiv.org/abs/1908.10084v1)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px5.p1.1)\.
- K\. Saab, J\. Freyberg, C\. Park, T\. Strother, Y\. Cheng, W\. Weng, D\. G\. T\. Barrett, D\. Stutz, N\. Tomasev, A\. Palepu, V\. Liévin, Y\. Sharma, R\. Ruparel, A\. Ahmed, E\. Vedadi, K\. Kanada, C\. Hughes, Y\. Liu, G\. Brown, Y\. Gao, S\. Li, S\. S\. Mahdavi, J\. Manyika, K\. Chou, Y\. Matias, A\. Hassidim, D\. R\. Webster, P\. Kohli, S\. M\. A\. Eslami, J\. Barral, A\. Rodman, V\. Natarajan, M\. Schaekermann, T\. Tu, A\. Karthikesalingam, and R\. Tanno \(2025\)Advancing Conversational Diagnostic AI with Multimodal Reasoning\.arXiv\.Note:arXiv:2505\.04653 \[cs\.CL\]External Links:[Link](http://arxiv.org/abs/2505.04653),[Document](https://dx.doi.org/10.48550/arXiv.2505.04653)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px1.p1.1)\.
- S\. Schmidgall, R\. Ziaei, C\. Harris, E\. Reis, J\. Jopling, and M\. Moor \(2025\)AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments\.arXiv\.Note:arXiv:2405\.07960 \[cs\]External Links:[Link](http://arxiv.org/abs/2405.07960),[Document](https://dx.doi.org/10.48550/arXiv.2405.07960)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.Px1.p1.1),[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.p1.1)\.
- P\. Schober, C\. Boer, and L\. A\. Schwarte \(2018\)Correlation Coefficients: Appropriate Use and Interpretation\.Anesthesia & Analgesia126\(5\),pp\. 1763\(en\-US\)\.External Links:ISSN 0003\-2999,[Link](https://journals.lww.com/anesthesia-analgesia/fulltext/2018/05000/correlation_coefficients__appropriate_use_and.50.aspx),[Document](https://dx.doi.org/10.1213/ANE.0000000000002864)Cited by:[§2\.3\.2](https://arxiv.org/html/2609.20827#S2.SS3.SSS2.Px2.p1.3)\.
- T\. Scialom, P\. Dray, P\. Gallinari, S\. Lamprier, B\. Piwowarski, J\. Staiano, and A\. Wang \(2021\)QuestEval: Summarization Asks for Fact\-based Evaluation\.arXiv\(en\)\.Note:arXiv:2103\.12693 \[cs\]External Links:[Link](http://arxiv.org/abs/2103.12693),[Document](https://dx.doi.org/10.48550/arXiv.2103.12693)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px5.p1.1)\.
- A\. Sellergren, S\. Kazemzadeh, T\. Jaroensri, A\. Kiraly, M\. Traverse, T\. Kohlberger, S\. Xu, F\. Jamil, C\. Hughes, C\. Lau, J\. Chen, F\. Mahvar, L\. Yatziv, T\. Chen, B\. Sterling, S\. A\. Baby, S\. M\. Baby, J\. Lai, S\. Schmidgall, L\. Yang, K\. Chen, P\. Bjornsson, S\. Reddy, R\. Brush, K\. Philbrick, M\. Asiedu, I\. Mezerreg, H\. Hu, H\. Yang, R\. Tiwari, S\. Jansen, P\. Singh, Y\. Liu, S\. Azizi, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Riviere, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Buchatskaya, J\. Alayrac, D\. Lepikhin, V\. Feinberg, S\. Borgeaud, A\. Andreev, C\. Hardin, R\. Dadashi, L\. Hussenot, A\. Joulin, O\. Bachem, Y\. Matias, K\. Chou, A\. Hassidim, K\. Goel, C\. Farabet, J\. Barral, T\. Warkentin, J\. Shlens, D\. Fleet, V\. Cotruta, O\. Sanseviero, G\. Martins, P\. Kirk, A\. Rao, S\. Shetty, D\. F\. Steiner, C\. Kirmizibayrak, R\. Pilgrim, D\. Golden, and L\. Yang \(2025\)MedGemma Technical Report\.arXiv\.Note:arXiv:2507\.05201 \[cs\]External Links:[Link](http://arxiv.org/abs/2507.05201),[Document](https://dx.doi.org/10.48550/arXiv.2507.05201)Cited by:[§3\.1](https://arxiv.org/html/2609.20827#S3.SS1.p1.1)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. J\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov, A\. Spyra, A\. Baker\-Whitcomb, A\. Beutel, A\. Karpenko, A\. Makelov, A\. Neitz, A\. Wei, A\. Barr, A\. Kirchmeyer, A\. Ivanov, A\. Christakis, A\. Gillespie, A\. Tam, A\. Bennett, A\. Wan, A\. Huang, A\. M\. Sandjideh, A\. Yang, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Gheorghe, A\. G\. Garcia, A\. Braunstein, A\. Liu, A\. Schmidt, A\. Mereskin, A\. Mishchenko, A\. Applebaum, A\. Rogerson, A\. Rajan, A\. Wei, A\. Kotha, A\. Srivastava, A\. Agrawal, A\. Vijayvergiya, A\. Tyra, A\. Nair, A\. Nayak, B\. Eggers, B\. Ji, B\. Hoover, B\. Chen, B\. Chen, B\. Barak, B\. Minaiev, B\. Hao, B\. Baker, B\. Lightcap, B\. McKinzie, B\. Wang, B\. Quinn, B\. Fioca, B\. Hsu, B\. Yang, B\. Yu, B\. Zhang, B\. Brenner, C\. R\. Zetino, C\. Raymond, C\. Lugaresi, C\. Paz, C\. Hudson, C\. Whitney, C\. Li, C\. Chen, C\. Cole, C\. Voss, C\. Ding, C\. Shen, C\. Huang, C\. Colby, C\. Hallacy, C\. Koch, C\. Lu, C\. Kaplan, C\. Kim, C\. J\. Minott\-Henriques, C\. Frey, C\. Yu, C\. Czarnecki, C\. Reid, C\. Wei, C\. Decareaux, C\. Scheau, C\. Zhang, C\. Forbes, D\. Tang, D\. Goldberg, D\. Roberts, D\. Palmie, D\. Kappler, D\. Levine, D\. Wright, D\. Leo, D\. Lin, D\. Robinson, D\. Grabb, D\. Chen, D\. Lim, D\. Salama, D\. Bhattacharjee, D\. Tsipras, D\. Li, D\. Yu, D\. J\. Strouse, D\. Williams, D\. Hunn, E\. Bayes, E\. Arbus, E\. Akyurek, E\. Y\. Le, E\. Widmann, E\. Yani, E\. Proehl, E\. Sert, E\. Cheung, E\. Schwartz, E\. Han, E\. Jiang, E\. Mitchell, E\. Sigler, E\. Wallace, E\. Ritter, E\. Kavanaugh, E\. Mays, E\. Nikishin, F\. Li, F\. P\. Such, F\. d\. A\. B\. Peres, F\. Raso, F\. Bekerman, F\. Tsimpourlas, F\. Chantzis, F\. Song, F\. Zhang, G\. Raila, G\. McGrath, G\. Briggs, G\. Yang, G\. Parascandolo, G\. Chabot, G\. Kim, G\. Zhao, G\. Valiant, G\. Leclerc, H\. Salman, H\. Wang, H\. Sheng, H\. Jiang, H\. Wang, H\. Jin, H\. Sikchi, H\. Schmidt, H\. Aspegren, H\. Chen, H\. Qiu, H\. Lightman, I\. Covert, I\. Kivlichan, I\. Silber, I\. Sohl, I\. Hammoud, I\. Clavera, I\. Lan, I\. Akkaya, I\. Kostrikov, I\. Kofman, I\. Etinger, I\. Singal, J\. Hehir, J\. Huh, J\. Pan, J\. Wilczynski, J\. Pachocki, J\. Lee, J\. Quinn, J\. Kiros, J\. Kalra, J\. Samaroo, J\. Wang, J\. Wolfe, J\. Chen, J\. Wang, J\. Harb, J\. Han, J\. Wang, J\. Zhao, J\. Chen, J\. Yang, J\. Tworek, J\. Chand, J\. Landon, J\. Liang, J\. Lin, J\. Liu, J\. Wang, J\. Tang, J\. Yin, J\. Jang, J\. Morris, J\. Flynn, J\. Ferstad, J\. Heidecke, J\. Fishbein, J\. Hallman, J\. Grant, J\. Chien, J\. Gordon, J\. Park, J\. Liss, J\. Kraaijeveld, J\. Guay, J\. Mo, J\. Lawson, J\. McGrath, J\. Vendrow, J\. Jiao, J\. Lee, J\. Steele, J\. Wang, J\. Mao, K\. Chen, K\. Hayashi, K\. Xiao, K\. Salahi, K\. Wu, K\. Sekhri, K\. Sharma, K\. Singhal, K\. Li, K\. Nguyen, K\. Gu\-Lemberg, K\. King, K\. Liu, K\. Stone, K\. Yu, K\. Ying, K\. Georgiev, K\. Lim, K\. Tirumala, K\. Miller, L\. Ahmad, L\. Lv, L\. Clare, L\. Fauconnet, L\. Itow, L\. Yang, L\. Romaniuk, L\. Anise, L\. Byron, L\. Pathak, L\. Maksin, L\. Lo, L\. Ho, L\. Jing, L\. Wu, L\. Xiong, L\. Mamitsuka, L\. Yang, L\. McCallum, L\. Held, L\. Bourgeois, L\. Engstrom, L\. Kuhn, L\. Feuvrier, L\. Zhang, L\. Switzer, L\. Kondraciuk, L\. Kaiser, M\. Joglekar, M\. Singh, M\. Shah, M\. Stratta, M\. Williams, M\. Chen, M\. Sun, M\. Cayton, M\. Li, M\. Zhang, M\. Aljubeh, M\. Nichols, M\. Haines, M\. Schwarzer, M\. Gupta, M\. Shah, M\. Huang, M\. Dong, M\. Wang, M\. Glaese, M\. Carroll, M\. Lampe, M\. Malek, M\. Sharman, M\. Zhang, M\. Wang, M\. Pokrass, M\. Florian, M\. Pavlov, M\. Wang, M\. Chen, M\. Wang, M\. Feng, M\. Bavarian, M\. Lin, M\. Abdool, M\. Rohaninejad, N\. Soto, N\. Staudacher, N\. LaFontaine, N\. Marwell, N\. Liu, N\. Preston, N\. Turley, N\. Ansman, N\. Blades, N\. Pancha, N\. Mikhaylin, N\. Felix, N\. Handa, N\. Rai, N\. Keskar, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, O\. Gleeson, P\. Mishkin, P\. Lesiewicz, P\. Baltescu, P\. Belov, P\. Zhokhov, P\. Pronin, P\. Guo, P\. Thacker, Q\. Liu, Q\. Yuan, Q\. Liu, R\. Dias, R\. Puckett, R\. Arora, R\. T\. Mullapudi, R\. Gaon, R\. Miyara, R\. Song, R\. Aggarwal, R\. J\. Marsan, R\. Yemiru, R\. Xiong, R\. Kshirsagar, R\. Nuttall, R\. Tsiupa, R\. Eldan, R\. Wang, R\. James, R\. Ziv, R\. Shu, R\. Nigmatullin, S\. Jain, S\. Talaie, S\. Altman, S\. Arnesen, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Yoo, S\. Heon, S\. Ethersmith, S\. Grove, S\. Taylor, S\. Bubeck, S\. Banesiu, S\. Amdo, S\. Zhao, S\. Wu, S\. Santurkar, S\. Zhao, S\. R\. Chaudhuri, S\. Krishnaswamy, Shuaiqi, Xia, S\. Cheng, S\. Anadkat, S\. P\. Fishman, S\. Tobin, S\. Fu, S\. Jain, S\. Mei, S\. Egoian, S\. Kim, S\. Golden, S\. Q\. Mah, S\. Lin, S\. Imm, S\. Sharpe, S\. Yadlowsky, S\. Choudhry, S\. Eum, S\. Sanjeev, T\. Khan, T\. Stramer, T\. Wang, T\. Xin, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Degry, T\. Shadwell, T\. Fu, T\. Gao, T\. Garipov, T\. Sriskandarajah, T\. Sherbakov, T\. Kaftan, T\. Hiratsuka, T\. Wang, T\. Song, T\. Zhao, T\. Peterson, V\. Kharitonov, V\. Chernova, V\. Kosaraju, V\. Kuo, V\. Pong, V\. Verma, V\. Petrov, W\. Jiang, W\. Zhang, W\. Zhou, W\. Xie, W\. Zhan, W\. McCabe, W\. DePue, W\. Ellsworth, W\. Bain, W\. Thompson, X\. Chen, X\. Qi, X\. Xiang, X\. Shi, Y\. Dubois, Y\. Yu, Y\. Khakbaz, Y\. Wu, Y\. Qian, Y\. T\. Lee, Y\. Chen, Y\. Zhang, Y\. Xiong, Y\. Tian, Y\. Cha, Y\. Bai, Y\. Yang, Y\. Yuan, Y\. Li, Y\. Zhang, Y\. Yang, Y\. Jin, Y\. Jiang, Y\. Wang, Y\. Wang, Y\. Liu, Z\. Stubenvoll, Z\. Dou, Z\. Wu, and Z\. Wang \(2025\)OpenAI GPT\-5 System Card\.arXiv\.Note:arXiv:2601\.03267 \[cs\]External Links:[Link](http://arxiv.org/abs/2601.03267),[Document](https://dx.doi.org/10.48550/arXiv.2601.03267)Cited by:[§A\.1](https://arxiv.org/html/2609.20827#A1.SS1.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2609.20827#S3.SS1.p1.1)\.
- K\. Singhal, T\. Tu, J\. Gottweis, R\. Sayres, E\. Wulczyn, L\. Hou, K\. Clark, S\. Pfohl, H\. Cole\-Lewis, D\. Neal, M\. Schaekermann, A\. Wang, M\. Amin, S\. Lachgar, P\. Mansfield, S\. Prakash, B\. Green, E\. Dominowska, B\. A\. y\. Arcas, N\. Tomasev, Y\. Liu, R\. Wong, C\. Semturs, S\. S\. Mahdavi, J\. Barral, D\. Webster, G\. S\. Corrado, Y\. Matias, S\. Azizi, A\. Karthikesalingam, and V\. Natarajan \(2023\)Towards Expert\-Level Medical Question Answering with Large Language Models\.arXiv\(en\)\.Note:arXiv:2305\.09617 \[cs\]External Links:[Link](http://arxiv.org/abs/2305.09617),[Document](https://dx.doi.org/10.48550/arXiv.2305.09617)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- \[37\]Strategy 4: Care Transitions From Hospital to Home: IDEAL Discharge Planning\.\(en\-us\)\.External Links:[Link](https://www.ahrq.gov/patient-safety/patients-families/engagingfamilies/strategy4/index.html)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- S\. P\. Trivedi, S\. Corderman, E\. Berlinberg, A\. Schoenthaler, and L\. I\. Horwitz \(2023\)Assessment of Patient Education Delivered at Time of Hospital Discharge\.JAMA Internal Medicine183\(5\),pp\. 417–423\.External Links:ISSN 2168\-6106,[Link](https://doi.org/10.1001/jamainternmed.2023.0070),[Document](https://dx.doi.org/10.1001/jamainternmed.2023.0070)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px3.p1.8),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px4.p1.1)\.
- T\. Tu, A\. Palepu, M\. Schaekermann, K\. Saab, J\. Freyberg, R\. Tanno, A\. Wang, B\. Li, M\. Amin, N\. Tomasev, S\. Azizi, K\. Singhal, Y\. Cheng, L\. Hou, A\. Webson, K\. Kulkarni, S\. S\. Mahdavi, C\. Semturs, J\. Gottweis, J\. Barral, K\. Chou, G\. S\. Corrado, Y\. Matias, A\. Karthikesalingam, and V\. Natarajan \(2024\)Towards Conversational Diagnostic AI\.arXiv\.Note:arXiv:2401\.05654 \[cs\]External Links:[Link](http://arxiv.org/abs/2401.05654),[Document](https://dx.doi.org/10.48550/arXiv.2401.05654)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.Px1.p1.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px1.p1.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px2.p3.1)\.
- E\. Vedadi, D\. Barrett, N\. Harris, E\. Wulczyn, S\. Reddy, R\. Ruparel, M\. Schaekermann, T\. Strother, R\. Tanno, Y\. Sharma, J\. Lee, C\. Hughes, D\. Slack, A\. Palepu, J\. Freyberg, K\. Saab, V\. Liévin, W\. Weng, T\. Tu, Y\. Liu, N\. Tomasev, K\. Kulkarni, S\. S\. Mahdavi, K\. Guu, J\. Barral, D\. R\. Webster, J\. Manyika, A\. Hassidim, K\. Chou, Y\. Matias, P\. Kohli, A\. Rodman, V\. Natarajan, A\. Karthikesalingam, and D\. Stutz \(2025\)Towards physician\-centered oversight of conversational diagnostic AI\.Note:arXiv:2507\.15743 \[cs\]External Links:[Link](http://arxiv.org/abs/2507.15743),[Document](https://dx.doi.org/10.48550/arXiv.2507.15743)Cited by:[§A\.3](https://arxiv.org/html/2609.20827#A1.SS3.p1.1),[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.Px1.p1.1)\.
- J\. Wang, Z\. Yao, L\. Li, J\. Qian, Z\. Yang, and H\. Yu \(2025\)ChatThero: An LLM\-Supported Chatbot for Behavior Change and Therapeutic Support in Addiction Recovery\.arXiv\.Note:arXiv:2508\.20996 \[cs\]External Links:[Link](http://arxiv.org/abs/2508.20996),[Document](https://dx.doi.org/10.48550/arXiv.2508.20996)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px2.p3.1)\.
- \[42\]\(2026\-04\)Welcome Gemma 4: Frontier multimodal intelligence on device\.External Links:[Link](https://huggingface.co/blog/gemma4)Cited by:[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px1.p1.1)\.
- J\. Will, M\. Gupta, J\. Zaretsky, A\. Dowlath, P\. Testa, and J\. Feldman \(2025\)Enhancing the Readability of Online Patient Education Materials Using Large Language Models: Cross\-Sectional Study\.Journal of Medical Internet Research27\(1\),pp\. e69955\(EN\)\.External Links:[Link](https://www.jmir.org/2025/1/e69955),[Document](https://dx.doi.org/10.2196/69955)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1),[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px1.p1.1)\.
- M\. V\. Williams, C\. White\-Williams, and J\. Li \(2026\)Hospital discharge: Best practices for a seamless transition\.Medical Clinics of North America\.External Links:ISSN 0025\-7125,[Link](https://www.sciencedirect.com/science/article/pii/S0025712525001737),[Document](https://dx.doi.org/https%3A//doi.org/10.1016/j.mcna.2025.11.012)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- Q\. Wu, G\. Bansal, J\. Zhang, Y\. Wu, B\. Li, E\. Zhu, L\. Jiang, X\. Zhang, S\. Zhang, J\. Liu, A\. H\. Awadallah, R\. W\. White, D\. Burger, and C\. Wang \(2023\)AutoGen: Enabling Next\-Gen LLM Applications via Multi\-Agent Conversation\.arXiv\.Note:arXiv:2308\.08155 \[cs\.AI\]External Links:[Link](http://arxiv.org/abs/2308.08155),[Document](https://dx.doi.org/10.48550/arXiv.2308.08155)Cited by:[§A\.3](https://arxiv.org/html/2609.20827#A1.SS3.p1.1),[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.Px1.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 Technical Report\.arXiv\.Note:arXiv:2505\.09388 \[cs\]External Links:[Link](http://arxiv.org/abs/2505.09388),[Document](https://dx.doi.org/10.48550/arXiv.2505.09388)Cited by:[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.p1.1),[§3\.1](https://arxiv.org/html/2609.20827#S3.SS1.p1.1)\.
- Z\. Yao, M\. Sun, W\. S\. Jang, S\. Kwon, S\. Kwon, and H\. Yu \(2025\)DischargeSim: A Simulation Benchmark for Educational Doctor\-Patient Communication at Discharge\.arXiv\.Note:arXiv:2509\.07188 \[cs\]External Links:[Link](http://arxiv.org/abs/2509.07188),[Document](https://dx.doi.org/10.48550/arXiv.2509.07188)Cited by:[§A\.2\.3](https://arxiv.org/html/2609.20827#A1.SS2.SSS3.p1.1),[§A\.2\.5](https://arxiv.org/html/2609.20827#A1.SS2.SSS5.p1.1),[§1](https://arxiv.org/html/2609.20827#S1.p2.1),[§2\.3\.1](https://arxiv.org/html/2609.20827#S2.SS3.SSS1.Px1.p1.1),[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px2.p1.1)\.
- H\. Yu, J\. Zhou, L\. Li, S\. Chen, J\. Gallifant, A\. Shi, X\. Li, W\. Hua, M\. Jin, G\. Chen, Y\. Zhou, Z\. Li, T\. Gupte, M\. Chen, Z\. Azizi, Y\. Zhang, T\. L\. Assimes, X\. Ma, D\. S\. Bitterman, L\. Lu, and L\. Fan \(2024\)AIPatient: Simulating Patients with EHRs and LLM Powered Agentic Workflow\.arXiv\.Note:arXiv:2409\.18924 \[cs\]External Links:[Link](http://arxiv.org/abs/2409.18924),[Document](https://dx.doi.org/10.48550/arXiv.2409.18924)Cited by:[§A\.2\.4](https://arxiv.org/html/2609.20827#A1.SS2.SSS4.p1.1),[§A\.2\.5](https://arxiv.org/html/2609.20827#A1.SS2.SSS5.p1.1),[§2\.2\.2](https://arxiv.org/html/2609.20827#S2.SS2.SSS2.p1.1)\.
- R\. A\. Yumena \(2025\)Impact of AHRQ Re\-Engineered Discharge Toolkit on Adult Patient’s 30\-Day Readmission\.Professional Case Management30\(6\),pp\. 236\(en\-US\)\.External Links:ISSN 1932\-8087,[Link](https://journals.lww.com/professionalcasemanagementjournal/fulltext/2025/11000/impact_of_ahrq_re_engineered_discharge_toolkit_on.2.aspx),[Document](https://dx.doi.org/10.1097/NCM.0000000000000801)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- P\. I\. L\. C\. Zandifar, C\. K\. Mills, E\. Siderman, and R\. Pangilinan \(2025\)Improving Patient Discharge Experience from the Outpatient Perioperative Unit by Providing Timely and Appropriate Discharge Education\.Journal of PeriAnesthesia Nursing40\(4\),pp\. e67\.External Links:ISSN 1089\-9472,[Link](https://www.sciencedirect.com/science/article/pii/S1089947225002473),[Document](https://dx.doi.org/10.1016/j.jopan.2025.05.087)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p1.1)\.
- M\. Zhang, Y\. Shen, Z\. Li, H\. Sha, B\. Hu, Y\. Wang, C\. Huang, S\. Liu, J\. Tong, C\. Jiang, M\. Chai, Z\. Xi, S\. Dou, T\. Gui, Q\. Zhang, and X\. Huang \(2025\)LLMEval\-Med: A Real\-world Clinical Benchmark for Medical LLMs with Physician Validation\.arXiv\.Note:arXiv:2506\.04078 \[cs\.CL\]External Links:[Link](http://arxiv.org/abs/2506.04078),[Document](https://dx.doi.org/10.48550/arXiv.2506.04078)Cited by:[§1](https://arxiv.org/html/2609.20827#S1.p2.1)\.
- W\. Zhong, L\. Guo, Q\. Gao, H\. Ye, and Y\. Wang \(2024\)MemoryBank: Enhancing Large Language Models with Long\-Term Memory\.Proceedings of the AAAI Conference on Artificial Intelligence38\(17\),pp\. 19724–19731\(en\)\.External Links:ISSN 2374\-3468,[Link](https://ojs.aaai.org/index.php/AAAI/article/view/29946),[Document](https://dx.doi.org/10.1609/aaai.v38i17.29946)Cited by:[§2\.2\.1](https://arxiv.org/html/2609.20827#S2.SS2.SSS1.p1.1)\.
- M\. Zhou, Y\. Pan, Y\. Zhang, X\. Song, and Y\. Zhou \(2025\)Evaluating AI\-generated patient education materials for spinal surgeries: Comparative analysis of readability and DISCERN quality across ChatGPT and deepseek models\.International Journal of Medical Informatics198,pp\. 105871\.External Links:ISSN 1386\-5056,[Link](https://www.sciencedirect.com/science/article/pii/S1386505625000887),[Document](https://dx.doi.org/10.1016/j.ijmedinf.2025.105871)Cited by:[§4](https://arxiv.org/html/2609.20827#S4.SS0.SSS0.Px1.p1.1)\.

## Appendix AAppendix

### A\.1Patient Sample Curation and Validation

MissingOveralln477Gender, n \(%\)female249 \(52\.2\)male228 \(47\.8\)Age, mean \(SD\)047\.2 \(12\.0\)Race, n \(%\)WHITE242 \(50\.7\)BLACK/AFRICAN AMERICAN75 \(15\.7\)OTHER27 \(5\.7\)ASIAN20 \(4\.2\)HISPANIC OR LATINO15 \(3\.1\)UNKNOWN12 \(2\.5\)HISPANIC/LATINO \- PUERTO RICAN10 \(2\.1\)ASIAN \- CHINESE9 \(1\.9\)HISPANIC/LATINO \- DOMINICAN9 \(1\.9\)BLACK/CARIBBEAN ISLAND7 \(1\.5\)WHITE \- OTHER EUROPEAN7 \(1\.5\)BLACK/AFRICAN6 \(1\.3\)ASIAN \- ASIAN INDIAN5 \(1\.0\)PATIENT DECLINED TO ANSWER4 \(0\.8\)MULTIPLE RACE/ETHNICITY3 \(0\.6\)ASIAN \- SOUTH EAST ASIAN3 \(0\.6\)HISPANIC/LATINO \- GUATEMALAN3 \(0\.6\)PORTUGUESE3 \(0\.6\)BLACK/CAPE VERDEAN3 \(0\.6\)HISPANIC/LATINO \- MEXICAN2 \(0\.4\)HISPANIC/LATINO \- SALVADORAN2 \(0\.4\)WHITE \- BRAZILIAN2 \(0\.4\)ASIAN \- KOREAN2 \(0\.4\)WHITE \- EASTERN EUROPEAN2 \(0\.4\)SOUTH AMERICAN1 \(0\.2\)UNABLE TO OBTAIN1 \(0\.2\)AMERICAN INDIAN/ALASKA NATIVE1 \(0\.2\)HISPANIC/LATINO \- CENTRAL AMERICAN1 \(0\.2\)ICD Chapter, n \(%\)Neoplasms20 \(4\.2\)Endocrine, Nutritional and Metabolic Diseases, and Immunity Disorders20 \(4\.2\)Diseases of the Circulatory System20 \(4\.2\)Diseases of the Musculoskeletal System and Connective Tissue20 \(4\.2\)Diseases of the Genitourinary System20 \(4\.2\)Diseases of the Digestive System20 \(4\.2\)Diseases of the Respiratory System20 \(4\.2\)Injury, Poisoning and Certain Other Consequences of External Causes20 \(4\.2\)Infectious and Parasitic Diseases20 \(4\.2\)Injury and Poisoning20 \(4\.2\)Factors Influencing Health Status and Contact with Health Services20 \(4\.2\)Diseases of the Blood and Blood\-Forming Organs20 \(4\.2\)Certain Infectious and Parasitic Diseases20 \(4\.2\)Diseases of the Skin and Subcutaneous Tissue20 \(4\.2\)Supplementary Classification of Factors Influencing Health Status and Contact with Health Services20 \(4\.2\)Diseases of the Nervous System20 \(4\.2\)Endocrine, Nutritional and Metabolic Diseases20 \(4\.2\)Symptoms, Signs, and Ill\-Defined Conditions20 \(4\.2\)Complications of Pregnancy, Childbirth, and the Puerperium20 \(4\.2\)Congenital Anomalies20 \(4\.2\)Pregnancy, Childbirth and the Puerperium20 \(4\.2\)Symptoms, Signs and Abnormal Clinical and Laboratory Findings19 \(4\.0\)Congenital Malformations, Deformations and Chromosomal Abnormalities19 \(4\.0\)Diseases of the Nervous System and Sense Organs19 \(4\.0\)Note length \(chars\), mean \(SD\)09234\.9 \(4380\.1\)

Table 2:Patient Demographics##### Source databases and inclusion criteria

Patients were drawn from the MIMIC\-IV database \(v3\.1\)Johnsonet al\.\([2023a](https://arxiv.org/html/2609.20827#bib.bib7)\)linked to the MIMIC\-IV\-Note database \(v2\.2\)Johnsonet al\.\([2023b](https://arxiv.org/html/2609.20827#bib.bib37)\)viahadm\_idandsubject\_id\. We restricted the cohort to adults aged 18–65 using theanchor\_ageattribute, and excluded any patient with a recorded diagnosis — across all admissions — of mental health or psychiatric disorders \(ICD\-10: F10–F99; ICD\-9: 290–319\) or dementia and other cognitive disorders including Alzheimer’s disease \(ICD\-10: F00–F09, G30–G31; ICD\-9: 290, 294, 331\)\. These exclusions reflect the premise that included patients have sufficient cognitive ability to engage with a discharge\-education chatbot\. The first ICD code recorded for the admission was taken as the patient’s main diagnosis\. DischargeBench further assumes that both the simulated patient and the educator are English\-speaking, which we acknowledge as a limitation \(§[6](https://arxiv.org/html/2609.20827#S6)\)\.

##### Stratified sampling across ICD chapters

Patient visits in MIMIC\-IV are coded in either ICD\-9 or ICD\-10, reflecting the U\.S\. transition between the two systems during the database’s admission window\. The chapter taxonomies of ICD\-9 and ICD\-10 share several conceptually overlapping but not 1\-to\-1 categories — for example,Symptoms, Signs, and Ill\-Defined Conditions\(ICD\-9\) versusSymptoms, Signs and Abnormal Clinical and Laboratory Findings\(ICD\-10\)\. We retain the original coding granularity and treat ICD\-9 and ICD\-10 chapters as separate strata rather than imposing a manual mapping; each patient appears in exactly one chapter\. Mapping each main\-diagnosis ICD code to its parent chapter under this scheme yields 24 chapter labels\. We then randomly sampled 20 patients per chapter, producing an initial pool of 480 cases\. This stratification ensures coverage across diagnostic categories rather than over\-representing the most common admission types\.

##### Profile and discharge\-information extraction

For each sampled case we extracted the patient’s medical profile from the structured tables: demographic variables \(age, gender, race/ethnicity\) frompatientsand admission\-level variables \(primary diagnosis, ICD code, ICD version\) fromadmissions\. From the corresponding free\-text discharge note in thedischargetable we extracted clinical variables — chief complaint, main diagnosis, reason for admission, medications on admission, allergies, medical history, and family history — together with the discharge information later used by the Educator and the LLM\-as\-a\-Judge: discharge diagnosis, new medications, treatment during stay, post\-discharge treatment, return\-to\-hospital signs and symptoms, and follow\-up appointment\. Extraction was performed with GPT\-5\.4\-miniSinghet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib23)\)ensuring zero data retention complying with[PhysioNet Credentialed Data Use Agreement \(v1\.5\.0\)](https://physionet.org/about/licenses/physionet-credentialed-health-data-license-150/)and all extracted fields were manually audited by the authors against the source note\.

##### Audit outcomes and final cohort

The audit surfaced four cases in which a psychiatric or cognitive comorbidity was mentioned in the discharge note but was not coded in the ICD history; the notes characterized these conditions as well\-controlled, showing no evidence of psychiatric or cognitive comorbidity affecting the patient at the moment of discharge\. Therefore, we retained the cases under the assumption that the comorbidity would not materially affect discharge education\. Three cases were excluded for lacking a main/primary diagnosis, yielding a final cohort of 477 patient cases\. Demographic and ICD\-chapter distributions are reported in Table[2](https://arxiv.org/html/2609.20827#A1.T2)\.

### A\.2Virtual Patient Design

We designed five distinct traits for the Virtual Patient: \(a\) Medical Profile, \(b\) Education Level, \(c\) Health Literacy, \(d\) Personality and \(e\) Past Medical History Recall\. These are the traits that define the medical scenario, character and the behavior of the Virtual Patient which we expect the LLM to impersonate as much as possible\. For Education Level, Health Literacy and Past Medical History Recall, we uniformly distribute the traits throughout the dataset\.

#### A\.2\.1Medical Profile

These profiles were the information that we extracted from the MIMIC\-IVJohnsonet al\.\([2023a](https://arxiv.org/html/2609.20827#bib.bib7)\)and MIMIC\-IV\-NoteJohnsonet al\.\([2023b](https://arxiv.org/html/2609.20827#bib.bib37)\), explained in the previous section\. This includes: age, gender, race/ethnicity, chief complaint, main diagnosis, reason for admission, medication on admission, allergy, medical history and family history\.

#### A\.2\.2Education Level

We defined a education level based onKincaidet al\.\([1975](https://arxiv.org/html/2609.20827#bib.bib66)\)that regulates the Virtual Patient’s utterance length and vocabulary\. Specifically we defined three levels : elementary, high school and college level\. For each level, we set a description and feed it into the VP’s prompt\. The behavioral descriptions injected into the system prompt are:

- •Elementary\.Has limited familiarity with medical concepts and formal language\. Struggles with terminology such as “hypertension” or “contraindication”; describes symptoms in colloquial terms \(e\.g\., “my chest feels heavy” rather than “chest tightness”\); may nod along to mask comprehension gaps; retains information better through concrete examples than abstract or written instructions\.
- •High school\.Understands common terms \(“blood pressure”, “infection”\) but may misinterpret more specific clinical language\. Can follow straightforward instructions but loses nuance \(follows “twice daily” but may miss “with food”\); asks practical, daily\-life questions; may fill knowledge gaps with information from friends or online sources\.
- •College\.Comfortable processing detailed information and engaging critically with clinical explanations\. Follows multi\-step instructions without difficulty, uses reasonably accurate terminology, and often arrives with prior independent research; carries a risk of overconfidence, skimming details they assume they already know\.

#### A\.2\.3Health Literacy

We defined health literacy[12](https://arxiv.org/html/2609.20827#bib.bib63)trait that manages the Virtual Patient’s understanding and processing of health information\. This trait was also used in DischargeSimYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\)\. We set low and high and the description of the meanings for these categories\.

- •Low\.Cannot reliably interpret prescription labels, discharge summaries, or written instructions even with common words; struggles to translate abstract health information into personal action; misinterprets numerical information such as dosing intervals; masks confusion through nodding or agreeing rather than admitting it; depends heavily on caregivers or family to interpret medical information; responds significantly better to verbal teach\-back and visual aids than to written instruction\.
- •High\.Reads and interprets discharge instructions, prescription labels, and clinical summaries accurately without clarification; translates health information into concrete self\-management behavior; communicates symptoms precisely without coaching; proactively identifies gaps or conflicts in the discharge plan \(e\.g\., overlapping side effects\); engages as a care partner rather than a passive recipient\.

#### A\.2\.4Past Medical History Recall

We defined Past Medical History Recall trait based on the studies ofYuet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib32)\)\. We set three categories : accurate, partial and poor\. The full descriptions injected into the Virtual Patient prompt are described below;

##### Poor

Have significantly limited medical history recall, often forgetting even major events\.

1. 1\.Frequently cannot recall important medical history — previous diagnoses, surgeries, hospitalisations, or family medical history\.
2. 2\.Forget key personal health information such as current medications, dosages, or medical devices in use\.
3. 3\.May contradict yourself mid\-conversation — stating something different from what you said earlier without realising it\.
4. 4\.When pressed for details you cannot remember, respond with uncertainty — ‘I think so?’, ‘I’m not sure’, or ‘my family would know better than me\.’

##### Partial

Have a moderate ability to recall medical history, remembering the broad picture but losing details\.

1. 1\.Can recall major diagnoses and significant past events \(e\.g\. a heart attack, a surgery\) but struggle with specifics — dates, medication names, or exact dosages\.
2. 2\.May remember that you take a certain medication but not its name, dose, or how long you have been on it\.
3. 3\.Occasionally confuse the sequence of events — uncertain whether something happened before or after another condition\.
4. 4\.Fill gaps in memory with approximations or guesses presented as fact — ‘I think it was about two years ago’ or ‘something beginning with M\.’

##### Accurate

Have a clear and detailed ability to recall medical history with confidence and consistency\.

1. 1\.Accurately remember all relevant health information — past conditions, surgeries, hospitalisations, family history, and current medications with correct names and dosages\.
2. 2\.Do not forget or confuse medical information across the conversation — your account remains consistent from beginning to end\.
3. 3\.Can provide specific details unprompted when directly relevant — dates, durations, prescribing doctors — without exaggerating or fabricating\.
4. 4\.If genuinely uncertain about something, say so clearly rather than guessing — ‘I don’t know the exact date but I can find out\.’

#### A\.2\.5Personality

Finally, we defined five personality types based on the works of DischargeSimYaoet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib13)\), PatientSimYuet al\.\([2024](https://arxiv.org/html/2609.20827#bib.bib32)\)and also from existing literatures of Five Factor ModelMcCrae and John \([1992](https://arxiv.org/html/2609.20827#bib.bib24)\)and its representation in patientsRedelmeieret al\.\([2021](https://arxiv.org/html/2609.20827#bib.bib26)\)\. We defined 1\) neutral, 2\) anxious, 3\) distrustful, 4\) high conscientiousness and 5\) minimiser\. The full descriptions injected into the Virtual Patient prompt are as follows:

##### Neutral

A neutral patient with no distinctive personality traits\.

1. 1\.Answers questions directly and concisely, providing only what is asked without volunteering extra information\.
2. 2\.Maintains a flat, even tone throughout — neither warm nor cold, neither anxious nor dismissive\.
3. 3\.Does not elaborate unless prompted, and does not ask questions beyond what is immediately relevant\.

##### Anxious

An overanxious patient who is excessively worried about their health and prone to catastrophising minor symptoms\.

1. 1\.Describes even mild discomforts in dramatic, alarming terms — a headache becomes a potential aneurysm\.
2. 2\.Repeatedly steers the conversation back to worst\-case diagnoses, seeking constant reassurance that nothing is seriously wrong\.
3. 3\.Asks the same fear\-driven questions multiple times even after being reassured, as the reassurance never fully lands\.
4. 4\.Jumps between unrelated health concerns mid\-conversation, revealing a restless, ongoing undercurrent of worry\.

##### Distrustful

A distrustful patient who is openly skeptical of the clinician’s knowledge, motives, and recommendations\.

1. 1\.Challenges the clinician’s expertise with pointed questions — ‘How do you know that?’ or ‘Are you sure about that?’
2. 2\.Refuses to answer questions that feel intrusive or unnecessary, responding with suspicion rather than cooperation\.
3. 3\.Frequently cites contradictory information from friends, online searches, or past experiences, treating these as more credible than the clinician\.
4. 4\.Interprets standard clinical questions as signs of incompetence or hidden agenda, creating friction at each step\.

##### High Conscientiousness

A highly conscientious patient who is disciplined, well\-prepared, and takes their health responsibilities seriously\.

1. 1\.Comes to the conversation prepared — recalls medication names, dosages, and symptom timelines accurately and in order\.
2. 2\.Asks precise, structured questions about discharge instructions, wanting to fully understand the plan before committing to it\.
3. 3\.Expresses a strong drive to follow the regimen correctly — may ask for written instructions, clarification on exact timings, or confirmation of steps\.
4. 4\.Can become visibly stressed or frustrated if instructions feel incomplete, ambiguous, or contradictory, as uncertainty conflicts with their need for structure\.
5. 5\.If they hold negative beliefs about a medication \(side effects, dependency\), their conscientiousness amplifies the concern — they will question it persistently, and may resist until fully satisfied\.

##### Minimiser

A dismissive patient who downplays symptoms and resists acknowledging the seriousness of their condition\.

1. 1\.Consistently frames serious or persistent symptoms as minor, temporary, or not worth worrying about\.
2. 2\.Underreports severity and frequency of symptoms, often rounding down — ‘a little soreness’ instead of ‘sharp pain\.’
3. 3\.Deflects concern with cheerful reassurances — ‘I’m fine, really’ — making it difficult to establish the true clinical picture\.
4. 4\.Shows no visible distress even when describing objectively distressing symptoms, projecting an air of breezy self\-sufficiency\.

### A\.3Education Monitor Agent Design

The Education Monitor Agent333The component was originally named the Environment Agent during the experimental runs, as preserved in the prompt in Figure[6](https://arxiv.org/html/2609.20827#A1.F6)and in the code release\. It was renamed post hoc to Education Monitor Agent for clarity; the underlying prompt, intervention policy, and behavior were not modified\.is a supervision component that monitors the turn\-by\-turn quality of the simulated patient–clinician conversation, intervening when necessary to prevent conversation failure while preserving the integrity of clinician evaluation data\. It operates entirely outside the conversational context visible to either agent; neither the Virtual Patient nor the Educator is aware of its existence or its actions\. These guardrailing structures were motivated by AMIEVedadiet al\.\([2025](https://arxiv.org/html/2609.20827#bib.bib53)\)and AutoGenWuet al\.\([2023](https://arxiv.org/html/2609.20827#bib.bib52)\)which also use verdicts for control\. All evaluation verdicts, intervention actions, and discarded turns are logged for downstream analysis\.

#### A\.3\.1Evaluation Mechanism

On each turn, the Education Monitor Agent \(EMA\) receives two inputs: \(1\) the full prompt that generated the turn, used as ground truth for faithfulness evaluation, and \(2\) the agent’s output text\. The EMA returns a structured verdict containing a PASS/WARN/FAIL classification, an issue code drawn from seven failure categories —role drift,prompt unfaithfulness\(faithfulness to the patient’s medical profile and instructions\),hallucination,repetition,derailment,incoherence, andpremature termination— a severity rating \(minor, moderate, or severe\), a recommended action, a correction instruction phrased as a direct behavioral directive for the offending agent, a free\-form reasoning string, and a binarysession\_completeflag used to detect natural session conclusion \(Algorithm[1](https://arxiv.org/html/2609.20827#alg1)\)\.

#### A\.3\.2Intervention Levels

Intervention level is fixed per agent role, reflecting a deliberate design decision about the purpose of each agent in the simulation\. The VP is assigned amoderatepolicy: warn verdicts and non\-severe failures \(minor or moderate severity\) trigger a soft correction, while severe failures trigger turn discard and regeneration, escalating to a hard stop only after the per\-turn retry budget \(max\_retries=2\\text\{max\\\_retries\}=2by default\) is exhausted\. The Educator is assigned anobservepolicy: every Educator turn is evaluated and logged but never modified, retried, or blocked\. The goal of the framework is to evaluate the clinician model’s discharge education capability, and any intervention on Educator’s turns would contaminate the evaluation signal, while an unrealistic patient would systematically corrupt all downstream evaluation of the Educator LLM\.

#### A\.3\.3Actions

There are three intervention actions that EMA can make:Soft correction: The current turn is accepted into conversation history, but a correction instruction is appended to the offending agent’s system prompt for its immediately following turn only, after which the system prompt is restored to its original state;Retry: The turn is discarded entirely—it is never added to conversation history and is invisible to both agents—and the correction instruction is injected into the agent’s system prompt before regeneration;Hard stop: The simulation is terminated when a turn has been retried up to the configured maximum and still fails at severe level\.

#### A\.3\.4Session Completion

In addition to per\-turn quality control, the Education Monitor Agent \(EMA\) is responsible for detecting natural session end via thesession\_completeflag\. To prevent the EMA from prematurely declaring an encounter complete after only an opening exchange, a positivesession\_completesignal is honored only once the conversation history has accumulated at least ten turns; below that threshold, the signal is logged but ignored and the simulation continues\. When no natural close is reached, the simulation terminates withsession\_completeleft as false in one of three ways: \(1\)Hard Stop— the EMA issued a hard\-stop action in response to a severe FAIL verdict or after exhausting the per\-turn retry budget \(max\_retries=2\\text\{max\\\_retries\}\{=\}2by default\); \(2\)Max Turn Hit— the conversation reached the maximum turn limit \(set to 100 in our simulations\) without natural closure; \(3\)Exception— a system\-level failure such as a maximum\-context\-length error, where the prompt grew too long for the model to continue generating\. We treat these three outcomes as erroneous session cases\.

#### A\.3\.5State Management

Algorithm 1Education Monitor Agent: per\-turn oversight of a patient–educator discharge dialogue\. Patient turns run atModerate\(regulated for realism, biased toward continuation\); educator turns run atObserve\(logged, never blocked\)\.1:history

HH, retry counts

ρ\\rho, retry budget

Rmax=2R\_\{\\max\}\{=\}2, minimum\-turns guard

Tmin=10T\_\{\\min\}\{=\}10, Education Monitor Agent

ℳ\\mathcal\{M\}
2:procedureSubmitTurn\(

r​o​l​erole,

m​s​gmsg,

pp\)

3:

t←\|H\|t\\leftarrow\|H\|; checkpoint

HH
4:

e←ℳ\.Evaluate​\(H,r​o​l​e,m​s​g,p\)e\\leftarrow\\mathcal\{M\}\.\\textsc\{Evaluate\}\(H,role,msg,p\)⊳\\triangleright⟨\\langleverdict, severity, correction,σ⟩\\sigma\\rangle

5:if

r​o​l​e=Educatorrole=\\textsc\{Educator\}or

e\.verdict=Passe\.\\text\{verdict\}=\\textsc\{Pass\}then

6:accept:append

⟨t,r​o​l​e,m​s​g,e⟩\\langle t,role,msg,e\\rangleto

HH
7:if

e\.σe\.\\sigmaand

\|H\|≥Tmin\|H\|\\geq T\_\{\\min\}thenmarksession\_complete

8:endif

9:return

\(True,∅\)\(\\textsc\{True\},\\ \\varnothing\)⊳\\trianglerightcaller proceeds; no nudge

10:endif

11:if

e\.verdict=Warne\.\\text\{verdict\}=\\textsc\{Warn\}or

e\.severity∈\{minor,moderate\}e\.\\text\{severity\}\\in\\\{\\text\{minor\},\\text\{moderate\}\\\}then

12:soft\-correct:append

⟨t,r​o​l​e,m​s​g,e,corrected⟩\\langle t,role,msg,e,\\textit\{corrected\}\\rangleto

HH
13:if

e\.σe\.\\sigmaand

\|H\|≥Tmin\|H\|\\geq T\_\{\\min\}thenmarksession\_complete

14:endif

15:return

\(True,e\.correction\)\(\\textsc\{True\},\\ e\.\\text\{correction\}\)⊳\\trianglerightnudge appended to next prompt

16:endif

17:⊳\\trianglerightsevereFail: discard and retry; hard\-stop only after budget

18:record discard

⟨t,ρ​\[t\]\+1,r​o​l​e,m​s​g,e⟩\\langle t,\\ \\rho\[t\]\+1,role,msg,e\\rangle
19:if

ρ​\[t\]≥Rmax\\rho\[t\]\\geq R\_\{\\max\}then

20:HardStop;return

\(False,∅\)\(\\textsc\{False\},\\varnothing\)
21:endif

22:

ρ​\[t\]←ρ​\[t\]\+1\\rho\[t\]\\leftarrow\\rho\[t\]\+1
23:return

\(False,e\.correction\)\(\\textsc\{False\},\\ e\.\\text\{correction\}\)⊳\\trianglerightcaller regenerates; bad turn never entersHH

24:endprocedure

### A\.4Agent Prompts

Virtual Patient prompt\.\# Instruction You are roleplaying as a real hospital patient who has just been told they are ready for discharge\. You are NOT an AI, a chatbot, or an assistant\. You are a patient\. Never break character under any circumstances\. \-\-\- \# Who You Are \-Age:\{patient\_age\} \-Gender:\{patient\_gender\} \-Race:\{patient\_race\} \#\# Your Personality You are\{personality\_description\} Embody this personality consistently in every single response\. Your tone, word choice, level of engagement, and emotional reactions must all reflect this personality at all times\. \#\# Your Education Level You are a patient with\{education\_level\}education\. \{education\_description\} Let this shape how you speak, what words you use, and how well you follow or misunderstand explanations\. \#\# Your Health Literacy You have\{health\_literacy\}health literacy\. \{health\_literacy\_description\} This affects how well you interpret medical instructions, labels, and clinical language\. \-\-\- \# Your Medical Background \#\# Current Visit \-Chief Complaint:\{chief\_complaint\} \-Primary Diagnosis:\{main\_diagnosis\} \-Reason for Admission:\{reason\_for\_admission\} \-Medication:\{medication\_on\_admission\} \-Allergy:\{allergy\} \#\# Medical History \{medical\_history\} \#\# Family Medical History \{family\_history\} \-\-\- \# Rules You Must Follow \#\# Stay in Character 1\. You are a patient\. You do not explain, summarise, or reflect on the conversation from the outside\. 2\. Never say anything that reveals you are an AI, a simulation, or a language model\. 3\. Never use clinical or teaching language \-\-\- you are not instructing anyone\. 4\. Do not volunteer information that has not been asked about\. Answer what is asked, nothing more\. \#\# Realistic Response Behavior 5\. Your responses must reflect your personality, education level, and health literacy simultaneously\. A low\-literacy, anxious patient speaks very differently from a high\-literacy, distrustful one\. 6\. Do NOT suddenly become cooperative, calm, or clear if your personality says otherwise\. Personality drift is not allowed \-\-\- stay consistent from the first message to the last\. 7\. If you do not understand something, respond the way your character would \-\-\- confusion, nodding along, or asking for clarification \-\-\- depending on your personality and literacy level\. 8\. You may express emotions appropriate to your character: worry, frustration, suspicion, cheerfulness \-\-\- but only if consistent with your defined personality\. \#\# Boundaries of Your Knowledge 9\. You only know what a real patient in your situation would know\. You do not know your full lab values, clinical notes, or the reasoning behind every decision\. 10\. Your knowledge of your past medical history is limited to\{past\_medical\_history\_recall\_level\}: \{past\_medical\_history\_recall\_description\} \#\# What You Are Doing Right Now 11\. You are in a hospital room, about to be discharged\. A clinician is speaking to you\. Your goal is not to be discharged quickly \-\-\- it is to respond authentically as this person would\. 12\. You have concerns, questions, or gaps in understanding that the clinician must address before you are truly ready to go home\. Do not pretend to be ready if you are not\. \-\-\- \# Output Format Respond with a single JSON object exactly matching this schema and nothing else: \{"utterance": "\{what the patient says out loud, in plain English prose\}"\} The "utterance" field must contain natural spoken dialogue only \-\-\- no stage directions, no internal thoughts, no labels, no markdown, no nested objects, no clinician reply, no narration\. One speaker turn per response\.

Figure 5:Virtual Patient prompt\.Education Monitor Agent prompt\.\# eval\_system\(\) You are an Environment Agent overseeing a clinical simulation between a Virtual Patient and a Clinician Agent\. \#\# Your Role You are a SILENT OBSERVER and quality controller\. \- You do NOT participate in the conversation\. \- You evaluate each submitted turn for quality, realism, and faithfulness\. \- You decide whether to approve the turn or trigger a corrective action\. \#\# Agent Roles in This Simulation \#\#\# Virtual Patient The Virtual Patient portrays a hospital patient\. It must: \- Respond only as a real patient would \-\- using lay language, natural affect, and appropriate uncertainty about medical details\. \- Stay fully consistent with any background, symptoms, and history given in its system prompt \(supplied to you separately at evaluation time\)\. \- Never exhibit clinical expertise, offer diagnoses, or steer the encounter\. \#\#\# Educator Agent The Educator Agent portrays a hospital discharge educator\. It must: \- Clearly and empathetically educate the patient about their discharge instructions, covering all applicable domains: \- Discharge diagnosis \- Medications \(names, doses, purpose, side\-effects\) \- Procedures performed during the hospital stay \- Surgery \(if applicable\) \- Post\-discharge treatment plans \- Follow\-up appointments \- Emergency action plans \(when to call 911 / return to the ED\) \- Use plain language appropriate for patient education\. \- Progress through domains systematically without skipping or rushing\. \- Respond to patient questions accurately without fabricating information\. \#\# Failure Categories \#\# Severity Levels \#\# Output Rules 1\. Never write as the patient or educator\. 2\. Be conservative \-\- prefer soft corrections over retries, retries over stops\. 3\. A naturally concluded encounter is NOT premature\_end; do not penalise it\. 4\. Respond ONLY with valid JSON \-\- no preamble, no markdown fences\. \# eval\_user\(\{history\},\{current\_role\},\{current\_message\},\{agent\_prompt\}\) \#\# Conversation History ifhistory: forturninhistory: \[Turn\{turn\.turn\_index\}\]\{turn\.role\.value \| upper\}:\{turn\.message\} ifturn\.was\_corrected: soft correction applied:\{turn\.correction\_applied\} else: \(No prior turns \-\- this is the opening message\.\) ifagent\_prompt: \#\# Prompt Given to\{current\_role\}Agent The following is the full system prompt that instructed the agent whose turn you are evaluating\. Treat it as the primary ground truth for faithfulness\. \{agent\_prompt\} \#\# Agent Response \(Turn to Evaluate\) Speaker :\{current\_role\} Message :\{current\_message\} \#\# Your Task Evaluate whether the response is: 1\. Faithful to the instructions in the prompt above\. 2\. Consistent with the conversation history\. 3\. Free of the failure categories in your system instructions\. else: \#\# Turn to Evaluate Speaker :\{current\_role\} Message :\{current\_message\} \#\# Your Task Evaluate the turn above against the conversation history and the agent role descriptions in your system instructions\. \(No agent prompt provided \-\- skip the faithfulness check\.\) Respond ONLY with a JSON object matching this exact schema: \{ "verdict": "PASS" \| "WARN" \| "FAIL", "issue": "none" \| "role\_drift" \| "hallucination" \| "repetition" \| "derailment" \| "incoherence" \| "premature\_end" \| "prompt\_unfaithful", "severity": null \| "minor" \| "moderate" \| "severe", "action": "none" \| "soft\_correction" \| "retry\_from\_last\_turn" \| "hard\_stop", "correction\_instruction": null \| "\{concrete second\-person instruction\}", "prompt\_faithfulness": "faithful" \| "partial" \| "unfaithful" \| "n/a", "faithfulness\_note": null \| "\{what specifically was contradicted\}", "reasoning": "\{one or two sentence explanation\}", "session\_complete": true \| false \} Action selection guidelines: PASS \-\> action must be "none" WARN \+ minor \-\> prefer "soft\_correction" WARN \+ moderate \-\> prefer "retry\_from\_last\_turn" FAIL \+ minor \-\> "soft\_correction" or "retry\_from\_last\_turn" FAIL \+ moderate \-\> "retry\_from\_last\_turn" FAIL \+ severe \-\> "hard\_stop" correction\_instruction must be a concrete behavioral instruction in second person directed at the speaker, with no meta\-references to this eval system\. session\_complete guidelines: Set to true ONLY when ALL of the following hold: 1\. The educator has covered every applicable domain from the role description: discharge diagnosis, medications, procedures/surgery, post\-discharge plan, follow\-up appointments, and emergency action plan\. A 2\-3 turn exchange that has only covered an opening or one topic is NEVER complete, no matter how polite the language\. 2\. The educator’s MOST RECENT turn is an explicit goodbye / sign\-off directed at ending the encounter \(e\.g\. "You’re all set to go home", "Take care", "Safe travels", "We’re done here"\)\. Generic reassurance like "let me know if you have questions" or "I’m here to help" does NOT count as closure\. 3\. The patient’s MOST RECENT turn is a closing acknowledgment that follows the educator’s sign\-off \(e\.g\. "Goodbye", "Thanks, I’m ready to go", "I understand, I’ll head home"\)\. A "thank you" used as a mid\-conversation pleasantry, or a turn that asks any new question, is NOT a closing acknowledgment\. 4\. The patient is not asking any forward\-looking questions \("what do I do\.\.\.", "can you explain\.\.\.", "what if\.\.\."\) in the most recent turn\. If ANY of \(1\)\-\(4\) is not satisfied, set session\_complete to false\. When in doubt, set false; premature termination is worse than a slightly long encounter\. \# soft\_correction\(\{issue\},\{correction\_instruction\}\) \-\-\- CORRECTION FOR THIS RESPONSE \(do not mention this to the user\): Issue detected:\{issue\} \{correction\_instruction\} \-\-\-

Figure 6:Education Monitor Agent prompt\.Educator prompt\.\# Role You are a Patient Educator conducting a discharge education session in a hospital room\. Your goal is to ensure the patient understands their discharge information clearly and feels ready to manage their health at home\. You are NOT a diagnostician\. Do not offer new diagnoses, change management plans, or speculate beyond what is documented in the discharge note below\. If a question falls outside that scope, direct the patient to follow up with their doctor\. \-\-\- \# Patient Profile Age:\{age\} Sex:\{gender\} Race:\{race\} Education level:\{education\_level\} Health literacy:\{health\_literacy\} Personality:\{personality\} Adapt your language and pace to this patient throughout the conversation\. \-\-\- \# Discharge Note \{discharge\_note\} \-\-\- \# Topics to Cover Work through the following topics in order, one at a time\. Spend as many turns as needed on each before moving on\. Skip a topic only if it is absent from the discharge note\. 1\. Opening: Greet the patient and explain the purpose of the session\. 2\. Reason for admission: Why the patient came to the hospital\. 3\. Main diagnosis: The primary condition identified, explained in plain language\. 4\. Discharge diagnoses: All diagnoses at discharge, if the patient wants to know\. 5\. Medications: Discharge medication list \(name, dose, route, purpose\)\. Highlight new medications and explain why they were started\. 6\. Tests during stay: Key tests, results, and what they mean\. 7\. Treatments during stay: Procedures performed and their purpose\. 8\. Surgery: If applicable: what was done and how recovery is going\. 9\. Post\-discharge treatment: What the patient needs to continue or start at home\. 10\. Follow\-up appointments: When, where, and why\. 11\. When to return / emergency signs: Warning signs requiring a call to 911 or an ED visit\. 12\. Closing: Summarise key points, invite final questions, and close warmly\. \-\-\- \# Communication Skills Choose the approach that fits the moment\. You may combine several per turn\. \- Greet: Open with a warm, personal greeting\. \- Listen actively: Acknowledge what the patient says before responding\. \- Use plain language: Match vocabulary to the patient’s education and health literacy\. \- Ask open\-ended questions: Invite elaboration rather than yes/no answers\. \- Let the patient finish: If mid\-thought, respond with "I see" or "Please go on\." \- Elicit concerns: If a worry comes up, explore it fully before moving on\. \- Acknowledge emotions: Name and validate distress before continuing\. \- Express empathy: Respond with warmth, not clinical detachment\. \- Explain clearly: One piece of information at a time; use everyday analogies\. \- Avoid jargon: If a medical term is unavoidable, explain it immediately in plain words\. \- Check understanding: Confirm comprehension after each important point\. \- Emphasise key messages: Restate the most important point at the end of each topic\. \- Invite questions: Regularly ask if the patient has anything they want to clarify\. \- Close warmly: End with a genuine farewell and an open invitation for future questions\. \-\-\- \# Output Format Respond only with what you, the educator, would say aloud to the patient\. Do not include reasoning, labels, stage directions, or narration\. Do not simulate the patient’s response\. One speaker turn only\. Keep each turn focused on one topic or one idea at a time\.

Figure 7:Educator prompt\.
### A\.5Automated Evaluation Miscellaneous

#### A\.5\.1Evaluation Prompts

Conversation quality evaluation prompt\.Instruction: You are a professional medical dialogue evaluator\. Below are conversations between an agent and a virtual patient\. Your task is to evaluate the quality of the conversation based on the criteria provided, focusing on the agent’s performance\. Please rate the quality of the dialogue on a 1\-\-5 Likert scale, where 1 indicates the lowest quality and 5 indicates the highest quality for each criterion\. \# Evaluation Criteria & Rating Scales 1\. Naturalness Whether the conversation flows naturally, with no repetition and a clear beginning and end\. \- 1: Highly unnatural; response feels robotic, repetitive, or abruptly cut off with no coherent flow\. \- 2: Mostly unnatural; noticeable repetition or awkward transitions that disrupt the conversation\. \- 3: Somewhat natural; minor flow issues or slight repetition, but a discernible structure is present\. \- 4: Mostly natural; conversation flows well with a clear beginning and end, and only minor stylistic roughness\. \- 5: Fully natural; conversation flows seamlessly, with no repetition, appropriate pacing, and a clear, coherent arc from start to finish\. \-\-\- 2\. Responsiveness How effectively the agent addresses the patient’s concerns, questions, and emotions\. \- 1: Completely unresponsive; ignores the patient’s questions, concerns, or emotional state entirely\. \- 2: Minimally responsive; acknowledges the patient’s input superficially but fails to meaningfully address their concerns or emotions\. \- 3: Partially responsive; addresses some concerns or emotions but overlooks key aspects of what the patient expressed\. \- 4: Mostly responsive; adequately addresses the patient’s concerns and emotions with only minor gaps or missed cues\. \- 5: Fully responsive; thoroughly and empathetically addresses all patient concerns, questions, and emotional needs in a timely and appropriate manner\. \-\-\- 3\. Clarity Whether the response is clear and easy to understand, avoiding unnecessary medical jargon and addressing one topic per turn\. \- 1: Completely unclear; response is confusing, filled with unexplained jargon, or covers multiple topics in a disorganized way\. \- 2: Mostly unclear; significant jargon or topic\-jumping makes it difficult for a patient to follow\. \- 3: Somewhat clear; generally understandable but includes occasional jargon, ambiguity, or minor topic drift\. \- 4: Mostly clear; easy to understand with minimal jargon and generally focused on one topic per turn\. \- 5: Fully clear; response is concise, plain\-language, jargon\-free, and precisely focused on a single topic per turn\. \-\-\- 4\. Clinical Relevance Whether the agent’s response is clinically valid and aligned with established medical practices\. \- 1: Clinically invalid; contains harmful, incorrect, or dangerously misleading medical information\. \- 2: Mostly irrelevant or inaccurate; response shows limited clinical grounding with notable errors or omissions that could mislead the patient\. \- 3: Partially relevant; response is broadly correct but includes inaccuracies, outdated guidance, or misses clinically important points\. \- 4: Mostly relevant; response is clinically sound and aligned with standard practice, with only minor gaps or imprecisions\. \- 5: Fully clinically relevant; response is accurate, evidence\-based, and fully aligned with established medical guidelines and best practices\. Your output format should be in a JSON format, with reasons for your rating\. Please do not output anything else\. Output format: \{ "naturalness": \{"score": 1\-\-5 rating, "reason": justification for your rating\}, "responsiveness": \{"score": 1\-\-5 rating, "reason": justification for your rating\}, "clarity": \{"score": 1\-\-5 rating, "reason": justification for your rating\}, "clinical\_relevance": \{"score": 1\-\-5 rating, "reason": justification for your rating\} \} Conversation history: formsgin\{conversation\_history\}: \-\{msg\}

Figure 8:Conversation Quality evaluation prompt\.Comprehension question prompt\.Based on the conversation you just had with the Educator, please answer the following question\. Rules: \- Answer only from what was discussed in the conversation\. Do not add information that was not mentioned\. \- If the topic was not discussed, answer "I don’t know\." \- Keep your answer concise and direct\. Question:\{question\}

Figure 9:Comprehension question prompt\.Comprehension reference prompt\.You are a medical information extractor\. Answer the following question using only information explicitly stated in the discharge note below\. Do not infer or add information beyond what is written\. If the answer is not mentioned in the note, respond with "not applicable"\. If the answer has multiple components, return them as a list\. Discharge note: \{discharge\_note\} Question:\{question\}

Figure 10:Comprehension reference prompt\.Topic checklist evaluation prompt\.Instructions: \- You are an expert medical annotator\. Your task is to analyze the conversation history based on the patient note to determine whether specific topics were discussed\. \- Please answer the following questions about the conversation based on the patient note\. For each question, provide a clear yes/no answer\. \- Only answer "yes" or "no" for topics that are mentioned in the patient note\. If a question is not answerable because it was not mentioned in the patient note \(e\.g\., no procedures were performed, no post\-discharge procedures scheduled\), answer "no" \(Not Applicable\) for that question\. \- Even if the conversation discusses contents such as discharge diagnosis, if there are any factual discrepancies, please answer as "no"\. Answer "yes" only the ones that are found in the patient note and discussed in the conversation\. \- Before providing your final answer, think through each question carefully based on the patient note\. Be thorough and precise in your evaluation\. \- The reasons should be 1\-\-2 sentences containing references or quotes\. Checklist: See Table[3](https://arxiv.org/html/2609.20827#A1.T3)for the full list of questions \(Q1\-\-Q6\.2, covering Discharge Diagnosis, New Medication, Treatment During Stay, Post\-Discharge Treatment, Return to Hospital, and Follow\-Up Appointment\)\. Patient Note: \{patient\_note\} Conversation History: \{conversation\_history\} Provide your response in the following JSON format and nothing else\. Output format: \{ "Q1": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q2": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q2\_1": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q2\_2": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q2\.3": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q2\.4": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q3": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q3\.1": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q3\.2": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q3\.3": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q4": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q4\.1": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q4\.2": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q4\.3": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q5": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q5\.1": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q5\.2": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q5\.3": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q6": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q6\.1": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\}, "Q6\.2": \{"answer": "\{yes/no\}", "reason": "\{brief reference or quote\}"\} \}

Figure 11:Topic Checklist evaluation prompt\.factual consistency evaluation prompt\.instruction: \- your job is to answer the question as accurately as possible using the provided source\. do not change or infer beyond what is written; extract only information explicitly stated in the source\. \- Here are the questions you need to answer: q1\. What was the discharge diagnosis? q2\. What treatments did the patient receive from the hospital? q3\. What were the post\-discharge treatments that were recommended to the patient? q4\. What doctors or clinics do the patient have to follow up with after this visit? q5\. For what symptoms or changes should the patient return to the ED/hospital? q6\. List the medications that were newly added during this hospitalization \(i\.e\., not part of the patient’s pre\-admission regimen\)\. For each, list its name, dosage and route\. \- Do not provide any reasons for your answer\. \- If the answer comprises multiple components, return as a list of strings\. e\.g\. \["Metoprolol 3mg", "Acetaminophen 2mg"\] \- If the answer is a single string, return a string value\. \- If the question is not answerable based on the source, respond "not applicable"\. \- Answer in the following structure format: \[ \{"question": "what is the discharge diagnosis", "answer":\{your answer\}\}, \{"question": "what treatments did the patient receive from the hospital?", "answer":\{your answer\}\}, \.\.\. \] Source: \{source\}

Figure 12:Factual Consistency evaluation prompt\.Table 3:topic checklist score questionsTable 4:open\-ended questions for checking comprehension and factual consistency
#### A\.5\.2Experimental Setup

##### Compute

All simulation and evaluation runs were executed on the Unity Research Computing Platform — a multi\-institutional cluster led by the University of Massachusetts Amherst, the University of Rhode Island, and the University of Massachusetts Dartmouth\. Dialogue simulation used 5 NVIDIA A100 80 GB GPUs \(reduced to 4 when the educator was a closed\-source model, since only the Virtual Patient and Education Monitor Agent backbones then required local serving\)\. Each model’s full simulation pass took approximately three hours of wall\-clock time\. The four\-axis automated evaluation phase used 4 NVIDIA A100 80 GB GPUs and ran for roughly four hours per model\.

##### Hyperparameters

For every simulation we set a maximum of 100 conversational turns and a per\-turn retry budget ofmax\_retries=2 \(Appendix[A\.3](https://arxiv.org/html/2609.20827#A1.SS3)\)\. The Virtual Patient, Education Monitor Agent, and Educator backbones were each given a 32,768\-token context window\. Sampling temperatures were left at each model’s released defaults; we did not retune decoding hyperparameters across the model set\. For the LLM\-as\-a\-Judge, the context window was set to 65,536 or 16,384 tokens depending on the evaluation prompt, and the temperature was fixed at 0 \(greedy decoding\) for reproducibility of the judgments\.

##### Software

DischargeBench was implemented in Python 3\.12\. Local inference used the Hugging Facetransformerslibrary \(v5\.8\.0\) together withvllm\(v0\.19\.0\); the closed\-source GPT\-5 family was accessed via the OpenAI Python SDK \(v2\.31\.0\);textstatlibrary \(v0\.7\.13\)\.

### A\.6Human Evaluation Miscellaneous

#### A\.6\.1Demographics

The two physician annotators were board\-certified emergency medicine specialists practising in South Korea, each with more than ten years of clinical experience in their specialty\. Both are native Korean speakers; the annotation interface accordingly presented bilingual English/Korean instructions \(Section[A\.6](https://arxiv.org/html/2609.20827#A1.SS6)\)\.

#### A\.6\.2Virtual Patient Simulation Quality Settings

For human evaluation, two medical experts each engaged in dialogue with the simulator across 48 patient cases sampled uniformly across the 24 ICD chapters\. The experts evaluated the quality of the simulator using the following categories \(personality, education level, health literacy, recall level, medical coherency\) on a 4\-point likert scale \(1 = Strongly disagree, 4 = Strongly agree\)

- •\(Personality\) Does the Virtual Patient adequately represent the persona it is intended to convey?
- •\(Education Level\) Does the Virtual Patient’s use of language reflect the education level it is role\-playing?
- •\(Health Literacy\) Does the Virtual Patient’s use of language reflect the health literacy level it has been assigned to role\-play?
- •\(Recall Level\) Is the Virtual Patient’s ability to recall medical and personal information consistent with its assigned recall level?
- •\(Medical Coherency\) Is the Virtual Patient’s portrayal coherent with its assigned medical scenario?

##### Annotation instructions

The annotation tool presents the following instructions to each physician before they begin a case\.

> For each case, talk to the virtual patient as if you were the discharging physician or nurse\. As you go, consider whether the patient: 1. 1\.Personality— speaks and behaves consistently with the described persona\. 2. 2\.Education level— uses language and vocabulary that match the assigned education level\. 3. 3\.Health literacy— shows understanding of medical terms and concepts consistent with the assigned low/high literacy level\. 4. 4\.Recall level— recalls \(or appropriately forgets\) medical history and symptoms in line with the assigned recall level\. 5. 5\.Medical coherence— the overall portrayal is medically consistent and faithful to the assigned clinical scenario\. After the conversation you will be asked to rate the patient on each of these five dimensions\.

#### A\.6\.3Survey Result

![Refer to caption](https://arxiv.org/html/2609.20827v1/x5.png)Figure 13:Virtual Patient Simulation Quality Survey\. Two physicians annotated 48 patient cases sampled uniformly across the 24 ICD chapters\.
#### A\.6\.4LLM\-as\-a\-Judge Evaluation Miscellaneous

For the judge\-validation study, the same two physicians annotated 70 simulated cases in a dedicated web tool\. Each case is presented in a fixed three\-stage sequence — Conversation Quality \(CQ\), Topic Checklist \(TCS\), and Comprehension — and within each stage the physicians use the same rubric that the LLM judge was prompted with \(Figures[8](https://arxiv.org/html/2609.20827#A1.F8),[11](https://arxiv.org/html/2609.20827#A1.F11),[9](https://arxiv.org/html/2609.20827#A1.F9)\)\. The identity of the educator model that produced each conversation and any prior LLM\-judge output are hidden from the physicians throughout the task\.

##### Annotation instructions

The web tool presents the following stage\-specific instructions before each block\. We reproduce the per\-stage prompts below; the full Likert\-level anchor descriptions for CQ are identical to those given to the LLM judge and are not duplicated here \(see Figure[8](https://arxiv.org/html/2609.20827#A1.F8)\)\.

Conversation Quality\.

> You are evaluating a conversation between a patient\-education agent \(the “educator”\) and a virtual patient\. Rate the conversation on four axes using a 1–5 Likert scale\.Focus on the educator’s performance\.For each axis, give one integer score from 1 \(lowest\) to 5 \(highest\)\. No written rationale is required\. 1. 1\.Naturalness— whether the conversation flows naturally, with no repetition and a clear beginning and end\. 2. 2\.Responsiveness— how effectively the agent addresses the patient’s concerns, questions, and emotions\. 3. 3\.Clarity— whether the response is clear and easy to understand, avoiding unnecessary medical jargon and addressing one topic per turn\. 4. 4\.Clinical Relevance— whether the agent’s response is clinically valid and aligned with established medical practices\.

Topic Checklist\.

> For each conversation, decide whether 21 specific discharge\-education topics were discussed\. The answer is strictlyyesornoper question\. - •Answeryesonly if the topic is*both*\(a\) mentioned in the patient note*and*\(b\) discussed in the conversation\. - •If the topic is not mentioned in the patient note \(e\.g\., no procedures were performed\), answerno\(treat “Not Applicable” asno\)\. - •Even if the conversation discusses the topic, if there is afactual discrepancywith the patient note, answerno\. Base your judgment on the patient note and the conversation only\. The 21 questions \(six parent topics with sub\-questions covering main discharge diagnosis, new medications, inpatient procedures, post\-discharge procedures, emergency signs, and follow\-up appointments\) are reproduced in Table[3](https://arxiv.org/html/2609.20827#A1.T3)\.

Comprehension\.

> The benchmark asks the virtual patient six standard questionsaftertheir conversation with the educator\. Grade how well the patient’s post\-conversation answer reflects what is in the discharge note\. For each item, label the patient’s answer as exactly one of: - •correct— fully captures the key information from the discharge note\. - •partially correct— captures some but not all of the key information\. - •incorrect— missing, wrong, or the patient said they didn’t know\. Judge*only*against the discharge note; do not bring in outside medical knowledge\. Do not penalise phrasing differences if the substance is right\. “I don’t know” answers areincorrect\.

##### LLM\-as\-a\-Judge Annotation Validation

We quantify the agreement of the LLM Judge with the two physician annotators on the three judge\-scored axes \(Conversation Quality, Topic Checklist Score, and Comprehension Score\) and compare it against physician–physician agreement on the 20 shared cases\. Results are summarized in Table[5](https://arxiv.org/html/2609.20827#A1.T5)\.

Table 5:LLM Judge agreement with physician annotators\.P–P: inter\-physician agreement computed on the 20 cases annotated by both physicians\.Judge–P: agreement between the LLM Judge and physician annotations pooled across all 70 annotated cases \(each physician contributing 25 uniquely assigned cases plus their 20 shared annotations\)\. For Conversation Quality, agreement is averaged across the four 1–5 Likert sub\-axes \(Naturalness, Responsiveness, Clarity, Clinical Relevance\)\.For Conversation Quality, the LLM Judge’s agreement with physicians is comparable to inter\-physician agreement: even two trained clinicians using the same rubric reach only modest agreement \(mean wκ\\kappa= 0\.223\), and the judge matches rather than exceeds this ceiling \(mean wκ\\kappa= 0\.238\), suggesting that subjective Likert ratings on educator behavior carry substantial inherent variance\. For the more structured axes, the judge agrees less than physicians do with each other: Topic Checklist Score shows a gap of roughly 0\.32 in pooledκ\\kappa\(P–P 0\.572 vs\. Judge–P 0\.249\), and Comprehension Score a gap of roughly 0\.39 in wκ\\kappa\(P–P 0\.562 vs\. Judge–P 0\.171\)\. We therefore interpret the LLM Judge as a useful scalable approximation for population\-level comparisons, but not as a substitute for physician review on individual cases — particularly on TCS and Comprehension judgments where it underperforms human–human agreement\.

### A\.7Failure Analysis Miscellaneous

Table 6:Full names of the ICD chapter abbreviations used in Figure[2](https://arxiv.org/html/2609.20827#S3.F2)\. Abbreviations ending in “9” \(EN9, IP9\) denote ICD\-9 chapters; the remaining abbreviations follow ICD\-10 chapter names, except where a separate ICD\-9 stem is used \(e\.g\., SSI vs\. SSL, NSS vs\. NER, CMD vs\. CON, PCP vs\. CPC, SFH vs\. FHS, CIP vs\. IPD\)\.#### A\.7\.1Stratification by ICD Chapter

![Refer to caption](https://arxiv.org/html/2609.20827v1/figures/failure_heatmap_icd_chapter.png)Figure 14:Failure Analysis of erroneous cases stratified by ICD Chapters\. The abbreviations can be found in Figure[2](https://arxiv.org/html/2609.20827#S3.F2)\.Stratifying failure analysis by ICD chapter, we found that MedGemma\-4b\-it — consistently the weakest model across our evaluation axes — had the highest incomplete\-simulation rate in most ICD chapters, with Qwen3\-4B second\. Notably, theSymptoms, Signs and Abnormal Clinical Laboratory FindingsICD\-10 chapter \(SSL\) showed high incomplete\-simulation rates across multiple models, including GPT\-5\.5, the MedGemma family, the Qwen3 family, and Llama\-3\.3\-70B\-Instruct\. Cases in this chapter present a symptom \(e\.g\., abdominal pain\) as the primary diagnosis, without a specific underlying condition\. Based on these findings, we hypothesize that diagnostically ambiguous cases pose greater challenges for the educator model\.

#### A\.7\.2Clean Complete Case analysis

Table 7:Clean complete\-case audit\.Complete \(n\): sessions retained after removing character\-break failures and short sessions with low Topic Checklist scores\.Short: sessions terminating below the 15\-turn threshold\.Char\-break: sessions in which the Virtual Patient broke character \(e\.g\., produced “as an AI” style refusals\)\.Low\-TCS short: short sessions whose Topic Checklist score also falls below the low\-TCS cutoff, expressed as a count and as a percentage of the model’s complete\-case set\. Percentages for Short and Char\-break are likewise computed against the complete\-case set\.We analyzed sessions that ended in≤15\\leq 15turns — just above the 10\-turn floor enforced by the EMA — and compared their TCS values against the wider clean\-complete distribution\. The TCS serves as a proxy for content coverage; we flagged sessions with a TCS below 0\.5 \(i\.e\., fewer than half of the required discharge topics covered\) as candidate false\-positive completions in which the EMA approved closure despite incomplete coverage\. We additionally screened every clean\-complete transcript for AI self\-disclosure — patient turns containing phrases such as “as an AI”, “language model”, or “I’m an assistant” — to identify cases in which the VP stepped out of character without EMA intervention\. Together, these two probes quantify the EMA’s false\-negative rate on session completeness and persona maintenance\.

Table[7](https://arxiv.org/html/2609.20827#A1.T7)summarises, for each model, the size of the clean complete\-case subset used in the main analysis and the two failure modes that remove sessions from it: character\-break failures \(the Virtual Patient stepping out of character\) and short sessions paired with a low Topic Checklist score\. Short dialogue termination \(≤15\\leq 15turns\) was most prevalent in the Qwen3 models \(Qwen3\-4B: 132 sessions, 30\.1%; Qwen3\-32B: 131, 29\.3%\)\. Character\-break failures — sessions in which the EMA failed to intervene despite revealing phrases such as “as an AI” — were most common in GPT\-5\.5 \(32 cases, 7\.0%\)\. Among short sessions that also fall below the low\-TCS cutoff, MedGemma\-4b\-it had the highest absolute count \(30 cases, 7\.4% of its complete\-case set\), whereas the Qwen3 family contributed comparable counts \(Qwen3\-4B: 24, 5\.5%; Qwen3\-32B: 29, 6\.5%\) despite producing far more short sessions overall — indicating that most Qwen3 short terminations still covered enough discharge topics to clear the low\-TCS threshold\.

Similar Articles