From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education

arXiv cs.CL Papers

Summary

This study examines how coded dialogues from GenAI virtual patients can provide teacher-interpretable process evidence of clinical reasoning in medical education, using learning analytics to analyze dialogue logs from medical students.

arXiv:2608.28619v1 Announce Type: new Abstract: Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue. However, these logs are educationally difficult to use directly. Complete transcripts are too detailed for routine teacher review, whereas final scores obscure whether learners followed up patient cues, checked uncertainty, or used summaries to guide later questioning. This study examined whether coded GenAI VP dialogues can provide teacher-interpretable process evidence of clinical reasoning. We analysed 1{,}030 GenAI VP dialogues from 210 second-year medical learners across five weeks chest-pain cases. Each consultation was teacher-scored using a rubric assessing the full history taking dialogue, and consultations were classified within each week as high- or low-rated using the weekly median score. To explain how rated performance was reflected in the dialogue process, we applied three analytic layers to the same coded dialogue data: behavioural prevalence, local co-occurrence using Epistemic Network Analysis, and sequential transition using Transition Network Analysis. High-rated consultations involved more history taking activity, but differences were not simply about volume. High rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, and synthesis. Summarising and organising moves more often led to verification or mechanism-oriented follow-up. These findings show how layered analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.
Original Article
View Cached Full Text

Cached at: 09/01/26, 11:52 AM

# From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education
Source: [https://arxiv.org/html/2608.28619](https://arxiv.org/html/2608.28619)
\[5,1\]\\fnmDragan\\surGašević \[2\]\\fnmYizhou\\surFan

1\]\\orgdivFaculty of Information Technology,\\orgnameMonash University,\\orgaddress\\cityMelbourne,\\postcode3800,\\countryAustralia

2\]\\orgdivGraduate School of Education,\\orgnamePeking University,\\orgaddress\\cityBeijing,\\postcode100871,\\countryChina

3\]\\orgdivDepartment of Clinical Skills Training Center,,\\orgnameShantou University Medical College,\\orgaddress\\cityShantou,\\postcode515041,\\countryChina

4\]\\orgdivOffice of Teaching Affairs,\\orgnameShantou University Medical College,\\orgaddress\\cityShantou,\\postcode515041,\\countryChina

5\]\\orgdivFaculty of Education & School of Computing and Data Science,\\orgnameThe University of Hong Kong,\\orgaddress\\cityHong Kong,\\countryChina

6\]\\orgdivSchool of Public Health,\\orgnameThe University of Hong Kong,\\orgaddress\\cityHong Kong,\\countryChina

7\]\\orgdivSchool of Public Health and Preventive Medicine,\\orgnameMonash University,\\orgaddress\\cityMelbourne,\\postcode3004,\\countryAustralia

###### Abstract

Medical history taking is a dialogue\-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds\. Generative AI\-powered virtual patients \(GenAI VPs\) make repeated history taking practice scalable and preserve full turn by turn dialogue\. However, these logs are educationally difficult to use directly\. Complete transcripts are too detailed for routine teacher review, whereas final scores obscure whether learners followed up patient cues, checked uncertainty, or used summaries to guide later questioning\. This study examined whether coded GenAI VP dialogues can provide teacher\-interpretable process evidence of clinical reasoning\. We analysed 1,030 GenAI VP dialogues from 210 second\-year medical learners across five weeks chest\-pain cases\. Each consultation was teacher\-scored using a rubric assessing the full history taking dialogue, and consultations were classified within each week as high\- or low\-rated using the weekly median score\. To explain how rated performance was reflected in the dialogue process, we applied three analytic layers to the same coded dialogue data: behavioural prevalence, local co\-occurrence using Epistemic Network Analysis, and sequential transition using Transition Network Analysis\. High\-rated consultations involved more history taking activity, but differences were not simply about volume\. High rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, and synthesis\. Summarising and organising moves more often led to verification or mechanism\-oriented follow\-up\. These findings show how layered analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process\-focused feedback in medical education\.

###### keywords:

Generative AI, Virtual Patients, Clinical Reasoning, History Taking, Learning Analytics, Human\-AI Interaction, Dialogue Trace Analysis

## 1Introduction

History taking is a central component of medical education because it requires learners to conduct a clinical conversation with a patient while simultaneously reasoning about the patient’s problem\. During a consultation, learners must gather relevant information, recognise cues, clarify uncertainty, adjust questioning, and build a coherent account of the patient’s condition\[[52](https://arxiv.org/html/2608.28619#bib.bib52),[61](https://arxiv.org/html/2608.28619#bib.bib61)\]\. Clinical reasoning therefore does not occur only when a diagnosis is stated at the end of a case; it unfolds throughout the consultation as learners judge which information matters, which symptoms require follow\-up, and whether the information gathered so far supports a candidate explanation\[[14](https://arxiv.org/html/2608.28619#bib.bib14)\]\. In this sense, history taking is a form of clinical reasoning carried out through dialogue\.

Assessing clinical reasoning as a process remains difficult\. Course assessment and clinical skills assessment often rely on final grades, task outcomes, checklist scores, diagnostic accuracy, or overall ratings\[[7](https://arxiv.org/html/2608.28619#bib.bib7),[5](https://arxiv.org/html/2608.28619#bib.bib5)\]\. These measures can indicate whether a learner met a standard, but they reveal less about how the learner reached that outcome\. In a chest pain case, for example, two learners may arrive at a similar final diagnosis, but differ substantially in whether they followed up pain characteristics, checked risk factors, clarified ambiguous answers, connected symptoms across domains, or summarised information at appropriate moments\[[6](https://arxiv.org/html/2608.28619#bib.bib6),[21](https://arxiv.org/html/2608.28619#bib.bib21)\]\. These in\-process behaviours are important because they make metacognitive regulation partly visible: the learner must monitor what is known, notice what remains uncertain, and decide how the next question should move the consultation forward\[[17](https://arxiv.org/html/2608.28619#bib.bib17)\]\.

Providing repeated opportunities to practise and assess such processes is challenging in medical education\. Standardised patient \(SP\) encounters and objective structured clinical examinations \(OSCEs\) offer authentic practice and structured assessment, but they require trained personnel, scheduling, room arrangements, and safeguards for rating consistency\[[18](https://arxiv.org/html/2608.28619#bib.bib18)\]\. These logistical demands limit how much complete\-consultation practice each learner can receive, especially in large cohorts\[[42](https://arxiv.org/html/2608.28619#bib.bib42)\]\. Even when SP or OSCE encounters are recorded, the recordings are rarely transformed into turn\-by\-turn evidence that educators can use for routine feedback on learners’ reasoning processes\[[60](https://arxiv.org/html/2608.28619#bib.bib60)\]\.

Generative AI\-powered virtual patients \(GenAI VPs\) offer a way to address limitations in existing practice and evidence generation in history taking education\. Unlike scripted or menu\-based virtual patients, a GenAI VP allows learners to ask questions freely in natural language and generates patient responses within the frame of a clinical case\[[65](https://arxiv.org/html/2608.28619#bib.bib65)\]\. As a result, the same case can unfold into different consultations depending on each learner’s questioning path\[[24](https://arxiv.org/html/2608.28619#bib.bib24)\]\. This makes repeated, large\-scale history taking practice more feasible and, importantly, preserves the full learner–GenAI VP dialogue as turn\-by\-turn data for later analysis\[[30](https://arxiv.org/html/2608.28619#bib.bib30)\]\.

This new source of dialogue data also creates an assessment problem\. The educational challenge is therefore not simply that GenAI VP systems generate large amounts of dialogue data\. The challenge is that the two most accessible forms of evidence are both limited for formative use\. A final score can indicate whether a consultation was judged successful, but it cannot show where the learner missed a cue, repeated already answered questions, failed to clarify uncertainty, or used a summary to redirect later questioning\. A full transcript contains this information, but it is too detailed for routine inspection by teachers in large cohorts and repeated practice tasks\[[55](https://arxiv.org/html/2608.28619#bib.bib55),[39](https://arxiv.org/html/2608.28619#bib.bib39)\]\. For GenAI VP practice to support clinical reasoning education, dialogue logs therefore need to be transformed into process evidence that is detailed enough to show how history taking unfolded, but structured enough for teachers to interpret\.

This need motivates the present study\. Rather than treating GenAI VP logs as raw transcripts or reducing them to final scores, we examined whether coded learner turns could reveal performance\-related patterns in history taking processes\. The comparison between high\- and low\-rated consultations was used as a performance\-anchored contrast\[[57](https://arxiv.org/html/2608.28619#bib.bib57),[20](https://arxiv.org/html/2608.28619#bib.bib20)\]\. It was not intended to classify learners as generally strong or weak\. Instead, teacher\-rated history taking scores were used to examine whether consultations judged to be of high or low quality within the same weekly case showed different dialogue processes\. This within\-week comparison was important because the five cases differed in clinical content and likely difficulty; comparing learners within the same weekly task allowed process differences to be interpreted against the same case and scoring context\.

To make these process differences interpretable, we applied a layered analytic approach to the same coded dialogue data\. Behavioural prevalence analysis examined which coded history taking behaviours occurred and how often they appeared\. Local co\-occurrence analysis examined which behaviours were connected within nearby learner turns\. Sequential transition analysis examined which behaviours tended to follow one another as the consultation unfolded\. These three layers address complementary educational questions: whether important behaviours were present, whether they were locally coordinated with other clinically relevant moves, and whether they guided the next step of the consultation\[[36](https://arxiv.org/html/2608.28619#bib.bib36),[51](https://arxiv.org/html/2608.28619#bib.bib51),[49](https://arxiv.org/html/2608.28619#bib.bib49)\]\. The present study analysed GenAI VP history taking tasks completed by 210 second\-year medical learners over five weeks\. The analysis drew on a coding scheme for GenAI VP questioning behaviours developed in earlier work\[[2](https://arxiv.org/html/2608.28619#bib.bib2)\]\.

## 2Literature Review

### 2\.1History taking as observable metacognitive regulation of clinical reasoning

History taking is an early and observable setting in which medical learners practise clinical reasoning\[[27](https://arxiv.org/html/2608.28619#bib.bib27)\]\. During history taking, learners gather information through patient dialogue, elicit symptom details, clarify ambiguity, follow clinically relevant cues, and organise information for later diagnostic and management decisions\[[32](https://arxiv.org/html/2608.28619#bib.bib32)\]\. History taking is therefore not a checklist of symptom questions\. It is a dialogue\-based inquiry practice in which learners decide what to ask, how to respond to patient answers, and when to reorganise collected information while the consultation is still unfolding\[[64](https://arxiv.org/html/2608.28619#bib.bib64)\]\.

Clinical reasoning theories explain why the organisation of history taking matters\[[9](https://arxiv.org/html/2608.28619#bib.bib9),[11](https://arxiv.org/html/2608.28619#bib.bib11)\]\. Hypothetico\-deductive accounts describe how early patient information triggers provisional explanations that later questions test and refine\[[53](https://arxiv.org/html/2608.28619#bib.bib53)\]\. Illness\-script and knowledge\-encapsulation accounts emphasise organised case\-based knowledge that helps clinicians connect symptoms, risk factors, mechanisms, and likely diagnoses\[[50](https://arxiv.org/html/2608.28619#bib.bib50)\]\. Across these accounts, competent clinical reasoning depends on metacognition: monitoring what is known, noticing what remains incomplete or uncertain, and adjusting inquiry accordingly\[[13](https://arxiv.org/html/2608.28619#bib.bib13)\]\. In history taking, these metacognitive demands can be reflected in observable behaviours such as systematic coverage, cue\-responsive follow\-up, clarification, hypothesis\-sensitive questioning, checking, and interim summarising\[[41](https://arxiv.org/html/2608.28619#bib.bib41)\]\.

The metacognitive character of history taking makes records of dialogue in patient consultations useful for process analysis\[[62](https://arxiv.org/html/2608.28619#bib.bib62)\]\. Dialogue records cannot directly reveal metacognitive states such as planning, monitoring, or uncertainty regulation\[[43](https://arxiv.org/html/2608.28619#bib.bib43)\]\. Dialogue records can, however, preserve visible counterparts of these functions: how learners organise questions, respond to cues, confirm ambiguity, integrate information, and move between routine questioning and reasoning\-oriented moves\[[62](https://arxiv.org/html/2608.28619#bib.bib62)\]\. History taking dialogue therefore provides observable behavioural traces of clinical reasoning, even though these traces are not direct measurements of cognition\[[58](https://arxiv.org/html/2608.28619#bib.bib58)\]\.

### 2\.2From simulated encounters to analysable GenAI VP dialogue corpora

Simulation\-based formats have long supported history taking practice and assessment in medical education\. SP encounters and OSCEs provide interactive clinical scenarios and structured opportunities to judge learner performance\[[35](https://arxiv.org/html/2608.28619#bib.bib35)\]\. For process analysis of history taking, however, the key issue is not whether these formats are educationally valuable, but whether they can routinely generate fine\-grained, comparable, and reusable process data at scale\. In many teaching settings, the interaction itself remains difficult to convert into turn\-by\-turn evidence for everyday feedback and research, even when performance ratings or recordings are available\[[38](https://arxiv.org/html/2608.28619#bib.bib38),[60](https://arxiv.org/html/2608.28619#bib.bib60)\]\.

Earlier virtual patient systems addressed some problems of standardisation and accessibility by allowing learners to work through repeatable clinical cases\[[15](https://arxiv.org/html/2608.28619#bib.bib15)\]\. These systems also made certain learner actions easier to record, such as menu selections, chosen pathways, time on task, and diagnostic decisions\[[28](https://arxiv.org/html/2608.28619#bib.bib28)\]\. Yet many earlier VP designs constrained learner history taking through predefined options or scripted branches\[[26](https://arxiv.org/html/2608.28619#bib.bib26)\]\. For instance, branched VP systems such as those used in the TAME project structured interactions around an "ideal" pathway of patient management decisions, with linear variants offering no deviation from scripted routes\[[63](https://arxiv.org/html/2608.28619#bib.bib63)\]\. Similarly, VPs commonly assessed clinical reasoning through multiple\-choice questions or discrete decision points, while the non\-linearity of clinical reasoning posed a persistent challenge for scoring and feedback that quantitative methods could not sufficiently capture\[[22](https://arxiv.org/html/2608.28619#bib.bib22)\]\. Such designs supported consistency, but they offered limited evidence about how learners formulate their own questions, redirect inquiry after patient responses, or reorganise information during an open\-ended consultation\[[33](https://arxiv.org/html/2608.28619#bib.bib33)\]\.

GenAI VPs introduce a different data affordance for history taking research and education\. Because learners can ask questions in natural language and receive case\-bounded patient\-like responses, the same clinical case can produce many learner\-generated consultation paths\[[65](https://arxiv.org/html/2608.28619#bib.bib65),[10](https://arxiv.org/html/2608.28619#bib.bib10)\]\. This changes the data analysis focus from a fixed pathway or final decision to a corpus of complete learner–VP dialogues\. These logs make it possible to compare how learners working on the same clinical problem initiate inquiry, pursue cues, check uncertainty, and reorganise information over time\[[24](https://arxiv.org/html/2608.28619#bib.bib24),[23](https://arxiv.org/html/2608.28619#bib.bib23)\]\. In this sense, GenAI VPs change not only the conditions of practice, but also the conditions under which history taking processes can be studied\[[8](https://arxiv.org/html/2608.28619#bib.bib8),[30](https://arxiv.org/html/2608.28619#bib.bib30)\]\.

### 2\.3Learning analytics from dialogue records to process evidence

Dialogue records do not automatically become usable evidence for assessment or feedback\[[55](https://arxiv.org/html/2608.28619#bib.bib55)\]\. A coded learner turn does not have a fixed educational meaning in isolation\. Its meaning depends partly on the utterances immediately before and after it within the same consultation\[[45](https://arxiv.org/html/2608.28619#bib.bib45)\]\. For example, a symptom\-specific question may introduce a new line of clinically relevant line of questioning, clarify an ambiguous patient answer, or repeat information that the patient has already provided\[[40](https://arxiv.org/html/2608.28619#bib.bib40)\]\. The same coded behaviour can therefore carry different meanings depending on the behaviour’s position in the dialogue\[[54](https://arxiv.org/html/2608.28619#bib.bib54)\]\. Because the evidence is sequential and interactional, coding individual utterances is only a first step\. Additional analytic methods are needed to examine how coded behaviours accumulate, combine, and unfold over time\. In this sense, learning analytics provides a methodological bridge between raw dialogue records and interpretable process evidence for teachers\[[59](https://arxiv.org/html/2608.28619#bib.bib59)\]\. Learning analytics is therefore needed to transform dialogue records into interpretable process evidence rather than treating raw transcripts as self\-explanatory\[[31](https://arxiv.org/html/2608.28619#bib.bib31)\]\.

Teacher\-rated history taking rubric scores provide a necessary performance anchor because they summarise the assessed quality of the full consultation dialogue\. However, such scores do not explain the dialogue process that produced the rating\[[57](https://arxiv.org/html/2608.28619#bib.bib57)\]\. A history taking score may reflect broader coverage, better organisation, more appropriate follow\-up to patient cues, or more effective use of summaries and checks\[[20](https://arxiv.org/html/2608.28619#bib.bib20)\]\. Similar scores may also conceal different history taking behaviour paths\[[44](https://arxiv.org/html/2608.28619#bib.bib44)\]\. For example, two learners may receive similar history taking scores because both covered the required domains\. Yet one learner may first characterise the chest pain, follow up a patient cue about exertion, summarise the emerging pattern, and then check risk factors, whereas another learner may ask the same domains as a checklist and repeat information already given\. The score may be similar, but the dialogue process differs\. Linking teacher\-rated performance to coded dialogue behaviour is therefore necessary if GenAI VP logs are to support formative feedback rather than only summative classification\[[36](https://arxiv.org/html/2608.28619#bib.bib36)\]\. Interaction coding provides one route into this linkage\[[46](https://arxiv.org/html/2608.28619#bib.bib46)\]\. Consultation frameworks such as the Calgary\-Cambridge Guide\[[29](https://arxiv.org/html/2608.28619#bib.bib29)\], the Roter Interaction Analysis System\[[46](https://arxiv.org/html/2608.28619#bib.bib46)\], classify communicative and structural features of medical encounters\. In addition, history taking assessment work has identified observable indicators of clinical reasoning, including recognising relevant information, specifying symptoms, asking pathophysiologically oriented questions, putting questions in a logical order, checking with the patient, and summarising\[[19](https://arxiv.org/html/2608.28619#bib.bib19),[16](https://arxiv.org/html/2608.28619#bib.bib16)\]

The unresolved problem is how to connect coded dialogue behaviours with assessed consultation quality without losing the sequential character of the consultation\. A count\-based view can identify whether behaviours such as symptom specification, checking, logical organisation, or summarising appear more often in consultations that teachers judge to be of higher quality\. This is educationally useful because assessment and communication frameworks treat these behaviours as relevant indicators of history taking quality and clinical reasoning during the encounter\[[29](https://arxiv.org/html/2608.28619#bib.bib29),[46](https://arxiv.org/html/2608.28619#bib.bib46),[19](https://arxiv.org/html/2608.28619#bib.bib19),[16](https://arxiv.org/html/2608.28619#bib.bib16)\]\. However, counts alone do not show how those behaviours are positioned around other learner moves, and prior learning analytics research has shown that process evidence requires attention to how actions are situated in time rather than only how often they occur\[[45](https://arxiv.org/html/2608.28619#bib.bib45),[36](https://arxiv.org/html/2608.28619#bib.bib36),[55](https://arxiv.org/html/2608.28619#bib.bib55)\]\. A routine question placed near checking, summarising, or logical organisation may contribute to a coherent clinical account, whereas a routine question placed mainly near further routine questions may reflect checklist\-like coverage\. Local coordination also does not show temporal direction\. A summary followed by checking or mechanism\-oriented questioning suggests a different consultation process from a summary followed by repeated questioning, even though both cases contain the same summary code\. ENA and sequence\-oriented learning analytics provide ways to examine these differences because they model, respectively, local connections among coded behaviours and the order in which coded behaviours unfold\[[51](https://arxiv.org/html/2608.28619#bib.bib51),[12](https://arxiv.org/html/2608.28619#bib.bib12),[48](https://arxiv.org/html/2608.28619#bib.bib48),[47](https://arxiv.org/html/2608.28619#bib.bib47),[56](https://arxiv.org/html/2608.28619#bib.bib56)\]\. These limitations create three related gaps for GenAI VP history taking research: limited evidence about which coded behaviours are associated with teacher\-rated consultation quality, limited evidence about how those behaviours are locally coordinated within short episodes of a consultation, and limited evidence about how one coded behaviour leads into the next step of the consultation\. Comparing high\- and low\-rated consultations is therefore useful when it is treated as a performance\-anchored contrast rather than as a claim about stable learner ability\. Such a contrast can show whether consultations judged under the same case and scoring context differ in behavioural prevalence, local co\-occurrence, and sequential transition\[[57](https://arxiv.org/html/2608.28619#bib.bib57),[20](https://arxiv.org/html/2608.28619#bib.bib20),[31](https://arxiv.org/html/2608.28619#bib.bib31)\]\.

### 2\.4Research Questions

Building on these gaps, the study used teacher\-rated history taking rubric scores as a performance anchor for analysing GenAI VP dialogue processes\. High\- and low\-rated consultations were compared within each weekly case, not to classify learners as generally strong or weak, but to examine whether consultations judged under the same case and scoring context differed in observable history taking processes\. This study asks the three following research questions:

- •RQ1 \(Behavioural prevalence\):Which coded history taking behaviours distinguish high\- and low\-rated GenAI VP history taking consultations?
- •RQ2 \(Local co\-occurrence\):How are coded history taking behaviours combined in nearby turns in high\- and low\-rated consultations?
- •RQ3 \(Sequential transitions\):How do high\- and low\-rated consultations differ in transitions from one coded history taking behaviour to another?

## 3Methods

### 3\.1Participants

A total of 210 second\-year medical students from an \[anonymised\] medical college in China participated from April to May 2024\. At the time of the study, these students were enrolled in theSymptomatology and History Takingcourse and had completed core instruction on foundational history taking knowledge \(e\.g\., key history domains, basic interviewing approaches, and selected symptom\-focused content\)\.

Analyses were conducted separately for each week using all valid dialogue logs available for the weekly task\. After integrity checks at the task level \(see Data preparation\), the analytic samples were: W1n=206n=206, W2n=207n=207, W3n=205n=205, W4n=206n=206, and W5n=206n=206\. In total, 197 students had complete dialogue logs for all five weeks \(W1 to W5\)\.

### 3\.2Study design and learning environment

#### 3\.2\.1Ethics

The study was approved by the institutional ethics committee at the participating medical college \(specific approval details anonymised for double\-blind peer review\)\. Participants received written information about study objectives, procedures, and participant rights, and provided informed consent prior to participation\. Participants were informed that they could withdraw at any time; upon withdrawal, their data would be removed from the study records\.

#### 3\.2\.2Learning environment and tasks

All tasks were completed on \[Anonymous system\], a Moodle integrated learning platform that hosted instructional resources \(medical knowledge, diagnostic principles, disease reasoning guidance, and evaluation criteria\) and integrated tools supporting consultation and submission workflows\[[3](https://arxiv.org/html/2608.28619#bib.bib3)\]\. The GenAI VP was implemented as a chatbot in \[Anonymous system\]\. The chatbot used a case\-specific prompt designed for GPT\-3\.5 that specified the patient role, the clinical history of the case, and response constraints intended to keep the virtual patient in role during the consultation\. The platform recorded the full learner\-VP dialogue for subsequent process analysis\. The five weekly tasks were chest pain cases representing spontaneous pneumothorax, stable angina, aortic dissection, acute pulmonary embolism, and acute pericarditis\. In each task, learners conducted a history taking encounter with the GenAI VP, could take notes during the encounter, and submitted a diagnostic conclusion at the end of the task\.

#### 3\.2\.3Task procedure

The procedure relevant to the present study comprised an orientation and training task followed by five consecutive history taking tasks\. Before the history taking tasks, participants reviewed task instructions and rubrics and watched an instructional demonstration video\. Each task followed the same structure: participants reviewed instructions and criteria, conducted the history taking with the GenAI VP while optionally taking notes and consulting instructional resources, and then drafted and submitted a diagnostic conclusion\. The analyses reported in this paper used the history taking dialogue logs and weekly teacher\-rated history taking scores\.

#### 3\.2\.4History taking performance scores

The performance variable used for grouping was the weekly history taking score assigned to each consultation dialogue\. The score was based on the full consultation dialogue rather than the submitted diagnosis alone\. The scoring rubric was developed using the Kalamazoo Essential Elements checklist for medical encounters\[[34](https://arxiv.org/html/2608.28619#bib.bib34)\], national medical licensing examination criteria for history taking, and course teaching requirements\. The rubric covered history information, including chief complaint and history of present illness, as well as communication techniques, such as use of medical terminology and continuity of questioning\[[37](https://arxiv.org/html/2608.28619#bib.bib37)\]\. It recorded two components: a history taking content score \(HT\-Content\) and a history taking technique score \(HT\-Technique\), which together formed the total history taking score \(HT\-Total\)\.

To examine scoring consistency, 20 participants were randomly selected and their five history taking exercises were scored independently by three raters: two course instructors and one fifth\-year medical student\. The raters worked independently and were blinded to one another’s ratings\. Because the rubric contained multiple ordinal or score\-based items, inter\-rater consistency was summarised with Cronbach’s alpha for the rubric components and total score rather than reporting every item separately\. For each weekly task, Cronbach’s alpha was calculated across the three raters’ scores for HT\-Content, HT\-Technique, and HT\-Total\. Across the five weekly tasks, the total history taking score showed high consistency \(α=\.926\\alpha=\.926–\.966\.966\); the content component ranged fromα=\.887\\alpha=\.887to\.964\.964, and the technique component ranged fromα=\.753\\alpha=\.753to\.929\.929\. The final diagnostic conclusion was treated separately as correct or incorrect and was not used to define the high and low performance groups in this paper\.

### 3\.3Dialogue dataset preparation

To protect privacy, identifiable information was removed and replaced with de\-identified user IDs\. Names, contact details, and other direct identifiers were excluded from the analytic dataset\. Data integrity checks were performed at the task level to ensure that each included record contained a complete consultation dialogue log required for process analysis\. Because analyses were conducted separately for each week, inclusion was determined week by week: a learner contributes to weekttif their dialogue log for weekttmeets completeness criteria\. The resulting analytic sample sizes varied slightly across weeks \(W1n=206n=206, W2n=207n=207, W3n=205n=205, W4n=206n=206, W5n=206n=206\)\. A subset of 197 learners had complete dialogue logs for all five weeks\.

To transform raw dialogue into interpretable indicators of clinical reasoning and history taking processes on the learner side, learner utterances were represented using the behavioural coding scheme developed and validated in prior work on GenAI VP history taking dialogues\[[2](https://arxiv.org/html/2608.28619#bib.bib2)\]\. The prior coding study developed a 12\-code scheme across three behavioural dimensions and established inter\-rater reliability through iterative calibration, with Cohen’sκ\\kappacalculated on independent pre\-discussion ratings and all codes reaching the predefined reliability threshold\. The coding scheme operationalised task\-relevant learner dialogue behaviours, such as questioning driven by hypotheses, follow\-up to patient cues, clinically grounded organisation, and integrative synthesis\. In the present study, we used the finalised coded dialogue dataset as the basis for process analysis\. The coding unit was the learner utterance\. Patient responses generated by the GenAI VP were retained as the interactional context for learner questioning but were not treated as behavioural states in the data analyses\. Table[1](https://arxiv.org/html/2608.28619#S3.T1)summarises the codebook used for the present analysis\.

Table 1:Codebook for identifying learners’ behaviours in history taking dialogues\.DimensionCode Name \(Abbreviation\)Operational definition \(brief\)ClinicalReasoningPathophysiological Question \(PQ\)Hypothesis\-driven or mechanism\-oriented questioning with clear diagnostic intent\.Relevant Response \(RR\)Follow\-up that recognises and probes diagnostically significant cues introduced by the patient\.Summarising & Integrating \(SI\)Synthesis and restructuring of collected information into a coherent summary that supports reasoning\.Logical Organisation \(LO\)Clinically coherent organisation of inquiry based on relevance to the chief complaint and differential value\.InformationGatheringSpecifying Symptoms \(SS\)Systematic probing of symptom characteristics \(e\.g\., onset, duration, severity, triggers\)\.Routine Question \(RQ\)Standard history taking questions \(e\.g\., basic history domains, routine openers\) resembling interview checklists\.Summarising & Restating \(SR\)Restating collected information without restructuring or diagnostic synthesis\.Checking \(CK\)Clarifying or confirming patient\-reported information \(e\.g\., resolving ambiguity or inconsistency\)\.Repeating Question \(RT\)Redundant or repeated questioning targeting the same information point\.Fuzzy Question \(FQ\)Repeated vague or overly broad open\-ended prompts within a domain \(beyond appropriate initial openers\)\.CommunicationFacilitative Communication \(CC\)Greetings, reassurance, brief explanations, transitional phrases, and other non\-diagnostic social talk\.Off\-topic Statement \(OS\)Illogical, incomplete, or case\-irrelevant utterances that cannot be meaningfully classified elsewhere\.The codes in Table[1](https://arxiv.org/html/2608.28619#S3.T1)operationalise the learner\-side behaviours through which clinical reasoning is enacted in a history taking dialogue\. As established in Section[2\.1](https://arxiv.org/html/2608.28619#S2.SS1), history taking is not treated in this study as a checklist of questions, but as a dialogue\-based clinical reasoning process in which learners gather information, recognise patient cues, clarify ambiguity, organise what has been established, and decide how later questions should develop the consultation\[[9](https://arxiv.org/html/2608.28619#bib.bib9),[11](https://arxiv.org/html/2608.28619#bib.bib11),[13](https://arxiv.org/html/2608.28619#bib.bib13)\]\. The coding scheme therefore provided an utterance\-level representation of observable history taking behaviours, including information gathering, symptom specification, checking, cue\-responsive follow\-up, logical organisation, mechanism\-oriented questioning, communication management, and summarising\[[19](https://arxiv.org/html/2608.28619#bib.bib19),[16](https://arxiv.org/html/2608.28619#bib.bib16)\]\.

The analyses did not recode these utterances as metacognitive states\. This distinction is important because dialogue logs record what learners said in the consultation, not what they were consciously planning, monitoring, or regulating\. For that reason, all statistical analyses used the original behavioural codes and their observed frequency, co\-occurrence, and transition patterns\. When interpreting the results, we related selected patterns involving codes such as LO, SI, CK, RR, and PQ to the clinical reasoning and metacognitive functions reviewed in Section[2\.1](https://arxiv.org/html/2608.28619#S2.SS1); however, these interpretations were treated as behavioural inferences from dialogue traces rather than direct measurements of learners’ cognition\[[62](https://arxiv.org/html/2608.28619#bib.bib62),[43](https://arxiv.org/html/2608.28619#bib.bib43)\]\.

### 3\.4Analysis plan

All analyses were conducted week by week \(W1 to W5\)\. This decision was aligned with the research questions, which asked whether consultations judged as high or low in quality within the same GenAI VP case showed different behavioural prevalence, local co\-occurrence, and sequential transition patterns\. The study did not aim to estimate longitudinal growth across the five weeks\. Instead, each weekly task was treated as a separate case context because the five chest\-pain cases differed in clinical content and likely difficulty\. Analysing weeks separately therefore kept each high–low contrast anchored to learners working on the same case and being judged within the same weekly scoring context\.

Within each week, learners were split into a high\-rated and a low\-rated group using that week’s teacher\-rated total history taking score\. Learners at or above the weekly median formed the high\-rated group, and learners below the weekly median formed the low\-rated group\. Because the grouping was performed separately by week, group membership could vary across weeks\. The groups should therefore be interpreted as consultation\-level performance contrasts within a weekly case, not as stable learner\-level ability groups\. This grouping strategy allowed the same performance\-anchored comparison to be used across the three analytic layers: behavioural prevalence for RQ1, local co\-occurrence for RQ2, and sequential transition for RQ3\. Prior to analysis, the Off Topic Statement \(OS\) category was excluded from the dialogue data because it was extraneous to the research objectives\.

#### 3\.4\.1Behavioural prevalence and overall behavioural composition \(RQ1\)

RQ1 examined which coded history taking behaviours distinguish high\-rated from low\-rated encounters\. For individual behaviours, we aggregated each learner’s raw code counts within each week and compared the high\-rated and low\-rated groups using two\-sided Mann\-WhitneyUUtests, while controlling for the false discovery rate \(FDR\) across codes within each week with the Benjamini\-Hochberg \(BH\) procedure\[[4](https://arxiv.org/html/2608.28619#bib.bib4)\]\. We reported rank\-biserial correlations for the Mann\-WhitneyUUtests, calculated asrr​b=2​U/\(nhigh​nlow\)−1r\_\{rb\}=2U/\(n\_\{\\mathrm\{high\}\}n\_\{\\mathrm\{low\}\}\)\-1, where positive values indicate higher counts in the high\-rated group\. Because raw counts are sensitive to dialogue length, we also checked total learner turns and repeated the code\-level comparisons using code rates, computed as each code count divided by the learner’s total number of coded turns in that week\. In addition to testing each code separately, RQ1 examined whether the overall behavioural composition of a consultation differed by performance group\. For this analysis, each learner dialogue was represented as an 11\-dimensional behavioural profile, where each dimension was the count of one retained code\. This profile\-level analysis was included because two consultations may not differ strongly on a single code but may differ in the overall combination of routine questioning, symptom specification, checking, organisation, communication, and summarising\. We computed Bray\-Curtis dissimilarities from each learner’s raw code count vector and ran PERMANOVA with 9,999 permutations per week\[[1](https://arxiv.org/html/2608.28619#bib.bib1)\], reportingR2R^\{2\}as the profile\-level effect size\. Rate\-profile PERMANOVA used the sameR2R^\{2\}effect\-size statistic and served as a length\-control check\.

Finally, because the main RQ1 analyses used weekly median splits to create performance groups, we conducted a continuous\-score sensitivity analysis\. The purpose was to examine whether the behavioural patterns observed in the high–low contrasts were consistent with score\-related associations when the full range of weekly HT\-Total scores was retained\. For each retained code, we fitted a separate mixed\-effects regression model with continuous HT\-Total score as the outcome\. The predictor of interest was the learner’s rate for that code in that week, calculated as the code count divided by the learner’s total number of coded turns\. Total coded learner turns was included as a covariate to account for consultation length, and week was included as a fixed effect\. Learner ID was included as a random intercept to account for repeated observations from the same learner across weeks\. The model was:

H​T​T​o​t​a​li​j=β0\+β1​C​o​d​e​R​a​t​ei​j\+β2​T​o​t​a​l​T​u​r​n​si​j\+γW​e​e​kj\+ui\+ϵi​j,HTTotal\_\{ij\}=\\beta\_\{0\}\+\\beta\_\{1\}CodeRate\_\{ij\}\+\\beta\_\{2\}TotalTurns\_\{ij\}\+\\gamma\_\{Week\_\{j\}\}\+u\_\{i\}\+\\epsilon\_\{ij\},
whereiiindexes learners andjjindexes weekly consultations\. The same model was fitted separately for each of the 11 retained codes\. Model estimates, standard errors, confidence intervals, unadjustedppvalues, BH\-adjustedqqvalues, and model fit indices are reported in Supplementary Table S13\.

#### 3\.4\.2Local co\-occurrence of coded history taking behaviours \(RQ2\)

RQ2 examined whether history taking behaviours were combined differently in nearby turns in high\- and low\-rated encounters\. We used Epistemic Network Analysis \(ENA\) to model local co\-occurrence among the 11 retained utterance states after filtering out OS \(i\.e\., SS, LO, RQ, PQ, RR, SI, SR, CK, FQ, RT, and CC were included\)\[[51](https://arxiv.org/html/2608.28619#bib.bib51),[12](https://arxiv.org/html/2608.28619#bib.bib12)\]\. ENA was built separately for each week, with each learner dialogue treated as one analytic unit\. In ENA, the moving window defines the local context within which coded behaviours are treated as meaningfully connected, and this window should reflect the temporal grain of the process being studied\[[51](https://arxiv.org/html/2608.28619#bib.bib51),[12](https://arxiv.org/html/2608.28619#bib.bib12)\]\. We therefore defined local co\-occurrences using a backward moving window of four coded learner turns and accumulated these co\-occurrences into one network representation for each dialogue\. This a priori choice was intended to represent short history taking episodes rather than single next\-step dependencies or whole\-topic segments\. Immediate next\-state dependencies were analysed separately with TNA, whereas a much wider ENA window would make local co\-occurrence less distinguishable from broader topic coverage across the consultation\[[48](https://arxiv.org/html/2608.28619#bib.bib48),[49](https://arxiv.org/html/2608.28619#bib.bib49)\]\. As a sensitivity check, we repeated the ENA projection and edge\-difference analyses with backward windows of three, five, and six turns\. High\- and low\-rated encounters were first compared visually using weekly ENA difference networks\. For statistical comparisons, each dialogue’s score on the first ENA projection dimension \(MR1\) was compared between groups using a two\-sided Mann\-WhitneyUUtest\. Effect sizes for the MR1 comparisons were reported as rank\-basedrr, calculated as the standardised Mann\-Whitney test statistic divided byN\\sqrt\{N\}\.

#### 3\.4\.3Sequential transitions between coded history taking behaviours \(RQ3\)

RQ3 examined how high\- and low\-rated encounters moved from one history taking behaviour to another\. We used Transition Network Analysis \(TNA\) to model each dialogue as a first\-order sequence over the 11 retained learner behavioural states and estimated conditional transition probabilities separately for each week and group\[[49](https://arxiv.org/html/2608.28619#bib.bib49)\]\. For each transition, the effect estimate was the raw conditional probability difference,Δ​p=P​\(high\-rated\)−P​\(low\-rated\)\\Delta p=P\(\\text\{high\-rated\}\)\-P\(\\text\{low\-rated\}\)\. This value can be interpreted as an unstandardised transition\-level effect size metric in probability\-point units; positive values indicate transitions more likely in high\-rated encounters, and negative values indicate transitions more likely in low\-rated encounters\. Edge\-level permutation tests evaluated whether the observedΔ​p\\Delta pfor each transition was larger than expected under permuted group labels\. The interpretation focused on transitions that were statistically supported and educationally interpretable in relation to the process functions discussed in Section[2\.1](https://arxiv.org/html/2608.28619#S2.SS1): information gathering and symptom specification \(RQ, SS\), logical organisation \(LO\), cue\-responsive follow\-up and mechanism\-oriented questioning \(RR, PQ\), checking and clarification \(CK\), summarising or restating information \(SI, SR\), facilitative communication \(CC\), and less productive questioning patterns such as repeated or vague questioning \(RT, FQ\)\. Full transition\-level results for all tested source–target pairs are reported in Supplementary Table S11\.

Table 2:Descriptive statistics for total coded learner turns by week and performance group\.

## 4Results

We report results in the order of the three research questions: behavioural prevalence and overall behavioural composition \(RQ1\), local co\-occurrence structure \(RQ2\), and sequential transitions \(RQ3\)\. Because performance groups were defined separately within each week, the results are interpreted as within\-week contrasts\.

### 4\.1Behavioural prevalence and overall behavioural composition \(RQ1\)

RQ1 examined consultation length, individual coded history taking behaviours, and overall behavioural composition\. Table[2](https://arxiv.org/html/2608.28619#S3.T2)reports total coded learner turns by week and performance group\. Descriptive statistics for the 11 retained code counts are reported in Supplementary Table S3, and descriptive statistics for code rates are reported in Supplementary Table S4\.

High\-rated consultations contained more coded learner turns than low\-rated consultations in every week\. The descriptive gap was largest in W1 and smallest in W2\. Full consultation\-length tests, including Mann–WhitneyUU, unadjustedppvalues, BH\-adjustedqqvalues, and rank\-biserial correlations, are reported in Supplementary Table S2\. Mann–WhitneyUUtests are reported with group sizes rather than degrees of freedom\.

Table 3:Summary of count\-based Mann–WhitneyUUtest results across weeks after BH within each week across codes\. Reported are the number of weeks with significant group differences \(q<\.05q<\.05\), the significant weeks, the mean BH\-adjustedqqvalue, and the mean effect size as rank\-biserial correlation \(rr​br\_\{rb\}\) across those significant weeks\. Full descriptive statistics are reported in Supplementary Tables S3–S5\.The raw\-count comparisons showed higher counts for several coded behaviours in high\-rated consultations \(Table[3](https://arxiv.org/html/2608.28619#S4.T3)\)\. SI and RQ differed in all five weeks\. LO, CC, and SS each differed in four weeks\. FQ, CK, and RR showed more limited differences, and RT, SR, and PQ did not show reliable raw\-count differences after BH correction\. Across significant contrasts, mean rank\-biserial correlations ranged from 0\.164 for RR to 0\.424 for LO\.

After normalising counts by total coded learner turns, fewer differences remained\. Significant rate\-based contrasts were found for LO in W3 and W5, SI in W1 and W3, CK in W1, and RQ in W3\. The rate\-based results therefore narrow the raw\-count interpretation: high\-rated consultations were longer and contained more coded behaviours, but proportional differences were concentrated in a smaller set of behaviours, especially organisation, summarising/integrating, and checking\.

The overall behavioural composition also differed between groups\. Raw\-count PERMANOVA was significant in all five weeks, with effect sizes ranging fromR2=\.0225R^\{2\}=\.0225in W2 toR2=\.1306R^\{2\}=\.1306in W1\. Rate\-profile PERMANOVA showed significant composition differences in W1 \(R2=\.012R^\{2\}=\.012,p=\.041p=\.041\) and W4 \(R2=\.013R^\{2\}=\.013,p=\.016p=\.016\), while W2, W3, and W5 were not significant and had small effect sizes \(R2=\.006R^\{2\}=\.006–\.009\.009\)\. Complete raw\-count and rate\-profile PERMANOVA outputs, including degrees of freedom, sums of squares, mean squares, pseudo\-FF, permutationppvalues,R2R^\{2\}, and number of permutations, are reported in Supplementary Tables S6 and S7\.

The continuous\-score sensitivity analysis partly supported the median\-split findings\. In separate mixed\-effects models using continuous HT\-Total as the outcome and controlling for week and total coded learner turns, LO \(β=1\.076\\beta=1\.076,p<\.001p<\.001\), SI \(β=1\.269\\beta=1\.269,p<\.001p<\.001\), and CC \(β=\.616\\beta=\.616,p=\.016p=\.016\) were positively associated with HT\-Total\. SS was negatively associated with HT\-Total \(β=−\.615\\beta=\-\.615,p=\.024p=\.024\), and RQ was not statistically reliable\. Full model estimates, standard errors, confidence intervals, unadjustedppvalues, BH\-adjustedqqvalues, and model fit indices are reported in Supplementary Table S13\. The SS result requires particular caution\. SS appeared more often in high\-rated consultations in raw counts, but its rate was negatively associated with continuous HT\-Total after controlling for week and total coded learner turns\. This suggests a length/composition distinction: high\-rated consultations may include more symptom\-specific questions in absolute terms because they contain more learner turns overall, but a higher proportion of SS may indicate that the consultation remains focused on symptom detailing rather than moving toward organisation, checking, or synthesis\.

![Refer to caption](https://arxiv.org/html/2608.28619v1/images/ena_diff_original_nopoints_221.png)Figure 1:Weekly ENA difference networks \(high\-rated minus low\-rated\)\. Nodes are coded learner behaviours from Table[1](https://arxiv.org/html/2608.28619#S3.T1)\. Blue edges are stronger in the high\-rated group, red edges are stronger in the low\-rated group, and thicker edges indicate larger group differences\. Layout and edge scaling are fixed across weeks; individual dialogue points and centroid confidence boxes are omitted to focus on co\-occurrence contrasts\.
### 4\.2Performance\-related differences in co\-occurrence structure \(RQ2\)

RQ2 examined how coded history taking behaviours were placed near one another\. Figure[1](https://arxiv.org/html/2608.28619#S4.F1)maps local pairings by week: blue edges mark pairings stronger in high\-rated consultation sessions, red edges mark pairings stronger in low\-rated consultation sessions, and thicker edges indicate larger group differences\. The high\-rated group scored significantly higher on the first ENA dimension \(MR1\) across all five weeks, indicating consistent group separation in local coordination\. Mann\-WhitneyUUtests do not have degrees of freedom, so the tests are reported with weekly group sizes and rank\-based effect sizes: W1,nhigh=109n\_\{\\mathrm\{high\}\}=109,nlow=97n\_\{\\mathrm\{low\}\}=97,U=6847\.0U=6847\.0,p<\.001p<\.001,r=\.255r=\.255; W2,nhigh=106n\_\{\\mathrm\{high\}\}=106,nlow=101n\_\{\\mathrm\{low\}\}=101,U=6526\.0U=6526\.0,p=\.006p=\.006,r=\.189r=\.189; W3,nhigh=107n\_\{\\mathrm\{high\}\}=107,nlow=98n\_\{\\mathrm\{low\}\}=98,U=6765\.0U=6765\.0,p<\.001p<\.001,r=\.251r=\.251; W4,nhigh=104n\_\{\\mathrm\{high\}\}=104,nlow=101n\_\{\\mathrm\{low\}\}=101,U=6645\.0U=6645\.0,p=\.001p=\.001,r=\.229r=\.229; and W5,nhigh=105n\_\{\\mathrm\{high\}\}=105,nlow=101n\_\{\\mathrm\{low\}\}=101,U=7057\.0U=7057\.0,p<\.001p<\.001,r=\.286r=\.286\. The window\-size sensitivity check showed the same MR1 group separation for windows of three, five, and six turns in every week \(allp<\.007p<\.007\)\. Edge\-difference vectors from the alternative windows were also strongly aligned with the four\-turn solution, with mean Pearson correlations of \.970, \.980, and \.952 for windows of three, five, and six turns, respectively\. The clearest pattern concerned the placement of RQ\. In W1, W2, and W4, the RQ\-CC edge was stronger in high\-rated consultation sessions\. This means that routine inquiry was more often placed near moves that oriented the patient, signalled a topic shift, or maintained the interaction, rather than appearing only as a checklist\-like progression\. In W3, RQ was more tightly connected with LO and SI, indicating that routine inquiry was locally tied to organising the case and pulling information together\. In W4, stronger edges around SI again placed synthesis close to ongoing inquiry\. SS showed a more case\-sensitive pattern\. The SS\-RQ pairing was stronger in low\-rated consultation sessions in W1 and W4, where local episodes were more dominated by symptom probing and routine inquiry, but it was stronger in high\-rated consultations in W5\. This suggests that probing symptom detail was not uniformly associated with high performance; its meaning depended on the surrounding moves and the weekly case\. Across the five tasks, the co\-occurrence evidence shows that high\-rated consultations differed most in how RQ, CC, LO, and SI were locally connected, while SS became useful when it was embedded in a broader inquiry structure rather than simply paired mainly with further routine questioning\.

### 4\.3Temporal sequencing and transition differences \(RQ3\)

RQ3 examined which history taking behaviour usually followed a given history taking behaviour\. The analysis separates the shared sequence structure of the weekly consultations from the points where high\- and low\-rated consultations diverged within the shared sequence structure\. Figure[2](https://arxiv.org/html/2608.28619#S4.F2)shows the full\-cohort transition routine for each week, and Figure[3](https://arxiv.org/html/2608.28619#S4.F3)marks transition contrasts within the weekly transition routine\. PositiveΔ​p\\Delta pvalues indicate transitions more likely in high\-rated encounters; negative values indicate transitions more likely in low\-rated encounters\. The reportedΔ​p\\Delta pvalues are raw conditional\-probability difference estimates and can be read as unstandardised transition\-level effect\-size metrics\.

![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w1_global_model_filtered.png)\(a\)W1
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w2_global_model_filtered.png)\(b\)W2
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w3_global_model_filtered.png)\(c\)W3
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w4_global_model_filtered.png)\(d\)W4
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w5_global_model_filtered.png)\(e\)W5

Figure 2:Full\-cohort weekly transition networks\. Nodes are coded learner behaviours from Table[1](https://arxiv.org/html/2608.28619#S3.T1); directed edges show common next\-state transitions in the full cohort\. Self loops indicate repeated use of the same behavioural state\.![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w1_permutation_high_low.png)\(a\)W1
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w2_permutation_high_low.png)\(b\)W2
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w3_permutation_high_low.png)\(c\)W3
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w4_permutation_high_low.png)\(d\)W4
![Refer to caption](https://arxiv.org/html/2608.28619v1/images/w5_permutation_high_low.png)\(e\)W5

Figure 3:Transition differences by week based on edge\-level permutation tests \(p<\.05p<\.05\)\. Nodes represent coded learner behaviours from Table[1](https://arxiv.org/html/2608.28619#S3.T1)\. Directed edges are labelled byΔ​p=P​\(high\)−P​\(low\)\\Delta p=P\(\\text\{high\}\)\-P\(\\text\{low\}\), the raw conditional\-probability difference estimate\. Green edges indicate transitions more likely in the high\-rated group, red edges indicate transitions more likely in the low\-rated group, and thicker edges indicate larger absoluteΔ​p\\Delta pvalues\.The full\-cohort networks show that all five weekly tasks shared a similar transition routine\. The most visible structure was repeated movement within RQ and SS, with additional repeated use of LO and CC\. Most learners, regardless of performance group, spent substantial portions of the encounter in routine inquiry, symptom detail work, organisation, or managing the conversation\.

The group contrasts were differences inside a shared routine rather than completely different sequence structures\. CC appeared in different parts of the shared routine\. In W2, high\-rated encounters more often moved from LO to CC \(LO→\\rightarrowCC; effect estimateΔ​p=\.039\\Delta p=\.039,p=\.002p=\.002\), whereas low\-rated consultations more often moved from RQ to CC \(RQ→\\rightarrowCC; effect estimateΔ​p=−\.018\\Delta p=\-\.018,p=\.008p=\.008\)\. In W5, high\-rated consultations more often moved from RR to CC \(RR→\\rightarrowCC; effect estimateΔ​p=\.079\\Delta p=\.079,p=\.001p=\.001\), whereas low\-rated consultations more often moved from CC back to RQ \(CC→\\rightarrowRQ; effect estimateΔ​p=−\.092\\Delta p=\-\.092,p=\.006p=\.006\)\. CC tended to follow organising the case or following up patient cues in high\-rated consultations, but it more often returned the learner to routine questioning in low\-rated consultations\. The second difference concerns what happened after SI\. In W3, high\-rated consultations more often moved from SI to PQ \(SI→\\rightarrowPQ; effect estimateΔ​p=\.045\\Delta p=\.045,p=\.001p=\.001\), and in W4 they more often moved from SI to CK \(SI→\\rightarrowCK; effect estimateΔ​p=\.030\\Delta p=\.030,p=\.004p=\.004\)\. By contrast, W1 low\-rated consultations more often moved from SI to RT \(SI→\\rightarrowRT; effect estimateΔ​p=−\.050\\Delta p=\-\.050,p=\.001p=\.001\)\. The same summary move could lead into mechanism\-focused questioning or confirming details, or it could be followed by repeated questions\. These transition patterns suggest that the same behaviour, especially SI or CC, carried different instructional meanings depending on what it led to next\.

## 5Discussion

### 5\.1Performance\-related process evidence in GenAI VP history taking

Section[2\.3](https://arxiv.org/html/2608.28619#S2.SS3)identified a problem in using GenAI VP dialogue logs for clinical reasoning education: teacher\-rated scores can anchor consultation quality, but they do not explain how the consultation unfolded\. The RQ1 results addressed the first part of this problem by linking teacher\-rated consultation quality to coded learner behaviours \(Section[4\.1](https://arxiv.org/html/2608.28619#S4.SS1)\)\. The clearest difference was behavioural volume\. High\-rated consultations contained more coded learner turns in every weekly case and higher raw counts of several behaviours\. This finding should not be dismissed as a nuisance effect, because longer consultations may give learners more opportunity to gather information, organise the case, check uncertainty, and summarise\. At the same time, the length\-adjusted analyses narrowed the interpretation\. Rate\-based comparisons were more selective, rate\-profile PERMANOVA effects were small, and the continuous\-score sensitivity analysis identified LO, SI, and CC as the clearest positive score\-related behaviours\. The contribution of RQ1 is therefore not a claim that performance differences were independent of consultation length\. Rather, RQ1 shows that raw behavioural volume was the most stable group difference, while length\-adjusted and continuous\-score analyses helped identify which behaviours remained most closely associated with assessed history\-taking quality\.

This distinction is especially important for interpreting symptom\-specific questioning\. In the raw\-count analyses, SS appeared more often in high\-rated consultations\. However, in the continuous\-score sensitivity analysis, SS was negatively associated with HT\-Total after controlling for week and total coded learner turns\. This is not merely an unstable result; it is a length–composition reversal\. High\-rated consultations may contain more symptom\-specific questions in absolute terms because they contain more learner turns overall, but a higher proportion of SS may indicate that the consultation remains concentrated on symptom detailing rather than moving toward organisation, checking, or synthesis\. SS should therefore not be interpreted as uniformly beneficial\. Its educational meaning depends on how symptom details are used in the wider consultation\.

The RQ2 results explain why the raw\-count results should not be read as a simple volume effect\. ENA showed that high\-rated consultations more often connected routine questioning and symptom\-specific questioning with communication, organisation, checking, and summarising \(Section[4\.2](https://arxiv.org/html/2608.28619#S4.SS2)\)\. This matters because RQ and SS are not inherently strong or weak behaviours\. A routine question placed near LO, CK, or SI can contribute to an organised clinical account; a routine question placed mainly near further routine questioning can remain checklist\-like\. This finding extends communication and history taking assessment frameworks by showing that the educational meaning of a coded behaviour depends partly on its local coordination with other behaviours\[[29](https://arxiv.org/html/2608.28619#bib.bib29),[46](https://arxiv.org/html/2608.28619#bib.bib46)\]\.

The RQ3 results added temporal direction to this interpretation\. TNA showed that SI had different meanings depending on what followed it \(Section[4\.3](https://arxiv.org/html/2608.28619#S4.SS3)\)\. SI followed by CK or PQ suggested that summarising helped the learner verify uncertainty or pursue mechanism\-oriented follow\-up\. SI followed by RT suggested that the learner restated information without using the summary to guide the next question\. CC also depended on sequence position: CC after LO or RR suggested that communication supported the flow of an organised consultation, whereas CC followed by RQ suggested a return to general information gathering\. These findings show why sequence evidence is needed in addition to prevalence and co\-occurrence evidence: temporal order helps distinguish behaviours that are merely present from behaviours that shape the next step of the consultation\[[47](https://arxiv.org/html/2608.28619#bib.bib47),[49](https://arxiv.org/html/2608.28619#bib.bib49)\]\.

The main contribution is therefore not that high\-rated learners used a separate set of behaviours\. Rather, high\-rated consultations showed stronger coordination among ordinary history taking behaviours\. This interpretation is consistent with a metacognition\-informed view of history taking, in which learners monitor what has been established, identify uncertainty, organise information, and decide how later questions should develop the consultation\[[62](https://arxiv.org/html/2608.28619#bib.bib62),[43](https://arxiv.org/html/2608.28619#bib.bib43)\]\. The interpretation remains cautious because the data record learner utterances, not learners’ conscious planning or monitoring\. A turn coded as LO, SI, CK, RR, or PQ is consistent with organising, synthesising, monitoring, cue response, or hypothesis testing, but it does not directly measure those cognitive processes\.

The patterns observed here are not specific to GenAI VPs\. Structuring the consultation, checking patient information, summarising, and following clinically relevant cues are also emphasised in clinical communication and history taking assessment frameworks\[[29](https://arxiv.org/html/2608.28619#bib.bib29),[46](https://arxiv.org/html/2608.28619#bib.bib46),[19](https://arxiv.org/html/2608.28619#bib.bib19),[16](https://arxiv.org/html/2608.28619#bib.bib16)\]\. The same analytic approach could be applied to SP or OSCE encounters if those encounters were recorded, transcribed, segmented into learner turns, and coded with the same behavioural scheme\. Prevalence analysis could compare how often learners used behaviours such as CK, SI, or SS; ENA could examine whether these behaviours were locally connected with LO, CC, or RQ; and TNA could examine whether summaries led to checking, mechanism\-oriented questioning, or repetition\. The practical difference is that GenAI VP systems already store complete turn\-by\-turn logs during routine practice, whereas SP and OSCE settings require additional recording, transcription, segmentation, and coding\.

### 5\.2Instructional uses of performance\-linked dialogue patterns

The results suggest three uses for instruction and system design\. First, teacher\-facing reports could help instructors locate transcript segments worth reviewing\. For example, a report could flag a long stretch of RQ or SS that is not followed by LO, CK, or SI\. The teacher could then discuss with the learner what information had already been established, what remained uncertain, and how the next question could have been guided by the case information already collected\.

Second, GenAI VP systems could use sequence patterns to support adaptive practice\. The transition results suggest that SI is a useful point for intervention because it can either redirect the consultation or fail to change the subsequent questioning\. If a learner summarises and then repeats an already answered question, the system could prompt the learner to clarify an uncertainty or ask a mechanism\-oriented follow\-up\. If a learner summarises and then moves to CK or PQ, the system could reinforce the use of summaries as a bridge to verification or reasoning\-oriented questioning\.

Third, aggregated dialogue patterns could inform curriculum review\. If many learners ask symptom\-specific questions but rarely connect them to checking, organisation, or summarising, the issue may not be lack of symptom coverage\. It may be difficulty using symptom information to structure the consultation\. Similarly, if CC often leads back to routine questioning, learners may need more guidance on using communication to maintain consultation flow rather than restarting general data gathering\.

These uses should not be treated as ready\-made scoring rules\. The present study identified performance\-linked dialogue patterns; it did not test whether reports, prompts, or teacher review based on these patterns improve later history taking\. Future studies should examine whether teachers find the patterns interpretable, whether learners can act on them, and whether instruction based on these patterns improves later GenAI VP, SP, or OSCE performance\.

### 5\.3Limitations and future work

Several design features shape the interpretation of the findings\. First, the performance anchor was the course\-based history taking score\. This anchor fits the teaching context, but future studies should test whether the same process signatures align with OSCE performance, expert panel ratings, diagnostic reasoning assessments, or later clinical interview performance\[[38](https://arxiv.org/html/2608.28619#bib.bib38)\]\. The present analyses also focused on weekly performance contrasts rather than individual growth; longitudinal models are still needed to examine whether the same learners develop stronger coordination and transition patterns across repeated GenAI VP encounters\[[39](https://arxiv.org/html/2608.28619#bib.bib39)\]\.

Second, the coded sequence focused on learner utterances\. GenAI VP responses were retained as interactional context but were not coded as behavioural states\. This learner\-only focus limits claims about the full learner–GenAI VP interaction, because a learner’s next move may depend on the VP’s cue, wording, ambiguity, or response quality\[[25](https://arxiv.org/html/2608.28619#bib.bib25)\]\. Coding the VP side would help distinguish learner uncertainty, missed patient cues, and limitations in the VP response\. The study was also conducted in one course, one GenAI VP environment, and five chest pain cases generated with GPT\-3\.5\. Replication with newer GenAI models, other symptoms, other institutions, and different VP designs is needed to determine which process signatures are stable features of coherent inquiry and which are case\-, system\-, or curriculum\-specific\.

Finally, dialogue logs were the sole source of process evidence\. Without think\-aloud protocols, retrospective interviews, eye\-tracking, or other cognitive measures, the metacognitive interpretation remains an inference from behaviour rather than direct observation\. Future work should combine GenAI VP logs with complementary process data and validated GenAI\-assisted coding pipelines\. Such work can test whether process\-signature feedback is interpretable to teachers, actionable for learners, and effective in improving later history taking performance\[[58](https://arxiv.org/html/2608.28619#bib.bib58)\]\.

## 6Conclusion

The present study analysed GenAI VP history taking dialogues to examine whether open\-ended consultation logs can be transformed into interpretable process evidence\. By comparing high\- and low\-rated consultations through behavioural prevalence, ENA, and TNA, the study showed that stronger performance was not explained only by asking more questions or producing longer dialogues\. Rather, high\-rated consultations showed how learners wove basic history questions, symptom exploration, conversation management, clarification, case structuring, and interim synthesis into an organised clinical account, then used that account to verify details or pursue reasoning\-oriented follow\-up\. These findings suggest that GenAI VP logs can make learners’ history taking processes visible at scale, offering learning analytics a way to connect teacher\-rated performance with concrete behavioural patterns that can inform future feedback and support for clinical reasoning practice\.

### Abbreviations

## Declarations

Identifying details \(author names, affiliations, funding sources, ethics approval numbers, and author contributions\) are anonymised in the submitted manuscript for double\-blind peer review and will be provided on the title page and in the final manuscript\.

### Acknowledgements

The authors thank the participating students and instructors for their contributions to the study\.

### Funding

Blinded for review\.

### Conflict of interest

The authors declare that they have no competing interests\.

### Ethics approval and consent to participate

The study was approved by the institutional ethics committee at the participating medical college \(specific approval details anonymised for double\-blind peer review and provided on the title page\)\. All participants received written information about study objectives, procedures, and participant rights, and provided written informed consent prior to participation\.

### Data Availability

The datasets used and analysed during the current study are available from the corresponding author on reasonable request\.

### Authors Contribution

Blinded for review\.

## References

- \\bibcommenthead
- Anderson \[\\APACyear2001\]\\APACinsertmetastaranderson2001new\{APACrefauthors\}Anderson, M\.J\.\\APACrefYearMonthDay2001\.\\BBOQ\\APACrefatitleA new method for non\-parametric multivariate analysis of variance A new method for non\-parametric multivariate analysis of variance\.\\BBCQ\\APACjournalVolNumPagesAustral ecology26132–46,\\PrintBackRefs\\CurrentBib
- Authors \[\\APACyear2026\\APACexlab\\BCnt1\]\\APACinsertmetastarchen2026developing\{APACrefauthors\}Authors\\APACrefYearMonthDay2026\\BCnt1\.\\BBOQ\\APACrefatitleAnonymized title Anonymized title\.\\BBCQ\\APACjournalVolNumPagesAnonymized venue,\\PrintBackRefs\\CurrentBib
- Authors \[\\APACyear2026\\APACexlab\\BCnt2\]\\APACinsertmetastarli2026flora\{APACrefauthors\}Authors\\APACrefYearMonthDay2026\\BCnt2\.\\BBOQ\\APACrefatitleAnonymized title Anonymized title\.\\BBCQ\\APACjournalVolNumPagesAnonymized venue,\\PrintBackRefs\\CurrentBib
- Benjamini\\BBAHochberg \[\\APACyear1995\]\\APACinsertmetastarbenjamini1995controlling\{APACrefauthors\}Benjamini, Y\.\\BCBT\\BBAHochberg, Y\.\\APACrefYearMonthDay1995\.\\BBOQ\\APACrefatitleControlling the false discovery rate: a practical and powerful approach to multiple testing Controlling the false discovery rate: a practical and powerful approach to multiple testing\.\\BBCQ\\APACjournalVolNumPagesJournal of the Royal statistical society: series B \(Methodological\)571289–300,\\PrintBackRefs\\CurrentBib
- Bennett \[\\APACyear2011\]\\APACinsertmetastarbennett2011formative\{APACrefauthors\}Bennett, R\.E\.\\APACrefYearMonthDay2011\.\\BBOQ\\APACrefatitleFormative assessment: A critical review Formative assessment: A critical review\.\\BBCQ\\APACjournalVolNumPagesAssessment in education: principles, policy & practice1815–25,\\PrintBackRefs\\CurrentBib
- Bickley\\BBASzilagyi \[\\APACyear2012\]\\APACinsertmetastarbickley2012bates\{APACrefauthors\}Bickley, L\.\\BCBT\\BBASzilagyi, P\.G\.\\APACrefYear2012\.\\APACrefbtitleBates’ guide to physical examination and history\-taking Bates’ guide to physical examination and history\-taking\.\\APACaddressPublisherLippincott Williams & Wilkins\.\\PrintBackRefs\\CurrentBib
- Black\\BBAWiliam \[\\APACyear1998\]\\APACinsertmetastarblack1998assessment\{APACrefauthors\}Black, P\.\\BCBT\\BBAWiliam, D\.\\APACrefYearMonthDay1998\.\\BBOQ\\APACrefatitleAssessment and classroom learning Assessment and classroom learning\.\\BBCQ\\APACjournalVolNumPagesAssessment in Education: principles, policy & practice517–74,\\PrintBackRefs\\CurrentBib
- Bond\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarbond2024meta\{APACrefauthors\}Bond, M\., Khosravi, H\., De Laat, M\., Bergdahl, N\., Negrea, V\., Oxley, E\.\\BDBLSiemens, G\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleA meta systematic review of artificial intelligence in higher education: A call for increased ethics, collaboration, and rigour A meta systematic review of artificial intelligence in higher education: A call for increased ethics, collaboration, and rigour\.\\BBCQ\\APACjournalVolNumPagesInternational journal of educational technology in higher education2114,\\PrintBackRefs\\CurrentBib
- Bowen \[\\APACyear2006\]\\APACinsertmetastarbowen2006educational\{APACrefauthors\}Bowen, J\.L\.\\APACrefYearMonthDay2006\.\\BBOQ\\APACrefatitleEducational strategies to promote clinical diagnostic reasoning Educational strategies to promote clinical diagnostic reasoning\.\\BBCQ\\APACjournalVolNumPagesNew England Journal of Medicine355212217–2225,\\PrintBackRefs\\CurrentBib
- Brügge\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarbrugge2024large\{APACrefauthors\}Brügge, E\., Ricchizzi, S\., Arenbeck, M\., Keller, M\.N\., Schur, L\., Stummer, W\.\\BDBLDarici, D\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleLarge language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial\.\\BBCQ\\APACjournalVolNumPagesBMC medical education2411391,\\PrintBackRefs\\CurrentBib
- Charlin\\BOthers\. \[\\APACyear2007\]\\APACinsertmetastarcharlin2007scripts\{APACrefauthors\}Charlin, B\., Boshuizen, H\.P\., Custers, E\.J\.\\BCBLFeltovich, P\.J\.\\APACrefYearMonthDay2007\.\\BBOQ\\APACrefatitleScripts and clinical reasoning Scripts and clinical reasoning\.\\BBCQ\\APACjournalVolNumPagesMedical education41121178–1184,\\PrintBackRefs\\CurrentBib
- Csanadi\\BOthers\. \[\\APACyear2018\]\\APACinsertmetastarcsanadi2018coding\{APACrefauthors\}Csanadi, A\., Eagan, B\., Kollar, I\., Shaffer, D\.W\.\\BCBLFischer, F\.\\APACrefYearMonthDay2018\.\\BBOQ\\APACrefatitleWhen coding\-and\-counting is not enough: Using epistemic network analysis \(ENA\) to analyze verbal data in CSCL research When coding\-and\-counting is not enough: Using epistemic network analysis \(ena\) to analyze verbal data in cscl research\.\\BBCQ\\APACjournalVolNumPagesInternational Journal of Computer\-Supported Collaborative Learning134419–438,\\PrintBackRefs\\CurrentBib
- Cutrer\\BOthers\. \[\\APACyear2017\]\\APACinsertmetastarcutrer2017fostering\{APACrefauthors\}Cutrer, W\.B\., Miller, B\., Pusic, M\.V\., Mejicano, G\., Mangrulkar, R\.S\., Gruppen, L\.D\.\\BDBLMoore Jr, D\.E\.\\APACrefYearMonthDay2017\.\\BBOQ\\APACrefatitleFostering the development of master adaptive learners: a conceptual model to guide skill acquisition in medical education Fostering the development of master adaptive learners: a conceptual model to guide skill acquisition in medical education\.\\BBCQ\\APACjournalVolNumPagesAcademic medicine92170–75,\\PrintBackRefs\\CurrentBib
- Eva \[\\APACyear2005\]\\APACinsertmetastareva2005every\{APACrefauthors\}Eva, K\.W\.\\APACrefYearMonthDay2005\.\\BBOQ\\APACrefatitleWhat every teacher needs to know about clinical reasoning What every teacher needs to know about clinical reasoning\.\\BBCQ\\APACjournalVolNumPagesMedical education39198–106,\\PrintBackRefs\\CurrentBib
- Fąferek\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarfaferek2024integrating\{APACrefauthors\}Fąferek, J\., Cariou, P\\BHBIL\., Hege, I\., Mayer, A\., Morin, L\., Rodriguez\-Molina, D\.\\BDBLKononowicz, A\.A\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleIntegrating virtual patients into undergraduate health professions curricula: a framework synthesis of stakeholders’ opinions based on a systematic literature review Integrating virtual patients into undergraduate health professions curricula: a framework synthesis of stakeholders’ opinions based on a systematic literature review\.\\BBCQ\\APACjournalVolNumPagesBMC Medical Education241727,\\PrintBackRefs\\CurrentBib
- Fürstenberg\\BOthers\. \[\\APACyear2020\]\\APACinsertmetastarfurstenberg2020assessing\{APACrefauthors\}Fürstenberg, S\., Helm, T\., Prediger, S\., Kadmon, M\., Berberat, P\.O\.\\BCBLHarendza, S\.\\APACrefYearMonthDay2020\.\\BBOQ\\APACrefatitleAssessing clinical reasoning in undergraduate medical students during history taking with an empirically derived scale for clinical reasoning indicators Assessing clinical reasoning in undergraduate medical students during history taking with an empirically derived scale for clinical reasoning indicators\.\\BBCQ\\APACjournalVolNumPagesBMC Medical Education201368,\\PrintBackRefs\\CurrentBib
- Goldowsky\\BBARencic \[\\APACyear2023\]\\APACinsertmetastargoldowsky2023self\{APACrefauthors\}Goldowsky, A\.\\BCBT\\BBARencic, J\.\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleSelf\-regulated learning and the future of diagnostic reasoning education Self\-regulated learning and the future of diagnostic reasoning education\.\\BBCQ\\APACjournalVolNumPagesDiagnosis10124–30,\\PrintBackRefs\\CurrentBib
- Hamilton\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarhamilton2024evolution\{APACrefauthors\}Hamilton, A\., Molzahn, A\.\\BCBLMcLemore, K\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleThe evolution from standardized to virtual patients in medical education The evolution from standardized to virtual patients in medical education\.\\BBCQ\\APACjournalVolNumPagesCureus1610,\\PrintBackRefs\\CurrentBib
- Haring\\BOthers\. \[\\APACyear2017\]\\APACinsertmetastarharing2017observable\{APACrefauthors\}Haring, C\.M\., Cools, B\.M\., van Gurp, P\.J\., van der Meer, J\.W\.\\BCBLPostma, C\.T\.\\APACrefYearMonthDay2017\.\\BBOQ\\APACrefatitleObservable phenomena that reveal medical students’ clinical reasoning ability during expert assessment of their history taking: a qualitative study Observable phenomena that reveal medical students’ clinical reasoning ability during expert assessment of their history taking: a qualitative study\.\\BBCQ\\APACjournalVolNumPagesBMC Medical education171147,\\PrintBackRefs\\CurrentBib
- Haring\\BOthers\. \[\\APACyear2020\]\\APACinsertmetastarharing2020validity\{APACrefauthors\}Haring, C\.M\., Klaarwater, C\.C\., Bouwmans, G\.A\., Cools, B\.M\., van Gurp, P\.J\., van der Meer, J\.W\.\\BCBLPostma, C\.T\.\\APACrefYearMonthDay2020\.\\BBOQ\\APACrefatitleValidity, reliability and feasibility of a new observation rating tool and a post encounter rating tool for the assessment of clinical reasoning skills of medical students during their internal medicine clerkship: a pilot study Validity, reliability and feasibility of a new observation rating tool and a post encounter rating tool for the assessment of clinical reasoning skills of medical students during their internal medicine clerkship: a pilot study\.\\BBCQ\\APACjournalVolNumPagesBMC Medical Education201198,\\PrintBackRefs\\CurrentBib
- Hasnain\\BOthers\. \[\\APACyear2001\]\\APACinsertmetastarhasnain2001historytaking\{APACrefauthors\}Hasnain, M\., Bordage, G\., Connell, K\.J\.\\BCBLSinacore, J\.M\.\\APACrefYearMonthDay2001\.\\BBOQ\\APACrefatitleHistory\-taking behaviors associated with diagnostic competence of clerks: an exploratory study History\-taking behaviors associated with diagnostic competence of clerks: an exploratory study\.\\BBCQ\\APACjournalVolNumPagesAcademic Medicine7610S14–S17,\\PrintBackRefs\\CurrentBib
- Hege\\BOthers\. \[\\APACyear2018\]\\APACinsertmetastarhege2018advancing\{APACrefauthors\}Hege, I\., Kononowicz, A\.A\., Berman, N\.B\., Lenzer, B\.\\BCBLKiesewetter, J\.\\APACrefYearMonthDay2018\.\\BBOQ\\APACrefatitleAdvancing clinical reasoning in virtual patients–development and application of a conceptual framework Advancing clinical reasoning in virtual patients–development and application of a conceptual framework\.\\BBCQ\\APACjournalVolNumPagesGMS journal for medical education351Doc12,\\PrintBackRefs\\CurrentBib
- Holderried, Stegemann\-Philipps, Herrmann\-Werner\\BCBL\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarholderried2024feedback\{APACrefauthors\}Holderried, F\., Stegemann\-Philipps, C\., Herrmann\-Werner, A\., Festl\-Wietek, T\., Holderried, M\., Eickhoff, C\.\\BDBLothers\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleA language model–powered simulated patient with automated feedback for history taking: Prospective study A language model–powered simulated patient with automated feedback for history taking: Prospective study\.\\BBCQ\\APACjournalVolNumPagesJMIR Medical Education101e59213,\\PrintBackRefs\\CurrentBib
- Holderried, Stegemann\-Philipps, Herschbach\\BCBL\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarholderried2024generative\{APACrefauthors\}Holderried, F\., Stegemann\-Philipps, C\., Herschbach, L\., Moldt, J\\BHBIA\., Nevins, A\., Griewatz, J\.\\BDBLMahling, M\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleA generative pretrained transformer \(GPT\)–powered chatbot as a simulated patient to practice history taking: prospective, mixed methods study A generative pretrained transformer \(gpt\)–powered chatbot as a simulated patient to practice history taking: prospective, mixed methods study\.\\BBCQ\\APACjournalVolNumPagesJMIR medical education101e53961,\\PrintBackRefs\\CurrentBib
- Järvelä\\BOthers\. \[\\APACyear2023\]\\APACinsertmetastarjarvela2023human\{APACrefauthors\}Järvelä, S\., Nguyen, A\.\\BCBLHadwin, A\.\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleHuman and artificial intelligence collaboration for socially shared regulation in learning Human and artificial intelligence collaboration for socially shared regulation in learning\.\\BBCQ\\APACjournalVolNumPagesBritish Journal of Educational Technology5451057–1076,\{APACrefDOI\}[https://doi\.org/10\.1111/bjet\.13325](https://doi.org/10.1111/bjet.13325)\\PrintBackRefs\\CurrentBib
- Jay\\BOthers\. \[\\APACyear2025\]\\APACinsertmetastarjay2025use\{APACrefauthors\}Jay, R\., Sandars, J\., Patel, R\., Leonardi\-Bee, J\., Ackbarally, Y\., Bandyopadhyay, S\.\\BDBLWilson, E\.\\APACrefYearMonthDay2025\.\\BBOQ\\APACrefatitleThe use of virtual patients to provide feedback on clinical reasoning: a systematic review The use of virtual patients to provide feedback on clinical reasoning: a systematic review\.\\BBCQ\\APACjournalVolNumPagesAcademic Medicine1002229–238,\\PrintBackRefs\\CurrentBib
- Keifenheim\\BOthers\. \[\\APACyear2015\]\\APACinsertmetastarkeifenheim2015teaching\{APACrefauthors\}Keifenheim, K\.E\., Teufel, M\., Ip, J\., Speiser, N\., Leehr, E\.J\., Zipfel, S\.\\BCBLHerrmann\-Werner, A\.\\APACrefYearMonthDay2015\.\\BBOQ\\APACrefatitleTeaching history taking to medical students: a systematic review Teaching history taking to medical students: a systematic review\.\\BBCQ\\APACjournalVolNumPagesBMC medical education151159,\\PrintBackRefs\\CurrentBib
- Kononowicz\\BOthers\. \[\\APACyear2019\]\\APACinsertmetastarkononowicz2019virtual\{APACrefauthors\}Kononowicz, A\.A\., Woodham, L\.A\., Edelbring, S\., Stathakarou, N\., Davies, D\., Saxena, N\.\\BDBLZary, N\.\\APACrefYearMonthDay2019\.\\BBOQ\\APACrefatitleVirtual patient simulations in health professions education: systematic review and meta\-analysis by the digital health education collaboration Virtual patient simulations in health professions education: systematic review and meta\-analysis by the digital health education collaboration\.\\BBCQ\\APACjournalVolNumPagesJournal of medical Internet research217e14676,\\PrintBackRefs\\CurrentBib
- Kurtz\\BOthers\. \[\\APACyear2003\]\\APACinsertmetastarkurtz2003marrying\{APACrefauthors\}Kurtz, S\., Silverman, J\., Benson, J\.\\BCBLDraper, J\.\\APACrefYearMonthDay2003\.\\BBOQ\\APACrefatitleMarrying content and process in clinical method teaching: Enhancing the Calgary\-Cambridge guides Marrying content and process in clinical method teaching: Enhancing the calgary\-cambridge guides\.\\BBCQ\\APACjournalVolNumPagesAcademic Medicine788802–809,\{APACrefDOI\}[https://doi\.org/10\.1097/00001888\-200308000\-00011](https://doi.org/10.1097/00001888-200308000-00011)\\PrintBackRefs\\CurrentBib
- D\. Li\\BBALebai Lutfi \[\\APACyear2026\]\\APACinsertmetastarli2026large\{APACrefauthors\}Li, D\.\\BCBT\\BBALebai Lutfi, S\.\\APACrefYearMonthDay2026\.\\BBOQ\\APACrefatitleLarge Language Model–Based Virtual Patient Systems for History\-Taking in Medical Education: Comprehensive Systematic Review Large language model–based virtual patient systems for history\-taking in medical education: Comprehensive systematic review\.\\BBCQ\\APACjournalVolNumPagesJMIR Medical Informatics14e79039,\{APACrefDOI\}[https://doi\.org/10\.2196/79039](https://doi.org/10.2196/79039)\{APACrefURL\}https://medinform\.jmir\.org/2026/1/e79039/\\PrintBackRefs\\CurrentBib
- T\. Li\\BOthers\. \[\\APACyear2023\]\\APACinsertmetastarli2023analytics\{APACrefauthors\}Li, T\., Fan, Y\., Tan, Y\., Wang, Y\., Singh, S\., Li, X\.\\BDBLothers\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleAnalytics of self\-regulated learning scaffolding: effects on learning processes Analytics of self\-regulated learning scaffolding: effects on learning processes\.\\BBCQ\\APACjournalVolNumPagesFrontiers in Psychology14,\\PrintBackRefs\\CurrentBib
- Mahbubani \[\\APACyear2023\]\\APACinsertmetastarmahbubani2023history\{APACrefauthors\}Mahbubani, K\.\\APACrefYear2023\.\\APACrefbtitleHistory Taking in Clinical Practice History taking in clinical practice\.\\APACaddressPublisherSpringer\.\\PrintBackRefs\\CurrentBib
- Maicher\\BOthers\. \[\\APACyear2023\]\\APACinsertmetastarmaicher2023artificial\{APACrefauthors\}Maicher, K\.R\., Stiff, A\., Scholl, M\., White, M\., Fosler\-Lussier, E\., Schuler, W\.\\BDBLothers\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleArtificial intelligence in virtual standardized patients: combining natural language understanding and rule based dialogue management to improve conversational fidelity Artificial intelligence in virtual standardized patients: combining natural language understanding and rule based dialogue management to improve conversational fidelity\.\\BBCQ\\APACjournalVolNumPagesMedical teacher453279–285,\\PrintBackRefs\\CurrentBib
- Makoul \[\\APACyear2001\]\\APACinsertmetastarmakoul2001essential\{APACrefauthors\}Makoul, G\.\\APACrefYearMonthDay2001\.\\BBOQ\\APACrefatitleEssential elements of communication in medical encounters: The Kalamazoo consensus statement Essential elements of communication in medical encounters: The kalamazoo consensus statement\.\\BBCQ\\APACjournalVolNumPagesAcademic Medicine764390–393,\{APACrefDOI\}[https://doi\.org/10\.1097/00001888\-200104000\-00021](https://doi.org/10.1097/00001888-200104000-00021)\\PrintBackRefs\\CurrentBib
- Malau\-Aduli\\BOthers\. \[\\APACyear2022\]\\APACinsertmetastarmalauaduli2022osce\{APACrefauthors\}Malau\-Aduli, B\.S\., Jones, K\., Saad, S\.\\BCBLRichmond, C\.\\APACrefYearMonthDay2022\.\\BBOQ\\APACrefatitleHas the OSCE met its final demise? Rebalancing clinical assessment approaches in the peri\-pandemic world Has the osce met its final demise? rebalancing clinical assessment approaches in the peri\-pandemic world\.\\BBCQ\\APACjournalVolNumPagesFrontiers in Medicine9825502,\{APACrefDOI\}[https://doi\.org/10\.3389/fmed\.2022\.825502](https://doi.org/10.3389/fmed.2022.825502)\\PrintBackRefs\\CurrentBib
- Matcha\\BOthers\. \[\\APACyear2019\]\\APACinsertmetastarmatcha2019analytics\{APACrefauthors\}Matcha, W\., Gašević, D\., Uzir, N\.A\., Jovanović, J\.\\BCBLPardo, A\.\\APACrefYearMonthDay2019\.\\BBOQ\\APACrefatitleAnalytics of learning strategies: Associations with academic performance and feedback Analytics of learning strategies: Associations with academic performance and feedback\.\\BBCQ\\APACrefbtitleProceedings of the 9th International Conference on Learning Analytics & Knowledge Proceedings of the 9th international conference on learning analytics & knowledge \(\\BPGS461–470\)\.\\PrintBackRefs\\CurrentBib
- Milota\\BOthers\. \[\\APACyear2019\]\\APACinsertmetastarmilota2019narrative\{APACrefauthors\}Milota, M\.M\., van Thiel, G\.J\.M\.W\.\\BCBLvan Delden, J\.J\.M\.\\APACrefYearMonthDay2019\.\\BBOQ\\APACrefatitleNarrative medicine as a medical education tool: A systematic review Narrative medicine as a medical education tool: A systematic review\.\\BBCQ\\APACjournalVolNumPagesMedical Teacher417802–810,\{APACrefDOI\}[https://doi\.org/10\.1080/0142159X\.2019\.1584274](https://doi.org/10.1080/0142159X.2019.1584274)\\PrintBackRefs\\CurrentBib
- Misra\\BBASuresh \[\\APACyear2024\]\\APACinsertmetastarmisra2024osce\{APACrefauthors\}Misra, S\.M\.\\BCBT\\BBASuresh, S\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleArtificial intelligence and objective structured clinical examinations: using ChatGPT to revolutionize clinical skills assessment in medical education Artificial intelligence and objective structured clinical examinations: using chatgpt to revolutionize clinical skills assessment in medical education\.\\BBCQ\\APACjournalVolNumPagesJournal of Medical Education and Curricular Development1123821205241263475,\{APACrefDOI\}[https://doi\.org/10\.1177/23821205241263475](https://doi.org/10.1177/23821205241263475)\\PrintBackRefs\\CurrentBib
- Molenaar\\BOthers\. \[\\APACyear2023\]\\APACinsertmetastarmolenaar2023measuring\{APACrefauthors\}Molenaar, I\., de Mooij, S\., Azevedo, R\., Bannert, M\., Järvelä, S\.\\BCBLGašević, D\.\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleMeasuring self\-regulated learning and the role of AI: Five years of research using multimodal multichannel data Measuring self\-regulated learning and the role of ai: Five years of research using multimodal multichannel data\.\\BBCQ\\APACjournalVolNumPagesComputers in Human Behavior139107540,\\PrintBackRefs\\CurrentBib
- Nendaz\\BOthers\. \[\\APACyear2006\]\\APACinsertmetastarnendaz2006beyond\{APACrefauthors\}Nendaz, M\.R\., Gut, A\.M\., Perrier, A\., Louis\-Simonet, M\., Blondon\-Choa, K\., Herrmann, F\.R\.\\BDBLVu, N\.V\.\\APACrefYearMonthDay2006\.\\BBOQ\\APACrefatitleBrief report: beyond clinical experience: features of data collection and interpretation that contribute to diagnostic accuracy Brief report: beyond clinical experience: features of data collection and interpretation that contribute to diagnostic accuracy\.\\BBCQ\\APACjournalVolNumPagesJournal of general internal medicine21121302–1305,\\PrintBackRefs\\CurrentBib
- Ng\\BOthers\. \[\\APACyear2025\]\\APACinsertmetastarng2025clinical\{APACrefauthors\}Ng, I\.K\.S\., Goh, W\.G\.W\., Teo, D\.B\., Chong, K\.M\., Tan, L\.F\.\\BCBLTeoh, C\.M\.\\APACrefYearMonthDay2025\.\\BBOQ\\APACrefatitleClinical reasoning in real\-world practice: a primer for medical trainees and practitioners Clinical reasoning in real\-world practice: a primer for medical trainees and practitioners\.\\BBCQ\\APACjournalVolNumPagesPostgraduate Medical Journal101119168–75,\{APACrefDOI\}[https://doi\.org/10\.1093/postmj/qgae079](https://doi.org/10.1093/postmj/qgae079)\\PrintBackRefs\\CurrentBib
- Plackett\\BOthers\. \[\\APACyear2022\]\\APACinsertmetastarplackett2022effectiveness\{APACrefauthors\}Plackett, R\., Kassianos, A\.P\., Mylan, S\., Kambouri, M\., Raine, R\.\\BCBLSheringham, J\.\\APACrefYearMonthDay2022\.\\BBOQ\\APACrefatitleThe effectiveness of using virtual patient educational tools to improve medical students’ clinical reasoning skills: a systematic review The effectiveness of using virtual patient educational tools to improve medical students’ clinical reasoning skills: a systematic review\.\\BBCQ\\APACjournalVolNumPagesBMC Medical Education221365,\{APACrefDOI\}[https://doi\.org/10\.1186/s12909\-022\-03410\-x](https://doi.org/10.1186/s12909-022-03410-x)\\PrintBackRefs\\CurrentBib
- Raković\\BOthers\. \[\\APACyear2023\]\\APACinsertmetastarrakovic2023harnessing\{APACrefauthors\}Raković, M\., Iqbal, S\., Li, T\., Fan, Y\., Singh, S\., Surendrannair, S\.\\BDBLothers\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleHarnessing the potential of trace data and linguistic analysis to predict learner performance in a multi\-text writing task Harnessing the potential of trace data and linguistic analysis to predict learner performance in a multi\-text writing task\.\\BBCQ\\APACjournalVolNumPagesJournal of Computer Assisted Learning393,\\PrintBackRefs\\CurrentBib
- Regehr\\BOthers\. \[\\APACyear1998\]\\APACinsertmetastarregehr1998comparing\{APACrefauthors\}Regehr, G\., MacRae, H\., Reznick, R\.K\.\\BCBLSzalay, D\.\\APACrefYearMonthDay1998\.\\BBOQ\\APACrefatitleComparing the psychometric properties of checklists and global rating scales for assessing performance on an OSCE\-format examination Comparing the psychometric properties of checklists and global rating scales for assessing performance on an osce\-format examination\.\\BBCQ\\APACjournalVolNumPagesAcademic Medicine739993–7,\\PrintBackRefs\\CurrentBib
- Reimann \[\\APACyear2007\]\\APACinsertmetastarreimann2007time\{APACrefauthors\}Reimann, P\.\\APACrefYearMonthDay2007\.\\BBOQ\\APACrefatitleTime is precious: Why process analysis is essential for CSCL \(and can also help to bridge between experimental and descriptive methods\) Time is precious: Why process analysis is essential for cscl \(and can also help to bridge between experimental and descriptive methods\)\.\\BBCQ\\PrintBackRefs\\CurrentBib
- Roter\\BBALarson \[\\APACyear2002\]\\APACinsertmetastarroter2002roter\{APACrefauthors\}Roter, D\.\\BCBT\\BBALarson, S\.\\APACrefYearMonthDay2002\.\\BBOQ\\APACrefatitleThe Roter interaction analysis system \(RIAS\): utility and flexibility for analysis of medical interactions The roter interaction analysis system \(rias\): utility and flexibility for analysis of medical interactions\.\\BBCQ\\APACjournalVolNumPagesPatient education and counseling464243–251,\\PrintBackRefs\\CurrentBib
- Saint\\BOthers\. \[\\APACyear2022\]\\APACinsertmetastarsaint2022temporally\{APACrefauthors\}Saint, J\., Fan, Y\., Gašević, D\.\\BCBLPardo, A\.\\APACrefYearMonthDay2022\.\\BBOQ\\APACrefatitleTemporally\-focused analytics of self\-regulated learning: A systematic review of literature Temporally\-focused analytics of self\-regulated learning: A systematic review of literature\.\\BBCQ\\APACjournalVolNumPagesComputers and Education: Artificial Intelligence3100060,\{APACrefDOI\}[https://doi\.org/10\.1016/j\.caeai\.2022\.100060](https://doi.org/10.1016/j.caeai.2022.100060)\\PrintBackRefs\\CurrentBib
- Saint\\BOthers\. \[\\APACyear2020\]\\APACinsertmetastarsaint2020combining\{APACrefauthors\}Saint, J\., Gašević, D\., Matcha, W\., Uzir, N\.A\.\\BCBLPardo, A\.\\APACrefYearMonthDay2020\.\\BBOQ\\APACrefatitleCombining analytic methods to unlock sequential and temporal patterns of self\-regulated learning Combining analytic methods to unlock sequential and temporal patterns of self\-regulated learning\.\\BBCQ\\APACrefbtitleProceedings of the 10th International Conference on Learning Analytics and Knowledge*\(LAK 2020\), 23–27 March 2020, Frankfurt, Germany*Proceedings of the 10th International Conference on Learning Analytics and Knowledge*\(LAK 2020\), 23–27 March 2020, Frankfurt, Germany*\(\\BPGS402–411\)\.\\APACaddressPublisherACM\.\\PrintBackRefs\\CurrentBib
- Saqr\\BOthers\. \[\\APACyear2024\]\\APACinsertmetastarsaqr2024sequence\{APACrefauthors\}Saqr, M\., López\-Pernas, S\., Helske, S\., Durand, M\., Murphy, K\., Studer, M\.\\BCBLRitschard, G\.\\APACrefYearMonthDay2024\.\\BBOQ\\APACrefatitleSequence analysis in education: principles, technique, and tutorial with R Sequence analysis in education: principles, technique, and tutorial with r\.\\BBCQ\\APACrefbtitleLearning analytics methods and tutorials: A practical guide using R Learning analytics methods and tutorials: A practical guide using r \(\\BPGS321–354\)\.\\APACaddressPublisherSpringer Nature Switzerland Cham\.\\PrintBackRefs\\CurrentBib
- Schmidt\\BBARikers \[\\APACyear2007\]\\APACinsertmetastarschmidt2007expertise\{APACrefauthors\}Schmidt, H\.G\.\\BCBT\\BBARikers, R\.M\.\\APACrefYearMonthDay2007\.\\BBOQ\\APACrefatitleHow expertise develops in medicine: knowledge encapsulation and illness script formation How expertise develops in medicine: knowledge encapsulation and illness script formation\.\\BBCQ\\APACjournalVolNumPagesMedical education41121133–1139,\\PrintBackRefs\\CurrentBib
- Shaffer\\BOthers\. \[\\APACyear2016\]\\APACinsertmetastarshaffer2016tutorial\{APACrefauthors\}Shaffer, D\.W\., Collier, W\.\\BCBLRuis, A\.R\.\\APACrefYearMonthDay2016\.\\BBOQ\\APACrefatitleA tutorial on epistemic network analysis: Analyzing the structure of connections in cognitive, social, and interaction data A tutorial on epistemic network analysis: Analyzing the structure of connections in cognitive, social, and interaction data\.\\BBCQ\\APACjournalVolNumPagesJournal of Learning Analytics339–45,\{APACrefDOI\}[https://doi\.org/10\.18608/jla\.2016\.33\.3](https://doi.org/10.18608/jla.2016.33.3)\\PrintBackRefs\\CurrentBib
- Shea\\BBAChan \[\\APACyear2023\]\\APACinsertmetastarshea2023clinical\{APACrefauthors\}Shea, G\.K\.\\BCBT\\BBAChan, P\.C\.\\APACrefYearMonthDay2023\.\\BBOQ\\APACrefatitleClinical reasoning in medical education: a primer for medical students Clinical reasoning in medical education: a primer for medical students\.\\BBCQ\\APACjournalVolNumPagesTeaching and Learning in Medicine364547–555,\{APACrefDOI\}[https://doi\.org/10\.1080/10401334\.2023\.2230201](https://doi.org/10.1080/10401334.2023.2230201)\\PrintBackRefs\\CurrentBib
- Si \[\\APACyear2022\]\\APACinsertmetastarsi2022strategies\{APACrefauthors\}Si, J\.\\APACrefYearMonthDay2022\.\\BBOQ\\APACrefatitleStrategies for developing pre\-clinical medical students’ clinical reasoning based on illness script formation: a systematic review Strategies for developing pre\-clinical medical students’ clinical reasoning based on illness script formation: a systematic review\.\\BBCQ\\APACjournalVolNumPagesKorean Journal of Medical Education34149–61,\{APACrefDOI\}[https://doi\.org/10\.3946/kjme\.2022\.219](https://doi.org/10.3946/kjme.2022.219)\\PrintBackRefs\\CurrentBib
- Sonnenberg\\BBABannert \[\\APACyear2015\]\\APACinsertmetastarsonnenberg2015discovering\{APACrefauthors\}Sonnenberg, C\.\\BCBT\\BBABannert, M\.\\APACrefYearMonthDay2015\.\\BBOQ\\APACrefatitleDiscovering the effects of metacognitive prompts on the sequential structure of SRL\-processes using process mining techniques Discovering the effects of metacognitive prompts on the sequential structure of srl\-processes using process mining techniques\.\\BBCQ\\APACjournalVolNumPagesJournal of Learning Analytics2172–100,\\PrintBackRefs\\CurrentBib
- Swiecki\\BOthers\. \[\\APACyear2022\]\\APACinsertmetastarswiecki2022assessment\{APACrefauthors\}Swiecki, Z\., Khosravi, H\., Chen, G\., Martinez\-Maldonado, R\., Lodge, J\.M\., Milligan, S\.\\BDBLGašević, D\.\\APACrefYearMonthDay2022\.\\BBOQ\\APACrefatitleAssessment in the age of artificial intelligence Assessment in the age of artificial intelligence\.\\BBCQ\\APACjournalVolNumPagesComputers and Education: Artificial Intelligence3100075,\\PrintBackRefs\\CurrentBib
- Tao\\BOthers\. \[\\APACyear2025\]\\APACinsertmetastartao2025exploring\{APACrefauthors\}Tao, L\., Song, Y\.\\BCBLFu, J\.\\APACrefYearMonthDay2025\.\\BBOQ\\APACrefatitleExploring students’ self\-regulated learning behavioural patterns and perceptions in an English speaking task within a generative AI\-supported immersive VR Exploring students’ self\-regulated learning behavioural patterns and perceptions in an english speaking task within a generative ai\-supported immersive vr\.\\BBCQ\\APACjournalVolNumPagesComputers & Education105515,\\PrintBackRefs\\CurrentBib
- Thampy\\BOthers\. \[\\APACyear2019\]\\APACinsertmetastarthampy2019assessing\{APACrefauthors\}Thampy, H\., Willert, E\.\\BCBLRamani, S\.\\APACrefYearMonthDay2019\.\\BBOQ\\APACrefatitleAssessing clinical reasoning: targeting the higher levels of the pyramid Assessing clinical reasoning: targeting the higher levels of the pyramid\.\\BBCQ\\APACjournalVolNumPagesJournal of general internal medicine3481631–1636,\\PrintBackRefs\\CurrentBib
- Veenman \[\\APACyear2007\]\\APACinsertmetastarveenman2007assessment\{APACrefauthors\}Veenman, M\.V\.\\APACrefYearMonthDay2007\.\\BBOQ\\APACrefatitleThe assessment and instruction of self\-regulation in computer\-based environments: a discussion The assessment and instruction of self\-regulation in computer\-based environments: a discussion\.\\BBCQ\\APACjournalVolNumPagesMetacognition and Learning22177–183,\{APACrefDOI\}[https://doi\.org/10\.1007/s11409\-007\-9017\-6](https://doi.org/10.1007/s11409-007-9017-6)\\PrintBackRefs\\CurrentBib
- Verbert\\BOthers\. \[\\APACyear2014\]\\APACinsertmetastarverbert2014learning\{APACrefauthors\}Verbert, K\., Govaerts, S\., Duval, E\., Santos, J\.L\., Van Assche, F\., Parra, G\.\\BCBLKlerkx, J\.\\APACrefYearMonthDay2014\.\\BBOQ\\APACrefatitleLearning dashboards: an overview and future research opportunities Learning dashboards: an overview and future research opportunities\.\\BBCQ\\APACjournalVolNumPagesPersonal and Ubiquitous Computing1861499–1514,\\PrintBackRefs\\CurrentBib
- Wagner\-Menghin\\BOthers\. \[\\APACyear2020\]\\APACinsertmetastarwagner2020communication\{APACrefauthors\}Wagner\-Menghin, M\., de Bruin, A\.B\.\\BCBLvan Merriënboer, J\.J\.\\APACrefYearMonthDay2020\.\\BBOQ\\APACrefatitleCommunication skills supervisors’ monitoring of history\-taking performance: an observational study on how doctors and non\-doctors use cues to prepare feedback Communication skills supervisors’ monitoring of history\-taking performance: an observational study on how doctors and non\-doctors use cues to prepare feedback\.\\BBCQ\\APACjournalVolNumPagesBMC medical education20136,\\PrintBackRefs\\CurrentBib
- Windish\\BOthers\. \[\\APACyear2005\]\\APACinsertmetastarwindish2005teaching\{APACrefauthors\}Windish, D\.M\., Price, E\.G\., Clever, S\.L\., Magaziner, J\.L\.\\BCBLThomas, P\.A\.\\APACrefYearMonthDay2005\.\\BBOQ\\APACrefatitleTeaching medical students the important connection between communication and clinical reasoning Teaching medical students the important connection between communication and clinical reasoning\.\\BBCQ\\APACjournalVolNumPagesJournal of General Internal Medicine20121108–1113,\\PrintBackRefs\\CurrentBib
- Winne \[\\APACyear2022\]\\APACinsertmetastarwinne2022modeling\{APACrefauthors\}Winne, P\.H\.\\APACrefYearMonthDay2022\.\\BBOQ\\APACrefatitleModeling self\-regulated learning as learners doing learning science: How trace data and learning analytics help develop skills for self\-regulated learning Modeling self\-regulated learning as learners doing learning science: How trace data and learning analytics help develop skills for self\-regulated learning\.\\BBCQ\\APACjournalVolNumPagesMetacognition and Learning173773–791,\\PrintBackRefs\\CurrentBib
- Woodham\\BOthers\. \[\\APACyear2019\]\\APACinsertmetastarwoodham2019virtual\{APACrefauthors\}Woodham, L\.A\., Round, J\., Stenfors, T\., Bujacz, A\., Karlgren, K\., Jivram, T\.\\BDBLPoulton, T\.\\APACrefYearMonthDay2019\.\\BBOQ\\APACrefatitleVirtual patients designed for training against medical error: Exploring the impact of decision\-making on learner motivation Virtual patients designed for training against medical error: Exploring the impact of decision\-making on learner motivation\.\\BBCQ\\APACjournalVolNumPagesPloS one144e0215597,\\PrintBackRefs\\CurrentBib
- Xu\\BOthers\. \[\\APACyear2021\]\\APACinsertmetastarxu2021methods\{APACrefauthors\}Xu, H\., Ang, B\.W\.G\., Soh, J\.Y\.\\BCBLPonnamperuma, G\.G\.\\APACrefYearMonthDay2021\.\\BBOQ\\APACrefatitleMethods to improve diagnostic reasoning in undergraduate medical education in the clinical setting: a systematic review Methods to improve diagnostic reasoning in undergraduate medical education in the clinical setting: a systematic review\.\\BBCQ\\APACjournalVolNumPagesJournal of General Internal Medicine3692745–2754,\{APACrefDOI\}[https://doi\.org/10\.1007/s11606\-021\-06916\-0](https://doi.org/10.1007/s11606-021-06916-0)\\PrintBackRefs\\CurrentBib
- Yi\\BBAKim \[\\APACyear2025\]\\APACinsertmetastaryi2025feasibility\{APACrefauthors\}Yi, Y\.\\BCBT\\BBAKim, K\\BHBIJ\.\\APACrefYearMonthDay2025\.\\BBOQ\\APACrefatitleThe feasibility of using generative artificial intelligence for history taking in virtual patients The feasibility of using generative artificial intelligence for history taking in virtual patients\.\\BBCQ\\APACjournalVolNumPagesBMC Research Notes18180,\\PrintBackRefs\\CurrentBib

Similar Articles

Synthesis and Evaluation of Long-term History-aware Medical Dialogue

arXiv cs.CL

This paper introduces a framework for synthesizing long-term medical dialogue datasets using LLMs, and creates MediLongChat with three benchmark tasks to evaluate healthcare agents' memory and reasoning capabilities. Experiments show that even state-of-the-art LLMs struggle with these tasks.

Teaching with AI

OpenAI Blog

OpenAI shares perspectives from educators on integrating AI tools like ChatGPT into teaching, including using AI for language support and teaching students to think critically about AI-generated information.