CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

arXiv cs.CL Papers

Summary

CliniCIRCA is a multi-stage LLM framework for reconstructing longitudinal mental health patient journeys from unstructured EHR narratives, achieving temporal classification without event-level timestamps and using clinician-in-the-loop evaluation for verification.

arXiv:2609.19585v1 Announce Type: new Abstract: In mental health care, reasoning over patient journeys is a key task for clinicians. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making temporal recovery challenging. We present CliniCIRCA, a multi-stage LLM framework for Calendar-anchored, Imprecision-aware Reconstruction of Clinical Annals. To our knowledge, CliniCIRCA is the first to temporally classify clinical events across unstructured discharge summaries without event-level timestamps. From 14,882 MIMIC-III mental health admissions, we first construct a benchmark of 52 discharge summaries on which CliniCIRCA produces 15,891 temporally tagged events. After correcting 629 errors based on a clinician-in-the-loop evaluation, we produce verified gold-standard labels. Finally, the corrected timelines drive a temporally grounded summarization stage that compresses each source 1.52 times into a date-grouped chronological record. We then scale the framework to generate 1,000 silver-standard timelines and evaluate them as training data. Compared with zero- and few-shot prompting, instruction tuning generally improves five open-weight models on event extraction, temporal tagging, and summarization across silver and clinician-verified evaluations.
Original Article
View Cached Full Text

Cached at: 09/18/26, 08:56 AM

# CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives
Source: [https://arxiv.org/html/2609.19585](https://arxiv.org/html/2609.19585)
Nimra IshfaqMohit ChandraSantiago Alvarez LesmesAffiliation:Georgia Institute of Technology University of Texas at Austin Northwell HealthAdam CosciaKhatiya Chelidze MoonAffiliation:Georgia Institute of Technology University of Texas at Austin Northwell HealthXiaohan DingMunmun De Choudhury

###### Abstract

In mental health care, reasoning over patient journeys is a key task for clinicians\. Yet these journeys, encompassing a longitudinal progression of biological, psychological, and social events, are often spread across disparate unstructured text narratives, making temporal recovery challenging\. We presentCliniCIRCA, a multi\-stage LLM framework forCalendar\-anchored,Imprecision\-awareReconstruction ofClinicalAnnals\. To our knowledge,CliniCIRCAis the first to temporally classify clinical events across unstructured discharge summaries without event\-level timestamps\. From14,88214,882MIMIC\-III mental health admissions, we first construct a benchmark of5252discharge summaries on whichCliniCIRCAproduces15,89115,891temporally tagged events\. After correcting629629errors based on a clinician\-in\-the\-loop evaluation, we produce verified gold\-standard labels\. Finally, the corrected timelines drive a temporally grounded summarization stage that compresses each source1\.521\.52times into a date\-grouped chronological record\. We then scale the framework to generate1,0001,000silver\-standard timelines and evaluate them as training data\. Compared with zero\- and few\-shot prompting, instruction tuning generally improves five open\-weight models on event extraction, temporal tagging, and summarization across silver and clinician\-verified evaluations\.

## 1Introduction

Every patient follows a clinical journey: a longitudinal progression of symptoms, diagnoses, treatments, life events, and responses to care that demonstrate how their health has evolved over time[Kraljevic et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib55);[Ruan et al\. \(2019\)](https://arxiv.org/html/2609.19585#bib.bib56)\. Clinicians draw on this progression to interpret a patient’s current presentation and guide treatment\.

Unfortunately, these journeys are often not documented as explicit timelines[Loftus et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib53);[Linhares et al\. \(2023\)](https://arxiv.org/html/2609.19585#bib.bib54);[Olex and McInnes \(2021\)](https://arxiv.org/html/2609.19585#bib.bib15)\. For example, discharge summaries are a rich source of longitudinal patient information\. They synthesize prior history, hospital course, and clinical outcomes in free\-form text, where events may be repeated, narrated nonsequentially, or linked only through relative temporal expressions and surrounding context[K et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib40);[Seinen et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib39);[Lee et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib41);[Kim et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib38)\.

Recovering temporal structure is thus a central challenge in clinical NLP for modeling patient journeys from clinical text[Amirahmadi et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib37);[Makarov et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib36)\. Patient journeys are often spread across disparate text sources, limiting the use of these narratives by computational systems that require structured longitudinal representations[Chaturvedi \(2024\)](https://arxiv.org/html/2609.19585#bib.bib57);[Moharasan and Ho \(2019\)](https://arxiv.org/html/2609.19585#bib.bib58)\. As a result, information fragmentation shifts the burden of reconstruction to clinicians and computational systems\. Clinicians must piece together the patient’s course under time and information constraints, while models must infer temporal relations that remain implicit[Linhares et al\. \(2023\)](https://arxiv.org/html/2609.19585#bib.bib54);[Gao et al\. \(2022\)](https://arxiv.org/html/2609.19585#bib.bib17);[Olex and McInnes \(2021\)](https://arxiv.org/html/2609.19585#bib.bib15)\.

![Refer to caption](https://arxiv.org/html/2609.19585v1/FIG1_V5.png)Figure 1:Methodology OverviewPrior work treated event reconstruction as a set of separable subtasks, which makes temporal alignment challenging\. For example, clinical summarization compresses lengthy records into coherent accounts, but often under\-preserves temporality and rarely represents complex or uncertain event chronology explicitly[Kruse et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib8);[Cui et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib22);[Croxford et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib14)\. Temporal information extraction preserves finer structure by extracting time expressions, events, and temporal relations, but most systems are built around predefined annotation schemas and pairwise relation classification rather than recovery of an open\-ended patient trajectory[Chaturvedi et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib9);[Gumiel et al\. \(2022\)](https://arxiv.org/html/2609.19585#bib.bib59);[Alfattni et al\. \(2020\)](https://arxiv.org/html/2609.19585#bib.bib60)\.

Recent LLM timeline generation comes closer to recovering full patient trajectories[Kumar et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib19);[Noroozizadeh et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib21);[Wang and Weiss \(2025\)](https://arxiv.org/html/2609.19585#bib.bib23);[Wang et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib4), but often relies on structured EHR timestamps, preselected note fragments, or curated case reports, and commonly represents time using a single relative offset\. What remains missing is a unified, text\-only reconstruction formulation that preserves event breadth, calendar timing, temporal uncertainty, and patient\-level coherence\.

We specifically focus on bridging the gap in mental health care, where a patient’s course cannot be understood only from diagnoses and procedures alone\. Engel’s framework emphasizes that illness emerges through the interaction of biological, psychological, and social factors[Engel \(1977\)](https://arxiv.org/html/2609.19585#bib.bib42); in mental health, these factors are often expressed through medication use, trauma, social history, changes in daily functioning, etc[Saha et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib6)\. Such information is part of the trajectory rather than peripheral context, yet the combination of events differs across patients\. Mental health records therefore expose the limitations of fixed schemas: a narrow schema would omit events that may define an individual cause, while an overly exhaustive one quickly becomes impractical\.

Hence, we formulate patient\-journey recovery asopen\-vocabulary, calendar\-anchored, and uncertainty\-aware temporal reconstruction from raw clinical narratives\. Specifically,journeys are reconstructed within a single discharge summary corresponding to one hospital admission, as a step toward multi\-encounter longitudinal modeling\. We introduceCliniCIRCA\(Figure[1](https://arxiv.org/html/2609.19585#S1.F1)\), a three\-stage LLM framework to convert discharge summaries to coherent temporal clinical trajectories\.CliniCIRCAproduces two linked representations: Output A, an inspectable event\-level timeline for computational use and auditing, and Output B, a clinician\-readable account derived from the same temporal substrate\. Our experiments demonstrate that patient journeys produced by a high\-capacity model \(Section[6](https://arxiv.org/html/2609.19585#S6)\) can provide effective supervision for substantially smaller models, particularly for atomic clinical event extraction, while temporal reconstruction remains reasoning\-intensive\. By transforming fragmented narratives into linked structured and readable representations,CliniCIRCAcould support longitudinal chart review, clinical handoffs, and follow\-up care while preserving an auditable connection to the underlying events\. We have provided code in a repository here111https://anonymous\.4open\.science/r/CliniCIRCA\-E5FF/\.

## 2Dataset

We use MIMIC\-III v1\.4[Johnson et al\. \(2016\)](https://arxiv.org/html/2609.19585#bib.bib1);[Johnson et al\. \(2015\)](https://arxiv.org/html/2609.19585#bib.bib2), a publicly available, de\-identified critical\-care database from the Beth Israel Deaconess Medical Center \(2001–2012\) in English\.CliniCIRCAreceives unprocessed discharge summaries without chunking, section parsing, or normalization\. Structured EHR fields are used to select the cohort and provide three temporal anchors: \(1\) date of birth, \(2\) admission date, and \(3\) discharge date\. We define the cohort at the admission level: an admission is retained if it contains at least one psychiatric diagnosis from the ICD\-9 Mental Disorders chapter \(codes 290–319\) and can be paired with a discharge summary\. This produces14,88214,882summaries from12,27312,273patients, each summary treated as a separate reconstruction instance\. Because admissions are flagged by the presence of any Mental Disorders diagnosis code rather than by principal diagnosis, the cohort includes both primary psychiatric admissions and medical/surgical admissions with psychiatric comorbidity\. More details can be found in Appendix[A](https://arxiv.org/html/2609.19585#A1)\.

## 3TheCliniCIRCAFramework

CliniCIRCAis a sequential framework comprising three LLM stages \(Figure[1](https://arxiv.org/html/2609.19585#S1.F1)\), each receiving the output of the previous stage and having a single main task\.Stage 1reads the raw discharge summary \(obtained from the MIMIC\-III Dataset\) without preprocessing and extracts a high\-recall list of atomic clinical events\.Stage 2resolves the timing of each event, tagging it with an ISO date and a certainty label anchored against temporal references; its output isOutput A, a timeline of \(ISO date and certainty tag, atomic event\) pairs\.Stage 3aggregates these tagged events by date intoOutput B, a date\-grouped chronological summary\. Decomposing the task in this way allows each stage to be prompted, evaluated, and improved independently, and confines each temporal decision to the stage best positioned to make it\.

## 4Model and Agent Evaluation Setup

### 4\.1Candidate Pool

ForCliniCIRCA, we select1111LLM candidates based on reported performance across general\-purpose and clinical benchmarks[Yu et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib50);[Sarvari and Al\-fagih \(2025\)](https://arxiv.org/html/2609.19585#bib.bib52);[Shi et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib51)\. For proprietary models, we evaluate Gemini 2\.5 Pro and 2\.5 Flash[Comanici et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib70); Claude Sonnet 4\.6[Anthropic \(2026\)](https://arxiv.org/html/2609.19585#bib.bib71)and Haiku 4\.5[Anthropic \(2025\)](https://arxiv.org/html/2609.19585#bib.bib72)\. For open\-weight models, we evaluate Llama 3\.3 70B[Meta AI \(2024\)](https://arxiv.org/html/2609.19585#bib.bib73), Qwen 3\.6 35B\-A3B[Qwen Team \(2026\)](https://arxiv.org/html/2609.19585#bib.bib74), Gemma 4 26B\-A4B[Gemma Team \(2026\)](https://arxiv.org/html/2609.19585#bib.bib76), Mistral Small 3\.2 24B\-A4B[Mistral AI \(2025\)](https://arxiv.org/html/2609.19585#bib.bib77), Phi\-4 Mini 3\.8B[Abouelenin and others \(2025\)](https://arxiv.org/html/2609.19585#bib.bib79), and medical\-domain LLMs \- MedGemma 27B[Sellergren et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib80)and Meditron 3 Qwen 2\.5 7B[OpenMeditron \(2025\)](https://arxiv.org/html/2609.19585#bib.bib78)\. Model specifications, generation settings, and computational details are provided in Appendix[B](https://arxiv.org/html/2609.19585#A2)\.

### 4\.2Implementation Setup

We construct a benchmark set of5252MIMIC\-III discharge summaries, averaging1,7741,774words each \(range:437437–3,5833,583; Appendix[A\.2](https://arxiv.org/html/2609.19585#A1.SS2)\)\. Of the 52 cases, 29 had a principal psychiatric diagnosis; the remaining 23 were medical or surgical admissions with psychiatric comorbidity\. For the benchmark, every candidate is run through Stages 1\-2, with no mixing of models across stages\. The size is determined by a power analysis \(Appendix[C](https://arxiv.org/html/2609.19585#A3)\) for a pairwiseχ2\\chi^\{2\}comparison \(large effect,α\\alpha=0\.050\.05,9595% power\), yieldingN=5252needed to reach statistical significance[Ki et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib31);[Card et al\. \(2020\)](https://arxiv.org/html/2609.19585#bib.bib46);[Dror et al\. \(2018\)](https://arxiv.org/html/2609.19585#bib.bib47)\.

Each candidate is run through Stages 1\-2 as a fixed pipeline configuration, producing572572model–document outputs evaluated using the clinician\-authored rubric described in Section[5\.1](https://arxiv.org/html/2609.19585#S5.SS1)\. We retain the selected model for Stage 3 rather than optimizing a separate model for each stage\. This avoids a combinatorial search over model combinations, reveals how one model’s capabilities propagate through theCliniCIRCAframework, and simplifies reproducibility and deployment governance[Umeton et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib29);[de Hond et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib30);[Meskó and Topol \(2023\)](https://arxiv.org/html/2609.19585#bib.bib28)\. We select based on Stage 2 because the summary stage depends on the accuracy of its event\-level temporal input\. Following prior health\-domain work that validates LLM judges against expert\-vetted human labels[Mittal et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib13), two human annotators and two clinicians, all with relevant annotation experience and high proficiency in English, form a feedback loop to evaluate the results\. We establish human\-evaluator agreement on both Stage\-2 and 3 outputs\.

## 5Constructing Patient Timelines from Discharge Summaries

### 5\.1Stage 1 & 2: From Raw Discharge Summary to Tagged Timeline

Faithful clinical timeline recovery requires both comprehensive event extraction and accurate temporal anchoring, yet extraction errors can propagate through temporal reasoning and timeline construction[Gumiel et al\. \(2022\)](https://arxiv.org/html/2609.19585#bib.bib59);[Alfattni et al\. \(2020\)](https://arxiv.org/html/2609.19585#bib.bib60)\. We therefore decompose the task into Stage 1 for event extraction and Stage 2 for temporal anchoring, allowing each component to be independently specified and evaluated\. Full generation prompts can be found in Appendix[D](https://arxiv.org/html/2609.19585#A4)Tables[A5](https://arxiv.org/html/2609.19585#A4.T5)\-[A6](https://arxiv.org/html/2609.19585#A4.T6)\.

Table 1:Stage\-2 model\-selection results on the 52\-sample benchmark\. The highlighted row indicates the model selected for the final pipeline\. Results are averaged over 52 samples\. Dimension A: Completeness; Dimension B: Semantic Faithfulness; and Dimension C: Temporal Tag Accuracy\. Dimensions A to C are Likert\-scale averages reported to two decimal places\. The best\-performing value for each metric is shown inbold, and the second\-best value isunderlined\. Tied values receive the same formatting\.Inclusive extraction and anchor\-based tagging\.Stage 1 reads the raw discharge summary, applying no chunking, section parsing, or normalization, and extracts events inclusively and faithfully\. Accordingly, Stage 2 then resolves each event against the admission date, discharge date, and date of birth, assigning exactly one of four temporal categories: 1\)\[DATE/RANGE\] \[EXACT\]: a date written in the event text itself\. 2\)\[DATE/RANGE\] \[APPROX\]: no written date, but the event resolves to an anchor or a relative offset from one\. 3\)\[PRE ADM\]: framed as prior to or continuous through this admission \(history, “status post,” demographics, undated chronic states\), and 4\)\[INDETERMINATE\]: a guardrail used only when none of the above applies and the timing genuinely cannot be determined\.

As specified above, to maximize precision, only explicitly written dates receive \[EXACT\]; any inferred or relative date resolves to \[APPROX\] at the coarsest defensible granularity \(e\.g\., year\-only when only a year is recoverable\)\. Stage 2’s output is a list of temporally classified clinical events\.

#### Operationalizing the notion of “faithfulness”\.

Because no reference timeline exists against which these tuples could be scored, a “good” timeline must itself be operationalized\. We define three clinician\-informed criteria for evaluating the Stage\-2 output \(full definitions in Appendix[E](https://arxiv.org/html/2609.19585#A5)\)\. A\) Completeness:whether all clinically relevant events in the note are captured\. B\) Semantic faithfulness:whether each extracted event preserves the meaning of the source text\. C\) Temporal classification accuracy:whether the assigned date and tag are correct\.

#### Calibrating the automated judge\.

Following established practice in LLM evaluation to use a strong later\-generation model[Zheng et al\. \(2023\)](https://arxiv.org/html/2609.19585#bib.bib18), we select Gemini 3\.1 Pro as the LLM evaluator, chosen to be at least as capable as any candidate it scores\. Prior to selection, the evaluator is calibrated against expert human scoring\. Only the calibrated LLM evaluator is used in the selection and scaling\.

Selecting the agent model\.Following the agent selection mechanism as described in Section[4\.2](https://arxiv.org/html/2609.19585#S4.SS2), Gemini 2\.5 Pro is selected as the strongest configuration, ranking among the top candidates on every dimension and achieving5\.005\.00on Completeness,5\.005\.00on Semantic Faithfulness, and4\.834\.83on Temporal Tag Accuracy in Table[1](https://arxiv.org/html/2609.19585#S5.T1)\. As a judge\-independent check, we also report MEDCON[Yim et al\. \(2023\)](https://arxiv.org/html/2609.19585#bib.bib34), the F1 overlap of UMLS medical concepts between each model’s extracted events and the source note; once again, Gemini 2\.5 Pro attains the highest MEDCON F1 \(Table[1](https://arxiv.org/html/2609.19585#S5.T1)\), corroborating its judge\-scored ranking\. We select based on Stage 2 alone and then freeze Gemini 2\.5 Pro as the only agent advancing to Stage 3\.

### 5\.2Validating the Stage 1\-2 Timeline

We validate the output of Stage 1\-2 through a human review with two psychiatric clinicians during 6 iterative annotation sessions\. The human review process was used both to evaluate the automated judge and to correct all confirmed errors, providing a clinician\-verified gold reference dataset \(N=52N=52\)\. Evaluation was conducted on an internally hosted, custom web\-based secure platform, demonstrated in Appendix[F](https://arxiv.org/html/2609.19585#A6)\.

We compare the calibrated Gemini 3\.1 Pro evaluator with the clinician\-verified gold annotation \(see Appendix[E](https://arxiv.org/html/2609.19585#A5)for complete rubrics\)\. Overall agreement is high, with Cohen’sκ=0\.655\\kappa=0\.655\. However, because only3\.96%3\.96\\%of events were labeled as errors, class imbalance likely depressesκ\\kappa[Gwet \(2008\)](https://arxiv.org/html/2609.19585#bib.bib12)\. This is reflected in a Gwet’s AC1 of0\.9720\.972and negative agreement of0\.9860\.986, indicating strong agreement on events judged correct\. Agreement is lower for the rare error class, with positive agreement of0\.6690\.669\.

### 5\.3Stage 3: From Timeline to Temporally\-Guided Summary

Next, Stage 3 tests if the verified event timeline can be rendered as a faithful, temporally grounded summary without re\-reading the discharge note or reconstructing its chronology\. Two clinicians defined a specification for the mental\-health cohort that retains demographics and social history, reason for admission, key in\-stay events, medication changes, discharge, and follow\-up\. Each event is grouped under its Stage 2 temporal label\.

Stage 3 is evaluated for faithfulness to the verified timeline and temporal accuracy using a separately calibrated Gemini 3\.1 Pro evaluator validated against clinician judgments\. Generation prompts and rubrics appear in Appendices[H](https://arxiv.org/html/2609.19585#A8)\-[I](https://arxiv.org/html/2609.19585#A9)\. Across the benchmark, Stage 3 reduces reading burden by roughly a third \(31\.5% mean and 33\.8% median per\-note; 34\.0% corpus\-wide\), a stable 1\.52 times compression that removes∼\\sim600 words per admission \(Table[2](https://arxiv.org/html/2609.19585#S5.T2)\)\.

Table 2:Reading\-burden reduction from Stage 3 summarization over the 52\-note benchmark\. Per\-note percentages are averaged across notes; corpus\-level reduction pools word counts across all notes\.
### 5\.4Validating the Stage 3 Summary

We compare the judge’s error counts against human adjudication across5252summaries \(Table[A21](https://arxiv.org/html/2609.19585#A11.T21)\)\. The judge flags262262errors, of which human annotators uphold3939: a6\.76\.7\-fold inflation, averaging4\.34\.3spurious errors per summary\. Agreement is substantial on whether a summary contains any error \(κ=0\.77\\kappa=0\.77, positive agreement = 0\.84\) and on the relative ordering of summaries by error count \(Spearmanρ\\rho=0\.770\.77\)\. With clinician guidance, we develop a two\-dimensional rubric \(Tables[A16](https://arxiv.org/html/2609.19585#A9.T16)\-[A17](https://arxiv.org/html/2609.19585#A9.T17)\) to validate if generated summaries: \(A\) faithfully retained relevant mental\-health and social events without distortion and \(B\) placed them accurately in time\. See Appendix[K](https://arxiv.org/html/2609.19585#A11)for full metrics and results, and Appendix[L](https://arxiv.org/html/2609.19585#A12)for the internally hosted, custom web\-based evaluation platform\.

### 5\.5Qualitative Analysis

We qualitatively investigate what this reconstructed timeline helps a reader \(e\.g\., a clinician\) understand\. To examine the clinical content captured by the timelines, we analyze three clinician\-verified gold cases using two World Health Organization frameworks: eight categories of mental disorders and associated symptoms[World Health Organization \(2025a\)](https://arxiv.org/html/2609.19585#bib.bib10), and three levels of risk factors: Individual, Family & Community, and Structural[World Health Organization \(2025b\)](https://arxiv.org/html/2609.19585#bib.bib11)\. We label a risk factor as co\-documented rather than causal unless the record explicitly links it to the disorder\. Across three exemplar cases, temporal reconstruction distinguishes long\-standing risk factors from acute in\-stay changes and diagnostic reassessment, while preserving whether their relationships are causal, co\-documented, or unresolved rather than collapsing them into a single causal narrative\. Full case analyses are provided in Appendix[M](https://arxiv.org/html/2609.19585#A13)\.

## 6Evaluation with Open\-Weight Models

Table 3:Task A measured by embedding\-based RM\-F1@0\.85\.Silver→\\rightarrowSilvertrains and evaluates on the silver split;Silver→\\rightarrowGoldtrains on silver and evaluates on all5252clinician\-verified examples;Golduses four\-fold cross\-validation on the gold set alone\. Results are point estimates±\\pmbootstrap SD\.Table 4:Task B temporal tagging, scored by macro\-averagedF​1\\mathrm\{F\}1over the four temporal tags with the atomic events supplied as oracle input\. Results are point estimates±\\pmbootstrap SD\.Table 5:Task C timeline summarization measured by ROUGE\-L\. Models receive temporally tagged events and the source discharge summary as oracle input\. Results are point estimates±\\pmbootstrap SD\.We evaluate whetherCliniCIRCA’s intermediate outputs can serve as reusable supervision under three training and evaluation settings\.

Silver→\\rightarrowSilvermeasures how well a model fits pipeline\-generated supervision\. We sample1,0001\{,\}000admissions disjoint from the5252gold examples, label them with the frozen Stage 1–2 pipeline, and split them into750750training,100100validation, and150150held\-out test documents\. This Stage 1 to 3 generation processed61\.8461\.84M input and output tokens in total, averaging61,83661,836tokens per admission across the complete framework \(Appendix[A\.3](https://arxiv.org/html/2609.19585#A1.SS3)\)\.

Silver→\\rightarrowGoldmeasures whether that supervision transfers to human\-verified annotations: models trained on the same750750silver documents are evaluated on all5252gold examples\.

Goldmeasures what a small clinician\-verified set alone can support\. Because5252examples cannot sustain a stable held\-out split, we run four\-fold cross\-validation and pool the out\-of\-fold predictions before computing metrics and bootstrap intervals\. Throughout, the gold examples are excluded from silver labeling, silver\-trained adapters, and few\-shot exemplar selection to prevent unintentional leakage\. Under theGoldsetting, each document is evaluated only by the fold\-specific model for which it was held out\.

### 6\.1Experimental Setup

#### Models\.

We evaluate five locally deployable open\-weight models spanning dense and mixture\-of\-experts architectures,77B–3535B parameter scales, and general\-purpose versus clinical specialization: Gemma4\-31B[Gemma Team \(2026\)](https://arxiv.org/html/2609.19585#bib.bib76), Mistral\-24B[Mistral AI \(2025\)](https://arxiv.org/html/2609.19585#bib.bib77), Qwen3\.6\-35B\-A3B[Qwen Team \(2026\)](https://arxiv.org/html/2609.19585#bib.bib74), Qwen3\-30B\-A3B[Qwen Team \(2025\)](https://arxiv.org/html/2609.19585#bib.bib75), and Meditron\-7B[OpenMeditron \(2025\)](https://arxiv.org/html/2609.19585#bib.bib78)\. Qwen3\-30B\-A3B and Gemma4\-31B were not evaluated during Stage 2 model selection, allowing us to test whether observed performance extends beyond the models previously selected\.

#### Prompting and tuning\.

For each task, we compare three settings under identical task prompts: zero\-shot prompting, few\-shot prompting with three training exemplars, and LoRA\-based instruction tuning with one adapter per model and task\. Instruction tuning is intended to align models to the expected output format and decision rules rather than to inject new clinical knowledge\.

#### Metrics\.

We evaluate Task A using relaxed\-match F1 with a cosine similarity threshold of0\.850\.85\(RM\-F1@0\.85\) based on prior research[Sharif et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib61);[Zhang et al\. \(2019\)](https://arxiv.org/html/2609.19585#bib.bib62)\. Task B uses macro\-F1 over four temporal categories, with a prediction considered correct when both the temporal label and resolved ISO date or range match the reference\. Task C is evaluated using ROUGE\-L; ROUGE\-1 and ROUGE\-2 are reported in Appendix[N](https://arxiv.org/html/2609.19585#A14)\.

#### Statistical comparison\.

We report point estimates on the full evaluation set with standard deviations from2,0002\{,\}000document\-level bootstrap resamples\. For comparisons on the same evaluation set, resamples are paired across systems, and differences are assessed using95%95\\%confidence intervals for the paired score difference\. We describe a difference as reliable when its confidence interval excludes zero; otherwise, we make no claim that one system outperforms the other\.

### 6\.2Results

We examine if silver supervision improves open\-weight models, transfers to clinician\-verified annotations, and provides greater utility than the gold set alone\. Note that zero\-shot prompting requires no training fold, so its GOLD\-condition performance is identical to the zero\-shot SILVER→\\rightarrowGOLD condition; we therefore omit the redundant GOLD zero\-shot column in Tables[3](https://arxiv.org/html/2609.19585#S6.T3)\-[5](https://arxiv.org/html/2609.19585#S6.T5)\. Unless otherwise noted, comparisons below describe point\-estimate rankings; we flag a difference as statistically reliable only where the paired bootstrap confidence interval excludes zero, and treat other orderings as suggestive rather than established\.

#### Task A: Atomic event extraction\.

UnderSilver→\\rightarrowGold, instruction tuning yields higher RM\-F1@0\.85 scores than zero\-shot and few\-shot prompting for all five models \(Table[3](https://arxiv.org/html/2609.19585#S6.T3)\)\. Mistral\-24B achieves the highest point estimate at0\.8470\.847, compared with0\.2840\.284zero\-shot and0\.5450\.545few\-shot\. Silver\-trained models also retain similar point estimates when evaluation shifts from the silver test set to clinician\-verified annotations\. For example, Mistral\-24B scores0\.8480\.848underSilver→\\rightarrowSilverand0\.8470\.847underSilver→\\rightarrowGold\. These results show that benefits of silver supervision extend beyond evaluation against the generated labels\.

#### Task B: Temporal tagging\.

Temporal tagging shows a less uniform effect of instruction tuning \(Table[4](https://arxiv.org/html/2609.19585#S6.T4)\)\. UnderSilver→\\rightarrowGold, instruction tuning yields higher scores than both prompting baselines for four of the five models\. Mistral\-24B and Qwen3\.6\-35B\-A3B reach0\.7340\.734and0\.7310\.731, respectively\. Gemma4\-31B is the exception, achieving its highest score with few\-shot prompting \(0\.8040\.804\), compared with0\.7990\.799zero\-shot and0\.7490\.749after instruction tuning\.

#### Task C: Timeline summarization\.

UnderSilver→\\rightarrowGold, instruction tuning yields higher ROUGE\-L scores than both prompting baselines for all five models \(Table[5](https://arxiv.org/html/2609.19585#S6.T5)\)\. Qwen3\.6\-35B\-A3B achieves the highest point estimate at0\.7540\.754, followed by Gemma4\-31B at0\.7490\.749and Mistral\-24B at0\.7380\.738\. Meditron\-7B shows the largest increase, rising from0\.2930\.293zero\-shot and0\.2440\.244few\-shot to0\.6630\.663after instruction tuning\.

#### Silver supervision is more effective than gold\-only tuning at the current scale\.

Across all three tasks and all five models, instruction tuning on the silver dataset yields higher gold\-set point estimates than training on the gold folds alone\. For Task A, Mistral\-24B achieves0\.8470\.847underSilver→\\rightarrowGold, compared with0\.7150\.715underGold; Qwen3\-30B\-A3B scores0\.7950\.795and0\.6060\.606, respectively\. The same pattern appears in Task B\.

The gold\-only results should not be interpreted as evidence that the silver annotations are of higher quality\. Rather, they show that at the evaluated scale, broader coverage and larger volume of the silver dataset provide sufficient supervision for downstream models to learn the target tasks\. Gold annotations remain essential as an independent benchmark for assessing whether models trained on silver supervision generalize to human\-verified outputs\.

## 7Related Work

### 7\.1Temporal Annotation in Clinical NLP

Clinical temporal NLP has traditionally focused on extracting typed events, time expressions, and pairwise temporal relations in benchmarks such as i2b2, THYME, and n2c2[Styler et al\. \(2014\)](https://arxiv.org/html/2609.19585#bib.bib26);[Sun et al\. \(2013\)](https://arxiv.org/html/2609.19585#bib.bib25);[Uzuner et al\. \(2011\)](https://arxiv.org/html/2609.19585#bib.bib24)\. Later work extended this paradigm toward absolute event timing and temporal intervals[Leeuwenberg and Moens \(2020\)](https://arxiv.org/html/2609.19585#bib.bib27);[Frattallone\-Llado et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib20)\. These methods established the foundations of clinical temporal reasoning, but generally assume predefined event types or annotated mentions[Wang et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib4)\.CliniCIRCAinstead reconstructs an open\-vocabulary, patient\-level timeline directly from narrative text, grounding events to calendar dates or ranges while preserving temporal uncertainty\.

### 7\.2Constructing Event Timelines & Summaries from Clinical Narratives

Recent work has used LLMs to generate event time sequences from clinical narratives through either single\-pass extraction or multi\-stage reconstruction[Kumar et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib19);[Noroozizadeh et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib21);[Wang and Weiss \(2025\)](https://arxiv.org/html/2609.19585#bib.bib23)\. Existing approaches make this task tractable by imposing different forms of structure\. MedTimeline focuses on a predefined schema of chemotherapy events and does not model temporally vague events[Wang et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib4)\.[Frattallone\-Llado et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib20)and[Kumar et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib19)rely on structured EHR data to improve or calibrate event timing\. MIMIC\-IV\-Ext\-22MCTS applies chunking and retrieval to discharge summaries before generating events with relative timestamps[Wang et al\. \(2025a\)](https://arxiv.org/html/2609.19585#bib.bib3)\. In contrast,[Noroozizadeh et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib21)construct timelines from published case reports, which are typically more organized and curated than routine clinical narratives\. Together, these studies demonstrate the feasibility and downstream value of narrative timeline generation, but they focus on settings that are narrower, more structured, or more curated than ours\.

Clinical summarization addresses a complementary goal: condensing lengthy records into coherent accounts for review and decision\-making[Luo et al\. \(2023\)](https://arxiv.org/html/2609.19585#bib.bib49);[Lee et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib41);[Croxford et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib14)\. Most summarization systems generate directly from the source note, combining event selection, temporal reasoning, and narrative generation in a single step that makes the underlying chronology difficult to inspect\.CliniCIRCAinstead decouples these decisions: it first constructs an event\-level, uncertainty\-aware timeline and then summarizes from that representation\. The summary is therefore not an independent compression of the original note, but a readable realization of an auditable temporal substrate\.

## 8Broader Implications

Beyond demonstrating the feasibility of patient\-journey construction from discharge summaries, our findings reveal practical implications for clinical NLP and usage in real\-world contexts\.

#### Reusable supervision\.

Clinical institutions often have abundant unstructured records but limited resources for expert manual annotation, secure computing resources, and access to large proprietary models[Balasubramanian et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib65);[Ntinopoulos et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib5);[Grothey et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib64)\.CliniCIRCAoffers a practical route from raw text to task\-specific supervision: instruction tuning on only1,0001\{,\}000silver\-labeled examples improved every model on extraction and summarization, and four of five on temporal tagging\. These results suggest that modest pipeline\-generated datasets can support smaller, locally deployable models without requiring expert labeling of an entire clinical corpus or complex schema development, as has been the case in prior research \(Section[7\.2](https://arxiv.org/html/2609.19585#S7.SS2)\)\.

#### The necessity of human\-in\-the\-loop oversight\.

Reliable evaluation is particularly important in clinical NLP because infrequent errors can remain consequential even when aggregate performance appears high[Jiang et al\. \(2026\)](https://arxiv.org/html/2609.19585#bib.bib66);[Hager et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib48)\. Prior work in psychiatric NLP similarly finds that LLM outputs may resemble expert responses on surface\-level properties such as tone while diverging on clinically consequential dimensions[Chandra et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib7);[Wang et al\. \(2025b\)](https://arxiv.org/html/2609.19585#bib.bib63)\. Our automated judge agrees strongly with mental health clinicians on the large majority of correct events, but is less reliable in identifying the nature of rare errors, particularly those involving pre\-admission medications and date granularity\. Producing a trustworthy reference set therefore requires exhaustive clinician\-guided adjudication, since expert review remains necessary to resolve ambiguous temporal decisions and determine which errors are clinically meaningful[Bavaresco et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib32);[Diekmann et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib67);[Chen et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib68)\.

#### Supporting longitudinal chart review\.

Reviewing longitudinal records is time\- and effort\-intensive, especially for multimorbid patients whose histories could be distributed across multiple clinical encounters and narrative documents[Van Veen et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib69)\. Our proposed Stage 3 renders the event\-level timeline as a readable patient journey while retaining its calendar anchors and temporal uncertainty\. Such summaries could provide clinicians with a traceable overview before time\-constrained follow\-up visits, especially in contexts in which manually reconstructing a complex history from several discharge summaries may be impractical\. That said, we underscore that they should complement rather than replace the source record, with individual claims remaining connected to inspectable event\-level evidence\.

## 9Conclusion and Future Work

CliniCIRCAshows that recovering a patient journey from clinical narrative is best treated not as a single generation problem, but as separable representational decisions: what happened, when it happened, and how it should be rendered for review\. This decomposition makes the resulting chronology inspectable, exposes where temporal ambiguity remains unresolved, and enables each stage to be evaluated independently\. Our evaluation further shows that automated judges are useful for locating problematic outputs, but not for replacing expert adjudication of rare temporal errors\. At the same time, the reconstructed journeys provide effective supervision for smaller open\-weight models, suggesting that structured intermediate representations can serve not only as outputs, but also as reusable training resources\. Future work will extendCliniCIRCAto multi\-document patient records, richer temporal representations, and additional note types, to support traceable longitudinal chart review and memory\-aware clinical systems\.

## Limitations

Although we present a framework with clinician\-designed and verified definitions and steps, we recognize several areas for improvement\.

Representational scope\.Our date\-centered schema supports chronological ordering but may impose artificial timing on atemporal facts or persistent states, such as sex, family history, and long\-term substance use\. Future work should compare alternatives, including atemporal labels, validity intervals, and separate state/event categories, and evaluate their effects on timeline fidelity, summarization, and borderline cases\.

Summarization salience\.Stage 3 follows a clinician\-authored specification for retaining, combining, and omitting events in the narrative summary\. Because the same two clinicians developed this specification and conducted the faithfulness and temporal\-accuracy evaluation described in Section[5\.4](https://arxiv.org/html/2609.19585#S5.SS4), the results may understate disagreement in salience or interpretation\. Independent evaluation would strengthen this validation\. The specification also reflects one view of clinical relevance, since priorities may differ across clinicians and care settings\. The resulting summaries should therefore not be treated as a single canonical account of a patient’s journey\. Future work should include more diverse evaluators and clinical settings, salience policies, and multiple reference summaries\.

Psychiatric symptom representation in Stage 3\.Our Stage 3 specification \(Appendix[H](https://arxiv.org/html/2609.19585#A8)\) excludes affect descriptions with vital signs and routine examination findings to reduce boilerplate\. As a result, mood and affect trajectories – a core component of psychiatric symptom tracking, as illustrated by the depression case in Section[5\.5](https://arxiv.org/html/2609.19585#S5.SS5)– are preserved in the auditable Output A timeline \(e\.g\., via diagnosis and treatment events\) but not always carried into the reader\-facing Output B narrative in their own right\. Revisiting this exclusion for the mental\-health setting specifically, for example by retaining clinically significant affect changes as a distinct category rather than folding them under general exam findings, is a natural next step for future research\.

Single\-document scope\.CliniCIRCAreconstructs the course described within one discharge summary per admission rather than multiple notes from the full longitudinal record\. This bounded setting isolates narrative reconstruction, but cannot reconcile duplicated, conflicting, or evolving information across admissions\. Prior patient\-journey methods are not directly comparable because they use different inputs, prompts, and context assumptions\. We therefore compare 11 candidate models under the same framework \(Section[4](https://arxiv.org/html/2609.19585#S4)\), but do not include a single\-pass end\-to\-end baseline\. Future work should ablate the modular design using the same LLM and extend it to multi\-document records through cross\-note entity resolution, temporal alignment, and representation of revisions in clinical understanding\.

Despite these limitations, our proposed frameworkCliniCIRCAprovides a clinician\-validated representation for open\-vocabulary temporal reconstruction from raw clinical narratives\.

## Ethical Considerations

We use the MIMIC\-III Clinical Database \(v1\.4\), accessed through PhysioNet under its Credentialed Health Data Use Agreement \(v1\.5\.0\) with the required CITI training completed for all personnel with direct access or analysis of the data\. LLMs were used in compliance with PhysioNet’s policy on responsible use of MIMIC data with LLMs and online services\. We accessed Claude models through Amazon Bedrock and Gemini models through Vertex AI because the data were neither externally retained nor used for model training\. All human evaluation was performed exclusively by the authors, without the recruitment of external human subjects\. As a secondary analysis of de\-identified records, this work is not human\-subjects research and required no additional IRB review\.

Because source notes may unevenly document social history, trauma, or substance use across patients – for instance by diagnosis, demographic group, or documentation era –CliniCIRCA’s extraction and summarization stages could reproduce such documentation asymmetries rather than correct them; auditing this is an important direction for future work\. We also note that pipeline errors are not equally consequential: a mis\-tagged or fabricated event involving suicidality, self\-harm, or medication timing carries substantially more clinical risk than a mis\-tagged laboratory value\. Given this, we emphasize that Stage 3 outputs \(Output B\) are intended to complement, not replace, the source record, and should remain linked to their inspectable Stage 1\-2 source events \(Output A\) in any deployment setting within the healthcare system, particularly for safety\-relevant content\.

## Acknowledgments

Zhang, Ding, and De Choudhury were partly supported through funds from The Children’s Healthcare of Atlanta Pediatric Technology Center at Georgia Tech\. We thank Jiawei Zhou, Rijul Magu, Shravika Mittal, Viet Cuong Nguyen, Pinxian Lu, Owen Xingjian Zhang, Zeyu Hua, and Zikang Leng for valuable technical ideation and suggestions throughout this work\. We are grateful to Franklin Ye Ruan for discussions of clinical and medical knowledge that informed the design of this study\. We thank Sonakshi Sisodiya for early\-stage annotation, and Andrew Zhao for technical and server support\. We also thank Kaike Ping for statistical guidance and support\.

## References

- Aboueleninet al\.\(2025\)A\. Aboueleninet al\.Phi\-4\-Mini Technical Report: compact yet powerful multimodal language models via mixture\-of\-loras\.arXiv preprint arXiv:2503\.01743\.Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.12.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Alfattniet al\.\(2020\)G\. Alfattni, N\. Peek, and G\. NenadicExtraction of temporal relations from clinical free text: A systematic review of current approaches\.Journal of Biomedical Informatics108,pp\. 103488\(en\)\.External Links:ISSN 15320464,[Link](https://linkinghub.elsevier.com/retrieve/pii/S1532046420301167),[Document](https://dx.doi.org/10.1016/j.jbi.2020.103488)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p4.1),[§5\.1](https://arxiv.org/html/2609.19585#S5.SS1.p1.1)\.
- Amirahmadiet al\.\(2025\)A\. Amirahmadi, F\. Etminani, J\. Björk, O\. Melander, and M\. OhlssonTrajectory\-Ordered Objectives for Self\-Supervised Representation Learning of Temporal Healthcare Data Using Transformers: Model Development and Evaluation Study\.JMIR Medical Informatics13,pp\. e68138\(en\)\.External Links:ISSN 2291\-9694,[Link](https://medinform.jmir.org/2025/1/e68138),[Document](https://dx.doi.org/10.2196/68138)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- Anthropic \(2025\)AnthropicClaude Haiku 4\.5 System Card\.Note:System cardExternal Links:[Link](https://www-cdn.anthropic.com/7aad69bf12627d42234e01ee7c36305dc2f6a970.pdf)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.5.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Anthropic \(2026\)AnthropicClaude Sonnet 4\.6 System Card\.Note:System cardExternal Links:[Link](https://www-cdn.anthropic.com/78073f739564e986ff3e28522761a7a0b4484f84.pdf)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.4.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Balasubramanianet al\.\(2025\)J\. B\. Balasubramanian, D\. Adams, I\. Roxanis, A\. B\. De Gonzalez, P\. Coulson, J\. S\. Almeida, and M\. García\-ClosasLeveraging large language models for structured information extraction from pathology reports\.Journal of Pathology Informatics19,pp\. 100521\(en\)\.External Links:ISSN 21533539,[Link](https://linkinghub.elsevier.com/retrieve/pii/S2153353925001075),[Document](https://dx.doi.org/10.1016/j.jpi.2025.100521)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px1.p1.1)\.
- Bavarescoet al\.\(2025\)A\. Bavaresco, R\. Bernardi, L\. Bertolazzi, D\. Elliott, R\. Fernández, A\. Gatt, E\. Ghaleb, M\. Giulianelli, M\. Hanna, A\. Koller, A\. Martins, P\. Mondorf, V\. Neplenbroek, S\. Pezzelle, B\. Plank, D\. Schlangen, A\. Suglia, A\. K\. Surikuchi, E\. Takmaz, and A\. TestoniLLMs instead of human judges? a large scale empirical study across 20 NLP evaluation tasks\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 238–255\.External Links:[Link](https://aclanthology.org/2025.acl-short.20/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-short.20),ISBN 979\-8\-89176\-252\-7Cited by:[Appendix E](https://arxiv.org/html/2609.19585#A5.p1.1),[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- Cardet al\.\(2020\)D\. Card, P\. Henderson, U\. Khandelwal, R\. Jia, K\. Mahowald, and D\. JurafskyWith Little Power Comes Great Responsibility\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),Online,pp\. 9263–9274\(en\)\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.745),[Document](https://dx.doi.org/10.18653/v1/2020.emnlp-main.745)Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p1.1)\.
- Centers for Medicare & Medicaid Services and National Center for Health Statistics \(2011\)Centers for Medicare & Medicaid Services and National Center for Health StatisticsICD\-9\-CM Official Guidelines for Coding and Reporting\.Note:Effective October 1, 2011U\.S\. Department of Health and Human ServicesExternal Links:[Link](https://www.cdc.gov/nchs/data/icd/icd9cm_guidelines_2011.pdf)Cited by:[§A\.1](https://arxiv.org/html/2609.19585#A1.SS1.p1.1),[Table A1](https://arxiv.org/html/2609.19585#A1.T1)\.
- Champely \(2020\)S\. ChampelyPwr: basic functions for power analysis\.Note:R package version 1\.3\-0External Links:[Link](https://cran.r-project.org/package=pwr)Cited by:[Appendix C](https://arxiv.org/html/2609.19585#A3.p9.1)\.
- Chandraet al\.\(2025\)M\. Chandra, S\. Sriraman, G\. Verma, H\. S\. Khanuja, J\. S\. Campayo, Z\. Li, M\. L\. Birnbaum, and M\. De ChoudhuryLived Experience Not Found: LLMs Struggle to Align with Experts on Addressing Adverse Drug Reactions from Psychiatric Medication Use\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),Albuquerque, New Mexico,pp\. 11083–11113\(en\)\.External Links:[Link](https://aclanthology.org/2025.naacl-long.553),[Document](https://dx.doi.org/10.18653/v1/2025.naacl-long.553)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- Chaturvediet al\.\(2025\)R\. Chaturvedi, P\. Baghershahi, S\. Medya, and B\. Di EugenioTemporal relation extraction in clinical texts: a span\-based graph transformer approach\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 25765–25788\.External Links:[Link](https://aclanthology.org/2025.acl-long.1251/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.1251),ISBN 979\-8\-89176\-251\-0Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p4.1)\.
- Chaturvedi \(2024\)R\. ChaturvediTemporal Knowledge Graph Extraction and Modeling across Multiple Documents for Health Risk Prediction\.InCompanion Proceedings of the ACM Web Conference 2024,Singapore Singapore,pp\. 1182–1185\(en\)\.External Links:ISBN 979\-8\-4007\-0172\-6,[Link](https://dl.acm.org/doi/10.1145/3589335.3651256),[Document](https://dx.doi.org/10.1145/3589335.3651256)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- Chenet al\.\(2024\)G\. H\. Chen, S\. Chen, Z\. Liu, F\. Jiang, and B\. WangHumans or LLMs as the Judge? A Study on Judgement Bias\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Miami, Florida, USA,pp\. 8301–8327\(en\)\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.474),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.474)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- Cohen \(1988\)J\. CohenStatistical power analysis for behavioral sciences\.2nd edition,Lawrence Erlbaum Associates,Hillsdale, NY\.Cited by:[Appendix C](https://arxiv.org/html/2609.19585#A3.p1.1)\.
- Comaniciet al\.\(2025\)G\. Comanici, E\. Bieber, M\. Schaekermann, I\. Pasupat, N\. Sachdeva, I\. Dhillon, M\. Blistein, O\. Ram, D\. Zhang, E\. Rosen, L\. Marris, S\. Petulla, C\. Gaffney, A\. Aharoni, N\. Lintz, T\. C\. Pais, H\. Jacobsson, I\. Szpektor, N\. Jiang, K\. Haridasan, A\. Omran, N\. Saunshi, D\. Bahri, G\. Mishra, E\. Chu, T\. Boyd, B\. Hekman, A\. Parisi, C\. Zhang, K\. Kawintiranon, T\. Bedrax\-Weiss, O\. Wang, Y\. Xu, O\. Purkiss, U\. Mendlovic, I\. Deutel, N\. Nguyen, A\. Langley, F\. Korn, L\. Rossazza, A\. Ramé, S\. Waghmare, H\. Miller, N\. Byrd, A\. Sheshan, R\. Hadsell, S\. Bhardwaj, P\. Janus, T\. Rissa, D\. Horgan, A\. Abdagic, L\. Belenki, J\. Allingham, A\. Singh, T\. Guidroz, S\. Srinivasan, H\. Schmit, K\. Chiafullo, A\. Elisseeff, N\. Jha, P\. Kolhar, L\. Berrada, F\. Ding, X\. Si, S\. B\. Mallick, F\. Och, S\. Erell, E\. Ni, T\. Latkar, S\. Yang, P\. Sirkovic, Z\. Feng, R\. Leland, R\. Hornung, G\. Wu, C\. Blundell, H\. Alvari, P\. Huang, C\. Yip, S\. Deur, L\. Liu, G\. Surita, P\. Duque, D\. Damen, J\. Jia, A\. Guez, M\. Mircea, A\. Sinha, A\. Magni, P\. Stradomski, T\. Marian, V\. Galić, W\. Chen, H\. Husain, A\. Singhal, D\. Grewe, F\. Aubet, S\. Song, L\. Blanco, L\. Rechis, L\. Ho, R\. Munoz, K\. Zheng, J\. Hamrick, K\. Mather, H\. Taitelbaum, E\. Rutherford, Y\. Lei, K\. Chen, A\. Shukla, E\. Moreira, E\. Doi, B\. Isik, N\. Shabat, D\. Rogozińska, K\. Kolipaka, J\. Chang, E\. Vušak, S\. Venkatachary, S\. Noghabi, T\. Bharti, Y\. Jun, A\. Zaks, S\. Green, J\. Challagundla, W\. Wong, M\. Mohammad, D\. Hirsch, Y\. Cheng, I\. Naim, L\. Proleev, D\. Vincent, A\. Singh, M\. Krikun, D\. Krishnan, Z\. Ghahramani, A\. Atias, R\. Aggarwal, C\. Kirov, D\. Vytiniotis, C\. Koh, A\. Chronopoulou, P\. Dogra, V\. Ion, G\. Tyen, J\. Lee, F\. Weissenberger, T\. Strohman, A\. Balakrishna, J\. Rae, M\. Velic, R\. de Liedekerke, O\. Elyada, W\. Yuan, C\. Liu, L\. Shani, S\. Kishchenko, B\. Alessio, Y\. Li, R\. Song, S\. Kwei, O\. Jankowski, A\. Pappu, Y\. Namiki, Y\. Ma, N\. Tripuraneni, C\. Cherry, M\. Ikonomidis, Y\. Ling, C\. Ji, B\. Westberg, A\. Wright, D\. Yu, D\. Parkinson, S\. Ramaswamy, J\. Connor, S\. H\. Yeganeh, S\. Grover, G\. Kenwright, L\. Litchev, C\. Apps, A\. Tomala, F\. Halim, A\. Castro\-Ros, Z\. Li, A\. Boral, P\. Sho, M\. Yarom, E\. Malmi, D\. Klinghoffer, R\. Lin, A\. Ansell, P\. K\. S, S\. Zhao, S\. Zuo, A\. Santoro, H\. Cheng, S\. Demmessie, Y\. Liu, N\. Brichtova, A\. Culp, N\. Braun, D\. Graur, W\. Ng, N\. Mehta, A\. Phillips, P\. Sundberg, V\. Godbole, F\. Liu, Y\. Katariya, D\. Rim, M\. Seyedhosseini, S\. Ammirati, J\. Valfridsson, M\. Malihi, T\. Knight, A\. Toor, T\. Lampe, A\. Ittycheriah, L\. Chiang, C\. Yeung, A\. Fréchette, J\. Rao, H\. Wang, H\. Srivastava, R\. Zhang, R\. Rhodes, A\. Brand, D\. Weesner, I\. Figotin, F\. Gimeno, R\. Fellinger, P\. Marcenac, J\. Leal, E\. Marcus, V\. Cotruta, R\. Cabrera, S\. Luo, D\. Garrette, V\. Axelrod, S\. Baltateanu, D\. Barker, D\. Chen, H\. Toma, B\. Ingram, J\. Riesa, C\. Kulkarni, Y\. Zhang, H\. Liu, C\. Wang, M\. Polacek, W\. Wu, K\. Hui, A\. N\. Reyes, Y\. Su, M\. Barnes, I\. Malhi, A\. Siddiqui, Q\. Feng, M\. Damaschin, D\. Pighin, A\. Steiner, S\. Yang, R\. S\. Boppana, S\. Ivanov, A\. Kandoor, A\. Shah, A\. Mujika, D\. Huang, C\. A\. Choquette\-Choo, M\. Patel, T\. Yu, T\. Creswell, Jerry, Liu, C\. Barros, Y\. Razeghi, A\. Roy, P\. Culliton, B\. Xiong, J\. Pan, T\. Strohmann, T\. Powell, B\. Seal, D\. DeCarlo, P\. Shyam, K\. Katircioglu, X\. Wang, C\. Hardin, I\. Odisho, J\. Broder, O\. Chang, A\. Nair, A\. Shtefan, M\. O’Brien, M\. Agarwal, S\. Potluri, S\. Goyal, A\. Jhindal, S\. Thakur, Y\. Stuken, J\. Lyon, K\. Toutanova, F\. Feng, A\. Wu, B\. Horn, A\. Wang, A\. Cullum, G\. Taubman, D\. Shrivastava, C\. Shi, H\. Tomlinson, R\. Patel, T\. Tu, A\. M\. Oflazer, F\. Pongetti, M\. Yang, A\. A\. Taïga, V\. Perot, N\. W\. Pierse, F\. Han, Y\. Drori, I\. Iturrate, A\. Chakrabarti, L\. Yeung, D\. Dopson, Y\. Chen, A\. Kulshreshtha, T\. Guo, P\. Pham, T\. Schuster, J\. Chen, A\. Polozov, J\. Xing, H\. Zhou, P\. Kacham, D\. Kukliansky, A\. Miech, S\. Yaroshenko, E\. Chi, S\. Douglas, H\. Fei, M\. Blondel, P\. Myla, L\. Madmoni, X\. Wu, D\. Keysers, K\. Kjems, I\. Albuquerque, L\. Yu, J\. D’sa, M\. Plantan, V\. Ionescu, J\. S\. Elias, A\. Gupta, M\. R\. Vuyyuru, F\. Alcober, T\. Zhou, K\. Ji, F\. Hartmann, S\. Puttagunta, H\. Song, E\. Amid, A\. Stefanoiu, A\. Lee, P\. Pucciarelli, E\. Wang, A\. Raul, S\. Petrov, I\. Tian, V\. Anklin, N\. Nti, V\. Gomes, M\. Schumacher, G\. Vesom, A\. Panagopoulos, K\. Bousmalis, D\. Andor, J\. Jacob, Y\. Zhang, B\. Rosgen, M\. Kecman, M\. Tung, A\. Belias, N\. Goodman, P\. Covington, B\. Wieder, N\. Saxena, E\. Davoodi, M\. Huang, S\. Maddineni, V\. Roulet, F\. Campbell\-Ajala, P\. G\. Sessa, Xintian, Wu, G\. Lai, P\. Collins, A\. Haig, V\. Sakenas, X\. Xu, M\. Giustina, L\. E\. Shafey, P\. Charoenpanit, S\. Garg, J\. Ainslie, B\. Severson, M\. G\. Arenas, S\. Pathak, S\. Rajayogam, J\. Feng, M\. Bakker, S\. Li, N\. Wichers, J\. Rogers, X\. Geng, Y\. Li, R\. Jagerman, C\. Jia, N\. Olmert, D\. Sharon, M\. Mauger, S\. Mariserla, H\. Ma, M\. Mohabey, K\. Kim, A\. Andreev, S\. Pollom, J\. Love, V\. Jain, P\. Agrawal, Y\. Schroecker, A\. Fortin, M\. Warmuth, J\. Liu, A\. Leach, I\. Blok, G\. P\. Girirajan, R\. Aharoni, B\. Uria, A\. Sozanschi, D\. Goldberg, L\. Ionita, M\. T\. Ribeiro, M\. Zlocha, V\. Birodkar, S\. Lachgar, L\. Yuan, H\. Choudhury, M\. Ginsberg, F\. Zheng, G\. Dibb, E\. Graves, S\. Lokhande, G\. Rasskin, G\. Muraru, C\. Quick, S\. Tata, P\. Sermanet, A\. Chawla, I\. Karo, Y\. Wang, S\. Zhang, O\. Keller, A\. Dragan, G\. Su, I\. Chou, X\. Liu, Y\. Tao, S\. Prabhakara, M\. Wilson, R\. Liu, S\. Wang, G\. Evans, D\. Du, A\. Castaño, G\. Prasad, M\. E\. Mahdy, S\. Gerlach, M\. Reid, J\. Kahn, A\. Zait, T\. S\. Pillai, T\. Ulrich, G\. Wang, J\. Wassenberg, E\. Farkash, K\. Yalasangi, C\. Wang, M\. Bauza, S\. Bucher, T\. Liu, J\. Yan, G\. Leung, V\. Sindhwani, P\. Barnes, A\. Singh, I\. Jurin, J\. Chang, N\. K\. Bhumihar, S\. Eiger, G\. Citovsky, B\. Withbroe, Z\. Li, S\. Xue, N\. D\. Santo, G\. Stoyanov, Y\. Raimond, S\. Zheng, Y\. Gao, V\. Listík, S\. Kwasiborski, R\. Saputro, A\. Ozturel, G\. Mallya, K\. Majmundar, R\. West, P\. Caron, J\. Wei, L\. Castrejon, S\. Vikram, D\. Ramachandran, N\. Dhawan, J\. Park, S\. Smoot, G\. van den Driessche, Y\. Blau, C\. Malik, W\. Liang, R\. Hirsch, C\. N\. dos Santos, E\. Weinstein, A\. van den Oord, S\. Lall, N\. FitzGerald, Z\. Jiang, X\. Yang, D\. Webster, A\. Elqursh, A\. Pope, G\. Rotival, D\. Raposo, W\. Zhu, J\. Dean, S\. Alabed, D\. Tran, A\. Gupta, Z\. Gleicher, J\. Austin, E\. Rosseel, M\. Umekar, D\. Das, Y\. Sun, K\. Chen, K\. Misiunas, X\. Zhou, Y\. Di, A\. Loo, J\. Newlan, B\. Li, V\. Ramasesh, Y\. Xu, A\. Chen, S\. Gandhe, R\. Soricut, N\. Gupta, S\. Hu, S\. El\-Sayed, X\. Garcia, I\. Brusilovsky, P\. Chen, A\. Bolt, L\. Huang, A\. Gurney, Z\. Zhang, A\. Pritzel, J\. Wilkiewicz, B\. Seybold, B\. K\. Shamanna, F\. Fischer, J\. Dean, K\. Gill, R\. Mcilroy, A\. Bhowmick, J\. Selier, A\. Yang, D\. Cheng, V\. Magay, J\. Tan, D\. Varma, C\. Walder, T\. Kocisky, R\. Nakashima, P\. Natsev, M\. Kwong, I\. Gog, C\. Zhang, S\. Dieleman, T\. Jimma, A\. Ryabtsev, S\. Brahma, D\. Steiner, D\. Du, A\. Žužul, M\. Žanić, M\. Raghavachari, W\. Gierke, Z\. Zheng, D\. Petrova, Y\. Dauphin, Y\. Liu, I\. Kessler, S\. Hand, C\. Duvarney, S\. Kim, H\. Lee, L\. Hussenot, J\. Hui, J\. Smith, D\. Jain, J\. Xia, G\. S\. Tomar, K\. Amiri, D\. Phan, F\. Fuchs, T\. Weyand, N\. Tomasev, A\. Cordell, X\. Liu, J\. Mallinson, P\. Joshi, A\. Crawford, A\. Suggala, S\. Chien, N\. Fernando, M\. Sanchez\-Vargas, D\. Williams, P\. Crone, X\. Luo, I\. Karpov, J\. Shan, T\. Thurk, R\. Strudel, P\. Voigtlaender, P\. Patil, T\. Dozat, A\. Khodaei, S\. Singla, P\. Ambroszczyk, Q\. Wu, Y\. Chang, B\. Roark, C\. Hegde, T\. Ding, A\. Filos, Z\. Wu, A\. S\. Pinto, S\. Liu, S\. Khanna, A\. Pandey, S\. Mcloughlin, Q\. Li, S\. Haves, A\. Zhou, E\. Buchatskaya, I\. Leal, P\. de Boursac, N\. Akazawa, N\. Anderson, T\. Chen, K\. Somandepalli, C\. Liang, S\. Goenka, S\. Winkler, A\. Grushetsky, Y\. Ding, J\. Smith, F\. Ye, J\. Pont\-Tuset, E\. Li, R\. Li, T\. Golany, D\. Wegner, T\. Jiang, O\. Barak, Y\. Shangguan, E\. Vértes, R\. Wong, J\. Bornschein, A\. Tudor, M\. Bevilacqua, T\. Schaul, A\. S\. Rawat, Y\. Zhao, K\. Axiotis, L\. Meng, C\. McLean, J\. Lai, J\. Beattie, N\. Kushman, Y\. Liu, B\. Kutzman, F\. Lang, J\. Ye, P\. Netrapalli, P\. Mishra, M\. Khan, M\. Goel, R\. Willoughby, D\. Tian, H\. Zhuang, J\. Chen, Z\. Tsai, T\. Kementsietsidis, A\. Khare, J\. Keeling, K\. Xu, N\. Waters, F\. Altché, A\. Popat, B\. Mittal, D\. Saxton, D\. E\. Badawy, M\. Mathieu, Z\. Zheng, H\. Zhou, N\. Ranka, R\. Shin, Q\. Duan, T\. Salimans, I\. Mihailescu, U\. Shaham, M\. Chang, Y\. Assael, N\. Dikkala, M\. Izzard, V\. Cohen\-Addad, C\. Graves, V\. Feinberg, G\. Chung, D\. Strouse, D\. Karmon, S\. Sharifzadeh, Z\. Ashwood, K\. Pham, J\. Blanton, A\. Vasiloff, J\. Barber, M\. Geller, A\. Zhou, F\. Zubach, T\. Huang, L\. Zhang, H\. Gupta, M\. Young, J\. Proskurnia, R\. Votel, V\. Gabeur, G\. Barcik, A\. Tripathi, H\. Yu, G\. Yan, B\. Changpinyo, F\. Pavetić, A\. Coyle, Y\. Fujii, J\. G\. Mendez, T\. Zhou, H\. Rajamani, B\. Hechtman, E\. Cao, D\. Juan, Y\. Tan, V\. Dalibard, Y\. Du, N\. Clay, K\. Yao, W\. Jia, D\. Vijaykumar, Y\. Zhou, X\. Bai, W\. Hung, S\. Pecht, G\. Todorov, N\. Khadke, P\. Gupta, P\. Lahoti, A\. Autef, K\. Duddu, J\. Lee\-Thorp, A\. Bykovsky, T\. Misiunas, S\. Flennerhag, S\. Thangaraj, J\. McGiffin, Z\. Nado, M\. Kunesch, A\. Noever, A\. Hertz, M\. Liang, V\. Stone, E\. Palmer, S\. Daruki, A\. Pramanik, S\. Põder, A\. Kyker, M\. Khan, E\. Sluzhaev, M\. Ritter, A\. Ruderman, W\. Zhou, C\. Nagpal, K\. Vodrahalli, G\. Necula, P\. Barham, E\. Pavlick, J\. Hartford, I\. Shafran, L\. Zhao, M\. Mikuła, T\. Eccles, H\. Shimokawa, K\. Garg, L\. Vilnis, H\. Chen, I\. Shumailov, K\. Lee, A\. Abdelhamed, M\. Xie, V\. Cohen, E\. Hlavnova, D\. Malkin, C\. Sitawarin, J\. Lottes, P\. Coquinot, T\. Yu, S\. Kumar, J\. Zhang, A\. Mahendru, Z\. Ahmed, J\. Martens, T\. Chen, A\. Boag, D\. Peng, C\. Devin, A\. Klimovskiy, M\. Phuong, D\. Vainstein, J\. Xie, B\. Ramabhadran, N\. Howard, X\. Yu, G\. Goswami, J\. Cui, S\. Shleifer, M\. Pinto, C\. Yeh, M\. Yang, S\. Javanmardi, D\. Ethier, C\. Lee, J\. Orbay, S\. Kotecha, C\. Bromberg, P\. Shaw, J\. Thornton, A\. G\. Rosenthal, S\. Gu, M\. Thomas, I\. Gemp, A\. Ayyar, A\. Ushio, A\. Selvan, J\. Wee, C\. Liu, M\. Majzoubi, W\. Yu, J\. Abernethy, T\. Liechty, R\. Pan, H\. Nguyen, Qiong, Hu, S\. Perrin, A\. Arora, E\. Pitler, W\. Wang, K\. Shivakumar, F\. Prost, B\. Limonchik, J\. Wang, Y\. Gao, T\. Cour, S\. Buch, H\. Gui, M\. Ivanova, P\. Neubeck, K\. Chan, L\. Kim, H\. Chen, N\. Goyal, D\. Chung, L\. Liu, Y\. Su, A\. Petrushkina, J\. Shen, A\. Joulin, Y\. Xu, S\. X\. Lin, Y\. Kulizhskaya, C\. Chelba, S\. Vasudevan, E\. Collins, V\. Bashlovkina, T\. Lu, D\. Fritz, J\. Park, Y\. Zhou, C\. Su, R\. Tanburn, M\. Sushkov, M\. Rasquinha, J\. Li, J\. Prendki, Y\. Li, P\. LV, S\. Sharma, H\. Fitoussi, H\. Huang, A\. Dai, P\. Dao, M\. Burrows, H\. Prior, D\. Qin, G\. Pundak, L\. L\. Sjoesund, A\. Khurshudov, Z\. Zhu, A\. Webson, E\. Kemp, T\. Tan, S\. Agrawal, S\. Sargsyan, L\. Cheng, J\. Stephan, T\. Kwiatkowski, D\. Reid, A\. Byravan, A\. H\. Michaely, N\. Heess, L\. Zhou, S\. Goenka, V\. Carpenter, A\. Levskaya, B\. Wang, R\. Roberts, R\. Leblond, S\. Chikkerur, S\. Ginzburg, M\. Chang, R\. Riachi, Chuqiao, Xu, Z\. Borsos, M\. Pliskin, J\. Pawar, M\. Lustman, H\. Kirkwood, A\. Anand, A\. Chaudhary, N\. Kalb, K\. Milan, S\. Augenstein, A\. Goldie, L\. Prince, K\. Raman, Y\. Sun, V\. Xia, A\. Cohen, Z\. Huo, J\. Camp, S\. Ellis, L\. Zilka, D\. V\. Torres, L\. Patel, S\. Arora, B\. Chan, J\. Adler, K\. Ayoub, J\. Liang, F\. Jamil, J\. Jiang, S\. Baumgartner, H\. Sun, Y\. Karov, Y\. Akulov, H\. Zheng, I\. Cai, C\. Fantacci, J\. Rubin, A\. R\. Acha, M\. Wang, N\. D’Souza, R\. Sathyanarayana, S\. Dai, S\. Rowe, A\. Simanovsky, O\. Goldman, Y\. Kuang, X\. Pan, A\. Rosenberg, T\. Rojas\-Esponda, P\. Dutta, A\. Zeng, I\. Jurenka, G\. Farquhar, Y\. Bansal, S\. Iqbal, B\. Roelofs, G\. Joung, P\. Beak, C\. Ryu, R\. Poplin, Y\. Wu, J\. Alayrac, S\. Buthpitiya, O\. Ronneberger, C\. Habtegebriel, W\. Li, P\. Cavallaro, A\. Wei, G\. Bensky, T\. Denk, H\. Ganapathy, J\. Stanway, P\. Joshi, F\. Bertolini, J\. Lo, O\. Ma, Z\. Charles, G\. Sampemane, H\. Sahni, X\. Chen, H\. Askham, D\. Gaddy, P\. Young, J\. Tan, M\. Eyal, A\. Bražinskas, L\. Zhong, Z\. Wu, M\. Epstein, K\. Bailey, A\. Hard, K\. Lee, S\. Goldshtein, A\. Ruiz, M\. Badawi, M\. Lochbrunner, J\. Kearns, A\. Brown, F\. Pardo, T\. Weber, H\. Yang, P\. Jiang, B\. Akin, Z\. Fu, M\. Wainwright, C\. Zou, M\. Gaba, P\. Manzagol, W\. Kan, Y\. Song, K\. Zainullina, R\. Lin, J\. Ko, S\. Deshmukh, A\. Jindal, J\. Svensson, D\. Tyam, H\. Zhao, C\. Kaeser\-Chen, S\. Baird, P\. Moradi, J\. Hall, Q\. Guo, V\. Tsang, B\. Liang, F\. Pereira, S\. Ganesh, I\. Korotkov, J\. Adamek, S\. Thiagarajan, V\. Tran, C\. Chen, C\. Tar, S\. Jain, I\. Dasgupta, T\. Bilal, D\. Reitter, K\. Zhao, G\. Vezzani, Y\. Gehman, P\. Mehta, L\. Beltrone, X\. Dotiwalla, S\. Guadarrama, Z\. Abbas, S\. Karp, P\. Georgiev, C\. Ferng, M\. Brockschmidt, L\. Peng, C\. Hirnschall, V\. Verma, Y\. Bi, Y\. Xiao, A\. Dabush, K\. Xu, P\. Wallis, R\. Parker, Q\. Wang, Y\. Xu, I\. Safarli, D\. Tewari, Y\. Zhang, S\. Kim, A\. Gesmundo, M\. Thomas, S\. Levi, A\. Chowdhury, K\. Rao, P\. Garst, S\. Conway\-Rahman, H\. Ran, K\. McKinney, Z\. Xiao, W\. Yu, R\. Agrawal, A\. Stjerngren, C\. Ionescu, J\. Chen, V\. Sharma, J\. Chiu, F\. Liu, K\. Franko, C\. Sanford, X\. Cai, P\. Michel, S\. Ganapathy, J\. Labanowski, Z\. Garrett, B\. Vargas, S\. Sun, B\. Gale, T\. Buschmann, G\. Desjardins, N\. Ghelani, P\. Jain, M\. Verma, C\. Asawaroengchai, J\. Eisenschlos, J\. Harlalka, H\. Kazawa, D\. Metzler, J\. Howland, Y\. Jian, J\. Ades, V\. Shah, T\. Gangwani, S\. Lee, R\. Ring, S\. M\. Hernandez, D\. Reich, A\. Sinha, A\. Sathe, J\. Kovac, A\. Gill, A\. Kannan, A\. D’olimpio, M\. Sevenich, J\. Whang, B\. Kim, K\. C\. Sim, J\. Chen, J\. Zhang, S\. Lall, Y\. Matias, B\. Jia, A\. Friesen, S\. Nasso, A\. Thapliyal, B\. Perozzi, T\. Yu, A\. Shekhawat, S\. Huda, P\. Grabowski, E\. Wang, A\. Sreevatsa, H\. Dib, M\. Hassen, P\. Schuh, V\. Milutinovic, C\. Welty, M\. Quinn, A\. Shah, B\. Wang, G\. Barth\-Maron, J\. Frye, N\. Axelsson, T\. Zhu, Y\. Ma, I\. Giannoumis, H\. Sedghi, C\. Ye, Y\. Luan, K\. Aydin, B\. Chandra, V\. Sampathkumar, R\. Huang, V\. Lavrenko, A\. Eleryan, Z\. Hong, S\. Hansen, S\. M\. Carthy, B\. Samanta, D\. Ćevid, X\. Wang, F\. Li, M\. Voznesensky, M\. Hoffman, A\. Terzis, V\. Sehwag, G\. Fidel, L\. He, M\. Cai, Y\. He, A\. Feng, M\. Nikoltchev, S\. Phatale, J\. Chase, R\. Lawton, M\. Zhang, T\. Ouyang, M\. Tragut, M\. H\. Manshadi, A\. Narayanan, J\. Shen, X\. Gao, T\. Bolukbasi, N\. Roy, X\. Li, D\. Golovin, L\. Panait, Z\. Qin, G\. Han, T\. Anthony, S\. Kudugunta, V\. Patraucean, A\. Ray, X\. Chen, X\. Yang, T\. Bhatia, P\. Talluri, A\. Morris, A\. Ražnatović, B\. Brownfield, J\. An, S\. Peng, P\. Kane, C\. Zheng, N\. Duduta, J\. Kessinger, J\. Noraky, S\. Liu, K\. Rong, P\. Veličković, K\. Rush, A\. Goldin, F\. Wei, S\. M\. R\. Garlapati, C\. Pantofaru, O\. Kwon, J\. Ni, E\. Noland, J\. D\. Trapani, F\. Beaufays, A\. G\. Roy, Y\. Chow, A\. Turker, G\. Cideron, L\. Mei, J\. Clark, Q\. Dou, M\. Bošnjak, R\. Leith, Y\. Du, A\. Yazdanbakhsh, M\. Nasr, C\. Kwak, S\. S\. Sheth, A\. Kaskasoli, A\. Anand, B\. Lakshminarayanan, S\. Jerome, D\. Bieber, C\. Chu, A\. Senges, T\. Shen, M\. Sridhar, N\. Ndebele, B\. Beyret, S\. Mohamed, M\. Chen, M\. Freitag, J\. Guo, L\. Liu, P\. Roit, H\. Chen, S\. Yan, T\. Stone, J\. Co\-Reyes, J\. Cole, S\. Scellato, S\. Azizi, H\. Hashemi, A\. Jin, A\. Iyer, M\. Valentine, A\. György, A\. Ahuja, D\. H\. Diaz, C\. Lee, N\. Clement, W\. Kong, D\. Garmon, I\. Watts, K\. Bhatia, K\. Gupta, M\. Miecnikowski, H\. Vallet, A\. Taly, E\. Loper, S\. Joshi, J\. Atwood, J\. Chick, M\. Collier, F\. Iliopoulos, R\. Trostle, B\. Gunel, R\. Leal\-Cavazos, A\. M\. Hrafnkelsson, M\. Guzman, X\. Ju, A\. Forbes, J\. Emond, K\. Chauhan, B\. Caine, L\. Xiao, W\. Zeng, A\. Moufarek, D\. Murphy, M\. Meng, N\. Gupta, F\. Riedel, A\. Das, E\. Lawal, S\. Narayan, T\. Sosea, J\. Swirhun, L\. Friso, B\. Neyshabur, J\. Lu, S\. Girgin, M\. Wunder, E\. Yvinec, A\. Pyne, V\. Carbune, S\. Rijhwani, Y\. Guo, T\. Doshi, A\. Briukhov, M\. Bain, A\. Hitron, X\. Wang, A\. Gupta, K\. Chen, C\. Du, W\. Zhang, D\. Shah, A\. Akula, M\. Dylla, A\. Kachra, W\. Kuo, T\. Zou, L\. Wang, L\. Xu, J\. Zhu, J\. Snyder, S\. Menon, O\. Firat, I\. Mordatch, Y\. Yuan, N\. Ponomareva, R\. Blevins, L\. Moore, W\. Wang, P\. Chen, M\. Scholz, A\. Dwornik, J\. Lin, S\. Li, D\. Antognini, T\. I, X\. Song, M\. Miller, U\. Kalra, A\. Raveret, O\. Akerlund, F\. Wu, A\. Nystrom, N\. Godbole, T\. Liu, H\. DeBalsi, J\. Zhao, B\. Liu, A\. Caciularu, L\. Lax, U\. Khandelwal, V\. Langston, E\. Bailey, S\. Lattanzi, Y\. Wang, N\. Kovelamudi, S\. Mondal, G\. Guruganesh, N\. Hua, O\. Roval, P\. Wesołowski, R\. Ingale, J\. Halcrow, T\. Sohn, C\. Angermueller, B\. Raad, E\. Stickgold, E\. Lu, A\. Kosik, J\. Xie, T\. Lillicrap, A\. Huang, L\. L\. Zhang, D\. Paulus, C\. Farabet, A\. Wertheim, B\. Wang, R\. Joshi, C\. Ko, Y\. Wu, S\. Agrawal, L\. Lin, X\. Sheng, P\. Sung, T\. Breland\-King, C\. Butterfield, S\. Gawde, S\. Singh, Q\. Zhang, R\. Apte, S\. Shetty, A\. Hutter, T\. Li, E\. Salesky, F\. Lebron, J\. Kanerva, M\. Paganini, A\. Nguyen, R\. Vallu, J\. Peter, S\. Velury, D\. Kao, J\. Hoover, A\. Bortsova, C\. Bishop, S\. Jakobovits, A\. Agostini, A\. Agarwal, C\. Liu, C\. Kwong, S\. Tavakkol, I\. Bica, A\. Greve, A\. GP, J\. Marcus, L\. Hou, T\. Duerig, R\. Moroshko, D\. Lacey, A\. Davis, J\. Amelot, G\. Wang, F\. Kim, T\. Strinopoulos, H\. Wan, C\. L\. Lan, S\. Krishnan, H\. Tang, P\. Humphreys, J\. Bai, I\. H\. Shtacher, D\. Machado, C\. Pang, K\. Burke, D\. Liu, R\. Aravamudhan, Y\. Song, E\. Hirst, A\. Singh, B\. Jou, L\. Bai, F\. Piccinno, C\. K\. Fu, R\. Alazard, B\. Meiri, D\. Winter, C\. Chen, M\. Zhang, J\. Heitkaemper, J\. Lambert, J\. Lee, A\. Frömmgen, S\. Rogulenko, P\. Nair, P\. Niemczyk, A\. Bulyenov, B\. Xu, H\. Shemtov, M\. Zadimoghaddam, S\. Toropov, M\. Wirth, H\. Dai, S\. Gollapudi, D\. Zheng, A\. Kurakin, C\. Lee, K\. Bullard, N\. Serrano, I\. Balazevic, Y\. Li, J\. Schalkwyk, M\. Murphy, M\. Zhang, K\. Sequeira, R\. Datta, N\. Agrawal, C\. Sutton, N\. Attaluri, M\. Chiang, W\. Farhan, G\. Thornton, K\. Lin, T\. Choma, H\. Nguyen, K\. Dasgupta, D\. Robinson, I\. Comşa, M\. Riley, A\. Pillai, B\. Mustafa, B\. Golan, A\. Zandieh, J\. Lespiau, B\. Porter, D\. Ross, S\. Rajayogam, M\. Agarwal, S\. Venugopalan, B\. Shahriari, Q\. Yan, H\. Xu, T\. Tobin, P\. Dubov, H\. Shi, A\. Recasens, A\. Kovsharov, S\. Borgeaud, L\. Dery, S\. Vasanth, E\. Gribovskaya, L\. Qiu, M\. Mahdieh, W\. Skut, E\. Nielsen, C\. Zheng, A\. Yu, C\. G\. Bostock, S\. Gupta, A\. Archer, C\. Rawles, E\. Davies, A\. Svyatkovskiy, T\. Tsai, Y\. Halpern, C\. Reisswig, B\. Wydrowski, B\. Chang, J\. Puigcerver, M\. H\. Taege, J\. Li, E\. Schnider, X\. Li, D\. Dena, Y\. Xu, U\. Telang, T\. Shi, H\. Zen, K\. Kastner, Y\. Ko, N\. Subramaniam, A\. Kumar, P\. Blois, Z\. Dai, J\. Wieting, Y\. Lu, Y\. Zeldes, T\. Xie, A\. Hauth, A\. Ţifrea, Y\. Li, S\. El\-Husseini, D\. Abolafia, H\. Zhou, W\. Ding, S\. Ghalebikesabi, C\. Guía, A\. Maksai, Á\. Weisz, S\. Arik, N\. Sukhanov, A\. Świetlik, X\. Jia, L\. Yu, W\. Wang, M\. Brand, D\. Bloxwich, S\. Kirmani, Z\. Chen, A\. Go, P\. Sprechmann, N\. Kannen, A\. Carin, P\. Sandhu, I\. Edkins, L\. Nooteboom, J\. Gupta, L\. Maggiore, J\. Azizi, Y\. Pritch, P\. Yin, M\. Gupta, D\. Tarlow, D\. Smith, D\. Ivanov, M\. Babaeizadeh, A\. Goel, S\. Kambala, G\. Chu, M\. Kastelic, M\. Liu, H\. Soltau, A\. Stone, S\. Agrawal, M\. Kim, K\. Soparkar, S\. Tadepalli, O\. Bunyan, R\. Soh, A\. Kannan, D\. Kim, B\. J\. Chen, A\. Halumi, S\. Roy, Y\. Wang, O\. Sercinoglu, G\. Gibson, S\. Bhatnagar, M\. Sano, D\. von Dincklage, Q\. Ren, B\. Mitrevski, M\. Olšák, J\. She, C\. Doersch, Jilei, Wang, B\. Liu, Q\. Tan, T\. Yakar, T\. Warkentin, A\. Ramirez, C\. Lebsack, J\. Dillon, R\. Mathews, T\. Cobley, Z\. Wu, Z\. Chen, J\. Simon, S\. Nath, T\. Sainath, A\. Bendebury, R\. Julian, B\. Mankalale, D\. Ćurko, P\. Zacchello, A\. R\. Brown, K\. Sodhia, H\. Howard, S\. Caelles, A\. Gupta, G\. Evans, A\. Bulanova, L\. Katzen, R\. Goldenberg, A\. Tsitsulin, J\. Stanton, B\. Schillings, V\. Kovalev, C\. Fry, R\. Shah, K\. Lin, S\. Upadhyay, C\. Li, S\. Radpour, M\. Maggioni, J\. Xiong, L\. Haas, J\. Brennan, A\. Kamath, N\. Savinov, A\. Nagrani, T\. Yacovone, R\. Kappedal, K\. Andriopoulos, L\. Lao, Y\. Li, G\. Rozhdestvenskiy, K\. Hashimoto, A\. Audibert, S\. Austin, D\. Rodriguez, A\. Ruoss, G\. Honke, D\. Karkhanis, X\. Xiong, Q\. Wei, J\. Huang, Z\. Leng, V\. Premachandran, S\. Bileschi, G\. Evangelopoulos, T\. Mensink, J\. Pavagadhi, D\. Teplyashin, P\. Chang, L\. Xue, G\. Tanzer, S\. Goldman, K\. Patel, S\. Li, J\. Wiesner, I\. Zheng, I\. Stewart\-Binks, J\. Han, Z\. Li, L\. Luo, K\. Lenc, M\. Lučić, F\. Xue, R\. Mullins, A\. Guseynov, C\. Chang, I\. Galatzer\-Levy, A\. Zhang, G\. Bingham, G\. Hu, A\. Hartman, Y\. Ma, J\. Griffith, A\. Irpan, C\. Radebaugh, S\. Yue, L\. Fan, V\. Ungureanu, C\. Sorokin, H\. Teufel, P\. Li, R\. Anil, D\. Paparas, T\. Wang, C\. Lin, H\. Peng, M\. Shum, G\. Petrovic, D\. Brady, R\. Nguyen, K\. Macherey, Z\. Li, H\. Singh, M\. Yenugula, M\. Iinuma, X\. Chen, K\. Kopparapu, A\. Stern, S\. Dave, C\. Thekkath, F\. Perot, A\. Kumar, F\. Li, Y\. Xiao, M\. Bilotti, M\. H\. Bateni, I\. Noble, L\. Lee, A\. Vázquez\-Reina, J\. Salazar, X\. Yang, B\. Wang, E\. Gruzewska, A\. Rao, S\. Raghuram, Z\. Xu, E\. Ben\-David, J\. Mei, S\. Dalmia, Z\. Zhang, Y\. Liu, G\. Bansal, H\. Pankov, S\. Schwarcz, A\. Burns, C\. Chan, S\. Sanghai, R\. Liang, E\. Liang, A\. He, A\. Stuart, A\. Narayanan, Y\. Zhu, C\. Frank, B\. Fatemi, A\. Sabne, O\. Lang, I\. Bhattacharya, S\. Settle, M\. Wang, B\. McMahan, A\. Tacchetti, L\. B\. Soares, M\. Hadian, S\. Cabi, T\. Chung, N\. Putikhin, G\. Li, J\. Chen, A\. Tarango, H\. Michalewski, M\. Kazemi, H\. Masoom, H\. Sheftel, R\. Shivanna, A\. Vadali, R\. Comanescu, D\. Reid, J\. Moore, A\. Neelakantan, M\. Sander, J\. Herzig, A\. Rosenberg, M\. Dehghani, J\. Choi, M\. Fink, R\. Hayes, E\. Ge, S\. Weng, C\. Ho, J\. Karro, K\. Krishna, L\. N\. Thiet, A\. Skerry\-Ryan, D\. Eppens, M\. Andreetto, N\. Sarma, S\. Bonacina, B\. K\. Ayan, M\. Nawhal, Z\. Shan, M\. Dusenberry, S\. Thakoor, S\. Gubbi, D\. D\. Nguyen, R\. Tsarfaty, S\. Albanie, J\. Mitrović, M\. Gandhi, B\. Chen, A\. Epasto, G\. Stephanov, Y\. Jin, S\. Gehman, A\. Amini, J\. Weber, F\. Behbahani, S\. Xu, M\. Allamanis, X\. Chen, M\. Ott, C\. Sha, M\. Jastrzebski, H\. Qi, D\. Greene, X\. Wu, A\. Toki, D\. Vlasic, J\. Shapiro, R\. Kotikalapudi, Z\. Shen, T\. Saeki, S\. Xie, A\. Cassirer, S\. Bharadwaj, T\. Kiyono, S\. Bhojanapalli, E\. Rosenfeld, S\. Ritter, J\. Mao, J\. G\. Oliveira, Z\. Egyed, B\. Bandemer, E\. Parisotto, K\. Kinoshita, J\. Pluto, P\. Maniatis, S\. Li, Y\. Guo, G\. Ghiasi, J\. Tarbouriech, S\. Chatterjee, J\. Jin, Katrina, Xu, J\. Palomaki, S\. Arnold, M\. Sewak, F\. Piccinini, M\. Sharma, B\. Albrecht, S\. Purser\-haskell, A\. Vaswani, C\. Chen, M\. Wisniewski, Q\. Cao, J\. Aslanides, N\. M\. Phu, M\. Sieb, L\. Agubuzu, A\. Zheng, D\. Sohn, M\. Selvi, A\. Andreassen, K\. Subudhi, P\. Eruvbetine, O\. Woodman, T\. Mery, S\. Krause, X\. Ren, X\. Ma, J\. Luo, D\. Chen, W\. Fan, H\. Griffiths, C\. Schuler, A\. Li, S\. Zhang, J\. Sarr, S\. Luo, R\. Patana, M\. Watson, D\. Naboulsi, M\. Collins, S\. Sidhwani, E\. Hoogeboom, S\. Silver, E\. Caveness, X\. Zhao, M\. Rodriguez, M\. Deines, L\. Bai, P\. Griffin, M\. Tagliasacchi, E\. Xue, S\. R\. Babbula, B\. Pang, N\. Ding, G\. Shen, E\. Peake, R\. Crocker, S\. S\. Raghvendra, D\. Swisher, W\. Han, R\. Singh, L\. Wu, V\. Pchelin, T\. Munkhdalai, D\. Alon, G\. Bacon, E\. Robles, J\. Bulian, M\. Johnson, G\. Powell, F\. T\. Ferreira, Y\. Li, F\. Benzing, M\. Velimirović, H\. Soyer, W\. Kong, Tony, Nguyên, Z\. Yang, J\. Liu, J\. van Amersfoort, D\. Gillick, B\. Sun, N\. Rauschmayr, K\. Zhang, S\. Zhan, T\. Zhou, A\. Frolov, C\. Yang, D\. Vnukov, L\. Rouillard, H\. Li, A\. Mandhane, N\. Fallen, R\. Venkataraman, C\. H\. Hu, J\. Brennan, J\. Lee, J\. Chang, M\. Sundermeyer, Z\. Pan, R\. Ke, S\. Tong, A\. Fabrikant, W\. Bono, J\. Gu, R\. Foley, Y\. Mao, M\. Delakis, D\. Bhaswar, R\. Frostig, N\. Li, A\. Zipori, C\. Hope, O\. Kozlova, S\. Mishra, J\. Djolonga, C\. Schiff, M\. A\. Merey, E\. Briakou, P\. Morgan, A\. Wan, A\. Hassidim, R\. Skerry\-Ryan, K\. Sengupta, M\. Jasarevic, P\. Kallakuri, P\. Kunkle, H\. Brennan, T\. Lieber, H\. Mansoor, J\. Walker, B\. Zhang, A\. Xie, G\. Žužić, A\. Chukwuka, A\. Druinsky, D\. Cho, R\. Yao, F\. Naeem, S\. Butt, E\. Kim, Z\. Jia, M\. Jordan, A\. Lelkes, M\. Kurzeja, S\. Wang, J\. Zhao, A\. Over, A\. Chakladar, M\. Prasetya, N\. Jha, S\. Ganapathy, Y\. Cong, P\. Shroff, C\. Saroufim, S\. Miryoosefi, M\. Hammad, T\. Nasir, W\. Xi, Y\. Gao, Y\. Maeng, B\. Hora, C\. Cheng, P\. Haghani, Y\. Lewenberg, C\. Lu, M\. Matysiak, N\. Raisinghani, H\. Wang, L\. Baugher, R\. Sukthankar, M\. Giang, J\. Schultz, N\. Fiedel, M\. Chen, C\. Lee, T\. Dey, H\. Zheng, S\. Paul, C\. Smith, A\. Ly, Y\. Wang, R\. Bansal, B\. Perz, S\. Ricco, S\. Blank, V\. Keshava, D\. Sharma, M\. Chow, K\. Lad, K\. Jalan, S\. Osindero, C\. Swanson, J\. Scott, A\. Ilić, X\. Li, S\. R\. Jonnalagadda, A\. S\. Soudagar, Y\. Xiong, B\. Batsaikhan, D\. Jarrett, N\. Kumar, M\. Shah, M\. Lawlor, A\. Waters, M\. Graham, R\. May, S\. Ramos, S\. Lefdal, Z\. Cankara, N\. Cano, B\. O’Donoghue, J\. Borovik, F\. Liu, J\. Grimstad, M\. Alnahlawi, K\. Tsihlas, T\. Hudson, N\. Grigorev, Y\. Jia, T\. Huang, T\. P\. Igwe, S\. Lebedev, X\. Tang, I\. Krivokon, F\. Garcia, M\. Tan, E\. Jia, P\. Stys, S\. Vashishth, Y\. Liang, B\. Venkatraman, C\. Gu, A\. Kementsietsidis, C\. Zhu, J\. Jung, Y\. Bai, M\. J\. Hosseini, F\. Ahmed, A\. Gupta, X\. Yuan, S\. Ashraf, S\. Nigam, G\. Vasudevan, P\. Awasthi, A\. M\. Gilady, Z\. Mariet, R\. Eskander, H\. Li, H\. Hu, G\. Garrido, P\. Schlattner, G\. Zhang, R\. Saxena, P\. Dević, K\. Muralidharan, A\. Murthy, Y\. Zhou, M\. Choi, A\. Wongpanich, Z\. Wang, P\. Shah, Y\. Xu, Y\. Huang, S\. Spencer, A\. Chen, J\. Cohan, J\. Wang, J\. Tompson, J\. Wu, R\. Haroun, H\. Li, B\. Huergo, F\. Yang, T\. Yin, J\. Wendt, M\. Bendersky, R\. Chaabouni, J\. Snaider, J\. Ferret, A\. Jindal, T\. Thompson, A\. Xue, W\. Bishop, S\. M\. Phal, A\. Sharma, Y\. Sung, P\. Radhakrishnan, M\. Shomrat, R\. Ingle, R\. Vij, J\. Gilmer, M\. D\. Istin, S\. Sobell, Y\. Lu, E\. Nottage, D\. Sadigh, J\. Willcock, T\. Zhang, S\. Xu, S\. Brown, K\. Lee, G\. Wang, Y\. Zhu, Y\. Tay, C\. Kim, A\. Gutierrez, A\. Sharma, Y\. Xian, S\. Seo, C\. Cui, E\. Pochernina, C\. Baetu, K\. Jastrzębski, M\. Ly, M\. Elhawaty, D\. Suh, E\. Sezener, P\. Wang, N\. Yuen, G\. Tucker, J\. Cai, Z\. Yang, C\. Wang, A\. Muzio, H\. Qian, J\. Yoo, D\. Lockhart, K\. R\. McKee, M\. Guo, M\. Mehrotra, A\. Mendonça, S\. V\. Mehta, S\. Ben, C\. Tekur, J\. Mu, M\. Zhu, V\. Krakovna, H\. Lee, A\. Maschinot, S\. Cevey, H\. Choe, A\. Bai, H\. Srinivasan, D\. Gasaway, N\. Young, P\. Siegler, D\. Holtmann\-Rice, V\. Piratla, K\. Baumli, R\. Yogev, A\. Hofer, H\. van Hasselt, S\. Grant, Y\. Chervonyi, D\. Silver, A\. Hogue, A\. Agarwal, K\. Wang, P\. Singh, F\. Flynn, J\. Lipschultz, R\. David, L\. Bellot, Y\. Yang, L\. Le, F\. Graziano, K\. Olszewska, K\. Hui, A\. Maurya, N\. Parotsidis, W\. Chen, T\. Oguntebi, J\. Kelley, A\. Baddepudi, J\. Mauerer, G\. Shaw, A\. Siegman, L\. Yang, S\. Shetty, S\. Roy, Y\. Song, W\. Stokowiec, R\. Burnell, O\. Savant, R\. Busa\-Fekete, J\. Miao, S\. Ghosh, L\. MacDermed, P\. Lippe, M\. Dektiarev, Z\. Behrman, F\. Mentzer, K\. Nguyen, M\. Wei, S\. Verma, C\. Knutsen, S\. Dasari, Z\. Yan, P\. Mitrichev, X\. Wang, V\. Shejwalkar, J\. Austin, S\. Sunkara, N\. Potti, Y\. Virin, C\. Wright, G\. Liu, O\. Riva, E\. Pot, G\. Kochanski, Q\. Le, G\. Balasubramaniam, A\. Dhar, Y\. Liao, A\. Bloniarz, D\. Shukla, E\. Cole, J\. Lee, S\. Zhang, S\. Kafle, S\. Vashishtha, P\. Mahmoudieh, G\. Chen, R\. Hoffmann, P\. Srinivasan, A\. D\. Lago, Y\. B\. Shalom, Z\. Wang, M\. Elabd, A\. Sharma, J\. Oh, S\. Kothawade, M\. Le, M\. Monteiro, S\. Yang, K\. Alarakyia, R\. Geirhos, D\. Mincu, H\. Garnes, H\. Kobayashi, S\. Mariooryad, K\. Krasowiak, Zhixin, Lai, S\. Mourad, M\. Wang, F\. Bu, O\. Aharoni, G\. Chen, A\. Goyal, V\. Zubov, A\. Bapna, E\. Dabir, N\. Kothari, K\. Lamerigts, N\. D\. Cao, J\. Shar, C\. Yew, N\. Kulkarni, D\. Mahaarachchi, M\. Joshi, Z\. Zhu, J\. Lichtarge, Y\. Zhou, H\. Muckenhirn, V\. Selo, O\. Vinyals, P\. Chen, A\. Brohan, V\. Mehta, S\. Cogan, R\. Wang, T\. Geri, W\. Ko, W\. Chen, F\. Viola, K\. Shivam, L\. Wang, M\. C\. Elish, R\. A\. Popa, S\. Pereira, J\. Liu, R\. Koster, D\. Kim, G\. Zhang, S\. Ebrahimi, P\. Talukdar, Y\. Zheng, P\. Poklukar, A\. Mikhalap, D\. Johnson, A\. Vijayakumar, M\. Omernick, M\. Dibb, A\. Dubey, Q\. Hu, A\. Suman, V\. Aggarwal, I\. Kornakov, F\. Xia, W\. Lowe, A\. Kolganov, T\. Xiao, V\. Nikolaev, S\. Hemingray, B\. Li, J\. Iljazi, M\. Rybiński, B\. Sandhu, P\. Lu, T\. Luong, R\. Jenatton, V\. Govindaraj, Hui, Li, G\. Dulac\-Arnold, W\. Park, H\. Wang, A\. Modi, J\. Pouget\-Abadie, K\. Greller, R\. Gupta, R\. Berry, P\. Ramachandran, J\. Xie, L\. McCafferty, J\. Wang, K\. Gupta, H\. Lim, B\. Bratanič, A\. Brock, I\. Akolzin, J\. Sproch, D\. Karliner, D\. Kim, A\. Goedeckemeyer, N\. Shazeer, C\. Schmid, D\. Calandriello, P\. Bhatia, K\. Choromanski, C\. Montgomery, D\. Dua, A\. Ramalho, H\. King, Y\. Gao, L\. Nguyen, D\. Lindner, D\. Pitta, O\. Johnson, K\. Salama, D\. Ardila, M\. Han, E\. Farnese, S\. Odoom, Z\. Wang, X\. Ding, N\. Rink, R\. Smith, H\. T\. Lehri, E\. Cohen, N\. Vats, T\. He, P\. Gopavarapu, A\. Paszke, M\. Patel, W\. V\. Gansbeke, L\. Loher, L\. Castro, M\. Voitovich, T\. von Glehn, N\. George, S\. Niklaus, Z\. Eaton\-Rosen, N\. Rakićević, E\. Jue, S\. Perel, C\. Zhang, Y\. Bahat, A\. Pouget, Z\. Xing, F\. Huot, A\. Shenoy, T\. Bos, V\. Coriou, B\. Richter, N\. Noy, Y\. Wang, S\. Ontanon, S\. Qin, G\. Makarchuk, D\. Hassabis, Z\. Li, M\. Sharma, K\. Venkatesan, I\. Kemaev, R\. Daniel, S\. Huang, S\. Shah, O\. Ponce, Warren, Chen, M\. Faruqui, J\. Wu, S\. Andačić, S\. Payrits, D\. McDuff, T\. Hume, Y\. Cao, M\. Tessler, Q\. Wang, Y\. Wang, I\. Rendulic, E\. Agustsson, M\. Johnson, T\. Lando, A\. Howard, S\. G\. S\. Padmanabhan, M\. Daswani, A\. Banino, M\. Kilgore, J\. Heek, Z\. Ji, A\. Caceres, C\. Li, N\. Kassner, A\. Vlaskin, Z\. Liu, A\. Grills, Y\. Hou, R\. Sukkerd, G\. Cheon, N\. Shetty, L\. Markeeva, P\. Stanczyk, T\. Iyer, Y\. Gong, S\. Gao, K\. Gopalakrishnan, T\. Blyth, M\. Reynolds, A\. Bhoopchand, M\. Bilenko, D\. Gharibian, V\. Zayats, A\. Faust, A\. Singh, M\. Ma, H\. Jiao, S\. Vijayanarasimhan, L\. Aroyo, V\. Yadav, S\. Chakera, A\. Kakarla, V\. Meshram, K\. Gregor, G\. Botea, E\. Senter, D\. Jia, G\. Kovacs, N\. Sharma, S\. Baur, K\. Kang, Y\. He, L\. Zhuo, M\. Kostelac, I\. Laish, S\. Peng, L\. O’Bryan, D\. Kasenberg, G\. R\. Rao, E\. Leurent, B\. Zhang, S\. Stevens, A\. Salazar, Y\. Zhang, I\. Lobov, J\. Walker, A\. Porter, M\. Redshaw, H\. Ke, A\. Rao, A\. Lee, H\. Lam, M\. Moffitt, J\. Kim, S\. Qiao, T\. Koo, R\. Dadashi, X\. Song, M\. Sundararajan, P\. Xu, C\. Kawamoto, Y\. Zhong, C\. Barbu, A\. Reddy, M\. Verzetti, L\. Li, G\. Papamakarios, H\. Klimczak\-Plucińska, M\. Cassin, K\. Kavukcuoglu, R\. Swavely, A\. Vaucher, J\. Zhao, R\. Hemsley, M\. Tschannen, H\. Ge, G\. Menghani, Y\. Yu, N\. Ha, W\. He, X\. Wu, M\. Song, R\. Sterneck, S\. Zinke, D\. A\. Calian, A\. Marsden, A\. C\. Ruiz, M\. Hessel, A\. Gueta, B\. Lee, B\. Farris, M\. Gupta, Y\. Li, M\. Saleh, V\. Misra, K\. Xiao, P\. Mendolicchio, G\. Buttimore, V\. Krayvanova, N\. Nayakanti, M\. Wiethoff, Y\. Pande, A\. Mirhoseini, N\. Lao, J\. Liu, Y\. Hua, A\. Chen, Y\. Malkov, D\. Kalashnikov, S\. Gupta, K\. Audhkhasi, Y\. Zhai, S\. Kopalle, P\. Jain, E\. Ofek, C\. Meyer, K\. Baatarsukh, H\. Strejček, J\. Qian, J\. Freedman, R\. Figueira, M\. Sokolik, O\. Bachem, R\. Lin, D\. Kharrat, C\. Hidey, P\. Xu, D\. Duan, Y\. Li, M\. Ersoy, R\. Everett, K\. Cen, R\. Santamaria\-Fernandez, A\. Taubenfeld, I\. Mackinnon, L\. Deng, P\. Zablotskaia, S\. Viswanadha, S\. Goel, D\. Yates, Y\. Deng, P\. Choy, M\. Chen, A\. Sinha, A\. Mossin, Y\. Wang, A\. Szlam, S\. Hao, P\. K\. Rubenstein, M\. Toksoz\-Exley, M\. Aperghis, Y\. Zhong, J\. Ahn, M\. Isard, O\. Lacombe, F\. Luisier, C\. Anastasiou, Y\. Kalley, U\. Prabhu, E\. Dunleavy, S\. Bijwadia, J\. Mao\-Jones, K\. Chen, R\. Pasumarthi, E\. Wood, A\. Dostmohamed, N\. Hurley, J\. Simsa, A\. Parrish, M\. Pajarskas, M\. Harvey, O\. Skopek, Y\. Kochinski, J\. Rey, V\. Rieser, D\. Zhou, S\. J\. Lee, T\. Acharya, G\. Li, J\. Jiang, X\. Zhang, B\. Gipson, E\. Mahintorabi, M\. Gelmi, N\. Khajehnouri, A\. Yeh, K\. Lee, L\. Matthey, L\. Baker, T\. Pham, H\. Fu, A\. Pak, P\. Gupta, C\. Vasconcelos, A\. Sadovsky, B\. Walker, S\. Hsiao, P\. Zochbauer, A\. Marzoca, N\. Velan, J\. Zeng, G\. Baechler, D\. Driess, D\. Jain, Y\. Huang, L\. Tao, J\. Maggs, N\. Levine, J\. Schneider, E\. Gemzer, S\. Petit, S\. Han, Z\. Fisher, D\. Zelle, C\. Biles, E\. Ie, A\. Fadeeva, C\. Liu, J\. V\. Franco, A\. Collister, H\. Zhang, R\. Wang, R\. Zhao, L\. Kieliger, K\. Shuster, R\. Zhu, B\. Gong, L\. Chan, R\. Sun, S\. Basu, R\. Zimmermann, J\. Hayes, A\. Bapna, J\. Snoek, W\. Yang, P\. Datta, J\. A\. Abdallah, K\. Kilgour, L\. Li, S\. Mah, Y\. Jun, M\. Rivière, A\. Karmarkar, T\. Spalink, T\. Huang, L\. Gonzalez, D\. Tran, A\. Nowak, J\. Palowitch, M\. Chadwick, E\. Talius, H\. Mehta, T\. Sellam, P\. Fränken, M\. Nicosia, K\. He, A\. Kini, D\. Amos, S\. Basu, H\. Jobe, E\. Shaw, Q\. Xu, C\. Evans, D\. Ikeda, C\. Yan, L\. Jin, L\. Wang, S\. Yadav, I\. Labzovsky, R\. Sampath, A\. Ma, C\. Schumann, A\. Siddhant, R\. Shah, J\. Youssef, R\. Agarwal, N\. Dabney, A\. Tonioni, M\. Ambar, J\. Li, I\. Guyon, B\. Li, D\. Soergel, B\. Fang, G\. Karadzhov, C\. Udrescu, T\. Trinh, V\. Raunak, S\. Noury, D\. Guo, S\. Gupta, M\. Finkelstein, D\. Petek, L\. Liang, G\. Billock, P\. Sun, D\. Wood, Y\. Song, X\. Yu, T\. Matejovicova, R\. Cohen, K\. Andra, D\. D’Ambrosio, Z\. Deng, V\. Nallatamby, E\. Songhori, R\. Dangovski, A\. Lampinen, P\. Botadra, A\. Hillier, J\. Cao, N\. Baddi, A\. Kuncoro, T\. Yoshino, A\. Bhagatwala, M\. Ranzato, R\. Schaeffer, T\. Liu, S\. Ye, O\. Sarvana, J\. Nham, C\. Kuang, I\. Gao, J\. Baek, S\. Mittal, A\. Wahid, A\. Gergely, B\. Ni, J\. Feldman, C\. Muir, P\. Lamblin, W\. Macherey, E\. Dyer, L\. Kilpatrick, V\. Campos, M\. Bhutani, S\. Fort, Y\. Ahmad, A\. Severyn, K\. Chatziprimou, O\. Ferludin, M\. Dimarco, A\. Kusupati, J\. Heyward, D\. Bahir, K\. Villela, K\. Millican, D\. Marcus, S\. Bahargam, C\. Unlu, N\. Roth, Z\. Wei, S\. Gopal, D\. Ghoshal, E\. Lee, S\. Lin, J\. Lees, D\. Lee, A\. Hosseini, C\. Fan, S\. Neel, M\. Wu, Y\. Altun, H\. Cai, E\. Piqueras, J\. Woodward, A\. Bissacco, S\. Haykal, M\. Bordbar, P\. Sundaram, S\. Hodkinson, D\. Toyama, G\. Polovets, A\. Myers, A\. Sinha, T\. Levinboim, K\. Krishnakumar, R\. Chhaparia, T\. Sholokhova, N\. B\. Gundavarapu, G\. Jawahar, H\. Qureshi, J\. Hu, N\. Momchev, M\. Rahtz, R\. Wu, A\. P\. S, K\. Dhamdhere, M\. Guo, U\. Gupta, A\. Eslami, M\. Schain, M\. Blokzijl, D\. Welling, D\. Orr, L\. Bolelli, N\. Perez\-Nieves, M\. Sirotenko, A\. Prasad, A\. Kar, B\. D\. B\. Pigem, T\. Terzi, G\. Weisz, D\. Ghosh, A\. Mavalankar, D\. Madeka, K\. Daugaard, H\. Adam, V\. Shah, D\. Berman, M\. Tran, S\. Baker, E\. Andrejczuk, G\. Chole, G\. Raboshchuk, M\. Mirzazadeh, T\. Kagohara, S\. Wu, C\. Schallhart, B\. Orlando, C\. Wang, A\. Rrustemi, H\. Xiong, H\. Liu, A\. Vezer, N\. Ramsden, S\. Chang, S\. Mudgal, Y\. Li, N\. Vieillard, Y\. Hoshen, F\. Ahmad, A\. Slone, A\. Hua, N\. Potikha, M\. Rossini, J\. Stritar, S\. Prakash, Z\. Wang, X\. Dong, A\. Nazari, E\. Nehoran, K\. Tekelioglu, Y\. Li, K\. Badola, T\. Funkhouser, Y\. Li, V\. Yerram, R\. Ganeshan, D\. Formoso, K\. Langner, T\. Shi, H\. Li, Y\. Yamamori, A\. Panda, A\. Saade, A\. S\. Scarpati, C\. Breaux, C\. Carey, Z\. Zhou, C\. Hsieh, S\. Bridgers, A\. Butryna, N\. Gupta, V\. Tulsyan, S\. Woo, E\. Eltyshev, W\. Grathwohl, C\. Parks, S\. Benjamin, R\. Panigrahy, S\. Dodhia, D\. D\. Freitas, C\. Sauer, W\. Song, F\. Alet, J\. Tolins, C\. Paduraru, X\. Zhou, B\. Albert, Z\. Zhang, L\. Shu, M\. Bansal, S\. Nguyen, A\. Globerson, O\. Xiao, J\. Manyika, T\. Hennigan, R\. Rong, J\. Matak, A\. Bakalov, A\. Sharma, D\. Sinopalnikov, A\. Pierson, S\. Roller, G\. Brown, M\. Gao, T\. Fukuzawa, A\. Ghafouri, K\. Vassigh, I\. Barr, Z\. Wang, A\. Korsun, R\. Jayaram, L\. Ren, T\. Zaman, S\. Khan, Y\. Lunts, D\. Deutsch, D\. Uthus, N\. Katz, M\. Samsikova, A\. Khalifa, N\. Sethi, J\. Sun, L\. Tang, U\. Alon, X\. Luo, D\. Yu, A\. Nayyar, B\. Petrini, W\. Truong, V\. Hellendoorn, N\. Chinaev, C\. Alberti, W\. Wang, J\. Hu, V\. Mirrokni, A\. Balashankar, A\. Aharon, A\. Mehta, A\. Iscen, J\. Kready, L\. Manning, A\. Mohananey, Y\. Chen, A\. Tripathi, A\. Wu, I\. Petrovski, D\. Hwang, M\. Baeuml, S\. Chandrakaladharan, Y\. Liu, R\. Coaguila, M\. Chen, S\. Ma, P\. Tafti, S\. Tatineni, T\. Spitz, J\. Ye, P\. Vicol, M\. Rosca, A\. Puigdomènech, Z\. Yahav, S\. Ghemawat, H\. Lin, P\. Kirk, Z\. Nabulsi, S\. Brin, B\. Bohnet, K\. Caluwaerts, A\. S\. Veerubhotla, D\. Zheng, Z\. Dai, P\. Petrov, Y\. Xu, R\. Mehran, Z\. Xu, L\. Zintgraf, J\. Choi, S\. A\. Hombaiah, R\. Thoppilan, S\. Reddi, L\. Lew, L\. Li, K\. Webster, K\. Sawhney, L\. Lamprou, S\. Shakeri, M\. Lunayach, J\. Chen, S\. Bagri, A\. Salcianu, Y\. Chen, Y\. Donchev, C\. Magister, S\. Nørly, V\. Rodrigues, T\. Izo, H\. Noga, J\. Zou, T\. Köppe, W\. Zhou, K\. Lee, X\. Long, D\. Eisenbud, A\. Chen, C\. Schenck, C\. M\. To, P\. Zhong, E\. Taropa, M\. Truong, O\. Levy, D\. Martins, Z\. Zhang, C\. Semturs, K\. Zhang, A\. Yakubovich, P\. Moreno, L\. McConnaughey, D\. Lu, S\. Redmond, L\. Weerts, Y\. Bitton, T\. Refice, N\. Lacasse, A\. Conmy, C\. Tallec, J\. Odell, H\. Forbes\-Pollard, A\. Socala, J\. Hoech, P\. Kohli, A\. Walton, R\. Wang, M\. Sazanovich, K\. Zhu, A\. Kapishnikov, R\. Galt, M\. Denton, B\. Murdoch, C\. Sikora, K\. Mohamed, W\. Wei, U\. First, T\. McConnell, L\. C\. Cobo, J\. Qin, T\. Avrahami, D\. Balle, Y\. Watanabe, A\. Louis, A\. Kraft, S\. Ariafar, Y\. Gu, E\. Rives, C\. Yoon, A\. Rusu, J\. Cobon\-Kerr, C\. Hahn, J\. Luo, Yuvein, Zhu, N\. Ahuja, R\. Benenson, R\. L\. Kaufman, H\. Yu, L\. Hightower, J\. Zhang, D\. Ni, L\. A\. Hendricks, G\. Wang, G\. Yona, L\. Jain, P\. Barrio, S\. Bhupatiraju, S\. Velusamy, A\. Dafoe, S\. Riedel, T\. Thomas, Z\. Yuan, M\. Bellaiche, S\. Panthaplackel, K\. Kloboves, S\. Jauhari, C\. Akbulut, T\. Davchev, E\. Gladchenko, D\. Madras, A\. Chuklin, T\. Hill, Q\. Yuan, M\. Madhavan, L\. Leonhard, D\. Scandinaro, Q\. Chen, N\. Niu, A\. Douillard, B\. Damoc, Y\. Onoe, F\. Pedregosa, F\. Bertsch, C\. Leichner, J\. Pagadora, J\. Malmaud, S\. Ponda, A\. Twigg, O\. Duzhyi, J\. Shen, M\. Wang, R\. Garg, J\. Chen, U\. Evci, J\. Lee, L\. Liu, K\. Kojima, M\. Yamaguchi, A\. Rajendran, A\. Piergiovanni, V\. K\. Rajendran, M\. Fornoni, G\. Ibagon, H\. Ragan, S\. M\. Khan, J\. Blitzer, A\. Bunner, G\. Sun, T\. Kosakai, S\. Lundberg, N\. Elue, K\. Guu, S\. Park, J\. Park, A\. Narayanaswamy, C\. Wu, J\. Mudigonda, T\. Cohn, H\. Mu, R\. Kumar, L\. Graesser, Y\. Zhang, R\. Killam, V\. Zhuang, M\. Giménez, W\. A\. Jishi, R\. Ley\-Wild, A\. Zhai, K\. Osawa, D\. Cedillo, J\. Liu, M\. Upadhyay, M\. Sieniek, R\. Sharma, T\. Paine, A\. Angelova, S\. Addepalli, C\. Parada, K\. Majumder, A\. Lamp, S\. Kumar, X\. Deng, A\. Myaskovsky, T\. Sabolić, J\. Dudek, S\. York, F\. de Chaumont Quitry, J\. Nie, D\. Cattle, A\. Gunjan, B\. Piot, W\. Khawaja, S\. Bang, S\. Wang, S\. Khodadadeh, R\. R, P\. Rawlani, R\. Powell, K\. Lee, J\. Griesser, G\. Oh, C\. Magalhaes, Y\. Li, S\. Tokumine, H\. N\. Vogel, D\. Hsu, A\. BC, D\. Jindal, M\. Cohen, Z\. Yang, J\. Yuan, D\. de Cesare, T\. Bruguier, J\. Xu, M\. Roy, A\. Jacovi, D\. Belov, R\. Arya, P\. Meadowlark, S\. Cohen\-Ganor, W\. Ye, P\. Morris\-Suzuki, P\. Banzal, G\. Song, P\. Ponnuramu, F\. Zhang, G\. Scrivener, S\. Zaiem, A\. R\. Rochman, K\. Han, B\. Ghazi, K\. Lee, S\. Drath, D\. Suo, A\. Girgis, P\. Shenoy, D\. Nguyen, D\. Eck, S\. Gupta, L\. Yan, J\. Carreira, A\. Gulati, R\. Sang, D\. Mirylenka, E\. Cooney, E\. Chou, M\. Ling, C\. Fan, B\. Coleman, G\. Tubone, R\. Kumar, J\. Baldridge, F\. Hernandez\-Campos, A\. Lazaridou, J\. Besley, I\. Yona, N\. Bulut, Q\. Wellens, A\. Pierigiovanni, J\. George, R\. Green, P\. Han, C\. Tao, G\. Clark, C\. You, A\. Abdolmaleki, J\. Fu, T\. Chen, A\. Chaugule, A\. Chandorkar, A\. Rahman, W\. Thompson, P\. Koanantakool, M\. Bernico, J\. Ren, A\. Vlasov, S\. Vassilvitskii, M\. Kula, Y\. Liang, D\. Kim, Y\. Huang, C\. Ye, D\. Lepikhin, and W\. HelmholzGemini 2\.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities\.External Links:2507\.06261,[Link](https://arxiv.org/abs/2507.06261)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.2.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Croxfordet al\.\(2025\)E\. Croxford, Y\. Gao, N\. Pellegrino, K\. Wong, G\. Wills, E\. First, M\. Schnier, K\. Burton, C\. Ebby, J\. Gorski, M\. Kalscheur, S\. Khalil, M\. Pisani, T\. Rubeor, P\. Stetson, F\. Liao, C\. Goswami, B\. Patterson, and M\. AfsharDevelopment and validation of the provider documentation summarization quality instrument for large language models\.Journal of the American Medical Informatics Association32\(6\),pp\. 1050–1060\(en\)\.External Links:ISSN 1067\-5027, 1527\-974X,[Link](https://academic.oup.com/jamia/article/32/6/1050/8125016),[Document](https://dx.doi.org/10.1093/jamia/ocaf068)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p4.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p2.1)\.
- Cuiet al\.\(2025\)H\. Cui, A\. Unell, B\. Chen, J\. A\. Fries, E\. Alsentzer, S\. Koyejo, and N\. H\. ShahTIMER: temporal instruction modeling and evaluation for longitudinal clinical records\.npj Digital Medicine8\(1\),pp\. 577\(en\)\.External Links:ISSN 2398\-6352,[Link](https://www.nature.com/articles/s41746-025-01965-9),[Document](https://dx.doi.org/10.1038/s41746-025-01965-9)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p4.1)\.
- de Hondet al\.\(2024\)A\. de Hond, T\. Leeuwenberg, R\. Bartels, M\. van Buchem, I\. Kant, K\. G\. Moons, and M\. van SmedenFrom text to treatment: the crucial role of validation for generative large language models in health care\.The Lancet Digital Health6\(7\),pp\. e441–e443\.Note:doi: 10\.1016/S2589\-7500\(24\)00111\-0External Links:ISSN 2589\-7500,[Link](https://doi.org/10.1016/S2589-7500(24)00111-0),[Document](https://dx.doi.org/10.1016/S2589-7500%2824%2900111-0)Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p2.1)\.
- Diekmannet al\.\(2025\)Y\. Diekmann, C\. Fensore, R\. Carrillo\-Larco, E\. Castejon Rosales, S\. Shiromani, R\. Pai, M\. Shah, and J\. HoLLMs as Medical Safety Judges: Evaluating Alignment with Human Annotation in Patient\-Facing QA\.InProceedings of the 24th Workshop on Biomedical Language Processing,Viena, Austria,pp\. 217–224\(en\)\.External Links:[Link](https://aclanthology.org/2025.bionlp-1.19),[Document](https://dx.doi.org/10.18653/v1/2025.bionlp-1.19)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- Droret al\.\(2018\)R\. Dror, G\. Baumer, S\. Shlomov, and R\. ReichartThe Hitchhiker’s Guide to Testing Statistical Significance in Natural Language Processing\.InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),Melbourne, Australia,pp\. 1383–1392\(en\)\.External Links:[Link](http://aclweb.org/anthology/P18-1128),[Document](https://dx.doi.org/10.18653/v1/P18-1128)Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p1.1)\.
- Engel \(1977\)G\. L\. EngelThe Need for a New Medical Model: A Challenge for Biomedicine\.Science196\(4286\),pp\. 129–136\(en\)\.External Links:ISSN 0036\-8075, 1095\-9203,[Link](https://www.science.org/doi/10.1126/science.847460),[Document](https://dx.doi.org/10.1126/science.847460)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p6.1)\.
- Frattallone\-Lladoet al\.\(2024\)G\. Frattallone\-Llado, J\. Kim, C\. Cheng, D\. Salazar, S\. Edakalavan, and J\. C\. WeissUsing Multimodal Data to Improve Precision of Inpatient Event Timelines\.InAdvances in Knowledge Discovery and Data Mining,D\. Yang, X\. Xie, V\. S\. Tseng, J\. Pei, J\. Huang, and J\. C\. Lin \(Eds\.\),Vol\.14648,pp\. 322–334\(en\)\.Note:Series Title: Lecture Notes in Computer ScienceExternal Links:ISBN 978\-981\-97\-2240\-2 978\-981\-97\-2238\-9,[Link](https://link.springer.com/10.1007/978-981-97-2238-9_25),[Document](https://dx.doi.org/10.1007/978-981-97-2238-9%5F25)Cited by:[§7\.1](https://arxiv.org/html/2609.19585#S7.SS1.p1.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p1.1)\.
- Gaoet al\.\(2022\)Y\. Gao, D\. Dligach, T\. Miller, D\. Xu, M\. M\. M\. Churpek, and M\. AfsharSummarizing patients’ problems from hospital progress notes using pre\-trained sequence\-to\-sequence models\.InProceedings of the 29th International Conference on Computational Linguistics,N\. Calzolari, C\. Huang, H\. Kim, J\. Pustejovsky, L\. Wanner, K\. Choi, P\. Ryu, H\. Chen, L\. Donatelli, H\. Ji, S\. Kurohashi, P\. Paggio, N\. Xue, S\. Kim, Y\. Hahm, Z\. He, T\. K\. Lee, E\. Santus, F\. Bond, and S\. Na \(Eds\.\),Gyeongju, Republic of Korea,pp\. 2979–2991\.External Links:[Link](https://aclanthology.org/2022.coling-1.264/)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- Gemma Team \(2026\)Gemma TeamGemma 4 Technical Report\.External Links:2607\.02770,[Link](https://arxiv.org/abs/2607.02770)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.8.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px1.p1.1)\.
- Grotheyet al\.\(2025\)B\. Grothey, J\. Odenkirchen, A\. Brkic, B\. Schömig\-Markiefka, A\. Quaas, R\. Büttner, and Y\. TolkachComprehensive testing of large language models for extraction of structured data in pathology\.Communications Medicine5\(1\),pp\. 96\(en\)\.External Links:ISSN 2730\-664X,[Link](https://www.nature.com/articles/s43856-025-00808-8),[Document](https://dx.doi.org/10.1038/s43856-025-00808-8)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px1.p1.1)\.
- Gumielet al\.\(2022\)Y\. B\. Gumiel, L\. E\. Silva E Oliveira, V\. Claveau, N\. Grabar, E\. C\. Paraiso, C\. Moro, and D\. R\. CarvalhoTemporal Relation Extraction in Clinical Texts: A Systematic Review\.ACM Computing Surveys54\(7\),pp\. 1–36\(en\)\.External Links:ISSN 0360\-0300, 1557\-7341,[Link](https://dl.acm.org/doi/10.1145/3462475),[Document](https://dx.doi.org/10.1145/3462475)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p4.1),[§5\.1](https://arxiv.org/html/2609.19585#S5.SS1.p1.1)\.
- Gwet \(2008\)K\. L\. GwetComputing inter\-rater reliability and its variance in the presence of high agreement\.British Journal of Mathematical and Statistical Psychology61\(1\),pp\. 29–48\.External Links:[Document](https://dx.doi.org/10.1348/000711006X126600)Cited by:[§5\.2](https://arxiv.org/html/2609.19585#S5.SS2.p2.1)\.
- Hageret al\.\(2024\)P\. Hager, F\. Jungmann, R\. Holland, K\. Bhagat, I\. Hubrecht, M\. Knauer, J\. Vielhauer, M\. Makowski, R\. Braren, G\. Kaissis, and D\. RueckertEvaluation and mitigation of the limitations of large language models in clinical decision\-making\.Nature Medicine30\(9\),pp\. 2613–2622\(en\)\.External Links:ISSN 1078\-8956, 1546\-170X,[Link](https://www.nature.com/articles/s41591-024-03097-1),[Document](https://dx.doi.org/10.1038/s41591-024-03097-1)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- Jianget al\.\(2026\)Z\. Jiang, H\. Chen, Y\. Wu, Y\. Qin, C\. Pei, D\. Zeng, B\. Sheng, and T\. Y\. WongBeyond multiple\-choice questions: Rethinking evaluation frameworks for large language models for clinical medicine\.Intelligent Medicine6\(2\),pp\. 109–115\(en\)\.External Links:ISSN 26671026,[Link](https://linkinghub.elsevier.com/retrieve/pii/S266710262600001X),[Document](https://dx.doi.org/10.1016/j.imed.2026.01.001)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- Jinet al\.\(2024\)Y\. Jin, M\. Chandra, G\. Verma, Y\. Hu, M\. De Choudhury, and S\. KumarBetter to ask in english: cross\-lingual evaluation of large language models for healthcare queries\.InProceedings of the ACM Web Conference 2024,WWW ’24,New York, NY, USA,pp\. 2627–2638\.External Links:ISBN 9798400701719,[Link](https://doi.org/10.1145/3589334.3645643),[Document](https://dx.doi.org/10.1145/3589334.3645643)Cited by:[Appendix B](https://arxiv.org/html/2609.19585#A2.p2.1)\.
- Johnsonet al\.\(2016\)A\. E\.W\. Johnson, T\. J\. Pollard, L\. Shen, L\. H\. Lehman, M\. Feng, M\. Ghassemi, B\. Moody, P\. Szolovits, L\. Anthony Celi, and R\. G\. MarkMIMIC\-III, a freely accessible critical care database\.Scientific Data3\(1\),pp\. 160035\(en\)\.External Links:ISSN 2052\-4463,[Link](https://www.nature.com/articles/sdata201635),[Document](https://dx.doi.org/10.1038/sdata.2016.35)Cited by:[§2](https://arxiv.org/html/2609.19585#S2.p1.1)\.
- Johnsonet al\.\(2015\)A\. Johnson, T\. Pollard, and R\. MarkMIMIC\-III Clinical Database\.PhysioNet\.External Links:[Link](https://physionet.org/content/mimiciii/1.4/),[Document](https://dx.doi.org/10.13026/C2XW26)Cited by:[§2](https://arxiv.org/html/2609.19585#S2.p1.1)\.
- Ket al\.\(2025\)K\. K, R\. Thirukovalluru, and D\. CarlsonClinStructor: AI\-Powered Structuring of Unstructured Clinical Texts\.InProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia\-Pacific Chapter of the Association for Computational Linguistics,Mumbai, India,pp\. 2822–2836\(en\)\.External Links:[Link](https://aclanthology.org/2025.ijcnlp-long.151),[Document](https://dx.doi.org/10.18653/v1/2025.ijcnlp-long.151)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1)\.
- Kiet al\.\(2025\)D\. Ki, M\. Carpuat, P\. McNamee, D\. Khashabi, E\. Yang, D\. Lawrie, and K\. DuhLinguistic nepotism: trading\-off quality for language preference in multilingual rag\.arXiv preprint arXiv:2509\.13930\.Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p1.1)\.
- Kimet al\.\(2024\)M\. K\. Kim, C\. Rouphael, J\. McMichael, N\. Welch, and S\. DasarathyChallenges in and Opportunities for Electronic Health Record\-Based Data Analysis and Interpretation\.Gut and Liver18\(2\),pp\. 201–208\(en\)\.External Links:ISSN 1976\-2283, 2005\-1212,[Link](http://gutnliver.org/journal/view.html?doi=10.5009/gnl230272),[Document](https://dx.doi.org/10.5009/gnl230272)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1)\.
- Kraljevicet al\.\(2024\)Z\. Kraljevic, D\. Bean, A\. Shek, R\. Bendayan, H\. Hemingway, J\. A\. Yeung, A\. Deng, A\. Balston, J\. Ross, E\. Idowu, J\. T\. Teo, and R\. J\. B\. DobsonForesight—a generative pretrained transformer for modelling of patient timelines using electronic health records: a retrospective modelling study\.The Lancet Digital Health6\(4\),pp\. e281–e290\(en\)\.External Links:ISSN 25897500,[Link](https://linkinghub.elsevier.com/retrieve/pii/S2589750024000256),[Document](https://dx.doi.org/10.1016/S2589-7500%2824%2900025-6)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p1.1)\.
- Kruseet al\.\(2025\)M\. Kruse, S\. Hu, N\. Derby, Y\. Wu, S\. Stonbraker, B\. Yao, D\. Wang, E\. M\. Goldberg, and Y\. GaoLarge language models with temporal reasoning for longitudinal clinical summarization and prediction\.Findings of ACL\. EMNLP\. Conference on Empirical Methods in Natural Language Processing2025,pp\. 20715 – 20735\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1128)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p4.1)\.
- Kumaret al\.\(2026\)S\. Kumar, S\. Noroozizadeh, J\. Kim, and J\. C\. WeissText Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval\-Augmented Multimodal Alignment\.arXiv\.Note:Version Number: 1Other Sayantan Kumar, Shahriar Noroozizadeh, Juyong Kim \(authors contributed equally\)External Links:[Link](https://arxiv.org/abs/2605.15168),[Document](https://dx.doi.org/10.48550/ARXIV.2605.15168)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p5.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p1.1)\.
- Leeet al\.\(2024\)C\. Lee, K\. A\. Vogt, and S\. KumarProspects for AI clinical summarization to reduce the burden of patient chart review\.Frontiers in Digital Health6,pp\. 1475092\.External Links:ISSN 2673\-253X,[Link](https://www.frontiersin.org/articles/10.3389/fdgth.2024.1475092/full),[Document](https://dx.doi.org/10.3389/fdgth.2024.1475092)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p2.1)\.
- Leeuwenberg and Moens \(2020\)A\. Leeuwenberg and M\. MoensTowards Extracting Absolute Event Timelines From English Clinical Reports\.IEEE/ACM Transactions on Audio, Speech, and Language Processing28,pp\. 2710–2719\.External Links:ISSN 2329\-9290, 2329\-9304,[Link](https://ieeexplore.ieee.org/document/9207839/),[Document](https://dx.doi.org/10.1109/TASLP.2020.3027201)Cited by:[§7\.1](https://arxiv.org/html/2609.19585#S7.SS1.p1.1)\.
- Liet al\.\(2023\)J\. Li, X\. Cheng, X\. Zhao, J\. Nie, and J\. WenHaluEval: a large\-scale hallucination evaluation benchmark for large language models\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6449–6464\.External Links:[Link](https://aclanthology.org/2023.emnlp-main.397/),[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.397)Cited by:[Appendix E](https://arxiv.org/html/2609.19585#A5.p1.1)\.
- Linhareset al\.\(2023\)C\. D\. G\. Linhares, D\. M\. Lima, J\. R\. Ponciano, M\. M\. Olivatto, M\. A\. Gutierrez, J\. Poco, C\. Traina, and A\. J\. M\. TrainaClinicalPath: A Visualization Tool to Improve the Evaluation of Electronic Health Records in Clinical Decision\-Making\.IEEE Transactions on Visualization and Computer Graphics29\(10\),pp\. 4031–4046\.External Links:ISSN 1077\-2626, 1941\-0506, 2160\-9306,[Link](https://ieeexplore.ieee.org/document/9779066/),[Document](https://dx.doi.org/10.1109/TVCG.2022.3175626)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1),[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- Loftuset al\.\(2024\)T\. J\. Loftus, J\. A\. Balch, J\. L\. Marquard, J\. M\. Ray, B\. S\. Alper, N\. Ojha, A\. Bihorac, G\. Melton\-Meaux, G\. Khanna, and C\. J\. TignanelliLongitudinal clinical decision support for assessing decisions over time: State\-of\-the\-art and future directions\.DIGITAL HEALTH10,pp\. 20552076241249925\(en\)\.External Links:ISSN 2055\-2076, 2055\-2076,[Link](https://journals.sagepub.com/doi/10.1177/20552076241249925),[Document](https://dx.doi.org/10.1177/20552076241249925)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1)\.
- Luoet al\.\(2023\)Z\. Luo, Y\. Ji, A\. Gupta, Z\. Li, A\. Frisch, and D\. HeTowards Accurate and Clinically Meaningful Summarization of Electronic Health Record Notes: A Guided Approach\.In2023 IEEE EMBS International Conference on Biomedical and Health Informatics \(BHI\),Pittsburgh, PA, USA,pp\. 1–5\.External Links:ISBN 979\-8\-3503\-1050\-4,[Link](https://ieeexplore.ieee.org/document/10313411/),[Document](https://dx.doi.org/10.1109/BHI58575.2023.10313411)Cited by:[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p2.1)\.
- Makarovet al\.\(2025\)N\. Makarov, M\. Bordukova, P\. Quengdaeng, D\. Garger, R\. Rodriguez\-Esteban, F\. Schmich, and M\. P\. MendenLarge language models forecast patient health trajectories enabling digital twins\.npj Digital Medicine8\(1\),pp\. 588\(en\)\.External Links:ISSN 2398\-6352,[Link](https://www.nature.com/articles/s41746-025-02004-3),[Document](https://dx.doi.org/10.1038/s41746-025-02004-3)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- Meskó and Topol \(2023\)B\. Meskó and E\. J\. TopolThe imperative for regulatory oversight of large language models \(or generative AI\) in healthcare\.npj Digital Medicine6\(1\),pp\. 120\(en\)\.External Links:ISSN 2398\-6352,[Link](https://www.nature.com/articles/s41746-023-00873-0),[Document](https://dx.doi.org/10.1038/s41746-023-00873-0)Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p2.1)\.
- Meta AI \(2024\)Meta AILlama 3\.3 70B Instruct model card\.Note:Hugging Face model cardAccessed July 27, 2026External Links:[Link](https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.6.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Mistral AI \(2025\)Mistral AIMistral\-Small\-3\.2\-24B\-Instruct\-2506 model card\.Note:Hugging Face model cardAccessed July 27, 2026External Links:[Link](https://huggingface.co/mistralai/Mistral-Small-3.2-24B-Instruct-2506)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.10.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px1.p1.1)\.
- Mittalet al\.\(2026\)S\. Mittal, E\. Kasson, L\. Paraboschi, E\. Laufenberg, J\. Zhou, P\. A\. Cavazos\-Rehg, T\. Mitra, and M\. De ChoudhuryFollow the rules \(or not\): community norms and AI\-generated support in online health communities\.InProceedings of the International AAAI Conference on Web and Social Media,Note:To appearExternal Links:[Link](https://arxiv.org/abs/2603.19093)Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p2.1)\.
- Moharasan and Ho \(2019\)G\. Moharasan and T\. HoExtraction of Temporal Information from Clinical Narratives\.Journal of Healthcare Informatics Research3\(2\),pp\. 220–244\(en\)\.External Links:ISSN 2509\-4971, 2509\-498X,[Link](http://link.springer.com/10.1007/s41666-019-00049-0),[Document](https://dx.doi.org/10.1007/s41666-019-00049-0)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- Noroozizadehet al\.\(2025\)S\. Noroozizadeh, S\. Kumar, G\. H\. Chen, and J\. C\. WeissTemporally annotated textual time series from PubMed Open Access clinical case reports\.Health Informatics\(en\)\.External Links:[Link](http://medrxiv.org/lookup/doi/10.1101/2025.11.01.25339297),[Document](https://dx.doi.org/10.1101/2025.11.01.25339297)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p5.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p1.1)\.
- Ntinopouloset al\.\(2025\)V\. Ntinopoulos, H\. Rodriguez Cetina Biefer, I\. Tudorache, N\. Papadopoulos, D\. Odavic, P\. Risteski, A\. Haeussler, and O\. DzemaliLarge language models for data extraction from unstructured and semi\-structured electronic health records: a multiple model performance evaluation\.BMJ Health & Care Informatics32\(1\),pp\. e101139\(en\)\.External Links:ISSN 2632\-1009,[Link](https://informatics.bmj.com/lookup/doi/10.1136/bmjhci-2024-101139),[Document](https://dx.doi.org/10.1136/bmjhci-2024-101139)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px1.p1.1)\.
- Olex and McInnes \(2021\)A\. L\. Olex and B\. T\. McInnesReview of Temporal Reasoning in the Clinical Domain for Timeline Extraction: Where we are and where we need to be\.Journal of Biomedical Informatics118,pp\. 103784\(en\)\.External Links:ISSN 15320464,[Link](https://linkinghub.elsevier.com/retrieve/pii/S1532046421001131),[Document](https://dx.doi.org/10.1016/j.jbi.2021.103784)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1),[§1](https://arxiv.org/html/2609.19585#S1.p3.1)\.
- OpenMeditron \(2025\)OpenMeditronMeditron3\-Qwen2\.5\-7B model card\.Note:Hugging Face model cardAccessed July 27, 2026External Links:[Link](https://huggingface.co/EPFLiGHT/Meditron3-Qwen2.5-7B)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.11.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px1.p1.1)\.
- Qwen Team \(2025\)Qwen TeamQwen3 technical report\.External Links:2505\.09388,[Link](https://arxiv.org/abs/2505.09388)Cited by:[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px1.p1.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.6\-35B\-A3B: agentic coding power, now open to all\.External Links:[Link](https://qwen.ai/blog?id=qwen3.6-35b-a3b)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.7.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1),[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px1.p1.1)\.
- Ruanet al\.\(2019\)T\. Ruan, L\. Lei, Y\. Zhou, J\. Zhai, L\. Zhang, P\. He, and J\. GaoRepresentation learning for clinical time series prediction tasks in electronic health records\.BMC Medical Informatics and Decision Making19\(S8\),pp\. 259\(en\)\.External Links:ISSN 1472\-6947,[Link](https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-019-0985-7),[Document](https://dx.doi.org/10.1186/s12911-019-0985-7)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p1.1)\.
- Sahaet al\.\(2026\)K\. Saha, D\. W\. Yoo, V\. Das Swain, and M\. De ChoudhuryLife events as predictors of wellbeing outcomes\.npj Digital Public Health1\(1\),pp\. 5\(en\)\.External Links:ISSN 3091\-2784,[Link](https://www.nature.com/articles/s44482-025-00005-3),[Document](https://dx.doi.org/10.1038/s44482-025-00005-3)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p6.1)\.
- Sarvari and Al\-fagih \(2025\)P\. Sarvari and Z\. Al\-fagihRapidly Benchmarking Large Language Models for Diagnosing Comorbid Patients: Comparative Study Leveraging the LLM\-as\-a\-Judge Method\.JMIRx Med6,pp\. e67661–e67661\(en\)\.External Links:ISSN 2563\-6316,[Link](https://xmed.jmir.org/2025/1/e67661),[Document](https://dx.doi.org/10.2196/67661)Cited by:[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Seinenet al\.\(2025\)T\. M\. Seinen, J\. A\. Kors, E\. M\. Van Mulligen, and P\. R\. RijnbeekUsing Structured Codes and Free\-Text Notes to Measure Information Complementarity in Electronic Health Records: Feasibility and Validation Study\.Journal of Medical Internet Research27,pp\. e66910\(en\)\.External Links:ISSN 1438\-8871,[Link](https://www.jmir.org/2025/1/e66910),[Document](https://dx.doi.org/10.2196/66910)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p2.1)\.
- Sellergrenet al\.\(2026\)A\. Sellergren, S\. Kazemzadeh, T\. Jaroensri, A\. Kiraly, M\. Traverse, T\. Kohlberger, S\. Xu, F\. Jamil, C\. Hughes, C\. Lau, J\. Chen, F\. Mahvar, L\. Yatziv, T\. Chen, B\. Sterling, S\. A\. Baby, S\. M\. Baby, J\. Lai, S\. Schmidgall, L\. Yang, K\. Chen, P\. Bjornsson, S\. Reddy, R\. Brush, K\. Philbrick, M\. Asiedu, I\. Mezerreg, H\. Hu, H\. Yang, R\. Tiwari, S\. Jansen, P\. Singh, Y\. Liu, S\. Azizi, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Riviere, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Buchatskaya, J\. Alayrac, D\. Lepikhin, V\. Feinberg, S\. Borgeaud, A\. Andreev, C\. Hardin, R\. Dadashi, L\. Hussenot, A\. Joulin, O\. Bachem, Y\. Matias, K\. Chou, A\. Hassidim, K\. Goel, C\. Farabet, J\. Barral, T\. Warkentin, J\. Shlens, D\. Fleet, V\. Cotruta, O\. Sanseviero, G\. Martins, P\. Kirk, A\. Rao, S\. Shetty, D\. F\. Steiner, C\. Kirmizibayrak, R\. Pilgrim, D\. Golden, and L\. YangMedGemma technical report\.External Links:2507\.05201,[Link](https://arxiv.org/abs/2507.05201)Cited by:[Table A4](https://arxiv.org/html/2609.19585#A2.T4.2.9.1.1.1),[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Sharifet al\.\(2025\)O\. Sharif, J\. Gatto, M\. Basak, and S\. M\. PreumREGen: a reliable evaluation framework for generative event argument extraction\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 12146–12168\.Cited by:[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px3.p1.1)\.
- Shiet al\.\(2024\)W\. Shi, R\. Xu, Y\. Zhuang, Y\. Yu, J\. Zhang, H\. Wu, Y\. Zhu, J\. C\. Ho, C\. Yang, and M\. D\. WangEHRAgent: Code Empowers Large Language Models for Few\-shot Complex Tabular Reasoning on Electronic Health Records\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Miami, Florida, USA,pp\. 22315–22339\(en\)\.External Links:[Link](https://aclanthology.org/2024.emnlp-main.1245),[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.1245)Cited by:[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Styleret al\.\(2014\)W\. F\. Styler, S\. Bethard, S\. Finan, M\. Palmer, S\. Pradhan, P\. C\. De Groen, B\. Erickson, T\. Miller, C\. Lin, G\. Savova, and J\. PustejovskyTemporal Annotation in the Clinical Domain\.Transactions of the Association for Computational Linguistics2,pp\. 143–154\(en\)\.External Links:ISSN 2307\-387X,[Link](https://direct.mit.edu/tacl/article/43304),[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00172)Cited by:[§7\.1](https://arxiv.org/html/2609.19585#S7.SS1.p1.1)\.
- Sunet al\.\(2013\)W\. Sun, A\. Rumshisky, and O\. UzunerEvaluating temporal relations in clinical text: 2012 i2b2 Challenge\.Journal of the American Medical Informatics Association20\(5\),pp\. 806–813\(en\)\.External Links:ISSN 1067\-5027, 1527\-974X,[Link](https://academic.oup.com/jamia/article-lookup/doi/10.1136/amiajnl-2013-001628),[Document](https://dx.doi.org/10.1136/amiajnl-2013-001628)Cited by:[§7\.1](https://arxiv.org/html/2609.19585#S7.SS1.p1.1)\.
- Umetonet al\.\(2024\)R\. Umeton, A\. Kwok, R\. Maurya, D\. Leco, N\. Lenane, J\. Willcox, G\. A\. Abel, M\. Tolikas, and J\. M\. JohnsonGPT\-4 in a Cancer Center — Institute\-Wide Deployment Challenges and Lessons Learned\.NEJM AI1\(4\) \(en\)\.External Links:ISSN 2836\-9386,[Link](https://ai.nejm.org/doi/10.1056/AIcs2300191),[Document](https://dx.doi.org/10.1056/AIcs2300191)Cited by:[§4\.2](https://arxiv.org/html/2609.19585#S4.SS2.p2.1)\.
- Uzuneret al\.\(2011\)Ö\. Uzuner, B\. R\. South, S\. Shen, and S\. L\. DuVall2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text\.Journal of the American Medical Informatics Association18\(5\),pp\. 552–556\(en\)\.External Links:ISSN 1527\-974X, 1067\-5027,[Link](https://academic.oup.com/jamia/article/18/5/552/830538),[Document](https://dx.doi.org/10.1136/amiajnl-2011-000203)Cited by:[§7\.1](https://arxiv.org/html/2609.19585#S7.SS1.p1.1)\.
- Van Veenet al\.\(2024\)D\. Van Veen, C\. Van Uden, L\. Blankemeier, J\. Delbrouck, A\. Aali, C\. Bluethgen, A\. Pareek, M\. Polacin, E\. P\. Reis, A\. Seehofnerová, N\. Rohatgi, P\. Hosamani, W\. Collins, N\. Ahuja, C\. P\. Langlotz, J\. Hom, S\. Gatidis, J\. Pauly, and A\. S\. ChaudhariAdapted large language models can outperform medical experts in clinical text summarization\.Nature Medicine30\(4\),pp\. 1134–1142\(en\)\.External Links:ISSN 1078\-8956, 1546\-170X,[Link](https://www.nature.com/articles/s41591-024-02855-5),[Document](https://dx.doi.org/10.1038/s41591-024-02855-5)Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px3.p1.1)\.
- Virtanenet al\.\(2020\)P\. Virtanen, R\. Gommers, T\. E\. Oliphant, M\. Haberland, T\. Reddy, D\. Cournapeau, E\. Burovski, P\. Peterson, W\. Weckesser, J\. Bright, S\. J\. van der Walt, M\. Brett, J\. Wilson, K\. J\. Millman, N\. Mayorov, A\. R\. J\. Nelson, E\. Jones, R\. Kern, E\. Larson, C\. J\. Carey, İ\. Polat, Y\. Feng, E\. W\. Moore, J\. VanderPlas, D\. Laxalde, J\. Perktold, R\. Cimrman, I\. Henriksen, E\. A\. Quintero, C\. R\. Harris, A\. M\. Archibald, A\. H\. Ribeiro, F\. Pedregosa, P\. van Mulbregt, and SciPy 1\.0 ContributorsSciPy 1\.0: Fundamental Algorithms for Scientific Computing in Python\.Nature Methods17,pp\. 261–272\.External Links:[Document](https://dx.doi.org/10.1038/s41592-019-0686-2)Cited by:[Appendix C](https://arxiv.org/html/2609.19585#A3.p9.1)\.
- Wanget al\.\(2025a\)J\. Wang, X\. Niu, J\. Kim, J\. Shen, T\. Zhang, and J\. WeissMIMIC\-IV\-Ext\-22MCTS: A 22 Millions\-Event Temporal Clinical Time\-Series Dataset with Relative Timestamp for Risk Prediction\.In Review\.External Links:[Link](https://www.researchsquare.com/article/rs-6347897/v1),[Document](https://dx.doi.org/10.21203/rs.3.rs-6347897/v1)Cited by:[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p1.1)\.
- Wang and Weiss \(2025\)J\. Wang and J\. C\. WeissA Large\-Language Model Framework for Relative Timeline Extraction from PubMed Case Reports\.AMIA Joint Summits on Translational Science proceedings\. AMIA Joint Summits on Translational Science2025,pp\. 598–606\(eng\)\.External Links:ISSN 2153\-4063Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p5.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p1.1)\.
- Wanget al\.\(2024\)L\. Wang, Q\. Lu, R\. Li, S\. Fu, and H\. LiuWonder at Chemotimelines 2024: MedTimeline: An End\-to\-End NLP System for Timeline Extraction from Clinical Narratives\.InProceedings of the 6th Clinical Natural Language Processing Workshop,Mexico City, Mexico,pp\. 483–487\(en\)\.External Links:[Link](https://aclanthology.org/2024.clinicalnlp-1.48),[Document](https://dx.doi.org/10.18653/v1/2024.clinicalnlp-1.48)Cited by:[§1](https://arxiv.org/html/2609.19585#S1.p5.1),[§7\.1](https://arxiv.org/html/2609.19585#S7.SS1.p1.1),[§7\.2](https://arxiv.org/html/2609.19585#S7.SS2.p1.1)\.
- Wanget al\.\(2025b\)X\. Wang, J\. Zhang, G\. Zhang, and H\. GuoFeel the difference? a comparative analysis of emotional arcs in real and LLM\-generated CBT sessions\.InFindings of the Association for Computational Linguistics: EMNLP 2025,C\. Christodoulopoulos, T\. Chakraborty, C\. Rose, and V\. Peng \(Eds\.\),Suzhou, China,pp\. 19999–20017\.External Links:[Link](https://aclanthology.org/2025.findings-emnlp.1089/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.1089),ISBN 979\-8\-89176\-335\-7Cited by:[§8](https://arxiv.org/html/2609.19585#S8.SS0.SSS0.Px2.p1.1)\.
- World Health Organization \(2025a\)World Health OrganizationMental disorders\.Note:[https://www\.who\.int/news\-room/fact\-sheets/detail/mental\-disorders](https://www.who.int/news-room/fact-sheets/detail/mental-disorders)Accessed: 2026\-07\-13Cited by:[§5\.5](https://arxiv.org/html/2609.19585#S5.SS5.p1.1)\.
- World Health Organization \(2025b\)World Health OrganizationMental health\.Note:[https://www\.who\.int/news\-room/fact\-sheets/detail/mental\-health\-strengthening\-our\-response](https://www.who.int/news-room/fact-sheets/detail/mental-health-strengthening-our-response)Accessed: 2026\-07\-14Cited by:[§5\.5](https://arxiv.org/html/2609.19585#S5.SS5.p1.1)\.
- Yimet al\.\(2023\)W\. Yim, Y\. Fu, A\. Ben Abacha, N\. Snider, T\. Lin, and M\. YetisgenACI\-bench: a novel ambient clinical intelligence dataset for benchmarking automatic visit note generation\.Nature Scientific Data\.External Links:[Link](https://www.nature.com/articles/s41597-023-02487-3)Cited by:[§5\.1](https://arxiv.org/html/2609.19585#S5.SS1.SSS0.Px2.p2.1)\.
- Yuet al\.\(2025\)H\. Yu, J\. Zhou, L\. Li, S\. Chen, J\. Gallifant, A\. Shi, J\. Sun, X\. Li, J\. He, W\. Hua, M\. Jin, G\. Chen, Y\. Zhou, Z\. Li, T\. Gupte, M\. Chen, Z\. Azizi, Q\. Dou, B\. P\. Yan, Y\. Xing, Y\. Zhang, T\. L\. Assimes, D\. S\. Bitterman, X\. Ma, L\. Lu, and L\. FanSimulated patient systems powered by large language model\-based AI agents offer potential for transforming medical education\.Communications Medicine6\(1\),pp\. 27\(en\)\.External Links:ISSN 2730\-664X,[Link](https://www.nature.com/articles/s43856-025-01283-x),[Document](https://dx.doi.org/10.1038/s43856-025-01283-x)Cited by:[§4\.1](https://arxiv.org/html/2609.19585#S4.SS1.p1.1)\.
- Zhanget al\.\(2019\)T\. Zhang, V\. Kishore, F\. Wu, K\. Q\. Weinberger, and Y\. ArtziBertscore: evaluating text generation with bert\.arXiv preprint arXiv:1904\.09675\.Cited by:[§6\.1](https://arxiv.org/html/2609.19585#S6.SS1.SSS0.Px3.p1.1)\.
- Zhenget al\.\(2023\)L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. StoicaJudging llm\-as\-a\-judge with mt\-bench and chatbot arena\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.Cited by:[§5\.1](https://arxiv.org/html/2609.19585#S5.SS1.SSS0.Px2.p1.1)\.

## Appendix ADataset Extraction and Profile

### A\.1Mental Health Admissions Subset

We operationalize mental\-health\-coded admissions using the International Classification of Diseases, Ninth Revision, Clinical Modification \(ICD\-9\-CM\): an admission qualifies if it contains at least one diagnosis code from Chapter 5, Mental Disorders \(290–319\) \(Centers for Medicare & Medicaid Services and National Center for Health Statistics, 2011\)[Centers for Medicare & Medicaid Services and National Center for Health Statistics \(2011\)](https://arxiv.org/html/2609.19585#bib.bib16)\. Table[A1](https://arxiv.org/html/2609.19585#A1.T1)breaks down the key categories\.

Our framework uses only free\-text discharge summaries from the MIMIC\-IIINOTEEVENTStable, which describe each patient’s hospital stay\. We link each summary to structured MIMIC\-III fields only to define the study cohort and provide date anchors: diagnosis codes fromDIAGNOSES\_ICD, date of birth fromPATIENTS, and admission and discharge dates fromADMISSIONS\. These structured fields are not used as model inputs or training labels\. The discharge summaries are provided to the framework in their original form, without chunking, section detection, or text normalization\.

Table A1:Subcategories within ICD\-9\-CM Chapter 5, Mental Disorders \(codes 290–319\)[Centers for Medicare & Medicaid Services and National Center for Health Statistics \(2011\)](https://arxiv.org/html/2609.19585#bib.bib16)\. An admission is included in the mental\-health\-coded cohort if it contains at least one ICD\-9\-CM diagnosis code within the Mental Disorders chapter \(290–319\)
### A\.2Benchmark Input Text Profile

Because both model evaluation and human validation require comparison against the complete source discharge summary, input length determines the amount of narrative context that must be processed and reviewed\. Table[A2](https://arxiv.org/html/2609.19585#A1.T2)summarizes the raw\-text lengths of the 52\-document benchmark\. The summaries contain1,7741,774words on average \(median:1,6241,624; SD:821821\), ranging from437437to3,5833,583words and totaling92,22292,222words\. Character counts show similar variation, with a mean of11,87111,871characters and a range of2,9522,952to24,80824,808\. These statistics reflect the original discharge summaries supplied to the framework without chunking, section parsing, or text normalization, and illustrate the long\-context processing and manual\-review burden of the benchmark\.

Table A2:Raw\-text profile of the 52 discharge summaries in the benchmark\.
### A\.3Token Usage for Silver\-Standard Generation

To quantify the inference required to construct the silver\-standard supervision used in Section[6](https://arxiv.org/html/2609.19585#S6), we recorded the API token\-usage metadata returned by Gemini 2\.5 Pro for the 1,000 admissions processed through the frozen three\-stage framework\. Across Stages 1 to 3, generation processed45\.1345\.13M input tokens and produced16\.7016\.70M output tokens, for a combined total of61\.8461\.84M tokens, or61,83661,836tokens per admission on average\. Stage 1 accounted for25\.6625\.66M tokens \(41\.541\.5% of the total\), Stage 2 for19\.5219\.52M \(31\.631\.6%\), and Stage 3 for16\.6616\.66M \(26\.926\.9%\)\. Table[A3](https://arxiv.org/html/2609.19585#A1.T3)provides the complete stage\-level breakdown\.

These counts quantify the inference used to generate the1,0001,000\-example silver corpus in Section[6](https://arxiv.org/html/2609.19585#S6)and illustrate the computational demands of processing complete discharge summaries and their intermediate representations through a multi\-stage long\-context pipeline\. The resulting corpus, however, can be reused across three downstream tasks and multiple open\-weight student models, amortizing the initial teacher\-generation cost\. Of the reported input tokens,16,177,25416,177,254were served from cached context; because cached tokens are a subset of input tokens, they are not added again to the total\.

Table A3:Token usage for generating the 1,000\-example silver\-standard corpus with the frozen Gemini 2\.5 Pro\. Total tokens are the sum of input and output tokens\. The grand total corresponds to an average of 61,836 tokens per admission across all three stages\. Of the input tokens, 16,177,254 were served from cached context and are already included in the input counts\.

## Appendix BAgent Candidate Selection

Following the selection of agents in Section[4\.2](https://arxiv.org/html/2609.19585#S4.SS2), a candidate is eligible if it supports a native context window of at least3232K tokens which is sufficient to fit our P9999input without truncation\. All candidates are evaluated with identical prompts and greedy decoding \(t=0t=0\)\. An open\-weight candidate additionally has to be \(i\) downloadable under a license permitting research and on\-premise inference and \(ii\) loadable on our deployment stack across four L40S GPUs at bfloat16 precision\. A proprietary candidate additionally has to be \(iii\) served through a managed endpoint compatible with our institutional data\-use agreement \(e\.g\., Amazon Bedrock or Google Vertex AI\) and \(iv\) cost\-feasible at corpus scale\. A complete table of agent specifications can be found in Table[A4](https://arxiv.org/html/2609.19585#A2.T4)\.

Prior work recommends low decoding temperatures for classification and labelling tasks to stabilize outputs[Jin et al\. \(2024\)](https://arxiv.org/html/2609.19585#bib.bib33)\. Because our stages demand both high precision and high recall, including sharp inclusion/exclusion at extraction and unambiguous tagging decisions downstream, we set temperaturet=0t=0\(greedy decoding\) for all three agentic stages\. All other generation settings were held constant across models, and for open\-weight models this produces byte\-identical outputs across runs\.

Table A4:Model agents included in the candidate pool\. Models vary in provider, domain specialization, serving environment, parameter count, and architecture\.
## Appendix CBenchmark Size Selection Via Power Analysis

We fix the number of manually evaluated documents using an a priori power analysis for a chi\-square test[Cohen \(1988\)](https://arxiv.org/html/2609.19585#bib.bib43)\. The required sample size is

N=λ∗w2,N=\\frac\{\\lambda^\{\*\}\}\{w^\{2\}\},\(1\)
wherewwis Cohen’s effect\-size index andλ∗\\lambda^\{\*\}is the noncentrality parameter at which the noncentral chi\-square distributionχ2​\(d​f,λ\)\\chi^\{2\}\(df,\\lambda\)reaches the target power relative to the critical valueχ1−α2​\(d​f\)\\chi^\{2\}\_\{1\-\\alpha\}\(df\)\.

For a pairwise comparison with one degree of freedom, we targeted a large effect size \(w=0\.5w=0\.5\), a significance level ofα=0\.05\\alpha=0\.05, and power of1−β=0\.951\-\\beta=0\.95\. This gives

λ∗=\(z1−α/2\+z1−β\)2=12\.99,\\lambda^\{\*\}=\\left\(z\_\{1\-\\alpha/2\}\+z\_\{1\-\\beta\}\\right\)^\{2\}=12\.99,\(2\)
and therefore

N=12\.990\.52=12\.990\.25=51\.98\.N=\\frac\{12\.99\}\{0\.5^\{2\}\}=\\frac\{12\.99\}\{0\.25\}=51\.98\.\(3\)
We round this value up toN=52N=52documents\. Applying a finite\-population correction for the full corpus of 14,882 documents produced an adjusted estimate ofNadj=51\.80N\_\{\\mathrm\{adj\}\}=51\.80, leaving the required sample size effectively unchanged\. AtN=52N=52, the achieved power is0\.9500\.950\.

The analysis is conducted using thepwrpackage[Champely \(2020\)](https://arxiv.org/html/2609.19585#bib.bib44)and independently verified using the noncentral chi\-square distribution implemented inSciPy[Virtanen et al\. \(2020\)](https://arxiv.org/html/2609.19585#bib.bib45)\. The sampling unit is the discharge\-summary document\. Thus, the independence assumption applies across sampled documents, and correlations among events within a document do not artificially increase the effective sample size\.

## Appendix DStage 1 & Stage 2 Generation Prompts

We present the concrete definitions for each of the sub\-dimensions of timeline\-robustness\.

We provide the complete prompts used for Stages 1 and 2 in Tables[A5](https://arxiv.org/html/2609.19585#A4.T5)and[A6](https://arxiv.org/html/2609.19585#A4.T6)\. Each stage uses a fixed system prompt and a document\-specific user\-prompt template\. The system prompts were applied unchanged across all candidate models\. At inference time, only the bracketed placeholders were replaced with the corresponding discharge summary, extracted events, and admission anchors\. No model\-specific prompt tuning or manual preprocessing was performed\.

### D\.1Stage 1: Atomic Clinical Event Extraction

The Stage 1 prompt instructs the model to extract all clinical information from the discharge summary as atomic, self\-contained events\. The prompt prioritizes coverage and source faithfulness: repeated mentions are retained, negation is preserved, and the model is instructed not to correct the note, infer unstated information, or add clinical interpretation\. Events are returned in their original document order as an unnumbered list\. The complete prompt is shown in Table[A5](https://arxiv.org/html/2609.19585#A4.T5)\.

### D\.2Stage 2: ISO Date and Certainty Tagging

The Stage 2 prompt assigns one temporal label to each event extracted in Stage 1\. The model receives the full discharge summary as global context, together with the admission date, discharge date, date of birth, patient age, and the ordered list of atomic events\. It is instructed to preserve the event text and assign exactly one of four labels:\[EXACT\],\[APPROX\],\[PRE ADM\], or\[INDETERMINATE\]\. The prompt follows a conservative decision rule: explicitly written dates receive\[EXACT\], dates resolved from anchors or relative expressions receive\[APPROX\], events framed as prior history receive\[PRE ADM\], and\[INDETERMINATE\]is used only when no defensible temporal resolution is available\. The model is explicitly prohibited from guessing dates\. The complete prompt is shown in Table[A6](https://arxiv.org/html/2609.19585#A4.T6)\.

Table A5:Prompt used for atomic clinical event extractionTable A6:Prompt used for date & certainty tagging

## Appendix EStage 1 & Stage 2 Evaluation Rubrics

Stage 2 is evaluated using four complementary dimensions\. Completeness measures whether the extracted timeline captures the source\-grounded clinical events in the discharge summary\. Semantic faithfulness assesses whether each event preserves the meaning of the source, including negation, uncertainty, quantities, and other qualifiers\. Temporal\-tag accuracy evaluates both the assigned certainty category and the resolved date or date range\. Anti\-hallucination identifies events that introduce claims unsupported by the source note\. It’s important to note that anti\-hallucination was evaluated only by human annotators\. Unlike the other dimensions, this criterion requires determining whether any claim in an extracted event lacks support anywhere in the source note\. Prior work shows that LLM\-as\-judge reliability varies substantially across tasks and evaluation properties, and that even strong LLMs remain imperfect at recognizing hallucinated content[Bavaresco et al\. \(2025\)](https://arxiv.org/html/2609.19585#bib.bib32);[Li et al\. \(2023\)](https://arxiv.org/html/2609.19585#bib.bib35)\. This limitation is especially consequential in the clinical domain, where subtle, plausible unsupported claims may be difficult to distinguish from source\-grounded information\. Because a missed hallucination would directly contaminate the clinician\-verified reference set, we reserved this criterion for exhaustive human review rather than treating an LLM’s factuality judgment as ground truth\.

All rubrics are included in Tables[A7](https://arxiv.org/html/2609.19585#A5.T7)to[A11](https://arxiv.org/html/2609.19585#A5.T11)\.

Table A7:Stage\-2 evaluation rubric for Dimension A: Completeness\.Table A8:Stage\-2 evaluation rubric for Dimension B: Semantic Faithfulness\.Table A9:Stage\-2 evaluation rubric for Dimension C: Temporal\-Tag Accuracy \(Part 1 of 2\)\.Table A10:Stage\-2 evaluation rubric for Dimension C: Temporal\-Tag Accuracy \(Part 2 of 2\)\.Table A11:Stage\-2 evaluation rubric for Dimension D: Anti\-Hallucination\.
## Appendix FStage 1 & Stage 2 Evaluation Platform

Human evaluation of Stage\-2 outputs was conducted through a purpose\-built local web application \(Figure[A1](https://arxiv.org/html/2609.19585#A6.F1)\)\. The tool runs as a self\-contained Python program with a static HTML interface, requires no external services or database, and is accessed through SSH port forwarding so protected health information remains on the host machine\. Model outputs are loaded into an ordered annotation queue, and annotators resume from their first unfinished document\.

The interface displays the full discharge summary alongside the model\-generated events in source order, including their temporal and certainty tags\. For each event, annotators independently assess temporal accuracy, certainty\-tag accuracy, semantic faithfulness, and hallucination, with criterion\-specific comment fields for flagged errors\. To capture omissions, the interface also records the number of missing events and the corresponding source passages\.

The platform includes progress tracking, bulk confirmation for all\-correct events, checks that prevent advancing with incomplete ratings, keyboard navigation, and immediate crash\-safe saving to per\-annotator JSON files\. Completed work can therefore be interrupted and resumed without loss\.

![Refer to caption](https://arxiv.org/html/2609.19585v1/stage2_platform.png)Figure A1:Web interface of the Stage\-2 event\-level human annotation platform\. A single subject is shown: the source discharge summary \(left\) beside the model’s Stage\-2 tagged event list \(right\), with per\-event Yes/No controls for the four rubric dimensions \(time tag, type tag, semantic faithfulness, hallucination\) and a footer for logging missed events\. All protected health information and note/event text are redacted \(gray bars\) for this figure\.
## Appendix GStage 1 & Stage 2 Evaluation Agreement and Error Analysis

Table[A12](https://arxiv.org/html/2609.19585#A7.T12)reports agreement between the Gemini 3\.1 Pro evaluator and clinician\-guided human annotations across 15,891 Stage\-2 events from 52 discharge summaries\. We report Cohen’sκ\\kappa, Gwet’s AC1, and positive and negative agreement\. Because confirmed errors are rare relative to correct events, Cohen’sκ\\kappais sensitive to class prevalence; Gwet’s AC1 and the class\-specific agreement measures provide complementary views of agreement under this imbalance\.

Overall agreement is high on the dominant non\-error class: negative agreement reaches0\.9860\.986, while Gwet’s AC1 is0\.9720\.972\. Agreement is lower on the rare error class, with positive agreement of0\.6690\.669\. This indicates that the evaluator and human reviewers agree strongly on which events are correct, but differ more often when identifying or characterizing errors\.

The disagreements are concentrated in temporal accuracy rather than semantic faithfulness\. Although human annotation and the automated judge report nearly identical total numbers of temporal errors \(608608versus602602\), an event\-level breakdown reveals a strong directional asymmetry: human review identified 175 over\-specified tags on pre\-admission medications and demographics, 169 of which the judge failed to flag, whereas judge\-only findings were concentrated in overlapping date\-range and date\-granularity codes \(Appendix[G\.1](https://arxiv.org/html/2609.19585#A7.SS1)\)\. Temporal\-accuracy agreement reachesκ=0\.617\\kappa=0\.617, whereas semantic\-faithfulness negative agreement is1\.0001\.000\. Manual review identified two recurring sources of temporal disagreement\. First, some models anchor home medications to the admission date using\[APPROX\], while clinicians classify them as\[PRE ADM\]; the automated evaluator does not consistently penalize this distinction\. Second, the evaluator sometimes accepts a coarse date range even when the event can be resolved to a more specific in\-stay anchor\. Thus, disagreement is concentrated in borderline cases of pre\-admission status and date granularity rather than being distributed uniformly across the task\.

The integrity of the gold set does not depend on evaluator agreement\. Human annotators exhaustively reviewed all 15,891 events under clinician guidance and corrected every confirmed error after escalating errors to discussion with clinicians, regardless of whether the automated evaluator detected it\. This process produced 52 human\-corrected pairs of discharge summaries and tagged timelines\.

### G\.1Temporal Error\-Type Analysis

To characterize the lower positive agreement on temporal accuracy, we compared the finalized event\-level error codes assigned through exhaustive human review with those assigned by the calibrated LLM judge across the 52\-note benchmark\.

Figure A2:Stage\-2 temporal\-error counts from human annotation and the calibrated LLM judge\. Despite similar totals \(608 versus 602\), human review surfaced 175 over\-specified tags on pre\-admission medications and demographics that the judge almost entirely missed, while judge\-only errors concentrated in overlapping date\-range and granularity categories\.As shown in Figure[A2](https://arxiv.org/html/2609.19585#A7.F2), the aggregate counts are similar: human annotation identified 608 temporal errors, compared with 602 identified by the judge\. Human review found at least one temporal error in 50 of the 52 notes, whereas the judge identified an error in 46\.

The clearest asymmetry concerns over\-specified temporal tags\. Human annotation identified 175 such errors involving pre\-admission medications \(151 events\) and patient demographics \(24 events\), of which the judge failed to flag 169 \(96\.6%\)\. These are not cases in which the human and judge applied different error codes to the same behavior: the judge returned no temporal error at all\. The pattern therefore represents a systematic false\-negative class in which the judge accepts admission\-anchored\[EXACT\]or\[APPROX\]labels for standing or pre\-admission facts that should receive\[PRE ADM\]\.

The reverse asymmetry is qualitatively different\. Of the 200 judge\-only errors, 80 were coded asdate\_range\_wrongand 56 asdate\_too\_coarse\. These codes describe closely related aspects of temporal granularity and may place the boundary between an incorrect range and an insufficiently precise date differently\. Thus, similar aggregate error counts do not indicate equivalent error detection: human review exposes a clinically coherent error class that the judge largely misses, whereas much of the judge’s excess is concentrated around rubric boundaries for range and granularity\.

### G\.2Other Error\-Types and Analysis

The2121non\-temporal errors are minor and mechanical, encompassing just0\.13%0\.13\\%of all15,89115,891events\. Unlike the temporal errors, almost none reflect clinical misunderstanding\. The dominant failure mode is over\-precision: the model quietly hardens hedged or masked source text into confident assertions\. Roughly half of the semantic errors convert uncertainty into fact \(e\.g\., “? <\*clinical event\*\>" becomes “possible <\*clinical event\*\>"\), or add unsupported specificity to medications \(such as "twice a day", "twice daily" appended to a drug the note never dosed\)\. The rest stem from misreading MIMIC’s de\-identification masks and OCR\-garbled tokens\. The 7 hallucinations are similar in spirit — mostly home/admission medications and lab values carrying detail not fully grounded in the source — and the single completeness miss \(1 missing patient date of birth event across 52 notes\) is negligible\. In short, on everything except temporal tagging the model is essentially faithful; its only residual weakness is a tendency to resolve ambiguity in the source rather than preserve it\.

Table A12:Agreement between the Gemini 3\.1 Pro and gold annotation, calculated over 52 discharge summaries producing 15,891 events\. AC1 denotes Gwet’s AC1; Pos\. and Neg\. denote positive and negative agreement on the rare error and dominant non\-error classes, respectively\. Anti\-hallucination is excluded because it was evaluated only by human annotators\.

## Appendix HStage 3 Generation Prompt

The Stage 3 prompt converts the tagged event timeline into a chronological prose summary\. It specifies how temporal tags should be rendered, which clinical information should be included or excluded, how duplicate or overlapping events should be combined, and how factual and temporal fidelity should be preserved\. Because of its length, the complete prompt is presented across three consecutive sub\-tables \(Table[A13](https://arxiv.org/html/2609.19585#A8.T13)\)\.

Table A13:Prompt used for temporal\-guided clinical summarization \(Part 1 of 3\)\.Table A14:Prompt used for temporal\-guided clinical summarization \(Part 2 of 3\)\.Table A15:Prompt used for temporal\-guided clinical summarization \(Part 3 of 3\)\.
## Appendix IStage 3 Evaluation Rubrics

Stage 3 is evaluated by both human annotators and the calibrated LLM evaluator using two shared dimensions\. Faithfulness to Source measures whether each clinical claim in the generated summary is supported by the human\-verified event timeline and preserves its original clinical meaning\. Temporal Accuracy measures whether included events are rendered under the correct date and certainty tag and whether only events with identical temporal tags are grouped together\. The human\-verified timeline serves as the reference of record; the original discharge summary is used only to clarify unclear wording\. Because selective content inclusion is expected in summarization, omitted events are not penalized under either dimension\. The complete rubrics are provided in Tables[A16](https://arxiv.org/html/2609.19585#A9.T16)and[A17](https://arxiv.org/html/2609.19585#A9.T17)\.

Table A16:Stage\-3 evaluation rubric for Dimension A: Faithfulness to Source\.Table A17:Stage\-3 evaluation rubric for Dimension B: Temporal Accuracy\.
## Appendix JStage 3 LLM and Auto\-Evaluator Results

Table A18:Evaluation results for summaries generated by Gemini 2\.5 Pro\. Faithfulness and temporal accuracy are reported as mean rubric bands and percentage scores\. MEDCON precision, recall, and F1 measure clinical\-concept overlap\.As shown in Table[A18](https://arxiv.org/html/2609.19585#A10.T18), faithfulness is saturated: all 52 summaries score band 5, with 5 distorted claims overall\.

Temporal accuracy shows the mean likert band at 4\.75 / 5, macro\-mean 95\.33% of events correctly placed \(micro 96\.17%, 257 errors overall; median 100%\)\.

MEDCON precision \(0\.7250\.725\) exceeds recall \(0\.5190\.519\), F10\.5990\.599\. Asserted content is largely grounded, consistent with the faithfulness band; roughly half of source concepts are omitted\. The judge scores only claims that were made and is therefore blind to this coverage gap, while MEDCON is insensitive to temporal placement\. The two measure disjoint failure axes and are reported together for that reason\.

## Appendix KStage 3 Evaluation Agreement

### K\.1Annotation Protocol and Unit Definitions

Table A19:Additional lexical evaluation for Task C timeline summarization\. We report ROUGE\-1 and ROUGE\-2 F1 underSilver→\\rightarrowSilver,Silver→\\rightarrowGold, andGold\. Results are mean±\\pmbootstrap SD over 2,000 document\-level resamples\. Higher is better\.Human annotators reviewed all 52 benchmark summaries alongside the judge’s output\. The review interface presents the judge’s findings as a list and allows the annotator to mark any finding not a mistake\.

Table A20:Agreement between the automated evaluator and human annotations\.PoP\_\{o\}denotes observed agreement\. Confidence intervals are reported at the 95% level\.
### K\.2Summary\-level Detection Agreement

Table[A20](https://arxiv.org/html/2609.19585#A11.T20)treats each summary as one unit and asks whether the judge and the annotator agree that it contains at least one error\. This framing avoids committing to a claim inventory as the denominator, and at a positive rate of 35% the prevalence index falls to 0\.40, so Cohen’sκ\\kappaand Gwet’s AC1 no longer diverge and both are interpretable\. The judge flagged 18 summaries; the annotator’s adjudication left at least one surviving flag in 13 of them and eliminated all flags in 5\.

### K\.3Error\-count concordance

Table A21:Error\-count concordance between the automated judge and human annotator\. Per\-summary agreement between the judge’s error count and the number of errors upheld by the annotator\. CCC denotes Lin’s concordance correlation coefficient; ICC\(A,1\) denotes the absolute\-agreement intraclass correlation coefficient\.Table[A21](https://arxiv.org/html/2609.19585#A11.T21)compares the number of errors the judge reports in each summary against the number the annotator upheld\. The two disagree sharply in magnitude and agree well in order\. Lin’s concordance correlation coefficient is 0\.103 and ICC\(A,1\) is 0\.105, both near the floor, because the judge over\-reports by a mean of 4\.3 errors per summary with 95% limits of agreement spanning−25\.7\-25\.7to\+34\.3\+34\.3– an interval wider than the error count of any but the most heavily flagged summary\. Spearman’sρ\\rhois 0\.768, however: the judge ranks summaries by error burden much as the annotator does\.

The gap between Spearman’sρ\\rho\(0\.768\) and Pearson’srr\(0\.511\) reflects a single influential summary, subject 7676, at 100 judge errors against 5 upheld; rank\-based statistics absorb it and moment\-based ones do not\. Taken together these results characterize the judge as a triage and ranking instrument rather than a measurement one\. Its absolute error counts should not be reported as validated error rates, and any downstream comparison between generation models should use the judge’s ordering rather than its magnitudes\.

## Appendix LStage 3 Evaluation Platform

For Stage 3, we built a self\-contained, locally\-run web application for human review and correction of Gemini 2\.5 Pro\-generated Stage 3 summaries\.

The review screen has three regions \(Figure[A3](https://arxiv.org/html/2609.19585#A12.F3)\)\. The left panel shows the source material, toggling between the original discharge summary and the Stage 2 gold event list, and is read\-only\. The right panel presents the model’s Stage 3 summary in an editable field: rather than only flagging problems, the reviewer corrects the summary in place to produce a gold summary\. Edits autosave, and an "edited" indicator with a revert\-to\-original control distinguishes corrected from untouched summaries; the correction is stored as a separatestage3\_gold\_summaryfield, preserving the model’s original output for comparison\. The bottom panel surfaces an LLM\-judge evaluation as a scaffold for verification\. The summary is scored on two rubric criteria as listed in Appendix[I](https://arxiv.org/html/2609.19585#A9)\.

![Refer to caption](https://arxiv.org/html/2609.19585v1/stage3_platform.png)Figure A3:The human review and correction platform\. For each subject, the reviewer compares the model’s Stage 3 summary against the source and corrects it in place\. Left: source panel, toggled between the discharge summary and Stage 2 gold events\. Right: the model’s Stage 3 summary in an editable field, saved as a separate field\. Patient\-text regions from the original MIMIC III source materials are redacted \(grey\)\.
## Appendix MQualitative Analysis Cases

In this section, we analyze three clinician\-verified gold cases using WHO categories of mental disorders, symptoms, and risk factors as mentioned in Section[5\.5](https://arxiv.org/html/2609.19585#S5.SS5), grounding each assignment in the gold timeline and treating unestablished links as co\-documented rather than causal\.

### M\.1Case A: Chronic Disability and an In\-Stay Depression

Case A comprises 306 gold events across a 3\-week admission\. The record describes roughly three decades of C5 quadriplegia, an Individual\-level risk factor that Stage 2 resolves to a year\-only\[APPROX\]date some thirty years before admission, and a marked deterioration in ventilatory status during the stay\. The note itself states the sequence, attributing the patient’s depressed mood to that deterioration, so the link is explicit rather than inferred by us\. Depression is documented as a diagnosis: pharmacotherapy for situational depression is started during the admission, and psychiatry is consulted on the 11th hospital day\. The symptoms recorded are suicidal thoughts and refusal of food; at discharge the note records restored hope, concentration and sleep, and continuation of an SSRI\. Several symptoms are therefore attested only through their documented reversal, an implicit form of expression that an extraction system keyed to symptom mentions alone would miss\.

What makes this case instructive is where the onset falls\. The background risk factor begins decades before admission, whereas the depressive episode emerges during the hospital stay, making their temporal relationship difficult to recover from the narrative alone\. Only a calendar\-anchored representation places them on one ordered chain while preserving that separation\. The distinction is also fragile in exactly the way our error analysis predicts: had the onset been tagged\[PRE ADM\]rather than resolved to the admission interval, the clinically salient fact that this depression arose during the stay would disappear from the timeline entirely\. This is the same boundary that accounts for the largest class of temporal disagreements in Section[G](https://arxiv.org/html/2609.19585#A7)\.

### M\.2Case B: Depression Across All Three WHO Risk Levels

Case B contains 494 gold events spanning a 1\.5\-month admission after an intentional overdose\. It is the only case in our sample with documented risk factors across all three WHO levels\. At the individual level, the record describes long\-standing alcohol misuse across the past medical history, social history, and discharge diagnoses\. At the Family & Community level, it notes limited social support, poor self\-care, and the recent loss of a living arrangement with a relative\. At the Structural level, it identifies multiple social stressors and a lack of insurance coverage\. These links are stated directly in the record, which cites them as reasons for involving psychiatry and social work\. The associated symptoms and behaviors include the presenting overdose, two earlier suicide attempts by different means, a discharge diagnosis noting multiple attempts, and a documented reduction in suicidal intent during the admission\.

This case highlights two important features of the reconstruction\. First, the structural risk factors are exactly the kind of information that a fixed clinical schema might omit, even though they are central to the explanation recorded by the care team\. This illustrates the need for the open\-vocabulary representation described in Section[1](https://arxiv.org/html/2609.19585#S1)\. Second, the record contains an unresolved contradiction: psychiatry documents that the patient was not currently depressed during the admission, while depression remains listed among the discharge diagnoses\.

### M\.3Case C: A Risk Cascade Under an Unsettled Diagnosis

Case C contains 227 gold events spanning a 12\-day admission and includes four documented risk factors, only one of which the source explicitly links to the diagnosis\. The record identifies more than two decades of alcohol misuse as clinically relevant to the patient’s diagnosis\. It also documents childhood sexual abuse by a caregiver, job loss attributed to drinking, and living alone after the end of a relationship about a year before admission\. Because the note does not connect these latter factors directly to the disorder, we treat them as co\-documented rather than causal\. Bipolar disorder appears as an established diagnosis, supported by mood\-stabilizing treatment at admission, repeated overdose\-related hospitalizations, a prior psychiatric admission after a suicide attempt, and multiple presentations following suicide threats\.

What distinguishes this case is that the record later revises its own clinical interpretation\. During the admission, psychiatry questions whether the bipolar diagnosis is accurate and considers whether long\-term alcohol use may have contributed to a misdiagnosis\. Most background risk factors are undated and therefore tagged\[PRE ADM\], whereas this diagnostic reconsideration is placed within the admission interval\. Temporal anchoring makes the progression visible: long\-standing background factors, an established diagnosis, and then in\-stay doubt about that diagnosis\. The case also shows that uncertainty is not only temporal; it also concerns causality and diagnosis\. Finally, it demonstrates the limits of surface matching\. During screening, a statement about episodic drinking was incorrectly matched to an eating\-disorder symptom term and was removed after clinician review, illustrating why lexical overlap alone cannot support reliable qualitative interpretation\.

## Appendix NAdditional Results for Evaluation

We present additional evaluation results for the open\-weight models \(Table[A19](https://arxiv.org/html/2609.19585#A11.T19)\)\.

Similar Articles