An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

arXiv cs.CL Papers

Summary

This study develops an ambiguity taxonomy to evaluate large language model performance on clinical registry abstraction from unprocessed EMR data, finding that LLM accuracy is significantly lower than human abstractors and declines as task ambiguity increases.

arXiv:2608.20373v1 Announce Type: new Abstract: Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry (ACC NCDR). In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers (501 pilot; 4,214 validation). In the pilot, candidate data sources per question averaged between 14.6 (SD 13.9) for demographics and 89.2 (SD 56.1) for history and risk factors. In validation, human inter-rater agreement was approximately 98\% while 87\% of LLM answers exactly matched consensus, 2\% partially, and 9\% did not. Mean question-level accuracy was 91.5\% (SD 13.4\%) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\% for Medication/Event Flag to 62\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:14 AM

# An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study
Source: [https://arxiv.org/html/2608.20373](https://arxiv.org/html/2608.20373)
Betsy CastilloCarta Healthcare, San Francisco, CAAndrew Y\. ShinPediatrics, Stanford University School of Medicine, Stanford, CASenior authors\. Correspondence to: David Scheinker,dscheink@stanford\.eduDavid ScheinkerPediatrics, Stanford University School of Medicine, Stanford, CAMedicine, Stanford University School of Medicine, Stanford, CASenior authors\. Correspondence to: David Scheinker,dscheink@stanford\.edu

###### Abstract

Objectives\.To evaluate large language model \(LLM\) performance on unprocessed electronic medical record \(EMR\) data for clinical registry abstraction\.

Methods\.We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry \(ACC NCDR\)\. In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question\-specific document sets\. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question\-specific document sets\. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing\.

Results\.The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers \(501 pilot; 4,214 validation\)\. In the pilot, candidate data sources per question averaged between 14\.6 \(SD 13\.9\) for demographics and 89\.2 \(SD 56\.1\) for history and risk factors\. In validation, human inter\-rater agreement was approximately 98% while 87% of LLM answers exactly matched consensus, 2% partially, and 9% did not\. Mean question\-level accuracy was 91\.5% \(SD 13\.4%\) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96% for Medication/Event Flag to 62% for Event Timing questions\.

Conclusions\.LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors, even with question\-specific document scoping\. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased\.

Keywords:clinical natural language processing; large language model evaluation; inter\-rater reliability; clinical data abstraction; registry quality; CathPCI; EPDI; multi\-institutional study

## Introduction

Large language models \(LLMs\) now perform at expert level on curated clinical vignettes and multiple\-choice medical examinations \[1,2\]\. Although these benchmarks have accelerated progress in clinical natural language processing, they differ substantially from real\-world documentation tasks: most rely on narrowly defined questions with prespecified answers drawn from a single clinical context and summarize performance with a single overall metric, such as accuracy or macro\-averaged F1 score \[3,4\]\. Such aggregates may obscure variation across clinically distinct tasks and may not reflect performance on unprocessed electronic medical record \(EMR\) data\. Two recent studies illustrate the gap: four leading LLMs generating ICD\-9, ICD\-10, and CPT codes at a single academic medical center all fell below 50% exact\-match accuracy against clinician\-assigned codes \[5\], and in a Fast Healthcare Interoperability Resources \(FHIR\)\-compliant virtual EHR environment with 300 clinically derived tasks, the best model achieved a 69\.7% success rate with substantial variation across task categories \[6\]\. The first study covered a single institution and task; the second’s tasks remain structured and its environment artificial\. Both point to the need to evaluate LLMs on real, unprocessed EMR documents, with their fragmented, redundant, and occasionally conflicting information\. Clinical registries provide a practical framework for such evaluation\. A growing body of work applies LLMs to registry\-style abstraction in oncology \[7\], pulmonary embolism \[8\], and cardiovascular report classification \[9\], typically reporting strong aggregate accuracy without explaining how and why performance varies across question types within a single registry\. The American College of Cardiology National Cardiovascular Data Registry \(ACC NCDR\) defines hundreds of structured questions that trained hospital abstractors extract directly from the EMR for quality reporting, research, and regulatory submission \[10\]; these data have underpinned landmark national assessments of cardiovascular care \[11\]\. The questions span a wide spectrum of clinical reasoning, from transcribing a date of birth to synthesizing notes, laboratory results, and narratives into coded answers that may reasonably vary between trained abstractors; registries manage this variability through dual independent abstraction, reconciliation of disagreements, and measurement of inter\-rater reliability \(IRR\), yielding an adjudicated reference standard well suited to evaluating LLMs on real\-world, multi\-source EMR tasks\. In this study we separated two sources of ambiguity as we measured LLM accuracy on unprocessed EMR data across categories of registry questions that demand different levels of clinical reasoning\. Documentation Ambiguity arises across EMR sources, where the same information may appear in multiple notes that may be inconsistent or contradictory; we measure it by tasking the LLM with finding all data sources potentially relevant to each question and identifying whether it can confidently answer from each\. Clinical Ambiguity arises from the question itself, which may permit several defensible interpretations even when the documentation is clear; we measure it with an ambiguity taxonomy in which experienced abstractors categorize questions, before the study, by the required level of clinical reasoning\. We call the relationship between a category’s position in the taxonomy and the model’s accuracy the ambiguity–performance gradient\.

## Methods

### Study Design

We conducted the study in two phases\. The pilot phase, at Institution A, used the ACC NCDR Electrophysiology Device Implant registry \(EPDI\) and included 168 registry questions extracted from 3 de\-identified patient records \(501 question\-patient observations\)\. The validation phase, at Institution B, used the ACC NCDR Cardiac Catheterization and Percutaneous Coronary Intervention registry \(CathPCI v5\.7\.1\) and included 232 registry questions across 25 de\-identified patient records \(4,214 question\-patient observations\)\. The two institutions use different EMR systems and documentation conventions, so the same concept may appear in different note types, under different labels, with varying structure\.

### Reference Standard

At each site, two practicing cardiac registry abstractors, credentialed through the ACC/NCDR certification program, independently answered each question; disagreements were reconciled through discussion to a single consensus answer, which served as the reference standard\. Input data consisted of unprocessed EMR documentation, including clinical, procedure, consultation, and nursing notes, diagnostic imaging reports, discharge and transfer summaries, and available structured data elements\.

### Question Category Classification

In the EPDI pilot, questions were grouped by registry section, a coarser scheme suited to the pilot’s document\-routing objective\. Before the validation study, the abstractors independently grouped each registry question by the level of ambiguity, and corresponding clinical reasoning, its answer demands, assigning each to one of six categories ordered from lowest to highest reasoning demand: Medication/Event Flag \(n = 62\), Binary Clinical Presence \(n = 41\), Administrative \(Transcription\) \(n = 18\), Quantitative Laboratory/Physiologic \(n = 10\), Clinical Interpretation \(n = 22\), and Event Timing \(n = 4\); full definitions are in the Supplementary Methods\. These counts reflect the 157 questions with at least 20 evaluable observations that constitute the analytic set \(Table 3\); all 232 validation questions were classified into the same six categories\. A second reviewer independently classified a random 10% sample of questions; agreement was 92% \(Cohen’sκ\\kappa= 0\.89\), and disagreements were resolved by consensus\.

### LLM Protocol and Document\-Context Targeting

The same model configuration was used in both phases: Claude Sonnet 4\.6 for more complex registry questions and Claude Haiku 3\.5 for simpler questions, accessed via application programming interface within the institution’s secure, HIPAA\-compliant cloud environment; no protected health information left institutional boundaries\. Each prompt contained the verbatim ACC NCDR question definition, including value sets where specified, and a standing instruction to answer when confident and decline when not; for questions with defined value sets, the model selected from the allowed options or explicitly abstained \(full prompt structure in the Supplementary Methods\)\. In the pilot phase, we prompted the LLM to identify all documentation sources within each record that could plausibly contain the information needed for each question and, for each source, whether it could use it to answer\. Across sequential iterations, these findings were developed into a structured mapping between question type and documentation context\. In the validation phase, in order to isolate the effect of question ambiguity, the model received only the documents in each question’s assigned context and produced a single answer for every patient record to which the question applied\. Answers were scored against the adjudicated consensus as exact matches, partial matches \(the answer contained the reference value\), mismatches, or declinations\.

### Outcomes and Statistical Analysis

The primary outcome was overall accuracy, the proportion of non\-abstained answers achieving an exact or partial match\. Secondary outcomes were mean question\-level accuracy among the 157 questions with at least 20 answers \(some questions do not apply to all patients\) and category\-level accuracy, the mean question\-level accuracy within each category, with sensitivity tested by repeating the analysis at a 10\-answer threshold\. Differences across categories were tested with a Kruskal\-Wallis H test and pre\-specified pairwise Mann\-Whitney U tests with rank\-biserial effect sizes\. The pilot, with 3 patients, is reported descriptively; to characterize what drives answerability, we fit linear models of percent answerable on the number of candidate sources, with and without registry section as a categorical term \(software details in the Supplementary Methods\)\.

## Results

### Pilot Study: Identifying Question\-Specific Document Subsets

The demographics section had the lowest average number of candidate documentation sources per question, 14\.6 \(SD 13\.9\), and history and risk factors the highest, 89\.2 \(SD 56\.1\) \(Table 1\)\. Record size, documentation density, and candidate\-source counts differed substantially across patients \(Supplementary Table S1\)\. The answerable rate, the proportion of identified sources from which the model could generate an answer, was lowest for Episode of Care questions, 40\.3% \(SD 28\.5%\), and highest for laboratory questions, 79\.0% \(SD 18\.5%\) \(Table 1\)\. Although the model identified many potentially relevant sources per question, a large fraction could not be used to answer; in several sections fewer than half of the candidate sources supported an answer \(Table 1\)\. A linear model of the association between the answerable rate and the number of candidate sources and the registry section explained 31% of the variation \(adjusted R2= 0\.31\) and found that section membership was strongly associated with answerability \(F = 18\.3; p ¡ 0\.0001\) while the number of sources was not\. This is reported descriptively given the three\-patient pilot \(Supplementary Figure S1\)\. These results were used to construct the document\-context targeting scheme for the validation study: for each group of questions, the abstractors identified the document set to which the LLM’s attention was restricted \(Table 2; full retrieval specification in Supplementary Table S2\)\.

### Validation Study: Overall Answer and Match Rates

In the validation study, the model generated answers for 99% of observations \(4,171 of 4,214\) and declined for 1% \(43\)\. Across all observations, 87% were exact matches to the adjudicated abstractor consensus, 2% were partial matches, 9% were mismatches, and 1% were abstentions\. The analytical sample comprised 9,430 answers generated by experienced clinical abstractors and reconciled to 4,715 consensus answers \(501 pilot, 4,214 validation\), which served as the reference standard for evaluating LLM performance\. The two human abstractors agreed on approximately 98% of questions before reconciliation; the remainder, plus a few questions flagged during quality review, were reconciled to a single consensus answer\. Weighted aggregate accuracy across all observations was 89\.6%\. Mean question\-level accuracy, counting each of the 232 questions equally, was 84\.0%, a difference of 5\.6 percentage points\. For the 157 questions with at least 20 evaluable observations, mean accuracy was 91\.5% \(SD 13\.4%; median 96\.0%; range 12\.0–100\.0%\): 61 questions \(39%\) had 100% accuracy, 18 \(11\.5%\) fell below 80%, and 6 \(3\.8%\) were at or below 60% \(Figure 1\)\. Relaxing the threshold to at least 10 observations \(165 questions\) left the distribution essentially unchanged \(mean 90\.6%; median 96\.0%; 22 questions, 13\.3%, below 80%\); the eight additional questions were sampled more sparsely and skewed slightly harder \(Supplementary Table S3 gives the full ranked distribution\)\. Mean question\-level accuracy fell steadily as category complexity rose, from 96\.1% for Medication/Event Flag questions to 62\.0% for Event Timing questions \(Table 3\), a gap of 34\.1 percentage points\. Cross\-category differences were highly significant \(Kruskal\-Wallis p ¡ 0\.0001\)\. Pairwise contrasts confirmed the gradient: Medication/Event Flag versus Clinical Interpretation \(Mann\-Whitney p ¡ 0\.0001; rank\-biserial r = 0\.81\) and Binary Clinical Presence versus Clinical Interpretation \(p ¡ 0\.0001\) showed large effects, while the two most complex categories did not differ from each other \(p = 0\.39\)\.

### Accuracy in Clinically Ambiguous Questions

Clinical Interpretation and Event Timing questions were over\-represented among questions with accuracy below 80%: the two categories represent 17% of the analytic set \(26 of 157 questions\) but accounted for 12 of the 18 such questions \(67%\), whereas among the 114 questions above 90% accuracy only 6 \(5%\) came from these categories, an approximately 13\-fold enrichment \(Figure 2\)\. Questions with perfect accuracy included binary demographics and comorbidities such as date of birth, diabetes, hypertension, and sex \(all 100% accuracy\), along with most medication\-administration and post\-procedure event flags \(Table 4\); one notable exception was history of myocardial infarction, a Binary Clinical Presence question on which models achieved 68% accuracy\. Of the six questions at or below 60%, five were interpretive or timing questions: cardiac cath lab visit indication \(12%\), arrival date/time \(20%\), the CSHA frailty scale and cardiovascular instability \(both 56%\), and discharge date/time \(60%\); the sixth was an administrative name question \(48%\)\.

### Abstention and Precision

The model’s rate of declining to answer rose with category complexity, from 0\.2% on Medication/Event Flag questions to 4\.9% on Event Timing questions\. However, on questions with accuracy below 80%, recall remained above 90% while precision fell below 60%\.

## Discussion

In the validation phase of the evaluation of frontier LLMs across two ACC/NCDR registries at two academic medical centers, overall LLM accuracy was 89\.6% on 4,214 questions, falling from 96\.1% on simple binary questions to 62\.0% on questions requiring the most clinical reasoning to resolve ambiguity\. The two most complex categories, Clinical Interpretation and Event Timing, were about thirteen times over\-represented among the questions with the lowest accuracy\. Although the model was instructed to abstain when unsure, it seldom did including on questions on which it achieved very low accuracy\.

### Unprocessed EMR data

Models evaluated on clinical vignettes are asked a clean question with one adjudicated answer from a single clinical context\. The medical record is not clean: the same fact can sit in a progress note, a discharge summary, and a nursing flowsheet, recorded differently under different labels across encounters months apart\. In our pilot, a single history or risk\-factor question drew on an average of 89 candidate sources per patient, fewer than half of which the model judged usable\. These findings align with recent reports that feeding an LLM more of the chart made its diagnoses worse, because models are sensitive to both the amount and order of information \[12\]\. Because the pilot scoped each question to a defined document set, the validation phase could examine performance as a function of complexity alone\.

### Implications for evaluating and interpreting LLM performance

Our findings suggest that model performance should be reported by category and by question, rather than in aggregate\. Our weighted accuracy exceeded our question\-mean accuracy by 5\.6 percentage points because the easy questions recur in many records, so improving on common, easy questions raised the weighted score more than improving on rare, clinically challenging questions\. This matches, and adds clinical context to, concerns previously raised about aggregating across questions for conversational tasks \[13\]\. The performance gradient is not peculiar to our registries: earlier studies found near\-perfect accuracy on left ventricular function but markedly lower accuracy on culprit vessel in the same angiography reports and concordance on thyroid pathology reports highest on simple binary and categorical questions and falling with textual interpretation \[9,14\]\. The agreement between model performance and the ambiguity taxonomy presented here suggests that it may serve as a useful framework for evaluation, and that other schemes should preserve the divide between direct extraction and interpretive synthesis\. The way the model erred is itself informative\. On the hard questions it answered anyway, often incorrectly, the strategy that training on single labels graded by accuracy rewards, since producing the most probable label beats conceding ignorance \[15\]\. A model trained to decline when unsure would instead let those questions be routed to a person\. The models fared worst on Event Timing questions\. In those, the right answer usually depends on reconciling several documents rather than reading any one, since a timestamp may differ across a triage note, an EMS handoff, a nursing intake, and a billing record\. Work on verifying facts against the EHR calls this a needle\-in\-a\-haystack problem, with evidence scattered through the chart and hard to localize even with retrieval \[16\]\. The ambiguity taxonomy is consistent with other recent work\. Medication extraction, with end\-to\-end F1 scores of 0\.69 and 0\.82 in French and English \[17\], is a Medication/Event Flag task checkable against structured records\. Systematic\-review abstract screening, where LLM sensitivity approaches 1\.0 \[18\], is a binary\-presence question with a constrained answer set and a single self\-contained source\. Some studies find that the hard questions are hard because they are fundamentally ambiguous, in the sense that experts disagree about the answer \[19,20\]\. The data argue against this, though they do not settle it: if the model’s 12% to 20% accuracy on the hardest questions reflected genuine disagreement about truth, the abstractors should have disagreed nearly as often, and they did not\. We think it more likely that the documentation on hard questions is multi\-source and sometimes contradictory, and reaching the answer takes reconciliation: the abstractors reconciled by thinking or talking the case through\. The 98% agreement suggests that this was possible for most questions, but it is an aggregate lifted by the many easy questions\. Measuring category\-specific and question\-specific human IRR is an important area for future work\. A model graded only on a single gold label learns that guessing beats abstaining, which is one explanation for why it answered confidently on the questions it was most likely to get wrong \[15\]\. Our findings align with a decade of work on training models to handle ambiguity and suggest a specific clinical application: allow models to measure their own uncertainty and index against human uncertainty, what we call IRR\-indexed loss\. Model IRR, or semantic entropy, is measured by sampling several answers to the same question, clustering them by meaning, and using the spread across clusters to estimate how unsure the model is\.\[21\] Currently, it is used primarily to flag potential hallucinations\. Along with human IRR, it could be used to direct training and evaluation\. When human IRR is high, the loss should push the model to spend more time on inference: read more carefully, reconcile across sources, and reason more deeply\. When human IRR is low, the question is ambiguous and the fiction of a single correct answer will encourage confident guessing or hallucination; the target should be a distribution over the defensible answers\.

### Limitations

The number of patients in both studies was relatively small \(3 pilot, 25 validation patients\), so per\-question estimates are wide\. For the human IRR, we rely on category\-level averages; per\-category IRR, a key follow\-up, was not recorded\. The reference standard reflects credentialed abstractors, so other registries and later models will need their own evaluation\.

## Conclusion

Across two ACC/NCDR registries at two institutions with 98% human IRR, LLM accuracy fell with question ambiguity, from 96% on the simplest category to 62% on the most complex\. Studies of LLM performance on clinical abstraction tasks should be reported and interpreted across categories of different levels of ambiguity, rather than in aggregate\.Acknowledgments:The authors thank the clinical abstractors of Carta Healthcare for the dual abstraction and adjudication that produced the reference standard, the Carta engineering team for the evaluation infrastructure, and the participating institutions for access to the de\-identified records used in this study\.

Funding:No funding

Conflict of interest:Authors JM and BC are employed by Carta Healthcare, the company that conducted the study\. Authors AS and DS advise Carta Healthcare\.

Data availability:Question\-level performance data and analysis code are available from the corresponding author upon reasonable request\.

Table 1:Candidate documentation sources per question and answerable rate\. Pilot study \(Institution A, EPDI registry, n = 3 patients\); values aggregated across patients by registry section\.![Refer to caption](https://arxiv.org/html/2608.20373v1/x1.png)Figure 1:Question\-level accuracy by category, ordered left to right by increasing clinical ambiguity \(157 CathPCI questions, n≥\\geq20 observations\)\. Boxes show interquartile range and median; whiskers extend to 1\.5×\\timesIQR; points show individual questions\. Dashed line marks the overall mean \(91\.5%\)\.Table 2:Document\-context targeting derived from the pilot\. Each row gives a set of registry questions, the documentation context to which the LLM was restricted, the EMR resources that context retrieves, and the number of CathPCI validation\-study questions routed to it\.Table 3:Accuracy by category \(CathPCI validation study, questions with at least 20 observations\), ordered from simplest to most complex\. Human inter\-rater reliability was approximately 98% in aggregate across the validation set; question\-level IRR by category was not separately recorded\.Table 4:Selected questions with very high and very low accuracy, with interpretive annotations\.![Refer to caption](https://arxiv.org/html/2608.20373v1/x2.png)Figure 2:Composition of each accuracy tier by question type\. Bars show the share of questions in each tier classified as Clinical Interpretation or Event Timing \(red\) versus the other four categories \(gray\), with counts inset\. These two categories make up 5% of the≥\\geq90% tier but 67% of the ¡80% tier, an approximately 13\-fold enrichment\.![Refer to caption](https://arxiv.org/html/2608.20373v1/x3.png)Figure 3:Supplementary Figure S1\.Section\-level mean answerable rate versus mean candidate sources per question in the EPDI pilot \(bubble area proportional to number of questions\)\. Answerable rate varies by section but shows no consistent relationship with the number of candidate sources, indicating that answerability is driven by question type rather than source volume\.
## References

1. \[1\]Singhal K, Tu T, Gottweis J, et al\. Toward expert\-level medical question answering with large language models\. Nat Med\. 2025;31\(3\):943–950\. doi:10\.1038/s41591\-024\-03423\-7\.
2. \[2\]Goh E, Gallo RJ, Strong E, et al\. GPT\-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial\. Nat Med\. 2025;31\(4\):1233–1238\. doi:10\.1038/s41591\-024\-03456\-y\.
3. \[3\]Wornow M, Xu Y, Lehman EP, et al\. The shaky foundations of large language models and foundation models for electronic health records\. NPJ Digit Med\. 2023;6\(1\):135\.
4. \[4\]Bedi S, Liu Y, Orr\-Ewing L, et al\. Testing and evaluation of health care applications of large language models: a systematic review\. JAMA\. 2025;333\(4\):319\-328\. doi:10\.1001/jama\.2024\.21700\.
5. \[5\]Soroush A, Glicksberg BS, Zimlichman E, et al\. Large language models are poor medical coders: benchmarking of medical code querying\. NEJM AI\. 2024;1\(5\):AIdbp2300040\. doi:10\.1056/AIdbp2300040\.
6. \[6\]Jiang Y, Black KC, Geng G, Park D, Ng AY, Chen JH\. MedAgentBench: a virtual EHR environment to benchmark medical LLM agents\. NEJM AI\. 2025;2\(2\):AIdbp2500144\. doi:10\.1056/AIdbp2500144\.
7. \[7\]Enikeev R, Moldovan M, Chu M, Amalraj A, Koli PP, Syed Abdul S, Sivaraj H, Iqbal U, Toh CK\. Privacy\-preserving large language model deployment for oncology registry abstraction: structure\-aware evaluation in a real\-world clinical setting\. medRxiv\. 2026:2026\.05\.18\.26353541\. doi:10\.64898/2026\.05\.18\.26353541\.
8. \[8\]Alwakeel M, et al\. Evaluating large language models for automated clinical abstraction in pulmonary embolism registries: performance across model sizes, versions, and parameters\. In: Wu S, Shabestari B, Xing L, eds\. Applications of Medical Artificial Intelligence \(AMAI 2025\)\. Lecture Notes in Computer Science, vol 16206\. Cham: Springer; 2026\. doi:10\.1007/978\-3\-032\-09569\-5\_21\.
9. \[9\]van der Loo W, van der Valk V, van den Broek T, Atsma D, Staring M, Scherptong R\. Large language models for structured cardiovascular data extraction: a foundation for scalable research and clinical applications\. Eur Heart J Digit Health\. 2026;7\(2\):ztaf127\. doi:10\.1093/ehjdh/ztaf127\.
10. \[10\]Brindis RG, Fitzgerald S, Anderson HV, et al\. The American College of Cardiology–National Cardiovascular Data Registry \(ACC\-NCDR\): building a national clinical data repository\. J Am Coll Cardiol\. 2001;37\(8\):2240–2245\. doi:10\.1016/S0735\-1097\(01\)01372\-9\.
11. \[11\]Chan PS, Patel MR, Klein LW, et al\. Appropriateness of percutaneous coronary intervention\. JAMA\. 2011;306\(1\):53–61\. doi:10\.1001/jama\.2011\.916\.
12. \[12\]Hager P, Jungmann F, Holland R, et al\. Evaluation and mitigation of the limitations of large language models in clinical decision\-making\. Nat Med\. 2024;30\(9\):2613–2622\. doi:10\.1038/s41591\-024\-03097\-1\.
13. \[13\]OpenAI\. HealthBench: evaluating large language models across real\-world healthcare conversations\. 2025\. https://openai\.com/index/healthbench/
14. \[14\]Lee D, Vaid A, Menon KM, Freeman R, Matteson DS, Marin ML, Nadkarni GN\. Using large language models to automate data extraction from surgical pathology reports: retrospective cohort study\. JMIR Form Res\. 2025;9:e64544\. doi:10\.2196/64544\.
15. \[15\]Kalai AT, Nachum O, Vempala SS, Zhang E\. Why language models hallucinate\. arXiv:2509\.04664 \[cs\.CL\]\. 2025\.
16. \[16\]Chung P, Swaminathan A, Goodell AJ, Kim Y, Momsen Reincke S, Han L, et al\. Verifying facts in patient care documents generated by large language models using electronic health records\. NEJM AI\. 2026;3\(1\):AIdbp2500418\.
17. \[17\]Fabacher T, Sauleau EA, Arcay E, et al\. Efficient extraction of medication information from clinical notes: an evaluation in 2 languages\. J Am Med Inform Assoc\. 2025;32\(12\):1855–1864\. doi:10\.1093/jamia/ocaf113\.
18. \[18\]Sanghera R, Thirunavukarasu AJ, El Khoury M, et al\. High\-performance automated abstract screening with large language model ensembles\. J Am Med Inform Assoc\. 2025;32\(5\):893–904\. doi:10\.1093/jamia/ocaf050\.
19. \[19\]O’Malley KJ, Cook KF, Price MD, Wildes KR, Hurdle JF, Ashton CM\. Measuring diagnoses: ICD code accuracy\. Health Serv Res\. 2005;40\(5 Pt 2\):1620–1639\.
20. \[20\]Reamaroon N, Sjoding MW, Lin K, Iwashyna TJ, Najarian K\. Accounting for label uncertainty in machine learning for detection of acute respiratory distress syndrome\. IEEE J Biomed Health Inform\. 2019;23\(1\):407–415\. doi:10\.1109/JBHI\.2018\.2810820\.
21. \[21\]Farquhar, S\., Kossen, J\., Kuhn, L\. and Gal, Y\., 2024\. Detecting hallucinations in large language models using semantic entropy\. Nature, 630\(8017\), pp\.625\-630\.

## Supplementary Appendix

An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi\-site prospective study

### Supplementary Methods

S1\. Registry Background\.The pilot phase used the ACC NCDR Electrophysiology Device Implant registry \(EPDI; formerly the Implantable Cardioverter\-Defibrillator registry\)\. The validation phase used the ACC NCDR Cardiac Catheterization and Percutaneous Coronary Intervention registry \(CathPCI v5\.7\.1\), which captures percutaneous coronary intervention data across more than 1,600 US hospitals\. In the EPDI pilot, questions were grouped by registry section: Demographics; Episode of Care; History and Risk Factors; Diagnostic Studies; Labs; Procedure Information; Device Implant/Explant; Lead Assessment; Intra/Post\-Procedure Events; and Discharge\. This coarser scheme was suited to the pilot’s document\-routing objective\.

S2\. Ambiguity Taxonomy: Full Category Definitions\.In the validation study, each registry question was assigned to one of six categories, ordered from lowest to highest reasoning demand\. Medication/Event Flag \(n = 62\) comprised questions asking whether a specific medication was given or a specific peri\-procedural event occurred, typically verifiable directly from structured documentation\. Binary Clinical Presence \(n = 41\) comprised yes/no clinical conditions or comorbidities defined by ACC/NCDR coding criteria\. Administrative/Transcription \(n = 18\) comprised demographic or identification data transcribed directly from the record without clinical interpretation\. Quantitative Laboratory/Physiologic \(n = 10\) comprised numeric laboratory or physiologic measurements, occasionally requiring limited disambiguation when multiple sources exist\. Clinical Interpretation \(n = 22\) required synthesis of multiple data sources and clinical judgment, such that reasonable disagreement between trained abstractors could occur\. Event Timing \(n = 4\) required determining the correct clinical date or timestamp by identifying the moment from documentation that may conflict across source types\.

S3\. Prompt Structure\.Each prompt consisted of three components: a general instruction simulating a clinical abstractor role, the verbatim ACC NCDR question definition \(including value sets where specified\), and a structured output format\. Across every record\-question pair, the model was given a single standing instruction, to answer when confident and to decline when not\. For questions with defined value sets, the model was required to select from the allowed options or explicitly abstain; free\-text substitutions were not permitted\.

S4\. Software\.Analyses used Python 3\.10 \(NumPy 1\.24, Pandas 2\.0, SciPy 1\.10\)\.

### Supplementary Material

Supplementary Table S1\. Pilot Study: Per\-Patient DetailPilot data \(Institution A, EPDI registry\) disaggregated by patient, showing how candidate\-source counts and answerable rates varied across the three patient records\. The cross\-patient variance illustrates why source\-data ambiguity must be controlled before measuring inherent task ambiguity: the same question can have very different document footprints across patients\.

PatientRegistry SectionN QuestionsTotal Answers \(Mean±\\pmSD\)Answerable % \(Mean±\\pmSD\)Pat 1A\. Demographics93\.0±\\pm0\.044\.4%±\\pm23\.6%Pat 1B\. Episode of Care93\.7±\\pm1\.636\.7%±\\pm33\.2%Pat 1C\. History and Risk Factors66157\.2±\\pm22\.154\.4%±\\pm14\.8%Pat 1D\. Diagnostic Studies1357\.0±\\pm79\.067\.4%±\\pm12\.9%Pat 1E\. Labs318\.7±\\pm0\.681\.8%±\\pm18\.1%Pat 1F\. Procedure Information86\.0±\\pm2\.549\.3%±\\pm15\.4%Pat 1G\. Device Implant/Explant237\.8±\\pm0\.460\.6%±\\pm22\.8%Pat 1H\. Lead Assessment117\.6±\\pm0\.554\.2%±\\pm23\.4%Pat 1I\. Intra/Post\-Procedure Events1514\.5±\\pm0\.652\.9%±\\pm12\.7%Pat 1J\. Discharge108\.3±\\pm7\.164\.7%±\\pm14\.3%Pat 2A\. Demographics933\.7±\\pm0\.754\.6%±\\pm18\.2%Pat 2B\. Episode of Care934\.7±\\pm1\.642\.5%±\\pm28\.2%Pat 2C\. History and Risk Factors6685\.0±\\pm12\.938\.9%±\\pm11\.0%Pat 2D\. Diagnostic Studies1365\.1±\\pm32\.077\.2%±\\pm20\.8%Pat 2E\. Labs333\.7±\\pm0\.692\.9%±\\pm12\.2%Pat 2F\. Procedure Information850\.8±\\pm24\.150\.4%±\\pm21\.9%Pat 2G\. Device Implant/Explant2367\.3±\\pm9\.563\.3%±\\pm20\.0%Pat 2H\. Lead Assessment1162\.8±\\pm12\.562\.9%±\\pm15\.7%Pat 2I\. Intra/Post\-Procedure Events1554\.6±\\pm0\.853\.6%±\\pm14\.1%Pat 2J\. Discharge1020\.9±\\pm23\.681\.4%±\\pm16\.9%Pat 3A\. Demographics97\.0±\\pm0\.052\.4%±\\pm16\.0%Pat 3B\. Episode of Care97\.1±\\pm1\.141\.7%±\\pm27\.0%Pat 3C\. History and Risk Factors6625\.3±\\pm3\.233\.2%±\\pm8\.5%Pat 3D\. Diagnostic Studies1323\.2±\\pm3\.747\.4%±\\pm18\.1%Pat 3E\. Labs323\.0±\\pm0\.062\.3%±\\pm13\.3%Pat 3F\. Procedure Information89\.6±\\pm3\.951\.2%±\\pm15\.7%Pat 3G\. Device Implant/Explant2312\.4±\\pm1\.257\.0%±\\pm21\.1%Pat 3H\. Lead Assessment1111\.9±\\pm1\.556\.0%±\\pm15\.7%Pat 3I\. Intra/Post\-Procedure Events1518\.0±\\pm0\.035\.2%±\\pm9\.5%Pat 3J\. Discharge108\.7±\\pm9\.079\.0%±\\pm27\.6%Supplementary Table S2\. Document\-Context Targeting: Detailed Retrieval SpecificationsFor each documentation context summarized in Table 2, this table gives the FHIR resources, document types, and example questions used by the LLM during the validation study\.

Document ContextFHIR Resources / Document Types Retrieved\# QuestionsExample QuestionsAll DocumentsFull patient record, no filters applied3Arrival Date/Time; Discharge Date/Time; Procedure Start Date/TimePatient Demographics & H&PPatient resource; Encounter; History & Physical; Consult notes; Discharge summaries16Patient Last/First Name; DOB; Sex; Race; Hispanic Origin; Health Insurance; ZIP CodeBroad Clinical NotesEncounter; Progress notes; Discharge summaries; H&P; Consult; Procedure; Nursing; Transfer summaries111Admission Source; Prior MI; Hypertension; Diabetes; Dyslipidemia; NYHA Class; Prior PCI; Tobacco Use; Stress Test Type/ResultLab ResultsLab Observations; Diagnostic reports \(CBC, metabolic panels, cath\-lab labs\)12Hemoglobin; Sodium; Creatinine \(pre/post\-procedure\); Potassium; BUN; Calcium Score AssessedMedicationsMedicationRequest; MedicationAdministration; MedicationStatement; Discharge summaries5Discharge Medication Code \(RxNorm\); Pre\-procedure / Procedure Medication Administration; GDMT Maximum DoseDischarge SummaryDischarge summary documents only10CABG Status; PCI Date; Discharge Status/Location/Hospice; Comfort Care; Cardiac Rehab ReferralEchocardiography & ImagingEchocardiography; cardiac ultrasound reports4Prior LVEF Assessed; Most Recent LVEF %/Date; ECG ResultsSupplementary Table S3\. Complete Ranked Question\-Level Performance \(n≥\\geq20\)Full 157\-question performance distribution from the CathPCI validation study, ordered by ascending accuracy\. PR Gap = Recall \- Precision\. Question identifiers are given as the registry element names\.

RankQuestion IDCategoryNAccuracyPrecisionRecallPR Gap1cath\_lab\_visit\_indicationClinical Interpretation2512\.0%12\.0%100\.0%88\.0%2arrival\_date\_timeEvent Timing2520\.0%20\.0%100\.0%80\.0%3adm\_l\_nameAdministrative2548\.0%50\.0%92\.3%42\.3%4csha\_scaleClinical Interpretation2556\.0%56\.0%100\.0%44\.0%5cv\_instabilityClinical Interpretation2556\.0%56\.0%100\.0%44\.0%6dc\_date\_timeEvent Timing2560\.0%60\.0%100\.0%40\.0%7pci\_indicationClinical Interpretation2568\.0%68\.0%100\.0%32\.0%8hx\_miBinary Clinical Presence2568\.0%68\.0%100\.0%32\.0%9pre\_proc\_lvef\_assessedClinical Interpretation2568\.0%68\.0%100\.0%32\.0%10cabg\_planned\_dcClinical Interpretation2470\.8%94\.4%73\.9%20\.5%11multi\_vessel\_dzClinical Interpretation2572\.0%72\.0%100\.0%28\.0%12pci\_statusClinical Interpretation2572\.0%72\.0%100\.0%28\.0%13ec\_assess\_methodClinical Interpretation2572\.0%72\.0%100\.0%28\.0%14weightQuantitative Lab/Physiologic2572\.0%78\.3%90\.0%11\.7%15d\_cath\_l\_nameAdministrative2475\.0%75\.0%100\.0%25\.0%16dyslipidemiaBinary Clinical Presence2576\.0%76\.0%100\.0%24\.0%17dc\_med\.41549009\.dc\_med\_adminMedication/Event Flag2378\.3%78\.3%100\.0%21\.7%18pre\_proc\_timiClinical Interpretation2479\.2%86\.4%90\.5%4\.1%19prior\_dx\_angio\_procBinary Clinical Presence2580\.0%80\.0%100\.0%20\.0%20post\_proc\_creatQuantitative Lab/Physiologic2080\.0%80\.0%100\.0%20\.0%21stenosis\_prior\_treatClinical Interpretation2580\.0%80\.0%100\.0%20\.0%22stress\_test\_resultClinical Interpretation2080\.0%84\.2%94\.1%9\.9%23procedure\_end\_date\_timeEvent Timing2580\.0%80\.0%100\.0%20\.0%24stress\_test\_typeClinical Interpretation2080\.0%84\.2%94\.1%9\.9%25dc\_med\.33252009\.dc\_med\_adminMedication/Event Flag2382\.6%82\.6%100\.0%17\.4%26dc\_med\_reconciledClinical Interpretation2382\.6%86\.4%95\.0%8\.6%27dc\_med\_recon\_completedClinical Interpretation2382\.6%82\.6%100\.0%17\.4%28dc\_card\_rehabClinical Interpretation2483\.3%83\.3%100\.0%16\.7%29ipp\_event\.385494008\.post\_proc\_occurredMedication/Event Flag2584\.0%84\.0%100\.0%16\.0%30pre\_proc\_med\.1191\.pre\_proc\_med\_adminMedication/Event Flag2584\.0%84\.0%100\.0%16\.0%31hgbQuantitative Lab/Physiologic2185\.7%90\.0%94\.7%4\.7%32dc\_med\.372913009\.dc\_med\_adminMedication/Event Flag2387\.0%87\.0%100\.0%13\.0%33ipp\_event\.1000142371\.post\_proc\_occurredMedication/Event Flag2588\.0%88\.0%100\.0%12\.0%34ipp\_event\.1000142419\.post\_proc\_occurredMedication/Event Flag2588\.0%88\.0%100\.0%12\.0%35hx\_cvdBinary Clinical Presence2588\.0%88\.0%100\.0%12\.0%36prior\_padBinary Clinical Presence2588\.0%88\.0%100\.0%12\.0%37pre\_proc\_med\.33252009\.pre\_proc\_med\_adminMedication/Event Flag2588\.0%88\.0%100\.0%12\.0%38pcil\_nameAdministrative2588\.0%88\.0%100\.0%12\.0%39cardiac\_ctaBinary Clinical Presence2588\.0%88\.0%100\.0%12\.0%40ipp\_event\.1000142440\.post\_proc\_occurredMedication/Event Flag2588\.0%88\.0%100\.0%12\.0%41proc\_systolic\_bpQuantitative Lab/Physiologic2588\.0%91\.7%95\.6%4\.0%42procedure\_start\_date\_timeEvent Timing2588\.0%88\.0%100\.0%12\.0%43family\_hx\_cadBinary Clinical Presence2588\.0%91\.7%95\.6%4\.0%44pre\_proc\_creatQuantitative Lab/Physiologic2290\.9%90\.9%100\.0%9\.1%45anti\_arrhy\_therapyClinical Interpretation2290\.9%95\.2%95\.2%0\.0%46dc\_med\.1116632\.dc\_med\_adminMedication/Event Flag2391\.3%91\.3%100\.0%8\.7%47pre\_proc\_med\.96302009\.pre\_proc\_med\_adminMedication/Event Flag2592\.0%92\.0%100\.0%8\.0%48calcium\_score\_assessedBinary Clinical Presence2592\.0%92\.0%100\.0%8\.0%49heightQuantitative Lab/Physiologic2592\.0%100\.0%92\.0%8\.0%50access\_siteBinary Clinical Presence2592\.0%92\.0%100\.0%8\.0%51pre\_proc\_med\.372913009\.pre\_proc\_med\_adminMedication/Event Flag2592\.0%92\.0%100\.0%8\.0%52pre\_proc\_med\.41549009\.pre\_proc\_med\_adminMedication/Event Flag2592\.0%92\.0%100\.0%8\.0%53ipp\_event\.22298006\.post\_proc\_occurredMedication/Event Flag2592\.0%92\.0%100\.0%8\.0%54ipp\_event\.84114007\.post\_proc\_occurredMedication/Event Flag2592\.0%92\.0%100\.0%8\.0%55concom\_procBinary Clinical Presence2592\.0%92\.0%100\.0%8\.0%56pci\_decisionClinical Interpretation2592\.0%92\.0%100\.0%8\.0%57hx\_hfBinary Clinical Presence2592\.0%92\.0%100\.0%8\.0%58dc\_med\.96302009\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%59dc\_med\.1191\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%60dc\_med\.100014161\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%61dc\_med\.1546356\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%62dc\_med\.1659152\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%63dc\_med\.613391\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%64dc\_med\.1665684\.dc\_med\_adminMedication/Event Flag2395\.6%95\.6%100\.0%4\.3%65dc\_hospiceBinary Clinical Presence2495\.8%95\.8%100\.0%4\.2%66proc\_med\.1116632\.proc\_med\_adminMedication/Event Flag2596\.0%96\.0%100\.0%4\.0%67proc\_med\.1000142427\.proc\_med\_adminMedication/Event Flag2596\.0%100\.0%96\.0%4\.0%68ecg\_resultsClinical Interpretation2596\.0%100\.0%96\.0%4\.0%69contrast\_volQuantitative Lab/Physiologic2596\.0%100\.0%96\.0%4\.0%70ca\_in\_hospBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%71current\_dialysisBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%72crossoverBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%73dc\_comfortBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%74ipp\_event\.410429000\.post\_proc\_occurredMedication/Event Flag2596\.0%96\.0%100\.0%4\.0%75left\_heart\_cathBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%76prior\_pciBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%77ipp\_event\.100014076\.post\_proc\_occurredMedication/Event Flag2596\.0%96\.0%100\.0%4\.0%78hx\_chronic\_lung\_diseaseBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%79venous\_accessBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%80v\_supportBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%81zip\_codeAdministrative2596\.0%96\.0%100\.0%4\.0%82race\_asianAdministrative2596\.0%100\.0%96\.0%4\.0%83proc\_med\.32968\.proc\_med\_adminMedication/Event Flag2596\.0%96\.0%100\.0%4\.0%84proc\_med\.96382006\.proc\_med\_adminMedication/Event Flag2596\.0%100\.0%96\.0%4\.0%85race\_whiteAdministrative2596\.0%100\.0%96\.0%4\.0%86segment\_idClinical Interpretation2596\.0%96\.0%100\.0%4\.0%87tobacco\_useBinary Clinical Presence2596\.0%100\.0%96\.0%4\.0%88race\_nat\_hawAdministrative2596\.0%100\.0%96\.0%4\.0%89proc\_med\.400610005\.proc\_med\_adminMedication/Event Flag2596\.0%100\.0%96\.0%4\.0%90ipp\_event\.89138009\.post\_proc\_occurredMedication/Event Flag2596\.0%96\.0%100\.0%4\.0%91fluoro\_dose\_dapQuantitative Lab/Physiologic2596\.0%96\.0%100\.0%4\.0%92graft\_stenosisBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%93health\_insAdministrative2596\.0%96\.0%100\.0%4\.0%94diag\_cor\_angioBinary Clinical Presence2596\.0%96\.0%100\.0%4\.0%95first\_nameAdministrative2596\.0%96\.0%100\.0%4\.0%96dominanceClinical Interpretation2596\.0%96\.0%100\.0%4\.0%97fluoro\_dose\_kermQuantitative Lab/Physiologic25100\.0%100\.0%100\.0%0\.0%98pci\_procBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%99last\_nameAdministrative25100\.0%100\.0%100\.0%0\.0%100nv\_stenosisBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%101ipp\_event\.422504002\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%102ipp\_event\.417941003\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%103ipp\_event\.74474003\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%104mid\_nameAdministrative23100\.0%100\.0%100\.0%0\.0%105hypertensionBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%106hipsBinary Clinical Presence24100\.0%100\.0%100\.0%0\.0%107hisp\_origAdministrative25100\.0%100\.0%100\.0%0\.0%108hosp\_interventionBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%109ipp\_event\.230713003\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%110ipp\_event\.230706003\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%111dissection\_segBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%112fluoro\_timeQuantitative Lab/Physiologic25100\.0%100\.0%100\.0%0\.0%113ca\_out\_hospitalBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%114cp\_sx\_assessBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%115ca\_transfer\_facBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%116enrolled\_studyAdministrative25100\.0%100\.0%100\.0%0\.0%117dobAdministrative25100\.0%100\.0%100\.0%0\.0%118dc\_statusBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%119device\_deployedBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%120diabetesBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%121dc\_med\.32968\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%122dc\_med\.1364430\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%123dc\_med\.1537034\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%124dc\_med\.1599538\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%125dc\_locationBinary Clinical Presence24100\.0%100\.0%100\.0%0\.0%126dc\_med\.10594\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%127dc\_med\.1114195\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%128dc\_med\.11289\.dc\_med\_adminMedication/Event Flag23100\.0%100\.0%100\.0%0\.0%129proc\_med\.1537034\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%130proc\_med\.15202\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%131proc\_med\.1364430\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%132proc\_med\.11289\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%133proc\_med\.1114195\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%134prior\_cabgBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%135pre\_proc\_med\.48698004\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%136perf\_segBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%137post\_proc\_timiClinical Interpretation24100\.0%100\.0%100\.0%0\.0%138post\_transfusionBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%139pre\_proc\_med\.100014162\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%140pre\_proc\_med\.100014161\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%141pre\_proc\_med\.112000000694\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%142pre\_proc\_med\.1656341\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%143pre\_proc\_med\.31970009\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%144pre\_proc\_med\.35829\.pre\_proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%145ipp\_event\.95549001\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%146ipp\_event\.35304003\.post\_proc\_occurredMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%147race\_blackAdministrative25100\.0%100\.0%100\.0%0\.0%148proc\_med\.1599538\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%149proc\_med\.321208\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%150proc\_med\.1656052\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%151proc\_med\.373294004\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%152proc\_med\.613391\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%153proc\_med\.1546356\.proc\_med\_adminMedication/Event Flag25100\.0%100\.0%100\.0%0\.0%154race\_am\_indianAdministrative25100\.0%100\.0%100\.0%0\.0%155pt\_restriction2Administrative25100\.0%100\.0%100\.0%0\.0%156sexAdministrative25100\.0%100\.0%100\.0%0\.0%157stress\_performedBinary Clinical Presence25100\.0%100\.0%100\.0%0\.0%

Similar Articles

Evaluating Chinese Ambiguity Understanding in Large Language Models

arXiv cs.CL

This paper introduces CHA-Gen, a Chinese ambiguity dataset grounded in Potential Ambiguity theory, and evaluates several LLMs on ambiguity detection, finding that models struggle but benefit from chain-of-thought prompting, and that instruction tuning induces overconfidence.