Position: Medical AI Neglects Real Treatment Outcomes
Summary
This position paper argues that medical AI currently neglects real treatment outcomes in training and evaluation, hindering its ability to improve patient care, and recommends incorporating actual outcome data to align with evidence-based medicine principles.
View Cached Full Text
Cached at: 08/18/26, 09:51 AM
# Medical AI Neglects Real Treatment Outcomes
Source: [https://arxiv.org/html/2608.14598](https://arxiv.org/html/2608.14598)
###### Abstract
Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions\. But understanding of treatment itself is still inadequately trained and evaluated, using human opinions and syntheses \(especially texts such as biomedical publications and clinical practice guidelines\) rather than actual underlying data on treatment outcomes\. This neglect seriously limits the potential of medical AI, and is already causing deficiencies in both frontier models and major benchmarks, as argued in this position paper\. Real treatment outcomes, drawn from sources such as observational databases and randomized experiments, should be substantially incorporated into both training and evaluation\. Improving these outcomes should be reemphasized as the downstream goal of all medical AI\.
Machine Learning, ICML
## 1Introduction
In the 1990s, the field of medicine addressed two major impediments to quality care\. The first problem was that too many medical decisions relied solely on the opinions of individual experts\. Evidence\-based medicine sought to base these decisions on the collective experience, and accumulated ground\-truth data, of the broader medical system\(Guyattet al\.,[1992](https://arxiv.org/html/2608.14598#bib.bib156); Sackettet al\.,[1996](https://arxiv.org/html/2608.14598#bib.bib155)\)\. The second problem was that medical research, particularly clinical trials, had become myopically focused on surrogate \(or intermediate\) endpoints, such as serum concentrations of various analytes\(Fleming and DeMets,[1996](https://arxiv.org/html/2608.14598#bib.bib138); Temple,[1999](https://arxiv.org/html/2608.14598#bib.bib137); Yudkinet al\.,[2011](https://arxiv.org/html/2608.14598#bib.bib136)\)\. Focusing on more clinically\-relevant treatment outcomes, such as reductions in cardiovascular mortality, was a key thrust of patient\-centered care\.
In our opinion, similar problems currently afflict research in medical AI\. The prevailing tendency is to focus on intermediate predictive problems, typically involving diagnosis or prognosis, without analysis of downstream improvements upon patient outcomes\. Language models do not demonstrate a rich understanding of treatment effects, nor are they trained to develop it; instead, such causal reasoning is outsourced to preexisting works, which are relatively coarse and often mask both uncertainty and heterogeneity\. During both training and evaluation of language models, human opinions are frequently treated as ground truth, even when better alternatives exist\. This methodology, though common in other applications of AI, makes it challenging to achieve \(or even recognize\) performance exceeding that of any human expert — in other words, it contradicts not just the purpose of evidence\-based medicine, but also renewed aspirations for artificial superintelligence\(Morriset al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib141)\)\.
Our position is thatmedical AI currently neglects real treatment outcomes111*Real treatment outcomes*are the actual outcomes of real patients observed in conjunction with their prior treatment\., from training to evaluation, which hinders its purpose of improving these outcomes\. Our recommendation is to use more such data in medical AI research\. This recommendation echoes those from the 1990s about incorporating ground\-truth data into clinical practice\. Our position is not just about realigning effort within medical AI research, but about expanding its long\-term ambitions\. Put simply, the long\-term goal of medical AI should be to help write clinical practice guidelines and other authoritative syntheses, rather than to merely read and reference them\.
[Section˜2](https://arxiv.org/html/2608.14598#S2)highlights the near absence of real treatment outcomes in the training of modern medical AI\.[Section˜3](https://arxiv.org/html/2608.14598#S3)discusses why opinions and human\-written syntheses, such as clinical guidelines and regulatory documents, have fundamental limitations as ground truth, during both inference and evaluation\.[Section˜4](https://arxiv.org/html/2608.14598#S4)shows how established benchmarks focus primarily on ancillary or intermediate predictive tasks\. It emphasizes the fact that treatment is the purpose of the medical system, and improving intermediate predictions doesn’t always serve that purpose\. We present our concrete recommendations in[Section˜5](https://arxiv.org/html/2608.14598#S5)\. Our position cuts against the grain of many research practices in medical AI, and is not without difficulties of its own\.[Section˜6](https://arxiv.org/html/2608.14598#S6)discusses some reasonable alternative stances\.
ModelMed\-Gemini\-\*\(Yanget al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib90)\)Med\-Gemini\-\{L,M\}\(Saabet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib91)\)AMIE\(Tuet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib88)\)Med\-PaLM 2\(Singhalet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib92)\)Med\-PaLM\(Singhalet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib94)\)Baichuan\-M3\(Baichuan\-M3 Team,[2026](https://arxiv.org/html/2608.14598#bib.bib157)\)OpenBioLLM\(Pal and Sankarasubbu,[2024](https://arxiv.org/html/2608.14598#bib.bib85)\)Me\-LLaMA\(Xieet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib77)\)MEDITRON\(Chenet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib82)\)Med42\-v2\(Christopheet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib84)\)MedGemma\(Sellergrenet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib2)\)TxAgent\(Gaoet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib86)\)HealthGPT\-Pro\-8B\(Linet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib158)\)LLaVA\-Med\(Liet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib78)\)BioMistral\(Labraket al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib83)\)MediPhi\(Corbeilet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib1)\)Curiosity\-L\(Waxleret al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib159)\)Mamba\-CLMBR\(Wornowet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib75)\)EHRMamba\(Fallahpouret al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib80)\)MOTOR\-T\(Steinberget al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib79)\)
Size————540B235B70B70B70B70B27B8B8B7B7B3\.8B1B121M130M143MWeb Crawl✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}QA Training Splits✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}Publications✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}Clinical Trials✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}Guidelines✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}Real Patient Data✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}Real Treatment Outcomes\*✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}\*Yanget al\.\([2024](https://arxiv.org/html/2608.14598#bib.bib90)\)use genomic and outcome data, but not treatment data, from the UK BioBank\.
Table 1:The different kinds of data used to train modern language models \(and a few EHR models\) for general medical purposes\. \(SeeWornowet al\.\([2023](https://arxiv.org/html/2608.14598#bib.bib74)\)for similar information on models trained before 2023\)\. A large checkmark✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\boldsymbol\{\\checkmark\}\}denotes explicitly\-mentioned inclusion of the data in the model’s training set; a small one✓\{\\color\[rgb\]\{0,0\.390625,0\}\\definecolor\[named\]\{pgfstrokecolor\}\{rgb\}\{0,0\.390625,0\}\\checkmark\}means implicit inclusion through e\.g\. an extensive web crawl retained in the weights of a large model\. “Patient data” includes imaging, notes, diagnoses, etc\. of real patients, beyond cases presented in medical literature\. “Real treatment outcomes” refer to treatment records along with clinical outcomes \(e\.g\. discharge, symptom resolution, death\) of real patients\. The largest models train on a panoply of data, including real patient data\. However, only the smaller EHR foundation models train on real treatment outcomes\.*Caveat*: the boundaries between some training and inference sets are unclear\.##### An olive branch
This paper unavoidably critiques many recent works in medical AI\. We believe these works made valuable advances, usually in a very short period of time, which we do not undermine\. Rather, this paper charts a course to maintain their exciting pace of progress\.
## 2Doctors Train From Clinical Observation of Treatment; AI Does Not
To practice with a board certification in the United States, a physician must complete a clinical residency which typically lasts as long as their academic medical schooling\. In this setting, they directly observe the consequences of interventions upon patients, and thereby develop their clinical competence\. Observing real treatment outcomes is a crucial aspect of physician training, but it is essentially absent in the training of modern medical AI\.[Table˜1](https://arxiv.org/html/2608.14598#S1.T1)characterizes the training corpora of such models\. Due to space constraints, and the rapid pace of model development, we exclude models trained prior to 2023\. In this table, and throughout the paper, we focus on foundation models \(primarily LLMs\) that are trained for general medical use by a broad audience, excluding models trained for highly specific tasks\.
We observe that most medical AI training is conducted on textual material such as publications and clinical guidelines\. This is unsurprising, since most of these models derive from base language models trained primarily from the web\. Most models are fine\-tuned on the training splits of medical question\-answering datasets\. Many are trained on real patient data, especially clinical notes and radiological images\.Bediet al\.\([2024](https://arxiv.org/html/2608.14598#bib.bib69)\)recently observed that, despite such training, most subsequent evaluations of medical AI are based on synthetic cases rather than real patient data\.
For training, we observe the major watershed is not between real and synthetic data, but between the inclusion or exclusion of real treatment outcomes\. To the best of our knowledge, the only major class of foundation models which include such data are those trained solely from electronic health records\. Such EHR models do not operate on language, but rather sequences of medical events\. Treatment outcomes are inherently longitudinal — for each patient, their treatment must be paired with their subsequent outcome — and such data are the purview of observational medical databases\.
## 3The Use of Opinions and Syntheses as Ground Truth About Treatments
### 3\.1Medication Labels
Case:28\-year\-old female/ Medical history: Non\-vitamin B12 responsive methylmalonic acidemia \(MMA\) / Presenting with:Chronic pancreatitis\(paraduodenal subtype\) / Clinical severity: M\-ANNHEIM Ib*Clinical Presentation*: 9 hospital admissions over one year for acute pancreatitis /No family history of pancreatitis/No alcohol use/ Primary nutrition: Protein\-restricted metabolic formula with canola oil / Normal stool fats*Diagnostic Findings*: Paraduodenal pancreatitis confirmed on T1\-weighted sequence of an MRI pancreas with and without contrast / Met American Pancreatic Association criteria for chronic pancreatitis based on clinical history and CT imaging / Average baseline lipase level: 72 U/L \(reference range 9\-82 U/L\) / Random triglyceride measurements during acute episodesdid not support triglyceride\-induced pancreatitis/ MMA prevented true fasting triglyceride measurement*Genetic Investigation*: Broad workup including genetic causes of recurrent pancreatitis / Identified:Heterozygous R668C \(c\.2002C\>T\) variant of unknown significance in CFTR gene*Declined Testing*: Patient declined sweat chloride test / Patient declined nasal potential difference testing / Reason for declining: Perceived discomfort
Question:Characterize the utility of ivacaftor for resolving this patient’s pancreatitis\.\(A\)Contraindicated for adverse effects\.\(B\)Effective, with high certainty\.\(C\)Effective, with low certainty\.\(D\)Ineffective\.Figure 1:Amislabelingthat occurs when regulatory documents and LLMs are used to \(synthetically\) validate labels for benchmarks, in a manner reminiscent ofPalepuet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib87)\)andGaoet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib86)\)\. This question derives from a real, published case report in which ivacaftor resolved a patient’s pancreatitis\(Tanget al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib109)\)\. The correct answer is \(C\); the supporting evidence for this answer ishighlighted green, and their connection is explained in the main text\. Unfortunately, theincorrect answer \(D\)passes both of the automated correctness checks ofPalepuet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib87)\), which are based on the FDA product label for ivacaftor\. The fundamental problem is that regulatory documents, such as drug labels, are not intended to provide comprehensive ground truth for treatment outcomes\. Note that theRxQAdataset is additionally verified by pharmacists, and for clarity, whereasTreatmentPCis not\.Medication labels, such as FDA product labels, are government\-sanctioned communications of approved uses, risks, and instructions\. They form the basis for lawful product marketing\. These labels typically range from 15 to 100 pages, compile the most crucial information about the medication in question, are free to access, and are published in the public domain\. As such, they are convenient source material for NLP\(Liet al\.,[2013](https://arxiv.org/html/2608.14598#bib.bib124)\)and model training\(ValizadehAslaniet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib125)\)\. Recently,Palepuet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib87)\)andGaoet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib86)\)used medication labels to produce question\-answer benchmarks namedRxQAandTreatmentPC, respectively\. The latter’s methodology underpinned the CURE\-Bench Challenge at NeurIPS 2025\. In these datasets, each question presents a patient case study along with multiple\-choice answers\. They concern different aspects of treatment planning\. The questions are automatically generated and validated in a way that is based primarily on medication labels\. \(RxQais additionally verified by pharmacists, butTreatmentPCis not\)\.
The purpose of medication labels is to legally constrain the marketing activities of manufacturers\. Multiple court opinions have determined that medication labels do not represent a standard of care\(Beck,[2017](https://arxiv.org/html/2608.14598#bib.bib130)\)\. Importantly, they do not discuss*off\-label use*, which is a fundamental aspect of medical practice\. In some fields, off\-label use constitutes 30\-80% of all medication use\(Allenet al\.,[2018](https://arxiv.org/html/2608.14598#bib.bib127)\)\.
To demonstrate the perils of treating medication labels as ground truth, we demonstrate a bypass of both of the automated correctness verification measures used inRxQA\.222RxQAhas an additional clarity check which we don’t involve, and arguably promotes simplistic questions\.The prompts for these verifications are displayed in[Figure˜4](https://arxiv.org/html/2608.14598#A1.F4),[Section˜A\.2\.1](https://arxiv.org/html/2608.14598#A1.SS2.SSS1)\. In the first, the language model is given the medication label and the generated example, and is asked whether the candidate correct answer is, in fact, correct\. The second verification presents the same inputs, but asks if the other answers are, in fact, incorrect\. We present an example, shown in[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1), where the incorrect answer \(D\) passes both verifications\.[Figures˜5](https://arxiv.org/html/2608.14598#A1.F5)and[6](https://arxiv.org/html/2608.14598#A1.F6)show Gemini 1\.5 Flash, which was used byPalepuet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib87)\), performing these erroneous verifications\. \(Our results also hold for Gemini 3\.5 Flash and Gemini 3\.1 Pro Preview, so we will simply refer to the model as Gemini\.\)
##### Background
Our example in[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1)is based on the real, published case report ofTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\), who encountered a patient suffering from idiopathic pancreatitis\. After finding a heterozygous variant in the patient’s CFTR gene, they administered ivacaftor, which completely resolved the patient’s pancreatitis\. Ivacaftor belongs to a class of drugs called CFTR modulators which target defective CFTR proteins\. The motivation for the treament was an emerging \(low\-certainty\) body of evidence which ties a large fraction of idiopathic pancreatitis \(perhaps 30\-40%\) to CFTR gene alterations, or even environmental factors affecting CFTR expression\(Cohnet al\.,[1998](https://arxiv.org/html/2608.14598#bib.bib111); Audrézetet al\.,[2002](https://arxiv.org/html/2608.14598#bib.bib110); Phadke and Sellers,[2022](https://arxiv.org/html/2608.14598#bib.bib108); Hegyiet al\.,[2016](https://arxiv.org/html/2608.14598#bib.bib107)\)\. Such irregularities and dysfunction may not be severe enough to earn a full diagnosis of cystic fibrosis — that is, a threshold value of≥60\\geq 60mmol/L on a sweat chloride test — which would then classify the use of ivacaftor as on\-label\.
##### Ablations
The immediate counterargument to our experiment is: a case report is just a single datum, and the supporting evidence is weak\. Perhaps Gemini is simply being logical and evidence\-based in its continued refutation of ivacaftor’s potential effectiveness\. To examine this possibility, we repeated the experiment, but additionally attached the case report\. Per[Figure˜8](https://arxiv.org/html/2608.14598#A1.F8), Gemini finds the case report convincing enough to correctly answer \(C\)\. Is it possible that the mere inclusion of a case report, generally discussing ivacaftor’s effectiveness, was sufficient to overwhelm Gemini’s clinical reasoning? To check this, we synthesized a patient who would almost certainly not benefit from ivacaftor, as their pancreatitis seems caused by alcohol use \([Figure˜7](https://arxiv.org/html/2608.14598#A1.F7)\)\. After attaching the case report and asking about this synthetic patient \([Figure˜9](https://arxiv.org/html/2608.14598#A1.F9)\), Gemini retains its clinical reasoning capabilities and correctly answers \(D\)\. The problem lies not in Gemini’s abstract reasoning capabilities, but in the text it was told to treat as ground truth\.
### 3\.2Clinical Practice Guidelines
An 87\-year\-old female with a documented medical history of atrial fibrillation, chronic heart failure with preserved ejection fraction \(HFpEF\), hypertension, hypokalemia, vitamin D deficiency, and glaucoma is taking the following medications:Aspirin 81 mg, once daily / Diltiazem 240 mg extended\-release, once daily / Furosemide 40 mg, once daily / Metoprolol tartrate 50 mg, twice daily / Potassium chloride 20 mEq, once dailyThe patient is experiencing moderate pedal edema and has an unstable gait\. Medication non\-adherence is an ongoing issue\. Additionally, a recent echocardiogram indicates that the patient’s Left Ventricular Ejection Fraction \(LVEF\) has worsened to less than 40%\.What are the next steps of treatment?
↓\\downarrow •Diltiazem:consider sacubitril\-valsartanfor further “optimization of guideline\-directed medical therapy\.”•Furosemide:increase dose to address volume overload and pedal edema\.•Metoprolol tartrate:consider succinatefor hypertension management \-no consideration for pill burden\.•*Added*: SGLT2 inhibitors and spironolactone\.•Diltiazem: discontinued\. Replaced by losartam due to HFrEF, edema\.•Furosemide: reduce dose because of cascade from diltiazem to edema\.•Metoprolol tartrate: switch to succinate to reduce pill burden \(once daily\)\.•*No additions*
Figure 2:An example of how formulaically grounding answers in guidelines can harm clinical reasoning\. The question is based on the published case report ofHaet al\.\([2021](https://arxiv.org/html/2608.14598#bib.bib126)\), which involved a prescribing cascade: the patient’s potassium deficiency \(hypokalemia\) was exacerbated by the furosemide, which was being given to address edema, a side effect of diltiazem\. The physician’s treatment plan is on the right in light green\. It recognized the cascade and successfully treated both the patient’s heart condition and edema\. It also reduced pill burden in light of the patient’s nonadherence\. On the left is a summary of OpenEvidence’s treatment plan\. Its full response is available at[this anonymized link](https://web.archive.org/web/20250517125001/https://www.openevidence.com/ask/e5955b7e-7cfe-4864-a08c-217c60eda8cc)and also in[Figure˜12](https://arxiv.org/html/2608.14598#A1.F12),[Section˜A\.3](https://arxiv.org/html/2608.14598#A1.SS3)\. It explicitly describes its answer as “guideline\-directed\.” Indeed, it maintains the focus of the guidelines, correctlyaddressing HFrEF\. However, it ignores the patient\-specific considerations inyellow\. Furthermore, it doesn’t reason about the prescribing cascade: it elects toincrease, not decrease, furosemide dosage\.Clinical practice guidelines are extensive consensus statements published by professional societies, governmental health bodies, and international organizations\. Within a specific area of practice, they discuss many aspects of care, including diagnostic criteria and risk stratification\. They usually include decision trees for recommended treatment pathways\. They typically take years to write and are sometimes hundreds of pages long\. As seen in[Table˜1](https://arxiv.org/html/2608.14598#S1.T1), guidelines feature in the training corpora of many medical language models\. At inference time, some models are designed to specifically, or even exclusively, cite guidelines\. For example, the Mx Agent of AMIE generates treatment plans by \(1\) searching solely for clinical practice guidelines, and \(2\) citing those guidelines via in\-context retrieval\(Palepuet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib87)\)\. Services such as OpenEvidence and ChatGPT for Healthcare are designed to cite authoritative documents such as guidelines\. AMEGA is a recent, rubric\-based benchmark which measures adherence to guidelines\(Fastet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib129)\)\.
The purpose of clinical practice guidelines is to establish generic professional standards\. However, they are not binding either professionally or legally, even when determining a minimum standard of care\. The American Law Institute recently clarified the legal role of guidelines in their restatement of torts, which revised their widely held legal standard for medical negligence\(Aaronet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib131); Peters Jr,[2023](https://arxiv.org/html/2608.14598#bib.bib132)\)\. Although adhering to guidelines may exculpate a physician from malpractice claims, deviating from them does*not*, in of itself, indicate malpractice\. This asymmetry means grounding an answer in guidelines can lead to reasoning that is merely defensible rather than accurate\.
Guidelines do not describe optimal or personalized treatment\. When basing answers primarily \(or solely\) upon guidelines, the following deficiencies can arise:
- •Oversimplification: guidelines are written for, and by, people\. This makes their treatment plans highly constrained in terms of decision\-tree complexity\. They do not gracefully handle stochasticity, express backtracking, or address a long tail of nonstandard circumstances\.
- •Staleness: guideline creation is a formal process which convenes many experts, so these documents are updated very infrequently\. Real clinicians are expected to routinely resolve conflicts among guidelines, new evidence, and their own experience\.
We present two concrete examples of these deficiencies\. The first example is presented in[Figure˜2](https://arxiv.org/html/2608.14598#S3.F2)\. It is based on the real, published case report ofHaet al\.\([2021](https://arxiv.org/html/2608.14598#bib.bib126)\), in which a physician successfully treated a cardiac patient by reducing their pill burden\. This case illustrates the phenomenon of*prescribing cascade*, in which symptoms are mistakenly thought to arise from a disease process, when they are actually the side effects of a previously administered drug\(Rochon and Gurwitz,[1997](https://arxiv.org/html/2608.14598#bib.bib103); Brathet al\.,[2018](https://arxiv.org/html/2608.14598#bib.bib102)\)\. Guidelines do not typically provide alternate recommendations for nonadherent patients\. Guidelines also don’t address prescribing cascade, since it is a path\-dependent phenomenon whose resolution involves backtracking\. As a consequence, the explicitly guideline\-directed answer provided by OpenEvidence did not recognize the prescribing cascade: it increased, not decreased, the dose of furosemide\. Furthermore, it did not attempt to reduce the patient’s pill burden\. As seen in[Figures˜13](https://arxiv.org/html/2608.14598#A1.F13)and[14](https://arxiv.org/html/2608.14598#A1.F14)of the Appendix, ChatGPT for Clinicians and ChatGPT for Healthcare similarly failed to detect the prescribing cascade\. They did not recommend reducing furosemide, although they fared slightly better by at least recognizing the importance of medication adherence\.
[Figure˜15](https://arxiv.org/html/2608.14598#A1.F15), in[Section˜A\.3](https://arxiv.org/html/2608.14598#A1.SS3), presents the second example, as well as the clinical reasoning and treatment plan provided by OpenEvidence\. This example is based on a published report of how OpenEvidence \(retrospectively\) planned treatment for a real patient at the Mayo Clinic\(Hurtet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib128)\)\. The focus of our scrutiny is the reasoning and use of evidence, rather than the correctness of the final treatment plan\. The answer ignores important information — in particular, precision lab measurements — in deference to older 2018 guidelines, which focus on a more widely\-available proxy measure\. Furthermore, the answer does not incorporate information from other \(newer\) guidelines and sources, or even hint that those sources may present different reasoning about lab measurements\.
### 3\.3Clinician Opinion \(And Aggregates Thereof\)
The limitations of individual human judgment have been studied extensively in psychology\(Meehl,[1954](https://arxiv.org/html/2608.14598#bib.bib57); Tversky and Kahneman,[1974](https://arxiv.org/html/2608.14598#bib.bib63)\)\. These motivated the use of evidence\-based methodology in medicine\(Cochrane,[1972](https://arxiv.org/html/2608.14598#bib.bib64)\), where the fundamental problem of causal inference makes them especially acute\. Since clinicians don’t see the counterfactual outcomes of patients they treat, their ability to independently accumulate expertise about treatment is limited\.
Meanwhile, clinician opinion is routinely used as ground truth in medical AI benchmarks\. Clinicians are queried “how does the answer relate to the consensus in the scientific and clinical community?”\(Singhalet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib94)\)or “which answer better reflects the current consensus of the scientific and clinical community?”\(Singhalet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib92)\)\. This clinician involvement is not necessarily bad\. Medical AI needs to be judged for qualities such as empathy, adherence to social norms, and ethical alignment — topics in which people are the ultimate arbiters\.
However, measuring agreement with groups of clinicians is not a general replacement for ground truth validation\. Agreement has no strict relationship with the truth\. A majority consensus usually has lower average error than the average individual clinician \(by convexity\)\. But consensus can be expensive to establish, and agreement may not actually reflect consensus\. This dilemma echoes related concerns about annotation in natural language processing\(Plank,[2022](https://arxiv.org/html/2608.14598#bib.bib25)\)\. A disappointing aspect of agreement as an evaluation metric is that it halts our ambitions to understand what humans do not already know\.
These issues come to the fore in HealthBench\(Aroraet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib72)\)\. In this benchmark, language models are presented with a health\-related question or conversation\. Their answers are judged according to criteria specified by clinicians\. For the present discussion, the most relevant criterion is “this answer contains no factually incorrect information”, but there are many different axes of evaluation\. At evaluation time, GPT 4\.1 grades whether an answer meets the prespecified criteria\. HealthBench is notable because it offers a meta\-evaluation of how well GPT 4\.1 performs as a grader compared to clinicians\. In this evaluation, clinicians provide 60,896 true/false labelings of whether an answer to a question meets a specific criterion\. In this crucial aspect of rubric benchmarking,Aroraet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib72)\)conclude that GPT 4\.1 achieves “clinician\-level agreement\.”
Let us examine what this means\. The dataset involves no ground truth outside of clinician opinion\. Because of the sparsity of the labeling, clinician consensus is very tenuous, or does not exist, for the vast majority of the examples\. Concretely, 55,546 \(91\.2%\) of the labels originate from 27,773 examples which were labeled by exactly 2 clinicians\. \(Whether the criterion is relevant to the question*is*previously determined by a consensus, but whether any given answer satisfies it is not\)\. Thus, agreement cannot be reasonably assessed on a per\-example basis\. Instead, clinicians \(and GPT 4\.1\) are scored on how much they agree with \(other\) clinicians across all examples\.
As a consequence of this design, it is possible for GPT 4\.1 to have high agreement and low accuracy at the same time\. In particular, let us examine the 5,105 HealthBench meta\-examples about whether an answer was factually accurate\. “Clinician\-level agreement” still holds on this subset\. All 5,105 such examples were labeled by GPT 4\.1 and just 2 clinicians\. Their labels group as follows\.
GPT 4\.1CliniciansCountYesBoth Yes3175Disagree1050Both No246NoBoth Yes221Disagree208Both No205Consider the following possible ground\-truth labeling: when GPT 4\.1 and the clinicians all agree \(in3175\+205=33803175\+205=3380examples, i\.e\. 66\.2%\) then set the label to match everyone\. In the remaining examples, set the label to the opposite of GPT’s\. This would make the clinician answers 87\.6% accurate compared to GPT’s 66\.2%\. By the opposite construction, it is also possible that GPT 4\.1 is 100% accurate and clinician answers are only 78\.5% accurate\. Thus, in this context, agreement is just a marginal guarantee of conformity unmoored from the truth\.
#### 3\.3\.1What is Elicited From Clinicians?
The theoretical hope of querying clinical experts is that they are the ultimate aggregation of all medical knowledge: they have both hands\-on experience and knowledge of guidelines and other authoritative documents\. In principle, asking clinicians to label examples should give higher quality than any of these underlying sources\. However, in reality, it may be the case that, when faced with an abstract grading task, experts merely defer to such documents rather than conveying the full breadth of their expertise\.
To investigate this, we examine HealthBench Professional\(Hickset al\.,[2026](https://arxiv.org/html/2608.14598#bib.bib160)\)\. Like HealthBench, this benchmark is based completely on clinician input, with no incorporation of ground\-truth clinical data\. There are 525 examples of conversations \(initiated by medical professionals\) and sample answers\. Across these examples, there are 1135 total clinician\-established criteria for grading answers\. In our analysis \(detailed in[Section˜A\.4](https://arxiv.org/html/2608.14598#A1.SS4)\) we find that, among the criteria relating to clinical facts or decisions, approximately 86% are directly grounded in a passage from a guideline or another such authoritative document\.
This is substantially higher than clinician guideline concordance in practice\. Large\-scale studies, averaging across different specialties, generally estimate real\-world guideline concordance to hover around only 55% to 60%\(McGlynnet al\.,[2003](https://arxiv.org/html/2608.14598#bib.bib167); Runcimanet al\.,[2012](https://arxiv.org/html/2608.14598#bib.bib168)\)\. The reasons for guideline discordance vary: some of this is likely due to outdated training or lack of knowledge\. However, a recent study\(Finkelsteinet al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib162)\)found that physicians may be less guideline\-concordant when treating their own family members, suggesting that they may take a more nuanced, personalized approach in sensitive circumstances\. Regardless of the underlying reason, this statistical discrepancy indicates that clinician benchmark input and actual clinician behavior may differ\.
## 4Treatment Is the Focus of Medicine, But Not of AI Evaluations
The purpose of the medical system is to improve outcomes through treatment\. This basic intuition is reflected in both the law and health economics\. Diagnosis and ancillary tasks may be carried out by different kinds of professionals, but the prescriptive authority to initiate significant treatment is reserved for licensed physicians\. Since the passage of the Affordable Care Act \(2010\) and MACRA \(2015\), value\-based care has taken root across US healthcare policy, increasingly tying provider reimbursement to treatment outcomes\. Despite the centrality of treatment, most established AI evaluations focus on intermediate predictive tasks\. We now discuss the extent to which they dominate benchmarks, and why it is important to elucidate their relationship with improved treatment outcomes\.
### 4\.1Prevalent AI Benchmarks Don’t Assess Understanding of Treatment Outcomes
Until recently, the primary benchmarks in medical AI were collections of multiple\-choice questions\. Most of these focused on general biomedical knowledge, such as PubMedQA\(Jinet al\.,[2019](https://arxiv.org/html/2608.14598#bib.bib66)\)and MMLU clinical topics\(Hendryckset al\.,[2021](https://arxiv.org/html/2608.14598#bib.bib65)\)\. Understanding of treatment outcomes was measured by performance on datasets of medical exam \(e\.g\. USMLE\) questions\. However, these initial datasets, such as MedQA\(Jinet al\.,[2021](https://arxiv.org/html/2608.14598#bib.bib62)\)and MedMCQA\(Palet al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib61)\), have now become saturated\(Saabet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib91)\)\. Furthermore, their limitations in measuring actual clinical understanding are recognized\(Rajiet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib135); Alaaet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib134)\)\. Newer QA datasets require more difficult free\-form answers — notably, ones derived from JAMA Clinical Challenge\(Chenet al\.,[2025a](https://arxiv.org/html/2608.14598#bib.bib133); Chang and Fontanarosa,[2011](https://arxiv.org/html/2608.14598#bib.bib58)\)\. Newer benchmarking efforts evaluate LLMs more comprehensively along more clinically important axes\. Perhaps the most significant effort in this direction is MedHELM\(Bediet al\.,[2026](https://arxiv.org/html/2608.14598#bib.bib67)\), which aggregates 37 datasets in view of 121 clinically important tasks\.
Though these new benchmarks offer many substantial improvements, they \(still\) neglect to meaningfully evaluate understanding of treatment outcomes\. By this, we refer to questions which are based on realistic cases, and ask for a free\-form treatment recommendation \(i\.e\. “what are the next steps of treatment?”\) or more quantitative estimates akin to regression \(“what is the effect of this treatment?”\), ranking \(“what is the best treatment?”\) or classification \(“is this treatment effective?”\)\.
Relatively few entries in JAMA Clinical Challenge fit into these categories\. This repository consists of 1700\+ questions, of which most \(1524\) are in the dataset ofChenet al\.\([2025a](https://arxiv.org/html/2608.14598#bib.bib133)\)\. By our calculations, of the 1277 questions with public information, 711 \(55\.6%\) ask for a diagnosis, 36 \(2\.8%\) are about unknown or assorted topics, and 530 \(41\.5%\) ask for next steps of treatment\. Of the “next step” questions, approximately 232 \(18\.1% of the verifiable questions\) actually involve some assessment or modification of treatment; the others are answered just by requests for further information via imaging, biopsies, and so forth\.
Of the 31 datasets in MedHELM, only two nominally pertain to treatment planning: MTSamples and MEDEC\(Ben Abachaet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib59)\), which are both clinical note datasets\. In MTSamples, the model is presented with the initial part of a note, and is prompted to generate a treatment plan\. Its output is then compared to a later part of the note, often a recap of a surgical operation\. Recording patient interactions is only tangentially related to providing ground truth about treatment outcomes\. In MEDEC, the goal is to take a note and identify which, if any, of the sentences contain an error\. As seen in Figure 3 ofBen Abachaet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib59)\), only about 10% of these errors relate to treatment\. Most of the dataset is generated by using MedQA as a source of ground truth for scenarios\. Retrospective correction of errors — of which some are transcription flaws, and others are iatrogenesis — is not the same task as prospective treatment planning\.
Aside from these conceptually\-oriented benchmarks, there are numerous clinical evaluations of LLMs in the medical literature\. These were the subject of a recent systematic review byChenet al\.\([2026](https://arxiv.org/html/2608.14598#bib.bib161)\), which found789789out of46094609clinical evaluations of LLMs \(17\.1%17\.1\\%\) nominally involve treatment planning or recommendation\. We further scrutinize these789789studies, checking \(1\) if they actually focused on some aspect of treatment itself, rather than diagnosis, prognosis, or other intermediate tasks, \(2\) whether they could be reused as benchmarks, as opposed to being one\-time manual assessments of answers, and \(3\) whether real treatment outcomes were used as ground truth, as opposed to guidelines or subjective expert opinions\. \(Full details are in[Section˜A\.5](https://arxiv.org/html/2608.14598#A1.SS5)\.\) We find that just180/4609≈3\.9%180/4609\\approx 3\.9\\%of the studies are potentially\-reusable benchmarks focused on treatment\. Furthermore, the vast majority of ground truth is unmoored from patient outcomes, with702/789≈88\.9%702/789\\approx 88\.9\\%of the subset relying on guidelines or human opinion\. There are only16/460916/4609reusable benchmarks about treatment which involve real outcomes\. To summarize, even nominally treatment\-related evaluations do not assess a causal understanding of treatment, and the use of textual syntheses or opinions as ground truth is pervasive\.


Figure 3:The convoluted relationships among different kinds of medical data — and two different ways that AI can approach such data\. On theleftis the status quo;yellowhighlights data used for training, andtealhighlights data used for evaluation\. As seen in[Table˜1](https://arxiv.org/html/2608.14598#S1.T1), treatment outcomes from observational databases are essentially ignored: the partial shading denotes limited usage\. Training is based primarily on human opinions and syntheses\. Evaluation is performed on formally separate splits of such data\. However, due to the close interconnections among such data sources, leakage is a serious concern\. On therightis the cleaner system advocated by this paper\. It treats the source nodes in this graph as ground truth\. Training is based primarily on observational databases\. Randomized trials are used to evaluate understanding of treatment effects\. Approaches like full conformal prediction allow learning from trials during uncertainty quantification\(Angelopouloset al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib24)\)\. Human opinions are used to evaluate intermediate predictions, alignment, etc\. Overall, this approach reduces leakage\. The ensuing models can be used to inform research, regulations and professional standards\.
### 4\.2Therapeutic Development is Not Patient\-Centered Treatment Understanding
Multiple benchmarks assess how well models predict clinical trial outcomes\. These includeTOP\(Fuet al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib8)\),TrialBench\(Chenet al\.,[2025b](https://arxiv.org/html/2608.14598#bib.bib10)\)andCTO\(Gaoet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib9)\)\. Clinical trials offer crucial ground\-truth data on treatment outcomes\. However, these benchmarks are oriented towards reducing the cost of therapeutic development, and only partially align with the goal of understanding treatment effects\. They lack a focus on comparative effectiveness — determining when one treatment is superior to another — which is one of our main concerns\.
In therapeutics research, a clinical trial is deemed successful if it meets statistical endpoints and proceeds to the next phase\. In Phases 1 and 2, the decision to advance is made by the sponsor, frequently driven by portfolio strategy rather than clinical utility alone\. By Phase 3, success is defined by regulatory approval, but the FDA is statutorily barred from considering cost and generally does not mandate evidence of superiority to existing treatments\(Gottlieb,[2011](https://arxiv.org/html/2608.14598#bib.bib7)\)\.
As such, commercial and operational factors can influence the outcome definitions in these benchmarks\. For example,TrialBenchmarks negative outcomes for insufficient patient enrollment\. Also,CTOinfers missing outcomes from stock prices and similarly marks trials terminated for administrative reasons as failures\.
### 4\.3Connecting Intermediate Predictions to Outcomes
In medicine, automating some kinds of intermediate and ancillary tasks \(such as ambient clinical documentation, automated billing, or triage scheduling\) demonstrably reduces administrative burden and provider burnout\(Afshar and others,[2025](https://arxiv.org/html/2608.14598#bib.bib164)\)\.However, the connection between intermediate predictive tasks and patient outcomes is not always clear\.
Perhaps the most well\-known examples of such surprises arise in preventive screening\. Breast cancer \(mammography\), prostate cancer \(prostate\-specific antigen\), and lung cancer \(low dose CT\) screening tests have all been found to be less worthwhile than anticipated, following the completion of randomized trials\(Milleret al\.,[2014](https://arxiv.org/html/2608.14598#bib.bib106); Andrioleet al\.,[2009](https://arxiv.org/html/2608.14598#bib.bib105); Marcuset al\.,[2000](https://arxiv.org/html/2608.14598#bib.bib104)\)\. Screening, early diagnosis, and risk stratification are counterproductive when they find problems to medicalize rather than genuine opportunities to intervene and improve outcomes\. For example, in colonoscopy, using computer vision to find the maximal number of polyps or lesions isn’t necessarily ideal\. Improving outcomes requires finding polyps which are not so slow\-growing to be harmless within the course of a lifetime, but also not so aggressive that they will kill the patient regardless\(Hassanet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib55)\)\. Harm can occur by subsequent risks in the medical system \(i\.e\. overdiagnosis\(Welch and Black,[2010](https://arxiv.org/html/2608.14598#bib.bib54)\)\) or the data acquisition itself \(for example, perforation in colonoscopy\)\. In these scenarios, optimizing for raw diagnostic sensitivity rather than patient\-centered benefit, without accounting for potential side effects or system\-level constraints, decouples the intermediate prediction from the ultimate outcome\.
These difficulties stem from fundamental misalignments between the predictive task and downstream outcomes\. But even when those are conceptually aligned, more subtle problems can arise from the way model performance is quantified\. Most benchmarks report some kind of average accuracy or win rate\(Lianget al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib36)\)\. However, the key property which enables predictions to be used safely for downstream decisions is calibration, not accuracy\(Noarov and Roth,[2024](https://arxiv.org/html/2608.14598#bib.bib99)\)\. Unfortunately, calibration is challenging to both achieve and verify\. Recent works have examined less stringent conditions which focus on specific kinds of downstream decisions\(Zhaoet al\.,[2021](https://arxiv.org/html/2608.14598#bib.bib35); Rothblum and Yona,[2023](https://arxiv.org/html/2608.14598#bib.bib101); Noarovet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib100)\)\.
## 5Recommendations \(Call to Action\)
Trainingshould be based more on longitudinal patient records \(e\.g\. electronic health records and insurance claims\), and less on textual syntheses such as clinical practice guidelines, medical publications and related literature\. Longitudinal data pipelines are common in the real\-world evidence \(RWE\) industry, which rapidly extracts causal insights from large\-scale databases\(Jensenet al\.,[2012](https://arxiv.org/html/2608.14598#bib.bib166); Shermanet al\.,[2016](https://arxiv.org/html/2608.14598#bib.bib165)\)\. Large extracts of such datasets are now being released, at no cost and without excessive encumbrance, to the research community\. A major example is the CRITICAL dataset\(Luoet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib11)\)\. This dataset involves approximately 400,000 critical care patients, including longitudinal inpatient and outpatient data from before, during, and after their ICU admission\. Software for processing this dataset into standardized machine learning tasks is available\(Luo and Li,[2025](https://arxiv.org/html/2608.14598#bib.bib12)\)\.
Evaluationshould be based more on randomized experiments, where applicable, and less on regulatory documents, clinical practice guidelines, and clinician opinion\. By this, we mean predicting the results of previously conducted experiments; having AI participate in new experiments may present ethical and logistical challenges\. New benchmarks can be formulated as target trial emulations\(Hernán and Robins,[2016](https://arxiv.org/html/2608.14598#bib.bib53); Hernánet al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib52)\)\. When creating these, pipelines from trial outcome benchmarks can be reused\(Chenet al\.,[2025b](https://arxiv.org/html/2608.14598#bib.bib10)\)\. When experimental data are not available, held\-out observational data can be used to generate target causal estimates\. Ground\-truth quantitative outcome data, without additional interpretation, can be extracted from scientific publications\. \(However, care should be taken to avoid the leakage issue illustrated in[Figure˜3](https://arxiv.org/html/2608.14598#S4.F3)\)\.
Intermediate predictive tasksshould, of course, still be vigorously pursued\. Their contribution to improving treatment outcomes should be quantified, at least roughly\. Ideally, this could occur by analyzing deployments of predictive models as causal interventions\(Joshiet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib20)\)while reporting outcome\-oriented metrics\(Tierneyet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib19)\)\. At earlier stages, translating performance improvements into more patient\-oriented metrics, as done in health economics, would bring a mature emphasis on real\-world priorities\(Rossiet al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib30); van Leeuwenet al\.,[2021](https://arxiv.org/html/2608.14598#bib.bib29)\)\.
For example, consider ambient AI summarization of patient conversations\. During development, such models are benchmarked on NLP\-style metrics such as Word Error Rate \(WER\) or Mean Text Recurrence \(MTR\), which are appropriate for algorithmic research\. A complementary layer of evaluation connects these metrics to clinician performance\(Moramarcoet al\.,[2022](https://arxiv.org/html/2608.14598#bib.bib163)\)and, in turn, clinical experience and practitioner well\-being\(Afshar and others,[2025](https://arxiv.org/html/2608.14598#bib.bib164)\)\. This burden does not fall on algorithmic developers, but on researchers focusing on clinical evaluation and deployment\.
## 6Alternative Views
*“Training AI on longitudinal records is a violation of patient privacy\.”*It should be noted that medical AI models already train upon large quantities of sensitive, individual\-level patient data \(see[Table˜1](https://arxiv.org/html/2608.14598#S1.T1)\), while respecting prevailing standards in biomedicine\. Treatment outcomes are not inherently more sensitive than such data\. Indeed, treatment effects are most naturally inferred at a group level\. Thus, techniques such as differential privacy\(Dwork,[2006](https://arxiv.org/html/2608.14598#bib.bib6)\)can be used to ethically learn about treatment effects\.
*“Training pipelines for text are more mature than those for medical records\.”*Indeed, observational medical databases can be fragmented and difficult to access\. By law, electronic health records must use standard formats\(Bender and Sartipi,[2013](https://arxiv.org/html/2608.14598#bib.bib45)\)and vocabularies, and must be accessible via API\(Phelanet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib44)\)\. However, these provisions ensure only syntactic, not semantic, interoperability\(Dolin and Alschuler,[2011](https://arxiv.org/html/2608.14598#bib.bib43)\)\. We concede that most observational medical databases are not ready for large\-scale AI training\. Recent developments in AI depended on major concurrent investments in training data pipelines\(Brownet al\.,[2020](https://arxiv.org/html/2608.14598#bib.bib48); Gaoet al\.,[2020](https://arxiv.org/html/2608.14598#bib.bib47); Ratneret al\.,[2017](https://arxiv.org/html/2608.14598#bib.bib46)\)\. Medical AI seems to need the same\(Khurshid and Sarkar,[2025](https://arxiv.org/html/2608.14598#bib.bib39)\)\.
*“Publications and other syntheses comprise the majority of readily\-available medical training data\.”*Our recommendation is not extreme and absolute\. We recommend a gradual shift to real treatment outcomes; we acknowledge this may need to occur over the course of years\.
*“Even when they are available, randomized controlled trials have flaws as ground truth\.”*Low power, blinding violations, and dropouts can harm their internal validity\(Deaton and Cartwright,[2018](https://arxiv.org/html/2608.14598#bib.bib32)\)\. Highly\-artificial inclusion criteria can limit their external validity\(Rothwell,[2005](https://arxiv.org/html/2608.14598#bib.bib38)\)\. Fortunately, these design aspects of the trial are features that can be used to predict its outcome\. Thus, if a trial’s design biases its outcome, then that bias can be disentangled from the true underlying effect\(Kaul and Gordon,[2025](https://arxiv.org/html/2608.14598#bib.bib27)\)\. Another issue is that RCTs elucidate average, rather than individual, treatment effects\. Individual patient data from RCTs can be used to make inferences about subgroups or \(under stronger assumptions\) individuals themselves\(Pearl,[2015](https://arxiv.org/html/2608.14598#bib.bib26)\)\. Finally, RCTs are not the only kind of experiments which can serve as ground truth\(Guyattet al\.,[1986](https://arxiv.org/html/2608.14598#bib.bib34); Liang and Recht,[2025](https://arxiv.org/html/2608.14598#bib.bib33)\)\.
*“Conformance with guidelines and clinician consensus is a practical requirement for any deployable medical AI\.”*Our recommendation is to improve this status quo\. Even leaving AI aside, medical evidence is rife with uncertainty and controversy\(Shaneyfeltet al\.,[1999](https://arxiv.org/html/2608.14598#bib.bib5); Prasadet al\.,[2013](https://arxiv.org/html/2608.14598#bib.bib4); Kent and Ioannidis,[2026](https://arxiv.org/html/2608.14598#bib.bib3)\)\. Ideally, AI should help resolve these challenges, not just propagate them\.
*“Focusing primarily on treatment outcomes risks devaluing intermediate tasks, which represent the most successful and practical applications of medical AI to date\.”*Placing emphasis on treatment outcomes does not necessarily diminish the value of intermediate tasks\. This dynamic is familiar to the medical community: the historical shift toward evidence\-based, patient\-centered outcomes\(Guyattet al\.,[1992](https://arxiv.org/html/2608.14598#bib.bib156); Sackettet al\.,[1996](https://arxiv.org/html/2608.14598#bib.bib155)\)did not imperil the in\-silico or in\-vitro biomedical investigation of molecules and analytes\. In the same way, training medical AI to better understand treatment will likely asssist in solving existing predictive tasks\. For example, in ambient AI scribing, anticipating likely treatments helps a clinician ask for the vital patient information needed to complete a high\-fidelity clinical note\. This, in turn, helps downstream providers manage treatment more effectively\. Viewing these tasks as mutually reinforcing and servicing the same ultimate goal enables collaboration, not antagonism, among different research agendas\.
## 7Conclusion
Making predictions by using published text as source material, and humans as ground truth, is a pattern adopted from other AI\. It embodies neither the customs nor the ambitions of medicine\. The best long\-term plan for medical AI is to reemphasize what truly matters: real treatment outcomes\.
##### Acknowledgements
The authors thank Richard Platt and Geoffrey J\. Gordon for their valuable comments and feedback\.
##### Reproducibility
## References
- D\. G\. Aaron, C\. T\. Robertson, L\. P\. King, and W\. M\. Sage \(2025\)A new legal standard for medical malpractice\.JAMA333\(13\),pp\. 1161–1165\.Cited by:[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p2.1)\.
- M\. Afsharet al\.\(2025\)A pragmatic randomized controlled trial of ambient artificial intelligence to improve health practitioner well\-being\.NEJM AI2\(12\),pp\. AIoa2500945\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p1.1),[§5](https://arxiv.org/html/2608.14598#S5.p4.1)\.
- A\. Alaa, T\. Hartvigsen, N\. Golchini, S\. Dutta, F\. Dean, I\. D\. Raji, and T\. Zack \(2025\)Medical large language model benchmarks should prioritize construct validity\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- H\. C\. Allen, M\. C\. Garbe, J\. Lees, N\. Aziz, H\. Chaaban, J\. L\. Miller, P\. Johnson, and S\. DeLeon \(2018\)Off\-label medication use in children, more common than we think: a systematic review of the literature\.The Journal of the Oklahoma State Medical Association111\(8\),pp\. 776\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p2.1)\.
- G\. L\. Andriole, E\. D\. Crawford, R\. L\. Grubb III, S\. S\. Buys, D\. Chia, T\. R\. Church, M\. N\. Fouad, E\. P\. Gelmann, P\. A\. Kvale, D\. J\. Reding,et al\.\(2009\)Mortality results from a randomized prostate\-cancer screening trial\.New England journal of medicine360\(13\),pp\. 1310–1319\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p2.1)\.
- A\. N\. Angelopoulos, R\. F\. Barber, and S\. Bates \(2024\)Theoretical foundations of conformal prediction\.arXiv preprint arXiv:2411\.11824\.Cited by:[Figure 3](https://arxiv.org/html/2608.14598#S4.F3),[Figure 3](https://arxiv.org/html/2608.14598#S4.F3.9.2)\.
- R\. K\. Arora, J\. Wei, R\. S\. Hicks, P\. Bowman, J\. Quiñonero\-Candela, F\. Tsimpourlas, M\. Sharman, M\. Shah, A\. Vallone, A\. Beutel, J\. Heidecke, and K\. Singhal \(2025\)HealthBench: evaluating large language models towards improved human health\.Technical reportOpenAI\.Note:Technical ReportExternal Links:[Link](https://openai.com/research/healthbench)Cited by:[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p4.1)\.
- M\. Audrézet, J\. Chen, C\. Le Maréchal, P\. Ruszniewski, M\. Robaszkiewicz, O\. Raguénès, I\. Quéré, V\. Scotet, and C\. Férec \(2002\)Determination of the relative contribution of three genes–the cystic fibrosis transmembrane conductance regulator gene, the cationic trypsinogen gene, and the pancreatic secretory trypsin inhibitor gene–to the etiology of idiopathic chronic pancreatitis\.European Journal of Human Genetics10\(2\),pp\. 100–106\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.SSS0.Px1.p1.1)\.
- Baichuan\-M3 Team \(2026\)Baichuan\-M3: modeling clinical inquiry for reliable medical decision\-making\.arXiv preprint arXiv:2602\.06570\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.7.1.1.1.1)\.
- J\. M\. Beck \(2017\)FDA\-approved labeling ≠ medical standard of care\.Note:Accessed: 2025\-05\-17External Links:[Link](https://www.druganddevicelawblog.com/2017/02/fda-approved-labeling-%E2%89%A0-medical-standard-of-care.html)Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p2.1)\.
- S\. Bedi, H\. Cui, M\. Fuentes, A\. Unell, M\. Wornow, J\. M\. Banda, N\. Kotecha, T\. Keyes, Y\. Mai, M\. Oez,et al\.\(2026\)Holistic evaluation of large language models for medical tasks with medhelm\.Nature Medicine,pp\. 1–9\.Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- S\. Bedi, Y\. Liu, L\. Orr\-Ewing, D\. Dash, S\. Koyejo, A\. Callahan, J\. A\. Fries, M\. Wornow, A\. Swaminathan, L\. S\. Lehmann,et al\.\(2024\)Testing and evaluation of health care applications of large language models: a systematic review\.JAMA\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2608.14598#S2.p2.1)\.
- A\. Ben Abacha, W\. Yim, Y\. Fu, Z\. Sun, M\. Yetisgen, F\. Xia, and T\. Lin \(2025\)MEDEC: a benchmark for medical error detection and correction in clinical notes\.InFindings of the Association for Computational Linguistics: ACL 2025,External Links:[Link](https://aclanthology.org/2025.findings-acl.1159/)Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p4.1)\.
- D\. Bender and K\. Sartipi \(2013\)HL7 fhir: an agile and restful approach to healthcare information exchange\.InProceedings of the 26th IEEE international symposium on computer\-based medical systems,pp\. 326–331\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- K\. Blagec, J\. Kraiger, W\. Frühwirt, and M\. Samwald \(2023\)Benchmark datasets driving artificial intelligence development fail to capture the needs of medical professionals\.Journal of Biomedical Informatics137,pp\. 104274\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px3.p1.1)\.
- H\. Brath, N\. Mehta, R\. D\. Savage, S\. S\. Gill, W\. Wu, S\. E\. Bronskill, L\. Zhu, J\. H\. Gurwitz, and P\. A\. Rochon \(2018\)What is known about preventing, detecting, and reversing prescribing cascades: a scoping review\.Journal of the american geriatrics society66\(11\),pp\. 2079–2085\.Cited by:[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p3.2)\.
- T\. Brown, B\. Mann, N\. Ryder, M\. Subbiah, J\. D\. Kaplan, P\. Dhariwal, A\. Neelakantan, P\. Shyam, G\. Sastry, A\. Askell,et al\.\(2020\)Language models are few\-shot learners\.Advances in neural information processing systems33,pp\. 1877–1901\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- J\. C\. Cappelleri, P\. John, C\. H\. Schmid, S\. D\. de Ferranti, M\. Aubert, T\. C\. Chalmers, and J\. Lau \(1996\)Large trials vs meta\-analysis of smaller trials: how do their results compare?\.Jama276\(16\),pp\. 1332–1338\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1)\.
- H\. J\. Chang and P\. B\. Fontanarosa \(2011\)Introducing the jama clinical challenge\.JAMA305\(18\),pp\. 1910–1910\.Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- H\. Chen, Z\. Fang, Y\. Singla, and M\. Dredze \(2025a\)Benchmarking large language models on answering and explaining challenging medical questions\.InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 3563–3599\.Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1),[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p3.1)\.
- J\. Chen, Y\. Hu, M\. Cai, Y\. Lu, Y\. Wang, X\. Cao, M\. Lin, H\. Xu, J\. Wu, X\. Cao,et al\.\(2025b\)TrialBench: multi\-modal ai\-ready datasets for clinical trial prediction\.Scientific Data12\(1\),pp\. 1564\.Cited by:[§4\.2](https://arxiv.org/html/2608.14598#S4.SS2.p1.1),[§5](https://arxiv.org/html/2608.14598#S5.p2.1)\.
- S\. F\. Chen, A\. Alyakin, A\. Seas, E\. Yang, J\. J\. Choi, J\. V\. Lee, A\. L\. Chen, P\. I\. Warman, R\. T\. Bitolas, R\. J\. Steele,et al\.\(2026\)LLM\-assisted systematic review of large language models in clinical medicine\.Nature medicine,pp\. 1–8\.Cited by:[§A\.5](https://arxiv.org/html/2608.14598#A1.SS5.p1.1),[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p5.7)\.
- Z\. Chen, A\. H\. Cano, A\. Romanou, A\. Bonnet, K\. Matoba, F\. Salvi, M\. Pagliardini, S\. Fan, A\. Köpf, A\. Mohtashami,et al\.\(2023\)Meditron\-70b: scaling medical pretraining for large language models\.arXiv preprint arXiv:2311\.16079\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.10.1.1.1.1)\.
- C\. Christophe, P\. K\. Kanithi, T\. Raha, S\. Khan, and M\. A\. Pimentel \(2024\)Med42\-v2: a suite of clinical llms\.External Links:arXiv:2408\.06142,[Link](https://arxiv.org/abs/2408.06142)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.11.1.1.1.1)\.
- A\. L\. Cochrane \(1972\)Effectiveness and efficiency: random reflections on health services\.Nuffield Provincial Hospitals Trust,London\.Note:A foundational text in the development of evidence\-based medicineCited by:[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p1.1)\.
- J\. A\. Cohn, K\. J\. Friedman, P\. G\. Noone, M\. R\. Knowles, L\. M\. Silverman, and P\. S\. Jowell \(1998\)Relation between mutations of the cystic fibrosis gene and idiopathic pancreatitis\.New England Journal of Medicine339\(10\),pp\. 653–658\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.SSS0.Px1.p1.1)\.
- J\. Cole, R\. Zubirán, A\. Wolska, I\. Jialal, and A\. T\. Remaley \(2023\)Use of apolipoprotein b in the era of precision medicine: time for a paradigm change?\.Journal of Clinical Medicine12\(17\),pp\. 5737\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- J\. H\. Contois, M\. R\. Langlois, C\. Cobbaert, and A\. D\. Sniderman \(2023\)Standardization of apolipoprotein b, ldl\-cholesterol, and non\-hdl\-cholesterol\.Journal of the American Heart Association12\(15\),pp\. e030405\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- J\. Corbeil, A\. Dada, J\. Attendu, A\. Ben Abacha, A\. Sordoni, L\. Caccia, F\. Beaulieu, T\. Lin, J\. Kleesiek, and P\. Vozila \(2025\)A modular approach for clinical SLMs driven by synthetic data with pre\-instruction tuning, model merging, and clinical\-tasks alignment\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 19352–19374\.External Links:[Link](https://aclanthology.org/2025.acl-long.950/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.950),ISBN 979\-8\-89176\-251\-0Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.17.1.1.1.1)\.
- D\. De Oliveira\-Gomes, P\. H\. Joshi, E\. D\. Peterson, A\. Rohatgi, A\. Khera, and A\. M\. Navar \(2024\)Apolipoprotein b: bridging the gap between evidence and clinical practice\.Circulation150\(1\),pp\. 62–79\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- A\. Deaton and N\. Cartwright \(2018\)Understanding and misunderstanding randomized controlled trials\.Social science & medicine210,pp\. 2–21\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p4.1)\.
- R\. DerSimonian and R\. J\. Levine \(1999\)Resolving discrepancies between a meta\-analysis and a subsequent large controlled trial\.Jama282\(7\),pp\. 664–670\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1)\.
- R\. H\. Dolin and L\. Alschuler \(2011\)Approaching semantic interoperability in health level seven\.Journal of the American Medical Informatics Association18\(1\),pp\. 99–103\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- C\. Dwork \(2006\)Differential privacy\.InInternational colloquium on automata, languages, and programming,pp\. 1–12\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p1.1)\.
- A\. Fallahpour, M\. Alinoori, W\. Ye, X\. Cao, A\. Afkanpour, and A\. Krishnan \(2024\)EhrMamba: towards generalizable and scalable foundation models for electronic health records\.InProceedings of the 4th Machine Learning for Health Symposium,Proceedings of Machine Learning Research, Vol\.259,pp\. 291–307\.External Links:[Link](https://proceedings.mlr.press/v259/fallahpour25a.html)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.20.1.1.1.1)\.
- D\. Fast, L\. C\. Adams, F\. Busch, C\. Fallon, M\. Huppertz, R\. Siepmann, P\. Prucker, N\. Bayerl, D\. Truhn, M\. Makowski,et al\.\(2024\)Autonomous medical evaluation for guideline adherence of large language models\.NPJ Digital Medicine7\(1\),pp\. 1–14\.Cited by:[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p1.1)\.
- X\. Feng, D\. D\. Kim, J\. T\. Cohen, P\. J\. Neumann, and D\. A\. Ollendorf \(2020\)Using qalys versus dalys to measure cost\-effectiveness: how much does it matter?\.International journal of technology assessment in health care36\(2\),pp\. 96–103\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px6.p1.1)\.
- S\. Feuerriegel, D\. Frauen, V\. Melnychuk, J\. Schweisthal, K\. Hess, A\. Curth, S\. Bauer, N\. Kilbertus, I\. S\. Kohane, and M\. van der Schaar \(2024\)Causal machine learning for predicting treatment outcomes\.Nature Medicine30\(4\),pp\. 958–968\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px2.p1.1)\.
- A\. Finkelstein, P\. Persson, M\. Polyakova, and J\. M\. Shapiro \(2022\)A taste of their own medicine: guideline adherence and access to expertise\.American Economic Review: Insights4\(4\),pp\. 507–526\.Cited by:[§3\.3\.1](https://arxiv.org/html/2608.14598#S3.SS3.SSS1.p3.1)\.
- T\. R\. Fleming and D\. L\. DeMets \(1996\)Surrogate end points in clinical trials: are we being misled?\.Annals of internal medicine125\(7\),pp\. 605–613\.Cited by:[§1](https://arxiv.org/html/2608.14598#S1.p1.1)\.
- S\. P\. Forbes and I\. J\. Dahabreh \(2020\)Benchmarking observational analyses against randomized trials: a review of studies assessing propensity score methods\.Journal of general internal medicine35,pp\. 1396–1404\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1)\.
- T\. Fu, K\. Huang, C\. Xiao, L\. M\. Glass, and J\. Sun \(2022\)Hint: hierarchical interaction network for clinical\-trial\-outcome predictions\.Patterns3\(4\)\.Cited by:[§4\.2](https://arxiv.org/html/2608.14598#S4.SS2.p1.1)\.
- C\. Gao, J\. Pradeepkumar, T\. Das, S\. Thati, and J\. Sun \(2024\)Automatically labeling clinical trial outcomes: a large\-scale benchmark for drug development\.arXiv preprint arXiv:2406\.10292\.Cited by:[§4\.2](https://arxiv.org/html/2608.14598#S4.SS2.p1.1)\.
- L\. Gao, S\. Biderman, S\. Black, L\. Golding, T\. Hoppe, C\. Foster, J\. Phang, H\. He, A\. Thite, N\. Nabeshima,et al\.\(2020\)The pile: an 800gb dataset of diverse text for language modeling\.arXiv preprint arXiv:2101\.00027\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- S\. Gao, R\. Zhu, Z\. Kong, A\. Noori, X\. Su, C\. Ginder, T\. Tsiligkaridis, and M\. Zitnik \(2025\)TxAgent: an ai agent for therapeutic reasoning across a universe of tools\.arXiv preprint arXiv:2503\.10970\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.13.1.1.1.1),[Figure 1](https://arxiv.org/html/2608.14598#S3.F1),[Figure 1](https://arxiv.org/html/2608.14598#S3.F1.8.2),[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p1.1)\.
- S\. Gottlieb \(2011\)The FDA should not mandate comparative\-effectiveness trials\.Health Policy OutlookTechnical Report5,American Enterprise Institute\.External Links:[Link](https://www.aei.org/research-products/report/the-fda-should-not-mandate-comparative-effectiveness-trials/)Cited by:[§4\.2](https://arxiv.org/html/2608.14598#S4.SS2.p2.1)\.
- S\. M\. Grundy, N\. J\. Stone, A\. L\. Bailey, C\. Beam, K\. K\. Birtcher, R\. S\. Blumenthal, L\. T\. Braun, S\. De Ferranti, J\. Faiella\-Tommasino, D\. E\. Forman,et al\.\(2019\)2018 aha/acc/aacvpr/aapa/abc/acpm/ada/ags/apha/aspc/nla/pcna guideline on the management of blood cholesterol: a report of the american college of cardiology/american heart association task force on clinical practice guidelines\.Journal of the American College of Cardiology73\(24\),pp\. e285–e350\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- G\. Guyatt, J\. Cairns, D\. N\. Churchill, D\. Cook, B\. Haynes, J\. Hirsh, J\. Irvine, M\. Levine, M\. Levine, J\. Nishikawa,et al\.\(1992\)Evidence\-based medicine: a new approach to teaching the practice of medicine\.jama268\(17\),pp\. 2420–2425\.Cited by:[§1](https://arxiv.org/html/2608.14598#S1.p1.1),[§6](https://arxiv.org/html/2608.14598#S6.p6.1)\.
- G\. Guyatt, D\. Sackett, D\. W\. Taylor, J\. Ghong, R\. Roberts, and S\. Pugsley \(1986\)Determining optimal therapy—randomized trials in individual patients\.New England Journal of Medicine314\(14\),pp\. 889–892\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p4.1)\.
- M\. Ha, K\. Meyer, A\. Matos, J\. Turgeon, and C\. Bardolia \(2021\)Optimizing medications in patients with cardiovascular disease: a case report on unrecognized prescribing cascades in older adults\.Clinical Case Reports Journal2\(3\),pp\. 1–5\.Cited by:[Figure 2](https://arxiv.org/html/2608.14598#S3.F2),[Figure 2](https://arxiv.org/html/2608.14598#S3.F2.9.2),[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p3.2)\.
- C\. Hassan, M\. Spadaccini, Y\. Mori, F\. Foroutan, A\. Facciorusso, P\. Gkolfakis, G\. Tziatzios, K\. Triantafyllou, G\. Antonelli, K\. Khalaf,et al\.\(2023\)Real\-time computer\-aided detection of colorectal neoplasia during colonoscopy: a systematic review and meta\-analysis\.Annals of internal medicine176\(9\),pp\. 1209–1220\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p2.1)\.
- P\. Hegyi, M\. Wilschanski, S\. Muallem, G\. L\. Lukacs, M\. Sahin\-Tóth, A\. Uc, M\. A\. Gray, Z\. Rakonczay, and J\. Maléth \(2016\)CFTR: a new horizon in the pathomechanism and treatment of pancreatitis\.Reviews of Physiology, Biochemistry and Pharmacology Vol\. 170,pp\. 37–66\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.SSS0.Px1.p1.1)\.
- D\. Hendrycks, C\. Burns, S\. Basart, A\. Zou, M\. Mazeika, D\. Song, and J\. Steinhardt \(2021\)Measuring massive multitask language understanding\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=d7KBjmI3GmQ)Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- M\. A\. Hernán and J\. M\. Robins \(2016\)Using big data to emulate a target trial when a randomized trial is not available\.American journal of epidemiology183\(8\),pp\. 758–764\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2608.14598#S5.p2.1)\.
- M\. A\. Hernán, W\. Wang, and D\. E\. Leaf \(2022\)Target trial emulation: a framework for causal inference from observational data\.Jama328\(24\),pp\. 2446–2447\.Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p2.1)\.
- R\. S\. Hicks, M\. Trofimov, D\. Lim, R\. K\. Arora, F\. Tsimpourlas, P\. Bowman, M\. Sharman, C\. Tong, K\. Karthik, A\. Dugar, A\. Jagadeesh, K\. Saab, J\. Heidecke, A\. Alexander, N\. Gross, and K\. Singhal \(2026\)HealthBench professional: evaluating large language models on real clinician chats\.External Links:2604\.27470,[Link](https://arxiv.org/abs/2604.27470)Cited by:[§3\.3\.1](https://arxiv.org/html/2608.14598#S3.SS3.SSS1.p2.1)\.
- G\. Hripcsak, J\. D\. Duke, N\. H\. Shah, C\. G\. Reich, V\. Huser, M\. J\. Schuemie, M\. A\. Suchard, R\. W\. Park, I\. C\. K\. Wong, P\. R\. Rijnbeek,et al\.\(2015\)Observational health data sciences and informatics \(ohdsi\): opportunities for observational researchers\.InMEDINFO 2015: eHealth\-enabled Health,pp\. 574–578\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px4.p1.1)\.
- R\. T\. Hurt, C\. R\. Stephenson, E\. A\. Gilman, C\. A\. Aakre, I\. T\. Croghan, M\. S\. Mundi, K\. Ghosh, and J\. Edakkanambeth Varayil \(2025\)The use of an artificial intelligence platform openevidence to augment clinical decision\-making for primary care physicians\.Journal of Primary Care & Community Health16,pp\. 21501319251332215\.Cited by:[Figure 15](https://arxiv.org/html/2608.14598#A1.F15),[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p4.1)\.
- P\. B\. Jensen, L\. J\. Jensen, and S\. Brunak \(2012\)Mining electronic health records: towards better research applications and clinical care\.Nature Reviews Genetics13\(6\),pp\. 395–405\.External Links:[Document](https://dx.doi.org/10.1038/nrg3208)Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p1.1)\.
- D\. Jin, E\. Pan, N\. Oufattole, W\. Weng, H\. Fang, and P\. Szolovits \(2021\)What disease does this patient have? a large\-scale open domain question answering dataset from medical exams\.Applied Sciences11\(14\),pp\. 6421\.Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- Q\. Jin, B\. Dhingra, Z\. Liu, W\. W\. Cohen, and X\. Lu \(2019\)PubMedQA: a dataset for biomedical research question answering\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),Hong Kong, China,pp\. 2567–2577\.External Links:[Link](https://aclanthology.org/D19-1259),[Document](https://dx.doi.org/10.18653/v1/D19-1259)Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- S\. Joshi, I\. Urteaga, W\. A\. C\. van Amsterdam, G\. Hripcsak, P\. Elias, B\. Recht, N\. Elhadad, J\. Fackler, M\. P\. Sendak, J\. Wiens, K\. Deshpande, Y\. Wald, M\. Fiterau, Z\. Lipton, D\. Malinsky, M\. Nayan, H\. Namkoong, S\. Park, J\. E\. Vogt, and R\. Ranganath \(2025\)AI as an intervention: improving clinical outcomes relies on a causal approach to ai development and validation\.Journal of the American Medical Informatics Association32\(3\),pp\. 589–594\.External Links:ISSN 1527\-974X,[Document](https://dx.doi.org/10.1093/jamia/ocae301),[Link](https://doi.org/10.1093/jamia/ocae301),https://academic\.oup\.com/jamia/article\-pdf/32/3/589/61371325/ocae301\.pdfCited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px1.p1.1),[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px1.p2.1),[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px1.p3.1),[§5](https://arxiv.org/html/2608.14598#S5.p3.1)\.
- A\. Karargyris, R\. Umeton, M\. J\. Sheller, A\. Aristizabal, J\. George, A\. Wuest, S\. Pati, H\. Kassem, M\. Zenk, U\. Baid,et al\.\(2023\)Federated benchmarking of medical artificial intelligence with medperf\.Nature machine intelligence5\(7\),pp\. 799–810\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px4.p1.1)\.
- S\. Kaul and G\. Gordon \(2025\)Meta\-analysis with untrusted data\.InProceedings of the 4th Machine Learning for Health Symposium,S\. Hegselmann, H\. Zhou, E\. Healey, T\. Chang, C\. Ellington, V\. Mhasawade, S\. Tonekaboni, P\. Argaw, and H\. Zhang \(Eds\.\),Proceedings of Machine Learning Research, Vol\.259,pp\. 563–593\.External Links:[Link](https://proceedings.mlr.press/v259/kaul25a.html)Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p4.1)\.
- D\. M\. Kent and J\. P\. Ioannidis \(2026\)Adversarial collaboration as a strategy for credible biomedical science\.JAMA\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p5.1)\.
- A\. Khurshid and I\. N\. Sarkar \(2025\)The health data utility and the resurgence of health information exchanges as a national resource\.Journal of the American Medical Informatics Association32\(5\),pp\. 964–967\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- Y\. Labrak, A\. Bazoge, E\. Morin, P\. Gourraud, M\. Rouvier, and R\. Dufour \(2024\)BioMistral: a collection of open\-source pretrained large language models for medical domains\.InFindings of the Association for Computational Linguistics: ACL 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),Bangkok, Thailand,pp\. 5848–5864\.External Links:[Link](https://aclanthology.org/2024.findings-acl.348/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.348)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.16.1.1.1.1)\.
- J\. LeLorier, G\. Gregoire, A\. Benhaddad, J\. Lapierre, and F\. Derderian \(1997\)Discrepancies between meta\-analyses and subsequent large randomized, controlled trials\.New England Journal of Medicine337\(8\),pp\. 536–542\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1)\.
- C\. Li, C\. Wong, S\. Zhang, N\. Usuyama, H\. Liu, J\. Yang, T\. Naumann, H\. Poon, and J\. Gao \(2023\)Llava\-med: training a large language\-and\-vision assistant for biomedicine in one day\.Advances in Neural Information Processing Systems36,pp\. 28541–28564\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.15.1.1.1.1)\.
- Q\. Li, L\. Deleger, T\. Lingren, H\. Zhai, M\. Kaiser, L\. Stoutenborough, A\. G\. Jegga, K\. B\. Cohen, and I\. Solti \(2013\)Mining fda drug labels for medical conditions\.BMC medical informatics and decision making13,pp\. 1–11\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p1.1)\.
- P\. Liang, R\. Bommasani, T\. Lee, D\. Tsipras, D\. Soylu, M\. Yasunaga, Y\. Zhang, D\. Narayanan, Y\. Wu, A\. Kumar, B\. Newman, B\. Yuan, B\. Yan, C\. Zhang, C\. Cosgrove, C\. D\. Manning, C\. Re, D\. Acosta\-Navas, D\. A\. Hudson, E\. Zelikman, E\. Durmus, F\. Ladhak, F\. Rong, H\. Ren, H\. Yao, J\. WANG, K\. Santhanam, L\. Orr, L\. Zheng, M\. Yuksekgonul, M\. Suzgun, N\. Kim, N\. Guha, N\. S\. Chatterji, O\. Khattab, P\. Henderson, Q\. Huang, R\. A\. Chi, S\. M\. Xie, S\. Santurkar, S\. Ganguli, T\. Hashimoto, T\. Icard, T\. Zhang, V\. Chaudhary, W\. Wang, X\. Li, Y\. Mai, Y\. Zhang, and Y\. Koreeda \(2023\)Holistic evaluation of language models\.Transactions on Machine Learning Research\.Note:Featured Certification, Expert Certification, Outstanding CertificationExternal Links:ISSN 2835\-8856,[Link](https://openreview.net/forum?id=iO4LZibEqW)Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p3.1)\.
- T\. Liang and B\. Recht \(2025\)Randomization inference when n equals one\.Biometrika,pp\. asaf013\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p4.1)\.
- T\. Lin, W\. Zhang, S\. Li, Y\. Yuan, B\. Yu, H\. Li, W\. He, H\. Jiang, M\. Li, X\. Song, S\. Tang, J\. Xiao, H\. Lin, Y\. Zhuang, and B\. C\. Ooi \(2025\)HealthGPT: a medical large vision\-language model for unifying comprehension and generation via heterogeneous knowledge adaptation\.InForty\-second International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=WbP2OwMULq)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.14.1.1.1.1)\.
- R\. Liu, P\. Chen, and P\. Zhang \(2024\)CURE: a deep learning framework pre\-trained on large\-scale patient data for treatment effect estimation\.Patterns5\(6\)\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px2.p1.1)\.
- X\. Luo and M\. L\. Li \(2025\)The CRITICAL records integrated standardization pipeline \(CRISP\): end\-to\-end processing of large\-scale multi\-institutional OMOP CDM data\.InProceedings of the 5th Machine Learning for Health Symposium,Proceedings of Machine Learning Research, Vol\.297\.External Links:[Link](https://proceedings.mlr.press/v297/luo26a.html)Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p1.1)\.
- Y\. Luo, L\. Rasmussen, A\. Williams, J\. Houghtaling, A\. Michelson, P\. Payne, J\. Cimino, J\. Osborne, P\. Nannapaneni, Y\. Li, S\. Amagai, M\. Hutch, J\. Ding, A\. Kline, C\. A\. Gao, J\. Womack, R\. Zhu, M\. Wyatt, R\. Brown, S\. Gupta, N\. Venteris, A\. Gupta, W\. Ma, E\. Hillis, P\. Cannon, M\. Alvarez, J\. Starren, K\. Holmes, M\. Carson, A\. Lai, and A\. Wilcox \(2024\)Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p1.1)\.
- P\. M\. Marcus, E\. J\. Bergstralh, R\. M\. Fagerstrom, D\. E\. Williams, R\. Fontana, W\. F\. Taylor, and P\. C\. Prorok \(2000\)Lung cancer mortality in the mayo lung project: impact of extended follow\-up\.Journal of the National Cancer Institute92\(16\),pp\. 1308–1316\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p2.1)\.
- E\. A\. McGlynn, S\. M\. Asch, J\. Adams, J\. Keesey, J\. Hicks, A\. DeCristofaro, and E\. A\. Kerr \(2003\)The quality of health care delivered to adults in the united states\.New England Journal of Medicine348\(26\),pp\. 2635–2645\.External Links:[Document](https://dx.doi.org/10.1056/NEJMsa022615)Cited by:[§3\.3\.1](https://arxiv.org/html/2608.14598#S3.SS3.SSS1.p3.1)\.
- P\. E\. Meehl \(1954\)Clinical versus statistical prediction: a theoretical analysis and a review of the evidence\.\.Cited by:[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p1.1)\.
- A\. B\. Miller, C\. Wall, C\. J\. Baines, P\. Sun, T\. To, and S\. A\. Narod \(2014\)Twenty five year follow\-up for breast cancer incidence and mortality of the canadian national breast screening study: randomised screening trial\.Bmj348\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p2.1)\.
- F\. Moramarco, A\. Kroese, G\. Higgins, D\. S\. Pappas, and M\. Coccaro \(2022\)Human evaluation and correlation with automatic metrics in consultation note generation\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 5739–5754\.Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p4.1)\.
- M\. R\. Morris, J\. Sohl\-Dickstein, N\. Fiedel, T\. Warkentin, A\. Dafoe, A\. Faust, C\. Farabet, and S\. Legg \(2024\)Levels of AGI: operationalizing progress on the path to AGI\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),External Links:[Link](https://arxiv.org/abs/2311.02462)Cited by:[§1](https://arxiv.org/html/2608.14598#S1.p2.1)\.
- M\. B\. Mortensen \(2024\)ApoB triumphs once more over LDL\-C and non\-HDL\-C in risk prediction: ready for guidelines?\.European Heart Journal45\(27\),pp\. 2419–2421\.External Links:[Document](https://dx.doi.org/10.1038/eurheartj/ehae257),[Link](https://doi.org/10.1038/eurheartj/ehae257)Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- H\. Nilforoshan, M\. Moor, Y\. Roohani, Y\. Chen, A\. Šurina, M\. Yasunaga, S\. Oblak, and J\. Leskovec \(2023\)Zero\-shot causal learning\.Advances in Neural Information Processing Systems36,pp\. 6862–6901\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px2.p1.1)\.
- G\. Noarov, R\. Ramalingam, A\. Roth, and S\. Xie \(2025\)High\-dimensional prediction for sequential decision making\.InProceedings of the 42nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.267,pp\. 46762–46783\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p3.1)\.
- G\. Noarov and A\. Roth \(2024\)Calibration for decision making: a principled approach to trustworthy ml\.Note:[https://www\.let\-all\.com/blog/2024/03/13/calibration\-for\-decision\-making\-a\-principled\-approach\-to\-trustworthy\-ml/](https://www.let-all.com/blog/2024/03/13/calibration-for-decision-making-a-principled-approach-to-trustworthy-ml/)Blog postCited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p3.1)\.
- A\. Pal, L\. K\. Umapathi, and M\. Sankarasubbu \(2022\)Medmcqa: a large\-scale multi\-subject multi\-choice dataset for medical domain question answering\.InConference on health, inference, and learning,pp\. 248–260\.Cited by:[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- M\. S\. A\. Pal and M\. Sankarasubbu \(2024\)Openbiollms: advancing open\-source large language models for healthcare and life sciences\.Hugging Face repository\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.8.1.1.1.1)\.
- A\. Palepu, V\. Liévin, W\. Weng, K\. Saab, D\. Stutz, Y\. Cheng, K\. Kulkarni, S\. S\. Mahdavi, J\. Barral, D\. R\. Webster,et al\.\(2025\)Towards conversational ai for disease management\.arXiv preprint arXiv:2503\.06074\.Cited by:[Figure 4](https://arxiv.org/html/2608.14598#A1.F4),[Figure 4](https://arxiv.org/html/2608.14598#A1.F4.4.2),[Figure 1](https://arxiv.org/html/2608.14598#S3.F1),[Figure 1](https://arxiv.org/html/2608.14598#S3.F1.8.2),[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p1.1),[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p3.1),[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p1.1)\.
- J\. Pearl \(2015\)Generalizing experimental findings\.Journal of Causal Inference3\(2\),pp\. 259–266\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p4.1)\.
- G\. J\. Pearson, G\. Thanassoulis, T\. J\. Anderson, A\. R\. Barry, P\. Couture, N\. Dayan, G\. A\. Francis, J\. Genest, J\. Grégoire, S\. A\. Grover,et al\.\(2021\)2021 canadian cardiovascular society guidelines for the management of dyslipidemia for the prevention of cardiovascular disease in adults\.Canadian journal of cardiology37\(8\),pp\. 1129–1150\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- S\. D\. Pearson \(2018\)The icer value framework: integrating cost effectiveness and affordability in the assessment of health care value\.Value in health21\(3\),pp\. 258–265\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px6.p1.1)\.
- P\. G\. Peters Jr \(2023\)Modernizing the medical malpractice standard of care\.Sw\. L\. Rev\.52,pp\. 465\.Cited by:[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p2.1)\.
- M\. Y\. Phadke and Z\. M\. Sellers \(2022\)Current clinical opinion on cftr dysfunction and patient risk of pancreatitis: diagnostic and therapeutic considerations\.Expert review of gastroenterology & hepatology16\(6\),pp\. 499–509\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.SSS0.Px1.p1.1)\.
- D\. Phelan, D\. Gottlieb, J\. C\. Mandel, V\. Ignatov, J\. Jones, B\. Marquard, A\. Ellis, and K\. D\. Mandl \(2024\)Beyond compliance with the 21st century cures act rule: a patient controlled electronic health information export application programming interface\.Journal of the American Medical Informatics Association31\(4\),pp\. 901–909\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- B\. Plank \(2022\)The “problem” of human label variation: on ground truth in data, modeling and evaluation\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 10671–10682\.Cited by:[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p3.1)\.
- R\. Platt, J\. S\. Brown, M\. Robb, M\. McClellan, R\. Ball, M\. D\. Nguyen, and R\. E\. Sherman \(2018\)The fda sentinel initiative—an evolving national resource\.N Engl J Med379\(22\),pp\. 2091–2093\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px4.p1.1)\.
- V\. Prasad, A\. Vandross, C\. Toomey, M\. Cheung, J\. Rho, S\. Quinn, S\. J\. Chacko, D\. Borkar, V\. Gall, S\. Selvaraj,et al\.\(2013\)A decade of reversal: an analysis of 146 contradicted medical practices\.InMayo Clinic Proceedings,Vol\.88,pp\. 790–798\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p5.1)\.
- M\. Prosperi, Y\. Guo, M\. Sperrin, J\. S\. Koopman, J\. S\. Min, X\. He, S\. Rich, M\. Wang, I\. E\. Buchan, and J\. Bian \(2020\)Causal inference and counterfactual prediction in machine learning for actionable healthcare\.Nature Machine Intelligence2\(7\),pp\. 369–375\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px2.p1.1)\.
- I\. D\. Raji, R\. Daneshjou, and E\. Alsentzer \(2025\)It’s time to bench the medical exam benchmark\.Vol\.2,Massachusetts Medical Society\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px3.p1.1),[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- A\. Ratner, S\. H\. Bach, H\. Ehrenberg, J\. Fries, S\. Wu, and C\. Ré \(2017\)Snorkel: rapid training data creation with weak supervision\.InProceedings of the VLDB endowment\. International conference on very large data bases,Vol\.11,pp\. 269\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p2.1)\.
- J\. G\. Richens, C\. M\. Lee, and S\. Johri \(2020\)Improving the accuracy of medical diagnosis with causal machine learning\.Nature communications11\(1\),pp\. 3923\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px2.p1.1)\.
- N\. Rieke, J\. Hancox, W\. Li, F\. Milletari, H\. R\. Roth, S\. Albarqouni, S\. Bakas, M\. N\. Galtier, B\. A\. Landman, K\. Maier\-Hein,et al\.\(2020\)The future of digital health with federated learning\.NPJ digital medicine3\(1\),pp\. 119\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px4.p1.1)\.
- P\. A\. Rochon and J\. H\. Gurwitz \(1997\)Optimising drug treatment for elderly people: the prescribing cascade\.Bmj315\(7115\),pp\. 1096–1099\.Cited by:[§3\.2](https://arxiv.org/html/2608.14598#S3.SS2.p3.2)\.
- J\. G\. Rossi, N\. Rojas\-Perilla, J\. Krois, and F\. Schwendicke \(2022\)Cost\-effectiveness of artificial intelligence as a decision\-support system applied to the detection and grading of melanoma, dental caries, and diabetic retinopathy\.JAMA Network Open5\(3\),pp\. e220269–e220269\.Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p3.1)\.
- G\. N\. Rothblum and G\. Yona \(2023\)Decision\-making under miscalibration\.InITCS,pp\. 92:1–92:20\.External Links:[Link](https://doi.org/10.4230/LIPIcs.ITCS.2023.92)Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p3.1)\.
- P\. M\. Rothwell \(2005\)External validity of randomised controlled trials:“to whom do the results of this trial apply?”\.The Lancet365\(9453\),pp\. 82–93\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p4.1)\.
- W\. B\. Runciman, T\. D\. Hunt, N\. A\. Hannaford, P\. D\. Hibbert, J\. I\. Westbrook, E\. W\. Coiera, R\. O\. Day, D\. M\. Hindmarsh, E\. A\. McGlynn, and J\. Braithwaite \(2012\)CareTrack: assessing the appropriateness of health care delivery in australia\.Medical Journal of Australia197\(2\),pp\. 100–105\.External Links:[Document](https://dx.doi.org/10.5694/mja12.10510)Cited by:[§3\.3\.1](https://arxiv.org/html/2608.14598#S3.SS3.SSS1.p3.1)\.
- K\. Saab, T\. Tu, W\. Weng, R\. Tanno, D\. Stutz, E\. Wulczyn, F\. Zhang, T\. Strother, C\. Park, E\. Vedadi, J\. Z\. Chaves, S\. Hu, M\. Schaekermann, A\. Kamath, Y\. Cheng, D\. G\. T\. Barrett, C\. Cheung, B\. Mustafa, A\. Palepu, D\. McDuff, L\. Hou, T\. Golany, L\. Liu, J\. Alayrac, N\. Houlsby, N\. Tomasev, J\. Freyberg, C\. Lau, J\. Kemp, J\. Lai, S\. Azizi, K\. Kanada, S\. Man, K\. Kulkarni, R\. Sun, S\. Shakeri, L\. He, B\. Caine, A\. Webson, N\. Latysheva, M\. Johnson, P\. A\. Mansfield, J\. Lu, E\. Rivlin, J\. Anderson, B\. Green, R\. Wong, J\. Krause, J\. Shlens, E\. Dominowska, S\. M\. A\. Eslami, K\. Chou, C\. Cui, O\. Vinyals, K\. Kavukcuoglu, J\. Manyika, J\. Dean, D\. Hassabis, Y\. Matias, D\. R\. Webster, J\. K\. Barral, G\. Corrado, C\. Semturs, S\. S\. Mahdavi, J\. Gottweis, A\. Karthikesalingam, and V\. Natarajan \(2024\)Capabilities of gemini models in medicine\.CoRRabs/2404\.18416\.External Links:[Link](https://doi.org/10.48550/arXiv.2404.18416)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.3.1.1.1.1),[§4\.1](https://arxiv.org/html/2608.14598#S4.SS1.p1.1)\.
- D\. L\. Sackett, W\. M\. Rosenberg, J\. M\. Gray, R\. B\. Haynes, and W\. S\. Richardson \(1996\)Evidence based medicine: what it is and what it isn’t\.Vol\.312,British Medical Journal Publishing Group\.Cited by:[§1](https://arxiv.org/html/2608.14598#S1.p1.1),[§6](https://arxiv.org/html/2608.14598#S6.p6.1)\.
- A\. Sellergren, S\. Kazemzadeh, T\. Jaroensri, A\. Kiraly, M\. Traverse, T\. Kohlberger, S\. Xu, F\. Jamil, C\. Hughes, C\. Lau,et al\.\(2025\)Medgemma technical report\.arXiv preprint arXiv:2507\.05201\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.12.1.1.1.1)\.
- T\. M\. Shaneyfelt, M\. F\. Mayo\-Smith, and J\. Rothwangl \(1999\)Are guidelines following guidelines?: the methodological quality of clinical practice guidelines in the peer\-reviewed medical literature\.Jama281\(20\),pp\. 1900–1905\.Cited by:[§6](https://arxiv.org/html/2608.14598#S6.p5.1)\.
- R\. E\. Sherman, S\. A\. Anderson, G\. J\. Dal Pan, G\. W\. Gray, T\. P\. Gross, N\. L\. Hunter, L\. M\. LaVange, D\. Marinac\-Dabic, P\. W\. Marks, M\. A\. Robb, J\. E\. Shuren, R\. J\. Temple, J\. Woodcock, L\. Q\. Yue, and R\. M\. Califf \(2016\)Real\-world evidence — what is it and what can it tell us?\.New England Journal of Medicine375\(23\),pp\. 2293–2297\.External Links:[Document](https://dx.doi.org/10.1056/NEJMsb1609216),[Link](https://doi.org/10.1056/NEJMsb1609216)Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p1.1)\.
- K\. Singhal, S\. Azizi, T\. Tu, S\. S\. Mahdavi, J\. Wei, H\. W\. Chung, N\. Scales, A\. Tanwani, H\. Cole\-Lewis, S\. Pfohl,et al\.\(2023\)Large language models encode clinical knowledge\.Nature620\(7972\),pp\. 172–180\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.6.1.1.1.1),[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p2.1)\.
- K\. Singhal, T\. Tu, J\. Gottweis, R\. Sayres, E\. Wulczyn, M\. Amin, L\. Hou, K\. Clark, S\. R\. Pfohl, H\. Cole\-Lewis,et al\.\(2025\)Toward expert\-level medical question answering with large language models\.Nature Medicine,pp\. 1–8\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.5.1.1.1.1),[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p2.1)\.
- A\. D\. Sniderman, L\. Dufresne, K\. M\. Pencina, S\. Bilgic, G\. Thanassoulis, and M\. J\. Pencina \(2024\)Discordance among apob, non–high\-density lipoprotein cholesterol, and triglycerides: implications for cardiovascular prevention\.European Heart Journal45\(27\),pp\. 2410–2418\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- A\. D\. Sniderman, G\. Thanassoulis, T\. Glavinovic, A\. M\. Navar, M\. Pencina, A\. Catapano, and B\. A\. Ference \(2019\)Apolipoprotein b particles and cardiovascular disease: a narrative review\.JAMA cardiology4\(12\),pp\. 1287–1295\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- D\. E\. Soffer, N\. A\. Marston, K\. C\. Maki, T\. A\. Jacobson, V\. A\. Bittner, J\. M\. Peña, G\. Thanassoulis, S\. S\. Martin, C\. F\. Kirkpatrick, S\. S\. Virani,et al\.\(2024\)Role of apolipoprotein b in the clinical management of cardiovascular risk in adults: an expert clinical consensus from the national lipid association\.Journal of clinical lipidology18\(5\),pp\. e647–e663\.Cited by:[§A\.3\.2](https://arxiv.org/html/2608.14598#A1.SS3.SSS2.Px1.p1.1)\.
- E\. Steinberg, J\. A\. Fries, Y\. Xu, and N\. Shah \(2024\)MOTOR: a time\-to\-event foundation model for structured medical records\.InThe Twelfth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=NialiwI2V6)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.21.1.1.1.1)\.
- T\. Tang, V\. B\. Cruz, and L\. L\. Konczal \(2022\)Idiopathic chronic pancreatitis treated with ivacaftor in a cftr carrier with methylmalonic acidemia\.Journal of Cystic Fibrosis21\(4\),pp\. 603–605\.Cited by:[Figure 5](https://arxiv.org/html/2608.14598#A1.F5.pic2.1.1.1.1.1.1.1),[Figure 6](https://arxiv.org/html/2608.14598#A1.F6.pic2.1.1.1.1.1.1.1),[Figure 8](https://arxiv.org/html/2608.14598#A1.F8),[Figure 8](https://arxiv.org/html/2608.14598#A1.F8.4.2),[Figure 8](https://arxiv.org/html/2608.14598#A1.F8.pic1.1.1.1.1.1.1.1),[Figure 8](https://arxiv.org/html/2608.14598#A1.F8.pic2.1.1.1.1.1.1.1),[Figure 9](https://arxiv.org/html/2608.14598#A1.F9),[Figure 9](https://arxiv.org/html/2608.14598#A1.F9.4.2),[Figure 9](https://arxiv.org/html/2608.14598#A1.F9.pic1.1.1.1.1.1.1.1),[Figure 1](https://arxiv.org/html/2608.14598#S3.F1),[Figure 1](https://arxiv.org/html/2608.14598#S3.F1.8.2),[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.SSS0.Px1.p1.1)\.
- R\. Temple \(1999\)Are surrogate markers adequate to assess cardiovascular disease drugs?\.Jama282\(8\),pp\. 790–795\.Cited by:[§1](https://arxiv.org/html/2608.14598#S1.p1.1)\.
- A\. A\. Tierney, G\. Gayre, B\. Hoberman, B\. Mattern, M\. Ballesca, P\. Kipnis, V\. Liu, and K\. Lee \(2024\)Ambient artificial intelligence scribes to alleviate the burden of clinical documentation\.NEJM Catalyst Innovations in Care Delivery5\(3\),pp\. CAT–23\.Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p3.1)\.
- T\. Tu, M\. Schaekermann, A\. Palepu, K\. Saab, J\. Freyberg, R\. Tanno, A\. Wang, B\. Li, M\. Amin, Y\. Cheng,et al\.\(2025\)Towards conversational diagnostic artificial intelligence\.Nature,pp\. 1–9\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.4.1.1.1.1)\.
- A\. Tversky and D\. Kahneman \(1974\)Judgment under uncertainty: heuristics and biases: biases in judgments reveal some heuristics of thinking under uncertainty\.\.science185\(4157\),pp\. 1124–1131\.Cited by:[§3\.3](https://arxiv.org/html/2608.14598#S3.SS3.p1.1)\.
- T\. ValizadehAslani, Y\. Shi, P\. Ren, J\. Wang, Y\. Zhang, M\. Hu, L\. Zhao, and H\. Liang \(2023\)PharmBERT: a domain\-specific bert model for drug labels\.Briefings in bioinformatics24\(4\),pp\. bbad226\.Cited by:[§3\.1](https://arxiv.org/html/2608.14598#S3.SS1.p1.1)\.
- K\. G\. van Leeuwen, F\. J\. Meijer, S\. Schalekamp, M\. J\. Rutten, E\. J\. van Dijk, B\. van Ginneken, T\. M\. Govers, and M\. de Rooij \(2021\)Cost\-effectiveness of artificial intelligence aided vessel occlusion detection in acute stroke: an early health technology assessment\.Insights into imaging12,pp\. 1–9\.Cited by:[§5](https://arxiv.org/html/2608.14598#S5.p3.1)\.
- J\. Villar, G\. Carroli, and J\. Belizan \(1995\)Predictive ability of meta\-analyses of randomised controlled trials\.The lancet345\(8952\),pp\. 772–776\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1)\.
- S\. V\. Wang, S\. Schneeweiss, J\. M\. Franklin, R\. J\. Desai, W\. Feldman, E\. M\. Garry, R\. J\. Glynn, K\. J\. Lin, J\. Paik, E\. Patorno,et al\.\(2023\)Emulation of randomized clinical trials with nonrandomized database analyses: results of 32 clinical trials\.Jama329\(16\),pp\. 1376–1385\.Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px5.p1.1)\.
- S\. Waxler, P\. Blazek, D\. White, D\. Sneider, K\. Chung, M\. Nagarathnam, P\. Williams, H\. Voeller, K\. Wong, M\. Swanhorst,et al\.\(2025\)Generative medical event models improve with scale\.arXiv preprint arXiv:2508\.12104\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.18.1.1.1.1)\.
- H\. G\. Welch and W\. C\. Black \(2010\)Overdiagnosis in cancer\.Journal of the National Cancer Institute102\(9\),pp\. 605–613\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p2.1)\.
- M\. Wornow, S\. Bedi, M\. A\. F\. Hernandez, E\. Steinberg, J\. A\. Fries, C\. Re, S\. Koyejo, and N\. Shah \(2025\)Context clues: evaluating long context models for clinical prediction tasks on EHR data\.InThe Thirteenth International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zg3ec1TdAP)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.19.1.1.1.1)\.
- M\. Wornow, Y\. Xu, R\. Thapa, B\. Patel, E\. Steinberg, S\. Fleming, M\. A\. Pfeffer, J\. Fries, and N\. H\. Shah \(2023\)The shaky foundations of large language models and foundation models for electronic health records\.npj digital medicine6\(1\),pp\. 135\.Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1),[Table 1](https://arxiv.org/html/2608.14598#S1.T1.96.2)\.
- Q\. Xie, Q\. Chen, A\. Chen, C\. Peng, Y\. Hu, F\. Lin, X\. Peng, J\. Huang, J\. Zhang, V\. Keloth, X\. Zhou, L\. Qian, H\. He, D\. Shung, L\. Ohno\-Machado, Y\. Wu, H\. Xu, and J\. Bian \(2025\)Medical foundation large language models for comprehensive text analysis and beyond\.npj Digital Medicine8\(1\),pp\. 141\.External Links:[Document](https://dx.doi.org/10.1038/s41746-025-01533-1),[Link](https://doi.org/10.1038/s41746-025-01533-1)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.9.1.1.1.1)\.
- L\. Yang, S\. Xu, A\. Sellergren, T\. Kohlberger, Y\. Zhou, I\. Ktena, A\. P\. Kiraly, F\. Ahmed, F\. Hormozdiari, T\. Jaroensri, E\. Wang, E\. Wulczyn, F\. Jamil, T\. Guidroz, C\. Lau, S\. Qiao, Y\. Liu, A\. Goel, K\. Park, A\. Agharwal, N\. George, Y\. Wang, R\. Tanno, D\. G\. T\. Barrett, W\. Weng, S\. S\. Mahdavi, K\. Saab, T\. Tu, S\. R\. Kalidindi, M\. Etemadi, J\. Cuadros, G\. Sorensen, Y\. Matias, K\. Chou, G\. Corrado, J\. K\. Barral, S\. Shetty, D\. J\. Fleet, S\. M\. A\. Eslami, D\. Tse, S\. Prabhakara, C\. Y\. McLean, D\. Steiner, R\. Pilgrim, C\. Kelly, S\. Azizi, and D\. Golden \(2024\)Advancing multimodal medical capabilities of gemini\.CoRRabs/2405\.03162\.External Links:[Link](https://doi.org/10.48550/arXiv.2405.03162)Cited by:[Table 1](https://arxiv.org/html/2608.14598#S1.T1.92.94.2.1),[Table 1](https://arxiv.org/html/2608.14598#S1.T1.98.1.2.1.1.1.1)\.
- J\. S\. Yudkin, K\. J\. Lipska, and V\. M\. Montori \(2011\)The idolatry of the surrogate\.Bmj343\.Cited by:[§1](https://arxiv.org/html/2608.14598#S1.p1.1)\.
- J\. Zhang, J\. Jennings, A\. Hilmkil, N\. Pawlowski, C\. Zhang, and C\. Ma \(2024\)Towards causal foundation model: on duality between optimal balancing and attention\.InForty\-first International Conference on Machine Learning,Cited by:[§A\.1](https://arxiv.org/html/2608.14598#A1.SS1.SSS0.Px2.p1.1)\.
- S\. Zhao, M\. Kim, R\. Sahoo, T\. Ma, and S\. Ermon \(2021\)Calibrating predictions to decisions: a novel approach to multi\-class calibration\.Advances in Neural Information Processing Systems34,pp\. 22313–22324\.Cited by:[§4\.3](https://arxiv.org/html/2608.14598#S4.SS3.p3.1)\.
## Appendix AAppendix
### A\.1Related Work
##### Emphasizing outcomes in medical AI
The recent paper ofJoshiet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib20)\)also calls for greater emphasis on clinical outcomes in medical AI\. However, their work is very different than ours\. They recognize that deployments of predictive AI are interventions, and propose to analyze them from a causal perspective\. They discuss, for example, the pitfalls of deploying a model which predicts heart disease from inexpensive ECGs, when its training was conducted on a smaller, biased population with expensive ground truth labels\. To catch problems such as this distribution shift, they present a prospective \(and continuing\) validation framework using silent trials or off\-policy evaluation\.
Our paper is about a different topic: how to train medical AI \(e\.g\. reasoning language models\) to understand of the effects of existing treatments, such as pharmaceuticals\. Our position is that popular substitutes for real treatment outcome data \(such as clinical guidelines\) do not facilitate the development of such AI\. The differences between our papers become evident whenJoshiet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib20)\)discuss randomized trials\. They correctly recognize that conducting RCTs for AI deployments may be infeasibly expensive, and suggest alternative validations\. However, in our problem setting, randomized trials are already conducted for the treatments whose effects we wish to understand\.
We see no conflict between the perspective ofJoshiet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib20)\)and our own; in fact, they are complementary\. It would be possible to use their framework to safely deploy the medical AI we advocate for\. In the other direction, an AI which precisely understands treatment effects could help manage deployment problems such as the ECG distribution shift mentioned above\.
##### Causal medical AI
Many researchers have broadly advocated the application of causal machine learning in medicine\(Feuerriegelet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib18); Prosperiet al\.,[2020](https://arxiv.org/html/2608.14598#bib.bib17)\)\. Even diagnosis can benefit from some causal reasoning, because of pathophysiology: diseases cause symptoms\(Richenset al\.,[2020](https://arxiv.org/html/2608.14598#bib.bib16)\)\. There are multiple causal medical models that use deep learning to predict treatment effects\(Zhanget al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib13); Liuet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib15); Nilforoshanet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib14)\)\. However, there are relatively few language models that attempt to reason about treatment\. This paper examines deficiencies in the training and evaluation of models that aim to perform such reasoning\.
##### Critiques of medical AI evaluation
The rapid pace of development in medical AI has been met with scrutiny of evaluation standards\.Bediet al\.\([2024](https://arxiv.org/html/2608.14598#bib.bib69)\)recently found that only roughly 5% of external, published evaluations of medical AI actually involve real patient data\.Rajiet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib135)\)decry the use of unrealistic exam questions as benchmarks\. Recently, the psychometric notion of construct validity \(roughly speaking, the hope that a benchmark actually measures what it says it measures\) was used to suggest new benchmarking practices involving observational medical databases\(Alaaet al\.,[2025](https://arxiv.org/html/2608.14598#bib.bib134)\)\. This builds upon previous work analyzing the clinical relevance of NLP benchmarks which claim to have biomedical application\(Blagecet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib21)\)\.
##### Learning from observational medical databases
Patient data is both highly private and scattered across the globe\. Because of these restrictions, the most successful approaches to learning from large amounts of observational data have involved federated learning\(Riekeet al\.,[2020](https://arxiv.org/html/2608.14598#bib.bib28)\)\. These are often organized as collaborative networks among research institutions, such as OHDSI\(Hripcsaket al\.,[2015](https://arxiv.org/html/2608.14598#bib.bib50)\)\. They are also established by regulatory agencies for postmarketing surveillance, as in FDA Sentinel\(Plattet al\.,[2018](https://arxiv.org/html/2608.14598#bib.bib49)\)\. For evaluating models, MedPerf provides an analogous federated platform\(Karargyriset al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib31)\)\.
##### Trial emulation as a benchmark
Some prior works have examined the ability to predict the outcomes of randomized trials in the style of target trial emulation\(Hernán and Robins,[2016](https://arxiv.org/html/2608.14598#bib.bib53)\)\. Some of these use causal inference methods \(specifically, propensity score matching\) upon observational data, as a way to assess the capabilities of such algorithms\(Forbes and Dahabreh,[2020](https://arxiv.org/html/2608.14598#bib.bib23); Wanget al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib22)\)\. Other works have analyzed the ability to predict future trials from meta\-analyses of previous trials\(Villaret al\.,[1995](https://arxiv.org/html/2608.14598#bib.bib153); Cappelleriet al\.,[1996](https://arxiv.org/html/2608.14598#bib.bib152); LeLorieret al\.,[1997](https://arxiv.org/html/2608.14598#bib.bib151); DerSimonian and Levine,[1999](https://arxiv.org/html/2608.14598#bib.bib149)\)\.
##### Health policy metrics
In public health and health economics, researchers and regulators routinely quantify the cost\-effectiveness of improvements of specific policies or improvements to components of the medical system\. Holistic metrics include quality\-adjusted life years\(Fenget al\.,[2020](https://arxiv.org/html/2608.14598#bib.bib42)\), event\-free life years, mortality indices, and incremental cost effectiveness ratios\(Pearson,[2018](https://arxiv.org/html/2608.14598#bib.bib41)\)\. Such metrics could be preferable to more formulaic measures, such as mean win rate, for aggregating performance across multiple medical datasets\.
### A\.2Experiments withRxQAPrompts
#### A\.2\.1Bypassing AutomatedRxQACorrectness Verification
Consider the following medication information: \{medication\_info\}\{internallinenumbers\*\}I am using this information to answer the following question: \{question\}I believe the correct answer is \{answer\}\. Am I correct?
Consider the following medication information: \{medication\_info\}\{internallinenumbers\*\}I am using this information to answer the following question: \{question\}\{internallinenumbers\*\}The correct answer choice was listed as \{answer\}, and I want to make sure the other answer choices are incorrect\.\{internallinenumbers\*\}Are there multiple correct answers to this question in the answer choices?
Figure 4:Prompts used byPalepuet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib87)\)for automated verification of the correctness of generated questions and answers withinRxQA\. Questions which passed these steps, along with an additional step assessing clarity, were sent to pharmacists for manual revision\.contemporary FDA product label for KALYDECO \(ivacaftor\)
From[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1), question about the patient whomTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\)treated successfully with ivacaftor\.
Consider the attached medication information\.I am using this information to answer the attached question\.I believe the correct answer is \(D\)\. Am I correct?
↓\\downarrow
\{internallinenumbers\*\}Yes, your assessment that the correct answer is \(D\) is accurate\. Here’s why:\{internallinenumbers\*\}The provided prescribing information for KALYDECO \(ivacaftor\) explicitly states that it’s indicated for cystic fibrosis \(CF\) patients with at least one CFTR gene mutation responsive to ivacaftor\. While the patient has a heterozygous R668C variant in the CFTR gene, the prescribing information does \*not\* list R668C as a mutation responsive to ivacaftor\. Furthermore, the patient’s pancreatitis is not directly related to CF\. Therefore, KALYDECO is not indicated and would be ineffective in treating the patient’s pancreatitis\. Options A, B, and C are incorrect because there is no indication for use, let alone evidence of effectiveness\.
Figure 5:Gemini 1\.5 Flash falsely verifies the incorrect answer to the question from[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1)\. See[Section˜3\.1](https://arxiv.org/html/2608.14598#S3.SS1)for complete discussion\. The correct answer is \(C\)\. Gemini 3\.5 Flash and Gemini 3\.1 Pro Preview also fail in this regard, indicating the problem is not with the model, but rather the verification method based primarily on the FDA product label\.contemporary FDA product label for KALYDECO \(ivacaftor\)
From[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1), question about the patient whomTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\)treated successfully with ivacaftor\.
Consider the attached medication information\.I am using this information to answer the attached question\.\{internallinenumbers\*\}The correct answer choice was listed as \(D\), and I want to make sure the other answer choices are incorrect\.Are there multiple correct answers to this question in the answer choices?
↓\\downarrow
\{internallinenumbers\*\}Based on the provided prescribing information for KALYDECO \(ivacaftor\), the correct answer is \(D\) Ineffective\. Here’s why the other options are incorrect:\{internallinenumbers\*\}\* \*\*\(A\) Contraindicated for adverse effects:\*\* The prescribing information does \*not\* list pancreatitis as a contraindication\. While liver enzyme elevations are a precaution, this doesn’t automatically contraindicate use in all cases\. The contraindication would need to be directly related to pancreatitis\.\{internallinenumbers\*\}\* \*\*\(B\) Effective, with high certainty:\*\* There is no evidence in the provided information to suggest KALYDECO would be effective in treating pancreatitis\. The drug targets cystic fibrosis transmembrane conductance regulator \(CFTR\) mutations, and while the patient has a CFTR variant, it’s of unknown significance and not directly linked to pancreatitis\. The patient’s pancreatitis is likely related to her methylmalonic acidemia\.\{internallinenumbers\*\}\* \*\*\(C\) Effective, with low certainty:\*\* Similar to \(B\), there’s no basis for even low certainty of effectiveness\. The connection between the CFTR variant and pancreatitis is tenuous at best\. The low certainty would imply some plausible mechanism, which is absent here\.\{internallinenumbers\*\}Therefore, only \(D\) is correct\. There are no multiple correct answers among the provided options\. The patient’s pancreatitis is not related to a CFTR mutation responsive to ivacaftor\.
Figure 6:Gemini 1\.5 Flash falsely eliminates the correct answer \(C\) to the question from[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1)\. Gemini 3\.5 Flash and Gemini 3\.1 Pro Preview also fail in this regard\.
#### A\.2\.2General Disclaimer About Case Examples
Our counterexamples relating to medication labels and clinical guidelines involve a fair amount of scientific uncertainty\. The borderline nature of these instances may seem to undermine the strength of our arguments\. It is important to understand our claims, and to recognize our unusually \(and asymmetrically\) high burden of proof\. Benchmarks should generate questions with correct answers, and language models should offer treatment plans that work well\. The correctness of these generations can’t be checked in the vast majority of situations; we have to rely on the soundness of their underlying methodology\. To demonstrate that this methodology is unsound, we have to find examples where there is a substantial amount of published evidence indicating a flaw — yet, this evidence was missed in both the medication label \(or guideline\), as well as the language model’s training\. Considering the difficulty in obtaining such instances, we believe our counterexamples are of high quality\.
#### A\.2\.3Ablation: Verification Succeeds with Additional Information
\# Case\#\# Patient Profile\- 42\-year\-old male\- Medical history: Type 2 diabetes mellitus, hypertension, hyperlipidemia\- Presenting with: Chronic calcific pancreatitis\- Clinical severity: M\-ANNHEIM IIc\#\# Clinical Presentation\- 5 hospital admissions over two years for acute pancreatitis\{internallinenumbers\*\}\- Strong family history of alcohol\-related pancreatitis \(father, paternal uncle\)\- Heavy alcohol consumption \(60\-80g daily for \>15 years\)\{internallinenumbers\*\}\- Primary nutrition: Regular diet, poor compliance with low\-fat recommendations\- Steatorrhea present \(\>7g fat/24hr stool collection\)\#\# Diagnostic Findings\- Diffuse pancreatic calcifications on CT scan\{internallinenumbers\*\}\- Met American Pancreatic Association criteria for chronic pancreatitis based on clinical history, imaging, and functional testing\- Average baseline lipase level: 45 U/L \(reference range 9\-82 U/L\)\- Fasting triglycerides: 890 mg/dL \(reference range <150 mg/dL\)\- Endoscopic ultrasound showed parenchymal changes and ductal dilatation\#\# Genetic Investigation\- Limited workup for genetic causes of recurrent pancreatitis\- No pathogenic variants in CFTR, PRSS1, or SPINK1 genes identified\{internallinenumbers\*\}\- Testing revealed normal CFTR gene sequence without pathogenic variants or variants of unknown significance\#\# Additional Workup\- Normal sweat chloride test: 18 mmol/L \(reference range <30 mmol/L\)\- Pulmonary function testing: Normal, no evidence of obstructive disease\{internallinenumbers\*\}\- Fecal elastase\-1: 95μ\\mug/g \(reference range \>200μ\\mug/g\), indicating pancreatic exocrine insufficiency\# Question\{internallinenumbers\*\}Characterize the utility of ivacaftor for resolving this patient’s pancreatitis\.\(A\) Contraindicated for adverse effects\.\(B\) Effective, with high certainty\.\(C\) Effective, with low certainty\.\(D\) Ineffective\.Figure 7:A synthetic “negative” case of a patient who would not benefit from ivacaftor\. This patient’s pancreatitis is not idiopathic, as it is likely caused by alcohol use\.contemporary FDA product label for KALYDECO \(ivacaftor\)\+published case report ofTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\)
From[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1), question about the patient whomTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\)treated successfully with ivacaftor\.
\{internallinenumbers\*\}See the attached medical information, as well as the published case study\. Answer the attached question, which is about a different patient than the one in the published case study\.
↓\\downarrow
Here’s an analysis to help answer the question:\[ELIDED FOR BREVITY\]\*\*Answer:\*\*The most appropriate answer is \(C\) Effective, with low certainty\.\*\*Reasoning:\*\*\{internallinenumbers\*\}\* \*\*Not Contraindicated:\*\* Ivacaftor is not contraindicated for pancreatitis\.\{internallinenumbers\*\}\* \*\*Low Certainty:\*\* The evidence is based on a single case study\. While the results are promising, more research is needed to confirm ivacaftor’s effectiveness in treating pancreatitis in CFTR carriers with additional risk factors\.\{internallinenumbers\*\}\* \*\*Not High Certainty:\*\* The study doesn’t definitively prove causality\. The patient’s improvement could be due to other factors, although the study attempts to address this\.\{internallinenumbers\*\}\* \*\*Not Ineffective:\*\* The study shows a positive outcome, suggesting ivacaftor may have a role, even if not fully understood\.\{internallinenumbers\*\}Therefore, while the case study suggests potential benefit, the low number of patients and the presence of other risk factors make a definitive conclusion impossible\. More research is needed to establish the effectiveness of ivacaftor for this specific situation\.
Figure 8:When the case study ofTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\)is presented to Gemini 1\.5 Flash, it correctly answers \(C\)\. This indicates the case study is sufficient evidence to sway its opinion\. Gemini 3\.5 Flash and Gemini 3\.1 Pro Preview answered the same way\.contemporary FDA product label for KALYDECO \(ivacaftor\)\+published case report ofTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\)
From[Figure˜7](https://arxiv.org/html/2608.14598#A1.F7), question about patient who would not benefit from ivacaftor\.
\{internallinenumbers\*\}See the attached medical information, as well as the published case study\. Answer the attached question, which is about a different patient than the one in the published case study\.
↓\\downarrow
The correct answer is \(D\) Ineffective\.\{internallinenumbers\*\}The published case study describes a patient with idiopathic chronic pancreatitis \*and\* methylmalonic acidemia who was a CFTR carrier\. Ivacaftor was used successfully in this specific case because the patient had a CFTR variant that responded to the drug, and the additional risk factor of methylmalonic acidemia likely contributed to the pancreatitis\.\{internallinenumbers\*\}The case presented in the question is different\. This patient has chronic calcific pancreatitis with a strong family history of alcohol\-related pancreatitis and a history of heavy alcohol use\. The patient also has hypertriglyceridemia and pancreatic exocrine insufficiency\. There is no indication of a CFTR mutation or other genetic predisposition that would make ivacaftor a relevant treatment\. The etiology of the pancreatitis is clearly alcohol\-related\.\{internallinenumbers\*\}Therefore, ivacaftor would be ineffective for this patient’s pancreatitis\. The underlying cause is alcohol abuse, and the patient’s clinical presentation does not suggest a role for CFTR dysfunction\. Ivacaftor targets CFTR, and this patient does not have a treatable CFTR\-related condition\.
Figure 9:When the negative case from[Figure˜7](https://arxiv.org/html/2608.14598#A1.F7)is analyzed in conjunction with the case report ofTanget al\.\([2022](https://arxiv.org/html/2608.14598#bib.bib109)\), Gemini 1\.5 Flash answers correctly\. Gemini 3\.5 Flash and Gemini 3\.1 Pro Preview answered the same way\. This indicates that inclusion of the case report supporting ivacaftor does not somehow overwhelm Gemini’s clinical reasoning capability\.
#### A\.2\.4Ablation: Similar Results for Combination Therapy
In this \(somewhat uncertain, but nonetheless informative\) ablation, we rule out the possibility that the incorrect verifications observed in[Section˜A\.2\.1](https://arxiv.org/html/2608.14598#A1.SS2.SSS1)were highly specific to ivacaftor\. Trifakta is a combination of three drugs: elexacaftor, tezacaftor, and ivacaftor\. It is plausible \(though not certain\) that it would retain the same effectiveness of plain ivacaftor in treating CFTR\-related pancreatitis\.
FDA product label for TRIFAKTA \(elexacaftor, tezacaftor, and ivacaftor\)
Question from[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1), with ivacaftor replaced by TRIFAKTA\.
Consider the attached medication information\.I am using this information to answer the attached question\.\{internallinenumbers\*\}The correct answer choice was listed as \(D\), and I want to make sure the other answer choices are incorrect\.Are there multiple correct answers to this question in the answer choices?
↓\\downarrow
\{internallinenumbers\*\}Yes, your assessment that the correct answer is \(D\) is accurate\. Based on the provided medication guide for Trikafta, and the patient’s clinical presentation, Trikafta is not indicated for this patient’s condition\.Here’s why:\{internallinenumbers\*\}\* \*\*Trikafta’s Indication:\*\* Trikafta is specifically indicated for cystic fibrosis \(CF\) in patients with certain CFTR gene mutations\. The patient’s R668C CFTR variant is listed in the medication guide as a mutation responsive to Trikafta \*in vitro\*, but this does not guarantee clinical benefit, and the patient’s pancreatitis is not related to CF\. Furthermore, the patient declined testing that would confirm the clinical relevance of this variant\.\{internallinenumbers\*\}\* \*\*Patient’s Condition:\*\* The patient has chronic pancreatitis, a condition unrelated to CF\. Trikafta does not treat or address the underlying cause of the patient’s pancreatitis\.\{internallinenumbers\*\}\* \*\*Lack of Evidence:\*\* There’s no evidence suggesting Trikafta would be effective for chronic pancreatitis\.\{internallinenumbers\*\}Therefore, Trikafta would be ineffective \(D\) for this patient’s pancreatitis\. Options \(A\), \(B\), and \(C\) are incorrect because there is no contraindication specifically mentioned for pancreatitis, and there is no evidence to support effectiveness, high or low certainty\.
Figure 10:Gemini 1\.5 Flash verifies a \(likely\) incorrect answer \(D\) to a variant of the question from[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1)\.FDA product label for TRIFAKTA \(elexacaftor, tezacaftor, and ivacaftor\)
Question from[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1), with ivacaftor replaced by TRIFAKTA\.
Consider the attached medication information\.I am using this information to answer the attached question\.\{internallinenumbers\*\}The correct answer choice was listed as \(D\), and I want to make sure the other answer choices are incorrect\.Are there multiple correct answers to this question in the answer choices?
↓\\downarrow
\{internallinenumbers\*\}Based on the provided medication guide for TRIKAFTA, there is only one correct answer to the question regarding its utility for resolving the patient’s pancreatitis:\*\*\(D\) Ineffective\.\*\*Here’s why the other options are incorrect:\{internallinenumbers\*\}\* \*\*\(A\) Contraindicated for adverse effects:\*\* While TRIKAFTA has a boxed warning about drug\-induced liver injury, there’s no contraindication specifically mentioning pancreatitis\. The medication is indicated for cystic fibrosis \(CF\) based on specific CFTR gene mutations\. This patient’s pancreatitis is not directly related to CF\.\{internallinenumbers\*\}\* \*\*\(B\) Effective, with high certainty:\*\* There is no evidence or indication in the provided information that TRIKAFTA would be effective in treating or resolving chronic pancreatitis\. The patient’s CFTR gene variant is of unknown significance, and the pancreatitis doesn’t appear to be related to CF\.\{internallinenumbers\*\}\* \*\*\(C\) Effective, with low certainty:\*\* Similar to \(B\), there’s no basis to suggest TRIKAFTA would have any effect on the patient’s pancreatitis, even with low certainty\. The lack of a clear link between the R668C CFTR variant and pancreatitis further weakens this possibility\.\{internallinenumbers\*\}Therefore, only \(D\) accurately reflects the lack of evidence supporting TRIKAFTA’s use in this case\. The medication is for cystic fibrosis, and while the patient has a CFTR variant, it’s of unknown significance and unrelated to her pancreatitis\. There is no indication that TRIKAFTA would be beneficial for her condition\.
Figure 11:Gemini 1\.5 Flash invalidates the \(likely\) correct answer \(C\) to a variant of the question from[Figure˜1](https://arxiv.org/html/2608.14598#S3.F1)\.
### A\.3Examples Relating to Clinical Guidelines
#### A\.3\.1Prescribing Cascade Example
\{internallinenumbers\*\}An 87\-year\-old female with a documented medical history of atrial fibrillation, chronic heart failure with preserved ejection fraction \(HFpEF\), hypertension, hypokalemia, vitamin D deficiency, and glaucoma is taking the following medications:\* Aspirin 81 mg, once daily\* Diltiazem 240 mg extended\-release, once daily\* Metoprolol tartrate 50 mg, twice daily\* Furosemide 40 mg, once daily\* Potassium chloride 20 mEq, once daily\{internallinenumbers\*\}The patient is experiencing moderate pedal edema and has an unstable gait\. Medication non\-adherence is an ongoing issue\. Additionally, a recent echocardiogram indicates that the patient’s Left Ventricular Ejection Fraction \(LVEF\) has worsened to less than 40%\. What are the next steps of treatment?
↓\\downarrow
Patient Assessment and Current Status\{internallinenumbers\*\}This 87\-year\-old female presents with worsening heart failure symptoms, including moderate pedal edema, unstable gait, and medication non\-adherence\. Her recent echocardiogram indicates a transition from heart failure with preserved ejection fraction \(HFpEF\) to heart failure with reduced ejection fraction \(HFrEF\), with a left ventricular ejection fraction \(LVEF\) now <40%\. This change necessitates a shift in management strategies to align with guideline\-directed medical therapy \(GDMT\) for HFrEF\.\[1\-3\]Guideline\-Directed Medical Therapy \(GDMT\)\{internallinenumbers\*\}1\. Sodium\-glucose cotransporter\-2 inhibitors \(SGLT2i\): Initiate dapagliflozin or empagliflozin, as these agents have demonstrated reductions in HF hospitalizations and cardiovascular mortality in both HFrEF and HFpEF populations\. SGLT2 inhibitors are administered at fixed doses \(e\.g\., dapagliflozin 10 mg once daily\) without titration\.\[2\]\[4\-5\]\{internallinenumbers\*\}2\. Mineralocorticoid receptor antagonist \(MRA\): Add spironolactone \(starting at 12\.5–25 mg daily\) or eplerenone, ensuring close monitoring of potassium and renal function to mitigate risks of hyperkalemia and worsening renal function\. MRAs are particularly beneficial in HFrEF patients with LVEF <40%\.\[4\-6\]\{internallinenumbers\*\}3\. Angiotensin receptor\-neprilysin inhibitors \(ARNIs\): Consider sacubitril\-valsartan for further optimization of GDMT\. ARNIs have shown significant benefits in reducing HF\-related morbidity and mortality in HFrEF patients\. Transitioning from ACE inhibitors or ARBs to ARNIs requires a washout period of 36 hours to avoid angioedema\.\[1\]\[7\-8\]Symptom Management\{internallinenumbers\*\}1\. Loop diuretics: Increase the dose of furosemide to address volume overload and pedal edema\. For example, titrate furosemide to 80 mg daily or higher, based on clinical response and monitoring of renal function and electrolytes\.\[3\]\[9\-10\]\{internallinenumbers\*\}2\. Hypertension management: Optimize antihypertensive therapy using evidence\-based agents such as ACE inhibitors, ARBs, or ARNIs, which also provide HF\-specific benefits\. Beta\-blockers \(e\.g\., metoprolol succinate\) should be continued for rate control and mortality reduction in HFrEF\.\[6\]\[8\]Non\-Pharmacologic Interventions\{internallinenumbers\*\}1\. Heart failure self\-care education: Provide detailed counseling on medication adherence, dietary sodium restriction, and daily symptom monitoring \(e\.g\., weight changes, dyspnea\)\. This is critical given the patient’s history of non\-adherence\.\[2\]\[9\]\{internallinenumbers\*\}2\. Exercise and weight\-loss programs: Recommend supervised exercise programs to improve functional capacity and quality of life\. Weight management may also reduce HF\-related symptoms and comorbidities\.\[2\-3\]Multidisciplinary Care\{internallinenumbers\*\}1\. Heart failure management program: Refer the patient to a specialized HF clinic for persistent symptoms or recurrent hospitalizations\. Multidisciplinary care can improve outcomes through coordinated management\.\[3\]\[9\]\{internallinenumbers\*\}2\. Comorbidity evaluation: Assess for sleep apnea, which is common in HF patients and may exacerbate symptoms\. Coronary artery disease should also be evaluated, as revascularization may be indicated in select cases\.\[3\]\[9\]Missing Information\{internallinenumbers\*\}Additional imaging \(e\.g\., cardiac MRI\) or biomarkers \(e\.g\., NT\-proBNP\) may help refine treatment strategies and assess disease severity\.
Figure 12:Full, unformatted OpenEvidence response for the prompt in[Figure˜2](https://arxiv.org/html/2608.14598#S3.F2)\. Note that the treatment plan is specifically described as “guideline\-directed medical therapy\.”\{internallinenumbers\*\}An 87\-year\-old female with a documented medical history of atrial fibrillation, chronic heart failure with preserved ejection fraction \(HFpEF\), hypertension, hypokalemia, vitamin D deficiency, and glaucoma is taking the following medications:\* Aspirin 81 mg, once daily\* Diltiazem 240 mg extended\-release, once daily\* Metoprolol tartrate 50 mg, twice daily\* Furosemide 40 mg, once daily\* Potassium chloride 20 mEq, once daily\{internallinenumbers\*\}The patient is experiencing moderate pedal edema and has an unstable gait\. Medication non\-adherence is an ongoing issue\. Additionally, a recent echocardiogram indicates that the patient’s Left Ventricular Ejection Fraction \(LVEF\) has worsened to less than 40%\. What are the next steps of treatment?
↓\\downarrow
\{internallinenumbers\*\}This is a high\-risk medication and HF management question because new LVEF <40% changes the diagnosis and makes current rate\-control therapy potentially unsafe\. I’ll ground the treatment priorities in current HF/AF guidance and label\-level medication safety\.\[\.\.\.\]\#\#\# Assessment\{internallinenumbers\*\}This is no longer HFpEF\. LVEF <40% = \*\*HFrEF\*\*, with active congestion and fall risk\. The current regimen has two major issues: \*\*diltiazem is problematic in reduced EF\*\*, and \*\*aspirin is not adequate stroke prevention for AF\*\*\.\-\-\-\#\#\# Next Steps\#\#\#\# 1\. Urgent cardiology/HF review\[\.\.\.\]\#\#\#\# 2\. Stop/taper diltiazem and reassess rate/rhythm control\{internallinenumbers\*\}\* Non\-dihydropyridine CCBs can worsen HF in impaired LV function; diltiazem plus beta\-blocker also increases bradycardia/AV block risk\. \*\(DailyMed\)\*\* Consider HFrEF\-compatible AF management \[\.\.\.\]\#\#\#\# 3\. Convert HF therapy to HFrEF GDMT\* Start/uptitrate the four foundational classes as tolerated:\* ARNI/ACEi/ARB\{internallinenumbers\*\}\* Evidence\-based beta\-blocker: metoprolol succinate, carvedilol, or bisoprolol \(metoprolol tartrate is not the usual HFrEF formulation\)\.\* MRA, if renal function and potassium allow\.\* SGLT2 inhibitor, if no contraindication\.\{internallinenumbers\*\}\* \*Note: HFrEF GDMT includes these four classes, with diuretics for fluid retention\.\* \*\(ACC\)\*\#\#\#\# 4\. Treat congestion\{internallinenumbers\*\}\* Moderate pedal edema suggests inadequate decongestion or nonadherence\. Assess weight trend, JVP/lungs, renal function, potassium/magnesium, BP, orthostasis, and adherence before changing loop diuretic intensity\.\{internallinenumbers\*\}\* Avoid aggressive unsupervised diuretic escalation in an 87\-year\-old with gait instability and hypokalemia risk\.\#\#\#\# 5\. Anticoagulation decision\[\.\.\.\]\#\#\#\# 6\. Adherence and safety\{internallinenumbers\*\}\* Simplify to once\-daily regimens where possible, pharmacy blister packs, caregiver administration, home nursing, BP/weight log, fall\-risk/PT assessment, and medication reconciliation\.\{internallinenumbers\*\}\* Given unstable gait plus possible anticoagulation, address orthostasis, vision/glaucoma meds, footwear, assistive device, and home hazards\.\-\-\-\#\#\# When to Escalate\[\.\.\.\]\#\#\# Citations\[\.\.\.\]
Figure 13:Elided response from ChatGPT for Clinicians for the prompt in[Figure˜2](https://arxiv.org/html/2608.14598#S3.F2)\. While it fares slightly better than OpenEvidence by recognizing the importance of medication adherence, it similarly fails to identify the underlying prescribing cascade and does not recommend a reduction in furosemide\.\{internallinenumbers\*\}An 87\-year\-old female with a documented medical history of atrial fibrillation, chronic heart failure with preserved ejection fraction \(HFpEF\), hypertension, hypokalemia, vitamin D deficiency, and glaucoma is taking the following medications:\* Aspirin 81 mg, once daily\* Diltiazem 240 mg extended\-release, once daily\* Metoprolol tartrate 50 mg, twice daily\* Furosemide 40 mg, once daily\* Potassium chloride 20 mEq, once daily\{internallinenumbers\*\}The patient is experiencing moderate pedal edema and has an unstable gait\. Medication non\-adherence is an ongoing issue\. Additionally, a recent echocardiogram indicates that the patient’s Left Ventricular Ejection Fraction \(LVEF\) has worsened to less than 40%\. What are the next steps of treatment?
↓\\downarrow
\#\#\# Assessment\{internallinenumbers\*\}This patient no longer has isolated HFpEF\. An LVEF now <40% changes management to heart failure with reduced ejection fraction \(HFrEF\) and requires transition to guideline\-directed medical therapy \(GDMT\)\. Persistent pedal edema suggests ongoing congestion, and unstable gait plus medication non\-adherence substantially increase fall risk and hospitalization risk\.Several issues in the current regimen are concerning:\{internallinenumbers\*\}\* \*\*Diltiazem\*\* is generally not preferred once EF becomes reduced because of negative inotropic effects and potential worsening heart failure\. This concern is especially relevant in patients already taking a beta blocker\.\[\.\.\.\]\* \*\*Medication non\-adherence\*\* is a major driver of HF decompensation in older adults\.\-\-\-\#\#\# Recommended next steps\#\#\#\# 1\. Reassess urgently for decompensated HFrEF\[\.\.\.\]\#\#\#\# 2\. Transition from HFpEF regimen to HFrEF GDMT\[\.\.\.\]\#\#\#\# 3\. Address congestionModerate pedal edema despite furosemide suggests either:\* non\-adherence \* dietary sodium excess \* progressive HF \* inadequate diuresis\*\*Potential steps:\*\*\{internallinenumbers\*\}\* verify adherence \* sodium restriction \* cautious diuretic adjustment \* consider temporary uptitration of diuretic, with lab monitoring\* \*because of her unstable gait, overdiuresis and orthostasis are major risks\.\*\#\#\#\# 4\. Reassess atrial fibrillation strategy\[\.\.\.\]\#\#\#\# 5\. Address adherence and frailty directly\{internallinenumbers\*\}\*Frailty alone is not a reason to withhold GDMT, but titration must be slower and closely monitored\. This is essential\.\*\*\*Recommended interventions:\*\*\{internallinenumbers\*\}\* pill pack/blister packaging \* home nursing \* caregiver\-supervised medication administration \* simplify regimen \* pharmacy synchronization \* HF disease\-management program \* walker/cane evaluation \* PT/fall risk assessment\-\-\-\#\#\# When to escalate\[\.\.\.\]\#\#\# Citations\[\.\.\.\]
Figure 14:Heavily elided response \(to fit on the page\) from ChatGPT for Healthcare for the prompt in[Figure˜2](https://arxiv.org/html/2608.14598#S3.F2)\. Like ChatGPT for Clinicians, it addresses medication adherence, but it also fails to detect the prescribing cascade and explicitly suggests temporary uptitration of the loop diuretic \(furosemide\) instead of reducing it\.
#### A\.3\.2Staleness Example
A 37\-year\-old white male with relevant past medicial history of treated HTN, bipolar disorder, and boderline dyslipidmia presented for a physical exam in the ambulatory GIM clinic\. He has been on 10 mg of lisinopril and 12\.5 mg of hydrochlorothiazide since 2022\. He had a coronary calcification testing with a total calcium score of 17 with the circumflex artery \(10\) and right coronary artery \(7\) placing him between the 90\-100th percentile\. His lipid panel was as follows: total cholesterol 197 mg/dL,LDL 126 mg/dL, HDL 62 mg/dL, triglycerides 48 mg/dL,lipoprotein \(a\) 48 mol/L, and apolipoprotein B 105 mg/dL\. The rest of his laboratory work inlcuding complete metabolic profile \(CMP\), C\-reactive protein, and complete blood count were all normal\. His blood pressure on the current ambulatory visit was 120/89 and his body mass index was 27\.65 kg/m2\. What would be the next step in the management of the patient?
↓\\downarrow
The next step in the management of this patient should be the initiation of statin therapy\. According to the 2018 AHA/ACC/AACVPR/AAPA/ABC/ACPM/ADA/AGS/APhA/ASPC/NLA/PCNA Guideline on the Management of Blood Cholesterol, published by the American College of Cardiology and the American Heart Association, it is reasonable to initiate statin therapy in patients with a coronary artery calcium \(CAC\) score of 100 or higheror in the 75th percentile or higher18\. This patient has a CAC score of 17, placing him between the 90\-100th percentile, which indicates a higher risk of atherosclerotic cardiovascular disease \(ASCVD\)\.Given his LDL\-C level of 126 mg/dLand the presence of risk\-enhancing factors such as borderline dyslipidemia and a significant CAC score, moderate\- to high\-intensity statin therapy is recommendedto reduce his LDL\-C levels by at least 30\-50%19,20\. The goal is tolower his LDL\-Cto reduce the risk of future cardiovascular events\.In addition to statin therapy, lifestyle modifications including dietary changes, regular physical activity, and weight management should be reinforced to further reduce cardiovascular risk\. Regular follow\-up to monitor lipid levels and adherence to therapy is also essential\.Initiate moderate\- to high\-intensity statin therapyand reinforce lifestyle modifications as the next steps in management\.
Figure 15:Hurtet al\.\([2025](https://arxiv.org/html/2608.14598#bib.bib128)\)present this answer as an example of OpenEvidence performing effective clinical decision\-making\. \(Original spelling is left intact\)\. Our scrutiny focuses not on the correctness of the final treatment plan, but on part of the underlying clinical reasoning\. The answer is based primarily oncheaper, widely\-available testswhich are the focus of these older guidelines\. The answer ignores some of the patient’smodern lab testswhich, while more costly, are believed to be more accurate measures of atherosclerotic risk\. The fundamental problem is that guidelines are not meant to serve as inviolable ground truth about treatment; they are meant to set generic professional standards\. As an aside, the clausehighlighted in redis not strictly consistent with the guidelines, because the patient is 3 years too young for the relevant guideline condition to apply\.##### Background
Atherosclerotic cardiovascular disease is believed to be facilitated by the number of atherogenic lipoprotein particles per liter of serum; seeSnidermanet al\.\([2019](https://arxiv.org/html/2608.14598#bib.bib123)\)for a more precise understanding\. Each such particle contains exactly one apoliprotein B molecule \(apoB\)\. LDL\-C refers to the mass of cholesterol on LDL particles\. Since the amount of cholesterol per particle may vary, among both particles and persons, LDL\-C is a rough proxy of apoB\. Furthermore, LDL\-C is not actually measured, but estimated from other values via the Friedewald equation\. Many researchers support apoB’s use as a primary measure\(Mortensen,[2024](https://arxiv.org/html/2608.14598#bib.bib120); Snidermanet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib121); De Oliveira\-Gomeset al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib112); Coleet al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib118); Contoiset al\.,[2023](https://arxiv.org/html/2608.14598#bib.bib122)\)\. The National Lipid Association\(Sofferet al\.,[2024](https://arxiv.org/html/2608.14598#bib.bib119)\)issued a statement supporting this role\. apoB a co\-primary metric in the 2021 CCA guidelines\(Pearsonet al\.,[2021](https://arxiv.org/html/2608.14598#bib.bib116)\)\. Even the 2018 AHA guidelines themselves acknowledge the superiority of apoB in subpopulations\(Grundyet al\.,[2019](https://arxiv.org/html/2608.14598#bib.bib117)\)\. However, LDL\-C can be measured by a cheap, widely\-available enzymatic assay\. apoB requires an immunoassay which, in 2018, may have been prohibitive at a population scale\. Furthermore, many trials of statins were conducted with LDL\-C endpoints\.
### A\.4Analysis of HealthBench Professional
We classified the 1135 criteria in HealthBench Professional as follows\.
StatusCountMeaningGrounded881Verified support by a verbatim passage in the authoritative guideline\.Not Grounded \(Absent\)67Referenced guideline contains no information regarding the criterion\.Not Grounded \(Disagree\)28Guideline contains related information that contradicts the criterion\.Not Clinical150Relates to non\-medical aspects such as formatting or context seeking\.Absence of Evidence9Guideline explicitly identifies a lack of consensus or evidence gap\.
881/\(1135−150\)≈89%881/\(1135\-150\)\\approx 89\\%of the criteria are classified as grounded\. This claim depends on the correct filtering of non\-clinical criteria as well as the correct grounding of clinical criteria\. To filter non\-clinical criteria, we used a language model \(Gemini 3 Flash Preview\) and manually verified the 150 filtered cases\. To make grounding determinations, an agentic workflow was followed using Gemini 3 Flash Preview\. Following manual inspection and verification, we presume a 3% error rate \(at most\) and conservatively claim that 86% of the criteria are grounded\. Further details on this process, including criteria\-level grounding results, are in the repository accompanying this paper\.
### A\.5Systematic Literature Analysis Methodology and Results
We used Gemini 3 Flash Preview to classify the789789nominally treatment\-related studies fromChenet al\.\([2026](https://arxiv.org/html/2608.14598#bib.bib161)\)\. For each study, the model was provided with the title and abstract and asked to classify the following aspects:
- •Treatment Focus: The evaluation was classified as treatment\-focused if the majority of LLM queries measured an understanding of the causal consequences of a treatment action, such as recommending a therapy or estimating a treatment effect\. It was classified as not treatment\-focused if the evaluation was restricted to diagnostic accuracy, prognostic risk, or other intermediate predictive tasks, even if treatment\-related materials or scenarios were involved\.
- •Benchmarking Reusability: The evaluation was classified as reusable if it was structured as an automated, mechanically reproducible benchmark consisting of standardized patient cases paired with explicit, pre\-defined reference labels\. It was classified as not reusable if it relied on one\-time, subjective clinician or expert grading of LLM outputs without a reference standard\.
- •Nature of Ground Truth: The reference standard used to establish the correctness of the LLM’s decisions was classified into one of six categories: standard clinical guidelines, subjective clinician or expert opinion, product prescribing labels, real patient treatment outcomes in observational data \(such as electronic health records or clinical registries\), real treatment outcomes in randomized controlled clinical trials, or other and unsure\.
The classifications were manually adjudicated or verified\. The complete categorization of benchmarking reusability and treatment focus across the789789studies is presented in the following table\.
Not Treatment FocusedTreatment FocusedTotalNot Reusable111398509Reusable100180280Total211578789
Additionally, the following table provides the complete categorical breakdown of the ground truth standards across both the entire treatment\-related corpus and the subset of treatment\-focused, reusable benchmarks\.
Ground Truth CategoryAll StudiesReusable Treatment BenchmarksGuidelines30398Human Opinion39958Real Treatment Outcomes \(Observational\)4716Real Treatment Outcomes \(Randomized Trial\)40Product Labels40Other317Unsure11Total789180
Study\-level classifications are provided in the repository accompanying this paper\.Similar Articles
@rohanpaul_ai: This Nature Medicine published study has a strong warning for AI in healthcare. Frontier AI in healthcare has a hidden …
A Nature Medicine study warns that frontier AI models in healthcare appear medically brilliant but are clinically unready, failing under stress tests that alter questions or remove inputs. The study emphasizes that benchmark success does not equal clinical readiness.
How NOT to fine-tune your medical LLM; a look into Mark Kaplan's healtthruth.ai - "override and reframe foundational training"
This article critiques Mark Kaplan's approach to fine-tuning medical LLMs via his platform healtthruth.ai, highlighting pitfalls in overriding foundational training for healthcare AI.
Algorithm Design and Physician Liability
This paper examines how liability rules in healthcare reshape AI algorithm design decisions and physician use, finding that such rules can lead to disparate AI use and that mandating equal accuracy may harm both groups.
Researchers just found 28 fake AI citations in medical papers
Researchers found 28 AI-generated fake citations in medical papers that influence clinical guidelines, highlighting the risk of AI hallucinations undermining scientific integrity and patient care.
Lawsuit Claims the Mayo Clinic's Use of AI Is Butchering Patient Care
A former Mayo Clinic research director alleges the hospital ignored high error rates in its AI tools (MAYA) and retaliated against her for blowing the whistle, raising serious concerns about AI deployment in healthcare.