Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset
Summary
This feasibility study compares a standalone LLM with a pre-specified agentic pipeline for explaining ICU mortality predictions, finding that the agentic approach improves guideline grounding and patient-specific detail but requires attribution-based checks for safety.
View Cached Full Text
Cached at: 08/28/26, 09:27 AM
# Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset Source: [https://arxiv.org/html/2608.26109](https://arxiv.org/html/2608.26109) Chen Xie1,bHaoyun ZhangcZihan WeidZiwei Wang\*,dJiazhao ShieZiyu WangfQiyang Xieg \(aSanta Clara University, Santa Clara, United States bUniversity of Massachusetts Amherst, Amherst, United States cUniversity of Pennsylvania, Philadelphia, United States dCarnegie Mellon University, Pittsburgh, United States eNew York University, Brooklyn, United States fWake Forest University, Winston\-Salem, United States gNortheastern University, Boston, United States 1These authors contributed equally\. \*Corresponding author: Ziwei Wang\.\) ###### Abstract Machine\-learning models can predict ICU mortality accurately, but feature\-attribution methods alone rarely provide the clinical narrative needed for bedside use\. Large language models \(LLMs\) may bridge this gap, and multi\-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation\. This revised feasibility study preserves the original standalone\-versus\-agentic comparison while making the main clinical findings more explicit\. Using the retained local eICU Demo artifact set \(2,353 ICU stays; 8\.1% mortality\), XGBoost achieved an AUROC of 0\.855 \(95% CI 0\.796–0\.906\) and an AUPRC of 0\.332 \(95% CI 0\.217–0\.494\)\. On a stratified 38\-case explanation subset, the standalone LLM produced 1 explanation with explicit outcome leakage, whereas the four\-step agentic pipeline produced none\. Among the 14 cases that overlapped with the SHAP review subset, the standalone LLM showed higher SHAP alignment \(mean Jaccard 0\.171 versus 0\.077\) and higher direction consistency \(92\.9% versus 78\.6%\), while the agentic pipeline showed higher guideline grounding \(0\.762 versus 0\.143\), higher value specificity \(0\.236 versus 0\.143\), and slightly higher plausibility \(0\.700 versus 0\.671\)\. Clinically, the results suggest that agentic decomposition may improve safety\-relevant grounding and patient\-specific detail, but it should be paired with attribution\-based checks before use in high\-stakes risk explanation\. ## 1Introduction Mortality prediction in the intensive care unit \(ICU\) remains a central benchmark for critical\-care risk modeling\. Modern machine\-learning models often exceed the discrimination of classical severity scores when applied to structured electronic health record data\[[1](https://arxiv.org/html/2608.26109#bib.bib1),[2](https://arxiv.org/html/2608.26109#bib.bib2)\]\. The practical barrier is not purely predictive performance; it is whether clinicians can understand and audit why a model has assigned high risk to a particular patient\[[3](https://arxiv.org/html/2608.26109#bib.bib3)\]\. Post\-hoc explainability methods such as SHAP are valuable because they identify the features that move a model prediction\[[4](https://arxiv.org/html/2608.26109#bib.bib4),[10](https://arxiv.org/html/2608.26109#bib.bib10)\]\. However, feature attribution is not equivalent to clinical explanation\. A ranked list of variables, even when mathematically faithful, does not by itself produce the pathophysiological narrative that clinicians use for verification, triage, and communication\[[12](https://arxiv.org/html/2608.26109#bib.bib12)\]\. This gap has motivated interest in large language models \(LLMs\), which can transform structured values into concise natural\-language reasoning grounded in clinical concepts\[[5](https://arxiv.org/html/2608.26109#bib.bib5),[6](https://arxiv.org/html/2608.26109#bib.bib6),[7](https://arxiv.org/html/2608.26109#bib.bib7)\]\. A single LLM prompt is an attractive baseline, but it may collapse multiple reasoning steps into one opaque response\. Agentic systems offer an alternative by decomposing the task into data interpretation, application of formal criteria, differential construction, and synthesis\[[14](https://arxiv.org/html/2608.26109#bib.bib14),[15](https://arxiv.org/html/2608.26109#bib.bib15)\]\. That decomposition is especially appealing in critical care, where thresholds, syndromic criteria, and multiple interacting organ systems matter\. Our original study design therefore asked three research questions: 1. RQ1How accurately can standard machine\-learning models predict ICU mortality from structured first\-24\-hour data? 2. RQ2Can a standalone LLM produce clinically plausible explanations that retain some alignment with SHAP attributions? 3. RQ3Does a structured agentic pipeline produce higher\-quality, more clinically grounded explanations than a standalone LLM? This revised manuscript preserves the same questions and underlying idea, but it reports only claims supported by the auditable local artifact set\. Specifically, RQ1 is addressed quantitatively on the held\-out cohort, whereas RQ2 and RQ3 are addressed on an audited explanation subset generated from the same versioned test snapshot\. ## 2Results ### 2\.1Cohort characteristics and predictive performance The study cohort contained 2,353 adult ICU stays with an in\-hospital mortality rate of 8\.1% \(n=191n=191\)\. Non\-survivors were older and showed higher heart rate, respiratory rate, lactate, and blood urea nitrogen values than survivors\. Baseline descriptive statistics are shown in Table[1](https://arxiv.org/html/2608.26109#S2.T1)\. Table 1:Baseline characteristics of the study cohortXGBoost outperformed logistic regression on the held\-out test set, although both models showed the typical precision\-recall constraints expected under low event prevalence\. XGBoost achieved an AUROC of 0\.855 \(95% CI 0\.796–0\.906\) and an AUPRC of 0\.332 \(95% CI 0\.217–0\.494\), whereas logistic regression achieved an AUROC of 0\.823 \(95% CI 0\.752–0\.886\) and an AUPRC of 0\.345 \(95% CI 0\.218–0\.506\) \(Table[2](https://arxiv.org/html/2608.26109#S2.T2); Figure[1](https://arxiv.org/html/2608.26109#S2.F1)\)\. These discrimination estimates are consistent with prior ICU mortality modeling studies based on structured EHR data\[[1](https://arxiv.org/html/2608.26109#bib.bib1),[2](https://arxiv.org/html/2608.26109#bib.bib2)\]\. Table 2:Bootstrapped discrimination metrics on the held\-out test set\.Figure 1:Receiver operating characteristic and precision\-recall curves for logistic regression and XGBoost on the held\-out test set\. ### 2\.2Global feature attribution with SHAP The SHAP summary remained clinically coherent\. Age, minimum SpO2, blood urea nitrogen, lactate, and respiratory rate were the most influential features in the retained XGBoost model, matching well\-established ICU mortality correlates such as advanced age, hypoxemia, renal dysfunction, and respiratory compromise\[[11](https://arxiv.org/html/2608.26109#bib.bib11)\]\. The summary plot is shown in Figure[2](https://arxiv.org/html/2608.26109#S2.F2)\. Figure 2:SHAP summary plot for the XGBoost model\. Each point represents one patient, and color denotes feature value\. ### 2\.3Standalone LLM audit and retained explanation quality The regenerated standalone explanation set contained 38 outputs\. A direct audit of explanation text found 1 explanation with explicit outcome leakage terms and removed it from the valid comparison set\. After this audit, 37 explanations remained, and 14 overlapped with the pre\-existing SHAP\-reviewed patient subset and could therefore be evaluated against feature attribution\. Within these 14 leakage\-free overlapping cases, explanation quality remained mixed \(Table[3](https://arxiv.org/html/2608.26109#S2.T3); Figure[3](https://arxiv.org/html/2608.26109#S2.F3)\)\. Mean SHAP alignment was 0\.171 \(95% CI 0\.075–0\.279\), and 50\.0% of cases showed any feature overlap at all\. Explanation plausibility was moderate \(mean 0\.671, 95% CI 0\.571–0\.764\), while direction consistency with the model\-predicted risk label remained high at 92\.9% \(95% CI 78\.6–100\.0\)\. Value specificity was low \(mean 0\.143, 95% CI 0\.029–0\.300\), and guideline grounding was limited \(mean 0\.143, 95% CI 0\.071–0\.238\)\. Taken together, these results suggest that the standalone LLM can produce concise and directionally coherent narratives, but those narratives remain only modestly aligned with the model’s top SHAP features\. Table 3:Outcome\-leakage audit and retained standalone explanation quality metrics\.Figure 3:Audited standalone explanation quality metrics with bootstrap confidence intervals\. Direction consistency is normalized to the 0–1 scale for display\. ### 2\.4Agentic pipeline comparison The same four\-step agentic design from the original manuscript was preserved in this refit because it is central to RQ3\. The pipeline comprises a clinical data interpreter, a guideline consultant, a differential reasoner, and a final synthesizer\. Each step was rewritten in the versioned scripts to avoid outcome leakage and to consume cleaned prompt\-facing values only\. The architecture is shown in Figure[4](https://arxiv.org/html/2608.26109#S2.F4)\. The regenerated agent run also produced 38 outputs, and none were flagged for explicit outcome leakage\. Fourteen cases overlapped with the SHAP\-reviewed subset and were therefore available for direct comparison with the regenerated standalone baseline \(Table[4](https://arxiv.org/html/2608.26109#S2.T4); Figure[5](https://arxiv.org/html/2608.26109#S2.F5)\)\. The comparative pattern was not monotonic\. The agentic pipeline had lower SHAP alignment than the standalone baseline \(mean Jaccard 0\.077, 95% CI 0\.018–0\.143, versus 0\.171, 95% CI 0\.075–0\.279\) and lower direction consistency \(78\.6% versus 92\.9%\)\. In contrast, it showed markedly stronger guideline grounding \(0\.762, 95% CI 0\.571–0\.929, versus 0\.143, 95% CI 0\.071–0\.238\), higher value specificity \(0\.236, 95% CI 0\.093–0\.407, versus 0\.143, 95% CI 0\.029–0\.300\), more frequent mention of patient\-specific data \(85\.7% versus 64\.3%\), and slightly higher plausibility \(0\.700, 95% CI 0\.557–0\.843, versus 0\.671, 95% CI 0\.571–0\.764\)\. These results suggest that task decomposition made the explanations more explicit about clinical criteria and supporting evidence, but did not move them closer to the dominant SHAP features of the underlying mortality model\. Table 4:Audited explanation\-quality comparison between the standalone LLM and the pre\-specified agentic pipeline\.Step 1: Clinical data interpreterInput:cleaned patient values plus reference rangesOutput:abnormality summary with patient\-specific values and clinical significance↓\\downarrowStep 2: Guideline consultantInput:patient data plus Step 1 summaryOutput:structured application of SIRS, SOFA, qSOFA, and KDIGO\-style criteria↓\\downarrowStep 3: Differential reasonerInput:patient data plus Steps 1–2 outputsOutput:ranked mortality\-driving conditions with evidence and mechanisms↓\\downarrowStep 4: Final synthesizerInput:patient data plus Steps 1–3 outputsOutput:structured JSON explanation in the same schema as the standalone baselineFigure 4:Pre\-specified outcome\-free agentic pipeline retained for RQ3\.Figure 5:Audited explanation\-quality comparison between the standalone LLM and the pre\-specified agentic pipeline\. Direction consistency is normalized to the 0–1 scale for display\. ## 3Discussion The revised analysis prioritizes three clinically meaningful findings\. First, the structured mortality model was sufficiently discriminative to justify explanation work, and its most influential features—age, oxygenation, blood urea nitrogen, lactate, and respiratory rate—are recognizable markers of physiologic instability in critical care\. Second, the standalone LLM usually produced a concise risk narrative in the correct direction, but the explanations were often weak in patient\-specific values and only modestly aligned with SHAP attributions\. Third, the pre\-specified agentic pipeline reduced explicit leakage and produced more guideline\-grounded, patient\-specific explanations, but this improvement came with lower SHAP alignment and lower direction consistency\. These findings support the original motivation while clarifying the practical tradeoff\. Feature attribution and natural\-language explanation can be complementary: SHAP helps identify what drove the model, whereas an LLM can translate abnormal values into a readable clinical narrative\. However, an agentic pipeline should not be assumed to be superior simply because it is more structured\. In this study, decomposition improved guideline use and bedside readability, but it did not automatically recover the features that most strongly drove the XGBoost prediction\. A clinically useful system may therefore need both components: an attribution check for model fidelity and a guideline\-grounded language layer for interpretability\. The revised analysis also highlights data\-quality issues that directly affect explanation reliability\. In the prompt\-facing test snapshot, mean arterial pressure was missing in 84\.7% of rows after invalid values below 20 mmHg were excluded, temperature was missing in 93\.4% of rows after one Fahrenheit\-style outlier was harmonized to Celsius, lactate was missing in 81\.5% of rows, and reconstructed GCS was unavailable in 13\.2% of rows\. These values are clinically important, so missingness should be surfaced rather than hidden; otherwise, an LLM may sound more certain than the available evidence supports\. Several limitations remain\. The eICU Demo dataset is small relative to full\-scale ICU cohorts\. Only 14 generated cases overlapped with the retained SHAP\-reviewed subset, so the confidence intervals around the comparative explanation metrics remain wide\. Both generators were run with a single local base model, so the observed differences reflect one particular implementation of standalone prompting and one particular implementation of task decomposition rather than an architecture\-independent truth\. The attribution comparison also depends on heuristic mapping from narrative factors to structured model features\. Finally, automated explanation metrics remain proxies for clinician judgment, not substitutes for prospective evaluation by critical\-care experts\. Recent work on retrieval\-grounded evaluation for conversational LLM\-based risk assessment similarly emphasizes that risk\-oriented LLM outputs should be judged by evidence grounding and task\-relevant reference information, rather than fluency or plausibility alone\[[16](https://arxiv.org/html/2608.26109#bib.bib16)\]\. Despite those limitations, the revised manuscript strengthens the original conference draft by making the evidence more transparent\. It adds an auditable head\-to\-head comparison, makes leakage explicit, separates prompt\-cleaning rules from downstream interpretation, and presents agentic decomposition as a measurable tradeoff rather than a categorical gain\. ## 4Methods ### 4\.1Data source and cohort We used the eICU Collaborative Research Database Demo v2\.0\.1\[[9](https://arxiv.org/html/2608.26109#bib.bib9)\]\. The retained cohort definition from the original study included adults aged at least 18 years, ICU length of stay of at least 4 hours, and non\-missing hospital discharge status\. The final cohort contained 2,353 ICU stays, and the primary outcome was in\-hospital mortality\. ### 4\.2Structured features and prompt\-facing cleaning Features were derived from the first 24 hours and covered demographics, vital signs, laboratory results, and APACHE\-related neurological variables\. The versioned refit preserved the original model inputs but added prompt\-facing cleaning rules for the explanation pipeline\. GCS was reconstructed only from valid APACHE eye, motor, and verbal components; negative sentinels were not treated as clinical values\. Temperature values above 45 were treated as Fahrenheit and converted to Celsius\. Mean arterial pressure values below 20 mmHg and respiratory\-rate values below 5 per minute were excluded from prompt\-facing summaries\. Missingness after these cleaning steps was tabulated in the versioned results tree\. ### 4\.3Mortality modeling The original predictive models were retained: L2\-regularized logistic regression and XGBoost\[[8](https://arxiv.org/html/2608.26109#bib.bib8)\]\. The scirep\_v1 refit added bootstrap 95% confidence intervals for AUROC and AUPRC on the held\-out test set by resampling the 471 test encounters with replacement\. ### 4\.4SHAP attribution The retained SHAP artifact set from the original study was reused to summarize global feature importance and to provide a per\-patient explanation reference subset\. SHAP top features for each reviewed patient were defined as the three largest absolute attributions\. ### 4\.5Standalone and agentic explanation designs The standalone baseline was defined as a single outcome\-free prompt that received cleaned first\-24\-hour patient data and the model\-predicted mortality probability, and returned a structured JSON explanation\. The agentic pipeline used the same cleaned patient representation but decomposed the task into four serial steps: data interpretation, explicit guideline application, differential reasoning, and synthesis\. For explanation generation, we selected a stratified subset of 38 held\-out cases spanning low\-, intermediate\-, and high\-risk predictions and ran both generators on the same cases using the same localollama\-servedllama3\.2:3bmodel\. Both versioned generator scripts wrote toresults/scirep\_v1/so reruns would not overwrite the original retained artifacts\. ### 4\.6Leakage audit and explanation scoring The scirep\_v1 audit searched explanation text for explicit outcome or survival language\. Any explanation containing terms such as “actual outcome”, “survived”, “survival”, or “expired” was flagged and excluded from valid explanation\-quality scoring\. Retained explanations were evaluated on SHAP alignment \(Jaccard overlap with top\-3 SHAP features\), clinical plausibility, direction consistency, value specificity, guideline grounding, and reasoning depth\. Head\-to\-head comparison between standalone and agentic outputs was restricted to the overlap between the generated subset and the retained SHAP\-reviewed patient subset\. Confidence intervals for mean quality metrics were estimated by non\-parametric bootstrap over the retained overlapping cases\. ## Data availability The data supporting the findings of this study are available from the corresponding author upon reasonable request\. The study also uses the publicly available eICU Collaborative Research Database Demo\[[9](https://arxiv.org/html/2608.26109#bib.bib9)\]\. ## Code availability The code used for the analyses in this study is available from the corresponding author upon reasonable request\. ## Author contributions Di Zhu and Chen Xie contributed equally to the review and analysis of the literature and to revision of the manuscript\. Ziwei Wang served as the corresponding author\. Chen Xie prepared the revised manuscript files\. All authors approved the final version of this draft\. ## Competing interests The authors declare no competing interests\. ## Acknowledgements This research used the eICU Collaborative Research Database Demo, made available by Philips Healthcare and the MIT Laboratory for Computational Physiology\. ## References - \[1\]A\. E\. W\. Johnson, T\. J\. Pollard, and R\. G\. Mark, “Reproducibility in critical care: a mortality prediction case study,” in*Proc\. Mach\. Learn\. Healthcare*, 2017\. - \[2\]H\. Harutyunyan, H\. Khachatrian, D\. C\. Kale, G\. Ver Steeg, and A\. Galstyan, “Multitask learning and benchmarking with clinical time series data,”*Sci\. Data*, vol\. 6, p\. 96, 2019\. - \[3\]S\. Tonekaboni, S\. Joshi, M\. D\. McCradden, and A\. Goldenberg, “What clinicians want: contextualizing explainable machine learning for clinical end use,” in*Proc\. Mach\. Learn\. Healthcare*, 2019\. - \[4\]S\. M\. Lundberg and S\.\-I\. Lee, “A unified approach to interpreting model predictions,” in*Adv\. Neural Inf\. Process\. Syst\.*, pp\. 4765–4774, 2017\. - \[5\]K\. Singhal*et al\.*, “Large language models encode clinical knowledge,”*Nature*, vol\. 620, pp\. 172–180, 2023\. - \[6\]H\. Nori, N\. King, S\. M\. McKinney, D\. Carignan, and E\. Horvitz, “Capabilities of GPT\-4 on medical challenge problems,”*arXiv:2303\.13375*, 2023\. - \[7\]P\. Lee, S\. Bubeck, and J\. Petro, “Benefits, limits, and risks of GPT\-4 as an AI chatbot for medicine,”*New Engl\. J\. Med\.*, vol\. 388, no\. 13, pp\. 1233–1239, 2023\. - \[8\]T\. Chen and C\. Guestrin, “XGBoost: a scalable tree boosting system,” in*Proc\. 22nd ACM SIGKDD Int\. Conf\. Knowl\. Discovery Data Mining*, pp\. 785–794, 2016\. - \[9\]T\. J\. Pollard, A\. E\. W\. Johnson, J\. D\. Raffa, L\. A\. Celi, R\. G\. Mark, and O\. Badawi, “The eICU Collaborative Research Database, a freely available multi\-center database for critical care research,”*Sci\. Data*, vol\. 5, p\. 180178, 2018\. - \[10\]S\. M\. Lundberg, G\. Erion, H\. Chen*et al\.*, “From local explanations to global understanding with explainable AI for trees,”*Nat\. Mach\. Intell\.*, vol\. 2, pp\. 56–67, 2020\. - \[11\]G\. Gutierrez, “Artificial intelligence in the intensive care unit,”*Crit\. Care*, vol\. 24, p\. 101, 2020\. - \[12\]M\. Ghassemi, L\. Oakden\-Rayner, and A\. L\. Beam, “The false hope of current approaches to explainable artificial intelligence in health care,”*Lancet Digit\. Health*, vol\. 3, no\. 11, pp\. e745–e750, 2021\. - \[13\]H\. Touvron*et al\.*, “LLaMA: open and efficient foundation language models,”*arXiv:2302\.13971*, 2023\. - \[14\]L\. Wang*et al\.*, “A survey on large language model based autonomous agents,”*Front\. Comput\. Sci\.*, vol\. 18, no\. 6, p\. 186345, 2024\. - \[15\]Z\. Xi*et al\.*, “The rise and potential of large language model based agents: a survey,”*arXiv:2309\.07864*, 2024\. - \[16\]Y\. Hu, “Toward retrieval\-grounded evaluation for conversational large language model\-based risk assessment,”*JMIR AI*, vol\. 5, e90759, 2026\. doi:10\.2196/90759\.
Similar Articles
RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
RealICU is a hindsight-annotated benchmark for evaluating LLMs in ICU settings, covering four physician-motivated tasks. Experiments reveal that existing LLMs struggle with recall-safety tradeoffs and anchoring bias, while a new structured-memory agent improves reasoning but not fully eliminate safety failures.
LLMs for Cardiovascular Risk Prediction from Structured Clinical Data
This paper presents a hybrid framework that combines structured clinical data with LLM-generated narratives for coronary artery disease prediction, achieving high fidelity in variable extraction and comparing ML models with LLM-based zero-shot and few-shot classification.
Patients-like-me: A Variational LM--GNN Framework for Explainable Clinical Prediction
This paper proposes Patients-like-me (PLM), a unified LM–GNN framework that integrates local patient semantics with global cohort structure for explainable clinical prediction. It introduces a Variational Expectation-Maximization algorithm and demonstrates state-of-the-art results on MIMIC-III and MIMIC-IV with reference-patient explanations.
Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis
This paper presents ClaMPAPP, a hybrid architecture that uses an LLM as an interface to extract features from clinical narratives, which are then passed to an XGBoost classifier for pediatric appendicitis diagnosis, demonstrating improved robustness and safety over end-to-end LLM baselines.
A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
Presents a lightweight knowledge-injection framework for zero-shot ICU delirium prediction that augments structured EHR data summaries with external clinical knowledge at inference time, improving AUROC by up to 8.57 percentage points on LLaMA models without fine-tuning.