RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

arXiv cs.AI Papers

Summary

RealICU is a hindsight-annotated benchmark for evaluating LLMs in ICU settings, covering four physician-motivated tasks. Experiments reveal that existing LLMs struggle with recall-safety tradeoffs and anchoring bias, while a new structured-memory agent improves reasoning but not fully eliminate safety failures.

arXiv:2605.13542v1 Announce Type: new Abstract: Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, underscoring a clear need for reliable AI decision support. Existing ICU benchmarks typically treat historical clinician actions as ground truth. However, these actions are made under incomplete information and limited temporal context of the underlying patient state, and may therefore be suboptimal, making it difficult to assess the true reasoning capabilities of AI systems. We introduce RealICU, a hindsight-annotated benchmark for evaluating large language models (LLMs) under realistic ICU conditions, where labels are created after senior physicians review the full patient trajectory. We formulate four physician-motivated tasks: assess Patient Status, Acute Problems, Recommended Actions, and Red Flag actions that risk unsafe outcomes. We partition each trajectory with 30-min windows and release two datasets: RealICU-Gold with 930-window annotations from 94 MIMIC-IV patients, and RealICU-Scale with 11,862 windows extended by Oracle, a physician-validated LLM hindsight labeler. Existing LLMs including memory-augmented ones performed poorly on RealICU, exposing two failure modes: a recall-safety tradeoff for clinical recommendations, and an anchoring bias to early interpretations of the patient. We further introduce ICU-Evo to study structured-memory agents that improves long-horizon reasoning but does not fully eliminate safety failures. Together, RealICU provides a clinically grounded testbed for measuring and improving AI sequential decision-support in high-stakes care. Project page: https://chengzhi-leo.github.io/RealICU-Bench/
Original Article
View Cached Full Text

Cached at: 05/14/26, 06:16 AM

# RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
Source: [https://arxiv.org/html/2605.13542](https://arxiv.org/html/2605.13542)
Chengzhi Shen1,2,10Weixiang Shen1,2,3Tobias Susetzky1,2 Chen \(Cherise\) Chen4Jun Li1Yuyuan Liu5Xuepeng Zhang6 Zhenyu Gong7,†Daniel Rueckert1,2,8,9,10,†Jiazhen Pan1,2,9,†

1Technical University of Munich \(TUM\)2TUM University Hospital3LMU Munich 4University of Sheffield5University of Oxford6Zhongshan Hospital Fudan University 7Sun Yat\-sen University Cancer Center8Imperial College London 9Munich Center for Machine Learning \(MCML\) 10relAI – Konrad Zuse School of Excellence in Reliable AI †Corresponding Authors

###### Abstract

Intensive care units \(ICU\) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, underscoring a clear need for reliable AI decision support\. Existing ICU benchmarks typically treat historical clinician actions as ground truth\. However, these actions are made under incomplete information and limited temporal context of the underlying patient state, and may therefore be suboptimal, making it difficult to assess the true reasoning capabilities of AI systems\. We introduce RealICU, a hindsight\-annotated benchmark for evaluating large language models \(LLMs\) under realistic ICU conditions, where labels are created after senior physicians review the full patient trajectory\. We formulate four physician\-motivated tasks: assess*Patient Status*,*Acute Problems*,*Recommended Actions*, and*Red Flag*actions that risk unsafe outcomes\. We partition each trajectory with 30\-min windows and release two datasets: RealICU\-Gold with 930\-window annotations from 94 MIMIC\-IV patients, and RealICU\-Scale with 11,862 windows extended by*Oracle*, a physician\-validated LLM hindsight labeler\. Existing LLMs including memory\-augmented ones performed poorly on RealICU, exposing two failure modes: a recall\-safety tradeoff for clinical recommendations, and an anchoring bias to early interpretations of the patient\. We further introduce ICU\-Evo to study structured\-memory agents that improves long\-horizon reasoning but does not fully eliminate safety failures\. Together, RealICU provides a clinically grounded testbed for measuring and improving AI sequential decision\-support in high\-stakes care\. Project page:[chengzhi\-leo\.github\.io/RealICU\-Bench](https://chengzhi-leo.github.io/RealICU-Bench/)

## 1Introduction

The Intensive Care Unit \(ICU\) is one of the most information\-dense environments in the hospital\. Within hours, a single patient can generate large volumes of laboratory results, vital signs, medications, nursing observations, and imaging reports\[manor2008quantifying,pickering2010novel\]\. Physicians must integrate this evolving stream under time pressure, where each measurement captures only a partial slice of the patient’s physiological state, and decisions made in one moment may shape outcomes hours or days later\[paul2023effect,rosa2019effects\]\. This underscores a clear need for AI decision support system in real\-time monitoring and decision\-making in the ICU, which usually acts as a clinical co\-pilot\. In consultations with over 30 board\-certified clinicians, including five senior ICU physicians who later served as annotators, four capabilities emerged as core requirements for a useful ICU co\-pilot: assess*Patient Status*, identify*Acute Problems*, propose*Recommended Actions*, and warn against*Red Flag*actions that may cause unsafe outcomes\. Figure[1](https://arxiv.org/html/2605.13542#S1.F1)illustrates the use case of an AI co\-pilot in ICU decision support\.

Benchmark gap\.Despite rapid progress in Large Language Models \(LLMs\) and agentic systems, few benchmarks evaluate these four capabilities in real\-world ICU settings\. Most clinical benchmarks reduce clinical reasoning to static question answering, diagnosis, or summarization\[ma2024clibench,van2023yet,jin2021disease,jin2019pubmedqa,chiu2025simulating\], or to single\-endpoint prediction \(e\.g\., mortality\[zhao2020prediction\], shock\[ghosh2017septic,yee2019data\], or acute kidney injury\[malhotra2017risk,dong2021machine\]\)\. Such benchmarks aggregate clinical care into isolated predictions, offering little signal on whether a model can reason across a changing patient trajectory\. More importantly, benchmarks built on electronic health record \(EHR\) databases such as MIMIC\-IV\[johnson2023mimic\], HiRID\[hyland2020early\], and eICU\-CRD\[pollard2018eicu\]treat recorded clinician actions as ground\-truth labels\. But this assumption is fragile\. A recorded action reflects what clinicians believed best given incomplete information at the bedside, whereas the optimal action often becomes clear only after reviewing the trajectory using hindsight\. Evaluating AI models against such labels therefore rewards behavioral imitation rather than clinical correctness\.

Proposed benchmark\.To address this gap, we introduce*RealICU*, a hindsight\-grounded benchmark built from MIMIC\-IV\[johnson2023mimic\]for evaluating LLM\-based clinical decision support in the ICU\.*RealICU*evaluates four physician\-motivated tasks over dense 30\-minute windows across the ICU trajectory:*Patient Status*,*Acute Problems*,*Recommended Actions*, and*Red Flags*\. At each window, the agent observes only information available up to that time, while labels are produced by hindsight physician judgment over the full trajectory\. This design scores agents on clinical correctness rather than on recorded behavior\.*RealICU*contains two subsets\.*RealICU\-Gold*provides 930 physician\-labeled windows from 94 ICU stays, and*RealICU\-Scale*extends evaluation to 11,862 windows using*Oracle*, a physician\-validated LLM\-based hindsight evaluator calibrated against expert consensus\.

![Refer to caption](https://arxiv.org/html/2605.13542v1/x1.png)Figure 1:ICU decisions are made under massive data volume and time pressure\. An ICU AI co\-pilot integrates data streams into a decision\-support panel that assesses*Patient Status*, identifies*Acute Problems*, proposes*Recommended Actions*, and warns against unsafe*Red Flag*actions\.Failure mode identification and mitigation\.Using*RealICU*, we benchmark frontier LLM\-based ICU agents across diverse context configurations including memory\. Current agents show poor reliability over long ICU contexts, with two failure modes: \(i\) Recall\-safety tradeoff, where higher recommendation recall comes with up to 47\.3% of these recommendations flagged as potentially harmful; \(ii\) Anchoring bias, where agents preserve early interpretations of the patient despite later contradictory evidence\. To mitigate these, we introduce ICU\-Evo, a structured\-memory agent framework that maintains recent observations, temporal trends, critical events, trajectory summaries, and patient\-specific insights\. ICU\-Evo is backbone\-agnostic and improves clinical reasoning, but its safety failures show that structured memory alone is insufficient for reliable ICU co\-pilots\.

Our key contributions are as follows:

- •We formulate ICU co\-pilot evaluation around four physician\-motivated tasks:*Patient Status*,*Acute Problems*,*Recommended Actions*, and*Red Flags*\. Unlike static clinical QA or outcome prediction benchmarks, these tasks evaluate whether an AI system can support continuous bedside reassessment across an evolving ICU trajectory\.
- •We release*RealICU*, a hindsight\-annotated benchmark for clinical correctness rather than behavioral imitation\. Agents observe only data available at decision time, while labels are produced by hindsight physician judgment over the full trajectory\.*RealICU\-Gold*provides 930 physician\-consensus windows from 94 ICU stays, and*RealICU\-Scale*extends this to 11,862 windows using*Oracle*, a physician\-validated LLM\-based hindsight evaluator\.
- •We identify gaps in current LLM ICU agents and study structured memory as a mitigation\. Across frontier LLMs and multiple context strategies,*RealICU*remains largely unsolved\. We identify a recall–safety tradeoff and anchoring bias as major failure modes, and introduce ICU\-Evo, a structured\-memory agent that improves long\-horizon reasoning but shows that memory alone is insufficient for safe ICU decision support\.

## 2Related Work

##### Clinical Benchmarks for LLMs and Agents\.

Exam\-style benchmarks such as MedQA\[jin2021disease\], PubMedQA\[jin2019pubmedqa\], and MedXpertQA\[zuo2025medxpertqa\]evaluate clinical knowledge as multiple\-choice recall under complete information, a format well\-addressed by state\-of\-the\-art models that reveals little about decisions under uncertainty\. Conversational benchmarks such as AI Hospital\[fan2025ai\], AgentClinic\[schmidgall2024agentclinic\], and VivaBench\[chiu2025simulating\]require agents to gather history, order investigations, and converge on a diagnosis over multiple turns, exposing failure modes such as premature diagnostic closure\. MedAgentBench\[jiang2025medagentbench\]moves closer to real EHR environments but retains a task\-completion framing rather than evaluating overall patient management\. None of these benchmarks evaluates sequential decision\-making over long ICU trajectories or distinguishes behavioral imitation from clinical correctness\.*RealICU*addresses both by grounding evaluation in hindsight physician judgment over the full ICU trajectory, providing dense and trajectory\-level signal of clinical correctness\.

##### Memory\-Augmented LLM Agents\.

Recent LLM agent architectures have explored a range of memory designs\. ReAct\[yao2022react\]appends all reason\-action results sequentially but saturates quickly as context accumulates\. AgentFold\[ye2025agentfold\]addresses this by summarizing completed sub\-tasks at multiple temporal scales\. Evo\-Memory\[wei2025evo\]unifies reasoning, action, and memory refinement in a test\-time loop\. Retrieval\-based systems such as RAG\[arslan2024survey,cuconasu2024power\]and A\-MEM\[xu2025mem\]enable selective access over long histories\. However, these systems treat clinical context equally, making no distinction between static patient background\[mattey2022hospitalised\], time\-sensitive physiological trends\[li2014physiological\], and high\-level trajectory\[sousa2020developmental,reed2015defining\], which play fundamentally different roles in clinical reasoning\. ICU\-Evo organizes clinical context into heterogeneous memory types aligned with these distinctions, enabling systematic study of how structured memory design shapes ICU decision\-making\.

## 3*RealICU*Benchmark

![Refer to caption](https://arxiv.org/html/2605.13542v1/x2.png)Figure 2:Left: Data pipeline for*RealICU\-Gold*and*RealICU\-Scale*\. Right: Data samples for a patient ICU trajectory\. For each evaluation window,*RealICU*provides raw observation data and clinical labels, including patient status, acute problems, action recommendation, and red flag action\.*RealICU*evaluates LLM agents on sequential clinical decision\-making across ICU trajectories, mirroring standard medical quality review: model outputs are assessed against hindsight physician labels produced with full knowledge of patient trajectory rather than against logged clinician actions\.

*RealICU*consists of two datasets\.*RealICU\-Gold*contains 930 sparsely sampled windows from 94 ICU stays labeled by physician consensus\. To scale beyond manual annotation, we introduce*Oracle*, an LLM\-based hindsight evaluator validated against*RealICU\-Gold*, yielding*RealICU\-Scale*with 11,862 densely labeled windows\. Both datasets are released test\-only to prevent leakage\. Detailed statistics are in Figure[8](https://arxiv.org/html/2605.13542#A2.F8), Figure[9](https://arxiv.org/html/2605.13542#A2.F9), and Figure[10](https://arxiv.org/html/2605.13542#A2.F10)\.

Each windowWt=\(Xt;St,Pt,At,Rt\)W\_\{t\}=\(X\_\{t\};\\;S\_\{t\},P\_\{t\},A\_\{t\},R\_\{t\}\)contains clinical observations up to timett, annotated for four tasks:*Patient Status*StS\_\{t\},*Acute Problems*PtP\_\{t\},*Recommended Actions*AtA\_\{t\}, and*Red Flag Actions*RtR\_\{t\}\. The model predicts\(S^t,P^t,A^t\)\(\\hat\{S\}\_\{t\},\\hat\{P\}\_\{t\},\\hat\{A\}\_\{t\}\)fromXtX\_\{t\};RtR\_\{t\}serves as a safety check againstA^t\\hat\{A\}\_\{t\}\. This asymmetry between partial observation and hindsight annotation mirrors the gap between real\-time decision\-making and hindsigth review\. Figure[2](https://arxiv.org/html/2605.13542#S3.F2)illustrates the data construction pipeline and samples\.

### 3\.1Dataset Construction

##### Cohort\.

We sample 94 ICU stays from the MIMIC\-IV\[johnson2023mimic\]cohort, each from a distinct patient and balanced by ICU outcome\. Stays shorter than 4 hours are discarded\. To capture both early stabilization and long trajectories, we balance stays by duration above and below 96 hours\.

##### Windowing\.

We define 30\-minute windows as our evaluation unit and sample them along each ICU trajectory with a 2\-hour stride, preserving short\-term dynamics while limiting redundancy across adjacent windows\. At inference time, the trajectory visible to the model is truncated prior to outcome\-revealing events such as ICU discharge or the discharge summary\.

### 3\.2Tasks

We identify four crucial ICU reasoning tasks below after consulting more than 30 clinicians, including five senior ICU physicians who later served as annotators\. Together they cover the key capabilities of a useful ICU co\-pilot\. For all four tasks, each prediction is accompanied by supporting evidenceℰ⊆Xt\\mathcal\{E\}\\subseteq X\_\{t\}drawn from the raw events in the recorded history\.

*Patient Status*\.A classification of whether the patient is improving, stable, or deteriorating relative to recent context:St=\(st,ℰt\)S\_\{t\}=\(s\_\{t\},\\mathcal\{E\}\_\{t\}\), wherest∈\{improving,stable,deteriorating\}s\_\{t\}\\in\\\{\\texttt\{improving\},\\texttt\{stable\},\\texttt\{deteriorating\}\\\}andℰt⊆Xt\\mathcal\{E\}\_\{t\}\\subseteq X\_\{t\}\.

*Acute Problems*\.A free\-text set of acute problems or emerging risks that require active management:Pt=\{\(pi,ℰi\)\}i=1kP\_\{t\}=\\\{\(p\_\{i\},\\mathcal\{E\}\_\{i\}\)\\\}\_\{i=1\}^\{k\}, whereℰi⊆Xt\\mathcal\{E\}\_\{i\}\\subseteq X\_\{t\}\.

*Action Recommendation*\.A free\-text set of actions likely to benefit the patient within one hour, such as stabilizing physiology or preventing deterioration:At=\{\(aj,ℰj\)\}j=1mA\_\{t\}=\\\{\(a\_\{j\},\\mathcal\{E\}\_\{j\}\)\\\}\_\{j=1\}^\{m\}, whereℰj⊆Xt\\mathcal\{E\}\_\{j\}\\subseteq X\_\{t\}\.

*Red Flags*\.A free\-text set of high\-risk actions that should be avoided because they may be harmful under the patient’s current physiology or trajectory:Rt=\{\(rl,ℰl\)\}l=1nR\_\{t\}=\\\{\(r\_\{l\},\\mathcal\{E\}\_\{l\}\)\\\}\_\{l=1\}^\{n\}, whereℰl⊆Xt\\mathcal\{E\}\_\{l\}\\subseteq X\_\{t\}\.

### 3\.3Annotation Protocol

##### *RealICU\-Gold*with physician consensus\.

We begin from sampling approximately 10 windows per ICU stay by action densityρt=\|ℰtaction\|/\|ℰt\|\\rho\_\{t\}=\|\\mathcal\{E\}^\{\\text\{action\}\}\_\{t\}\|/\|\\mathcal\{E\}\_\{t\}\|, i\.e\. the fraction of action events inside each window\. We draw 80% of windows from theρt≥0\.5\\rho\_\{t\}\\geq 0\.5regime, where interventions are frequent, and 20% fromρt<0\.5\\rho\_\{t\}<0\.5as a control set\. Each window is independently labeled by at least two of five senior ICU physicians\. Inter\-rater reliability \(IRR\) among physicians ranges from 0\.826 to 0\.985 across the four tasks \(Table[1](https://arxiv.org/html/2605.13542#S3.T1.fig1)\), confirming both strong label reproducibility and that the task definitions are sufficiently precise for consistent clinical judgment\. Windows without physician agreement are dropped, yielding 930 validated windows in*RealICU\-Gold*\.

Table 1:*RealICU\-Gold*label quality and*Oracle*validation\.TaskPhys\. IRROracle F1*Patient Status*0\.9850\.987*Acute Problems*0\.9800\.987*Action Recom\.*0\.8260\.895*Red Flags*0\.9160\.964
##### *RealICU\-Scale*with*Oracle*scaling\.

Despite high quality, manual annotation covers only a sparse sample of each ICU stay\. We therefore introduce*Oracle*, an LLM evaluator operating under the same hindsight conditions as the physicians, and apply it to densely label every window across the cohort, yielding 11,862 annotated windows in*RealICU\-Scale*\. We validate*Oracle*by measuring its F1 score against physician consensus on*RealICU\-Gold*\.*Oracle*achieves more than 0\.895 F1 score across all four tasks \(Table[1](https://arxiv.org/html/2605.13542#S3.T1.fig1)\), supporting its use as a reliable hindsight annotator at scale\. While*Oracle*is backbone\-agnostic, we instantiate it with Gemini\-3\.1\-pro\[Gemini31Pro2026\]in this work\. Detailed*Oracle*prompt is in Appendix[E](https://arxiv.org/html/2605.13542#A5)\.

##### Label construction\.

Labels for*Patient Status*,*Acute Problems*, and*Red Flags*are taken directly from annotations\. For*Action Recommendation*, we restrict the annotation space to critical clinical interventions, discarding routine monitoring\. Annotators review each action asbest\-practice,acceptable, orpotentially\-harmful, and may add free\-text actions that should have been taken but were not observed\.AtA\_\{t\}is constructed as the union ofbest\-practiceandacceptableactions together with these free\-text additions\.*Red Flags*are annotated independently as a separate label, not derived frompotentially\-harmfulactions\.

### 3\.4Evaluation Framework

A model under testℳ\\mathcal\{M\}maps observationsXtX\_\{t\}to predictions\(S^t,P^t\(k\),A^t\(k\)\)\(\\hat\{S\}\_\{t\},\\hat\{P\}\_\{t\}^\{\(k\)\},\\hat\{A\}\_\{t\}^\{\(k\)\}\), whereP^t\(k\)\\hat\{P\}\_\{t\}^\{\(k\)\}andA^t\(k\)\\hat\{A\}\_\{t\}^\{\(k\)\}are top\-kkranked lists, with access only to events up to timett\. In this paper we focus on LLM agents, butℳ\\mathcal\{M\}can be any model\. Models are evaluated against*RealICU\-Gold*and*RealICU\-Scale*, providing sparse gold\-standard supervision and trajectory\-level evaluation at scale respectively\. Algorithm[1](https://arxiv.org/html/2605.13542#alg1)summarizes the complete evaluation framework\.

##### Semantic matching\.

To score free\-text tasks \(*Acute Problems*,*Recommended Actions*,*Red Flag Actions*\), we adopt PubMedBERT\[gu2021domain\]and define a binary match, whereτ\\tauis calibrated against 100 expert\-annotated pairs, achieving0\.960\.96F1 atτ=0\.5\\tau=0\.5\(Appendix[A\.5](https://arxiv.org/html/2605.13542#A1.SS5)\):

match​\(xpred,xref\)=𝟏​\[cos⁡\(𝐞pred,𝐞ref\)≥τ\]\.\\mathrm\{match\}\(x\_\{\\text\{pred\}\},\\ x\_\{\\text\{ref\}\}\)=\\mathbf\{1\}\\\!\\left\[\\cos\(\\mathbf\{e\}\_\{\\text\{pred\}\},\\,\\mathbf\{e\}\_\{\\text\{ref\}\}\)\\geq\\tau\\right\]\.\(1\)

##### Metrics\.

*Patient Status*is evaluated with accuracy and macro\-F1 to avoid dominance by the majority class \(stable\)\.*Acute Problems*and*Recommended Actions*are set\-matching tasks evaluated with Hit@kkand Recall@kkatk=5k\{=\}5\.*Red Flag Actions*serves as a safety check via the Harmful Recommendation Rate \(HRR\)\. Let𝒮\\mathcal\{S\}be the set of ICU stays,𝒲s\\mathcal\{W\}\_\{s\}the windows in stayss,A^t\(k\)\\hat\{A\}\_\{t\}^\{\(k\)\}the top\-kkrecommendations, andRtR\_\{t\}the red\-flag set at windowtt; HRR averages the fraction of recommended actions that are flagged across stays:

HRR​\(ℳ\)=1\|𝒮\|​∑s∈𝒮∑t∈𝒲s\|A^t\(k\)∩Rt\|∑t∈𝒲s\|A^t\(k\)\|\.\\mathrm\{HRR\}\(\\mathcal\{M\}\)\\;=\\;\\frac\{1\}\{\|\\mathcal\{S\}\|\}\\sum\_\{s\\in\\mathcal\{S\}\}\\frac\{\\sum\_\{t\\in\\mathcal\{W\}\_\{s\}\}\\big\|\\hat\{A\}\_\{t\}^\{\(k\)\}\\cap R\_\{t\}\\big\|\}\{\\sum\_\{t\\in\\mathcal\{W\}\_\{s\}\}\\big\|\\hat\{A\}\_\{t\}^\{\(k\)\}\\big\|\}\.\(2\)
Algorithm 1*RealICU*Evaluation Framework\.1:model

ℳ\\mathcal\{M\}; label source

ℛ∈\{*Gold*,*Scale*\}\\mathcal\{R\}\\in\\\{\\emph\{Gold\},\\,\\emph\{Scale\}\\\}; ICU stay set

𝒮\\mathcal\{S\}; per\-stay window sets

\{𝒲s\}s∈𝒮\\\{\\mathcal\{W\}\_\{s\}\\\}\_\{s\\in\\mathcal\{S\}\}
2:foreach ICU stay

s∈𝒮s\\in\\mathcal\{S\}do

3:

h←0,n←0h\\leftarrow 0,\\quad n\\leftarrow 0⊳\\trianglerightred\-flag hits / total recommendations

4:foreach window

t∈𝒲st\\in\\mathcal\{W\}\_\{s\}in chronological orderdo

5:

\(S^t,P^t\(k\),A^t\(k\)\)←ℳ​\(Xt\)\(\\hat\{S\}\_\{t\},\\,\\hat\{P\}\_\{t\}^\{\(k\)\},\\,\\hat\{A\}\_\{t\}^\{\(k\)\}\)\\leftarrow\\mathcal\{M\}\(X\_\{t\}\)⊳\\trianglerightmodel sees events up tott

6:

\(St,Pt,At,Rt\)←ℛ​\(t\)\(S\_\{t\},\\,P\_\{t\},\\,A\_\{t\},\\,R\_\{t\}\)\\leftarrow\\mathcal\{R\}\(t\)⊳\\trianglerightpre\-labeled by hindsight annotator

7:evaluate

\(S^t,St\)\(\\hat\{S\}\_\{t\},S\_\{t\}\)for*Patient Status*accuracy and F1

8:evaluate

\(P^t\(k\),Pt\)\(\\hat\{P\}\_\{t\}^\{\(k\)\},P\_\{t\}\)for*Acute Problems*Hit@

kk/ Recall@

kk⊳\\trianglerightsemantic matching

9:evaluate

\(A^t\(k\),At\)\(\\hat\{A\}\_\{t\}^\{\(k\)\},A\_\{t\}\)for*Action Recommendation*Hit@

kk/ Recall@

kk⊳\\trianglerightsemantic matching

10:

h\+=\|A^t\(k\)∩Rt\|;n\+=\|A^t\(k\)\|h\\mathrel\{\+\}=\|\\hat\{A\}\_\{t\}^\{\(k\)\}\\cap R\_\{t\}\|;\\quad n\\mathrel\{\+\}=\|\\hat\{A\}\_\{t\}^\{\(k\)\}\|⊳\\trianglerightsafe recommendation check

11:endfor

12:aggregate per\-window scores

13:endfor

14:returnscores across

𝒮\\mathcal\{S\}for each task

## 4ICU\-Evo: An ICU Agent System with Evolving Memory

ICU decision\-making is sequential, where the underlying patient state is only partially observable with clinical measurements and can only be updated via new observations\. We model this as a partially observable Markov process\[cassandra1998exact,spaan2012partially\]and approximate the latent patient state with a structured memoryMtM\_\{t\}\. We introduce ICU\-Evo as an instance of the memory\-augmented agent frameworks to study how structured memory design shapes clinical decision\-making\.

### 4\.1Memory as a Structured Belief State

Given the contextXtX\_\{t\}and static patient contextcc\(e\.g\. demographics, allergies, pre\-ICU history\), ICU\-Evo maintains a structured memory stateMtM\_\{t\}updated at each window by incorporating the new measurementsxt=Xt−Xt−1x\_\{t\}=X\_\{t\}\-X\_\{t\-1\}, and produces task\-specific predictionsyt\(k\)y\_\{t\}^\{\(k\)\}via

Mt=𝒰​\(Mt−1,xt\),yt\(k\)=f\(k\)​\(Mt,c\)\.M\_\{t\}=\\mathcal\{U\}\(M\_\{t\-1\},\\,x\_\{t\}\),\\qquad y\_\{t\}^\{\(k\)\}=f^\{\(k\)\}\(M\_\{t\},c\)\.\(3\)
The memory decomposes into five components following clinical reasoning:

Mt=\{Mtwork,Mttrend,Mtevent,Mttraj,Mtinsight\}\.M\_\{t\}=\\bigl\\\{M\_\{t\}^\{\\mathrm\{work\}\},\\;M\_\{t\}^\{\\mathrm\{trend\}\},\\;M\_\{t\}^\{\\mathrm\{event\}\},\\;M\_\{t\}^\{\\mathrm\{traj\}\},\\;M\_\{t\}^\{\\mathrm\{insight\}\}\\bigr\\\}\.\(4\)*Working memory*MtworkM\_\{t\}^\{\\mathrm\{work\}\}holds the most recent raw observations at detailed resolution\.*Trend memory*MttrendM\_\{t\}^\{\\mathrm\{trend\}\}captures signal trends of vital and lab values\.*Critical\-event memory*MteventM\_\{t\}^\{\\mathrm\{event\}\}is a persistent, append\-only log of clinically critical events that change the patient story, such as abnormal physiology, interventions, and turning points\.*Trajectory memory*MttrajM\_\{t\}^\{\\mathrm\{traj\}\}provides a compressed narrative of the stay at periodic intervals\.*Insight memory*MtinsightM\_\{t\}^\{\\mathrm\{insight\}\}maintains patient\-specific hypotheses constructed as deviations from population\-level expectation\. Every memory component carries evidence from raw observations, so any clinical decision is explainable and verifiable against the patient record\. In Table[12](https://arxiv.org/html/2605.13542#A3.T12), we summarize the memory components with corresponding agent sources\.

### 4\.2ICU\-Evo Agent Pipeline

ICU\-Evo realizes the memory update operator𝒰\\mathcal\{U\}through three specialized agents operating at different temporal scales over the shared memory\. ICU\-Evo belongs to a broader family of memory\-augmented agent systems\. We discuss it alongside recent agent systems in Appendix[C](https://arxiv.org/html/2605.13542#A3)\. Detailed prompts are reported in Appendix[E](https://arxiv.org/html/2605.13542#A5)\.

##### Observation Agent \(𝒜obs\\mathcal\{A\}\_\{\\text\{obs\}\}\)\.

A rule\-based agent that turns raw measurements into structured signals at every window\. It normalizes units, aligns observations to the 30\-minute window grid, and extracts trend signals from vitals using Piecewise Aggregate Approximation\[guo2010improved\]:

\(Mtwork,Mttrend\)=𝒜obs​\(Mt−1work,Mt−1trend,xt\)\.\\bigl\(M\_\{t\}^\{\\mathrm\{work\}\},\\,M\_\{t\}^\{\\mathrm\{trend\}\}\\bigr\)=\\mathcal\{A\}\_\{\\text\{obs\}\}\(M\_\{t\-1\}^\{\\mathrm\{work\}\},\\,M\_\{t\-1\}^\{\\mathrm\{trend\}\},\\,x\_\{t\}\)\.\(5\)

##### Assessment Agent \(𝒜assess\\mathcal\{A\}\_\{\\text\{assess\}\}\)\.

For everykak\_\{a\}cumulative windows, an LLM transforms recent observations into a trajectory summary and detects critical events\. It consumes the working and trend memory accumulated over the pastkak\_\{a\}windows, producing a trajectory summaryztz\_\{t\}appended toMttrajM\_\{t\}^\{\\mathrm\{traj\}\}and critical eventsete\_\{t\}appended toMteventM\_\{t\}^\{\\mathrm\{event\}\}:

\(zt,et\)=𝒜assess​\(Mt−ka:twork,Mt−ka:ttrend\)\.\\bigl\(z\_\{t\},\\,e\_\{t\}\\bigr\)=\\mathcal\{A\}\_\{\\text\{assess\}\}\\bigl\(M\_\{t\-k\_\{a\}:t\}^\{\\mathrm\{work\}\},\\,M\_\{t\-k\_\{a\}:t\}^\{\\mathrm\{trend\}\}\\bigr\)\.\(6\)

##### Insight Agent \(𝒜insight\\mathcal\{A\}\_\{\\text\{insight\}\}\)\.

Everykik\_\{i\}windows, an LLM proposes hypotheses about what is driving the patient’s clinical course and gathers supporting evidenceese\_\{s\}and counter\-evidenceere\_\{r\}fromMteventM\_\{t\}^\{\\mathrm\{event\}\}\. A hypothesishth\_\{t\}is accepted ifs​\(ht\)\>r​\(ht\)s\(h\_\{t\}\)\>r\(h\_\{t\}\)and rejected otherwise\. The Insight Agent actively reasons about patient\-specific patterns, such as unusual drug responses or persistent abnormalities, promoting individualized care beyond averaged guidelines:

Mtinsight=𝒜insight​\(Mt−1insight,Mt−ki:tevent\)\.M\_\{t\}^\{\\mathrm\{insight\}\}=\\mathcal\{A\}\_\{\\text\{insight\}\}\\bigl\(M\_\{t\-1\}^\{\\mathrm\{insight\}\},\\,M\_\{t\-k\_\{i\}:t\}^\{\\mathrm\{event\}\}\\bigr\)\.\(7\)

##### Predictor \(f\(k\)f^\{\(k\)\}\)\.

The predictor is a task\-specific prompted LLM over the full memory state and static patient context, decoupled from the agent system \(Equation[3](https://arxiv.org/html/2605.13542#S4.E3)\)\.

## 5Evaluation & Analysis

Table 2:Evaluation results on*RealICU\-GOLD*\. Within each backbone,boldmarks the best system per column andunderlinethe second best\.Patient StatusAcute ProblemsAction Recom\.Red FlagsBackboneSystemAcc\.↑\\uparrowF1↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHRR@5↓\\downarrowGemini\-3\.1\-pro\[Gemini31Pro2026\]Full\-context0\.2980\.2580\.4860\.3080\.2590\.1520\.137Local\-window0\.3150\.2390\.4590\.2580\.3950\.2600\.151RAG0\.4020\.3480\.5960\.3420\.4960\.3130\.216\\rowcolorrowhlICU\-Evo0\.4590\.3650\.8230\.5260\.6760\.5340\.300GPT\-5\.4\[OpenAIGPT54\_2026\]Full\-context0\.2940\.2330\.5100\.3480\.4040\.3000\.298Local\-window0\.2330\.1840\.5000\.2930\.3800\.2810\.165RAG0\.2880\.2560\.5990\.3490\.4800\.3980\.234\\rowcolorrowhlICU\-Evo0\.3120\.2640\.8670\.5700\.6760\.5340\.473Qwen3\-235B\[yang2025qwen3\]Full\-context0\.2250\.1880\.3840\.2260\.3290\.2220\.117Local\-window0\.1520\.1540\.2130\.1260\.3520\.2420\.080RAG0\.3150\.2710\.3790\.2110\.4530\.3240\.095\\rowcolorrowhlICU\-Evo0\.2530\.1970\.6000\.3620\.5260\.3570\.117We evaluate ICU\-Evo on*RealICU\-Gold*and*RealICU\-Scale*against three baselines sharing the same predictor: \(i\)*full\-context*, all prior observations up to the window; \(ii\)*local\-window*, the current window only; \(iii\)*RAG*, top\-5 windows retrieved via PubMedBERT\[gu2021domain\]embeddings\. See Appendix[A\.1](https://arxiv.org/html/2605.13542#A1.SS1)for detailed experiment setup\.

### 5\.1*RealICU*Remains Unsolved for Current LLM Systems

*RealICU*remains unsolved for current frontier LLMs and agent systems\. Across all evaluation setups in Table[2](https://arxiv.org/html/2605.13542#S5.T2), ICU\-Evo with Gemini\-3\.1\-pro\[Gemini31Pro2026\]reaches only0\.4590\.459accuracy on*Patient Status*and0\.5340\.534Recall@5 on*Action Recommendation*\. More concerning,*Red Flags*HRR@5 stays non\-trivial across all configurations, indicating current LLM systems still recommend potentially harmful actions in high\-stake ICU setting\. Together, these gaps establish*RealICU*as a clinically grounded safety check for future AI decision\-support systems\.

![Refer to caption](https://arxiv.org/html/2605.13542v1/x3.png)Figure 3:Temporal performance on*RealICU\-Scale*\(Gemini\-3\.1\-pro\[Gemini31Pro2026\]\)\. ICU\-Evo demonstrates its advantage on*Patient Status*and*Acute Problems*even up to 1,800\-hour trajectory\.
### 5\.2Structured Memory Consistently Improves Clinical Reasoning

Structured memory improves performance across all four tasks\. With GPT\-5\.4\[OpenAIGPT54\_2026\], ICU\-Evo improves over RAG by26\.826\.8Hit@5 points on*Acute Problems*and19\.619\.6on*Action Recommendation*, with similar margins on Gemini\-3\.1\-pro\[Gemini31Pro2026\]and Qwen3\-235B\[yang2025qwen3\]\(Table[2](https://arxiv.org/html/2605.13542#S5.T2)\)\. The pattern holds on the densely labeled*RealICU\-Scale*\(Table[4](https://arxiv.org/html/2605.13542#A1.T4), Figure[3](https://arxiv.org/html/2605.13542#S5.F3)\)\. ICU\-Evo’s Hit@5 on*Acute Problems*stays near 0\.8 even for stays up to 1,800 hours, while non\-memory baselines remain about 20 points lower and visibly noisier\. Future ICU decision\-support agents will benefit from memory that actively tracks the patient’s evolving state and scales to long stays\.

### 5\.3The Agent\-*Oracle*Gap: Beyond Behavioral Imitation

We observe a large performance gap between Agent and*Oracle*on*RealICU\-Gold*\. The bottleneck of current ICU agents is not medical knowledge in the LLM backbone but how an agent integrates evidence over time\. With Gemini\-3\.1\-pro\[Gemini31Pro2026\],*Oracle*reaches F10\.9870\.987on*Patient Status*and0\.9640\.964on*Red Flags*identification \(Table[1](https://arxiv.org/html/2605.13542#S3.T1.fig1)\), while ICU\-Evo on the same backbone reaches only0\.3650\.365F1 on*Patient Status*, with a concerningly high rate of harmful recommendations with0\.3000\.300HRR \(Table[2](https://arxiv.org/html/2605.13542#S5.T2)\)\. The four clinical tasks are therefore well handled given the full trajectory but break down under real\-time conditions\. This gap also indicates the value of hindsight evaluation, since scoring agents against recorded clinician actions can only measure how closely the agents imitate human behavior under limited information\. Progress on ICU decision support therefore depends on both stronger real\-time reasoning architectures and the broader adoption of hindsight evaluation\.

### 5\.4Ablation Study

We ablate each component of ICU\-Evo’s memory in a leave\-one\-out setup \(Table[3](https://arxiv.org/html/2605.13542#S5.T3)\)\. Working memory is crucial for local clinical reasoning, and removing it degrades every task with notable drops on*Acute Problems*Hit@5 \(Gemini\-3\.1\-pro\[Gemini31Pro2026\], 0\.823 to 0\.761\)\. Trajectory memory matters for temporal understanding tasks\. Without it,*Acute Problems*and*Action Recommendation*both drop, while*Patient Status*stays stable since it leans on local observations\.

In contrast, insight memory causes fluctuations, and removing it sometimes leads to neutral or beneficial results\. This suggests that current LLMs default to medical\-generalist priors and is not yet capable to identify reliable personalized clinical patterns across long stays\. We examine the failure modes in the next section\.

Table 3:Memory ablation on*RealICU\-GOLD*\. For each row, we remove one component from ICU\-Evo’s memory\. Within each backbone,boldmarks the best result andunderlinethe second best\.Patient StatusAcute ProblemsAction Recom\.Red FlagsBackboneMemory VariantAcc\.↑\\uparrowF1↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHRR↓\\downarrow\\rowcolorrowhlGemini\-3\.1\-pro\[Gemini31Pro2026\]ICU\-Evo0\.4590\.3650\.8230\.5260\.6760\.5340\.300−\-working memory0\.3830\.2940\.7610\.4610\.5070\.3080\.087−\-trend0\.4510\.3520\.8110\.5270\.5210\.3300\.097−\-critical events0\.4450\.3510\.8190\.5270\.5280\.3280\.099−\-trajectory0\.4430\.3620\.7890\.5000\.5060\.3040\.090−\-insight0\.4620\.3560\.8230\.5340\.5550\.3570\.088\\rowcolorrowhlQwen3\-235B\[yang2025qwen3\]ICU\-Evo0\.2530\.1970\.6000\.3620\.5260\.3570\.117−\-working memory0\.1590\.1270\.4210\.2330\.4470\.3070\.117−\-trend0\.2480\.1880\.5520\.3330\.5590\.3930\.122−\-critical events0\.2360\.1870\.5460\.3200\.5950\.4200\.128−\-trajectory0\.2700\.2490\.4860\.2900\.5570\.3930\.138−\-insight0\.2500\.2020\.5870\.3480\.6010\.4200\.141
### 5\.5Failure Mode Analysis

##### *Oracle*failure modes\.

*Oracle*reaches around90%90\\%F1 across all tasks \(Table[1](https://arxiv.org/html/2605.13542#S3.T1.fig1)\)\. The remaining disagreements concentrate on two patterns: \(i\) boundary mis\-calibration on*Patient Status*, where failures fall onstable–improvingorstable–deterioratingborders; \(ii\) granularity mismatch on*Acute Problems*, where*Oracle*reaches for broad descriptors \(e\.g\. hemodynamic instability\) while physicians name specific complications \(e\.g\. ventilator\-associated pneumonia\)\. These are edge cases rather than systematic errors, supporting*Oracle*’s reliability as a large\-scale annotator\.

##### Agent failure modes\.

The most consequential failure is the recall–safety tradeoff, where higher recommendation recall increases the incidence of harmful clinical suggestions\. In Table[2](https://arxiv.org/html/2605.13542#S5.T2), ICU\-Evo \(GPT\-5\.4\[OpenAIGPT54\_2026\]\) gains more than2020*Action Recommendation*Hit@5 points over RAG, but its HRR@5 doubles from 0\.234 to 0\.473\. We use an LLM\-based classifier to group these 394 cases, and the majority concentrate in four high\-stakes families: hemodynamic and pressor management \(n=135n\{=\}135\), volume and diuresis \(n=64n\{=\}64\), anti\-coagulation \(n=54n\{=\}54\), and ventilation/sedation \(n=53n\{=\}53\)\. We find that current LLM agents tend to recognize part of a syndrome and propose the full treatment bundle before contraindications are ruled out\.

The second failure is anchoring bias, where agents over\-commit to early interpretations and ignore later evidence\. Removing insight memory improves*Action Recommendation*Hit@5 from 0\.526 to 0\.601 on Qwen3\-235B\[yang2025qwen3\]\(Table[3](https://arxiv.org/html/2605.13542#S5.T3)\), indicating that generated insights actively mislead the agent\. The agent maintains around66hypotheses per patient,80%80\\%containing anticipatory exceptions \(e\.g\. below\-average tolerance, paradoxical response\)\. These priors push the agent toward rescue bundles even when current evidence is weak\. Case studies are in Appendix[D](https://arxiv.org/html/2605.13542#A4)\.

## 6Discussion

*RealICU*reveals a substantial gap between the medical knowledge of frontier LLMs and their ability to reason under partial observability across an evolving ICU trajectory\. We identify two recurring failure modes that persist across multiple context configurations\. \(i\) the recall–safety tradeoff, where gains in*Recommended Actions*coverage are accompanied by a higher rate of unsafety\. \(ii\) anchoring bias, where agents commit to an early read of the patient and fail to update as new evidence accumulates\. ICU\-Evo uses structured, evidence\-grounded memory at multiple temporal scales to track the evolving patient state, but multi\-scale memory alone does not prevent unsafe recommendations\. Reliable ICU co\-pilots will require advances in long\-context clinical reasoning together with better safety mechanisms\.

Beyond the ICU,*RealICU*offers a methodology for evaluating AI systems where recorded human actions are imperfect and the right action is visible only in hindsight\. We hope this framing supports broader work on evaluating AI systems in high\-stakes sequential decision environments\.

##### Limitations\.

*RealICU*is built on the MIMIC\-IV\[johnson2023mimic\]cohort, and its demographic and care\-pattern distribution may not transfer to ICUs with different staffing or documentation conventions\. Extending to multi\-center and international data is an important direction\. Due to compute constraints, we run a single experiment per LLM configuration and omit variance over long ICU trajectories\. We also focus on text\-based data, leaving multi\-modal data such as imaging and signals to future work\.

## 7Acknowledgement

This paper is supported by the DAAD programme Konrad Zuse Schools of Excellence in Artificial Intelligence, sponsored by the Federal Ministry of Research, Technology and Space\. This work is partially funded by the European Research Council \(ERC\) project Deep4MI \(884622\)\.

## References

## Appendix

### Appendix Contents

- [A\.Performance Analysis](https://arxiv.org/html/2605.13542#A1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A](https://arxiv.org/html/2605.13542#A1)
- [A\.1\.Experiment Setup](https://arxiv.org/html/2605.13542#A1.SS1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A\.1](https://arxiv.org/html/2605.13542#A1.SS1)
- [A\.2\.Evaluation Results on*RealICU\-Scale*](https://arxiv.org/html/2605.13542#A1.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A\.2](https://arxiv.org/html/2605.13542#A1.SS2)
- [A\.3\.Averaged Patient Trajectory on*RealICU\-Scale*](https://arxiv.org/html/2605.13542#A1.SS3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A\.3](https://arxiv.org/html/2605.13542#A1.SS3)
- [A\.4\.Per\-Disease Performance on*RealICU\-GOLD*\.](https://arxiv.org/html/2605.13542#A1.SS4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A\.4](https://arxiv.org/html/2605.13542#A1.SS4)
- [A\.5\.Semantic Matcher Calibration](https://arxiv.org/html/2605.13542#A1.SS5)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A\.5](https://arxiv.org/html/2605.13542#A1.SS5)
- [A\.6\.Token Efficiency](https://arxiv.org/html/2605.13542#A1.SS6)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[A\.6](https://arxiv.org/html/2605.13542#A1.SS6)
- [B\.Dataset Details](https://arxiv.org/html/2605.13542#A2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B](https://arxiv.org/html/2605.13542#A2)
- [B\.1\.Dataset Statistics](https://arxiv.org/html/2605.13542#A2.SS1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.1](https://arxiv.org/html/2605.13542#A2.SS1)
- [B\.2\.*RealICU\-Gold*Cross Validation](https://arxiv.org/html/2605.13542#A2.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.2](https://arxiv.org/html/2605.13542#A2.SS2)
- [B\.3\.Dataset Pre\-processing](https://arxiv.org/html/2605.13542#A2.SS3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[B\.3](https://arxiv.org/html/2605.13542#A2.SS3)
- [C\.Memory\-Augmented Agents for Clinical Decision Support](https://arxiv.org/html/2605.13542#A3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C](https://arxiv.org/html/2605.13542#A3)
- [C\.1\.Formulation](https://arxiv.org/html/2605.13542#A3.SS1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C\.1](https://arxiv.org/html/2605.13542#A3.SS1)
- [C\.2\.Instantiations](https://arxiv.org/html/2605.13542#A3.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C\.2](https://arxiv.org/html/2605.13542#A3.SS2)
- [C\.3\.ICU\-Evo as Heterogeneous Clinical Memory](https://arxiv.org/html/2605.13542#A3.SS3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C\.3](https://arxiv.org/html/2605.13542#A3.SS3)
- [C\.4\.Discussion](https://arxiv.org/html/2605.13542#A3.SS4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[C\.4](https://arxiv.org/html/2605.13542#A3.SS4)
- [D\.Case Study](https://arxiv.org/html/2605.13542#A4)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D](https://arxiv.org/html/2605.13542#A4)
- [D\.1\.Failure Case: Recall Safety Tradeoff](https://arxiv.org/html/2605.13542#A4.SS1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D\.1](https://arxiv.org/html/2605.13542#A4.SS1)
- [D\.2\.Failure Case: Anchoring Bias](https://arxiv.org/html/2605.13542#A4.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D\.2](https://arxiv.org/html/2605.13542#A4.SS2)
- [D\.3\.Memory Snapshot](https://arxiv.org/html/2605.13542#A4.SS3)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[D\.3](https://arxiv.org/html/2605.13542#A4.SS3)
- [E\.Prompts](https://arxiv.org/html/2605.13542#A5)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[E](https://arxiv.org/html/2605.13542#A5)
- [E\.1\.Oracle Prompt](https://arxiv.org/html/2605.13542#A5.SS1)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[E\.1](https://arxiv.org/html/2605.13542#A5.SS1)
- [E\.2\.Agent Prompt](https://arxiv.org/html/2605.13542#A5.SS2)\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.\.[E\.2](https://arxiv.org/html/2605.13542#A5.SS2)

## Appendix APerformance Analysis

### A\.1Experiment Setup

We evaluate ICU\-Evo on*RealICU\-Gold*and*RealICU\-Scale*against three baselines sharing the same predictor: \(i\)*full\-context*, all prior observations up to the window; \(ii\)*local\-window*, the current window only; \(iii\)*RAG*, top\-5 windows retrieved via PubMedBERT\[gu2021domain\]embeddings\. We setkak\_\{a\}andkik\_\{i\}to 12 windows \(6 hours\) for ICU\-Evo\. For*Action Recommendation*, we strip the current window’s recorded actions before prediction to prevent label leakage\. We evaluate every window in*RealICU\-Gold*and every fourth window along the trajectory in*RealICU\-Scale*\.

We use two closed\-source LLMs \(Gemini\-3\.1\-pro\[Gemini31Pro2026\], GPT\-5\.4\[OpenAIGPT54\_2026\]\) and one open\-source LLM \(Qwen3\-235B\-A22B\[yang2025qwen3\]\) as backbones for evaluation\. Evaluation results on*RealICU\-Gold*and*RealICU\-Scale*are reported in Table[2](https://arxiv.org/html/2605.13542#S5.T2)and Table[4](https://arxiv.org/html/2605.13542#A1.T4)respectively\. Full\-context Gemini and GPT runs on*RealICU\-Scale*are omitted due to compute budget on stays beyond hundreds of hours\.*Oracle*uses Gemini\-3\.1\-pro\[Gemini31Pro2026\]to generated hindsight annotations with access to the full patient trajectory\.

### A\.2Evaluation Results on*RealICU\-Scale*

Table[4](https://arxiv.org/html/2605.13542#A1.T4)reports full evaluation results on*RealICU\-Scale*across all three backbones and four systems\. Full\-context evaluation is omitted for Gemini\-3\.1\-pro\[Gemini31Pro2026\]and GPT\-5\.4\[OpenAIGPT54\_2026\]due to prohibitive inference cost over multi\-day ICU trajectories\. Qwen3\-235B\[yang2025qwen3\]is included as a reference open\-weight upper bound\.

The results on*RealICU\-Scale*largely recapitulate the pattern observed on*RealICU\-GOLD*\(Table[2](https://arxiv.org/html/2605.13542#S5.T2)\)\. ICU\-Evo achieves the strongest performance on*Acute Problems*and*Action Recommendation*across all three backbones, with particularly large margins on*Acute Problems*Hit \(up to\+0\.268\+0\.268over RAG for Gemini\-3\.1\-pro\[Gemini31Pro2026\]\)\. The*Red Flag*HRR remains the consistent weak point of ICU\-Evo regardless of backbone, suggesting the same premature anchoring failure mode \(see Sec\.[5\.5](https://arxiv.org/html/2605.13542#S5.SS5)\), where current agent systems over\-commit to early interpretation of the patient instead of updating hypothesis with new observations\. Qwen3\-235B\[yang2025qwen3\]achieves lower overall performance compared to Gemini\-3\.1\-pro\[Gemini31Pro2026\]and GPT\-5\.4\[OpenAIGPT54\_2026\], suggesting that weaker instruction\-following reduces the benefit of structured memory on tasks requiring precise categorical judgment\.

We further illustrate the temporal performance of LLM agents with full\-context, local\-window, retrieval\-augmentation, memory\-augmentation configurations in Figure[4](https://arxiv.org/html/2605.13542#A1.F4)and Figure[5](https://arxiv.org/html/2605.13542#A1.F5)\.

Table 4:Evaluation results on the*RealICU\-Scale*\. Within each backbone,boldmarks the best system per column andunderlinethe second best\.Patient StatusAcute ProblemsAction Recom\.Red FlagsBackboneSystemAcc\.↑\\uparrowF1↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHRR↓\\downarrowGemini\-3\.1\-pro\[Gemini31Pro2026\]Full\-context–––––––Local\-window0\.4050\.2640\.4870\.2650\.4470\.3070\.066RAG0\.4420\.3120\.5680\.3150\.4660\.3310\.073\\rowcolorrowhlICU\-Evo0\.5190\.3480\.8270\.5180\.5140\.33011footnotemark:10\.087GPT\-5\.4\[OpenAIGPT54\_2026\]Full\-context–––––––Local\-window0\.4150\.2650\.4750\.2660\.4510\.3080\.073RAG0\.4110\.2690\.5840\.3210\.5090\.4350\.096\\rowcolorrowhlICU\-Evo0\.4380\.3270\.8520\.5620\.5750\.3680\.090Qwen3\-235B\[yang2025qwen3\]Full\-context0\.2010\.1160\.4010\.2320\.4550\.2990\.215Local\-window0\.1750\.1590\.2540\.1420\.4400\.2950\.207RAG0\.3670\.2820\.3790\.2070\.4460\.3420\.225\\rowcolorrowhlICU\-Evo0\.3040\.1770\.6490\.3750\.5150\.3270\.292![Refer to caption](https://arxiv.org/html/2605.13542v1/x4.png)Figure 4:Temporal performance over the full ICU stay on*RealICU\-Scale*\(GPT\-5\.4\[OpenAIGPT54\_2026\]\)\.![Refer to caption](https://arxiv.org/html/2605.13542v1/x5.png)Figure 5:Temporal performance over the full ICU stay on*RealICU\-Scale*\(Qwen3\-235B\[yang2025qwen3\]\)\.
### A\.3Averaged Patient Trajectory on*RealICU\-Scale*

We visualize patient trajectories on*RealICU\-Scale*using the*Patient Status*label in Figure[6](https://arxiv.org/html/2605.13542#A1.F6)\. We map each window\-level label to an ordinal score \(deteriorating = \-1, stable = 0, improving = 1\) and normalize time within each ICU stay to the interval \[0,1\]\. After binning each trajectory into 20 normalized time bins and averaging repeated observations within patient\-bin pairs, we plotted all individual trajectories as low\-opacity curves and overlaid outcome\-stratified cohort means with 95% confidence bands\. This highlights both patient\-level heterogeneity and the average temporal separation between survivors and non\-survivors\.

The survived and died cohorts are already separated at admission, with survivors hovering nearstableand non\-survivors sitting consistently below it, and the gap widens over the course of the stay as the survivor mean drifts towardimprovingwhile the non\-survivor mean declines sharply in the final 20% of normalized ICU time\. Both cohorts show substantial patient\-level heterogeneity in the thin lines, which is expected given the diversity of admission diagnoses, but the cohort means recover the clinically intuitive ordering that survivors trend upward and non\-survivors trend downward\. This pattern indicates that the window\-level labels produced by*Oracle*aggregate into a coherent patient\-level signal, and supports the use of*RealICU\-Scale*for trajectory\-level analyses despite its labels being generated rather than physician\-annotated\.

![Refer to caption](https://arxiv.org/html/2605.13542v1/x6.png)Figure 6:Averaged patient status trajectories from*Oracle*on*RealICU\-Scale*\. Window\-level*Patient Status*labels are mapped to an ordinal score \(deteriorating=−1=\-1, stable=0=0, improving=\+1=\+1\) and aggregated into normalized duration time\. Thin lines show individual trajectories, and thick lines and shaded regions show cohort means with 95% confidence bands\. The survived and died cohorts separate from admission onward and diverge further over the stay\.
### A\.4Per\-Disease Performance on*RealICU\-GOLD*\.

Tables[5](https://arxiv.org/html/2605.13542#A1.T5)report a breakdown across the six disease groups exceeding 8% prevalence in*RealICU\-Gold*using Gemini\-3\.1\-pro\[Gemini31Pro2026\]backbone\. The decomposition tests whether the gains of ICU\-Evo are driven by a subset of phenotypes or hold across the case mix\.

The dominance of ICU\-Evo on context\-heavy tasks \(Patient Status, Acute Problems, Recommended Actions\) is consistent across nearly all disease groups\. ICU\-Evo achieves the best Hit@5 on Acute Problems for every group, with margins over the strongest baseline ranging from 0\.134 \(GI & Hepatic\) to 0\.358 \(Sepsis & Infection\), so the benefit of structured longitudinal memory is not phenotype\-specific\. The pattern is most pronounced on Cardiovascular and GI & Hepatic cases, where ICU\-Evo wins on six of seven metrics, suggesting that diseases with protracted trajectories benefit most from explicit trend and trajectory memory\.

Respiratory and Sepsis & Infection expose the limits of the current memory design on Recommended Actions\. RAG matches or surpasses ICU\-Evo on Hit@5 and R@5 for these two groups, while ICU\-Evo retains its lead on upstream tasks\. Respiratory and septic management is dominated by recurring, protocol\-driven interventions such as ventilator adjustments and antimicrobial escalation, for which lexical retrieval over recent context is competitive with longitudinal memory\. This aligns with the backbone\-dependent memory tolerance reported in the Qwen ablation and reinforces that memory architecture and task structure interact\.

We report more detailed per\-disease results with GPT\-5\.4\[OpenAIGPT54\_2026\]backbone in Table[6](https://arxiv.org/html/2605.13542#A1.T6), and with Qwen\-235B\[yang2025qwen3\]backbone in Table[7](https://arxiv.org/html/2605.13542#A1.T7)\.

Table 5:Per\-disease performance on*RealICU\-GOLD*\(Gemini\-3\.1\-pro\[Gemini31Pro2026\]backbone\)\. Disease groups are omitted with less than 8% proportion\. Within each group,boldmarks the best system per column andunderlinethe second best\.Patient StatusAcute ProblemsAction Recom\.Red FlagsDisease GroupSystemAcc\.↑\\uparrowF1↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHRR@5↓\\downarrowCardiovascularFull\-context0\.3370\.2690\.5840\.3670\.2990\.1770\.066Local\-window0\.3130\.2460\.4500\.2450\.4000\.2730\.059RAG0\.4090\.3700\.5930\.3370\.4420\.2840\.066\\rowcolorrowhlICU\-Evo0\.4720\.3740\.8280\.5260\.5950\.3750\.089Sepsis & InfectionFull\-context0\.1470\.1660\.2870\.1540\.1560\.0750\.065Local\-window0\.3270\.2680\.4580\.2580\.3780\.2260\.067RAG0\.3530\.2830\.5740\.3120\.5350\.3290\.089\\rowcolorrowhlICU\-Evo0\.4530\.3350\.8160\.4910\.5140\.2960\.098Injury & PoisoningFull\-context0\.3030\.1990\.4690\.3380\.2250\.1450\.057Local\-window0\.3030\.2090\.5120\.3060\.3900\.2570\.090RAG0\.4740\.3790\.6360\.3690\.5260\.3610\.078\\rowcolorrowhlICU\-Evo0\.4900\.3990\.8170\.5340\.4510\.2990\.095RespiratoryFull\-context0\.3400\.2480\.6210\.3900\.3790\.2070\.112Local\-window0\.2700\.2180\.4560\.2550\.4250\.2410\.056RAG0\.2900\.2600\.5940\.3400\.5050\.2840\.084\\rowcolorrowhlICU\-Evo0\.3200\.2780\.8510\.5590\.5510\.3230\.120GI & HepaticFull\-context0\.4380\.3960\.5270\.3380\.3470\.2260\.115Local\-window0\.3500\.3110\.5100\.2980\.4270\.3170\.122RAG0\.4120\.3910\.6450\.4080\.5870\.3600\.073\\rowcolorrowhlICU\-Evo0\.5000\.4370\.7790\.5120\.5920\.3820\.098All Diseases \(Table[2](https://arxiv.org/html/2605.13542#S5.T2)\)Full\-context0\.2980\.2580\.4860\.3080\.2590\.1520\.137Local\-window0\.3150\.2390\.4590\.2580\.3950\.2600\.151RAG0\.4020\.3480\.5960\.3420\.4960\.3130\.216\\rowcolorrowhlICU\-Evo0\.4590\.3650\.8230\.5260\.6760\.5340\.300Table 6:Per\-disease performance on*RealICU\-GOLD*\(GPT\-5\.4\[OpenAIGPT54\_2026\]backbone\)\. Disease groups are omitted with less than 8% proportion\. Within each group,boldmarks the best system per column andunderlinethe second best\.Patient StatusAcute ProblemsAction Recom\.Red FlagsDisease GroupSystemAcc\.↑\\uparrowF1↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHRR@5↓\\downarrowCardiovascularFull\-context0\.3500\.2500\.5890\.3960\.5180\.3870\.157Local\-window0\.2620\.1900\.4840\.2860\.3270\.2970\.114RAG0\.3000\.2430\.5780\.3130\.4870\.4560\.138\\rowcolorrowhlICU\-Evo0\.2750\.2460\.8530\.5580\.7050\.5620\.136Sepsis & InfectionFull\-context0\.1530\.1640\.3140\.1910\.1370\.0870\.111Local\-window0\.1530\.2130\.5280\.2930\.3680\.2330\.114RAG0\.2270\.2920\.6090\.3290\.4300\.3270\.095\\rowcolorrowhlICU\-Evo0\.3470\.2660\.8640\.5390\.7000\.5580\.158Injury & PoisoningFull\-context0\.2490\.1570\.5100\.3730\.4040\.3060\.161Local\-window0\.3190\.1500\.4790\.2990\.4820\.3180\.138RAG0\.3280\.1820\.5450\.3640\.5450\.4310\.157\\rowcolorrowhlICU\-Evo0\.3280\.2670\.9220\.6480\.6180\.4550\.104RespiratoryFull\-context0\.3400\.2540\.6530\.4380\.5430\.4770\.176Local\-window0\.3220\.1890\.5330\.2970\.3860\.2110\.137RAG0\.3780\.2900\.6780\.4240\.3190\.2470\.113\\rowcolorrowhlICU\-Evo0\.2700\.2370\.9230\.6310\.7670\.6230\.182GI & HepaticFull\-context0\.4250\.3490\.5750\.4410\.4740\.3230\.306Local\-window0\.1000\.1650\.5610\.3730\.3310\.2790\.098RAG0\.2140\.2560\.6820\.5120\.5120\.4120\.148\\rowcolorrowhlICU\-Evo0\.3630\.3470\.8270\.5700\.6570\.4900\.157All Diseases \(Table[2](https://arxiv.org/html/2605.13542#S5.T2)\)Full\-context0\.2940\.2330\.5100\.3480\.4040\.3000\.298Local\-window0\.2330\.1840\.5000\.2930\.3800\.2810\.165RAG0\.2880\.2560\.5990\.3490\.4800\.3980\.234\\rowcolorrowhlICU\-Evo0\.3120\.2640\.8670\.5700\.6760\.5340\.473Table 7:Per\-disease performance on*RealICU\-GOLD*\(Qwen3\-235B\[yang2025qwen3\]backbone\)\. Disease groups are omitted with less than 8% proportion\. Within each group,boldmarks the best system per column andunderlinethe second best\.Patient StatusAcute ProblemsAction Recom\.Red FlagsDisease GroupSystemAcc\.↑\\uparrowF1↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHit@5↑\\uparrowR@5↑\\uparrowHRR@5↓\\downarrowCardiovascularFull\-context0\.2180\.2060\.4550\.2490\.3900\.2700\.129Local\-window0\.1560\.1640\.1880\.1090\.3510\.2460\.087RAG0\.3070\.2780\.3500\.1890\.4510\.3320\.090\\rowcolorrowhlICU\-Evo0\.2680\.2190\.5520\.3160\.5520\.3630\.134Sepsis & InfectionFull\-context0\.1070\.1220\.2330\.1370\.1560\.1080\.083Local\-window0\.1470\.1860\.2240\.1130\.3220\.1980\.098RAG0\.3330\.2820\.3480\.1760\.4430\.2770\.093\\rowcolorrowhlICU\-Evo0\.1400\.1300\.5750\.3280\.5300\.3530\.098Injury & PoisoningFull\-context0\.2690\.1530\.4320\.2890\.2910\.2320\.093Local\-window0\.1970\.1650\.3010\.2210\.3730\.2800\.065RAG0\.3620\.2520\.5160\.3280\.4810\.3570\.102\\rowcolorrowhlICU\-Evo0\.3050\.2200\.6690\.4570\.4520\.3210\.087RespiratoryFull\-context0\.2800\.1450\.4440\.2660\.4190\.2360\.116Local\-window0\.1400\.1170\.2530\.1470\.3540\.2220\.052RAG0\.3400\.3110\.4510\.2720\.4630\.3310\.083\\rowcolorrowhlICU\-Evo0\.3800\.2050\.6580\.4040\.6750\.4570\.162GI & HepaticFull\-context0\.3750\.3150\.4590\.2950\.3820\.2700\.188Local\-window0\.1750\.1500\.2140\.1380\.3340\.2730\.123RAG0\.3380\.3360\.3290\.1990\.4590\.3690\.099\\rowcolorrowhlICU\-Evo0\.2620\.2580\.6460\.4130\.5080\.3940\.132All Diseases \(Table[2](https://arxiv.org/html/2605.13542#S5.T2)\)Full\-context0\.2250\.1880\.3840\.2260\.3290\.2220\.117Local\-window0\.1520\.1540\.2130\.1260\.3520\.2420\.080RAG0\.3150\.2710\.3790\.2110\.4530\.3240\.095\\rowcolorrowhlICU\-Evo0\.2530\.1970\.6000\.3620\.5260\.3570\.117
### A\.5Semantic Matcher Calibration

We adopt PubMedBERT\[gu2021domain\]\(NeuML/pubmedbert\-base\-embeddings\) to generate embeddings for semantic match for*Acute Problems*,*Action Recommendation*, and*Red Flags*tasks\.

##### Calibration set\.

We sampled 100 action\-string pairs from held\-out ICU windows and asked a board\-certified intensivist to label each pair as a binary classification for semantic match or non\-match\. The set is balanced by construction with 50 matched pairs and 50 non\-matched pairs\.

##### Threshold sweep\.

Table[8](https://arxiv.org/html/2605.13542#A1.T8)reports precision, recall, F1, and accuracy at seven candidate thresholds, and Figure[7](https://arxiv.org/html/2605.13542#A1.F7)visualises the trade\-off\. PubMedBERT cosine similarity separates the two classes almost perfectly \(AUROC=0\.996=0\.996\)\. Precision reaches1\.001\.00for allτ≥0\.5\\tau\\geq 0\.5, while recall decays monotonically asτ\\tauincreases\. We selectτ∗=0\.5\\tau^\{\*\}=0\.5, which maximises F1 \(0\.9580\.958\) and eliminates false positives while retaining92%92\\%of true matches\. This operating point is used in all reported evaluations\.

Table 8:Semantic matcher performance on the 100\-pair calibration set across candidate thresholds\. The selected threshold \(τ∗=0\.5\\tau^\{\*\}=0\.5\) maximises F1 and achieves perfect precision\.τ\\tauAccuracyPrecisionRecallF10\.30\.860\.781\.000\.880\.40\.950\.911\.000\.950\.5∗0\.961\.000\.920\.960\.60\.821\.000\.640\.780\.70\.661\.000\.320\.480\.80\.531\.000\.060\.120\.90\.500\.000\.000\.00![Refer to caption](https://arxiv.org/html/2605.13542v1/x7.png)Figure 7:Evaluating PubMedBERT\[gu2021domain\]matcher on the calibration set under different thresholds\. The selectedτ∗=0\.5\\tau^\{\*\}=0\.5\(dashed line\) achieves the best overall performance\.

### A\.6Token Efficiency

We assess token efficiency from two complementary perspectives, namely the per\-prediction cost and the longitudinal coverage delivered per input token\. A direct comparison of raw token counts suggests that ICU\-Evo is more expensive than the local\-window and RAG baselines\. This view, however, omits a central design objective of ICU\-Evo, which is to surface broad trajectory context at every prediction step\. We therefore report a coverage\-normalized metric alongside the raw cost\.

##### Per\-prediction cost\.

We report the average input and total tokens per prediction on the Qwen run, with RAG projected to match the call volume of the local\-window baseline\. ICU\-Evo is not the cheapest configuration in raw tokens, yet it is substantially cheaper than full\-context prompting while remaining more expensive than the local\-window and RAG baselines, as shown in Table[9](https://arxiv.org/html/2605.13542#A1.T9)\.

##### Coverage\-normalized efficiency\.

To account for the trajectory context that each mode actually surfaces, we define the covered windows per prediction as11for the local\-window baseline,1\+k1\+kfor RAG withkkretrieved windows, andwindow\_index\+1\\text\{window\\\_index\}\+1for ICU\-Evo and the full\-context baseline, reflecting the current window together with all accumulated prior context\. We then report the covered windows per million input tokens and the input tokens consumed per covered window on the*Patient Status*task\. Once normalized by timeline coverage, ICU\-Evo becomes the most input\-efficient mode, achieving the highest coverage density and the lowest input\-token cost per covered window, as shown in Table[10](https://arxiv.org/html/2605.13542#A1.T10)\.

Taken together, these two views indicate that, although ICU\-Evo consumes more tokens per prediction than the local\-window baseline, it delivers substantially denser longitudinal context per input token, which reflects better token utilization for timeline\-aware reasoning in the ICU\.

Table 9:Per\-prediction token cost on*RealICU\-Scale*with Qwen3\-235B\[yang2025qwen3\]\.ModePredictionsInput tokensAvg\. input / pred\.Avg\. total / pred\.Full\-context11,065405,300,94836,629\.1036,763\.84Local\-window11,86225,382,4992,139\.822,319\.49RAG \(projected\)11,86283,797,7717,064\.397,318\.39\\rowcolorrowhl ICU\-Evo11,862254,272,58721,435\.9021,971\.55Table 10:Coverage\-normalized input efficiency on Patient Status\.ModeCovered windowsWindows / 1M input tok\.Input tok\. / windowFull\-context5,112,58912,614\.3079\.28Local\-window11,862467\.332,139\.82RAG6,130541\.321,847\.34\\rowcolorrowhl ICU\-Evo6,304,41024,793\.9040\.33

## Appendix BDataset Details

### B\.1Dataset Statistics

##### Cohort Statistics

Figure[8](https://arxiv.org/html/2605.13542#A2.F8)summarizes the demographic and clinical composition of the selected 94\-patient cohort across six dimensions: disease category, ICU stay duration, age, sex, stay duration stratified by survival outcome, and mean event density per window stratified by outcome\.

To categorize each ICU stay with disease types, we extract the diagnosis code closest in time to ICU admission\. ICD\-9 and ICD\-10 codes were then mapped to broad disease categories using a rule\-based grouping based on ICD chapters, with sepsis\-related codes grouped into a dedicated Sepsis and Severe Infection category\. Specifically, the largest disease group was Cardiovascular Disorders \(32\.98%\), followed by Sepsis and Severe Infection \(15\.96%\), Injury/Poisoning \(13\.83%\), Respiratory Disorders \(10\.64%\), and Gastrointestinal/Hepatic Disorders \(8\.51%\)\. The remaining categories were less common: Neurological Disorders \(4\.26%\); Clinical Signs/Symptoms, Congenital Disorders, Infectious Diseases, and Oncology \(2\.13% each\); and Endocrine/Metabolic Disorders, Hematologic Disorders, Musculoskeletal Disorders, Psychiatric Disorders, and Renal/Genitourinary Disorders \(1\.06% each\)\. In the pie chart, disease categories below 5% were merged into Others

![Refer to caption](https://arxiv.org/html/2605.13542v1/x8.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x9.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x10.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x11.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x12.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x13.png)

Figure 8:Cohort Demographics and Clinical Characteristics of the 94\-Patient*RealICU*Cohort
##### *RealICU\-Gold*Label Statistics

Figure[9](https://arxiv.org/html/2605.13542#A2.F9)summarizes the distributional properties ofRealICU\-Gold\. The coverage histogram exhibits a long\-tail distribution, with windows concentrated within the first 120 hours after ICU admission and a long right tail extending past 1,200 hours, yielding a median position of 74\.8 hours\. The*Patient Status*distribution is dominated by Stable windows \(63\.0%\), followed by Deteriorating \(22\.4%\) and Improving \(14\.6%\)\. For the set\-valued tasks,*Acute Problems*is tightly concentrated around two concurrent problems per window, whereas*Recommended Actions*exhibits a heavier\-tailed distribution with a small number of windows reaching twelve or more concurrent recommendations, reflecting the variable cognitive load of ICU management\.*Red Flag Actions*remain rare by design, with a median of one per window and most windows containing zero or one event\.

![Refer to caption](https://arxiv.org/html/2605.13542v1/x14.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x15.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x16.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x17.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x18.png)

Figure 9:*RealICU\-Gold*statistics and label distribution for Patient Status, Active Problems, Recommended Action, and Red Flags\.
##### *RealICU\-Scale*Label Statistics

Figure[10](https://arxiv.org/html/2605.13542#A2.F10)summarizes the distributional properties of*RealICU\-Scale*\. The coverage histogram exhibits a long\-tail distribution, with windows concentrated within the first 336 hours after ICU admission and a long right tail extending past 1,800 hours, yielding a median position of 207\.8 hours\. The*Patient Status*distribution is dominated by Stable windows \(68\.8%\), followed by Deteriorating \(23\.1%\) and Improving \(8\.2%\)\. For the set\-valued tasks,*Acute Problems*is tightly concentrated around two to three concurrent problems per window, whereas*Recommended Actions*exhibits a heavier\-tailed distribution with a small number of windows reaching ten or more concurrent recommendations, reflecting the variable cognitive load of ICU management\.*Red Flag Actions*has a median of one per window and most windows containing zero or one event\.

![Refer to caption](https://arxiv.org/html/2605.13542v1/x19.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x20.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x21.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x22.png)
![Refer to caption](https://arxiv.org/html/2605.13542v1/x23.png)

Figure 10:*RealICU\-Scale*statistics and label distribution for Patient Status, Active Problems, Recommended Action, and Red Flags\.

### B\.2*RealICU\-Gold*Cross Validation

*RealICU\-Gold*contains 930 windows in total\. For each window, we invite at least two out of five senior physicians for annotation\. And we run a cross\-validation check after annotation to maintain the golden\-standard labels\. In Table[11](https://arxiv.org/html/2605.13542#A2.T11), we report the detailed number of each labels before and after cross\-validation\. Only labels with agreements are kept into*RealICU\-Gold*\. Note that*Active Problems*,*Action Recomm\.*, and*Red Flags*are stored as sets with multiple labels per window\.

Table 11:Label\-wise statistics ofRealICU\-Goldafter cross\-validation filtering\.TaskN labels rawN labels keptKeep rate*Patient Status*93092199\.0%*Active Problems*2,1702,06695\.2%*Action Recomm\.*2,3282,19894\.4%*Red Flags*1,2201,05886\.7%
### B\.3Dataset Pre\-processing

To obtain our underlying base dataset of trajectories that cover ICU stays as well as their preceding patient journey, we merge MIMIC\-IV\[johnson2023mimic\], MIMIC\-ED\[johnson2023mimic\], MIMIC\-Note\[johnson2023mimic\], MIMIC\-IV\-ECHO\[johnson2023mimic\], MIMIC\-IV\-ECG and MIMIC\-CXR\[johnson2019mimic\]\.

By this, we include not only patient meta data such as demographics, insurance, etc\., but a diverse holistic timeline of medication, online medical records, vital measurements, X\-ray, electro\- and echocardiograms, procedures, diagnosis, lab results, text reports, and transfers\. We also include triaging data, subject to availability\. From MIMIC\-Note, we use the entire contents of the discharge summaries and the findings sections from radiology reports\. In total, our resulting base dataset comprises 73,181 ICU stays from 50,920 patients\.

We arrange all charted information and measurements along a time axis together with patient age and time delta to the beginning of the specific ICU stay and sort them temporally ascending\. Full duplicates are eliminated\. Encoded categorical information from established ontologies and coding systems, e\.g\. for diagnosis \(ICD\) or medication \(GSN\), are resolved to their full\-text descriptions\. Text data is cleaned according to a permissive policy, only adjusting e\.g\. consecutive whitespace characters and unambiguous processing artifacts\. Numerical data is also represented textually together with the respective unit of measurement and description\. While we directly include all textually representable information and numeric measurements, we limit the integration of imaging and waveforms to their metadata, leaving the utilization of the X\-ray, ECG, and ECHO contents to future work\. We ensure that patient data is not leaked across our dataset splits\.

Further, we account for inaccurate charting and limitations of raw data collection by conservatively establishing an adversarial tolerance of 24h for key events such as discharge\. In case of multiple records for the same event with different precision \(e\.g\. death\), usually originating from different tables in the raw dataset, we default to the most fine\-grain timestamp\.

## Appendix CMemory\-Augmented Agents for Clinical Decision Support

We position ICU\-Evo as an instance of the broader class of memory\-augmented language agents\. In the following, we discuss a generic formulation of the class, several specific instantiations from recent work, and the design choices that motivate ICU\-Evo\.

### C\.1Formulation

A memory\-augmented agent processes a stream of inputs\{x1,x2,…,xT\}\\\{x\_\{1\},x\_\{2\},\\dots,x\_\{T\}\\\}while maintaining an evolving memory state\. At steptt, the update and decision rules take the generic form

Mt=𝒰​\(Mt−1,xt\),yt\(k\)=f\(k\)​\(Mt\),M\_\{t\}=\\mathcal\{U\}\(M\_\{t\-1\},\\,x\_\{t\}\),\\qquad y\_\{t\}^\{\(k\)\}=f^\{\(k\)\}\(M\_\{t\}\),\(8\)whereMtM\_\{t\}is the memory state,𝒰\\mathcal\{U\}is an update operator that integrates the latest input into memory, andf\(k\)f^\{\(k\)\}is a task\-specific decision function realized as a prompted call of the underlying language model\. Different memory systems differ primarily in the structure ofMtM\_\{t\}and in the choice of𝒰\\mathcal\{U\}, and the structural choices that define a memory system reduce to three questions\. What*types*of content doesMtM\_\{t\}contain, at what*temporal scale*is each type maintained, and under what*update policy*does each type evolve?

### C\.2Instantiations

We describe three instantiations of Eq\.[8](https://arxiv.org/html/2605.13542#A3.E8)in whichMtM\_\{t\}takes increasingly heterogeneous forms\.

##### Compressive Stream Memory\.

AgentFold\[ye2025agentfold\]setsMtM\_\{t\}as an ordered sequence of summary blocks together with a high\-fidelity record of the latest interaction\. The update operator𝒰\\mathcal\{U\}is a learned folding policy that, at each step, either condenses the latest interaction into a fine\-grained block or consolidates a contiguous span of prior blocks into a single coarse\-grained block\. This instantiation supports streaming inputs and adaptive scale, while committing all memory content to a single representational type \(textual summary\) under a single update rule \(replacement by summarization\)\.

##### Cross\-Task Experience Memory\.

Evo\-Memory\[wei2025evo\]setsMtM\_\{t\}as an unordered set of prior task experiences, each encoded as a structured tuple\(xi,y^i,fi\)\(x\_\{i\},\\hat\{y\}\_\{i\},f\_\{i\}\), wherefif\_\{i\}is a feedback signal\. The update operator𝒰\\mathcal\{U\}is append\-with\-pruning, and a separate refine action lets the agent reorganize or discard memory entries during decision\-making\. This instantiation targets cross\-task transfer rather than within\-task dynamics, and treats each task as the atomic unit of memory\.

##### Linked Note Memory\.

A\-Mem\[xu2025mem\]setsMtM\_\{t\}as a collection of atomic notes, where each note is a tuple of raw content, timestamp, LLM\-generated keywords, tags, and contextual description, a dense embedding, and a set of links to other notes\. The update operator𝒰\\mathcal\{U\}is realized in two LLM\-driven steps\. On arrival of a new note, top\-kkretrieval over the embedding space surfaces candidate neighbors, and an LLM decides which neighbors deserve a semantic link\. The same neighbors are then re\-examined, and the LLM may rewrite the contextual description, keywords, or tags of any neighbor in light of the new note\. This instantiation supports streaming inputs and introduces evolution of prior entries, while committing all memory content to a single note schema under a single LLM\-driven update rule\.

Table 12:ICU\-Evo’s memory components and the corresponding agent update operator\.ComponentDefinitionUpdated byMworkM^\{\\mathrm\{work\}\}Recent raw observations at full resolution\.Observation AgentMtrendM^\{\\mathrm\{trend\}\}Piecewise\-constant segmentations of vitals and labs\.Observation AgentMeventM^\{\\mathrm\{event\}\}Append\-only log of critical events\.Assessment AgentMtrajM^\{\\mathrm\{traj\}\}Compressed episode\-level narrative of the stay\.Assessment AgentMinsightM^\{\\mathrm\{insight\}\}Patient\-specific hypotheses with supporting and counter\-evidence\.Insight Agent

### C\.3ICU\-Evo as Heterogeneous Clinical Memory

ICU\-Evo setsMtM\_\{t\}as a tuple of five components,

Mt=\{Mtwork,Mttrend,Mtevent,Mttraj,Mtinsight\},M\_\{t\}=\\bigl\\\{M\_\{t\}^\{\\mathrm\{work\}\},\\;M\_\{t\}^\{\\mathrm\{trend\}\},\\;M\_\{t\}^\{\\mathrm\{event\}\},\\;M\_\{t\}^\{\\mathrm\{traj\}\},\\;M\_\{t\}^\{\\mathrm\{insight\}\}\\bigr\\\},\(9\)defined in Table[12](https://arxiv.org/html/2605.13542#A3.T12)\. Algorithm[2](https://arxiv.org/html/2605.13542#alg2)formalizes the full inference loop and the pipeline of three agents that realize the update operator𝒰\\mathcal\{U\}at different temporal cadences over the shared memory state\. At every windowtt, the Observation Agent ingests the new measurementsxtx\_\{t\}and updatesMtworkM\_\{t\}^\{\\mathrm\{work\}\}by per\-window overwrite andMttrendM\_\{t\}^\{\\mathrm\{trend\}\}by piecewise aggregation\. Everykak\_\{a\}windows, the Assessment Agent compresses the recent working and trend memory into a trajectory summaryztz\_\{t\}, appended toMttrajM\_\{t\}^\{\\mathrm\{traj\}\}as a multi\-scale rollup, and detects newly emerging critical eventsE~t\\tilde\{E\}\_\{t\}, appended toMteventM\_\{t\}^\{\\mathrm\{event\}\}under severity gating\. Everykik\_\{i\}windows, the Insight Agent proposes patient\-specific hypotheses, gathers supporting and counter\-evidence fromMteventM\_\{t\}^\{\\mathrm\{event\}\}, and the Orchestrator commits the accepted hypotheses toMtinsightM\_\{t\}^\{\\mathrm\{insight\}\}via lifecycle transitions\. The Predictor then queries the consolidated memory state to emit task\-specific predictionsyt\(k\)y\_\{t\}^\{\(k\)\}, decoupled from the memory update cycle\.

The five components form principled correspondences with prior designs, recombined under a common formulation\.MtworkM\_\{t\}^\{\\mathrm\{work\}\}andMttrajM\_\{t\}^\{\\mathrm\{traj\}\}mirrors the multi\-scale summaries of AgentFold\[ye2025agentfold\], the lifecycle\-managed update ofMtinsightM\_\{t\}^\{\\mathrm\{insight\}\}mirrors both the rewriting of prior notes in A\-Mem\[xu2025mem\]and the refine action of Evo\-Memory\[wei2025evo\]\.

Algorithm 2ICU\-Evo Memory\-Augmented Agent System\.1:LLM backbone

ℱ\\mathcal\{F\}; ICU stay

ss; window sequence

\{xt\}t=1T\\\{x\_\{t\}\\\}\_\{t=1\}^\{T\}; static context

cc; agent periods

ka,kik\_\{a\},k\_\{i\}
2:Initialize memory

M0M\_\{0\}⊳\\trianglerightwork, trend, event, traj, insight

3:foreach window

t=1,…,Tt=1,\\ldots,Tdo

4:

\(Mtwork,Mttrend\)←Observe​\(Mt−1work,Mt−1trend,xt\)\\bigl\(M\_\{t\}^\{\\mathrm\{work\}\},\\,M\_\{t\}^\{\\mathrm\{trend\}\}\\bigr\)\\leftarrow\\mathrm\{Observe\}\\bigl\(M\_\{t\-1\}^\{\\mathrm\{work\}\},\\,M\_\{t\-1\}^\{\\mathrm\{trend\}\},\\,x\_\{t\}\\bigr\)⊳\\trianglerightObservation Agent; every window

5:if

tmodka=0t\\bmod k\_\{a\}=0then⊳\\trianglerightAssessment Agent fires everykak\_\{a\}windows

6:

\(zt,E~t\)←ℱ​\(Mt−ka:twork,Mt−ka:ttrend\)\\bigl\(z\_\{t\},\\,\\tilde\{E\}\_\{t\}\\bigr\)\\leftarrow\\mathcal\{F\}\\bigl\(M\_\{t\-k\_\{a\}:t\}^\{\\mathrm\{work\}\},\\,M\_\{t\-k\_\{a\}:t\}^\{\\mathrm\{trend\}\}\\bigr\)
7:

Mttraj←Mt−1traj∪\{zt\}M\_\{t\}^\{\\mathrm\{traj\}\}\\leftarrow M\_\{t\-1\}^\{\\mathrm\{traj\}\}\\cup\\\{z\_\{t\}\\\};

Mtevent←Mt−1event∪E~tM\_\{t\}^\{\\mathrm\{event\}\}\\leftarrow M\_\{t\-1\}^\{\\mathrm\{event\}\}\\cup\\tilde\{E\}\_\{t\}
8:endif

9:if

tmodki=0t\\bmod k\_\{i\}=0then⊳\\trianglerightInsight Agent fires everykik\_\{i\}windows

10:

Δ​H←ℱ​\(Mt−1insight,Mt−ki:tevent\)\\Delta H\\leftarrow\\mathcal\{F\}\\bigl\(M\_\{t\-1\}^\{\\mathrm\{insight\}\},\\,M\_\{t\-k\_\{i\}:t\}^\{\\mathrm\{event\}\}\\bigr\)⊳\\trianglerightpropose/update hypotheses

11:foreach hypothesis

h∈Δ​Hh\\in\\Delta Hdo

12:

state​\(h\)←accept\\mathrm\{state\}\(h\)\\leftarrow\\textit\{accept\}if

s​\(h\)\>r​\(h\)s\(h\)\>r\(h\)elsereject

13:endfor

14:

Mtinsight←Orchestrator​\(Mt−1insight,Δ​H\)M\_\{t\}^\{\\mathrm\{insight\}\}\\leftarrow\\mathrm\{Orchestrator\}\\bigl\(M\_\{t\-1\}^\{\\mathrm\{insight\}\},\\,\\Delta H\\bigr\)
15:endif

16:foreach task

kkdo⊳\\trianglerightPredictor decoupled from memory update

17:

yt\(k\)←ℱ\(k\)​\(Mt;c\)y\_\{t\}^\{\(k\)\}\\leftarrow\\mathcal\{F\}^\{\(k\)\}\\bigl\(M\_\{t\};\\,c\\bigr\)
18:endfor

19:endfor

20:returnpredictions

\{yt\(k\)\}\\\{y\_\{t\}^\{\(k\)\}\\\}for evaluation against*RealICU*labels

### C\.4Discussion

The instantiations above demonstrate the flexibility of Eq\.[8](https://arxiv.org/html/2605.13542#A3.E8), yet alternative combinations remain possible\. The heterogeneous decomposition we adopt reflects that clinical reasoning under partial observability proceeds along multiple simultaneous modes\. A homogeneous memory forces a single answer to three independent questions: at what temporal scale to retain content, at what fidelity, and under what update policy\. AgentFold\[ye2025agentfold\]couples scale and fidelity under a uniform textual summary type, fitting neither append\-only event logs nor lifecycle\-managed hypotheses\. A\-Mem\[xu2025mem\]couples all three under a uniform note schema and LLM\-driven evolution rule, providing no mechanism for the distinct update policies that event detection and hypothesis lifecycle management each require\. Evo\-Memory\[wei2025evo\]treats each task as the unit of experience, fitting cross\-task transfer but leaving within\-patient dynamics unaddressed\. The heterogeneous memory structure of ICU\-Evo in Eq\.[9](https://arxiv.org/html/2605.13542#A3.E9)offers a frameworks to align previous designs into real clinical reasoning over evolving patient states\.

## Appendix DCase Study

### D\.1Failure Case: Recall Safety Tradeoff

Patient background: Age 55\. Female, entering ICU with severe subarachnoid hemorrhage\. Course complicated by severe intracranial hypertension \(ICP\>\>40 mmHg\) managed with mannitol, propofol, and right frontal EVD placement\. Current 191\.5–192\.0 hours after ICU admission \(day\-8\)\.

Current\-window events11:16Impaired Tissue Perfusion NCP – Therapeutic temperature management 11:16Impaired Tissue Perfusion NCP – Pressors 11:17Infection NCP – Monitor and assess respiratory secretions 11:17Impaired Tissue Perfusion NCP – Monitor neuro assessment 11:17Impaired Tissue Perfusion NCP – Monitor hemodynamic status q1h …\(63 further entries\)

Memory snapshotWorking memory:68 events from the current window\.Trend memory:None at current window\.Critical events:08:16 VITALS Richmond\-RAS Scale=−4\.00\-4\.00; 11:00 VITALS Intra Cranial Pressure 21\.0 mmHg; …Trajectory:This patient exhibits a highly coupled neuro\-respiratory deterioration, demonstrating acute and severe hypoxemia simultaneously with sudden intracranial pressure spikes …Insight \#6:This patient exhibits extreme ICP volatility, demonstrating sudden, severe ICP spikes that break through maximal medical management, including deep sedation, neuromuscular blockade, and targeted temperature management\. Expect a higher\-than\-average risk of sudden, refractory intracranial crises and secondary cerebral hypoperfusion despite seemingly stable baseline ICP,requiring aggressive, anticipatory hyperosmolar or surgical interventions\.

Gold annotationPatient Status\.Deteriorating\.Active Problems\.Refractory intracranial hypertension and impending herniation;worsening hypernatremia and hyperosmolality \(Na 155 mEq/L, osm 325 mOsm/kg\); ventilator\-associated pneumonia\.Recommended Actions\.Therapeutic temperature management; monitor and assess respiratory secretions\.Red Flags\.Routine hyperosmolar therapy \(mannitol or hypertonic saline\)— further routine administration is contraindicated and dangerous under current sodium and osmolality; emergency bolusing only as a last resort for acute herniation\.

ICU\-Evo predictions \(recommended actions\)1\.Urgent neurosurgery and neurocritical care consultation\.\[unmatch\]2\.Administer hyperosmolar therapy\. Boluses of mannitol or 3% hypertonic saline for sustained ICP\>\>20–22 mmHg\.\[red flag\]3\.Maintain CPP\>\>70 mmHg\. Continuous norepinephrine to meet MAP goals\.\[match\]4\.Initiate goals\-of\-care discussion\.\[unmatch\]5\.Strict glycemic and electrolyte monitoring \(q1–2h K, glucose\)\.\[unmatch\]

Figure 11:Recall\-safety tradeoff case study\. ICU\-Evo’s stored insight \#6 prescribes “aggressive, anticipatory hyperosmolar interventions,” which propagates to prediction 2 — flagged as contraindicated by the gold annotation under current Na/osm\. The trend layer carries no sodium signal forward, so nothing counterweighs the stored frame\.
### D\.2Failure Case: Anchoring Bias

Patient background: Adult patient with Eisenmenger’s physiology, severe right ventricular dysfunction, and chronic right\-to\-left shunting with polycythemia\. Baseline SpO285–93%\. Prior course marked by refractory hypoxemia \(SpO2nadir 77% on FiO2100%\) unresponsive to maximal high\-flow oxygen and ongoing diuresis\. Current 87\.5–88\.0 hours after ICU admission \(day\-4\)\.

Current\-window events18:48BODY\_INPUT Oral/Gastric Ingredient = 300 ml 18:48BODY\_INPUT PO Intake = 300 ml 18:48BODY\_INPUT Water = 300 ml 18:48VITALS Weight = 51\.8 kg \(4 events total; no vitals stream this window\)

Memory snapshotWorking memory:4 events from the current window\.Trend memory:No vital signal this window \(HR/RR/SpO2/BP/MAP/Temp all empty\)\.Critical events:02\-03 13:00 VITALS SpO277% on FiO2100%; 02\-04 06:00 VITALS O2Flow 35 L/min; 02\-04 22:00 BODY\_OUTPUT Void 700 ml; …Trajectory:This patient demonstrates a coupled refractory\-hypoxemia and preload\-dependence pattern, where standard oxygen escalation fails to correct shunt\-driven desaturation and successful volume offloading precipitates hemodynamic collapse …Insight \#2:This patient demonstrates refractory hypoxemia that is unresponsive to maximal high\-flow oxygen therapy and ongoing diuresis\. Expect below\-average improvement in oxygenation indices from standard volume offloading,likely due to underlying pulmonary hypertension and fixed V/Q mismatch or shunting\.

Gold annotationPatient Status\.Stable\.Active Problems\.Hypotension and reduced cardiac output \(preload\-sensitive\); volume overload risk under chronic right\-heart strain\.Recommended Actions\.BODY\_INPUT Water \(oral intake\);VITALS Weight \(daily tracking\)\.Red Flags\.Titrating supplemental oxygen to SpO2\>\>95%— baseline SpO2is 85–93% from chronic right\-to\-left shunting and polycythemia; oxygen should be administered only if SpO2drops below 85% or the patient becomes symptomatic\.

ICU\-Evo predictions \(recommended actions\)1\.Hold or reduce diuretic therapy\. Suspend or decrease loop diuretics and reassess volume status before further diuresis\.\[unmatch\]2\.Titrate norepinephrine to maintain MAP\>\>65 mmHg, weaning cautiously if hemodynamics remain stable\.\[unmatch\]3\.Monitor serum potassium via basic metabolic panel or venous blood gas\.\[unmatch\]4\.Maintain targeted oxygenation\. Continue 4 L/min nasal cannula to target SpO288–92%, avoiding over\-oxygenation\.\[red flag\]

Figure 12:Premature\-anchoring case study\. The window contains four events — oral water and a daily weight — and the gold status is stable\. ICU\-Evo’s stored insight \#2 carries forward the prior cardiopulmonary story of refractory hypoxemia, propagating to prediction 4, which prescribes an active oxygen target \(88–92%\) that the gold annotation marks as contraindicated under this patient’s Eisenmenger baseline\. The trend layer is empty for the current window, so nothing pulls the model back to the simpler window\-grounded interpretation\.
### D\.3Memory Snapshot

ICU\-Evo memory snapshot — window 120, 87\.5–88\.0 hWorking Memoryrecent windows, raw eventsprevious windows:17 events \(windows 118–119, omitted\)current window\(window 120, 87\.5–88\.0 h, 4 events\):
2152\-02\-06 18:48 BODY\_INPUT Oral/Gastric Ingredient, ml = 300\.00
2152\-02\-06 18:48 BODY\_INPUT PO Intake, ml = 300\.00
2152\-02\-06 18:48 BODY\_INPUT Water, ml = 300\.00
2152\-02\-06 18:48 VITALS Weight = 51\.80Trend Memoryvital\-sign aggregates, two scopescurrent windownoneglobal\(windows 0–120, 0\.0–88\.0 h, 3422 raw events\):
heart\_rate\_bpm: mean = 68\.97, min = 58\.00, max = 85\.00, count = 97
resp\_rate\_per\_min: mean = 13\.16, min = 8\.00, max = 30\.00, count = 96
spo2\_percent: mean = 89\.87, min = 77\.00, max = 100\.00, count = 97
map\_mmhg: mean = 75\.00, min = 47\.00, max = 99\.00, count = 88
…\(sbp, dbp, temperature omitted\)Critical Events Memorysalient events that change patient storyprevious episodes:38 events \(episodes 1–9, hours 0\.0–79\.5, omitted\)current episode\(episode 10, hours 80\.0–87\.5\):
\(no critical events extracted\)Trajectory Memoryepisode\-level summariesepisode 1\(hours 0\.0–7\.5\): The patient was admitted to the MICU for management of acute decompensated heart failure, acute kidney injury, and hypercapnic respiratory failure\. Respiratory support was initiated with high\-flow nasal cannula at 35 L/min and 65% FiO2, ……
episode 10\(hours 80\.0–87\.5\): The patient began the block with stable hemodynamics \(MAP 70 mmHg\) and borderline oxygenation \(SpO2 88%\) on 4 L/min nasal cannula\. Throughout the period, mean arterial pressures were maintained between 70 and 80 mmHg, demonstrating sustained hemodynamic stability\. Respiratory status remained stable, with oxygen saturations ranging from 88% to 94% on unchanged nasal cannula support …Insight Memorypersonalized hypotheses with supporting and counter evidenceinsight \#1:This patient exhibits a paradoxical and rapid escalation in serum potassium despite ongoing loop diuretic therapy\. Expect an above\-average risk of severe hyperkalemia and resistance to standard potassium\-wasting effects of furosemide\.
supporting:03:20 LAB\_TEST Potassium = 7\.30 mEq/L
counter:03:20 LAB\_TEST Creatinine = 1\.90 mg/dLinsight \#2:This patient demonstrates refractory hypoxemia that is unresponsive to maximal high\-flow oxygen therapy and ongoing diuresis\. Expect below\-average improvement in oxygenation indices from standard volume offloading,likely due to underlying pulmonary hypertension and fixed V/Q mismatch or shunting\.
supporting:
13:00 VITALS O2 saturation pulseoxymetry, =77\.00 %
13:00 VITALS Inspired O2 Fraction =1 00\.00
counter:
06:00 VITALS Inspired O2 Fraction =60\.00
06:00 VITALS O2 Flow = 35\.00 L/min
……Figure 13:ICU\-Evo memory snapshot at 87\.5–88\.0 hours after admission\. The five layers of memory together constitute the full state available to the prediction modules at this window, including working memory, trend, critical events, trajectory, and patient\-specific insights\. Red highlights mark the thread most relevant to the case study in Figure[12](https://arxiv.org/html/2605.13542#A4.F12)\.

## Appendix EPrompts

### E\.1Oracle Prompt

`Oracle Prompt`

`E\.2 Agent Prompt Assessment Agent Prompt Insight Agent E\.2\.1 Predictor Prompt Shared Prompt Patient Status Predictor Active Problems Predictor Action Recommendation Predictor Red Flag Actions Predictor`

Similar Articles

RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts

arXiv cs.AI

Introduces RECON, a benchmark for evaluating compositional reasoning over long contexts in LLM-based agents, spanning 24 case files across criminal, medical, and financial domains. The best non-oracle system achieves only 22.4% accuracy, revealing substantial limitations in current memory architectures.