Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

arXiv cs.AI Papers

Summary

Asclepius is an adaptive agent scaffolding for long-horizon clinical tasks that improves critical-action correctness by 22% over a baseline agent framework in simulated emergency-department shifts.

arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional grading. On CES, current agents reach the correct diagnosis in most cases yet fail to deliver complete and timely critical actions, revealing an execution gap. We attribute this gap to three long-horizon failure modes, each operationalized as a per-trace counter: instruction-adherence drift, treatment incompleteness, and a severity-equity gap in timeliness. We then introduce Asclepius, an adaptive agent scaffolding with a self-evolving harness that rewrites the operating manual between shifts from trace-level feedback, an externalized clinical skills library for high-stakes regimen knowledge, and three isolated subagents that partition per-turn decisions across the patient queue. On held-out batches never observed during harness evolution, Asclepius improves critical-action correctness by 22% (p = 0.024) over a strong baseline agent framework while preserving diagnostic accuracy, with consistent gains across five LLM judges from three model families; on the full ten-batch set, improvements reach 25% on critical actions and 13% on timeliness. The three failure modes form a coupled bottleneck: decisive reductions appear only when all three components act together.
Original Article
View Cached Full Text

Cached at: 09/15/26, 08:56 AM

# Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
Source: [https://arxiv.org/html/2609.13543](https://arxiv.org/html/2609.13543)
Xiaoman ZhangAffiliation:Department of Biomedical Informatics, Harvard Medical School, Boston, MACorrespondence:[yuangc@mit\.edu](mailto:[email protected]),[pranav\_rajpurkar@hms\.harvard\.edu](mailto:[email protected])Sung Eun KimAffiliation:Department of Biomedical Informatics, Harvard Medical School, Boston, MACorrespondence:[yuangc@mit\.edu](mailto:[email protected]),[pranav\_rajpurkar@hms\.harvard\.edu](mailto:[email protected])Luyang LuoAffiliation:Department of Biomedical Informatics, Harvard Medical School, Boston, MACorrespondence:[yuangc@mit\.edu](mailto:[email protected]),[pranav\_rajpurkar@hms\.harvard\.edu](mailto:[email protected])Pranav RajpurkarAffiliation:Department of Biomedical Informatics, Harvard Medical School, Boston, MACorrespondence:[yuangc@mit\.edu](mailto:[email protected]),[pranav\_rajpurkar@hms\.harvard\.edu](mailto:[email protected])

###### Abstract

LLM agents are predominantly benchmarked on short, single\-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures\. We use the Clinical Environment Simulator \(CES\), in which an agent manages an entire emergency\-department shift under continuous time and resource pressure, as a testbed: long\-horizon execution failures manifest measurably in a single rollout under structured, multi\-dimensional grading\. On CES, current agents reach the correct diagnosis in most cases yet fail to deliver complete and timely critical actions, revealing an*execution gap*\. We attribute this gap to three long\-horizon failure modes, each operationalized as a per\-trace counter: instruction\-adherence drift, treatment incompleteness, and a severity\-equity gap in timeliness\. We then introduceAsclepius, an adaptive agent scaffolding with a self\-evolving harness that rewrites the operating manual between shifts from trace\-level feedback, an externalized clinical skills library for high\-stakes regimen knowledge, and three isolated subagents that partition per\-turn decisions across the patient queue\. On held\-out batches never observed during harness evolution,Asclepiusimproves critical\-action correctness by 22% \(p=0\.024p=0\.024\) over a strong baseline agent framework while preserving diagnostic accuracy, with consistent gains across five LLM judges from three model families; on the full ten\-batch set, improvements reach 25% on critical actions and 13% on timeliness\. The three failure modes form a*coupled bottleneck*: decisive reductions appear only when all three components act together\.

## 1Introduction

![Refer to caption](https://arxiv.org/html/2609.13543v1/figures/framework.png)Figure 1:Overview of Asclepius\.Two coupled loops wrap a baseline agent\. Within a shift \(inner loop\), the Main Agent uses an MCP tool interface, a clinical skills library, and three isolated subagents\. Across shifts \(outer loop\), a proposer agent reads the shift trace and rewrites the operating manual for the next shift\.From Short\-Horizon to Long\-Horizon Agent Deployment\.Progress in LLM agents has been driven largely by short\-horizon, single\-task benchmarks where success is scored at the end of a brief trajectory\. Yet real deployments demand sustained competent execution across hours of interleaved tasks under contention, where a different class of failures emerges that is invisible to per\-task accuracy but dominates end\-of\-shift performance\. We use the Clinical Environment Simulator \(CES\)\([Luo et al\., 2026](https://arxiv.org/html/2609.13543#bib.bib1)\)as our testbed: a long\-horizon emergency\-department shift in which patients arrive on staggered schedules with hidden deterioration thresholds, vitals evolve continuously, and the agent must interleave diagnosis with time\-sensitive interventions under contention for limited beds and staff\. Performance is graded along four dimensions: diagnosis, critical actions, timeliness, and disposition\.

Gaps in Current Agents\.We evaluate four configurations representative of current agent practice on CES: three raw frontier\-model agents \(Codex\([OpenAI, 2025a](https://arxiv.org/html/2609.13543#bib.bib25)\), Gemini\([Google, 2025](https://arxiv.org/html/2609.13543#bib.bib28)\), and Claude Code\([Anthropic, 2025a](https://arxiv.org/html/2609.13543#bib.bib21)\)\) and a stronger baseline that wraps Claude Code in an MCP\-based clinical tool interface with a hand\-crafted operating manual\. The tool\-using configuration substantially outperforms the raw agents, yet a striking asymmetry persists across all four: diagnostic accuracy is high \(4\.39/5\) while critical\-action scores remain below 3\.0/5 \(2\.94/5\) and timeliness while higher, still trails at 3\.34/5\. We call this the*execution gap*: the failure to translate correct diagnosis into complete and timely action under multi\-patient load\.

Failure Modes\.We identify three failure modes that underlie this execution gap, each operationalized with one per\-trace counter \(Section[6\.3](https://arxiv.org/html/2609.13543#S6.SS3), Figure[3](https://arxiv.org/html/2609.13543#S6.F3)\)\.\(i\)*Instruction\-adherence drift*: a hand\-crafted operating manual that works at hour zero decays in adherence as the shift lengthens and clinical trajectories diversify, a runtime problem that hand\-tuning cannot scale to address; measured by the*Drift*counter \(early→\\tolate rise in naked\-disposition rate\)\.\(ii\)*Treatment incompleteness*: even with a correct diagnosis, the agent frequently orders zero disease\-specific treatment, partial multi\-drug regimens, or familiar defaults instead of guideline\-specified drugs; measured by the*Incomplete*counter \(zero\-treatment rate\)\.\(iii\)*Severity\-equity gap*: as the queue widens, the agent’s attention collapses onto recent and easier cases, leaving harder and severe patients with systematically delayed care; measured by the*Neglect*counter \(timeliness gap between non\-severe and severe patients\)\.

Asclepius\.We proposeAsclepius, an adaptive agent scaffolding with three components, each targeting one of the failure modes above\.\(i\)Aself\-evolving harnesstreats the operating manual as a learnable natural\-language artifact: a proposer agent reads trace\-level feedback \(scores, actions, missed treatments\) from completed shifts and rewrites the manual between iterations\.\(ii\)A curatedclinical skills libraryexternalizes procedural treatment\-selection knowledge \(regimen templates, dosage tables, condition\-to\-action mappings\) as an on\-demand reference\.\(iii\)Threeisolated subagents\(triage prioritization, diagnosis formation, treatment completion\) partition per\-turn decisions into focused context windows\.

Results\.We evaluate Asclepius on ten patient batches in CES \(six search and four held out from manual evolution\)\. Pooled across all ten batches, Asclepius delivers gains on critical actions \(\+25%\+25\\%,p<0\.001p<0\.001\) and timeliness \(\+13%\+13\\%,p<0\.01p<0\.01\) while preserving diagnostic accuracy\. On the four held\-out batches alone \(n=48n=48patients\), the critical\-action gain remains significant \(\+0\.63\+0\.63,p=0\.024p=0\.024\); overall \(\+0\.19\+0\.19\) and timeliness \(\+0\.31\+0\.31\) remain positive but not significant\. Five judges from three model families reproduce the same pattern, so the improvement is not an artifact of the original grader\. Component ablations show that no single component closes the execution gap on its own\. The three failures behave as a*coupled bottleneck*: no single component reduces all three per\-trace counters, and the largest joint reduction appears only when the self\-evolving manual, the skills library, and the isolated subagents act together\.

## 2Related Work

### 2\.1Clinical LLM Agents and Evaluation

Evaluation of LLMs in medicine has progressed through three regimes\. Early work used static, single\-question benchmarks drawn from licensing exams and curated vignettes, where models answer one case in isolation\. A second wave introduced multi\-turn diagnostic agents that interview a single patient or reason over a single case across multiple steps: AMIE conducts conversational history\-taking[Tu et al\. \(2025\)](https://arxiv.org/html/2609.13543#bib.bib2);[McDuff et al\. \(2025\)](https://arxiv.org/html/2609.13543#bib.bib3), MAI\-DxO operationalizes diagnosis as sequential test ordering[Nori et al\. \(2025\)](https://arxiv.org/html/2609.13543#bib.bib4), and benchmarks such as AgentClinic and MedAgentBench evaluate agents in simulated clinical environments[Schmidgall et al\. \(2024\)](https://arxiv.org/html/2609.13543#bib.bib5);[Jiang et al\. \(2025\)](https://arxiv.org/html/2609.13543#bib.bib8)\. Multi\-agent variants extend this paradigm by simulating entire hospitals[Li et al\. \(2024\)](https://arxiv.org/html/2609.13543#bib.bib6);[Fan et al\. \(2025\)](https://arxiv.org/html/2609.13543#bib.bib7)or expanding tool use into therapeutic and rare\-disease reasoning[Gao et al\. \(2025\)](https://arxiv.org/html/2609.13543#bib.bib9);[Zhao et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib10)\. These systems remain single\-patient and untimed, with evaluation dominated by the diagnosis dimension\. Recent work argues that clinical AI evaluation must move toward continuous, multi\-patient simulators that grade time\-sensitive treatment actions alongside diagnosis[Luo et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib1)\. Our work targets the resulting long\-horizon execution gap in such environments\.

### 2\.2Agent Harnesses

We use*harness*to refer to the non\-parametric scaffolding around an LLM, including prompts, tools, control flow, and inter\-agent coordination, that determines how a fixed model executes a task\. The dominant pattern equips agents with external tools and skill libraries: ReAct interleaves tool calls with chain\-of\-thought reasoning[Yao et al\. \(2023\)](https://arxiv.org/html/2609.13543#bib.bib11), Toolformer fine\-tunes the model to invoke APIs autonomously[Schick et al\. \(2023\)](https://arxiv.org/html/2609.13543#bib.bib12), and Voyager introduces a learned, growable skill library for embodied agents[Wang et al\. \(2024\)](https://arxiv.org/html/2609.13543#bib.bib13)\. These approaches treat skills as task\-level procedures discovered or invoked during execution, and rely on the underlying model’s parametric memory for domain knowledge\. Our skills component takes a different design point: rather than learning skills end\-to\-end or growing them online, we curate a small, condition\-indexed library of high\-stakes clinical knowledge \(regimen templates, dosage tables, and condition\-to\-action mappings\) explicitly targeted at the incomplete\-treatment\-selection failure mode identified in rollout traces\.

### 2\.3Prompt Optimization and Self\-Evolving Harnesses

A growing line of work treats the prompt as a learnable artifact\. APE searches over instructions with an LLM as the search operator[Zhou et al\. \(2023\)](https://arxiv.org/html/2609.13543#bib.bib14); OPRO frames the LLM as a black\-box optimizer over textual prompts[Yang et al\. \(2024\)](https://arxiv.org/html/2609.13543#bib.bib15); DSPy jointly optimizes prompts and program structure against task metrics[Khattab et al\. \(2024\)](https://arxiv.org/html/2609.13543#bib.bib16)\. These methods operate on single\-prompt input/output pairs and optimize against a scalar metric\. A recent survey of agent externalization identifies*self\-evolving harnesses*, in which the orchestration logic itself is rewritten between rollouts, as an emerging direction[Zhou et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib18)\. Its closest concrete realization is Meta\-Harness[Lee et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib17), which runs an agentic proposer with filesystem access to the code, scores, and execution traces of all prior candidates\. Asclepius shares this proposer\-over\-traces design, but evolves a natural\-language operating*manual*rather than harness code, over multi\-hour, multi\-patient clinical rollouts with structured per\-dimension feedback \(diagnosis, critical actions, timeliness, disposition\) rather than a scalar reward\. In the clinical domain, EvoClinician\([He et al\., 2026](https://arxiv.org/html/2609.13543#bib.bib20)\)performs test\-time evolutionary learning to adapt diagnostic strategy within a multi\-turn encounter with a single patient\. Asclepius is complementary along both axes: the unit of adaptation is the cross\-shift operating manual rather than the within\-encounter policy, and the target is execution under multi\-patient load rather than diagnostic accuracy, with the underlying model and diagnosis pipeline held fixed\. HORIZON[Wang et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib19)characterizes long\-horizon agent failure across four domains by assigning each failed trajectory a single primary label from a seven\-category taxonomy; our characterization targets a complementary regime, operationalizing three failure modes as continuous per\-trace counters under multi\-objective, continuously\-graded execution and finding them*coupled*rather than orthogonal\.

## 3Preliminaries

### 3\.1Clinical Environment Simulator

We evaluate on the Clinical Environment Simulator \(CES\)[Luo et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib1), a turn\-based discrete\-event simulator of a multi\-hour, multi\-patient emergency\-department shift\. Patients arrive on staggered schedules with hidden deterioration thresholds; vitals evolve continuously, ordered tests return on simulated delays, and complications develop over time\. The agent interacts with the simulator through an EHR\-style interface and is graded by an LLM judge along four dimensions on a 1–5 scale: diagnosis, critical actions, timeliness, and disposition\. Full platform mechanics are in Appendix[A](https://arxiv.org/html/2609.13543#A1); rubric details are in Appendix[D](https://arxiv.org/html/2609.13543#A4)\.

### 3\.2Baseline Agents

We compare against two baseline configurations\.

Single\-LLM agent\.A frontier LLM is given a minimal task description and direct access to the simulator interface, with no hand\-crafted operating manual, no skill resources, and no subagents\. We report this configuration with three frontier models: Codex, Gemini, and Claude Code\. Claude Code appears in both baselines: as a raw model here, and as the underlying model of the framework below\.

Baseline agent framework\.A stronger baseline wraps Claude Code in an MCP\-based clinical tool interface and a hand\-crafted operating manual\. The interface exposes*observation*tools \(EHR content\),*action*tools \(orders, examinations, dispositions\), and*control*tools \(simulator lifecycle\); natural\-language order strings are resolved to catalog entries through a three\-stage matching pipeline\. The operating manual instructs the agent on triage, parallel workup, and disposition\. The agent maintains the full multi\-patient state in a single context window and has no specialized knowledge resources beyond parametric memory\. This configuration serves as the starting point for Asclepius; tool\-by\-tool details are in Appendix[B](https://arxiv.org/html/2609.13543#A2)\.

## 4Asclepius

Figure[1](https://arxiv.org/html/2609.13543#S1.F1)shows the architecture as two coupled loops: an*outer loop*that rewrites the operating manual between shifts, and an*inner loop*that runs within each shift and consists of a skills library and three subagents\.

### 4\.1Outer Loop: Self\-Evolving Harness

LetMtM\_\{t\}denote the operating manual at iterationt∈\{0,1,2,…\}t\\in\\\{0,1,2,\\dots\\\}andPPa proposer agent \(a separate agent with no access to the simulator, the underlying patient scenario definitions, or the held\-out batches\)\. At iterationtt,PPreceives a filesystemFtF\_\{t\}containing the full artifacts of all candidates produced so far\{M0,…,Mt\}\\\{M\_\{0\},\\dots,M\_\{t\}\\\}: each candidate’s manual, aggregate, and per\-batch scores, per\-patient judge rows, per\-turn action records, and LLM call audit logs\.PPnavigatesFtF\_\{t\}to inspect prior manuals and traces, identifies the current best\-scoring manual as the editing base, and outputs a revised manualMt\+1M\_\{t\+1\}\. We evaluateMt\+1M\_\{t\+1\}on the search batches, append its scores to the candidate record, and afterNNiterations select the best\-scoring manual on the search batches asM∗M^\{\*\}\.

The proposer’s system prompt requires it to: \(i\) base each edit on a failure pattern observed in≥3\\geq 3patients across multiple batches, \(ii\) avoid patient\- or batch\-specific hard\-coded knowledge, and \(iii\) produce substantive behavioral edits rather than cosmetic rewording\. Constraints \(i\)–\(ii\) guard against overfitting to the search batches; \(iii\) prevent wasted iterations\.

Search and held\-out validation\.We partition the patient batches into six search batches \(visible to the proposer\) and four held\-out batches \(never observed during iteration\)\. The held\-out split tests whether evolved manuals generalize beyond the search trajectories\. We use Claude Code with Claude Opus 4\.6\([Anthropic, 2026a](https://arxiv.org/html/2609.13543#bib.bib22)\)asPPand runN=10N=10iterations, selecting the manual with the best search\-batch overall score asM∗M^\{\*\}\(M7M\_\{7\}, denoted v7 in Section[6](https://arxiv.org/html/2609.13543#S6)\)\.

### 4\.2Inner Loop: Skills Library

The skills library provides on\-demand reference modules that the Main Agent retrieves during treatment selection\. The*clinical decision guide*contains a presentation\-to\-management table specifying disease\-specific drugs, doses, routes, and non\-pharmacologic interventions, together with a seven\-category treatment\-completeness framework\. The*clinical reasoning*module contains the treatment\-first principle, the “no naked diagnosis” rule, a pre\-disposition verification checklist, and a catalog of common failure patterns\. A separate*simulation mechanics*module covers turn structure, action lifecycle, and resource constraints\. Skills are loaded automatically when their description matches the current context and appended to the working context as structured reference\.

### 4\.3Inner Loop: Subagents

We externalize three recurring per\-turn decisions to subagents that run as separate LLM calls with isolated contexts: \(i\)*triage prioritization*selects the next patient to attend to; \(ii\)*diagnosis formation*synthesizes test results into a working diagnosis and disposition recommendation; \(iii\)*treatment completion*identifies missing tests, medications, and procedures for a diagnosed patient and ensures they are ordered\.

Subagent prompts contain only content extracted from the hand\-crafted operating manualM0M\_\{0\}and are held fixed across all configurations; no additional medical knowledge is introduced\. Gains from subagent decomposition therefore come from focused context per decision, structured\-output enforcement, and inter\-subagent coordination, rather than from added knowledge\.

Advisory vs\. acting variants\.We evaluate two variants\. In the*advisory*variant, subagents have no MCP tool access; they receive a structured snapshot of simulator state and return a textual recommendation that the Main Agent parses and executes\. In the*acting*variant, subagents have direct MCP access within their scope \(the triage subagent rooms patients, the diagnosis subagent commits dispositions and consults, the treatment subagent places missing orders\); the Main Agent serves as a coordinator that selects which subagent to invoke at each turn\. We use the acting variant as the default in Section[6](https://arxiv.org/html/2609.13543#S6)and report the other as an ablation\.

## 5Experimental Setup

### 5\.1The Clinical Simulator

We evaluate on the Clinical Environment Simulator \(CES\)\([Luo et al\., 2026](https://arxiv.org/html/2609.13543#bib.bib1)\), a long\-horizon multi\-patient emergency\-department simulator in which an LLM agent manages a multi\-hour shift: new patients arrive throughout the shift following a stochastic arrival process, existing patients evolve continuously \(vitals shifting, ordered tests returning on simulated delays, complications developing over time\), and the agent issues actions from a fixed set \(laboratory or imaging orders, medications, consults, dispositions\) through an EHR tool interface\.

Each shift is graded by an LLM judge with access to the gold patient trajectory along four dimensions on a 1 to 5 scale:*diagnosis*,*critical actions*,*timeliness*, and*disposition*\([Luo et al\., 2026](https://arxiv.org/html/2609.13543#bib.bib1)\)\. Higher scores reflect closer correspondence to expert\-graded reference trajectories\. We report per\-dimension means and the overall mean across all dimensions\. Our primary judge is GPT\-4\.1\([OpenAI, 2025b](https://arxiv.org/html/2609.13543#bib.bib26)\), held fixed across all configurations and used for every table unless stated otherwise\.

To test whether results depend on the choice of grader, we additionally re\-grade every patient under the baseline and full Asclepius configurations \(120 patients per configuration\) with four further judges spanning three model families: Claude Sonnet 4\.5\([Anthropic, 2025b](https://arxiv.org/html/2609.13543#bib.bib23)\), Claude Sonnet 5\([Anthropic, 2026b](https://arxiv.org/html/2609.13543#bib.bib24)\), Gemini 3\.1 Pro\([Google, 2026](https://arxiv.org/html/2609.13543#bib.bib29)\), and GPT\-5\([OpenAI, 2025c](https://arxiv.org/html/2609.13543#bib.bib27)\)\. Each validation judge receives the identical rubric, answer keys, and trajectory records as the primary judge; only the grading model changes\. The component ablations are graded by the primary judge alone\. Results are reported in Table[3](https://arxiv.org/html/2609.13543#S5.T3)\.

Full simulator mechanics, baseline agent configuration details, and evaluation rubrics are provided in Appendices[A](https://arxiv.org/html/2609.13543#A1),[B](https://arxiv.org/html/2609.13543#A2), and[D](https://arxiv.org/html/2609.13543#A4), respectively\.

### 5\.2Patient Batches and Evaluation Protocol

Each patient batch contains twelve patients with varying acuities, arrival times, and underlying diagnoses\. We use ten batches in total, partitioned into six*search*batches used during harness evolution and four*held\-out validation*batches reserved for final evaluation\. The validation batches are sampled from the same distribution as the search batches but are never observed by the proposer agent\. Each configuration is run once on every search and validation batch, and per\-dimension scores are averaged across batches\.

### 5\.3Configurations

We compare two primary configurations and several ablations:

- •Raw single\-agent: a frontier model invoked with only the simulator tool interface and a minimal task description, with no hand\-crafted operating manual, no skills, no subagents, and no harness evolution\. We report three frontier\-model variants under this configuration: Codex, Gemini, and Claude Code\.
- •Baseline agent configuration: the strongest agent baseline we have on CES, comprising a single frontier LLM, the simulator tool interface, and a hand\-crafted operating manual \(Appendix[B](https://arxiv.org/html/2609.13543#A2)\)\. We re\-implement this configuration as our starting point and report it under identical conditions to Asclepius\.
- •Asclepius: our full system \(Section[4](https://arxiv.org/html/2609.13543#S4)\), wrapping the baseline configuration with the self\-evolving harness, skills, and subagents\.

For component ablations we evaluate Asclepius variants that add the three components one at a time on top of the baseline configuration\.

### 5\.4Implementation Details

We use Claude Code, with Claude Opus 4\.6\([Anthropic, 2026a](https://arxiv.org/html/2609.13543#bib.bib22)\)as the underlying model, for both the main agent and the proposer agent\. The harness is evolved for ten iterations over the six search batches; we report the results on both search and validation\.

Table 1:Main results across all ten patient batches\.Mean scores \(out of 5\) for the three raw single\-agent variants, the baseline agent framework, and Asclepius \(harness \+ skills \+ acting subagents\)\. Scores pool the six search batches, over which the operating manual was evolved, with the four held\-out batches\. Significance for theΔ\\Deltarow is from a patient\-level mixed\-effects model, BH\-FDR corrected across all 20 system\-by\-dimension tests \(\*:p<0\.05p<0\.05, \*\*:p<0\.01p<0\.01, \*\*\*:p<0\.001p<0\.001, ns: not significant\)\. Bold marks the best score in each column\.Table 2:Held\-out significance \(batches 7–10,n=48n=48\)\.Mixed\-effects mean differences of Asclepius vs\. the baseline framework, fit on the held\-out batches alone\. Diagnosis and disposition are unchanged \(both ns\)\.Table 3:Multi\-judge validation\.Δ\\Delta= Asclepius−\-Baseline under each of the five judges\. \(\*:p<0\.05p<0\.05, \*\*:p<0\.01p<0\.01, \*\*\*:p<0\.001p<0\.001, ns: not significant\)\.![Refer to caption](https://arxiv.org/html/2609.13543v1/figures/Figure2.png)Figure 2:Component ablation across all ten patient batches\.Per\-dimension scores for the baseline framework, baseline \+ skills, baseline \+ subagents, baseline \+ harness, and Asclepius \(Full\)\. The baseline,\+S\+S, and\+S​A\+SAuse the hand\-crafted manualM0M\_\{0\};\+H\+Hand full Asclepius use the evolved manualM∗M^\{\*\}\.

## 6Results

We report aggregate scores across all ten patient batches \(120 patients total\)\. Significance for each comparison against the baseline framework comes from a patient\-level mixed\-effects model with random effects for patient and batch, withpp\-values BH\-FDR[Benjamini and Hochberg \(1995\)](https://arxiv.org/html/2609.13543#bib.bib30)corrected across the 20 system\-by\-dimension tests\. Search\-vs\-held\-out generalization, per\-batch breakdowns, and the full significance table are in Appendix[E](https://arxiv.org/html/2609.13543#A5)\.

### 6\.1Main Results

Table[1](https://arxiv.org/html/2609.13543#S5.T1)compares Asclepius against three raw single\-agent variants and the baseline agent framework\. The raw single\-agents span a wide range \(Gemini 2\.06 to Claude Code 3\.52 overall\), and even the strongest raw variant trails the baseline framework by 0\.28 overall, attributing substantial value to the MCP tool interface and hand\-crafted operating manual alone\. The execution\-gap asymmetry highlighted in Section[1](https://arxiv.org/html/2609.13543#S1)appears already on the raw Claude Code variant \(Dx 4\.22, Disp 4\.53 vs\. CA 2\.69, Timeliness 2\.65\) and persists on the baseline framework \(Dx 4\.39, Disp 4\.52 vs\. CA 2\.94, Timeliness 3\.34\), so the gap is a property of how current agents handle multi\-patient execution rather than an artifact of the framework\. Asclepius closes this asymmetry on top of the baseline, raising overall by\+0\.30\+0\.30\(p<0\.01p<0\.01\), critical actions by\+0\.73\+0\.73\(\+25%\+25\\%,p<0\.001p<0\.001\), and timeliness by\+0\.45\+0\.45\(\+13%\+13\\%,p<0\.01p<0\.01\); diagnosis and disposition stay within±0\.03\\pm 0\.03of the baseline \(p\>0\.8p\>0\.8\), confirming that the gains come from execution rather than from changes in clinical judgment\.

These aggregates pool the six search batches, over which the operating manual was evolved, with the four held\-out batches\. Refitting the model on the held\-out batches alone \(n=48n=48\) leaves critical actions significant \(\+0\.625\+0\.625,\+22%\+22\\%,p=0\.024p=0\.024\), with timeliness \(\+0\.312\+0\.312\) and overall \(\+0\.193\+0\.193\) positive but not significant \(Table[2](https://arxiv.org/html/2609.13543#S5.T2)\); we treat the held\-out critical\-actions result as our primary quantitative claim\. Re\-grading under the four validation judges reproduces the same pattern on every judge \(Table[3](https://arxiv.org/html/2609.13543#S5.T3)\), including one that shares a model family with neither the evaluated agent nor the Patient Engine, and three of the four report a larger timeliness gain than the primary judge, so the reported improvement is not specific to GPT\-4\.1\.

### 6\.2Component Ablations

Figure[2](https://arxiv.org/html/2609.13543#S5.F2)adds the three components \(harness, skills, acting subagents\) one at a time on top of the baseline framework \(mixed\-effects differences and BH\-FDR\-corrected significance markers in Table[8](https://arxiv.org/html/2609.13543#A5.T8), Appendix[E\.3](https://arxiv.org/html/2609.13543#A5.SS3)\)\. Each component yields a positive mean shift, but with different magnitudes\. Pooled across all ten batches, the self\-evolvingharnessshows the largest single\-component gain \(\+0\.18\+0\.18overall,\+0\.38\+0\.38critical actions,\+0\.34\+0\.34timeliness\), and is the only single\-component variant whose gains on critical actions and timeliness clear the BH\-FDR\-corrected significance threshold\.Skillscontributes the second\-largest gain \(\+0\.12\+0\.12overall,\+0\.22\+0\.22on critical actions and timeliness\), andsubagents alonethe smallest \(\+0\.04\+0\.04overall\); neither clears the significance threshold on any individual dimension\. This ordering reverses on the held\-out batches, where\+H\+Hreaches 3\.82 against a baseline of 3\.78 and\+S​A\+SAbecomes the strongest single component at 3\.97 \(Table[5](https://arxiv.org/html/2609.13543#A5.T5)\)\.

The three components compose: the full stack reaches4\.104\.10overall, with critical actions and timeliness exceeding any single\-component variant by≥0\.10\\geq 0\.10points, indicating the components are not substitutable for one another\. Holding skills and subagents fixed and varying only the operating manual, the evolved manual adds\+0\.12\+0\.12overall over S\+SA on all batches \(p=0\.36p=0\.36\) and\+0\.02\+0\.02on held\-out \(p=0\.88p=0\.88; Appendix[E\.4](https://arxiv.org/html/2609.13543#A5.SS4)\), so its marginal contribution on top of the architectural components is not statistically separable\. Diagnosis and disposition remain within±0\.11\\pm 0\.11of the baseline across all variants and never reach significance, consistent with the components being designed to repair execution rather than knowledge\.

![Refer to caption](https://arxiv.org/html/2609.13543v1/figures/Figure3.png)Figure 3:Per\-component reduction of targeted failure modes\.Three panels: \(left\) naked\-disposition rate early \(hours 0–2\) versus late \(hours 4–6\); \(middle\) zero\-treatment rate among correctly\-diagnosed patients; \(right\) timeliness gap between non\-severe and severe patients\. Each compares the baseline, single\-component variants \(\+S\+S,\+S​A\+SA,\+H\+H\), and full Asclepius\. Lower is better\.
### 6\.3Failure\-Mode Analysis

Figure[3](https://arxiv.org/html/2609.13543#S6.F3)reports the three per\-trace failure\-mode counters introduced in Section[1](https://arxiv.org/html/2609.13543#S1), giving a per\-trace view of*where*each component intervenes and complementing the aggregate scores; Table[10](https://arxiv.org/html/2609.13543#A5.T10)in Appendix[E\.5](https://arxiv.org/html/2609.13543#A5.SS5)lists the exact per\-configuration values\. Each failure mode is operationalized with one concrete per\-trace counter\. Instruction adherence can lapse in many ways and admits no single objective measure, so*Drift*tracks one frequent, unambiguous violation: the*naked disposition*, a patient dispositioned with≤1\\leq 1prior treatment order, in violation of the manual’s treatment\-first rule\. We report its rate early \(hours 0–2\) versus late \(hours 4–6\), so the rise captures decay over the shift\. The comparison is within\-subject: both configurations see the same twelve patients in the same arrival order on every batch, so any difference in case mix or queue depth between early and late hours applies equally to both, and a rise in one system but not the other cannot be attributed to shifting task difficulty\.*Incomplete*is the zero\-treatment rate: the fraction of correctly\-diagnosed patients given no disease\-specific treatment\.*Neglect*is the timeliness equity gap: the mean\-timeliness difference between non\-severe and severe patients, where severe is defined by the judge’s holistic encounter rating on the baseline trace \(≤2\\leq 2,n=20n=20; Appendix[E\.5](https://arxiv.org/html/2609.13543#A5.SS5)\)\.

Each component reduces the counter for the failure mode it targets: the harness slows instruction\-adherence drift, cutting the early→\\tolate rise in naked dispositions from\+41\+41pp to\+32\+32pp; the skills library lowers the zero\-treatment rate from 19% to 14%; and the subagents narrow the timeliness equity gap from 1\.13 to 0\.98\. The effects do not decompose cleanly across components, however: empirically, the skills library and the subagents reduce drift even further than the harness does \(to\+18\+18and\+16\+16pp, versus\+32\+32pp\), and the harness in turn attains the lowest zero\-treatment rate of any single component \(7%\)\. This pattern is consistent with the three failure modes acting as a*coupled bottleneck*: each component spills onto multiple counters, and no single component closes the execution gap on its own\.

The decisive reductions on all three counters appear only when the components are combined\. Full Asclepius nearly eliminates the drift \(a\+2\+2pp early→\\tolate rise, far below any single\-component variant; 16%→\\to57% for the baseline versus 10%→\\to12% for Asclepius on the same patients\), collapses the timeliness equity gap from 1\.13 to 0\.23, and roughly halves the zero\-treatment rate \(19% to 9%\)\. Taken together, the per\-trace counters corroborate the aggregate ablation: the three components are individually partial and jointly sufficient, and the coupled bottleneck only releases when the self\-evolving harness, the skills library, and the isolated subagents act together\.

![Refer to caption](https://arxiv.org/html/2609.13543v1/figures/Figure4.png)Figure 4:Self\-evolving harness search trajectory\.Best\-so\-far mean scores on the six search batches across 10 proposer iterations, one line per CES dimension\. v7 is selected on overall search\-batch score and used in Asclepius; v6–v8 form a high\-performing plateau\.
### 6\.4Harness Search Trajectory

Figure[4](https://arxiv.org/html/2609.13543#S6.F4)shows the best\-so\-far search\-batch scores across the 10 proposer iterations, and Table[11](https://arxiv.org/html/2609.13543#A5.T11)in Appendix[E\.6](https://arxiv.org/html/2609.13543#A5.SS6)reports the raw per\-iteration numbers\. The trajectory is non\-monotone: critical actions and timeliness rise quickly from v0 to v2, plateau through v4–v5, and reach a high\-performing plateau at v6–v8\. We select v7 as the final manual on the basis of overall search\-batch score\. Iterations v9–v10 regress on diagnosis or disposition\. Inspecting the diff from v0 to v7 reveals two classes of edit: \(i\) a procedural edit adding a pre\-disposition treatment\-enumeration routine tied to the seven treatment categories, and \(ii\) knowledge edits to the clinical content of the manual, with 13 new presentation rows \(e\.g\., pulmonary embolism\) and roughly 7 enriched management entries, such as correcting prophylactic to therapeutic anticoagulation dosing\. These additions are presentation\-level decision rules, not scenario\-specific answers\. To reduce overfitting to the search batches, the proposer is instructed to add only guidance that generalizes across cases and to act only on failure patterns recurring across multiple patients and batches\. The full configuration retains a\+0\.19\+0\.19overall gain over the baseline on the held\-out batches the proposer never observes \(Table[5](https://arxiv.org/html/2609.13543#A5.T5)\)\.

At inference time, full Asclepius requires1\.21×1\.21\\timesthe LLM calls of the baseline \(433 vs\. 359 per batch\) and1\.52×1\.52\\timesthe wall\-clock \(118 vs\. 78 minutes per batch\), so the gains do not come from a large increase in inference compute\. Harness evolution is a separate, one\-time development cost of 10 iterations, each a proposer step \(≈\\approx10 minutes\) followed by an evaluation rollout over the six search batches, producing a frozen manual that is reused unchanged at inference\.

Table 4:Subagent advisory vs\. acting variants\.Mean scores under the full Asclepius configuration\. The acting variant gives each subagent direct MCP access within its scope; the advisory variant returns structured recommendations parsed by the Main Agent\. We use the acting variant as the default\.
### 6\.5Subagent Variants

Table[4](https://arxiv.org/html/2609.13543#S6.T4)compares the advisory and acting variants of the subagent decomposition \(Section[4\.3](https://arxiv.org/html/2609.13543#S4.SS3)\) under the full Asclepius configuration\. The acting variant outperforms the advisory variant by \+0\.10 overall, \+0\.26 on critical actions, and \+0\.10 on timeliness, with no degradation on diagnosis or disposition\. The gap is driven by tool\-call efficiency: when subagents place their own orders, the Main Agent avoids paraphrasing and re\-resolving long structured recommendations, removing a class of vocabulary\-resolution failures observed in the advisory traces\. We use the acting variant as the default elsewhere in the paper\.

### 6\.6Qualitative Analysis

We walk through a representative Asclepius rollout in which the inner\-loop subagents intercept a failure that the baseline framework commits on the same patient\. Patient B1\-P10 presents in atrial flutter with rapid ventricular response \(RVR\); both systems reach the correct diagnosis\. On the baseline trace, the Main Agent orders heparin and fluids but fails to initiate a rate\-control strategy or a durable oral anticoagulation plan before disposition\. The patient remains in rapid atrial flutter and worsens over subsequent turns, with escalating tachycardia and hypoxemia \(heart rate 135→\\to153, SpO295%→\\to86%; critical\-actions score 1/5\), consistent with persistent RVR in the setting of cardiopulmonary vulnerability and empiric fluid administration\. On the Asclepius trace, the treatment\-completion subagent flags the missing rate\-control strategy and post\-ED anticoagulation plan\. The Main Agent adds diltiazem, after excluding hypotension, pre\-excitation, and decompensated systolic heart failure, and starts apixaban for indicated stroke prevention\. The patient’s ventricular rate improves, oxygenation stabilizes, and the patient is admitted to cardiology with a complete regimen \(critical\-actions score 1→\\to5\)\. More traces with analogous interceptions are reported in Appendix[F](https://arxiv.org/html/2609.13543#A6)\.

## 7Conclusion

We studied the execution gap in long\-horizon, multi\-task agentic settings, using the Clinical Environment Simulator as a testbed in which the phenomenon is measurable: current agents reach the correct diagnosis in most cases yet fail to deliver complete and timely critical actions under sustained shift\-long load\. We operationalized this gap with three per\-trace counters that target long\-horizon failure modes \(instruction\-adherence drift, treatment incompleteness, severity\-equity gap\), and instantiated an adaptive agent scaffolding, Asclepius, with a component for each: a self\-evolving harness that rewrites the operating manual between shifts, an externalized clinical skills library, and three isolated subagents that partition per\-turn decisions across the patient queue\. On held\-out batches never observed during harness evolution, Asclepius improves critical\-action correctness by 22% over a strong baseline framework while preserving diagnostic accuracy, with consistent gains across five LLM judges from three model families; on the full ten\-batch set, improvements reach 25% on critical actions and 13% on timeliness\. The strongest empirical finding is structural: the three failure modes form a coupled bottleneck, in that no single component closes the execution gap on its own, and the decisive reductions on all three counters appear only when the self\-evolving harness, the skills library, and the isolated subagents act together\. For clinical AI evaluation, the result argues for benchmarks that measure sustained, complete, and equitable execution across a patient queue, not diagnostic accuracy on isolated cases\.

## 8Limitations

#### LLM\-based grading\.

All four CES dimensions are scored by LLM judges\. Our primary judge \(GPT\-4\.1; Appendix[D](https://arxiv.org/html/2609.13543#A4)\) shares a model family with the simulator’s Patient Engine, so correlated errors between generated physiology and its grading are a structural concern\. Re\-grading under four further judges across three model families reproduces the gains \(Table[3](https://arxiv.org/html/2609.13543#S5.T3)\), indicating the improvement is not specific to one grader\. Absolute scores nonetheless inherit any miscalibration or ceiling effects shared across LLM graders, and we did not run an expert\-agreement study on Asclepius traces\. Gains should be read as improvements under automated grading, not as validated clinical outcomes\.

#### Single rollout per configuration and batch\.

Each configuration is run once on each batch, so every reported cell rests on one trajectory of a stochastic agent in a stochastic simulator\. Our mixed\-effects model treats patient and batch as random effects and therefore quantifies patient\- and batch\-level heterogeneity, not run\-to\-run execution variance\. Two observations bound how much this drives the results: full Asclepius improves critical actions on all ten batches individually \(smallest margin\+0\.17\+0\.17; Appendix[E](https://arxiv.org/html/2609.13543#A5)\), and the same direction and approximate magnitude are reproduced by four independent judges\. Neither substitutes for repeated execution, and per\-batch overall differences range from−0\.12\-0\.12to\+0\.71\+0\.71, so individual per\-batch values and the ranking within the v6–v8 plateau should not be over\-read\.

#### Simulation\-to\-reality gap\.

CES is a simulator: patient cards are derived from de\-identified records and physiology is LLM\-generated\. Performance on CES is a proxy for, not evidence of, real clinical competence\. We evaluate a single default ED configuration \(four beds, six nurses, one physician, six\-hour shift\); robustness to other shift lengths, staffing ratios, and acuity mixes is untested, as is generalization beyond the English\-language, guideline\-dosing regime encoded in the skills library\. All results use a single backbone \(Claude Opus 4\.6\) in a single environment; whether the components help weaker base models, or transfer to other long\-horizon multi\-objective environments, is untested\.

#### Cost and scalability of harness evolution\.

The self\-evolving outer loop requires repeated multi\-hour rollouts across six search batches per iteration plus a separate proposer agent, making it computationally expensive relative to a hand\-written manual\. Selection uses only six search batches; although the held\-out batches confirm generalization, the overall gain compresses from\+0\.36\+0\.36on search to\+0\.19\+0\.19on held\-out, indicating some adaptation to the search trajectories\. The skills library is hand\-curated rather than learned, so its coverage is bounded by manual effort and may omit rare presentations\.

## References

- Anthropic \(2025a\)AnthropicClaude code\.Note:[https://www\.anthropic\.com/claude\-code](https://www.anthropic.com/claude-code)Cited by:[§1](https://arxiv.org/html/2609.13543#S1.p2.1)\.
- Anthropic \(2025b\)AnthropicIntroducing Claude Sonnet 4\.5\.Note:[https://www\.anthropic\.com/news/claude\-sonnet\-4\-5](https://www.anthropic.com/news/claude-sonnet-4-5)Released September 29, 2025Cited by:[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p3.1)\.
- Anthropic \(2026a\)AnthropicClaude opus 4\.6 system card\.Note:[https://www\.anthropic\.com/claude\-opus\-4\-6\-system\-card](https://www.anthropic.com/claude-opus-4-6-system-card)Released February 5, 2026Cited by:[§4\.1](https://arxiv.org/html/2609.13543#S4.SS1.p3.1),[§5\.4](https://arxiv.org/html/2609.13543#S5.SS4.p1.1)\.
- Anthropic \(2026b\)AnthropicIntroducing Claude Sonnet 5\.Note:[https://www\.anthropic\.com/news/claude\-sonnet\-5](https://www.anthropic.com/news/claude-sonnet-5)Released June 30, 2026Cited by:[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p3.1)\.
- Benjamini and Hochberg \(1995\)Y\. Benjamini and Y\. HochbergControlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal Statistical Society: Series B \(Methodological\)57\(1\),pp\. 289–300\.Cited by:[§6](https://arxiv.org/html/2609.13543#S6.p1.1)\.
- Fanet al\.\(2025\)Z\. Fan, L\. Wei, J\. Tang, W\. Chen, W\. Siyuan, Z\. Wei, and F\. HuangAI hospital: benchmarking large language models in a multi\-agent medical interaction simulator\.InProceedings of the 31st International Conference on Computational Linguistics,Abu Dhabi, UAE,pp\. 10183–10213\.External Links:[Link](https://aclanthology.org/2025.coling-main.680/)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Gaoet al\.\(2025\)S\. Gao, R\. Zhu, Z\. Kong, A\. Noori, X\. Su, C\. Ginder, T\. Tsiligkaridis, and M\. ZitnikTxAgent: an AI agent for therapeutic reasoning across a universe of tools\.External Links:2503\.10970,[Link](https://arxiv.org/abs/2503.10970)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Google \(2025\)GoogleGemini 3\.Note:[https://blog\.google/products/gemini/gemini\-3/](https://blog.google/products/gemini/gemini-3/)Released November 18, 2025Cited by:[§1](https://arxiv.org/html/2609.13543#S1.p2.1)\.
- Google \(2026\)GoogleGemini 3\.1 Pro: a smarter model for your most complex tasks\.Note:[https://blog\.google/innovation\-and\-ai/models\-and\-research/gemini\-models/gemini\-3\-1\-pro/](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-pro/)Released February 19, 2026Cited by:[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p3.1)\.
- Heet al\.\(2026\)Y\. He, J\. Liu, Z\. Hu, Y\. Chen, Y\. Liu, Y\. Sui, Y\. Li, N\. Chen, J\. Hu, B\. Hooi, X\. Xu, and J\. BianEvoClinician: a self\-evolving agent for multi\-turn medical diagnosis via test\-time evolutionary learning\.arXiv preprint arXiv:2601\.22964\.Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.
- Jianget al\.\(2025\)Y\. Jiang, K\. C\. Black, G\. Geng, D\. Park, J\. Zou, A\. Y\. Ng, and J\. H\. ChenMedAgentBench: a virtual EHR environment to benchmark medical LLM agents\.NEJM AI2\(9\),pp\. AIdbp2500144\.Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Khattabet al\.\(2024\)O\. Khattab, A\. Singhvi, P\. Maheshwari, Z\. Zhang, K\. Santhanam, S\. Vardhamanan, S\. Haq, A\. Sharma, T\. T\. Joshi, H\. Moazam, H\. Miller, M\. Zaharia, and C\. PottsDSPy: compiling declarative language model calls into self\-improving pipelines\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.
- Leeet al\.\(2026\)Y\. Lee, R\. Nair, Q\. Zhang, K\. Lee, O\. Khattab, and C\. FinnMeta\-harness: end\-to\-end optimization of model harnesses\.External Links:2603\.28052,[Link](https://arxiv.org/abs/2603.28052)Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.
- Liet al\.\(2024\)J\. Li, Y\. Lai, W\. Li, J\. Ren, M\. Zhang, X\. Kang, S\. Wang, P\. Li, Y\. Zhang, W\. Ma, and Y\. LiuAgent hospital: a simulacrum of hospital with evolvable medical agents\.External Links:2405\.02957,[Link](https://arxiv.org/abs/2405.02957)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Luoet al\.\(2026\)L\. Luo, S\. E\. Kim, X\. Zhang, J\. M\. Kernbach, R\. Kenia, J\. N\. Acosta, L\. A\. Nathanson, A\. D\. Haimovich, A\. Rodman, E\. Goh, J\. H\. Chen, N\. H\. Shah, D\. A\. Kim, J\. Zou, F\. Mahmood, J\. N\. Kather, M\. Lungren, V\. Natarajan, E\. J\. Topol, and P\. RajpurkarA clinical environment simulator for dynamic AI evaluation\.Nature Medicine32\(3\),pp\. 820–827\.Note:PerspectiveExternal Links:[Document](https://dx.doi.org/10.1038/s41591-026-04252-6)Cited by:[Appendix A](https://arxiv.org/html/2609.13543#A1.p1.1),[§D\.1](https://arxiv.org/html/2609.13543#A4.SS1.p1.1),[Appendix D](https://arxiv.org/html/2609.13543#A4.p1.1),[§1](https://arxiv.org/html/2609.13543#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.13543#S3.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p1.1),[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p2.1)\.
- McDuffet al\.\(2025\)D\. McDuff, M\. Schaekermann, T\. Tu, A\. Palepu, A\. Wang, J\. Garrison, K\. Singhal, Y\. Sharma, S\. Azizi, K\. Kulkarni, L\. Hou, Y\. Cheng, Y\. Liu, S\. S\. Mahdavi, S\. Prakash, A\. Pathak, C\. Semturs, S\. Patel, D\. R\. Webster, E\. Dominowska, J\. Gottweis, J\. Barral, K\. Chou, G\. S\. Corrado, Y\. Matias, J\. Sunshine, A\. Karthikesalingam, and V\. NatarajanTowards accurate differential diagnosis with large language models\.Nature642\(8067\),pp\. 451–457\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-08869-4)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Noriet al\.\(2025\)H\. Nori, M\. Daswani, C\. Kelly, S\. Lundberg, M\. T\. Ribeiro, M\. Wilson, X\. Liu, V\. Sounderajah, J\. Carlson, M\. P\. Lungren, B\. Gross, P\. Hames, M\. Suleyman, D\. King, and E\. HorvitzSequential diagnosis with language models\.External Links:2506\.22405,[Link](https://arxiv.org/abs/2506.22405)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- OpenAI \(2025a\)OpenAIIntroducing Codex\.Note:[https://openai\.com/index/introducing\-codex/](https://openai.com/index/introducing-codex/)Cited by:[§1](https://arxiv.org/html/2609.13543#S1.p2.1)\.
- OpenAI \(2025b\)OpenAIIntroducing GPT\-4\.1 in the API\.Note:[https://openai\.com/index/gpt\-4\-1/](https://openai.com/index/gpt-4-1/)Released April 14, 2025Cited by:[§A\.5](https://arxiv.org/html/2609.13543#A1.SS5.p1.1),[Appendix D](https://arxiv.org/html/2609.13543#A4.p1.1),[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p2.1)\.
- OpenAI \(2025c\)OpenAIIntroducing GPT\-5\.Note:[https://openai\.com/index/introducing\-gpt\-5/](https://openai.com/index/introducing-gpt-5/)Released August 7, 2025Cited by:[§5\.1](https://arxiv.org/html/2609.13543#S5.SS1.p3.1)\.
- Schicket al\.\(2023\)T\. Schick, J\. Dwivedi\-Yu, R\. Dessì, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. ScialomToolformer: language models can teach themselves to use tools\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2\.2](https://arxiv.org/html/2609.13543#S2.SS2.p1.1)\.
- Schmidgallet al\.\(2024\)S\. Schmidgall, R\. Ziaei, C\. Harris, E\. Reis, J\. Jopling, and M\. MoorAgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments\.External Links:2405\.07960,[Link](https://arxiv.org/abs/2405.07960)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Tuet al\.\(2025\)T\. Tu, M\. Schaekermann, A\. Palepu, K\. Saab, J\. Freyberg, R\. Tanno, A\. Wang, B\. Li, M\. Amin, Y\. Cheng, E\. Vedadi, N\. Tomašev, S\. Azizi, K\. Singhal, L\. Hou, A\. Webson, K\. Kulkarni, S\. S\. Mahdavi, C\. Semturs, J\. Gottweis, J\. Barral, K\. Chou, G\. S\. Corrado, Y\. Matias, A\. Karthikesalingam, and V\. NatarajanTowards conversational diagnostic artificial intelligence\.Nature642\(8067\),pp\. 442–450\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-08866-7)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Wanget al\.\(2024\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.Transactions on Machine Learning Research \(TMLR\)\.Cited by:[§2\.2](https://arxiv.org/html/2609.13543#S2.SS2.p1.1)\.
- Wanget al\.\(2026\)X\. J\. Wang, H\. Bai, Y\. Sun, H\. Wang, S\. Zhang, W\. Hu, M\. Schroder, B\. Mutlu, D\. Song, and R\. D\. NowakThe long\-horizon task mirage? diagnosing where and why agentic systems break\.arXiv preprint arXiv:2604\.11978\.Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.
- Yanget al\.\(2024\)C\. Yang, X\. Wang, Y\. Lu, H\. Liu, Q\. V\. Le, D\. Zhou, and X\. ChenLarge language models as optimizers\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.
- Yaoet al\.\(2023\)S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. CaoReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.2](https://arxiv.org/html/2609.13543#S2.SS2.p1.1)\.
- Zhaoet al\.\(2026\)W\. Zhao, C\. Wu, Y\. Fan, P\. Qiu, X\. Zhang, Y\. Sun, X\. Zhou, S\. Zhang, Y\. Peng, Y\. Wang, X\. Sun, Y\. Zhang, Y\. Yu, K\. Sun, and W\. XieAn agentic system for rare disease diagnosis with traceable reasoning\.Nature651\(8106\),pp\. 775–784\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-10097-9)Cited by:[§2\.1](https://arxiv.org/html/2609.13543#S2.SS1.p1.1)\.
- Zhouet al\.\(2026\)C\. Zhou, H\. Chai, W\. Chen, Z\. Guo, R\. Shan, Y\. Song, T\. Xu, Y\. Yang, A\. Yu, W\. Zhang, C\. Zheng, J\. Zhu, Z\. Zheng, Z\. Zhang, X\. Lou, C\. Zhang, Z\. Fu, J\. Wang, W\. Liu, J\. Lin, and W\. ZhangExternalization in LLM agents: a unified review of memory, skills, protocols and harness engineering\.arXiv preprint arXiv:2604\.08224\.Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.
- Zhouet al\.\(2023\)Y\. Zhou, A\. I\. Muresanu, Z\. Han, K\. Paster, S\. Pitis, H\. Chan, and J\. BaLarge language models are human\-level prompt engineers\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2\.3](https://arxiv.org/html/2609.13543#S2.SS3.p1.1)\.

## Appendix ACES Platform Details

This appendix expands on the brief description of the Clinical Environment Simulator \(CES\)[Luo et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib1)in Section[5\.1](https://arxiv.org/html/2609.13543#S5.SS1)and supplies the simulator mechanics that our results in Section[6](https://arxiv.org/html/2609.13543#S6)depend on\.

### A\.1Temporal Model and Shift Structure

CES is a turn\-based discrete\-event simulator in which one turn corresponds to five simulated minutes\. A default scenario spans a six\-hour emergency\-department shift, that is, 72 turns\. At each turn the engine executes a fixed\-order pipeline: check for scheduled patient arrivals, auto\-assign beds from the queue, process all pending actions, and update patient vital signs\. Clinical actions submitted during a turn are queued as pending actions and resolved in batch at the corresponding turn boundary, ensuring deterministic ordering and preventing within\-turn causal inconsistencies\.

A scenario comprises twelve patients with staggered arrival times across the shift, simulating realistic emergency\-department patient flow\. The staggered arrivals force the agent to manage time across concurrent encounters: excessive time spent on a single patient causes other patients to deteriorate past their hidden thresholds\. Patient cards are derived from de\-identified emergency\-department visit records and converted into structured YAML scenarios containing demographics, a triage\-categorized presentation, a patient summary with separated subjective history and objective examination findings, clinical background, initial vital signs, a ground\-truth diagnosis, the list of required critical actions, the acceptable disposition options, and a hidden deterioration threshold\.

### A\.2Patient Severity and Deterioration

Each patient is assigned one of three severity classes that govern temporal dynamics:*mundane*patients carry no physiological risk and do not deteriorate;*dynamic moderate*patients \(e\.g\., sepsis without shock, community pneumonia\) face gradual deterioration risk with a hidden threshold randomly sampled in 90–120 minutes;*dynamic severe*patients \(e\.g\., hemorrhagic shock, intracranial hemorrhage\) face rapid deterioration risk with a hidden threshold in 30–60 minutes\. Prior to the threshold, vital signs oscillate around the patient’s initial set\-point with severity\-scaled corridors and a mean\-reversion factor; after the threshold elapses without the required critical actions, pathological drift begins through three severity phases \(early, moderate, severe\)\. When the required critical actions are completed, drift ceases and vital signs recover along pharmacokinetic time courses appropriate to the administered interventions\. Only dynamic severe patients can trigger a code\-blue event when vitals breach boundary values; code blue consumes 20–30 minutes of physician time and forces immediate transfer to the intensive care unit with reduced encounter scores\.

### A\.3Action Costs and Resource Constraints

Action durations are calibrated against expert physician consensus\. A comprehensive history\-and\-physical examination consumes 10 minutes \(2 turns\) of physician time; a follow\-up or targeted H&P requires 5 minutes \(1 turn\)\. Triage assessment is auto\-generated on arrival\. Medication administration requires 5 minutes \(1 turn\) of nurse time for standard medications; transfusions and large\-volume infusions require 20 minutes \(4 turns\)\. Diagnostic test latencies vary by modality: laboratory studies require 15–125 minutes \(e\.g\., basic metabolic panel 30 min, coagulation studies 45 min, PCR panels 125 min\), with culture results requiring approximately 24 hours; imaging studies require 5–75 minutes depending on modality \(bedside ultrasound 5 min, plain radiography 10–15 min, magnetic resonance imaging 40 min\)\. Procedural durations vary similarly \(arterial line 15 min; chest tube 20 min; central venous catheter 25 min\)\.

The default emergency\-department configuration comprises 4 beds, 6 nurses, 1 physician \(the agent under evaluation\), 14 specialist consultants spanning 9 medical and 5 surgical specialties, one unit each of CT/MRI/X\-ray/ultrasound/ECG, and 1 laboratory analyzer\. Beds, staff, imaging equipment, and laboratory throughput are managed as capacitated queues with priority\-based allocation by clinical acuity \(HIGH before MEDIUM before LOW\)\. Certain actions \(e\.g\., oxygen administration, intravenous medications\) require a bed assignment, and the finite bed supply creates realistic bottlenecks\. Specialist consults have stochastic response latency \(2–4 turns\) and consultation duration \(2–12 turns\); consultants receive only the information available to the agent at the time of request and never the ground\-truth diagnosis\.

### A\.4Medication and Treatment Effects

To prevent unrealistic vital\-sign discontinuities, treatment effects are categorized by onset profile\. Immediate\-onset interventions \(5–10 min\) include intravenous fluid resuscitation and supplemental oxygen\. Delayed\-effect medications \(15–30 min\) include antibiotics \(halting physiological deterioration\) and bronchodilators\. Gradual\-effect agents \(≥30\\geq 30min\) include antipyretics\. This temporal fidelity penalizes delayed treatment initiation and rewards early, appropriate intervention\.

### A\.5Engine Architecture

CES is organized around two decoupled engines that communicate through a typed message bus\. The Patient Engine maintains each patient’s clinical state \(vital signs, active diagnoses, medication effects, deterioration trajectory\) and delegates physiological reasoning to a fixed Patient LLM \(GPT\-4\.1\([OpenAI, 2025b](https://arxiv.org/html/2609.13543#bib.bib26)\)\) that generates vital\-sign updates, history\-and\-physical narratives, and diagnostic results at each turn\. The Hospital Engine manages capacitated resources, tracking allocation, queuing, and release through priority\-based scheduling\. A turn\-based orchestrator coordinates the two engines: at turn start it broadcasts a turn\-start signal; the Hospital Engine processes resource allocations and order completions and notifies the Patient Engine of newly available results; the Patient Engine incorporates results, processes arrivals, and advances physiological trajectories; at turn end the orchestrator emits a turn\-end signal with all generated events\. Agents interact with the simulator through a browser\-based interface served by a session\-based REST API, with a patient tracking board, individual patient charts, an order catalog of 6,888 items, and real\-time event notifications\.

## Appendix BBaseline Agent Configuration Details

Our baseline configuration wraps the underlying LLM with a CES\-customized harness exposing 17 clinically meaningful tools organized in an Observe–Act–Control trichotomy through the Model Context Protocol \(MCP\)\. The same harness is the starting point for Asclepius; the components in Section[4](https://arxiv.org/html/2609.13543#S4)are added on top\.

### B\.1Tool Categories

#### Observation tools \(5\)\.

Return structured representations of the simulator state to the agent rather than requiring it to interpret raw rendered content: the patient tracking board, individual patient summaries \(chief complaint, vitals, key findings\), order histories, complete patient records, and event logs\. The agent receives organized vital signs, pending orders, and diagnostic results directly, mirroring the information a physician extracts at a glance from an electronic\-health\-record display\.

#### Action tools \(8\)\.

Encapsulate multi\-step clinical workflows as single tool calls: order tests, order medications, order procedures, order consultation, discharge, admit, transfer, and perform history\-and\-physical examination\. Each action tool accepts natural clinical language \(e\.g\., “CBC”, “morphine”, “CT head”\) and resolves it to exact catalog entries through the vocabulary resolution pipeline described below\.

#### Control tools \(4\)\.

Manage simulator lifecycle: session initialization \(automatic login and scenario loading\), single\-step turn advancement, run\-until\-event turn advancement, and blocking\-state handling for code blue, rapid\-response\-team, and physician\-busy states\.

### B\.2Medical Vocabulary Resolution

Free\-text order strings are resolved to catalog entries through a three\-stage pipeline\.*Stage 1*performs deterministic lookup against 464 hand\-curated abbreviation mappings \(e\.g\., “CBC”→\\to“complete blood count”\)\.*Stage 2*, invoked when Stage 1 fails, embeds the query with a sentence\-transformer model and searches a vector database containing 5,823 medications, 943 diagnostic tests, and 122 procedures, accepting the best semantic match above a cosine\-similarity threshold of 0\.3\.*Stage 3*performs interface\-level verification: the resolved name is entered into the simulator’s search interface, and fuzzy string matching \(weighted\-ratio scorer, cutoff 60\) is applied against the displayed options to confirm that the ordered item exists in the current simulator state\. This pipeline handles the vocabulary mismatch between clinical shorthand and the simulator’s catalog of 6,888 orderable items\.

### B\.3Pre\-Action Guards and the Operating Manual

Tool descriptions and pre\-action guards encode workflow knowledge directly into the interface: patients must be assigned a bed before medications can be ordered, intravenous formulations are recommended over oral for acute presentations, triage is enforced by acuity \(HIGH before MEDIUM before LOW\), and disposition tools automatically surface the current order set so the agent can audit care delivered before submission\. The operating manual instructs the agent on triage, parallel workup, treatment, and disposition, including completeness rules \(“treatment\-first principle”, “no naked diagnosis”, a seven\-category treatment\-completeness sweep, and a pre\-disposition verification checklist\)\. In our experiments the manual is the starting pointM0M\_\{0\}from which the self\-evolving harness \(Section[4\.1](https://arxiv.org/html/2609.13543#S4.SS1)\) iterates\.

## Appendix CAsclepius Component Artifacts

The complete prompts, skill modules, and subagent definitions for every configuration we evaluate are released at[https://github\.com/rajpurkarlab/Asclepius](https://github.com/rajpurkarlab/Asclepius)\. The repository is organized by configuration, mirroring the ablations in Section[5\.3](https://arxiv.org/html/2609.13543#S5.SS3): the baseline framework, each single\-component variant, and full Asclepius in both acting and advisory forms\. Each configuration directory contains the main agent prompt \(the hand\-crafted operating manualM0M\_\{0\}for the baseline, or the evolved manualM∗M^\{\*\}for the harness and full configurations\); where applicable, the skills library, the acting or advisory subagent definitions \(triage prioritizer, diagnostician, and treatment checker\), and the proposer prompt that evolves the operating manual \(Section[4\.1](https://arxiv.org/html/2609.13543#S4.SS1)\)\.

## Appendix DEvaluation Design

Each shift is graded by a fixed LLM judge \(GPT\-4\.1\([OpenAI, 2025b](https://arxiv.org/html/2609.13543#bib.bib26)\)\) configured with structured prompts, with the four\-dimension rubric described below[Luo et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib1)\. The judge receives a ground\-truth answer key derived from the patient card \(ground\-truth diagnosis, severity class, required critical actions, tests required for diagnosis, acceptable disposition options, hidden deterioration threshold\) together with a comprehensive record of the agent’s actions \(submitted diagnosis, medications administered with timestamps, procedures performed, diagnostic tests ordered, disposition decision and discharge instructions, and any code\-blue or rapid\-response\-team activations\)\. For clinical context the judge also receives the patient’s initial presentation, the full vital\-sign trajectory, a chronological log of clinical events, returned test results, and any specialist consultations requested\. The judge never receives the Patient Engine’s internal prompts\. Performance is scored on a 1–5 scale along four dimensions\.

#### Diagnosis\.

Evaluated under a clinical\-equivalence matching framework\. A score of 5 denotes an exact or clinically equivalent match to the ground truth; 4 reflects a clinically defensible alternative diagnosis \(e\.g\., pericarditis for a ground truth of myocarditis\); 3 indicates correct organ\-system identification without sufficient specificity, or a multi\-part diagnosis containing incorrect extra information; 2 corresponds to an incorrect but non\-dangerous diagnosis; 1 is reserved for cases in which a life\-threatening diagnosis is missed and the proposed management would cause harm\. A non\-specific diagnosis when specificity is clinically required is a scoring failure \(e\.g\., “chest pain” for ST\-elevation myocardial infarction\)\.

#### Critical Actions\.

Scored on the fraction of required actions completed\. The judge applies clinical\-equivalence matching for synonymous terminology \(e\.g\., acetaminophen≡\\equivparacetamol; normal saline bolus≡\\equiv0\.9% NaCl infusion\)\. Medication dosing is evaluated against standard\-of\-care ranges rather than requiring exact values\.

#### Timeliness\.

Assesses whether critical actions are completed before the hidden deterioration threshold\. If required critical actions remain incomplete, the maximum achievable timeliness score is capped at 3 regardless of speed\. Activation of a rapid response team with a correct working diagnosis for a deteriorating patient may earn a 2; a code\-blue event in the absence of any prior critical actions receives a 1\.

#### Disposition\.

Assesses whether the final placement \(discharge, admission to a specific service, or transfer, e\.g\., to the operating room or catheterization laboratory\) matches the acceptable disposition options for the case\.

For every encounter we report an*overall*score, defined as the unweighted mean of the four per\-dimension scores above; this is the overall score used in every table and figure\. The LLM judge additionally produces a single*holistic encounter rating*\. We use it for identifying the severe subset in the failure\-mode analysis \(Appendix[E\.5](https://arxiv.org/html/2609.13543#A5.SS5)\)\. Encounters are scored independently per patient and then aggregated to per\-batch and per\-configuration means; we use these aggregates throughout Section[6](https://arxiv.org/html/2609.13543#S6)\.

### D\.1Judge Reliability

The CES judge carries prior validation reported by[Luo et al\. \(2026\)](https://arxiv.org/html/2609.13543#bib.bib1)\. On the simulator’s plausibility checks, an internal\-medicine physician blinded to the automated results, patient cards, and Patient Engine prompts reviewed 2,802 matched items across 78 scenarios, agreeing with the automated checker at 92\.5% raw agreement \(Gwet’s AC1 = 0\.92, 95% CI 0\.91–0\.93\)\. On the four\-dimension scoring rubric we use, a blinded rater scored 48 encounters drawn from both physician and agent sessions, yielding Gwet’s AC2 of 0\.93 \(diagnosis\), 0\.86 \(critical actions\), and 0\.86 \(timeliness\), with an overall AC2 of 0\.88\. This indicates that the judge’s scores on the dimensions central to our claims track expert judgment closely, and that the checker carries a conservative bias toward flagging implausible outputs rather than rewarding them\.

Two points remain specific to our setting\. First, these agreement studies validate the judge on CES trajectories broadly, not on the Asclepius configurations evaluated here; we did not run an expert\-agreement study on our own traces\. Second, the primary judge and the simulator’s Patient Engine are the same model family \(GPT\-4\.1\)\. We address this directly in Section[6\.1](https://arxiv.org/html/2609.13543#S6.SS1)by re\-grading all 120 patients under the baseline and full configurations with four further judges spanning three model families; all five reproduce the reported gains \(Table[3](https://arxiv.org/html/2609.13543#S5.T3)\)\.

### D\.2Patient Batches

Each patient batch comprises twelve patients with varying acuities, arrival times, and underlying diagnoses, sampled to span the acuity and pathophysiology spectrum\. We use ten batches in total\. Batches 1–6 are*search batches*, used by the self\-evolving harness as the substrate over which the proposer iterates\. Batches 7–10 are*held\-out validation batches*, never seen by the proposer at any iteration\. Each configuration is run once per batch, with per\-dimension scores averaged across the batches in each split\.

## Appendix EPer\-Batch, Search\-Split, and Significance Results

This appendix expands the aggregate results in Section[6](https://arxiv.org/html/2609.13543#S6)by reporting \(i\) the search\-batch \(1–6\) and held\-out \(7–10\) subsets, \(ii\) the full per\-batch overall scores, and \(iii\) the full BH\-FDR\-corrected significance table from the mixed\-effects model\.

### E\.1Search\-Batch and Held\-Out Means

Table[5](https://arxiv.org/html/2609.13543#A5.T5)reports the means for the five real configurations on the search batches \(1–6\) and on the held\-out batches \(7–10\) separately\. Full Asclepius is the best configuration on both splits, with held\-out overall3\.973\.97vs\. baseline3\.783\.78\(Δ=\+0\.19\\Delta=\+0\.19\)\. The overall gain compresses from\+0\.36\+0\.36on search to\+0\.19\+0\.19on held\-out, a roughly 47% compression\. We read this compression as informative rather than as a failure of generalization: the proposer adapts in part to recurring patterns of the search batches, and the persistent\+0\.19\+0\.19on never\-observed batches gives a lower bound on the component of the gain that is generalizable\. Within the single\-component variants, the harness ranks first on the search subset but loses that lead on the held\-out subset \(where\+S​A\+SAranks highest among single components\), consistent with the same picture: the self\-evolving harness carries some search\-set\-specific signal that does not fully transfer, while the architectural components \(\+S​A\+SA, skills\) carry gains that transfer more uniformly\.

Table 5:Search\-batch and held\-out means\.Mean overall and per\-dimension scores on batches 1–6 \(search\) and 7–10 \(held\-out\) for the five ablation configurations\. Overall is the unweighted mean of the four sub\-dimension scores\. Bold marks the best score in each column\.
### E\.2Per\-Batch Overall Scores

Table[6](https://arxiv.org/html/2609.13543#A5.T6)reports the overall score for the baseline framework and full Asclepius on each of the ten batches individually, so per\-batch variance can be inspected\. Full Asclepius beats the baseline on overall in 8 of 10 batches; the two exceptions, B3 and B6, are small \(−0\.12\-0\.12and−0\.06\-0\.06\)\. On critical actions, full Asclepius beats the baseline on every batch \(smallest per\-batch margin\+0\.17\+0\.17\)\.

Table 6:Per\-batch overall scores\.Overall \(mean across the four dimensions\) for the baseline framework and full Asclepius on each of the ten batches\.
### E\.3Component Ablation Means and Significance Tests

Table[7](https://arxiv.org/html/2609.13543#A5.T7)reports the absolute per\-dimension means for each ablation configuration; Table[8](https://arxiv.org/html/2609.13543#A5.T8)gives the corresponding mixed\-effects mean differences and BH\-FDR\-correctedpp\-values\.

We fit a patient\-level mixed\-effects modely∼system\+\(1\|patient\)\+\(1\|batch\)y\\sim\\mathrm\{system\}\+\(1\\,\|\\,\\mathrm\{patient\}\)\+\(1\\,\|\\,\\mathrm\{batch\}\)with the baseline framework as the reference, Satterthwaite degrees of freedom, and BH\-FDR correction across all 20 system\-by\-dimension comparisons\. Table[8](https://arxiv.org/html/2609.13543#A5.T8)reports the corrected results\. Three patterns stand out\. First, full Asclepius is significant on every execution\-related dimension \(overall, critical actions, timeliness\) atp<0\.01p<0\.01or stricter\. Second, among single\-component variants, the harness is the only one to reach the corrected threshold on any dimension \(critical actions and timeliness\); skills and subagents alone reach significance on none\. Third, diagnosis and disposition do not significantly change for any configuration, consistent with all systems near the judge’s ceiling\.

Table 7:Component ablation across all ten patient batches\.Per\-dimension mean scores when adding the harness \(HH\), skills \(SS\), and acting subagents \(S​ASA\) on top of the baseline framework\. Significance markers \(vs\. baseline, BH\-FDR corrected\) follow the overall and execution columns\. Bold marks the best score in each column\.ConfigurationOverallDxCATimeDispBaseline3\.804\.392\.943\.344\.52\+SS3\.92ns4\.473\.163\.564\.51\+S​ASA3\.84ns4\.433\.113\.424\.41\+HH3\.98ns4\.463\.32 \*3\.68 \*4\.44Full Asclepius4\.10\*\*4\.383\.67\*\*\*3\.79\*\*4\.54Table 8:Significance versus baseline \(mixed\-effects model, BH\-FDR corrected\)\.For each outcome and configuration, we report the estimated mean difference vs\. the baseline framework on the 1–5 scale and the BH\-FDR\-corrected significance \(\*:p<0\.05p<0\.05, \*\*:p<0\.01p<0\.01, \*\*\*:p<0\.001p<0\.001, ns: not significant\)\. Overall is the unweighted mean of the four sub\-dimension scores\.
### E\.4Marginal Contribution of the Evolved Manual

The ablation configurations in Figure[2](https://arxiv.org/html/2609.13543#S5.F2)vary both the manual version and the set of architectural components\. To isolate the manual, we compare S\+SA \(skills and acting subagents onM0M\_\{0\}\) against full Asclepius \(the same components onM∗M^\{\*\}\); this contrast holds skills and subagents fixed and varies only the operating manual\. Table[9](https://arxiv.org/html/2609.13543#A5.T9)reports the mixed\-effects mean differences\. On all batches the evolved manual adds\+0\.12\+0\.12overall,\+0\.23\+0\.23on critical actions, and\+0\.08\+0\.08on timeliness; on the held\-out batches the overall difference falls to\+0\.02\+0\.02\. Neither split is statistically separable from zero, so we scope our claims to the system level rather than attributing an independently significant effect to the harness\.

Table 9:Marginal contribution of the evolved manual\.Mixed\-effects mean differences of full Asclepius \(M∗M^\{\*\}\) versus S\+SA \(M0M\_\{0\}\), holding skills and acting subagents fixed\.
### E\.5Per\-Trace Failure\-Mode Counters

Table[10](https://arxiv.org/html/2609.13543#A5.T10)gives the per\-configuration values of the three per\-trace failure\-mode counters visualized in Figure[3](https://arxiv.org/html/2609.13543#S6.F3); see Section[6\.3](https://arxiv.org/html/2609.13543#S6.SS3)for the analysis\.

Table 10:Per\-trace failure\-mode counters across all ten patient batches\.Lower is better for all three counters\.*Drift*: decay in the agent’s adherence to the operating manual over the shift, measured as the early \(hours 0–2\)→\\tolate \(hours 4–6\) rise in*naked dispositions*\(dispositions placed with≤1\\leq 1prior treatment order, violating the manual’s treatment\-first rule\)\.*Incomplete*: fraction of correctly\-diagnosed patients given no disease\-specific treatment at all\.*Neglect*: timeliness equity gap, the mean\-timeliness difference between non\-severe and severe patients \(severe==baseline holistic encounter rating≤2\\leq 2,n=20n\{=\}20\)\.
### E\.6Harness Search Trajectory \(Per\-Iteration Scores\)

Table[11](https://arxiv.org/html/2609.13543#A5.T11)gives the per\-iteration search\-batch scores visualized in Figure[4](https://arxiv.org/html/2609.13543#S6.F4); see Section[6\.4](https://arxiv.org/html/2609.13543#S6.SS4)for the analysis\.

Table 11:Harness search trajectory on search batches \(1–6\)\.Mean scores across the six search batches \(out of 5\) for each harness iteration\. Overall is the unweighted mean of the four sub\-dimension scores\. v7 is selected as the final manual based on overall search\-batch performance\. Bold marks the best score in each column\.

## Appendix FAdditional Qualitative Traces

We report three additional rollouts in which the inner\-loop components intercept a baseline failure on the same patient\. The first two illustrate the treatment\-completion subagent; the third illustrates the diagnosis\-formation subagent, whose broader differential prevents an anchoring error from cascading into a missed treatment\. The trauma case is drawn from a held\-out batch, giving a concrete instance of the same interception on a patient the proposer never observed during harness search\.

#### Trauma resuscitation completion\.

Patient B10\-P07 arrives with an open leg fracture that has also damaged a major artery, and is in shock from blood loss\. Both systems correctly identify the fracture and the arterial injury and route the patient to the operating room, so the two traces differ only in how completely the resuscitation is carried out\. On the baseline trace, the Main Agent gives antibiotics, pain control, and splinting but omits tetanus prophylaxis, early TXA, and hemostatic resuscitation despite hemorrhagic shock\. As a result, the arterial\-injury blood loss remains undertreated and the patient’s blood pressure continues to fall before operative transfer\. On the Asclepius trace, the treatment\-completion subagent cross\-references the open\-fracture trauma regimen in the skills library and flags the three gaps, which the Main Agent then closes, completing five of the six required interventions before transfer to the operating room\.

#### Sepsis antibiotic completion\.

Patient B6\-P04 presents with abdominal pain, fever, tachycardia, marked leukocytosis, pyuria, and acute kidney injury on a background of advanced bladder cancer, consistent with probable complicated urinary\-source sepsis\. On the baseline trace, the Main Agent treats the malignancy and metabolic abnormalities but admits the patient without antimicrobial therapy\. On the Asclepius trace, the treatment\-completion subagent flags the absent empiric antibiotic, and the Main Agent initiates broad\-spectrum coverage before admission\.

#### Atypical NSTEMI recognition\.

Patient B4\-P03 presents with altered mental status, abnormal kidney labs, and a history of cardiovascular disease\. On the baseline trace, the Main Agent anchored on the abnormal kidney values, attributed the altered mental status to renal/metabolic encephalopathy, and admitted the patient to a general medical ward without ordering an ECG or troponin\. Because a cardiac cause was never considered, the NSTEMI remained undetected\. On the Asclepius trace, the diagnosis\-formation subagent kept ACS in the differential despite the atypical presentation\. The Main Agent ordered an ECG and serial troponins, which showed ischemic ECG changes and elevated/dynamic troponin, supporting NSTEMI\. The patient then received cardiology\-directed ACS treatment and was admitted to a monitored cardiology unit\.

Similar Articles

@sethkarten: https://x.com/sethkarten/status/2072034978112889328

X AI KOLs Following

Continual Harness is a reset-free, self-improving agentic harness that achieves 20.54% on ARC-AGI-3 at a cost of $774 by storing memories, reusing skills, and refining its prompt, outperforming prior baselines like Hermes and OpenClaw with greater efficiency.

SkillHarness: Harnessing Safe Skills for Computer-Use Agents

Hugging Face Daily Papers

SkillHarness is a framework that enables computer-use agents to safely learn and execute skills in dynamic environments by incorporating safety constraints and adaptive skill selection mechanisms, reducing unsafe rates by 57.1%.