Safe to Stop? Risk-Constrained Stopping for Sequential Clinical Diagnosis Agents
Summary
This paper introduces Cros, a risk-constrained stopping layer for sequential clinical diagnosis agents that ensures finite-sample guarantees for diagnostic error and coverage, with evaluation on a benchmark derived from MIMIC.
View Cached Full Text
Cached at: 09/11/26, 08:39 AM
# Safe to Stop? Risk-Constrained Stoppingfor Sequential Clinical Diagnosis Agents
Source: [https://arxiv.org/html/2609.09678](https://arxiv.org/html/2609.09678)
Yuexin WuAffiliation:Department of Computer ScienceAffiliation:University of MemphisEmail:[ywu10@memphis\.edu](mailto:)Vasile RusAffiliation:Department of Computer ScienceAffiliation:University of MemphisEmail:[vrus@memphis\.edu](mailto:)
###### Abstract
Clinical diagnosis agents must decide not only what test to request next, but also when to diagnose or defer\. Existing agent benchmarks largely evaluate accuracy after fixed or unconstrained interaction, leaving autonomous stopping reliability implicit\. We presentCros, a risk\-constrained stopping layer combining state\-wise error ranking, policy design on disjoint development splits, and LTT\-style exact tests of selective diagnostic error and minimum autonomous coverage for complete sequential policies\. Its finite\-sample guarantee requires the candidate family, testing rule, and any randomization to be frozen before calibration labels are accessed\. On a 1,834\-episode MIMIC\-derived abdominal\-pain benchmark, the full ranker achieves exploratory state\-error AUROC 0\.853, compared with 0\.715 for maximum class probability and 0\.552 for the backbone’s native stop score\. On the previously viewed 367\-episode evaluation split, analytically averaging over the frozenCros\-Mix weights yields 16\.9% selective error at 78\.8% coverage, cost 5\.57, and 0\.68 tests, versus 30\.8% error at 100% coverage, cost 8\.14, and 1\.53 tests under native stopping\. Forced continuation is non\-monotone: error is 28\.3% with HPI alone and 34\.3% after full workup\. However, the uniform\-weight mixture ablation is cheaper on this viewed split despite missing the locked development margins, andCros\-Mix nominally satisfies the joint criterion in only 6 of 20 development resplits\. Because evaluation labels were inspected during earlier development, these findings provide exploratory feasibility and audit evidence, not a confirmatory safety certificate\.
## 1Introduction
Language\-model agents can generate differential diagnoses, request tests, read new findings, and revise their hypotheses\. Recent systems and benchmarks have made this interaction increasingly realistic\([Hager et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib1);[Schmidgall et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib3);[Liu et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib4);[Bani\-Harouni et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib2)\)\. Yet a central decision remains weakly specified: when should the agent stop acquiring information and commit to an autonomous diagnosis? Stopping too early can miss a consequential disease; stopping too late wastes tests and may expose patients to avoidable procedures\. Uncalibrated confidence thresholds provide no finite\-sample joint risk–coverage guarantee, while unconstrained empirical cost minimization can overfit the selection sample\([Angelopoulos et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib8);[Laufer\-Goldshtein et al\., 2023](https://arxiv.org/html/2609.09678#bib.bib16)\)\.
We formulate sequential diagnosis as selective, risk\-constrained stopping\. At each stage the agent observes the history, proposes a diagnosis and next test, and the stopping layer either accepts the diagnosis, continues the shared acquisition trajectory, or defers\. The desired policy minimizes resource cost subject to two population constraints: conditional diagnostic error among autonomous decisions is at mostα\\alpha, and autonomous coverage is at leastγ\\gamma\. This differs from ordinary selective classification\([Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.09678#bib.bib10);[Geifman and El\-Yaniv, 2019](https://arxiv.org/html/2609.09678#bib.bib11)\): prediction quality and the information state both evolve over an agent trajectory\.
Our method,Cros\(Clinical Risk\-constrained Optimal Stopping\), separates representation, design, and calibration\. Here, “optimal” means cost\-minimizing within a finite, selection\-frozen family of threshold\-and\-horizon policies;Crosdoes not solve unrestricted Bellman optimal stopping or optimize the backbone’s test\-acquisition policy\. First, a risk ranker scores the likelihood that the backbone’s current diagnosis is wrong\. Second, a selection split freezes the candidate family and, optionally, a randomized mixture chosen to minimize expected cost\. Third, a prospectively fresh calibration split can supply exact binomialpp\-values for error and coverage\. Multiple testing procedures from learn\-then\-test \(LTT\)\([Angelopoulos et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib8)\)convert candidate\-wise tests into a finite\-sample guarantee\. The stopping layer is backbone\-agnostic and auditable: each action, observation, risk score, stop/continue/defer outcome, and calibration statistic is logged\. In the present study, previously opened labels mean that the same calculations are exploratory calibration checks, not a realized prospective guarantee\.
We make four contributions:
1. 1\.We adapt and operationalize LTT\-style joint testing of selective diagnostic error and minimum autonomous coverage for complete sequential stopping policies, obtaining an exact finite\-sample guarantee when deterministic or episode\-wise randomized policies are frozen before calibration\.
2. 2\.We instantiate episode\-wise randomized mixtures of deterministic stopping policies; Proposition[1](https://arxiv.org/html/2609.09678#Thmproposition1)shows that an optimal basic feasible mixture requires at most three component policies and remains compatible with exact calibration when randomized independently across episodes\.
3. 3\.We construct an auditable retrospective benchmark with 1,834 MIMIC\-derived ED episodes, 12 nonuniform\-cost actions, recorded\-result missingness, and a common\-backbone comparison that isolates stopping from diagnosis and test\-proposal quality\.
4. 4\.We report both favorable and negative evidence: the full ranker out\-ranks simple scores but weakens under diagnosis\-language masking, forced continuation is non\-monotone, optimized mixing does not dominate uniform mixing on evaluation, 20\-resplit feasibility is unstable, and aggregate control does not ensure subgroup safety\.
## 2Related Work
#### Clinical decision agents\.
MIMIC\-CDM evaluates LLMs on sequential clinical decision making and exposes limitations in diagnostic reasoning and tool use\([Hager et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib1)\)\. AgentClinic provides a multimodal simulated clinical environment\([Schmidgall et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib3)\); MedChain emphasizes interactive, sequential clinical benchmarking\([Liu et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib4)\); and DxChain uses panoramic profiling and adversarial debate\([Lv et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib5)\)\. LA\-CDM trains hypothesis and decision agents with reinforcement learning to choose tests and update diagnoses\([Bani\-Harouni et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib2)\)\. These systems motivate the same acquisition loop as our benchmark\. Our goal is complementary: we hold the diagnostic backbone and its action trajectory fixed when possible, then evaluate whether a stopping controller can provide a testable population guarantee\.
#### Resource\-aware sequential diagnosis\.
MAI\-DxO evaluates diagnostic accuracy jointly with the cost of adaptively requested tests\([Nori et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib25)\)\. ACTMED uses Bayesian experimental design to select the next test\([Ruhrberg Estévez et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib26)\), cost\-sensitive reinforcement learning learns adaptive test\-panel policies\([Yu et al\., 2023](https://arxiv.org/html/2609.09678#bib.bib30)\), and latent diagnostic trajectory learning trains planning and diagnostic agents to acquire evidence along learned paths\([Shen et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib27)\)\. These methods can change which tests are selected and therefore change the trajectory\.Crosis not an end\-to\-end acquisition method: its common\-path design holds backbone diagnoses, test proposals, and forced\-continuation trajectories fixed to isolate whether the controller stops, continues with the backbone\-proposed test, or defers\.
#### Selective prediction and risk control\.
Selective classifiers abstain on uncertain examples to trade coverage for conditional error\([Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.09678#bib.bib10);[Geifman and El\-Yaniv, 2019](https://arxiv.org/html/2609.09678#bib.bib11)\)\. Distribution\-free risk\-controlling prediction sets and conformal risk control extend calibration beyond marginal coverage\([Bates et al\., 2021](https://arxiv.org/html/2609.09678#bib.bib12);[Angelopoulos et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib9)\)\. LTT turns risk constraints into hypothesis tests and controls the probability of selecting an invalid procedure from a finite family\([Angelopoulos et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib8)\)\. Geometry\-Calibrated Conformal Abstention gives finite\-sample guarantees for both participation and correctness of emitted open\-ended language\-model responses\([Xu et al\., 2026b](https://arxiv.org/html/2609.09678#bib.bib23)\), while SCoRE uses conformal e\-values to control a general bounded risk among selected outputs\([Bai and Jin, 2026](https://arxiv.org/html/2609.09678#bib.bib24)\)\. Thus neither selective risk control nor participation guarantees are new in isolation\. We adapt these ideas to a stateful clinical stopping problem whose complete, frozen policy is jointly tested for conditional diagnostic error and a lower autonomous\-coverage bound, with resource costs used for policy design and comparison rather than included in the validity claim\. The exact guarantee concerns the policy chosen before calibration, not the accuracy of the learned risk score\.
#### Clinical stopping and abstention\.
Uncertainty\-aware abstention has been studied for static medical\-text prediction\([Vazhentsev et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib29)\), while Safe\-Psych evaluates diagnose, clarify, and abstain decisions as psychiatric evidence is revealed incrementally\([Presacan et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib28)\)\. MediQ studies interactive clinical diagnosis in which a model refrains from diagnosing under insufficient information and asks follow\-up questions, but does not provide finite\-sample joint control of selective diagnostic error and autonomous coverage\([Li et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib31)\)\. Foo and Chang formulate staged clinical prediction as an expected\-loss optimal\-stopping problem with explicit decision and testing costs and Bellman recursion\([Foo and Chang, 2026](https://arxiv.org/html/2609.09678#bib.bib21)\)\. Their objective optimizes whether expected decision value justifies further testing; it does not provideCros’s exact joint test of conditional diagnostic error and minimum autonomous coverage\.Crosinstead freezes a complete sequential policy selected on disjoint development data and then subjects that policy to the joint test\. This distinction is methodological, not a claim that Bellman stopping, interactive clinical diagnosis, or clinical abstention is new\.
#### Calibrated sequential stopping and acquisition\.
LTT\-style finite\-sample risk control is not new to this work\. Pareto Testing combines multi\-objective design and multiple testing\([Laufer\-Goldshtein et al\., 2023](https://arxiv.org/html/2609.09678#bib.bib16)\), while accumulated\-accuracy\-gap control gives distribution\-free stopping rules for early time classification\([Ringel et al\., 2024](https://arxiv.org/html/2609.09678#bib.bib17)\)\. MiCP allocates error budgets across turns of retrieval, tool\-use, and reasoning workflows, enabling adaptive early stopping with overall conformal coverage while reducing turns and inference cost\([Zhou et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib22)\)\. Its guarantee concerns coverage of the final prediction set in multi\-turn reasoning;Crosinstead tests selective diagnostic error and minimum autonomous coverage for a frozen stop–continue–defer controller on a common clinical trajectory\. Other recent work studies selective conformal risk control\([Xu et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib18)\), inference\-time reasoning under a compute budget\([Wang et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib19)\), and post\-acquisition recalibration when additional evidence can be requested\([Xu et al\., 2026a](https://arxiv.org/html/2609.09678#bib.bib20)\)\. These studies already establish important forms of sequential calibration, joint utility–risk design, or cost\-aware acquisition\. Our narrower contribution is their integration into sequential clinical diagnosis: a state\-wise clinical risk ranker, a lower autonomous\-coverage constraint, a common\-path MIMIC benchmark, and episode\-wise mixtures\. The sparsity of the mixture is a standard linear\-program consequence rather than a new optimization theorem\.
## 3Risk\-Constrained Sequential Diagnosis
### 3\.1Problem setup
An episodeZ=\(X0,Y,O1:H\)Z=\(X\_\{0\},Y,O\_\{1:H\}\)contains an initial presentationX0X\_\{0\}, a reference diagnosisY∈\{1,…,K\}Y\\in\\\{1,\\ldots,K\\\}, and potential recorded observations along a maximum horizonHH\. Here,HHis the maximum number of test\-acquisition stages available in an episode \(H=12H=12in our benchmark\)\. A policy\-specific horizonh≤Hh\\leq Hdetermines the latest stage at which that policy must stop or defer\. At stagett, the backbone has historySt=\(X0,A1:t,O1:t\)S\_\{t\}=\(X\_\{0\},A\_\{1:t\},O\_\{1:t\}\), produces a diagnosisY^t\\widehat\{Y\}\_\{t\}, and proposes the next testAt\+1∈𝒜A\_\{t\+1\}\\in\\mathcal\{A\}\. A stopping controllerπ\\pimaps the observed history and the backbone’s current proposal tostop,continue, ordefer\. Continuing reveals the next recorded result along the backbone\-proposed trajectory;Croscontrols whether that proposal is executed but never substitutes a different test\. This common\-path interaction is summarized in Figure[1](https://arxiv.org/html/2609.09678#S3.F1)\. LetTπT\_\{\\pi\}denote the terminal stage at which the controller either accepts a diagnosis or defers, and letDπ\(Z\)=1D\_\{\\pi\}\(Z\)=1if it returns an autonomous diagnosis andDπ\(Z\)=0D\_\{\\pi\}\(Z\)=0if it defers\. Define
ℛ\(π\)\\displaystyle\\mathcal\{R\}\(\\pi\)=Pr\(Y^Tπ≠Y∣Dπ=1\),\\displaystyle=\\Pr\\\!\\left\(\\widehat\{Y\}\_\{T\_\{\\pi\}\}\\neq Y\\mid D\_\{\\pi\}=1\\right\),\(1\)𝒞\(π\)\\displaystyle\\mathcal\{C\}\(\\pi\)=Pr\(Dπ=1\),\\displaystyle=\\Pr\(D\_\{\\pi\}=1\),\(2\)𝒥\(π\)\\displaystyle\\mathcal\{J\}\(\\pi\)=𝔼\[∑t<Tπc\(At\+1\)\+cdef\(1−Dπ\)\]\.\\displaystyle=\\mathbb\{E\}\\\!\\left\[\\sum\_\{t<T\_\{\\pi\}\}c\(A\_\{t\+1\}\)\+c\_\{\\mathrm\{def\}\}\(1\-D\_\{\\pi\}\)\\right\]\.\(3\)These quantities separate diagnostic reliability, autonomous participation, and resource use\. Specifically,ℛ\(π\)\\mathcal\{R\}\(\\pi\)is the diagnostic error rate among cases that the controller handles autonomously, rather than among all episodes\.𝒞\(π\)\\mathcal\{C\}\(\\pi\)is the population fraction receiving an autonomous diagnosis; the remaining fraction is deferred\. Finally,𝒥\(π\)\\mathcal\{J\}\(\\pi\)is the expected cumulative cost of requested actions, plus a downstream\-review penaltycdefc\_\{\\mathrm\{def\}\}whenever the case is deferred\.
We seek the least costly policy in the candidate familyΠ\\Piwhile requiring its selective diagnostic error to be at mostα\\alphaand its autonomous coverage to be at leastγ\\gamma:
minπ∈Π𝒥\(π\)s\.t\.ℛ\(π\)≤α,𝒞\(π\)≥γ\.\\min\_\{\\pi\\in\\Pi\}\\ \\mathcal\{J\}\(\\pi\)\\quad\\text\{s\.t\.\}\\quad\\mathcal\{R\}\(\\pi\)\\leq\\alpha,\\qquad\\mathcal\{C\}\(\\pi\)\\geq\\gamma\.\(4\)Thus,α\\alphaspecifies the maximum tolerated error rate among autonomous diagnoses, whereasγ\\gammaprevents the controller from achieving low error merely by deferring most cases\. Because conditional risk is undefined when𝒞\(π\)=0\\mathcal\{C\}\(\\pi\)=0, the positive coverage constraint also excludes the degenerate always\-defer policy\. The defer penalty affects policy design and cost comparisons, but it is not part of the subsequent binomial tests of risk and coverage\.
Initial presentationand triage→\\rightarrowBackbone proposesdiagnosis and test→\\rightarrowLogged resultupdates state→\\rightarrowCrosrisk check:stop, continue, or defer
Figure 1:A sequential episode\. The backbone generates the common forced\-continuation trajectory and proposes each test;Croscontrols whether to stop, continue with that proposal, or defer\. It does not select the test identity\.
### 3\.2Risk\-ranked stopping policies
We train an auxiliary estimatorrθ\(St\)∈\[0,1\]r\_\{\\theta\}\(S\_\{t\}\)\\in\[0,1\]for the eventY^t≠Y\\widehat\{Y\}\_\{t\}\\neq Y\. Features include theKKclass probabilities, maximum probability, probability margin, entropy, stage, fraction of missing results, cumulative resource cost, latency, and the backbone’s native stop score\. Training uses episode\-wise out\-of\-fold predictions so that multiple states from one patient never cross folds\.
For a horizonhhand thresholdτ\\tau, the deterministic policyπh,τ\\pi\_\{h,\\tau\}stops at the firstt≤ht\\leq hsatisfyingrθ\(St\)≤τr\_\{\\theta\}\(S\_\{t\}\)\\leq\\tauand otherwise defers athh\. A disjoint selection set freezes the estimator, a finite candidate listΠ0=\{π1,…,πL\}\\Pi\_\{0\}=\\\{\\pi\_\{1\},\\ldots,\\pi\_\{L\}\\\}, candidate order, and all design hyperparameters before calibration labels are accessed\. The learned ranker may be misspecified; validity below depends only on a fresh exchangeable calibration sample\.
### 3\.3Exact joint tests and multiplicity control
After a candidate policyπj\\pi\_\{j\}has been frozen, calibration asks whether it satisfies both population requirements: selective diagnostic error at mostα\\alphaand autonomous coverage at leastγ\\gamma\. Whenπj\\pi\_\{j\}is applied tonncalibration episodes, let
Mj=∑i=1nDπj\(Zi\),Ej=∑i=1nDπj\(Zi\)\{Y^Tπj,i≠Yi\}\.M\_\{j\}=\\sum\_\{i=1\}^\{n\}D\_\{\\pi\_\{j\}\}\(Z\_\{i\}\),\\qquad E\_\{j\}=\\sum\_\{i=1\}^\{n\}D\_\{\\pi\_\{j\}\}\(Z\_\{i\}\)\\mathbf\{1\}\\\!\\left\\\{\\widehat\{Y\}\_\{T\_\{\\pi\_\{j\}\},i\}\\neq Y\_\{i\}\\right\\\}\.\(5\)Here,MjM\_\{j\}is the number of episodes receiving an autonomous diagnosis, whileEjE\_\{j\}is the number of errors among those autonomous diagnoses\. Thus,Mj/nM\_\{j\}/nis the empirical autonomous coverage and, whenMj\>0M\_\{j\}\>0,Ej/MjE\_\{j\}/M\_\{j\}is the empirical selective diagnostic error\.
A candidate is invalid if either its population risk exceedsα\\alphaor its population coverage falls belowγ\\gamma\. We therefore test the union null
Hj:\{ℛ\(πj\)\>α\}∪\{𝒞\(πj\)<γ\}\.H\_\{j\}:\\\{\\mathcal\{R\}\(\\pi\_\{j\}\)\>\\alpha\\\}\\ \\cup\\ \\\{\\mathcal\{C\}\(\\pi\_\{j\}\)<\\gamma\\\}\.\(6\)RejectingHjH\_\{j\}requires evidence against both failure modes: the policy must have sufficiently few autonomous errors and sufficiently many autonomous diagnoses\.
LetFBin\(k,m,p\)F\_\{\\mathrm\{Bin\}\}\(k;m,p\)denote the probability that aBinomial\(m,p\)\\mathrm\{Binomial\}\(m,p\)random variable is at mostkk\. The corresponding one\-sided exact componentpp\-values are
pR,j=FBin\(Ej,Mj,α\),pC,j=1−FBin\(Mj−1,n,γ\)\.p\_\{R,j\}=F\_\{\\mathrm\{Bin\}\}\(E\_\{j\};M\_\{j\},\\alpha\),\\qquad p\_\{C,j\}=1\-F\_\{\\mathrm\{Bin\}\}\(M\_\{j\}\-1;n,\\gamma\)\.\(7\)The riskpp\-value becomes small when the policy makes unusually few errors relative to the boundary rateα\\alpha, whereas the coveragepp\-value becomes small when it autonomously diagnoses unusually many episodes relative to the boundary rateγ\\gamma\. We setpR,j=1p\_\{R,j\}=1whenMj=0M\_\{j\}=0, because an always\-defer policy provides no evidence about conditional diagnostic risk\. Because both constraints must be supported, the intersection\-union test uses
pj=max\(pR,j,pC,j\)\.p\_\{j\}=\\max\(p\_\{R,j\},p\_\{C,j\}\)\.\(8\)This combined value is small only when both componentpp\-values are small\. For a single pre\-frozen policy, no multiplicity adjustment is needed\. When several candidates are tested, we use fixed\-sequence LTT, Holm’s step\-down procedure\([Holm, 1979](https://arxiv.org/html/2609.09678#bib.bib13)\), or Bonferroni to control the probability of certifying any invalid candidate\. In a prospective study, a rejected candidate is certified under Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1); here, the same event is called an exploratory calibration pass because the labels are not untouched\.
###### Theorem 1\(Finite\-sample joint control\)\.
Assume the calibration episodes are i\.i\.d\. \(or exchangeable with the future population\), and the complete candidate policies and multiple\-testing rule are fixed independently of calibration outcomes\. Then eachpjp\_\{j\}in Eq\. \([8](https://arxiv.org/html/2609.09678#S3.E8)\) is super\-uniform underHjH\_\{j\}\.\. If the testing rule controls family\-wise error atδ\\delta, the probability that any certified policy violates eitherℛ\(π\)≤α\\mathcal\{R\}\(\\pi\)\\leq\\alphaor𝒞\(π\)≥γ\\mathcal\{C\}\(\\pi\)\\geq\\gammais at mostδ\\delta\.
The proof is in Appendix[A](https://arxiv.org/html/2609.09678#A1)\. The statement is finite\-sample and makes no assumption thatrθr\_\{\\theta\}is calibrated\.
### 3\.4Randomized sparse mixtures
A finite threshold\-and\-horizon grid may contain no single deterministic policy that achieves the desired risk–coverage trade\-off at minimum cost\. We therefore allow episode\-wise randomization over the frozen candidate familyΠ0=\{π1,…,πL\}\\Pi\_\{0\}=\\\{\\pi\_\{1\},\\ldots,\\pi\_\{L\}\\\}\. Letwjw\_\{j\}be the probability of selecting policyπj\\pi\_\{j\}, with weights in the probability simplexΔL:=\{w∈ℝ\+L:∑j=1Lwj=1\}\\Delta\_\{L\}:=\\\{w\\in\\mathbb\{R\}\_\{\+\}^\{L\}:\\sum\_\{j=1\}^\{L\}w\_\{j\}=1\\\}\. The resulting randomized policyπw\\pi\_\{w\}independently samplesJ∼Categorical\(w\)J\\sim\\operatorname\{Categorical\}\(w\)at the beginning of each episode and appliesπJ\\pi\_\{J\}throughout that episode\.
On the selection split, let𝒥^j\\widehat\{\\mathcal\{J\}\}\_\{j\}be the empirical mean cost ofπj\\pi\_\{j\},c^j=Pr^\(Dπj=1\)\\widehat\{c\}\_\{j\}=\\widehat\{\\Pr\}\(D\_\{\\pi\_\{j\}\}=1\)its empirical autonomous coverage, andq^j=Pr^\(Dπj=1,Y^≠Y\)\\widehat\{q\}\_\{j\}=\\widehat\{\\Pr\}\(D\_\{\\pi\_\{j\}\}=1,\\widehat\{Y\}\\neq Y\)its empirical error mass\. Whenc^j\>0\\widehat\{c\}\_\{j\}\>0, its empirical selective diagnostic error isq^j/c^j\\widehat\{q\}\_\{j\}/\\widehat\{c\}\_\{j\}\. Because the component is sampled independently for each episode, the mixture’s expected cost, error mass, and coverage are∑jwj𝒥^j\\sum\_\{j\}w\_\{j\}\\widehat\{\\mathcal\{J\}\}\_\{j\},∑jwjq^j\\sum\_\{j\}w\_\{j\}\\widehat\{q\}\_\{j\}, and∑jwjc^j\\sum\_\{j\}w\_\{j\}\\widehat\{c\}\_\{j\}, respectively\. We choose the least costly mixture satisfying the selection\-stage risk and coverage targets:
minw∈ΔL\\displaystyle\\min\_\{w\\in\\Delta\_\{L\}\}∑j=1Lwj𝒥^js\.t\.\\displaystyle\\sum\_\{j=1\}^\{L\}w\_\{j\}\\widehat\{\\mathcal\{J\}\}\_\{j\}\\quad\\text\{s\.t\.\}\\quad∑j=1Lwj\(q^j−αdesc^j\)≤0,∑j=1Lwjc^j≥γdes\.\\displaystyle\\sum\_\{j=1\}^\{L\}w\_\{j\}\\left\(\\widehat\{q\}\_\{j\}\-\\alpha\_\{\\mathrm\{des\}\}\\widehat\{c\}\_\{j\}\\right\)\\leq 0,\\qquad\\sum\_\{j=1\}^\{L\}w\_\{j\}\\widehat\{c\}\_\{j\}\\geq\\gamma\_\{\\mathrm\{des\}\}\.\(9\)The first constraint is equivalent to requiring the mixture’s empirical selective diagnostic error,\(∑jwjq^j\)/\(∑jwjc^j\)\(\\sum\_\{j\}w\_\{j\}\\widehat\{q\}\_\{j\}\)/\(\\sum\_\{j\}w\_\{j\}\\widehat\{c\}\_\{j\}\), to be at mostαdes\\alpha\_\{\\mathrm\{des\}\}; the denominator is positive because the second constraint requires coverage of at leastγdes\>0\\gamma\_\{\\mathrm\{des\}\}\>0\.
We use the stricter selection\-design targets\(αdes,γdes\)=\(0\.20,0\.80\)\(\\alpha\_\{\\mathrm\{des\}\},\\gamma\_\{\\mathrm\{des\}\}\)=\(0\.20,0\.80\), while the subsequent calibration tests use\(α,γ\)=\(0\.25,0\.70\)\(\\alpha,\\gamma\)=\(0\.25,0\.70\)\. These design margins provide a buffer against selection\-sample variation but are not themselves a statistical certificate\.
###### Proposition 1\(Sparsity and validity\)\.
If the linear program in Eq\. \([9](https://arxiv.org/html/2609.09678#S3.E9)\) is feasible, it admits an optimal basic feasible solution in which at most three weightswjw\_\{j\}are positive\. If the mixture weights and episode\-wise randomization mechanism are frozen before calibration, the induced randomized controller is a single frozen policy whose risk and coverage can be tested under Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1), subject to the theorem’s sampling assumptions\.
Proposition[1](https://arxiv.org/html/2609.09678#Thmproposition1)implies that, although the optimization considersLLdeterministic candidates, an optimal mixture uses at most three of them\. This bound follows from the simplex equality and the two risk and coverage design constraints in Eq\. \([9](https://arxiv.org/html/2609.09678#S3.E9)\)\. At deployment, component sampling is independent across episodes and is never conditioned on patient characteristics or intermediate observations\. Equivalently, the random seed can be treated as part of each i\.i\.d\. episode\. A single component draw reused for all calibration episodes would instead introduce shared randomness and would not justify the same exact binomial test\. A proof is provided in Appendix[A](https://arxiv.org/html/2609.09678#A1)\.
## 4Experimental Design
### 4\.1Benchmark and study splits
We construct a retrospective sequential\-diagnosis benchmark by linking credentialed\-access MIMIC\-IV\-ED v2\.2\([Johnson et al\., 2023](https://arxiv.org/html/2609.09678#bib.bib6)\)with MIMIC\-IV\-Ext\-CDS v1\.0\.2\([Gaber and Akalin, 2025](https://arxiv.org/html/2609.09678#bib.bib7)\)\. Retaining the earliest eligible ED stay per patient yields 1,834 patient\-level episodes across nine abdominal\-pain diagnosis classes\. The initial state includes the deidentified history of present illness, chief complaint, demographics, and triage measurements\.
The environment exposes 12 action groups spanning repeated vital signs, laboratory studies, electrocardiography, imaging, and microbiology\. Continuing along an episode reveals the recorded result of the test proposed by the backbone\. Because most discharge\-note test snippets lack reliable acquisition timestamps, this environment is a logged retrospective benchmark rather than a causal simulator of alternative testing decisions\. An unavailable result is represented byNO\_RECORDED\_RESULTand retains its prespecified cost\. The complete action definitions and cost vector are reported in Appendix[B](https://arxiv.org/html/2609.09678#A2)\.
Patients are divided into 1,100 development, 367 calibration, and 367 evaluation episodes, with no patient overlap\. The development cohort is further separated into 935 episodes for backbone and risk\-ranker fitting and 165 episodes for policy selection\. All model fitting, candidate construction, policy ordering, and mixture optimization use only these development partitions\. The primary population targets are selective diagnostic errorα=0\.25\\alpha=0\.25, autonomous coverageγ=0\.70\\gamma=0\.70, and family\-wise error levelδ=0\.05\\delta=0\.05\.
#### Exploratory status\.
Calibration and evaluation labels had been accessed during earlier method development\. Consequently, all reported calibration passes,pp\-values, and confidence bounds are interpreted descriptively rather than as realized prospective certificates\. The current freeze prevents further outcome\-dependent modification but cannot restore statistical independence; applying Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1)confirmatorily requires a new untouched cohort\.
### 4\.2Model instantiation and comparators
The diagnostic backbone is Qwen2\.5\-7B\-Instruct\([Qwen Team, 2024](https://arxiv.org/html/2609.09678#bib.bib15)\), adapted with LoRA\([Hu et al\., 2022](https://arxiv.org/html/2609.09678#bib.bib14)\)using an official\-code\-derived LA\-CDM training pipeline\. Training uses development data only\. Because our implementation transfers the official architecture and training structure to a fixed common\-path environment, it should not be interpreted as a prompt\-equivalent reproduction of LA\-CDM’s original free\-form rollouts\. Model configuration, optimization, prompts, seeds, and provenance checks are provided in Appendix[H](https://arxiv.org/html/2609.09678#A8)\.
A cross\-fitted histogram gradient\-boosting ranker estimates state\-level diagnostic error from backbone probabilities, uncertainty summaries, trajectory state, accumulated cost, missingness, and the native stopping score\. Crossing 10 target coverages with 13 policy horizons produces 130 threshold\-and\-horizon candidates\. The disjoint policy\-selection split freezes a tested family of 12 deterministic policies and, separately, one optimized randomized mixture\. Candidate construction, ranker hyperparameters, and threshold estimation are detailed in Appendix[H](https://arxiv.org/html/2609.09678#A8)\.
Every stopping method receives the same backbone probabilities, diagnoses, proposed actions, and forced\-continuation trajectory\. This common\-path protocol holds diagnosis and acquisition behavior fixed, isolating the decision to stop, continue, or defer\. Comparators include HPI\-only and full\-workup endpoints, fixed\-stage and confidence\-threshold stopping, the native LA\-CDM\-style rule with and without confidence deferral, and empirical cost minimization\. We additionally evaluate calibration and multiplicity procedures, score and history ablations, and a mixture\-weight ablation\. These labels describe experimental roles rather than additionalCrosmethods\. Complete comparator definitions are provided in Appendix[H](https://arxiv.org/html/2609.09678#A8)\.
#### Method nomenclature and policy construction\.
Crosis the overall risk\-constrained stopping framework and is instantiated here through exactly two named controllers\.Cros\-Det is a deterministic threshold\-and\-horizon policy selected from the locked development family, andCros\-Mix is an episode\-wise randomized, LP\-optimized mixture over that same frozen deterministic family\. All other experimental labels denote fixed\-information, confidence\-based, native\-agent, or empirical baselines; calibration and multiplicity procedures; ranker ablations; a policy\-design ablation; or evaluation modes ofCros\-Mix\. They are not additionalCrosmethods\.
Specifically,Cros\-Det is the lowest\-cost member of the full\-ranker 130\-policy grid that satisfies the locked selection\-split design margins, risk at most 20% and coverage at least 80%; in the present run it is\(h=3,τ=0\.3286\)\(h=3,\\tau=0\.3286\)\. The best deterministic component within the optimized mixture support is identical toCros\-Det in this run and is therefore not displayed separately\.Cros\-Mix always denotes the same three support policies and the same nonnegative LP\-optimized weights, fitted using only the policy\-selection split to minimize expected cost subject to the 20%/80% design margins\.Cros\-Mix, analytic expectation integrates component contributions for each episode, whereasCros\-Mix, realized draw uses one frozen independent episode\-wise component draw and yields integer autonomous/error counts for exact binomial testing\. These are two evaluation modes of one controller\. The Uniform\-weight mixture uses the same frozen support with equal rather than optimized weights and is a policy\-design ablation; it is not guaranteed to satisfy the selection margins\.
The Single\-candidate test, Fixed\-sequence LTT, Holm testing, and Bonferroni testing are calibration or multiplicity procedures applied to the same 12 selection\-frozen deterministic candidates\. They change the testing and return rule, not the risk ranker, backbone, or basic threshold\-and\-horizon policy\. In this run, Holm testing returns the same deterministic controller as Fixed\-sequence LTT, while Bonferroni testing returns a different deterministic controller\. Empirical ERM instead minimizes selection\-split cost without requiring the joint risk–coverage criterion\.
Myopic value\-of\-information and free\-form multi\-agent systems are not included in the primary paired comparison because they select different actions and therefore induce different trajectories\. Appendix[H](https://arxiv.org/html/2609.09678#A8)specifies how such systems would enter a future end\-to\-end confirmatory study\.
### 4\.3Evaluation
The primary outcomes are selective diagnostic error, autonomous coverage, total relative resource cost, and number of requested tests\. We also report error mass,Pr\(D=1,Y^≠Y\)\\Pr\(D=1,\\widehat\{Y\}\\neq Y\), which measures the population fraction receiving an incorrect autonomous diagnosis and therefore does not decrease merely because the controller defers additional cases\.
Uncertainty estimates respect the patient\-level sampling unit\. We report exact one\-sided Clopper–Pearson bounds for risk and coverage and paired patient\-level bootstrap intervals for method contrasts\. Randomized mixtures are evaluated both by analytically averaging component contributions and by repeated episode\-wise realizations to assess randomization stability\. Prespecified sensitivity analyses examine deferral penalties, action costs, recorded\-result missingness, diagnosis\-language masking, and development\-split stability\. Full metric definitions, resampling procedures, and sensitivity settings are reported in Appendices[B](https://arxiv.org/html/2609.09678#A2)–[H](https://arxiv.org/html/2609.09678#A8)\.
## 5Results
### 5\.1Exploratory joint calibration
Table[1](https://arxiv.org/html/2609.09678#S5.T1)reports realized calibration outcomes and descriptive evaluation performance\. All returned policies have evaluation point estimates below 25% risk and above 70% coverage\. The prespecifiedCros\-Mix realization is the least costly displayed policy, attaining 16\.3% error at 78\.5% coverage with cost 5\.68 and 0\.68 requested actions\.
The Single\-candidate test and theCros\-Mix realized draw are each tested once, Fixed\-sequence LTT follows its frozen order, and Bonferroni testing tests 12 candidates; Holm testing returns the same deterministic controller as Fixed\-sequence LTT in this run\. Appendix[E](https://arxiv.org/html/2609.09678#A5)gives the component tests, exact bounds, weights, and seed audit\. Later mixture comparisons use theCros\-Mix analytic expectation \(0\.169 risk, 0\.788 coverage, 5\.57 cost, and 0\.679 actions\), rather than the realization in Table[1](https://arxiv.org/html/2609.09678#S5.T1)\.
Table 1:Exploratory calibration and evaluation results \(ncal=neval=367n\_\{\\mathrm\{cal\}\}=n\_\{\\mathrm\{eval\}\}=367\)\. “Auto/err” denotes autonomous diagnoses/errors\. The table reports raw jointpp\-values before multiplicity adjustment\. The first three rows are testing procedures applied to the same frozen deterministic family; the final row is one realized draw ofCros\-Mix\. Holm testing returns the same controller as fixed\-sequence LTT and is not duplicated\.
### 5\.2Risk\-ranker and stopping ablation
The full ranker achieves state\-error AUROC 0\.853, compared with 0\.715 for the maximum\-probability ranker and 0\.552 for the Native\-score ranker \(Table[2](https://arxiv.org/html/2609.09678#S5.T2)\)\. Among these alternatives, only the full\-ranker controller satisfies the locked selection margins and has an exploratory calibration pass\. The maximum\-probability ranker is cheaper but lacks comparable calibration evidence\. The No\-history ranker nearly matches the full ranker \(AUROC 0\.852\), so this experiment does not isolate a material benefit from stage, cost, latency, and missing\-history features\. The supported contribution is therefore improved risk ranking and calibration power, not unconditional cost dominance\.
Table 2:Risk\-ranker ablation\. “Sel\.” indicates the locked selection margins\. AUROC intervals use 10,000 patient\-level resamples; calibrationpp\-values are descriptive\.
### 5\.3Policy and matched comparisons
Cros\-Mix, analytic expectation is 0\.903 cost units cheaper thanCros\-Det \(95% CI\[−1\.135,−0\.682\]\[\-1\.135,\-0\.682\]\) and requests 0\.195 fewer actions; its risk difference is unresolved and its coverage is 0\.021 lower \(Table[3](https://arxiv.org/html/2609.09678#S5.T3)\)\. The Uniform\-weight mixture is another 0\.977 units cheaper on evaluation, but misses both locked selection margins\. Thus, optimization enforces the development constraints rather than guaranteeing the lowest future cost\.
Compared with confidence thresholding and native stopping with deferral,Cros\-Mix, analytic expectation reduces risk by 0\.088 and 0\.063 and cost by 12\.54 and 4\.84 units, respectively; coverage differences are unresolved\. ERM is cheaper but has higher risk and fails exploratory calibration\. Appendix[D](https://arxiv.org/html/2609.09678#A4)reports the complete paired intervals\. Cost advantages persist in the primary, imaging\-sensitive, and missing\-result\-sensitive scenarios, but not under every deferral penalty \(Appendix[F](https://arxiv.org/html/2609.09678#A6)\)\.
Table 3:Policy\-design ablation using analytic mixture expectations\. “Sel\.” indicates the frozen 20% risk and 80% coverage design constraints, not an evaluation guarantee\. The fixed\-sequence LTT row reports the deterministic controller returned by that testing procedure\.Against the confidence\-threshold controller,Cros\-Mix, analytic expectation reduces selective risk by 0\.088 and cost by 12\.54 units; against native stopping with deferral, it reduces risk by 0\.063 and cost by 4\.84 units, with unresolved coverage differences in both comparisons\. ERM remains less costly but has higher risk and fails exploratory calibration\. Complete patient\-paired intervals and error\-mass comparisons are reported in Appendix[D](https://arxiv.org/html/2609.09678#A4)\.
The cost advantage over the principal calibrated and confidence\-based comparators persists under the primary, imaging\-sensitive, and missing\-result\-sensitive cost scenarios, but not under every deferral penalty\. Appendix[F](https://arxiv.org/html/2609.09678#A6)reports the complete frozen\-decision sensitivity analysis\.
### 5\.4Trajectory and robustness audits
Forced\-continuation error is 28\.3% with HPI alone, 27\.5% after one action, 35\.7% at stage 8, and 34\.3% after full workup, while cost and recorded\-result missingness increase along the trajectory \(Appendix Figure[2](https://arxiv.org/html/2609.09678#A4.F2)\)\. Thus, this backbone does not show monotone gains from additional acquisition\. Diagnosis\-language masking lowers ranker AUROC by approximately 0\.014 and modestly reduces coverage and increases cost; risk changes are unresolved \(Appendix Table[8](https://arxiv.org/html/2609.09678#A3.T8)\)\. Action availability is also highly structured and may encode clinicians’ historical ordering\. Across 20 development resplits,Cros\-Det nominally passes in 15, whereasCros\-Mix is feasible in 15 and passes in only six\. These dependent audits demonstrate design sensitivity, not independent confirmation\. Class\-level failures further preclude any subgroup safety claim \(Appendix[G](https://arxiv.org/html/2609.09678#A7)\)\.
## 6Discussion and Limitations
The results support a prospectively frozen follow\-up, not deployment or a clinical safety claim\. The full ranker outperforms simple scores, joint tests have apparent power for several frozen policies, andCros\-Mix reduces cost relative toCros\-Det\. However, the No\-history ranker nearly matches the full ranker, the Uniform\-weight mixture is cheaper despite missing the selection margins, forced continuation is non\-monotone, and mixture feasibility varies across development splits\.
The claims are limited by previously accessed labels; a single\-center, note\-derived benchmark without reliable non\-vital timestamps or counterfactual outcomes; informative missingness; unadjudicated labels, proxy misses, and costs; marginal rather than subgroup control; one backbone\-training seed; and dependent resplit audits\. Common\-path evaluation isolates stopping but does not measure end\-to\-end gains when agents choose different actions\.
A confirmatory study should lock a never\-viewed temporal or external cohort, prespecify labels, costs, subgroup hypotheses, sample size, and one testing graph, and obtain blinded clinical adjudication\. It should evaluate both common\-path stopping and end\-to\-end acquisition without changes after calibration access\.
## 7Conclusion
Crosmakes sequential diagnostic stopping auditable through joint testing of selective error and autonomous coverage\. The current retrospective results support feasibility and a lower\-cost deterministic–mixture trade\-off, but label reuse, structured missingness, resplit instability, and subgroup failures require a new prospectively locked study before any safety claim\.
## References
- A\. N\. Angelopoulos, S\. Bates, E\. J\. Candès, M\. I\. Jordan, and L\. LeiLearn then test: calibrating predictive algorithms to achieve risk control\.The Annals of Applied Statistics19\(2\),pp\. 1641–1662\.External Links:[Document](https://dx.doi.org/10.1214/24-AOAS1998)Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px9.p1.1),[§1](https://arxiv.org/html/2609.09678#S1.p1.1),[§1](https://arxiv.org/html/2609.09678#S1.p3.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Angelopouloset al\.\(2024\)A\. N\. Angelopoulos, S\. Bates, A\. Fisch, L\. Lei, and T\. SchusterConformal risk control\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=33XGfHLtZg)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Bai and Jin \(2026\)T\. Bai and Y\. JinConformal selective prediction with general risk control\.arXiv preprint arXiv:2603\.24704\.External Links:[Link](https://arxiv.org/abs/2603.24704)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Bani\-Harouniet al\.\(2026\)D\. Bani\-Harouni, C\. Pellegrini, E\. Özsoy, N\. Navab, and M\. KeicherLanguage agents for hypothesis\-driven clinical decision making with reinforcement learning\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=Pn8ETgEHW7)Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px6.p1.1),[§1](https://arxiv.org/html/2609.09678#S1.p1.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px1.p1.1)\.
- Bateset al\.\(2021\)S\. Bates, A\. Angelopoulos, L\. Lei, J\. Malik, and M\. JordanDistribution\-free, risk\-controlling prediction sets\.Journal of the ACM68\(6\),pp\. 1–34\.External Links:[Document](https://dx.doi.org/10.1145/3478535)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Foo and Chang \(2026\)H\. Foo and Y\. I\. ChangOptimal stopping in sequential clinical prediction\.arXiv preprint arXiv:2604\.22216\.External Links:[Link](https://arxiv.org/abs/2604.22216)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px4.p1.1)\.
- Gaber and Akalin \(2025\)M\. M\. Gaber and A\. AkalinMIMIC\-IV\-Ext Clinical Decision Support for Referral, Triage and Diagnosis \(version 1\.0\.2\)\.Note:PhysioNetExternal Links:[Document](https://dx.doi.org/10.13026/stnm-qx35),[Link](https://physionet.org/content/mimic-iv-ext-cds/1.0.2/)Cited by:[Appendix B](https://arxiv.org/html/2609.09678#A2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.09678#S4.SS1.p1.1)\.
- Geifman and El\-Yaniv \(2017\)Y\. Geifman and R\. El\-YanivSelective classification for deep neural networks\.InAdvances in Neural Information Processing Systems,Vol\.30,pp\. 4878–4887\.Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.09678#S1.p2.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Geifman and El\-Yaniv \(2019\)Y\. Geifman and R\. El\-YanivSelectiveNet: a deep neural network with an integrated reject option\.InProceedings of the 36th International Conference on Machine Learning,Vol\.97,pp\. 2151–2159\.Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px5.p1.1),[§1](https://arxiv.org/html/2609.09678#S1.p2.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Hageret al\.\(2024\)P\. Hager, F\. Jungmann, R\. Holland, K\. Bhagat, I\. Hubrecht, M\. Knauer, J\. Vielhauer, M\. Makowski, R\. Braren, G\. Kaissis, and D\. RueckertEvaluation and mitigation of the limitations of large language models in clinical decision\-making\.Nature Medicine30,pp\. 2613–2622\.External Links:[Document](https://dx.doi.org/10.1038/s41591-024-03097-1)Cited by:[§1](https://arxiv.org/html/2609.09678#S1.p1.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px1.p1.1)\.
- Holm \(1979\)S\. HolmA simple sequentially rejective multiple test procedure\.Scandinavian Journal of Statistics6\(2\),pp\. 65–70\.Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px9.p1.1),[§3\.3](https://arxiv.org/html/2609.09678#S3.SS3.p3.3)\.
- Huet al\.\(2022\)E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.International Conference on Learning Representations\.External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.09678#S4.SS2.p1.1)\.
- Johnsonet al\.\(2023\)A\. E\. W\. Johnson, L\. Bulgarelli, L\. Shen, A\. Gayles, A\. Shammout, S\. Horng, T\. J\. Pollard, S\. Hao, B\. Moody, B\. Gow, L\. H\. Lehman, L\. A\. Celi, and R\. G\. MarkMIMIC\-IV\-ED \(version 2\.2\)\.Note:PhysioNetExternal Links:[Document](https://dx.doi.org/10.13026/5ntk-km72),[Link](https://physionet.org/content/mimic-iv-ed/2.2/)Cited by:[Appendix B](https://arxiv.org/html/2609.09678#A2.SS0.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.09678#S4.SS1.p1.1)\.
- Laufer\-Goldshteinet al\.\(2023\)B\. Laufer\-Goldshtein, A\. Fisch, R\. Barzilay, and T\. S\. JaakkolaEfficiently controlling multiple risks with pareto testing\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=cyg2YXn_BqF)Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px9.p1.1),[§1](https://arxiv.org/html/2609.09678#S1.p1.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px5.p1.1)\.
- Liet al\.\(2024\)S\. S\. Li, V\. Balachandran, S\. Feng, J\. S\. Ilgen, E\. Pierson, P\. W\. Koh, and Y\. TsvetkovMediQ: question\-asking llms and a benchmark for reliable interactive clinical reasoning\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 28858–28888\.External Links:[Document](https://dx.doi.org/10.52202/079017-0908),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/32b80425554e081204e5988ab1c97e9a-Paper-Conference.pdf)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px4.p1.1)\.
- Liuet al\.\(2024\)J\. Liu, W\. Wang, Z\. Ma, G\. Huang, Y\. Su, K\. Chang, W\. Chen, H\. Li, L\. Shen, and M\. LyuMedChain: bridging the gap between LLM agents and clinical practice with interactive sequence\.arXiv preprint arXiv:2412\.01605\.External Links:[Link](https://arxiv.org/abs/2412.01605)Cited by:[§1](https://arxiv.org/html/2609.09678#S1.p1.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px1.p1.1)\.
- Lvet al\.\(2026\)Z\. Lv, D\. Tu, J\. Li, M\. Zhao, H\. Zhu, W\. Li, and S\. K\. ZhouThinking like a clinician: a cognitive AI agent for clinical diagnosis via panoramic profiling and adversarial debate\.arXiv preprint arXiv:2604\.23605\.External Links:[Link](https://arxiv.org/abs/2604.23605)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px1.p1.1)\.
- Noriet al\.\(2025\)H\. Nori, M\. Daswani, C\. Kelly, S\. Lundberg, M\. T\. Ribeiro, M\. Wilson, X\. Liu, V\. Sounderajah, J\. Carlson, M\. P\. Lungren, B\. Gross, P\. Hames, M\. Suleyman, D\. King, and E\. HorvitzSequential diagnosis with language models\.arXiv preprint arXiv:2506\.22405\.External Links:[Link](https://arxiv.org/abs/2506.22405)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px2.p1.1)\.
- Presacanet al\.\(2026\)O\. Presacan, A\. Grama, L\. Irimină, A\. Nik, J\. Ojha, V\. Thambawita, C\. I\. Băcilă, B\. Ionescu, and M\. A\. RieglerAsk before you diagnose: Safe\-Psych, a sequential evaluation benchmark for LLMs in psychiatry\.arXiv preprint arXiv:2607\.13036\.External Links:[Link](https://arxiv.org/abs/2607.13036)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px4.p1.1)\.
- Qwen Team \(2024\)Qwen TeamQwen2\.5 technical report\.arXiv preprint arXiv:2412\.15115\.External Links:[Link](https://arxiv.org/abs/2412.15115)Cited by:[Appendix H](https://arxiv.org/html/2609.09678#A8.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2609.09678#S4.SS2.p1.1)\.
- Ringelet al\.\(2024\)L\. Ringel, R\. Cohen, D\. Freedman, M\. Elad, and Y\. RomanoEarly time classification with accumulated accuracy gap control\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 42584–42600\.External Links:[Link](https://proceedings.mlr.press/v235/ringel24a.html)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px5.p1.1)\.
- Ruhrberg Estévezet al\.\(2025\)S\. Ruhrberg Estévez, N\. Astorga, and M\. van der SchaarTimely clinical diagnosis through active test selection\.InAdvances in Neural Information Processing Systems,External Links:[Link](https://arxiv.org/abs/2510.18988)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px2.p1.1)\.
- Schmidgallet al\.\(2024\)S\. Schmidgall, R\. Ziaei, C\. Harris, E\. Reis, J\. Jopling, and M\. MoorAgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments\.arXiv preprint arXiv:2405\.07960\.External Links:[Link](https://arxiv.org/abs/2405.07960)Cited by:[§1](https://arxiv.org/html/2609.09678#S1.p1.1),[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px1.p1.1)\.
- Shenet al\.\(2026\)X\. Shen, H\. Liu, D\. Song, and M\. R\. MinUncertainty\-guided latent diagnostic trajectory learning for sequential clinical diagnosis\.arXiv preprint arXiv:2604\.05116\.External Links:[Link](https://arxiv.org/abs/2604.05116)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px2.p1.1)\.
- Vazhentsevet al\.\(2025\)A\. Vazhentsev, I\. Sviridov, A\. Barseghyan, G\. Kuzmin, A\. Panchenko, A\. Nesterov, A\. Shelmanov, and M\. PanovUncertainty\-aware abstention in medical diagnosis based on medical texts\.arXiv preprint arXiv:2502\.18050\.External Links:[Link](https://arxiv.org/abs/2502.18050)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px4.p1.1)\.
- Wanget al\.\(2026\)X\. Wang, A\. Suresh, A\. Zhang, R\. More, W\. Jurayj, B\. Van Durme, M\. Farajtabar, D\. Khashabi, and E\. NalisnickConformal thinking: risk control for reasoning on a compute budget\.arXiv preprint arXiv:2602\.03814\.External Links:[Link](https://arxiv.org/abs/2602.03814)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px5.p1.1)\.
- Xuet al\.\(2026a\)J\. Xu, Y\. Wu, D\. Zeng, J\. Paisley, and Q\. ZhaoLook again before you abstain: budgeted conformal evidence acquisition for reliable vision\-language model\.arXiv preprint arXiv:2606\.16667\.External Links:[Link](https://arxiv.org/abs/2606.16667)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px5.p1.1)\.
- Xuet al\.\(2026b\)R\. Xu, Y\. Chen, S\. Xie, and H\. XiongGeometry\-calibrated conformal abstention for language models\.arXiv preprint arXiv:2604\.27914\.External Links:[Link](https://arxiv.org/abs/2604.27914)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px3.p1.1)\.
- Xuet al\.\(2025\)Y\. Xu, W\. Guo, and Z\. WeiSelective conformal risk control\.arXiv preprint arXiv:2512\.12844\.External Links:[Link](https://arxiv.org/abs/2512.12844)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px5.p1.1)\.
- Yuet al\.\(2023\)Z\. Yu, Y\. Li, J\. C\. Kim, K\. Huang, Y\. Luo, and M\. WangDeep reinforcement learning for cost\-effective medical diagnosis\.InThe Eleventh International Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=0WVNuEnqVu)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px2.p1.1)\.
- Zhouet al\.\(2026\)X\. Zhou, H\. Nguyen, B\. Yu, C\. Liu, and L\. ChengAdaptive stopping for multi\-turn LLM reasoning\.arXiv preprint arXiv:2604\.01413\.External Links:[Link](https://arxiv.org/abs/2604.01413)Cited by:[§2](https://arxiv.org/html/2609.09678#S2.SS0.SSS0.Px5.p1.1)\.
## Appendix AProofs
###### Proof of Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1)\.
Fix a candidateπj\\pi\_\{j\}independently of the calibration outcomes\. Conditional onMj=m\>0M\_\{j\}=m\>0, the number of autonomous errors isEj∼Binomial\(m,rj\)E\_\{j\}\\sim\\mathrm\{Binomial\}\(m,r\_\{j\}\)under i\.i\.d\. sampling, whererj=ℛ\(πj\)r\_\{j\}=\\mathcal\{R\}\(\\pi\_\{j\}\)\. For the nullrj\>αr\_\{j\}\>\\alpha, the lower\-tail statisticFBin\(Ej,m,α\)F\_\{\\mathrm\{Bin\}\}\(E\_\{j\};m,\\alpha\)is largest at the boundary in the rejection\-relevant direction; its discreteness makes it super\-uniform\. ThuspR,jp\_\{R,j\}is valid forHR,j:rj\>αH\_\{R,j\}:r\_\{j\}\>\\alpha\. SettingpR,j=1p\_\{R,j\}=1whenm=0m=0preserves validity\.
Marginally,Mj∼Binomial\(n,cj\)M\_\{j\}\\sim\\mathrm\{Binomial\}\(n,c\_\{j\}\)withcj=𝒞\(πj\)c\_\{j\}=\\mathcal\{C\}\(\\pi\_\{j\}\)\. UnderHC,j:cj<γH\_\{C,j\}:c\_\{j\}<\\gamma, the upper\-tail value1−FBin\(Mj−1,n,γ\)1\-F\_\{\\mathrm\{Bin\}\}\(M\_\{j\}\-1;n,\\gamma\)is super\-uniform, again with the boundary least favorable in the rejection direction\. The candidate null is the unionHj=HR,j∪HC,jH\_\{j\}=H\_\{R,j\}\\cup H\_\{C,j\}\. The intersection\-union test rejects only if both component tests reject\. Therefore, for any distribution inHjH\_\{j\}, at least one component null is true and
Pr\{max\(pR,j,pC,j\)≤u\}≤Pr\{pk,j≤u\}≤u,\\Pr\\\{\\max\(p\_\{R,j\},p\_\{C,j\}\)\\leq u\\\}\\leq\\Pr\\\{p\_\{k,j\}\\leq u\\\}\\leq u,wherekkindexes a true component null\. Hencepj=max\(pR,j,pC,j\)p\_\{j\}=\\max\(p\_\{R,j\},p\_\{C,j\}\)is super\-uniform\. Applying any valid family\-wise\-error procedure at levelδ\\deltato the finite, pre\-frozen family implies that the probability of rejecting at least one trueHjH\_\{j\}is at mostδ\\delta\. A certified policy violates a desired constraint exactly when its union null is true, proving the result\. ∎
###### Proof of the sparsity and randomized\-policy proposition\.
Write the linear program in standard form after adding slack variables\. The policy\-weight vector obeys one simplex equality\. At a nondegenerate extreme point, at most two independent design inequalities can be active in addition to the simplex equality\. Therefore at most three policy weights need be basic and positive; a degenerate optimum has no larger support, and an optimal basic feasible solution can always be selected\.
For validity, augment each episode withUi∼Uniform\(0,1\)U\_\{i\}\\sim\\mathrm\{Uniform\}\(0,1\), independently across episodes and independent of the clinical variables\. The frozen mixture mapsUiU\_\{i\}to a deterministic component and then applies it toZiZ\_\{i\}\. Thus\(Zi,Ui\)\(Z\_\{i\},U\_\{i\}\)are i\.i\.d\. and the mixture is a single fixed randomized policy\. Its autonomous indicator and error indicator satisfy the same binomial conditioning argument as in Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1)\. The conclusion fails if mixture weights are estimated on calibration outcomes or if one shared random component is drawn for the entire calibration sample\. ∎
## Appendix BBenchmark, Costs, and Ranker Configuration
#### Cohort construction\.
We link credentialed\-access MIMIC\-IV\-ED v2\.2\[[Johnson et al\., 2023](https://arxiv.org/html/2609.09678#bib.bib6)\]with MIMIC\-IV\-Ext\-CDS v1\.0\.2\[[Gaber and Akalin, 2025](https://arxiv.org/html/2609.09678#bib.bib7)\]\. Cohort construction retains the earliest qualifying emergency\-department stay for each subject before any data partitioning, so each patient contributes exactly one episode\. The resulting benchmark contains 1,834 episodes assigned by the frozen label\-mapping code to nine abdominal\-pain diagnosis classes: appendicitis, biliary disease, bowel obstruction, diverticulitis, gastroenteritis/colitis, nonspecific abdominal pain, pancreatitis, renal colic, and urinary infection\.
The initial presentation includes a deidentified history of present illness, chief complaint, demographic variables, and triage measurements\. These fields constitute the stage\-zero information available before any action is requested\. The executable cohort query, label mapping, and episode identifiers are versioned as part of the frozen benchmark artifacts\.
#### Patient\-level partitions\.
Patients are partitioned into 1,100 development, 367 calibration, and 367 evaluation episodes\. The development set is further divided into 935 ranker\-training and 165 policy\-selection episodes\. All states from one episode remain in the same partition and, during ranker cross\-fitting, in the same fold\. The four resulting patient sets are therefore disjoint\. The 935\-episode subset is used to fit the state\-error ranker, whereas thresholds, deterministic candidates, candidate order, and mixture weights are designed only on the 165\-episode selection subset\. The calibration and evaluation subsets are used for the analyses described in the main text\. As discussed in Section[6](https://arxiv.org/html/2609.09678#S6), prior access to their labels makes the present results exploratory rather than confirmatory\.
Table 4:Benchmark summary\. Here,δ\\deltais the family\-wise error level used by the joint testing procedure\.
#### Retrospective action environment\.
The action space comprises rechecking vital signs, complete blood count, metabolic panel, hepatic panel, lipase, urinalysis, electrocardiogram, X\-ray, ultrasound, computed tomography, magnetic resonance imaging, and microbiology\. The maximum horizonH=12H=12specifies the largest number of action opportunities considered in a forced\-continuation episode; it does not imply that all 12 results are available in the record\.
For each action group, recorded findings are extracted from discharge\-note test snippets\. Except for vital signs, these snippets generally lack acquisition timestamps sufficiently reliable to reconstruct the clinical order in which results became available\. Consequently, an observation represents information found in the retrospective record for the requested action group, rather than the counterfactual result of ordering that action prospectively\. When no corresponding result is available, the environment returnsNO\_RECORDED\_RESULT\. This missing\-result token is retained as an observed outcome and may affect subsequent backbone predictions\.
The benchmark should therefore be interpreted as a logged\-acquisition environment for comparing stopping rules on a shared backbone trajectory\. It does not estimate how ordering a different test would alter subsequent care, physiology, documentation, or diagnostic outcomes\. In particular,Croscontrols whether the next backbone\-proposed action is executed, but it does not replace that action or simulate an alternative trajectory\.
Table 5:Frozen relative action\-cost vector shared by every controller\. The full\-workup total is the sum of the 12 action costs; the deferral penalty is separate\.Action groupCostAction groupCostRecheck vitals0\.2Electrocardiogram1\.5Complete blood count1\.0X\-ray4\.0Metabolic panel1\.2Ultrasound6\.0Hepatic panel1\.2Computed tomography12\.0Lipase1\.0Magnetic resonance imaging20\.0Urinalysis0\.8Microbiology3\.0All 12 actions51\.9Deferral penalty10\.0
#### Cost interpretation\.
An attempted action incurs its Table[5](https://arxiv.org/html/2609.09678#A2.T5)cost even when the retrospective record returnsNO\_RECORDED\_RESULT\. The reported cost is a normalized resource index intended to support controlled comparisons among stopping policies\. It is not a hospital bill, reimbursement amount, radiation dose, patient utility, or estimate of clinical harm\. Likewise, the deferral penalty represents a stylized downstream\-review cost rather than a measured clinical or monetary quantity\.
The frozen\-decision sensitivity analysis varies the deferral penalty over\{5,10,15,20,25,30\}\\\{5,10,15,20,25,30\\\}and separately applies multipliers\{0\.5,1,2\}\\\{0\.5,1,2\\\}to laboratory, imaging, and attempted\-but\-missing action costs according to the frozen category mapping\. These analyses retain the originally frozen stopping decisions and mixture weights; they recalculate costs but do not refit the ranker, reconstruct thresholds, or reoptimize the policies\.
#### Risk\-ranker inputs\.
Each stage\-level ranker target is the binary indicator𝟏\{Y^t≠Y\}\\mathbf\{1\}\\\{\\widehat\{Y\}\_\{t\}\\neq Y\\\}, which records whether the backbone’s current diagnosis is incorrect\. The ranker uses the backbone’sKKdiagnosis probabilities together with their maximum, the margin between the two largest probabilities, predictive entropy, the current stage, the fraction of previously requested actions with no recorded result, cumulative relative cost, the recorded latency feature, and the backbone’s native stopping score\. It does not select the next action or directly modify the backbone diagnosis\.
The 935 ranker\-training episodes yield 12,155 stage\-level records\. Because states from the same patient are correlated, cross\-fitting and all resampling operations are performed at the episode level rather than at the state level\.
Table 6:Frozen histogram gradient\-boosting risk\-ranker configuration\. Automatic early stopping follows scikit\-learn’s implementation\.
#### Cross\-fitting and final ranker\.
The 935 ranker\-training episodes are divided into five stratified outer folds\. For each fold, the estimator is fitted using the other four folds and generates predictions for all states belonging to the held\-out episodes\. This construction prevents states from the same patient from appearing on both sides of an outer\-fold fit\. The resulting out\-of\-fold scores are used for ranker\-development diagnostics\.
Each outer\-fold training set contains fewer than 10,000 state records, so scikit\-learn’s automatic early\-stopping condition is not activated and all five estimators reach the 160\-iteration limit\. After cross\-fitting, the final ranker is refitted on all 12,155 state records\. Because this fit exceeds the automatic early\-stopping sample threshold, it uses the internal 10% validation fraction and stops after 149 iterations\. This final frozen estimator produces the risk scores used on the disjoint policy\-selection, calibration, and evaluation episodes\.
#### Coverage\-indexed threshold construction\.
Thresholds are constructed only on the 165\-episode policy\-selection split\. For horizonh∈\{0,…,12\}h\\in\\\{0,\\ldots,12\\\}and selection episodeii, define the best\-so\-far risk score and the relevant one\-indexed order statistic as
mi\(h\)=min0≤t≤hrθ\(Sit\),kq=⌈q\(165−1\)⌉\+1\.m\_\{i\}\(h\)=\\min\_\{0\\leq t\\leq h\}r\_\{\\theta\}\(S\_\{it\}\),\\qquad k\_\{q\}=\\left\\lceil q\(165\-1\)\\right\\rceil\+1\.For each targetq∈\{0\.72,0\.75,0\.78,0\.80,0\.82,0\.85,0\.88,0\.90,0\.93,0\.95\}q\\in\\\{0\.72,0\.75,0\.78,0\.80,0\.82,0\.85,0\.88,0\.90,0\.93,0\.95\\\}, the 165 valuesmi\(h\)m\_\{i\}\(h\)are sorted andτh,q\\tau\_\{h,q\}is set to theirkqk\_\{q\}th order statistic\. This is equivalent tonumpy\.quantile\(\.\.\., method="higher"\)\. The ten\(q,kq\)\(q,k\_\{q\}\)pairs are\(\.72,120\)\(\.72,120\),\(\.75,124\)\(\.75,124\),\(\.78,129\)\(\.78,129\),\(\.80,133\)\(\.80,133\),\(\.82,136\)\(\.82,136\),\(\.85,141\)\(\.85,141\),\(\.88,146\)\(\.88,146\),\(\.90,149\)\(\.90,149\),\(\.93,154\)\(\.93,154\), and\(\.95,157\)\(\.95,157\)\.
These targets are coverage indices used to construct the grid, not claims about calibration or population coverage\. In particular, they do not define ten global score cutoffs: every pair\(h,q\)\(h,q\)has its own horizon\-specific threshold\. Policy\(h,q\)\(h,q\)stops at the first staget≤ht\\leq hsatisfyingrθ\(Sit\)≤τh,qr\_\{\\theta\}\(S\_\{it\}\)\\leq\\tau\_\{h,q\}and otherwise defers athh\. By construction, its empirical autonomous coverage on the selection split is at leastqq, with possible excess coverage when risk scores are tied\.
Crossing 13 horizons with 10 coverage indices yields 130 deterministic candidates\. Every threshold is then reused unchanged on calibration and evaluation episodes\. Using only the selection split, the frozen pipeline additionally determines the tested 12\-policy family, its testing order, and the randomized\-mixture weights\. Candidate and mixture design use\(αdes,γdes\)=\(0\.20,0\.80\)\(\\alpha\_\{\\mathrm\{des\}\},\\gamma\_\{\\mathrm\{des\}\}\)=\(0\.20,0\.80\), whereas the subsequent joint tests use\(α,γ,δ\)=\(0\.25,0\.70,0\.05\)\(\\alpha,\\gamma,\\delta\)=\(0\.25,0\.70,0\.05\)\.
## Appendix CDetailed Baseline and Masking Results
Table 7:Calibration and evaluation results for common\-path controllers\. Jointppis the raw intersection–union value;Cros\-Mix, realized draw is shown using the prespecified frozen episode\-wise assignment, while analytic expectations are used in the paired tables\.All controllers in Table[7](https://arxiv.org/html/2609.09678#A3.T7)receive the same backbone diagnoses, action proposals, native stopping scores, and forced\-continuation observations\. Their differences therefore reflect stopping and deferral rather than alternative test acquisition\. The baseline calibrationpp\-values are descriptive and do not imply that each baseline belongs to the frozen multiplicity\-controlledCrosfamily\. ERM is inexpensive because its selected policy requests no additional action on evaluation, but it fails the joint calibration criterion\.
#### Diagnosis\-language masking\.
An aggregate regex screen identified explicit diagnostic and future\-information language in a subset of the initial presentations\. We therefore froze two ontology\-wide, label\-independent masking conditions before rerunning the backbone\. The diagnosis\-name condition removes prespecified names and synonyms for all nine diagnosis classes\. The expanded condition additionally removes prespecified diagnostic and future\-information phrases\. These conditions match 79/367 and 116/367 evaluation HPIs, respectively\.
No backbone parameter, ranker, threshold, horizon, candidate order, mixture weight, cost, or stopping policy is tuned on the masked calibration or evaluation outcomes\. Every controller within a masking condition receives the same condition\-specific common path\.
Table 8:Complete frozen evaluation masking outputs\. Every controller uses the same condition\-specific common path\.Cros\-Mix is reported in its analytic\-expectation evaluation mode\.Relative to the original inputs, diagnosis\-name and expanded masking changeCros\-Mix analytic\-expectation selective risk by\+\.015\+\.015\(95% CI\[−\.001,\.033\]\[\-\.001,\.033\]\) and\+\.013\+\.013\[−\.004,\.032\]\[\-\.004,\.032\], coverage by−\.018\-\.018\[−\.036,−\.0004\]\[\-\.036,\-\.0004\]and−\.020\-\.020\[−\.040,−\.001\]\[\-\.040,\-\.001\], and cost by\+\.504\+\.504\[\.066,\.968\]\[\.066,\.968\]and\+\.632\+\.632\[\.150,1\.130\]\[\.150,1\.130\], respectively\. The corresponding ranker\-AUROC changes are−\.014\-\.014\[−\.026,−\.005\]\[\-\.026,\-\.005\]and−\.014\-\.014\[−\.026,−\.003\]\[\-\.026,\-\.003\]\. Full\-workup error changes by\+\.022\+\.022\[\.003,\.044\]\[\.003,\.044\]and\+\.027\+\.027\[\.005,\.049\]\[\.005,\.049\]\.
The evaluation ranker AUROCs are \.853, \.839, and \.840 for the original, diagnosis\-name, and expanded masking conditions; the corresponding calibration AUROCs are \.840, \.842, and \.842\. All masked\-minus\-original intervals use 10,000 paired patient resamples\. Explicit\-language removal therefore weakens ranking and efficiency without collapsing the frozen controller\. However, deterministic masking cannot remove every implicit cue or establish that no diagnostic leakage remains\. The versioned artifacts include the regex list, replacement rules, split\-level match counts, and immutable input/output hashes\.
## Appendix DExpanded Paired Bootstrap and Trajectory Details
All paired entries reportCros\-Mix, analytic expectation minus the named comparator\. Intervals are 2\.5–97\.5 percentile intervals from 10,000 patient\-level paired bootstrap resamples\. All states and outcomes from one patient are resampled together; states are never resampled independently\. Replicates with no autonomous diagnoses are omitted only when the selective\-risk difference is undefined\.
Table 9:Cros\-Mix, analytic expectation minus comparator, using 10,000 patient\-level resamples: selective risk, coverage, and error mass\.Table 10:Cros\-Mix, analytic expectation minus comparator: total relative cost, requested actions, and proxy\-miss mass\.These paired comparisons support a narrow cost claim\.Cros\-Mix, analytic expectation is less costly thanCros\-Det, the deterministic controller returned by Fixed\-sequence LTT, confidence thresholding, and native stopping with deferral, but it is more costly than the Uniform\-weight mixture and ERM\. Its selective\-risk difference fromCros\-Det, the deterministic controller returned by Fixed\-sequence LTT, and the Uniform\-weight mixture is unresolved\. ERM is less costly but has higher selective risk and does not pass the exploratory joint calibration criterion\.
Figure 2:Forced\-continuation evaluation audit on the same 367 episodes at every stage\. Bands are patient\-bootstrap 95% intervals; diagnostic error becomes non\-monotone as cost and recorded\-result missingness accumulate\.#### Complete forced\-continuation trajectory\.
Table[11](https://arxiv.org/html/2609.09678#A4.T11)reports the numerical trajectory underlying Figure[2](https://arxiv.org/html/2609.09678#A4.F2)\. Every row contains the same 367 patients\. “New\-action missing” is the fraction of actions requested at that stage that returnNO\_RECORDED\_RESULT\.
Table 11:Forced\-continuation evaluation trajectory\. All metrics have patient\-bootstrap intervals in the released CSV\.Diagnostic error reaches its minimum after one requested action and subsequently becomes non\-monotone\. Because patient composition is identical across stages, the pattern is not caused by different patients remaining at later horizons\. The degradation occurs alongside increasing cumulative cost and high recorded\-result missingness\. It therefore characterizes the frozen backbone and retrospective benchmark rather than establishing that clinical testing is generally harmful\.
## Appendix EExact Bounds and Frozen Frontiers
#### Exact component and jointpp\-values\.
Table[12](https://arxiv.org/html/2609.09678#A5.T12)reports the risk and coverage components of the intersection–union test\. The joint value ispj=max\(pR,j,pC,j\)p\_\{j\}=\\max\(p\_\{R,j\},p\_\{C,j\}\)and is small only when both requirements receive sufficient evidence\.
Table 12:Exact component and joint calibrationpp\-values, computed from unrounded episode counts\.The Single\-candidate test and the separately frozenCros\-Mix draw are each tested once at level \.05\. Fixed\-sequence LTT follows its prespecified order and passes the reported candidate when that candidate is reached\. Holm testing returns the same deterministic controller as Fixed\-sequence LTT in this run\. Bonferroni testing uses threshold\.05/12=\.00417\.05/12=\.00417; its returned controller has unrounded raw joint value 0\.00044157, giving adjustedp=min\{1,12p\}=0\.00530p=\\min\\\{1,12p\\\}=0\.00530\. These are exploratory calibration calculations because the calibration labels are not prospectively untouched\.
#### Mixture weights and randomization stability\.
The frozen mixture places weights 0\.322, 0\.044, and 0\.633, rounded to three decimal places, on\(h=1,τ=0\.3805\)\(h=1,\\tau=0\.3805\),\(h=1,τ=0\.4506\)\(h=1,\\tau=0\.4506\), and\(h=3,τ=0\.3286\)\(h=3,\\tau=0\.3286\), respectively\. The printed weights sum to 0\.999 because of rounding; sampling and analytic integration use the full\-precision normalized weights\.
The displayedCros\-Mix realized draw produces 280 autonomous diagnoses with 51 errors on calibration and 288 autonomous diagnoses with 47 errors on evaluation\. Exact testing uses these realized episode\-level outcomes\. Analytic mixture quantities instead integrate each episode’s component contributions over the frozen weights\.
Table 13:Two evaluation modes ofCros\-Mix and stability across 1,000 episode\-wise mixture seeds\. Brackets in the final row give 2\.5–97\.5 percentiles\.All 1,000 evaluation realizations have point estimates below the numerical risk target and above the numerical coverage target\. This post\-freeze analysis measures sensitivity to the episode\-wise random draws; it is not used to select a favorable seed and does not create 1,000 independent calibration experiments\.
#### One\-sided exact bounds\.
The following bounds use the realized episode\-level outcomes for randomized policies\. Evaluation bounds are descriptive because the evaluation labels were previously viewed\.
Table 14:One\-sided 95% Clopper–Pearson bounds\. The support\-restricted deterministic choice equalsCros\-Det and is not duplicated\.The bounds concern marginal population risk and coverage and do not imply disease\-class or demographic control\.
#### Frozen evaluation frontiers\.
Figure[3](https://arxiv.org/html/2609.09678#A5.F3)displays the complete candidate grids for the supported risk\-score families\. The marked operating points and matched\-coverage policies are selected using development or selection data rather than visually favorable evaluation outcomes\.
Figure 3:Exploratory frozen frontiers\. Curves contain the complete candidate grid for the five supported score families; marked operating points and the 0\.80 matched\-coverage policies were selected on development or selection data\.
## Appendix FCost, Missingness, and Resplit Details
#### Frozen\-decision cost sensitivity\.
The cost analysis changes the accounting vector while retaining the original backbone trajectories, stopping decisions, deferral outcomes, thresholds, and mixture weights\. Risk, coverage, and requested actions therefore remain unchanged\.
Table 15:Frozen\-decision scenario costs\. Low review, primary, and high review use deferral penalties 5, 10, and 25, respectively\. Imaging2×2\\timesand missing2×2\\timesdouble the corresponding action\-cost components\.ForCros\-Mix, analytic expectation, doubling laboratory costs gives total cost 5\.77\. Across the full deferral\-penalty grid\{5,10,15,20,25,30\}\\\{5,10,15,20,25,30\\\}, the break\-even penalty against native stopping is approximately 22\.15\.Cros\-Mix, analytic expectation remains less costly thanCros\-Det, the deterministic controller returned by Fixed\-sequence LTT, confidence thresholding, native stopping, and native stopping with deferral in the primary, imaging\-sensitive, and missingness\-sensitive scenarios\. Native stopping becomes less costly under the deferral penalty of 25, while ERM remains less costly throughout but fails exploratory calibration\.
Figure 4:Frozen\-decision sensitivity\. Positive bars in \(b\) favorCros\-Mix\. Risk, coverage, and requested actions remain fixed; only cost accounting changes\.
#### Action availability and informative missingness\.
Action availability varies because an action is marked available only when a corresponding result is found in the retrospective record\.
Table 16:Evaluation action availability and final common\-path diagnosis error\. Small available or missing\-result denominators should not be overinterpreted\.A logistic model using only action\-availability indicators has diagnosis macro\-AUROC 0\.597 and state\-error AUROC 0\.535\. Missingness therefore contains weak information about diagnosis and backbone error\. This is consistent with the record encoding clinicians’ historical ordering behavior, although it does not identify the causal mechanism producing the missingness\.
#### Development\-resplit audit\.
For each of 20 development\-only resplits, the 935/165 division is recreated, the ranker is refitted, and the policies are redesigned using cached backbone trajectories\. The outer calibration and evaluation cohorts remain unchanged\.
Table 17:Twenty development\-only repeated splits\. Wilson intervals are reported for rates; metric ranges summarize all available policies\. Infeasible mixture runs remain failures in the reported rates\.TheCros\-Mix LP is infeasible for seeds 20260911, 20260913, 20260915, 20260923, and 20260924\. Failures remain in the denominator rather than being discarded\. The full run\-level distribution, including unsuccessful runs, accompanies the aggregate artifacts\. Because the same outer calibration and evaluation cohorts are reused, these scores are dependent exploratory checks rather than 20 independent confirmations\. The audit was not used to replace or modify the primary frozen policy\.
## Appendix GDisease\-Class Audit
Table[18](https://arxiv.org/html/2609.09678#A7.T18)reports the complete disease\-class results for the prespecifiedCros\-Mix realized\-draw evaluation seed\. The values are descriptive, are not multiplicity\-controlled, and are not class\-conditional certificates\.
Table 18:Complete nine\-class audit forCros\-Mix, realized draw\.Coverage is below 70% for biliary disease, diverticulitis, and nonspecific abdominal pain\. Selective error is 0\.500 for diverticulitis and 1\.000 for the small nonspecific\-abdominal\-pain subgroup\. The denominators are too small to support simultaneous class\-level certificates atα=\.25\\alpha=\.25while maintaining useful coverage\.
These results do not contradict the marginal population calculation: Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1)concerns aggregate selective risk and coverage unless subgroup constraints are explicitly incorporated into the frozen testing family\. A future confirmatory design should first certify marginal risk, then test prespecified clinically meaningful or high\-harm strata with an allocated multiplicity budget\. Groups lacking adequate sample size should be reported as unsupported rather than safe\.
## Appendix HFrozen Protocol and Confirmatory Extension
#### Backbone implementation\.
The diagnostic backbone is Qwen2\.5\-7B\-Instruct\[[Qwen Team, 2024](https://arxiv.org/html/2609.09678#bib.bib15)\]with LoRA adaptation\[[Hu et al\., 2022](https://arxiv.org/html/2609.09678#bib.bib14)\]using rank 8, scaling parameter 16, and dropout 0\.1\. Following the official LA\-CDM implementation at frozen repository commit3f435a1, adaptation proceeds through decision\-agent GRPO, hypothesis\-agent supervised fine\-tuning, and confidence\-calibration GRPO\. Training runs for four epochs and 3,740 optimizer steps with seed 269\.
The rehashed training manifest references only the 935\-episode ranker\-training partition and the 165\-episode development\-selection partition; the canonical calibration and evaluation files are not loaded by the backbone\-training pipeline\. The resulting model produces the diagnosis distribution, diagnosis proposal, next\-action proposal, native stopping score, and other logged quantities subsequently used by all stopping controllers\. Our implementation is an official\-code\-derived transfer of the LA\-CDM training structure and prompt conventions\. It should not be interpreted as a prompt\-equivalent reproduction of LA\-CDM’s original free\-form trajectories\.
#### Common\-path stopping protocol\.
For each episode, the frozen backbone is first evaluated along a maximum forced\-continuation trajectory of 12 action opportunities\. At each stage, the trace records the current information state, diagnosis probabilities, proposed diagnosis, native stopping score, proposed next action, returned observation, and accumulated relative cost\. Continuing reveals the result associated with the next backbone\-proposed action, includingNO\_RECORDED\_RESULTwhen no corresponding result is found\. No stopping controller substitutes a different action\.
All compared controllers are then applied to the same stored backbone trajectory\. A controller determines only whether to accept the current diagnosis, continue with the backbone proposal, or defer\. Its requested\-action count and cost are truncated at its terminal stage, with the deferral penalty added when applicable\. Thus paired differences among the primary controllers and comparators isolate stopping and deferral decisions while holding the backbone’s diagnoses, action proposals, and potential recorded observations fixed\. They do not compare alternative test\-acquisition strategies\.
#### Stopping comparators\.
The common\-path comparison separates the two formalCroscontrollers from baselines, testing procedures, and ablations as follows\.
#### Fixed\-information comparators\.
HPI\-only accepts the stage\-zero diagnosis without requesting an action\. Full workup continues through all 12 action opportunities before accepting the terminal diagnosis\. The fixed\-stage comparator accepts at one common stage chosen on the development selection split\. These study\-defined controls compare fixed information budgets and are not implementations of external methods\.
#### Confidence\-based comparators\.
The confidence\-threshold controller stops when the backbone’s maximum diagnosis probability crosses its selection\-frozen threshold and otherwise defers at its horizon\. The maximum\-probability ranker directly uses maximum class probability as the state\-ranking score, while the entropy\-margin ranker uses predictive entropy and the gap between the two largest diagnosis probabilities\. These are study\-implemented confidence heuristics inspired by standard selective prediction and abstention\[[Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.09678#bib.bib10),[Geifman and El\-Yaniv, 2019](https://arxiv.org/html/2609.09678#bib.bib11)\]; they are not exact reproductions of those papers’ training procedures\.
#### Native\-agent comparators\.
LA\-CDM native follows the backbone’s own stopping output\. Native plus defer applies an additional selection\-frozen confidence requirement to the diagnosis produced at the native stopping stage; cases failing that requirement are deferred\. The native component is an official\-code\-derived transfer of LA\-CDM\[[Bani\-Harouni et al\., 2026](https://arxiv.org/html/2609.09678#bib.bib2)\]to the fixed common\-path setting, not a prompt\-equivalent reproduction of its free\-form trajectories\. Native plus defer is our derivative baseline combining that native stop with confidence\-based deferral; it is not a named method from the LA\-CDM paper\.
#### Empirical cost minimization\.
Empirical ERM selects the policy with lowest empirical mean cost on the policy\-selection split without requiring the joint selective\-risk and autonomous\-coverage criterion used byCros\. It tests whether empirical cost optimization alone sacrifices risk control\.
#### FormalCroscontrollers\.
Cros\-Det is the least\-cost deterministic threshold\-and\-horizon candidate satisfying the selection\-stage design constraints\.Cros\-Mix is the LP\-optimized episode\-wise mixture defined by Eq\. \([9](https://arxiv.org/html/2609.09678#S3.E9)\) over the frozen deterministic family\. In the primary freeze, the best component in the optimized support is identical toCros\-Det and is not displayed separately\.
#### Calibration and multiplicity procedures\.
The Single\-candidate test evaluates one policy fixed before calibration\. Fixed\-sequence LTT applies the prespecified LTT order\[[Angelopoulos et al\., 2025](https://arxiv.org/html/2609.09678#bib.bib8)\]; Holm testing applies Holm’s step\-down rule\[[Holm, 1979](https://arxiv.org/html/2609.09678#bib.bib13)\]; and Bonferroni testing applies the classical Bonferroni threshold to the same frozen 12\-policy family\. Holm testing returns the same deterministic controller as fixed\-sequence LTT in this run, whereas Bonferroni testing returns a different deterministic controller\. These procedures alter which candidate is tested or returned, not its ranker, backbone, or basic stopping rule\. Pareto Testing is related multi\-risk methodology\[[Laufer\-Goldshtein et al\., 2023](https://arxiv.org/html/2609.09678#bib.bib16)\], but is not a separate controller in the reported table\.
#### Ablations\.
The No\-history ranker removes stage, missing\-result fraction, cumulative cost, and latency while retaining probability\-derived features; it tests whether history features improve state\-error ranking and stopping\. The Native\-score ranker directly uses the backbone’s native stopping score; it tests whether the learned ranker improves on that score\. The Uniform\-weight mixture assigns equal probability to the same frozen support asCros\-Mix; it tests whether optimized weights improve on support selection alone\. These ablations neither retrain nor alter the diagnostic backbone\.
#### Mixture evaluation modes and legacy identifiers\.
Cros\-Mix, analytic expectation integrates each episode’s support contributions, whereasCros\-Mix, realized draw uses one frozen episode\-wise component assignment\. They are evaluations of one controller, not separate methods\. Released machine\-readable files retainBESTDET,SUPPORTBEST,OPTIMIZEDMIX, andUNIFORMMIXfor artifact compatibility; their display meanings areCros\-Det, the support component identical toCros\-Det in this run,Cros\-Mix, and the Uniform\-weight mixture, respectively\.
The exact thresholds, horizons, 12\-policy testing family, candidate order, and mixture weights are frozen using only the designated development\-selection data and are stored in the run manifest\. Comparators that select a threshold, stage, horizon, or policy use the same development\-selection partition\.
#### Episode\-wise mixture randomization\.
For a frozen mixture with weightsww, one component policy is sampled independently at the beginning of each episode and is retained for the entire episode\. Sampling is not conditioned on patient characteristics, backbone predictions, intermediate observations, or calibration outcomes\. The base policy seed is 20260902; the displayed calibration and evaluation realizations use the frozen split\-specific seeds 20260904 and 20260905, respectively\.
The exact binomial test treatsCros\-Mix, realized draw as one frozen policy and uses its realized autonomous and error indicators\.Cros\-Mix, analytic expectation instead averages each episode’s contributions over the frozen component weights\. It reduces Monte Carlo noise in descriptive cost and paired comparisons, but is not substituted for the realized Bernoulli outcomes in the exact calibration test\. The separate 1,000\-seed analysis measures randomization stability and is not used to select a favorable seed\. Drawing one global component and applying it to every calibration episode would create shared dependence and would not justify the same binomial calculation\.
#### Metrics and uncertainty\.
For a controller producingMMautonomous diagnoses andEEerrors amongnnepisodes, the primary diagnostic quantities are selective errorE/ME/M, autonomous coverageM/nM/n, and error massE/nE/n\. Selective error is undefined whenM=0M=0\. Error mass is reported alongside selective error so that changes in the number of deferred cases do not conceal the absolute frequency of autonomous mistakes\.
Resource outcomes are total relative cost and the number of attempted actions\. An attempted action is counted even when it returnsNO\_RECORDED\_RESULT\. Proxy\-miss mass is the proportion of episodes with a wrong autonomous prediction whose reference label is one of the prespecified disease\-specific classes other than gastroenteritis/colitis or nonspecific abdominal pain\. This endpoint is descriptive and is not a clinician\-adjudicated measure of harm or diagnostic urgency\.
Paired contrasts use 10,000 patient\-level bootstrap resamples\. All states and outcomes belonging to the same patient are resampled together; state\-level resampling is never used\. Replicates with no autonomous diagnoses are excluded only from contrasts for which selective error is mathematically undefined\. Analytic mixture comparisons integrate each patient’s component\-level contributions before patient\-level resampling\. A separate 1,000\-seed analysis describes the variability of finite mixture realizations\.
For deterministic policies and the prespecifiedCros\-Mix realized draw, we also report one\-sided 95% Clopper–Pearson upper bounds for selective risk and lower bounds for coverage\. Evaluation bounds remain descriptive in the present study because those labels were previously viewed\. Frozen\-decision sensitivity analyses change the cost vector without rerunning the backbone, refitting the ranker, reconstructing thresholds, or changing controller decisions\.
#### Freeze order\.
The recorded pipeline proceeds in the following order: backbone adaptation on development data; episode\-wise cross\-fitting and final risk\-ranker fitting on the 935\-patient ranker\-training subset; threshold, candidate\-family, candidate\-order, and mixture design on the disjoint 165\-patient policy\-selection subset; and finally calibration and evaluation\. Candidate ranking and the 12\-policy multiple\-testing order use selection data only within the recorded pipeline\.
The policy, ranker, thresholds, candidate order, mixture weights, cost vector, evaluation code, and randomization mechanism are fixed before the reported calibration and evaluation computations under the current freeze\. This procedural freeze makes the reported analyses reproducible, but it does not restore prospective independence after labels have previously been accessed\.
#### Recorded provenance\.
The run stores model and adapter identifiers, repository commit, training and generation configuration, dataset and split hashes, label code, action\-cost vector, ranker configuration, all candidate policies, the 12\-policy testing family and its order, mixture weights, base seed 20260902, split\-specific realization seeds, evaluator version, and per\-episode traces\. Thresholds are stored numerically rather than recomputed during calibration or evaluation\.
Two independently implemented analysis scripts reproduce the reported summary files and verify split disjointness\. The audit recomputes zero patient intersections among the 935 ranker\-training, 165 policy\-selection, 367 calibration, and 367 evaluation partitions\. Because cohort construction retains the earliest qualifying stay before partitioning, each episode corresponds to one patient\. Derived analysis artifacts contain no patient identifiers\.
#### Separate post\-freeze exploratory branch\.
A separately constructed cohort branch uses partitions of 543/85/190/199 episodes and is not pooled with the canonical cohort, used to modify its policies, or included in its primary tables\. In that branch, Primary\-RobustDet is point\-estimate dominated by the branch\-specificCros\-Mix in selective risk, coverage, and cost\. Full\-Ensemble\-UCB has AURC 0\.1290, compared with 0\.1289 for the No\-history ranker, and therefore does not improve the ranking result\. Among the corresponding contrasts, only the cost contrast is statistically resolved\. We did not continue tuning either variant after observing these results\. This branch is reported as negative exploratory evidence, not as an independent confirmation of the canonical analysis\.
#### Required confirmatory freeze\.
Before accessing a new confirmatory cohort, a signed and timestamped manifest must fix the cohort dates, inclusion and exclusion rules, one\-episode\-per\-patient construction, diagnosis\-label mapping, timestamp interpretation, handling of missing or unavailable tests, adjudication procedures, harm and fairness strata, action and deferral costs, model and adapter hashes, generation settings, ranker, thresholds, candidate family and order, mixture weights and episode\-wise randomization mechanism,\(α,γ,δ\)\(\\alpha,\\gamma,\\delta\), multiplicity graph, sample size, stopping rule for data collection, and executable analysis code\.
Calibration labels must remain inaccessible to model fitting, candidate construction, policy ordering, mixture optimization, seed selection, and sample\-size revision\. Any amendment made after calibration outcomes are accessed creates a new exploratory analysis and requires another untouched cohort for a confirmatory claim\. Freezing code after examining labels does not recreate the independence assumed by Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1)\.
#### Confirmatory decision rule\.
The confirmatory calibration set should be opened once and evaluated using the preregistered candidate family and testing graph\. A policy is certified only if its union null is rejected by the frozen family\-wise\-error procedure at levelδ\\delta\. If no candidate is rejected, the confirmatory outcome is that no policy is certified; the investigators should not replace it with the empirically best failed candidate\.
Any subsequent evaluation cohort should be used only after the calibration decision has been finalized\. Its role is to estimate performance, subgroup behavior, resource use, and distribution shift, rather than to revise the certified policy or repeat the certificate test\.
#### Sample\-size planning\.
Power must be determined before the confirmatory freeze using development estimates and prespecified worst\-case margins\. Planning should simulate the joint distribution of autonomous diagnoses and autonomous errors because the selective\-risk denominator is itself random\. It should also reproduce the intended candidate order, intersection\-union tests, multiplicity rule, mixture randomization, and any subgroup or harm constraints\.
Althoughn=367n=367is sufficient to reject several candidate nulls in the present exploratory calculations, it does not establish adequate power under smaller risk or coverage margins, temporal shift, multiplicity, subgroup constraints, or clinician\-adjudicated harm endpoints\. Sample size must follow the complete testing graph rather than repeated inspection of certificate outcomes\. Early stopping, sample\-size extension, or repeated testing requires a separately valid sequential design\.
#### End\-to\-end agent study\.
A secondary experiment should allow each agent to choose its own actions rather than restricting all methods to a common trajectory\. The comparison should include fixed\-stage, confidence, ERM, myopic value of information, official LA\-CDM, LTT,Cros\-Det,Cros\-Mix, and compatible released implementations of AgentClinic, MedChain, and DxChain\. It should report invalid or unavailable actions, repeated actions, token use, latency, diagnostic accuracy, selective error, error mass, autonomous coverage, resource cost, and trace validity\. Clinician adjudication should additionally assess whether requested tests and terminal decisions are medically appropriate\.
Because independently acting agents generate different histories, their outcomes cannot be interpreted as paired stopping comparisons on a common information path\. Such an experiment evaluates the combined acquisition, reasoning, stopping, and deferral system and therefore addresses a broader question than the primary experiment\.
#### Scope of the theoretical claim\.
Theorem[1](https://arxiv.org/html/2609.09678#Thmtheorem1)controls the probability of certifying an invalid frozen policy only under the stated exchangeability and independent\-calibration assumptions\. It does not establish causal clinical benefit, correctness of the reference labels, realism of the relative cost vector, optimality outside the frozen candidate family, subgroup or hospital\-level robustness, or deployment safety\.
Repeated episodes from the same patient, clinician, or institution would require the sampling unit and exchangeability assumptions to be redefined, together with a corresponding cluster\-aware calibration procedure\. Settings in which an action changes patient physiology, future observability, clinical management, or documentation require a causal or interactive environment rather than the present logged trajectory\. These limitations define the boundary of the theoretical and empirical claims\.Similar Articles
CDPR: Counterfactual Advantage-based Credit Assignment for Cost-Aware Sequential Medical Diagnosis
CDPR is a counterfactual advantage-based credit assignment method for training cost-aware sequential medical diagnosis models using reinforcement learning, improving accuracy while reducing examination costs and number.
Unmasking Toxic Mimicry in Medical Offline Reinforcement Learning for ICU Sepsis Management via Counterfactual Clinical Audits
This paper introduces the Counterfactual Clinical Audit (CCA) framework to evaluate offline reinforcement learning agents for ICU sepsis management, exposing 'toxic mimicry' where agents replicate harmful treatment patterns that standard metrics miss. Using MIMIC-III data, it shows a Medical Decision Transformer fails to escalate vasopressors under rising lactate, while a causal transformer performs safely.
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
MedDDC-Eval introduces a diagnosis-decoupled evaluation testbed for multi-turn medical consultation agents, isolating the policy-elicited conversation history from diagnosis generation to enable cleaner measurement of evidence acquisition and diagnostic usefulness.
ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science
ClinLens is a new benchmark of 200 executable clinical data-science tasks over five linked MIMIC resources, evaluating long-horizon coding agents on longitudinal multimodal data. Results show strong code execution but poor clinical analysis correctness, highlighting a gap between runnable submissions and valid analyses.
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Introduces Safe-Psych, a sequential benchmark for evaluating how large language models handle diagnostic uncertainty in psychiatry, revealing that even strong models often fail to abstain or seek clarification when clinical evidence is incomplete.