Certifying when decision-time information justifies adaptive experimentation

arXiv cs.LG 论文

摘要

This paper introduces Opal (Opportunity-aware Policy Authorization for Laboratories), a framework that certifies whether adaptive experimentation should be enabled by precommitting to non-trivial adaptation, controlled target risk, and positive executed value after cost. It establishes an impossibility boundary and demonstrates the method on a Cell Painting dataset, achieving risk control and positive value.

arXiv:2607.27651v1 Announce Type: new Abstract: Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted. We introduce Opportunity-aware Policy Authorization for Laboratories (\OPAL{}), a framework that decides whether adaptation should be enabled at all. \OPAL{} uses a precommitted contract to require non-trivial adaptation, controlled target risk and positive executed value after cost. We establish an impossibility boundary: source outcomes and unlabelled target covariates cannot uniformly support non-trivial authorization under unrestricted conditional outcome shift, and derive a target-calibrated recovery. Applied to an unseen 11,265-compound Cell Painting partition, the frozen gate selected 595 compounds, captured 384 positive opportunities and achieved strictly positive executed value under least-favourable completion; its 5.18\% false-activation upper bound remained below a 7.5\% limit. Among six methods, only \OPAL{} combined non-zero activation with this risk control. Locked pharmacogenomic and finite-campaign studies distinguish policy misalignment from non-certifiability, establishing authorization as a distinct layer for safe adaptive science.
查看原文
查看缓存全文

缓存时间: 2026/07/31 10:04

# Certifying when decision-time information justifies adaptive experimentation
Source: [https://arxiv.org/html/2607.27651](https://arxiv.org/html/2607.27651)
\[1\]\\fnmJia\\surBi \[3\]\\fnmChenyang\\surZhu

\[1\]\\orgdivScientific Computing Department,\\orgnameScience and Technology Facilities Council,\\orgaddress\\streetRutherford Appleton Laboratory,\\cityDidcot,\\postcodeOX11 0QX,\\countryUK 2\]\\orgdivDiamond Light Source,\\orgaddress\\streetHarwell Science and Innovation Campus,\\cityDidcot,\\postcodeOX11 0QX,\\countryUK \[3\]\\orgdivSchool of Electronics and Computer Science,\\orgnameUniversity of Southampton,\\orgaddress\\citySouthampton,\\postcodeSO17 1BJ,\\countryUK

###### Abstract

Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted\. We introduce Opportunity\-aware Policy Authorization for Laboratories \(Opal\), a framework that decides whether adaptation should be enabled at all\.Opaluses a precommitted contract to require non\-trivial adaptation, controlled target risk and positive executed value after cost\. We establish an impossibility boundary: source outcomes and unlabelled target covariates cannot uniformly support non\-trivial authorization under unrestricted conditional outcome shift, and derive a target\-calibrated recovery\. Applied to an unseen 11,265\-compound Cell Painting partition, the frozen gate selected 595 compounds, captured 384 positive opportunities and achieved strictly positive executed value under least\-favourable completion; its 5\.18% false\-activation upper bound remained below a 7\.5% limit\. Among six methods, onlyOpalcombined non\-zero activation with this risk control\. Locked pharmacogenomic and finite\-campaign studies distinguish policy misalignment from non\-certifiability, establishing authorization as a distinct layer for safe adaptive science\.

###### keywords:

autonomous laboratories, adaptive experimentation, value of information, safe policy learning, finite\-population inference, scientific machine learning

## 1Introduction

Autonomous laboratories combine instruments, robotics and machine learning to choose measurements during an experiment\[hase2019,burger2020,szymanski2023,dai2024robots,harris2025oversight,scheurer2025human\]\. Most adaptive algorithms begin from an implicit premise: the option to adapt has already been granted\. At large scientific facilities, however, instrument modes, calibration, staffing and experimental capacity may be committed before the evidence needed to personalize those choices arrives\[noack2021,ament2021\]\. The prior decision is therefore whether the available evidence justifies enabling an adaptive branch at all\.

Existing methods solve important parts of this problem\. Active learning, Bayesian optimization and experimental design choose the next measurement\[lindley1956,chaloner1995,shahriari2016,rainforth2024\]; safe exploration and policy\-improvement methods constrain departures from a baseline\[wu2016conservative,sui2015safeopt,thomas2015,thomas2019seldonian,laroche2019spibb,cho2025cspimt\]; and value\-of\-information, selective prediction and calibrated\-risk methods quantify evidence value or risk–coverage trade\-offs\[fenwick2020voi,chow1970,elyaniv2010selective,angelopoulos2025,angelopoulos2024crc\]\. None alone determines whether scientific opportunity, decision\-time identifiability, finite\-campaign evidence, executed value, non\-trivial use and cost jointly warrant authorization\.Opalmakes that joint decision the primary object of a single precommitted contract\.

Source\-to\-target shift makes this distinction unavoidable\. Source outcomes and unlabelled target covariates do not identify risk or executed value on the target activated region without target outcomes or a transport restriction\[bendavid2010domains\]\. We formalize this boundary for non\-trivial authorization and derive a target\-calibrated recovery that separates labelled target development, pre\-outcome assignment and one\-shot final evaluation\.

Our contribution is fourfold\. First,Opaldefines information\-indexed opportunity and states when delayed commitment or an additional probe can change no optimal action\. Second, it supplies simultaneous fallback protection and an exact campaign\-size–error\-budget boundary for certifying the unobserved complement of a finite experiment\. Third, it requires risk control, a positive executed\-value lower bound and a prespecified minimum activation or sensitivity, thereby excluding the otherwise safe but uninformative always\-fallback policy\. Fourth, it links operational gain to information and reuse cost, so a technically safe policy is not automatically an adoptable one\.

The empirical studies follow this authorization chain and distinguish failure modes rather than pooling them into one score\. A target\-calibrated Cell Painting gate recovered a non\-trivial, positive\-value active branch and was the only activating method in a six\-method comparison to remain below the false\-activation limit\. The contract nevertheless withheld authorization because one confidence bound missed its target\. A separately locked pharmacogenomic study shows measured opportunity missed by a source\-trained score, while exact finite\-campaign and held\-out simulator studies separate non\-certifiability from risk\-controlled activation\. All external evaluations are offline\.

## 2Results

We first report the end\-to\-end target\-calibrated test, then use four diagnostic studies to isolate where authorization can fail: executed value, finite\-campaign certifiability, activation\-risk calibration or source\-to\-target policy alignment \(Supplementary TableLABEL:SI\-tab:study\_design\)\.

### 2\.1Adaptive capacity as a certification problem

![Refer to caption](https://arxiv.org/html/2607.27651v1/x1.png)Figure 1:Certifying adaptive opportunity under asynchronous evidence and commitment\.Figure[1](https://arxiv.org/html/2607.27651#S2.F1)separates four questions that a single performance score cannot: whether heterogeneous actions have value, decision\-time measurements expose that value, the available evidence can certify the deployment population, and the executed gain repays information and implementation cost\. Failure at any link returns the experiment to its study\-specific best fixed capacityK0K\_\{0\}, not to “no experimentation”\.

The propositions mirror this chain: information bounds attainable gain, simultaneous inference protects fallback, two distinct impossibility results separate target non\-identifiability from finite\-campaign non\-certifiability, andH⋆H^\{\\star\}separates operational improvement from adoption\. The empirical layers below identify which link fails\.

### 2\.2Target calibration recovers a valuable active branch

Target\-labelled development recovered the active branch that source\-only transfer had not\. In the cpg0012 Cell Painting archive, the decision was whether one initial compound profile justified acquiring two more; an additional replicate provided evaluation\-only verification\[wawer2014profiling,bray2017cellpaintingdata\]\. The gate could use only structure, initial morphology and pre\-imaging metadata\. After development on 13,748 labelled compounds, one rule and all final assignments were sealed before anyFINAL\_LOCKoutcome was accessed \(Supplementary SectionLABEL:SI\-sec:sr\_cpg0012\)\.

The frozen gate produced a sparse, high\-precision active branch on the previously unopened 11,265\-compound final partition\. It activated 595 compounds \(5\.28%\), of which 384 \(64\.5%\) were measured POSITIVE opportunities\. Among compounds counted as NULL under the least\-favourable failure convention, the false\-activation upper bound was 5\.18%, below its 7\.5% limit\. The gate recovered 5\.88% of POSITIVE compounds, with a 5\.41% lower bound above a 5% non\-degeneracy floor\. That floor was defined using the two labelled pre\-final development partitions and then frozen in the amended v5 contract before anyFINAL\_LOCKoutcome was accessed\. At this archive scale, it required recovery of hundreds of positives while allowing the contract to prioritize false\-activation control over broad recall\. After assigning every missing active outcome its registered lower value, the exact final\-archive mean gain was1\.948×10−3\>01\.948\\times 10^\{\-3\}\>0; its compound\-bootstrap BCa lower endpoint was1\.229×10−3\>01\.229\\times 10^\{\-3\}\>0\(Fig\.[6](https://arxiv.org/html/2607.27651#S2.F6)c\)\.

The complementary active\-set question was stricter: what fraction of activations could be false? Here the false\-discovery point estimate was206/595=34\.62%206/595=34\.62\\%, below the internal 35% target, but its one\-sided 95% upper bound was 37\.97%\. Six of seven component checks therefore passed, yet the conjunctive rule withheld the certificate\. This is the central governance result: non\-trivial activation, controlled false activation and positive executed value were observed, but a favourable point estimate could not replace its prespecified uncertainty bound\.

Always\-fallback had no sensitivity or value, whereas forced activation, a simple uncertainty rule, expected net benefit and a risk\-only gate all violated the same false\-activation limit\.Opalwas the only activating method among the six evaluated to remain below that limit \(Fig\.[6](https://arxiv.org/html/2607.27651#S2.F6)d; Supplementary TableLABEL:SI\-tab:cpg0012\_comparators\)\. No method met the complete contract\.

### 2\.3Execution audit bounds adaptive value

![Refer to caption](https://arxiv.org/html/2607.27651v1/x2.png)Figure 2:Correct execution, baseline identity and information cost determine apparent adaptive value\.Correcting the estimand overturned the apparent adaptive gain\. A label\-based surrogate suggested a5\.549×10−35\.549\\times 10^\{\-3\}improvement against its score\-construction reference,Kscore,−f=1K\_\{\\mathrm\{score\},\-f\}=1\. Once the actual policy was executed and compared with the best fixed action,K0=2K\_\{0\}=2, the contrast became−8\.14×10−4\-8\.14\\times 10^\{\-4\}\(Fig\.[2](https://arxiv.org/html/2607.27651#S2.F2)a,b and Extended Data Fig\.[1](https://arxiv.org/html/2607.27651#S5.F1)\)\. The dominant correction was baseline identity \(−12\.862×10−3\-12\.862\\times 10^\{\-3\}\); the smaller value\-source and execution terms completed the accounting identity\. Thus the initial positive number described a surrogate comparison, not deployable value\.

None of the six thresholds applied to the common cross\-fitted score had a simultaneous lower bound above zero\. The safe wrapper therefore used its intended active safeguard and returned control toK0K\_\{0\}, rather than selecting the least unfavourable learned rule\. A separate full\-state audit found a positive opportunity envelope of2\.690×10−32\.690\\times 10^\{\-3\}, so the result cannot be read as an absence of heterogeneous action value\. Because the observable\-information optima were not identified, the remaining gap cannot be assigned uniquely to insufficient decision\-time information or model error\. With no positive executed contrast, no finite reuse boundaryH⋆H^\{\\star\}existed\.

#### Probe and expert diagnostics\.

The architecture diagnostics explained why the active stack added no value\. Cross\-validation selected no residual expert in any fold, and only 10 of 810 banks purchased a probe; two probes changed capacity\. After information cost, the realized probe contribution was−5\.86×10−4\-5\.86\\times 10^\{\-4\}, with a simultaneous interval ending at zero \(Fig\.[2](https://arxiv.org/html/2607.27651#S2.F2)c,d and Extended Data Fig\.[2](https://arxiv.org/html/2607.27651#S5.F2)\)\. This is the decision\-theoretic signature of a low\-value probe: it rarely changes the optimizer, so its information cost dominates\. The audit therefore rejects unsupported complexity rather than manufacturing a positive result from a different baseline or value source\.

### 2\.4Finite campaigns create an exact certification boundary

![Refer to caption](https://arxiv.org/html/2607.27651v1/x3.png)Figure 3:Exact finite\-campaign inference exposes a certification boundary rather than an estimator failure\.Opportunity was present in the mechanism population, but the 16\-bank campaign could not certify a switch for the units that would receive it\. We inverted the complete three\-class likelihood for every registered pilot size, noise level and possible signal\-count vector\. Across all 1,113 configurations, a least\-favourable campaign composition retained non\-positive action gain\. No observation was actionable, the uncosted certification envelope reached only zero and the costed envelope remained negative \(Fig\.[3](https://arxiv.org/html/2607.27651#S2.F3)a\)\. This is a property of the registered evidence population and decision rule, not an optimization failure\.

The general boundary explains how the design can change\. Even an all\-actionable pilot may leave all counterexamples in the unobserved complement\. Proposition[7](https://arxiv.org/html/2607.27651#Thmproposition7)gives the exact\(N,q,ρ,α\)\(N,q,\\rho,\\alpha\)condition under which that most favourable observation can first be certified\. For a strict\-majority claim at the registeredα=0\.02/7\\alpha=0\.02/7,N=16N=16is excluded; the first campaign size not excluded by the noiseless binary condition isN=25N=25, withq=14q=14\(Fig\.[4](https://arxiv.org/html/2607.27651#S2.F4)\)\. This necessary boundary does not guarantee that the noisy three\-class procedure will succeed, but it converts a negative small\-campaign result into sample\-size guidance\.

![Refer to caption](https://arxiv.org/html/2607.27651v1/x4.png)Figure 4:A strict\-majority slice of the general finite\-population certifiability boundary\.Separately implemented enumeration verified exact coverage over all finite\-support campaign–pilot pairs\. In 7,442 sequential campaigns the rule never activated, and fallback lost exactly the information bill already spent \(Fig\.[3](https://arxiv.org/html/2607.27651#S2.F3)b,c; Extended Data Figs\.[3](https://arxiv.org/html/2607.27651#S5.F3)and[4](https://arxiv.org/html/2607.27651#S5.F4)\)\. Yet full\-state and pilot\-observable opportunity remained positive in the actionable fixtures \(Fig\.[3](https://arxiv.org/html/2607.27651#S2.F3)d\)\. The mismatch isolates non\-certifiability: a mechanism can have adaptive value in expectation while a small finite campaign cannot support the required claim about its specific unobserved complement\.

### 2\.5Held\-out calibration controls simulator activation risk

![Refer to caption](https://arxiv.org/html/2607.27651v1/x5.png)Figure 5:Prospectively held\-out calibration controls simulator false activation while retaining measurable sensitivity\.We next asked whether a population\-level gate could recover an active branch while controlling erroneous activation\. A NULL\-only tuning set selectedτ=0\.0180806\\tau=0\.0180806; the threshold was then frozen and passed an independent risk lock without a second attempt \(Extended Data Fig\.[5](https://arxiv.org/html/2607.27651#S5.F5)a,b\)\.

On 900 previously unseen mechanism fixtures, the gate activated none of 400 NULL fixtures\. Its one\-sided 99\.5% false\-activation upper bound was 1\.316%, well below the registered 5% limit\. It also activated 342 of 500 positive fixtures, giving sensitivity 68\.4% \(95% confidence interval, 64\.1–72\.5%; Fig\.[5](https://arxiv.org/html/2607.27651#S2.F5)a–c\)\. Risk control therefore did not require an always\-fallback policy in this population\.

The gate did not increase mean value detectably over forced activation, but it changed the failure profile\. Forced activation was harmful in 13\.6% of positive fixtures; gating removed that negative tail while declining some small positive effects\. The paired mean difference was8\.60×10−58\.60\\times 10^\{\-5\}\(retrospective 95% interval,−6\.78×10−4\-6\.78\\times 10^\{\-4\}to8\.75×10−48\.75\\times 10^\{\-4\}; Fig\.[5](https://arxiv.org/html/2607.27651#S2.F5)d\)\. Post\-result threshold ablations attained higher sensitivity, but the most permissive rule exceeded the 5% risk limit \(Extended Data Fig\.[6](https://arxiv.org/html/2607.27651#S5.F6)\)\. These comparisons explain a risk–sensitivity trade\-off; only the originally frozen rule carries the binding simulator claim\.

This layer is intentionally diagnostic\. Its campaign scoreSevalS^\{\\mathrm\{eval\}\}uses common\-random\-number potential outcomes and is not available at decision time\. The result shows that the authorization logic can support non\-trivial activation under known simulator truth; it does not validate a deployable physical\-instrument score\. The separate Causal Chambers analysis establishes constructibility of an archive\-backed proxy only \(Extended Data Fig\.[8](https://arxiv.org/html/2607.27651#S5.F8)\)\.

### 2\.6Measured opportunity exceeds the locked policy score

The locked CTRP study asked a different question: can a source\-trained, observable\-only score recognize opportunity that is visible after complete outcome measurement? CTRP v2 provides cancer\-cell\-line responses to small\-molecule perturbations\[seashoreludlow2015\]\. Before response access, we fixed a four\-compound panel and actionsK∈\{1,2,4\}K\\in\\\{1,2,4\\\}, whereKKdenotes the size of an offline response bundle ranked from development data\. Every retained family had outcomes for every compound, so each action value was an exact measured lookup rather than an imputed counterfactual\. Families were hash\-partitioned into development, calibration and locked evaluation \(Supplementary SectionLABEL:SI\-sec:sr\_ctrp\)\.

![Refer to caption](https://arxiv.org/html/2607.27651v1/x6.png)Figure 6:Measured\-outcome studies separate source\-policy misalignment from target\-calibrated authorization\.a, CTRP calibration rejectedτ=0\\tau=0because its one\-sided 99\.5% NULL false\-activation upper bound exceeded the 5% limit, and selectedτ=0\.5\\tau=0\.5\.b, All 254 locked CTRP scores remained belowτ=0\.5\\tau=0\.5, including the seven outcome\-defined POSITIVE families\.c, In the cpg0012 final test, circles are point estimates, triangles are one\-sided 95% confidence bounds and short black ticks are the internally frozen requirements\. The false\-activation UCB was 5\.18%, below its 7\.5% limit\. The FDP point estimate was34\.62%<35%34\.62\\%<35\\%, whereas its 95% UCB was37\.97%\>35%37\.97\\%\>35\\%; the sensitivity bound passed, the fixed\-archive worst\-case mean gain was1\.948×10−31\.948\\times 10^\{\-3\}, and its compound\-bootstrap BCa lower endpoint was1\.229×10−3\>01\.229\\times 10^\{\-3\}\>0\.d, Comparator methods evaluated on identical final compounds, decision\-time observables, action and cost\. Each point is one complete method; the dashed line is the 7\.5% false\-activation UCB limit\.Opalwas the only nonzero\-activation method below that limit, but no method met the complete joint contract\.Risk calibration materially changed the policy\. The unconstrained threshold failed the 5% NULL\-risk limit, whereasτ=0\.5\\tau=0\.5passed calibration and was frozen before locked responses were opened \(Fig\.[6](https://arxiv.org/html/2607.27651#S2.F6)a\)\. Every locked score subsequently fell below that threshold\. The final rule therefore made no false activations among 244 NULL families, but it also activated none of the seven POSITIVE families and had zero executed gain\.

The absence of activation did not mean that the panel lacked opportunity\. POSITIVE families had mean panel\-relative measured opportunity 0\.2054, while the score–opportunity rank correlation was only9×10−39\\times 10^\{\-3\}\(Fig\.[6](https://arxiv.org/html/2607.27651#S2.F6)b; Extended Data Fig\.[7](https://arxiv.org/html/2607.27651#S5.F7)b–d\)\. CTRP therefore identifies*observable\-policy misalignment*: this frozen source\-trained scorer did not rank measured opportunity in the locked target strata\. It is neither a negative drug\-efficacy result nor proof that every representation of the decision\-time observables must fail\.

The counts0/2440/244,0/70/7and zero paired gain are exact properties of the fixed archive\. Repeated\-sampling confidence bounds are secondary summaries under independent\-family exchangeability and do not imply transport to future cell lines\. The action also remains an offline compound\-bundle abstraction, not a native laboratory control\.

## 3Discussion

Opalmakes adaptive opportunity, target identifiability, finite\-sample certifiability and executed value distinct properties\. Passing any one does not imply the others, so authorization cannot be reduced to predictive accuracy or a risk threshold\. The framework changes the order of evaluation: first establish that adaptation can be identified, certified, exercised non\-trivially and valued after cost; only then optimize the active policy\.

The Cell Painting study demonstrates why this ordering matters\. In cpg0012, target\-labelled development produced selective activation with controlled false activation and a positive executed\-value lower bound\. Under identical observables, actions and costs, no other activating method in the comparison met the false\-activation requirement\. YetOpaldid not award its own certificate because the false\-discovery confidence bound exceeded the frozen limit, although the point estimate did not\. The value of the contract is therefore visible both in the active branch it permits and in the favourable result it refuses to overinterpret\.

Together, CTRP and cpg0012 illustrate the two sides of the target\-identification boundary: source\-trained transfer in CTRP abstained despite measured opportunity, whereas labelled target development in cpg0012 recovered a non\-trivial gate before a single held\-out final evaluation\. The remaining studies locate different breaks in the chain\. The finite\-population result shows that opportunity can remain uncertifiable at a given campaign size, and the simulator confirms that false\-activation control can trade sensitivity for protection against harmful activation\. Aligning the executed comparator and charging probe cost further show why positive surrogates need not survive deployment\. These diagnoses prescribe different responses—change the representation, acquire target labels, enlarge the campaign or retainK0K\_\{0\}—rather than one generic model update\.

Together, the studies establish an auditable authorization method and general certifiability boundaries in offline settings\. CTRP uses author\-defined compound\-bundle actions, and cpg0012 uses an adaptive\-replicate abstraction with normalized cost and an author\-controlled final separation\. A prospective physical campaign with native actions, measured operational costs and an independently sealed final partition is the next test of deployment readiness\.

## 4Methods

### 4\.1Decision setting and information value

Consider decision unitsiinested within campaignscc\. A campaign contains a finite set of units that share one frozen policy and commitment contract; in the simulations these units are called banks, whereas in the external archive they are cell\-line families\. Unitiihas a payoff\-relevant latent stateMiM\_\{i\}, observable precommitment informationXiX\_\{i\}, an optional probe resultZi​rZ\_\{ir\}, and a finite capacity actionK∈𝒦K\\in\\mathcal\{K\}\. The stateMiM\_\{i\}contains the mechanism variables needed to determine conditional expected utility; it need not contain independent survey or probe noise\. Capacity levels are ordered units of experimental capability, and their physical realization is application specific\. Utility

Ui​\(K\)=Yi​\(K\)−Ciop​\(K\)U\_\{i\}\(K\)=Y\_\{i\}\(K\)\-C\_\{i\}^\{\\mathrm\{op\}\}\(K\)combines a normalized scientific\-success valueYi​\(K\)Y\_\{i\}\(K\)with action\-dependent operational cost on one registered scale\. Two baseline identities are kept distinct\.Kscore,−fK\_\{\\mathrm\{score\},\-f\}is the fold\-specific reference used to construct a cross\-fitted selector score;K0K\_\{0\}is the best executed fixed policy used as the operational comparator and, when verified, the fallback of the cohort\-level safe wrapper\. The retrospective selector\-frontier audit hadKscore,−f=1K\_\{\\mathrm\{score\},\-f\}=1andK0=2K\_\{0\}=2; the finite\-campaign certification validation used one frozenK0=2K\_\{0\}=2throughout\.

The decision timeline is:

train→observe​X→optionally acquire​Zr→reserve​K→execute→settle costs\.\\text\{train\}\\rightarrow\\text\{observe \}X\\rightarrow\\text\{optionally acquire \}Z\_\{r\}\\rightarrow\\text\{reserve \}K\\rightarrow\\text\{execute\}\\rightarrow\\text\{settle costs\}\.Every policy is constrained to the information available at its decision time\. Oracle mechanism variables and evaluation outcomes are prohibited from deployable features\. The same scoring function, pilot size and information set must be used during calibration and deployment; we call this the*deployment\-isomorphism invariant*\(Supplementary SectionLABEL:SI\-sec:deployment\_isomorphism\)\. A score that uses potential outcomes is therefore an evaluation score for a simulator, not a deployable diagnostic\.

Unless stated otherwise, “prespecified”, “frozen” and “locked” denote internally timestamped, SHA\-256\-bound artefacts created before the relevant outcome stage; they do not imply registration in an external public registry\.

#### Information opportunity\.

For an information setℐ\\mathcal\{I\}, define

V​\(ℐ\)=𝔼​\[maxk∈𝒦⁡𝔼​\{U​\(k\)∣ℐ\}\],V0=maxk∈𝒦⁡𝔼​\{U​\(k\)\}\.V\(\\mathcal\{I\}\)=\\mathbb\{E\}\\\!\\left\[\\max\_\{k\\in\\mathcal\{K\}\}\\mathbb\{E\}\\\{U\(k\)\\mid\\mathcal\{I\}\\\}\\right\],\\qquad V\_\{0\}=\\max\_\{k\\in\\mathcal\{K\}\}\\mathbb\{E\}\\\{U\(k\)\\\}\.\(1\)The routing opportunity available fromℐ\\mathcal\{I\}is

G​\(ℐ\)=V​\(ℐ\)−V0\.G\(\\mathcal\{I\}\)=V\(\\mathcal\{I\}\)\-V\_\{0\}\.\(2\)We use

GX=G​\{σ​\(X\)\},GX​Z=G​\{σ​\(X,Zr\)\},GM=G​\{σ​\(M\)\},G\_\{X\}=G\\\{\\sigma\(X\)\\\},\\qquad G\_\{XZ\}=G\\\{\\sigma\(X,Z\_\{r\}\)\\\},\\qquad G\_\{M\}=G\\\{\\sigma\(M\)\\\},whereGMG\_\{M\}is simulator truth and is never a deployable feature\. We do not assumeσ​\(X\)⊆σ​\(M\)\\sigma\(X\)\\subseteq\\sigma\(M\): noisy measurements need not be sub\-σ\\sigma\-fields of the latent state\. Instead, the information channels satisfy the payoff\-sufficiency condition

𝔼​\{U​\(k\)∣M,X,Zr\}=𝔼​\{U​\(k\)∣M\},k∈𝒦\.\\mathbb\{E\}\\\{U\(k\)\\mid M,X,Z\_\{r\}\\\}=\\mathbb\{E\}\\\{U\(k\)\\mid M\\\},\\qquad k\\in\\mathcal\{K\}\.\(3\)ThusXXandZrZ\_\{r\}may contain independent noise, but cannot reveal payoff information that is absent from the payoff\-relevant state\. This is the relevant Blackwell\-garbling condition\[blackwell1953\]\.

###### Proposition 1\(Opportunity monotonicity and existence\)\.

For a finite action set, common population and common utility scale, under Eq\. \([3](https://arxiv.org/html/2607.27651#S4.E3)\),

0≤GX≤GX​Z≤GM\.0\\leq G\_\{X\}\\leq G\_\{XZ\}\\leq G\_\{M\}\.\(4\)Moreover,GM=0G\_\{M\}=0if and only if at least one fixed action maximizes𝔼​\{U​\(k\)∣M\}\\mathbb\{E\}\\\{U\(k\)\\mid M\\\}almost surely\. Consequently, whenGM=0G\_\{M\}=0, no router using a deployable information channel whose payoff information is Blackwell\-dominated byMMcan improve on the best fixed action\.

Proof sketch\.RefiningXXto\(X,Zr\)\(X,Z\_\{r\}\)cannot reduce the optimal conditional value\. Payoff sufficiency makes the latter information channel a garbling ofMMfor every action, which gives the upper inequality\. Non\-negativity follows because fixed actions remain feasible, and equality at zero requires a common statewise\-optimal action\. Supplementary PropositionLABEL:SI\-prop:opportunitygives the complete measurability assumptions, equality conditions and proof\.

The opportunity hierarchy is an upper\-bound diagnostic rather than a claim that an oracle is deployable\. Truth strata are defined fromGMG\_\{M\}and deployable observabilityη=GX​Z/GM\\eta=G\_\{XZ\}/G\_\{M\}: NULL opportunity, weak/ambiguous opportunity and positive actionable opportunity\. Only NULL and positive strata carry operating\-characteristic targets; the ambiguous stratum is descriptive\.

#### Delayed commitment\.

Delayed commitment has the logic of a real option: evidence is valuable only insofar as waiting can change an action before an irreversible or costly commitment\[dixitpindyck1994\]\. Here that idea is expressed directly on the registered scientific\-utility scale\.

For a candidate probe typerr, the gross conditional option value is

Ωrgross​\(X\)=𝔼​\[maxk⁡𝔼​\{U​\(k\)∣X,Zr\}∣X\]−maxk⁡𝔼​\{U​\(k\)∣X\}\.\\Omega\_\{r\}^\{\\mathrm\{gross\}\}\(X\)=\\mathbb\{E\}\\\!\\left\[\\max\_\{k\}\\mathbb\{E\}\\\{U\(k\)\\mid X,Z\_\{r\}\\\}\\mid X\\right\]\-\\max\_\{k\}\\mathbb\{E\}\\\{U\(k\)\\mid X\\\}\.\(5\)Resource and latency chargescrc\_\{r\}andcrlatc\_\{r\}^\{\\mathrm\{lat\}\}produce

Ω⋆​\(X\)=max⁡\[0,maxr⁡\{Ωrgross​\(X\)−cr−crlat\}\]\.\\Omega^\{\\star\}\(X\)=\\max\\\!\\left\[0,\\ \\max\_\{r\}\\\{\\Omega\_\{r\}^\{\\mathrm\{gross\}\}\(X\)\-c\_\{r\}\-c\_\{r\}^\{\\mathrm\{lat\}\}\\\}\\right\]\.\(6\)The explicit zero is the no\-probe action\. Individual net probe values may be negative\.

###### Proposition 2\(Value and disappearance of delayed commitment\)\.

Ωrgross​\(X\)≥0\\Omega\_\{r\}^\{\\mathrm\{gross\}\}\(X\)\\geq 0\. Conditional onXX, equality holds if and only if there is an action that is optimal for almost every possibleZrZ\_\{r\}, up to ties\. Thus a probe that improves prediction but never changes an optimal action has zero gross decision value; it has negative net value whenever it has positive cost\.

Proof sketch\.Conditional Jensen applied to the pointwise maximum gives non\-negativity\. Equality requires the conditional support to lie in a common face of that maximum, which for finitely many actions is equivalent to a common optimizer across probe outcomes, allowing ties\. Supplementary PropositionLABEL:SI\-prop:optiongives the full conditional statement, tie conditions and proof\.

The deployed probe gate purchasesqqonly when a cross\-fitted action\-change model and a net\-gain lower bound both pass\. Development uses randomized forced exploration with known propensities; policy evaluation uses disjoint campaigns\. Candidate\-level resource, latency, gross value, net value and action changes are retained even when no probe is bought\.

### 4\.2Adaptive policy and safe fallback

One cross\-fitted value model producesV^i​\(k\)\\widehat\{V\}\_\{i\}\(k\)and paired uncertaintys^i​\(k,Kscore,−f\)\\widehat\{s\}\_\{i\}\(k,K\_\{\\mathrm\{score\},\-f\}\)\. A frozen family\{πλ:λ∈Λ\}\\\{\\pi\_\{\\lambda\}:\\lambda\\in\\Lambda\\\}changes only the evidence threshold:

πλ​\(Xi\)=\{arg⁡maxk⁡V^i​\(k\),λ=0,arg⁡maxk:V^i​\(k\)−V^i​\(Kscore,−f\)\>λ​s^i​\(k,Kscore,−f\)⁡V^i​\(k\),λ\>0,Kscore,−f,if no deviation qualifies\.\\pi\_\{\\lambda\}\(X\_\{i\}\)=\\begin\{cases\}\\displaystyle\\arg\\max\_\{k\}\\widehat\{V\}\_\{i\}\(k\),&\\lambda=0,\\\\\[3\.0pt\] \\displaystyle\\arg\\max\_\{k:\\,\\widehat\{V\}\_\{i\}\(k\)\-\\widehat\{V\}\_\{i\}\(K\_\{\\mathrm\{score\},\-f\}\)\>\\lambda\\widehat\{s\}\_\{i\}\(k,K\_\{\\mathrm\{score\},\-f\}\)\}\\widehat\{V\}\_\{i\}\(k\),&\\lambda\>0,\\\\ K\_\{\\mathrm\{score\},\-f\},&\\text\{if no deviation qualifies\}\.\\end\{cases\}\(7\)Non\-finite or degenerate standard errors forceKscore,−fK\_\{\\mathrm\{score\},\-f\}\. Ties preferKscore,−fK\_\{\\mathrm\{score\},\-f\}, then the smaller capacity\. The reference is recomputed inside each outer training fold\. The completeλ\\lambdagrid is frozen before any frontier outcome is examined\. This bank\-level score fallback is not a safety guarantee against a different executed comparator\.

Let

Δr​\(λ\)=Vr​\(πλ\)−Vr​\(K0\)\\Delta\_\{r\}\(\\lambda\)=V\_\{r\}\(\\pi\_\{\\lambda\}\)\-V\_\{r\}\(K\_\{0\}\)in truth regimerr, after charging policy\-specific operational cost\. Because everyπλ\\pi\_\{\\lambda\}uses only deployable information,

Vr​\(k\)=𝔼r​\{U​\(k\)\},V0,r=maxk∈𝒦⁡Vr​\(k\),br=V0,r−Vr​\(K0\)≥0,V\_\{r\}\(k\)=\\mathbb\{E\}\_\{r\}\\\{U\(k\)\\\},\\qquad V\_\{0,r\}=\\max\_\{k\\in\\mathcal\{K\}\}V\_\{r\}\(k\),\\qquad b\_\{r\}=V\_\{0,r\}\-V\_\{r\}\(K\_\{0\}\)\\geq 0,and therefore

supλ∈ΛΔr​\(λ\)≤GX​Z​\(r\)\+br≤GM​\(r\)\+br\.\\sup\_\{\\lambda\\in\\Lambda\}\\Delta\_\{r\}\(\\lambda\)\\leq G\_\{XZ\}\(r\)\+b\_\{r\}\\leq G\_\{M\}\(r\)\+b\_\{r\}\.\(8\)When the executed comparator is the population\-optimal fixed action on the same population and utility scale,Vr​\(K0\)=V0,rV\_\{r\}\(K\_\{0\}\)=V\_\{0,r\}, sobr=0b\_\{r\}=0and Eq\. \([8](https://arxiv.org/html/2607.27651#S4.E8)\) reduces to the information\-opportunity ceiling\. Otherwisebrb\_\{r\}records the fixed\-baseline gap and cannot be attributed to routing information\. The first inequality can be strict because of estimation error, restricted policy classes and information cost\. The frontier reports value, switch rate, correct and incorrect switches, no\-difference switches and per\-switch harm at everyλ\\lambda; no empirically favourableλ\\lambdais selected\.

#### Global model and residual experts\.

The primary value model is global\. A secondary layer, related to classical adaptive and hierarchical mixtures of experts\[jacobs1991,jordan1994\], adds softly gated residual experts to the global prediction rather than replacing it:

V^i​\(k\)=V^global,i​\(k\)\+𝟏​\{E\>0\}​∑e=1Ewi​e​r^e,i​\(k\),∑e=1Ewi​e=1​when​E\>0,wi​e≥0\.\\widehat\{V\}\_\{i\}\(k\)=\\widehat\{V\}\_\{\\mathrm\{global\},i\}\(k\)\+\\mathbf\{1\}\\\{E\>0\\\}\\sum\_\{e=1\}^\{E\}w\_\{ie\}\\widehat\{r\}\_\{e,i\}\(k\),\\qquad\\sum\_\{e=1\}^\{E\}w\_\{ie\}=1\\ \\text\{when \}E\>0,\\quad w\_\{ie\}\\geq 0\.\(9\)Expert count, temperature and shrinkage are selected inside each outer training fold by nested cross\-validation and a one\-standard\-error rule\. Unsupported or out\-of\-distribution beliefs fall back to the global model\.E=0E=0is an intended model\-selection outcome, not a failed fit\. This architecture tests whether stable local residual structure exists without forcing a high\-variance router to create heterogeneity\.

#### Cohort\-level safety\.

The cohort wrapper is related to high\-confidence policy improvement and calibrated risk control\[thomas2015,angelopoulos2025\], but its guarantee is stated for the frozen finite policy family and the executed comparator used here\.

LetΠ=\{πλ:λ∈Λ\}∪\{K0\}\\Pi=\\\{\\pi\_\{\\lambda\}:\\lambda\\in\\Lambda\\\}\\cup\\\{K\_\{0\}\\\}, and letL​\(π\)L\(\\pi\)be a simultaneous lower confidence bound forΔ​\(π\)=V​\(π\)−V​\(K0\)\\Delta\(\\pi\)=V\(\\pi\)\-V\(K\_\{0\}\)over this finite class\. The safe selector chooses a preregistered member withL​\(π\)≥−ϵL\(\\pi\)\\geq\-\\epsilon; if none exists it returnsK0K\_\{0\}\.

###### Proposition 3\(Safe data\-dependent selection\)\.

Forϵ≥0\\epsilon\\geq 0, if

Pr⁡\{Δ​\(π\)≥L​\(π\)​for all​π∈Π\}≥1−α,\\Pr\\\{\\Delta\(\\pi\)\\geq L\(\\pi\)\\ \\text\{for all \}\\pi\\in\\Pi\\\}\\geq 1\-\\alpha,then the selected policyπ^safe\\widehat\{\\pi\}\_\{\\mathrm\{safe\}\}satisfies

Pr⁡\{V​\(π^safe\)≥V​\(K0\)−ϵ\}≥1−α\.\\Pr\\\{V\(\\widehat\{\\pi\}\_\{\\mathrm\{safe\}\}\)\\geq V\(K\_\{0\}\)\-\\epsilon\\\}\\geq 1\-\\alpha\.\(10\)

Proof sketch\.On the simultaneous\-coverage event, every eligible policy hasΔ​\(π\)≥L​\(π\)≥−ϵ\\Delta\(\\pi\)\\geq L\(\\pi\)\\geq\-\\epsilon; an empty eligible set returnsK0K\_\{0\}\. The guarantee therefore holds throughout an event of probability at least1−α1\-\\alpha\. Supplementary PropositionLABEL:SI\-prop:safegives the full proof and the registered decompositionϵ=rmax​δharm\+Cdecision\+Cprobe\\epsilon=r\_\{\\max\}\\delta\_\{\\mathrm\{harm\}\}\+C\_\{\\mathrm\{decision\}\}\+C\_\{\\mathrm\{probe\}\}used here\.

#### Joint authorization and nontriviality\.

A risk bound alone admits the degenerate always\-K0K\_\{0\}rule\. We therefore separate safety from evidence that the active branch is both exercised and valuable\. Conditional on development and calibration, letURU\_\{R\}be an upper confidence bound for NULL false\-activation riskR0R\_\{0\}, letLμL\_\{\\mu\}be a lower confidence bound for executed mean valueμ=𝔼​\{U​\(π​\(X\)\)−U​\(K0\)\}\\mu=\\mathbb\{E\}\\\{U\(\\pi\(X\)\)\-U\(K\_\{0\}\)\\\}, and letLTL\_\{T\}lower\-bound one preregistered nontriviality endpointTT, such as activation coverage or POSITIVE sensitivity\. The endpoint identity and a minimumtmin\>0t\_\{\\min\}\>0must be fixed before evaluation; choosing among endpoints after opening the lock is not permitted\.

###### Proposition 4\(Conjunctive risk–value–nontriviality authorization\)\.

LetCjC\_\{j\},j=1,…,mj=1,\\ldots,m, denote prespecified component claims and letPjP\_\{j\}be a level\-αj\\alpha\_\{j\}pass event for the complementary nullCjcC\_\{j\}^\{c\}, conditional on all development data\. If authorization occurs only when everyPjP\_\{j\}occurs, then

supθ∈∪jCjcPrθ⁡\(authorize\)≤maxj⁡αj\.\\sup\_\{\\theta\\in\\cup\_\{j\}C\_\{j\}^\{c\}\}\\Pr\_\{\\theta\}\(\\mathrm\{authorize\}\)\\leq\\max\_\{j\}\\alpha\_\{j\}\.No independence among the component statistics is required\. By contrast, the probability that every reported marginal confidence bound covers simultaneously is at least1−∑jαj1\-\\sum\_\{j\}\\alpha\_\{j\}by the union bound\.

Proof sketch\.If the conjunction is false, at least one component nullCjcC\_\{j\}^\{c\}is true, and authorization is a subset of that component’s pass event\. Its probability is therefore at mostαj\\alpha\_\{j\}\. Supplementary PropositionLABEL:SI\-prop:joint\_authorizationgives the complete statement\. For the basic contract the pass events areUR≤r0U\_\{R\}\\leq r\_\{0\},Lμ\>0L\_\{\\mu\}\>0andLT≥tminL\_\{T\}\\geq t\_\{\\min\}\. Always\-K0K\_\{0\}has zero paired gain and active coverage, so it cannot pass\.

#### Source\-only target\-outcome identification boundary\.

###### Proposition 5\(Source\-only target authorization boundary\)\.

Letℱpre\\mathcal\{F\}\_\{\\mathrm\{pre\}\}contain the source outcomes, unlabelled target covariates and frozen gate available before target outcomes are observed\. Without a restriction on the target conditional outcome law, every favourable target extension supporting a non\-trivial joint claim has an adverse extension with the same distribution onℱpre\\mathcal\{F\}\_\{\\mathrm\{pre\}\}, and hence the same assignments, for which that claim is false\. Anyℱpre\\mathcal\{F\}\_\{\\mathrm\{pre\}\}\-measurable certificate that controls false authorization at levelα\\alphauniformly over adverse extensions therefore authorizes with probability at mostα\\alphaunder an observationally indistinguishable favourable extension\.

Proof sketch\.Hold the target covariate law and gate fixed, and alter only outcomes on the activated region to violate risk, value or non\-triviality\. The two laws are indistinguishable before outcomes are opened\. The complete construction is Supplementary PropositionLABEL:SI\-prop:target\_shift\. This is a uniform identification boundary, not a claim that every source\-only policy fails under every target law\.

The corresponding target\-calibrated specialization uses labelled target development to select a pointwise rule, a pre\-outcome seal of nonzero final assignments, and one untouched final outcome release \(Supplementary SectionLABEL:SI\-sec:target\_calibrated\_recovery\)\. The ideal three\-way design retains a separate certification sample; the realized cpg0012 study used its first two labelled partitions as a 13,748\-compound model\-and\-threshold development population, then froze one rule before the still\-unopened final partition\. It therefore supplies a single\-held\-out final test; the labelled development population is not presented as a separate confirmation\.

### 4\.3Finite\-campaign and risk certification

For a finite campaign ofNNbanks, a pilotPqP\_\{q\}of sizeqqis sampled without replacement and the policy acts onDq=\{1,…,N\}∖PqD\_\{q\}=\\\{1,\\ldots,N\\\}\\setminus P\_\{q\}\. The primary estimand is

Δcamp​\(q,Iq\)=1N​\[∑i∈Dq\{Ui​\(πIq\)−Ui​\(K0\)\}−Cinfo,total​\(q\)\]\.\\Delta\_\{\\mathrm\{camp\}\}\(q,I\_\{q\}\)=\\frac\{1\}\{N\}\\left\[\\sum\_\{i\\in D\_\{q\}\}\\\{U\_\{i\}\(\\pi\_\{I\_\{q\}\}\)\-U\_\{i\}\(K\_\{0\}\)\\\}\-C\_\{\\mathrm\{info,total\}\}\(q\)\\right\]\.\(11\)Define the pre\-information\-cost deployment\-complement gain

GA,D​\(q,Iq\)=1N−q​∑i∈Dq\{Ui​\(πIq\)−Ui​\(K0\)\}\.G\_\{A,D\}\(q,I\_\{q\}\)=\\frac\{1\}\{N\-q\}\\sum\_\{i\\in D\_\{q\}\}\\\{U\_\{i\}\(\\pi\_\{I\_\{q\}\}\)\-U\_\{i\}\(K\_\{0\}\)\\\}\.ThenΔcamp=\{\(N−q\)​GA,D−Cinfo,total\}/N\\Delta\_\{\\mathrm\{camp\}\}=\\\{\(N\-q\)G\_\{A,D\}\-C\_\{\\mathrm\{info,total\}\}\\\}/N\. The denominator remains the original campaign size\. Pilot observations, future deployment\-complement value and a new\-bank superpopulation mean are different estimands; they are not interchangeable\.

For the finite\-campaign boundary study, latent capacity classes are categorical, pilot signals follow a registered misclassification channel, and the finite composition is inferred by exact multivariate\-hypergeometric likelihood inversion\. Let𝒞q\\mathcal\{C\}\_\{q\}be a simultaneous confidence set for composition and any observation\-channel nuisance\. The lower bound is

Δ¯q=infϑ∈𝒞qΔcamp​\(q,Iq;ϑ\)\.\\underline\{\\Delta\}\_\{q\}=\\inf\_\{\\vartheta\\in\\mathcal\{C\}\_\{q\}\}\\Delta\_\{\\mathrm\{camp\}\}\(q,I\_\{q\};\\vartheta\)\.\(12\)The policy changes the deployment\-complement action only ifΔ¯q\>0\\underline\{\\Delta\}\_\{q\}\>0\. Otherwise it retainsK0K\_\{0\}\. Full construction details are in Supplementary SectionLABEL:SI\-sec:finite\_algorithm\.

###### Proposition 6\(Costed finite\-campaign safety\)\.

If the confidence sets cover the finite\-campaign state simultaneously over all registered looks with probability at least1−α1\-\\alpha, then every activated switch has non\-negativeΔcamp\\Delta\_\{\\mathrm\{camp\}\}with probability at least1−α1\-\\alpha\. If the procedure abstains after observingqqpilot banks, its incremental loss relative to immediate fixed deployment is exactly the registered information bill divided byNN, provided the pilot banks themselves remain onK0K\_\{0\}\.

Proof sketch\.Simultaneous coverage places the true finite state in every confidence set, so the infimum in Eq\. \([12](https://arxiv.org/html/2607.27651#S4.E12)\) lower\-bounds the costed contrast even at a data\-dependent look\. Under fallback, the two policies take the same actions and their utilities cancel, leaving only the information bill\. Supplementary PropositionLABEL:SI\-prop:finite\_safesupplies the full sequential formulation and proof\.

#### Exact qualification and sample\-complexity boundary\.

For binary actionable status, letXXbe the actionable count in a simple random pilot of sizeq<Nq<N\. Writed=N−qd=N\-q, and let the registered claim be that the actionable fraction in the specific unobserved complement exceedsρ∈\[0,1\)\\rho\\in\[0,1\)\. The most favourable observation isX=qX=q\. The boundary\-null composition then containsB0=q\+⌊ρ​d⌋B\_\{0\}=q\+\\lfloor\\rho d\\rflooractionable units, for which

h​\(N,q,ρ\)=PrB0⁡\(X=q\)=\(q\+⌊ρ​\(N−q\)⌋q\)\(Nq\)\.h\(N,q,\\rho\)=\\Pr\_\{B\_\{0\}\}\(X=q\)=\\frac\{\\binom\{q\+\\lfloor\\rho\(N\-q\)\\rfloor\}\{q\}\}\{\\binom\{N\}\{q\}\}\.\(13\)
###### Proposition 7\(Exact best\-case finite\-population boundary\)\.

A non\-randomized monotone level\-α\\alphacertificate, where activation probability is non\-decreasing in the observed actionable count, based on the binary pilot count is nontrivial at lookqqif and only if

h​\(N,q,ρ\)≤α\.h\(N,q,\\rho\)\\leq\\alpha\.\(14\)Ifh​\(N,q,ρ\)\>αh\(N,q,\\rho\)\>\\alpha, even an all\-actionable pilot cannot be certified\. Any randomized level\-α\\alpharule can certify on that event with probability at mostα/h​\(N,q,ρ\)<1\\alpha/h\(N,q,\\rho\)<1\. Accordingly,

q†​\(N,ρ,α\)=min⁡\{q:h​\(N,q,ρ\)≤α\},q^\{\\dagger\}\(N,\\rho,\\alpha\)=\\min\\\{q:h\(N,q,\\rho\)\\leq\\alpha\\\},withq†q^\{\\dagger\}undefined when the registered look set contains no qualifyingqq, is the exact best\-case sample\-complexity boundary\.

Proof sketch\.Under the boundary\-null composition, the all\-actionable pilot is a false certificate with probabilityhh, which proves necessity and the randomization bound\. Conversely, the monotone rule that certifies only atX=qX=qhas maximal false\-certification probabilityhh, proving sufficiency\. Supplementary PropositionLABEL:SI\-prop:impossibilitygives the full proof, the finite\-NNsaw\-tooth qualification surface and the limith​\(N,q,ρ\)→ρqh\(N,q,\\rho\)\\to\\rho^\{q\}\. For the registered strict\-majority targetρ=1/2\\rho=1/2,N=16N=16andα=0\.02/7\\alpha=0\.02/7, no registered look qualifies; the first non\-excluded finite design isN=25,q=14N=25,q=14, whereas the large\-NNbest\-case limit is⌈log⁡\(α\)/log⁡\(ρ\)⌉=9\\lceil\\log\(\\alpha\)/\\log\(\\rho\)\\rceil=9\. This binary noiseless boundary is necessary design guidance; it is not a feasibility proof for the complete noisy three\-class rule\.

###### Proposition 8\(Full\-likelihood non\-certifiability for the registered finite\-campaign rule\)\.

Let𝐒=\(S0,S1,S2\)\\mathbf\{S\}=\(S\_\{0\},S\_\{1\},S\_\{2\}\)be the observed pilot signal counts and let𝒞q,p​\(𝐒\)\\mathcal\{C\}\_\{q,p\}\(\\mathbf\{S\}\)be the exact registered confidence set of campaign/pilot composition pairs\(𝐌,𝐇\)\(\\mathbf\{M\},\\mathbf\{H\}\)under identity\-mixture channel parameterpp, where

Pr⁡\(S=s∣k⋆=j,p\)=p​1​\(s=j\)\+\(1−p\)/3\.\\Pr\(S=s\\mid k^\{\\star\}=j,p\)=p\\,1\(s=j\)\+\(1\-p\)/3\.For the registeredN=16N=16,q∈\{2,4,…,14\}q\\in\\\{2,4,\\ldots,14\\\},p∈\{0\.70,0\.80,0\.92\}p\\in\\\{0\.70,0\.80,0\.92\\\}, prior support, value table, information bill and decision rule, define the least\-favourable pair

\(𝐌†,𝐇†\)∈arg⁡min\(𝐌,𝐇\)∈𝒞q,p​\(𝐒\)⁡\{V𝐌−𝐇​\(aq,p​\(𝐒\)\)−V𝐌−𝐇​\(K0\)\}\.\(\\mathbf\{M\}^\{\\dagger\},\\mathbf\{H\}^\{\\dagger\}\)\\in\\arg\\min\_\{\(\\mathbf\{M\},\\mathbf\{H\}\)\\in\\mathcal\{C\}\_\{q,p\}\(\\mathbf\{S\}\)\}\\\{V\_\{\\mathbf\{M\}\-\\mathbf\{H\}\}\(a\_\{q,p\}\(\\mathbf\{S\}\)\)\-V\_\{\\mathbf\{M\}\-\\mathbf\{H\}\}\(K\_\{0\}\)\\\}\.For every possible\(q,p,𝐒\)\(q,p,\\mathbf\{S\}\), this minimum is non\-positive\. Hence the registered actionable\-state condition fails and the costed lower bound is strictly negative\. The complete rule cannot certify or activate a switch for any possible pilot observation in this finite experiment\.

Proof sketch\.The support is finite\. We enumerate all 21\(q,p\)\(q,p\)cells, every weak composition ofqqinto three signal counts, every retained\(𝐌,𝐇\)\(\\mathbf\{M\},\\mathbf\{H\}\)pair, and the exact Bayes action\. Across 1,113 signal compositions, the maximum least\-favourable gain is0, all state labels areINDETERMINATE, and the largest costed lower bound is−0\.06875\-0\.06875\. A separately implemented validator reconstructs the support, least\-favourable pairs and cost identity\. The complete derivation and certificate are Supplementary PropositionLABEL:SI\-prop:full\_likelihood\_impossibility\. This proposition concerns the frozen finite likelihood and decision contract; it is not a universal impossibility result for every three\-class diagnostic\.

#### Constrained\-risk evaluation\.

When conservative estimation of continuous opportunity is infeasible,Opaltreats activation as a binary risk\-control problem\. For a scoreSS, a frozen rule producesDτ=𝟏​\{S\>τ\}D\_\{\\tau\}=\\mathbf\{1\}\\\{S\>\\tau\\\}\. For the NULL mechanism population,

R0​\(τ\)=Pr⁡\(Dτ=1∣NULL\)R\_\{0\}\(\\tau\)=\\Pr\(D\_\{\\tau\}=1\\mid\\mathrm\{NULL\}\)is the false\-activation risk\. Candidate thresholds are the order statistics of independent NULL tuning scores together with the two abstention endpoints\. The procedure selects the lowest threshold whose one\-sided exact Clopper–Pearson upper bound is at mostr0r\_\{0\}, applies that single threshold to an independent NULL risk\-lock set, and falls back to always\-K0K\_\{0\}if locking fails\. A final independent NULL set evaluates the binding claim; positive\-mechanism campaigns estimate sensitivity with no pass/fail minimum\.

Two information roles must be distinguished\. A deployable diagnostic requiresS=Sdep​\(X,Zpilot\)S=S^\{\\mathrm\{dep\}\}\(X,Z\_\{\\mathrm\{pilot\}\}\)to be computable from observations available at activation time\. The prospective simulation study considered here instead uses an evaluation score

Seval=1N​∑i=1N\{Ui​\(K^i\)−Ui​\(K0\)\},S^\{\\mathrm\{eval\}\}=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\\{U\_\{i\}\(\\widehat\{K\}\_\{i\}\)\-U\_\{i\}\(K\_\{0\}\)\\\},\(15\)whereK^i\\widehat\{K\}\_\{i\}is selected from the observable survey and frozen value model, but the score itself is evaluated using common\-random\-number potential outcomes for all actions\. Equation \([15](https://arxiv.org/html/2607.27651#S4.E15)\) is available only in the simulator\. It measures the operating characteristics of opportunity recognition under known truth; it does not validate a real\-time facility diagnostic\. A physical deployment claim would require a separately registered observable pilot estimator and its own coverage analysis\.

###### Proposition 9\(Held\-out constrained\-risk validation\)\.

Conditional on all tuning and locking data, if the held\-out validation fixtures and campaigns are mutually independent andDτD\_\{\\tau\}is Bernoulli with riskR0​\(τ\)R\_\{0\}\(\\tau\), then the one\-sided Clopper–Pearson upper boundUαU\_\{\\alpha\}satisfies

Pr⁡\{R0​\(τ\)≤Uα\}≥1−α\.\\Pr\\\{R\_\{0\}\(\\tau\)\\leq U\_\{\\alpha\}\\\}\\geq 1\-\\alpha\.Consequently, observingUα≤r0U\_\{\\alpha\}\\leq r\_\{0\}certifies the false\-activation risk at levelr0r\_\{0\}\. If locking fails, always\-K0K\_\{0\}has false\-activation risk zero\.

Proof sketch\.Conditioning on tuning and locking fixesτ\\tau; exact binomial coverage then bounds the held\-out Bernoulli risk\. The always\-K0K\_\{0\}rule has zero false activation by construction\. Supplementary PropositionLABEL:SI\-prop:riskgives the exact construction and proof\. The result applies to either score type when the registered independent\-unit assumptions hold, but it does not establish that the score is observable\. In the simulation study the primary unit is one distinct continuous fixture with one fresh campaign, not multiple campaigns pooled within a fixture\. Tuning, locking, NULL validation and positive validation use mutually disjoint fixture and seed namespaces\.

### 4\.4Value accounting and evidence design

Letϕ∈Φ\\phi\\in\\Phiindex a preregistered commitment contract\. For policy actionaaand observable contextxx, its operational charge is

Cϕ​\(a;x\)=Cϕreserve\+Cϕsetup\+Cϕidle\+Cϕcancel\+Cϕspot\+Cϕdelay\+Cϕinformation\.C\_\{\\phi\}\(a;x\)=C\_\{\\phi\}^\{\\mathrm\{reserve\}\}\+C\_\{\\phi\}^\{\\mathrm\{setup\}\}\+C\_\{\\phi\}^\{\\mathrm\{idle\}\}\+C\_\{\\phi\}^\{\\mathrm\{cancel\}\}\+C\_\{\\phi\}^\{\\mathrm\{spot\}\}\+C\_\{\\phi\}^\{\\mathrm\{delay\}\}\+C\_\{\\phi\}^\{\\mathrm\{information\}\}\.\(16\)The terms are read from an auditable ledger; absent terms are zero rather than absorbed into scientific success\. The operational contrast is

Δop​\(ϕ\)=𝔼​\[Y​\(π\)−Y​\(K0\)−Cϕ​\(π;X\)\+Cϕ​\(K0;X\)\]\.\\Delta\_\{\\mathrm\{op\}\}\(\\phi\)=\\mathbb\{E\}\[Y\(\\pi\)\-Y\(K\_\{0\}\)\-C\_\{\\phi\}\(\\pi;X\)\+C\_\{\\phi\}\(K\_\{0\};X\)\]\.\(17\)Hereϕ\\phiindexes a fully specified commitment\-and\-cost scenario; every reported value uses the complete contract in Eq\. \([16](https://arxiv.org/html/2607.27651#S4.E16)\)\. No unexercised scalar\-friction approximation is used\.

LetCincC\_\{\\mathrm\{inc\}\}be shared incremental training and characterization cost\. If one frozen model servesHHdeployments before retraining,

Δlife​\(ϕ,H\)=Δop​\(ϕ\)−CincH\.\\Delta\_\{\\mathrm\{life\}\}\(\\phi,H\)=\\Delta\_\{\\mathrm\{op\}\}\(\\phi\)\-\\frac\{C\_\{\\mathrm\{inc\}\}\}\{H\}\.\(18\)ForΔop​\(ϕ\)\>0\\Delta\_\{\\mathrm\{op\}\}\(\\phi\)\>0, the adoption boundary is

H⋆​\(ϕ\)=CincΔop​\(ϕ\)\.H^\{\\star\}\(\\phi\)=\\frac\{C\_\{\\mathrm\{inc\}\}\}\{\\Delta\_\{\\mathrm\{op\}\}\(\\phi\)\}\.\(19\)No finite positive break\-even exists when the operational contrast is non\-positive\. When its lower confidence limit is non\-positive, the upper confidence limit forH⋆H^\{\\star\}is right\-censored at infinity\. WithCinc∈\[CL,CU\]C\_\{\\mathrm\{inc\}\}\\in\[C\_\{L\},C\_\{U\}\]and a positive simultaneous bandΔop∈\[L,U\]\\Delta\_\{\\mathrm\{op\}\}\\in\[L,U\],

H⋆∈\[CL/U,CU/L\]\.H^\{\\star\}\\in\[C\_\{L\}/U,\\ C\_\{U\}/L\]\.The boundary is a set in\(ϕ,H\)\(\\phi,H\); a single phase\-transition point is reported only under a registered single\-crossing condition\. If deployment count is random, the expected charge isCinc​𝔼​\(1/H\)C\_\{\\mathrm\{inc\}\}\\mathbb\{E\}\(1/H\), notCinc/𝔼​\(H\)C\_\{\\mathrm\{inc\}\}/\\mathbb\{E\}\(H\)\(Supplementary PropositionLABEL:SI\-prop:adoption\)\.

The primary adoption object isH⋆​\(ϕ\)H^\{\\star\}\(\\phi\), not a favourable assumed reuse count\. Scenario definitions and facility\-scale anchors are reported in Supplementary SectionLABEL:SI\-sec:cost\.

#### Mechanism population\.

The prospective simulator samples six continuous mechanism axes and three\-class capacity\-demand proportions\. Numerical ranges are author\-specified stress envelopes, not estimated facility distributions\. The constrained\-risk population is conditional on mechanisms whose oracle\-optimal fixed action isK0=1K\_\{0\}=1, and each inferential unit is one new fixture with one 200\-bank campaign\. NULL and positive strata are defined before policy outcomes are computed\. Ranges, value functions and overlap checks are in Supplementary SectionsLABEL:SI\-sec:populationandLABEL:SI\-sec:study3\_design\.

#### Inference and design certification\.

Value models and fixed baselines are trained out of fold\. Comparisons use common random numbers within campaigns and aggregate to the declared inferential unit\. Cross\-fitting reduces reuse bias\[chernozhukov2018\]; one resample plan and multiplier/max\-ttcritical value protect each finite policy family\[chernozhukov2013\]\. Binomial risks use exact one\-sided Clopper–Pearson bounds\[clopper1934\]; finite campaigns use without\-replacement likelihoods\[serfling1974\]; and sequential looks use simultaneous validity\[howard2021\]\. Exact allocations and independently certified simulator sample sizes are given in Supplementary SectionsLABEL:SI\-sec:inferenceandLABEL:SI\-sec:study3\_design\.

#### Data provenance and evidence roles\.

The three internal layers use author\-generated mechanistic simulations\. Oracle variables are isolated from deployable policy features, except in the explicitly evaluation\-only simulator score in Eq\. \([15](https://arxiv.org/html/2607.27651#S4.E15)\)\. Facility documents provide scheduling or reuse context, not policy outcomes\. Causal Chambers contributes only a post\-selection constructibility audit of author\-defined archive\-backed proxies; it supplies neither native interventions nor causal policy value\[gamella2025,zhang2023causal\]\. The complete evidence\-role and source registry is Supplementary TableLABEL:SI\-tab:data\_sources\.

The locked pharmacogenomic exact\-outcome evaluation used measured dose–response summaries from CTRP v2\[seashoreludlow2015\]\. A metadata\-only audit selected a four\-compound panel before response access and required complete measured coverage for every retained cell\-line family\. Hash\-based, disjoint partitions assigned 125 families to development, 244 to calibration and 254 to locked evaluation\. Observable features were restricted to pre\-response cell\-line and assay metadata\. The frozen actionsK∈\{1,2,4\}K\\in\\\{1,2,4\\\}selected the topKKcompounds under a development\-trained ranker; action value was the best measured normalized response in the selected bundle minus the registered bundle cost\. Unlike logged\-policy evaluation, which generally requires action support plus weighting or outcome\-model assumptions\[swaminathan2015,jiang2016\], complete panel coverage made every retained action value a measured\-outcome lookup rather than an imputed counterfactual\.

The primary external estimand is finite and archive specific: the activation rate and paired value among the 254 hash\-assigned locked families, with NULL and POSITIVE subsets defined after outcome opening for evaluation\. Their observed counts are exact properties of this locked archive\. The protocol also reports Clopper–Pearson bounds as secondary repeated\-sampling summaries under independent\-family exchangeability; those bounds do not quantify uncertainty about the already observed finite archive or guarantee transport to future cell lines\. Truth strata and oracle headroom use locked responses only after assignment and are evaluation quantities, not policy inputs\. The exact partition, compound identifiers, score, threshold calibration and inference are in Supplementary SectionLABEL:SI\-sec:sr\_ctrp\.

The target\-calibrated morphology study used the public cpg0012 Cell Painting profiles\[wawer2014profiling,bray2017cellpaintingdata\]\. A design\-only audit reconstructed compounds and planned wells before profile access, yielding 25,013 eligible compound clusters\. Each cluster supplied one deterministically assigned initial replicate, two potentially authorized additional replicates and one held\-out verifier\. The utility of the initial or augmented estimate was half\-scaled cosine similarity to the verifier; net gain subtractedc=0\.02c=0\.02\. NULL, AMBIGUOUS and POSITIVE were respectively defined by net gain≤0\\leq 0, between 0 and5×10−35\\times 10^\{\-3\}, and≥5×10−3\\geq 5\\times 10^\{\-3\}\. Decision\-time inputs comprised the initial morphology, stateless hashed structure features and pre\-imaging plate metadata\.

The model family and gate thresholds were selected using 13,748 labelled development compounds, with the fitted multi\-head ExtraTrees predictor trained on the target\-tuning subset only\. The final rule was frozen before access to 11,265 final outcomes\. It required predicted NULL probability≤0\.395\\leq 0\.395, predicted POSITIVE probability≥0\.633333\\geq 0\.633333, predicted gain≥0\.053625\\geq 0\.053625and OOD percentile≤1\\leq 1\. Final base profiles were scored first; 595 assignments and all relevant hashes were sealed before the additional and verifier profiles were decrypted once\.

The four inferential endpoints formed a conjunctive intersection–union authorization rule: each marginal one\-sided bound was evaluated at 95%, and all four inequalities plus the finite\-frame activation, coverage and missingness guards had to pass\. Under valid component tests, this controls false authorization of the conjunction at 0\.05; it does not make the four marginal intervals a simultaneous 95% confidence set\. The archived “fixed\-sequence” field denoted evaluation and reporting order; no alpha recycling was used\.

Missing outcomes were recoded separately in the least\-favourable direction for each endpoint\. This produced false activation206/4453206/4453, sensitivity384/6534384/6534and an all\-compound worst\-case mean gain of1\.948×10−31\.948\\times 10^\{\-3\}\. Clopper–Pearson bounds and the 20,000\-resample compound\-level BCa value bound\[efron1987bca\]have a repeated\-sampling interpretation only under the declared exchangeable\-compound sampling frame; plate or batch dependence was not modelled\. The BCa lower bound was1\.229×10−31\.229\\times 10^\{\-3\}, whereas the secondary distribution\-free empirical\-Bernstein stress bound was−1\.099×10−3\-1\.099\\times 10^\{\-3\}\. Thus the positive archive\-wide worst\-case mean is exact for this final partition, and the positive population lower confidence bound is bootstrap\-based rather than a distribution\-free finite\-sample guarantee\.

The internal limits were 7\.5% false activation, 35% false discovery, 5% sensitivity, strictly positive executed value, 100 activations, 5% coverage and 5% final\-outcome missingness\. A separate preoutcome readiness screen also required base\-feature missingness below 5%; it is not one of the seven final components\. This is an author\-controlled, procedurally self\-blinded archive evaluation; model and threshold selection use development data, and the one\-shot final partition is the only scientific test\. The earlier, stricter three\-way assurance scenario and the amended single\-held\-out contract are distinguished in Supplementary SectionLABEL:SI\-sec:sr\_cpg0012\.

The simulation studies use author\-generated synthetic campaign ledgers\. The external CTRP study uses publicly released measured cancer\-cell\-line responses and is separately labelled as an offline pharmacological capacity abstraction\. The cpg0012 study uses publicly released Cell Painting profiles and is separately labelled as an offline adaptive\-replicate abstraction\. Figure source data, frozen evidence tables, validation reports, manifests and derived audit artefacts are included in the revised OPAL evidence and source\-data record at Zenodo \([https://doi\.org/10\.5281/zenodo\.21694517](https://doi.org/10.5281/zenodo.21694517)\); this includes the cpg0012 configuration, protocol amendments, final assignment seal, result\-bearing tables and 46\-check audit\. Third\-party raw files are not redistributed\.

The audited external repositories were[Olympus](https://github.com/aspuru-guzik-group/olympus)\(pinned commit440b6b58ebfcaa2391cff7e94b570fb4fda98d68\) and[Causal Chambers](https://github.com/juangamella/causal-chamber)\(pinned commit6da4e646cf62bc902f87911c866298a8c98ad2eb\)\. These were not core experimental outcome sources\. Fourteen scalar Causal Chambers archives were downloaded and verified against pinned checksums; 13 entered only the post\-selection Supplementary constructibility analysis\. Raw third\-party files are not redistributed\. The research archive includes the derived compatibility\-audit manifests, proxy\-mapping registry, prefix\-only replay audit, descriptive session\-level summaries and claim\-boundary metadata\. These physical archives contribute measured trajectories to a constructibility and archival\-prediction audit, but no causal policy outcome or independent risk\-control result\.

CTRP v2 data are available from the[Broad Institute CTRP portal](https://portals.broadinstitute.org/ctrp.v2.1/)and the[NCI CTD2Data Portal](https://ocg.cancer.gov/programs/ctd2/data-portal); the resource and assay are described by Seashore–Ludlow*et al\.*\[seashoreludlow2015\]\. The pinned archiveCTRPv2\.0\_2015\_ctd2\_ExpandedDataset\.ziphad SHA\-2568f62b3b5ed70cfd367cf52ce0a99884dd0a674d1a8c301474b707648689bdee3; the normalized reference\-AUC table used by the frozen parser had SHA\-25600e6b8645cd75fe7a62a742fbfda04ae22fcfbb56b0f0fad3d7fb802f15303bc\. Raw CTRP files are not redistributed\. Olympus contributes no outcome\. No private facility log or patient/human\-participant dataset was used; CTRP contains established cell\-line assay data rather than human\-participant records\.

The cpg0012 profiles are from the Cell Painting Gallery targetcpg0012\-wawer\-bioactivecompoundprofiling, originally described by Wawer*et al\.*and Bray*et al\.*\[wawer2014profiling,bray2017cellpaintingdata\]\. The frozen source object wascpg0012\-wawer\-bioactivecompoundprofiling/broad/workspace/gigascience\_profiles/CDRP\.tar\.gz, with recorded size, ETag and manifest hashes retained in the evidence bundle\. The study used per\-well profiles, published plate maps and compound annotations, not patient or human\-participant records\.

The complete generator, reducers, exact finite\-population enumeration, design\-certification programs, CTRP parsers, cpg0012 frozen\-assignment and final evaluators, validators and tests required for end\-to\-end regeneration are available at[https://github\.com/JIABI/mard\_autonomous](https://github.com/JIABI/mard_autonomous)\. Every executable stage records its configuration, source\-tree hash, seed namespace and output manifest\. The frozen evidence and figure source data are archived separately at[Zenodo 10\.5281/zenodo\.21694517](https://doi.org/10.5281/zenodo.21694517), including the cpg0012 configuration, protocol amendments, final assignment seal, final result and read\-only audit\.

This work was supported by the CNPC Innovation Fund \(grant no\. 2024DQ02\-0501\), the Royal Society \(grant no\. IECNSFC233444\), Intelligent Manufacturing Longcheng Laboratory \(grant no\. CJ20254004\) and the Youth Science and Technology Talent Promotion Project of Jiangsu Province \(grant no\. JSTJ\-2025\-137\)\. We are grateful to colleagues at Diamond Light Source for discussions that helped frame the motivating operational problem: how a scientific facility should authorize adaptive experimentation when instrument configurations, staffing and experimental capacity must be committed before all decision\-relevant evidence becomes available\.

\\bmhead

Competing interests The authors declare no competing interests\.

## 5Extended Figures

![[Uncaptioned image]](https://arxiv.org/html/2607.27651v1/x7.png)

Figure 1:Extended Data Figure 1\. Estimand and comparator reconciliation in the retrospective selector\-frontier audit\.a, K\-only surrogate and executed contrasts over the same frozenλ\\lambdafamily\. The surrogate uses label\-derived value againstKscore,−f=1K\_\{\\mathrm\{score\},\-f\}=1; the executed estimand uses operational utility againstK0=2K\_\{0\}=2\.b, Four intermediate contrasts that isolate value\-source, policy\-rule and comparator effects\.c, Distribution of the reconciliation quantity across the registered fold/cell aggregation\. The retrospective selector\-frontier audit is retrospective and descriptive\.
![Refer to caption](https://arxiv.org/html/2607.27651v1/x8.png)Figure 2:Extended Data Figure 2\. Model\-complexity and action\-change diagnostics\.a, Candidate residual\-expert countsE=0,…,6E=0,\\ldots,6in each outer fold\. Filled diamonds mark the nested\-cross\-validation choice under the one\-standard\-error rule; all five folds selected the global modelE=0E=0\.b, Confusion matrix for the learned capacity and information oracle over 720 in\-support banks\. The selector choseK=1K=1139 times, of which three matched the oracle\.c, Probe outcomes over 810 banks\. Counts are displayed on a symmetric\-log scale: 800 no\-purchase decisions, eight purchases without an action change and two action\-changing purchases\.![Refer to caption](https://arxiv.org/html/2607.27651v1/x9.png)Figure 3:Extended Data Figure 3\. Exact finite\-campaign coverage audit\.a, Empirical joint\-coverage shortfall per 1,000 for the exact composition–channel confidence set over 21 combinations of pilot size and channel accuracy\. The dashed line is the registered maximum shortfall corresponding to coverage0\.997140\.99714; observed coverage was0\.99820\.9982–1\.00001\.0000\.b, Numbers of campaigns contributing to each operating\-characteristic endpoint as a function ofqq\. Counts vary because truth strata are assigned from the finite realized campaign and deployable opportunity\.![Refer to caption](https://arxiv.org/html/2607.27651v1/x10.png)Figure 4:Extended Data Figure 4\. Finite\-campaign information cost and lifecycle scenarios\.a, Best lifecycle contrast and its range over all registered pilot sizes as reuse populationHHincreases\. The entire range remains below zero because operational contrast is already non\-positive before shared cost is added\.b, Information bill per original 16\-bank campaign\. No finite positiveH⋆H^\{\\star\}exists at any displayedqq\. TheH=100,000H=100\{,\}000column is a unit\-relaxed stress scenario, not a campaign\-level adoption estimate\.![Refer to caption](https://arxiv.org/html/2607.27651v1/x11.png)Figure 5:Extended Data Figure 5\. NULL\-only risk calibration and descriptive positive\-population heterogeneity\.a, One\-sided 99\.5% Clopper–Pearson upper risk bound across order\-statistic thresholds computed from 140 NULL tuning fixtures\. The horizontal line is the 0\.05 target and the vertical line is the selectedτ=0\.0180806\\tau=0\.0180806\.b, Tuning and independent risk\-lock scores with the frozen threshold\. Jitter is used only for display\. The risk\-lock set contained 500 disjoint fixtures and produced zero activations; no positive or validation outcome entered threshold selection\. Positive\-validation fixtures were divided into quartiles ofc, full\-state opportunityGMG\_\{M\};d, survey separability; ande, routing\-cost multiplier\. Points are activation fractions and error bars are two\-sided 95% exact Clopper–Pearson intervals\. The dotted line is the overall sensitivity0\.6840\.684\. These post hoc descriptive strata carry no binding claim and were not used to modifyτ\\tau\.![Refer to caption](https://arxiv.org/html/2607.27651v1/x12.png)Figure 6:Extended Data Figure 6\. Descriptive shared\-score gate ablation for the held\-out simulator\.a, NULL false\-activation rate against positive\-mechanism sensitivity for seven rules evaluated on the same 400 NULL and 500 positive validation fixtures\. Error bars are two\-sided 95% exact Clopper–Pearson intervals; the vertical line is the original 0\.05 risk limit\. The horizontal axis uses a symmetric\-log transform to retain zero\-rate rules and the always\-activate endpoint\.b, Paired sensitivity differences from the risk\-controlled rule with unadjusted 95% fixture\-bootstrap intervals\. The right column gives the corresponding false\-activation\-rate difference\. The oracle label is non\-deployable\. This ablation reuses one oracle\-assisted evaluation score and changes only its threshold; it is not a full nearest\-neighbour method comparison\. It was specified after the primary held\-out result was known and is therefore retrospective and descriptive; it does not modify the original binding claim\.![Refer to caption](https://arxiv.org/html/2607.27651v1/x13.png)Figure 7:Extended Data Figure 7\. Measured opportunity and frozen\-policy diagnostics in CTRP\.The locked pharmacogenomic exact\-outcome evaluation is an offline, prospectively partitioned evaluation of an author\-defined capacity abstraction; it is not an online treatment policy or physical\-facility deployment\.a, One\-sided 99\.5% exact NULL false\-activation upper bounds on the calibration partition\. Atτ=0\\tau=0, 6/236 NULL families activated \(upper bound0\.06500\.0650\); the frozen thresholdτ=0\.5\\tau=0\.5produced 0/236 \(upper bound0\.02220\.0222\) and was selected before locked\-response access\.b, Locked observable scores versus full\-state oracle headroom\. Colours denote outcome\-defined truth strata, which are evaluation labels and are unavailable to the policy\. The maximum score was0\.3454<0\.50\.3454<0\.5, including all seven POSITIVE families\.c, Distributions of measured full\-state headroom and executed gain over 254 locked families\. Complete four\-compound response coverage makes both quantities exact measured\-outcome lookups under the frozen bundle definition; no response model or propensity weighting is used\. Executed gain was zero because every assignment fell back to fixedK=2K=2\.d, Exact locked\-evaluation results\. NULL false activation was 0/244 \(one\-sided 99\.5% upper bound0\.0215<0\.050\.0215<0\.05\); positive\-family sensitivity was 0/7 \(two\-sided 95% interval0–0\.4100\.410\); and the one\-sided 97\.5% lower bound for paired gain was zero\. The binding joint claim required both risk control and a strictly positive value lower bound and was therefore not supported\.![Refer to caption](https://arxiv.org/html/2607.27651v1/x14.png)Figure 8:Extended Data Figure 8\. Post\-selection constructibility on public physical\-system archives\.This analysis uses author\-constructed archive\-backed predictive measurement\-capacity proxies; it is not an independent, causal or risk\-controlled physical\-system validation\.a, Numbers of archived experiment sessions in 13 scalar Causal Chambers datasets\. Dark bars denote datasets with experiment\-level splits eligible for session\-level descriptive summaries; light bars denote descriptive time splits\.b, Fractions of sessions with a finite pilot\-prefix score, an active proxy action and an out\-of\-distribution flag\. All 13 schemas admitted score–action–fallback execution, but activity is not a physical treatment effect\.c, Mean normalized archival predictive loss over 30 locked descriptive sessions\. Lines are session\-bootstrap intervals\. The archive oracle uses suffix outcomes and is explicitly non\-deployable\. Intervals overlap broadly and no descriptive predictive advantage overK0K\_\{0\}was observed\.d, Evidence boundary\. Suffix mutation produced zero score, action, fallback\-reason or deterministic\-replay mismatches across 196 sessions, establishing a property of the code data flow\. None of the 13 proxy mappings was a documented physicalKKaction, and no archive met the independent\-session plus propensity gate for formal prospective physical\-policy validation\.
## 6Extended Tables

Table 1:Extended Data Table 1\. Selector\-frontier estimand and baseline crosswalk\.Internal analysis identifiers are replaced by their scientific definitions\.EstimandOutcome/value sourceComparatorMean contrastScientific roleK\-only surrogateLabel\-derived per\-KKvalueKscore,−f=1K\_\{\\mathrm\{score\},\-f\}=1\+5\.549×10−3\+5\.549\\times 10^\{\-3\}Surrogate decision value; not adoptionExecuted operationalExecuted no\-probe policy utilityK0=2K\_\{0\}=2−0\.814×10−3\-0\.814\\times 10^\{\-3\}Correct operational comparisonIncremental\-cost lifecycleExecuted operational plus incremental shared costK0=2K\_\{0\}=2No finiteH⋆H^\{\\star\}Fixed\-selection cost not identified in frozen ledgerCommon\-cost lifecycleExecuted operational with common\-cost cancellationK0=2K\_\{0\}=2No finiteH⋆H^\{\\star\}Lifecycle contrast equals non\-positive operational contrastTable 2:Extended Data Table 2\. Constrained\-risk calibration and prospectively held\-out simulator validation\.PartitionFixturesEventsExact interval/boundRoleNULL tuning1400 above selected maximumSelectedτ=0\.0180806\\tau=0\.0180806Threshold construction onlyNULL risk lock5000 activations99\.5% upper bound 0\.01054Independent lock; passedNULL validation4000 activations99\.5% upper bound 0\.01316Binding claim; risk limit 0\.05Positive validation500342 activations95% CI 0\.641–0\.725Sensitivity estimate 0\.684; no pass/fail target
## Figure legends

Figure[1](https://arxiv.org/html/2607.27651#S2.F1)\|\|Certifying adaptive opportunity under asynchronous evidence and commitment\.a, A sample or condition bank is measured on a scientific instrument, but capacity for later branching must be reserved from metadata and an early pilot signal\.K0K\_\{0\}is the study\-specific fixed baseline; lower and higher alternatives are generic resource levels rather than literal facility products\. Evidence arriving after the commitment deadline can change the run only if the corresponding capability was reserved\.b, A policy can use only the information available at its decision time\. Survey opportunityGXG\_\{X\}and survey\-plus\-probe opportunityGX​ZG\_\{XZ\}are bounded by full\-state oracle opportunityGMG\_\{M\}, so0≤GX≤GX​Z≤GM0\\leq G\_\{X\}\\leq G\_\{XZ\}\\leq G\_\{M\}\. Bar widths encode this conceptual ordering, not empirical magnitudes\.c, In a finite campaign, the observed pilotPqP\_\{q\}is the input to a certificate whose inference target is the unobserved deployment complementDqD\_\{q\};DqD\_\{q\}is not observed by the decision rule\.Opaluses deployment\-isomorphic scoring and exact finite\-population inference, activating only with a risk\-valid certificate and otherwise returning the frozen capacityK0K\_\{0\}\. Information cost and net value retain the original campaign denominatorNN\.d, Adoption requires controlled false activation, positive executed operational gainΔop\\Delta\_\{\\mathrm\{op\}\}and a reuse populationHHsufficient to amortize incremental shared costCincC\_\{\\mathrm\{inc\}\}\. The decreasing boundaryH⋆=Cinc/ΔopH^\{\\star\}=C\_\{\\mathrm\{inc\}\}/\\Delta\_\{\\mathrm\{op\}\}separates combinations that can cover that cost \(above the curve\) from those that cannot \(below the curve\)\.

Figure[2](https://arxiv.org/html/2607.27651#S2.F2)\|\|Correct execution, baseline identity and information cost determine apparent adaptive value\.This figure is a retrospective descriptive reanalysis of frozen development data; it is not a confirmatory policy comparison\.a, Executed operational value of the six\-member selector family relative to executed fixedK=2K=2\(n=720n=720in\-support banks across 45 cells\)\. Points are means and error bars are simultaneous max\-ttintervals from one shared 2,000\-replicate, cell\-stratified bootstrap family\. Everyλ\\lambdachanges only the critical threshold applied to one cross\-fitted score; no model or resample family is re\-selected at a frontier point\.b, Exact reconciliation from the K\-only surrogate to the executed estimand\. Bars after the surrogate show signed changes; correcting the baseline fromKscore,−f=1K\_\{\\mathrm\{score\},\-f\}=1toK0=2K\_\{0\}=2determines the reversal\. Rounding explains the displayed10−610^\{\-6\}\-scale difference from the exact final value−0\.814×10−3\-0\.814\\times 10^\{\-3\}\.c, Executed3×33\\times 3factorial ablation\. Cells give operational contrasts against fixedKKwith no probe \(n=720n=720\)\. Soft\-expert rows equal global learned\-KKrows because nested cross\-validation selectedE=0E=0in every outer fold\. Full\-panel cells are saturated in colour but retain their numerical values\.d, Distribution of gross and net decision\-focused probe option values over 810 banks\. White boxes show the interquartile range, central lines are medians, violins show the distribution and points are a reproducible subsample for visibility\. The arrow marks the reduction after probe and latency costs\. Ten probes were purchased, two changedKK, and the realized in\-support mean contribution was−5\.86×10−4\-5\.86\\times 10^\{\-4\}, with simultaneous interval\[−1\.172×10−3,0\]\[\-1\.172\\times 10^\{\-3\},0\]\.

Figure[3](https://arxiv.org/html/2607.27651#S2.F3)\|\|Exact finite\-campaign inference exposes a certification boundary rather than an estimator failure\.a, Full three\-class certification envelope\. At each pilot size, the line is the largest costed lower bound over every possible signal\-count vector and registered identity\-mixture channel; the narrow purple region shows the range across channel parametersp∈\{0\.70,0\.80,0\.92\}p\\in\\\{0\.70,0\.80,0\.92\\\}\. The dotted zero line is the maximum uncosted gain envelope\. All 1,113 observations retained a least\-favourable composition with non\-positive gain and were classifiedINDETERMINATE\.b, Simultaneous exact upper probability bounds over pilot size\. The actionable series is the upper bound on sensitivity after zero observed activations\. The NULL and nonactionable series are upper bounds on erroneous activation, obtained as one minus the simultaneous lower bounds for specificity and abstention\. This common error\-probability scale separates endpoints whose empirical rates otherwise coincide at zero or one\. Denominators vary withqqand truth stratum and are reported in Supplementary TableLABEL:SI\-tab:sr\_finite\_oc\.c, Positive cost magnitudes displayed side by side\. Because the rule always retainedK0K\_\{0\}, the observed fallback loss exactly equals the registered information bill\(0\.1\+0\.5​q\)/16\(0\.1\+0\.5q\)/16at every pilot size\.d, Full\-state opportunityGMG\_\{M\}\(open circles\) and the maximum expected deployable opportunityGA,DG\_\{A,D\}over registered pilot sizes \(filled diamonds\) for each fixture\. Lines join quantities from the same fixture\. NULL, weak and positive\-unobservable fixtures haveGA,D=0G\_\{A,D\}=0, whereas transition and actionable fixtures retain some deployable opportunity\. Despite this opportunity, none of 7,442 sequential campaigns activated and every campaign reached the maximumq=14q=14\.

Figure[4](https://arxiv.org/html/2607.27651#S2.F4)\|\|A strict\-majority slice of the general finite\-population certifiability boundary\.a, For deployment\-complement targetρ=1/2\\rho=1/2, the minimum pilot fractionq†/Nq^\{\\dagger\}/Nfor which an all\-actionable simple\-random pilot is not excluded by the level\-α\\alphaboundary\-null condition\. Beige cells have noq∈\{1,…,N−1\}q\\in\\\{1,\\ldots,N\-1\\\}that escapes this impossibility condition; the marker is the registered pointN=16,α=0\.02/7N=16,\\alpha=0\.02/7\.b, Minimum pilot size not excluded by the binary condition overNNfor four error levels\. At0\.02/70\.02/7, the first such size isN=25,q†=14N=25,q^\{\\dagger\}=14;N=24,q=15N=24,q=15appears only after rounding to3×10−33\\times 10^\{\-3\}\. The sawtooth reflects integer parity in⌊\(N−q\)/2⌋\\lfloor\(N\-q\)/2\\rfloor\. This is a necessary boundary only: noise, cost and the complete three\-class rule can still prevent certification\.

Figure[5](https://arxiv.org/html/2607.27651#S2.F5)\|\|Prospectively held\-out calibration controls simulator false activation while retaining measurable sensitivity\.a, Empirical cumulative distributions of the evaluation\-only score for four disjoint partitions: NULL tuning \(n=140n=140\), NULL risk lock \(n=500n=500\), NULL validation \(n=400n=400\) and positive validation \(n=500n=500\)\. The vertical dashed line is the frozenτ=0\.0180806\\tau=0\.0180806\.b, Evaluation score against oracle full\-state opportunity for the 900 held\-out simulator validation fixtures\. Each point is one distinct mechanism fixture with one fresh 200\-bank campaign\. The horizontal line is the frozen activation threshold;GMG\_\{M\}andSevalS^\{\\mathrm\{eval\}\}are evaluation\-only\.c, False\-activation counts and one\-sided 99\.5% exact upper confidence bounds in the independent risk lock and final NULL validation\. Error bars extend from the observed rate of zero to the Clopper–Pearson upper limit; the dashed line is the registered 0\.05 risk limit\.d, Retrospective descriptive value decomposition on the 500 positive simulator fixtures\. Bars report paired value relative toK0K\_\{0\}for the gated policy over all fixtures, activated fixtures, executed fallback fixtures and a force\-activate comparator; error bars are percentile intervals from 10,000 fixture bootstrap resamples\. The first three summarize the frozen gate, whereas force activation applies the learned bank\-level action to every positive fixture\. The paired gated\-minus\-forced difference was8\.60×10−58\.60\\times 10^\{\-5\}\(95% interval\[−6\.78×10−4,8\.75×10−4\]\[\-6\.78\\times 10^\{\-4\},8\.75\\times 10^\{\-4\}\]\)\. BecauseSevalS^\{\\mathrm\{eval\}\}is computed from potential outcomes, these are evaluation\-only diagnostics, not deployable value estimates\.

Figure[6](https://arxiv.org/html/2607.27651#S2.F6)\|\|Measured\-outcome studies separate source\-policy misalignment from target\-calibrated authorization\.a, CTRP calibration rejectedτ=0\\tau=0because its one\-sided 99\.5% NULL false\-activation upper bound exceeded the 5% limit, and selectedτ=0\.5\\tau=0\.5\.b, All 254 locked CTRP scores remained belowτ=0\.5\\tau=0\.5, including the seven outcome\-defined POSITIVE families\. This is evidence of score–policy misalignment under the registered representation, not an absence of pharmacological opportunity\.c, In the cpg0012 final test, circles are point estimates, triangles are one\-sided 95% confidence bounds and short black ticks are the internally frozen requirements\. The false\-activation UCB was 5\.18%, below its 7\.5% limit\. The FDP point estimate was34\.62%<35%34\.62\\%<35\\%, whereas its 95% UCB was37\.97%\>35%37\.97\\%\>35\\%\. The sensitivity bound passed; the fixed\-archive worst\-case mean gain was1\.948×10−31\.948\\times 10^\{\-3\}, and its compound\-bootstrap BCa lower endpoint was1\.229×10−3\>01\.229\\times 10^\{\-3\}\>0\.d, Comparator methods were evaluated on identical final compounds, decision\-time observables, action and cost\. Each point is one complete method; the dashed line is the 7\.5% false\-activation UCB limit\.Opalwas the only nonzero\-activation method below that limit, but no method met the complete joint contract\.

## References

相似文章

基于局部披露的具有策略性主体的离线策略评估

arXiv cs.AI

本文研究当决策主体(智能体)为了回应策略而策略性地修改其协变量时的离线策略评估(OPE)。该方法利用事后解释进行局部披露,以揭示智能体的前策略协变量,并构建策略价值的双重稳健估计量。

PACE:自进化代理的任意有效验收测试

arXiv cs.AI

PACE 为自进化代理引入了一种任意有效的提交门,它用序贯假设检验替代贪婪接受,控制错误提交概率,减少震荡,同时保持性能且方差更低。