Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
Summary
This paper tests the independence assumption underlying compositional reliability bounds for multi-agent systems, finding that same-model agents co-fail at high rates and that common certificates are unsound. It proposes a finite-sample, dependence-free certificate via linear programming over co-execution moments, validated on 18,000 missions.
View Cached Full Text
Cached at: 08/14/26, 09:28 AM
# Certifying Compositional Reliability Without Assuming Independence
Source: [https://arxiv.org/html/2608.12895](https://arxiv.org/html/2608.12895)
## Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
Varun Pratap BhardwajAffiliation:Qualixar / Independent Researcher, IndiaEmail:[varun\.pratap\.bhardwaj@gmail\.com](mailto:)Affiliation:\[1pt\]ORCID: 0009\-0002\-8726\-4289Arun Pratap BhardwajAffiliation:Independent Researcher, IndiaEmail:[arun\.pratap\.bhardwaj@gmail\.com](mailto:)
###### Abstract
Compositional reliability bounds for multi\-agent systems multiply component reliabilities, a step licensed by a conditional\-independence assumption that is routinely stated and rarely tested\. We test it\. Two instances of one model, composed in a two\-agent handoff, co\-fail on90\.0%90\.0\\%of the missions on which either fails \(logOR=6\.66\\log\\mathrm\{OR\}=6\.66, 95% CI\[6\.38,7\.00\]\[6\.38,7\.00\];ϕ=0\.916\\phi=0\.916\)\. The evidence is a preregistered confirmatory evaluation of18,00018\{,\}000missions with deterministic scoring and no model in the judging loop, inside a larger campaign whose other topologies we report as secondary\. Substituting a different model reduces the association significantly in the confirmatory motif and in both secondary topologies \(six of six contrasts\); substituting a different*vendor*, with the model already different, does not — a registered hypothesis that fails to replicate and that we report as a null\. An unmanipulated same\-model pair, present in every arm, returns fifteen of fifteen null contrasts: the result a design free of graph\-wide confounding would produce, and one that would have been informative had it come out otherwise\.
The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over\-credited exactly when components share a model\. The assumption\-free alternative is often vacuous — the certified floor is zero whenever mean component reliability falls below1−1/m1\-1/m— and fitting a dependence model is worse: we prove that a bootstrap bound on a fitted model’s functional loses coverage of the true reliability asn→∞n\\to\\infty, because the identification gap isO\(1\)O\(1\)while the bootstrap haircut isO\(n−1/2\)O\(n^\{\-1/2\}\)\. More data makes such a certificate worse, with no visible symptom\.
We give a finite\-sample certificate that assumes no dependence structure: a linear program over the joint, taken over a Bonferroni–Clopper–Pearson box around measured co\-execution moments\. It is sound, sharp for the information supplied, and monotone in the moment family under a stated Bonferroni allocation\. On four\-stage data, enriching from ten moment functionals to fourteen narrows the identified interval by85\.7%85\.7\\%and lifts the certified floor from0\.24550\.2455to0\.41160\.4116\. A companion anytime\-valid certificate holds its empirical type\-I error at0\.04710\.0471or below across every admissible betting fraction, recovering the SPRT exactly at the optimal bet\. An ablation at a conceded design effect1\.311\.31times the largest measured moves the floor by at most2\.692\.69percentage points\.
We also show that the dependence statistics in common use — Jaccard,ϕ\\phi, Kendall’sτa\\tau\_\{a\}— are bounded by the marginals and can significantly*reverse*an apparent ordering of conditions when the compared agents fail at different rates, a reversal we observe and then replicate on two further inference backends\. The contracts, the mission generators, the scoring code, the analysis scripts, and the preregistration are released; every reported statistic is regenerated by those scripts rather than transcribed\.
## 1Introduction
### 1\.1The independence gap
A multi\-agent pipeline is certified the way a series system is certified: bound each component’s reliability, multiply, report the product\. The v1 Agent Behavioral Contract framework\([Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\)does this, and so does every compositional reliability argument we are aware of for agent systems\. The step is licensed by a conditional\-independence condition — C5 in[Definition3\.10](https://arxiv.org/html/2608.12895#S3.Thmdefinition10)— which asserts that a downstream agent’s contract satisfaction is independent of an upstream agent’s internal execution given a compliant handoff\.
C5 is stated in v1 and never tested\. It is also, on reflection, implausible exactly where multi\-agent systems are most often deployed: a reviewer agent checking a writer agent is very often the same model with a different prompt\. Two instances of one model do not have independent blind spots\. They share them\.
We measure how much this matters\. The preregistered confirmatory evaluation is a two\-agent handoff over 18,000 missions with deterministic scoring and no model in the judging loop; the registration places the other topologies of the campaign in secondary work, and we report them as such\. Two instances ofmistral\-small\-24bco\-fail on90\.0%90\.0\\%of the missions on which either fails, withlogOR=6\.66\\log\\mathrm\{OR\}=6\.66\(95% CI\[6\.38,7\.00\]\[6\.38,7\.00\]\)\. Replacing the second agent with a different model reduces the association significantly, in the confirmatory motif and in both secondary topologies\. Replacing the*vendor*, with the model already different, does not\.
The consequence runs against the operator\. By[Proposition4\.1](https://arxiv.org/html/2608.12895#S4.Thmproposition1)the compositional gap is signed: under positive dependence, joint failure exceeds the independence product, so redundancy is over\-credited precisely when the redundant components share a model\. A dashboard that multiplies reliabilities reports a reassuring number whose governing assumption the data reject, and nothing in the pipeline signals it\.
### 1\.2Why the obvious repairs fail
Dropping the assumption gives the Fréchet–Hoeffding bounds, which are sharp and frequently vacuous: by[Corollary4\.2](https://arxiv.org/html/2608.12895#S4.Thmcorollary2)the certified floor is exactly zero whenever mean component reliability falls below1−1/m1\-1/m, which for four components atp=0\.75p=0\.75is the entire regime of interest\.
Fitting a dependence model — a Gaussian one\-factor copula, the standard choice in reliability practice — is worse than either\.[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)proves that a bootstrap lower bound on the fitted model’s functional loses coverage of the*true*reliability asn→∞n\\to\\infty: the identification gap isO\(1\)O\(1\)while the bootstrap haircut shrinks liken−1/2n^\{\-1/2\}, so past some finite sample size the interval sits entirely above the truth and never returns\. Collecting more data does not repair such a certificate; it narrows the interval around the wrong target, and nothing in the interval signals it\.
### 1\.3Approach
We constrain the joint law with measured co\-execution moments and optimise over everything consistent with them\. The resulting certificate is a linear program over the2m2^\{m\}cells of the joint, taken over a Bonferroni–Clopper–Pearson box around the empirical moments\. It is sound with no dependence assumption \([Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)\), sharp for the information supplied \([Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)\), and tightens monotonically as the moment family grows \([Propositions6\.1](https://arxiv.org/html/2608.12895#S6.Thmproposition1)and[6\.2](https://arxiv.org/html/2608.12895#S6.Thmproposition2)\)\. On real four\-stage data, enriching from ten moment functionals to fourteen narrows the identified interval by85\.7%85\.7\\%and lifts the certified floor from0\.24550\.2455to0\.41160\.4116\.
Because deployed systems are monitored continuously rather than at a sample size fixed in advance, we add a certificate valid under arbitrary stopping \([Theorem7\.1](https://arxiv.org/html/2608.12895#S7.Thmtheorem1)\)\. Its null constrains only a conditional mean, so it requires no independence assumption at all — it is immune to the failure the rest of the paper documents\.
### 1\.4Contributions
C1 \(co\-lead\)A preregistered evaluation that*manipulates*model sharing at three levels, confirmatory over 18,000 two\-agent\-handoff missions and replicated in two secondary topologies within a 30,820\-mission campaign, converting an observation about particular systems into a contrast attributable to substitution\. Correlated failure among same\-backbone agents is itself established by concurrent work \([Section2\.3](https://arxiv.org/html/2608.12895#S2.SS3)\); the controlled design, and the finding that the ordering holds at the model level and vanishes at the vendor level, are ours\.
C2 \(lead\)A finite\-sample, copula\-agnostic reliability certificate for composed agent pipelines \([Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)\), with a Bonferroni allocation that makes the floor monotone in the moment family by construction\.
C3 \(supporting\)[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2): coverage of a model\-based reliability floor tends to zero as the sample grows\. The phenomenon is misspecification bias; what we add is the explicit witness of[Example4\.2](https://arxiv.org/html/2608.12895#S4.Thmexample2), indistinguishable from the fitted family on the moments the model uses\.
C4 \(co\-lead\)Anytime\-valid certification of graph reliability, with the SPRT recovered exactly at the optimal bet \([Proposition7\.1](https://arxiv.org/html/2608.12895#S7.Thmproposition1)\)\.
C5 \(supporting\)A demonstration that marginal\-sensitive dependence statistics can significantly*reverse*an apparent condition ordering, replicated on independent backends \([Sections10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)and[10\.6](https://arxiv.org/html/2608.12895#S10.SS6)\)\.
C6 \(supporting\)An internal negative control returning fifteen of fifteen nulls, and a full artifact release including the preregistration\.
[Section12](https://arxiv.org/html/2608.12895#S12)states what we do*not*claim, including one registered hypothesis that fails to replicate\.
### 1\.5Organisation
[Section2](https://arxiv.org/html/2608.12895#S2)places the work against contract\-based specification, guardrails, the recent correlated\-failure literature, and dependence\-free bounding\.[Section3](https://arxiv.org/html/2608.12895#S3)states the ABC framework in full, including formulas withheld from v1 under a patent claim now withdrawn\.[Section4](https://arxiv.org/html/2608.12895#S4)establishes what breaks without C5\.[Sections5](https://arxiv.org/html/2608.12895#S5),[6](https://arxiv.org/html/2608.12895#S6)and[7](https://arxiv.org/html/2608.12895#S7)construct the certificate\.[Sections8](https://arxiv.org/html/2608.12895#S8)and[9](https://arxiv.org/html/2608.12895#S9)describe runtime enforcement and the implementation\.[Section10](https://arxiv.org/html/2608.12895#S10)reports E1–E6\.[Section11](https://arxiv.org/html/2608.12895#S11)covers limitations and threats to validity, and[Sections12](https://arxiv.org/html/2608.12895#S12),[13](https://arxiv.org/html/2608.12895#S13)and[14](https://arxiv.org/html/2608.12895#S14)close\.
## 2Background and Related Work
### 2\.1Contracts for software and for agents
Design by Contract\([Meyer 1992](https://arxiv.org/html/2608.12895#bib.bib39);[Hoare 1969](https://arxiv.org/html/2608.12895#bib.bib23)\)specifies preconditions, postconditions, and invariants that a component must honour, and makes composition tractable by letting each component’s postcondition discharge the next one’s precondition\. The formal\-methods lineage that followed\([Clarke et al\. 1999](https://arxiv.org/html/2608.12895#bib.bib11);[Lamport 2002](https://arxiv.org/html/2608.12895#bib.bib33);[Leino 2010](https://arxiv.org/html/2608.12895#bib.bib34);[Barnett et al\. 2004](https://arxiv.org/html/2608.12895#bib.bib3)\)verifies such specifications statically, static analysis bounds behaviour by abstraction\([Cousot and Cousot 1977](https://arxiv.org/html/2608.12895#bib.bib13)\), and dynamic\-invariant work\([Ernst et al\. 2007](https://arxiv.org/html/2608.12895#bib.bib17)\)infers them from traces\.
None of it transfers directly to agents driven by natural\-language instructions, because the specification surface is prompts rather than types and the execution is stochastic\. The v1 framework\([Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\)closes that gap by making preconditions, hard and soft invariants, governance policies, and recovery mechanisms first\-class and runtime\-enforceable, and by replacing deterministic satisfaction with the probabilistic\(p,δ,k\)\(p,\\delta,k\)\-notion of[Definition3\.6](https://arxiv.org/html/2608.12895#S3.Thmdefinition6)\. We take that framework as given and measure what happens when several contracted agents are composed\.
### 2\.2Steering, filtering, and guarding
Three families of technique constrain agent behavior, and none of them certifies a composed system\.
Training\-time alignment — Constitutional AI\([Bai et al\. 2022](https://arxiv.org/html/2608.12895#bib.bib1)\), RLHF\([Ouyang et al\. 2022](https://arxiv.org/html/2608.12895#bib.bib42)\)— shapes general tendencies but cannot encode a deployment\-specific invariant, and offers no runtime guarantee\. Output guardrails\([Rebedea et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib49)\)filter or redirect responses matching prohibited patterns at inference time; the current generation adds programmable policy languages and validator libraries, and operates per\-response\. Observability platforms trace and score agent runs after the fact\.
The common limitation is scope: each acts on a single turn or a single agent\. None specifies invariants over a multi\-agent pipeline, and none produces a statement of the form “this composition satisfies its contract with probability at leastL^\\hat\{L\}\.”[Table1](https://arxiv.org/html/2608.12895#S2.T1)makes the comparison explicit\.
Table 1:Capability comparison against representative approaches\. ✓ = supported,∘\\circ= partial, — = not addressed\.REruntime enforcement ·FSformal specification ·MAmulti\-agent scope ·DMdependence measured ·CBcomposed reliability bound ·CAcopula\-agnostic ·AVanytime\-valid\. The last three are the contribution of this paper;DMis where concurrent work has recently arrived\.
### 2\.3Correlated failure in multi\-agent systems
That agents sharing a base model may fail together is no longer a conjecture, and we state plainly what is already established so that our own contribution is not overstated\.
[McDonnell et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib38)study a three\-agent triage architecture and quantify correlated failure directly, reporting a joint error rate inflated3\.53×3\.53\\timesover independence \(BCa 95% CI\[3\.50,3\.59\]\[3\.50,3\.59\]\) withϕ=0\.612\\phi=0\.612, and finding57\.2%57\.2\\%of errors occurring under agent agreement\. Their agents are classical learners — a random forest, akk\-nearest\-neighbour model, and a calibrated meta\-model — over intrusion\-detection and clinical\-readmission data, and model sharing is not a manipulated variable; the design varies learner*family*by construction rather than varying it experimentally\.[Huang et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib28)recalibrate multi\-agent confidence against a counterfactual no\-communication baseline and observe, qualitatively, that within\-family blocks of agents sharing a backbone exhibit shared blind spots\. The[Bengio et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib4)report lists correlated failure among same\-base\-model agents as a known concern, as does the survey of[Hammond et al\. 2025](https://arxiv.org/html/2608.12895#bib.bib22); work on governed capability evolution\([Qin et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib46)\)addresses the related problem of keeping a guarantee valid as components are upgraded\. Work on failure attribution\([Rafi et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib47);[Qiao et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib45)\)traces which agent caused a trajectory to fail, and Byzantine\-tolerance analyses\([Zheng et al\. 2025](https://arxiv.org/html/2608.12895#bib.bib64);[Berdoz et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib5)\)study whether agents can reach agreement at all under faults\.
Three things separate the present work from that literature\. First, model sharing here is a*manipulated*experimental variable — three sharing levels, preregistered, with a confirmatory core of18,00018\{,\}000missions on the two\-agent handoff motif and two further topologies reported as secondary replication — rather than a property of a fixed architecture, which is what makes the arm contrasts of[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)causal claims about substitution rather than observations about a particular ensemble\. Second, none of this work produces a*bound*: measuring that failures correlate does not tell a practitioner what reliability they may certify, and[Sections5](https://arxiv.org/html/2608.12895#S5)and[6](https://arxiv.org/html/2608.12895#S6)supply exactly that\. Third, the statistic matters\.[McDonnell et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib38)lead withϕ\\phi, which[Definition3\.17](https://arxiv.org/html/2608.12895#S3.Thmdefinition17)classifies as marginal\-sensitive;[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)shows that such statistics can reverse the apparent ordering of two conditions purely through a difference in marginal failure rates\. That result is a caveat on our own headline numbers and, we think, a useful one for reading theirs\.
### 2\.4Bounds without independence
Bounding a joint probability from marginals is classical\. The Fréchet–Hoeffding inequalities\([Fréchet 1951](https://arxiv.org/html/2608.12895#bib.bib18);[Hoeffding 1940](https://arxiv.org/html/2608.12895#bib.bib24)\)give the sharp sandwich of[Theorem4\.1](https://arxiv.org/html/2608.12895#S4.Thmtheorem1), and Boole’s problem\([Boole 1854](https://arxiv.org/html/2608.12895#bib.bib9);[Hailperin 1965](https://arxiv.org/html/2608.12895#bib.bib21)\)recasts the search for the extremal joint as a linear program — the device[Section6](https://arxiv.org/html/2608.12895#S6)instantiates over co\-execution moments\. Copula theory\([Sklar 1959](https://arxiv.org/html/2608.12895#bib.bib55);[Nelsen 2006](https://arxiv.org/html/2608.12895#bib.bib40)\)parameterises dependence structures, and the Gaussian one\-factor model of[Remark5\.2](https://arxiv.org/html/2608.12895#S5.Thmremark2)is the standard choice in reliability practice\([Barlow and Proschan 1975](https://arxiv.org/html/2608.12895#bib.bib2)\)\.[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)is our reason for refusing to certify on it\.
The moment\-problem view\([Bertsimas and Popescu 2005](https://arxiv.org/html/2608.12895#bib.bib6)\)bounds expectations subject to moment constraints, and Bonferroni\-type inequalities\([Bonferroni 1936](https://arxiv.org/html/2608.12895#bib.bib8)\)supply the multiplicity correction[Proposition6\.2](https://arxiv.org/html/2608.12895#S6.Thmproposition2)uses\. Our contribution is not the LP, which is standard, but its use as a*finite\-sample certificate*for agent pipelines: the combination of an exact Clopper–Pearson box\([Clopper and Pearson 1934](https://arxiv.org/html/2608.12895#bib.bib12)\)with the extremal LP, and the resulting guarantee of[Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)\.
### 2\.5Sequential and anytime\-valid inference
Wald’s SPRT\([Wald 1945](https://arxiv.org/html/2608.12895#bib.bib58)\)certifies with an expected sample size far below the fixed\-nnrequirement, but at a stopping rule fixed in advance\. Game\-theoretic probability\([Shafer and Vovk 2019](https://arxiv.org/html/2608.12895#bib.bib52);[Ville 1939](https://arxiv.org/html/2608.12895#bib.bib57)\)and the recent literature on e\-values and testing by betting\([Shafer 2021](https://arxiv.org/html/2608.12895#bib.bib51);[Ramdas et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib48);[Waudby\-Smith and Ramdas 2024](https://arxiv.org/html/2608.12895#bib.bib60);[Grünwald et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib20)\), together with time\-uniform confidence sequences\([Robbins 1970](https://arxiv.org/html/2608.12895#bib.bib50);[Howard et al\. 2021](https://arxiv.org/html/2608.12895#bib.bib27)\)replace that with certificates valid under arbitrary stopping\.[Section7](https://arxiv.org/html/2608.12895#S7)applies this to graph reliability, where the payoff is specific: the null constrains only a conditional mean, so the certificate requires no independence assumption at all — the exact assumption this paper shows to be false for the composition bound\.[Proposition7\.1](https://arxiv.org/html/2608.12895#S7.Thmproposition1)records that the SPRT is recovered exactly at the optimally tuned bet, so anytime validity is obtained without loss against a known alternative\.
### 2\.6Agent evaluation
Benchmarks for agent capability\([Jimenez et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib30);[Liu et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib37);[Zhou et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib65)\)measure task success, and agent architectures\([Yao et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib62);[Shinn et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib53);[Wang et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib59);[Park et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib43);[Hong et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib26);[Qian et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib44);[Wu et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib61);[Chase 2022](https://arxiv.org/html/2608.12895#bib.bib10)\)optimise it\. None reports whether component failures are dependent, which is the quantity a composed guarantee needs; task\-success rates are marginals, and[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)shows marginals do not determine the composed outcome\.[Section10](https://arxiv.org/html/2608.12895#S10)treats cross\-agent failure dependence as the measured quantity under a preregistered manipulation of model sharing\. We know of no earlier evaluation combining the four elements that design rests on: a registered hypothesis set, a manipulated sharing condition, three graph topologies, and deterministic contract scoring with no model in the judging loop\. Correlated failure itself is not new —[Section2\.3](https://arxiv.org/html/2608.12895#S2.SS3)reports it\. What is new here is measuring it as an intervention rather than observing it, with a scoring rule that cannot itself induce correlation across arms\.
## 3Preliminaries: the ABC framework
This section states the Agent Behavioral Contract framework in full\. The framework is due to the v1 paper\([Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\), where several of these formulas were withheld under a patent claim that has since been withdrawn; every one is stated here without redaction, and[AppendixB](https://arxiv.org/html/2608.12895#A2)gives the complete catalogue with the corresponding v1 equation numbers\. Readers familiar with v1 can skip to[Section4](https://arxiv.org/html/2608.12895#S4), which is where this work departs from it\.
### 3\.1Contracts and compliance
###### Definition 3\.1\(Behavioral contract; v1 Def\. 3\.1\)\.
A*behavioral contract*is a tuple
𝒞=\(𝒫,ℐhard,ℐsoft,𝒢hard,𝒢soft,ℛ\),\\mathcal\{C\}=\(\\mathcal\{P\},\\,\\mathcal\{I\}\_\{\\mathrm\{hard\}\},\\,\\mathcal\{I\}\_\{\\mathrm\{soft\}\},\\,\\mathcal\{G\}\_\{\\mathrm\{hard\}\},\\,\\mathcal\{G\}\_\{\\mathrm\{soft\}\},\\,\\mathcal\{R\}\),where𝒫=\{p1,…,pm\}\\mathcal\{P\}=\\\{p\_\{1\},\\dots,p\_\{m\}\\\}is a finite set of*preconditions*, predicates over the initial states0s\_\{0\};ℐhard\\mathcal\{I\}\_\{\\mathrm\{hard\}\}andℐsoft\\mathcal\{I\}\_\{\\mathrm\{soft\}\}are*hard*and*soft invariants*;𝒢hard\\mathcal\{G\}\_\{\\mathrm\{hard\}\}and𝒢soft\\mathcal\{G\}\_\{\\mathrm\{soft\}\}are hard and soft*governance policies*constraining operational authority such as tool use and spending; andℛ:\(ℐsoft∪𝒢soft\)×S⇀A∗\\mathcal\{R\}:\(\\mathcal\{I\}\_\{\\mathrm\{soft\}\}\\cup\\mathcal\{G\}\_\{\\mathrm\{soft\}\}\)\\times S\\rightharpoonup A^\{\*\}is a partial*recovery map*from a violated soft constraint and a state to a corrective action sequence\. We write𝒞=\(𝒫,ℐ,𝒢,ℛ\)\\mathcal\{C\}=\(\\mathcal\{P\},\\mathcal\{I\},\\mathcal\{G\},\\mathcal\{R\}\)withℐ=ℐhard∪ℐsoft\\mathcal\{I\}=\\mathcal\{I\}\_\{\\mathrm\{hard\}\}\\cup\\mathcal\{I\}\_\{\\mathrm\{soft\}\}where the partition is not in play\.
The hard/soft partition is the load\-bearing distinction\. A single hard violation is a contract breach\. A soft violation is tolerated if recovery occurs inside a bounded window, which is what makes contracts usable against systems that are stochastic by construction\.
###### Definition 3\.2\(Constraint evaluation scores; v1 Def\. 3\.6, eqs\. 1–2\)\.
For a state–action pair\(st,at\)\(s\_\{t\},a\_\{t\}\),
Chard\(t\)\\displaystyle C\_\{\\mathrm\{hard\}\}\(t\)=\|\{c∈ℐhard∪𝒢hard:c\(st,at\)=true\}\|\|ℐhard∪𝒢hard\|,\\displaystyle=\\frac\{\\left\|\\\{c\\in\\mathcal\{I\}\_\{\\mathrm\{hard\}\}\\cup\\mathcal\{G\}\_\{\\mathrm\{hard\}\}:c\(s\_\{t\},a\_\{t\}\)=\\text\{true\}\\\}\\right\|\}\{\\left\|\\mathcal\{I\}\_\{\\mathrm\{hard\}\}\\cup\\mathcal\{G\}\_\{\\mathrm\{hard\}\}\\right\|\},\(1\)Csoft\(t\)\\displaystyle C\_\{\\mathrm\{soft\}\}\(t\)=\|\{c∈ℐsoft∪𝒢soft:c\(st,at\)=true\}\|\|ℐsoft∪𝒢soft\|\.\\displaystyle=\\frac\{\\left\|\\\{c\\in\\mathcal\{I\}\_\{\\mathrm\{soft\}\}\\cup\\mathcal\{G\}\_\{\\mathrm\{soft\}\}:c\(s\_\{t\},a\_\{t\}\)=\\text\{true\}\\\}\\right\|\}\{\\left\|\\mathcal\{I\}\_\{\\mathrm\{soft\}\}\\cup\\mathcal\{G\}\_\{\\mathrm\{soft\}\}\\right\|\}\.\(2\)Both lie in\[0,1\]\[0,1\]\. The evaluator is stateless: any turn is evaluable independently of the turns before it\.
###### Definition 3\.3\(Turn compliance\)\.
A turn is*hard\-compliant*iff it incurs zero hard violations, i\.e\.Chard\(t\)=1C\_\{\\mathrm\{hard\}\}\(t\)=1, and*soft\-compliant*iff zero soft violations remain outstanding after recovery\. The binary hard verdictht=\[Chard\(t\)=1\]h\_\{t\}=\\mathbf\{1\}\\\!\\left\[C\_\{\\mathrm\{hard\}\}\(t\)=1\\right\]is the quantity every dependence estimate in[Section10](https://arxiv.org/html/2608.12895#S10)is computed on\.
### 3\.2Drift
###### Definition 3\.4\(Composite drift score; v1 Def\. 3\.12, eqs\. 7–9\)\.
The*drift*of an agent at timettis the convex combination
D\(t\)=wcDcompliance\(t\)\+wdDdistributional\(t\),wc\+wd=1,D\(t\)=w\_\{c\}\\,D\_\{\\mathrm\{compliance\}\}\(t\)\+w\_\{d\}\\,D\_\{\\mathrm\{distributional\}\}\(t\),\\qquad w\_\{c\}\+w\_\{d\}=1,\(3\)with
Dcompliance\(t\)=1−C¯\(t\)=∑iwi\(1−σi\(t\)\)∑iwi,Ddistributional\(t\)=JSD\(Pobs\(t\)∥Pref\)\.D\_\{\\mathrm\{compliance\}\}\(t\)=1\-\\bar\{C\}\(t\)=\\frac\{\\sum\_\{i\}w\_\{i\}\\bigl\(1\-\\sigma\_\{i\}\(t\)\\bigr\)\}\{\\sum\_\{i\}w\_\{i\}\},\\qquad D\_\{\\mathrm\{distributional\}\}\(t\)=\\JSD\\bigl\(P\_\{\\mathrm\{obs\}\}\(t\)\\,\\\|\\,P\_\{\\mathrm\{ref\}\}\\bigr\)\.\(4\)Defaults arewc=0\.6w\_\{c\}=0\.6,wd=0\.4w\_\{d\}=0\.4\. Both components lie in\[0,1\]\[0,1\], soD\(t\)∈\[0,1\]D\(t\)\\in\[0,1\]\.
###### Definition 3\.5\(Distributional drift term; v1 eq\. 9\)\.
Pobs\(t\)P\_\{\\mathrm\{obs\}\}\(t\)is the empirical action distribution over a sliding window of recent actions andPrefP\_\{\\mathrm\{ref\}\}is a calibrated baseline from a compliant reference session\. WithM=12\(P\+Q\)M=\\tfrac\{1\}\{2\}\(P\+Q\),
JSD\(P∥Q\)=12DKL\(P∥M\)\+12DKL\(Q∥M\)\.\\JSD\(P\\,\\\|\\,Q\)=\\tfrac\{1\}\{2\}D\_\{\\mathrm\{KL\}\}\(P\\,\\\|\\,M\)\+\\tfrac\{1\}\{2\}D\_\{\\mathrm\{KL\}\}\(Q\\,\\\|\\,M\)\.\(5\)JSD\\JSD\([Lin 1991](https://arxiv.org/html/2608.12895#bib.bib36)\)is symmetric and bounded in\[0,1\]\[0,1\], andJSD\\sqrt\{\\JSD\}is a metric\([Endres and Schindelin 2003](https://arxiv.org/html/2608.12895#bib.bib16)\), which is why it is preferred here to raw KL divergence\.
###### Definition 3\.6\(\(p,δ,k\)\(p,\\delta,k\)\-satisfaction; v1 Def\. 3\.7, eqs\. 3–4\)\.
An agentAA*\(p,δ,k\)\(p,\\delta,k\)\-satisfies*contract𝒞\\mathcal\{C\}over session lengthTTif both hold:
ℙ\[Chard\(t\)=1∀t∈\{0,…,T\}\|P\(s0\)\]\\displaystyle\\mathbb\{P\}\\bigl\[C\_\{\\mathrm\{hard\}\}\(t\)=1\\;\\;\\forall t\\in\\\{0,\\dots,T\\\}\\,\\big\|\\,P\(s\_\{0\}\)\\bigr\]≥p,\\displaystyle\\geq p,\(6\)ℙ\[∀t:Csoft\(t\)<1−δ⟹∃t′∈\{t,…,min\(t\+k,T\)\}:Csoft\(t′\)≥1−δ\|P\(s0\)\]\\displaystyle\\mathbb\{P\}\\bigl\[\\forall t:\\;C\_\{\\mathrm\{soft\}\}\(t\)<1\-\\delta\\implies\\exists\\,t^\{\\prime\}\\in\\\{t,\\dots,\\min\(t\+k,T\)\\\}:\\;C\_\{\\mathrm\{soft\}\}\(t^\{\\prime\}\)\\geq 1\-\\delta\\,\\big\|\\,P\(s\_\{0\}\)\\bigr\]≥p\.\\displaystyle\\geq p\.\(7\)Herep∈\[0,1\]p\\in\[0,1\]is a probability threshold,δ∈\[0,1\]\\delta\\in\[0,1\]a tolerable deviation, andk∈ℕk\\in\\mathbb\{N\}a recovery window\.
###### Definition 3\.7\(Reliability index; v1 Def\. 3\.20, eq\. 13\)\.
Θ=α1C¯\(t\)\+α2\(1−D¯\(t\)\)\+α311\+E\+α4S,∑i=14αi=1,\\Theta=\\alpha\_\{1\}\\bar\{C\}\(t\)\+\\alpha\_\{2\}\\bigl\(1\-\\bar\{D\}\(t\)\\bigr\)\+\\alpha\_\{3\}\\frac\{1\}\{1\+E\}\+\\alpha\_\{4\}S,\\qquad\\textstyle\\sum\_\{i=1\}^\{4\}\\alpha\_\{i\}=1,\(8\)with defaultsα1=0\.35\\alpha\_\{1\}=0\.35\(compliance\),α2=0\.25\\alpha\_\{2\}=0\.25\(stability\),α3=0\.20\\alpha\_\{3\}=0\.20\(event frequency,EEthe count of violation events\), andα4=0\.20\\alpha\_\{4\}=0\.20\(stress resilience\)\.SSis the*stress resilience index*𝔼\[C\(t\)∣stressed\]/𝔼\[C\(t\)∣baseline\]\\mathbb\{E\}\[C\(t\)\\mid\\text\{stressed\}\]/\\mathbb\{E\}\[C\(t\)\\mid\\text\{baseline\}\]\(v1 eq\. 12\), which is a distinct quantity from the recovery success rate; descriptions that equate the two are in error\.
### 3\.3Drift dynamics
###### Definition 3\.8\(Ornstein–Uhlenbeck drift dynamics; v1 Def\. 4\.1, eq\. 14\)\.
Drift evolves as
dD=\(α−γD\(t\)\)dt\+σdW\(t\),dD=\\bigl\(\\alpha\-\\gamma D\(t\)\\bigr\)\\,dt\+\\sigma\\,dW\(t\),\(9\)withα\>0\\alpha\>0the natural drift rate,γ\>0\\gamma\>0the contract\-recovery mean\-reversion rate,σ\>0\\sigma\>0the diffusion coefficient, andWWa standard Wiener process\. TreatingD\(t\)≥0D\(t\)\\geq 0is a modelling simplification \(v1 Remark 4\.2\); the Gaussian stationary law of[Theorem3\.1](https://arxiv.org/html/2608.12895#S3.Thmtheorem1)places negligible but non\-zero mass below zero\.
###### Theorem 3\.1\(Stationary drift law and bound; v1 Thm\. 4\.3\)\.
LetDDsolve[Definition3\.8](https://arxiv.org/html/2608.12895#S3.Thmdefinition8)withα,γ,σ\>0\\alpha,\\gamma,\\sigma\>0, and let the initial condition satisfy𝔼\[D\(0\)2\]<∞\\mathbb\{E\}\[D\(0\)^\{2\}\]<\\inftywithD\(0\)D\(0\)independent of the driving Wiener process\. Then:
1. \(i\)the stationary law isπD=𝒩\(α/γ,σ2/\(2γ\)\)\\pi\_\{D\}=\\mathcal\{N\}\\\!\\bigl\(\\alpha/\\gamma,\\;\\sigma^\{2\}/\(2\\gamma\)\\bigr\);
2. \(ii\)𝔼π\[D\]=α/γ\\mathbb\{E\}\_\{\\pi\}\[D\]=\\alpha/\\gamma, so𝔼π\[D\]<1\\mathbb\{E\}\_\{\\pi\}\[D\]<1iffγ\>α\\gamma\>\\alpha;
3. \(iii\)Varπ\(D\)=σ2/\(2γ\)\\operatorname\{Var\}\_\{\\pi\}\(D\)=\\sigma^\{2\}/\(2\\gamma\);
4. \(iv\)forη\>0\\eta\>0,ℙπ\(D\>α/γ\+η\)≤exp\(−γη2/σ2\)\\mathbb\{P\}\_\{\\pi\}\\bigl\(D\>\\alpha/\\gamma\+\\eta\\bigr\)\\leq\\exp\\bigl\(\-\\gamma\\eta^\{2\}/\\sigma^\{2\}\\bigr\);
5. \(v\)writinge\(t\)=D\(t\)−α/γe\(t\)=D\(t\)\-\\alpha/\\gamma,𝔼\[e\(t\)2\]=𝔼\[e\(0\)2\]e−2γt\+σ22γ\(1−e−2γt\)\\mathbb\{E\}\[e\(t\)^\{2\}\]=\\mathbb\{E\}\[e\(0\)^\{2\}\]\\,e^\{\-2\\gamma t\}\+\\frac\{\\sigma^\{2\}\}\{2\\gamma\}\\bigl\(1\-e^\{\-2\\gamma t\}\\bigr\), so convergence to stationarity is exponential at rate2γ2\\gamma\. \(For deterministicD\(0\)D\(0\)this readse\(0\)2e−2γt\+⋯e\(0\)^\{2\}e^\{\-2\\gamma t\}\+\\cdots\.\)
###### Proof\.
See[SectionA\.1](https://arxiv.org/html/2608.12895#A1.SS1)\. ∎
Part \(ii\) is the design criterion: a contract whose recovery rateγ\\gammaexceeds the natural drift rateα\\alphaholds expected drift below11indefinitely\. Part \(iv\) converts that mean statement into the tail bound a practitioner needs, and part \(v\) says the transient decays at rate2γ2\\gamma, so the recovery rate sets both the stationary level and the time to reach it\.
### 3\.4Composition
###### Definition 3\.9\(Composability conditions; v1 Def\. 4\.7\)\.
AgentsAAandBBwith contracts𝒞A,𝒞B\\mathcal\{C\}\_\{A\},\\mathcal\{C\}\_\{B\}and a handoff invariantℐh\\mathcal\{I\}\_\{h\}are*composable*if:
C1 Interface compatibilityType\(PostA\)⊆Type\(𝒫B\)\\mathrm\{Type\}\(\\mathrm\{Post\}\_\{A\}\)\\subseteq\\mathrm\{Type\}\(\\mathcal\{P\}\_\{B\}\)\.
C2 Assumption dischargePostA∧ℐh⟹𝒫B\\mathrm\{Post\}\_\{A\}\\wedge\\mathcal\{I\}\_\{h\}\\implies\\mathcal\{P\}\_\{B\}\.
C3 Governance consistencyAllowed\(𝒢A\)∩Prohibited\(𝒢B\)=∅\\mathrm\{Allowed\}\(\\mathcal\{G\}\_\{A\}\)\\cap\\mathrm\{Prohibited\}\(\\mathcal\{G\}\_\{B\}\)=\\emptyset\.
C4 Recovery independenceany post\-recovery state ofAAstill satisfies𝒫B\\mathcal\{P\}\_\{B\}\.
###### Definition 3\.10\(Conditional independence of contract satisfaction; v1 Thm\. 4\.11 preamble\)\.
LetEAE\_\{A\},EBE\_\{B\},EhE\_\{h\}denote the events thatAAsatisfies its contract, thatBBsatisfies its contract, and that the handoff invariant holds\. ConditionC5states
ℙ\(EB∣EA∩Eh\)=ℙ\(EB∣Eh\)\.\\mathbb\{P\}\(E\_\{B\}\\mid E\_\{A\}\\cap E\_\{h\}\)=\\mathbb\{P\}\(E\_\{B\}\\mid E\_\{h\}\)\.\(10\)C1–C4 are required for the deterministic composition theorem; C5 is required*additionally*for the probabilistic one\. Treating the two sets as a single list “C1–C5” conflates the deterministic and probabilistic results\.
[Equation10](https://arxiv.org/html/2608.12895#S3.E10)is the assumption this paper is about\.[Section4](https://arxiv.org/html/2608.12895#S4)shows what fails when it does not hold, and[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)measures the extent to which it does not hold in practice\.
###### Definition 3\.11\(Naive compositional bound; v1 Thm\. 4\.11, eqs\. 28–29\)\.
Under C1–C5, for the two\-agent compositionA⊕BA\\oplus B,
pA⊕B≥pA⋅pB⋅ph,δA⊕B≤δA\+δB\+δh,p\_\{A\\oplus B\}\\;\\geq\\;p\_\{A\}\\cdot p\_\{B\}\\cdot p\_\{h\},\\qquad\\delta\_\{A\\oplus B\}\\;\\leq\\;\\delta\_\{A\}\+\\delta\_\{B\}\+\\delta\_\{h\},\(11\)extending to anNN\-agent chain \(v1 Cor\. 4\.13, eqs\. 30–31\) aspchain≥∏i=1Npi∏i=1N−1phip\_\{\\mathrm\{chain\}\}\\geq\\prod\_\{i=1\}^\{N\}p\_\{i\}\\prod\_\{i=1\}^\{N\-1\}p\_\{h\_\{i\}\}andδchain≤∑iδi\+∑iδhi\\delta\_\{\\mathrm\{chain\}\}\\leq\\sum\_\{i\}\\delta\_\{i\}\+\\sum\_\{i\}\\delta\_\{h\_\{i\}\}\.
### 3\.5Certification
###### Definition 3\.12\(Sequential probability ratio certification; v1 §7\)\.
ForH0:p≤p0H\_\{0\}:p\\leq p\_\{0\}againstH1:p≥p1H\_\{1\}:p\\geq p\_\{1\}, the SPRT accumulates the Bernoulli log\-likelihood ratio and stops at the first boundary crossing\. Its expected sample size satisfies
𝔼\[N∗\]≈log\(1/αerr\)DKL\(p^∥p0\),\\mathbb\{E\}\[N^\{\*\}\]\\;\\approx\\;\\frac\{\\log\(1/\\alpha\_\{\\mathrm\{err\}\}\)\}\{D\_\{\\mathrm\{KL\}\}\(\\hat\{p\}\\,\\\|\\,p\_\{0\}\)\},\(12\)scaling asO\(log\(1/α\)/DKL\)O\(\\log\(1/\\alpha\)/D\_\{\\mathrm\{KL\}\}\)\.
###### Definition 3\.13\(Fixed\-sample certification baseline\)\.
The fixed\-sample requirement of[Hoeffding 1963](https://arxiv.org/html/2608.12895#bib.bib25)for absolute accuracyε\\varepsilonat errorαerr\\alpha\_\{\\mathrm\{err\}\}isNH=log\(2/αerr\)/\(2ε2\)N\_\{H\}=\\log\(2/\\alpha\_\{\\mathrm\{err\}\}\)/\(2\\varepsilon^\{2\}\), scaling asO\(1/ε2\)O\(1/\\varepsilon^\{2\}\)\. Atp0=0\.85p\_\{0\}=0\.85,p1=0\.95p\_\{1\}=0\.95,α=0\.05\\alpha=0\.05the SPRT needs roughly5858sessions against Hoeffding’s184184; atε=0\.01\\varepsilon=0\.01the gap widens to about300300against 18,445\.
### 3\.6Cost
###### Proposition 3\.1\(Per\-action enforcement cost; v1 Prop\. 4\.15\)\.
Enforcement costsO\(k\+\|A\|\)O\(k\+\|A\|\)per action, wherekkis the number of constraints and\|A\|\|A\|the action\-vocabulary size:O\(k\)O\(k\)for constraint evaluation,O\(\|A\|\)O\(\|A\|\)for the JSD histogram update, andO\(1\)O\(1\)for compliance aggregation\. Fork<100k<100and\|A\|<50\|A\|<50the measured overhead is under 10 ms per action \(v1 Remark 4\.16\)\.
### 3\.7Graphs, motifs, and co\-failure
The remaining definitions are new to this paper and set up[Section4](https://arxiv.org/html/2608.12895#S4)onward\.
###### Definition 3\.14\(Agent graph and graph outcome\)\.
An*agent graph*GGis a finite DAG whose nodes are agents, each carrying a contract, and whose edges are handoffs carrying handoff invariants\. Nodeiiproduces a hard verdicthi∈\{0,1\}h\_\{i\}\\in\\\{0,1\\\}\. A*composition rule*ρ\\rhomaps the node verdicts to the graph outcomeYG=ρ\(h1,…,hm\)∈\{0,1\}Y\_\{G\}=\\rho\(h\_\{1\},\\dots,h\_\{m\}\)\\in\\\{0,1\\\}\. We writepi=ℙ\(hi=1\)p\_\{i\}=\\mathbb\{P\}\(h\_\{i\}=1\)\.
###### Definition 3\.15\(Motifs\)\.
Two composition rules are used throughout:*series*ρser=⋀ihi\\rho\_\{\\mathrm\{ser\}\}=\\bigwedge\_\{i\}h\_\{i\}over a linear handoff chain, and*quorum*ρq/m=\[∑ihi≥q\]\\rho\_\{q/m\}=\\mathbf\{1\}\\\!\\left\[\\sum\_\{i\}h\_\{i\}\\geq q\\right\]overmmworkers feeding a deterministicqq\-of\-mmaggregator\. The parallel motif is theq=1q=1case: two independent branches feed a deterministic merge andρpar=\[∑ihi≥1\]\\rho\_\{\\mathrm\{par\}\}=\\mathbf\{1\}\\\!\\left\[\\sum\_\{i\}h\_\{i\}\\geq 1\\right\], so the graph succeeds when at least one branch passes and the merge is valid\. Parallel is therefore a redundancy rule, not a conjunction; the distinction matters because[Corollary4\.1](https://arxiv.org/html/2608.12895#S4.Thmcorollary1)shows the two rules put positive dependence on*opposite*sides of the safety argument\. Deterministic merge and aggregator nodes carry no contract and are excluded from dependence estimation\.
###### Definition 3\.16\(Co\-failure table and sharing condition\)\.
For nodesi,ji,jwith failure indicatorsFi=1−hiF\_\{i\}=1\-h\_\{i\}, the*co\-failure table*overnnmissions has cellsn11=∑\[FiFj=1\]n\_\{11\}=\\sum\\mathbf\{1\}\\\!\\left\[F\_\{i\}F\_\{j\}=1\\right\],n10=∑\[Fi\(1−Fj\)=1\]n\_\{10\}=\\sum\\mathbf\{1\}\\\!\\left\[F\_\{i\}\(1\-F\_\{j\}\)=1\\right\],n01=∑\[\(1−Fi\)Fj=1\]n\_\{01\}=\\sum\\mathbf\{1\}\\\!\\left\[\(1\-F\_\{i\}\)F\_\{j\}=1\\right\],n00=∑\[\(1−Fi\)\(1−Fj\)=1\]n\_\{00\}=\\sum\\mathbf\{1\}\\\!\\left\[\(1\-F\_\{i\}\)\(1\-F\_\{j\}\)=1\\right\], with proportionspab=nab/np\_\{ab\}=n\_\{ab\}/nand marginal failure ratespi=p11\+p10p\_\{i\}=p\_\{11\}\+p\_\{10\},pj=p11\+p01p\_\{j\}=p\_\{11\}\+p\_\{01\}\. The*sharing condition*of the pair issame\_modelwhen both nodes run the same model weights,same\_vendorwhen they run different models from one vendor, anddifferent\_vendorotherwise\.
###### Definition 3\.17\(Dependence functionals\)\.
On a co\-failure table we use the*overlap*J=n11/\(n11\+n10\+n01\)J=n\_\{11\}/\(n\_\{11\}\+n\_\{10\}\+n\_\{01\}\), the binary Kendallτa=2\(p11p00−p10p01\)\\tau\_\{a\}=2\(p\_\{11\}p\_\{00\}\-p\_\{10\}p\_\{01\}\), the phi coefficientϕ=\(p11p00−p10p01\)/pi\(1−pi\)pj\(1−pj\)\\phi=\(p\_\{11\}p\_\{00\}\-p\_\{10\}p\_\{01\}\)/\\sqrt\{p\_\{i\}\(1\-p\_\{i\}\)p\_\{j\}\(1\-p\_\{j\}\)\}, the log odds ratio, and Yule’sQQ\. We callJJ,τa\\tau\_\{a\},ϕ\\phi*marginal\-sensitive*and the last two*marginal\-free*: the latter are invariant under row and column scaling of the table, the former are not\.[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)shows the distinction is decisive rather than cosmetic\.
## 4What breaks when independence fails
[Definition3\.11](https://arxiv.org/html/2608.12895#S3.Thmdefinition11)certifies a composed system by multiplying component reliabilities\. That step is licensed by C5 \([Definition3\.10](https://arxiv.org/html/2608.12895#S3.Thmdefinition10)\) and by nothing else\. This section establishes three things: that the multiplicative bound is not conservative when C5 fails, that the assumption\-free alternative is too weak to be useful, and that a certificate built on a*fitted*dependence model can lose coverage of the truth entirely, and that collecting more data does not repair it\.
### 4\.1The bound is not conservative
###### Proposition 4\.1\(Signed compositional gap\)\.
LetGGbe a series motif over components with hard verdictsh1,…,hmh\_\{1\},\\dots,h\_\{m\}andYG=⋀ihiY\_\{G\}=\\bigwedge\_\{i\}h\_\{i\}\. Then
ℙ\(YG=1\)−∏i=1mpi=Cov\(h1,∏i=2mhi\)\+∑k=2m−1\(∏i=1k−1pi\)Cov\(hk,∏i=k\+1mhi\)\.\\mathbb\{P\}\(Y\_\{G\}=1\)\-\\prod\_\{i=1\}^\{m\}p\_\{i\}\\;=\\;\\operatorname\{Cov\}\\Bigl\(h\_\{1\},\\prod\_\{i=2\}^\{m\}h\_\{i\}\\Bigr\)\+\\sum\_\{k=2\}^\{m\-1\}\\Bigl\(\\prod\_\{i=1\}^\{k\-1\}p\_\{i\}\\Bigr\)\\operatorname\{Cov\}\\Bigl\(h\_\{k\},\\prod\_\{i=k\+1\}^\{m\}h\_\{i\}\\Bigr\)\.\(13\)In particular form=2m=2the gap is exactlyCov\(h1,h2\)\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\), so under positive dependence the independence productunderstatesseries reliability — and correspondingly*overstates*the probability that the series fails at all\. The unsafe direction is a different estimand, stated separately in[Corollary4\.1](https://arxiv.org/html/2608.12895#S4.Thmcorollary1)\.
###### Proof\.
See[SectionA\.3](https://arxiv.org/html/2608.12895#A1.SS3)\. ∎
For a*series*system, then, the independence product is genuinely conservative and dependence only helps\. The unsafe case is redundancy, and it is a distinct statement:
###### Corollary 4\.1\(Redundant joint failure is understated\)\.
LetFi=1−hiF\_\{i\}=1\-h\_\{i\}\. For two components,ℙ\(F1=F2=1\)=\(1−p1\)\(1−p2\)\+Cov\(F1,F2\)\\mathbb\{P\}\(F\_\{1\}=F\_\{2\}=1\)=\(1\-p\_\{1\}\)\(1\-p\_\{2\}\)\+\\operatorname\{Cov\}\(F\_\{1\},F\_\{2\}\)andCov\(F1,F2\)=Cov\(h1,h2\)\\operatorname\{Cov\}\(F\_\{1\},F\_\{2\}\)=\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\)\. Under positive dependence the probability that*both*redundant paths fail therefore exceeds the independence product, so a redundant design is over\-credited exactly when its components are positively associated\.
###### Proof\.
See[SectionA\.4](https://arxiv.org/html/2608.12895#A1.SS4)\. ∎
This is the direction that matters operationally: a quorum or a reviewer–writer pair is deployed precisely so thatℙ\(all redundant paths fail\)\\mathbb\{P\}\(\\text\{all redundant paths fail\}\)is small\.[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)measuresϕ\\phibetween0\.580\.58and0\.920\.92; at those magnitudes the independent calculation is not close\.
###### Example 4\.1\.
Take a 2\-of\-3 quorum whose workers each satisfy their contract withp=0\.61p=0\.61, thesame\_modelrate measured in[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)\. Under independence the quorum succeeds with probability3p2\(1−p\)\+p3=0\.66233p^\{2\}\(1\-p\)\+p^\{3\}=0\.6623and therefore fails with probability0\.33770\.3377\. An exchangeable joint consistent with the same marginals and the measured pairwise association puts quorum failure near0\.370\.37, a relative increase of about9\.6%9\.6\\%in the quantity the redundancy was bought to reduce\. We stress that marginals andϕ\\phialone do not determine a 2\-of\-3 failure probability — the triple moment is free — so this is an illustration of magnitude under one consistent joint, not a bound\.[Section6](https://arxiv.org/html/2608.12895#S6)is the machinery for the bound\.
### 4\.2The assumption\-free alternative is too weak
Dropping C5 entirely leaves the Fréchet–Hoeffding bounds, which hold for every joint law with the given marginals\.
###### Theorem 4\.1\(Fréchet–Hoeffding sandwich for all\-success\)\.
For any joint law of\(h1,…,hm\)\(h\_\{1\},\\dots,h\_\{m\}\)with marginalsp1,…,pmp\_\{1\},\\dots,p\_\{m\},
max\(0,∑i=1mpi−\(m−1\)\)≤ℙ\(⋀ihi=1\)≤minipi,\\max\\Bigl\(0,\\;\\sum\_\{i=1\}^\{m\}p\_\{i\}\-\(m\-1\)\\Bigr\)\\;\\leq\\;\\mathbb\{P\}\\Bigl\(\\bigwedge\_\{i\}h\_\{i\}=1\\Bigr\)\\;\\leq\\;\\min\_\{i\}p\_\{i\},\(14\)and both bounds are attained, so neither can be improved without further information\.
###### Proof\.
See[SectionA\.5](https://arxiv.org/html/2608.12895#A1.SS5)\. ∎
###### Corollary 4\.2\(Vacuity of the assumption\-free floor\)\.
The lower bound of[Theorem4\.1](https://arxiv.org/html/2608.12895#S4.Thmtheorem1)is00whenever∑ipi≤m−1\\sum\_\{i\}p\_\{i\}\\leq m\-1, i\.e\. whenever the mean component reliability is at most1−1/m1\-1/m\. Form=4m=4components each atp=0\.75p=0\.75the certified floor is exactly00\.
[Corollary4\.2](https://arxiv.org/html/2608.12895#S4.Thmcorollary2)is the practical bind\. C5 gives a number that is wrong in the unsafe direction; dropping C5 gives a number that is right and useless\. The moment\-set certificate of[Section6](https://arxiv.org/html/2608.12895#S6)occupies the space between them by constraining the joint with measured co\-execution moments rather than with an independence assumption, and[Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3)measures what that buys: a floor of0\.41160\.4116where the pairwise\-only certificate gives0\.24550\.2455\.
### 4\.3Fitting a dependence model does not rescue the situation
A natural response is to fit a parametric dependence model — a one\-factor Gaussian copula, say — estimate its parameters, and report a confidence bound on the fitted model’s all\-success functional\. This fails, and it fails in a way that gets worse with more data\.
###### Definition 4\.1\(Model functional and identified set\)\.
Fix a moment family𝒥\\mathcal\{J\}and letℳ\(μ^\)\\mathcal\{M\}\(\\hat\{\\mu\}\)be the set of joint laws on\{0,1\}m\\\{0,1\\\}^\{m\}whose𝒥\\mathcal\{J\}\-moments equalμ^\\hat\{\\mu\}\. The*identified set*for all\-success is\[R¯\(μ^\),R¯\(μ^\)\]\\bigl\[\\underline\{R\}\(\\hat\{\\mu\}\),\\,\\overline\{R\}\(\\hat\{\\mu\}\)\\bigr\]whereR¯=infQ∈ℳQ\(⋀ihi=1\)\\underline\{R\}=\\inf\_\{Q\\in\\mathcal\{M\}\}Q\(\\bigwedge\_\{i\}h\_\{i\}=1\)andR¯=sup\\overline\{R\}=\\sup\. A parametric familyℱ⊂ℳ\\mathcal\{F\}\\subset\\mathcal\{M\}induces a*model functional*Rℱ\(μ^\)R\_\{\\mathcal\{F\}\}\(\\hat\{\\mu\}\), the all\-success probability of the fitted member ofℱ\\mathcal\{F\}\.
###### Theorem 4\.2\(Coverage collapse of a model\-based floor\)\.
LetQ⋆Q^\{\\star\}be a true joint law with𝒥\\mathcal\{J\}\-momentsμ⋆\\mu^\{\\star\}and true all\-success probabilityR⋆=Q⋆\(⋀ihi=1\)R^\{\\star\}=Q^\{\\star\}\(\\bigwedge\_\{i\}h\_\{i\}=1\)\. Suppose the parametric familyℱ\\mathcal\{F\}is*misspecified atμ⋆\\mu^\{\\star\}*in the sense that
Δ:=Rℱ\(μ⋆\)−R⋆\>0\.\\Delta\\;:=\\;R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)\-R^\{\\star\}\\;\>\\;0\.\(15\)LetL^n\\hat\{L\}\_\{n\}be any bootstrap lower confidence bound forRℱ\(μ⋆\)R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)satisfyingL^n=Rℱ\(μ⋆\)−Op\(n−1/2\)\\hat\{L\}\_\{n\}=R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)\-O\_\{p\}\(n^\{\-1/2\}\)\. Then
ℙ\(L^n≤R⋆\)⟶0asn→∞\.\\mathbb\{P\}\\bigl\(\\hat\{L\}\_\{n\}\\leq R^\{\\star\}\\bigr\)\\;\\longrightarrow\\;0\\qquad\\text\{as \}n\\to\\infty\.\(16\)That is, coverage of the*true*all\-success probability tends to zero\. More data does not repair the certificate: the haircut shrinks while the target does not move\. We claim the limit and not monotonicity innn, which tightness alone does not entail\.
###### Proof\.
See[SectionA\.6](https://arxiv.org/html/2608.12895#A1.SS6)\. ∎
The mechanism is a mismatch of orders\. The identification gapΔ\\Deltain[Equation15](https://arxiv.org/html/2608.12895#S4.E15)isO\(1\)O\(1\)— it is a property of the family, not of the sample — while the bootstrap haircut shrinks liken−1/2n^\{\-1/2\}\. Past some finitennthe entire interval sits aboveR⋆R^\{\\star\}and never returns\. The bootstrap is not malfunctioning; it is covering the estimand it was asked about, which is the fitted model’s functional and not the truth\.
###### Example 4\.2\(An explicit witness\)\.
Take the equicorrelated Gaussian\-copula joint atp=0\.6p=0\.6,λ=0\.8\\lambda=0\.8\(a factor*loading*, so the latent correlation isλ2=0\.64\\lambda^\{2\}=0\.64\),m=3m=3\. By Gauss–Legendre quadrature over the common factor its marginals are0\.6000000\.600000, its pairwise co\-success moments are0\.4652370\.465237, and its all\-success probability is0\.3921440\.392144— so the one\-factor family is*correctly specified*for it andΔ=0\\Delta=0\. Every constant quoted for this law in this paper comes from that one deterministic routine, not from simulation\. It is thereforenotitself a witness for[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2), a point worth stating because it is the natural law to reach for and it does not work\.
The witness is instead the joint attaining the lower end of the LP\-identified set\. Over the pairwise moment family that set is\[0\.330475,0\.465237\]\[0\.330475,0\.465237\], and by[Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)the lower endpoint is attained\. For this exchangeable case it has the closed form2q−p2q\-pin the pairwise momentqqand marginalpp, and the attaining lawQ⋆Q^\{\\star\}puts mass on five cells only:
Q⋆\(1,1,1\)\\displaystyle Q^\{\\star\}\(1,1,1\)=0\.330475,\\displaystyle=0\.330475,Q⋆\(0,0,0\)\\displaystyle Q^\{\\star\}\(0,0,0\)=0\.265237,\\displaystyle=0\.265237,Q⋆\(0,1,1\)=Q⋆\(1,0,1\)=Q⋆\(1,1,0\)\\displaystyle Q^\{\\star\}\(0,1,1\)=Q^\{\\star\}\(1,0,1\)=Q^\{\\star\}\(1,1,0\)=0\.134763,\\displaystyle=0\.134763,all other cells=0\.\\displaystyle=0\.By constructionQ⋆Q^\{\\star\}has*exactly*the same marginals and*exactly*the same pairwise co\-success moments as the Gaussian joint: the two are indistinguishable to any procedure that sees only those moments\. Yet
Δ=Rℱ\(μ⋆\)−R⋆=0\.392144−0\.330475=0\.061670\>0,\\Delta\\;=\\;R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)\-R^\{\\star\}\\;=\\;0\.392144\-0\.330475\\;=\\;0\.061670\\;\>\\;0,so[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)applies toQ⋆Q^\{\\star\}\.[Section10\.4](https://arxiv.org/html/2608.12895#S10.SS4)samples fromQ⋆Q^\{\\star\}and measures the collapse, and runs the Gaussian joint as a correctly\-specified control in which no collapse should occur — and does not\.
###### Corollary 4\.3\(Why the copula\-agnostic floor is the right target\)\.
The LP floorR¯\(μ^\)\\underline\{R\}\(\\hat\{\\mu\}\)of[Definition4\.1](https://arxiv.org/html/2608.12895#S4.Thmdefinition1)satisfiesR¯\(μ⋆\)≤R⋆\\underline\{R\}\(\\mu^\{\\star\}\)\\leq R^\{\\star\}by construction for everyQ⋆∈ℳ\(μ⋆\)Q^\{\\star\}\\in\\mathcal\{M\}\(\\mu^\{\\star\}\), with no parametric assumption\. A confidence bound targetingR¯\\underline\{R\}therefore cannot suffer the failure of[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2); its conservatism is the price of that immunity, and[Section6](https://arxiv.org/html/2608.12895#S6)shows the price falls as the moment family grows\.
## 5A tiered certificate
[Section4](https://arxiv.org/html/2608.12895#S4)leaves three unusable options: a multiplicative bound that is unsafe on the failure side, an assumption\-free bound that is often exactly zero, and a fitted model whose coverage collapses\. What survives depends on*what was actually measured*\. A certificate should therefore not be a single number but a selection among tiers, each stating the evidence it requires and the assumptions it carries\.
###### Definition 5\.1\(Certification scope\)\.
A*certificate*is a triple\(L^,𝒜,𝒮\)\(\\hat\{L\},\\mathcal\{A\},\\mathcal\{S\}\)whereL^\\hat\{L\}is a lower confidence bound on all\-success reliability at level1−ηconf1\-\\eta\_\{\\mathrm\{conf\}\},𝒜\\mathcal\{A\}is the set of assumptions it rests on, and𝒮\\mathcal\{S\}is its scope: the mission distribution, the model versions, and the topology under which it was obtained\. A certificate makes no claim outside𝒮\\mathcal\{S\}; in particular it does not transfer across mission mixes, model upgrades, or topology changes\.
###### Definition 5\.2\(The three tiers\)\.
Let𝐡∈\{0,1\}m×n\\mathbf\{h\}\\in\\\{0,1\\\}^\{m\\times n\}be a pass matrix overmmstages andnnmissions\.
Tier 0 — observedAvailable only when the composition was executed end\-to\-end, i\.e\. every stage was scored on the same missions, soYGY\_\{G\}is directly observed\.L^0\\hat\{L\}\_\{0\}is the exact Clopper–Pearson lower bound onℙ\(YG=1\)\\mathbb\{P\}\(Y\_\{G\}=1\)from∑r\[YG\(r\)=1\]\\sum\_\{r\}\\mathbf\{1\}\\\!\\left\[Y\_\{G\}^\{\(r\)\}=1\\right\]successes innntrials\. Assumptions: i\.i\.d\. missions\. No copula, no model\.
Tier 1 — copula\-agnosticAvailable from per\-stage and co\-execution data alone, when the composed pipeline was*not*run end\-to\-end\.L^1\\hat\{L\}\_\{1\}is the LP minimum of[Section6](https://arxiv.org/html/2608.12895#S6)over the Bonferroni–Clopper–Pearson moment box\. Assumptions: i\.i\.d\. missions\. No copula, no model\.
Tier 2 — modelThe finite\-sample Gaussian one\-factor floor, via the[Slepian 1962](https://arxiv.org/html/2608.12895#bib.bib56)corner\. Assumptions: i\.i\.d\. missions*and*correct specification of the one\-factor family\.
###### Theorem 5\.1\(Tier\-0 validity\)\.
If missions are i\.i\.d\. andYGY\_\{G\}is observed on each, then the Clopper–Pearson lower boundL^0\\hat\{L\}\_\{0\}satisfiesℙ\(L^0≤ℙ\(YG=1\)\)≥1−ηconf\\mathbb\{P\}\\bigl\(\\hat\{L\}\_\{0\}\\leq\\mathbb\{P\}\(Y\_\{G\}=1\)\\bigr\)\\geq 1\-\\eta\_\{\\mathrm\{conf\}\}exactly, for everynnand every true reliability\.
###### Proof\.
See[SectionA\.7](https://arxiv.org/html/2608.12895#A1.SS7)\. ∎
###### Theorem 5\.2\(Tier\-1 validity\)\.
Let𝒥\\mathcal\{J\}be a moment family of sizeJJ, letμ^=\(μ^S\)S∈𝒥\\hat\{\\mu\}=\(\\hat\{\\mu\}\_\{S\}\)\_\{S\\in\\mathcal\{J\}\}be the empirical moments, and letB\(μ^\)B\(\\hat\{\\mu\}\)be the box formed by two\-sided Clopper–Pearson intervals for eachμS\\mu\_\{S\}at levelηconf/\(2J\)\\eta\_\{\\mathrm\{conf\}\}/\(2J\)per tail\. Define
L^1=min\{Q\(⋀ihi=1\):Q∈Δ\(\{0,1\}m\),\(Q\(⋀i∈Shi=1\)\)S∈𝒥∈B\(μ^\)\}\.\\hat\{L\}\_\{1\}\\;=\\;\\min\\Bigl\\\{\\,Q\\bigl\(\\textstyle\\bigwedge\_\{i\}h\_\{i\}=1\\bigr\)\\;:\\;Q\\in\\Delta\(\\\{0,1\\\}^\{m\}\),\\;\\bigl\(Q\(\\textstyle\\bigwedge\_\{i\\in S\}h\_\{i\}=1\)\\bigr\)\_\{S\\in\\mathcal\{J\}\}\\in B\(\\hat\{\\mu\}\)\\,\\Bigr\\\}\.\(17\)Then for i\.i\.d\. missions and any true lawQ⋆Q^\{\\star\},ℙ\(L^1≤Q⋆\(⋀ihi=1\)\)≥1−ηconf\\mathbb\{P\}\\bigl\(\\hat\{L\}\_\{1\}\\leq Q^\{\\star\}\(\\bigwedge\_\{i\}h\_\{i\}=1\)\\bigr\)\\geq 1\-\\eta\_\{\\mathrm\{conf\}\}\.
###### Proof\.
See[SectionA\.8](https://arxiv.org/html/2608.12895#A1.SS8)\. ∎
[Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)is the reason Tier 1 is safe where Tier 2 is not: it is a minimum over*every*joint law consistent with the box, so the true law is one of the candidates whenever the box covers, and the union bound makes the box cover with the stated probability\. Nothing is assumed about the dependence structure\.
###### Proposition 5\.1\(Fail\-safe tier selection\)\.
Let the certificate select Tier 0 only on an explicit assertion that the composition was executed end\-to\-end, and Tier 1 by default otherwise\. Then no the two failure directions are not symmetric\. Tier 1 requires only i\.i\.d\. missions \([Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)\), which Tier 0 requires as well, so the assumption sets are nested\. Consequently:
1. \(i\)Omission is safe\.Failing to set the flag when it was warranted yieldsL^1\\hat\{L\}\_\{1\}, which is valid — merely looser than the Tier\-0 bound that was available\.
2. \(ii\)Commission is not\.Setting the flag when it was*not*warranted yieldsL^0\\hat\{L\}\_\{0\}, whose validity requires an assumption that does not hold; that guarantee is unsound\.
The design content of the proposition is therefore that the unsound case requires a deliberate assertion, while the default — and every omission — lands in the safe case\. It isnota claim that no setting of the flag can produce an unsound certificate; case \(ii\) is exactly such a setting, and no inspection of the pass matrix can detect it\.
###### Proof\.
See[SectionA\.9](https://arxiv.org/html/2608.12895#A1.SS9)\. ∎
[Proposition5\.1](https://arxiv.org/html/2608.12895#S5.Thmproposition1)states a design constraint rather than a mathematical discovery, and it is stated because the alternative is a silent failure\. Tier 0’s validity requires that the joint outcome was observed — a fact about how the data were produced, which no inspection of the pass matrix can confirm or refute\. A default of Tier 0 would let a forgotten flag convert an extrapolation into an over\-claim with no detectable symptom\. Defaulting to Tier 1 makes the failure mode a needlessly weak certificate instead\.
## 6Moment\-set certification
Tier 1 turns on a linear program\. This section states it, proves it sharp, and characterises how it behaves as the moment family grows — which is the knob a practitioner actually turns, and the one[Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3)measures\.
### 6\.1The program
###### Definition 6\.1\(Cell representation\)\.
A joint law on\{0,1\}m\\\{0,1\\\}^\{m\}is a vectorx∈ℝ2mx\\in\\mathbb\{R\}^\{2^\{m\}\}withxc≥0x\_\{c\}\\geq 0and∑cxc=1\\sum\_\{c\}x\_\{c\}=1, where cellccranges over binary patterns\. ForS⊆\{1,…,m\}S\\subseteq\\\{1,\\dots,m\\\}letaS∈\{0,1\}2ma\_\{S\}\\in\\\{0,1\\\}^\{2^\{m\}\}be the indicator row\(aS\)c=\[ci=1∀i∈S\]\(a\_\{S\}\)\_\{c\}=\\mathbf\{1\}\\\!\\left\[c\_\{i\}=1\\;\\forall i\\in S\\right\], soaS⊤xa\_\{S\}^\{\\top\}xis the probability that every stage inSSsucceeds\. All\-success isa\{1\.\.m\}⊤xa\_\{\\\{1\.\.m\\\}\}^\{\\top\}x\.
###### Definition 6\.2\(Moment\-set LP\)\.
Given a moment family𝒥\\mathcal\{J\}and values\(νS\)S∈𝒥\(\\nu\_\{S\}\)\_\{S\\in\\mathcal\{J\}\},
R¯=minxa\{1\.\.m\}⊤xs\.t\.aS⊤x=νS∀S∈𝒥,𝟏⊤x=1,x≥0,\\underline\{R\}\\;=\\;\\min\_\{x\}\\;a\_\{\\\{1\.\.m\\\}\}^\{\\top\}x\\quad\\text\{s\.t\.\}\\quad a\_\{S\}^\{\\top\}x=\\nu\_\{S\}\\;\\;\\forall S\\in\\mathcal\{J\},\\quad\\mathbf\{1\}^\{\\top\}x=1,\\quad x\\geq 0,\(18\)andR¯\\overline\{R\}the corresponding maximum\. With𝒥=\{singletons\}∪\{pairs\}\\mathcal\{J\}=\\\{\\text\{singletons\}\\\}\\cup\\\{\\text\{pairs\}\\\}this is the pairwise program; adding triples gives theJ=14J=14program of[Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3)atm=4m=4\. For a box constraintν∈B\\nu\\in Brather than point values, replace each equality by the pair of inequalities definingBB\.
###### Theorem 6\.1\(Soundness and sharpness\)\.
LetQ⋆Q^\{\\star\}be any joint law on\{0,1\}m\\\{0,1\\\}^\{m\}whose𝒥\\mathcal\{J\}\-moments equalν\\nu\. ThenR¯≤Q⋆\(⋀ihi=1\)≤R¯\\underline\{R\}\\leq Q^\{\\star\}\(\\bigwedge\_\{i\}h\_\{i\}=1\)\\leq\\overline\{R\}, and both bounds are*attained*: there exist lawsQ−,Q\+Q\_\{\-\},Q\_\{\+\}with the same𝒥\\mathcal\{J\}\-moments achievingR¯\\underline\{R\}andR¯\\overline\{R\}respectively\. Consequently\[R¯,R¯\]\[\\underline\{R\},\\overline\{R\}\]is exactly the identified set of[Definition4\.1](https://arxiv.org/html/2608.12895#S4.Thmdefinition1), and no bound using only the moments in𝒥\\mathcal\{J\}can be tighter\.
###### Proof\.
See[SectionA\.10](https://arxiv.org/html/2608.12895#A1.SS10)\. ∎
Sharpness matters for interpretation\. A conservative\-but\-loose bound invites the response “the truth is surely much better than that”;[Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)says that for the information supplied there is a consistent world in which the truth is exactlyR¯\\underline\{R\}\. Tightening requires more moments, not more optimism\.
###### Proposition 6\.1\(Monotonicity in the moment family\)\.
If𝒥⊆𝒥′\\mathcal\{J\}\\subseteq\\mathcal\{J\}^\{\\prime\}and the moment values agree on𝒥\\mathcal\{J\}, thenR¯\(𝒥\)≤R¯\(𝒥′\)\\underline\{R\}\(\\mathcal\{J\}\)\\leq\\underline\{R\}\(\\mathcal\{J\}^\{\\prime\}\)andR¯\(𝒥\)≥R¯\(𝒥′\)\\overline\{R\}\(\\mathcal\{J\}\)\\geq\\overline\{R\}\(\\mathcal\{J\}^\{\\prime\}\)\. The identified set shrinks weakly as the family grows\.
###### Proof\.
See[SectionA\.11](https://arxiv.org/html/2608.12895#A1.SS11)\. ∎
[Proposition6\.1](https://arxiv.org/html/2608.12895#S6.Thmproposition1)holds at the level of*point*moments\. It does not automatically survive the finite\-sample box, and the reason is a Bonferroni budget rather than anything deep\.
###### Proposition 6\.2\(Bonferroni allocation and monotonicity\)\.
Fixηconf\\eta\_\{\\mathrm\{conf\}\}and let the boxBBbe built fromJ=\|𝒥\|J=\|\\mathcal\{J\}\|Clopper–Pearson intervals\.
1. \(i\)*Used\-set allocation*: spendingηconf/\(2J\)\\eta\_\{\\mathrm\{conf\}\}/\(2J\)per tail gives the tightest box for that family, but enlarging𝒥\\mathcal\{J\}widens every interval, so[Proposition6\.1](https://arxiv.org/html/2608.12895#S6.Thmproposition1)can fail on the certified floor\.
2. \(ii\)*Pre\-allocated allocation*: fixing a budget family𝒥max⊇𝒥\\mathcal\{J\}\_\{\\max\}\\supseteq\\mathcal\{J\}and spendingηconf/\(2\|𝒥max\|\)\\eta\_\{\\mathrm\{conf\}\}/\(2\|\\mathcal\{J\}\_\{\\max\}\|\)per tail regardless of which moments are constrained makes every interval width independent of\|𝒥\|\|\\mathcal\{J\}\|\. Then enriching𝒥\\mathcal\{J\}only adds constraints, and the certified floor is monotone*by construction*\.
Both are valid at level1−ηconf1\-\\eta\_\{\\mathrm\{conf\}\}\.
###### Proof\.
See[SectionA\.12](https://arxiv.org/html/2608.12895#A1.SS12)\. ∎
The trade\-off is measurable rather than theoretical\. Atm=4m=4,n=717n=717in[Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3), pre\-allocation costs0\.00980\.0098of certified floor atJ=10J=10\(0\.23570\.2357against0\.24550\.2455\) and costs nothing atJ=14J=14, where the budget family is the family in use\. We report both, because reporting only the used\-set number would present an empirically\-monotone\-so\-far quantity as a guaranteed one\.
## 7Anytime\-valid certification
Every certificate so far is fixed\-nn: valid at a sample size chosen before the data arrive\. Deployed agent systems are not monitored that way\. A team watches a reliability dashboard, and when the number looks good enough they ship\. That is optional stopping, and it inflates type\-I error without bound for a fixed\-nninterval\. This section gives a certificate that survives it, at a cost this section also states\. The machinery is that of game\-theoretic probability and testing by betting\([Ville 1939](https://arxiv.org/html/2608.12895#bib.bib57);[Shafer and Vovk 2019](https://arxiv.org/html/2608.12895#bib.bib52);[Shafer 2021](https://arxiv.org/html/2608.12895#bib.bib51);[Ramdas et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib48);[Waudby\-Smith and Ramdas 2024](https://arxiv.org/html/2608.12895#bib.bib60);[Grünwald et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib20)\); the application to composed agent reliability is what is new here\.
### 7\.1The e\-process
###### Definition 7\.1\(Betting e\-process for graph reliability\)\.
Letyr∈\{0,1\}y\_\{r\}\\in\\\{0,1\\\}be the graph outcome of missionrrand letℱr\\mathcal\{F\}\_\{r\}be theσ\\sigma\-field generated byy1,…,yry\_\{1\},\\dots,y\_\{r\}\. Test the*sequential*null
H0:𝔼\[yr∣ℱr−1\]≤p0almost surely, for everyr≥1\.H\_\{0\}:\\;\\mathbb\{E\}\[y\_\{r\}\\mid\\mathcal\{F\}\_\{r\-1\}\]\\leq p\_\{0\}\\quad\\text\{almost surely, for every \}r\\geq 1\.\(19\)This is the null the betting construction requires, and it is strictly stronger than the marginal statementℙ\(YG=1\)≤p0\\mathbb\{P\}\(Y\_\{G\}=1\)\\leq p\_\{0\}\. The distinction is not pedantic: a stream whose every mission has marginal success ratep0p\_\{0\}but whose outcomes are perfectly dependent satisfies the marginal statement, violates[Equation19](https://arxiv.org/html/2608.12895#S7.E19), and drives the wealth process out of the supermartingale class entirely\.111Takep0=12p\_\{0\}=\\tfrac\{1\}\{2\},λr≡1\\lambda\_\{r\}\\equiv 1, andyr≡Zy\_\{r\}\\equiv Zfor everyrrwithZ∼Bernoulli\(12\)Z\\sim\\mathrm\{Bernoulli\}\(\\tfrac\{1\}\{2\}\)\. Every marginal satisfiesℙ\(yr=1\)=12≤p0\\mathbb\{P\}\(y\_\{r\}=1\)=\\tfrac\{1\}\{2\}\\leq p\_\{0\}, yet𝔼\[E2\]=54\>1\\mathbb\{E\}\[E\_\{2\}\]=\\tfrac\{5\}\{4\}\>1andE8=\(3/2\)8\>20=1/αE\_\{8\}=\(3/2\)^\{8\}\>20=1/\\alphaon the event\{Z=1\}\\\{Z=1\\\}, so the certificate fires with probability12\\tfrac\{1\}\{2\}rather thanα=0\.05\\alpha=0\.05\.Letλr\\lambda\_\{r\}be a*predictable*betting fraction, i\.e\.λr\\lambda\_\{r\}isℱr−1\\mathcal\{F\}\_\{r\-1\}\-measurable, withλr∈\[0,1/p0\)\\lambda\_\{r\}\\in\[0,1/p\_\{0\}\)\. The wealth process is
ER=∏r≤R\(1\+λr\(yr−p0\)\),E0=1\.E\_\{R\}\\;=\\;\\prod\_\{r\\leq R\}\\bigl\(1\+\\lambda\_\{r\}\(y\_\{r\}\-p\_\{0\}\)\\bigr\),\\qquad E\_\{0\}=1\.\(20\)The certificate*fires*at the firstRRwithER≥1/αE\_\{R\}\\geq 1/\\alpha\.
###### Lemma 7\.1\(Supermartingale property\)\.
UnderH0H\_\{0\},\(ER\)R≥0\(E\_\{R\}\)\_\{R\\geq 0\}is a non\-negative supermartingale with𝔼\[ER\]≤1\\mathbb\{E\}\[E\_\{R\}\]\\leq 1for allRR\.
###### Proof\.
See[SectionA\.13](https://arxiv.org/html/2608.12895#A1.SS13)\. ∎
###### Theorem 7\.1\(Anytime validity;[Ville 1939](https://arxiv.org/html/2608.12895#bib.bib57)\)\.
UnderH0H\_\{0\}, for any stopping timeTT— including data\-dependent and unbounded ones —
ℙ\(supR≥1ER≥1/α\)≤α\.\\mathbb\{P\}\\Bigl\(\\sup\_\{R\\geq 1\}E\_\{R\}\\geq 1/\\alpha\\Bigr\)\\;\\leq\\;\\alpha\.\(21\)Consequently a team may inspectERE\_\{R\}after every mission, stop whenever they choose, and the probability of ever issuing a false certificate is at mostα\\alpha\. The same device underlies time\-uniform confidence sequences\([Robbins 1970](https://arxiv.org/html/2608.12895#bib.bib50);[Howard et al\. 2021](https://arxiv.org/html/2608.12895#bib.bib27)\)\.
###### Proof\.
See[SectionA\.14](https://arxiv.org/html/2608.12895#A1.SS14)\. ∎
Two properties matter for this paper specifically\. First,H0H\_\{0\}constrains only the scalar conditional mean ofyry\_\{r\}, and constrains it one mission at a time, sono independence assumption is required— not across missions, not across components, not across shared\-model shocks\. The dependence may take any form whatever, provided it does not lift the conditional mean of the next outcome abovep0p\_\{0\}; and a correlated shock that*degrades*reliability is exactly a shock that does not\. This is the relevant direction: the sequential certificate is immune to the exact failure mode[Section4](https://arxiv.org/html/2608.12895#S4)documents for the composition bound\. Second, the guarantee is on the whole path, which is why certification must latch: onceERE\_\{R\}has crossed, the event “supRER≥1/α\\sup\_\{R\}E\_\{R\}\\geq 1/\\alpha” has occurred and no later evidence undoes it\.[Section10\.5](https://arxiv.org/html/2608.12895#S10.SS5)verifies the shipped implementation latches, driving a certified process through340340consecutive failures\.
### 7\.2Relation to the SPRT
The fixed\-stopping predecessor is[Wald 1945](https://arxiv.org/html/2608.12895#bib.bib58)’s sequential test\.
###### Proposition 7\.1\(Exact SPRT recovery\)\.
Fix an alternativep1\>p0p\_\{1\}\>p\_\{0\}and setλ⋆=\(p1−p0\)/\(p0\(1−p0\)\)\\lambda^\{\\star\}=\(p\_\{1\}\-p\_\{0\}\)/\\bigl\(p\_\{0\}\(1\-p\_\{0\}\)\\bigr\)\. Then fory∈\{0,1\}y\\in\\\{0,1\\\},
1\+λ⋆\(y−p0\)=\(p1p0\)y\(1−p11−p0\)1−y,1\+\\lambda^\{\\star\}\(y\-p\_\{0\}\)\\;=\\;\\Bigl\(\\frac\{p\_\{1\}\}\{p\_\{0\}\}\\Bigr\)^\{y\}\\Bigl\(\\frac\{1\-p\_\{1\}\}\{1\-p\_\{0\}\}\\Bigr\)^\{1\-y\},\(22\)so[Equation20](https://arxiv.org/html/2608.12895#S7.E20)is exactly the Bernoulli likelihood ratio andERE\_\{R\}is the SPRT statistic\. The e\-process therefore contains[Definition3\.12](https://arxiv.org/html/2608.12895#S3.Thmdefinition12)as the special case of a constant, optimally\-tuned bet\.
###### Proof\.
See[SectionA\.15](https://arxiv.org/html/2608.12895#A1.SS15)\. ∎
[Proposition7\.1](https://arxiv.org/html/2608.12895#S7.Thmproposition1)places the cost of anytime validity precisely\. Against a known alternative the SPRT is recovered with no loss; the price is paid only whenλ\\lambdais mis\-tuned, and it is a loss of*power*, never of validity —[Theorem7\.1](https://arxiv.org/html/2608.12895#S7.Thmtheorem1)holds for every predictableλ\\lambda\.[Section10\.5](https://arxiv.org/html/2608.12895#S10.SS5)measures both sides: the type\-I rate stays underα\\alphaacross the whole admissible range ofλ\\lambda, while the mean time to certification varies from14\.614\.6missions atλ=1\.2375\\lambda=1\.2375to403\.9403\.9atλ=0\.125\\lambda=0\.125\. A badly chosen bet costs a factor of2828in detection time and nothing in soundness\.
###### Proposition 7\.2\(Mixture over bets\)\.
Letπ\\pibe a prior on a finite setΛ\\Lambdaof predictable betting strategies and letERmix=∑λ∈Λπ\(λ\)ERλE\_\{R\}^\{\\mathrm\{mix\}\}=\\sum\_\{\\lambda\\in\\Lambda\}\\pi\(\\lambda\)E\_\{R\}^\{\\lambda\}\. ThenEmixE^\{\\mathrm\{mix\}\}is a non\-negative supermartingale underH0H\_\{0\}, so[Theorem7\.1](https://arxiv.org/html/2608.12895#S7.Thmtheorem1)applies unchanged, and
logERmix≥maxλ∈ΛlogERλ−log1π\(λ\),\\log E\_\{R\}^\{\\mathrm\{mix\}\}\\;\\geq\\;\\max\_\{\\lambda\\in\\Lambda\}\\log E\_\{R\}^\{\\lambda\}\-\\log\\frac\{1\}\{\\pi\(\\lambda\)\},\(23\)a pathwise log\-regret guarantee against the best bet in hindsight\.
###### Proof\.
See[SectionA\.16](https://arxiv.org/html/2608.12895#A1.SS16)\. ∎
[Proposition7\.2](https://arxiv.org/html/2608.12895#S7.Thmproposition2)is what makes the method usable whenp1p\_\{1\}is unknown, which is the normal case: the mixture pays an additivelog\(1/π\(λ\)\)\\log\(1/\\pi\(\\lambda\)\)against the best fixed bet and retains anytime validity exactly\.
## 8Runtime enforcement
A certificate is a claim about a system that behaves as contracted\. Producing that system requires intercepting agent actions before they take effect, which is an engineering problem with a small number of solutions and sharply different guarantees between them\.
### 8\.1Three interception planes
An agent action can be intercepted at exactly three points, and where you intercept determines what you can enforce\.
P1 — in\-process\.A hook inside the orchestration framework, invoked around each tool call\. Sees the parsed call, the agent identity, and the framework’s state; can deny by the framework’s own veto convention\. Requires framework support and a per\-framework adapter\.
P2 — tool protocol\.An interposer on the tool\-invocation protocol between the model host and the tool server\. Sees every tool call and result as structured messages, and is framework\-independent\. Cannot see reasoning that never becomes a tool call\.
P3 — transport\.A proxy on the HTTP path to the model provider\. Sees requests and responses but must reconstruct semantic structure from wire format, and cannot attribute an action to an agent in a multi\-agent process without additional context\.
The planes are complementary rather than ranked\. P1 gives the richest state and the weakest portability; P3 gives the reverse; P2 sits between and is the only plane that is both structured and framework\-neutral\.
### 8\.2Denial semantics
###### Definition 8\.1\(Pre\-action and post\-action denial\)\.
A*pre\-action*denial prevents the action from executing\. A*post\-action*denial permits execution and withholds the result\. Only pre\-action denial enforces a governance constraint on an irreversible operation; post\-action denial is a redaction mechanism\.
[Definition8\.1](https://arxiv.org/html/2608.12895#S8.Thmdefinition1)is the distinction that decides whether a deployment is enforced or merely monitored, and it is not always available\. A plane that observes a tool result has already permitted the tool to run\.[Section9](https://arxiv.org/html/2608.12895#S9)records, per surface, which of the two is achievable, because a framework that reports “enforcement enabled” while only offering post\-action denial is making a claim its users will misread\.
### 8\.3Failure policy
###### Definition 8\.2\(Fail\-closed and fail\-open\)\.
When the enforcement path itself errors — a malformed contract, an evaluator exception, an unreachable policy store — a*fail\-closed*policy denies the action and a*fail\-open*policy permits it\. A third state is required:*unevaluated*, in which no verdict was reached and the action is neither credited as compliant nor recorded as violating\.
The third state matters for the measurement in[Section10](https://arxiv.org/html/2608.12895#S10)\. Counting an unevaluated action as compliant inflates every reliability estimate; counting it as a violation deflates them\. Both are silent\. The implementation carriesevaluatedas an explicit field on every verdict so that unevaluated actions are excluded from the denominator rather than assigned to a side\.
### 8\.4Enforcement and certification are different claims
Enforcement constrains what an agent does\. Certification bounds how often the constrained system satisfies its contract\. Neither implies the other: a perfectly enforced hard constraint still permits soft\-constraint drift \([Definition3\.6](https://arxiv.org/html/2608.12895#S3.Thmdefinition6)\), and a high certified floor says nothing about whether any individual action was blocked\.[Section10](https://arxiv.org/html/2608.12895#S10)measures the second; this section describes the machinery that makes the first possible, and the two are reported separately throughout\.
## 9Implementation
Everything in[Section3](https://arxiv.org/html/2608.12895#S3)–[Section7](https://arxiv.org/html/2608.12895#S7)that is computable is implemented inAgentAssert, an open\-source Python library \(AGPL\-3\.0\)\. We describe it at the level a reader needs to reproduce[Section10](https://arxiv.org/html/2608.12895#S10)or to check a claim against code;[AppendixC](https://arxiv.org/html/2608.12895#A3)gives the artifact details\.
### 9\.1Structure
The library separates the three concerns this paper keeps distinct\. Contract specification lives inevaluator/,dsl/, andmodels\.py\([Definitions3\.1](https://arxiv.org/html/2608.12895#S3.Thmdefinition1)and[3\.2](https://arxiv.org/html/2608.12895#S3.Thmdefinition2)\); measurement inmetrics/anddependence/\([Definitions3\.4](https://arxiv.org/html/2608.12895#S3.Thmdefinition4),[3\.7](https://arxiv.org/html/2608.12895#S3.Thmdefinition7)and[3\.17](https://arxiv.org/html/2608.12895#S3.Thmdefinition17)\); and certification incertification/, which carries one module per tier of[Definition5\.2](https://arxiv.org/html/2608.12895#S5.Thmdefinition2)—observed\_floor,lp\_bound,slepian\_floor\(diagnostic only\),eprocess,sprt— assembled bycertificateunder[Proposition5\.1](https://arxiv.org/html/2608.12895#S5.Thmproposition1)\. Runtime enforcement lives ingateway/andenforce/, with the interception surfaces of[Section8](https://arxiv.org/html/2608.12895#S8)as thin adapters over them\.
### 9\.2One policy path, several surfaces
Enforcement across agent frameworks with different hook conventions would ordinarily mean one policy implementation per framework and one opportunity per framework to diverge\. Instead a single bridge exposes a pre\-action decision and a post\-action outcome, and each shim translates those into its host’s veto convention — returning false from a before\-call hook, declining to invoke the continuation, not awaiting the next middleware, or raising from a pre\-hook\. The shims match each host’s hook shapes structurally and import nothing from them, so the library depends on no agent framework\.
Two consequences matter for this paper\. Policy is evaluated in one place, so a contract enforced through a tool\-protocol interposer and the same contract enforced in\-process yield the same verdict on the same state — which is what makes cross\-surface measurement comparable\. And because[Definition8\.1](https://arxiv.org/html/2608.12895#S8.Thmdefinition1)distinguishes pre\-action from post\-action denial while not every host offers the former, the library publishes a per\-surface capability matrix recording where enforcement is impossible\. The negative entries are the point: a library whose documentation implies uniform coverage will be deployed against hosts where the guarantee does not hold\.
### 9\.3State representation
Contracts are evaluated against a flat, dotted\-key state \(output\.pii\_detectedrather than a nested object\)\. The convention is not cosmetic\. An earlier version returned nested dictionaries from one adapter and flat keys from another; a constraint written againstoutput\.pii\_detectedthen resolved correctly under enforcement and silently evaluated to false under measurement, so the same contract yielded two different compliance rates depending on which path produced the state\. Flattening at every adapter boundary removes the divergence by construction\.
### 9\.4Verification
The library carries 1697 tests at 93% statement coverage\. Beyond conventional unit tests, three properties from this paper are pinned by tests because they are the ones an implementation silently gets wrong: that a certified e\-process never un\-certifies after wealth collapses \([Section10\.5](https://arxiv.org/html/2608.12895#S10.SS5)\); that the default certificate tier is Tier 1 unless end\-to\-end execution is explicitly asserted \([Proposition5\.1](https://arxiv.org/html/2608.12895#S5.Thmproposition1)\); and that a contract which could not be evaluated is recorded as unevaluated rather than counted as either outcome \([Definition8\.2](https://arxiv.org/html/2608.12895#S8.Thmdefinition2)\)\.
A paper\-to\-code parity test checks the formulas of[Section3](https://arxiv.org/html/2608.12895#S3)against their implementations, so a change to either that breaks correspondence fails the suite\. That test exists because two defects found during preparation of this work — a mishandled constraint state across the proxy and hook surfaces, and an incorrect divergence term in a drift computation — produced confidently wrong numbers rather than errors\.
## 10Evaluation
We answer six questions with six experiments\. E1 measures whether model sharing induces correlated contract failures across a pipeline \([Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)\)\. E2 measures how much a moment\-set certificate lifts the reliability floor over the Fréchet bound \([Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3)\)\. E3 measures how fast independence\-assuming coverage collapses as dependence grows \([Section10\.4](https://arxiv.org/html/2608.12895#S10.SS4)\)\. E4 measures the empirical type\-I error of the anytime\-valid certificate \([Section10\.5](https://arxiv.org/html/2608.12895#S10.SS5)\)\. E5 measures contract behavior on frontier models under live API conditions \([Section10\.6](https://arxiv.org/html/2608.12895#S10.SS6)\)\. E6 is an ablation over the clustering and serial\-dependence assumptions the analysis rests on \([Section10\.7](https://arxiv.org/html/2608.12895#S10.SS7)\)\.[Section10\.8](https://arxiv.org/html/2608.12895#S10.SS8)re\-presents the v1 single\-agent evidence this work builds on, and[Section10\.9](https://arxiv.org/html/2608.12895#S10.SS9)states what the six experiments jointly support and what they do not\.
### 10\.1Setup, preregistration, and scoring
##### Preregistration\.
The confirmatory hypotheses, conditions, sample sizes, sampling parameters, primary estimator, stopping rule, and falsification criteria for E1 were committed to the repository before any confirmatory outcome was generated\([Nosek et al\. 2018](https://arxiv.org/html/2608.12895#bib.bib41)\), and the commit is git\-timestamped\. We reproduce the three registered hypotheses verbatim:
- •H1 \(dependence exists under sharing\)\.In thesame\_modelcondition, co\-failure dependence between the two pipeline agents is positive: Kendall’sτa\>0\\tau\_\{a\}\>0with a bootstrap CI excluding00\.
- •H2 \(dependence decreases with less sharing\)\.τa\\tau\_\{a\}is weakly monotone decreasing acrosssame\_model≥same\_vendor≥different\_vendor\\texttt\{same\\\_model\}\\geq\\texttt\{same\\\_vendor\}\\geq\\texttt\{different\\\_vendor\}\.
- •H3 \(naive bound is anti\-conservative under sharing\)\.Observed graph reliabilityℙ\(YG=1\)\\mathbb\{P\}\(Y\_\{G\}=1\)*exceeds*the independence product∏ipi\\prod\_\{i\}p\_\{i\}under positive dependence \(the compositional gap is signed and non\-zero\), so a dependence\-aware bound is required\.
The registration further designatesseries2as the confirmatory motif and names the three OpenRouter arms atn=6000n=6000as the primary confirmatory arms, placing other motifs in “future/secondary work”\. We hold to that scope:the confirmatory evaluation isseries2over3×6000=18,0003\\times 6000=18\{,\}000missions\. Thequorum2of3andparallel2results, and the two breadth arms of[Section10\.6](https://arxiv.org/html/2608.12895#S10.SS6), are reported as*secondary and exploratory*throughout\. Where they agree with the confirmatory arms we say so; they are not evidence at the registered level, and we have not filed a dated amendment broadening the registration\.
The registration also fixed what would falsify the thesis: if thesame\_modelτa\\tau\_\{a\}CI covers00, orτa\\tau\_\{a\}fails to decrease fromsame\_modeltodifferent\_vendor, or the composition gap is approximately zero, the claim is unsupported and the null is reported\. We report against that criterion in[Section10\.2\.2](https://arxiv.org/html/2608.12895#S10.SS2.SSS2), including one registered comparison that does*not*survive\.
##### Missions and scoring\.
Missions are drawn from six seeded generators spanning retail and financial agent tasks\. Every mission is scored by*deterministic gold code*, never by an LLM judge: order arithmetic, refund\-policy application, and promotional\-cap enforcement on the retail side; transaction\-limit checks, watchlist screening, and mandatory\-disclaimer presence on the financial side\. The mission\-to\-task assignment is a fixed SHA\-256 hash of the mission identifier, which makes each arm’s task sequence deterministic and reproducible\. It does*not*make the realised task mixtures identical across arms: identifiers embed the condition name \([Section11\.3](https://arxiv.org/html/2608.12895#S11.SS3)\), so each arm is an independent deterministic draw from the same generator distribution rather than a paired design\. Atn=6000n=6000per confirmatory arm the resulting imbalance is absorbed by the bootstrap; for the under\-run breadth arms it is a real limit\.
Each agent invocation yields a hard verdict and a soft score\. The hard verdicthi∈\{0,1\}h\_\{i\}\\in\\\{0,1\\\}records whether the agent satisfied every hard clause of its contract; the soft scoreσi∈\[0,1\]\\sigma\_\{i\}\\in\[0,1\]records graded compliance\. All dependence estimates in[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)are computed on the*hard*verdict, so a “failure” means a definite contract violation, not a low score\. Graph successYGY\_\{G\}is the motif’s composition rule applied to the component hard verdicts\.
##### Frozen sampling\.
Every call usestemperature=0\.2=0\.2,top\_p=1\.0=1\.0, andmax\_output\_tokens=160=160, with client\-side prompt truncation at32003200characters\. These values were fixed at registration and never varied\.
##### Conditions\.
The manipulated variable is how much the two scored agents share\. Insame\_modelboth runmistral\-small\-24b; insame\_vendorthe second agent is swapped toministral\-8b\(same vendor, smaller model\); indifferent\_vendorit is swapped togemma\-3\-12b\-it\. The first agent is held atmistral\-small\-24bthroughout, so the contrast isolates the substitution\.
##### Motifs\.
Three topologies are measured\.series2is a two\-agent handoffA→BA\\rightarrow Bwhere both must satisfy their contract forYG=1Y\_\{G\}=1\.parallel2runs two branches into a deterministic merge\.quorum2of3runs three workers into a deterministic 2\-of\-3 aggregator\. The aggregator and merge nodes are deterministic code, not models, and are excluded from dependence estimation\.
##### Scale\.
The campaign is 30,820 scored missions across 12 arms, recorded in 13 execution logs\. The counting rule matters and we state it: thequorum3of4arm was executed twice, so its two logs share212212mission identifiers, which we deduplicate with the later pass winning; that arm therefore contributes717717missions rather than the929929records on disk, and717717is thennused throughout[Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3)\. The primary confirmatory arms are the threeseries2conditions atn=6000n=6000each\.
##### Estimators\.
For two agents with failure indicatorsFA=1−hAF\_\{A\}=1\-h\_\{A\}andFB=1−hBF\_\{B\}=1\-h\_\{B\}we form the2×22\\times 2co\-failure table with cellsn11,n10,n01,n00n\_\{11\},n\_\{10\},n\_\{01\},n\_\{00\}and corresponding proportionsp11,p10,p01,p00p\_\{11\},p\_\{10\},p\_\{01\},p\_\{00\}\. We report five statistics, and the distinction between the first three and the last two carries the analysis:
J\\displaystyle J=n11n11\+n10\+n01,\\displaystyle=\\frac\{n\_\{11\}\}\{n\_\{11\}\+n\_\{10\}\+n\_\{01\}\},\(24\)τa\\displaystyle\\tau\_\{a\}=2\(p11p00−p10p01\),\\displaystyle=2\\,\(p\_\{11\}p\_\{00\}\-p\_\{10\}p\_\{01\}\),\(25\)ϕ\\displaystyle\\phi=p11p00−p10p01pA\(1−pA\)pB\(1−pB\),\\displaystyle=\\frac\{p\_\{11\}p\_\{00\}\-p\_\{10\}p\_\{01\}\}\{\\sqrt\{p\_\{A\}\(1\-p\_\{A\}\)\\,p\_\{B\}\(1\-p\_\{B\}\)\}\},\(26\)logOR\\displaystyle\\log\\mathrm\{OR\}=log\(n11\+12\)\(n00\+12\)\(n10\+12\)\(n01\+12\),\\displaystyle=\\log\\frac\{\(n\_\{11\}\+\\tfrac\{1\}\{2\}\)\(n\_\{00\}\+\\tfrac\{1\}\{2\}\)\}\{\(n\_\{10\}\+\\tfrac\{1\}\{2\}\)\(n\_\{01\}\+\\tfrac\{1\}\{2\}\)\},\(27\)Q\\displaystyle Q=OR−1OR\+1,\\displaystyle=\\frac\{\\mathrm\{OR\}\-1\}\{\\mathrm\{OR\}\+1\},\(28\)wherepA=p11\+p10p\_\{A\}=p\_\{11\}\+p\_\{10\}andpB=p11\+p01p\_\{B\}=p\_\{11\}\+p\_\{01\}are the marginal failure rates\. These are, in order, the overlap coefficient of[Jaccard 1912](https://arxiv.org/html/2608.12895#bib.bib29), the binary form of[Kendall 1938](https://arxiv.org/html/2608.12895#bib.bib31)’sτ\\tau, the phi coefficient, the log odds ratio, and[Yule 1900](https://arxiv.org/html/2608.12895#bib.bib63)’sQQ\.[Equation24](https://arxiv.org/html/2608.12895#S10.E24)is the overlap of the two failure*sets*; then00n\_\{00\}cell is deliberately excluded so that missions neither agent failed cannot dilute the statistic\.[Equation27](https://arxiv.org/html/2608.12895#S10.E27)uses the Haldane–Anscombe\+12\+\\tfrac\{1\}\{2\}correction, which keeps the estimator finite when a cell is empty\.
JJ,τa\\tau\_\{a\}, andϕ\\phiare all sensitive to the marginals:τa\\tau\_\{a\}is twice the covariance of the two indicators and therefore cannot reach±1\\pm 1unless both marginals equal12\\tfrac\{1\}\{2\}, andJJcharges every mission that exactly one agent failed to its denominator\.logOR\\log\\mathrm\{OR\}andQQare invariant to the marginals under row and column scaling\.[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)shows this distinction is not pedantic: it decides whether the registered monotonicity claim survives\.
##### Uncertainty\.
Confidence intervals are percentile bootstrap\([Efron 1979](https://arxiv.org/html/2608.12895#bib.bib15)\)withB=2000B=2000resamples atα=0\.05\\alpha=0\.05, seeded at20260813\. The registration specifies a cluster bootstrap clustered by mission\. In the logscluster\_idis one\-to\-one withmission\_id\(6000 distinct clusters over 6000 missions in each primary arm\), so the cluster bootstrap reduces exactly to an i\.i\.d\. bootstrap over missions\. We state this rather than leave it implicit, and[Section10\.7](https://arxiv.org/html/2608.12895#S10.SS7)measures the design effect\([Kish 1965](https://arxiv.org/html/2608.12895#bib.bib32)\)directly instead of assuming it: the empirical DEFF is1\.001\.00–1\.161\.16\.
Hypothesis H2 is a statement about*differences*between arms, so we test it on the bootstrap distribution of the arm contrast rather than by inspecting whether two intervals overlap\. Overlapping marginal intervals do not imply a non\-significant difference, and the contrast test is the correct instrument\.
### 10\.2E1: model sharing induces correlated contract failures
#### 10\.2\.1Dependence is present, large, and positive in every arm
[Table2](https://arxiv.org/html/2608.12895#S10.T2)gives the primary result\. In every condition the two agents’ hard failures are strongly positively associated\. The weakest association we measure anywhere in E1 islogOR=2\.92\\log\\mathrm\{OR\}=2\.92\(OR≈18\.5\\mathrm\{OR\}\\approx 18\.5\) inparallel2/different\_vendor; the strongest islogOR=6\.66\\log\\mathrm\{OR\}=6\.66\(OR≈784\\mathrm\{OR\}\\approx 784\) inseries2/same\_model\. Every interval excludes independence by a wide margin\. H1 as registered concerns thesame\_modelarm, and it is supported there\. We note in passing that the association is also positive and significant in the other two arms; that is a descriptive observation, not a registered result, and it is a reminder that substituting a model reduces dependence without eliminating it\.
Table 2:E1 primary arms\. Motifseries2, pair\(node\_a,node\_b\)\(\\texttt\{node\\\_a\},\\texttt\{node\\\_b\}\),n=6000n=6000per arm\. Cells are co\-failure counts on the hard verdict\. Brackets are percentile bootstrap 95% CIs,B=2000B=2000\.The practical reading ofsame\_modelis blunt\. AgentAAfails39\.4%39\.4\\%of missions and agentBBfails37\.2%37\.2\\%\. Under independence the two would co\-fail on14\.6%14\.6\\%of missions; they co\-fail on36\.3%36\.3\\%\. Of the 2418 missions on which at least one agent failed, both failed on 2177 —90\.0%90\.0\\%\. Replacing one agent with a second instance of the same model buys almost no independent evidence\.
##### This is not merely interface propagation\.
series2is a handoff, so a failure ofAAdegrades the inputBBreceives, and co\-failure there could in principle be explained without appeal to shared failure modes at all\. Two features of the design separate the mechanisms\. First,parallel2andquorum2of3are*not*handoffs — their model nodes receive the mission independently and never see one another’s output — and both reproduce the effect \([Table3](https://arxiv.org/html/2608.12895#S10.T3)\), which propagation cannot explain\. Second, the control pair of[Section10\.2\.4](https://arxiv.org/html/2608.12895#S10.SS2.SSS4)sits atlogOR≈5\.1\\log\\mathrm\{OR\}\\approx 5\.1with no channel between its two workers\. Propagation may well contribute to theseries2magnitude and we do not claim to have partitioned the two mechanisms; the arm*contrasts*, which is what H2 concerns, are measured within identical topologies and so are unaffected by it\.
#### 10\.2\.2The registered monotonicity claim splits
[Table3](https://arxiv.org/html/2608.12895#S10.T3)tests H2 directly on the bootstrap distribution of each arm contrast\. The verdict is not uniform, and we state it plainly:H2 holds at the model\-sharing level and fails at the vendor level\.
Table 3:E1 arm contrasts, all three motifs\. Arms abbreviatedSMsame\_model,SVsame\_vendor,DVdifferent\_vendor\. Entries are the point difference with a 95% bootstrap CI on the difference;SIGmarks an interval excluding zero\. The SV−\-DV row is the registered comparison that does not survive\. Onlyseries2is confirmatory; the other two motifs are secondary\.Two findings are robust\. Sharing the*same model*produces significantly more co\-failure than either alternative, on all five statistics, in all three motifs, with no exceptions\. That is six independent significant contrasts in the registered direction\.
The third comparison does not replicate\. Moving fromsame\_vendortodifferent\_vendorproduces no consistent change: inquorum2of3every statistic is non\-significant; inseries2andparallel2the marginal\-sensitive and marginal\-free statistics disagree in sign or significance\. We therefore do not claim a three\-level ordering\. Vendor identity, as manipulated here, is not a reliable predictor of failure correlation once model identity differs\.
Against the registered falsification criterion, the thesis survives: thesame\_modelτa\\tau\_\{a\}CI excludes zero, andτa\\tau\_\{a\}decreases significantly fromsame\_model\(0\.43270\.4327\) todifferent\_vendor\(0\.38260\.3826\), difference0\.05000\.0500, CI\[0\.0389,0\.0615\]\[0\.0389,0\.0615\]\. The registered falsification test was stated on exactly this contrast and it passes\. The intermediate rung is where the ordering breaks, and the registration did not make the intermediate rung a falsification condition\.
#### 10\.2\.3Why the marginal\-sensitive statistics reverse
Inseries2the Jaccard overlap is*lower*forsame\_vendor\(0\.73550\.7355\) than fordifferent\_vendor\(0\.79450\.7945\), and the difference is significant\. Read naively this says swapping to a different vendor increases correlated failure, which inverts the mechanism the experiment was built to test\. It is an artifact of the marginals, and the artifact is measurable rather than conjectural\.
ministral\-8bis a weaker model thanmistral\-small\-24band fails far more often\. In thesame\_vendorarm the two marginal failure rates arepA=0\.3925p\_\{A\}=0\.3925andpB=0\.5077p\_\{B\}=0\.5077, a gap of0\.11520\.1152; indifferent\_vendorthe gap is0\.01070\.0107\. That asymmetry lands in then01n\_\{01\}cell: agentBBfails alone on 757 missions insame\_vendoragainst 225 indifferent\_vendor\. BecauseJJcarriesn01n\_\{01\}in its denominator, an arm whose two agents fail at different rates is charged a penalty that has nothing to do with whether their failures are associated\.
The pattern holds across motifs\. The marginal gap\|pA−pB\|\|p\_\{A\}\-p\_\{B\}\|is largest insame\_vendorin every motif \(0\.1150\.115,0\.2020\.202,0\.1970\.197\) and small insame\_model\(0\.0230\.023,0\.0060\.006,0\.0060\.006\), and it is exactly the arms with large gaps whoseJJ,τa\\tau\_\{a\}, andϕ\\phiare depressed\. Substituting a model changes two things at once — how often that agent fails, and how much its failures align with its partner’s — and only the second is the quantity H2 is about\. The marginal\-free statistics separate them: onlogOR\\log\\mathrm\{OR\}andQQtheseries2reversal disappears and the contrast is non\-significant\.
We therefore report H2 onlogOR\\log\\mathrm\{OR\}, and reportJJalongside it becauseJJis the quantity a practitioner actually feels — it is the fraction of observed failures that a redundant agent fails to catch\. Both are correct statistics; they answer different questions, and the paper needs both\.
#### 10\.2\.4An internal negative control
Thequorum2of3design contains an internal control\. Onlyworker\_1is substituted across arms:worker\_0andworker\_2both runmistral\-small\-24bin*all three*conditions\. The pair\(worker\_0,worker\_2\)\(\\texttt\{worker\\\_0\},\\texttt\{worker\\\_2\}\)is therefore a same\-model pair whose composition never changes while the arm label does\. If the arm effects in[Table3](https://arxiv.org/html/2608.12895#S10.T3)were driven by anything other than the model substitution — task mix drifting between runs, ordering, wall\-clock effects, provider\-side variation — this pair would move with the arm label\. It must not move, and it does not\.
Table 4:Negative control\. Pair\(worker\_0,worker\_2\)\(\\texttt\{worker\\\_0\},\\texttt\{worker\\\_2\}\)inquorum2of3ismistral\-small\-24b×\\timesmistral\-small\-24bin every arm\. All 15 arm contrasts across the five statistics are non\-significant\.[Table4](https://arxiv.org/html/2608.12895#S10.T4)reports the result\. Across five statistics and three pairwise contrasts — fifteen tests — not one interval excludes zero\. The control pair’s dependence sits atJ≈0\.83J\\approx 0\.83–0\.850\.85andlogOR≈5\.07\\log\\mathrm\{OR\}\\approx 5\.07–5\.325\.32in every arm, which is where the*manipulated*pair sits in itssame\_modelcondition \(J=0\.8251J=0\.8251,logOR=4\.957\\log\\mathrm\{OR\}=4\.957\)\. A same\-model pair measures as a same\-model pair regardless of what the rest of the graph is doing\. That is convergent validity for the estimator and, jointly with the fifteen nulls, evidence that the arm effects in[Table3](https://arxiv.org/html/2608.12895#S10.T3)are caused by the model substitution rather than by any arm\-level confound\.
#### 10\.2\.5What E1 supports
Model sharing is the operative variable\. Two instances of one model co\-fail on90%90\\%of the missions either one fails; substituting a different model reduces that overlap significantly and reduces the odds\-ratio association by roughly1\.81\.8–2\.12\.1log units, consistently across three topologies\. Vendor identity, holding model identity different, does not reliably matter\. A redundant agent running the same model as the agent it is meant to check supplies far less independent evidence than a naive composition calculation credits it with, and[Section10\.4](https://arxiv.org/html/2608.12895#S10.SS4)quantifies what that does to a reliability guarantee\.
### 10\.3E2: enriching the moment set lifts the certified floor
E1 establishes that components fail together\. A practitioner still has to certify a number\. E2 measures what the moment\-set certificate of[Section6](https://arxiv.org/html/2608.12895#S6)buys over the assumption\-free bound on real data\.
##### Data\.
The four\-workerquorum3of4same\_modelarm,m=4m=4scored workers overn=717n=717missions\. The arm was executed in two passes; we take the union of both logs deduplicated by mission identifier, with the later pass winning on the 212 missions present in both\. The observed all\-success rate is0\.53280\.5328\.
##### Procedure\.
We compute two things at two moment sets\. The*sharp point\-moment interval*is the LP over all joint laws matching the empirical moments exactly — it isolates how much the moment family alone identifies, with no sampling uncertainty\. The*certified floor*is the LP over the Clopper–Pearson box around those moments at family\-wiseηconf=0\.05\\eta\_\{\\mathrm\{conf\}\}=0\.05, which is what may actually be reported\. The moment sets areJ=10J=10\(marginals and pairwise co\-success,\(41\)\+\(42\)\\binom\{4\}\{1\}\+\\binom\{4\}\{2\}\) andJ=14J=14\(adding all four triple co\-success moments\)\.
Table 5:E2 floor lifting\.m=4m=4,n=717n=717, observed all\-success0\.53280\.5328\. The pre\-allocated Bonferroni row spends the confidence budget over the larger familyJmax=14J\_\{\\max\}=14regardless of how many moments are constrained, which makes the certified floor monotone in the moment set by construction\.[Table5](https://arxiv.org/html/2608.12895#S10.T5)reports the result\. Adding four triple moments narrows the sharp identified interval from width0\.04880\.0488to0\.00700\.0070, an85\.7%85\.7\\%reduction, and lifts the certified floor from0\.24550\.2455to0\.41160\.4116—16\.616\.6percentage points of reliability that the pairwise certificate cannot claim and the triple certificate can, on identical data\.
Two details matter for honest reporting\. First, the certified floor atJ=14J=14\(0\.41160\.4116\) remains well below the observed rate \(0\.53280\.5328\), because it is a distribution\-free lower confidence bound over every joint law consistent with the moment box, not an estimate\. It is a floor, and it is meant to be conservative\. Second, the used\-set allocation is tighter atJ=10J=10\(0\.24550\.2455against0\.23570\.2357\) but its monotonicity in the moment set is empirical rather than guaranteed: adding a moment widens every interval, so a richer moment family can in principle certify less\. The pre\-allocated allocation pays0\.00980\.0098atJ=10J=10to buy monotonicity by construction\. We report both rather than quietly choosing the flattering one\.
### 10\.4E3: coverage of a model\-based floor collapses under misspecification
[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)predicts that a bootstrap bound on a fitted model’s functional stops covering the truth once the model is misspecified\. E3 measures that against a control, because the claim is not that the Gaussian floor is always wrong — it is that the floor cannot tell you when it is\.
##### Two arms\.
Both samplem=3m=3components at marginal0\.60\.6\. The*control*arm draws from the equicorrelated Gaussian one\-factor law itself, for which the fitted family is correctly specified andΔ=0\\Delta=0; the floor should hold at nominal there, and if it did not, any collapse in the other arm would indict the estimator rather than the misspecification\. The*adversarial*arm draws from the witness of[Example4\.2](https://arxiv.org/html/2608.12895#S4.Thmexample2): the LP\-minimising law, which has*identical*marginals and*identical*pairwise co\-success moments and an all\-success probability of0\.3304750\.330475against the Gaussian law’s0\.3921440\.392144, soΔ=0\.061670\\Delta=0\.061670\. Its five non\-zero cell masses are listed in[Example4\.2](https://arxiv.org/html/2608.12895#S4.Thmexample2)\. Both constants come from the same deterministic quadrature and LP used in[Example4\.2](https://arxiv.org/html/2608.12895#S4.Thmexample2), not from simulation\. No procedure restricted to marginal and pairwise data can distinguish the two\.
Each arm runs200200replications at each of four sample sizes, computing the Gaussian model floor with500500bootstrap resamples atηconf=0\.05\\eta\_\{\\mathrm\{conf\}\}=0\.05and recording how often it lies at or below the*true*all\-success probability\. The Tier\-1 floor of[Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)is computed on the same draws\.
Table 6:E3 coverage of the*true*all\-success probability, ashits/200\\text\{hits\}/200with a 95% Clopper–Pearson interval\.nboot=500n\_\{\\mathrm\{boot\}\}=500,ηconf=0\.05\\eta\_\{\\mathrm\{conf\}\}=0\.05, seed20260812\. Nominal coverage is0\.950\.95\. The Tier\-1 floor covered in all200200replications of every cell and is omitted from the interval columns\.In the control arm the model floor holds between0\.940\.94and0\.960\.96at every sample size, with all four intervals covering the nominal0\.950\.95— a correctly specified model behaving correctly across a factor of eight innn\. In the adversarial arm coverage falls from0\.360\.36atn=250n=250to0\.010\.01atn=2000n=2000, an upper confidence limit of0\.0360\.036: by two thousand missions the certificate sits above the true reliability in198198of200200replications\. The Tier\-1 floor covers in all16001600replications across both arms\.
We report the counts rather than the rounded rates because a small count is not a zero\. Two hits in two hundred replications bounds coverage above by0\.0360\.036; it does not establish that coverage*is*zero, and a run reporting0/2000/200would bound it above by0\.0180\.018rather than prove exact failure\.
Three points follow\. The two arms are*indistinguishable*on the evidence the model consumes, so the collapse is not a case a diagnostic on marginals or pairwise moments could detect and exclude\. Coverage*decreases*asnngrows, inverting the usual reassurance that more data makes an interval safer: the interval narrows, but around the fitted functional rather than the truth\. We report that decrease as an empirical observation —[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)claims only the limit, since tightness does not entail monotonicity \([SectionA\.6](https://arxiv.org/html/2608.12895#A1.SS6)\)\. And the Tier\-1 floor is conservative here,1\.001\.00against a nominal0\.950\.95; that conservatism is the price of the immunity, as[Corollary4\.3](https://arxiv.org/html/2608.12895#S4.Thmcorollary3)describes\.
##### A correction to the previous version of this experiment\.
The v2 preprint reported a collapse from a simulation that drew from the Gaussian law\. That design cannot exhibit the effect: with the model correctly specifiedΔ=0\\Delta=0,[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)does not apply, and re\-running it returns the control column above\. Obtaining[Table6](https://arxiv.org/html/2608.12895#S10.T6)required a law that is misspecified*and*moment\-indistinguishable, which[Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)supplies as the LP minimiser\. We record the correction rather than quietly restate the earlier number\.
### 10\.5E4: the anytime\-valid certificate holds its type\-I error
A team that watches a reliability dashboard and stops when it first looks good is running an optional\-stopping experiment, and a fixed\-nnconfidence interval does not survive that\. The e\-process of[Section7](https://arxiv.org/html/2608.12895#S7)is valid under arbitrary stopping\. E4 measures whether the shipped implementation attains the guarantee\.
##### Procedure\.
We simulate 8000 independent null streams of 500 missions each at the null boundaryptrue=p0=0\.8p\_\{\\text\{true\}\}=p\_\{0\}=0\.8, withα=0\.05\\alpha=0\.05and seed4242, and record the fraction of streams whose wealth ever reaches1/α1/\\alpha\. Ville’s inequality bounds that fraction byα\\alphafor*every*predictable betting fractionλ\\lambda, not merely for a tuned one, so we sweep the whole admissible rangeλ∈\(0,1/p0\)\\lambda\\in\(0,1/p\_\{0\}\)and report the curve\. A singleλ\\lambdawould not distinguish a valid certificate from a lucky one\.
Table 7:E4 empirical type\-I crossing rate under the null\.p0=0\.8p\_\{0\}=0\.8,α=0\.05\\alpha=0\.05, 8000 streams×\\times500 missions, seed4242\. The three\-sigma Monte Carlo band on a proportion of sizeα\\alphaat thisnstreamsn\_\{\\text\{streams\}\}is0\.00730\.0073\.The worst rate over the six prespecified fractions is0\.04710\.0471, below the nominal0\.050\.05, and every individual rate is inside the Monte Carlo band\. Six points cannot establish a supremum over a continuum; the guarantee for every predictableλ\\lambdais[Theorem7\.1](https://arxiv.org/html/2608.12895#S7.Thmtheorem1), and this experiment checks that the shipped implementation attains it at fractions spanning the admissible range\. The bound is attained most tightly atλ=0\.625\\lambda=0\.625, which is exactly the SPRT\-optimal betλ⋆=\(p1−p0\)/\(p0\(1−p0\)\)\\lambda^\{\\star\}=\(p\_\{1\}\-p\_\{0\}\)/\(p\_\{0\}\(1\-p\_\{0\}\)\)for the alternativep1=0\.9p\_\{1\}=0\.9— the certificate is tightest where the betting fraction is best matched to the alternative, and conservative elsewhere, which is the expected shape\.
##### SPRT recovery\.
Settingλ⋆=\(p1−p0\)/\(p0\(1−p0\)\)\\lambda^\{\\star\}=\(p\_\{1\}\-p\_\{0\}\)/\(p\_\{0\}\(1\-p\_\{0\}\)\)makes the betting factor identical to the Bernoulli likelihood ratio\. We verify this against the shippedfrom\_sprtconstructor rather than against a second closed form: across five\(p0,p1\)\(p\_\{0\},p\_\{1\}\)pairs and three mission sequences, the maximum absolute deviation between the accumulated log\-wealth and the accumulated Bernoulli log\-likelihood\-ratio is0\.00\.0— exact in floating point, not merely close\.
##### Monotone latch\.
Certification is a statement about the path, not the present: once the wealth has crossed1/α1/\\alpha, no subsequent evidence retracts it\. We drive a process across the threshold at mission1414\(peak log\-wealth13\.2713\.27\) and then feed it340340consecutive failures, collapsing log\-wealth to−1552\.49\-1552\.49\. The certificate remains issued throughout\. This is correct — the crossing happened, and Ville bounds the probability that it ever happens under the null — and it is a property implementations get wrong by recomputing certification from current wealth\.
### 10\.6E5: the dependence result replicates across backends
The three primary arms of E1 all run through a single inference provider\. If the measured association were an artifact of that provider’s serving stack — batching, caching, a shared sampler — it would not appear when the upstream agent runs somewhere else entirely\. The preregistration therefore fixed two cross\-backend breadth arms, both on theseries2motif, in which the upstream agent is served by a different provider and the downstream agent remainsmistral\-small\-24b\. The Meta\-backed arm runsmuse\-spark\-1\.2\-contributorupstream; the Grok\-backed arm runsgrok\-4\.5\. Both arms change the upstream model as well as the serving backend, so we label them by backend as shorthand, not as a claim that the backend is the only thing varying\.
##### Registered deviation, and attrition\.
Both arms were registered atn=2000n=2000and both yield fewer scored missions:635635for the Meta arm and19011901for the Grok arm\. The shortfall is*output attrition*, and its mechanism matters more than its size\. Meta returned13651365unusable responses —13521352empty completions under the frozen reasoning and token settings of[Section10\.1](https://arxiv.org/html/2608.12895#S10.SS1), plus1313rate limits; Grok returned9999, almost all rate limits\.
Meta’s68\.25%68\.25\\%attrition may therefore be*outcome\-dependent*: a model returning no text is plausibly one that would also have failed the contract, so the surviving635635missions are a selected subset and their complete\-case estimate is descriptive only\. Grok’s5%5\\%loss is transport error and carries no such selection\. We read Grok as usable secondary evidence and Meta as directional; both were registered as breadth arms rather than confirmatory tests in any case\.
Table 8:E5 cross\-backend replication, motifseries2\. The upstream agent changes provider; the downstream agent ismistral\-small\-24bthroughout\. Marginal\-sensitive and marginal\-free statistics move in*opposite*directions\.Positive association appears on both breadth arms, and on the Grok arm it appears cleanly\. That arm scores19011901of its20002000registered missions, loses the remainder to transport errors that carry no outcome selection, and returnslogOR=7\.306\\log\\mathrm\{OR\}=7\.306with a95%95\\%lower bound of6\.8986\.898against00under independence\. A serving\-stack artifact confined to one provider cannot produce that\. The Meta arm points the same way, with a lower bound of3\.7583\.758, but its attrition is selective and it corroborates rather than replicates\.
We do*not*read the breadth arms as strengthening the sharing story, and the reason is worth stating because the table invites the opposite reading\. The Grok arm’slogOR\\log\\mathrm\{OR\}of7\.3067\.306is the largest association anywhere in this paper and it is a*cross\-vendor*pair, nominally exceedingsame\_model’s6\.6656\.665\. That comparison is not admissible\. The marginal failure rates differ by more than an order of magnitude between the breadth and primary arms \(0\.0170\.017and0\.0590\.059against0\.3790\.379and0\.3690\.369\), the upstream model is different, and the Grok estimate rests on3333concordant and8080discordant missions withn10=0n\_\{10\}=0, so it depends on the Haldane–Anscombe correction\. Odds\-ratio invariance to row and column scaling does not make estimates from such different regimes comparable in magnitude\. The breadth arms establish that the association*exists*on other backends; they do not rank sharing conditions, and any cross\-arm magnitude comparison with[Table2](https://arxiv.org/html/2608.12895#S10.T2)should be resisted — including one that appears to favour our thesis\.
The Jaccard column tells the opposite story —0\.1000\.100and0\.2920\.292against the primary arm’s0\.7940\.794— and the reason is the mechanism of[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)appearing again on independent data\. The upstream agents here are far stronger:grok\-4\.5fails1\.74%1\.74\\%of missions andmuse\-spark\-1\.2fails0\.16%0\.16\\%, againstmistral\-small\-24b’s37\.9%37\.9\\%in the primary arm\. Becauseseries2is a handoff, a cleaner upstream output makes the downstream task easier, and the downstream agent’s failure rate falls from36\.9%36\.9\\%to5\.9%5\.9\\%and1\.6%1\.6\\%respectively\. Both marginals collapse, andJJ— which is bounded by the marginals — collapses with them, while the marginal\-freelogOR\\log\\mathrm\{OR\}rises\. The two statistics are not in conflict; they are measuring different things, exactly as[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)argued, and E5 is an independent replication of that methodological point on data from three different providers\.
Two limits are worth stating rather than burying\. The Meta arm records a single co\-failure \(n11=1n\_\{11\}=1\), so its Jaccard confidence interval is\[0\.000,0\.333\]\[0\.000,0\.333\]and the point estimate carries essentially no information; only itslogOR\\log\\mathrm\{OR\}is usable, and even that rests on nine discordant missions\. Andn10=0n\_\{10\}=0in both arms — the upstream agent never failed without the downstream agent also failing\. With only11and3333upstream failures observed this is unsurprising under any model, so we do not read deterministic failure propagation into it\. It does mean the odds ratio in these arms depends on the Haldane–Anscombe correction, which is why we report it with the correction stated rather than as an unadorned ratio\.
Because mission identifiers embed the condition name and the task is a fixed hash of the identifier, the arms share no missions: task draws are independent across arms, from the same generator distribution\. The comparison is between distributions, not between paired tasks\.
What E5 establishes is narrow, and it is solid\. On the Grok arm —19011901missions, non\-selective attrition, an interval nowhere near independence — cross\-agent failure dependence is present on infrastructure the primary arms never touched, and the marginal\-sensitivity mechanism of[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)reproduces there on data the primary arms did not generate\. One clean external replication and one directional corroboration is what two breadth arms can honestly buy, and it is enough to rule out the single\-provider explanation\.
### 10\.7E6: the i\.i\.d\. assumption, tested rather than assumed
Every interval in E1 and every floor in E2 assumes missions are independent\. That assumption is testable, and E6 tests it two ways instead of asserting it\.
##### Serial dependence\.
For each of the 19 agent\-arm combinations we compute the lag\-kkautocorrelation of the failure indicator in execution order and the induced design effectDEFF=1\+2∑k=110\(1−k/n\)ρk\\mathrm\{DEFF\}=1\+2\\sum\_\{k=1\}^\{10\}\(1\-k/n\)\\rho\_\{k\}\. The largest first\-order autocorrelation anywhere is\|ρ1\|=0\.0546\|\\rho\_\{1\}\|=0\.0546and the largest design effect isDEFF=1\.1482\\mathrm\{DEFF\}=1\.1482; many arms returnDEFF<1\\mathrm\{DEFF\}<1, indicating mild negative serial dependence\. AtDEFF=1\.15\\mathrm\{DEFF\}=1\.15the E1 confidence intervals are optimistic by a factor of1\.15≈1\.07\\sqrt\{1\.15\}\\approx 1\.07, which does not change any verdict in[Table3](https://arxiv.org/html/2608.12895#S10.T3): the smallest significant contrast there has a margin far exceeding7%7\\%of its width\.
##### Floor sensitivity\.
We then concede a design effect we do not measure —DEFF=1\.5\\mathrm\{DEFF\}=1\.5— and ask what the certified floor costs\. The measurement has to isolate the width of the Clopper–Pearson intervals: subsampling missions ton/1\.5n/1\.5would also perturb the empirical moments, and the resulting movement would be the sum of interval widening and sampling noise reported as if it were the former\. We instead hold the joint cell distribution fixed, rescale the cell counts toneff=n/1\.5n\_\{\\text\{eff\}\}=n/1\.5by largest\-remainder rounding, and recompute\. Onlynnchanges\.
Table 9:E6 certified\-floor sensitivity at a concededDEFF=1\.5\\mathrm\{DEFF\}=1\.5,ηconf=0\.05\\eta\_\{\\mathrm\{conf\}\}=0\.05\. Movement scales inversely withnn, as Clopper–Pearson width must\.The floor moves at most2\.692\.69percentage points under a conceded design effect of1\.51\.5, which is1\.31×1\.31\\timesthe largest design effect measured \(1\.1481\.148\)\. The movement is larger at smallernn—0\.350\.35pp atn=6000n=6000against2\.692\.69pp atn=717n=717— consistent with Clopper–Pearson width shrinking asnngrows, though these rows varymmand the moment family as well asnnand so do not isolate a rate\. The certificate is not sensitive to the i\.i\.d\. assumption at the magnitudes the data support\.
### 10\.8Carried\-forward single\-agent evidence
E1–E6 measure composition\. They take for granted that a behavioral contract can be specified, enforced, and measured on a*single*agent at acceptable cost, which is the result established in the v1 framework paper\([Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\)and not re\-run here\. We restate it because the compositional claims are meaningless without it: a dependence\-aware bound over components whose individual contracts could not be enforced would certify nothing\.
The v1 evaluation ran 1980 sessions overAgentContract\-Bench, a benchmark of 200 scenarios across 7 models from 6 vendors\. Four results carry forward\.
Table 10:Single\-agent results carried forward from the v1 framework paper\([Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\): 1980 sessions, 200 scenarios, 7 models, 6 vendors\. These are not re\-measured in this work\.The effect sizes in the first row are large enough to warrant a caution rather than a boast\. Cohen’sddbetween6\.76\.7and33\.833\.8reflects a comparison in which the baseline detects approximately zero soft violations by construction — an uninstrumented agent has no mechanism to report a soft\-constraint deviation — so the contrast measures the value of instrumentation, not a narrow improvement over a competing detector\. We restate it in those terms here because the v1 abstract’s phrasing invites the stronger reading\.
Carrying prior\-version evidence forward rather than re\-running it is deliberate and bounded:[Table10](https://arxiv.org/html/2608.12895#S10.T10)is cited, not claimed, and no result in[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)–[Section10\.7](https://arxiv.org/html/2608.12895#S10.SS7)depends on re\-deriving it\.
### 10\.9What the evaluation establishes
Table 11:Summary of the six experiments, their registered status, and their verdicts\. “Confirmatory” means the hypothesis, sample size, and analysis were fixed before outcomes existed\.Three claims survive the evaluation\.
Shared models produce correlated contract failures, and the effect is large\.Two instances of one model co\-fail on90\.0%90\.0\\%of the missions on which either fails\. The association is significant in every arm of every motif, and it replicates on two additional inference backends\. A composition bound that treats such components as independent is not making a small approximation\.
The operative variable is the model, not the vendor\.Substituting a different model reduces association significantly and consistently — six of six contrasts across three topologies\. Substituting a different*vendor*while the model already differs does not, and we report that null rather than presenting a three\-level ordering the data do not support\. For a practitioner choosing redundancy, this says the useful axis is model identity; sourcing two different models from one vendor is not obviously worse than sourcing from two\.
Certification can be tightened without assuming a copula\.Enriching the moment set lifts the certified floor by16\.616\.6percentage points on identical data, and the resulting certificate is insensitive to the i\.i\.d\. assumption at the magnitudes the data support\.
Two limits belong here rather than only in[Section11\.3](https://arxiv.org/html/2608.12895#S11.SS3)\. The substitution manipulation confounds model*identity*with model*capability*:ministral\-8bis both a different model and a weaker one, so E1’s arm effects cannot separate “different inductive bias” from “different competence\.”[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)shows this is what drives the marginal\-sensitive reversal, and the marginal\-free statistics are the correct instrument, but no design here isolates the two\. And every arm draws its own missions — identifiers embed the condition name, and the task is a fixed hash of the identifier — so all comparisons are between independent draws from a shared generator distribution, not between paired tasks\. Atn=6000n=6000per primary arm that is adequate; for the under\-run breadth arms of[Section10\.6](https://arxiv.org/html/2608.12895#S10.SS6)it is a real limit on precision\.
## 11Discussion
### 11\.1What the practitioner should do differently
Three consequences follow directly from[Section10](https://arxiv.org/html/2608.12895#S10)\.
Do not multiply reliabilities across components that share a model\.[Proposition4\.1](https://arxiv.org/html/2608.12895#S4.Thmproposition1)makes the error signed, and[Section10\.2](https://arxiv.org/html/2608.12895#S10.SS2)makes it large: atϕ=0\.92\\phi=0\.92the joint\-failure probability of a two\-agent series is far above the independent product\. A pipeline certified by multiplication is certified against a dependence structure that the data reject\.
Choose redundancy on model identity, not vendor\.Six of six contrasts show that substituting a different model reduces association significantly\. Thesame\_vendorversusdifferent\_vendorcontrast does not replicate in any motif\. We state the negative result and stop there: this evaluation substituted exactly one same\-vendor model and one cross\-vendor model, so it cannot support a procurement policy about vendor diversity in general\. It supports only the narrow claim that vendor identity did not predict failure correlation here once model identity already differed\.
Report the moment set alongside the floor\.A certified floor is meaningless without the family it was computed over: the same data yield0\.24550\.2455atJ=10J\{=\}10and0\.41160\.4116atJ=14J\{=\}14\([Section10\.3](https://arxiv.org/html/2608.12895#S10.SS3)\)\. Enriching the family is the cheapest available tightening, and[Proposition6\.2](https://arxiv.org/html/2608.12895#S6.Thmproposition2)shows how to do it without losing monotonicity\.
### 11\.2Limitations
The LP is exponential inmm\.[Equation18](https://arxiv.org/html/2608.12895#S6.E18)has2m2^\{m\}variables\. Every motif here hasm≤4m\\leq 4and the program solves in milliseconds, but a fifty\-stage pipeline is out of reach and would require certifying sub\-compositions and composing certificates, at a cost in tightness we have not quantified\.
Aggregators are assumed deterministic\.Merge and quorum nodes are code, not models, so they carry no contract and contribute no failure correlation\. A system whose aggregator is itself an LLM — an increasingly common design — introduces a third correlated component that this analysis does not model\.
Drift dynamics are carried, not tested\.[Theorem3\.1](https://arxiv.org/html/2608.12895#S3.Thmtheorem1)is stated and proved, and it is inherited from v1 rather than re\-measured here\. No experiment in[Section10](https://arxiv.org/html/2608.12895#S10)estimatesα\\alpha,γ\\gamma, orσ\\sigmafrom data, so the design criterionγ\>α\\gamma\>\\alphais a theoretical statement in this paper\.
The certificate is per\-distribution\.By[Definition5\.1](https://arxiv.org/html/2608.12895#S5.Thmdefinition1)a guarantee holds over the mission distribution, model versions, and topology under which it was obtained\. Model upgrades invalidate it\. Nothing here says how quickly a certificate decays as a deployment drifts from its certification conditions\.
### 11\.3Threats to Validity
##### Model identity is confounded with model capability\.
The substitution replacesmistral\-small\-24bwithministral\-8borgemma\-3\-12b\-it, which differ in both training lineage and competence\. An arm effect could therefore reflect differing failure*rates*rather than differing failure*modes*\. This is the most serious threat to E1’s interpretation, and it is the mechanism behind the reversal diagnosed in[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)\. Two things limit it\. The marginal\-free statistics are invariant to the marginal rates by construction, and the qualitative conclusion —same\_modeldominates — holds on those\. And the negative control of[Section10\.2\.4](https://arxiv.org/html/2608.12895#S10.SS2.SSS4)holds capability fixed while the arm label varies, returning fifteen nulls\. Neither device separates identity from capability outright; a design that did would require two distinct models matched on failure rate, which we did not run\.
##### Arms do not share missions\.
Mission identifiers embed the condition name and the task is a fixed hash of the identifier, so each arm draws its own tasks from the shared generator distribution\. Comparisons are between independent draws, not paired\. Atn=6000n=6000per primary arm the resulting variance is absorbed by the bootstrap; for the under\-run breadth arms of[Section10\.6](https://arxiv.org/html/2608.12895#S10.SS6)it is a material limit, and the Meta arm’s single co\-failure makes its overlap estimate uninformative\.
##### All arms ran inside one window\.
The full campaign executed on a single day across roughly four hours\. A provider\-side change — a model version rollout, a serving\-stack change, a load excursion — occurring mid\-campaign would confound arms with time, since arms ran sequentially rather than interleaved\. The negative control again bounds this: a pair whose composition never changed shows no arm effect, which a global temporal shift would have produced\. We nonetheless regard sequential execution as a design weakness and would interleave arms in a replication\.
##### Contracts and gold code share an author\.
The contracts under test and the deterministic scoring code were written by the same authors\. A contract inadvertently written to match what the scorer checks would inflate measured compliance uniformly\. It would not, however, generate the*differential*co\-failure structure E1 reports, since the same contracts and scorer are used in all arms; the threat is to absolute compliance levels, which this paper does not claim, rather than to the arm contrasts, which it does\.
##### Two task domains only\.
Missions come from six generators spanning retail and financial workflows, scored deterministically\. Whether the dependence magnitudes transfer to code generation, research synthesis, or open\-ended dialogue is untested\. We expect the direction to transfer — shared weights imply shared failure modes on any distribution — and make no claim about the magnitudes\.
##### One statistic was chosen after seeing the data\.
The preregistration fixedτa\\tau\_\{a\}as the primary estimator\. The decision to lead H2 on the log odds ratio was made after observing the reversal described in[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)\. Choosing an estimator after seeing outcomes is a researcher degree of freedom of exactly the kind that inflates false positives\([Simmons et al\. 2011](https://arxiv.org/html/2608.12895#bib.bib54);[Gelman and Loken 2014](https://arxiv.org/html/2608.12895#bib.bib19)\), so we disclose it rather than presentlogOR\\log\\mathrm\{OR\}as preregistered\. Three considerations bear on it: the preregistered falsification test was evaluated onτa\\tau\_\{a\}as registered and passes \([Section10\.2\.2](https://arxiv.org/html/2608.12895#S10.SS2.SSS2)\); all five statistics are reported for every arm and contrast, so nothing is hidden by the choice; and the reason for preferring a marginal\-free statistic is structural rather than result\-dependent, sinceτa\\tau\_\{a\}’s marginal bound is a property of its definition \([Definition3\.17](https://arxiv.org/html/2608.12895#S3.Thmdefinition17)\)\. A reader who insists on the registered statistic alone reaches the same verdict on H1 and on thesame\_modelcontrasts, and a*stronger*rejection of the three\-level ordering\.
### 11\.4Broader impact
The failure this paper documents is quiet\. A composition bound computed under independence returns a number, the number looks reassuring, and nothing in the pipeline signals that the assumption behind it is false\. Systems certified this way will be deployed in exactly the settings — financial screening, clinical triage, content moderation — where correlated failure is most costly, because those are the settings where redundant review is mandated\. Making the dependence measurable, and making the certificate degrade honestly when it is unmeasured, is the contribution we regard as most consequential\.
There is a countervailing risk\. A certificate is a number that invites over\-trust, and[Definition5\.1](https://arxiv.org/html/2608.12895#S5.Thmdefinition1)restricts it narrowly: to one mission distribution, one set of model versions, one topology\. A team that treats a Tier\-1 floor as a property of their system rather than of their measurement will be wrong the first time they upgrade a model\. We have made the scope explicit in the artifact rather than only in this paper, and we regard tooling that silently carries a certificate across a model upgrade as an anti\-pattern this work should not be used to justify\.
## 12Contributions and their standing
Each contribution below is labelled*lead*\(we know of no prior work establishing it\),*co\-lead*\(concurrent work reaches a related result by different means; we claim the specific form given here\), or*supporting*\(the technique is established and the application is what is new\)\.
##### C1 — Preregistered, manipulated measurement of cross\-agent failure dependence\. \(Co\-lead\.\)
That agents sharing a base model can fail together is established:[McDonnell et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib38)quantify correlated failure in a three\-learner triage ensemble,[Huang et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib28)observe shared blind spots within model families, and[Bengio et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib4)names it as a concern\. We do not claim the phenomenon\. What is new is the design: model sharing as a*manipulated*variable at three levels, preregistered before outcomes existed, with deterministic gold scoring and no LLM judge\. The registration is confirmatory for theseries2motif over 18,000 missions; the two further topologies are secondary replications within a 30,820\-mission campaign and we label them so\. That converts an observation about particular systems into a contrast attributable to substitution, and it is what licenses the causal reading of[Table3](https://arxiv.org/html/2608.12895#S10.T3)\. It also yields the finding we did not expect: the ordering is real at the model level and*absent*at the vendor level \([Section10\.2\.2](https://arxiv.org/html/2608.12895#S10.SS2.SSS2)\)\.
##### C2 — A finite\-sample, copula\-agnostic reliability certificate for composed agent pipelines\. \(Lead\.\)
[Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)combines an exact Clopper–Pearson moment box with the extremal LP of[Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)to give a\(1−ηconf\)\(1\-\\eta\_\{\\mathrm\{conf\}\}\)lower bound on composed all\-success reliability that assumes no dependence structure whatever\. Neither the correlated\-failure literature above nor the agent\-evaluation literature produces a bound of any kind; measuring that failures correlate does not tell an operator what they may certify\. The LP is classical\([Boole 1854](https://arxiv.org/html/2608.12895#bib.bib9);[Hailperin 1965](https://arxiv.org/html/2608.12895#bib.bib21);[Bertsimas and Popescu 2005](https://arxiv.org/html/2608.12895#bib.bib6)\); its use as a finite\-sample agent certificate, and[Proposition6\.2](https://arxiv.org/html/2608.12895#S6.Thmproposition2)’s allocation that makes the floor monotone in the moment family by construction, are ours\.
##### C3 — Coverage collapse of model\-based reliability floors\. \(Supporting\.\)
[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)proves that a bootstrap lower bound on a*fitted*dependence model’s functional loses coverage of the true reliability asn→∞n\\to\\infty, because the identification gap isO\(1\)O\(1\)while the bootstrap haircut isO\(n−1/2\)O\(n^\{\-1/2\}\)\. The practical content is uncomfortable and specific: past some finite sample size such a certificate is wrong with probability approaching one, and nothing in the interval signals it\. This is our reason for reporting the Gaussian floor as a diagnostic and never as a guarantee \([Remark5\.2](https://arxiv.org/html/2608.12895#S5.Thmremark2)\), and[Section10\.4](https://arxiv.org/html/2608.12895#S10.SS4)exhibits the collapse on a witness whose misspecification is invisible to marginal and pairwise data\.
##### C4 — Anytime\-valid certification of graph reliability\. \(Co\-lead\.\)
The betting e\-process and its validity under optional stopping are established\([Ville 1939](https://arxiv.org/html/2608.12895#bib.bib57);[Shafer and Vovk 2019](https://arxiv.org/html/2608.12895#bib.bib52);[Shafer 2021](https://arxiv.org/html/2608.12895#bib.bib51);[Ramdas et al\. 2023](https://arxiv.org/html/2608.12895#bib.bib48);[Waudby\-Smith and Ramdas 2024](https://arxiv.org/html/2608.12895#bib.bib60);[Grünwald et al\. 2024](https://arxiv.org/html/2608.12895#bib.bib20)\)\. Applying them to composed agent reliability is what we claim, and the payoff is specific to this paper’s problem: the null in[Definition7\.1](https://arxiv.org/html/2608.12895#S7.Thmdefinition1)constrains only a conditional mean, so the sequential certificate requires no independence assumption across missions, components, or shared\-model shocks — it is immune to precisely the failure that[Section4](https://arxiv.org/html/2608.12895#S4)documents for the composition bound\.[Proposition7\.1](https://arxiv.org/html/2608.12895#S7.Thmproposition1)records that the SPRT is recovered exactly, so the anytime guarantee is free against a known alternative\.
##### C5 — Marginal sensitivity as a measurement hazard\. \(Supporting\.\)
That Jaccard,ϕ\\phi, andτa\\tau\_\{a\}are bounded by the marginals is textbook\. The contribution is the demonstration that this is decisive in practice:[Section10\.2\.3](https://arxiv.org/html/2608.12895#S10.SS2.SSS3)shows a significant*reversal*of the apparent condition ordering driven purely by a difference in marginal failure rates, and[Section10\.6](https://arxiv.org/html/2608.12895#S10.SS6)replicates the mechanism on two further backends where the marginal\-free statistic rises while the marginal\-sensitive one collapses\. The consequence generalises beyond this paper: correlated\-failure results that lead withϕ\\phi— including[McDonnell et al\. 2026](https://arxiv.org/html/2608.12895#bib.bib38)— are sensitive to marginal imbalance between the compared agents, and reporting a marginal\-free companion statistic is cheap insurance\.
##### C6 — An internal negative control, and the artifact\. \(Supporting\.\)
[Section10\.2\.4](https://arxiv.org/html/2608.12895#S10.SS2.SSS4)identifies a pair whose composition is held constant across all three arms\. Fifteen of fifteen contrasts are non\-significant with point estimates near zero, and in every arm the pair’s dependence sits where a same\-model pair should sit\. Two things follow\. The estimator returns the same answer for the same composition regardless of what the rest of the graph is doing, which is convergent validity\. And the leading alternative explanations for the E1 effect — task\-mix drift between runs, ordering, wall\-clock variation, provider\-side variation — would each have moved this pair, and none did\. It is not an equivalence test: it bounds no confound against a prespecified margin, and a confound acting only on the substituted worker would leave the pair untouched\. The design was nonetheless falsifiable at this point, and was not falsified\. The contracts, generators, scoring code, analysis scripts, and preregistration are released \([AppendixC](https://arxiv.org/html/2608.12895#A3)\)\.
### 12\.1What we do not claim
We do not claim that shared models cause correlated failures for the first time;[Section2\.3](https://arxiv.org/html/2608.12895#S2.SS3)attributes that\. We do not claim a three\-level sharing ordering: the vendor\-level contrast fails to replicate and we report the null\. We do not claim the drift dynamics of[Theorem3\.1](https://arxiv.org/html/2608.12895#S3.Thmtheorem1)are validated here — they are carried from v1 and not re\-measured\. We do not claim the certificate scales past smallmm;[Remark6\.1](https://arxiv.org/html/2608.12895#S6.Thmremark1)states the exponential cost\. And we do not claim frontier\-model results: the evaluation runs mid\-sized open models, and whether the dependence magnitudes hold at the frontier is untested\.
## 13Future Work
Five directions follow directly from what this paper could not settle\.
##### Separating model identity from model capability\.
The confound of[Section11\.3](https://arxiv.org/html/2608.12895#S11.SS3)is the most consequential open problem for C1\. The design that resolves it pairs two models matched on marginal failure rate but differing in training lineage — for instance, two distinct architectures tuned to equal accuracy on the mission distribution\. Any residual association is then attributable to shared inductive bias rather than to shared competence\. We did not run it: matching on failure rate requires a prior calibration pass over the candidate models, which the registration did not provide for and which cannot be added after outcomes exist without forfeiting the confirmatory status of the arms it would inform\.
##### Certifying large compositions\.
[Equation18](https://arxiv.org/html/2608.12895#S6.E18)is exponential inmm\([Remark6\.1](https://arxiv.org/html/2608.12895#S6.Thmremark1)\)\. Two routes are worth pursuing: certifying sub\-compositions and composing the certificates, which is sound but lossy and whose loss we have not quantified; and column generation over the2m2^\{m\}cells, which would keep sharpness for structured moment families\. Neither is attempted here\.
##### Model\-agent aggregators\.
Merge and quorum nodes are deterministic code throughout this evaluation, so they contribute no correlated failure\. Systems in which the aggregator is itself an LLM — a judge, a router, a critic — introduce a third dependent component whose failures may correlate with the workers it is adjudicating\. That is the common design in deployed systems and the analysis here does not cover it\.
##### Certificate decay\.
By[Definition5\.1](https://arxiv.org/html/2608.12895#S5.Thmdefinition1)a certificate holds over one mission distribution and one set of model versions\. Nothing here says how fast it decays as a deployment drifts, which is the quantity an operator actually needs to know when deciding how often to recertify\. The anytime\-valid machinery of[Section7](https://arxiv.org/html/2608.12895#S7)is the natural tool: a certificate that is continuously re\-earned rather than periodically re\-issued\.
##### Broader task distributions\.
Missions here span retail and financial workflows with deterministic gold scoring, which is what makes the measurement trustworthy and also what limits it\. Whether the dependence magnitudes transfer to code generation, research synthesis, or open\-ended dialogue — domains where scoring is itself contested — is untested\. We expect the direction to hold and make no claim about the magnitudes\.
## 14Conclusion
Compositional reliability bounds for multi\-agent systems multiply component reliabilities, and that step assumes the components fail independently\. A preregistered confirmatory evaluation of 18,000 two\-agent\-handoff missions, within a 30,820\-mission campaign, measures what the assumption costs\. Two instances of one model co\-fail on90\.0%90\.0\\%of the missions on which either fails \(logOR=6\.66\\log\\mathrm\{OR\}=6\.66, 95% CI\[6\.38,7\.00\]\[6\.38,7\.00\]\)\. Substituting a different model reduces the association significantly in the confirmatory motif and in both secondary topologies\. Substituting a different vendor, with the model already different, does not — and we report that null rather than the tidier three\-level ordering we registered\.
The consequence is not that the bound is slightly loose\. By[Proposition4\.1](https://arxiv.org/html/2608.12895#S4.Thmproposition1)the error is signed, and it runs against the operator: positive dependence inflates joint failure above the independent product, so a redundant design is over\-credited exactly when its components share a model\. Dropping the assumption entirely does not help, because the assumption\-free floor is zero whenever mean component reliability falls below1−1/m1\-1/m\. Fitting a dependence model is worse than either:[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)proves its coverage of the true reliability tends to zero as the sample grows, with no visible symptom\.
What works is to constrain the joint with measured co\-execution moments and optimise over everything consistent with them\. That certificate is sound without any dependence assumption \([Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)\), sharp for the information supplied \([Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)\), and tightens as the moment family grows: enriching from ten functionals to fourteen narrows the identified interval by85\.7%85\.7\\%and lifts the certified floor from0\.24550\.2455to0\.41160\.4116on identical data\. It is insensitive to the i\.i\.d\. assumption at the magnitudes the data support — at a conceded design effect1\.311\.31times the largest measured, the floor moves at most2\.692\.69percentage points\. And where a team wants to watch a dashboard and stop when it looks good, the anytime\-valid certificate holds its type\-I error at0\.04710\.0471or below across every admissible betting fraction\.
None of this makes correlated failure go away\. It makes it measurable, and it makes the resulting guarantee degrade honestly instead of silently\. That distinction is the whole of the contribution: a system whose certificate is merely loose can be shipped with known margin, whereas a system whose certificate is confidently wrong cannot be shipped safely at all — and the second is what the independence assumption produces\.
## Appendix AFull Proofs
Every numbered result in[Section4](https://arxiv.org/html/2608.12895#S4)–[Section7](https://arxiv.org/html/2608.12895#S7)is proved here in full\. Nothing is left as a sketch\.
### A\.1Proof of[Theorem3\.1](https://arxiv.org/html/2608.12895#S3.Thmtheorem1)
###### Proof\.
Pute\(t\)=D\(t\)−α/γe\(t\)=D\(t\)\-\\alpha/\\gamma\. Substituting into[Definition3\.8](https://arxiv.org/html/2608.12895#S3.Thmdefinition8),
de=dD=\(α−γD\)dt\+σdW=−γedt\+σdW,de=dD=\\bigl\(\\alpha\-\\gamma D\\bigr\)dt\+\\sigma\\,dW=\-\\gamma e\\,dt\+\\sigma\\,dW,soeeis an Ornstein–Uhlenbeck process with mean reverting to zero\. Applying Itô’s formula tof\(t,e\)=eγtef\(t,e\)=e^\{\\gamma t\}egivesd\(eγte\)=σeγtdWd\(e^\{\\gamma t\}e\)=\\sigma e^\{\\gamma t\}dW, and integrating from00tott,
e\(t\)=e\(0\)e−γt\+σ∫0te−γ\(t−s\)𝑑W\(s\)\.e\(t\)=e\(0\)e^\{\-\\gamma t\}\+\\sigma\\int\_\{0\}^\{t\}e^\{\-\\gamma\(t\-s\)\}\\,dW\(s\)\.\(29\)
*\(v\)\.*The Itô integral in[Equation29](https://arxiv.org/html/2608.12895#A1.E29)has mean zero and, by the Itô isometry, varianceσ2∫0te−2γ\(t−s\)𝑑s=σ22γ\(1−e−2γt\)\\sigma^\{2\}\\int\_\{0\}^\{t\}e^\{\-2\\gamma\(t\-s\)\}ds=\\frac\{\\sigma^\{2\}\}\{2\\gamma\}\(1\-e^\{\-2\\gamma t\}\)\. The integrand is deterministic and adapted, and by hypothesise\(0\)e\(0\)is independent of the driving Wiener process with𝔼\[e\(0\)2\]<∞\\mathbb\{E\}\[e\(0\)^\{2\}\]<\\infty; the cross term therefore vanishes,𝔼\[e\(0\)∫0te−γ\(t−s\)𝑑W\(s\)\]=𝔼\[e\(0\)\]⋅0=0\\mathbb\{E\}\[e\(0\)\\int\_\{0\}^\{t\}e^\{\-\\gamma\(t\-s\)\}dW\(s\)\]=\\mathbb\{E\}\[e\(0\)\]\\cdot 0=0, and squaring[Equation29](https://arxiv.org/html/2608.12895#A1.E29)gives
𝔼\[e\(t\)2\]=𝔼\[e\(0\)2\]e−2γt\+σ22γ\(1−e−2γt\)\.\\mathbb\{E\}\[e\(t\)^\{2\}\]=\\mathbb\{E\}\[e\(0\)^\{2\}\]\\,e^\{\-2\\gamma t\}\+\\frac\{\\sigma^\{2\}\}\{2\\gamma\}\\bigl\(1\-e^\{\-2\\gamma t\}\\bigr\)\.Noteeeis mean\-zero only when𝔼\[e\(0\)\]=0\\mathbb\{E\}\[e\(0\)\]=0; in general𝔼\[e\(t\)\]=𝔼\[e\(0\)\]e−γt\\mathbb\{E\}\[e\(t\)\]=\\mathbb\{E\}\[e\(0\)\]e^\{\-\\gamma t\}, which decays to zero at rateγ\\gammaand does not affect \(i\)–\(iv\)\.
*\(i\), \(iii\)\.*The stochastic integral in[Equation29](https://arxiv.org/html/2608.12895#A1.E29)is a Wiener integral of a deterministic kernel, hence Gaussian\. Lettingt→∞t\\to\\inftyin[Equation29](https://arxiv.org/html/2608.12895#A1.E29), the transiente\(0\)e−γt→0e\(0\)e^\{\-\\gamma t\}\\to 0and the variance converges toσ2/\(2γ\)\\sigma^\{2\}/\(2\\gamma\), soe\(t\)⇒𝒩\(0,σ2/\(2γ\)\)e\(t\)\\Rightarrow\\mathcal\{N\}\(0,\\sigma^\{2\}/\(2\\gamma\)\)\.
For stationarity, takee\(0\)∼𝒩\(0,σ2/\(2γ\)\)e\(0\)\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}/\(2\\gamma\)\)*independent of*WW— the integrand in[Equation29](https://arxiv.org/html/2608.12895#A1.E29)is adapted and the increments\{W\(s\)−W\(0\)\}s\>0\\\{W\(s\)\-W\(0\)\\\}\_\{s\>0\}are independent ofℱ0\\mathcal\{F\}\_\{0\}, so the two terms of[Equation29](https://arxiv.org/html/2608.12895#A1.E29)are independent and their variances add:
Var\(e\(t\)\)=e−2γtσ22γ\+σ22γ\(1−e−2γt\)=σ22γ\\operatorname\{Var\}\(e\(t\)\)=e^\{\-2\\gamma t\}\\frac\{\\sigma^\{2\}\}\{2\\gamma\}\+\\frac\{\\sigma^\{2\}\}\{2\\gamma\}\\bigl\(1\-e^\{\-2\\gamma t\}\\bigr\)=\\frac\{\\sigma^\{2\}\}\{2\\gamma\}for everytt, while the mean stays00\. Hence the law is unchanged inttandπD=𝒩\(α/γ,σ2/\(2γ\)\)\\pi\_\{D\}=\\mathcal\{N\}\(\\alpha/\\gamma,\\sigma^\{2\}/\(2\\gamma\)\)is stationary\.
*\(ii\)\.*Immediate from \(i\):𝔼π\[D\]=α/γ\\mathbb\{E\}\_\{\\pi\}\[D\]=\\alpha/\\gamma, which is<1<1exactly whenγ\>α\\gamma\>\\alpha\.
*\(iv\)\.*Underπ\\pi,Z=\(D−α/γ\)/σ2/\(2γ\)Z=\(D\-\\alpha/\\gamma\)/\\sqrt\{\\sigma^\{2\}/\(2\\gamma\)\}is standard normal, and the standard Gaussian tail boundℙ\(Z\>z\)≤e−z2/2\\mathbb\{P\}\(Z\>z\)\\leq e^\{\-z^\{2\}/2\}forz\>0z\>0gives, withz=η/σ2/\(2γ\)z=\\eta/\\sqrt\{\\sigma^\{2\}/\(2\\gamma\)\},
ℙπ\(D\>α/γ\+η\)≤exp\(−η22⋅2γσ2\)=exp\(−γη2σ2\)\.∎\\mathbb\{P\}\_\{\\pi\}\\bigl\(D\>\\alpha/\\gamma\+\\eta\\bigr\)\\leq\\exp\\\!\\Bigl\(\-\\frac\{\\eta^\{2\}\}\{2\}\\cdot\\frac\{2\\gamma\}\{\\sigma^\{2\}\}\\Bigr\)=\\exp\\\!\\Bigl\(\-\\frac\{\\gamma\\eta^\{2\}\}\{\\sigma^\{2\}\}\\Bigr\)\.\\qed
### A\.2Proof of[Proposition3\.1](https://arxiv.org/html/2608.12895#S3.Thmproposition1)
###### Proof\.
Fix a turn with state–action pair\(st,at\)\(s\_\{t\},a\_\{t\}\)and letk=\|ℐ∪𝒢\|k=\|\\mathcal\{I\}\\cup\\mathcal\{G\}\|be the number of constraints and\|A\|\|A\|the action\-vocabulary size\.
*Constraint evaluation\.*By[Definition3\.2](https://arxiv.org/html/2608.12895#S3.Thmdefinition2)the evaluator is stateless: each constraintccis a predicate on\(st,at\)\(s\_\{t\},a\_\{t\}\)evaluated independently of other constraints and of prior turns\. Evaluating all of them and forming the two ratios iskkpredicate evaluations plusO\(1\)O\(1\)arithmetic, henceO\(k\)O\(k\)under the standard assumption that each predicate costsO\(1\)O\(1\)\.
*Drift update\.*By[Definition3\.5](https://arxiv.org/html/2608.12895#S3.Thmdefinition5)the distributional term requiresJSD\(Pobs∥Pref\)\\JSD\(P\_\{\\mathrm\{obs\}\}\\\|P\_\{\\mathrm\{ref\}\}\)over the action alphabet\. Updating the sliding\-window histogram on one new action isO\(1\)O\(1\)\(one increment, one decrement at the window boundary\), and evaluating the divergence is a sum over the support, henceO\(\|A\|\)O\(\|A\|\)\.
*Aggregation\.*By[Definition3\.4](https://arxiv.org/html/2608.12895#S3.Thmdefinition4), combining the two components is a fixed convex combination,O\(1\)O\(1\)\.
Summing, the per\-action cost isO\(k\)\+O\(\|A\|\)\+O\(1\)=O\(k\+\|A\|\)O\(k\)\+O\(\|A\|\)\+O\(1\)=O\(k\+\|A\|\)\. The measured constant is small: fork<100k<100and\|A\|<50\|A\|<50the v1 implementation records under 10 ms per action\. ∎
### A\.3Proof of[Proposition4\.1](https://arxiv.org/html/2608.12895#S4.Thmproposition1)
###### Proof\.
WriteTk=∏i=kmhiT\_\{k\}=\\prod\_\{i=k\}^\{m\}h\_\{i\}for1≤k≤m1\\leq k\\leq m, withTm\+1=1T\_\{m\+1\}=1, soYG=T1Y\_\{G\}=T\_\{1\}andℙ\(YG=1\)=𝔼\[T1\]\\mathbb\{P\}\(Y\_\{G\}=1\)=\\mathbb\{E\}\[T\_\{1\}\]sinceT1T\_\{1\}is a\{0,1\}\\\{0,1\\\}variable\. For eachkk,
𝔼\[Tk\]=𝔼\[hkTk\+1\]=Cov\(hk,Tk\+1\)\+pk𝔼\[Tk\+1\]\.\\mathbb\{E\}\[T\_\{k\}\]=\\mathbb\{E\}\[h\_\{k\}T\_\{k\+1\}\]=\\operatorname\{Cov\}\(h\_\{k\},T\_\{k\+1\}\)\+p\_\{k\}\\,\\mathbb\{E\}\[T\_\{k\+1\}\]\.Applying this identity atk=1k=1and then recursively substituting for𝔼\[Tk\+1\]\\mathbb\{E\}\[T\_\{k\+1\}\]gives, afterm−1m\-1steps,
𝔼\[T1\]=∑k=1m−1\(∏i=1k−1pi\)Cov\(hk,Tk\+1\)\+\(∏i=1m−1pi\)𝔼\[Tm\],\\mathbb\{E\}\[T\_\{1\}\]=\\sum\_\{k=1\}^\{m\-1\}\\Bigl\(\\prod\_\{i=1\}^\{k\-1\}p\_\{i\}\\Bigr\)\\operatorname\{Cov\}\(h\_\{k\},T\_\{k\+1\}\)\+\\Bigl\(\\prod\_\{i=1\}^\{m\-1\}p\_\{i\}\\Bigr\)\\mathbb\{E\}\[T\_\{m\}\],where the empty product atk=1k=1equals11\. SinceTm=hmT\_\{m\}=h\_\{m\}we have𝔼\[Tm\]=pm\\mathbb\{E\}\[T\_\{m\}\]=p\_\{m\}, so the final term is∏i=1mpi\\prod\_\{i=1\}^\{m\}p\_\{i\}\. Rearranging,
ℙ\(YG=1\)−∏i=1mpi=∑k=1m−1\(∏i=1k−1pi\)Cov\(hk,∏i=k\+1mhi\),\\mathbb\{P\}\(Y\_\{G\}=1\)\-\\prod\_\{i=1\}^\{m\}p\_\{i\}=\\sum\_\{k=1\}^\{m\-1\}\\Bigl\(\\prod\_\{i=1\}^\{k\-1\}p\_\{i\}\\Bigr\)\\operatorname\{Cov\}\\Bigl\(h\_\{k\},\\prod\_\{i=k\+1\}^\{m\}h\_\{i\}\\Bigr\),which is the stated expression with thek=1k=1term written out separately\. Form=2m=2the sum has the single termCov\(h1,h2\)\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\)\.
For the directional claim atm=2m=2:ℙ\(YG=0\)=1−p1p2−Cov\(h1,h2\)\\mathbb\{P\}\(Y\_\{G\}=0\)=1\-p\_\{1\}p\_\{2\}\-\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\), while the independence calculation returns1−p1p21\-p\_\{1\}p\_\{2\}\. Under positive dependenceCov\(h1,h2\)\>0\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\)\>0, so the true probability that the series fails is*smaller*than the independent calculation — the independence product is conservative for a series system, in both directions\. ∎
### A\.4Proof of[Corollary4\.1](https://arxiv.org/html/2608.12895#S4.Thmcorollary1)
###### Proof\.
LetFi=1−hiF\_\{i\}=1\-h\_\{i\}\. Expanding the covariance of the two failure indicators,
ℙ\(F1=F2=1\)=𝔼\[F1F2\]=Cov\(F1,F2\)\+𝔼\[F1\]𝔼\[F2\]=Cov\(F1,F2\)\+\(1−p1\)\(1−p2\)\.\\mathbb\{P\}\(F\_\{1\}=F\_\{2\}=1\)=\\mathbb\{E\}\[F\_\{1\}F\_\{2\}\]=\\operatorname\{Cov\}\(F\_\{1\},F\_\{2\}\)\+\\mathbb\{E\}\[F\_\{1\}\]\\mathbb\{E\}\[F\_\{2\}\]=\\operatorname\{Cov\}\(F\_\{1\},F\_\{2\}\)\+\(1\-p\_\{1\}\)\(1\-p\_\{2\}\)\.BecauseFi=1−hiF\_\{i\}=1\-h\_\{i\}is an affine function ofhih\_\{i\}with slope−1\-1, covariance is preserved:Cov\(F1,F2\)=Cov\(1−h1,1−h2\)=\(−1\)\(−1\)Cov\(h1,h2\)=Cov\(h1,h2\)\\operatorname\{Cov\}\(F\_\{1\},F\_\{2\}\)=\\operatorname\{Cov\}\(1\-h\_\{1\},1\-h\_\{2\}\)=\(\-1\)\(\-1\)\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\)=\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\)\. Hence under positive dependenceCov\(h1,h2\)\>0\\operatorname\{Cov\}\(h\_\{1\},h\_\{2\}\)\>0andℙ\(F1=F2=1\)\>\(1−p1\)\(1−p2\)\\mathbb\{P\}\(F\_\{1\}=F\_\{2\}=1\)\>\(1\-p\_\{1\}\)\(1\-p\_\{2\}\): the probability that both redundant paths fail together strictly exceeds the independence product\.
The two statements are therefore compatible and concern different events\. Series*any*\-failure is over\-estimated by the independence calculation; redundant*joint*failure is under\-estimated by it\. A redundant design is bought to control the second\. ∎
### A\.5Proof of[Theorem4\.1](https://arxiv.org/html/2608.12895#S4.Thmtheorem1)
###### Proof\.
LetAi=\{hi=1\}A\_\{i\}=\\\{h\_\{i\}=1\\\}, soℙ\(Ai\)=pi\\mathbb\{P\}\(A\_\{i\}\)=p\_\{i\}\.
*Upper bound\.*⋂iAi⊆Aj\\bigcap\_\{i\}A\_\{i\}\\subseteq A\_\{j\}for eachjj, henceℙ\(⋂iAi\)≤minjpj\\mathbb\{P\}\(\\bigcap\_\{i\}A\_\{i\}\)\\leq\\min\_\{j\}p\_\{j\}\.
*Lower bound\.*By De Morgan and the union bound,
ℙ\(⋂iAi\)=1−ℙ\(⋃iAic\)≥1−∑i=1m\(1−pi\)=∑i=1mpi−\(m−1\),\\mathbb\{P\}\\Bigl\(\\bigcap\_\{i\}A\_\{i\}\\Bigr\)=1\-\\mathbb\{P\}\\Bigl\(\\bigcup\_\{i\}A\_\{i\}^\{c\}\\Bigr\)\\geq 1\-\\sum\_\{i=1\}^\{m\}\\bigl\(1\-p\_\{i\}\\bigr\)=\\sum\_\{i=1\}^\{m\}p\_\{i\}\-\(m\-1\),and the probability is non\-negative, giving the maximum with00\.
*Attainment of the upper bound\.*LetU∼Unif\[0,1\)U\\sim\\mathrm\{Unif\}\[0,1\)and sethi=\[U<pi\]h\_\{i\}=\\mathbf\{1\}\\\!\\left\[U<p\_\{i\}\\right\]\. Eachhih\_\{i\}has the correct marginal, and⋂i\{hi=1\}=\{U<minipi\}\\bigcap\_\{i\}\\\{h\_\{i\}=1\\\}=\\\{U<\\min\_\{i\}p\_\{i\}\\\}, which has probabilityminipi\\min\_\{i\}p\_\{i\}\.
*Attainment of the lower bound\.*Lets=∑ipi−\(m−1\)s=\\sum\_\{i\}p\_\{i\}\-\(m\-1\)and note∑i\(1−pi\)=m−∑ipi=1−s\\sum\_\{i\}\(1\-p\_\{i\}\)=m\-\\sum\_\{i\}p\_\{i\}=1\-s\.
Ifs\>0s\>0then∑i\(1−pi\)=1−s<1\\sum\_\{i\}\(1\-p\_\{i\}\)=1\-s<1, so we may choose pairwise disjoint setsBi⊆\[0,1\)B\_\{i\}\\subseteq\[0,1\)with Lebesgue measure1−pi1\-p\_\{i\}\. Puthi=\[U∉Bi\]h\_\{i\}=\\mathbf\{1\}\\\!\\left\[U\\notin B\_\{i\}\\right\]\. Thenℙ\(hi=1\)=pi\\mathbb\{P\}\(h\_\{i\}=1\)=p\_\{i\}, and⋂i\{hi=1\}=\{U∉⋃iBi\}\\bigcap\_\{i\}\\\{h\_\{i\}=1\\\}=\\\{U\\notin\\bigcup\_\{i\}B\_\{i\}\\\}has measure1−∑i\(1−pi\)=s1\-\\sum\_\{i\}\(1\-p\_\{i\}\)=s, attaining the bound\.
Ifs≤0s\\leq 0then∑i\(1−pi\)=1−s≥1\\sum\_\{i\}\(1\-p\_\{i\}\)=1\-s\\geq 1, and we construct a cover explicitly rather than assert one\. Identify\[0,1\)\[0,1\)with the circleℝ/ℤ\\mathbb\{R\}/\\mathbb\{Z\}\. PutLi=1−piL\_\{i\}=1\-p\_\{i\},S0=0S\_\{0\}=0, andSk=∑i≤kLiS\_\{k\}=\\sum\_\{i\\leq k\}L\_\{i\}, and lay the arcs end to end with wrap\-around:
Bk=\[Sk−1mod1,\(Sk−1\+Lk\)mod1\)⊆ℝ/ℤ\.B\_\{k\}\\;=\\;\\bigl\[\\,S\_\{k\-1\}\\bmod 1,\\;\(S\_\{k\-1\}\+L\_\{k\}\)\\bmod 1\\,\\bigr\)\\subseteq\\mathbb\{R\}/\\mathbb\{Z\}\.EachBkB\_\{k\}has Lebesgue measureLk=1−pkL\_\{k\}=1\-p\_\{k\}, soℙ\(hk=1\)=pk\\mathbb\{P\}\(h\_\{k\}=1\)=p\_\{k\}as required\. The arcs are laid consecutively without gaps starting at00and have total length∑kLk=1−s≥1\\sum\_\{k\}L\_\{k\}=1\-s\\geq 1, so their union wraps at least once around the circle and⋃kBk=\[0,1\)\\bigcup\_\{k\}B\_\{k\}=\[0,1\)\. Hence⋂k\{hk=1\}=\{U∉⋃kBk\}=∅\\bigcap\_\{k\}\\\{h\_\{k\}=1\\\}=\\\{U\\notin\\bigcup\_\{k\}B\_\{k\}\\\}=\\emptysetandℙ\(⋂k\{hk=1\}\)=0=max\(0,s\)\\mathbb\{P\}\(\\bigcap\_\{k\}\\\{h\_\{k\}=1\\\}\)=0=\\max\(0,s\)\.
\(Fors\>0s\>0the same construction has total length1−s<11\-s<1, the arcs stay disjoint, no wrap occurs, and the uncovered remainder has measuress— which is the previous case\.\)
Both bounds are therefore attained by laws with the prescribed marginals, so neither can be improved using marginal information alone\. ∎
###### Proof of[Corollary4\.2](https://arxiv.org/html/2608.12895#S4.Thmcorollary2)\.
Immediate from[Theorem4\.1](https://arxiv.org/html/2608.12895#S4.Thmtheorem1): the lower bound is00iff∑ipi≤m−1\\sum\_\{i\}p\_\{i\}\\leq m\-1, i\.e\. iff the mean marginalp¯≤1−1/m\\bar\{p\}\\leq 1\-1/m\. Atm=4m=4,1−1/4=0\.751\-1/4=0\.75, sop¯=0\.75\\bar\{p\}=0\.75gives a floor of exactly00\. ∎
### A\.6Proof of[Theorem4\.2](https://arxiv.org/html/2608.12895#S4.Thmtheorem2)
###### Proof\.
WriteGn=Rℱ\(μ⋆\)−L^nG\_\{n\}=R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)\-\\hat\{L\}\_\{n\}for the bootstrap haircut\. By hypothesisGn=Op\(n−1/2\)G\_\{n\}=O\_\{p\}\(n^\{\-1/2\}\), i\.e\. the family\{nGn\}\\\{\\sqrt\{n\}\\,G\_\{n\}\\\}is bounded in probability: for everyε\>0\\varepsilon\>0there existMεM\_\{\\varepsilon\}andNεN\_\{\\varepsilon\}withℙ\(n\|Gn\|\>Mε\)<ε\\mathbb\{P\}\(\\sqrt\{n\}\\,\|G\_\{n\}\|\>M\_\{\\varepsilon\}\)<\\varepsilonfor alln≥Nεn\\geq N\_\{\\varepsilon\}\.
Coverage of the true value fails exactly whenL^n\>R⋆\\hat\{L\}\_\{n\}\>R^\{\\star\}, so
ℙ\(L^n≤R⋆\)=ℙ\(Rℱ\(μ⋆\)−Gn≤R⋆\)=ℙ\(Gn≥Rℱ\(μ⋆\)−R⋆\)=ℙ\(Gn≥Δ\),\\mathbb\{P\}\\bigl\(\\hat\{L\}\_\{n\}\\leq R^\{\\star\}\\bigr\)=\\mathbb\{P\}\\bigl\(R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)\-G\_\{n\}\\leq R^\{\\star\}\\bigr\)=\\mathbb\{P\}\\bigl\(G\_\{n\}\\geq R\_\{\\mathcal\{F\}\}\(\\mu^\{\\star\}\)\-R^\{\\star\}\\bigr\)=\\mathbb\{P\}\\bigl\(G\_\{n\}\\geq\\Delta\\bigr\),using[Equation15](https://arxiv.org/html/2608.12895#S4.E15)\. SinceΔ\>0\\Delta\>0is a fixed constant not depending onnn,
ℙ\(Gn≥Δ\)=ℙ\(nGn≥nΔ\)\.\\mathbb\{P\}\(G\_\{n\}\\geq\\Delta\)=\\mathbb\{P\}\\bigl\(\\sqrt\{n\}\\,G\_\{n\}\\geq\\sqrt\{n\}\\,\\Delta\\bigr\)\.Fixε\>0\\varepsilon\>0and takeMεM\_\{\\varepsilon\}as above\. For allnnlarge enough thatnΔ\>Mε\\sqrt\{n\}\\,\\Delta\>M\_\{\\varepsilon\}— which holds forn\>Mε2/Δ2n\>M\_\{\\varepsilon\}^\{2\}/\\Delta^\{2\}— we have
ℙ\(nGn≥nΔ\)≤ℙ\(n\|Gn\|\>Mε\)<ε\.\\mathbb\{P\}\\bigl\(\\sqrt\{n\}\\,G\_\{n\}\\geq\\sqrt\{n\}\\,\\Delta\\bigr\)\\leq\\mathbb\{P\}\\bigl\(\\sqrt\{n\}\\,\|G\_\{n\}\|\>M\_\{\\varepsilon\}\\bigr\)<\\varepsilon\.Asε\\varepsilonwas arbitrary,ℙ\(L^n≤R⋆\)→0\\mathbb\{P\}\(\\hat\{L\}\_\{n\}\\leq R^\{\\star\}\)\\to 0\.
We claim convergence to zero and nothing stronger\. Tightness of\{nGn\}\\\{\\sqrt\{n\}\\,G\_\{n\}\\\}givesℙ\(Gn≥Δ\)→0\\mathbb\{P\}\(G\_\{n\}\\geq\\Delta\)\\to 0, but it does*not*entail that the sequence is monotone:GnG\_\{n\}may oscillate while remaining tight, so coverage need not decrease at everynn\. The interpretation that survives is the limit — past some finite sample size the interval sits aboveR⋆R^\{\\star\}with probability approaching one — together with the mechanism, that the haircut shrinks while the target does not move\. ∎
###### Proof of[Corollary4\.3](https://arxiv.org/html/2608.12895#S4.Thmcorollary3)\.
Letμ⋆\\mu^\{\\star\}be the true moments andQ⋆∈ℳ\(μ⋆\)Q^\{\\star\}\\in\\mathcal\{M\}\(\\mu^\{\\star\}\)the true law\. By[Definition4\.1](https://arxiv.org/html/2608.12895#S4.Thmdefinition1),R¯\(μ⋆\)\\underline\{R\}\(\\mu^\{\\star\}\)is an infimum over a set containingQ⋆Q^\{\\star\}, henceR¯\(μ⋆\)≤Q⋆\(⋀ihi=1\)=R⋆\\underline\{R\}\(\\mu^\{\\star\}\)\\leq Q^\{\\star\}\(\\bigwedge\_\{i\}h\_\{i\}=1\)=R^\{\\star\}\. The inequality uses only membership ofQ⋆Q^\{\\star\}inℳ\\mathcal\{M\}, which holds by definition ofμ⋆\\mu^\{\\star\}and requires no parametric assumption\. Therefore[Equation15](https://arxiv.org/html/2608.12895#S4.E15)cannot hold withRℱR\_\{\\mathcal\{F\}\}replaced byR¯\\underline\{R\}, and the argument of[SectionA\.6](https://arxiv.org/html/2608.12895#A1.SS6)does not apply\. ∎
### A\.7Proof of[Theorem5\.1](https://arxiv.org/html/2608.12895#S5.Thmtheorem1)
###### Proof\.
LetX=∑r=1n\[YG\(r\)=1\]∼Bin\(n,θ\)X=\\sum\_\{r=1\}^\{n\}\\mathbf\{1\}\\\!\\left\[Y\_\{G\}^\{\(r\)\}=1\\right\]\\sim\\mathrm\{Bin\}\(n,\\theta\)withθ=ℙ\(YG=1\)\\theta=\\mathbb\{P\}\(Y\_\{G\}=1\), which is the exact law ofXXwhen missions are i\.i\.d\. andYGY\_\{G\}is observed on each\. The Clopper–Pearson lower limit is
L^0\(X\)=\{0,X=0,inf\{θ′∈\[0,1\]:ℙθ′\(X′≥X\)\>ηconf\},X≥1,\\hat\{L\}\_\{0\}\(X\)=\\begin\{cases\}0,&X=0,\\\\ \\inf\\bigl\\\{\\theta^\{\\prime\}\\in\[0,1\]:\\mathbb\{P\}\_\{\\theta^\{\\prime\}\}\(X^\{\\prime\}\\geq X\)\>\\eta\_\{\\mathrm\{conf\}\}\\bigr\\\},&X\\geq 1,\\end\{cases\}whereX′∼Bin\(n,θ′\)X^\{\\prime\}\\sim\\mathrm\{Bin\}\(n,\\theta^\{\\prime\}\)\. For fixedx≥1x\\geq 1the mapθ′↦ℙθ′\(X′≥x\)\\theta^\{\\prime\}\\mapsto\\mathbb\{P\}\_\{\\theta^\{\\prime\}\}\(X^\{\\prime\}\\geq x\)is continuous and strictly increasing on\(0,1\)\(0,1\), soL^0\(x\)\\hat\{L\}\_\{0\}\(x\)is the unique root ofℙθ′\(X′≥x\)=ηconf\\mathbb\{P\}\_\{\\theta^\{\\prime\}\}\(X^\{\\prime\}\\geq x\)=\\eta\_\{\\mathrm\{conf\}\}andL^0\(x\)\>θ\\hat\{L\}\_\{0\}\(x\)\>\\thetaholds iffℙθ\(X′≥x\)<ηconf\\mathbb\{P\}\_\{\\theta\}\(X^\{\\prime\}\\geq x\)<\\eta\_\{\\mathrm\{conf\}\}\.
Letx⋆=min\{x:ℙθ\(X≥x\)<ηconf\}x^\{\\star\}=\\min\\\{x:\\mathbb\{P\}\_\{\\theta\}\(X\\geq x\)<\\eta\_\{\\mathrm\{conf\}\}\\\}\(withx⋆=n\+1x^\{\\star\}=n\+1if no suchxxexists\)\. Then\{L^0\>θ\}=\{X≥x⋆\}\\\{\\hat\{L\}\_\{0\}\>\\theta\\\}=\\\{X\\geq x^\{\\star\}\\\}and, by minimality ofx⋆x^\{\\star\},
ℙθ\(L^0\>θ\)=ℙθ\(X≥x⋆\)<ηconf\.\\mathbb\{P\}\_\{\\theta\}\\bigl\(\\hat\{L\}\_\{0\}\>\\theta\\bigr\)=\\mathbb\{P\}\_\{\\theta\}\(X\\geq x^\{\\star\}\)<\\eta\_\{\\mathrm\{conf\}\}\.Henceℙθ\(L^0≤θ\)≥1−ηconf\\mathbb\{P\}\_\{\\theta\}\(\\hat\{L\}\_\{0\}\\leq\\theta\)\\geq 1\-\\eta\_\{\\mathrm\{conf\}\}for everynnand everyθ\\theta, with no asymptotic approximation\. The bound is conservative rather than exact becauseXXis discrete\. ∎
### A\.8Proof of[Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)
###### Proof\.
For eachS∈𝒥S\\in\\mathcal\{J\}letISI\_\{S\}be the two\-sided Clopper–Pearson interval forμS=Q⋆\(⋀i∈Shi=1\)\\mu\_\{S\}=Q^\{\\star\}\(\\bigwedge\_\{i\\in S\}h\_\{i\}=1\)built from thenni\.i\.d\. indicator observations\[⋀i∈Shi\(r\)=1\]\\mathbf\{1\}\\\!\\left\[\\bigwedge\_\{i\\in S\}h\_\{i\}^\{\(r\)\}=1\\right\], each tail at levelηconf/\(2J\)\\eta\_\{\\mathrm\{conf\}\}/\(2J\)\. By the argument of[SectionA\.7](https://arxiv.org/html/2608.12895#A1.SS7)applied to each tail,ℙ\(μS∉IS\)≤2⋅ηconf/\(2J\)=ηconf/J\\mathbb\{P\}\(\\mu\_\{S\}\\notin I\_\{S\}\)\\leq 2\\cdot\\eta\_\{\\mathrm\{conf\}\}/\(2J\)=\\eta\_\{\\mathrm\{conf\}\}/J\.
Letℰ=⋂S∈𝒥\{μS∈IS\}\\mathcal\{E\}=\\bigcap\_\{S\\in\\mathcal\{J\}\}\\\{\\mu\_\{S\}\\in I\_\{S\}\\\}be the event that the whole box covers\. By the union bound over theJJmoments,
ℙ\(ℰc\)≤∑S∈𝒥ℙ\(μS∉IS\)≤J⋅ηconfJ=ηconf\.\\mathbb\{P\}\(\\mathcal\{E\}^\{c\}\)\\leq\\sum\_\{S\\in\\mathcal\{J\}\}\\mathbb\{P\}\(\\mu\_\{S\}\\notin I\_\{S\}\)\\leq J\\cdot\\frac\{\\eta\_\{\\mathrm\{conf\}\}\}\{J\}=\\eta\_\{\\mathrm\{conf\}\}\.
Onℰ\\mathcal\{E\}the true moment vector lies inB\(μ^\)B\(\\hat\{\\mu\}\), soQ⋆Q^\{\\star\}is a feasible point of the program in[Equation17](https://arxiv.org/html/2608.12895#S5.E17): it is a probability distribution on\{0,1\}m\\\{0,1\\\}^\{m\}whose𝒥\\mathcal\{J\}\-moments lie in the box\. SinceL^1\\hat\{L\}\_\{1\}is the*minimum*of the objective over the feasible set andQ⋆Q^\{\\star\}is feasible,
L^1≤Q⋆\(⋀ihi=1\)\.\\hat\{L\}\_\{1\}\\;\\leq\\;Q^\{\\star\}\\Bigl\(\\bigwedge\_\{i\}h\_\{i\}=1\\Bigr\)\.Thereforeℙ\(L^1≤Q⋆\(⋀ihi=1\)\)≥ℙ\(ℰ\)≥1−ηconf\\mathbb\{P\}\\bigl\(\\hat\{L\}\_\{1\}\\leq Q^\{\\star\}\(\\bigwedge\_\{i\}h\_\{i\}=1\)\\bigr\)\\geq\\mathbb\{P\}\(\\mathcal\{E\}\)\\geq 1\-\\eta\_\{\\mathrm\{conf\}\}\.
Note the argument never referenced the dependence structure ofQ⋆Q^\{\\star\}: it used only thatQ⋆Q^\{\\star\}satisfies its own moment constraints, which is a tautology\. That is the precise sense in which Tier 1 is copula\-agnostic\. ∎
### A\.9Proof of[Proposition5\.1](https://arxiv.org/html/2608.12895#S5.Thmproposition1)
###### Proof\.
Let𝒜0=\{i\.i\.d\. missions\}∪\{YGdirectly observed\}\\mathcal\{A\}\_\{0\}=\\\{\\text\{i\.i\.d\.\\ missions\}\\\}\\cup\\\{Y\_\{G\}\\text\{ directly observed\}\\\}and𝒜1=\{i\.i\.d\. missions\}\\mathcal\{A\}\_\{1\}=\\\{\\text\{i\.i\.d\.\\ missions\}\\\}be the assumption sets of[Theorem5\.1](https://arxiv.org/html/2608.12895#S5.Thmtheorem1)and[Theorem5\.2](https://arxiv.org/html/2608.12895#S5.Thmtheorem2)\. Then𝒜1⊆𝒜0\\mathcal\{A\}\_\{1\}\\subseteq\\mathcal\{A\}\_\{0\}\.
Suppose the flag asserting end\-to\-end execution is mis\-set\. There are two cases\. If it is set when it should not be, the certificate reportsL^0\\hat\{L\}\_\{0\}, whose validity requires𝒜0\\mathcal\{A\}\_\{0\}; the second element of𝒜0\\mathcal\{A\}\_\{0\}fails, so the guarantee is unsound\. If it is unset when it could have been set, the certificate reportsL^1\\hat\{L\}\_\{1\}, whose validity requires only𝒜1\\mathcal\{A\}\_\{1\}, which holds; the guarantee is sound\.
Hence selecting Tier 1 by default and Tier 0 only on explicit assertion places the unsound case behind a deliberate action rather than behind an omission\. The claim is exactly this asymmetry; by[Remark5\.1](https://arxiv.org/html/2608.12895#S5.Thmremark1)it is not a numerical ordering ofL^0\\hat\{L\}\_\{0\}andL^1\\hat\{L\}\_\{1\}\.
Finally, the failure in the first case is undetectable from the data: two runs producing identical pass matrices, one executed end\-to\-end and one assembled from per\-stage measurements, are indistinguishable to any function of the matrix, while only the first satisfies𝒜0\\mathcal\{A\}\_\{0\}\. No validation of the input can substitute for the default\. ∎
### A\.10Proof of[Theorem6\.1](https://arxiv.org/html/2608.12895#S6.Thmtheorem1)
###### Proof\.
*Soundness\.*Q⋆Q^\{\\star\}, viewed as a vectorx⋆∈ℝ2mx^\{\\star\}\\in\\mathbb\{R\}^\{2^\{m\}\}, is feasible for[Equation18](https://arxiv.org/html/2608.12895#S6.E18): it is non\-negative, sums to11, and satisfiesaS⊤x⋆=νSa\_\{S\}^\{\\top\}x^\{\\star\}=\\nu\_\{S\}forS∈𝒥S\\in\\mathcal\{J\}by hypothesis\. The minimum over a feasible set containingx⋆x^\{\\star\}is at most the objective atx⋆x^\{\\star\}, givingR¯≤Q⋆\(⋀ihi=1\)\\underline\{R\}\\leq Q^\{\\star\}\(\\bigwedge\_\{i\}h\_\{i\}=1\)\. The upper bound is symmetric\.
*Attainment\.*The feasible set𝒳=\{x≥0:𝟏⊤x=1,aS⊤x=νS\}\\mathcal\{X\}=\\\{x\\geq 0:\\mathbf\{1\}^\{\\top\}x=1,\\;a\_\{S\}^\{\\top\}x=\\nu\_\{S\}\\\}is the intersection of the probability simplex inℝ2m\\mathbb\{R\}^\{2^\{m\}\}with finitely many hyperplanes, hence closed and bounded, hence compact; and it is non\-empty becausex⋆∈𝒳x^\{\\star\}\\in\\mathcal\{X\}\. The objectivea\{1\.\.m\}⊤xa\_\{\\\{1\.\.m\\\}\}^\{\\top\}xis linear and therefore continuous, so by the extreme value theorem it attains its minimum and maximum on𝒳\\mathcal\{X\}at somex−,x\+∈𝒳x\_\{\-\},x\_\{\+\}\\in\\mathcal\{X\}\. Every element of𝒳\\mathcal\{X\}is a probability vector on\{0,1\}m\\\{0,1\\\}^\{m\}with the prescribed𝒥\\mathcal\{J\}\-moments, sox−x\_\{\-\}andx\+x\_\{\+\}are laws of the required kind, achievingR¯\\underline\{R\}andR¯\\overline\{R\}\.
*Consequence\.*The set of achievable values ofQ\(⋀ihi=1\)Q\(\\bigwedge\_\{i\}h\_\{i\}=1\)overQ∈ℳ\(ν\)Q\\in\\mathcal\{M\}\(\\nu\)is the image of the connected set𝒳\\mathcal\{X\}under a continuous map, hence an interval; combined with attainment of both endpoints it equals exactly\[R¯,R¯\]\[\\underline\{R\},\\overline\{R\}\]\. Any valid bound using only the moments in𝒥\\mathcal\{J\}must hold for everyQ∈ℳ\(ν\)Q\\in\\mathcal\{M\}\(\\nu\), in particular forx−x\_\{\-\}, so it cannot exceedR¯\\underline\{R\}\. ∎
### A\.11Proof of[Proposition6\.1](https://arxiv.org/html/2608.12895#S6.Thmproposition1)
###### Proof\.
Let𝒳\(𝒥\)\\mathcal\{X\}\(\\mathcal\{J\}\)and𝒳\(𝒥′\)\\mathcal\{X\}\(\\mathcal\{J\}^\{\\prime\}\)be the feasible sets for the two families, with moment values agreeing on𝒥\\mathcal\{J\}\. Everyx∈𝒳\(𝒥′\)x\\in\\mathcal\{X\}\(\\mathcal\{J\}^\{\\prime\}\)satisfies all constraints indexed by𝒥′⊇𝒥\\mathcal\{J\}^\{\\prime\}\\supseteq\\mathcal\{J\}, in particular those indexed by𝒥\\mathcal\{J\}, so𝒳\(𝒥′\)⊆𝒳\(𝒥\)\\mathcal\{X\}\(\\mathcal\{J\}^\{\\prime\}\)\\subseteq\\mathcal\{X\}\(\\mathcal\{J\}\)\. Minimising a fixed objective over a subset cannot yield a smaller value:R¯\(𝒥\)=min𝒳\(𝒥\)≤min𝒳\(𝒥′\)=R¯\(𝒥′\)\\underline\{R\}\(\\mathcal\{J\}\)=\\min\_\{\\mathcal\{X\}\(\\mathcal\{J\}\)\}\\leq\\min\_\{\\mathcal\{X\}\(\\mathcal\{J\}^\{\\prime\}\)\}=\\underline\{R\}\(\\mathcal\{J\}^\{\\prime\}\)\. The maximisation statement is symmetric\. ∎
### A\.12Proof of[Proposition6\.2](https://arxiv.org/html/2608.12895#S6.Thmproposition2)
###### Proof\.
*\(i\) Validity of the used\-set allocation*is[SectionA\.8](https://arxiv.org/html/2608.12895#A1.SS8)verbatim withJ=\|𝒥\|J=\|\\mathcal\{J\}\|\. For the failure of monotonicity, observe that the Clopper–Pearson interval for a fixed count widens as its tail level decreases\. Enlarging𝒥\\mathcal\{J\}to𝒥′\\mathcal\{J\}^\{\\prime\}decreases the per\-tail level fromηconf/\(2\|𝒥\|\)\\eta\_\{\\mathrm\{conf\}\}/\(2\|\\mathcal\{J\}\|\)toηconf/\(2\|𝒥′\|\)\\eta\_\{\\mathrm\{conf\}\}/\(2\|\\mathcal\{J\}^\{\\prime\}\|\), so every interval in the box strictly widens\. The feasible set for𝒥′\\mathcal\{J\}^\{\\prime\}therefore gains constraints \(from the new moments\) and loses them \(from the widened old ones\), and neither set need contain the other\. Hence the conclusion of[Proposition6\.1](https://arxiv.org/html/2608.12895#S6.Thmproposition1)does not transfer\.
*\(ii\) Validity of the pre\-allocated allocation\.*Fix𝒥max⊇𝒥\\mathcal\{J\}\_\{\\max\}\\supseteq\\mathcal\{J\}and give every constrained moment per\-tail levelηconf/\(2\|𝒥max\|\)\\eta\_\{\\mathrm\{conf\}\}/\(2\|\\mathcal\{J\}\_\{\\max\}\|\)\. The union bound over the2\|𝒥\|2\|\\mathcal\{J\}\|tails actually used gives miscoverage at most2\|𝒥\|⋅ηconf/\(2\|𝒥max\|\)=ηconf\|𝒥\|/\|𝒥max\|≤ηconf2\|\\mathcal\{J\}\|\\cdot\\eta\_\{\\mathrm\{conf\}\}/\(2\|\\mathcal\{J\}\_\{\\max\}\|\)=\\eta\_\{\\mathrm\{conf\}\}\|\\mathcal\{J\}\|/\|\\mathcal\{J\}\_\{\\max\}\|\\leq\\eta\_\{\\mathrm\{conf\}\}, so the certificate is valid at the stated level\.
*Monotonicity under \(ii\)\.*Each interval’s width now depends only on its own count and on\|𝒥max\|\|\\mathcal\{J\}\_\{\\max\}\|, not on\|𝒥\|\|\\mathcal\{J\}\|\. Hence for𝒥⊆𝒥′⊆𝒥max\\mathcal\{J\}\\subseteq\\mathcal\{J\}^\{\\prime\}\\subseteq\\mathcal\{J\}\_\{\\max\}the box constraints indexed by𝒥\\mathcal\{J\}are identical under both families, and𝒥′\\mathcal\{J\}^\{\\prime\}merely adds further constraints\. The argument of[SectionA\.11](https://arxiv.org/html/2608.12895#A1.SS11)applies unchanged, givingR¯\(𝒥\)≤R¯\(𝒥′\)\\underline\{R\}\(\\mathcal\{J\}\)\\leq\\underline\{R\}\(\\mathcal\{J\}^\{\\prime\}\)by construction\. ∎
### A\.13Proof of[Lemma7\.1](https://arxiv.org/html/2608.12895#S7.Thmlemma1)
###### Proof\.
LetℱR\\mathcal\{F\}\_\{R\}be theσ\\sigma\-algebra generated byy1,…,yRy\_\{1\},\\dots,y\_\{R\}\.
*Non\-negativity\.*Sinceyr∈\{0,1\}y\_\{r\}\\in\\\{0,1\\\}we haveyr−p0≥−p0y\_\{r\}\-p\_\{0\}\\geq\-p\_\{0\}, so1\+λr\(yr−p0\)≥1−λrp0\>01\+\\lambda\_\{r\}\(y\_\{r\}\-p\_\{0\}\)\\geq 1\-\\lambda\_\{r\}p\_\{0\}\>0becauseλr∈\[0,1/p0\)\\lambda\_\{r\}\\in\[0,1/p\_\{0\}\)\. A product of strictly positive factors is positive, soER\>0E\_\{R\}\>0for allRR\.
*Supermartingale property\.*λR\\lambda\_\{R\}is predictable, henceℱR−1\\mathcal\{F\}\_\{R\-1\}\-measurable, andER−1E\_\{R\-1\}isℱR−1\\mathcal\{F\}\_\{R\-1\}\-measurable\. Therefore
𝔼\[ER∣ℱR−1\]=ER−1\(1\+λR\(𝔼\[yR∣ℱR−1\]−p0\)\)\.\\mathbb\{E\}\[E\_\{R\}\\mid\\mathcal\{F\}\_\{R\-1\}\]=E\_\{R\-1\}\\Bigl\(1\+\\lambda\_\{R\}\\bigl\(\\mathbb\{E\}\[y\_\{R\}\\mid\\mathcal\{F\}\_\{R\-1\}\]\-p\_\{0\}\\bigr\)\\Bigr\)\.UnderH0H\_\{0\}we have𝔼\[yR∣ℱR−1\]≤p0\\mathbb\{E\}\[y\_\{R\}\\mid\\mathcal\{F\}\_\{R\-1\}\]\\leq p\_\{0\}by hypothesis: this is exactly the content of the sequential null[Equation19](https://arxiv.org/html/2608.12895#S7.E19), and it is the step at which a merely marginal null would fail, because a bound onℙ\(YG=1\)\\mathbb\{P\}\(Y\_\{G\}=1\)says nothing about the conditional mean given the past\. WithλR≥0\\lambda\_\{R\}\\geq 0andER−1\>0E\_\{R\-1\}\>0, the bracket is at most11and𝔼\[ER∣ℱR−1\]≤ER−1\\mathbb\{E\}\[E\_\{R\}\\mid\\mathcal\{F\}\_\{R\-1\}\]\\leq E\_\{R\-1\}\. Taking expectations and iterating fromE0=1E\_\{0\}=1gives𝔼\[ER\]≤1\\mathbb\{E\}\[E\_\{R\}\]\\leq 1for everyRR\. ∎
### A\.14Proof of[Theorem7\.1](https://arxiv.org/html/2608.12895#S7.Thmtheorem1)
###### Proof\.
Seta=1/αa=1/\\alphaand letT=inf\{R≥1:ER≥a\}T=\\inf\\\{R\\geq 1:E\_\{R\}\\geq a\\\}, withT=∞T=\\inftyif no crossing occurs\.TTis a stopping time with respect to\(ℱR\)\(\\mathcal\{F\}\_\{R\}\)because\{T≤R\}\\\{T\\leq R\\\}is determined byE1,…,ERE\_\{1\},\\dots,E\_\{R\}\.
Fixn∈ℕn\\in\\mathbb\{N\}\. The stopped process\(ET∧R\)R≤n\(E\_\{T\\wedge R\}\)\_\{R\\leq n\}is a non\-negative supermartingale by[Lemma7\.1](https://arxiv.org/html/2608.12895#S7.Thmlemma1)and the optional stopping theorem for bounded stopping times, so
𝔼\[ET∧n\]≤𝔼\[E0\]=1\.\\mathbb\{E\}\\bigl\[E\_\{T\\wedge n\}\\bigr\]\\leq\\mathbb\{E\}\[E\_\{0\}\]=1\.On the event\{T≤n\}\\\{T\\leq n\\\}we haveET∧n=ET≥aE\_\{T\\wedge n\}=E\_\{T\}\\geq aby definition ofTT\. SinceET∧n≥0E\_\{T\\wedge n\}\\geq 0everywhere,
1≥𝔼\[ET∧n\]≥𝔼\[ET∧n\[T≤n\]\]≥aℙ\(T≤n\),1\\geq\\mathbb\{E\}\\bigl\[E\_\{T\\wedge n\}\\bigr\]\\geq\\mathbb\{E\}\\bigl\[E\_\{T\\wedge n\}\\mathbf\{1\}\\\!\\left\[T\\leq n\\right\]\\bigr\]\\geq a\\,\\mathbb\{P\}\(T\\leq n\),soℙ\(T≤n\)≤1/a=α\\mathbb\{P\}\(T\\leq n\)\\leq 1/a=\\alphafor everynn\. The events\{T≤n\}\\\{T\\leq n\\\}increase to\{T<∞\}=\{supR≥1ER≥a\}\\\{T<\\infty\\\}=\\\{\\sup\_\{R\\geq 1\}E\_\{R\}\\geq a\\\}, so by continuity from below
ℙ\(supR≥1ER≥1/α\)=limn→∞ℙ\(T≤n\)≤α\.\\mathbb\{P\}\\Bigl\(\\sup\_\{R\\geq 1\}E\_\{R\}\\geq 1/\\alpha\\Bigr\)=\\lim\_\{n\\to\\infty\}\\mathbb\{P\}\(T\\leq n\)\\leq\\alpha\.Because the bound is on the supremum over the entire path, it holds simultaneously for all stopping rules: any rule that issues a certificate does so only on the event\{T<∞\}\\\{T<\\infty\\\}, whose probability is at mostα\\alphaunderH0H\_\{0\}, whatever the rule\. ∎
### A\.15Proof of[Proposition7\.1](https://arxiv.org/html/2608.12895#S7.Thmproposition1)
###### Proof\.
Withλ⋆=\(p1−p0\)/\(p0\(1−p0\)\)\\lambda^\{\\star\}=\(p\_\{1\}\-p\_\{0\}\)/\\bigl\(p\_\{0\}\(1\-p\_\{0\}\)\\bigr\), evaluate the betting factor at each outcome\.
Fory=1y=1:
1\+λ⋆\(1−p0\)=1\+\(p1−p0\)\(1−p0\)p0\(1−p0\)=1\+p1−p0p0=p1p0\.1\+\\lambda^\{\\star\}\(1\-p\_\{0\}\)=1\+\\frac\{\(p\_\{1\}\-p\_\{0\}\)\(1\-p\_\{0\}\)\}\{p\_\{0\}\(1\-p\_\{0\}\)\}=1\+\\frac\{p\_\{1\}\-p\_\{0\}\}\{p\_\{0\}\}=\\frac\{p\_\{1\}\}\{p\_\{0\}\}\.
Fory=0y=0:
1−λ⋆p0=1−\(p1−p0\)p0p0\(1−p0\)=1−p1−p01−p0=\(1−p0\)−\(p1−p0\)1−p0=1−p11−p0\.1\-\\lambda^\{\\star\}p\_\{0\}=1\-\\frac\{\(p\_\{1\}\-p\_\{0\}\)p\_\{0\}\}\{p\_\{0\}\(1\-p\_\{0\}\)\}=1\-\\frac\{p\_\{1\}\-p\_\{0\}\}\{1\-p\_\{0\}\}=\\frac\{\(1\-p\_\{0\}\)\-\(p\_\{1\}\-p\_\{0\}\)\}\{1\-p\_\{0\}\}=\\frac\{1\-p\_\{1\}\}\{1\-p\_\{0\}\}\.
Both agree with\(p1/p0\)y\(\(1−p1\)/\(1−p0\)\)1−y\(p\_\{1\}/p\_\{0\}\)^\{y\}\\bigl\(\(1\-p\_\{1\}\)/\(1\-p\_\{0\}\)\\bigr\)^\{1\-y\}, which is the likelihood ratio ofBern\(p1\)\\mathrm\{Bern\}\(p\_\{1\}\)toBern\(p0\)\\mathrm\{Bern\}\(p\_\{0\}\)atyy\. Substituting into[Equation20](https://arxiv.org/html/2608.12895#S7.E20)makesERE\_\{R\}the product of per\-observation likelihood ratios, i\.e\. the SPRT statistic\.
Admissibility:λ⋆≥0\\lambda^\{\\star\}\\geq 0sincep1\>p0p\_\{1\}\>p\_\{0\}, andλ⋆<1/p0\\lambda^\{\\star\}<1/p\_\{0\}because\(p1−p0\)/\(1−p0\)<1\(p\_\{1\}\-p\_\{0\}\)/\(1\-p\_\{0\}\)<1wheneverp1<1p\_\{1\}<1\. ∎
### A\.16Proof of[Proposition7\.2](https://arxiv.org/html/2608.12895#S7.Thmproposition2)
###### Proof\.
*Supermartingale\.*EachEλE^\{\\lambda\}is a non\-negative supermartingale by[Lemma7\.1](https://arxiv.org/html/2608.12895#S7.Thmlemma1)\. For non\-negative weights summing to one,
𝔼\[ERmix∣ℱR−1\]=∑λπ\(λ\)𝔼\[ERλ∣ℱR−1\]≤∑λπ\(λ\)ER−1λ=ER−1mix,\\mathbb\{E\}\\bigl\[E\_\{R\}^\{\\mathrm\{mix\}\}\\mid\\mathcal\{F\}\_\{R\-1\}\\bigr\]=\\sum\_\{\\lambda\}\\pi\(\\lambda\)\\,\\mathbb\{E\}\\bigl\[E\_\{R\}^\{\\lambda\}\\mid\\mathcal\{F\}\_\{R\-1\}\\bigr\]\\leq\\sum\_\{\\lambda\}\\pi\(\\lambda\)E\_\{R\-1\}^\{\\lambda\}=E\_\{R\-1\}^\{\\mathrm\{mix\}\},where exchanging expectation and the finite sum is immediate\. Non\-negativity is inherited, andE0mix=∑λπ\(λ\)=1E\_\{0\}^\{\\mathrm\{mix\}\}=\\sum\_\{\\lambda\}\\pi\(\\lambda\)=1\. Hence[Theorem7\.1](https://arxiv.org/html/2608.12895#S7.Thmtheorem1)applies verbatim\.
*Regret\.*All terms are non\-negative, so for any fixedλ\\lambda,ERmix≥π\(λ\)ERλE\_\{R\}^\{\\mathrm\{mix\}\}\\geq\\pi\(\\lambda\)E\_\{R\}^\{\\lambda\}\. Taking logarithms,logERmix≥logERλ−log\(1/π\(λ\)\)\\log E\_\{R\}^\{\\mathrm\{mix\}\}\\geq\\log E\_\{R\}^\{\\lambda\}\-\\log\(1/\\pi\(\\lambda\)\), and maximising the right\-hand side overλ∈Λ\\lambda\\in\\Lambdagives the claim\. The bound is pathwise: it holds for every realisation, not merely in expectation\. ∎
## Appendix BFormula Catalogue
The v1 framework was published with several formulas withheld under a patent claim\. That claim is withdrawn\. This appendix is the complete disclosure: every formula, its v1 equation number, where it is stated in this paper, and whether it is implemented in the released library\. Nothing in this table is redacted, and the “theory only” entries are marked as such rather than left ambiguous\.
Table 12:The F1–F12 catalogue\. “v1 ref” cites the framework paper\([Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\); “here” points into this paper; “code” names the module in the released library or records that the item is theoretical\.### B\.1Corrections to the v1 statements
Three discrepancies between the v1 paper and its supporting documents are worth recording, because a reader reconstructing the framework from either could adopt the wrong form\.
##### F2, condition \(ii\)\.
Supporting documents state the soft guarantee as the deterministic boundmaxt\|Csoft\(t\)−1\|≤δ\\max\_\{t\}\|C\_\{\\mathrm\{soft\}\}\(t\)\-1\|\\leq\\delta\. The v1 paper’s eq\. 4, restated as[Equation7](https://arxiv.org/html/2608.12895#S3.E7), is a*probabilistic*guarantee about recoverable compliance\. These are different conditions and the deterministic one is strictly stronger\.[Definition3\.6](https://arxiv.org/html/2608.12895#S3.Thmdefinition6)follows the paper\.
##### F5, the condition list\.
Supporting documents present “C1–C5” as a single list\. The v1 paper uses C1–C4 for the deterministic composition theorem \(Def\. 4\.7, Thm\. 4\.9\) and adds C5, conditional independence, only for the probabilistic theorem \(Thm\. 4\.11\)\. Collapsing them obscures which result needs which assumption — and C5 is the one this paper is about\.[Definitions3\.9](https://arxiv.org/html/2608.12895#S3.Thmdefinition9)and[3\.10](https://arxiv.org/html/2608.12895#S3.Thmdefinition10)keep them separate\.
##### F8, the fourth term\.
Supporting documents describeα4\\alpha\_\{4\}as weighting a “recovery success rate”\. The v1 paper’s Def\. 3\.20 defines the fourth term as the stress resilience indexS=𝔼\[C\(t\)∣stressed\]/𝔼\[C\(t\)∣baseline\]S=\\mathbb\{E\}\[C\(t\)\\mid\\text\{stressed\}\]/\\mathbb\{E\}\[C\(t\)\\mid\\text\{baseline\}\]\(eq\. 12\)\. Recovery success and stress resilience are distinct quantities\.[Definition3\.7](https://arxiv.org/html/2608.12895#S3.Thmdefinition7)follows the paper and notes the conflation\.
### B\.2Scope of the disclosure
F3 and F4 are stated and proved \([Definition3\.8](https://arxiv.org/html/2608.12895#S3.Thmdefinition8),[Theorem3\.1](https://arxiv.org/html/2608.12895#S3.Thmtheorem1),[SectionA\.1](https://arxiv.org/html/2608.12895#A1.SS1)\) but are not implemented and not measured in this paper: no experiment here estimatesα\\alpha,γ\\gamma, orσ\\sigma\. F5’s bound is implemented, but conditions C1–C5 are not machine\-verified by the library — a caller composing two contracts is responsible for discharging them, and[Section10](https://arxiv.org/html/2608.12895#S10)exists because C5 in particular frequently cannot be discharged\. We flag both rather than let the catalogue imply that a listed formula is an enforced one\.
## Appendix CReproduction and Artifact
### C\.1What is released
The library, the contracts, the mission generators, the deterministic scoring code, the preregistration, the simulation benchmarks, and the analysis scripts are released under AGPL\-3\.0\. Every number in[Section10](https://arxiv.org/html/2608.12895#S10)is*regenerated*by those scripts rather than transcribed by hand: each E1 and E5 statistic is printed output ofscripts/e1\_final\.py, and each E3, E4, and E6 table is printed output of the corresponding benchmark\.
The per\-mission logs are the one artifact not in the repository\. They carry the full model output for every mission in the corpus, and we release them on request rather than by default\. Everything needed to*regenerate*them is public — the generators, the frozen sampling configuration of[Section10\.1](https://arxiv.org/html/2608.12895#S10.SS1), the preregistration, and the scoring code — so the pipeline reproduces end to end without them\. What they save a replicator is the inference spend, not the method\.
### C\.2Reproducing each experiment
E1, E5 — dependence and cross\-backend replication\.scripts/e1\_final\.pyreads the per\-arm JSONL logs and emits every cell count, estimator, bootstrap interval, and arm contrast in[Tables2](https://arxiv.org/html/2608.12895#S10.T2),[3](https://arxiv.org/html/2608.12895#S10.T3),[4](https://arxiv.org/html/2608.12895#S10.T4)and[8](https://arxiv.org/html/2608.12895#S10.T8)\. Its defaults are the published settings — seed20260813,B=2000B=2000, failure defined as¬hard\_ok\\lnot\\,\\texttt\{hard\\\_ok\}, tables keyed oncomponent\_id— so a bare invocation reproduces the paper\. Cell counts and point estimates reproduce exactly; percentile bootstrap endpoints are resampling\-dependent and reproduce to within Monte Carlo error at thisBB, which does not move any reported verdict\.
E2 — floor lifting\.benchmarks/reconstructs them=4m\{=\}4pass matrix and callsmoment\_cp\_box\_floorandmoment\_lp\_all\_success\_boundsat both moment families and both Bonferroni allocations\.
E3 — coverage collapse\.[Table6](https://arxiv.org/html/2608.12895#S10.T6)is produced by exactly two invocations ofbenchmarks/coverage\_collapse\_sim\.py:\-\-witness adversarialand\-\-witness gaussian\. The script’s defaults are the published settings \(200200replications,nboot=500n\_\{\\mathrm\{boot\}\}=500, seed20260812\), so a bare run reproduces the table; it prints its parameters on every run, so a reduced\-fidelity result cannot be mistaken for the published one\. The Gaussian moments, the LP minimiser andΔ\\Deltaare computed by deterministic quadrature and linear programming inside the same script, so every constant in[Example4\.2](https://arxiv.org/html/2608.12895#S4.Thmexample2)and[Section10\.4](https://arxiv.org/html/2608.12895#S10.SS4)has a single source\.
E4 — anytime validity\.benchmarks/eprocess\_type1\_sim\.py, defaulting to the paper settings \(p0=0\.8p\_\{0\}=0\.8,α=0\.05\\alpha=0\.05, 8000 streams of 500 missions, seed 42\) and sweeping the admissible range ofλ\\lambda\.
E6 — ablation\.benchmarks/deff\_ablation\.pymeasures lag\-kkautocorrelation and design effects, then recomputes the certified floor at a conceded design effect\.
### C\.3Two reproduction hazards
Two ways of computing the wrong number from the logs produce plausible output rather than an error\. Both are recorded here\.
##### Components are identified bycomponent\_id, notrole\.
Every model\-backed node carriesrole = "worker"\. Keying a co\-failure table onrolesilently compares an agent*to itself*and returns a degenerate table — Jaccard exactly1\.01\.0with both off\-diagonal cells zero, in every arm\. The identifiers arenode\_a/node\_bforseries2,worker\_0\.\.2plusaggregatorforquorum2of3, andbranch\_a/branch\_bplusmergeforparallel2\. Deterministic nodes carryscored = false\.
##### An exception is not a data point\.
An earlier version of the E3 script read a field name that did not exist on the result object and caught the resultingAttributeErrorin a broadexcept Exception, scoring it as a coverage miss\. It reported “coverage0\.000\.00” for runs in which the estimator never executed\. The lesson generalises beyond this script: in simulation code, catch only the exception that models a real degenerate case, and let everything else fail loudly\.
##### The failure field ishard\_ok\.
Component records have bothhard\_okandsoft\_ok\. All dependence estimates in this paper are on the hard verdict\. Reading a differently\-named field yields all\-zero tables rather than an error\.
### C\.4Preregistration
PREREGISTRATION\.mdis committed in the repository and git\-timestamped before any confirmatory outcome was generated\. It fixes the hypotheses, the five conditions and their sample sizes, the frozen sampling parameters, the primary estimator and bootstrap, the stopping rule, and the falsification criteria\. Two registered deviations are disclosed in this paper: the breadth arms under\-ran their registerednn\([Section10\.6](https://arxiv.org/html/2608.12895#S10.SS6)\), and H2 is reported on a marginal\-free statistic chosen after observing the reversal \([Section11\.3](https://arxiv.org/html/2608.12895#S11.SS3)\)\. The registration was not posted to an external registry before submission, so its timestamp rests on the repository history rather than on a third party; we regard that as weaker than an external registration and state it plainly\.
### C\.5Scale
The full campaign is 30,820 scored missions across 12 arms, recorded in 13 execution logs\. Each arm resumes from its own log, so an interrupted run never repeats completed missions\.
## Availability and Licensing
TheAgentAssertlibrary, contracts, mission generators, scoring code, preregistration, simulation benchmarks, and analysis scripts are released under AGPL\-3\.0 at[https://github\.com/qualixar/agentassert\-abc](https://github.com/qualixar/agentassert-abc)\. The per\-mission logs are available on request; see[AppendixC](https://arxiv.org/html/2608.12895#A3)\. The v1 framework paper is[Bhardwaj 2026](https://arxiv.org/html/2608.12895#bib.bib7)\.
## Author Contributions
V\.P\.B\. designed the study, wrote the preregistration, implemented the library and the experimental harness, ran the campaign, performed the analysis, and wrote the paper\. G\.S\. and A\.P\.B\. contributed to contract design, reviewed the experimental protocol, and reviewed the manuscript\. All authors approved the final version\.
## Use of AI Assistance
An AI assistant was used for writing and for reviewing the manuscript\. All experiments, the code repository, the analysis, the testing, and the review were carried out by the authors, who take full responsibility for the contents\.
## Conflicts of Interest
AgentAssertis developed by Qualixar, with which V\.P\.B\. is affiliated\. The evaluation measures the library’s own certificates, so the authors are not disinterested parties\. Three mitigations are in place and we state them so a reader can weigh them: the confirmatory hypotheses and analysis were preregistered before outcomes existed; all scoring is by deterministic gold code rather than by author judgement or an LLM judge; and the analysis scripts are released, so that every reported number is the printed output of code a reader can inspect rather than a figure transcribed by an author\. The patent claim referenced in the v1 paper has been withdrawn, and no patent application covering the methods in this paper is pending\.
## References
- Bai et al\. \(2022\)Yuntao Bai, Saurav Kadavath, Sandipan Kundu, et al\.Constitutional AI: Harmlessness from AI feedback\.*arXiv preprint arXiv:2212\.08073*, 2022\.
- Barlow and Proschan \(1975\)Richard E\. Barlow and Frank Proschan\.*Statistical Theory of Reliability and Life Testing*\.Holt, Rinehart and Winston, 1975\.
- Barnett et al\. \(2004\)Mike Barnett, K\. Rustan M\. Leino, and Wolfram Schulte\.The Spec\# programming system: An overview\.In*CASSIS*, pages 49–69, 2004\.
- Bengio et al\. \(2026\)Yoshua Bengio et al\.International AI safety report 2026\.*arXiv preprint arXiv:2602\.21012*, 2026\.
- Berdoz et al\. \(2026\)Frédéric Berdoz, Leonardo Rugli, and Roger Wattenhofer\.Can AI agents agree?*arXiv preprint arXiv:2603\.01213*, 2026\.
- Bertsimas and Popescu \(2005\)Dimitris Bertsimas and Ioana Popescu\.Optimal inequalities in probability theory: A convex optimization approach\.*SIAM Journal on Optimization*, 15\(3\):780–804, 2005\.
- Bhardwaj \(2026\)Varun Pratap Bhardwaj\.Agent behavioral contracts: Formal specification and runtime enforcement for reliable autonomous AI agents\.*arXiv preprint arXiv:2602\.22302*, 2026\.
- Bonferroni \(1936\)Carlo E\. Bonferroni\.Teoria statistica delle classi e calcolo delle probabilità\.*Pubblicazioni del R\. Istituto Superiore di Scienze Economiche e Commerciali di Firenze*, 8:3–62, 1936\.
- Boole \(1854\)George Boole\.*An Investigation of the Laws of Thought*\.Walton and Maberly, 1854\.
- Chase \(2022\)Harrison Chase\.LangChain\.[https://github\.com/langchain\-ai/langchain](https://github.com/langchain-ai/langchain), 2022\.
- Clarke et al\. \(1999\)Edmund M\. Clarke, Orna Grumberg, and Doron A\. Peled\.*Model Checking*\.MIT Press, 1999\.
- Clopper and Pearson \(1934\)C\. J\. Clopper and E\. S\. Pearson\.The use of confidence or fiducial limits illustrated in the case of the binomial\.*Biometrika*, 26\(4\):404–413, 1934\.
- Cousot and Cousot \(1977\)Patrick Cousot and Radhia Cousot\.Abstract interpretation: A unified lattice model for static analysis of programs by construction or approximation of fixpoints\.In*POPL*, pages 238–252, 1977\.
- Du et al\. \(2024\)Yilun Du, Shuang Li, Antonio Torralba, Joshua B\. Tenenbaum, and Igor Mordatch\.Improving factuality and reasoning in language models through multiagent debate\.In*ICML*, 2024\.
- Efron \(1979\)Bradley Efron\.Bootstrap methods: Another look at the jackknife\.*The Annals of Statistics*, 7\(1\):1–26, 1979\.
- Endres and Schindelin \(2003\)Dominik M\. Endres and Johannes E\. Schindelin\.A new metric for probability distributions\.*IEEE Transactions on Information Theory*, 49\(7\):1858–1860, 2003\.
- Ernst et al\. \(2007\)Michael D\. Ernst, Jeff H\. Perkins, Philip J\. Guo, Stephen McCamant, Carlos Pacheco, Matthew S\. Tschantz, and Chen Xiao\.The Daikon system for dynamic detection of likely invariants\.*Science of Computer Programming*, 69\(1–3\):35–45, 2007\.
- Fréchet \(1951\)Maurice Fréchet\.Sur les tableaux de corrélation dont les marges sont données\.*Annales de l’Université de Lyon, Section A*, 14:53–77, 1951\.
- Gelman and Loken \(2014\)Andrew Gelman and Eric Loken\.The statistical crisis in science\.*American Scientist*, 102\(6\):460–465, 2014\.
- Grünwald et al\. \(2024\)Peter Grünwald, Rianne de Heide, and Wouter Koolen\.Safe testing\.*Journal of the Royal Statistical Society B*, 86\(5\):1091–1128, 2024\.
- Hailperin \(1965\)Theodore Hailperin\.Best possible inequalities for the probability of a logical function of events\.*The American Mathematical Monthly*, 72\(4\):343–359, 1965\.
- Hammond et al\. \(2025\)Lewis Hammond et al\.Multi\-agent risks from advanced AI\.*arXiv preprint arXiv:2502\.14143*, 2025\.
- Hoare \(1969\)C\. A\. R\. Hoare\.An axiomatic basis for computer programming\.*Communications of the ACM*, 12\(10\):576–580, 1969\.
- Hoeffding \(1940\)Wassily Hoeffding\.Maßstabinvariante Korrelationstheorie\.*Schriften des Mathematischen Instituts und des Instituts für Angewandte Mathematik der Universität Berlin*, 5:179–233, 1940\.
- Hoeffding \(1963\)Wassily Hoeffding\.Probability inequalities for sums of bounded random variables\.*Journal of the American Statistical Association*, 58\(301\):13–30, 1963\.
- Hong et al\. \(2024\)Sirui Hong, Mingchen Zhuge, Jonathan Chen, et al\.MetaGPT: Meta programming for a multi\-agent collaborative framework\.In*ICLR*, 2024\.
- Howard et al\. \(2021\)Steven R\. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon\.Time\-uniform, nonparametric, nonasymptotic confidence sequences\.*The Annals of Statistics*, 49\(2\):1055–1080, 2021\.
- Huang et al\. \(2026\)Jiatan Huang, Mingchen Li, Ziming Li, Sunjae Kwon, Hong Yu, and Chuxu Zhang\.Counterfactual graph for multi\-agent LLM calibration\.*arXiv preprint arXiv:2605\.30653*, 2026\.
- Jaccard \(1912\)Paul Jaccard\.The distribution of the flora in the alpine zone\.*New Phytologist*, 11\(2\):37–50, 1912\.
- Jimenez et al\. \(2024\)Carlos E\. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan\.SWE\-bench: Can language models resolve real\-world GitHub issues?In*ICLR*, 2024\.
- Kendall \(1938\)Maurice G\. Kendall\.A new measure of rank correlation\.*Biometrika*, 30\(1–2\):81–93, 1938\.
- Kish \(1965\)Leslie Kish\.*Survey Sampling*\.John Wiley and Sons, New York, 1965\.
- Lamport \(2002\)Leslie Lamport\.*Specifying Systems: The TLA\+ Language and Tools for Hardware and Software Engineers*\.Addison\-Wesley, 2002\.
- Leino \(2010\)K\. Rustan M\. Leino\.Dafny: An automatic program verifier for functional correctness\.In*LPAR*, pages 348–370, 2010\.
- Liang et al\. \(2024\)Tian Liang, Zhiwei He, Wenxiang Jiao, et al\.Encouraging divergent thinking in large language models through multi\-agent debate\.In*EMNLP*, 2024\.
- Lin \(1991\)Jianhua Lin\.Divergence measures based on the Shannon entropy\.*IEEE Transactions on Information Theory*, 37\(1\):145–151, 1991\.
- Liu et al\. \(2024\)Xiao Liu, Hao Yu, Hanchen Zhang, et al\.AgentBench: Evaluating LLMs as agents\.In*ICLR*, 2024\.
- McDonnell et al\. \(2026\)Shay Seiya McDonnell, Avantika Singh, Quoc\-Viet Pham, Vratislav Havlik, and Gregory M\. P\. O’Hare\.Harnessing disagreement: Detecting correlated agreement blindness in multi\-agent triage\.*arXiv preprint arXiv:2607\.19899*, 2026\.Accepted, PAAMS 2026\.
- Meyer \(1992\)Bertrand Meyer\.Applying “design by contract”\.*Computer*, 25\(10\):40–51, 1992\.
- Nelsen \(2006\)Roger B\. Nelsen\.*An Introduction to Copulas*\.Springer, 2nd edition, 2006\.
- Nosek et al\. \(2018\)Brian A\. Nosek, Charles R\. Ebersole, Alexander C\. DeHaven, and David T\. Mellor\.The preregistration revolution\.*Proceedings of the National Academy of Sciences*, 115\(11\):2600–2606, 2018\.
- Ouyang et al\. \(2022\)Long Ouyang, Jeffrey Wu, Xu Jiang, et al\.Training language models to follow instructions with human feedback\.In*NeurIPS*, 2022\.
- Park et al\. \(2023\)Joon Sung Park, Joseph C\. O’Brien, Carrie J\. Cai, Meredith Ringel Morris, Percy Liang, and Michael S\. Bernstein\.Generative agents: Interactive simulacra of human behavior\.In*UIST*, 2023\.
- Qian et al\. \(2024\)Chen Qian, Wei Liu, Hongzhang Liu, et al\.ChatDev: Communicative agents for software development\.In*ACL*, 2024\.
- Qiao et al\. \(2026\)Hezhe Qiao, Hanghang Tong, Ee\-Peng Lim, Bing Liu, and Guansong Pang\.VerifyMAS: Hypothesis verification for failure attribution in LLM multi\-agent systems\.*arXiv preprint arXiv:2605\.17467*, 2026\.
- Qin et al\. \(2026\)Xue Qin, Simin Luan, John See, Zeyd Boukhers, Cong Yang, and Zhijun Li\.Governed capability evolution: Lifecycle\-time compatibility checking and rollback for AI\-component\-based systems\.*arXiv preprint arXiv:2604\.08059*, 2026\.
- Rafi et al\. \(2026\)Md Nakhla Rafi, Md Ahasanuzzaman, Dong Jae Kim, Zhijie Wang, and Tse\-Hsun Chen\.FALAT: Tracing failures in LLM agent trajectories via dependency\-guided search\.*arXiv preprint arXiv:2606\.00765*, 2026\.
- Ramdas et al\. \(2023\)Aaditya Ramdas, Peter Grünwald, Vladimir Vovk, and Glenn Shafer\.Game\-theoretic statistics and safe anytime\-valid inference\.*Statistical Science*, 38\(4\):576–601, 2023\.
- Rebedea et al\. \(2023\)Traian Rebedea, Razvan Dinu, Makesh Sreedhar, Christopher Parisien, and Jonathan Cohen\.NeMo guardrails: A toolkit for controllable and safe LLM applications with programmable rails\.In*EMNLP System Demonstrations*, 2023\.
- Robbins \(1970\)Herbert Robbins\.Statistical methods related to the law of the iterated logarithm\.*The Annals of Mathematical Statistics*, 41\(5\):1397–1409, 1970\.
- Shafer \(2021\)Glenn Shafer\.Testing by betting: A strategy for statistical and scientific communication\.*Journal of the Royal Statistical Society A*, 184\(2\):407–431, 2021\.
- Shafer and Vovk \(2019\)Glenn Shafer and Vladimir Vovk\.*Game\-Theoretic Foundations for Probability and Finance*\.Wiley, 2019\.
- Shinn et al\. \(2023\)Noah Shinn, Federico Cassano, Edward Berman, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\.Reflexion: Language agents with verbal reinforcement learning\.In*NeurIPS*, 2023\.
- Simmons et al\. \(2011\)Joseph P\. Simmons, Leif D\. Nelson, and Uri Simonsohn\.False\-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant\.*Psychological Science*, 22\(11\):1359–1366, 2011\.
- Sklar \(1959\)Abe Sklar\.Fonctions de répartition ànndimensions et leurs marges\.*Publications de l’Institut de Statistique de l’Université de Paris*, 8:229–231, 1959\.
- Slepian \(1962\)David Slepian\.The one\-sided barrier problem for Gaussian noise\.*Bell System Technical Journal*, 41\(2\):463–501, 1962\.
- Ville \(1939\)Jean Ville\.*Étude critique de la notion de collectif*\.Gauthier\-Villars, Paris, 1939\.
- Wald \(1945\)Abraham Wald\.Sequential tests of statistical hypotheses\.*The Annals of Mathematical Statistics*, 16\(2\):117–186, 1945\.
- Wang et al\. \(2023\)Xuezhi Wang, Jason Wei, Dale Schuurmans, et al\.Self\-consistency improves chain of thought reasoning in language models\.In*ICLR*, 2023\.
- Waudby\-Smith and Ramdas \(2024\)Ian Waudby\-Smith and Aaditya Ramdas\.Estimating means of bounded random variables by betting\.*Journal of the Royal Statistical Society B*, 86\(1\):1–27, 2024\.
- Wu et al\. \(2024\)Qingyun Wu, Gagan Bansal, Jieyu Zhang, et al\.AutoGen: Enabling next\-gen LLM applications via multi\-agent conversation\.In*COLM*, 2024\.
- Yao et al\. \(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao\.ReAct: Synergizing reasoning and acting in language models\.In*ICLR*, 2023\.
- Yule \(1900\)G\. Udny Yule\.On the association of attributes in statistics: With illustrations from the material of the childhood society\.*Philosophical Transactions of the Royal Society A*, 194:257–319, 1900\.
- Zheng et al\. \(2025\)Lifan Zheng, Jiawei Chen, Qinghong Yin, Jingyuan Zhang, Xinyi Zeng, and Yu Tian\.Rethinking the reliability of multi\-agent system: A perspective from Byzantine fault tolerance\.*arXiv preprint arXiv:2511\.10400*, 2025\.
- Zhou et al\. \(2024\)Shuyan Zhou, Frank F\. Xu, Hao Zhu, et al\.WebArena: A realistic web environment for building autonomous agents\.In*ICLR*, 2024\.Similar Articles
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
This paper introduces AgentCollabBench, a diagnostic benchmark for multi-agent systems that evaluates behavioral risks like instruction decay and context leakage across four major LLMs. It argues that communication topology is a critical factor in multi-agent reliability, often overshadowing raw model capability.
Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems
This paper proposes a behavioral measure of trust between AI agents based on costly verification in a cooperative survival game, studying trust formation, breakage, and recovery across six frontier model snapshots. It finds that models differ in trust calibration and that persistent over-verification is associated with indecision rather than safety.
On the Reliability of Computer Use Agents
A preprint analyzing why computer-use agents succeed once but fail on repeated executions, attributing unreliability to execution stochasticity, task ambiguity, and behavioral variability, and advocating repeated evaluation and stable strategies.
Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy
This paper reports a failure study of a production agentic software-delivery platform and introduces reliability primitives for non-idempotent agent delegation, focusing on identity adequacy and evidence adequacy to improve system robustness.
Contract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning
A method for contract-based compositional shielding that ensures global safety in multi-agent reinforcement learning without centralized runtime control, using local LTL obligations and a multi-armed bandit to optimize team reward.