Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning

arXiv cs.AI 论文

摘要

This paper proves that local verification in agentic AI is structurally incomplete, using cohomology theory to show that harmonic evidence conflicts (non-transportability) are undetectable by any local checks. It proposes Ksetra, a method that gates abstention on harmonic energy, and provides a statistical test for global claims.

arXiv:2608.11252v1 Announce Type: new Abstract: Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local verification: checking at each step that the entity is representable in the chosen tool, that parameters are compatible, and that outputs cohere with the plan. We prove this class of safeguard is structurally incomplete. Modelling a covering of context space by its nerve and evidence by a real-valued 1-cochain, an agent chaining evidence performs path integration: its conclusion is path-independent if and only if the cochain is exact, and disagreement between valid reasoning paths is exactly the holonomy of a first Cech cohomology class. Hodge decomposition partitions evidence conflict into a gradient part (calibration), a curl part (local inconsistency, visible at triple overlaps) and a harmonic part. Our central result is that no family of simplex-supported consistency checks can distinguish omega from omega+h for harmonic h, which nonetheless generates non-zero disagreement between valid paths; detection requires a statistic on a cycle basis. The resulting procedure, Ksetra, estimates by coboundary projection and gates abstention on the harmonic component, which we give a mechanism: it arises from effect modification combined with overlap-specific population composition, and vanishes to machine precision when effect modification is absent. The degrees of freedom of an evidence network partition into calibration, coherence and transport, yielding an exact F-test for the existence of a global claim; we quantify its distortion under unequal precision and supply the precision-whitened form that restores exactness. Foreign exchange, where the arbitrage-free null makes the cochain exactly a coboundary, serves as a calibration bench: the test is correctly sized, fires on loop arbitrage, and ignores triangular arbitrage.
查看原文
查看缓存全文

缓存时间: 2026/08/13 15:24

# Local verification cannot detect non-transportability Cohomological limits of context preservation in agentic reasoning
Source: [https://arxiv.org/html/2608.11252](https://arxiv.org/html/2608.11252)
###### Abstract

Agentic AI systems increasingly transport conclusions across biological, clinical and commercial contexts\. The current state of the art addresses this by*local verification*: checking, at each step, that the entity under discussion is representable in the selected tool, that parameters are compatible, and that outputs cohere with the plan\. We prove that this class of safeguard is structurally incomplete\. Modelling a covering of context space by its nerve and evidence by a real\-valued11\-cochain, we show that an agent chaining evidence performs path integration, that its conclusion is path\-independent if and only if the evidence cochain is exact, and that the ambiguity between reasoning paths is exactly the holonomy of the first Čech cohomology class\[ω\]∈H1\[\\omega\]\\in H^\{1\}\. Hodge decomposition then partitions evidence conflict into three orthogonal parts with distinct operational meanings: a gradient part \(per\-context calibration, removable\), a curl part \(local inconsistency, detectable at triple overlaps\), and a harmonic part \(structural non\-transportability\)\. Our central negative result is that*every verification scheme whose checks are supported on simplices of the nerve is blind to the harmonic part*, which nonetheless generates non\-zero disagreement between valid reasoning paths\. We proposeKṣetra\(ASCII: Ksetra\), which estimates by coboundary projection and gates abstention on harmonic energy alone, and we show that a non\-zero class is not merely a warning but a data\-collection instruction: refine the covering\. We further derive an exactFF\-test for the existence of a global claim, resting on the observation that the degrees of freedom of an evidence network partition into calibration, local coherence and transport\. We give the harmonic component a mechanism: it is generated by effect modification combined with overlap\-specific population composition, and vanishes to machine precision \(1\.5×10−151\.5\\times 10^\{\-15\}\) when effect modification is absent—Simpson’s paradox on a network\. In simulations where holonomy is emergent rather than injected, harmonic energy predicts the irreducible error of the optimal estimator \(ρ=0\.37\\rho=0\.37,p<10−48p<10^\{\-48\}\), retaining independent value after conditioning on the other two components \(partialρ=0\.19\\rho=0\.19,p<10−13p<10^\{\-13\}\); curl energy is comparably predictive \(ρ=0\.33\\rho=0\.33\) because the same latent cause generates both, so the components are distinguished by their*remedy*, not by their predictive value\. Gating on harmonic energy rather than total conflict lowers area under the risk–coverage curve by0\.0290\.029and0\.0390\.039\(paired bootstrap,p<0\.001p<0\.001\)—a modest gain we report as corroboration, not as the contribution\. Exactness of theFF\-test requires isotropic precision; we quantify the distortion when it fails \(size0\.0710\.071at nominal0\.050\.05under a fourfold spread of edge precisions\) and supply the precision\-whitened generalisation that restores it\. We conclude with the real\-data validation required to move these results off simulation\.

## 1Introduction

A therapeutic hypothesis, a pricing rule and a credit cut\-off share a structural property: each is a claim indexed by a context, and each is routinely applied outside the context in which it was established\. Recent work on agentic reasoning in biology has made this the organising problem\.Medea\[[1](https://arxiv.org/html/2608.11252#bib.bib1)\]identifies three failure modes of tool\-using agents—loss of biological context over long horizons, absent verification of intermediate steps, and unreconciled conflict across evidence sources—and addresses them with context verification against a tool space, pre\- and post\-execution checks, and a multi\-round consensus panel that may abstain\. The reported gains are substantial, and the validation against a previously unpublished genome\-wide yeast screen is the strongest available evidence that such gains reflect biology rather than benchmark leakage\.

We take that architecture as given and ask a different question:*what class of failure can such verification detect, and what class can it not?*

Our answer is negative and, we believe, sharp\. All of the verification in\[[1](https://arxiv.org/html/2608.11252#bib.bib1)\]—plan\-time context checks, run\-time provenance auditing, literature relevance screening, panel deliberation—is*local*in a precise sense: each check is supported on a bounded piece of context space \(one context, one pair of contexts being bridged, one triple being cross\-checked\)\. We show that evidence conflict decomposes orthogonally into three parts, that local checks see exactly two of them, and that the third—the harmonic part, the first cohomology of the nerve of the context covering—is invisible to every local check while being precisely the part that makes an agent’s conclusion depend on which reasoning path it happened to take\. An agent can pass every check at every step and still return an answer that is an artefact of its route through the evidence\.

This is not an exotic possibility\. It is the formal content of a phenomenon clinical evidence synthesis has recognised for two decades under the name*loop inconsistency*in network meta\-analysis\[[6](https://arxiv.org/html/2608.11252#bib.bib6),[8](https://arxiv.org/html/2608.11252#bib.bib8),[7](https://arxiv.org/html/2608.11252#bib.bib7)\]: direct and indirect evidence around a closed loop of comparisons can each be individually sound and jointly incoherent\. Our contribution is to identify loop inconsistency asH1H^\{1\}of a context site, to extend it from treatment networks to arbitrary context coverings, to prove that it bounds what local verification can achieve, and to convert it into an abstention criterion and a data\-collection instruction for AI agents\.

#### Contributions\.

1. 1\.Theory\.We model context space as a site, evidence as a11\-cochain on the nerve of a covering, and agentic transport as path integration\. Theorem[5](https://arxiv.org/html/2608.11252#Thmtheorem5)establishes path\-invariance⇔\\iffexactness and identifies inter\-path disagreement with holonomy\. Theorem[7](https://arxiv.org/html/2608.11252#Thmtheorem7)gives the operational meaning of the Hodge components\. Corollary[9](https://arxiv.org/html/2608.11252#Thmtheorem9)is the impossibility result for local verification\. Proposition[12](https://arxiv.org/html/2608.11252#Thmtheorem12)shows refinement is the unique remedy\.
2. 2\.Method\.Kṣetra: coboundary\-projection estimation, harmonic\-gated abstention, and a three\-way triage of evidence conflict into*recalibrate*,*re\-measure*, and*refine the covering*\.
3. 3\.Cross\-industry evidence\.A controlled study over2,4002\{,\}400context\-transport problems in two structurally different domains, one with a contractible attribute space and one whose attribute space is a circle—the business cycle—and therefore carries obstruction by construction\.
4. 4\.Decision economics\.A transparent mapping from selective\-risk geometry to portfolio outcomes, with the threshold harm:benefit ratio above which cohomological gating earns its cost\.

## 2Related work

Agentic reasoning with context preservation\.Medea\[[1](https://arxiv.org/html/2608.11252#bib.bib1)\]is the immediate antecedent\. ItsResearchPlanningmodule verifies that a requested entity is representable in the selected tool; itsAnalysismodule performs pre\-run compatibility and post\-run provenance checks; itsMultiRoundDiscussionmodule adapts ReConcile\-style panel deliberation\[[2](https://arxiv.org/html/2608.11252#bib.bib2)\]over a ReAct control loop\[[3](https://arxiv.org/html/2608.11252#bib.bib3)\]\. Our results apply to this whole family: we characterise what such architectures can and cannot see\.

Transportability\.Pearl and Bareinboim give necessary and sufficient conditions for transporting causal effects across populations using selection diagrams\[[4](https://arxiv.org/html/2608.11252#bib.bib4),[5](https://arxiv.org/html/2608.11252#bib.bib5)\]\. That theory answers “may I transport between*these two*populations?” Our contribution is orthogonal and global: given a network of contexts with pairwise bridging evidence, it asks whether the pairwise answers can be glued into a coherent global assignment at all\.

Evidence inconsistency\.Loop inconsistency in network meta\-analysis, node splitting and design\-by\-treatment interaction models\[[6](https://arxiv.org/html/2608.11252#bib.bib6),[7](https://arxiv.org/html/2608.11252#bib.bib7),[8](https://arxiv.org/html/2608.11252#bib.bib8)\]are, in our language, tests for a non\-zero11\-cocycle on a treatment network\. We generalise the object and reinterpret the test\.

Combinatorial Hodge theory\.The decomposition we use is that of Jiang, Lim, Yao and Ye for statistical ranking\[[9](https://arxiv.org/html/2608.11252#bib.bib9)\], with antecedents in applied topology\[[10](https://arxiv.org/html/2608.11252#bib.bib10),[11](https://arxiv.org/html/2608.11252#bib.bib11)\]and sheaf theory\[[13](https://arxiv.org/html/2608.11252#bib.bib13),[12](https://arxiv.org/html/2608.11252#bib.bib12)\]\. To our knowledge the harmonic component has not previously been proposed as an abstention criterion for reasoning systems\.

Selective prediction\.Risk–coverage analysis follows El\-Yaniv and Wiener\[[14](https://arxiv.org/html/2608.11252#bib.bib14)\]and Geifman and El\-Yaniv\[[15](https://arxiv.org/html/2608.11252#bib.bib15)\]\. Our departure is that the gating statistic is derived from the topology of the evidence graph rather than from model confidence\.

## 3Theory

### 3\.1The context site and the evidence cochain

###### Definition 1\(Context site\)\.

LetXXbe a space of decision\-relevant situations and𝒰=\{Ui\}i∈I\\mathcal\{U\}=\\\{U\_\{i\}\\\}\_\{i\\in I\}a finite covering by*contexts*: subsets on which a claim is asserted\. LetN​\(𝒰\)N\(\\mathcal\{U\}\)be the nerve: verticesII, an edgei​jijwheneverUi∩Uj≠∅U\_\{i\}\\cap U\_\{j\}\\neq\\emptyset\(a population, cohort or stratum on which the two contexts may be bridged\), a trianglei​j​kijkwheneverUi∩Uj∩Uk≠∅U\_\{i\}\\cap U\_\{j\}\\cap U\_\{k\}\\neq\\emptyset\.

Writeδ0:C0→C1\\delta^\{0\}:C^\{0\}\\to C^\{1\},\(δ0​x\)i​j=xj−xi\(\\delta^\{0\}x\)\_\{ij\}=x\_\{j\}\-x\_\{i\}andδ1:C1→C2\\delta^\{1\}:C^\{1\}\\to C^\{2\},\(δ1​y\)i​j​k=yj​k−yi​k\+yi​j\(\\delta^\{1\}y\)\_\{ijk\}=y\_\{jk\}\-y\_\{ik\}\+y\_\{ij\}, soδ1​δ0=0\\delta^\{1\}\\delta^\{0\}=0\.

###### Definition 2\(Evidence cochain\)\.

An*evidence cochain*ω∈C1​\(N​\(𝒰\);ℝ\)\\omega\\in C^\{1\}\(N\(\\mathcal\{U\}\);\\mathbb\{R\}\)assigns to each overlapi​jijthe measured contrast in the quantity of interest between contextsiiandjj, estimated onUi∩UjU\_\{i\}\\cap U\_\{j\}\. A*claim assignment*isθ∈C0\\theta\\in C^\{0\}\. Evidence is*coherent*withθ\\thetaifω=δ0​θ\\omega=\\delta^\{0\}\\theta\.

### 3\.2Context drift is a non\-functorial move

###### Proposition 4\(Silent coarsening\)\.

Letc⊑c′c\\sqsubseteq c^\{\\prime\}be a refinement of contexts, with restrictionres:F​\(c′\)→F​\(c\)\\mathrm\{res\}:F\(c^\{\\prime\}\)\\to F\(c\)on the evidence presheafFF\. Answering a query posed atccwith a section defined atc′c^\{\\prime\}is not the application of any morphism ofFF; the induced error equals the within\-c′c^\{\\prime\}variation of the section,θc−𝔼c′′⊑c′​\[θc′′\]\\theta\_\{c\}\-\\mathbb\{E\}\_\{c^\{\\prime\\prime\}\\sqsubseteq c^\{\\prime\}\}\[\\theta\_\{c^\{\\prime\\prime\}\}\], and is unbounded by any quantity computed atc′c^\{\\prime\}\.

This is the formal content of an agent collapsing naïve CD4\+α​β\\alpha\\betaT cells into “CD4\+T cells”\[[1](https://arxiv.org/html/2608.11252#bib.bib1)\]: the answer is not wrong atc′c^\{\\prime\}, it is*unanswered*atcc, and no diagnostic available atc′c^\{\\prime\}reveals the gap\. Section[5\.5](https://arxiv.org/html/2608.11252#S5.SS5)quantifies the cost in isolation\.

### 3\.3Agentic transport is path integration

An agent asked for the claim at targettt, anchored on established knowledge ataa, proceeds by chaining: it recalls the value ataa, bridges to a neighbouring context, bridges again, and arrives attt\. Formally, for a pathγ\\gammafromaatottinN​\(𝒰\)N\(\\mathcal\{U\}\)with signed edges, the transported estimate is

θ^tγ=θa\+∑e∈γ±ωe=θa\+⟨ω,γ⟩\.\\hat\{\\theta\}^\{\\gamma\}\_\{t\}\\;=\\;\\theta\_\{a\}\\;\+\\;\\sum\_\{e\\in\\gamma\}\\pm\\,\\omega\_\{e\}\\;=\\;\\theta\_\{a\}\+\\langle\\omega,\\gamma\\rangle\.
###### Theorem 5\(Path invariance and holonomy\)\.

\(i\)θ^tγ\\hat\{\\theta\}^\{\\gamma\}\_\{t\}is independent ofγ\\gammafor every pair\(a,t\)\(a,t\)if and only ifω∈im⁡δ0\\omega\\in\\operatorname\{im\}\\delta^\{0\}\. \(ii\) For two pathsγ,γ′\\gamma,\\gamma^\{\\prime\}with common endpoints,θ^tγ−θ^tγ′=⟨ω,γ−γ′⟩\\hat\{\\theta\}^\{\\gamma\}\_\{t\}\-\\hat\{\\theta\}^\{\\gamma^\{\\prime\}\}\_\{t\}=\\langle\\omega,\\gamma\-\\gamma^\{\\prime\}\\rangle, which depends only on the class\[ω\]∈H1​\(N​\(𝒰\);ℝ\)\[\\omega\]\\in H^\{1\}\(N\(\\mathcal\{U\}\);\\mathbb\{R\}\)wheneverδ1​ω=0\\delta^\{1\}\\omega=0\.

###### Proof\.

z=γ−γ′z=\\gamma\-\\gamma^\{\\prime\}is a11\-cycle\. For anyx∈C0x\\in C^\{0\},⟨δ0​x,z⟩=⟨x,∂z⟩=0\\langle\\delta^\{0\}x,z\\rangle=\\langle x,\\partial z\\rangle=0, so exact cochains pair trivially with cycles; hence the difference depends onω\\omegaonly moduloim⁡δ0\\operatorname\{im\}\\delta^\{0\}, giving \(ii\)\. For \(i\), path\-independence for all pairs is equivalent to⟨ω,z⟩=0\\langle\\omega,z\\rangle=0for every cyclezz, i\.e\.ω⟂Z1\\omega\\perp Z\_\{1\}\. Overℝ\\mathbb\{R\},C1=im⁡δ0⊕\(im⁡δ0\)⟂C^\{1\}=\\operatorname\{im\}\\delta^\{0\}\\oplus\(\\operatorname\{im\}\\delta^\{0\}\)^\{\\perp\}and\(im⁡δ0\)⟂=Z1\(\\operatorname\{im\}\\delta^\{0\}\)^\{\\perp\}=Z\_\{1\}under the standard inner product, soω⟂Z1⇔ω∈im⁡δ0\\omega\\perp Z\_\{1\}\\iff\\omega\\in\\operatorname\{im\}\\delta^\{0\}\. ∎

### 3\.4Hodge triage and the blindness of local verification

###### Theorem 7\(Triage\)\.

C1=imδ0⊕ℋ1⊕im\(δ1\)⊤C^\{1\}=\\operatorname\{im\}\\delta^\{0\}\\ \\oplus\\ \\mathcal\{H\}^\{1\}\\ \\oplus\\ \\operatorname\{im\}\(\\delta^\{1\}\)^\{\\\!\\top\}orthogonally, whereℋ1=kerδ1∩ker\(δ0\)⊤≅H1\(N\(𝒰\);ℝ\)\\mathcal\{H\}^\{1\}=\\ker\\delta^\{1\}\\cap\\ker\(\\delta^\{0\}\)^\{\\\!\\top\}\\cong H^\{1\}\(N\(\\mathcal\{U\}\);\\mathbb\{R\}\)\. Writingω=g\+h\+c\\omega=g\+h\+c:

1. \(a\)g=δ0​βg=\\delta^\{0\}\\betais exactly the contamination produced by per\-context instrument offsetsβ\\betaand is removed by recalibration against any single anchored context\.
2. \(b\)c∈im\(δ1\)⊤c\\in\\operatorname\{im\}\(\\delta^\{1\}\)^\{\\\!\\top\}is the unique component withδ1​c≠0\\delta^\{1\}c\\neq 0; it is precisely what a triple\-overlap coherence check detects\.
3. \(c\)hhsatisfiesδ1​h=0\\delta^\{1\}h=0, so it passes every triple\-overlap check, yeth∉im⁡δ0h\\notin\\operatorname\{im\}\\delta^\{0\}, so by Theorem[5](https://arxiv.org/html/2608.11252#Thmtheorem5)it produces non\-zero disagreement between some pair of reasoning paths\.

###### Proof\.

Orthogonality ofim⁡δ0\\operatorname\{im\}\\delta^\{0\}andim\(δ1\)⊤\\operatorname\{im\}\(\\delta^\{1\}\)^\{\\top\}follows fromδ1​δ0=0\\delta^\{1\}\\delta^\{0\}=0:⟨δ0​x,\(δ1\)⊤​z⟩=⟨δ1​δ0​x,z⟩=0\\langle\\delta^\{0\}x,\(\\delta^\{1\}\)^\{\\top\}z\\rangle=\\langle\\delta^\{1\}\\delta^\{0\}x,z\\rangle=0\. The remaining summand isker\(δ0\)⊤∩kerδ1\\ker\(\\delta^\{0\}\)^\{\\top\}\\cap\\ker\\delta^\{1\}, which is the space of harmonic11\-cochains, isomorphic toH1H^\{1\}by the discrete Hodge theorem\. \(a\) is immediate\. \(b\) holds becauseδ1\\delta^\{1\}annihilates the other two summands\. \(c\) combinesh∈ker⁡δ1h\\in\\ker\\delta^\{1\}withh⟂im⁡δ0h\\perp\\operatorname\{im\}\\delta^\{0\}and Theorem[5](https://arxiv.org/html/2608.11252#Thmtheorem5)\(i\)\. ∎

###### Definition 8\(Simplex\-supported consistency check\)\.

A statisticTTis a*consistency check on the simplexσ\\sigma*if \(i\)TTis a function ofω\\omegaonly through its restriction to the star ofσ\\sigma, and \(ii\)T​\(ω\)=0T\(\\omega\)=0whenever that restriction is exact\. Context representability, parameter compatibility, pairwise bridging plausibility, triple\-overlap coherence and single\-execution provenance audits are all of this form\.

###### Corollary 9\(Local verification is incomplete\)\.

No family of simplex\-supported consistency checks can distinguishω\\omegafromω\+h\\omega\+hforh∈ℋ1h\\in\\mathcal\{H\}^\{1\}\. DetectingH1H^\{1\}obstruction requires a statistic supported on a cycle basis ofN​\(𝒰\)N\(\\mathcal\{U\}\), i\.e\. a genuinely global computation\.

Corollary[9](https://arxiv.org/html/2608.11252#Thmtheorem9)is the paper’s central claim\. It says that the verification strategy embodied in current context\-preserving agents, however carefully executed, has a blind spot that is not a matter of implementation quality\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F1_hodge.png)Figure 1:Hodge decomposition of one evidence 1\-cochain\. The three components are orthogonal and carry distinct operational meanings; only the third is invisible to every local verification check\.
### 3\.5Abstention as a cohomological quantity

Non\-zero\[ω\]\[\\omega\]has a substantive interpretation: the covering is*too coarse*\. If a coherent global claim existed at this granularity, evidence would be exact up to noise\. Persistent holonomy certifies that the nominal contexts still contain latent strata whose claims differ\. We therefore adopt:

###### Proposition 11\(Error floor\)\.

Suppose \(i\) the estimator is the coboundary projectionθ^=arg⁡minθ⁡‖δ0​θ−ω‖2\\hat\{\\theta\}=\\arg\\min\_\{\\theta\}\\\|\\delta^\{0\}\\theta\-\\omega\\\|\_\{2\}, and \(ii\) the residual within\-context heterogeneity unlocked by a coarse covering has scale increasing in‖h‖\\\|h\\\|\. Then the risk of any decision taken at the nominal granularity is bounded below by a monotone function of‖h‖\\\|h\\\|, and is asymptotically insensitive to‖g‖\\\|g\\\|and‖c‖\\\|c\\\|, which the estimator projects out\.

Assumption \(ii\) is substantive and is the object of our empirical test in §[5\.2](https://arxiv.org/html/2608.11252#S5.SS2): it predicts that harmonic energy, and*only*harmonic energy, should correlate with the irreducible error of the optimal estimator\.

###### Proposition 12\(Refinement is the remedy\)\.

Let𝒱\\mathcal\{V\}refine𝒰\\mathcal\{U\}\. Classes inH1​\(N​\(𝒰\)\)H^\{1\}\(N\(\\mathcal\{U\}\)\)that arise from unresolved within\-context heterogeneity are filled inN​\(𝒱\)N\(\\mathcal\{V\}\); in the colimit over refinements the Čech obstruction to gluing vanishes for a sheaf\. Operationally: a non\-zero class is an instruction to*stratify further and collect there*, not an instruction to gather more evidence at the same granularity\.

###### Proposition 13\(Bridge to transportability, informal\)\.

IfN​\(𝒰\)N\(\\mathcal\{U\}\)is generated by a selection diagram and each edgei​jijisSS\-admissible in the sense of\[[4](https://arxiv.org/html/2608.11252#bib.bib4)\], then eachωi​j\\omega\_\{ij\}identifiesθj−θi\\theta\_\{j\}\-\\theta\_\{i\}andω\\omegais exact up to sampling error\. Conversely a non\-zero class certifies thatSS\-admissibility fails somewhere on the corresponding loop, without identifying where\. Cohomology thus provides a*global detector*for a condition the do\-calculus checks*locally*\.

## 4TheKṣetraprocedure

1. 1\.Site construction\.Elicit the attribute factorisation of context space; buildN​\(𝒰\)N\(\\mathcal\{U\}\)from the overlaps for which bridging evidence exists\. ReportdimH1\\dim H^\{1\}\.
2. 2\.Evidence assembly\.Populateω\\omegafrom tool outputs, bridging studies and literature contrasts, one entry per usable overlap\.
3. 3\.Decomposition\.Computeω=g\+h\+c\\omega=g\+h\+cby least squares againstδ0\\delta^\{0\}and\(δ1\)⊤\(\\delta^\{1\}\)^\{\\top\}\.
4. 4\.Triage\.Dominantgg⇒\\Rightarrowrecalibrate instruments and proceed\. Dominantcc⇒\\Rightarrowlocal inconsistency; re\-measure or re\-run the deliberation\. Dominanthh⇒\\Rightarrowstructural obstruction\.
5. 5\.Estimate and gate\.Reportθ^\\hat\{\\theta\}by coboundary projection; abstain iff‖h‖/\|E\|\\\|h\\\|/\\sqrt\{\|E\|\}exceeds a threshold calibrated to the decision tolerance\. On abstention, emit the cycle carrying the largest holonomy as the recommended stratification target\.

## 5Simulation study

### 5\.1Design

We generate context sites as products of two ordinal attributes on a4×44\\times 4grid, with adjacent contexts sharing population \(edges\) and a minority admitting genuine three\-way overlap \(triangles\), then delete a random fraction of overlaps to represent evidence gaps\. Evidence is generated asω=δ0​\(θ\+β\)\+h\+c\+ε\\omega=\\delta^\{0\}\(\\theta\+\\beta\)\+h\+c\+\\varepsilonwithβ\\betaper\-context instrument offsets,hhdrawn in the harmonic subspace with scaleη\\eta,ccdrawn inim\(δ1\)⊤\\operatorname\{im\}\(\\delta^\{1\}\)^\{\\top\}with per\-replication magnitude, andε\\varepsilonsampling noise\. Consistent with Proposition[11](https://arxiv.org/html/2608.11252#Thmtheorem11), the realised outcome at the target carries additional dispersion proportional toη\\eta, representing the latent stratum the coarse covering failed to separate\.

#### Two instantiations\.

*Pharma RWE*— contexts are \(CKD stage×\\timescare setting\); the attribute space is contractible, andH1H^\{1\}arises only from unfilled squares and evidence gaps\.*Consumer credit*— contexts are \(region×\\timesmacroeconomic regime\) with the regime axis*cyclic*: expansion→\\tolate cycle→\\tocontraction→\\torecovery→\\toexpansion\. Because the business cycle is a circle, this site carriesdimH1≥1\\dim H^\{1\}\\geq 1by construction, independent of any data problem\. Realised means:dimH1=5\.84\\dim H^\{1\}=5\.84\(pharma\) and9\.759\.75\(credit\) over1616contexts\.

#### Agents\.

Pooledignores context\.Single\-chainfollows one plausible reasoning path and always commits\.Panel debatesamples three paths and abstains on spread—our model of aMedea\-style consensus module\.Residual\-gateduses the coboundary\-projection estimator and abstains on*total*residual conflict: the strongest baseline a careful engineer would build\.Kṣetrauses the same estimator and abstains on harmonic energy alone\.

1,2001\{,\}200replications per domain; thresholds swept to trace full risk–coverage curves so that agents are compared at matched coverage rather than at hand\-picked operating points\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F3_obstruction_error.png)Figure 2:Harmonic energy governs both the spread across valid reasoning paths \(orange\) and the error floor of the optimal path\-invariant estimator \(purple\)\.

### 5\.2Do the theoretical claims hold?

Table 1:Spearman correlations,n=1,200n=1\{,\}200per domain\.p∗⁣∗∗<10−4\{\}^\{\*\*\*\}p<10^\{\-4\}; all unmarked entriesp\>0\.16p\>0\.16\.

Claim A confirms Theorem[5](https://arxiv.org/html/2608.11252#Thmtheorem5): both non\-exact components generate path\-dependence, gradient does not\. Claim B is the discriminating result\. Only harmonic energy predicts the error that survives optimal estimation; curl and gradient are statistically indistinguishable from zero\. Claim C follows: an agent that gates on total conflict is using a signal diluted by two components that carry no information about irreducible risk\.

### 5\.3Selective performance

Table 2:Main results\. AURC is area under the risk–coverage curve \(lower is better\); harmful==error beyond the decision tolerance\. Agents without a conflict signal cannot abstain and are shown at full coverage\.Paired bootstrap over2,0002\{,\}000resamples,Kṣetraminus the strongest baseline:Δ​AURC=−0\.029\\Delta\\mathrm\{AURC\}=\-0\.029\[−0\.045,−0\.014\]\[\-0\.045,\-0\.014\],p<0\.001p<0\.001\(pharma\) and−0\.039\-0\.039\[−0\.052,−0\.027\]\[\-0\.052,\-0\.027\],p<0\.001p<0\.001\(credit\); harmful rate at75%75\\%coverageΔ=−0\.032\\Delta=\-0\.032\[−0\.052,−0\.012\]\[\-0\.052,\-0\.012\],p=0\.003p=0\.003and−0\.043\-0\.043\[−0\.060,−0\.026\]\[\-0\.060,\-0\.026\],p<0\.001p<0\.001\. The two agents share an estimator and differ only in the gating statistic, so the entire gap is attributable to Theorem[7](https://arxiv.org/html/2608.11252#Thmtheorem7)\.

#### Two further observations\.

First, the single\-chain agent is*worse*than the context\-blind pooled agent \(MAE1\.821\.82vs0\.830\.83\)\. Chaining evidence compounds variance along the path; a longer, more sophisticated reasoning trajectory over a noisy evidence graph is actively harmful unless the trajectory is aggregated correctly\. This is a mechanism for the observed brittleness of long\-horizon agents that does not require appeal to hallucination\. Second, panel debate improves on single\-chain but remains far behind projection\-based estimation, consistent with the remark following Theorem[5](https://arxiv.org/html/2608.11252#Thmtheorem5): sampling paths estimates the holonomy but does not remove it\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F2_risk_coverage.png)Figure 3:Risk–coverage\. The two right\-hand curves share an estimator and differ only in the abstention statistic; the gap is the empirical content of Theorem[7](https://arxiv.org/html/2608.11252#Thmtheorem7)\.![Refer to caption](https://arxiv.org/html/2608.11252v1/F4_eta_sweep.png)Figure 4:Degradation as structural non\-transportabilityη\\etaincreases, at full coverage\. Projection\-based estimation degrades gracefully; path\-chaining does not\.

### 5\.4Conflict triage

Kṣetraclassifies each query\. Pharma:41\.0%41\.0\\%calibration offset \(reconcile and proceed\),34\.4%34\.4\\%local inconsistency \(re\-measure\),24\.6%24\.6\\%H1H^\{1\}obstruction \(refine the covering\)\. Credit:39\.4%39\.4\\%/31\.6%31\.6\\%/29\.0%29\.0\\%\. Roughly one query in four is structurally unanswerable at the granularity posed—and in the cyclic credit site the fraction is higher, as the topology predicts\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F5_harm_triage.png)Figure 5:Left: harmful commitments at matched 75% coverage\. Right:Kṣetraconflict triage—recalibrate, re\-measure, or refine the covering\.![Refer to caption](https://arxiv.org/html/2608.11252v1/F6_drift.png)Figure 6:Cost of silent coarsening in isolation, with every other error source removed\.
### 5\.5The cost of drift alone

Isolating Proposition[4](https://arxiv.org/html/2608.11252#Thmtheorem4)with all other error sources removed: silent coarsening to the parent context on30%30\\%of queries yields MAE0\.1340\.134and a9\.0%9\.0\\%harmful rate; at50%50\\%,0\.2060\.206and16\.0%16\.0\\%\. Drift is expensive even when every other component of the system is perfect\.

## 6Simulated real\-world impact

We translate selective\-risk geometry into portfolio outcomes\.*All unit costs are declared parameters, not findings*; the scientific content is the mapping and its sensitivity\. A deploying organisation must substitute audited figures\.

#### Scenario A: target nomination\.

500500context\-specific nominations screened per year; a false commitment consumes one preclinical validation campaign \(cost11\); a correct nomination returns1\.61\.6; an abstention costs0\.150\.15in expert routing\. At75%75\\%coverage:134\.2134\.2harmful commitments under residual gating versus122\.1122\.1underKṣetra—1212campaigns per year redirected, a net gain of31\.431\.4units\.

#### Scenario B: credit transport\.

100,000100\{,\}000scored applications per quarter; a mis\-transported cut\-off costs11, a correctly priced account returns0\.550\.55, a manual referral costs0\.080\.08\. At75%75\\%coverage:23,83323\{,\}833versus20,66720\{,\}667harmful decisions—3,1673\{,\}167per quarter—for a net gain of4,9084\{,\}908units\. Optimising coverage rather than fixing it at75%75\\%raisesKṣetrato9,0549\{,\}054at65%65\\%coverage against3,4833\{,\}483for the baseline, a factor of2\.62\.6\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F7_impact.png)Figure 7:Decision economics\. Left: value–coverage frontier\. Centre and right: the outcome mix at 75% coverage, harmful commitments to the right of the axis and abstentions to the left\.
#### When does this pay?

Sweeping the harm:benefit ratio from0\.250\.25to1616, the advantage of cohomological gating grows monotonically \(24→32924\\to 329units in pharma;2,177→29,6082\{,\}177\\to 29\{,\}608in credit\)\. Below a ratio of roughly11, the pharma optimum sits near full coverage and the abstention machinery earns little\. This is a genuine boundary condition on the method and we state it plainly:*cohomological abstention is worth its complexity only where committing wrongly costs materially more than committing rightly gains*—which is the regime of regulated therapeutic and credit decisions, and is not the regime of exploratory analysis\.

## 7A finance test case: a domain where the null is exact

Biology cannot falsify this theory cleanly, because the null—“a coherent global claim exists”—is never known independently of the data\. Finance can\. We therefore use it as a metrology bench\.

### 7\.1An exact test for the existence of a global claim

![Refer to caption](https://arxiv.org/html/2608.11252v1/F8_finance_test.png)Figure 8:Left: theFF\-test is exactly calibrated under the null over4,0004\{,\}000replications\. Centre: power\. Right: harmonic monitoring detects loop arbitrage at11bp that total\-residual monitoring misses until44bp\.The degrees of freedom of an evidence network partition exactly:

\|E\|⏟comparisons=\(\|V\|−k\)⏟calibration\+rank⁡δ1⏟local coherence\+β1⏟transport,k=\#​components\.\\underbrace\{\|E\|\}\_\{\\text\{comparisons\}\}=\\underbrace\{\(\|V\|\-k\)\}\_\{\\text\{calibration\}\}\+\\underbrace\{\\operatorname\{rank\}\\delta^\{1\}\}\_\{\\text\{local coherence\}\}\+\\underbrace\{\\beta\_\{1\}\}\_\{\\textsc\{transport\}\},\\qquad k=\\\#\\text\{components\}\.UnderH0H\_\{0\}\(evidence coherent, homoscedastic noise\) the harmonic and curl energies are independentχ2\\chi^\{2\}variates on their respective dimensions, so

F=‖h‖2/β1‖c‖2/rank⁡δ1∼F​\(β1,rank⁡δ1\)\\boxed\{\\;F=\\frac\{\\\|h\\\|^\{2\}/\\beta\_\{1\}\}\{\\\|c\\\|^\{2\}/\\operatorname\{rank\}\\delta^\{1\}\}\\;\\sim\\;F\(\\beta\_\{1\},\\operatorname\{rank\}\\delta^\{1\}\)\\;\}is an*exact*test that a global section exists\. It is an analysis of variance in which the residual is partitioned by topology rather than by design factors\.

Over4,0004\{,\}000null replications on random nerves the test is correctly sized:0\.10050\.1005,0\.04700\.0470and0\.00770\.0077at nominal0\.100\.10,0\.050\.05and0\.010\.01, with Kolmogorov–SmirnovD=0\.018D=0\.018,p=0\.135p=0\.135against uniformity\. Power atα=0\.05\\alpha=0\.05with noiseσ=0\.20\\sigma=0\.20reaches0\.620\.62atη=0\.30\\eta=0\.30and0\.880\.88atη=0\.60\\eta=0\.60, crossing80%80\\%nearη≈0\.47\\eta\\approx 0\.47\.

### 7\.2Foreign exchange: the null is not estimated, it is known

In an arbitrage\-free market the log\-quote cochain is*exactly*a coboundary, sincelog⁡Si​j=vj−vi\\log S\_\{ij\}=v\_\{j\}\-v\_\{i\}for a numeraire valuevv\. Triangular arbitrage is the curl component\. Loop arbitrage on cycles that no*simultaneously executable*triangle fills is the harmonic component\. This identification is the practical content of Theorem[7](https://arxiv.org/html/2608.11252#Thmtheorem7)for a trading desk:*arbitrage is holonomy*\.

Two claims must be kept apart here\. That an arbitrage\-free log\-quote cochain is exactly a coboundary is a fact about the market\. The topology below is a*constructed illustration*: we build a1515\-currency complex with a G10 tier \(USD hub plus liquid crosses\) and a deliverable\-restricted tier of five Asian currencies whose regional crosses are stipulated to print in a fixing window disjoint from the New York window of their USD legs, so the USD triangulation of a regional cross is not simultaneously executable\. The fragmentation mechanism is plausible but stylised; a desk should rebuild the nerve from its own executability constraints before drawing conclusions\.

Table 3:Topology of the quote network\.3030quoted pairs,1919closed triangles, of which1414are simultaneously executable\.Two consequences\. First, the G10 complex is cohomologically trivial: every cycle is filled by an executable triangle, which is a*proof*that triangular\-arbitrage monitoring is complete there—standard practice is correct, and now for a stated reason\. Second, the naive audit of the full complex also reportsdimH1=0\\dim H^\{1\}=0and concludes there is nothing to see; only the executability\-aware nerve reveals five independent loops carrying P&L that no triangle check can detect\. This is Corollary[9](https://arxiv.org/html/2608.11252#Thmtheorem9)in a domain with money attached\.

Table 4:400400simulated books per scenario, quoting noise0\.350\.35bp\.The test fires on loop arbitrage and correctly ignores triangular arbitrage, which lives in the curl space\. The fourth row is a genuine weakness and we report it plainly: when curl energy is large it inflates the denominator and*masks*the harmonic signal, so theFFratio loses power\. The remedy is to use harmonic energy directly as a detection statistic with a resampled null\. Doing so, and asking the detector to find loop arbitrage hidden behind44bp of triangular arbitrage, the harmonic detector reaches AUROC0\.8350\.835at0\.50\.5bp and0\.9970\.997at11bp, whereas total\-residual monitoring is still at0\.5150\.515at11bp and does not reach0\.950\.95until44bp\.*Harmonic monitoring sees at one basis point what conventional residual monitoring cannot see until four\.*

### 7\.3IFRS 9 and SR 11\-7: transport across a cyclic regime space

A PD model inventory over five retail segments and four macroeconomic regimes, where the regime axis is a circle: expansion→\\tolate cycle→\\tocontraction→\\torecovery→\\toexpansion\. Joint\-validation samples spanning three cells exist only in the stable regimes\. The resulting site has2020contexts,4242pairwise validations, six joint\-validation triples, and

42=19⏟calibration\+6⏟coherence\+𝟏𝟕⏟transport\.42\\;=\\;\\underbrace\{19\}\_\{\\text\{calibration\}\}\\;\+\\;\\underbrace\{6\}\_\{\\text\{coherence\}\}\\;\+\\;\\underbrace\{\\mathbf\{17\}\}\_\{\\textsc\{transport\}\}\.Seventeen independent loops in the inventory can carry an incoherent PD story that no amount of pairwise or triple validation will detect\. This is a statement about the*governance architecture*, available before any data is examined\.

Expressing the consequence as expected\-credit\-loss misstatement in basis points of exposure \(EAD $100m per cell, LGD45%45\\%, base PD3\.0%3\.0\\%\):

Table 5:ECL misstatement, bp of exposure\. Abstention is gated on significance*and*materiality \(\>10\>10bp\), because significance alone is not a control\.Chaining a single validated route—the natural thing for both a human validator and an agent to do—degrades roughly linearly in the incoherence, reaching7979bp\. Projection is flat at55bp throughout: it is immune by construction, which is Theorem[5](https://arxiv.org/html/2608.11252#Thmtheorem5)\. The abstention rate is the diagnostic that tells the institution which is which\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F9_finance_impact.png)Figure 9:Left: ECL misstatement under each transport policy\. Right: the model\-inventory cycle audit—loops where the PD story does not close\.#### The deliverable\.

Kṣetraemits an artefact a validation function can act on: loops ranked by holonomy, expressed as the unclosed PD gap\. In the audited inventory the worst loops carry0\.680\.68log\-odds of holonomy, a290290bp gap in PD that does not close when the story is walked around the cycle\. Each such loop names the segments and regimes whose joint stratification would fill it\. Under SR 11\-7\[[16](https://arxiv.org/html/2608.11252#bib.bib16)\]this is a materially different artefact from a confidence score: it is a finding, with a location and a remediation\.

## 8Where the obstruction comes from

An earlier version of this work injected the harmonic component directly and then reported that it predicted error\. That is circular, and the objection is fatal, so we remove the injection entirely and let holonomy arise from a mechanism\.

Let each coarse contextccbe a union of fine stratass, with true effect

θc,s=μc\+λs\+m​γc,s,\\theta\_\{c,s\}=\\mu\_\{c\}\+\\lambda\_\{s\}\+m\\,\\gamma\_\{c,s\},whereγ\\gammais a context\-by\-stratum interaction \(effect modification\) andmmits strength\. An overlapi​jijcan only estimate the contrast on the case mixwi​jw\_\{ij\}it actually contains, so the bridging evidence isωi​j=∑swi​j,s​\(θj,s−θi,s\)\\omega\_\{ij\}=\\sum\_\{s\}w\_\{ij,s\}\\,\(\\theta\_\{j,s\}\-\\theta\_\{i,s\}\)\.

###### Proposition 14\(Mechanism\)\.

Ifm=0m=0thenωi​j=μj−μi\\omega\_\{ij\}=\\mu\_\{j\}\-\\mu\_\{i\}for every overlap regardless ofwi​jw\_\{ij\}, soω\\omegais exact and the obstruction vanishes\. Ifm≠0m\\neq 0the contrast depends on which strata the overlap happened to sample, and generically\[ω\]≠0\[\\omega\]\\neq 0\.

Non\-transportability is therefore*effect modification combined with overlap\-specific population composition*: Simpson’s paradox on a network rather than on a single table\. Numerically, atm=0m=0the harmonic energy is1\.52×10−151\.52\\times 10^\{\-15\}—machine zero, not merely small—and it grows monotonically inmmthereafter \(Figure[10](https://arxiv.org/html/2608.11252#S8.F10), left\)\.

### 8\.1A correction to a claim we previously made

Under this mechanistic model, one of our earlier findings does not survive, and we report the correction rather than the original\.

Table 6:Spearman correlations with error,n=1,500n=1\{,\}500, holonomy emergent\.We previously reported that curl energy carries no information about irreducible error\.*That was an artefact of drawing the curl and harmonic components independently\.*In a mechanistic model a single latent cause—effect modification—generates both, so they co\-vary, and curl is nearly as predictive as harmonic\. What survives is narrower and, we think, more defensible: harmonic energy retains independent predictive value after conditioning on the other two components, and the operational triage of §[3\.4](https://arxiv.org/html/2608.11252#S3.SS4)is untouched because it concerns*remedy*rather than prediction\. Curl is repairable by re\-measuring the triple; harmonic is not repairable at that granularity by any amount of data\. Two components may both predict error while demanding different responses, and it is the response that the decomposition is for\.

### 8\.2Exactness of theFF\-test under unequal precision

Bridging estimates differ in precision\. TheFF\-test of §[7](https://arxiv.org/html/2608.11252#S7)is exact only under isotropic noise, and our initial expectation—that ignoring precision would be conservative—was wrong\.

Table 7:Empirical size at nominal0\.050\.05;2,5002\{,\}500null replications per cell\.The unweighted test is mildly anti\-conservative—a42%42\\%inflation of the false\-positive rate at fourfold spread\. The remedy is a change of metric\. WithD=diag⁡\(1/σe\)D=\\operatorname\{diag\}\(1/\\sigma\_\{e\}\), sendω↦D​ω\\omega\\mapsto D\\omega,δ0↦D​δ0\\delta^\{0\}\\mapsto D\\delta^\{0\},δ1↦δ1​D−1\\delta^\{1\}\\mapsto\\delta^\{1\}D^\{\-1\}\. Thenδ1​D−1​D​δ0=δ1​δ0=0\\delta^\{1\}D^\{\-1\}D\\delta^\{0\}=\\delta^\{1\}\\delta^\{0\}=0, so the whitened complex is a genuine cochain complex and every result above applies verbatim\. Whitening improves size and power together \(0\.6210\.621against0\.5830\.583atη=0\.30\\eta=0\.30\)\. Residual distortion persists at eightfold spread \(KS​p=0\.010\\mathrm\{KS\}\\ p=0\.010\); for heterogeneity beyond that we recommend a permutation null rather than theFFreference\.

### 8\.3Evidence gaps hide obstruction rather than manufacturing it

We had listed the confounding of gaps with obstruction as a limitation\. The sign is the opposite of what we assumed\. Deleting an overlap destroys cycles faster than it destroys the triangles that fill them, so on a2525\-context site with trueβ1=21\\beta\_\{1\}=21, observedβ1\\beta\_\{1\}falls to18\.318\.3,12\.012\.0and8\.98\.9at10%10\\%,30%30\\%and40%40\\%missing overlaps\.

###### Corollary 15\.

dimH1\\dim H^\{1\}computed on an incomplete nerve is a*lower bound*on the obstruction\. A sparse evidence base looks more coherent than it is\.

This is the dangerous direction, and it means the prevalence figures reported in the inconsistency literature—already acknowledged there as underpowered—are understatements for a second, structural reason\.

![Refer to caption](https://arxiv.org/html/2608.11252v1/F15_council_fixes.png)Figure 10:Left: holonomy emerges from effect modification, vanishing to machine precision when it is absent\. Centre: emergent obstruction predicts irreducible error for both estimators\. Right: ignoring unequal precision costs power as well as validity\.

## 9Discussion

#### What this changes\.

Current context\-preserving agents treat abstention as a confidence heuristic emitted by a deliberation panel\. We give it a definition: abstain when the evidence cochain carries a non\-zero cohomology class relative to the decision tolerance\. This is computable, auditable, reproducible, and—unlike a panel verdict—accompanied by an instruction \(which stratification to refine\)\. For regulated deployment, under GxP or SR 11\-7\[[16](https://arxiv.org/html/2608.11252#bib.bib16)\], an abstention that comes with a cycle witness is a materially different artefact from an abstention that comes with a confidence score\.

#### Why cross\-industry\.

The credit instantiation is not decoration\. It demonstrates that obstruction can be a property of the*context space*, not of the data: because the macroeconomic regime axis is a circle, any model transported around a full cycle is exposed to holonomy by construction\. We suspect this is why through\-the\-cycle credit models are recurrently surprised at regime turns, and we offer it as a testable hypothesis rather than a claim\.

#### Limitations\.

\(1\) There is no real data in this paper\. The simulations are controlled validations of a theory, and while §[8](https://arxiv.org/html/2608.11252#S8)removes the circularity of the original design by letting holonomy emerge from effect modification rather than injecting it, an emergent mechanism in a model we wrote is still a model we wrote\. The paper should be read as theory with simulation support, and §[9](https://arxiv.org/html/2608.11252#S9.SS0.SSS0.Px4)states the real\-data study that would change that\. \(2\) Correlated errors across overlaps are not handled: precision whitening \(§[8](https://arxiv.org/html/2608.11252#S8)\) addresses unequal variances but not covariance, which would require the full inverse\-covariance metric and an estimate of it\. \(3\) Evidence gaps and obstruction interact, but not symmetrically—gaps*hide*obstruction, so our estimates are conservative; distinguishing “no bridge” from “incoherent bridge” still needs the sheaf\-theoretic treatment with non\-constant stalks\. \(4\) We treat contexts as given; eliciting the right factorisation is the hard applied problem and has no theory here\. \(5\) Real\-valued claims only; categorical and ordinal claims require cohomology with other coefficients, where the Hodge decomposition is unavailable\.

#### Real\-data validation\.

The fastest credible real\-data test does not require any new collection\. Published network meta\-analyses report contrast estimates and their standard errors; reconstructing the nerve from a published network, applying the precision\-whitenedFF\-test and comparing against the node\-splitting results the original authors reported would place this framework directly against the established method on its own ground, using only numbers already in print\. A second test uses existing open assets\.Medea’s benchmarks, tools and the released yeast E\-MAP screen\[[1](https://arxiv.org/html/2608.11252#bib.bib1)\]permit a direct test: construct the nerve over the 29 cell type×\\times5 disease contexts, populateω\\omegafrom PINNACLE and TranscriptFormer contrasts, and test whether harmonic energy predicts \(a\) disagreement amongMultiRoundDiscussionpanellists and \(b\) which of the agent’s confident answers are wrong\. Prediction: the79\.1%79\.1\\%abstention rate of the literature\-only ablation and the1\.8%1\.8\\%rate of the LLM\-only ablation bracket a harmonic\-gated rate that is both lower than the former and better targeted than the latter\. If harmonic energy fails to predict panel disagreement on real agent traces, the theory is wrong in the way that matters\.

## References

- \[1\]P\. Sui, M\. M\. Li, B\. P\. Munson, S\. Gao, W\. Shen, V\. Giunchiglia, A\. Shen, Y\. Huang, Z\. Kong, K\. Licon, T\. Ideker, M\. Zitnik\. Medea: An AI agent for therapeutic reasoning across biological contexts\.*bioRxiv*2026\.01\.16\.696667 \(2026\)\.
- \[2\]J\. C\.\-Y\. Chen, S\. Saha, M\. Bansal\. ReConcile: Round\-table conference improves reasoning via consensus among diverse LLMs\.*ACL*\(2024\)\.
- \[3\]S\. Yao et al\. ReAct: Synergizing reasoning and acting in language models\.*ICLR*\(2023\)\.
- \[4\]J\. Pearl, E\. Bareinboim\. External validity: From do\-calculus to transportability across populations\.*Statistical Science*29\(4\):579–595 \(2014\)\.
- \[5\]E\. Bareinboim, J\. Pearl\. Causal inference and the data\-fusion problem\.*PNAS*113\(27\):7345–7352 \(2016\)\.
- \[6\]G\. Lu, A\. E\. Ades\. Assessing evidence inconsistency in mixed treatment comparisons\.*JASA*101\(474\):447–459 \(2006\)\.
- \[7\]S\. Dias, N\. J\. Welton, D\. M\. Caldwell, A\. E\. Ades\. Checking consistency in mixed treatment comparison meta\-analysis\.*Statistics in Medicine*29\(7–8\):932–944 \(2010\)\.
- \[8\]J\. P\. T\. Higgins, D\. Jackson, J\. K\. Barrett, G\. Lu, A\. E\. Ades, I\. R\. White\. Consistency and inconsistency in network meta\-analysis\.*Research Synthesis Methods*3\(2\):98–110 \(2012\)\.
- \[9\]X\. Jiang, L\.\-H\. Lim, Y\. Yao, Y\. Ye\. Statistical ranking and combinatorial Hodge theory\.*Mathematical Programming*127:203–244 \(2011\)\.
- \[10\]R\. Ghrist\.*Elementary Applied Topology*\. \(2014\)\.
- \[11\]M\. Robinson\.*Topological Signal Processing*\. Springer \(2014\)\.
- \[12\]J\. Curry\. Sheaves, cosheaves and applications\. PhD thesis, University of Pennsylvania \(2014\)\.
- \[13\]S\. Mac Lane, I\. Moerdijk\.*Sheaves in Geometry and Logic*\. Springer \(1992\)\.
- \[14\]R\. El\-Yaniv, Y\. Wiener\. On the foundations of noise\-free selective classification\.*JMLR*11:1605–1641 \(2010\)\.
- \[15\]Y\. Geifman, R\. El\-Yaniv\. Selective classification for deep neural networks\.*NeurIPS*\(2017\)\.
- \[16\]Board of Governors of the Federal Reserve System\. SR 11\-7: Guidance on Model Risk Management \(2011\)\.

相似文章

程序验证的智能体证明

arXiv cs.AI

本文在Clever基准的程序验证任务中,采用智能体证明框架评估Claude Code,在规范生成和端到端验证方面取得了超过98%的成功率,揭示出现有基准可能不足以评估现代智能体证明器的能力。