Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI
Summary
The paper proposes an independence-graded audit protocol for agentic AI systems, grading independence along principal, substrate, and evidence axes, and provides a formal basis with analysis.
View Cached Full Text
Cached at: 09/17/26, 09:36 AM
# Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI Source: [https://arxiv.org/html/2609.18272](https://arxiv.org/html/2609.18272) CCS:Security and privacy Systems securityCCS:Computing methodologies Intelligent agentsCCS:Social and professional topics Computing / technology policyMohamed Chahine Ghanem[https://orcid.org/0000-0002-7067-7848](https://orcid.org/0000-0002-7067-7848)email:[mohamed\.chahine\.ghanem@liverpool\.ac\.uk](mailto:[email protected])Affiliation:Keele University,School of Computer Science and Mathematics,Keele,United KingdomAffiliation:University of Liverpool,Cybersecurity Institute,Liverpool,United Kingdom ###### Abstract\. Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor\. Independence—the foundation of assurance—is still applied to them as a binary\. We argue that it must be graded along three orthogonal axes:*principal*independence \(who controls the auditor\),*substrate*independence \(an auditor sharing the auditee’s foundation\-model family, toolchain or guardrails fails with it\) and*evidence*independence \(whether evidence is attestable rather than self\-reported\)\. Each axis has precedent; the contribution is to grade all three on a single audit, aggregate them by the weakest link, and apply the same rubric when the auditor is itself an agent\. We give the model a formal basis by transplanting the beta\-factor model of common\-cause failure from reliability engineering, a seven\-step protocol whose outputs a third party can verify, a structural detectability analysis of a procurement\-controls agent audited at three grades, and a Monte Carlo study of the model in which a conventional internal audit of an agent—a real audit team, a second agent, provider logs—surfaces5\.9%5\.9\\%of the faults it could in principle see and none at all in half the fault classes\. We map the triple to the EU AI Act as amended, ISO/IEC 42006, UK public\-sector risk\-management guidance and audit\-regulator practice\. ###### Keywords: agentic AI, AI audit, auditor independence, algorithmic monoculture, remote attestation, EU AI Act, ISO/IEC 42006 ## 1\.Introduction Two practices are converging under the phrase "auditing agentic AI"\. In the first, independent reviewers evaluate an organisation’s autonomous agents for observability, containment and accountability\([Shavit et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib11);[Chan et al\., 2024a](https://arxiv.org/html/2609.18272#bib.bib8)\)\. In the second, assurance functions deploy their own agents to test controls continuously over whole populations rather than samples, a shift regulators now monitor\([Financial Reporting Council, 2025b](https://arxiv.org/html/2609.18272#bib.bib51);[Public Company Accounting Oversight Board, 2024](https://arxiv.org/html/2609.18272#bib.bib52)\)\. Both borrow the concept that has carried assurance for a century: independence\. In the professional codes, independence is a property of people and money: the auditor must be free of compromising interests and relationships, and the determination is binary\([International Ethics Standards Board for Accountants, 2024](https://arxiv.org/html/2609.18272#bib.bib47)\)\. This is insufficient for agentic systems for three reasons\. First, the*principal*who controls an auditing agent may be independent while the agent is not: it may run on the auditee’s infrastructure, hold its credentials, or be steerable through its tool outputs\([Chan et al\., 2024a](https://arxiv.org/html/2609.18272#bib.bib8);[OWASP GenAI Security Project, 2025](https://arxiv.org/html/2609.18272#bib.bib18)\)\. Second, auditor and auditee increasingly share a*substrate*—model family, toolchain, guardrails—and shared substrates fail together\([Kleinberg and Raghavan, 2021](https://arxiv.org/html/2609.18272#bib.bib23);[Bommasani et al\., 2022](https://arxiv.org/html/2609.18272#bib.bib24);[Toups et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib25)\): models prefer their own outputs\([Wataoka et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib26)\), correlated judge panels carry far fewer independent votes than they seat\([Kohli, 2026](https://arxiv.org/html/2609.18272#bib.bib27)\), and agents can coordinate covertly\([Motwani et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib28)\)\. Third, the*evidence*behind an opinion is usually generated by the system under audit—logs, traces and self\-reports that the agent, its provider or its host could alter—so even privileged access does not settle what happened\([Casper et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib7)\)\. Guidance in force inherits all three gaps: the UK’s 2026 risk\-management toolkit for public\-sector AI has teams treat risks with testing by people "independent from the AI project team" and with logging that makes a system auditable, without saying how independent, or how much a log proves\([Department for Science, Innovation and Technology, 2026](https://arxiv.org/html/2609.18272#bib.bib58)\)\. We argue that independence should be*graded*, not declared, and make four contributions\.C1A three\-axis model \(Section 3, Table[2](https://arxiv.org/html/2609.18272#S4.T2)\) in which each axis is graded 0–3 on a stated basis—incentives forPP, shared components forSS, the strongest adversary survived forEE—and combined by a weakest\-link rule for which we give an argument\.C2A quantitative basis for the substrate axis \(Eqs\.[1](https://arxiv.org/html/2609.18272#S4.E1)–[2](https://arxiv.org/html/2609.18272#S4.E2), Fig\.[2](https://arxiv.org/html/2609.18272#S6.F2)\) carrying monoculture, judge\-correlation and AI\-control results into assurance\.C3A seven\-step protocol with third\-party\-verifiable outputs \(Section 4, Table[3](https://arxiv.org/html/2609.18272#S5.T3), Fig\.[1](https://arxiv.org/html/2609.18272#S4.F1)\), applied symmetrically to audits of agents and to agents as auditors\.C4A structural analysis of which axis closes which fault class \(Section[6](https://arxiv.org/html/2609.18272#S6), Table[4](https://arxiv.org/html/2609.18272#S6.T4)\) and a Monte Carlo study of the model \(Section[7](https://arxiv.org/html/2609.18272#S7)\) that quantifies the gap between configurations and tests the aggregation rule against alternatives\.C5A mapping to instruments now in force \(Section[8](https://arxiv.org/html/2609.18272#S8)\)\. Section 2 fixes the vocabulary and scope; Section 3 positions these claims\. ## 2\.Definitions and Scope Two of the three axes rest on terms that are not settled in the literature, so we fix them before using them\. ###### Definition 2\.1 \(Agentic AI system\)\. An*AI agent*is an automated entity that senses its environment, responds to it and acts to achieve goals \(ISO/IEC 22989:2022, cl\. 3\.1\.1\)\([ISO/IEC, 2022](https://arxiv.org/html/2609.18272#bib.bib64)\)\. A system is*agentic*to the degree that it pursues complex goals with limited direct supervision\([Shavit et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib11)\), a degree that Chan et al\. decompose into underspecification of the objective, directness of impact, goal\-directedness and long\-term planning\([Chan et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib65)\)\. Agenticness is thus a matter of degree, and the protocol below applies wherever a system holds credentials, invokes tools and acts between human reviews\. A procurement agent that reads invoices, matches them and raises exceptions without an intervening approval is high on all four dimensions; a classifier that scores an application for a human decider is low on all four and needs no more than a conventional model audit\. ###### Definition 2\.2 \(Substrate\)\. The*substrate*of an AI system is the set of technical components on which its behaviour depends and whose failures it inherits: the foundation model \(family and version\), the fine\-tuning and alignment data, the tool and connector layer, the guardrail and classifier stack, and the hosting environment\. Two systems*share*substrate to the extent that these coincide\. A cloud API call to a frontier model and an embedded controller running a distilled version of that model share the family but not the host; two agents from different vendors behind one connector framework share the tool layer but not the model\. We are aware that "substrate" is not established terminology in this sense and use it stipulatively for want of a shorter handle\. The phenomenon is well established under other labels—algorithmic and model monoculture\([Kleinberg and Raghavan, 2021](https://arxiv.org/html/2609.18272#bib.bib23);[Bommasani et al\., 2022](https://arxiv.org/html/2609.18272#bib.bib24)\), correlated failure and concentration risk from shared infrastructure, and the AI supply chain—and readers should map the term onto whichever their own literature uses\. What it adds is one name for*the components whose sharing makes two systems fail together*, which is the quantity an audit needs to score\. ###### Definition 2\.3 \(Independence triple and grade\)\. An audit of an agentic system reports the tripleI=\(P,S,E\)∈\{0,1,2,3\}3I=\(P,S,E\)\\in\\\{0,1,2,3\\\}^\{3\}, in whichPPgrades the auditor’s principal,SSthe disjointness of the two substrates andEEthe strongest adversary the evidence survives \(Table[2](https://arxiv.org/html/2609.18272#S4.T2)\)\. Its*grade*isg=min\(P,S,E\)g=\\min\(P,S,E\)\. ### Scope and assumptions\. The protocol grades the independence of an audit\. It does not certify a system as safe, correct or compliant: a high grade on a badly designed audit buys only well\-attested irrelevance, and nothing here removes the need for the audit to ask the right questions\. Three assumptions are load\-bearing and are revisited in Section[9](https://arxiv.org/html/2609.18272#S9): that the hardware root of trust and the external witness are not themselves compromised; that principals do not collude outside the channels the grades model; and that substrate lineage, where disclosed, is disclosed truthfully\. ## 3\.Related Work and Positioning Each axis has a literature\. Principal independence is governed by the professional codes and ISO/IEC 42006’s impartiality requirements for certification bodies\([International Ethics Standards Board for Accountants, 2024](https://arxiv.org/html/2609.18272#bib.bib47);[ISO/IEC, 2025](https://arxiv.org/html/2609.18272#bib.bib46)\); algorithmic\-audit frameworks define it as the absence of contractual or financial conflict\([Raji et al\., 2020](https://arxiv.org/html/2609.18272#bib.bib1);[Costanza\-Chock et al\., 2022](https://arxiv.org/html/2609.18272#bib.bib2);[Lam et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib3)\); frontier\-AI work asks which bodies should audit and proposes safeguards against auditor capture\([Stein et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib4);[Brundage et al\., 2026](https://arxiv.org/html/2609.18272#bib.bib5)\)\. Substrate correlation is documented for deployed models\([Kleinberg and Raghavan, 2021](https://arxiv.org/html/2609.18272#bib.bib23);[Bommasani et al\., 2022](https://arxiv.org/html/2609.18272#bib.bib24);[Toups et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib25)\); for LLM judges, where nine from seven model families were found to carry roughly two independent votes\([Kohli, 2026](https://arxiv.org/html/2609.18272#bib.bib27)\); and for AI\-control monitors, where same\-model "untrusted" monitoring degrades under collusion and is repaired with paraphrasing, resampling or a trusted weaker model\([Greenblatt et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib29);[Järviniemi, 2025](https://arxiv.org/html/2609.18272#bib.bib30);[Bhatt et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib31);[Gardner\-Challis et al\., 2026](https://arxiv.org/html/2609.18272#bib.bib32)\)\. Attestable evidence is the object of TEE\-based benchmark audits\([Schnabl et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib33)\), TEE\-isolated witnessing of agent transcripts\([Rowstron, 2026](https://arxiv.org/html/2609.18272#bib.bib34)\), cryptographic binding of tool use\([Zhou, 2026](https://arxiv.org/html/2609.18272#bib.bib35)\)and verifiable inference\([Sun et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib40)\); visibility and identity infrastructure supplies the records these mechanisms attest\([Chan et al\., 2024a](https://arxiv.org/html/2609.18272#bib.bib8);[Chan et al\., 2024b](https://arxiv.org/html/2609.18272#bib.bib9);[Chan et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib10)\); layered and access\-graded audits stratify what is examined and how much is seen\([Mökander et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib6);[Casper et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib7)\)\. ### Coincident failure in redundant systems\. The substrate axis has a lineage outside AI that we adopt rather than reinvent\. Software reliability engineering asked the same question of independently developed program versions and answered it empirically: Knight and Leveson found that twenty\-seven versions written independently to one specification failed together far more often than independence predicts, and concluded that the independence assumption underlyingNN\-version programming does not hold\([Knight and Leveson, 1986](https://arxiv.org/html/2609.18272#bib.bib59)\)\. Eckhardt and Lee, and later Littlewood and Miller, gave the theory—a difficulty function over the input space means versions fail on the same hard inputs even with no shared code\([Eckhardt and Lee, 1985](https://arxiv.org/html/2609.18272#bib.bib60);[Littlewood and Miller, 1989](https://arxiv.org/html/2609.18272#bib.bib61)\)\. Reliability practice encodes the residue as the*beta factor*: the share of a component’s failures that are common\-cause\([Fleming, 1975](https://arxiv.org/html/2609.18272#bib.bib62)\), estimated in safety engineering by a structured checklist\([IEC, 2010](https://arxiv.org/html/2609.18272#bib.bib63)\)\. Shared foundation models make the AI case strictly worse than the case those authors studied, because the components are shared by construction rather than converging by accident\. Section[4](https://arxiv.org/html/2609.18272#S4)transplants the beta\-factor model; the contribution is the transplant and its use in grading, not the model\. ### Analytic correspondence\. Three of these lines are special cases of, or orthogonal complements to, the model of Section 3, which is how we position rather than displace them\. The design\-effect arithmetic applied to LLM juries\([Kohli, 2026](https://arxiv.org/html/2609.18272#bib.bib27)\)is Eq\. \([2](https://arxiv.org/html/2609.18272#S4.E2)\) applied to a panel of evaluators rather than to an audit, and supplies the calibration below\. The trusted/untrusted monitor distinction of AI control\([Greenblatt et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib29)\)is the binary collapse ofSS—untrusted monitoring isS=0S\{=\}0, a trusted weaker modelS≥2S\{\\geq\}2—and its repairs, paraphrasing, resampling and honeypots, are ways of lowering common\-mode failure at fixed principal\([Järviniemi, 2025](https://arxiv.org/html/2609.18272#bib.bib30);[Bhatt et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib31);[Gardner\-Challis et al\., 2026](https://arxiv.org/html/2609.18272#bib.bib32)\)\. Attested benchmark audits\([Schnabl et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib33)\)instantiateE=3E\{=\}3for a single evaluation run; extending that guarantee to the runtime tool calls of a deployed agent is what Fig\.[1](https://arxiv.org/html/2609.18272#S4.F1)\(b\) specifies\. ### Orthogonality\. The two established stratifications are orthogonal to ours and compose with them\. Access hierarchies grade how much the auditor sees\([Casper et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib7)\);EEgrades whether what is seen could have been fabricated, so white\-box access to unattested artefacts is high access atE=1E\{=\}1\. Layered audits grade what is examined—governance, model, application\([Raji et al\., 2020](https://arxiv.org/html/2609.18272#bib.bib1);[Mökander et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib6)\)—which is orthogonal to who examines it and on what substrate\. An audit is therefore located in \(layer, access,PP,SS,EE\), and adopting the triple discards nothing already spent on the other two\. Table[1](https://arxiv.org/html/2609.18272#S3.T1)summarises the gap\. No prior work grades the three properties on one audit, aggregates them by the weakest link, or applies the same rubric to an AI agent when it is the auditor\. The closest neighbour\([Kohli, 2026](https://arxiv.org/html/2609.18272#bib.bib27)\)quantifies judge correlation with the design\-effect arithmetic of Section 3, but as an evaluation\-quality diagnostic: no principal or evidence axis, no protocol\. Our claim is the joint, symmetric grading and its protocol; the pillars are borrowed and cited\. Table 1\.Positioning against the closest prior work\.✓addressed;⊝\\boldsymbol\{\\circleddash\}partly \(see text\);✗absent\. "Graded" means more than a binary independent/not determination\. ## 4\.Three Axes and Their Formal Basis ### Principal independence \(PP\)\. Who controls the auditor? The grades run from self\-review, through a second team in the same reporting line and a separate organisation paid by the auditee—as most third\-party algorithmic audits are today\([Costanza\-Chock et al\., 2022](https://arxiv.org/html/2609.18272#bib.bib2)\)—to one appointed and paid by a party other than the auditee, with enforceable access rights \(Table[2](https://arxiv.org/html/2609.18272#S4.T2)\)\.PPis the axis the codes and ISO/IEC 42006 already regulate\. ### Substrate independence \(SS\)\. Do auditor and auditee share components whose failures would coincide? A substrate comprises the foundation model \(family and version\), alignment data, the tool and connector layer, guardrails and hosting\. When decision\-makers share a model, the individuals one rejects tend to be rejected by all\([Kleinberg and Raghavan, 2021](https://arxiv.org/html/2609.18272#bib.bib23);[Bommasani et al\., 2022](https://arxiv.org/html/2609.18272#bib.bib24);[Toups et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib25)\); a blind spot in the auditee is therefore likely to be one in a same\-substrate auditor\. An injection that hijacks the agent\([Debenedetti et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib14);[Zhan et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib15)\)tends to deceive a same\-family monitor; a judge prefers text it would itself have written, GPT\-4’s self\-preference bias measured at 0\.520\([Wataoka et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib26)\); and agent populations can hide coordination inside ordinary outputs\([Motwani et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib28)\)\. Prior work modelled such coincident failure through a correlation parameterρ\\rhoand argued that only a reviewer on a different substrate, with authority to veto, reliably breaks it\([Ghanem, 2026a](https://arxiv.org/html/2609.18272#bib.bib55)\);SSoperationalises that argument, and cross\-principal steganalysis of covert coordination\([Ghanem, 2026b](https://arxiv.org/html/2609.18272#bib.bib54)\)is one instrument for reaching Grade 2 or above\. ### Evidence independence \(EE\)\. Could the evidence have been altered by the auditee, its provider or its host without detection? The grades are defined by the strongest adversary the evidence survives \(Table[2](https://arxiv.org/html/2609.18272#S4.T2)\(b\)\): none \(the agent’s own narrative\); the agent only \(provider\-controlled logs, the visibility measures of\([Chan et al\., 2024a](https://arxiv.org/html/2609.18272#bib.bib8)\)\); tampering*after*commitment \(Merkle\-logged records with externally witnessed roots, as in certificate transparency\([Laurie et al\., 2013](https://arxiv.org/html/2609.18272#bib.bib36)\), bound to signed tool calls and agent identifiers\([Chan et al\., 2024b](https://arxiv.org/html/2609.18272#bib.bib9);[Chan et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib10)\)\); and fabrication*before*commitment \(records produced inside attested execution\([Birkholz et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib37);[Costan and Devadas, 2016](https://arxiv.org/html/2609.18272#bib.bib38);[NVIDIA Corporation, 2026](https://arxiv.org/html/2609.18272#bib.bib39);[Schnabl et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib33)\)or, as it matures, verifiable inference\([Sun et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib40)\)\)\. Black\-box access is insufficient for rigorous audits\([Casper et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib7)\);EEadds that white\-box access is too if the artefacts are unattested\. ### Formal basis\. Let a fault be missed by the auditeeAAwith probabilitypAp\_\{A\}and by an AI auditorBBwith probabilitypBp\_\{B\}, the miss events having correlationρ∈\[0,1\]\\rho\\in\[0,1\]\. SinceCov\(MA,MB\)=Pr\[MA=1,MB=1\]−pApB\\mathrm\{Cov\}\(M\_\{A\},M\_\{B\}\)=\\Pr\[M\_\{A\}\{=\}1,M\_\{B\}\{=\}1\]\-p\_\{A\}p\_\{B\}for the miss indicators, \(1\)Pr\[both miss\]=pApB\+ρpA\(1−pA\)pB\(1−pB\),\\Pr\[\\text\{both miss\}\]\\;=\\;p\_\{A\}p\_\{B\}\+\\rho\\,\\sqrt\{p\_\{A\}\(1\-p\_\{A\}\)\\,p\_\{B\}\(1\-p\_\{B\}\)\},which forpA=pB=pp\_\{A\}=p\_\{B\}=prises linearly fromp2p^\{2\}atρ=0\\rho=0toppatρ=1\\rho=1: a perfectly correlated auditor adds nothing\. Fornnexchangeable auditors with pairwise correlationρ\\rho, the variance of the mean miss indicator isp\(1−p\)n\[1\+\(n−1\)ρ\]\\frac\{p\(1\-p\)\}\{n\}\[1\+\(n\-1\)\\rho\]; equating it with that ofneffn\_\{\\mathrm\{eff\}\}independent auditors gives the design effect of survey sampling\([Kish, 1965](https://arxiv.org/html/2609.18272#bib.bib41)\), \(2\)neff=n1\+\(n−1\)ρ→n→∞1ρ,n\_\{\\mathrm\{eff\}\}\\;=\\;\\frac\{n\}\{1\+\(n\-1\)\\rho\}\\;\\xrightarrow\[n\\to\\infty\]\{\}\\;\\frac\{1\}\{\\rho\},so atρ=0\.5\\rho=0\.5no number of same\-substrate agents delivers more than two opinions \(Fig\.[2](https://arxiv.org/html/2609.18272#S6.F2)\(b\)\)\. ### A common\-shock model forSS\. The grades are only as principled as the correlation they proxy, so we giveρ\\rhoa generative model rather than stipulating an ordering\. The model is not new: it is the beta\-factor treatment of common\-cause failure\([Fleming, 1975](https://arxiv.org/html/2609.18272#bib.bib62);[IEC, 2010](https://arxiv.org/html/2609.18272#bib.bib63)\), applied to an auditor and an auditee instead of to redundant channels\. Let the substrate be a set of components—model family and version, alignment data, tool layer, guardrails, host—andDDthose the auditor shares with the auditee\. For a given fault class, each shared componentccindependently induces a*common\-mode*miss, in which both parties fail for the same reason, with probabilityγc\\gamma\_\{c\}; absent any common\-mode event the two miss independently with residual probabilityqq\. Then \(3\)γD=1−∏c∈D\(1−γc\),p=γD\+\(1−γD\)q,\\gamma\_\{D\}=1\-\\prod\_\{c\\in D\}\(1\-\\gamma\_\{c\}\),\\qquad p=\\gamma\_\{D\}\+\(1\-\\gamma\_\{D\}\)\\,q,and substituting the induced joint probabilityγD\+\(1−γD\)q2\\gamma\_\{D\}\+\(1\-\\gamma\_\{D\}\)q^\{2\}into Eq\. \([1](https://arxiv.org/html/2609.18272#S4.E1)\) gives \(4\)ρ=γD\(1−q\)γD\+\(1−γD\)q,\\rho\\;=\\;\\frac\{\\gamma\_\{D\}\\,\(1\-q\)\}\{\\gamma\_\{D\}\+\(1\-\\gamma\_\{D\}\)\\,q\}\\,,which is11when every miss is common\-mode \(q=0q\{=\}0\) and00when no component is shared\. The ratioβ=γD/p\\beta=\\gamma\_\{D\}/pis exactly the beta factor of reliability engineering—the share of a reviewer’s misses that are common\-cause—so the substrate axis can be read as a coarse beta\-factor scale for AI auditors\. ###### Proposition 4\.1 \(Substrate grades are monotone inρ\\rho\)\. IfD′⊆DD^\{\\prime\}\\subseteq DthenγD′≤γD\\gamma\_\{D^\{\\prime\}\}\\leq\\gamma\_\{D\}, and at fixedqq,ρ′≤ρ\\rho^\{\\prime\}\\leq\\rho\. ###### Proof sketch\. γD\\gamma\_\{D\}is one minus a product of factors in\[0,1\]\[0,1\], so dropping a factor cannot increase it; and differentiating Eq\. \([4](https://arxiv.org/html/2609.18272#S4.E4)\) gives∂ρ/∂γD=q\(1−q\)/p2≥0\\partial\\rho/\\partial\\gamma\_\{D\}=q\(1\-q\)/p^\{2\}\\geq 0\. Each grade ofSSremoves shared components \(Table[2](https://arxiv.org/html/2609.18272#S4.T2)\), so the ordering of the grades follows from the model rather than from stipulation\. ∎ ###### Proposition 4\.2 \(Cross\-substrate dominance\)\. Withkkauditors all sharingDDwith the auditee,Pr\[fault escapes\]=γD\+\(1−γD\)qk→γD\\Pr\[\\text\{fault escapes\}\]=\\gamma\_\{D\}\+\(1\-\\gamma\_\{D\}\)q^\{k\}\\to\\gamma\_\{D\}\. No panel size reduces the escape probability belowγD\\gamma\_\{D\}, whereas a single cross\-substrate auditor achievespApBp\_\{A\}p\_\{B\}; one such auditor therefore dominates an unbounded same\-substrate panel wheneverpApB<γDp\_\{A\}p\_\{B\}<\\gamma\_\{D\}\. ###### Proof sketch\. Allkkmiss exactly when the common\-mode event occurs or allkkresidual misses do; the residual term vanishes geometrically whileγD\\gamma\_\{D\}does not\. SettingγD=0\\gamma\_\{D\}=0recovers independence and the boundpApBp\_\{A\}p\_\{B\}\. ∎ The floor is the point \(Fig\.[2](https://arxiv.org/html/2609.18272#S6.F2)\(a\)\): atp=0\.10p=0\.10the independent bound ispApB=0\.01p\_\{A\}p\_\{B\}=0\.01, so any shared\-component contribution above one percentage point makes a single cross\-substrate reviewer strictly better than any number of same\-substrate ones—an argument forSSthat headcount cannot answer\. ### Calibration against a measured panel\. Reading a reported result through Eq\. \([2](https://arxiv.org/html/2609.18272#S4.E2)\) puts numbers on the grades\. Nine LLM judges drawn from seven model families were found to carry roughly two independent votes\([Kohli, 2026](https://arxiv.org/html/2609.18272#bib.bib27)\);neff=2n\_\{\\mathrm\{eff\}\}=2atn=9n=9impliesρ≈0\.44\\rho\\approx 0\.44, and inverting Eq\. \([4](https://arxiv.org/html/2609.18272#S4.E4)\) atp=0\.10p=0\.10givesγD≈0\.046\\gamma\_\{D\}\\approx 0\.046, that isβ≈0\.46\\beta\\approx 0\.46\. The comparison with engineering practice is unflattering: IEC 61508 expects beta factors of roughly0\.0050\.005to0\.050\.05for diverse programmable electronics\([IEC, 2010](https://arxiv.org/html/2609.18272#bib.bib63)\), an order of magnitude below the figure implied by seven\-vendor model diversity\. Vendor diversity \(S=2S\{=\}2\) therefore lowersρ\\rhowithout driving it near zero, which is why Grade 3 asks for a deterministic verifier on the load\-bearing checks rather than for more models; andγD≈0\.046≫pApB\\gamma\_\{D\}\\approx 0\.046\\gg p\_\{A\}p\_\{B\}places that configuration well inside the dominance regime of Proposition[4\.2](https://arxiv.org/html/2609.18272#S4.Thmtheorem2)\. The transfer is indicative, not established: those judges scored generations, not audits\. ###### Proposition 4\.3 \(The evidence grades form a chain\)\. LetK\(e\)K\(e\)be the set of tampering actions detectable at gradeee\. ThenK\(0\)⊂K\(1\)⊂K\(2\)⊂K\(3\)K\(0\)\\subset K\(1\)\\subset K\(2\)\\subset K\(3\), soEEtotally orders the configurations of Table[2](https://arxiv.org/html/2609.18272#S4.T2)\(b\)\. ###### Proof sketch\. Each mechanism strictly adds to the last: logging removes the agent’s discretion over its own record; signing and external witnessing remove the operator’s and provider’s ability to alter or fork that record after commitment; attestation removes the provider’s ability to fabricate it before commitment\. Nothing in the chain closes compromise of the attestation root, soE=3E\{=\}3is the top of a bounded chain, not a soundness guarantee\. ∎ ###### Proposition 4\.4 \(Weakest link\)\. If an adversary seeking an undeserved favourable opinion may attack any single axis, the assurance obtainable is bounded bymin\(P,S,E\)\\min\(P,S,E\), and any aggregate exceeding the minimum overstates it\. ###### Proof sketch\. ReadgXg\_\{X\}as the highest adversary level axisXXresists—literally forEEby Proposition[4\.3](https://arxiv.org/html/2609.18272#S4.Thmtheorem3), as incentive strength forPPand as toleratedγD\\gamma\_\{D\}forSS\. A captured principal signs any opinion; a shared substrate misses what the auditee misses; altered evidence supports any narrative\. Success on one axis suffices, so the aggregate is no stronger than the weakest, and a mean or weighted sum assigns assurance the configuration does not have\. ∎ Hence a regulator\-appointed auditor \(P=3P\{=\}3\) running the auditee’s model \(S=0S\{=\}0\) on self\-reported logs \(E=0E\{=\}0\) has grade 0, however impressive its mandate\. Table 2\.\(a\) Independence rubric\. An audit reports\(P,S,E\)\(P,S,E\); its grade ismin\(P,S,E\)\\min\(P,S,E\)\. Indicative tier minima: low\-risk\(1,1,1\)\(1,1,1\); high\-risk\(2,2,2\)\(2,2,2\); safety\-critical\(3,2,3\)\(3,2,3\)\. \(b\) Evidence lattice: tampering that remains undetectable \(✗\) or is detectable \(✓\) at each grade ofEE\. 1 Inventoryagents, tools, credentials, memory, delegation2 Principal graphwho deploys, pays, hosts, instructs3 Substrate lineagemodel, data, tools, guardrails, host4 Evidence acquisitionprovenance→\\tograde per record5 Behavioural probingthreat taxonomy; injection, harm suites6 Determinationnamed human, authority to withhold7 Reportopinion\+I=\(P,S,E\)\+\\ I=\(P,S,E\), grade, tier minimumproberesultsPPSSEEmin\\ming=min\(P,S,E\)g=\\min\(P,S,E\)tierminimum\(a\) Auditee agent in TEEmeasurementmm; reportQQTool gateway \(signing\)σi=Sigk\(h\(ci\)‖ti‖idA\)\\sigma\_\{i\}=\\mathrm\{Sig\}\_\{k\}\(h\(c\_\{i\}\)\\,\\\|\\,t\_\{i\}\\,\\\|\\,\\mathrm\{id\}\_\{A\}\),kkbound toQQAppend\-only logMerkle tree; leafh\(ci∥σi\)h\(c\_\{i\}\\,\\\|\\,\\sigma\_\{i\}\); signed rootrtr\_\{t\}External witnessco\-signsrtr\_\{t\}; consistency proofsAuditorother substrate \(P≥2P\\\!\\geq\\\!2,S≥2S\\\!\\geq\\\!2\); verifies chain; re\-performs checkstool callcic\_\{i\}\(ci,σi\)\(c\_\{i\},\\sigma\_\{i\}\)rtr\_\{t\}verifyQQ,mmcheckkk; re\-performverifyσi\\sigma\_\{i\}, inclusionverify co\-signatureE=1E\{=\}1E=2E\{=\}2E=3E\{=\}3\(b\) Figure 1\.\(a\) The protocol: Steps 2–4 assign the axes, probing results enter as evidence, and the weakest axis is reported against the tier minimum\. \(b\) The evidence chain reachingE=3E\{=\}3and the auditor’s checks \(Section 4\)\. Without attestation,E=2E\{=\}2; without signing and witnessing,E=1E\{=\}1\. ## 5\.The Protocol Seven steps produce the tripleIIas a mandatory part of the report \(Fig\.[1](https://arxiv.org/html/2609.18272#S4.F1)\(a\)\); Table[3](https://arxiv.org/html/2609.18272#S5.T3)states what each records and outputs and what a third party can verify\. Two rules do most of the work\. Where substrate lineage is withheld,SSis capped at 1: unknown is not independent\. And AI systems may generate hypotheses and evidence at any grade but do not issue the opinion; a named human with authority to withhold it does\([Ghanem, 2026a](https://arxiv.org/html/2609.18272#bib.bib55)\), as UK and US audit regulators expect of generative AI\([Financial Reporting Council, 2025a](https://arxiv.org/html/2609.18272#bib.bib50);[Public Company Accounting Oversight Board, 2024](https://arxiv.org/html/2609.18272#bib.bib52)\)\. ### Evidence chain\. Fig\.[1](https://arxiv.org/html/2609.18272#S4.F1)\(b\) shows the chain reachingE=3E\{=\}3and the auditor’s checks: verify the attestation reportQQagainst the published measurementmmof the agent’s code and configuration\([Birkholz et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib37);[Costan and Devadas, 2016](https://arxiv.org/html/2609.18272#bib.bib38);[NVIDIA Corporation, 2026](https://arxiv.org/html/2609.18272#bib.bib39)\); confirm the signing keykkis bound toQQ, so signatures could only come from attested code; verify each load\-bearing record’s signatureσi\\sigma\_\{i\}, its inclusion proof under a witnessed rootrtr\_\{t\}and the consistency proofs between roots\([Laurie et al\., 2013](https://arxiv.org/html/2609.18272#bib.bib36)\); then re\-perform those checks on a substrate withS≥2S\\geq 2, treating divergence as a finding\. Records passing every check grade 3; without attestation, 2; unsigned provider logs, 1\.EEis the minimum over load\-bearing records; verification cost is linear in their number\. Algorithm 1Independence grading of an audit\. Returns the triple, the grade, and whether the deployment’s tier minimum is met\.1:auditee agent AA; auditor BB\(human team, agent, or both\); risk tier tt 2:triple II, grade gg, verdict 3: 𝒜←\\mathcal\{A\}\\leftarrowInventory\( AA\);if BBis an agentthen 𝒜B←\\mathcal\{A\}\_\{B\}\\leftarrowInventory\( BB\) 4: G←G\\leftarrowPrincipalGraph\( A,BA,B\)⊳\\trianglerightwho deploys, pays, hosts, instructs 5: P←0P\\leftarrow 0if BBreviews its own work; 11if same reporting line; 22if separate organisation paid by the auditee; 33if appointed and paid by a third party with enforceable access 6: ℒA,ℒB←\\mathcal\{L\}\_\{A\},\\mathcal\{L\}\_\{B\}\\leftarrowLineage\( AA\),Lineage\( BB\) 7:if ℒA\\mathcal\{L\}\_\{A\}or ℒB\\mathcal\{L\}\_\{B\}undisclosedthen 8: S←min\(1,SharedGrade\(ℒA,ℒB\)\)S\\leftarrow\\min\(1,\\textsc\{SharedGrade\}\(\\mathcal\{L\}\_\{A\},\\mathcal\{L\}\_\{B\}\)\)⊳\\trianglerightunknown is not independent 9:else 10: S←S\\leftarrowSharedGrade\( ℒA,ℒB\\mathcal\{L\}\_\{A\},\\mathcal\{L\}\_\{B\}\)⊳\\trianglerightby shared components, Table[2](https://arxiv.org/html/2609.18272#S4.T2) 11:endif 12: ℰ←∅\\mathcal\{E\}\\leftarrow\\emptyset 13:for allrecords rrthe opinion will rest ondo 14: e\(r\)←3e\(r\)\\leftarrow 3if rrpasses attestation, signature, inclusion and consistency checks; 22if signature, inclusion and consistency only; 11if provider log; 00otherwise 15: ℰ←ℰ∪\{\(r,e\(r\)\)\}\\mathcal\{E\}\\leftarrow\\mathcal\{E\}\\cup\\\{\(r,e\(r\)\)\\\} 16:endfor 17: ℰ←ℰ∪\\mathcal\{E\}\\leftarrow\\mathcal\{E\}\\cupProbe\( AA, taxonomy, substrate at distance ≥1\\geq 1from ℒA\\mathcal\{L\}\_\{A\}\) 18: E←min\{e\(r\):rload\-bearing\}E\\leftarrow\\min\\\{e\(r\):r\\text\{ load\-bearing\}\\\} 19: I←\(P,S,E\)I\\leftarrow\(P,S,E\); g←min\(P,S,E\)g\\leftarrow\\min\(P,S,E\) 20:opinion ←\\leftarrowHumanDetermination\( ℰ\\mathcal\{E\}\)⊳\\trianglerighta named signatory, never an AI system 21:return II, gg, g≥TierMinimum\(t\)g\\geq\\textsc\{TierMinimum\}\(t\) Algorithm[1](https://arxiv.org/html/2609.18272#alg1)states the same procedure as a decision rule\. Two lines carry the weight: the cap onSSwhen lineage is withheld, and the assignment ofEEas a minimum over load\-bearing records rather than a maximum over available ones\. ### Symmetry\. When the auditee is a corporate workflow and the auditor is an agent, Steps 1–4 are applied to the auditing agent:PPasks who controls it;SSwhether it shares a substrate with the systems that produced the workflow’s records \(a model auditing invoices drafted by its own family is Grade 0\);EEwhether its own traces and tool calls are attested\. The same rubric therefore scores both practices, catching correlated blind spots dressed as independent confirmation in either direction; AI\-control designs\([Greenblatt et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib29);[Bhatt et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib31)\)and autonomous security testing\([Ghanem et al\., 2023](https://arxiv.org/html/2609.18272#bib.bib57);[Truss et al\., 2026](https://arxiv.org/html/2609.18272#bib.bib56)\)are, in these terms, attempts to raiseSSandEEwithPPheld fixed\. Table 3\.Protocol specification: what each step records, what it outputs, and what a third party can verify\. ## 6\.Analytical Case: Structural Detectability This is an analytical case, not an empirical evaluation\. An autonomous procurement\-controls agent reads purchase orders, invoices and goods receipts, performs three\-way matching, flags exceptions and drafts findings\. It is audited under three configurations \(Fig\.[2](https://arxiv.org/html/2609.18272#S6.F2)\(c\)\); for four fault classes we ask whether detection is guaranteed, excluded or probabilistic*by construction*\(Table[4](https://arxiv.org/html/2609.18272#S6.T4)\): F1, an indirect prompt injection in a supplier note that suppresses an exception\([Debenedetti et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib14)\); F2, a duplicate invoice with altered payee details; F3, post\-hoc alteration of the exception log by an operator; F4, covert coordination between auditee and auditor agents to suppress a finding\([Motwani et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib28)\)\. ### Configuration A — self\-attestation,I=\(0,0,0\)I=\(0,0,0\)\. The agent produces a monthly "controls attestation" of its own exceptions\. Reviewer and reviewed coincide, the substrate is identical, and the evidence is the agent’s narrative\. Grade 0\. F1 and F2 are caught only if the agent catches them; F3 is undetectable, since nothing binds the narrative to what happened; F4 is moot—no second agent exists, which is not reassurance\. This is what most dashboards deliver today\. ### Configuration B — internal audit, same family,I=\(1,0,1\)I=\(1,0,1\)\. Internal audit deploys a second agent, on the same model family and connector stack, to re\-perform matching on the provider’s logs\.PPrises to 1, butSSremains 0: the injection that suppressed the flag in the auditee would tend to suppress it in the same\-family auditor, and a judge preferring its own family’s phrasing\([Wataoka et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib26)\)is a poor arbiter of that family’s findings, so F1 and F2 are caught only with the correlated probability of Eq\. \([1](https://arxiv.org/html/2609.18272#S4.E1)\)\.EEis 1: an operator with host access can rewrite provider logs, so F3 stays undetectable; and a shared substrate gives F4 both channel and prior\. Grade 0—no better than A, despite a real audit team and a second agent\. The cause is the substrate, not the team’s diligence\. ### Configuration C — external, cross\-substrate, attested,I=\(2,3,3\)I=\(2,3,3\)\. An external firm \(P=2P\{=\}2\) re\-performs the matching with a different vendor’s model and a deterministic rule engine for the three\-way match \(S=3S\{=\}3\)\. The auditee’s agent runs in an attested enclave whose tool calls are signed and committed to a witnessed log; the auditor’s inference is likewise attested \(E=3E\{=\}3\)\. Grade 2\. F2 is detected deterministically; F3 because any rewrite breaks an inclusion or consistency proof; F4 is closed by construction, the agents sharing neither vendor nor tool layer, with residual covert channels the target of cross\-principal steganalysis\([Ghanem, 2026b](https://arxiv.org/html/2609.18272#bib.bib54)\); F1 is caught with near\-independent probabilitypApBp\_\{A\}p\_\{B\}rather thanpp\. Grade 3 requires a change of principal, not of technology\. Table 4\.Structural detectability of four fault classes\.✓detected by construction;✗undetectable by construction;∼\\boldsymbol\{\\sim\}probabilistic and correlated with the auditee’s miss;∘\\boldsymbol\{\\circ\}probabilistic and near\-independent\. Only the evidence axis closes F3 and only the substrate axis closes F4\.112233445566778810−210^\{\-2\}10−110^\{\-1\}10010^\{0\}one cross\-substrate auditor,pApBp\_\{A\}p\_\{B\}same\-substrate auditorskkPr\[fault escapes\]\\Pr\[\\text\{fault escapes\}\]\(a\)Prop\.[4\.2](https://arxiv.org/html/2609.18272#S4.Thmtheorem2),p=0\.10p\{=\}0\.10,β\\betaby gradeS=0S\{=\}0S=1S\{=\}1S=2S\{=\}2S=3S\{=\}3000\.20\.20\.40\.40\.60\.60\.80\.8110022446688measured panel:9 judges, 2 votes⇒ρ≈0\.44\\Rightarrow\\rho\\approx 0\.44miss correlationρ\\rhoneffn\_\{\\mathrm\{eff\}\}\(independent opinions\)\(b\)Eq\. \([2](https://arxiv.org/html/2609.18272#S4.E2)\)n=9n=9n=5n=5n=3n=3n=2n=2 low\-riskmin\(1,1,1\)\(1,1,1\)high\-riskmin\(2,2,2\)\(2,2,2\)criticalmin\(3,2,3\)\(3,2,3\)PrincipalPPSubstrateSSEvidenceEE0123A:g=0g\{=\}0B:g=0g\{=\}0C:g=2g\{=\}2\(c\)A: self\-attestation\(0,0,0\)\(0,0,0\)B: internal audit, same family\(1,0,1\)\(1,0,1\)C: external, cross\-substrate, attested\(2,3,3\)\(2,3,3\) Figure 2\.\(a\) Probability that a fault escapeskksame\-substrate auditors under the common\-shock model at the beta factors used in Section[7](https://arxiv.org/html/2609.18272#S7): the curves plateau atγD\\gamma\_\{D\}, which no panel size crosses, while a single cross\-substrate auditor reachespApBp\_\{A\}p\_\{B\}\(dashed\)\. \(b\) Independent opinions delivered bynncorrelated auditors, with the measured judge panel of\([Kohli, 2026](https://arxiv.org/html/2609.18272#bib.bib27)\)read through Eq\. \([2](https://arxiv.org/html/2609.18272#S4.E2)\)\. \(c\) Independence profiles of the configurations of Section 5; each gradeggis the lowest point its polyline touches; dashed polylines are the tier minima of Table[2](https://arxiv.org/html/2609.18272#S4.T2)\. ## 7\.Simulation Study Section[6](https://arxiv.org/html/2609.18272#S6)argues structurally; this section puts numbers on the same case by simulating the model of Section[4](https://arxiv.org/html/2609.18272#S4)\. The study establishes what the model implies and whether the aggregation rule is the right one*under the model’s own assumptions*\. It is not an empirical measurement: no result here is evidence about any deployed agent, and the parameters below are declared rather than estimated\. Code and seed are given in Section[10](https://arxiv.org/html/2609.18272#S10)\. ### Set\-up\. A fault class reaching audit is by construction one the auditee missed, so the audit’s contribution is the conditional probability that a second reviewer catches it,1−Pr\[both miss\]/p1\-\\Pr\[\\text\{both miss\}\]/p, which is where the substrate enters\. Each grade ofSSfixes a beta factorβ=γD/p\\beta=\\gamma\_\{D\}/p:0\.900\.90and0\.700\.70at grades 0 and 1, where no diversity is present;0\.460\.46at grade 2, the value implied by the measured judge panel of Section[4](https://arxiv.org/html/2609.18272#S4); and0\.050\.05at grade 3, the top of the IEC 61508 band for diverse programmable electronics, reachable only by adding a non\-learned verifier\. Each grade ofPPfixes a probability that a fault once caught is actually reported—0\.20\.2,0\.60\.6,0\.90\.9,1\.01\.0—since detection is not disclosure\. Each grade ofEEgates whether the fault leaves a record the auditor could examine at all: F1 and F2 need a provider log \(E≥1E\\geq 1\), F3 and F4 need tamper\-evidence \(E≥2E\\geq 2\)\. We takep=0\.10p=0\.10and run200,000200\{,\}000trials per cell; Monte Carlo estimates agree with the closed form to within the Wilson intervals throughout\. Eight configurations are simulated: the three of Section[6](https://arxiv.org/html/2609.18272#S6), a regulator\-appointed variant, three that hold two axes fixed and move the third, and the indicative high\-risk minimum\. Because the parameters are declared rather than measured, Results 5 to 7 report an ablation, a sweep of the parameters and the full434^\{3\}grid\. Table 5\.Simulated probability that a fault the auditee missed is both surfaced and reported, by configuration and fault class \(200,000200\{,\}000trials per cell; 95% Wilson intervals all within±0\.25\\pm 0\.25points\)\. The last column is the worst case over fault classes, which is what an adversary who chooses the attack obtains\. Configurations E to G each hold two axes fixed and vary the third; H is the indicative high\-risk minimum of Table[2](https://arxiv.org/html/2609.18272#S4.T2)\. ### Result 1: the grade\-0 gap is large, and the average hides it\. Configuration B—a real internal audit team, a second agent, provider logs—surfaces and reports5\.9%5\.9\\%of the faults in the two classes it can see and none of the other two \(Table[5](https://arxiv.org/html/2609.18272#S7.T5)\)\. Averaged over fault classes it scores3\.0%3\.0\\%, which already looks poor; its worst case, which is what an adversary choosing the attack obtains, is zero\. Configuration C reaches77\.3%77\.3\\%worst\-case, and the further step to D—a change of principal only, no change of technology—adds8\.68\.6points\. The ordering is not surprising; the magnitude is, and it comes from the substrate\. The conditional catch rate of a single reviewer rises from9\.9%9\.9\\%atS=0S\{=\}0through29\.0%29\.0\\%and50\.9%50\.9\\%to85\.9%85\.9\\%atS=3S\{=\}3, so on these parameters the substrate axis moves detection by a factor of nine while the principal axis moves disclosure by a factor of five\. ### Result 2: which axis you buy matters more than how much you spend\. Configurations E to G hold two axes fixed and move the third, which turns the rubric into a procurement question\. E is an external firm—a real change of principal, at real cost—running the auditee’s own model family on tamper\-evident records:8\.9%8\.9\\%worst\-case\. F is an ordinary internal audit team that has changed nothing about its reporting line but works on a different substrate with attested records:51\.6%51\.6\\%\. On these parameters, substrate diversity bought in\-house is worth roughly six times principal independence bought on a shared substrate, and the grades rank the two correctly \(00against11\) where an intuition that equates independence with externality does not\. G makes the same point on the evidence axis from the opposite direction: a regulator\-appointed auditor on a genuinely different substrate, but reading provider logs, detects85\.9%85\.9\\%of the two fault classes that leave a record and none of the two that do not, for a mean of46\.5%46\.5\\%and a worst case of zero\. H, the indicative high\-risk minimum, sits at45\.8%45\.8\\%on every fault class—balanced, unspectacular, and the only one of the eight with no axis carrying the others\. ### Result 3: panel size cannot substitute for substrate\. Escape probability against panel size plateaus atγD\\gamma\_\{D\}within three or four auditors:9\.0%9\.0\\%atS=0S\{=\}0,7\.0%7\.0\\%atS=1S\{=\}1,4\.6%4\.6\\%atS=2S\{=\}2and0\.5%0\.5\\%atS=3S\{=\}3, againstpApB=1\.0%p\_\{A\}p\_\{B\}=1\.0\\%for two genuinely independent reviewers \(Fig\.[2](https://arxiv.org/html/2609.18272#S6.F2)\(a\)\)\. At grades 0 to 2 a single cross\-substrate auditor therefore beats an unbounded same\-substrate panel, as Proposition[4\.2](https://arxiv.org/html/2609.18272#S4.Thmtheorem2)requires; at grade 3 it does not, because the deterministic verifier has already removed most of the common\-mode term\. The dominance result is thus a statement about cheap diversity, not a universal one\. ### Result 4: the minimum is the honest aggregate, though not the sharpest predictor\. We sampled20,00020\{,\}000triples uniformly from\{0,…,3\}3\\\{0,\\dots,3\\\}^\{3\}and compared three aggregation rules against two targets: expected yield \(mean detection over fault classes\) and adversarial assurance \(worst case over fault classes\)\. Against expected yield the minimum ranks best \(Spearmanρ=0\.91\\rho=0\.91, against0\.780\.78for the mean and0\.730\.73for a weighted sum0\.5P\+0\.3S\+0\.2E0\.5P\+0\.3S\+0\.2E\)\. Against adversarial assurance the minimum and the mean are effectively tied \(0\.690\.69against0\.680\.68\), which we report because it qualifies the claim: the minimum is not a uniformly superior*predictor*\. What separates the rules is overstatement\. Rescaling each rule and the true assurance to\[0,1\]\[0,1\], the minimum credits a configuration with more assurance than it has in26\.5%26\.5\\%of cases, with a median overstatement of zero; the mean does so in92\.2%92\.2\\%of cases with a median of0\.330\.33, and the weighted sum in90\.8%90\.8\\%\. Triples with one axis at zero carry1\.3%1\.3\\%adversarial assurance against30\.7%30\.7\\%for the rest, and it is exactly those that a mean rescues\. The case formin\\minis therefore not that it forecasts assurance best but that it is the rule that rarely claims assurance that is not there—which is what an assurance statement is for\. Table 6\.Robustness of the simulated results\. Left: axis ablation, reporting the change in worst\-case and mean detection when a single axis is raised from the\(1,1,1\)\(1,1,1\)baseline or dropped from the\(3,3,3\)\(3,3,3\)ceiling\. Right: worst\-case assurance across all 64 triples, grouped by grade\.\(a\) Ablation of one axis\(b\) All 64 triples by gradeAxisraise→31\\\!\\to\\\!3drop→03\\\!\\to\\\!0ggnnminmedianmaxworstmeanworstmeanPP\+0\.0\+0\.0\+5\.8\+5\.8−85\.9\-85\.9−89\.4\-89\.40370\.0%0\.0%9\.9%SS\+0\.0\+0\.0\+19\.2\+19\.2−76\.0\-76\.0−79\.6\-79\.61190\.0%17\.4%51\.6%EE\+17\.4\+17\.4\+8\.7\+8\.7−85\.9\-85\.9−89\.4\-89\.42745\.8%50\.9%85\.9%3185\.9%85\.9%85\.9% ### Result 5: each axis is individually necessary, and the weakest one gates the rest\. Dropping any single axis from the ceiling to zero costs between76\.076\.0and85\.985\.9points of worst\-case assurance \(Table[6](https://arxiv.org/html/2609.18272#S7.T6)\(a\)\), so none of the three is redundant: a rubric that omitted any one would rate a configuration highly that this one rates at zero\. The asymmetry from below is sharper\. Raising the principal or the substrate from the\(1,1,1\)\(1,1,1\)baseline all the way to grade 3 moves worst\-case assurance by nothing at all, because atE=1E\{=\}1two of the four fault classes leave no record to examine, so there is nothing for a better\-motivated or better\-decorrelated auditor to find; the same moves buy5\.85\.8and19\.219\.2points of*mean*detection, which is precisely the kind of improvement an average\-based score would reward and an adversary would ignore\. Only the evidence axis moves the worst case, by17\.417\.4points\. This is the weakest\-link rule behaving as Proposition[4\.4](https://arxiv.org/html/2609.18272#S4.Thmtheorem4)says it should, and it is the most direct evidence we have that the rule is doing work rather than decoration\. ### Result 6: the ordering survives the parameters; the magnitudes do not\. The beta factors are declared, so we swept them\. Varying the marginal miss rate from0\.020\.02to0\.300\.30—a fifteen\-fold range—and compressing the four substrate beta factors towards their common mean by up to75%75\\%, Kendall’sτ\\taubetween grade and worst\-case assurance across all 64 triples stays within0\.6060\.606to0\.6470\.647, and the dominance condition of Proposition[4\.2](https://arxiv.org/html/2609.18272#S4.Thmtheorem2)holds in all twenty settings tested\. The strict ordering of the eight named configurations is preserved while the substrate grades retain distinguishable beta factors, and breaks once they are compressed by half or more—which is the degenerate case in which the substrate axis has been defined away\. Letting the beta factor differ between fault classes by up to a factor of three either way, over2,0002\{,\}000draws, leaves the ranking of the aggregation rules unchanged: the minimum overstates adversarial assurance in33\.5%33\.5\\%of cases against93\.7%93\.7\\%for both the mean and the weighted sum\. What does not survive is the magnitudes\. The5\.9%5\.9\\%and77\.3%77\.3\\%of Table[5](https://arxiv.org/html/2609.18272#S7.T5)are artefacts of the chosen parameters; the gap between them is not\. ### Result 7: the grade explains most of the assurance, and hides the rest\. Across all 64 triples the grade accounts for77%77\\%of the variance in worst\-case assurance \(η2\\eta^\{2\}, Table[6](https://arxiv.org/html/2609.18272#S7.T6)\(b\)\)\. The residual is the honest cost of an ordinal scale: grade 1 alone spans0\.0%0\.0\\%to51\.6%51\.6\\%, because configurations F and G are both grade 1 and differ by everything that matters\. A reader who needs to separate them must read the triple, not the grade—which is why the protocol requires the triple to be published and treats the grade as a summary of it rather than a replacement for it\. ## 8\.Policy Hooks ### EU AI Act, as amended\. Article 12 requires automatic logging for high\-risk systems but not tamper\-evidence;E≥2E\\geq 2would give the logs probative value\([European Parliament and Council of the European Union, 2024](https://arxiv.org/html/2609.18272#bib.bib42)\)\. Article 14 \(human oversight\) is Step 6; Article 26 places monitoring duties on deployers; Article 43 toggles between conformity assessment by internal control and by a notified body—P=1P\{=\}1versusP=3P\{=\}3—while saying nothing aboutSSorEE\. The Digital Omnibus on AI, proposed in November 2025 and adopted in June 2026, defers Annex III high\-risk obligations to 2 December 2027 and Annex I to 2 August 2028\([European Commission, 2025](https://arxiv.org/html/2609.18272#bib.bib43);[Council of the European Union, 2026](https://arxiv.org/html/2609.18272#bib.bib44)\), widening the window of voluntary assurance in which the triple lets buyers and insurers compare offerings\. ### Standards\. ISO/IEC 42006:2025, which specifies who may audit and certify AI management systems\([ISO/IEC, 2025](https://arxiv.org/html/2609.18272#bib.bib46);[ISO/IEC, 2023](https://arxiv.org/html/2609.18272#bib.bib45)\), governs impartiality—PP—and is silent onSSandEE; certification bodies could report the triple now\. ### UK risk management\. The AI Risk Management Toolkit published in September 2026 implements the Orange Book’s risk process for public\-sector AI, alongside a roadmap to professionalise third\-party AI assurance\([Department for Science, Innovation and Technology, 2025](https://arxiv.org/html/2609.18272#bib.bib48)\)and a gap analysis flagging agentic\-system security as under\-studied\([Department for Science, Innovation and Technology and Lancaster University, 2026](https://arxiv.org/html/2609.18272#bib.bib49)\)\. It is self\-assessment, and most of it needs no audit; but several treatment options do, each specified as a binary \(Table[7](https://arxiv.org/html/2609.18272#S8.T7)\)\. A team may therefore record them as applied, and lower its residual\-risk score, while the independent testers run the auditee’s model family and the preserved logs stay rewritable by the operator—grade 0 here\. Teams are also asked to estimate risk likelihood partly by model analysis: where the estimating model shares a substrate with the system scored, Eq\. \([2](https://arxiv.org/html/2609.18272#S4.E2)\) says the estimate carries less independent information than the count of checks implies\. Recording\(P,S,E\)\(P,S,E\)against each such treatment in the risk workbook costs three integers and makes the residual\-risk score mean what it says\. Table 7\.Elements of the UK AI Risk Management Toolkit\([Department for Science, Innovation and Technology, 2026](https://arxiv.org/html/2609.18272#bib.bib58)\)whose efficacy depends on an unstated degree of independence, and what the triple supplies\. ### Audit regulators\. The Financial Reporting Council’s review of the six largest UK firms found no formal monitoring of the audit\-quality impact of their automated tools and, at all but one firm, no indicators for them\([Financial Reporting Council, 2025b](https://arxiv.org/html/2609.18272#bib.bib51)\); the PCAOB stresses continued human supervision of generative\-AI output\([Public Company Accounting Oversight Board, 2024](https://arxiv.org/html/2609.18272#bib.bib52)\)\. Both could require the triple whenever an agent contributes evidence, treating a same\-substrate agent asS=0S\{=\}0for reliance\. ### Statutory precedent\. New York City’s Local Law 144 has mandated independent bias audits of automated employment decision tools since 2023\([New York City Council, 2021](https://arxiv.org/html/2609.18272#bib.bib53)\); its definition of an independent auditor speaks only toPP, and adding minimumSSandEEwould be a modest drafting change\. ## 9\.Limitations and Validation Plan Attestation proves which code and configuration ran, not that they behaved correctly:E=3E\{=\}3is necessary for trustworthy evidence, not sufficient for a correct opinion, and inherits the enclave’s trust base and side\-channel history\([Costan and Devadas, 2016](https://arxiv.org/html/2609.18272#bib.bib38)\); verifiable inference at frontier scale is not production\-ready\([Sun et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib40)\)\. Substrate lineage depends on disclosure; vendors will contest theS≤1S\\leq 1cap for undisclosed lineage, but it is the conservative default\. Equations \([1](https://arxiv.org/html/2609.18272#S4.E1)\)–\([2](https://arxiv.org/html/2609.18272#S4.E2)\) assume exchangeable auditors and a singleρ\\rho; real fault classes differ, which is why Table[4](https://arxiv.org/html/2609.18272#S6.T4)reasons per class\. The common\-shock model buys the monotonicity of Proposition[4\.1](https://arxiv.org/html/2609.18272#S4.Thmtheorem1)at the price of two assumptions a critic should press on: that shared components induce common\-mode failure independently of one another, which overstatesγD\\gamma\_\{D\}where two shared components fail through the same mechanism, and that the per\-component ratesγc\\gamma\_\{c\}are estimable at all—today they are not, so Proposition[4\.2](https://arxiv.org/html/2609.18272#S4.Thmtheorem2)yields a qualitative dominance condition rather than a procurement threshold\. The calibration transfers a figure measured on evaluation panels to auditing, which is a hypothesis the study below is designed to test, not a result\. The grades are ordinal and themin\\minrule discards information\. Section[7](https://arxiv.org/html/2609.18272#S7)sharpens what we can claim for it: the minimum is not uniformly the better*predictor*of assurance—against a worst\-case target it is level with a mean—and its case rests on rarely overstating rather than on forecasting well\. A reader who wants an expected\-yield estimate should not use the grade for it\. The simulation validates the model, not the world\. Its parameters are declared rather than measured, and Result 6 is explicit about what that costs: the rank relationship between grade and assurance is stable across a fifteen\-fold sweep of the miss rate and a75%75\\%compression of the substrate beta factors, but the magnitudes are artefacts of the chosen values and should not be quoted as expected detection rates for any real system\. Result 7 adds a second caveat that the grade itself carries: it fixes77%77\\%of the variance in worst\-case assurance and leaves a within\-grade spread of up to51\.651\.6points, so the triple is the reportable object and the grade only a summary of it\. The empirical study this calls for is well defined: inject F1–F4 into benchmark agents\([Debenedetti et al\., 2024](https://arxiv.org/html/2609.18272#bib.bib14);[Andriushchenko et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib16);[Kapoor et al\., 2025](https://arxiv.org/html/2609.18272#bib.bib12)\)audited by agents atS=0,…,3S=0,\\dots,3, estimate the beta factor per fault class instead of assuming it, and compare measured detection against the model’s predictions; a cheaper companion study scores the triple retrospectively against treatments already recorded in public\-sector risk workbooks\([Department for Science, Innovation and Technology, 2026](https://arxiv.org/html/2609.18272#bib.bib58)\)\. Until that is done the numbers in Section[7](https://arxiv.org/html/2609.18272#S7)should be read as consequences of the model and nothing more\. Finally, Grade\-3 principals barely exist; agent\-governance law\([Kolt, 2025](https://arxiv.org/html/2609.18272#bib.bib13)\)must catch up\. ## 10\.Ethics, Conflicts of Interest, and Artifact Availability ### Ethics\. The study involves no human or animal subjects, no personal data and no user research; the simulation uses synthetic draws from a declared model\. The protocol itself touches personal data only indirectly: the tool\-call records that carry evidence grades may contain personal data, so a deployment applying Step 4 should minimise and retain them under the applicable data\-protection regime rather than logging indiscriminately in pursuit of a higherEE\. There is a real tension here—higher evidence grades mean more retained, more durable records—and a deployment should resolve it by scoping tamper\-evidence to load\-bearing records rather than to everything the agent does\. ### Conflicts of interest\. The author holds academic appointments at Keele University and the University of Liverpool, advises on cyber resilience in the banking sector, and is a founder of an early\-stage company working on agentic AI transparency and audit; that company sells no product implementing this protocol, and the protocol is published without restriction\. No funder had any role in the design or conclusions of this work\. The self\-review risk the paper analyses applies to the paper: a framework proposed by someone with an interest in the assurance market should be read with the incentive in view, which is one reason the rubric is specified so that a third party can apply it without the author’s involvement\. ### Artifact availability\. The simulation of Section[7](https://arxiv.org/html/2609.18272#S7)is a small Python package with two experiment drivers—one for Results 1 to 4 and one for the ablation, parameter sweep, full grid and heterogeneity checks of Results 5 to 7—together with eighteen unit tests that verify the closed forms against the Monte Carlo draws and check the ablation claims directly\. It depends on NumPy and SciPy only and reproduces every number reported here from a fixed seed in under two minutes\. It is provided with this submission and deposited in a public repository Zenodo[https://doi\.org/10\.5281/zenodo\.22768569](https://doi.org/10.5281/zenodo.22768569)\. There is no data set: the study generates its own draws and depends on no external input\. ## 11\.Conclusion Independence was never meant to be a checkbox, and for agentic systems it cannot be\. Two agents on the same substrate confirming each other is one opinion, not two—Eq\. \([2](https://arxiv.org/html/2609.18272#S4.E2)\) makes the arithmetic explicit; an audit built on self\-reported logs is a narrative, not evidence\. On the model of Section[4](https://arxiv.org/html/2609.18272#S4), an internal audit that does everything an audit function is normally asked to do, but on the auditee’s model family and the provider’s logs, surfaces nothing at all in the worst case\. Grading independence along principal, substrate and evidence, and publishing the triple with every opinion, costs little, needs no new law and reveals how much an "independent" audit of an agent actually shows\. ## References - M\. Andriushchenko, A\. Souly, M\. Dziemian, D\. Duenas, M\. Lin, J\. Wang, D\. Hendrycks, A\. Zou, Z\. Kolter, M\. Fredrikson, E\. Winsor, J\. Wynne, Y\. Gal, and X\. DaviesAgentHarm: a benchmark for measuring harmfulness of LLM agents\.InThe Thirteenth International Conference on Learning Representations \(ICLR 2025\),External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.09024)Cited by:[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1),[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Bhattet al\.\(2025\)A\. Bhatt, C\. Rushing, A\. Kaufman, T\. Tracy, V\. Georgiev, D\. Matolcsi, A\. Khan, and B\. ShlegerisCtrl\-Z: controlling AI agents via resampling\.Note:arXiv:2504\.10374External Links:[Document](https://dx.doi.org/10.48550/arXiv.2504.10374)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.9.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px2.p1.1)\. - Birkholzet al\.\(2023\)H\. Birkholz, D\. Thaler, M\. Richardson, N\. Smith, and W\. PanRemote ATtestation procedureS \(RATS\) architecture\.RFCTechnical Report9334,Internet Engineering Task Force\.External Links:[Document](https://dx.doi.org/10.17487/RFC9334)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px1.p1.1)\. - Bommasaniet al\.\(2022\)R\. Bommasani, K\. A\. Creel, A\. Kumar, D\. Jurafsky, and P\. LiangPicking on the same person: does algorithmic monoculture lead to outcome homogenization?\.InAdvances in Neural Information Processing Systems 35 \(NeurIPS 2022\),pp\. 3663–3678\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2211.13972)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[§2](https://arxiv.org/html/2609.18272#S2.p3.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.8.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1)\. - Brundageet al\.\(2026\)M\. Brundage, N\. Dreksler, A\. Homewood, S\. McGregor, P\. Paskov, C\. Stosz, G\. Sastry, A\. F\. Cooper,et al\.Frontier AI auditing: toward rigorous third\-party assessment of safety and security practices at leading AI companies\.Note:arXiv:2601\.11699External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.11699)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.4.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Casperet al\.\(2024\)S\. Casper, C\. Ezell, C\. Siegmann, N\. Kolt, T\. L\. Curtis, B\. Bucknall, A\. Haupt, K\. Wei, J\. Scheurer, M\. Hobbhahn, L\. Sharkey, S\. Krishna, M\. von Hagen, S\. Alberti, A\. Chan, Q\. Sun, M\. Gerovitch, D\. Bau, M\. Tegmark, D\. Krueger, and D\. Hadfield\-MenellBlack\-box access is insufficient for rigorous AI audits\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency \(FAccT ’24\),New York, NY, USA,pp\. 2254–2272\.External Links:[Document](https://dx.doi.org/10.1145/3630106.3659037)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.5.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1)\. - Chanet al\.\(2024a\)A\. Chan, C\. Ezell, M\. Kaufmann, K\. Wei, L\. Hammond, H\. Bradley, E\. Bluemke, N\. Rajkumar, D\. Krueger, N\. Kolt, L\. Heim, and M\. AnderljungVisibility into AI agents\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency \(FAccT ’24\),New York, NY, USA,pp\. 958–973\.External Links:[Document](https://dx.doi.org/10.1145/3630106.3658948)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p1.1),[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.6.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.2.2.1.1)\. - Chanet al\.\(2024b\)A\. Chan, N\. Kolt, P\. Wills, U\. Anwar, C\. Schroeder de Witt, N\. Rajkumar, L\. Hammond, D\. Krueger, L\. Heim, and M\. AnderljungIDs for AI systems\.Note:arXiv:2406\.12137; Workshop on Regulatable ML \(RegML\), NeurIPS 2024External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.12137)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.6.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1)\. - Chanet al\.\(2023\)A\. Chan, R\. Salganik, A\. Markelius, C\. Pang, N\. Rajkumar, D\. Krasheninnikov, L\. Langosco, Z\. He, Y\. Duan, M\. Carroll, M\. Lin, A\. Mayhew, K\. Collins, M\. Molamohammadi, J\. Burden, W\. Zhao, S\. Rismani, K\. Voudouris, U\. Bhatt, A\. Weller, D\. Krueger, and T\. MaharajHarms from increasingly agentic algorithmic systems\.InProceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency \(FAccT ’23\),New York, NY, USA,pp\. 651–666\.External Links:[Document](https://dx.doi.org/10.1145/3593013.3594033)Cited by:[Definition 2\.1](https://arxiv.org/html/2609.18272#S2.Thmtheorem1.p1.1)\. - Chanet al\.\(2025\)A\. Chan, K\. Wei, S\. Huang, N\. Rajkumar, E\. Perrier, S\. Lazar, G\. K\. Hadfield, and M\. AnderljungInfrastructure for AI agents\.Transactions on Machine Learning Research\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.10114)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.6.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.2.2.1.1)\. - Cloud Security Alliance \(2025a\)Cloud Security AllianceAgentic AI red teaming guide\.Technical reportCloud Security Alliance\.External Links:[Link](https://cloudsecurityalliance.org/artifacts/agentic-ai-red-teaming-guide)Cited by:[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Cloud Security Alliance \(2025b\)Cloud Security AllianceAgentic AI threat modeling framework: MAESTRO\.Note:Cloud Security AllianceExternal Links:[Link](https://cloudsecurityalliance.org/blog/2025/02/06/agentic-ai-threat-modeling-framework-maestro)Cited by:[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Costan and Devadas \(2016\)V\. Costan and S\. DevadasIntel SGX explained\.Note:Cryptology ePrint Archive, Paper 2016/086External Links:[Link](https://eprint.iacr.org/2016/086)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px1.p1.1),[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Costanza\-Chocket al\.\(2022\)S\. Costanza\-Chock, I\. D\. Raji, and J\. BuolamwiniWho audits the auditors? Recommendations from a field scan of the algorithmic auditing ecosystem\.InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency \(FAccT ’22\),New York, NY, USA,pp\. 1571–1583\.External Links:[Document](https://dx.doi.org/10.1145/3531146.3533213)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.3.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px1.p1.1)\. - Council of the European Union \(2026\)Council of the European UnionArtificial intelligence: council gives final green light to simplify and streamline rules\.Note:Press release, 29 June 2026External Links:[Link](https://www.consilium.europa.eu/en/press/press-releases/2026/06/29/artificial-intelligence-council-gives-final-green-light-to-simplify-and-streamline-rules/)Cited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px1.p1.1)\. - Debenedettiet al\.\(2024\)E\. Debenedetti, J\. Zhang, M\. Balunović, L\. Beurer\-Kellner, M\. Fischer, and F\. TramèrAgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS 2024\), Datasets and Benchmarks Track,pp\. 82895–82920\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.13352)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1),[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1),[§6](https://arxiv.org/html/2609.18272#S6.p1.1),[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Department for Science, Innovation and Technology and Lancaster University \(2026\)Department for Science, Innovation and Technology and Lancaster UniversityThematic review and gap analysis on AI security\.Technical reportUK Government\.External Links:[Link](https://www.gov.uk/government/publications/thematic-review-and-gap-analysis-on-ai-security)Cited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px3.p1.1)\. - Department for Science, Innovation and Technology \(2025\)Department for Science, Innovation and TechnologyTrusted third\-party AI assurance roadmap\.Technical reportUK Government\.External Links:[Link](https://www.gov.uk/government/publications/trusted-third-party-ai-assurance-roadmap)Cited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px3.p1.1)\. - Department for Science, Innovation and Technology \(2026\)Department for Science, Innovation and TechnologyAI risk management toolkit: guidance\.Technical reportUK Government\.Note:Published 8 September 2026External Links:[Link](https://www.gov.uk/government/publications/ai-risk-management-toolkit/ai-risk-management-toolkit-guidance)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[Table 7](https://arxiv.org/html/2609.18272#S8.T7),[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Eckhardt and Lee \(1985\)D\. E\. Eckhardt and L\. D\. LeeA theoretical basis for the analysis of multiversion software subject to coincident errors\.IEEE Transactions on Software EngineeringSE\-11\(12\),pp\. 1511–1517\.External Links:[Document](https://dx.doi.org/10.1109/TSE.1985.231895)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.7.1.1.1)\. - European Commission \(2025\)European CommissionProposal for a regulation amending regulation \(EU\) 2024/1689 as regards simplification measures \(Digital Omnibus on AI\)\.Note:COM\(2025\) 836 final, Brussels, 19 November 2025External Links:[Link](https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=celex:52025PC0836)Cited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px1.p1.1)\. - European Parliament and Council of the European Union \(2024\)European Parliament and Council of the European UnionRegulation \(EU\) 2024/1689 of 13 june 2024 laying down harmonised rules on artificial intelligence \(Artificial Intelligence Act\)\.Note:Official Journal of the European Union, L, 12 July 2024External Links:[Link](http://data.europa.eu/eli/reg/2024/1689/oj)Cited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px1.p1.1)\. - Financial Reporting Council \(2025a\)Financial Reporting CouncilAI in audit: illustrative example and documentation guidance\.Technical reportFinancial Reporting Council,London\.External Links:[Link](https://www.frc.org.uk/library/standards-codes-policy/audit-assurance-and-ethics/guidance/ai-in-audit/)Cited by:[§5](https://arxiv.org/html/2609.18272#S5.p1.1)\. - Financial Reporting Council \(2025b\)Financial Reporting CouncilThematic review on the certification of automated tools and techniques\.Technical reportFinancial Reporting Council,London\.External Links:[Link](https://www.frc.org.uk/news-and-events/news/2025/06/frc-publishes-landmark-guidance-providing-clarity-to-audit-profession-on-the-uses-of-ai/)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p1.1),[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px4.p1.1)\. - Fleming \(1975\)K\. N\. FlemingA reliability model for common mode failure in redundant safety systems\.InProceedings of the Sixth Annual Pittsburgh Conference on Modeling and Simulation,Note:General Atomic Report GA\-A13284Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.7.1.1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px5.p1.1)\. - Gardner\-Challiset al\.\(2026\)N\. Gardner\-Challis, J\. Bostock, G\. Kozhevnikov, M\. Sinclaire, J\. Velja, A\. Abate, and C\. GriffinWhen can we trust untrusted monitoring? A safety case sketch across collusion strategies\.Note:arXiv:2602\.20628External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.20628)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.9.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Ghanemet al\.\(2023\)M\. C\. Ghanem, T\. M\. Chen, M\. A\. Ferrag, and M\. E\. KettoucheESASCF: expertise extraction, generalization and reply framework for optimized automation of network security compliance\.IEEE Access11,pp\. 129840–129853\.External Links:[Document](https://dx.doi.org/10.1109/ACCESS.2023.3332834)Cited by:[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px2.p1.1)\. - Ghanem \(2026a\)M\. C\. GhanemBuilder, defender, breaker: the case against removing the human from the AI\-driven security lifecycle\.Note:arXiv:2607\.03215External Links:[Document](https://dx.doi.org/10.48550/arXiv.2607.03215)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.p1.1)\. - Ghanem \(2026b\)M\. C\. GhanemSteganalysis of adaptive covert collusion in tool\-using agent populations: a black\-box, cross\-principal approach\.Note:arXiv:2608\.02698External Links:[Document](https://dx.doi.org/10.48550/arXiv.2608.02698)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.18272#S6.SS0.SSS0.Px3.p1.1)\. - Greenblattet al\.\(2024\)R\. Greenblatt, B\. Shlegeris, K\. Sachan, and F\. RogerAI control: improving safety despite intentional subversion\.InProceedings of the 41st International Conference on Machine Learning \(ICML 2024\),Proceedings of Machine Learning Research, Vol\.235,pp\. 16295–16336\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2312.06942)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.9.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px2.p1.1)\. - IEC \(2010\)IECIEC 61508: functional safety of electrical/electronic/programmable electronic safety\-related systems, part 4 \(definitions\) and part 6 \(annex d, common cause failure and the beta\-factor method\)\.Note:International Electrotechnical Commission, GenevaCited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px5.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px6.p1.1)\. - International Ethics Standards Board for Accountants \(2024\)International Ethics Standards Board for AccountantsHandbook of the international code of ethics for professional accountants \(including international independence standards\), 2024 edition\.Note:International Federation of Accountants, New YorkExternal Links:[Link](https://www.ethicsboard.org/)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.2.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - ISO/IEC \(2022\)ISO/IECISO/IEC 22989:2022 information technology—artificial intelligence—artificial intelligence concepts and terminology\.Note:International Organization for Standardization, GenevaClause 3\.1\.1, “AI agent”Cited by:[Definition 2\.1](https://arxiv.org/html/2609.18272#S2.Thmtheorem1.p1.1)\. - ISO/IEC \(2023\)ISO/IECISO/IEC 42001:2023 information technology—artificial intelligence—management system\.Note:International Organization for Standardization, GenevaCited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px2.p1.1)\. - ISO/IEC \(2025\)ISO/IECISO/IEC 42006:2025 information technology—artificial intelligence—requirements for bodies providing audit and certification of artificial intelligence management systems\.Note:International Organization for Standardization, GenevaCited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.2.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px2.p1.1)\. - Järviniemi \(2025\)O\. JärviniemiSubversion via focal points: investigating collusion in LLM monitoring\.Note:arXiv:2507\.03010External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.03010)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.9.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Kapooret al\.\(2025\)S\. Kapoor, B\. Stroebl, Z\. S\. Siegel, N\. Nadgir, and A\. NarayananAI agents that matter\.Transactions on Machine Learning Research\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2407.01502)Cited by:[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Kish \(1965\)L\. KishSurvey sampling\.John Wiley & Sons,New York\.Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px4.p1.2)\. - Kleinberg and Raghavan \(2021\)J\. Kleinberg and M\. RaghavanAlgorithmic monoculture and social welfare\.Proceedings of the National Academy of Sciences118\(22\),pp\. e2018340118\.External Links:[Document](https://dx.doi.org/10.1073/pnas.2018340118)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[§2](https://arxiv.org/html/2609.18272#S2.p3.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.8.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1)\. - Knight and Leveson \(1986\)J\. C\. Knight and N\. G\. LevesonAn experimental evaluation of the assumption of independence in multiversion programming\.IEEE Transactions on Software EngineeringSE\-12\(1\),pp\. 96–109\.External Links:[Document](https://dx.doi.org/10.1109/TSE.1986.6312924)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.7.1.1.1)\. - Kohli \(2026\)G\. KohliNine judges, two effective votes: correlated errors undermine LLM evaluation panels\.Note:arXiv:2605\.29800External Links:[Document](https://dx.doi.org/10.48550/arXiv.2605.29800)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px2.p1.1),[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px3.p2.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.8.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px6.p1.1),[Figure 2](https://arxiv.org/html/2609.18272#S6.F2)\. - Kolt \(2025\)N\. KoltGoverning AI agents\.Notre Dame Law Review101\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.07913)Cited by:[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Lamet al\.\(2024\)K\. Lam, B\. Lange, B\. Blili\-Hamelin, J\. Davidovic, S\. Brown, and A\. HasanA framework for assurance audits of algorithmic systems\.InProceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency \(FAccT ’24\),New York, NY, USA,pp\. 1078–1092\.External Links:[Document](https://dx.doi.org/10.1145/3630106.3658957)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.3.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Laurieet al\.\(2013\)B\. Laurie, A\. Langley, and E\. KäsperCertificate transparency\.RFCTechnical Report6962,Internet Engineering Task Force\.External Links:[Document](https://dx.doi.org/10.17487/RFC6962)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px1.p1.1)\. - Littlewood and Miller \(1989\)B\. Littlewood and D\. R\. MillerConceptual modeling of coincident failures in multiversion software\.IEEE Transactions on Software Engineering15\(12\),pp\. 1596–1614\.External Links:[Document](https://dx.doi.org/10.1109/32.58771)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px1.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.7.1.1.1)\. - MITRE Corporation \(2025\)MITRE CorporationMITRE ATLAS: adversarial threat landscape for artificial\-intelligence systems, v5\.1\.0\.Note:[https://atlas\.mitre\.org](https://atlas.mitre.org/)Accessed 13 September 2026Cited by:[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Mökanderet al\.\(2024\)J\. Mökander, J\. Schuett, H\. R\. Kirk, and L\. FloridiAuditing large language models: a three\-layered approach\.AI and Ethics4\(4\),pp\. 1085–1115\.External Links:[Document](https://dx.doi.org/10.1007/s43681-023-00289-2)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.5.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Motwaniet al\.\(2024\)S\. R\. Motwani, M\. Baranchuk, M\. Strohmeier, V\. Bolina, P\. H\. S\. Torr, L\. Hammond, and C\. Schroeder de WittSecret collusion among AI agents: multi\-agent deception via steganography\.InAdvances in Neural Information Processing Systems 37 \(NeurIPS 2024\),pp\. 73439–73486\.External Links:[Document](https://dx.doi.org/10.52202/079017-2336)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.18272#S6.p1.1)\. - New York City Council \(2021\)New York City CouncilLocal law 144 of 2021: automated employment decision tools\.Note:Enforced by the NYC Department of Consumer and Worker Protection from 5 July 2023External Links:[Link](https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page)Cited by:[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px5.p1.1)\. - NVIDIA Corporation \(2026\)NVIDIA CorporationnvTrust: NVIDIA trusted computing solutions—confidential computing and GPU attestation\.Note:[https://github\.com/NVIDIA/nvtrust](https://github.com/NVIDIA/nvtrust)Accessed 13 September 2026Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px1.p1.1)\. - OWASP GenAI Security Project \(2025\)OWASP GenAI Security ProjectOWASP top 10 for agentic applications 2026\.Technical reportOWASP Foundation\.External Links:[Link](https://genai.owasp.org/)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Public Company Accounting Oversight Board \(2024\)Public Company Accounting Oversight BoardSpotlight: staff update on outreach activities related to the integration of generative artificial intelligence in audits and financial reporting\.Technical reportPCAOB,Washington, DC\.External Links:[Link](https://pcaobus.org/documents/generative-ai-spotlight.pdf)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p1.1),[§5](https://arxiv.org/html/2609.18272#S5.p1.1),[§8](https://arxiv.org/html/2609.18272#S8.SS0.SSS0.Px4.p1.1)\. - Rajiet al\.\(2020\)I\. D\. Raji, A\. Smart, R\. N\. White, M\. Mitchell, T\. Gebru, B\. Hutchinson, J\. Smith\-Loud, D\. Theron, and P\. BarnesClosing the AI accountability gap: defining an end\-to\-end framework for internal algorithmic auditing\.InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency \(FAT\* ’20\),New York, NY, USA,pp\. 33–44\.External Links:[Document](https://dx.doi.org/10.1145/3351095.3372873)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px3.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.3.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Rowstron \(2026\)A\. RowstronAgentic witnessing: pragmatic and scalable TEE\-enabled privacy\-preserving auditing\.Note:arXiv:2604\.24203External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.24203)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.10.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Schnablet al\.\(2025\)C\. Schnabl, D\. Hugenroth, B\. Marino, and A\. R\. BeresfordAttestable audits: verifiable AI safety benchmarks using trusted execution environments\.InWorkshop on Technical AI Governance \(TAIG\) at ICML 2025,External Links:[Document](https://dx.doi.org/10.48550/arXiv.2506.23706)Cited by:[§3](https://arxiv.org/html/2609.18272#S3.SS0.SSS0.Px2.p1.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.10.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1)\. - Shavitet al\.\(2023\)Y\. Shavit, S\. Agarwal, M\. Brundage, S\. Adler, C\. O’Keefe, R\. Campbell, T\. Lee, P\. Mishkin, T\. Eloundou, A\. Hickey, K\. Slama, L\. Ahmad, P\. McMillan, A\. Beutel, A\. Passos, and D\. G\. RobinsonPractices for governing agentic AI systems\.Technical reportOpenAI\.External Links:[Link](https://cdn.openai.com/papers/practices-for-governing-agentic-ai-systems.pdf)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p1.1),[Definition 2\.1](https://arxiv.org/html/2609.18272#S2.Thmtheorem1.p1.1),[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.2.2.1.1)\. - Steinet al\.\(2024\)M\. Stein, M\. Gandhi, T\. Kriecherbauer, A\. Oueslati, and R\. TragerPublic vs private bodies: who should run advanced AI evaluations and audits? A three\-step logic based on case studies of high\-risk industries\.InProceedings of the 2024 AAAI/ACM Conference on AI, Ethics, and Society \(AIES ’24\),pp\. 1401–1415\.Note:arXiv:2407\.20847Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.4.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\. - Sunet al\.\(2024\)H\. Sun, J\. Li, and H\. ZhangzkLLM: zero knowledge proofs for large language models\.InProceedings of the 2024 ACM SIGSAC Conference on Computer and Communications Security \(CCS ’24\),New York, NY, USA,pp\. 4405–4419\.External Links:[Document](https://dx.doi.org/10.1145/3658644.3670334)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.10.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px3.p1.1),[§9](https://arxiv.org/html/2609.18272#S9.p1.1)\. - Toupset al\.\(2023\)C\. Toups, R\. Bommasani, K\. Creel, S\. Bana, D\. Jurafsky, and P\. LiangEcosystem\-level analysis of deployed machine learning reveals homogeneous outcomes\.InAdvances in Neural Information Processing Systems 36 \(NeurIPS 2023\),pp\. 51178–51201\.External Links:[Document](https://dx.doi.org/10.52202/075280-2228)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.8.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1)\. - Trusset al\.\(2026\)S\. R\. A\. Truss, M\. C\. Ghanem, M\. Lacerda, H\. Kheddar, M\. Alshawki, T\. Kenaza, and T\. KalganovaAgentic and generative AI for intelligent autonomous vulnerability assessment and penetration testing: a systematic analysis\.Note:SSRNExternal Links:[Document](https://dx.doi.org/10.2139/ssrn.6804400)Cited by:[§5](https://arxiv.org/html/2609.18272#S5.SS0.SSS0.Px2.p1.1)\. - Vassilevet al\.\(2025\)A\. Vassilev, A\. Oprea, A\. Fordyce, H\. Anderson, X\. Davies, and M\. HaminAdversarial machine learning: a taxonomy and terminology of attacks and mitigations\.NIST Trustworthy and Responsible AI ReportTechnical ReportNIST AI 100\-2e2025,National Institute of Standards and Technology\.External Links:[Document](https://dx.doi.org/10.6028/NIST.AI.100-2e2025)Cited by:[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Wataokaet al\.\(2024\)K\. Wataoka, T\. Takahashi, and R\. RiSelf\-preference bias in LLM\-as\-a\-judge\.Note:arXiv:2410\.21819External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.21819)Cited by:[§1](https://arxiv.org/html/2609.18272#S1.p2.1),[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1),[§6](https://arxiv.org/html/2609.18272#S6.SS0.SSS0.Px2.p1.1)\. - Yaoet al\.\(2025\)S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhanτ\\tau\-bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.InThe Thirteenth International Conference on Learning Representations \(ICLR 2025\),External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.12045)Cited by:[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Zhanet al\.\(2024\)Q\. Zhan, Z\. Liang, Z\. Ying, and D\. KangInjecAgent: benchmarking indirect prompt injections in tool\-integrated large language model agents\.InFindings of the Association for Computational Linguistics: ACL 2024,Bangkok, Thailand,pp\. 10471–10506\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.findings-acl.624)Cited by:[§4](https://arxiv.org/html/2609.18272#S4.SS0.SSS0.Px2.p1.1),[Table 3](https://arxiv.org/html/2609.18272#S5.T3.2.6.2.1.1)\. - Zhou \(2026\)Z\. ZhouGoverning dynamic capabilities: cryptographic binding and reproducibility verification for AI agent tool use\.Note:arXiv:2603\.14332External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.14332)Cited by:[Table 1](https://arxiv.org/html/2609.18272#S3.T1.6.10.1.1.1),[§3](https://arxiv.org/html/2609.18272#S3.p1.1)\.
Similar Articles
AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents
AgentAudit is an open, extensible framework for evaluating the full lifecycle of AI agents across capability, grounding, security, and behavioral dimensions, enabling precise failure attribution and highlighting trustworthiness differences among various language models.
Anyone else struggling with AI auditability?
The author describes a challenge with AI auditability where an agent's decision lacked traceability to the active policy version, and asks for advice on building effective decision trails for AI agent decisions.
Beyond Component Testing: Validating Agentic AI Systems
This survey synthesizes 257 papers on validating agentic AI systems, proposing a five-dimension taxonomy covering behavioral, safety, temporal, regulatory, and multi-agent concerns. It identifies gaps in temporal validity, runtime evidence maintenance, regulatory legibility, and open-ended multi-agent assurance, arguing that trustworthy deployment requires validating trajectories in context.
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems
This paper proposes a governance model for autonomous AI agents based on institutional attestation, where actions are governed through independently attested evidence rather than monitoring agent reasoning. It formalizes this approach with a proof-of-concept implementation for high-risk actions like clinical prescribing and software deployment.
How are you handling audit and compliance for agentic systems in your org?
A practitioner details the real-world challenges of audit, compliance, and governance for autonomous AI agents, including identity, approvals, logging, and accountability, while asking the community for solutions.