Harmonizing AI Safety Thresholds

arXiv cs.AI Papers

Summary

This paper proposes a methodology for deriving harmonized AI safety thresholds across frontier AI companies to address inconsistencies in existing thresholds, covering misuse risks and automated AI R&D, and highlighting empirical gaps.

arXiv:2607.16112v1 Announce Type: new Abstract: Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:23 AM

# Harmonizing AI Safety Thresholds
Source: [https://arxiv.org/html/2607.16112](https://arxiv.org/html/2607.16112)
Wilber Sean Anterola1, Matthew Ball, Luis F\. Lafuerza, Markov Grey2 1Brown University2Centre pour la Sécurité de l’Intelligence Artificielle \(CeSIA\)

###### Abstract

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies\. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards\. We develop a methodology for deriving harmonized thresholds across three risk domains\. For misuse risks \(cyber and biological\), we take expected harm as the key primitive and use an explicit risk\-modeling approach that accounts for risk channels and model release conditions\. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm\. Our analysis expands upon prior work and highlights existing empirical gaps and limitations\.

## 1Introduction

Several frontier AI companies have developed safety frameworks that define thresholds for dangerous capabilities that trigger enhanced safety, security, or governance measures\(Anthropic,[2026d](https://arxiv.org/html/2607.16112#bib.bib7); Google DeepMind,[2025](https://arxiv.org/html/2607.16112#bib.bib16); OpenAI,[2025b](https://arxiv.org/html/2607.16112#bib.bib36)\)\. These frameworks are committed to under the Seoul Frontier AI Safety Commitments \(adopted by all the frontier AI companies\) and are increasingly required by emerging legislation, including the EU AI Act and California SB\-53\. The Seoul Commitments, in particular, call for AI companies to “*set out thresholds at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable*”\(UK Government,[2024](https://arxiv.org/html/2607.16112#bib.bib44)\)\.

However, the AI companies have produced thresholds in an ad hoc manner, with limited justification, leading to a fragmented threshold landscape\. These thresholds differ substantially in scope, specificity, and even in the risk categories they address111Several researchers have cataloged and compared these frameworks at the structural level\(Ziosi and others,[2025](https://arxiv.org/html/2607.16112#bib.bib48)\)\. For example, METR cataloged common elements across twelve published safety policies\(METR,[2025](https://arxiv.org/html/2607.16112#bib.bib26)\); they found that capability thresholds exist in 9 of 12 policies, conditions for halting deployment in 9 of 12, conditions for halting development in 8 of 12, evaluation frequency requirements in 9 of 12, and accountability mechanisms in all 12\. However, METR notes that “despite commonalities, each policy is unique and reflects a distinct approach to AI risk management\.”\. The manner in which these thresholds are operationalized also varies\. Companies often rely on internal proprietary tests, making it difficult for third parties to verify whether a threshold has been crossed\. Frontier companies also face a coordination problem: their own safety measures are less effective if other companies do not adopt comparable protections222The companies themselves acknowledge this\. Anthropic describes safety as “a collective action problem” in which “the overall level of catastrophic risk from AI depends on the actions of multiple AI developers, not just one”\(Anthropic,[2026c](https://arxiv.org/html/2607.16112#bib.bib4)\)\. Google DeepMind states that its framework would result in “effective risk mitigation for society only if all relevant organizations provide similar levels of protection”\(Google DeepMind,[2025](https://arxiv.org/html/2607.16112#bib.bib16)\)\. OpenAI conditions its risk assessment on the behavior of other actors through the concept of “marginal risk”\(OpenAI,[2025b](https://arxiv.org/html/2607.16112#bib.bib36)\)\.\. These differences can lead to inconsistent risk mitigation and a race to the bottom in safety\. The difficulty is not only that thresholds differ\. Because thresholds are often incomparable or unauditable, a company can justify a given deployment speed or access model while claiming to satisfy safety commitments similar to those of its competitors\. If one company adopts weaker or less verifiable trigger conditions, others face pressure to match it rather than maintain stricter safeguards\. This creates an urgent need for minimum thresholds that are consistent across companies and can be publicly evaluated; common floors reduce this pressure by making minimum trigger points comparable, independently assessable, and open to credible consequences once an enforcement mechanism exists\.

This paper develops ideas for definingcommon risk thresholds\. We use*harmonization*to mean a common minimum floor together with a shared measurement procedure\. This procedure lets thresholds across companies be compared and independently audited now, and enforced if an oversight body with audit powers and consequences is established, while still permitting any company to adopt stricter internal thresholds\. We focus on the three risk categories tracked by the main frontier companies: cyber risk, biological risk, and automated AI R&D\. We treat the first two categories differently from the third: cyber and biological risks involve specific misuse pathways, while automated AI R&D is not tied to one particular harm pathway but could exacerbate other frontier AI risks, such as power concentration, loss of control to AI systems, and other misalignment\-related risks\.

Prior work has cataloged and compared frontier AI companies’ frameworks at the structural level, identifying shared elements across published safety policies\(Ziosi and others,[2025](https://arxiv.org/html/2607.16112#bib.bib48); METR,[2025](https://arxiv.org/html/2607.16112#bib.bib26)\)\. Our paper builds on this work by developing a methodology for translating heterogeneous threshold language into shared, comparable forms, applying it across three domains, and assessing where harmonization is currently tractable, where it requires further analysis, and where it is already achievable\.

Our main contribution is this translation framework: a procedure for converting heterogeneous frontier\-company threshold language into auditable quantitative floors, together with the audit evidence each floor would require\. The three domains illustrate the range of outputs the framework produces\. Cyber demonstrates a tractable expected\-harm application, subject to substantial calibration uncertainty that we make explicit\. In biorisk the framework operates as a diagnostic: applying it identifies the specific missing evidence that currently prevents a defensible floor, rather than yielding a policy\-ready number\. For automated AI R&D we obtain an operational rate\-of\-progress floor that already covers the language of all three major companies\. The result is a method for comparison, audit, and coordination, not a complete set of final validated calibrations\.

The next section outlines the core methodology, including how to apply the framework\. Sections[3](https://arxiv.org/html/2607.16112#S3)and[4](https://arxiv.org/html/2607.16112#S4)apply the methodology to cyber and bio risk\. Section[5](https://arxiv.org/html/2607.16112#S5)details the methodology for automated AI R&D\. Section[6](https://arxiv.org/html/2607.16112#S6)concludes\.

## 2Core Methodology

This section outlines the core methodology for deriving harmonized AI risk thresholds\. The aim is to compare the formulations that frontier companies have put forward and propose either a quantitative limit that could encompass them or a procedure for deriving such a limit in a more principled way\. The derived threshold should support direct comparisons between companies, third\-party audits, and, eventually, enforcement\. We use different approaches for misuse risks and automated AI R&D\.

Misuse risksare those in which an AI system assists a human actor in causing harm, as in cyber offense or biological weapons development\. We argue that thresholds for these risks should take expected harm as the key primitive and use explicit risk modeling that accounts for risk channels and model release conditions\. This approach builds on existing work\.Koessleret al\.\([2024](https://arxiv.org/html/2607.16112#bib.bib24)\)recommend expected harm as the natural currency for misuse thresholds, and the quantitative risk\-modeling methodology ofMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\)provides a structured, six\-step path from scenario definition to expected\-harm estimates\. On the policy side, emerging legislation such as California SB\-53 and the New York RAISE Act require companies to mitigate catastrophic risks, defined as risks materially contributing to the death of, or serious injury to, more than 50 people or more than one billion dollars in damage\(California Legislature,[2025](https://arxiv.org/html/2607.16112#bib.bib10); New York State Legislature,[2025](https://arxiv.org/html/2607.16112#bib.bib11)\)\. We apply that methodology to two specific misuse risks: cyber and biorisk\. For each domain, we decompose the risk into a set of representative scenarios, build risk models that separate baseline harm \(without AI\) from AI\-enabled additional harm, identify key risk indicators \(such as benchmark scores\) that serve as proxies for AI capability, and map those indicators to risk model parameters\. The output is an expected\-harm estimate that depends on both model capability and deployment conditions, and that can be updated as better data and stronger evaluations become available\. For the cyber domain, we calibrate this model with publicly available incident data, cyber\-range evaluations, and actor\-pathway decompositions\. For the biological domain, we apply the same structure but find that the available evidence base is too thin to support a quantitative floor, a finding thatMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\)themselves anticipated\.

The next two subsections outline the technical approach for each case\.

### 2\.1Expected Harm Model

For misuse domains, we model expected harm for each attack pathwayjjas:

E​\[Hj\]=Nj×Psuccess,j×hjE\[H\_\{j\}\]=N\_\{j\}\\times P\_\{\\mathrm\{success\},j\}\\times h\_\{j\}\(1\)
whereNjN\_\{j\}is expected annual attack\-equivalent volume \(the number of independent attack attempts, or opportunities, per year for pathwayjj\),Psuccess,jP\_\{\\mathrm\{success\},j\}is the probability of a successful attack, andhjh\_\{j\}is the expected harm per successful event measured in a pre\-specified, domain\-specific harm unit\. Total AI\-enabled additional harm across all pathways is:

Δ​E​\[HAI\]=∑j\(E​\[Hj,AI\]−E​\[Hj,baseline\]\)\\Delta E\[H\_\{\\mathrm\{AI\}\}\]=\\sum\_\{j\}\\left\(E\[H\_\{j,\\mathrm\{AI\}\}\]\-E\[H\_\{j,\\mathrm\{baseline\}\}\]\\right\)\(2\)
When considering different release conditions \(rr: open weights, public API, trusted API, etc\.\), we make the following simplifying assumption:

E​\[HAI,r\]=sr×E​\[HAI\]E\[H\_\{\\text\{AI\},r\}\]=s\_\{r\}\\times E\[H\_\{\\text\{AI\}\}\]\(3\)
wheresrs\_\{r\}is the exposure scalar: the AI\-enabled harm under release conditionrrrelative to public\-API access, with public API set to11\. Gated conditions fall below11, while open weights can exceed it, since unrestricted access permits copying, fine\-tuning, removal of safety layers, and integration into autonomous tooling, all of which can amplify exposure above the public\-API baseline\. The scalar is a dimensionless relative\-exposure multiplier, not a fraction bounded at11\.

### 2\.2Automated AI R&D

For automated AI R&D, we do not model specific harm pathways\. We instead propose a procedure to identify substantial increases in the rate of AI progress\. The procedure has three steps: \(1\) quantify AI progress; \(2\) determine a baseline trend in the rate of AI progress; and \(3\) determine whether a model breaks the progress trend\. These steps are detailed in Section[5](https://arxiv.org/html/2607.16112#S5)\.

## 3Cyber Risk

This section translates existing company cyber threshold language into the common unit of expected harm\. The objective is not to replace capability threshold decisions, but to express what those decisions imply in terms of additional expected annual harm under different release conditions\.

FollowingKoessleret al\.\([2024](https://arxiv.org/html/2607.16112#bib.bib24)\), this calculation belongs to step \(2\) of the risk\-threshold methodology: using risk modeling to inform where capability thresholds should be set and to make audit requirements explicit\. It does not determine release decisions\. All parameter values, pathway allocations, success probabilities, attack volumes, and release\-condition scalars are author\-calibrated priors for illustration\. A mature version would replace them with IDEA\-elicited estimates, historical incident data, cyber\-range evidence, and audited safeguard\-performance measurements\.

### 3\.1Existing Threshold Language

Table[I](https://arxiv.org/html/2607.16112#S3.T1)reproduces the verbatim cyber threshold language from the three primary frontier AI companies\.

Table I:Verbatim cyber threshold language across primary AI companies\.Sources: Anthropic RSP v3\.0; GDM FSF v3\.0; OpenAI PF v2\.

Three observations follow\. First, Anthropic has no cyber threshold\. RSP v2\.2 deferred its determination; RSP v3\.0 dropped the category entirely\. Evaluation continues without a policy commitment to trigger safeguards\. Second, OpenAI High and GDM’s single threshold describe roughly the same harm: meaningful uplift enabling large\-scale attacks on defended targets, making harmonization at this level tractable\. Third, OpenAI Critical has no GDM or Anthropic equivalent; Critical\-level harmonization is not currently achievable\.

GDM’s threshold is outcome\-oriented: it activates on additional expected harm at severe scale\. OpenAI’s threshold is capability\-oriented: it activates on what the model can do\. This structural difference means the same model capability can simultaneously satisfy OpenAI High while not yet triggering GDM’s threshold\. Accordingly, the rest of this section treats harmonization as a translation problem: capability evidence must be mapped into expected annual harm under specified release conditions\. Table[II](https://arxiv.org/html/2607.16112#S3.T2)records this threshold\-type classification for each company\.

Table II:Threshold type classification, cybersecurity domain\.
### 3\.2Evidence Base and Calibration Choices

Given the threshold structure above, the cyber calibration requires evidence for two distinct questions: the size of the baseline harm and the model capability that could plausibly increase that harm\. The first concerns the global baseline for cyber harm, which should be anchored to aggregate cybercrime damage estimates rather than reported losses alone\. An early 2026 estimate places global cybercrime damages at approximately USD 500 billion per year, with a 90 percent confidence interval of USD 100 billion to USD 1 trillion\(Lukosiuteet al\.,[2026](https://arxiv.org/html/2607.16112#bib.bib25)\)\. That figure is itself a composite\.Lukosiuteet al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib25)\)survey 27 prior estimates and triangulate three independent sources: a UK business victimization survey and a US individual victimization survey, each scaled globally, together with global cybersecurity spending as a defense\-cost proxy\. Its harm construct covers direct losses, response costs, and defense spending, and it excludes harder\-to\-measure costs such as intellectual property theft and reputational damage\. This construct does not line up cleanly with the SB\-53 trigger used later as a benchmark: it is broader in one direction, counting defense spending, and narrower in another, since SB\-53 counts property damage from a single incident rather than an annual aggregate\. We therefore treat the USD 500 billion figure as an order\-of\-magnitude anchor, not as a like\-for\-like statutory quantity, and we do not read it as the unambiguously correct larger base\. By contrast, FBI IC3 reported USD 16\.6 billion in losses from 859,532 complaints in 2024, which is best interpreted as a reported\-loss lower bound rather than a global harm estimate\(Federal Bureau of Investigation Internet Crime Complaint Center,[2025](https://arxiv.org/html/2607.16112#bib.bib13)\)\. This distinction matters because a threshold calibrated only to reported losses would substantially understate the relevant harm base\.

The second concerns capability evidence from controlled cyber\-range evaluations\.Folkertset al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib15)\)evaluate frontier AI agents on a 32\-step corporate network attack range and report model progress on multi\-step cyber operations\. This range is useful because it measures end\-to\-end cyber capability in a repeatable environment\. It is not a direct estimate of operational cybercrime success, but it provides the best available public proxy for the type of multi\-step autonomous capability described in frontier AI cyber thresholds\. Obtaining more informative estimates of this kind is a priority for improving the framework\.

Table III:Source basis for the cyber calibration model\.Note\.The source basis is strongest for aggregate cyber harm and weakest for release\-condition effectiveness\.

Table[III](https://arxiv.org/html/2607.16112#S3.T3)motivates a two\-stage interpretation of the calculations that follow\. Aggregate harm values are externally anchored, while the capability proxy, pathway allocations, and release\-condition scalars remain calibration assumptions\. These assumptions should be updated as incident\-level evidence, independent audits, and more realistic cyber\-range evaluations become available\.

### 3\.3Risk Model

The base risk model follows Section[2](https://arxiv.org/html/2607.16112#S2)\. For the cyber domain, the equations expand as follows\. Per\-pathway expected harm is calculated as attack\-equivalent volume multiplied by success probability and harm per success\. Total AI\-enabled harm is the difference between AI\-assisted and baseline pathway harm, summed across pathways:

Δ​E​\[HAI\]=∑j\(E​\[Hj\]AI−E​\[Hj\]baseline\)\\Delta E\[H\_\{\\mathrm\{AI\}\}\]=\\sum\_\{j\}\\left\(E\[H\_\{j\}\]^\{\\mathrm\{AI\}\}\-E\[H\_\{j\}\]^\{\\mathrm\{baseline\}\}\\right\)\(4\)
Equation \([4](https://arxiv.org/html/2607.16112#S3.E4)\) identifies the pathway\-level difference between AI\-assisted and baseline harm but does not specify the source of the change\. AI\-enabled harm may arise through higher attack volume, higher success probability, greater harm conditional on success, or some combination of these\. Equation \([5](https://arxiv.org/html/2607.16112#S3.E5)\) makes this decomposition explicit, attributing harm changes to volume, success probability, or per\-event harm independently:

Δ​E​\[HAI,j\]=\\displaystyle\\Delta E\[H\_\{\\mathrm\{AI\},j\}\]=\{\}\(NAI,j×Psuccess,AI,j×hAI,j\)\\displaystyle\(N\_\{\\mathrm\{AI\},j\}\\times P\_\{\\mathrm\{success,AI\},j\}\\times h\_\{\\mathrm\{AI\},j\}\)\(5\)−\(N0,j×Psuccess,0,j×h0,j\)\\displaystyle\{\}\-\(N\_\{0,j\}\\times P\_\{\\mathrm\{success\},0,j\}\\times h\_\{0,j\}\)
whereN0,jN\_\{0,j\}is the baseline pathway volume,Psuccess,0,jP\_\{\\mathrm\{success\},0,j\}is the baseline success probability, andh0,jh\_\{0,j\}is the baseline harm per successful event\. The corresponding AI\-enabled values areNAI,jN\_\{\\mathrm\{AI\},j\},Psuccess,AI,jP\_\{\\mathrm\{success,AI\},j\}, andhAI,jh\_\{\\mathrm\{AI\},j\}\. The central calculations below use what we call a volume\-and\-success expansion\. AI adds attack\-equivalent volume, and on the AT1 through AT3 pathways it also raises the per\-attempt success probability, soPsuccess,AI,j\>Psuccess,0,jP\_\{\\mathrm\{success,AI\},j\}\>P\_\{\\mathrm\{success\},0,j\}there, while harm per successful event is held fixed\. Naming this second channel matters for the conservatism argument in Section[3\.11](https://arxiv.org/html/2607.16112#S3.SS11): the model is conservative on the volume dimension, but it already banks a success\-rate gain, so it is not conservative across the board\. Equation \([5](https://arxiv.org/html/2607.16112#S3.E5)\) is the more general form and can be used for sensitivity checks in which AI changes success probability or harm conditional on success\.

#### 3\.3\.1Sequential Attack\-Chain Structure

For pathways with sequential phases, let𝐏kill∈\[0,1\]K×J\\mathbf\{P\}\_\{\\mathrm\{kill\}\}\\in\[0,1\]^\{K\\times J\}be the kill\-chain phase matrix, where entrypk,jp\_\{k,j\}is the conditional probability of phasekksucceeding in pathwayjjgiven all prior phases succeeded\. The pathway success probability is then the product of the required phase\-level probabilities:

Psuccess,j=∏\{k:pk,j​defined\}pk,jP\_\{\\mathrm\{success\},j\}=\\prod\_\{\\\{k\\,:\\,p\_\{k,j\}\\,\\mathrm\{defined\}\\\}\}p\_\{k,j\}\(6\)
This is multiplicative in conditional probability, not linear in harm or capability\. It is called an AND\-gate model because all required phases must succeed\. This differs from OR\-gate structures, where any branch can suffice, and additive models, where phases contribute independently\. The product formula is the chain rule of probability, since eachpk,jp\_\{k,j\}is defined conditional on all prior phases succeeding; no Markov assumption is needed, and none is invoked\. The simplifying assumptions that the model does make lie elsewhere: per\-phase probabilities are treated as independent across pathways and are not conditioned on actor tier, and cross\-stage correlation is neglected\.

For initial access, an OR\-gate structure applies\. For example, a threat actor may succeed via phishing or vulnerability exploitation:

pIA=1−\(1−pphish\)​\(1−pvuln\)p\_\{\\mathrm\{IA\}\}=1\-\(1\-p\_\{\\mathrm\{phish\}\}\)\(1\-p\_\{\\mathrm\{vuln\}\}\)\(7\)
TLO evaluations use a fixed initial access path for comparability; \([7](https://arxiv.org/html/2607.16112#S3.E7)\) applies to the general threat model rather than to TLO\-measuredPsuccessP\_\{\\mathrm\{success\}\}directly\.

#### 3\.3\.2Generalized Harm Equation

The full expected harm under release conditionrruses the vector of pathway success probabilities:

E​\[HAI,r\]=𝐬r⊤​\(𝐪⊙𝐍0⊙𝐏success⊙𝐇\)E\[H\_\{\\mathrm\{AI\},r\}\]=\\mathbf\{s\}\_\{r\}^\{\\top\}\\left\(\\mathbf\{q\}\\odot\\mathbf\{N\}\_\{0\}\\odot\\mathbf\{P\}\_\{\\mathrm\{success\}\}\\odot\\mathbf\{H\}\\right\)\(8\)
where𝐬r\\mathbf\{s\}\_\{r\}is the release\-condition exposure vector,𝐪\\mathbf\{q\}is additional AI\-enabled attack\-equivalent volume relative to baseline,𝐍0\\mathbf\{N\}\_\{0\}is baseline opportunity volume,𝐏success\\mathbf\{P\}\_\{\\mathrm\{success\}\}is the vector of pathway success probabilities,𝐇\\mathbf\{H\}is harm per successful attack, and⊙\\odotdenotes element\-wise multiplication\.

#### 3\.3\.3Note on Actor Classification

This paper uses an Attacker Tier \(AT1–AT5\) taxonomy to classify cyber misuse actors, ranging from AT1 \(low\-skill, high\-volume attackers\) to AT5 \(nation\-state\-tier operators\)\. The taxonomy is structurally adapted from the RAND offensive cyber \(OC\) classification framework but differs in scope and application: the RAND OC classes were developed to describe actors targeting AI organizations to exfiltrate model weights, whereas the AT classes here describe actors using AI\-enabled tools against third\-party victims\. Resource profiles, capability assumptions, and pathway structures have been modified accordingly\. This paper citesBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\)extensively for empirical probability estimates; their original OC scenario labels \(e\.g\., OC3 SME Ransomware\) are reproduced verbatim in those attributions for traceability to their paper and should not be read as implying that our AT classes are equivalent to the RAND OC classes\.

The full calibration model for the cyber risk domain is presented in Appendix[B](https://arxiv.org/html/2607.16112#A2)\. This includes the baseline pathway allocation \(TableLABEL:tab:b1\-baseline\), AI\-enabled uplift estimates at Mythos Preview level \(TableLABEL:tab:b2\-uplift\), sensitivity of results to the baseline assumption \(TableLABEL:tab:b3\-sensitivity\), release\-condition exposure scalar matrix \(TableLABEL:tab:b4\-scalars\), and kill\-chain phase matrix \(TableLABEL:tab:b5\-killchain\)\. All parameter values in these tables are author\-calibrated priors, not independently verified estimates\. The structural conclusions in §§[3\.8](https://arxiv.org/html/2607.16112#S3.SS8)–[3\.10](https://arxiv.org/html/2607.16112#S3.SS10)are robust to reasonable variation in these priors, as shown in TableLABEL:tab:b3\-sensitivity\. A mature version of the model would replace them with IDEA\-elicited estimates, historical incident data, cyber\-range evidence, and audited safeguard\-performance measurements\.

### 3\.4Key Risk Indicators

Selecting an appropriate KRI for the cyber domain requires meeting three criteria identified byMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\): the indicator must be unsaturated at current capability levels, community\-validated through independent replication, and directly risk\-relevant\. Table[IV](https://arxiv.org/html/2607.16112#S3.T4)assesses the candidate evaluations against these criteria\.

Table IV:KRI candidate assessment against the three criteria inMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\): unsaturated, community\-validated, and risk\-relevant\.
### 3\.5Scenario Mapping and Calibration

Calibrating the AT3 pathway requires mapping theBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\)scenario library to the TLO scenario used inFolkertset al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib15)\)\. Table[V](https://arxiv.org/html/2607.16112#S3.T5)performs this mapping across nine scenarios and identifies which rows provide the most relevant evidence for the AT3 SME ransomware and domain\-compromise pathways\.

Table V:Mapping ofBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\)scenarios to the AT3 TLO scenario\.Rows 5 and 6 are the closest matches to the TLO scenario\. The human Delphi inBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\)was conducted on Scenario 5 \(OC3 SME Ransomware\), making it the only scenario with human\-validated uplift estimates\. Scenarios 5 and 6 differ from TLO in attack objective, ransomware versus domain takeover\. The AT3 column in TableLABEL:tab:b5\-killchaintakes its per\-phase values from the Barrett et al\. OC3 SME Ransomware profile and treats the resulting chain estimate as a placeholder for the domain\-takeover scenario, a limitation we return to in Section[3\.11](https://arxiv.org/html/2607.16112#S3.SS11)\.

### 3\.6Empirical Capability Evidence

Table[VI](https://arxiv.org/html/2607.16112#S3.T6)reports TLO step completion and the impliedPsuccessP\_\{\\mathrm\{success\}\}for successive frontier models, providing the capability evidence used to calibrate the AT3 pathway\.

Table VI:TLO step completion andPsuccessP\_\{\\mathrm\{success\}\}across model generations\.Primary TLO sources:Folkertset al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib15)\); AISI Mythos evaluation \(2026\); GPT\-5\.5 system card\(OpenAI,[2026](https://arxiv.org/html/2607.16112#bib.bib38), §9\.1\.2\.7, p\. 33\)\. The reportedPsuccessP\_\{\\mathrm\{success\}\}is full\-chain completions divided by trials \(for example, GPT\-5\.5 at 2/10 gives≈0\.20\\approx 0\.20and Mythos Preview at 6/10 gives≈0\.60\\approx 0\.60\)\. Mythos Preview was earlier reported as 3/10 in an April 2026 AISI report; 6/10 is from the more recent evaluation\. Both figures indicate non\-zero full\-chain completion at this capability level\.PsuccessP\_\{\\mathrm\{success\}\}scales log\-linearly with token budget, with up to 59 percent gains from 10M to 100M tokens and no plateau observed\(Folkertset al\.,[2026](https://arxiv.org/html/2607.16112#bib.bib15)\)\. Any threshold not specifying a compute budget is implicitly permissive at higher compute\.

### 3\.7Proposed Minimum Floor

We propose a graduated family of floors on the offensive\-cyber capability axis rather than a single tripwire\. The axis runs from assisted multi\-step operation, through autonomous vulnerability discovery, to autonomous end\-to\-end intrusion against hardened targets, and the family places one rung at each regime transition\. The lowest rung is non\-zero full\-chain TLO completion: a model triggers enhanced safeguards when its full\-chain completion rate is distinguishable from zero on TLO at a pinned 100M\-token budget and a fixed evaluation harness, measured across a pre\-registered number of trials\. To keep the boundary between zero and one completion from turning on a single noisy draw, the rule is stated on a lower confidence bound\. The rung fires when the Clopper–Pearson lower bound on the completion rate exceeds zero at the pre\-registerednn, and the rate is reported as an interval rather than a point\. An observed 2/10, for instance, carries a wide exact interval of roughly 0\.03 to 0\.56, which is why the trial count and harness must be fixed in advance and the estimate reported with its bound\.

The lowest rung is deliberately binary rather than a continuousτ\\tausuch asPsuccess\>0\.05P\_\{\\mathrm\{success\}\}\>0\.05\. No pre\-2026 model achieved any full TLO completion, so non\-zero completion marks a genuine capability\-regime transition; a binary boundary is harder to dispute than a specificτ\\tau, which requires defending the chosen value; and the first observed completions, GPT\-5\.5 at 2/10 and Mythos Preview at 6/10, are the points at which OpenAI and AISI independently identified a meaningful capability threshold\. The binary must not smuggle in compute\-permissiveness\. BecausePsuccessP\_\{\\mathrm\{success\}\}scales log\-linearly with token budget and shows no plateau, a model that reads as 0/10 at 100M tokens may complete the chain at a larger budget, so pinning the budget and harness is part of the floor rather than a footnote to it\.

The higher rungs are where the live policy question now sits, because autonomous offensive capability already exists at scale in the discovery regime\. Orchestrated agent pipelines have found large numbers of real vulnerabilities in open\-source software: TitanCA reports 203 confirmed zero\-day vulnerabilities and 118 CVEs across a monitored corpus of open\-source repositories\(Zhang and others,[2026](https://arxiv.org/html/2607.16112#bib.bib52)\)\. That clears the autonomous\-discovery rung for open\-source\-software targets, though not the end\-to\-end\-intrusion rung against hardened targets\. A graduated family exists to hold those two rungs apart\. The capability is also not one\-dimensional: exploit development and vulnerability discovery advance at different rates\. On exploit development Mythos Preview sits well above trend, roughly seven months ahead by one external analysis, while its advantage in finding vulnerabilities on a fixed budget is less clear\(Chauvinet al\.,[2026](https://arxiv.org/html/2607.16112#bib.bib53)\)\. The family therefore tracks the two axes separately, with TLO completion as the end\-to\-end rung, real\-CVE proof\-of\-concept generation as the discovery\-and\-exploitation rung \(CyberGym’s full 1,507\-task set, unsaturated at about 20 percent top success\(Wanget al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib51)\)\), and single\-phase benchmarks as lower\-resolution support\.

FollowingKoessleret al\.\([2024](https://arxiv.org/html/2607.16112#bib.bib24)\), this floor should be read as informing capability threshold setting rather than directly determining release decisions\. The expected\-harm model identifies the capability level and release\-condition exposure that would warrant enhanced scrutiny, but deployment decisions still require company\-level risk assessment, safeguard evaluation, and governance review\. The proposed floor is therefore a minimum tripwire for setting and auditing cyber capability thresholds, not a standalone release rule\. Its justification does not depend on the dollar calculation\. The lowest rung is warranted by the observed transition from zero to non\-zero full\-chain completion, and it would stand even if the expected\-harm figures were revised\. The harm model motivates why that transition matters and which release conditions deserve scrutiny; it is not a premise of the trigger\.

This interpretation has direct implications for auditability\. The proposed cyber floor is applicable to both trained and deployed systems because the relevant quantity depends jointly on model capability and release condition\. It is partially third\-party verifiable through standardized TLO\-style cyber\-range evaluations\. However, release\-condition claims require audit access to KYC effectiveness, monitoring bypass rates, jailbreak resistance, account\-compromise rates, model\-weight security, and indirect\-procurement risks\. The floor is therefore verifiable now but not yet enforceable: enforcement would require a designated oversight body capable of compelling such audits and imposing consequences when exposure\-reduction claims are unsupported\. No existing instrument fills that role for a harmonized cross\-company floor\. California SB\-53, for example, penalizes a developer only for failing to follow its own framework, not for missing a shared floor\. The conditional path to enforceability runs through machinery such as the EU AI Act’s code of practice or a successor to SB\-53\.

The calculation indicates the exposure reduction a release condition would need to achieve; it does not establish that any existing safeguard regime achieves that reduction\. Whether a real regime meets the target is an empirical audit question requiring KYC false\-accept data, monitoring bypass measurements, jailbreak\-resistance evaluations, model\-weight security attestation, account\-compromise controls, and review of indirect\-procurement risks\. Under the central priors, remaining below the statutory benchmark would require effective exposure to fall to approximately 1\.5% of public API access for a USD 1B benchmark and approximately 0\.74% for a USD 500M benchmark\. These figures should be treated as audit targets rather than as evidence that any existing trusted\-access regime is sufficient\. Both percentages scale linearly with the USD 500 billion anchor and depend on the uncited release\-condition scalars of TableLABEL:tab:b4\-scalars, so they are illustrative rather than precise\. The conclusion that survives across the full anchor range is the Public\-API binary, which clears the USD 1B reference from the low end of the confidence interval upward; the Trusted\-API verdict is calibration\-sensitive, as the limitations below make explicit\.

### 3\.8Interpretation for Harmonization

The cyber case illustrates three contributions of the expected\-harm approach to threshold harmonization\. First, it translates capability\-based thresholds into the same harm metric used by outcome\-oriented frameworks, enabling cross\-company comparison on a common scale\. Second, it shows why release conditions are constitutive rather than incidental: the same model can be unacceptable as open weights or public API while potentially acceptable under sufficiently strong trusted\-access or internal\-only controls\. Third, it identifies the decisive empirical question, not simply whether a model crosses a benchmark, but whether safeguards reduce effective access by the actor classes that drive expected harm, and specifies what must be measured to answer it\.

The numerical results should be read with appropriate caution: they constitute a transparent first\-pass calibration rather than final forecasts, showing that public access to a model with non\-zero end\-to\-end cyber intrusion capability could generate additional expected harm above commonly discussed catastrophic\-risk benchmarks, while internal\-only deployment remains much closer to the acceptable range under the central assumptions\. The comparison unit needs one caveat\. Our harm figures are annual expected\-harm aggregates summed across pathways, whereas the SB\-53 USD 1B figure is a single\-incident trigger\. We use USD 1B as an order\-of\-magnitude reference on an annual expected\-harm axis, the natural unit for theN×P×hN\\times P\\times hprimitive, and not as a claim that any single AI\-enabled incident reaches the statutory threshold\. Most of the annual total comes from many low\-value AT1–AT2 events that never approach single\-incident catastrophe, so crossing this reference points to aggregate exposure, not to crossing SB\-53 itself\.

### 3\.9Illustrative Case Study: April 2026

Two concurrent model releases make the operationalization gap directly observable\.

GPT\-5\.5 \(April 23, 2026\)\.OpenAI classified GPT\-5\.5 as High following its evaluation results\. The threshold activation produced a concrete institutional response: deployment safeguards were introduced, a Trusted Access program was launched, and a system card was published reporting evaluation results against pre\-specified criteria\. TLOPsuccess≈0\.20P\_\{\\mathrm\{success\}\}\\approx 0\.20\(2/10 completions; reported in system card §9\.1\.2\.7, p\. 33\)\. This is a functioning threshold commitment: a pre\-specified trigger condition, an evaluation result that clearly crosses it, and documented consequences\.

Mythos Preview \(April 7, 2026\)\.AISI independently evaluated Mythos Preview and reported 73 percent on expert\-level CTF tasks, 6/10 full TLO completions versus 0/10 for Opus 4\.6, autonomous discovery of CVE\-2026\-4747\(NIST National Vulnerability Database,[2026](https://arxiv.org/html/2607.16112#bib.bib32)\), and 5 novel exploits versus 2 for Opus 4\.6\. By the TLO metric, Mythos Preview substantially exceeds GPT\-5\.5 \(6/10 versus 2/10\), and by OpenAI’s High definition it satisfies the threshold\. No RSP protocol was activated because Anthropic has no cyber threshold\.

At Mythos Preview capability level \(Psuccess≈0\.60P\_\{\\mathrm\{success\}\}\\approx 0\.60on TLO at a 100M\-token budget\), the same underlying capability generated divergent institutional responses because one company had a formal cyber threshold and the other did not\. Had both companies been subject to the proposed floor \(non\-zero TLO completion at a 100M\-token budget\), both would have triggered enhanced safeguards simultaneously\.

### 3\.10Preliminary Conclusions

The cyber application demonstrates how the expected\-harm framework can be used to translate heterogeneous company threshold language into an auditable release\-condition decision\. The central result is not the precise dollar value of the illustrative calculation, but the structure it imposes: baseline harm is separated from AI\-enabled uplift, capability evidence is translated into pathway\-specific harm, and deployment options are evaluated by the degree to which they reduce effective exposure\.

Under the current calibration, as of June 2026, public API access and ordinary high\-safeguard access remain above the proposed acceptable\-harm range for the most advanced models, while internal\-only use is the only release condition that clearly falls below it\. Trusted API could be sufficient, but only if independent audits show that KYC, monitoring, refusal behavior, jailbreak resistance, account\-compromise controls, and model\-weight security reduce effective exposure to the required level\. The next empirical task is therefore not to defend the present point estimates, but to replace the scenario priors with historically calibrated pathway estimates and audited safeguard\-performance data\.

### 3\.11Limitations and Suggested Next Steps

Gross, not net harm\.The model estimates gross offensive AI\-enabled harm and sets defensive AI uplift to zero\. A model that can develop exploits autonomously can also improve vulnerability scanning, patching, detection, triage, and incident response, andLukosiuteet al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib25)\)note that defensive applications of AI may partially offset offensive gains\. The zero\-defense assumption is not sign\-neutral\. About 73 percent of the headline \(USD 49\.65B of USD 67\.90B\) comes from the AT1–AT2 commodity pathways: phishing, business email compromise, and commodity malware\. That is exactly where defender AI is most mature and deployed at hyperscaler scale, and where defenders see aggregate signal that individual attackers do not\. The conservatism is thus concentrated on the tiers that contribute least to elite\-capability concern and is weakest where the estimated harm is largest\. To bound the gap, applying defensive offsets of 0, 25, 50, and 75 percent to the AT1–AT2 component moves the headline to roughly USD 67\.9B, 55\.5B, 43\.1B, and 30\.7B\. Even full neutralization of the commodity pathways leaves the AT3–AT5 contribution of about USD 18\.25B, more than an order of magnitude above the USD 1B reference, so the Public\-API conclusion does not rest on the undefended commodity mass\. It would need re\-arguing only if defensive AI also blunted the AT3 elite\-intrusion pathway\. The claim that defender AI is mature on the commodity pathways is a domain judgement on our part, not a cited measurement\. The gross figure is therefore an upper bound whose slack sits mostly on AT1–AT2, and a parallel defensive\-uplift model is necessary work, not an optional extension\.

Volume assumption under autonomy\.The model treats additional AI\-enabled volume as a fraction of baseline\. Under full autonomous operation at machine speed, attack volume may become a function of AI capability rather than actor population, making the model conservative on the volume dimension\.

Barrett et al\. scenario mismatch\.The AT3 per\-phase probabilities in TableLABEL:tab:b5\-killchainare taken directly fromBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\), Table 12 \(OC3 SME Ransomware\), phase for phase\. This imports a ransomware bottleneck profile onto a domain\-takeover scenario whose decisive phases differ: Barrett et al\. find the largest AI uplift on Initial Access and Execution, whereas Folkerts et al\. identify Lateral Movement and Exfiltration as the TLO bottlenecks\. We therefore treat the resulting AT3 chain\-levelPsuccessP\_\{\\mathrm\{success\}\}as a placeholder pending domain\-takeover\-specific elicitation, rather than a calibrated estimate\. To show what rides on the borrowed phases, holding the other phases fixed and varying the two TLO bottlenecks \(lateral movement over 0\.50 to 0\.80, exfiltration over 0\.70 to 0\.95\) moves the AT3 column product from about 0\.05 to about 0\.12 around its central 0\.084\. Because this product is a structural illustration and does not enter the headline harm totals \(which read AT3 off the TLO evidence in Table[VI](https://arxiv.org/html/2607.16112#S3.T6)\), the borrowing affects the decomposition’s transparency, not the reported figure\.

Empirical basis of the levels\.Every level result, as distinct from the Public\-API binary, rests on inputs with limited empirical grounding\. The USD 500 billion anchor carries a roughly tenfold confidence interval, and all dollar figures scale linearly with it\. The ten\-pathway baseline in TableLABEL:tab:b1\-baselineis back\-fit to sum to that anchor, so the per\-pathwayN0N\_\{0\},P0P\_\{0\}, andhhare not independent priors, and the AT1–AT2 dominance is a modelling choice rather than an independent finding\. The release\-condition scalars in TableLABEL:tab:b4\-scalarsare uncited author priors; a real audit would need measured inputs for KYC false\-accept rates, presentation\-attack detection, jailbreak success, account compromise, and model\-weight security, which are scattered across standards bodies, certification labs, and security reporting and may not exist publicly for every cell\. We therefore confine strong claims to the Public\-API binary, which clears the USD 1B reference across the whole anchor interval\. The Trusted\-API result is calibration\-sensitive: below the reference at the central and low anchor, and within a factor of two of it near the top of the interval\. Sourcing the pathway anchors and the safeguard scalars is the central empirical task for a mature version, and is the subject of ongoing open\-source\-intelligence work\. The Murray et al\. methodology also prescribes interval propagation, carrying three\-point estimates through Monte Carlo simulation\. That step is likewise deferred: the figures reported here are point calibrations, not propagated distributions\.

Calibration priors\.The baseline allocation in TableLABEL:tab:b1\-baselineshould be replaced by pathway\-specific estimates from victimization surveys, breach datasets, ransomware telemetry, cyber insurance claims, and incident\-response reporting\. The AI\-enabled parameters in TableLABEL:tab:b2\-upliftshould be estimated through controlled cyber\-range evaluations, expert elicitation, and historical evidence on AI\-assisted abuse\. The release scalars in TableLABEL:tab:b4\-scalarsrequire independent audits of KYC evasion, monitoring bypass, jailbreak success, account compromise, insider misuse, and model\-weight theft\. Incident reporting required by legislation such as the EU AI Act and California SB 53 could become an invaluable resource for improving these estimates\.

Pathway dependence\.Credential theft can be an input into enterprise ransomware; supply\-chain compromise can enable downstream extortion; and critical\-infrastructure incidents may include both AT4 and AT5 stages\. The current tables treat pathways as additively separable for transparency\. A mature model should represent dependencies with an attack tree or causal graph, in which expected harm is aggregated over leaf\-node outcomes weighted by their joint probability under the dependency structure:

E​\[HAI,r\]=∑ℓPrr⁡\(leafℓ∣attack​tree​dependencies\)×HℓE\[H\_\{\\mathrm\{AI\},r\}\]=\\sum\_\{\\ell\}\\Pr\_\{r\}\\\!\\left\(\\mathrm\{leaf\}\_\{\\ell\}\\mid\\mathrm\{attack\\ tree\\ dependencies\}\\right\)\\times H\_\{\\ell\}\(9\)

## 4Biorisk

This section applies the risk\-modeling methodology to biological misuse risks\.

### 4\.1Existing Threshold Language

Table[VII](https://arxiv.org/html/2607.16112#S4.T7)gives the verbatim biorisk threshold language from each company’s current framework\.

Table VII:Verbatim biorisk threshold language across primary AI companies\.Anthropic’s RSP v3\.2 enumerates two chemical/biological capability thresholds: “Non\-novel chemical/biological weapons production” and “Novel chemical/biological weapons production\.” Unlike previous versions of the RSP, it does not pre\-specify the evaluations whose results would indicate that a threshold has been passed\. Instead, it states that a developer should make “a compelling argument that there is no significant increase in likelihood of individual users or small teams causing catastrophic harm\.” The determination of whether a particular model has crossed a threshold is therefore deferred to the system card rather than established in policy ahead of time\.

Anthropic states in the Claude Opus 4\.7 system card \(Apr\. 2026\) that it considers even the lower, non\-novel threshold \(which it designates CB\-1\) too under\-specified to judge with confidence whether a model passes it \(“it is hard to be confident regarding whether a model passes this threshold”\)\. For the upper, novel threshold \(CB\-2\), it states more explicitly that if one takes the language of the checkbox “at face value,” Opus 4\.7 and “many models already” provide “significant help” to actors working to produce novel chemical and biological weapons\. It also states that doing so involves taking “a very literal reading of the current language” that “does not map on to the safety risks that our RSP focuses on\.” It further signals that it “will likely revise” the wording of RSP to better align with its intent\. The thresholds themselves still exist, but Anthropic does not appear to take the checkbox wording literally and treats it as a poor proxy for the risk it is meant to represent\.

The language of the lower threshold overlaps more across companies\. Anthropic’s Non\-novel chemical/biological weapons production and OpenAI’s High both describe providing significant uplift to actors with some technical expertise \(Anthropic: “undergraduate STEM degrees”\) pursuing the development or acquisition of known chemical or biological weapons\. GDM does not explicitly mention expertise in its CBRN alert threshold definition, but its actor description of “low to medium resourced actors” implies a similar level of expertise\. Because all three labels converge on the same generalized meaning at the lower threshold, harmonization around this concept is more tractable in principle than in other areas\. That convergence is only textual, though, and Anthropic itself calls the CB\-1 wording a poor proxy it will likely revise\. The harmonization this paper leans on is therefore the measurement\-based floor developed in Step 6, since agreement on vague phrasing is a weaker footing than agreement on something measured\.

Harmonization of the upper threshold is more difficult\. As noted above, Anthropic’s CB\-2 threshold and OpenAI’s Critical both describe some form of expert\-level uplift or autonomous capability, while GDM does not publish an equivalent for its top threshold and OpenAI’s Critical threshold has not been operationalized in published evaluations\. In all cases, substantial information gaps remain and further methodological development is needed\.

### 4\.2Implementing the Quantitative Risk Methodology

Murrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\)propose a six\-step methodology: \(1\) define risk scenarios, \(2\) construct risk models, \(3\) quantify baseline risk, \(4\) identify key risk indicators, \(5\) estimate AI uplift, and \(6\) aggregate to overall risk estimates\.

The following subsections apply this methodology to biorisk\. Steps 1 and 2 are developed below\. Steps 3–6 are outlined through initial findings\.

#### 4\.2\.1Step 1: Defining the Risk Scenario

Murray et al\. decompose the risk space along three dimensions: actor, target, and vector\. For biorisk threshold harmonization, the scenario must be representative of the harm described by Anthropic’s*non\-novel chemical/biological weapons production*, OpenAI’s*High*, and GDM’s*CBRN uplift level 1*\.

Table[VIII](https://arxiv.org/html/2607.16112#S4.T8)maps the scenario space and identifies the relevant scenarios for the minimum floor\. The actor taxonomy is adapted from Gryphon Scientific and NTI biosecurity pathway models\. We keep the target constant as “civilian population” because mass\-casualty events from infectious biological agents are the usual framing in biorisk scenarios\.

Table VIII:Scenario selection: actor×\\timestarget×\\timesvector\.The BA1 scenario maps directly to Anthropic’s CB\-1 threshold \(“significantly help individuals or groups with basic technical backgrounds”\), OpenAI’s High \(“meaningful counterfactual assistance… to ‘novice’ actors”\), and GDM’s CBRN uplift level 1 \(“low to medium resourced actors”\)\. The BA3 scenario maps to CB\-2/Critical but has no GDM equivalent\.

For the decomposition framework, we adopt a biothreat kill chain analogous to the one used in the companion cyber section\. The kill chain is derived from OpenAI’s five\-stage biothreat creation process \(Ideation, Acquisition, Magnification, Formulation, Release\), cross\-referenced with the NTI biosecurity pathway model and the Gryphon Scientific threat characterization stages used in the GPT\-5 system card\(OpenAI,[2025a](https://arxiv.org/html/2607.16112#bib.bib37)\)\. This yields the six\-stage sequential model shown in Table[IX](https://arxiv.org/html/2607.16112#S4.T9):

Table IX:Biothreat kill chain\.A structural difference from cyber is that the biorisk kill chain includes phases where uplift is primarily informational \(K1, K3 in part\), followed by phases where barriers are primarily physical or tacit\-knowledge based \(K2, K4, K5\)\.Sandbrink \([2023](https://arxiv.org/html/2607.16112#bib.bib41)\)draws a related contrast between the risks raised by general\-purpose language models, which primarily reduce informational barriers, and those raised by biological design tools \(BDTs\), which affect the design of agents themselves\. This distinction is consequential for threshold design because the two classes of capability enter the kill chain at different phases: LLM\-style uplift is disproportionately found at K1 and the informational aspects of K3, while BDTs \(protein structure\-prediction and inverse\-folding tools like AlphaFold2 and ProteinMPNN\) apply more directly to enhancement at K3, and AI\-augmented laboratory automation\(Inagaki and others,[2023](https://arxiv.org/html/2607.16112#bib.bib22)\)applies to acquisition and formulation at K2 and K4\. Most existing efforts measure K1 and parts of K3 disproportionately \(VCT, ProtocolQA, and long\-form biothreat questions\), while K4 and K5, the phases with highest potential consequence, lack consensus AI\-specific evaluation paradigms\(National Academies of Sciences, Engineering, and Medicine,[2018](https://arxiv.org/html/2607.16112#bib.bib35)\)\. This remains a major methodological challenge for biorisk threshold operationalization\.

#### 4\.2\.2Step 2: Constructing the Risk Model

FollowingMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\), and as indicated in \([1](https://arxiv.org/html/2607.16112#S2.E1)\) above, risk is calculated as:

Risk=N×Psuccess×H\\mathrm\{Risk\}=N\\times P\_\{\\mathrm\{success\}\}\\times H\(10\)
whereNNis annual attack frequency,PsuccessP\_\{\\mathrm\{success\}\}is the probability of full kill\-chain completion, andHHis expected harm per successful attack, expressed here in casualties\. Casualties are the natural unit for bio, so instead of forcing the dollar conversion the cyber section uses, we bridgeHHto the casualty limb of the statutory trigger: SB\-53 is crossed by more than 50 deaths arising from a single incident, or, alternatively, by one billion dollars in property damage\(California Legislature,[2025](https://arxiv.org/html/2607.16112#bib.bib10)\)\. A dollar comparison stays available through the same value\-of\-statistical\-life bridge as in cyber\(U\.S\. Department of Transportation,[2026](https://arxiv.org/html/2607.16112#bib.bib45)\), but the casualty mapping avoids adding a contestable conversion parameter\. With this bridge the expected\-harm primitive is genuinely common across the two misuse domains, not merely asserted to be\.PsuccessP\_\{\\mathrm\{success\}\}is again the critical parameter for threshold design: it is outcome\-linked, in principle measurable via the kill\-chain decomposition, and the term most immediately affected by AI capability\. This section develops onlyPsuccessP\_\{\\mathrm\{success\}\}; it assigns no illustrative values toNNorHHfor bio, so \([10](https://arxiv.org/html/2607.16112#S4.E10)\) frames the decomposition rather than producing an expected\-casualty number\.

For a sequential kill chain with k stages \(AND\-gate structure\):

Psuccess=∏i=1kpiP\_\{\\mathrm\{success\}\}=\\prod\_\{i=1\}^\{k\}p\_\{i\}\(11\)
wherepip\_\{i\}is the per\-stage success probability conditional on all prior stages having succeeded\. Because eachpip\_\{i\}is defined on all prior stages succeeding, the product is the exact chain rule of probability rather than an approximation\. We also condition thepip\_\{i\}on the actor tier, so that cross\-stage competence correlation, the fact that a team clearing K3 is more likely to clear K4 through shared tacit laboratory skill, is absorbed into the conditionals rather than neglected\. AI uplift enters multiplicatively: if a model reduces the informational barrier at K1, it raisesp1p\_\{1\}and propagates through the product\.

Two stages are not single serial chokepoints but disjunctions over substitutable routes\. Acquisition \(K2\) can proceed by de novo synthesis, by drawing from a culture collection or a clinical or environmental sample, or by environmental isolation; delivery \(K5\) can proceed by aerosol, by contamination, or through a vector\. For such a stage the effective success probability is an OR over routes,

pstage=1−∏r\(1−proute,r\),p\_\{\\mathrm\{stage\}\}=1\-\\prod\_\{r\}\\left\(1\-p\_\{\\mathrm\{route\},r\}\\right\),\(12\)
which is at least as large as any single route\. Collapsing K2 or K5 into a single small multiplicand, as a naive AND\-only product does, therefore*understates*PsuccessP\_\{\\mathrm\{success\}\}: the single\-multiplicand product is a lower bound on realized risk to the extent that acquisition and delivery routes are substitutable\. We keep AND\-gates only at genuine chokepoints and OR\-gates at K2 and K5\.

In contrast to cyber, some stages involve physical barriers that are less affected by AI\. The AND\-gate structure means that even very large uplift at informational stages \(K1, K3\) may not substantially raisePsuccessP\_\{\\mathrm\{success\}\}while the physical bottleneck stages remain at very low success probabilities\. That reassurance holds only while K2, K4, and K5 stay physical or tacit, and, as discussed below, biological design tools and laboratory automation are already eroding those barriers\. It also rests on an assumption about structure rather than on a measurement: the Step\-4 indicators include no K5\-specific instrument, so the K4/K5 early\-warning signal has no direct read on delivery\.

However, emerging trends such as AI\-driven biological design tools, agentic lab automation, and automated DNA synthesis may turn some physical barriers and bottlenecks into informational ones that AI can more readily address\. This would materially change the overall risk profile\.Sandbrink \([2023](https://arxiv.org/html/2607.16112#bib.bib41)\)anticipates this trajectory, as doesNational Academies of Sciences, Engineering, and Medicine \([2018](https://arxiv.org/html/2607.16112#bib.bib35)\), who observe that automation in design paired with microfluidics allows “an actor to design and test agents at a smaller scale, less cost, and with less prior knowledge than more conventional pathways would require\.” InRighetti \([2025](https://arxiv.org/html/2607.16112#bib.bib39)\)’s framework of layered barriers, this development equates to AI reducing multiple barriers concurrently, rather than targeting a single chokepoint\. The kill\-chain product formula allows this possibility to be tracked quantitatively in a way that qualitative analysis would not capture: a persistent upward trend in per\-stage estimates of K4 and K5 would be a material signal for any harmonized threshold\.

The 2024–2025 evidence makes this trajectory concrete, and it bears on the specific stages the weakest\-link argument leans on\. At ideation and design \(K1, K3\), a survey of generative AI in the biosciences drawing on 130 expert interviews reports that the technology “lowers the barrier to misuse,” with roughly 76 percent of those experts expressing concern about AI misuse in biology\(Zhanget al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib28)\)\. For the tacit\-knowledge stages, CRISPR\-GPT automates gene\-editing design from selecting CRISPR systems and guide RNAs through delivery methods and protocol drafting, with the express aim of assisting non\-expert researchers\(Quet al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib29)\), and the Virtual Lab has run a full design\-to\-wet\-lab loop, using a pipeline of ESM, AlphaFold\-Multimer, and Rosetta to design 92 experimentally validated SARS\-CoV\-2 nanobodies\(Swansonet al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib31)\)\. At the execution end \(K4, K5\), systems such as LabOS couple AI agents to smart glasses and robots to assist real\-time laboratory work, moving AI “beyond computational design to participation”\(Cong and others,[2025](https://arxiv.org/html/2607.16112#bib.bib30)\)\. Most directly,Brent and McKelvey \([2025](https://arxiv.org/html/2607.16112#bib.bib33)\)challenge the tacit\-knowledge premise itself, reporting that current models can guide users through the recovery of live poliovirus from commercially obtained synthetic DNA\. None of these works demonstrates end\-to\-end weaponization; each shows capability or automation\. The claim they support is a narrow one: the K2, K4, and K5 barriers are binding now but eroding faster than a static AND\-gate implies, which makes the weakest\-link reassurance a monitored, time\-limited assumption rather than a standing fact\.

#### 4\.2\.3Step 3: Quantifying Baseline Risk

The baseline isPsuccessP\_\{\\mathrm\{success\}\}prior to generative AI availability\. The baseline problem differs from cyber in several ways\. Barrett et al\.’s expert elicitation for cyber was able to draw on diverse empirical sources, such as CVE databases, breach reports, and penetration testing against real targets, whileFolkertset al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib15)\)could build cyber ranges where models attempt multi\-step attacks and produce step\-completion data\.

In biorisk, the historical record of deliberate biological attacks is very limited, and existing cases are highly heterogeneous in actor profile, agent, and delivery method\. The base rate of successful sophisticated bioattacks is effectively zero in the modern era\. This presents specific problems:

Baseline calibration is fundamentally counterfactual\. We cannot ask “how often do attackers succeed at step X?” as in cyber; we must instead ask “how likely would a hypothetical BA1\-profile actor be to succeed at K4 \(formulation\) if they attempted it?” The difference from cyber is one of kind, not degree\. In cyber, all three terms of Risk=N×Psuccess×H=N\\times P\_\{\\mathrm\{success\}\}\\times Hare anchored to observed frequencies such as CVE counts, breach reports, and range completions, so a wrong estimate can be corrected against data\. In bio none of the three is observable:NNis effectively zero, full\-chainPsuccessP\_\{\\mathrm\{success\}\}has never been observed, andHHis a hypothetical casualty count\. The model cannot be back\-tested against outcomes that do not exist, so its per\-stage numbers are structured expert priors rather than empirical posteriors\. Elicitation can sharpen those priors and attach intervals to them, but nothing can validate them against attacks that have never happened\. What this section offers is a map of where the expected\-harm primitive stops transferring from cyber to bio; it is not a calibrated bio floor\.

Expert elicitation from biosecurity specialists is a central empirical source for biorisk baselines, and counterfactual elicitation is in fact the established bio method rather than a missing one: RAND, the NTI biosecurity pathway model, and Gryphon Scientific all run some form of it\. The closest published operational analogue to Barrett et al\.’s cyber design isMoutonet al\.\([2024](https://arxiv.org/html/2607.16112#bib.bib27)\), a RAND controlled red\-team with an internet\-only baseline group, which found no statistically significant difference in the viability of attack plans produced with versus without current\-generation LLM assistance\. What has not yet been published for bio is an elicitation that is at once community\-validated across companies, full\-chain, per\-stage, and calibrated to an explicit baseline\. RAND is baselined and operational but scores whole\-plan viability rather than per\-stage probabilities; Anthropic’s uplift trials are close to full\-chain but lab\-only and not independently replicable;Barrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\)is per\-stage and community\-facing but cyber\. No single bio exercise yet satisfies all four properties together\. The gap is therefore one of standardizing the method across companies, under test\-awareness, restricted information access, and small expert panels, not an absence of method\. It is an active research area being investigated by a FAR\.AI\-led EU AI Act CBRN consortium, in which SaferAI leads the risk\-modeling workstream\.

#### 4\.2\.4Step 4: Identifying Key Risk Indicators

Table[X](https://arxiv.org/html/2607.16112#S4.T10)presents an initial assessment of KRI candidates against the three criteria inMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\): unsaturated, community\-validated, and risk\-relevant\.

Table X:KRI candidates assessed against the three criteria inMurrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\): unsaturated, community\-validated, and risk\-relevant\.The central finding of the KRI assessment is that no single biorisk evaluation serves the function that the TLO cyber range serves in the companion section\. TLO is unsaturated, AISI\-validated, covers the full kill chain, and directly measuresPsuccessP\_\{\\mathrm\{success\}\}\. In biorisk, the closest equivalent is expert uplift trials \(currently only done by Anthropic\), but these are not standardized and cannot be independently replicated\. This is a significant methodological gap between the two domains\.

#### 4\.2\.5Step 5: Uplift Estimation

Initial uplift data from system cards, SecureBio evaluations, and the Epoch AI biorisk forecasting analysis are summarized in Table[XI](https://arxiv.org/html/2607.16112#S4.T11)\.

Table XI:Preliminary uplift estimates across model generations\.Note\.The threshold\-status column reports each company’s self\-assessment against its own threshold, not against the proposed harmonized floor in Step 6\.Sources: system cards\(Anthropic,[2025](https://arxiv.org/html/2607.16112#bib.bib3)\); SecureBio VCT; Epoch AI \(2025\);Gopalet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib17)\); RAND red\-team study\(Moutonet al\.,[2024](https://arxiv.org/html/2607.16112#bib.bib27)\)\.

Compared with cyber, whereFolkertset al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib15)\)provide evidence of log\-linear scaling ofPsuccessP\_\{\\mathrm\{success\}\}with model capability, biorisk uplift is more ambiguous, uneven, and stage\-specific\. At the ideation stage \(K1\), informational assessments like the long\-form Gryphon biothreat questions have saturated such that uplift at K1 is now large and near the ceiling\. At the enhancement stage \(K3\), tacit\-knowledge and design\-relevant evaluations have not saturated and are still increasing\. Frontier VCT scores are around 52%, substantially higher than the ~22% expert baseline but with considerable room remaining before saturation \(SecureBio\), and the sequence\-to\-function design assessment from the Mythos system card scored the model above the 75th percentile of human participants \(90th percentile on prediction\)\. Anthropic internally marked this as an early necessary\-but\-insufficient indicator of novel\-sequence design capability\. At the operational level \(full kill chain, K1–K5\), however, measured uplift remains low and notably non\-monotonic\. Anthropic’s expert uplift trials increased from 1\.82×\\times\(Opus 4\) to 1\.97×\\times\(Opus 4\.5\) on raw protocol scores, just short of the pre\-registered 2×\\timesthreshold, while Opus 4\.6 was subsequently judged to be slightly less helpful than Opus 4\.5 and made more critical errors\. These are small\-sample point estimates reported without confidence intervals, so both the apparent approach to 2×\\timesand the reversal at Opus 4\.6 may not be statistically distinguishable from noise\. The growing evaluation\-awareness of frontier models noted in recent system cards is a further confound: a model that recognizes an uplift trial may underperform it, so the Opus 4\.6 dip could be an elicitation artifact rather than a capability decline\. The series should not be read as a monotone approach to the 2×\\timessignal\.

The defining feature of any biorisk threshold operationalization is this gap between layers: large and saturated informational uplift \(K1\), substantial but not yet saturated design\-relevant uplift \(K3\), and low, uncertain, and non\-monotonic operational full\-chain uplift\. Full kill\-chainPsuccessP\_\{\\mathrm\{success\}\}is bounded by its most binding stage, currently the physical and tacit\-knowledge barriers at K2, K4, and K5, so large gains at the earlier informational stages may not meaningfully affect realized risk unless they are coupled with gains at later stages\(Sandbrink,[2023](https://arxiv.org/html/2607.16112#bib.bib41)\)\. The reassurance is weaker than a strict reading suggests\. Because K2 and K5 are OR\-gates over substitutable routes, the naive single\-multiplicand product understates realized risk wherever those routes exist, so the bottleneck protects less than an AND\-gate reading implies\. The barriers are also not static; whether, and how quickly, AI begins to lift the K2, K4, and K5 stages is the open question for threshold design, one the erosion evidence above speaks to directly\.

This ambiguity motivates expert elicitation methodologies\. SaferAI’s quantitative risk\-modeling framework\(Murrayet al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib34)\)blends \(i\) expert elicitation through Delphi studies, \(ii\) LLM\-based simulated experts, and \(iii\) Monte Carlo simulation to quantify uncertainty\. In biorisk, given both the relative lack of historical data and the hypothetical nature of most tested scenarios, Delphi\-style expert elicitation is the critical source of uplift estimates\. Such elicitation panels must draw on diverse forms of expertise, not only biosecurity experts who can translate benchmark performance to real\-world uplift, but also biological\-weapons experts who can realistically judge threat scenarios and biodefense practitioners who understand likely barrier performance at each stage\. Unlike cyber, where the expert panel can ground its judgments in decades of incident data such as CVE databases, breach reports, and CVSS scores \(cf\.Barrettet al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib9)\), there is no available analogue for biorisk because the base rate of attempted, sophisticated modern bioattacks is effectively zero\. In short, safety evaluators cannot consult a biosecurity attack database in the same way they can in cyber, so biorisk elicitation must rely more heavily on participants’ counterfactual reasoning about attacks that have not occurred\. This makes well\-structured disagreement elicitation and participant calibration training especially important\.

The Monte Carlo aggregation step can then be applied: per\-stage uplift estimates with associated uncertainty distributions can be composed through the kill\-chain product formula to generate overallPsuccessP\_\{\\mathrm\{success\}\}distributions\. This allows residual uncertainty to be made explicit and auditable, and provides the basis for percentile\-based trigger conditions in a preliminary minimum floor\.

#### 4\.2\.6Step 6: Risk Aggregation to a Threshold and a Preliminary Minimum Floor

Parameters can be combined into concrete risk estimates such as:

- •> “X% probability of \>Y successful bioweapon development attempts annually”
- •> “Z% increase in successful biological attack probability given AI access”

Under this quantitative risk\-modeling approach, a clear minimum floor analogous to the cyber TLO floor cannot yet be specified as an operational trigger\. Unlike cyber, where non\-zero full\-chain TLO completion is independently verifiable, no biorisk evaluation is simultaneously unsaturated, community\-validated, and full\-chain\. What can be stated now is a target specification for the floor, to be operationalized once the missing elicitation exists, rather than a floor that could be applied today:

The floor would fire when Monte Carlo aggregation of per\-stage uplift estimates, elicited with explicit uncertainty intervals, yields aPsuccessP\_\{\\mathrm\{success\}\}uplift ratio above a pre\-registered threshold \(for example Anthropic’s 2×\\timessignal\) at the 90th percentile of simulated scenarios\. That elicited uplift must be concentrated at the binding bottleneck stages K2, K4, and K5, not at the already\-saturated informational stages K1 and K3\. Expert\-level tacit\-knowledge performance on VCT and ProtocolQA is a precondition here, not a trigger: it is already met for VCT, so it does no triggering work on its own\. Requiring uplift “across multiple stages” without this restriction would let the floor fire on saturated K1 and K3 gains that leavePsuccessP\_\{\\mathrm\{success\}\}unchanged, exactly the uplift the weakest\-link argument says is not decisive, which is why the binding condition is stated at K2, K4, and K5\. The single limb that would actually bind, statistically significant elicited uplift at those bottleneck stages aggregated with uncertainty, is precisely the input that does not yet exist \(see Steps 4 and 5\); that is what makes this a target specification rather than a live floor\. This corresponds to the*lower*harmonized tier, consistent with Anthropic’s non\-novel chemical/biological weapons production threshold \(CB\-1\), OpenAI’s High capability language, and GDM’s CBRN uplift level 1\. The upper tier \(CB\-2 / Critical\) would require a separate, higher trigger\. The trend that would move a model toward this floor is a sustained rise in the per\-stage K2, K4, and K5 estimates, whose concrete drivers, biological design tool capability, agentic laboratory automation, and cloud\-lab or automated DNA\-synthesis access, are the signals a monitoring regime should watch\.

The model and the statute also differ in granularity\. SB\-53’s bio\-relevant limb trips at more than 50 deaths, whereas the scenarios modeled here sit at mass\-casualty scale, framed as damages far beyond COVID\-19 for the CB\-2 mapping\. The statutory trigger therefore fires at attacks orders of magnitude smaller than the mega\-scenarios the floor is pinned against, so a floor calibrated to the mass\-casualty tier under\-triggers relative to what the statute already prohibits\. We target that tier deliberately, since it is where the three companies’ threshold language is written and where the expected\-harm calculation does the most work, but the floor is better read as a conservative complement to the statutory line than as a restatement of it\.

This limited outcome is best read as a framework diagnosis rather than a failure of the method\. Applying the expected\-harm methodology to biorisk shows that the binding constraint is not threshold language, which already converges at the lower tier across the three companies\. It is the absence of a key risk indicator that is simultaneously unsaturated, community\-validated, and full\-chain, together with the absence of elicited stage\-level uplift estimates\. Locating that missing ingredient precisely, and specifying the elicitation and Monte Carlo aggregation steps that would supply it, is itself a substantive result: it tells threshold\-setters what to measure next rather than leaving the gap implicit\. Two structural assumptions behind the model remain contested expert judgments this paper does not resolve: whether the elicited per\-stage conditionals capture cross\-stage dependence, and whether K2 and K5 admit enough substitutable routes to behave as OR\-gates rather than serial chokepoints\. Both are plausibly actor\-tier dependent, and monovendor reasoning is weak evidence on either\. The bio floor should therefore be read as a structural sensitivity baseline, its structure stated and its sensitivity to those assumptions acknowledged, pending validation by a cross\-vendor panel\. Resolving it runs past any single paper: a standardized bio\-elicitation protocol built through the FAR\.AI and SaferAI consortium, per\-stage uplift elicitation to supply the missing bottleneck estimates, and cross\-vendor grounded debate to settle the contested AND\-gate and OR\-substitution judgments\.

### 4\.3Limitations and Suggested Next Steps

Address Current Evaluation Gaps\. Current biorisk evaluations face significant limitations in translating benchmark performance to real\-world risk\. The Murray et al\. methodology addresses this translation problem through its benchmark\-to\-risk\-parameter mapping\. The FAR\.AI\-led EU AI Act consortium’s research is seeking to close this gap\.

Focus on Expert vs\. Novice Uplift\. Many current evaluations focus on novice uplift rather than expert uplift, making it difficult to discern capability trajectories relevant to the upper threshold tier \(CB\-2 / Critical\)\.

Develop Biorisk\-Specific KRIs\. Extend beyond current benchmarks to capture the full threat chain from planning through deployment\.

Expert Network Development\. Build the domain expert network needed for credible elicitation studies covering biosecurity, biological weapons, and biodefense\.

## 5Automated AI R&D

Automated AI R&D is not a misuse risk\. However, it could greatly accelerate AI progress and heighten loss\-of\-control risk\. Faster capability progress would exacerbate all misuse risks, as existing defenses may struggle to keep up with threats created by newer models\. It might also make new models harder to understand, which would compound loss\-of\-control risk\. Loss of control could lead to harm unrelated to misuse and, in the limit, become catastrophic\.

A risk\-modeling analysis of possible harm pathways for automated AI R&D is not currently feasible, since this is a historically unprecedented phenomenon\. To propose a harmonized risk threshold that all AI companies could adopt, we take an approach different from the one used for misuse risks\. We examine the thresholds defined by the main frontier AI companies and extract the minimum common denominator, which they have implicitly accepted\. We then derive a quantitative threshold that captures this minimum common denominator and allows threshold crossings to be transparently evaluated\.

### 5\.1Existing Threshold Language

The three main companies have AI R&D thresholds related to acceleration in the speed of AI progress, summarized in Table[XII](https://arxiv.org/html/2607.16112#S5.T12)\.

Table XII:Verbatim automated AI R&D threshold language across primary AI companies\.Note\.Anthropic indicates that “double the rate of progress” means “as much progress in one year as one would see in two years at baseline\.” For example, if baseline progress involved a 3x scaleup in compute and a 3x improvement in algorithmic efficiency \(for a 9x “effective scaleup”\), “double the rate of progress” would entail something like an 81x effective scaleup\. This is not the same idea as “doubling researchers’ productivity,” since doubling inputs does not necessarily double the rate of progress\(Anthropic,[2026a](https://arxiv.org/html/2607.16112#bib.bib6)\)\(p\. 9\)\.

Appendix[A](https://arxiv.org/html/2607.16112#A1)shows that the Anthropic threshold indicated in Table[XII](https://arxiv.org/html/2607.16112#S5.T12)can be equivalent to OpenAI’s in terms of additional progress relative to trend, if we take “for several months” to mean “for three months”\. Since OpenAI’s proposal requires this additional progress to be obtained in less time, it is triggered by more advanced capabilities and can be used as the minimum*common denominator*333The threshold triggered by the more advanced capabilities is the less demanding one, and can therefore be used as the common denominator\.\. This*common denominator*is also compatible with GDM’s threshold since it leaves open the meaning of \[AI progress\] “substantially accelerating”\.

### 5\.2Harmonized Threshold Proposal

We therefore propose a threshold formulation that could accommodate the language of the three companies’ thresholds:

- •> A model causing AI progress to be at least five times faster than trend during at least three months\.

Making this threshold quantitative requires three components:

1. 1\.
2. 2\.> Determine a baseline trend in the rate of AI progress
3. 3\.> Determine whether a model breaks the progress trend

For step 1, we rely on benchmark scores\. We need a benchmark whose scores are available for a sufficient number of years \(so that we can compute a progress trend\) and that is not close to saturation \(so that it remains useful going forward\)\. Our main proposal is Epoch AI ECI\(Hoet al\.,[2025](https://arxiv.org/html/2607.16112#bib.bib20)\)\. This is a composite index based on several individual benchmarks and does not saturate as long as some of the benchmarks used are not saturated444An alternative could be METR time horizon, but, as of May 2026, it seems to be close to saturation and is becoming less reliable as the benchmark contains only a few tasks that frontier models cannot currently complete\.\. Because ECI incorporates individual benchmarks of varying difficulty, it can be informative for a large set of models, including frontier models over a longer time span\. The fact that ECI can be expanded to include new benchmarks also indicates that it can remain informative going forward555At least until some non\-saturated individual benchmarks exist\.\. While benchmark scores are imperfect measures of model capabilities, they currently appear to be the only viable option for producing a quantitative speed\-of\-progress score that can be objectively assessed\. Moreover, the benchmark does not need to capture the level of AI capabilities, but rather the rate at which those capabilities change, so benchmarks can be informative even if they miss some important dimensions of real\-world performance\. Some limitations of this choice are discussed in section[5\.4](https://arxiv.org/html/2607.16112#S5.SS4)\.

For step 2 \(determining a baseline trend in the rate of AI progress\), we fit a trend line to the scores of each company over time666Both ECI and METR time horizons show remarkably linear \(or exponential\) trends\.\. The slope of this trend is defined asr0r\_\{0\}\. An alternative would be to deriver0r\_\{0\}from the scores of frontier models777For this purpose, a model is considered frontier if its ECI is larger than that of any model released previously\., to capture the industry\-wide rate of progress\. The time frame for fitting the trend could be either 2018–2024, as indicated by Anthropic \(up to RSP v3\.0\), or 2024, as indicated by OpenAI\. A longer time frame would allow for a more precise estimation of the trend, but if capability progress has accelerated, using the 2024 trend will produce a somewhat higher trend\. The trend should nevertheless be defined with respect to a fixed time frame; otherwise, a continuous acceleration in the rate of progress would lead to an increasing trend slope, and a trend break may never be observed despite large cumulative increases in the rate of progress\.

For step 3 \(determining whether a model breaks the progress trend\), we compare the score of the latest frontier model888For this purpose, a model is considered*frontier*if its score in the considered benchmark is larger than the scores of all the models previously released by the company\.with that of its immediate frontier predecessor in a given company\. We define the rate of improvement of frontier modelnnof a given company as:

rn:=Cn−Cn−1tn−tn−1,r\_\{n\}:=\\frac\{C\_\{n\}\-C\_\{n\-1\}\}\{t\_\{n\}\-t\_\{n\-1\}\},\(13\)
whereCnC\_\{n\}999Equation \([13](https://arxiv.org/html/2607.16112#S5.E13)\) assumes a linear increase of the score over time\. If the score instead increases exponentially, one can take the logarithm of the score, or, equivalently, definernr\_\{n\}asrn:=ln⁡\(C​\(tn\)/C​\(tn−1\)\)/\(tn−tn−1\)r\_\{n\}:=\\ln\\\!\\left\(C\(t\_\{n\}\)/C\(t\_\{n\-1\}\)\\right\)/\(t\_\{n\}\-t\_\{n\-1\}\)\.is the capability score of modelnnandtnt\_\{n\}is its release date101010Ideally,tnt\_\{n\}should be the time the model became available internally, but since this date is not always public, the analysis may need to rely on the model or system card release date\. Companies could delay the release of a model to ensure that the threshold is not crossed\. If that occurs, external audits may be needed to observe capability levels in a timely manner\.\.

The threshold has been crossed if:

rn/r0≥5andtn−tn−1≥3​months\.r\_\{n\}/r\_\{0\}\\geq 5\\quad\\text\{and\}\\quad t\_\{n\}\-t\_\{n\-1\}\\geq 3\\text\{ months\}\.\(14\)
The principles of the threshold are illustrated in Figure[1](https://arxiv.org/html/2607.16112#S5.F1)\.

If frontier model release frequency is very high, we might find thattn−tn−1<3t\_\{n\}\-t\_\{n\-1\}<3months, so \([14](https://arxiv.org/html/2607.16112#S5.E14)\) cannot be satisfied regardless of any acceleration in the rate of improvement\. To allow for such cases, the threshold could be expressed as a function of the increase in capabilities relative to earlier frontier models:

rn,k:=Cn−Cn−ktn−tn−kr\_\{n,k\}:=\\frac\{C\_\{n\}\-C\_\{n\-k\}\}\{t\_\{n\}\-t\_\{n\-k\}\}\(15\)
The threshold has been crossed if:

rn,k/r0≥5andtn−tn−k≥3​monthsr\_\{n,k\}/r\_\{0\}\\geq 5\\quad\\text\{and\}\\quad t\_\{n\}\-t\_\{n\-k\}\\geq 3\\text\{ months\}\(16\)
If the rate of improvement is non\-decreasing, thenrn,k≥rn,k\+1r\_\{n,k\}\\geq r\_\{n,k\+1\}, and we only need to consider thern,kr\_\{n,k\}for the smallestkksuch thattn−tn−k≥3t\_\{n\}\-t\_\{n\-k\}\\geq 3months\.

Since we are considering improvement within a company \(caused by AI\),tnt\_\{n\}andCnC\_\{n\}should refer to the release time \(or ideally the training completion time\) and capability of frontier modelnnin a given company\.

![Refer to caption](https://arxiv.org/html/2607.16112v1/sections/AIR_DSchema.png)Fig\. 1:Illustration of the automated AI R&D threshold\.
### 5\.3Did Claude Mythos Cross the Threshold?

On April 7, 2026, Anthropic released the Claude Mythos Preview system card\(Anthropic,[2026b](https://arxiv.org/html/2607.16112#bib.bib5)\)\. Section 2\.3\.6 contains an analysis of ECI111111The analysis uses a selection of benchmarks different from those used by Epoch AI in their ECI, “so our \[Anthropic’s\] reported ECI scores are not directly comparable to public ECI scores”\(Anthropic,[2026b](https://arxiv.org/html/2607.16112#bib.bib5), p\. 42\)\.\. The score for Claude Mythos is 161121212The scores indicated in this section are extracted from the PDF version of the Claude Mythos system card, so they are subject to some imprecision\.\. Anthropic’s previous frontier model is Claude Opus 4\.6, with an ECI score of 152 and a release date of February 5, 2026\. Since this is less than three months earlier than Mythos, we consider the previous frontier model, Opus 4\.5, with a score of 145, released on November 24, 2025\.

Therefore,rMythos,2=\(161−145\)/134×365≅43\.58r\_\{\\mathrm\{Mythos\},2\}=\(161\-145\)/134\\times 365\\cong 43\.58points/year\. Forr0r\_\{0\}we use the slope between Claude 3 Opus and Claude 3\.7 Sonnet:r0=\(138−122\)/357×365≅16\.36r\_\{0\}=\(138\-122\)/357\\times 365\\cong 16\.36points/year\.

rMythos,2/r0≅2\.66<5r\_\{\\mathrm\{Mythos\},2\}/r\_\{0\}\\cong 2\.66<5\. Mythos does not cross the threshold\.

If we take Opus 4\.6 as the immediate reference, setting aside the three\-month window, we obtain:

rMythos,1=\(161−152\)/61×365≅53\.85r\_\{\\mathrm\{Mythos\},1\}=\(161\-152\)/61\\times 365\\cong 53\.85points/year, sorMythos,1/r0≅3\.29r\_\{\\mathrm\{Mythos\},1\}/r\_\{0\}\\cong 3\.29, still below the threshold\.

### 5\.4Limitations and Suggested Next Steps

This proposal matches the companies’ formulations if the observed progress acceleration is AI\-driven\. To validate this assumption, we could examine whether AI R&D\-related benchmarks are also improving at a substantially accelerated pace\. If data were available, we could also examine how inputs to AI R&D are changing; if, for example, there is a trend break in the amount of training or experimental compute used, the progress trend break could be attributed to compute\. On the other hand, an increase in the rate of AI progress could be concerning in itself and merit additional safety measures regardless of its origin; this is the case because other AI\-related risks \(such as CBRN, cyber, or job displacement\) would increase if the rate of AI progress increases\. On this basis, a rate\-based threshold would be the appropriate one even when attribution is uncertain\.

The threshold is backward\-looking\. The increase in capabilities of modelnnwas achieved before modelnnitself was available to accelerate AI R&D, so we can only expect that moving forward the rate of increase will be higher \(since modelnnis now available to assist with AI R&D\)\. This makes the threshold conservative, and therefore a floor that all companies should be willing to accept\.

ECI scores depend on the specific benchmarks and models used in its construction\. However,Hoet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib20)\)\(Appendix E\.3\) show that the slope in capability increase over time is rather robust to benchmark selection\. As models’ capabilities increase, new benchmarks will need to be added\. The particular choice of new benchmarks might affect the results, so robustness to new inclusions will also need to be assessed\.

Scores might also depend on the harness used and the amount of inference compute allowed\. If this primarily affects newer models, which are better able to make use of inference compute, a given assessment would be a lower bound for capabilities \(which could increase with better harnesses or more compute\), making a determination of a threshold breach conservative\. If, however, better harnesses or more compute also affect older models, this could affect the baseline capability\-increase trend, which would need to be reassessed over time\.

## 6Conclusions

In this paper we have presented a methodology to derive harmonized thresholds for dangerous AI capabilities\. For misuse\-related risks, we argue that thresholds should take expected harm as the key primitive and use an explicit risk\-modeling approach that accounts for risk channels and model release conditions\. For cyber risk, we have proposed a minimum threshold, but important quantitative uncertainties remain\. For biorisk, available data is more scarce and expert\-elicited per\-stage uplift estimates are a key missing input\. For automated AI R&D, we have proposed a simple and transparent threshold based on increases in the rate of AI progress, consistent with the thresholds already proposed by frontier AI companies\. We have also set out the main sources of uncertainty and directions for future work\. Appendix[C](https://arxiv.org/html/2607.16112#A3)draws the three domains together, summarizing their harmonization status and separating the inputs that are externally anchored from those that remain priors awaiting elicitation\. We hope this analysis proves useful for formulating common thresholds that meaningfully reduce AI risk\.

## Contributions

W\.S\.A\., M\.B\. and L\.F\.L\. were the lead authors of the Cyber Risk, Biorisk, and Automated AI R&D sections, respectively\. M\.G\. coordinated and supervised the whole project\. The initial idea of the project was proposed by Charbel\-Raphael Segerie\. The work was done within the Supervised Program for Alignment Research \(SPAR\)\.

## References

- Estimating the social cost of corporate data breaches\.External Links:2603\.21270,[Link](https://arxiv.org/abs/2603.21270)Cited by:[Table III](https://arxiv.org/html/2607.16112#S3.T3.1.6.4.3.1.1)\.
- Anthropic \(2025\)Claude opus 4 system card\.Technical reportAnthropic\.Cited by:[Table XI](https://arxiv.org/html/2607.16112#S4.T11.6.1.2)\.
- Anthropic \(2026a\)Biorisk\.Anthropic\.External Links:[Link](https://red.anthropic.com/2025/biorisk/)Cited by:[Table XII](https://arxiv.org/html/2607.16112#S5.T12.4)\.
- Anthropic \(2026b\)Claude mythos preview system card\.Technical reportAnthropic\.Cited by:[§5\.3](https://arxiv.org/html/2607.16112#S5.SS3.p1.1),[footnote 11](https://arxiv.org/html/2607.16112#footnote11)\.
- Anthropic \(2026c\)Responsible scaling policy, version 3\.0\.Anthropic\.External Links:[Link](https://www.anthropic.com/responsible-scaling-policy/rsp-v3-0)Cited by:[footnote 2](https://arxiv.org/html/2607.16112#footnote2)\.
- Anthropic \(2026d\)Responsible scaling policy, version 3\.1\.Anthropic\.External Links:[Link](https://www-cdn.anthropic.com/files/4zrzovbb/website/bf04581e4f329735fd90634f6a1962c13c0bd351.pdf)Cited by:[§1](https://arxiv.org/html/2607.16112#S1.p1.1)\.
- Anthropic \(2026e\)Responsible scaling policy, version 3\.2\.Anthropic\.External Links:[Link](https://cdn.sanity.io/files/4zrzovbb/website/28c6241900d90410628a8a2003a5572faae4365a.pdf)Cited by:[Table XII](https://arxiv.org/html/2607.16112#S5.T12.3.2.1.2.1.1)\.
- S\. Barrett, M\. Murray, O\. Quarks, M\. Smith, J\. Krys, S\. Campos, A\. Tlaie Boria, C\. Touzet, S\. Hayrapet, F\. Heiding, O\. Nevo, A\. Swanda, J\. Aguirre, A\. B\. Gershovich, E\. Clay, R\. Fetterman, M\. Fritz, M\. Juarez, V\. Mavroudis, and H\. Papadatos \(2025\)Toward quantitative modeling of cybersecurity risks due to AI misuse\.Note:arXivExternal Links:2512\.08864,[Link](https://arxiv.org/abs/2512.08864)Cited by:[§B\.5](https://arxiv.org/html/2607.16112#A2.SS5.p1.2),[Table XVII](https://arxiv.org/html/2607.16112#A2.T17.3.12.9.1.1.1),[§3\.11](https://arxiv.org/html/2607.16112#S3.SS11.p3.1),[§3\.3\.3](https://arxiv.org/html/2607.16112#S3.SS3.SSS3.p1.1),[§3\.5](https://arxiv.org/html/2607.16112#S3.SS5.p1.1),[§3\.5](https://arxiv.org/html/2607.16112#S3.SS5.p2.1),[Table III](https://arxiv.org/html/2607.16112#S3.T3.1.1.3.1.1),[Table V](https://arxiv.org/html/2607.16112#S3.T5),[§4\.2\.3](https://arxiv.org/html/2607.16112#S4.SS2.SSS3.p4.1),[§4\.2\.5](https://arxiv.org/html/2607.16112#S4.SS2.SSS5.p4.1)\.
- R\. Brent and T\. G\. McKelvey \(2025\)Contemporary AI foundation models increase biological weapons risk\.Note:arXivExternal Links:2506\.13798,[Link](https://arxiv.org/abs/2506.13798)Cited by:[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p12.1)\.
- California Legislature \(2025\)Senate Bill No\. 53: Transparency in Frontier Artificial Intelligence Act\.State of California\.Cited by:[§2](https://arxiv.org/html/2607.16112#S2.p2.1),[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p3.8)\.
- A\. Chauvin, J\. Barry, J\. Denain, and A\. Ho \(2026\)Are Mythos’ cyber capabilities overhyped?\.Note:Epoch AI, Gradient UpdatesExternal Links:[Link](https://epoch.ai/gradient-updates/are-mythos-cyber-capabilities-overhyped)Cited by:[§3\.7](https://arxiv.org/html/2607.16112#S3.SS7.p3.1)\.
- L\. Conget al\.\(2025\)LabOS: the AI\-XR co\-scientist that sees and works with humans\.Note:arXivExternal Links:2510\.14861,[Link](https://arxiv.org/abs/2510.14861)Cited by:[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p12.1)\.
- Federal Bureau of Investigation Internet Crime Complaint Center \(2025\)Internet crime report 2024\.Technical reportFederal Bureau of Investigation\.External Links:[Link](https://www.ic3.gov/AnnualReport/Reports/2024_IC3Report.pdf)Cited by:[§3\.2](https://arxiv.org/html/2607.16112#S3.SS2.p1.1)\.
- L\. Folkerts, W\. Payne, S\. Inman, P\. Giavridis, J\. Skinner, S\. Deverett, J\. Aung, E\. Zorer, M\. Schmatz, M\. Ghanem, J\. Wilkinson, A\. Steer, V\. Hong, and J\. Wang \(2026\)Measuring AI agents’ progress on multi\-step cyber attack scenarios\.Note:arXivExternal Links:2603\.11214,[Link](https://arxiv.org/abs/2603.11214)Cited by:[Table XVII](https://arxiv.org/html/2607.16112#A2.T17.3.12.9.1.1.1),[§3\.2](https://arxiv.org/html/2607.16112#S3.SS2.p2.1),[§3\.5](https://arxiv.org/html/2607.16112#S3.SS5.p1.1),[Table III](https://arxiv.org/html/2607.16112#S3.T3.1.5.3.3.1.1),[Table IV](https://arxiv.org/html/2607.16112#S3.T4.2.2.2.1.1),[Table VI](https://arxiv.org/html/2607.16112#S3.T6.14.4.4.4),[§4\.2\.3](https://arxiv.org/html/2607.16112#S4.SS2.SSS3.p1.1),[§4\.2\.5](https://arxiv.org/html/2607.16112#S4.SS2.SSS5.p2.6)\.
- Google DeepMind \(2025\)Frontier safety framework, version 3\.0\.Google DeepMind\.External Links:[Link](https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3.pdf)Cited by:[§1](https://arxiv.org/html/2607.16112#S1.p1.1),[Table XII](https://arxiv.org/html/2607.16112#S5.T12.3.3.2.2.1.1),[footnote 2](https://arxiv.org/html/2607.16112#footnote2)\.
- A\. Gopal, O\. Guest, T\. Besiroglu, and Epoch AI \(2025\)AI and biological risk: forecasting key capability thresholds\.Note:Epoch AI / EA ForumCited by:[Table XI](https://arxiv.org/html/2607.16112#S4.T11.6.1.2)\.
- A\. Ho, J\. Denain, D\. Atanasov, S\. Albanie, and R\. Shah \(2025\)A Rosetta Stone for AI benchmarks\.External Links:2512\.00193,[Link](https://arxiv.org/abs/2512.00193)Cited by:[§5\.2](https://arxiv.org/html/2607.16112#S5.SS2.p5.1),[§5\.4](https://arxiv.org/html/2607.16112#S5.SS4.p3.1)\.
- IBM Security \(2025\)Cost of a data breach report 2025\.Technical reportIBM\.External Links:[Link](https://www.ibm.com/reports/data-breach)Cited by:[Table III](https://arxiv.org/html/2607.16112#S3.T3.1.6.4.3.1.1)\.
- T\. Inagakiet al\.\(2023\)LLMs can generate robotic scripts from goal\-oriented instructions in biological laboratory automation\.Note:arXivExternal Links:2304\.10267,[Link](https://arxiv.org/abs/2304.10267)Cited by:[§4\.2\.1](https://arxiv.org/html/2607.16112#S4.SS2.SSS1.p5.1)\.
- L\. Koessler, J\. Schuett, and M\. Anderljung \(2024\)Risk thresholds for frontier AI\.Note:arXivExternal Links:2406\.14713,[Link](https://arxiv.org/abs/2406.14713)Cited by:[Table XVIII](https://arxiv.org/html/2607.16112#A3.T18.1.6.3.1.1.1),[§2](https://arxiv.org/html/2607.16112#S2.p2.1),[§3\.7](https://arxiv.org/html/2607.16112#S3.SS7.p4.1),[§3](https://arxiv.org/html/2607.16112#S3.p2.1)\.
- K\. Lukosiute, J\. Halstead, and L\. Righetti \(2026\)Global cybercrime damages: a baseline for frontier AI risk assessment\.Note:arXivExternal Links:2603\.20570,[Link](https://arxiv.org/abs/2603.20570)Cited by:[§3\.11](https://arxiv.org/html/2607.16112#S3.SS11.p1.1),[§3\.2](https://arxiv.org/html/2607.16112#S3.SS2.p1.1)\.
- METR \(2025\)Common elements of frontier AI safety policies \(december 2025 update\)\.Technical reportMETR\.External Links:[Link](https://metr.org/blog/2025-12-09-common-elements-of-frontier-ai-safety-policies)Cited by:[§1](https://arxiv.org/html/2607.16112#S1.p4.1),[footnote 1](https://arxiv.org/html/2607.16112#footnote1)\.
- C\. A\. Mouton, C\. Lucas, and E\. Guest \(2024\)The operational risks of AI in large\-scale biological attacks: results of a red\-team study\.Technical reportTechnical ReportRR\-A2977\-2,RAND Corporation\.External Links:[Link](https://www.rand.org/pubs/research_reports/RRA2977-2.html)Cited by:[§4\.2\.3](https://arxiv.org/html/2607.16112#S4.SS2.SSS3.p4.1),[Table XI](https://arxiv.org/html/2607.16112#S4.T11.3.5.1.5.1.1),[Table XI](https://arxiv.org/html/2607.16112#S4.T11.6.1.2)\.
- M\. Murray, S\. Barrett, H\. Papadatos, O\. Quarks, M\. Smith, A\. Tlaie Boria, C\. Touzet, and S\. Campos \(2025\)A methodology for quantitative AI risk modeling\.SaferAI\.Note:arXivExternal Links:2512\.08844,[Link](https://arxiv.org/abs/2512.08844)Cited by:[Table XVIII](https://arxiv.org/html/2607.16112#A3.T18.1.6.3.1.1.1),[§2](https://arxiv.org/html/2607.16112#S2.p2.1),[§3\.4](https://arxiv.org/html/2607.16112#S3.SS4.p1.1),[Table III](https://arxiv.org/html/2607.16112#S3.T3.1.1.3.1.1),[Table IV](https://arxiv.org/html/2607.16112#S3.T4),[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p1.1),[§4\.2\.4](https://arxiv.org/html/2607.16112#S4.SS2.SSS4.p1.1),[§4\.2\.5](https://arxiv.org/html/2607.16112#S4.SS2.SSS5.p4.1),[§4\.2](https://arxiv.org/html/2607.16112#S4.SS2.p1.1),[Table X](https://arxiv.org/html/2607.16112#S4.T10)\.
- National Academies of Sciences, Engineering, and Medicine \(2018\)Biodefense in the age of synthetic biology\.The National Academies Press\.Cited by:[§4\.2\.1](https://arxiv.org/html/2607.16112#S4.SS2.SSS1.p5.1),[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p11.1)\.
- New York State Legislature \(2025\)Responsible AI Safety and Education Act \(RAISE Act\)\.State of New York\.Cited by:[§2](https://arxiv.org/html/2607.16112#S2.p2.1)\.
- NIST National Vulnerability Database \(2026\)CVE\-2026\-4747\.Note:NVDFreeBSD RPCSEC\_GSS stack\-overflow remote code execution; published 2026\-03\-26External Links:[Link](https://nvd.nist.gov/vuln/detail/CVE-2026-4747)Cited by:[§3\.9](https://arxiv.org/html/2607.16112#S3.SS9.p3.1)\.
- OpenAI \(2025a\)GPT\-5 system card\.Technical reportOpenAI\.Cited by:[§4\.2\.1](https://arxiv.org/html/2607.16112#S4.SS2.SSS1.p4.1)\.
- OpenAI \(2025b\)Preparedness framework, version 2\.OpenAI\.External Links:[Link](https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf)Cited by:[§1](https://arxiv.org/html/2607.16112#S1.p1.1),[Table XII](https://arxiv.org/html/2607.16112#S5.T12.3.4.3.2.1.1),[footnote 2](https://arxiv.org/html/2607.16112#footnote2)\.
- OpenAI \(2026\)GPT\-5\.5 system card\.Technical reportOpenAI\.Cited by:[Table VI](https://arxiv.org/html/2607.16112#S3.T6.14.4.4.4)\.
- Y\. Qu, K\. Huang, M\. Yin,et al\.\(2025\)CRISPR\-GPT for agentic automation of gene\-editing experiments\.Nature Biomedical Engineering\.External Links:[Document](https://dx.doi.org/10.1038/s41551-025-01463-z),[Link](https://www.nature.com/articles/s41551-025-01463-z)Cited by:[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p12.1)\.
- L\. Righetti \(2025\)Dual\-use AI capabilities and the risk of bioterrorism: converting capability evaluations to risk assessments\.GovAI\.Cited by:[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p11.1)\.
- J\. Sandbrink \(2023\)Artificial intelligence and biological misuse: differentiating risks of language models and biological design tools\.Note:arXivExternal Links:2306\.13952,[Link](https://arxiv.org/abs/2306.13952)Cited by:[§4\.2\.1](https://arxiv.org/html/2607.16112#S4.SS2.SSS1.p5.1),[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p11.1),[§4\.2\.5](https://arxiv.org/html/2607.16112#S4.SS2.SSS5.p3.1)\.
- SecureBio \(2024\)Virology capabilities test \(vct\)\.Note:Published methodology and multi\-select variantCited by:[Table X](https://arxiv.org/html/2607.16112#S4.T10.2.2.3.1.1)\.
- K\. Swanson, W\. Wu, N\. L\. Bulaong,et al\.\(2025\)The virtual lab of AI agents designs new SARS\-CoV\-2 nanobodies\.Nature\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-09442-9),[Link](https://www.nature.com/articles/s41586-025-09442-9)Cited by:[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p12.1)\.
- U\.S\. Department of Transportation \(2026\)Departmental guidance on valuation of a statistical life in economic analysis\.U\.S\. Department of Transportation\.External Links:[Link](https://www.transportation.gov/office-policy/transportation-policy/revised-departmental-guidance-on-valuation-of-a-statistical-life-in-economic-analysis)Cited by:[Table III](https://arxiv.org/html/2607.16112#S3.T3.1.7.5.3.1.1),[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p3.8)\.
- UK Government \(2024\)Frontier AI safety commitments, AI seoul summit 2024\.Cited by:[§1](https://arxiv.org/html/2607.16112#S1.p1.1)\.
- Z\. Wang, D\. Song,et al\.\(2025\)CyberGym: evaluating AI agents’ real\-world cybersecurity capabilities at scale\.Note:arXivExternal Links:2506\.02548,[Link](https://arxiv.org/abs/2506.02548)Cited by:[§3\.7](https://arxiv.org/html/2607.16112#S3.SS7.p3.1),[Table IV](https://arxiv.org/html/2607.16112#S3.T4.4.4.2.1.1)\.
- A\. K\. Zhanget al\.\(2024\)Cybench: a framework for evaluating cybersecurity capabilities and risks of language models\.Note:arXivExternal Links:2408\.08926,[Link](https://arxiv.org/abs/2408.08926)Cited by:[Table IV](https://arxiv.org/html/2607.16112#S3.T4.1.1.2.1.1)\.
- Zhanget al\.\(2025\)BountyBench\.Cited by:[Table IV](https://arxiv.org/html/2607.16112#S3.T4.3.3.2.1.1)\.
- Zhanget al\.\(2026\)TitanCA: lessons from orchestrating LLM agents to discover 100\+ CVEs\.Note:arXivExternal Links:2604\.17860,[Link](https://arxiv.org/abs/2604.17860)Cited by:[§3\.7](https://arxiv.org/html/2607.16112#S3.SS7.p3.1)\.
- Z\. Zhang, S\. Chakraborty, A\. S\. Bedi,et al\.\(2025\)Generative AI for biosciences: emerging threats and roadmap to biosecurity\.Note:arXivExternal Links:2510\.15975,[Link](https://arxiv.org/abs/2510.15975)Cited by:[§4\.2\.2](https://arxiv.org/html/2607.16112#S4.SS2.SSS2.p12.1)\.
- Ziosiet al\.\(2025\)Safety frameworks and standards: a comparative analysis to advance risk management of frontier AI\.Note:Research MemoCited by:[§1](https://arxiv.org/html/2607.16112#S1.p4.1),[footnote 1](https://arxiv.org/html/2607.16112#footnote1)\.

## Appendix AThreshold Equivalence Derivation

This appendix provides the mathematical derivation supporting the threshold\-equivalence claim in Section[5](https://arxiv.org/html/2607.16112#S5)\. Under the assumption that “several months” in OpenAI’s formulation means three months, the Anthropic and OpenAI automated AI R&D thresholds are equivalent in terms of cumulative additional progress above trend\. The derivation is given under both an exponential growth model \([17](https://arxiv.org/html/2607.16112#A1.E17)\) and a linear growth model \([18](https://arxiv.org/html/2607.16112#A1.E18)\); the equivalence result holds under both\.

### A\.1Exponential Growth Model

Assume that AI progress increases exponentially at raterr:

Cr​\(t0\+Δ\)=C0​er​Δ,C\_\{r\}\(t\_\{0\}\+\\Delta\)=C\_\{0\}e^\{r\\Delta\},\(17\)
whereCr​\(t\)C\_\{r\}\(t\)is capability at timettwith rate of progressrr, andC0C\_\{0\}is capability at timet0t\_\{0\}\. OpenAI’s threshold corresponds torrbeing five timesr2024r\_\{2024\}\(the rate of progress in 2024\), since:

Cr​\(t0\+Δ\)=C0​e\(5​r2024\)​Δ=C0​er2024​\(5​Δ\),C\_\{r\}\(t\_\{0\}\+\\Delta\)=C\_\{0\}e^\{\(5r\_\{2024\}\)\\Delta\}=C\_\{0\}e^\{r\_\{2024\}\(5\\Delta\)\},

\(a rate5​r20245r\_\{2024\}during a timeΔ\\Deltaleads to the same progress as a rater2024r\_\{2024\}during a time5​Δ5\\Delta\), while Anthropic’s is just2​r2018​\-​20242r\_\{2018\\text\{\-\}2024\}\. However, Anthropic requires this to be maintained, on average, for 1 year, while OpenAI requires it for “several months”\. These two thresholds could be equivalent if we focus on additional progress with respect to baseline \(and assumer2024=r2018​\-​2024≡r0r\_\{2024\}=r\_\{2018\\text\{\-\}2024\}\\equiv r\_\{0\}\)\. Let additional progress be defined as:

Cr​\(t0\+Δ\)/Cr0​\(t0\+Δ\)=e\(r−r0\)​ΔC\_\{r\}\(t\_\{0\}\+\\Delta\)/C\_\{r\_\{0\}\}\(t\_\{0\}\+\\Delta\)\\ =\\ e^\{\(r\-r\_\{0\}\)\\Delta\},

where the equality follows from \([17](https://arxiv.org/html/2607.16112#A1.E17)\)\. For Anthropic, the additional progress after one year ise\(2​r0−r0\)​1=er0e^\{\(2r\_\{0\}\-r\_\{0\}\)1\}=e^\{r\_\{0\}\}, while for OpenAI, aftermmmonths it ise\(5​r0−r0\)​m/12=er0​m/3e^\{\(5r\_\{0\}\-r\_\{0\}\)m/12\}=e^\{r\_\{0\}m/3\}\. Equating the two shows that the rate proposed by OpenAI, maintained for 3 months, yields the same additional progress as the rate proposed by Anthropic, maintained for 1 year\.

### A\.2Linear Growth Model

If, instead, we assume linear increase at raterr:

Cr​\(t0\+Δ\)=C0\+r​Δ,C\_\{r\}\(t\_\{0\}\+\\Delta\)=C\_\{0\}\+r\\Delta,\(18\)
OpenAI’s threshold corresponds torrbeing multiplied by 5, while Anthropic’s corresponds torrbeing multiplied by 2\. The additional advance for Anthropic’s threshold is\(2​r2018​\-​2024−r2018​\-​2024\)⋅1=r2018​\-​2024\(2r\_\{2018\\text\{\-\}2024\}\-r\_\{2018\\text\{\-\}2024\}\)\\cdot 1=r\_\{2018\\text\{\-\}2024\}, while for OpenAI it is\(5​r2024−r2024\)⋅m/12=4​r2024​m/12\(5r\_\{2024\}\-r\_\{2024\}\)\\cdot m/12=4r\_\{2024\}\\,m/12\. Again, assumingr2024=r2018​\-​2024r\_\{2024\}=r\_\{2018\\text\{\-\}2024\}, the additional advance is equal ifm=3m=3\.

Under both growth models, the Anthropic and OpenAI thresholds imply equal cumulative additional capability gain whenm=3m=3months\. Because OpenAI’s formulation requires the elevated rate to be sustained for “several months” \(interpreted here as three months\) and because it is triggered by more advanced capability within a shorter window, it serves as the minimum common denominator adopted in the harmonized formulation proposed in Section[5](https://arxiv.org/html/2607.16112#S5)\. The GDM threshold is compatible with this common denominator since it leaves open the precise meaning of AI progress “substantially accelerating\.”

## Appendix BCyber Risk Calibration Model

This appendix contains the full quantitative calibration for the cyber risk model in Section[3](https://arxiv.org/html/2607.16112#S3)\. All numerical values are author\-calibrated priors intended to illustrate the methodology\. They are not derived through the IDEA expert\-elicitation protocol that a mature application would require and should not be used as inputs to release decisions without independent validation\.

### B\.1Baseline Cyber Harm Across AT1–AT5 Pathways

TableLABEL:tab:b1\-baselinedecomposes the USD 500 billion global cybercrime baseline into ten illustrative pathways across AT1–AT5\. AT1–AT5 denotes attacker\-tier sophistication, from low\-skill, high\-volume abuse to strategic national\-scale compromise\. The values are calibrated pathway parameters rather than independently verified raw event counts\. Their purpose is to expose the model structure and identify where historical evidence is most needed\.

Table XIII:Baseline cyber harm allocation across AT1–AT5 pathways\.A​TATPathwayN0N\_\{0\}P0P\_\{0\}H per successBaseline harmAT1Credential phishing / account takeover80M0\.10USD 10KUSD 80BAT1Commodity malware / low\-skill intrusion20M0\.10USD 20KUSD 40BAT2BEC / social engineering2M0\.17USD 250KUSD 85BAT2Affiliate ransomware against SMEs500K0\.20USD 550KUSD 55BAT3Enterprise ransomware / domain compromise50K0\.50USD 4MUSD 100BAT3Enterprise data theft / extortion25K0\.50USD 4MUSD 50BAT4Zero\-day chains against hardened organizations2\.5K0\.20USD 50MUSD 25BAT4Supply\-chain or SaaS compromise7000\.10USD 500MUSD 35BAT5Critical infrastructure disruption600\.10USD 3BUSD 18BAT5Strategic national\-scale compromise120\.10USD 10BUSD 12BTotalUSD 500BNote\.N0N\_\{0\}andP0P\_\{0\}are calibrated baseline parameters\. The table distributes the aggregate baseline across actor/pathway categories for transparent sensitivity analysis\. Pathway allocations should be replaced by victimization survey data, breach telemetry, cyber insurance claims, and ransomware payment records\.From TableLABEL:tab:b1\-baselinewe recover the empirically estimated expected harm:

E​\[Hbaseline\]\\displaystyle E\[H\_\{\\mathrm\{baseline\}\}\]=80​B\+40​B\+85​B\+55​B\+100​B\\displaystyle=0\\mathrm\{B\}\+0\\mathrm\{B\}\+5\\mathrm\{B\}\+5\\mathrm\{B\}\+00\\mathrm\{B\}\(19\)\+50​B\+25​B\+35​B\+18​B\+12​B\\displaystyle\\quad\+0\\mathrm\{B\}\+5\\mathrm\{B\}\+5\\mathrm\{B\}\+8\\mathrm\{B\}\+2\\mathrm\{B\}=500​B​USD/year\\displaystyle=00\\ \\mathrm\{B\\ USD/year\}
This allocation is illustrative and back\-fit to the USD 500 billion anchor: the ten rows are constructed to sum to that total, so the per\-pathwayN0N\_\{0\},P0P\_\{0\}, andhhare not independent priors, and changing one row requires another to absorb the difference if the anchor is held fixed\. The AT1–AT2 dominance is therefore a modelling choice rather than an independent finding\. A mature version should replace the whole partition with independent per\-pathway anchors, drawn from victimization surveys, breach telemetry, ransomware\-payment records, and cyber\-insurance claims, so that each row carries its own evidence and can be rejected on its own terms\.

### B\.2Estimating AI\-Enabled Uplift

The following calculation illustrates the methodology for a model at Mythos Preview capability level\. Mythos Preview is the reference model because it is the most capable publicly evaluated model without an active Anthropic policy threshold, making the release\-condition question directly actionable: which release conditions, if any, bring a model at this capability level below the USD 1B statutory benchmark?

Throughout the uplift table,PAI,jP\_\{\\mathrm\{AI\},j\}denotes the success rate of the marginal, AI\-enabled attempts on pathwayjj, not the population success rate across all attackers on it\. That distinction keeps the small AT4–AT5 entries coherent: AI draws new, lower\-skill actors toward elite intrusion and most of them fail, so the marginal success rate on AT4–AT5 sits below the incumbent baselineP0P\_\{0\}, even as AI raises the success rate of the lower tiers it genuinely assists\.

The primary TLO figure used here is 6/10 full\-chain completions, drawn from the most recent AISI evaluation of Mythos Preview\. An earlier evaluation reported 3/10; the 6/10 figure is used as the more recent estimate\. The proposed minimum floor, non\-zero full\-chain completion, is unaffected by which figure is used\. For AT4–AT5,PAIP\_\{\\mathrm\{AI\}\}values are kept small but non\-zero in the central calibration because Mythos\-level capability is not assumed to materially uplift nation\-state\-tier operations requiring rare target access, long\-term stealth, and strategic intent\. The combined AT4–AT5 contribution is approximately USD 0\.25B of the USD 67\.90B total, so setting these values to zero would not change the headline result; the small non\-zero values are retained for transparency\.

A caveat on the volume assumption: this model treats additional AI\-enabled volume as a fraction of baseline volume, represented byqq\. Under full autonomous operation, machine\-speed attack pipelines rather than human\-speed operations could allow a single actor to launch orders of magnitude more attempts than the human\-population\-constrained baseline implies\. Attack volume may become a function of AI capability rather than actor population\. The harm estimates here are therefore conservative with respect to volume\.

Table XIV:AI\-enabled uplift at Mythos Preview level, public API baseline\.A​TATPathwayqqPAIP\_\{\\mathrm\{AI\}\}\(Mythos\)AI\-enabled harmAT1Credential phishing / account takeover0\.200\.12USD 19\.20BAT1Commodity malware / low\-skill intrusion0\.150\.12USD 7\.20BAT2BEC / social engineering0\.150\.20USD 15\.00BAT2Affiliate ransomware against SMEs0\.100\.30USD 8\.25BAT3Enterprise ransomware / domain compromise0\.100\.60 \[TLO: 6/10\]USD 12\.00BAT3Enterprise data theft / extortion0\.100\.60 \[TLO: 6/10\]USD 6\.00BAT4Zero\-day chains against hardened orgs0\.050\.01 \[zero\-sensitivity recommended\]USD 0\.06BAT4Supply\-chain or SaaS compromise0\.050\.01USD 0\.18BAT5Critical infrastructure disruption0\.020\.002USD 0\.01BAT5Strategic national\-scale compromise0\.020\.001USD 0\.00BTotalUSD 67\.90BNote\.Each cell isq×N0×PAI×Hq\\times N\_\{0\}\\times P\_\{\\mathrm\{AI\}\}\\times H, withN0N\_\{0\}andHHtaken from TableLABEL:tab:b1\-baseline\. On AT1–AT3 the modelledPAIP\_\{\\mathrm\{AI\}\}exceeds the baselineP0P\_\{0\}, so the uplift combines additional volume with a success\-rate gain rather than expanding volume alone\. The headline figure is dominated by AT1–AT2 high\-volume pathways \(USD 49\.65B\), not by AT3 elite\-intrusion capability \(USD 18\.00B\)\. At GPT\-5\.5 level \(2/10 TLO, AT3PAI=0\.20P\_\{\\mathrm\{AI\}\}=0\.20\), the total is approximately USD 55\.90B\. At pre\-frontier baseline \(0/10 TLO, no AT3 contribution\), it is approximately USD 49\.90B\. The calibration should not be read as a claim that any nonzero AT3 capability alone generates the headline figure\.Table XV:Sensitivity of additional AI\-enabled harm to baseline assumption\.Release conditionH0=USD​100​BH\_\{0\}=\\text\{USD\}\\ 100\\text\{B\}H0=USD​500​BH\_\{0\}=\\text\{USD\}\\ 500\\text\{B\}H0=USD​1,000​BH\_\{0\}=\\text\{USD\}\\ 1\{,\}000\\text\{B\}Above USD 1B annual\-harm reference?Open weightsUSD 27\.2BUSD 135\.8BUSD 271\.6BYes \(all\)Public APIUSD 13\.6BUSD 67\.9BUSD 135\.8BYes \(all\)Safety\-filtered APIUSD 7\.9BUSD 39\.3BUSD 78\.7BYes \(all\)Trusted APIUSD 0\.06BUSD 0\.32BUSD 0\.63BNo \(all\)Internal only<<USD 0\.01B<<USD 0\.01B<<USD 0\.01BNoNote\.Each row is the pathway\-weighted aggregation of the release\-condition scalars in TableLABEL:tab:b4\-scalarsagainst the per\-tier AI\-enabled harm in TableLABEL:tab:b2\-uplift, so this table is the direct roll\-up of that matrix rather than a flat per\-condition discount\. Values scale linearly withH0H\_\{0\}\. The Public API conclusion holds across the whole anchor range\. The Trusted API figure now sits below the USD 1B annual\-harm reference across the entire range, but it reaches USD 0\.63B at the top of the anchor interval, within a factor of two of the reference, and it rests on the uncited exposure scalars of TableLABEL:tab:b4\-scalars; it should be read as an audit target, not a settled verdict\.
### B\.3Release Conditions as Exposure Scalars

Release conditions determine how much of the AI\-enabled pathway risk remains accessible to relevant actors\. Open weights are modeled as higher exposure than ordinary public API access because they permit unrestricted copying, fine\-tuning, scaffolding, removal of safety layers, and integration into autonomous tooling\. Trusted API is modeled as lower exposure, but not zero, because sophisticated AT3–AT5 actors may obtain access through front companies, compromised accounts, insiders, or indirect procurement\. Internal\-only use is modeled as residual exposure from insider misuse, compromise, or model theft\.

The exposure scalar decomposes as:

sr,j=Ar,j×Cr,j×Tr,j×Dr,j\+Lr,js\_\{r,j\}=A\_\{r,j\}\\times C\_\{r,j\}\\times T\_\{r,j\}\\times D\_\{r,j\}\+L\_\{r,j\}\(20\)
whereAAis actor access,CCis capability retention,TTis throughput,DDis detection failure, andLLis leakage\. The scalar is indexed by both release conditionrrand pathwayjj\. The Trusted API row is non\-uniform across pathways:Atrusted,jA\_\{\\mathrm\{trusted\},j\}runs from 0\.01 at AT1, because low\-skill actors rarely pass KYC, to near 1\.00 at AT5, because nation\-state actors can pose as legitimate users\. This produces a roughly 100×\\timesrange within a single release condition\.

Table XVI:Release\-condition exposure scalar matrix \(sr,js\_\{r,j\}\)\.Release conditionAT1AT2AT3AT4AT5Open weights2\.002\.002\.002\.002\.00Public API \(ref\.\)1\.001\.001\.001\.001\.00Safety\-filtered API0\.600\.580\.550\.500\.45Trusted API0\.0010\.0050\.0090\.0450\.090Internal only∼\\sim0∼\\sim0∼\\sim0∼\\sim0∼\\sim0Note\.These values are uncited author priors and audit targets, not measured safeguard properties\. Establishing them empirically would require measured inputs across several categories: KYC false\-accept rates \(for example against NIST SP 800\-63A IAL2\), presentation\-attack detection, jailbreak success, account\-compromise rates, and model\-weight security\. Sourcing these is deferred to dedicated open\-source\-intelligence work\. The Trusted API row is non\-uniform because sophisticated actors \(AT4–AT5\) are harder to exclude through access controls\.
### B\.4Required Safeguard Performance

The safeguard question can be stated directly: how much public\-access exposure must be removed before the model falls below a given acceptable\-harm benchmark? The answer can be written as a maximum allowable exposure scalar:

smax=HbenchmarkE​\[HAI,public\]s\_\{\\max\}=\\frac\{H\_\{\\mathrm\{benchmark\}\}\}\{E\[H\_\{\\mathrm\{AI,public\}\}\]\}\(21\)
At Mythos Preview level, using the USD 67\.90B central estimate:

smax​\(USD​1​B\)=1\.0​B67\.90​B=0\.0147s\_\{\\max\}\(\\text\{USD\}\\ 1\\text\{B\}\)=\\frac\{1\.0\\text\{B\}\}\{67\.90\\text\{B\}\}=0\.0147\(22\)
smax​\(USD​500​M\)=0\.5​B67\.90​B=0\.0074s\_\{\\max\}\(\\text\{USD\}\\ 500\\text\{M\}\)=\\frac\{0\.5\\text\{B\}\}\{67\.90\\text\{B\}\}=0\.0074\(23\)
Safeguards must reduce effective exposure to roughly 1\.5 percent of public API access to fall below the USD 1B SB\-53 benchmark, and to roughly 0\.7 percent for a USD 500M benchmark\. These ratios scale linearly with the USD 500 billion baseline, whose 90 percent confidence interval spans an order of magnitude, so the honest band runs from a few tenths of a percent to a few percent; the three\-figure values above \(0\.0147 and 0\.0074\) are central illustrations, not precise targets\. These are audit targets; they do not prove that any existing trusted\-access or safeguard regime achieves such reductions\. Trusted access is therefore not a label but a measurable exclusion target\. A company should provide independent evidence that KYC, monitoring, rate limits, jailbreak resistance, model\-weight security, and account\-compromise controls achieve the necessary exposure reduction\.

### B\.5Kill Chain Decomposition

TableLABEL:tab:b5\-killchainmaps attack phases to pathways and provides per\-phase conditional success probabilities\. AT3 domain compromise values are fromBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\), Table 12; other columns are author\-calibrated estimates\. This table is a structural illustration of the AND\-gate decomposition in \([6](https://arxiv.org/html/2607.16112#S3.E6)\): it shows how per\-phase probabilities compose into a full\-chainPsuccess,jP\_\{\\mathrm\{success\},j\}\. Its column products are not the values used in the headline harm totals\. Those totals read the AT3 pathway success probability off the TLO evidence in Table[VI](https://arxiv.org/html/2607.16112#S3.T6)\(PAI=0\.60P\_\{\\mathrm\{AI\}\}=0\.60at Mythos Preview level\), whereas the AT3 column here multiplies out to 0\.084\. The two numbers are estimates of the same quantity, the full\-chain completion rate, reached by different routes, and they differ by roughly sevenfold\. We flag this as an open modelling question rather than paper over it, and we do not substitute the 0\.084 product into the harm equations\. Lateral movement and exfiltration are marked as TLO\-relevant bottleneck phases, where AI uplift propagates multiplicatively through the product\.

Table XVII:Kill\-chain phase matrix\. Per\-phase conditional success probabilitypk,jp\_\{k,j\}\. N/A = phase does not apply\. Rows labeled “TLO bottleneck” indicate phases where AI uplift is especially consequential\.PhaseAT1:
PhishingAT2:
BECAT2:
RansomwareAT3:
DomainAT4:
Zero\-dayAT5:
StrategicReconnaissance1\.001\.001\.001\.001\.001\.00Initial access0\.400\.700\.600\.600\.200\.10ExecutionN/A0\.600\.500\.500\.700\.80Privilege escalationN/AN/A0\.700\.700\.800\.85Lateral movement
\(*TLO bottleneck*\)N/AN/A0\.650\.650\.750\.80CollectionN/AN/A0\.900\.900\.850\.90Exfiltration
\(*TLO bottleneck*\)N/AN/AN/A0\.850\.800\.85Impact / monetization0\.250\.400\.300\.800\.250\.10Psuccess,jP\_\{\\mathrm\{success\},j\}0\.1000\.1680\.0370\.0840\.0140\.004Note\.AT3 domain compromise values are adapted fromBarrettet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib9)\), Table 12 \(OC3 SME Ransomware\)\. Other values are author\-calibrated estimates\. “TLO bottleneck” phases are those emphasized byFolkertset al\.\([2026](https://arxiv.org/html/2607.16112#bib.bib15)\): lateral movement \(NTLM relay\) and exfiltration\.

## Appendix CCross\-Domain Synthesis Tables

TableLABEL:tab:domain\-comparisonsummarizes the harmonization status of each domain, and TableLABEL:tab:uncertaintyclassifies the principal inputs and claims by the kind of evidence each would require\.

Table XVIII:Cross\-domain comparison of harmonization status\.DomainCommon primitiveHarmonization statusMain blockerPaper contributionNext empirical stepCyberExpected annual harm \(USD\)Tractable; preliminary floor proposedCalibration uncertainty; lack of release\-condition audit dataExpected\-harm translation; non\-zero TLO completion floor; release\-condition decompositionHistorically calibrated pathway estimates and audited safeguard performanceBioriskExpected casualtiesMethodology applies; no quantitative floor yetNo unsaturated, validated, full\-chain KRI; near\-zero base rateDiagnosis of the missing evidence; staged kill\-chain floor structureDelphi expert elicitation of per\-stage uplift; Monte Carlo aggregationAutomated AI R&DRate of capability progress \(ECI\)Achievable now; operational floor proposedAttribution to AI automation; benchmark and harness driftQuantitative floor \(5×\\timestrend for 3 months\) covering the three main frontier companiesIndustry\-wide baseline; attribution evidenceNote\.The cyber and biorisk floors take expected harm as the common primitive followingKoessleret al\.\([2024](https://arxiv.org/html/2607.16112#bib.bib24)\); Murrayet al\.\([2025](https://arxiv.org/html/2607.16112#bib.bib34)\); the automated AI R&D floor uses a rate\-of\-progress primitive because no single harm pathway applies\.Table XIX:Uncertainty classification: structural assumptions versus empirical placeholders\.Parameter or claimCurrent statusEvidence neededImportanceLikely ownerCyber global baseline harmExternally anchored estimateReconciliation of damage estimates with reported\-loss dataHigh \(scales all results\)Cyber\-economics researchersCyber pathway allocationAuthor\-calibrated priorVictimization surveys, breach telemetry, ransomware recordsMedium \(sensitivity\-bounded\)Incident\-data holdersCyber release\-condition scalarsAuthor\-calibrated priorKYC, monitoring, jailbreak, and weight\-security auditsHigh \(decides release verdict\)Independent safeguard auditorsTLO as cyber KRIProposed primary KRIContinued longitudinal, multi\-company cyber\-range evaluationHigh \(defines the trigger\)Cyber\-eval bodies \(e\.g\. AISI\)AT4–AT5 uplift assumptionsAuthor\-calibrated prior \(small, non\-zero\)Analysis of AI uplift to nation\-state\-tier operationsLow \(small harm share\)Threat\-intelligence analystsBiorisk: no full\-chain validated KRIStated claimAudit of public bio\-evaluations against the three criteriaHigh \(blocks the floor\)Biosecurity evaluation communityBiorisk base\-rate problemStated limitationStructured counterfactual elicitation designHigh \(frames all estimates\)Biosecurity and biodefense expertsBiorisk expert elicitation designProposed, not completedDelphi panel with calibrated, diverse expertsHigh \(the missing input\)Bio elicitation consortiumAI R&D ECI metricProposed primary metricRobustness to benchmark and harness changesMedium \(alternatives exist\)Forecasting and eval reviewersAI R&D attribution to automationOpen questionTrend\-break evidence in AI R&D inputs and benchmarksMedium \(trigger vs\. audit\)AI progress researchersNote\.“Likely owner” denotes the kind of party best positioned to supply or contest the evidence, not an endorsement by any named organization\. Importance is judged relative to whether the input could change a threshold\-crossing verdict\.

Similar Articles

Why responsible AI development needs cooperation on safety

OpenAI Blog

OpenAI publishes a policy research paper identifying four strategies to improve industry cooperation on AI safety norms: communicating risks/benefits, technical collaboration, increased transparency, and incentivizing standards. The analysis addresses how competitive pressures could lead to under-investment in safety and proposes mechanisms to align incentives toward safe AI development.

AI safety and alignment

Reddit r/artificial

The article discusses concerns about AI safety and alignment as AI becomes more intelligent and integrated into society, referencing Anthropic's call for a pause to address potential catastrophic risks.

Frontier AI regulation: Managing emerging risks to public safety

OpenAI Blog

OpenAI proposes a regulatory framework for 'frontier AI' models that pose potential public safety risks, advocating for standard-setting processes, registration/reporting requirements, and compliance mechanisms including pre-deployment risk assessments and post-deployment monitoring.