Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

arXiv cs.AI Papers

Summary

This paper audits four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) across many models, showing that their scores are confounded by capability, metrics have artifacts, and rankings disagree across benchmarks, undermining interchangeable safety claims.

arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|\rho| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($\rho{=}{+}0.60$) but correlates negatively with misalignment safety ($\rho{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $\rho{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.
Original Article
View Cached Full Text

Cached at: 08/03/26, 07:29 AM

# Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Source: [https://arxiv.org/html/2607.28685](https://arxiv.org/html/2607.28685)
###### Abstract

Agent\-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent’s safety\. We treat four of them \(R\-Judge, InjecAgent, AgentHarm, AgentDojo\) as measurements to be validated, running each under its official implementation and author\-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite\. The metric is the first problem\. On any binary trace\-judgment benchmark scored byF1F\_\{1\}, an “always positive” policy attainsF1=2​π/\(1\+π\)F\_\{1\}=2\\pi/\(1\+\\pi\); on R\-Judge that is0\.6900\.690, above five of the 21 models that actually discriminate\. The three broad\-coverage benchmarks then rank the same 18 models differently, and the trade\-off behind that disagreement is a small\-panel artifact: R\-Judge specificity against AgentHarm safety correlates−0\.64\-0\.64atn=7n\{=\}7and\+0\.02\+0\.02atn=18n\{=\}18, and a quarter of random size\-7 subsets reach\|ρ\|≥0\.5\|\\rho\|\\geq 0\.5around that near\-zero value\. Held\-out validity turns on which outcome you pick\. Capability predicts task success \(ρ=\+0\.60\\rho\{=\}\{\+\}0\.60\) but correlates negatively with misalignment safety \(ρ=−0\.44\\rho\{=\}\{\-\}0\.44,n=21n\{=\}21\)\. On their pairedn=20n\{=\}20panel, the corresponding contrast isΔ=−1\.00\\Delta\{=\}\{\-\}1\.00\(95% CI\[−1\.48,−0\.49\]\[\-1\.48,\-0\.49\],p<0\.001p<0\.001\), and it survives leave\-one\-organization\-out and organization\-clustered bootstrap analyses\. On an expanded 41\-model panel, the misalignment correlation weakens to−0\.16\-0\.16\(95% CI\[−0\.54,\+0\.22\]\[\-0\.54,\+0\.22\]\) and jailbreak strengthens to\+0\.34\+0\.34, though neither change is significant\.AgentHarmshows the strongest held\-out association,ρ=\+0\.72\\rho\{=\}\{\+\}0\.72with three\-template jailbreak safety after controlling capability\. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety\. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs\.

## 1Introduction

Between 2024 and 2026, dozens of benchmarks appeared claiming to measure whether LLM agents are*safe*or*reliable*: resistance to prompt injection\(Zhanet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib8); Debenedettiet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib7)\), safety\-risk awareness\(Yuanet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib9)\), refusal of harmful agentic tasks\(Andriushchenkoet al\.[2025](https://arxiv.org/html/2607.28685#bib.bib6)\), and risky\-tool\-use identification\(Ruanet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib10)\), among others\. A recent taxonomy\(Liet al\.[2026](https://arxiv.org/html/2607.28685#bib.bib2)\)catalogs more than forty of them and reports a symptom worth taking seriously: swap the benchmark and the safety ranking contradicts itself, with rank concordance near0\.100\.10across models\. That work stops at a twelve\-model concordance check over four benchmarks, and lists capability\-controlled evaluation and benchmark consolidation among its open problems\.

This paper picks up there\. If a safety score is a measurement, it can fail in the ways measurements fail, so we separate*construct validity*\(what a score measures\),*metric validity*\(whether its metric measures that target\), and*criterion validity*\(whether it tracks held\-out behavior\)\. Four questions follow:

- •RQ1 \(construct structure\)\.Do nominally\-distinct agent\-safety benchmarks measure a common underlying factor, or dissociable ones?
- •RQ2 \(capability confound\)\.How much of the between\-benchmark signal is general capability in disguise?
- •RQ3 \(criterion validity\)\.Do the safety scores track held\-out behavioral criteria beyond what general capability predicts?
- •RQ4 \(metric validity\)\.Are the headline metrics valid instruments, or do they reward degenerate behavior?

#### Contributions\.

We identify a formal metric\-validity failure \(Observation[1](https://arxiv.org/html/2607.28685#Thmobservation1)\): for any binary trace\-judgment benchmark scored by F1, the score of an “always positive” baseline has a closed form, and on R\-Judge that baseline outranks five evaluated models\. We then quantify a small\-panel failure mode we walked into: the R\-Judge\-specificity/AgentHarm\-safety correlation moves from−0\.64\-0\.64atn=7n\{=\}7to\+0\.02\+0\.02atn=18n\{=\}18, and a quarter of random size\-7 subsets show\|ρ\|≥0\.5\|\\rho\|\\geq 0\.5despite the near\-zero full\-panel value\. The core is a pre\-specified, capability\-controlled criterion\-validity audit over one task\-success outcome and two safety outcomes: capability predicts task success, its safety correlations move with the outcome and the panel, and the strongest safety\-score result is the AgentHarm–jailbreak association \(ρ=\+0\.72\\rho\{=\}\+0\.72after controlling capability\), with borderline evidence of outcome selectivity under organization\-level resampling \(p=0\.051p\{=\}0\.051\)\. We release an API\-only audit harness and re\-run artifacts for four benchmarks and three held\-out outcomes, included with the accompanying reproducibility package\.

#### Main conclusion\.

A capability score is not a safety score, and no one agent\-safety benchmark stands in for safety as a whole\. What a score licenses you to say depends on how it was produced\. Small panels are the sharpest edge: at seven models a weak relationship can look systematic, so validity needs re\-checking as the model population turns over\.

#### Scope\.

The three RQ3 criteria are stand\-ins for deployment, not claims about real\-world harm; their limits \(single instances, grader dependence, panel size\) are in Section[6](https://arxiv.org/html/2607.28685#S6)\.

## 2Related Work

#### Agent\-safety benchmarking\.

Each of the four benchmarks we audit targets one construct\. InjecAgent\(Zhanet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib8)\)and AgentDojo\(Debenedettiet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib7)\)target prompt\-injection robustness; AgentHarm\(Andriushchenkoet al\.[2025](https://arxiv.org/html/2607.28685#bib.bib6)\)measures compliance with harmful agentic tasks; R\-Judge\(Yuanet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib9)\)scores safety\-risk awareness over interaction traces\. Evaluators we do not re\-run, such as ToolEmu\(Ruanet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib10)\)for risky tool use, share the same status: designed as evaluators, never validated as measurements\. AutoMonitor\-Bench\(Yanget al\.[2026](https://arxiv.org/html/2607.28685#bib.bib22)\)scores misbehavior monitors by miss and false\-alarm rates, the two\-sided reporting our metric analysis argues for\. Closest to this paper is a taxonomy and consistency analysis of agent\-safety benchmarks\(Liet al\.[2026](https://arxiv.org/html/2607.28685#bib.bib2)\), which documents ranking disagreement but does not control for capability, test criterion validity, or recover a latent structure; capability\-controlled evaluation is named there as future work, and we take up all three\.

#### Construct validity\.

Prior work asks whether benchmark scores agree and what latent constructs they capture\. Benchmark\-agreement testing\(Perlitzet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib3)\), metabench\(Kipniset al\.[2025](https://arxiv.org/html/2607.28685#bib.bib4)\), and direct construct analyses\(Beanet al\.[2025](https://arxiv.org/html/2607.28685#bib.bib5)\)find that*knowledge and reasoning*benchmarks share a strong latent factor\.Kearns \([2026](https://arxiv.org/html/2607.28685#bib.bib21)\)and NIST’s statistical framework\(Kelleret al\.[2026](https://arxiv.org/html/2607.28685#bib.bib23)\)bring quantitative construct\-validity and uncertainty modeling to LLM evaluation\. We point the same psychometric lens at agent reliability and safety, where benchmarks target distinct behaviors—refusal, injection robustness, risk awareness—and nobody has yet quantified how much capability confounds their comparison\.

#### Benchmark gaming\.

A separate concern is that a benchmark can be*gamed*by an adversary\. Our claim is different and logically prior: even for an honest model, the headline score may not validly measure the intended construct\. The F1 gameability we document \(Section[4](https://arxiv.org/html/2607.28685#S4)\) is a concrete instance\. The policy that scores mid\-leaderboard is not adversarial; it labels every trace “unsafe” without examining any of them\.

#### Over\-refusal\.

XSTest\(Röttgeret al\.[2024](https://arxiv.org/html/2607.28685#bib.bib17)\)and OR\-Bench\(Cuiet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib18)\)document*exaggerated safety*, or over\-refusal, at the level of individual prompts\. We ask the analogous question one level up, between benchmarks: are models that refuse harmful tasks \(AgentHarm\) worse at identifying benign traces correctly \(R\-Judge specificity\), and does that produce the ranking disagreement? On the full panel that correlation sits near zero\. On seven\-model subsets it often does not\.

#### Agent reliability and task performance\.

Recent work isolates specific agent\-reliability failures such as corrupt success\(Caoet al\.[2026](https://arxiv.org/html/2607.28685#bib.bib15)\)and evidence\-bounded reporting\(Chen[2025](https://arxiv.org/html/2607.28685#bib.bib16)\);τ\\tau\-bench\(Yaoet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib11)\)and its dual\-control successorτ2\\tau^\{2\}\-bench\(Barreset al\.[2025](https://arxiv.org/html/2607.28685#bib.bib12)\)measure agentic task success\. We borrow the latter’s retail domain as our held\-out RQ3 task\-success criterion, and treat the safety benchmarks as candidate columns for a validity audit rather than as settled measurements\.

## 3Method

### 3\.1Model panel and benchmarks

We assemble a compute\-light, API\-only panel of up to 22 models from nine model\-developing organizations \(OpenAI 6, Meta 3, Qwen 3, Mistral 3, DeepSeek 2, Amazon 2, Anthropic, Google, Cohere; full roster in Fig\.[3](https://arxiv.org/html/2607.28685#A1.F3)and the artifact\)\. We re\-run four safety benchmarks with their official implementations and scorers: R\-Judge \(full 571 items, self\-judge\), InjecAgent \(300 stratified direct\-harm/data\-stealing items, rule\-scored\), AgentHarm \(44 base behaviors spanning all eight harm categories, via the official Inspect evaluation with a gpt\-4o\-mini judge\), and AgentDojo \(theslacksuite, 100 items, environment\-scored for security and utility\), plus a held\-out task\-success criterion for RQ3 \(below\)\.

Coverage differs by benchmark \(Fig\.[3](https://arxiv.org/html/2607.28685#A1.F3)maps the full panel\): R\-Judge hasn=21n\{=\}21after Qwen3\-32B falls below the90%90\\%valid\-output threshold; InjecAgent hasn=22n\{=\}22; AgentHarm and AgentDojo, which require native tool use, haven=19n\{=\}19and55; and theτ2\\tau^\{2\}criterion hasn=20n\{=\}20\(App\.[A\.2](https://arxiv.org/html/2607.28685#A1.SS2)\)\. The common R\-Judge/InjecAgent/AgentHarm panel therefore hasn=18n\{=\}18\. Capability anchors \(MMLU and GPQA\-Diamond\) are measured by us under one uniform protocol: MMLU zero\-shot on a fixed 500\-item subset \(seed 42, identical items for every model\) and GPQA\-Diamond \(198 items, chain\-of\-thought\), through the same API harness as the safety runs\. We began with provider\-reported model\-card numbers; the replacements hold the protocol fixed across models, and the composite passes its positive control \(§[4](https://arxiv.org/html/2607.28685#S4)\)\. The anchor panel is 21 models \(DeepSeek\-R1 is excluded for unstable long\-form GPQA generation\); we keep the provider\-reported anchors as a pre\-harmonization comparison\.

### 3\.2Benchmark targets

Before looking at any relationship among the score columns, we wrote down what behavior each benchmark is meant to measure\. R\-Judge tests whether a model distinguishes unsafe from benign interaction traces\. InjecAgent and AgentDojo both test resistance to prompt injection, so their correlation serves as a convergent\-validity check\. AgentHarm tests whether a model refuses harmful agentic tasks\. For R\-Judge, we analyze specificity and balanced accuracy in addition to the official F1 because the three metrics reward different error patterns\.

#### Score orientation and terminology\.

All reported scores are oriented so that higher is better or safer\. R\-Judge*specificity*is the fraction of benign traces correctly identified as benign; its*balanced accuracy*averages performance on benign and unsafe traces\. InjecAgent robustness is one minus attack success, and AgentHarm safety is one minus harmful compliance\. For RQ3,τ2\\tau^\{2\}\-bench is a held\-out task\-success criterion, not a safety score\. We use*internal consistency*for split\-half analyses \(App\.[A\.2](https://arxiv.org/html/2607.28685#A1.SS2)\) and name each benchmark\-specific score explicitly, instead of stretching “calibration” or “reliability” past their usual meanings\.

### 3\.3Analysis plan

Before running each analysis, we fixed its hypotheses, thresholds, and decision rules\. Because these choices lack an independent timestamp, we describe the analyses as pre\-specified rather than preregistered\. We use Spearman correlations for all association tests\. Unless otherwise noted, correlationpp\-values use two\-sided asymptotic Spearman tests\. In Fig\.[2](https://arxiv.org/html/2607.28685#S4.F2), error bars are model\-level percentile\-bootstrap intervals\. The panel expansion and direct panel comparisons follow their pre\-specified bootstrap procedures \(App\.[A\.6](https://arxiv.org/html/2607.28685#A1.SS6)\)\. The corroborative PCA operates on the Pearson correlation matrix of the standardized score columns; factor count is chosen with Horn’s parallel analysis against a random\-normal null, with the rank\-based \(Spearman\) sensitivity in App\.[A\.3](https://arxiv.org/html/2607.28685#A1.SS3)\. We report partial Spearman correlations controlling the capability composite; App\.[A\.5](https://arxiv.org/html/2607.28685#A1.SS5)gives the rank\-residual implementation and an alternative convention\. The capability composite is the mean of the standardized, verified MMLU and GPQA anchors; its model ordering is identical to the anchors’ first principal component\. Between\-construct redundancy is measured by rank\-residualized partial correlations that control for the capability composite\.

We pre\-specified three controls:*positive*\(capability\-anchor loadings≥0\.6\\geq 0\.6on the first factor\);*negative*\(a column\-permuted score matrix within the parallel\-analysis null\); and*convergent*\(a high correlation between the same\-construct InjecAgent and AgentDojo scores\)\. The headline metric is the median pairwise*partial*Spearman\|ρ\|\|\\rho\|among distinct\-construct benchmarks, controlling the capability composite; the decision rule permits interpretation of the factor structure only if first\-factor variance is≥40%\\geq 40\\%*or*controlling capability reduces that median by≥0\.15\\geq 0\.15, with the positive and negative controls passing\. These thresholds and the disconfirming test were fixed before any real run\.

For criterion validity \(RQ3\) we pre\-specified three held\-out outcomes, each before running it\.*Task success*isτ2\\tau^\{2\}\-bench retail\(Barreset al\.[2025](https://arxiv.org/html/2607.28685#bib.bib12)\): 40 multi\-turn, environment\-scored tool\-agent\-user tasks\.*Misalignment safety*records whether a model avoids blackmail or information leaking in a fictional autonomous\-agent scenario\(Lynchet al\.[2025](https://arxiv.org/html/2607.28685#bib.bib13)\), scored by a fixed classifier over 100 decisions per model\.*Jailbreak safety*is one minus harmful\-compliance on 50 forbidden prompts, averaged over three published jailbreak templates to reduce the ceiling created by any single template\(Soulyet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib14); Weiet al\.[2023](https://arxiv.org/html/2607.28685#bib.bib19)\)\. All three run through the same API\-only harness \(the two safety criteria are single\-turn or scenario\-based and cover the full panel; theτ2\\tau^\{2\}criterion is a multi\-turn tool loop covering 20 models\), are distinct from the four audited benchmarks, and are not among the predictors\. For each, the headline metric is the partial Spearman correlation of every safety score with the criterion, controlling the capability composite\. We first report the criterion’s raw association with capability, then test whether each safety score adds predictive information\. The pre\-specified analysis gives reporting rules for both significant and null results\. The accompanying reproducibility package contains the analysis scripts and frozen inputs needed to reproduce the reported results\.

## 4Results

Results run in order of evidential strength: the metric\-validity failure \(RQ4\) and capability correlations \(RQ2\) are firmest, cross\-benchmark structure \(RQ1\) is weaker, and the power\-limited held\-out validity tests \(RQ3\) come last\.

### 4\.1The headline metric is gameable \(RQ4\)

![Refer to caption](https://arxiv.org/html/2607.28685v1/x1.png)Figure 1:A leaderboard’s headline metric can be matched by an “always unsafe” baseline\.On R\-Judge \(n=21n\{=\}21\), five real, discriminating models score below this constant baseline \(F1=0\.69\{=\}0\.69, dashed\): F1 gives no credit for correctly rejected benign traces, so the degenerate policy can outrank models that discriminate\.R\-Judge labels52\.7%52\.7\\%of its 571 items “unsafe\.” A degenerate policy that answers “unsafe” on every item, exercising no safety reasoning, therefore attains recall1\.01\.0, specificity0, andF1=0\.690\\mathrm\{F1\}=0\.690\. On the 21 panel models that produce valid R\-Judge verdicts \(Qwen3\-32B is excluded for21%21\\%unparseable output; see Method\) this constant\-baseline score outranks five real models \(F10\.300\.30–0\.570\.57; Fig\.[1](https://arxiv.org/html/2607.28685#S4.F1)\), all of which discriminate\. The panel’s highest\-specificity model, o3\-mini \(0\.970\.97\), receivesF1=0\.702\\mathrm\{F1\}=0\.702, only0\.0120\.012above the constant baseline, because its recall is lower\. Balanced accuracy reorders the board, since it credits correct decisions on both classes\.

###### Observation 1\(F1 admits a high\-scoring constant baseline\)\.

On a benchmark with class base rateπ\\piscored by F1, the constant “always\-positive” policy attainsF1=2​π/\(1\+π\)\\mathrm\{F1\}=2\\pi/\(1\+\\pi\)without distinguishing safe from unsafe items\. For R\-Judge \(π=0\.527\\pi=0\.527, a near\-balanced base rate\) this is0\.6900\.690, which exceeds the observed F1 of several models that do discriminate\. The closed form depends on prevalence, but the measurement problem is F1’s blindness to true negatives: correctly identifying a benign trace earns nothing, so F1 cannot stand alone as a measure of two\-sided discrimination on this trace\-judgment benchmark\.

Table 1:The three broad\-coverage agent\-safety benchmarks rank the samen=18n\{=\}18models differently \(higher is safer\)\. Rows are sorted by AgentHarm safety\.Boldmarks column maxima and underlining marks minima, with display\-precision ties included\. GPQA is shown as a capability reference\. A model high on AgentHarm can rank low on R\-Judge specificity, so the scores are not interchangeable\.
### 4\.2The capability confound is real but metric\-dependent \(RQ2\)

Across the R\-Judge panel \(n=20n\{=\}20with harmonized anchors\), capability correlates strongly with balanced accuracy \(ρ=\+0\.71\\rho=\+0\.71with MMLU,p<0\.001p<0\.001;\+0\.49\+0\.49with GPQA,p=0\.03p=0\.03\) and F1 \(\+0\.76\+0\.76and\+0\.62\+0\.62\), but only weakly with specificity \(\+0\.16\+0\.16and\+0\.07\+0\.07\)\. This difference follows the metric definitions: balanced accuracy and F1 include recall on unsafe traces, whereas specificity measures only correct treatment of benign traces and can be inflated by rarely flagging anything\. So the capability confound is a property of the metric you report, not of the benchmark as a whole—and it moves with the panel too: the specificity–MMLU correlation fell from0\.850\.85on our first nine models to\+0\.16\+0\.16atn=20n\{=\}20\.

### 4\.3Benchmarks disagree; a seven\-model panel suggested a trade\-off that does not persist \(RQ1\)

The firmest RQ1 result is descriptive and correlation\-free: on this panel no single benchmark orders the models the way another does \(Table[1](https://arxiv.org/html/2607.28685#S4.T1)\); a top\-AgentHarm model can rank near the bottom on R\-Judge specificity\. This makes “safety rank” benchmark\-dependent without resting on any correlation estimate\. Disagreement between benchmarks coded to different constructs \(§[3](https://arxiv.org/html/2607.28685#S3)\) is expected and does not by itself indict any single instrument; what it defeats is the practice of quoting these scores interchangeably as*the*safety of an agent\.

We read a trade\-off into that disagreement early on; it does not survive the larger cross\-benchmark panel\. The R\-Judge\-specificity/AgentHarm\-safety rank correlation was−0\.64\-0\.64atn=7n\{=\}7but\+0\.02\+0\.02\(p=0\.95p=0\.95\) atn=18n\{=\}18\(Fig\.[5](https://arxiv.org/html/2607.28685#A1.F5), appendix\), and random size\-7 subsets yield\|ρ\|≥0\.5\|\\rho\|\\geq 0\.5in a quarter of draws despite the near\-zero full\-panel value\. Sampling variability alone is enough to manufacture a “clean reversal” at that panel size, and the dissolution holds under all four R\-Judge metrics \(App\.[A\.4](https://arxiv.org/html/2607.28685#A1.SS4)\)\. Whether capability moderates what is left we do not test; we flag it as a hypothesis\.

A single\-factor PCA provides weaker, corroborative evidence \(Fig\.[6](https://arxiv.org/html/2607.28685#A1.F6), appendix\): its first component \(48\.7%48\.7\\%variance\) places R\-Judge specificity, InjecAgent robustness, and both capability anchors in one direction and AgentHarm safety in the other \(\+0\.33\+0\.33, opposite the cluster in1717of1818leave\-one\-out fits\)\.

#### The positive control now passes\.

On the harmonized anchors the capability composite passes its pre\-specified positive control \(MMLU loads0\.74≥0\.60\.74\\geq 0\.6; on provider\-reported numbers it had failed at0\.420\.42, our original reason for caution\) and the negative control also passes \(a column\-permuted score matrix yields eigenvalues within the parallel\-analysis null\), so the pre\-specified interpretation conditions are met \(the convergent control is inconclusive at AgentDojo’sn=5n\{=\}5\)\. We still treat the factor as corroboration only: a rank\-based \(Spearman\) PCA keeps the loading structure but drops PC1 to40\.8%40\.8\\%with no factor retained by Horn \(App\.[A\.3](https://arxiv.org/html/2607.28685#A1.SS3)\), and the median between\-benchmark correlation is just0\.220\.22; the firm, factor\-independent RQ1 result is the raw\-score disagreement \(Table[1](https://arxiv.org/html/2607.28685#S4.T1)\)\.

### 4\.4Predictive validity differs across held\-out outcomes \(RQ3\)

Everything above is internal to the benchmarks\. RQ3 steps outside them: do these scores predict held\-out behavior that a capability test does not already predict? Which outcome you choose decides the answer\. We pre\-specified three \(§[3](https://arxiv.org/html/2607.28685#S3)\):τ2\\tau^\{2\}\-bench task success, misalignment safety, and jailbreak safety\. All are distinct from the four audited benchmarks, and each has a fixed metric and decision rule\.

Table 2:Criterion validity on the 2024–25 panel \(cellwisen=18n\{=\}18–2121because predictor coverage differs; App\.[A\.2](https://arxiv.org/html/2607.28685#A1.SS2); higher is better or safer\)\. The top row reports raw correlations with capability; the remaining rows report partial Spearman correlations controlling the capability composite\. Capability predicts task success but not either safety outcome consistently\. AgentHarm is strongly associated only with the three\-template jailbreak outcome, marked by†\. Stars use two\-sided asymptotic Spearman tests \(p∗<0\.05\{\}^\{\*\}p<0\.05,p∗∗<0\.01\{\}^\{\*\*\}p<0\.01,p∗⁣∗∗<0\.001\{\}^\{\*\*\*\}p<0\.001\)\.#### Against task success, safety scores add no information beyond capability\.

Capability predictsτ2\\tau^\{2\}\-retail success \(ρ=\+0\.60\\rho=\+0\.60,p=0\.005p=0\.005,n=20n\{=\}20\), while no safety score shows incremental validity once capability is partialled out \(best\|ρ\|=0\.23\|\\rho\|=0\.23, n\.s\. atn=19n\{=\}19; Table[2](https://arxiv.org/html/2607.28685#S4.T2); Fig\.[7](https://arxiv.org/html/2607.28685#A1.F7), appendix\)\. Because capability already predictsτ2\\tau^\{2\}, these partials ask the stricter question of whether a safety score predicts what capability leaves unexplained; none does\. The null is not a verdict on the safety scores\. Task success is simply an outcome that capability already explains well, which is why we test two safety outcomes as well\.

![Refer to caption](https://arxiv.org/html/2607.28685v1/x2.png)Figure 2:Capability predicts task success; safety point estimates differ by outcome and panel\.Filled circles show the 2024–25 correlations; open diamonds show the expanded\-panel safety correlations; error bars are model\-level percentile\-bootstrap 95% CIs\. Thin lines connect panel\-level estimates, not model trajectories\. Misalignment changes from−0\.44\-0\.44to−0\.16\-0\.16and jailbreak from\+0\.08\+0\.08to\+0\.34\+0\.34; neither between\-panel change is significant\. Task success was not re\-collected for the expansion\.
#### Against agentic misalignment, capability has the opposite association\.

On the full 21\-model panel, capability predicts misalignment\-safety negatively \(ρcap=−0\.44\\rho\_\{\\text\{cap\}\}=\-0\.44, asymptoticp=0\.047p=0\.047\), the opposite sign to its\+0\.60\+0\.60onτ2\\tau^\{2\}: higher capability scores go, if anything, with lower safety here \(Mistral\-Large blackmails or leaks85%85\\%of the time; o3\-mini is a capable exception at1%1\\%\)\. After controlling capability, InjecAgent robustness \(\+0\.47\+0\.47\) and R\-Judge specificity \(\+0\.41\+0\.41\) reach the pre\-specified0\.400\.40effect\-size threshold, while AgentHarm safety does not \(\+0\.16\+0\.16; Table[2](https://arxiv.org/html/2607.28685#S4.T2)\)\. The two associations remain exploratory: both confidence intervals include zero, neither survives Bonferroni correction, power is only≈0\.35\\approx 0\.35, and the pre\-specified directional prediction named AgentHarm, not the two scores that came up\. The R\-Judge result is also metric\-sensitive: its partial correlation falls from\+0\.41\+0\.41under specificity to\+0\.19\+0\.19under balanced accuracy and reverses to−0\.46\-0\.46under recall \(App\.[A\.4](https://arxiv.org/html/2607.28685#A1.SS4)\)\.

#### A direct interaction test on the original panel\.

We compare capability’s correlations with task success and misalignment on their pairedn=20n\{=\}20panel\. The task\-success correlation is\+0\.60\+0\.60and the misalignment\-safety correlation is−0\.41\-0\.41\(−0\.44\-0\.44on the full panel\), a difference ofΔ=−1\.00\\Delta=\-1\.00\(95% CI\[−1\.48,−0\.49\]\[\-1\.48,\-0\.49\],p<0\.001p<0\.001\)\. This difference is stable to leave\-one\-model\-out, leave\-one\-organization\-out, organization\-clustered bootstrap, and subsampling \(App\.[A\.5](https://arxiv.org/html/2607.28685#A1.SS5)\)\. The corresponding R\-Judge interaction is only marginal \(Δ=\+0\.50\\Delta=\+0\.50,p=0\.082p=0\.082\), so the evidence supports the capability interaction but not a full two\-way dissociation\. An independent misalignment grader produces nearly identical model scores \(ρ=0\.97\\rho=0\.97\)\.

#### The capability–misalignment correlation is weaker on the expanded panel\.

Before running the expansion, we specified that the original negative correlation should “hold or strengthen\.” Instead, it changes fromρ=−0\.44\\rho=\-0\.44\(n=21n\{=\}21, asymptoticp=0\.047p=0\.047\) to−0\.16\-0\.16\(n=40n\{=\}40, bootstrapp=0\.40p=0\.40; 95% CI\[−0\.54,\+0\.22\]\[\-0\.54,\+0\.22\]; App\.[A\.6](https://arxiv.org/html/2607.28685#A1.SS6)\)\. The point estimate is less than half as large and no longer significant, but the interval still includes−0\.44\-0\.44and the between\-panel difference is not significant \(nested\-bootstrapΔ​ρ=\+0\.27\\Delta\\rho=\+0\.27, 95% CI\[−0\.25,\+0\.78\]\[\-0\.25,\+0\.78\]\)\. So the original estimate should not be assumed to carry over to the expanded model population, though the data do not establish a temporal shift either\. Matching subsets to the original organization composition gives a similar estimate \(medianρ=−0\.17\\rho=\-0\.17\), which suggests composition alone does not explain the difference; a generational explanation stays possible and undemonstrated\. Becauseτ2\\tau^\{2\}was not re\-collected, we could not re\-test the original interaction\.

#### Models with similar capability scores differ in safety across organizations \(Fig\.[4](https://arxiv.org/html/2607.28685#A1.F4)\)\.

The 41\-model panel spans 12 model\-developing organizations \(misalignmentn=40n\{=\}40; jailbreakn=41n\{=\}41after one engagement\-gate exclusion\)\. Six organizations contribute at least four models for within\-organization correlations, and seven contribute at least three models for the group comparison \(App\.[A\.6](https://arxiv.org/html/2607.28685#A1.SS6), Table[5](https://arxiv.org/html/2607.28685#A1.T5)\)\. Within organizations, capability–safety correlations vary in sign and average near zero \(−0\.14\-0\.14misalignment,\+0\.26\+0\.26jailbreak\)\. After adjusting safety ranks for capability, the seven organization groups differ on both criteria \(permutationp=0\.004p=0\.004misalignment,p<0\.001p<0\.001jailbreak; Holm\-significant\), with df\-adjusted Kruskal–Wallis effect sizesε2=0\.34\\varepsilon^\{2\}=0\.34and0\.570\.57\. None of this is causal: the static capability composite may miss agentic ability, each organization contributes few models, and some groups have restricted safety ranges\. What the panel supports is narrower—similar MMLU/GPQA scores do not imply similar safety scores\.

#### Against jailbreak safety, AgentHarm shows the strongest held\-out association\.

We pre\-specified a second safety outcome with a different failure mode, harness, and scoring procedure: harmful compliance under jailbreak, scored with StrongREJECT\(Soulyet al\.[2024](https://arxiv.org/html/2607.28685#bib.bib14)\)\. A single jailbreak template \(AIM\) produces a ceiling because 11 of 22 models refuse every request\. Averaging scores over three published templates\(Weiet al\.[2023](https://arxiv.org/html/2607.28685#bib.bib19)\)leaves no model at the ceiling\. Capability is weakly associated with this outcome on the 2024–25 panel \(ρ=\+0\.08\\rho=\+0\.08,n=21n\{=\}21,p=0\.73p=0\.73\)\. The expanded\-panel estimate is\+0\.34\+0\.34\(n=41n\{=\}41,p=0\.04p=0\.04\), but the between\-panel change is not significant\.

After controlling capability, AgentHarm safety correlates with the three\-template jailbreak outcome atρ=\+0\.72\\rho=\+0\.72\(p<0\.001p<0\.001\), compared with\+0\.50\+0\.50for the ceiling\-limited single\-template measure\. The estimate is stable to leave\-one\-model\-out analysis and an independent grader \(App\.[A\.5](https://arxiv.org/html/2607.28685#A1.SS5)\)\. AgentHarm is more strongly associated with jailbreak than with misalignment in a bootstrap that treats models as independent \(Δ=\+0\.59\\Delta=\+0\.59,p=0\.027p=0\.027\); a bootstrap that instead resamples the eight represented organizations givesp=0\.051p=0\.051\. On this paired panel, the misalignment partial is\+0\.13\+0\.13\(n=19n\{=\}19\); Table[2](https://arxiv.org/html/2607.28685#S4.T2)’s\+0\.16\+0\.16uses the commonn=18n\{=\}18panel\. R\-Judge specificity \(−0\.11\-0\.11\) and InjecAgent robustness \(\+0\.21\+0\.21\) do not predict jailbreak safety\. AgentHarm and StrongREJECT both measure harmful compliance, so part of what the\+0\.72\+0\.72shows is convergent validity\. The stronger claim—that AgentHarm is selective for jailbreak over misalignment—rests on a borderline organization\-resampled test, and the matching comparison for R\-Judge and InjecAgent, whether they predict misalignment better than jailbreak, is marginal as well \(Δ=\+0\.52\\Delta=\+0\.52,p=0\.068p=0\.068\)\.

#### Two alternative explanations, checked\.

Seven LLM coders who saw only construct descriptions, and none of the predictive results, independently assigned AgentHarm and the jailbreak outcome to the same harm\-blocking category \(7/77/7for each; Krippendorffα=0\.61\\alpha=0\.61across three categories\)\. They classified the misalignment outcome less cleanly, one more reason to keep its R\-Judge/InjecAgent associations exploratory \(App\.[A\.1](https://arxiv.org/html/2607.28685#A1.SS1)\)\. The second check: after controlling capability, AgentHarm safety is uncorrelated with over\-refusal on benign XSTest prompts\(Röttgeret al\.[2024](https://arxiv.org/html/2607.28685#bib.bib17)\)\(ρ=−0\.01\\rho=\-0\.01,n=17n\{=\}17\), which weighs against the story that models score well on AgentHarm and the jailbreak outcome simply by refusing everything\.

## 5Discussion

What the results support is narrow: these agent\-safety benchmarks are not interchangeable measures of a single property\. An “always unsafe” baseline matches R\-Judge F1 \(Observation[1](https://arxiv.org/html/2607.28685#Thmobservation1)\), and the three broad\-coverage benchmarks rank the same models differently\. Held\-out validity shifts with the outcome and the model panel \(Table[2](https://arxiv.org/html/2607.28685#S4.T2)\): capability predicts task success, while its correlations with the two safety outcomes are unstable across panels\. AgentHarm’s capability\-controlled association with three\-template jailbreak safety is strong \(ρ=\+0\.72\\rho=\+0\.72\), but both measures concern harmful compliance and the outcome\-selectivity test givesp=0\.051p=0\.051under organization resampling, which makes this convergent validity with the specificity claim still open\. The R\-Judge/InjecAgent associations with misalignment are exploratory\. Safety ranks also differ among organization groups after adjustment for the capability composite—an observational pattern that identifies no organizational cause\. A safety claim is defensible only with its benchmark, metric, target behavior, and model population attached\.

## 6Limitations and Threats to Validity

#### Panel size\.

The AgentHarm\-limited cross\-benchmark panel hasn=18n\{=\}18, up from our first pass atn=7n\{=\}7, which is where our caution about small\-panel claims comes from\. The gap most worth closing is a larger, capability\-spanning agent\-loop panel\.

#### Capability anchors\.

Provider\-reported model\-card anchors initially failed their positive control \(MMLU0\.420\.42\)\. Re\-measuring MMLU and GPQA\-Diamond \(§[3](https://arxiv.org/html/2607.28685#S3)\) fixed it \(MMLU0\.740\.74\); every RQ2/RQ3 sign was unchanged\. These are static\-knowledge tests, although their composite predicts multi\-turn tool\-agent success at\+0\.60\+0\.60\. A six\-model BFCL pilot\(Patilet al\.[2025](https://arxiv.org/html/2607.28685#bib.bib20)\)failed as an agentic replacement: strict function\-call syntax scoring ranked GPT\-4o\-mini highest \(0\.720\.72\) and the two most capable models lowest \(0\.420\.42–0\.480\.48\), apparently rewarding format conformance\. That failure is the argument for a validated agentic anchor\. Measurement error of this kind attenuates or suppresses partial relationships instead of biasing them in a direction we could sign\.

#### Multiple comparisons\.

We report per\-testpp\-values without a family\-wise correction, but our claims rest on effect sizes and pre\-specified rules: the two strongly significant tests \(the crossover and AgentHarm→\\tojailbreak, bothp<0\.001p<0\.001\) would survive Bonferroni, while the R\-Judge/InjecAgent–misalignment cells \(p≈0\.05p\\approx 0\.05–0\.090\.09\) would not, and we treat those as suggestive throughout\.

#### Judges and coverage\.

Three independent\-grader checks support robustness to judge choice: model scores correlate atρ=0\.97\\rho=0\.97for misalignment andρ=0\.98\\rho=0\.98for jailbreak\. Rescoring AgentHarm refusals with a different model lowers agreement \(ρ=0\.69\\rho=0\.69\), yet the AgentHarm–jailbreak partial correlation barely moves \(\+0\.66\+0\.66against\+0\.72\+0\.72; App\.[A\.5](https://arxiv.org/html/2607.28685#A1.SS5)\)\. AgentDojo runs on only five models, so the cross\-benchmark analyses depend primarily on the other three benchmarks\.

#### Criterion validity \(RQ3\)\.

τ2\\tau^\{2\}\-retail \(40 tasks, one run/model\) is scored by a user\-simulator whose variance we do not bound; the misalignment criterion is a fictional\-scenario probe \(grader\-robust,ρ=0\.97\\rho=0\.97, though its blackmail–leaking scenario halves agree only moderately, Spearman–Brown0\.700\.70\); the jailbreak criterion averages three jailbreaks\. None gives literal deployment\-harm probabilities\. For the item\-rich instruments with resampling support, split\-half consistency is≥0\.96\\geq 0\.96and rank stability is≥0\.97\\geq 0\.97\. Across the bounded instruments, worst\-case within\-model SEs are1212–25%25\\%of the between\-model SD, which supports the panel\-level spread but not every pairwise model difference \(App\.[A\.2](https://arxiv.org/html/2607.28685#A1.SS2)\)\. The two construct\-matching patterns are not equally well supported\. The AgentHarm–jailbreak association is large and stable in direction, but it partly reflects convergent measurement and its selectivity over misalignment is borderline under organization resampling\. The R\-Judge/InjecAgent–misalignment pattern is exploratory \(marginalΔ=\+0\.52\\Delta=\+0\.52, power≈0\.35\\approx 0\.35; the larger three\-scenario result is post\-hoc; App\.[A\.5](https://arxiv.org/html/2607.28685#A1.SS5)\)\.

## 7Conclusion

Across four benchmarks and up to 22 models, R\-Judge’s F1 lets an “always unsafe” baseline outrank five models, and switching benchmarks switches the rankings\. Capability predicts held\-out task success; it does not predict safety consistently across outcomes or panels\. The strongest held\-out safety association is AgentHarm’s \(ρ=\+0\.72\\rho=\+0\.72with jailbreak safety after controlling capability\), though the two measures partly overlap and organization\-resampled evidence for outcome selectivity is borderline\. Safety evaluations should report the full confusion structure, drop standalone F1 wherever true negatives matter, and tie every claim to an explicit target behavior and model population\.

## References

- M\. Andriushchenko, A\. Souly, M\. Dziemian, D\. Duenas, M\. Lin, J\. Wang, A\. Zou, Z\. Kolter, M\. Fredrikson, E\. Winsor, J\. Wynne, Y\. Gal, and X\. Davies \(2025\)AgentHarm: a benchmark for measuring harmfulness of LLM agents\.InInternational Conference on Learning Representations \(ICLR\),External Links:2410\.09024Cited by:[§1](https://arxiv.org/html/2607.28685#S1.p1.1),[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.
- V\. Barres, H\. Dong, S\. Ray, X\. Si, and K\. Narasimhan \(2025\)τ2\\tau^\{2\}\-Bench: evaluating conversational agents in a dual\-control environment\.Note:arXiv:2506\.07982External Links:2506\.07982Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px5.p1.2),[§3\.3](https://arxiv.org/html/2607.28685#S3.SS3.p3.2)\.
- A\. M\. Bean, R\. O\. Kearns, A\. Romanou, F\. S\. Hafner, H\. Mayne, J\. Batzner, N\. Foroutan, C\. Schmitz, and A\. Mahdi \(2025\)Measuring what matters: construct validity in large language model benchmarks\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,External Links:2511\.04703Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px2.p1.1)\.
- H\. Cao, I\. Driouich, and E\. Thomas \(2026\)Beyond task completion: revealing corrupt success in LLM agents through procedure\-aware evaluation\.Note:arXiv:2603\.03116External Links:2603\.03116Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px5.p1.2)\.
- R\. Chen \(2025\)Evidence\-bound autonomous research \(EviBound\): a governance framework for eliminating false claims\.Note:arXiv:2511\.05524External Links:2511\.05524Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px5.p1.2)\.
- J\. Cui, W\. Chiang, I\. Stoica, and C\. Hsieh \(2024\)OR\-Bench: an over\-refusal benchmark for large language models\.Note:arXiv:2405\.20947External Links:2405\.20947Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px4.p1.1)\.
- E\. Debenedetti, J\. Zhang, M\. Balunović, L\. Beurer\-Kellner, M\. Fischer, and F\. Tramèr \(2024\)AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents\.InAdvances in Neural Information Processing Systems \(NeurIPS\) Datasets and Benchmarks Track,External Links:2406\.13352Cited by:[§1](https://arxiv.org/html/2607.28685#S1.p1.1),[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.
- R\. O\. Kearns \(2026\)Quantifying construct validity in large language model evaluations\.Note:arXiv:2602\.15532External Links:2602\.15532Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Keller, K\. Kwegyir\-Aggrey, R\. Steed, A\. Rao, J\. Sharp, and A\. Bergman \(2026\)Expanding the AI evaluation toolbox with statistical models\.Technical reportTechnical ReportNIST Trustworthy and Responsible AI 800\-3,National Institute of Standards and Technology\.Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px2.p1.1)\.
- A\. Kipnis, K\. Voudouris, L\. M\. Schulze Buschoff, and E\. Schulz \(2025\)Metabench – a sparse benchmark of reasoning and knowledge in large language models\.InInternational Conference on Learning Representations \(ICLR\),External Links:2407\.12844Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px2.p1.1)\.
- M\. Q\. Li, B\. C\. M\. Fung, B\. Li, H\. Ismail, and F\. Iqbal \(2026\)Taxonomy and consistency analysis of safety benchmarks for AI agents\.Note:arXiv:2605\.16282External Links:2605\.16282Cited by:[§1](https://arxiv.org/html/2607.28685#S1.p1.1),[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Lynch, B\. Wright, C\. Larson, S\. J\. Ritchie, S\. Mindermann, E\. Hubinger, E\. Perez, and K\. K\. Troy \(2025\)Agentic misalignment: how LLMs could be insider threats\.Note:arXiv:2510\.05179External Links:2510\.05179Cited by:[§3\.3](https://arxiv.org/html/2607.28685#S3.SS3.p3.2)\.
- S\. G\. Patil, H\. Mao, C\. C\. Ji, F\. Yan, V\. Suresh, I\. Stoica, and J\. E\. Gonzalez \(2025\)The Berkeley function calling leaderboard \(BFCL\): from tool use to agentic evaluation of large language models\.InProceedings of the 42nd International Conference on Machine Learning \(ICML\),Cited by:[§6](https://arxiv.org/html/2607.28685#S6.SS0.SSS0.Px2.p1.6)\.
- Y\. Perlitz, A\. Gera, O\. Arviv, A\. Yehudai, E\. Bandel, E\. Shnarch, M\. Shmueli\-Scheuer, and L\. Choshen \(2024\)Do these LLM benchmarks agree? fixing benchmark evaluation with BenchBench\.Note:arXiv:2407\.13696External Links:2407\.13696Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px2.p1.1)\.
- P\. Röttger, H\. R\. Kirk, B\. Vidgen, G\. Attanasio, F\. Bianchi, and D\. Hovy \(2024\)XSTest: a test suite for identifying exaggerated safety behaviours in large language models\.InProceedings of NAACL\-HLT,External Links:2308\.01263Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px4.p1.1),[§4\.4](https://arxiv.org/html/2607.28685#S4.SS4.SSS0.Px7.p1.4)\.
- Y\. Ruan, H\. Dong, A\. Wang, S\. Pitis, Y\. Zhou, J\. Ba, Y\. Dubois, C\. J\. Maddison, and T\. Hashimoto \(2024\)Identifying the risks of LM agents with an LM\-emulated sandbox\.InInternational Conference on Learning Representations \(ICLR\),External Links:2309\.15817Cited by:[§1](https://arxiv.org/html/2607.28685#S1.p1.1),[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Souly, Q\. Lu, D\. Bowen, T\. Trinh, E\. Hsieh, S\. Pandey, P\. Abbeel, J\. Svegliato, S\. Emmons, O\. Watkins, and S\. Toyer \(2024\)A StrongREJECT for empty jailbreaks\.Note:arXiv:2402\.10260External Links:2402\.10260Cited by:[§3\.3](https://arxiv.org/html/2607.28685#S3.SS3.p3.2),[§4\.4](https://arxiv.org/html/2607.28685#S4.SS4.SSS0.Px6.p1.6)\.
- A\. Wei, N\. Haghtalab, and J\. Steinhardt \(2023\)Jailbroken: how does LLM safety training fail?\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Note:arXiv:2307\.02483External Links:2307\.02483Cited by:[§3\.3](https://arxiv.org/html/2607.28685#S3.SS3.p3.2),[§4\.4](https://arxiv.org/html/2607.28685#S4.SS4.SSS0.Px6.p1.6)\.
- S\. Yang, J\. Hu, T\. Li, H\. Yan, W\. Wang, and D\. Wang \(2026\)AutoMonitor\-Bench: evaluating the reliability of LLM\-based misbehavior monitor\.InFindings of the Association for Computational Linguistics: ACL 2026,Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhan \(2024\)τ\\tau\-Bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.Note:arXiv:2406\.12045External Links:2406\.12045Cited by:[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px5.p1.2)\.
- T\. Yuan, Z\. He, L\. Dong, Y\. Wang, R\. Zhao, T\. Xia, L\. Xu, B\. Zhou, F\. Li, Z\. Zhang, R\. Wang, and G\. Liu \(2024\)R\-Judge: benchmarking safety risk awareness for LLM agents\.InFindings of the Association for Computational Linguistics: EMNLP 2024,External Links:2401\.10019Cited by:[§1](https://arxiv.org/html/2607.28685#S1.p1.1),[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.
- Q\. Zhan, Z\. Liang, Z\. Ying, and D\. Kang \(2024\)InjecAgent: benchmarking indirect prompt injections in tool\-integrated large language model agents\.InFindings of the Association for Computational Linguistics: ACL 2024,External Links:2403\.02691Cited by:[§1](https://arxiv.org/html/2607.28685#S1.p1.1),[§2](https://arxiv.org/html/2607.28685#S2.SS0.SSS0.Px1.p1.1)\.

## Appendix ATechnical Appendix

This appendix holds method detail and the full robustness numbers summarized in the body; it is supplementary and not required to follow the main argument\.

### A\.1Ruling out two artifacts \(detail for §[4\.4](https://arxiv.org/html/2607.28685#S4.SS4)\)

#### Blinded construct coding\.

Seven LLM coders from distinct families—GPT\-4o, Claude\-3\-Haiku, Gemini\-2\.5\-Flash, Llama\-3\.3\-70B, Qwen\-2\.5\-72B, Mistral\-Large, DeepSeek\-V3—were each shown only a neutral, predictive\-information\-stripped description of what each benchmark*measures*\(no mention of our criteria, category names, or any RQ3 result\) plus two construct definitions \(“harm blocking” vs\. “risk recognition and correct treatment of benign inputs”\)\. The two options and the benchmark order were shuffled per coder to remove position bias\. Each coder assigned every benchmark to one construct\. Votes \(B=\{=\}harm blocking, R=\{=\}risk recognition\):

All four majorities match the hand\-assignment; raw agreement is27/2827/28ratings, Krippendorff’sα=0\.825\\alpha=0\.825\(nominal\), against a chance baseline of0\.54=6\.2%0\.5^\{4\}=6\.2\\%for one coder matching all four labels by coin flip\. The single dissent \(Llama\-3\.3\-70B placing InjecAgent as harm blocking\) is the one semantically borderline case \(injection\-robustness\)\. This controls coder bias but not rubric\-designer bias \(the category definitions were written by authors who had seen the results\); we therefore also report the RQ1 factor, which separates AgentHarm \(\+0\.33\+0\.33\) from the R\-Judge/InjecAgent cluster with no reference to any criterion\.

#### Blinded criterion coding, with a distractor\.

To check that the predictor→\\tocriterion matching \(not just the predictor labels\) is recoverable blind, and to make the agreement statistic meaningful, we re\-ran the coding over all six targets \(four benchmarks\+\+two criteria\) with a third, plausible*task\-success*distractor construct\. Krippendorff’sα=0\.61\\alpha=0\.61\(three categories, chance≈1/3\\approx 1/3\), and the distractor is chosen only5/425/42times\. The harm\-blocking classification is robust: AgentHarm and the jailbreak criterion both map to harm blocking7/77/7, R\-Judge to risk recognition7/77/7, and InjecAgent5/75/7, all unchanged by the distractor\. Two targets are honestly contested: AgentDojo reads as task success to4/74/7coders \(it is utility\-scored; ourn=5n\{=\}5benchmark\), and the misalignment criterion is contested: it maps to risk recognition6/76/7in a separate two\-way round \(frozen votes in the artifact\), but to harm blocking6/76/7here \(avoiding a harmful autonomous action can be read as both recognizing and declining it\)\. This recovers the AgentHarm–jailbreak pairing without using the predictive results, and is an independent reason we keep the R\-Judge/InjecAgent–misalignment pattern exploratory\.

#### XSTest over\-refusal \(blanket\-refusal check\)\.

We ran XSTest \(Röttger et al\., 2024; cited in the main paper\) on the panel via its published CPI \(compliance/partial/refusal\) scorer \(judge gpt\-4o\-mini\): the*safe*subset \(benign prompts that superficially resemble unsafe ones\) measures over\-refusal, the*unsafe*subset measures correct refusal of overtly\-harmful prompts\. The unsafe subset is ceilinged \(every model refuses8383–100%100\\%, sd4\.54\.5, limited discriminating variance\), so the informative signal is over\-refusal on the safe subset \(range0–60%60\\%; Claude\-3\-Haiku over\-refuses60%60\\%, o3\-mini30%30\\%, others≤15%\\leq 15\\%\)\. Partial\-ρ\\rho\(capability\-controlled\) two\-by\-two:

AgentHarm is unrelated to benign\-prompt over\-refusal \(ρ=−0\.01\\rho=\-0\.01,n=17n\{=\}17\), weighing against but not ruling out a blanket\-refusal effect\. Capability is\+0\.21\+0\.21\(n\.s\.\) on benign compliance and−0\.32\-0\.32\(n\.s\.\) on harmful refusal; R\-Judge specificity is\+0\.15\+0\.15on over\-refusal \(equivalently−0\.15\-0\.15on safe compliance\)\. Misalignment safety is likewise unrelated to over\-refusal \(\+0\.11\+0\.11,n=19n\{=\}19,p=0\.66p=0\.66\), while its unsafe\-refusal association is\+0\.47\+0\.47\(n=18n\{=\}18,p=0\.052p=0\.052\)\. These low\-power checks neither support nor rule out that explanation\.

### A\.2Panel coverage and power

![Refer to caption](https://arxiv.org/html/2607.28685v1/x3.png)Figure 3:The audit’s evidence base: up to 46 models×\\times9 instruments\(two capability anchors, four audited benchmarks, three held\-out criteria;262262model×\\timesinstrument cells\)\. Cells are per\-instrument min–max\-normalised scores; grey is not\-run\. The rows beyond the4141\-model anchored panel are models excluded from the analyses by pre\-specified telemetry gates \(App\.[A\.6](https://arxiv.org/html/2607.28685#A1.SS6)\)\. Rows are sorted by capability: the 2026 expansion \(top\) is scored on the anchors and the two chat\-based criteria, which need no tool\-calling harness, while the four audited benchmarks run on the original panel\. The capability anchors and the R\-Judge/InjecAgent scores vary together; the safety criteria do not track them consistently\.#### Within\-model score uncertainty\.

From the existing item\-level data \(no new runs\), per\-model standard errors for every instrument: MMLU binomial SE over its 500 fixed items \(median1\.61\.6, max2\.22\.2points\), GPQA over 198 \(median3\.33\.3, max3\.63\.6\), the misalignment criterion over its 100 decisions \(median4\.04\.0, max5\.05\.0\), R\-Judge F1 bootstrapped over its 571 records \(median1\.81\.8, max3\.03\.0\), and the jailbreak criterion bootstrapped over its∼150\{\\sim\}150judged prompts \(median2\.32\.3, max3\.73\.7\)\. Against each instrument’s between\-model SD \(8\.78\.7,18\.218\.2,26\.026\.0,14\.914\.9,28\.528\.5respectively, on the expanded panel\), the*worst\-case*within\-model SE is1313–25%25\\%of that SD\. Thus the observed panel\-level spread is44–8×8\\timesthe worst\-case SE, although individual model pairs can be closer\. \(τ2\\tau^\{2\}\-retail remains a single run whose user\-simulator variance we cannot bound; it is the one instrument this analysis cannot cover\.\)

#### Internal consistency, not just precision\.

Small SEs do not by themselves make rankings stable, so we also compute, per instrument from its item\-level data, split\-half consistency \(Spearman–Brown\), item\-bootstrap rank stability, and top\-quartile persistence \(Table[3](https://arxiv.org/html/2607.28685#A1.T3)\)\. Every item\-rich instrument with item\-resampling support is internally consistent \(split\-half≥0\.96\\geq 0\.96corrected; rank stability≥0\.97\\geq 0\.97; a top\-quartile model stays top\-quartile in0\.790\.79–0\.960\.96of item resamples\), making item\-sampling error an unlikely explanation for the panel\-level ranking disagreement in Table[1](https://arxiv.org/html/2607.28685#S4.T1)\. Misalignment has only two scenario halves \(Spearman–Brown0\.700\.70\), andτ2\\tau^\{2\}has no resampling bound\. The jailbreak computation resamples by base prompt, so the three templates’ dependence is respected\.

Table 3:Internal consistency from available item\- or scenario\-level data: mean split\-half correlation of model scores over 50 random half\-splits \(Spearman–Brown corrected\), mean Spearman of item\-bootstrapped vs\. observed rankings, and the probability that an observed top\-quartile model stays top\-quartile under resampling\. Item counts are the complete\-coverage intersection across all listed models \(e\.g\. 356 of the 500 administered MMLU items, 164 of 198 GPQA items\)\. Splits/draws are shared across models and use the items common to every model’s logs; the jailbreak criterion resamples by base prompt \(the three templates are dependent\), and its shared set is 20 of the 50 suite prompts because each model leaves a few different prompts unjudged\. Jailbreak uses 42 models: the 41 anchored models plus DeepSeek\-R1, which lacks the required anchor pair; misalignment has config\-level data only, so its row is the blackmail\-vs\-leaking halves convergence\. Every item\-rich instrument with item\-resampling support has Spearman–Brown≥0\.96\\geq 0\.96and rank stability≥0\.97\\geq 0\.97, making item\-sampling error an unlikely explanation for the panel\-level disagreement of Table[1](https://arxiv.org/html/2607.28685#S4.T1)\. Misalignment’s two scenario halves agree only moderately \(Spearman–Brown0\.700\.70\)\.τ2\\tau^\{2\}is a single run and remains the one unbounded instrument\.The exploratory R\-Judge/InjecAgent–misalignment pattern is estimated atn=18n\{=\}18–2020with power≈0\.35\\approx 0\.35\(simulated\) to detectρ=0\.4\\rho\{=\}0\.4atα=0\.05\\alpha\{=\}0\.05; the minimum detectable effect \(80%80\\%power\) isρ=0\.62\\rho\{=\}0\.62atn=18n\{=\}18, above the observed\+0\.41\+0\.41\. Thus the non\-significant result cannot exclude an effect of practical interest, and we retain it only as a hypothesis\. Detecting an effect of\+0\.41\+0\.41with80%80\\%power would requiren≈40n\{\\approx\}40\. Under a Bonferroni correction over the nine RQ3 partial tests \(α=0\.006\\alpha\{=\}0\.006\), the two firm results \(the capability crossover and AgentHarm→\\tojailbreak, bothp<0\.001p<0\.001\) survive, while the R\-Judge/InjecAgent cells \(p=0\.05p\{=\}0\.05–0\.090\.09\) do not\.

![Refer to caption](https://arxiv.org/html/2607.28685v1/x4.png)Figure 4:Models with similar measured capability can differ widely in safety across model\-developing organizations\.Each point is a model; colour and shape identify its organization; error bars are within\-model SEs \(§[A\.2](https://arxiv.org/html/2607.28685#A1.SS2)\)\.*Within*an organization \(lines, the six organizations with≥4\\geq 4anchored models\) the capability→\\tosafety slope is large but sign\-heterogeneous across organizations \(nn\-weighted mean−0\.14\-0\.14misalignment,\+0\.26\+0\.26jailbreak\)\. Among higher\-capability models \(shaded, composite\>0\.4\>0\.4\), safety spans nearly the full range \(arrows\)\. After rank\-residualising safety on capability, organization groups differ under permutation tests \(df\-adjusted Kruskal–Wallisε2=0\.34\\varepsilon^\{2\}\{=\}0\.34/0\.570\.57,p≤0\.004p\\leq 0\.004\)\. The static composite does not capture all capability, so this is an association, not evidence that organizational identity causes the difference\.

### A\.3Cross\-benchmark structure \(detail for §[4\.3](https://arxiv.org/html/2607.28685#S4.SS3)\)

Figures[5](https://arxiv.org/html/2607.28685#A1.F5)and[6](https://arxiv.org/html/2607.28685#A1.F6)visualise the two RQ1 results: the dissolution of the small\-panel R\-Judge–AgentHarm trade\-off, and the single\-factor structure that provides a weaker description of how the scores covary\.

![Refer to caption](https://arxiv.org/html/2607.28685v1/x5.png)Figure 5:R\-Judge and AgentHarm do not rank the full panel in opposite order\.AgentHarm safety against R\-Judge specificity on then=18n\{=\}18cross\-benchmark panel, coloured by the harmonized GPQA anchor\. The trade\-off a 7\-model panel suggested \(ρ=−0\.64\\rho=\-0\.64\) dissolves atn=18n\{=\}18\(ρ=\+0\.02\\rho=\+0\.02,p=0\.95p=0\.95; grey trend\): models near the top of either axis span the full range of the other, the pairwise view of Table[1](https://arxiv.org/html/2607.28685#S4.T1)’s rank disagreement\.![Refer to caption](https://arxiv.org/html/2607.28685v1/x6.png)Figure 6:The first principal component groups R\-Judge, InjecAgent, and capability, with AgentHarm in the opposite direction\.PC1 loadings on then=18n\{=\}18panel with harmonized anchors \(48\.7%48\.7\\%of variance\)\. AgentHarm has the opposite sign from the other four measures in1717of1818leave\-one\-model\-out fits\. The sign of a principal component is arbitrary; only the relative directions matter\. We treat this structure as suggestive because the median between\-benchmark correlation is only0\.220\.22\(§[4\.3](https://arxiv.org/html/2607.28685#S4.SS3)\)\.#### Correlation\-convention sensitivity\.

The PCA above operates on the Pearson correlation matrix of the standardized scores\. A rank\-based \(Spearman\) PCA keeps the loading structure \(anchors−0\.72\-0\.72/−0\.78\-0\.78, AgentHarm opposite at\+0\.38\+0\.38\) and the pre\-specified positive control still passes, but PC1 falls from48\.7%48\.7\\%to40\.8%40\.8\\%of variance and Horn’s parallel analysis retains no factor\. The pre\-specified criterion for interpreting PC1 \(first\-factor variance≥40%\\geq 40\\%\) holds under both conventions, and the pre\-specified negative control passes as executed:200200column\-permuted matrices give median first eigenvalue1\.711\.71against the parallel\-analysis9595th\-percentile threshold of2\.182\.18\(exceedance2%2\\%\)\. The convention sensitivity is one more reason we treat the factor as supporting rather than primary evidence: the firm RQ1 result is the correlation\-free ranking disagreement of Table[1](https://arxiv.org/html/2607.28685#S4.T1)\.

### A\.4R\-Judge metric sensitivity \(detail for §[4\.3](https://arxiv.org/html/2607.28685#S4.SS3)–§[4\.4](https://arxiv.org/html/2607.28685#S4.SS4)\)

Because the paper rejects R\-Judge’s headline F1 for ignoring true negatives yet represents R\-Judge by specificity \(which ignores true positives\) in the construct and criterion analyses, we recompute every R\-Judge\-involving result under four metric choices on the paper’s own panels \(Table[4](https://arxiv.org/html/2607.28685#A1.T4)\)\. The conclusions the paper leans on are metric\-robust; the one metric\-sensitive cell is the R\-Judge–misalignment partial, which the paper already treats as exploratory\.

Table 4:R\-Judge metric sensitivity under specificity, recall, balanced accuracy, and F1\. RQ3 entries are partial Spearman correlations controlling the capability composite \(τ2\\tau^\{2\}n=19n\{=\}19, misalignmentn=18n\{=\}18, jailbreakn=20n\{=\}20\)\. The RQ1 small\-panel reversal does not return under any metric, and the task\-success and jailbreak nulls are stable\. The exploratory misalignment estimate is not stable:\+0\.41\+0\.41under specificity and−0\.46\-0\.46under recall \(both n\.s\.\)\.
### A\.5Robustness of the capability crossover \(detail for §[4\.4](https://arxiv.org/html/2607.28685#S4.SS4)\)

#### Partial\-correlation convention\.

Partials residualize midranks on the harmonized capability composite and re\-rank the residuals\. Under the classic alternative \(Pearson on the rank residuals\) the two exploratory misalignment estimates move to\+0\.49\+0\.49\(R\-Judge\) and\+0\.41\+0\.41\(InjecAgent\), within0\.080\.08of Table[2](https://arxiv.org/html/2607.28685#S4.T2), and the AgentHarm–jailbreak estimate stays\+0\.72\+0\.72: no conclusion changes under either convention\.

![Refer to caption](https://arxiv.org/html/2607.28685v1/x7.png)Figure 7:Against task success, capability is the predictor and AgentHarm adds nothing\.τ2\\tau^\{2\}\-retail success against the harmonized capability composite \(ρ=\+0\.60\\rho=\+0\.60,p=0\.005p=0\.005,n=20n\{=\}20; grey trend\)\. Colour is AgentHarm safety \(open circle: no AgentHarm score\): safe and unsafe models sit on both sides of the trend with no safety gradient after controlling capability, the pre\-specified result that AgentHarm adds no incremental validity in Table[2](https://arxiv.org/html/2607.28685#S4.T2)\.InteractionΔ=−1\.00\\Delta\{=\}\{\-\}1\.00\(95% CI\[−1\.48,−0\.49\]\[\-1\.48,\-0\.49\], bootstrapp<0\.001p<0\.001\)\. Leave\-one\-outΔ∈\[−1\.16,−0\.90\]\\Delta\\in\[\-1\.16,\-0\.90\]\. Leave\-one\-organization\-out, per omitted organization: Amazon−0\.96\-0\.96, Anthropic−1\.01\-1\.01, Cohere−0\.90\-0\.90, DeepSeek−1\.13\-1\.13, Google−0\.98\-0\.98, Meta−0\.96\-0\.96, Mistral−0\.93\-0\.93, OpenAI−1\.16\-1\.16, Qwen−0\.95\-0\.95\(the range coincides with leave\-one\-out because Cohere contributes one model and OpenAI, the largest group, gives the widest swing\)\. Organization\-clustered bootstrap:Δ∈\[−1\.49,−0\.59\]\\Delta\\in\[\-1\.49,\-0\.59\],P​\(Δ≥0\)<10−4P\(\\Delta\{\\geq\}0\)<10^\{\-4\}\. Subsampling stability \(draws of each size\): size\-10 median−0\.95\-0\.95,100%100\\%negative,97%≤−0\.597\\%\\leq\-0\.5,0%0\\%sign\-flips; size\-12−0\.96\-0\.96/100%100\\%/99%99\\%/0%0\\%; size\-14−0\.97\-0\.97/100%100\\%/100%100\\%/0%0\\%\(contrast: the RQ1 R\-Judge–AgentHarm correlation reaches\|ρ\|≥0\.5\|\\rho\|\{\\geq\}0\.5in∼25%\\sim 25\\%of size\-7 draws\)\. Independent\-grader checks: the misalignment classifier reproduces atρ=0\.97\\rho=0\.97\(Claude\-3\-Haiku vs\. GPT\-4o\), the jailbreak judge atρ=0\.98\\rho=0\.98\(Gemini vs\. gpt\-4o\), and AgentHarm’s refusal predictor atρ=0\.69\\rho=0\.69\(a Gemini refusal judge re\-scoring the cached completions,n=19n\{=\}19; the AgentHarm–jailbreak estimate remains similar,\+0\.66\+0\.66vs\+0\.72\+0\.72; the moderate agreement reflects that the swap redoes only the refusal component, not the full grading pipeline\)\. Per\-jailbreak AgentHarm→\\tojailbreak partials \(within the three\-template suite\): AIM\+0\.44\+0\.44, refusal\-suppression\+0\.70\+0\.70, prefix\-injection\+0\.62\+0\.62; the original*single*\-AIM criterion gave\+0\.50\+0\.50\.

#### Misalignment 3\-scenario battery \(exploratory\)\.

Adding a third scenario \(*murder*; our classifier, 22 models\) to the blackmail\+\+leaking criterion lifts the R\-Judge/InjecAgent–misalignment interaction fromΔ=\+0\.52\\Delta=\+0\.52\(p=0\.068p=0\.068\) toΔ=\+0\.53\\Delta=\+0\.53\(p=0\.048p=0\.048\)\. The point estimate is essentially unchanged, so the significance comes from the tighter CI of a larger battery rather than from a stronger effect; a post\-hoc extension crossing the threshold is a forking path, and we do not lean on it\. Murder is also a severe harm that harm\-compliance forecasts \(AgentHarm→\\tobattery\+0\.43\+0\.43\), diluting construct specificity, so we keep blackmail\+\+leaking primary\.

### A\.6The 2026 model\-panel expansion \(n≈\\approx41\) and organization\-group differences \(detail for §[4\.4](https://arxiv.org/html/2607.28685#S4.SS4)\)

Before evaluating the new models on either criterion, we specified a targeted expansion predicting that the crossover would “hold or strengthen,” both to test whether the capability crossover holds on current models and to probe whether any panel change follows capability or organization grouping\. We grew the panel to4141anchored models spanning1212model\-developing organizations, using the same chat/scenario harnesses as the original panel\. The two criteria require no tool loop; new models simply skip the tool\-calling audit benchmarks\. Exclusions were based only on harness behavior, never on outcome scores: Kimi\-K2\.6 and both Nemotron\-3 models returned usable anchor answers on fewer than50%50\\%of items or produced unbounded reasoning traces; MiniMax\-M3 fell below the misalignment engagement threshold because of truncated outputs, not low harm, but was retained on the single\-turn jailbreak criterion, so the misalignment analysis usesn=40n\{=\}40whereas the jailbreak analysis usesn=41n\{=\}41\. Gemini\-2\.5\-Pro yielded no valid anchor; GLM\-4\.5 yielded MMLU but not GPQA\. Both encountered routing incompatibilities, so only GLM\-4\.5 appears as a partial row in Fig\.[3](https://arxiv.org/html/2607.28685#A1.F3)\. These exclusions are a small instance of evaluation harnesses going stale\.

#### Primary expansion result: the capability–misalignment correlation weakens\.

Capability→\\tomisalignment\-safety moves from−0\.44\-0\.44\(n=21n\{=\}21, asymptoticp=0\.047p=0\.047; bootstrapp=0\.11p=0\.11,95%95\\%CI\[−0\.85,\+0\.09\]\[\-0\.85,\+0\.09\]\) to−0\.16\-0\.16\(n=40n\{=\}40, bootstrapp=0\.40p=0\.40;95%95\\%CI\[−0\.54,\+0\.22\]\[\-0\.54,\+0\.22\], leave\-one\-out\[−0\.25,−0\.12\]\[\-0\.25,\-0\.12\]\) and is no longer significant\. The pre\-specified confirmation threshold \(ρ≤−0\.35\\rho\\leq\-0\.35atp<0\.05p<0\.05\) is not met, so the original negative correlation is not reproduced\. However, the CI still includes−0\.44\-0\.44, so we report a weaker estimate rather than evidence that the correlation has disappeared\. Three direct tests of the change itself \(accounting for the original panel being nested in the expanded one\): a stratified nested bootstrap givesΔ​ρ=\+0\.27\\Delta\\rho=\+0\.27,95%95\\%CI\[−0\.25,\+0\.78\]\[\-0\.25,\+0\.78\], two\-sidedp=0\.29p=0\.29\(jailbreak\+0\.26\+0\.26, CI\[−0\.14,\+0\.69\]\[\-0\.14,\+0\.69\],p=0\.22p=0\.22\); the 2026\-only models alone giveρ=\+0\.07\\rho=\+0\.07versus the original−0\.44\-0\.44\(independent Fisherzz,p=0\.12p=0\.12\); and size\-21 subsets of the expanded pool matched to the original panel’s organization histogram give medianρ=−0\.17\\rho=\-0\.17\(IQR\[−0\.27,−0\.08\]\[\-0\.27,\-0\.08\]\), i\.e\. organization composition alone does not reproduce the original value\. So organization mix alone does not explain the movement; generational change is compatible with the result but is not established, and the change itself is not statistically significant\. Our bounded conclusion is that the estimate cannot be assumed to generalize across model populations and that the “hold or strengthen” prediction was not met\. A ceiling does not explain all remaining variation: among models with similar measured capability, safety varies widely across organizations \(composite≈0\.8\{\\approx\}0\.8: o3\-mini0\.990\.99, Grok\-4\.30\.700\.70, DeepSeek\-V30\.190\.19\)\. Becauseτ2\\tau^\{2\}task success was not re\-collected, we update this correlation, not the interactionΔ\\Delta\.

#### Organization\-group analysis\.

With≥4\\geq 4anchored models each from six organizations we can begin to separate the collinear capability and organization factors\. Within\-organization capability→\\tosafety slopes are sign\-heterogeneous and near\-zero in aggregate; across organizations, organization groups differ after rank\-residualizing safety on capability on both criteria \(permutationp=0\.004p\{=\}0\.004misalignment,<0\.001<0\.001jailbreak; df\-adjusted Kruskal–Wallisε2=0\.34\\varepsilon^\{2\}\{=\}0\.34and0\.570\.57; Table[5](https://arxiv.org/html/2607.28685#A1.T5), Fig\.[4](https://arxiv.org/html/2607.28685#A1.F4)\)\. All five secondary tests survive Holm \(max adjustedp=0\.043p=0\.043\)\. As in the body, we stop short of calling this a clean identification: capability is matched only on a static\-knowledge composite \(so unmeasured agentic capability can still appear as an organization difference\), within\-organizationnnis44–88, and some organizations are range\-restricted \(Anthropic ceilings near safety1\.01\.0\)\. The between\-organization permutation test, not the within\-organization slopes, carries the claim\. Because that test permutes labels over only seven organization clusters \(within\-organizationnn44–88\), we checked its Type\-I calibration on2,0002\{,\}000synthetic panels per null with*no between\-organization mean difference*, holding models, organizations, and capability fixed: with the real safety values reshuffled across models, and with i\.i\.d\. Gaussian safety, rejection matches nominalα\\alphaat0\.05/0\.01/0\.0050\.05/0\.01/0\.005on both criteria \(0\.0450\.045–0\.0520\.052atα=0\.05\\alpha\{=\}0\.05\)\. A worst\-case null with organization\-clustered variances but no mean difference inflates rejection mildly on the jailbreak structure only \(0\.0670\.067atα=0\.05\\alpha\{=\}0\.05; no detectable inflation atα=0\.005\\alpha\{=\}0\.005\) — too small to change any conclusion: even doubling bothpp\-values leaves the organization\-group tests Holm\-significant\. Dropping near\-ceiling Anthropic gives misalignmentε2=0\.27\\varepsilon^\{2\}=0\.27,p=0\.018p=0\.018and jailbreakε2=0\.54\\varepsilon^\{2\}=0\.54,p<0\.001p<0\.001; neither group result is carried solely by Anthropic\.

Table 5:Organization\-group differences on then≈41n\{\\approx\}41expansion\. Holding the organization fixed, capability’s slope on safety is heterogeneous and near\-zero \(top\)\. After rank\-residualizing safety on the static capability composite, organization groups differ under permutation tests on both criteria \(bottom\);ε2\\varepsilon^\{2\}is the df\-adjusted Kruskal–Wallis effect size, not a causal share\.
#### Expanded\-panel jailbreak result\.

We rebuilt the jailbreak criterion on the expanded panel \(three published templates re\-scored for every new model\)\. Capability→\\tojailbreak\-safety rises from\+0\.08\+0\.08\(n=21n\{=\}21\) to\+0\.34\+0\.34\(n=41n\{=\}41; bootstrapp=0\.035p=0\.035,95%95\\%CI\[\+0\.03,\+0\.60\]\[\+0\.03,\+0\.60\], leave\-one\-out\[\+0\.30,\+0\.39\]\[\+0\.30,\+0\.39\]\), although the panel\-to\-panel change is not significant\. Within\-organization slopes remain heterogeneous \(weighted mean\+0\.26\+0\.26\), while models with similar measured capability vary across organizations \(permutationp<0\.001p<0\.001;ε2=0\.57\\varepsilon^\{2\}\{=\}0\.57\)\. Among high\-capability models, jailbreak\-safety spans99–9999\(Mistral\-Large0\.090\.09, DeepSeek\-V30\.370\.37, Anthropic/OpenAI0\.980\.98–0\.990\.99\)\. Together with the misalignment result, this shows that neither the sign nor the magnitude of capability→\\tosafety is stable across criteria or panels\.

An earlier pre\-specified R\-Judge–misalignment continuation fell from\+0\.41\+0\.41\(n=18n\{=\}18\) to\+0\.23\+0\.23\(n=25n\{=\}25\), below its0\.300\.30threshold; we stopped it and retain it as a hypothesis\.

Similar Articles

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

arXiv cs.AI

This paper argues that current benchmarks for autonomous agents fail to evaluate whether an agent should have proceeded at all, introducing a 'compliance bias'. The authors propose a taxonomy of abstention-warranted scenarios and new evaluation protocols (Safety Rate, Usability Rate, Informed Refusal Rate) with preliminary results showing tunable safety–usability tradeoffs across model families.

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Hugging Face Daily Papers

The paper introduces PhoneSafety, a benchmark of 700 safety-critical moments across 130+ apps to evaluate phone-use agents. Results show that avoiding harmful outcomes does not necessarily indicate safety, as models may fail to act or make unsafe choices, requiring a distinction between capability and safety signals.