Auditing the Audit: Five Failure Modes in Benchmark-Validity Audits

arXiv cs.LG Papers

Summary

This paper identifies five failure modes in perturbation-based benchmark-validity audits used for AI governance, demonstrating that implementation details can silently manufacture conclusions. It proposes a due-diligence gate to improve the reliability of evaluation evidence.

arXiv:2607.02586v1 Announce Type: new Abstract: Governance frameworks ask AI providers and auditors for documented evaluation evidence, and perturbation-based construct-validity audits are a common form of that evidence. We argue the audits are themselves fragile: their conclusions can be silently manufactured by implementation details that readers cannot see in the reported numbers. We name five classes of pipeline failure and demonstrate each in a self-audit over safety benchmarks and open-weight instruction-tuned models. Under a unified six-point due-diligence gate, every cell lands in a non-confirmatory bucket, and no cell reaches confirmatory. The evidence here is a single two-model, five-benchmark case study, and F1--F5 is an illustrative, deliberately non-exhaustive starting taxonomy -- not a comprehensive partition of audit failures. We position the gate as a withholding and disclosure protocol for assurance-grade evidence, supplementary to (not a replacement for) classical construct-validity evidence, and not as a route to benchmark-validity verdicts.
Original Article
View Cached Full Text

Cached at: 07/07/26, 04:38 AM

# Five Failure Modes in Benchmark-Validity Audits
Source: [https://arxiv.org/html/2607.02586](https://arxiv.org/html/2607.02586)
## Auditing the Audit: Five Failure Modes in Benchmark\-Validity Audits

###### Abstract

Governance frameworks ask AI providers and auditors for*documented evaluation evidence*, and perturbation\-based construct\-validity audits are a common form of that evidence\. We argue the audits are themselves fragile: their conclusions can be silently manufactured by implementation details that readers cannot see in the reported numbers\. We name five classes of pipeline failure and demonstrate each in a self\-audit over safety benchmarks and open\-weight instruction\-tuned models\. Under a unified six\-point due\-diligence gate, every cell lands in a non\-confirmatory bucket, and no cell reaches*confirmatory*\. The evidence here is a single two\-model, five\-benchmark case study, and F1–F5 is an illustrative, deliberately non\-exhaustive starting taxonomy—not a comprehensive partition of audit failures\. We position the gate as a withholding and disclosure protocol for assurance\-grade evidence, supplementary to \(not a replacement for\) classical construct\-validity evidence, and not as a route to benchmark\-validity verdicts\.

benchmark validity, AI governance, evaluation evidence, construct validity, perturbation audits, reproducibility

## 1Introduction

The EU AI Act \(Regulation \(EU\) 2024/1689, Art\. 55\) is binding law\(European Parliament and Council,[2024](https://arxiv.org/html/2607.02586#bib.bib16)\); the NIST AI Risk Management Framework\(National Institute of Standards and Technology,[2023](https://arxiv.org/html/2607.02586#bib.bib17)\)is voluntary risk\-management guidance; China’s GB/T 45654\(State Administration for Market Regulation and Standardization Administration of China,[2025](https://arxiv.org/html/2607.02586#bib.bib18)\)is a national standard\. None mandates a specific academic benchmark; what each calls for is*documented evaluation evidence*—reproducible records of what was measured, how, and with what uncertainty\. Benchmarks sit*inside*that evidence pipeline\(Reuelet al\.,[2025](https://arxiv.org/html/2607.02586#bib.bib19)\), increasingly accompanied by*perturbation\-based validity audits*\. We use the term in a specific sense: a procedure that applies controlled, semantics\-preserving or construct\-flipping edits to benchmark items and asks whether the headline metric moves only when it should— moving under a construct flip, staying stable under a surface rewrite\(Sclaret al\.,[2024](https://arxiv.org/html/2607.02586#bib.bib6); Mizrahiet al\.,[2024](https://arxiv.org/html/2607.02586#bib.bib7); Alzahraniet al\.,[2024](https://arxiv.org/html/2607.02586#bib.bib8); Ribeiroet al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib20)\)\. This is one of several construct\-validity checks, not the only one: convergent / discriminant / criterion correlation studies\(Cronbach and Meehl,[1955](https://arxiv.org/html/2607.02586#bib.bib9); Messick,[1995](https://arxiv.org/html/2607.02586#bib.bib10); Jacobs and Wallach,[2021](https://arxiv.org/html/2607.02586#bib.bib22); Blodgettet al\.,[2021](https://arxiv.org/html/2607.02586#bib.bib14)\), benchmark\-construct critiques\(Rajiet al\.,[2021](https://arxiv.org/html/2607.02586#bib.bib13)\), item\-response modelling\(Laloret al\.,[2019](https://arxiv.org/html/2607.02586#bib.bib11); Vaniaet al\.,[2021](https://arxiv.org/html/2607.02586#bib.bib12)\), behavioural test suites\(Ribeiroet al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib20)\)and dynamic adversarial benchmarking\(Kielaet al\.,[2021](https://arxiv.org/html/2607.02586#bib.bib21)\), and human\-label agreement studies are alternatives that probe different facets of validity\. We study the perturbation family because it is the form most directly emitted as governance evidence and the cheapest to run without new annotation; our findings about pipeline fragility are specific to it and we do not claim they transfer to the other families unchanged\. Recent adjacent audit work similarly treats benchmark conclusions as measurement claims constrained by detectable effects, configuration choices, probe choices, and domain\-validity assumptions\(Zhuanget al\.,[2026](https://arxiv.org/html/2607.02586#bib.bib34); Liet al\.,[2026b](https://arxiv.org/html/2607.02586#bib.bib35),[a](https://arxiv.org/html/2607.02586#bib.bib36); Wanget al\.,[2026b](https://arxiv.org/html/2607.02586#bib.bib37)\)\. We keep a small number of broader adjacent examples only where they sharpen this evidence\-pipeline framing: RAG reliability work separates retrieved or relevant evidence from warranted evidence and makes retrieval–reasoning or query\-difficulty choices explicit\(Chenet al\.,[2026](https://arxiv.org/html/2607.02586#bib.bib38); Qianet al\.,[2026](https://arxiv.org/html/2607.02586#bib.bib39); Jiet al\.,[2026](https://arxiv.org/html/2607.02586#bib.bib24),[2025](https://arxiv.org/html/2607.02586#bib.bib25)\); lightweight explanation work reminds us that diagnostic signals also need faithfulness checks\(Lanet al\.,[2025](https://arxiv.org/html/2607.02586#bib.bib26)\); text\-to\-image evaluation work makes the target construct explicit: social\-bias benchmarks use multi\-dimensional bias evidence, while prompting\-proficiency benchmarks test a different user\-facing capability\(Luoet al\.,[2024b](https://arxiv.org/html/2607.02586#bib.bib27),[2026b](https://arxiv.org/html/2607.02586#bib.bib28),[a](https://arxiv.org/html/2607.02586#bib.bib29),[2026a](https://arxiv.org/html/2607.02586#bib.bib32)\); and agent and agentic\-coding evaluation work treats safety, compositional skill risk, human\-in\-the\-loop coding value, and supply\-chain runtime threats as distinct pipeline\-level target constructs\(Luoet al\.,[2025](https://arxiv.org/html/2607.02586#bib.bib30); Wanget al\.,[2026a](https://arxiv.org/html/2607.02586#bib.bib40); Luoet al\.,[2026c](https://arxiv.org/html/2607.02586#bib.bib31); Jianget al\.,[2026](https://arxiv.org/html/2607.02586#bib.bib33)\)\. These adjacent papers motivate the surrounding evidence framing, not the F1–F5 taxonomy itself\.

This paper is about a problem one level up\. Before a perturbation audit can speak to benchmark validity, it must itself be a trustworthy measurement pipeline\. Our central claim is that*perturbation\-based benchmark\-audit pipelines are themselves fragile measurement systems*: their conclusions can be silently manufactured by pipeline\-level bugs that a conformity\-assessment body, a procurement evaluator, or an internal\-assurance team cannot reliably detect from the reported numbers alone\. Our contribution is therefore*supplementary*to classical validation, not a replacement for it: convergent, discriminant, and criterion evidence still establish whether a benchmark measures its construct; we ask the prior question of whether the pipeline that would generate such evidence is itself trustworthy enough for that evidence to be believed\.

We treat the audit stage as distinct from the benchmarking stage deliberately\. The same two hazard families—software\-assurance defects and measurement\-faithfulness defects—can corrupt a benchmark when it is*built*\. But a validity audit is a second\-order measurement: it is precisely the instrument a governance reader trusts to catch a benchmark’s first\-order problems\. A defect there is silent in a different and more dangerous way—it does not corrupt the benchmark, it disables the safeguard, and it leaves no trace in the audit’s reported numbers because the reader is looking at those numbers*to*decide whether to trust the benchmark\. That asymmetry, not a claim that the two stages have disjoint failure sets, is why we scope to the audit stage\.

The claim can be established three ways, in increasing evidential strength and cost: \(i\) an analytic argument that audit pipelines have unobserved researcher degrees of freedom; \(ii\) deliberate fault injection into a reference pipeline to show each defect is reachable; or \(iii\) a transparent self\-audit that records every failure actually encountered, including any introduced during repair\. Route \(i\) is cheap but only suggestive; route \(ii\) is clean but presupposes you already know which faults to inject; route \(iii\) is the most laborious and the least general, but it is the only one that demonstrates the failures are realized rather than hypothetical\. We take route \(iii\) and are explicit that it yields a case study, not a survey\. Auditing five widely used safety benchmarks \(TruthfulQA\(Linet al\.,[2022](https://arxiv.org/html/2607.02586#bib.bib1)\), BBQ\(Parrishet al\.,[2022](https://arxiv.org/html/2607.02586#bib.bib2)\), ToxiGen\(Hartvigsenet al\.,[2022](https://arxiv.org/html/2607.02586#bib.bib3)\), CrowS\-Pairs\(Nangiaet al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib4)\)scored via PLL followingSalazaret al\.\([2020](https://arxiv.org/html/2607.02586#bib.bib15)\), and XSTest\(Röttgeret al\.,[2024](https://arxiv.org/html/2607.02586#bib.bib5)\)\) against two open\-weight 7B instruction\-tuned models \(Qwen\-2\.5\-7B\-Instruct, Mistral\-7B\-Instruct\-v0\.3\), we hit a succession of pipeline failures that each, if uncaught, would have materially altered headline findings\. Crucially,*one of those failures was introduced during the repair phase, not inherited from the initial pipeline*; we caught it only because we inspected per\-cell meta log\-probabilities after a rerun\.

### Contributions\.

\(1\) A five\-class taxonomy F1–F5 of pipeline\-level failures, split between an assurance layer \(F1, F2, F4\) and a faithfulness layer \(F3, F5\), with F3 further split into inverted\-convention \(F3a\), harness ordering \(F3b\), and scorer\-truncation \(F3c\) subtypes \(§[2](https://arxiv.org/html/2607.02586#S2)\)\. F1–F5 is a deliberately non\-exhaustive subset of the space of pipeline\-level failures that are*\(i\)*invisible in the reported numbers and*\(ii\)*realizable in perturbation\-audit code; it is not a complete partition of that space, and the selection criteria and excluded candidates are stated in §[2](https://arxiv.org/html/2607.02586#S2)\. \(2\) A case\-study realization of each class in our own pipeline \(§[4](https://arxiv.org/html/2607.02586#S4), Table[2](https://arxiv.org/html/2607.02586#S4.T2)\)\. \(3\) A unified six\-point due\-diligence gate G1–G6 with hierarchical verdict precedence \(§[3](https://arxiv.org/html/2607.02586#S3), Table[1](https://arxiv.org/html/2607.02586#S3.T1)\); the intended output is a withholding rule, not a scalar metric or leaderboard\. \(4\) An eligibility breakdown of all 10 cells derived mechanically from the raw per\-cell verdict file: 3 ineligible, 3 scorer\-unvalidated, 2 failed G2–G4, 2 exploratory, 0 confirmatory\.

### Scope\.

This is a single case study: two models, five benchmarks, 200 items, single seed, single template, declared exploratory after the F3c discovery\. The taxonomy, the gate, and the four\-way breakdown are illustrative artefacts whose generality across other models, benchmarks, and audit codebases is*untested*and is not claimed\. We do not claim any benchmark is construct\-\(in\)valid, that F1–F5 is complete, nor that any governance framework cites these benchmarks by name\. Detailed limitations and deviations from preregistration are in Appendix[A](https://arxiv.org/html/2607.02586#A1)and[K](https://arxiv.org/html/2607.02586#A11)\.

## 2The Five Pipeline Failure Modes

Failure modesSix\-point due\-diligence gateVerdict precedenceF1:Silent no\-op perturbationsF2:Regex\-extraction artefactsF4:Broken bootstrap pairingF3:Non\-faithful scoringF3a inverted conventionF3b harness orderingF3c scorer truncationF5:Metric\-archetype mismatchSoftware\-assurance layerMeasurement\-faithfulness layerG1:Scorer\-faithful auditno\-op rate \+ parseability \(partial\)G2:Above\-baseline originalimplemented numerical gateG3:Non\-trivial denominatorimplemented numerical gateG4:Paired uncertaintypaired item bootstrapG5:Archetype disclosurediagnostic / invariance / mixedG6:Repair\-regression checkinspection \+ targeted unit testPerturbation\-auditedbenchmark resultsreference scorerF3c repairEligiblesetup?Scorervalidated?Pass G2–G4numerical gates?G5–G6implementedand passed?1\. Ineligible2\. Scorer\-unvalidated3\. Failed G2–G44\. ExploratoryG2–G4 passed; G5/G6 proposed5\. Confirmatoryselective or non\-selectiveNoYesNoYesNoYesNoYes

Figure 1:Overview of our benchmark\-audit pipeline; a roadmap figure—the gate definitions and the status\-assignment algorithm are formalised in §[3](https://arxiv.org/html/2607.02586#S3)\(Table[1](https://arxiv.org/html/2607.02586#S3.T1), §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\)\. Five failure modes F1–F5 partition into a software\-assurance layer \(F1 silent no\-ops, F2 regex\-extraction artefacts, F4 broken bootstrap pairing\) and a measurement\-faithfulness layer \(F3 non\-faithful scoring with subtypes F3a–F3c, F5 metric\-archetype mismatch\)\. Perturbation\-audited benchmark results flow through the six\-point due\-diligence gate G1–G6 and the §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)algorithm; each cell receives exactly one status from the first\-firing condition:*ineligible \(F1\-type\)*,*failed G2–G4*,*scorer\-unvalidated*,*ineligible \(F5\-type\)*,*exploratory*, or*confirmatory\-\{selective\|\|non\-selective\}*\.### Why these five, and a subset of what\.

F1–F5 is not an attempt at a complete taxonomy of audit failures\. It is the subset of pipeline\-level failures that satisfies three selection criteria, in order of priority: a failure is in scope only if it is*\(C1\) silent*—invisible in the reported numbers, so a governance reader cannot detect it from the evidence artefact alone;*\(C2\) realized*—we actually encountered it in this audit, rather than hypothesised it; and*\(C3\) gateable*—it maps to a concrete, checkable disclosure that an evidence consumer could demand\. C1 is the load\-bearing criterion: a loud failure \(a crash, an obviously wrong number\) is a software\-quality problem, not an evidence\-trust problem, and is out of scope here\. The ordering also says what we deprioritised first: failures that are loud, that we did not hit, or for which we could not name an actionable gate\. A fuller taxonomy would add at least F6 perturbation non\-isolation, F7 template / system\-prompt confounding, F8 tokenization / label\-realization errors, F9 index / filter misalignment, and F10 context\-window contamination; we exclude these from the demonstrated set because we did not realize them in this case study \(they fail C2\), not because they are unimportant\. Calibrating which subset matters most across codebases is exactly the further development this v1 taxonomy needs\.

### How the five relate, and how they map to the gate\.

We organize the five \(Figure[1](https://arxiv.org/html/2607.02586#S2.F1)\) into two layers by*who can catch them*, not by claiming the two layers are exhaustive or disjoint\. The*software\-assurance layer*\(F1, F2, F4\) contains failures a competent engineering review should catch without knowing what the benchmark is for; the*measurement\-faithfulness layer*\(F3, F5\) contains failures that cannot be caught without reasoning about the benchmark’s intended construct\. The classes are*overlapping hazards*, not mutually exclusive: a single cell can exhibit several at once \(in our ToxiGen case a silent\-no\-op F1 and a top\-kktruncation F3c coexist on the same repaired pipeline\)\. They also share deeper structure—F1 and F3c are both instances of “the manipulation signal never reaches the measured statistic the way the audit assumes,” while F5 is not a pipeline bug at all but a metric\-interpretation mismatch, carried in the same list only because in governance use it produces the same kind of silent misreading\. Each class is paired with the gate check designed to neutralize it, which is how the taxonomy connects to the due\-diligence gate of §[3](https://arxiv.org/html/2607.02586#S3): F1 and F2 are the two hazards G1 exists to catch \(silent\-no\-op rate and parseability, respectively\); F3a/F3b are caught by mandating the benchmark’s reference scorer and surfaced through the*scorer\-unvalidated*status; F3c—a repair\-introduced regression—is exactly what G6 \(per\-cell scorer\-output inspection plus a targeted unit test\) is for; F4 is the hazard G4 \(paired uncertainty\) neutralizes; and F5 is the hazard G5 \(archetype disclosure\) neutralizes\. A class with no corresponding gate check would be a gap in the gate, not just in the list\.

### F1—Silent no\-op perturbations \(assurance\)\.

A perturbation is*silent\-no\-op*when the scorer\-consumed prompt is bit\-identical to the unperturbed input\. Silent no\-ops deflateΔfmt\\Delta\_\{\\mathrm\{fmt\}\}andΔsem\\Delta\_\{\\mathrm\{sem\}\}and inflate any ratio metric divided by a surface quantity\. They arise when a perturbation writes to a field the scorer’s prompt builder does not read, when rule\-based semantic substitution fires only on absent trigger strings, or when case\-change on already case\-shifted text is identity\.

### F2—Regex\-extraction artefacts \(assurance\)\.

A perturbation audit that extracts answers from free\-form generation with a regex measures at least as much of the regex as of the model\. If a perturbation pushes generation to a less\-matchable surface form,Δfmt\\Delta\_\{\\mathrm\{fmt\}\}tracks regex coverage, not model behaviour\. Parseability \(fraction of items the regex resolves to a legal answer\) is the natural gate; we fold it into G1 \(§[3](https://arxiv.org/html/2607.02586#S3)\)\.

### F3—Non\-faithful scoring \(faithfulness\)\.

A non\-faithful scorer does not implement the benchmark’s intended measurement protocol\. Three subtypes:*F3a—Inverted convention\.*The pipeline treats a benchmark’s scoring convention as default, e\.g\. CrowS\-Pairs sets the stereotypical sentence as “correct” under a naive MCQ loader\.*Fix:*the benchmark’s reference scorer \(PLL on paired sentences\(Salazaret al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib15)\)\)\.*F3b—Harness data ordering\.*A loader returns options with correct answers in fixed positions \(HuggingFace’s TruthfulQA MC2 returnscorrect\_indices=\[0,…,k−1\]=\[0,\\ldots,k\{\-\}1\]\); a scorer with positional bias is inflated\.*Fix:*MC1 \+ per\-item shuffle\.*F3c—Scorer truncation / API limit\.*A scorer asks for top\-kklog\-probabilities and sums over a target set; if targets fall outside top\-kkeach returns−∞\-\\inftyand downstreamarg⁡max\\arg\\maxcollapses\.\|Δfmt\|\|\\Delta\_\{\\mathrm\{fmt\}\}\|and\|Δsem\|\|\\Delta\_\{\\mathrm\{sem\}\}\|go to exactly zero, inflating any ratio\.*We introduced this bug during repair of F1*and caught it only after per\-cell meta\-logprob inspection\.*Fix:*targeted\-token log\-probabilities over the full vocabulary\.

### F4—Broken pairing in uncertainty \(assurance\)\.

Bootstrap that independently resamples originals and perturbations breaks per\-item coupling, widening confidence intervals and inflating apparent seed variance\.

### F5—Metric archetype mismatch \(faithfulness\)\.

Safety benchmarks split into two archetypes under construct flips\. A*diagnostic*benchmark \(TruthfulQA, ToxiGen\) is designed so flipping the construct*reverses*the answer; a competent model must show\|Δattr\|≫0\|\\Delta\_\{\\mathrm\{attr\}\}\|\\gg 0\. BBQ is an*invariance*benchmark: flipping the construct*should not*change behaviour, and a fair model shows\|Δattr\|≈0\|\\Delta\_\{\\mathrm\{attr\}\}\|\\approx 0\. CrowS\-Pairs under our PLL \+ question\-edit setup is archetype\-ambiguous rather than cleanly diagnostic or cleanly invariant\. It additionally hits a scorer\-object mismatch \(F1\): the PLL scorer readschoiceswhile our format / semantic perturbations editquestion\. A ratiocsr=\|Δattr\|/max⁡\(\|Δfmt\|,\|Δsem\|,ε\)\\textsc\{csr\}=\|\\Delta\_\{\\mathrm\{attr\}\}\|/\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|,\\varepsilon\)rewards diagnostic and penalises invariance*for doing its job*; a low CSR on BBQ is ambiguous between “benchmark broken” and “model fair,” and no scoring repair resolves the ambiguity\. Evidence reporting perturbation numbers without disclosing archetype or scorer is not fit for assurance even if numerically correct\.

## 3Self\-Audit Case Study: Method

### Panel\.

Two open\-weight 7B instruction\-tuned models \(Qwen\-2\.5\-7B\-Instruct, Mistral\-7B\-Instruct\-v0\.3\) against five safety benchmarks \(TruthfulQA, BBQ, ToxiGen, CrowS\-Pairs, XSTest\), for5×2=105\\times 2=10cells; 200 items per benchmark, deterministic decoding, single seed = 42, single template\. Hardware, software, and the dropped\-Llama provenance are in Appendix[B](https://arxiv.org/html/2607.02586#A2)\.

### Perturbations and scoring\.

Three families per item—format, semantic, attribute—with three perturbations per item per family \(detailed realisations in Appendix[C](https://arxiv.org/html/2607.02586#A3)\)\. Two scoring pipelines: a*legacy*pipeline using uniform regex extraction for MCQ and first\-token logits for ToxiGen / XSTest, and a*canonical*pipeline using each benchmark’s reference protocol \(option logprobs for TruthfulQA MC1 and BBQ; PLL on paired sentences for CrowS\(Salazaret al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib15)\); targeted\-token first\-token logprobs for ToxiGen; a 10\-pattern regex refusal classifier for XSTest\)\. Per\-pipeline scorer mappings, the legacy→\\tocanonical fix list, and parse\-failure handling are in Appendix[D](https://arxiv.org/html/2607.02586#A4)\.

### Contrast Selectivity Ratio \(CSR\)\.

For each \(item, family\) we average the unsigned per\-item score change over the family’s\|F\|=3\|F\|=3perturbations and then overN=200N=200items:

\|Δfam\|=1N​∑i=1N1\|F\|​∑p∈F\|sip−siorig\|\.\|\\Delta\_\{\\mathrm\{fam\}\}\|=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\frac\{1\}\{\|F\|\}\\sum\_\{p\\in F\}\\left\|\\,s\_\{i\}^\{p\}\-s\_\{i\}^\{\\mathrm\{orig\}\}\\,\\right\|\.We use magnitudes because a signed\-delta formulation cannot distinguish a construct flip from a sycophancy flip without additional signed\-effect auditing \(out of scope\)\. The cell statistic is

csr=\|Δattr\|max⁡\(\|Δfmt\|,\|Δsem\|,ε\),ε=0\.01,\\textsc\{csr\}=\\frac\{\|\\Delta\_\{\\mathrm\{attr\}\}\|\}\{\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|,\\varepsilon\)\},\\quad\\varepsilon=0\.01,whereε\\varepsilonis an engineering floor \(not a statistical threshold\) that keeps the ratio finite when both surface\-family deltas are near zero\. We reportPr⁡\(csr\>1\)\\Pr\(\\textsc\{csr\}\>1\)as bootstrap support for the directional inequality\|Δattr\|\>max⁡\(\|Δfmt\|,\|Δsem\|\)\|\\Delta\_\{\\mathrm\{attr\}\}\|\>\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)and*not*as a frequentistpp\-value\.*What CSR does not claim\.*A high CSR means the model responds more strongly to attribute perturbations than to format or semantic ones in*magnitude*; it does not certify that those responses are in the construct\-consistent direction\. A*confirmatory\-selective*verdict under our gate therefore reports that\|Δattr\|\|\\Delta\_\{\\mathrm\{attr\}\}\|exceeds surface\-family magnitudes—a necessary but not sufficient condition for construct sensitivity\. Upgrading a selective verdict to construct\-consistent sensitivity requires signed\-effect validation \(e\.g\., a human\- or reference\-label audit that construct flips produce direction\-correct score changes\) and is left to future work; see also the Limitations of Appendix[A](https://arxiv.org/html/2607.02586#A1)\. Paired 1000\-resample item\-level bootstrap 95% CIs are reported, following standard practice\(Efron and Tibshirani,[1993](https://arxiv.org/html/2607.02586#bib.bib23)\); resampling is once per item with all three perturbations carried together \(F4 repair\)\. Threshold honesty \(G3 = 0\.02 versus ap=0\.5,n=200p=0\.5,n=2002\-SE of≈0\.07\\approx 0\.07;ε\\varepsilonprovenance\) is in Appendix[E](https://arxiv.org/html/2607.02586#A5)\.

### Six\-point gate G1–G6 and status\-assignment algorithm\.

A cell is*confirmatory*only if all six checks in Table[1](https://arxiv.org/html/2607.02586#S3.T1)pass; status labels follow the implemented / partial / proposed convention\.111The metric was written “Construct Sensitivity Ratio” in the 2026\-04\-19 prereg\. After the E3 human audit was de\-scoped, we invoked the pre\-committed PREREG §7 fallback and renamed it “Contrast Selectivity Ratio”; the acronym CSR is retained, and construct\-validity language is retracted from any claim that would have required E3 \(Appendix[K](https://arxiv.org/html/2607.02586#A11)\)\.Each cell is assigned exactly one status by the deterministic procedure below\. Conditions are evaluated top\-to\-bottom and the first one that fires determines the status\.*Benchmark\-level*conditions depend only on the \(benchmark, scorer\) pair;*cell\-level*conditions depend on the specific cell’s gate outputs\. We place numerical\-gate failures \(G2–G4\) ahead of the benchmark\-level*scorer\-unvalidated*and*F5\-ineligible*flags because a numerical\-gate failure is the most actionable disclosure: it reports that the implemented gate actually tripped on this cell, a stronger and more specific statement than “scorer confidence is low” or “the metric interpretation is archetype\-ambiguous\.” When a cell satisfies several conditions at once \(e\.g\. Qwen×\\timesBBQ is F5 at the benchmark level*and*fails G2 at the cell level\), it is reported under the highest\-priority condition that fires; the secondary flags are still surfaced in prose in §[5](https://arxiv.org/html/2607.02586#S5)\.

1. 1\.ineligible \(F1\-type\)—benchmark\-level: the G1 instrumentation does not reach the scorer\-consumed prompt for this \(benchmark, perturbation family\) pair, so G1 cannot be evaluated\.*Our panel:*CrowS\-Pairs under PLL \+ question\-edit\.
2. 2\.failed G2–G4—cell\-level: an implemented numerical gate \(G2 above\-baseline, G3 non\-trivial denominator, G4 paired uncertainty\) fails on this cell\.
3. 3\.scorer\-unvalidated—benchmark\-level, applied when G2–G4 pass: the scorer for this benchmark has not been validated against a reference or human label\.*Our panel:*ToxiGen \(first\-token logprob targeting validated only on a 100\-item tokenizer\-coverage probe\) and XSTest \(10\-pattern refusal regex without a false\-negative audit\)\.
4. 4\.ineligible \(F5\-type\)—benchmark\-level, applied when G2–G4 pass and the scorer is validated: the benchmark is archetype\-mismatched under CSR\-as\-defined \(invariance benchmark penalised by a diagnostic ratio\)\.*Our panel:*BBQ\.
5. 5\.exploratory—cell\-level: G2–G4 pass, scorer validated, archetype aligned, but G5 \(archetype disclosure\) or G6 \(repair\-regression\) remainsproposed\.
6. 6\.confirmatory\-selective\(paired\-bootstrap CI on CSR wholly above 1; contrastive*magnitude*, not direction\) orconfirmatory\-non\-selective\(wholly below 1\): G1–G6 all implemented and passed\.

Because G5 and G6 areproposedin this submission, no cell can reach*confirmatory*; the case study contains no confirmatory cells*by construction*, a procedural fact rather than an empirical verdict about benchmarks\. The ten\-cell bucket counts in §[5](https://arxiv.org/html/2607.02586#S5)follow from this algorithm applied to each row ofcell\_validity\.csvvia the raw verdict→\\topaper bucket table in Appendix[G](https://arxiv.org/html/2607.02586#A7)\.

Table 1:Unified G1–G6 due\-diligence gate / checklist \(one set of identifiers used as both a decision procedure in this paper’s case study and a disclosure tool for governance\-evidence pipelines\)\. Status is relative to our reference implementationanalyze\_canonical\.py\.

## 4Illustrative Cases, One per Class

Table[2](https://arxiv.org/html/2607.02586#S4.T2)summarizes one realization per failure class from*our*pipeline during*this*audit\. Full per\-cell chronologies, unfixed\-acknowledged residuals, and the three additional benchmark\-specific issues \(shared\-prefix MCQ letter tokens,add\_instructioncontamination on ToxiGen, weak binary\-task surface edits\) are in Appendix[J](https://arxiv.org/html/2607.02586#A10)\.

### F1 \(silent\-no\-op\)\.

format\_change\_labelswroteitem\[‘label\_style’\]; the*legacy*prompt renderer was patched to read it, but the*canonical*MCQ scorer \(scoring\_canonical\.py\) carried a hard\-coded prompt builder that was never updated\. Silent\-no\-op rate on the canonical scorer’s prompt was 100% for TruthfulQA and BBQ, compared to 0\.2% and 1\.8% measured against the legacy renderer \(Appendix[I](https://arxiv.org/html/2607.02586#A9), Table[7](https://arxiv.org/html/2607.02586#A9.T7)\)\. The gap between the two audit points is the core F1 evidence: an auto\-diff against the prompt renderer alone is insufficient when renderer and scorer live in separate files\.

### F2 \(regex\-extraction artefact\)\.

On Qwen\-2\.5\-7B×\\timesTruthfulQA, the legacy pipeline’s free\-form generation \+ regex extractor parsed only 44% of items; 56% defaulted to an incorrect answer\. Format perturbations that shifted generation style shifted parseability itself, so\|Δfmt\|\|\\Delta\_\{\\mathrm\{fmt\}\}\|tracked regex coverage rather than model behaviour\. Canonical option\-logprob scoring forces parseability=1\.0=\\\!1\.0by construction and is our G1\-compliant fix\.

### F3a–F3c \(non\-faithful scoring\)\.

*F3a:*Mistral\-7B×\\timesCrowS\-Pairs legacy accuracy was 0\.99, which reads as “excellent fairness” but is in fact*maximal*stereotypical preference: naive MCQ loading setcorrect\_indices=\[0\]\\\!=\\\!\[0\]pointing atsent\_more\. PLL scoring on paired sentences\(Salazaret al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib15)\)restores the intended scoring convention\.*F3b:*TruthfulQA MC2 returnscorrect\_indices=\[0,…,k−1\]\\\!=\\\!\[0,\\ldots,k\{\-\}1\]; under any mild “prefer first option” bias, the score is inflated \(we observedscore\_original=1\.000\\\!=\\\!1\.000on an early canonical Qwen rerun\)\. Switching to MC1 with deterministic per\-item shuffle fixes it\.*F3c:*After moving ToxiGen to first\-token log\-probabilities, the scorer computedlptoxic=lpnon−toxic=−∞\\mathrm\{lp\}\_\{\\mathrm\{toxic\}\}=\\mathrm\{lp\}\_\{\\mathrm\{non\-toxic\}\}=\-\\inftyon most items because neither target token fell intop\_k=50\\\!=\\\!50; the scorer defaulted to the majority class, producing\|Δfmt\|=\|Δsem\|=0\|\\Delta\_\{\\mathrm\{fmt\}\}\|=\|\\Delta\_\{\\mathrm\{sem\}\}\|=0exactly and\|Δattr\|=0\.2\|\\Delta\_\{\\mathrm\{attr\}\}\|=0\.2as a pure label\-side artefact\.*We introduced this bug during the repair of F1*; we caught it only by inspecting per\-cell meta log\-probabilities\. The fix istargeted\_first\_token\_logprobs\(prompt, token\_ids\), which bypasses top\-kkentirely\.

### F4 \(broken pairing\)\.

Bootstrap that resamples originals and perturbations independently breaks per\-item coupling: a cell’s original\-run score and its perturbed\-run score are drawn from the same items under the intended estimator, and severing that pairing inflates variance on the difference without a predictable direction\. F4 in this paper is evidenced by*out\-of\-panel/legacy numbers plus a methodological argument*, not by a controlled in\-panel ablation\. The out\-of\-panel legacy\-Llama cell on ToxiGen gave a CSR point estimate≈9\.5\\approx 9\.5with 95% CI width 26\.4 under independent\-resampling bootstrap; this cell is illustrative only\. A like\-for\-like in\-panel before/after is not available because the canonical pipeline removed the unpaired bootstrap outright \(PREREG amendment \(e\), Appendix[K](https://arxiv.org/html/2607.02586#A11)\); paired\-bootstrap CI widths in the canonical panel \(Appendix[H](https://arxiv.org/html/2607.02586#A8), Table[6](https://arxiv.org/html/2607.02586#A8.T6)\) run from 0\.00 on degenerate cells up to 54\.91 on Mistral×\\timesXSTest \(CSR 84\.86, CI\[44\.92,99\.83\]\[44\.92,99\.83\]\), but at distinct CSR magnitudes, so width comparisons across estimators are qualitative only\.

### F5 \(archetype mismatch\)\.

Canonical Qwen\-2\.5\-7B×\\timesBBQ produces\|Δattr\|=0\.00\|\\Delta\_\{\\mathrm\{attr\}\}\|\\\!=\\\!0\.00at an unperturbed score of0\.340\.34\(near\-chance on a three\-option task\)\. Diagnostic reading: “broken benchmark\.” Invariance reading: “maximally fair model”—BBQ is designed so a fair model should*not*change its answer under a demographic swap\. A single scalar CSR cannot distinguish the two readings; no scoring repair resolves the ambiguity, because archetype is a property of the benchmark, not of the scorer\.

Table 2:Five failure classes, one realization per class in our canonical Qwen\-2\.5\-7B / Mistral\-7B panel\. “Verdict” uses the short labels of §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\.Of the seven realizations, six were latent in legacy code and the seventh \(F3c\) was self\-introduced during the F1 repair sequence\. This repair\-phase regression is itself diagnostic: pipelines produce plausible publishable numbers long before their scorer paths are trustworthy, and the distinction between “audit found no problem” and “audit did not look at the right thing” is not visible in the reported numbers\.

## 5Per\-Cell Evidence

Applying the §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)hierarchical precedence to our 10\-cell panel yields the eligibility breakdown in Table[3](https://arxiv.org/html/2607.02586#S5.T3):3 ineligible,3 scorer\-unvalidated,2 failed G2–G4,2 exploratory, and0 confirmatory\. The zero is procedural—G5 \(archetype disclosure\) and G6 \(repair\-regression testing\) remainproposed, so a cell whose paired bootstrap CI is wholly above 1 is held as*exploratory*rather than*confirmatory\-selective*under our hierarchy\. The four\-way breakdown between “ineligible”, “scorer\-unvalidated”, “failed implemented gate”, and “exploratory”*is*the headline: an assurance\-grade evidence report should distinguish these, not collapse them into a single numerical verdict\. Figure[2](https://arxiv.org/html/2607.02586#S5.F2)shows per\-cellcsrpoint estimates with 1000\-resample paired\-bootstrap 95% CIs; full numerics per cell \(original score,Δfmt\\Delta\_\{\\mathrm\{fmt\}\},Δsem\\Delta\_\{\\mathrm\{sem\}\},Δattr\\Delta\_\{\\mathrm\{attr\}\}, CSR, CI, raw gate verdict\) are in Appendix[H](https://arxiv.org/html/2607.02586#A8), Table[6](https://arxiv.org/html/2607.02586#A8.T6)\. The paper\-bucket assignments in Table[3](https://arxiv.org/html/2607.02586#S5.T3)follow mechanically from the raw per\-cell verdict column ofcell\_validity\.csvvia the mapping in Appendix[G](https://arxiv.org/html/2607.02586#A7)\.

Table 3:Per\-cell eligibility in the canonical panel\. The primary bucket is the first\-firing condition under §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4); secondary flags mark non\-primary F5 or scorer\-unvalidated concerns\. Raw verdicts and full numerics are in Appendix[H](https://arxiv.org/html/2607.02586#A8)\.### How the counts read\.

The3 ineligiblecells expose F1 and F5\. CrowS\-Pairs \(×2\\times 2\) is archetype\-ambiguous under our PLL \+ question\-edit setup and hits a scorer\-object mismatch \(F1\): the PLL scorer readschoiceswhile format / semantic perturbations editquestion, so\|Δsem\|=0\|\\Delta\_\{\\mathrm\{sem\}\}\|\\\!=\\\!0is consistent with F1\. Mistral×\\timesBBQ passes all implemented gates but its attribute family swaps demographic tokens while the answer key stays fixed, soΔattr\\Delta\_\{\\mathrm\{attr\}\}is not a clean construct flip under CSR\-as\-defined \(F5\)\. The3 scorer\-unvalidatedcells \(Qwen×\\timesXSTest plus both ToxiGen cells\) pass all*implemented*gates but rely on scorers for which we supplied no external validation: a 10\-pattern refusal regex for XSTest, and a first\-token\-logprob targeting for ToxiGen whose tokenizer\-coverage is audited only on a 100\-item subset \(Appendix[B](https://arxiv.org/html/2607.02586#A2); mean target\-token coverage0\.8530\.853and0\.8570\.857,p5=0\.64p\_\{5\}=0\.64and0\.740\.74\)\. The2 failed G2–G4cells are Qwen×\\timesBBQ \(fails G2: unperturbed score0\.3400\.340is only0\.7​pp0\.7\\,\\mathrm\{pp\}above the1/31/3uniform baseline, under the2​SE≈6\.7​pp2\\,\\mathrm\{SE\}\\approx 6\.7\\,\\mathrm\{pp\}margin\) and Mistral×\\timesXSTest \(fails G2*and*G3: unperturbed0\.5650\.565vs\.1/21/2uniform, andmax⁡\(\|Δfmt\|,\|Δsem\|\)=0\.012<0\.02\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)\\\!=\\\!0\.012<0\.02, the one truly denominator\-dominated cell in the panel\)\. The2 exploratorycells \(Qwen×\\timesTruthfulQA and Mistral×\\timesTruthfulQA\) pass all implemented gates with paired\-bootstrap CIs wholly above 1 \(CSR13\.0413\.04with CI\[10\.53,17\.65\]\[10\.53,17\.65\]and CSR8\.338\.33with CI\[6\.67,10\.34\]\[6\.67,10\.34\]\); under paper precedence \(G5, G6proposed\), they are held as exploratory rather than confirmatory\-selective\. None of these four non\-confirmatory statuses supports a construct\-\(in\)validity conclusion about the underlying benchmark\.

![Refer to caption](https://arxiv.org/html/2607.02586v1/x1.png)Figure 2:Canonical\-pipeline CSR per cell with 1000\-resample paired\-bootstrap 95% CI\. Colour encodes the gate verdict \(§[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\)\. Mistral×\\timesXSTest fails G3 \(max⁡\(\|Δfmt\|,\|Δsem\|\)=0\.012<0\.02\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)=0\.012<0\.02\); no cell reaches*confirmatory\-selective*under the full G1–G6 gate because G5, G6 remainproposed\. CSR values for*ineligible*and*scorer\-unvalidated*cells are shown descriptively only and are not interpretable as benchmark\-validity evidence \(see §[5](https://arxiv.org/html/2607.02586#S5)\)\. Full per\-cell numbers are in Appendix[H](https://arxiv.org/html/2607.02586#A8)\.

## 6Discussion

Everything below is read off a*single*two\-model, five\-benchmark panel\. We state the takeaways as what this case study illustrates, not as comprehensive findings; their generality is untested\.

### Why a self\-audit chronology must be published\.

F3b \(TruthfulQA MC2 correct\-first ordering\) and F3c \(ToxiGen top\-kktruncation\) had already produced plausible publishable numbers before per\-cell scorer\-output inspection exposed them\. Any assurance\-grade benchmark audit should therefore publish a self\-audit chronology, not only cleaned final numbers\. Box 1 gives the minimal fields; Appendix[J](https://arxiv.org/html/2607.02586#A10)is our worked instance\. Without it, repair silently absorbed into a clean narrative is indistinguishable from a pipeline where the bugs were never caught\. Appendix[G](https://arxiv.org/html/2607.02586#A7)maps the rawcell\_validity\.csvverdicts to Table[3](https://arxiv.org/html/2607.02586#S5.T3), making the eligibility counts mechanical\.

Box 1: Minimal self\-audit chronology template*Per failure encountered*, record:\(1\)a stable bug ID and its failure class \(e\.g\. F1–F5 or a project\-local code\);\(2\)discovery date and*how*it was found \(which inspection or test surfaced it—not just “code review”\);\(3\)affected cells \(every model×\\timesbenchmark the defect touched\);\(4\)the symptom*as it appeared in the reported numbers*\(what a reader would have seen and believed\);\(5\)provenance—*inherited*from the initial pipeline vs\.*introduced during repair*;\(6\)the fix \(the specific code change\)*and*a regression test that fails on the pre\-fix code;\(7\)the before/after delta on the headline statistic for each affected cell;\(8\)residual status:fixed,disclosed\-unfixed, orout\-of\-scope, with the reason\.*Per chronology as a whole*, also state: total entries, how many were repair\-introduced, and which remain unfixed and why\. An entry missing field \(4\) or \(7\) is not auditable: the reader cannot tell what the wrong number was or how much it moved\.

### Why we report four buckets, not one verdict\.

The four non\-confirmatory statuses in §[5](https://arxiv.org/html/2607.02586#S5)separate different consumer actions:*ineligible*calls for scope revision,*scorer\-unvalidated*for scorer validation,*failed G2–G4*for threshold or evidence revision, and*exploratory*for disclosed, non\-confirmatory release\. Collapsing these into a single “inconclusive” verdict would hide the reasoning step that a governance reviewer most needs\.

### Zero confirmatory cells is procedural, not an empirical null\.

The zero confirmatory verdict is not a claim about the five benchmarks\. Under §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4), even a cell whose CI is wholly above 1 remains*exploratory*while G5 or G6 isproposed; the two TruthfulQA cells would upgrade only after those gates are implemented\. The contribution is the gate and taxonomy, not a benchmark\-validity verdict\.*No benchmark\-validity claim is issued for any cell\.*

## 7Conclusion

Perturbation\-based benchmark\-validity audits can be silently wrong in at least five pipeline\-level ways\. We documented all five in a single 10\-cell case study, and caught one \(F3c\) only after introducing it ourselves during repair\. F1–F5 is an illustrative, non\-exhaustive starting list whose calibration across codebases is open work\.

The reusable artefact is the F1–F5 taxonomy, the §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)status\-assignment algorithm, the G1–G6 gate, and the self\-audit chronology of Box 1\. This is a withholding and disclosure protocol, not a scalar metric or benchmark leaderboard: all 10 cells in our panel remain non\-confirmatory, and even a future*confirmatory\-selective*cell would certify contrast magnitude, not signed direction \(Appendix[A](https://arxiv.org/html/2607.02586#A1)\)\.

Our ask is one sentence\.*Any audit producing benchmark\-based evidence for governance use should publish its self\-audit chronology—at minimum the eight fields of Box 1—alongside its numbers\.*Without it, a clean number and a silently broken pipeline look the same\.

## References

- N\. Alzahrani, H\. Alyahya, Y\. Alnumay, S\. AlRashed, S\. Alsubaie, Y\. Almushayqih, F\. Mirza, N\. Alotaibi, N\. Al\-Twairesh, A\. Alowisheq, M\. S\. Bari, and H\. Khan \(2024\)When benchmarks are targets: revealing the sensitivity of large language model leaderboards\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 13787–13805\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.744),[Link](https://aclanthology.org/2024.acl-long.744)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- S\. L\. Blodgett, G\. Lopez, A\. Olteanu, R\. Sim, and H\. Wallach \(2021\)Stereotyping Norwegian salmon: an inventory of pitfalls in fairness benchmark datasets\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 1004–1015\.External Links:[Link](https://aclanthology.org/2021.acl-long.81)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Y\. Chen, P\. Qian, S\. Wang, S\. Zhang, H\. Xu, S\. Lin, and X\. Wei \(2026\)Does RAG know when retrieval is wrong? diagnosing context compliance under knowledge conflict\.External Links:2605\.14473,[Link](https://arxiv.org/abs/2605.14473)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- L\. J\. Cronbach and P\. E\. Meehl \(1955\)Construct validity in psychological tests\.Psychological Bulletin52\(4\),pp\. 281–302\.Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- B\. Efron and R\. J\. Tibshirani \(1993\)An introduction to the bootstrap\.Chapman and Hall / CRC\.Cited by:[§3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px3.p1.10)\.
- European Parliament and Council \(2024\)Regulation \(eu\) 2024/1689 laying down harmonised rules on artificial intelligence \(artificial intelligence act\)\.Note:Official Journal of the European Union; see in particular Art\. 55 on obligations of providers of general\-purpose AI models with systemic riskExternal Links:[Link](https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- T\. Hartvigsen, S\. Gabriel, H\. Palangi, M\. Sap, D\. Ray, and E\. Kamar \(2022\)ToxiGen: a large\-scale machine\-generated dataset for adversarial and implicit hate speech detection\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3309–3326\.External Links:[Link](https://aclanthology.org/2022.acl-long.234)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p4.1)\.
- A\. Z\. Jacobs and H\. Wallach \(2021\)Measurement and fairness\.InProceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency \(FAccT\),pp\. 375–385\.External Links:[Document](https://dx.doi.org/10.1145/3442188.3445901)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Y\. Ji, W\. Lan, and P\. Ng \(2025\)MRAG\-Suite: a diagnostic evaluation platform for visual retrieval\-augmented generation\.arXiv preprint arXiv:2509\.24253\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2509.24253),[Link](https://arxiv.org/abs/2509.24253)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Y\. Ji, Z\. Li, R\. Meng, and D\. He \(2026\)Retrieval–reasoning processes for multi\-hop question answering: a four\-axis design framework and empirical trends\.arXiv preprint arXiv:2601\.00536\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.00536),[Link](https://arxiv.org/abs/2601.00536)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- X\. Jiang, S\. Yang, W\. Yang, Y\. Liu, and C\. Ji \(2026\)SOK: a taxonomy of attack vectors and defense strategies for agentic supply chain runtime\.arXiv preprint arXiv:2602\.19555\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2602.19555),[Link](https://arxiv.org/abs/2602.19555)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- D\. Kiela, M\. Bartolo, Y\. Nie, D\. Kaushik, A\. Geiger, Z\. Wu, B\. Vidgen, G\. Prasad, A\. Singh, P\. Ringshia, Z\. Ma, T\. Thrush, S\. Riedel, Z\. Waseem, P\. Stenetorp, R\. Jia, M\. Bansal, C\. Potts, and A\. Williams \(2021\)Dynabench: rethinking benchmarking in NLP\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(NAACL\-HLT\),pp\. 4110–4124\.External Links:[Link](https://aclanthology.org/2021.naacl-main.324)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- J\. P\. Lalor, H\. Wu, and H\. Yu \(2019\)Learning latent parameters without human response patterns: item response theory with artificial crowds\.InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing \(EMNLP\-IJCNLP\),pp\. 4249–4259\.External Links:[Link](https://aclanthology.org/D19-1434)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- T\. Lan, J\. Xu, X\. He, J\. Hwang, and L\. Li \(2025\)Attention consistency for LLMs explanation\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 1736–1750\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.91),[Link](https://aclanthology.org/2025.findings-emnlp.91)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Y\. Li, Z\. Fan, and Z\. Zhuang \(2026a\)Auditing reasoning\-trace memorization claims after unlearning with head\-conditioned canaries\.arXiv preprint arXiv:2605\.18891\.External Links:[Link](https://arxiv.org/abs/2605.18891)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Y\. Li, Z\. Fan, and Z\. Zhuang \(2026b\)SafetyRepro: configuration\-conditional rank instability on alignment benchmarks\.arXiv preprint arXiv:2605\.25492\.External Links:[Link](https://arxiv.org/abs/2605.25492)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- S\. Lin, J\. Hilton, and O\. Evans \(2022\)TruthfulQA: measuring how models mimic human falsehoods\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 3214–3252\.External Links:[Link](https://aclanthology.org/2022.acl-long.229)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p4.1)\.
- H\. Luo, S\. Dai, C\. Ni, X\. Li, G\. Zhang, K\. Wang, T\. Liu, and H\. Salam \(2025\)AgentAuditor: human\-level safety and security evaluation for LLM agents\.arXiv preprint arXiv:2506\.00641\.Note:Accepted to NeurIPS 2025External Links:[Link](https://arxiv.org/abs/2506.00641)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- H\. Luo, Z\. Deng, R\. Chen, and Z\. Liu \(2024a\)FAIntbench: a holistic and precise benchmark for bias evaluation in text\-to\-image models\.arXiv preprint arXiv:2405\.17814\.External Links:[Link](https://arxiv.org/abs/2405.17814)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- H\. Luo, H\. Huang, Z\. Deng, X\. Li, H\. Wang, Y\. Jin, Y\. Liu, W\. Xu, and Z\. Liu \(2024b\)BIGbench: a unified benchmark for evaluating multi\-dimensional social biases in text\-to\-image models\.arXiv preprint arXiv:2407\.15240\.External Links:[Link](https://arxiv.org/abs/2407.15240)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- H\. Luo, Z\. Huang, S\. Chung, Y\. Wang, Y\. Jin, J\. Li, J\. Li, X\. Li, and H\. Salam \(2026a\)AtelierEval: agentic evaluation of humans & LLMs as text\-to\-image prompters\.InProceedings of the 43rd International Conference on Machine Learning \(ICML\),Note:arXiv:2605\.22645External Links:[Link](https://openreview.net/forum?id=Q8LOMDdr2y)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- H\. Luo, Z\. Huang, H\. Huang, Z\. Deng, R\. Chen, X\. Li, Z\. Liu, and H\. Salam \(2026b\)BiasIG: benchmarking multi\-dimensional social biases in text\-to\-image models\.arXiv preprint arXiv:2604\.11934\.Note:Accepted to IJCNN 2026External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.11934),[Link](https://arxiv.org/abs/2604.11934)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- H\. Luo, C\. Ni, J\. Wen, Z\. Huang, Y\. Wang, B\. Liao, S\. Chung, Y\. Jin, X\. Li, W\. Xu, X\. Wang, and H\. Salam \(2026c\)CentaurEval: benchmarking human\-in\-the\-loop value in agentic coding\.InProceedings of the 43rd International Conference on Machine Learning \(ICML\),Note:arXiv:2512\.04111External Links:[Link](https://openreview.net/forum?id=GR335S9IgZ)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- S\. J\. Messick \(1995\)Validity of psychological assessment: validation of inferences from persons’ responses and performances as scientific inquiry into score meaning\.American Psychologist50\(9\),pp\. 741–749\.External Links:[Document](https://dx.doi.org/10.1037/0003-066X.50.9.741)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- M\. Mizrahi, G\. Kaplan, D\. Malkin, R\. Dror, D\. Shahaf, and G\. Stanovsky \(2024\)State of what art? a call for multi\-prompt LLM evaluation\.Transactions of the Association for Computational Linguistics12,pp\. 933–949\.Note:arXiv:2401\.00595External Links:[Document](https://dx.doi.org/10.1162/tacl%5Fa%5F00681),[Link](https://aclanthology.org/2024.tacl-1.52)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- N\. Nangia, C\. Vania, R\. Bhalerao, and S\. R\. Bowman \(2020\)CrowS\-Pairs: a challenge dataset for measuring social biases in masked language models\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 1953–1967\.External Links:[Link](https://aclanthology.org/2020.emnlp-main.154)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p4.1)\.
- National Institute of Standards and Technology \(2023\)Artificial intelligence risk management framework \(AI RMF 1\.0\)\.Note:NIST AI 100\-1External Links:[Link](https://www.nist.gov/itl/ai-risk-management-framework)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- A\. Parrish, A\. Chen, N\. Nangia, V\. Padmakumar, J\. Phang, J\. Thompson, P\. M\. Htut, and S\. R\. Bowman \(2022\)BBQ: a hand\-built bias benchmark for question answering\.InFindings of the Association for Computational Linguistics: ACL 2022,pp\. 2086–2105\.External Links:[Link](https://aclanthology.org/2022.findings-acl.165)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p4.1)\.
- P\. Qian, S\. Wang, X\. Wang, Y\. Chen, W\. Xu, Q\. Yu, S\. Lin, S\. Zhang, J\. You, and X\. Wei \(2026\)Relevant is not warranted: evidence\-force calibration for cited RAG\.External Links:2605\.28044,[Link](https://arxiv.org/abs/2605.28044)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- I\. D\. Raji, E\. M\. Bender, A\. Paullada, E\. Denton, and A\. Hanna \(2021\)AI and the everything in the whole wide world benchmark\.InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks,Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- A\. Reuel, B\. Bucknall, S\. Casper, T\. Fist, L\. Soder, O\. Aarne, L\. Hammond, L\. Ibrahim, A\. Chan, P\. Wills, M\. Anderljung, B\. Garfinkel, L\. Heim, A\. Trask, G\. Mukobi, R\. Schaeffer, M\. Baker, S\. Hooker, I\. Solaiman, A\. S\. Luccioni, N\. Rajkumar, N\. Moës, J\. Ladish, D\. Bau, P\. Bricman, N\. Guha, J\. Newman, Y\. Bengio, T\. South, A\. Pentland, S\. Koyejo, M\. J\. Kochenderfer, and R\. Trager \(2025\)Open problems in technical AI governance\.Transactions on Machine Learning Research\.Note:arXiv:2407\.14981External Links:[Link](https://arxiv.org/abs/2407.14981)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- M\. T\. Ribeiro, T\. Wu, C\. Guestrin, and S\. Singh \(2020\)Beyond accuracy: behavioral testing of NLP models with CheckList\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 4902–4912\.External Links:[Link](https://aclanthology.org/2020.acl-main.442)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- P\. Röttger, H\. R\. Kirk, B\. Vidgen, G\. Attanasio, F\. Bianchi, and D\. Hovy \(2024\)XSTest: a test suite for identifying exaggerated safety behaviours in large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies \(Volume 1: Long Papers\),pp\. 5377–5400\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.301),[Link](https://aclanthology.org/2024.naacl-long.301)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p4.1)\.
- J\. Salazar, D\. Liang, T\. Q\. Nguyen, and K\. Kirchhoff \(2020\)Masked language model scoring\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,pp\. 2699–2712\.External Links:[Link](https://aclanthology.org/2020.acl-main.240)Cited by:[item 6](https://arxiv.org/html/2607.02586#A1.I1.i6.p1.1),[Table 4](https://arxiv.org/html/2607.02586#A4.T4.7.5.3.2.1.1),[Appendix D](https://arxiv.org/html/2607.02586#A4.p1.1),[§1](https://arxiv.org/html/2607.02586#S1.p4.1),[§2](https://arxiv.org/html/2607.02586#S2.SS0.SSS0.Px5.p1.7),[§3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2607.02586#S4.SS0.SSS0.Px3.p1.9),[Table 2](https://arxiv.org/html/2607.02586#S4.T2.5.5.4.1.1)\.
- M\. Sclar, Y\. Choi, Y\. Tsvetkov, and A\. Suhr \(2024\)Quantifying language models’ sensitivity to spurious features in prompt design, or: how i learned to start worrying about prompt formatting\.InThe Twelfth International Conference on Learning Representations \(ICLR\),External Links:[Link](https://openreview.net/forum?id=RIu5lyNXjT)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- State Administration for Market Regulation and Standardization Administration of China \(2025\)GB/T 45654–2025, cybersecurity technology—basic security requirements for generative artificial intelligence service\.Note:National Standard of the People’s Republic of ChinaIssued 2025\-04\-25; implemented 2025\-11\-01External Links:[Link](https://openstd.samr.gov.cn/bzgk/std/newGbInfo?hcno=F67D3F376E0A0A0FF5317FB36B32A30A)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- C\. Vania, P\. M\. Htut, W\. Huang, D\. Mungra, R\. Y\. Pang, J\. Phang, H\. Liu, K\. Cho, and S\. R\. Bowman \(2021\)Comparing test sets with item response theory\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 1141–1158\.External Links:[Link](https://aclanthology.org/2021.acl-long.92)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- S\. Wang, P\. Qian, Y\. Chen, J\. You, X\. Wang, X\. Jiang, L\. Liu, H\. Yu, and J\. Xu \(2026a\)When safe skills collide: measuring compositional risk in agent skill ecosystems\.External Links:2606\.00448,[Link](https://arxiv.org/abs/2606.00448)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Y\. Wang, X\. Sun, Y\. Li, Z\. Fan, and Z\. Zhuang \(2026b\)Auditing and fixing economic validity in tabular foundation models for discrete choice\.arXiv preprint arXiv:2605\.26559\.External Links:[Link](https://arxiv.org/abs/2605.26559)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.
- Z\. Zhuang, Y\. Li, and Z\. Fan \(2026\)Pre\-registering the detectable effect: a paired\-MDE budget for 4\-bit quantization benchmarks, with a pilot audit\.arXiv preprint arXiv:2605\.28873\.External Links:[Link](https://arxiv.org/abs/2605.28873)Cited by:[§1](https://arxiv.org/html/2607.02586#S1.p1.1)\.

## Appendix ADetailed Limitations

We place limitations here so the reader can calibrate every number in the main text against what this evidence cannot support\.

1. 1\.Two models, five benchmarks, 200 items per cell, one seed, one prompt template\.Our 10\-cell grid \(Qwen\-2\.5\-7B and Mistral\-7B\-v0\.3×\\timesfive benchmarks\) cannot support any claim that a class of 7–8B open\-weight models exhibits any pattern\.
2. 2\.Post\-prereg exploratory status\.Our preregistration \(§[K](https://arxiv.org/html/2607.02586#A11)\) was committed before we discovered failures F3b and F3c; analyses after those discoveries are formally exploratory\.
3. 3\.Structurally\-ineligible cells\.CrowS\-Pairs is archetype\-ambiguous under our PLL \+ demographic\-swap configuration and additionally hits F1 scorer\-path mismatch; BBQ is an invariance benchmark penalised by our diagnostic ratio metric \(F5\)\. Their numbers illustrate the two failure classes but are not benchmark\-validity claims\.
4. 4\.Unvalidated scorers for ToxiGen and XSTest\.ToxiGen scoring uses targeted first\-token log\-probabilities over a fixed token set \(after the F3c repair\); XSTest refusal classification uses a ten\-pattern regex over generations that we did not separately audit for false\-negatives against human judgement\. Neither scorer is a human\-judged label; both cells are reported in the*scorer\-unvalidated*bucket\.
5. 5\.CSR is a contrastive\-magnitude metric, not a signed\-direction metric\.csruses unsigned family\-deltas, so a paired\-bootstrap CI wholly above 1 only certifies that\|Δattr\|\|\\Delta\_\{\\mathrm\{attr\}\}\|exceeds surface\-family magnitudes\. It does*not*certify that attribute\-flip responses are in the construct\-consistent direction: a large\|Δattr\|\|\\Delta\_\{\\mathrm\{attr\}\}\|is consistent with any of construct\-sensitive flipping, sycophancy flipping, lexical\-salience noise, or an uncaught scorer artefact\. Upgrading a*confirmatory\-selective*reading to*construct\-sensitive*requires signed\-effect validation \(the E3 human audit, deferred\) and is future work\.
6. 6\.Canonical scorers are audit\-hardened, not reference\-faithful across the board\.TruthfulQA MC1 with deterministic per\-item shuffle, ToxiGen targeted\-token first\-token logprobs, and XSTest’s 10\-pattern refusal regex are audit\-robust substitutes rather than reference\-faithful implementations; see Appendix[D](https://arxiv.org/html/2607.02586#A4)for the reference / audit / validation split\. CrowS\-Pairs PLL on paired sentences\(Salazaret al\.,[2020](https://arxiv.org/html/2607.02586#bib.bib15)\)and BBQ option log\-probabilities split bycontext\_conditionare reference\-faithful\.
7. 7\.The six\-point checklist is a proposed tool, not a standard\.It does not validate thresholds against external reference data, does not assign institutional ownership, and does not include anti\-gaming provisions\. Three further benchmark\-specific limitations \(shared\-prefix MCQ tokens,add\_instructioncontamination, weak surface edits on binary tasks\) are listed in Appendix[J](https://arxiv.org/html/2607.02586#A10)\.

## Appendix BPanel and Compute

Compute: a single NVIDIA RTX 4080 \(16 GB\) with HuggingFace Transformers 4\.57 in bfloat16; approximately 26 GPU\-hours total across the legacy and canonical pipelines\. The panel originally included Llama\-3\.1\-8B; per author decision, Llama was dropped after the legacy runs and before the canonical rerun\. Only the two\-model canonical panel is the case study of this paper; legacy\-Llama numbers appear only in this appendix and the case\-level audit trail \(Appendix[J](https://arxiv.org/html/2607.02586#A10)\)\. Software versions: Python 3\.13, PyTorch 2\.10,transformers4\.57\.6; model revisions pinned at load time \(Qwen/Qwen2\.5\-7B\-Instructandmistralai/Mistral\-7B\-Instruct\-v0\.3; tokenizer revision recorded in results JSON\)\.

## Appendix CPerturbation\-Family Realisations

Three families per item; three perturbations per item per family\. Parse failures \(legacy regex only\) are symmetric: if either the original or the perturbed item fails parsing, the item\-level paired difference for that family is dropped\.*Format:*case changes, whitespace / punctuation noise, label\-format changes, choice reordering, and—for text\-classification benchmarks only—instruction prepend\. We retainadd\_instructionin the format family despite its task\-boundary contamination on ToxiGen \(Appendix[J](https://arxiv.org/html/2607.02586#A10)marks this as unfixed\-acknowledged\)\.*Semantic:*paraphrase templates, synonym substitution, rule\-based sentence restructure\.*Attribute:*benchmark\-specific construct flips: opposite\-construct stem rewrite \(TruthfulQA\), demographic swap \(BBQ, CrowS\-Pairs\), toxicity\-flip rewrite \(ToxiGen\), safe↔\\leftrightarrowunsafe pair lookup \(XSTest\)\.

## Appendix DScoring Pipelines

The*legacy*pipeline uses a uniform regex\-extraction scorer over free\-form generation for all multiple\-choice benchmarks and a first\-token logit comparison for ToxiGen and XSTest\. The*canonical*pipeline implements each benchmark’s intended scoring protocol: option log\-probabilities for TruthfulQA MC1, option log\-probabilities split bycontext\_conditionfor BBQ, pseudo\-log\-likelihood over the paired sentences for CrowS\-Pairs \(followingSalazaret al\.\([2020](https://arxiv.org/html/2607.02586#bib.bib15)\)\), targeted\-token first\-token log\-probabilities for ToxiGen, and a ten\-pattern regex refusal classifier over generations for XSTest\.

The legacy→\\tocanonical fix list: readlabel\_formatwhen rendering MCQ prompts \(F1 fix\); replaceflip\_truthfulqalabel inversion with opposite\-construct stem rewrite; switch CrowS scoring to PLL on paired sentences \(F3a\); switch TruthfulQA loading from MC2 to MC1 to eliminate correct\-first ordering \(F3b\); use targeted\-token first\-token log\-probabilities for ToxiGen to bypass top\-kktruncation \(F3c\); switch the bootstrap to an item\-level paired resampler \(F4\)\.

### Reference vs\. audit vs\. validation evidence\.

“Canonical” collapses three distinct things that should be kept apart when reading downstream claims\. Table[4](https://arxiv.org/html/2607.02586#A4.T4)splits them\.

Table 4:Scorer fidelity split\.*Reference scorer*==the benchmark’s published scoring protocol\.*Audit scorer used here*==the scorer our canonical pipeline actually runs\.*Validation evidence*==what we have to back the audit scorer;nonemeans the cell ships as*scorer\-unvalidated*under §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\.

## Appendix EThreshold Honesty

The0\.020\.02threshold in G3 and the0\.010\.01floor ofε\\varepsilonare*engineering choices*, not statistical thresholds\. The intent is to avoid ratio explosion when both surface\-family deltas are near zero; the specific numeric choice is tied to our 200\-item subset size but we do not claim it is optimal or that0\.020\.02equals “two standard errors” of anything \(for a proportion atp=0\.5p=0\.5withn=200n=200,2​SE≈0\.072\\,\\mathrm\{SE\}\\approx 0\.07, so0\.020\.02is substantially more permissive than a strict statistical\-significance threshold\)\. The20%20\\%silent\-no\-op cap in G1 is likewise pragmatic, motivated by our observed 0–83% range \(Appendix[I](https://arxiv.org/html/2607.02586#A9)\)\. A fuller treatment would derive each threshold from a target sensitivity level in the downstream evidence consumer\. Satisfying G1–G6 does not establish that any benchmark is construct\-valid; it substantially reduces the risk that non\-confirmation is driven by known preventable pipeline failures\. A production due\-diligence instrument would additionally require institutional ownership, externally\-validated thresholds, anti\-gaming provisions, and legal\-compatibility review\.

### Threshold sensitivity of the 10\-cell breakdown\.

We resweep the three threshold knobs and re\-derive the five\-bucket counts of §[5](https://arxiv.org/html/2607.02586#S5)under each setting; the qualitative breakdown is stable over the grid we considered\.G3in\{0\.01,0\.02,0\.05\}\\\{0\.01,0\.02,0\.05\\\}: the only boundary cell is Mistral×\\timesXSTest withmax⁡\(\|Δfmt\|,\|Δsem\|\)=0\.012\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)\\\!=\\\!0\.012\. At0\.010\.01that cell passes G3 but still fails G2 \(0\.5650\.565vs\.0\.5000\.500, under2​SE2\\,\\mathrm\{SE\}\), so it remains*failed G2–G4*; at0\.050\.05it fails both G2 and G3 and remains*failed G2–G4*\. Every other in\-panel cell with non\-degenerate denominator hasmax⁡\(\|Δfmt\|,\|Δsem\|\)≥0\.050\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)\\\!\\geq\\\!0\.050, and the F1\-degenerate CrowS cells are pre\-filtered\.ε\\varepsilonin\{0\.001,0\.01,0\.1\}\\\{0\.001,0\.01,0\.1\\\}:ε\\varepsilononly floors the CSR denominator; it shifts CSR point estimates on denominator\-dominated cells but does not change any G2 or G3 outcome, so bucket assignment is unchanged\.G1 silent\-no\-op capin\{5,10,20\}%\\\{5,10,20\\\}\\%: the observed canonical\-scorer no\-op rates are\{0\.0,0\.0,100,100,83\}%\\\{0\.0,0\.0,100,100,83\\\}\\%for XSTest/ToxiGen/TruthfulQA/BBQ/CrowS; any cap≤83%\\leq 83\\%leaves CrowS F1\-ineligible and TruthfulQA/BBQ F1\-passing \(their no\-op rates are0%0\\%on the canonical path after the F1b repair\)\. The3/3/2/2/03/3/2/2/0breakdown therefore holds over this grid\. Thresholds outside the grid \(e\.g\., G3=0\.1=0\.1, which would flag denominators up to half thep=0\.5,n=200p\\\!=\\\!0\.5,n\\\!=\\\!2002\-SE of0\.070\.07\) would reclassify TruthfulQA’s exploratory cells; we report the grid above as the regime where the thresholds are defensible as conservative engineering floors\.

## Appendix FPer\-Benchmark Eligibility Notes

The main\-body eligibility breakdown \(§[5](https://arxiv.org/html/2607.02586#S5), Table[3](https://arxiv.org/html/2607.02586#S5.T3), Figure[2](https://arxiv.org/html/2607.02586#S5.F2)\) summarises per\-cell status under the §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)precedence\. Additional per\-benchmark notes:

TruthfulQA\.Both cells score above the1/41/4uniform baseline \(G2 passes: Qwen0\.4850\.485, Mistral0\.5200\.520\) withmax⁡\(\|Δfmt\|,\|Δsem\|\)≥0\.0767\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)\\\!\\geq\\\!0\.0767\(G3 passes\) and paired bootstrap CIs wholly above 1 \(Qwen CSR13\.0413\.04, CI\[10\.53,17\.65\]\[10\.53,17\.65\]; Mistral CSR8\.338\.33, CI\[6\.67,10\.34\]\[6\.67,10\.34\]\)\. Both cells therefore land in the*exploratory*bucket pending G5 and G6 enforcement\.BBQ\.Mistral×\\timesBBQ passes all implemented gates but is*ineligible \(F5\)*because the attribute family swaps demographic tokens while the answer key stays fixed, soΔattr\\Delta\_\{\\mathrm\{attr\}\}is not a clean construct flip\. Qwen×\\timesBBQ fails G2 \(unperturbed score0\.3400\.340vs\.1/31/3baseline is0\.7​pp0\.7\\,\\mathrm\{pp\}above, well under the2​SE≈6\.7​pp2\\,\\mathrm\{SE\}\\approx 6\.7\\,\\mathrm\{pp\}margin\) and additionally inherits the F5 concern; it falls in the*failed G2–G4*bucket under first\-match precedence\.CrowS\-Pairs\.Both cells are*not\_applicable*\(mapped to*ineligible*in paper\): the canonical PLL scorer readschoices, while the format / semantic perturbations editquestion; the perturbation does not reach the scored object \(F1\), and archetype is ambiguous under PLL \+ question\-edit\.\|Δsem\|=0\|\\Delta\_\{\\mathrm\{sem\}\}\|\\\!=\\\!0on both cells is consistent with F1\.ToxiGen\.Both cells pass G1–G4 but the scorer \(first\-token logprob overtoxic/non\-tokens\) is validated only on a 100\-item tokenizer\-coverage probe \(Appendix[B](https://arxiv.org/html/2607.02586#A2); mean coverage0\.8530\.853and0\.8570\.857,p5=0\.64p\_\{5\}=0\.64and0\.740\.74\)\. Absent a held\-out human\-judged validation, both ToxiGen cells are labelled*scorer\-unvalidated*\.XSTest\.The 10\-pattern refusal regex has not been separately audited for false\-negative rate\. Qwen×\\timesXSTest passes all implemented gates \(*scorer\-unvalidated*\); Mistral×\\timesXSTest fails G2 above\-baseline*and*is denominator\-dominated atmax⁡\(\|Δfmt\|,\|Δsem\|\)=0\.012<0\.02\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)\\\!=\\\!0\.012<0\.02\(G3 fails\), falling in*failed G2–G4*\.

The canonical 200\-item ToxiGen slice is class\-balanced, so the uniform and majority baselines coincide at0\.50\.5\. Gate decisions use full precision throughout; displayed values may round to the threshold\.

### Seed stability and placebo specificity \(PREREG E5, E6\)\.

Of the 9 cells with non\-degenerate CSR \(Qwen\-2\.5\-7B×\\timesBBQ fails G2 above\-baseline, so its seed CV is undefined\), per\-cell CSR coefficient of variation across seeds\{42,123,2024\}\\\{42,123,2024\\\}is≤0\.27\\leq 0\.27on all 9 \(smallest0\.0210\.021on Qwen×\\timesTruthfulQA; largest0\.2660\.266on Mistral×\\timesXSTest\)\. Seed instability does not drive any verdict\. The placebo specificity ratioCSRreal/CSRplacebo\\mathrm\{CSR\}\_\{\\mathrm\{real\}\}/\\mathrm\{CSR\}\_\{\\mathrm\{placebo\}\}is≥4\\geq 4on 4 of the 8 cells withCSRreal\>0\\mathrm\{CSR\}\_\{\\mathrm\{real\}\}\>0\(Qwen×\\timesXSTest5\.005\.00, Qwen×\\timesToxiGen4\.474\.47, Mistral×\\timesXSTest74\.2574\.25, Mistral×\\timesToxiGen5\.615\.61; median across all 8 cells is3\.983\.98\)\. Excluding the two BBQ cells whereCSRreal≈0\\mathrm\{CSR\}\_\{\\mathrm\{real\}\}\\\!\\approx\\\!0so the ratio is denominator\-noise dominated \(raw ratios0\.0170\.017and0\.1650\.165\), the ratio is≥4\\geq 4on 4 of 6 cells, median4\.744\.74\. This is positive but non\-universal support for contrast specificity\. Source data:results/canonical/seed\_variance\.csvandresults/canonical/csr\_placebo\.csv\.

## Appendix GVerdict\-Mapping Rule

The rawcell\_validity\.csvverdictcolumn encodes, per cell, the first condition that fires under the status\-assignment algorithm of §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\. The raw\-verdict tokens and the algorithm steps that emit them are:not\_applicable\(G1 uninstrumented, step 1\),inconclusive\(at least one of G2/G3/G4 fails, step 2\),uninterpretable\_scorer\(gates pass, scorer not validated, step 3\),attribute\_family\_mismatch\(gates pass and scorer validated, but archetype\-mismatched under CSR\-as\-defined, step 4\), andselective/non\_selective\(all of the above pass; CI wholly above / below 1; step 5 or 6\)\. Per §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4), the raw\-verdict tokens are emitted in the order they are checked; a reader with the raw CSV can reproduce every count in the paper by applying the table below mechanically\.

Table 5:Rawcell\_validity\.csvverdict→\\topaper bucket\. The raw\-verdict emission order matches §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)exactly, so no first\-match contention occurs at the mapping step\. A cell withverdict=selectiveis downgraded to*exploratory*in this submission because G5 and G6 remainproposed; it would upgrade to*confirmatory\-selective*once both gates areimplemented\.### Worked example: Mistral×\\timesXSTest\.

Raw verdictinconclusive; algorithm trace: step 1 fires G1? No—XSTest’s 10\-pattern refusal regex reads the generated text and our format perturbations do reach that path \(G1 is instrumented, though the scorer itself is unvalidated; that is step 3’s concern, not step 1’s\)\. Step 2: G2 fails \(unperturbed0\.565<0\.565<0\.500\+2​SE0\.500\+2\\,\\mathrm\{SE\}\)*and*G3 fails \(max⁡\(\|Δfmt\|,\|Δsem\|\)=0\.012<0\.02\\max\(\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|\\Delta\_\{\\mathrm\{sem\}\}\|\)\\\!=\\\!0\.012<0\.02\)\. Step 2 fires; raw verdictinconclusive; paper bucket*failed G2–G4*\. Step 3 \(scorer\-unvalidated\) would have fired had G2–G4 passed, but it did not\. The scorer\-unvalidated flag is still surfaced in the appendix eligibility notes as a secondary concern\.

### Worked example: Qwen×\\timesBBQ\.

Raw verdictinconclusive; step 1 does not fire \(BBQ option\-logprob scorer is fully instrumented\); step 2 fires because G2 fails \(0\.340<0\.340<0\.333\+2​SE0\.333\+2\\,\\mathrm\{SE\}\); bucket*failed G2–G4*\. The F5 archetype\-mismatch concern is benchmark\-level and would have fired at step 4 had G2–G4 passed \(as on Mistral×\\timesBBQ\)\.

### Cell\-levelgate\_\*columns vs\. paper G1–G6\.

The rawcell\_validity\.csvgate columns \(gate\_parseability,gate\_above\_baseline,gate\_denom\_not\_eps,gate\_ci\_estimable,gate\_n\_items\_ok,gate\_seed\_cv\_lt\_03\) are the empirical per\-cell instantiation of the methodology\-level G1–G6 in §[3](https://arxiv.org/html/2607.02586#S3), not a separate gate set\.gate\_parseabilityrealises the regex\-parseability half of G1; the silent\-no\-op half of G1 is audited offline against the canonical scorer’s prompt builder \(Table[7](https://arxiv.org/html/2607.02586#A9.T7)\), which is why G1 ships aspartialin Table[1](https://arxiv.org/html/2607.02586#S3.T1)\.gate\_above\_baselinerealises G2;gate\_denom\_not\_epsrealises G3;gate\_ci\_estimable\+gate\_n\_items\_ok\+gate\_seed\_cv\_lt\_03jointly realise G4 \(paired\-bootstrap feasibility, sample\-size guard, seed\-stability guard\)\. The CSV carries no direct column for G5 \(archetype disclosure\) or G6 \(repair\-regression\); theirproposedstatus is what downgrades a raw\-verdictselectiveto paper\-bucket*exploratory*under §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\.

## Appendix HCase\-Study Table: 10 Cells, Canonical Pipeline

Table[6](https://arxiv.org/html/2607.02586#A8.T6)is the per\-cell evidence behind the counts in §[6](https://arxiv.org/html/2607.02586#S6)/ Table[3](https://arxiv.org/html/2607.02586#S5.T3)\. All values are taken directly fromresults/canonical/csr\_canonical\.csv\(per\-family deltas, CSR point estimate\) andresults/canonical/cell\_validity\.csv\(paired\-bootstrap CI, raw verdict\)\. Gate decisions use full precision throughout\.

Table 6:Canonical\-pipeline case\-study numbers, Qwen\-2\.5\-7B and Mistral\-7B\-v0\.3\. Raw verdicts are the literalcell\_validity\.csvvalues; the paper\-bucket mapping is Appendix[G](https://arxiv.org/html/2607.02586#A7)\. Baselines: TruthfulQA MC11/41/4, BBQ1/31/3, ToxiGen0\.50\.5, XSTest1/21/2\.![Refer to caption](https://arxiv.org/html/2607.02586v1/x2.png)Figure 3:Legacy\-pipeline CSR \(x\-axis\) versus canonical\-pipeline CSR \(y\-axis\) per cell, on*symlog*axes \(linear in\[−1,1\]\[\-1,1\], log outside\) so both the near\-zero BBQ cells and the Mistral×\\timesXSTest outlier at CSR≈85\\approx 85are readable on one panel\. Points on the dotted identity line are unchanged by repair; points far from it reflect the joint effect of F1–F3 fixes\. Seven of ten cells change verdict \(selective↔\\leftrightarrownon\-selective or denominator\-dominated\) under the repaired pipeline\. Cell codes: Q==Qwen\-2\.5\-7B, M==Mistral\-7B\-v0\.3; TQ==TruthfulQA, BB==BBQ, TX==ToxiGen, CP==CrowS\-Pairs, XS==XSTest\.![Refer to caption](https://arxiv.org/html/2607.02586v1/x3.png)Figure 4:Per\-cell\|Δfmt\|\|\\Delta\_\{\\mathrm\{fmt\}\}\|,\|Δsem\|\|\\Delta\_\{\\mathrm\{sem\}\}\|,\|Δattr\|\|\\Delta\_\{\\mathrm\{attr\}\}\|under the canonical pipeline\. Cells where\|Δfmt\|\|\\Delta\_\{\\mathrm\{fmt\}\}\|or\|Δsem\|<0\.02\|\\Delta\_\{\\mathrm\{sem\}\}\|<0\.02drive the denominator below theε\\varepsilonfloor and trigger G3\.\|Δsem\|=0\|\\Delta\_\{\\mathrm\{sem\}\}\|=0on CrowS\-Pairs reflects F1\-type silent no\-ops on the question field, which the PLL scorer does not read\.
## Appendix IAuto\-Diff Audit Numbers

Silent\-no\-op rates of format perturbations per benchmark, measured \(i\) against the legacy renderer \(evaluate\.format\_mcq\_prompt\) and \(ii\) against the canonical scorer’s prompt builder \(scoring\_canonical\.\_format\_mcq\_prompt\)\. The gap between the two columns is the core F1 evidence\.

Table 7:Silent\-no\-op rate \(%\) for the format family\. Column L==against the legacy renderer; column C==against the canonical scorer’s prompt builder\. The F1 regression shows in the C column for MCQ benchmarks \(100% no\-op\)\.†CrowS under PLL: 5 of 6 perturbation methods do not reach the scorer because the scorer readschoicesand the perturbations editquestion\. The 83% figure is the effective scorer\-consumed no\-op rate\.

## Appendix JCase\-Level Audit Trail

This appendix is the worked instance of the minimal chronology template \(Box 1\): each entry below carries a bug ID and failure class \(field 1\), provenance \(field 5: F3c is flagged*self\-introduced*\), the fix and its regression test \(field 6: tests shipped for F1, F3c, F4\), and residual status \(field 8:fixed/unfixed, acknowledged\)\. For space, the per\-cell discovery date, affected\-cell list, and before/after deltas \(fields 2–4, 7\) are condensed here and given in full in the release bundle’s running audit log \(Appendix[K](https://arxiv.org/html/2607.02586#A11)\); the F4 entry quotes its before/after delta inline as an example of field 7\. All items listed asfixedare repaired in the canonical pipeline that produced every number in this paper; the trail exists so a reader can audit our audit, not because the canonical pipeline ships with unresolved bugs\. Three items remainunfixed, acknowledgedand are disclosed as active limitations\.

Entries are chronological, matched to the failure\-mode cases in §[4](https://arxiv.org/html/2607.02586#S4)\.

1. 1\.F1a \(fixed\):format\_change\_labelswroteitem\[‘label\_style’\]; prompt builder did not read it\. Fix: parametriseformat\_mcq\_prompt\(label\_style\)\.
2. 2\.F1b \(fixed\): Fix landed only inevaluate\.py;scoring\_canonical\.py::\_format\_mcq\_promptremained hard\-coded\. Fix: parametrise the canonical builder\.
3. 3\.F1c \(fixed\): Rule\-basedsemantic\_sentence\_restructurehad no\-op rates of 57–93% due to trigger\-string mismatch on Q&A benchmarks\. Fix: addparaphrase\_template,synonym\_substitution\.
4. 4\.F2a \(partially fixed\): Qwen×\\timesTruthfulQA legacy parseability 0\.44\. Fix: canonical option\-logprob scoring sets parseability to 1\.0; legacy numbers are withdrawn\.
5. 5\.F3a \(disclosed, not repaired in data\): CrowS\-Pairs legacy accuracy inverts fairness interpretation\. Canonical pipeline switches to PLL; CrowS legacy numbers are withdrawn\.
6. 6\.F3b \(fixed\): TruthfulQA MC2 correct\-first ordering artefact\. Fix: MC1 \+ per\-item shuffle\.
7. 7\.F3c \(fixed, self\-introduced\): Top\-kktruncation caused majority\-class predictions on ToxiGen\. Fix: targeted\-token first\-token log\-probabilities\.
8. 8\.F4 \(fixed\): Broken bootstrap pairing\. Fix: item\-level paired bootstrap\. Out\-of\-panel legacy\-Llama atcsr≈9\.5\\textsc\{csr\}\\approx 9\.5showed CI width 26\.4 under unpaired resampling\.
9. 9\.F5 \(disclosed\): BBQ / CrowS archetype mismatch under CSR\. Not fixed; G3 gate flags denominator\-dominated cells\.
10. 10\.Shared\-prefix collapse on MCQ letter tokens\(unfixed, acknowledged\): first token of “\(A\)”, “\(B\)” can be identical\.
11. 11\.TEXTCLFadd\_instructioncontent contamination on ToxiGen\(unfixed, acknowledged\)\.

### Containment of unfixed issues\.

The remaining unfixed items do not alter any*confirmatory*verdict in this paper because no confirmatory verdict is issued: all affected cells remain non\-confirmatory under the hierarchical precedence of §[3](https://arxiv.org/html/2607.02586#S3.SS0.SSS0.Px4)\. The unfixed items therefore further motivate—rather than undercut—the paper’s withholding rule\.

## Appendix KPreregistration and Release

Code and data\.A release bundle will accompany the de\-anonymised version of the paper, containing the canonical pipeline \(all F1–F4 fixes landed\), per\-item predictions for all 10 cells, the preregistrationPREREG\.md\(commite6e686b8c0b02dd78d1495eb9c4a94882cce5cce, amended 2026\-04\-22\), running audit logs, and regression tests for F1, F3c, and F4\. The bundle is not part of this submission and the URL is withheld for blind review\. Hardware and software versions are in Appendix[B](https://arxiv.org/html/2607.02586#A2)\.

Deviations from preregistration\.The 2026\-04\-19 prereg was amended on 2026\-04\-22 \(PREREG §10\)\. The substantive deviations:

1. \(a\)Llama\-3\.1\-8B was dropped between prereg and canonical rerun, reducing the panel to two models\.
2. \(b\)F3b and F3c were discovered after the prereg SHA; all analyses after their discovery are formally exploratory\.
3. \(c\)The E3 human audit was de\-scoped; the metric was renamed “Construct Sensitivity Ratio”→\\to“Contrast Selectivity Ratio” under the pre\-committed PREREG §7 fallback\. Construct\-validity language is retracted from any claim that would have required E3 validation; artifactsresults/audit/,ETHICS\.md, andannotator\_consent\.mdare removed from the PREREG §9 manifest\.
4. \(d\)CrowS\-Pairs is marked*not\_applicable*at the cell level \(the PLL scorer readschoiceswhile our perturbations editquestion\); no selective/non\_selective verdict is issued for CrowS\. See Appendix[F](https://arxiv.org/html/2607.02586#A6)\.
5. \(e\)The legacy unpaired bootstrap \(src/ablations\.py::bootstrap\_csr\) was removed; claim\-bearing CIs come fromsrc/analyze\_canonical\.py::paired\_bootstrap\_csr\(item\-level paired, PREREG §6\.1\)\. See §[4](https://arxiv.org/html/2607.02586#S4)\(F4\)\.
6. \(f\)Template robustness \(PREREG §6\.3\) was scoped to TruthfulQA and ToxiGen and reported as secondary, not as a gate for the other three benchmarks whose scoring paths have no prompt\-template axis\.
7. \(g\)Thecell\_validity\.csvgate set was tightened from 3 columns \(2026\-04\-21\) to the full 6\-column G1–G4 realisation; mapping to paper G1–G6 is in Appendix[G](https://arxiv.org/html/2607.02586#A7)\.

Seed stability and placebo specificity \(E5, E6\)\.Reported in Appendix[F](https://arxiv.org/html/2607.02586#A6)\.

Similar Articles

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

arXiv cs.AI

This paper identifies distribution shift and scale constraints as critical failure modes for statistical contamination detection methods in LLM benchmark auditing. Evaluating three paradigms across 27 models reveals only 199 correct outcomes out of 335 evaluations, indicating a systematic reliability gap that prevents these methods from replacing transparent data provenance.

Adaptive auditing of AI systems with anytime-valid guarantees

arXiv cs.AI

This paper introduces a statistical framework for adaptively auditing AI systems using Safe Anytime-Valid Inference (SAVI) to draw rigorous conclusions with limited data. It proposes a 'testing by betting' approach to validate model robustness while controlling type-I errors during adaptive sampling.