What Was That Again? Certified Robustness for Automatic Speech Recognition

arXiv cs.LG Papers

Summary

This paper presents a certification-inspired mechanism for automatic speech recognition that uses a dual-gate diagnostic pipeline (Two-Sided Atomic Audit and Rank-Based Tournament) to provide certified robustness and achieve up to a 55% relative reduction in word error rate across diverse architectures.

arXiv:2606.27698v1 Announce Type: new Abstract: Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription. We demonstrate that employing a certification-inspired mechanism can significantly decrease WER, increase recall, and decrease the Spearman correlation between confidence and WER. We achieve this through a dual-gate diagnostic pipeline: a Two-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank-Based Tournament that selects the winning sequence. Our evaluations across four diverse architectures demonstrate up to a 55% relative reduction in Word Error Rate, while also providing granular word- and sentence-level certifications to enhance acoustic security.
Original Article
View Cached Full Text

Cached at: 06/29/26, 05:25 AM

# What Was That Again? Certified Robustness for Automatic Speech Recognition
Source: [https://arxiv.org/html/2606.27698](https://arxiv.org/html/2606.27698)
Andrew C\. Cullen University of Melbourne andrew\.cullen@unimelb\.edu\.au &Neil Marchant University of Melbourne &Jiani Xie University of Melbourne Paul Montague DST Group, Adelaide &Benjamin I\. P\. Rubinstein University of Melbourne

###### Abstract

Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations\. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle knowledge of the true transcription\. We demonstrate that employing a certification\-inspired mechanism can significantly decrease WER, increase recall, and decrease the Spearman correlation between confidence and WER\. We achieve this through a dual\-gate diagnostic pipeline: a Two\-Sided Atomic Audit that accumulates statistical wealth to certify both token existence and adversarial exclusion, and a Rank\-Based Tournament that selects the winning sequence\. Our evaluations across four diverse architectures demonstrate up to a55%55\\%relative reduction in Word Error Rate, while also providing granular word\- and sentence\-level certifications to enhance acoustic security\.

## 1Introduction

For all their transformative utility, neural networks in speech processing remain notoriously sensitive to infinitesimal input perturbations\. These perturbations—known as*adversarial examples*—demonstrate that model decision boundaries often lack the semantic alignment required for safety\-critical deployments\. While many approaches have been proposed to mitigate the risk associated with these adversarial examples, the most conceptually promising solution is a family of guarantees producing*Certified Robustness*\(Lecuyeret al\.,[2019](https://arxiv.org/html/2606.27698#bib.bib6); Cohenet al\.,[2019](https://arxiv.org/html/2606.27698#bib.bib5); Cullenet al\.,[2022](https://arxiv.org/html/2606.27698#bib.bib7)\)\.

These robustness mechanisms are rigorous frameworks for mathematically guaranteeing that a model’s prediction remains invariant to perturbations\. In the case of a classifierFF, the guarantee is oftentimes a radiusrrsuch that we can guaranteeF​\(x\)=F​\(x′\)F\(x\)=F\(x^\{\\prime\}\)for allx′x^\{\\prime\}in theℓp\\ell\_\{p\}\-ballBp​\(x,r\)=\{x′:‖x′−x‖p≤r\}B\_\{p\}\(x,r\)=\\\{x^\{\\prime\}:\\\|x^\{\\prime\}\-x\\\|\_\{p\}\\leq r\\\}\. While several approaches have been proposed, Randomised Smoothing \(RS\) has proven particularly effective, as it introduces no additional architectural or infrastructure burdens on the modelFF\(Cohenet al\.,[2019](https://arxiv.org/html/2606.27698#bib.bib5)\)\.

However, extending the guarantees of CR to non\-categorical, high\-dimensional sequence outputs—such as those encountered in Natural Language Processing or Automatic Speech Recognition \(ASR\)—presents a combinatorial challenge\(Huanget al\.,[2023](https://arxiv.org/html/2606.27698#bib.bib62),[2024](https://arxiv.org/html/2606.27698#bib.bib64)\)\. In ASR, if sentences are treated as discrete classes, then the output space explodes in size as noise levels increase\. Traditional RS workflows, which rely upon finding a sentence representing the*majority class*fail, because the probability mass of any single transcription collapses under such conditions\(Olivier,[2023](https://arxiv.org/html/2606.27698#bib.bib63)\)\. Resolving this limitation as crucial, as it is incredibly difficult to audit deployed acoustic models to ensure that they are producing safe outputs when exposed to untrusted inputs\.

This failure mode reveals a fundamental dual\-layer opportunity for sequence certification, in that one must not only certify the*atomic content*, the specific words present in the signal, but also their*structural arrangement*, representing the relative ordering and grammatical coherence of these atomic elements\. Prior works in this space have attempted to address this through numerically expensive sequence alignment; however the resulting outputs have typically focused upon either providing high\-confidence word inclusions at the cost of structural fragmentation, or attempting holistic sequence certification that is hamstrung by the problem space’s inherent combinatorial scaling\.

In this work, we propose a novel framework to address both forms of certification through E\-value tournaments to produceRank\-Based Sentence Certification\(Shafer and Vovk,[2019](https://arxiv.org/html/2606.27698#bib.bib15); Ramdaset al\.,[2023](https://arxiv.org/html/2606.27698#bib.bib16)\)\. Our approach bridges the gap between atomic and structural guarantees through the use of a dual\-gate certification pipeline, that eschews the need for sentence alignment, resulting in a stable certification recall \(40\.5%–90\.3%\) even at high noise levels, where alternate baselines collapse to<1%<1\\%\. Our new approach yields non\-vacuous safety radii, and significant WER improvements on standard benchmarks, where traditional smoothing fails to certify\. Our contributions are:

1. 1\.Anytime\-valid Certification for Sequence\-to\-Sequence tasks: We demonstrate how Ville’s Inequality and E\-values can be deployed to provide valid safety radii at any point in the sampling process\. This provides a framework for reducing the computational cost of certifications\.
2. 2\.Dual Gate and Two\-Sided Adversarial Certifications: We introduce a dual\-gate pipeline that performs a two\-sided audit—simultaneously proving the existence of safe tokens and the rigorous exclusion of adversarial hallucinations\. This mechanism allows the system to prune the search space before the final rank\-based tournament, which supplements the final certification with a measure of robustness to sentence\-level changes\.
3. 3\.Dynamic ASR Auditing: Demonstrating how our dual\-phase certifications can be employed to audit the performance of ASR models on a Part\-of\-Speech framework on LibriSpeech and Common Voice, even in extreme noise\.

## 2Related Work

#### ASR

Modern Automatic Speech Recognition \(ASR\) systems can be categorized in terms of their architecture\. The most commonly employed approaches include CTC\-based architectures such as DeepSpeechHannunet al\.\([2014](https://arxiv.org/html/2606.27698#bib.bib43)\); self\-supervised Transformer models such as Wav2Vec 2\.0Baevskiet al\.\([2020a](https://arxiv.org/html/2606.27698#bib.bib44)\); and large\-scale encoder\-decoders such as WhisperRadfordet al\.\([2023a](https://arxiv.org/html/2606.27698#bib.bib45)\)\. However, in each of these, the model is susceptible to adversarial perturbations that can substantially degrade transcription accuracyCarlini and Wagner \([2018](https://arxiv.org/html/2606.27698#bib.bib46)\); Qinet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib47)\); Olivier and Raj \([2022](https://arxiv.org/html/2606.27698#bib.bib48)\)\.

While these perturbations exist within the long, established, mature history of adversarial manipulations against computer vision tasks\(Goodfellowet al\.,[2014](https://arxiv.org/html/2606.27698#bib.bib2); Madryet al\.,[2018](https://arxiv.org/html/2606.27698#bib.bib3); Cullenet al\.,[2024](https://arxiv.org/html/2606.27698#bib.bib58)\)—manipulating models to change predictions—the audio domain presents distinct challenges\. Speech is a time\-varying physical signal in which small waveform changes do not necessarily correspond to small changes in perceived sound, and commonly usedℓp\\ell\_\{p\}\-norm distances can correlate poorly with human judgment of audibilityQinet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib47)\); Schönherret al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib49)\)\.

With this said, the most common metrics for assessing acoustic systems are the Signal\-to\-Noise Ratio \(SNR\) and the Word Error Rate \(WER\)\. The SNR serves as the fundamental measure of input perturbation, representing the log\-ratio of signal amplitude to noise amplitude by

S​N​Rd​B=20​log10⁡\(‖x‖p‖ϵ‖p\)SNR\_\{dB\}=20\\log\_\{10\}\\left\(\\frac\{\\\|x\\\|\_\{p\}\}\{\\\|\\epsilon\\\|\_\{p\}\}\\right\)\(1\)wherexxis the clean audio signal,ϵ\\epsilonis the additive perturbation, andppis the norm, which is typicallyp∈\{2,∞\}p\\in\\\{2,\\infty\\\}\. In adversarial contexts, attackers seek to minimize the perturbation magnitude \(maximizing SNR\) while inducing catastrophic failures in the model’s output\.

Conversely, the WER serves as the standard metric for transcription quality, calculated as the normalized Levenshtein distance

W​E​R=S\+D\+INWER=\\frac\{S\+D\+I\}\{N\}\(2\)whereS,D,S,D,andIIrepresent the number of substitutions, deletions, and insertions required to align the hypothesis with a ground truth of lengthNN\. While it is ubiquitous in acoustic \(and textual\) contexts, it must be emphasized that the WER can often be distorted by the failure modes commonly seen within such systems\. In recognition of the issues inherent to these metrics, alternative approaches have considered acoustic performance through a psychoacoustic lensSzurley and Kolter \([2019](https://arxiv.org/html/2606.27698#bib.bib53)\); Sunet al\.\([2024](https://arxiv.org/html/2606.27698#bib.bib50)\), although it has been noted that it is fundamentally difficult to define consistent notions of imperceptibilityHussainet al\.\([2021](https://arxiv.org/html/2606.27698#bib.bib54)\)\.

#### Certified Robustness

While originally built upon foundations of Differential Privacy, the current state of the art in RS applies the Neyman\-Pearson lemma\. At their core, all RS based certifications transform a modelffinto a smoothed counterpartgg, subject to provableℓp\\ell\_\{p\}margin guarantees\. As established byCohenet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib5)\), for a noise levelσ\\sigma, the radiusrris a function of the probabilitypAp\_\{A\}of the*most likely class*111The original formulation ofCohenet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib5)\)was derived in terms of the two most likely class probabilitiespAp\_\{A\}andpBp\_\{B\}, however nearly all implementations reduce to the variant described in this paper\.\. To achieve this, such certifications employ an independent two\-phase approach, where Phase I employs an initial batch to establish the target class, while Phase II repeatedly samples under noise by way ofpA=ℙϵ∼𝒩​\(0,σ2​I\)​\[f​\(x\+ϵ\)=cA\]p\_\{A\}=\\mathbb\{P\}\_\{\\epsilon\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\}\[f\(x\+\\epsilon\)=c\_\{A\}\]\. This probability denotes the probability that the smoothed classifier predicts the target classcAc\_\{A\}under additive Gaussian noise\. This ultimately yields a certification formulation ofr=σ​Φ−1​\(pA\)\.r=\\sigma\\Phi^\{\-1\}\(p\_\{A\}\)\.

WhilepAp\_\{A\}represents the theoretical expectation over the noise distribution, practical calculation of this quantity is impossible\. While this can be practically estimated through Monte\-Carlo sampling, the need to construct a conservative certification—to control the risk of the certification being violated—requires obtaining a high\-probability lower boundpA¯\\underline\{p\_\{A\}\}such thatP​\(pA<pA¯\)≤αP\(p\_\{A\}<\\underline\{p\_\{A\}\}\)\\leq\\alpha, which is typically estimated through the Clopper\-Pearson mechanism\(Clopper and Pearson,[1934](https://arxiv.org/html/2606.27698#bib.bib9)\)\.

The need to produce tight lower bounds is the primary driver of RS’s temporal cost, which typically requires tens of thousands of samples\.This is not to say that all inputs require such a volume of samples, but the frequentist approach of Clopper\-Pearson introduces apeeking problem: if a practitioner monitors the empirical mean and stops sampling once an online estimate looks appropriate for certification, the Type\-I error rate is no longer bounded byα\\alpha\(Johariet al\.,[2017](https://arxiv.org/html/2606.27698#bib.bib42)\)as peeking introduces a multiple comparison problem\. Thus setting the number of samplesnnis an important\-but\-wasteful part of RS—setting it too high results in inputs far from the decision boundary requiring thousands of redundant model calls; while setting it too low it may lead to a total failure to certify\.

As an alternative to the computational limitations of frequentist frameworks, E\-values provide an expectation\-constrained hypothesis testing measure that is immune to the peeking problem\. An E\-value for the null hypothesisH0H\_\{0\}is a non\-negative random variableEEsuch that𝔼H0​\[E\]≤1\\mathbb\{E\}\_\{H\_\{0\}\}\[E\]\\leq 1\. For the purposes of controlling the Type\-I error associated with certifications,EEcan be viewed as a multiplier on wealth in a fair game\. While such a context renders wealth a martingale, expected wealth cannot grow under the null hypothesis\. Define the accumulated wealthWt=∏i=1tEiW\_\{t\}=\\prod\_\{i=1\}^\{t\}E\_\{i\}\. Crucially, this wealth\-testing tramework is inherently anytime\-valid by Ville’s inequality\(Ville,[1939](https://arxiv.org/html/2606.27698#bib.bib11); Doob,[1940](https://arxiv.org/html/2606.27698#bib.bib10)\), which states that

P\(∃t:Wt≥1α\)≤α,P\\left\(\\exists t:W\_\{t\}\\geq\\frac\{1\}\{\\alpha\}\\right\)\\leq\\alpha\\kern 5\.0pt,\(3\)thus allowing for optimal stopping with zero penalty\. As long as the accumulated wealthWtW\_\{t\}surpasses1/α1/\\alpha, we can be confident that the probability of a Type\-I error is strictly bounded byα\\alpha\.

The application of E\-values to RS was pioneered byVoráček \([2024](https://arxiv.org/html/2606.27698#bib.bib8)\); however, their approach is rooted in classical hypothesis testing, rather than standard certification practices\. Under the Voracek approach—as defined in their Definition 3\.1—the tested null hypothesis is restricted to a pre\-defined radiusr0r\_\{0\}, leading to evaluation of the null hypothesis

H0:pA≤Φ​\(r0σ\)\.H\_\{0\}:p\_\{A\}\\leq\\Phi\\left\(\\frac\{r\_\{0\}\}\{\\sigma\}\\right\)\\kern 5\.0pt\.\(4\)While this yields a binary certification of whether a sample is robust to a radius ofr0r\_\{0\}, it fails to identify the*maximum certified radius*\. Naively extending this framework would require running an infinite number of independent hypothesis tests, for each possibler0∈\[0,∞\)r\_\{0\}\\in\[0,\\infty\)\. As such testing is impossible, we instead will employ the Method of Mixtures to provide E\-value based certifications, allowing infinite hypothesis testing in finite computational time\.

#### Certifications Beyond Classifiers

The baseline for sequence certifications, as pioneered byOlivier and Raj \([2021](https://arxiv.org/html/2606.27698#bib.bib57)\), relies on performing a task known as multiple sequence alignment to reduce the high\-dimensional transcription space into a structured voting network\. To achieve this, the system maintains a consensus backboneℬ=\{S1,S2,…,SL\}\\mathcal\{B\}=\\\{S\_\{1\},S\_\{2\},\\ldots,S\_\{L\}\\\}, where each elementSiS\_\{i\}is a dictionary that acts as a frequency counter of words observed at that relative temporal position\. For each new sample drawn under noiseϵ\\epsilon, the algorithm performs a string\-to\-graph alignment via a Word Transition Network to find the optimal mapping between the new tokens and the existing dictionaries\. If a token cannot be aligned to an existing slot within a defined similarity threshold, the backbone is expanded, creating a new slot to represent this new element\.

However, as we will show within this work, such an approach is fragile to the very kinds of low SNR regimes that would likely be seen within an adversarial ASR context\. Under these conditions, ASR models produce high variance outputs typified by small but frequent structural changes \(manifesting as repeated syllables, misspellings, or substituted words\), which rapidly induce failures within the sequence alignment algorithm\. When the alignment matcher fails to recognize two semantically identical tokens as the same slot, it defaults to creating a new insertion slot\. This creates a catastrophic feedback loop: as the number of slotsLLincreases, the Bonferroni\-corrected budgetα/L\\alpha/Lbecomes increasingly stringent, making it statistically impossible to certify any individual slot\. Under testing, 18 ground\-truth words expanded to over thousands of slots after only a few thousand samples\. The resulting sequence becomes an interleaved concatenation of fragments from different samples—e\.g\., if Sample 1 isA B Cand Sample 2 isD E F, the misaligned output becomesA D B E C F\)\. This structural collapse leads to rapid growth in the WER and vacuous certifications\.

## 3Problem Definition

We aim to extend certified robustness to sequence\-valued predictions by constructing a*certified transcription*Y^\\hat\{Y\}of an audio signalxx, such thatY^\\hat\{Y\}is invariant to adversarial perturbations within anℓ2\\ell\_\{2\}radiusRRin the input space\. Formally, let𝒳\\mathcal\{X\}denote the global universe of tokens, corresponding to the model’s vocabulary, and let𝒳⋆=⋃L=0∞𝒳L\\mathcal\{X\}^\{\\star\}=\\bigcup\_\{L=0\}^\{\\infty\}\\mathcal\{X\}^\{L\}denote the set of all finite\-length sequences over𝒳\\mathcal\{X\}, where𝒳0=\{∅\}\\mathcal\{X\}^\{0\}=\\\{\\emptyset\\\}\. The base ASR modelf:ℝd→𝒳⋆f:\\mathbb\{R\}^\{d\}\\to\\mathcal\{X\}^\{\\star\}thus maps an input signal to a transcription sequencey=\(w1,…,wL\)y=\(w\_\{1\},\\ldots,w\_\{L\}\)wherewi∈𝒳w\_\{i\}\\in\\mathcal\{X\}\. FollowingCohenet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib5)\), a standard smoothed predictorF:ℝd→𝒳⋆F:\\mathbb\{R\}^\{d\}\\to\\mathcal\{X\}^\{\\star\}would ideally select the most probable transcription under Gaussian noise:

F​\(x\)=arg⁡maxy∈𝒳⋆⁡ℙ​\(f​\(x\+ϵ\)=y\),ϵ∼𝒩​\(0,σ2​I\)\.F\(x\)=\\arg\\max\_\{y\\in\\mathcal\{X\}^\{\\star\}\}\\mathbb\{P\}\\big\(f\(x\+\\epsilon\)=y\\big\),\\quad\\epsilon\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\\kern 5\.0pt\.\(5\)However, in practice, the sequence space is so fragmented that the probability of any single transcription collapses toward zero under significant noise\. As such, we define the smoothed predictor through a generalized hierarchical aggregatorG​\(⋅\)G\(\\cdot\)

F​\(x\)=G​\(f​\(x\+ϵ\)\),ϵ∼𝒩​\(0,σ2​I\),F\(x\)=G\\big\(f\(x\+\\epsilon\)\\big\),\\quad\\epsilon\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)\\kern 5\.0pt,\(6\)the form of which we will now introduce, with the aim of constructing the pair\(Y^,R\)\(\\hat\{Y\},R\)such that

ℙ​\(∃δ∈ℝd​s\.t\.​‖δ‖2<R​and​F​\(x\+δ\)≠Y^\)≤α,\\mathbb\{P\}\\Big\(\\exists\\,\\delta\\in\\mathbb\{R\}^\{d\}\\;\\;\\text\{s\.t\.\}\\;\\;\\\|\\delta\\\|\_\{2\}<R\\;\\;\\text\{and\}\\;\\;F\(x\+\\delta\)\\neq\\hat\{Y\}\\Big\)\\leq\\alpha\\kern 5\.0pt,\(7\)given some global error budgetα∈\(0,1\)\\alpha\\in\(0,1\)\. To achieve this, we decomposeG​\(⋅\)G\(\\cdot\)into two distinct stages with a partitioned error budgetα=αa​t​o​m​i​c\+αt​o​u​r​n\\alpha=\\alpha\_\{atomic\}\+\\alpha\_\{tourn\}\.

1. 1\.Token\-Level Atomic Certification: We first identify a finiteCandidate Vocabulary𝒱⊂𝒳\\mathcal\{V\}\\subset\\mathcal\{X\}through an initial discovery phase, and then certify words based upon this Candidate Vocabulary\. Specifically, ifw∈Vw\\in V, then with probability at least1−αa​t​o​m​i​c1\-\\alpha\_\{atomic\}, the true probability ofwwappearing in a noisy transcriptionp​\(w\)\>0\.5p\(w\)\>0\.5\. Conversely, ifw∉𝒱w\\notin\\mathcal\{V\}andwwis not ambiguous, we certifyp​\(w\)<0\.5p\(w\)<0\.5\.
2. 2\.Sentence\-Level Structural Certificate: A sequential tournament is conducted among a set of candidate transcriptions𝒞⊂𝒱⋆\\mathcal\{C\}\\subset\\mathcal\{V\}^\{\\star\}, which have been filtered by the atomic gate to ensure structural and statistical validity\.

The final system radiusR=min⁡\(Ra​t​o​m​i​c,Rt​o​u​r​n\)R=\\min\(R\_\{atomic\},R\_\{tourn\}\)is the minimum radius required to violate either of these hierarchical guarantees\.

### 3\.1Atomic Certifications

The atomic gate identifies and verifies constituent tokens in two phases: discovery, in which the candidate vocabulary𝒱\\mathcal\{V\}is identified; and audit, in which certifications are performed\.

#### Discovery Phase

We drawN1N\_\{1\}i\.i\.d\. samplesYi=f​\(x\+ϵi\)Y\_\{i\}=f\(x\+\\epsilon\_\{i\}\), whereϵi∼𝒩​\(0,σ2​I\)\\epsilon\_\{i\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)to identify the Candidate Vocabulary𝒱⊂𝒳\\mathcal\{V\}\\subset\\mathcal\{X\}by𝒱=⋃i=1N1\{w:w∈Yi\}\\mathcal\{V\}=\\bigcup\_\{i=1\}^\{N\_\{1\}\}\\\{w:w\\in Y\_\{i\}\\\}\. The intent of this stage is to prune the search space to tokens with non\-trivial support in the local noise distribution\. These samples are subsequently discarded to ensure statistical independence\. However, it must be noted that any word that is not present in the firstN1N\_\{1\}samples will be unable to be verified—even if it would otherwise present as a high frequency inclusion—and as such will be filtered out from all subsequent stages\.

#### Audit Phase

Following the discovery of the candidate vocabulary𝒱\\mathcal\{V\}, we verify the existence of each token using a fresh stream ofNCN\_\{C\}samples\. For any tokenw∈𝒱w\\in\\mathcal\{V\}, letpw=ℙ​\(w∈f​\(x\+ϵ\)\)p\_\{w\}=\\mathbb\{P\}\(w\\in f\(x\+\\epsilon\)\)denote its marginal probability of appearance\. We accumulate two\-sided wealth via non\-negative martingalesEp​o​s​\(w\)E\_\{pos\}\(w\)andEn​e​g​\(w\)E\_\{neg\}\(w\)to test the null hypothesesHp​o​s:pw≤0\.5H\_\{pos\}:p\_\{w\}\\leq 0\.5andHn​e​g:pw≥0\.5H\_\{neg\}:p\_\{w\}\\geq 0\.5respectively\. These hypotheses will allow us to comprehensively certify the inclusion or exclusion within the observation set\.

To achieve this, for each noisy samplett, letWw,t∈\{0,1\}W\_\{w,t\}\\in\\\{0,1\\\}be the indicator that tokenwwis present in the transcription\. To ensure numerical stability and support high\-throughput GPU vectorization, we perform wealth compounding in log\-space\. For a betting parameterλ∈\(0,2\]\\lambda\\in\(0,2\], the accumulated wealth at timeTTis defined as

ln⁡Ep​o​s,T​\(w\)=∑t=1Tln⁡\(1\+λ​\(Ww,t−0\.5\)\)l​n​En​e​g,T​\(w\)=∑t=1Tln⁡\(1\+λ​\(0\.5−Ww,t\)\),\\ln E\_\{pos,T\}\(w\)=\\sum\_\{t=1\}^\{T\}\\ln\\big\(1\+\\lambda\(W\_\{w,t\}\-0\.5\)\\big\)\\qquad lnE\_\{neg,T\}\(w\)=\\sum\_\{t=1\}^\{T\}\\ln\\big\(1\+\\lambda\(0\.5\-W\_\{w,t\}\)\\big\)\\kern 5\.0pt,\(8\)whereEi,0=1E\_\{i,0\}=1fori∈\{p​o​s,n​e​g\}i\\in\\\{pos,neg\\\}\. Under the null hypothesispw=0\.5p\_\{w\}=0\.5, the expectation of the multiplierEt=1\+λ​\(Ww,t−0\.5\)E\_\{t\}=1\+\\lambda\(W\_\{w,t\}\-0\.5\)should be exactly11, renderingETE\_\{T\}a martingale\.

As was discussed in Section[2](https://arxiv.org/html/2606.27698#S2), by Ville’s inequality\(Ville,[1939](https://arxiv.org/html/2606.27698#bib.bib11); Doob,[1940](https://arxiv.org/html/2606.27698#bib.bib10)\), the probability that the wealth ever exceeds the safety threshold is bounded throughP\(∃T:ET≥1/α\)≤αP\(\\exists T:E\_\{T\}\\geq 1/\\alpha\)\\leq\\alpha\. We thus define theCertified Vocabulary𝒱​c​e​r​t\\mathcal\{V\}\{cert\}andExcluded Vocabulary𝒱​e​x​c​l\\mathcal\{V\}\{excl\}as the sets of tokens whose wealth crosses the significance threshold1/αa​t​o​m​i​c1/\\alpha\_\{atomic\}\.

#### Confidence Sequences and Certifications

To translate these measures of statistical wealth into certified radii, we must identify the set of success probabilitiesppthat are consistent with the observed evidence\. Under the framework ofConfidence Sequences\(Ramdaset al\.,[2023](https://arxiv.org/html/2606.27698#bib.bib16)\), we obtain anytime\-valid bounds by inverting the E\-value process\. For tokens in𝒱c​e​r​t\\mathcal\{V\}\_\{cert\}or𝒱e​x​c​l\\mathcal\{V\}\_\{excl\}, the anytime\-valid bounds are the extrame probabilities consistent with the wealth

p¯w,T=inf\{p≥0\.5:LT​\(p\)<αa​t​o​m​i​c−1\},p¯w,T=sup\{p≤0\.5:LT​\(p\)<αa​t​o​m​i​c−1\}\\underline\{p\}\_\{w,T\}=\\inf\\\{p\\geq 0\.5:L\_\{T\}\(p\)<\\alpha\_\{atomic\}^\{\-1\}\\\},\\quad\\overline\{p\}\_\{w,T\}=\\sup\\\{p\\leq 0\.5:L\_\{T\}\(p\)<\\alpha\_\{atomic\}^\{\-1\}\\\}\(9\)whereLT​\(p\)=∏t=1TBernoulli​\(Ww,t;p\)Bernoulli​\(Ww,t;0\.5\)L\_\{T\}\(p\)=\\prod\_\{t=1\}^\{T\}\\frac\{\\text\{Bernoulli\}\(W\_\{w,t\};p\)\}\{\\text\{Bernoulli\}\(W\_\{w,t\};0\.5\)\}is the likelihood ratio martingale, andp¯w\\underline\{p\}\_\{w\}andp¯w\\overline\{p\}\_\{w\}are the anytime\-valid lower and upper bounds respectively\. FollowingCohenet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib5)\), we map these probability bounds to a a safety radiusrwr\_\{w\}in the inputℓ2\\ell\_\{2\}space\. This radius represents the minimum perturbation required to alter the inclusion or exclusion of tokenwwby way of

rw=\{σ​Φ−1​\(p¯w,T\)for​w∈𝒱​c​e​r​tσ​Φ−1​\(1−p¯w,T\)for​w∈𝒱​e​x​c​l\.r\_\{w\}=\\begin\{cases\}\\sigma\\Phi^\{\-1\}\(\\underline\{p\}\_\{w,T\}\)&\\text\{for \}w\\in\\mathcal\{V\}\{cert\}\\\\ \\sigma\\Phi^\{\-1\}\(1\-\\overline\{p\}\_\{w,T\}\)&\\text\{for \}w\\in\\mathcal\{V\}\{excl\}\\kern 5\.0pt\.\\end\{cases\}\(10\)The global radiusRa​t​o​m​i​c=minw∈𝒱c​e​r​t∪𝒱e​x​c​l⁡rwR\_\{atomic\}=\\min\_\{w\\in\\mathcal\{V\}\_\{cert\}\\cup\\mathcal\{V\}\_\{excl\}\}r\_\{w\}represents the smallest perturbation required to alter the inclusion or exclusion of any token in the atomic gate\. This ensures that any perturbationδ\\deltawith‖δ‖2<Ra​t​o​m​i​c\\\|\\delta\\\|\_\{2\}<R\_\{atomic\}is guaranteed, with probability1−αa​t​o​m​i​c1\-\\alpha\_\{atomic\}, to leave the certified and excluded vocabularies unchanged, preserving the downstream structural tournaments candidate pool\.

#### Limitations: Multiplicity and Global Control

It may be notable that our process applies a fixed threshold of1/αa​t​o​m​i​c1/\\alpha\_\{atomic\}to each token independently\. This provides a rigorouslocalguarantee for each word but does not strictly control the Family\-Wise Error Rate over the entire vocabulary\. The true Type I error rate associated with the atomic gate is notαa​t​o​m​i​c\\alpha\_\{atomic\}, but rather it scales with the size of the vocabulary, to having a true confidence of\|𝒱\|​αa​t​o​m​i​c\|\\mathcal\{V\}\|\\alpha\_\{atomic\}\. As such, it is possible that an adversary may be able to insert or remove a word into the certified vocabulary𝒱c​e​r​t\\mathcal\{V\}\_\{cert\}with a smaller perturbation than the one calculated—which would be required to survive tournament certification\.

While more conservative global control can be achieved via the e\-Benjamini\-Hochberg \(e\-BH,Benjamini and Hochberg \([1995](https://arxiv.org/html/2606.27698#bib.bib59)\); Wang and Ramdas \([2022](https://arxiv.org/html/2606.27698#bib.bib60)\)\) procedure or Bonferroni corrections, we emphasize that the subsequentSentence Tournamentacts as a robust secondary gate\. Even if a low\-probability hallucination passes the atomic certification, it must still prove its structural and statistical validity in the tournament to appear in the final output\.

For reference, the e\-BH procedure would offer a more sophisticated mechanism for managing the global False Discovery Rate\. By sorting the E\-values\{E\(1\)≥E\(2\)​⋯≥E\(\|𝒱\|\)\}\\\{E\_\{\(1\)\}\\geq E\_\{\(2\)\}\\dots\\geq E\_\{\(\|\\mathcal\{V\}\|\)\}\\\}and finding the largestkksuch that1\|𝒱\|​∑i=1kE\(i\)≥1/αa​t​o​m​i​c\\frac\{1\}\{\|\\mathcal\{V\}\|\}\\sum\_\{i=1\}^\{k\}E\_\{\(i\)\}\\geq 1/\\alpha\_\{atomic\}, e\-BH provides a rigorous guarantee that is robust to the unknown dependence structures common in ASR output distributions\. We exclude e\-BH from our implementation because its dynamic, data\-dependent threshold is difficult to map back to stable safety radii, and the sort\-and\-sum operation introduces a GPU synchronization bottleneck\.

### 3\.2Tournament Certification

Up to this point, we have a certified vocabulary𝒱c​e​r​t\\mathcal\{V\}\_\{cert\}, which can be used to understand the components of a sentence that may be more \(or less\) vulnerable to adversarial manipulation\. While this can provide valuable security insights in high stakes transcription environments, it still represents a bag\-of\-words that is unable to recreate the desired output of an ASR model: a clean, certified transcription\. We bridge this gap by constructing a set of candidate transcriptions𝒞\\mathcal\{C\}restricted to the certified vocabulary𝒱c​e​r​t\\mathcal\{V\}\_\{cert\}\.

#### Nomination via Filtered Mapping

To achieve this, we draw a new set ofN3N\_\{3\}samples under the same gaussian noise distribution as above, and apply a filtering mappingΠ𝒱c​e​r​t:𝒳⋆→𝒱c​e​r​t⋆\\Pi\_\{\\mathcal\{V\}\_\{cert\}\}:\\mathcal\{X\}^\{\\star\}\\to\\mathcal\{V\}\_\{cert\}^\{\\star\}, which converts anyYi=\(wi,1,…,wi,Li\)Y\_\{i\}=\(w\_\{i,1\},\\dots,w\_\{i,L\_\{i\}\}\)to

Y~i=\(wi,j∈Yi​s\.t\.​wi,j∈𝒱c​e​r​t\),i=1,…,NT\.\\tilde\{Y\}\_\{i\}=\\big\(w\_\{i,j\}\\in Y\_\{i\}\\;\\;\\text\{s\.t\.\}\\;\\;w\_\{i,j\}\\in\\mathcal\{V\}\_\{cert\}\\big\),\\quad i=1,\\dots,N\_\{T\}\\kern 5\.0pt\.\(11\)By removing tokens that were unable to be verified by the atomic gate, we prune hallucinations and temporal jitter\. The top\-KKmost frequent unique cleaned sequences form the Structural Candidate Pool𝒞=\{C1,…,CK\}\\mathcal\{C\}=\\\{C\_\{1\},\\dots,C\_\{K\}\\\}\.

#### The Competitive Structural Tournament

Following the filtered sampling stage, our goal is to estimate the most likely transcription under the induced distribution over filtered sequences

Y^=arg⁡maxy∈𝒱c​e​r​t⋆⁡ℙ​\(Π𝒱c​e​r​t​\(f​\(x\+ϵ\)\)=y\)\.\\hat\{Y\}=\\arg\\max\_\{y\\in\\mathcal\{V\}\_\{cert\}^\{\\star\}\}\\mathbb\{P\}\\big\(\\Pi\_\{\\mathcal\{V\}\_\{cert\}\}\(f\(x\+\\epsilon\)\)=y\\big\)\\kern 5\.0pt\.\(12\)Accordingly, the tournament procedure constructs a Monte Carlo estimatorY^\\hat\{Y\}of the smoothed predictorF​\(x\)F\(x\)by approximating the maximizer of the induced distribution over filtered sequences\. The tournament stage maintains competing hypotheses over𝒞\\mathcal\{C\}and performs an anytime\-valid sequential test\. For a fresh batch ofNTN\_\{T\}transcriptionsYtY\_\{t\}, we updateKKparallel E\-values using a competitive betting function

Ei,t=Ei,t−1×\(1\+λ​\(𝕀​\(i=arg⁡minj⁡WER​\(Cj,Yt\)\)−1K\)\),E\_\{i,t\}=E\_\{i,t\-1\}\\times\\left\(1\+\\lambda\\left\(\\mathbb\{I\}\(i=\\arg\\min\_\{j\}\\text\{WER\}\(C\_\{j\},Y\_\{t\}\)\)\-\\frac\{1\}\{K\}\\right\)\\right\)\\kern 5\.0pt,\(13\)whereWER​\(⋅,⋅\)\\text\{WER\}\(\\cdot,\\cdot\)denotes the word error rate\. Under the null hypothesis that no candidate is dominant, eachEiE\_\{i\}is a martingale\.

###### Lemma 1\(The Structural Multiplicity Subsidy\)

The average wealthE¯​t​o​u​r​n=1K​∑i=1K​Ei\\bar\{E\}\{tourn\}=\\frac\{1\}\{K\}\\sum\{i=1\}^\{K\}E\_\{i\}is a non\-negative martingale\. We stop the tournament at any timeτ\\tauwhereE¯τ≥1/αt​o​u​r​n\\bar\{E\}\_\{\\tau\}\\geq 1/\\alpha\_\{tourn\}\. This allows the winning candidate’s wealth to subsidize the testing debt of other elements, enabling sequence\-level certification withO​\(1\)O\(1\)multiplicity scaling relative to the candidate pool size\.

The structural radius is one again produced through aCohenet al\.\([2019](https://arxiv.org/html/2606.27698#bib.bib5)\)style certification\. However, as the tournament exists overKK\-classes \(corresponding to theKKhighest frequency observations\), the associated certified radius—expressed in terms of the probability boundsp¯\\underline\{p\}andp¯\\overline\{p\}—becomes

Rt​o​u​r=σ2​\(Φ−1​\(p¯w​i​n​n​e​r\)−Φ−1​\(p¯r​u​n​n​e​r−u​p\)\)\.R\_\{tour\}=\\frac\{\\sigma\}\{2\}\(\\Phi^\{\-1\}\(\\underline\{p\}\_\{winner\}\)\-\\Phi^\{\-1\}\(\\overline\{p\}\_\{runner\-up\}\)\)\\kern 5\.0pt\.\(14\)
###### Theorem 1\(End\-to\-End Certified Transcription\)

LetY^\\hat\{Y\}be the output of the hierarchical aggregation procedure described above, with error budgetα=αa​t​o​m​i​c\+αt​o​u​r​n\\alpha=\\alpha\_\{atomic\}\+\\alpha\_\{tourn\}\. Then, with probability at least1−α1\-\\alpha, the predictionY^\\hat\{Y\}is invariant under all perturbationsδ\\deltasatisfying‖δ‖2<R\\\|\\delta\\\|\_\{2\}<R, where

R=min⁡\(Ra​t​o​m​i​c,Rt​o​u​r​n\)\.R=\\min\(R\_\{atomic\},R\_\{tourn\}\)\.\(15\)

#### Discussion: Avoidance of Alignment

The traditional approach to sequence certification relies on reassembling transcriptions via sequence alignment algorithms such as Recognizer Output Voting Error Reduction \(ROVER\)\(Fiscus,[1997](https://arxiv.org/html/2606.27698#bib.bib55); Haihuaet al\.,[2009](https://arxiv.org/html/2606.27698#bib.bib61)\)or confusion network lattices\(Manguet al\.,[2000](https://arxiv.org/html/2606.27698#bib.bib56)\), and applying union bounds over the resulting paths\(Olivier and Raj,[2021](https://arxiv.org/html/2606.27698#bib.bib57)\)\. However, this reassembly introduces a fundamental*Multiplicity Bottleneck*\. To certify a sequence ofNNtokens, one must implicitly or explicitly prove the correctness of the underlying total order\. Because a total order is uniquely defined by the set of all\(N2\)\\binom\{N\}\{2\}pairwise relations, maintaining a global confidence level1−α1\-\\alphaunder a standard union bound requires each pairwise test to satisfy an error probability ofO​\(α/N2\)O\(\\alpha/N^\{2\}\)\. This quadratic decay in statistical power rapidly results in vacuous radii\.

Moreover, alignment based mechanisms have the potential to produce transcriptions that are locally robust, but globally incoherent based upon the combination of words from multiple exclusive linguistic paths\. By contrast, our framework treats the sequence as the fundamental unit of competition\. By restricting the candidate pool𝒞\\mathcal\{C\}to model\-generated sequences passing through the Atomic Gate, we ensure that the tournament is over linguistically coherent sentences grounded in certified evidence\.

Our categorical tournament approach transforms the sequence certification problem from a combinatorial search over an alignment lattice into a competitive wealth redistribution task\. A key property of E\-values is that the statistical evidence required to crown a structural winner does not scale with the number of possible word orderings, but rather with the relative dominance of the winning candidate over its primary competitors\. This allows us to produce rigorous, sentence\-level certificates in high\-noise environments\.

## 4Analysis

To assess the performance our approach, we considered experiments using LibriSpeech and Common Voice exposed to the Whisper\-v3\-Large, Whisper\-Small, wav2vec 2\.0 Large and HuBERT architectures \(further detailed in Appendix[A](https://arxiv.org/html/2606.27698#A1)\. At a high level, the empirical results contained within Tables[1](https://arxiv.org/html/2606.27698#S4.T1)and[2](https://arxiv.org/html/2606.27698#S4.T2)provide strong evidence for the effectiveness of our hierarchical E\-value framework to produce diagnostic markers of robustness\.

Perhaps most crucially, Figure[1](https://arxiv.org/html/2606.27698#S4.F1)demonstrates that there is a clear correlation between the constructed Certifications and the measured WER\. For the purposes of a system in production, the WER requires knowing the ground truth transcription—something that is not possible in practice\. That the Certified Radius correlates with this, without requiring access to the ground truth transcription, highlights how our approach can be used to guide a confident view of the performance of an ASR system exposed to untrusted data\. Figure[3](https://arxiv.org/html/2606.27698#A3.F3)further expands on this, demonstrating that there is a notable performance advantage induced by our mechanism\. This is achieved through establishing consensus through an alignment lattice–a process that becomes statistically and computationally prohibitive as noise increases, due to the growth in hallucinated textual observations—we instead focus upon realizable computational utility through our approach\. Through our atomic gate, we prune the hypothesis space before reassembly occurs, reducing sentence diversity to a managable level\.

As we noted in the Limitations section of Section[3\.1](https://arxiv.org/html/2606.27698#S3.SS1), it is important to consider these certifications as a diagnostic marker, rather than a measure of true robustness, as the atomic certifications underestimate the true Type I error rate associated with the individual word certifications\. However, even with this limitation, our two phase approach provides significant utility, in that it allows granular word\-by\-word certifications to be produced, analyzed, and used in concert with broader sentence level certifications\. These have the potential to provide significant downstream utility, as it provides a mechanism for analyzing the stability of the ultimate utterance, as well as all its components\.

These results are expanded upon in Table[1](https://arxiv.org/html/2606.27698#S4.T1), which demonstrates that our certified pipeline induces a significant reduction in the WER across all tested SNRs\. In the highest\-noise regimes, OVER \(Olivier\) recall collapses to0\.0–2\.1%in extreme noise\. In contrast, our framework maintains a stable certification recall of40\.5–74\.0%at the same noise level\. Furthermore, for the SOTA Whisper\-Large\-v3 model, our framework repairs the raw prediction from0\.273 to 0\.126 WER—a54%54\\%relative improvement in robust ASR\. Notably, the observed Average Radius is inversely correlated with the SNR, validating that our anytime\-valid bonds adapt to the underlying quality of the signal\. That our aggregate consistently reduces the raw WER demonstrates that our approach is not merely a filter, but a robust aggregator that extracts semantic coherence from noisy model outputs\.

The comparison between our Tournament framework, ROVER \(Olivier\), and Naive Cohen randomized smoothing reveals two critical statistical phenomena that justify the utility of anytime\-valid sequence certification\. As shown in Table[1](https://arxiv.org/html/2606.27698#S4.T1), baseline methods \(ROVER and Cohen\) occasionally exhibit more negative Spearman correlation coefficients \(ρ\\rho\) than our approach at high SNR levels \(e\.g\., ROVER’s \-0\.825 for Whisper\-Large at 10dB\)\. However, correlation is meaningless in the presence of low recall\. For example, at 10dB SNR, our framework achieves significantly higher Recall \(73%\) than ROVER \(59%\) while maintaining a highly informative radius \(ρ=−0\.310\\rho=\-0\.310\)\. Crucially, our Certified Radius provides a trust score for every sample, not just the certified subset, providing actionable information across the entire dataset\.

We also stress that as the level of induced noise increases to SNR \-5\.0, the utility of baseline certificates collapses entirely\. That baseline recall collapses towards 0% in noise \(e\.g\., 0% for HuBERT and Wav2Vec2\), their certificates become a constant vector of zeros\. In contrast, our Tournament maintains stable, non\-zero variance and significant informative power even at \-5dB, proving it is the only viable path for trust\-scoring in extreme environments\. Even relative to the raw WER, our tournament approach yields a55\.1%55\.1\\%reduction in the WER\.

Table 1:Comprehensive Evaluation over all SNR levels\.Recallis 99% confidence certification success\.Corr\(ρ\\rho\) is the Spearman correlation between the method’s confidence metric and the resulting WER\. Note that baselines vanish in noise while our framework remains informative\. Models are Hubert \(H\), Whisper\-Large \(W\.L\), wav2vec Large \(W2\.L\) and Whisper\-Small \(W\.S\)\.#### Content Fragility

One advantage of considering ASR transcriptions on both a sentence level and as an assembly of vocabularies is that we can consider the performance of specific textual components within the overall robustness framework\. To achieve this, we employed the spaCy Natural Language Processing framework\(Honnibalet al\.,[2020](https://arxiv.org/html/2606.27698#bib.bib70)\)to perform Part\-of\-Speech \(POS\) tagging over the transcription corpus\. Table[2](https://arxiv.org/html/2606.27698#S4.T2)’s audit reveals a crucial linguistic insight:robustness is not uniformly distributed across word classes\. Content\-heavy tokens such as nouns \(NOUN\) and verbs \(VERB\) exhibit significantly lower raw accuracy and smaller certification margins, as compared to more functional tokens \(like coordinating conjunctions CCONJ and determiners DET\)\. This is an important insight for ASR validation, that is likely explained by the relative paucity of specific nouns and proper nouns \(PROPN\) in the ASR training corpora\.

The Certified Accuracy column demonstrates the utility of our tournament stage, in that the system can successfully recover these forms of content words with high accuracy—including reaching92\.2%92\.2\\%accuracy for nouns, a relative improvement of20\.5%20\.5\\%\. This demonstrates that anchoring the structural selection to a robust functional skeleton can significantly improve the performance of ASR systems within complex acoustic environments\.

Further results can be found in Appendix[C](https://arxiv.org/html/2606.27698#A3), covering computational efficiency and generalization\.

Table 2:Linguistic Fragility and Accuracy by POS \(see Appendix[B](https://arxiv.org/html/2606.27698#A2)for details\)\. Comparison of raw model recall vs\. Certified System Recall\.![Refer to caption](https://arxiv.org/html/2606.27698v1/x1.png)Figure 1:Observed WER as a function of Certified Radius: demonstrating the broad correlation between these two quantities\. Left: LibriSpeech\. Right: Common Voice\.

## 5Conclusion

In this work, we have demonstrated that acoustic robustness for sequence\-to\-sequence systems can be achieved through flexible, computationally efficient statistical mechanisms\. By replacing combinatorial sequence alignment with a hierarchy of E\-value tournaments, we demonstrate that it is possible to achieve superior performance as compared to alternate approaches, maintaining a stable Precision\-Recall balance even as environmental noise increases\.

Rather than strictly optimizing for mathematically rigorous, worst\-case certifications—which often scale poorly—our Tournament approach leverages these statistical mechanics to repair model transcriptions and generate dynamic safety markers\. Crucially, we demonstrate that these markers strongly correlate with Word Error Rate\. This correlation is invaluable for deployed systems, as ground\-truth transcripts are unavailable at runtime\. This observation effectively nullifies the utility of the WER as a live diagnostic tool\. By instead producing a product for transcription accuracy, our framework bridges this gap\. Ultimately, this work provides the foundations for anytime\-valid safety evaluations of high\-dimensional discrete outputs, ith potential for extension to machine translation, code generation, and autonomous command\-and\-control ecosystems where reliability is paramount\.

## Acknowledgments

This work was supported by the Australian Defence Science and Technology \(DST\) Group via the Advanced Strategic Capabilities Accelerator \(ASCA\) program\.

## Impact Statement

This work explores the potential for enhancing robustness in adversarially exposed acoustic systems, specifically Automatic Speech Recognition models, and exists within the oeuvre of Adversarial Machine Learning\. While defensive works in this space are typically assessed as eliciting no harms, we feel it is important to emphasize two key societal concerns which may be of note\.

The first of which is that there are some applications where a lack of robustness in a model may be positive\. In a world where broad scale surveillance is increasingly normalized, it may well be the case that adversarial attacks may induce privacy, creating a net public good\.

The second relates to how works like this position risk and harm\. A common precept within the Adversarial Machine Learning community is to assume a particular threat model, with the nature of academic comparisons often incentivizing us to then follow in the footsteps of those who came before us\. However, in doing so, we inadvertently create—and, crucially, present—\-a myopic view of the risk landscape\. In essence, we portray to practitioners that risk is concentrated within the areas in which we act as a community, when our investigations may be more motivated by historic alignment to academic norms and mathematical convenience\. This work considersℓ2\\ell\_\{2\}perturbations, which, while aligned with classical acoustic threat models, still represent a restriction relative to the overall threat landscape\. Such a consideration is especially important, given the limitations associated withαa​t​o​m​i​c\\alpha\_\{atomic\}, as discussed within the work\.

We emphasize the above points not just for the risks of erroneous portrayals of risk to practitioners, but also because our focus on these spaces inherently biases real attacker behavior away from these threat models\. After all, if an attacker understands that anℓ2\\ell\_\{2\}threat model is likely defended against, they’re naturally incentivized to consider an alternative pathway for model manipulation\.

With these points made, we still believe that research into defences, and in particular certified defences, induce a net societal gain\. Improving robustness to natural or adversarial perturbations will improve the performance of systems that are already one of the dominant access portals for AI within the community\. Moreover, voice\-based systems provide significant accessibility dividends to members of the community who do not have the ability to employ textual impacts, presenting an additional accessibility dividend\.

## References

- R\. Ardila, M\. Branson, K\. Davis, M\. Kohler, J\. Meyer, M\. Henretty, R\. Morais, L\. Saunders, F\. Tyers, and G\. Weber \(2020\)Common Voice: A Massively\-Multilingual Speech Corpus\.InProceedings of the twelfth language resources and evaluation conference,pp\. 4218–4222\.Cited by:[§A\.1](https://arxiv.org/html/2606.27698#A1.SS1.p1.1)\.
- Wav2Vec 2\.0: A Framework for Self\-Supervised Learning of Speech Representations\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Baevski, Y\. Zhou, A\. Mohamed, and M\. Auli \(2020b\)Wav2vec 2\.0: A Framework for Self\-Supervised Learning of Speech Representations\.Advances in neural information processing systems33,pp\. 12449–12460\.Cited by:[§A\.1](https://arxiv.org/html/2606.27698#A1.SS1.p1.1)\.
- Y\. Benjamini and Y\. Hochberg \(1995\)Controlling the false discovery rate: a practical and powerful approach to multiple testing\.Journal of the Royal Statistical Society: Series B \(Methodological\)57\(1\),pp\. 289–300\.Cited by:[§3\.1](https://arxiv.org/html/2606.27698#S3.SS1.SSS0.Px4.p2.1)\.
- N\. Carlini and D\. Wagner \(2018\)Audio Adversarial Examples: Targeted Attacks on Speech\-to\-Text\.InIEEE Security and Privacy Workshops \(SPW\),Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p1.1)\.
- C\. J\. Clopper and E\. S\. Pearson \(1934\)The Use of Confidence or Fiducial Limits Illustrated in the case of the Binomial\.Biometrika26\(4\),pp\. 404–413\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px2.p2.3)\.
- J\. Cohen, E\. Rosenfeld, and Z\. Kolter \(2019\)Certified Adversarial Robustness via Randomized Smoothing\.InInternational Conference on Machine Learning,pp\. 1310–1320\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p1.1),[§1](https://arxiv.org/html/2606.27698#S1.p2.7),[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px2.p1.9),[§3\.1](https://arxiv.org/html/2606.27698#S3.SS1.SSS0.Px3.p1.9),[§3\.2](https://arxiv.org/html/2606.27698#S3.SS2.SSS0.Px2.p2.4),[§3](https://arxiv.org/html/2606.27698#S3.p1.13),[footnote 1](https://arxiv.org/html/2606.27698#footnote1)\.
- A\. C\. Cullen, S\. Liu, P\. Montague, S\. M\. Erfani, and B\. I\. Rubinstein \(2024\)Et Tu Certifications: Robustness Certificates Yield Better Adversarial Examples\.InProceedings of the 41st International Conference on Machine Learning,pp\. 9745–9761\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p2.1)\.
- A\. C\. Cullen, P\. Montague, S\. Liu, S\. M\. Erfani, and B\. I\.P\. Rubinstein \(2022\)Double Bubble, Toil and Trouble: Enhancing Certified Robustness through Transitivity\.Advances in Neural Information Processing Systems35,pp\. 19099–19112\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p1.1)\.
- J\. L\. Doob \(1940\)Regularity Properties of Certain Families of Chance Variables\.Transactions of the American Mathematical Society47\(3\),pp\. 455–486\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px2.p4.5),[§3\.1](https://arxiv.org/html/2606.27698#S3.SS1.SSS0.Px2.p3.4)\.
- J\. G\. Fiscus \(1997\)A Post\-Processing System to Yield Reduced Word Error Rates: Recognizer Output Voting Error Reduction \(ROVER\)\.In1997 IEEE workshop on automatic speech recognition and understanding proceedings,pp\. 347–354\.Cited by:[§3\.2](https://arxiv.org/html/2606.27698#S3.SS2.SSS0.Px3.p1.4)\.
- I\. J\. Goodfellow, J\. Shlens, and C\. Szegedy \(2014\)Explaining and Harnessing Adversarial Examples\.arXiv preprint arXiv:1412\.6572\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p2.1)\.
- X\. Haihua, Z\. Jie, and G\. Wu \(2009\)An efficient multistage ROVER Method for Automatic Speech Recognition\.In2009 IEEE International Conference on Multimedia and Expo,pp\. 894–897\.Cited by:[§3\.2](https://arxiv.org/html/2606.27698#S3.SS2.SSS0.Px3.p1.4)\.
- A\. Hannun, C\. Case, J\. Casper, B\. Catanzaro, G\. Diamos, E\. Elsen, R\. Prenger, S\. Satheesh, S\. Sengupta, A\. Coates, and A\. Y\. Ng \(2014\)Deep Speech: Scaling up End\-to\-End Speech Recognition\.arXiv preprint\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p1.1)\.
- M\. Honnibal, I\. Montani, S\. Van Landeghem, A\. Boyd,et al\.\(2020\)spaCy: Industrial\-strength Natural Language Processing in Python\.Cited by:[Appendix B](https://arxiv.org/html/2606.27698#A2.p2.1),[§4](https://arxiv.org/html/2606.27698#S4.SS0.SSS0.Px1.p1.1)\.
- W\. Hsu, B\. Bolte, Y\. H\. Tsai, K\. Lakhotia, R\. Salakhutdinov, and A\. Mohamed \(2021\)HuBERT: Self\-Supervised Speech Representation Learning by Masked Prediction of Hidden Units\.IEEE/ACM transactions on audio, speech, and language processing29,pp\. 3451–3460\.Cited by:[§A\.1](https://arxiv.org/html/2606.27698#A1.SS1.p1.1)\.
- Z\. Huang, N\. G\. Marchant, K\. Lucas, L\. Bauer, O\. Ohrimenko, and B\. Rubinstein \(2023\)RS\-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized Deletion\.Advances in Neural Information Processing Systems36,pp\. 18676–18711\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p3.1)\.
- Z\. Huang, N\. G\. Marchant, O\. Ohrimenko, and B\. I\. Rubinstein \(2024\)CERT\-ED: Certifiably Robust Text Classification for Edit Distance\.InFindings of the Association for Computational Linguistics: EMNLP 2024,pp\. 10813–10835\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p3.1)\.
- S\. Hussain, P\. Neekhara, S\. Dubnov, J\. McAuley, and F\. Koushanfar \(2021\)WaveGuard: Understanding and Mitigating Audio Adversarial Examples\.In30th USENIX Security Symposium \(USENIX Security 21\),pp\. 2273–2290\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p4.3)\.
- R\. Johari, P\. Koomen, L\. Pekelis, and D\. Walsh \(2017\)Peeking at A/B Tests: Why it matters, and what to do about it\.InProceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining,pp\. 1517–1525\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px2.p3.2)\.
- M\. Lecuyer, V\. Atlidakis, R\. Geambasu, D\. Hsu, and S\. Jana \(2019\)Certified Robustness to Adversarial Examples with Differential Privacy\.In2019 IEEE Symposium on Security and Privacy \(SP\),pp\. 656–672\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p1.1)\.
- A\. Madry, A\. Makelov, L\. Schmidt, D\. Tsipras, and A\. Vladu \(2018\)Towards Deep Learning Models Resistant to Adversarial Attacks\.InInternational Conference on Learning Representations,Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p2.1)\.
- L\. Mangu, E\. Brill, and A\. Stolcke \(2000\)Finding Consensus in Speech Recognition: Word Error Minimization and Other Applications of Confusion Networks\.Computer Speech & Language14\(4\),pp\. 373–400\.Cited by:[§3\.2](https://arxiv.org/html/2606.27698#S3.SS2.SSS0.Px3.p1.4)\.
- R\. Olivier and B\. Raj \(2021\)Sequential Randomized Smoothing for Adversarially Robust Speech Recognition\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 6372–6386\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px3.p1.3),[§3\.2](https://arxiv.org/html/2606.27698#S3.SS2.SSS0.Px3.p1.4)\.
- R\. Olivier \(2023\)Assessing and Enhancing Adversarial Robustness in Context and Applications to Speech Security\.Ph\.D\. Thesis,Carnegie Mellon University, USA\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p3.1)\.
- R\. Olivier and B\. Raj \(2022\)Fooling Whisper with Adversarial Examples\.arXiv preprint\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p1.1)\.
- V\. Panayotov, G\. Chen, D\. Povey, and S\. Khudanpur \(2015\)Librispeech: An ASR Corpus based on Public Domain Audio Books\.In2015 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 5206–5210\.Cited by:[§A\.1](https://arxiv.org/html/2606.27698#A1.SS1.p1.1)\.
- Y\. Qin, N\. Carlini, I\. Goodfellow, G\. Cottrell, and C\. Raffel \(2019\)Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition\.InProceedings of the 36th International Conference on Machine Learning \(ICML\),Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p2.1)\.
- A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever \(2023a\)Robust Speech Recognition via Large\-Scale Weak Supervision\.InProceedings of the 40th International Conference on Machine Learning \(ICML\),pp\. 28448–28493\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever \(2023b\)Robust speech recognition via large\-scale weak supervision\.InInternational conference on machine learning,pp\. 28492–28518\.Cited by:[§A\.1](https://arxiv.org/html/2606.27698#A1.SS1.p1.1)\.
- A\. Ramdas, P\. Grünwald, V\. Vovk, and G\. Shafer \(2023\)Game\-Theoretic Statistics and Safe Anytime\-Valid Inference\.Statistical Science38\(4\),pp\. 576–601\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p5.1),[§3\.1](https://arxiv.org/html/2606.27698#S3.SS1.SSS0.Px3.p1.3)\.
- L\. Schönherr, K\. Kohls, S\. Zeiler, T\. Holz, and D\. Kolossa \(2019\)Adversarial Attacks against Automatic Speech Recognition Systems via Psychoacoustic Hiding\.InNetwork and Distributed System Security Symposium \(NDSS\),External Links:[Document](https://dx.doi.org/10.14722/ndss.2019.23288)Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p2.1)\.
- G\. Shafer and V\. Vovk \(2019\)Game\-Theoretic Foundations for Probability and Finance\.John Wiley & Sons\.Cited by:[§1](https://arxiv.org/html/2606.27698#S1.p5.1)\.
- Q\. Sun, S\. Chen, Y\. Zhai, Y\. Liu, and Z\. Zhong \(2024\)CommanderUAP: Practical and Transferable Universal Adversarial Perturbations on ASR Systems\.Cybersecurity\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p4.3)\.
- J\. Szurley and J\. Z\. Kolter \(2019\)Perceptual Based Adversarial Audio Attacks\.arXiv preprint\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px1.p4.3)\.
- J\. Ville \(1939\)Etude Critique de la Notion de Collectif\.Vol\.3,Gauthier\-Villars Paris\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px2.p4.5),[§3\.1](https://arxiv.org/html/2606.27698#S3.SS1.SSS0.Px2.p3.4)\.
- V\. Voráček \(2024\)Treatment of Statistical Estimation Problems in Randomized Smoothing for Adversarial Robustness\.Advances in Neural Information Processing Systems37,pp\. 133464–133486\.Cited by:[§2](https://arxiv.org/html/2606.27698#S2.SS0.SSS0.Px2.p5.1)\.
- R\. Wang and A\. Ramdas \(2022\)False discovery rate control with e\-values\.Journal of the Royal Statistical Society Series B: Statistical Methodology84\(3\),pp\. 822–852\.Cited by:[§3\.1](https://arxiv.org/html/2606.27698#S3.SS1.SSS0.Px4.p2.1)\.

## Appendix AAlgorithm

We present our full ASR certification pipeline in Algorithm[1](https://arxiv.org/html/2606.27698#alg1)\. In practice, we setτ=0\.5\\tau=0\.5, corresponding to majority occurrence under the smoothing distribution\.

Algorithm 1Certified ASR Pipeline1:Input:Audio signal

xx, target SNR \(dB\), confidence level

α=αa​t​o​m​i​c\+αt​o​u​r​n\\alpha=\\alpha\_\{atomic\}\+\\alpha\_\{tourn\}, base model

ff, threshold

τ∈\(0,1\)\\tau\\in\(0,1\), step size

λ∈\(0,1\]\\lambda\\in\(0,1\]
2:S0 \(Noise Calibration\):

3:Compute signal power

Ps=1d​∑i=1dxi2P\_\{s\}=\\frac\{1\}\{d\}\\sum\_\{i=1\}^\{d\}x\_\{i\}^\{2\}, noise power

Pϵ=Ps/10SNR/10P\_\{\\epsilon\}=P\_\{s\}/10^\{\\text\{SNR\}/10\}, and set

σ=Pϵ\\sigma=\\sqrt\{P\_\{\\epsilon\}\}\.

4:S1 \(Candidate Vocabulary Construction\):

5:for

i=1i=1to

N1N\_\{1\}do

6:Sample

ϵi∼𝒩​\(0,σ2​I\)\\epsilon\_\{i\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)
7:

Yi←f​\(x\+ϵi\)Y\_\{i\}\\leftarrow f\(x\+\\epsilon\_\{i\}\)//Yi∈𝒱⋆Y\_\{i\}\\in\\mathcal\{V\}^\{\\star\}

8:endfor

9:

𝒱←⋃i=1N1\{w:w∈Yi\}\\mathcal\{V\}\\leftarrow\\bigcup\_\{i=1\}^\{N\_\{1\}\}\\\{w:w\\in Y\_\{i\}\\\}
10:S2 \(Atomic Certification\):

11:for

t=1t=1to

NCN\_\{C\}do

12:Sample

ϵt∼𝒩​\(0,σ2​I\)\\epsilon\_\{t\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)
13:

Yt←f​\(x\+ϵt\)Y\_\{t\}\\leftarrow f\(x\+\\epsilon\_\{t\}\)
14:foreach

w∈𝒱w\\in\\mathcal\{V\}do

15:

Zw,t←𝕀​\{w∈Yt\}Z\_\{w,t\}\\leftarrow\\mathbb\{I\}\\\{w\\in Y\_\{t\}\\\}
16:

Ep​o​s​\(w\)←Ep​o​s​\(w\)⋅\(1\+λ​\(Zw,t−τ\)\)E\_\{pos\}\(w\)\\leftarrow E\_\{pos\}\(w\)\\cdot\(1\+\\lambda\(Z\_\{w,t\}\-\\tau\)\)
17:

En​e​g​\(w\)←En​e​g​\(w\)⋅\(1\+λ​\(τ−Zw,t\)\)E\_\{neg\}\(w\)\\leftarrow E\_\{neg\}\(w\)\\cdot\(1\+\\lambda\(\\tau\-Z\_\{w,t\}\)\)
18:endfor

19:endfor

20:

𝒱c​e​r​t←\{w∈𝒱:Ep​o​s​\(w\)≥1/αa​t​o​m​i​c\}\\mathcal\{V\}\_\{cert\}\\leftarrow\\\{w\\in\\mathcal\{V\}:E\_\{pos\}\(w\)\\geq 1/\\alpha\_\{atomic\}\\\}
21:

𝒱e​x​c​l←\{w∈𝒱:En​e​g​\(w\)≥1/αa​t​o​m​i​c\}\\mathcal\{V\}\_\{excl\}\\leftarrow\\\{w\\in\\mathcal\{V\}:E\_\{neg\}\(w\)\\geq 1/\\alpha\_\{atomic\}\\\}
22:S3 \(Filtered Transcription Sampling\):

23:for

i=1i=1to

N3N\_\{3\}do

24:Sample

ϵi∼𝒩​\(0,σ2​I\)\\epsilon\_\{i\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)
25:

Yi←f​\(x\+ϵi\)Y\_\{i\}\\leftarrow f\(x\+\\epsilon\_\{i\}\)
26:

Y~i←\(w∈Yi​s\.t\.​w∈𝒱c​e​r​t\)\\tilde\{Y\}\_\{i\}\\leftarrow\(w\\in Y\_\{i\}\\;\\text\{s\.t\.\}\\;w\\in\\mathcal\{V\}\_\{cert\}\)
27:endfor

28:Select top\-

KKmost frequent unique sequences from

\{Y~i\}\\\{\\tilde\{Y\}\_\{i\}\\\}as candidates

𝒞\\mathcal\{C\}
29:S4 \(Tournament Certification\):

30:Initialize wealth

Ei←1E\_\{i\}\\leftarrow 1for each

Ci∈𝒞C\_\{i\}\\in\\mathcal\{C\}
31:while

maxi⁡Ei<1/αt​o​u​r​n\\max\_\{i\}E\_\{i\}<1/\\alpha\_\{tourn\}and

t<N3t<N\_\{3\}do

32:Sample

ϵt∼𝒩​\(0,σ2​I\)\\epsilon\_\{t\}\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}I\)
33:

Yt←f​\(x\+ϵt\)Y\_\{t\}\\leftarrow f\(x\+\\epsilon\_\{t\}\)
34:

Y~t←\(w∈Yt​s\.t\.​w∈𝒱c​e​r​t\)\\tilde\{Y\}\_\{t\}\\leftarrow\(w\\in Y\_\{t\}\\;\\text\{s\.t\.\}\\;w\\in\\mathcal\{V\}\_\{cert\}\)
35:

i∗←arg⁡mini⁡WER​\(Ci,Y~t\)i^\{\*\}\\leftarrow\\arg\\min\_\{i\}\\text\{WER\}\(C\_\{i\},\\tilde\{Y\}\_\{t\}\)
36:for

i=1i=1to

KKdo

37:

Ei←Ei⋅\(1\+λ​\(𝕀​\(i=i∗\)−1/K\)\)E\_\{i\}\\leftarrow E\_\{i\}\\cdot\(1\+\\lambda\(\\mathbb\{I\}\(i=i^\{\*\}\)\-1/K\)\)
38:endfor

39:endwhile

40:Output:

Y^=arg⁡maxi⁡Ei,R=min⁡\{Ra​t​o​m​i​c,Rt​o​u​r​n\}\\hat\{Y\}=\\arg\\max\_\{i\}E\_\{i\},\\quad R=\\min\\\{R\_\{atomic\},R\_\{tourn\}\\\}

### A\.1Configuration

Our evaluation was performed across four ASR architectures, to demonstrate the model\-agnostic utility of our approach\. These include Whisper \(Large\-v3 & Small, MIT License,Radfordet al\.\([2023b](https://arxiv.org/html/2606.27698#bib.bib67)\)\), HuBERT\-Large \(MIT License,Hsuet al\.\([2021](https://arxiv.org/html/2606.27698#bib.bib68)\)\) and Wav2Vec2\-Large \(MIT License,Baevskiet al\.\([2020b](https://arxiv.org/html/2606.27698#bib.bib69)\)\)\. Of these, the Whisper is a transformer\-based encoder\-decoded model; HuBERT is a self\-supervised hidden unit BERT model fine\-tuned for CTC\-based ASR; and Wav2Vec is a self\-supervised framework\. Evaluations are performed zero\-shot \(inference\-only\) using thetest\-cleanandtest\-othersubsets of theLibriSpeechdataset\(Panayotovet al\.,[2015](https://arxiv.org/html/2606.27698#bib.bib65)\)\(licensed under Creative Commons Attribution 4\.0 International\) and the Englishtestsplit of theCommon Voice 17\.0variant\(Ardilaet al\.,[2020](https://arxiv.org/html/2606.27698#bib.bib66)\), which employs the Mozilla Public License 2\.0\. For each permutation, we process 100 random sentences, resulting in a comprehensive matrix of 3,200 unique certification trials\.

All experiments were conducted using the parameters outlined in Table[3](https://arxiv.org/html/2606.27698#A1.T3), on a H100 Tensor Core GPU with 80GB VRAM, using 16\-bit precision for Whisper\-Large\-v3\. Total experimental time was88GPU\-days\.

Table 3:Hyperparameter Configuration for Tournament Framework\.StageParameterValueDiscovery \(S1\)Sample Count \(NS​1N\_\{S1\}\)50Certification \(S2\)Sample Count \(NS​2N\_\{S2\}\)1000Nomination \(S3\)Candidate Count \(KK\)5Tournament \(S4\)Max Sample Budget \(NS​4N\_\{S4\}\)250MartingaleConfidence Level \(α\\alpha\)0\.01BettingInstance Count Parameter \(λi​n​s​t\\lambda\_\{inst\}\)0\.50BettingTournament Parameter \(λt​o​u​r​n​e​y\\lambda\_\{tourney\}\)0\.20#### Noise Augmentation

Our certification framework is built upon Additive White Gaussian Noise \(AWGN\) at four Signal\-to\-Noise Ratio \(SNR\) levels:\{10,5,0,−5\}\\\{10,5,0,\-5\\\}dB\. Given a clean audio signalxx, the noisy realizationx′x^\{\\prime\}is generated asx′=x\+ηx^\{\\prime\}=x\+\\eta, whereη∼𝒩​\(0,σ2\)\\eta\\sim\\mathcal\{N\}\(0,\\sigma^\{2\}\)andσ2\\sigma^\{2\}is scaled to achieve the target SNR\. All audio is standardized to a single\-channel 16kHz format before transcription\.

## Appendix BLinguistic Fragility

In a linguistic context, robustness of the final transcription may not be the only goal\. In fact, in security conscious environments, it may also be important to audit the robustness of individual words\. Such an audit framework also provides key information to those developing and deploying ASR models, as the results can be leveraged to understand weak points in systemic performance\.

Table 4:Definitions and examples for Part\-of\-Speech \(POS\) categories used in the linguistic fragility audit\.To achieve this audit, we employed spaCy Natural Language Processing framework\(Honnibalet al\.,[2020](https://arxiv.org/html/2606.27698#bib.bib70)\), where each word was tagged to a Part\-of\-Speech \(POS\) category \(see Table[4](https://arxiv.org/html/2606.27698#A2.T4)\) using spaCy’sen\_core\_web\_smmodel\. We then calculated theCertification Recall—defined as the percentage of tokens within a given linguistic class that successfully accumulated enough statistical wealth to be formally certified\.

Our audit indicates that robustness is not uniformly distributed across word classes\. Functional tokens, such as Pronouns \(PRON, 21\.7% recall\), Conjunctions \(CCONJ, 16\.7%\), and Adpositions \(ADP, 13\.1%\), are significantly easier to certify in noise\. In stark contrast, content\-heavy tokens, such as Nouns \(NOUN, 2\.8% recall\), Verbs \(VERB, 3\.8%\), and Proper Nouns \(PROPN, 1\.2%\), exhibit extreme fragility, with recall rates falling drastically\.

This discrepancy likely stems from the relative semantic density of these content words in the training corpora for these models\. In noisy environments, an ASR model may hallucinate numerous phonetically similar but semantically distinct nouns \(e\.g\., "cat", "cap", "bat"\)\. This high entropy disperses the probability mass, preventing any single content token from accumulating the statistical wealth required for certification under a strict E\-value martingale\. Functional words, however, belong to smaller, more closed linguistic classes and are often highly predictable given the surrounding context, allowing them to rapidly gain consensus\. This audit highlights the need for "linguistically\-aware" certification frameworks that can prioritize or differentially weight evidence accumulation for critical content words to improve system\-level trust\.

![Refer to caption](https://arxiv.org/html/2606.27698v1/x2.png)Figure 2:Relative Certification performance across all approaches
## Appendix CAdditional Results

The performance deltas visualized in Figures[3](https://arxiv.org/html/2606.27698#A3.F3)and[2](https://arxiv.org/html/2606.27698#A2.F2)confirm that the utility of the Tournament framework extends beyond mere certification\. By aggregating evidence across a massive search space, the system consistently produces a certified transcription that is significantly more accurate than any individual transcription under noise\. As detailed in Table[1](https://arxiv.org/html/2606.27698#S4.T1), this multiplicity subsidy is most pronounced in high\-noise environments, where our framework achieves an absolute WER reduction of up to 10\.6% at SNR \-5dB on the LibriSpeech dataset, and 8\.1% on Common Voice\. This indicates that the aggregation mechanism is particularly valuable precisely when the underlying model begins to fail catastrophically\. We note that the relative performance differences seen between the CTC\-based architectures \(HuBERT, wav2vec\) and the auto\-regressive models of Whisper are not distributed evenly \(see Figure[4](https://arxiv.org/html/2606.27698#A3.F4)\)\. We believe that this is a product of both the relative balance of the models computational cost to the certifications, and, perhaps more importantly, a decreased sensitivity to noise, which manifests as significantly smaller transcription sets for the ROVER based mechanism to aggregate ov

A critical requirement for practical ASR safety is real\-time or near real\-time feasibility\. Table[5](https://arxiv.org/html/2606.27698#A3.T5)demonstrates that for CTC\-based architectures \(HuBERT, Wav2Vec2\), our anytime\-valid framework is17–20% fasterthan traditional ROVER alignment\. While auto\-regressive models like Whisper\-Large incur a higher absolute cost due to their decoding mechanisms, the anytime\-stopping rule provides a28% compute savingrelative to a fixed\-budget approach, exiting the tournament as soon as statistical confidence is reached\. This proves that martingale\-based certification is a practical path for sequence\-level guarantees\.

![Refer to caption](https://arxiv.org/html/2606.27698v1/Figures/fig1_wer_scaling.png)Figure 3:Relationship between SNR and WER\. Solid lines: Certified Transcriptions, Dashed lines: Uncertified Transcriptions Under NoiseTable 5:Computational Efficiency \(RTF\)\. Tournament anytime\-stopping maintain competitive scaling even under heavy noise\.![Refer to caption](https://arxiv.org/html/2606.27698v1/x3.png)Figure 4:Relationship between SNR and the Real Time Factor \(RTF\) for different models\.

Similar Articles