Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Summary
This paper introduces a gradient-based speech-to-text alignment method applicable to any differentiable ASR model, including CTC, transducer, attention-based encoder-decoder, and speech large language models, requiring no training or model modification.
View Cached Full Text
Cached at: 07/09/26, 07:48 AM
# Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Source: [https://arxiv.org/html/2607.06831](https://arxiv.org/html/2607.06831)
###### Abstract
Speech\-to\-text alignment means finding the temporal boundaries of each word in the audio\. Some models provide such an alignment directly and others do not\. Connectionist temporal classification \(CTC\) and transducer models have an alignment by construction, whereas attention\-based encoder\-decoders \(AED\) and speech large language models \(LLMs\) do not, and their word timings are usually read off the attention weights instead\. All of these signals live on the encoder frame grid, which bounds their temporal precision\. We study a generic*gradient\-based alignment*that applies to any differentiable ASR model\. We take the gradient of each teacher\-forced token log probability with respect to the input, reduce it to a per\-frame saliency, and decode the resulting matrix into word boundaries with a single dynamic\-programming pass\. The method needs no training, no model modification and no alignment heads, works across all model families including the speech LLMs, and aligns on the input grid rather than on the coarser encoder grid\. We evaluate it on sixteen models from four families, on read \(TIMIT\) and spontaneous \(Buckeye\) speech, each against the model’s own native or attention\-based alignment\. We find that the gradient yields a usable alignment for every model, that it is usually somewhat behind a strong native aligner but better where the native alignment is weak, as for the streaming models, and that its main disadvantage is the cost of one backward pass per token\.
## IIntroduction & Related Work
Speech\-to\-text alignment assigns each transcribed word a start and end time in the audio\. It is classically produced by finding the best path through a Gaussian\-mixture hidden Markov model \(GM\-HMM\) which is still the most accurate aligner for read speech\[[13](https://arxiv.org/html/2607.06831#bib.bib62),[32](https://arxiv.org/html/2607.06831#bib.bib58),[20](https://arxiv.org/html/2607.06831#bib.bib56),[18](https://arxiv.org/html/2607.06831#bib.bib32),[35](https://arxiv.org/html/2607.06831#bib.bib41)\]\. Current models, however, differ in what alignment they provide on their own\[[29](https://arxiv.org/html/2607.06831#bib.bib34)\]: connectionist temporal classification \(CTC\)\[[15](https://arxiv.org/html/2607.06831#bib.bib7),[18](https://arxiv.org/html/2607.06831#bib.bib32)\]and transducer models\[[16](https://arxiv.org/html/2607.06831#bib.bib63)\]have one by construction, whereas attention\-based encoder\-decoders \(AED\)\[[9](https://arxiv.org/html/2607.06831#bib.bib10),[7](https://arxiv.org/html/2607.06831#bib.bib11),[47](https://arxiv.org/html/2607.06831#bib.bib87)\]and speech large language models \(LLMs\)\[[48](https://arxiv.org/html/2607.06831#bib.bib78),[10](https://arxiv.org/html/2607.06831#bib.bib79),[37](https://arxiv.org/html/2607.06831#bib.bib88)\]do not, and their word timings are usually read off the attention weights instead\. All of these signals live on the 20 to 80 ms encoder frame grid, which bounds their temporal precision\. For AED, reading the timing off the cross\-attention is established practice: Whisper emits coarse segment\-level timestamp tokens\[[33](https://arxiv.org/html/2607.06831#bib.bib52)\], word\-level timestamps follow from a dynamic time warping \(DTW\) of the decoder cross\-attention\[[6](https://arxiv.org/html/2607.06831#bib.bib51)\]that fine\-tuning can sharpen\[[49](https://arxiv.org/html/2607.06831#bib.bib53)\], and the alignment\-bearing heads can even be selected by an unsupervised heuristic and refined with character\-level teacher forcing without any training\[[46](https://arxiv.org/html/2607.06831#bib.bib54)\]\. Speech LLM self\-attention carries a similar alignment, used to constrain text\-to\-speech synthesis\[[43](https://arxiv.org/html/2607.06831#bib.bib49)\]and as a policy for simultaneous translation\[[24](https://arxiv.org/html/2607.06831#bib.bib48)\]; forced alignment has separately been cast as a trained LLM task\[[22](https://arxiv.org/html/2607.06831#bib.bib50)\]\.
We study a generic*gradient\-based alignment*that applies to any differentiable speech recognition model\. For each transcript token we take the gradient of its teacher\-forced log probability with respect to the input, reduce it to a per\-frame saliency\[[40](https://arxiv.org/html/2607.06831#bib.bib64)\], and decode the resulting token\-by\-frame matrix into word boundaries with a single dynamic\-programming pass \([Fig\.1](https://arxiv.org/html/2607.06831#S1.F1)\)\. As it uses the gradient w\.r\.t\. the input signal, it aligns on the input grid rather than on the coarser encoder grid, and it can correct temporal shifts of the encoder: the encoder is often so powerful that it can displace the signal in time \(e\.g\. with streaming models\) or even reverse the time dimension\[[36](https://arxiv.org/html/2607.06831#bib.bib46)\]111Arguably reversing the time dimension will not happen for CTC, though\., which degrades the native forced alignment, while the gradient should still give a meaningful alignment\. It needs no training and no model modification and it applies to all model families\. Unlike the attention\-based alignment, which first has to find which attention head carries the alignment \(hand\-picked for Whisper, and unpublished for the other models, so that we have to select it on a labeled development set\), the gradient uses one canonical signal, the saliency of the token’s own log probability\. In prior work, the same method has been applied to speech recognition AED models\[[36](https://arxiv.org/html/2607.06831#bib.bib46),[46](https://arxiv.org/html/2607.06831#bib.bib54)\]and machine\-translation models\[[11](https://arxiv.org/html/2607.06831#bib.bib55)\]\.
This work is not a proposal of gradient alignment as the best aligner\. A strong native aligner is usually still somewhat better, and the gradient is considerably more expensive\. We rather provide a broad and fair analysis of what the gradient alignment is, how well it works across all the model families, and where it wins and loses\. Our contributions are:
- •We show that the gradient\-based alignment applies to all ASR model families: CTC \(prefix scores\), transducers \(RNN\-T and TDT, also via prefix scores\), AED, and speech LLMs\.
- •Improved alignment path scoring, where the best alignment path is found via dynamic programming\. The cross\-attention dynamic time warping \(DTW\) used by Whisper and CrisperWhisper is a special case of it\. Our version improves on that DTW both on the attention and on the gradient signal\.
- •For the speech LLMs, we read the alignment off their self\-attention, the analog of the encoder\-decoder cross\-attention\.
- •A comprehensive and fair comparison of sixteen models across the four families, on read \(TIMIT\) and spontaneous \(Buckeye\) speech, each against its own native or attention\-based alignment, including the tokenization granularity, the input\-grid resolution, the recognition\-mode alignment, and the compute cost\.
- •Public source code to reproduce all results including the whole pipeline\.
Figure 1:Posteriorslog\(pt\(y=as∣x1T′\)\)s,t∈ℝS×T\\log\\left\(p\_\{t\}\(y\{=\}a\_\{s\}\\mid x\_\{1\}^\{T^\{\\prime\}\}\)\\right\)\_\{s,t\}\\in\\mathbb\{R\}^\{S\\times T\}, gradient scoreslogsoftmaxT\(G\)∈ℝS×T\\log\\operatorname\{softmax\}\_\{T\}\(G\)\\in\\mathbb\{R\}^\{S\\times T\}, and log self\-attention weights \(inℝS×T\\mathbb\{R\}^\{S\\times T\}\), all without energy weighting here, each with word boundaries, in comparison to the reference segmentation \(silence in white, words in blue, with word boundaries\)\.
## IIAlignments via Gradients
We can calculate the gradient of the log probabilityp\(as∣a1s−1,x1T′\)p\(a\_\{s\}\\mid a\_\{1\}^\{s\-1\},x\_\{1\}^\{T^\{\\prime\}\}\)of some target labelas∈𝒜a\_\{s\}\\in\\mathcal\{A\}\(the label vocabulary\) in label positionss, given the inputx1T′x\_\{1\}^\{T^\{\\prime\}\}and the historya1s−1∈𝒜s−1a\_\{1\}^\{s\-1\}\\in\\mathcal\{A\}^\{s\-1\}of previous labels, w\.r\.t\. an input framextx\_\{t\}\. The transcripta1Sa\_\{1\}^\{S\}hasSSlabels and the inputT′T^\{\\prime\}frames; the saliency matrix and the alignment below haveTTtime frames:T=T′T=T^\{\\prime\}when the gradient is taken at the input, or the number of encoder frames when it is taken at an encoder layer\. Comparing the norm of these gradients over the time framesttwill give us an indication of the importance of each frame for this specific output labelasa\_\{s\}\. Specifically, we calculate the log norm222We found that the log norm was better conditioned than the norm, and yielded better results\. We also tested differentpp\-norms, and foundp=2p=2in most cases to perform best\.
Gs,t:=log∥∇xtlogp\(as∣a1s−1,x1T′\)∥p∈ℝ\.G\_\{s,t\}:=\\log\\left\\\|\\nabla\_\{x\_\{t\}\}\\log p\(a\_\{s\}\\mid a\_\{1\}^\{s\-1\},x\_\{1\}^\{T^\{\\prime\}\}\)\\right\\\|\_\{p\}\\in\\mathbb\{R\}\.\(1\)The matrixlogsoftmaxTG\\log\\operatorname\{softmax\}\_\{T\}Gin[Fig\.1](https://arxiv.org/html/2607.06831#S1.F1)shows the alignment clearly\.
Note thatlogp\(as∣a1s−1,x1T′\)\\log p\(a\_\{s\}\\mid a\_\{1\}^\{s\-1\},x\_\{1\}^\{T^\{\\prime\}\}\)is straightforward to compute for an AED model or speech LLM \(we exclude the EOS label here\), and was done in a similar way in\[[36](https://arxiv.org/html/2607.06831#bib.bib46),[46](https://arxiv.org/html/2607.06831#bib.bib54)\]\. It is possible for CTC and transducer as well, using the prefix scores\[[17](https://arxiv.org/html/2607.06831#bib.bib18)\]which can be calculated efficiently using dynamic programming\.
To use this to get some alignment, we need to define which alignment label topology we allow \(mappinga1Sa\_\{1\}^\{S\}toy1Ty\_\{1\}^\{T\}\) and how to score one particular alignmenty1Ty\_\{1\}^\{T\}such that we can search for the one with the highest score\.
For the label topology, we mapa1Sa\_\{1\}^\{S\}toy1Ty\_\{1\}^\{T\}over𝒴=𝒜∪\{ϵ\}\\mathcal\{Y\}=\\mathcal\{A\}\\cup\\\{\\epsilon\\\}, letting each real label repeat over consecutive frames withϵ\\epsilon\(blank/silence\) labels in between\. We write this as a finite state automaton with enumerated statesY12S\+1=\(ϵ,1,ϵ,2,…,S,ϵ\)Y\_\{1\}^\{2S\+1\}=\(\\epsilon,1,\\epsilon,2,\\dots,S,\\epsilon\), whoseS\+1S\+1blank states sit before the first label, between adjacent labels, and after the last\. A topology is fixed by which of these blank states it permits, the rest being disallowed\. Our default is a*word\-level*topology: a blank only at word boundaries \(before the first word, after the last, and between adjacent words\), forbidding blanks inside a word, as we evaluate word boundaries\. The*full*, CTC\-like topology instead permits every blank \(unlike CTC, we do not force aϵ\\epsilonbetween two equal labelsas=as\+1a\_\{s\}=a\_\{s\+1\}\), and a variant with*no interior silence*keeps only the leading and trailing blank, so adjacent tokens follow each other directly\. We decode all topologies with the standard time\-synchronous Viterbi search \(as in CTC\), advancing one frame per step, so every token spans at least one frame\. Whisper and CrisperWhisper instead align the cross\-attention by DTW, which additionally allows vertical transitions \(advancing the token within a frame\), so adjacent tokens may overlap by up to one frame; with no interior silence and without the log of[eq\.3](https://arxiv.org/html/2607.06831#S2.E3), our decoder reduces to that DTW up to those vertical transitions \([TableVIII](https://arxiv.org/html/2607.06831#S2.T8)\)\. We search for an allowed state sequencer1Tr\_\{1\}^\{T\},rt∈\{1,…,2S\+1\}r\_\{t\}\\in\\\{1,\\dots,2S\+1\\\}, compatible witha1Sa\_\{1\}^\{S\}, that maximizes
GradScore\(r1T\)=∑t=1T\{GYrt,t′,Yrt≠ϵ,βt,Yrt=ϵ,\\operatorname\{GradScore\}\(r\_\{1\}^\{T\}\)=\\sum\_\{t=1\}^\{T\}\\begin\{cases\}G^\{\\prime\}\_\{Y\_\{r\_\{t\}\},t\},&Y\_\{r\_\{t\}\}\\neq\\epsilon,\\\\ \\beta\_\{t\},&Y\_\{r\_\{t\}\}=\\epsilon,\\end\{cases\}\(2\)with a token scoreG′G^\{\\prime\}and a silence scoreβt\\beta\_\{t\}defined below\. The bestr1Tr\_\{1\}^\{T\}is found via dynamic programming \([Fig\.2](https://arxiv.org/html/2607.06831#S2.F2)\), and the alignmenty1Ty\_\{1\}^\{T\}read off withyt=aYrty\_\{t\}=a\_\{Y\_\{r\_\{t\}\}\}for non\-blank states andyt=ϵy\_\{t\}=\\epsilonotherwise\.
ϵ\\epsilona1a\_\{1\}ϵ\\epsilona2a\_\{2\}ϵ\\epsilona3a\_\{3\}ϵ\\epsilonno intra\-wordϵ\\epsilontimettstayskipϵ\\epsilonFigure 2:Decoding the saliency into an alignment, for two words\(a1a2\)\(a3\)\(a\_\{1\}a\_\{2\}\)\(a\_\{3\}\)\. The DP runs a time\-synchronous Viterbi over the FSA statesY12S\+1Y\_\{1\}^\{2S\+1\}\(rows\): per frame it stays, advances one state \(\+1\+1\), or skips a blank\. The word\-level topology shades out the intra\-word blank \(betweena1a\_\{1\}anda2a\_\{2\}\), forcing a skip there, while keeping the word\-boundary blank \(state55\)\. DTW instead allows a vertical move \(advancing the token within a frame\), so adjacent tokens may overlap by up to one frame\.The token score is the over\-time log\-softmax of the \(optionally energy\-weighted\) saliency,
Gs,t′=logsoftmaxT\(G\+ρlogE\)s,t,G^\{\\prime\}\_\{s,t\}=\\log\\operatorname\{softmax\}\_\{T\}\(G\+\\rho\\log E\)\_\{s,t\},\(3\)whereρ≥0\\rho\\geq 0is a weighting \(ρ=0\.5\\rho=0\.5by default;ρ=0\\rho=0disables it\) andEt∈\[0,1\]E\_\{t\}\\in\[0,1\]a smoothed audio\-energy envelope on the frame grid: from the waveformxxwe form the windowed root\-mean\-square energyE~=w∗x2\\tilde\{E\}=\\sqrt\{w\*x^\{2\}\}, withwwa normalized2525ms Hann window \(∑iwi=1\\sum\_\{i\}w\_\{i\}=1\) and∗\*convolution, sample it at the frame centersctc\_\{t\}, and normalize by its maximum,Et=E~ct/maxt′E~ct′E\_\{t\}=\\tilde\{E\}\_\{c\_\{t\}\}/\\max\_\{t^\{\\prime\}\}\\tilde\{E\}\_\{c\_\{t^\{\\prime\}\}\}\. The energy weight suppresses the gradient’s spurious response in silent frames\.
The silence scoreβt\\beta\_\{t\}admits the constant blank of CTC, but also a self\-calibrating blank derived from the per\-frame statistics of the token scores\. Letμt\\mu\_\{t\}andσt\\sigma\_\{t\}be the mean and standard deviation of\{Gs,t′\}s\\\{G^\{\\prime\}\_\{s,t\}\\\}\_\{s\}over the labelsss\(taken atρ=0\\rho=0, so the blank calibrates against the raw saliency\)\. We consider
βt=\{γ,\(constant\)μt\+κσt,\(z\-score\)μt−λz\(Et\)σt,\(energy\)\\beta\_\{t\}=\\begin\{cases\}\\gamma,&\\text\{\(constant\)\}\\\\ \\mu\_\{t\}\+\\kappa\\,\\sigma\_\{t\},&\\text\{\(z\-score\)\}\\\\ \\mu\_\{t\}\-\\lambda\\,z\(E\_\{t\}\)\\,\\sigma\_\{t\},&\\text\{\(energy\)\}\\end\{cases\}\(4\)withz\(Et\)=\(Et−E¯\)/std\(E\)z\(E\_\{t\}\)=\(E\_\{t\}\-\\bar\{E\}\)/\\operatorname\{std\}\(E\)the z\-score of the energy over theTTframes \(E¯\\bar\{E\}andstd\(E\)\\operatorname\{std\}\(E\)its mean and standard deviation\), and hyperparametersγ\\gamma\(constant level\),κ\\kappaandλ\\lambda\. The energy blank is a VAD\-style silence emission: in a low\-energy framez\(Et\)<0z\(E\_\{t\}\)<0raisesβt\\beta\_\{t\}above the token mean, so the frame goes to the blank; in speech framesβt\\beta\_\{t\}drops below the token scores\.
TABLE I:Word\-boundary error and accuracy per model, gradient alignment vs\. each model’s native or attention aligner, on TIMIT\-test and Buckeye, grouped by family\. The attention and posterior aligners use the model’s native subword units; the gradient uses characters for AED and speech\-LLM and native subwords otherwise \([TableIV](https://arxiv.org/html/2607.06831#S2.T4)\)\. Starred Whisper\-large\-v3 gradient row \(\*\) is taken at optimal encoder depth \(encoder 3/4,[TableX](https://arxiv.org/html/2607.06831#S2.T10)\)\.ModelAlignmethodTIMITBuckeyeTypeNameWBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowWBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowGM\-HMMMFALikelihood191992\.392\.3323290\.090\.0CTCMMS\-FAGradients494964\.364\.312112158\.258\.2Posteriors373771\.571\.5464669\.969\.9XLS\-R\(Phoneme\)Gradients494965\.765\.7818166\.466\.4Posteriors303080\.780\.7444470\.470\.4ParakeetCTCGradients949440\.840\.811711737\.137\.1Posteriors777739\.439\.4999932\.232\.2OWSM\-CTCGradients14214228\.428\.414514530\.430\.4Posteriors12012020\.620\.611011025\.225\.2FastConformer\(streaming\)Gradients12712732\.032\.015815830\.130\.1Posteriors3573571\.11\.13663662\.22\.2Transd\.ParakeetRNN\-TGradients12512531\.531\.514714729\.129\.1Posteriors797939\.339\.3939335\.435\.4ParakeetTDTGradients12812832\.932\.915715727\.527\.5Posteriors808041\.241\.2909038\.638\.6Emformer\(streaming\)Gradients15915924\.824\.813913930\.430\.4Posteriors3943940\.30\.33993990\.50\.5FastConformer\(streaming\)Gradients15315327\.327\.318918924\.024\.0Posteriors3073073\.63\.63313313\.93\.9AEDWhisper\-baseGradients515159\.359\.3757552\.752\.7Cross\-att\.515162\.362\.3626266\.366\.3Whisper\-large\-v3Gradients555557\.857\.8535366\.266\.2Gradients\*333376\.976\.9393980\.080\.0Cross\-att\.424266\.766\.7494969\.169\.1Crisper\-WhisperGradients606054\.354\.3595959\.259\.2Cross\-att\.333380\.780\.7474780\.180\.1OWLS\-1BGradients15315331\.731\.718018032\.932\.9Cross\-att\.353576\.776\.7515169\.169\.1SpeechLLMVoxtralGradients777745\.645\.6747451\.051\.0Self\-att\.535361\.161\.1525265\.565\.5Phi\-4\-MMGradients848443\.043\.010410443\.143\.1Self\-att\.686846\.646\.6797939\.839\.8Canary\-QwenGradients909040\.640\.6999940\.640\.6Self\-att\.13113128\.028\.021121116\.316\.3
TABLE II:Hypothesis\-mode alignmenton Buckeye, where each model aligns its own recognition\. Identity\-gated F1 at a5050ms collar and Levenshtein\-matched WBE against the reference\.ModelRef\-WER\[%\]↓\\downarrowAlignmethodMatched\-F1TypeNamematch\[%\]↑\\uparrowWBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowCTCParakeetCTC92\.392\.311\.911\.9Gradients11611614\.014\.0Posteriors999911\.111\.1OWSM\-CTC89\.289\.214\.314\.3Gradients1421429\.99\.9Posteriors1141145\.45\.4FastConformer\(streaming\)86\.586\.516\.816\.8Gradients1641649\.79\.7Posteriors3633630\.20\.2Transd\.ParakeetRNN\-T89\.889\.813\.613\.6Gradients1561567\.67\.6Posteriors929213\.213\.2ParakeetTDT91\.891\.812\.812\.8Gradients1551557\.87\.8Posteriors868617\.117\.1Emformer\(streaming\)72\.172\.130\.530\.5Gradients1551557\.17\.1Posteriors3913910\.00\.0FastConformer\(streaming\)87\.187\.115\.915\.9Gradients2062065\.95\.9Posteriors3183180\.40\.4AEDWhisper\-base85\.585\.518\.018\.0Gradients858525\.525\.5Cross\-att\.777735\.835\.8Whisper\-large\-v388\.188\.115\.315\.3Gradients666640\.440\.4Cross\-att\.666640\.440\.4Crisper\-Whisper92\.092\.012\.712\.7Gradients565633\.833\.8Cross\-att\.444459\.659\.692\.692\.612\.212\.2Official464643\.543\.5OWLS\-1B85\.785\.7208\.0208\.0Gradients102410243\.83\.8Cross\-att\.131213127\.87\.8SpeechLLMVoxtral88\.388\.315\.315\.3Gradients878725\.325\.3Self\-att\.676737\.437\.4Phi\-4\-MM92\.892\.811\.711\.7Gradients10210219\.019\.0Self\-att\.808016\.016\.0Canary\-Qwen93\.193\.111\.411\.4Gradients959517\.117\.1Self\-att\.2122123\.93\.9
TABLE III:Signed boundary offsetson Buckeye, gradient alignment vs\. the model’s alternative aligner\. WBE is the mean absolute word\-boundary error, start/end off\. the signed mean boundary offset \(positive = late\), and width err\. the signed word\-width error \(0 = correct duration\)\. Parakeet RNN\-T is the offline contrast to the streaming transducers\.ModelAlignmethodWBE\[ms\]↓\\downarrowStart off\.\[ms\]End off\.\[ms\]Width err\.\[ms\]TypeNameCTCMMS\-FAGradients121121−93\-93−91\-91\+2\+2Posteriors4646\+47\+47−14\-14−60\-60XLS\-R\(Phoneme\)Gradients8181−72\-72−47\-47\+25\+25Posteriors4444\+46\+46−14\-14−60\-60ParakeetCTCGradients117117−33\-33−54\-54−21\-21Posteriors9999\+53\+53−51\-51−104\-104OWSM\-CTCGradients145145−89\-89−107\-107−18\-18Posteriors110110\+125\+125−10\-10−136\-136FastConformer\(streaming\)Gradients158158−62\-62−97\-97−35\-35Posteriors366366\+415\+415\+310\+310−106\-106Transd\.ParakeetRNN\-TGradients147147−22\-22−42\-42−20\-20Posteriors9393\+42\+42−70\-70−113\-113Emformer\(streaming\)Gradients139139\+67\+67\+32\+32−35\-35Posteriors399399\+484\+484\+308\+308−176\-176FastConformer\(streaming\)Gradients189189−85\-85−123\-123−37\-37Posteriors331331\+386\+386\+260\+260−126\-126AEDWhisper\-baseGradients7575\+4\+4−4\-4−9\-9Cross\-att\.6262−38\-38−19\-19\+19\+19
TABLE IV:Character vs\. subword targets, one representative model per family, Buckeye\. Each model shows its gradient alignment and its model\-native alternative\. The character target is a valid re\-segmentation of the transcript, one token per character with a word\-boundary token between words\.ModelAlignmethodCharSubwordTypeNameWBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowWBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowCTCParakeet CTCGradients416741672\.82\.811711737\.137\.1Posteriors17317326\.126\.1999932\.232\.2Transd\.Parakeet RNN\-TGradients38438422\.122\.114714729\.129\.1Posteriors1255125514\.814\.8939335\.435\.4AEDWhisper\-large\-v3Gradients535366\.266\.210810841\.241\.2Cross\-att\.575775\.275\.2494969\.169\.1SpeechLLMVoxtralGradients747451\.051\.011811838\.038\.0Self\-att\.12912946\.046\.0525265\.565\.5
TABLE V:Grad\-score ablationacross families, Buckeye\. The feature\-axis reduction \(L0\.5 / L1 / L2 / sum\), and whether the gradient is multiplied by the input \(the right group\)\. The signed sum columns need special handling, since their score can be negative: normalized score, constant blank score, on plain CTC topology\.ModelWBE\[ms\]↓\\downarrowTypeName∇\\nabla\(gradient\)∇×\\nabla\\timesinputL0\.5L1L2sumL2sumCTCMMS\-FA121121121121121121197197120120164164Parakeet CTC118118117117117117172172116116166166Transd\.Parakeet RNN\-T147147147147147147181181146146188188Emformer\(streaming\)142142140140139139184184138138190190AEDWhisper\-base7474747475751101107676122122SpeechLLMVoxtral7373747474741121127676121121Phi\-4\-MM107107106106104104138138103103119119Canary\-Qwen10010010010099991071079999113113
TABLE VI:Blank\-scoring schemeacross signals, Buckeye, word\-level topology, energy weighting fixed atρ=0\.5\\rho\{=\}0\.5\. The constant blankγ\\gamma, the energy\-aware silenceβt=μt−λz\(Et\)σt\\beta\_\{t\}=\\mu\_\{t\}\-\\lambda\\,z\(E\_\{t\}\)\\,\\sigma\_\{t\}, and the z\-scoreβt=μt\+κσt\\beta\_\{t\}=\\mu\_\{t\}\+\\kappa\\,\\sigma\_\{t\}\. WBE for the gradient signal and for the native attention signal\.BlankOptsWBE\[ms\]↓\\downarrowGradientsCross\-att\.Self\-att\.MMS\-FAWhi\-sperOWLSVox\-tralWhi\-sperOWLSVox\-tralconstantγ=−3\\gamma=\-317617612912965865818318310010074746363γ=−5\\gamma=\-51311318787577577131131444465655353γ=−8\\gamma=\-814114165652372379494656570706565energyλ=0\.5\\lambda=0\.513113154541931937777636365655959λ=1\\lambda=112712753531881887575616160605656λ=2\\lambda=212112153531801807474494951515252λ=3\\lambda=311711754541741747575484848485555z\-scoreκ=0\.5\\kappa=0\.513013054541921927777636365655959κ=1\\kappa=112712753531861867575616159595656κ=2\\kappa=212312356561811817878545451516060
TABLE VII:Audio\-energy token weighting, Buckeye, word\-level topology\. The token scores are weighted by \(audio energy\)ρbefore the DP, and we sweepρ\\rho\(ρ=0\\rho\{=\}0disables it\) at each blank scheme\. This complements[TableVI](https://arxiv.org/html/2607.06831#S2.T6), which sweeps the blank scheme at fixedρ=0\.5\\rho\{=\}0\.5\.Blankρ\\rhoWBE\[ms\]↓\\downarrowGradientsCross\-att\.Self\-att\.MMS\-FAWhi\-sperOWLSVox\-tralWhi\-sperOWLSVox\-tralconstantγ=−5\\gamma=\-50\.014714799998848841741745454777759590\.2514014089897007001411414848717156560\.513113187875775771311314444656553530\.7512612689895125121311314545606052521\.01241249292485485137137474757575252energyλ=1\\lambda=10\.0141141707026326394946464676762620\.25136136585821221281816363656559590\.5127127535318818875756161606056560\.75119119535317717775755757575754541\.011611656561711717777525253535353z\-scoreκ=1\\kappa=10\.0142142707026326394946565666664640\.25137137575721021081816363646459590\.5127127535318618675756161595956560\.75120120545417517576765757555554541\.011711757571731737979525252525454
TABLE VIII:Whisper’s cross\-attention DTW vs\. our aligner\(Whisper\-base cross\-attention or grad\. scores, Buckeye\)\. Top row reproduces Whisper’sfind\_alignment\. Columns: alignment\-head set \(Whisper’s curated heads, our gold\-tuned top\-kk, or single best\), token z\-norm, median filter, pre\-log of the score matrix \(*log*\), log\-softmax over time \(*log sm*; DP sums these log\-scores\), mono \(✓=\{\}=\{\}our monotonic DP forbidding vertical step;×=\\times\{\}=\{\}Whisper DTW\), silence topology \(*word*=\{\}=\{\}blank only between words;*none*=\{\}=\{\}no blank\), and energy weighting \(*en*\)\.Configurationheadsz\-normmed\.filt\.loglogsmmonosil\.en\.WBE\[ms\]↓\\downarrowCross\-attnWhisper \(faithful\)Whi\-sper✓✓×\\times×\\times×\\timesnone×\\times8585\+\+energy✓8989\+\+mono DP✓×\\times8484\+\+silenceword8484\+\+no median\-filter×\\times×\\timesnone8585\+\+no z\-norm×\\times✓8383\+\+log softmax✓8181\+\+mono DP✓✓8383\+\+silenceword7676\+\+our headsours✓×\\times×\\times×\\timesnone7979\+\+single best head1\-best9595ours \(full\)ours×\\times×\\times✓✓✓word✓7171\+\+DTW×\\times×\\timesnone7373\+\+no energy×\\times7171\+\+no silence✓✓✓117117Gradours \(full\)–––✓✓✓word✓7575\+\+DTW, word\-topo×\\times7676\+\+no word\-toponone138138\+\+no energy×\\times152152
TABLE IX:MMS\-FA \(wav2vec 2\.0\) CTC gradient alignment across the internal level the gradient is taken at, Buckeye\. The levels run from the convolutional feature encoder through the feature projection to the raw waveform at several pooling rates\. The feat\-proj linear row equals the MMS\-FA gradient number reported in[TableI](https://arxiv.org/html/2607.06831#S2.T1)\.Gradient w\.r\.t\.ms/frameGrid\[Hz\]WBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowConv 00\.31320013713748\.748\.7Conv 10\.62160013313349\.549\.5Conv 21\.2580013313349\.549\.5Conv 32\.540013213250\.050\.0Conv 4520013013050\.850\.8Conv 51010012812852\.352\.3Feat\-proj LayerNorm205012212255\.355\.3Feat\-proj Linear205012112158\.258\.2Raw waveform, pool 320205013513552\.152\.1Raw waveform, pool 80520014014048\.448\.4Raw waveform, pool 161100014214247\.547\.5Raw waveform, pool 10\.06251600013413447\.747\.7
TABLE X:Encoder depth the gradient is taken atfor gradient alignment\. Buckeye\. Rows run from log\-mel input through encoder to its output\. Whisper\-large\-v3 uses char targets \(32 layers; 1/4 etc\. = L8/16/24\) and FastConformer\-CTC subword targets \(17 layers; L4/9/13\)\. Off\. is mean signed boundary offset \(positive = late\)\.Gradientw\.r\.t\.Whisper\-large\-v3 \(char\)FastConformer\-CTCGrid\[Hz\]WBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowOff\.\[ms\]Grid\[Hz\]WBE\[ms\]↓\\downarrow≤\\leq50ms\[%\]↑\\uparrowOff\.\[ms\]Log\-mel in100100535366\.266\.2−10\-1010010015815830\.130\.1−80\-80Encoder in5050444474\.674\.6−9\-912\.512\.515215230\.030\.0−63\-63Encoder 1/4424276\.376\.3−9\-914514529\.429\.4−33\-33Encoder 1/2414178\.278\.2−9\-914614623\.923\.9\+73\+73Encoder 3/4393980\.080\.0−9\-922922912\.512\.5\+196\+196Encoder out424279\.079\.0−13\-133053054\.94\.9\+274\+274
TABLE XI:OWSM\-CTC alignment from each inter\-CTC block\(6/12/15/21\) and the final block \(27\)\. Gradient alignment vs\. posterior alignment\.Emit blockWBE\[ms\]↓\\downarrowTIMITBuckeyeGrad\.Post\.Grad\.Post\.614214212012014514511011012150150124124167167118118151601601251251831831191192116816812512519419412012027 \(final\)172172125125198198121121
TABLE XII:Cost of gradient alignment, single GPU, subset of TIMIT, as real\-time factor \(RTF\)×103\\times 10^\{3\}\(ms compute per second of audio\)\. Stages: Forward \(encoder→\\rightarrowlogits, shared with native methods\), prefix\-sum \(CTC/transducer only\), and backward \(encoder VJP, batched\)\. Grad\. align extra cost is prefix\-sum\+\+backward; Shared DP align is excluded\.ModelCompute time\[RTF×103\\times 10^\{3\}\]TypeNameModelPrefix\-sumFwdBwdFwdBwdCTCMMS\-FA66575733212212AEDWhisper\-large\-v32424277277n/aSpeech LLMPhi\-4\-MM2626216216
## IIIExperimental Setup
We evaluate sixteen models from the four families\. For reference we run two dedicated aligners, a GM\-HMM via MFA\[[20](https://arxiv.org/html/2607.06831#bib.bib56)\], and MMS\-FA\[[30](https://arxiv.org/html/2607.06831#bib.bib57),[19](https://arxiv.org/html/2607.06831#bib.bib86)\], a wav2vec 2\.0\[[5](https://arxiv.org/html/2607.06831#bib.bib70)\]CTC model built specifically for forced alignment\. For CTC we further use XLS\-R on phonemes\[[27](https://arxiv.org/html/2607.06831#bib.bib74),[4](https://arxiv.org/html/2607.06831#bib.bib73)\], Parakeet CTC\[[34](https://arxiv.org/html/2607.06831#bib.bib75)\], OWSM\-CTC\[[25](https://arxiv.org/html/2607.06831#bib.bib81),[26](https://arxiv.org/html/2607.06831#bib.bib80)\], and a streaming FastConformer\-CTC\[[23](https://arxiv.org/html/2607.06831#bib.bib76)\]\. For the transducers we use Parakeet RNN\-T\[[34](https://arxiv.org/html/2607.06831#bib.bib75)\]and TDT\[[44](https://arxiv.org/html/2607.06831#bib.bib77)\], and the streaming FastConformer\-RNN\-T\[[23](https://arxiv.org/html/2607.06831#bib.bib76)\]and Emformer\[[38](https://arxiv.org/html/2607.06831#bib.bib71),[45](https://arxiv.org/html/2607.06831#bib.bib85)\]\. The AED models are Whisper base and large\-v3\[[33](https://arxiv.org/html/2607.06831#bib.bib52)\], CrisperWhisper\[[49](https://arxiv.org/html/2607.06831#bib.bib53)\], and OWLS\-1B\[[8](https://arxiv.org/html/2607.06831#bib.bib83)\], and the speech LLMs are Voxtral\[[21](https://arxiv.org/html/2607.06831#bib.bib84)\], Phi\-4\-multimodal \(Phi\-4\-MM\)\[[1](https://arxiv.org/html/2607.06831#bib.bib47)\], and Canary\-Qwen\[[31](https://arxiv.org/html/2607.06831#bib.bib82)\]\. We align on TIMIT \(test\)\[[14](https://arxiv.org/html/2607.06831#bib.bib59)\]and on Buckeye spontaneous speech\[[28](https://arxiv.org/html/2607.06831#bib.bib60)\], both with gold word boundaries\. For Buckeye we use a55h subset stratified by speaker \(all speakers represented proportionally\), split at every≥1\\geq 1s inter\-word silence, then split any piece still longer than1818s at its largest internal gap333\[[18](https://arxiv.org/html/2607.06831#bib.bib32)\]drops long utterances, but we keep them by splitting\.\. We report the word\-boundary error \(WBE\), i\.e\. the mean per\-word start and end absolute error, averaged over all words in the corpus444\[[35](https://arxiv.org/html/2607.06831#bib.bib41)\]averages first per utterance and then over utterances\., and the accuracy at a collar, i\.e\. the fraction of all boundaries within5050ms over the corpus\. We compare each model against its own alignment, i\.e\. the CTC or transducer forced alignment on the model’s own emission \(the “posteriors”\), the AED cross\-attention DTW\[[6](https://arxiv.org/html/2607.06831#bib.bib51)\], and, for the speech LLMs, the analogous text\-to\-audio self\-attention, decoded by the same dynamic program as the gradient scores\. We select the attention heads by per\-head WBE on TIMIT development gold and average the top 8 heads\. In hypothesis mode, where the recognition differs from the reference, we match the words by identity and report an identity\-gated F1 at a5050ms collar together with the matched WBE\.
## IVResults
### Alignment quality across model families
[TableI](https://arxiv.org/html/2607.06831#S2.T1)compares the gradient alignment versus each model’s own native or attention aligner, across CTC, AED, transducer and speech\-LLM families\. Gradient alignment is competitive on every family, beats the native alignment where it is weak \(the streaming models, Canary\-Qwen\), and at the starred best encoder depth it also beats the Whisper\-large\-v3 cross\-attention\. The streaming posteriors’ large error is largely a systematic emission\-delay bias rather than scatter \([TablesIII](https://arxiv.org/html/2607.06831#S2.T3)and[1](https://arxiv.org/html/2607.06831#S1.F1)\); the same table shows that gradient alignment shifts both boundaries together \(a positional lead\) rather than shrinking the word like the posteriors do\.
[TableII](https://arxiv.org/html/2607.06831#S2.T2)aligns each model’s own recognition \(identity\-gated F1 at a5050ms collar and matched WBE\), showing that gradient alignment remains usable when the hypothesis differs from the reference\.
### Tokenization and the attention baseline
[TableIV](https://arxiv.org/html/2607.06831#S2.T4)justifies the tokenization we use throughout: the gradient alignment uses character targets for the AED and speech\-LLM families and subword targets for CTC and transducers, while every attention or posterior baseline uses the model’s native subword tokens\. Character targets improve the gradient markedly for the autoregressive families \(on Whisper\-large\-v3 and Voxtral the char gradient is well below the subword gradient\), but they do not help the model\-native attention aligners, and for the non\-autoregressive CTC and transducer, which have no per\-character acoustic emission, char\-level collapses: char targets are consistently worse than subword for both the gradient and the posteriors, in the worst cases by an order of magnitude\.
### Ablations
The grad\. score per\-token reduction \([TableV](https://arxiv.org/html/2607.06831#S2.T5)\) moves WBE by only a few ms across thepp\-norms, while the plain sum is clearly worse, and gradient versus gradient\-×\\times\-input\[[39](https://arxiv.org/html/2607.06831#bib.bib66),[3](https://arxiv.org/html/2607.06831#bib.bib67)\]is within∼\\sim1–3 ms with neither winning consistently\. Multi\-pass attribution \(SmoothGrad\[[41](https://arxiv.org/html/2607.06831#bib.bib68)\], VarGrad\[[2](https://arxiv.org/html/2607.06831#bib.bib69)\], Integrated\[[42](https://arxiv.org/html/2607.06831#bib.bib65)\]and Expected Gradients\[[12](https://arxiv.org/html/2607.06831#bib.bib72)\]\) does not beat the single\-pass gradient \(numbers omitted here\)\. Most decoding options behave consistently \([TablesVI](https://arxiv.org/html/2607.06831#S2.T6),[VII](https://arxiv.org/html/2607.06831#S2.T7)and[VIII](https://arxiv.org/html/2607.06831#S2.T8)\): the gradient is active even in silence and needs the energy\-aware \(λ≈2\\lambda\\approx 2\) or z\-score blank, whereas the attention is already silence\-aware and a constant blank suffices; the z\-score blank is the energy\-free option that matches the energy one for gradients\. The audio\-energy token\-weighting exponent matters little around our default and behaves consistently across the blank schemes \([TableVII](https://arxiv.org/html/2607.06831#S2.T7)\)\. For the cross\-attention DTW, the dominant gain over the original Whisper heuristic is the gold\-tuned head selection \(its curated set includes some near\-useless heads\); the Whisper\-DTW setting is exactly our decoder with the log\-compression and the blank state removed, and energy weighting helps the gradient but not the cross\-attention DTW\.
### Where the gradient is taken
For MMS\-FA, which is a wav2vec 2\.0 model\[[5](https://arxiv.org/html/2607.06831#bib.bib70)\]\([TableIX](https://arxiv.org/html/2607.06831#S2.T9)\), the gradient localizes best at the feature\-projection output \(20 ms encoder grid\); finer levels \(the convolutional feature encoder, or the raw waveform pooled down to sample resolution\) do not help, and the raw\-waveform gradient is noisier\.[TableX](https://arxiv.org/html/2607.06831#S2.T10)sweeps the depth for Whisper\-large\-v3 and the FastConformer\-CTC, from the log\-mel input through the encoder to its output \(the activation just before the decoder / CTC linear\): the sharpest gradient is at an intermediate depth \(three quarters of the Whisper encoder, the L24 of the starred row in[TableI](https://arxiv.org/html/2607.06831#S2.T1); a quarter for the FastConformer\-CTC\), not the input, and in the streaming FastConformer\-CTC the encoder’s time shift accumulates with depth, so its output gradient ends as delayed as its posteriors \([Fig\.1](https://arxiv.org/html/2607.06831#S1.F1)\)\.
### Speech\-LLM prompt influence
Gradient alignment is robust to the prompt used across all the speech LLMs and across the gradient\-vs\-self\-attention signal \(numbers omitted\)\.
### Inter\-CTC
OWSM\-CTC has intermediate CTC losses\[[25](https://arxiv.org/html/2607.06831#bib.bib81),[26](https://arxiv.org/html/2607.06831#bib.bib80)\], which we can use to take the gradient at each block of the encoder \([TableXI](https://arxiv.org/html/2607.06831#S2.T11)\)\. Interestingly, the earlier blocks are better than the later ones, both for the gradient and for the native CTC alignment\.
### Time stretch
Under audio time\-stretch \(numbers omitted\), the same upsampling mechanism acts with opposite sign: it hurts a fine\-grid model \(MMS\-FA\) but helps a coarse\-grid one \(Voxtral\) at mild factors\.
### Compute cost
The only extra cost over the native methods is the per\-token backward \([TableXII](https://arxiv.org/html/2607.06831#S2.T12)\); the forward and the DP decode are shared, and for the CTC the per\-token prefix\-score lattice backward dominates\.
## VConclusion
Gradient alignment produces a usable word alignment for every model and family we tried, including the speech LLMs that have no built\-in word aligner\. It is usually a little behind a strong native aligner, but clearly better where that alignment is weak, as on the streaming models and Canary\-Qwen, and on Whisper\-large\-v3 it even beats its cross\-attention at the best encoder depth\. Our decoder also improves on Whisper’s cross\-attention DTW itself \([TableVIII](https://arxiv.org/html/2607.06831#S2.T8)\)\. Its input\-grid resolution lets finer character targets improve accuracy, which the coarse attention grid cannot exploit\. The per\-token backward makes it much more expensive than a single forward, so we do not propose it as a practical aligner\. The result is rather that one training\-free input saliency aligns every differentiable ASR model across all four families, in some cases even better than the model’s own alignment\. Beyond alignment, the gradient also serves as an analysis tool\.
## Acknowledgment
This work was partially supported by NeuroSys, which as part of the initiative “Clusters4Future” is funded by the Federal Ministry of Research, Technology and Space BMFTR \(funding IDs 03ZU2106DA and 03ZU2106DD\), and by the project RESCALE within the programAI Lighthouse Projects for the Environment, Climate, Nature and Resourcesfunded by the Federal Ministry for the Environment, Nature Conservation, Nuclear Safety and Consumer Protection \(BMUV\), funding ID: 67KI32006A\. The authors gratefully acknowledge the computing time provided to them at the NHR Center NHR4CES at RWTH Aachen University \(project number p0023999\)\. This is funded by the Federal Ministry of Education and Research, and the state governments participating on the basis of the resolutions of the GWK for national high performance computing at universities \([www\.nhr\-verein\.de/unsere\-partner](https://arxiv.org/html/2607.06831v1/www.nhr-verein.de/unsere-partner)\)\.
## AI\-Generated Content Disclosure
We used AI assistants \(large language models\) substantially in this work, to help set up and run the experiments, to debug and write parts of the code, and to draft and edit parts of this paper\. All results were verified by the authors, who take full responsibility for the content\.
## References
- \[1\]A\. Abouelenin, A\. Ashfaq, A\. Atkinson, H\. Awadalla, N\. Bach, J\. Bao, A\. Benhaim, M\. Cai, V\. Chaudhary, C\. Chen, D\. Chen, D\. Chen, J\. Chen, W\. Chen, Y\. Chen, Y\. Chen, Q\. Dai, X\. Dai, R\. Fan, M\. Gao, M\. Gao, A\. Garg, A\. Goswami, J\. Hao, A\. Hendy, Y\. Hu, X\. Jin, M\. Khademi, D\. Kim, Y\. J\. Kim, G\. Lee, J\. Li, Y\. Li, C\. Liang, X\. Lin, Z\. Lin, M\. Liu, Y\. Liu, G\. Lopez, C\. Luo, P\. Madan, V\. Mazalov, A\. Mitra, A\. Mousavi, A\. Nguyen, J\. Pan, D\. Perez\-Becker, J\. Platin, T\. Portet, K\. Qiu, B\. Ren, L\. Ren, S\. Roy, N\. Shang, Y\. Shen, S\. Singhal, S\. Som, X\. Song, T\. Sych, P\. Vaddamanu, S\. Wang, Y\. Wang, Z\. Wang, H\. Wu, H\. Xu, W\. Xu, Y\. Yang, Z\. Yang, D\. Yu, I\. Zabir, J\. Zhang, L\. L\. Zhang, Y\. Zhang, and X\. Zhou\(2025\)Phi\-4\-mini technical report: compact yet powerful multimodal language models via mixture\-of\-loras\.Note:ArXiv 2503\.01743External Links:[Link](https://arxiv.org/abs/2503.01743)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[2\]J\. Adebayo, J\. Gilmer, M\. Muelly, I\. Goodfellow, M\. Hardt, and B\. Kim\(2018\)Sanity checks for saliency maps\.InProc\. NeurIPS,S\. Bengio, H\. Wallach, H\. Larochelle, K\. Grauman, N\. Cesa\-Bianchi, and R\. Garnett \(Eds\.\),Vol\.31,pp\.\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2018/file/294a8ed24b1ad22ec2e7efea049b8737-Paper.pdf)Cited by:[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px3.p1.4)\.
- \[3\]M\. Ancona, E\. Ceolini, C\. Öztireli, and M\. Gross\(2018\)Towards better understanding of gradient\-based attribution methods for deep neural networks\.InProc\. ICLR,External Links:[Link](https://openreview.net/forum?id=Sy21R9JAW)Cited by:[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px3.p1.4)\.
- \[4\]A\. Babu, C\. Wang, A\. Tjandra, K\. Lakhotia, Q\. Xu, N\. Goyal, K\. Singh, P\. von Platen, Y\. Saraf, J\. Pino, A\. Baevski, A\. Conneau, and M\. Auli\(2022\)XLS\-R: self\-supervised cross\-lingual speech representation learning at scale\.InProc\. Interspeech,pp\. 2278–2282\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2022-143),ISSN 2958\-1796Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[5\]A\. Baevski, Y\. Zhou, A\. Mohamed, and M\. Auli\(2020\)wav2vec 2\.0: a framework for self\-supervised learning of speech representations\.InProc\. NeurIPS,H\. Larochelle, M\. Ranzato, R\. Hadsell, M\.F\. Balcan, and H\. Lin \(Eds\.\),Vol\.33,pp\. 12449–12460\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2020/file/92d1e1eb1cd6f9fba3227870bb6d7f07-Paper.pdf)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5),[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px4.p1.1)\.
- \[6\]M\. Bain, J\. Huh, T\. Han, and A\. Zisserman\(2023\)WhisperX: time\-accurate speech transcription of long\-form audio\.InProc\. Interspeech,pp\. 4489–4493\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2023-78),ISSN 2958\-1796Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[7\]W\. Chan, N\. Jaitly, Q\. V\. Le, and O\. Vinyals\(2016\)Listen, attend and spell: a neural network for large vocabulary conversational speech recognition\.InProc\. IEEE ICASSP,pp\. 4960–4964\.Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[8\]W\. Chen, J\. Tian, Y\. Peng, B\. Yan, C\. H\. Yang, and S\. Watanabe\(2025\)OWLS: scaling laws for multilingual speech recognition and translation models\.InProc\. ICML,External Links:[Link](https://openreview.net/forum?id=xnPW7yYomF)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[9\]J\. K\. Chorowski, D\. Bahdanau, D\. Serdyuk, K\. Cho, and Y\. Bengio\(2015\)Attention\-based models for speech recognition\.InNIPS,pp\. 577–585\.Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[10\]Y\. Chu, J\. Xu, X\. Zhou, Q\. Yang, S\. Zhang, Z\. Yan, C\. Zhou, and J\. Zhou\(2023\)Qwen\-Audio: advancing universal audio understanding via unified large\-scale audio\-language models\.Note:arXiv:2311\.07919Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[11\]S\. Ding, H\. Xu, and P\. Koehn\(2019\-08\)Saliency\-driven word alignment interpretation for neural machine translation\.InProceedings of the Fourth Conference on Machine Translation \(Volume 1: Research Papers\),O\. Bojar, R\. Chatterjee, C\. Federmann, M\. Fishel, Y\. Graham, B\. Haddow, M\. Huck, A\. J\. Yepes, P\. Koehn, A\. Martins, C\. Monz, M\. Negri, A\. Névéol, M\. Neves, M\. Post, M\. Turchi, and K\. Verspoor \(Eds\.\),Florence, Italy,pp\. 1–12\.External Links:[Link](https://aclanthology.org/W19-5201/),[Document](https://dx.doi.org/10.18653/v1/W19-5201)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p2.1)\.
- \[12\]G\. Erion, J\. D\. Janizek, P\. Sturmfels, S\. M\. Lundberg, and S\. Lee\(2021\-07\)Improving performance of deep learning models with axiomatic attribution priors and expected gradients\.Nature Machine Intelligence3\(7\),pp\. 620–631\.External Links:ISSN 2522\-5839,[Document](https://dx.doi.org/10.1038/s42256-021-00343-w)Cited by:[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px3.p1.4)\.
- \[13\]M\. Gales and S\. Young\(2008\-01\)The application of hidden Markov models in speech recognition\.Found\. Trends Signal Process\.1\(3\),pp\. 195–304\.External Links:ISSN 1932\-8346,[Document](https://dx.doi.org/10.1561/2000000004)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[14\]J\. S\. Garofolo, L\. F\. Lamel, W\. M\. Fisher, J\. G\. Fiscus, D\. S\. Pallett, N\. L\. Dahlgren, and V\. Zue\(1993\)TIMIT acoustic\-phonetic continuous speech corpus\.Note:Linguistic Data Consortium, Philadelphia, LDC93S1External Links:[Document](https://dx.doi.org/10.35111/17gk-bn40)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[15\]A\. Graves, S\. Fernández, F\. Gomez, and J\. Schmidhuber\(2006\)Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks\.InProceedings of the 23rd international conference on Machine learning,IDSIA,pp\. 369–376\.Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[16\]A\. Graves\(2012\)Sequence transduction with recurrent neural networks\.Note:ArXiv:1211\.3711, ICML Representation Learning WorkshopCited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[17\]T\. Hori, S\. Watanabe, and J\. Hershey\(2017\-07\)Joint CTC/attention decoding for end\-to\-end speech recognition\.InProc\. ACL,R\. Barzilay and M\. Kan \(Eds\.\),Vancouver, Canada,pp\. 518–529\.External Links:[Link](https://aclanthology.org/P17-1048),[Document](https://dx.doi.org/10.18653/v1/P17-1048)Cited by:[§II](https://arxiv.org/html/2607.06831#S2.p2.1)\.
- \[18\]R\. Huang, X\. Zhang, Z\. Ni, L\. Sun, M\. Hira, J\. Hwang, V\. Manohar, V\. Pratap, M\. Wiesner, S\. Watanabe, D\. Povey, and S\. Khudanpur\(2024\)Less peaky and more accurate CTC forced alignment by label priors\.InProc\. IEEE ICASSP,pp\. 11831–11835\.Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[footnote 3](https://arxiv.org/html/2607.06831#footnote3)\.
- \[19\]J\. Hwang, M\. Hira, C\. Chen, X\. Zhang, Z\. Ni, G\. Sun, P\. Ma, R\. Huang, V\. Pratap, Y\. Zhang, A\. Kumar, C\. Yu, C\. Zhu, C\. Liu, J\. Kahn, M\. Ravanelli, P\. Sun, S\. Watanabe, Y\. Shi, and Y\. Tao\(2023\)TorchAudio 2\.1: advancing speech recognition, self\-supervised learning, and audio processing components for PyTorch\.InProc\. IEEE ASRU,Vol\.,pp\. 1–9\.External Links:[Document](https://dx.doi.org/10.1109/ASRU57964.2023.10389648)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[20\]M\. McAuliffe, M\. Socolof, S\. Mihuc, M\. Wagner, and M\. Sonderegger\(2017\)Montreal Forced Aligner: trainable text\-speech alignment using Kaldi\.InProc\. Interspeech,pp\. 498–502\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2017-1386),ISSN 2958\-1796Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[21\]Mistral AI\(2025\)Voxtral\.Note:arXiv:2507\.13264Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[22\]B\. Mu, X\. Shi, X\. Wang, H\. Liu, J\. Xu, and L\. Xie\(2026\)LLM\-ForcedAligner: a non\-autoregressive and accurate llm\-based forced aligner for multilingual and long\-form speech\.Note:ArXiv 2601\.18220External Links:2601\.18220,[Link](https://arxiv.org/abs/2601.18220)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[23\]V\. Noroozi, S\. Majumdar, A\. Kumar, J\. Balam, and B\. Ginsburg\(2024\)Stateful conformer with cache\-based inference for streaming automatic speech recognition\.InProc\. IEEE ICASSP,Vol\.,pp\. 12041–12045\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP48485.2024.10446861)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[24\]S\. Papi and L\. Bentivogli\(2026\)DOA: training\-free decoder\-only attention policy for long\-form simultaneous translation with SpeechLLMs\.Note:ArXiv 2605\.31432External Links:2605\.31432,[Link](https://arxiv.org/abs/2605.31432)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[25\]Y\. Peng, S\. Muhammad, Y\. Sudo, W\. Chen, J\. Tian, C\. Lin, and S\. Watanabe\(2025\)OWSM v4: improving open whisper\-style speech models via data scaling and cleaning\.InProc\. Interspeech,Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5),[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px6.p1.1)\.
- \[26\]Y\. Peng, Y\. Sudo, M\. Shakeel, and S\. Watanabe\(2024\-08\)OWSM\-CTC: an open encoder\-only speech foundation model for speech recognition, translation, and language identification\.InProc\. ACL,pp\. 10192–10209\.External Links:[Link](https://aclanthology.org/2024.acl-long.549),[Document](https://dx.doi.org/10.18653/v1/2024.acl-long.549)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5),[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px6.p1.1)\.
- \[27\]V\. Phy\(2022\)Automatic phoneme recognition on TIMIT dataset with Wav2Vec 2\.0\.Hugging Face\.External Links:[Document](https://dx.doi.org/10.57967/hf/0125),[Link](https://huggingface.co/vitouphy/wav2vec2-xls-r-300m-timit-phoneme)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[28\]M\. A\. Pitt, K\. Johnson, E\. Hume, S\. Kiesling, and W\. Raymond\(2005\)The Buckeye corpus of conversational speech: labeling conventions and a test of transcriber reliability\.Speech Communication45\(1\),pp\. 89–95\.External Links:ISSN 0167\-6393,[Document](https://dx.doi.org/10.1016/j.specom.2004.09.001)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[29\]R\. Prabhavalkar, T\. Hori, T\. N\. Sainath, R\. Schlüter, and S\. Watanabe\(2023\)End\-to\-end speech recognition: a survey\.IEEE/ACM Trans\. Audio, Speech, and Language Processing32,pp\. 325–351\.External Links:[Document](https://dx.doi.org/10.1109/TASLP.2023.3328283)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[30\]V\. Pratap, A\. Tjandra, B\. Shi, P\. Tomasello, A\. Babu, S\. Kundu, A\. Elkahky, Z\. Ni, A\. Vyas, M\. Fazel\-Zarandi, A\. Baevski, Y\. Adi, X\. Zhang, W\. Hsu, A\. Conneau, and M\. Auli\(2024\-01\)Scaling speech technology to 1,000\+ languages\.Journal of Machine Learning Research25\(1\)\.External Links:ISSN 1532\-4435Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[31\]K\. C\. Puvvada, P\. Żelasko, H\. Huang, O\. Hrinchuk, N\. R\. Koluguri, K\. Dhawan, S\. Majumdar, E\. Rastorgueva, Z\. Chen, V\. Lavrukhin, J\. Balam, and B\. Ginsburg\(2024\)Less is more: accurate speech recognition & translation without web\-scale data\.InProc\. Interspeech,pp\. 3964–3968\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2024-2294),ISSN 2958\-1796Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[32\]L\. R\. Rabiner\(1989\)A tutorial on hidden Markov models and selected applications in speech recognition\.Proc\. of the IEEE77\(2\),pp\. 257–286\.External Links:[Document](https://dx.doi.org/10.1109/5.18626)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[33\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. Mcleavey, and I\. Sutskever\(2023\-23–29 Jul\)Robust speech recognition via large\-scale weak supervision\.InProc\. ICML,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 28492–28518\.External Links:[Link](https://proceedings.mlr.press/v202/radford23a.html)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[34\]D\. Rekesh, N\. R\. Koluguri, S\. Kriman, S\. Majumdar, V\. Noroozi, H\. Huang, O\. Hrinchuk, K\. Puvvada, A\. Kumar, J\. Balam, and B\. Ginsburg\(2023\)Fast conformer with linearly scalable attention for efficient speech recognition\.InProc\. IEEE ASRU,Vol\.,pp\. 1–8\.External Links:[Document](https://dx.doi.org/10.1109/ASRU57964.2023.10389701)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[35\]R\. Rousso, E\. Cohen, J\. Keshet, and E\. Chodroff\(2024\)Tradition or innovation: a comparison of modern asr methods for forced alignment\.InProc\. Interspeech,pp\. 1525–1529\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2024-429),ISSN 2958\-1796Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[footnote 4](https://arxiv.org/html/2607.06831#footnote4)\.
- \[36\]R\. Schmitt, A\. Zeyer, M\. Zeineldeen, R\. Schlüter, and H\. Ney\(2025\-04\)The conformer encoder may reverse the time dimension\.InProc\. IEEE ICASSP,Hyderabad, India\.Note:Preprint ArXiv:2501\.04521Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p2.1),[§II](https://arxiv.org/html/2607.06831#S2.p2.1)\.
- \[37\]R\. Schmitt, A\. Zeyer, M\. Zeineldeen, R\. Schlüter, and H\. Ney\(2026\-03\)LLMs and speech: integration vs\. combination\.Note:Arxiv:2603\.15045External Links:[Link](https://arxiv.org/abs/2603.15045)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[38\]Y\. Shi, Y\. Wang, C\. Wu, C\. Yeh, J\. Chan, F\. Zhang, D\. Le, and M\. Seltzer\(2021\)Emformer: efficient memory transformer based acoustic model for low latency streaming speech recognition\.InProc\. IEEE ICASSP,Vol\.,pp\. 6783–6787\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP39728.2021.9414560)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[39\]A\. Shrikumar, P\. Greenside, and A\. Kundaje\(2017\-08\)Learning important features through propagating activation differences\.InProc\. ICML,D\. Precup and Y\. W\. Teh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.70,pp\. 3145–3153\.External Links:[Link](https://proceedings.mlr.press/v70/shrikumar17a.html)Cited by:[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px3.p1.4)\.
- \[40\]K\. Simonyan, A\. Vedaldi, and A\. Zisserman\(2014\)Deep inside convolutional networks: visualising image classification models and saliency maps\.Note:ArXiv:1312\.6034; ICLR Workshop TrackCited by:[§I](https://arxiv.org/html/2607.06831#S1.p2.1)\.
- \[41\]D\. Smilkov, N\. Thorat, B\. Kim, F\. Viégas, and M\. Wattenberg\(2017\)SmoothGrad: removing noise by adding noise\.Note:arXiv:1706\.03825Cited by:[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px3.p1.4)\.
- \[42\]M\. Sundararajan, A\. Taly, and Q\. Yan\(2017\-08\)Axiomatic attribution for deep networks\.InProc\. ICML,D\. Precup and Y\. W\. Teh \(Eds\.\),Proceedings of Machine Learning Research, Vol\.70,pp\. 3319–3328\.External Links:[Link](https://proceedings.mlr.press/v70/sundararajan17a.html)Cited by:[§IV](https://arxiv.org/html/2607.06831#S4.SS0.SSS0.Px3.p1.4)\.
- \[43\]H\. Wang, C\. Du, Y\. Guo, S\. Wang, X\. Chen, and K\. Yu\(2024\)Attention\-constrained inference for robust decoder\-only text\-to\-speech\.Note:ArXiv 2404\.19723External Links:2404\.19723,[Link](https://arxiv.org/abs/2404.19723)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[44\]H\. Xu, F\. Jia, S\. Majumdar, H\. Huang, S\. Watanabe, and B\. Ginsburg\(2023\-23–29 Jul\)Efficient sequence transduction by jointly predicting tokens and durations\.InProc\. ICML,A\. Krause, E\. Brunskill, K\. Cho, B\. Engelhardt, S\. Sabato, and J\. Scarlett \(Eds\.\),Proceedings of Machine Learning Research, Vol\.202,pp\. 38462–38484\.External Links:[Link](https://proceedings.mlr.press/v202/xu23g.html)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[45\]Y\. Yang, M\. Hira, Z\. Ni, A\. Astafurov, C\. Chen, C\. Puhrsch, D\. Pollack, D\. Genzel, D\. Greenberg, E\. Z\. Yang, J\. Lian, J\. Hwang, J\. Chen, P\. Goldsborough, S\. Narenthiran, S\. Watanabe, S\. Chintala, and V\. Quenneville\-Bélair\(2022\)TorchAudio: building blocks for audio and speech processing\.InProc\. IEEE ICASSP,Vol\.,pp\. 6982–6986\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP43922.2022.9747236)Cited by:[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.
- \[46\]S\. Yeh, Y\. Meng, and H\. Tang\(2025\)Whisper has an internal word aligner\.InProc\. IEEE ASRU,Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[§I](https://arxiv.org/html/2607.06831#S1.p2.1),[§II](https://arxiv.org/html/2607.06831#S2.p2.1)\.
- \[47\]A\. Zeyer, K\. Irie, R\. Schlüter, and H\. Ney\(2018\-09\)Improved training of end\-to\-end attention models for speech recognition\.InInterspeech,Hyderabad, India\.Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[48\]D\. Zhang, S\. Li, X\. Zhang, J\. Zhan, P\. Wang, Y\. Zhou, and X\. Qiu\(2023\-12\)SpeechGPT: empowering large language models with intrinsic cross\-modal conversational abilities\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 15757–15773\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.1055/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.1055)Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1)\.
- \[49\]M\. Zusag, L\. Wagner, and B\. Thallinger\(2024\)CrisperWhisper: accurate timestamps on verbatim speech transcriptions\.InProc\. Interspeech,pp\. 1265–1269\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2024-731),ISSN 2958\-1796Cited by:[§I](https://arxiv.org/html/2607.06831#S1.p1.1),[§III](https://arxiv.org/html/2607.06831#S3.p1.5)\.Similar Articles
Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models
This paper introduces an instruction-free alignment-only method for building large audio-language models by freezing the LLM and audio encoder, training only a lightweight projector on self-generated data, achieving competitive performance with less data than traditional multi-stage pipelines.
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
This paper proposes a training-free pipeline for pronunciation transcription that integrates lexical resources and acoustic models to achieve low error rates and high efficiency, outperforming baselines including commercial multimodal LLMs.
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
This paper introduces LoRA-GA2, a fine-tuning algorithm that leverages multi-step gradient information to improve the performance of Low-Rank Adaptation for large language models, achieving better results on benchmarks while preserving efficiency.
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
This paper investigates the distributional gap between synthetic and real speech in LLM-based ASR systems, identifies where the LLM separates them, and proposes using layer-selection and RIR augmentation to match real-data baselines with less real data.
Streaming Speech-to-Text Translation with a SpeechLLM
Presents a SpeechLLM architecture for streaming speech-to-text translation that adaptively decides when to output tokens based on audio, achieving 1-2 second latency with quality close to non-streaming baselines.