Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Hugging Face Daily Papers Papers

Summary

This paper introduces Pruned CTC, a method that reduces memory usage in Connectionist Temporal Classification training for large-vocabulary automatic speech recognition by restricting alignment computation, proving equivalence to full-vocabulary CTC and demonstrating strong performance with LLM adaptations in offline and streaming settings.

Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1times with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10times faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
Original Article
View Cached Full Text

Cached at: 09/29/26, 04:11 PM

Paper page - Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Source: https://huggingface.co/papers/2609.33645 Authors:

,

,

,

,

,

,

,

,

,

,

,

Abstract

Connectionisttemporalclassification(CTC)naturallysupportsofflineandstreamingspeechrecognitionwithutterance-levelsupervision,butconventionalimplementationsmaterializeframe-by-vocabularyactivationsinmemory,makingCTCtrainingwithnativeLLMvocabulariesprohibitivelymemory-intensive.AkeyobservationisthateveryvalidCTCalignmentusesonlytargettokensandblank,andtheirunionacrossabatchtypicallyformsasmallsubsetofthefullvocabulary.WeintroducePrunedCTC,whichrestrictsalignmentcomputationtothissubsetwhileretainingfull-vocabularynormalization.Weprovethatthisvocabularyreductionisexactlyequivalenttofull-vocabularyCTCinlossandgradients.Head-and-lossactivationmemorynolongerscaleslinearlywithvocabularysize.Wefurtherapplyfinite-beamalignmentpruning.BuildingonPrunedCTC,wedevelopLLM-CTC,whichadaptspretrainedLLMsfornon-autoregressiveASRwhileretainingcausalattentionandnativevocabularies,andextendittobounded-historystreaming,avoidingchunk-levelspeech--textalignments.Experimentsshowthat,withZipformer-Mencoderand180Kvocabulary,PrunedCTCreducesfull-stepmemoryby5.1timeswithonly17%step-timeoverhead.Acrossthreecorpora,itmatchesstandardCTCaccuracy.OnGigaSpeech,acrosssixQwen3modelsizesfrom0.6Bto32B,LLM-CTCremainswithin7%relativeWERofLLM-CEwith7to10timesfasterrecognition;whenfine-tuningQwen3-ASRforbounded-historystreaming,LLM-CTCremainswithin3%relativeWERofmatchedofflinemodelsonthetestset.Together,theseresultsestablishPrunedCTCasascalablesequenceobjectivefornative-vocabularyLLMASRacrossofflineandstreamingsettings.

View arXiv pageView PDFGitHub21Add to collection

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2609.33645 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2609.33645 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2609.33645 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Small LLMs: Pruning vs. Training from Scratch

arXiv cs.LG

This paper empirically compares pruning vs. training small language models from scratch, finding that pruning provides a strong advantage under limited token budgets but that the advantage diminishes as training scales, especially with coarse pruning.