Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Summary
This paper introduces Pruned CTC, a method that reduces memory usage in Connectionist Temporal Classification training for large-vocabulary automatic speech recognition by restricting alignment computation, proving equivalence to full-vocabulary CTC and demonstrating strong performance with LLM adaptations in offline and streaming settings.
View Cached Full Text
Cached at: 09/29/26, 04:11 PM
Paper page - Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Source: https://huggingface.co/papers/2609.33645 Authors:
,
,
,
,
,
,
,
,
,
,
,
Abstract
Connectionisttemporalclassification(CTC)naturallysupportsofflineandstreamingspeechrecognitionwithutterance-levelsupervision,butconventionalimplementationsmaterializeframe-by-vocabularyactivationsinmemory,makingCTCtrainingwithnativeLLMvocabulariesprohibitivelymemory-intensive.AkeyobservationisthateveryvalidCTCalignmentusesonlytargettokensandblank,andtheirunionacrossabatchtypicallyformsasmallsubsetofthefullvocabulary.WeintroducePrunedCTC,whichrestrictsalignmentcomputationtothissubsetwhileretainingfull-vocabularynormalization.Weprovethatthisvocabularyreductionisexactlyequivalenttofull-vocabularyCTCinlossandgradients.Head-and-lossactivationmemorynolongerscaleslinearlywithvocabularysize.Wefurtherapplyfinite-beamalignmentpruning.BuildingonPrunedCTC,wedevelopLLM-CTC,whichadaptspretrainedLLMsfornon-autoregressiveASRwhileretainingcausalattentionandnativevocabularies,andextendittobounded-historystreaming,avoidingchunk-levelspeech--textalignments.Experimentsshowthat,withZipformer-Mencoderand180Kvocabulary,PrunedCTCreducesfull-stepmemoryby5.1timeswithonly17%step-timeoverhead.Acrossthreecorpora,itmatchesstandardCTCaccuracy.OnGigaSpeech,acrosssixQwen3modelsizesfrom0.6Bto32B,LLM-CTCremainswithin7%relativeWERofLLM-CEwith7to10timesfasterrecognition;whenfine-tuningQwen3-ASRforbounded-historystreaming,LLM-CTCremainswithin3%relativeWERofmatchedofflinemodelsonthetestset.Together,theseresultsestablishPrunedCTCasascalablesequenceobjectivefornative-vocabularyLLMASRacrossofflineandstreamingsettings.
View arXiv pageView PDFGitHub21Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.33645 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.33645 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.33645 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
This paper introduces CATS, a cascaded adaptive tree speculation framework designed to accelerate LLM inference on memory-constrained edge devices by optimizing memory usage while maintaining high token acceptance rates.
Batch-wise Adaptive Pruning: Periodic Neuron Activation-Aware Weight Pruning for Language Reasoning Model
This paper proposes a training-free adaptive pruning method for large reasoning models during batched inference, using periodic top-k selection and activation memory to improve accuracy and computational efficiency.
Small LLMs: Pruning vs. Training from Scratch
This paper empirically compares pruning vs. training small language models from scratch, finding that pruning provides a strong advantage under limited token budgets but that the advantage diminishes as training scales, especially with coarse pruning.
Correlation-Aware Structured Pruning for Large Language Models
The paper proposes a correlation-aware structured pruning method for large language models, explicitly modeling cross-unit dependencies to improve pruning decisions and achieve competitive accuracy-efficiency trade-offs.
Constrained CTC Decoding for Efficient Diacritic Restoration
This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.