Tag
This paper investigates the effects of phonetic versus character targets and selective state-space models (Mamba) for intracortical brain-to-text decoding, finding that a GRU-based recurrent decoder remains the strongest performer on the Brain-to-Text '25 benchmark.
This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.
This paper introduces a gradient-based speech-to-text alignment method applicable to any differentiable ASR model, including CTC, transducer, attention-based encoder-decoder, and speech large language models, requiring no training or model modification.