Tag
Audio8 ASR Infinite is a native streaming speech recognition model with selectable audio clock and configurable transcription delay, supporting unlimited-length audio transcription without drifting.
NetEase Youdao AI is open-sourcing Confucius4-R2T2, a 1.7B frontier real-time streaming ASR model designed for voice agents, featuring low latency, high accuracy, and configurable decoding chunks.
This paper introduces X2-Turn, a frame-synchronous dual-head model that jointly performs streaming ASR and turn state prediction on shared representations, improving turn-taking accuracy and latency in spoken dialogue systems.
This paper presents an engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin, achieving 42.97% and 33.98% WER on internal sets for Kikuyu and Dholuo, respectively, through data-centric techniques including corpus auditing, normalization, and streaming evaluation.
This paper investigates the impact of data scale versus latency on cross-lingual transfer for streaming ASR, finding that multilingual initialization benefits are data-limited, not latency-limited, and diminish as target-language data increases.
Proposes a non-autoregressive scoring method for punctuation restoration in streaming ASR that preserves the input transcript and outperforms prompt-based and fine-tuned baselines under a limited lookahead budget.
A routing-based approach for real-time multilingual ASR that uses smaller monolingual models with a rollback mechanism to handle language switches, achieving ~13% WER on inter-utterance code-switching and open-sourcing the system.