Tag
Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.
NetEase Youdao has released two AI models: a streaming ASR model with 2B parameters and a Chinese-English simultaneous translation model with 14B parameters, both emphasizing real-time performance.
This paper proposes a practical recipe for semi-supervised federated ASR using online pseudo-labels with server update stabilization, demonstrating significant improvements over prior methods in both in-domain and cross-domain settings.
The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.
The Bairong system presents a dynamic question-aware evidence routing approach for multilingual conversational spoken question answering, achieving improved performance by effectively integrating transcript and audio evidence in the MLC-SLM 2026 Challenge.
The article highlights how aggregate Word Error Rate (WER) metric fails to capture critical errors in voice agents for production, such as misrecognized bank codes and multilingual speech, necessitating per-field tracking for real-world applications like banking and collections.
R2T2 is a low-latency and high-accuracy real-time speech recognition model that processes audio in small chunks and commits text without revision, suitable for applications like live captioning and translation.
The article reflects on the need for stability over speed in ASR for voice agents, discussing the 'Confucius r2t2' model that integrates wait/commit mechanisms to enhance reliability.
The paper proposes T-SANDHI, a Tone Sandhi-aware Adaptive Network for low-resource Taiwanese Hokkien speech recognition, which explicitly decouples tonal variations to improve accuracy on top of a Whisper backbone.
NetEase Youdao open-sourced Confucius4-R2T2, a streaming ASR model for voice agents that incrementally processes speech and emits only committed text to prevent state corruption.
Nari Labs introduces Qwen3-TTS and Qwen3-ASR models, providing high accuracy, low latency, and cost-effectiveness in a free public beta, alongside optimized APIs and services for production deployment.
This paper introduces ASCIL, a post-ASR framework that adapts to user feedback and contextual signals to correct false wake-up activations in AI assistants, achieving significant error reduction with low latency.
The paper reports human annotation of Quran recitation transcripts to distinguish mistakes from repetitions and repairs, with an evaluator scoring labels and positions, and preliminary evaluations of coding agents showing high performance.
Clean English speech is commoditized, but in 2026, ASR faces challenges such as overlapping speech recognition, real-time transcription, dialect and terminology handling, and optimizing the speech tower to avoid burdening Speech LLM.
Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet, using fitted Gabor kernels to replace encoder filters, outperforming the base model on most benchmarks.
This paper introduces the first benchmark for automatic lyric transcription in Greek songs, demonstrating that adapting Whisper through multitask learning and two-stage training achieves a 27.2% Word Error Rate, significantly improving over zero-shot baselines.
Confucius4-R2T2 is a low-latency, high-accuracy real-time speech recognition model developed by NetEase Youdao, featuring configurable chunking and stable output for applications like live captioning.
This paper introduces DasanCallDial, a large-scale Korean benchmark dataset for dialogue-level ASR error correction, and proposes the DCSC framework, achieving state-of-the-art performance in text-only post-editing.
This paper proposes Dual-Form ASR, a framework that integrates spoken-form automatic speech recognition with semantics-aware written-form inverse text normalization using paired supervision and a sequence-level objective, improving performance on Chinese speech recognition tasks.
Microsoft has released VibeVoice-ASR-Streaming, a unified streaming ASR model that transcribes who said what with support for customized hotwords and 10 languages.