Tag
Shinji Watanabe announces his Google Scholar h-index has reached 100, crediting decades of work on speech processing research and open-source tooling to his students and collaborators. The post highlights landmark works like ESPnet, Deep Clustering, and SUPERB that underpin modern speech recognition.
The team behind .wave introduces an inference engine using a Wave Persistent Kernel compiler to serve NVIDIA Nemotron 3.5 ASR Streaming at $0.00045/minute, sustaining 4,800 concurrent streams on one H100 with p99 frame completion under 106.3 ms — roughly 20x NVIDIA's published stream capacity and a fraction of competing ASR pricing.
FermionResearch released Phonon-2, the most accurate open English speech recognition model under 900 MB, averaging 5.21% word error on the Open ASR Leaderboard while matching its 2.5 GB teacher through ~2.1-bit quantization.
The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.
The paper introduces target-speaker unlearning ASR (TSU-ASR) and proposes a novel Enrollment-Conditioned Gating module for dynamic opt-out of speakers during inference in LLM-based ASR, enhancing privacy in online conferencing.
This paper systematically evaluates pseudo-labeling for adapting pretrained ASR models like Whisper and Qwen3-ASR to noisy police audio, introducing an LLM-as-a-judge filtering method and a cross-model paradigm to reduce word error rates.
This paper introduces Pruned CTC, a method that reduces memory usage in Connectionist Temporal Classification training for large-vocabulary automatic speech recognition by restricting alignment computation, proving equivalence to full-vocabulary CTC and demonstrating strong performance with LLM adaptations in offline and streaming settings.
BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.
PTC-Bias is a two-stage framework for speech large language models that uses phoneme-level temporal competition to improve rare-word recognition by efficiently retrieving and correcting bias words, with experiments showing significant gains on LibriSpeech.
Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.
NetEase Youdao has released two AI models: a streaming ASR model with 2B parameters and a Chinese-English simultaneous translation model with 14B parameters, both emphasizing real-time performance.
This paper proposes a practical recipe for semi-supervised federated ASR using online pseudo-labels with server update stabilization, demonstrating significant improvements over prior methods in both in-domain and cross-domain settings.
The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.
The Bairong system presents a dynamic question-aware evidence routing approach for multilingual conversational spoken question answering, achieving improved performance by effectively integrating transcript and audio evidence in the MLC-SLM 2026 Challenge.
The article highlights how aggregate Word Error Rate (WER) metric fails to capture critical errors in voice agents for production, such as misrecognized bank codes and multilingual speech, necessitating per-field tracking for real-world applications like banking and collections.
R2T2 is a low-latency and high-accuracy real-time speech recognition model that processes audio in small chunks and commits text without revision, suitable for applications like live captioning and translation.
The article reflects on the need for stability over speed in ASR for voice agents, discussing the 'Confucius r2t2' model that integrates wait/commit mechanisms to enhance reliability.
The paper proposes T-SANDHI, a Tone Sandhi-aware Adaptive Network for low-resource Taiwanese Hokkien speech recognition, which explicitly decouples tonal variations to improve accuracy on top of a Whisper backbone.
NetEase Youdao open-sourced Confucius4-R2T2, a streaming ASR model for voice agents that incrementally processes speech and emits only committed text to prevent state corruption.
Nari Labs introduces Qwen3-TTS and Qwen3-ASR models, providing high accuracy, low latency, and cost-effectiveness in a free public beta, alongside optimized APIs and services for production deployment.