asr

Tag

Cards List
#asr

@shinjiw_at_cmu: My Google Scholar h-index just reached 100! My first paper was in 1999, and behind these 100 are students and collabora…

X AI KOLs Following ↗ · 19h ago Cached

Shinji Watanabe announces his Google Scholar h-index has reached 100, crediting decades of work on speech processing research and open-source tooling to his students and collaborators. The post highlights landmark works like ESPnet, Deep Clustering, and SUPERB that underpin modern speech recognition.

0 favorites 0 likes
#asr

Why real-time ASR is so expensive to serve — and how to fix it

Reddit r/AI_Agents ↗ · 23h ago

The team behind .wave introduces an inference engine using a Wave Persistent Kernel compiler to serve NVIDIA Nemotron 3.5 ASR Streaming at $0.00045/minute, sustaining 4,800 concurrent streams on one H100 with p99 frame completion under 106.3 ms — roughly 20x NVIDIA's published stream capacity and a fraction of competing ASR pricing.

0 favorites 0 likes
#asr

FermionResearch/Phonon-2

Hugging Face Models Trending ↗ · 6d ago Cached

FermionResearch released Phonon-2, the most accurate open English speech recognition model under 900 MB, averaging 5.21% word error on the Open ASR Leaderboard while matching its 2.5 GB teacher through ~2.1-bit quantization.

0 favorites 0 likes
#asr

I-Parakeet: Integer-Only Conformer ASR on Mobile NPU

arXiv cs.CL ↗ · 6d ago Cached

The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.

0 favorites 0 likes
#asr

Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition

arXiv cs.CL ↗ · 6d ago Cached

The paper introduces target-speaker unlearning ASR (TSU-ASR) and proposes a novel Enrollment-Conditioned Gating module for dynamic opt-out of speakers during inference in LLM-based ASR, enhancing privacy in online conferencing.

0 favorites 0 likes
#asr

Pretrained ASR Pseudo-labeling for Noisy Police Audio

arXiv cs.AI ↗ · 6d ago Cached

This paper systematically evaluates pseudo-labeling for adapting pretrained ASR models like Whisper and Qwen3-ASR to noisy police audio, introducing an LLM-as-a-judge filtering method and a cross-model paradigm to reduce word error rates.

0 favorites 0 likes
#asr

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Hugging Face Daily Papers ↗ · 2026-09-27 Cached

This paper introduces Pruned CTC, a method that reduces memory usage in Connectionist Temporal Classification training for large-vocabulary automatic speech recognition by restricting alignment computation, proving equivalence to full-vocabulary CTC and demonstrating strong performance with LLM adaptations in offline and streaming settings.

0 favorites 0 likes
#asr

@omarsar0: Recommended benchmark. I expect voice to become one of the main ways people interact with robots and physical AI. That …

X AI KOLs Timeline ↗ · 2026-09-25 Cached

BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.

0 favorites 0 likes
#asr

PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

arXiv cs.CL ↗ · 2026-09-25 Cached

PTC-Bias is a two-stage framework for speech large language models that uses phoneme-level temporal competition to improve rare-word recognition by efficiently retrieving and correcting bias words, with experiments showing significant gains on LibriSpeech.

0 favorites 0 likes
#asr

@SamuelZengML: Speech recognition is easy—until you ask it to listen forever. Today we’re open-sourcing Audio8 ASR Infinite: Ultra-low…

X AI KOLs Following ↗ · 2026-09-23 Cached

Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.

0 favorites 0 likes
#asr

@AdinaYakup: 2 models from NetEase Youdao are trending today Streaming ASR model: - 2B + custom NetEase license - No output revision…

X AI KOLs Timeline ↗ · 2026-09-23 Cached

NetEase Youdao has released two AI models: a streaming ASR model with 2B parameters and a Chinese-English simultaneous translation model with 14B parameters, both emphasizing real-time performance.

0 favorites 0 likes
#asr

A Practical Recipe for Semi-Supervised Federated ASR: Online Pseudo-Labels with Server Update Stabilization

arXiv cs.LG ↗ · 2026-09-23 Cached

This paper proposes a practical recipe for semi-supervised federated ASR using online pseudo-labels with server update stabilization, demonstrating significant improvements over prior methods in both in-domain and cross-domain settings.

0 favorites 0 likes
#asr

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Hugging Face Daily Papers ↗ · 2026-09-23 Cached

The paper introduces a method to prune encoder layers in the Whisper ASR model, reducing encoder size by 18.5% and recovering performance through unlabeled data distillation, with code and a pre-trained model released for adoption.

0 favorites 0 likes
#asr

The Bairong System for MLC-SLM 2026: Dynamic Question-Aware Evidence Routing for Multilingual Conversational Speech Understanding

arXiv cs.CL ↗ · 2026-09-22 Cached

The Bairong system presents a dynamic question-aware evidence routing approach for multilingual conversational spoken question answering, achieving improved performance by effectively integrating transcript and audio evidence in the MLC-SLM 2026 Challenge.

0 favorites 0 likes
#asr

Aggregate WER is a useless metric for voice agents in production

Reddit r/AI_Agents ↗ · 2026-09-21

The article highlights how aggregate Word Error Rate (WER) metric fails to capture critical errors in voice agents for production, such as misrecognized bank codes and multilingual speech, necessitating per-field tracking for real-world applications like banking and collections.

0 favorites 0 likes
#asr

@svpino: This new model does something really cool: It turns speech into text as you speak. This is different from every other a…

X AI KOLs Timeline ↗ · 2026-09-20 Cached

R2T2 is a low-latency and high-accuracy real-time speech recognition model that processes audio in small chunks and commits text without revision, suitable for applications like live captioning and translation.

0 favorites 0 likes
#asr

Stability vs. speed: Rethinking ASR for the age of voice agents

Reddit r/artificial ↗ · 2026-09-18

The article reflects on the need for stability over speed in ASR for voice agents, discussing the 'Confucius r2t2' model that integrates wait/commit mechanisms to enhance reliability.

0 favorites 0 likes
#asr

T-SANDHI: Tone Sandhi-aware Adaptive Network with Decoupled Hybrid Injection for Low-resource Taiwanese Hokkien Speech Recognition

arXiv cs.CL ↗ · 2026-09-17 Cached

The paper proposes T-SANDHI, a Tone Sandhi-aware Adaptive Network for low-resource Taiwanese Hokkien speech recognition, which explicitly decouples tonal variations to improve accuracy on top of a Whisper backbone.

0 favorites 0 likes
#asr

@rohanpaul_ai: Voice agents should consume speech incrementally but only act on committed text, because a fast transcript that mutates…

X AI KOLs Following ↗ · 2026-09-16 Cached

NetEase Youdao open-sourced Confucius4-R2T2, a streaming ASR model for voice agents that incrementally processes speech and emits only committed text to prevent state corruption.

0 favorites 0 likes
#asr

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

Hacker News Top ↗ · 2026-09-14 Cached

Nari Labs introduces Qwen3-TTS and Qwen3-ASR models, providing high accuracy, low latency, and cost-effectiveness in a free public beta, alongside optimized APIs and services for production deployment.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback