word-error-rate

Tag

Cards List
#word-error-rate

@svpino: Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement. Voice isolation is the …

X AI KOLs Timeline ↗ · yesterday Cached

Krisp released an open benchmark and dataset showing that voice isolation reduces word error rates in speech-to-text models by 73%, with significant improvements across workplace and call-center recordings.

0 favorites 0 likes
#word-error-rate

Canto: A speech model built for the real world

Hacker News Top ↗ · 2026-09-17 Cached

Wispr Flow introduces Canto, a speech model for real-time dictation that excels in noisy and challenging real-world conditions, achieving the lowest word error rate in evaluations against competitors.

0 favorites 0 likes
#word-error-rate

@rohanpaul_ai: Love this, another huge release from Meta. Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest…

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Meta has released Muse Voice Transcribe, a real-time speech-to-text model with a 3.1% word error rate and adaptive streaming capabilities, making it suitable for voice agents.

0 favorites 0 likes
#word-error-rate

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

arXiv cs.CL ↗ · 2026-09-01 Cached

This preregistered ablation study tests prompt-level context in a production speech transcription tool and finds no detectable change in side-level word error rate, contradicting earlier reports of gains from prompt conditioning.

0 favorites 0 likes
#word-error-rate

Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper presents a controlled benchmark comparing six multilingual pre-trained ASR models on Nepali speech, finding Whisper-Large-v3-Turbo and IndicWav2Vec perform best, while CTC decoders offer up to 29x faster inference. It provides the first standardized efficiency-aware reference numbers for Nepali ASR.

0 favorites 0 likes
#word-error-rate

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

arXiv cs.CL ↗ · 2026-08-07 Cached

This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.

0 favorites 0 likes
#word-error-rate

Room reverberation and low SNR hurt STT accuracy far more than model size

Reddit r/ArtificialInteligence ↗ · 2026-07-30

Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.

0 favorites 0 likes
#word-error-rate

Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor

Hacker News Top ↗ · 2026-07-13 Cached

Apple's new SpeechAnalyzer API significantly outperforms both its predecessor SFSpeechRecognizer and OpenAI's Whisper models in accuracy and speed for English on-device transcription, as benchmarked on an M2 Pro machine. The new API achieves a 2.12% word error rate on clean speech, compared to 3.74% for Whisper Small, and runs three times faster.

0 favorites 0 likes
#word-error-rate

Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition

arXiv cs.CL ↗ · 2026-07-08 Cached

This paper revisits the classic relation between language model perplexity and ASR word error rate in the context of modern end-to-end ASR systems, finding that while external LMs still improve WER, the log-log linear relation still holds but is affected by internal language modeling in encoder-decoder models.

0 favorites 0 likes
#word-error-rate

What Was That Again? Certified Robustness for Automatic Speech Recognition

arXiv cs.LG ↗ · 2026-06-29 Cached

This paper presents a certification-inspired mechanism for automatic speech recognition that uses a dual-gate diagnostic pipeline (Two-Sided Atomic Audit and Rank-Based Tournament) to provide certified robustness and achieve up to a 55% relative reduction in word error rate across diverse architectures.

0 favorites 0 likes
#word-error-rate

Voice agents in noisy environments

Reddit r/AI_Agents ↗ · 2026-06-16

A speech company trained a model that cancels noise and identifies the primary speaker, achieving 50% lower word error rate on leading ASR models in noisy environments.

0 favorites 0 likes
#word-error-rate

Beyond Single Ground Truth: Reference Monism as Epistemic Injustice in ASR Evaluation

arXiv cs.CL ↗ · 2026-05-11 Cached

This paper critiques the use of single-reference ground truth in ASR evaluation, arguing it causes epistemic injustice for speakers with aphasia. It proposes a new metric, Epistemic Injustice Distance, and advocates for WER-Range to account for diverse transcription conventions.

0 favorites 0 likes
← Back to home

Submit Feedback