Tag
The paper proposes SAMA-ASR, a multimodal adapter that improves automatic speech recognition for low-resource languages by using semantic anchors from translations and acoustic anchors from speech, with experiments on Taiwanese Hokkien and Hakka showing effectiveness over baselines.
VoiceCodeBench is a new benchmark for evaluating exact structured-token recovery in automatic speech recognition, showing that traditional WER metrics are insufficient for production voice workflows.
This preregistered ablation study tests prompt-level context in a production speech transcription tool and finds no detectable change in side-level word error rate, contradicting earlier reports of gains from prompt conditioning.
This paper introduces PromptKWS, a novel prompt-guided open-vocabulary keyword spotting framework that uses prompt embeddings and cross-attention to improve accuracy, achieving over 10% improvement in wakeup rate and over 15% in accuracy compared to baseline systems.
Sunbird AI announces the expansion of Sunflower to 67 African languages with real-time speech and offline access, hosting a live webinar on September 24, 2026.
The Open ASR Leaderboard is a benchmarking tool from Hugging Face for evaluating Automatic Speech Recognition models, accompanied by a technical blog post and linked to the Voice Arena Leaderboard.
FrankenWhisper, an open-source iOS app with speaker ID and noise reduction, has been approved by Apple and is available for free.
Timnit Gebru criticizes the dominant AI paradigm of building monolithic 'machine gods' and advocates for smaller, specialized, community-owned AI tools that are efficient, ethical, and context-specific.
Google DeepMind announces improvements to their AI model, enhancing its ability to understand complex phone numbers, postal codes, and order IDs in noisy environments, along with features like removing filler words and recognizing custom vocabulary.
Google releases Gemini 3.5 Transcribe, an AI model that improves audio transcription by automatically removing filler words, supporting over 85 languages, and offering custom vocabularies and speaker attribution.
IBM releases two new compact AI models, Granite Speech 5.0 Turbo CTC, for extremely fast and accurate English speech transcription, achieving over 12,600 RTFx on NVIDIA H200 GPU.
small-int8 and paraformer-zh are lightweight AI models, each around 200MB, ideal for daily interactions with Agents, offering fast speed and timely responses.
Loqua is a product that enables users to speak naturally to transform ideas into writing, understand screen content, and automate workflows through voice commands, enhancing productivity by reducing typing and context switching.
This article discusses research on measuring benchmark optimization in speech recognition, where some ASR models may optimize for test benchmarks rather than real-world performance, and introduces tests to quantify this phenomenon.
This paper quantifies how high-performing ASR models optimize for benchmarks in ways that inflate scores without improving real-world transcription, using behavioral probes to reveal benchmark-conditioned behaviors.
The article explores the ethical balance between personal voice data ownership and collective benefit in developing AI language technology for speech recognition.
Wispr, an AI dictation startup, raised $280 million in Series B funding at a $2 billion valuation and launched a new model called Canto to improve speech understanding accuracy.
This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.
Talkify is a free, open-source macOS dictation app built on Apple's SpeechAnalyzer for instant, on-device voice transcription with low latency and multi-language support, living in the notch.
This paper presents a controlled benchmark comparing six multilingual pre-trained ASR models on Nepali speech, finding Whisper-Large-v3-Turbo and IndicWav2Vec perform best, while CTC decoders offer up to 29x faster inference. It provides the first standardized efficiency-aware reference numbers for Nepali ASR.