speech-recognition

Tag

Cards List
#speech-recognition

Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

arXiv cs.CL ↗ · 2026-09-01 Cached

The paper proposes SAMA-ASR, a multimodal adapter that improves automatic speech recognition for low-resource languages by using semantic anchors from translations and acoustic anchors from speech, with experiments on Taiwanese Hokkien and Hakka showing effectiveness over baselines.

0 favorites 0 likes
#speech-recognition

VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition

arXiv cs.CL ↗ · 2026-09-01 Cached

VoiceCodeBench is a new benchmark for evaluating exact structured-token recovery in automatic speech recognition, showing that traditional WER metrics are insufficient for production voice workflows.

0 favorites 0 likes
#speech-recognition

No Detectable Change in Side-Level WER from Prompt-Level Context: A Preregistered Ablation on a Production Oral-History Corpus

arXiv cs.CL ↗ · 2026-09-01 Cached

This preregistered ablation study tests prompt-level context in a production speech transcription tool and finds no detectable change in side-level word error rate, contradicting earlier reports of gains from prompt conditioning.

0 favorites 0 likes
#speech-recognition

PromptKWS: A Novel Prompt-Guided Open-Vocabulary Keyword Spotting Framework

arXiv cs.CL ↗ · 2026-09-01 Cached

This paper introduces PromptKWS, a novel prompt-guided open-vocabulary keyword spotting framework that uses prompt embeddings and cross-attention to improve accuracy, achieving over 10% improvement in wakeup rate and over 15% in accuracy compared to baseline systems.

0 favorites 0 likes
#speech-recognition

@SunbirdAI: Most AI systems serve a small fraction of the world's languages. Across Africa, that gap has meant limited access to te…

X AI KOLs Following ↗ · 2026-08-31 Cached

Sunbird AI announces the expansion of Sunflower to 67 African languages with real-time speech and offline access, hosting a live webinar on September 24, 2026.

0 favorites 0 likes
#speech-recognition

@rohanpaul_ai: Open ASR Leaderboard - https://huggingface.co/spaces/hf-audio/open_asr_leaderboard… Technical Release Blog - https://hu…

X AI KOLs Timeline ↗ · 2026-08-28 Cached

The Open ASR Leaderboard is a benchmarking tool from Hugging Face for evaluating Automatic Speech Recognition models, accompanied by a technical blog post and linked to the Voice Arena Leaderboard.

0 favorites 0 likes
#speech-recognition

@doodlestein: The FrankenWhisper app was finally approved by Apple! Get it here, it’s totally free (zero ads, zero in-app purchases!)…

X AI KOLs Timeline ↗ · 2026-08-27 Cached

FrankenWhisper, an open-source iOS app with speaker ID and noise reduction, has been approved by Apple and is available for free.

0 favorites 0 likes
#speech-recognition

They're trying to build a machine god | Timnit Gebru

Reddit r/ArtificialInteligence ↗ · 2026-08-27 Cached

Timnit Gebru criticizes the dominant AI paradigm of building monolithic 'machine gods' and advocates for smaller, specialized, community-owned AI tools that are efficient, ethical, and context-specific.

0 favorites 0 likes
#speech-recognition

@GoogleDeepMind: Here’s what’s new: It’s better at understanding complex phone numbers, postal codes, and order IDs – even in noisy envi…

X AI KOLs ↗ · 2026-08-26 Cached

Google DeepMind announces improvements to their AI model, enhancing its ability to understand complex phone numbers, postal codes, and order IDs in noisy environments, along with features like removing filler words and recognizing custom vocabulary.

0 favorites 0 likes
#speech-recognition

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

The Verge ↗ · 2026-08-26 Cached

Google releases Gemini 3.5 Transcribe, an AI model that improves audio transcription by automatically removing filler words, supporting over 85 languages, and offering custom vocabularies and speaker attribution.

0 favorites 0 likes
#speech-recognition

Granite Speech 5.0 Turbo CTC: Extremely Fast and Accurate Transcription

Reddit r/LocalLLaMA ↗ · 2026-08-25 Cached

IBM releases two new compact AI models, Granite Speech 5.0 Turbo CTC, for extremely fast and accurate English speech transcription, achieving over 12,600 RTFx on NVIDIA H200 GPU.

0 favorites 0 likes
#speech-recognition

@iamcheyan: Both small-int8 and paraformer-zh are only about 200MB. For daily use, paraformer-zh is perfectly sufficient for communicating with Agents. The speed is also very fast. For just one or two sentences, it's basically press the button to speak, release and get the result instantly.

X AI KOLs Following ↗ · 2026-08-22 Cached

small-int8 and paraformer-zh are lightweight AI models, each around 200MB, ideal for daily interactions with Agents, offering fast speed and timely responses.

0 favorites 0 likes
#speech-recognition

Loqua

Product Hunt ↗ · 2026-08-21 Cached

Loqua is a product that enables users to speak naturally to transform ideas into writing, understand screen content, and automate workflows through voice commands, enhancing productivity by reducing typing and context switching.

0 favorites 0 likes
#speech-recognition

Measuring benchmark optimization in speech recognition

Hugging Face Blog ↗ · 2026-08-21 Cached

This article discusses research on measuring benchmark optimization in speech recognition, where some ASR models may optimize for test benchmarks rather than real-world performance, and introduces tests to quantify this phenomenon.

0 favorites 0 likes
#speech-recognition

Towards Quantifying Benchmark Optimization in ASR Models

Hugging Face Daily Papers ↗ · 2026-08-20 Cached

This paper quantifies how high-performing ASR models optimize for benchmarks in ways that inflate scores without improving real-world transcription, using behavioral probes to reveal benchmark-conditioned behaviors.

0 favorites 0 likes
#speech-recognition

Who owns your voice?

Reddit r/ArtificialInteligence ↗ · 2026-08-19 Cached

The article explores the ethical balance between personal voice data ownership and collective benefit in developing AI language technology for speech recognition.

0 favorites 0 likes
#speech-recognition

Wispr raises $280M at $2B valuation as it looks beyond dictation

TechCrunch AI ↗ · 2026-08-17 Cached

Wispr, an AI dictation startup, raised $280 million in Series B funding at a $2 billion valuation and launched a new model called Canto to improve speech understanding accuracy.

0 favorites 0 likes
#speech-recognition

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.

0 favorites 0 likes
#speech-recognition

@tornikegomareli: I wanted voice dictation on macOS to feel instant and native, and I built Talkify an 8.2 MB macOS dictation app with 12…

X AI KOLs Timeline ↗ · 2026-08-15 Cached

Talkify is a free, open-source macOS dictation app built on Apple's SpeechAnalyzer for instant, on-device voice transcription with low latency and multi-language support, living in the notch.

0 favorites 0 likes
#speech-recognition

Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper presents a controlled benchmark comparing six multilingual pre-trained ASR models on Nepali speech, finding Whisper-Large-v3-Turbo and IndicWav2Vec perform best, while CTC decoders offer up to 29x faster inference. It provides the first standardized efficiency-aware reference numbers for Nepali ASR.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback