asr

Tag

Cards List
#asr

Easper: An Accessible ASR Pipeline for Language Documentation

arXiv cs.CL · 4d ago Cached

This paper presents Easper, an open-source no-code ASR pipeline that lets field linguists fine-tune models like Whisper from ELAN annotations, and evaluates data selection strategies for bootstrapping ASR on low-resource Vanuatu languages.

0 favorites 0 likes
#asr

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

arXiv cs.CL · 5d ago Cached

This paper argues that single-run evaluations in low-resource ASR are unreliable and demonstrates with a new multi-seed Garhwali ASR benchmark that many reported gains vanish under seed-level testing, while standard CTC with w2v-BERT 2.0 remains the most robust approach.

0 favorites 0 likes
#asr

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

arXiv cs.AI · 5d ago Cached

Presents an age-aware multi-task learning method for phoneme recognition from children's speech, enabling a lightweight 94M-parameter model to outperform larger models and run on edge devices like phones.

0 favorites 0 likes
#asr

parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM

Reddit r/LocalLLaMA · 2026-08-07

parakeet.wgsl enables fast, accurate NVIDIA Parakeet TDT 0.6B V2 speech transcription entirely in the browser using raw WebGPU compute shaders and SIMD WASM, with a live demo and open-source library.

0 favorites 0 likes
#asr

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

arXiv cs.CL · 2026-08-07 Cached

This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.

0 favorites 0 likes
#asr

What STT API are you using for production voice agents, and what broke first?

Reddit r/AI_Agents · 2026-08-05

A developer asks what STT APIs people use in production voice agents, comparing Deepgram, AssemblyAI, and Smallest AI Pulse, and highlighting common failure points like endpointing, latency, and barge-in.

0 favorites 0 likes
#asr

@Azure: Now in Microsoft Foundry: GPT-transcribe and GPT-live-transcribe. Build with high-accuracy transcription for recorded a…

X AI KOLs Timeline · 2026-07-30 Cached

OpenAI and Microsoft released GPT-transcribe and GPT-live-transcribe in Microsoft Foundry, offering high-accuracy asynchronous transcription and low-latency streaming transcription for recorded and live audio, with features like background noise handling, accent robustness, and alphanumeric perception.

0 favorites 0 likes
#asr

@SamuelZengML: Open-source ASR just took the top spot. Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a …

X AI KOLs Timeline · 2026-07-30 Cached

Audio8's open-source ARK-ASR-3B claims the #1 spot on the Hugging Face Open ASR Leaderboard with a 4.76 Mean WER, while their 0.6B model also ranks in the Top 5, showcasing a competitive open speech stack.

0 favorites 0 likes
#asr

Voice Memory for Agentic Speech Recognition

arXiv cs.CL · 2026-07-30 Cached

Voice Memory introduces an inference-only scheme for agentic speech recognition where a frozen corrector uses a per-domain memory file to decide per utterance whether to act or abstain, reducing word error rate across multiple domains without weight updates.

0 favorites 0 likes
#asr

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

arXiv cs.CL · 2026-07-29 Cached

This paper presents the first systematic study of federated training for SpeechLLM-based end-to-end ASR systems, evaluating on English and Italian tasks and achieving competitive word error rates with reduced communication costs.

0 favorites 0 likes
#asr

microsoft/VibeVoice-ASR-BitNet

Reddit r/LocalLLaMA · 2026-07-28 Cached

Microsoft releases VibeVoice-ASR-BitNet, a compressed multilingual ASR model for real-time CPU inference. It achieves 1.6-2.3x faster inference than Whisper.cpp with real-time capability on as few as 3 CPU threads.

0 favorites 0 likes
#asr

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

arXiv cs.CL · 2026-07-28 Cached

MoLGE assigns dedicated expert modules to clusters of similar languages in a mixture-of-experts framework for large-scale multilingual ASR, achieving improvements across 495 languages with minimal parameter increase.

0 favorites 0 likes
#asr

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

arXiv cs.CL · 2026-07-28 Cached

Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.

0 favorites 0 likes
#asr

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

arXiv cs.CL · 2026-07-27 Cached

MEUSLI is an open-source multilingual projector family that connects Whisper encoder with multilingual LLMs, enabling end-to-end ASR in 28 European languages and extending to speech translation and topic identification.

0 favorites 0 likes
#asr

@DanKornas: Speech recognition work gets messy when training code, pretrained models, and deployment runtimes live in separate plac…

X AI KOLs Timeline · 2026-07-24 Cached

WeNet is an open-source end-to-end speech recognition toolkit that unifies training code, pretrained models, and deployment runtimes in one project, supporting multiple model architectures and runtime backends.

0 favorites 0 likes
#asr

Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

arXiv cs.CL · 2026-07-22 Cached

This paper compares state-of-the-art ASR systems to human listeners on recognizing diverse Dutch speech, finding that ASR systems match or exceed human performance in some cases, with Google Telephony leading. It highlights the impact of speaker age, regional accents, and test set selection on benchmarking conclusions.

0 favorites 0 likes
#asr

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

arXiv cs.CL · 2026-07-22 Cached

This paper addresses the uncontrolled latent variable of transcription style in ASR models by introducing a method using coverage-aware decoder task tokens to enable controllable verbatim transcription with accurate word-level timing, achieving high disfluency detection F1 via zero-shot cross-lingual transfer.

0 favorites 0 likes
#asr

@MSFTResearch: We’ve significantly expanded coverage to support: ● 22 additional languages, now reaching communities across 38 African…

X AI KOLs Following · 2026-07-21 Cached

Microsoft Research announced expanded coverage for their ASR model, adding 22 languages across 38 African countries, with new datasets and test samples.

0 favorites 0 likes
#asr

@MSFTResearch: Today, we announce the second release of PazaBench, our benchmark for evaluating Automatic Speech Recognition (ASR) mod…

X AI KOLs Timeline · 2026-07-21 Cached

Microsoft Research announces the second release of PazaBench, a benchmark for evaluating automatic speech recognition models across African languages.

0 favorites 0 likes
#asr

Running a 13M ASR conformer on a microcontroller

Reddit r/LocalLLaMA · 2026-07-20

Details a method to run a 13 million parameter ASR Conformer model directly on a microcontroller, highlighting advances in edge AI deployment.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback