asr

Tag

Cards List
#asr

parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM

Reddit r/LocalLLaMA · 2d ago

parakeet.wgsl enables fast, accurate NVIDIA Parakeet TDT 0.6B V2 speech transcription entirely in the browser using raw WebGPU compute shaders and SIMD WASM, with a live demo and open-source library.

0 favorites 0 likes
#asr

How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs

arXiv cs.CL · 2d ago Cached

This paper compares context biasing methods and speech LLMs for recognizing rare and new words in automatic speech recognition, reporting trade-offs in word error rate across read and non-read speech.

0 favorites 0 likes
#asr

What STT API are you using for production voice agents, and what broke first?

Reddit r/AI_Agents · 4d ago

A developer asks what STT APIs people use in production voice agents, comparing Deepgram, AssemblyAI, and Smallest AI Pulse, and highlighting common failure points like endpointing, latency, and barge-in.

0 favorites 0 likes
#asr

@Azure: Now in Microsoft Foundry: GPT-transcribe and GPT-live-transcribe. Build with high-accuracy transcription for recorded a…

X AI KOLs Timeline · 2026-07-30 Cached

OpenAI and Microsoft released GPT-transcribe and GPT-live-transcribe in Microsoft Foundry, offering high-accuracy asynchronous transcription and low-latency streaming transcription for recorded and live audio, with features like background noise handling, accent robustness, and alphanumeric perception.

0 favorites 0 likes
#asr

@SamuelZengML: Open-source ASR just took the top spot. Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a …

X AI KOLs Timeline · 2026-07-30 Cached

Audio8's open-source ARK-ASR-3B claims the #1 spot on the Hugging Face Open ASR Leaderboard with a 4.76 Mean WER, while their 0.6B model also ranks in the Top 5, showcasing a competitive open speech stack.

0 favorites 0 likes
#asr

Voice Memory for Agentic Speech Recognition

arXiv cs.CL · 2026-07-30 Cached

Voice Memory introduces an inference-only scheme for agentic speech recognition where a frozen corrector uses a per-domain memory file to decide per utterance whether to act or abstain, reducing word error rate across multiple domains without weight updates.

0 favorites 0 likes
#asr

SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

arXiv cs.CL · 2026-07-29 Cached

This paper presents the first systematic study of federated training for SpeechLLM-based end-to-end ASR systems, evaluating on English and Italian tasks and achieving competitive word error rates with reduced communication costs.

0 favorites 0 likes
#asr

microsoft/VibeVoice-ASR-BitNet

Reddit r/LocalLLaMA · 2026-07-28 Cached

Microsoft releases VibeVoice-ASR-BitNet, a compressed multilingual ASR model for real-time CPU inference. It achieves 1.6-2.3x faster inference than Whisper.cpp with real-time capability on as few as 3 CPU threads.

0 favorites 0 likes
#asr

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

arXiv cs.CL · 2026-07-28 Cached

MoLGE assigns dedicated expert modules to clusters of similar languages in a mixture-of-experts framework for large-scale multilingual ASR, achieving improvements across 495 languages with minimal parameter increase.

0 favorites 0 likes
#asr

Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages

arXiv cs.CL · 2026-07-28 Cached

Indic DiarBench is a multilingual joint diarization and ASR benchmark covering all 22 scheduled languages of India with 108 hours of human-corrected multi-speaker audio, capturing conversational nuances like code-mixing and speaker overlap.

0 favorites 0 likes
#asr

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

arXiv cs.CL · 2026-07-27 Cached

MEUSLI is an open-source multilingual projector family that connects Whisper encoder with multilingual LLMs, enabling end-to-end ASR in 28 European languages and extending to speech translation and topic identification.

0 favorites 0 likes
#asr

@DanKornas: Speech recognition work gets messy when training code, pretrained models, and deployment runtimes live in separate plac…

X AI KOLs Timeline · 2026-07-24 Cached

WeNet is an open-source end-to-end speech recognition toolkit that unifies training code, pretrained models, and deployment runtimes in one project, supporting multiple model architectures and runtime backends.

0 favorites 0 likes
#asr

Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results

arXiv cs.CL · 2026-07-22 Cached

This paper compares state-of-the-art ASR systems to human listeners on recognizing diverse Dutch speech, finding that ASR systems match or exceed human performance in some cases, with Google Telephony leading. It highlights the impact of speaker age, regional accents, and test set selection on benchmarking conclusions.

0 favorites 0 likes
#asr

Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing

arXiv cs.CL · 2026-07-22 Cached

This paper addresses the uncontrolled latent variable of transcription style in ASR models by introducing a method using coverage-aware decoder task tokens to enable controllable verbatim transcription with accurate word-level timing, achieving high disfluency detection F1 via zero-shot cross-lingual transfer.

0 favorites 0 likes
#asr

@MSFTResearch: We’ve significantly expanded coverage to support: ● 22 additional languages, now reaching communities across 38 African…

X AI KOLs Following · 2026-07-21 Cached

Microsoft Research announced expanded coverage for their ASR model, adding 22 languages across 38 African countries, with new datasets and test samples.

0 favorites 0 likes
#asr

@MSFTResearch: Today, we announce the second release of PazaBench, our benchmark for evaluating Automatic Speech Recognition (ASR) mod…

X AI KOLs Timeline · 2026-07-21 Cached

Microsoft Research announces the second release of PazaBench, a benchmark for evaluating automatic speech recognition models across African languages.

0 favorites 0 likes
#asr

Running a 13M ASR conformer on a microcontroller

Reddit r/LocalLLaMA · 2026-07-20

Details a method to run a 13 million parameter ASR Conformer model directly on a microcontroller, highlighting advances in edge AI deployment.

0 favorites 0 likes
#asr

@YichiZ03: https://x.com/YichiZ03/status/2078588932191895976

X AI KOLs Timeline · 2026-07-18 Cached

MOSS-TD, a speaker-aware ASR system, is optimized within the SGLang-Omni serving stack, enabling 38-minute multi-speaker audio to be transcribed in about 49 seconds on a single H100, with concurrent processing of 16 meetings.

0 favorites 0 likes
#asr

@Open_MOSS: MOSS-Transcribe-Diarize is among the top trending models on @huggingface . A few weeks after release, the model has bee…

X AI KOLs Timeline · 2026-07-16 Cached

MOSS-Transcribe-Diarize, an open-source ASR model with multi-speaker diarization and hotword biasing, trends on Hugging Face after release.

0 favorites 0 likes
#asr

@thesupermanmx: NVIDIA open-sourced a 600M model that transcribes 40 languages in real-time at 80ms latency and it costs $0. that's fas…

X AI KOLs Timeline · 2026-07-16 Cached

NVIDIA open-sourced a 600M parameter model that transcribes 40 languages in real-time with 80ms latency, supporting multiple languages from a single checkpoint with built-in punctuation and capitalization.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback