speech-recognition

Tag

Cards List
#speech-recognition

Leading-Silence Augmentation and Multi-Stage Synthetic Supervision for the Second MLC-SLM Challenge

arXiv cs.CL ↗ · 2026-08-17 Cached

This paper introduces techniques for the second MLC-SLM Challenge, including random leading-silence cropping and synthetic data generation to enhance multilingual conversational speech tasks, achieving improved accuracy and reduced error rates.

0 favorites 0 likes
#speech-recognition

@tornikegomareli: I wanted voice dictation on macOS to feel instant and native, and I built Talkify an 8.2 MB macOS dictation app with 12…

X AI KOLs Timeline ↗ · 2026-08-15 Cached

Talkify is a free, open-source macOS dictation app built on Apple's SpeechAnalyzer for instant, on-device voice transcription with low latency and multi-language support, living in the notch.

0 favorites 0 likes
#speech-recognition

Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition

arXiv cs.CL ↗ · 2026-08-14 Cached

This paper presents a controlled benchmark comparing six multilingual pre-trained ASR models on Nepali speech, finding Whisper-Large-v3-Turbo and IndicWav2Vec perform best, while CTC decoders offer up to 29x faster inference. It provides the first standardized efficiency-aware reference numbers for Nepali ASR.

0 favorites 0 likes
#speech-recognition

I fit a complete offline voice agent into 1.2 GB on Android

Reddit r/ArtificialInteligence ↗ · 2026-08-13

A developer demonstrates a complete offline voice-to-action agent on Android using small local models (Silero VAD, Parakeet-EOU STT, FunctionGemma 270M, Pocket TTS), running in ~1.2 GB with no cloud dependencies.

0 favorites 0 likes
#speech-recognition

Easper: An Accessible ASR Pipeline for Language Documentation

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper presents Easper, an open-source no-code ASR pipeline that lets field linguists fine-tune models like Whisper from ELAN annotations, and evaluates data selection strategies for bootstrapping ASR on low-resource Vanuatu languages.

0 favorites 0 likes
#speech-recognition

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces DonorRank, a learning-to-rank framework for selecting effective donor languages in low-resource cross-lingual speech recognition, evaluated on Indic and African language corpora. It demonstrates improved donor selection over genetic-similarity and high-resource heuristics, and provides insights into transfer patterns for multilingual ASR.

0 favorites 0 likes
#speech-recognition

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper argues that single-run evaluations in low-resource ASR are unreliable and demonstrates with a new multi-seed Garhwali ASR benchmark that many reported gains vanish under seed-level testing, while standard CTC with w2v-BERT 2.0 remains the most robust approach.

0 favorites 0 likes
#speech-recognition

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.

0 favorites 0 likes
#speech-recognition

Show HN: Vocal Slice – Cut audio by selecting text, fully on-device

Hacker News Top ↗ · 2026-08-10 Cached

Vocal Slice is a desktop application for audio editing that uses on-device Whisper transcription to allow users to select and export clips by highlighting text in the transcript, designed for podcasters and voice professionals.

0 favorites 0 likes
#speech-recognition

VoiceGecko

Product Hunt ↗ · 2026-08-09

VoiceGecko is an open-source, local voice-to-text tool being showcased on Product Hunt.

0 favorites 0 likes
#speech-recognition

parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM

Reddit r/LocalLLaMA ↗ · 2026-08-07

parakeet.wgsl enables fast, accurate NVIDIA Parakeet TDT 0.6B V2 speech transcription entirely in the browser using raw WebGPU compute shaders and SIMD WASM, with a live demo and open-source library.

0 favorites 0 likes
#speech-recognition

🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp

Reddit r/LocalLLaMA ↗ · 2026-08-06

NVIDIA's entire speech stack—ASR, TTS, and codec—is now quantized to GGUF and runs locally on-device via NeMo-Speech.cpp, with new model releases for Magpie-TTS, Nemotron Speech Streaming, and Parakeet.

0 favorites 0 likes
#speech-recognition

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper introduces MERaLiON-GR, a speech gender recognition model for English and Southeast Asian languages, fine-tuned from MERaLiON-SpeechEncoder-2 with LoRA and an ECAPA-TDNN head, achieving state-of-the-art performance across multilingual benchmarks.

0 favorites 0 likes
#speech-recognition

Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P]

Reddit r/MachineLearning ↗ · 2026-08-05

LiveTranscriber is an open-source iOS app that runs Whisper, Qwen3-ASR, Nemotron, MOSS, and Qwen3 fully offline on iPhone, offering speech transcription, multi-speaker support, summaries, and real-time translation. The developer shares the engineering challenges and invites feedback from ASR and on-device AI communities.

0 favorites 0 likes
#speech-recognition

@MSFTResearch: Small language models learn to negotiate with SocialRL, PazaBench V2 expands speech AI evaluation across African langua…

X AI KOLs Following ↗ · 2026-08-03 Cached

Microsoft Research highlights new research on SocialRL for small language model negotiation, PazaBench V2 for African language speech evaluation, EvoLib for agent experience learning, improved A/B testing methods, and AI-driven precision oncology.

0 favorites 0 likes
#speech-recognition

Zen Whisper

Product Hunt ↗ · 2026-07-31

Zen Whisper is a Mac app offering on-device dictation that can type into any application.

0 favorites 0 likes
#speech-recognition

@SamuelZengML: Open-source ASR just took the top spot. Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a …

X AI KOLs Timeline ↗ · 2026-07-30 Cached

Audio8's open-source ARK-ASR-3B claims the #1 spot on the Hugging Face Open ASR Leaderboard with a 4.76 Mean WER, while their 0.6B model also ranks in the Top 5, showcasing a competitive open speech stack.

0 favorites 0 likes
#speech-recognition

Voice Memory for Agentic Speech Recognition

arXiv cs.CL ↗ · 2026-07-30 Cached

Voice Memory introduces an inference-only scheme for agentic speech recognition where a frozen corrector uses a per-domain memory file to decide per utterance whether to act or abstain, reducing word error rate across multiple domains without weight updates.

0 favorites 0 likes
#speech-recognition

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper evaluates forced alignment for Hindi-English code-mixed speech using the Montreal Forced Aligner, demonstrating that bootstrapping strategies and code-mixed training data achieve a tenfold improvement in alignment accuracy over monolingual alternatives.

0 favorites 0 likes
#speech-recognition

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

arXiv cs.CL ↗ · 2026-07-28 Cached

MoLGE assigns dedicated expert modules to clusters of similar languages in a mixture-of-experts framework for large-scale multilingual ASR, achieving improvements across 495 languages with minimal parameter increase.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback