speech-recognition

Tag

Cards List
#speech-recognition

I fit a complete offline voice agent into 1.2 GB on Android

Reddit r/ArtificialInteligence ↗ · 2026-08-13

A developer demonstrates a complete offline voice-to-action agent on Android using small local models (Silero VAD, Parakeet-EOU STT, FunctionGemma 270M, Pocket TTS), running in ~1.2 GB with no cloud dependencies.

0 favorites 0 likes
#speech-recognition

Easper: An Accessible ASR Pipeline for Language Documentation

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper presents Easper, an open-source no-code ASR pipeline that lets field linguists fine-tune models like Whisper from ELAN annotations, and evaluates data selection strategies for bootstrapping ASR on low-resource Vanuatu languages.

0 favorites 0 likes
#speech-recognition

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

arXiv cs.CL ↗ · 2026-08-13 Cached

This paper introduces DonorRank, a learning-to-rank framework for selecting effective donor languages in low-resource cross-lingual speech recognition, evaluated on Indic and African language corpora. It demonstrates improved donor selection over genetic-similarity and high-resource heuristics, and provides insights into transfer patterns for multilingual ASR.

0 favorites 0 likes
#speech-recognition

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper argues that single-run evaluations in low-resource ASR are unreliable and demonstrates with a new multi-seed Garhwali ASR benchmark that many reported gains vanish under seed-level testing, while standard CTC with w2v-BERT 2.0 remains the most robust approach.

0 favorites 0 likes
#speech-recognition

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper analyzes multimodal systems for the CHiME-9 MCoRec cocktail-party scenario, comparing design strategies such as audio-visual target speech separation, improved recognition, and LLM-based conversational grouping, finding that speech overlap alone does not explain performance differences.

0 favorites 0 likes
#speech-recognition

Show HN: Vocal Slice – Cut audio by selecting text, fully on-device

Hacker News Top ↗ · 2026-08-10 Cached

Vocal Slice is a desktop application for audio editing that uses on-device Whisper transcription to allow users to select and export clips by highlighting text in the transcript, designed for podcasters and voice professionals.

0 favorites 0 likes
#speech-recognition

VoiceGecko

Product Hunt ↗ · 2026-08-09

VoiceGecko is an open-source, local voice-to-text tool being showcased on Product Hunt.

0 favorites 0 likes
#speech-recognition

parakeet.wgsl – Fast, accurate ASR in the browser, via raw WebGPU & SIMD WASM

Reddit r/LocalLLaMA ↗ · 2026-08-07

parakeet.wgsl enables fast, accurate NVIDIA Parakeet TDT 0.6B V2 speech transcription entirely in the browser using raw WebGPU compute shaders and SIMD WASM, with a live demo and open-source library.

0 favorites 0 likes
#speech-recognition

🟩 NVIDIA's whole speech stack just went local. ASR + TTS + codec, quantized to GGUF, running on-device via NeMo-Speech.cpp

Reddit r/LocalLLaMA ↗ · 2026-08-06

NVIDIA's entire speech stack—ASR, TTS, and codec—is now quantized to GGUF and runs locally on-device via NeMo-Speech.cpp, with new model releases for Magpie-TTS, Nemotron Speech Streaming, and Parakeet.

0 favorites 0 likes
#speech-recognition

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper introduces MERaLiON-GR, a speech gender recognition model for English and Southeast Asian languages, fine-tuned from MERaLiON-SpeechEncoder-2 with LoRA and an ECAPA-TDNN head, achieving state-of-the-art performance across multilingual benchmarks.

0 favorites 0 likes
#speech-recognition

Running Whisper, Qwen3-ASR, Nemotron & MOSS completely offline on iPhone [P]

Reddit r/MachineLearning ↗ · 2026-08-05

LiveTranscriber is an open-source iOS app that runs Whisper, Qwen3-ASR, Nemotron, MOSS, and Qwen3 fully offline on iPhone, offering speech transcription, multi-speaker support, summaries, and real-time translation. The developer shares the engineering challenges and invites feedback from ASR and on-device AI communities.

0 favorites 0 likes
#speech-recognition

@MSFTResearch: Small language models learn to negotiate with SocialRL, PazaBench V2 expands speech AI evaluation across African langua…

X AI KOLs Following ↗ · 2026-08-03 Cached

Microsoft Research highlights new research on SocialRL for small language model negotiation, PazaBench V2 for African language speech evaluation, EvoLib for agent experience learning, improved A/B testing methods, and AI-driven precision oncology.

0 favorites 0 likes
#speech-recognition

Zen Whisper

Product Hunt ↗ · 2026-07-31

Zen Whisper is a Mac app offering on-device dictation that can type into any application.

0 favorites 0 likes
#speech-recognition

@SamuelZengML: Open-source ASR just took the top spot. Audio8 ARK-ASR-3B now ranks #1 on the @huggingface Open ASR Leaderboard with a …

X AI KOLs Timeline ↗ · 2026-07-30 Cached

Audio8's open-source ARK-ASR-3B claims the #1 spot on the Hugging Face Open ASR Leaderboard with a 4.76 Mean WER, while their 0.6B model also ranks in the Top 5, showcasing a competitive open speech stack.

0 favorites 0 likes
#speech-recognition

Voice Memory for Agentic Speech Recognition

arXiv cs.CL ↗ · 2026-07-30 Cached

Voice Memory introduces an inference-only scheme for agentic speech recognition where a frozen corrector uses a per-domain memory file to decide per utterance whether to act or abstain, reducing word error rate across multiple domains without weight updates.

0 favorites 0 likes
#speech-recognition

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

arXiv cs.CL ↗ · 2026-07-29 Cached

This paper evaluates forced alignment for Hindi-English code-mixed speech using the Montreal Forced Aligner, demonstrating that bootstrapping strategies and code-mixed training data achieve a tenfold improvement in alignment accuracy over monolingual alternatives.

0 favorites 0 likes
#speech-recognition

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition

arXiv cs.CL ↗ · 2026-07-28 Cached

MoLGE assigns dedicated expert modules to clusters of similar languages in a mixture-of-experts framework for large-scale multilingual ASR, achieving improvements across 495 languages with minimal parameter increase.

0 favorites 0 likes
#speech-recognition

Show HN: Yap – OSS on-device voice dictation for macOS with no model to download

Hacker News Top ↗ · 2026-07-27 Cached

Yap is an open-source macOS app for blazing-fast on-device voice dictation using Apple's Speech framework, requiring no model download, API key, or internet connection.

0 favorites 0 likes
#speech-recognition

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

arXiv cs.CL ↗ · 2026-07-27 Cached

MEUSLI is an open-source multilingual projector family that connects Whisper encoder with multilingual LLMs, enabling end-to-end ASR in 28 European languages and extending to speech translation and topic identification.

0 favorites 0 likes
#speech-recognition

Whisper Live - A nearly-live implementation of Open AI's Whisper, free & open-source

Reddit r/artificial ↗ · 2026-07-25 Cached

WhisperLive is an open-source real-time transcription tool using OpenAI's Whisper, supporting multiple backends like faster-whisper and TensorRT for live speech-to-text.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback