Tag
The developer released LokalBot 0.9.2, a free, open-source Mac meeting notetaker that runs entirely on-device with a 6.8 GB stack of five local models (Qwen3-ASR for transcription, Nemotron 3 for speaker diarization, Qwen3.5 4B for notes, Harrier for embeddings, and LFM2.5 for autocomplete), along with benchmark improvements on an M4 Max.
Oído is an open-source speech recognition system from Lokutor that runs NVIDIA's Conformer-CTC Small (13M params, int8) on a $5 ESP32-S3 microcontroller, outperforming Whisper tiny.en on LibriSpeech and noisy benchmarks with no GPU or NPU.
getcta.store is an AI teleprompter for macOS that positions under the MacBook notch, offering features like script generation, voice-following highlighting, and hidden operation during video calls.
Zumbo is an open-source local AI voice-to-text tool for Mac that provides private speech-to-text conversion with high accuracy, multiple vocabulary packs, and offline functionality.
The paper presents I-Parakeet, an integer-only Conformer ASR model optimized for mobile NPUs, enabling efficient on-device speech recognition.
This paper investigates whether audio large language models can detect unreliable transcriptions and proposes a lightweight predictor using audio encoder representations to trigger clarification requests.
The paper introduces target-speaker unlearning ASR (TSU-ASR) and proposes a novel Enrollment-Conditioned Gating module for dynamic opt-out of speakers during inference in LLM-based ASR, enhancing privacy in online conferencing.
This paper introduces Pruned CTC, a method that reduces memory usage in Connectionist Temporal Classification training for large-vocabulary automatic speech recognition by restricting alignment computation, proving equivalence to full-vocabulary CTC and demonstrating strong performance with LLM adaptations in offline and streaming settings.
BRIDGE ASR 2.0 is a benchmark that evaluates speech recognition models on real two-person conversations in 21 languages, focusing on code-switching and multilingual performance to enhance voice interaction in AI and robotics.
Agentic-GER proposes an LLM-based agent for correcting domain-specific terminology in long-form speech transcripts using global context and selective re-transcription, achieving significant improvements in ASR accuracy for Chinese and English.
This paper investigates how post-training weight compression of Whisper ASR models widens demographic disparities in transcription errors, leading to increased correction time for marginalized speakers.
This survey paper reviews developments in brain-to-language decoding, translating neural activity into linguistic outputs for communication restoration and scientific study, covering tasks, methods, evaluation, and future directions.
Ruby-ASR presents a new supervision method for Japanese automatic speech recognition that binds orthographic spans to their lexical readings, improving reading recognition while maintaining transcription accuracy.
Open-sourcing Audio8 ASR Infinite, a speech recognition tool with ultra-low latency, unlimited audio support, 24/7 transcription, and built-in semantic turn detection, claimed to be new state-of-the-art for streaming ASR.
This paper proposes a practical recipe for semi-supervised federated ASR using online pseudo-labels with server update stabilization, demonstrating significant improvements over prior methods in both in-domain and cross-domain settings.
NetEase Youdao's open-source AI models R2T2 and T3PO have topped Hugging Face leaderboards for speech recognition and translation, outperforming major competitors with impressive real-time performance and stability.
Speechka is a real-time voice translation tool that mimics the user's voice across 44 languages, available on macOS, Windows, and browsers.
Audio8 ASR Infinite is a native streaming speech recognition model with selectable audio clock and configurable transcription delay, supporting unlimited-length audio transcription without drifting.
The article highlights how aggregate Word Error Rate (WER) metric fails to capture critical errors in voice agents for production, such as misrecognized bank codes and multilingual speech, necessitating per-field tracking for real-world applications like banking and collections.
The paper proposes optimizations to the Viterbi algorithm using the Hirschberg algorithm and constrained random walk, reducing memory usage from 140 GB to 5 MB and improving speed, enabling forced alignment to run on end-user devices for better scalability in speech processing.