Tag
Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.
OpenAI releases two new transcription models: GPT Live Transcribe for low-latency and GPT Transcribe for batch workloads, with up to 41% lower error rates and improved semantic accuracy using context.
Wispro is a voice typing tool that allows users to dictate text instead of typing, converting speech into written content.
Wisprkey is a Mac AI assistant that lets you talk to any app, enabling hands-free voice input.
This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.
Elon Musk announces Grok Build, a new feature that allows users to talk to Grok like a person using speech-to-text for task completion, including coding tasks.
Speech To Markdown is a product that converts speech into markdown format using local AI, designed for note-taking.
The author built callitdone.today, an AI voice agent that makes real phone calls, navigates IVR menus, waits on hold, and speaks with humans, sharing key technical challenges and lessons learned.
Simulstream is an open-source framework for evaluating and demonstrating streaming speech-to-text translation systems, supporting both incremental and re-translation decoding on long-form speech with fine-grained logging and an interactive web interface.
This article introduces local large model hardware configurations from $2,000 to $40,000, including detailed setups from dual RTX 3090 to quad RTX PRO 6000, covering PCIe switches, GPU communication, Docker configuration, and speech-to-text.
The article argues that voice agent STT should be evaluated on entity accuracy (e.g., phone numbers, dates) rather than general word error rate, because missing critical fields can break workflows. It mentions testing with HubSpot fields and notes Smallest AI Pulse as an interesting tool for capturing workflow-critical entities in real time.
EdgeSpeak officially launched, a local-first, privacy-preserving accurate transcription tool, supporting semantic segmentation and timestamps, compatible with OpenAI Audio API, etc. It will later add speaker labeling and speech generation features.
A new speech-to-text tool claims to rival Dragon Professional on Windows, offering a competitive alternative for voice recognition.
Describes a medical speech-to-text system that runs locally on a MacBook, enabling streaming transcription without cloud dependency.
An evaluation of leading STT models on 1000+ noisy real-world clips reveals most perform poorly in noisy environments, with DG Nova performing best. Applying noise cancellation significantly improves accuracy.
The author argues that for live voice agents, STT latency and real-time behavior are more critical than raw transcription accuracy, and proposes a different evaluation scorecard.
Explores whether OpenAI's Whisper remains the top choice for real-time speech-to-text applications, considering alternatives and performance trade-offs.
A guide to building a fully local voice assistant using Platypush on a Raspberry Pi, covering hotword detection, speech-to-text, text-to-speech, and home automation integration.
Google AI Edge Eloquent now supports Mac as a fully local Wispr Flow alternative, offering real-time voice transcription and voice command text editing based on the latest Gemma model. Free, no subscription, and fully private locally.
Mutter AI Dictation is a private AI dictation tool that operates offline.