Tag
Krisp released an open benchmark and dataset showing that voice isolation reduces word error rates in speech-to-text models by 73%, with significant improvements across workplace and call-center recordings.
Splitting a fused real-time voice AI stack into separate STT, LLM, and TTS stages cut costs by about 14x but increased latency, with an unexpected benefit of better inspectability for content guardrails.
The paper introduces ZENDAYA, a closed-form bandwidth dial for streaming multimodal decoders that dynamically adjusts input reading to improve real-time text generation performance. It demonstrates that reducing input consumption can enhance quality and efficiency in streaming settings across video and audio benchmarks.
Grok Voice Transcribe 2.0 is a new speech-to-text model from x.ai that doubles the accuracy of its predecessor and excels in multilingual and real-world audio transcription, ranking first on public leaderboards.
RayFernando1337 announces the launch of SayRay, a local AI-powered tool that converts speech to text on-device. It aims to provide an affordable alternative to overpriced subscriptions for basic accessibility needs.
Shall We Talk is an open-source voice dictation tool for iPhone and Mac that converts speech to clean text, provides speaker-labeled transcripts, and includes cleanup features while preserving user wording.
Meta introduces Muse Voice Transcribe, a real-time audio perception model that excels in streaming ASR and diarization with multilingual support, topping public benchmarks.
The article emphasizes that low latency is critical for voice AI agents to maintain natural human-like conversations, as delays in the processing chain can make interactions feel artificial, especially in customer support and sales applications.
Voiskey is an AI voice typing tool that converts speech into context-appropriate text, offering 5x faster input than typing across multiple platforms and languages.
Ink-2 ranked #2 on the new VoiceCodeBench benchmark, demonstrating its suitability for real-time consumer apps, with GPT Live Transcribe taking the #1 spot.
Vercel announces that Google's Gemini 3.5 Transcribe model is now available on AI Gateway, supporting live and recorded audio transcription in over 85 languages with automatic language detection and custom vocabulary.
Google has announced Gemini 3.5 Transcribe, an AI model for speech-to-text that improves speed and accuracy by removing fillers like 'ums' and supporting 85 languages, rolling out across its ecosystem.
Gemini 3.5 Transcribe is announced as the most precise speech-to-text model yet, with discussion links provided on Product Hunt.
Google DeepMind announces Gemini 3.5 Transcribe, a new speech-to-text model for precise and intelligent transcriptions.
Google introduces Gemini 3.5 Transcribe, a new AI model for precise and intelligent real-time speech-to-text transcription, available via APIs for developers.
superwhisper/s1-mini is a 0.6B-parameter text normalizer fine-tuned from Qwen3-0.6B to clean speech-to-text transcripts by removing fillers, correcting errors, and applying punctuation and formatting, achieving 94.8% accuracy on English data.
S1-mini is a 600M-parameter LLM that cleans up speech-to-text transcripts by removing errors and adding punctuation, running locally in the browser via WebGPU and Transformers.js.
Speko is a platform that optimizes and routes voice AI model stacks (STT, LLM, TTS) based on user constraints, with public benchmarks and an open-source gateway.
This paper introduces MSRT, a framework with a resource-aware Mixture of Speech Encoders (MoSE) to overcome the curse of multilinguality in many-to-many speech-to-text translation. The 4B-parameter model achieves state-of-the-art results across 45 languages, particularly improving low-resource speech translation with only 10 hours of paired data per language.
A guide on evaluating speech-to-text APIs by ranking transcript mistakes based on their actual impact on users, rather than raw accuracy metrics.