speech-to-text

Tag

Cards List
#speech-to-text

@svpino: Cutting down noise before sending the audio to a speech-to-text model makes a huge improvement. Voice isolation is the …

X AI KOLs Timeline ↗ · yesterday Cached

Krisp released an open benchmark and dataset showing that voice isolation reduces word error rates in speech-to-text models by 73%, with significant improvements across workplace and call-center recordings.

0 favorites 0 likes
#speech-to-text

Un-fused our realtime voice stack (STT -> LLM -> TTS) and cut cost ~14x. the tradeoff is latency, plus one upside i didn't expect

Reddit r/AI_Agents ↗ · 3d ago

Splitting a fused real-time voice AI stack into separate STT, LLM, and TTS stages cut costs by about 14x but increased latency, with an unexpected benefit of better inspectability for content guardrails.

0 favorites 0 likes
#speech-to-text

Reading Less While Writing: A Closed-Form Bandwidth Dial for Streaming Multimodal Decoders

arXiv cs.CL ↗ · 5d ago Cached

The paper introduces ZENDAYA, a closed-form bandwidth dial for streaming multimodal decoders that dynamically adjusts input reading to improve real-time text generation performance. It demonstrates that reducing input consumption can enhance quality and efficiency in streaming settings across video and audio benchmarks.

0 favorites 0 likes
#speech-to-text

Introducing Grok Voice Transcribe 2.0 (3 minute read)

TLDR AI ↗ · 5d ago Cached

Grok Voice Transcribe 2.0 is a new speech-to-text model from x.ai that doubles the accuracy of its predecessor and excels in multilingual and real-world audio transcription, ranking first on public leaderboards.

0 favorites 0 likes
#speech-to-text

@RayFernando1337: Subscriptions are dead...local models are smart enough to do things on device and most people are overpaying to do basi…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

RayFernando1337 announces the launch of SayRay, a local AI-powered tool that converts speech to text on-device. It aims to provide an affordable alternative to overpriced subscriptions for basic accessibility needs.

0 favorites 0 likes
#speech-to-text

Shall We Talk

Product Hunt ↗ · 2026-09-14 Cached

Shall We Talk is an open-source voice dictation tool for iPhone and Mac that converts speech to clean text, provides speaker-labeled transcripts, and includes cleanup features while preserving user wording.

0 favorites 0 likes
#speech-to-text

Meta's Muse Voice Transcribe (4 minute read)

TLDR AI ↗ · 2026-09-02 Cached

Meta introduces Muse Voice Transcribe, a real-time audio perception model that excels in streaming ASR and diarization with multilingual support, topping public benchmarks.

0 favorites 0 likes
#speech-to-text

Latency in Voice AI: Why milliseconds decide whether a call feels human

Reddit r/AI_Agents ↗ · 2026-09-01

The article emphasizes that low latency is critical for voice AI agents to maintain natural human-like conversations, as delays in the processing chain can make interactions feel artificial, especially in customer support and sales applications.

0 favorites 0 likes
#speech-to-text

Voiskey

Product Hunt ↗ · 2026-08-31 Cached

Voiskey is an AI voice typing tool that converts speech into context-appropriate text, offering 5x faster input than typing across multiple platforms and languages.

0 favorites 0 likes
#speech-to-text

@krandiash: Super cool to see we're #2 on this new benchmark out of the box (#1 soon) This makes Ink-2 a great fit for consumer app…

X AI KOLs Timeline ↗ · 2026-08-28 Cached

Ink-2 ranked #2 on the new VoiceCodeBench benchmark, demonstrating its suitability for real-time consumer apps, with GPT Live Transcribe taking the #1 spot.

0 favorites 0 likes
#speech-to-text

@vercel_dev: Gemini 3.5 Transcribe is live on AI Gateway with automatic detection for 85+ languages. Live audio + recordings: • 𝚐𝚘…

X AI KOLs Following ↗ · 2026-08-26 Cached

Vercel announces that Google's Gemini 3.5 Transcribe model is now available on AI Gateway, supporting live and recorded audio transcription in over 85 languages with automatic language detection and custom vocabulary.

0 favorites 0 likes
#speech-to-text

Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text

Ars Technica ↗ · 2026-08-26 Cached

Google has announced Gemini 3.5 Transcribe, an AI model for speech-to-text that improves speed and accuracy by removing fillers like 'ums' and supporting 85 languages, rolling out across its ecosystem.

0 favorites 0 likes
#speech-to-text

Gemini 3.5 Transcribe

Product Hunt ↗ · 2026-08-26

Gemini 3.5 Transcribe is announced as the most precise speech-to-text model yet, with discussion links provided on Product Hunt.

0 favorites 0 likes
#speech-to-text

@GoogleDeepMind: Gemini 3.5 Transcribe is our latest speech-to-text model for precise and intelligent transcriptions.

X AI KOLs ↗ · 2026-08-26 Cached

Google DeepMind announces Gemini 3.5 Transcribe, a new speech-to-text model for precise and intelligent transcriptions.

0 favorites 0 likes
#speech-to-text

Intelligent transcription with Gemini 3.5 Transcribe

Google DeepMind Blog ↗ · 2026-08-26 Cached

Google introduces Gemini 3.5 Transcribe, a new AI model for precise and intelligent real-time speech-to-text transcription, available via APIs for developers.

0 favorites 0 likes
#speech-to-text

superwhisper/s1-mini (8 minute read)

TLDR AI ↗ · 2026-08-20 Cached

superwhisper/s1-mini is a 0.6B-parameter text normalizer fine-tuned from Qwen3-0.6B to clean speech-to-text transcripts by removing fillers, correcting errors, and applying punctuation and formatting, achieving 94.8% accuracy on English data.

0 favorites 0 likes
#speech-to-text

S1-mini by Superwhisper running 100% locally in the browser on WebGPU with Transformers.js

Reddit r/LocalLLaMA ↗ · 2026-08-19

S1-mini is a 600M-parameter LLM that cleans up speech-to-text transcripts by removing errors and adding punctuation, running locally in the browser via WebGPU and Transformers.js.

0 favorites 0 likes
#speech-to-text

Launch HN: Speko (YC S26) – OpenRouter for Voice AI

Hacker News Top ↗ · 2026-08-17 Cached

Speko is a platform that optimizes and routes voice AI model stacks (STT, LLM, TTS) based on user constraints, with public benchmarks and an open-source gateway.

0 favorites 0 likes
#speech-to-text

Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper introduces MSRT, a framework with a resource-aware Mixture of Speech Encoders (MoSE) to overcome the curse of multilinguality in many-to-many speech-to-text translation. The 4B-parameter model achieves state-of-the-art results across 45 languages, particularly improving low-resource speech translation with only 10 hours of paired data per language.

0 favorites 0 likes
#speech-to-text

Before choosing an STT API, rank which transcript mistakes would actually hurt users.

Reddit r/AI_Agents ↗ · 2026-08-02

A guide on evaluating speech-to-text APIs by ranking transcript mistakes based on their actual impact on users, rather than raw accuracy metrics.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback