speech-to-text

Tag

Cards List
#speech-to-text

Room reverberation and low SNR hurt STT accuracy far more than model size

Reddit r/ArtificialInteligence ↗ · 2026-07-30

Room reverberation and low-frequency noise from the environment hurt speech-to-text accuracy far more than the choice of model size; front-end audio preprocessing like adaptive spectral subtraction can recover masked phonemes and reduce word error rate more effectively than upgrading the model backend.

0 favorites 0 likes
#speech-to-text

@reach_vb: Two new transcription models are now available in the API! > GPT Live Transcribe for low-latency live transcription > G…

X AI KOLs Following ↗ · 2026-07-28 Cached

OpenAI releases two new transcription models: GPT Live Transcribe for low-latency and GPT Transcribe for batch workloads, with up to 41% lower error rates and improved semantic accuracy using context.

0 favorites 0 likes
#speech-to-text

Wispro

Product Hunt ↗ · 2026-07-22

Wispro is a voice typing tool that allows users to dictate text instead of typing, converting speech into written content.

0 favorites 0 likes
#speech-to-text

Wisprkey

Product Hunt ↗ · 2026-07-22

Wisprkey is a Mac AI assistant that lets you talk to any app, enabling hands-free voice input.

0 favorites 0 likes
#speech-to-text

Constrained CTC Decoding for Efficient Diacritic Restoration

arXiv cs.CL ↗ · 2026-07-22 Cached

This paper proposes a non-autoregressive CTC-based approach for speech-to-text diacritic restoration in Arabic, incorporating hard constraints during decoding to improve efficiency and reduce error rates.

0 favorites 0 likes
#speech-to-text

@elonmusk: You can talk to Grok like a person to accomplish tasks via Grok Build http://X.ai/cli

X AI KOLs Timeline ↗ · 2026-07-22 Cached

Elon Musk announces Grok Build, a new feature that allows users to talk to Grok like a person using speech-to-text for task completion, including coding tasks.

0 favorites 0 likes
#speech-to-text

Speech To Markdown

Product Hunt ↗ · 2026-07-19

Speech To Markdown is a product that converts speech into markdown format using local AI, designed for note-taking.

0 favorites 0 likes
#speech-to-text

Building an AI agent that makes real phone calls (hold music, IVRs, angry humans), here’s what I learned so far

Reddit r/AI_Agents ↗ · 2026-07-10

The author built callitdone.today, an AI voice agent that makes real phone calls, navigates IVR menus, waits on hold, and speaks with humans, sharing key technical challenges and lessons learned.

0 favorites 0 likes
#speech-to-text

Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems

arXiv cs.CL ↗ · 2026-07-09 Cached

Simulstream is an open-source framework for evaluating and demonstrating streaming speech-to-text translation systems, supporting both incremental and re-translation decoding on long-form speech with fine-grained logging and an interactive web interface.

0 favorites 0 likes
#speech-to-text

@PandaTalk8: A hardware and configuration guide for running large models locally. The author shares local LLM setups ranging from about $2,000 to $40,000: the budget option uses dual RTX 3090s to run Qwen and local speech-to-text; the high-end option uses 4 RTX PRO 6000 cards with 384GB…

X AI KOLs Timeline ↗ · 2026-07-04 Cached

This article introduces local large model hardware configurations from $2,000 to $40,000, including detailed setups from dual RTX 3090 to quad RTX PRO 6000, covering PCIe switches, GPU communication, Docker configuration, and speech-to-text.

0 favorites 0 likes
#speech-to-text

My voice agent sounded smart until one phone number was transcribed wrong.

Reddit r/AI_Agents ↗ · 2026-07-03

The article argues that voice agent STT should be evaluated on entity accuracy (e.g., phone numbers, dates) rather than general word error rate, because missing critical fields can break workflows. It mentions testing with HubSpot fields and notes Smallest AI Pulse as an interesting tool for capturing workflow-critical entities in real time.

0 favorites 0 likes
#speech-to-text

@FeitengLi: Next week, after adding speaker labeling and speech generation, it won't be this cheap early bird price anymore.

X AI KOLs Timeline ↗ · 2026-07-03 Cached

EdgeSpeak officially launched, a local-first, privacy-preserving accurate transcription tool, supporting semantic segmentation and timestamps, compatible with OpenAI Audio API, etc. It will later add speaker labeling and speech generation features.

0 favorites 0 likes
#speech-to-text

STT That Can Challenge Dragon Professional on Windows

Reddit r/LocalLLaMA ↗ · 2026-06-27

A new speech-to-text tool claims to rival Dragon Professional on Windows, offering a competitive alternative for voice recognition.

0 favorites 0 likes
#speech-to-text

Streaming medical STT running locally on a MacBook

Reddit r/LocalLLaMA ↗ · 2026-06-26

Describes a medical speech-to-text system that runs locally on a MacBook, enabling streaming transcription without cloud dependency.

0 favorites 0 likes
#speech-to-text

I evaluated top STT models on large Real-World data

Reddit r/AI_Agents ↗ · 2026-06-25

An evaluation of leading STT models on 1000+ noisy real-world clips reveals most perform poorly in noisy environments, with DG Nova performing best. Applying noise cancellation significantly improves accuracy.

0 favorites 0 likes
#speech-to-text

Best STT API for voice agents? I’d test latency before accuracy

Reddit r/AI_Agents ↗ · 2026-06-25

The author argues that for live voice agents, STT latency and real-time behavior are more critical than raw transcription accuracy, and proposes a different evaluation scorecard.

0 favorites 0 likes
#speech-to-text

Is Whisper still the best default for speech-to-text if the app needs to be real time?

Reddit r/AI_Agents ↗ · 2026-06-23

Explores whether OpenAI's Whisper remains the top choice for real-time speech-to-text applications, considering alternatives and performance trade-offs.

0 favorites 0 likes
#speech-to-text

A fully local voice assistant setup

Lobsters Hottest ↗ · 2026-06-22 Cached

A guide to building a fully local voice assistant using Platypush on a Raspberry Pi, covering hotword detection, speech-to-text, text-to-speech, and home automation integration.

0 favorites 0 likes
#speech-to-text

@iluciddreaming: Google just killed another startup... Google AI Edge Eloquent now supports Mac, a fully local Wispr Flow alternative. Based on the latest Gemma model, supports real-time voice transcription + voice commands to edit text. Free, no subscription, no...

X AI KOLs Timeline ↗ · 2026-06-22 Cached

Google AI Edge Eloquent now supports Mac as a fully local Wispr Flow alternative, offering real-time voice transcription and voice command text editing based on the latest Gemma model. Free, no subscription, and fully private locally.

0 favorites 0 likes
#speech-to-text

Mutter AI Dictation

Product Hunt ↗ · 2026-06-19

Mutter AI Dictation is a private AI dictation tool that operates offline.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback