Realtime voice models compounds on cost (and forgets)- "Flowcat" fixed both (4x cheaper, 7x more context)
Summary
Flowcat addresses the high cost and limited context of realtime voice models, achieving 4x lower cost and 7x more context.
Similar Articles
Why real-time AI models cost up to 56× more than they should to serve, and how to fix it
dotwave.ai shares technical write-ups showing their inference engine can serve 56 concurrent real-time sessions on a single H100 (vs 1 for NVIDIA's reference stack) by keeping the model resident on the GPU and batching all sessions per tick, cutting per-conversation GPU cost by ~98% with zero missed 160ms audio deadlines.
Un-fused our realtime voice stack (STT -> LLM -> TTS) and cut cost ~14x. the tradeoff is latency, plus one upside i didn't expect
Splitting a fused real-time voice AI stack into separate STT, LLM, and TTS stages cut costs by about 14x but increased latency, with an unexpected benefit of better inspectability for content guardrails.
One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
A universal speech enhancement model that allows configurable control over both algorithmic and computational latency via parallel convolutions and early-exit mechanisms, enabling a single model to serve diverse real-time applications without retraining.
@tarat_211: I love Wispr Flow, but it's $12/month and still doesn't give my agent the visual context it needs. So I built BetterVoi…
A developer built BetterVoice, an open-source tool that transcribes speech locally and captures screen context for AI agents, as an alternative to Wispr Flow.
I made a sourced comparison of realtime speech-to-speech models
The author compiled a public, sourced comparison of realtime speech-to-speech AI models, detailing features like interruption behavior, pricing, and integrations with verified links.