Why real-time ASR is so expensive to serve — and how to fix it

Reddit r/AI_Agents Products

Summary

The team behind .wave introduces an inference engine using a Wave Persistent Kernel compiler to serve NVIDIA Nemotron 3.5 ASR Streaming at $0.00045/minute, sustaining 4,800 concurrent streams on one H100 with p99 frame completion under 106.3 ms — roughly 20x NVIDIA's published stream capacity and a fraction of competing ASR pricing.

Hey everyone, we built .wave, an inference engine for models that run continuously. We’re now using it to serve NVIDIA Nemotron 3.5 ASR Streaming 0.6B at $0.00045/minute. I wanted to share how the engine works, why it matters for streaming workloads, and what we measured. When you’re building a voice agent, transcription becomes part of an ongoing conversation. Audio keeps arriving, each stream carries context, and adding concurrent calls means keeping all of those streams moving. A useful starting point is understanding what limits your GPU: computation, memory traffic, or time spent waiting for the next operation. Small operations can leave substantial capacity unused, even on powerful hardware. This is why profiling the whole serving path matters. A fast encoder benchmark doesn’t tell you how the encoder, decoder, scheduling, and transport behave together under load. At the core of our engine is WPK — Wave Persistent Kernel. We express the model in Python Tensor IR and reuse a library of CUDA operators. The compiler resolves dependencies, assigns work to GPU workers, and plans memory reuse. It produces a profile for a specific model and stream capacity. That changes execution in three ways: The GPU program stays resident. Instead of repeatedly handing control back to the CPU between operations, GPU workers execute the compiled phases. The host supplies inputs and collects outputs. Compatible streams advance together. A weight tile can serve multiple streams in a cohort before execution moves on. Their individual computation still exists, but they share weight reads. Streaming state stays resident. Each stream keeps its own cache between audio chunks. The memory layout is prepared ahead of execution, reducing dynamic allocation and coordination in the recurring path. Batching and CUDA Graphs can already reduce some overhead. WPK makes the cross-operation schedule explicit in a persistent program. For ASR, we compile the audio front end, encoder, and decoder together. The economic benefit follows from usable concurrency: GPU cost per stream-hour ≈ GPU hourly cost ÷ concurrent streams sustained This approximation assumes a full GPU at steady load. Operating the service also involves networking, CPU resources, and utilization between traffic peaks. Our highest tested configuration achieved: 4,800 concurrent streams on one H100 SXM 80GB 80 ms audio chunks 7.8 million frames completed per run, across three consecutive runs 106.3 ms worst-stream p99 completion interval, below the 120 ms bound Median frame-release-to-completion latency was 207–211 ms, including an 80 ms jitter buffer. The 80 ms chunk size describes the input cadence; it isn’t the measured end-to-end latency. Measurements used authenticated WebSockets over loopback on the serving host, excluding WAN latency. NVIDIA publishes 240 streams for its 80 ms/H100 configuration, giving a 20× capacity ratio. That comparison uses its published figure, not a matched NIM run. Transcript disagreement against an fp32 reference averaged 0.23% over 23 recordings. This measures reference parity, not accuracy against human transcripts. If you’re optimizing your own streaming stack, I’d investigate three things: Gaps between GPU operations: where does execution wait for the host? Repeated work across streams: are compatible operations actually batched? Latency under concurrency: do slower streams maintain their cadence as you add sessions? For a voice-agent application, stable transcription under load helps keep the rest of the pipeline moving. It also allows more calls to share the same GPU allocation. Our hosted service supports 32 languages, an OpenAI-compatible Realtime interface, and a Deepgram-compatible endpoint usable with existing LiveKit and Pipecat Deepgram plugins. At $0.00045/minute, the published rate is: One tenth of Together AI’s $0.0045/minute for the same model. About one seventeenth of Deepgram Flux Multilingual’s $0.0078/minute pay-as-you-go rate. These compare minute rates; accuracy, features, and billing behavior still need evaluation for your application. We bill from the first audio until the session closes. The service is in public beta. We haven’t published per-language accuracy evaluations yet, and there are 230H in signup credits, with no card required, to test your own audio. I’ve put the technical explanation, benchmark methodology, integration examples, and pricing sources in the first comment. Happy to discuss the compiler, GPU execution, or how you measure streaming performance in your own stack.
Original Article

Similar Articles

Why real-time AI models cost up to 56× more than they should to serve, and how to fix it

Reddit r/ArtificialInteligence

dotwave.ai shares technical write-ups showing their inference engine can serve 56 concurrent real-time sessions on a single H100 (vs 1 for NVIDIA's reference stack) by keeping the model resident on the GPU and batching all sessions per tick, cutting per-conversation GPU cost by ~98% with zero missed 160ms audio deadlines.

nvidia/nemotron-3.5-asr-streaming-0.6b

Hugging Face Models Trending

NVIDIA releases Nemotron 3.5 ASR, a 600M parameter multilingual streaming speech recognition model supporting 40 language-locales with a Cache-Aware FastConformer-RNNT architecture for low-latency transcription. The model supports configurable chunk sizes and is ready for commercial use under the OpenMDW-1.1 license.

@kwindla: https://x.com/kwindla/status/2062544580105359686

X AI KOLs Timeline

NVIDIA released Nemotron 3.5 ASR, an open-source multilingual speech-to-text model with the lowest latency tested, available in multilingual and English-only variants, ideal for voice agents and self-hosted deployments.