Ultrafast Qwen3-TTS at 34 ms Time-to-First-Audio, Handling 10 Requests Per Second [OSS]
Summary
Nari Qwen3-TTS is a high-performance serving implementation for the Qwen3-TTS model, achieving sub-50ms time-to-first-audio and handling 10 requests per second on a single H100 GPU.
View Cached Full Text
Cached at: 08/21/26, 07:07 PM
nari-labs/nari-qwen3-tts
Source: https://github.com/nari-labs/nari-qwen3-tts
Nari Qwen3-TTS
TL;DR
Nari Qwen3-TTS is a high-performance, single-H100 serving implementation of Qwen3-TTS 1.7B CustomVoice. It exposes streaming and non-streaming speech generation over HTTP, plus WebSocket-based input streaming for incremental text input.
It achieves 10 requests per second (RPS) and sub-50 ms p95 time-to-first-audio (TTFA) while maintaining real-time playback on a single NVIDIA H100 SXM. Even at 20 RPS, it sustains sub-80 ms p95 TTFA.
Below is a performance comparison with popular serving engines.

Methodology
Read our blog post for the full methodology. For details on the benchmark, see the benchmark repository.
Run with Docker
The container requires an NVIDIA H100, the NVIDIA Container Toolkit, and a driver compatible with CUDA 13.0. The published image supports Linux x86_64 (linux/amd64) H100 hosts. The engine has been tested only with English as the primary language.
docker run --rm --gpus all \
-p 8000:8000 \
-e HF_TOKEN \
-e QWEN3_TTS_PROFILE=ttfa \
-v nari-qwen3-tts-cache:/home/nari/.cache \
ghcr.io/nari-labs/nari-qwen3-tts:latest
Model files and compiled kernels are cached in the named volume. Once model loading and CUDA Graph capture finish, check readiness with:
curl --fail http://127.0.0.1:8000/ready
To build the image locally instead:
docker build \
--build-arg VCS_REF="$(git rev-parse HEAD)" \
-t nari-qwen3-tts:local .
Run with uv
Python and uv versions are pinned in .python-version and pyproject.toml.
The checked-in uv.lock defines the complete environment.
Install the required system packages (Debian/Ubuntu):
sudo apt-get install -y build-essential libsndfile1 sox
uv sync --frozen --extra codec --extra cuda --extra serving
uv run --frozen nari-qwen3-tts-server --profile ttfa
Add --local-files-only after the model is cached to prevent downloads at
startup. --frozen is intentional: an out-of-date lockfile fails instead of
silently resolving a different CUDA or Python environment.
The distribution name uses hyphens, while Python imports use underscores:
from nari_qwen3_tts import ModelAssetConfig, open_model
Profiles
ttfa: prioritizes time to first audio with latency-oriented scheduling and smaller initial Codec chunks.balanced: the default container profile, balancing first-audio latency and sustained request throughput.throughput: uses larger Codec chunks and batches to prioritize aggregate throughput under load.
Select a Docker profile with QWEN3_TTS_PROFILE, or pass --profile to
nari-qwen3-tts-server. You can also set QWEN3_TTS_MODEL_CACHE_DIR and
QWEN3_TTS_LOCAL_FILES_ONLY=1 in Docker.
For advanced tuning, apply a strict partial YAML overlay:
nari-qwen3-tts-server \
--local-files-only \
--engine-config /path/to/engine.yaml
The overlay may name its packaged base with extends: ttfa, balanced, or
throughput; alternatively, pass --profile and omit extends. Unknown keys,
invalid capture lists, and profile/base mismatches fail before model loading.
The fully resolved config and its SHA-256 are printed at startup.
API and architecture
The service exposes:
GET /healthGET /readyGET /v1/modelsPOST /v1/audio/speechWS /v1/audio/speech/ws
See WebSocket speech API for the live-text protocol and client example.
HTTP client example
POST /v1/audio/speech follows the OpenAI Audio Speech request shape. Nari Qwen3-TTS
supports only the model listed above, wav and pcm output, and speed: 1.0.
It also accepts Nari-specific controls such as language.
curl http://127.0.0.1:8000/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice",
"input": "Hello from Nari Labs.",
"voice": "ryan",
"language": "english",
"response_format": "wav",
"stream": false
}' \
--output speech.wav
Readiness remains false until CUDA Graph capture and a warm-up TTS request have both completed.
Thanks and references
Built by Nari Labs. Thanks to the Qwen3-TTS authors for releasing Qwen3-TTS, and to these projects for their work on high-performance multimodal and speech serving:
Similar Articles
Qwen3 TTS is seriously underrated - I got it running locally in real-time and it's one of the most expressive open TTS models I've tried
Developer shows how to run Qwen3 TTS locally in real-time with streaming, quantization, word-level alignment, and custom voice fine-tuning for an expressive open-source TTS pipeline.
Qwen3-TTS Technical Report
The Qwen3-TTS technical report introduces a series of advanced multilingual text-to-speech models with voice cloning and controllable generation, featuring a dual-track LM architecture and specialized tokenizers for low-latency streaming.
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Alibaba's Qwen team releases Qwen3-TTS-12Hz-1.7B-CustomVoice, a powerful text-to-speech model supporting 10 languages with low-latency streaming, instruction-based voice control, and robust contextual understanding.
How fast can I get a voice assistant to respond without a GPU? Qwen3-ASR and Kokoro-TTS ONNX on CPU.
Explores the performance of running a voice assistant with Qwen3-ASR and Kokoro-TTS ONNX models on CPU, measuring response times without a GPU.
I pushed Qwen3.8-27B to 99 tps single request and 1150 tps with a batch request on a RTX 3090
The author optimized the Qwen3.8-27B model inference on an RTX 3090 GPU, achieving up to 99 tokens per second for single requests and 1150 tps with batch processing through various quantization and optimization techniques, and released the updated code on GitHub.