Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Hugging Face Daily Papers Papers

Summary

This paper presents Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained on synthetic speech from a voice-cloning teacher, achieving high accuracy and prosody for on-device use without reference audio.

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.
Original Article
View Cached Full Text

Cached at: 09/14/26, 02:36 PM

Paper page - Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Source: https://huggingface.co/papers/2609.03502

Abstract

A compact Thai text-to-speech model is distilled from a large voice-cloning teacher using synthetic data from brief voice references, achieving strong on-device accuracy and prosody with minimal errors.

In low-resource settings, deploying TTS typically requires choosing between a largevoice-cloningmodel with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a largevoice-cloningmodel as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compactfixed-voice studenttrained entirely onsynthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries,lexical tone, names and loanwords, numeric verbalization, and Thai-Englishcode-switching. We study howtext preparation, synthetic generation,quality filtering,rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluateCER, Challenge-Set Keyword Accuracy,Prosody Pause Accuracy,speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-deviceThai TTSwithout reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1%CERon Thai and English, respectively. We open-source the model and evaluation framework forThai TTSdevelopment.

View arXiv pageView PDFAdd to collection

Get this paper in your agent:

hf papers read 2609\.03502

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### wayu-ai/wayu-paxa-tts-edge Text-to-Speech• Updated11 days ago • 63 • 1

Datasets citing this paper1

#### wayu-ai/thai-tts-keyword-bench Viewer• Updated11 days ago • 1.53k • 94

Spaces citing this paper2

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

kyutai-labs/pocket-tts

GitHub Trending (daily)

Kyutai releases Pocket TTS, a lightweight text-to-speech model that runs efficiently on CPUs with 100M parameters, low latency, and voice cloning, supporting multiple languages.