Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Summary
This paper presents Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained on synthetic speech from a voice-cloning teacher, achieving high accuracy and prosody for on-device use without reference audio.
View Cached Full Text
Cached at: 09/14/26, 02:36 PM
Paper page - Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
Source: https://huggingface.co/papers/2609.03502
Abstract
A compact Thai text-to-speech model is distilled from a large voice-cloning teacher using synthetic data from brief voice references, achieving strong on-device accuracy and prosody with minimal errors.
In low-resource settings, deploying TTS typically requires choosing between a largevoice-cloningmodel with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a largevoice-cloningmodel as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compactfixed-voice studenttrained entirely onsynthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries,lexical tone, names and loanwords, numeric verbalization, and Thai-Englishcode-switching. We study howtext preparation, synthetic generation,quality filtering,rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluateCER, Challenge-Set Keyword Accuracy,Prosody Pause Accuracy,speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-deviceThai TTSwithout reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1%CERon Thai and English, respectively. We open-source the model and evaluation framework forThai TTSdevelopment.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2609\.03502
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper1
#### wayu-ai/wayu-paxa-tts-edge Text-to-Speech• Updated11 days ago • 63 • 1
Datasets citing this paper1
#### wayu-ai/thai-tts-keyword-bench Viewer• Updated11 days ago • 1.53k • 94
Spaces citing this paper2
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech
This paper presents a method to build a compact fixed-voice Thai TTS system using synthetic speech from a larger model, evaluating its performance and introducing an 82M-parameter model for on-device deployment.
Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care
This paper introduces a synthetic Bengali speech dataset of 10,000 audio-text pairs for telecom customer care scenarios, generated using OmniVoice voice-cloning, and evaluates it with an ASR model, achieving low word error rates.
FastThaiG2P: Lightning-fast Thai Grapheme-to-phoneme Conversion for Voice Agent Pipelines
This paper presents FastThaiG2P, a sub-millisecond Thai grapheme-to-phoneme conversion tool for TTS pipelines, achieving 0.15 ms average latency on a 27k-utterance benchmark. The authors demonstrate it by training a StyleTTS 2 Thai TTS model on a phonemized 20-hour open dataset.
Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions
Poly-InstructTTS is a text-to-speech system that learns expressive speech from open-ended natural language instructions using a large-scale multi-modal dataset, improving instruction adherence and expressiveness in TTS models.
kyutai-labs/pocket-tts
Kyutai releases Pocket TTS, a lightweight text-to-speech model that runs efficiently on CPUs with 100M parameters, low latency, and voice cloning, supporting multiple languages.