FreyaTTS Technical Report
Summary
FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model based on a non-autoregressive conditional flow-matching Diffusion Transformer, achieving state-of-the-art performance with a fraction of the parameters of larger systems and released under Apache-2.0.
View Cached Full Text
Cached at: 07/13/26, 07:58 AM
# FreyaTTS Technical Report
Source: [https://arxiv.org/html/2607.09530](https://arxiv.org/html/2607.09530)
###### Abstract
We introduce FreyaTTS, a compact, tokenizer\-free, Turkish\-first text\-to\-speech model designed for highly reliable and efficient conversational synthesis\. FreyaTTS is a183\.2183\.2M\-parameter non\-autoregressive conditional flow\-matching Diffusion Transformer \(DiT\) that operates in the frozen continuous\-latent space of the frozen AudioVAE2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\), reused unmodified \(1616kHz encode,4848kHz decode\); holding the codec fixed lets the model devote all of its capacity to the text\-to\-latent map while inheriting4848kHz reconstruction for free\. We advance the framework across three key dimensions: \(i\)Rule\-Free End\-to\-End Modeling, by driving generation end\-to\-end from a9292\-symbol Turkish character vocabulary with no phonemizer, grapheme\-to\-phoneme frontend, or discrete speech tokenizer, so that agglutinative morphology, vowel harmony, and the spoken form of numbers and acronyms are learned directly from audio; \(ii\)Non\-Autoregressive Parallel Denoising, which avoids the left\-to\-right error accumulation of autoregressive decoders by predicting and denoising the entire latent sequence in parallel over a predicted duration; and \(iii\)Production\-Hardening Post\-Training, utilizing a two\-stage post\-training recipe–a single\-speaker voice lock that stabilizes speaker identity \(collapsing cross\-generationF0F\_\{0\}standard deviation from74\.974\.9Hz to5\.05\.0Hz\) followed by short\-utterance coverage to resolve isolated\-token and short\-phrase failures\. On our Freya\-TR\-Eval benchmark, FreyaTTS achieves a band\-matched WER of8\.08\.0% and CER of3\.03\.0%, outperforming larger open\-source systems at a fraction of their parameter size\. With a real\-time factor of0\.110\.11on consumer GPUs and the ability to run faster than real time on a laptop CPU, the model is highly optimized for resource\-constrained edge deployment\. We release the model weights, the training and inference code, and the evaluation benchmark under the Apache\-2\.0 license\.
Contents
## 1Introduction
Text\-to\-speech \(TTS\) has advanced from producing merely intelligible speech toward generating natural, expressive, and controllable audio\(Shenet al\.,[2018](https://arxiv.org/html/2607.09530#bib.bib34); Renet al\.,[2020](https://arxiv.org/html/2607.09530#bib.bib2)\), driven by the large\-language\-model paradigm of framing synthesis as sequence modeling over discrete audio tokens\(Borsoset al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib7); Chenet al\.,[2025a](https://arxiv.org/html/2607.09530#bib.bib6)\)and, more recently, by continuous\-latent flow\-matching generators\(Shenet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib16); Leet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib21); Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20)\)\. This progress, however, is concentrated in high\-resource English and Chinese\. Mid\-resource languages such as Turkish, for which public corpora such as FLEURS\-tr\(Conneauet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib112)\)and Common\-Voice\-tr\(Ardilaet al\.,[2020](https://arxiv.org/html/2607.09530#bib.bib113)\)offer only tens of hours of validated read speech, remain underserved: large multilingual systems\(Casanovaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib13); Duet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib24)\)treat Turkish as one language among dozens rather than as a first\-class target\. A second gap is computational\. The strongest open systems reach their quality through multi\-billion\-parameter backbones trained on hundreds of thousands to millions of hours of speech \(for example, CosyVoice 3 on roughly one million hours and Qwen3\-TTS on over five million\)\(Duet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib24); Huet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib31)\), an operating point that excludes single\-GPU serving and on\-device deployment\. Turkish additionally stresses the text frontend: agglutinative morphology, vowel harmony, and the spoken\-form expansion of numbers, dates, and acronyms make hand\-built grapheme\-to\-phoneme and normalization pipelines brittle and incomplete\. These pressures argue for a compact, Turkish\-first system that learns pronunciation end\-to\-end from audio rather than dictating it by rules\.
Two modeling paradigms dominate contemporary TTS\. Discrete\-token systems represent speech as codec tokens and inherit LLM\-style scaling and in\-context learning\(Borsoset al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib7); Chenet al\.,[2025a](https://arxiv.org/html/2607.09530#bib.bib6); Duet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib24)\), but quantization discards fine acoustic detail and typically forces a multi\-stage pipeline\. Continuous\-latent systems instead model speech representations directly with denoising or flow\-matching objectives, spanning non\-autoregressive diffusion models\(Shenet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib16); Leet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib21); Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20); Eskimezet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib33)\)and diffusion\-autoregressive hybrids\(Liet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib36); Jiaet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib37)\)\. VoxCPM\(Zhouet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib90)\)unified these lines through a differentiable finite\-scalar\-quantization bottleneck\(Mentzeret al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib47)\)inside a single continuous\-latent backbone, and its successor VoxCPM2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\)scales the design to a 2B\-parameter, 48 kHz, 30\-language foundation model whose asymmetric AudioVAE encodes at 16 kHz and reconstructs at 48 kHz\. FreyaTTS takes a different route: we treat its AudioVAE2 as a*frozen*, Apache\-2\.0 continuous\-latent codec, a fixed2525Hz,6464\-dimensional latent space, and train a compact183\.2183\.2M\-parameter generator entirely inside it\. Because high\-fidelity4848kHz reconstruction is already solved and held fixed, every parameter and training signal is spent on the text\-conditional prior rather than on waveform modeling, and a from\-scratch Turkish pretraining run over a large, high\-quality corpus yields a strong Turkish\-first model at a fraction of the size of the multilingual foundation models above\.
The second design decision concerns*how*the latent sequence is generated\. VoxCPM and VoxCPM2 decode autoregressively, one acoustic patch at a time, atop a MiniCPM\-4 backbone\(Teamet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib45)\)\. Autoregression is expressive, but it accumulates errors over long horizons and is prone to number and word garbling, a benign glitch in a story reader but a critical failure wherever a misread digit changes meaning\. FreyaTTS therefore adopts a non\-autoregressive formulation: a conditional flow\-matching Diffusion Transformer predicts an entire utterance’s latent sequence in parallel over a separately predicted duration, sidestepping left\-to\-right error accumulation by construction\. Text conditions the latents through cross\-attention over ConvNeXt\-refined character features from a9292\-symbol Turkish character vocabulary, with no byte\-pair encoding, phonemizer, or grapheme\-to\-phoneme step, so the model learns the pronunciation of in\-context numbers, currencies, and acronyms directly from audio\. What limits digit rendering in the NAR design is not decoding order but duration allocation, so digit strings are expanded to their spoken form at the text frontend, the standard input contract for character\-level synthesis, while isolated bare tokens are handled by post\-training and a thin inference layer \([sections˜3\.5](https://arxiv.org/html/2607.09530#S3.SS5)and[4](https://arxiv.org/html/2607.09530#S4)\)\. A two\-stage post\-training recipe then specializes the from\-scratch pretrained prior into a single\-voice production model\.
The main contributions of FreyaTTS are as follows\.
1. 1\.A tokenizer\-free non\-autoregressive Turkish TTS model\.FreyaTTS is a 183\.2M\-parameter conditional flow\-matching Diffusion Transformer that synthesizes 48 kHz Turkish speech in the frozen continuous\-latent space of AudioVAE2, from a 92\-symbol character vocabulary with no phonemizer, grapheme\-to\-phoneme frontend, or discrete speech tokenizer, to our knowledge the first openly released tokenizer\-free NAR TTS model pretrained from scratch as a Turkish\-first system, whereas prior open Turkish baselines relied on phonemizers or autoregressive architectures\(Pratapet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib115); Casanovaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib13)\)\. We release the single\-voice production model’s weights publicly\.
2. 2\.A pretrain\-to\-SFT production\-hardening recipe\.A full\-parameter*voice\-lock*stage on a single consented speaker writes the voice into the weights, collapsing cross\-generationF0F\_\{0\}standard deviation from74\.974\.9to5\.05\.0Hz, and a*short\-utterance\-coverage*stage over forced\-aligned one\-word and two\-word segments repairs the isolated\-token failures a sentence\-only corpus structurally cannot\.
3. 3\.A public benchmark and an efficiency\-oriented evaluation\.We releaseFreya\-TR\-Eval, a 495\-sentence, domain\-neutral Turkish benchmark, together with its seeded build scripts and our full evaluation harness, so that every number in this report is reproducible without redistributing third\-party audio\. Under an identical band\-matched Whisper WER/CER protocol, FreyaTTS achieves a WER of 8\.0 % and a CER of 3\.0 %\. The model also supports efficient batched non\-autoregressive inference, delivering high throughput with a small memory footprint on both consumer GPUs and laptop CPUs\.
The remainder of this report is organized as follows\.[Section˜2](https://arxiv.org/html/2607.09530#S2)situates FreyaTTS within large\-scale and low\-resource TTS foundation models and continuous\-latent flow\-matching generation\.[Section˜3](https://arxiv.org/html/2607.09530#S3)presents the FreyaTTS system: the tokenizer\-free non\-autoregressive architecture over the frozen VoxCPM2 latent space, the flow\-matching and duration objectives, the from\-scratch Turkish pretraining stage, and the two\-stage voice\-lock and short\-utterance\-coverage fine\-tuning recipe\.[Section˜4](https://arxiv.org/html/2607.09530#S4)reports the Turkish evaluation against openly available systems together with inference\- and serving\-efficiency measurements\.[Section˜5](https://arxiv.org/html/2607.09530#S5)discusses limitations and future directions\.
## 2Related Work
We situate FreyaTTS along two axes: the generative paradigm it instantiates, non\-autoregressive flow matching over a continuous latent space \([section˜2\.1](https://arxiv.org/html/2607.09530#S2.SS1)\), and the niche it fills, a compact Turkish\-first model in a field of large multilingual generalists \([section˜2\.2](https://arxiv.org/html/2607.09530#S2.SS2)\)\.
### 2\.1Speech Synthesis Paradigms
Contemporary TTS is organized by how speech is represented and generated\. We review the three families that frame FreyaTTS: discrete\-token language models, non\-autoregressive generation over continuous features, and diffusion\-autoregressive hybrids\.
#### Discrete\-token codec language models\.
The dominant paradigm casts synthesis as autoregressive language modeling over discrete acoustic tokens from a neural audio codec\(Défossezet al\.,[2022](https://arxiv.org/html/2607.09530#bib.bib10); Kumaret al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib11)\), inheriting the scaling and in\-context learning of large language models\. VALL\-E\(Chenet al\.,[2025a](https://arxiv.org/html/2607.09530#bib.bib6)\)cast zero\-shot cloning as prompt continuation over residual\-quantized tokens, and the CosyVoice family\(Duet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib22),[2025](https://arxiv.org/html/2607.09530#bib.bib24)\)pairs a language model over supervised semantic tokens with a separate flow\-matching decoder that restores acoustic detail\. The route scales well but pays two structural prices: quantization discards fine acoustic information, and the semantic\-token and acoustic\-decoder stages cannot be optimized jointly end to end\.
#### Non\-autoregressive generation over continuous features\.
A second family drops discrete tokens and models continuous acoustic features directly and in parallel, under a diffusion or conditional\-flow\-matching objective\. NaturalSpeech 2\(Shenet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib16)\)applies latent diffusion with explicit duration and pitch predictors, and Voicebox\(Leet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib21)\)frames TTS as text\-conditioned flow matching over mel frames guided by an external alignment and a duration model\. This alignment\-and\-duration template runs through diffusion TTS: Grad\-TTS\(Popovet al\.,[2021](https://arxiv.org/html/2607.09530#bib.bib123)\)couples a score\-based decoder with monotonic\-alignment search and a duration predictor, and Matcha\-TTS\(Mehtaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib122)\)recasts it under optimal\-transport conditional flow matching\. Nearby, StyleTTS 2\(Liet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib124)\)adds style diffusion and adversarial training, NaturalSpeech 3\(Juet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib125)\)factorizes a codec latent for attribute\-wise diffusion, and end\-to\-end VAE\-GAN models of the VITS\(Kimet al\.,[2021](https://arxiv.org/html/2607.09530#bib.bib126)\)family, on which the MMS Turkish baseline is built\(Pratapet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib115)\), learn alignment and waveform jointly\. A distinct branch removes the alignment and duration modules altogether: E2\-TTS\(Eskimezet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib33)\)pads the character sequence with filler tokens to the target length and lets a single Transformer learn the text\-to\-audio correspondence implicitly through self\-attention, and F5\-TTS\(Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20)\)keeps this filler\-padded formulation while adding a ConvNeXt text refiner and sway\-sampling inference; compact successors such as ZipVoice\(Zhuet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib88)\)push the same objective toward small, fast models\. FreyaTTS belongs to this family but conditions through cross\-attention over character features rather than filler\-padded implicit alignment, and it generates a frozen VoxCPM2 latent rather than a mel\-spectrogram\.
#### Diffusion\-autoregressive hybrids\.
A third family autoregresses over continuous latents while rendering each step with a local diffusion head, combining language\-model\-style planning with continuous acoustic fidelity\(Menget al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib35)\)\. The mechanism originates in image generation, where MAR\(Liet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib36)\)replaced the categorical softmax of a token\-based model with a small per\-token diffusion loss, showing that vector quantization is not required for autoregressive generation\. DiTAR\(Jiaet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib37)\)ported this to speech through a local diffusion transformer over patches of continuous latents, and VoxCPM\(Zhouet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib90)\)realized a semantic\-acoustic hierarchy inside a single continuous\-latent backbone via a differentiable finite\-scalar\-quantization bottleneck\(Mentzeret al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib47)\)atop a MiniCPM\-4 language model\(Teamet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib45)\)\. These hybrids are expressive but retain an autoregressive decode loop, with the exposure bias and error accumulation it carries over long or number\-dense utterances\.
#### Positioning FreyaTTS\.
FreyaTTS keeps VoxCPM2’s continuous\-latent representation but discards its autoregressive stack\. It reuses only the frozen AudioVAE, the2525Hz,6464\-dimensional,1616kHz\-in /4848kHz\-out codec of VoxCPM2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\), under its Apache\-2\.0 license and never trained, so representation learning is decoupled from generative modeling and4848kHz super\-resolution comes for free from the frozen decoder\. In place of the hierarchical TSLM/FSQ/RALM stack and the autoregressive loop, a single183\.2183\.2M\-parameter flow\-matching transformer predicts an utterance duration and denoises all latent frames in parallel \(a3232\-step Euler ODE at inference\)\. The system is tokenizer\-free on both ends: no discrete audio codec, and a9292\-symbol character vocabulary with no phonemizer or grapheme\-to\-phoneme frontend, so that the pronunciation of numbers, acronyms, and code\-switched terms is learned from audio rather than dictated by a rule\-based frontend\.
### 2\.2Turkish, Multilingual, and Efficient TTS
#### The mid\-resource regime: Turkish beyond nominal coverage\.
Progress in large\-scale TTS is measured almost exclusively on English and Chinese\. Training corpora are dominated by read and audiobook speech in those two languages, as in Emilia\(Heet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib48)\), and the zero\-shot benchmarks that define the state of the art, notably Seed\-TTS\-Eval\(Anastassiouet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib51)\), inherit the same distribution\. Massively multilingual systems nominally widen this coverage to dozens or hundreds of languages\(Casanovaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib13); Duet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib24); Zhuet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib30)\), and Turkish typically appears among them\. Nominal coverage, however, is not specialization: a single checkpoint must amortize a fixed capacity budget across every language it serves, and the supervision it receives for each is generic read speech\. The consequence is sharpest for languages that are neither high\- nor low\-resource\. Turkish occupies precisely this mid\-resource band, where public read\-speech corpora such as FLEURS\-tr\(Conneauet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib112)\)and Common\-Voice\-tr\(Ardilaet al\.,[2020](https://arxiv.org/html/2607.09530#bib.bib113)\)together supply only tens of hours, an order of magnitude below their English or Chinese counterparts, and where no openly available corpus targets a conversational domain\. A language in this regime is well served neither by the multilingual generalist, whose Turkish is a marginal slice of a shared budget, nor by the low\-resource literature, whose methods assume a scarcity that does not hold\. FreyaTTS is designed for this regime: a Turkish\-first model trained from scratch on a large in\-domain Turkish corpus, rather than one language slot in a multilingual grid\.
#### Compact and deployable synthesis\.
The strongest open systems obtain their quality from multi\-billion\-parameter backbones that preclude single\-GPU serving and on\-device use\. A parallel line pursues efficiency instead: VITS\-based on\-device voices such as Piper and the multilingual MMS\-TTS\(Pratapet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib115)\), and compact flow\-matching models such as ZipVoice\(Zhuet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib88)\), trade breadth for a small, fast footprint\. FreyaTTS shares this goal but reaches it differently\. A frozen high\-fidelity codec absorbs waveform reconstruction, leaving a183\.2183\.2M\-parameter non\-autoregressive generator that serves faster than real time on a consumer GPU and runs on a laptop CPU, while carrying a single production voice in its weights rather than cloning one from a reference prompt\.
## 3Methodology
### 3\.1Overview
FreyaTTS is a tokenizer\-free, non\-autoregressive \(NAR\) text\-to\-speech model realized as a conditional flow\-matching Diffusion Transformer \(DiT\) over the continuous latent space of a*frozen*AudioVAE2\. It shares VoxCPM’s core premise, modeling speech in a continuous latent space rather than over discrete codec tokens\(Zhouet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib90)\), but departs sharply from its generative backbone\. VoxCPM and VoxCPM2 render audio autoregressively: a MiniCPM\-4 language model\(Teamet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib45)\)predicts semi\-discrete FSQ semantic states\(Mentzeret al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib47)\)that a local diffusion head decodes patch by patch\. We instead discard the entire autoregressive stack \(the text\-semantic LM, the FSQ bottleneck, the residual acoustic LM, and the local diffusion transformer\) and keep only AudioVAE2, replacing the backbone with a single NAR DiT that denoises all latent frames of an utterance in parallel \([fig\.˜1](https://arxiv.org/html/2607.09530#S3.F1)\)\.
Figure 1:Overall architecture of FreyaTTS\. Character\-level Turkish text \(a9292\-symbol vocabulary; no byte\-pair encoding, phonemizer, or grapheme\-to\-phoneme frontend\) is embedded and refined by a ConvNeXt\-1d encoder into a character\-feature sequencecc, which \(i\) drives a duration head that predicts the total latent lengthT^\\hat\{T\}and \(ii\) serves as the key/value memory for the cross\-attention layers of a non\-autoregressive Diffusion Transformer \(DiT\)\. Conditioned on the flow\-matching timestepttthrough adaLN\-zero modulation, the DiT denoises a length\-T^\\hat\{T\}sequence of6464\-dimensional latent frames drawn from Gaussian noise; the resulting clean latents are rendered to4848kHz audio by the*frozen*AudioVAE2 decoder\. The autoregressive TSLM/FSQ/RALM backbone is discarded; only AudioVAE2 is reused\.This choice is dictated by reliability\. Conversational Turkish is dense with digit strings, dates, and amounts, and autoregressive latent prediction can accumulate error across such sequences, garbling numbers mid\-utterance\. Predicting the whole latent sequence in parallel removes the left\-to\-right dependency; it also decouples the number of*sequential*generation steps from utterance length \(a fixed ODE\-step budget replaces a length\-proportional decode loop, though each step still attends over allTTframes\) and pairs naturally with flow matching\(Leet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib21); Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20); Anastassiouet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib51)\), in contrast to the multi\-stage LM\-plus\-vocoder pipelines of discrete\-token systems\(Duet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib24)\)\.
#### Formulation\.
Letx1∈ℝT×64x\_\{1\}\\in\\mathbb\{R\}^\{T\\times 64\}be the clean latent sequence produced by the frozen AudioVAE2 encoder andc=TextEnc\(y\)c=\\mathrm\{TextEnc\}\(y\)the character\-feature memory of the input textyy\. We adopt a conditional flow\-matching objective with a linear probability path, which interpolates on a straight line between a Gaussian priorx0∼𝒩\(0,I\)x\_\{0\}\\sim\\mathcal\{N\}\(0,I\)and the datax1x\_\{1\}and trains a velocity fieldvθv\_\{\\theta\}to transport one to the other:
xt\\displaystyle x\_\{t\}=\(1−t\)x0\+tx1,x0∼𝒩\(0,I\),t∼𝒰\[0,1\],\\displaystyle=\(1\-t\)\\,x\_\{0\}\+t\\,x\_\{1\},\\qquad x\_\{0\}\\sim\\mathcal\{N\}\(0,I\),\\quad t\\sim\\mathcal\{U\}\[0,1\],\(1\)ℒ\\displaystyle\\mathcal\{L\}=𝔼t,x0,x1\[‖m⊙\(vθ\(xt,t,c\)−\(x1−x0\)⏞v⋆\)‖22\]⏟masked flow\-matching\+λdur\(logT^−logT\)2\.\\displaystyle=\\underbrace\{\\mathbb\{E\}\_\{t,\\,x\_\{0\},\\,x\_\{1\}\}\\Big\[\\big\\\|\\,m\\odot\\big\(v\_\{\\theta\}\(x\_\{t\},t,c\)\-\\overbrace\{\(x\_\{1\}\-x\_\{0\}\)\}^\{v^\{\\star\}\}\\big\)\\big\\\|\_\{2\}^\{2\}\\Big\]\}\_\{\\text\{masked flow\-matching\}\}\\;\+\\;\\lambda\_\{\\mathrm\{dur\}\}\\,\\big\(\\log\\hat\{T\}\-\\log T\\big\)^\{2\}\.Herev⋆=x1−x0v^\{\\star\}=x\_\{1\}\-x\_\{0\}is the target velocity, which is*constant*along the straight path so that a single network evaluation supervises the entire trajectory;m∈\{0,1\}Tm\\in\\\{0,1\\\}^\{T\}masks padding frames, and the flow\-matching term is averaged over the‖m‖1\\\|m\\\|\_\{1\}real frames;T^\\hat\{T\}is the length predicted by the duration head \([section˜3\.3](https://arxiv.org/html/2607.09530#S3.SS3)\); andλdur=0\.1\\lambda\_\{\\mathrm\{dur\}\}=0\.1\. Becausex0x\_\{0\}andx1x\_\{1\}are sampled independently rather than optimally coupled, this is the linear\-path,*rectified\-flow*instance of conditional flow matching\(Lipmanet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib119); Tonget al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib127)\), not a minibatch optimal\-transport coupling\. There is no adversarial loss, no autoregression, and no discrete\-token cross\-entropy: only the masked flow\-matching regression and the auxiliary duration term\. Because the model consumes raw characters and no discrete speech tokenizer intervenes, the pronunciation of in\-context numbers, currencies, and acronyms is learned end\-to\-end from audio rather than delegated to a text frontend; isolated tokens are harder and are addressed by post\-training and inference\-time guards \([section˜3\.5](https://arxiv.org/html/2607.09530#S3.SS5)\)\.
### 3\.2Frozen AudioVAE2 Latent Space
We build on AudioVAE2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\), reused*frozen*under its Apache\-2\.0 license and never trained\. Its encoder maps a1616kHz waveform to6464\-dimensional latent frames at2525Hz \(one frame per4040ms\) and its decoder reconstructs at a4848kHz sample rate\. We stress that this1616kHz\-in /4848kHz\-out asymmetry fixes the output*sample rate*, not the acoustic bandwidth: the usable bandwidth of a synthesis is still bounded by the encoder’s1616kHz input and by the band of the training audio, so the narrowband single\-speaker corpus of[section˜3\.5](https://arxiv.org/html/2607.09530#S3.SS5)yields a4848kHz\-rate signal that retains a telephony\-band fidelity ceiling \([section˜5](https://arxiv.org/html/2607.09530#S5)\)\. Holding this codec fixed is what lets us train a compact generator from scratch: high\-fidelity reconstruction is already solved, so all capacity goes to the text\-conditional prior, and modeling the*continuous*latent rather than discrete tokens keeps fine acoustic detail and leaves the pipeline tokenizer\-free\(Shenet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib16); Leet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib21)\)\. We retain only this continuous codec; the semi\-discrete FSQ bottleneck\(Mentzeret al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib47)\)that gives VoxCPM its semantic\-acoustic factorization lives in the autoregressive backbone we discard, and our NAR DiT operates directly on the native2525Hz frame grid with no patch grouping, so the sequence lengthTTis simply the number of4040ms frames\.
### 3\.3The FreyaTTS Diffusion Transformer
The velocity fieldvθv\_\{\\theta\}is a DiT that takes the noisy latent sequencext∈ℝT×64x\_\{t\}\\in\\mathbb\{R\}^\{T\\times 64\}, the timesteptt, and the character memorycc, and predicts a velocity at every frame\.[Table˜1](https://arxiv.org/html/2607.09530#S3.T1)lists the configuration\.
#### Text encoder and duration head\.
The text encoder embeds the character\-level Turkish vocabulary \(9292symbols in total: the2929\-letter alphabet including ç, ğ, ı, ö, ş, ü, together with digits, punctuation, and whitespace\) and refines it with four ConvNeXt\-1d blocks\(Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20); Liuet al\.,[2022](https://arxiv.org/html/2607.09530#bib.bib121)\)into the feature sequencecc, shared by the cross\-attention memory and the duration head\. Operating on raw characters means that Turkish orthographic phenomena \(vowel harmony, the dotted/dotless ı distinction, and the spoken form of acronyms\) are absorbed into the weights rather than handed to a grapheme\-to\-phoneme frontend\. Digit strings are the deliberate exception: their orthographic and phonetic lengths diverge so sharply that we expand them to spoken form at the text frontend, since the bottleneck lies in the duration predictor rather than in the acoustic model; isolated acronyms and bare tokens remain a documented failure mode addressed in[sections˜3\.5](https://arxiv.org/html/2607.09530#S3.SS5)and[4](https://arxiv.org/html/2607.09530#S4)\. Since generation is non\-autoregressive, the target length must be known before denoising: a small MLP over the mask\-averaged mean ofccregresseslogT^\\log\\hat\{T\}, supervised by the auxiliary term of[eq\.˜1](https://arxiv.org/html/2607.09530#S3.E1)and used to size the noise tensor at inference\.
#### DiT block\.
Each of the1616layers updates the frame hidden statesh∈ℝT×dh\\in\\mathbb\{R\}^\{T\\times d\}\(d=640d\{=\}640\) through three residual sub\-layers \(RoPE self\-attention, text cross\-attention, and a SwiGLU feed\-forward network\), each modulated by the flow\-matching timestep through adaLN\-zero:
\(β1,γ1,g1,β2,γ2,g2,β3,γ3,g3\)\\displaystyle\(\\beta\_\{1\},\\gamma\_\{1\},g\_\{1\},\\;\\beta\_\{2\},\\gamma\_\{2\},g\_\{2\},\\;\\beta\_\{3\},\\gamma\_\{3\},g\_\{3\}\)=MLPzero\(emb\(t\)\),\\displaystyle=\\mathrm\{MLP\}\_\{\\mathrm\{zero\}\}\\big\(\\mathrm\{emb\}\(t\)\\big\),\(2\)h\\displaystyle h←h\+g1⊙SelfAttnRoPE\(γ1⊙LN\(h\)\+β1\),\\displaystyle\\leftarrow h\+g\_\{1\}\\odot\\mathrm\{SelfAttn\}\_\{\\mathrm\{RoPE\}\}\\big\(\\gamma\_\{1\}\\odot\\mathrm\{LN\}\(h\)\+\\beta\_\{1\}\\big\),h\\displaystyle h←h\+g2⊙CrossAttn\(γ2⊙LN\(h\)\+β2,c\),\\displaystyle\\leftarrow h\+g\_\{2\}\\odot\\mathrm\{CrossAttn\}\\big\(\\gamma\_\{2\}\\odot\\mathrm\{LN\}\(h\)\+\\beta\_\{2\},\\;c\\big\),h\\displaystyle h←h\+g3⊙SwiGLU\(γ3⊙LN\(h\)\+β3\),\\displaystyle\\leftarrow h\+g\_\{3\}\\odot\\mathrm\{SwiGLU\}\\big\(\\gamma\_\{3\}\\odot\\mathrm\{LN\}\(h\)\+\\beta\_\{3\}\\big\),whereemb\(t\)\\mathrm\{emb\}\(t\)is a sinusoidal timestep embedding and⊙\\odotapplies per\-channel shift/scale/gate broadcast over frames\. The nine modulation vectors are produced by a per\-layer MLP whose final projection is*zero\-initialized*\(adaLN\-zero\): every block therefore begins as the identity and departs from it only as training warrants, a stabilizer standard in diffusion transformers\(Peebles and Xie,[2023](https://arxiv.org/html/2607.09530#bib.bib120); Liet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib36); Jiaet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib37)\)\. Self\-attention carries rotary position embeddings\(Suet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib105)\)over the frame axis; in contrast to the NoPE\(Kazemnejadet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib69)\)residual LM of VoxCPM2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\), we retain RoPE because our sequence is generated in one shot and benefits from explicit positional structure\. InCrossAttn\\mathrm\{CrossAttn\}the queries are the modulated frame states while the keys and values are the character featurescc: this dedicated cross\-attention is the sole pathway through which text conditions acoustics, following the alignment\-based diffusion decoders of Grad\-TTS, Matcha\-TTS, and NaturalSpeech 2\(Popovet al\.,[2021](https://arxiv.org/html/2607.09530#bib.bib123); Mehtaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib122); Shenet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib16)\)\.
Table 1:FreyaTTS configuration\. AudioVAE2 is reused frozen; only the DiT and text encoder \(183\.2183\.2M parameters\) are trained\.ComponentSettingLatent space \(frozen AudioVAE2\)6464\-dim @2525Hz;1616kHz encode /4848kHz decodeText vocabularycharacter\-level Turkish,9292symbols \(no BPE, no G2P\)Text encoderConvNeXt\-1d×4\\times\\,4DiT width / depth / headsd=640d\{=\}640/1616/1010DiT feed\-forwardSwiGLU, hidden20482048Self\-attention position encodingRoPETime conditioningadaLN\-zero \(99\-way modulation\)Text conditioningcross\-attention \(frame queries→\\tocharacter keys/values\)Duration objective weightλdur\\lambda\_\{\\mathrm\{dur\}\}0\.10\.1\(pretraining\)Trainable parameters183\.2183\.2MInference solver3232\-step Euler ODE
### 3\.4Pretraining from Scratch
The DiT and text encoder are trained from random initialization \(AudioVAE2 contributes no gradients\) on a large\-scale, high\-quality internal corpus of multi\-speaker Turkish speech, pre\-encoded offline into AudioVAE2 latents so that the frozen encoder never runs during training, and spanning number\-string reads, name\-list reads, conversational sentences, and general speech\. We optimize[eq\.˜1](https://arxiv.org/html/2607.09530#S3.E1)in bf16 with AdamW \(learning rate5×10−45\\\!\\times\\\!10^\{\-4\}, weight decay0\.010\.01, gradient clipping at1\.01\.0,2,0002\{,\}000\-step linear warmup then cosine decay\) for150150k steps at batch size6464, in a single training run with a fixed seed\. Training runs on H100 and H200 GPUs with a memory\-safe loader over the pre\-encoded fp16 latents\.
#### Pretraining behavior\.
The flow\-matching loss descends smoothly to convergence over the150150k steps \([fig\.˜2\(a\)](https://arxiv.org/html/2607.09530#S3.F2.sf1)\), and re\-transcribing generations with Whisper\-large\-v3\(Radfordet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib111)\)yields intelligible, well\-aligned Turkish, with in\-context numbers read in their full spoken form\. The pretrained model has, however,*no speaker identity*: with no speaker field in the corpus and no speaker conditioning in the network, the Gaussian priorx0x\_\{0\}seeds a fresh speaker on each generation\. This is expected behavior for an unconditioned prior, and it directly motivates the voice\-lock post\-training below;[section˜4\.3](https://arxiv.org/html/2607.09530#S4.SS3)quantifies the effect\.
### 3\.5Post\-training
Two supervised fine\-tuning stages turn the speaker\-agnostic pretrained model into a deployable single\-voice product\.
#### SFT v1: single\-speaker voice lock\.
Initialized from the150150k\-step checkpoint, we full\-parameter fine\-tune \(AudioVAE2 still frozen\) on an internal single\-speaker corpus of a single consented professional voice talent, recorded at1616kHz over a narrowband channel matching the deployment domain\. Training uses multi\-GPU DDP, bf16, and learning rate1×10−41\\\!\\times\\\!10^\{\-4\}with cosine decay; the flow\-matching loss falls from0\.9180\.918to≈\\approx0\.620\.62within roughly one epoch\. The speaker identity is written into the weights: cross\-generation F0 standard deviation collapses from74\.974\.9to5\.05\.0Hz and content similarity reaches0\.9230\.923, while domain acronyms stabilize, pronounced as words rather than spelled letter by letter\. We selectstep 1000: continued training over\-fits and destabilizes the voice, with the voice\-lock standard deviation drifting from5\.05\.0back to1111Hz by33k steps\.
#### SFT v2: short\-utterance coverage\.
The single\-speaker corpus contains*no*clips of two words or fewer, which we identified as the structural root cause of a collapse on isolated acknowledgments, bare numbers, and lone acronyms\. We therefore continue fine\-tuning from the v1 step\-10001000checkpoint on the same corpus augmented with mined short single\-speaker segments \(one\-word and two\-word phrases, numbers, and acknowledgments\) that we extracted from the same speaker by MMS forced alignment \(torchaudioMMS\_FAwith a Turkish\-to\-roman character map\)\. This continuation stage uses learning rate5×10−55\\\!\\times\\\!10^\{\-5\}andλdur=0\.2\\lambda\_\{\\mathrm\{dur\}\}\{=\}0\.2, and we again select step 1000\. Isolated acknowledgments are rendered once, cleanly, rather than entering a repetition loop; in\-context numbers are correct; and full\-sentence quality and the locked voice are preserved\.SFT v2 is the shipped production model\.A thin inference wrapper completes the system: clause\-level chunking for long inputs, a duration floor for very short ones, and a voicing\-based retry on the rare degenerate draw\.
### 3\.6Inference
Given input text, the text encoder producesccand the duration head predicts the latent lengthT^\\hat\{T\}\. We drawx0∼𝒩\(0,I\)x\_\{0\}\\sim\\mathcal\{N\}\(0,I\)of lengthT^\\hat\{T\}and integrate the flow\-matching ODEdx/dt=vθ\(xt,t,c\)\\mathrm\{d\}x/\\mathrm\{d\}t=v\_\{\\theta\}\(x\_\{t\},t,c\)fromt=0t\{=\}0tot=1t\{=\}1with a fixed\-step Euler solver \(3232steps\), evaluating the DiT once per step withccsupplied through cross\-attention\. The resulting clean latent sequence is decoded by the frozen AudioVAE2 into a4848kHz waveform\. Because generation is non\-autoregressive, the whole sequence is denoised in parallel: the number of*sequential*solver steps is fixed at3232regardless of utterance length \(each step still attending over allTTframes\), and there is no left\-to\-right error accumulation\. Unlike the autoregressive sampler of VoxCPM2, which leans on classifier\-free guidance, sway sampling\(Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20)\), and CFG\-Zero∗\(Fanet al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib92)\), we find a plain3232\-step deterministic Euler integrator sufficient; the inference wrapper re\-drawsx0x\_\{0\}only on the rare voicing\-check failure\. Empirically the shipped FreyaTTS model attains a mean per\-utterance real\-time factor of≈\\approx0\.140\.14on an H100, which we detail in[section˜4\.4](https://arxiv.org/html/2607.09530#S4.SS4)\.
\(a\)Training loss\.
\(b\)Voice\-lock \(F0F\_\{0\}\)\.
\(c\)WER vs\. length\.
Figure 2:Left:flow\-matching loss over pretraining \(log\-scale steps\) with the single\-speaker SFT plateau\.Center:medianF0F\_\{0\}across regenerations of a fixed prompt; the multi\-speaker pretrained model scatters across an octave \(std74\.974\.9Hz\) whereas the voice\-locked SFT model is stable \(std5\.05\.0Hz\)\.Right:word error rate as a function of input length on a length\-swept probe set; long\-horizon drift in the pretrained model is mitigated, though not removed, by inference\-time clause chunking\.
## 4Experiments and Results
We evaluate FreyaTTS along three axes: main results against the field of openly\-available, sub\-billion\-parameter Turkish TTS systems on a single held\-out benchmark \([sections˜4\.1](https://arxiv.org/html/2607.09530#S4.SS1)and[4\.2](https://arxiv.org/html/2607.09530#S4.SS2)\); speaker consistency and the voice\-lock stage \([section˜4\.3](https://arxiv.org/html/2607.09530#S4.SS3)\); and inference efficiency \([section˜4\.4](https://arxiv.org/html/2607.09530#S4.SS4)\)\.
### 4\.1Experimental Setup
#### Benchmark\.
We evaluate on a single, purpose\-built benchmark,Freya\-TR\-Eval:495495natural, domain\-neutral Turkish sentences of everyday conversational register \(statements,*mI*\-questions, and exclamations\), length33–1313words\. It merges native Turkish sentence texts sampled from Common Voice 17\-tr\(Ardilaet al\.,[2020](https://arxiv.org/html/2607.09530#bib.bib113)\)and CoVoST2\-tr\(Wanget al\.,[2021](https://arxiv.org/html/2607.09530#bib.bib114)\)with everyday\-conversational lines generated bygemini\-3\.1\-pro\-previewunder fixed topic and sentence\-type quotas \(2222everyday topics\), filtered for length, full Turkish grapheme coverage \(ç ğ ı i ö ş ü\), and de\-duplication\. FreyaTTS is trained only on internal corpora \([sections˜3\.4](https://arxiv.org/html/2607.09530#S3.SS4)and[3\.5](https://arxiv.org/html/2607.09530#S3.SS5)\) disjoint from the benchmark sources, so every item probes generalization; the exact sentence list and the seeded build scripts are released\.111Released asfreyavoice/freya\-tr\-evalon HuggingFace\.Because FreyaTTS is a single\-voice synthesizer rather than a zero\-shot cloner, we synthesize the reference text of each item and score the output; the two reference\-prompt baselines \(XTTS\-v2, F5\-TTS\) are given one fixed shared reference clip \(cloner scores can be sensitive to this choice\)\.
#### Comparison systems\.
We position FreyaTTS, the shipped183\.2183\.2M single\-voice model, against the field of*openly available, sub\-billion\-parameter*Turkish TTS systems, the models a practitioner can download and self\-host:MMS\-TTS\-tr\(Pratapet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib115)\)\(3636M, Meta’s VITS\-based multilingual TTS\),Coqui GlowTTS\-tr\(∼\\sim2828M, flow\-based\),Piper\-tr\(∼\\sim1616M, VITS, on\-device\),SpeechT5\-tr\(Aoet al\.,[2022](https://arxiv.org/html/2607.09530#bib.bib128)\)\(∼\\sim144144M, fine\-tuned on Turkish Common Voice\),XTTS\-v2\(Casanovaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib13)\)\(∼\\sim470470M, multilingual zero\-shot cloner\), andF5\-TTS\-Turkish\(Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20)\)\(∼\\sim336336M, a flow\-matching zero\-shot cloner fine\-tuned for Turkish\)\.
#### Metrics\.
The primary metrics areband\-matched WER and CER: every system’s output is downsampled to88kHz before transcription with Whisper\-large\-v3\(Radfordet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib111)\), so that FreyaTTS’s telephony\-band4848kHz output \([section˜3\.2](https://arxiv.org/html/2607.09530#S3.SS2)\) is not penalized relative to natively wideband baselines, and reference and transcript pass the same Turkish text normalization before scoring; the protocol is applied identically to every system\. We also reportreal\-time factor\(RTF\): per\-utterance synthesis time over audio duration at batch size11, including AudioVAE2 decoding, averaged over the set, on an H100\. Naturalness is measured with amean opinion score\(MOS\) listening study:2424native Turkish raters scored randomized, system\-blind samples of every system on a55\-point scale \(25662566ratings in total,360360–375375per system\); we report per\-system means with95%95\\%confidence intervals\.Voice consistency\([section˜4\.3](https://arxiv.org/html/2607.09530#S4.SS3)\) is theF0F\_\{0\}standard deviation across regenerations of the same text\.Content similarityis11minus the normalized character\-level edit distance between the input text and the Whisper\-large\-v3 transcript of the synthesis\.
### 4\.2Main Results
Table 2:Conversational Turkish TTS, sub\-1B open models\.All systems evaluated on the full Freya\-TR\-Eval set \(495495everyday, general\-purpose Turkish sentences\) under an identical protocol:WER/CERfrom Whisper\-large\-v3 with88kHz band\-matching \(in %\), and naturalnessMOSfrom a listening study with2424native Turkish raters \(25662566ratings,360360–375375per system,55\-point scale,95%95\\%confidence intervals\)\. Best per column inbold; systems ordered by size\.SystemParamsWER↓\\downarrowCER↓\\downarrowMOS↑\\uparrowPiper \(tr, dfki\)16M4\.41\.13\.47±0\.223\.47\\pm 0\.22Coqui GlowTTS \(tr\)28M12\.13\.32\.53±0\.192\.53\\pm 0\.19MMS\-TTS \(tr\)36M6\.81\.73\.58±0\.203\.58\\pm 0\.20SpeechT5 \(tr\)144M83\.445\.51\.59±0\.141\.59\\pm 0\.14FreyaTTS \(ours\)183\.2M8\.03\.03\.68±0\.223\.68\\pm 0\.22F5\-TTS \(tr\)336M24\.310\.93\.63±0\.233\.63\\pm 0\.23XTTS\-v2 \(multi\)470M11\.13\.93\.82±0\.19\\textbf\{3\.82\}\\pm 0\.19#### FreyaTTS outperforms the larger open systems in its field\.
[Table˜2](https://arxiv.org/html/2607.09530#S4.T2)reports band\-matched Whisper WER/CER for every system on the same495495\{\}conversational sentences\. FreyaTTS reaches WER 8\.0 % and CER 3\.0 %,outperforming both larger open systems in the field, the zero\-shot cloners XTTS\-v2 \(∼\\sim470470M, WER 11\.1 %\) and F5\-TTS \(∼\\sim336336M, WER 24\.3 %\), each1\.81\.8–2\.6×2\.6\\timesits parameter count, despite their larger backbones and far greater training data: cloning a target voice from a short reference generalizes poorly to arbitrary sentences, whereas FreyaTTS carries a single voice natively in its weights\. It also comfortably beats the weaker same\-class models \(GlowTTS, SpeechT5\)\. Piper \(∼\\sim1616M, WER 4\.4 %\) and MMS\-TTS \(∼\\sim3636M, WER 6\.8 %\) are compact phonemizer\-driven VITS baselines whose highly regular pronunciation maximizes recognizer agreement; FreyaTTS supplies the single, consistent*learned*voice they lack \([section˜4\.3](https://arxiv.org/html/2607.09530#S4.SS3)\)\. The listening study bears this out on naturalness: FreyaTTS reaches MOS3\.683\.68, above every same\-class system including the two low\-WER VITS baselines \(Piper3\.473\.47, MMS\-TTS3\.583\.58\) and F5\-TTS \(3\.633\.63\), and second overall only to the2\.6×2\.6\\times\-larger XTTS\-v2 \(3\.823\.82\); the95%95\\%intervals of the top four systems overlap, so we read the MOS column as a ranking rather than pairwise significance\. That a183183M model trained from scratch outperforms every larger open system in this field is the paper’s central empirical claim\.
#### Statistical robustness, calibration, and the inference layer\.
Three checks contextualize the headline numbers\. First, a sentence\-level bootstrap \(10,00010\{,\}000resamples over the495495items, computed on an independent re\-synthesis of the full set\) gives WER7\.9%7\.9\\%with a95%95\\%confidence interval of\[6\.8,9\.1\]\[6\.8,9\.1\]and CER2\.9%2\.9\\%\[2\.4,3\.4\]\[2\.4,3\.4\], reproducing[table˜2](https://arxiv.org/html/2607.09530#S4.T2)within run\-to\-run noise\. Second, two anchors calibrate the protocol’s absolute scale: real human recordings \(FLEURS\-tr test\(Conneauet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib112)\)\) transcribe at9\.7%9\.7\\%WER under the identical band\-matched pipeline, and re\-encoding held\-out*real*recordings of the target speaker through the frozen AudioVAE alone costs2\.6%2\.6\\%WER – the transcription floor of the latent space itself\. FreyaTTS’s8\.0%8\.0\\%is therefore comparable in magnitude to the recognizer’s error on real human speech, though the two are measured on different corpora and Whisper WER is sensitive to text distribution, and sits within5\.45\.4points of the latent space’s floor\. Third, ablating the entire inference layer of[section˜3\.6](https://arxiv.org/html/2607.09530#S3.SS6)\(normalization, clause chunking, duration floor, voicing retry\) moves WER by at most0\.30\.3points \(paired bootstrap95%95\\%CI of the difference\[−0\.3,0\.0\]\[\-0\.3,0\.0\]\); the voicing retry fires on11of495495sentences and clause chunking on77of495495\. The guards are tail insurance against rare failure modes, not a crutch the headline numbers depend on\. Error does climb with input length on a length\-swept probe \([fig\.˜2\(c\)](https://arxiv.org/html/2607.09530#S3.F2.sf3)\): long\-horizon drift on long utterances is mitigated, though not eliminated, by this clause chunking\.
### 4\.3Speaker Consistency and the Voice\-Lock
Because FreyaTTS carries no speaker field and no reference\-prompt pathway, the pretrained prior has no notion of speaker identity: the Gaussian priorx0x\_\{0\}seeds a fresh, uncontrolled speaker on every generation\.[Figure˜2\(b\)](https://arxiv.org/html/2607.09530#S3.F2.sf2)makes this concrete\. Regenerating a fixed text many times \(a114114\-generation voice\-consistency probe\), the pretrained model’s per\-generation medianF0F\_\{0\}spans roughly7070–347347Hz, with a standard deviation of74\.974\.9Hz and a60%60\\%rate of crossing the165165Hz male/female pitch boundary between successive draws\. Intelligibility metrics are blind to this \(the content is correct on each draw\), yet for a deployed single\-voice product it is a categorical failure: the listener hears a different person each turn\.
The voice\-lock stage writes a single identity into the weights\. Full\-parameter fine\-tuning on the consented target speaker \(SFT\-v1;[section˜3\.5](https://arxiv.org/html/2607.09530#S3.SS5)\) collapses the cross\-generationF0F\_\{0\}standard deviation from74\.974\.9Hz to5\.05\.0Hz and effectively eliminates the gender flips, while the content\-similarity metric of[section˜4\.1](https://arxiv.org/html/2607.09530#S4.SS1)reaches0\.9230\.923: the model now emits a deterministic, consistent target voice regardless of the noise draw \([fig\.˜2\(b\)](https://arxiv.org/html/2607.09530#S3.F2.sf2)\), measured on sentence\-length in\-distribution text\. A speaker\-verification embedding confirms the lock at the identity level: the ECAPA\-TDNN\(Desplanqueset al\.,[2020](https://arxiv.org/html/2607.09530#bib.bib117)\)cosine similarity between FreyaTTS syntheses and the target\-speaker centroid is0\.873±0\.0290\.873\\pm 0\.029, within one standard deviation of the0\.883±0\.0260\.883\\pm 0\.026same\-speaker ceiling measured on disjoint*real*recordings of the speaker, and far above the0\.1400\.140different\-speaker floor\. The identity is a property of the parameters, not of a prompt: no reference audio is supplied at inference\. We select the step\-10001000checkpoint because prolonged fine\-tuning over\-specializes and destabilizes the voice, with the consistency standard deviation drifting back from5\.05\.0Hz toward1111Hz by step33k\. SFT also stabilizes pronunciation: all\-caps acronyms are spoken natively as words rather than spelled out, and isolated short utterances, covered by SFT\-v2’s mined one\-word and two\-word segments \([section˜3\.5](https://arxiv.org/html/2607.09530#S3.SS5)\), are rendered cleanly\. This pretrain\-to\-SFT contrast, an unconditioned multi\-speaker prior turned into a locked single voice, is the mechanism by which the foundation\-style generative prior becomes a deployable single\-voice product\.
### 4\.4Inference Efficiency
FreyaTTS generates4848kHz audio non\-autoregressively: given the duration head’s predicted lengthT^\\hat\{T\}, a single fixed\-step Euler integrator runs3232ODE steps from noise to a clean latent sequence, which the frozen AudioVAE decodes in one pass \([section˜3\.6](https://arxiv.org/html/2607.09530#S3.SS6)\)\. The number of*sequential*solver steps is fixed at3232regardless ofT^\\hat\{T\}, there is no left\-to\-right error accumulation, and a plain deterministic integrator suffices, without classifier\-free guidance or sway sampling\. At183\.2183\.2M parameters the model fits in1\.51\.5GB of VRAM\.
Table 3:Inference efficiency on a single RTX 4090\(fp32/bf16 as released, batch size11, end\-to\-end including vocoder/VAE decoding; median of1010runs after33warmups on a fixed Turkish prompt set of three length buckets \(five sentences each; released with our code\)\)\.TTFT= time to first audio on the long bucket \(2222–2828words\): native streaming first\-chunk for XTTS\-v2, first synthesized clause for FreyaTTS, full\-utterance latency for the offline systems\.RTF= synthesis time over audio duration\.Throughput= aggregate audio\-seconds per wall\-second over a4545\-utterance queue at the best concurrency level \(CCengines in separate processes\)\. Best per column inbold\.†Serving\-style rows: VoxCPM2 through the nanovllm continuous\-batching engine \(chunk streaming,1616concurrent sequences, GPU\-memory utilization0\.50\.5\), and FreyaTTS through a static\-batching ODE server \(padded batch, one masked3232\-step solve; TTFT column reports its per\-batch latency atB=8B\{=\}8\)\. All other rows are the systems’ released library stacks\.SystemParamsVRAM \(GB\)TTFTlong\{\}\_\{\\text\{long\}\}\(s\)RTFmed\{\}\_\{\\text\{med\}\}RTFlong\{\}\_\{\\text\{long\}\}ThroughputFreyaTTS \(ours\)183M1\.50\.520\.110\.109\.4 \(C=4C\{=\}4\)F5\-TTS \(tr\)∼\\sim336M0\.80\.740\.130\.058\.4 \(C=2C\{=\}2\)XTTS\-v2 \(multi\)∼\\sim470M2\.20\.300\.330\.333\.4 \(C=2C\{=\}2\)Spark\-TTS∼\\sim0\.5B5\.414\.01\.081\.041\.4 \(C=2C\{=\}2\)VoxCPM2∼\\sim2B5\.54\.20\.370\.371\.7 \(C=1C\{=\}1\)VoxCPM2 \(serving engine\)†∼\\sim2B11\.40\.140\.150\.1324\.8 \(C=16C\{=\}16\)FreyaTTS \(batched ODE\)†183M3\.80\.54––65\.2\(B=8B\{=\}8\)[Table˜3](https://arxiv.org/html/2607.09530#S4.T3)measures FreyaTTS against four larger open systems, including the∼\\sim22B VoxCPM2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\)whose AudioVAE2 we reuse, the autoregressive Spark\-TTS\(Wanget al\.,[2025](https://arxiv.org/html/2607.09530#bib.bib55)\), XTTS\-v2\(Casanovaet al\.,[2024](https://arxiv.org/html/2607.09530#bib.bib13)\), and F5\-TTS\(Chenet al\.,[2025b](https://arxiv.org/html/2607.09530#bib.bib20)\), each run end\-to\-end through its released inference stack with default sampling on an RTX 4090\. Among the released library stacks, FreyaTTS attains the best medium\-sentence RTF \(0\.110\.11\) and the highest throughput \(9\.49\.4audio\-seconds per wall\-second with four concurrent engines\) at the second\-smallest memory footprint\. Against the production\-class VoxCPM2 it is3\.4×3\.4\\timesfaster in medium\-bucket RTF,8×8\\timesfaster to first audio on long inputs,3\.7×3\.7\\timessmaller in VRAM, and5\.5×5\.5\\timeshigher in throughput; against the autoregressive Spark\-TTS the gaps widen to roughly10×10\\timesin RTF and27×27\\timesin TTFT\. The optimized F5\-TTS runtime is the one system in the same speed class: its fused bf16 kernel path reaches a lower long\-form RTF, while FreyaTTS, running unoptimized fp32 eager PyTorch, retains lower TTFT, higher throughput, and, per[table˜2](https://arxiv.org/html/2607.09530#S4.T2), a three\-times\-lower Turkish WER with no reference clip at inference\. XTTS\-v2’s native chunk streaming yields the best library\-stack TTFT \(0\.300\.30s\) at roughly3×3\\timesthe RTF\. We additionally measure VoxCPM2 through a dedicated continuous\-batching serving engine \(nanovllm\) rather than its released library stack: chunk streaming brings its TTFT to0\.140\.14s and batching lifts throughput to24\.824\.8audio\-seconds per wall\-second at1616concurrent sequences, at11\.411\.4GB of reserved VRAM\. A serving engine is thus worth roughly an order of magnitude on both axes for the22B model\. The same treatment favors the non\-autoregressive design even more: a static\-batching ODE server for FreyaTTS \(pad the batch, run one masked3232\-step solve, decode, trim; released with our inference tooling\) reaches65\.265\.2audio\-seconds per wall\-second at batch size88with a0\.540\.54s per\-batch latency and3\.83\.8GB of VRAM,2\.6×2\.6\\timesthe22B engine’s continuous\-batching throughput at one\-third of its memory, before any length\-bucketing or chunk\-streaming refinements\. On an H100, FreyaTTS’s benchmark\-set mean RTF is≈\\approx0\.140\.14\.
The compact size also admits deployment classes the systems of[table˜3](https://arxiv.org/html/2607.09530#S4.T3)do not target\. On a laptop CPU \(Apple M3, four threads, fp32, no quantization\) FreyaTTS synthesizes at or faster than real time: RTF0\.700\.70on medium\-length sentences and0\.940\.94–1\.001\.00on the short and long buckets, with a5\.75\.7s first\-clause TTFT on long inputs\. Compiled to Core ML at fp16, the DiT’s full3232\-step loop renders4\.44\.4s of audio in0\.2050\.205s on the laptop’s neural engine \(RTF0\.0470\.047\) and0\.720\.72s on CPU only \(RTF0\.1650\.165\), with AudioVAE decoding adding0\.330\.33s on CPU, for an end\-to\-end real\-time factor of≈\\approx0\.120\.12on consumer Apple silicon\. A model competitive with the fastest GPU systems of its field while synthesizing faster than real time on a laptop CPU is an operating point that pairs the foundation\-style generative prior with genuinely edge\-deployable inference\.
## 5Conclusion and Future Work
In this work, we presented FreyaTTS, a tokenizer\-free, non\-autoregressive flow\-matching model for Turkish conversational speech synthesis\. Built as a compact183\.2183\.2M\-parameter conditional flow\-matching Diffusion Transformer inside the*frozen*2525Hz continuous\-latent space of the AudioVAE2\(Zhouet al\.,[2026](https://arxiv.org/html/2607.09530#bib.bib129)\), and driven by a9292\-symbol character vocabulary with no phonemizer, grapheme\-to\-phoneme frontend, or discrete speech tokenizer, FreyaTTS is trained from scratch on a large\-scale Turkish speech corpus and denoises an entire utterance in parallel in a fixed3232\-step ODE\. A two\-stage pretrain\-to\-SFT recipe, a single\-speaker voice lock \(F0F\_\{0\}standard deviation74\.9→5\.074\.9\\\!\\rightarrow\\\!5\.0Hz\) followed by short\-utterance coverage, converts the speaker\-agnostic prior into a reliability\-hardened single\-voice production model\. Evaluated under a band\-matched Whisper WER/CER protocol on the releasedFreya\-TR\-Evalbenchmark, FreyaTTS \(WER 8\.0 %, CER 3\.0 %\) outperforms both larger open systems, XTTS\-v2 and F5\-TTS, at a real\-time factor of≈\\approx0\.140\.14and roughly4040–55%55\\%of their parameter count\.
While FreyaTTS delivers strong results for its size, several challenges remain\. Two compact phonemizer\-driven VITS baselines still attain lower error on clean conversational text; digit\-dense input requires spoken\-form expansion at the text frontend, because the character\-level duration predictor under\-allocates frames for compact digit strings; and the shipped voice inherits a narrowband fidelity ceiling from its telephony\-band source\. Future work will focus on a jointly learned CTC\-monotonic aligner, already prototyped, that targets the residual word\-skip and long\-horizon\-drift failures a single global length prediction cannot prevent; preference\-based post\-training\(Rafailovet al\.,[2023](https://arxiv.org/html/2607.09530#bib.bib118)\)from intelligibility and voicing rewards; extending the voice\-lock recipe from one speaker to multiple voices and natural\-language voice design; and explicit bandwidth extension so that a narrowband\-trained speaker can be rendered in true wideband\. We hope the open release of FreyaTTS, the model weights, the Freya\-TR\-Eval benchmark, and the inference tooling provides a solid foundation for Turkish and low\-resource speech research\.
## Contributors
Ahmet Erdem Pamuk, Ömer Yentür, Ahmet Tunga Bayrak, Yavuz Alp Sencer Öztürk, Mustafa Yavuz\.
## References
- P\. Anastassiou, J\. Chen, J\. Chen, Y\. Chen, Z\. Chen, Z\. Chen, J\. Cong, L\. Deng, C\. Ding, L\. Gao,et al\.\(2024\)Seed\-tts: a family of high\-quality versatile speech generation models\.arXiv preprint arXiv:2406\.02430\.Cited by:[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p2.1)\.
- J\. Ao, R\. Wang, L\. Zhou, C\. Wang, S\. Ren, Y\. Wu, S\. Liu, T\. Ko, Q\. Li, Y\. Zhang, Z\. Wei, Y\. Qian, J\. Li, and F\. Wei \(2022\)SpeechT5: unified\-modal encoder\-decoder pre\-training for spoken language processing\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(ACL\),pp\. 5723–5738\.Cited by:[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px2.p1.12)\.
- R\. Ardila, M\. Branson, K\. Davis, M\. Henretty,et al\.\(2020\)Common voice: a massively\-multilingual speech corpus\.InLanguage Resources and Evaluation Conference \(LREC\),Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px1.p1.4)\.
- Z\. Borsos, R\. Marinier, D\. Vincent, E\. Kharitonov, O\. Pietquin, M\. Sharifi, D\. Roblek, O\. Teboul, D\. Grangier, M\. Tagliasacchi,et al\.\(2023\)Audiolm: a language modeling approach to audio generation\.IEEE/ACM transactions on audio, speech, and language processing31,pp\. 2523–2533\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p2.4)\.
- E\. Casanova, K\. Davis, E\. Gölge, G\. Göknar, I\. Gulea, L\. Hart, A\. Aljafari, J\. Meyer, R\. Morais, S\. Olayemi,et al\.\(2024\)Xtts: a massively multilingual zero\-shot text\-to\-speech model\.arXiv preprint arXiv:2406\.04904\.Cited by:[item 1](https://arxiv.org/html/2607.09530#S1.I1.i1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px2.p1.12),[§4\.4](https://arxiv.org/html/2607.09530#S4.SS4.p2.26)\.
- S\. Chen, C\. Wang, Y\. Wu, Z\. Zhang, L\. Zhou, S\. Liu, Z\. Chen, Y\. Liu, H\. Wang, J\. Li,et al\.\(2025a\)Neural codec language models are zero\-shot text to speech synthesizers\.IEEE Transactions on Audio, Speech and Language Processing\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px1.p1.1)\.
- Y\. Chen, Z\. Niu, Z\. Ma, K\. Deng, C\. Wang, J\. JianZhao, K\. Yu, and X\. Chen \(2025b\)F5\-tts: a fairytaler that fakes fluent and faithful speech with flow matching\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6255–6271\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p2.1),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px1.p1.5),[§3\.6](https://arxiv.org/html/2607.09530#S3.SS6.p1.17),[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px2.p1.12),[§4\.4](https://arxiv.org/html/2607.09530#S4.SS4.p2.26)\.
- A\. Conneau, M\. Ma, S\. Khan, Y\. Zhang,et al\.\(2023\)FLEURS: few\-shot learning evaluation of universal representations of speech\.InIEEE Spoken Language Technology Workshop \(SLT\),Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2607.09530#S4.SS2.SSS0.Px2.p1.18)\.
- A\. Défossez, J\. Copet, G\. Synnaeve, and Y\. Adi \(2022\)High fidelity neural audio compression\.arXiv preprint arXiv:2210\.13438\.Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px1.p1.1)\.
- B\. Desplanques, J\. Thienpondt, and K\. Demuynck \(2020\)ECAPA\-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification\.InInterspeech,Cited by:[§4\.3](https://arxiv.org/html/2607.09530#S4.SS3.p2.11)\.
- Z\. Du, Q\. Chen, S\. Zhang, K\. Hu, H\. Lu, Y\. Yang, H\. Hu, S\. Zheng, Y\. Gu, Z\. Ma,et al\.\(2024\)Cosyvoice: a scalable multilingual zero\-shot text\-to\-speech synthesizer based on supervised semantic tokens\.arXiv preprint arXiv:2407\.05407\.Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px1.p1.1)\.
- Z\. Du, C\. Gao, Y\. Wang, F\. Yu, T\. Zhao, H\. Wang, X\. Lv, H\. Wang, C\. Ni, X\. Shi,et al\.\(2025\)Cosyvoice 3: towards in\-the\-wild speech generation via scaling\-up and post\-training\.arXiv preprint arXiv:2505\.17589\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px1.p1.1),[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p2.1)\.
- S\. E\. Eskimez, X\. Wang, M\. Thakker, C\. Li, C\. Tsai, Z\. Xiao, H\. Yang, Z\. Zhu, M\. Tang, X\. Tan,et al\.\(2024\)E2 tts: embarrassingly easy fully non\-autoregressive zero\-shot tts\.In2024 IEEE spoken language technology workshop \(SLT\),pp\. 682–689\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1)\.
- W\. Fan, A\. Y\. Zheng, D\. Zhu, Y\. Ma, N\. Liu, Z\. Wang, and D\. Liu \(2025\)CFG\-zero\*: improved classifier\-free guidance for flow matching models\.arXiv preprint arXiv:2503\.18886\.Cited by:[§3\.6](https://arxiv.org/html/2607.09530#S3.SS6.p1.17)\.
- H\. He, Z\. Shang, C\. Wang, X\. Li, Y\. Gu, H\. Hua, L\. Liu, C\. Yang, J\. Li, P\. Shi,et al\.\(2024\)Emilia: an extensive, multilingual, and diverse speech dataset for large\-scale speech generation\.In2024 IEEE Spoken Language Technology Workshop \(SLT\),pp\. 885–890\.Cited by:[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1)\.
- H\. Hu, X\. Zhu, T\. He, D\. Guo, B\. Zhang, X\. Wang, Z\. Guo, Z\. Jiang, H\. Hao, Z\. Guo,et al\.\(2026\)Qwen3\-tts technical report\.arXiv preprint arXiv:2601\.15621\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1)\.
- D\. Jia, Z\. Chen, J\. Chen, C\. Du, J\. Wu, J\. Cong, X\. Zhuang, C\. Li, Z\. Wei, Y\. Wang,et al\.\(2025\)DiTAR: diffusion transformer autoregressive modeling for speech generation\.InInternational Conference on Machine Learning,pp\. 27255–27270\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- Z\. Ju, Y\. Wang, K\. Shen, X\. Tan,et al\.\(2024\)NaturalSpeech 3: zero\-shot speech synthesis with factorized codec and diffusion models\.International Conference on Machine Learning \(ICML\)\.Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1)\.
- A\. Kazemnejad, I\. Padhi, K\. Natesan Ramamurthy, P\. Das, and S\. Reddy \(2023\)The impact of positional encoding on length generalization in transformers\.Advances in Neural Information Processing Systems36,pp\. 24892–24928\.Cited by:[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- J\. Kim, J\. Kong, and J\. Son \(2021\)Conditional variational autoencoder with adversarial learning for end\-to\-end text\-to\-speech\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1)\.
- R\. Kumar, P\. Seetharaman, A\. Luebs, I\. Kumar, and K\. Kumar \(2023\)High\-fidelity audio compression with improved rvqgan\.Advances in Neural Information Processing Systems36,pp\. 27980–27993\.Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px1.p1.1)\.
- M\. Le, A\. Vyas, B\. Shi, B\. Karrer, L\. Sari, R\. Moritz, M\. Williamson, V\. Manohar, Y\. Adi, J\. Mahadeokar,et al\.\(2023\)Voicebox: text\-guided multilingual universal speech generation at scale\.Advances in neural information processing systems36,pp\. 14005–14034\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p2.1),[§3\.2](https://arxiv.org/html/2607.09530#S3.SS2.p1.12)\.
- T\. Li, Y\. Tian, H\. Li, M\. Deng, and K\. He \(2024\)Autoregressive image generation without vector quantization\.Advances in Neural Information Processing Systems37,pp\. 56424–56445\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px3.p1.1),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- Y\. A\. Li, C\. Han, V\. Raghavan, G\. Mischler, and N\. Mesgarani \(2023\)StyleTTS 2: towards human\-level text\-to\-speech through style diffusion and adversarial training with large speech language models\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1)\.
- Y\. Lipman, R\. T\. Q\. Chen, H\. Ben\-Hamu, M\. Nickel, and M\. Le \(2023\)Flow matching for generative modeling\.International Conference on Learning Representations \(ICLR\)\.Cited by:[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.SSS0.Px1.p1.13)\.
- Z\. Liu, H\. Mao, C\. Wu, C\. Feichtenhofer, T\. Darrell, and S\. Xie \(2022\)A ConvNet for the 2020s\.IEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\)\.Cited by:[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px1.p1.5)\.
- S\. Mehta, R\. Tu, J\. Beskow, É\. Székely, and G\. E\. Henter \(2024\)Matcha\-TTS: a fast TTS architecture with conditional flow matching\.InIEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- L\. Meng, L\. Zhou, S\. Liu, S\. Chen, B\. Han, S\. Hu, Y\. Liu, J\. Li, S\. Zhao, X\. Wu,et al\.\(2025\)Autoregressive speech synthesis without vector quantization\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1287–1300\.Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px3.p1.1)\.
- F\. Mentzer, D\. Minnen, E\. Agustsson, and M\. Tschannen \(2024\)Finite scalar quantization: vq\-vae made simple\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2607.09530#S3.SS2.p1.12)\.
- W\. Peebles and S\. Xie \(2023\)Scalable diffusion models with transformers\.InIEEE/CVF International Conference on Computer Vision \(ICCV\),Cited by:[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- V\. Popov, I\. Vovk, V\. Gogoryan, T\. Sadekova, and M\. Kudinov \(2021\)Grad\-TTS: a diffusion probabilistic model for text\-to\-speech\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- V\. Pratap, A\. Tjandra, B\. Shi,et al\.\(2024\)Scaling speech technology to 1,000\+ languages\.Journal of Machine Learning Research\.Cited by:[item 1](https://arxiv.org/html/2607.09530#S1.I1.i1.p1.1),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px2.p1.12)\.
- A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever \(2023\)Robust speech recognition via large\-scale weak supervision\.InInternational Conference on Machine Learning \(ICML\),Cited by:[§3\.4](https://arxiv.org/html/2607.09530#S3.SS4.SSS0.Px1.p1.2),[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px3.p1.11)\.
- R\. Rafailov, A\. Sharma, E\. Mitchell, C\. D\. Manning, S\. Ermon, and C\. Finn \(2023\)Direct preference optimization: your language model is secretly a reward model\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§5](https://arxiv.org/html/2607.09530#S5.p2.1)\.
- Y\. Ren, C\. Hu, X\. Tan, T\. Qin, S\. Zhao, Z\. Zhao, and T\. Liu \(2020\)FastSpeech 2: fast and high\-quality end\-to\-end text to speech\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1)\.
- J\. Shen, R\. Pang, R\. J\. Weiss, M\. Schuster, N\. Jaitly, Z\. Yang, Z\. Chen, Y\. Zhang, Y\. Wang, R\. Skerrv\-Ryan,et al\.\(2018\)Natural tts synthesis by conditioning wavenet on mel spectrogram predictions\.In2018 IEEE international conference on acoustics, speech and signal processing \(ICASSP\),pp\. 4779–4783\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1)\.
- K\. Shen, Z\. Ju, X\. Tan, E\. Liu, Y\. Leng, L\. He, T\. Qin, J\. Bian,et al\.\(2023\)NaturalSpeech 2: latent diffusion models are natural and zero\-shot speech and singing synthesizers\.InThe Twelfth International Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p1.1),[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§3\.2](https://arxiv.org/html/2607.09530#S3.SS2.p1.12),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- J\. Su, M\. Ahmed, Y\. Lu, S\. Pan, W\. Bo, and Y\. Liu \(2024\)Roformer: enhanced transformer with rotary position embedding\.Neurocomputing568,pp\. 127063\.Cited by:[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7)\.
- M\. Team, C\. Xiao, Y\. Li, X\. Han, Y\. Bai, J\. Cai, H\. Chen, W\. Chen, X\. Cong, G\. Cui,et al\.\(2025\)Minicpm4: ultra\-efficient llms on end devices\.arXiv preprint arXiv:2506\.07900\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p3.1),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p1.1)\.
- A\. Tong, K\. Fatras, N\. Malkin,et al\.\(2024\)Improving and generalizing flow\-based generative models with minibatch optimal transport\.Transactions on Machine Learning Research \(TMLR\)\.Cited by:[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.SSS0.Px1.p1.13)\.
- C\. Wang, A\. Wu, and J\. Pino \(2021\)CoVoST 2 and massively multilingual speech translation\.InInterspeech,Cited by:[§4\.1](https://arxiv.org/html/2607.09530#S4.SS1.SSS0.Px1.p1.4)\.
- X\. Wang, M\. Jiang, Z\. Ma, Z\. Zhang, S\. Liu, L\. Li, Z\. Liang, Q\. Zheng, R\. Wang, X\. Feng,et al\.\(2025\)Spark\-tts: an efficient llm\-based text\-to\-speech model with single\-stream decoupled speech tokens\.arXiv preprint arXiv:2503\.01710\.Cited by:[§4\.4](https://arxiv.org/html/2607.09530#S4.SS4.p2.26)\.
- Y\. Zhou, G\. Zeng, X\. Liu, X\. Li, R\. Yu, J\. Gui, J\. Wu, Z\. Wang, X\. Shen, R\. Ye, Z\. Zhang, J\. Zhou, B\. Bai, W\. Sun, M\. Deng, Q\. Shi, Z\. Wu, and Z\. Liu \(2026\)VoxCPM2 technical report\.arXiv preprint arXiv:2606\.06928\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px4.p1.8),[§3\.2](https://arxiv.org/html/2607.09530#S3.SS2.p1.12),[§3\.3](https://arxiv.org/html/2607.09530#S3.SS3.SSS0.Px2.p1.7),[§4\.4](https://arxiv.org/html/2607.09530#S4.SS4.p2.26),[§5](https://arxiv.org/html/2607.09530#S5.p1.10)\.
- Y\. Zhou, G\. Zeng, X\. Liu, X\. Li, R\. Yu, Z\. Wang, R\. Ye, W\. Sun, J\. Gui, K\. Li,et al\.\(2025\)Voxcpm: tokenizer\-free tts for context\-aware speech generation and true\-to\-life voice cloning\.arXiv preprint arXiv:2509\.24650\.Cited by:[§1](https://arxiv.org/html/2607.09530#S1.p2.4),[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2607.09530#S3.SS1.p1.1)\.
- H\. Zhu, W\. Kang, Z\. Yao, L\. Guo, F\. Kuang, Z\. Li, W\. Zhuang, L\. Lin, and D\. Povey \(2025\)Zipvoice: fast and high\-quality zero\-shot text\-to\-speech with flow matching\.arXiv preprint arXiv:2506\.13053\.Cited by:[§2\.1](https://arxiv.org/html/2607.09530#S2.SS1.SSS0.Px2.p1.1),[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px2.p1.1)\.
- H\. Zhu, L\. Ye, W\. Kang, Z\. Yao, L\. Guo, F\. Kuang, Z\. Han, W\. Zhuang, L\. Lin, and D\. Povey \(2026\)OmniVoice: towards omnilingual zero\-shot text\-to\-speech with diffusion language models\.arXiv preprint arXiv:2604\.00688\.Cited by:[§2\.2](https://arxiv.org/html/2607.09530#S2.SS2.SSS0.Px1.p1.1)\.Similar Articles
dots.tts Technical Report
dots.tts presents a 2B-parameter continuous autoregressive TTS model trained on multilingual data, achieving state-of-the-art performance on benchmarks like Seed-TTS-Eval with low-latency streaming via CFG-aware MeanFlow distillation. The model, code, and checkpoints are released under Apache 2.0.
TontaubeV1 - Open TTS model release for local long-form generation
TontaubeV1 is an open-weight text-to-speech model released for local long-form generation, supporting English and German with zero-shot voice cloning and low-latency inference on GPUs.
Aratako/Irodori-TTS-500M-v3
Irodori-TTS-500M-v3 is a Japanese TTS model based on Rectified Flow Diffusion Transformer, supporting zero-shot voice cloning and unique emoji-based style/sound effect control.
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
WavTTS presents the first raw waveform generative text-to-speech model using flow matching and Diffusion Transformer, achieving performance comparable to latent-space diffusion models while avoiding information loss from compressed representations.
BreezeBlue/Breeze-TTS-2
Breeze TTS 2 is an open-weight, real-time text-to-speech model that supports voice design and bilingual speech, ranking first among open-weight models on the Artificial Analysis TTS leaderboard.