@NFTCPS: Fraud call centers have a new weapon — voice cloning has been pushed to new heights again. LuxTTS, a lightweight TTS model, after seeing it I can only say: truly insane. Fast: 150x real-time on a single GPU, even runs faster than real speech on CPU. Clear: 48kHz directly, most models are still stuck at 24kHz…
Summary
LuxTTS is a lightweight voice cloning TTS model, supporting 48kHz high-fidelity output, achieving 150x real-time speed on a single GPU, requiring only 1GB VRAM for local operation, with performance comparable to models ten times its size.
View Cached Full Text
Cached at: 07/05/26, 06:36 PM
The telecom fraud park has new toys again — voice cloning has been pushed to a whole new level. LuxTTS, a lightweight TTS model, left me with just three words: absolutely insane.
- Fast: 150x real-time on a single GPU, even runs faster than human speech on CPU.
- Clear: Direct 48kHz output, while most models are still stuck at 24kHz.
- Efficient: Only 1GB VRAM needed; your old GPU can handle it.
- Cloning quality: Dares to challenge models 10x its size, runs locally without begging for cloud APIs.
https://github.com/ysharma3501/LuxTTS…
ysharma3501/LuxTTS
Source: https://github.com/ysharma3501/LuxTTS
LuxTTS
LuxTTS is an lightweight zipvoice based text-to-speech model designed for high quality voice cloning and realistic generation at speeds exceeding 150x realtime.
https://github.com/user-attachments/assets/a3b57152-8d97-43ce-bd99-26dc9a145c29
The main features are
- Voice cloning: SOTA voice cloning on par with models 10x larger.
- Clarity: Clear 48khz speech generation unlike most TTS models which are limited to 24khz.
- Speed: Reaches speeds of 150x realtime on a single GPU and faster then realtime on CPU’s as well.
- Efficiency: Fits within 1gb vram meaning it can fit in any local gpu.
Usage
You can try it locally, colab, or spaces.
Open In Colab (https://colab.research.google.com/drive/1cDaxtbSDLRmu6tRV_781Of_GSjHSo1Cu?usp=sharing)
Open in Spaces (https://huggingface.co/spaces/YatharthS/LuxTTS)
Simple installation:
git clone https://github.com/ysharma3501/LuxTTS.git cd LuxTTS pip install -r requirements.txt
Load model:
``python
from zipvoice.luxvoice import LuxTTS
load model on GPU
lux_tts = LuxTTS(‘YatharthS/LuxTTS’, device=‘cuda’)
load model on CPU
lux_tts = LuxTTS(‘YatharthS/LuxTTS’, device=‘cpu’, threads=2)
load model on MPS for macs
lux_tts = LuxTTS(‘YatharthS/LuxTTS’, device=‘mps’)
``
Simple inference
``python
import soundfile as sf
from IPython.display import Audio
text = “Hey, what’s up? I’m feeling really great if you ask me honestly!”
change this to your reference file path, can be wav/mp3
prompt_audio = ‘audio_file.wav’
encode audio(takes 10s to init because of librosa first time)
encoded_prompt = lux_tts.encode_prompt(prompt_audio, rms=0.01)
generate speech
final_wav = lux_tts.generate_speech(text, encoded_prompt, num_steps=4)
save audio
final_wav = final_wav.numpy().squeeze()
sf.write(‘output.wav’, final_wav, 48000)
display speech
if display is not None:
display(Audio(final_wav, rate=48000))
``
Inference with sampling params:
``python
import soundfile as sf
from IPython.display import Audio
text = “Hey, what’s up? I’m feeling really great if you ask me honestly!”
change this to your reference file path, can be wav/mp3
prompt_audio = ‘audio_file.wav’
rms = 0.01 ## higher makes it sound louder(0.01 or so recommended)
t_shift = 0.9 ## sampling param, higher can sound better but worse WER
num_steps = 4 ## sampling param, higher sounds better but takes longer(3-4 is best for efficiency)
speed = 1.0 ## sampling param, controls speed of audio(lower=slower)
return_smooth = False ## sampling param, makes it sound smoother possibly but less cleaner
ref_duration = 5 ## Setting it lower can speedup inference, set to 1000 if you find artifacts.
encode audio(takes 10s to init because of librosa first time)
encoded_prompt = lux_tts.encode_prompt(prompt_audio, duration=ref_duration, rms=rms)
generate speech
final_wav = lux_tts.generate_speech(text, encoded_prompt, num_steps=num_steps, t_shift=t_shift, speed=speed, return_smooth=return_smooth)
save audio
final_wav = final_wav.numpy().squeeze()
sf.write(‘output.wav’, final_wav, 48000)
display speech
if display is not None:
display(Audio(final_wav, rate=48000))
``
Tips
- Please use at minimum a 3 second audio file for voice cloning.
- You can use return_smooth = True if you hear metallic sounds.
- Lower t_shift for less possible pronunciation errors but worse quality and vice versa.
Community
- Lux-TTS-Gradio (https://github.com/NidAll/LuxTTS-Gradio): A gradio app to use LuxTTS.
- OptiSpeech (https://github.com/ycharfi09/OptiClone): Clean UI app to use LuxTTS.
- LuxTTS-Comfyui (https://github.com/DragonDiffusionbyBoyo/BoyoLuxTTS-Comfyui.git): Nodes to use LuxTTS in comfyui.
- FalAI hosting (https://fal.ai/models/fal-ai/lux-tts): Luxtts demo hosted by FalAI
- Lux-TTS ONNX (https://github.com/ningyos/luxtts-onnx): Onnx code to use luxtts
Thanks to all community contributions!
Info
Q: How is this different from ZipVoice?
A: LuxTTS uses the same architecture but distilled to 4 steps with an improved sampling technique. It also uses a custom 48khz vocoder instead of the default 24khz version.
Q: Can it be even faster?
A: Yes, currently it uses float32. Float16 should be significantly faster(almost 2x).
Roadmap
- Release model and code
- Huggingface spaces demo
- Release MPS support (thanks to @builtbybasit)
- Release LuxTTS v1.5
- Release code for float16 inference
Acknowledgments
- ZipVoice (https://github.com/k2-fsa/ZipVoice) for their excellent code and model.
- Vocos (https://github.com/gemelo-ai/vocos.git) for their great vocoder.
Final Notes
The model and code are licensed under the Apache-2.0 license. See LICENSE for details.
Stars/Likes would be appreciated, thank you.
Email: [email protected]
Similar Articles
@Honcia13: Open-source TTS is going crazy! New weapons for industrial park scams? Tsinghua OpenBMB just released VoxCPM2: 20 billion parameters + 2 million hours of multilingual data training, 48kHz studio-quality sound! The most intense part is—no Tokenizer needed at all, performing diffusion autoregression directly in continuous latent space, maximizing detail retention!
Tsinghua University's OpenBMB has released VoxCPM2, an open-source multilingual TTS model with 20 billion parameters. It supports continuous latent space diffusion autoregressive generation without a Tokenizer, offering 48kHz studio-quality audio and powerful voice cloning and design capabilities.
@FeitengLi: A 99M parameter TTS runs on CPU, faster than a 2B model on A100. Supertone's newly open-sourced supertonic-3 with ONNX Runtime, fully local, can run in browser, on phone, and even on Raspberry Pi.
Supertone released Supertonic 3, an open-source TTS model with 99M parameters that runs faster on CPU than a 2B model on A100, supporting 31 languages and ONNX Runtime for fully local inference.
@Gorden_Sun: NetEase Youdao open-sources Confucius4-TTS, a 1.3B TTS model, supports multilingual, supports voice cloning, good results, very fast. Github: https://github.com/netease-youdao/Confucius4-TTS… Online demo: …
NetEase Youdao open-sourced the 1.3B parameter Confucius4-TTS model, supporting zero-shot voice cloning and cross-lingual speech synthesis in 14 languages, fast and with excellent results.
@NFTCPS: 4GB VRAM running 70B large model? It actually works! AirLLM did a clever trick — layered inference, not loading the whole model into VRAM at once, but layer by layer, compute and discard, squeezing the giant into a small GPU. The best part: 100% open source, freebie warning https://github.com/0xSo…
AirLLM is a fully open-source tool that uses layered inference (loading and releasing VRAM layer by layer) to enable 70B large language models to run on GPUs with only 4GB VRAM, without quantization, distillation, or pruning. It already supports running Llama3.1 405B on 8GB VRAM.
@lxfater: NetEase Youdao open-sourced ZiYue 4 model, within 27B parameters, SOTA in math and science. But what really interests me is its voice feature!! Cloning a voice is nothing new, ElevenLabs could do it long ago. But they all share a common flaw: cross-language accent. Take your Chinese voice and use it to speak Japanese — it has a Chinese accent, you can tell it's a foreigner struggling...
NetEase Youdao open-sourced the ZiYue 4 model with 27B parameters, achieving SOTA in math and science; its voice feature supports 3-second cross-language voice cloning across 14 languages with no accent issue, along with open-sourcing the all-scenario intelligent agent 'Longxia' (Lobster).