Tag
Qwen3.8-LiveTranslate is an AI model that reduces translation lag and improves quality using an interleaved audio-text architecture, featuring real-time speaker separation and voice cloning in 60 languages.
The tweet showcases ten trending open-source GitHub repositories focused on AI agents, browser automation, voice cloning, and local model execution, indicating a growing infrastructure for AI agent development.
A solo developer has open-sourced a free, local AI tool that serves as an ElevenLabs replacement, enabling voice cloning, video dubbing in 646 languages, and more without any data leaving the user's machine.
OpenVoice is an open-source AI tool for voice cloning that uses a short reference audio clip to replicate timbre, enabling multilingual and emotional adjustments, and has gained significant traction on GitHub.
Steven Johnson of Google Labs explains how AI can clone a writer's voice without permission, highlighting challenges for copyright protection and the freelance economy.
LoudKit is an open-source local TTS tool with voice cloning, supporting 10 languages, and available as SDKs for Python, Swift, Go, Rust, and TypeScript, designed to run efficiently on edge devices.
Tencent has open-sourced a 1.5B parameter AI model called AuK that can replace multiple audio tools, handling tasks like TTS, voice cloning, and denoising via natural language instructions.
The article discusses three converging toolchains in neural voice synthesis—vocal identity, signal manipulation, and performance realism—focusing on voice cloning and timbre transfer to transform vocal performances while preserving original phrasing and dynamics.
This paper presents a method to build a compact fixed-voice Thai TTS system using synthetic speech from a larger model, evaluating its performance and introducing an 82M-parameter model for on-device deployment.
This paper presents Wayu-Paxa-TTS-Edge, an 82M-parameter Thai TTS model trained on synthetic speech from a voice-cloning teacher, achieving high accuracy and prosody for on-device use without reference audio.
VoiceStudio is an open-source local voice studio tool offering voice cloning, dubbing, transcription, and more in 646 languages, providing an alternative to ElevenLabs.
Introducing the local voice cloning tool VoiceStudio, which only requires 8 seconds of recording to highly replicate the voice, runs completely locally to protect privacy, and supports multiple languages and platforms.
A lightweight TTS implementation, replicating Audio8's training with 2000 hours of data, achieving a SIM metric of 0.72 in under 10 hours of training on H200.
TontaubeV1 is an open-weight text-to-speech model released for local long-form generation, supporting English and German with zero-shot voice cloning and low-latency inference on GPUs.
Sopro V2 is an open-source, fast, on-device text-to-speech model with voice cloning, specifically targeting European Portuguese and other languages for low-latency and private communication.
The article reviews and ranks AI voice cloning models from 2026, including CosyVoice 3 and VibeVoice, based on their accuracy in replicating the author's voice through fine-tuning and zero-shot methods.
This paper introduces a synthetic Bengali speech dataset of 10,000 audio-text pairs for telecom customer care scenarios, generated using OmniVoice voice-cloning, and evaluates it with an ASR model, achieving low word error rates.
VoiceStudio is a locally running AI voice tool, offering features such as voice cloning, sound design, video dubbing, etc., supporting 646 languages, with no need for network connection or subscription, and all data processed locally.
FireRedAudio and FireRedTTS3 are AI models for unified audio understanding and generation, offering ASR, TTS, voice cloning, editing, and multilingual support with competitive benchmarks.
Audio8 TTS Preview 0.1b is a compact zero-shot text-to-speech model with approximately 170M parameters for the main model, supporting voice cloning and multiple languages.