OpenVoice: Versatile Instant Voice Cloning

Papers with Code Trending Papers

Summary

OpenVoice is a versatile voice cloning method that enables flexible voice style control and zero-shot cross-lingual cloning with high efficiency using a single audio clip.

We introduce OpenVoice, a versatile voice cloning approach that requires only a short audio clip from the reference speaker to replicate their voice and generate speech in multiple languages. OpenVoice represents a significant advancement in addressing the following open challenges in the field: 1) Flexible Voice Style Control. OpenVoice enables granular control over voice styles, including emotion, accent, rhythm, pauses, and intonation, in addition to replicating the tone color of the reference speaker. The voice styles are not directly copied from and constrained by the style of the reference speaker. Previous approaches lacked the ability to flexibly manipulate voice styles after cloning. 2) Zero-Shot Cross-Lingual Voice Cloning. OpenVoice achieves zero-shot cross-lingual voice cloning for languages not included in the massive-speaker training set. Unlike previous approaches, which typically require extensive massive-speaker multi-lingual (MSML) dataset for all languages, OpenVoice can clone voices into a new language without any massive-speaker training data for that language. OpenVoice is also computationally efficient, costing tens of times less than commercially available APIs that offer even inferior performance. To foster further research in the field, we have made the source code and trained model publicly accessible. We also provide qualitative results in our demo website. Prior to its public release, our internal version of OpenVoice was used tens of millions of times by users worldwide between May and October 2023, serving as the backend of MyShell.
Original Article
View Cached Full Text

Cached at: 08/24/26, 10:08 AM

Paper page - OpenVoice: Versatile Instant Voice Cloning

Source: https://huggingface.co/papers/2312.01479

Abstract

OpenVoice enables flexible voice style control and zero-shot cross-lingual voice cloning with high efficiency using a single audio clip.

We introduce OpenVoice, a versatilevoice cloningapproach that requires only a shortaudio clipfrom the reference speaker to replicate their voice and generate speech in multiple languages. OpenVoice represents a significant advancement in addressing the following open challenges in the field: 1) Flexible Voice Style Control. OpenVoice enables granular control over voice styles, includingemotion,accent,rhythm,pauses, andintonation, in addition to replicating thetone colorof the reference speaker. Thevoice stylesare not directly copied from and constrained by the style of the reference speaker. Previous approaches lacked the ability to flexibly manipulatevoice stylesafter cloning. 2) Zero-Shot Cross-LingualVoice Cloning. OpenVoice achieves zero-shot cross-lingualvoice cloningfor languages not included in the massive-speaker training set. Unlike previous approaches, which typically require extensive massive-speaker multi-lingual (MSML) dataset for all languages, OpenVoice can clone voices into a new language without any massive-speaker training data for that language. OpenVoice is also computationally efficient, costing tens of times less than commercially available APIs that offer even inferior performance. To foster further research in the field, we have made thesource codeandtrained modelpublicly accessible. We also provide qualitative results in our demo website. Prior to its public release, our internal version of OpenVoice was used tens of millions of times by users worldwide between May and October 2023, serving as the backend of MyShell.

View arXiv pageView PDFGitHub37.3kautoAdd to collection

Get this paper in your agent:

hf papers read 2312\.01479

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper4

#### rsxdalv/OpenVoiceV2 #### ameerazam08/Udiff UpdatedDec 16, 2023 #### IrieDinamik/OpenVoice UpdatedMay 12 #### ayodkay-hf/soninho-voice Updated18 days ago

Datasets citing this paper5

#### tsinghua-ee/QualiSpeech Viewer• UpdatedAug 4, 2025 • 14.6k • 571 • 24 #### dlxjj/Openvoice Viewer• UpdatedMay 8, 2025 • 10 • 217 #### Pendrokar/open_tts_tracker Viewer• UpdatedFeb 20 • 60 • 175 • 30 #### purdueviperlab/diffssd Viewer• UpdatedSep 30, 2024 • 283k • 152 • 2 Browse 5 datasets citing this paper### Spaces citing this paper8

Collections including this paper2

Similar Articles

k2-fsa/OmniVoice

Hugging Face Models Trending

OmniVoice is a massively multilingual zero-shot text-to-speech model supporting over 600 languages, built on a diffusion language model architecture with fast inference and voice cloning capabilities.