Open source : Turning vocal imitations into sound effects. (New UX for sound generation)
Summary
An open-source AI model that generates sound effects from vocal imitations and text descriptions, addressing the challenge of searching for specific sounds.
Similar Articles
@Ryrenz: So powerful—OpenVoice just needs a short reference audio clip to clone that voice timbre, and it can even switch langua…
OpenVoice is an open-source AI tool for voice cloning that uses a short reference audio clip to replicate timbre, enabling multilingual and emotional adjustments, and has gained significant traction on GitHub.
@CycleDecoded: Shanghai AI Laboratory (OpenMMLab) has completely demolished "sound creation" this time. Their open-source Amphion is simply an "all-purpose arsenal" for the audio-visual creation world. From speech synthesis to AI singing voice conversion, this thing can beat the vast majority of paid voice software. Basic Info Project Name:…
OpenMMLab has open-sourced Amphion, an audio generation toolbox supporting TTS, singing voice conversion, sound effect generation, and more. It is completely free for commercial use and supports local deployment.
Scenema Audio: Zero-shot expressive voice cloning and speech generation [N]
Scenema AI releases Scenema Audio, an open-source diffusion-based model for zero-shot expressive voice cloning and speech generation, separating emotional performance from voice identity to allow any voice to perform any emotion.
@FakeMaidenMaker: Explosive! This open-source project converts text to human-like voice for free, can clone anyone's voice, and adjust timbre with text! GitHub has garnered 30K stars, from Mianbao Intelligent OpenBMB, VoxCPM previously topped both GitHub and HuggingFace charts. Do...
VoxCPM2 is an open-source speech synthesis model from OpenBMB, using a tokenizer-free diffusion autoregressive architecture, supporting 30 languages, voice design, and controllable voice cloning. It can clone a voice with just one sentence, or create a brand new voice using text, outputting 48kHz high-quality audio, and is commercially usable.
Ultimate List: Best Open Models for Coding, Chat, Vision, Audio & More
Curated list of top open-source models for coding, chat, vision, audio, TTS, voice cloning, music, image and video generation with links and brief performance notes.