Open source : Turning vocal imitations into sound effects. (New UX for sound generation)

Reddit r/LocalLLaMA Models

Summary

An open-source AI model that generates sound effects from vocal imitations and text descriptions, addressing the challenge of searching for specific sounds.

Hello guys I want to introduce my new project! Have you ever needed a specific sound while making a video or a game? You know exactly what it sounds like in your head, but have no idea how to search for it. That’s why sound design meetings at game studios often turn into people making noises with their mouths. “Not pewpew… more like pew↘︎pew↘︎.” That’s what inspired this project! It’s a model that lets you imitate a sound with your voice, then uses that vocal imitation together with text as input to generate the sound you actually want. repo: [https://github.com/thxxx/VTS](https://github.com/thxxx/VTS) *(You’ll get a better sense of it if you check out the demo in the repo. Would love to hear your feedback in the comments.)*
Original Article

Similar Articles

@CycleDecoded: Shanghai AI Laboratory (OpenMMLab) has completely demolished "sound creation" this time. Their open-source Amphion is simply an "all-purpose arsenal" for the audio-visual creation world. From speech synthesis to AI singing voice conversion, this thing can beat the vast majority of paid voice software. Basic Info Project Name:…

X AI KOLs Timeline

OpenMMLab has open-sourced Amphion, an audio generation toolbox supporting TTS, singing voice conversion, sound effect generation, and more. It is completely free for commercial use and supports local deployment.

@FakeMaidenMaker: Explosive! This open-source project converts text to human-like voice for free, can clone anyone's voice, and adjust timbre with text! GitHub has garnered 30K stars, from Mianbao Intelligent OpenBMB, VoxCPM previously topped both GitHub and HuggingFace charts. Do...

X AI KOLs Timeline

VoxCPM2 is an open-source speech synthesis model from OpenBMB, using a tokenizer-free diffusion autoregressive architecture, supporting 30 languages, voice design, and controllable voice cloning. It can clone a voice with just one sentence, or create a brand new voice using text, outputting 48kHz high-quality audio, and is commercially usable.