@rohanpaul_ai: Fish Audio just made S2.1 Pro free for a month. Here’s everything you need to know about it - Clones any voice from 10 …
Summary
Fish Audio has made its S2.1 Pro voice cloning service free for a month, featuring 10-15 second voice cloning, ~90ms response time, support for 83 languages, word-level control, and open-weight models at 1/6th the cost of ElevenLabs.
View Cached Full Text
Cached at: 07/31/26, 12:43 AM
Fish Audio just made S2.1 Pro free for a month.
Here’s everything you need to know about it
- Clones any voice from 10 to 15 seconds of audio.
- ~90ms response, fast enough for real conversation
- 83 languages, one model.
- Word-level control over pronunciation and pauses
- Also has open weighted models
Under the hood it uses a Dual-AR design: a 4B-parameter model works out what to say and how it should feel, and a 400M-parameter model fills in the fine sound detail. That split is why it’s both fast and expressive.
It’s built from Fish Speech, their open-source project with tens of thousands of GitHub stars, and they still ship open-weight models you can self-host.
Price: roughly 1/6th of ElevenLabs for the same output.
The real pitch isn’t a clean 10-second clip. It’s a voice that survives a whole conversation, interruptions, corrections, laughter and language switches included.
Similar Articles
Fish Audio launches S2.1 Pro with support for 83 languages (2 minute read)
Fish Audio launches S2.1 Pro, a production voice model with 90ms latency, support for 83 languages, voice cloning from short samples, and multi-speaker dialogue, available via API with a free tier for development.
@driscollis: Fish Audio's AI can clone a voice in 5 seconds! I can see this being a really useful tool for content creation, especia…
Fish Audio announced its S2.1 Pro model, which can clone a voice from just 5 seconds of audio, alongside a $52M seed funding round.
Fish Audio raises $52M seed to build AI voice models for creators and enterprises
Fish Audio has raised $52 million in seed funding to develop AI voice models for creators and enterprises. The startup, which generates $21M in ARR and has 8 million users, offers open-source and paid voice generation models, including its latest S2.1 Pro API.
this new Moss tts 1.5 is damn good with voice cloning
MOSS TTS 1.5 is a new text-to-speech model with voice cloning capabilities, offered via a Hugging Face Space, and is considered better than Fish Audio S2 Pro due to open licensing.
@svpino: Here is a new open-weight audio model you can integrate with your app. I'm a huge sucker for open models that you can h…
Fish Audio S2 is a new open-weight audio model available on Hugging Face, offering two models for timing and acoustic details, with fast inference and a hosted version S2.1 Pro supporting 83 languages at lower cost than ElevenLabs.