@svpino: Here is a new open-weight audio model you can integrate with your app. I'm a huge sucker for open models that you can h…
Summary
Fish Audio S2 is a new open-weight audio model available on Hugging Face, offering two models for timing and acoustic details, with fast inference and a hosted version S2.1 Pro supporting 83 languages at lower cost than ElevenLabs.
View Cached Full Text
Cached at: 07/29/26, 03:56 AM
Here is a new open-weight audio model you can integrate with your app.
I’m a huge sucker for open models that you can host to keep your data secure and away from Big AI.
The model is Fish Audio S2.
You’ll find the weights, the code to fine-tune it, and a streaming inference engine in Hugging Face. (Link below).
Under the hood, it’s two models: a 4B one to handle timing and meaning, and a 400M one to fill in the acoustic details.
The model is really fast!
If you host it on an H200, it will start returning audio in about 100ms.
If you don’t want to host it, Fish Audio offers a hosted version: the S2.1 Pro.
This model is even faster, with around 90ms latency, and it supports 83 languages.
The S2.1 Pro is also ~83% more affordable than ElevenLabs: you pay around $15 for around 12 hours of transcribed audio, which is incredible.
HuggingFace links below.
Similar Articles
Fish Audio launches S2.1 Pro with support for 83 languages (2 minute read)
Fish Audio launches S2.1 Pro, a production voice model with 90ms latency, support for 83 languages, voice cloning from short samples, and multi-speaker dialogue, available via API with a free tier for development.
Fish Audio S2 Technical Report
Fish Audio S2 is an open-source text-to-speech system featuring multi-speaker capabilities, multi-turn generation, and instruction-following control, backed by a production-ready inference engine with low latency.
@rohanpaul_ai: Fish Audio just made S2.1 Pro free for a month. Here’s everything you need to know about it - Clones any voice from 10 …
Fish Audio has made its S2.1 Pro voice cloning service free for a month, featuring 10-15 second voice cloning, ~90ms response time, support for 83 languages, word-level control, and open-weight models at 1/6th the cost of ElevenLabs.
Fish Audio raises $52M seed to build AI voice models for creators and enterprises
Fish Audio has raised $52 million in seed funding to develop AI voice models for creators and enterprises. The startup, which generates $21M in ARR and has 8 million users, offers open-source and paid voice generation models, including its latest S2.1 Pro API.
@HuggingApps: NVIDIA Nemotron just dropped an audio-native model that hears the world, not just words transcription, translation, sou…
NVIDIA released Nemotron, an audio-native model capable of transcription, translation, sound recognition, audio Q&A, TTS, and full speech-to-speech, with open weights in 2B and 30B sizes.