Term
Fish Audio S2
Fish Audio S2 is an open-source text-to-speech model from the Fish Speech family. With inline tags it controls emotion and emphasis at word level, supports more than 80 languages and delivers very low latency.
Fish Audio S2 — explained in more detail
Fish Audio S2 is the second-generation text-to-speech model from Fish Audio and part of the Fish Speech family. Both inference code and model weights are open-source and available on GitHub and Hugging Face, so S2 can be run and fine-tuned on your own infrastructure. The model was trained on more than 10 million hours of audio across roughly 80 languages.
A distinctive feature is fine-grained control via inline tags: natural-language instructions are embedded directly in the text and act down to word level, for example for emotion and emphasis. S2 offers native multi-speaker support and very low latency. On the Artificial Analysis Speech Arena leaderboard, Fish Audio S2 Pro led the field of open-weight models and narrowed the gap to proprietary models.
Example / In practice
For a multilingual explainer video, Fish Audio S2 lets you specify per passage, via an inline tag, which word is emphasized or with what emotion a sentence is spoken, while operation runs fully self-hosted and free of charge.
Distinction from similar terms
Like Chatterbox, Fish Audio S2 is an open-weight speech model and thus self-hostable, in contrast to the closed ElevenLabs v3. As a TTS model it generates speech from text; its emphasis lies on fine-grained tag control at word level and broad language coverage.