Skip to main content
Audio modelOpen weights

Fish Audio S2 Pro

Fish Audio

Released
March 9, 2026
Data date
October 3, 2026
Model class

Fish Audio S2 Pro is an open speech synthesis model controlled through prosody and emotion tags. It supports multi-speaker dialogue, voice cloning, and broad language coverage, while the released S2 line is presented for API and local-weights use.

Fish Audio cites about 100 milliseconds to first audio. The official S2 sources cite roughly 50 to 80 languages.

Specifications and access

SpecificationValue and source
Model class
Speech synthesisSource
Access
API and local weightsSource
Input
Text with prosody and emotion tagsSource
Output
Multilingual speechSource
Checkpoint
fishaudio/s2-proSource
Languages
About 50 to 80 languages depending on the documentation variantSource
Latency
About 100 milliseconds to first audioSource
Features
Inline prosody, emotion, multi-speaker dialogue, and voice cloningSource