Data as of Aug 16, 2026 · Based on 3,131,739 AI responses across 10,525 prompts · See how Parse measures this
Fish Audio S2 is a text-to-speech system trained on over 10 million hours of audio across 50 languages, using a Dual Autoregressive architecture and reinforcement learning to generate natural, emotionally rich speech. It supports fine-grained inline control of prosody and emotion via natural language tags, as well as native multi-speaker and multi-turn generation.
Sources
bentoml.com shapes more of what AI says about Fish Speech than any other source, at 43% of its citations.
reddit.com · github.com · youtube.com
The market map
Speech AI APIs and Services →Where AI ranks Fish Speech
Excerpts where Fish Speech appeared in the AI's answer

Fish Speech / Fish Audio: A massive multimodal open-source project (fish-speech ) supporting fine-grained emotional and prosodic control

Fish Speech (S2-Pro): Considered top-tier for naturalness, often used for high-fidelity voice cloning.