AudioFreemium
Fish Audio logo

Fish Audio

The most expressive, emotionally controllable real-time voice model

Overview

About Fish Audio

Fish Audio is what happens when an open source speech model grows a business around it. The company published Fish Speech, a text to speech model that now carries more than 32,000 GitHub stars, then built a hosted platform on the traffic that model already attracted. The result covers text to speech, voice cloning from roughly fifteen seconds of reference audio, speech to text, voice conversion and audio separation, sold both as a consumer web studio and as a low latency streaming API that other AI companies build on top of.

The differentiator the company leads with is emotional control. Most synthesised speech is flat, and the usual workaround is to re-record or hand tune prosody afterwards. Fish Audio exposes emotion and effect tags inline in the text itself, so whispering, excitement, sighing, laughing and explicit pauses become directives you write rather than artifacts you hope for. Around that sits a community library of more than two million user uploaded voices spanning 30 or more languages. Named API customers include HeyGen, Retell AI, LiveKit, Telnyx and Sanas, which positions the company as infrastructure underneath other AI products rather than a rival to them. The models are published under a custom research licence rather than a standard open source one, so read the licence before shipping on it.

Pricing

$0 - $999/mo

Free tier
  • Free Tier$0/mo
  • Plus$15/mo
  • Pro$100/mo
  • Max$999/mo
  • EnterpriseContact sales

From the vendor pricing page, 2026-09-19

Details

GitHub Stars 32,275
Forks 2,781
Users 8M+
Rating 4.9/5 from 2,847 reviews
Data from: GitHub • Website•Updated: Aug 19, 2026
voicetext-to-speechttsspeechvoice-cloning