
Fish Audio
The most expressive, emotionally controllable real-time voice model
About Fish Audio
Fish Audio is what happens when an open source speech model grows a business around it. The company published Fish Speech, a text to speech model that now carries more than 32,000 GitHub stars, then built a hosted platform on the traffic that model already attracted. The result covers text to speech, voice cloning from roughly fifteen seconds of reference audio, speech to text, voice conversion and audio separation, sold both as a consumer web studio and as a low latency streaming API that other AI companies build on top of.
The differentiator the company leads with is emotional control. Most synthesised speech is flat, and the usual workaround is to re-record or hand tune prosody afterwards. Fish Audio exposes emotion and effect tags inline in the text itself, so whispering, excitement, sighing, laughing and explicit pauses become directives you write rather than artifacts you hope for. Around that sits a community library of more than two million user uploaded voices spanning 30 or more languages. Named API customers include HeyGen, Retell AI, LiveKit, Telnyx and Sanas, which positions the company as infrastructure underneath other AI products rather than a rival to them. The models are published under a custom research licence rather than a standard open source one, so read the licence before shipping on it.
Start in the web studio, where you either pick a voice from the community library or clone one from a reference clip of about fifteen seconds. Write your text and place emotion tags inline where you want the delivery to shift, then generate. Output is metered in credits, and the platform documents that a minute of generation costs roughly 600 to 625 credits, so a tier translates directly into minutes rather than an abstract quota. Developers move the same workflow to the REST API, which ships Python and TypeScript SDKs and supports streaming for conversational agents where latency decides whether a voice feels live. Professional voice slots on paid tiers hold your own cloned voices privately rather than publishing them to the shared library. Enterprise agreements add zero data retention, on premise deployment and SOC2 compliance for teams that cannot send audio to a shared service.
- •Inline Emotion Tags - Direct delivery with tags for angry, sad, embarrassed, whispering, soft, breathy and excited rather than accepting flat output
- •Non-verbal Effect Tags - Laughing, sighing, panting and explicit pauses written directly into the script
- •Fifteen Second Voice Cloning - Enhanced voice cloning listed on every tier including the free one
- •Two Million Voice Library - A browsable community library of user uploaded voices across 30 or more languages
- •Low Latency Streaming API - Documented Python and TypeScript SDKs built for conversational agents, used in production by HeyGen, LiveKit, Retell AI and Telnyx
- •Speech to Text - Multispeaker transcription offered as its own product surface alongside generation
- •Voice Changer and Audio Separation - Convert an existing recording into a different voice, or split a mixed track into components
- •Story Studio - Multi-voice narrated audio pieces assembled from a single script
- •Enterprise Controls - Zero data retention, on premise deployment and SOC2 compliance on custom agreements
Developers building voice agents are the core audience, particularly anyone who has hit the latency ceiling of a general purpose text to speech API or the price ceiling of the incumbent. At the entry paid tier the per minute cost undercuts the established players by a wide margin, which matters most for products where every user session generates audio. Creators producing narration, audiobooks or character work get the emotion tags, which is the single thing most complained about in synthetic speech. Teams in regulated environments have a path through on premise deployment and zero data retention. Anyone whose use case involves cloning a voice they do not own should read the licence terms first, because this is the most legally exposed corner of generative AI and the permissive posture here is a business risk as much as a feature.
Pricing
$0 - $999/mo
- Free Tier$0/mo
- Plus$15/mo
- Pro$100/mo
- Max$999/mo
- EnterpriseContact sales
From the vendor pricing page, 2026-09-19














