Fish Audio
Clone a voice from a 10-second clip and generate lifelike speech via API or playground, now on the S2.1 Pro model with emotion tags, fresh off a $52M seed round.
Alternatives
Overview
Fish Audio builds the text-to-speech models that quietly power voice features inside HeyGen, Retell, Sanas, and other AI products, and it also ships that same technology as a standalone playground and API. Its newer flagship model, S2.1 Pro, adds in-text emotion tags, marking a line as angry, whispering, or laughing, ultra-low latency, and support for 30-plus languages, built on top of the S1 lineage that was trained on more than 2 million hours of audio and hit a 0.8% word error rate on the Seed TTS benchmark. Voice cloning needs as little as 10 to 30 seconds of reference audio with no separate training step.
A $52M seed round in late July 2026 confirmed real traction: 8M-plus users and roughly $21M in annualized revenue at seed stage, unusually strong numbers for a company still raising its first institutional round. The API is pay-as-you-go from the first call, priced at $15 per million UTF-8 bytes of input text with no seat fees or minimum commitment, and a free-tier model variant exists for testing and prototyping without production guarantees. For teams that need to self-host, Fish Audio's core models ship as open weights under a paid commercial license (three of its five released models are fully open-source), deployable in a private VPC, on-prem, or air-gapped environment.
The consumer-facing playground adds credit-based subscription plans with unlimited public voice slots on every tier, making it accessible for solo creators as well as engineering teams shipping voice into their own product.
Our Take
Fish Audio's actual story is that it's the TTS infrastructure already running inside HeyGen and Retell in production, available directly as a standalone API and playground. That makes it the logical pick for developers who want voice cloning or speech synthesis without building on top of a consumer-facing product that abstracts away the underlying model. Voice cloning works from a 10-second clip, API pricing starts at $15 per million characters with no minimum commitment, and the genuinely usable free tier covers experimentation. Commercial voice cloning rights require a paid plan, so confirm licensing before deploying.
Key Features
- Zero-shot voice cloning from 10-30 seconds of reference audio
- S1 flagship model with open-domain emotion and tone markers
- Pay-as-you-go API with WebSocket streaming and Python/TypeScript SDKs
- Open-weight self-hosting option under a commercial license
- Multilingual, cross-lingual speech with no phoneme dependency
- • Same model quality already trusted in production by HeyGen and Retell
- • Genuinely usable free tier plus API pricing with no minimum commitment
- • Voice cloning quality holds up well even from short reference samples
- • Credit-to-minute math on the consumer plans takes a minute to understand
- • Best commercial voice cloning rights require a paid plan, not the free tier
- • Self-hosting the open-weight models is an enterprise-level engagement, not a quick setup
People Also Use
Other Audio tools builders reach for alongside Fish Audio.
Suno
Turn a prompt into a finished track: vocals, instruments, and full production in seconds. Advanced Split now rebuilds individual stems, and Suno Studio exports MIDI.
Krisp
Strip background noise and accents out of calls in real time, with AI meeting notes and call-center agent assist layered on top.
Adobe Podcast
Strip background noise and echo from raw recordings to get studio-quality audio, plus record, caption, and transcribe podcasts directly in the browser.
Otter.ai
Transcribe and summarize meetings in real time, then chat with an AI across your meeting history and CRM. The new SDR Agent runs autonomous, personalized video calls with website visitors.
Lalal.ai
Separate vocals, drums, bass, and other instruments from a track into up to 10 individual stems for remixing, mastering, or karaoke use.
Murf
Generate studio-quality voiceovers for videos and presentations from text, with commercial usage rights included from the Creator tier up.