Back to Directory
Fish Audio logo

Fish Audio

Clone a voice from a 10-second clip and generate lifelike speech via API or playground, now on the S2.1 Pro model with emotion tags, fresh off a $52M seed round.

Audio
4.6freemium

Alternatives

Overview

Fish Audio builds the text-to-speech models that quietly power voice features inside HeyGen, Retell, Sanas, and other AI products, and it also ships that same technology as a standalone playground and API. Its newer flagship model, S2.1 Pro, adds in-text emotion tags, marking a line as angry, whispering, or laughing, ultra-low latency, and support for 30-plus languages, built on top of the S1 lineage that was trained on more than 2 million hours of audio and hit a 0.8% word error rate on the Seed TTS benchmark. Voice cloning needs as little as 10 to 30 seconds of reference audio with no separate training step.

A $52M seed round in late July 2026 confirmed real traction: 8M-plus users and roughly $21M in annualized revenue at seed stage, unusually strong numbers for a company still raising its first institutional round. The API is pay-as-you-go from the first call, priced at $15 per million UTF-8 bytes of input text with no seat fees or minimum commitment, and a free-tier model variant exists for testing and prototyping without production guarantees. For teams that need to self-host, Fish Audio's core models ship as open weights under a paid commercial license (three of its five released models are fully open-source), deployable in a private VPC, on-prem, or air-gapped environment.

The consumer-facing playground adds credit-based subscription plans with unlimited public voice slots on every tier, making it accessible for solo creators as well as engineering teams shipping voice into their own product.

Our Take

Fish Audio's actual story is that it's the TTS infrastructure already running inside HeyGen and Retell in production, available directly as a standalone API and playground. That makes it the logical pick for developers who want voice cloning or speech synthesis without building on top of a consumer-facing product that abstracts away the underlying model. Voice cloning works from a 10-second clip, API pricing starts at $15 per million characters with no minimum commitment, and the genuinely usable free tier covers experimentation. Commercial voice cloning rights require a paid plan, so confirm licensing before deploying.

Key Features

  • Zero-shot voice cloning from 10-30 seconds of reference audio
  • S1 flagship model with open-domain emotion and tone markers
  • Pay-as-you-go API with WebSocket streaming and Python/TypeScript SDKs
  • Open-weight self-hosting option under a commercial license
  • Multilingual, cross-lingual speech with no phoneme dependency
Pros
  • Same model quality already trusted in production by HeyGen and Retell
  • Genuinely usable free tier plus API pricing with no minimum commitment
  • Voice cloning quality holds up well even from short reference samples
Cons
  • Credit-to-minute math on the consumer plans takes a minute to understand
  • Best commercial voice cloning rights require a paid plan, not the free tier
  • Self-hosting the open-weight models is an enterprise-level engagement, not a quick setup

Other Audio tools builders reach for alongside Fish Audio.