Fish Audio raises $52M seed to build AI voice models for creators and enterprises

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane July 28, 2026 2 min read
Fish Audio raises $52M seed to build AI voice models for creators and enterprises

Fish Audio has raised $52 million to expand its library of AI voice models for creators and businesses. The Palo Alto startup announced the seed round on Tuesday, led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

Business figures

Since launching last year, the company now has more than 8 million users across its open-source and hosted versions. Annual recurring revenue stands at $21 million. The startup currently offers five models: four for speech generation and one for speech-to-text. Three of the speech generation models are open-source, while the latest S2.1 Pro model is available only via a paid API.

Origins and growth

Shijia Liao, a former NVIDIA researcher, started the project after becoming frustrated with the lack of expressiveness in synthetic voices available at the time. He trained a voice generation model on a single GPU and released the code. The Fish Speech repository on GitHub now holds more than 31,000 stars. Indie developers, video game designers, and creators currently use the tools.

Enterprise and creator plans

Monthly subscriptions provide a set number of generation minutes plus voice cloning features. An enterprise version of the API and platform is also available. HeyGen and Sanas are among the organisations currently using the service. Rissa Cao, Fish Audio’s CEO and co-founder, noted that different clients require different capabilities. HeyGen prioritises realism for AI avatars, while gaming studios seek expressive voices for characters. LiveKit, a voice agent company, needs low-latency, natural-sounding output suitable for calls.

The company builds its voice library by asking users to submit samples and compensating them for usage. This approach caused trouble a few months ago when some creators claimed their voices were uploaded without consent. A DMCA process existed to address these issues, but removals took too long. Fish Audio has since automated the take-down procedure. Creators can submit a short voice sample or contract to prove ownership, and removals now happen in under three minutes.

Automated removals do not stop anyone from uploading an artist’s voice without their knowledge initially. The voice remains on the platform until the artist discovers the issue and files a request. Osuke Honda, a partner at Coreline Ventures, stated that a community-driven model only works when creators trust the platform. He argued that consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. He believes the industry needs verified voice ownership, clear licensing terms, easy reporting, and revenue-sharing models where creators benefit financially from commercial use.

Cao explained that the startup previously ran efficiently as an open-source project without needing capital. The decision to seek funding came from a desire to develop more advanced models and accommodate enterprise demand as investor interest grew.

Future roadmap

Plans include releasing an audio understanding model this year. The company is also building a speech-to-speech model.

Competition

The speech generation market is crowded. Competitors include ElevenLabs, WellSaid, Cartesia, Speechify, Async, and Krisp. Rico Mallozzi, a partner at 359 Capital, noted that fine-grained controls for developers and cost-efficient training will help Fish Audio compete with larger AI labs. He described the team’s ability to build state-of-the-art models with high technical acumen as incredible, particularly in closing the gap between artificial and human-like voices.

Scroll to Top