Powered by Smartsupp

Fish Audio Launches 15,000+ Natural Language Controls for Expressive, Steerable AI Voice Models



By admin | Jul 28, 2026 | 3 min read


Fish Audio Launches 15,000+ Natural Language Controls for Expressive, Steerable AI Voice Models

The market for AI-generated voice models has grown immensely, driven by two major demands. Creative professionals need voices that convey genuine emotion and expression, while enterprises automating customer support and sales operations require models that are highly controllable and predictable. Fish Audio, based in Palo Alto, aims to serve both sides with a library of over 15,000 natural language controls. Since its launch last year, the startup has attracted more than 8 million users across its open-source and hosted versions, and now generates $21 million in annual recurring revenue.

To build on this momentum, Fish Audio announced on Tuesday that it has raised $50 million in a seed round led by Coreline Ventures and Capital Today. Additional investors include 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The company began as a small project by former NVIDIA researcher Shijia Liao, who grew frustrated with the lack of expressiveness in synthetic voices. He trained a voice generation model on a single GPU and open-sourced it. That repository, Fish Speech, now has over 31,000 stars on GitHub and is widely used by indie developers, video game designers, and content creators.

Over the past year, Fish Audio has released five models: four for speech generation and one for speech-to-text. Three of its speech generation models are open-source, but its latest S2.1 Pro model is available only through a paid API. The company offers monthly subscription plans for creators and teams, which include a set number of generation minutes and voice cloning features. It also provides an enterprise version of its APIs and platform, with companies like HeyGen, Sanas, and Plaud already using its technology.

“Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls,” said Cao, a company representative.

One way Fish Audio built its voice library was by asking users to submit their own voices for training, compensating them if their voices were used. However, this approach led to controversy a few months ago when some creators alleged their voices were uploaded without consent. The startup had a DMCA takedown process in place, but removals were slow. Now, creators can submit a short voice sample or a contract to prove ownership, and their voice will be removed from the platform in under three minutes. Still, this system does not prevent unauthorized uploads from happening in the first place, and a voice remains on the platform until the creator discovers it and files a takedown request.

Oskue Honda, a partner at Coreline Ventures, emphasized that a community-driven model only succeeds when trust is established. “A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially,” he said.

Cao noted that when Fish Audio was solely an open-source project with creator-focused plans, it operated efficiently without needing external funding. But as the company sought to develop more advanced models and accommodate growing enterprise interest, it decided to raise capital. Looking ahead, Fish Audio plans to release an audio understanding model later this year and is also building a speech-to-speech model.

The speech generation market is crowded, with competitors like ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp all vying for creators and enterprise budgets. According to Rico Mallozzi, a partner at 359 Capital, Fish Audio’s fine-grained controls for developers and cost-efficient model training could give it an edge over larger AI labs. “I think what they’ve been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible,” he said.




Comments

Please log in to leave a comment.

No comments yet. Be the first to comment!