Fish Audio Raises $52M Seed to Revolutionize AI Voice Models for Creators and Enterprises


Source: Ivan Mehta / techcrunch.com

Revolutionizing AI Voice Models

The market for AI-generated voice models is rapidly growing, with various use cases requiring distinct characteristics. On one hand, creative applications necessitate expressive voice models, while enterprises seeking to automate customer support and sales operations demand steerable voice models.

Palo Alto-based Fish Audio is poised to cater to these diverse needs with its extensive library of over 15,000 natural language controls. Since its inception last year, the startup has witnessed significant traction, with over 8 million users leveraging its open-source or hosted models. Moreover, Fish Audio has achieved an impressive annual recurring revenue of $21 million.

To further solidify its position in the market, Fish Audio has secured $52 million in a seed round, led by Coreline Ventures and Capital Today. The funding round also saw participation from prominent investors, including 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

Fish Audio’s journey began as a small project by former Nvidia researcher Shijia Liao, who was dissatisfied with the non-expressive synthetic voices available in the market. Liao trained a voice-generation model on a single GPU and subsequently open-sourced it, giving birth to the Fish Speech repository on GitHub. This repository has garnered over 31,000 stars and is utilized by indie developers, video game designers, and creators.

In the past year, Fish Audio has launched five models, including four speech-generation models and one speech-to-text model. While three of its speech-generation models are open-sourced, its latest S2.1 Pro model is available exclusively through its paid API.

Fish Audio offers paid monthly plans tailored for creators and teams, which unlock a set number of minutes of generation along with voice-cloning features. The company also provides an enterprise version of its APIs and platform, with notable clients such as HeyGen and Sanas.

According to Fish Audio’s CEO and co-founder Rissa Cao, every enterprise has distinct use cases and preferences. For instance, companies like HeyGen, which utilize Fish Audio’s voices to power AI avatars, demand realism in voices. In contrast, a gaming studio would require expressive voices for their characters, while voice agent companies like LiveKit need more natural-sounding and low-latency voices that are expressive enough for calls.

One of the key ways Fish Audio has built its library of voices is by soliciting users to submit their own voices for training its models, and compensating them if their voices are used. However, this approach led to some challenges a few months ago, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA takedown process in place to address such concerns, but the takedowns themselves took a considerable amount of time.

Now, Fish Audio has automated the takedown process, enabling creators to submit a short voice sample or a contract to prove that an uploaded voice belongs to them. This will result in the voice being taken off the platform within less than three minutes.

Osuke Honda, a partner at Coreline Ventures, emphasized the importance of a community-driven approach in AI voice models, stating that creators’ trust is crucial for the platform’s success. He advocated for verified voice ownership, clear licensing terms, easy reporting, and takedown processes, as well as revenue-sharing models where creators benefit financially when their voices are licensed or used commercially.

Cao mentioned that when Fish Audio was only offering its product as an open-source project with plans for creators, it was running efficiently and didn’t require outside capital. However, the company wanted to develop more advanced models and accommodate enterprises as investor interest was ramping up, prompting it to seek capital.

In the near future, Fish Audio plans to release an audio understanding model and is also working on a speech-to-speech model. The startup aims to continue pushing the boundaries of AI voice models and solidify its position in the market.

The speech-generation market is increasingly crowded, with companies like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp vying for creators and enterprises’ budgets. Rico Mallozzi, a partner at 359 Capital, believes that fine-grained controls for developers and cost-efficient model training will help Fish Audio compete better with big AI labs.