2 Sources
[1]
Fish Audio raises $50M seed to build AI voice models for creators and enterprises
The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable. Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup today has more than 8 million people using the open-source or hosted versions of its models, and now generates annual recurring revenue of $21 million. To continue building on that traction, the startup on Tuesday said it has raised $50 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. Fish Audio started as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars, and is used by indie developers, video game designers, and creators. The company has launched five models in the last year: four speech generation models and one speech-to-text model. It has open-sourced three of its speech generation models, but its latest S2.1 Pro model is available only through its paid API. Fish Audio offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation, plus voice cloning features. The company also offers an enterprise version of its APIs and platform, and says organizations like HeyGen, Sanas and Plaud are already using it. "Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls," Cao said. One way the startup has built its library of voices is by simply asking users to submit their own voices for training its models, and compensating them if their voices are used. That resulted in some trouble a few months ago, however, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA content take-down process in place to address such concerns, but the take-downs themselves took a long time. Fish Audio's CEO and co-founder Rissa Cao told TechCrunch that the company has now automated the take-down process. Creators can easily submit a short voice sample or a contract to prove that an uploaded voice belongs to them, and their voice will be taken off the startup's platform in less than 3 minutes, she said. Still, that doesn't prevent anyone from uploading an artist's voice without their knowledge. And until the artist finds out, their voice will continue to be used on the platform until they file for it to be taken down. Oskue Honda, a partner at Coreline Ventures, said a community-driven model only works when creators trust the platform. "A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially," he said. Cao said when the startup was only offering its product as an open-source project with plans for creators, it was running efficiently and didn't need money. But it wanted to develop more advanced models, and also wanted to accommodate enterprises as investor interest was ramping up, which led it to seek capital. Looking ahead, Fish Audio plans to release an audio understanding model this year. It's also building a speech-to-speech model. The speech generation market is crowded, with companies like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp competing for creators and enterprises' wallets. According to Rico Mallozzi, a partner at 359 Capital, fine-grained controls for developers and cost-efficient model training will help Fish Audio compete better with big AI labs. "I think what they've been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices," Mallozzi told TechCrunch over a call.
[2]
Fish Audio makes a splash after raising $52M seed funding for AI voices
Fish Audio, an ambitious artificial intelligence startup that began life as a weekend project in its founder's bedroom, said today it has raised an impressive $52 million in seed funding to try and establish voice as the default interface for every AI model. Coreline Ventures and Capital Today led the round, which also saw participation from a host of other backers, including 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners and a number of unnamed angels. The startup, which is officially known as Hanabi AI Inc., emerged from a familiar developer origin story. Its co-founder and Chief Scientist Shijia Liao, who formerly worked as a video researcher at Nvidia Corp. and is a lifelong fan of Japanese anime, explained that he grew tired of having to listen to the flat and monotonously robotic synthetic voices of early AI models, and decided that he needed to do something about it. Liao decided to start training his own voice AI models, and set out to do so with nothing more than a single graphics processing unit housed in the laptop in his bedroom. Despite the limited compute available to him, Liao managed to transform the initial text-to-speech and voice cloning models he developed into Fish Speech, an open-source project that quickly gained rapid traction on GitHub. It amassed more than 31,000 stars as it caught the attention of indie developers, content creators and video games designers who were desperate for livelier and more expressive voice generation tools. Since its launch in 2023, Fish Audio has evolved to become one of the most comprehensive voice AI platforms in the business, used by developers to quickly build powerful text-to-speech and voice cloning systems as well as voice agents. The platform was designed specifically to eliminate the lifeless synthetic audio that characterized early AI models by providing granular, word-level emotion controls that are driven by over 15,000 natural language prompts. It allows developers to fine-tune the exact tone, inflection and pacing of their AI-generated voices. Fish Audio's platform is powerful. It claims to be able to clone a voice from a mere five-second audio clip in less than 15 seconds. It boasts native support for 83 languages too, and its current flagship model S2.1 Pro was able to outperform its top competitors in a series of blind listening tests. According to those tests, 67% of listeners preferred Fish Audio's voice outputs ahead of those from other models. Those numbers help to explain why its user base has grown to over eight million and its annual recurring revenue now exceeds $21 million. While initially targeted at video games developers and content creators, Fish Audio now caters to organizations in regulated industries such as healthcare and financial services, offering secure on-premises deployments with zero-data retention and HIPAA compliance to avoid compromising customer privacy. Today's round positions Fish Audio as a formidable challenger to better known voice AI startups such as ElevenLabs Inc., which recently raised $500 million in a round that valued it at a staggering $11 billion. Fish Audio Chief Executive Rissa Cao said he started the company along with Liao because he wanted to make AI voices that sound human to be accessible to everyone. "We make high-quality, human-sounding voices available to every user, from beginner creatives to million-dollar enterprises, so communication is not only more efficient, but more trustworthy," he insisted. "We've always believed that if we kept making the models better, people would notice. Eight million of them did." Following today's massive cash infusion, the startup is now looking to expand its model lineup beyond its initial text-to-speech capabilities. It intends to build a full audio-native stack that's composed of voice-native large language models and real-time speech-to-speech translation tools. At the same time, it also plans to invest in an enterprise sales team and expand its developer tooling with more application programming interface integrations through partners such as Retell AI Inc. and LiveKit Inc.. Fish Audio also wants to accelerate the adoption of its platform, and to that end it's planning to make its flagship S2.1 Pro model available to every developer free of charge through its official API, starting at the end of August. Coreline Ventures' managing partner Osuke Honda said he invested in Fish Audio because he believes voice is rapidly becoming the default interface for AI systems. "In its short history, Fish Audio has built an unbeatable track record of pushing the envelope on performance, multilingual support, emotional expression and cost," he said. "All factors that have quickly made Fish Audio the default choice for creators, developers, and now enterprises globally, and we expect them to continue to lead the way."
Share
Copy Link
Fish Audio, which started as a weekend project by a former NVIDIA researcher, has raised $50 million in seed funding led by Coreline Ventures and Capital Today. The Palo Alto-based startup now serves over 8 million users with its library of more than 15,000 natural language controls and generates $21 million in annual recurring revenue.
Fish Audio has secured $50 million in seed funding led by Coreline Ventures and Capital Today, marking a dramatic evolution from its origins as a weekend project. The Palo Alto-based startup, which began when former NVIDIA researcher Shijia Liao trained a voice generation model on a single GPU in his bedroom, now serves over 8 million users and generates $21 million in annual recurring revenue
1
2
. The round also attracted participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0, positioning the company to challenge established players like ElevenLabs in the increasingly competitive voice AI market.
Source: TechCrunch
What sets Fish Audio apart in the crowded AI voice generation landscape is its library of more than 15,000 natural language controls that enable developers to create both expressive and steerable voices. The voice AI platform allows users to fine-tune exact tone, inflection, and pacing at the word level, addressing a critical gap in the market where creative use cases demand more expressive outputs while enterprises need steerable systems for customer support and sales automation
1
. The company's flagship S2.1 Pro model demonstrated its technical prowess in blind listening tests, with 67% of listeners preferring its outputs over competitors2
. Fish Audio can clone a voice from just a five-second audio clip in under 15 seconds, with native support for 83 languages.
Source: SiliconANGLE
Fish Audio's open-source strategy has proven instrumental in its growth trajectory. The Fish Speech repository on GitHub has amassed over 31,000 stars, attracting indie developers, video game designers, and content creators desperate for human-like voice synthesis capabilities
1
. Since launching in 2023, the startup has released five AI voice models—four speech generation models and one speech-to-text model—with three speech generation models available as open-source. However, the latest S2.1 Pro model remains exclusive to paid API users, though the company plans to make it available free to all developers starting at the end of August2
. Organizations like HeyGen, Sanas, and Plaud already use Fish Audio's enterprise APIs, with each requiring different voice characteristics—from realistic AI avatars to expressive game characters and natural-sounding, low-latency voice agents.Related Stories
Fish Audio's community-driven approach to building its voice library has encountered friction around consent and creator rights. The startup asks users to submit their own voices for training AI voice models, compensating them when their voices are used. However, several months ago, creators alleged their voices were uploaded to Fish Audio without consent, exposing vulnerabilities in the platform's DMCA takedown process
1
. CEO and co-founder Rissa Cao told TechCrunch the company has now automated takedowns, allowing creators to submit a voice sample or contract to prove ownership and remove their voice in less than three minutes. Yet this reactive approach doesn't prevent unauthorized uploads in the first place—voices remain available until artists discover the infringement and file for removal. Osuke Honda, a partner at Coreline Ventures, emphasized that "a community-centric approach can only become a durable advantage if creators trust the platform," advocating for verified voice ownership, clear licensing terms, and revenue-sharing models where creators benefit financially from commercial use1
.The seed funding will fuel Fish Audio's expansion beyond its creator-focused roots into regulated industries. The company now offers HIPAA-compliant, on-premises deployments with zero-data retention for healthcare and financial services organizations
2
. Cao explained that while the startup was running efficiently with its open-source project and creator plans, it needed capital to develop more advanced models and accommodate enterprise customers as investor interest intensified1
. Looking ahead, Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model to create a full audio-native stack with voice-native large language models and real-time translation tools1
2
. The company will also invest in an enterprise sales team and expand developer tooling through API integrations with partners like Retell AI and LiveKit. Rico Mallozzi, a partner at 359 Capital, noted that fine-grained natural language controls for developers and cost-efficient model training will help Fish Audio compete with larger AI labs, praising the team's ability to build state-of-the-art models and close the gap between artificial-sounding and human-like voices with limited resources1
.Summarized by
Navi
31 Jan 2025•Business and Economy

02 Dec 2025•Startups

15 Dec 2025•Startups

1
Technology

2
Policy and Regulation

3
Technology
