Fish Audio raises $50M seed to build expressive AI voice models for creators and enterprises

2 Sources

Share

Fish Audio, which started as a weekend project by a former NVIDIA researcher, has raised $50 million in seed funding led by Coreline Ventures and Capital Today. The Palo Alto-based startup now serves over 8 million users with its library of more than 15,000 natural language controls and generates $21 million in annual recurring revenue.

From Bedroom Project to $50M Voice AI Contender

Fish Audio has secured $50 million in seed funding led by Coreline Ventures and Capital Today, marking a dramatic evolution from its origins as a weekend project. The Palo Alto-based startup, which began when former NVIDIA researcher Shijia Liao trained a voice generation model on a single GPU in his bedroom, now serves over 8 million users and generates $21 million in annual recurring revenue

1

2

. The round also attracted participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0, positioning the company to challenge established players like ElevenLabs in the increasingly competitive voice AI market.

Source: TechCrunch

Source: TechCrunch

Building Expressive and Steerable Voices at Scale

What sets Fish Audio apart in the crowded AI voice generation landscape is its library of more than 15,000 natural language controls that enable developers to create both expressive and steerable voices. The voice AI platform allows users to fine-tune exact tone, inflection, and pacing at the word level, addressing a critical gap in the market where creative use cases demand more expressive outputs while enterprises need steerable systems for customer support and sales automation

1

. The company's flagship S2.1 Pro model demonstrated its technical prowess in blind listening tests, with 67% of listeners preferring its outputs over competitors

2

. Fish Audio can clone a voice from just a five-second audio clip in under 15 seconds, with native support for 83 languages.

Source: SiliconANGLE

Source: SiliconANGLE

Open-Source Roots Drive Rapid Adoption

Fish Audio's open-source strategy has proven instrumental in its growth trajectory. The Fish Speech repository on GitHub has amassed over 31,000 stars, attracting indie developers, video game designers, and content creators desperate for human-like voice synthesis capabilities

1

. Since launching in 2023, the startup has released five AI voice models—four speech generation models and one speech-to-text model—with three speech generation models available as open-source. However, the latest S2.1 Pro model remains exclusive to paid API users, though the company plans to make it available free to all developers starting at the end of August

2

. Organizations like HeyGen, Sanas, and Plaud already use Fish Audio's enterprise APIs, with each requiring different voice characteristics—from realistic AI avatars to expressive game characters and natural-sounding, low-latency voice agents.

Navigating Voice Consent and Creator Trust

Fish Audio's community-driven approach to building its voice library has encountered friction around consent and creator rights. The startup asks users to submit their own voices for training AI voice models, compensating them when their voices are used. However, several months ago, creators alleged their voices were uploaded to Fish Audio without consent, exposing vulnerabilities in the platform's DMCA takedown process

1

. CEO and co-founder Rissa Cao told TechCrunch the company has now automated takedowns, allowing creators to submit a voice sample or contract to prove ownership and remove their voice in less than three minutes. Yet this reactive approach doesn't prevent unauthorized uploads in the first place—voices remain available until artists discover the infringement and file for removal. Osuke Honda, a partner at Coreline Ventures, emphasized that "a community-centric approach can only become a durable advantage if creators trust the platform," advocating for verified voice ownership, clear licensing terms, and revenue-sharing models where creators benefit financially from commercial use

1

.

Enterprise Expansion and Future Model Development

The seed funding will fuel Fish Audio's expansion beyond its creator-focused roots into regulated industries. The company now offers HIPAA-compliant, on-premises deployments with zero-data retention for healthcare and financial services organizations

2

. Cao explained that while the startup was running efficiently with its open-source project and creator plans, it needed capital to develop more advanced models and accommodate enterprise customers as investor interest intensified

1

. Looking ahead, Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model to create a full audio-native stack with voice-native large language models and real-time translation tools

1

2

. The company will also invest in an enterprise sales team and expand developer tooling through API integrations with partners like Retell AI and LiveKit. Rico Mallozzi, a partner at 359 Capital, noted that fine-grained natural language controls for developers and cost-efficient model training will help Fish Audio compete with larger AI labs, praising the team's ability to build state-of-the-art models and close the gap between artificial-sounding and human-like voices with limited resources

1

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved