Modulate Raises $25M to Expand Audio-Native AI Models for Deepfake Detection & Voice Analysis

4 Sources

Share

Boston-based voice AI startup Modulate has raised $25 million in funding led by Future Ventures to expand its audio-native AI models. The company's Velma platform uses over 100 specialized models to analyze raw audio signals for emotion, intent, and deepfake detection, processing more than 10 million hours of audio monthly across fraud prevention, customer experience, and trust and safety applications.

Modulate Raises $25M to Scale Audio-Native AI Models

Boston-based voice intelligence startup Modulate has secured $25 million in new funding led by Future Ventures, with participation from Hyperplane and Lakestar

1

2

. This brings the company's total funding to $60 million

3

. The funding arrives as voice AI becomes a primary interface for artificial intelligence applications, creating challenges that transcription alone cannot solve. Founded in 2017 by MIT physics undergrads Mike Pappas and Carter Huffman, Modulate initially focused on voice modulation for gaming before pivoting to voice-based moderation and now concentrates on detecting AI-generated audio and analyzing conversational intent

1

.

Source: The Next Web

Source: The Next Web

Velma Platform Processes Raw Audio Signals for Deeper Analysis

Modulate's flagship product, the Velma platform, analyzes raw audio signals rather than working from text transcripts, extracting emotion, tone, intent, emphasis, and synthetic speech determination

2

. The platform employs what Modulate calls an Ensemble Listening Model architecture, which selects and combines more than 100 smaller, specialized audio models for each task. This approach proves up to 1,000 times more efficient than using a single large model, according to the company

2

3

. Carter Huffman, co-founder and CEO, explained that emotional analysis goes beyond basic sentiment: "Many times people will be polite even to AI agents or bots and they won't come across as angry, but they'll be very dissatisfied"

1

. Because the company runs smaller AI models, it doesn't require specialized hardware and extensive compute resources, which becomes crucial as token bills increase

1

.

Source: SiliconANGLE

Source: SiliconANGLE

Deepfake Detection Leads Industry Benchmarks

Modulate's deepfake detection capabilities have achieved significant recognition on industry benchmarks. The company's Velma Deepfake Detect model has ranked first on Hugging Face's Speech Deepfake Arena since March, with an equal error rate of 1.1% and 98.9% accuracy on public benchmark data

2

3

. In July, Modulate's transcription models took first place out of 88 entries on Hugging Face's Open ASR Leaderboard

2

. The company states Velma is twice as accurate as general-purpose large language models at identifying actual problems while generating seven times fewer false alarms

2

. Healthcare institutions use these models to defend against deepfake callers impersonating staff, addressing a growing threat as AI-related fraud complaints exceeded 22,000 with losses of more than $893 million according to the FBI's 2025 Internet Crime Report

4

.

Fraud Prevention and Customer Experience Applications

Modulate's AI models now process more than 10 million hours of audio monthly, with over 600 million hours analyzed in total

2

3

. The voice analysis platform serves varied applications across fraud prevention, AI agent supervision, trust and safety, and customer experience

4

. For call centers, Modulate specializes in alerting organizations to possible scams and monitors how AI agents respond to customers to assess call quality while ensuring compliance with regulatory requirements

1

. The company also detects child grooming in voice conversations and offers voice masking to protect staff in high-risk roles

2

. Research from PYMNTS Intelligence shows that 58% of companies with upwards of $1 billion in annual revenue reported encountering AI-generated documents or deepfake-related attacks in the past year

4

.

Source: PYMNTS

Source: PYMNTS

Expansion Plans Target Developer Adoption

The new funding will accelerate Modulate's push to court developers with new SDKs, APIs, and partner integrations, enabling companies building voice products to purchase audio understanding capabilities rather than train their own models

2

. Developers can already access batch transcription through the existing API for $0.03 per hour

2

. The company currently employs 40-45 people and plans to add 10 more employees in coming months to bolster model building, with hiring focused on research, engineering, and developer relations

1

3

. Modulate is also working on increasing its on-premises and on-device deployment capabilities for enhanced privacy

1

. Steve Jurvetson, co-founder of Future Ventures, noted that Modulate has "gained a significant technical lead in audio-native AI" and that demand is spreading beyond gaming into AI agents, security, and customer experience

2

. Huffman emphasized that developers "shouldn't have to rebuild the audio intelligence layer every time they create a new voice experience," positioning Modulate to address the growing need for sophisticated voice AI capabilities across industries

3

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved