2 Sources
[1]
Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human
While AI agents are increasingly capable of solving customer support problems, most people can still tell immediately when they're talking to a machine instead of a human. Smallest.ai, a startup founded in late 2024, is betting the next leap in voice agents will not come from making large language models faster, but from using smaller, specialized models built for human conversation. Simply put, the company wants to make speaking to an AI agent indistinguishable from talking to a human. To do so, it's developing a small voice model designed to mimic how humans process information by listening, thinking, and speaking simultaneously. "While I'm speaking to you, you're already thinking, and you might interrupt me if I talk for too long," Sudarshan Kamath (pictured left), founder and CEO of Smallest.ai, told TechCrunch, adding that this is exactly how the startup's model is designed to work. To fuel this mission, Smallest.ai has raised $13 million in a Series A round, led by Seligman Ventures with participation from Sierra Ventures and 3one4 Capital. The fresh capital brings the startup's total funding to over $21 million. "The way an LLM works is you give it an entire prompt, and then it starts thinking," Kamath said. While that latency is acceptable in a text chat, in a voice conversation, even a short pause feels unnatural. "If you think about how we are talking, I'm not giving you like a large clipping of my audio, and then you start thinking." The startup's model serves as a real-time intelligence layer that enables natural customer conversations on specific topics, with virtually zero response lag. But if the model encounters a subject outside its limited knowledge base, Smallest.ai hands off the query to a large foundational model, briefly placing the customer on hold to "research" the issue -- just as a real human would do. Kamath believes that all AI agents will soon rely on two models: a small voice model for real-time interaction, and an "offline" LLM that is called upon as needed to solve complex problems. Unlike large foundational models, Smallest.ai focuses strictly on voice-specific nuances, such as handling diverse accents, supporting dozens of languages, and operating in noisy environments. The startup's existing customers include companies in the voice space, including RingCentral and Truecaller. Kamath said that any customer support company, including newer ones like Sierra and Decagon, is a potential customer for the startup. When asked why a well-funded AI customer support company wouldn't build its own voice model, Kamath said that for customer support startups, becoming "extremely good at doing voice is a distraction from their core business." Smallest.ai competes with voice AI leader ElevenLabs, as well as Cartesia and regional players like Sarvam that focus on local languages. While some competitors apply voice AI to use cases, like audio dubbing and podcasting, Smallest.ai focuses strictly on real-time conversational voice agents for its enterprise customers. "We want our models to break the Turing test," Kamath said. "You should speak to our model and not know it's AI or human. That's the sole focus of the company."
[2]
Smallest.ai raises $13M to accelerate the development of its asynchronous voice AI architecture
The momentum behind voice artificial intelligence is accelerating with Smallest.ai becoming the latest startup in this emerging niche to secure more funding. Officially known as Smallest Inc., it said today it has closed on a $13 million Series A investment led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital, which were the main backers in its $8 million seed funding round in October. The round brings Smallest.ai's total amount raised to date to over $21 million. Even more importantly, it chose the occasion to debut a new asynchronous speech-to-speech model called Hydra that's based on its most advanced "Voice 4.0" architecture, designed to make AI conversations more natural, responsive and scalable. The startup highlights the enormous potential of a global voice AI industry that's currently valued at just $2.4 billion annually. According to a study by Market.US, that market is expected to grow to more than $47.5 billion by the end of 2034, at which point it's likely that a substantial percentage of AI interactions will be enabled through voice. But we're not there yet, with AI currently accounting for less than 1% of the world's voice interactions, and there are good reasons for that. Smallest.ai argues that existing voice AI systems simply aren't able to handle the complexity of real-world interactions, which is why we really only see them deployed in very narrow customer service use cases and little else. The problem is that AI voices are still too robotic and struggle with noticeable latency, limiting its usefulness to only a few applications where vast amounts of training data exist. Founder and Chief Executive Sudarshan Kamath told SiliconANGLE that the deficiencies of voice AI stem from the architectural design of speech models. They're built on a chained stack of separate technologies, including speech recognition, large language model processing, orchestration layers, memory systems, text-to-speech engines and guardrails, which must all be cobbled together so that everything can execute, one after another. It's a disjointed process that results in both latency and interactions that feel unmistakably artificial. Smallest.ai's solution to this is Hydra, the foundational model that sits at the heart of Voice 4.0. Kamath said voice AI has undergone a number of evolutionary steps over the years, with the initial wave of Voice 1.0 models enabling rigid, interactive voice response trees that were used in early customer service applications. They were followed by Voice 2.0, which introduced machine learning-powered voice bots that were more flexible but still too rigid and robotic. Then, with Voice 3.0, we saw the first generative AI agents that could interact using voice as a medium, but still struggle with a lack of authenticity and low latency. With Voice 4.0, Kamath said Smallest.ai is ushering in a paradigm shift for the voice AI industry. It's an asynchronous AI architecture that's uniquely able to process listening, reasoning, take actions and respond in parallel, rather than do everything in a sequential way. Hydra allows these functions to take place simultaneously to support real-time conversational flows, natural interruptions and mid-conversation tool use. Besides enabling human-to-machine conversations, it can be paired with the company's earlier speech-to-text models to support more rapid transcription, with latency measured in milliseconds. "Humans don't wait for someone to finish speaking before they begin thinking. We listen, think, and respond simultaneously," Kamath said. "Voice AI needs to work the same way. By rethinking the stack instead of simply scaling models, we're reducing latency to the point where voice interactions feel genuinely human." Hydra is the latest addition to Smallest.ai's growing technology stack. It has also developed speech-to-text models such as Pulse STT Pro and Lightning V3.1, which consistently rank among the highest voice AI systems on the Artificial Analysis benchmark. When the original Lightning model was released last year to coincide with the company's seed funding round, it was described as the fastest text-to-speech model on the market, abe to generate 10 seconds of speech in 100 milliseconds, which corresponds to just a tenth of a second. Smallest.ai has since expanded Lightning to support 38 languages and enrich its capabilities with emotion detection, speaker diarization, data redaction and noise reduction features. It has been deployed by customers including RingCentral Inc., Truecaller AB, Kogtal Financial Ltd. and Readymode Inc. to help reduce customer support costs by as much as 80% in some cases. Today's funding round and launch will help Smallest.ai to keep pace with its rivals in an increasingly crowded field of specialized voice AI startups. Earlier this week, a rival called Fish Audio raked in $52 million from investors to expand adoption of its open-source platform that developers can use to train and fine-tune speech models. But by far and away the best-funded voice AI startup is ElevenLabs Inc., which raised $500 million in February to build out its agentic voice AI platform. Although Smallest.ai's competitors have raised significantly more cash, Seligman Ventures' Ashish Kakran said the potential is so big that there's plenty of room for others to shine, especially if they can make life easier. "Developers now increasingly talk to their machines instead of typing code," he said. "Smallest.ai is taking a fundamentally different approach to the category by rethinking architecture itself. Customers get an efficient vertically integrated stack and don't need to waste time stitching models together."
Share
Copy Link
Smallest.ai secured $13 million in Series A funding led by Seligman Ventures to develop voice AI that sounds genuinely human. The startup unveiled Hydra, an asynchronous speech-to-speech model built on Voice 4.0 architecture, designed to eliminate latency and enable real-time conversational flows. Unlike traditional systems, Hydra processes listening, reasoning, and responding simultaneously rather than sequentially.
Smallest.ai has raised $13 million in Series A funding led by Seligman Ventures, with participation from Sierra Ventures and 3one4 Capital
1
2
. This brings the startup's total funding to over $21 million since its founding in late 20241
. The fresh capital will accelerate development of voice AI models designed to make conversations with AI agents indistinguishable from speaking with humans. Founded by Sudarshan Kamath, Smallest.ai is betting that the next leap in voice AI won't come from making large language models faster, but from using smaller, specialized models built specifically for human conversation1
.
Source: TechCrunch
Coinciding with the funding announcement, Smallest.ai unveiled Hydra, an asynchronous speech-to-speech model built on its most advanced Voice 4.0 architecture
2
. This represents a fundamental shift in how voice AI systems operate. Traditional voice AI relies on a chained stack of separate technologies including speech recognition, large language model processing, orchestration layers, memory systems, text-to-speech engines, and guardrails that must execute sequentially2
. This disjointed process creates noticeable latency and interactions that feel unmistakably artificial. Hydra's asynchronous voice AI architecture processes listening, reasoning, taking actions, and responding in parallel rather than sequentially2
. "Humans don't wait for someone to finish speaking before they begin thinking. We listen, think, and respond simultaneously," Kamath explained. "Voice AI needs to work the same way"2
.
Source: SiliconANGLE
Smallest.ai's approach fundamentally differs from conventional AI agents. The startup's model serves as a real-time intelligence layer that enables natural customer conversations on specific topics with virtually zero response lag
1
. When the model encounters a subject outside its limited knowledge base, it hands off the query to a large foundational model, briefly placing the customer on hold to "research" the issue, just as a real human would do1
. "While I'm speaking to you, you're already thinking, and you might interrupt me if I talk for too long," Kamath told TechCrunch, explaining how the startup's model mimics human information processing1
. Kamath believes all AI agents will soon rely on two models: a small voice model for real-time interaction and an "offline" LLM called upon as needed to solve complex problems1
.Unlike large foundational models, Smallest.ai focuses strictly on voice-specific nuances including handling diverse accents, supporting dozens of languages, and operating in noisy environments
1
. The company addresses a critical problem: while latency is acceptable in text chat, even a short pause in voice conversation feels unnatural1
. Smallest.ai has developed speech-to-text models such as Pulse STT Pro and Lightning V3.1, which consistently rank among the highest voice AI systems on the Artificial Analysis benchmark2
. The original Lightning model, released last year, was described as the fastest text-to-speech model on the market, able to generate 10 seconds of speech in 100 milliseconds2
. Lightning has since expanded to support 38 languages and been enriched with emotion detection, speaker diarization, data redaction, and noise reduction features2
.Related Stories
Smallest.ai's existing customers include companies in the voice space such as RingCentral and Truecaller, along with Kogtal Financial and Readymode
1
2
. These deployments have helped reduce customer support costs by as much as 80% in some cases2
. Kamath said any customer support company, including newer ones like Sierra and Decagon, represents a potential customer for the startup1
. When asked why well-funded AI customer support companies wouldn't build their own voice models, Kamath argued that for customer support startups, becoming "extremely good at doing voice is a distraction from their core business"1
. Smallest.ai competes with voice AI leader ElevenLabs, as well as Cartesia and regional players like Sarvam that focus on local languages1
. While some competitors apply voice AI to use cases like audio dubbing and podcasting, Smallest.ai focuses strictly on real-time conversational voice agents for enterprise customers1
.The global voice AI industry is currently valued at just $2.4 billion annually but is expected to grow to more than $47.5 billion by the end of 2034, according to a study by Market.US
2
. Despite this enormous potential, AI currently accounts for less than 1% of the world's voice interactions2
. Existing voice AI systems simply aren't able to handle the complexity of real-world interactions, which is why they're only deployed in very narrow customer service use cases2
. "We want our models to break the Turing test," Kamath stated. "You should speak to our model and not know it's AI or human. That's the sole focus of the company"1
. Watch for Smallest.ai to expand its language support and deploy Hydra across more enterprise customer support scenarios as it pursues this ambitious goal of creating truly indistinguishable human-like interactions.Summarized by
Navi
31 Jan 2025•Business and Economy

28 Jul 2026•Startups

02 Dec 2025•Startups
