The startup is building voice models designed to make AI phone calls pass the Turing test.

While AI agents are increasingly capable of solving customer support problems, most people can still tell immediately when they’re talking to a machine instead of a human. Smallest.ai, a startup founded in late 2024, is betting the next leap in voice agents will not come from making large language models faster, but from using smaller, specialized models built for human conversation. Simply put, the company wants to make speaking to an AI agent indistinguishable from talking to a human. To do so, it’s developing a small voice model designed to mimic how humans process information by listening, thinking, and speaking simultaneously. “While I’m speaking to you, you’re already thinking, and you might interrupt me if I talk for too long,” Sudarshan Kamath (pictured left), founder and CEO of Smallest.ai, told TechCrunch, adding that this is exactly how the startup’s model is designed to work. To fuel this mission, Smallest.ai has raised $13 million in a Series A round, led by Seligman Ventures with participation from Sierra Ventures and 3one4 Capital. The fresh capital brings the startup’s total funding to over $21 million. “The way an LLM works is you give it an entire prompt, and then it starts thinking,” Kamath said. While that latency is acceptable in a text chat, in a voice conversation, even a short pause feels unnatural. “If you think about how we are talking, I’m not giving you like a large clipping of my audio, and then you start thinking.” The startup’s model serves as a real-time intelligence layer that enables natural customer conversations on specific topics, with virtually zero response lag. But if the model encounters a subject outside its limited knowledge base, Smallest.ai hands off the query to a large foundational model, briefly placing the customer on hold to “research” the issue — just as a real human would do. Kamath believes that all AI agents will soon rely on two models: a small voice model for real-time interaction, and an “offline” LLM that is called upon as needed to solve complex problems.