mission
manifesto
Voice agents have become much better at sounding human. But sounding human is not the same as holding a human conversation. Teams building real-world agents told us that, on their existing voice stacks, more than 40% of people ended the call within the first 30 seconds.
The failures often seem small: an awkward pause, a missed interruption, the agent talking over the person, a correction it fails to catch, or a background voice that derails the call. Together, these moments compound quickly and create friction that causes the person to disengage or end the conversation.
Most voice agents combine separate systems for listening, reasoning, speaking, and turn-taking. That architecture assumes clean, sequential turns. Human conversation is different: it is a continuous, two-way exchange shaped by interruptions, overlap, pauses, and corrections.
MetaVoice brings those capabilities into one end-to-end duplex speech model. It listens while it speaks, allowing it to handle interruptions and background voices naturally. It reasons directly over speech, including how something was said and what is happening around the speaker. And because these capabilities live in one model, there is no separate ASR, TTS, turn detector or dialogue harness to assemble. Together, this brings voice AI closer to how people actually converse.
Developers define the workflow, tools, guardrails, and personality. MetaVoice handles the conversation.
We are starting with revenue calls, where better conversations drive revenue.
Our ambition is larger: to make voice the preferred interface between people and AI.
our team
Built & commercialised frontier AI at
Researched at
our values
01Make something people want
02First-principles thinking
03Velocity over certainty
04Relentless iteration
05Radical honesty, low ego
backed by