START: Aoden Teo, CEO & Co-Founder, Miso Labs: "The most emotive foundation models for voice"

By
·
August 27, 2026

Voice AI can pass a Turing test. For about a minute.

That's a generated clip, though. Have a human actually talk back and the number collapses to six or seven seconds, roughly where generated voice sat three years ago.

One reason, per Aoden Teo of Miso Labs: real conversation isn't turn-based. Around 20% of the time more than one person is speaking, and laughter drives a lot of that overlap, since you laugh at a joke while it's still being told. We also adjust our pacing toward whoever we're talking to without noticing we're doing it.

Voice models struggle with all of this. Full-duplex voice, where a model listens and speaks at the same time, is still extremely early.

So an agent can know your joke is funny and still have to wait until you've finished before it laughs, by which point the timing has killed it.

Aoden describes a second consequence: agents get pushed toward almost "psychotically emotive" behavior. If they can only talk once you've stopped, they need some other way to show they were listening. You finish your sentence, and the thing goes "Hmm?" You've heard it.

Underneath that sits an architecture problem. Voice models have to respond fast, which constrains how large they can be, and fast means something different here than it does in text. Working with an LLM like Claude, Aoden points out, you care how quickly it finishes your code, more than how quickly it starts.

Voice inverts that. Nobody needs 10 hours of audio generated in two seconds, because nobody can listen to 10 hours of audio in two seconds; what matters is reaction time. Most architectural decisions trade latency against throughput, and Aoden expects voice to keep moving away from LLM-style designs toward ones built around very low latency.

Miso is already pushing on it. Miso-1 got 3,000 stars on GitHub and 5 million views on Twitter, and they record data in their own LA studio because the internet doesn't contain every kind of audio a voice model might need. Nobody has released a podcast of someone reading millions and millions of email addresses, and people still want voice models that can read email addresses aloud, so teams end up generating some very strange training data themselves.

The clip isn't the hard part. The hard part starts when you talk back.}

🎙️Aoden Teo, CEO & Co-Founder, Miso Labs on Fondo START  

1:03 Miso-1: 3K+ GitHub stars + 5M X views

1:59 Why emotiveness matters for games, UGC + interactive products

3:06 Measuring progress in voice AI with longer Turing tests

4:01 Why interactive conversation is harder than generating convincing clips

5:08 Full-duplex voice, interruptions + why laughter matters

6:04 Latency vs. throughput - and why voice differs from LLMs

7:09 Miso's LA recording studio + the challenge of voice training data

9:02 Talking teddy bears, UGC, anime + unexpected voice AI use cases

10:19 From serious chess player to math obsession to building @MisoLabsAI

12:11 The surprise YC interview


Check out
misolabs.ai