
Can AI really hear a call the way a person does?
About this episode
PolyAI just launched Dialog-RSN-1, its first audio-native model, and your host Nikola Mrkšić sat down with its builder, Matt Henderson, to unpack why it’s a game-changer for building voice agents.
Most voice AI either flattens a call into a transcript and loses the audio, or goes fully speech-to-speech and gives up control of the voice. Dialog-RSN-1 does neither. It hears the raw audio directly, decides when to speak, and keeps text-to-speech separate so the voice stays under your control, all in under 300 milliseconds.
Nikola and Matt get into what audio-native really means, how auto-reasoning keeps it fast, and why it beats every other real-time model on quality and quickness. Hear the full episode, and see how PolyAI builds dialog agents that hear the whole call at https://poly.ai?utm_source=youtube&utm_medium=podcast&utm_campaign=podcast&utm_content=podcast
Get every episode summarized
Each time Deep Learning with PolyAI publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Deep Learning with PolyAI

Who's coordinating your army of AI agents?
Deep Learning with PolyAI

Why should CX leaders care about MCP?
Deep Learning with PolyAI

Is word error rate just a vanity metric?
Deep Learning with PolyAI

Is AI the end of large engineering teams?
Deep Learning with PolyAI