
040: The multilayered anatomy of a generative AI voice assistant
About this episode
It's getting technical up in here! Kylie welcomes PolyAI CTO & Co-Founder Shawn Wen to the show to discuss the complexities of building high-quality voice AI systems. From the limitations of large language models (LLMs) to the nuances across each layer of the tech stack, they'll cover setting up audio channels, streaming speech recognition, voice activity detection, and the uniquely global challenge of handling different accents and languages. Shawn gives away a bit of our secret sauce, while underscoring the importance of adapting technology to user characteristics and enterprise demands.
Get every episode summarized
Each time Deep Learning with PolyAI publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Deep Learning with PolyAI

What happens if your AI agents get 1% better every day?
Deep Learning with PolyAI

Who's coordinating your army of AI agents?
Deep Learning with PolyAI

Why should CX leaders care about MCP?
Deep Learning with PolyAI

Can AI really hear a call the way a person does?
Deep Learning with PolyAI