
About this episode
Neel and his team are trying to do something phenomenally difficult: understand an intelligence that didn't come with a manual. Together, they explore the cutting-edge "neuroscience" of artificial intelligence—revealing the surprising, elegant structures being discovered inside these networks (like spare autoencoders), the inherent limits of looking under the hood, and why interpretability is absolutely essential if we are to build safe, aligned and trustworthy AI as we move towards AGI. Learn more about this area of research via https://deepmind.google/
Timecodes
- 00:00 Introduction
- 02:41 Motivation for interpretability research
- 04:01 Mechanistic interpretability
- 08:14 Chain of thought monitoring
- 18:14 Interpretability techniques
- 35:00 Auditing models for safety
- 48:53 What comes next for interpretability
Please leave us a review on Spotify or Apple Podcasts if you enjoyed this episode. We always want to hear from our audience whether that's in the form of feedback, new idea or a guest recommendation!
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Get every episode summarized
Each time Google DeepMind: The Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Google DeepMind: The Podcast

The mathematics of AI uncertainty
Google DeepMind: The Podcast
When millions of AI agents meet
Google DeepMind: The Podcast

10 Years of AlphaGo: The Turning Point for AI | Thore Graepel & Pushmeet Kohli
Google DeepMind: The Podcast

The Future of Intelligence with Demis Hassabis (Co-founder and CEO of DeepMind)
Google DeepMind: The Podcast