
technologySep 16, 202558:26pending
Is It Time to Rethink LLM Pre-Training? with Aditi Raghunathan - #747
About this episode
Today, we're joined by Aditi Raghunathan, assistant professor at Carnegie Mellon University, to discuss the limitations of LLMs and how we can build more adaptable and creative models. We dig into her ICML 2025 Outstanding Paper Award winner, “Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction,” which examines why LLMs struggle with generating truly novel ideas. We dig into the "Roll the dice" approach, which encourages structured exploration by injecting randomness at the start of generation, and the "Look before you leap" concept, which trains models to take "leaps of thought" using alternative objectives to create more diverse and structured outputs. We also discuss Aditi’s papers exploring the counterintuitive phenomenon of "catastrophic overtraining," where training models on more data improves benchmark performance but degrades their ability to be fine-tuned for new tasks, and dig into her lab's work on creating more controllable and reliable models, including the concept of "memorization sinks," an architectural approach to isolate and enable the targeted unlearning of specific information.
The complete show notes for this episode can be found at https://twimlai.com/go/747.
Get every episode summarized
Each time The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

How Capital One Delivers Multi-Agent Systems with Rashmi Shetty - #765
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Apr 16, 202654:18failed

The Race to Production-Grade Diffusion LLMs with Stefano Ermon - #764
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Mar 26, 20261:03:18failed

Agent Swarms and Knowledge Graphs for Autonomous Software Development with Siddh...
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Mar 10, 20261:16:14failed

AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka...
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Feb 26, 20261:18:55pending