
technologyMar 3, 202549:29pending
Inside s1: An o1-Style Reasoning Model That Cost Under $50 to Train with Niklas Muennighoff - #721
About this episode
Today, we're joined by Niklas Muennighoff, a PhD student at Stanford University, to discuss his paper, “S1: Simple Test-Time Scaling.” We explore the motivations behind S1, as well as how it compares to OpenAI's O1 and DeepSeek's R1 models. We dig into the different approaches to test-time scaling, including parallel and sequential scaling, as well as S1’s data curation process, its training recipe, and its use of model distillation from Google Gemini and DeepSeek R1. We explore the novel "budget forcing" technique developed in the paper, allowing it to think longer for harder problems and optimize test-time compute for better performance. Additionally, we cover the evaluation benchmarks used, the comparison between supervised fine-tuning and reinforcement learning, and similar projects like the Hugging Face Open R1 project. Finally, we discuss the open-sourcing of S1 and its future directions.
The complete show notes for this episode can be found at https://twimlai.com/go/721.
Get every episode summarized
Each time The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

How Capital One Delivers Multi-Agent Systems with Rashmi Shetty - #765
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Apr 16, 202654:18failed

The Race to Production-Grade Diffusion LLMs with Stefano Ermon - #764
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Mar 26, 20261:03:18failed

Agent Swarms and Knowledge Graphs for Autonomous Software Development with Siddh...
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Mar 10, 20261:16:14failed

AI Trends 2026: OpenClaw Agents, Reasoning LLMs, and More with Sebastian Raschka...
The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)
Feb 26, 20261:18:55pending