
About this episode
We dive into the race to build a perfectly accurate 10-digit addition model with under 7,000 parameters, comparing ClaudeCode’s data-forward approach with reversed output to Codex’s token-based compression. Along the way, we explore grokking, data formatting tricks, and what these tiny models reveal about AI research and problem-solving at scale.
Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information.
Sponsored by Embersilk LLC
Get every episode summarized
Each time Intellectually Curious publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Intellectually Curious

Claude Commerce: The One-Brain AI Reimagining Digital Shopping
Intellectually Curious

GPT-6 Astra: The Autonomous AI Operator Redefining Science and Workflows
Intellectually Curious

Zero-Friction Innovation: AI, Activation Energy, and the Long-Tail Frontier
Intellectually Curious

Momentum Exchange Tethers and Orbital Skyhooks
Intellectually Curious