
scienceOct 13, 20255:55pending
TD-Gammon: Self-Taught Reinforcement Learning and the Backgammon Breakthrough
About this episode
Gerald Tesoro’s TD-Gammon (early 1990s, IBM) proved that reinforcement learning could reach world-class backgammon by learning from self‑play alone. A small neural network used temporal-difference learning to bootstrap its way toward better play, training on roughly 1.5 million self‑played games with a 3-layer architecture (198 inputs, ~80–160 hidden units, 4 outputs predicting White/Black win with or without a gammon). It barely lost to top players and, in doing so, shifted human strategy (notably the 2-1 opening) and helped spark modern RL breakthroughs that culminated in Deep Q‑Networks and AlphaGo/AlphaZero. The TD error signal also draws a provocative parallel to dopamine-based learning in the brain, suggesting universal principles behind intelligence that transcend systems.
Note: This podcast was AI-generated, and sometimes AI can make mistakes. Please double-check any critical information.
Sponsored by Embersilk LLC
Get every episode summarized
Each time Intellectually Curious publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes
More from Intellectually Curious

GPT-6 Astra: The Autonomous AI Operator Redefining Science and Workflows
Intellectually Curious
Sep 5, 20266:02completed

Claude Commerce: The One-Brain AI Reimagining Digital Shopping
Intellectually Curious
Sep 5, 20265:58completed

Zero-Friction Innovation: AI, Activation Energy, and the Long-Tail Frontier
Intellectually Curious
Sep 4, 20266:51completed

Momentum Exchange Tethers and Orbital Skyhooks
Intellectually Curious
Sep 3, 20266:16completed