Skip to content
TrackPodcasts
scienceMar 18, 20263:54

The Engines of Our Ingenuity 3363: Reinforcement Learning

About this episode

Episode: 3363 Richard Sutton and reinforcement learning.  Today, reinforcement learning.

Get every episode summarized

Each time Engines of Our Ingenuity publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Transcript ready

40 searchable segments. Every word is indexed and playable.

The Engines of Our Ingenuity 3363: Reinforcement Learning

Engines of Our Ingenuity

0:00
3:54

Full transcript

Engines of Our IngenuityThe Engines of Our Ingenuity 3363: Reinforcement Learning. Machine-transcribed; use the interactive transcript above to jump the player to any line.

This programming is sponsored by the Dan L. Duncan Comprehensive Cancer Care Center at Baylor St. Luke's Medical Center, offering comprehensive cancer care that is compassionate, personalized, and driven by clinical research. More at stluxhealth.org slash cancer. This is the engines of our ingenuity made possible by the friends of KUHF Houston. Today, reinforcement learning. The University of Houston presents the series about the machines that make our civilization run and the people who's ingenuity created them. How do we learn new skills? We try, fail, and adjust. We remember what works and we try it again. We avoid repeating what didn't work. We learn from the consequences of our actions. Psychologists call this reinforcement learning. And this idea has been central to building

machines that can teach themselves. The first reinforcement learning algorithms learned only slowly from experience. They tried random actions and get rewarded only when they succeeded. But this approach has a fatal flaw. It makes it difficult to determine what action should be given credit for success. In 1951, Marvin Minsky built snark, a maze-solving machine that learned through trial and error. It worked, but barely. The problem was timing. Consider teaching a computer to play chess. If the computer wins after 50 moves, which moves decided the game? Crediting moves closer to the game's end ignores the early moves that may have paved the way to victory. Crediting all moves creates too much noise to learn anything useful. Enter Richard Sutton, a psychology PhD student who noticed something crucial about animal behavior. Animals don't wait for final outcomes. They get excited when they expect rewards. Pavlov's dogs salivated when they heard

a dinner bell, not just when eating. Sutton asked if machines could learn in the same way. In his 1984 dissertation Sutton proposed temporal difference learning. The idea was simple. Don't wait for the final outcome, but instead learn from your own changing expectations. At every moment, predict how much future reward you will receive. When something happens that makes the prediction jump, say a chess position goes from 50% chance of winning to 80% chance of winning, use that jump as the learning signal. Thus, you do not need to wait until the chess match ends to evaluate a move. But skeptics ask why learning from your own flawed predictions works better than learning from actual outcomes. The answer came in the 90s when Gerald Tisoro built TD Gammon using Sutton's method. This backgammon program learned purely through self-play and reached a level that rivaled the world's best players. But the deepest validation came from

biology itself. Researchers discovered that cells and monkey brains work as Sutton suggested. These cells signal not when animals receive rewards, but when their predictions of a reward suddenly jumps up. Evolution has solved the credit assignment problem millions of years ago. Today's self-driving cars, game-playing AIs and recommendation systems all trace their lineage to Sutton's insight. And that insight was powerful. To learn, you do not need to know whether you'll succeed. You just need to know that you're headed in the right direction. This is Karesha Yosech at the University of Houston, where we are interested in the way it worked. This programming is sponsored by Metro, introducing the Ride Metro Fair Card. The Ride Metro Fair System replaces the Metro Q Fair Card for bus and rail fares. Details at ridemetro.org

Slash Fair System.

More episodes

More from Engines of Our Ingenuity

View all episodes →