Skip to content
TrackPodcasts
technologyDec 11, 202412:40pending

(Voiceover) OpenAI's Reinforcement Finetuning and RL for the masses

Interconnects

About this episode

Original post:

https://www.interconnects.ai/p/openais-reinforcement-finetuning

Chapters

00:00 Introduction

04:19 The impact of reinforcement finetuning’s existence

07:29 Hypotheses on reinforcement finetuning’s implementation

Figures

Fig. 1, Yann’s Cake

Fig. 2, Grader config

Fig. 3, RLVR learning curves



This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit www.interconnects.ai/subscribe

Get every episode summarized

Each time Interconnects publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

(Voiceover) OpenAI's Reinforcement Finetuning and RL for the masses

Interconnects

0:00
12:40

More episodes

More from Interconnects

View all episodes →