Skip to content
TrackPodcasts
scienceSep 30, 202523:30pending

Language Models Can Learn from Verbal Feedback Without Scalar Rewards

About this episode

🤗 Upvotes: 48 | cs.CL, cs.AI, cs.LG

Authors:
Renjie Luo, Zichen Liu, Xiangyan Liu, Chao Du, Min Lin, Wenhu Chen, Wei Lu, Tianyu Pang

Title:
Language Models Can Learn from Verbal Feedback Without Scalar Rewards

Arxiv:
http://arxiv.org/abs/2509.22638v1

Abstract:
LLMs are often trained with RL from human or AI feedback, yet such methods typically compress nuanced feedback into scalar rewards, discarding much of their richness and inducing scale imbalance. We propose treating verbal feedback as a conditioning signal. Inspired by language priors in text-to-image generation, which enable novel outputs from unseen prompts, we introduce the feedback-conditional policy (FCP). FCP learns directly from response-feedback pairs, approximating the feedback-conditional posterior through maximum likelihood training on offline data. We further develop an online bootstrapping stage where the policy generates under positive conditions and receives fresh feedback to refine itself. This reframes feedback-driven learning as conditional generation rather than reward optimization, offering a more expressive way for LLMs to directly learn from verbal feedback. Our code is available at https://github.com/sail-sg/feedback-conditional-policy.

Get every episode summarized

Each time Daily Paper Cast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

No transcript yet

This episode has not been transcribed. Request it and it moves to the front of the queue.

Language Models Can Learn from Verbal Feedback Without Scalar Rewards

Daily Paper Cast

0:00
23:30

More episodes

More from Daily Paper Cast

View all episodes →