
Get every episode summarized
Each time Seventy3 publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
About this episode
今天的这篇的主题是: Not All LLM Reasoning is Visible in the Chain-of-Thought
Seventy3 是一档借助 NotebookLM 解读前沿论文的播客——不是念摘要,也不是泛泛而谈,而是让 AI 帮你把论文掰开揉碎,说人话。我们蹲在人工智能、大模型、机器人算法、crypto 这几个领域,每期挑一篇值得关注的工作,用对话的方式聊给你听。你可以在通勤路上听,做实验的时候听,也可以当背景音放着,让新知识自然地长进脑子里。
如果你正在做有意思的研究,想让更多人看到——不管你是博士生、独立研究者还是实验室的博士后——把论文发给我们,我们帮你用 AI 做一期深度解读,让你的工作被更多同路人发现。
联系小助手微信:eccstartup(加群/投稿论文)
Summary
AI 安全面临的一个关键问题在于,语言模型是否在其输出 token 中表达了其全部的推理过程。我们展示了一种具体的失效模式(failure mode):前沿模型通过利用在语义上无关的填充 token(filler tokens)来提升在合成推理任务上的性能,从而表现出了“隐形推理”(invisible reasoning)。
我们在三项任务上评估了 13 个前沿语言模型,发现许多模型能从填充 token 中显著获益,准确率提升最高可达 13 个百分点。这种收益取决于使用了哪些 token,且在不同模型间存在差异。我们进一步表明,填充 token 能够使 Claude Opus 4.5 在不牺牲其主要任务准确率的前提下,满足一项隐藏的同余算术(modular arithmetic)约束,这证明了隐形推理可以服务于对思维链(CoT)监控而言完全不可见的目标。
强化学习使 Qwen3-235B 对填充 token 的内容产生了强烈的偏好,但无论是强化学习还是监督微调,都未能产生在测试阶段持续存在的填充 token 收益。我们的结果表明,前沿模型已经在进行不带任何可解释输出 token 痕迹的重大计算(consequential computation)。
原文链接:https://arxiv.org/abs/2607.22925
前往小宇宙评论区与主播互动
Get every episode summarized
Each time Seventy3 publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.
Email me new episodesFree for 3 shows. No card needed.
Hosts & guests
No transcript yet
This episode has not been transcribed. Request it and it moves to the front of the queue.
More episodes



