Richard Sutton 认为强化学习是 AI 的本质,而大语言模型缺乏对世界的真正理解和正确行动的定义。
Richard Sutton argues that reinforcement learning is the essence of AI, while large language models lack a true understanding of the world and a definition of right action.
要点 · TL;DR
从经验中强化学习是智能的核心,而非被动模仿。 Reinforcement learning from experience is the core of intelligence, not passive imitation.
大语言模型缺乏真正的目标和世界模型,限制了其智能。 LLMs lack true goals and world models, limiting their intelligence.
超级智能不可避免;我们应拥抱设计后继者。 Superintelligence is inevitable; we should embrace designing successors.
核心观点 · Key points
强化学习是基本的 AI 范式,通过目标从经验中学习。 Reinforcement learning is the basic AI paradigm, learning from experience with goals.
智能需要目标;下一个词预测不是实质性目标。 Intelligence requires a goal; next-token prediction is not a substantive goal.
真正的学习是主动试错,而非被动模仿或监督学习。 True learning is active trial and error, not passive imitation or supervised learning.
深度学习中的泛化理解不足,常需人工设计。 Generalization in deep learning is poorly understood and often requires human engineering.
由于缺乏全球共识和对超级智能的追求,AI 向数字智能的接替不可避免。 AI succession to digital intelligence is inevitable due to lack of global consensus and pursuit of superintelligence.
反共识 · Contrarian takes
大型语言模型没有世界模型;它们模仿人类文本而不预测真实结果。 Large language models do not have world models; they mimic human text without predicting real outcomes.
模仿学习不是自然学习过程;动物通过预测和试错学习。 Imitation learning is not a natural learning process; animals learn via prediction and trial-and-error.
人类首先是动物;理解松鼠几乎就能理解人类智能。 Humans are animals first; understanding a squirrel gets us almost all the way to human intelligence.
扩展大型语言模型并非苦涩教训;真正的扩展来自从经验中学习。 Scaling large language models is not the bitter lesson; true scaling comes from learning from experience.
我们应为催生设计智能这一重大宇宙转变感到自豪。 We should be proud of giving rise to designed intelligences as a major universal transition.
本期章节 · Chapters(共 18)
引言:RL 与 LLM 视角对比Introduction and RL vs LLM perspectives
LLM 缺乏目标与世界预测LLMs lack goals and world prediction
LLM 与苦涩教训LLMs and the Bitter Lesson
学习目标:模仿 vs 经验Goals in learning: imitation vs. experience
文化知识与模仿Cultural Knowledge and Imitation
Labelbox 广告Labelbox Ad
经验范式Experiential Paradigm
长期奖励学习与 TD 学习Learning from Long-Term Rewards and TD Learning
世界模型与迁移学习World models and transfer learning
AI 进展中的意外Surprises in AI progress
苦涩教训与后 AGI 研究Bitter lesson and post-AGI research
AGI 与超人智能AGI and Superhuman Intelligence
苦涩教训与 AI 协作The Bitter Lesson and AI Collaboration
量化公司与 AI 实验室的保密文化Culture of Secrecy in Quant Firms and AI Labs
超级智能的必然性与继承Inevitability of superintelligence and succession
控制未来 vs 局部目标Control over the future vs. local goals
养育子女的类比Analogy with raising children
设计未来与自愿改变Designing the future and voluntary change