Jerry Torque discusses the limits and future of scaling pre-training and reinforcement learning in AI.
要点 · TL;DR
扩展预训练和强化学习可预测地提升模型,但泛化能力仍然有限。 Scaling pre-training and RL predictably improves models, but generalization remains limited.
持续学习是实现 AGI 的必要缺失环节,而不仅仅是扩展强化学习。 Continual learning is a necessary missing piece for AGI, not just scaling RL.
AI 最大的风险是退入虚拟世界的反乌托邦,而非灭绝。 The biggest AI risk is a dystopian retreat into virtual worlds, not extinction.
核心观点 · Key points
Scaling(规模扩张)预训练和强化学习可预测地提升模型,但超出训练数据的泛化仍然有限。 Scaling pre-training and RL predictably improves models, but generalization beyond training data remains limited.
强化学习适用于有明确反馈的技能,但难以用于写书等主观任务。 Reinforcement learning works well for skills with clear feedback, but hard for subjective tasks like writing books.
当前模型缺乏从失败中自我修正的能力,这是实现 AGI(通用人工智能)的关键缺失环节。 Current models lack the ability to self-correct from failures, a key missing piece for AGI.
持续学习需要稳健的训练过程,不会崩溃,这与当前脆弱的深度学习不同。 Continual learning requires robust training processes that don't collapse, unlike current fragile deep learning.
专注解释了大多数实验室的成功;Anthropic 在编码方面的领先源于对该领域的持续专注。 Focus explains most lab successes; Anthropic's coding lead comes from sustained focus on that domain.
AI 最大的社会风险不是灭绝,而是退入虚拟世界的反乌托邦。 The biggest societal risk from AI is not extinction but a dystopian retreat into virtual worlds.
反共识 · Contrarian takes
仅靠 Scaling(规模扩张)强化学习可能无法实现 AGI(通用人工智能);持续学习是必要元素。 Scaling RL alone may not lead to AGI; continual learning is a necessary element.
数据驱动的改进导致专业化,但研究突破可以同时提升所有领域。 Data-driven improvement leads to specialization, but research breakthroughs can improve all domains at once.
成功的 AI 应用公司最终应该训练自己的模型,而不仅仅是使用 API。 Successful AI application companies should eventually train their own models, not just use APIs.
优秀的 AI 研究者需要系统工程和理论能力,以及敢于反共识的勇气。 Being a great AI researcher requires both systems engineering and theory, plus the courage to be contrarian.
与编码智能体合作的最佳方式类似于管理初级工程师:深入理解但下放控制权。 The best way to work with coding agents is like managing junior engineers: understand deeply but delegate control.
机器人领域将在 2-3 年内迎来类似 ChatGPT 的时刻,比大多数人预期的要早。 Robotics will have a ChatGPT-like moment in 2-3 years, sooner than most expect.