Jared Kaplan 讨论预训练和强化学习中的扩展定律如何推动 AI 模型的可预测改进,迈向人类级 AI。
Jared Kaplan discusses how scaling laws in pre-training and reinforcement learning drive predictable improvements in AI models, leading towards human-level AI.
要点 · TL;DR
在预训练和强化学习中扩展计算量可预测地提升 AI 性能。 Scaling compute in pre-training and RL predictably improves AI performance.
AI 任务时长每 7 个月翻倍,使得更长的任务成为可能。 AI task horizon doubles every 7 months, enabling longer tasks.
通用人工智能的关键要素包括组织知识、记忆和监督。 Key ingredients for AGI include organizational knowledge, memory, and oversight.
核心观点 · Key points
在预训练和强化学习中扩展算力可预测地提升 AI 性能。 Scaling compute in pre-training and RL predictably improves AI performance.
AI 任务时间跨度大约每 7 个月翻倍,使更长任务成为可能。 AI task horizon doubles roughly every 7 months, enabling longer tasks.
实现 AGI 的关键要素包括组织知识、记忆和可扩展监督。 Key ingredients for AGI include organizational knowledge, memory, and oversight.
缩放定律已在五个数量级上成立,表明进步将持续。 Scaling laws have held over five orders of magnitude, suggesting continued progress.
人机协作目前是高级任务中最有趣的领域。 Human-AI collaboration is currently the most interesting area for advanced tasks.
AI 跨领域的知识广度为研究提供了独特价值。 AI's breadth of knowledge across domains offers unique value for research.
反共识 · Contrarian takes
缩放定律失效很可能意味着训练出错,而非缩放本身终结。 Scaling laws failing likely means we screwed up training, not that scaling is over.
大部分价值可能来自前沿模型,而非更便宜但能力较弱的模型。 Most value may come from frontier models, not cheaper less capable ones.
AI 的判断与生成能力比人类更接近。 AI's judgment and generative capability are closer than in humans.
由于 AI 快速进步,建议构建尚未完全可行的产品。 Building products that don't quite work yet is recommended due to rapid AI improvement.
AI 集成的瓶颈是发展速度,而非能力不足。 AI integration bottleneck is speed of development, not lack of capability.
由于 AI 持续快速进步,均衡状态可能永远不会到来。 The equilibrium state may never arrive as AI keeps improving rapidly.
本期章节 · Chapters(共 21)
引言与背景Introduction and Background
AI 训练的两个阶段Two Phases of AI Training
预训练的缩放定律Scaling Laws for Pre-training
强化学习的缩放定律Scaling Laws for Reinforcement Learning
RL 与预训练中的缩放Scaling in RL and Pre-training
AI 能力的两个维度Two Axes of AI Capabilities
未来趋势:长任务与组织工作Future Trajectory: Longer Tasks and Organizational Work
人类级 AI 的要素:知识、记忆、监督Ingredients for Human-Level AI: Knowledge, Memory, Oversight
为未来准备:构建、整合、采用Preparing for the Future: Build, Integrate, Adopt
引言与 Claude 4 发布Introduction and Claude 4 Release
激动人心的功能与记忆Exciting Features and Memory
人机协作与自动化Human-AI Collaboration and Automation
人机循环的愿景Vision for Human-AI Loop
利用 AI 的广博知识Leveraging AI's breadth of knowledge
物理背景与缩放定律Physics background and scaling laws
缩放定律变化的实证迹象Empirical signs of scaling law changes
缩放定律与失败Scaling Laws and Failures
给早期职业听众的建议Advice for Early-Career Audience
关于缩放定律的听众提问Audience Question on Scaling Laws
通过 RL 与验证信号进行缩放Scaling via RL and verification signals
为 RL 创建任务:AI vs 人类Creating tasks for RL: AI vs humans