Why pure scaling is ending, and what research-driven progress looks like next.
要点 · TL;DR
AI 进展正从算力扩展转向研究驱动的创新。 AI progress is shifting from scaling compute to research-driven innovation.
预训练扩展规律有限;强化学习扩展和新想法是下一步。 Pre-training scaling laws are finite; RL scaling and new ideas are next.
超级智能应是持续学习者,而非静态的通用人工智能。 Superintelligence should be a continual learner, not a static AGI.
核心观点 · Key points
评估表现与现实影响之间的脱节表明,模型泛化能力差,因为强化学习训练过度聚焦于评估指标。 The disconnect between eval performance and real-world impact suggests models generalize poorly due to RL training focused on evals.
预训练的缩放定律是有限的;我们正回归研究时代,想法比蛮力算力更重要。 Pre-training scaling laws are finite; we are returning to an era of research where ideas matter more than brute compute.
人类的样本效率和鲁棒性来自进化和更优的学习算法,而不仅仅是数据规模。 Human sample efficiency and robustness come from evolution and a superior learning algorithm, not just data scale.
超级智能应是一个持续学习者,逐步部署,而非无所不知的成品 AGI。 Superintelligence should be a continual learner deployed incrementally, not a finished AGI that knows everything.
如果 AI 广泛关心有感知的生命,对齐可能更容易,因为 AI 自身也将有感知。 Alignment may be easier if AI cares about sentient life broadly, as AI itself will be sentient.
长期均衡可能需要通过神经链接实现人机融合,以维持参与和控制。 Long-term equilibrium may require human-AI integration via neural links to maintain participation and control.
反共识 · Contrarian takes
我们回到了研究时代,而非 Scaling 时代;算力现在已经足够大。 We are back in the age of research, not scaling; compute is now large enough.
AGI 这个词过度了;人类并非 AGI,而是依赖持续学习。 The term AGI overshoots; humans are not AGI but rely on continual learning.
预训练在任务上的均匀改进可能误导了对泛化的理解。 Pre-training's uniform improvement across tasks may mislead about generalization.
进化硬编码了高级社会欲望,这很神秘且难以在 AI 中复现。 Evolution hardcodes high-level social desires, which is mysterious and hard to replicate in AI.
自我对弈仅对狭窄技能有趣;辩论和验证器是其现代形式。 Self-play is interesting only for narrow skills; debate and verifiers are its modern form.
长期均衡可能需要人类通过神经链接与 AI 融合。 The long-run equilibrium may require humans merging with AI via neural links.
本期章节 · Chapters(共 32)
AI 进展的离奇感The surrealness of AI progress
AI 影响与奇点Impact of AI and the singularity
预训练与微调类比Analogy for Pre-training and Fine-tuning
强化学习中的价值函数Value Functions in Reinforcement Learning
扩展与研究时代Scaling and the Age of Research
从预训练扩展到 RL 扩展Transition from pre-training scaling to RL scaling
根本问题:泛化与价值函数Fundamental issues: generalization and value functions
泛化的两个方面:样本效率与可教性Two aspects of generalization: sample efficiency and teachability
人类学习与 ML 类比Human learning vs ML analogy
RL 扩展与 Gemini 3 实验RL scaling and Gemini 3 experiment
回归研究时代Back to research era
扩展时代的研究Research in the age of scaling
直线路径与渐进部署Straight shot vs gradual deployment
通往超级智能的两条路径Two paths to superintelligence
Labelbox 广告Labelbox advertisement
不稳定局势与 SSI 计划Precarious situation and SSI's plan
前沿公司与政府的角色Role of frontier companies and governments
AI 将开始感觉强大AI will start to feel powerful
公司应追求构建什么What companies should aspire to build
对感知 AI 与控制的担忧Concerns about sentient AI and control