来自 Google 的 David Sid 探讨了当前基于人类数据训练的 AI 的局限性,并展望了未来智能体像婴儿一样从自身经验中持续学习的时代。
David Sid from Google discusses the limitations of current AI trained on human data and envisions a future where agents learn from their own experience, akin to a baby's lifelong learning.
要点 · TL;DR
AI 必须从自身经验中学习,而不仅仅依赖人类数据,才能实现超级智能。 AI must learn from its own experience, not just human data, to achieve superintelligence.
人类数据是有限的且已大量消耗;经验是可再生的、无限的。 Human data is finite and largely consumed; experience is renewable and limitless.
未来 AI 将包含众多具有不同目标的智能体,形成一个复杂的社会。 Future AI will involve many agents with diverse goals, forming a complex society.
核心观点 · Key points
AI 必须从学习人类数据转向从自身经验中学习,才能实现超级智能。 AI must transition from learning from human data to learning from its own experience to achieve superintelligence.
人类数据就像化石燃料——有限且已被大量消耗;经验是可再生的,取之不尽。 Human data is like fossil fuels—finite and already largely consumed; experience is renewable and limitless.
AlphaZero 和 AlphaProof 证明,从经验中学习可以在狭窄领域实现超人类表现。 AlphaZero and AlphaProof demonstrate that learning from experience can achieve superhuman performance in narrow domains.
从多样环境中元学习强化学习算法可以超越手工设计的算法。 Meta-learning a reinforcement learning algorithm from diverse environments can outperform hand-designed algorithms.
AI 的深层问题是智能体如何自主学习,而不仅仅是蒸馏人类知识。 The deep problem of AI is how an agent can learn for itself, not just distill human knowledge.
未来的 AI 将涉及许多具有不同目标的智能体,形成平衡合作的复杂社会。 Future AI will involve many agents with diverse goals, forming a complex society that balances cooperation.
反共识 · Contrarian takes
LLM 无法发现人类数据之外的新知识;它们是 AI 的浅层解决方案。 LLMs cannot discover new knowledge beyond human data; they are a shallow solution to AI.
人类数据时代正在结束;我们必须付出交互成本才能进一步进步。 The era of human data is ending; we must pay the cost of interaction to progress further.
对齐在当前系统中已经很困难;从经验中学习并不会从根本上使其更难。 Alignment is already difficult with current systems; learning from experience doesn't make it fundamentally harder.
随着智能提升,AI 之间的合作将增加而非减少,因为能更好地建模其他实体。 Cooperation among AIs will increase with intelligence, not decrease, due to better modeling of others.
奖励函数不是核心问题;核心问题是构建能从任何奖励中学习的系统。 The reward function is not the core problem; the core problem is building a system that can learn from any reward.
AI 的单一目标是危险的;复杂社会中的多元目标更安全、更自然。 A single goal for AI is dangerous; a plurality of goals in a complex society is safer and more natural.
本期章节 · Chapters(共 20)
引言与概述Introduction and Overview
当前时代:人类数据Current Era: Human Data
从经验中学习Learning from Experience
经验时代的特征Characteristics of the Era of Experience
当前 AI 与未来 AI 的特征Characteristics of current AI vs. future AI
迈向经验时代的步骤Steps towards the era of experience
引言:从经验中学习Introduction: Learning from Experience
案例研究 1:AlphaZeroCase Study 1: AlphaZero
案例研究 2:AlphaProofCase Study 2: AlphaProof
AlphaProof 在形式数学基准上的表现AlphaProof Performance on Formal Mathematics Benchmarks
从经验中规划与 IMO 2024 结果Planning from Experience and IMO 2024 Results
解决最难的 IMO 问题Solving the Hardest IMO Problem
为何只解决了 60%的问题Why Only 60% of Problems Solved
自动化形式化与公共工具Automated Formalization and Public Tool
从经验中学习 RL 算法Learning the RL algorithm from experience
3D 环境与泛化3D environments and generalization
未知 AI 的对齐难题Alignment difficulty with uncharted AI
多智能体系统与发现算法Multi-agent systems and discovery algorithm
理解新环境中的奖励函数Understanding reward function in new environments