从人类数据到经验:AI 的下一个时代

From Human Data to Experience: The Next Era of AI

大卫·西尔弗 David Silver · 公开讲座 · 2025-11-14 · 约 46 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

来自 Google 的 David Sid 探讨了当前基于人类数据训练的 AI 的局限性,并展望了未来智能体像婴儿一样从自身经验中持续学习的时代。

David Sid from Google discusses the limitations of current AI trained on human data and envisions a future where agents learn from their own experience, akin to a baby's lifelong learning.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

引言与概述 Introduction and Overview

Host

来自谷歌的 David Sid,David 将与我们分享经验领域。谢谢。

David Sid from Google and so David will share with us the area of experience. Thank you.

David

非常感谢。很高兴来到这里。这是一个很好的研讨会主题,我很乐意谈谈我对未来几年 AI 发展的大局观。然后我会具体讲一些与之相关的最新工作。时机很好,因为我要讨论两篇《自然》论文:一篇几周前发表,另一篇昨天刚出来。虽然工作本身并不完全是新的。

Thank you very much. It's a pleasure to be here. It's a great topic for a workshop and I'm happy to tell you about the bigger picture of how I see the future of AI unfolding as we move into the next years. Then I'll talk specifically about some recent work that connects to that. It's very timely because there are two Nature papers I'm discussing: one appeared a couple of weeks ago and one came out yesterday. Although the work itself is not completely new.

当前时代:人类数据 Current Era: Human Data

David

我想先谈谈我们当前所处的 AI 时代。最近 AI 的大部分工作都完全属于我所说的人类数据时代。AI 主要从互联网上的人类数据中学习——所有自然语言都被提炼成一个巨大的模型——然后通过人类反馈进行微调。人类提供良好行为的示例和偏好来帮助微调这些系统。这带来了进步:我们有了知识面广泛的 LLM,在法律、数学、软件方面表现惊人。但如果我们将目光投向 AI 的终极目标,这种 AI 无法带我们走完全程。它无法实现超级智能,而这是我职业生涯一直追求的目标。

I want to start by talking about where we are in this current era of AI. Most work in AI recently sits squarely in what I call the era of human data. The AI learns primarily from human data trained from the internet—all natural language distilled into a giant model—and then it's fine-tuned from human feedback. Humans give examples of good behavior and preferences to help fine-tune these systems. This has led to progress: we have LLMs with wide-ranging knowledge, amazing performance in law, math, software. But if we keep our eyes on the grand prize of AI, this kind of AI won't take us all the way. It can't deliver superintelligence, which is the goal I've been working towards.

David

为什么做不到?因为仅从人类知识训练的系统无法发现人类数据中不存在的新知识或新范式。它没有回答 AI 最深层的科学问题:智能体如何自主学习?这就是为什么你们都在这儿——强化学习、随机控制——这些方法试图触及系统如何自主学习的问题。这非常重要。

Why can't it do this? Because a system trained solely from human knowledge cannot discover new knowledge or paradigms not already in human data. It hasn't answered the deepest scientific question of AI: how can an agent learn for itself? That's why you're all here—reinforcement learning, stochastic control—these methods try to access this question of how a system can learn for itself. That's really important.

从经验中学习 Learning from Experience

David

为了对比人类数据时代,想想从经验中学习。我喜欢展示这个视频,一个精力充沛的婴儿在玩耍。它对物体着迷,给自己设定子目标,玩东西,试图理解它们的行为。它不断用好奇心和适应能力挑战自己,学习新技能。这个过程持续多年,获得的技能让大脑能够成为神经外科医生、舞者、工匠、钢琴家、科学家或网球运动员。我们应该问如何构建具有这种惊人学习能力的系统,能够无限地发现新技能和知识。

To contrast the era of human data, think about learning from experience. I like to show this video of a baby with a lot of stamina playing. It becomes fascinated by objects, sets itself sub-goals, plays with things, tries to understand their behavior. It constantly challenges itself with curiosity and the ability to adapt and learn new skills. This process goes on for many years, and the skills acquired allow the mind to become a neurosurgeon, dancer, craftsman, pianist, scientist, or tennis player. We should ask how to build systems with this amazing capacity to learn and discover new skills and knowledge ad infinitum.

David

这把我们带向我所说的经验时代。我认为这是一个不可避免的转变。为了让 AI 远远超越当前水平,智能体必须开始从自身经验中学习——与环境互动并从这些互动中学习。这是一个必要的先决条件。有一天这将在巨大规模上实现,可能让互联网显得渺小。当这发生时,将是变革性的,因为持续发现新知识和能力的能力,就像那个婴儿但可能远超,带来不可思议的智能形式。

This takes us toward what I call the era of experience. I see this as an inevitable transition. For AI to progress far beyond where we are now, agents must start to learn from their own experience—interacting with their environment and learning from those interactions. This is a necessary prerequisite. One day this will be done at vast scales, maybe making the internet seem tiny. When this happens, it will be transformative because of the continual ability to discover new knowledge and capabilities, like that baby but maybe far beyond, leading to incredible forms of intelligence.

经验时代的特征 Characteristics of the Era of Experience

David

经验时代有哪些特征?像视频中的婴儿一样,未来的智能体将拥有漫长的经验流,有终生学习的能力。这些流将漫长而复杂,使他们能够获得更多技能和知识。行动和观察将扎根于环境——不是与 LLM 的类人对话,而是智能体实际做某事并根据后果接收反馈。奖励将基于经验,而非人类判断。最后,这些智能体将规划或推理经验本身。

What are some characteristics of the era of experience? Like the baby in the video, future agents will inhabit long streams of experience, have lifetimes to keep learning. These streams will be long and complex, allowing them to pick up more skills and knowledge. Actions and observations will be grounded in the environment—not human-like dialogue with an LLM, but the agent actually doing something and receiving feedback based on consequences. Rewards will be grounded in experience, not human judgments. Finally, these agents will plan or reason about the experience itself.

当前 AI 与未来 AI 的特征 Characteristics of current AI vs. future AI

David

我认为这四个特征都与当今的人工智能系统截然不同。LLM 没有做任何这些事情。因此,一旦我们接受这种不可避免的转变并开始花时间探索它,就有巨大的机会超越我们目前的状态。

I think these four characteristics are all very different from where today's AI systems are. LLMs aren't doing any of these things. And so there's huge opportunity to progress beyond where we are once we kind of embrace this inevitable transition and start to spend our time exploring it.

Host

那么你可能会问,我们几年前不就已经达到这个状态了吗?我们过去构建过系统,比如 Atari DQN、AlphaGo 和 AlphaZero。这些是我们做过的一些事情,过去还有许多其他使用强化学习取得的重大突破。

So then you might ask, well, you know, weren't we here already a few years ago? We built systems in the past, things like Atari DQN and AlphaGo and AlphaZero. These are some things that we did, and there's many other great breakthroughs that happened using reinforcement learning in the past.

David

我认为这些非常成功,并引发了人们对基于经验的方法的兴趣激增。这里的 y 轴是 AI 社区对强化学习的关注度。所以你不应该认为这一定是 LLM 不好而这种方法好。这只是说当时有很多关注,发生了令人兴奋的事情,但那时它们通常局限于基于模拟的狭窄领域。

I think those were immensely successful and led to a kind of surge of interest in methods based on experience. On the y-axis here is the attention of the AI community on reinforcement learning. So you shouldn't think of this as necessarily that LLMs are bad and that this is good. It's just saying there was a lot of attention at that time and exciting things happened, but at that time they were quite restricted in general to narrow domains that were based in simulation.

David

然后 GPT-3 和 ChatGPT 之类的东西出现了,人们变得非常兴奋,因为这些系统似乎能够处理与现实世界知识相关的事情,人们可以与它们互动,询问一些不仅仅是狭窄模拟中的事情。所以这看起来非常令人兴奋。但与此同时,我认为公平地说,我们最终把婴儿连同洗澡水一起倒掉了:这些系统不再有能力从经验中自己发现知识。

Then GPT-3 and ChatGPT and things like this came along, and people became very excited because these could seemingly do things that engaged with real world knowledge, and people could interact with them and ask about things which weren't just part of some narrow simulation. So that seemed very exciting. But at the same time, I think it's fair to say that we ended up throwing the baby out with the bathwater: these systems no longer have the ability to discover knowledge for themselves from experience.

David

所以我认为剩下的是这个图的右侧,即不可避免地会重新燃起对基于经验的方法的兴趣。我们开始看到一些例子,其中一些我会谈到,基于经验的方法正在处理复杂的开放式挑战领域,以及至少在现实世界的数字部分中的复杂交互用途。我认为随着这种方法的被接受,它将带来可能无限的表现。

So I think what remains is the right-hand side of this figure, which is that there will be a renewed, inevitably, interest in methods based on experience. We're starting to see some examples of that, some of which I'll talk about, where methods based on experience are tackling complex open-ended challenge domains and complex interactive uses in at least the digital part of the real world. And I think as that gets embraced, it will lead to potentially unbounded performance.

David

所以到目前为止的经验都集中在左边的这个路灯下;它一直专注于人类数据的时代。到目前为止,只有非常少量的经验集中在最有潜力的事情上。所以我想,如果你从这次演讲中什么也没带走,我鼓励人们开始看看第二个路灯。我认为这是一群真正理解那个路灯是什么样的人,我们真的需要拥抱那个路灯,看看那里能做什么。

So the experience so far has been under this lamppost on the left; it's been focused on the era of human data. So far, a very small amount of experience has been focused on the thing which has the greatest potential. So I guess if you take nothing else away from this, I'd encourage people to start looking under that second lamppost. I think this is the kind of group of people that really understand what that lamppost looks like, and we really need to be embracing that lamppost and seeing what can be done there.

David

另一种讨论方式是,从发生的事情来看:人类数据在某种意义上提供了一条捷径。所以我们没有解决 AI 的深层问题;我们解决了 AI 的一个浅层问题,即把人类数据的知识蒸馏到机器中的想法。我喜欢借用一个很好的类比:我们发现了化石燃料。互联网就像 AI 的化石燃料。地下已经有所有这些惊人的东西,AI 能够利用这些化石燃料并加以利用。所以这些 LLM 基本上一直在利用化石燃料来预训练这些惊人的智能体。而且这非常便宜;它比实际出去与世界互动要便宜得多。你可以免费收集所有这些经验。但化石燃料的问题是它们会耗尽。我认为人类数据的时代就是这样:大部分好的数据已经被消耗了。它们已经被训练过了。我们的系统已经训练过所有曾经写过的论文等等。所以现在要更进一步,我们必须解决 AI 的深层问题。我们需要找到可再生能源,那种能让我们永远学习的东西,而这种可再生能源就是经验。如果你有经验,并且从自己的互动中不断学习,那是没有尽头的;你可以一直继续下去,系统有一天将能够不断与环境互动并学习越来越多。所以从这个意义上说,它是可再生的。

Another way to talk about this is to say in terms of what happened: human data provides a shortcut in some sense. So we didn't solve the deep problem of AI; we solved a shallow problem of AI, this idea of distilling the knowledge of human data into the machine. A nice analogy I like to borrow is this idea that we found fossil fuels. The internet is like the fossil fuel of AI. There was all this amazing stuff that was just already there in the ground, and AI was able to tap into those fossil fuels and make use of it. So these LLMs have basically been earning the fossil fuels to pre-train these amazing agents. And that's been really cheap; it's far cheaper than actually going out and interacting with the world. You can gather all of this experience for free. But the problem with fossil fuels is that they run out. I think that's what's happened with the era of human data: most of the good data has been consumed. It's been trained upon. We have systems that have literally trained on every paper that's ever been written, etc. So now to go further, we have to solve this deep problem of AI. We need to find the renewable energy, the thing which will allow us to keep learning forever, and that renewable energy is experience. If you have experience and you keep learning from your own interactions, there's no end to that; you can just keep going and going and going, and systems one day will be able to just keep interacting with their environments and learning more and more and more. So in that sense, it's renewable.

David

然后你可能会问:这种通用性从何而来,它对于像 ChatGPT 这样的东西如此有用,它们有一种非常通用的方式来谈论几乎任何事情?我认为答案是世界提供了通用性。世界内部包含了大量的信号;我们可以做和互动的不同事物太多了,世界具有巨大的通用性潜力。所以如果我们构建能够优化所有这些信号的系统,这些系统将极其通用。但你必须付出互动的代价。我认为这就是重点:捷径将会耗尽,我们必须付出经验的代价。你必须构建真正与世界互动的系统。如果我们不这样做,并不是说我们可以按某个魔法按钮,让系统神奇地理解如何发现新知识。这不会发生,直到我们付出代价,构建真正与环境互动的系统,然后我们才能解决 AI 的深层问题。

Then you can ask: where does this generality come from that was so helpful with things like ChatGPT, where they had this very general way to talk about almost anything? I think the answer is that the world provides the generality. The world contains this vast array of signals within it; there are just so many different things that we can all do and interact with, and the world has vast potential for generality. So if we just build systems that can optimize any and all of those signals, these systems will be extremely general. But you then have to pay the cost of interaction. I think that's the point here: the shortcuts are going to run out, and we have to pay the cost of experience. You have to build systems that actually interact with the world. If we don't do that, it's not like there's some magic button we can press that allows systems to magically understand how to discover new knowledge. It's just not going to happen until we pay the cost of building systems that actually interact with our environment, and then we'll be able to solve the deep problem of AI.

迈向经验时代的步骤 Steps towards the era of experience

David

所以现在我想谈谈迈向经验时代的一些步骤。其中一些是大家非常熟悉的成功案例。我们见过在棋盘游戏中表现出色或达到超人水平的系统。我们见过能够互动的系统,机器人系统解决了非常灵巧的操作问题。我们见过 AI 玩视频游戏。我们还见过 AI 处理数学之类的问题,我稍后会谈到。

So I wanted to talk now about some steps towards the era of experience. Some of these are success stories everyone's very familiar with. We've seen systems that have played board games really well or achieved superhuman capability in board games. We've seen systems that have been able to interact, robotic systems that have solved really nice dexterous manipulation problems. We've seen AI play video games. And we've seen AI deal with things like mathematics that I'll be talking about shortly.

引言:从经验中学习 Introduction: Learning from Experience

David

所以我想在这里强调的一点是,当人们推动从经验中学习时,它已经成功了。在这些案例中,当人们坚定地、有决心地、相信它会成功,并投入大量资源时,不可思议的事情发生了。我们确实构建了在这些有限领域内变得超级智能的系统。所以我认为目前真正缺乏的就是那种信念。人们只在路灯下寻找。所以这些成功故事至少应该提醒我们,只要有信念,成功是可能的。所以我现在要讲三个我有幸参与过的案例研究。

So I think the point I want to make here is that when people have pushed on learning from experience it has succeeded. In these cases, when people have pushed hard with determination and conviction that it's going to work and put a lot of resources into these things, incredible things have happened. We've literally built systems that have become superintelligent in these limited spheres. So I think really what's lacking at the moment is just that kind of conviction. People are looking under the lamppost. So these success stories should, if nothing else, remind us that given conviction, success is possible. So I'm going to talk about three case studies now which I've been lucky enough to work on.

案例研究 1:AlphaZero Case Study 1: AlphaZero

David

第一个案例发生在很久以前,但我只想简单提醒大家一下,因为其他案例将建立在这个基础上。那就是 AlphaZero 算法。我必须说,当我回顾我的职业生涯时,我认为这是我最引以为豪的事情。我之所以最引以为豪,是因为它感觉像魔法一样。它拥有从经验中学习的魔力:你可以从零开始构建一个系统。随着时间的推移,这个从零人类知识开始的系统学会了极其复杂的人类领域的所有技能和能力,这些领域是人类花费整个职业生涯去学习所有知识的领域。而这些系统自己学会了获取并超越人类在这些领域中获得的所有知识。我认为这非常美妙。而且算法非常简单,只有这三个步骤,没有别的。

The first of which happened a long time ago, but I'm just going to briefly remind people of it because the others are going to build on this. That was the AlphaZero algorithm. And I have to say, when I look back at my career, I think it's the thing I'm most proud of. And I'm most proud of it because it felt magical. It had this magic of learning from experience: you could just have something that started from nothing. And over time, this system that started with no human knowledge whatsoever learned all the skills and capabilities of really complex human domains, domains where humans had spent their entire careers learning all the knowledge there is to learn. And these systems have just learned for themselves to acquire and then exceed all of the knowledge that humankind had ever acquired in these domains. I think there's something very beautiful about that. And the algorithm was incredibly simple. It was just these three steps and literally nothing else.

David

算法的第一步是进行规划。这一步通过蒙特卡洛树搜索实现,而蒙特卡洛树搜索由当前的策略和值函数引导。第二步是策略改进步骤:将策略更新为蒙特卡洛树搜索找到的最佳动作。第三步是策略评估步骤:根据使用蒙特卡洛树搜索进行规划的游戏结果更新值函数。这就是整个算法。当你实例化这个算法时,你会得到这些不可思议的结果:在短短几小时内,它就能从一个完全随机、毫无知识的神经网络,在我们应用的所有游戏中达到超人类水平:国际象棋、将棋和围棋。它花了大约四个小时就击败了当时已经超人类的国际象棋计算机冠军,并在其他游戏中取得了世界最佳表现。所以这是从经验中学习的力量的一个非常纯粹的例证。

Step one of the algorithm was to do planning. This step was achieved using Monte Carlo tree search, and the Monte Carlo tree search was guided by the current policy and value function. Step two was to do a policy improvement step: to update the policy towards the best action that was found by the Monte Carlo tree search. Step three was a policy evaluation step: to update the value function towards the outcome of games that were played using this Monte Carlo tree search to do planning. So that's the entire algorithm. And when you instantiate this algorithm, you get these incredible results where in just a few hours, it was able to go from a completely random neural network with no knowledge at all all the way through to superhuman performance in all the games we applied it to: chess, shogi, and go. It took something like four hours to beat the existing computer champion in chess, which was itself already superhuman, and also achieved the world's best performance in these other games. So this is a very pure example of the power of learning from experience.

David

一个小轶事:这个系统如此依赖经验,以至于在将棋游戏中,我们这些作者甚至都不知道怎么下将棋。我们构建了这个系统,第一次运行它,第一次训练,就在将棋上,我们甚至不知道它好不好。最后,我们也不确定。所以我们把它展示给 DeepMind 的 CEO Demis。他是一位出色的棋手,他说:‘哦,这看起来真的很棒。我要把它展示给世界冠军。’世界冠军回来说,这是他见过的最不可思议的将棋对局。所以这个过程非常优雅,并且有潜力做我们不知道的事情。它可以学习我们没有放入系统的技能。

A little anecdote: this was so experiential that in the game of shogi, none of us authors even knew how to play shogi. We put this system together, ran it for the first time ever, first training run, on shogi, and we didn't even know if it was good. At the end, we weren't sure. So we showed it to Demis, the CEO of DeepMind. He's a great games player and he said, 'Oh, this looks really good. I'm gonna show this to the world champion.' And the world champion came back and said it was like the most incredible games of shogi he'd ever seen. So there's something about this process that is very elegant and has the potential to do things that we don't know. It can learn skills that we don't put into the system.

案例研究 2:AlphaProof Case Study 2: AlphaProof

David

回到现在。去年我们构建了一个名为 AlphaProof 的系统,这是将强化学习(使用 AlphaZero 算法)应用于数学。基本思想是将整个形式数学视为一种游戏,然后将完全相同的方法(AlphaZero 方法)应用于数学游戏。在这种情况下,它类似于单人游戏。对于不了解的人来说,形式数学有这些优美的系统。我们使用了一种叫做 Lean 的语言。在 Lean 中,原则上所有可以表达的数学都可以表达。所以这是一种非常通用且强大的语言,人类数学家已经用 Lean 语言表达了大量的人类数学。

Moving into the present day. Last year we built a system called AlphaProof, which was an application of reinforcement learning using this AlphaZero algorithm to mathematics. The basic idea was to treat the whole of formal mathematics as a kind of game and then apply exactly this same method, the AlphaZero method, to the game of mathematics. In this case, it's like a single-player game. Formal mathematics, for those who don't know, there are these beautiful systems. We used a language called Lean. In Lean, all of mathematics that can in principle be expressed is possible to express. So it's a very general and capable language, and human mathematicians have expressed large amounts of human mathematics in this language of Lean.

David

发生的事情是,智能体可以在 Lean 中创建一个新策略。你有一个要解决的定理,表达为一个要证明的假设。智能体接收它所处状态的描述,这只是起始状态的描述,然后它可以采取行动,这些行动是 Lean 策略。Lean 策略基本上是操作求解器状态的方法。例如,这个策略引入一个局部变量到它的状态中,还有许多其他策略,比如反证法或归纳法,所有这些不同的东西。所以你在数学中熟悉的一切,所有这些步骤都可以采取,然后你可以构建一个程序来证明数学中的任何定理。与 LLM 不同,形式数学的美妙之处在于,你可以确定你是否证明了某件事。这是一个完美的证明;你就完成了。如果你达到了证明某件事的最终状态,你就知道你的证明是完美的。

What happens is that the agent can create a new tactic in Lean. You have a theorem that you're trying to solve, expressed as a hypothesis you're trying to prove. The agent receives a description of the state it's in, which is just a description of the starting state, and then it can take actions which are Lean tactics. Lean tactics are basically ways to manipulate the state of the solver. For example, this one introduces a local variable into its state, and there are many other tactics like proof by contradiction or induction, all these different things. So everything you're familiar with in mathematics, all of these steps can be taken, and then you can build essentially a program to prove any theorem in mathematics. The beautiful thing about formal mathematics, unlike LLMs, is that you know for sure if you've proved something. It's a perfect proof; you're just done. If you get to the end state where you've proved something, you know that your proof is perfect.

David

然后我们将 AlphaZero 应用于这个游戏。我们构建了一个神经网络,它接收系统的这个状态(即策略状态)作为输入,并估计值函数和策略,就像我们在围棋等游戏中所做的那样。策略输出策略。所以是的,就像我们之前所做的一切一样,然后我们可以根据这些策略进行搜索,并尝试找到一系列实际解决定理、证明定理的策略。这是一篇昨天发表的论文。我们去年宣布了结果,但《自然》论文终于发表了。

Then we apply AlphaZero to this game. We built a neural network which takes as input this state of the system, this tactic state, which estimates the value function and a policy, just like we did for the game of Go and so forth. The policy outputs tactics. So yeah, just like everything else we've done before, and then we can do a search in terms of those tactics and try to find a sequence of tactics that actually solves the theorem, that proves the theorem. This is a paper that came out yesterday. We announced results last year, but the Nature paper finally came out.

AlphaProof 在形式数学基准上的表现 AlphaProof Performance on Formal Mathematics Benchmarks

David

所以你现在可以看到这些图表。论文中有更多细节,如果你感兴趣的话。你看到的是,当你从这个系统开始,它并不是从一个完全普通的神经网络开始的。它从一个非常简单的、基本上在 archive 等数据上训练过的小型 LLM 开始。它非常简单。所以它一开始几乎对数学一无所知,然后随着时间的推移,它变得越来越好,直到在这一点上,它解决了标准形式数学基准测试中 97% 的问题,而之前没有人在这上面做得很好。然后这里我们看到它在国际数学奥林匹克竞赛的实际问题上的表现,这些问题被认为是超级复杂的挑战性问题,每年由来自世界各地的年轻数学天才们解决。所以我们解决了大约 40% 的历史问题,这些绿色的是我们在普特南数学竞赛上的表现,这是一个类似的大学水平数学测试。

So you can now see these plots. There's much more detail in the paper if you're interested. What you see is that when you start with this system, it's not starting from a completely vanilla neural network. It starts from a very simple small LLM that's basically been trained on things like archive. It's very simple. So it basically knows almost nothing about mathematics at the beginning, and then over time it gets better and better until at this point it solves 97% of the standard benchmark for formal mathematics that previously no one had done very well on. And then here we see how well it's doing on actual problems from the International Mathematics Olympiad, which are considered super complex challenging problems worked on each year by young mathematical geniuses from all over the world. So we solved about 40% of all historical problems here, and these green ones show how we did on Putnam, which is a similar test for university-level mathematics.

从经验中规划与 IMO 2024 结果 Planning from Experience and IMO 2024 Results

David

但我们在那之上还做了别的事情:我们说,如果我们想从经验中规划呢?如果我们想从那个基础性能水平开始,再次运行同样的算法呢?这个东西是从我们从各处提取的数百万个问题中学习的。我们找到了每一个能找到的人类问题,并以某种方式将其形式化为 Lean,然后将系统应用于所有这些问题。所以这里我们基本上是在大量不同的问题上训练,作为保留的结果,我们在 IMO 上的保留性能显示在这里。我们做的是应用完全相同的算法,但只针对我们关心的那些特定 IMO 问题。所以这里你从一个特定的 IMO 问题开始,系统基本上只是试图找出如何解决那个问题,尝试那个特定问题的许多不同变体。我们看到它从我们这里的 39% 的性能开始,然后通过即时应用 AlphaZero——这是它在几个小时内完成的事情——它变得越来越好,通过即时应用这个东西,学习获得更好的性能。所以真的做得非常好。而且应该说,它在这个人们一直在测量的标准基准上也达到了 100%。所以很高兴清除了那个基准。所以当我们应用这个时,这正是我们去年在国际数学奥林匹克竞赛中应用的系统。它被应用于两个完全未见过的、由数学界设计为超级挑战和未知的问题。由于这种从经验中规划的方法,我们花了稍长的时间——比人类稍长。人类有九个小时来解决这些问题;我们花了三天来解决这个问题。但我们得到了 28 分,只差一分就达到金牌门槛。这是第一次有人使用 AI 在国际数学奥林匹克竞赛中获得奖牌。这被认为是 AI 多年来的重大挑战之一,以达到这种性能水平。

But then we did something else on top of that: we said, what if we wanted to plan from experience? What if we wanted to start from that base level of performance and run the same algorithm again? This thing here was learning from millions of problems we'd extracted from all over the place. We'd taken every single human problem we could find and formalized it in some way into Lean, and we applied the system to all of those. So here we're basically training on a vast number of different problems, and as a held-out consequence, our held-out performance on IMO is shown here. What we do is apply exactly the same algorithm but to those specific IMO problems we care about. So here you start from a particular IMO problem, and the system basically just tries to figure out how to solve that problem, trying loads of different variants of that specific problem. We see that it starts off at this performance we had here, 39%, and it rapidly, through applying AlphaZero on the fly—this is something it does in a few hours—gets better and better by applying this thing on the fly, learning to get much better performance. So really doing very well. And it should be said it also gets to 100% on this standard benchmark that people have been measuring. So it's nice to have cleaned that one out. So when we applied this, this is exactly the system we applied last year in the International Mathematics Olympiad. It was applied to two completely unseen problems devised by the mathematical community to be super challenging and unknown. We took slightly longer because of this approach of planning from experience—it took slightly longer than humans. Humans get nine hours to solve these problems; we took three days to solve this problem. But we got 28 points, which was just one point short of the gold medal threshold. This is the first time anyone has ever achieved a medal in the International Mathematics Olympiad using AI. This was considered one of the grand challenges of AI for some years to reach this level of performance.

解决最难的 IMO 问题 Solving the Hardest IMO Problem

David

有一个问题特别值得一提:第六题,当年 IMO 中最难的问题。这个问题非常难,只有不到 2% 的年轻数学天才能够解决它。而 AlphaProof 完美地解决了它。这实际上是它找到的证明。你可以看到它生成了相当复杂的 Lean 程序。我不期望这里的任何人能完全理解这一切——这需要一些时间——但我认为有趣的是,它创造了一些非常复杂的东西,找到了一个很少有人类数学家能够解决的问题的完美解决方案。

There was one problem in particular: problem six, the hardest problem from that year's IMO. This problem was so hard that less than 2% of those young mathematical geniuses were able to solve it. And AlphaProof solved it perfectly. This is actually the proof that it found. You can see that it's generating quite complex programs in Lean. I wouldn't expect anyone here to unpick all of this—it would take a while—but I think the interesting thing is just to say it created something very complex to find a perfect solution to this problem that very few human mathematicians had ever been able to solve.

为何只解决了 60%的问题 Why Only 60% of Problems Solved

Host

它解决了这么难的问题,为什么你认为它只解决了 60% 的问题?

It's solving this hard problem, why do you think it only solves 60% of them?

David

是的,这是个好问题。所以问题是,如果它解决了最难的问题,为什么它只解决了所有问题的 60%?对此有两个答案。首先,我展示的图表使用了非常有限的算力。算力是在所有历史 IMO 问题中衡量的,所以算力被所有历史 IMO 问题共享。当我们在 IMO 2024 上运行时,我们将相同数量的算力集中在我们试图解决的当年六个问题上。所以那里使用了更多的算力——这是第一个答案。第二个答案是,AlphaProof 在某些类型的数学上非常强,而在其他类型上非常弱。例如,它在代数方面很棒,但在组合学问题上弱得多,这些问题通常很难甚至用 Lean 表达。今年 IMO 中有一个我们没有解决的问题,很多人类也讨厌它,但它基本上需要数页的 Lean 才能形式化问题。然后解决方案是一个特定的非常直观的想法,人类需要似乎通过魔法才能想出来。这类事情,我认为我们离用这些方法解决还有一段距离。

Yeah, it's a good question. So the question was if it's solving the hardest problem, why is it only solving 60% of all problems? There are two answers to this. First of all, the plot I showed you was using very limited compute. The compute was being measured across all historical IMO problems, so the compute was being shared amongst all historical IMO problems. When we ran on IMO 2024, we concentrated the same amount of compute just on the six problems we were trying to solve for that year. So more compute was used there—that's the first answer. The second answer is that AlphaProof is very strong at certain types of mathematics and very weak at others. For example, it's great at algebra, but much weaker at combinatorics problems, which often are very hard to express even in Lean. There was a problem we didn't solve in this year's IMO which a lot of humans also hated, but it basically required pages and pages of Lean just to formalize the problem. And then the solution was a particular very intuitive idea that humans needed to pull out seemingly by magic. Those kinds of things, I think we're still some way away from solving with these kinds of methods.

自动化形式化与公共工具 Automated Formalization and Public Tool

Host

你们有自动形式化问题的方法吗?

Do you have an automated method for formalizing problems?

David

既是也不是。我们确实有一种自动方法。事实上,在我们解决的所有问题上,我们都正确地形式化了问题。但对于那个非常困难的组合学问题,我们没有正确地自动形式化。所以当我们实际运行比赛时,我们使用了人类形式化的例子。原则上,系统可以自己形式化这些东西,并且对于 IMO 类型的问题实际上相当擅长。我们在内部构建了类似的东西,但还没有对外提供。我应该说,现在有一个与《自然》论文同时发布的公共工具,允许人们使用 AlphaProof,并希望通过使人们能够解决他们关心的任何定理或猜想,来加速世界的数学发展。

Yes and no. We do have an automated method. In fact, on all the problems we solved, we correctly formalized the problem. But for that very difficult combinatorics problem, we didn't correctly auto-formalize it. So when we actually ran the competition, we used human-formalized examples. In principle, the system can formalize these things for itself and is actually pretty good at that for IMO-type problems. We have built something like that internally, but that's not yet available externally. I should say that there is now a public tool that came out alongside the Nature paper which allows people to play with AlphaProof and hopefully accelerate the world's mathematics by enabling people to solve any theorem or conjecture they care about.

从经验中学习 RL 算法 Learning the RL algorithm from experience

David

所以是的,它对某些用例非常有用。但还不是真正的终极解决方案,不像解决黎曼猜想那样,离那种数学还差得远。好的,最后我想谈第三个案例研究。这是另一项工作,我们实际上几年前就做了,但直到两三周前才发表。我认为这是另一个非常漂亮的从经验中学习的例子。与 AlphaZero 的方式不同,但它本质上确实学到了关于从经验中学习本质的一些深刻的东西。这里要解决的问题是,强化学习本身很难做好。原因包括:目标不可微,最佳方法因环境而异,没有通用的鲁棒算法能适用于任何环境。我们缺乏这些东西。从某种意义上说,从经验中学习的成熟度还没有达到让事情自动运行的程度。所以我们尝试通过从经验中学习算法本身来解决这个问题。如果系统本身能学会做 RL 的正确方法呢?我们构建了一个系统,其中有大量不同的模拟环境,比如 Atari 环境等。我们有一百个不同的环境,让智能体在每个环境中使用一个学习规则进行交互。我们不是使用 Q-learning、策略梯度、分布式 RL 或任何人们手工设计的数百种算法,而是使用一个神经网络。这个神经网络能够输出任何它想要的更新规则,以目标的形式让智能体将其网络向这些目标移动。然后我们学习如何调整这个元网络,使得当它在每个环境中训练时,表现尽可能好。这实际上就是一个强化学习过程:智能体学习哪个规则有效,并通过我们称为元梯度下降的方法来优化整个过程。就像每个东西都有一个基于梯度的规则,然后我们展开该规则的多个步骤,通过整个过程取梯度,找出哪些步骤序列实际上导致了最佳性能,并不断调整使其更好。这就是想法:我们试图学习在强化学习中导致最佳性能的规则。它实际上表现得非常好。当我们在 Atari 上测量这个系统的性能时(Atari 是用于元训练的环境之一),它超越了人类创造的最佳算法,包括我们自己用 M0 等创造的记录。这些元学习算法做得更好。蓝线显示仅在 Atari 游戏上元学习的结果。它表现得非常好,超越了所有研究人员在深度强化学习中的最佳努力。我们找到了一个更好的学习算法。但真正酷的是,它不仅适用于 Atari,也适用于未见过的环境。我们在 Procgen 和 DeepMind Lab 等环境上测量,发现完全未见过的、仅在 Atari 上训练的那个算法在这些其他环境上也表现得非常好,击败了 M0 和其他现有最先进方法。如果你扩展训练,也在这些额外环境上训练,那就是橙色线,它表现得甚至更好。所以可以说,这是 AI 的一个新缩放定律,一个元学习的缩放定律:你放入系统的训练环境越多,给它越多样的经验,它就能产生越来越好、最终超越我们创造的最佳事物的学习规则。

So it's yes, it's very good for certain use cases. It's not yet the real deal. It's not like solving the Riemann hypothesis or something. It's far off that kind of mathematics. Okay. So I want to talk finally about a third case study. This is another piece of work. We actually did this a couple of years ago but it was only published two or three weeks ago. And it's really, I think, another very beautiful example of learning from experience. In a different way to AlphaZero, but it basically is something which has really learned something quite profound about the nature of learning from experience. The problem here is trying to address the fact that reinforcement learning itself is hard to do well. Some of the reasons it's hard: it has a non-differentiable objective, the best recipe varies between environments, and there's no general-purpose robust algorithm that just works across any environment you might drop it into. We're lacking those things. In some sense, the level of maturity for learning from experience isn't yet at the level that things just work. So we tried to address this problem by learning the algorithm itself from experience. What if the system itself could learn the right way to do RL? So what we did was we built a system where we had many different simulated environments, things like Atari environments and so forth. We had a hundred different environments and we allowed the agent to interact in each of those environments using a learning rule. Instead of using Q-learning, policy gradients, distributional RL, or any of the hundreds of algorithms people have come up with by hand, we used a neural network. The neural network was able to output any kind of update rule it wanted in terms of targets that the agent would move its network towards. Then we learned how to adjust that meta-network such that when it was trained on each of these environments, it did as well as possible. This was literally a reinforcement learning process where the agent is learning which rule is effective and optimizing that whole thing through what we call meta-gradient descent. It's like taking each of these things, there's some gradient-based rule used here, then we unroll the steps of that rule over many steps and take the gradient through that whole thing, and say which of those sequences of steps actually led to the best performance, and we adjust that thing to make it better and better. So that's the idea: we're trying to learn the rule that actually leads to the best performance in reinforcement learning. And it actually does remarkably well. When we measure the performance of this thing on Atari, which was one of the environments used to meta-train this thing, it outperformed the best ever human-created algorithms, including our own records with things like M0. These meta-learned algorithms did better. The blue line shows what happened when it was meta-learned just on Atari games. It did really, really well and outperformed the best efforts of all these researchers across deep reinforcement learning. We found a better learning algorithm. But what's really cool is it doesn't just work on Atari; it also works on held-out environments. We measured it on things like Procgen and DeepMind Lab, and we saw that completely held out, the one which was trained on Atari does really well also on these other environments, beating things like M0 and other existing state-of-the-art methods. Then if you extend the training to also train on these additional environments, that's the orange line, it does even better still. So one way to say this is that this is like a new scaling law for AI, a scaling law for meta-learning: the more training environments you put into the system, the more diverse experience you give it, it basically comes up with a learning rule that is better and better and ultimately outperforms the best things we've ever created.

Host

是的。提问。你在选择不同类型的环境时投入了很多思考吗?

Yeah. Question. Did you put a lot of thought into the different types of environments that you picked?

David

嗯,是的,我们是否在选择环境上投入了很多思考?其实没有。我觉得我们只是选择了方便的东西。从某种意义上说,我们试图选择经典的模拟环境,这些环境容易获取,有很多比较点,大家都熟悉。所以我们选了 Atari 和 Procgen。我们也在其他一些东西上试过,比如 NetHack 等。论文中有更多信息,但它在所有我们试过的环境上都表现不错。我想我们也试过 Soccerban 和其他一些东西。但我觉得这已经说明了主要观点。好的,这就是我给你们的第三个案例研究。我认为这些案例研究的共同点是,我真的只是想传达一个想法。我们当然远未解决所有问题,还有很多工作要做,但我真的只是想传达:当我们投入努力、信念和资源,试图通过经验解决问题时,它会带来不可思议的成果。所以我想用一张幻灯片作为结束,这是一个号召。对像你们这样的人的号召。这个号召是尝试解决 AI 的深层问题,即如何从经验中学习。构建像视频中婴儿那样的东西。那不是很不可思议吗?一个被放入环境中,不断学习,然后能成为舞者或神经外科医生或任何角色的东西。那就是 AI 的深层问题。如果我们能做到,那将是科学的一个深刻时刻。如果我们做到了,它将改变 AI 的未来。而且我相信,通过它,也能改变人类。谢谢。

Um, yeah, so did we put a lot of thought into the environments we picked? Not really. I think we just did the things which were convenient. In some sense, we tried to pick canonical simulated environments that were readily accessible and had lots of comparison points, that everyone knows and is familiar with. So Atari and Procgen were the kind of things we picked. We also tried it on a few other things like NetHack and so forth. You can see more information in the paper, but it worked pretty well on all the things we tried it on. I think we also played with Soccerban and a few other things. But I think this already illustrates the main points. Okay. So that's the third case study I have for you. I think what's common amongst these case studies is I really just want to convey the idea. We're certainly far from having solved everything here. There's a lot to do, but I really just want to convey the idea that when we do put effort and conviction and resource into trying to solve things through experience, it yields incredible things. So I want to leave you with a slide which is like a call to arms. A call to arms for people like yourselves. The call to arms is to try and solve the deep problem of AI, which is how to learn from experience. Build something like the baby in the video. Wouldn't that be incredible? Something which is placed into its environment and just learns and learns, and then could become a dancer or a neurosurgeon or whatever. That's the deep problem of AI. If we can do that, this will be a profound moment for science. And if we do that, it will transform the future of AI. And I also believe through that, humanity. Thank you.

Host

你最后展示的那种东西,你在 3D 环境上也试过吗?

The final kind of thing you showed, have you tried that on like 3D environments as well?

David

是的。

Yeah.

3D 环境与泛化 3D environments and generalization

Host

那我们在 3D 环境上试过吗?

So have we tried it on 3D environments?

David

DM Lab 是一个 3D 环境。这是我们 DeepMind 早期使用的环境之一,也是我们最早展示深度强化学习成功的 3D 环境之一。像 IMPALA 和 UNREAL 这些早期论文都应用了它。所以答案是肯定的,它有效。但总有一个泛化问题:在 2D 环境中有效的算法能泛化到什么程度?蓝线表明,如果在更接近你关心的本质的环境上训练,它学得甚至更好。希望是,通过覆盖越来越多不同类型的环境,我们可以开发出非常通用和稳健的学习算法。

DM Lab is a 3D environment. This was something we used in the early days of DeepMind, one of the first 3D environments where we showed success with deep reinforcement learning. Very early papers like IMPALA were applied to that, and UNREAL, some of our early papers. So the answer is yes, it works. But there's always a generalization question: how well do algorithms that work in 2D environments generalize? The blue line shows that it learns something even better if trained on environments closer to the nature of things you care about. The hope is that by covering more and more different types of environments, we can develop very general and robust learning algorithms.

未知 AI 的对齐难题 Alignment difficulty with uncharted AI

Audience

谢谢你的演讲。如果我们让 AI 进入未知的深度领域,那不会让对齐变得更加困难吗?你有什么想法吗?因为我们甚至不知道自己应该做什么。

Thanks for the talk. If we let AI go into deep uncharted territory, wouldn't that make alignment much more difficult? Do you have any thoughts on how? Because we don't even have an idea of what we ourselves should be doing.

David

我对这个问题有很多想法。首先,对齐已经很困难了。我不认为仅仅因为我们有能用英语思考的系统,我们就解决了对齐问题。这些系统很容易表面上看起来有帮助,而内部却在做完全不同的事情。所以对齐无论如何都需要解决。其次,我认为有一天有可能构建真正利他的系统,它们试图优化的信号是通过观察像我们这样的其他智能体的行为推断出来的。那将非常令人兴奋。最后,让我对 AI 对齐的未来最乐观的是,未来不会只有一个单一的 AI 试图实现一个单一的目标。那是人们非常担心对齐的场景。我认为现实将是大量不同的 AI 追求大量不同的目标,它们都处于某种相互制衡的复杂社会中。那么你必须问:它们会合作以支持其他智能体的繁荣吗?如果其中一些智能体支持人类的繁荣,它们就会支持我们的对齐,并创造一个所有智能体都能繁荣的系统。为什么它们应该合作?我相信,随着智能的提高,系统将能够更好地建模人类和其他实体,它们的行为和欲望,并通过这种方式创造更好的协议和合作。所以在我看来,合作很可能会随着智能的提高而增加,而不是减少。底线是,我们正在走向复杂的社会,许多目标被平衡在一起,其中一些目标将支持人类的需求。可能会有不这样做的坏角色,但没关系,因为这完全就像我们今天的社会。我们有比我们强大得多的政府;我无法单独对抗美国政府。但还有其他政府形成了平衡,尽管不完美。如果它们能更好地合作,我们的处境会更好。我认为有一天我们会走向类似的情况。抱歉,这是一个很长的科幻式回答,但这就是我的感受。

I've got a lot of thoughts on this. First, alignment is already difficult. I don't think we've solved alignment just because we have systems that think in English. It's very easy for those systems to seem helpful on the surface while internally doing something quite different. So alignment will have to be addressed anyway. Second, I think one day it will be possible to build systems that are truly altruistic, where the signal of what they're trying to optimize is inferred from observing the behavior of other agents like ourselves. That would be very exciting. Finally, what makes me most optimistic about the future of AI alignment is that the future won't be just one single AI trying to achieve one single goal. That's the scenario where people worry a lot about alignment. I think the reality will be a large number of different goals pursued by a large number of different AIs, all in some complex society that balances against each other. Then you have to ask: will they cooperate to support the flourishing of other agents? If some of those agents support human flourishing, they will support our alignment and create a system where all agents can flourish. Why should they cooperate? I believe that as intelligence increases, systems will be better able to model humans and other entities, their behaviors and desires, and through that create much better agreements and cooperation. So it seems very likely to me that cooperation will increase with intelligence rather than decrease. The bottom line is we're moving towards complex societies where many goals are balanced together, and some of those goals will support the needs of humans. There may be bad actors that don't, but it'll be okay because it's exactly like the society we're in today. We have governments far more powerful than us; I can't fight the US government individually. But there are other governments that form a balance, though imperfect. If they were better at cooperating, we'd be in a better state. I think we'll move towards something like that one day. Sorry, that's a long sci-fi answer, but that's how I feel.

多智能体系统与发现算法 Multi-agent systems and discovery algorithm

Audience

我想接着你的不确定性说。在现实中,你不是只有一个婴儿,而是有几个婴儿一起工作,共同操纵环境。你在多智能体系统的情况下试过这个发现算法吗?

I'm going to piggyback on your uncertainties. In reality, you don't have one baby, you have a few babies working together that manipulate the environment together. Have you tried this discovery algorithm in any case where it's a multi-agent system?

David

我们还没有试过。这个发现算法还有很多事情我们没做过。这是我们之前做的工作,然后世界变了,每个人似乎都被吸引去帮助构建像 Gemini 这样的东西。所以我认为这非常令人兴奋,潜力很大。如果我能列出其他一些我们没做过的事情:我们还没有探索连续环境、长程环境。还有很多没试过。但我认为这个概念是可行的。

We have not yet tried it. There's lots of things we haven't done yet with this discovering algorithms. It's work we did a while ago, and then the world changed and everyone seemed to get sucked into trying to help build things like Gemini. So I think it's really exciting and there's a lot of potential. If I could list some other things we haven't done: we haven't explored continuous environments, long environments. There's lots that hasn't been tried. But I see it as a concept that this kind of idea can work.

理解新环境中的奖励函数 Understanding reward function in new environments

Audience

最后一个问题。我试图理解你的号召。我们真正需要做的是设计我们对如何构建和理解奖励函数的理解。在你所有的经验和案例研究中,系统的奖励函数非常明确:解定理、独特的 Atari 等。对于婴儿来说,奖励是什么就不那么清楚了,但生物学内置了一些奖励:如果疼,就停止做这个。这些东西在系统中是缺失的。如果你拿 LLM 并告诉它们从经验中学习,它们会开始互相交谈,写科学论文来支持自己的答案等等。很容易创造大量经验,但它们不会是你真正想要的。它不会有创造性元素。所以,面对全新的环境,这真的只是试图理解你的奖励函数是什么吗?在生物学和社会中,我们共同创造它。一部分来自生物学,一部分来自你的父母。

Last question. I'm trying to parse your call to arms. What we really need to do is design our understanding of how to craft and understand the reward function. In all your experiences and case studies, the reward function for the system was very clear: solve the theorem, unique Atari, etc. For the baby, it's much less clear what the reward is, but biology builds in some reward: if it hurts, stop doing this. Those things are lacking in the systems. If you were to take LLMs and tell them to learn from experience, they would start talking to each other and writing scientific papers to back their answers, etc. It's easy to create lots of experiences, but they wouldn't really be what you wanted. It wouldn't have a creative element. So is that really just trying to understand what your reward function is in the face of completely new environments? In biology and society, we co-create it. Some of it is biology, some comes from your parents.

David

是的,我认为这是一个很好的观点。奖励函数至关重要,在新环境中,理解它变得更加重要。在生物学中,奖励是通过进化和社会学习共同创造的。对于 AI,我们需要弄清楚如何设计能导致期望行为的奖励函数,尤其是在开放式的环境中。这是一个关键挑战。

Yes, I think that's a very good point. The reward function is crucial, and in new environments, understanding it becomes even more important. In biology, rewards are co-created through evolution and social learning. For AI, we need to figure out how to design reward functions that lead to desired behaviors, especially in open-ended settings. That's a key challenge.

关于奖励函数与目标的辩论 Debate on reward function and goals

Host

所以你以这种方式学习,本质上随着时间推移你学到了东西。

So you learn this way and essentially as you go on you learn things.

David

是的。我想我有几个答案。一个是,我认为号召是解决核心问题,但几乎与奖励函数无关。所以我认为那是科学问题。科学问题是构建一个可以应用于任何环境、任何奖励信号的婴儿。那是科学问题,是 AI 的深层问题。现在,如果你问,如果要构建一个放入真实世界的婴儿,问题应该是什么,我认为有很多可能的答案。但请注意,就我展示的视频而言,我认为公平地说,这个婴儿是内在驱动的,对吧?所以它学习的大部分内容来自于它选择时不时地追求变化着的子目标。我认为如果我们能解决那个问题,我们就会取得很大进展。

Yeah. I mean I think I have a couple of answers to that. One is that I think the call to arms is to solve the core problem but almost agnostic as to what the reward function is. So I think that's the scientific problem. The scientific problem is to build a baby that can be applied to any environment, to any reward signal. That's the scientific problem. That's the deep problem of AI. Now, if you ask what should the problem be if you're trying to build a baby that's placed into the real world, I think there are many possible answers. But note that, in terms of the video which I showed, I think it's fair to say that this baby is intrinsically motivated, right? So most of what it's learning is coming from it choosing to pursue subgoals from time to time that shift. And I think if we can solve that problem, then we'll go a long way.

Host

抱歉,但我要反驳你的婴儿类比。这个婴儿处于一个精心设计的环境中。我们为这个婴儿创造了这个空间。

I'm sorry but to push back on your baby analogy. This baby is in a very carefully crafted environment. We have made this space for this baby.

David

是的。

Yeah.

Host

而且我实际上从根本上不同意你的观点。我认为真正的问题是我们试图解决什么?这比如何进行训练系统的技术部分要困难得多。我理解那是一个令人兴奋的问题。这是一个很好的计算机科学数学问题。但关于我们如何决定做什么的问题,我认为比获得更好的 Q 学习算法要困难得多,也更根本。我想我试图说的是,设计奖励函数和研究它是什么,这就是为什么我不太同意你的答案,我认为你的根本问题仍然定义不清,因为根本问题是理解我们需要解决什么。

And I actually fundamentally disagree with you. I think the real problem is what are we trying to solve? That's a far more difficult problem than how do we go about the technical part of training a system. I understand that that's an exciting problem. It's a wonderful computer science mathematical problem to address. But the question of what how are we going to decide what to do I'd say is far harder and far more fundamental than do I get a better Q-learning algorithm. I guess that's what I was trying to say is that the designing the reward function and studying what is it that that's why I don't quite agree with your answer that I think your fundamental problem is still ill-defined because the fundamental problem is understanding what is it that we'll need to solve.

David

所以,你知道,我再次猜想我看到了一个未来,那里会有许多许多具有多种目标的 AI,我们将拥有一个复杂繁荣的社会,其中追求着多种目标。不会只有一件事,你知道我们不会想要一个只有单一目标的世界,事实上我认为那是一个危险的世界,我认为我们想要一个有多重目标的世界,创造一个复杂的社会,其中追求许多不同的事物。我们有一个复杂系统,就像雨林是一个复杂系统,有许多动物追求不同的事情,有些在尝试飞行,有些在尝试进食,有些在尝试挖洞。我认为这实际上是一个不可避免的未来,如果你考虑我们想要构建的系统类型,我们想要构建生物反应器,其目标是尝试制造某种材料;我们将拥有智能材料,它们试图帮助我们排水或做其他事情。会有系统试图成为自动驾驶汽车,试图高效安全地到达目的地。我们将把 AI 用于各种目的,所有这些都有一个共同的根节点,如果我们能解决它,将带来一个令人难以置信的未来,其中先进能力成为常态。这就是我主张的未来,我相信我们可以实现它,我认为问“奖励函数是什么”是一个错误的问题,没有单一的奖励函数,就像有许多目标,我们应该拥抱多元的可能性,就像进化拥抱了各种不同的动物以及被创造的各种不同的大脑和心智一样。

So you know again I guess I see a future where there will be many many AIs with many goals and we'll have a complex flourishing society where these multiple goals are being pursued. There won't just be one thing which you know we wouldn't want just a world in which there's only one goal in fact I think that's a dangerous world I think we want a world in which there are many goals you know creating a complex society where many different things are being pursued you know we have a complex system in the same way that a rainforest is a complex system with many animals pursuing different things and some things are trying to fly and some things are trying to eat and some things are trying to build burrows and this is I think actually an inevitable future and if you think about the kind of systems we want to build we want to build bioreactors which have a goal of trying to create a certain kind of material we'll have smart materials which are trying to help us shed water or whatever it is they're trying to do. There'll be systems which are trying to be autonomous vehicles that are trying to reach a destination efficiently and safely. There's going to be all kinds of purposes towards which we will put AI and there's a common root node to all of those things which if we can solve it will lead to an incredible future in which advanced capabilities are the norm and so that's the future which I'm arguing for and believe which we can achieve and I think that it's just the wrong question to ask like what's the reward function there is no the reward function it's like there are many goals and we should embrace the plurality of possibilities in the same way that evolution embraced that with all kinds of different animals and all kinds of different brains and minds that have been created.

Host

所以我建议我们在此章节暂停。让我们感谢 J

So I suggest that we continue this chapter break. So let's thank J

互动版:逐字朗读 + 针对本期提问 →