里奇·萨顿:这个领域才是奇怪的,不是我

Rich Sutton: The Field Is Weird, Not Me

理查德·萨顿 Richard Sutton · Training Data · 2026-08-18 · 约 54 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

里奇·萨顿讨论了他关于持续学习的非传统观点,以及他创立 Oak Lab 的历程。

Rich Sutton discusses his unconventional views on continual learning and his journey founding Oak Lab.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 14)

全文 · Full transcript(中英对照)

开场与介绍 Opening and Introduction

Host

今天我们很荣幸请到了伟大的 Rich Sutton 来到节目。Rich,你发明了强化学习,你写了那本开创性的教科书,你是这个领域的关键学生之一。像 Dave Silver 这样的学生,你写了那篇《苦涩的教训》,我相信那是这个领域的圣经。你一直是推动这个领域前进的伟人之一。所以,感谢你今天抽出时间加入我们。嗯,Rich 身边还有他的联合创始人、阿尔伯塔大学的前学生 Kuram Javeed。嗯,你们两位创办了 Oak Lab。我今天非常兴奋能和你们聊这个。那么,在今天的环节中,我们首先要聊《苦涩的教训》、我们今天所知的世界的状态、LLM 是否能带我们到达那里,然后我们会过渡到聊你们的资源议程和 Oak 的计划。

We are honored to have the great Rich Sutton with us here today. Rich, you invented reinforcement learning. You wrote the seminal textbook. You're one of the key students in the field. Folks like Dave Silver, you wrote the essay, the bitter lesson, that I believe is the Bible of the field. And you have just been one of the greats in propelling the field forward. So, thank you for taking the time to join us today. Um, Rich is joined by Kuram Javeed, his co-founder, uh, and former student from the University of Alberta. Um, the two of you have set off to found Oak Lab. I'm very excited to talk to you about that today. So, for today's session, we're going to start talking about the bitter lesson, the state of the world as we know it today, whether LLM will get us there or not, and then we're going to transition to start talking about your resource agenda and your plan for Oak.

Host

Rich,也许带我们回到过去。我本来想从《苦涩的教训》开始,但我其实想更早一点开始。几十年前,你决定把职业生涯奉献给强化学习,特别是深度强化学习,并把阿尔伯塔大学建立成那个领域的堡垒,当时我认为这个领域还处于萌芽阶段。是什么给了你这样的信念?

Rich, maybe take us back. I was going to start with the bitter lesson, but I actually want to start earlier than that. Decades ago, you decided to dedicate your career to reinforcement learning, to deep reinforcement learning in particular, and you established the University of Alberta as a bastion of that back when I think the field was very much in its infancy. What gave you the conviction to do that?

Richard

嗯,你还能做什么呢?我们试图理解心智,而学习是心智的核心部分,有目标也是心智的核心部分,是智能的核心部分。是的。所以,我只是在我一直思考的事情上加倍下注。

Well, what else are you going to do? We're trying to figure out the mind and learning is a central part of the mind and having a goal is a central part of the mind, central part of intelligence. Yeah. So, I was just doubling down on what I was always thinking.

Host

当时人们觉得你疯了吗?

Did people think you were crazy at the time?

Richard

嗯,那是一个冬天。那是 AI 的冬天。

Um, it was a winter. It was an AI winter.

Host

那是哪一年?

What year was this?

Richard

那是 2003 年。

It was in 2003.

Host

好的。

Okay.

Richard

而且实际上有点疯狂,真相是,我当时病得很重。我快死了。2003 年我其实因癌症濒临死亡,嗯,但我还没完全死掉,你知道。我努力了好几年,但我没死。我处于又一次缓解期。所以我说:“嗯,我没死。我没能成功死掉,所以我不如,你知道,这事拖得够久了。我不如再找份工作。”所以我去了阿尔伯塔,开始在那里教书,最后我没死。这有点神奇。因为,是的,就是这样。我现在是开玩笑,但当时很严重。嗯,这是一个更令人心酸的问题。为什么我继续做研究,当我知道自己只有几个月时间?我总是想起,我认为本杰明·富兰克林说过,如果你想知道为什么某人做某事,几乎总是两个原因之一:要么是习惯,要么是虚荣。好吧,所以我认为这可能是真的。也许是我的习惯,继续做我一直做的事,或者也许是虚荣。我不知道。我认为更像是习惯,因为我快死了。

And it's kind of crazy actually, the truth, because I was really sick. I was dying. I was actually dying of cancer in 2003, and uh, but I wasn't quite dead, you know. I've been trying for a number of years and I wasn't dead. I was in another remission. And so I said, "Well, I'm not dying. I haven't succeeded in dying, so I might as well, you know, it's going on long enough. I might as well just try to get another job." And so I went to Alberta and started teaching there, and then in the end I didn't die. It's kind of amazing. It's because, yeah, it's like that. I'm joking about it now, but it was quite serious. And um, it's an even more poignant question. Why did I continue to work on this research stuff when I was, you know, only had a few months? I would always keep reminded what I think it's Benjamin Franklin is supposed to have said that, you know, if you ever wonder why someone is doing something, uh, it's almost always one of two things: it's either habit or vanity. Okay, so I think that's probably true. Maybe it was my habit, just kept doing what I always been doing, or maybe it was vanity. I don't know. I think it was more like habit, because I was dying.

Host

哇。天意。

Wow. Divine intervention.

Richard

嗯,对我来说保持坚定一直很容易。嗯,我要在这个回答上说得更长。

Um, it's always been easy for me to keep being very determined. Um, and I'm going to go even longer on this answer.

Host

请讲。

Please do.

Richard

呃,人们认为我有一个激进的观点。有时他们开始提问,说我的想法和所有人都不一样。但我完全不这么看。我觉得我是在以普通的方式思考。只是其他所有人想得有点奇怪。我的意思是,你知道,只是最近人们想得奇怪。如果你回顾人们对心智的看法,哪怕只是十年前,你会发现这样的想法:学习很重要。你必须有一个目标。嗯,你知道,感知很重要,我们是低级生物。我们以高速生成动作和感知数据,然而我们必须更高层次地思考。你知道,回到 AI 疯狂之前的几十年。你不必说持续学习,因为谈论非持续的学习毫无意义。所有学习都是持续的。你知道,这不是一个特殊阶段。我们总是行动和学习。这就是正常的思考方式。我不奇怪。这个领域才奇怪。他们需要称之为持续学习。那只是学习。

Uh, people think I have a radical point of view. Sometimes they start questions saying how what I'm thinking is so different from everyone else. But I don't see it that way at all. I see it as like I'm thinking the ordinary way. It's just everyone else that's thinking a bit weird. And I mean that like, you know, it's just the recent times people are thinking weird. If you look back at what people thought about the mind, you know, even just a decade, you'll find the kind of thoughts that learning is important. You've got to have a goal. Um, and you know, perception is important, and we are low-level beings. We are generating actions and perceiving data at a fast speed, and yet we have to think at higher levels. And you know, go back a few decades before there was all this AI craziness. You wouldn't have to say continual learning, because it wouldn't make any sense to talk about learning that wasn't continual. All learning is continual. You know, it's not a special phase. We always act and we learn. That's just the normal way of thinking. I'm not weird. The field is weird. They need to call it continual learning. It's just learning.

Host

我不奇怪。其他人才奇怪。这是一个很好的生活格言。我们得为此发一篇檄文。我们很高兴你活了下来,这个领域很高兴你活了下来,感谢你推动 AI 的前沿。

I'm not weird. Everybody else is. That's a good motto to live by. We're gonna have to send out an expost about that. We're very happy that you lived on, where the field is happy that you lived on, and thank you for pushing the frontier of AI.

Richard

感谢你推动 AI 的前沿。我真的很高兴,感谢这一切和推动,你似乎也培养了很多推动前沿的学生。你是怎么挑选他们的?在过去二三十年里?

Thank you for pushing the frontier of AI. I'm sure really happy and thank you for all of that and push, and you've like been able to sort of educate a lot of students who push the frontier as well. How'd you pick them? How'd you over the last 20, 30 years?

Host

哦,好吧,你给了我谦卑的机会,我喜欢谦卑,并指出所有这些伟大的决定都是偶然发生的,这就是我对学生的感觉。我不觉得我选得很好。我只是有时幸运,有时不幸。我不觉得我特别擅长挑选学生。我现在看着 Kerm。我觉得有时你最终会遇到真正优秀的人。David 选了我,David Silver 选了我。

Oh well, you are giving me opportunities to be humble, like I like to be humble and point out how all these great decisions are just happen, and it's that's the way I feel about students. I don't feel that I choose them very well. I've just sometimes I'm lucky, sometimes I'm unlucky. I don't feel I'm particularly good at picking my students. I'm looking at Kerm now. I think sometimes you end up with really great ones. David picked David Silver picked me.

Richard

是的。我是怎么得到你的,K?

Yeah. How is it that I got you, K?

Kuram

是的,那也是,我完成硕士不是跟你,呃,我本来打算进入工业界,然后我们开始合作一个项目,也是有机地开始的,比如我做过一些工作,Rich 在一次会议上,然后他们提到我做过那个。所以我被拉进来,我们开始合作,进展很顺利,我对那次合作感到非常高兴,也感觉非常好,然后六个月后我们取得了一些进展,把它转化为论文提案就顺理成章了。

Yeah, that was also so I finished my masters not with you, uh, and I was planning to join industry, and then we were collaborating on a project which also just started organically, like there was something I worked on that Rich was in a meeting, then they mentioned that I worked on it. So I got pulled into it, we started collaborating, it went really well, like I felt so happy with that collaboration, which also felt really good about it, and then 6 months down the road we had made some progress and it just made sense to convert that into a thesis proposal.

苦涩教训的起源 The Bitter Lesson's Origin

Richard

所以我从未申请过,也从未问过你是否愿意做我的博士导师。我们一起工作,然后我们觉得这会是一篇不错的论文,之后我才申请了博士。

So at no point did I apply, at no point did I ask, should you be my PhD adviser. We worked together, then we decided this would be a pretty good thesis, and then after that I applied for PhD.

Host

生活总是出人意料。带我们回到 2019 年,你写了《苦涩的教训》,它已成为现代经典。2019 年写这篇文章的时机很有趣,因为 ImageNet 是 2009 年的,AlphaGo 是 2015 年的。是什么促使你在 2019 年反思并写下这篇文章?因为那时大语言模型的 Scaling 范式尚未起飞,但深度学习已经证明了自身。

Life works in unexpected ways. Take us to 2019. You wrote 'The Bitter Lesson,' which has become the modern tome. 2019 was a funny time to be writing that piece because ImageNet was 2009, AlphaGo was 2015. What caused you in 2019 to reflect and write that? Because it was before the current scaling paradigm around large language models had taken off, but it was after deep learning had really proven itself.

Richard

嗯,这是酝酿已久的,你知道,正如《苦涩的教训》所表达的。这是你可以观察很长时间、几十年的东西。而且它至少同样源于我所经历的那一轮符号人工智能。关键在于不要被试图注入人类知识所分心,而只关注问题需要什么以及如何用算力进行 Scaling。我知道我在至少一年前就写过它的几个版本,我还做过演讲,一年前就做过一次演讲,这并不是对那个时刻的特定回应。这是对我长期经验的特定回应,不同的人试图以不同的方式思考如何制造智能系统。

Well, it was a long time coming, you know, as 'The Bitter Lesson' expresses. It's something that you can observe for a long time, many decades. And it's definitely at least as much due to the round of symbolic AI, which I lived through. It's all about not getting distracted by trying to put in your human knowledge and just paying attention to what the problem needs and how you can scale with computation. I know I wrote versions of it at least a year before, and I gave talks, I gave a talk a year before, and it wasn't a particular response to the moment. It was a particular response to my long experience, different people trying to think in different ways about how you can make smart systems.

苦涩教训的本质 Essence of the Bitter Lesson

Host

苦涩的教训的本质是什么?

What is the essence of the bitter lesson?

Richard

是的。你知道,最近我在会议中最常听到的短语可能就是“苦涩的教训”。

Yes. And you know, it may be the phrase that I hear used the most in my meetings these days is 'bitter lesson'.

Host

我想,鉴于这个短语的流行,它可能已经被以各种你最初未预料到的方式曲解和误用了。那么你认为它的本质是什么,你认为人们在理解它时哪里容易出错?

I would imagine, given the popularity of the phrase, it's probably been tortured and misused in different ways that you didn't originally intend it. So what do you think is the essence of it, and where do you think people go wrong in their attempt to understand it?

Richard

是的,你现在让我想起了 X。在我最近的一篇帖子中,我尝试用 26 个字概括苦涩的教训。大意是:不要像传统人工智能多次那样被人类知识分心。相反,专注于能随算力扩展的学习方法,比如搜索和学习。所以它真正关乎的是专注于算法和改进。这并不是说不需要花哨的算法。你需要花哨的算法,但你想要的是能随算力扩展的花哨算法。

Yeah, you're making me think about X now. In my recent post, I tried to do the bitter lesson in 26 words. It goes something like: don't be distracted by human knowledge as AI traditionally has been many times. Instead, focus on learning methods that will scale with computation, like search and like learning. So it's really all about focusing on algorithms and improvements. It's not saying you don't need fancy algorithms. You need fancy algorithms, but you want fancy algorithms that will scale with computation.

Host

而不是随数据扩展。

Rather than scaling with data.

Richard

而不是随人类输入扩展。是的。

Rather than scaling with human input. Yeah.

Host

那么接下来一个问题,如果我能预料的话,大语言模型呢?它们与你的观点一致还是不一致?

And then the question, if I can anticipate, what about large language models? Are they consistent or inconsistent with your...?

Richard

是的。我考虑过这个问题,我认为还有另一篇相关的文章,但结论是它既是苦涩的教训的正面例子,也是反面例子。首先,大语言模型实现了随算力的巨大 Scaling,你可以尽情吸收互联网内容并大幅扩展。所以这是一种通过可扩展的方法获得更强大系统的途径。然后,随着进一步发展,它最终会受到信息的限制。互联网是有限的,很难获得更多示例,而世界是巨大的,世界比我们存储在互联网上的一切都要大得多。所以最终,这似乎可能是,我想那会是一个正面例子,说明当我们过度依赖人类知识时,它最终会阻碍我们。

Yeah. And I thought about this, and I think there's another expost about it, but the conclusion is that it's both a positive example and a negative example of the bitter lesson. First, large language models enabled enormous scaling with computation, and you could just drink in the internet and scale so much. So it was a way of getting much more capable systems just by methods that scale. Then after that, as you go on further, it eventually gets limited by that information. The internet is finite, and it's hard to get more examples, and the world is big, and the world is massively bigger than everything we stored on the internet. So in the end, it seems like it could be, I guess that would be a positive example of when we relied too much on human knowledge and it eventually holds us back.

合成数据与大世界假说 Synthetic Data and the Big World Hypothesis

Host

我能稍微深入一下吗?看起来很多基础模型实验室现在正在做的事情是合成数据生成,以便让我们超越现有的人类互联网这一“化石燃料”。作为 LLM Scaling 范式的一部分,合成数据生成是一种利用算力的通用方法吗?

Can I just push on this a little bit? It seems like a lot of what the foundation model labs are working on right now is synthetic data generation in order to kind of get us beyond the fossil fuel that is the existing human internet. Is synthetic data generation, as part of this LLM scaling paradigm, a general method that leverages computation?

Richard

不,那是个大错误。

No, that's just a big mistake.

Host

为什么?

Why?

Richard

嗯,这是个很大的问题,也许这是下一个大教训。它在阿尔伯塔已经流传了五到十年了。

Well, it's such a big, maybe it's the next big lesson. It's been floating around Alberta for five or ten years.

Host

好的。

Okay.

Richard

我们称之为“大世界视角”或“大世界假说”。库拉姆最终把它写成了一篇论文,有一篇小论文叫《大世界假说》。大世界是指世界是无限大的。有无限多的事物需要学习,你可以让人们生成这些合成数据集,但总会有更多的东西需要学习。正因为如此,如果你能从经验中学习,如果你能把人类从循环中移除,那么你就会有能做任何事的系统,因为你知道世界是巨大的,有很多任务我们希望它们去做,它们就能通过从自身经验中学习来做任何事。

And we call it the big world perspective or the big world hypothesis. Kuram, who eventually wrote it up as a paper, there's a little paper called 'The Big World Hypothesis.' The big world is that the world is infinitely big. There are infinitely many things to learn, and you can have people generating these synthetic datasets, but there would always be more things to learn. And because of that, if you could just learn from experience, if you could remove the humans from the loop, then you would have systems that can do everything, because you know the world is big, there are many tasks we want them to do, and they would be able to do anything by learning from their own experience.

Host

回到合成数据的问题,谁来决定好的合成数据和坏的合成数据?因为我可以写一个程序输出大量会损害程序的合成数据。

Going back to the synthetic data question too, who decides what's good synthetic data versus bad synthetic data? Because I can write a program that can output a lot of synthetic data which would hurt programs.

Richard

现在,我会说是人类决定,这就是瓶颈。你可以让人类决定如何生成这些数据集,但你需要知道什么数据集好、什么数据集坏的人类专家,这种方法才能扩展。所以它受限于人类。

Right now, I would say humans decide, and that's the bottleneck. You can have humans deciding how to generate these datasets, but you need human experts who know what's a good dataset and what's a bad dataset for that approach to scale. So it is bottlenecked by humans.

Host

难道我的损失曲线不能决定吗?比如用这个数据集比用那个数据集我进步了多少?

Doesn't my loss curve decide, like how much better did I get with this dataset versus that dataset?

Richard

对吧?但如果 OpenAI、Anthropic 或所有大型实验室的工程师都去度假了,谁来生成合成数据?这就是问题所在。它不是来自智能体的经验。它不是智能体自己生成的东西。必须有人类来决定生成什么样的合成数据是正确的,这需要人类专业知识。例如,如果你想让一个系统做从物理学角度来说非常具有挑战性的事情,比如你想要一架像蝙蝠一样用回声定位飞行的无人机,那么正确的合成数据是什么?我认为你需要聘请领域专家去弄清楚什么是正确的数据并生成它,然后也许你就能从中学习。但领域专家必须首先存在。所以在那一点上我们受限于人类专业知识。

Right? But if all the engineers at OpenAI, Anthropic, or all the big labs, or engineers went on vacation, who would generate the synthetic data? That's the question. It doesn't come from the agent's experience. It's not something that the agent is generating itself. Some human has to decide what is the right synthetic data to generate, and that requires human expertise. So for example, if you want a system to do something very challenging from a physics point of view, maybe you want a drone that flies with echolocation like a bat, for example, what's the right synthetic data for that? I think you would need to hire domain experts to go figure out what is the right data and generate it, and then maybe you would be able to learn from that. But the domain expert has to exist first. So we are bottlenecked by human expertise at that point.

Host

但你可以有无限的合成世界。现有世界是有限的。

But you can have infinite synthetic worlds. The existing world is finite.

Richard

但让我们回到回声定位的事情,对吧?那是我想要的。我想要一架能通过回声定位来定位自己并移动的无人机。那是我的目标。机器人正在生成自己的经验。所以它完全可以从中学习,但它做不到。不管你生成多少合成数据。不管你生成覆盖 50 个不同宇宙的合成数据。没有人类先弄清楚,它就无法完成那个任务。

But let's go back to the echolocation thing, right? That's what I want. I want a drone that can localize itself and move with echolocation. That's my goal. The robot is generating its own experience. So it could totally learn from its own experience, but it wouldn't be able to. Doesn't matter how much synthetic data you generate. Doesn't matter if you generate synthetic data that captures 50 different universes. It will not allow you to do that task without humans figuring it out first.

Host

但首先,就说合成数据是错的。我的意思是,它不会是正确的。

But first, just say it's the synthetic data is wrong. I mean, it won't be correct.

合成数据的局限 Synthetic Data Limitations

Richard

那将是一个合成世界,不是真实世界,而且这很重要。世界极其复杂。如果你写一个小程序,因为它将是一个生成合成数据的小程序,那将是一个非常小的世界。

It'll be a synthetic world. Won't be the real world and it will matter. The world is incredibly complex. If you write a little program, because it's going to be a little program that'll generate the synthetic data, it'll be a very small world.

Host

所以,比如,对我来说重要的是你现在脑子里在想什么。好吗?为什么你说“我为什么不弄点合成数据?告诉我别人脑子里在想什么。”不,我们不可能有关于别人思想的合成数据。而且别人的思想对我们很重要。你知道,我今天跟你们聊了投资,所以我在乎你们在想什么。我怎么能得到关于这种事情的数据呢?

So, for example, what's important to me is what's going on in your mind right now. Okay? And why you're saying, "Why don't I get some synthetic data? Tell me what's going on in other people's minds." No, there's no way we can have synthetic data for other people's minds. And other people's minds matter to us. You know, like I talked to you guys about investing today, so I care what's going on in your minds. And how could I get synthetic data on such a thing?

Richard

真的,你甚至无法获得任何事物的合成数据。你无法获得关于无人机如何在物理世界中与环境互动的合成数据,包括摩擦、机器人电机的磨损。世界是无限复杂的,任何对它的模拟都像是微观的。大世界假说,简单说就是,世界比你的思想、比任何智能体都要复杂得多。这很明显,因为世界包含许多其他智能体。所以,因为世界极其复杂,你不可能做任何声称最优或完美的事情。你将会是不完美的,你必须进行近似,而且这些近似将是严重的。因此,这最终是我们必须继续学习的原因。如果你想把它当作一个理由,我们必须继续学习,因为我们会遇到这个巨大世界的某个特定部分,我们必须学习一个针对我们所处世界部分的近似,而不是针对我们不在的其他所有部分。

Really, you can't even get synthetic data on anything. You can't get synthetic data on how the drone is going to interact with its environment in the physical world, and the friction, and the wear in the motors of this robot. The world is infinitely complex and any simulation of it is like microscopic. The big world hypothesis, let's say what it is, is that the world is massively more complex than your mind, than any agent. And this is obvious because the world contains many other agents. So because the world is massively complex, you can in no way do anything that might claim to be optimal or perfect. You're going to be imperfect and you have to have approximations, and those approximations will be severe. And so because of that, that is ultimately the reason why we have to continue learning. If you want to think of it as a reason, we have to continue learning because we'll encounter some particular part of this immense world and we'll have to learn an approximation that's tuned to the part of the world we're in, not to all the other parts that we're not in.

Host

是的,我要再追问一次。抱歉我为了争论而争论,但我想理解。

Yeah, I'm going to push on this one more time. And sorry I'm being argumentative for the sake of being argumentative, but I'm trying to understand.

Host

我的理解是,最新一代的自动驾驶汽车公司,很多主要在模拟环境中训练,然后他们,你知道,必须做一些后训练,我猜,以确保它们在真实世界中工作,但这是一个非常有效的流程。

My understanding is that the newest cohort of self-driving car companies, many of them were primarily trained in sim and then they, you know, have to do some sort of post-training, I guess, to make sure they work in the real world, but that it's been a very effective pipeline.

Richard

是的。所以我认为这里要问的重要问题是,有多少工程师参与了构建那个模拟,我们是否准备好说,唯一值得解决的问题是那些我们可以雇佣一大队工程师先做模拟的问题?而且我确信他们必须进行多次迭代,他们制作模拟,在其中学习,意识到存在不可接受的模拟到现实的差距,然后修复它。所以有这种人在循环中修复模拟,就像他们从现实世界获得反馈,人类,然后他们修复模拟。为什么我们不能移除人类,让智能体自己做呢?

Yeah. So I think the important question to ask here is how many engineers were involved in building that simulation and are we ready to say that the only problem worth solving are those where we can hire a large team of engineers to first make a simulation? And I'm sure they had to do multiple iterations where they made the simulation, they learned in it, they realized there was a sim-to-real gap that was not acceptable, then they fixed it. So there is this human in the loop fixing the simulation, like they're getting feedback from the real world, humans, and then they're fixing the simulation. Why can't we just remove the human and let the agent do it itself?

Host

然后当它真正再次驾驶时,意想不到的事情会发生。

And then when it actually drives again, something unexpected will happen.

Richard

那就是他们想要从经验中学习的时候。

And that's when they want to learn from experience.

Host

所以你的观点是,从经验中得来的数据将比人类策划创造的数据多得多。

So your point is there's just so much more data that's going to come from experience than there possibly can be from humans curating creating data.

Richard

是的。而且我认为从模拟中学习显然有价值,而且有一种方法可以做到。智能体可以从自己的经验中学习模型,当智能体学习时,它会更好,因为如果模型不正确,它可以通过持续学习来修复它。如果人类在制作模拟器,那么模型只有在人类发现有问题时才会更新。所以是的,规划很重要。智能体应该从模拟器中学习。但模拟器是他们自己制造的。

Yeah. And I think there is obviously value in learning from simulation and there is a way of doing it. The agent can learn a model from its own experience, and when the agent learns it, it's much better because if the model is incorrect, it can fix it by continuous learning. If the humans are making a simulator, then the model only gets updated when the humans figure out that something is wrong. So yes, planning is important. The agents should learn from simulators. But simulators they make themselves.

去除人类知识 Removing Human Knowledge

Host

好的,我想转到苦涩教训的另一部分:从你的文章中移除人类知识。引用:“为了在短期内取得显著改进,研究人员寻求利用他们对领域的人类知识,但从长远来看,唯一重要的是算力的利用。”而且,你知道,你的前学生戴夫,用 AlphaGo 和 AlphaZero,对我来说,那是移除人类先验的胜利的例子。那个结果让你惊讶吗?我想为什么或为什么不?

Okay, I want to move to another part of the bitter lesson: removing human knowledge from your essay. Quote, "Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation." And if I, you know, your former student Dave with AlphaGo and AlphaZero, for me, that was an example of a triumph of removing human priors. Did that result surprise you? I guess why or why not?

Richard

当然,这让我非常高兴,让我感到被证实了。你知道,结果本来可能走向任何一边。不是因为先验知识有帮助。你知道,先验知识没有错。我在苦涩教训的开头就说了。我说没有理由先验知识和学习知识之间必须有冲突。你可以放一些先验进去,然后开始学习。原则上没有理由这些必须对立。事实上,它们都是关于知识的。生活就是获得知识和拥有知识。为什么这些,你知道,天性和教养变成了敌人?但实际上,先验学习是你已经拥有的,然后你学更多,它们应该是朋友。但正如我在苦涩教训开头所说,在实践中,它们一直是敌人。在实践中,那些对现有人类知识有感情的人最终希望它获胜。所以他们想最小化或否定学习。所以现在我相信你对我的印象是,我是一个热爱学习并想否定先验知识的人。但你知道,我真的是一个对心智感兴趣的人。心智是你有先验知识,然后你得到更多,一旦你得到更多,那又成为你的先验知识,你得到越来越多,这两者协同工作。我最终看起来主要对学习感兴趣,因为世界上其他所有人都在说,你只需要足够的知识,不需要学习。你知道,大型语言模型是,“哦,我们要把所有知识放进系统,大型语言模型在运行时不会学习。”它在与人交谈,在互动,但权重绝对从未改变。所以你知道,我不是奇怪的人。你们才是奇怪的人,认为那是可能的,你们可能,你知道,他们声称能从根本不再学习的东西中制造出博士水平的经验和专业知识。所以你知道,我不是奇怪的人。

Of course, it made me very happy and made me feel vindicated. You know, it could have gone either way. It wasn't that, because prior knowledge can help. You know, there's nothing wrong with prior knowledge. And I say this right at the very beginning of the bitter lesson. I say there's no reason why there has to be a conflict between prior knowledge and learning knowledge. You can put some prior in there and then start learning. There's no reason in principle why these have to be opposed. In fact, they're all about knowledge. Life is gaining knowledge and having knowledge. And why are these somehow, you know, nature and nurture became enemies? But really, prior learning is what you already have, then you learn more, and they should be friends. But as I say at the beginning of the bitter lesson, in practice, they have been enemies. In practice, people who had an affection for existing human knowledge ended up wanting that to win. And so they wanted to minimize or dismiss learning. And so now I'm sure your sense of me is that I'm someone who loves learning and wants to dismiss prior knowledge. But you know, I'm really someone who's interested in the mind. The mind is you have prior knowledge and then you get more, and then once you've gotten more, then that becomes your prior knowledge as you get more and more and more, and these two things work together. I end up appearing to be someone who's interested in learning primarily because all the rest of the world is talking about all you need is enough knowledge. You don't need to learn. You know, large language models are, "Oh, we're going to put all this knowledge into the system and the large language model will not learn when it runs." It's talking to people, it's interacting, but absolutely the weights never change. So you know, I am not the weird one. It's you guys that are the weird one that think that that's possible, that you could possibly, you know, they claim they could make a PhD level experience and expertise out of something that doesn't learn at all anymore. So you know, I'm not the weird one.

Host

所以,你的建议是让算法运行更长更长的时间,然后再喂入

So, your recommendation is let the algorithms run for a much much longer period of time before feeding

Richard

持续学习

continually learn

Host

在你喂它数据之前

before you feed it data

Richard

或先验数据。

or prior data.

Host

所以,沿途滴灌先验数据,先验知识。

So, drip or drip the prior data along the way, prior knowledge.

Richard

所以,两者都很重要。

So, both are important.

智能机器人学习的未来 Future of Learning in Intelligent Robots

Host

但从长远来看,你必须获取并组织新知识的获取。从长远来看,这才是最重要的。当你这样做的时候,会有一些你之前已经获得的知识。比如,当我们展望未来,拥有智能机器人时,这会如何运作?我们是让它们都从头学习,还是复制它们,并让它们从当前状态继续学习?

But in the long run, you've got to gain and structure the gaining of new knowledge. That's what matters in the long run. As you are doing this, there'll be some that you got previously. Like, how will it work if we look into the future when we have intelligent robots? Will we have them all learn from scratch, or will we copy them and ask them to keep learning from wherever they are?

Richard

我的意思是,它们将是数字化的,复制起来很容易。因此,与其花费巨额资金从互联网上重新训练它们,不如直接复制智能体,然后从那里继续学习。所以从某种意义上说,先验知识应该被丢弃,因为你只需要从之前的机器人那里复制它。

I mean, they'll be digital and it'll be easy to copy them. And so, instead of having this huge thing where we're spending zillions of dollars to retrain them from the internet, we'll just copy the agent and keep learning from there. And so in some sense, the prior knowledge should be dismissed because you're just going to copy it from the previous robot.

Host

那么,你能否为我们描述一下,你认为一个从经验中学习的机器或计算机是什么样的?

So why don't you just describe for us what you think a machine or a computer that learns from experience looks like?

Richard

嗯,它可能看起来像一个机器人。有可能,但它也可能完全存在于互联网上。例如,你可以通过互联网路由数据包,并以一种对经验敏感、随时间推移变得更好的方式来做。或者你可以成为与人们交互的用户界面,比如在手机或电脑上,它随时间推移变得更好。是的,就像一个智能助手必须随时间变得更好,必须知道你想要什么。

Well, it could look like a robot. It could, but it also could live entirely on the internet. You could, for example, route packets through the internet and do that in a way that's sensitive to experience and becomes better over time. Or you can be the user interface that's interacting with people, like on your phone or on your computer, and it becomes better over time. Yeah, like an intelligent assistant has to become better over time, has to know what you want.

Host

你是否认为当前流行的基于大语言模型的助手范式——你是否认为这些不是经验学习者或持续学习者?如果是,根本差距是什么?

Would your contention be that the current paradigm of the popular LLM-based assistants—would your contention be that these are not experiential learners or continual learners? And if so, what is the fundamental gap?

Richard

你是认真的吗?我是说,显然它们有记忆;它们学习关于我的记忆。它们在做一些上下文学习。

Are you serious? I mean, obviously they have memory; they learn memories about me. They're doing some in-context learning.

Host

它们的权重从不改变。

Their weights never change.

Richard

顺便问一下,少量权重改变是否足够,还是需要所有权重都改变?

And by the way, is a small number of the weights changing sufficient, or do you need all the weights to be changing?

Richard

嗯,想想创造大语言模型时所有的结构化和新概念生成。那都是权重学习。你希望能够继续这样做。你不希望那只发生一次。

Well, so think of all the structuring and generation of new concepts that went into creating the large language models. All that is the weight learning. And you want to continue to be able to continue doing that. You don't want that to happen just once.

Host

换种说法,我们在发布模型之前做了太多的预训练和后训练?之后它们就不再学习了。

Is another way of saying we do too much pre-training and post-training before we launch the models? They don't learn after that.

Richard

最大的分歧点就是我们不让它们在之后学习。

The only point of big disagreement is we don't let them learn after that.

Host

是的,我们不让它们学习。

Yeah, we don't let them learn.

Richard

我们可以做尽可能多的预训练。那没问题。后训练也可以。但当我使用模型时,它就停止学习了。你可以给它更多上下文;你可以通过给它更多上下文来改变模型的状态。所以它已经学会了,如果状态不同,如果状态有新的内容,那么它会利用它来做下一个预测。但模型并没有在学习。

We can do as much pre-training as we want. That's okay. Post-training is fine. But then when I'm using the model, it stops learning. You can give it more context; you can change the state of the model by giving it more context. So it has already learned that if the state is different, if the state says something new, then it will use that to make the next prediction. But the model is not learning.

Host

Cursor 的 Tab 自动补全模型——它确实会根据那些模型的权重变化而更新。

Cursor's tab autocomplete model—it does get updated based on those models' weights changing.

Richard

那些权重会变化。这是两个例子,比如 Cursor 的 Tab 和 Composer,它们也在更新。是的,这是持续学习的两个例子,但它可以做得更好。据我所知,它们的做法是,很多人使用 Tab。他们从数百万或数千用户那里收集所有这些数据,然后从这批数据中对策略进行一次更新。所以想象这可以工作,但想象我想教这个模型一些特定的东西。我不想和十万个其他人争论他们想教他们的模型什么。我想教我的模型一些非常具体的东西,而且我想在我的模型版本上做。我不在乎模型从其他人那里获得的共享知识。所以这是一种非常低效的方式。

Those weights change. Those are two examples, like Cursor's Tab and I think the Composer, they were also updating. Yeah, those are two examples of continual learning, but it can be much better. So the way they do it, as far as I understand, is a lot of people are using Tab. They collect all this data from millions of users or thousands of users, and then they do one update of the policy from this batch data. So imagine this could work, but imagine I want to teach this model something specific. I don't want to fight with 100,000 other people about what they want to teach their models. I want to teach my model something very specific, and I want to do it to my version of the model. I don't care about the shared knowledge that the model has coming from other people. And so it's a very inefficient way of doing it.

Host

似乎目前的做法是,有一些基本技能可能是在权重中学习的,对每个人都通用,然后个性化以上下文的形式发生。对吧。

It seems like the way that this is currently done is that there are fundamental skills maybe that are learned in the weights that are common to everybody, and then there's personalization that happens in the form of context. Right.

Richard

是的。是的。

Yeah. Yeah.

Host

这不是学习应该如何运作的正确心智模型吗?所有的上下文都应该存在于权重本身吗?

Is that not the right mental model for how learning should work? Should all the context live in the weights themselves?

Richard

所以上下文也可以存在于状态中;可以两者都有。但你仍然需要能够更新权重。所以如果我给你一个例子,一些非常好的案例研究是关于人类残疾的。当人类经历一些改变他们思维的事情,或者一些传感器,你可以看到他们适应。例如,我们有本体感觉;我们有内部传感器告诉我们身体的位置,我们用它来走路。有些情况下,人们完全失去这种能力,然后他们根本无法走路,因为那真的是他们走路策略的基础。它根植于大脑,但经过两三年,他们可以通过看自己的脚来重新学习走路。所以通过视觉反馈。所以大脑在某种意义上具有惊人的可塑性,它可以学习很多东西。一个已经成立 20 年的事实,当它不再成立时,它可以去更新并摆脱它。我认为这种能力非常有用,我们希望在我们的系统中拥有它。

So context can be in the state too; it could be both. But you still need to be able to update the weights. So if I give you an example, some really good case studies are with human disabilities. When humans go through something that changes their mind, or some sensors, you can see them adapt. For example, we have proprioception; we have internal sensors that tell us how the body is positioned, and we use this for walking. There are cases where people lose this ability completely, and then they can't walk at all because that is literally the foundation of their walking policies. It is ingrained in the brain, but then over the course of two or three years, they can learn to walk again by looking at their feet. So visual feedback through that. So the brain is insanely plastic in the sense that it can learn a lot of things. Something that has been true for 20 years, when it stops being true, it can go and update that and get rid of that. And that is the capability I think is extremely useful and we would want in our systems.

Host

从人类婴儿或动物学习的方式中,我们能学到什么?你从中汲取了多少灵感?

What is there for us to learn from how human babies or animals learn, and how much inspiration do you take from that?

Richard

嗯,我们汲取了很多灵感。我们不把它作为 AI 必须像自然系统(如婴儿、人或动物)那样行为的要求。但它是灵感的来源,是灵感而非来自动物学习的约束。

Well, we take a lot of inspiration. We don't take it as a requirement that the AI has to behave like the natural system, like babies or people or animals. But it's a source of inspiration, inspiration but not constraint from animal learning.

Host

要与苦涩的教训保持一致。

Be consistent with the bitter lesson.

Richard

是的。

Yeah.

Host

你认为我们应该从生物学习的方式中汲取灵感,而这种方式在当今系统中并不存在?

Where do you think we should most seek to draw inspiration from the way that biological learning works that is not present in today's systems?

Richard

我觉得我现在只是在发表意见,但它们只是显而易见的意见。所以我认为很明显,没有动物通过监督学习来学习,因为我们不会得到肌肉应该如何抽搐的例子,而那是我们的输出。

I feel like I'm just giving opinions now, but they're just obvious opinions. So I think it's apparent that no animal learns by supervised learning because we don't get examples of how our muscles should twitch, and that's our output.

Host

但所有的学校教育都是监督学习。

But all of school is supervised learning.

Richard

不,绝对不是。但即使如此,学校只是我们所学的一小部分。比如我们学会看,学会走路,学会世界如何运作。但即使在学校,也没有人告诉我们肌肉应该如何抽搐。

No, absolutely not. But even if it was, school is like a tiny fraction of what we learn. Like we learn to see and we learn to walk, and we learn how the world works. But even in school, no one tells us how we should twitch our muscles.

Host

我获得的知识和技能来自学校的监督学习。

The knowledge and skills I acquired were from supervised learning in school.

Richard

我不想说向他人学习、从他人那里传递知识不重要。它极其重要,语言也极其重要。

I don't want to say that learning from others, transmission from others, is not important. It's extremely important, and language is extremely important.

从经验学习与学校教育 Learning from Experience vs. School

Richard

但我们缺的是什么?没有监督学习,没有给我们的目标。你听到正确答案——比如法国的首都是什么,我们知道答案是巴黎。但没人告诉我该怎么发音“巴黎”。你说答案是巴黎,我听你说,听到你的话,然后我会做其他肌肉动作来产生答案。巴黎。这不是字面意义上的监督学习。总之,是的。所以我认为这确实是真的。首先,学校是无关紧要的。松鼠不去学校,也不那样学习。动物不那样学习。学校是非常特殊的东西,甚至我们直到几百年前才有的。它不是智能的本质,把它当作学习的主要例子是一种干扰——这是我们作为动物不做的事情。

But what are we missing? There is no supervised learning, there are no targets given to us. You hear the right answer—like what's the capital of France—and we know the answer is Paris. But no one tells me how I should pronounce Paris. You say the answer is Paris, I listen to you, I hear your words, and I make some other muscle motions to produce the answer. Paris. It's not literally supervised learning. Anyway, yeah. So I think it's really true. The first thing is school is irrelevant. Squirrels don't go to school and learn that way. Animals don't learn that way. School is a very special thing that even we didn't have until a few hundred years ago. It's not part of the essence of intelligence, and it's a distraction to think of that as your primary example of learning—this thing which we didn't do as animals.

Host

我希望你当时能告诉我父母,他们强迫我好好上学,应付所有的条条框框。

I wish you had been around to tell my parents that forced me to have good go to school and deal with all the structure.

Richard

问题是,松鼠擅长从树上跳下来,但松鼠不能证明数学定理。如果我想学习如何证明数学定理,我就去上学。

The thing is, squirrels are wonderful at jumping off trees, but squirrels can't prove math theorems. And if I want to learn how to prove a math theorem, I go to school.

Host

是的。它们也没有 DVD 和 iPod,你知道,很多东西。它们能做我们做不到的事。但数学定理——是的,它们也不下棋。这更像是一个悖论。这些是我们认为非常智能的高级事物,但对计算机来说却很容易,而所有这些普通的事情却很难,比如移动、用注意力看东西等等。

Yeah. They also don't have DVDs and iPods, you know, a lot of things. They can do things that we can't do. But math theorems—yeah, and they don't play chess. It's more of a paradox. These are advanced things we think of as really intelligent, but they're sort of easy for computers to do, as opposed to all these regular things that are hard, like moving and seeing with attention and everything.

Richard

我认为监督学习是好事。我喜欢寻找显而易见的东西。没有人通过给我们例子来告诉我们如何抽动肌肉,因为他们不可能做到——我们必须自己摸索。

I think supervised learning is a good thing. I like to look for obvious things. No one tells us how to twitch our muscles by giving us examples, because they couldn't possibly—we've had to figure that out.

Host

是的。

Yeah.

Richard

而且他们的答案会是错的,对吧?所以如果我以和 Rich 完全一样的方式移动我的嘴巴、舌头和声带来说“巴黎”,我肯定发出的声音会非常不同。所以从某种意义上说,Rich 或任何人都不知道用我的身体发出声音的正确方式。只有我知道。

And their answer would be wrong, right? So if I moved my mouth and my tongue and my vocal cords exactly the same way that Rich does to pronounce Paris, I'm sure a very different sound would come out. So in some sense, Rich or no one knows the right way of producing a sound with my body. Only I know that.

Host

是的,在我看来,许多最原始的感官运动能力,尤其是与物理世界中的运动相关的能力,我同意你的看法,这似乎是天生从经验中学习的。但在我看来,还有更高层次的抽象让我们更接近人类的伟大之处,而其中很多并不存在于这种低层次的感官运动学习中。你的世界模型,我想,是否涵盖了从感官运动学习一直到更高层次?

Yeah, it seems to me that many of the most raw sensory motor capabilities, especially related to movement in the physical world, I agree with you that that seems something inherently learned from experience. It seems to me though that there are higher levels of abstraction that bring us closer to what makes humans great, and much of that doesn't live in this low level of sensory motor learning. Does your world model, I guess, span sensory motor learning all the way up?

Richard

是的,这就是雄心。绝对如此。顺便说一句,松鼠能做很多极其抽象的事情。

Yeah, that's the ambition. Absolutely. And squirrels, by the way, can do some enormously abstract things.

Host

松鼠能做的最酷的事情是什么?

What's the coolest thing a squirrel can do?

Richard

嗯,它总能进入你的喂鸟器,无论你设置什么障碍。它能找到新的跳跃和攀爬方式,做很多事情。

Well, it can always get into your bird feeder, no matter what obstacles you put in the way. It can find new ways to jump and climb and do lots of things.

Host

好的。

Okay.

Richard

它们计算轨迹相当好。动物很擅长理解物理世界,而不需要我们认为自己在做的那种心理计算,当我们考虑把自己发射到太空时。

You calculate trajectories pretty well. Animals are pretty good at understanding the physical world without the mental calculations that we think we are doing when we think about launching ourselves into space.

Host

摔倒时,它们能实时以正确的方式防止受伤。

Breaking a fall, they can do it in real time the right way to prevent injuries.

Richard

好的,有道理。我认为这只是程度问题——我喜欢认为其他动物与人类非常接近。我认为试图强调我们做得不同、我们与动物的不同之处是傲慢的。最好看到共同点,我认为我们只是程度问题。这是程度,当然社会和文化给了我们巨大的优势。语言给了我们巨大的优势。

Okay, fair enough. I think it's just a question of degree between—I like to think that other animals are very close to humans. I think it's hubristic to try to emphasize what we do differently, how we're different from animals. It's better to see the commonalities, and I think we are just a question of degree. It's degree, and of course society and culture give us big advantages. Language gives us big advantages.

Host

我能就其中一些再追问一下吗?

Can I just push on some of this?

Richard

好的,请讲。

Yeah, go ahead.

Host

因为我想支持 Sonia 的观点。所以,我相信动物和儿童从经验中学习,并从经验中学习做不可思议的事情。当我儿子两三岁或四岁时,我想,“哇,这真的很有趣,我儿子能学到这些东西,没有人真正教他怎么做。”但同时,Sonia 说的是,让人类独一无二的是——能够去外太空,建造火箭——这些不是 100% 从经验中学到的。因为在你发射火箭之前,你实际上必须抽象地思考,这种方式不是从经验中学到的,因为你不知道它是否会成功。你必须想象它。我们如何教机器想象以前不存在的东西?这可能就是我们试图推动的事情,因为我们还不太理解。

Because I want to back up, Sonia. So, I believe animals and children learn from experience and do incredible things learning from experience. And when my son was two or three or four, I'm like, "Wow, this is really interesting that my son can learn these things without nobody really teaching him how to do these things." But at the same time, what Sonia is saying is like what makes human uniquely human—to be able to go to outer space, build a rocket—those are not things that are learned 100% from experience. Because before you launch the rocket, you actually have to abstract thinking through it in a way that is not learned from experience, because you don't know if it's going to work or not. You have to imagine it. How do we teach a machine to imagine things that were not available before? That's probably the thing that we're trying to push on, because we're not quite understanding that.

Richard

我们绝对同意你的观点。你必须能够计划。你必须能够想象。

We're absolutely going to agree with you there. You have to be able to plan. You have to be able to imagine.

Host

是的。

Yeah.

Richard

你会说一千年前的人类,在他们完成我们谈论的大部分事情之前,他们是否同样聪明?例如,如果那个时代的人接触到这种新文化,他们能否获得同样的技能并开始做有用的事情?

Would you say that humans a thousand years ago, before they had done most of the things we're talking about, were they as intelligent? If, for example, someone from that era was exposed to this new culture, would they be able to get the same skills and start doing useful things?

Host

即使在过去的 1 万年里,我认为人类大脑并没有进化那么多,因为从根本上说是同一台机器。

Even over the last 10,000 years, I don't think the human brain has evolved that much, because fundamentally the same machine.

Richard

从根本上说是同一台机器,但我们已经积累了 1 万年的知识。

Fundamentally the same machine, but we've built up 10,000 years of knowledge.

Host

是的。

Yes.

Richard

而且我通过上学、通过监督学习,学到了 1 万年的知识。而且我学到的速度比通过经验学习要快得多。所以我认为你完全正确。我们希望我们的系统从经验中学习,而它们经验的一部分将是接触我们的文化,然后了解我们的文化。它们应该从中学习。这都很好。但让我们谈谈当有人做出范式转变的事情时。每个人都举爱因斯坦的例子,但我认为还有很多例子。学习也是那样——比如看待学习与编程事物。所以当这些范式转变发生时,我会说是一个积累了所有这些知识的人,然后从他们的经验中,他们构建新的抽象,用它们进行规划,然后发现新知识。而那种提出新抽象、然后学习世界模型并用它们进行规划的能力——那个问题——那个能力在我们当前的系统中完全缺失。

And I get to learn 10,000 years of knowledge by going to school through supervised learning. And I get all that much faster than trying to learn through experience. So I think you're totally right. We would want our systems to learn from experience, and part of their experience would be getting exposed to our culture and then learning about our culture. They should learn from that. That's all good. But let's talk about when someone goes and does paradigm-shifting things. Everyone gives the example of Einstein, but I think there are many examples. Learning is that too—like looking at learning versus programming things. So when these paradigm shifts happen, I would say it's a human who has accumulated all this knowledge and then from their experience they're building new abstractions, they're planning with them, and then they're discovering new knowledge. And that skill of coming up with new abstractions and then learning world models and planning with them—that problem—that skill is totally missing in our current systems.

范式转变与学习模型 Paradigm Shifts and Learning Models

Host

而且你可以在人类知识的边缘暴露这一点,但你也可以在感觉运动流层面研究这个问题。所以我们不是反对这个原则。我们需要形成抽象,以便在高层次上推理。你们快要做到我说过我们绝不应该做的事情了,那就是争论先验知识重要还是获取知识重要。你知道这就是你们刚才说的。你说你仍然需要学习东西。而你在说,“哦,我可以从我的文化和先验知识中获得东西,但这些不应该互相争斗。”

And you can expose this at the edge of human knowledge, but you can also study this problem at the sensory motor stream level. So we're not arguing with the principle. We need to form abstractions so we can reason at a high level. You guys are coming close to doing that thing that I said we should never do, which is argue is prior knowledge important or gaining knowledge important. You know that's what you guys just said. You said it's a you're still going to have to learn things. And you're saying, "Oh, I can get things from my culture and from prior knowledge, but these should not fight for each other."

Host

关于范式转变的确切问题,我们如何创造一台机器,让它理解何时该转变范式?

On the exact thing around paradigm shifts, how do we create a machine that understands when to shift the paradigm?

Richard

是的,我认为是通过它的经验,对吧?所以,它必须通过自己的经验。它不能依赖人类知识,因为我们假设人类看到一种范式,而我们想要一种不同的看待事物的方式。所以通过它的经验,它必须找到更好的东西。也许它泛化得更好,做出更好的预测。也许它在其他方面更好,但它必须通过自己的经验。我们领域中看不到的重大挑战,我们尚未看到的能力,是学习一个模型然后用该模型进行规划的能力。我们可以做数学的事情,我们可以做 AlphaGo,因为游戏我们知道模型。我们知道走法如何运作,在数学中我们知道运算符是什么。你知道我们知道 Lean 会带我们从一种知识状态到证明状态再到下一种状态。但如果我们必须学习模型,就没有——我要说出来。这可能是一个奇怪的例子,但也是一个反例,但我看到在我们领域中没有学习模型然后用模型进行规划的例子,至少没有像自我发现的抽象那样。所以有人说我只需要学习下一秒或下一毫秒发生什么的模型。但这不是我们模型的工作方式。我们的模型更抽象。我们的模型相当不同。

Yeah, I think through its experience, right? So, it would have to through its own experience. It can't rely on human knowledge because we're assuming the humans see one paradigm and we want a different way of looking at things. And so through its experience, it has to find something that is better. Maybe it generalizes better and makes better predictions. Maybe it's better in some other ways, but it has to be through its own experience. The big challenge that we don't see in our field, the ability we don't see in our field yet is the ability to learn a model and then plan with a model. We can do the math things and we can do AlphaGo because the games we know the model. We know how the moves work and in math we know what the operators are. You know we know Lean will take us from one state of knowledge to the state of the proof to the next state. But if we have to learn the models there are no—I'm going to say it. It's probably maybe a weird example but a counter example but I can see there's no instances of learning the model and then planning with the model in our field at least not with like self-discovered abstractions. So there are people who say I'm just going to learn a model of what happens in the next second or next millisecond. But that's not how our models work. Our models are more abstract. Our models are quite different.

Host

所以,我喜欢你们在这里做的事情之一是你们不只是坐在那里高谈阔论或哀叹世界的现状。你们非常注重行动。这就是你们创办公司的原因。那么,让我们开始谈谈这个。2022 年,Rich,你制定了一个非常具体的 12 点计划,即艾伯塔人工智能研究计划。也许跟我们说说这个。

So, one of the things I like about what you're doing here is you're not just sitting around pontificating or lamenting the state of the world as it is. You're very action-oriented. It's why you've started a company. So, let's start talking about that a bit. In 2022, Rich, you laid out a very specific 12-point plan, the Alberta plan for AI research. Maybe tell us about that.

Richard

艾伯塔计划之所以出现,是因为我们确实有总体想法,但也需要将它们转化为更小的部分。所以这 12 个步骤是试图具体化特定的部分。有一个非常重要的早期步骤二,即持续深度学习,我们认为这几乎是最重要的,因为它解锁了其他一切。如果你能做持续深度学习,你就能不断更新你的世界模型。然后如果你知道如何做对抽象,后半部分步骤都是关于如何让抽象正确的。

So the Alberta plan came about because we did have general ideas but we also needed to convert them into smaller chunks. So the 12 steps are an attempt to crystallize particular chunks. There's a very important early step two which is continual deep learning and we think that one is like almost the most important because it unlocks everything else. If you could do continual deep learning, you could then continually update your model of the world. And then if you knew how to do the abstraction rights and like the second half of the steps are all about how to get the abstractions right.

Richard

所以,不仅仅是把抽象做对,我是什么意思?我不是说得到正确的抽象,因为没有人能说正确的抽象是什么。这取决于你所处的世界。你的智能体必须为它所处的任何世界学习正确的抽象。所以你知道,也许那些是两个关键的事情:你必须找到正确的抽象,然后你必须能够做持续深度学习。我认为这个领域的很多人意识到我们需要模型,我们需要用它们来规划,但抽象告诉我们模型应该以什么为条件。所以模型应该预测什么?你将要做什么,然后会发生什么,更重要的是,那从哪里来?所以我非常喜欢精英运动员的例子。如果你问精英运动员他们如何做某些事情,他们会有奇怪的专业术语来做非常具体的事情。他们就像,你知道,我做这件事,他们会给它起个名字。如果他们交流,有时他们甚至没有名字,如果他们只是独自做。那么,他们是如何想出那些抽象的呢?在某种意义上,这是缺失的关键东西,艾伯塔计划的后半部分回答了这个问题。

So, and not only by abstractions right, what do I mean by I don't mean get the right abstractions because no one can say what the right abstractions are. That depends on the world that you're in. Your agent would have to learn the correct abstractions for whatever world it's in. And so you know if you maybe those are the two key things you have to find the right abstractions and then you have to be able to do continual deep learning. I think that a lot of the people in the field realize that we need models we need to plan with them but the abstractions tell us what the model should be conditioned on. So what should the model predict? What are you going to do and then something is going to happen and more importantly where would that come from? So I really like the example of elite athletes. If you ask elite athletes about how they do certain things, they would have weird niche terminologies for doing very specific things. They were like, you know, I do this thing and they would have a name for it. If they communicate, sometimes they don't even have a name for it if they're just doing it alone. So, how did they come up with those abstractions? That's in some sense a crucial thing that's missing that the later half of Alberta plan answers.

Host

我们能谈谈持续深度学习这部分吗?它是今天存在的算法差距,还是仅仅是实际部署基础设施数据隐私差距?因为如果我想做所谓的基于用户交互的朴素权重更新,我今天就可以做到,对吧?那么在你看来,我们实现持续深度学习最缺少的是什么?

Can we talk about the continual deep learning part? Is it an algorithmic gap that exists today or is it a just a practical deployment infrastructure data privacy gap because if I wanted to do call it naive updating of weights based on user interaction I can do that today right and so what in your opinion is the biggest thing that we're missing to kind of get to continual deep learning?

Richard

是的,这绝对是算法差距。你可以做朴素的事情,但你会看到各种问题。例如,如果你说我要取一个样本,然后用那个样本更新我的整个模型。你会遇到这个问题,现在模型中的所有先前知识都会受到负面影响。目前我们绕过这个问题的方式正是 Cursor 所做的。他们不使用一个例子。他们使用来自大量用户的大批量。所以在你可以拥有那种情况的用例中,你可以做持续学习。但大多数用例你没有那种情况。大多数用例你有一个单一的数据流,然后如果你应用朴素的方法,它会以非常破坏性的方式完全摧毁你的先验知识。

Yeah, so it's absolutely an algorithmic gap. You can do the naive thing but then you'll see all sorts of problems. So for example if you say I'm going to take one sample and then I'm going to update my whole model with that one sample. You will run into this problem that now all of the previous knowledge in the model it's impacted negatively. And the way currently we get around this is exactly what Cursor does. They don't use one example. They use a large batch coming from a lot of users. So in use cases where you can have that you can do continual learning. Well most use cases you don't have that. Most use cases you have a single stream of data and then if you apply it do the naive thing it just completely destroys your prior knowledge in a very destructive way.

Host

灾难性遗忘。

Catastrophic forgetting.

Richard

是的,就是那个。

Yeah, that is the yeah.

Host

但有了正确的算法,它是完全可以治愈的。

But it's totally curable to have the right algorithm.

Richard

是的,完全正确。那是什么呢?

Yeah, exactly. What is it?

Richard

嗯,你知道首先你需要做我们所说的步长优化,这意味着你网络中的每个权重都必须有一个单独的步长。所以有些会移动得快,有些会移动得慢,你将不得不为每个权重元学习步长。你的大部分网络将具有微小步长的权重。所以当你训练一个新例子时,它们不会被破坏。

Well, you know first you need to do what we call step size optimization and that means every weight in your network has to have a separate step size. So some will move fast, some will move slow and you will have to meta-learn the step sizes for each weight. Most of your network will have weights that have tiny step sizes. So then when you train on a new example, they don't get destroyed.

Host

只发生在正确的地方。

Happens just to the right places.

Richard

然后其次,你必须使用某种形式的生成和测试,这是在特征空间中。所以你提出新的特征或新的单元,而不遵循梯度,因为梯度是一个非常缓慢的过程。你只在你知道它是有帮助的方向上移动,这总是非常缓慢,并且不会给你一条路径来变得越来越复杂并持续学习。你需要有某种东西,它只是提出一堆新的单元,然后从那里开始,我想。

And then secondly, you have to use some form of generate and test which is in feature space. So you come up with new features or new units and without following gradients because gradients is a very slow process. You only move in a direction if you know it's the helpful one and that's always going to be very slow and doesn't give you a path to grow more and more complex and to have sustained learning. You need to have something that just proposes a bunch of new units and then goes from there I guess.

持续反向传播与新算法 Continual Backprop and New Algorithms

Richard

所以我可以举一个具体的例子,让它至少变得具体一点:我们有一种叫做“持续反向传播”的算法,几年前发表在《自然》杂志上。它和反向传播几乎一样,但你会不断植入新的单元种子,这些单元用随机权重重新初始化。反向传播只在初始时刻有随机权重,随着时间推移,那些随机性、由随机性带来的多样性都会被消耗殆尽。而持续反向传播会不断注入一点随机性,一点“生成并测试”的机制。反向传播的操作就是测试者,所以你需要这些。如果你把这些很好地结合起来,我认为你会迎来新一代、性能大幅提升的持续深度学习。这就是我们未来几年希望做的事情。

So there is a specific thing I can say to make it at least concrete, which is to say we have this algorithm called continual backprop, which we published in Nature a couple of years ago. It is exactly like backprop, but you also plant new seeds of units that are newly initialized with random weights. Backprop only has random weights at the beginning of time, and then as you go on, all that randomness, all that variety from the randomness, gets used up. With continual backprop, you keep injecting a bit of randomness, a bit of generate and test. The operation of backprop is the tester, so you need that. And if you put those together really well, I think you'll have a new generation of massively superior continual deep learning. And that's what we hope to do in the next couple of years.

Host

太棒了。你认为这些算法能应用到当前人们 Scaling(规模扩张)大语言模型、并试图让它们在没有灾难性遗忘的情况下进行持续学习的现状中吗?

Wonderful. Do you think that these algorithms can be applied to the current state of affairs with people scaling LLMs and trying to get them to do continual learning without catastrophic forgetting?

Richard

是的,绝对可以。我认为……我不认为你可以拿一个现有模型说“我就开始用这些算法更新它”,因为这些算法是元学习如何学习的。所以实际上你必须说“我要从头开始学”。比如说,学习一个新的基础模型,但我要用这些新算法来学。这些新算法除了学习知识,还会学习如何学习未来的东西。所以它们同时在学习两件事。然后我认为你就能在没有灾难性遗忘的情况下学习新东西了。

Yeah, absolutely. I think it's... I don't think that you could take an existing model and say, 'I'm going to just start updating it with these algorithms,' because these algorithms meta-learn how to learn. So really you have to say, 'I'm going to learn from scratch.' So let's say learn a new foundation model, but I'm going to learn with these new algorithms. These new algorithms, in addition to learning the knowledge, they're also going to learn how to learn future things. So they're learning two things at the same time. And then I think you would be able to learn new things without catastrophic forgetting.

Host

从当前状态来看,你们公司试图同时做这两件事,这是不是最激进的事情?

Is that the most radical thing that you're trying to do in your company, in terms of from the current state of affairs, to try to do these two things at the same time?

Richard

最激进的事情是……我觉得那是……

Most radical thing is... I think that's...

Host

这又回到了“我没疯,是别人都疯了”那句话。

This goes back to 'I'm not crazy, everyone else is crazy.'

Richard

是的,那也许不算完全激进。在 2016 到 2018 年期间,很多人都在大量探索这些想法。他们是在一个更受限的环境下做的。所以他们会说“我们有一个问题分布”,然后在这个特定案例中我们会这么做,而我们想从单一的经验流中做到这一点。所以我们的方法应该更通用。所以我认为很多人已经探索过这个,但没有人探索过这种通用设置,即最终算法能适用于所有地方。

Yeah, that's perhaps not totally radical. There was a point in like 2016 to 2018 where a lot of people were exploring these ideas quite a bit. They were doing it in a much more limited setting. So they would say, 'We have a distribution of problems,' and then in this specific case we'll do it, whereas we want to do it from a single stream of experience. So our method should be more generally applicable. So I think many people have explored this, but no one has explored this in the general setting where the resulting algorithm would be applicable everywhere.

Host

那么你们的新公司试图做的最激进的事情是什么,是别人没在做的?

So what would be the most radical thing that your company, your new company, is trying to do that other people are not doing?

Richard

最有雄心的东西是什么?记住,我不觉得自己奇怪,所以我不想说……最有雄心的东西是什么?我认为是试图拥有完整的知识谱系,既包括微小的事物,也包括宏大的事物。你知道,比如思考如何把一架飞机从一个城市带到另一个城市——那是一个非常宏大的……它更像你那个太空飞行的例子,但只是更符合常识,因为我们很多人都坐飞机,而且我们所有人都在生活的方方面面使用抽象,甚至松鼠也用抽象。所以拥有从微小到宏大的知识谱系,并以统一的方式对待它,并且能够让它自我维持。你知道,大问题总是:你有一个基于知识的系统,是什么让其中的知识保持正确?那么,是什么让大语言模型中的知识保持正确?嗯,人们做了很多后训练,他们确保它是正确的,然后在那之后冻结它。所以那就是让它保持正确的东西。但事实上,我们的心智——我们总是在改变事物,然而有某种东西让它保持组织性和连贯性,并让它回到一个好的状态,而不是漂移到疯狂之地。我认为我们最大的雄心就是:拥有一个自洽的心智,能够持续训练自己,并让它保持连贯。

What's the most ambitious thing? Remember, I don't think I'm weird, so I don't want to say... What's the most ambitious thing? I think it is to try to have the full spectrum of knowledge, both about the tiny things and about the big things. You know, like thinking about how you take an airplane from one city to another—that's a very big... it's more like your space flight example, but it's just kind of more commonsensical to think about because we all—many of us take airplanes—and all of us use abstractions in all kinds of our life, and even the squirrels use abstractions. So to have that spectrum of knowledge, from the small to the big, and to treat it in a uniform way, and to be able to have it self-maintaining. You know, the big question is always: you have your knowledge-based system, and what keeps the knowledge in it correct? Well, what keeps the knowledge correct in a large language model is—well, people did a lot of post-training, and they made sure it was correct, and then they freeze it after that. So that's what keeps it correct. But really, our minds—we are always changing things, and yet something keeps it organized and coherent and settling back into a good place rather than drifting off into crazy land. That is, I think, our biggest ambition: to have a mind that is self-consistent and can keep training itself and making it coherent.

Host

连贯。

Coherent.

Richard

我喜欢这个。

I love that.

Host

我能问一下吗——这几乎看起来是一个如此雄心勃勃的愿景,而且所有这些都能统一到一个单一心智中的想法是如此雄心勃勃。

Can I ask—it almost seems that it's such an ambitious vision, and the idea that all these things can be unified into a single mind is so ambitious.

Richard

这是触手可及的。我认为这是触手可及的。现在是 2026 年。是的。而且我们的计算机如此之快。

It's within reach. I think it's within reach. It's here, it's 2026. Yeah. And our computers are so fast.

Host

你知道,它是雄心勃勃到遥不可及,还是我们已经隐约知道所有步骤该如何完成?

You know, is it so ambitious it's out of reach, or do we already have inklings of how all the steps can be done?

Richard

我认为我们有愿景和隐约的想法。我不认为这是不合适的。

I think we have a vision and inklings. I don't think it's inappropriate.

Host

你的愿景涉及一个 20 瓦的万亿参数模型。这看起来相当雄心勃勃。

Your vision involves a trillion parameter model with 20 watts. That seems pretty ambitious.

Richard

在某种意义上,这是雄心勃勃的。以当前技术,我会说这也是不可能的。比如,仅将一万亿参数存储在内存中,用现有的内存技术可能就会消耗超过 20 瓦的能量。但我们真正在思考的是——好吧,事情在变好,计算变得更便宜,能效也在提高。那么 5 到 10 年后我们会处于什么位置?我认为 5 到 10 年,有了正确的算法,我们完全可能处于一个这样的世界成为可能的状态。

That is ambitious in some sense. With current technology, I would say it's also impossible. Like just storing a trillion parameters in memory would probably use more than 20 watts of energy with memory technologies. But we are really thinking of—okay, things are getting better, computation is getting cheaper, it is getting more energy efficient. So where would we be in 5 to 10 years? And I think 5 to 10 years with the right algorithms, we can totally be in a world where this would be possible.

Host

所以 5 到 10 年是摩尔定律的两个数量级。这是一个标准的改进。如果我们每 18 个月翻一番,10 年就能带来两个数量级的提升。所以为了让 Kerm 的说法合理,那么今天你应该能用——多少?20 瓦,两个数量级——2000 瓦。如果你今天能用 2000 瓦做到,那么 10 年后你就能用 20 瓦做到。

So 5 to 10 years is two orders of magnitude of Moore's law. It's a standard improvement. If we double every 18 months, 10 years would give you two orders of magnitude. And so for Kerm's statement to be plausible, then today you should be able to do it for—what? 20 watts, two orders of magnitude—2,000 watts. If you can do it with 2,000 watts today, then in 10 years you'll be able to do it for 20 watts.

Richard

你认为你能用 2000 瓦做到吗?研究实验室里有很多人能获得远超这个的资源。

You think you can do it for 2,000 watts? You have lots of people at research labs that have access to way more than that.

Host

是的,我认为即使现在,有了正确的算法,我们也能比那更高效。

Yeah, I think we can be more efficient than that even now with the right algorithm.

Richard

如果我们能比那更高效,那为什么我们没做到?并不是说人们就想花光所有的钱,花光我们所有的钱。

If we can be more efficient than that, then why aren't we? It's not like people just want to spend all their money, spend all our money.

Host

有时候看起来他们确实想。

Sometimes it seems like they want to.

Richard

所有的钱。有时候看起来他们确实想。

All money. Sometimes it seems like they want to.

Host

是的,我……

Yeah, I...

Richard

不是吗?

Doesn't it?

Host

我觉得他们就是用大量能源来展示自己是真男人。

I think that's how they show they're real men, by using lots of energy.

Richard

至少当我看到不同的研究团队时,我甚至看不到有人相信这是可能的。我认为如果你不相信它,你就不会去解决那些技术问题并攻克它们。

At least when I look at different research groups, I don't even see anyone believing it's possible. And I think if you don't believe in it, you're just not going to work on the technical problems and work through them.

Host

是它不可能,还是系统中有太多浪费?到底是哪一个?有没有人知道如何高效地做到?

Is it that it's not possible, or is it that there's so much waste in the system? Like, which one is it? Is there someone who knows how to do it efficiently?

Richard

是的。然后同一个实验室里有 10 倍数量的人在忙其他事情。所以从某种意义上说,十个人里有九个在浪费。我的看法是,我们陷入了一个局部最小值。

Yeah. And then there's 10 times the number of people in the same lab doing all these other things. And so nine out of 10 people are wasting, in some sense. The way I think about it is that we are stuck in a local minimum.

向新范式过渡 Transition to New Paradigms

Richard

所以如果我们想迈向这种新算法,几乎不可能不先经历一段变糟的时期。当我们开始探索这些新方向时,你不会从第一天就得到最先进的性能。但这是因为这是一个不同的范式。但那条路会以更高的能量达到类似的性能,而这些大型实验室如此执着于他们喜欢的产品,他们不可能去走一条会变糟的路。

So if we want to move towards this new kind of algorithms, it is almost impossible that things will not get worse before they get better. So when we start exploring these new directions, you're not going to get state-of-the-art performance from day one. But it is because it is a different paradigm. But that path leads to similar performance at a higher energy and these big labs they are so locked into a product that they like it, it is not possible for them to pursue a path where things get worse.

Host

因为他们当前的范式允许他们继续 Scaling(规模扩张),而这个新范式他们必须下注,而且必须解决一些困难的技术问题。

Because their current paradigm allows them to keep scaling and this new paradigm they have to take a bet and they have to figure out some technical things that are difficult.

Richard

我们已经思考了很多年。我们认识一些思考这些问题多年的人,当我与他们交谈时,我意识到这是可行的,但你需要长时间思考这些挑战。

We have thought about it for many years. We know people who have thought about these things for many years and when I talk to them it makes sense that it's doable but you need to think about those challenges for a long period of time.

公司愿景 Vision for the Company

Host

那么如果一切顺利,Oaks 这家公司,你在打造什么样的公司?

So if everything goes right with Oaks with the company, what kind of company are you building?

Richard

如果一切顺利,我们实现这个架构。我们能拥有真正的持续学习,能形成抽象,从而进行规划和推理,我们有点像真正的智能,然后你知道,很难想象到那时会发生什么,但我认为人类将变得无关紧要。

If everything goes right, we implement the architecture. We can have genuine continual learning and we can form abstractions so that we can do planning and reasoning and we sort of like true intelligence and then you know it's hard to imagine just exactly what will happen by then but I think humans will become irrelevant.

Host

我完全不这么认为。

I don't think that's true at all.

Richard

我们也不这么认为。我认为世界会变得令人兴奋,对人类来说甚至更加激动人心和有趣,但特别是,我认为你不得不担心大型语言模型,当这一切最终发生时,它们可能面临风险。

We don't either. I think the world becomes exciting and even more exciting and interesting for humans but in particular I think there are the you have to wonder about the large language models they might be at risk when this eventually happens.

Host

你知道,我确信它们会有一段好时光,它们已经有过一段好时光了,你知道,它们非常成功。

You know, I'm sure they'll get a good run, they've already had a good run, you know, they've been very successful.

Richard

让我明确说一下,大型语言模型是一个惊人的科学突破,是神经网络熟练运用语言的突破。这完全出乎意料。你知道,语言一直是符号方法的地盘,而它们彻底改变了现在的看法。

Let me say just for clarity that large language models are an amazing scientific breakthrough, a breakthrough in the skillful use of language by neural networks. It's totally unanticipated. You know, it was always a holdout for symbolic methods in language and they have totally changed how that's thought about now.

Host

是的,这是一个重大突破。

Yeah, it's a big breakthrough.

Richard

所以,让我沮丧的是,我们不能仅仅庆祝我们在 AI 问题的这个子集上取得的巨大进步并享受它。相反,它必须假装是全部 AI。所有智能不仅仅是流畅、熟练的语言运用。还有更多。这是一个重要的部分。你知道,它大约占智能的 20% 或四分之一。还有更多。

So, it's frustrating to me that we have to not just celebrate that we've made this great progress in this subset of the problem of AI and enjoy that. Instead, it has to pretend to be all of AI. All of intelligence is not fluid, capable use of language. There's so much more. It's an important part. You know, it's like 20% or a quarter of intelligence. There's more.

Host

是的。

Yeah.

Richard

我们还没完。

We're not done.

Host

是的。如果一切顺利,你是否想象有一个单一的思维可以做所有事情,从学习如何在树枝上荡来荡去,到制造宇宙飞船,到你知道我们今天讨论的各种事情。它是一个单一的思维,一组单一的权重可以做所有这些事情,还是……

Yeah. If everything goes right, are you imagining that there's a single mind that can do everything from learn how to swing from tree branches to make a spaceship, to you know all these various things we've talked about today. Is it a single mind and is it a single set of weights that can do all these things or is it...

Richard

这是一个单一的设计。

It's a single design.

Host

好的。

Okay.

Richard

许多不同的,会有许多思维。

Many different, there'll be many minds.

Host

好的。好的,所以这是一个对不同环境做出反应的单一设计。

Okay. Okay, so it's a single design that reacts to different environments.

Richard

那个思维的不同版本会学到不同的东西,因为它们有不同的经历。这有点回到大世界假说,即有无穷多的事情要学。所以一个系统无法学习无穷多的事情,而且我认为 Rich 已经提到过这一点,但如果你有两个这样的系统,如果你有两个世界上最大的系统,那么它们显然无法相互建模,因为它们同样复杂。所以一个单一系统永远无法达到能学习一切的地步。总是会有多个系统从它们自己的经验中学习。

And different versions of that mind would learn different things because they have different experience. This sort of goes back to the big world hypothesis that there are infinitely many things to learn. So one system cannot learn infinitely many things and like I think Rich already mentioned this but if you have two of these systems, if you have two of the largest systems in the world, then it is trivial that they cannot model each other because they are equally complex. So a single system would never be able to get to a point where it can learn everything. It would always be multiple systems that are learning from their own experience.

招聘与团队 Hiring and Team

Host

你们在招人。你们在找什么样的人?

You guys are hiring. What kind of people are you looking for?

Richard

我们在招人。最初的团队大部分我们心里已经有了。所以这些人是长期思考这些想法的人,我们将采取略有不同的方法,因为这是一个不同的范式。快速变大没有意义,因为在某种意义上,我们雇的每个人都必须来看我们所看到的,不是每个人都看到。所以我们将从小处开始,慢慢增长到也许几个或两三个,然后从那里开始。我们希望超级对齐。

We are hiring. The initial team most of it we already have in our mind. So these are people who have thought about these ideas for a long time and we are going to take a slightly different approach because this is a different paradigm. It doesn't make sense to become large very quickly because in some sense everyone we hire has to come to see what we see and not everyone sees that. So we're going to start small, slowly grow to maybe a handful or two or three, and then go from there. We want to be super aligned.

Host

希望超级对齐,这样我们可以非常高效地合作并 Scaling(规模扩张)进展。

Want to be super aligned so that we can be very productive working together and scaling the progress.

Richard

绝对。

Absolutely.

结束语 Closing Remarks

Host

非常非常酷。太棒了。我很喜欢这次对话。感谢你抽出时间分享你在做的事情。你知道,你是一位对强化学习和算法设计走向有着异常深刻思考的人,今天能和你一起探索真是莫大的荣幸。所以,谢谢你。

Very very cool. Wonderful. I loved this conversation. Thank you for taking the time to share what you're up to. You are, you know, an unusually deep thinker about where reinforcement learning and algorithmic design will go and it was a true pleasure to get to explore it together with you today. So, thank you.

Richard

非常感谢。

Thank you very much.

Host

我们的荣幸。

Our pleasure.

互动版:逐字朗读 + 针对本期提问 →