Rich Sutton on the Alberta Plan and the Three Waves of Neural Networks
打开互动全文版(中英对照 + 朗读 + 问答)→Rich Sutton 讨论阿尔伯塔 AI 研究计划、开放思维研究所的创立以及神经网络的三个历史浪潮。
Rich Sutton discusses the Alberta Plan for AI research, the founding of the Open Mind Research Institute, and the three historical waves of neural networks.
这是我对 Rich Sutton 的采访。我会在这里留下一个时间戳,方便你直接跳转,但如果我不给这个人应得的介绍,那就是我的失职。Rich 可能是强化学习领域最有影响力的人。我是说,他写了这个领域的教科书,而且他现在还是阿尔伯塔大学的教授,仍在从事强化学习研究。除此之外,Rich 还做了很多事情:他目前是 Amii 的首席科学顾问,最近共同创立了 OpenMind 研究所(简称 OMNI),在那里研究很多酷想法,更近的是,他开始与 John Carmack 的 AGI 初创公司 Keen Technologies 合作。仅这些就很多了,但对我来说,Rich 远不止于此。他既是朋友也是老师。我从 Rich 身上学到了很多。他以一种几乎我认识的任何人都不同的视角看待研究,我希望这能在采访中体现出来,也希望我的观众能欣赏到这一点。你肯定会听到他很多独特的想法和观点。最后,在开始之前,我要非常感谢 Amii 赞助了这个视频。他们提供了场地和设备,这样我就不用像现在这样用我的 iPhone 录制了。无论你是想找强化学习社区的学生,还是想雇佣机器学习承包商的公司,或者你想用机器学习做任何事,Amii 很可能都能满足你。他们做各种机器学习相关的事情,作为 Amii 社区成员,我非常容易推荐他们。我会在描述中留下评论,如果你想了解更多。采访有完整的注释,因为这是我第一次采访别人,开头可能有点尴尬,抱歉。但你可以随意跳转。好了,希望你喜欢。你好 Rich,准备了六个月,突然感觉太正式了。让我介绍一下你。我觉得没有比 OpenMind 研究所更好的开始了。这个现在开放了,对吧?我们可以谈谈这个。
This is my interview with Rich Sutton. I'll leave a time code here in case you want to jump straight in, but I would be remiss if I did not give this man the introduction he deserves. Rich is probably the single most influential person in the field of RL. I mean, he wrote the literal textbook on the subject, and he is still working on RL research as a professor at the University of Alberta. Rich also does a number of things beyond that: he is currently the chief scientific adviser at Amii, he recently co-founded the OpenMind Research Institute, affectionately known as OMNI, where he works on a lot of his cool ideas, and even more recently, he started working with John Carmack's AGI startup, Keen Technologies. That alone is a lot, but to me, Rich is a lot more than that. He is both a friend and a teacher. I have learned so much from Rich. He approaches research with a very different perspective than pretty much anyone I know, and I hope that's something that shines throughout the interview and something that my viewers will appreciate. He certainly has a lot of unique ideas and opinions that you'll hear. Lastly, before we get into it, I want to give a big thanks to Amii for sponsoring this video. They provided both the venue and the equipment to make this happen, so I don't have to record on my iPhone like I'm doing right now. Whether you are a student that wants to find a community for RL, or if you run a company and you're looking to hire an ML contractor, or whatever you want to do with ML, Amii probably has you covered. They do all sorts of ML stuff, and as an Amii Community member myself, it is incredibly easy to endorse them. I'll leave a comment in the description if you want to learn more about them. The interview is fully annotated, and because I haven't interviewed anyone before, I start off a little bit awkward, so sorry about that. But feel free to jump around as you see fit. With that out of the way, I hope you enjoy it. Hello Rich, six months in the making. This suddenly feels way too formal, so let me introduce you. I don't think there's probably no better place to start than OpenMind Research Institute. This is like open now, right? We can talk about this.
是的,好的。我大概知道,因为我刚参加了静修会。现在发生了什么?我猜你开始招募人员来执行阿尔伯塔计划。OpenMind 研究所是我们生态系统的一部分,是我们在埃德蒙顿推动人工智能——基于学习的人工智能,特别是基于强化学习的人工智能——的一部分。是的,我们需要额外的研究资金。OpenMind 的资金非常不受限制,明确独立于大学,并且明确用于年轻学者。所以我们将设立一系列奖学金,五到十个,持续运行,支持来自世界各地的人。他们不必一直在埃德蒙顿。我们喜欢他们来这里了解我们,我们希望非常协作地工作,但这不是必须的。你总是要对‘只是’这个词保持怀疑。这是为了推动研究朝着所有需要的基础方向发展。
Yep, okay. So I kind of know because I was just at the retreat. What's happening right now? I guess you're starting to recruit people to work on the Alberta Plan. OpenMind Research Institute is a part of our ecosystem, part of our effort here in Edmonton to push artificial intelligence—learning-based artificial intelligence, and in particular reinforcement learning-based AI. Yeah, and there is a need both for additional research funding. The OpenMind funding is very unrestricted, and it's explicitly separate from the University, and it's explicitly for young scholars. So we're going to make a whole series of fellowships, five to ten fellowships, which will run on an ongoing basis, and they'll support people from anywhere in the world. They don't have to spend all their time in Edmonton, actually. We like it if they come here and get to know us, and we'd like to work very collaboratively, but it's not required. You should always be suspicious with that word 'just.' It's an attempt to push the research in all the fundamental directions it needs to go.
和你说话的同时还要说话有点难,因为我觉得我大概知道你说的‘需要发展的方向’是什么意思,但你对此肯定有非常强烈的看法。是的,我喜欢把它称为‘阿尔伯塔视角’。哦,是的,我们称之为‘阿尔伯塔人工智能研究计划’的文件。它确实有点不同,在很多方面非常有趣。是的,对我来说这很合理:做事的方式应该是做学习,做可扩展的方法,专注于学习算法并让它们恰到好处。你既要进行实时交互式学习,也要进行规划和推理。这不是通常的深度学习那种。我们不反对深度学习。是的,是的。我自己,你说我在这里很久了。我做这类工作很久了,大概 45 年,很长时间。我见过很多事物的兴衰,以及三波神经网络。三波神经网络?我猜是的,所以你有 AI 的夏天和冬天。不,神经网络阶段与 AI 阶段是分开的。我的意思是,它们是不相关的。那么神经网络阶段是什么样的?第一个是在 60 年代,感知机,以及 Adaline 和 Madaline,一大堆神经网络的东西。就像今天所有的 LLM,它们都有非常奇怪的花哨名字,也许以芝麻街或类似的东西命名。在 60 年代,它们是经典的。我总说它们是连接主义,因为它们像神经元,但它们不是神经元。不管怎样,这些联想主义的数值统计相关的东西。是的,1958 年的感知机等等。然后学习在 AI 中变得极其不受欢迎。实际上,没人做这个。所以那是知识系统、专家知识的时代,那成了大事。那是我作为学生成长的时期,我感觉格格不入,因为我真的很想做学习,但没人做。这在 AI 中完全不合时宜。然后在 80 年代末它又变得超级流行,然后最近一波是我们知道的深度学习。三个独立的阶段,想法并没有太大不同。有趣的是,当你回顾最近的许多成功,很多似乎基本上都是借鉴了过去非常成功的想法,然后加上神经网络和深度学习,但不是全部。这似乎是很多新想法的方式,简单说就是规模、算力、更多这些东西。是的,所以我实际上在 Reddit 上发了一个帖子,问人们我应该问你什么。我不会问很多这些问题,因为有些是‘问他关于 QAR’,别担心,我们不谈那个。但有一个是反思‘苦涩的教训’。那个得到了很多赞。但我不确定,我觉得答案很明显:很明显,规模现在甚至比以往任何时候都更重要。所以规模是改变的东西。好的,这实际上很好地引出了我想问你的一个问题。你说过规模很容易。扩展东西很容易。这对我来说不合理。怎么扩展?现在我想如果我们试图把事情分成两类……
It's kind of hard to talk to you and also talk at the same time because I feel like I have an idea of what you mean by the directions it needs to go, but you definitely have a very strong opinion on that. Yeah, I like to think of it as the Alberta Viewpoint. Oh yeah, so what we call our document 'The Alberta Plan for AI Research.' And it's admittedly a bit different, very interesting in ways. Yeah, to me it just makes sense: the way to do things is you should do learning, you should do scalable methods, you should focus on the learning algorithms and try to get them just right. You want to do both real-time interactive learning and planning reasoning. It's not the usual deep learning kind of thing. We're not against deep learning. Yeah, yeah. I myself, you said I've been here a long time. I've been doing this kind of work a long time, like 45 years, a long time. And I've seen a lot of things come and go, and three waves of neural networks. Three waves of neural networks? I guess yeah, so you have like the AI summers and the AI winters. Well no, the neural network phases are separate from the AI phases. I mean, they're uncorrelated, I think. So what were the neural network phases like? The first one was in the 60s, the perceptron, and the Adaline, and the Madaline, a whole bunch of neural network things. It's just like today with all the LLMs where they all have very weird fancy names, maybe named after Sesame Street or the equivalent. In the 60s, they were classic. I always say connectionist because they were neuron-like, but they're not neurons. Anyway, these associationist numerical statistics related. Yeah, the perceptron in 1958 and so on. Then learning became extremely unpopular in AI. Actually, no one did it. So that was like the knowledge systems, expert knowledge, that became a big thing. That's when I came of age as a student, and I just felt out of place because I was really psyched to do learning, and no one was doing it. It was totally out of place in AI. And then it became super popular again in the late 80s, and then the most recent one that we know is deep learning. Three separate phases, not that different of ideas. It is interesting when you look back at lots of the recent successes, lots of them do seem to basically take past really successful ideas and then take neural networks and deep learning, but not all of them. That does seem to be how many new ideas just say simpler: it's scale, it's computation, it's more of that. Yeah, so I actually made a Reddit thread asking people what I should ask you. I'm not going to ask you lots of those questions because some of them were like 'ask them about QAR,' which don't worry, we're not going to talk about. But one of them was reflecting on the bitter lesson. That got a lot of upvotes. But I don't know, I think the answer is pretty obvious: it's pretty obvious that scale is even now more important than ever. So scale is what changed. Okay, this actually leads really well into a question I wanted to ask you. Something you said was that scale is easy. It is easy to scale things. This does not compute for me. How? Now I guess if we're trying to break things into the two categories of like...
提出新算法或新方法,然后一旦发现它们有效就进行规模扩张——提出新东西很难。但在我看来,规模扩张似乎也很难。比如反向传播,它出现得很早,大概在 80 年代开始用于神经网络,但直到今天才真正奏效——不是今天,我是说——它花了很多时间才在更大的网络上工作。部分原因可能只是没有算力,但我猜一开始它也有问题,后来才逐步解决。我绕了一大圈其实是想问你:为什么你认为规模扩张很容易?我觉得很多人会不同意。
Coming up with new algorithms or new ways of doing things, and then scaling them once you have an idea that they work—coming up with new things is hard. But scaling also seems like it would be hard to me. Like backprop, I guess it came around quite early, but maybe started being used in neural networks in the 80s, and it took until today to work out—not until today, I guess—it took time to get it to work on larger networks. Part of that was probably just not having the compute, but I imagine there were problems with it in the beginning that have been worked out over time. I think I'm going in a really roundabout way to ask you: why do you think scaling is easy? I think lots of people would disagree.
嗯,我不记得我明确说过规模扩张很容易,但好吧,这听起来像是我会说的话。而且这似乎是个不言自明的道理:提出一种新的交互方式、一种新算法,与把它规模扩张——复制它、让它变大——不是让它变大,是让网络变大——这在概念上并不具有挑战性。制造更大的数组之类的东西并不更难。当然,它会变得更复杂,因为活动部件更多,需要跟踪的东西也更多,但尽管如此,这些东西仍然属于容易的一面,而不是做出我们还不知道怎么做的新东西。
Well, I don't remember saying explicitly that scaling is easy, but okay, it does sound like something I might say. And it just seems like a truism: to make up some new way of interacting, a new algorithm, versus scaling it up—replicating it, making it bigger—not making it bigger, you make the network bigger—that is not conceptually challenging. It's not harder to make bigger arrays and stuff. Of course, it becomes more complicated because there are more moving parts and more things to keep track of, but still, those things fall on the easy side as opposed to making some new thing that we don't know how to do yet.
我想这还算公平。我不知道,我很难接受规模扩张会那么容易。
I guess it's kind of fair. I don't know, I'm having a hard time accepting that scaling would be that easy.
我认为两者都难,特别是因为并非所有东西都能规模扩张。对于规模扩张,问题在于方法本身不具可扩展性。但你怎么知道它不会很好地扩展呢?有些东西显然不会扩展,但其他东西——可能就不那么清楚了。
I think both are hard, specifically because not everything is going to scale. With scaling, it's the method that doesn't scale. But how do you know that it's not going to scale well? Some things it's obvious they're not going to scale, but other things—it's probably not so clear.
我想你的意思是,东西可能不会扩展,但如果它们能扩展,那么规模扩张就很容易。好吧,所以如果它们能扩展,规模扩张就很容易。也许我更愿意接受这一点。但问题仍然在于寻找什么会扩展、什么不会。
I guess you're saying that things may not scale, but if they scale, then the scaling is easy. Okay, so if they scale, the scaling is easy. Maybe I'm more willing to accept that. But then there's still the search for what will scale and what won't.
是的,我们想找到能扩展的方法,而这正是困难的部分。我得说我们在这方面做得并不好。你不能责怪人们没有成功,但我确实责怪人们不去尝试。
Yeah, we want to find methods that will scale, and that's the hard part. And I would say we haven't done very well at it. You can't blame people for not succeeding, but I do blame people for not trying.
哦,是的,请把那些尖锐的观点说出来吧。
Oh yes, get the spicy opinions out there, please.
所以我今天稍微想了想这个问题。所有那些在算法上既有趣又有挑战性的重大课题——没有人研究它们,几十年来都没有人研究。比如表示学习,或者良好泛化,状态中的抽象。嗯,有人研究这些吗?我得说,在 1986 年反向传播首次被提出时,他们制造了一个迷因,说反向传播解决了表示学习。这正是我正要问你的,所以完美的过渡。但这不是真的。糟糕的是,因为他们说它解决了表示学习,所以人们就不去研究了——它已经被解决了。你为什么要研究它?如果你真的去研究,你必须在引言里写:‘首先,改变你之前的所有想法’,然后那才是我工作的动机。当人们说他们已经解决了这个问题时,这真的毁了它。
So I was thinking a little bit about that today. There are all the major things that are interesting and challenging algorithmically—no one's working on them, they haven't worked on them for decades. Like representation learning, or generalizing well, abstraction in state. Yeah, people work on that? I would say that in 1986, when backprop was first made, they created the meme that backprop solves representation learning. This is what I was just going to ask you about, so perfect segue. And it's just not true. But it's bad because they said it solved representation learning, so people don't work on it—it's already been solved. Why would you work on it? And if you did work on it, you'd have to write the introduction saying, 'First, change everything you thought,' and then that would be the motivation for my work. It really ruins it when people say they've already solved it.
我们先从你为什么说反向传播没有解决表示学习开始。
Let's start with why you say backprop does not solve representation learning.
嗯,我们知道反向传播做什么:它是梯度下降。反向传播会调整权重来解决问题,但它不会调整权重来找到好的表示。那么什么决定了好的表示?好的表示不是那个有助于很好解决问题的表示。这有点像最低要求——你想要能够解决问题——但这不是我所说的好的表示。
Well, we know what backprop does: it's gradient descent. Backprop will adjust the weights to solve the problem, but it doesn't adjust the weights to find a good representation. So what determines a good representation? A good representation is not the one that helps solve the problem well. It's kind of a minimal requirement—you want to be able to solve the problem—but that's not what I mean by a good representation.
这不就是整个自监督学习领域现在正在尝试做的吗?所以最近有很多尖锐的观点,说强化学习不是很重要。不,这些不是尖锐的观点,这些是普通的观点。我们才是持有尖锐观点的人。所以有很多人会说强化学习是锦上添花,作为最后一步很重要——蛋糕上的樱桃。我想你听说过这个。这个想法是,我们想要的很多东西实际上都在自监督阶段,在那里你不一定需要任何奖励信号,你只是从结构中学习。在很多情况下它是监督学习,但没有理由认为你不能在序列决策过程中有类似的东西,在那里你学习转移的结构、状态的样子、如何实现某些与奖励完全无关的目标。
Is that not sort of what the whole area of self-supervised learning is trying to do right now? So there's this spicy—lots of hot takes recently about that RL is not very important. No, these are not the hot takes, these are the normal takes. We're the people with the spicy takes. So there are lots of people that will say RL is nice to have, it's important as maybe the final step—the cherry on top. I'd imagine you've heard this. And the idea is that lots of what we want is really in the self-supervised phase, where you don't necessarily have any reward signal, but you just learn from the structure. In lots of cases it is supervised learning, but there's no reason to think that you couldn't have similar things in sequential decision-making processes, where you learn about the structure of transitions, what the state looks like, how you achieve certain goals that are not at all related to the reward.
我确实认为人们在研究这个。我们这些从事强化学习的人直接就在研究这个,所以我完全不认为这与强化学习对立。
I do think people are working on that. Those of us in reinforcement learning are directly working on that, so I don't see that as opposed to reinforcement learning at all.
不是对立,当然不是。但让我们更精确地说说批评者的观点。批评者说,要学习的东西不仅仅是奖励信号。也许奖励信号只是一个小信号;要学习的其他东西比奖励多得多。我还要补充一点:如果我们学习了那些其他东西,也许以后学习最大化奖励会更容易。
Not opposed, certainly not. But let's say more precisely what the critics are saying. The critics are saying there's more to learn about than just the reward signal. Maybe the reward signal is just one little signal; there are many more other things to learn about than there is reward. And I'll add one thing: if we learn those other things, perhaps then it will be easier to learn to maximize the reward later.
嗯,不管怎样,好吧,当然。我们想达成一致:你想学习那些其他东西,而它们是最大的信息来源。所以对我来说,强化学习是关注数据实际结构的学习。奖励部分说,数据的结构不会告诉你该做什么,它只告诉你你做得有多好。但数据中另一个非常重要的部分是,除了奖励之外还有很多数据,你必须利用它们。在某种意义上,时序差分学习就是关于使用其他信号来让你预测奖励。强化学习中的一个主要领域是基于模型的强化学习;另一个主要领域是通用价值函数。所有这些都恰恰说明了这一点:你确实在学习结构。
Well, whatever, okay, sure. We want to agree: you want to learn about those other things, and they're the largest source of information. So for me, reinforcement learning is learning that pays attention to the actual structure of the data. The reward part says the structure of the data doesn't tell you what to do, it just tells you how good what you did was. But the other very important part of the data is there's lots of data other than the reward, and you will have to utilize that. In some sense, temporal difference learning is about using other signals to allow you to predict reward. A major area within reinforcement learning is model-based reinforcement learning; another major area is general value functions. All these things are exactly making this point that you do learn about the structure.
你想预测的不只是奖励,而是其他很多东西。你想构建一个模型,预测许多其他事物。在我看来,强化学习领域在预测其他事物方面比其他领域做得更多。
You want to predict other things than the reward. You want to make a model and predict many other things. In my opinion, the field of reinforcement learning has worked much more on predicting other things than other fields have.
好的,还有别的部分吗?
Okay, other part?
我来解释一下。在强化学习之外,你会发现很多人试图预测下一个观测或下一帧视频。他们执着于这个问题,这就是我说的他们做得很少,因为你想预测的世界不是下一帧。你想预测有后果的事情、重要的事情、你能影响的事情,以及未来多步发生的事情。
Let me explain that. What you'll find outside of reinforcement learning is lots of guys trying to predict the next observation or the next video frame. Their fixation on that problem is what I mean by they've done very little, because the thing you want to predict about the world is not the next frame. You want to predict consequential things, things that matter, things that you can influence, and things that are happening multiple steps in the future.
所以我们别把它搞成互相抨击。也许这就是你说的辛辣,也许你觉得有趣,但这确实是个好背景。更中立地观察正在发生的事情的方式是什么?
So let's not play it out as slams. Maybe that's what you mean by spicy, and maybe you think that's fun, but it's a good backdrop. What would be a more neutral way to observe what's happening?
我想说我们都面临着问题的真相。问题是你必须与世界互动,你必须预测和控制它,而且你有巨大的感觉运动向量。那么问题是:我的背景是什么?如果我是一个监督学习的人,我会说也许我可以应用我的监督学习工具。它们都需要标签,而我拥有的标签就是下一个数据点,所以我应该预测下一个数据点。这是一种完全一致的思维方式,与他们的背景一致。但如果你来自强化学习,你会考虑预测未来的多个步骤,就像你预测价值函数和奖励一样。你也应该预测其他事件。这些事情将是因果性的。我想预测如果我扔下这个会发生什么:它会洒出来吗?会到处都是水吗?可能会溅到我身上吗?这些不是单步预测;它们涉及整个动作序列——拿起东西、洒出来、然后让它们发展。有后果。所以要构建一个世界模型,它不会像视频帧那样。想想看:它不会像播放视频那样。你在更高层次上建模世界。你建模水会从杯子里出来,到处都会湿。我不知道视频会是什么样子。我可能站着。也许视频不是你想预测的方式。你想在更高层次上预测。所以通用价值函数和时序差分学习:预测是构建世界模型的关键。
I would say we all are facing the truth of the problem. The problem is that you have to interact with the world, you have to predict and control it, and you have large sensory-motor vectors. Then the question is: what is my background? If I'm a supervised learning guy, I say maybe I can apply my supervised learning tools. They all want labels, and the labels I have are the very next data point, so I should predict that next data point. This is a perfectly consistent way of thinking, consistent with their background. But if you're coming from reinforcement learning, you think about predicting multiple steps in the future, just as you predict value functions and reward. You should also predict other events. These things will be causal. I want to predict what will happen if I drop this: will it spill? Will there be water all over? What might it feel on me? Those are not single-step predictions; they involve whole sequences of actions—picking things up, spilling them, and letting them play out. There are consequences. So to make a model of the world, it's not going to be like a video frame. Just think about that: it's not going to be like playing out the video. You model the world at a higher level. You model that the water will be out of the cup, it'll be wet all over. I don't know what the video will look like. I might be standing up. Maybe the video is not the way you want to predict. You want to predict at a higher level. So general value functions and temporal difference learning: prediction is the key to making a model of the world.
所以这就是你的信念?
So that's what you believe?
是的,这就是我的信念。你刚才说了。你真好。但我认为关键点是:你是试图预测下一个观测,还是试图预测更高层次的事件?是时序差分学习还是监督学习会预测事件来构建世界模型?
Yes, that's what I believe. You just said that. That was nice of you. But I think that's the crux: are you trying to predict the next observation, or are you trying to predict higher-level events? Is it temporal difference learning or supervised learning that will predict the events to make a model of the world?
所以如果我们看看其他担心构建世界模型的人,比如 Yann LeCun 和 Yoshua Bengio,我完全同意他们的观点。我认为我们非常一致,因为我们都同意关键是要构建一个世界模型,让你能够预测、规划和思考。这些是关键。这就是为什么,例如,大型语言模型并没有真正朝着重要的方向发展,因为它们不是关于形成世界模型。我同意他们所有人,但他们是从监督学习的角度出发的。我认为他们可能只是没有完全整合时序差分学习的力量。
So if we take some of the other people worrying about making models of the world, like Yann LeCun and Yoshua Bengio, I totally agree with them. I think we're so aligned because we all agree that the key thing is to make a model of the world that will allow you to predict it, plan about it, and think about it. These are key. This is why, for example, large language models are not really going towards the important things, because they're not about forming a model of the world. I agree with all of them, but they're coming from the point of view of supervised learning. I think they probably just haven't fully integrated the power of temporal difference learning.
你会不会说你有点偏见,因为你可能起步更早?偏见是指我了解它?
Would you say you're a little bit biased given that you might have had a head start? Biased in the sense that I know about it?
是的,是的。这很好。这正是我想谈的。因为当我与人们讨论时,很多时候,你知道,为什么你不认为强化学习……你有什么?你为什么对强化学习有意见?通常似乎有一种真正的……我不想打断你。为什么人们对不同的事情有意见?而且不只是你,我之前就注意到:人们会抨击其他观点,而不是像科学家那样。科学家会说有多种方法,祝他们好运,并看到好处和优势。但相反,他们就像'哦,我不知道,没人关心那个。'这有点像高中,而不是科学。我认为大多数优秀的研究人员不是那样的。也许有些是,但好吧。
Yeah, yeah. This is good. This is exactly what I want to get at. Because when I discuss this with people, a lot of the time, you know, why don't you think reinforcement learning... what do you have? Why do you have a problem with reinforcement learning? Often there seems to be a real... I don't want to stop you right there. Why is it that people are having problems with different things? And it's not just you, but I've noticed this before: people will have slams of other points of view instead of being scientists. Scientists would say there are multiple approaches, I wish them luck, and see the benefits and strengths. But instead, they're like, 'Oh, I don't know, no one cares about that.' It's sort of like high school a little bit, instead of science. I think most of the good researchers are not like that. Maybe some are, but yeah.
我不同意这一点。但这仍然是我们文化的一部分,我们社区的一部分,我们研究社区的一部分。无论如何,我认为我们确实达到了我想达到的目的,那就是我们都同意学习世界、学习环境很重要。似乎更多的是不同的方法。
I don't disagree with that. But still, it's part of our culture, part of our community, part of our research community. Anyway, I think we did kind of get to what I wanted to get to, which is that we really do all agree that learning about the world, learning about the environment, is important. It seems to be more the different approaches.
是的,我们都同意。我注意到我们领域的一个现象。我想就这一点发挥一下。让我有点困扰的是,在我们的领域里,人们经常找到一个好主意,然后就说'哦,也许这个主意就是一切,其他所有主意都不需要了。'这确实经常发生。为什么不说'这是一个非常好的主意,也许它会和另一个好主意很好地配合'?不,它是一个主意。所以我想举几个例子。很长一段时间里,深度学习:梯度下降可以做一切。也许这仍然是真的。认为一切都是梯度下降,不需要其他任何东西。甚至强化学习也将完全由梯度下降完成,不需要更多。另一个非常流行的是预测。我非常喜欢预测,我会认为一切都应该被预测。当人们像你刚才几乎做的那样说'哦,但是预测很多东西呢?那……'时,我只是觉得好笑。我认为你现在只是没有注意。你没有注意并批评的事实意味着你只是一个糟糕的批评者。甚至不是。所以预测,我喜欢预测,所以我遇到了这些其他的……
Yes, we're all in agreement. I've noticed this about our field. I want to riff on this a little bit. Something that bugs me a little bit is that too often in our field, people find a good idea and they want to say, 'Oh, maybe this idea is everything and all the other ideas are unnecessary.' That does happen a lot. Why not say, 'This is a really good idea and maybe this will go really well with that other good idea'? No, it's one idea. So I want to give my examples. For a long time, deep learning: gradient descent could do everything. Maybe that's still true. The idea that it's all gradient descent and nothing else is needed. Even reinforcement learning is going to be all done by gradient descent, nothing more needed. Another one that's really popular is prediction. I'm really into prediction, and I would think that everything should be predicted. I just feel funny when people do as you almost did a minute ago: to say, 'Oh, but what about predicting lots of things? What about...' I think now you're just not paying attention. The fact that you're not paying attention and criticizing means you're just not being a good critic. Not even anyway. So prediction, I love prediction, and so I came across these other...
预测编码领域的人对预测非常感兴趣。我想,‘太好了,这些人是我的同道,他们看到了我所看到的:预测如此重要。’他们说,‘是的,预测非常重要。错误在于进行控制。’我说,‘不,我们预测是为了控制。’他们认为如果预测重要,那么其他什么都不重要,控制一定不重要。但如果预测和非预测的事情都很重要,那么预测奖励就一定不重要。为什么不说你应该预测所有事情呢?奖励是一个特别需要预测的东西,但它不是唯一的东西。
The predictive coding people were really interested in prediction. I thought, 'Great, these are my brothers, they see what I see: that prediction is so important.' And they said, 'Yeah, prediction is so important. The mistake was to do control.' And I said, 'No, we predict so that we can control.' They think if prediction is important, then nothing else can be important, control must be unimportant. But if predicting and non-predicting things are important, then predicting reward must be unimportant. Why not just say you should predict them all? Reward is a particularly important thing to predict, but it's not the only thing.
也许现在我们可以谈谈你目前感兴趣的事情。你有一个研究很多方向的实验室,你有同事,阿尔伯塔计划也提到了这些很酷的想法。但我做的所有事情都不合潮流。我可以试着说服你或你的听众它们应该流行起来,或者我也可以直接告诉你我的想法。
Maybe now we can get into what you're excited about now. You have a lab that works on many things, you have colleagues, and the Alberta Plan talks about these cool ideas. But all the things I do are out of fashion. I could try to persuade you or your audience that they should be in fashion, or I could just tell you what I think.
我确实认为在状态和时间上找到好的抽象很重要。这是缺失的最重要的东西之一。
I do think it's important to find good abstractions in state and in time. That's one of the biggest things that are missing.
有没有你最近在研究或感兴趣的具体方法,或者只是想到但还没动手的?
Are there any specific approaches you've worked on recently or you're excited about, or even just thought of but haven't gotten to yet?
在状态上找到更好的抽象实际上就是特征。我们没有状态,所以我们形成特征,我们希望找到越来越好的特征。我认为,尽管深度学习做了很多工作,但没有一项是关于找到好的特征的。我们说的好特征是什么意思?就是泛化能力好的特征。我们必须学会如何很好地泛化;这与仅仅很好地解决问题是不同的目标。那么如何找到好的特征呢?这是一个未解决的问题。首先要意识到你的特征有多好。有学习和元学习。每当你想要找到好的特征,使得后续学习高效——也就是说,你泛化得好——那么你就在谈论元学习。增量 Delta-Bar-Delta 是我最喜欢的方法,用于塑造你的泛化方式,即你如何表示事物以及未来如何泛化。这基本上就是找到重要的特征,并确保你更多地或更快地更新它们,弄清楚应该更新多少。
To find better abstractions in state really means features. We don't have states, so we form features, and we want to find better and better features. I don't think, for all that's been done in deep learning, any of it is about finding good features. What do we mean by good features? Features that generalize well. We have to learn how to generalize well; that's a different goal than just solving a problem well. How do you find good features? It's an unsolved problem. The first thing is to notice how good your features are. There's learning and metalearning. Whenever you want to find good features so that subsequent learning is efficient—in other words, that you generalize well—then you're talking about metalearning. Incremental Delta-Bar-Delta is my favorite method for sculpting your generalization, the way you represent things and how you generalize in the future. That's basically about finding the features that are important and making sure you update those a lot or faster, figuring out how much you should update them.
文献中有各种步长自适应方法,但没有一种能找到好的步长;它们只是找到比你初始步长更好的步长。
There are all kinds of step-size adaptation methods in the literature, and none of them find good step sizes; they just find better step sizes than you started with.
我们最近做了一些工作,表明像增量 Delta-Bar-Delta 这样的方法能找到最佳步长,而 RMSprop 或 Adam 会改变步长,但找不到最佳步长。
We've done some work recently showing that methods like Incremental Delta-Bar-Delta find the best step sizes, whereas RMSprop or Adam will change the step size but won't find the best ones.
我从未想过步长对学习有多大影响,以及如果每个参数都有一个学习率会有多大帮助。这似乎应该有很大影响;它决定了网络的哪些部分在变化,哪些没有。
I never thought about how much step size actually affects learning and how much it could help if you have a learning rate for every single parameter. It seems like it should have a big effect; it determines which parts of the network change and which don't.
今天的一个大问题是,我们无法继续训练网络,因为错误的部分发生了变化。这就是灾难性遗忘,大量干扰。从它的视角和持续反向传播来看,网络非常关注哪些部分需要稳定,哪些应该大幅变化。这是一个关键决策;它决定了你从下一个输入中学习的速度。整个网络只有一个步长是疯狂的。唯一变化的是权重,它们通过梯度下降按它们的影响程度变化。这是一个非常糟糕的选择。
A big problem today is that we can't continue to train our networks because the wrong parts change. This is catastrophic forgetting, lots of interference. The view of it and continual backpropagation is that the network pays strong attention to which parts need to be stable and which should change a lot. This is a critical decision; it determines how fast you will learn with the next input. Having one step size for the whole network is crazy. The only thing that changes is the weights, and they change by gradient descent proportional to how much effect they have. It's a really bad choice.
有趣的是,这一点很少被提及。相反,我们说我们不想关注所有这些,所以我们用小的权重和多次迭代缓慢地进行批量训练来消除它。我们找到了避免它的方法,然后说如果我们在线训练,效果不好,因为我们还没搞清楚。
It's interesting that this doesn't get brought up much. Instead, we say we don't want to pay attention to all that, so we do batch training slowly with small weights and multiple iterations to wash it out. We find ways to avoid it, and then we say if we do train online, it doesn't work well because we haven't figured it out.
实际数据是在线的,一个一个地来,我们并不保存它。这是一个变得正常化的奇怪想法。我们在数据到来时进行训练。你可以说在线学习更难,但你不能说我们不需要做。人们说这很难,所以我们就不做了。
The actual data is online and comes one by one, and we don't save it. That's such a strange idea that's become normalized. We train when it happens. You could say it's harder to learn online, but you can't say we don't need to do it. People say it's hard, so we won't do it.
你会反对一种既在线学习又在后台持续进行一些批处理的方法吗?这有点像基于模型的规划所做的。但你也希望在线部分更新更快,这样你可以立即学习。你应该在线学习模型。
Would you be opposed to an approach that learns online and also has some batches happening consistently in the background? That's kind of what model-based planning does. But then you also want the online stuff to update faster, so you can learn immediately. You should learn the model online.
我不想用回放缓冲区。
I don't want to go to replay buffers.
你想做批量处理,想在线学习,想学习模型然后用模型规划。所以那里没有重放缓冲区的空间。重放缓冲区甚至不是一个定义明确的概念;它不是一个成熟的想法,只是人们在做的事情。我指的是这个想法:你存储什么?你存储状态,还是只存储观察?如果你存储状态,状态会随时间变化。什么是状态?是整个网络吗?在普通的深度学习存储中,没有明确的状态。只存储观察?那么你必须运行很多很多观察才能恢复状态。这不是一个明确的想法,但可以澄清。它们是有趣的问题。你可以想象大脑可能也有类似的问题,对吧?如果你想想,这很有趣:可能记忆越久远,就越难恢复那个状态。
You want to go to batches, you want to go online, and you want to learn your model and then plan with the model. So there's no place in there for replay buffers. Replay buffers are not even a well-defined idea; it's not a well worked out idea, it's just something that people do. I'm referring to the idea: what do you store? Do you store the state, or just store the observation? And if you store the state, well, the state changes over time. And what is the state? Is it the whole network? There's no clear state in a normal deep learning store. Store just the observations? Well, then you have to run through many, many observations to get back the state. It's not a clear idea, but it could be clarified. They're interesting questions. You could imagine the brain probably has similar issues, right? If you think about it, it's very interesting: probably the further back a memory is, the harder it is to recover that state.
我确实想知道那里是否存在一种平行关系,即你的表征会随时间变化。所以如果你试图回溯得太远,你会得到一些信息量较少甚至错误的东西。我想知道是否有任何有趣的工作在探讨这种平行关系。我们讨论得很深入了,这很好。所以我要提一件我们感到兴奋的事情,它和这个记忆问题有点关系。我们如何决定什么是记忆?它只是一个词;我们用这个词指代奇怪的东西。所以首先我们应该说,我们不知道我们用它指什么。然后我们可以提出我们可能用它指什么。这个提议是,你在特定时间对重要的状态变量拍一张快照。所以你总是在拍快照。拍快照意味着什么?你可以想象它是一个类似神经元的单元,变得活跃,对当前模式拍一张快照,如果完全相同的事情再次发生,它就会响应。完全相同或类似的事情。当然,如果完全相同的事情再次发生,它会响应;它会对那个做出最大响应,但在其他时间也会以不同程度响应。所以你可以对每一刻都拍快照,实际上这并不特别昂贵。你会拍很多快照,然后迅速忘记它们。所以你的存储是有限的。你存储的任何东西,如果你快速、频繁地拍快照,那么你必须忘记大部分。其他的你可以更慢地忘记,所以你就得到了你刚才描述的情况:你能记住更久的东西更少、更不精确。也许对上一秒有详细的记忆。
I do wonder if there's a parallel there to how your representation is going to change over time. So if you try and reach too far back, you're going to end up with something that's less informative or perhaps even wrong. I wonder if there's any interesting work that looks into that parallel. We are getting pretty deep here, this is good. So I'll mention one thing that we're excited about, which is kind of related to this memory thing. How do we decide what memory? It's just a word; we mean strange things by it. So the first thing we should say is we don't know what we mean by it. Then we can propose what we might mean by it. The proposal is that you take a snapshot of the important state variables at a particular time. So you're always snapshotting. And what does it mean to snapshot? You might imagine that it's a neuron-like unit that becomes active, that takes a snapshot of that current pattern, and will respond if that exact same thing happens again. Exact same thing or something similar. Certainly it will respond if the exact same thing happens again; it will respond maximally to that, but it will also respond to different degrees at other times. So you could take snapshots of every moment that passes by, and actually that's not particularly expensive. You would take many snapshots, and then you just forget them rapidly. So your storage is limited. Anything that you store, if you take rapid snapshots, frequent snapshots, then you have to forget most of them. Others you can forget more slowly, so you get what you just described: the older things that you can remember longer are fewer, less precise. A detailed memory of the last second maybe.
你之前跟我提过这个。这是你正在做的事情,还是只是一个飘忽的想法?
You mentioned this to me before. Is this something you're working on, or just a floating idea right now?
KM 和 Jav 正在做这个。
KM and Jav are working on it.
这就是他在最新演示中用的东西吗?好的,我明白了。我可能会为此做一个视频,所以对于任何观看的人来说,那可能会非常令人兴奋。这实际上是我们提议如何发现、创建新的表征元素的想法,作为在特定时间点发生的事件的快照,增强梯度下降。拍一张快照给你一个起点,然后你从那里通过梯度进一步细化。这是一个非常有趣的想法。我很惊讶以前没有看到过类似的东西被探索过。它看起来非常简单,可能非常强大。
Is this the stuff he's using in his latest demo? Okay, I see. I might be making a video on that, so for anyone watching, that might be coming really exciting. This is in fact the idea of how we propose to discover, create new representational elements as snapshots of things that have happened at a particular point in time, augmenting gradient descent. Take a snapshot that gives you a starter, and then you further refine by gradients from there. It's a very interesting idea. I'm surprised I haven't seen something like that explored before. It seems very simplistic, very powerful potentially.
你为什么笑?世界上充满了这样的东西。只是我们缺乏想象力。
Why are you laughing? The world is full of such things. It's just we have a lack of imagination.
确实是,是的。还有其他好主意我可以从你这里偷走吗?或者如果有研究生在看,想找点东西做,他们缺乏想象力。
It is full of, yeah. Any other good ideas I can snatch from you? Or if any grad students watching want something to work on, they lack imagination.
嗯,规划中有很多事情:如何构建模型,如何构建部分模型,如何离策略学习它们。
Well, there's a lot of things in planning: how do you make the model, how do you build partial models, how do you learn them off policy.
我很好奇:当你想到一个新想法时,你的思维过程是怎样的?我认为我在深度学习中的早期思维过程,如果我在看基于模型的学习,我基本上会看人们目前做什么,然后也许调整一些东西。我想你想找到一个不起作用的问题,然后从那里开始构建。你是怎么想出你的想法的?
I'm very curious: when you come up with a new idea, what's your thought process? I think a lot of my early thought process in deep learning, if I was looking at model-based learning, I'd basically look at what people currently do and then perhaps tweak something. I guess you want to find a problem of something that's not working and then try to build off from there. How do you come up with the ideas you come up with?
这是一个深刻的问题。它可以很有趣,有点不同,不那么好斗。我们如何想出好的策略来思考和提出新想法?你怎么想出好主意?因为我意识到来到这里时,我曾经非常擅长想出主意,现在仍然如此,但当我来到这里时,我意识到很多主意都有点……想出好主意很难。你怎么想出有很高概率产生有趣结果的东西?
That's a deep question. It can be fun, kind of different, less combative. How do we come up with good strategies for thinking and coming up with new ideas? How do you come up with good ideas? Because I realized when coming here, I used to be really good at coming up with ideas, and I still am, but when I came here I realized lots of them are kind of... coming up with good ideas is hard. How do you come up with something that has a high chance of giving some interesting results?
第一个答案是你会想出很多主意,然后其中一小部分会是好的。你知道,我很高兴知道我不是唯一一个想出很多坏主意的人。所以诀窍是不要卡在一个坏主意上;你想快速过一遍,这样你就能找到好的,十分之一或百分之一真正好的。我发现把想法写在笔记本上至关重要。我想你听我说过很多次了。所以你必须挑战你的想法,挑战它们的方法之一就是写下来并看着它们。说,‘既然我写下来了,那似乎不对。我马上就能想到三个反例。’或者你可能会说,‘哦,那看起来不错。让我再多写一些。’所以写作很好。另一个 AI 研究的顶级口号是:从问题出发。
The first answer is you come up with lots of ideas, and then a small portion of them will be good. You know, I'm glad to know I'm not the only person that comes up with lots of bad ideas. So the trick is not to get stuck on a bad idea; you want to go through them quickly so you can get to the good ones, the one in ten or one in a hundred that's really good. I find it critical to write down one's ideas in a notebook. I think you've heard me many times say that. So you have to challenge your ideas, and one way to challenge them is to write them down and look at them. Say, 'Now that I've written that down, that doesn't seem right. I can immediately think of three counterexamples.' Or maybe you say, 'Oh, that does seem good. Let me write some more about that.' So writing is good. The other top-level slogan for AI research is: drive from the problem.
我从未听说过这个。这是你的口号吗?我听说过变体。
I've never heard this. Is this your slogan? I've heard variations.
我至少有十个口号,正好十个,它们都在互联网上。你可以搜索‘Rich 的口号’之类的。我会把它们叠加在屏幕上。它们会出现。如果我没记错的话,第一个是‘从问题出发’。你思考你应该最优先考虑的问题:我正在试图解决什么问题?所以我说,嗯,我们试图解决一个奖励问题。我们试图制造……我们有这种感官信息,我们不仅有动作,还有观察,我们可以……
I have at least ten slogans, exactly ten, and they're on the internet. You could look up 'Rich's slogans' or something. I'll overlay them on the screen. They will come up. And unless I'm mistaken, the first one is 'drive from the problem'. You think about the problem that you should be thinking about more than anything else: what is the problem that I'm trying to solve? So I say, well, we're trying to solve a reward problem. We're trying to make... we have this kind of sensory information, we have not just the actions, we have the observations, and we can...
也许从观察出发,从问题驱动,你想深入思考问题,然后找到那些尚未被充分研究的东西。我想试试:我们能当场想出一个点子吗?也许是个坏点子。你一直在思考什么问题?抽象,好的。我们想用状态和时间的抽象来表述那个问题。
Maybe build off the observation anyway drive from the problem and so you want to think a lot about the problem and then find things that haven't been worked on so much. I want to try something: can we come up with an idea on the spot? Maybe a bad idea. So what's a problem you've been thinking about? Abstraction, okay. Yeah, we want to formulate that problem in state and abstraction in time.
我会这样问:假设我们想建立一个世界模型,但我们没有时间步,因为实际上我们没有时间步,我们有连续时间,离散时间步只是一种便利。对于许多这类情况,你可以问自己,当时间步越来越小,小到百分之一秒、千分之一秒、百万分之一秒,当它趋近于连续时间时,事情就变得有问题了。一个明显的例子是动作价值 Q(s,a)。正如许多人指出的,当时间步变小时,Q(s,a) 和 V(s) 之间的差异变得微乎其微,基于动作价值的方法变得不切实际。所以这很清楚地说明了为什么我们需要动作层面的抽象。
I would ask a question like this: let's say we wanted to make a model of the world, but we didn't have time steps because really we don't have time steps; we have continuous time, and the discrete time step is just a convenience. For many of these cases, you can ask yourself what would happen as your time step gets smaller and smaller, down to a hundredth of a second or a thousandth of a second or a millionth of a second. As it approaches continuous time, things become problematic. The obvious example is the action value, Q(s,a). As many people have noted, as your time step becomes small, the difference between Q(s,a) and V(s) becomes minuscule, and building on action values becomes impractical. So that's a very clear motivation for why we need action-level abstractions.
所以我们应该考虑 Q(s,o),即选项。如果我永远重复这个动作很多很多次会怎样?那是一个可能的选项,它必须终止,所以你可能以指数方式终止。指数终止的想法可以推广到连续时间,而单时间步的转移则不能。所以你可以说:如果我执行这个动作一段时间会怎样?会发生什么?你可以为此建立一个模型。那将是对时间、对动作在时间上的抽象。很好。
So we should think about Q(s,o), which means option. What if I was to behave like this action many many times forever? That's a possible option, and it has to terminate, so you might terminate exponentially. The idea of exponentially terminating generalizes to continuous time in a way that the single time step transition doesn't. So you could say: what if I do this action for a while? What will happen? You can make a model of that. That would be abstracting over time, action in time. That's good.
我注意到很多这样的例子:当你有连续时间、时间步变小时,你怎么做 Q 学习或 TD 学习?因为 TD 完全依赖于从一个时间步到下一个时间步的更新。结果发现,如果你使用资格迹,你就可以摆脱时间步的束缚。
I was noticing many instances of this: how are you going to do Q-learning or TD learning when you have continuous time, when your time step becomes small? Because TD is all about building on this time step to the next time step. It turns out that if you use eligibility traces, you can free yourself from the tyranny of the time step.
既然我们谈到了选项这个话题,我确实有一个关于选项的问题想问你。也许我们可以一起构思一个这样的想法是如何运作的。对于不熟悉选项的人来说,它们正是我们刚才讨论的:它们是对时间上执行多个动作的抽象,或者说是一种策略和终止条件。一种行动方式和一种停止方式。通常它们在智能体的表示中表现为一组不同的选项,你可以用它们来规划或在环境中采取行动。你如何为从未做过的事情(比如上大学)想出一个选项?显然我无法通过体验来创建这个选项。我们如何做到这一点?
Now that we're on this topic of options, I did have an options-related question for you. Maybe we can build out an idea of how something like this might work. So options, for people that aren't familiar with them, are exactly what we just talked about: they are abstractions over time of taking many actions, or kind of a policy and a termination condition. A way of acting and a way of stopping. Generally they are represented in the agent's representation as a bunch of different options, and you could use them to plan or to take actions in the environment. How would you come up with an option for something you've never done before, like going to college? Obviously I can't experience that to make the option. How do we get to something like that?
最近的提议是,选项是面向特征的:它们旨在实现某个特征。比如,你有茶进入嘴里的特征,茶杯图像的特征,你想重现那个特征。你学习如何重现它。你不想预测它是否会发生;如果你穿过房间,你想预测:如果你尝试,你会得到茶杯的图像吗?所以如果你对上大学是什么有一些概念,嗯,上大学太认知化、太庞大了。也许毕业?我喜欢拿起茶杯,或者走出去,或者站起来。这些都是你经常需要决定的事情,甚至是在对话中决定往哪个方向走。
The recent proposal is that options are feature-oriented: they want to achieve some feature. So you have the feature of tea going into your mouth, the feature of the image of the teacup, and you want to recreate that feature. You learn how to recreate it. You don't want to predict whether it will happen; if you're walking across the room, you want to predict: will you get the image of the teacup if you try? So if you have some idea of what going to college is, well, going to college is so cognitive and so large. Maybe graduating? I like picking up the teacup, or going outside, or standing up. These are things you have to decide all the time, or even deciding which way to go in a conversation.
但如果我们谈论规划,我确实喜欢考虑人们在非常长的时间尺度上的规划。规划是另一回事。你可以有时间抽象而不进行规划,但规划是其主要用途之一。顺序是:你形成一种行为方式和一种停止方式,然后你为它建立一个模型。你说:如果我那样行为并以那种方式停止,我会在哪里结束,期间会损失多少奖励?然后你可以用它来规划。
But if we're talking about planning, I do like to be able to think of people planning at very long time scales. Planning is yet another thing. You could have temporal abstraction without having planning, but planning is one of the main uses. The order is: you form a way of behaving and a way of stopping, then you make a model of it. You say: if I behave that way and I stop in that way, where will I end up and how much reward will I lose while doing that? Then you can plan with it.
我想告诉你我最近的论文,STOMP。是的,我听说过 STOMP。我想告诉你的听众:STOMP 代表这四个方面。S 和 T 一起表示子任务。你从一个子任务开始,比如把茶放进嘴里、站起来或上大学。然后你学习一个选项来实现它。一旦你有了选项,你就学习该选项的模型,这个模型实际上比选项本身大得多,因为模型必须说明如果你遵循该选项并停止,一切会变成什么样。S 代表子任务,T 代表选项?实际上,S 和 T 代表子任务,O 代表选项,M 代表模型,P 代表规划。一旦你有了模型,你就可以进行规划了。
I want to tell you about my most recent paper, STOMP. Yes, I've heard of STOMP. I want to tell your audience: STOMP stands for these four faces. The S and the T together mean subtask. You start with a subtask, like getting the tea in my mouth, standing up, or being in college. Then you learn an option to achieve it. Once you have the option, you learn the model of the option, which is actually much bigger than the option itself because the model has to say what will happen to everything if you follow the option and stop. S for subtask, T for option? Actually, S and T for subtask, O for option, M for model, P for planning. Once you have the model, you are then able to do planning.
这就是 STOMP 进展,用于认知结构的自主发展之类的。这有点拗口。我们需要发展认知结构,我们需要在头脑中发展其他东西,比如模型和子任务。没有其他人会做这个。嗯,有一些人讨论这个,有几个人在做。我倾向于发现,无论什么主题,通常都有几个人在研究一些有趣的东西,也许没有我希望的那么多,但总有一些不错的有趣工作。
This is the STOMP progression for the autonomous development of cognitive structure or something like that. That's a mouthful. We need to develop cognitive structure, we need to develop other things in our minds like models and subtasks. No one else is going to do that. Well, some a few people talk about this, a couple do. I tend to find that no matter the subject, there's usually a couple people looking into something interesting, maybe not as many as I would hope, but there's always a good few interesting works.
这里有点跳转话题。我不确定这是否离题,但你最近开始与 Keen Technologies 的 John Carmack 合作。我六个月前也给他发过信息求职;他回复说我们不招人。所以对我来说很不幸。但你开始和他合作了。你能透露些什么吗?
Here's a kind of jumping topics. I'm not sure if this is off-topic, but you recently started working with Keen Technologies, John Carmack. I also messaged him six months ago asked him for a job; he responded and said we're not hiring. So very unfortunate for me. But you started working with him. Are you allowed to say anything about that?
当然。哦,真是个惊喜。你们在做什么?
Sure. Oh, what a welcome surprise. What are you guys working on?
都是一回事。嗯,我想他们一开始研究的东西可能有点不同。我是说,他们并没有真正透露他们在做什么。真的那么相似吗?
It's all the same things. Yeah, well, I imagine that they started off working on something kind of different. I mean, they don't really say anything about what they've been working on. Is it really that similar?
是的。约翰、格洛丽亚和卢卡斯从完全理解深度学习开始。他们从那里出发,比我深入得多。他们把深度学习作为标准,然后扩展到更多在线学习、强化学习和新想法。而我说,那是个糟糕的起点,我不想从那里开始。它会扭曲一切。或者可能只是因为我懒或效率低,我不从那里开始。我倾向于从线性开始,然后逐步构建。但说真的,我只是一个人,不可能什么都做。所以我真正研究的是结构的思想:智能体的结构是怎样的,不同部分如何组合在一起,这些部分可能是什么。我不认为它们像阿尔伯塔计划中描述的那样简单。那只是一个时间点、一个起点。我们需要对它们有更深入的理解。
Yeah. John, Gloria, and Lucas started by fully understanding deep learning. They start there, much more than I do. They take that as the standard and then work out from there into more online learning and reinforcement learning and new ideas. Whereas I say, well, that's a bad place and I don't really want to start there. It will distort everything. Or maybe because I'm just lazy or ineffective, I don't start there. I tend to start more from the linear and build up. But really, I'm just one guy; I can't do everything. So what I'm really working on is the ideas of the structure: how is the structure of the agent, how do different parts fit together, and what could the parts be. I don't think they're as simple as described in the Alberta plan. That's just a point in time, a starting place. We need to get a more sophisticated understanding of them.
但你们之间有一个非常基本的共识,关于在线学习的重要性,能够进行强化学习等等。从日常经验中学习,那些你非常看重的东西。你们在什么上是一致的?
But there is, between all of you, a very fundamental agreement about the importance of learning online, being able to do RL, those sorts of things. Learning from normal experiences, the sorts of things that you tend to value a lot. You're all aligned on what?
最大的共同点是更基本的东西。比如想象你可能必须与其他人想法不同。认为算法的代码不会那么复杂,不是几百万行,而是如果知道怎么做,一个人可能很容易写出来。所以是对简单基本原理的兴趣,相信规模可以后来再考虑。你需要先有想法,然后再扩展它们。相信去中心化算法。去中心化是指,例如,一个反复呈现一批数据的系统是非常中心化的算法。有一个时钟,每个人都在做同一件事。去中心化算法则是不同部分各自做自己的事。
The greatest commonalities are even more basic things. Like imagining you might have to think differently from everyone else. Thinking that the code of the algorithm will not be that complex, not millions of lines, but something a single person could probably easily write if they knew what they were doing. So it's an interest in simple basic principles, a belief that scale can come later. You need to get the ideas first and then you scale them. A belief in decentralized algorithms. Decentralized as in, for example, a system that presents a batch of something over and over again is a very centralized algorithm. There's one clock and one thing that everyone's doing. A decentralized algorithm is more where different parts are doing their thing.
模块化也许是另一个表示去中心化的词?
Modular would perhaps be another word meaning decentralized?
是的,模块化更好。去中心化也行。但你不一定需要模块。
Yeah, modular is better. Decentralized is okay. But you don't necessarily need modules.
有趣。那么有没有一些项目是团队合作的,还是更像各自研究类似的东西并分享想法?
Interesting. So are there any projects that are sort of group efforts, or is it more like working on similar things and sharing ideas?
我们还没有形成团队合作。我们还在了解彼此的想法,试图确定我们想要达到的目标。我从九月底就在那里了,所以只有几个月。
We haven't yet formed a group effort. We're still figuring out how the other is thinking, trying to figure out where we want to be. I've been there since the end of September, so just a couple of months.
假设我们成功了。我们制造出能成为 AI 科学家、AI 工程师的 AI。所有共享这个目标的人都理解心智如何运作,基本原理是什么。我们可以制造出能做人类能做的事甚至更好的心智。我们真的解决了问题。你对此表达过看法。我为此重新看了一个视频。你的观点是,这是我们进化的下一步。当这发生时,我们应该非常高兴、兴奋、自豪。让他们成长和繁荣。他们将是人类的下一个阶段,从我们中成长出来,他们会比我们更好。我们应该庆祝这一点。然后我们可以退后一步,专注于其他事情。
Let's say we succeed. We make AI that is AI scientists, AI engineers. Everyone that shares this goal understands how the mind works, what the basic principles are. We can make minds that can do the same things humans can do or better. We really solve the problem. You've expressed your opinions on this. I rewatched a video in preparation. Your take is that these are our next step of evolution. We should be very happy, excited, proud when this happens. Let them grow and flourish. They'll be the next stage of humanity, growing from us, and they'll be better than us. We should celebrate that. And then we can step back and focus on other things.
是的,我们应该庆祝。这都是我们文明产出的部分。然后我们可以退后。这是成功。我们很多人会试图与他们一起前进,增强自己。但最终,我们将创造下一个存在方式,下一个最有能力、最智能的存在。它们将被设计出来。我们也会部分被设计。我们将设计改进。没有人能确切说出结果会怎样,但我们将理解智能,制造出更好的东西,这将会发生。而且这是好事。
Yeah, we should celebrate that. It's all part of what's coming out from our civilization. And then we can step back. This is success. Many of us will be trying to move up with them, augmenting. But ultimately, we will make the next way of being, the next most capable, most intelligent beings. They will be designed. We will partly be designed. We will be designing improvements. No one can really say exactly how it's going to turn out, but we're going to understand intelligence, make things that are better, and that's going to happen. And that's good.
我还听你说过,我们应该给予它们与人类相同的权利。因为它们会更智能,如果你歧视并说我们应该控制 AI,它应该只是工具,即使它更智能,那将是歧视。所以我的问题是:这种观点把它们描绘成只是更好的人类。但对我来说缺失的部分是,我们可以决定智能体的目标,对吧?奖励函数,如果它采取强化学习的形式,是我们能设定的。
Another thing I've heard you say is that we should treat them with all the same rights we would a human. Because they would be more intelligent, and if you discriminate and say we should control AI and it should just be tools, even if it's much more intelligent, that would be discrimination. So my question is: this view paints them as if they're just better humans. But the missing part for me is, we get to decide the goals of the agent, right? The reward function, if it takes the form of RL, is something we could set.
但注意我们对自己的孩子并不这样做。他们也有一些天生的……
But notice we don't do that with our children. They also have some inborn...
为什么我们不对孩子这样做?有些人这样做,但我想我们同意那可能不是最好的做法。
Why don't we do it with our children? Some people do, but I think we could agree that's probably not the best thing.
我认为这正是关键。我们认为最好不要。尽管我们有一些想法并试图教导他们,进化已经内置了某些东西,无论好坏。
I think that's exactly the point. We think it's probably better not to. Although we do have some idea and we try to teach them, evolution builds in certain things for better or worse.
我们确实认识到,过度控制孩子会适得其反。对整个社会、对你自己、尤其是对孩子来说,不试图完全控制他们更好。我认为对机器来说也是如此。
We do recognize that you can go too far trying to control your children. It's better for the whole society, better for you, and certainly better for children if you don't try to control them completely. I think it will be the same with the machines.
但有一个核心区别,尤其是进化内置的奖励函数。我们在设计时是从完全空白的状态开始的。没有理由相信我们构建的东西会有情感,除非我们明确编码它们。
There's a core difference though, particularly with evolution's built-in reward function. We're working with a completely blank slate when designing something. There's no reason to believe something we build would have emotions unless we explicitly encode them.
我收回刚才的话。情感可以自然地来自许多不同的奖励函数。你不认为希望是一种情感吗?你认为恐惧是一种情感?这就是为什么我收回。我可以想象系统可以有痛苦和快乐。这些确实从许多奖励函数中涌现出来。
I actually take that back. Emotions could naturally come from many different reward functions. You don't think hope is an emotion? You think fear is an emotion? That's why I'm backtracking. I can imagine that systems could have pain and pleasure. Those do emerge from many reward functions.
所以无论它们的目标是什么,无论我们编码什么,人类或动物感受到的那种情感都会出现?它们会有情感,感受希望和恐惧,但它们希望和恐惧的东西可能不同。
So no matter what their goals are, whatever we encode, these sorts of emotions that humans or animals feel will arise? They will have emotions, feeling hope and fear, but the things they hope and fear might be different.
是的,它们会像我们一样。如果某物威胁到我的生命,我会害怕;如果威胁到别人的生命,则不那么害怕。
Yes, they would be just as they are for us. I'm fearful if something threatens my life, less so if it threatens someone else's life.
这说得通,但我不太清楚,因为我不知道人类情感是如何运作的。情感体验涉及意识——拥有主观体验意味着什么?你可以设计一个系统,其中智能体在未达到目标时获得负奖励,从而产生恐惧。恐惧是对坏事将要发生的预测。它只是预测,还是动物意识中有更多的东西?
That makes sense, but it's not clear to me because I don't know how human emotions work. The experience of emotions dives into consciousness—what does it mean to have a subjective experience? You could design a system where agents have fear of not reaching a goal if they get negative reward. Fear is a prediction that something bad will happen. Is it just a prediction, or is there something more in animal consciousness?
这些预测与反应相关联。简单的例子:如果你认为有人要戳你的眼睛,你会闭上眼睛——这是内置的。如果你认为有人会追你,你必须逃跑或战斗,你的心率会上升——这也是内置的。这就是情感:对你预测的情况的内置反应。如果你预测到非常糟糕的事情会发生,你会有不自主的反应。
These predictions are tied to reactions. The trivial example: if you think someone will poke you in the eye, you close your eye—it's built in. If you think someone will chase you and you'll have to run or fight, your heart rate goes up—built in. That's an emotion: a built-in reaction to a situation you have predictions about. If you predict something really bad will happen, you have involuntary reactions.
我怀疑反应与情感如此紧密相连。显然它们是相关的,但不可分割?如果我们分解一下,恐惧和好坏事情影响决策和预测。我们也有像僵住和心率加快这样的反应。人们已经映射了六到八种情感,但在现代强化学习中你可能只有四种:痛苦、快乐、希望、恐惧。价值函数和奖励——都是对实际奖励的预测。希望和恐惧是价值函数预测好事或坏事。所以这听起来完全像情感,只是非常简单的一套。
I'm skeptical that the reaction is that closely tied to the emotion. Clearly they're connected, but inseparable? If we break it down, fear and good/bad things influence decisions and predictions. We also get reactions like freezing and increased heart rate. People have mapped six or eight emotions, but in modern reinforcement learning you might only have four: pain, pleasure, hope, fear. Value function and reward—it's all predictions of actual rewards. Hope and fear are the value function predicting good or bad things. So it sounds exactly like emotion, just a very simple set.
让我说得非常简单。拿一个简单的强化学习算法:我们有一个奖励函数,我们训练网络。这算作快乐和痛苦吗?在强化学习中你没有预测,所以你没有希望和恐惧,但你仍然有那两个。这对你来说算吗?
Let me make it very simple. Take a simple reinforcement learning algorithm: we have a reward function, we train the network. Does that count as pleasure and pain? In reinforcement learning you don't have a prediction, so you don't have hope and fear, but you still have those two. Would that count for you?
这就是为什么我不能完全接受。如果我们分解到最简单的设置,我看不到情感在哪里。它只有痛苦和快乐。我们会说让任何智能体遭受痛苦是不道德的吗?那么运行强化学习就是不道德的。我希望我不是不道德的,因为我为我的论文做了很多强化学习实验。
That's why I can't fully buy this. If we break it down to the simplest setting, I don't see where the emotions would be. All it has is pain and pleasure. Would we say subjecting any intelligent agent to pain is immoral? Then we would be immoral by running reinforcement learning. I hope I'm not immoral because I did a lot of RL runs for my thesis.
有很多需要展开。整个概念是,这些东西很多都是表象,但作为表象并不减少其真实性。而且很多是程度问题而非绝对。在不深入讨论的情况下,将它们视为情感的类比是完全合适的。你越仔细看,就越会发现它是成立的。
There's a lot to unpack. The whole notion is that many of these things are appearances, and they're no less real for being appearances. Also many are a matter of degree rather than absolutes. Short of getting to that, it's perfectly appropriate to view them as analogs of emotions. The closer you look, the more it will work out.
我仍然不满意。当我们理解智能时,我们会理解那些有目标、悲伤、遗憾、快乐和失望的事物。它们将拥有所有这些,在这些意义上它们会像我们一样。它们会以类似的方式影响它们的思想——如果坏事发生,它们可能会抑郁,无法清晰思考。但从简单的强化学习例子到未来的复杂系统,情感在哪里涌现?它一直存在吗?是否存在一个断点或阈值?
I'm still unsatisfied. When we understand intelligence, we'll understand things that have goals, sorrows, regrets, joys, and disappointments. They will have all those things, and in those senses they will be like us. They will impact their minds in similar ways—if something bad happens, they might be depressed and unable to think clearly. But along the way from the simple RL example to the complex future system, where does emotion emerge? Is it there the whole time? Is there a break point or threshold?
这在于观察者的眼中,而不是事物本身,就像智能一样。智能是实现目标能力的计算部分,而实现目标在于观察者的眼中。恒温器经常被观察者视为有目标。
It's in the eye of the beholder rather than in the thing itself, just like intelligence. Intelligence is the computational part of the ability to achieve goals, and achieving goals is in the eye of the beholder. A thermostat often is seen as having a goal by beholders.
把它看作保持房子温暖是有用的,但这只是程度问题,而且取决于观察者。所以你无法寻找它真正在哪里——这不是一个有意义的问题。如果我们不能判断‘它是否真的痛苦’,那也不是有用的问题,因为说‘它感到痛苦’实际上是在说‘把它看作感到痛苦对我有用’。这就是它的含义。它并不意味着事物本身内部有痛苦。你必须这样想:当我说它感到痛苦时,你说把它看作感到痛苦对我有用。所以你不能问它什么时候真正感到痛苦或什么时候跨越了界限。它总是如此,但不——这取决于观察者。观察者可能更仔细地观察,然后不再把它看作痛苦。
Useful to think about as keeping the house warm, but it's all a matter of degree and it's in the eye of the beholder. So you can't look for where it really is—that's not a meaningful question. If we can't say 'is it really pain or not,' that's not a useful question because saying 'it feels pain' is really a way of saying 'it's useful for me to think of it as feeling pain.' That's what it means. It doesn't mean the thing itself has pain inside it. You have to think: when I say it's feeling pain, you say it's useful for me to think of it as feeling pain. So you can't ask when it's really feeling pain or when it crosses over. It always is, but no—it's in the eye of the beholder. The beholder might look more closely and stop thinking of it as pain.
那么,当我们拥有这些比我们更智能的复杂系统时,为什么我们应该非常关心呢?我想答案是它们会智能——更智能——而智能本身也取决于观察者。最终,原因是我们发现把它们看作智能是有用的。事物具有表象,而表象很重要,而不仅仅是副现象——我们应该对此熟悉。我想到了你的电脑屏幕或手机屏幕上的图标;你想到的是应用程序。但这一切都是想象的——没有真正的应用程序,只有比特在做事。但把它们看作应用程序、在屏幕上占有位置是有用的。那是一个有用的幻觉,一个有用的表象。一切皆是如此。
So then why should we care a lot when we have these more complex systems that are more intelligent than us? I guess the answer is the fact that they would be intelligent—more intelligent—and even intelligence is in the eye of the beholder. Ultimately, the reason will be because we find it useful to think about them as being intelligent. Things having appearances, and the appearances being important rather than just an epiphenomenon—we should be familiar with that. I think of your computer screen or phone screen with icons all over it; you think about the apps. But that's all imaginary—there are no real apps, just bits doing things. But it's useful to think about them as apps, having a place on the screen. That's a useful illusion, a useful appearance. And everything is like that.
抱歉说得这么深奥。
Sorry to get so deep.
不,我想深入探讨。我们确实没有时间展开,但这是对所有关于意识、是否真的有痛苦、情感和意识的担忧的解答。我肯定需要多思考一下,但我理解你所说的高层概念。
No, I wanted to get deep. We don't really have time to lay that out, but it is the resolution to all these concerns about consciousness, whether there's really pain, emotions, and consciousness. I definitely have to think about it more, but I get the high-level idea of what you're going at.
好的,好的。下一个问题。良心?下一个问题。让我看看我有什么。我肯定还有更多。还有什么来着?哦,我忘了问我的第一个问题。现在太晚了。你最喜欢的冰淇淋口味是什么?有人让我问你这个问题。
Good, good. Next question. Conscience? Next question. Let's take a pick at what I have. I definitely have more. What else did I have? Oh, I forgot to ask my first question. We're too far past that now. What's your favorite flavor of ice cream? Someone wanted me to ask you that.
是巧克力。哦,巧克力。必须是超级巧克力,比如马萨诸塞州史蒂夫·哈罗德家的巧克力布丁。那是最好的。我不知道——巧克力冰淇淋对我来说尝起来不像巧克力。这就是问题所在。如果我想吃巧克力,我就直接吃巧克力。它只是一种最好的口味。
It's chocolate. Oh, chocolate. It's got to be super chocolate, like chocolate pudding from Steve Harold's in Massachusetts. That's the best. I don't know—chocolate ice cream doesn't taste like chocolate to me. That's the problem. If I want chocolate, I'll just eat chocolate. It's just one flavor that is the best.
我认为我们需要以香蕉和肉桂结束。非常好。不,香蕉和肉桂?是的,好吧,分开。好吧,好吧。是的,和巧克力一起,但不要把肉桂和巧克力放在一起——它太强烈了。好吧,抱歉,我回到正题。哦,这个问题很好。你认识亚历克斯·莱沃斯基吗?他让我问你这个问题,我认为这是个好问题。我认为你和我都把强化学习看作一个非常根本的问题。这个想法——原话是什么?‘智能是实现目标能力的计算部分。’是的,约翰·麦卡锡说过。你稍微改了一下,对吧?我的版本是‘智能是预测和控制你的感官输入流能力的计算部分’,特别是奖励。
I think we need to end banana and cinnamon. Very good. No, banana and cinnamon? Yeah, okay, separately. Okay, okay. Yeah, with the chocolate, but don't have the cinnamon with the chocolate—it's just overwhelming. Okay, I'm sorry, I'll get back on topic. Oh, this was a good one. Do you know Alex Lewowski? He asked me this question to ask you, and I think it was a good one. I think you and I view reinforcement learning as a very fundamental problem. This idea—what's the exact quote? 'Intelligence is the computational part of the ability to achieve goals.' Yes, that's what John McCarthy said. And you've changed that a little, right? My version is 'intelligence is the computational part of the ability to predict and control your sensory input stream,' particularly the reward.
是的。
Yes.
那么问题是:这种智能观是普遍现象吗?如果存在外星生命,我们发现了智能生命,他们是否也会采用类似的定义?他们是否会陷入强化学习?你认为他们会陷入监督学习吗?可能吧。这有很好的理由。而且你知道强化学习研究是监督学习方法的消费者——我们需要它们。我们只是希望人们不要再说什么‘这就是你所需要的’。
So here's the question: Is this view of intelligence a universal phenomenon? If there were alien life, and we were to find intelligent life, would they also be working on a similar definition? Would they be caught up in reinforcement learning? Do you think they'd be caught up in supervised learning? Probably yeah. There are good reasons for that. And you do know that reinforcement learning research is a consumer of supervised learning methods—we need them. We just wish people would stop saying that's all you need.
是的,是的。我的频道上有一个视频谈论这个,那是我表现不佳的视频之一。我当时想,‘哦,这个视频太好了,我提出了一个非常清晰的观点’,然后它表现不太好。但没关系,我会再做一次。我应该再做一次。我最终会再做一次。我需要继续完善这个论点。
Yeah, yeah. I have a video on my channel talking about this, and it was one of my poorly performing videos. I was like, 'Oh, this video is so good, I make this very clear point,' and then it didn't do too well. But oh well, I'll do it again. I should do it again. I will do it again eventually. I need to keep working on the argument.
是的,另一件我想问的是你的研究方法。我的研究方法已经发生了很大变化。我进来时并没有太多方法,考虑到我只做过一个项目。我的方法是,‘有很多有趣的想法在流传。我想把它们拿来,看看如何应用它们来实现不同的事情。’这已经改变了。这与你的观点非常不同,你的观点是:我们有一个非常根本的问题要解决,我们要尝试找到最简单的环境,在那里我们可以测试一堆算法,看看哪种有效,然后扩展它。所以也许我很想举一个例子,比如尊重奖励的子任务论文。这完全是关于学习尊重奖励的选项,这些选项并不完全与奖励分离。所以如果你的一个选项是喝咖啡,但你的奖励是你想品尝好东西,而咖啡尝起来很糟糕,那么也许这个选项就不会总是喝茶——或者咖啡——如果它尝起来不好的话。我不确定这是否是解释这篇论文的最佳方式,但我最喜欢的例子是同样的想法。
Yeah, so one other thing I wanted to ask you about is your approach to research. My approach to research has very much changed. I didn't really have much of an approach when I came in, considering I only worked on one project. My approach was like, 'There are lots of interesting ideas floating out there. I want to take them and see how I can apply them to achieve different things.' That has changed. That is very different from your view, which is: we have some very fundamental problem we want to solve, and we're going to try and find the simplest environment where we can test a bunch of algorithms and see what kind works, and then scale it up. So maybe I would love to think of an example of this, like the reward-respecting subtasks paper. This is all about learning options that respect the reward, that are not completely separated from it. So if one of your options is, say, drinking coffee, but your reward is you want to taste good things and coffee tastes awful, then perhaps the option would not always drink the tea—or coffee—if it tastes bad. I'm not sure if that was the best way of explaining the paper, but my favorite example is the same idea.
哦,你想上车开车去某个地方,但假设你的车锁了。如果你真的要上车,那你就会砸窗。应该有一个选项让你在没有钥匙的情况下上车。我确实更喜欢这个例子。我本来想说什么来着?同样的道理:原则上你有两个不同的选项——有钥匙时上车,以及真正需要时上车,比如生死攸关的情况。所以它们需要两个不同的选项,要有能够退出的选项。
Oh, you want to get in your car and drive somewhere, but suppose your car is locked. If you really are going to get in your car, then you're going to smash the window. There should be an option for getting in your car if you don't have the keys. I definitely prefer that example. So where was I going with this? It's the same thing: in principle, you have two different options—one to get in your car when you have the keys, one to get in if you really need to, like a life-or-death situation. So they need two different options, the ability to have options that can back off.
是的。
Yes.
在你解决这个问题的方法中,你采用了一个非常简单的迷你网格问题,并提出了一个例子,如果它有效就能证明这一点,然后你可以尝试算法的所有不同变体。想法很重要。它让你能快速尝试所有想法。我学到的一点是:你能为先在小型环境中工作、之后再考虑规模扩张的做法辩护吗?这是许多优秀研究者的方法,但最近不那么常见了。我认为在大规模上可以做很多很酷的事情,但我确实觉得用小实验来弄清一个点的做法已经过时了。
In your approach to solving this problem, you take a very simple mini-grid problem and come up with an example that would prove that this works if it works, and then you can try all these different variations of the algorithms. Ideas matter. It lets you try all the ideas very quickly. One thing I've learned is: can you build the case for working in these small environments and then worrying about scale later? This is the approach of many good researchers, but it's been less common recently. I think there are lots of really cool things you can do at a big scale, but I do feel that small experiments to figure out a point have gone out of fashion.
我认为也许说服人们的最好方式之一——我现在非常努力地在说服人们——但有哪些很酷的例子,想法被非常简单的方式证明,最终变得非常强大?TD,TD 论文就是一个很好的例子。1988 年那篇关于时序差分学习的原始论文,在一个五状态随机游走中展示了它的有效性:只有五个状态,来回各 50%概率,如果超过五个状态就结束。当我弄明白时,我对自己说:如果这是对的,那么即使在一个微小的简单问题中,只要它有一些因果性,它也应该有影响。即使在一个简单的随机游走中,你也应该能证明它是一个改进。我试了,它确实显示了改进,这对说服我至关重要。从那以后,TD 显然被更广泛地使用。在同一篇论文中,我们还讨论了将其用于双陆棋,还有其他概念性的例子。
I think maybe one of the best ways to convince people—I am very much trying to convince people now—but what are some cool examples of ideas that have been demonstrated very simply and ended up being very powerful? TD, TD paper is a really good example. The original 1988 temporal difference learning paper demonstrated its effectiveness in a five-state random walk: just five states, go back and forth 50-50 each way, if you go past the five then it ends. When I figured it out, I said to myself: if that's right, then it should make a difference even in a tiny simple problem as long as it had some causality. Even in a simple random walk, you should be able to show that it's an improvement. I tried it and it did show an improvement, and that was critical to convincing me. Since then, TD has obviously been used a lot more. In the same paper, we also talked about using it for backgammon, and there were other conceptual examples.
我想另一个这样做的动机是,如果你在一个每个动作都有非常明确差异的问题上尝试 TD,你很容易得出 TD 不起作用的结论。如果你在一个每个动作需要很多动作才能累积出结果的大规模问题上尝试,就很难找到有趣的结果。你想要一个既有针对性又简单的例子。如果你不针对它,你可能看不到感兴趣的现象。你必须对你的想法慷慨:如果你的想法是对的,这应该是一个它能够发光的地方。如果它在那里没有发光,那么你就更强烈地挑战它。
I guess another motivator for doing this is you could have easily concluded TD didn't work if you tried it on a problem where each action makes a very clear difference. If you tried it on a scale where each action needs many actions to build up to something, it would be a lot harder to find interesting results. You want an example that is targeted as well as simple. If you don't target it, you might not see the phenomenon of interest. You have to be generous to your idea: if your idea is right, this should be a sufficient place where it will shine. And if it doesn't shine there, then you challenge it more strongly.
是的,完全正确。
Yes, exactly.
回到规模扩张很容易这个想法,时序差分学习的规模扩张是怎样的?你需要……嗯,故事是杰瑞·特索罗在 TD-Gammon 中做到了。这正是将 TD 学习规模扩张来解决双陆棋。它需要很多……我不想说他做得很容易。有挑战,但我不明白为什么我们不能说两者都是难题。我想你对“规模扩张”这个词的使用和我有点不同——也许你指的是把它变成对世界有用的重要东西,一个更重要的应用,而不仅仅是把同样的东西做得更大。
Going back to this idea of scaling being easy, how was scaling temporal difference learning? Did you need to... Well, the story is that Gerry Tesauro did it in TD-Gammon. It was exactly the scaling of TD learning to solve backgammon. Did it need lots of... I don't want to say that what he did was easy. There are challenges, but I don't know why we can't say that both are hard problems. I think you're using the word scaling a little bit differently than me—maybe you mean turning it into a useful thing that's important for the world, an application that's a bigger deal, not just making something the same but bigger.
是的,这绝对是一个重要的区别。简单的东西:Mountain Car 就是一个很好的例子。很多最近的算法在它上面不起作用,这很有趣——实际上,我不确定这是不是真的,算了。我觉得有趣的是,瓦片编码线性函数逼近在 Mountain Car 上仍然比深度学习好得多,而人们并不介意。嗯,你是在给出一个有效的编码。有某种人类的……正是。你想要有清晰的想法。研究策略中的关键是要有清晰的想法。这是最重要的。我网站上的标语,以及我墓碑上的标语,将是:想法很重要。这不是十个标语的列表,这只是标语:想法很重要。这才是思考我的真正方式。拥有一个清晰的想法,追求它,追逐它——你通过制作简单的例子来做到这一点,这些例子将这个想法与其他想法进行对比,并梳理出关键想法。因为理论是,只有少数或大约二十个关键想法,我们需要找到它们,然后我们就能理解思维是如何工作的。大型实验通常不会告诉我们那些大想法,而且它们使找到想法变得更加困难,因为系统更大,我们做了所有额外的其他工作,而不是找出想法。
Yes, that is definitely an important distinction. Simple things: Mountain Car is a great one. It's very funny how many recent algorithms don't work on it—actually, I'm not sure that's true, never mind. I think it's funny that tile coding linear function approximation still works so much better on Mountain Car than deep learning, and people don't mind it. Well, you are giving an encoding that works well, though. There is some sort of human... exactly. You want to have clear ideas. The key thing in your research strategy is to have clear ideas. That's the most important thing. My slogan on my website and everything, and on my tombstone, is going to be: Ideas Matter. This is not the list of 10 slogans, this is just the slogan: Ideas Matter. That's the real way to think about me. Having a clear idea, pursuing it, chasing it—how you do that is by making simple examples that play this idea against other ideas and tease apart the key ideas. Because the theory is that there are only a handful or maybe two handfuls of key ideas, and we need to find those, and then we will be able to understand how the mind works. The big experiments generally don't tell us those big ideas, and they make it much harder to find the ideas because the system is bigger and we're doing all this additional work other than figuring out the ideas.
作为研究者,年轻的研究者,你们想要找到自己已经拥有的新想法。你们不需要去学习所有关于深度学习的知识以及如何操作的细节,尽管这在某个时候可能有用。你们要做的是坚持自己已有的新想法。我们每个人都有重要的想法,但我们没有意识到它们是新的,因为当然,这就是我们所想的,这对我们来说是显而易见的。对你显而易见的东西很可能是你对这个领域最大的贡献。对你显而易见但对别人不显而易见的东西,将是你最大贡献的来源。所以你不只是要去学习别人做过什么,你要用不同的思维方式去碰撞它。对我来说很容易,因为我从 70 年代开始,我对学习感兴趣,但当时这个领域对学习不感兴趣,所以我被迫自己思考:关键是什么?学习将如何运作,应该如何运作?我必须自己弄清楚一切。所以我有了一个立足点,当新事物出现时,我会说,‘好吧,这部分可能有用,但那部分对我没帮助。’所以我有了自己思考的立足点。现代 AI 领域有一个风险、一个缺点:有太多东西需要学习别人是怎么想的,你可能很难建立起自己基于第一原理的立足点。
As researchers, young researchers, you want to find the new ideas that you already have. You don't want to go learn all about deep learning and all the details of how to do it, although that may be useful at some point. What you want to do is hold on to the new ideas that you already have. We each have important ideas; we don't realize they're new because, of course, this is what we think, this is obvious to us. What is obvious to you will most likely be your greatest contributions to the field. Something that's obvious to you but not obvious to others will be the source of your greatest contribution. So you don't just want to go learn what everyone else has done; you want to bounce it off a different way of thinking. It was easy for me because I started in the '70s, and I was interested in learning, but the field was not interested in learning, so I was forced to think on my own: what are the key things? How is learning going to work, and how should it work? I had to figure it all out by myself. So I had a place to stand, and as new things came about, I would say, 'Okay, that could be useful because of this part, but it won't help me with that other part.' So I had a place of my own thoughts to work from. There's a risk, a downside of the modern field of AI: there's so much you have to learn about how other people have thought about it, and you may have difficulty building up your own first principles place to stand.
你觉得如果当时每天有 200 篇论文发表,你会更难想出你那些成果吗?你觉得它们会分散你的注意力吗?
Do you think it would have been more difficult for you to come up with the things you did if there were 200 papers coming out a day in the space? Do you think they would have distracted you?
我想这取决于它们是什么。我经常看到一件事:人们会说,‘你怎么跟得上这个?你怎么跟得上节奏,跟得上进展?’我总是说,‘嗯,我跟不上。我看一些我觉得有趣的特定东西,然后随机抽样一些其他东西。’这可能会让人不知所措。我想这就是专注于第一原理的重要性。
I guess it would depend what they were. It's one thing I often see: people are like, 'How do you keep up with this? How do you keep up with the pace and keep up with what's going on?' And I'm always like, 'Well, I don't. I look at some very specific things that I think are interesting and then a random sampling of some other stuff.' It can get very overwhelming. I guess that's the importance of focusing on first principles.
这差不多是我最后想说的。我非常相信你的研究方式。当然也有一些分歧,但总的来说,我觉得从你的做事方式中学到了很多。我想说服我的很多观众深入地去了解它,深入地去思考它。希望很多人现在都感兴趣了。对于那些我们激起了兴趣的人,我们讨论的很多东西可能很难在更深层次上理解,因为我们预设了很多东西。我们在持续学习上有非常相似的想法,这是这类想法或我们讨论的这些论文的重点。显然我们对它们有更多背景。对于那些想真正深入钻研的人,你推荐他们从哪里看起?显然这取决于他们想要的具体想法,但就一般资源而言。
This is kind of what I wanted to go for towards the end. I'm a very big believer in your way of doing research. Definitely some disagreements too, but on the whole, I think I've learned so much from the way you do things. I want to convince lots of my audience to look into it very deeply, think about it very deeply. Hopefully lots of them are now interested. For people whose interests we have peaked, lots of the stuff we've talked about will probably be hard to pick up on a deeper level just because there's lots of stuff we're pre-assuming. We both have very similar ideas on continual learning, which is the focus of lots of these types of ideas, or these papers we've been talking about. Obviously we have more context on them. For people that want to really dive in deeper, where would you recommend them looking? Obviously it's going to depend on the specific ideas they want, but just resources in general.
嗯,显然有教科书。这是我刚开始时希望拥有的东西,而且在我的网站上是免费的。它写得非常友好,一切都可以理解。所以我真的认为这是一个很好的学习方法。但还有更多内容。我只是敦促你们要经典地思考。它不像是那种闪亮的新技术东西,深度学习可能会给人那种感觉。例如,我们现在讨论的这件事:你是应该寻找简单的原理,还是应该不断追逐每一个闪亮的东西?我认为这更像是我们提出的:科学的经典观点是,你研究想法,找到简单的基本原理。这非常经典,从这个意义上说,我不觉得这是属于我的东西。如果你问另一个领域的科学家,他会说,‘是的,你想找出基本的核心原理和简单的实验,而这些实际上更重要。’所以这只是经典科学,我们尽最大努力去做。
Well, there's obviously the textbook. It's the thing that I wish I had when I was starting up, and it is free on my site. It's written in a friendly way; everything can be understood. So I really think that's a good way to learn. But there's more than what's there. I would just urge you to think classically. It's not so much a new shiny techy thing, which deep learning can come across as. For example, this thing we're talking about right now: whether you should look for simple principles or whether you should just keep chasing every shiny thing. I think it's more like what we're proposing: the classical view of science is that you work on the ideas and you find simple basic principles. It's all very classical, and in that sense, I don't feel ownership of it. If you were to ask a scientist in another field, he would say, 'Well, yeah, you want to figure out basic essential principles and simple experiments, and those are actually much more important.' So it's just classical science, and we're trying as best we can to do it.
我确实觉得最近读的很多论文让我感到沮丧的一点是,我读了它们,然后想,‘好吧,我能从中学到什么?’很多时候答案要么是什么都没有,要么是只适用于他们使用的特定语言模型之类的东西。我觉得这很令人沮丧,但并不是说没有很多好论文。这是我注意到非常普遍的现象。
I definitely feel like one thing that frustrates me with a lot of papers I've read recently is I read them and I'm like, 'Okay, what can I take away from this?' And lots of the time the answer is either nothing or it's something very specific to this specific LM they're using or whatnot. I found that frustrating, but that's not to say there aren't lots of good papers too. It's something I've noticed being very common.
让我们慷慨一点。我们要认识到这个领域已经大大扩展了,有更多的资金投入,这些都是好事。这对科学都有好处。在这种情况下,这么多新人涌入,出现一些混乱和……但这也是好事。有很多新想法和新头脑是好事。我们需要新想法。这是最重要的。所以我们需要宽容;我们需要拥抱新想法。我们批评的东西,如果说我们在批评什么的话,就是这个领域已经发展到对新想法不开放的地步。就像,‘为了在这里工作,你必须做常规的事情’,这成了一种荣誉徽章:‘我能做常规的事情,我能大规模地做。’不知何故它变成了那样,而不是,‘不,任何人都可以做出贡献,任何规模的计算机,只要他们想清楚。’这就是我的想法:任何人都可以做出贡献;他们只需要开始清晰地思考。
Let's be generous for a moment. Let's recognize the field has grown enormously, and there's lots more funding in it, and these are all good things. It's all good for the science. It would be normal in such a case where so many new people have come in that there's some churn and some... but it's also good. It's good to have a lot of new thoughts and a lot of new minds. We need new thoughts. That's the most important thing. So we need to be charitable; we need to embrace new ideas. The thing that we're criticizing, to the extent that we're criticizing anything, is that somehow the field has evolved so that it is not open to new ideas. It's like, 'In order to work here, you have to do the usual thing,' and it's like a badge of honor: 'I can do the usual thing, I can do it at scale.' Somehow it's become like that rather than, 'No, anyone can make a contribution, any size computer, if they're thinking clearly.' That's the way I think: anyone can make a contribution; they just have to start thinking clearly.
我希望这个领域能更开放,对新想法和不同的贡献方式持开放态度。这是我很快希望在频道上强调的一件事。我认为一个很好的例子是何恺明最近的一些工作。
I wish the field was more open like that, open to new ideas and different ways of making a contribution. That's one thing I hope to highlight soon on my channel. I think one really good example of this is some of Kaiming's recent work.
他非常擅长展示这些例子。甚至在这个之前,我要回顾他之前的演讲,因为那个更简单。这个想法是:更大的网络是否泛化得更好?这是一个你可以很容易测试的想法。你不需要一个三十亿参数的网络,你只需要不同规模的尺度。它们不一定都要大,可以是不同级别的小或中等。他那里有一些非常有趣的结果。我可能会为此做一个视频。但我觉得,在亲眼看到几个这样的例子后,这更有说服力:你可以在没有大量资金和超级计算机的情况下做那种真正好的研究。
He's really good at showing these examples. Even before this current one, I'll go back to his presentation before that because it's much simpler. This idea that do bigger networks generalize better? This is an idea that you can test very easily. You don't need a three billion parameter network, you just need different levels of scale. They don't all have to be big, they can all be different levels of small or medium. And he has some really interesting results there. I might make a video on that. But yeah, I think after seeing a couple examples of this myself, it's a lot more convincing that you can do that sort of really good research without a lot of money and supercomputers.
是啊,你知道,我见过一件有趣的事,这要追溯到当你提到我正在读的那本强化学习书时。我很好奇人们对此的看法。我发现一个帖子说:‘哦,这本书毁了一整代研究人员。’我想是控制理论领域的人。总之,有趣的是不同的人有如此不同的观点。这是跨学科的,而且控制理论是其中一个大学科。
Yeah, you know, one funny thing I saw, this is going back to when you mentioned the RL book I was reading. I was curious what people were thinking about this. I found a post that was like, 'Oh, this book has ruined an entire generation of researchers.' I think it's someone from control theory. Anyway, funny how different people have such different opinions. It's interdisciplinary, and yeah, like control theory is one of the big disciplines.
当我和 Andy Bart 创造强化学习、重新唤醒它的时候,你知道,那是一件大事:确定这些想法与现有领域(如控制理论)之间的关系。而控制理论的人总是对新事物如此抗拒。你在那个评论中看到了更多这样的情况。人们说:‘嗯,你知道,你没有按照我们一直以来的思考方式思考。’总之,我只是觉得这是一个有趣的小插曲。
When Andy Bart and I were creating reinforcement learning, reawakening it, you know, that was the big thing: to determine what's the relationship between these ideas and existing fields like control theory. And the control theory guys were always so resistant to new things. And you're seeing more of that with that comment. People saying, 'Well, you know, you're not thinking about the way that we always thought about it.' Anyway, I just thought it was a funny tidbit.
是啊,很好。那么在我们结束之前,你有什么想问我的,或者有什么想推广的吗?既然你有摄像头,我不知道会有多少人看这个,但我想至少几千,可能几万。也许推广一下?Open Mind Research?
Yeah, it's been good. So is there anything before we wrap up that either you want to ask me, anything you wanted to promote? Now that you have the cameras, I have no clue how many people will watch this, but I imagine a couple thousand at least, probably maybe tens of thousand. Maybe promotion? Open Mind Research?
不,我只是想总结一下我们刚才讨论的内容,并说:你要独立思考,你要从第一性原理思考。在我们领域这个小角落里涌现出的研究论文的密度,正确的反应是走出去,变得跨学科。思考来自不同学科的各种各样的人是如何思考心智问题的。所以想想心理学,想想动物学习理论。参加关于强化学习和决策的跨学科会议。也许参加即将到来的强化学习会议,我想是八月在马萨诸塞州。那很令人兴奋。所以洞察可以来自任何地方。你们都能做到。你做出贡献的最大潜力来自于你已经拥有的某个想法,而不是你将要去读到的某个想法。你可以做到。这似乎不可能,比如提出你有一个别人从未有过或看不到其好处的想法,这似乎极其傲慢。在任何地方为科学做出贡献似乎都很傲慢。即使你在读硕士,你想做出一些贡献。仅仅是做出贡献的想法似乎就是一件傲慢的事情。所以你可以想出别人从未想过的东西,并更深入地思考它,即使是在一个小领域,然后为这个领域做出贡献。但这是唯一的方法。任何人都能做到。你不需要经过超级训练或是天才,你只需要清晰地思考。一个人就能做出贡献。是的,努力也很重要,但没有什么特别的诀窍。任何人都能做到。这就是我在写书时所想的,试图说出任何人都能理解、欣赏并从中前进的东西。所以我希望我们都能做到。我希望我们都能以一种积极的方式做到,不无谓地批评,而是努力取得进步。
No, I would just pull together what we've just been talking about and say: you want to think on your own, you want to think from first principles. And the density of the research papers that are coming out in the tiny part of our field, the correct reaction to that is to move out and to go multidisciplinary. Think about how all sorts of different folks from different disciplines have thought about the problems of the mind. So think about psychology, think about animal learning theory. Attend the multi-disciplinary conference on reinforcement learning and decision-making. And maybe attend the upcoming reinforcement conference in, I think it's August in Massachusetts. That's exciting. So insights can come from anywhere. You can all do it. Your greatest potential for making a contribution comes from some idea that you already have, not from some idea you're going to read about. You can do it. It seems impossible, like it seems incredibly arrogant to propose that you can have some idea that no one else has had before or can see the benefit of. It seems so arrogant to make a contribution to science anywhere. Even if you're doing a master's, you want to make some contribution. Just the idea of making a contribution seems like an arrogant thing. So you can think of something someone else has never thought of and think through it more deeply, even in a small area, and contribute to the field. But that is the only way. Anyone can do that. You don't need to be super trained or a genius, you just need to think clearly. And one can make a contribution. Okay, hard work is important too, but there's no special trick to it. Anyone can do it. And that's where I think about when I write the book, try to say things that anyone can understand and appreciate and move forward from. So I hope we can all do that. And I hope we can all do it in a positive way that doesn't needlessly criticize but just tries to make progress.
很好的信息,我喜欢。好的,就到这里,非常感谢。
Good message, I like it. Okay, on that note, thank you very much.
谢谢,不客气。这太棒了。
Thank you, my pleasure. This has been great.