人工智能与机器人:从实验室到现实世界——对话 Pieter Abbeel

AI and Robotics: From Lab to Real World with Pieter Abbeel

彼得·阿贝尔 Pieter Abbeel · TWIML AI 播客 · 2021-04-19 · 约 66 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Pieter Abbeel 分享他在学术界和工业界的双重角色,专注于将 AI 机器人带入现实世界,并结合无监督学习与强化学习。

Pieter Abbeel discusses his dual roles in academia and industry, focusing on bringing AI robotics into the real world and combining unsupervised learning with reinforcement learning.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 18)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Background

Host

好的,各位,我非常激动能请到 Pieter Abbeel。Pieter 是加州大学伯克利分校的教授,也是伯克利人工智能研究实验室(BAIR)的联合主任,同时还是 Covariant 的联合创始人、总裁兼首席科学家。Pieter,欢迎回到 TWIML AI 播客。

All right, everyone, I am super excited to be here with Pieter Abbeel. Pieter is a professor at UC Berkeley, where he's co-director of the Berkeley Artificial Intelligence Research Lab, or BAIR, as well as co-founder, president, and chief scientist at Covariant. Pieter, welcome back to the TWIML AI podcast.

Pieter

Sam,很高兴再次见到你。谢谢你邀请我。

Sam, really good to see you again. Thanks for having me.

Host

真不敢相信距离你上次上节目已经快四年了。你的采访是第 28 期,现在我们快做到 500 期了。真是一段旅程。

It is really hard to believe that it's been just under four years since you were on this show. Your interview was number 28, and we're approaching 500 now. What a journey.

Pieter

真是一段旅程。

What a journey.

Host

你做了很多很酷的事情,包括创办了自己的播客。我们稍后会聊到。但既然过了这么久,你又做了这么多事,不如花几分钟跟我们的听众分享一下你的背景和主要研究兴趣吧?

You've been up to a lot of cool stuff, including starting your own podcast. We'll talk a little bit about that. But since it's been so long and you've been up to so much, why don't you take a few minutes and share with our audience a little bit of background and what your primary research interests are?

Pieter

当然。是的,四年前,那是很久以前了。从那以后发生了很多事。我最近主要在想的事情有两方面。一方面,戴着 Covariant 的帽子,作为联合创始人,我们想把 AI 机器人带进现实世界——把它从实验室、从模拟中带出来,让机器人在现实世界里为我们做事。那里有很多进展,Sam,我很期待和你多聊聊这个。另一方面,我在伯克利的另一顶帽子,我们在做 AI 和机器人领域的学术研究。我想说,我们最近最兴奋的一些工作是处于无监督学习和强化学习交汇点的研究。当然,无监督学习是在没有任何数据标注的情况下进行学习,这很好,因为这样效率可以高很多——不需要标注。但强化学习,这种试错学习,可以让机器人从自己的经验、自己的练习中获取技能。所以这两者的结合正是我们目前在伯克利花很多时间的地方。

Sure. Yeah, so four years ago, a long time. A lot has happened since. The things I'm mostly thinking about these days are kind of two-fold. One is with my Covariant hat on, as co-founder at Covariant, we're trying to bring AI robotics into the real world—take it out of the lab, out of simulation, and make robots do things for us in the real world. So a lot happening there, and I'm looking forward to talking more about that with you, Sam. And then my other hat at Berkeley, we're doing academic research in AI and robotics. I would say some of the things we're most excited about these days are work at the junction of unsupervised learning and reinforcement learning. Of course, unsupervised learning is where the learning happens without any annotation of the data, which is nice because it can be a lot more efficient that way—no annotation needed. But then reinforcement learning, the trial-and-error learning, can allow a robot to acquire skills from its own experience, its own practice. So the combination of those two is where we're spending a lot of time right now at Berkeley.

Host

太棒了。我们稍后会深入探讨这两个领域。但在此之前,对于那些四年前没听过你播客或不知道你故事的人,你是怎么进入机器人领域的?

Awesome. And we'll dig into both of those areas in detail. But before we do, for those who didn't catch your podcast four years ago or don't know your story, how did you get into robotics?

Pieter

当然。是的,对我来说,从某种程度上说,它并非始于机器人本身,而是始于人工智能——机器人背后的那种大脑。对我来说,那是在我本科学习各种东西的时候。当时似乎一切都有趣,但似乎要做好,你得选一个专攻的方向。所以问题是,好吧,什么才是我最值得花时间的有趣事情?对我来说,变得很清楚的是,我们人类能够思考这件事太迷人了。这可以说是让我们区别于其他动物的原因,我们比大多数动物思考得更深、更努力。所以这对我来说太迷人了。当然,你首先可能会想,好吧,那就研究神经科学,因为那是研究大脑如何工作的。但那个学科,我的意思是,很迷人,但似乎很难取得进展——解剖大脑并理解它如何工作,太难了。所以我想,试着去工程化一个能思考的东西。这似乎是取得进展并更接近我们人类思考时所作所为的更自然的方式。

Sure. Yeah, so for me, in some ways, it didn't start in robotics per se. It started with artificial intelligence—the kind of brain behind the robots. For me, it was when I was in undergrad studying all kinds of things. It just seemed everything was interesting, but it also seemed that to do well, you got to pick something to specialize in. So the question was, okay, what is going to be the most interesting thing to spend my time on? And for me, it became pretty clear that it's so intriguing that we as humans can think. It's kind of what sets us apart, I would say mostly from other animals, is that we can think deeper, harder than most animals. So that to me was so fascinating. And of course, the first thing you might think is, okay, then study neuroscience, because that's studying how the brain works. But that discipline, I mean, fascinating, but it just seems so hard to make progress on—to dissect the brain and understand how it works, so difficult. So I figured, try to engineer something that thinks. It seems more natural as a way to make some progress and get closer to what we're doing as humans when we're thinking.

Host

你会关注神经科学研究吗?它会启发你吗?

Do you follow neuroscience research at all? Does that inspire you?

Pieter

哦,神经科学中有很多东西启发了我。我觉得它真的很有趣。这些天最启发我的可能是关于人脑有多么通用的发现。几年前的一些研究结果显示了——你可能以前听说过,Sam——它表明,例如,你可以把电极模式放在一个人的舌头上,这个人可以蒙上眼睛,可以是盲人。如果你舌头上的电模式以与你用眼睛看到的图像相同的模式被激活,你的大脑可以学会处理它并看见。他们有一个演示,有人实际上在通过舌头看东西的同时攀岩。所以这种事情让我震惊。大脑如此通用——你用来接收味觉输入的东西可以用来视觉。所以是的,我认为朝着通用性的方向前进真的很鼓舞人心。

Oh, there's a lot of things in neuroscience that have inspired me. I think it's really interesting. The thing that maybe inspires me most these days is these findings about how general the human brain is. There are findings from studies several years back now where it was shown that—you probably heard this before, Sam—it was shown that you can, for example, put an electrode pattern on a person's tongue, and this person can be blindfolded, can be blind. And if the electrical pattern on your tongue is activated in the same kind of pattern as an image that you would see with your eyes, your brain can learn to process that and see. And they had this demonstration where someone was effectively rock climbing while seeing through their tongue. So this kind of thing to me blows my mind. The brain is so general—what you use for taste input can be used to see. So yeah, I think it's really inspiring to go in that direction of generality.

Host

绝对,绝对。我们也许可以从聊聊你在 Covariant 的工作开始。我想我们上次谈话时,如果我没记错的话,你大概还有一年就要创立 Covariant 了。

Absolutely, absolutely. Let's maybe start with talking about what you're up to at Covariant. I think the last time we spoke, if I remember correctly, you were a year or so out from founding Covariant.

Pieter

是的,所以上次我们——嗯,是的,大概在创立 Covariant 前一年。当时我们开始看到一个趋势,似乎 awesome 和许多其他人在构建更智能的机器人方面已经取得了很大进展。而那似乎总是缺失的一块。如果你看看世界上那些为我们做有用事情的机器人——你去一家汽车工厂,环顾四周,太神奇了,这么多机器人,对吧?它们在造车,那太神奇了。但如果你看细节,这些机器人实际上是在一遍又一遍地重复同样的动作,非常勤奋,但也需要大量的结构。大多数问题不能仅仅通过重复动作来解决。所以三四年前似乎时机到了,我们可以让机器人更聪明。已经取得了足够的研究进展,可以开始把它带到现实世界,给机器人看、思考和对环境做出反应的能力,而不是仅仅遵循预编程的动作。那似乎会开辟比以往更多的机会。所以对我们来说,那真的是这个想法的触发点:好吧,我们认为我们可以开始构建能够看、反应、思考,并且能做比预编程机器人更多事情的机器人。

Yeah, so last time we were—well, yeah, still about a year before we started Covariant. What happened is we started seeing this trend that it seemed a lot of progress had been made by awesome and by many others on building more intelligent robots. And that always seemed the missing piece. If you look at robots out in the world that are doing useful things for us—you go to a car factory and you look around, it's amazing, so many robots, right? And they're building cars, and that's amazing. But if you look at the details, these robots are doing the same motion effectively over and over and over, very diligent, but also it requires a lot of structure. Most problems cannot be solved by just repeated motion. So it seemed three, four years ago that maybe time had come where we can make robots smarter. Enough research progress had been made to start taking on bringing this to the real world, to give robots the ability to think, see, think, and react to their environment, rather than just following pre-programmed motions. And that seemed like it would open up so many more opportunities than what was possible until then. So that was for us really the trigger to this notion: okay, we think we can start building robots that can see, react, think, and do many more things than pre-programmed robots can do.

Host

现在,能思考的机器人的机会显然非常广泛。你们是在追求特定的任务或问题吗?这是经典的抓取和放置类机器人问题,还是你们已经超越了?你们如何看待问题领域?

Now, the opportunity for robots that can think is obviously quite broad. Are there specific tasks or problems that you're going after? Is this the classical pick-and-place kind of robotics problem, or are you beyond that? How do you think about the problem domain?

Pieter

你提到这个真有趣,因为你说抓取和放置,对吧?而这正是我们在看的,但我们并不是这样起步的。所以我们开始时说,让我们全面调查一下,当我们与制造公司、仓储公司、农业公司以及建筑等公司交谈时,最紧迫的问题是什么。我们花了 Covariant 的前半年时间与大约 200 家不同的公司交谈,了解基本上,如果你能有一个真正智能的机器人,考虑到它的物理形态——你知道,一个标准的工业机器人——但更聪明,它能为你做什么?人们想要机器人做很多很多事情,但也很清楚的是,在物流、仓储、配送中心、电子商务……

It's so interesting you bring that up, because you say pick and place, right? And it's exactly what we're looking at, but it's not how we started out. So we started out and we said, let's do a full investigation of what are the most pressing problems when we talk with manufacturing companies, with warehousing companies, with agricultural firms, and so forth, construction. And we spent the first half year of Covariant talking with about 200 different companies and getting a sense for essentially, if you could have a really smart robot that, given its physical form factor—you know, a standard industrial robot—but smarter, what would it be able to do for you? And many, many things people wanted robots to do, but it became also very clear that in logistics, slash warehousing, distribution centers, e-commerce...

初期聚焦仓库自动化 Initial Focus on Warehouse Automation

Pieter

履约环节,那里大家都在拼命寻求帮助。他们希望机器人来帮忙。他们其实已经自动化了那些靠腿脚完成的工作——机器人在仓库里跑来跑去。很多人知道亚马逊的 Kiva 机器人,它们能钻到货架底下,把货架搬到人面前。还有其他系统,但思路都一样:跑腿的活儿已经自动化了。但那是结构化的自动化问题,跟我们如今讨论的 AI 不是一回事。而手上的活儿——人们用手做的事情——根本没有自动化,他们渴望自动化。他们想完成这些仓库的自动化进程。六个月后,我们非常清楚,这就是我们需要专注的问题,至少是第一步。希望我们能从那里泛化出去,我们相信可以。但我们的初始焦点就是为仓储、物流、履约等提供有效的抓取和放置解决方案。

Fulfillment, that's where everybody was just hurting for help. They wanted robots to come help. They had already automated what effectively is done with legs—they're running around the warehouse. A lot of people know about the Kiva robots at Amazon that can go under shelves and bring shelves to people. There are other systems too, but the same idea: the legwork has been automated. But that's a structured automation problem; it's not an AI problem in the same way that we're looking at AI these days. But the handwork—what people do with their hands—there just was no automation for it, and they wanted it. They wanted to complete that automation process of these warehouses. After six months, it became so clear to us that that's the problem we need to focus on, at least first. And hopefully we can generalize from there, and we think we will. But that's our initial focus: effectively pick-and-place type solutions for warehousing, logistics, fulfillment, and so forth.

适应机器人vs通用智能 Adapting Robots vs. General Intelligence

Host

我想象中,这个领域迄今的很多进展,都是让机器人去适应问题中非常具体的形态因素。比如我们有这些箱子,我们要以这种方式对齐这些箱子,这样机器人就能扮演一个较小的角色。也许你试图做的一部分,就是通过让机器人在真实世界中的操作能力变得更聪明、更灵活,来放宽这些约束。

I'm imagining that a lot of the progress in that domain to date has been adapting the robots to the very specific form factors of the problem that someone is trying to solve. We've got these boxes, we're going to align these boxes in this way, and so the robot can play some smaller part. Maybe a part of what you're trying to do is loosen those constraints by making the robots smarter and more agile in their ability to manipulate in the real world.

Pieter

完全正确。当你走进一个仓库,那里已经有很多传送带、移动机器人等,把东西带到大致需要的位置。但它们不搬运单个物品,它们往往搬运货架或箱子。下一步就是让工人或机器人查看那个箱子,挑出履行下一个订单所需的那个物品。要做到这一点,你实际上需要一种非常通用的能力。我想这对许多不熟悉仓库的人来说可能有点惊讶。许多仓库有数百万种不同的 SKU,即库存单位,而且这些 SKU 周转很快,包装各异等等。所以你不能——仅仅说“哦,这些是 SKU,我只要建立一个针对这些特定 SKU 的系统”是不够的。你实际上需要以某种方式构建一个 AI 系统,它理解物体的通用概念,能查看一个箱子、箱子里的物品——尤其是箱子本身很容易,但箱子里的物品——并理解:“哦,那里有物品,那边有一个,那边还有一个,那里有重叠的物品,那是我要捡的,我要这样捡:我的吸盘或夹爪要放在哪里,我要如何把它从箱子里拿出来,而不会把其他东西也带出来。”对我们人类来说,这些事情非常简单,你可以不假思索地完成。你甚至可以在做这件事的时候听播客,思考播客里讲的内容。但要让机器人做到这一点,真的非常非常难。这种事情,当然,达到 80% 的成功率不难,但要达到商业环境中可信赖所需的自主水平——几个九的可靠性,99.9 及以上——突然就变得非常非常难,因为这个系统现在需要理解它遇到的几乎所有情况。虽然大多数人可能对仓库不太熟悉,但这与自动驾驶汽车非常相似。我们展示自动驾驶汽车已经 10 年甚至更久了,今天的演示如果只看一两分钟,看起来不一定有多大不同,但它们变得更可靠了。但要达到完美或接近完美的可靠性仍然非常困难。当然,这就是挑战:一个长尾,总是有新的事物出现,无论是自动驾驶汽车在那个领域,还是仓库机器人,你总是看到新物品、新包装,你需要以某种方式理解这些,并可靠地抓取和放置。

That's exactly right. So when you go into a warehouse, there is already a lot of conveyors and mobile robots and so forth that bring things where they generally need to be. But they don't bring individual items; they tend to bring shelves or boxes. And then the next step is for either a worker or a robot to look into that box and pick out the one item that's needed to fulfill the next order. To do that, you actually need a very general kind of capability. And I think that's probably a bit surprising to many people who are not familiar with warehouses. Many of these warehouses have millions of different SKUs, which are stock keeping units, and these SKUs also turn over very quickly; they have different packaging and so forth. So you cannot—it's not enough to say, 'Oh, these are the SKUs, I'm just going to set up a system that's ready for these specific SKUs.' You actually need to somehow build an AI system that understands the general notion of object, that can look at a box, items in the box—especially the box itself is easy, but the items in the box—and understand, 'Oh, there are items there, and there's an item over there, another one over there, there's overlapping items there, and that's the one I want to pick, and here's how I'm going to pick it: where I'm going to place my suction cup or my gripper, and this is how I'm going to maneuver it out of that box without flinging anything else out of the box too.' For us as humans, these things are very simple to do; you can do them without thinking. You can be listening to a podcast or something while doing this and thinking about what's being talked about. But to get a robot to do that, it's actually really, really hard. It's one of those things where sure, it's not too hard to get 80% success, but to get the levels of autonomy that you need to be trustworthy in a commercial setting—several nines of reliability, 99.9 and above—all of a sudden becomes really, really hard, because this system now needs to understand pretty much any situation it encounters. And while most people are probably not so familiar with warehouses, it's very similar to self-driving cars. We've had demos of self-driving cars for 10 years or longer, and the demos today don't necessarily look all that different if you watch a one-minute, two-minute demo of a self-driving car, but they've become more reliable. But it remains very hard to get to perfect or near-perfect reliability. Of course, that's the challenge: a long tail of always new things you can encounter, either as a self-driving car in that domain or as a warehouse robot, where you're always seeing new items, new packagings, and somehow you need to make sense of that and reliably pick and place.

受限环境与自动驾驶类比 Constrained Environments and Self-Driving Analogy

Host

这可能是一个有趣的类比。我想起这个领域的许多人——也许吴恩达是最著名的之一——他们谈过这样的观点:自动驾驶汽车将首先在约束环境中实现,比如机场的自动驾驶公交专用道,而不是通用的自动驾驶能力。我们在工业领域已经做了很多这样的工作。你谈到了传送带机器人,以及我们如何真正约束了那个环境。这是否说明,即使对自动驾驶汽车来说,那也可能不够?还是说这个类比推得太远了?

It's maybe an interesting analogy. I'm thinking of many folks in the field—maybe Andrew Ng is one of the most well-known—who have talked about this view that self-driving cars will first happen when we constrain their environment, maybe the autonomous bus lane in an airport as opposed to general self-driving ability. We've already done a lot of that work in the industrial realm. You've talked about the conveyor robots and how we've really constrained that environment. Does that say that maybe that's not going to be enough even for self-driving cars, or is that pushing the analogy too far?

Pieter

是的,我认为这是一个非常好的问题。在自动驾驶汽车方面,在约束环境中创造真正价值是否可能?例如,最自然的就是自动驾驶汽车专用车道。这可能会让事情变得容易得多——很可能容易得多。但与此同时,我的意思是,那些车道在哪里?他们标记了吗?火车当然就是这样。我的意思是,火车——你知道,不要踩到铁轨上,因为那很危险,那不是你应该待的地方。所以,我绝对可以看到一种版本,如果你走得那么远,真正划定这只能是自动驾驶汽车的领域,我怀疑今天的技术可能会创造很多价值,但代价是大量房地产被那些专用轨道等占用。可能这就是我看到的情况。但当然,我个人并不是推动自动驾驶汽车的人,但我有很多朋友在做这个。高速公路似乎是足够受约束的,因为通常人们不会在那里走动,而且一切都是汽车和摩托车。所以那里有不同的环境。

Yeah, I think it's a really good question. On the self-driving car side, is it possible to create real value within constrained environments? For example, the natural one would be lanes dedicated to self-driving cars. It might make it much easier—probably would be much easier. But at the same time, I mean, where are those lanes? Do they mark them? Trains are like that, of course. And I mean, trains—you know, not to step on the railroad tracks because that's dangerous, that's not where you're supposed to be. So definitely, I could see a version of if you go that far and really delineate this is self-driving car territory only, I suspect today's technology would probably create a lot of value, but at the cost of a lot of real estate being taken up by those dedicated tracks and so forth. Probably that's what I see happening there. But of course, I'm not the one pushing self-driving cars personally, but I have a lot of friends doing that. It seems highways are sufficiently constrained because typically people aren't walking around there, and everything is cars and motorcycles. So there are different environments there.

工厂经验对自动驾驶的启示 Lessons from Factories for Self-Driving

Host

我想知道的是——这可能有点跑题——你在工厂的经验中,有没有学到一些东西,可以指导你如何处理自动驾驶汽车问题和整个环境约束问题,如果你明天决定去尝试解决那个问题的话?

I think what I'm wondering is—and this is maybe a bit of a tangent—are there things that you've learned in your experience in factories that inform how you might approach the self-driving car problem and this whole constraining the environment, if you were to decide to go try to solve that problem tomorrow?

Pieter

在那个方面没有直接的教训,但我认为在另一个方面有一个非常强的类比,那就是追逐事件的长尾。无论是物品还是物品的配置,还是交通状况,达到 80% 的覆盖率并不难——通过相对快速的努力是非常可行的。但要达到你实际良好运行所需的可靠性,就有这个长尾,你几乎从未见过的东西。每一个你都几乎没见过,但当它们加在一起时,数量就很多了。所以你需要以某种方式观察所有这些罕见事件。

No direct lessons in that way, but I think there's a very strong analogy in a different way, which is the chasing the long tail of events. Whether it's items or configurations of items, or it's traffic situations, getting 80% coverage isn't too hard—that's very feasible with a relatively fast effort. But getting to the reliability that you need to actually function well, there's this long tail of things that you barely ever see. Each one of them you barely ever see, but when you add them all up together, there's a lot of it. So you need to somehow observe all these rare events.

罕见事件与泛化 Rare events and generalization

Pieter

这些事件加在一起,构成了并非罕见的概率质量。这就是为什么它如此重要。如果只是一个百万年一遇的罕见事件,谁在乎呢?但如果有上百万个这样的百万年一遇的事件,突然之间,你每年都会遇到一个事件,对吧?而这就是正在发生的事情。有一个非常罕见的事件的长尾,你需要以某种方式训练你的神经网络去应对。在大多数情况下,我想自动驾驶的人也会训练神经网络,但对我们 Covariant 来说肯定如此——训练你的神经网络,让它们对物体、物体的形状、可抓取性、物体在料箱中的相互作用、当你试图取出一个物体时、或者当你放置物体时物体之间如何相互作用,有一个更通用的理解。你不可能在没有能力针对每个场景进行训练的情况下,理解你可能遇到的每一个场景。你需要训练出能泛化的东西。我认为这在两种情况下都是如此,无论是自动驾驶还是仓库拣选。你都需要那种泛化能力。

Events that together add up to non-rare probability mass. That's why it's so important. If it was just one rare event that happens once in a million years, who cares? But if there are a million of those events that happen once in a million years, all of a sudden every year you have an event, right? And that's kind of what's happening. There is this long tail of very rare things that you somehow need to train your neural networks on. In most cases, I imagine self-driving people train neural networks too, but definitely for us at Covariant, train your neural networks to have a more general understanding of objects, their shapes, their graspability, their interactions when they're in a bin, when you try to pick one out, or when you place, how that interacts with each other. You don't understand that for every scenario you might encounter without having the ability to train on every scenario. You need to train something that generalizes. And I think that that is kind of the case in both situations, whether it's self-driving or warehouse picking. You need that generalization ability.

Host

我记得我们上次对话时,你非常坚定地支持端到端深度学习。我想知道,你在现实世界和工业领域解决问题的经验是否改变了这一点。你知道,我们谈过端到端、集成学习,以及融入物理知识之类的东西。你现在还那么强烈地坚持单一的端到端深度学习模型吗?

I remember from our last conversation you were very staunchly for end-to-end deep learning. And I wonder if your experience solving problems in the real world and the industrial domain has changed that at all. You know, we talked about end-to-end, ensembles, and incorporating physical knowledge of the world, that kind of thing. Do you still feel as strongly about a single end-to-end deep learning model?

Pieter

所以我觉得,是的,你对我们上次聊的内容的观察完全正确。也许我稍微调整了一些东西,但我仍然非常接近上次的立场。首先,我想说有两件不同的事情需要考虑。一件是研究,学术研究,当你写论文的时候。你认为大多数新颖性、惊喜、新事物会出现在哪里?我的感觉是,另一方面,当我们把东西放到现实世界中时,通常用你从物理学等获得的先验知识来补充它,可以给你一个更强大的系统。从纯粹的学术研究角度来看,我的感觉是,那些更经典的方法已经被研究得很透彻了,在新颖性和惊喜方面,以及那些让你觉得“哇,没想到这也能行,现在计算机、AI、机器人能做到这个”的事情上,已经没有太多空间去突破边界了。所以在学术领域,我非常主张尽可能纯粹地依靠学习,因为那里有最多的新颖性。在更实际的方面,你当然——扔掉任何有保证、保证有效的东西是愚蠢的。如果你知道你对环境有 3D 理解,那么,环境是 3D 的,你可以构建它的 3D 地图,或者至少是深度图或任何你想构建的东西,不使用这种概念,只说“嘿,我们永远不会明确告诉神经网络 3D 很重要”,这似乎很愚蠢。这似乎是浪费步骤。说出来听起来很疯狂,是的。

So I think, yeah, you're absolutely right with your observation about what we chatted about last time. And maybe I've adjusted things a little bit, but I'm still very close to where I was last time. So first of all, I would say there are two different things to think about. One thing to think about is research, academic research, when you write papers. Where do you think most of the novelty, surprises, new things are going to be? And my sense is that while on the other hand, when we put things in the real world, often complementing it with prior knowledge you have from physics and so forth can give you a stronger system. From a purely research academic point of view, my sense is that those more classical approaches have been researched so well already that there is less room to push the boundary in terms of novelty and surprises and things that you know, 'oh wow, didn't expect this to be possible, and now a computer, AI, robot can do this.' And so in the academic realm, I'm really big on going as purely on learning as possible, because there is the most novelty there. On the more practical thing, which you're of course, it's stupid to throw away anything that has guarantees and is guaranteed to work. If you know that you have maybe a 3D understanding of the environment, well, the notion that the environment is 3D and you can build a 3D map of it, or at least a depth map or whatever you want to build, it seems stupid to not use that kind of notion and just say, 'Hey, we'll never tell the neural net explicitly that 3D matters.' It seems a waste of steps. Sounds crazy saying it, yeah.

Host

话虽如此,我们发现了一件事,我认为这与这有关,那就是编码先验知识的棘手之处在于,很难对先验知识做到非常精确。所以每当你认为你在编码先验知识时,你可能包含了在某些情况下实际上略微不真实、略微不正确的东西。

That said, here's one thing that we found, I think that ties into this, which is the tricky thing with coding in prior knowledge is that it's very hard to be super precise about prior knowledge. And so whenever you think you're encoding prior knowledge, you might be including something that's actually slightly untrue, slightly incorrect in some situations.

Pieter

你能举个例子吗?

Can you give an example?

Host

那么,举个简单的例子。假设你说:“嘿,用吸盘抓取物体的一个好方法是去一个平坦区域的中间。去平坦区域的中间。”这看起来是一个很好的启发式方法,然后说:“好的,那就是你应该去的地方。”显然,如果你从第一天起就把它编码进去,它会比你说“嘿,只要学习在哪里抓取”效果更好,因为现在它必须弄清楚中间可能比边缘更好等等。但是,硬编码“中间是好位置”的缺点是,现在突然之间,也许你有重叠的物体,也许中间不是一个好的到达点,或者你需要以另一种方式滑出东西,或者也许是一个重心离中间很远的物体,或者也许有一种方法可以将多个吸盘偏离中心放置,这实际上比把东西放在正中间更稳健。那么,如果你硬编码,你说:“好吧,我们只是硬编码,一旦我们理解场景,你就去中间。”那么,突然之间,你如何围绕这一点进行特例处理?你有点卡住了。所以我喜欢的总体理念是,如果你有先验知识,不要把它硬编码到你的系统中。相反,用它来有效地生成数据。任何先验知识通常都可以转化为数据增强或数据生成方案。所以,如果我们考虑抓取,你也许能通过这种方式生成一堆示例抓取,它们是好的训练数据,因为它们很好。它们不一定完美,但它们很好,它们会帮助你。但我认为,随着你遇到更多的边缘情况,这最终会让你不断改进,数据驱动的方法会给你一个例外,得到一个特殊的例外,而无需 if-then-else、if-then-else。只是更多的数据会说明一切,新网络会吸收它。所以我认为,从长远来看,这真的是我倾向于的看法:利用所有先验知识来帮助你训练神经网络,因为先验知识可以帮助你为神经网络生成大量数据。

So let's say, here's the simple example. Let's say you say, 'Hey, a good way to pick up an object is to go to the middle of a flat region with a suction cup. Go to the middle of a flat region.' That might seem like a good heuristic, and say, 'Okay, that's where you should go.' And obviously, if you code that in from day one, it'll work better than if you say, 'Hey, just learn where to pick something,' because now it has to figure out that the middle might be better than an edge and so forth. But the downside of hard coding in that the middle is a good spot is now all of a sudden maybe you have overlapping objects, and maybe the middle is not a great spot to reach, or you need to kind of slide things out another way, or maybe it's an object that has a center of gravity that's not even close to the middle, or maybe there is a way to place multiple suction cups off center in a way that is actually more robust than placing things right in the middle. And so what then happens if you hard code and you say, 'Well, we're just gonna hard code that once we understand the scene, you go for the middle.' Well, all of a sudden, how do you now special case around that? And you're kind of stuck. And so the general philosophy I like is this notion that if you have prior knowledge, don't hard code it into your system. Instead, use it for effectively data generation. Any prior knowledge can usually be turned into a data augmentation or a data generation scheme. And so you might be able to generate a bunch of example grasps that way, if we're thinking about grasping, and they're good training data because they're good. They're not necessarily perfect, but they're good, and they'll help you. But I think that's ultimately going to let you keep improving as you encounter more corner cases, where a data-driven approach will give you an exception to get an exceptional exception without if-then-else, if-then-else. It's just the more data speaks for itself, and the new net will absorb that. And so I think that that's really the way I tend to see it in the long term: use all that prior knowledge to help you train the neural network, because that prior knowledge can help you generate so much data for a neural network.

Host

这是统一这两个想法的一个非常有趣的方式。当你处理你正在处理的那类问题时,想象一下,最终你试图构建一个尽可能通用的产品或平台,但每个具体问题都有很多类似定制咨询的定制化。我很好奇你能谈谈这个,以及你是如何应对的。

That's a really interesting way to unify the two ideas. When you approach the types of problems you're approaching, imagining that ultimately you're trying to build a product or a platform that is as general as possible, but there's a lot of kind of custom consulting-ish customization for each individual problem. I'm curious if you could speak to that and how you've approached it.

Pieter

是的,我认为如果你试图建立一家可持续发展的公司,这正是要避免的陷阱。如果你陷入大量的一次性咨询工作中,你就不是在构建产品,你不是在构建可以重复交付的东西。但与此同时,当你被要求向客户交付某样东西时,它必须是他们需要的。你不能只说:“嘿,我们构建的不是你想要的,不是你需要的。”因为那样他们只会说:“是啊,那我们不想要了。”所以,对,这不会正是你需要的,但它也会是 10 个客户都不完全需要的,没错。所以那里有一种自然的张力。我们一直在思考的方式是,首先进行大量的市场研究。当我们做市场研究时,我们看到了实际上很大的……

Yeah, I think that is exactly the trap to avoid if you try to build a sustainable company. If you get into a ton of one-off consulting efforts, you're just not building a product, you're not building something that you can repeatedly deliver. But at the same time, whenever you're asked to deliver something to a customer, it has to be what they need. You can't just say, 'Hey, we built not what you want, not what you need,' because then they'll just be like, 'Yeah, well then we don't want it.' So right, this isn't going to be exactly what you need, but it's going to be not exactly what 10 customers need too, exactly. So there's a natural tension there. And the way we've been thinking about this is by first of all doing a lot of market research. And so as we did the market research, we saw effectively big...

应用概览 Overview of applications

Pieter

从图像上看,我们看到了一个反复出现的主题:机器人需要更高的智能。如果我们能造出让机器人理解环境并做出反应的东西,那就能驱动很多不同的解决方案。但聚焦到我们目前交付技术的滩头阵地,在仓储领域,我们决定主要服务三个应用。第一个是订单拣选:机器人面对存放在仓库里的物品,需要一次一个地拣出来放进运输容器,然后发货。这就是我们拣选的内容。第二个是分拣墙,有时也叫“put-to-light”。在分拣墙场景中,很多订单进入仓库,有人或机器人已经在仓库里转了一圈,为所有这些订单收集了物品——可能是一百个甚至更多订单——全部放在一个购物车或料箱里,没有按订单分类。他们只是收集了完成那一百个订单所需的所有东西。然后这些物品出现在分拣墙前,机器人的任务就是一次一个地从购物车里取出物品,扫描,然后放到与该特定订单关联的临时存储位置,直到该订单从购物车里完全配齐。之后分拣墙后面的人会取走、打包并发货。这就是分拣墙。第三个是包裹单件分离、导入和分拣。这些设施里有很多包裹在流转,它们也需要被单件分离,因为有一大批包裹进来。你能可靠地把它们一个一个放到传送带上,让它们通过扫描仪等设备,被分拣到正确的出库卡车吗?所以我们首先聚焦这三个应用,我们可以构建非常标准化的东西。任何需要分拣墙、订单拣选或导入分拣的客户,我们都能交付,而不必为每个特定部署做定制化咨询。

Picture-wise, we saw a recurring theme of more intelligence for robots. If we can build something that can make robots understand their environment and react to it, that's going to power a lot of different solutions. But zooming in on the beachheads where we're delivering our technology right now, in warehousing there are really three main applications we decided to cater to. The first is order picking: a robot is presented with items that were in storage and needs to pick them one at a time to put into a shipping container, and then things get shipped off. That's what we're picking. The second is put-walling, or sometimes put-to-light. In the put wall, many orders come into a warehouse, and somebody or a robot has gone around the warehouse and collected the items for all these orders—maybe a hundred orders or more—and it's all collected in, say, a shopping cart or a bin, not sorted by order. They just collected everything needed to fulfill those hundred orders. That shows up at a put wall, and now the robot's job is to take one item at a time out of that shopping cart, scan it, and place it into a temporary storage associated with that specific order, until that order has been completely filled from that shopping cart. Then somebody behind the put wall will take it, pack it up, and ship it off. So that's put-walling. The third is parcel singulation, induction, and sorting. There are a lot of parcels going around these facilities, and they need to be singulated because there's a bulk of parcels coming in. Can you reliably place them one at a time onto a conveyor so they can go through scanners and so forth to be sorted to the right outbound truck? So these three are the ones we're focused on first, where we can build something very channel. Anybody who needs a put wall, order picking, or induction sorting, we can deliver to that without having to do one-off consulting for each specific deployment.

受限问题定义 Constrained problem definitions

Host

这些问题的定义是否足够受限,以至于在分拣墙这类场景中,你能覆盖全部或绝大多数用例?还是说它们还会被进一步约束,比如单个 SKU 必须是盒装的,而不是散装的螺栓之类的东西?打击区需要多精细?

Are those problem definitions constrained enough that within put-walling, for example, you can cover all or the vast majority of use cases? Or are they further constrained by, for example, the individual SKUs having to be in boxes as opposed to loose bolts or things like that? How granular does the strike zone need to be?

Pieter

是的,你问得真有意思。我很高兴你问这个,因为我本来就应该提到这一点。我们构建这个系统最有趣的地方在于——至少让我特别兴奋的是——我们在所有三个应用领域背后构建了一个单一系统。同一个神经网络同时为分拣墙、订单拣选和导入分拣训练。你可能会想,为什么?我的意思是,这些是有些不同的问题,但核心是一样的。你看一堆物品——当然,包裹对分拣墙通常是小的物品,对订单拣选则是各种尺寸——但根本上你还是在看一堆物品,确保你可靠地一次一个地拣取,可能还要扫描,然后放置。所以它在三个领域都能泛化。在它们内部,如你所说,有些设施可能专注化妆品,有些专注药品,有些专注服装,还有些可能专注电气用品等等。同样,所有这些领域都用同一个神经网络。我认为这实际上是最好的做法,因为即使在单一领域内,你也不可能覆盖所有东西。这些 SKU 及其包装的周转率太高了,你不可能覆盖一切。你不能说,这些是我们需要处理的 100 个物品,让我们为它们做点专门的东西。你需要真正能泛化的东西。通过以比化妆品更通用的方式构建,你实际上在神经网络中注入了更多关于它倾向于看到什么的专业知识,什么构成一个单独的物体,它看到的场景的 3D 情况是什么,接近、抓取、移除的最佳方式是什么,然后从 A 点到 B 点的最快轨迹是什么,等等。

Yeah, it's really interesting you ask that. I'm glad you're asking because I should have mentioned that anyway. What's so interesting about how we're building this is that—at least what I'm so excited about—is that we're building a single system behind all three application domains. The same neural networks are trained for put-walling, order picking, and induction sorting. You might think, why? I mean, those are somewhat different problems, but the core is still the same. You look at a bunch of items—sure, parcels versus put-wall often small items versus order picking a range of sizes—but you're still at the fundamental level looking at a bunch of items and making sure you're picking reliably one at a time, possibly scanning, and then placing. So it generalizes across all three. And within them, as you said, some facilities might focus on cosmetics, others on pharmaceuticals, others on apparel, yet others on electrical supplies, and so forth. Again, it's all the same neural network for all those domains. I think that's actually the best way to do it, because even within a single domain you cannot cover everything. There's so much turnover on all these SKUs and their packaging that you cannot cover everything. You cannot say, these are the 100 items we need to do, let's have something specific for that. You need something that really generalizes. By building it in an even more general way than just cosmetics, you actually build more expertise into the neural networks about what it tends to see, what makes for an individual object, what is the 3D of the situation it's looking at, what's the best way to approach, to grasp, to remove, and then what's the fastest trajectory to bring from point A to point B, and so forth.

模型范围 Scope of models

Host

明确一下,你们的模型主要是聚焦于操作任务,还是它的某种泛化?我假设一个给定的机器人可能有一组工具需要选择,你需要确定最佳接触点,以及一系列与所谓“操作”相关的因素。然后我想在操作之外,还有其他部分更像传统优化,不一定需要神经网络。但我想这正是我要请你纠正我的地方:在你们构建的这三个问题领域内,模型的范畴是什么?

To be clear, are your models primarily focused on the manipulation task or whatever the generalization of that is? I'm assuming a given robot might have some set of tools it needs to select, and you need to identify the best contact point and a variety of factors associated with, let's call it, manipulation. And then I imagine beyond manipulation there are other parts of the problem that are more like traditional optimization and not necessarily places where you require neural networks. But I guess that's what I'm asking you to correct me on: what's the scope within these three problem domains of the models that you're building?

Pieter

是的,你提到了几个有趣的点。一个是确实不是每个物品都能用同一个末端执行器来抓取。有所谓的工具快换器,机器人可以在飞行中更换末端执行器,比如,哦,我现在应该用夹爪,或者用双吸盘夹爪,或者六吸盘末端执行器。你可以根据需要切换。但事实证明——最初有点意外,现在不意外了——在大多数仓储设施中,吸盘能覆盖非常非常多的场景。还有一点很有趣:神经网络在思考如何抓取和放置这些物体时是同一个,无论用吸盘还是夹爪,也无论用多少个吸盘。这有点深入了,但最简单的理解方式是:想象你有一个单一的神经网络,网络的主体负责思考抓取。你不只是给它它要抓取的场景;我们还给它末端执行器的配置。基于场景和末端执行器的组合,它决定怎么做。这样,你又能泛化到原本不可能的范围,这真的很酷。所以你不会为新末端执行器做一次性训练。不会,所有部署和所有类型的末端执行器都共享学习成果。

Yeah, so there are a couple of interesting things you're touching upon. One is that indeed not every item can be picked with the same end effector. There are things called tool changers, where a robot can change the end effector on the fly and say, oh, actually I should use a gripper now, or a two-suction-cup gripper, or maybe a six-suction-cup end effector. You can switch that as needed. Though it turns out—somewhat surprisingly initially, not surprising anymore now—suction cups go very, very far in most warehouse facilities. What's also interesting is that the neural network is the same when it's thinking about how to grasp and place these objects, whether it's a suction cup or gripper, and independent of the number of suction cups it's using. This is going a little bit deeper, but the simplest way to think of it is: imagine you have a single neural network, the main body of the network that thinks about grasping. You don't just give it the scene it's looking at to grasp in; we also give it the configuration of the end effector. Based on the combination of scene and effector, it decides what to do. That way, again, you can generalize beyond what's possible otherwise, which is really cool. So you don't have one-off training for new end effectors. No, it's all shared learnings across every deployment and every type of end effector.

输入数据 Input data

Host

这里插个问题:你们是否也会给它与产品相关的表格化数据,比如 SKU 尺寸之类的?还是主要依赖视觉或传感器输入?

Just punching in with a question here: are you also giving it tabular-style data associated with the product, like SKU dimensions, that kind of thing? Or are you primarily focused on visual or sensor inputs?

Pieter

这些数据通常可以获得,但并非总是如此。我们倾向于不依赖它,因为依赖它可能是一个很强的假设。但我的意思是,有些地方可以提供这些数据,但我们更倾向于不依赖它。

So that data can often be made available, but not always. We prefer not to rely on it, just because it can be a pretty strong assumption to rely upon. But I mean, there are places where it can be made available, but we prefer not to rely on it.

Host

好的。是的,你提到的另一件事是某种优化。

Okay. Yeah, and the other thing you brought up is kind of optimization.

仓库机器人的优化与可靠性 Optimization and Reliability in Warehouse Robotics

Pieter

是的,优化非常重要。事实证明,速度——我的意思是,可靠性是关键,但一旦可靠了,速度也很重要。如果你的机器人以一半的速度工作,它创造的价值就减半。在仓库或工厂这样的设施里,如果你跟不上节奏,可能会拖累整个流程。所以保持所需的速度非常重要。有很多方法可以优化机器人轨迹,但一旦你拿着可能会甩动的物体——因为它们又重又软——事情就变得有趣了。突然间,你不能像只考虑机器人本身那样用分析的方法来处理了。而学习——你原本可能觉得对运动无关紧要——突然之间,学习又能发挥作用,让你做得比不学习更好。

Yes, optimization matters a lot. It turns out speed is—I mean, reliability is key, but then once it's reliable, speed also matters. If your robot works at half the speed, it's creating half the value. And in a facility like a warehouse or factories, if you're not keeping the pace, you might be bottlenecking the entire process. So keeping the required pace is really important. There are many methods to optimize robot trajectories that can help, but things get really interesting once you're holding objects that might be flinging because they're heavy and floppy, and so forth. All of a sudden, you can't go about it as analytically as you would when you just think about the robot itself. And all of a sudden, learning—which you thought maybe doesn't matter for motion—all of a sudden, learning can matter again to do better than you can without learning.

Host

有意思。继续刚才那个物体甩动的例子——你知道,吸盘偶尔会掉东西。你觉得这算不算超出范围,然后每天晚上有个人类到处捡起掉落的东西重新放好?还是说你们也在尝试把智能构建到机器人里,让它去补救、捡起掉落的物品,或者别的什么方案?

Interesting. Kind of continuing that example of objects flinging—you know, there's a suction cup, every once in a while something's gonna drop. Do you consider that kind of out of scope, and every night a human goes around and picks up all the dropped things and re-bends them? Or are you also trying to build intelligence into the robots to remediate that, pick up the items that fall, or some other scenario?

Pieter

是的,我觉得你问的这个问题,核心正是为什么仓库机器人是一个绝佳的领域。你需要非常可靠,但如果偶尔需要人工干预,那也没关系。比如一个人每天结束时花 10 分钟清理机器人留下的一些小问题,这没问题。显然你不希望机器人弄坏任何东西,但如果它只是把东西放在箱子旁边,然后在轨迹中掉落——当然我们不希望发生这种情况,我们希望它永远 100% 可靠。但这里的美妙之处在于,一旦你达到比如 99.9% 的可靠性——这已经是非常高的标准,很难达到——你就能创造大量价值。但 99.9% 不是 100%,而 99.9% 可能不适合自动驾驶。你可能不希望一辆 99.9% 可靠性的自动驾驶汽车上路,因为这意味着,你知道,每过一千次路口——如果我们这样计算的话——它就会出一次事故。那就是 99.9% 的过路口可靠性,这不够好,你永远不会部署它。但如果你想想——这正是商业化的关键,我们详细思考的地方——机器人何时开始创造价值。从我们的角度来看,以及我们在客户那里看到的一切,价值创造大约从 99.9% 这个点开始。因为一个机器人通常每小时执行 500 到 2000 次操作,取决于工作站类型。有些工作站需要更快,有些工作站有更多参与,比如额外扫描等等,所以会慢一些。但即使在一个慢速工作站——每小时 500 次——在 99.9% 的可靠性下,意味着你每两小时才犯一次错误。这意味着没有人需要不断管理或照看这个机器人;每两小时一次,当然,我们可以快速修复一下。这就是在创造真正的价值。在我看来,这就是基准。99.9% 当然是——我们希望能更进一步,别误会——但这就是机器人变得有用,而不是需要保姆的时刻。

Yeah, so what you're getting at there, I think, is at its core why warehouse robotics is such a great place to be. You need to be very reliable, but if just every now and then something needs a human intervention, that's okay. If a person needs to spend 10 minutes, let's say, at the end of the day to clean up a couple of things that the robot did, it's no problem. Obviously you don't want the robot to break anything, but if it just put something next to a bin and drops it in trajectory, of course we don't want it to happen—we want it to be always 100% reliable. But the beauty here is that you can create a lot of value once you are, let's say, 99.9% reliable, which is a really high bar and very hard to meet. But 99.9% is not 100%, and 99.9% maybe wouldn't cut it for self-driving. You might not want a 99.9% self-driving car on the road, because it means, you know, one in a thousand times it crosses an intersection—if that's how we count it—it gets into an accident. That would be 99.9% reliability on crossing intersections, and that's not good enough; you would never deploy that. But if you think about—and that's really where commercialization is, what we think about in great detail—is when does a robot start creating value. The value creation in our perspective, and everything we've seen within customers, starts roughly at that 99.9% mark. Because a robot would typically do anywhere from 500 to 2000 operations per hour, depending on the type of station. Some stations have to be faster, some stations have more involvement like extra scanning and so forth, so they're a bit slower. But even at a slow station—a 500 per hour station—at 99.9%, it means you're making a mistake once every two hours. That means nobody has to constantly manage or babysit this robot; it's once every two hours, sure, let's quickly fix something. And that's creating real value. In my mind, that's kind of the benchmark. 99.9% is, of course—we're looking to go further than that, don't get me wrong—but that's where it becomes a robot that is helpful, as opposed to a robot that needs a babysitter.

无监督与强化学习的交汇 Junction of Unsupervised and Reinforcement Learning

Host

是的,我可以就工业 AI 和机器人聊上整整一期节目、整个采访。我写过一些关于工业 AI 的文章,觉得这是一个迷人的话题,但我也想谈谈你之前提到的一些研究兴趣。特别是你谈到了无监督学习和强化学习的交汇点,以及你在那里做的一些工作。无监督学习从研究角度来看一直是一个目标,因为它具有数据效率,但最近在自然语言处理中它真正大放异彩。请谈谈那个交汇点,以及你觉得它有趣的地方。

Yeah, so I could continue talking about industrial AI and robotics for the entire episode, the entire interview. I've written a bit about industrial AI and I find it a fascinating topic, but I also want to touch on some of the research interests that you mentioned earlier. In particular, you talked about this junction of unsupervised and reinforcement learning and some of the work you're doing there. Unsupervised learning is something that has been a goal from a research perspective for a while, for its data efficiency, but it's really played out in big ways in natural language processing recently. Talk a little bit about that junction and where you see it being interesting.

Pieter

是的,正如你所知,Sam,我研究强化学习已经很久了。强化学习就是试错学习,这是我们作为人类非常熟悉的东西。当我们看到孩子学会爬、学会跑,那就是强化学习。当我们训练狗坐下,那也是强化学习:我们说“坐下”,它坐下了,我们给它零食;如果它不坐下,我们可能会呵斥它。从这些反馈中,它明白了“坐下”是什么意思,并获得了这个技能——倾听和执行。这就是强化学习。到目前为止,我们看到强化学习的许多重大成功都是在模拟环境中。可以说最著名的成功是 AlphaGo,来自 DeepMind 的最强计算机围棋程序,击败了最强的人类棋手。强化学习在其中扮演了非常重要的角色,因为它通过反复与自己对抗来变得更好,但这一切都在模拟中。许多视频游戏、许多模拟机器人结果也是如此。但问题是,当你试图从模拟过渡到现实世界时,你可以做几件事。当然,你可以说,嘿,也许我的模拟器可以完美匹配现实世界,这会有帮助。但最终,你想要更高的数据效率。你不想只从奖励中学习;你想以其他方式学习。你不想只是做某事一分钟,然后在最后知道对错。不,你想要更密集、信息量更大的反馈。自然,强化学习在其原始形式下并不能真正做到这一点。但正如你所说,我们在自然语言处理中看到了无监督学习,即仅从海量文本中学习。打个比方:在自然语言处理中,你可能想对情感进行分类——这是正面评论还是负面评论,或者这是一个正确的英语句子还是一个破碎的英语句子,诸如此类。你想学习分类。与其生成正面和负面文章的例子——因为那需要大量标注训练——你不如先在所有你能找到的互联网文本上训练,只预测句子中的下一个词。一旦你在这方面做得很好,你必然在神经网络内部内化了一些关于语言及其意义的东西,这样现在从几个例子中你就能理解正面或负面的情感分类。所以问题是,我们能在强化学习中做同样的事情吗?我们能否拥有这种辅助……

Yeah, so as you know, Sam, I've been working on reinforcement learning for a very long time. Reinforcement learning is trial-and-error learning, and it's something we're very familiar with as humans. When we see a child learn to crawl, learn to run, that's reinforcement learning. When we train a dog to sit, that's reinforcement learning: when we say sit and it sits, we give it a treat; if it doesn't sit, we might yell at it. From that feedback, it figures out what it means to sit and acquires that skill—listening and executing. So that's reinforcement learning. A lot of the big successes in reinforcement learning that we've seen so far have been in simulation. Arguably the most famous success is AlphaGo, the best computer Go player, beating the best human players, out of DeepMind. Reinforcement played a very big role in that because it was playing against itself over and over to become better, but it's all in simulation. Same with many video games, a lot of simulated robotics results. But the question is, when you try to transition from simulation to the real world, there are a few things you can do. Of course, you can say, hey, maybe my simulator can be perfectly matched with the real world, and that can help you. But ultimately, you want to be more data-efficient. You want to not just learn from reward; you want to learn in other ways. You want to not just do something for a minute and at the end know whether it was right or wrong. No, you want feedback that's much denser, much more informative. Naturally, reinforcement learning doesn't really do that in its vanilla form. But we've seen, as you said, in natural language processing, we've seen unsupervised learning, which is learning from just vast amounts of text. To make the analogy: in natural language processing, you might want to classify sentiment—is this a positive or a negative review, or is this a correct English sentence or a broken English sentence, things like that. You want to learn to classify that. Instead of just generating examples of positive and negative articles—because that would take a lot of annotated training—you would first train on just predicting the next word in a sentence, on all the text you can find on the internet. Once you're really good at that, you must have internalized something inside the neural network about language and its meaning, such that now from a few examples you understand positive or negative sentiment classification. So the question is, can we do the same thing in reinforcement learning? Can we have this auxiliary

用对比学习弥合差距 Bridging the gap with contrastive learning

Pieter

一个无监督任务,它包含大量信号,虽然不是直接的强化信号,但仍然是神经网络可以学习的信号,以便随后快速习得新技能。过去一年半里我们在这方面做了不少工作。我们看到的——而且不只是我们,纽约大学、蒙特利尔的一些人也是——我们实际上已经能够弥合这种差距:比如一个机器人要学习某项技能,比如跑步或爬行,当它能够随时完全访问自身全部状态——所有关节角度、位置、朝向——那是完全访问、完整状态访问,这样学习总是很快。但如果它只能看到自己的视频流,那可能是一张 100×100 的图像,也就是 1 万个像素,这比简洁的机器人状态描述要高维得多,学习就会很慢。所以我们看到的是,从状态访问学习(相当高效)与仅从图像输入学习之间的巨大效率差距。问题是如何弥合这个差距?我们实际上是用无监督学习做到的。当机器人进行试错时,它不只是关注奖励,根据到目前为止的好坏表现来改进,它同时也在对看到的图像进行无监督学习。具体来说,我们做的是对比学习。对比学习的思路是:你想知道图像里有什么,但如果没人告诉你图像里有什么,你怎么学?想法是这样的:假设你下载了两张图像,你不知道里面是什么,完全不知道。但现在对其中一张图像,你做一个副本但加以修改,比如你把一张图像用两种不同方式裁剪,于是得到同一张图像的两个不同裁剪,然后还有另一张图像。现在,同一张图像的两个不同裁剪,即使你不知道里面是什么,你也知道它们包含相同的内容,而且很可能与另一张图像的内容不同。如果你随机下载两张图像,很可能另一张图像是别的东西。这就是信号的来源。然后你训练神经网络去理解:同一张图像的两个裁剪,神经网络应该知道它们是相同的。它不知道是什么,但应该把它们嵌入到某个空间中,让这两个靠得很近,而另一张图像应该被推得远远的。这个想法结果非常非常强大。Jeff 及其合作者,以及这里和其他地方的合作者,已经证明它在图像上效果很好,之后再进行图像识别训练,非常非常高效。你可以在基于图像的强化学习中做同样的事情。所以应用同样的想法:机器人现在看到的与它在不同时间看到的。现在看到的,你取两个不同的裁剪,那仍然是相同的东西,神经网络学到这是同一事物的两个视图,并且与另一事物不同。概念上非常简单,但非常非常强大。这样训练,你突然就能几乎像直接访问状态一样高效地从图像输入进行训练。我们在标准的 DeepMind 控制套件模拟机器人环境上评估了这一点。

An unsupervised task that has a lot of signal that's not directly the reinforcement signal but still signal to learn from for the neural network to then quickly acquire a new skill. So we've been working on that quite a bit in the last year, year and a half. And what we've seen—and it's not just us, also some people at NYU, Montreal, and so forth—we've actually been able to bridge the gap between what happens when you learn, let's say, a robot has to learn some skill, let's say running or maybe crawling and so forth, when the robot has to learn it with having full access to its entire configuration at all times—all its joint angles, its position, its orientation—that's full access, full state access. That always learned pretty fast. But it would learn slow if all it gets to see is a video stream of itself, because that's maybe a hundred by a hundred image, so 10,000 pixels. It's much higher dimensional than the succinct state description of the robot. And so what we saw is that massive gap in learning efficiency when learning from access to state, which is quite efficient, compared to learning with access to image input only. It was just this massive gap. And so the question is, how do we bridge this? And we did it with unsupervised learning effectively. So as the robot is doing its trial and error, instead of only paying attention to rewards and trying to become better based on what is a good or a bad run so far, it is also doing unsupervised learning on the images it's seeing. And what it means specifically in our case is we did something called contrastive learning. So in contrastive learning, what you do is you want to learn what's in an image, but how do you learn it if nobody tells you what's in an image? Well, here's the idea. Imagine you just download two images and you don't know what's in them, you have no idea what's in them. But now for one of the images, you make a duplicate but a modification as you duplicate it. So maybe you have an image and you crop it in two different ways. So now I have two different crops of the same image, and then there's the other one. Now, the two different crops of the same image, even though you don't know what's in it, you know they have the same thing in it, and it's different from what's in the other image most likely. If you randomly download two images, most likely the other image has something else. And that's really where the signal comes from. Now you train your neural network to understand that these two crops of that same first image—the neural network should know that's the same. It doesn't know what it is, but it should embed it somewhere in a space where it puts those two close together, and the other image should be put far away from it. That idea turns out really, really powerful. It's been shown to work really well by Jeff and collaborators, as well as others here and collaborators, on doing that on images followed by image recognition training. And it's very, very efficient. You can do the same thing in reinforcement learning from images. So you apply the same idea: what the robot is seeing now versus what it's seeing at a different time. Well, the thing is seeing now, you take two different crops and that is still the same thing, and the neural network learns that there's two views of the same thing and different from the other thing. Very simple idea conceptually, but very, very powerful. You train that way, all of a sudden you can train almost as efficiently from image inputs as with direct access to state. This one we evaluated on the standard DeepMind control suite, simulated robotics environments.

多任务学习与无监督辅助任务 Multitask learning and unsupervised auxiliary tasks

Host

我在想,你看到的其中一部分效果,是否只是给网络一些别的东西去学,而不是无监督任务对核心强化学习目标有所贡献。

I'm wondering if part of what you're seeing is the effect of just giving the network something else to learn, as opposed to the unsupervised task contributing to the core reinforcement learning thing that it's supposed to learn.

Pieter

你说得完全对,多任务通常有帮助。我的意思是,这就是为什么,比如在 Covariant,我们跨许多应用领域训练同一个神经网络,因为用同一个网络跨所有这些领域训练会有帮助,而不是为每个领域专门化。这里的研究也一样:当我们为多个任务训练一个神经网络时,它往往能做得更好。但引入无监督损失的美妙之处在于,另一个——你可以把无监督损失看作多任务,其中额外的多任务部分不需要你做任何标注,不需要你提供任何奖励信号。所以就人力而言,这是实现多任务最便宜的方式。

You're absolutely right in that multitask tends to help. I mean, it's why when, for example, at Covariant, we train the same neural network across many application domains, because it'll help to train across all those domains using the same network rather than specialize to each one of them. Same in the research here: when we train a neural network for multiple tasks, often it can do better. But the beauty about bringing in the unsupervised loss is that the other—you can think of the unsupervised loss as multitask where the multitask additional thing doesn't require you to do any annotation, doesn't require you to give any reward signals. So it's like the cheapest way to achieve multitask in terms of human effort.

Host

不过还是需要大量算力,只是说清楚。

Still a lot of compute required, just to be clear.

Pieter

是的,还是涉及大量算力。

Still, yeah, a lot of compute involved.

Host

我想我想到的类比是,在 NLP 中,当你训练语言模型之类的东西时,你对问题有一个无监督的表述。你在很多方面试图解决的问题直接与构建语言模型相关——你知道,把词遮住就是在教你语言。我在想,第二个任务——它恰好是一个无监督任务——与它和你要解决的核心强化学习任务之间的关系,是否存在区别。我不确定我把问题说清楚了,但想法是:无监督任务是核心的——它是否为核心强化学习过程和那个损失提供信息——还是它只是一个辅助的其他东西,而正是多任务方面或其他任务创造了价值或提升了性能?

I think the analogy that I was coming from is like in NLP when you're training something like a language model, and you have this unsupervised formulation of the problem. The problem that you're trying to solve in a lot of ways is directly related to building the language model—you know, blanking out the words is teaching you language. And I'm wondering if there's a distinction between the second task—it just happens to be an unsupervised task—and the relationship between that and the core reinforcement learning task that you're trying to solve. I'm not sure that I'm completely articulating the question in a clear way, but the idea is, is the unsupervised task kind of core to—does it inform the core reinforcement learning process and that loss—or is it just an ancillary other thing, and it's the multi-task aspect or the other task that is what creates value or causes you to increase performance?

Pieter

你实际上触及了我们正在研究的东西。让我多说一点。直接回答你的问题,我认为当我们对给定图像与不同时间的图像使用对比损失时,我们并没有把它与运动任务(比如爬行、坐下或推物体)直接联系起来。它与那关系并不密切。我想——我倾向于把它看作是在学习机器人的视觉系统。它是在学习看和理解:在这两种情况下看到的是相同的东西,而在另一种情况下看到的是别的东西。但它不是在学世界如何运作。这正是你所说的。而在强化学习中,你试图实现目标。你的机器人应该从 A 点到 B 点,或者用面前的物体实现某件事,或者在游戏中获得高分,等等。所以,实现目标和世界如何运作这整个概念,我认为是一个自然的补充。

So you're actually getting at something that we're working on right now. So let me say a bit more about that. To very directly answer your question, I think when we use the contrastive loss on a given image compared to an image at a different time, we're not tying it very directly to the locomotion task versus crawling versus maybe sit or maybe push an object. It's not very closely tied to that. I would think—I'm inclined to think of it as it's like learning the vision system of the robot. It's learning to see and understand that it's seeing the same thing in these two situations versus seeing something else there. But it's not learning about how the world works. And that's kind of what you're getting at. And in reinforcement learning, you try to achieve goals. Your robot's supposed to get from point A to point B or achieve something with the objects in front of it, or it could be in a game you're supposed to get a high score in the game, and so forth. And so this whole notion of achieving goals and how the world works, I think, is a natural complementary thing.

视觉与世界模型 Vision and World Models

Host

所以想象一下,在“不意外”的意义上,你已经能够补充你的强化学习,训练出一个视觉系统。现在你还需要一个“世界如何运作”的系统,这样当我作为一个机器人被要求在这个世界上做某事时,我已经大致知道世界是如何运作的,我不必以无数种方式乱挥手臂,而且我也不必学习“如果我不碰物体,它就不会移动”这类事情。

So imagine you in the unsurprised sense you've been able to complement your reinforcement to get a bit of a vision system trained now what you also want is effectively a how does the world work system such that when I'm asked to do something in this world as a robot I already kind of know how the world works I don't have to flail my arms a gazillion different ways and you know I don't have to learn that if I don't touch an object it's not going to move kind of thing.

Pieter

是的,我知道世界是通过接触力等运作的。所以如果我想把 A 块放在 B 块上,那么底部的 B 块必须先放在桌子上。而这些知识并不来自我刚才描述的对比学习,你需要一个时间维度。所以我认为这是重要的下一步之一。这需要更多的算力,这可能就是为什么它在整体研究进展中出现得稍晚一些。但那里的概念是,你能理解什么是自然的视频序列吗?如果你下载大量视频,它们通常看起来像什么?这是一个非常困难的问题,因为视频是高维的。我的意思是,训练神经网络进行视频预测或理解哪些视频彼此更相关或更不相关,是一个计算密集型的问题。所以这类问题,我相信你很清楚,Yann LeCun 已经谈论了很长时间。当他给出蛋糕类比时,强化学习是蛋糕上的樱桃,即奖励信号,但无监督学习是蛋糕的大部分,对吧?而 icing 是监督学习。无监督,大部分,他想到的是视频。至少在机器人技术的背景下,我认为视频是机器人观看了大量视频,从而知道世界如何运作。那里还有很多研究要做,但我认为那将是真正重要的第三部分。

Yeah, I know that the world works with contact forces and so forth. So if I want block A on block B, well block B that's on the bottom has to be on the table first. And then those things are things that don't come from what I just described, the contrastive learning I just described. You need a temporal aspect to it. And so that I think is one of the important next steps. It requires more compute, and that's probably why it's coming a bit later in the progression of research overall. But the notion there would be, can you understand what are natural video sequences? So if you download a lot of video, what does it tend to look like? And that's a very hard problem because video is very high dimensional. I mean, training neural networks for video prediction or understanding which videos are more related or less related to each other is computationally an intense problem. So that's the kind of problem that, I'm sure you're well aware, Jan LeCun's been talking about for a very long time. When he gives the cake analogy, reinforcement learning is the cherry on the cake, the reward signal, but the unsupervised is the most of the cake, right? And the icing is the supervised learning. Unsupervised, most of it, what he thinks of is video. In the context of robotics, at least, I think of video as robots that have watched so many videos that they know how the world works. A lot of research still has to be done there, but I think that'll be really important third part.

Pieter

所以理解你所看到的是第一部分,我们在这方面取得了很大进展。然后是视频理解,行为理解,即世界如何运作。第三部分是让机器人拥有自己的技能,这更接近你所说的,也是我们正在积极研究的,即能否让机器人独自度过时间?机器人独自在房间里或其他地方,能否让它有意义地、有效地度过时间,即玩耍?对于孩子,我们会称之为玩耍。你说,哦,孩子只是在玩,但实际上,是的,孩子只是在玩。你可能会说,如果它做些家务或什么的就好了,但实际上孩子只是在玩,它实际上在学习世界。孩子现在理解了,尤其是年幼的孩子,他们理解物体如何互动,他们通过互动理解世界如何运作。如果我们考虑机器人技术的长期未来,这种玩耍是另一个重要组成部分。如果我们想要减少监督,减少将一切训练到机器人中的需要,就让它们自己尝试。

So there's the understanding what you're looking at is part one, we made a lot of progress on that. Then video understanding, behavior understanding in that sense, how the world works. The third part is for the robot to have its own skills, and that's where it gets much closer to what you're talking about, and something we're actively working on, which is can you let the robot just spend time on its own? Just the robot's just on its own in a room or wherever it is, and can you make it spend its time meaningfully, effectively play? What for children we would call play. You say, oh kid is just playing, but actually yeah, kid is just playing. You might say, well, be nice if it did some chores or whatever, but actually the kid is just playing, it's actually learning about the world. It's the kid now understands, especially young kids, they understand how objects interact, they understand how the world works from interacting with it. And that kind of play is another important component if we think about the long-term future of robotics. If we want to get to less supervision, less need to train everything into the robots, just let them try things on their own.

Host

为什么这种玩耍会不同于我们今天在强化学习中做的目标导向探索?

Why would that kind of play be different from the goal-directed exploration that we do today in RL?

Pieter

所以,在强化学习中已经有大量工作基本上是在让这种玩耍浮现出来。目标导向探索就是一个很好的例子。还有其他类似的工作,但被称为好奇心或类似的东西。基本上,你为体验以前没有体验过的东西给予奖励。我认为在某种程度上封闭的环境中,这已经非常有效。我的意思是,如果你考虑一个 Atari 游戏,这是一个流行的强化学习基准环境,通常有一种正确的游戏方式,而且你能做的其他事情不多,否则你就会死。这就是我所说的封闭环境。如果你好奇并尝试体验新事物,而你在游戏中已经死了很多次,那就不再是新奇的了。所以新奇的事物是那些能让你在游戏中进入下一关的东西。因此,在新奇性和我们关心的实际任务之间有一个非常自然的对齐。这就是为什么好奇心、目标导向探索等取得了巨大成功。但在现实世界中,这行不通,因为你可以做太多事情,甚至在一些游戏如 Minecraft 中,它正成为一个新的基准,人们为此兴奋不已,正是因为这个原因。在 Minecraft 中,你可以建造许多不同的东西,所以仅仅体验新奇的事物,你可以永远这样做,却永远不会学到有趣的东西。我认为那是那里的下一步。也许这就是我认为儿童玩耍有点不同的地方,因为似乎他们对我们当前 AI 系统更有直觉,知道什么是有趣的玩耍,而不仅仅是新奇但实际并不那么有趣的东西。

So there is a good amount of work that's already happening in RL on essentially getting this kind of play to surface. So goal-directed exploration is a great example. There is other work that, you know, kind of similar, but it's called curiosity or something like that. Essentially you give rewards for experiencing something that you haven't experienced before. And I think that's worked really, really well in environments that are, I would say, somewhat closed. Meaning if you think about, let's say, an Atari game, which is a popular reinforcement learning benchmark environment, usually there's a right way to play the game, and there's not too many other things you can do, or you die. And that's what I mean with a kind of closed environment. If you're curious and you try to experience new things, while you've died many times in the game, that's not new anymore. So the new things are the things that take you to the next level in the game. And so there's a very natural alignment between novelty and the actual task that we care about. And that's why curiosity, goal-oriented exploration, and so forth have had a great amount of success. But that breaks down in the real world where there are so many things you could do, or even in some games like Minecraft, which is becoming a new benchmark that people get excited about for this exact reason. In Minecraft, you can build so many different things, and so just experiencing something novel, you can keep doing that forever and never learn something interesting. And I think that's kind of the next step there. And that's maybe where I think of things like children's play as being a bit different, because it seems like somehow they have a bit more intuition than our current AI systems about what's interesting play, as opposed to just what's novel but actually not that interesting.

预训练Transformer作为通用计算引擎 Pretrained Transformers as Universal Computation Engines

Host

有趣。我想换个话题,谈谈你最近在 arXiv 上发布的一篇论文《预训练 Transformer 作为通用计算引擎》。请给我们讲讲这项工作,以及你希望在那里实现什么。

Interesting. I want to switch gears to a paper that you recently posted up on arXiv, 'Pretrained Transformers as Universal Computation Engines'. Tell us a little bit about that work and what you're looking to achieve there.

Pieter

是的,这篇论文对我来说是我们做过的最令人惊讶的事情之一。通常,我觉得当我们写一篇论文时,我们有一个相当清晰的直觉,我们提前知道:好吧,这个直觉应该以这种方式在算法中利用,因此它应该有效,我们应该能够把事情提升到下一个水平。这是一个相当典型的研究进展。当然,这些迭代中,直觉可能是错误的,但随后你完善你的直觉并改进算法。但这对我来说有点令人惊讶,因为在这里我们并没有真正提出一个新算法,这真的是一项调查。所以我们做的是研究预训练语言模型。如今非常流行的是,在大量文本上训练一个 Transformer 模型,这是一种特定的神经网络架构,用于在文档中进行下一个词预测。如果你这样做,事实证明它在其他语言任务上也表现得非常好,这已经众所周知,OpenAI、Google、Facebook 等都有许多相关结果。但我们想知道的是,如果它在所有这些并非直接训练的任务上表现如此出色,那么它是否学到了更一般的东西?当它被训练在互联网上如此多的文本上预测下一个词,通过看到前面的文本预测接下来会发生什么,是否有一种更一般的推理机制被内化在这种神经网络内部,超越了语言?所以我们测试的方式是,我们说,好吧,让我们看看如果我们只训练一个语言模型,不训练其他任何东西……

Yeah, so this paper for me was one of the most surprising things we've done. Usually, I feel like when we write a paper, we have a pretty clear intuition that we know ahead of time: okay, this intuition should be leveraged this way in the algorithm, and as a consequence it should work, and we should be able to take things to the next level. That's a fairly typical research kind of progression. And sure, those iterations, intuition might be wrong, but then you refine your intuition and you improve the algorithm. But this is kind of surprising to me because here we didn't really come up with a new algorithm, it was really an investigation. So what we did is we looked at pre-trained language models. So what's very popular these days is to on massive amount of text train a transformer model, which is a specific architecture of a neural network, to do next-token prediction in a document. And if you do this, it turns out it works really well on other language tasks, and that had been known, and OpenAI, Google, Facebook as well, many many results around this. But the thing we were wondering is, well, if it's so good at doing all these tasks it wasn't really directly trained for, could there be something more general it has learned? When it's trained to predict the next token on so much text on the internet, by seeing the previous text predict what comes next, might there be a more general reasoning mechanism that has been internalized inside this kind of neural network beyond language? And so the way we tested this, we said, okay, let's see if we take a model trained on language only, not trained on anything else...

在非语言任务上测试语言模型 Testing language models on non-language tasks

Pieter

然后我们把它用在图像分类、蛋白质序列结合位点预测,或者像计算一串比特的异或这样的简单数学问题上。显然,这些都不是语言任务。你必须——你不能直接把语言模型放在前面,因为语言模型只接受语言输入。你不能给它一张图像,它不知道该怎么办。所以它们也不是像 GPT-3 那样的生成式任务,比如从文本生成网页,但那些任务都有一个共同的生成式、基于文本的特性。没错,它们倾向于生成你之前见过的东西,对吧?文本。所以这个模型完全是在文本上训练的。但我们相信,它有可能以更通用的方式进行推理,神经网络中有某种东西让它对输入中的对象进行推理,然后在输出中得出结论。也许那里有通用的推理模式,如果我们利用那个推理引擎,把它放在图像前面,它也能对图像中的内容以及这些内容如何相互作用进行推理,等等。

And then let's put it to use to now classify images or do a prediction about a protein sequence binding sites or a simple math kind of problem like compute the XOR of a sequence of bits. And none of these are language tasks, obviously. You have to do— you can't just put the language model in front of it because the language model only takes in language. You can't give it an image; it doesn't know what to do. So they're also not generative tasks like we've seen GPT-3 apply to lots of different areas, generating web pages from text, but they all share this common generative, text-based property. Exactly, they tend to generate something like the things you have seen before, right? Text. So this model is all trained on text. But we believed that there was a chance that it actually reasons in a more general way, that there's some kind of thing in the neural network that makes it reason about objects that are on its input and then draw conclusions on its output, and that maybe there are general reasoning patterns in there that if we use that reasoning engine and put it in front of an image, it'll also reason about what's in the image and how these things might interact, and so forth.

Host

当然,你要做一点阻抗匹配。我们拿到图像,必须做一个嵌入,就是一个线性嵌入层,这一层对神经网络来说几乎做不了什么工作,对吧?这正是我们的想法,我们希望预训练语言模型来完成所有工作。所以就是一个线性嵌入层,预训练语言模型,然后再一个线性输出层,因为这次我们不想输出文字,我们想输出一个决策:图像里有什么,或者这是 X 还是 0 还是 1,或者这里会不会有蛋白质结合。所以只是用单个线性层改变输入和输出,然后还有一件事我们必须做:我们必须处理 Transformer 网络内部的归一化,也就是层归一化。我们必须特别重新训练它,以确保它对于传入的数据处于正确的尺度。仅仅这样做就足以在这些其他任务上获得惊人的好性能。

Of course you do a little bit of impedance matching, so we take the image, we have to do an embedding, which is just a linear embedding layer, which is one layer that can do almost no work for neural networks, right? That's the whole idea, that we want that pre-trained language model to do all the work. So just a linear embedding layer, pre-trained language model, and then a linear output layer again, because we don't want to output words this time, we want to output a decision: what's in an image, or is this X or a zero or a one, or will there be a bind here for the protein or not. So just changing the output and input with just linear single layer, and then one other thing we had to do: we had to do the normalization that happens inside the transformer network, so there's layer normalization. We had to especially retrain that to make sure it's on the right scale for the data that comes through. Now just doing that was enough to get surprisingly good performance on these other tasks.

Pieter

所以对我来说,这非常令人惊讶,因为它证实了当你用语言训练这个神经网络时,它实际上根本不是专门针对语言的。如果你训练足够的语言,它实际上在内化一些更通用的推理模式。当然,我们还没有完全理解这一点,但我们有这个观察:它内化了某种非常普遍可重用的东西。当然,我们也做了测试:我们说,如果我们在中间放一个随机 Transformer 呢?和语言模型一样,但随机初始化,所以没有在语言上训练,同样的架构,对吧?随机初始化。它实际上做了一些令人惊讶的事情,它也有一定的能力,但不如预训练语言模型好。所以看起来,架构本身就有一些力量,在总体上惊人地强大,然后你在语言上学到的东西转移到这些其他领域,还有额外的力量。

And so to me that was very surprising, because it confirmed that when you train this neural network on language, it's actually not that specialized for language at all. If you train enough language, it's actually internalizing some more general reasoning pattern. Of course we don't fully understand this yet, but we have this observation here that it has internalized something that's very generally reusable. And of course we tested: we said what if we just put a random transformer in the middle, same as the language model but randomly initialized, so not trained on language, same architecture, right? Randomly initialized. And it actually does something which is kind of surprising, that it's also kind of capable of doing something, but not as well as the pre-trained language model. So it seems like there's both some power in the architecture being surprisingly powerful in general, and then there's additional power in what you learn on language that transfers over to these other domains.

Pieter

这又回到了我在最开始提到的事情。我们谈到了我对什么感到兴奋,我提到了这个观念——嗯,人脑是如此通用,似乎大脑中通常用于某种推理的部分可以用于其他事情。来自舌头的信号处理显然可以用于视觉处理。人们已经看到了这一点。还有盲人,通常用于视觉处理的大脑部分,对盲人来说可以用于音频处理。所以这种可重用性、通用性——我们正在看到,但和人脑完全不同,要说明白。记住,这完全没有被理解,而且比我们在这里做的任何事情都要先进得多。但是——你知道,我们试图朝着更通用的推理方向取得进展,而不是针对特定领域的专用目的。

Which goes back to something that I mentioned in the very beginning. We talked about what am I excited about, and I mentioned this notion of— well, the human brain is so general, and it seems like some part of the brain that's normally used for one kind of reasoning could be used for other things. Processing of signal from your tongue can apparently do visual processing. People have seen this. And blind people also, that what's normally part of brain used for visual processing, for blind people can be used for audio processing part of it. And so this kind of reusability, generality— we're kind of seeing, nothing like human brain, just to be clear. Remember, it's completely not understood and way more advanced than anything we're working on here. But it's— you know, tried to make progress in that direction of something that's more general reasoning, not special purpose to a specific domain.

Host

当你谈到随机模型和预训练模型,夹在这些线性层之间时,你是冻结 Transformer,然后以监督方式为特定问题微调或训练线性层吗?

When you talk about the random model and the pre-trained model, kind of sandwiched between these linear layers, are you freezing the transformers and then fine-tuning or training the linear layers in a supervised manner for the specific problem?

Pieter

就是这个想法,没错。所以一旦有了监督任务,你就冻结 Transformer,除了层归一化参数和线性输入层、线性输出层。这些会被重新训练。我认为这大约占总参数的 0.1% 左右,所以就是这样训练的。

That's the idea, correct. So once you have the supervised task, you freeze the transformer except for the layer norm parameters and the linear input and linear output layer. So those get retrained. I think it's about 0.1% or something of the overall parameters, so it's being trained that way.

Host

然后你说“惊人的好性能”,那是指在某些任务上达到最先进水平,还是我们惊讶于它居然能工作,但并不是特别有用或最先进的?

And then when you say surprisingly good performance, does that mean state-of-the-art on some task, or we're surprised that it worked at all but it's not particularly useful or state-of-the-art?

Pieter

它不是最先进的。我的意思是,在这种研究类型中,它是最先进的,我指的是这种研究类型,你不允许在你关心的任务上训练,或者训练很少,只有那些线性层。在这个意义上,是的,绝对是最先进的。但如果你说,我想要世界上最好的图像分类器,对吧?对。我是不是要先在语言上训练,然后只有输入和输出的线性层可以用?不,那还不是——还不是,或者也许永远不会,我不知道。给我们世界上最好的图像分类器。我只是想确保我没有对“惊人地好”的含义做出假设。

It wasn't state-of-the-art. I mean, it was state-of-the-art in terms of this kind of research, I mean, in terms of this kind of research where you're not allowed to train on the task you care about, or not much, only those linear layers. In that sense, yes, absolutely state-of-the-art. But in terms of if you say I want the best in world image classifier, right? Right. Am I going to first train on language and then only have a linear layer in the input and the output to work with? No, that's not— not yet, or maybe we'll never, I don't know. Give us the best in world image classifier. I just want to make sure I was not making assumptions on what surprisingly good meant.

Pieter

是的,所谓“惊人地好”,是指它居然——你知道,它比随机猜测好得多得多。是的。

Yeah, what's surprisingly good is surprising that it even— you know, that it does much, much better than chance. Yeah.

Host

太棒了。你认为这条特定的研究路线会走向何方?下一步是什么?

Awesome. And where do you see this particular line of research going? What are the next steps?

Pieter

所以我认为在多模态数据研究方面有很多机会。当然,我在这里提到的工作只是一个。另一个突出的工作,我相信你见过,是 OpenAI 的 CLIP 工作,他们专门同时训练语言和图像。所以它是同时在语言和图像上训练的,不只是在一个上训练然后也能在另一个上工作。但情况往往如此:你通常可以访问多种数据模态,它们在某种意义上不一定完全对齐,但你知道它们明显相关,你可以共同训练它们。

So I think there is a lot of opportunity in research on multimodal data. And of course, the work I mentioned here is one work. Another work that stands out, that I'm sure you've seen, is the CLIP work by OpenAI, where they specifically trained at the same time on language and images. So it was trained simultaneously on language and images, not just trained on one and then it also works on the other. But that's often the case: often you have access to multiple data modalities that are in some sense not perfectly aligned necessarily, but you know that are clearly related, and you can co-train them.

Pieter

我认为那里有很多机会。我认为有很多——我的意思是,甚至可以是同时训练的音频和视频。可以是文本和图像。也许可以——如果你想想机器人技术,同样的事情。可以是音频、视频,但也许有一天也可以是气味,如果我们有更好的人工嗅觉传感器,并将其与其他感知结合起来。我认为学习统一表征有很多机会,可能会走得更远。

I think there's a lot of opportunity there. I think there's a lot of— I mean, even this could be audio and video that could be trained on at the same time. This could be text and images. This could maybe— and if you think about robotics, the same things. It could be audio, video, but it could maybe also be scent someday, if we have better olfactory artificial sensors and combine that with other percepts. I think there's a lot of opportunity to learn unified representations that could probably go further.

结束语与播客推广 Closing Remarks and Podcast Promotion

Host

太好了,彼得,和你聊天非常愉快。感谢你慷慨地抽出时间,分享你最近在做的事情。非常酷。

Great, well Peter, it's been wonderful catching up with you. Appreciate you being so generous with your time and sharing a bit about what you've been up to. Very cool stuff.

Pieter

嗯,山姆,非常感谢。这非常有趣。实际上,在我们结束之前,我想我应该提前提一下,你已经加入了播客兄弟姐妹的行列,你知道的。我们不多谈这个,但你知道,你的播客是什么,你想做什么,人们应该去哪里找到它?

Well Sam, thank you so much. This was a lot of fun. And actually, I guess before we close out, I should mention this up front, but you've joined the podcaster brotherhood and sisterhood, you know. Let's not talk too much about it, but you know, what's your podcast, what are you trying to do, and where should folks go to find it?

Host

是的,几周前我开始了自己的播客,叫《机器人脑》(The Robot Brains),你可以在 Spotify、Apple 等平台找到它。就叫 The Robot Brains。我玩得很开心,遇到……我们真正关注的是那些试图弥合 AI 研究与将 AI 带入现实世界之间差距的嘉宾。这大致就是主题。我认为这是一个非常激动人心的时刻,看到 AI 在现实世界的许多地方转型,所以很多讨论都围绕这一点展开。但你知道,它也会延伸到其他话题,通常是 AI 研究、机器人研究和应用。

Yeah, so just a few weeks ago I started my own podcast called The Robot Brains, and you can find it on Spotify, Apple, and so forth. Just The Robot Brains. And having a lot of fun meeting... what we're really focused on is guests who try to bridge the gap between AI research and bringing AI into the real world. That's kind of the general theme. I think it's a really exciting time to see AI transition in many, many places in the real world, and so a lot of the discussions are centered around that. But you know, it also goes from there to other topics generally: AI research, robotics research, and applications.

Pieter

非常酷。假设它可以在 Spotify、Apple、Google 等所有常见平台上找到,对吧?只要搜索 The Robot Brains 就行。

Very cool. Assume it can be found on Spotify, Apple, Google, all the usual places, right? Just look for The Robot Brains.

Host

太棒了。我们会在节目说明中附上链接。再次感谢你,彼得。

Awesome. We'll link to it in the show notes. Thanks once again, Peter.

Pieter

谢谢你,山姆。非常感谢你邀请我。这非常有趣。谢谢。

Thank you, Sam. Thank you so much for having me. This was a lot of fun. Thanks.

互动版:逐字朗读 + 针对本期提问 →