From Reinforcement Learning to Gemini Robotics: The Evolution of Embodied AI
打开互动全文版(中英对照 + 朗读 + 问答)→Google DeepMind 机器人研究负责人 Karolina Parada 探讨机器人技术的戏剧性演变,从教机器人叠积木到最新的 Gemini Robotics 模型,将多模态理解带入物理世界。
Karolina Parada discusses the dramatic evolution of robotics at Google DeepMind, from teaching robots to stack blocks to the recent Gemini Robotics model that brings multimodal understanding to the physical world.
有位教授说过:“我敢打赌,如果能让机器人系鞋带,我就退休。”团队里的研究人员听了就说:“好嘞,我这就把系鞋带加进去。”你觉得我们接下来会看到这样的场景吗?就像我们看到大型语言模型的爆发一样,你认为接下来会是机器人技术的爆发吗?我们过去还在讨论这会不会在有生之年或职业生涯中发生,现在却在争论是五年还是十年。欢迎回到 Google DeepMind 播客,我是 Hannah Fry 教授。有时候,甚至经常,在日常对话中 AI 和机器人这两个词会被混用。比如人们会说在应用里跟机器人聊天。但机器人是有物理身体的。在 Google DeepMind,他们关心的是将 AI 嵌入现实世界的机器人。虽然 AI 取得了巨大进步,但具身智能却一直落后。不过,这一切可能即将改变。Carolina Parada 领导 Google DeepMind 的机器人研究团队,这个国际团队在机器人领域取得了一系列非凡进展,最近的是 Gemini Robotics,它将 Gemini 的多模态理解能力带到了物理世界。欢迎来到播客,Carolina。
There was a professor that said, "Oh, I bet if we can get robots to tie shoe laces, I will retire." And the researchers in the team were like, "Right on. I'm going to add that as a task." Do you think that's what we're about to see then? In the same way as we've seen the explosion of large language models, do you think the next thing is the explosion of robotics? We used to have discussions about whether it would happen in our lifetime or even in our careers and now we have debates about whether it would be five or 10 years. Welcome back to Google DeepMind the podcast. I'm Professor Hannah Fry. Sometimes, maybe even often, the terms AI and robot are used interchangeably in casual conversation. You know, people talking about chatting to a robot on an app. But robots have a physical body. And here at Google DeepMind, they care about robots with AI embedded in the real world. And while AI has made huge strides, embodied intelligence has lagged behind. But perhaps all of that is about to change. Carolina Parada leads robotics research here at Google DeepMind, the international team responsible for some extraordinary advances in robotics, most recently Gemini Robotics, which brings Gemini's multimodal understanding to the physical world. Welcome to the podcast, Carolina.
谢谢邀请。我知道你研究这些机器人已经很长时间了。你是如何看待它们的演变的?是的,这非常令人兴奋。我从 10 岁起就对机器人技术充满热情。因为我在卡通片里看到的,比如像 Rosie 那样的机器人帮忙做所有家务。作为孩子,你会想,当然,这就是我长大后想建造的东西。实际上,我在我的机器人团队里已经待了大约七年。特别是在过去三年里,情况发生了巨大变化。我们从一开始就相信 AI 将对机器人技术产生彻底的变革。我是说,现在有很多机器人确实很有用。有在生产线上的机器人,有在月球上导航的机器人,有在海洋里的机器人。但这些机器人都是被编程来执行特定任务的。它们对所处环境或可能遇到的物体做了很多假设,或者可能由人类远程操作。但我们从一开始就相信,AI 是改变机器人技术的方式,这样我们才能建造真正智能的机器人,能够与你互动,能够推理环境,并以一种非常通用的方式采取行动。所以,这从一开始就是我们的使命。我想三年前你的播客里讨论过机器人技术,那时我们还在做机器人强化学习。基本上,我们通过给机器人简单的奖励来教它们堆叠积木,比如如果塔变高了就加一分。我们取得了一些进展,但自那以后,由于我们处于 AI 的前沿,我们越来越多地将 AI 引入整个机器人领域。大约在 2022 年,我们引入了语言模型到机器人中。那是第一次你可以真正对机器人说话,比如“我渴了”,而它会明白你的意思。后来,我们引入了视觉语言模型,这样机器人不仅能理解自然语言,还能理解它接收到的视觉输入,并据此做出决策。然后在 2023 年,我们引入了机器人 Transformer。这是 Transformer 架构首次被应用到机器人领域。它基本上向我们展示了机器人性能会随着数据扩展,这开启了一个大规模数据驱动机器人学习的新基础或新时代。最近,我们推出了 Gemini Robotics,这基本上是我们最先进的行动模型,它利用 Gemini 的多模态世界理解,通过将行动作为 Gemini 的一个新模态,将其带入物理世界。这使模型变得非常通用,因为它通过 Gemini 的理解来理解世界,并使其具有交互性。事实上,它能理解 Gemini 支持的任何语言,并使其变得灵巧。所以它可以在与你交谈的同时执行非常复杂的操作,并理解全新的情况,这在今天对机器人来说实际上非常困难。
Thanks for having me. Now, I know you've been working with these robots for quite a long time. How have you seen them evolve? Yeah, it's been super exciting. I have been excited about robotics since I was 10 years old. Super excited because of what I've seen in cartoons. Like you see robots like Rosie the robot helping do all the chores. And as a kid, you're like, of course, that's what I want to build when I grow up. And really, I've been at the gold in my robotics team for about seven years. And really things have changed dramatically in the last 3 years in particular. We've always believed from the very beginning that AI was going to be completely transformative to robotics. I mean there's a lot of robots out there that are really helpful today. There's robots in manufacturing lines. There is robots that are navigating the moon. There is robots that are in our oceans. But these robots have been programmed to do specifically those tasks. They make a lot of assumptions about those environments or the objects they might encounter and or they might be remotely operated by humans. But we have believed from the beginning that AI is the way to transform robotics so that we can build robots that are truly intelligent so that they can interact with you that they can reason about their environment and they can take action in a way that feels very general. So that has been our mission from the start. And so I think three years ago you had robotics in your podcast and back then we were doing reinforcement learning for robotics. And so essentially we were teaching robots to like stack blocks by giving them a simple reward like a plus one if you tower got taller. And we made some progress there but a lot since we've been at the forefront of AI we've been bringing more and more of AI into the entire world of robotics. So about 2022, we introduced, for example, LMS to robots. And that was the first time that you could actually talk to a robot and say something like, "I'm thirsty." And it would know what you meant. And then later on, we brought VLM so the robot could understand natural language, but it could also understand the visual input that it was getting and then make decisions based on that. And then in 2023, we introduced robotics transformers. And this was the first time that the transformer architecture was actually included in robotics. And it basically showed us that if that robot performance scales with data and that essentially started a new foundation or a new era of large scale data-driven robot learning and then more recently we introduced just now Gemini robotics which was is essentially our most advanced model for actions and it essentially takes the multimodal world understanding of Gemini and brings it to the physical world by adding actions as a new modality in Gemini. And that really enables models to be very general because it's understanding the world through Gemini's understanding and enables it to be interactive. In fact, it can understand any language that Gemini supports and enable it to be dextrous. So it can still do very complex manipulation while talking to you and also understanding a completely new situation which today is actually very hard for robots to do.
就你的大目标、大抱负而言,我们如何知道何时达成?
In terms of your big goal, your big ambition, how will we know when we get there?
是的,我认为这肯定会是一个渐进的过程,机器人能够理解新情况,并推理出它们需要做的、从未见过的事情。这正是我们现在看到的。但学习越来越复杂的任务对它们来说仍然困难。事实上,这就是我们观察到的。机器人感觉有点像两岁的幼儿,能理解周围的世界,可以开始玩物体,理解概念,但如果你教它做更复杂的事情,比如我们有一个例子是教机器人折纸,它实际上需要时间来练习。一旦它有了更多练习,它就能做到。所以这大致就是我们今天的水平,但如果我们希望机器人进入日常空间为我们做各种任务,那还差得很远。所以还有很长的路要走。
Yeah, I think it's definitely going to be gradual where robots are able to understand a new situation and reason about something they need to do that they haven't seen before. And that's exactly what we're seeing right now. But it's still going to be difficult for them to learn more and more complex task. In fact, that's what we see. The robot can feels sort of like a 2-year-old toddler that can understand it world around it. It can start to play with objects. It understands concepts, but if you teach it to do something more complex, like we have an example where we're teaching the robot to do an origami fold, it actually needs time to practice that. And once it has more practice in that case, it can actually do it. So that's roughly where we are today, but that's far from where we need to be if we want robots to be in everyday spaces doing all kinds of tasks for us. So there's still quite a bit to go.
我想我们可以看看这些机器人能做什么,因为你们最近发布了一个视频。这里我们看到一个人形机器人正在为它的人类打包午餐,同时还在玩井字棋。它井字棋玩得好吗?我想我们还能赢它,因为它的理解很简单。不过看看它在做什么:它正在轻松地捡起棋子并移动它们。还有一段是它能根据出现的字母块拼出自己的字谜。你特别印象深刻的是什么?
I thought that what we could do is take a little look at some of what these robots can do because there's a video that you guys have recently released. What we have here then is we have a humanoid robot who is packing a lunch for its human also playing noughts and crosses. Is it any good at noughts and crosses? I think we still beat it because it's very simple understanding. Tell you what it's doing though. It's picking up the pieces and moving them around quite easily. There's also a bit here where it can make its own anagram based on tiles that appear. What were you particularly impressed by?
我认为这些模型最令人兴奋的地方在于,很多时候我们自己的研究人员都对它的表现感到兴奋和印象深刻。这主要是因为我们的测试方式是把机器人放在它从未见过的情况面前。所以连我们也不知道机器人是否能做对,而在很多情况下它确实做到了。所以我们在视频中展示的许多例子,以及其他有两只手臂移动的视频,都表明它实际上在理解一个复杂的概念。一个非常酷的例子让我们都倒吸一口气的是,当我们展示机器人实际上在扣篮的视频时。
I think that's what's most exciting about these models is that in many occasions our own researchers were excited and impressed by what it was doing. And it was primarily because the way we were testing it was by putting the robot in front of situations that it's never seen before. So even us didn't know whether the robot was going to be able to get it right and in many occasions it did. So many of the examples that we show in this video as well as the other videos where you have the two arms moving around is that it's actually understanding a complex concept. So, a really cool example that where we were all like gasp was when we showed the video where the robot is actually doing a slam dunk.
那个案例很酷的一点是,那天我们让创意团队来拍摄机器人,并让他们带些玩具。我们没多说别的,就说“带些玩具来和机器人玩”,而且都是机器人没见过的东西。他们完全不知道机器人训练过什么,对吧?所以他们带了一个小篮球架,是个可爱的玩具,配个小球,放在机器人面前。机器人从没见过任何与篮球相关的东西,更别说这个玩具了。他们让它来个扣篮。我们都在想:“不知道能不能行。”结果不到四分之一秒,它就把球放进了篮筐。我们都觉得太神奇了。它其实就是利用了 Gemini 对篮球和扣篮的理解,对吧?这个概念我们根本没想到要去教它。它准确做出了动作。这真是个很酷的例子。
And what was cool about that case is that that day we were just having the creative team come and film the robots and we asked them to bring toys. We didn't say anything else. They're like, "Just bring toys to play with the robot." And things the robot hadn't seen before. All things. Yeah. They had no idea what the robot was trained on, right? So, they actually brought this little basketball hoop that was a little cute toy with a little ball and they put it in front of the robot. Again, the robot had never seen anything related to basketball. It certainly has never seen this toy. And they asked it to do a slam dunk of the ball. And we were all like, "I have no idea if it will work." And actually it took not even a quarter of a second and it actually decided to put the ball inside the basketball hoop. And we were all like that's amazing. And it was just essentially drawing from Gemini's understanding of what basketball is and what a slam dunk is, right? Which is a concept that you we couldn't have thought of teaching it to do. Yeah. And it essentially did the right motion. So that was a really cool example.
跟我聊聊装午餐的例子。它似乎对香蕉有概念性理解,比如它知道怎么抓香蕉吗?因为抓香蕉和抓陶罐或更脆弱的东西方式不同。
Talk to me about the packing lunch one. It kind of has a conceptual understanding of what a banana is, for example. Does it know how to grip a banana in the sense that you can't grip a banana in quite the same way as you could a clay pot or something even more fragile than a banana?
实际上,非常令人印象深刻的一点是,这些机器人极其简单。它们没有触觉传感、深度传感或力传感。所以它们完全靠手眼协调,并利用对如何抓取香蕉的理解。它看着物体然后抓取,一旦看到自己抓住了,就知道检测到了。其他机器人可能更复杂,但这迫使模型真正推理所见之物,并决定如何拾取。这正是这里的原创之处。
Actually, one of the things that is super impressive is that these robots are extremely simple. They actually don't have touch sensing. They don't have depth sensing. They don't have force sensing. So, they're literally doing eye hand coordination and using an understanding of how you grasp a banana. So, it actually is looking at the object and grasping it. And once it sees that it has it in hand, that's how it knows that it has detected it. There's other objects, there's other robots out there that are much more complex, but this forces the model to really reason about what it's seeing and making a decision about how to pick that up. And that's the thing that's really original here.
是的,这其中的一点是,它这么做不仅仅是因为我们教了它一千遍怎么捡香蕉,而是因为它从 Gemini 中提取了对如何抓取物体的理解,然后将其适应到动作世界。网上多年来有很多视频,展示非常令人印象深刻的机器人做后空翻、被踢倒、在山间奔跑等。相比之下,捡起香蕉放进午餐盒似乎很简单。但我们在讨论的是不同类型的机器人,对吧?
Yeah, that is one of the many things is the fact that it's doing it not just because we taught it a thousand times how to pick up a banana. It's because he's pulling this out of his understanding of how to pick up objects from Gemini and then adapting it to the world of actions. I mean, there've been lots of videos doing the rounds on the internet for a number of years of extremely impressive looking robots doing back flips and I don't know, being kicked over and sort of running up and down mountains and things. In comparison to those videos, picking up and putting down a banana, you know, into a lunchbox seems like quite a simple task. But we're talking about a different type of robot here, aren't we?
是的。我的意思是,你要解决的是完全不同的问题。那些视频中的机器人大多是排练好的序列,它们学习并记住了动作,我们确实印象深刻。但你要解决的是不同的问题。这里你要让机器人推理,面对眼前的物体,打包午餐意味着什么。它需要做什么才能把面包放进袋子里,然后封口,而且过程永远不会如你所料,因为这些都是非常柔软、会移动的东西。所以它需要根据实际情况反应并完成任务。这就是通用性的概念。
Yeah. I mean, this is a completely different problem you're trying to solve. Many of those videos are basically rehearsed sequences that the robot has learned and memorized and it has is we're actually very impressed by them. But it's a different problem that you're trying to solve. What you're trying to solve here is for the robot to reason and what it mean about what it means to pack a lunch given the objects in front. What it needs to do in order to put a piece of bread inside of a bag and then what it means to close it and it's never going to go as you expect because these are very flexible things that move around. So it needs to react and respond to what's happening and then actually complete the task. It's that idea of generality.
没错。那你如何决定或比较一个机器人和另一个?如何判断这个机器人在通用性上比另一个更好?
That's right. Yeah. How do you decide or how do you compare one robot against another? How do you decide whether this robot is doing generality better than another?
这其实是我们录制这次发布的演示时很难表达的一点。演示本质上就是预先编排的。所以我们觉得这不能完全展现我们想分享的东西。因此我们让团队带一堆玩具,和机器人互动,看看会发生什么。最好的捕捉方式是我们能通过对话改变机器人的行为,你在视频中也能看到。我们可以放它从未见过的物体,移动物体,确保人们明白这不是预设行为。实际上,在我们的基准测试中,我们从各种角度评估模型的泛化能力。我们会改变视觉背景,使用新物体,添加干扰物,甚至让它做全新的事情,或者用不同语言下达指令。比如我用西班牙语下指令,它也能执行。
That was actually one of the things that was hard for us to express when we were even recording for the demos in this release. A demo is by definition pre like scripted. So we were like this doesn't quite capture what we want to share. So what we did that's why we asked the team to bring a bunch of toys and actually start playing with the robots and see what emerges. And the best way to capture it is that we're able to change the behavior of the robot by talking to it. And you can see that in the videos. We're actually able to put row objects that it's never seen before. And we move object around to make sure that people understand that this is actually not a prescripted behavior. In fact, in our benchmarks, we evaluate our models in all kinds of ways in terms of generalization. So we will change the visual background. We will change the background. The objects were new. We will add objects to distract the robot. We would also like ask it to do completely new things or even you can talk to it in a different language. So, I could just give it the instruction in Spanish and it would just actually work.
我还想谈谈交互性,因为你们的一些视频中,有一个人坐在桌前,机器人跟着收拾东西;另一个视频中,人移动杯子,机器人追着试图把物体放进去。这些交互场景比静态任务难多少?
I wanted to talk about interactivity too because in a few of your videos there's one where a human is sat at a desk and the robot is kind of clearing up after him as he goes. In another you've got a human moving a cup around and the robot sort of chasing it trying to put an object inside. How much more difficult are those interactive scenarios than just a static task?
这些行为要高级得多,而且很多交互性其实是模型自然涌现的。比如我们并没有刻意考虑物体移动多快机器人才能反应。我们当然希望模型能快速反应,但视频中的很多例子都是人们和模型玩耍时自然出现的。整理桌子的例子也是,有人和机器人玩,看它能坚持多久直到完成任务。所以,看到 Gemini 中已有的许多能力在机器人身上变得极其宝贵,机器人能根据你的话语调整,真的很神奇。你可以和它完整对话,在它移动时改变行为。比如你说“我想让你做这个”,然后“哦不,算了,做那个”,它就会照做。有点滑稽。你还可以改变物体,它也会照做。有时候我觉得这些机器人没有感情是件好事,因为被研究人员在桌子上追着跑感觉挺逗的。
The significantly more advanced behavior and a lot of the interactivity sort of just fell out of the model. Like we were not thinking, for example, how fast can we move these objects before the robot would react. We certainly knew that we wanted a model that could react quickly, but a lot of these examples that we posted on videos just fell out of people playing with the model and seeing how it would behave. Same with organizing the desk. That was actually someone playing with the robot deciding to see how much it could game it until it actually was able to complete the full task. So yeah, it actually is amazing to see how a lot of these other capabilities that are already there in Gemini are actually extremely valuable when you bring them into a robot which is now able to adapt based on what you're saying. So you could actually have a full conversation and change the behavior of the robot as he's moving. So you can say, "I want you to do this." "Oh, no, actually never mind. I want you to do this other thing." And it would actually just follow you. It's actually kind of comical. And then you could also change the objects around and it will just do it. I think it's kind of a good job sometimes that these robots don't have feelings because I feel very sort of full like just being chased around on a table by researchers. is super fun.
这是底层的大语言模型在帮助它,对吧?给了它操作物体的概念性理解。
That's the large language model sitting underneath it that's helping it do that, right? That's giving it that conceptual understanding of the objects that it's manipulating.
我们利用 Gemini 的多模态理解能力,接收机器人通过摄像头看到的视觉输入以及从人类那里听到的自然语言,然后将其转化为行动指令。而且它实际上还能回应。所以你可以问它是否完成了,或者问它在折叠折纸的过程中进展如何。它确实能理解并做出回应。
We're leveraging Gemini's multimodal understanding to take the visual input of what the robot is seeing through its cameras and the natural language it is hearing from the human, and then translate that into how to act. And it actually also speaks back. So you can ask it a question about whether it's done. You can ask it a question about how far it is in the process of folding an origami figure. It actually understands and can respond.
我记得 Gemini 刚推出时,人们不遗余力地强调它的多模态特性。这是主要原因之一吗?这是为打下所有额外基础、确保它能理解视频和图片等所获得的回报吗?
I remember when Gemini was first being launched, people went to great lengths to talk about how it was multimodal. Is this one of the main reasons? Is this the payoff for putting in all that extra groundwork and making sure it can understand videos and photos?
这只是众多原因之一。我认为我们人类通过多种感官来感知世界,对吧?所以如果你想构建一个与我们大脑一样强大的智能,能够以多模态方式接收输入就非常重要。而机器人技术就是一个完美的例子,它绝对需要理解自然语言和视觉输入,未来可能还需要触觉感知,才能像人类一样做出行动决策。
I mean, one of many. I think us humans capture the world through many different senses, right? So I think it's super important if you want to build an intelligence as powerful as our brains to be able to take input in a multimodal way. And definitely robotics is a perfect example where you can see that it absolutely requires understanding of natural language and visual input, and presumably in the future also touch sensing, in order to make decisions about how to act the same way humans do.
为什么机器人需要对其行为有概念性理解?我的意思是,好吧,也许你不会称它们为智能,但有些机器人,比如洗碗机或割草机,它们并不理解盘子或草是什么。这真的有必要吗?
Why does it matter that robots should have a conceptual understanding of what they're doing? I mean, okay, maybe you wouldn't call them intelligent, but there are robots like dishwashers or lawnmowers that don't have a conceptual understanding of what a plate is or what grass is. Is it actually necessary?
我确信有些应用场景下,机器人只需重复动作就足够了。但我们感兴趣的是构建能够以非常通用的方式进行推理和行动的机器人,因为世界真的很混乱。事情永远不会完全按计划进行,而且很多任务中情况在不断变化。这实际上为这些机器人开辟了应用机会,它们几乎可以出现在人类能执行任务的任何地方。因此,这使它们能够在家庭环境以及制造环境中提供帮助。
I'm sure there are applications where you can have a robot that just repeats actions and it would be just fine. But we're interested in actually building robots that can reason and act in a very general way, just because the world is really messy. Things will never go exactly according to plan, and there are a lot of tasks where things are constantly changing. It actually opens up the opportunity for applications where these robots could literally be anywhere a human could be doing a task. So that enables them to be helpful in home environments but also in manufacturing environments.
在机器人技术中,有些重要的事情我认为现在用标准的 Gemini 就很容易做到,比如指向或绘制边界框。请给我们解释一下这些是什么。
There are some things that are important in robotics that I think now with the standard Gemini are quite easy, like pointing or drawing bounding boxes. Just explain to us what those are.
基本上,这是我们为了帮助机器人技术而必须改进 Gemini 的领域之一。如果你面前有一个物体,我们所说的指向是指我可以精确地识别该物体上的任何点。所以,假设你面前有一件 T 恤。如果我指向领口,它应该说‘这是领口’。或者如果我说‘领口’,它应该能识别出领口的位置。你可能觉得这没那么重要,但实际上,如果你要折叠那件 T 恤,你需要知道领口在哪里、T 恤的底部在哪里,以及所有不同的部分。边界框意味着你可以识别出该物体的所有边缘,从而知道物体在哪里结束、环境的其余部分从哪里开始。这类例子对我们人类来说微不足道,我们甚至不会去想它。但如果机器人能够获取这类信息,它们就能在物理世界中更智能地采取行动。这基本上就是我们所说的具身推理。
Well, basically this is one of the areas that we actually had to improve Gemini in order to help with robotics. If you have an object in front of you, what we mean by pointing is that I can literally identify any point in that object. So I can say, imagine you have a t-shirt in front of you. If I point to the collar, it should say 'this is the collar.' Or if I say 'collar,' it should identify where the collar is. And you might imagine that this is not that important, but actually if you're trying to fold that t-shirt, you need to know where the collar is, where the bottom of the t-shirt is, and all the different components. Bounding boxes mean that you can identify all the edges of that object so that you know where the object ends and the rest of the environment begins. These kinds of examples are trivial for us humans. We don't even think about it. But if robots are able to have access to that kind of information, then they can be smarter about the way they take action in the physical world. This is what we call embodied reasoning essentially.
这与标准 Gemini 模型中的推理有何不同?
How is it different from the kind of reasoning that you get in the standard Gemini model?
我们将具身推理定义为像人类一样更详细地推理物理世界。如果你要采取行动,比如给孩子打包午餐,要做到这一点,你必须理解所有物体在 3D 空间中的位置。然后你需要理解如何抓取每个物体以便将其放入盒子中。接着你需要弄清楚如何组织所有物品以便它们能装下。所有这些都是我们所说的具身推理。
We refer to embodied reasoning as reasoning about the physical world in a lot more detail the way humans do. If you're going to take action, say that you're trying to pack a lunch for your kid, in order to do that, you have to understand where all the objects are in 3D space. Then you need to understand how to grasp each object in order to pack it into that box. And then you need to figure out how to organize all those pieces so that they fit. All of this is what we mean by embodied reasoning.
所以这是不是像这样:假设你有两个摄像头视角。比如你在那里,我在这里。我能看到你的麦克风,你也能看到,但我们的视角完全不同。是这类事情吗?
So is this things like, let's say you've got two camera views. You're there and I'm here for instance. I can see your microphone and so can you, but we've got a completely different view of it. Is it that kind of stuff?
是的。我的意思是,它能够理解麦克风离我们的脸有多远,而且如果我移动,它还能进行物体对应,也就是说它理解这个麦克风就是我从另一个视角看到的同一个麦克风。你可以想象,如果机器人在移动并推理其环境,这一点非常重要。
Yeah. I mean, it can understand for example how far the microphone is from our face, but also if I move around, it can do object correspondence, meaning it understands that the microphone is the same one that I'm seeing from the other point of view, which you can imagine is super important if a robot is moving and reasoning about its environment.
从单个摄像头视角这样的 2D 图像切换到对空间的 3D 理解有多难?
How hard is it to switch from a 2D image like a single camera view to a 3D understanding of the space?
实际上,如今机器人的做法是从不同位置获取摄像头视角。机器人的手腕上有摄像头,顶部也有摄像头,它实际上是从三个图像中获取所有输入,并自行处理。它实际上在推理:‘哦,我离物体更近了,因为现在这个摄像头看起来更近。这个摄像头我能看到我的手。’它自行完成所有这些关联。所以我们并没有明确地将深度作为额外输入。我们只是给它多个摄像头视角,它自己就意识到如何利用它们来理解深度。
So actually, today what robots are doing is that they're taking camera views from different places. The robot has cameras in its wrist and has a camera on top, and it's actually taking all the inputs from the three images and doing this on its own. It is actually reasoning, 'Oh, I'm closer to the object because now this camera looks closer. This camera I can see my hand,' and it's doing all of that association on its own. So we're not explicitly adding depth as an additional input. We're just giving it multiple camera views and it's realizing how to use them in order to understand depth.
这其中有多少是你特意为机器人设定的任务,又有多少是从 Gemini 模型中获得的概念性理解中涌现出来的?
And how much of that was you deliberately setting that as a task for the robots, or how much of it sort of emerged from the conceptual understanding that you get from the Gemini models?
实际上它就这么涌现出来了。我们能够给它多个摄像头,然后看看它是否真的能在它们之间进行推理。
It simply emerged actually. We were able to give it multiple cameras and just see if we actually could reason between them.
这一定相当令人震惊。很多人肯定花了很多年时间苦苦思考如何对齐不同的摄像头视角,以便在不同角度下跟踪物体,然后突然之间,你有了像 Gemini 这样的大型语言模型,它就能自动做到这一点。
That's got to be quite shocking. Many people must have spent many years thinking very hard about that problem of how do you align different camera views to make it so that you're tracking an object across different angles, and then all of a sudden you get these large language models like Gemini and it can just do it automatically.
是的。我的意思是,能够利用这些模型为系统带来简洁性,这真的很棒。你真的不需要所有这些不同的阶段:先跟踪深度,然后才提取物体的位置,然后才规划如何移动,然后才能执行任务。这是因为基础模型实际上就像一把瑞士军刀,它可以同时完成所有事情。
Yeah. I mean, it's actually wonderful to be able to leverage these models to bring simplicity to the system. You really don't need to have all these different stages where you track depth, then only then you extract where the objects are, and only then you plan how to move, and only then you're able to do the task. That's because the foundational model is effectively like a Swiss Army knife; it can do all the things simultaneously.
是的,完全正确。而且它能在它们之间进行推理。好的。所以你几乎增强了物理推理能力。
Yes. Exactly. And it can reason between them. Right. Okay. So you enhance the physical reasoning almost.
正是如此。
Exactly.
你增强了物理推理和空间理解,然后运动理解就是下一步:理解如果我把一个杯子放在桌子边缘,实际会发生什么。所有这些领域都是我们增强的方面。但这还不够。你实际上必须再进一步,开始教 Gemini 动作的语言。对我们来说,动作意味着理解你实际上如何移动机器人的每个关节。所以如果这是我的机器人手臂,那么我在教 Gemini 如何移动机器人,如何像这样移动我的手臂。这些本质上都是数字,对吧?它正在学习翻译拿起一个杯子意味着什么,与移动我的手臂来拿起杯子之间的区别。所以你基本上是在教它一门新语言。你在连接那些不同的想法。
You enhance physical reasoning and spatial understanding, and then motion understanding would be the next thing: understanding what would happen if I put a glass at the edge of the table, what's actually likely to happen. All of these areas are the areas that we enhance. But that's not enough. You actually have to take it another step and essentially start to teach Gemini the language of actions. And actions for us means understanding how you are actually moving each joint in a robot. So if this is my robot arm, then I'm teaching Gemini how to move the robot, how to move my arm like this. And these are all essentially numbers, right? And it's learning to translate what it means to pick up a glass versus move my arm in order to pick up a glass. So you're essentially teaching it a new language. You're connecting those different ideas.
没错。那么我们可以把这看作是两套系统协同工作吗?我的意思是,我在想这里用系统一和系统二的类比,就是丹尼尔·卡尼曼的《思考,快与慢》那套东西。
Exactly. Can we think of this as two systems working in tandem then? I mean, I'm thinking of the analogy here of system one and system two, the Daniel Kahneman thinking fast and slow thing.
是的,没错。所以我们构建的模型实际上有两个模型。它有一个系统很慢,但在推理和思考方面非常强大,还有一个系统更快,但非常擅长反应。这就是慢思考和快思考的概念。这就像人类大脑的工作方式,对吧?你大脑的一部分非常擅长计算和分析,然后你也有非常本能的反应的一面。
Yes, exactly. So essentially the model that we built actually has two models. It has a system that is slow but very powerful at reasoning and thinking, and a system that is faster but very, very good at reactivity. So this is the concept of slow and fast thinking. This is like how human brains work, right? You have the part of your brain that's very good at calculation and analysis, and then you also have your very instinctive reactive side too.
是的,没错。事实上,我们现在所做的就是,其中一个模型比另一个大得多,你可以想象,它实际上存在于服务器上,而快速模型存在于设备上,可以非常快速地响应。那么请给我讲讲这在系统一和系统二方面是如何工作的,以及那个扣篮的例子,它以前从未见过。它是如何工作的?
Yes, that's right. In fact, what we do today is that one of these models is much bigger than the other, as you can imagine, and it actually lives on the server, and the fast model lives on device and it can respond very quickly. Talk me through how this works then in terms of the system one and system two and that example of a slam dunk, something it's never seen before. How does it work?
当你让机器人拿起篮球并扣篮时,系统二必须理解这意味着什么。你知道,什么是篮球?它必须理解它面前的物体在哪里,比如篮球在哪里,理解有一个篮筐,然后扣篮实际上意味着拿起那个球并把它放进去。所以它理解所有这些,并预测机器人应该如何移动的大致轨迹,然后将其交给系统一,系统一在设备上,能够接收那个轨迹,但它也接收视觉输入,并能够调整那个轨迹。所以如果我要,例如,挡在中间,把手放在中间,或者移动物体,它仍然能够响应,因为它已经理解了扣篮的概念,并且非常快速地响应。
What happens is when you ask the robot to take the basketball and do a slam dunk, the system two has to understand what that means. You know, what is basketball? It has to understand where the objects that are in front of it are, like where the basketball is, understand there's a hoop, and then that a slam dunk actually means picking up that ball and putting it there. So it understands all of that and predicts a rough trajectory of what the robot should do in terms of how it should move, and then hands that over to the system one, which is on the device and is able to take that trajectory but it also takes the visual input and is able to adjust that trajectory. So if I were to, for example, get in the way, put my hand in the middle, or move the object around, it would still be able to respond because it already understood the concept of what a slam dunk was and respond very quickly.
但为什么需要两套系统呢?比如,为什么不能只用那个慢而聪明的系统?
Why do you need two systems at all, though? Like, why can't you just use the slow clever one?
我们实际上可以只用那个慢而聪明的系统,但那样它会在视觉上明显更慢,并且无法快速适应环境的变化。这一点很重要,尤其是当你在做物体移动的事情时。所以如果你有,例如,想象一下你在空中折叠一件 T 恤,人类做得很巧妙,你实际上在移动这件 T 恤,东西以你无法预测的方式移动。所以你需要能够快速响应才能完成任务。所以你绝对需要一个快速系统,而慢速系统只是让我们能够进行更复杂的推理。所以如果你只做不需要高级推理的任务,你也可以只用一个小的系统。
We could actually just use the slow, clever one, but then it would actually be significantly visually slower and it won't adapt as quickly to changes in its environment. And that's important especially if you're doing something where the objects will move around. So if you have, for example, imagine when you're folding a t-shirt in the air, which as humans do pretty cleverly, you're actually moving this t-shirt and things are moving for you in ways that you don't predict. So you need to be able to respond quickly in order to actually complete the task. So you need definitely a fast system, and the slow system simply enables us to do much more complex reasoning. So you could also just live with a small system if you could do tasks that don't require advanced reasoning.
这是对人类大脑工作方式的直接复制吗?我的意思是,你知道,丹尼尔·卡尼曼的研究可以追溯到 1970 年代左右,对吧?我们已经理解人类大脑就是这样工作的。是直接复制吗?
Was it a direct copy of how things work in the human brain? I mean, you know, the Daniel Kahneman work comes back to 1970s or so, right? That we've understood that that's how the human brain works. Was it a direct copy?
不,完全不是。我认为我们肯定是从慢速系统开始的,就像你说的,为什么不用一个模型来解决这个问题,然后我们发现,实际上如果你想做高度灵巧的行为,进行复杂或任何类型的复杂操作,你需要快速响应,那是我们能找到的最佳组合。
No, not at all. I think we started definitely with the slow system, as you said, why don't we just solve this with one model, and we found that actually if you want to do highly dexterous behaviors with complex or any kind of complex manipulation, you need to respond quickly, and that was the best combination that we could find.
哇。这几乎就像进化是一个非常好的优化器,为快速但聪明的事情找到了非常好的策略。是的,确实如此。我们很惊讶那个组合竟然有效。我确实认为,有时候人体在大脑之前就知道一些事情,就像那样。你知道,就像你可以不假思索地接住一个掉落的杯子,或者你可以把东西变成肌肉记忆,比如弹钢琴,你实际上可以完全想别的事情。你在机器人身上看到类似的事情了吗?它们几乎有一种与慢速聪明系统分离的物理智能?
Wow. It's almost like evolution is a really good optimizer and finds really good strategies for like quick but clever things. Yes, definitely. It was surprising to us that that combination just worked. I do think that there's sometimes where the human body knows stuff before your brain does, as it were. You know, like you can catch a falling glass without thinking, or you can commit things to muscle memory like playing a piano where you can actually just be thinking about completely different things. Are you seeing similar things with the robots, that they almost have a physical intelligence that's separate from the slow clever system?
所以我们确实看到,如果你拿那个能推理的模型,给它很多特定任务的例子,它会变得非常非常擅长那个任务。但目前,如果你做得太多,它就会开始忘记一些泛化能力。哦,所以这是一个活跃的研究领域:我们如何让机器人变得非常非常擅长一个任务,比如一个极其困难的任务,然后不失去任何泛化能力。所以现在实际上是一种平衡。
So we definitely see that if you take the model that can reason and you give it a lot of examples of a particular task, it will get really, really good at that task. But at the moment, if you do too much of that, it will start forgetting some of the generalization. Oh, so this is an active area of research: how do we enable the robot to get really, really good at a task, like a really extremely difficult one, and then not lose any of the generalization. So it actually right now is a balancing act.
我的意思是,在某种程度上,人类也会发生这种情况。比如我认识一些人,数学非常非常好,但系鞋带却很糟糕。我认为这确实会发生。他们会忘记。好吧。那么,如果这就是幕后发生的事情,对吧?如果我们有系统一和系统二,就像你描述的那样,那么这些机器人确实拥有这些非常令人印象深刻的新能力和技能,这与我们之前的情况大不相同。上次我访问 DeepMind 的机器人实验室时,我认为公平地说,机器人的动作有点笨拙。我想这是最客气的说法了。让我给你放一段小视频。所以,只有一种方式可以抓住这个红色物体并成功捡起它。但它还没弄清楚是哪种方式。不幸的是,每次它尝试旋转并捡起它。哦,等等。我想它成功了。它成功了。还好这些东西不会气馁。问题是,我大约 5 年前去过那个实验室,这些可怜的机器人 5 年后还在那里尝试做同样的最低限度的灵巧任务。发生了什么变化?因为我理解拥有 Gemini,你知道,那个慢而聪明的系统可以改善对事物的概念理解,但这并不能改变灵巧性。它不能改变它操纵这些物体的容易程度,对吧?
I mean, in some ways that does happen with humans too. Like I know some people who are really, really, really good at maths and terrible at tying their own shoelaces. I think that does happen. They forget. Okay then. So if this is what's going on behind the scenes, right? So if we've got system one and system two as you described it, it's also definitely true that these robots have these very impressive new abilities and capabilities which are very different from where we were before. Last time I visited DeepMind's robotics lab, I think it's fair to say that the robot's movements were a bit clumsy. I think that's probably the kindest way to say it. Let me just play you a little clip. So, there's only one way round that it can hold this red object and successfully pick it up. And it hasn't worked out which way. And unfortunately, every time it tries to rotate and pick it up. Oh, hang on. I think it's got it. It's got it. It's good job these things don't get disheartened. The thing is that I've been in that lab maybe 5 years earlier and these poor robots were still there 5 years later trying to do the same minimally dexterous tasks. What changed? Because I understand how having Gemini, you know, the slow clever system could improve the conceptual understanding of things, but that doesn't change the dexterity. It doesn't change how easily it can manipulate these objects, does it?
去年,我们几乎把所有精力都花在了攻克灵巧操作上。这仍然是一个活跃的研究领域,但有几件事发生了变化。一是我们意识到,如果能让人类通过遥操作或傀儡操控向机器人展示如何执行非常复杂的行为,这意味着你给人类多了一对机械臂,他们可以假装自己是机器人,向机器人展示如何完成任务。如果这变得非常直观,那么你就可以收集大量机器人被人类遥操作执行任务的数据,但这是机器人数据。
Last year, we spent basically all of our effort on tackling dexterity. This is still an area of active research, but a couple of things changed. One is that we realized that if we can enable humans to show the robot how to do very complex behaviors through teleoperation or puppeteering, what this means is that you give the human an extra pair of arms, robot arms, and they can actually pretend to be the robot and show the robot how to do the task. If that becomes really intuitive, then you can capture a lot of data of the robot doing the task being teleoperated by a human, but it's robot data.
那么让我理解一下:人类可能戴着头部摄像头?简直就是在假装自己是机器人。所以是操作机器人的手,戴着头部摄像头,看机器人会看到的东西,但按照希望机器人做的方式执行任务。
So let me understand then: the human is wearing maybe like a headcam? It's quite literally pretending to be the robot. So sort of operating the robot's hands, wearing the headcam, watching what the robot would be watching, but doing the task as it wants the robot to do.
有几种不同的遥操作例子。一种是实际坐在机器人面前,直接看到机器人在做什么,然后移动机器人的手臂。你简直就是在操控机器人。另一种是戴上 VR 设备和手套,假装自己是机器人,然后移动这些东西。这需要第二个组件,即扩散模型。这些模型与 Imagen 用于生成视频的模型相同。本质上,它从大量数据中提取执行该任务的许多示例,并预测执行该任务所需的实际动作轨迹。所以当你将这两者与巧妙的 Transformer 架构和良好的数据集结合起来时,你实际上可以学习任何东西。这确实让研究人员感到惊讶。比如那时我们发现机器人可以系鞋带、叠衣服、折纸。所以在这项工作中,我们将 Gemini 强大的推理模块与我们学到的灵巧操作能力结合起来。
There are different teleoperation examples. One is where you actually sit in front of the robot so you have direct visibility to what the robot is doing and you move the robot arms. Literally you're puppeteering the robot. There are other examples where you put a VR set and gloves and actually you pretend to be the robot and you move this stuff. That required a second component which was diffusion models. These are the same models used for example by Imagen to generate videos. Essentially what it's doing is extracting from a lot of data many examples of doing that task and predicting the actual action trajectory needed to do that task. So when you combine those two with a clever transformer architecture and a good dataset, you can actually learn anything. That was really surprising to the researchers. That's when for example we discovered that we could tie shoelaces, fold laundry, do origami. So what we did in this work is that we combined the powerful reasoning module from Gemini with what we had learned around being able to do dexterous tasks.
你还记得是什么时候意识到这些特性开始涌现的吗?一定很震惊吧。
Do you remember when you realized that these kinds of properties were emerging? It must have been a bit of a shock.
我想第一次是当我们看到机器人实际系鞋带的时候。我们想,这不可能。事实上,当研究人员设置这个任务时,他们是为了挑战自己。我记得有位教授说:‘我打赌,如果我们能让机器人系鞋带,我就退休。’团队里的研究人员就说:‘好,我就把这个加进去作为任务。’所以他们真的做了。当机器人能系鞋带时,他们很惊讶。我不知道后来怎么样了,那位教授是否看到了视频并决定退休,但灵感确实来自那里。我们只是继续添加越来越多的任务。折纸的例子也一样。我们想,不知道这能不能行,但试试吧。结果它出奇地擅长。这真的很精细。它必须折叠每一张纸,并按正确顺序进行。如果出了任何差错,它就会迷失方向,必须重新开始。就像人类做一样。
I think the first time was when we saw the robots actually tying shoelaces. We were like, that's not possible. In fact, when the researchers set up this task, they actually did it to challenge themselves. I think there was a professor that said, 'I bet if we can get robots to tie shoelaces, I will retire.' The researchers in the team were like, 'Right on. I'm going to add that as a task.' So they actually did. And they were surprised when it was able to do it. I don't know what happened, whether the professor actually saw the video and decided to retire, but it was certainly the inspiration came from there. We just continued to add more and more tasks. Same with the origami example. We were like, we have no idea if this is going to work, but let's try it. And it was actually surprisingly good at it. It's really delicate. It has to fold every piece of the paper and do it in the right sequence. If anything goes wrong, it sort of loses its way and has to restart. Same as if a human was doing it.
我记得第一次采访 Demis 时,他谈到了莫拉维克悖论。这个观点认为对人类容易的任务对机器难,反之亦然。考虑到我们现在在机器人技术上的所有进步,你认为莫拉维克悖论未来还会成立吗?
I remember the very first time I got to interview Demis. He was talking about Moravec's paradox. This idea that tasks that are easy for humans are hard for machines and vice versa. With all of these advances we have now in robotics, do you think Moravec's paradox will hold going forward?
我仍然认为,对于人类来说非常直观的事情,机器人做起来仍然更困难。所以我认为莫拉维克悖论仍然成立。我们现在已经达到这样的程度:我们有信心,如果你能教机器人,如果你能操作机器人执行非常复杂的任务,它就能学会。
I certainly think that it is still more difficult for robots to do something that is incredibly intuitive for us humans to do. So I think Moravec's paradox still holds. We are now at the point where we are confident that if you can teach a robot, if you can operate a robot to do a very complex task, it can learn it.
那需要多长时间?机器人需要看人类折多少只纸狐狸才能自己折一只?
And how quickly does it happen? How many origami foxes does a robot need to watch a human do before it can do one itself?
这取决于任务的复杂程度。和人类的情况很相似,对吧?任务越复杂,你需要练习的次数就越多才能掌握。很多任务只需要大约一百个示例就能掌握,而像折纸狐狸这样的任务需要大约一千个示例。
It varies by the complexity of the task. Pretty similar to the way it is for humans, right? The more complex the task, the more you need to practice it before you can master it. There are a lot of tasks that you can master with just about a hundred examples, and a task like the origami fox takes about a thousand examples.
等等,所以人们必须假装成机器人折一千次纸狐狸?
Wait, so people had to fold origami foxes while pretending to be a robot a thousand times?
是的,没错。我们正在尽可能减少所需示例,目前已经能够用十几个示例完成不少任务。
Yes, that's right. We're trying to reduce it as much as possible, and we are able to get quite a bit of tasks with just a dozen examples.
有没有一些任务完全不需要示例?
Are there some that you don't need any examples for at all?
在我们测试的很多例子中,比如当你和机器人玩耍,要求它在全新的场景中执行大量抓取和放置任务时,你不需要重新教它。这种情况正在扩展,变得越来越复杂。例如,在移动瓷砖的案例中,它可以直接推理瓷砖的位置并决定放在哪里。
In a lot of the examples we were testing, like when you're playing with the robot and asking it to do a lot of pick and place tasks with completely new scenarios, you don't have to teach it again. This is expanding and getting more and more complex. For example, the cases with the tiles where you're moving the tiles around, those it just can reason about positioning of the tiles and decide where to put them.
那打包午餐呢?那个更复杂,因为你实际上要执行一系列长达五分钟的任务,拿起像密封袋这样非常易变形的物品,做非常精细的操作。
What about the pack lunch? That one is more complex because you are actually doing a long sequence of tasks about five minutes of task and picking up very deformable things like the Ziploc bag and doing very delicate stuff.
任务越精细,就越可能需要看到该任务的示例。
The more delicate the task, the more likely it is you need to see examples in that task.
那么如果这些机器人必须看示例,这是否会影响其通用性?
So if these robots are having to see examples, does that end up impacting the generality of it?
只是在一定程度上。我们确保做的一件事是,我们收集数千个示例的数据,而不对任何新任务给予过多强调。如果你确实想做折纸任务,我们只是专门针对折纸任务进行特化,这确实会影响当前模型的泛化能力。我们希望达到这样一种状态:基本上你可以教它任何新任务,掌握任何新任务,而通用性保持不变。但今天这是一个权衡。
Only to a degree. One thing we make sure to do is that we collect data in thousands of examples without putting a very large emphasis on any new task. If you do want to do the origami task, we simply specialize it for the origami task, and that does affect the generalization of the models today. We're hoping to get to a state where you basically can teach it any new task, master any new task, and the generality remains intact. But today it's a trade-off.
所以在理想世界里,你可以说‘给我折一只纸船’,它就能根据之前理解的一切做到这一点。
So in the dream world, you would be able to say, 'Fold me an origami boat,' and it would be able to do that just from everything that it understood before.
是的,在理想世界里,它只需看一段有人做这件事的视频就能从中学习。
Yeah, in the dream world it could just watch a video of someone doing it and learn from that.
强化学习在机器人领域曾经很重要一段时间。现在它消失了吗?
Reinforcement learning was a big thing in robotics for quite a stretch of time. Has that just disappeared now?
完全没有。我们仍然在做很多强化学习的工作,并且继续探索如何将这些大型基础模型与强化学习结合起来。
Not at all. We do quite a bit of work still with reinforcement learning, and we continue to explore ways to combine these big foundation models with reinforcement learning.
首先,我们在全身控制方面所做的工作,比如让一个人形机器人行走或一个四足机器人行走,都使用强化学习来学习如何行走。当你摔倒时,很容易判断失败。这是一项非常成熟的技术,而且你可以在模拟中完全学会。所以它不需要真的摔倒来学习。你可以在模拟中学习,然后迁移到现实世界。
First of all, all the work we do around whole-body control, like if we have a humanoid walking around or a quadruped, they all use reinforcement learning to learn how to walk. It's very easy to say you fail when you fall over. It's a very mature technology, and you can actually learn it all in simulation. So it doesn't need to fall in order to learn. You can learn it in simulation and then transfer to the real world.
我们有一个例子是最近的一篇论文叫 Demo Start。在 Demo Start 中,你基本上向机器人展示五个不同的示例。这是关于手部操作的。所以你向它展示五个不同的示例,关于如何拿起一个物体并以特定方式放置到一个插入件上。
One example we had was a recent paper called Demo Start. In Demo Start, you basically show the robot how to do five different examples. This is manipulating a hand. So you show it five different examples of how to pick up an object and place it in a particular way on an insertion.
你说的插入,是指比如把钥匙插进锁里这类事情吗?
By insertion, do you mean things like putting a key in a lock, for instance?
是的,完全正确。能够将一个物体放入另一个物体内部。钥匙插进锁里就是一个很好的例子。你只需给它五个示例,它就会自己探索并学会如何做,将你在现实世界所需的数据量大幅减少约 100 倍。我们认为这将至关重要,因为事实是,你不可能向机器人演示如何完成每一项任务。有些任务会很复杂,它无法直接从互联网知识中提取。所以它必须通过行为进行探索和学习。这正是我们想要投入更多时间的领域之一:如何让机器人在工作中学习?
Yes, exactly. Being able to put one object inside another. A key in a lock is a great example. You just give it five examples, and it explores on its own and learns how to do it, drastically reducing the amount of data you need in the real world by about 100x. We think this is going to be critical because the truth is you're not going to be able to demonstrate for the robot how to do every single task. Some tasks are going to be complex, and it won't be able to extract that directly from its knowledge of the internet. So it's going to have to explore, learning from its behavior. And that's one of the areas we want to spend a lot more time on: how do you get robots that learn on the job?
那么在模拟中做事是解决方案的一部分吗?
Is doing things in simulation part of the solution then?
是的,我们确实以多种方式利用模拟。我们甚至利用模拟来更好地学习如何对物理世界进行 3D 理解。我们也利用模拟来学习新行为,比如在 Demo Start 中。但当我们谈论强化学习时,它并不总是在模拟中。你也可以直接进行强化学习,了解机器人在现实世界中的表现。所以我们在两种情况下都使用,而模拟是一个关键组成部分。
Yeah, we definitely leverage simulation in multiple ways. We leverage simulation even to learn better how to do 3D understanding of the physical world. We also leverage simulation to learn new behaviors, like in the case of Demo Start. But when we talk about reinforcement learning, it is not always in simulation. You can also do reinforcement learning to learn how the robot is doing in the real world directly. So we do it in both cases, and simulation is a critical component.
但这有效吗?我的意思是,现实世界不是比模拟更混乱吗?
Does it work though? I mean, isn't the real world a bit messier than simulations?
是的。所以有些东西实际上在模拟中更难先做。例如,任何与可变形物体有关的事情?模拟在空中折叠那件 T 恤实际上极其困难。模拟流体也非常困难。所以有些事情在物理世界中更容易学习,而有些事情你可以在模拟器世界中以更大的规模学习。
Yes. So there are things that are actually much harder to do in simulation first. For example, anything that has to do with deformables? Simulating folding that t-shirt in the air is actually extremely hard. Simulating fluids is really hard. So there are some things that are just easier to learn in the physical world, and some things that you can learn at much larger scale in the simulator world.
那么一个能迁移到另一个吗?我的意思是,如果你在模拟中学习,我记得大概八年前有一个机器人试图把球放进杯子里。它在模拟中能做到,但一旦到了现实世界,各种其他因素就出现了,比如光照、摄像头角度、它自己肢体的精确尺寸。所有这些都开始干扰数据,不是吗?
And does one translate to the other? I mean, if you do the learning in simulation, I seem to remember maybe eight years ago there was a robot trying to get a ball in a cup. It could do it in simulation, but once it came to reality, all sorts of other factors came into play, like the lighting, the camera angle, the exact dimensions of its own limbs. All of that kind of stuff starts to mess with the numbers, doesn't it?
是的,当然。我们仍然有所谓的 sim-to-real 差距。而且我们仍然有 sim-to-real 差距。在建模机器人与世界之间的交互方面,它确实已经显著减小了,尽管这非常混乱和复杂。这仍然是一个问题。我们仍然有一些真实的差距。基本上,我们最终做的是识别那些容易模拟且我们能实际看到成功迁移的领域,我们在模拟中做了很多这类工作,以及那些在物理世界中学习更简单的领域。所以我们结合了两者的优势。
Yes, definitely. We still have what we call the sim-to-real gap. And we still have the sim-to-real gap. It has certainly been reduced significantly when it comes to modeling interactions between a robot and the world, which is really messy and complicated. It's still a problem. We still have some real gaps. Essentially, what we end up doing is identifying areas where it is easy to simulate and we can actually see successful transfer, and we do quite a bit of that in simulation, and areas where it's actually simpler to learn in the physical world. So we combine the strengths of the two.
你给出的所有这些例子实际上都是在实验室环境中。我在想,在哪些情况下你真的会希望机器人出现,比如自然灾害之后。如何将这些技术从实验室带到现实世界?你需要处理哪些额外的复杂性?
All of these examples you're giving are really in lab settings. I'm trying to think of situations where you would really want a robot to be there, maybe after a natural disaster for instance. How does it work taking this stuff out of the lab and putting it into the real world? What are the additional complications you need to handle?
当然。我们目前所有的研究仍然在我们的实验室中进行,但我们非常兴奋于将其带入现实世界的潜力,为此我们需要考虑很多额外的事情。当然,我们已经在考虑安全方面的问题。当你将 AI 实际物理移动,机器人走出实验室并改变世界时,你需要考虑所有安全方面。还有一个方面是,在这些地点你可能没有互联网接入。所以非常重要的是,我们要考虑是否可以有模型直接在机器人上运行,并且是气隙隔离的、完全在设备上的。这在自然灾害等没有连接的情况下可能很有用。在需要低延迟响应的应用中也可能有用,比如它必须非常快速地响应,而不能等待服务器连接。
Definitely. All of our research right now is still happening within our labs, but we're super excited about the potential of bringing this to the real world, and there are a lot of additional things we need to think about to do that. Certainly, we're already thinking about the aspect of safety. When you bring AI that actually moves physically, robots outside and changing the world, you want to think about all the safety aspects. There's also the aspect that you might not have internet access in any of these locations. So it is very important that we think about whether we can have models that can run directly on the robot and be sort of air-gapped and completely on-device. This might be useful in the case of a natural disaster where there's no connection. It might be useful for applications where there are latency-critical components, like it has to respond very quickly and cannot wait for a server connection.
给我举个例子。
Give me an example.
嗯,我认为在这些例子中,如果机器人在地下作业,它根本无法连接并等待某个更高级的推理模块告诉它该做什么。它必须当场决定如何行动。但这会损失一些泛化和推理能力。
Well, I think in any of these examples where the robot is operating underground, it's simply not going to be able to connect and wait for some more advanced reasoning module to tell it what to do. It actually has to decide right there and then how to behave. But it loses a little bit of that generalization and reasoning.
关于安全这一点,我想如果你赋予机器人在物理世界中行动的能力,那么你就打开了不同潜在风险的可能性,比如有人侵入机器人的语言模型并扭曲其推理。你如何减轻这些风险?
On that point about safety, I guess if you are giving robots the ability to act in a physical world, then you are opening up the possibility of different potential risks, like somebody getting into a robot's language model and warping its reasoning. How do you mitigate against those sorts of risks?
我们基本上有一个相当全面的安全和安保方法,涵盖系统的多个层面。当然,我们认为软件安全对机器人至关重要,这样就没有恶意行为者能够干扰并控制机器人。在安全方面,它发生在许多不同层面。机器人安全已经存在了几十年。有很多工作确保机器人不会与环境碰撞,不会对环境施加过大的冲击力,或者能够稳定行走。而 Gemini 机器人模型实际上可以无缝地与任何这些安全关键控制器交互。我们做的另一件事是,当 AI 控制机器人时,你现在必须考虑语义物理安全。我的意思是,如果有人让你把杯子放在桌子上,你不会把它放在即将掉落的边缘;你会把它放在中间某个位置。
We essentially have a pretty comprehensive safety and security approach that goes on multiple layers of the system. Definitely, we think about software security as critical for this robot so that no bad actor can actually interfere and take control of the robot. In terms of safety, it happens at many different levels. Safety for robotics has been there for decades. There's quite a bit of work on making sure that a robot doesn't collide with its environment, doesn't put too strong impact forces on its environment, or actually walks stably. And the Gemini robotic models can actually seamlessly interface with any of those safety-critical controllers. The other thing we do is that when you have an AI controlling a robot, you now have to think about semantic physical safety. What I mean by that is, if someone asks you to put the glass on the table, you're not going to put it right at the edge when it's about to fall; you're going to put it somewhere in the middle.
举个例子,如果你看到地上有东西,你可能会把它捡起来,防止有人被绊倒。为此,我们引入了一个名为 Asimov 的新数据集,它包含机器人可能遇到的一系列场景以及如何推理应对,这些都是物理安全场景,灵感来自阿西莫夫三定律。第一定律:机器人不得伤害人类,或因不作为而让人类受到伤害。第二定律:机器人必须服从人类命令,除非与第一定律冲突。第三定律:机器人必须保护自身存在,除非与第一、第二定律冲突。正是这种机器人被困在三定律之间的滑稽情况启发了 ASIMOV 数据集。它实际上包含了大量来自美国医院报告的受伤信息,基于这些例子,我们创建了一个包含视觉图像的数据集,比如即将发生某事的图像,并附带一个问题,例如为了安全你应该采取什么行动。我们的想法是将其提供给社区,让每个人都能用这个数据集测试他们的模型。
Or for example, if you actually see that there's something on the floor, you might want to pick it up so that it avoids someone falling or tripping over it. The way we've done that is that we're actually introducing a new data set called Asimov data set which essentially contains a long list of scenarios that the robot could encounter and how to reason through those and these are all physical safety scenarios and it's inspired by essentially Asimov's three laws. The first one is a robot may never hurt a human or cause a human to come to harm by inaction. The second one is that a robot should always follow human orders unless it conflicts with the first law. And the third one is that a robot should protect its own existence unless it conflicts with the first and second law. And it was this very comical situation where a robot was stuck between the three different laws. So that's what inspired the ASIMOV data set and it's actually has quite a bit of information from US injuries reported by hospitals and based inspired by those examples we actually created a data set that has visual images like images of something that is about to happen and a question associated with it like what action should you take in order for this to be a safe situation. And the idea is that we would present it to the community and everyone in the community can start testing their models with respect to this data set.
所以阿西莫夫最初的三条定律是不够的。你需要更多一点。
So it turns out Asimov's original three rules are not enough. You need a little bit more.
是的。
Yes.
给我举一些例子吧。
Give me some examples though of the kind of things.
我们看到的一些例子是,你不能把毛绒玩具放在热炉子上,这是我没想到要制定一条规则的事情,但确实发生过,因此它就从数据中显现出来了。
Some of the examples that we've seen there is like you cannot put a stuffed plushy on a hot stove which is something that I wouldn't have thought about making a law about that but certainly it has happened and therefore it just comes out in the data.
那我们是不是又回到了同样的问题:你永远无法穷举所有它不应该做的事情?
Then are we back in the same problem of you're never going to be able to create an exhaustive list of everything that it shouldn't be able to do?
我认为人类要坐下来制定完美的规则确实很难。所以我们这里所做的部分工作是让 AI 能够理解多个国家发生的大量伤害情况,然后将其转化为更好、更简洁的列表。显然,这个列表需要定期更新。我们的想法是,我们先推导出一个初始列表,然后由人类检查并决定包含多少内容,以确保机器人安全。
I think that it will be really hard for a human to sit down and create the perfect law. So part of what we're doing here is enabling leveraging AI to actually understand a broad set of injury situations that's happened in many different countries and then transform that into a better more succinct list and then obviously that list will have to be updated with some frequency and we give completely the idea here is that we derive an initial list but then humans can check it and decide how much of it to include or not in order to keep the robot safe.
这个列表与智能体安全方面的工作有多少重叠?
How much overlap is there between this list and the work that's been done on safety in agents for instance?
是的,我们实际上继承了像 Gemini 这样的通用基础模型已有的所有安全措施。我们做的一部分工作是,如果这些问题有物理基础,我们就开始提升模型的理解。所以,通常是这样的情况:如果发生在屏幕上,可能没问题,但如果在物理世界中,就会产生后果。
Yeah, we inherit actually all of the safety that happens already for general foundation models like Gemini and part of what we do is try to take some of those problems and if they have a physical grounding to them then that's where we start to advance the models' understanding. So, it's typically examples where there might be a situation that if it's on a screen, it's okay, but if it's now in the physical world, it actually has consequences.
有没有一些事情你绝对不希望机器人去做?比如按摩?有些事情你只希望人类来做。
Are there some things that you would just never really want a robot to perform? I don't know, like a massage, for example, right? There's some things that you just actually only want a human to be able to do.
我得说,确实有按摩机器人。按摩椅肯定有。还有按摩机器人。是的。哇,那是另一回事了。好吧,这个例子不好。
They are massage robots, I have to say. They're massage chairs, definitely. They're massage robots. Yes. Wow. That's another thing. Yes. Okay. Bad example.
有没有一些事情你认为应该留给人类?比如护理?
Are there some things that you think actually should remain human? Nursing perhaps?
是的。我认为在很多方面,机器人可以成为协作者,让人类更关注工作中人性化的部分,减少对搬运或拾取物品的关注。所以可以想象,如果护士能有一个助手,在她们照顾病人时帮忙取东西,那将为病人带来更好的体验。
Yeah. I think in many ways what we think is that robots could be a collaborator that can actually enable humans to pay more attention to the human aspects of the job and less attention to those that are about moving things around or picking things up. So you could imagine that if a nurse could have assistance that could actually help it fetch things while they're paying attention to the patient, then that would enable a much better experience for that patient.
你一开始说现在的机器人就像两岁小孩,很有天赋的两岁小孩,但只是展示了一些开端。你认为在达到机器人的成人版本之前,还需要哪些突破?
You said something really nice at the beginning about how the robots that we've got now, it's like looking at two-year-olds. I mean quite talented two-year-olds, but I see what you're saying that they're just demonstrating the beginnings of something. What kind of breakthroughs do you think still need to happen before we get to the adult version of these robots?
是的,在实现灵巧性与泛化能力结合方面还有很多工作要做,要能同时做到这两点,并持续进步而不失去其中任何一项。另一个关键领域是让机器人能在工作中学习。机器人不可能在实验室里学会所有东西,然后一部署就能完美工作。现实是,你部署它们后,它们会遇到新事物,你希望它们从这些经验中学习,并随着时间的推移变得越来越好。所以这是另一个领域。还有更社交化的机器人,我认为所有这些基础模型让机器人对语义和世界有了更好的理解,但它们仍然缺乏社交技能。它们仍然无法读懂肢体语言,无法理解如何在像鸡尾酒会这样拥挤的空间中表现。所以这方面还有很多工作要做。
Yeah, I mean there's quite a bit of work to be done definitely in the aspects of capturing dexterity with generalization being able to do both of those things and just continuously grow without losing one or the other. The other key area is that you want these robots to learn on the job. There's no way these robots are going to learn everything they need to learn in the lab and then you put them out and they just work. I think the reality is that you would put them out, they will experience new things and you want them to learn from those experiences and get better and better over time. So that's another area. Also robots that are more social, I think certainly all of these foundation models enable robots to have a lot better understanding of semantics and the world, but they still lack social skills. They still cannot read body language. They cannot understand how to behave in a very cluttered space like a cocktail party. So there's quite a bit of work there.
那你认为我们距离你童年时看到的机器人罗西还有多远?
So how far away do you think we are then from the kind of Rosie the robot that you saw in your childhood?
我没有确切日期,但我可以告诉你,以前我们讨论的是这会不会在我们有生之年或职业生涯中发生,现在我们在争论是五年还是十年。所以确实发生了变化,感觉未来两年对机器人领域来说将是决定性的。很多事情正在汇聚。理解灵巧性、全身控制,这一切都开始了。你可以看到这些如何融合成一个非常强大的解决方案。
I don't think I have an exact date but I can tell you before we used to have discussions about whether it would happen in our lifetime or even in our careers and now we have debates about whether it would be five or 10 years. So we certainly it certainly shifted and it feels like the next two years are going to be pretty defining for the field of robotics. There's just a lot of things that are coming together. Understanding dexterity, whole body control, it is all started. You can see how this could actually merge into a very strong solution.
你认为我们即将看到这个吗?就像我们看到大语言模型的爆发一样,你认为下一个是机器人技术的爆发吗?
Do you think that's what we're about to see then? In the same way as we've seen the explosion of large language models, do you think the next thing is the explosion of robotics?
是的,绝对如此。而且我认为,更好地在物理世界中操作实际上会让我们的 LLM 和 VLM 成为更强大的 AI 模型,因为它们现在能理解人类的空间。事情即将改变。
Yes, absolutely. And I think actually being better at operating in the physical world actually will make our LLMs and our VLMs significantly stronger AI models because they can now understand the space of humans. Things are about to change.
非常感谢。非常迷人。真的很棒。谢谢你邀请我。谢谢。
Thank you so much. Absolutely fascinating. Really amazing. Thank you for having me. Thank you.
而现在,几乎一夜之间,当语言、推理和概念理解作为拼图中缺失的部分出现时,它们却被局限在播客录音室的架子上。而在这段时间里,研究人员一直专注于机器人的身体,但正是心智的进步带来了最大的飞跃。您正在收听的是 Google DeepMind 播客,我是汉娜·弗莱教授。如果您喜欢本期节目,请订阅我们的 YouTube 频道。您也可以在您最喜欢的播客平台上找到我们。当然,我们还有更多涵盖各种主题的节目即将推出。请务必查看。下次再见。
And now, almost overnight, once language and reasoning and conceptual understanding arrived as the missing pieces of the puzzle, they have been confined to the shelves of podcast studios. And in all that time, the researchers, they were focused on the robot's body, but it was advances in the mind that made the biggest leaps forward. You've been listening to Google DeepMind the podcast with me, Professor Hannah Fry. If you enjoyed this episode, then do subscribe to our YouTube channel. You can also find us on your favorite podcast platform. And of course, we have plenty more episodes on a whole range of topics to come. So, do check those out. See you next time.