Building Foundation Models for Robotics with Physical Intelligence
打开互动全文版(中英对照 + 朗读 + 问答)→Carol 和 Toby 解释为何经典机器人方法有误,以及端到端强化学习如何让部署成为可能。
Carol and Toby explain why the classical approach to robotics was wrong and how end-to-end learning with reinforcement learning is making deployment possible.
就像整个这东西能运作这件事本身,就挺令人震撼的。
Just like the fact that this whole thing works, it's kind of mind-blowing.
是啊。
Yeah.
对。你构建了这个大致受大脑启发的、拥有非常通用学习算法的东西。你喂给它数据,它不知怎么就学会了,而且学得比我们以前任何东西都好。这适用于机器人,也适用于视觉、语言、声音以及各种其他领域。如果你停下来想一想它是如何运作的,以及它居然能运作,这绝对令人震撼。
Right. You build this loosely brain-inspired thing that has a very general-purpose learning algorithm. You feed it data and it somehow gets it, and gets it way better than anything we've ever had before. And this applies to robots and it applies to vision and language and sound and all kinds of other things. If you stop for a second and just think about how it works and that it works, it's absolutely mind-blowing.
在本期节目中,我们与 Physical Intelligence 的 Carol 和 Toby 坐下来聊聊,这家公司正在为机器人构建基础模型。Carol 和 Toby 解释了为什么将机器人技术分解为感知、规划和控制这种经典方法从根本上就是错误的,以及端到端学习结合强化学习如何最终让部署成为可能。你将听到他们如何实现稳健的真实世界性能,让机器人连续 13 小时煮咖啡,以及这些模型如何以我们尚未完全理解的方式,在从手术机器人到无人机飞行等截然不同的任务中泛化。我们还讨论了 Pi Star 0.6 背后的技术见解,这是 Physical Intelligence 的最新模型,它通过强化学习从经验中学习。请享受节目。
In this episode, we sit down with Carol and Toby of Physical Intelligence, a company building foundation models for robotics. Carol and Toby explain why the classical approach of breaking robotics down into perception, planning, and control was fundamentally wrong, and how end-to-end learning with reinforcement learning is finally making deployment possible. You'll hear how they achieved robust real-world performance, getting robots to make coffee for 13 hours straight, and how these models generalize across radically different tasks from surgical robots to drone flying in ways that we don't fully understand. We also talk about the technical insights behind Pi Star 0.6, which is Physical Intelligence's newest model that learns from experience using reinforcement learning. Enjoy the show.
Carl,Toby,非常感谢你们今天来到这里。
Carl, Toby, thank you so much for joining us here today.
谢谢邀请。很兴奋能谈论关于物理智能、通用机器人等一切。也许在我们深入之前,先为我们的听众介绍一下,Physical Intelligence 是什么,以及你们追求的使命?
Thank you for having us. Excited to talk everything physical intelligence, general robotics, etc. Maybe before we get into it, just for our audience, can you share a little bit about what Physical Intelligence is and the mission that you're after?
是的,在 Physical Intelligence,我们正在构建机器人基础模型。这些模型原则上应该能够让任何机器人执行任何任务。在过去大约一年半的时间里,我们开始构建正确的构建模块,展示这些模型如何扩展。我们已经证明它们能够控制许多不同的机器人形态,许多不同类型的机器人。我们还证明了它们能够泛化,所以你可以把它带到全新的环境中,以及它们泛化需要什么条件。而我们刚刚发布的这个最新版本,叫做 Pi Star 0.6,我们也想多告诉你一些,它展示了我们如何让它们获得良好性能,从而开始变得可部署。
Yeah, so at Physical Intelligence, we are building robotic foundation models. These are models that in principle should be able to have any robot do any task. Over the past one and a half years or so, we started building the right building blocks that show how these models could scale. We've shown that they're able to control many different robotic form factors, many different types of robots. We've also shown that they're able to generalize, so you can bring it to completely new environments and what it takes for them to generalize. And this last release that we just had called Pi Star 0.6, that we also wanted to tell you more about, shows how we can get them to good performance so that they're starting to become deployable.
这对我们来说非常重要,因为我们希望看到这项技术真正部署在现实世界中,但也因为我们没有互联网上免费数据的好处。没有机器人动作的数据。所以我们需要自己创建数据集。所以我们追求的是物理智能的问题,是为机器人创建基础模型的问题。我们已经取得了相当大的进展。
And this is really important to us because we want to see this technology actually deployed in the real world, but also because we don't have the benefit of having the free data on the internet. There is no data of robot actions. So we need to create the data sets ourselves. So we are after the problem of physical intelligence, after the problem of creating foundation models for robots. And we've made quite a lot of progress.
太好了。我能问一下为什么决定构建基础模型,而不是像现在有些公司那样构建完全垂直整合的机器人产品?你知道,上个月的 Sunday 发布我还记得。你可以买一个可爱的小机器人助手来帮忙做家务。有公司在做烹饪机器人。显然还有做人形机器人的公司。为什么构建基础模型而不是自己造机器人?
Wonderful. And can I ask why the decision to build foundation models as opposed to, you know, there are companies that are building fully vertically integrated robotic products right now. You know, the Sunday launch last month is in the back of my head. You can buy a cute little robot helper for your household. There's companies working on cooking robots. There's obviously the humanoid companies. Why build a foundation model versus build a robot yourselves?
是的。所以我认为如果你看看机器人学的历史,对我来说非常清楚,而且我认为对许多机器人学家来说,我们一直受限于智能。我们已经有能够做不可思议事情的机器人,无论是在家庭还是工业环境中。我们十多年前就看到过机器人,如果远程操作,它们可以打扫整个房子。而真正重要的条件是‘如果远程操作’。所以如果背后有人类思维,很明显硬件能够做很多不同的事情。在很长一段时间里,机器人公司都是按照你描述的方式构建的,你基本上是在考虑创建一个专门为单一任务或单一应用设计的特定机器人。相反,我们认为真正能帮助这个领域的是专注于智能这个瓶颈。所以我们创建了一家公司来专注于这个瓶颈,因为我们认为如果我们解决了这个瓶颈,我们就能真正让机器人成为现实。如果你用其他方式,你基本上没有在瓶颈上取得尽可能多的进展。所以我们认为我们应该直接针对这个问题。专注于智能,如果我们能做到,那将导致许多不同的垂直产品。它将导致机器人在家庭、工业环境,基本上任何地方部署。
Yeah. So I think if you look at the history of robotics, it's very clear to me and I think to many roboticists that we've been always bottlenecked on intelligence. We've had robots that are capable of doing incredible things whether it's in the home or in industrial settings. We've seen robots more than a decade ago that if teleoperated, they can clean the entire house. And the really important caveat is 'if teleoperated'. So if there is a human mind behind it, it's clear that the hardware is capable of doing lots of different things. And for a very long time, robotics companies have been structured the way you described, where you kind of think of creating a specific robot that's designed to do just a single task or a single application. Instead, what we thought would really help the field is to focus on the bottleneck on the intelligence. So we created a company to focus on that bottleneck because we think that if we address that bottleneck, we can actually make robots happen. And if you do it any other way, you're basically not making as much progress on the bottleneck as you could be. So we thought we would just target this problem head-on. Focus on the intelligence, and if we can do that, that would lead to many different vertical products. It will lead to robots being deployed in the home, in industrial settings, basically anywhere.
我能稍微施压,测试一下吗?在硬件方面,我看到了最新的视频,比如 Optimus 的手。它很精致。是一件艺术品。而我之前没看过十年前人们远程操作机器人打扫房子的视频,但我在想,是否有一组任务现在正处于可能实现的边缘,比如烹饪,或者能够削皮切洋葱,这些在目前的硬件之前是无法做到的。所以你认为硬件在‘为什么是现在’这个问题上占多大比重?
Can I just pressure that, test that a little bit? So on the hardware side, I've seen the latest videos, for example, of the Optimus hand. It's exquisite. It's a piece of art. And I hadn't seen the videos of people teleoperated robots cleaning houses 10 years ago, but I'm wondering if there's a set of tasks that's maybe now just on the cusp of becoming possible, for example, cooking or being able to peel and dice an onion that you couldn't have done with hardware prior to where we currently are. So how much of a 'why now' do you think hardware is or isn't?
所以,硬件方面有很多进步,尤其是在人形硬件上,比如灵巧手,正如你提到的,它们现在比几年前好多了。
So, there's a lot of progress in hardware, especially in humanoid hardware, like dexterous hands for instance, as you mentioned, they're much better now than they were even a few years ago.
是的。
Yeah.
但这仍然没有解决瓶颈问题。我们以前甚至可以用简单的夹爪让机器人切菜或做饭。问题是我们没有智能来操作这些机器人。而且硬件越复杂,并不能真正解决那个瓶颈,对吧?它可能让你做更多事情,但你仍然受限于机器人不够智能这个根本挑战。
But that still doesn't address the bottleneck. We could have had robots operating, you know, chopping vegetables or doing cooking even with simple grippers before. The problem is that we don't have the intelligence to operate these robots. And the more complex the hardware is, it doesn't really resolve that bottleneck, right? It allows you to do more potentially, but you're still bottlenecked by the fundamental challenge of robots not being intelligent enough.
我明白了。所以硬件可能提高了你能做的事情的上限,但在能力下限上我们甚至还没到那一步。
I see. So hardware may raise the ceiling on what you're able to do, but with the capability floor we're not even there yet.
没错。所以即使使用简单的机器人,我们还没有达到人类操作员的水平。
That's right. So even with simple robots we are not yet at the level of a human operator.
所以限制在于智能层。发展智能的限制是什么?是收集数据吗?是低成本地做这件事吗?因为你已经分解了问题。我们会继续问你为什么,并进一步深挖。那么下一层是‘好吧,解决智能泛化的瓶颈是什么?’
So the limit being the intelligence layer. What's the limit to developing the intelligence? Is that collecting data? Is it doing it cheaply? Because you've broken down the problem. We're going to keep asking you why and just drill down further. So what's the next layer of the 'okay what's the bottleneck for solving intelligence generalization?'
这是个好问题。所以我们从三个因素来考虑。
It's a good question. So we thought about it in terms of three factors.
我们称之为能力、泛化和性能。关于能力,我们的想法是,只要你能为某个任务或机器人收集数据,就应该有一个模型能够复制并自动化该任务。
We refer to them as capability, generalization and performance. With capability, our idea was that we want to get to the point where as long as you can collect data for a task or for a robot, you should have a model that can replicate that and automate that task.
这一点我们很快就实现了。大约一年前我们发布了 PI zero,证明了如果你能为任何任务、任何机器人收集数据,你就能自动化它,并且它能够学习。
This is something that we've gotten to fairly quickly. This was our PI zero release around a year ago or so, showing that it's basically possible that if you can collect data for any task, for any robot, you should be able to automate it and all should be able to learn it.
下一个挑战是泛化,这仍然是一个未解决的问题。我们希望达到这样的程度:机器人可以零样本工作,你把它带到一个新家,它就知道如何在那里操作。这是一个非常困难的问题。如果你把机器人放在一个新家,它需要理解不同物品的位置,台面看起来不同,光线也不同,等等。我不会说这个问题已经解决了,但我认为我们开始掌握如何解决它以及它如何扩展。在机器学习中,我们知道的泛化的唯一答案是通过数据的多样性。所以如果你看到很多不同的数据集,你应该能够泛化到类似的环境。这是我们在今年 4 月发布的 PIO5 中看到的。我们达到了这样的程度:我们可以把机器人带到一个它从未去过的新家,它能够在那里操作。虽然还不完美,但至少它有一些常识,知道如何完成简单的任务,比如清理厨房。
The next challenge is around generalization and this is still an open challenge. We wanted to get to the point where robots can just work zero shot and you can bring them to a new home and they should know how to operate in that home. This is a really difficult problem. If you put a robot in a new home, it needs to understand where different items are, that the counters look different, the lighting is different, and so on. I wouldn't say this problem is solved, but I think we start to get a handle on how to solve it and how it scales. The only answer to generalization that we know in machine learning is through diversity of data. So if you see a lot of diverse data sets, you should be able to generalize to a setting that is similar. This is something we've seen with our PIO5 release in April of this year. We got to the point where we can bring a robot to a new home it's never been to before and it's able to operate in that home. It's not perfect yet, but at least it has some common sense on how to go about simple tasks like cleaning up the kitchen.
最后一个尚未完全解决的挑战是性能。我们如何让这些模型的性能足够好,以便我们能够实际部署它们?部署非常重要,因为我们也需要收集数据。我认为这将是收集数据的最可扩展的方式,因为机器人会在外面执行有经济价值的任务,而数据收集的成本基本上是负的。你部署这项技术的范围越广,获得的数据就越多。最终,这将是最大的数据来源,比互联网数据还要大得多。
The last challenge that is also not fully solved yet is performance. How can we get these models to the point where the performance is good enough so we can actually deploy them? Deployments are really important because we also need to gather data. I think that is going to be the most scalable way of collecting data because you'll have robots out there doing economically valuable tasks, and the cost of that data collection is basically negative. The more broadly you can deploy this technology, the more data you'll be getting. In the limit, that will be the biggest source of data, much bigger than internet data for instance.
你认为我们离泛化或达到某种性能水平还有多远?也许是在受控环境中,也许是在家庭或办公室的一般环境中,但不是整个世界。如果你能限制范围,你认为在我们可以部署这类机器人之前,泛化性能需要达到什么水平?
How far away do you think we are from generalization or from a performance level that maybe it's a controlled environment, maybe it's a general environment in homes or offices but not the whole world? If you could limit that, where do you think generalization performance will need to be before we can deploy these kind of robots?
我认为我们实际上已经非常接近部署这些机器人了。我们自己已经开始部署了。我们原本以为这需要大约 5 年时间才能达到技术真正准备好将机器人部署到商业环境中并执行有价值任务的程度。但我们在大约两个月前就做到了。所以我认为我们现在正在接近那个门槛:模型足够有用,性能足够好,并且能够完成足够多样的任务,从而真正有用。这是一个非常激动人心的时刻。我认为我们刚刚跨过了那个门槛。至于我们可以在多大范围内部署,还有待确定。有些任务失败可能是灾难性的,也许这些还不是部署的最佳任务。有些任务需要大量的泛化,比如在家庭中部署,或者涉及隐私或安全问题。也许这些还不是部署的最佳场所。但我认为随着我们收集更多数据,随着这些模型变得更好,我们可以将它们部署到越来越多的环境中。所以我认为我们正在开始达到那个目标。
I think we are actually fairly close to deploying these robots. We started deploying them ourselves already. We thought this was something that was going to take something like 5 years to get to the point where the technology is actually ready to deploy a robot in a commercial setting and have it do something valuable. But we've done it I think two months ago or something like that. So I think we're now getting to that threshold that the models are useful enough, they're performant enough, and they can do enough variety of tasks to be actually useful. That's a really exciting moment. I think we just crossed that threshold. It's still to be determined how wide the aperture is of where we can deploy. There are some tasks where the failure can be really catastrophic. Maybe these are not the best tasks to deploy just yet. There are some tasks that require a ton of generalization like deploying in homes or that have privacy concerns or safety concerns. Maybe these are not the best places to deploy just yet. But I think the aperture is growing as we collect more data, as these models get better, we can deploy them in more and more settings. So I think we're starting to get there.
你们目前部署的范围是什么?
Where is the current aperture that you're deploying right now?
这是一个很难回答的问题,因为对于这些基础模型,有时你并不完全了解。类似于大型语言模型,你训练这个模型,你在内部精心调试,尽力做到最好,最后你得到一个产物,但你无法真正预测这个产物会有多好。你只能测试它。我们现在的模型也是如此。例如,我们开源它们,这样就不是只有我们在测试,我们也不是了解它们能力的瓶颈。通过开源,我们看到它们被应用于比我们想象的更多的应用,比如驾驶、手术机器人、农业等等。所以我对范围没有一个很好的估计。我认为它比我预期的要广。
This is a really difficult question to answer because with these foundation models sometimes you don't fully know. Similarly to how with large language models, you train this model, you kind of cook it in house, you try to make the best job possible, and then at the very end you get this artifact and you can't really predict how good the artifact is going to be. You kind of have to test it. That's where we are with these models as well. For instance, we open source them so that we are not the only ones testing it and we're not the bottleneck in knowing what their capabilities are. By open sourcing them, we see them being applied to many more applications that we could have imagined, things like driving or surgical robots or agriculture and places like that. So I don't have a very good estimate of what the aperture is. I think it's wider than what I had expected.
而且我认为它也会随着时间的推移而增长。这些模型获得的数据越多,它们就越成熟。我认为范围将继续扩大。
And I think it will also be growing over time. The more data these models get, the more mature they get. I think the aperture will continue to grow.
我想补充一点,在性能层面,正如你所说,这个窗口可能更宽,起点比我们想象的要宽。但同时,如果你真的希望每个应用场景的起点都能达到人们愿意将其作为日常业务驱动的水平,那么在性能方面可能还有相当多的爬坡工作要做。所以,在我们稍后会讨论的这个版本中,我猜是 pi star,我们在从经验数据中学习、反馈并让模型在部署后变得更好方面取得了进展。但对于很多我能天真想象到的事情,仍然会有很多场景存在非常长的尾部分布,可能出错或遇到的情况,我们还没有完全掌握如何解决。
I would add maybe on the performance level, as you said, the aperture is probably wider, the starting point is wider than we thought. But at the same time, of course, if you actually want each of those starting points for each of those applications to be at a level where people would want to use this as a day-to-day driver for their businesses, there's probably still quite a bit of hill climbing to do in terms of performance. So with this release we're going to talk about a bit later, I guess the pi star, we've made progress on learning from experience data, getting that back, and making the models better when they are deployed. It's still for a lot of things that I can naively imagine, there would be lots of scenarios where there's a really long tail of things that can go wrong or that you can encounter, that we don't yet have a great grasp on how to completely solve.
而且你们在发布结果时非常透明,开源发布。所以,在你愿意分享的范围内,你能谈谈你们的整体技术架构吗?你认为达到这个理想境界的架构是否已经基本定型,只是在我们现有主题上做些变化,然后我们需要收集大量数据?还是说架构仍在探索之中?
And you guys have been really great about publishing your results with a lot of transparency, releasing open source. So whatever you're comfortable sharing, can you talk about what your overall technical architecture, so to speak, is? And do you think that the architecture to kind of get to this promised land is pretty much baked and it'll be variations on the theme of where we are and we just need to collect a ton of data? Or do you think that the architecture is still being figured out?
我想我们可以先讨论一下我们现在的位置,然后再深入探讨可能的变化。目前,这个架构与你们日常可能接触到的视觉语言模型(VLM)非常相似,对吧?输入文字,放入图像,然后让它读取图像上的内容等等。我们也是从同样的起点开始的:有一个在互联网规模数据上训练的模型,它吸收了图像数据和文本,然后我们加入了所有机器人数据。实际上,我们的训练现在主要基于机器人数据,是我们自己收集的数据。我们混合了一点点互联网数据,但大部分是机器人数据。这个架构是一种视觉语言模型,我们在旁边添加了一个所谓的动作模型,即动作专家,这部分模型实际上负责驱动机器人,它查看图像和收到的指令,然后执行任务,向机器人发送命令。所以,大致上它是一个 Transformer 模型,目前相当大,有几十亿参数,我们在机器人数据和互联网数据上预训练它。它最初主要从人类演示数据中训练。Carol 之前提到过一点,我们有这种演示数据,即人类通过远程操作让机器人做事的遥操作数据。这就是目前的架构,我们获得的 Scaling(规模扩张)主要来自数据规模的扩大,我们使用的模型类似于 VLM 领域的模型。至于这如何变化,我认为是一个开放性问题。我认为有很多机会为这些模型增加更多能力,我们也在探索。对吧?你可以想象,你可能希望这些模型有更多上下文,你可能希望机器人上安装更多摄像头,模型需要能够使用它们。你可能希望更好地理解物理世界,比如确切知道房间里有什么,什么东西容易碎,什么东西容易移动等等。所以,我认为在这些能力方面还有很多工作要做,架构也会随之改变。如果五六年后再回头看,我们可能会说,‘哦,也许当时我们使用的模型主干,也就是目前来自 VLM 领域的东西,已经变了。也许我们已经向前发展,使用了稍微不同的东西。’我认为这会随着时间演变,但数据和如何将其引入模型的基础可能会保持这样。
I would say we can maybe start with a little bit discussing where we're at now, and then we can go into the details of how that might change. So at the moment, the architecture is very analogous to how VLMs are built that probably most of you interact with on a day-to-day basis, right? Type something in, put the image in, and ask it to read what's on the image, and so on. And we've kind of started from the same standpoint of there's a model that's trained on internet-scale data, and it's ingested image data and text, and we're adding all this robotics data. Our training actually predominantly now is on robotics data, on data that we have collected ourselves. We have a little bit of that internet data in the mix, but the majority of it is robotics data. The architecture is kind of this vision language model, and we add something on the side which is what we call the action model, the action expert, the part of the model that actually has to drive the robot, right? That basically looks at the image and the instruction it's getting and has to perform the task, has to send commands to the robot. So broadly, it's a transformer model that is a fairly large model, up to like some billion parameters at this point, that we pre-train on our robotics data and on internet data. And it is trained largely initially from human demonstration data. Carol mentioned this earlier a little bit, and we have this demonstration data, teleoperated data of humans trying to get the robot to do stuff. So that's the architecture that looks like now, and roughly the scaling we're getting is from scaling our data, and we use models similar to what comes from the VLM world. How that might change, I think, is an open question. I think there's lots of opportunities in adding more capabilities to these models that we're also exploring. Right? You can imagine that you might want more context in these models. You might want more cameras added to the robots that the model then needs to be able to use. You might want to have a better understanding of the physical world in the sense of understanding exactly what's in the room, what can break, what is easily movable, and so on. So there's lots to be done, I think, in those capabilities and also changing the architecture around, and I wouldn't be surprised if in like five or six years we look back and we say, 'Oh, you know, maybe the backbone of the model that we used at the time, which currently comes from this VLM land, has changed. Maybe we've moved on and we use something slightly different.' I think that will evolve over time, but I think the foundation of like the data and how we bring it into the model will probably stay like this.
明白了。我应该把它理解为像素或信号输入,然后动作输出吗?就像一个单一的大神经网络?
Got it. And should I think about it as it's pixels or signals in and then actions out? Is that like a single big neural net?
它是一个大模型。是的。基本上就是图像输入、文本输入、文本输出和动作输出。
It's one big model. Yeah. It's really just basically images in, text in, text out, and actions out at this point. Yeah.
那么,你们是否有独立的移动和操作堆栈?也许现在是时候谈谈机器人学的历史演变、不同的学习浪潮,以及它们如何与你们的堆栈相关。
And are you, I guess, do you have a separate kind of locomotion versus manipulation stack? Maybe this might be a good time to talk about kind of just the historical evolution in robotics and the various different waves of learning and how it pertains to your stack.
是的。很长一段时间,甚至在机器学习出现之前,人们认为机器人学是这样一个问题:如果你投入足够多的人,足够多的工程师,他们可以非常努力地思考,最终写出代码,让机器人做世界上任何事情。人们确实非常努力地尝试过这种方式,但结果发现世界实在太复杂了。对吧?你不可能写出你在现实世界中遇到的每一种情况。所以那行不通。而且,当我们试图解决那个版本的问题时,人们做了他们通常做的事情。他们试图将这个问题分解成更小的子问题。所以,不是解决整个机器人学问题,而是说问题有感知方面、控制方面、规划部分。这几乎发展成了不同的社区。有规划社区、控制社区,他们有自己的会议、自己的问题等等。然后,当我们意识到不可能手工编写所有这些规则时,人们认为我们应该学习它们。我们应该从数据中学习,这似乎是个好主意,对吧?我们也是这样学习的。但结果他们开始分别学习每个组件,这些分解后的组件,通过单独学习。所以你会有一个完全学习的感知层,也许有一个学习的控制层,也许有一个学习的规划器。这取得了一些进展,比以前好。
Yeah. So for a long time, even before learning arrived here, people thought that robotics is one of these problems where you can, if you put enough people on it, enough engineers, they can think really hard about it and eventually write the code that will have the robot do anything in the world. And people have tried really, really hard to do it this way, and then it turned out that the world is just way too complex. Right? Like you can't just write every single case you'll encounter in the real world. So that doesn't work. And also, as we were trying to work on that version of the problem, what ended up happening is people did what they usually do. They try to break down this problem into smaller sub-problems. So rather than working on the full robotics problem, you would say there's a perception aspect of the problem, there's a control aspect of the problem, there's the planning part of the problem. And this almost grew into different communities. There's a planning community, there's controls community, they have their own conferences, their own problems, and all of that. So then as we realized that it's not really possible to handwrite all of these rules, people thought that we should learn them. We should learn them from data, which seems like a really good idea, right? This is how we learn too. But what ended up happening is that they started learning each one of those components, these broken-down components, separately through learning separately. So you would have a perception layer that is fully learned. Maybe you'll have a control layer that is learned. Maybe you'll have a planner that is learned. And that showed some progress. It was better than what we had before.
但结果发现,把这个问题分解成这些子组件恰恰是行不通的。因为当我试图拿起这个杯子时,我不会从感知、规划、控制的角度去思考。我就是直接去做,拿起杯子,整个过程非常自然。所以事实证明,这种流水线方法——预先定义好接口,感知给出物体位置,规划器给出轨迹,控制执行——这些接口才是崩溃的环节。所以我们之前以为我们了解自身运作方式的一切,其实都是错的。
But then it turned out that breaking down this problem into these subcomponents actually is the piece that doesn't work. Because when I try to pick up this glass, I don't think about it in terms of perception, then planning, then control. I just go for it. I just pick up the glass, and it's all very natural. So it turned out that this pipeline approach, where you have these predefined interfaces—perception gives you the position of the object, the planner gives you the trajectory, and control executes it—those interfaces are the pieces that broke down. So everything that we thought we knew about how we work was always wrong.
于是我们进入了下一阶段,我们说,也许从一开始就把这个问题分解就是个坏主意。
So then we arrived at the next stage, where we said, maybe breaking down this problem was a bad idea to begin with.
对。
Yeah.
所以我们就端到端地训练整个系统。把感官输入作为网络的输入,动作作为输出。这就是我们所说的端到端方法,尝试直接从像素到动作,让网络或学习算法自己去弄清楚如何分解成这些不同组件——如果可能的话。
So let's just train the whole thing end to end. We'll take the sensory inputs as input to the network and have actions as the output. That's what we refer to as the end-to-end approach, where you try to go straight from pixels to actions, and we'll have the network figure out, or the learning algorithm figure out, how to split it into these different components if it's even possible.
对。而在做这件事的过程中,我们发现这实际上需要海量数据。而且当需要某种常识时,它常常会失败。通过第一人称动作数据集来收集这些常识非常非常困难,因为你需要体验世界上的每一件事才能做到。
Yeah. And while we were doing that, we figured that it actually requires a ton of data to do this. And often it breaks when it requires some kind of common sense. To gather that common sense through first-person action datasets is really, really hard because you would need to experience every single thing in the world to do this.
对。这就是我们偶然发现视觉-语言-动作模型的地方。我们可以使用那些在互联网数据上预训练过的模型,它们已经对世界如何运作有了相当好的理解。我们可以利用这些知识,这样就不需要亲身体验一切。只需在上面添加一些动作组件,拥有一个共同的世界理解,并将其与如何在世界中实际执行任务连接起来。
Yeah. And that's where we stumbled upon vision-language-action models, where we can use models that were pre-trained on internet data that already have a pretty good understanding of how the world works. We can utilize that knowledge so that we don't need to experience everything firsthand. You can just add some action components on top of it and have a common world understanding and connect it to how to actually perform things in the world.
我明白了。
I see.
这大致就是我们今天所处的阶段。
And that's more or less where we're at today.
我明白了。在 Physical Intelligence,我们还发现了一些其他问题。那么如何扩展这些模型?如何让它们泛化?如何让它们表现得更好?如何让它们移动得更快?如何让它们达到可以开始部署的程度?但我认为我们大体上仍处于这样一个时代:如何从互联网预训练中引入一些常识知识?如何让这些模型非常通用,以便它们能在任何机器人上工作并执行动作?
I see. Now at Physical Intelligence, we figured a few other things. So how do you scale these models? How do you get them to generalize? How do you get them to perform much better? How do you have them move much faster? How do you get them to the point where you can start deploying them? But I think largely we're still in this era of how do you bring some of the common sense knowledge from the internet pre-training? How do you make these models very general so that they can work on any robot and perform motions?
我能问一下推理方面吗?大语言模型领域的推理方面有很多进展。你们是否从作为 VLA 骨干的模型中获得了这些好处?在端到端训练过程中,推理是否作为结果涌现出来?或者,LLM 领域的一些进展是否对你们有益?
Can I ask about reasoning? There's so much happening in the reasoning side of the large language model space. Do you get the benefits of that as part of your VLA backbone? Does reasoning emerge as a consequence of what you're doing as you train these end to end? Or can I think about some of the benefits of what's happening in the LLM world? Do they benefit you or not?
我认为我们今天的模型确实已经在规划动作,不仅仅是即时动作,而是接下来 50 件需要做的事情。比如接下来的 50 个时间步——从某种意义上说,这是一个非常短的时间跨度,50 步意味着大约一两秒。而且它已经在语言空间中将任务分解为子任务。所以当我们要求它“清理厨房”时,它可能选出的第一个子任务是“我需要开到柜台,然后拿起杯子,把杯子放进水槽”。所以它在某种意义上已经具备了这些方面。它将任务分解为子任务,因为它给自己设定子任务,并预测一小段动作的走向。所以其中一些已经存在了。我认为未来可能会有更多。我完全期待所有关于推理的强化学习训练的进展也会进入机器人领域。
I think definitely the models we have today are already planning actions, not just at the immediate action but kind of what are the next 50 things I need to do. So like the next 50 time steps—in some sense it's a very short horizon, 50 steps means like a second or two. And it also additionally decomposes tasks into subtasks in language space already. So when we ask it, 'clean the kitchen,' the first subtask it might pick out is like, 'I have to drive to the counter, then I have to pick up the glass, move the glass into the sink.' So it already has those aspects in some sense. It decomposes tasks into subtasks because it gives itself its own subtask and it predicts a little bit of a horizon of how actions go. So some of it is already there. I think in the future there will probably be more of it. I do totally expect that all the advances on RL training for reasoning will also make their way into robotics.
对。我觉得这很有趣,因为它可能和人们做的数学问题的强化学习有点不同。我认为那些问题对我们人类来说很容易被当作文本问题来思考——你在脑子里用文字思考。“好吧,如果我这样改变这个公式,我会得到这个结果,”等等。而对于物理智能部分,可能不止于此。当你尝试学习一项新运动时,情况会有所不同。例如,我最近开始学打网球,我不会在脑子里想“我现在需要抓住球拍,把它移到这儿,然后做这个挥拍动作”。而是更像你在脑子里思考动作本身。你思考你的身体如何移动。也许你在脑子里以某种方式规划周围物体的轨迹。所以我认为随着时间的推移,我们会看到这些方面更多地融入模型中。
Yeah. I think it's interesting to think about because it's maybe a little different than the RL for math problems that people do. For example, I think those are very easy for us humans to think of as textual problems—you think through them in your head in text. 'Okay, if I change this formula this way, I will get this outcome,' and so on. And I think for the physical intelligence part of it, it will probably be a bit more than that. It's going to be a little bit different when you try to learn a new sport, for example. When I recently started to try to learn how to play tennis, I don't think through in my head, 'I need to now grab the racket, I need to move it here and I need to do this swing.' But it's more like you think through the motion itself. You think about how your body moves. Maybe you plan in some sense trajectories of objects around you in your head. And so those things I think we'll see come into the models more over time.
对,我怀疑随着时间的推移会这样。现在我们处于一个从视觉-语言模型中受益颇多的阶段。我认为很可能这种情况会逆转——我们今天在 LLM 中看到的许多缺点都是因为专注于文本问题,比如数学和编码问题。而我认为机器人技术将提供一条新途径,你需要重新思考如何推理。推理可能应该发生在某种抽象空间中,你可以在文本中推理一点,在图像中推理一点,也许你可以在轨迹中推理,或者在各种不同的空间中推理,以得出答案。机器人技术提供了一个非常好的测试平台,因为它扎根于物理世界。目前还没有那么多数据,所以你需要处理随之而来的一些困难。但我认为它将提供新的发现,然后这些发现会重新应用到 LLM 领域。
Yeah, I suspect that over time. Right now we're in a place where we benefit quite a bit from vision-language models. I think it's very likely that that's going to reverse—that a lot of the shortcomings we see in LLMs today are baked in because we are focused on the text problem, on problems like math and coding. And I think robotics will offer this new avenue where you need to rethink how to think about reasoning. Reasoning should probably happen in some kind of abstract space where you can reason a little bit in text, a little bit in images, maybe you can reason in trajectories or in all kinds of different spaces to arrive at the answer. And robotics provides this really nice test bed where you're grounded in the physical world. There is not that much data yet, so you kind of need to deal with some of the difficulties that come with that. But I think it will provide new findings that will then be reapplied to the LLM world.
说到数据,请给我们一个概念:你们如何衡量已经收集的数据量,以及明年希望收集多少?当然越多越好,但我们说的是什么量级?
Speaking about data, give us a sense of how you measure the magnitude of data you've already collected and how much you would like to collect in the next year. I'm sure more is better, but what is the magnitude we're talking about?
对,数据实际上是相当微妙的。不仅仅是数量的问题。
Yeah, data is one of those things that's actually fairly nuanced. It's not just a matter of quantity.
对。
Yeah.
质量显然很重要,但多样性也很重要。即使你考虑机器人数据的质量或多样性,这些术语的定义也并不严格。例如,如果你用 10 种不同的方式完成同一任务,这算多样数据吗?或者与用 10 种不同杯子收集的数据相比,多样性如何?我认为我们整个社区还没有完全理解:如何表征数据,如何描述多样性,如何描述数据质量,如何使其非常严谨。我们还发现数据的某些方面确实至关重要。例如,如果你想在某个任务上达到一定性能,仅仅增加已有数据的数量是不够的。我们在 Pi Star 0.6 版本中研究了三个不同的任务,很早就注意到,如果继续以相同方式收集更多数据,性能会停滞不前。你不会持续变好。所以你需要找到新的收集方法,或者开始思考什么样的数据能带来更好的性能。这就是强化学习这类方法能真正发挥作用的地方。
Quality obviously matters, but also things like diversity. Even when you think about the quality or diversity of robot data, these are not very strictly defined terms. For example, if you go for the same tasks in 10 different ways, is this diverse data or not? How do you compare it to the diversity of the data if you go for 10 different glasses? This is something I don't think we as a community fully understand: how to characterize the data, how to describe diversity, how to describe the quality of the data, how to make it very rigorous. We're also finding that there are some aspects of the data that really matter. For instance, if you want to get to a certain performance on a task, you're not going to get there by just increasing the quantity of the data you already have. We've been working on these three different tasks for the Pi Star 0.6 release and noticed fairly early on that if we just keep collecting more data the same way we've been collecting so far, the performance plateaus. You're not going to keep getting better. So you need to find either new ways of collecting it or start thinking about what kind of data will result in better performance. This is where things like reinforcement learning can really help.
我们来谈谈强化学习和 Pi Star 0.6。星号是致敬 Q*吗?
Let's talk reinforcement learning and let's talk Pi Star 0.6. Is the star a nod to Q*?
实际上是试图达到策略最优,即策略星号。
Effectively trying to get to policy star, actually optimal.
策略星号。好的,太棒了。你能简单说说你们用 Pi Star 0.6 在做什么,然后我们再深入探讨强化学习对你们领域意味着什么吗?
Policy star. Okay. Wonderful. Can you maybe say a word on what you guys are doing with Pi Star 0.6, and then we can dive into what RL means for your world?
当然。与我们之前讨论的主要区别在于,在此之前,所有机器人基础模型学习都基于演示数据——远程操作数据输入模型,模型只是模仿这些数据。而现在,通过这个新模型 Pi Star 0.6,我们使用强化学习,机器人通过实际运行策略自己收集经验。我们从演示训练的策略开始,然后部署它。我们让机器人尝试解决任务,它还会从人类那里获得奖励信号,并可以接受纠正。人类干预说:‘实际上,这不对,我们换个方式做。’这些数据被收集并反馈回来。模型利用这些数据判断哪些数据应该强化(多做),哪些应该少做,从而随时间自我改进。这是最大的区别。拥有这种真实数据流是缺失的一环,它让我们能够摆脱之前遇到的性能瓶颈。
For sure. The main difference from what we talked about earlier is that up to that point, all of the robotics foundational model learning we had done was basically demonstration data—teleoperated data going into the model. The model was trained to just imitate that data. Now with this new model Pi Star 0.6, we are using RL from experience that the robot collects itself by actually running a policy. We start with an initial policy that is this demonstration-trained policy, then deploy it. We try to have the robot solve the task, and it additionally gets reward signals from humans and can also get corrections. The human intervenes and says, 'Actually, this is not right, let's do this a little differently.' That data gets collected and comes back in. The model uses that data to figure out which data to reinforce—do more of—and which to do less of, and basically improves itself over time. That's the big distinction. Having that stream of real data coming in is the missing piece that allows us to escape the plateau we were otherwise hitting.
在我看来,强化学习就是在奖励信号上爬山。那么,当你在这些特定任务上爬山时,如何确保你在泛化?
In my brain, RL is hill climbing on your reward signal. So how do you make sure you're generalizing as you hill climb on these specific tasks?
对于这个具体问题,我们的思考方式是:你有一个通用模型,它达到的性能并不好。你的首要目标其实不是进一步泛化,而是先解决这个特定任务。所以我们部署它,我们选了三个或四个任务。它必须在任务间泛化,但方法本身必须泛化。当你实际部署并开始强化学习过程时,你真正关心的是确保你搞定这个任务,并且以一种能从不同位置解决它、处理所有会遇到的长尾失败的方式搞定它。从某种角度看,泛化和性能在这里似乎矛盾,你会想:‘等等,你现在只是在做这一个任务。’但归根结底,我们想做的是用同样的方法、同样的流程,部署到每个任务上,把性能提上去,然后我们就能拥有所有这些任务的数据,并把数据带回来。所以从这个意义上说,它们并不矛盾。
The way we're thinking about this for this specific problem is: you have this sort of general model and it achieves some performance that isn't great. Your first goal actually isn't to further generalize; you want to solve this specific task first. So we deploy it and we've picked three or four tasks. It has to generalize across tasks, but the method has to generalize. When you're actually deploying it and trying to start this RL process, you really care about making sure you nail down this task, and nail it down in a way where you can solve it from many different positions and deal with all the long tail of failures you will encounter. In some sense, generalization and performance here may seem at odds when you look at it like, 'Oh wait, but now you're just doing this one task.' But at the end of the day, what we want to do is have the same method, the same process, that deploys to each of these tasks, gets the performance high, and then we can have all of that data across all of these tasks and bring that data back. So in that sense, it's not actually at odds.
你们做了多少强化学习?听起来这是真实世界的强化学习。你能谈谈在仿真和真实环境中分别做了多少强化学习吗?
How much of the RL are you doing? It sounds like this is in real life RL. Can you talk a little bit about the approach to how much RL you're doing in sim versus in real life?
我们采取了非常注重真实世界的方法,而不是使用仿真。当然,我们也在探索仿真作为研究工具。但 Pi Star 0.6 论文中所有的强化学习实际上都是在真实世界的真实系统上完成的。原因是,部署时遇到的长尾失败很难建模。我可以从这次发布的任务中举很多例子,有些失败模式如果你只做仿真,可能根本看不到。例如,有一个任务是搭建盒子——一个实际的部署任务,我们搭建小纸板盒来装巧克力,以便包装和发货。起初搭建盒子很顺利,但后来新一批盒子以平板纸板的形式运来,这些纸板穿孔不完美,粘在一起。机器人抓起它们,放在桌子上试图搭建盒子,结果突然桌子上有两个盒子。这在仿真中不会发生,如果你写了一个好的仿真器,只会得到单个纸板并折叠它们。所以你必须处理这个问题。如果你在仿真中学习一切然后部署,你不会遇到它。
We have taken a quite real-world-first approach as opposed to using sim. We are exploring sim as a research tool, of course. But all the RL we've done for the Pi Star 0.6 paper is actually on real systems in the real world. The reason is that it's really hard to model the long tail of failures you see when you do deployments. I can give you a lot of examples from the tasks we've looked at for this release where there were failure modes that if you had just done a simulation, you might not have seen them. For example, we have one task where you have to build a box—an actual deployment task where we build little cardboard boxes to put chocolate into, so they can be packaged and sent out. Building this box initially worked great, but then new shipments of boxes came in as flattened sheets of cardboard, and these cardboards were not perfectly perforated—they were sticking together. The robot would grab them, put them on the table to build the box, and suddenly have two boxes on the table. This is something that wouldn't happen in sim if you had written a nice simulator where you just get individual cardboards and fold them. So you have to deal with this problem. If you learn everything in sim and then try to deploy, you wouldn't encounter it.
所以我们遇到它,然后我们的方法可以弄清楚,实际上我需要做的是分开这个,然后把第二块移回来,基本上就是搭盒子。我们看到很多 ARL 在仿真中应用并迁移到现实世界的成功案例,尤其是在运动控制方面。
So we encounter it and then our method can figure out that actually what I need to do is separate this and move that second piece back and build the box basically. And we see a lot of successes for ARL being applied in sim and transferred to the real world, especially in locomotion.
是的。
Yeah.
而我们还没有真正看到这类方法在操作任务上取得同样的成功。
And we haven't really seen that kind of success in manipulation for these kind of methods.
对于这类方法。
For these kind of methods.
我认为其中一个原因可能是,对于运动控制,试图四处移动,问题最大的部分似乎是建模自己的身体。所以如果你能弄清楚如何将自己建模为一个机器人,你基本上就快成功了。你可以只做一次这个建模仿真练习,因为你只需要为你自己、为这一个机器人做,然后基本上就完成了。如果你做得非常好,它应该能迁移。
And I think maybe one reason for that is that with locomotion, with trying to move around, it seems that the biggest part of the problem is modeling your own body. So if you can figure out how to model yourself as a robot, you're basically almost there. So you can do this modeling simulation exercise once because you only have to do it for yourself, for this one robot, and then you're basically done. If you do it really, really well, it should transfer.
然而,对于操作任务,问题不在于你如何移动自己的身体,而在于世界如何对它做出反应。你实际上是在改变你周围的世界。
With manipulation, however, the problem is not how you move your own body, it's how the world reacts to it. You're actually changing the world around you.
弄清楚如何将手从 A 点移动到 B 点并不难。难的是弄清楚这如何影响你正在交互的物体。现在问题不再只是建模你自己的机器人。你必须建模整个世界,对吧?就像你可能交互的每一个物体,你能想到的每一个任务。这就是我们看到规模问题的地方。
It's not difficult to figure out how to move your hand from A to B. It's difficult to figure out how this affects the objects you're interacting with. And now the problem is no longer just modeling your own robot. You have to model the entire world, right? Like every single object that you might be interacting with, every single task you can think of. And that's where we see scaling problems.
我认为这就是为什么我们还没有看到这类方法在操作任务中同样有效。
And that's I think why we haven't seen those kind of methods be as effective in manipulation.
Pi Zero 0.6 的结果标题是什么?在你关心的测试中,经过强化学习后,模型达到了什么水平?你认为这对你未来的整体训练方案意味着什么?
What was the headline of the results from Pi Zero 0.6? And where do you see the model get after RL on the test that you cared about, and what do you think that means about your overall training recipe going forward?
是的,所以对我来说,最令人印象深刻的事情,老实说,就是看到这些模型一次运行数小时,从各种不同的失败中恢复,基本上一直持续下去,同时以比我们初始模型快得多的速度完成任务。所以标题数据是,我们在这三个任务上将策略的吞吐量提高了两倍以上。其中一个任务是我已经提到的搭盒子任务,一个是使用真正的工业级意式咖啡机煮咖啡,另一个是叠衣服。
Yeah, so I think for me the most impressive thing honestly for me personally to see was just have these models run for hours at a time, recover from lots of different failures, and basically just keep going, and at the same time do that at a rate that is actually much better than the initial model that we started with. Right. So the headline figures were we increased the throughput of the policies by over 2x on these three tasks. So there's one task was this box building task I already talked about. One was making coffee with an actual kind of industrial scale espresso machine, and the other one was folding laundry.
对于每个任务,我们都成功地将仅从演示中训练的基础策略变得更快,并且使其能够更好地从失败中恢复。所以当你亲眼看到实际运行,如果你去我们的网站,你可以看视频,我们有机器人连续 13 小时煮咖啡,或者连续 4 小时叠衣服等等。亲眼看到这些会改变你对这些模型的看法。
And so for each of them we managed to make the base policy that was trained just from demonstrations much, much faster. And also make it be able to recover from failures much, much better. And so seeing that actually in action, when you sit there, right, we have, if you go to our website, you can look at the videos, we have the robot serve coffee for 13 hours in a row or fold laundry for 4 hours, things like that. Actually seeing that live changes the way you think about these models.
它改变了我的看法,至少我认为,我们实际上可以部署它们,并且不是仅仅展示一次的玩具演示,而是真正地完全执行实际任务。这在机器人领域一直是一个挑战,我认为很多人没有意识到这一点。
It changes the way, at least I think about it, actually being realistic that we can deploy them, and do it in a way where it's not just a toy demo which is shown once, but is actually doing the real thing fully. And that's been really a challenge in robotics that I don't think many people are aware of.
是的。
Yeah.
就像你知道的,你看到很多机器人做酷事的视频,我们也发布这些视频。基本上任何你想让机器人做的事情,可能都已经有机器人做过的视频了。
Like you know, you see so many videos of robots doing cool things, and we post these videos too. There's basically like anything you want a robot to do, there's probably already a video of a robot doing that.
是的。
Yeah.
但你可以拍很多次,你可以一直录直到得到完美的镜头。我认为每个人都会遇到的问题就是这些模型的可靠性。它们的性能如何,它们完成任务的速度有多快,你能实际部署它们多久而不失败。
But you can take as many takes as you want. You can keep on recording until you get the perfect shot. And the problem that I think everybody encounters is the reliability of these models. How performant they are, how fast they can go about the task, how long you can actually deploy them without failure.
我认为这是在现实世界中部署这些模型的最大瓶颈,因为如果它们每两次试验就失败一次,它们就不是真正可部署的,对吧?
And I think this is the biggest bottleneck in terms of deploying these models in the real world, because if they break every other trial, they're not really deployable, right?
我认为这就是 Pi Zero 0.6 版本对我们来说最重要的突破,我们实际上可以开始达到它们可部署的状态。
And this is I think the most important breakthrough for us with this Pi Zero 0.6 release, that we can actually start getting to a place where they are deployable.
是的。
Yeah.
我们在办公室使用这些机器人给我们煮咖啡,或者我们可以把它们交给 PI 的人在家里叠衣服,或者我们可以部署它们,让它们真正地折叠盒子。这真的非常令人兴奋。
Where we use these robots in our office to serve us coffee, or we can give them to people at PI to fold laundry in their home, or we can deploy them and have them fold boxes for real. And that is really, really exciting.
我们应该把你们用强化学习做的事情主要看作是客户部署可靠性的提升吗?比如你现在可以确保在客户现场可靠地部署煮咖啡模型,它会足够快,不会在长时间内失败。所以这更像是客户部署创新,而不是基础能力创新,还是两者兼有?
Should we think about what you guys are doing with reinforcement learning as primarily a customer deployment reliability point? Like you can now make sure that you can reliably deploy the coffee making model on a customer site and it's going to be fast enough, it's not going to fail over long time horizons. So it's more of a customer deployment innovation versus a fundamental capability innovation, or is it both?
我认为两者兼有。我想,卡尔,你之前也说过一点。我认为在某种程度上,我们真正想要的机器人,对吧?你希望在家里有一个机器人,可以洗衣服、洗碗、做饭、开车,还有那些小企业想要的机器人,可能解决他们不想用传统方式自动化的问题,因为太贵了,比如搭巧克力盒。这些都需要机器人可靠。它必须性能好,并且有能力完成在初始训练阶段未见过的任务。
I think it's both. I think, I mean Carl, you said this a little bit earlier. I think to some extent the robots that we really really want, right? The robot that you want at home which can do your laundry, do your dishes, cook for you, drive around, and also the robot that people want in these smaller businesses, maybe solving a real problem that they have that they don't want to automate in a classical way because it's too expensive, like building a chocolate box. Those are things where the robot has to be reliable. It has to be good and it has to have the capability to do a new task that it hasn't seen in initial training stages.
我认为我们假设可以仅仅通过越来越多的人类数据收集、越来越大,这是不现实的。我们会这样做,但你能获得的数据质量和数量以及初始策略的好坏总是有限度的。所以我认为,正如你所说,如果我们想要部署,我们需要这个,但我也认为在接下来的几年里,我预计我们会看到这些部署,并且这些数据实际上会成为预训练的有价值来源,用于改进我们自己的模型。我们将越来越依赖自主数据收集,至少这是我的预测,在接下来的几年里,建立那个数据主体,即我们希望机器人最终完成的所有任务的凸包,这样模型就能吸收这些数据,并擅长执行和插值。
I think it's unrealistic for us to assume that we can just go with more and more human data collection, go bigger bigger bigger. We will do that, but there is always going to be a limit to how good and how much data you can get and how good the initial policy is going to be. So I think it is what you said in terms of if we want deployments we need this, but also I think increasingly over the next years, I expect we will see that we will do these deployments and that data will actually become really valuable as a source for pre-training, for making our models better themselves. And we'll rely more and more on autonomous data collection, is my prediction at least, over the next coming years to kind of build that host of data, that convex hull of all the tasks that we want robots eventually to do, such that the model ingests this and becomes good at doing them and interpolating.
我认为这是一种新能力。
And I think of it as a new capability.
我们至今还没弄清楚如何从自身经验中学习,虽然有很多尝试,但我认为我们还没看到大规模实现,达到能让人信服并部署的程度。
We haven't so far figured out how to learn from your own experience or there's been many attempts but I don't think we've seen it done at scale to the extent that actually shows a convincing result that allows you to deploy something.
是的。这就是为什么这个结果对我们来说非常重要。我们想达到让它们能从自身经验中学习的程度。
Yeah. And this is why this result was really really important to us. We wanted to get to the point where they can learn from their own experience.
是的,因为就像我们学习一样,你可以从看视频、练习和向他人学习中学到一点,但到了某个点,你需要在实际工作中学习。你需要亲自尝试,需要看到你的行动如何影响你真正想达成的目标。
Yeah, because similarly to how we learn, you can learn a little bit from watching videos and practicing and maybe learning from others, but at some point you need to learn on the job. You need to try the thing yourself. You need to see how your actions impact what you actually want to achieve.
是的。然后自己得出结论,并以此方式学习。我认为这是迈向那一步的第一步。
Yeah. And make your own conclusions and try to learn that way. And I think this is the first step towards that.
你让我想起了今年 Rich Sutton 的《经验时代》论文。我觉得它非常深刻。你认为这为你们在机器人领域解锁了某种持续学习吗?这会成为其中的一部分吗?
You're reminding me of the Rich Sutton 'Age of Experience' paper this year. I thought it was very profound. Do you think this unlocks kind of continual learning in robotics for y'all? Will this be part of that?
这取决于人们对持续学习的定义。我认为它肯定比我们过去做的更持续,过去你有一个大的预训练混合集,可能还有一个后训练混合集,然后你坐下来,非常努力地工作,最终得到一个成品,就这样了。
It kind of depends what people mean by continual learning. I think it's definitely more continual than what we've done in the past where you have like a big pre-training mixture and maybe a post-training mixture and you sit down, work really hard, and then come up with an artifact and that's it.
对,成品完成了,没什么能改变它的了。
Right, the artifact is done and there's not much you can do to change it.
现在,这更像是一个活的东西,对吧?我们从类似的过程开始,但然后你部署它,它持续学习,对吧?所以从这个意义上说,它更持续:它尝试新事物,从自身经验中学习,并不断变得更好。
Now, this is a much more of a living thing, right? We start with a process similar to this, but then you deploy it and it keeps on learning, right? So it's much more continual in that sense that it tries new things, it tries to learn from its own experience, and it keeps on getting better.
是的。
Yeah.
现在我认为还有空间让它更持续,比如它能以此方式获得新技能,或者做得更快。它可能在整个过程中进行推理。所以我认为在“在职学习”程度上有一个谱系,这非常有前景,因为它表明你能做到,但我认为我们可以让它变得更好。
Now I think there is still room for it to be much more continual where it can acquire new skills that way or it can be even much faster in doing this. It can probably reason throughout this process. So I think there's a spectrum of how much you can learn on the job and this is really promising because it shows that you can do it, but I think we can make it much much better.
是的,我同意。我会说我们正处于这个过程的起点,对吧?这肯定不是人们传统意义上认为的持续学习,比如数据流然后整个系统转变,最终一路通向 AGI 之类的,但这是第一步。我会说我们正朝着正确的方向前进,还有很多工作要做。而且我要说,即使从这个发布中,我个人也印象深刻,甚至有些震惊,这些模型在吸收你放回数据中的小细节方面有多好。我很惊讶,即使只是人类的纠正,比如压粉——压粉是制作浓缩咖啡的一个特定步骤,对吧?你把咖啡豆放进去,然后必须把咖啡粉压平。所以我们的机器人一开始压得太用力了,因为最初的人类演示只是确保咖啡粉平整以便放入。然后机器人压得非常用力,几乎把自己从桌子上抬起来。我们看着它,心想‘哇,有点过了’。所以仅仅通过 30 到 50 个回合,人类做了一小部分纠正,我们把数据反馈回去,模型实际上开始变得更轻柔,做正确的事情。我对此真的很惊讶,因为你认为这个模型已经在数百万个回合上预训练过了,现在你只做一点小纠正,它居然有效。所以看到这种情况发生,我认为它指向了持续学习这部分,我觉得令人印象深刻。
Yeah, I would agree. I would say we're at the very beginning of this, right? And it's definitely not continuous learning in the classical sense that people would have thought about, like data streams and then the whole thing turns and ultimately leads all the way to AGI or something like this yet, but it's a first step. I would say we're moving in the right direction and there's lots more to be done. And I will say from even this release, I was personally impressed and to some extent shocked how good these models actually are at picking up little things that you put back into the data. I was surprised that even with just human corrections, for example, for tamping—tamping is a specific part of making an espresso, right? You put the beans and you have to tamp down the coffee before. So our robot in the beginning tamped way too hard because it just happened to be the case that the initial human demonstrations were just making sure the coffee grounds are flat so we can put it in. And then the robot was tamping really hard and almost lifting itself off the table. We looked at it and were like, 'Whoa, that's a bit much.' And so with just 30 to 50 episodes, a really small range of corrections that humans did, we fed that data back and the model actually started being much more gentle and doing the correct thing. I was really surprised by that because you think this model has been pre-trained on millions and millions of episodes and now you're just doing a little correction and that actually works. So seeing that happen was a thing that I think is pointing towards this continual learning part which I find impressive.
不过我能问一下吗,我仍然纠结的是泛化。所以,当我学会更好地压粉,这是否让我更擅长折叠盒子?
Can I ask though, the thing I'm still hung up on is generalization. So, as I learn how to tamp better, does that make me better at folding boxes or not?
呃,在这个具体案例中,不会。但机制是相同的,你也可以用它来修复,比如,我面前有两个盒子卡在一起,我需要把它们分开,对吧?因为你可以为压粉部分得到 30 次纠正。你为分开盒子得到 30 次纠正。你为盒子没有整齐折叠得到 30 次纠正。所有这些积累在一起,然后给你带来更泛化的改进,我会这么说。
Uh, in this specific case, no. But the mechanism is the same that you can also employ to fix the, oh, I have two boxes in front of me that are sort of stuck together and I need to pull them apart, right? Because you can get 30 corrections for the tamping part. You get 30 corrections for the pulling boxes apart bit. You get 30 corrections for, oh, this box wasn't neatly folded together. And all of this accumulates together to then give you this more generalized improvement, I would say.
好的。所以这是一个可重复的配方,但它们不一定相互促进。
Okay. So it's a repeatable recipe, but they don't necessarily cross-pollinate.
是的。我的意思是,我预计随着我们扩大规模,如果任务之间有相似的动作,我们可能会看到一些东西从 A 迁移到 B。但在这个阶段,是的,我会说它更像一个重复的配方。
Yeah. I mean, I would expect that as we scale this up, we might see also things actually kind of transfer from A to B if there are motions that are kind of similar across tasks. But at this point, yeah, I would say it's more like a repeated recipe.
是的。我们从预训练中看到了很多泛化,你训练越来越多的任务,越来越多的数据。你会发现任何新任务都更容易上手,或者出现你之前没预料到的零样本任务。而且这不断改进。我们按一定节奏启动预训练运行,每次我们都看到模型变得更好,因为有更多数据输入,我们对预训练过程做了更多改进等等。我还怀疑,随着我们部署越来越多的模型执行各种不同任务,它们也会带回数据。我认为一个我相当确定我们会看到更多泛化的方式来自这个过程:当你部署这些模型时,数据回来,模型变得更好,你可以部署更多,然后模型变得更好,你可以部署更多,如此循环。
Yeah. And we see a lot of generalization from pre-training where you train on more and more tasks, more and more data. You see that it's much easier to onboard any new task or you see tasks that appear zero shot that you didn't expect before. And this keeps on improving. We kick off a pre-training run at a certain cadence and every single time we start seeing that the model keeps on getting better because there's more data being fed in, there's more improvements that we're making to the pre-training process and so on. I also suspect that as we have more and more of these models deployed doing all kinds of different tasks, they also bring data back in. And I think one way where I'm quite certain we'll see more generalization is from that process: as you deploy these models, the data comes back, the models get better, you can deploy them more, then the models get better, you can deploy them more, and so on.
是的。我认为你提出的这一点值得讨论。我们还没有真正谈到这个配方的一个关键细节方面,那就是模型有两个部分。一个是策略,它试图通过纠正和强化学习反馈来改进。
Yeah. And I think maybe it's worthwhile for this point that you brought up. We haven't really talked about one crucial detail aspect of this recipe, which is that the model has kind of two parts. One is the policy that is trying to improve, via corrections and RL feedback.
另一个问题是如何真正获得强化学习反馈。我们之前聊到过人类可能会纠正,那是人类纠正的部分。强化学习反馈部分则有些不同,它已经包含了一些我认为你在寻找的泛化特性。我们的做法是:首先让人类告诉我们,某个尝试——比如做咖啡或整理盒子——是否成功。这些片段会有人类标注,然后我们训练一个所谓的价值函数,来预测从当前任务状态出发,我是否可能成功或失败。这个价值函数被用作基线,来决定对于这个数据点,我是应该上调还是下调,取决于我预期自己是更接近成功还是更可能失败。
And the other part is how do you actually get this RL feedback? So we've talked a little bit, I've mentioned humans might correct, and that's the human correction part. The RL feedback part is a little different and already has some aspects of generalization that I think you're searching for. The way we do this is we first get humans to tell us whether a specific attempt of making the coffee or doing the box was successful or not. So there will be human labels provided with these episodes, and then we train something called a value function to try and predict, from my given point in the task, whether I will likely succeed or fail. This value function is then used as a baseline to decide whether for this data point I should bump that up or bump that down, depending on whether I expect that I will be moving towards success or more likely towards failure.
我们在训练这些价值函数时发现——它们基于相同的骨干网络和模型,但在实际执行任务的策略之前预训练——添加更多来自不同任务的数据确实有帮助。模型开始变得非常擅长,至少在特定任务上,能在我明显察觉之前就预知失败。例如,当我观看它尝试将咖啡手柄插入咖啡机的视频时,它似乎知道角度不对,在失败发生前 30 到 40 步,价值函数的预测就会下降,并表明:‘这个片段不太好,我不应该采用这个数据。’
One thing we saw when we trained these value functions—they are trained from the same kind of backbone on the same kind of model, but pre-trained before the actual policy that runs the task—is that adding more data from different tasks actually helps. The model starts being really quite good, at least for certain tasks, at knowing when it will fail beforehand, before it is obvious to me. For example, when I look at a video of it trying to insert the porter filter into the coffee machine, it kind of knows that it doesn't have the right angle before that happens. So 30 or 40 steps before that actually happens, the value function's prediction drops and says, 'Oh, this is not good in this specific episode, so I shouldn't include this data.'
有意思。这跟 Karpathy 那种‘用吸管吸比特’的说法形成了有趣的对比,对吧?因为你并不是在等待最后那一点信息,而是在过程中就获得了大量信号。
Interesting. And so this is an interesting counterpoint to the Karpathy-like 'slurping bits from a straw' thing, right? Because you're not waiting for that final bit at the end. You're actually getting a lot of signal along the way.
我认为强化学习是一个广阔的领域,有很多不同的方法。人们常常将强化学习与策略梯度方法或非常特定的在策略学习方法联系起来。但强化学习更像是一个问题定义,有很多方法可以绕过你提到的那个问题——即只在最后获得奖励,这对于非常长周期的任务来说不可扩展。比如价值函数和时间差分学习,它们通过持续进行顺序预测来绕过这个问题。
I think RL is such a vast field with many different approaches. People often associate RL with policy gradient methods or very specific on-policy learning approaches. But RL is more of a problem definition, and there are many approaches that get around the problem you're referring to—that you only get the reward at the very end, which isn't scalable for very long horizon tasks. There are things like value functions and temporal difference learning that try to get around this problem by constantly making predictions in a sequential way.
这可能是另一个我认为机器人技术能真正帮助更广泛 AI 社区的地方,因为我们没有完美的语言模拟器可以随意运行无数模拟。相反,我们必须在现实世界中操作,所以需要更高效的方法。因此我们需要学习价值函数之类的东西,我认为这些方法在任何地方都会非常有价值。
And this is maybe another one of those things where I think robotics can really help the broader AI community, because we don't have the advantage of a perfect language simulator where you can run as many simulations as you like. Instead, you need to do it in the real world, so you need more efficient methods. Therefore we need to learn value functions and things like this, and I think these will be really valuable everywhere.
好的。我能再追问一下吗?我很想了解:互联网视频似乎是配方的一部分,但据我观察目前并不是重点。你认为互联网视频中还有金矿可挖吗?另外,看看现在视频模型——世界模型——的发展,你认为这会在多大程度上成为模型能力的非连续跃升,以及你模型流程中的重要组成部分?
Yeah. Can I push a little bit? I'd love to understand: internet video seems like it's part of the recipe but not a huge focus right now as I see it. Do you think that there's gold left to be mined in internet video? And then, if you look at what's happening in video models right now—world models—to what extent do you think that's going to be a discontinuous jump in model capabilities and an important part of your model pipeline?
嗯,我认为这里有两个问题。一个是关于数据——如何自举到可以开始部署的程度。另一个是关于视频模型和世界模型方面。在数据方面,我认为我们现在处于自举阶段,基本上什么方法都可以尝试。无论你能想到什么方式给模型带来好处,我觉得都是好的——无论是添加模拟数据、人类视频、某种手持设备数据、还是人类遥操作数据。这其实不重要;你只需要找到某种方式自举到可以部署这些模型的程度。因为我认为长期来看,会有这个自举阶段,但之后会是部署阶段,而部署阶段提供的数据将远远超过自举阶段所能获得的任何数据。所以我们目前处于一个奇怪的阶段,尝试很多不同的东西,看看哪些有效,只是为了达到部署的门槛。一旦能够部署,我认为那将远远超过之前能做的一切。这也是我们正在冲刺的目标。这就是为什么我们想要开始部署这些模型,在多种不同的环境中执行多种不同的任务,这样我们就能拥有一个非常强大的数据引擎。
Yeah, I think there are two questions there. One is about the data—how do you bootstrap yourself to the point where you can start deploying. The other question is about video models and the world model aspects. On the data point, I think we are now in this bootstrap phase where basically anything goes. Whatever you can figure out how to add to the model to its benefit, I think it's good—whether you can add sim, human videos, some kind of handheld devices, human teleoperations. It kind of doesn't matter; you just need to figure out some way to bootstrap yourself to the point where you can deploy these models. Because I think in the long term, there's going to be this bootstrap phase, but then there's going to be the deployment phase, and I think the deployment phase will provide much more data than anything you could do in the bootstrap phase. So we're in this kind of weird spot right now where we try many different things to see what sticks, just to get us to the deployment threshold. And once you can deploy, I think that will vastly be greater than anything you can do before. So that's also what we are sprinting towards. That's why we want to start deploying these models, with many different tasks in many different environments, so that we can have a very powerful data engine.
在世界模型方面,我认为世界模型和我们的方法都针对同一个问题:反事实问题,或者说信用分配问题。如何找出哪些行动真正对成功至关重要,以及如果采取了不同的行动,世界会如何演变?一种方法是通过预测可能发生的情况——生成完整的视频,比如如果我把这个咖啡手柄稍微放得不同,最终会怎样,是失败还是成功?或者你可以通过强化学习来实现,它通过稍微不同的机制,更隐式地处理,但根本上针对的是非常相似的问题。我们正在探索所有这些方法,试图找到真正解决反事实问题的方法。我认为目前还没有答案,但我们从强化学习中看到了很多进展,比如我们刚刚展示的 Pi Star 和 Pi Star 6。但我认为可能还有很多其他方法的空间。
Now, on the world modeling side, I think world models and our approaches are kind of targeting the same problem: the problem of counterfactuals, or a credit assignment problem. How do you figure out which actions were the ones that actually matter for your success, and how would the world have evolved had you taken a different action? One way you can do this is by predicting what would have happened—rolling out a full video of, if I put this porta filter a little bit differently, where would I end up, and would this be a failure or a success? Or you can do this through reinforcement learning, which does it through a slightly different mechanism, a little more implicitly, but it fundamentally targets a very similar problem. We are exploring all of those approaches and trying to see how to really solve the counterfactual problem. I don't think there is an answer yet, but we see a lot of progress with reinforcement learning that we've just shown with Pi Star, with Pi Star 6. But I think there is probably room for many other approaches too.
太棒了。我们能不能聊聊,一旦你们度过了那个启动阶段?我们来谈谈客户部署。你们给客户提供什么?你们卖给他们什么?然后你们想象这会如何随时间演变?比如,你们是卖给他们一个完全垂直整合的机器人解决方案?还是卖给他们一个模型,让他们自己想办法整合到运营中?这一切是怎么运作的?
Awesome. Can we talk about once you guys get past that bootstrap phase? Let's talk about customer deployments a little bit. What do you bring to a customer? What do you sell them? And then how do you imagine that's going to evolve over time? Like are you selling them a fully vertically integrated robotic solution? Are you selling them a model that they have to figure out how to integrate into their operations? How does this all work?
真正的答案是,我们还不知道。
The real answer is we don't know yet.
嗯。
Yeah.
我们还在摸索中。
We are still figuring that out.
嗯。
Yeah.
我们在技术上还处于非常早期的阶段,你可以看出来。我们才刚刚开始达到可以部署这些东西的门槛。所以我们相信应该先专注于技术,想办法让它变得容易部署,并扩大我们最初讨论的这个范围。机器人技术——机器人初创公司的历史通常是,你花一段时间开发技术。你从一个宏大的愿景开始,想象它能实现什么,它会有多通用,但一旦你选了一个要应用的具体场景,你就卡住了。你开始走捷径。你开始为这个应用找出非常专门的解决方案,很快你就变成了一家只专注于比如仓库拣选和放置机器人的应用公司,仅此而已。我们真的想避免那种未来。我们认为我们有机会真正解决物理智能,而这样做的收益将远远超过我们现在能专注的任何单一应用。所以我们希望确保技术尽可能通用,尽可能容易部署。这个范围尽可能宽,然后我们才会开始想办法商业化。就像你说的,可能有很多不同的方式。可能有些方式我们现在还想不到,因为它们取决于技术的发展。无论是成为模型提供商、完全垂直的解决方案,还是销售机器人或其他什么。但我认为现在回答这个问题还为时过早。这会让你很安心,你知道,就像选我们中的一个。
We are still quite early in the technology, as you can tell. We are just starting to even get to the threshold where we can start deploying these things. So we believe we should focus on the technology first to figure out how to get it to the point where it's actually easy to deploy and expand this aperture that we're talking about initially. And robotics—the history of robotics startups is very often that you develop a technology for some period of time. You start with this grand vision of what it should be able to enable, how general purpose it will be, and as soon as you pick an application that you want to apply it to, you're kind of stuck. You start cutting corners. You start figuring out very special purpose solutions just for this application, and very quickly you become an application company that just focuses on, let's say, warehouse pick and place robots and that's it. We really want to avoid that future. We think we have a chance to really solve physical intelligence, and the benefits of doing this will far outweigh any single applications that we can focus on now. So we want to make sure that the technology is as general as possible, as easily deployable as possible. This aperture is as wide as possible, and then we'll start figuring out how to commercialize it. And as you said, there could be many different ways of doing this. There are probably ways that we can't think of just yet because they will depend on how the technology goes. Whether it's being a model provider, a fully vertical solution, or you sell robots or whatever else. But I think it's a little too premature to answer this question. It will give you a lot of comfort, you know, just to like pick one of us.
让阿尔弗雷德很安心。
Give Alfred a lot of comfort.
是的,阿尔弗雷德会对我们满意。但我认为现在太早了。
Yeah, Alfred will be happy with us. But I think it's just too early.
不,你们有一个宏大的愿景。所以,感谢你们研究物理智能。这对 Pi Star06 来说是一个美妙的进步。这是一个巨大的突破,祝贺你们取得的成功。
No, you guys have a grand vision. So, thank you for working on physical intelligence. It's a wonderful improvement just for Pi Star06. It's just a huge sort of breakthrough and so congratulations on all the success you've had.
谢谢。我能接着问一个尖锐的问题吗?
Thank you. Can I follow up with a spicy question?
当然。
Sure.
所以,正如你所说,这个愿景如此宏大,如此广阔,你们在做所有这些不同的事情。我相信你们研究过所有之前的机器人技术努力,它们大多像你说的,一个应用接一个应用,变得越来越窄。大型应用中最成功的案例之一是自动驾驶,Waymo 或特斯拉做得非常好。但如果我要回顾历史,我是在 Sebastian Thrun 站在 TED 舞台上时了解到自动驾驶的,我想是 2009 年或 2010 年,他谈到了他们在 2007 年赢得 DARPA 挑战赛的事情。而现在是 2025 年,这东西几乎只能从旧金山开到这儿。他们现在勉强能做到,但走的是地方道路,甚至上不了高速公路。如果你做这样一个通用化的工作,你认为实现通用化和性能的跑道或时间线有多长?
So, as you said, this vision is so grand, so broad, you're doing all these different things. I'm sure you've studied all previous robotics efforts and they've largely, as you said, applied an application to an application and they get narrower and narrower. One of the most successful cases of a large application is self-driving, and Waymo or Tesla have done enormously well. But if I had to go back in history, I learned about self-driving when Sebastian Thrun was on the stage of TED in, I think, 2009 or 2010, and he talked about the thing where they won the DARPA challenge that was 2007. And we're in 2025 and the thing barely goes from San Francisco down here. They kind of can do it now, but they take local roads. They can't even get on the freeway. If you do such a generalized job, how long is the runway or the timeline that you're thinking about to build for generalization and performance?
是的。所以,这个问题有些方面比自动驾驶更容易,有些则更难。让它更容易的一点是,我们不需要等到它 100% 可靠才部署,对吧?有很多任务,即使你只有 95% 的可靠性,也完全没问题。如果你家里有一个机器人帮你叠衣服,每 100 件中有一次没叠好,你完全不会在意。
Yeah. So, there are some aspects of the problem that make it easier than self-driving and some that make it harder. One thing that makes it easier is that we don't need to deploy it only when it's 100% reliable, right? There are many tasks out there that even if you're at 95% reliability, you're totally fine. If you have a robot in your home folding your laundry and every 100 items, it doesn't fold it perfectly, you'll be totally fine.
你只需叫你的孩子去叠……
You just call your child to go fold the...
没错。我们仍然需要家务。
That's right. We still need chores.
是的,没错。
Yeah, exactly.
而对于自动驾驶来说,情况并非如此,对吧?如果你每 100 次中灾难性地失败一次,那是个大问题。所以我认为在部署这项技术方面,它可能更容易。现在我们还受益于这是一个不同的技术时代。我们处于视觉语言模型、具有常识的基础模型的时代,我们在 2009 年到 2025 年之间学到了很多经验教训,我们可以利用所有这些。所以我认为这也非常有帮助,而且这些解决方案比我们过去拥有的要通用得多。与此同时,有些事情会非常具有挑战性,对吧?不只是单一应用。这是一个非常通用的解决方案,可以应用于驾驶,也可以应用于操作、移动、飞行以及各种其他事情。我认为这有多难还有待观察。到目前为止,根据我们的经验,老实说,它似乎并没有那么难。似乎如果你从一开始就以非常通用的心态来处理这个问题,结果发现它可以相当好地泛化。而且物理智能中有一些我们尚未完全理解的东西,使得这些模型能够在驾驶、煮咖啡、驾驶无人机和操作手术机器人之间进行泛化,尽管它们看起来彼此相距甚远。而且看起来这些应该是不同的模型和不同的应用,但这些模型不知何故能够理解所有这些数据。这给了我很大的希望,也许这个问题并没有那么难,实际上可能更容易。所以我认为这是一个合理的问题,但我也不想从自动驾驶中得出错误的结论。
And with self-driving, that's not the case, right? If you fail every hundredth time catastrophically, that's a big problem. So I think in terms of deploying this technology, it might be easier. Now we also benefit from the fact that this is a different era of technology. We are at the era of vision language models, of foundation models that have some common sense, and we learn a lot of lessons between what was it 2009 and 2025, and we can benefit from all of those. So I think that also really helps, and these are much more general purpose solutions than what we had in the past. At the same time, there are some things that will be very challenging, right? There isn't just a single application. This is a very general purpose solution that can be applied to driving but also to manipulation and locomotion and flying and all kinds of other things. And I think it's to be seen how much harder this is. So far, based on what we've experienced, it doesn't seem to be that much harder to be honest. It seems that if you tackle this with a very general purpose kind of mindset from the get-go, it turns out that it can generalize fairly well. And there is something about physical intelligence that we don't fully understand that allows these models to generalize between driving and making coffee and flying a drone and operating a surgical robot, even though they seem so far apart from each other. And it seems that these should be all different models and different applications, but these models somehow can make sense out of all of that data. And that gives me a lot of hope that maybe the problem is not that much harder and it might be actually easier. So I think it's a fair question, but I also don't want to draw the wrong conclusions from what we've seen from self-driving.
太棒了。恭喜。除了你们自己的成果,什么结果给你印象最深?这是个好问题。
That's beautiful. Congratulations. What result impressed you the most outside of results you think? That's a great question.
是的,这确实是个好问题。
Yeah, it's a good question actually.
我先来。视频模型给我留下了深刻印象,就是你之前提到的。我几年前见过它们。我几年前研究过它们的一些方面,我没想到这个改进轨迹会如此陡峭。它们现在基本上与现实无法区分,而且能做出令人难以置信的事情。所以这真的让我印象深刻,也让我非常惊讶。
I can start. I've been really impressed by the video models, what you mentioned earlier. I saw them a few years ago. I worked on aspects of them a few years ago, and I didn't expect this trajectory to be so steep. They're basically indistinguishable right now from reality and they can do incredible things. So that's been really impressive and really surprising to me.
是的,我仍然有些惊叹,我们竟然走到了这一步,似乎从单纯的下一个词预测中就能得到具有通用智能的模型,这程度我当初真没预料到。我至今仍感到惊奇,每一项小小的进步,比如赢得国际数学奥赛挑战,或者应用于科学发现新事物,都让我惊叹。今年有很多事情让我觉得,哇,尽管年初时感觉大语言模型的预训练可能有点后劲不足,但进步空间依然很大。是的,意识到这几乎是第二波新鲜空气的到来。
Yeah, I would say I'm still in awe to some extent that we've gotten to this place where we do seem to get models that do seem generally intelligent to a level that I really didn't foresee coming out of just next token prediction. I'm still amazed with this and like every little advance that I see, you know, winning IMO math challenges or applying it to finding new stuff in science to me. Yeah, there are so many things this year where I thought like, wow, there's still a lot of progress to be made even though it felt like at the beginning of the year maybe this whole pre-training business of LLMs is kind of petering out a bit. Yeah, realizing that there's like this whole almost second breath of fresh air basically coming in.
是的,我想补充一点,就是整个这东西居然能行,这太令人震撼了。
Yeah, I would maybe add to this just like the fact that this whole thing works, it's kind of mind-blowing.
是的,我觉得我们并没有完全意识到这有多离谱,对吧?你构建了一个松散地受大脑启发的东西,它有一个非常通用的学习算法。你给它数据,它不知怎么就学会了,而且学得比我们以前拥有的任何东西都好。这适用于机器人,也适用于视觉、语言、声音以及各种其他领域。如果你停下来想一想它是如何工作的,以及它居然能工作,那绝对令人震撼。比如,我们能有机器人,把它放在一个家里,它就知道在一个从未去过的家里该做什么,或者它能连续 13 小时煮咖啡之类的。而这都来自这个非常通用的东西,它完全端到端训练,我们并不完全理解它,但它似乎开始掌握了。这对我来说就是令人震撼。
Yeah, I don't think we fully realize how ridiculous this is, right? You build this loosely brain inspired thing that has a very general purpose learning algorithm. You feed it data and it somehow gets it and gets it way better than anything we've ever had before. And this applies to robots and it applies to vision and language and sound and all kinds of other things. And I think if you stop for a second and just think about how it works and that it works, it's just absolutely mind-blowing. Like the fact that we can have robots, you can put it in a home and it kind of knows what to do in a home that it's never been to before or it can make coffee for 13 hours straight or things like that. And this is from this very general purpose thing that trains fully end to end that we don't fully understand, but it seems to start to get it. That to me is just mind-blowing.
我们活在模拟中。这是索尼娅相信的,我们活在模拟中。但这很有趣,对吧?在科学中,他们教你把一个大的问题分解成越来越小的问题。然后基本上有人意识到,这可能不是训练机器或任何类型机器人的最佳方式。
We're in a simulation. It's what Sonia believes that we're living in a simulation. But it is interesting, right? Like in science, they teach you to take a big problem and break it up into smaller and smaller problems. And then basically somebody realized that that's maybe not the best way to train machines or robots of any kind.
老实说,整个机器学习和人工智能领域在某种程度上也犯了同样的错误,对吧?很长一段时间里,人们都在非常深入地解决单个问题,对吧?然后随着时间的推移,出现了一种观念:哦,如果我们能把所有东西放在一起,比如做多任务学习,如果我们能做得非常好,我们就会做得好得多。但后来,这一切之所以发生,仅仅是因为我们转向了这个通用的预训练目标,然后一切就自然而然地出现了,这才是令人惊讶的部分,对吧?
And to be honest, the whole machine learning AI field made that same mistake actually to some extent, right? We were working for a long time, people were working on solving individual problems very deeply basically, right? And then over time there is this notion of oh if we can put it all together like do multitask learning if we could do that really really well we do much better and then but then the fact that that all happened just because we switched to this general pre-training objective and then it just all falls out that's the part that is the surprising bit right.
你认为这像手风琴一样吗?我们从一个框架到另一个框架,把大问题分解成越来越小的问题,这些方法在一段时间内有效,然后失效了,我们就说:‘好吧,让我们回到大问题,尝试更一般地解决它’,然后来回切换。
Do you think it's like an accordion where we go from one framework to the other framework we take big problems break them up into smaller and smaller ones that work for a period of time then it stopped working and we're like, 'All right, let's go back to the big problem and try to solve it more generally and go back and forth.'
我不认为我们会回去。
I don't see us going back.
是的,我不认为我们会回去。我认为有很多方法,或者很多人说你需要两全其美,你需要某种方式融入我们已经知道的规则,比如牛顿物理学。你不需要学习那个。我们已经知道它是如何工作的。所以,你能把它放到权重里吗?但根据我们目前所见,这行不通。如果你试图这样做,你会在某种程度上限制学习新事物的能力。我不认为有两全其美。我认为我们只能一路学习到底。而且有趣的是,这与我们学习的方式相似。你会想,如果有办法预先烘焙所有智能,进化早就想出来了。你生来就会知道一切。我们在其他物种身上也看到了这一点,对吧?比如鹿,它们出生时基本上就是它们一生中最聪明的状态了。它们一生中并没有学到太多东西。但对于像人类这样聪明的物种,还有乌鸦,它们有童年期、青春期,一开始并不聪明,但必须从自己的经验中学习,而不是预先设定好的。你必须自己去争取。我认为这其中有道理。你需要体验世界并从中学习。我认为这也是我们在机器学习和人工智能中学到的教训:我们以为我们知道我们是如何思考的,但实际上我们并不知道。我们只需要让算法从数据中学习。
Yeah, I don't see us going back. I think there's a lot of approaches or a lot of people saying that you need the best of both worlds and you need some kind of way of incorporating the rules that we already know about like Newtonian physics. You don't need to learn that. We already know how it works. So, can you just like put it somehow into the weights? But, from what we've seen so far, it doesn't work. If you try to do this, you kind of limit the ability to learn new things. And I don't think there's the best of both worlds. I think we just go all the way learning. And it's kind of interesting to, you know, how similarly to how we learn. You would think that if there was a way to pre-bake all of the intelligence, the evolution would have figured this out. You would have just been born knowing everything there is to know. And we see this with some other species, right? Like I think deer when they get born they're basically as smart as they will ever be. They don't really learn much throughout their lifetime. But for intelligent species like humans but also I think crows for instance they have these childhood periods, the adolescence period where they're not very smart to begin with but they have to learn from their own experience and it doesn't come pre-baked. You kind of have to earn it on your own. And I think there is something to that. You need to just experience the world and learn from that. And I think that's the lesson we're learning in machine learning as well in AI that we think we know how we think but we actually don't. And we just need to let the algorithm learn it from data.
养孩子也一样。我以为我知道我儿子在想什么,但其实我不知道。
Same thing with raising a child. I think I know how my son is thinking but I don't.
是的。我有一个小女儿,是的,太令人惊讶了,她们学得这么快。
Yeah. I have a small daughter and yeah, it's just so surprising like they learn so fast.
她们学得这么快,你不知道她们是从哪里学来的。
They learn so fast and you don't know where they get it from.
希望是从父母那里。
Hopefully from the parents.
希望如此。
Hopefully.
她肯定知道一些他们没有教她的东西。
She definitely knows some things that they didn't teach her.
非常感谢你们。你们正在构建的使命真的很美好。感谢你们来分享。
Thank you guys so much. It's a really beautiful mission you're building after. Thank you for coming to share.
谢谢。感谢邀请我们。感谢邀请我们。
Thank you. Thanks for having us. Thanks for having us.