开发通用机器人:从数据到物理智能

Developing General Purpose Robots: From Data to Physical Intelligence

切尔西·芬恩 Chelsea Finn · Y Combinator · 2025-07-22 · 约 45 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

探讨如何构建能在任何环境中执行任何任务的通用机器人,从语言模型和大规模数据收集中汲取经验。

Exploring how to build general-purpose robots that can perform any task in any environment, drawing lessons from language models and large-scale data collection.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 25)

全文 · Full transcript(中英对照)

引言:机器人技术的问题 Introduction: The Problem with Robotics

Chelsea

大家好。我非常兴奋能谈论开发通用机器人,以及我们如何真正开发并将智能带入物理世界。首先,我想谈谈这个问题:如果你想真正解决一个机器人应用,你基本上需要围绕这个应用建立一整个公司。你需要为物流、湿实验室自动化、厨房机器人、手术机器人等建立不同的公司。这非常困难,因为那家公司需要制造新硬件、开发定制软件、为该应用设计独特的运动基元、处理边缘情况等等。如果你想解决一个机器人问题,你必须从头开始做所有这些。结果,很多机器人公司在将机器人成功带入我们的日常生活方面并不成功。

Hi everyone. I'm really excited to talk about developing general purpose robots and how we might truly develop and bring intelligence into the physical world. To start off, I'd like to talk about this problem: if you want to truly solve a robotics application, you essentially need to build an entire company around that application. You need to build a different company for logistics, for wet lab automation, for robots in kitchens, for surgical robots, and so on. This is really hard because that company needs to make new hardware, develop custom software, design unique movement primitives for that application, handle edge cases, and so on. You have to do all of that from scratch if you want to solve a robotics problem. As a result, a lot of robotics companies haven't been very successful in actually bringing robots into the physical world in our daily lives.

Chelsea

我联合创立了一家名为 Physical Intelligence 的公司,试图解决这个问题。具体来说,我们正在开发一个通用模型,可以让任何机器人在任何环境中完成任何任务。我们认为这种通才模型可能比专用模型效果更好、更容易使用,就像我们在语言和其他应用的基础模型开发中看到的那样。例如,如果你想构建一个编程助手,你不会专门为编程开发一些东西,而是基于在大量数据上训练的模型,而不仅仅是代码。本质上,这就是试图开发这些基础模型,并将这种智能带入物理世界,而不是它们目前主要所在的数字世界。

I co-founded a company called Physical Intelligence that's trying to solve this problem. In particular, we're trying to develop a general purpose model that can enable any robot to do any task in any environment. We think this sort of generalist model may work better and be easier to use than purpose-built models, just like we've seen in the development of foundation models for language and other applications. For example, if you want to build a coding assistant, you don't develop something specifically for coding; you build on models trained on large amounts of data, not just on code. Essentially, this is the problem of trying to develop these foundation models and bring this intelligence into the physical world rather than the digital world where they largely are today.

规模与数据源的作用 The Role of Scale and Data Sources

Chelsea

那么我们如何做到这一点?在这次演讲中,我想谈谈我们如何着手做这件事。如果我们从语言模型中吸取教训,我们知道语言模型教会了我们 Scaling(规模扩张)的重要性。一个可能的结论是,Scaling(规模扩张)是开发这些模型的最重要因素。如果这个结论成立,你可能会寻找某些大规模数据源。例如,我们可以看看工业自动化的数据,这些数据提供了大量机器人反复执行任务的数据。但这种数据不会让机器人进入灾区、做三明治或装杂货。这种大规模缺乏解决这个通用问题所需的行为多样性。

So how do we do this? In this talk, I'd like to talk about how we go about doing this. If we take a lesson from language models, we know that language models have taught us the importance of scale. One possible conclusion is that scale is the most important ingredient for developing these models. If this conclusion is true, you might look to certain data sources for large-scale data. For example, we might look at data from industrial automation, which gives tons of data of robots doing tasks over and over again. But this sort of data isn't going to allow robots to go into disaster zones, make a sandwich, or bag groceries. This massive scale doesn't have the diversity of behaviors we need to solve this general problem.

Chelsea

或者,也许我们可以看看 YouTube 的数据,这也是一个庞大的数据源,有许多人类执行任务的视频,可能对训练机器人有用。但与此同时,我们不会通过看别人写字来学习写字,也不会通过看温网来成为网球专家。即使这里有大规模数据,使用起来也非常具有挑战性,而且机器人和人类的具身之间存在差距。最后,我们可能看看模拟数据,它也可以提供大规模数据,但这些数据缺乏真实感,并且与现实存在差距。我认为这里的教训是,Scaling(规模扩张)对于开发这些能在开放世界条件下泛化的模型是必要的,但它从属于实际解决问题。所以你需要规模,但它不足以解决整个问题。

Alternatively, maybe we look at data from YouTube, which is also a massive data source with many videos of humans doing tasks that could be useful for training robots. But at the same time, we don't learn how to write by watching other people write, and we don't become expert tennis players by watching Wimbledon. Even though there's massive scale here, it's very challenging to use, and there's a gap between the embodiment of robots and humans. Lastly, we might look at data from simulation, which can also provide massive scale, but this data lacks realism and has a gap from reality. I think the lesson here is that scale is necessary for developing these models that can generalize in open world conditions, but it's subordinate to actually solving the problem. So you need scale, but it's not sufficient for the entire problem.

物理智能方法:真实机器人数据 Physical Intelligence's Approach: Real Robot Data

Chelsea

在 Physical Intelligence,我们一直在收集数据片段。这是一个例子,为了纪念几个月前我们的一周年。这里你可以看到一位远程操作员亲自操作一些引导臂来控制机器人划火柴并点燃蜡烛。通过这类数据,我们可以训练机器人执行各种任务。我想谈谈我们最近在尝试用大规模真实机器人数据开发物理智能方面的一些成果。我要说明,按照今天的机器人标准,这算是大规模,但与我们未来几年应该拥有的数据量相比,可以说是微不足道。具体来说,我们将研究机器人是否能完成各种灵巧的长周期任务,机器人是否能在从未去过的地方成功,机器人是否能响应开放式提示和插话。即使你对机器人技术不感兴趣,我认为我们试图解决这些问题所获得的经验教训也适用于物理世界之外。

At Physical Intelligence, we've been collecting data episodes. This is an example in honor of our first anniversary a few months ago. Here you can see a teleoperator in person operating some leader arms to control the robot to light a match and light a candle. With this sort of data, we can train robots to do a variety of tasks. I'd like to talk about some of our recent results in trying to develop physical intelligence with large-scale real robot data. I should mention this is large scale by today's robot standards and arguably a minuscule amount of data compared to what we should have in the years to come. In particular, we'll be looking at whether robots can do a variety of dexterous long-horizon tasks, whether robots can succeed in places they've never been, whether robots can respond to open-ended prompts and interjections. Even if you're not excited about robotics, I think the lessons we've learned from trying to address these problems are applicable outside the physical world.

案例研究:叠衣机器人 Case Study: Laundry Folding Robot

Chelsea

我们能否开发出能完成灵巧长周期任务的机器人?在第一部分,我想谈谈我们如何训练一个 pi zero 基础模型来执行这个任务:从烘干机中取出衣物并叠好。到目前为止,我认为这是我在物理世界中见过的最令人印象深刻的机器人行为。这非常困难。这是一个极其困难的问题。你可以看到它并不完美;它会犯一些错误。但这真的很难,因为你需要处理衣物的可变性以及它们可能被放置和皱褶的方式,并且能够处理所有这些事情。当机器人执行这个任务时,大约需要 10 分钟,有很多机会灾难性地失败,例如,把东西掉在地上,这很难恢复。你必须能够从即使很小的错误中恢复。我个人与 Michael 和 Siraj 一起在这个叠衣机器人上做了很多工作,当然还有整个 Physical Intelligence 团队的支持。

Can we develop robots that can complete dexterous long-horizon tasks? In this first part, I'd like to talk about how we trained a pi zero foundation model to do this task: unload a dryer and fold laundry. To date, I think this is the most impressive thing I've seen a robot do in the physical world. It's really hard. This is an incredibly difficult problem. You can see it's not perfect; it makes some mistakes. But it's really hard because you have to deal with the variability in clothes and the way they might be positioned and crumpled, and be able to handle all those things. As you're doing this task, which takes about 10 minutes for the robot, there are many opportunities to fail catastrophically, for example, dropping things on the ground, which is hard to recover from. You have to be able to recover from even small mistakes. I was personally working quite a bit on this laundry folding robot along with Michael and Siraj, of course supported by the whole Physical Intelligence team.

Chelsea

那么你如何解决这个问题?这对机器人来说是一件非常困难的事情。我们做的是从简单开始。我们从:机器人能否折叠一件单一尺寸、单一品牌的衬衫?以及机器人能否动态展平一件衬衫,同样是单一品牌、单一尺寸?如果你从简单开始,这会让问题容易得多。我们通过远程操作收集了一些数据,并用模仿学习训练了一个策略。我们的模型大约有 1 亿个参数,从机器人摄像头的图像映射到机器人手臂的目标关节位置。我们在机器人上以 50 赫兹的频率进行这种控制。我们在 2024 年 3 月中旬成立了公司。几个月后,在我们设置好一切之后,我们得到了一个能够相当可靠地折叠单一尺寸、单一品牌衬衫的策略。你可以看到我在这里测试这个策略。

So how do you approach this problem? It's a really hard thing for a robot to do. What we did is we started simple. We started with: can a robot fold a single size, single brand shirt? And can a robot dynamically flatten one shirt, again single brand, single size? If you start simple, this makes the problem quite a bit easier. We collected some data with teleoperation and trained a policy with imitation learning. Our model had around 100 million parameters, mapping from images from the robot's cameras to target joint positions on the robot arms. We do this source of control at 50 hertz on the robot. We founded the company in mid-March of 2024. A couple months later, after we had set everything up, we were able to get a policy that could fairly reliably fold a single size, single brand shirt. You can see that I'm testing the policy right here.

初始叠衣测试 Initial Laundry Folding Tests

Chelsea

我们还想要测试一些动态动作,因为要完成这类动态动作,你需要精确匹配控制频率。这是我们针对叠衣服问题的一些初步测试。然后我们逐步增加难度。我们不再从衬衫平铺在桌上开始,而是从皱成一团的状态开始。这实际上让任务难了很多。这里有一些我们最初尝试训练机器人叠衬衫的视频。机器人很挣扎。它做了一些看起来还算合理的事情,但总体上无法取得进展。在多次测试中,我们经常得到 0%的成功率,真的很难取得进展。这引入了处理衬衫在桌上可能出现的各种皱褶状态的挑战。

We also wanted to test some dynamic motions because you need to match the control frequency accurately to do these sorts of dynamic motions. These were some of our initial tests addressing the laundry folding problem. Then we made the problem incrementally harder. Instead of starting with the shirt flat on the table, we started in a crumpled position. This actually makes it a lot harder. Here are some videos of our initial attempts to train the robot to fold these shirts. The robot struggles. It does some things that look somewhat sensible but generally isn't able to make progress. With many tests, we frequently got 0% success rate and really struggled to make progress. This introduces the challenge of handling the variability in how shirts might be crumpled on the table.

Chelsea

去年 6 月底,我们取得了一些初步进展。在这个案例中,机器人能够把衬衫弄平。它也能从那个初始状态相当好地叠好衬衫。但仍然不完美。如你所见,这花了相当长的时间。这个视频是 4 倍速播放的,所以不是你有耐心等待的事情。有了初步进展但成功率很低,我们开始转向一个稍微难一点的版本,衣服从洗衣篮里开始。我们还引入了不同尺寸的衬衫和短裤。同样,机器人非常挣扎。在许多测试中,我们全面得到 0%的成功率,真的很难让机器人学会这些任务。

We had some initial signs of life in late June of last year. In this case, the robot was able to make progress on flattening the shirt. It was also able to fold the shirt decently well from that initial state. Still not perfect. As you can see, it takes quite a while. This video was sped up 4x, so not something you might have patience for. With some initial signs of life but a very low success rate, we started to transition to a slightly harder version where the laundry starts in a laundry basket. We also introduced variable size shirts and shorts. Again, the robot really struggled. In many tests, we got 0% success rate across the board, and we really struggled to get the robots to learn these tasks.

Chelsea

此时,我们考虑了很多事情。我们想也许机器人需要记忆,需要某种历史信息。也许我们需要更长时间地训练模型。也许我们应该在末端执行器空间而不是关节空间进行控制。也许我们的编码器有校准问题,需要更一致的校准。也许我们需要在模型上附加更多关于数据的信息。也许我们需要层次结构,因为这是一个长周期任务,需要分解成子任务。也许我们需要更高分辨率的图像。也许我们需要在数据收集中引入干预。我们尝试了其中很多方法。我们经历了大约两到三个月的失败,没有任何方法真正奏效。但后来我们取得了突破:我们发现了一件事确实带来了改变。从语言建模中汲取灵感,我们不是只在所有数据上训练一个策略,而是在所有数据上预训练,然后在精心策划的、一致的、高质量的演示数据集上微调。当我们这样做时,机器人能够取得进展,更可靠地叠衣服。

At this point, we considered many things. We thought maybe the robot needs memory, needs history in some way. Maybe we need to train our models for longer. Maybe we should do control in end-effector space rather than joint space. Maybe our encoders had calibration issues and we need that calibration to be more consistent. Maybe we need to condition the model on more information about the data. Maybe we need hierarchy because this is a long-horizon task that needs to be broken down into subtasks. Maybe we need higher resolution images. Maybe we need to introduce interventions in data collection. We tried many of these things. We had about two to three months of failure where nothing really worked. But then we had a breakthrough: we found one thing that really made a difference. Taking inspiration from language modeling, instead of just training a policy on all our data, we pre-trained on all the data and then fine-tuned on a curated, consistent, high-quality set of demonstration data. When we did this, the robot was able to make progress and fold articles of clothing more reliably.

Chelsea

我想这个视频是第一个机器人能够连续叠五件衣服并堆叠起来的视频。那天我回家非常兴奋。那是 2024 年 9 月,距离我们最初的测试已经好几个月了。这远非完美。叠五件衣服需要 20 分钟。但这表明这个配方解锁了机器人叠衣服的能力。你可以看到这里的失败。在这个案例中,它尝试叠蓝色衬衫大约七次,最终才成功。还有其他失败模式。这里有一个例子,机器人把堆叠推到桌子角落,摆弄了一下,然后把它滑下桌子,接着若无其事地继续。

I think this video was the first where the robot was able to fold five items in a row and stack them. I went home very excited that day. This was in September 2024, multiple months after our initial tests. It's far from perfect. It takes 20 minutes to fold five items. But it suggested that this recipe unlocked the robot's capability to fold these articles. You can see failures here. In this case, it attempted to fold the blue shirt around seven times before eventually figuring it out. There are other failure modes too. Here's an example where the robot pushes the stack to the corner of the table, fiddles with it, then slides it off the table, and proceeds as if nothing happened.

Chelsea

我们继续迭代这个配方。我们选择并改进了我们的策划策略,以策划更高质量的演示数据集。我们将叠五件衣服的时间从 20 分钟缩短到 12 分钟。这是我们评估机器人系统的方式。它仍然会犯错。折叠质量仍然有变化,但比我们之前的策划配方好得多。此时,我们仍然主要只在洗衣数据上预训练和微调模型,没有利用社区中的预训练模型。Physical Intelligence 的一些人正在开发一个在所有机器人数据上训练的预训练模型。然后我们开始将这些模型引入我们的配方。

We continued to iterate on this recipe. We selected and worked on our curation strategy for curating a higher quality set of demonstration data. We got the time from 20 minutes down to 12 minutes for these five items. This is how we evaluated our robot system. It still makes mistakes. The fold quality still varies, but it's significantly better than our previous curation recipe. At this point, we were still training models largely pre-training and fine-tuning only on laundry data, and we weren't leveraging pre-trained models from the community. Some folks at Physical Intelligence were working on developing a pre-trained model trained on all robot data. We then started to introduce these models into our recipe.

Chelsea

我们采用了一个开源视觉语言模型,一个 30 亿参数的模型,叫做 PaliGemma。之前我们使用的是 1 亿到 3 亿参数的模型。这个模型以机器人的图像和语言指令作为输入,并有一个扩散头,关注视觉语言模型的所有内部值,结合关节角度,预测未来 50 个动作的块,大约 1 秒的动作步长。我们使用流匹配,一种扩散的变体,来输出连续动作。我们采用这个预训练模型,不再只在洗衣数据上预训练,而是在我们收集的所有机器人数据上预训练。然后我们用之前开发的相同后训练配方进行微调,没有使用视觉语言模型。当我们这样做时,机器人继续变得更好。在左边的视频中,它 9 分钟完成五件衣服,比之前的 12 分钟快。在右边的视频中,我们用新的衣物物品测试,发现它相当高效地连续叠多件物品。我们还看到使用这个大约大 10 倍、见过更多机器人数据的模型,折叠质量更加一致。

We took an open-source vision language model, a 3 billion parameter model called PaliGemma. Previously, we were using models with 100 to 300 million parameters. This model takes as input images from the robot and a language command, and has a diffusion head that attends to all internal values of the vision language model, and with joint angles, predicts a chunk of 50 actions into the future, about 1 second of action steps. We use flow matching, a variant of diffusion, to output continuous actions. We took this pre-trained model and instead of pre-training only on laundry, we pre-trained on all robot data we had collected. Then we fine-tuned it with the same post-training recipe we had developed without using vision language models. When we did this, the robot continued to get better. In the left video, it does five items in 9 minutes, faster than the 12 minutes before. In the right videos, we tested with novel clothing items and found it was quite efficient at folding multiple items in a row. We also saw more consistent fold quality by using this model that was about 10 times larger and had seen more robot data.

Chelsea

这里有一些亮点。这是一条机器人从未见过的短裤。

Here are a few highlights. Here's a pair of shorts that the robot hasn't seen before.

叠衣结果 Folding Laundry Results

Chelsea

这是一个有点棘手的场景:要把它弄平,机器人实际上需要伸到短裤底部下面。它做到了。它意识到应该伸到短裤左侧下面,最终把它弄平。一旦成功弄平,它就能顺利折叠。有时叠衬衫也需要类似操作。在这种情况下,它需要把衬衫对折,这可能会让衬衫更皱,但能让它找到衬衫的角,然后折叠。就像我提到的,它还能处理没见过的衣物。这里有个例子:一件 V 领衬衫,即使后训练数据集里没有 V 领输入,它也能折叠。它还能叠带纽扣的衬衫。所以它对不同衣物有一定程度的泛化能力。最后,由于这个策略是一个神经网络,以当前图像为输入,它能处理干扰。这里,Michael 在干扰机器人,机器人意识到在叠另一件衬衫时应该先把这件衬衫收起来。Michael 展开了一边,机器人做出反应。Michael 再次干扰,机器人犯了些错误但能恢复。Michael 再次搞乱。这就是机器人能做到的一些结果。

And this is a kind of tricky scenario where to flatten it, it actually needs to reach under the bottom of the shorts. And it's able to do that. It figures out that it should reach under the left part of the shorts in order to eventually flatten it. And then once it successfully flattens it, it's able to fold it successfully. It also has to do something similar at times to fold shirts. So in this case, it needs to fold the shirt over on itself, which puts it in a more crumpled state arguably, but allows it to find the corners of the shirt and then fold it. And like I mentioned, it also handles unseen clothing items. Here's an example of a shirt with a V-neck that it can fold even though the post-training dataset didn't have any V-necks as input. It also folds shirts with buttons. So it has some degree of generalization to different clothing items. And lastly, because this policy is a neural network taking the current image as input, it handles interruptions. Here, Michael is messing with the robot, and the robot figures out it should put the shirt away while trying to fold the other shirt. Michael unfolds one side and the robot reacts. Michael goes in again, the robot makes some mistakes but recovers. Michael messes it up again. So those are some results of what the robot can do.

预训练与后训练的定量评估 Quantitative Evaluation of Pre-training and Post-training

Chelsea

我提到过预训练和后训练这个方案非常重要。我们可以定量衡量它,并确认这确实是带来改进的原因。所以我们比较了这个预训练和后训练方案与两种变体:不使用任何预训练,只在精选数据集上训练;以及没有后训练,即在所有数据上训练而不是在精选数据集上微调。我们根据任务进展来评估这些模型:从篮子中取出(最简单的部分)算部分进展,弄平、折叠和堆叠算进一步进展。我们发现预训练和后训练方案比省略预训练或省略后训练的性能高得多。值得注意的是,两者都省略基本上只能把物品取出篮子,之后几乎没有进展。而结合预训练和精选后训练则性能高得多,能可靠地弄平和折叠物品。

Now I talked about this pre-training and post-training recipe being really important. We can quantitatively measure that and make sure this is what's leading to improvement. So we compared this pre-training and post-training recipe to not using any pre-training and only training on the curated dataset, versus no post-training where you train on all the data rather than fine-tuning on the curated dataset. We evaluated these models in terms of their progress on the task: partial progress for getting it out of the bin (the easiest part), and further progress for flattening, folding, and stacking. We see that the pre-training and post-training recipe achieves far higher performance than omitting pre-training or omitting post-training. Notably, omitting both basically only gets it out of the bin with little progress after that. Whereas combining pre-training and curated post-training gives far higher performance, reliably flattening and folding objects.

泛化到其他任务 Generalization to Other Tasks

Chelsea

最后我要提的是,这个方案中没有任何部分是专门针对洗衣的。所以我们采用了相同的方案,在其他任务上进行了微调。这里的任务是清理桌子。尽管我们主要迭代的是洗衣任务,但机器人成功完成了这个任务。它还能把咖啡豆舀进咖啡研磨机。这个任务相当困难。它需要组装纸板箱的底部,这需要相当的灵巧性。最后,用火柴自动点燃蜡烛,同样使用了这个预训练和后训练方案。这指向了我之前提到的基础模型的好处:要完成这些不同的任务,你不必从头开始。你可以利用跨多个机器人和多个任务的预训练。我们还将相同的方案应用于其他公司的机器人。这是一个我从未亲眼见过的机器人。他们收集了数据,发送给我们,我们在他们的数据上微调了我们的模型。我们甚至不知道模型是如何被控制的,也不知道他们动作的具体表示。但通过在这个新机器人上微调模型,模型能够控制机器人,在这个例子中制作了一杯咖啡。

And the last thing I'll mention on this note is that nothing in this recipe is specific to laundry. So we took the same recipe and fine-tuned on other tasks. Here the task is to clean up a table. The robot successfully does this task despite us primarily iterating on laundry. It also scoops coffee beans into a coffee grinder. This task is pretty hard. It has to construct the bottom part of a cardboard box, which requires quite a bit of dexterity, and lastly autonomously lighting a candle with a match, again with this same pre-training and post-training recipe. This points to the benefit of foundation models I alluded to before: to do these different tasks, you don't have to start from scratch. You can leverage pre-training across multiple robots and multiple tasks. We also apply that same recipe to robots at other companies. This is a robot I've never seen in person. They collected data, sent it to us, we fine-tuned our model on their data. We didn't even know exactly how the model was being controlled or the exact representation of their actions. But by fine-tuning the model on this new robot, the model controls the robot to make a cup of coffee in this case.

机器人基础模型 Takeaways and Limitations

Chelsea

这部分的一些要点:我们能够独立开发后训练和预训练,解耦问题,最终兼得两者之长。我们发现,在所有数据上训练对复杂任务不起作用,而这种在精选数据上的预训练和后训练能带来更好的性能。我们通过从折叠单件衬衫开始,逐步过渡到更复杂的版本,分解了折叠衣物这个难题。现在有很多局限性。我想指出一个局限性:这些机器人是在它们被测试的环境中训练的。这意味着原则上你可以用这些方法在一个环境中收集大量数据,并在该环境中部署。但最终,环境会发生变化,我们希望将这些机器人应用于它们从未见过的环境。那么机器人如何在从未去过的地方成功呢?从机器学习其他地方学到的教训是收集多样化的数据。所以我们开始收集在许多不同环境中整理卧室和厨房的数据。这里有一个数据样本。我们在旧金山的家庭以及多样化的模拟厨房和卧室中收集了机器人数据。总共有超过 100 个独特的房间出现在数据集中,这些数据最终成为更大预训练混合的一部分。我们在这些多样化的移动操作数据上进行了训练,包括低级动作预测和预测如何完成任务的高级子任务指令。但我们也在之前收集的相当多样化的静态操作数据上进行了训练:在我们的办公室和实验室收集的静态操作数据,以及网络数据和高层指令数据。我应该指出,整理卧室和厨房的移动操作数据只占整个预训练混合的 2.4%。这里的教训是,你可以启动一个新任务和一个全新的机器人,而无需重做所有的数据收集。混合数据中的其余部分没有包含这个特定移动操作器的任何移动操作数据,但我们能够建立在之前所做的一切基础上。

So some takeaways for this part: we were able to independently develop post-training and pre-training and decouple the problem, and eventually get the best of both. We found that training on all the data doesn't work for complex tasks, and this sort of pre-training and post-training on curated data leads to far better performance. And we broke up this really hard problem of folding laundry by gradually starting with folding single shirts and going to more complex versions. Now there are a number of limitations. One limitation I'd like to point out is that these robots were trained in the environments they were tested in. This means in principle you could use these methods to collect a lot of data in one environment and deploy in that environment. But ultimately, things change about an environment, and we want to apply these robots to environments they've never seen before. So how can robots succeed in places they've never been? The lesson from machine learning elsewhere is to collect diverse data. So we started collecting data of tidying bedrooms and kitchens in many different environments. Here's a sample of that data. We collected robot data in homes across San Francisco and in diverse mock kitchens and bedrooms. In total, we had more than 100 unique rooms represented in the dataset that ended up being part of a bigger pre-training mixture. We trained on this diverse mobile manipulation data, including low-level action prediction and predicting high-level subtask commands for how to complete the task. But we also trained on previously collected static manipulation data that was fairly diverse: static manipulation data collected in our office and labs, as well as web data and high-level instructional data. I should point out that the mobile manipulation data of tidying bedrooms and kitchens only accounted for 2.4% of the overall pre-training mix. The lesson here is that you can spin up a new task and an entirely new robot without redoing all the data collection. The rest of the mixture didn't have any mobile manipulation data with this particular mobile manipulator, but we're able to build upon everything that had been done before.

新环境中的训练与测试 Foundation Models for Robotics

Chelsea

这又是基础模型的故事——它们让启动新问题、新应用变得更容易,无需从头开始。但这并不完全容易。我们遇到了几个挑战。其中一个挑战是,这个模型会忽略语言指令。比如我们让它拿起切菜板,它却选择了拿起盘子。我们再次让它拿起切菜板,机器人却自作主张拿起了盘子。然后我们让它把盘子放进水槽。最终它离开切菜板后,还是决定拿起切菜板。所以在模型早期开发中,我们发现它经常忽略语言。为了解决这个问题,我们思考了视觉语言模型如何很好地遵循语言。也许有一种方法可以在处理这个任务时保留预训练模型的固有能力。我们做的就是用这个 PI zero 架构,这个使用扩散的动作头是随机初始化的。这实际上会破坏视觉语言模型中存在的预训练知识。我们发现如果能防止这种破坏,就能获得更好的语言遵循能力。我们想出的方法在某些方面非常相似,但我们要预测分词化的动作。然后当有扩散头时,我们会阻止来自随机初始化扩散头的梯度,以防止它破坏 VLM 骨干的语言遵循能力。我们发现这首先导致训练更快,因为分词化动作是更直接的监督信号。其次,它遵循语言的能力好得多。遵循率从 20%提高到 80%。这表明我们能够保留视觉语言模型骨干中的预训练知识。

And it's kind of this same story of foundation models being able to make it easier to spin up a new problem, a new application without starting from scratch. Now this wasn't completely easy. We had a couple challenges. One of the challenges we ran into is that naively this model can ignore language instructions. So we had in this case asked it to pick up the cutting board and it chose to pick up the plate instead. We're again asking it to pick up the cutting board, and instead the robot had a mind of its own decided to pick up the plate. And then we tell it to put the plate in the sink. And eventually it decides that after moving away from the cutting board, it eventually decided that it would actually pick up the cutting board. So in the early development of our model, we found that it often ignored language. And to solve this, we thought about how vision language models actually follow language well. And so maybe there's a way to preserve the inherent abilities of the pre-trained models when addressing this task. And so what we did is with this PI zero architecture, this action head that's using diffusion is randomly initialized. And this ends up actually deteriorating the pre-trained knowledge that's present in the vision language model. And we found that if we can prevent this deterioration, we might be able to get better language following. And so the recipe that we came up with was actually in some ways fairly similar, but instead we're going to be predicting tokenized actions. And then when we have the diffusion head, we'll be stopping the gradient from the randomly initialized diffusion head to prevent it from deteriorating the language following abilities of the VLM backbone. And we found that this first led to faster training because the tokenized actions are a more direct supervision signal. And second, it also followed language far better. An 80% follow rate rather than a 20% follow rate. Which suggests that we're able to preserve the pre-training in the vision language model backbone.

定量结果与数据多样性 Training and Testing in Novel Environments

Chelsea

所以我们把这些拼凑起来。我们采用这个方法,在所有数据上预训练,包括移动操作数据。我们在各种环境中对移动操作数据进行微调。然后我们在从未去过的地方测试模型。我们租了三个从未去过的 Airbnb。我们把机器人放在那些家里,比如厨房,我让它关柜门。我让它收拾碗碟。它从未见过这些碗碟或叉子等物品。尽管从未到过这里,机器人还是成功了。有不同的台面、家具、物品等等。最后,我让它清理溢出物,机器人照做了,擦干净溢出物,最后把海绵放进水槽。它也能在卧室这样做。劳拉让它清理卧室,它把衣物放进去,扔掉垃圾,然后整理床铺,把枕头放在床头,整理毯子或被子。

So we put those pieces together. We took that recipe and trained it pre-trained it on all of our data, including the mobile manipulation data. We fine-tuned it on mobile manipulation data in a variety of environments. And then we tested the model in places it's never been in before. So we rented three Airbnbs that we had never been to before. We put the robot in those homes, in this case, in the kitchen, and I asked it to close the cabinet. I asked it to put away the dishes. It has also never seen these dishes or these forks, these objects. And the robot's able to succeed even though it's never been here before. There's different countertops, different furniture, different objects, and so forth. Lastly, I asked it to clean up the spill, and the robot is able to oblige and wipe down the spill and eventually put the sponge into the sink. It's also able to do this for bedrooms. So Laura asked it in this case just clean the bedroom and it puts articles of clothing in. It throws away the trash and then is able to tidy the bed by putting the pillow at the top of the bed and tidying the blanket or the comforter of the bed.

失败模式与未来工作 Quantitative Results and Data Diversity

Chelsea

YC 下一批正在接受申请。你有创业想法吗?在 y combinator.com/apply 申请。永远不嫌早,填写申请会提升你的想法。好了,回到视频。定量来看,我提到混合数据中只有 2.7%左右,那么其他数据到底有多大帮助?我们能不能只训练那 2.7%?我们发现右边这些柱状图排除了来自实验室静态机器人等环境的数据,性能显著下降。排除这些数据后,在新家评估时性能降至不到 60%,而使用完整预训练混合数据则高出 20%以上。最后我们还研究了数据多样性是否有帮助?是否重要?我们增加了来自这些环境的数据量来测试。做直观评估固然好,但实际测量效果更有帮助。我们发现如果增加数据中代表的家和地点数量,性能会提升,这很好,而且实际上达到了与在目标环境数据上训练相同的性能水平。这意味着我们基本上缩小了泛化差距,并表明这类任务的瓶颈不在于收集更多样化的数据,而在于提高可靠性和性能。

YC's next batch is now taking applications. Got a startup in you? Apply at y combinator.com/apply. It's never too early and filling out the app will level up your idea. Okay, back to the video. So, quantitatively, I talked about how there's only 2.7% or something of the mixture and so how much does that other data actually help? Could we actually just train on that kind of 2.7%? And we find that these bars on the right which are excluding data from static robots in labs and environments and so forth reduces performance significantly. So the performance goes down to less than 60% when you exclude that data when evaluated in novel homes compared to if you use the full pre-training mixture it has more than 20% higher performance. Lastly we also looked at is the diversity of data helpful? Is it important? And so we increase the amount of data from these environments to test this. It's always good to like you can kind of do vibe eval but it's really helpful to actually measure how well these things work and so this is what this is measuring and we find that if we actually increase the amount of homes the amount of locations that are represented in the data the performance increases which is great and it actually gets to the same level of performance as if we train on data from that target environment and so it means we're actually mostly closing the generalization gap and suggest that the bottlenecks at this point for this sort of task lie not in collecting more diverse data but in actually getting higher reliability and higher performance.

结论与未来方向 Failure Modes and Future Work

Chelsea

现在我还应该提到,存在这样的失败模式,成功率大约 80%。还有很多改进空间。这里有几个失败模式的例子。这里让它把物品放进抽屉。它能放进去,但物品没有完全放入,它就认为完成了,然后继续做下一件事。这里机器人需要把衣服放进洗衣篮。它碾过衬衫,然后卡住了,无法抬起。这里我们让它把碗碟放进水槽,它成功放了一些,但很难拿起切菜板,因为切菜板很薄且紧贴台面。最后一个案例,可能是我最喜欢的,让它把锅铲放进抽屉,它觉得烤箱很像抽屉,于是打开烤箱试图放进去。除此之外,还有速度、部分可观测性、长期规划等挑战。所以还有很多工作要做。

Now I should also mention that there's failure modes like this the success rate was around 80%. There's lots of room for improvement. Here are a couple examples of those failure modes. So here it's told to put the items in the drawer. It is able to put it in the drawer but the item isn't fully in the drawer at the end and it decides that it's done and kind of moves on to the next thing. Here the robot needs to put the clothes in the laundry basket. It drives over the shirt and then it gets stuck and it's not able to lift it up. Here we asked it to put the dishes in the sink and it successfully is able to put a number of the dishes in the sink but it struggles to pick up the cutting board in this particular case because it's very thin and it's flush against the surface of the countertop. And in the last case, my probably my favorite case, it's told to put the spatula into a drawer and it decides that the oven looks a lot like a drawer and so it opens the oven and tries to put it in there. And beyond this, there's also challenges with regard to speed, partial observability, long-term planning and so yeah, lots of work to do still.

分层视觉-语言-动作模型用于开放式机器人控制 Conclusion and Future Directions

Chelsea

所以要点是,通过多样化数据,机器人可以在从未去过的环境中遵循各种指令。这比许多机器人场景(在测试场景中训练)有了很大进步。最后我想谈的是,这个模型的指令集相当有限。它只能遵循特定的一组命令。如果我们思考其他形式的人工智能技术是如何部署的,人们真的很喜欢定制,并告诉机器人他们想要什么,或者告诉系统他们想要从这些模型中得到什么。

So the takeaway here is that with diverse data, robots can follow a variety of instructions in environments that the robot has never been in before. Which is a big step up from a lot of robotic scenarios where they're trained in the scenarios that they are being tested. Now the last kind of bit I'd like to talk about is this model has a fairly limited instruction set. It can only follow kind of a certain set of commands. And if we think about how other forms of AI technology have been deployed, people really like to customize and actually tell the robot what they want or tell the system what they want from these kinds of models.

结束语与问答 Hierarchical Vision-Language-Action Models for Open-Ended Robot Control

Chelsea

就像我们给语言模型提示词一样,我们能否让机器人响应开放式提示和开放式打断?为了做到这一点,实际上也是延续之前的工作,我们采用了分层视觉-语言-动作模型。我们有一个高层策略,将提示分解为中间语言回应和中间原子语言指令。例如,高层提示可能是“你能给我做个三明治吗?”,这个高层策略会将其分解为拿起一片面包的子任务。这被传递给一个低层模型,该模型实际执行并预测目标关节角度,以完成拿起一片面包的低层指令。但仅靠这个,它无法遵循各种提示,处理开放式语言其实相当棘手,因为收集大量包含真实机器人在环的人机交互数据很困难,而且也很难规模化。所以我们做的是,利用所有现有的机器人数据,并为其生成合成数据。具体来说,我们使用语言模型重新标注并生成假设的人类提示,对应机器人所处的场景。例如,我们有一段数据:一个视频,下一个技能是拿起一条奇巧巧克力,因为这是机器人下一步要做的低级标注。然后对于机器人即将拿起奇巧巧克力的场景,我们可以问视觉语言模型:人类可能会给出什么样的假设提示,导致这个特定场景和机器人选择拿起奇巧巧克力?然后我们在这些合成提示上训练高层策略,用可能导致不同情况的各种人类交互来扩充机器人数据。结果,我们能够让机器人遵循各种不同的提示。左边我们问:“嗨,机器人,你能给我做一个火腿奶酪三明治吗?”机器人回答:“当然,我先拿面包,然后加火腿和奶酪。”它能够将这个任务分解为各个子任务:拿起一片面包,放在切菜板上,拿起一片奶酪,放在面包上,拿起一些火腿,等等。它还能遵循更复杂的提示,比如:“嗨,机器人,你能给我做一个素食三明治吗?不过我不喜欢泡菜。”在这种情况下,它分解任务并决定在三明治里加生菜和番茄,不加泡菜、奶酪和肉。除了提示,我们还能训练机器人处理不同的打断。这里有一个不同提示的例子:左边我们训练机器人清理桌子——把垃圾收走,把盘子放进垃圾桶。右边我们要求机器人只清理垃圾,不清理盘子。机器人理解这意味着什么,并将其与低层动作联系起来,只收走垃圾,并在垃圾全部收走后完成。最后,它还能处理打断和情境纠正。在这个例子中,机器人为用户取物品。用户打断说:“给我拿点甜的东西,但不在篮子里。”就在机器人把奇巧巧克力放进篮子之后。机器人说:“好的,我给你拿些彩虹糖。”并通过基本推理来满足用户的要求,能够响应这些发生在机器人所处世界中的纠正。你可能想知道,是否有些现有的基础模型可以作为机器人的高层规划器,进行这种高层推理,而无需训练单独的模型。我们也评估了这一点,发现蓝色显示的遵循指令和任务进展的性能远低于我们系统的性能(绿色显示)。总的来说,我们发现这些前沿模型在机器人相关的视觉理解方面表现不佳,这很合理,因为这些模型并非针对许多物理应用,而且物理世界的数据非常少。

And so just like we prompt language models, can we allow robots to respond to open-ended prompts and open-ended interjections? To do this and actually to do the past work, we're leveraging hierarchical vision-language-action models. So we have a high-level policy that breaks down the prompt into intermediate verbal responses and intermediate atomic language commands. For example, the high-level prompt might be "Can you make me a sandwich?" and this high-level policy will break it down into the subtask of picking up one slice of bread. This is passed to a low-level model that actually executes and predicts target joint angles to fulfill the low-level command of picking up one slice of bread. Now, on its own, this isn't going to be able to follow all sorts of prompts, and it's actually fairly tricky to handle open-ended language because it's challenging to collect a large number of human-robot interactions with the real robot in the loop. This is also fairly hard to scale. So what we did is we took all of our existing robot data and we can actually generate synthetic data for it. In particular, we use language models to relabel and generate hypothetical human prompts for the scenarios that the robots are in. So what this looks like is we take data that says: here's a video, and the next skill is to pick up a Kit Kat because that's what the robot does next in terms of basic low-level annotation. Then for this scenario where the robot is about to pick up the Kit Kat, we can ask a vision-language model: what is a hypothetical prompt that a human might have asked that led to this particular scenario and the robot choosing to pick up a Kit Kat? Then we train our high-level policy on these synthetic prompts to augment the robot data with various human interactions that might have led to those different situations. As a result, we're able to allow robots to follow a variety of different prompts. On the left, we ask, "Hi robot, can you make me a ham and cheese sandwich?" The robot says, "Sure, I'll start with the bread and add ham and cheese next." And it's able to break down this task into the various subtasks: picking up a slice of bread, putting it on the cutting board, picking up a slice of cheese, putting it on the bread, picking up some ham, and so on. It can also follow more complicated prompts like, "Hi robot, can you make me a vegan sandwich? I don't like pickles, though." In this case, it breaks down and decides to add lettuce and tomatoes to the sandwich, and not add pickles, cheese, or meat. In addition to prompts, we're also able to train the robot to handle different interjections. Here's a case of a different kind of prompt: on the left we train the robot to clean tables—put trash away and put dishes into the bin. On the right we ask the robot to clean up only the trash but not the dishes. The robot understands what that means and connects it to its low-level actions, only putting away the trash and completing when the trash is all put away. Lastly, it's able to handle interjections and situated corrections. In this case, the robot is getting items for a user. The user interjects and says, "Get me something sweet that's not in the basket," right after it had put a Kit Kat into the basket. The robot says, "Sure, let me get you some Skittles," and reasons through how to fulfill the user's request, responding to those kinds of corrections situated in the world. Now you might wonder if some existing foundation models could serve as a high-level planner for robots and do this sort of high-level reasoning without training a separate model. We evaluated that and found that in blue, the performance at following instructions and making progress on the task was substantially lower than the performance of our system, shown in green. In general, we found that these frontier models struggle with visual understanding as it pertains to robotics, which makes sense because these models aren't really targeting many physical applications and have very little data in the physical world.

机器人强化学习 Closing Remarks and Q&A

Chelsea

好了,开始总结,然后我们会有一些时间提问。我谈到了机器人如何通过预训练和后训练完成各种灵巧的长周期任务,如何在从未到过的地方成功,以及如何利用我们收集的机器人数据之上的语言模型合成数据来响应开放式提示和打断。最后几点:在这次演讲中,我们看到了几个不同场景,通用机器人可能比专用机器人更成功,因为我们基本上可以为现实世界中的物理智能建立更广泛的基础,而不是为每个应用从头开始。我们还看到,现实世界中的大规模数据对开发这些东西非常有帮助,我们发现这对于物理智能来说是必要但不充分的。还有很多挑战,我们需要更多的研究——我们自己以及通过开源贡献——才能让机器人真正准备好应对开放世界。我还想提一下,Physical Intelligence 正在招聘多个职位。如果你对我们讨论的内容感兴趣,可以在 PI 网站上看到空缺职位列表。太好了。很高兴回答问题。从左边开始。

Okay, so to start to wrap up, and then we'll have some time for questions. I talked a bit about how robots can do a variety of dexterous long-horizon tasks with pre-training and post-training. How robots can succeed in places they've never been, and how they can respond to open-ended prompts and interjections by leveraging synthetic data from language models on top of the robot data we collected. Now with some closing notes: we've seen a few different scenarios in this talk where general-purpose robots might be more successful than specialist robots because we can essentially build upon a much broader foundation for physical intelligence in the real world, rather than starting from scratch for every single application. We also saw that large-scale data in the real world is really helpful for developing these things, and we found that it's necessary but not sufficient for physical intelligence. There are a lot of challenges, and we need more research—ourselves and through open-source contributions—before robots will be truly ready to tackle the open world. I'd also like to mention that at Physical Intelligence we're hiring a number of roles. If you're excited about some of the things we talked about, you can see a list of open roles on the PI website. Awesome. Happy to take some questions. Let's start on the left.

Host

嗨,Chelsea。首先,我想感谢你在机器人学习方面的所有工作,都非常令人印象深刻。我主要有两个问题,特别是关于你提到的后训练部分。第一,你提到在后训练中,最重要的是拥有高质量的动作数据。我想知道它的组成部分是什么。第二个问题是,你认为强化学习在后训练中会扮演什么角色?

Hi Chelsea. First, I want to say thank you for all your work on robot learning. They're all really impressive. And so mainly I have two questions, especially regarding the post-training part you mentioned. The first thing is you mentioned that in post-training, the most important part is to have high-quality action data. So I'm wondering what the components of that would be. And then the second question is what do you think RL will play into the part of post-training?

Chelsea

当然。我认为它的不同组成部分——很大程度上归结为数据的一致性和所遵循的策略,以及数据是否高效且以可靠的策略完成任务。至于第二个问题,我认为强化学习在后训练中可以发挥非常大的作用。

Yeah, absolutely. So I think that the different components of it—a lot of it comes down to consistency of the data and the strategy being followed, and whether the data completes the task efficiently and with a reliable strategy. And then on the second question, I think that reinforcement learning can play a very large role in post-training.

资金与应用 Reinforcement learning for robots

Chelsea

我认为,通过强化学习获得的机器人在线数据,可以让机器人拥有更高的成功率,并且比仅通过模仿学习训练更快。

I think that online data from the robots which reinforcement learning allows you to use can allow robots to have a much higher success rate and also be faster than if they're just trained with imitation learning.

世界模型与 VLA 集成 Funding and applications

Host

非常感谢你的演讲。你的工作非常迷人,毫无疑问未来会产生很大影响。但在这个阶段,我想问你是如何找到资金的?因为老实说,我无法想象说服人们投资一个叠衣服和洗碗的机器人有多难。

Thank you so much for your talk. Your work is really fascinating and there is no doubt that it will have a lot of impact in the future. But can I ask you at this stage how can you find the funding? Because honestly I can't imagine how hard it can be to convince people to invest in a robot that folds clothes and deals with the dishes.

Chelsea

这是个好问题。我认为我们不仅仅专注于家庭应用。我们真正想要解决的是物理智能这个更广泛的问题,我们从那些容易取得进展的应用开始。但我们也一直在做像插入以太网电缆(我在演讲中提到过)以及组装纸板箱这样的任务。总的来说,我认为这类问题有巨大潜力,能在各个领域产生影响,而不仅仅是家务。即使在家庭任务中,我认为这项技术也有巨大的市场。我们自己并没有在融资上遇到太多挑战,而且我认为最近很多机器人公司也做得很好,发现人们对这类技术其实非常兴奋,因为事情真的开始奏效了。我十多年前就开始研究这项技术,那时它真的行不通。所以我认为,这种兴奋正在成熟,并真正准备好应用于现实世界。还有很多工作要做,但总体看来,很多人对这项技术感到兴奋,并渴望投入资金。

It's a good question. I think that we aren't just focused on applications in the home. We really want to solve this broader problem of physical intelligence and we've been starting with those applications because they're ones that are kind of easy to make progress on. But we've also been doing tasks like inserting an Ethernet cable which I put in the talk as well as constructing a cardboard box. Generally, I think that this sort of problem has a ton of potential for making impact in all sorts of realms, not just domestic tasks. And even in domestic tasks, I think there's a huge market for this kind of technology. We ourselves haven't had a lot of challenge with fundraising and I think that a lot of robotics companies recently have also done a great job and found that there's actually a lot of excitement around this sort of technology because things are actually starting to work. I started working on this technology more than 10 years ago and things really weren't working then. So I think that there's a lot of excitement that is starting to mature and actually be ready for the real world. There's a lot more work to do, but generally it seems like there's a lot of people excited about this technology and eager to put funds behind it.

模型规模与检索 World models and VLA integration

Host

非常感谢。我有两个问题,一个更宽泛,一个更技术性。技术问题是:VLA,至少在我看来,是一个与世界模型有些分离的框架,我想知道两者将如何相互作用,以及你是否计划将它们一起使用。就目前而言,VLA 更像是一种策略,可以从世界模型中获益良多。从更广泛的角度来看,我想知道哪些基础设施层最有用,比如可解释性、可追溯性或一般安全性,以便在现实世界中部署这样的模型。

Thank you so much. I have two questions, one more broad and one more technical. The technical one: VLA is, at least to my understanding, a framework that is a bit separate from world modeling, and I wonder how the two will interplay and whether you have planned to use them together. As I see right now, VLA is more of a policy that could benefit a lot from world modeling. And from a broader perspective, I wonder which infrastructure layers could be most useful to work on, such as explainability, traceability, or safety in general to deploy such models in the real world.

Chelsea

好问题。关于第一点,实际上有相当自然的方法可以将世界模型的目标整合到视觉语言动作模型中。我们做过一些工作,不是只预测下一个动作,而是预测一些中间子目标图像,比如为了完成任务未来应该发生什么,然后从那里预测动作。我们看到了一些迹象,表明这似乎很有前景。所以我认为有办法将这两种范式合并。同时,世界模型也面临很多挑战,比如你输入的数据不一定反映你将如何使用它。你可能会用成功完成任务的数据来训练它,然后评估它,试图用它来评估那些并非最优完成任务的行动。这时,即使你提供的输入行动实际上不会导致好的结果,世界模型也会幻觉出一个成功完成任务的视频。所以有一些挑战需要克服,但也有办法将其整合到 VLA 范式中。你能提醒我你的第二个问题吗?

Great question. On the first point, there are actually fairly natural ways to incorporate world model objectives into vision language action models. We've done some work where instead of only predicting the next action, you predict some intermediate subgoal image, like what should happen in the future in order to accomplish the task, and then predict an action from there. We've seen some signs of life that this seems quite promising. So I think there are ways to merge the two paradigms. At the same time, there are a lot of challenges with world modeling regarding the ways in which the data you put into it is not necessarily reflective of how you're going to use it. You might train it on demonstration data of successful task completion and then evaluate it to try to use it to evaluate actions that are not optimally completing the task. Then the world model will hallucinate a video of completing the task successfully even if the actions you provide as input wouldn't actually lead to a good outcome. So there are challenges to overcome, but there are also ways to integrate it into the VLA paradigm. Could you remind me of your second question?

Host

在最短时间内,你希望研究哪些基础设施层,以最大程度地改进这些模型在机器人上的实际运行?

What are the infrastructure layers you want to work on in the shortest term to bring the most improvements to actually run these models on robots?

Chelsea

要在机器人上实际运行这些模型,你需要一个实时系统,它需要达到一定的频率才能成功执行动作。如果系统有延迟,就会带来各种挑战。因此,考虑快速推理和机器人上的基础设施是我们软件团队工作的重要部分。同时,还要考虑大规模机器学习基础设施、训练大型模型、摄取大量数据。我们拥有的数据与典型数据集不同,因为它非常多模态:视频、动作、语言片段以及其他各种组件。所以,在机器人端和模型训练端都有有趣的基础设施问题。

To actually run these models on robots, you need a real-time system that needs to hit a certain frequency to execute actions successfully. If there is lag in that system, it introduces all sorts of challenges. So thinking about fast inference and infrastructure for that on the robot is a big part of what our software team does. Also, thinking about large-scale machine learning infrastructure, training large models, ingesting large amounts of data. The data we have is different from typical datasets because it's very multimodal: videos, actions, language segments, and various other components. So there are interesting infrastructure problems both on the robot side and on the model training side.

引言与致谢 Model size and retrieval

Host

嗨,我是弗雷德里克。我有一个关于模型大小的问题。我认为我们现在看到的是,更大的模型尺寸会带来更好的准确性。例如,在你的实验中,或者 OpenAI、Anthropic 等公司用他们的 LLM 所做的。然而,也有一种方法是使用一个非常小的模型,然后将世界知识外包给一个模型可以与之交互的数据库。你怎么看?你认为这是一种有效的方法,还是认为将所有世界知识封装在模型内部更好?

Hi, I'm Frederick. I have a question about model sizes in general. I think what we're seeing right now is that larger model sizes lead to better accuracy. For example, in your experiments, or also what OpenAI and Anthropic and others are doing with their LLMs. However, there is also the approach of using a quite small model and then outsourcing the world knowledge into a database with which the model can interact. What is your take on that? Do you think that's a valid approach, or do you think encapsulating all the world knowledge inside the model is better?

Chelsea

这是个有趣的问题。根据我从事基于检索的系统的经验,首先,要弄清楚哪些应该外包,哪些应该由模型实际完成,这其实有点棘手;其次,有时模型会忽略检索到的内容,试图自己生成一些东西。要让它在技术上完全按照你的意愿工作似乎非常棘手。我认为这可能取决于应用和用例,看是否有意义,但根据我的经验,弄清楚分工最终会相当棘手。即使是模型部分也需要一定程度的智能才能实际利用检索到的信息。

It's an interesting question. In my experience working on retrieval-based systems, it is actually a bit tricky to first figure out what should be offloaded versus actually done by the model, and second, sometimes the model will ignore the retrieved content and try to generate something itself. It seems to be very tricky to get that technically to work exactly the way you want. I think it's probably going to depend on the application and the use case in terms of whether that might make sense, but in my experience, it ends up being quite tricky to figure out the division of labor. Even the model part will need to have some degree of intelligence to actually make use of the retrieved information.

物理智能领域建设者的机遇 Introduction and appreciation

Host

嗨,Chelsea。我叫 Charu Thomas。首先,非常感谢你的演讲,非常精彩,从元学习开始我就是你工作的忠实粉丝。

Hi, Chelsea. My name is Charu Thomas. First off, really appreciate the talk. It was really fascinating and have been a big fan of your work since metalearning.

机器人合成数据 Opportunities for builders in physical intelligence

Host

当你思考软件和硬件将如何继续发展时,对于你提出的物理智能愿景,今天的建设者最大的机会是什么?

When you think about how software and hardware are going to continue to evolve, what are the biggest opportunities for builders today for your vision of physical intelligence?

Chelsea

我认为有很多不同的机会可以让事情变得更好,也有很多悬而未决的问题。正如我之前提到的,思考在机器人端更好的基础设施方式。这类事情的开源代码不多,但有很多机会改善机器人基础设施,而且没有太多人在做这方面的工作。此外,我喜欢人工智能和计算机科学的一点是庞大的开源社区。有大量的机会去做开源工作,并为更广泛的社区做出贡献,这个社区正在努力收集数据、开源模型、修复模型中的错误、微调模型,并找出微调的新方法。所以各种各样的问题,尤其是在开源领域的研究方面。

I think there are lots of different opportunities to make things work a lot better and a lot of open questions. As I mentioned before, thinking about better ways of having infrastructure on the robot side. There isn't a lot of open source code for that sort of thing, but there are many opportunities to make robot infrastructure better, and not a lot of people are working on that aspect. Also, one of the things I love about AI and computer science is the large open source community. There is a ton of opportunity to do open source work and contribute to a broader community that is trying to collect data, open source models, fix bugs on those models, fine-tune those models, and figure out new recipes for fine-tuning. So all sorts of questions, especially on the research side in the open source realm.

学术界与工业界的机器人研究 Synthetic data for robotics

Host

嗨,Chelsea。我最近读了很多你们团队的工作,特别喜欢 Siraj 的博士论文。它教会了我很多关于用数据扩展现实世界机器人的知识。我的问题是,你认为合成数据未来将如何为机器人技术扩展?正如我们在语言模型中看到的,我们已经从人类收集的数据转向创建更多的合成数据,并进行大量过滤和自我评分。那么,你认为使用生成式合成数据来创建环境或奖励模型将如何影响机器人技术?

Hi, Chelsea. I've been reading through a lot of your group's work recently and particularly enjoyed reading Siraj's PhD thesis. It taught me a lot about scaling real world robotics with data. A question I have is how do you think synthetic data will scale for robotics in the future? As we've seen with language models, we've moved away from human collected data into more creating synthetic data with a lot of filtering and self-grading. So, how do you think using generative synthetic data for creating environments or reward models will impact robotics?

Chelsea

我对这个话题有很多想法。我认为归根结底,真实数据是无法替代的。大量的真实机器人数据将是任何以可泛化方式工作的系统的必要组成部分。同时,我确实认为模拟和合成数据等工具可能在评估方面发挥作用。评估一个模型泛化到许多环境的效果非常棘手,因为你需要把机器人带到那些环境或构建它们。而在模拟中,这变得容易得多。所以我非常看好模拟和合成数据在这方面的应用。我还应该提到,语言模型中合成数据的类比不一定是机器人技术中的模拟,而更接近于强化学习之类的东西。很多合成数据是由试图完成任务并通过不同方式推理的模型生成的。类比是机器人尝试完成任务并从自身尝试中学习。来自模型的这种在线数据也将在后训练中发挥关键作用,这是我们正在大量研究的内容。所以我认为这非常重要且有用。

I have many thoughts on this topic. I think that at the end of the day, there will be no replacement for real data. Large amounts of real robot data will be a necessary component of any system that works in a generalizable way. At the same time, I do think that tools like simulation and synthetic data can potentially play a role on the evaluation side. It's very tricky to evaluate how well a model generalizes to many environments because you need to bring the robot to those environments or construct them. In simulation, that gets a lot easier. So I'm really excited about simulation and synthetic data for that use case. I should also mention that the analog of synthetic data in language models is not necessarily simulation in robotics but closer to something like reinforcement learning. A lot of synthetic data is generated by the model trying to do the task and reasoning through different ways. The analogy there is a robot that tries to attempt the task and learn from its own attempts. That sort of online data from the model will also play a critical role in post-training, something we are working on quite a bit. So I think that is really important and helpful.

架构限制与分词 Academia vs industry for robotics research

Host

看到你作为麻省理工学院 EECS 校友,现在从事非常酷的机器人研究,并和我们谈论机器人和创业,真是太棒了。我一直想知道涉及硬件组件的机器人研究在学术界和工业界有何不同。通常一个环境是否比另一个有更多的资源、更少的限制或更广泛的应用?你认为什么样的人或目标更适合每条道路?

It's super cool to see you as an MIT EECS alumni now working in a really cool robotics and talking to us about robotics and entrepreneurship. I've been wondering how robotics research that involves hardware components plays out differently in academia versus industry. Are there typically more resources, fewer constraints, or broader applications in one setting over the other? And what kind of people or goals do you think might be better suited for each path?

Chelsea

这是个有趣的问题。我仍然喜欢初创公司、学术和工业环境。它们各有优缺点。通常,学术环境在数据收集吞吐量、评估吞吐量和算力方面不如初创公司和工业实验室资源丰富。但与此同时,有很多问题不需要大量资源就能解决,我们需要在算法方面弄清楚。所以那里有很多有趣的工作要做。在工业和初创公司中,研究大型模型、扩展数据以及观察大规模下发生的事情非常棒。我认为两者都有其位置。差距并不像人们通常认为的那么大。通常工业界的人希望有更多的算力。你总是希望有更多资源。有时当你拥有大量资源时,你不会仔细考虑要运行什么,最终会浪费更多的算力。所以根据我的经验,拥有更多资源也有缺点。

It's an interesting question. I still love both startup, academic, and industry environments. They all have various pros and cons. Generally, academic environments aren't quite as well resourced in terms of data collection throughput, eval throughput, and compute as startups and industry labs. But at the same time, there are many problems you can solve without large amounts of resources that we need to figure out on the algorithm side. So there is a lot of interesting work to be done there. In industry and startups, doing research on big models, scaling up data, and seeing what happens at large scales is really great. I think there is a place for both. The gap isn't as large as often people make it seem. Often people in industry wish they had more compute. You always wish you had more resources. Sometimes when you have a lot of resources, you don't think as carefully about what runs you're going to do, and you end up being more wasteful of compute. So there are downsides to having more resources in my experience.

Architecture limits and tokenization Architecture limits and tokenization

Host

非常抱歉。我能问一个关于架构的快速问题吗?我知道缩放定律在基于 Transformer 的架构上效果很好。你目前是否看到基于 VLM 的架构存在限制,这种架构是为文本 token 设计的,因为它们没有物理感知模块?你如何处理这个问题?

I'm really sorry. Can I just ask one quick question on architecture? I know that the scaling laws have worked well for transformer based architectures. Do you see currently limits in VLM based architecture which are kind of made for text tokens because they don't have modules for physical awareness? And how do you deal with that?

Chelsea

我们将动作进行了 token 化。我建议你看看我们发表的快速 token 化论文,这是一种实现方式。我们就到这里吧。谢谢大家,希望你们喜欢这次活动。

We tokenized the actions. I'd encourage you to take a look at the fast tokenizer paper that we put out as a way to accomplish that. And we should wrap up there. Thanks everyone and hope you enjoy the event.

互动版:逐字朗读 + 针对本期提问 →