通用机器人在现实世界的实现路径

The Path to General-Purpose Robots in the Real World

切尔西·芬恩 Chelsea Finn · Y Combinator · 2026-08-12 · 约 58 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

探讨如何开发能在现实世界中自主可靠运行的通用机器人,并从 AI 应用历史中汲取经验。

Exploring how to develop general-purpose robots that can operate autonomously and reliably in the real world, drawing lessons from AI's production history.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 39)

全文 · Full transcript(中英对照)

开场与公司进展 Introduction and Company Progress

Chelsea

大家好,今天我要讲的是物理智能(physical intelligence)的前沿进展。特别是,两年前我创立了一家名为 Physical Intelligence 的公司。我们非常感兴趣的是,如何能够基本上开发出任何机器人,让任何机器人在现实世界中完成任何任务。我去年也在这个活动上发过言。去年,我分享了我们在公司取得的一些进展,当时我们已经能完成非常复杂的任务,比如折叠、取出并折叠衣物。我还谈到了我们首次展示了机器人如何在从未去过的房间等环境中完成有用任务。

Everyone, today I'm going to be talking about the state-of-the-art of physical intelligence. And in particular, two years ago, I founded a company called Physical Intelligence. And we're really interested in how we can basically develop any robot, allow any robot to do any task in the real world. And I actually spoke at this event a year ago. Last year, at the event, I shared some of our progress at the company, where we could do really complicated tasks like folding, unloading, and folding laundry. And I also talked about how, for the first time, we showed how robots can do useful tasks in environments, in rooms they've never been in before.

Chelsea

自那以后,也就是一年来,我们让机器人做了很多其他很酷的事情。比如,我们让机器人能够清洗右上角的油腻平底锅,或者削胡萝卜皮(下面视频里),或者做烤芝士三明治(下面视频里),或者切西葫芦,等等。但我今天真正想重点讲的不是机器人做各种事情的酷炫视频,而是让机器人在现实世界中有用真正需要什么。具体来说,我们如何开发出在现实世界中有用的通用机器人?

Now, since then, since one year ago, we have gotten robots to do a lot of other really cool things. So, for example, we've gotten robots to be able to wash a greasy pan in the top right, or peel a carrot in the video below that, or make a grilled cheese sandwich in the video below that, or slice a zucchini, and so forth. But what I'd really like to focus on today isn't cool videos of robots doing lots of different things, but what it actually takes to get robots to be useful in the real world. And specifically, how can we develop general-purpose robots that are useful in the real world?

Chelsea

这有两个方面。第一是通用性:我们如何开发通用模型。第二是真正把这些模型带到现实世界,让它们能产生实际影响、对人们有用。在第一部分,我会讲在现实世界中有用的问题。

Now there are two aspects of this. The first is general purpose: how we can develop general-purpose models. And the second is actually bringing those models to the real world so that they can actually have an impact and be useful to people. And in the first part, I'll talk about being useful in the real world.

AI在生产中的历史背景 Historical Context of AI in Production

Chelsea

所以,为了真正把技术带到现实世界,我认为我们需要弄清楚——回顾人们过去如何将 AI 带入现实世界会很有帮助。如果我们看看利用机器学习等技术的重大产品发布的时间线,我们会看到这样一条时间线。我认为最早真正在现实世界中应用机器学习的例子是产品推荐和广告排序。五年后,我们开始看到不仅使用机器学习,还使用深度学习来做类似的应用。这是一个非常激动人心的进步,因为深度学习是一种可以直接应用于复杂输入输出场景的算法,而且更容易迁移到其他应用。

So to actually bring a technology to the real world, I think we need to figure out—it's helpful to actually look at what people have done in the past to bring AI into the real world. And if we look at a timeline of major production launches that are leveraging technology like machine learning, we can see a timeline like this. So I think the really first early examples of machine learning being used for real in the real world were for things like product recommendations and ad ranking. And then five years later, we started to see not just machine learning being used but deep learning being used for the same sorts of applications. This was a really exciting advance because deep learning is an algorithm that you can really apply out of the box to scenarios that involve really complex inputs and outputs, and it makes it easier to translate to other applications.

Chelsea

但在此基础上,我认为在机器学习和 AI 生产应用中,我们看到的更激动人心的时刻是 2022 年 ChatGPT 的发布。这是我们第一次看到通用模型真正被现实世界中的许多人使用。五天内,ChatGPT 就达到了一百万用户。当然,最近我们还看到像 Claude Code 这样的工具在现实世界中非常有用,希望对我们许多人来说也是如此,还有其他编码智能体。

But from there, I think that an even more exciting moment in time that we saw in terms of machine learning and AI in production was in 2022 with the launch of ChatGPT. And this was the first time where we saw a general-purpose model truly being used by many different people in the real world. Within five days, ChatGPT had reached a million users. And then of course, more recently, we've seen things like Claude Code also be incredibly useful, hopefully to many of us in the real world, and other coding agents.

Chelsea

现在,如果我们看看 AI 在现实世界中的使用情况,我认为我们可以得出几个不同的结论。第一,通用模型越来越多地被用于解决现实世界的问题。所以我们确实看到能够做很多很多不同事情的通用 AI 模型在现实世界中被使用,我们看到了从左到右的转变。但我也认为,从这些应用中我们可以得出一个更微妙的观察,特别是如果我们看看所有这些应用,其中机器学习在现实世界中确实有用且盈利等等,在所有这些应用中,客户或多或少都是根据 AI 模型的推荐来做决策。

Now if we look at how AI has been used in the real world and kind of look at this, I think there are a few different takeaways we could make. The first is that generalist models are increasingly being used for real-world problems. So we're actually seeing general-purpose, generalist AI models that can do many, many different things actually be used in the real world, and we see that transition from the left to the right. But I also think that there's a more nuanced observation that we can make from looking at these applications, and in particular, if we look at all of these different applications that are used where machine learning has actually been useful in the real world and actually been profitable and so forth, in all of these applications, the customer is making a decision based off of the recommendation of the AI model more or less.

Chelsea

这意味着,如果最终是客户在做决策,那么即使系统犯错也没关系,因为通常人能够识别错误,或者即使在有错误的情况下也能决定该怎么做。所以即使这类系统并不完美,它们对不同的人来说仍然非常有用,而且它们不必完全完美的压力也更小。

And this means that if the customer is ultimately making the decision, this means that if the system makes a mistake, that's okay because usually the person can recognize that or decide what to do even despite that mistake. And so even when these sorts of systems aren't perfect, they're still incredibly useful to different people, and there's less pressure on them to be completely perfect.

Chelsea

我认为物理 AI 和机器人技术实际上与此非常不同,如果我们考虑真正在物理世界中运行的物理 AI,它们必须直接做出影响物理世界的决策。这意味着当它们完全自主运行时,它们会更有用。因此,这要求我们开发出比迄今为止部署的机器学习系统错误率低得多的物理 AI 系统。

And I think that actually physical AI and robotics is pretty different from this, where if we think about physical AI that are actually operating in the physical world, they have to be directly making decisions that affect the physical world. And this means that they're going to be far more useful when they're operating fully autonomously. And as a result, this requires us to develop physical AI systems that make far fewer mistakes than the machine learning systems that have been deployed thus far.

Chelsea

现在,最近发生的一件非常令人兴奋的事情是,一年前,Waymo 每周自动驾驶出行次数超过了 25 万次,这表明开发一个基于机器学习的系统,能够在物理世界中以可信赖和自主的方式运行,确实是可能的。我认为这为在物理世界中其他 AI 领域做同样的事情带来了很多希望和乐观情绪。

Now, one really exciting thing to highlight that has happened recently is that a year ago, Waymo passed the quarter of a million weekly autonomous rides, suggesting that it is really possible to develop a machine learning based system that can operate in a trustworthy and autonomous way directly in the physical world. And I think that brings a lot of hope and optimism for actually doing the same with the rest of AI in the physical world.

长期自主性与浓缩咖啡示例 Long-Term Autonomy and the Espresso Example

Chelsea

所以,如果我们想在现实世界中开发通用机器人,我认为我们需要思考如何让它们长时间自主运行,这样它们才真正有用,而不是让人类基于模型的预测来做决策。为了思考长期自主性,我想用一个具体例子来说明,假设我们希望机器人做浓缩咖啡。如果我们希望它对我们真正有用,我们需要可靠地制作浓缩咖啡,这样我们就不必频繁地照看机器人,它才能帮助提供饮品。

So if we want to develop general-purpose robots in the real world, I think we need to think about how we're going to make them autonomous for long periods of time so that they're actually useful, rather than having them be something where a human is basing decisions on the predictions of the model. So to think about long-term autonomy, I want to ground this in a specific example and say that we wanted a robot to make espresso. If we wanted it to be actually useful for us, we needed to make espresso reliably so that we don't have to babysit the robot very frequently in order for it to help serve drinks.

Chelsea

即使单独看,这个任务也非常困难。实际操作手柄需要非常精确和有力的控制才能正确插入。它还需要平稳地处理装有液体的杯子,并且不洒出来。它还需要有准确的时间感,这通常在其他机器学习领域并不是问题。而且我们不仅想做这个相当有挑战性的任务,我们还希望以超过 90% 的可靠性完成它。那么我们该怎么做呢?

Now even on its own, this task is really difficult. So actually operating the portafilter requires very precise and forceful control to insert it appropriately. It also needs to smoothly handle cups with liquid in it and not spill those cups. And it also needs to have an accurate sense of timing, which often isn't actually an issue in other areas of machine learning. And not only do we want to do this pretty challenging task, we want to do it with over 90% reliability. So how can we do this?

Chelsea

所以机器学习的第一步总是收集一些数据集,训练一个模型,并评估你的模型有多好。不幸的是,这很少能在第一次尝试时就可靠地工作。

So the first step in machine learning is always to collect some data set, train a model, and evaluate how good your model is. And unfortunately, this rarely works reliably on the very first try.

模型迭代与可靠性扩展 Iterating on Models and Scaling Reliability

Chelsea

在实践中,更好的做法是对你开发的模型进行迭代,尝试收集更多数据、提高数据集中标签的质量、让标签更详细、收集更多关于边缘情况和模型表现不佳场景的数据、调整数据集的平衡等等。虽然这通常能提高模型的可靠性,但人们最终会感到疲惫,而且靠人工手动调整很难达到真正的高可靠性。因此,更好的做法是让 AI 系统本身在需要更高可靠性的场景中迭代,自动寻找需要更多数据和更多监督的地方。如果我们能因为自动化而非人工而进行更多次迭代,那么这可能是让物理 AI 系统达到 99% 以上可靠性的途径。这就是我们将采取的方法,它看起来很像强化学习算法,尝试完成任务,从失败中学习,并自行改进。

In practice, it's a bit better to iterate on the model you've developed, where you try to collect more data, improve the quality of the labels, make the labels more detailed, collect more data on edge cases and scenarios where it's not working well, adjust the balancing of the data set, and so forth. While this generally improves the reliability of the model, people eventually get tired, and it's hard to get really high reliability with a person manually tuning this. So what would be even better is if the AI system itself can iterate on the scenario where you want it to have higher reliability, automatically seeking out places where it needs more data and more supervision. If we can do this for many more iterations because it's automatic rather than a person doing it, then this might be the way to get really high, like 99 plus percent reliability from physical AI systems. This is the approach we'll take, and it looks a lot like a reinforcement learning algorithm that tries to attempt the task, learn from its failures, and get better on its own.

为机器人开发可扩展RL Developing Scalable RL for Robotics

Chelsea

那么,我们如何为机器人技术开发可扩展的强化学习配方呢?在语言模型中,我们有像 PPO 和 GRPO 这样的算法,它们已经扩展到大型语言模型,并实现了非常复杂的推理。但将其应用于机器人技术存在挑战:这些算法通过扩展算力,已经用数百万次尝试(有时甚至数千万次)进行了训练,因为每次尝试只是在数据中心运行语言模型,使用算力。如果我们非常粗略地将其转化为机器人技术,假设我们不是数百万或数千万次,而只有一百万条一分钟机器人任务的轨迹——甚至比我提到的浓缩咖啡任务还短——这将对应 700 个机器人日才能获得该任务的高可靠性。也许这并非完全不可能,但这将相当具有挑战性,因为计算方式有所不同。我们不仅仅是在运行算力来优化用例;我们实际上是在现实世界中运行机器人,使用硬件,并在现实世界中尝试任务。因此,我们希望有一种算法能够更高效地迭代。

So how do we develop a scalable reinforcement learning recipe for robotics? In language models, we have algorithms like PPO and GRPO, which have scaled to large language models and enabled really complex reasoning. But there's a challenge in applying this to robotics: these algorithms have been trained with millions of attempts, sometimes tens of millions, by scaling up compute, because each attempt is simply running the language model in a data center using compute. If we were to translate this very approximately to robotics, say we had maybe not millions or tens of millions but just one million trajectories of a one-minute robot task—even shorter than the espresso task I mentioned—this would correspond to 700 robot days to get high reliability for that task. Maybe this isn't completely out of the question, but it would be quite challenging because the calculus is a bit different. We're not just running compute to optimize for a use case; we're actually running the robot in the real world, using the hardware, and attempting the task in the real world. So we'd like an algorithm that can iterate much more efficiently.

提升RL效率 Improving RL Efficiency

Chelsea

实际上,有办法让这些算法更加高效。这些用于语言模型的强化学习算法存在一些低效之处,比如很大的低效。首先,它们会在死胡同轨迹上花费大量时间。如果只是花费算力,这可能没问题,但在物理世界中会代价高昂。让我们看一个具体例子。假设我们希望机器人构建纸板箱并将它们堆叠在右侧。在这个轨迹中,机器人意外地抓住了两个紧贴在一起的箱子。如果我们让它继续,它只会继续尝试折叠那个箱子,而不是分开两个箱子。尝试将两个箱子折叠在一起并不是有用的数据,无法教会模型如何更好地完成任务。因此,这会在机器人尝试错误路径上浪费大量时间。与其花大量时间尝试该任务,我们会让人类干预并展示机器人该做什么以及如何从这种情况中恢复。你可以看到,这里有人正在远程操作并干预机器人,向它展示要从这种情况中恢复,它需要基本上尝试分开两个箱子。然后它把夹爪伸进去,看看机器人能否自主恢复——它不能——所以这个人再次干预,帮助它回到正轨,这样我们就能高效地利用机器人上的数据。

There are actually ways to make these algorithms a lot more efficient. There are a couple of inefficiencies, like large inefficiencies in these reinforcement learning algorithms for language models. The first is that they spend a lot of time on dead-end trajectories. Maybe this is okay if you're just spending compute on it, but it would cost a lot in the physical world. Let's look at a concrete example. Say we want a robot to construct cardboard boxes and stack them on the right. In this trajectory, the robot accidentally grabbed two boxes that are flush against each other. If we let it continue, it would just continue to try to fold that box rather than separate the two boxes. Trying to fold two boxes together isn't useful data that will teach the model how to get better at the task. So that would be wasting a lot of time on the robot attempting to go down the wrong path. Instead of spending a lot of time trying to do that task, we'll have a human intervene and show the robot what to do and how to recover from that situation. You can see here a human is teleoperating and intervening with the robot, showing it that to recover from this situation, it needs to essentially try to separate the two boxes. It then puts its gripper in, sees if the robot could autonomously recover—it doesn't—so the person intervenes again to help it get back on the right track, so we're efficiently using the data on the robot.

摊销价值估计 Amortizing Value Estimation

Chelsea

我们能做的第二件事是,PPO 和 GRPO 这类算法会对单个提示进行多次尝试。根据算法的不同,它们基本上是在估计这些不同响应中哪些是好的,哪些是坏的。即使对于单个提示,它们也会对该提示进行大约 10 次或 50 次的展开。它们这样做是因为试图估计这些不同尝试的价值,以便提高好事物的可能性,降低坏事物的可能性。但实际上,我们可以摊销这种成本,而不是试图为单个提示收集大量尝试。我们可以跨不同提示进行摊销,学习一个更通用的价值估计,判断什么好什么坏,并利用它通过自主经验进行改进。这看起来像是我们在大量机器人经验视频上训练一个通用价值函数。它可以学习诸如在尝试折叠衬衫时意外展开衬衫是坏事,并且是负向进展——用红色显示。或者如果它在取得正向进展,它也能识别出来。同一个价值函数还可以估计完全不同的场景中的好坏,比如从冰箱中取物品。这种通用价值模型,基本上预测成功所需的时间,可以显著减少从经验中学习改进所需的尝试次数。

The second thing we can do is that PPO and GRPO and these kinds of algorithms make many attempts at a single prompt. Depending on the algorithm, they're essentially trying to estimate for these different responses what is good and what is bad. Even for an individual prompt, they're going to roll out like 10 or 50 times for that prompt. They're doing this because they're trying to estimate the value of these different attempts, to then upweight or increase the likelihood of good things and decrease the likelihood of bad things. But we can actually amortize this cost rather than trying to collect a lot of attempts for a single prompt. We can amortize this across different prompts and learn a much more general value estimate of what's good and bad, and use this to improve with our autonomous experience. What this looks like is we can train a general-purpose value function on lots of videos of the robot experience. This can learn things like if it accidentally unfolds a shirt when it's trying to fold it, that's bad and making negative progress—shown in red. Or if it's making forward progress, it recognizes that as well. The same value function can also estimate what's good and bad for a completely different scenario, in this case retrieving an item from a fridge. This kind of general-purpose value model, which is basically predicting the time to success, can significantly reduce the number of attempts needed to learn how to improve from experience.

通用改进算法 General Improvement Algorithm

Chelsea

通过这两项对强化学习系统的改进,我们有了一个通用改进算法:在多样化数据上训练基础模型,然后从该模型收集经验,在必要时由人类干预以防止死胡同轨迹,然后训练一个通用的好坏估计——价值函数——并利用它来改进模型。通过这种改进,我们能够将基础模型微调到更高的性能水平。在制作拿铁的任务中,在这种情况下,我们将与人合作制作拿铁,机器人负责制作浓缩咖啡,人负责蒸牛奶。这就是模型的样子。

With these two improvements to a reinforcement learning system, we have a general improvement algorithm that trains a foundation model on diverse data, then collects experience from that with a human intervening as necessary to help prevent dead-end trajectories, and then trains a general-purpose estimate of what's good and bad—the value function—and uses that to improve the model. With this sort of improvement, we're able to fine-tune a foundation model to higher degrees of performance. In the task of making a latte, in this case we'll be making a latte in collaboration with a person, where the robot is in charge of making the espresso and the person is in charge of steaming the milk. This is what the model looks like.

机器人演示与可靠性 Robot Demonstrations and Reliability

Chelsea

模型直接利用机器人摄像头捕捉的图像来控制机器人的关节。我们可以看到,模型能够完成相当有挑战性的任务:插入 PA 滤器,等待适当的时间让浓缩咖啡流出,将蒸好的牛奶倒入杯中。而这项任务的最后一步实际上是最具挑战性的,它需要拿起一杯满满的拿铁,然后转移到杯垫上。这里展示的正是机器人直接看到的画面。你可以看到,策略非常精细,能够恰当而平稳地保持杯子平衡,确保拿铁不会洒出来。这让你感受到这类任务的难度。

The model is directly controlling the joints of the robot using the images from the robot's cameras as input. And we can see that the model is able to do the pretty challenging task of inserting the PA filter, waiting the appropriate amount of time for the espresso to dispense, pouring the steamed milk into the cup. And then the last part of this task is actually the most challenging, where it needs to take a very full latte cup and transfer that over to the coaster. And so here's actually the observation that the robot sees directly. And you can see that the policy is super delicate and able to balance the cup appropriately and smoothly so that the latte doesn't spill. So this gives you a sense of the difficulty of this kind of task.

Chelsea

回到可靠性这个问题,我们采用了这个策略,并且不止运行一次,而是连续运行了 13 个小时。我们主要想评估:这个策略不仅擅长制作一次拿铁,还能可靠地做到在现实世界中真正有用的程度吗?这是该过程的延时摄影。我们确实发现,机器人足够可靠,能够长时间工作而不频繁出错。

And kind of going back to this reliability question, we took this policy and we ran it not just once, but we ran it for 13 hours straight. And we basically wanted to evaluate: is this policy not only good at making a latte once, but can it do so reliably to the extent that it would be needed to be useful in the real world? And so here's a time lapse of that process. And indeed we found that the robot was reliable enough to be useful for long stretches of time without making mistakes frequently.

Chelsea

当然,同样的算法并不专门用于制作拿铁,因此我们也将其应用于其他场景。Dandelion 巧克力工厂离我们办公室只有几个街区。我们选取了他们通常由人工完成的工作流程,即制作这些纸板箱、贴标签并堆叠起来。我们训练机器人基本按照他们的真实工作流程操作,并使用我之前提到的强化学习算法进行训练,从而获得一个在制作、贴标和堆叠这些箱子方面更加可靠的策略。

Now the same algorithm isn't specific for making lattes of course, and so we also applied this to other applications as well. Dandelion Chocolate Factory is a few blocks from our office. And so we took a workflow that they typically have a person do, which is to construct these cardboard boxes, label them, and stack them. And we trained our robot to basically do exactly their real workflow, and trained it with the reinforcement learning algorithm that I talked about to get a policy that is far more reliable at constructing, labeling, and stacking these boxes.

Chelsea

然后我们也将该算法应用于折叠和关闭操作。在这种情况下,我们不仅想测试模型在一个环境中完成一项任务的能力,还想测试它在多种环境中的表现。这些是机器人从未见过的衣物,在一个从未见过的家中。它能够做到这一点,并长时间自主行动。

And then we also applied this algorithm to folding and closing as well. And we wanted in this case to not just test how well the model could do one task in one environment, but to do it in many environments. And so these are clothing items that the robot has never seen before in a home that's never seen before. And it's able to do so and act autonomously for an extended period of time.

Chelsea

视频并不总能展示一切。因此,我们也定量测量了这些模型的可靠性。我们既关心可靠性,也关心速度,比如每小时能制作多少个箱子。因此,我们将测量吞吐量,它结合了成功率和速度。我们发现,在从预训练到类似 SFT 阶段再到 RL 后训练阶段的训练过程中,成功率和吞吐量都大幅提升,尤其是仅从 RL 阶段就实现了约 2 倍的吞吐量提升,这表明我们如何通过强化学习获得更高的可靠性。

Now, videos don't always show everything. And so we also quantitatively measured the reliability of these models. We care both about the reliability as well as the speed, like how many boxes can it build per hour. And so we're going to measure throughput, which kind of couples both success rate and speed. And we find that over the phases of training from pre-training to like an SFT-like stage to an RL post-training stage, we see a drastic increase in success rate and indeed in throughput, and specifically around a 2x throughput just from the RL stage itself, showing how we can get much greater reliability from reinforcement learning.

Chelsea

对于浓缩咖啡任务,如果专门看成功率,我们在制作浓缩咖啡上实现了 90% 或超过 90% 的成功率。因此,这部分的关键结论是,我们可以开发出一种可扩展的方法,实现复杂机器人操作任务的高可靠性。在这个案例中,我们看到了通过经验和干预实现了 2 倍的吞吐量提升。但最重要的是,我们看到了如何在现实世界中人们真正关心的实际工作流程中实现长期自主性。我认为,这正是机器人在现实世界中有用所需的条件。

And for the espresso task, if you look specifically at the success rate, we achieved a 90% or over 90% success rate on making espresso. So the takeaways for this part is that we can develop a scalable recipe for high reliability of complex robotic manipulation tasks. And we saw in this case a 2x higher throughput from using experience and interventions. But most importantly, we saw how we can achieve long-term autonomy in real workflows that people actually care about in the real world. And this is what it's going to take, I think, for robots to be useful in the real world.

Chelsea

现在,还有更多的工作和机会。我们实际上只运行了几次算法改进的迭代,通过更多迭代,我们应该能够看到更大的改进或更高的可靠性。即使有了这种改进,机器人仍然会犯错,而且仍然比人慢。因此,在开发更强大的方法方面还有很大的改进空间。

Now, there's also a lot more work and a lot more opportunities. We actually only ran a few iterations of improvement of this algorithm, and with more iterations, we should be able to see even greater improvement or even greater reliability. And even with this improvement, the robot still makes mistakes. It's also still slower than people. And so, there's a ton of room for improvement for developing even more powerful recipes.

长期任务记忆 Memory for Long-Horizon Tasks

Chelsea

所以,我们已经看到了这些不同工作流程的长期自主性。但还有另一个要素我想谈谈,它能让机器人在长时间内自主且有用,那就是记忆。你可能会惊讶地发现,大多数最先进的机器人基础模型都没有记忆或上下文。它们只是基于当前的传感器观测、当前的摄像头读数进行操作,并据此预测动作。实际上,没有记忆你也可以完成短期的运动技能和重复性任务。我之前展示的视频也没有任何上下文。但如果你想完成一个涉及多个不同步骤的长期任务,那么记忆对于跟踪你已完成步骤的进度至关重要。

So, we've seen long-term autonomy for these different workflows. But there's actually one more ingredient that I'd like to talk about for enabling robots to be autonomous and useful for long periods of time. And that ingredient is memory. So, you might be surprised to hear that most state-of-the-art foundation models for robotics have no memory or no context. They're just operating on the current sensor observations, the current camera readings, and predicting actions based off of that. And you actually can do short motor skills, you can do repetitive tasks without memory. The videos that I showed before didn't have any context either. But if you want to do a long task that involves multiple different steps in sequence, then memory is critical for tracking progress of the steps that you've completed so far.

Chelsea

如果记忆对于这类长时程任务至关重要,那么为什么这些模型没有任何上下文或记忆呢?这有几个技术原因,我将讨论其中一个:如果你天真地处理记忆,试图将上下文(比如视频)输入机器人基础模型,假设你只输入 10 秒的视频。也许这 10 秒的视频以 50 赫兹采样,这是机器人技术中常见的控制频率,并且你输入机器人的所有四个摄像头流,每张图像使用大约 256 个 token。这相当于向模型输入 50 万个 token,这是大量的 token。而试图实时将这么多 token 输入模型目前是相当困难的。即使你降采样到每秒一帧,你仍然要向模型输入 10,000 个 token,至少目前对这些模型来说成本过高,而且这仍然只有 10 秒的记忆。

So if it's critical for doing these kinds of long-horizon tasks, then why don't these models have any context or memory? There's a couple reasons for this that are technical, and I'll talk through one of them, which is that if you naively approach memory and try to feed in context like pass video to a robot foundation model, say that you would just pass in 10 seconds of video. Maybe this 10 seconds of video is sampled at 50 hertz, which is a common control frequency in robotics, and you feed in all four camera streams on the robot and you use around 256 tokens per image. This corresponds to passing in half a million tokens into your model, which is a lot of tokens. And trying to do that in real time into your model right now is quite challenging. Even if you subsample to one frame per second, you're still going to be passing in 10,000 tokens into your model, which at least right now is prohibitively expensive for these models, and that's still only 10 seconds of memory.

Chelsea

我没有时间详细讨论我们在这里具体做了什么的技术细节,但我们也为这个上下文问题开发了一个解决方案,具体来说,我们开发了一个具有多时间尺度记忆的系统。第一个是短期视频记忆,大约有 10 秒的视频记忆,但它的计算方式比天真地将其输入模型要高效得多。对于更长的记忆,跨越几分钟或几小时的记忆,我们不一定需要过去历史的精确视频。因此,我们以文本形式表示这些部分的记忆,总结发生的事情,然后将过去 10-15 分钟内发生的事情的压缩文本摘要也纳入模型。通过这种多时间尺度的记忆,我们能够让机器人完全自主地执行持续 10 或 15 分钟的任务。

So I don't have time to go into the technical details of exactly what we did here, but we also developed a solution for this context problem, and specifically we developed a system that has memory at multiple time scales. The first is a short-term video memory that has about 10 seconds of video memory but is done so and computed much more efficiently than naively passing it into the model. And then for longer memory, for memory that spans multiple minutes or multiple hours, we don't necessarily need video of exactly what happened in that past history. And so instead, we represent memory for those parts in text where we summarize what happened in text space and then incorporate that much more compressed textual summary of what happened over the past 10-15 minutes into the model as well. And with this sort of memory at multiple different time scales, we're able to enable robots to do tasks that can operate for 10 or 15 minutes at a time completely autonomously.

长期自主性任务 Long-term autonomy task

Chelsea

与上一张幻灯片不同的是,这个任务不是重复性的。这是一个需要 10 到 15 分钟的任务,涉及清洁厨房。机器人不是一遍又一遍地重复做浓缩咖啡。它需要用海绵擦拭台面,用纸巾擦干台面,扔掉纸巾。接下来,它会把芥末放进冰箱。然后,它会把盘子放进橱柜,清洗水槽里的一些脏盘子,等等。通过整合记忆,它能够完成一个需要跟踪所有这些不同步骤的任务,以清洁厨房,并完全自主地成功运行 10 到 15 分钟。

And what's different from the previous slide is that this task isn't repetitive. This is a 10 to 15 minute task that involves cleaning a kitchen. The robot isn't just repeatedly making espresso over and over again. It involves wiping the counter with a sponge, drying the counter with a paper towel, throwing away the paper towel. Next, it's going to put away the mustard into the fridge. Then, it will put the dishes away into the cabinet, wash some of the dirty dishes in the sink, and so on. By incorporating memory, it's able to do a task that requires keeping track of all of these different steps, done to clean the kitchen, and successfully operate for 10 to 15 minutes, completely autonomously.

Host

太棒了。

Great.

长期自主性的要素 Ingredients for long-term autonomy

Chelsea

这些是长期自主性的几个要素。现在,我想在此基础上,将这些要素整合到一个通用模型中,这个模型既能完成我之前展示的所有任务,也能在单一模型中实现,并且还能做其他一些事情。要思考开发这样一个通用模型,我认为将机器人技术置于通用 AI 发展的其他时间线中来看是很有帮助的。

So those were a couple ingredients for long-term autonomy. Now, I'd like to build on that and actually take those ingredients and put it into a general purpose model that can do everything that I showed before, but also can do that in a single model and can do some other things as well. And to think about developing such a general purpose model, I think it's really helpful to contextualize where robotics is at within the timeline of other developments in generalist AI.

通用AI的演进 Evolution of generalist AI

Chelsea

如果我们思考通用 AI 系统在过去 15 年是如何演进的,我认为第一个重要里程碑是在 2012 年,当时我们看到一个从头训练的深度学习系统在外部基准测试中登顶,而该基准测试的所有先前方法都是专门为该应用设计的。具体来说,这是 ImageNet 基准,专为图像分类设计。这是深度学习系统首次超越那些更专门的系统。这是一个更通用的算法,并非专门为图像识别设计。

If we think about how generalist AI systems have evolved over the past 15 years, I think the first major milestone was in 2012 when we saw that a deep learning system trained from scratch topped an external benchmark, and all of the previous methods for that benchmark were specifically designed for that application. Specifically, this was the ImageNet benchmark, designed for image classification. This was the first time that a deep learning based system outperformed those more specialist systems. This is a much more general algorithm that wasn't specifically designed for image recognition.

Chelsea

几年后,我们发现我们不只是从头训练算法,而是能够获得预训练模型,这些模型对下游任务的微调很有用。采用在 ImageNet 上预训练的模型,然后在下游任务上进行微调成为常态。我们确实看到使用像 BERT 或 ImageNet 预训练模型这样的预训练模型能获得更好的性能。

Then just a couple years later, we found that we weren't just training algorithms from scratch, but we were able to get pre-trained models that are useful for fine-tuning to downstream tasks. It became the norm to take a model pre-trained on ImageNet and then fine-tune it on a downstream task. We actually saw better performance from using that pre-trained model like BERT or an ImageNet pre-trained model.

Chelsea

从那时起,通用 AI 模型的下一个重大阶段和转变不是使用预训练模型,而是从预训练微调机制转向直接使用通用模型的机制。这始于像 GPT-2 这样的模型,当然,如今我们交互的几乎所有模型都无需微调即可开箱即用,至少大多数消费级模型是这样。实际上还有其他模型仍然大量使用微调。

From there, the next big phase and transition in generalist AI models wasn't using pre-trained models but moving from a pre-training fine-tuning regime to a regime where we're just using generalist models out of the box. This was with models like the start of GPT-2, and of course almost all the models that we interact with today work just out of the box without fine-tuning, or at least most of the consumer models. There are actually other models that still use a lot of fine-tuning.

Chelsea

我想强调的另一个里程碑是在 2021 年,我们看到了这些模型中组合泛化的初步迹象。一个具体的例子是 DALL-E,我将在后面的幻灯片中详细讨论。这就是通用 AI 在过去 15 年中的发展方式。

Another milestone I want to highlight was in 2021, where we saw the first signs of compositional generalization in these models. One specific instance of that was with DALL-E, and I'll talk a little bit more about that in a later slide. So this is how generalist AI has advanced over the past 15 years.

物理AI时间线 Physical AI timeline

Chelsea

与此同时,如果我们考虑物理 AI,即使在 2021 年,也就是三年前,从事机器人技术的人通常都会为单个项目从头收集定制数据集,并从头开始训练。这类似于从头收集 ImageNet 并训练,或者在你刚刚从头收集的数据集上训练。如果你想开发一个通用模型,但你必须为每个项目从头收集数据集,那么你可能不会取得太大进展。

Meanwhile, if we think about physical AI, even just three years ago in 2023, it was extremely common for people working on robotics to collect a bespoke data set from scratch for an individual project and train from scratch on that data set. This is analogous to collecting ImageNet from scratch and training on it, or training on the data set that you just collected from scratch. If you want to develop a general purpose model, and you have to collect the data set from scratch for every single project, you're probably not going to make a lot of progress.

Chelsea

直到几年前,我认为我们在这个时间线上还处于相当靠左的位置。最近,我认为我们处于 2014 年的阶段,拥有一些好的预训练模型,但我们还没有真正进入右侧的机制。那么我们如何达到右侧的机制呢?具体来说,我们如何开发一个开箱即用且具有组合泛化能力的单一通用模型?

Until just a few years ago, I think we were pretty far on the left of this timeline. More recently, I think we've been in the 2014 phase where we have some good pre-trained models, but we haven't really been truly in the regime on the right. So how do we get to that regime on the right? Specifically, how do we develop a single general purpose model that works out of the box and also shows compositional generalization?

两大目标:开箱即用与组合泛化 Two goals: out-of-the-box and compositional generalization

Chelsea

这有两个目标。第一个是开箱即用的模型。这类似于从 BERT 到 GPT 的转变。目前,最好的机器人性能,如果你想让你的模型在给定任务上表现最佳,总是需要微调。我一开始展示的一些视频是经过微调的模型,用于解锁等任务。我们做的其他关于衡量人机迁移的工作也需要微调才能获得最佳性能。当然,我展示的所有带有 RL 后训练的视频也都是针对单个任务进行微调,以在制作浓缩咖啡等任务上获得最佳性能。但如果你必须微调模型,你实际上并没有得到一个通用模型,因为你必须为每个单独的任务进行微调。我们的第一个目标是朝着一个能够真正完成所有你希望它做的事情的单一通用模型迈进。

This has two goals. The first is out-of-the-box models. This is analogous to going from BERT to GPT. Right now, the best robot performance, if you want your model to perform the best it can on a given task, always requires fine-tuning. Some of the videos I showed at the beginning were fine-tuned models to do things like unlocking a lock. Other work we've done on measuring human-to-robot transfer also needed fine-tuning to get the best performance. Of course, all the videos I showed with RL post-training were also fine-tuning on an individual task to get the best performance on something like making espresso. But if you have to fine-tune a model, you actually aren't getting a general purpose model for the things you want it to do, because you have to fine-tune it for each individual thing. Our first goal is to move towards a single general purpose model that can actually do all of the things you want it to do.

Chelsea

我提到的第二个目标是组合泛化。这受到 2021 年 DALL-E 结果的启发。我认为这是一个非常重要且令人兴奋的里程碑,因为它实现了组合泛化。具体来说,当你拥有组合泛化能力时,当你能够基本上将牛油果和椅子的概念联系起来,并表明你可以将这两者结合起来,这意味着模型至少对牛油果是什么和椅子是什么有某种概念性理解,以至于它可以将它们组合成同时体现这两个概念的东西。其次,这意味着你拥有一定程度的数据效率,你的数据不需要覆盖数据中表示的所有可能的概念组合。你不需要在数据集中包含牛油果椅子的图片就能生成这样的东西。或者你不需要包含你可能在部署时要求模型做的其他事物的组合。

The second goal I mentioned is compositional generalization. This is inspired by the DALL-E result from 2021. I think this was a really important and exciting milestone because of the compositional generalization it achieved. Specifically, when you have compositional generalization, when you can basically bridge the concept of an avocado and a chair and show that you can combine those two, it means that the model has at least some kind of conceptual understanding of what an avocado is and what a chair is, to the point that it can combine them into something that exhibits both concepts at the same time. Second, it means that you have some degree of data efficiency where your data doesn't need to cover all possible combinations of concepts represented in your data. You don't need pictures of avocado chairs in your data set in order to generate something like this. Or you don't need combinations of other things that you might ask the model to do when it's deployed.

Chelsea

即使在 2021 年,它并不完美,但这些组合泛化的迹象对于展示模型的这两个属性来说确实令人兴奋。所以我们有两个目标:开箱即用的模型和组合泛化。

Even back in 2021, it wasn't perfect, but these signs of compositional generalization were really exciting for demonstrating these two attributes of the model. So we have these two goals that we'd like to do: an out-of-the-box model and compositional generalization.

训练配方 Training Recipe

Chelsea

开发这类模型的成熟配方是,首先获取足够大且多样化的数据集,其次训练一个容量足够的模型。所以我们要做的就是这两件事。我们会尝试使用所有可用的数据,包括非常多样的机器人演示数据,甚至质量很低的演示数据。还会包括策略 rollout 数据,也就是机器人尝试执行任务的数据。基本上,之前用于强化学习的所有训练数据都会纳入训练配方。我们还会加入人类视频,以及来自网络的数据。基本上,我们拥有的所有数据都会用上。

Now the tried and tested recipe for developing this kind of model is to first take a sufficiently large and diverse data set and second, train a model with sufficient capacity. So what we're going to do is we're going to do that. We're going to try to use all of the data that we have available. This includes really diverse robot demonstration data, including really low-quality demonstration data. It's also going to include policy rollout data—basically, attempts from the robot of doing the task. All of the training data that was used for reinforcement learning for the previous tasks will be included in the training recipe. We're also going to include videos of humans, and we're also going to include data from the web. Basically, all of the data that we have.

Chelsea

然后,为了训练一个容量足够的模型,我们当然会训练一个足够大的模型,但要拟合如此异构的数据,我们发现特别重要的是给模型提供它预测动作所需的全部上下文提示。我们发现这个想法是使用这类数据和这种异构程度数据的关键解锁点。

Then, to train a model with sufficient capacity, of course we'll train a model that's large enough, but to fit data that's so heterogeneous, we also find it particularly important to prompt the model with all of the context that it needs in order to predict actions. We found that this idea was really the key unlock to using this kind of data and data of this degree of heterogeneity.

Chelsea

具体来说,我们会训练一个基础模型,输入包括我之前提到的记忆、要做什么的指令,但也会输入子任务指令,即接下来应立即做什么。它还会输入元数据,指示数据质量、回合长度等。这些元数据为模型预测下一个动作提供了更多信息。然后,我们还会选择性地用子目标图像作为提示来训练模型,这基本上是在说,几秒后你应该尝试达到类似这张图像的状态。

Specifically, what this looks like is we're going to train a foundation model that takes as input the memory that I mentioned before, an instruction of what to do, but it's also going to take as input a subtask instruction of what the next immediate thing it should do is. It'll also take as input metadata that indicates the quality of the data, the length of the episodes, and so forth. This metadata gives it a lot more information about how it should predict the next action. Then optionally, we'll also train the model with a subgoal image as a prompt, essentially saying that a few seconds from now you should try to reach something that looks like this image.

Chelsea

通过这种详细的提示,我们发现模型确实能利用更多异构数据,稍后我会展示一些对比,真正说明这一点的重要性。

With this detailed prompting, we find that the model can really make use of much more heterogeneous data, and I'll show some comparisons later that really show how important it is.

部署与结果 Deployment and Results

Chelsea

现在,要实际部署这个模型,我们需要提供子任务构建和子目标图像等。这样,我们可以训练一个高层策略来预测子任务指令,比如下一步做什么,对于清理厨房的任务,下一个子任务是什么。我们还会额外训练一个世界模型,为机器人下一步该做什么生成图像,作为子目标图像条件。

Now, to actually deploy this model, we need to provide things like subtask construction and subgoal images. With that, we can train a high-level policy that predicts the subtask instruction—like what to do next, what is the next subtask for the task of cleaning the kitchen. We'll additionally train a world model to generate images for what the robot should do next as subgoal image conditioning.

Chelsea

这样,我们会在所有可用的多样化数据上训练一个具有这些属性的单一模型。以下是一些该单一模型能做什么的例子。所有这些视频都来自同一个模型,具体来说,我们称之为 PIO7 模型。左边可以看到它在折叠一件有领衬衫。右上方,它在执行一个非常精确的组装步骤,需要将螺丝插入并钻入机器人手臂。右下方,机器人在更换垃圾桶中的垃圾袋。

With this, we'll train a single model with those attributes on all of the diverse data that we had available. Here are some examples of what that single model can do. All of these videos are from a single model, specifically a model that we called the PIO7 model. On the left, you can see it doing things like folding a collared shirt. On the top right, it's doing a really precise assembly step where it needs to insert a screw and drill that screw into a robot arm. On the bottom right, the robot is replacing a trash bag in a trash can.

开箱即用性能 Out-of-the-Box Performance

Chelsea

我们一开始有两个目标。第一个是迈向开箱即用的模型。即使那些视频也表明,开箱即用时模型就能做很多事情。但真正关键的问题是,这个预训练模型与我之前提到的专门为制作咖啡、搭建箱子而训练的专业模型相比如何。

We had two goals at the start of this. The first was to move towards an out-of-the-box model. Even those videos showed that out of the box, the model is able to do quite a bit. But really, the key question here is how does this pre-trained model compare to the specialists that were trained specifically for coffee making, specifically for box building that I talked about previously.

Chelsea

如果我们衡量这个单一 PIO7 模型与微调后的 PIO6 模型的吞吐量和成功率,我们会看到,总体而言,单一的预训练 PIO7 模型匹配或超过了那些通过强化学习后训练为下游任务开发的专业模型。所以我们看到它能够匹配专业模型的性能。这也适用于 SFT 专业模型,而不仅仅是 RL 后训练模型,这表明我们确实拥有一个单一模型,能够开箱即用地以非常高的性能完成许多不同任务。

If we measure the throughput and the success rate of this single PIO7 model versus the fine-tuned PIO6 model, we see that across the board, the single pre-trained PIO7 model matches or outperforms the fine-tuned specialists that were developed with reinforcement learning post-training for those downstream tasks. So we see that it's able to match the performance of specialists. It also holds for SFT specialists, not just RL post-trained models as well, suggesting that we do indeed have a single model that can do a lot of different tasks with a really high degree of performance out of the box.

组合泛化 Compositional Generalization

Chelsea

好,这就是开箱即用模型的第一个目标。第二个目标是组合泛化。有几种不同的衡量方式,在机器人技术中你可能尝试多种方式组合概念。我们想做的第一个测试是看机器人能否与像空气炸锅这样相当罕见的电器互动。

Okay, so that was the first goal of out-of-the-box models. The second goal is compositional generalization. There are a few different ways to measure this, and many different ways you might try to combine concepts in robotics. The first test that we wanted to do was to see if a robot could interact with an appliance that's quite rare, like an air fryer.

Chelsea

这是一个例子。我们基本上想看看它能否打开空气炸锅,把红薯放进去,然后关闭空气炸锅。我们选择这个是因为我们认为数据集中没有任何空气炸锅。我们没有故意收集任何包含空气炸锅的训练数据。在对数据集进行分析后,我们实际上发现我们的数据集非常多样化,确实包含三个包含空气炸锅的回合。我们预计这些数据可能没有影响,即使我们不包含那三个回合,它可能仍然有效。

This is an example. We basically wanted to see if it could open an air fryer, put a sweet potato in the air fryer, and close the air fryer. We picked this because we thought that the data set didn't have any air fryers in it. We didn't intentionally collect any training data with air fryers. After we did some analysis on the data set, we actually found that our data set was so diverse that it did actually have three episodes with air fryers in it. We expect that they likely weren't having an impact, and that even if we didn't include those exact three episodes, it likely would still work.

Chelsea

我们总体上发现,机器人能够与训练数据集中几乎没有代表的电器互动,并将与之互动的技能(如打开、关闭等)与这个它从未见过的物体结合起来。在像 Lucy 那样指示它之后,我们可以训练一个高层策略来完全自主地完成这个任务。你可以在这个视频中看到机器人这样做。

What we found generally is that the robot was able to interact with an appliance that was hardly represented at all in the training data set, and combine the skill of interacting with it—like opening it, closing it, and so forth—with this object that it hasn't seen before. After instructing it like Lucy did, we can train a high-level policy to do this task fully autonomously. You can see the robot doing that in this video.

跨平台泛化 Cross-Platform Generalization

Chelsea

这是组合泛化的第一种形式。我们想研究的第二个组合泛化测试是,我们能否在任务和机器人平台之间进行组合泛化。我们想采用一个名为 barm 机器人的平台。它实际上是一个非常大的工业机器人平台。我们想看看它能否折叠衣物,尽管我们没有在这个机器人平台上收集任何折叠数据。

That's the first form of compositional generalization. The second compositional generalization test that we wanted to look at is whether we could compositionally generalize between tasks and robot platforms. We wanted to take a robot platform called a barm robot. It's actually a very large, kind of industrial robot platform. We wanted to see if it could fold clothes, despite the fact that we didn't collect any folding data on this robot platform.

Chelsea

具体来说,我们有在左边所示的机器人平台上折叠衣物(如折叠衬衫)的数据。然后我们想看看,开箱即用,在右边这个非常不同的机器人平台上没有收集任何折叠数据的情况下,机器人能否成功完成任务。我们在这个视频中看到的是,我们确实看到它以这种方式进行了组合泛化。

Specifically, we had data of folding clothes like folding a shirt on the robot platform pictured here on the left. Then we wanted to see, out of the box, without collecting any folding data on this very different robot platform on the right, could the robot successfully do the task. What we see in this video is that we indeed saw that it kind of compositionally generalized in this manner.

Chelsea

我们第一次看到机器人这样做时,都惊呆了,因为这项任务没有训练数据。这里的机器人与其他机器人非常不同,不仅在尺寸上,而且在机器人连杆的长度、机器人关节的配置等方面都不同。

The first time we saw the robot do this, we were floored, because there was no training data for this task. The robot here is quite different from the other robot, not just in size but also in the lengths of the linkages of the robot, in the configuration of the joints of the robot, and so forth.

机器人折叠演示 Robot Folding Demo

Chelsea

这是一个 1 倍速的视频,所以不是最快的。显然,如果模型没有见过某些训练数据,比如这真的是机器人第一次叠衬衫,那它可能需要尝试几次,但最终会叠好。你还可以在左上角看到生成的子目标图像。这些基本上是模型试图生成能推进折叠任务的图像,然后这些图像会作为输入传给模型。现在我们看到叠好的衬衫了。我觉得它最后会做几个小修正,让衬衫更平整一些。

And this is a 1x speed video, so it's not the fastest thing. And obviously, if you haven't seen any training data on something, you might not—if it's literally the robot's first time folding a shirt—it might take a few attempts, but eventually it will get to the folded shirt. You can also see the generated subgoal images on the top left. Those are basically the model's attempts to generate images that will make progress on the folding task, and then those are passed as input to the model. And we see the folded shirt here. I think it's going to make a couple of small corrections at the end to try to make it a little bit smoother.

Host

酷。

Cool.

Chelsea

所以,这里的要点是,无论是在语言与物体的交互,还是在任务与机器人的交互方面,我们都看到了这个模型表现出强烈的组合泛化迹象。

So, the takeaway here is that both in terms of language-object interactions and in terms of task-robot interactions, we see strong signs of compositional generalization in this model.

Host

好的。

Okay.

定量结果 Quantitative Results

Chelsea

然后从定量上看,我们也看到,当我们使用像 PIO7 这样更先进的模型时,在这个未见过的平台上叠毛巾和叠衬衫的性能显著提升,甚至接近人类遥操作的水平,尽管我们没有针对叠衣服的机器人专用训练数据。

And then quantitatively, we also see that as we get to these more advanced models like the PIO7 model, the performance of folding towels and folding shirts on this platform that hasn't been seen before increases dramatically, and it even approaches the performance of human teleop, despite the fact that we didn't have any robot-specific training data for folding clothes.

消融研究 Ablation Studies

Chelsea

然后对于我们这里做的最后一个实验,我认为这可能是最有趣的实验。我们想测试我提到的两个要素有多重要:多样化的数据有多重要,以及这种能力或详细提示对我展示的结果有多重要。如果我们从模型训练中移除最多样化的数据,如灰色所示,我们会发现在保留任务上的性能急剧下降。而如果我们只是随机取出 20% 的多样性较低的数据,性能只下降一点点。这表明,拥有真正多样化的数据在使其泛化到新任务方面起着重要作用。

And then for the last experiment that we did here, I think this is perhaps the most interesting experiment. We wanted to test how important are the two ingredients that I mentioned: how important is diverse data, and how important is this sort of capacity or detailed prompting for the kinds of results that I showed. And so if we remove the most diverse data from the model training, shown in the grayish color, we find that the performance on held-out tasks decreases dramatically. Whereas if we just take out a random 20% of the data that's less diverse than the most diverse subset, the performance only decreases a little bit. And so this suggests that actually having really diverse data plays an important role in enabling it to generalize to new tasks.

Chelsea

然后我们还尝试消融了用元数据提示模型这一因素。在这个实验中,我们比较了有和没有元数据提示的情况。有提示的用黄色显示,没有提示的用灰色显示。有提示时,效果显著提升。但最有趣的是,当你添加越来越多的数据,特别是越来越多的低质量数据时,性能如何变化。在没有元数据提示的情况下,当从 80% 的数据增加到 100% 的数据时,性能实际上下降了,这也许并不太令人惊讶,因为你是在向数据混合中添加低质量数据。而有了元数据提示,当添加这些低质量数据时,性能反而提升了,这表明有了这种提示,即使从低质量数据中也能榨出更多价值。

And then we tried to also ablate the fact that we are prompting the model with metadata. For this experiment, we looked at with and without prompting with metadata. With prompting is shown in yellow, and without prompting is shown in the gray color. And with prompting, it helps kind of significantly. But the most interesting thing is if you look at when you add more and more data, and specifically as you add more and more low-quality data, what is the performance. And without metadata prompting, when you add lower-quality data from 80% data to 100% data, the performance actually decreases, which is perhaps not too surprising because you're adding low-quality data to your data mixture. Whereas with the metadata prompting, the performance actually increases when you add that low-quality data, suggesting that it's actually able to get a lot more juice out of even low-quality data when you include this kind of prompting.

Host

酷。

Cool.

关键要点 Key Takeaways

Chelsea

所以这里的要点是,我们发现我们能够训练一个单一的模型来控制机器人,其性能达到或超过专门的后期训练模型。有点像从类似 BERT 的预训练模型,走向像 GPT 那样真正开箱即用的模型。我们还看到了类似 Dolly 的强烈组合泛化迹象。例如,在组合泛化技能应用于家电和新机器人方面,这些方式在训练数据中从未见过。

So the takeaways here are that we found that we're able to train a single model to control the robots that matches or exceeds the performance of specialized post-trained models. Kind of like going from a BERT-like pre-trained model to a model that really works out of the box like GPT. We also saw strong signs of compositional generalization in a Dolly-like way. For example, in compositionally generalizing skills applied to appliances and skills applied to new robots in ways that weren't seen in the training data.

Chelsea

然后我展示的所有视频和实验都只是对模型进行开箱即用的评估,没有任何后期训练。论文和在线技术报告中有更多的实验和细节。

And then all the videos and experiments that I showed were just evaluating the model out of the box without any post-training. And the paper and the technical report online have a lot more experiments and a lot more details.

当前状态与部署 Current Status and Deployment

Host

好的,我们已经讨论了长期自主性。然后我们展示了如何在一个通用模型中发展这一点。我们现在处于什么阶段?

Okay, so we've talked about long-term autonomy. We then showed how we can kind of develop that in a single general-purpose model. Where are we at now?

Chelsea

我要提到的第一点是,如果我们回到通用人工智能的时间线,我认为我们现在已经坚定地将物理智能放在了这条时间线的右侧。对于机器人和物理智能来说,我们坚定地处于类似 GPT 和 Dolly 的时代,这非常令人兴奋,而且我认为我们仅仅几年就达到了这个阶段。

The first thing that I'll mention is if we go back to the timeline of generalist AI, I think that we now kind of firmly have physical intelligence on the right side of this timeline. We're kind of firmly more in like a GPT and Dolly-like era for robotics and physical intelligence, which is really exciting, and I think that we've kind of went there in just a few years.

Chelsea

最后,我们还有这些模型实际上已部署在现实环境中。所以顶部的两个视频实际上是两家 YC 公司,Ultra 和 Weave,它们采用了 PI 模型,并进行了后期训练,以在部署中执行诸如折叠衣物和仓库包装等任务。左下角的视频是我之前展示过的,这种模型适用于非常多样化的机器人形态。顶部和左侧的是更标准的双臂平台。但它也可以适应无人机、四旋翼、手术机器人,以及右下角的拖拉机等。所以这真正展示了物理智能不仅能在演示和研究等方面产生影响,还能在现实部署中产生影响。

And lastly, we also have these models that are actually deployed in real-world circumstances. So the two videos on the top are actually two YC companies, Ultra and Weave, that have taken PI models and post-trained them to do in deployment tasks like folding laundry and packaging in a warehouse. The video on the bottom left is the video that I showed previously, and this kind of model works for a really diverse set of robot embodiments. The ones on the top and the left are a more standard bi-manual platform. But it also can be adapted to things like drones, quadcopters, surgical robots, and on the bottom right for things like tractors. And so this is really truly showing how physical intelligence can make an impact not just in demos and research and so forth, but actually in real-world deployment.

Chelsea

而且我认为,随着我们前进,我们将开始看到越来越多的机器人真正部署在物理世界中,这得益于过去几年我们看到的所有进步。

And I think that as we go, we'll start to see more and more robots actually deployed for real in the physical world with all the advances that we've been seeing over the past few years.

Host

太棒了。

Awesome.

招聘与问答 Hiring and Q&A

Chelsea

所以最后我要厚着脸皮提一下,我们 Physical Intelligence 正在招聘。如果你对我谈到的某些内容感到兴奋,我们鼓励你看看一些开放的职位并申请。而且,我们肯定有时间回答问题,也很乐意听取大家的想法。谢谢。

So the last thing that I'll mention shamelessly is that we are hiring at Physical Intelligence. So if you're excited about some of the stuff that I talked about, we encourage you to take a look at some of the open roles and apply. And yeah, definitely have time for questions and happy to get all your thoughts. Thanks.

Host

好的,第一个问题是,我们距离机器人领域的 ChatGPT 时刻还有多远,那会是什么样子?

Okay, so the first question is how far away are we from a ChatGPT moment for robotics, and what will that look like?

Chelsea

实际上,我会先回答第二部分,那就是我不确定它是否会像我们在语言模型中看到的 ChatGPT 时刻那样,比如 ChatGPT 在五天内就获得了一百万新用户。我认为物理模型的分发渠道将会更慢,不幸的是,因为你确实需要一台物理机器人。而且我认为,对于像 Waymo 这样的东西,我们看到它的推广确实令人难以置信,但实际在物理设备上部署仍然需要时间。所以我不知道我们是否会有一个像 ChatGPT 那样具有分发规模的单一时刻。

So I'll start with the second part actually, which is that I'm not sure it will really look like the ChatGPT moment that we saw in language models, which is that with something like ChatGPT, we saw it pass like a million new users in five days. I think that the distribution channel for physical models is going to be slower, unfortunately, because you actually need a physical robot there. And I think that we've seen for something like Waymo, the rollout has been incredible to see, but it still takes time to actually deploy things on physical devices. So I don't know if we'll have a single moment that has the distribution that ChatGPT had.

模型能力与视野 Model Capabilities and Horizon

Chelsea

与此同时,就这些模型的能力而言,我认为我们确实开始达到它们在现实世界中真正有用的程度。而达到类似 ChatGPT 的能力,在接下来的几年里非常有望实现。

At the same time, in terms of the capabilities of these models, I think we're really starting to get to the point where they're actually useful in the real world. And getting to the kind of capabilities of ChatGPT is very much on the horizon in the next few years.

向通用策略过渡 Transition to Generalist Policies

Host

好的。第二个问题是:小团队应该何时从扩展特定场景模型转向通用策略?这种转变实际上是什么样的?哪些信号表明时机已到?

Cool. Well, the second question is: when should a small team switch from scaling per-site models to a generalist policy, and what does that transition actually look like? What signals tell you it's time?

Chelsea

这是个好问题。至少我认为,即使一开始就直接采用通用策略并进行微调,也可能非常有效。幸运的是,许多通用策略实际上非常强大且开源。例如,PI Zero 和 PIO5 模型就是开源的。我们已经看到很多人从这些模型中获益良多。我们还在与许多合作伙伴合作,比如拖拉机公司、Ultra 和 Weave,利用我们最新的模型,为他们的应用榨取更多价值,使模型更强大。所以即使一开始,我认为你也可以使用它们。唯一不会使用它们的情况是环境确实受限。我和一些从事手术机器人的人聊过,他们在没有互联网连接、GPU 很差的地下室手术室里工作。有时使用更大的模型确实很困难。但你仍然可以在工作站上用这些模型进行本地推理。所以我认为,直接采用 PIO5 或你喜欢的模型并进行微调是可行的。我们会看到很多小公司这样做。在让这些机器人在现实世界中真正应用这项技术方面,还有很多工作要做。

This is a good question. At the very least, I actually think that just starting with a generalist policy and then fine-tuning it, even right off the bat, can be really effective. Fortunately, a lot of generalist policies are actually really powerful and open source. For example, the PI Zero and PIO5 models are open source. We've seen a lot of people get a lot of use out of those models already. We're also working with many partners, like the tractor company, Ultra, and Weave, to take our most recent models and get even more juice out of them, making them even more powerful for their own applications. So even right off the bat, I think you can use them. The only scenario where I wouldn't use them is if you're really in a constrained environment. I've talked to folks working on surgical robots in an operating room in the basement with no internet connection and a really bad GPU. Sometimes it's just really hard to use a larger model. But you can still do local inference on a workstation with these models. So I think right away, taking PIO5 or your favorite model and fine-tuning it is the way to go. We'll see lots of small companies doing this. There's so much work to do in actually getting these robots to work with this technology in the real world.

博士与业界 PhD vs Industry

Host

太好了。下一个问题是:鉴于机器人技术在工业界发展如此迅速,如今攻读博士学位有哪些真正的优势和劣势,尤其是对于之后想进入工业界的人?

Great. The next question is: given how fast robotics is moving in industry, what are the real advantages and drawbacks of doing a PhD today, especially for someone who wants to go into industry afterwards?

Chelsea

我原本没打算读博士。我一直计划直接进入工业界。我的父母是工程师,在工业界工作,我认为产生影响的方式是去公司。我爸爸甚至说他不会雇佣有博士学位的人,所以我想如果找不到工作,也许不该读博。但他是在不同的领域,土木工程。与此同时,我认为博士学位是一个绝佳的机会,我很喜欢我的博士经历。这很大程度上取决于导师和你所做的事情。但博士学位是学习如何处理不确定性、如何选择好的研究问题的绝佳机会。在研究中,没有人给你问题;你必须自己选择,而且你不知道所选问题在六个月、两年还是十年内能否取得进展。你学会应对这种不确定性。这非常有用。这也是一个做出色研究的机会,并且有很大的自由去研究你最感兴趣的内容。所以今天,它仍然是学习不确定性的绝佳机会,这在初创环境和 AI 前沿非常有用,因为没有人知道让模型更强大的最佳途径。与此同时,工业界也有许多绝佳的机会。开发我展示的一切不仅仅是研究;还有一整套软件栈需要在机器人上可靠运行,还有硬件、机器学习基础设施、数据基础设施等等。很多工程工作并不需要博士学位。在研究方面,也常常有机会参与,而且如今很多研究也是工程性的。所以这取决于个人;这是一个非常个人化的决定。即使在今天,回顾过去,我可能还是会想读博士,只是为了学习如何处理不确定性和做研究,因为我喜欢站在前沿思考挑战性问题。但两条道路都有很多绝佳的机会。

I was not planning to do a PhD. I was always planning to go straight to industry. My parents are engineers and worked in industry, and I thought the way to have impact was to go to a company. My dad even told me he wouldn't hire someone with a PhD, so I thought maybe I shouldn't get one if I wouldn't be able to get a job. But he's in a different field, civil engineering. At the same time, I think a PhD is an incredible opportunity, and I love my PhD. It depends a lot on the adviser and what you'd be doing. But the PhD is an incredible opportunity to learn how to handle uncertainty and how to pick good problems to work on. In research, no one gives you the problem; you have to pick it, and you don't know if it's achievable to make progress in six months, two years, or ten years. You learn to deal with that uncertainty. That's really useful. It's also an opportunity to do amazing research and have a lot of freedom to work on what you find most exciting. So today, it's still an amazing opportunity to learn about uncertainty, which is useful in startup environments and at the frontier of AI because no one knows the best route to make models more powerful. At the same time, there are incredible opportunities in industry. Developing everything I showed isn't just research; there's a whole software stack that needs to run reliably on the robot, hardware, machine learning infrastructure, data infrastructure, and so on. You don't need a PhD for a lot of that engineering work. On the research side, there are often opportunities to get involved, and a lot of research is engineering these days. So it depends; it's a very personal decision. Even today, retrospectively, I'd probably want to do a PhD just to learn how to handle uncertainty and do research because I love being at the frontier and thinking about challenging problems. But there are amazing opportunities in both paths.

机器人数据等价物 Robotics Data Equivalent

Host

好的。下一个问题是:大型语言模型从互联网学习,但机器人并没有互联网规模的物理经验数据集。机器人领域的对应物是什么?我们如何获得它?

Okay. The next question is: large language models learn from the internet, but robots don't really have an internet-scale dataset of physical experience. What's the robotics equivalent and how do we get it?

Chelsea

在机器人领域,嗯,也许先从语言模型说起,互联网上的数据是语言数据,并非全部高质量。但有些数据信息量很大且有用。它反映了很多你希望模型做的事情:预测文本、补全文本、回答问题等等。互联网上有大量被回答的问题和被补全的文本。一般来说,在机器学习中,你希望训练与测试匹配。所以训练数据应该反映你之后要求模型做的事情。机器人领域的对应物是机器人在真实世界环境中运行的数据。我们在 Physical Intelligence 的方法是收集数据,即机器人执行各种任务的体验。你可以通过远程操作收集初始数据,让机器人做有用的事情。但从长远来看,它也会包含大量机器人部署尝试的自主经验。就像我们在语言模型中看到的,很多时间花在通过运行模型并让其思考来生成合成数据上。

In robotics, well, maybe in language models to start off, the data on the web is language data, and not all of it is high quality. But some of it is really informative and useful. It reflects a lot of what you want a model to do: predict text, complete text, answer questions, and so forth. There's a lot of questions being answered and text being completed on the internet. In general, with machine learning, you want training to match testing. So you want the data you train on to reflect what you'll ask the model to do later. The equivalent in robotics is data of robots operating in real-world circumstances. The way we approach it at Physical Intelligence is to collect data, robot experience, of robots doing all sorts of tasks. You can collect this with teleoperation to get initial data of robots doing useful things. But in the long run, it will also contain a lot of autonomous experience of robots deployed attempting things. Just like we see in language models, where a lot of time is spent generating synthetic data by running the model and having it think through things.

机器人数据 Data for Robotics

Chelsea

我认为未来机器人领域的大量数据,将是机器人在众多真实世界场景中尝试完成大量任务。所以,我觉得这就是未来的样子。我还认为,还有其他对模型训练非常有用的信息来源,比如人类做事的视频,像 YouTube、网络数据和带字幕的图像,这些能告诉你,比如这是一个厨房,冰箱在水槽右边,等等。所有这些数据,我认为对于开发一种前沿多模态模型非常有用,这种模型可以控制机器人做事,推理如何完成长任务,并控制机器人执行这些任务。我认为机器人的自身经验无可替代。你不能仅仅通过看人类做事就学会,比如我看罗杰·费德勒打网球,并不意味着我就能打得和他一样好,很遗憾。同样,机器人也不能看了人做事就直接学会自己怎么做。它们确实需要在自己平台上的亲身经验才能有效学习。我认为我们将需要大规模数据集。但这并不意味着人类视频没有用。看费德勒打网球是有用的。但机器人平台上的实际经验将是开发机器人领域类似数据集的关键组成部分。

I think a lot of the data in the future in robotics is going to be the robot attempting to do lots of tasks in lots of real world circumstances. And so, yeah, I think that that's kind of what it looks like. I also think that there are other possible sources of information that's really useful for model training, like videos of people doing things, like YouTube, web data and captioned images, that tell you like this is a kitchen that has a fridge on the right of the sink and so forth. And all of that data, I think, can be really useful for developing a kind of frontier multimodal model that can control robots to do things, reason through how to do a long task, and also control the robot to do those tasks. I think that there's no substitute for the robot experience itself. You can't just, if you watch a human do something, like if I watch Roger Federer play tennis, doesn't mean I can play tennis as well as him, unfortunately. And likewise, robots can't just watch a person doing something and then figure out how to do it themselves directly. They really need their experience on their own platform to learn effectively. And I think that we will need large data sets. That doesn't mean the human video isn't useful. It's useful to watch Roger Federer play tennis. But the actual experience on robot platforms will be a critical component of developing an analogous data set for robotics.

开源与民主化 Open Source and Democratization

Host

下一个问题是,通用机器人模型是否有可能像大型语言模型那样通过开源实现民主化,还是具身数据和硬件的成本会让最好的模型集中在少数资源充足的实验室?

The next question is, is it possible that general purpose robotics models get democratized via open source the way that large language models did, or will the cost of embodied data and hardware keep the best models concentrated in a few well-resourced labs?

Chelsea

所以,我认为这是个好问题。我确实认为具身数据和硬件的成本可能会让情况大不相同,因为我认为即使是为了蒸馏模型,也很难像在互联网上那样轻易获取数据。我也看到过相当大的数据集被开源,相当强大的模型也被开源。我认为很难准确预测会发生什么。所以,是的,我不知道。我想说的一点是,对于语言模型,即使不谈蒸馏,即使是为了获得真正达到最先进水平的模型,即使是那些专注于闭源模型的公司,也在做很多开源工作。比如 Gemma,还有 GPT 开源等等。我认为这些公司喜欢支持开源,因为这实际上有助于围绕他们正在构建的东西建立生态系统。所以我猜想,我对此持乐观态度,无论如何都会有一个强大的开源社区,但我不确定它是否会完全像语言模型那样发展。

So, I think this is a good question. I do think the cost of embodied data and hardware could very much make this look different, because I think that it's harder to get data even to distill a model, for example, just readily on the internet. I also think that we've seen pretty large data sets get open sourced as well, and pretty powerful models get open sourced. I think it's really hard to say exactly what will happen. And so, yeah, I don't know. The one thing that I will say is that with language models, even aside from distillation, like getting models that really perform at the state-of-the-art, even then, companies that are focusing a lot on closed source models are also doing a lot of open sourcing. So there are like Gemma, for example, and the GPT open source and so forth. I think these companies like to support open source because it actually helps build the ecosystem around the things that they're building. So I imagine there being, I guess I'm optimistic that there will be a strong open source community regardless, but I don't know if it will exactly play out the way that language models played out.

输出级别:关节位置与力矩 Output Level: Joint Positions vs Torques

Host

好的,下一个问题是,模型是直接输出原始电机指令,还是输出目标手部位置,让控制器求解关节角度,以及为什么这是合适的学习层级?

Okay, the next question is, does the model output raw motor commands directly, or does it output a target hand position and let a controller solve for the joint angles, and what makes that the right level to learn at?

Chelsea

所以,我展示的所有模型都是输出目标关节位置。比如,这个关节的角度是多少?那个关节的角度是多少?然后有一个控制器,比如 PD 控制器,试图达到这些关节的目标位置。模型实际上也经过训练来预测目标夹持器位置,比如我的夹持器在 3D 空间中应该在哪里?你也可以利用这一点,反推出关节位置。你还可以直接输出电机扭矩,或者电压或力。不同选项各有优缺点。我们发现控制关节和控制夹持器的 3D 空间都效果不错。所以各有优缺点。我认为直接输出到电压的一个好处是,你可以得到更硬或更软的输出。而使用固定控制器,你就无法让模型控制那个方面。所以有不同的利弊。我认为我们正在做的工作似乎有效。它似乎不是瓶颈。我通常喜欢关注那些似乎是瓶颈的事情,而不是那些不是瓶颈的事情。

So, all the models that I showed were outputting target joint positions. So like, what is the angle of this joint? What is the angle of that joint, and so forth, that you want to hit? And then there's a controller, like a PD controller, that is trying to then hit that target position for those joints. The model actually is also trained to predict target gripper positions, like where in 3D space should my gripper be? And you could also use that as well and back out the joint positions. Another thing you could do is you could go directly to motor torques, or to voltages or efforts. There are pros and cons of different options. We have found controlling joints and controlling in the 3D space of the gripper to both work well. So there are pros and cons. I think that one thing that would be nice about going directly to the voltages is that you could also get a more stiff output or a less stiff output. Whereas with the controller, if you have a fixed controller, then you're not letting your model control that aspect. So yeah, there's different pros and cons. I think that what we're working on seems to work. It doesn't seem to be a bottleneck. And I often like to focus on the things that seem to be bottlenecks versus things that don't seem to be bottlenecks.

机器人中的想象力 Imagination in Robotics

Host

好的。下一个问题是,机器人是否需要某种想象力,即预想接下来应该发生什么的能力,才能变得真正有用?

Okay. Next question is, do robots need something like imagination, the ability to picture what should happen next, before they can become truly useful?

Chelsea

所以,我展示的 Pi0.7 模型就有类似的能力,它可以想象未来的图像应该是什么样子,然后努力实现它。我们发现这带来了改进,在叠衬衫的例子中,我们看到使用这种想象力相比不使用有数量上的提升。同时,我认为没有这种能力,模型的表现也出奇地好。我们实际上考虑过写一篇完整的技术报告,专门介绍该模型的这种能力。但没有这种能力的模型表现太好了,以至于我们觉得需要让这种能力在故事中扮演更重要的角色,因为它似乎在取得真正强大的结果方面确实发挥了作用。所以这似乎是一个设计选择。我认为很难说它是否会成为关键组成部分。这类模型的好处是,如果你开发了一个好的数据集,你可以进行实验,并利用你拥有的数据集有效地继续测试。我还认为,与预测未来动作相比,预测未来似乎是一个非常相关的目标。所以,我想这应该有助于从你所有可用的数据中学习。所以,是的,很难说它是否一定会成为关键组成部分。从经验上看,到目前为止它似乎有帮助,尽管可能没有你预期的那么多,而且即使没有这种想象力,机器人也能做非常了不起的事情。

So, the Pi0.7 model that I showed has something like this, where it can kind of imagine what a future image should look like and then try to accomplish that. We found that that leads to improvement, and we saw in the shirt folding example a quantitative bump from using that sort of imagination compared to not using it. At the same time, I think that the model actually performed surprisingly well without that as well. And we were actually thinking about writing an entire technical report just about that capability in that model. But the model without that was so good that we felt like we needed to actually have that play a bigger part of the story, because that seemed like it was really delivering in terms of actually getting really strong results. So it seems like one design choice. I think it's hard to say if it's going to be a critical component or not. The good news with these kinds of models is that if you develop a good data set, you can run experiments and continue to test things with the data set that you have quite effectively. I also think that being able to predict the future seems like a very relevant objective compared to predicting future actions. So that should, I would imagine, help in terms of learning from all the data that you have available to you. So, yeah, hard to say if it'll necessarily be a critical component or not. It seems like empirically so far it seems to help, although perhaps not as much as you might expect, and even without that imagination, the robot can do pretty incredible things.

提升速度 Improving Speed

Host

好的。接下来是,目前机器人似乎在以非常缓慢的方式完成惊人的任务。要提高速度需要什么?

Okay. Next is, right now it seems that robots are doing amazing tasks, but in a very slow manner. What is needed to improve the speed?

Chelsea

我对提高速度感到非常兴奋。我们确实从强化学习中看到了速度的提升。我们还有另一个发布,叫做 RL token,我们展示了更快的速度,实际上比人类远程操作还要快。

I'm really excited about improving the speed. We did see speed improvements from reinforcement learning. We also have another release called the RL token, where we showed actually even faster speed, and actually faster speed than human teleop.

机器人数据收集瓶颈 Bottlenecks in Robot Data Collection

Chelsea

我认为瓶颈之一是,当你通过远程操作让机器人做事时——这是教机器人做事最简单的方式——人远程操作机器人的速度很慢。我们有几个正在进行的项目,在获取快速策略方面有很有前景的结果。所以这方面还会有更多进展。我认为要么你需要想办法让数据更快,要么你需要想办法比数据更快。我们已经看到了能够比数据稍快一点的证据。至于下一步,要么更进一步,要么让数据更快。

I think one of the bottlenecks is that when you teleoperate robots to do things, which is the easiest way to teach a robot to do something, people are kind of slow at teleoperating the robot. We have a couple projects in the pipeline that I think have really promising results in terms of getting fast policies. So more to come there. I think it's either you need to figure out how to make the data faster or you need to figure out how to be faster than the data. We've seen evidence of being able to be a little bit faster than the data. In terms of next steps, it's either to go even further than that or make the data faster.

令人惊讶的机器人任务与涌现能力 Surprising Robot Task and Emergent Capabilities

Host

酷。你最近看到机器人完成的最令人惊讶的任务是什么?你希望它接下来做什么?

Cool. What's the most surprising task you've seen a robot complete recently? What do you want to see it do next?

Chelsea

最令人惊讶的事情其实不是某个任务,而是我们在做 PIO7 时,我亲自训练了一个策略,用于组装这个纸风车的一些初步测试。在我训练它组装纸风车时,真正让我惊讶的是,在我们所有的数据中,我们都仔细控制了组装纸风车的策略,基本上就是拿起预先裁好的纸,拿一个小别针,然后把别针插入纸上的孔里。在所有数据中,我们都是用右手拿起别针,左手拿起纸,然后插入。机器人开始这样做,然后它实际上犯了一个错误,纸到了右边,别针到了左边。机器人做的是:它用左手抓取器拿起纸和别针,然后用左手抓取器把别针插入纸中,用右手。它从未见过用左手抓取器插入别针的数据。这表明,即使在后训练数据中完全没有,甚至在预训练中也没有。机器人本质上学会了这种左右手之间的等变性,所以它实际上可以将行为从一只手转移到另一只手,尽管数据中从未出现过。那真是一个很酷的时刻。我不知道当我和一些人分享时,他们是否和我一样兴奋。但这有点展示了这些模型中我从未见过的涌现能力。

The most surprising thing was not really a task but when we were working on PIO7, I personally trained one of the policies for some of the initial tests for assembling this pinwheel. When I was training it to construct the pinwheel, one thing that really surprised me was that in all the data we carefully controlled the strategy for how to assemble the pinwheel, where you basically take the pre-cut piece of paper and take a little pin and insert the pin into a hole in the paper. In all the data, we picked up the pin with the right hand and picked up the paper with the left hand and inserted it. The robot started doing that, and then it actually made a mistake and the paper ended up on the right side and the pin ended up on the left side. What the robot did is it picked up the paper and picked up the pin with its left gripper, and it put the pin with its left gripper and inserted it into the paper with its right. It had never seen data of inserting the pin with its left gripper. It showed that even that wasn't in the post-training data at all, and it wasn't even in pre-training either. The robot essentially had learned this sort of equivariance between its left hand and its right hand, so it could actually transfer behaviors from one hand to another, despite the fact that that was never in the data. That was a really cool moment. I don't know if other people were as excited about it as I was when I shared it with some people. But it kind of shows this emergent capability in these models that I hadn't seen before.

Host

那你希望它接下来做什么?

And then what do you want to see it do next?

Chelsea

我不知道。我喜欢看机器人做任何事情。我认为在机器人能够长时间执行任务的可靠性方面,还有很长的路要走。我不太考虑单个任务,而是更多考虑能力,以及如何从这些模型中获得下一个能力。机器人做任何事情总是让我兴奋,即使是从未做过的事情。我们最近在做的一件事是让机器人用刀切蔬菜。我认为一旦你能安全地使用刀,你可以做很多事情,这是我们最近完成的一件事。

I don't know. I love seeing robots do anything. I think there's still a long way to push in terms of reliability for robots being able to do tasks for really long periods of time. I don't necessarily think much about individual tasks, but more about capabilities and how to get the next capability from these models. A robot doing anything always gets me excited, even if it's something that hasn't been done before. One thing that we've been doing recently is having robots use knives to slice vegetables. I think there's a lot that you can do there once you actually can use knives safely, which is one thing that we've done recently.

从软件工程转入机器人领域 Breaking into Robotics from Software Engineering

Host

好的。最后一个问题是,有软件工程背景的人如何进入机器人领域?

Okay. And then the last question is how can someone break into robotics from a software engineering background?

Chelsea

很好。我认为首先,机器人领域有很多软件工程,所以你可以尝试以软件工程师的身份加入机器人公司。另一件我想提到的事情,我确实见过有人走这条路,就是现在在 Physical Intelligence 工作的一个人,她叫 Jenny。她做过一段时间的算法交易。然后她在 Harvey 工作,做法律相关的事情。她对机器人非常兴奋,所以她买了一个便宜的机器人,基本上在她的卧室里摆弄它,尝试微调一个开源模型,试图让它做点什么。然后她分享了她所做的事情,给我发了一封冷邮件,说:‘嘿,我能在你的实验室工作吗?’她的背景看起来很有希望,而且她真的去尝试了,做了,而且她对此非常兴奋。现在她在 Physical Intelligence 工作。我认为只要让自己沉浸其中,尝试新事物,从经验中学习,然后利用这些经验与人分享,把它写在简历上等等,是一个很好的方式。幸运的是,有很多开源的东西可以让你开始做这些事情。

Great. I think first, there's a lot of software engineering in robotics, so you could try joining a robotics company as a software engineer. Another thing that I would mention, and I've actually seen someone take this path, is someone who now works at Physical Intelligence, her name is Jenny. She worked in algorithmic trading for a while. Then she worked at Harvey and was doing legal stuff. She was really excited about robots, so she bought a cheap robot and basically in her bedroom played around with it, tried fine-tuning an open source model and trying to get it to do something. Then she shared what she had done and sent me a cold email and was like, 'Hey, can I work in your lab?' It seemed like her profile was promising and that she actually got out there and tried it and done it and that she was really excited about that. Now she works at Physical Intelligence. I think just getting your feet wet, trying stuff out, and learning from that experience, and then using that experience to share with people, have it on your resume, and so forth, is a great way to do stuff. Fortunately, there's a lot of open source stuff out there that can allow you to get started on those kinds of things.

结束 Closing

Host

太好了。这是最后一个问题。感谢大家的收听。

Great. That was the last question. Thanks everyone for listening.

互动版:逐字朗读 + 针对本期提问 →