NVIDIA Jim Fan 专访:从 OpenAI 实习生到具身智能先驱

Inside NVIDIA's Jim Fan: From OpenAI Intern to Embodied AI Pioneer

范麟熙 Jim Fan · Radical Ventures · 2026-02-11 · 约 45 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Jim Fan 分享他从 OpenAI 首位实习生到 NVIDIA 具身智能领军人物的历程,揭示其开创性研究背后的原则。

Jim Fan shares his journey from OpenAI's first intern to leading embodied AI at NVIDIA, revealing the principles behind his groundbreaking research.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 14)

全文 · Full transcript(中英对照)

引言 Introduction

Host

欢迎收听《Radical Talks》特别版。我是 Molly Welch,Radical Ventures 的合伙人。今天带大家走进我们的年度 Radical AI 创始人大师课。这是一个独家项目,AI 研究者和技术创业者可以直接向行业先驱学习,从他们的经历中汲取灵感和实用见解。在这个系列中,我们将分享 2025 年课程中的对话,让你也能接触到这些帮助 AI 创始人打造创新公司、塑造未来的讨论。今天,我们带来的是我与 NVIDIA AI 总监、杰出科学家 Jim Fan 的对话。这次对话非常有趣。让我们开始吧。

Welcome to a special edition of Radical Talks. I'm Molly Welch and I'm a partner here at Radical Ventures. Today we're taking you inside our annual Radical AI founders master class. It's an exclusive program where AI researchers and technical entrepreneurs learn directly from industry pioneers, drawing inspiration and practical insights from their journeys. In this miniseries, we'll share these conversations from our 2025 sessions, giving you access to the discussions helping AI founders build innovative companies that shape the future. Today, we are featuring a conversation between myself and NVIDIA's director of AI and distinguished scientist, Jim Fan. This one was a lot of fun. So, let's dive in.

Host

我是 Molly Welch,Radical Ventures 的合伙人。我们是一家专注于 AI 的风险投资基金,非常高兴大家能来参加我们的年度大师课。今天有一场特别的环节。我们有幸邀请到 Jim Fan,他是 AI 智能体和机器人研究领域的领军人物。Jim 刚从 NVIDIA GTC 大会回来。我想他现在还在华盛顿特区,所以今天我们有很多有趣的话题可以聊。Jim 是 NVIDIA 的董事兼杰出研究科学家,在那里他共同领导了用于人形机器人的 Isaac Groot 平台的开发,并负责通用具身智能体研究团队(GEAR)。Jim 在斯坦福视觉实验室获得博士学位,导师是李飞飞。他的开创性工作包括 Voyager——第一个通过终身学习掌握《我的世界》的 AI 智能体;Mind Dojo——获得了 NeurIPS 2022 杰出论文奖;以及 Eureka——一个教机器人手完成转笔等复杂灵巧任务的项目。Jim 还有一个独特的经历,他是 OpenAI 在 2016 年的第一位实习生,当时共同开发了 World of Bits,我们稍后会详细聊这个。Jim,我们非常高兴你的到来。欢迎来到 Radical AI 创始人大师课。

My name is Molly Welch and I'm a partner here at Radical Ventures. We're an AI focused venture fund and we're thrilled to have you all here for our annual master class. We have a special session today. We have the privilege of speaking with Jim Fan, a leading voice in AI agents and robotics research. Jim is fresh off of Nvidia GTC. I think he's still in DC, so we're going to have a lot of fun things to speak about today. Jim is a director and a distinguished research scientist at NVIDIA where he co-leads development of the Isaac Groot platform for humanoid robotics and heads the gear or generalist embodied agent research team. Jim earned his PhD from the Stanford vision lab where he was advised by Fei-Fei Li. His groundbreaking work includes Voyager, the first AI agent to master Minecraft through lifelong learning, Mind Dojo, which won the NeurIPS 2022 outstanding paper award, and Eureka, a project teaching robot hands complex, dextrous tasks like pen spinning. Jim has also the unique distinction of being OpenAI's very first intern back in 2016 where he co-developed World of Bits, and we're going to hear a little bit more about that later. Jim, we're thrilled to have you. Welcome to the Radical AI Founders Master Class.

Jim

非常感谢你的邀请,也谢谢 Molly 这么友好的介绍。

Thank you so much for having me and thanks for the very kind intro, Molly.

Host

我们真的很高兴你能来。Jim,我想从你这段不可思议的经历开始聊起,从 OpenAI、斯坦福视觉实验室,到现在在 NVIDIA 领导具身智能体研究。请带我们回顾一下你的旅程,以及你在每次转变时的决策过程。特别是,我们想了解是什么吸引你从学术界走向这些行业角色。

We're really pleased to have you. Jim, I want to start with this incredible journey that you've been on from OpenAI, Stanford's Vision Lab, now leading embodied agents at NVIDIA. Walk us through your journey, your decision-making processes at each transition. And in particular, we want to hear about what drew you from academia to some of these roles in industry.

Jim

哇,这是个很大的问题。让我拆开来讲。我觉得,回想起来,这真的是机缘巧合,加上跟随一些好的直觉。我把我过去大约 10 年做的研究叫做“直觉研究”。我跟随的直觉基本上有两件事:一是选择热门问题,二是寻求简单的解决方案。我觉得我抓住的很多机会、做的很多项目,都遵循这两个原则。我最早对深度学习产生兴趣是在 2012 年,那时 AlexNet 刚出来。我被它的优雅深深吸引。在 2012 年之前,计算机视觉是一个非常复杂的流程:你要手工设计特征,对像素做各种几何分析,然后用 SVM 和一堆统计方法来完成图像分类。而 AlexNet 所做的只是把像素直接映射到类别。比如一千个不同的物体,你从一千个中选一个,就是这样。非常简单,非常优雅。所以我的第一个深度学习研究项目其实是在哥伦比亚大学计算机视觉实验室做的一个天体物理学项目。目标是从天体物理图像中估计引力透镜参数,那大约是 2013 年夏天到 2014 年。那时还是相当传统的流程,比如用 SIFT 特征和 SVM 做分类。我决定实现类似卷积神经网络的东西,但当时是用 MATLAB,然后我发现它真的把所有传统方法都碾压了。这是一个非常简单的方法:只要有一个小的训练集,你做梯度下降,系统就自我编程了。那是我第一次接触深度学习。然后在 2015 年,我加入了 Andrew 的实验室,参与一个叫 Deep Speech 的项目。顾名思义,它是做语音识别的。同样,当时有很多传统系统,先做音频分割,然后音素分类、波形特征提取,最后是一大堆概率评分模型。但在当时,对于 Deep Speech,我们相信端到端。基本上你输入原始音频波形,然后输出文本,一个文本序列,就是这样。那时我们甚至还没有 Transformer。所以我们用的是非常老的序列模型,叫 LSTM,而且我们在 2015 年就讨论了数据管道,讨论了扩展数据和模型。所以这是一个非常简单、优雅的方法,结果它在语音识别上表现得非常好,远远好于传统系统。然后在 2016 年,那是我博士开始前的最后一个暑期实习,我听说了一家新公司,一家有很多杰出人才的新创公司,那家公司就是 OpenAI。我联系了他们,然后他们说他们没有实习项目,但也许是时候创建一个了。所以我很荣幸地在那个夏天加入了 OpenAI,我参与的项目叫 World of Bits。我的导师正是伟大的 Andrej Karpathy。所以这是他的想法,他的愿景,我很高兴能成为这个愿景的一部分。World of Bits 的想法是,我们想教一个 AI 智能体通过观察浏览器渲染的像素来学习使用网页浏览器,然后直接输出键盘和鼠标控制。这实际上是 OpenAI 当时一个更大的计划的一部分,叫做 Universe。它更通用:基本上你在电脑上做的任何事情,原则上都可以通过这种从屏幕像素到键盘鼠标动作的映射接口来实现。所以你可以玩游戏,可以用 Photoshop、Microsoft Word 这样的软件,方式完全一样。所以那个项目实际上是关于为智能体构建一个通用接口。一旦你有了这个通用接口,你就可以实现端到端的智能体,从像素一直映射到动作。我觉得在当时这是革命性的。我从未见过类似的东西。当时的 AI 连玩非常简单的游戏都很吃力,更不用说解决通用的计算机使用问题了,而且我们用的是强化学习。那时还没有基础模型,一切都是从头训练的。所以在 2016 年,这远远超前于它的时代。

Wow, that's a huge question. So, let me break this out. I think, in retrospect, it's really serendipity and then following some good vibes. So, I'm calling it vibe research that I did over the last like 10 years. So, I would say the vibe I was following is basically two things. One is picking hot problems and the second is to seek simple solutions. And I feel like a lot of the opportunities that I hopped on, a lot of the projects, are all following these two principles. I got interested in deep learning first in 2012, and that's when AlexNet first came out. And I was really attracted by its elegance. Before 2012, computer vision is a very complex pipeline: you have hand-engineered features, all the geometric analysis on pixels, and then SVM and a bunch of statistical methods to reach a classification of the image. And all AlexNet did was mapping pixels directly to a category. So like a thousand different objects and then you pick one out of thousand, that's it. Very simple, very elegant. So my first deep learning research project was actually an astrophysics project at a Columbia computer vision lab. The goal is to estimate gravitational lens parameters from astrophysics images, and that was around summer of 2013 to 2014. At that time, it was still very kind of traditional pipelines, like you have SIFT features and SVM to do classification. I decided to implement something like a convolutional neural net, but at that time in MATLAB, and then I found that it actually blew everything out of the water. It's a very simple method: if you have a small training set, you do gradient descent and the system programs itself. So that was kind of my first contact with deep learning. Then in 2015, I joined Andrew's lab to work on a project called Deep Speech. As the name suggests, it's for speech recognition. Again, at that time there are a lot of traditional systems that first do audio segmentation, then phoneme classification, wave feature extraction, and finally a ton of probabilistic scoring models. But at that time for Deep Speech, we believed in end-to-end. Basically you take a raw audio wave in and then output text, a sequence of text, and that's it. At that time we didn't even have transformers. So it was like very old sequence models called LSTM, and we talked about having a data pipeline. We talked about scaling up data and the model back in 2015. So it was a very simple, elegant approach, and it turns out that it solves speech recognition really, really well, much better than traditional systems. Then 2016, that was the last summer internship before I started my PhD, and I heard about a new company, a new startup with a bunch of brilliant people, and that startup turns out to be OpenAI. I reached out, and then they said they don't have an intern program, but maybe it's time to create one. So I had the great honor to join OpenAI that summer, and the project I worked on was called World of Bits. I was mentored by none other than the great Andrej Karpathy. So it was his idea, his vision, and I was very gladly part of that vision. The idea of World of Bits was that we wanted to teach an AI agent to learn to use the web browser by observing the pixels rendered from the browser and then outputting keyboard and mouse controls directly. It's actually part of a much larger initiative at OpenAI at that time called the Universe. And it's more general: basically anything that you do on a computer can in principle be this interface of mapping pixels from the screen to keyboard and mouse actions. So you can play games, you can use software like Photoshop, Microsoft Word in exactly the same way. So that project is really about building a general-purpose interface for the agents. And once you have that general-purpose interface, then you can enable end-to-end agents that map from pixels all the way to actions. I think at that time it was revolutionary. I have never seen anything like that before. The AIs at that time were struggling to play very simple games, let alone solving general-purpose computer use, and we were using reinforcement learning. There were no foundation models, everything was kind of trained from scratch. So in 2016, that was way, way ahead of its time.

引言与早期生涯 Introduction and Early Career

Jim

但你知道,2025 年,9 年之后,我们有点看到历史重演,或者说不是完全重演,而是押韵。我们开始用现代大语言模型和强化学习来解决比特世界的问题。但这次不是从零开始,对吧?是在大语言模型之上进行推理,然后我们有了像计算机使用智能体、MCP 这样的东西,可以使用工具。我们终于开始解决这个问题了。所以即使我们在 2016 年没能完全解决,我认为那个夏天的实习真正塑造了我的研究品味,那就是数据管道可以非常复杂,但模型接口需要非常简单。因为只有简单优雅的方法才能利用各种数据源,只有简单的方法才能最好地随算力扩展。这真的是一个如何看待这些难题的视角。

But you know, this year 2025, 9 years later, we kind of see history repeats itself, or really it doesn't repeat exactly but it rhymes. We are starting to solve the world of bits using modern LLMs and also reinforcement learning. But this time it's not from scratch, right? It's reasoning on top of LLMs, and then we have things like computer-using agents, MCPs, that can use tools. We're finally starting to solve that. So even though we were not able to completely solve it in 2016, I think that summer's internship really shaped my research taste, which is that the data pipeline can be very complex, but the model interface needs to be very simple. Because only simple and elegant methods can leverage all kinds of data sources, and only simple methods can scale the best with compute. It's really a perspective on how to approach these hard problems.

Jim

所以那是那个夏天,然后我在斯坦福大学跟随李飞飞教授开始了我的博士研究。众所周知,飞飞不需要介绍。她是世界知名的计算机视觉科学家,以 ImageNet 闻名。当时,计算机视觉是关于图像分割基准、分类、视频标注等等。我们称之为静态计算机视觉。你有一个数据集,有静态图像和视频,然后映射到某种标签。现在,飞飞的愿景是静态计算机视觉正在碰壁,下一个重大革命将是具身视觉。意思是计算机视觉应该通过智能体来解决。所以你可以有模拟和现实世界中的视觉智能体,而现实世界中的视觉智能体基本上就是机器人技术。

So that was the summer, and then I started my PhD with Professor Fei-Fei Li at Stanford. As everyone knows, Fei-Fei doesn't need an introduction. She's the world-renowned scientist on computer vision. She was known for ImageNet. At that time, computer vision was about image segmentation benchmarks, classification, video annotation, and so on. We call those static computer vision. You have a dataset, you have static images and videos, and you map that to some sort of labels. Now, Fei-Fei's vision was that static computer vision is hitting a wall, and the next big revolution would be embodied vision. Meaning that computer vision should be solved with agency. So you can have visual agents in simulation and in the real world, and the visual agents in the real world is basically robotics.

Jim

所以从 2016 年到 2021 年,那五年里,我在实验室。我基本上见证了飞飞实验室从静态视觉到具身视觉的转变,然后有了一个新的北极星,即解决这类机器人技术和视觉理解,但要在主动感知循环中。所以那非常鼓舞人心,而且那是对我研究品味的第二次重大重塑。然后我毕业去了 NVIDIA。NVIDIA 仍然是我第一份也是唯一一份全职工作,我一直在把我博士工作中的很多想法带到构建通用机器人和机器人基础模型中。抱歉,这是一个非常非常长的介绍,但我确实希望回顾过去,看到所有这些线索是如何联系在一起的。这是氛围研究。

So from 2016 to 2021, over those five years, I was at the lab. I basically witnessed Fei-Fei Lab's transition from static to embodied vision, and then to have a new north star of solving these kinds of robotics and visual understanding, but in an active perceptive loop. So that was very inspiring, and again, that's a kind of second major reshaping of my research taste. Then I graduated and went to NVIDIA. NVIDIA is still my first and only full-time job, and I've been carrying on a lot of ideas from my PhD work into building general-purpose robotics and robot foundation models. Sorry, that's a very, very long intro, but I do hope to look back and see how all these threads tie together. It's vibe research.

Host

多么精彩的框架。基于氛围的研究,还有这么多精彩的线索,我很期待在我们今天的对话中,在你职业生涯的这几个激动人心的阶段中,逐一展开。也许先谈谈第一个阶段,你在 OpenAI 的实习,感觉有点难以置信,OpenAI 在最早期的这种研究,正如你所说,今天依然如此相关。而且非常符合时代精神,这些同样的雄心现在成为可能,只有在前沿模型和高算力强化学习的条件下才可能实现,正如你描述的那样。那么带我们回到 2016 年在 OpenAI 的第一次实习。那是在大语言模型真正成为 OpenAI 工作核心之前的 World of Bits。那段经历是怎样的?它如何塑造了你在 AI 研究和更广泛行业上的视角?

What a wonderful framing. Vibes-based research and so many great threads that I'm excited to pull on in our conversation today over the several exciting acts of your career so far. And maybe to touch on this first act, your internship at OpenAI, which feels kind of insane to conceptualize, OpenAI in the earliest days also this type of research, which as you said is so relevant today. And so much in the Zeitgeist, these same ambitions are now possible, uniquely possible with frontier models and high-compute RL as you described. So take us back to this first internship at OpenAI 2016. This is World of Bits before LLMs become really central to the work at OpenAI. What was that experience like and how did it shape your perspective on AI research and the broader industry?

Jim

当然。当时在 OpenAI,我参与的是 OpenAI Universe 项目。还有几个其他令人兴奋的项目。一个是 Dota 项目,叫做 OpenAI Five,基本上是训练智能体来解决像 Dota 这样的游戏。当时还有一个实际的机器人项目,使用高自由度的 Shadow 手来解决魔方。稍后我们会谈到一些像 Eureka 这样的研究,实际上重新实现了这些想法,然后训练灵巧的手做更多事情。所以我认为当时有很多构建模块,即使早在 2016 年,最重要的就是强化学习和序列模型。但直到 2017 年和 2018 年,Transformer 架构出现,它变成了一个数据海绵,你可以在数据和算力上真正扩展,你会看到预训练带来的魔法。所以在那之后,我们开始进入基础模型的预训练时代,然后强化学习回归,正如我提到的,历史不会重演,但会押韵。强化学习回归,并从这个强大的基础真正起飞。基础需要能够进行推理,需要能够进行上下文学习,所有这些能力,然后强化学习可以把它推向新的前沿。

Absolutely. At that time, in OpenAI, there was the OpenAI Universe that I worked on. There were also a few other exciting projects. One is the Dota project, which was called OpenAI Five, basically to train agents to solve games like Dota. There was also an actual robotics project at that time, to use high-degree-of-freedom shadow hands to solve Rubik's cube. Later we're going to talk about some research like Eureka that actually reinstantiate some of these ideas, and then train dexterous hands to do a lot more. So I think at that time there were a lot of building blocks even as early as 2016, the most significant ones being reinforcement learning and sequence models. But it wasn't until 2017 and 2018 where the Transformer architecture comes, and it becomes this data sponge where you can really scale on data and compute, and you will see magic happening from pre-training. So after that moment, we started to get into the pre-training age of foundation models, and then RL comes back, and as I mentioned, history doesn't repeat, it rhymes. RL comes back and really takes off from this strong foundation. The foundation needs to be able to do reasoning and needs to be able to do in-context learning, all of those capabilities, and then RL can push it to the next frontier.

Jim

所以那次在 OpenAI 的经历也激发了我 NVIDIA 最早的项目之一。那是你一开始提到的两个项目,MineDojo 和 Voyager,我们想让这些具身智能体解决 Minecraft。我认为那个项目也是我在工作中玩游戏的借口之一。这都是为了科学,对吧?我在收集数据,对吧?我没有偷懒。我在收集数据。

So that experience at OpenAI also inspired one of my first projects at NVIDIA. That was a pair of projects you mentioned at the beginning, MineDojo and Voyager, where we wanted to have these embodied agents to solve Minecraft. I think for that project, it was also one of my excuses to play games at work. It's all for science, right? I'm collecting data, right? I'm not slacking off. I'm collecting data.

Host

都是为了数据收集的精神。都是为了 AI。都是为了推动 AI 的前沿。

All in the spirit of data collection. It's all for AI. It's all for pushing the front of AI.

Jim

所以,我认为那个项目有趣的地方在于 Minecraft 是一个开放式的环境。它也是一个模拟环境,意味着你基本上可以在 Minecraft 中生成无限多样性的无限数据。我见过人们在这个沙盒游戏中做疯狂的事情,比如一块一块地建造霍格沃茨城堡,甚至用 Minecraft 中的基本构建块构建一个功能正常的 CPU 电路,这实际上使游戏变得图灵完备。所以这个游戏有无限的多样性,但人类玩得很好,而 AI 却远远落后。它甚至不能做一些基本的事情。所以当我们开始 Voyager 项目时,恰逢 GPT-4 的推出。那是第一个能非常好地编写代码的基础模型。所以我们如何将 GPT-4 应用到 Minecraft 中,我们给它一个 API,这样 GPT-4 可以阅读 Minecraft API 的指令,然后编写代码,然后我们在 Minecraft 模拟器中执行这些代码,驱动智能体去挖矿、探索、旅行、挖掘、制作新物品。所以基本上我们把游戏变成了一个文本环境,它通过代码进行交流。然后我们发现 GPT-4 不仅擅长上下文学习,还非常擅长一种叫做自我反思的能力。所以它可以迭代,对吧?没有人能在开始时写出完美的程序,即使是经验丰富的程序员也不行。你必须迭代。这就是我们在 GPT-4 身上观察到的。

So, I think what was fun about that project was Minecraft is this open-ended environment. It's also a simulation, meaning that you can basically generate infinite data of almost infinite diversity in Minecraft. I've seen people doing crazy things inside this sandbox game, like building the Hogwarts castle brick by brick, or even building a functioning CPU circuit using those fundamental building blocks in Minecraft, which actually makes the game Turing complete. So this game has infinite diversity, yet humans play it very well, but for AI it was lagging way behind. It couldn't even do some basic things. So at the time when we started the Voyager project, that coincided with the introduction of GPT-4. That was the first foundation model that can do coding extremely well. So how we apply GPT-4 to Minecraft is that we give it an API so that GPT-4 can read the instructions of the Minecraft API and then write code, and then we execute that code in the Minecraft simulator to drive the agent to mine blocks, to explore, to travel, to dig, to craft new items. So basically we turn the game into a text environment and it communicates by code. And then we found that GPT-4 is not only good at in-context learning, it's also very good at a capability called self-reflection. So it can iterate, right? No one can write a perfect program at the beginning, not even experienced programmers. You got to iterate. And that's what we observed with GPT-4.

Voyager的上下文终身学习 Voyager's In-Context Lifelong Learning

Jim

它会写代码、测试代码,代码出错时,它看到 bug,就调试,再试一次,一旦写出好的代码,就会提交到记忆里。所以我们称之为“即时代码仓库”。Voyager 会把写好的部分代码存到这个仓库里,以后遇到类似情况就能检索出来。这样一来,我们其实没法微调 Voyager 的参数,但它仍然会随着时间学习,因为它会提交代码、检索代码、不断试错,而且在这个过程中积累上下文。所以这是一种上下文内的终身学习。我觉得这个想法特别吸引人,因为当时人们真的在尝试探索这些编码前沿模型的能力边界,而 Voyager 提供了一种在它之上构建非常复杂的智能体系统的方法。两年后的今天,智能体已经强大得多。自我反思大家都不提了,已经成了标配。像 Claude 和 Cursor 这些,它们调试能力超强,能进行 30 分钟的会话来实现一个拉取请求,然后自动运行,通常完成得相当不错。它还能调用各种工具,比如网页浏览、连接 MCP。所以我觉得这个领域已经远远超越了 Minecraft。这真的很令人兴奋。当时还欠缺的一点是,我们没有解决运动控制问题,也没有解决感知问题。我们本质上是通过把 Minecraft 变成一套代码 API 和代码执行环境来简化问题。但后来我意识到,最困难的部分可能不一定是推理方面。我觉得这个领域已经在很好地解决推理了。具身智能体最困难的部分其实是底层控制,也就是与底层物理交互的部分。在 Minecraft 里是游戏物理,但在现实世界里是我们触摸的实际物理现实,这让机器人问题变得困难得多。

It writes code, tests it, and the code breaks, it sees the bug, it debugs, it tries again, and once it writes a good piece of code, it would commit it to a memory. So we call this kind of an on-the-fly code repo. So Voyager will save some of the code it wrote to this repo so later it can retrieve it if it sees a similar situation. And in this way, we couldn't actually fine-tune the parameters for Voyager, but it still learns over time because it's committing code, retrieving code, and doing trial and error. And it accumulates the context as it goes. So it's a kind of in-context lifelong learning. And I find that idea really fascinating because at that time, people were really trying to experiment and see what are the capabilities of these coding frontier models, and Voyager provided one way to build a very complex agentic system on top. And then two years later, these days, the agents are a lot more powerful. Self-reflection people don't even talk about it, it's kind of a given. All the Claude and Cursor of the world, they can do debugging super well, they can go on like a 30-minute session to implement a pull request, and then it just goes on autopilot and it typically does the job pretty well. And it's also able to call all kinds of tools like web browsing, connecting to MCP. So I think the field has gone way beyond Minecraft. So that was really exciting to see. Now, what was lacking a bit was that we didn't solve the motor control problem and we didn't solve the perception problem. We essentially simplified the problem by turning Minecraft into a set of code API and code execution environment. But later I have come to understand that really the most difficult part may not necessarily be the reasoning side. I think the field is already solving reasoning quite well. The most difficult part of embodied agent is actually the low-level control, the part where you interact with the low-level physics. In Minecraft it's the game physics, but in the real world it's the actual physical reality that we're touching, and that makes the robotics problem a lot harder.

Host

非常酷,而且极其巧妙。即使技术已经进步,这个方法依然极其巧妙。谢谢你为我们讲解。我很好奇,MineDojo 和 Voyager 代表了这种向自主或更自主通过探索学习的智能体的演进。你认为创造真正自我改进的智能体,剩下的关键瓶颈是什么?

Very cool, and incredibly clever. The approach remains incredibly clever even as the technology has progressed. So thank you for talking us through it. I'm curious, you know, MineDojo and Voyager represent this progression towards agents that learn autonomously or more autonomously through exploration. What do you think are the key remaining bottlenecks for creating truly self-improving agents?

Jim

我觉得有两件事。心理学里有一本著名的书叫《思考,快与慢》。基本上是说,所有像人类这样的智能生物都有两个系统。系统二是缓慢、深思熟虑、负责推理的;系统一是快速、反应式、直觉式的。所以智能体式 AI 采用双系统方法。我觉得现在系统二进展得相当不错。它不完美,但有很多方法可以改进系统二,比如更多数据、更多可以接入前沿模型的环境,让它们能做强化学习。如今前沿模型真正受限于环境的多样性,而不仅仅是预训练数据。我们正处于这种强化学习后训练的新时代。所以这是系统二。而系统一仍然具有挑战性,因为计算机视觉还没解决。而且为系统一收集数据也困难得多。我觉得这也是一个很好的过渡,来谈谈 Project GR00T 和机器人技术。

I think two things. So there's a popular work called Thinking Fast and Slow from psychology. Basically it says that all intelligent beings like humans have two systems. System two is slow and deliberate and does reasoning, and system one is fast, reactive, and intuitive. So the two-system approach to agentic AI. I think right now system two is progressing quite well. It's not perfect, but there are many ways to improve system two, like more data, more environments that we can plug into the frontier so they can do reinforcement learning. These days the frontier models are really bottlenecked by the diversity of environments, not just the pre-training data. We are kind of in this new era of RL post-training. So that's system two. And system one still remains challenging because computer vision is not solved. It's also much harder to collect data for system one. And I think that's also a good segue to maybe talk about Project GR00T and robotics.

Host

这是一个完美的过渡,正是我接下来想谈的。你领导着 GEAR 团队的物理 AI 研究,也就是通用机器人 Project GR00T,正如你描述的。你认为机器人技术的重大挑战是什么?为什么这么难?为什么对我们来说如此容易的系统一,系统二却是困难的部分?为什么系统一在机器人领域这么难?

It's a perfect segue, which is what I want to talk about next. So you're leading physical AI research at the GEAR team, Project GR00T for general-purpose robotics, as you described. What do you think is the grand challenge for robotics? Why is it so hard? Why is system one, which is so easy for us, system two is the hard piece? Why is system one so hard in robotics?

Jim

好问题,Molly。首先,对不熟悉的观众来说,Project GR00T 是英伟达解决通用机器人 AI 的登月计划。我认为要解决这个问题,首先需要定义一个里程碑,对吧?比如这个重大挑战是什么?所以这是我的提议。假设周日晚上你举办了一个大型黑客松派对,把房子搞得一团糟,对吧?到处都是披萨,完全一片狼藉。然后你周一去上班,回家时发现房子干净了,还有烛光晚餐,你分不清是人还是机器人来过。我称之为“物理图灵测试”。这是对真正的图灵测试的致敬,对吧?就像图灵提出的那样。我们社区或多或少已经解决了图灵测试。当你和聊天机器人聊天时,现在很容易骗过人类。但物理图灵测试呢?我觉得它看似简单,对吧?听起来那么简单、那么直观,但它可能是 AI 的下一个或最后一个重大挑战。而且 Molly,正如你暗示的,这对人类来说很容易,但我们正处于莫拉维克悖论的深处,即对人类容易的事情对机器来说却极其困难。为什么会这样?我觉得主要问题是数据问题。我听了伊利亚的演讲,他有一句名言说,大语言模型的预训练正在撞墙,因为我们快没有数据了,对吧?互联网,我们只有一个互联网,而互联网是 AI 的化石燃料。但我想说,你们大语言模型研究者太幸运了,因为你们至少有互联网可以训练,而我们机器人研究者不幸连这个都没有。所以让我共享屏幕,我很想给观众看看。机器人技术是非常视觉化的东西。我觉得一段视频胜过千言万语。让我共享一下,Molly 告诉我如果你看不到。

That's a great question, Molly. Well, first, for the audience who's not familiar, Project GR00T is Nvidia's moonshot initiative to solve general-purpose robot AI. And I think to solve this, first we need to define a milestone, right? Like what is this grand challenge? So here's my proposal. Suppose on Sunday night you hosted a big hackathon party and you wrecked the house, right? Like pizza everywhere. It's a complete mess. Now you go to work on Monday and you come back home to a clean house and a candle-lit dinner, and you couldn't tell if a human or robot had been there. And I'm calling this the physical Turing test. And it's a shout-out to the actual Turing test, right? Like that Turing proposed. And we as a community more or less solved the Turing test. When you're chatting to a chatbot, it's very easy to fool humans now. But what about the physical Turing test? I think it's deceptively simple, right? It sounds so simple, so intuitive, but it's perhaps the next or the last grand challenge in AI. And Molly, as you hinted, it's so easy for humans, but really we're at the thick of the Moravec's paradox, where things that are easy for humans are actually incredibly hard for the machine. And why is that? I think the main problem is the data problem. I listened to Ilya's talk, and he famously said that LLM's pre-training is hitting a wall because we're running out of data, right? Like the internet, we only have one internet, and the internet is the fossil fuel of AI. But I would say that you LLM researchers are so spoiled because you at least have the internet to train on, and we roboticists unfortunately we don't even have that. So let me share screen. I would love to show the audience. Robotics is a very visual thing. I think one video is worth 10,000 words. So let me share, and Molly tell me if you can't see it.

Host

我们能看见。我们看到英伟达总部的杂货店。

We can see it. We see the grocery in Nvidia HQ.

Jim

英伟达总部的杂货店。这是我们为机器人收集的一个典型数据片段。这是机器人的输入,实际上这是输出。这就是机器人关节控制或所有运动控制的样子。这是一个高自由度、高维度的连续信号,不幸的是,你无法从维基百科上抓取这种信号。所以我们必须以某种方式收集数据。我们通过一个称为“遥操作”的过程来做到这一点。这里这位研究员戴着 VR 设备,实时将她的手部姿态传输到实际机器人上。

Grocery in Nvidia HQ. So this is a typical data episode that we collect for robotics. And this is the input to the robot, and actually this is the output. That's what the robot joint control or all the motor control looks like. It's a high degree of freedom, high dimensional continuous signal, and unfortunately you cannot scrape this signal from Wikipedia. So we have to collect the data in some ways. So what we do is through a process that we call teleoperation. So here this researcher is wearing a VR device where it's streaming her hand pose in real time to the actual robot.

遥操作与数据收集 Teleoperation and Data Collection

Jim

所以机器人实际上在物理上镜像她的动作,通过这种方式我们可以控制机器人完成很多任务,比如把蜂蜜倒到吐司上。研究人员基本上通过这个第一人称摄像头看到机器人的动作。不幸的是,这个过程非常手动。你可以看到它非常繁琐,而且它达到了每台机器人每天 24 小时的物理极限。但那是上限。机器人很脆弱,我们一直要照看它们。它们会发脾气,如果机器人上帝那天发慈悲,我们可能能收集到最多四个小时的高质量数据。所以它永远达不到 24 小时。实际上,这是一种我们为机器人收集的人类燃料。这就像 Ilya 所说的所有模型都在预训练时使用的化石燃料。所以问题是,我们如何获得机器人的核燃料?什么是清洁能源?有什么方法可以生成几乎无限的数据来喂养我们的流水线?

So the robot is physically mirroring her movements, and in this way we can control the robot to do a lot of tasks, like pouring honey over toast. The researcher basically sees the robot's actions through this egocentric camera. Unfortunately, this process is very manual. You can see it's very tedious, and it hits the physical limit of 24 hours per robot per day. But that's the upper bound. The robots are fragile; we babysit them all the time. They throw tantrums, and if the robot god is merciful on that day, we may collect up to four hours of high-quality data. So it never hits 24 hours. Really, it's a kind of human fuel that we are collecting for the robot. And that's compared to the fossil fuel that Ilya said all models are pre-training on. So the question is, how can we get the nuclear fuel for robotics? What is the clean energy? What's a way to generate almost infinite amounts of data to feed our pipelines?

Jim

这里有一个快速金字塔:真实数据在最上面,数量最少。网络数据量巨大,而合成数据将是未来。原则上它是无限的,但它受到博士脑循环的限制,因为我们需要想出新颖的方法来实际生成数据,同时也受到算力的限制。买得越多,省得越多。这条消息已经得到我老板的批准。

Here's a quick pyramid: real data is on top and is the smallest amount. Web data is huge, and synthetic data will be the future. It's infinite in principle, but it's limited by PhD brain cycles because we need to come up with novel methods to actually generate the data, and also by compute. The more you buy, the more you save. This message has been approved by my boss.

合成数据与强化学习 Synthetic Data and Reinforcement Learning

Jim

我想展示一些我们在合成数据方面的新研究。Molly,开头你提到了 Eureka。我想展示这个:这是一种技术,让我们能够训练一个高自由度的机器人手,以人类水平做转笔。我想承认我是一个不合格的人类。我从来不会真正转笔,而且我小时候早就放弃了。所以看到我的 AI 为我糟糕的技能复仇,真的很高兴。

I want to show some new research we're doing on synthetic data. Molly, at the beginning you mentioned Eureka. I want to show this: it's a technique that allows us to train a high-degree-of-freedom robot hand to do pen spinning at human level. I want to admit that I am a subpar human. I could never really do pen spinning, and I gave up a long time ago in my childhood. So it's really glad to see my AI avenging my poor skills.

Jim

我们如何获得数据来训练这个?答案是实际上没有监督数据集。这一切都是通过在 GPU 上进行大规模并行模拟的强化学习完成的,因为幸运的是,所有物理方程大多只是矩阵乘法。所以你可以在 CUDA 上运行这个,模拟现实比实时快 10,000 倍,然后你可以将其转移到现实世界。我们怎么做呢?通过一个称为域随机化的过程,在 10,000 个不同的环境中,你可以改变重力、摩擦力和重量,然后你遵循模拟原则。所以假设一个 AI 模型能够掌握具有不同物理参数的百万个现实,那么它极有可能零样本泛化到第一百万零一个现实,也就是我们自己的物理世界。

How do we get the data to train this? The answer is there is actually no supervised dataset. It's all done by reinforcement learning in massively parallel simulations on the GPU, because luckily all the physical equations are mostly just matrix multiplications. So you can run this on CUDA, simulated reality at 10,000 times faster than real time, and then you can transfer this to the real world. How do we do that? Through a process called domain randomization, where for 10,000 different environments you can vary the gravity, friction, and weights, and then you follow the simulation principle. So suppose an AI model is able to master a million realities with different physical parameters, then it's highly likely that it's going to zero-shot the million-and-first reality, which happens to be our own physical world.

Jim

为了展示一些快速视频,我们可以运行一个机器人行走模拟,然后将其转移到现实世界。我们可以控制一个四指手来操纵一个立方体。而这个最有趣:我们有一个机器狗在瑜伽球上行走并试图保持平衡。实际上,瑜伽球很难模拟。它有弹性,很软。但我们甚至不模拟那个;我们只是用一个硬球体。但我们做了很多域随机化,然后我们是机器人专家,我们很奇怪,所以我们直接转移并在街上运行。这是一个真正的机器狗在街上跑瑜伽球。有时我觉得自己是《黑镜》剧集的导演。太奇怪了。我的朋友看到这个视频后,告诉我他试着让他的狗玩瑜伽球,但做不到。所以从某种意义上说,我们实现了超级狗的性能。

Just to show some quick videos, we can run a robot walking simulation and then transfer it to the real world. We can control a four-finger hand to manipulate a cube. And this one is most interesting: we have a robot dog walking on a yoga ball and trying to stay balanced. Actually, the yoga ball is very hard to simulate. It's bouncy, it's soft. But we don't even simulate that; it's just a hard sphere that we do. But we do a lot of domain randomization, and then we're roboticists, we're weird, so we just transfer and run it in the streets. It's an actual robot dog running on a yoga ball in the streets. Sometimes I feel like I'm the director of a Black Mirror episode. It's just so weird. My friend, after seeing this video, told me he tried a yoga ball with his own dog, and it couldn't do it. So in a sense, we have achieved super-dog performance.

Jim

当然,我们可以将其扩展到更复杂的具身形态,比如模拟中的一支人形机器人军队。这些人形机器人在仅仅两个小时的模拟时间内经历了相当于 10 年的高强度训练。有了这么多数据,我们可以训练一个 300 万参数的神经网络——不是数十亿,只有 300 万参数——来控制机器人身体保持平衡、行走、跑步、跳跃和跳舞。所以基本上,我们在模拟中训练的这个小型神经网络可以捕捉到我们人类一直在做的那种潜意识运动协调,并且它在板载 GPU 上以毫秒级运行。

Of course, we can scale this to a lot more complex embodiments, like an army of humanoids in sim. These humanoids are going through 10 years' worth of intense training in only two hours of simulation time. With that much data, we can train a 3-million-parameter neural network—not billions, only 3 million parameters—to control the robot body to stay balanced, walk, run, jump, and dance. So basically, this small neural network that we train in simulation can capture this kind of subconscious motor coordination that we as humans do all the time, and it runs at millisecond scale on a GPU on board.

Jim

这实际上是在 NVIDIA 总部 Voyager 大楼拍摄的。如果你参观 Voyager,你可能会遇到这些小客人。别见外——挥手,向跑来跑去的机器人打招呼。别忘了参观我们的机器人医院。这些小家伙和我们的研究人员一样努力工作。拍摄期间没有机器人受到伤害。

This is actually shot at the NVIDIA headquarters, the Voyager building. If you visit Voyager, you might run into these little guests. Don't be a stranger—wave, say hi to the robots running around. And don't forget to visit our robot hospital. These little guys work as hard as our researchers. No robots were harmed during filming.

主持人评论与世界模型讨论 Host Comment and World Models Discussion

Host

太棒了。硅谷的又一天,你看到机器狗在瑜伽球上,你看到机器人在 NVIDIA 总部做心肺复苏。

Too good. Another day in Silicon Valley where you see a robot dog on a yoga ball. You see a robot getting CPR at NVIDIA HQ.

Jim

是的,我们生活在一个神奇的世界。

Yes, we live in an amazing world.

Host

这只会发生在硅谷。我们是最奇怪的。

It only happens in Silicon Valley. We are the weirdest.

Jim

没错。我们非常奇怪。

Exactly. We're very weird.

Host

带我们进入另一个但相关的方向。我们谈到了合成数据,我们谈到了遥操作。我想听听你对世界模型的看法,这是今天谈到的一些人开创的新范式,可能在帮助解决机器人数据问题方面发挥作用。我很想听你谈谈世界模型,你认为它们在这个整体图景中处于什么位置,它们是否能在数据方面提供帮助。

Taking us in another but related direction. We talked about synthetic data, we talked about teleop. I want to hear your view on world models, which is this newer paradigm pioneered by some of the folks we've talked about today, and may have a role in helping solve the data problem for robotics. I'd love to hear you talk a little bit about world models, where you think they fit into this overall picture, whether they can help on the data side.

Jim

当然。首先,让我定义世界模型。我认为这个词对不同研究者意味着不同的东西。对我们来说,世界模型意味着这种预训练的视频基础模型,它可以将动作作为输入,然后以像素形式预测未来。这是我日常使用的世界模型的操作性定义。那么,为什么这些模型如此强大和有用?我喜欢做一个类比:我真的很喜欢把视频世界模型想象成奇异博士。它们从数十亿个互联网视频中看到了几乎每一种可能的物理现象。模型基本上学习了一种视觉现实的叠加。就像大语言模型从它们在互联网上阅读的所有小说和文章中学习各种不同人格的叠加一样。所以视频模型是视觉现实的叠加,当你提示这些模型时,你基本上将无限可能性坍缩成一个选定的未来。这样,这些视频生成模型实际上是反事实模拟引擎。它们已经看到了所有物理,并且知道在给定某些动作的情况下事物应该如何表现。

Absolutely. First, let me define world models. I think the term means different things to different researchers. To us, world model means this kind of pre-trained video foundation model that can take as input actions and then predict the future in pixels. That's an operational definition of world model that I work with on a daily basis. Now, why are these models so powerful and useful? I like to make an analogy: I really like to think of video world models as kind of like Doctor Strange. They have seen almost every possible physical phenomenon from billions of internet videos. The model basically learns a superposition of visual realities. It's just like LLMs learn a superposition of all kinds of different personalities from all the novels and articles they read on the internet. So the video models are a superposition of visual realities, and when you prompt these models, you basically collapse the infinite possibilities into one chosen future. In this way, these video generation models are actually counterfactual simulation engines. They have seen all the physics and they know how things should act given certain actions.

机器人世界模型 World Models for Robotics

Jim

比如你打翻一瓶水、打翻一杯水,它就会知道水会洒出来。它具备这种预测能力,而这对机器人来说极其重要,因为人类在做出某些动作时,内心会有一个关于动作会产生什么结果的心理模型,然后你会隐式地规划以优化动作。所以我们将世界模型视为一种替代方案,也是一种非常强大的新方式,既可以用来生成数据,甚至可能成为机器人大脑的骨干。我想再次共享屏幕,展示我们实验室最近的工作。让我共享屏幕。我们几分钟前刚展示过这段视频,其实我骗了你们所有人。这里面没有一个真实像素,完全是由视频世界模型生成的。我们把这个想法和算法叫做 Groot Dreams,对吧?就像人形机器人会梦见像素羊吗?名字的灵感就来自这里。它的工作原理是,你拿一个预训练的视频生成模型,然后在收集到的机器人数据上进行微调。我们发现,由于在数十亿互联网视频上的预训练,这个世界模型实际上能够学习机器人的物理约束,也就是机器人的所有机械参数,然后它就能在少量后训练数据下模拟机器人的行为。所以左边和右边其实是完全相同的初始帧,唯一区别是你给它不同的提示词,它就会模拟出不同的未来。左边是“拿起苹果放到平底锅里”,右边是“拿起罐子放到平底锅里”,模型知道如何遵循这些语言指令,而且所有这些像素都是生成的。另一个好处是,因为它是端到端学习的模型,所以它不在乎场景有多复杂。我认为这是它相对于传统图形渲染管线的最大优势,传统管线中场景越复杂,艺术家需要投入的精力就越多,而且还需要算力去做光线追踪等等。所以这里是一个更复杂的场景,视频世界模型不在乎,它照样能模拟得很好。然后我们就能在更多样化的场景上运行它。你可以看到银器上所有复杂的反射,它都能做到,而无需显式实现光线追踪算法。利用这个流程,我们可以生成大量合成数据,我们称之为神经轨迹。我们可以把这些数据当作真实机器人数据,一起用于训练,这样在机器人任务上就能表现得更好。所以我认为世界模型仍处于早期阶段,我们有一些初步的想法让它工作,但还需要做更多研究。

So let's say you drop a bottle of water, you drop a cup of water, it will know that the water will spill. It has this predictive ability, and for robotics this ability is super important because for humans, when you do certain actions, you have this mental model of what kind of results the action will yield, and then you plan implicitly to optimize for actions. So we see this world model as an alternative and a very powerful new way to either generate data or even become the backbone for a robot brain. And I would like to share screen again. I want to show very recent work from our lab. So let me share screen. We just showed this video a few minutes ago, and I tricked all of you. There is not a real pixel in this. It's actually generated entirely by a video world model. So we call this idea and algorithm Groot Dreams, right? Like, does a humanoid robot dream of pixelated sheep? That's where this name is inspired by. So how it works is you take a pre-trained video generation model and then you can fine-tune it on the robot data that you collected. And we found that because of all the pre-training on billions of internet videos, this world model is actually able to learn the physical constraints of a robot, all the mechanical parameters of the robot, and then it's able to simulate how the robot will behave given a small amount of post-training data. So here on left and right it's actually the exact same initial frame, and the only difference is you give it different prompts, and then it will simulate different futures. So on the left is pick up the apple and place on the pan, and on the right pick up the can and place on the pan, and the model knows how to follow these language commands, and all these pixels are generated. And also the other good thing is because it's an end-to-end learned model, it doesn't care how complex the scene is. And I think that's its strongest advantage over traditional graphics pipelines, where the more complex the scene is, the more effort artists need to put in, and also the compute you need to do ray tracing and so on. So here is a much more complex scene. The video world model doesn't care. It's going to simulate just as well. And then we're able to run this on a lot more diverse scenes. And you can see all the complex reflections on the silverware, it's able to do that without implementing a ray tracing algorithm explicitly. And then using this pipeline we can generate lots and lots of synthetic data that we call neural trajectories. And we can add this as if it's real robot data and then train on those data together, and we'll be able to perform a lot better on the robot task. So I would say that world model is still at its early stage. We have some initial ideas to make it work, but a lot more research needs to be done.

Host

太酷了,感谢你的分享,再次看到这些演示真的很棒。你提到了机器人大脑。我想把视角拉远一点,更广泛地谈谈机器人基础模型。构建机器人基础模型的原则是什么?目前机器人基础模型的发展处于什么阶段?

Very cool. Thank you for sharing that, and again, amazing to see the demos. You mentioned robotic brains. I want to zoom out and talk a little more broadly about robotic foundation models. What are the principles to build robotic foundation models, and where are we in the development of robotic foundation models today?

Jim

当然。我认为我非常遵循最初把我带到这里的那个理念。我之前提到,我基本上是在选择难题,然后选择简单优雅的方法,因为它们最能随算力和数据扩展。同样的原则我也应用到了机器人领域。我们可以用四个字来概括:数据极多、模型极简。我提到收集数据有很多不同的方式,对吧?我们可以通过遥操作收集,可以用 Isaac Lab 这样的模拟器生成数据,还可以用世界模型生成数据。所以数据管线可以非常复杂,但模型应该尽可能干净、尽可能端到端,因为这样你就能利用所有不同的数据源,而无需为模型构建单独的训练配方和损失函数。它输入像素,直接输出你看到的电机控制信号,也就是高维连续信号。今年早些时候,我们开源了一个名为 GR00T N1 的模型,将这种方法付诸实践。GR00T N1 模型有系统一和系统二,对吧?系统二是做推理的 VLM,系统一实际上是一个扩散模型,以超过 100 赫兹的频率渲染动作。尽管它采用了这种模块化方法,但它实际上是一个通过反向传播完全端到端训练的模型。这就是我们的方法。我们正在构建一个生态系统,我们开源一切,开源训练配方、模型参数。如果你在网上搜索 GR00T,今天就可以下载模型并试用。

Absolutely. And I think I'm very much following the vibe that brought me here at this moment in the first place. So I mentioned that I was basically choosing hard problems and then selecting for simple and elegant methods because they scale the best with compute and data. And it's the same principle that I'm applying to robotics. And we can summarize that in four words: data maximalist and model minimalist. I mentioned there are many different ways to collect the data, right? We can collect through teleoperation, we can generate data using simulators like Isaac Lab, and then we also generate data using world models. So the data pipeline can be very complex, but the model should be as clean as possible, as end-to-end as possible, because in this way you can leverage all the different data sources without building separate training recipes and loss functions for the models. So it takes as input pixels and then directly outputs the motor control signals that you saw, like the high-dimensional continuous signals. And earlier this year we open-sourced a model called GR00T N1 that operationalizes this approach. So the GR00T N1 model has System One and System Two, right? System Two is a VLM that does reasoning. System One is actually a diffusion model that renders the actions at more than 100 hertz. And despite this kind of modular approach, it's actually an end-to-end fully trained model through backpropagation. So that's how we approach it. And we're building an ecosystem. So we open-source everything. We open-source the training recipe, the model parameters. And if you search for GR00T online, you can download the model and try it out today.

Host

这是对生态系统的巨大宝贵贡献。感谢你的分享,也鼓励大家去看看。吉姆,我们之前聊过机器人领域的 GPT-4 时刻。

It's a hugely valuable contribution to the ecosystem. So thank you for sharing, and encourage others to check it out as well. We've talked a little bit before, Jim, about the GPT-4 moment in robotics.

Jim

是的。

Yes.

Host

你的预测是什么?我们什么时候会迎来机器人领域的 GPT-4 时刻?在那之前需要解决哪些研究和商业问题?

What's your prediction? When will we get the GPT-4 moment in robotics, and what are the research and commercial problems that need to be solved before that happens?

Jim

我确实经常问这个问题,非常感谢莫莉提出这个问题。我认为整个社区总体上倾向于高估短期、低估长期。我想说,我不确定机器人领域的 GPT-4 时刻具体何时到来,但我确定它一定会到来。这是我的推理,对吧?今年是 2025 年,AlexNet 是在 2012 年提出的。AlexNet 被视为一个里程碑,对吧?巨大的革命性时刻。但观众们知道 AlexNet 有多烂吗?那是个糟糕的模型。AlexNet 做的是物体分类,对吧?比如一千个类别中的一个,比如猫对狗对飞机。它的准确率只有 62.5%,意味着它超过三分之一的时间会混淆猫、狗和飞机。它完全不可用,对吧?尽管从方法论角度看它是革命性的,但性能非常差。然后快进到 2025 年,我们今天有什么?我们有通过图灵测试的大语言模型。最新的编程模型和推理模型在国际数学奥林匹克竞赛中赢得金牌,AlphaFold 刚刚获得诺贝尔奖。这就是 2025 年发生的事。过去 12 年发生的事情太疯狂了。而且技术不是线性发展的,而是指数级的。

I do ask this question a lot, and thank you so much, Molly, for bringing this up. So I think the community tends to, in general, overestimate the short term and underestimate the long term. I would say I don't know for sure exactly when the GPT-4 moment for robotics will happen, but I know for sure when it will certainly happen. So here's my reasoning, right? So this year is 2025, and AlexNet was introduced in 2012. AlexNet was held as this milestone, right? Huge revolutionary moment. But does the audience know how crappy AlexNet was? It's a horrible model. What AlexNet did is classifying objects, right? Like one out of a thousand categories, like cats versus dogs versus airplanes. And its accuracy was only 62.5%. Meaning that it confuses cats and dogs and airplanes more than a third of the time. It's completely unusable. Right? Even though it was a revolution from a methodology point of view, performance was super bad. And then fast forward to 2025, what do we have today? We have LLMs that passed the Turing test. We got the latest coding models and the reasoning models winning gold medals at IMO, and then AlphaFold just won a Nobel Prize. And that's what's happening in 2025. It's crazy what happened in the last 12 years. And then it so happens that technology doesn't progress linearly. It's exponential.

机器人数量预测 Robot count prediction

Jim

如果我们把 2025 年作为中点,再加上 13 年,那就是 2038 年,正好处在 AlexNet 和 2038 年的中间。而博士生们会拖延,所以我们四舍五入再多给他们两年。到 2040 年,我相信世界上智能机器人的数量将超过 iPhone 的数量。我说这话有超过 95% 的把握,这件事一定会发生。

So if we take 2025 as a middle point and add 13 years from now, that will be 2038, right where we're in the middle between AlexNet and 2038. And then PhDs procrastinate, so let's round up and give them two more years. By 2040, I believe the number of intelligent robots in the world will be more than the number of iPhones. And I say that with more than 95% certainty that it's going to happen.

Host

这个说法很有挑衅性,我想暂时把眼光放短一点。当你思考今天的模型和未来两三年时,很多人都在想自己旅程的下一章会是什么样子,无论是创业还是加入一家公司。鉴于模型当前的能力,你认为哪些应用领域兼具市场成熟度和技术可行性,能在未来两三年内真正落地?

That's a provocative framing, and I want to think shorter term for a moment. When you think about models today and the next two to three years, many folks are thinking about what the next chapter of their journey looks like, whether that's starting a company or joining one. Given models' current capabilities, which application areas do you think have the right combination of market readiness and technical feasibility to be really tractable in the next two to three years?

短期视野与突破 Short-term horizon and breakthroughs

Jim

这是个很好的问题。让我也稍微聚焦一下更短的时间范围。再说一遍,今年是 2025 年。如果我们回顾 2020 年 5 月,那是 GPT-3 首次推出的时候。从 GPT-3 到在 IMO 赢得金牌,我认为大约需要三次迭代,对吧?有 SFT,然后有 RLHF,即基于人类反馈的强化学习,这是两大突破,然后还有像 o1 和 o3 那样的强化学习,用于推理的强化学习。所以大约三次重大的模型能力突破把我们带到了 2020 到 2025 年。如果我们再加上五年,那就是 2030 年。所以我认为在方法论方面,机器人基础模型的突破大概需要三到五年才能实现。

That is a great question. Let me also zoom in a little bit on the shorter time horizon. Again, this year is 2025. If we look back to May 2020, that was when GPT-3 was first introduced. From GPT-3 to winning a gold medal at IMO, I think that roughly takes three iterations, right? There's SFT, then there's RLHF, reinforcement learning from human feedback, as two big breakthroughs, and then reinforcement learning like the o1 and o3 style, RL for reasoning. So roughly three major model capability breakthroughs carried us from 2020 to 2025. If we add five more years, that will be 2030. So I think in terms of methodology, the breakthrough in robot foundation models probably takes three to five years to get us there.

应用领域:可编程工厂 Application areas: programmable factories

Jim

那么在接下来的五年时间范围内,从应用的角度来看,如果我们有了机器人基础模型的 GPT-3 时刻,我认为以下应用领域是我最兴奋的。第一个是可编程工厂。我认为工厂和工业用例可能是最先被攻克的,因为它们更规律、更结构化,对安全性的要求也更低。你不需要人类在工厂里围着机器人转,对吧?你可以让工厂基本上封闭起来,然后机器人自己操作。所以它不需要非常严格的安全标准。而且制造过程被理解得很透彻,规格也很明确,所以它不需要太多的泛化能力。它确实需要一些泛化,但不需要太多。如果我们有一个 GPT-3 或 GPT-4 级别的机器人基础模型,那么我们就能快速重新编程一组智能机器人,让它们几乎一夜之间就能制造新产品。你基本上可以给机器人看几个演示,然后就能对它们进行编程,对吧?你通过演示、通过数据来编程,而不是手工编程。这就是为什么现在的机器人很难重新编程,因为人们是手工编程的,而大型工厂真的很难重新调整。你需要数月甚至数年的专业知识才能重新编程。但一旦我们有了这些基础模型,工厂的可编程性就会大大增强。所以我认为这是最先的领域之一。

So now, in the next five-year horizon, in terms of applications, if we have this GPT-3 moment for robot foundation models, I think the following application areas are the ones I'm most excited about. The first is programmable factories. I think factories and industrial use cases are probably the first to be tackled because they are more regular, more structured, and less demanding on safety. You don't need humans to be around robots inside a factory, right? You can have the factory basically walled off and then the robots operate themselves. So it doesn't need very strict safety criteria. Also, the manufacturing process is quite well understood with very good specifications, so it doesn't need too much generalization. It does need some generalization, but not too much. And if we have a GPT-3 or GPT-4 level robot foundation model, then we can quickly reprogram a set of intelligent robots so they can manufacture a new product almost overnight. You can basically show the robots a few demonstrations and then you can program them, right? You program by demonstration, by data, instead of programming by hand. That's why robots are very difficult to reprogram these days, because people do it by hand, and big factories are really hard to repivot. You need months, if not years, of expertise to reprogram something. But once we have these foundation models, the factories will be a lot more programmable. So I think that's one of the first areas.

应用领域:自动驾驶湿实验室 Application areas: self-driving wet labs

Jim

我兴奋的第二个领域,同样非常规律、结构化,对安全性要求也不高,是自动驾驶湿实验室。很多科学实验,比如化学或药物发现中的实验,仍然由人类完成,对吧?即使有特殊的化学混合设备,人们不再手工操作了,但想想那些生物学实验——过程中仍然有很多人使用那些设备。我们可以用自动化的机器人流水线完全取代这些。它会看起来很像可编程工厂,但这次是应用于科学发现。所以你可以通过演示来收集数据并编程整条流水线,而不是通过硬编码。

The second area I'm excited about, also really regular, very structured, and not that demanding on safety, is self-driving wet labs. A lot of scientific experiments, like in chemistry or drug discovery, are still carried out by humans, right? Even though there are special equipment for chemical mixing that people don't do by hand anymore, think about those biology experiments—there are still a lot of humans in the process using those equipments. We can completely replace that with an automated robot pipeline. It's going to look very similar to a programmable factory, but this time applied to scientific discoveries. So you can collect data and program this whole line by demonstration instead of by hard coding.

应用领域:智能体集群 Application areas: agentic fleets

Jim

然后是智能体舰队。我们讨论了很多关于单个机器人的话题,但如果我们有一整支机器人舰队协同工作,形成这种蜂巢思维呢?然后我们可以运行这些智能体流程,这些由中央大脑控制的多智能体流程,将任务分解并分配给整个机器人舰队,这样它们就能更好地支持工厂装配线,以及多阶段的科学实验。

Then there is the agentic fleet. We discuss a lot about single robots, but what about having this kind of hive mind of a whole fleet of robots all working together? Then we can run these agentic flows, these multi-agent flows controlled by a central brain, to break down and distribute a task to a whole fleet of robots so they can better support these factory assembly lines and also multi-stage scientific experimentation.

应用领域:家庭使用 Application areas: home use

Jim

最后是家庭使用。我认为家庭使用要完全自动化难得多。它更加无结构——每个家庭的杂乱方式都不一样。而且你肯定会有很多安全顾虑,对吧?有孩子、老人,而这些机器人很重,让任何一台机器人摔倒并砸到人可不好玩,对吧?所以我认为家庭使用,还有隐私等所有这些问题,我会说家庭使用还有点遥远。但我真心希望我们能在五年内看到。如果五年内不行,那到 2040 年肯定可以。

And finally, home use. I think home use is just a lot harder to fully automate. It's way more unstructured—every house is messy in a different way. And you for sure have a lot of safety concerns, right? There are kids, elderly people, and these robots weigh a lot, and it's not fun having any of those robots fall down and fall on a person, right? So I think home use, and there's also privacy and all those concerns, I would say home use is still a bit out there. But I really wish that we can see that in five years. If not five years, definitely by 2040.

结束语 Closing

Host

非常鼓舞人心的例子和应用领域。谢谢你,Jim。

Very inspiring examples and application areas. Thank you, Jim.

Host

您收听的是 Radical AI 创始人大师课系列的 Radical Talks 特别节目。特别感谢 NVIDIA 的 AI 总监兼杰出科学家 Jim。如果您喜欢我们的节目,请务必订阅并给我们留言。感谢收听。

You've been listening to a special episode of Radical Talks from the Radical AI Founders Masterclass series. A special thank you to NVIDIA's director of AI and distinguished scientist, Jim. If you like what you hear, please be sure to subscribe and leave us a comment. Thanks for listening.

互动版:逐字朗读 + 针对本期提问 →