From AlexNet to OpenAI: A Decade of AI Milestones
打开互动全文版(中英对照 + 朗读 + 问答)→Jim 分享了他从 2012 年研究 AlexNet 到在 OpenAI 工作的历程,强调了人工智能历史上的关键时刻以及在对的时间出现在对的地方的重要性。
Jim shares his journey from studying AlexNet in 2012 to working at OpenAI, highlighting key moments in AI history and the importance of being at the right place at the right time.
Jim,感谢你来做客。
Jim, thank you for joining us.
谢谢邀请。看到这么多观众我真的很激动。谢谢大家。
Thanks for having me. And I am overwhelmed by such a great crowd. Thanks, thank you all for having me.
是啊,他说‘我没想到会有这么多人’。我说‘这大概只有申请者的 20%’。确实如此。你在 Twitter 上有 15 万粉丝吧?在圈内小有名气。我们很荣幸能请到你。谢谢。好了,我简单介绍了一下你的背景,但那远不能概括你的成就。所以我想从基础开始,你能给我们讲讲你迄今为止的职业生涯轨迹吗?
Yes, he was like, 'I didn't expect this many people.' I was like, 'This is like 20% of the people who applied.' Yes, it is. And you have like what, 150,000 followers on Twitter, something like that? A little well-known in the community. So we're just very lucky to have you here. Thanks. All right, so I gave a quick bio, but that doesn't really do justice to all you've done. So I want to start out by just laying the foundation and just could you walk us through the arc of your career so far?
好的,我先从我的职业生涯讲起,因为我觉得它很好地串联了今天的主题,也串联了 AI 历史上的许多里程碑。这也能让你了解我是如何走到今天,做现在这些事的。乔布斯说过,我们只能向后看才能把点连起来,不能向前看。现在回顾我在 AI 领域的十年,我觉得有很多神奇的时刻想和大家分享。总的来说,我觉得自己超级幸运,因为很多时候我只是恰好在正确的时间出现在正确的地方,通常不是出于选择或设计,而是命运的随机性。我 2012 年在哥伦比亚大学读本科,2012 年是深度学习元年。AlexNet 在 2012 年问世,我是最早读到那篇论文的人之一。我至今记得那种兴奋,因为在 AlexNet 之前,计算机视觉是一个巨大的、极其混乱的流水线,有很多特征和统计方法。AlexNet 出现后,它用一个单一的神经网络直接把像素映射到输出,就这么简单,然后在大量数据上训练。那对我来说是非常鼓舞人心的时刻,我心想,我一定要学习 AI、机器学习,特别是神经网络。后来我加入了哥伦比亚大学的一个本科生研究小组,任务是语音识别。但那时深度学习还不是主流,他们做语音识别的方法是用音素分类。很多人没听说过音素,这正是我要说的,那东西很奇怪。然后做音素分类,再用一些老式的序列模型,经典编程那一套。整个过程非常繁琐。我对那并不感冒,但那是我第一次认真的研究尝试,不过不够优雅。后来我听说吴恩达去了百度担任首席科学家。2015 年夏天,我申请了百度在森尼维尔的实习。实际上,森尼维尔办公室是吴恩达团队的,他在构建一个叫 Deep Speech 的系统。从名字就能看出,Deep Speech 是一个神经网络。吴恩达的想法是,去掉所有音素和语言学的东西,只用单个神经网络从原始音频波形直接映射到句子,然后在大量数据上扩展。那时我被分配了一个导师,不是我自己选的,是 Dario Amodei。这个名字你可能听说过,他现在是 Anthropic 的 CEO。那个夏天我和 Dario 紧密合作,我们一起结对编程。Dario 当时在谈论一种叫 Scaling(规模扩张)的东西,他说我们的神经网络不够深、不够大,我们需要更深更胖的网络。老实说,我当时没意识到它的重要性,我并没有完全相信 Scaling 的故事,我持怀疑态度,不是反对,而是怀疑。Dario 坚持要做出更大更深的神经网络。那是 2015 年,还没有 Transformer,只有循环神经网络之类的老方法。但即使对这些方法,你也可以扩展,增加参数。那是我第一次接触工业级深度学习系统和 Scaling。后来 Dario 离开了百度,2016 年夏天他对我说:‘Jim,我们要创办一家新公司,会很激动人心。我们在百度度过了一个很棒的夏天,你想加入吗?’我说:‘当然,Dario,你去哪儿我就去哪儿。’那家公司就是 OpenAI。Dario 和 Andrej Karpathy、Ilya Sutskever 等人共同创立了 OpenAI。那时的 OpenAI 非常神奇。看看他们的技术团队名单,全是机器学习领域的大牛。比如 Andrej Karpathy、发明 GAN 的 Ian Goodfellow、发明 Adam 优化器(几乎所有神经网络都依赖它)的 Diederik Kingma、后来共同发明扩散模型的 Jonathan Ho、还有 Peter B.,以及发明 PPO 的 John Schulman。2016 年夏天他们都在那里。我当时简直是房间里最笨的人,大概有 35 个人。我 2016 年的实习项目是 OpenAI 最早尝试 AGI 的项目之一。OpenAI 早在 2016 年就在谈论 AGI。那个项目叫 OpenAI Universe。想法很简单:训练一个 AI 智能体,能读取电脑屏幕上的像素,然后控制键盘和鼠标。想想看,这是智能体最通用的接口,因为我们在数字空间做的所有事情——写邮件、玩游戏、浏览网页、用 Photoshop、用任何软件——都可以通过读取像素和生成键盘鼠标动作来表达。这听起来很吸引人,那是 OpenAI 启动的一个大项目。我负责的部分,顺便说一句,当时我的实习导师是 Andrej Karpathy,Andrej 和我,还有后来来自斯坦福的 Tim,负责一个叫 World of Bits 的子项目。这个名字是 Andrej 起的,他很会起名字。World of Bits 是 OpenAI Universe 的一部分,专门控制网页浏览器:读取浏览器像素,用键盘鼠标完成网页任务,比如浏览 Expedia、搜索信息、玩一些简单的网页游戏等等。那是我实习项目的早期尝试,但最终没有成功,因为那时没有 Transformer,没有基础模型,我们只用强化学习,而强化学习在那个规模上根本没有希望。
Yes, so I will start with the story of my career because I think it ties together nicely today's topics and also many of the kind of historic milestones in AI. And that also provides some perspective on how I got here today and doing what I'm doing right now. Now, you know, Steve Jobs said right, like we can only connect the dots looking backwards, we cannot connect dots looking forward. And now kind of looking backward at my, you know, 10 years of history in AI, I think there are just many kind of magical moments I want to share with you. And also just to sum it up, I think I was like super lucky because many times I was just happen to be at the right place at the right time, and typically not by choice but not by design, but you know, just because of randomness of fate. So I started my undergraduate at Columbia University in 2012, and 2012 was the first year of deep learning. So AlexNet was dropping in 2012, and I was among kind of the first crowd reading that paper. And I still remember the excitement because before AlexNet, right, computer vision was this giant pipeline, super messy pipeline, you have a lot of features and statistical methods. And AlexNet came and just blew everything out of the water because it's a single neural network that maps pixels directly to the output, and that's it. And you train it on lots and lots of data. So that was a very inspiring moment for me, and I was like, I need to study AI, machine learning, neural networks in particular. And then I joined an undergrad research group at Columbia University, and my task was speech recognition. But back then, you know, deep learning wasn't a conventional wisdom, so the way they did speech recognition was you do things like phoneme classification. Many of you haven't heard about phoneme, and that's exactly my point, right, it's such a weird thing. And then you do phoneme classification and you do some sequence models that are like old school, you know, classical programming stuff. And then you go through a very hairy pipeline. So I wasn't impressed by that, but I think that was still, you know, my first kind of serious research effort. But it wasn't elegant enough. And then I heard that Andrew Ng went to Baidu to be the chief scientist. So in the summer of 2015, I applied to Baidu's internship at Sunnyvale. Actually, the Sunnyvale office was kind of open for Andrew's group, and he was building a system called Deep Speech. So you can tell from this name, right, like Deep Speech is a neural network. And Andrew's idea was, let's get rid of all the phonemes, all the linguistic stuff, and you know, the old stuff. Let's just have one neural network that maps from raw audio wave to the sentence, to the spoken sentence, and that's it. And we will just scale it on a lot of data. And then at that time, I was assigned, again not by my choice, but I was assigned to an intern mentor called Dario Amodei. And that name may ring a bell, he's now the Anthropic CEO. And back then, for that summer, I worked very closely with Dario. And we did like pair programming together. So Dario was talking about something called scaling up, and he was like, we got to have like our neural networks are not deep enough, are not big enough. We just need deeper and fatter neural networks. So to be very honest with the audience, I did not recognize the significance of it. I wasn't, you know, completely buying into this scaling up story. I was skeptical. I wasn't against it, but I was skeptical. And Dario pushed for just making bigger and deeper neural networks. And back then it was 2015, there was no Transformer, it was like recurrent neural networks, some older methods. But even for those methods, you can scale it up, you can throw in more parameters. So that was kind of my first exposure to like industrial deep learning system and scaling up. And then Dario left Baidu, and in the summer of 2016 he said, 'Jim, like we are starting a new company. It's going to be exciting. And we had a great summer last summer at Baidu, so do you want to join me in this new effort?' I'm like, 'Hell yeah, Dario, wherever you go, I will follow.' And that company is called OpenAI. So Dario co-founded OpenAI with people like Andrej Karpathy, Ilya Sutskever. And back then, OpenAI was quite a magical place. Basically, if you look at their list of staff, you know, staff of their technical team, these are all big names, the biggest names you can see in machine learning. So for example, Andrej Karpathy was there, Ian Goodfellow who invented GAN, Diederik Kingma who invented the Adam Optimizer which basically every neural network relies on, Jonathan Ho who later was the co-inventor of diffusion models, and Peter B., and you know, John Schulman who invented PPO. They were all there in the summer of 2016. And I was like literally the dumbest person in the room at that time. There were like 35 of them, around that number. So my internship project in 2016 was one of OpenAI's earliest attempts at AGI. OpenAI has always been talking about AGI way back in 2016. And that project is called OpenAI Universe. So the idea is simple: we want to train one AI agent that can read the pixels on a computer screen and then control the keyboard and mouse. Right, if you think about it, this is the most general interface you can have for an agent because basically everything that we do in a digital space, be it writing emails, playing games, browsing the web, using Photoshop, using any software, can be expressed in reading pixels and then generating actions in keyboard and mouse. It sounds very appealing, and that was a huge initiative that OpenAI started. And the part I was responsible for, and by the way at that time my intern mentor was Andrej Karpathy, and Andrej and myself, and then later Tim from Stanford, we were responsible for a part of it called the World of Bits. And that's a great name started by Andrej, Andrej is great at naming things. So the World of Bits project was part of OpenAI Universe, and that was to control the web browser specifically: reading pixels from the web browser and control keyboard and mouse to do tasks on the web, like browsing Expedia, you know, searching for information, playing some you know simple web-based games, and so on and so forth. So that was an early attempt at my internship project. But it ultimately didn't quite work out because there was no Transformer, there were no foundation models back then, and we only used reinforcement learning, and there's no hope that reinforcement learning would work at that scale.
在这个极其复杂的空间中学习可以零样本泛化到任何事物,这真的很难。所以 Open Universe 没有成功,然后我在夏天后离开了 OpenAI,开始在斯坦福大学攻读博士学位,导师是李飞飞教授。你们中有些人可能不知道她,她是 ImageNet 的发明者,这要追溯到 2012 年。ImageNet 让 AlexNet 成为可能,因为网络需要大量数据来训练。在 ImageNet 之前,都是小数据集,而 ImageNet 有 120 万张图像,按那个时代的标准来说是巨大的。我再次加入了李飞飞实验室,时机很好,因为计算机视觉领域正在经历转型。想想 AlexNet,它是最早成功的计算机视觉神经网络之一。它拍摄一张静态快照,一张世界的截图,然后告诉你其中的物体,无论是猫、狗、卡车还是飞机。这有点奇怪,因为人类视觉系统并非如此运作。我们是具身的,生活在三维物理世界中,感知着连续不断的像素流。我们不是通过看一堆不相关的图像来学习的,而 AlexNet 正是这样训练的。所以当时,在李飞飞实验室,具身智能体成为一个主题,我们想要的不只是计算机视觉,而是具身计算机视觉——一个具身于物理世界的视觉系统,能够采取行动,观察行动的后果,并利用这些信号来进一步改进自身。这就像视频,源源不断的像素流。这正是我们人类发展视觉系统的方式。所以在那段时间,具身智能体的想法在我心中扎根。此外,在 OpenAI 的 AI 智能体方面,Open Universe 是数字空间中的另一个智能体。然后在 2018 年,我碰巧主持了斯坦福 AI 沙龙。他们每次轮换主持人。斯坦福 AI 沙龙是一个有趣的活动,每周五举行。基本上,有一群人,他们会邀请一些大牌嘉宾。不允许带电子设备,不允许带手机。大家会进行闭门讨论。我碰巧主持了那次活动,他们邀请了 Pedro Domingos 和另一个穿皮夹克的人,那就是黄仁勋。所以 Jensen 和 Pedro 来到斯坦福,我是他们的主持人。结果发现他们不需要主持人,他们基本上是自我主持。Jensen 几乎像个单口喜剧演员,他非常出色,非常有魅力,Pedro Domingos 也是如此。他们进行了讨论。我想,‘好吧,你们自己来吧,让我的工作变得非常轻松。你们搞定了。’之后,我和 Jensen 聊了聊。我告诉他关于具身智能体的事情。他对此非常兴奋。他说,‘AI 智能体,机器人技术将是未来,你应该在 NVIDIA 研究做你毕生的工作。’所以我在 2020 年去那里实习,然后在 2021 年毕业后加入了 NVIDIA 研究。我一直待在那里。所以现在我是一名高级科学家。我在那里工作了两年。在我们继续之前,还有一个里程碑:去年在新闻发布会上,我展示我的 Dojo,还有另一个 AI 的里程碑时刻,那就是遇到了 Kring 和 Josh。那时我仍然是一般智能,对吧?你们展示了 Avalon 的工作,我对此感到惊讶。我们参加了同一个小组讨论,因为 Mind Dojo 和 Avalon 共享很多共同的想法和理念,但也有不同的地方。我们有非常有趣和独特的设计选择。我们可以深入探讨,但这基本上是我经历的快速概述。
Learning in this vastly complex space can generalize to anything zero-shot, it's just really hard. So Open Universe didn't quite work out, and then I left OpenAI after the summer and started my PhD at Stanford with Professor Fei-Fei Li. For those of you who might not know her, she's the inventor of ImageNet, which ties back to 2012. ImageNet was what made AlexNet possible because networks need a lot of data to train. Before ImageNet, there were all small datasets, and ImageNet had 1.2 million images, huge by that era's standard. I joined the Fei-Fei Lab again at a pretty good timing because the field of computer vision was seeing a transition. If we think about AlexNet, it was one of the first successful computer vision neural networks. It takes a static snapshot, a screenshot of the world, and tells you the objects in it, be it cat, dog, truck, or airplane. It's kind of weird because the human vision system doesn't quite work that way. We are embodied, we're living in this 3D physical world, and we perceive a constant continuous stream of pixels. We don't learn by seeing a bag of unrelated images, which is how AlexNet was trained. So at that time, in the Fei-Fei Lab, embodied agents became a theme where we want to have not just computer vision but embodied computer vision, having a visual system that's embodied in the physical world, that can take actions, observe the consequences of the actions, and use that as a signal to improve itself even more. It's all like video, a constant stream of pixels coming in. That's how we as humans develop our visual system. So during that time, the idea of the embodied agent was planted in me. Also, in addition to AI agents at OpenAI, the Open Universe was another agent in the digital space. Then in 2018, I happened to host the Stanford AI Salon. They rotate hosts every time. The Stanford AI Salon is an interesting event, a weekly event on Friday. Basically, there are a group of people, and they invite some really big-name guests. No electronics allowed, no phones allowed. Everybody will have a closed-door discussion. I happened to be the host of that particular event where they invited Pedro Domingos and another guy in a leather jacket, and that's Jensen Huang. So Jensen and Pedro came to Stanford, and I was their host. It turns out that they don't need hosts; they're basically self-hosting. Jensen is almost like a stand-up comedian, he's so good and so charismatic, and so was Pedro Domingos. They had this discussion. I'm like, 'Okay, you guys just go ahead, you're making my job very easy. You got this.' Afterwards, I talked to Jensen. I told him about embodied agents and stuff. He was very excited by this. He said, 'AI agents, robotics will be the future, and you should do your life's work at NVIDIA Research.' So I did an internship there in 2020, and then after graduating in 2021, I joined NVIDIA Research. I've stayed there ever since. So now I'm a senior scientist. I've been there for two years. Then just one more milestone before we move on: last year at the news conference, I was presenting my Dojo, and there was another milestone moment in AI, and that is meeting Kring and Josh. Back then, I was still generally intelligent, right? And you guys were presenting the Avalon work, and I was amazed by it. We were on the same panel together because Mind Dojo and Avalon share a lot of common ideas and philosophies, but also differ in different ways. We had very interesting and very unique design choices. We can get into that, but that's basically a quick overview of my history.
非常有趣。所有这些时刻。我好奇的是,你非常热衷于具身智能体,我认为在这一点上这实际上是一个有争议的观点。所以我很好奇,具身性仍然吸引你的是什么?我认为今天在座的很多构建智能体的人会说,‘哦,我们可以使用语言模型或者多模态模型,这就足以让智能体工作了。’你感兴趣的是什么?为什么?
Super interesting. All of these moments. Something I'm curious about is you're really into embodied agents, and I think at this point that's actually a controversial view. So I'm curious, what is it about embodiment that still compels you? I think a lot of people building agents in the audience today are like, 'Oh, we can use language models or maybe multimodal models, and that's like enough for getting agents to work.' What are you interested in? Why?
是的,所以我认为大型语言模型和多模态模型将是拼图的关键部分,但它们还不够。我认为具身性不仅有用,因为我们可以用具身智能体做机器人技术,并讨论应用的影响。它们不仅有用,而且我认为它们是在通往高级智能的关键路径上。因为如果我们再想想,就像 AlexNet 一样,语言模型以一种非常奇特的方式学习,一种非常不像人类、外星人的方式。语言模型是如何工作的?它们是如何训练的?抓取大量互联网数据,然后通过预测下一个词来暴力学习。但想想人类婴儿是如何学习的。当你出生时,你没有被塞在电脑屏幕前,然后阅读一百万篇维基百科页面作为第一课,对吧?你不会那样做。人类婴儿做的第一件事是用他们的手、脚、耳朵、嘴巴,只是实验,与物体互动,实验,听,看,基本上沉浸在这个世界中,在这个三维世界中,然后进行实验,试错,在这个过程中理解物理。所以有一种说法是,所有婴儿都是摇篮里的科学家,我非常喜欢这个说法,因为我觉得这是 LLM 进化的下一步:给它们具身性,这样它们的知识就不再是空中楼阁,抽象且仅存在于阅读维基百科的统计数据中,而是有根基的。它们通过在物理世界中采取行动、观察反馈并从中学习,使知识变得可执行。在这个循环中,你理解因果关系,理解我们所谓的直觉物理。例如,如果我打翻这瓶水,我的大脑无法计算每个水分子的精确轨迹,但我大致知道它会朝这个方向洒出来,会弄得一团糟,让两个洞生气。我大致知道,对吧?这些就是我们所说的直觉物理。我认为 LLM 产生幻觉的部分原因是它们没有根基。它们没有这种具身性,没有我们拥有的经验,所以它们不知道什么是不可能的。它们基本上只是计算统计数据,因为它们没有我们拥有的生活经验。我认为这至少部分解释了为什么它们会产生那么多幻觉,说出对我们来说明显错误、违背常识的话。所以,是的,具身性是关键。
Yes, so I think large language models and multimodal models will be a key piece to the puzzle, but they are not enough. I think embodiment is not just useful because we can use embodied agents to do robotics and we can discuss the implications of the applications. They are not just useful, but I think they are on the critical path towards unlocking high levels of intelligence. Because if we think about it again, just like AlexNet, language models learn in a very peculiar way, a very unhuman-like, alien way. How does a language model work? How are they trained? Lots and lots of internet data you scrape, and then you brute force by predicting the next word. But think about how human babies learn. When you were born, you were not stuffed in front of a computer screen and then reading a million Wikipedia pages as your first class, right? You don't do that. The first things that human babies do are using their hands, feet, ears, mouth, and just experiment, interact with objects, experiment, listen, see, basically immerse in this world, in this 3D world, and then do experiments in it, trial and error, and understand physics in this process. So there is a saying that all babies are scientists in the crib, and I really like this saying because I feel that is the next step in the evolution of LLMs: to give them embodiment so that their knowledge is no longer castles in the air, abstract and only existing from reading statistics on Wikipedia, but it's grounded. It makes their knowledge executable by taking actions in this physical world, observing the feedback, and learning from that. In this loop, you understand causality, you understand what we call intuitive physics. For example, if I spill this bottle of water, my brain cannot compute the exact trajectory of every water molecule, but I roughly know that it's going to spill in this direction, it's going to make a mess, and make both holes angry. I kind of know that, right? These are what we call intuitive physics. I think partially the reason why LLMs hallucinate is because they are ungrounded. They don't have this embodiment, they don't have the experience that we have, so they don't know what is impossible. They basically just compute statistics because they don't have the life experience that we have. I think that at least partially explains why they hallucinate so much, saying things that are obviously wrong to us, that are against common sense. So yeah, embodiment is the key.
当你说具身智能体时,为了明确起见,你指的是什么?定义是什么?
When you say embodied agents, just to be clear, what do you mean by that? What is the definition?
具身,我指的是能够控制身体并与世界交互的智能体。而且这个世界不一定是物理的,也可以是虚拟的具身智能体,比如在《我的世界》中,那就是一个在模拟环境中的具身智能体,基本上是在 3D 沉浸式现实中。AGI 可以控制身体并进行交互。不过,有人可能会反驳说,在互联网上像比特机器人那样行动也算具身,因为你可以操作浏览器并运行实验。我好奇的一点是,语言模型可能不会做实验——很难让它们像科学家那样去实验,因为它们的训练方式和下一个词预测机制。所以,语言模型之所以不是科学家,可能不是具身的问题,而是某种能产生实验和探索行为的目标函数。你认为具身对于这种行为是必要的吗?
Embodiment, I mean agents that can control a body and then interact with a world. And now that world does not have to be physical; it could be like virtual embodied agents, you know, in Minecraft, for example. That is an embodied agent in a simulation, basically in a 3D kind of immersive reality. And the AGI can control a body, interact. Not to get too into the details, but a devil's advocate might say something like, okay, acting on the internet like a robot of bits counts as embodiment because I can act on the browser and run experiments. Maybe one thing I'm curious about is with language models, maybe they don't experiment. I mean, it's kind of hard to prompt them to experiment like a scientist would, because of the way that they're trained and the next-token prediction. And so maybe the issue of why language models are not scientists is not embodiment but rather some kind of objective function that would result in this kind of experimentation, exploration behavior. Do you think embodiment is necessary for that behavior?
我认为仅通过文本进行探索和利用有一些弱形式。但我感觉具身可能只是更快的方式,因为如果……你的意思是任何形式的行动,能够采取行动是更快的方式。基本上,我认为是的,行动是关键。不过,具身的一个好处是你能获得多感官反馈。所以除了训练文本推理,你还能得到像素和音频。不是像训练多模态模型那样在静态快照上做图像标注,而是在这个具身世界中训练。至少对于计算机视觉来说——我不认为推理也一定如此,推理可能只靠读维基百科就够了——但对于视觉,这绝对是最有效的方式。
I think there are some weak forms of being able to explore and exploit just using text. But I feel that having embodiment is maybe just a faster way to do it because if there's... and you mean like literally any action-taking, being able to take actions is a faster way to do it. Basically, I think yes, taking actions is the key word. Yeah. But one thing about embodiment is you get this multisensory feedback. So in addition to training textual reasoning, you also get pixels and audio. Instead of training multimodal models on static snapshots where you take a bunch of images and caption them, you do it in this embodied world. I think at least for computer vision—I wouldn't say this about reasoning necessarily; for reasoning you may get away with just reading lots of Wikipedia—but for vision, I think that is definitely the most effective way.
具身是个棘手的问题,因为计算机在某种意义上是串行的,而人类不是。在推理中,推理是串行的,所以我们才能用语言进行推理。但我同意,在串行环境中处理多模态输入——我还不知道怎么做。
Embodiment is a tricky thing because computers are serial in a way that humans are not. In reasoning, reasoning is serial, and that's why we can get away with reasoning and language. But I agree that dealing with this multimodal input in a serial environment—I don't know how to do that yet.
是的,还有一个论点:当你具身在一个足够丰富和开放的世界中,并且智能体能够很好地探索并保持好奇心,那么你基本上就能获得无限的数据,对吧?数据是当今基础模型的瓶颈。我们都听说过 OpenAI 的 token 快用完了,对吧?因为所有这些大型基础模型创建者都在抓取互联网,我们正在耗尽互联网上有用的高质量 token。公开可抓取的 token 就那么多。但与具身世界交互会产生无限的 token。这是一个持续的经验流,你创造自己的经验。你不仅拥有无限的 token,而且质量非常高,因为你控制着想要获取什么样的 token。这就是我们所说的科学家:婴儿对他们最好奇的事物采取行动。现在想想什么是好奇心:只有当你无法预测某件事的结果时,你才会好奇。如果我对日常事务非常确定,我就不会好奇刷牙是否会带来很多惊喜。不会,因为你确切知道会发生什么。但如果你去参加某个活动,比如音乐会或第一次去火人节,你会好奇,因为你从未经历过,无法预测会发生什么。就像我走进这个活动时,我没有预料到会有这么多人。现在我很好奇为什么人们感兴趣,因为这很难预测。这就是好奇心的来源。如果你做类似好奇心驱动的探索,基本上就是采取行动来收集当前心智状态下最有用的信息,这是一种极其有效的学习方式。
Yeah, and there's also another argument: when you are embodied in a world that is rich enough and open-ended enough, and if the agent is able to explore very well and keep its curiosity, then you basically get infinite data, right? Data is the bottleneck for today's foundation models. We have all heard about OpenAI running out of tokens, right? Because basically all these big foundation model creators scrape the internet, and we're kind of exhausting the useful tokens and high-quality tokens on the internet. There are just so many that you can publicly scrape. But interacting with the embodied world produces infinite tokens. It is a constant stream of experience, and you make your own experience. You not only have infinite tokens, but you also have very high-quality ones because you control what kind of tokens you want to get. That's what we mean by scientists in the right: babies take actions in things that they are most curious about. Now think about what curiosity is: you are only curious when you cannot predict the outcome of a certain thing. If I am so certain about my daily routine, I'm no longer curious about whether brushing my teeth will bring me a lot of curiosity today. No, because you know exactly what's going to happen. But if you're going to some event, like a concert or Burning Man for the first time, you are curious because you have never experienced it; you cannot predict what's going to happen. Much like when I walked into this event, I did not predict there would be a huge crowd here. Now I'm getting curious why people are interested, because it's hard to predict. That is where curiosity is. And if you do something like curiosity-driven exploration, basically you're taking actions to gather the information that's most useful given your current mental state, and that is an extremely effective way to learn.
我想稍微转向用例。一旦我们有了这些具身智能体,你最兴奋的是什么?它使哪些用例成为可能?
I want to transition a little bit to use cases. So once we have these embodied agents, what are you most excited about? The use cases that it makes possible.
是的,我喜欢这个问题。我一直在谈哲学,但现在该落地了,谈一些实际有经济价值的东西。太好了。我认为 AI 智能体应用分为三个领域:一是软件,我认为 Ken 可以多谈谈这个——你正在这个领域构建 AI 智能体。游戏是第二个主要应用,机器人基本上是物理世界中的 AI 智能体。那么先谈谈软件。OpenAI Universe 是早期尝试之一,让数字智能体运行,但使用强化学习不是正确的方法。我确实认为今天的大模型解决了许多 OpenAI Universe 在 2016 年无法解决的问题。现在你可以生成代码,执行这些动作。但它们仍然不够鲁棒,无法处理高风险任务。推理有时还不够强,还有对齐问题——它们是否完全按照你的意愿行事?所以我认为这些是更好的语言模型能够解决的问题。
Yeah, I love this question. I've been talking about philosophy, but now it's time to ground ourselves a bit, like actually economically useful things. Love it. So I think of AI agent applications divided into three areas: one is software, and I think Ken can speak more to this—you're building AI agents around this space. Gaming is a second major application, and robotics basically—AI agents in the physical world. So let's talk about software. OpenAI Universe was one of the first attempts at doing this, having this digital agent, but using reinforcement learning wasn't the right approach. I do think today's large models solve many of the problems that OpenAI Universe couldn't solve back in 2016. Now you get code generation, you can execute those actions. But they are still not quite robust enough to actually do high-stakes tasks. The reasoning is still not strong enough sometimes, and also the alignment—are they exactly doing what you want? So I think these are some of the problems that better language models will be able to address.
Ken,你有什么想说的吗?我本来想问:你看到的 AI 智能体领域的问题中,哪些你觉得今天已经解决了,哪些仍然未解决?给在座正在构建智能体的创始人一些建议。
Do you have a comment, Ken? I was going to ask: what were the problems you saw in the world of AI agents that you feel like are actually solved today, and then what are the remaining unsolved things? For the founders building agents in the audience.
我认为语言模型实现的是零样本行为。这对于构建任何需要泛化到野外查询的东西都至关重要。基本上,你的客户来了,他们可以做任何事,你希望模型至少能合理地适应任何情况。现在,我提到的其他问题都还没有解决,但至少我们在进步。安全性、鲁棒性——如果你给它代码执行能力,安全性就成了大问题。你需要保证它不会意外删除你的数据库。你需要设置护栏,这些护栏不一定是神经网络,可能由传统软件强制执行。但我们需要重新思考围绕 LLM 的整个技术栈。有一个我很喜欢的类比,我借用一下……
So I think what language models have enabled are zero-shot behaviors. And that is critical for building anything that needs to generalize to in-the-wild queries. Basically, your customer comes in and they can literally do anything, and you want the model to adapt to anything, at least reasonably. Now, I think the rest of the things I mentioned are all not solved, but at least we're making progress. Safety, robustness—safety in terms of if you're giving it code execution capability, that safety becomes a big issue. You need guarantees that it wouldn't accidentally delete your database. You need to have guardrails around it, and these guardrails may not necessarily be a neural network; they could be enforced by traditional software. But we need to rethink the whole stack around LLMs. There's one analogy that I really like, and I borrow from...
Andre Kari 本人认为语言模型就是操作系统。我非常喜欢这个类比,我认为它在将语言模型应用于 AI 智能体时非常贴切。语言模型就是操作系统。为什么它们是操作系统?因为它们有一个全新的界面,这个界面就是语言。你可以以工具的形式安装软件,也就是工具使用。OpenAI 的 ChatGPT 应用商店就是一个语言模型使用工具的完美例子,对吧?它有硬盘,也就是向量数据库。现在你不再保存显式的东西,而是保存嵌入,之后再进行检索,对吧?这就是硬盘。而且还有不同的文件类型出现,这些就是多模态语言模型。你有图像、文本、音频、3D。多模态正在到来,它们就像不同的文件类型。同时你也会得到操作系统的缺点,也就是安全性。现在你会感染病毒,对吧?就像 Windows 有病毒一样,语言模型的病毒就是提示注入攻击,对吧?所有那些我们在 Twitter 上看到的奇怪又有趣的提示注入攻击。这些就是新型病毒。而且也有防御措施。有一个非常活跃的研究领域,比如对抗性提示防御,这是一个新兴领域,五年前我甚至不会想到它会存在。如果你五年前问我,什么是对抗性提示?这完全说不通。但现在它完全说得通,因为有了这个新的操作系统,对吧?所以问题是,现在语言模型操作系统还非常早期,因为你可以把它看作一台单线程计算机,可能只有 10 赫兹,对吧?它每秒生成 10 个 token,所以慢得离谱,而且无法并行化。但我认为这个类比在未来会非常有意义。
Andre Kari himself is that LMs are operating systems. I just love this analogy. I think it works super well in this application of L for AI agents is that LMs are operating systems. And why are they operating systems? They have a very new interface, and the interface is language. And you can install software in the form of tools, of tool use. And OpenAI's ChatGPT App Store is a perfect example of an LM using tools, right? It has hard disk in terms of vector databases. Now instead of saving explicit things, you save embeddings and later you retrieve, right? There's hard disk. And there's also different file types coming, and these are multimodal language models. You have images, text, audio, 3D. Multimodal is coming, and these are like correspond to different file types. And you also get the downsides of operating systems that are security. Now you get virus, right? Like Windows has virus, and the virus for LMs are prompt hacking, right? All those really kind of weird and funny prompt hacking we saw on Twitter. Like these are the new type of virus. And there are also defenses against there. There is a very active research field like adversarial prompt defense, and that is an emerging field that I wouldn't even think would happen five years ago. If you ask me five years ago, just what the hell is adversarial prompt? It's like doesn't make any sense. But now it makes total sense because of this new operating system, right? So the thing is, right now it's very early in LMs operating system because you can think of it as a single-threaded computer that's only maybe 10 Hertz, right? It's generating 10 tokens per second, so it's kind of ridiculously slow and not parallelizable. But I think this analogy is gonna go a very long way in the future.
我很好奇。在座的很多创始人都在构建智能体。如果你是一位投资人,我来向你推销说‘嘿,我是做智能体的创始人’,你会对什么持怀疑态度?比如,你觉得哪些问题五年内都解决不了,不应该围绕它来创业?
I'm really curious. So there are a lot of founders in the audience building agents. If you were an investor and I came and pitched you, 'Hey, I'm a founder building agents,' what would you be skeptical of? Like, you know, where you think, 'Ah, these problems aren't going to be solved for five years, you shouldn't be building a company around that.'
是的,我认为现在构建一个演示很容易。构建一个酷炫的演示异常简单,但构建一个稳健的产品却异常困难,对吧?比如演示,你只需精心挑选一个提示,然后运行几次,直到它产生一些酷炫的东西,然后在 Twitter 上展示,获得大量关注。但要为所有用例和客户稳健地构建它真的很难。所以我会问的问题是,它在实际使用中运行得有多稳健?不仅仅是那些你认为客户会用的用例,而是客户实际在提示框中输入的用例,对吧?他们输入了什么?嗯,这是一方面。另一方面是关于长尾问题,对吧?有很多客户带着特定需求而来,而且非常依赖具体情况。我们知道语言模型擅长处理主流用例,一些最常见的,但对于长尾问题,它们往往会失败。那么我们如何解决这些问题?也许使用检索。有很多方法。嗯,但我认为这些是应该关注的问题。
Yeah, I think it's easy to build a demo these days. It's like exceptionally easy to build a cool demo, exceptionally hard to make a robust product, right? Like a demo, you just pick cherry-pick one of those prompts and then you run it a couple times until it produces some cool stuff, and then you show it on Twitter and get a lot of attention. But it's really hard to build this robustly for all the use cases and customers. So I think the question I would ask is how robust it's working for in-the-wild uses, right? Not just the ones that you think your customer will use, but the use cases that customers actually are entering into the prompt box, right? Like what are they typing? Um, and that's one thing. And the other thing is about long tail, right? There are many kind of customers coming in with specific needs, and it's very much case by case. And we know that LMs are good at addressing kind of the mainstream use case, some of the most common ones, but for the long tails they kind of falter. So how do we address those problems? Maybe using retrieval. There are like many methods around it. Um, but I think those are the questions that one should pay attention to.
你认为哪些方法对稳健性最有前景?我有点把它看作,你知道,我们曾经发明了纠错,或者像我们发明了 TCP/IP 来处理电路和有线信道中的稳健性问题。这里的抽象是什么?
Are there methods that you think are most promising for robustness? I kind of think of it as, you know, we invented error correction at some point, or like we invented TCP/IP to deal with robustness issues in circuits and wire channels. What are the abstractions here?
一个是检索。如果你能够检索数据并进行引用,那肯定会让语言模型更符合事实,减少幻觉。另一个是围绕语言模型构建一个软件栈,对吧?如果你给语言模型一个代码解释器,那么你很可能需要让这个代码解释器安全,这样它就不会意外删除你的数据库。你需要一个专门为语言模型定制的解释器,对吧?所以你是在为这个新操作系统开发一个专门的程序。所以我认为所有这些都需要从头开始重新思考。
One is retrieval. So if you're able to kind of retrieve data and do citations, that will definitely make LMs more factual and hallucinate less. And the other is just having like a software stack around LMs, right? If you give an LM a code interpreter, then you may very well make this code interpreter secure so it doesn't delete your database accidentally. You need to have a custom interpreter just for the LM, right? So you're kind of developing a specialized program for this new operating system. So I think like all of these things need to be rethought from the ground up.
我想退一步,我们很快就要进入提问环节了。嗯,你和许多研究人员不同,你非常活跃,尤其是在 Twitter 上。嗯,这很令人惊讶。这并不……你拥有相当多的粉丝,比大多数人都多得多。我很好奇这是怎么发生的?你是如何变得直言不讳并积累了大量粉丝的?这对你的职业生涯有什么改变吗?
I want to take a step back and we're gonna have to go into questions soon. Um, you unlike many researchers, you're very out there, especially very much on Twitter. Um, that is surprising. That is not... you have quite the following, much much bigger following than most. I'm curious how did that happen? How did you become outspoken and amass this huge following? And also has it changed anything for you in your career?
是的,我喜欢分享想法。嗯,而且自从 ChatGPT 问世以来,人们对 AI 的兴趣激增,我突然感到震惊。所以我一直在从事 AI 工作,当 ChatGPT 出现时,我也尝试了一下。我当然印象深刻,但不像……我没有预见到 ChatGPT 引发的地震。嗯,然后社区里有很多噪音,对吧?你知道,各种酷炫的演示和过度宣传,非常多。所以我确实想帮助社区提高信噪比。嗯,我从中获得了很多乐趣和愉悦,对吧?穿透噪音,并试图加入我的观点,不仅仅是分享论文和方法,还加入我个人的看法。这并不总是正确的。我也一直在更新自己的信念,对吧?嗯,但我认为这个过程对我自己来说也是一个学习过程。
Yeah, so I like to share ideas. Um, and I think since ChatGPT hit the shelf and there was an explosion in the interest in AI, and I was suddenly stunned. So I've always been working in AI, and when ChatGPT came, I also tried it out. I was impressed for sure, but not like... I did not foresee the earthquake that ChatGPT resulted in. Um, and then there's a lot of noise in the community, right? You know, kind of cool demos and overclaims, just lots and lots of them. So I do want to kind of help the community increase the signal-to-noise ratio. And um, I get a lot of fun and pleasure doing this, right? Kind of cutting through the noise and trying to also put my perspective, not just sharing papers, sharing methods, but also putting my personal perspective on it. And it isn't always correct. I've always been updating my own belief as well, right? Um, but I think you know this process is a learning process for myself as well.
那么直言不讳、获得大量反馈、让很多人看到你的观点,这有没有改变你什么,比如你看待这个世界的方式、你的职业道路等等?
And has being outspoken, getting a lot of feedback, having a lot of people see what you have to say, has that changed anything for you, how you look at this world, your career path, anything like that?
这就是我今天在这里的原因。确实如此。非常感谢,Twitter。
That's the reason why I'm here today. That's true. Thank you so much, Twitter.
我不认为那是你在这里的原因。你在这里的原因,我认为是 Voyager 非常有趣。Jim 构建的 Voyager 是一个智能体,它通过构建自己的工具来解决 Minecraft 中的问题。嗯,我认为有一件事非常有趣……好吧,我不告诉你它有趣在哪里,但你对游戏非常感兴趣。你认为在游戏领域接下来最引人注目的事情是什么?如果你要在未来五年内在游戏领域做点什么。
I don't think that's the reason you're here. The reason you're here, I think, is Voyager is really interesting. So Voyager, which Jim built, is an agent that builds its own tools in order to solve Minecraft. Um, and I think one thing that's really interesting... well, I won't tell you what's interesting about it, but you're really interested in gaming. What do you think is the most compelling thing to do next in gaming? If you were to do something in gaming today over the next five years.
我非常喜欢这个问题。我提到了 AI 智能体的三个应用:游戏和机器人。嗯,我会快速概述一下我对这两个行业的看法。嗯,首先是游戏。嗯,游戏整个行业绝对巨大。实际上,按市值计算,游戏比整个好莱坞和整个音乐产业加起来还要大。游戏就是这么大。它实际上是娱乐行业中最大的一个领域,而且遥遥领先。所以有巨大的经济价值。现在,我认为游戏的发展方向,顺便说一下,在完成 Voyager 之后,Voyager 就像 AI 一样,所以我们有一个 AI 来玩 Minecraft,它能够连续玩这个游戏几个小时,并且自己发现新东西。所以那是……
I absolutely love this question. So I mentioned three applications of AI agents: gaming and robotics. Um, I will quickly kind of gloss over how I think of these two industries. Um, so first about gaming. Um, gaming is like the whole industry is absolutely insanely enormous. Actually, gaming in terms of market cap is bigger than the entire Hollywood and the entire music industry combined. That's how big gaming is. It is literally the top one biggest sector in the entertainment industry, and by far the biggest one. So huge economic value there. Now, where I think gaming is going, the kind of after doing Voyager, by the way, Voyager is like AI so we have an AI to play Minecraft, and it's able to play this game, play Minecraft for like hours on end and kind of discover new things all by itself. So that's a...
快速回顾一下 V 的功能。在 Voyager 之后,我意识到这就是游戏的发展方向:我们在开放世界中有智能 NPC,每次你与他们交谈时,他们不仅像基于文本的聊天那样回复,还会通过行动来回应。你与这些 NPC 的互动会影响他们的未来、心理状态以及他们在游戏中的行为。如果你将这种效果指数级放大,考虑到游戏中有许多 NPC,你可以想象一个游戏变得真正活生生的世界,就像《头号玩家》一样。我只是想看到《头号玩家》成为现实,而且我实际上认为这将在未来三到五年内发生。嗯,可能不是全身部分,因为那是 AR 和 VR,但只是关于智能 NPC 和一个有这么多 AI 相互交互的世界。每次你打开游戏,即使是同一个软件,你也会有不同的体验。每个游戏都变得无限可重玩和个性化,每次玩都是一次独特的体验。想想这类游戏会带来的参与度和它产生的涌现行为。
Quick recap of what V does. After Voyager, I realized that's where gaming is going: we have intelligent NPCs in open-ended worlds, and every time you talk to them, they not only reply as in text-based chat but also with actions. Your interaction with those NPCs affects their future, their mental states, and how they act in the game. If you propagate this effect exponentially, given you have many NPCs in the game, you can imagine a world where games become truly alive, something like Ready Player One. I just want to see Ready Player One becoming a reality, and I actually think that's going to happen in the next three to five years. Well, maybe not the full-body part because that's AR and VR, but just about the intelligent NPCs and a world where there are so many AIs interacting with each other. Every time you open a game, even though it's the same software, you have a different experience. Every game becomes infinitely replayable and personalized, and every time you play it, it's a unique experience. Just think about the engagement this type of game will draw and the emergent behaviors it will produce.
前几天我和朋友聊天,他在玩《艾尔登法环》。《艾尔登法环》是一个大作,而且是一个美丽的游戏,对吧?它是开放世界,有这么多 NPC,这么多不同的事情可以做,还有很多彩蛋。团队在《艾尔登法环》上投入的工作量令人难以置信。但你还是会感到无聊,对一些人来说可能两周,对另一些人可能两个月,但最终你会感到无聊,因为你已经用尽了所有 NPC,打败了所有 Boss,你回去时,你已经杀了 Boss,他们回应时仍然说“我要踢你的屁股”,但他们不记得你已经和他们互动过。游戏是死的。但想象一下,如果在像《艾尔登法环》、《我的世界》或《塞尔达传说》这样的游戏中,有智能 NPC,它们每次都能适应,有长期记忆,也许它们记得你曾经背叛过它们,现在有机会背叛你。想想这些游戏会有多复杂。
The other day I was talking to my friend, and he was playing Elden Ring. Elden Ring was a big title, and it's a beautiful game, right? It's open world, so many NPCs, so many different things you can do, and lots of Easter eggs. It's crazy the amount of work that the team put into Elden Ring. But you still get bored after maybe two weeks for some people, maybe two months, but eventually you get bored because you exhausted all the NPCs, you beat all the bosses, and you go back like you already killed the boss and they respond and still say 'I'm gonna kick your ass,' but they don't remember that you've already interacted with them. The game is dead. But just imagine if you have intelligent NPCs in games like Elden Ring or Minecraft or Legend of Zelda that adapt every time, that have long-term memory, that maybe they remember you betrayed them once and now it's time for an opportunity to betray you. Just think about how complex these games will be.
我想强调的另一项研究来自斯坦福,叫做 Stanford Smallville,他们在一个小世界中实例化了 25 个智能体,让它们在模拟中相互交互。这些智能体不知道自己在模拟中,所以有点像数字版的《西部世界》。智能体互动,它们一起工作,我记得其中两个甚至相爱了,计划了一个情人节派对。这是一项非常鼓舞人心的研究,我认为整个游戏领域都将朝着这个方向发展。但关于那篇论文的一点是,因为它基于 ChatGPT,ChatGPT 太礼貌了,所以看起来不那么有趣。里面有家庭,一个父亲和一个儿子,儿子说“你好吗,父亲?”父亲说“很好,你好吗,儿子?”“我很好,谢谢你的关心。”这就是 ChatGPT 做的。我们可以做得更好,而且我们需要做得更好。
Another research work I want to highlight is from Stanford called the Stanford Smallville, where they instantiate 25 agents in a little world and have them interact with each other in the simulation. The agents don't know they're in a simulation, so it's kind of like a digital Westworld. Agents interact, they go to work together, and two of them I remember even fell in love, planning a Valentine's party. It's a very inspiring research, and I think that's where the whole field of gaming will move towards. But the thing about that particular paper is because it's based on ChatGPT, ChatGPT is way too polite, so it's not that fun to watch. There are families in it, a father and a son, and the son says 'How are you doing, father?' and the father says 'Great, how are you doing, son?' 'I'm doing okay, thank you so much for asking.' That's what ChatGPT does. We can do better, and we need to do better.
就此,第一轮掌声送给 Jim。非常感谢你的分享。谢谢。现在我们将进入问答环节。有一个问题:AI 工程师需要哪些尚不存在的新工具或抽象?
On that note, first round of applause for Jim. Thank you so much for sharing. Thanks. And now we'll move over into Q&A. So one question came: what new tools or abstractions will AI engineers need that don't exist yet?
我认为有一些非常好的新工具,由 LangChain、LlamaIndex 等开源工具以及向量数据库提供支持。让我们回到操作系统类比,因为我认为这个问题与之非常契合。操作系统需要什么?它需要软件,需要与核心计算引擎交互的 API。它还需要更好的磁盘,所以可能在向量数据库和检索方面还有更多工作要做。此外,多模态正在到来,我认为拥有像素、图像和视频也将催生一些新工具,也许还有围绕它们的新界面。因此,一些关于 AI 原生用户界面的开源库也可能非常有用。
I think there are really good new tools enabled by LangChain, LlamaIndex, all of these open-source tools, and vector databases as well. Let's still go back to this OS analogy because I think this question ties very nicely to it. What does an OS need? It needs software, it needs APIs to interact with the core computing engine. It also needs better disk, so maybe there are more things to be done on the vector database and retrieval. Also, multimodal is coming, and I think having pixels, images, and videos will also enable some new tools, maybe new interfaces around them. So some open-source libraries on AI-native UIs may also be very useful.
明白了。Sioban,这里有人吗?你能拿到左边后面吗?你认为强化学习作为当今 AI 智能体和未来具身智能体的良好问题框架如何?你还提到了 Pedro Domingos,他对 RL 持批评态度,所以我很好奇优缺点,有什么细微差别吗?
Got it. Sioban, someone here? Can you grab back left? What do you think about reinforcement learning as a good problem framing for AI agents today and embodied agents moving forward? You mentioned Pedro Domingos as well, he's notably been critical of RL, so I'm curious about pros and cons, any nuances there.
这是一个很好的问题。问题是关于强化学习在具身智能体中的作用以及它未来是否有用。强化学习在语言模型中已经非常有用,比如基于人类反馈的强化学习(RLHF)。强化学习是一种将语言模型与人类意图对齐的方式,所以它已经很有用了。就具身智能体而言,我认为强化学习和语言模型是两种强大的工具,各自做不同的事情。最好用系统一与系统二思维来框架化,这来自《思考,快与慢》这本书。系统一是高层次的深思熟虑的规划和推理,系统二是快速自动的事情,你无需有意识思考,比如我如何控制手中的每个马达来举起这个瓶子,我如何做早餐或刷牙。这些是非常复杂的动作,但你不会在大脑中分配算力给它们。我认为语言模型非常适合系统一:深思熟虑,你可以进行推理、思维链和高层次规划。但强化学习对于低层次控制将非常有用,因为对于许多这样的控制,比如你如何控制手来抓取物体,你无法用语言表达。这是通过强化学习、试错来完成的,这是训练这些灵巧控制器的最佳方式。所以在机器人学中,我认为语言模型和强化学习的这种结合将是未来:高层次规划由语言模型完成,所有这些低层次技能由强化学习习得。
That's a great question. The question is about the role of reinforcement learning in embodied agents and whether it will be useful in the future. Reinforcement learning is already very useful in language models in terms of reinforcement learning from human feedback, RLHF. Reinforcement learning is a way to align language models towards human intentions, so it's already useful. In terms of embodied agents, I think reinforcement learning and language models are two powerful tools doing separate things. It's best framed in system one versus system two thinking, from the book Thinking Fast and Slow. System one is high-level deliberate planning and reasoning, and system two is rapid, automatic things you do without conscious thought, like how I control each motor in my hand to hold up this bottle, how I cook breakfast or brush my teeth. These are very complex motions but you don't allocate compute in your brain to them. I think language models are awesome for system one: deliberate, you can do reasoning, chain of thought, and high-level planning. But reinforcement learning will be very useful for low-level controls because for many of these, like how you control your hand to grasp objects, you cannot express that in language. It's done by reinforcement learning, trial and error, and that is the best way to train these dexterous controllers. So in robotics, I think this combination of language models and reinforcement learning will be the future: high-level planning done by a language model and all these low-level skills learned by reinforcement learning.
Rick Lamers,所以具身智能体,这实际上是上一个问题的延伸。我认为具身智能体通常通过强化学习训练。强化学习面临样本效率的挑战,因此可能需要很长时间才能收敛。有人可能会争辩说,人类由于出生时大脑的结构而带来了许多有用的先验。当你出生时,你带着一个大脑,有证据表明大脑带来了许多有用的先验,以加速学习过程,从“摇篮中的科学家”的角度来看。你对如何使用预训练和基础模型来改进有什么想法吗?
Rick Lamers, so embodied agents and this is actually an extension of the previous question. I think embodied agents are often trained on reinforcement learning. Reinforcement learning has the challenge of sample efficiency, so it can take a very long time to converge. One could argue humans bring many useful priors due to the structure in their brain at birth. When you are born, you are born with a brain, and there's some evidence that brains bring a lot of useful priors to accelerate the learning process from the 'scientist in a crib' perspective. Do you have any thoughts on how to use pre-training and foundational models as a way to improve?
样本效率和强化学习,这是个很技术性的问题,我喜欢。
Sample efficiency and reinforcement learning, that's a very technical question. I like it.
是的,让我快速解释一下。样本效率指的是强化学习智能体达到一定性能需要多少数据。这里的数据是由智能体自己生成的,因为它与环境交互,具身于世界中,通过交互收集数据。这就是样本效率的含义。是的,强化学习在样本效率方面出了名的差。举个例子,2019 年,OpenAI 训练了 OpenAI Five,它在 DOTA 中达到了世界冠军水平,几乎同时,DeepMind 制作了一个叫 AlphaStar 的智能体,在星际争霸中达到了世界冠军水平。这些东西的样本效率极低。你训练它们好几个月,基本上相当于数千年甚至数万年的经验,然后它们才能与人类匹敌。但人类显然没有一万条命,所以这是这些算法的一个问题。我认为强化学习应该只用于微调。这里我想引用一个比喻:Yann LeCun 提出了所谓的 LeCun 蛋糕。这是一个普通的蛋糕,但他贴上了标签。蛋糕的主体是无监督学习,糖霜是监督学习,顶上的樱桃是强化学习。所以无监督学习完成了大部分工作。这个术语字面意思是在没有监督、没有人工标注的情况下学习。监督学习是糖霜,即你有人工标签,并用它们来监督智能体。然后强化学习是自我生成的,与环境交互等。ChatGPT 就是这样训练的。ChatGPT 的大部分计算在无监督学习中完成,因为它只是预测下一个词,你不需要人类来预测下一个词,数据中已经存在。监督学习是监督微调:他们雇佣人类编写数据集,给定指令,这是合适的响应,然后在此基础上微调。然后你做强化学习:让模型自我生成,并使用人类反馈作为奖励信号。我相信同样的范式也适用于具身智能体。我们应该做大量的无监督学习,这些无监督学习可以来自视频、文本、预训练语言模型或预训练多模态模型。然后你做监督学习,即某种模仿学习。例如,对于机器人,我们会让人类远程操作机器人做某些事情,收集数据,然后在此基础上微调。最后一步是强化学习:让智能体在现实世界中自由活动,与环境交互,进行实验,保持好奇,基本上所有有趣的事情。所以 AI 将在蛋糕上的樱桃阶段开始享受乐趣。
Yeah, so let me quickly explain a little bit. Sample efficiency means how much data is needed for a reinforcement learning agent to reach a certain performance. The data here is generated by the agent itself because it's interacting, it's embodied in a world, interacting and collecting data. So that's what sample efficiency means. And yes, reinforcement learning is notoriously bad for sample efficiency. Just to give some perspective, back in 2019, OpenAI trained OpenAI Five that plays DOTA at world champion level, and almost concurrently, DeepMind made an agent called AlphaStar that plays StarCraft at world champion level. Those things are massively sample inefficient. You train them for many months, basically thousands if not tens of thousands of years of experience, and then they can match humans. But humans obviously don't have 10,000 lifetimes, so that is a problem in these algorithms. I think reinforcement learning should only be used for fine-tuning. Here I want to invoke a mental image: Yann LeCun proposed something called the LeCun cake. So it's a regular cake, but he put labels on it. The body of the cake is unsupervised learning, the frosting is supervised learning, and there's a cherry on top, which is reinforcement learning. So unsupervised learning does most of the work. This term literally means learning without supervision, without human manual annotations. Supervised learning is the frosting, where you have human labels and you use them to supervise the agent. Then reinforcement learning is self-generated, interacting with the world, etc. ChatGPT is trained this way. ChatGPT does most of the computation in unsupervised learning because it just predicts the next word, and you don't need a human to predict the next word, it's already there in the data. Supervised learning is supervised fine-tuning: they hire humans to write datasets, given this instruction, this is the appropriate response, and you fine-tune on top of that. Then you do reinforcement learning: you have the model self-generate and use human feedback as a reward signal. I believe that's the same paradigm going into embodied agents. We should do a lot of unsupervised learning, and this unsupervised learning can come from videos, text, pre-trained language models, or pre-trained multimodal models. Then you do supervised learning, which is some kind of imitation learning. For example, for robotics, we'll have a human teleoperate a robot to do certain things, collect that data, and fine-tune on top of that. The final step is reinforcement learning: set the agent free in the wild, have it interact with the world, do experimentation, be curious, basically all the fun stuff. So AI will start to have fun at the stage of the cherry on the cake.
谢谢,听你讲话很愉快。有什么是人们想得不够多,而你希望更多人思考或研究的问题或可能性吗?有什么没有得到足够关注的事情?
Thank you, it's been wonderful listening to you. What's something that people aren't thinking about enough that you wish more people were thinking about or working on, either a problem or a possibility? Something that hasn't gotten enough attention.
我认为实际上具身智能体没有得到足够的关注,这就是为什么我在谈论并倡导它。因为如果我们考虑世界上所有的算力,比如用于 AI 的 GPU 算力,我认为现在可能 99%都用于语言模型,而用于智能体、具身等方面的肯定远低于 1%。就这个比例而言,我认为研究领域应该花更多时间在那 1%上。因为我觉得那 1%实际上会弥合通往更高智能的差距。所以这是一件事。另一件我们开始做但整个行业做得还不够的事情是模拟。因为具身智能体需要一个世界来具身。当然你可以做物理机器人,让它们在物理世界中自由活动,但这可能很危险。当它们进行探索时,可能会伤害自己,甚至更糟,伤害物理世界中周围的人。而且,这非常慢。想想那些机器人移动得有多慢;它们要花很长时间才能学到任何有趣的东西。训练具身智能体更好的方法是在模拟中。而且模拟不能是随便的模拟;它们需要是逼真的或具有非常复杂事物的模拟。它应该是开放式的,支持无限可能的事情,并且非常非常快。我将快速提到 NVIDIA 的一个项目,我的一些研究依赖于它。在成为 AI 公司之前,NVIDIA 是一家图形公司,所以图形和模拟对公司来说非常重要。NVIDIA 有一个名为 Isaac Sim 的项目。Isaac Sim 是一个 GPU 加速的模拟引擎。你基本上可以在单个 GPU 上同时并行化 10,000 个环境,这意味着你可以将现实加速 10,000 倍。通过这种方式,你可以大规模扩展好奇智能体可以探索的数据量。回到关于样本效率的问题,这实际上是解决样本效率的另一种方法:让样本生成速度非常非常快。这是解决它的一种方法,因为如果你能将现实加速 10,000 倍,你就可以迭代更多,有更多数据可用,并且可以实验更多。所以我认为这是一个非常有前途的未来方向:拥有这些极度加速的模拟,并且还有逼真度、GPU 加速的光线追踪等,这样智能体可以在模拟中开发出逼真的视觉系统。我相信社区应该花更多时间并更加关注模拟和具身智能体。
I think actually embodied agents are not getting enough attention, that's why I'm talking about and advocating for it. Because if we think about all of the world's compute, like GPU compute for AI, I think maybe 99% of it is dedicated towards language models right now, and definitely much less than 1% is dedicated to agents, embodiment, all of that. In terms of this percentage, I think the research field should dedicate more time towards that 1%. Because I feel that 1% will actually bridge the gap towards higher levels of intelligence. So that is one thing. The other thing that we are starting to do but not doing enough as an entire industry is simulation. Because embodied agents need a world to be embodied in. Of course you can do physical robots and set them free in the physical world, but that could be dangerous. When they're doing exploration, they could hurt themselves, or even worse, hurt people around them in the physical world. Also, it's extremely slow. Just think about how slow those robots are moving; it's going to take forever for them to learn anything interesting. A much better way to train embodied agents is in simulation. And the simulations cannot be any simulation; they need to be simulations with, let's say, photorealism or with very complex things. It should be open-ended, supporting infinite things that you can do, and also be very, very fast. I'll quickly mention one effort at NVIDIA that I am relying on for some of my research. Before being an AI company, NVIDIA was a graphics company, so graphics and simulation are very dear to the company. NVIDIA has this project called Isaac Sim. Isaac Sim is a GPU-accelerated simulation engine. You can basically parallelize 10,000 environments on a single GPU at the same time, which means you can accelerate reality 10,000 times. In this way, you can massively scale up the amount of data exploration that a curious agent can do. Moving back to the question about sample efficiency, this is actually another way to solve sample efficiency: just make sample generation speed very, very fast. That is one way to solve it, because if you can accelerate reality 10,000 times, you can iterate a lot more, you have a lot more data to work with, and you can experiment a lot more. So I think this is a very promising direction moving to the future: having these extremely accelerated simulations, and also having photorealism, GPU-accelerated ray tracing, all of that, so the agent can have realistic vision systems developed in the simulations. I believe the community should spend more time and pay more attention to simulations and embodied agents.
好的,我们最后问两个问题。感谢你精彩的分享,也祝贺你获得基金。别用那个词。我有一个与之前回答非常相关的问题:从现在起两三年后,你对于如何将具身智能体扩展到互联网世界(当今大多数商业和人类活动所在的地方)有什么愿景?
Alright, we'll do two last questions. Thank you for the wonderful panel and also congratulations on the fund. Don't use that word. I have a very related question to the previous answer: three years from now, two years from now, what's your vision in terms of how do you scale up embodied agents into, let's say, the internet world, where most of the commerce, most of human activity today lives?
谢谢。那么我认为那将是数字智能体,因为它不再是具身的了,对吧?互联网更像是一个数字世界。所以我们需要能够在该数字世界中运行的智能体,与网站、API 等进行交互。这是一种不同的具身形式,但仍然是一个智能体。我认为同样的原则适用:从大量数字数据中进行无监督学习,从人类任务演示中进行监督学习,以及用于探索和优化的强化学习。关键是构建数字世界的模拟器,比如网络环境,智能体可以在其中安全地大规模学习。所以我的愿景是,我们将看到物理模拟中的具身智能体和网络模拟中的数字智能体融合,所有这些都由相同的基础技术驱动。
Thanks. So then I think that would be like digital agents, because it's not embodied anymore, right? The internet is more like a digital world. So we need agents that can operate in that digital world, interacting with websites, APIs, and so on. That's a different kind of embodiment, but still an agent. I think the same principles apply: unsupervised learning from large amounts of digital data, supervised learning from human demonstrations of tasks, and reinforcement learning for exploration and optimization. The key is to build simulators of the digital world, like web environments, where agents can learn safely and at scale. So my vision is that we will see a convergence of embodied agents in physical simulations and digital agents in web simulations, all powered by the same underlying techniques.
比如说,用键盘和鼠标控制之类的,做一个互联网智能体。是的,所以我认为我提到的许多原则仍然适用:首先做无监督学习,再做一点监督学习,最后在上面加一点强化学习。但一个危险是,如果你让智能体在互联网上自由活动,又会出现安全问题。它可能会做危险的事情,甚至可能违法,我不知道,这是有可能的。所以我们需要给它设置护栏,就像让物理机器人在物理空间中探索一样,你必须要有这些护栏。但除此之外,许多原则仍然适用。
Talking about, let's say, keyboard and mouse control, something like that, to be like the internet agent. Yeah, so I think many of the principles I mentioned still apply: first you do unsupervised learning, you do a bit of supervised learning, and finally have a little bit of reinforcement learning on top. But one danger is, if you set the agent free on the internet, again there are safety issues. It might do things that are dangerous, that might be illegal, I don't know, it's possible. So we need to put guardrails around it, just like having a physical robot exploring in the physical space, you've got to have these guardrails. But otherwise, many of these principles apply.
关于在互联网上与人类互动,我认为这是一个有趣的话题。我还没有看到大规模与人类互动的 AI 智能体,但我想象可能会有一些涌现行为,不仅是在智能体这边,也在人类这边。当我们提示 ChatGPT 时,ChatGPT 也在提示我们。当 ChatGPT 刚出来时,我们在 Twitter 和 Reddit 上看到了很多非常有创意的提示。我认为 ChatGPT 在某种程度上也激发了人类的创造力。我确实认为,让互联网智能体与人类互动会在双方都引发有趣的涌现行为。
And regarding interacting with humans on the internet, I think it is an interesting topic. I have not seen an AI agent interacting with humans at scale, but I would imagine there could be some emergent behaviors, not just on the agent side but also on the human side. When we are prompting ChatGPT, ChatGPT is also prompting us. When ChatGPT first came out, we saw a lot of very creative prompts shared on Twitter and Reddit. I think ChatGPT also sparked our human creativity in some ways. And I do think having internet agents interacting with humans will induce interesting emergent behaviors on both sides.
明白了。好,最后一个问题:你个人如何决定当前和未来机器学习研究中什么是重要的?你如何判断 Twitter 和其他地方哪些是信号、哪些是噪音?
Got it. All right, one last question: how do you personally decide what's important in present and future ML research? How do you decide what's signal and what's noise on Twitter and elsewhere?
我来谈谈我如何决定自己的研究议程。我经历过博士阶段,博士非常痛苦——大量的痛苦和折磨,不管你在哪里读,都是痛苦和折磨。顺便说一句,PhD 代表永久性脑损伤。我认为,通过博士五年的痛苦和折磨,我学到了很多工程和实际研究技能等等。但我不认为那是我从博士中获得的最重要的东西。我认为最重要的是我们所说的研究品味:基本上你不仅学会了如何做事,还学会了如何选择哪些事情值得做。我认为'做什么'实际上比'怎么做'更难。这就是为什么 CEO 有时比工程师更难当——你需要决定下一步做什么,你需要优先排序,这需要大量经验,而且很难从教科书中学到,因为每个人都不一样。你需要经历这些痛苦和折磨才能达到那个境界。
I'll talk about how I decide my own research agenda. I have been through a PhD, and PhD is extremely painful—a lot of pain and suffering, doesn't matter where you do it, it's just pain and suffering. PhD by the way stands for permanent head damage. And I think through this pain and suffering over the five years of PhD, I learned a lot about engineering and actual research skills and all that. But I don't think that is the most important thing I took away from a PhD. I think the most important thing is what we call research taste: basically you not only learn how to do things but also learn how to pick what things are worth doing. And I think the 'what' is actually harder than the 'how'. That's why being a CEO is sometimes harder than being an engineer—you need to decide what to do next, you need to prioritize, and that takes a lot of experience and it's really hard to learn from a textbook because everyone's different. You need to go through this pain and suffering to get there.
现在在 AI 领域,我的研究品味,由于我参与的所有事件——在 BYU 与 Dario 合作,在 OpenAI 工作,看到 Scaling(规模扩张)的推进,还有模拟和视频模拟的 Scaling(规模扩张),经历了所有这些——我的研究品味倾向于简单而优雅的解决方案。这其实一点也不明显,至少在我读博期间不是这样。当我读博时,我自己写论文,也看到很多同龄人写非常复杂的论文,有 20 个模块相互作用来解决一个非常具体的问题。我们这样做是因为你可以在基准测试上得到更高的分数。学术界的同行评审几乎都是关于基准测试:你是否在图像基准或分割基准上比之前的方法提高了 1%的准确率?所以我们本能地追求复杂性。我花了很长时间来摆脱这种习惯,并培养出对简单优雅事物的研究品味。
Now in AI, my research taste, because of all the events I was in—working with Dario at BYU, working at OpenAI, seeing scaling up, also simulation and video simulation scaling up, going through all of this—my research taste is towards simple and elegant solutions. And that is actually not at all obvious, at least not during my PhD time. When I was doing my PhD, I myself wrote papers and also saw lots of my peers writing papers that were very complicated, with 20 modules interacting to deliver a solution to a very specific problem. The reason we do this is because you can get a higher number on a benchmark. Peer review in academia is all about benchmarking: have you beaten the prior method by 1% accuracy on an image benchmark or segmentation benchmark? So we reach for complexity as the first instinct. It took me a very long time to unlearn that and to develop a research taste towards simple and elegant things.
我认为我从 FA 以及观察 OpenAI 如何构建他们的系统中学到了很多。我仍然非常清楚地记得 CLIP 和 DALL·E 首次发布的那一天——DALL·E 1 和 CLIP。OpenAI 在同一天发布了两个模型。CLIP 是一个视觉语言模型,学习语言和图像之间的关联,DALL·E 从语言生成图像。他们在同一天发布了这两个模型,那一天改变了整个视觉语言模型研究领域。视觉语言模型领域曾经极其复杂——复杂的方法主导了该领域。你必须使用复杂的语言编码器,以不同方式分割图像,并弄清楚语言和图像如何交互。想想有多少设计点,每个点都可以成就一个博士生涯。但 OpenAI 的做法极其优雅。对于 DALL·E,你只需要大量的互联网文本和图像,对图像进行分词,把文本和图像变成一个巨大的序列,然后预测下一个词元。它如此简单、如此优雅,让我震惊。我当时想,'哇,你可以用这种方式做视觉语言系统?'我简直不敢相信,但它效果非常好,你只需要 Scaling(规模扩张)它。简单优雅的解决方案还有一个好处,就是它们具有可扩展性,因为你不需要大量的工程流水线和复杂的东西。你喂给它大量数据,你在互联网上能找到的所有数据都与这个简单而富有表现力的接口兼容:输入文本和图像词元,输出文本和图像词元。一切都可以用这种方式表达。你不需要针对特定类型的图像设计特定的流水线。所以这只是我如何培养这种研究品味的一个例子,今天我仍在追求同样的事情。
I think I took a lot from FA and also from observing how OpenAI builds their systems. I still remember very clearly the day when CLIP and DALL·E first came out—DALL·E 1 and CLIP. OpenAI released two models on that day. CLIP is a vision-language model that learns the association between language and image, and DALL·E generates images from language. They dropped both models on the same day, and that day changed the entire research field of vision-language models. Vision-language models as a field was extremely complicated—complex methods dominated the field. You had to go through complex language encoders, segment the image in different ways, and figure out how language and image interact. Just think about how many design points there are, and each point could make a PhD career. But what OpenAI did was extremely elegant. For DALL·E, you just have lots of internet text and images, you tokenize the image, you make text and image a big fat sequence, and you predict the next token. It's so simple and so elegant, it blew my mind. I was like, 'Wow, you can do vision-language systems in this way?' I did not believe it, but it works so well, and you just scale it up. What's also great about simple and elegant solutions is that they are scalable, because you don't need a lot of engineering pipelines and complex things. You feed it a lot of data, and all the data you can find on the internet are compatible with this simple and expressive interface: text and image tokens in, text and image tokens out. Everything can be expressed that way. You don't need specific pipelines for specific types of images. So this is just one example of how I developed my research taste towards that, and today I'm still pursuing the same thing.
我们能否有一天拥有一个单一的智能体——我称之为基础智能体——它可以泛化到不同的技能、不同的任务,甚至不同的现实?无论是游戏还是机器人,机器人是在物理世界中。我不会为这个特定的游戏或这个特定的物理世界和任务设计特定的方法。我不会为此设计模块。我只有一个巨大的 Transformer 或其他什么,只有一个跨现实工作的流水线。然后基本上每个任务和智能体都只是对同一个基础智能体的不同提示。这就是我正在努力的方向。我认为我从痛苦和折磨中获得了这种研究品味,但它引导我构建这些优雅的系统。
Can we have, someday, a single agent—I'm calling it a foundation agent—that can generalize to different skills, different tasks, and even different realities? Be it games or robotics, robotics is in the physical world. I don't design specific methods for this particular game or this particular physical world and task. I don't design modules for that. I just have a single big fat Transformer or whatever, just one pipeline that works across realities. And then basically every task and agent is just different prompts to the same foundation agent. That is what I'm working towards. And I think I got that research taste from the pain and suffering, but it guides me towards building these elegant systems.
太棒了。好的,就此,让我们为 Jim 热烈鼓掌。
Amazing. All right, on that note, huge round of applause for Jim.