Jobs, robots, and why world models are next
打开互动全文版(中英对照 + 朗读 + 问答)→「AI 教母」谈空间智能,以及以人为本的未来。
The godmother of AI on spatial intelligence and a human-centered future.
今天的嘉宾是李飞飞博士,她被誉为 AI 教母。飞飞一直处于许多重大突破的核心,这些突破引发了当前我们正在经历的 AI 革命。她主导创建了 ImageNet,这基本上是她意识到 AI 需要大量干净标注数据才能变得更聪明。那个数据集成为了突破,引领了当前构建和扩展 AI 模型的方法。她曾是谷歌云的首席 AI 科学家,许多早期重大技术突破都出自那里。她是斯坦福人工智能实验室(SAIL)的主任,许多顶尖 AI 人才都来自那里。她还是斯坦福以人为本 AI 研究所的联合创始人,该所在 AI 发展方向上扮演着关键角色。她曾担任推特董事会成员。她被《时代》杂志评为 AI 领域最具影响力的 100 人之一。她也是联合国顾问委员会的成员。我还可以继续说。在我们的对话中,飞飞分享了 AI 如何发展到今天的一段简史,包括一个令人震惊的提醒:9 到 10 年前,称自己为 AI 公司基本上是在给品牌敲丧钟,因为没人相信 AI 真的能行。今天完全不同了,每家公司都是 AI 公司。我们还聊了她如何看待 AI 未来对人类的影响,当前技术能带我们走多远,她为什么对构建世界模型如此热情,以及世界模型到底是什么。最令人兴奋的是,世界上第一个大型世界模型 Marble 刚刚发布,就在这期播客上线的时候。任何人都可以在 marble.worldlabs.ai 上体验它。太疯狂了,一定要去看看。飞飞非常了不起,但她的影响力却远远被低估了。所以,我非常高兴能邀请她来,把她的智慧传播给更多人。非常感谢 Ben Horowitz 和 Condoleezza Rice 为这次对话建议话题。如果你喜欢这期播客,别忘了在你最喜欢的播客应用或 YouTube 上订阅和关注。接下来,在简短赞助商广告之后,有请李飞飞博士。
Today, my guest is Dr. Fei-Fei Li, who's known as the godmother of AI. Fei-Fei has been responsible for and at the center of many of the biggest breakthroughs that sparked the AI revolution that we are currently living through. She spearheaded the creation of ImageNet, which was basically her realizing that AI needed a ton of clean labeled data to get smarter. And that data set became the breakthrough that led to the current approach to building and scaling AI models. She was chief AI scientist at Google Cloud, which is where some of the biggest early technology breakthroughs emerged from. She was director at SAIL, Stanford's artificial intelligence lab, where many of the biggest AI minds came out of. She's also co-creator of Stanford's Human-Centered AI Institute, which is playing a vital role in the direction that AI is taking. She's also been on the board of Twitter. She was named one of Time's 100 most influential people in AI. She's also on the United Nations Advisory Board. I could go on. In our conversation, Fei-Fei shares a brief history of how we got to today in the world of AI, including this mind-blowing reminder that 9 to 10 years ago, calling yourself an AI company was basically a death knell for your brand, because no one believed that AI was actually going to work. Today, it's completely different. Every company is an AI company. We also chat about her take on how she sees AI impacting humanity in the future, how far current technologies will take us, why she's so passionate about building a world model, and what exactly world models are. And most exciting of all, the launch of the world's first large world model, Marble, which just came out as this podcast comes out. Anyone can go play with this at marble.worldlabs.ai. It's insane. Definitely check it out. Fei-Fei is incredible and way too under the radar for the impact that she's had on the world. So, I am really excited to have her on and to spread her wisdom with more people. A huge thank you to Ben Horowitz and Condoleezza Rice for suggesting topics for this conversation. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. With that, I bring you Dr. Fei-Fei Li after a short word from our sponsors.
飞飞,非常感谢你来到这里,欢迎来到播客。
Fei-Fei, thank you so much for being here and welcome to the podcast.
我很兴奋能来这里,Lenny。
I'm excited to be here, Lenny.
我更兴奋能邀请到你。能和你聊天真是一种享受。我想聊的太多了。你一直处于当前 AI 爆发的中心,已经很久了。我们会聊很多历史,我觉得很多人甚至不知道这一切是怎么开始的。但让我先读一段 Wyatt 关于你的话,让大家有个概念,在开场白里我会分享你做的所有其他了不起的事情,但我觉得这是设定背景的好方法。飞飞是少数科学家之一,这个群体小到可能围坐在一张厨房桌子旁,他们负责了 AI 最近显著的进步。很多人称你为 AI 教母。与许多 AI 领袖不同,你是一个 AI 乐观主义者。你不认为 AI 会取代我们,不认为它会夺走我们所有的工作,不认为它会杀死我们。所以,我觉得从这里开始会很有趣。你的观点是什么?AI 将如何随着时间的推移影响人类?
I'm even more excited to have you here. It is such a treat to get to chat with you. There's so much that I want to talk about. You've been at the center of this AI explosion that we're seeing right now for so long. We're going to talk about a bunch of the history that I think a lot of people don't even know about how this whole thing started. But let me first read a quote from Wyatt about you just so people get a sense and in the intro I'll share all of the other epic things you've done but I think this is a good way to just set context. Fei-Fei is one of a tiny group of scientists, a group perhaps small enough to fit around a kitchen table, who are responsible for AI's recent remarkable advances. A lot of people call you the godmother of AI. And unlike a lot of AI leaders, you're an AI optimist. You don't think AI is going to replace us. You don't think it's going to take all our jobs. You don't think it's going to kill us. So, I thought it'd be fun to start there. Just what's your perspective on how AI is going to impact humanity over time?
好的。那么,Lenny,让我说清楚。
Yeah. Okay. So, Lenny, let me be very clear.
很多人称你为 AI 教母。你的工作实际上是让我们走出 AI 冬天的火花。在 2015 年中到 2016 年中,一些科技公司避免使用 AI 这个词,因为他们不确定 AI 是不是一个脏词。大约 2017 年,公司开始称自己为 AI 公司。我记得你在国会演讲时说过一句话:AI 没有什么人工的。它受人类启发,由人类创造,最重要的是,它影响人类。
A lot of people call you the godmother of AI. The work you did actually was the spark that brought us out of AI winter. In the middle of 2015, middle of 2016, some tech companies avoided using the word AI because they were not sure if AI was a dirty word. 2017ish was the beginning of companies calling themselves AI companies. There's this line, I think this was when you were presenting to Congress, there's nothing artificial about AI. It's inspired by people. It's created by people. And most importantly, it impacts people.
我并不是认为 AI 对工作或人类没有影响。事实上,我相信 AI 现在或将来做什么都取决于我们,取决于人类。我确实相信技术对人类是净正面的。但我认为每项技术都是一把双刃剑。如果我们作为社会、作为个人不做好事,我们也会搞砸。
It's not like I think AI will have no impact on jobs or people. In fact, I believe that whatever AI does currently or in the future is up to us. It's up to the people. I do believe technology is a net positive for humanity. But I think every technology is a double-edged sword. If we're not doing the right thing as a society, as individuals, we can screw this up as well.
你有一个突破性的洞察:我们可以训练机器像人类一样思考,但它缺少人类作为孩子学习时所拥有的数据。
You had this breakthrough insight of just okay we can train machines to think like humans but it's just missing the data that humans have to learn as a child.
我选择通过视觉智能的视角来看待人工智能,因为人类是深度视觉动物。我们需要用尽可能多的物体图像信息来训练机器。但物体非常非常难学。一个单一的物体在图像中可以有无穷的可能性。为了用成千上万个物体概念训练计算机,你真的需要向它展示数百万个例子。
I chose to look at artificial intelligence through the lens of visual intelligence because humans are deeply visual animals. We need to train machines with as much information as possible on images of objects. But objects are very, very difficult to learn. A single object can have infinite possibilities that are shown on an image. In order to train computers with tens and thousands of object concepts, you really need to show it millions of examples.
我不是乌托邦主义者。所以,我并不认为 AI 不会对就业或人们产生影响。事实上,我是一个人本主义者。我相信 AI 现在或将来做什么,取决于我们,取决于人。所以我确实相信技术对人类是净正面的。如果你看文明的漫长历程,我认为我们从根本上是一个创新的物种。从几千年前的文字记录到现在,人类一直在创新自我和创新工具,并以此让生活更美好,让工作更美好,建设文明,我相信 AI 是其中的一部分。这就是乐观的来源。但我认为每一项技术都是一把双刃剑。如果我们作为物种、社会、社区、个人没有做正确的事情,我们也可能搞砸。
I'm not a utopian. So, it's not like I think AI will have no impact on jobs or people. In fact, I'm a humanist. I believe that whatever AI does currently or in the future is up to us. It's up to the people. So I do believe technology is a net positive for humanity. If you look at the long course of civilization, I think we are fundamentally an innovative species. From the written record thousands of years ago to now, humans just kept innovating ourselves and innovating our tools, and with that we make lives better, we make work better, we build civilization, and I do believe AI is part of that. So that's where the optimism comes from. But I think every technology is a double-edged sword. And if we're not doing the right thing as a species, as a society, as communities, as individuals, we can screw this up as well.
我记得你在国会演讲时说过一句话:「AI 中没有任何人工的东西。它受人类启发,由人类创造,最重要的是它影响人类。」我这不是问题,但这句话说得真好。
There's this line I think this was when you were presenting to Congress. 'There's nothing artificial about AI. It's inspired by people. It's created by people and most importantly it impacts people.' I don't have a question there but what a great line.
是的,我对此感受很深。我 25 年前开始研究 AI,过去二十年一直在带学生。几乎每个毕业的学生,我都会提醒他们,你们这个领域叫做人工智能,但其中没有任何人工的东西。
Yeah, I feel pretty deeply. I started working in AI two and a half decades ago and I've been having students for the past two decades. Almost every student who graduates, I remind them when they graduate from my lab that your field is called artificial intelligence but there's nothing artificial about it.
回到你刚才说的,这一切走向何方取决于我们。你认为我们需要做对什么?如何让事情走上正轨?我知道这个问题很难回答,但你的建议是什么?
Coming back to the point you just made about how it's kind of up to us about where this all goes. What is it you think we need to get right? How do we set things on a path? I know this is a very difficult question to answer but just what's your advice?
嗯。
Yeah.
我们有多少小时?
How many hours do we have?
我们如何对齐 AI?就这样吧。我们来解决它。
How do we align AI? There we go. Let's solve it.
另外,我认为无论我们做什么,人们都应该是负责任的个体。这是我们教给孩子的,也是我们作为成年人需要做的。无论你参与 AI 开发、部署还是应用的哪个环节,而且很可能我们许多人,尤其是技术人员,都处于多个环节,我们都应该像负责任的个体一样行事,并关心这件事,实际上非常关心。我认为今天每个人都应该关心 AI,因为它将影响你的个人生活、你的社区、社会和未来一代。作为一个负责任的人去关心它,是第一步,也是最重要的一步。
Also, I think people should be responsible individuals no matter what we do. This is what we teach our children and this is what we need to do as grown-ups as well. No matter which part of the AI development or AI deployment or AI application you are participating in, and most likely many of us especially as technologists are in multiple points, we should act like responsible individuals and care about this, actually care a lot about this. I think everybody today should care about AI because it is going to impact your individual life, your community, the society, and the future generation. And caring about it as a responsible person is the first but also the most important step.
好的。那么,让我退一步,回到 AI 的起点。大多数人开始听说并关心 AI,就是今天所谓的 AI。就像几年前 ChatGPT 出现时一样。大概是三年前吧。
Okay. So, let me actually take a step back and kind of go to the beginning of AI. Most people started hearing and caring about AI is what it's called today. Just like I don't know a few years ago when ChatGPT came out. Maybe it was like three years ago.
三年前。差不多再一个月就满三年了。三年前。
Three years ago. Almost one more month. Three years ago.
哇。好的。那就是 ChatGPT 问世。那是你心目中的里程碑吗?好的。酷。这正是我所想的。但很少有人知道,人们为此工作了很长一段时间。那时它被称为机器学习,还有其他术语,现在一切都成了 AI。有一段很长的时期,很多人在研究它,然后出现了所谓的 AI 寒冬,人们几乎放弃了,认为这个想法行不通。然后你做的研究实际上是把我们带出 AI 寒冬的火花,直接导致了我们现在所处的世界——AI 成了我们谈论的一切。正如你所说,它将影响我们所做的一切。所以,我觉得听你讲讲 ImageNet 之前的世界是什么样的,你创建 ImageNet 的工作,为什么它如此重要,以及之后发生了什么,会非常有趣。
Wow. Okay. That was ChatGPT coming out. Is that the milestone that you have in mind? Okay. Cool. That's exactly how I saw it. But very few people know there was a long history of people working on it. It was called machine learning back then and there's other terms, and now it's just everything's AI. And there was kind of like a long period of just a lot of people working on it, and then there's this what people refer to as the AI winter where people just gave up, almost people did, and just okay this idea isn't going anywhere. And then the work you did actually was essentially the spark that brought us out of AI winter and is directly responsible for the world we're in now of just AI is all we talk about. As you just said, it's going to impact everything we do. So, I thought it'd be really interesting to hear from you just kind of like the brief history of what the world was like before ImageNet, then just the work you did to create ImageNet, why that was so important, and then just what happened after.
对我来说,很难记住 AI 对每个人来说都是如此新鲜。当我整个职业生涯都生活在 AI 中时,我内心有一部分感到非常满足,看到一种个人好奇心——我几乎刚出青少年时期就开始的——现在已经成为我们文明的变革力量。它基本上是一种文明级别的技术。所以这段旅程大约 30 年或 20 多年,非常令人满足。那么,这一切从哪里开始呢?我甚至不是第一代 AI 研究者。第一代真正可以追溯到 50 年代和 60 年代。艾伦·图灵在 40 年代就超前地提出了一个大胆的问题:我们能有会思考的机器吗?当然,他有一种特定的方式来测试这个思考机器的概念,即一个对话聊天机器人,按照他的标准,我们现在已经有了思考机器。但这只是一个轶事般的灵感。这个领域真正始于 50 年代,当时计算机科学家聚集在一起,研究如何利用计算机程序和算法来构建能够完成只有人类认知才能完成的事情的程序。这就是开始。创始人们,1956 年的达特茅斯研讨会,我们有约翰·麦卡锡教授,他后来来到斯坦福,创造了「人工智能」这个术语。在 50 年代、60 年代、70 年代和 80 年代之间,是 AI 探索的早期。我们有逻辑系统、专家系统,还有早期对神经网络的探索。然后到了 80 年代末、90 年代和 21 世纪初。那大约 20 年的时间实际上是机器学习的开始。它是计算机编程和统计学习的结合。这种结合给 AI 带来了一个非常关键的概念:纯粹基于规则的程序无法涵盖我们想象中计算机能够完成的巨大认知能力。所以我们必须让机器来学习模式。一旦机器能够学习模式,它就有希望做更多事情。例如,如果你给它三只猫,希望不仅仅是机器能识别这三只猫,而是机器能识别第四只、第五只、第六只以及所有其他的猫。这是一种学习能力,对人类和许多动物来说都是基础。我们作为这个领域意识到我们需要机器学习。所以一直到 21 世纪初。我进入 AI 领域是在 2000 年,那是我在加州理工学院开始博士研究的时候。所以我是第一代机器学习研究者之一,我们已经在研究机器学习的概念,尤其是神经网络。
It is for me hard to keep in mind that AI is so new for everybody. When I lived my entire professional life in AI, there's a part of me that is just so satisfying to see a personal curiosity that I started barely out of teenagehood and now has become a transformative force of our civilization. It generally is a civilizational level technology. So that journey is about 30 years or 20 something, 20 plus years, and it's just very satisfying. So, where did it all start? Well, I'm not even the first generation AI researcher. The first generation really dates back to the 50s and 60s. And Alan Turing was ahead of his time in the 40s by asking the daring question: can we have thinking machines? And of course he had a specific way of testing this concept of a thinking machine, which is a conversational chatbot, which to his standard we now have a thinking machine. But that was just an anecdotal inspiration. The field really began in the 50s when computer scientists came together and looked at how we can use computer programs and algorithms to build programs that can do things that have only been capable by human cognition. And that was the beginning. The founding fathers, the Dartmouth workshop in 1956, we have Professor John McCarthy who later came to Stanford, who coined the term artificial intelligence. And between the 50s, 60s, 70s, and 80s, it was the early days of AI exploration. We had logic systems, we had expert systems. We also had early exploration of neural networks. And then it came to around the late 80s, the 90s, and the very beginning of the 21st century. That stretch of about 20 years is actually the beginning of machine learning. It's the marriage between computer programming and statistical learning. And that marriage brought a very critical concept into AI: that purely rule-based programs are not going to account for the vast amount of cognitive capabilities that we imagine computers can do. So we have to use machines to learn the patterns. Once the machines can learn the patterns, it has a hope to do more things. For example, if you give it three cats, the hope is not just for the machines to recognize these three cats. The hope is the machines can recognize the fourth cat, the fifth cat, the sixth cat, and all the other cats. And that's a learning ability that is fundamental to humans and many animals. And we as a field realized we need machine learning. So that was up till the beginning of the 21st century. I entered the field of AI literally in the year of 2000. That's when my PhD began at Caltech. And so I was one of the first generation machine learning researchers and we were already studying this concept of machine learning, especially neural networks.
我记得那是我在加州理工学院最早的一门课之一,叫做神经网络,但那门课非常痛苦。当时仍处于所谓的 AI 寒冬之中,公众对此关注不多,资金也不充足,但很多想法在流动。我认为有两件事让我的职业生涯如此接近现代 AI 的诞生。第一是我选择通过视觉智能的视角来研究人工智能,因为人类是深度视觉动物。我们可以稍后再谈,但我们的智能很大一部分建立在视觉、感知和空间理解之上,而不仅仅是语言本身。我认为它们是互补的。所以我选择了研究视觉智能。在我的博士和早期教授生涯中,我和我的学生非常致力于一个北极星问题:解决物体识别问题,因为它是感知世界的基石。对吧?我们在世界中解读、推理和互动,基本上是在物体层面。我们不会在分子层面与世界互动。我们很少,比如,如果你想拿起一个茶壶,你不会说「好吧,这个茶壶由一百块瓷器组成,让我来处理这一百块。」你会把它看作一个物体并与之互动。所以物体非常重要。因此,我是最早将这个问题确定为北极星问题的研究人员之一。
I remember that was one of my first courses at Caltech, called neural networks, but it was very painful. It was still smack in the middle of the so-called AI winter, meaning the public didn't look at this too much. There wasn't that much funding, but there was also a lot of ideas flowing around. And I think two things happened to myself that brought my own career so close to the birth of modern AI. One is that I chose to look at artificial intelligence through the lens of visual intelligence, because humans are deeply visual animals. We can talk a little more later, but so much of our intelligence is built upon visual, perceptual, spatial understanding, not just language per se. I think they're complementary. So I chose to look at visual intelligence. And my PhD and my early professor years, my students and I were very committed to a north star problem: solving the problem of object recognition, because it's a building block for the perceptual world. Right? We go around the world interpreting, reasoning, and interacting with it more or less at the object level. We don't interact with the world at the molecular level. We don't interact with the world as we sometimes do, but we rarely, for example, if you want to lift a teapot, you don't say, 'Okay, the teapot is made of a hundred pieces of porcelain, let me work on these hundred pieces.' You look at it as one object and interact with it. So object is really important. So I was among the first researchers to identify this as a north star problem.
但我觉得,作为 AI 的学生和后来的研究者,我一直在研究各种数学模型,包括神经网络、贝叶斯网络等等。有一个痛点:这些模型没有数据可以训练。整个领域都过于关注模型,但我意识到人类学习和进化实际上是一个大数据学习过程。人类通过不断积累经验来学习。而进化,如果你看时间尺度,动物通过体验世界而进化。所以我和我的学生推测,一个被严重忽视的让 AI 活起来的关键因素是数据。于是我们在 2006、2007 年开始了 ImageNet 项目。我们非常雄心勃勃:想要获取互联网上所有物体的图像数据。当然,当时的互联网比现在小得多,所以我觉得这个野心至少不算太疯狂。现在想想,几个研究生和一个教授做这件事简直是妄想。但我们做到了。我们精心整理了互联网上的 1500 万张图像,创建了一个包含 22000 个概念的分类体系,借鉴了其他研究者的工作,比如语言学家在 WordNet 上的工作,这是一种特殊的词典编纂方式,我们将其与 ImageNet 结合。我们将其开源给研究社区,并举办了年度 ImageNet 挑战赛,鼓励大家参与。我们继续自己的研究。
But I think what happened is that as a student of AI and then a researcher of AI, I was working on all kinds of mathematical models, including neural networks, including Bayesian networks, including many, many models. And there was one singular pain point: these models don't have data to be trained on. And as a field, we were so focusing on these models, but it dawned on me that human learning as well as evolution is actually a big data learning process. Humans learn with so much experience, constantly. And evolution, if you look at time, animals evolve with just experiencing the world. So I think my students and I conjectured that a very critically overlooked ingredient of bringing AI to life is big data. And then we began this ImageNet project in 2006, 2007. We were very ambitious: we wanted to get the entire internet's image data on objects. Now granted, the internet was a lot smaller than today, so I felt like that ambition was at least not too crazy. Now it's totally delusional to think a couple of graduate students and a professor can do this. But that's what we did. We curated very carefully 15 million images on the internet. Created a taxonomy of 22,000 concepts, borrowing other researchers' work, like a linguist's work on WordNet, which is a particular way of dictionarying words, and we combined that into ImageNet. And we open-sourced that to the research community. We held an annual ImageNet challenge to encourage everybody to participate in this. We continued to do our own research.
但 2012 年被认为是深度学习或现代 AI 诞生的时刻,因为多伦多大学 Geoff Hinton 教授领导的一组研究人员参加了 ImageNet 挑战赛,利用 ImageNet 的大数据和 Nvidia 的两块 GPU,成功创建了第一个神经网络算法,虽然没有完全解决物体识别问题,但取得了巨大进展。这种三要素组合——大数据、神经网络和 GPU——成了现代 AI 的黄金配方。再后来,AI 的公众时刻,也就是 ChatGPT 时刻,如果你看 ChatGPT 的要素,技术上仍然使用了这三个要素。现在是互联网规模的数据,主要是文本;比 2012 年复杂得多的神经网络架构,但仍然是神经网络;以及更多的 GPU,但仍然是 GPU。所以这三个要素仍然是现代 AI 的核心。
But 2012 was the moment that many people think was the beginning of deep learning or the birth of modern AI, because a group of Toronto researchers led by Professor Geoff Hinton participated in the ImageNet challenge, used the ImageNet big data and two GPUs from Nvidia, and created successfully the first neural network algorithm that didn't fundamentally solve, but made a huge progress towards solving the problem of object recognition. And that combination of the trio technology—big data, neural network, and GPU—was kind of the golden recipe for modern AI. And then fast forward, the public moment of AI, which is the ChatGPT moment, if you look at the ingredients of what brought ChatGPT to the world, technically still uses these three ingredients. Now it's internet-scale data, mostly texts; a much more complex neural network architecture than 2012, but it's still neural network; and a lot more GPUs, but it's still GPUs. So these three ingredients are still at the core of modern AI.
太不可思议了。我从未听过完整的故事。我喜欢那两块 GPU。现在可能是几十万块了吧?性能高出几个数量级。那两块 GPU 是直接买的吗?就像游戏 GPU,他们直接去游戏店买的,对吧?就是人们用来玩游戏的。
Incredible. I have never heard that full story before. I love that it was two GPUs. And now it's, I don't know, hundreds of thousands, right? That are orders of magnitude more powerful. And those two GPUs, were they just bought? They were like gaming GPUs, they just went to like the game store, right? That people use for playing games.
正如你所说,这在很大程度上仍然是模型变得更聪明的方式。现在世界上一些增长最快的公司,我大多在播客上请过他们——Merkore、Surge 和 Scale——他们就是这样做的,继续为实验室提供更多他们最感兴趣的标注数据。
As you said, this continues to be in a large way the way models get smarter. Some of the fastest growing companies in the world right now, I've had them all mostly on the podcast—Merkore and Surge and Scale—like they do this, they continue to do this for labs: just give them more and more labeled data of the things they're most excited about.
是的,我记得 Scale 早期的 Alex Wang。我可能还保留着他创办 Scale 时的邮件。他非常友善,一直给我发邮件说 ImageNet 如何启发了 Scale。我很高兴看到这一点。
Yeah, I remember Alex Wang from Scale very early days. I probably still have his emails when he was starting Scale. He was very kind. He kept sending me emails about how ImageNet inspired Scale. I was very pleased to see that.
从你刚才分享的内容中,我另一个最喜欢的收获就是这样一个高能动性和「只管去做」的例子。这在 Twitter 上有点像梗:「你只管去做。」你就觉得,「好吧,这可能是推动 AI 所必需的。」当时它叫机器学习,对吧?那是大多数人用的术语吗?
One of my other favorite takeaways from what you just shared is just such an example of high agency and just doing things. That's kind of a meme on Twitter: 'You can just do things.' You're just like, 'Okay, this is probably necessary to move AI.' And it was called machine learning back then, right? Was that the term most people used?
我认为它们是互换使用的。确实如此。我记得科技公司——我不点名——但在早期的一次对话中,大概是 2015 年中到 2016 年中,一些科技公司避免使用 AI 这个词,因为他们不确定 AI 是不是一个脏词。我记得我实际上鼓励大家使用 AI 这个词,因为对我来说,这是人类在科技探索中提出的最大胆的问题之一,我为这个词感到自豪。但没错,一开始有些人并不确定。
I think it was interchangeably. It's true. I do remember the tech companies—I'm not going to name names—but I was in a conversation in one of the early days, I think in the middle of 2015, middle of 2016, some tech companies avoided using the word AI because they were not sure if AI was a dirty word. And I remember I was actually encouraging everybody to use the word AI, because to me that is one of the most audacious questions humanity has ever asked in our quest for science and technology, and I feel very proud of this term. But yes, at the beginning some people were not sure.
那大概是哪一年 AI 开始发展的?我想是 2016 年,不到十年前。那是一个转折点。有些人开始称它为 AI,但如果你看硅谷科技公司,追溯他们的营销术语,我认为 2017 年左右是公司开始自称 AI 公司的开始。
What year was that roughly when AI was developed? 2016 I think that was less than 10 years ago. That was the changing point. Some people started calling it AI, but I think if you look at Silicon Valley tech companies, if you trace their marketing term, I think 2017ish was the beginning of companies calling themselves AI companies.
太不可思议了,世界变化真大。现在你不可能不称自己为 AI 公司。
That's incredible, just how the world has changed. Now you can't not call yourself an AI company.
我知道。仅仅九年左右。
I know. Just nineish years later.
是的。
Yeah.
天哪。好吧。在我们讨论你工作的未来方向之前,关于早期历史,还有什么你认为人们不知道但很重要的东西吗?
Oh man. Okay. Is there anything else around the history, that early history, that you think people don't know, that you think is important, before we chat about where things are going in the work that you're doing?
我认为,就像所有历史一样,我深知自己因成为历史的一部分而被认可,但还有那么多英雄和研究者。
I think as all histories, I'm keenly aware that I am recognized for being part of the history, but there are so many heroes and so many researchers.
我们谈到了几代研究者。在我自己的领域里,有很多人启发了我,我在书里也提到过。但我确实觉得我们的文化,尤其是硅谷,倾向于把成就归功于一个人。虽然我认为这有其价值,但只是为了被记住。人工智能这个领域已经有 70 年了,我们经历了许多代。没有人能独自走到今天。
We're talking about generations of researchers there. In my own world, there are so many people who have inspired me, which I talked about in my book. But I do feel our culture, especially Silicon Valley, tends to assign achievements to a single person. While I think it has value, it's just to be remembered. AI is a field that is now 70 years old and we have gone through many generations. No one could have gotten here by themselves.
好的。那让我问你这个问题。感觉我们总是处在 AGI 的边缘。这是个模糊的词,人们随口就说。AGI 要来了。它会接管一切吗?你觉得我们离 AGI 还有多远?你认为我们沿着当前的轨迹能到达吗?我们需要更多突破吗?你认为当前的方法能带我们到那里吗?
Okay. So let me ask you this question. It feels like we're always on this precipice of AGI. This kind of vague term people throw around. AGI is coming. Is it going to take over everything? How far do you think we might be from AGI? Do you think we're going to get there on the current trajectory? Do we need more breakthroughs? Do you think the current approach will get us there?
是的,这是一个非常有趣的术语,Lenny。我不知道是否有人真正定义过 AGI。有很多不同的定义,从机器拥有某种超能力,到机器能否成为社会中经济上可行的智能体。换句话说,就是赚取薪水生活。这是 AGI 的定义吗?作为一名科学家,我非常认真地对待科学,我进入这个领域是因为受到这个大胆问题的启发:机器能否像人类一样思考和做事?对我来说,这始终是 AI 的北极星。从这个角度看,我不知道 AI 和 AGI 有什么区别。我认为我们在实现部分目标方面做得很好,包括对话式 AI,但我不认为我们已经完全征服了 AI 的所有目标。我认为我们的先驱,比如艾伦·图灵,如果今天他在世,你问他 AI 与 AGI 的区别,他可能只是耸耸肩说,我在 1940 年代就问过同样的问题。所以我不想陷入定义 AI 与 AGI 的兔子洞。我觉得 AGI 更像是一个营销术语,而不是科学术语。作为科学家和技术专家,AI 是我的北极星,是我领域的北极星,我很高兴人们随便怎么称呼它。
Yeah, this is a very interesting term, Lenny. I don't know if anyone has ever defined AGI. There are many different definitions, including some kind of superpower for machines all the way to whether machines can become economically viable agents in society. In other words, making salaries to live. Is that the definition of AGI? As a scientist, I take science very seriously and I entered the field because I was inspired by this audacious question of can machines think and do things in the way that humans can do. For me, that's always the northstar of AI. From that point of view, I don't know what's the difference between AI and AGI. I think we've done very well in achieving parts of the goal, including conversational AI, but I don't think we have completely conquered all the goals of AI. I think our founding fathers like Alan Turing, I wonder if Alan Turing is around today and you ask him to contrast AI versus AGI, he might just shrug and say, well, I asked the same question back in the 1940s. So I don't want to get into a rabbit hole of defining AI versus AGI. I feel AGI is more a marketing term than a scientific term. As a scientist and technologist, AI is my northstar, my field's northstar, and I'm happy people call it whatever name they want.
那让我换个方式问:你描述了从 ImageNet 和 AlexNet 到今天的这些组件。GPU、数据、标签、模型算法。Transformer 感觉是这条轨迹上的重要一步。你觉得这些相同的组件能带我们走向一个聪明 10 倍的模型,改变整个世界,还是我们需要更多突破?我知道我们要讨论世界模型,我认为这是其中的一部分,但还有其他东西吗?你觉得会达到平台期,还是只需要更多数据、更多算力、更多 GPU 就能继续前进?
So let me ask you maybe this way: you described there are these components that from ImageNet and AlexNet took us to where we are today. GPUs, data, labels, the algorithm of the model. The transformer feels like an important step in that trajectory. Do you feel like those are the same components that'll get us to a 10 times smarter model, something life-changing for the entire world, or do you think we need more breakthroughs? I know we're going to talk about world models, which I think is a component of this, but is there anything else that you think will plateau or will take us further with just more data, more compute, more GPUs?
哦不,我绝对认为我们需要更多创新。我认为 Scaling 定律——更多数据、更多 GPU 和更大的当前模型架构——还有很多工作要做。但我绝对认为我们需要更多创新。人类历史上没有一个深刻的科学学科到达过说「我们创新完了」的地步。而 AI 是人类文明中在科学和技术方面最年轻的学科之一。我们仍在触及表面。例如,就像我说的,我们今天要过渡到世界模型。你拿一个模型,让它看一段几个办公室的视频,然后问模型数一下有多少把椅子。这是幼儿都能做的事,或者小学生都能做,但 AI 做不到。所以今天 AI 还有很多不能做的事,更不用说思考像艾萨克·牛顿这样的人如何观察天体的运动并推导出支配所有物体运动的方程。那种创造力、外推、抽象,我们今天无法让 AI 做到。然后我们看看情商。如果一个学生走进老师的办公室,谈论动机、热情、学什么、真正困扰你的问题是什么,那种对话,尽管今天的对话机器人很强大,但今天的 AI 没有那种情感认知智能。所以我们还有很多可以改进的地方。我不相信我们已经创新完了。
Oh no, I definitely think we need more innovations. I think scaling laws of more data, more GPUs and bigger current model architecture, there's still a lot to be done there. But I absolutely think we need to innovate more. There is not a single deeply scientific discipline in human history that has arrived at a place that says we're done innovating. And AI is one of the youngest disciplines in human civilization in terms of science and technology. We're still scratching the surface. For example, like I said, we're going to segue into world models today. You take a model and run it through a video of a couple of office rooms and ask the model to count the number of chairs. This is something a toddler could do, or maybe an elementary school kid could do, and AI could not do that. So there's just so much AI today cannot do, let alone thinking about how someone like Isaac Newton looked at the movements of celestial bodies and derived an equation or a set of equations that governs the movement of all bodies. That level of creativity, extrapolation, abstraction, we have no way of enabling AI to do that today. And then let's look at emotional intelligence. If you look at a student coming into a teacher's office and having a conversation about motivation, passion, what to learn, what's the problem that's really bothering you, that conversation, as powerful as today's conversational bots are, you don't get that level of emotional cognitive intelligence from today's AI. So there's a lot we can do better. And I do not believe we're done innovating.
Demis 最近在 DeepMind Google 有一个非常有趣的采访,有人问他你觉得我们离 AGI 有多远?它实现时是什么样子?他有一个非常有趣的思路:如果我们给最前沿的模型所有到 20 世纪末的信息,看它能否得出爱因斯坦的所有突破。到目前为止,我们离那还很远。
Demis had this really interesting interview recently from DeepMind Google where someone asked him just like what do you think how far are we from AGI? What does it look like when it's through there? He had a really interesting way of approaching it: if we were to give the most cutting-edge model all the information until the end of the 20th century, see if it could come up with all the breakthroughs Einstein had. And so far we're nowhere near that.
不,我们没有。事实上,甚至更糟。让我们给 AI 所有数据,包括牛顿没有的现代仪器观测的天体数据,然后让 AI 创建 17 世纪关于物体运动定律的方程组。今天的 AI 做不到。
No, we're not. In fact, it's even worse. Let's give AI all the data including modern instruments data of celestial bodies which Newton did not have, and give it to that and just ask AI to create the 17th century set of equations on the laws of bodily movements. Today's AI cannot do that.
好吧,我听到的是我们还有很长的路要走。
All right, we're a ways away is what I'm hearing.
是的。
Yeah.
好的,那我们谈谈世界模型。对我来说,这又是一个你领先于他人的绝佳例子。你早就提出「我们需要大量干净数据让 AI 和神经网络学习」。你谈论世界模型这个想法已经很久了。你创办了一家公司来构建语言模型。但这是不同的东西。这是世界模型。我们来谈谈它是什么。而且在我准备这次访谈时,Elon 在谈论世界模型,Jensen 在谈论世界模型。我知道 Google 也在做这个。你在这方面已经很久了。而且你实际上刚刚发布了一些东西,就在这个播客播出之前。谈谈什么是世界模型?为什么它如此重要?
Okay, so let's talk about world models. This is to me just another really amazing example of you being ahead of where people end up. So you were way ahead on 'we just need a lot of clean data for AI and neural networks to learn.' You've been talking about this idea of world models for a long time. You started a company to build essentially language models. This is a different thing. This is a world model. We'll talk about what that is. And now as I was preparing for this, Elon's like talking about world models. Jensen's talking about world models. I know Google's working on this stuff. You've been at this for a long time. And you're actually just launched something that we're going to talk about right before this podcast airs. Talk about what is a world model? Why is it so important?
我很兴奋看到越来越多的人在谈论世界模型,比如 Elon,比如 Jensen。我一生都在思考如何推动 AI 前进。过去几年从研究界和 OpenAI 等涌现出的大语言模型,即使对我这样的研究者来说也极具启发性。我记得 GPT2 发布的时候,我想是在 2020 年底。
I'm very excited to see that more and more people are talking about world models like Elon, like Jensen. I have been thinking about how to push AI forward all my life. The large language models that came out of the research world and then OpenAI and all this for the past few years were extremely inspiring even for a researcher like me. I remembered when GPT2 came out and that was in I think late 2020.
我曾是联合主任,现在依然是,但当时我是斯坦福大学以人为本人工智能研究所的全职联合主任。我记得公众当时还没有意识到大语言模型的威力,但作为研究人员,我们已经看到了未来。我和自然语言处理领域的同事,比如 Percy Liang 和 Chris Batting,进行了很长的对话,讨论这项技术将有多么关键。斯坦福的 AI 研究所,以人为本人工智能研究所,是第一个建立基础模型完整研究中心的机构。Percy Liang 和许多研究人员领导了第一篇关于基础模型的学术论文。所以这对我来说非常鼓舞人心。当然,我来自视觉智能领域,我在想,我们可以在语言之外推动很多进展,因为人类利用空间智能和世界理解能力做了很多超越语言的事情。想象一个非常混乱的急救现场,无论是火灾、交通事故还是自然灾害。如果你置身其中,思考人们如何组织起来救人、阻止灾害蔓延、灭火——其中很多是动作,是对物体、世界、情境的自发理解。语言是其中的一部分,但很多情况下语言无法帮你灭火。所以我思考了很多,同时我也在做很多机器人研究。我突然意识到,连接语言之外的额外智能、连接具身 AI(即机器人)以及连接视觉智能的关键,就是这种理解世界的空间智能。那是在 2024 年,我发表了一场关于空间智能和世界模型的 TED 演讲。基于我的机器人和计算机视觉研究,我开始形成这个想法。有一点非常明确:我渴望与最优秀的技术专家合作,尽快将这项技术变为现实。于是我们创立了这家名为 World Labs 的公司。你可以看到公司名字里有「世界」这个词,因为我们非常相信世界模型和空间智能。
I was co-director, I still am, but I was at that time full-time co-director of Stanford's Human-Centered AI Institute. I remember the public was not aware of the power of the large language model yet, but as researchers we were seeing the future. I had pretty long conversations with my natural language processing colleagues like Percy Liang and Chris Batting, talking about how critical this technology is going to be. Stanford's AI Institute, Human-Centered AI Institute, was the first to establish a full research center on foundation models. Percy Liang and many researchers led the first academic paper on foundation models. So it was very inspiring for me. Of course, I come from the world of visual intelligence, and I was thinking there's so much we can push forward beyond language because humans have used our sense of spatial intelligence and world understanding to do so many things that are beyond language. Think about a very chaotic first responder scene, whether it's fire, a traffic accident, or a natural disaster. If you immerse yourself in that scene and think about how people organize themselves to rescue people, stop further disasters, put down fires—a lot of that is movements, spontaneous understanding of objects, worlds, situational awareness. Language is part of that, but a lot of those situations language cannot get you to put down the fire. So I was thinking a lot, and in the meantime I was doing a lot of robotics research. It dawned on me that the lynchpin connecting additional intelligence beyond language and connecting embodied AI—which is robotics—and connecting visual intelligence is this sense of spatial intelligence about understanding the world. That's when, I think it was 2024, I gave a TED talk about spatial intelligence and world models. I started formulating this idea based on my robotics and computer vision research. One thing that is really clear to me is that I really want to work with the brightest technologists and move as fast as possible to bring this technology to life. That's when we founded this company called World Labs. You can see the word 'world' is in the title because we believe so much in world modeling and spatial intelligence.
人们已经习惯了聊天机器人,那是一个大语言模型。所以理解世界模型的一个简单方式是,你描述一个场景,它就会生成一个可以无限探索的世界。我们会链接到你推出的那个东西,我们会讨论它,但这是一个简单的理解方式吗?
People are so used to just chatbots and that's a large language model. So the simple way to understand a world model is you basically describe a scene and it generates an infinitely explorable world. We'll link to the thing you launch which we'll talk about, but is that a simple way to understand it?
这是其中的一部分,Lenny。我认为理解世界模型的一个简单方式是,这个模型可以让任何人通过提示(无论是图像还是句子)在脑海中创造任何世界,并且能够在这个世界中互动——无论是浏览、行走、拿起物体、改变事物,还是在这个世界中推理。例如,如果消费世界模型输出的是一个机器人,它应该能够规划路径并帮助整理厨房。所以世界模型是一个基础,你可以用它来推理、互动和创造世界。
That's part of it, Lenny. I think a simple way to understand a world model is that this model can allow anyone to create any worlds in their mind's eye by prompting, whether it's an image or a sentence, and also be able to interact in this world—whether you're browsing and walking, picking objects up, changing things, as well as to reason within this world. For example, if the person consuming the output of the world model is a robot, it should be able to plan its path and help tidy the kitchen. So world model is a foundation that you can use to reason, to interact, and to create worlds.
太好了。是的。所以机器人感觉像是 AI 研究者的下一个大焦点,就像对世界的影响一样。你在这里说的是,这是让机器人在现实世界中真正工作的一个关键缺失部分——理解世界如何运作。
Great. Yeah. So robots feels like that's potentially the next big focus for AI researchers and just like the impact on the world. And what you're saying here is this is a key missing piece of making robots actually work in the real world. Understanding how the world works.
是的。首先,我认为除了机器人之外,还有更多令人兴奋的东西。但我同意你所说的一切。我认为世界模型和空间智能是具身 AI 的一个关键缺失部分。我也认为我们不要低估人类作为具身智能体,可以从世界模型和空间智能模型中受益的程度,就像机器人一样。就像今天人类是语言动物,但我们在完成语言任务(包括软件工程)时很大程度上得到了 AI 的增强。我认为我们不应该低估,或者说我们往往不谈论人类作为具身智能体如何能从世界模型和空间智能模型中受益如此之多。所以最大的突破点:机器人,这非常重要。如果这成功了,想象一下我们每个人都拥有机器人帮我们做各种事情。它们帮助我们应对灾难等等。游戏显然是一个非常酷的例子。就像可以无限玩的游戏,你从脑子里发明出来。然后创造力感觉就像玩得开心、有创意、想象狂野的新世界和环境。
Yeah. Well, first of all, I do think there's more than robots that's exciting. So, but I agree with everything you just said. I think world modeling and spatial intelligence is a key missing piece of embodied AI. I also think let's not underestimate that humans are embodied agents and humans can be augmented by AI's intelligence just like today humans are language animals but we're very much augmented by AI when helping us do language tasks including software engineering. I think that we shouldn't underestimate, or maybe it's we tend not to talk about how humans as embodied agents can actually benefit so much from world models and spatial intelligence models as well as robots can. So the big unlocks here: robots, which is a huge deal. If this works out, imagine each of us has robots doing a bunch of stuff for us. They help us with disasters, things like that. Games obviously is a really cool example. Just like infinitely playable games that you just invent out of your head. And then creativity feels like just having fun, being creative, thinking of wild new worlds and environments.
还有设计。人类设计从机器到建筑到住宅,还有科学发现,有很多。我喜欢用 DNA 结构发现的例子。如果你看 DNA 发现史上最重要的部分之一,就是 Rosalind Franklin 拍摄的 X 射线衍射照片。那是一张扁平的 2D 照片,结构看起来像一个带有衍射的十字。你可以谷歌这些照片。但凭借那张 2D 平面照片,人类,尤其是两位重要人物 James Watson 和 Francis Crick,结合其他信息,能够在 3D 空间中推理,推导出高度三维的 DNA 双螺旋结构。那个结构不可能是 2D 的。你不能用 2D 思维推导出那个结构。你必须用 3D 空间思维,运用人类的空间智能。所以我认为即使在科学发现中,空间智能或 AI 辅助的空间智能也至关重要。
And also design. Humans design from machines to buildings to homes and also scientific discovery right there is so much. I like to use the example of the discovery of the structure of DNA. If you look at one of the most important pieces in DNA's discovery history, it's the X-ray diffraction photo captured by Rosalind Franklin. It was a flat 2D photo of a structure that looks like a cross with diffractions. You can Google those photos. But with that 2D flat photo, humans, especially two important humans, James Watson and Francis Crick, in addition to their other information, were able to reason in 3D space and deduce a highly three-dimensional double helix structure of the DNA. That structure cannot possibly be 2D. You cannot think in 2D and deduce that structure. You have to think in 3D spatial, use human spatial intelligence. So I think even in scientific discovery, spatial intelligence or AI-assisted spatial intelligence is critical.
这正是一个例子,我记得是 Chris Dixon 说过,下一个大事件一开始会感觉像个玩具。当 ChatGPT 刚出来时,我记得 Sam Altman 发推文说「这是我们正在玩的一个很酷的东西。来看看吧。」现在它成了历史上增长最快的产品,改变了世界。
This is such an example of I think it was Chris Dixon that had this line that the next big thing is going to start off feeling like a toy. When ChatGPT just came out, I remember Sam Altman just tweeted like 'here's a cool thing we're playing with. Check it out.' Now it's the fastest growing product in history, changed the world.
是的。
Yeah.
而且往往是那些看起来「好吧,这很酷」,玩起来很有趣的东西,最终最能改变世界。
And it's oftentimes the things that just look like 'okay this is cool', that it's fun to play with, end up changing the world most.
是的。
Yeah.
本期节目由 Cinch 赞助,Cinch 是客户通信云。关于数字客户通信,无论你是发送营销活动、验证码还是账户提醒,你都需要它们可靠地到达用户。这就是 Cinch 的用武之地。超过 15 万家企业,包括全球前十大科技公司中的八家,使用 Cinch 的 API 将消息、电子邮件和通话功能集成到他们的产品中。
This episode is brought to you by Cinch, the customer communications cloud. Here's the thing about digital customer communications. Whether you're sending marketing campaigns, verification codes, or account alerts, you need them to reach users reliably. That's where Cinch comes in. Over 150,000 businesses, including eight of the top 10 largest tech companies globally, use Cinch's API to build messaging, email, and calling into their products.
我联系了本·霍洛维茨,他很欣赏你所做的事情。他是你的忠实粉丝。我相信他们是投资者。
I reached out to Ben Horowitz who loves what you're doing. A big fan of yours. They're investors I believe.
是的,我们认识很多年了,但没错,现在他们是 World Labs 的投资者。
Yeah, we've known each other for many years, but yes, right now they are investors of World Labs.
太棒了。好的。所以我问他我应该问你什么,他建议问你为什么单靠苦涩教训不太可能适用于机器人。所以首先解释一下 AI 历史上苦涩教训是什么,然后为什么它不能让我们达到机器人的目标。
Amazing. Okay. So I asked him what I should ask you about and he suggested ask you why is the bitter lesson alone not likely to work for robots. So first of all just explain what the bitter lesson was in the history of AI and then just why that won't get us to where we want to be with robots.
嗯,首先,有很多苦涩教训,但大家常说的苦涩教训是理查德·萨顿写的一篇论文,他最近获得了图灵奖,做了很多强化学习。理查德说过,如果你回顾历史,尤其是 AI 的算法发展,最终总是简单模型加上大量数据胜出,而不是复杂模型加上少量数据。实际上,这篇论文是在 ImageNet 之后多年才发表的。对我来说,那不是苦涩,而是甜蜜的教训。这就是为什么我构建了 ImageNet,因为我相信大数据扮演了那个角色。那么为什么苦涩教训单独在机器人领域行不通呢?首先,我认为我们需要承认我们现在的处境。机器人技术还处于实验的早期阶段。研究远没有语言模型那么成熟。所以很多人仍在尝试不同的算法,其中一些算法是由大数据驱动的。所以我确实认为大数据将继续在机器人领域发挥作用。但机器人领域难在哪里?有几个方面。一是获取数据更难。获取数据要困难得多。你可能会说,有网络数据。最新的机器人研究正在使用网络视频,我认为网络视频确实有作用。但如果你思考是什么让语言模型成功,作为一个从事计算机视觉、空间智能和机器人研究的人,我非常羡慕语言领域的同事,因为他们有一个完美的设置:训练数据是文字,最终是词元,然后他们生成一个输出文字的模型。所以在你希望得到的东西(我们称之为目标函数)和你的训练数据之间有着完美的对齐。但机器人不同。甚至空间智能也不同。你希望从机器人那里得到动作。但你的训练数据缺乏 3D 世界中的动作。而这就是机器人必须做的,对吧?3D 世界中的动作。所以你必须找到不同的方法来把方钉塞进圆孔。我们有的是大量的网络视频。所以我们必须开始讨论补充数据,比如遥操作数据或合成数据,这样机器人才能按照苦涩教训的假设(即大量数据)进行训练。我认为仍有希望,因为即使我们在世界建模方面所做的工作也将真正为机器人解锁大量信息。但我们必须小心,因为我们还处于早期阶段,苦涩教训仍有待检验,因为我们还没有完全弄清楚数据。对于机器人苦涩教训的另一部分,我认为我们应该非常现实地认识到,与语言模型甚至空间模型相比,机器人是物理系统。所以机器人更接近自动驾驶汽车,而不是大型语言模型。认识到这一点非常重要。这意味着要让机器人工作,我们不仅需要大脑,还需要物理身体,还需要应用场景。如果你看看自动驾驶汽车的历史,我的同事塞巴斯蒂安·特龙在 2006 年或 2005 年驾驶斯坦福的汽车赢得了第一届 DARPA 挑战赛。从那个能在内华达沙漠行驶 130 英里的自动驾驶汽车原型到今天旧金山街头的 Waymo,已经过去了 20 年,我们甚至还没有完成。还有很多工作要做。所以这是一个 20 年的旅程。而自动驾驶汽车是更简单的机器人。它们只是在二维表面上运行的金属盒子。目标是不要碰到任何东西。机器人是在三维世界中运行的三维物体,目标是触摸东西。所以这个旅程将涉及许多方面和元素。当然,有人可能会说,自动驾驶汽车的早期算法是深度学习时代之前的。所以深度学习正在加速大脑的发展,我认为这是真的。这就是为什么我从事机器人研究。这就是为什么我从事空间智能研究,并且对此感到兴奋。但与此同时,汽车行业非常成熟,产品化也涉及成熟的用例、供应链和硬件。所以我认为现在是研究这些问题的非常有趣的时期。但确实,本是对的。我们可能仍然要经历许多苦涩教训。
Well, first of all, there are many bitter lessons, but the bitter lesson everybody refers to is a paper written by Richard Sutton who won the Turing Award recently and he does a lot of reinforcement learning. Richard has said that if you look at the history, especially the algorithmic development of AI, it turns out simpler models with a ton of data always win at the end of the day instead of the more complex model with less data. I mean, that was actually this paper came years after ImageNet. To me, that was not bitter, it was a sweet lesson. That's why I built ImageNet because I believe that big data plays that role. So why can bitter lesson work in robotics alone? Well, first of all, I think we need to give credit to where we are today. Robotics is very much in the early days of experimentation. The research is not nearly as mature as say language models. So many people are still experimenting with different algorithms and some of those algorithms are driven by big data. So I do think big data will continue to play a role in robotics. But what is hard for robotics? There are a couple of things. One is that it's harder to get data. It's a lot harder to get data. You can say, well, there is web data. This is where the latest robotics research is using web videos, and I think web videos do play a role. But if you think about what made language models work, as someone who does computer vision and spatial intelligence and robotics, I'm very jealous of my colleagues in language because they had this perfect setup where their training data are in words, eventually tokens, and then they produce a model that outputs words. So you have this perfect alignment between what you hope to get, which we call objective function, and what your training data looks like. But robotics is different. Even spatial intelligence is different. You hope to get actions out of robots. But your training data lacks actions in 3D worlds. And that's what robots have to do, right? Actions in 3D worlds. So you have to find different ways to fit a square in a round hole. What we have is tons of web videos. So then we have to start talking about supplementing data such as teleoperation data or synthetic data so that the robots are trained with this hypothesis of bitter lesson, which is large amount of data. I think there's still hope because even what we are doing in world modeling will really unlock a lot of this information for robots. But I think we have to be careful because we're at the early days of this and bitter lesson is still to be tested because we haven't fully figured out the data. For another part of the bitter lesson of robotics, I think we should be so realistic about is again compared to language models or even spatial models, robots are physical systems. So robots are closer to self-driving cars than a large language model. And that's very important to recognize. That means that in order for robots to work, we not only need brains, we also need the physical body, we also need application scenarios. And if you look at the history of self-driving car, my colleague Sebastian Thrun took Stanford's car to win the first DARPA challenge in 2006 or 2005. It's 20 years since that prototype of a self-driving car being able to drive 130 miles in the Nevada desert to today's Waymo on the street of San Francisco and we're not even done yet. There's still a lot. So that's a 20 year journey. And self-driving cars are much simpler robots. They're just metal boxes running on 2D surfaces. And the goal is not to touch anything. Robot is 3D things running in 3D world and the goal is to touch things. So the journey is going to have many aspects and elements. And of course one could say well the self-driving car early algorithms were pre-deep learning era. So deep learning is accelerating the brains and I think that's true. That's why I'm in robotics. That's why I'm in spatial intelligence and I'm excited by it. But in the meantime, the car industry is very mature and productizing also involves the mature use cases, supply chains, the hardware. So I think it's a very interesting time to work in these problems. But it's true Ben is right. We might still be subject to a number of bitter lessons.
做这项工作,你是否曾对大脑的工作方式以及它为我们所做的一切感到敬畏?仅仅是让机器四处走动而不撞到东西或摔倒的复杂性。这是否让你对我们已有的东西更加敬畏?
Doing this work, do you ever just feel awe for the way the brain works and is able to do all of this for us? Just the complexity just to get a machine to just walk around and not hit things and fall. Does it just give you more spec for what we've already got?
完全同意。我们以大约 20 瓦的功率运行。这比我此刻所在房间里的任何灯泡都要暗。然而我们能做这么多。所以实际上,我在 AI 领域工作得越多,就越尊重人类。
Totally. We operate on about 20 watts. That's dimmer than any light bulb in the room I'm in right now. And yet we can do so much. So I think actually the more I work in AI, the more I respect humans.
我们来谈谈你刚刚推出的这个产品。它叫 Marble。一个非常可爱的名字。谈谈这是什么,为什么重要。我一直在玩它。太不可思议了。我们会提供链接让大家去看看。Marble 是什么?
Let's talk about this product you just launched. It's called Marble. A very cute name. Talk about what this is, why this important. I've been playing with it. It's incredible. We'll link to it for folks to check it out. What is Marble?
是的,我非常兴奋。首先,Marble 是 World Labs 推出的首批产品之一。World Labs 是一家基础前沿模型公司。我们由四位拥有深厚技术背景的联合创始人创立。我的联合创始人是贾斯汀·约翰逊、克里斯托夫·拉斯纳和本·米尔登霍尔。我们都来自 AI、计算机图形学、计算机视觉的研究领域。我们相信空间智能和世界建模与语言模型同样重要,甚至更重要,并且是对语言模型的补充。
Yeah, I'm very excited. So first of all, Marble is one of the first products that World Labs has rolled out. World Labs is a foundation frontier model company. We are founded by four co-founders who have deep technical history. My co-founders Justin Johnson, Christoph Lassner, and Ben Mildenhall. We all come from the research field of AI, computer graphics, computer vision. And we believe that spatial intelligence and world modeling is as important if not more to language models and complementary to language models.
所以我们想抓住这个机会,创建一个深度科技研究实验室,将前沿模型与产品连接起来。Marvel 是一个基于我们前沿模型构建的应用。我们花了一年多的时间构建了世界上第一个能够生成真正 3D 世界的生成模型。这是一个非常非常困难的问题。过程也非常艰难。我们拥有一支由杰出技术专家组成的创始团队。大约一两个月前,我们第一次看到,只需输入一个句子、一张图片或多张图片,就能创建出可以自由导航的世界。如果你戴上我们提供的可选眼镜,甚至可以四处走动。尽管我们已经开发了相当长一段时间,但这仍然令人敬畏,我们想把它交到需要的人手中。我们知道很多创作者、设计师、考虑机器人模拟的人、考虑可导航、可交互、沉浸式世界不同用例的人,以及游戏开发者都会觉得它有用。所以我们开发了 Marvel 作为第一步。它仍然非常早期,但它是世界上第一个做到这一点的模型,也是世界上第一个允许人们通过提示词生成世界的产品,我们称之为「提示到世界」。
So we wanted to seize this opportunity to create a deep tech research lab that can connect the dots between frontier models with products. So, Marvel is an app that's built upon our frontier models. We've spent a year and plus building the world's first generative model that can output genuinely 3D worlds. That's a very, very hard problem. And it was a very hard process. We have a team of incredible founding team of incredible technologists from incredible teams. And then around just a month or two ago, we saw the first time that we can just prompt with a sentence and an image and multiple images and create worlds that we can just navigate in. If you put it on a goggle, which we have an option to let you do that, you can even walk around. So it was, even though we've been building this for quite a while, it was still just awe-inspiring and we wanted to get it into the hands of people who need it. And then we know that so many creators, designers, people who are thinking about robotic simulation, people who are thinking about different use cases of navigable, interactable, immersive worlds, game developers will find this useful. So we developed Marvel as a first step. It's still very early, but it's the world's first model doing this and it's the world's first product that allows people to just prompt, we call it prompt to worlds.
嗯,我一直在玩它。太疯狂了。就像你可以有一个小世界,基本上可以无限地在中土世界漫步,虽然还没有人,但太疯狂了。你可以去任何地方。还有反乌托邦世界。我刚刚看了所有这些例子。
Well, I've been playing around with it. It is insane. Like you could just have a little sh world where you just infinitely walk around Middle Earth basically and there's no one there yet but it's insane. You just go anywhere. There's like dystopian world. I'm just looking at all these examples.
是的。实际上我最喜欢的部分,我不知道这是功能还是 bug,你可以在世界实际渲染所有纹理之前看到它的点阵。我就喜欢窥探这个模型在创建时到底发生了什么。
Yes. And my favorite part actually, I don't know if this is a feature or bug, you can see the dots of the world before it actually renders with all the textures. And I just love to get a glimpse into what is going on with this model basically create.
听到这太酷了,因为作为一名研究者,我学到了东西。那些引导你进入世界的点阵是一个有意的可视化功能。它不是模型的一部分。模型实际上只是生成世界。我们试图找到一种引导人们进入世界的方法,许多工程师尝试了不同版本,但我们最终采用了点阵。很多人,不只是你,告诉我们这种体验多么令人愉快。听到这个有意的可视化功能——不仅仅是那个大型核心模型——实际上让用户感到愉悦,我们非常满意。
That's so cool to hear because this is where as a researcher I'm learning because the dots that lead you into the world was an intentional feature visualization. It is not part of the model. It's the model actually just generates the world. We were trying to find a way to guide people into the world and a number of engineers worked on different versions but we converged on the dot and so many people, you're not the only one, told us how delightful that experience is and it was really satisfying for us to hear that this intentional visualization feature that's not just the big hardcore model actually has delighted our users.
哇。所以你添加这个是为了让人类更能理解发生了什么,更令人愉快。哇,太有趣了。这让我想到语言模型,它们谈论自己在想什么、在做什么的方式。
Wow. So, you add that to make it more like to have humans understand what's going on more, get more delightful. Wow, that is hilarious. It makes me think about LM and the way they talk about what they're thinking and what they're doing.
是的,确实如此。
Yes, it is. It is.
这也让我想到《黑客帝国》。就像,完全是《黑客帝国》的体验。我不知道这是不是你的灵感来源。
It also makes me think about just the Matrix. Like, it's exactly the Matrix experience. I don't know if that was your inspiration.
嗯,就像我说的,很多工程师参与了那个。可能是他们的灵感。在他们的潜意识里。
Well, like I said, a number of engineers worked on that. It could be their inspiration. It's in their subconscious.
好的。那么,对于可能想玩一玩、用一用的人,今天有哪些应用可以开始使用?你这次发布的目标是什么?
Okay. So, just for folks that may want to play around with this, maybe use it. What are some applications today that folks can start using today? What's your goal with this launch?
是的。我们相信世界建模是非常横向的,但我们已经看到一些非常令人兴奋的用例。电影虚拟制作,因为他们需要的是可以与摄像机对齐的 3D 世界,这样当演员表演时,他们可以很好地定位摄像机并拍摄片段,我们已经看到了不可思议的应用。事实上,我不知道你是否看过我们的 Marvel 发布视频,它是由一家虚拟制作公司制作的。我们与索尼合作,他们用 Marvel 的东西拍摄了那些视频。所以我们与那些技术艺术家和导演合作,他们说这已经将他们的制作时间缩短了 40 倍。确实如此。实际上,因为我们只有一个月的时间来完成这个项目,而且他们要拍摄的东西很多。所以使用 Marvel 确实大大加速了 VFX 和电影的虚拟制作。这是一个用例。我们已经看到用户将我们的 Marvel 场景导出网格并放入游戏中,无论是 VR 游戏还是他们开发的普通游戏。我们展示了一个机器人模拟的例子,因为当我——我仍然是一名从事机器人训练的研究人员——最大的痛点之一是为训练机器人创建合成数据。这些合成数据需要非常多样化,需要来自不同的环境,有不同的物体可以操作。一条路径是让计算机模拟。否则,人类必须为机器人构建每一个资产,那将花费更长的时间。所以已经有研究人员联系我们,希望使用 Marvel 来创建这些合成环境。我们还有意想不到的用户联系,关于他们想如何使用 Marvel。例如,一个心理学团队联系我们,想用 Marvel 做心理学研究。原来他们研究的一些精神病人需要了解他们的大脑如何对不同特征的沉浸式场景做出反应。例如,凌乱的场景或干净的场景,等等。研究人员很难获得这类沉浸式场景,创建它们需要太长时间和太多预算。而 Marvel 几乎可以即时地将这么多实验环境交到他们手中。所以,目前我们看到了多个用例,但 VFX、游戏开发者、模拟开发者以及设计师都非常兴奋。
Yeah. So, we do believe that world modeling is very horizontal, but we're already seeing some really exciting use cases. Virtual production for movies because what they need are 3D worlds that they can align with the camera so when the actors are acting on it they can position the camera and shoot the segments really well and we're already seeing incredible use. In fact, I don't know if you have seen our launch video showing Marvel, it was produced by a virtual production company. We collaborated with Sony and they used Marvel things to shoot those videos. So we were collaborating with those technical artists and directors and they were saying this has cut our production time by 40x. In fact it has. Yes. In fact, I had to because we only had one month to work on this project and there were so many things they were trying to shoot. So using Marvel really significantly accelerated the production of virtual production for VFX and movies. That's one use case. We are already seeing our users taking our Marvel scene and taking the mesh export and putting it into games, whether it's games on VR or games just fun games that they have developed. We have had, we were showing an example of robotic simulation because when I was, I mean I'm still am a researcher doing robotic training. One of the biggest pain points is to create synthetic data for training robots. And these synthetic data needs to be very diverse. They need to come from different environments with different objects to manipulate. And one path to it is to ask computers to simulate. Otherwise, humans have to build every single asset for robots. That's just going to take a lot longer. So we already have researchers reaching out and wanting to use Marvel to create those synthetic environments. We also have unexpected user outreach in terms of how they want to use Marvel. For example, a psychologist team called us to use Marvel to do psychology research. It turned out some of the psychiatric patients they study, they need to understand how their brain responds to different immersive scenes of different features. For example, messy scenes or clean scenes or whatever you name it. And it's very hard for researchers to get their hands on these kind of immersive scenes and it will take them too long and too much budget to create. And Marvel is a really almost instantaneous way of getting so many of these experimental environments into their hands. So, we're seeing multiple use cases at this point, but the VFX, the game developers, the simulation developers as well as designers are very excited.
这非常符合 AI 领域的工作方式。我在播客中采访过其他 AI 领导者,他们总是说尽早把东西推出去,以发现大的用例在哪里。ChatGPT 的负责人告诉我,当他们第一次推出 ChatGPT 时,他就在 TikTok 上扫描,看人们如何使用它,谈论什么,这说服了他们应该在哪里发力,帮助他们看到人们实际想如何使用它。我喜欢最后一个用例,比如用于治疗。
This is very much the way things work in AI. I've had other AI leaders on the podcast and it's always like put things out there early as soon as you can to discover where the big use cases are. The head of ChatGPT told me how when they first put out ChatGPT, he was just scanning TikTok to see how people were using it and all the things they were talking about and that's what convinced them where to lean in and help them see how people actually want to use it. I love this last use case of like for therapy.
我就在想人们面对恐高、蛇或蜘蛛的场景。太棒了。昨晚我一个朋友打电话跟我说他恐高,还问我是否应该用 Marble。太神奇了,你一下子就想到那儿了。
I'm just imagining people dealing with heights or snakes or spiders. It's amazing. A friend of mine last night literally called me and talked about his height scare and asked me if marble should be used. That's amazing. You went straight there.
那是因为我在想所有那些暴露疗法。这对此会非常有用。太酷了。
That's because I'm imagining all the exposure therapy stuff. This could be so good for that. That is so cool.
好的,那么我想问:这与 V3 和其他视频生成模型有何不同?对我来说很清楚,但我觉得解释一下它与人们见过的所有视频 AI 工具有何不同可能会有所帮助。
Okay, so let me ask you this: how does this differ from things like V3 and other video generation models? It's pretty clear to me, but I think it might be helpful to explain how this is different from all the video AI tools people have seen.
World Lab 的核心理念是空间智能从根本上非常重要。空间智能不仅仅是关于视频。事实上,世界并不是被动地观看视频流逝。我喜欢用柏拉图的洞穴寓言来描述视觉。他说想象一个囚犯被绑在洞穴里的椅子上,看着面前的全息剧场,但真正的演员表演的剧场在他背后。灯光使得动作投影在洞穴的墙壁上。这个囚犯的任务是弄清楚发生了什么。这是一个极端的例子,但它描述了视觉的本质:从二维中理解三维或四维世界。所以对我来说,空间智能比创造那个平坦的二维世界更深刻。空间智能是创造、推理、互动和理解深层空间世界的能力,无论是二维、三维还是四维,包括动态。World Lab 专注于这一点。创造视频本身的能力可以是其中的一部分。事实上,就在几周前,我们推出了世界上首个在单个 H100 GPU 上实时可演示的视频生成。所以我们的技术包括这一点。但我认为 Marble 非常不同,因为我们真正希望创作者、设计师和开发者拥有一个能够提供具有三维结构世界的模型,以便他们用于工作。这就是 Marble 如此不同的原因。
World Lab's thesis is that spatial intelligence is fundamentally very important. Spatial intelligence is not just about videos. In fact, the world is not passively watching videos passing by. I love Plato's allegory of the cave to describe vision. He said imagine a prisoner tied to his chair in a cave watching a full life theater in front of him, but the actual live theater with actors acting is behind his back. It was lit so that the projection of the action is on a wall of the cave. The task of this prisoner is to figure out what's going on. It's an extreme example, but it describes what vision is about: to make sense of the 3D world or 4D world out of 2D. So spatial intelligence to me is deeper than creating that flat 2D world. Spatial intelligence is the ability to create, reason, interact, and make sense of a deeply spatial world, whether it's 2D, 3D, or 4D, including dynamics. World Lab is focusing on that. The ability to create videos per se could be part of this. In fact, just a couple of weeks ago we rolled out the world's first real-time demoable video generation on a single H100 GPU. So part of our technology includes that. But I think Marble is very different because we really want creators, designers, and developers to have a model that can give them worlds with 3D structure so they can use it for their work. That's why Marble is so different.
在我看来,它是一个充满机遇的平台。正如你所说,视频只是一次性的,非常有趣和酷,然后就这样了,你继续前进。
The way I see it, it's a platform for a ton of opportunity to do stuff. As you described, videos are just a one-off video that's very fun and cool, and that's it, you move on.
顺便说一句,在 Marble 中,我们可以允许人们以视频形式导出。所以你可以进入一个世界。比如说一个霍比特人洞穴。作为创作者,你可以按照导演脑海中的特定轨迹移动摄像机,然后从 Marble 导出为视频。
By the way, in Marble we could allow people to export in video form. So you could go into a world. Let's say it's a hobbit cave. You can, especially as a creator, have a specific way of moving the camera in a trajectory in the director's mind, and then you can export that from Marble into a video.
创造这样的东西需要什么?团队有多大?你们用多少 GPU?有什么可以分享的吗?
What does it take to create something like this? How big is the team? How many GPUs are you working with? Anything you can share?
这需要大量的脑力。我们说的是每个大脑 20 瓦。从这个角度看,数字很小,但实际上是五亿年的进化才赋予我们这种能力。我们现在有大约 30 人的团队,主要是研究人员和研究工程师,但也有设计师和产品人员。我们坚信要创建一家以空间智能深度技术为核心的公司,同时也在构建严肃的产品。所以我们有研发和产品化的整合。当然,我们使用了大量的 GPU。
It takes a lot of brain power. We talk about 20 watts per brain. From that point of view, it's a small number, but it's actually an incredible half billion years of evolution to give us that power. We have a team of about 30 people now, predominantly researchers and research engineers, but we also have designers and product people. We really believe in creating a company anchored in the deep tech of spatial intelligence, but we are building serious products. So we have this integration of R&D and productization. And of course we use a ton of GPUs.
我很高兴听到这些。祝贺发布。我知道这是一个巨大的里程碑,付出了很多努力。所以我想向你和你的团队表示祝贺。
I'm so happy to hear. Well, congrats on the launch. I know this is a huge milestone. I know this took a ton of work. So I just want to say congrats to you and your team.
我们来谈谈你的创始人经历。你创办这家公司多久了?几年前?
Let me talk about your founder journey. You started this company how many years ago? A couple of years ago?
哦,一年前。18 个月。
Oh, a year ago. 18 months.
好的。有什么事情是你希望自己在开始之前就知道的,可以悄悄告诉 18 个月前的自己?
Okay. What's something you wish you knew before you started that you could whisper into the ear of Fei-Fei of 18 months ago?
我一直希望自己能预知技术的未来。我认为这是我们创始优势之一:我们比大多数人更早看到未来。但尽管如此,未知和即将到来的事物仍然令人兴奋和惊叹。但我知道你问这个问题不是为了技术未来。你可能更感兴趣的是,我 20 岁时并没有创办这种规模的公司。我 19 岁时开了一家干洗店,但规模小得多。后来我创立了 Google Cloud AI 和斯坦福的一个研究所,但那些是不同的事情。我觉得作为创始人,我比可能 20 岁的创始人更有准备应对艰辛的旅程。但我仍然感到惊讶,有时甚至偏执,因为 AI 领域的竞争异常激烈,无论是模型技术本身还是人才。当我创办公司时,我们并没有听说过某些人才要价那么高的故事。所以这些事情一直让我惊讶,我必须保持高度警惕。
I continue to wish I knew the future of technology. I think that's one of our founding advantages: we see the future earlier than most people. But still, this is so exciting and amazing, what's unknown and what's coming. But I know the reason you're asking is not about the future of technology. You're probably more interested in the fact that I did not start a company of this scale at 20 years old. I started a dry cleaner when I was 19, but that's a smaller scale. Then I founded Google Cloud AI and an institute at Stanford, but those are different beasts. I felt I was a little more prepared as a founder for the grinding journey compared to maybe 20-year-old founders. But I'm still surprised and sometimes paranoid about how intensely competitive the AI landscape is, from the model technology itself to talent. When I founded the company, we did not have these incredible stories of how much certain talents would cost. So these are things that continue to surprise me, and I have to be very alert.
所以你所说的竞争是人才竞争,以及事情发展的速度。
So the competition you're talking about is the competition for talent, the speed at which things are moving.
是的。
Yeah.
你提到了这一点,我想再谈一下:纵观你的职业生涯,你参与了所有那些促成今天许多突破的主要人类集合。显然我们谈到了 ImageNet,斯坦福也是很多工作发生的地方,Google Cloud 也是很多突破发生的地方。
You mentioned this point that I want to come back to: if you look over the course of your career, you were at all of the major collections of humans that led to so many breakthroughs today. Obviously we talked about ImageNet, also Stanford is where a lot of the work happened, Google Cloud where a lot of the breakthroughs happened.
是什么把你带到那些地方的?对于那些希望推进职业生涯、站在未来中心的人来说,有没有一条主线,把你从一个地方带到另一个地方,带入那些团队,这可能对人们有帮助?
What brought you to those places? For people looking for how to advance in their career, be at the center of the future, is there a throughline of what pulled you from place to place and into those groups that might be helpful for people to hear?
是的,这其实是个很好的问题,Lenny,因为我确实思考过。显然我们谈过好奇心和热情把我带到了 AI 领域。那更像是一个科学上的北极星,对吧?我不在乎 AI 是不是个东西。所以那是其中一部分。但我最终如何选择了我工作的特定地方,包括创办 World Labs,我认为我非常感激自己,或者也许是我父母的基因。我是一个在智力上非常无畏的人。我不得不说,当我招聘年轻人时,我会寻找这一点,因为我认为如果一个人想要有所作为,这是一个非常重要的品质。当你想要有所作为时,你必须接受你在创造新事物或深入人们未曾做过的新事物。如果你有这种自我意识,你几乎必须允许自己无畏和勇敢。所以当我来到斯坦福时,在学术界,我离所谓的终身教职很近,那是在普林斯顿永远拥有工作。但我选择来到斯坦福,因为我爱普林斯顿,那是我的母校。只是那时斯坦福有非常出色的人,硅谷生态系统也非常棒,所以我愿意冒风险重新开始我的终身教职计时,成为 SAIL 的第一位女性主任。实际上,相对而言,我当时是一位非常年轻的教员,我想这么做是因为我在乎那个社区。我没有花太多时间思考所有失败的情况。显然,我很幸运有更资深的教员支持我,但我只是想有所作为。然后去谷歌也是类似的。我想和 Jeff Dean、Geoff Hinton 以及所有那些了不起的人一起工作。所以 World Labs 也一样。我有这种热情,我也相信有相同使命的人可以做出不可思议的事情。这就是它如何指引我的人生。我不会过度思考所有可能出错的事情,因为那太多了。
Yeah, this is actually a great question, Lenny, because I do think about it. Obviously we talked about curiosity and passion that brought me to AI. That is more a scientific northstar, right? I did not care if AI was a thing or not. So that was one part. But how did I end up choosing the particular places I work in, including starting World Labs, is I think I'm very grateful to myself or maybe to my parents' genes. I'm an intellectually very fearless person. And I have to say when I hire young people, I look for that because I think that's a very important quality if one wants to make a difference. When you want to make a difference, you have to accept that you're creating something new or diving into something new that people haven't done. And if you have that self-awareness, you almost have to allow yourself to be fearless and to be courageous. So when I came to Stanford, in the world of academia, I was very close to this thing called tenure, which is having the job forever at Princeton. But I chose to come to Stanford because I love Princeton, it's my alma mater. It's just at that moment there are people who are so amazing at Stanford and the Silicon Valley ecosystem was so amazing that I was okay to take a risk of restarting my tenure clock, going to become the first female director of SAIL. I was actually relatively speaking a very young faculty at that time and I wanted to do that because I care about that community. I didn't spend too much time thinking about all the failure cases. Obviously, I was very lucky that the more senior faculty supported me, but I just wanted to make a difference. And then going to Google was similar. I wanted to work with people like Jeff Dean, Geoff Hinton, and all these incredible people. So the same with World Labs. I have this passion and I also believe that people with the same mission can do incredible things. So that's how it guided my life. I don't overthink all possible things that can go wrong because that's too many.
我觉得这其中的一个重要元素是不关注负面,更多地关注人和使命。什么让你兴奋?
I feel like that's an important element of this is not focusing on the downside, focusing more on the people, the mission. What gets you excited?
我确实想对所有 AI 领域的年轻人才、工程师和研究人员说一句话,因为你们中有些人申请了 World Labs。我感到非常荣幸你们考虑了 World Labs。我确实发现今天很多年轻人在决定工作时会考虑方程式的每一个方面。也许那是他们想要的方式。但有时我确实想鼓励年轻人关注重要的事情,因为我发现自己和求职者谈话时总是处于指导模式。不一定是招聘或不招聘,只是指导模式。当我看到一个了不起的年轻人才过度关注考虑工作的每一个微小维度时,也许最重要的事情是你的热情在哪里?你是否与使命一致?你是否相信并信任这个团队?然后专注于你能产生的影响以及你能与之共事的工作和团队。
I do want to say one thing to all the young talents in AI, the engineers, the researchers out there, because some of you apply to World Labs. I feel very privileged you considered World Labs. I do find many of the young people today think about every single aspect of an equation when they decide on jobs at some point. Maybe that's the way they want to do it. But sometimes I do want to encourage young people to focus on what's important because I find myself constantly in mentoring mode when I talk to job candidates. Not necessarily recruiting or not recruiting, but just in mentoring mode. When I see an incredible young talent who is overfocusing on every minute dimension and aspect of considering a job, when maybe the most important thing is where's your passion? Do you align with the mission? Do you believe and have faith in this team? And just focus on the impact you can make and the kind of work and team you can work with.
是的,这很难。现在对 AI 领域的人来说很难。有太多东西冲着他们来,太多新闻,太多事情发生,太多错失恐惧症。
Yeah, it's tough. It's tough for people in the AI space now. There's so much at them, so much news, so much happening, so much FOMO.
确实如此。
That's true.
我能看到压力。所以,我认为那个建议非常重要。就像什么才能真正让你在工作中感到满足,而不仅仅是哪家公司增长最快?谁会赢?我不知道。我想确保我问你关于你今天在斯坦福 HAI 的工作。我想是 HAI,以人为本的 AI 研究所。
I could see the stress. And so, I think that advice is really important. Just like what will actually make you feel fulfilled in what you're doing, not just where's the fastest growing company? Where's the who's going to win? I don't know. I want to make sure I ask you about the work you're doing today at Stanford at the HAI. I think it's HAI, Human-Centered AI Institute.
是的,HAI,以人为本的 AI 研究所,是由我和一组教员如 John Hennessy 教授、James Landay 教授、Chris Manning 教授在 2018 年共同创立的。我当时实际上在谷歌度过最后一个学术休假。这对我来说是一个非常重要的决定,因为我本可以留在工业界,但我在谷歌的时间教会了我一件事:AI 将是一项文明级的技术。我意识到这对人类有多么重要,以至于我在 2018 年写了一篇《纽约时报》文章,谈论需要有一个指导框架来开发和应用 AI,这个框架必须植根于人类福祉,即以人为本。我觉得斯坦福,作为硅谷中心的世界顶尖大学,孕育了从英伟达到谷歌等重要公司,应该成为思想领袖,创建这个以人为本的 AI 框架,并在我们的研究、教育、政策和生态系统工作中体现出来。所以我创立了 HAI。六七年之后,它已经成为世界上最大的 AI 研究所,从事以人为本的研究、教育、生态系统拓展和政策影响。它涉及斯坦福所有八个学院的数百名教员,从医学到教育到可持续发展到商业到工程到人文到法律。我们支持研究人员,特别是在跨学科领域,从数字经济到法律研究到政治科学到新药发现,到超越 Transformer 的新算法。我们还非常重视政策,因为当我们启动 HAI 时,我意识到硅谷不与华盛顿特区或布鲁塞尔或世界其他地区对话,而鉴于这项技术的重要性,我们需要让每个人都参与进来。所以我们创建了多个项目,从国会训练营到 AI 指数报告到政策简报,我们特别参与了政策制定,包括倡导一项国家 AI 研究云法案,该法案在特朗普第一届政府期间通过,并参与了州级 AI 监管讨论。所以我们做了很多,我继续担任领导者之一,尽管我在运营上参与少了很多,因为我不仅关心我们创造这项技术,还关心我们以正确的方式使用它。
So yes, HAI, Human-Centered AI Institute, was co-founded by me and a group of faculty like Professor John Hennessy, Professor James Landay, Professor Chris Manning back in 2018. I was actually finishing my last sabbatical at Google. And it was a very important decision for me because I could have stayed in industry but my time at Google taught me one thing: AI is going to be a civilizational technology. And it dawned on me how important this is to humanity, to the point that I actually wrote a piece in the New York Times that year 2018 to talk about the need for a guiding framework to develop and to apply AI, and that framework has to be anchored in human benevolence, is human-centeredness. And I felt that Stanford, one of the world's top universities in the heart of Silicon Valley that gave birth to important companies from Nvidia to Google, should be a thought leader to create this human-centered AI framework and to actually embody that in our research, education, and policy, and in ecosystem work. So I founded HAI. Fast forward after six, seven years, it has become the world's largest AI institute that does human-centered research, education, ecosystem outreach, and policy impact. It involves hundreds of faculty across all eight schools at Stanford, from medicine to education to sustainability to business to engineering to humanities to law. And we support researchers especially in the interdisciplinary area, from digital economy to legal studies to political science to discovery of new drugs, to new algorithms that go beyond transformers. We also actually put a very strong focus on policy because when we started HAI, I realized that Silicon Valley did not talk to Washington DC or Brussels or other parts of the world, and given how important this technology is, we need to bring everybody on board. So we created multiple programs from congressional boot camp to AI Index report to policy briefing, and we especially participated in policymaking, including advocating for a national AI research cloud bill that was passed in the first Trump administration, and participating in state-level regulatory AI discussions. So there's a lot we did, and I continue to be one of the leaders even though I'm much less involved operationally because I care not only that we create this technology but we use it in the right way.
哇。我之前不知道你做的所有这些其他工作。你说话时,我想起查理·芒格有句话:把一个简单的想法非常认真地对待。
Wow. I was not aware of all that other work you were doing. As you were talking, I was reminded Charlie Munger had this quote, take a simple idea and take it very seriously.
我觉得你在很多不同的方面都做到了这一点,并且坚持了下来,这些年来你在这么多方面产生的影响令人难以置信。我要跳过快速问答环节,直接问你最后一个问题。你还有什么想分享的吗?还有什么想留给听众的吗?
I feel like you've done that in so many different ways and stayed with it, and it's unbelievable the impact that you've had in so many ways over the years. I'm going to skip the lightning round and just ask you one last question. Is there anything else that you wanted to share? Anything else you want to leave listeners with?
我对 AI 感到非常兴奋,Lenny。我想回答一个我在世界各地旅行时每个人都会问我的问题:如果我是一名音乐家、一名教师、一名中学老师、一名护士、一名会计师、一名农民,我在 AI 中是否有角色,还是 AI 只会接管我的生活或工作?我认为这是 AI 最重要的问题。我发现,在硅谷,我们往往不会与像我们一样或不像我们的人推心置腹地交谈,而是像我们所有人一样,我们只是抛出诸如无限生产力、无限休闲时间或无限权力之类的词。但归根结底,AI 是关于人的。当人们问我这个问题时,答案是响亮的「是」。每个人在 AI 中都有角色。这取决于你做什么以及你想要什么。但任何技术都不应剥夺人的尊严,人的尊严和自主权应该处于每项技术的开发、部署以及治理的核心。所以,如果你是一位年轻的艺术家,你的热情在于讲故事,那就拥抱 AI 作为工具。事实上,拥抱 Marble。我希望它成为你的工具。因为你讲故事的方式是独一无二的,世界仍然需要它。但如何讲述你的故事,如何使用最不可思议的工具以最独特的方式讲述你的故事,这一点很重要,而且那个声音需要被听到。如果你是一位即将退休的农民,AI 仍然重要,因为你是一名公民。你可以参与你的社区。你应该在 AI 如何使用、如何应用方面拥有发言权。你与人们一起工作,你可以鼓励所有人使用 AI 来让生活更轻松。如果你是一名护士,我希望你知道,至少在我的职业生涯中,我在医疗保健研究方面做了很多工作,因为我觉得我们的医疗工作者应该得到 AI 技术的极大增强和帮助,无论是智能摄像头提供更多信息,还是机器人辅助,因为我们的护士工作过度、过度疲劳。随着社会老龄化,我们需要更多的帮助来照顾人们。所以 AI 可以扮演这个角色。我只想说,即使像我这样的技术人员也真诚地认为每个人在 AI 中都有角色,这一点非常重要。
I'm very excited by AI, Lenny. I want to answer one question that when I travel around the world, everybody asks me: if I'm a musician, if I'm a teacher, middle school teacher, if I'm a nurse, if I'm an accountant, if I'm a farmer, do I have a role in AI, or is AI just going to take over my life or my work? And I think this is the most important question of AI. I find that in Silicon Valley, we tend not to speak heart-to-heart with people like us and not like us in Silicon Valley, but like all of us, we tend to just toss around words like infinite productivity or infinite leisure time or infinite power or whatever. But at the end of the day, AI is about people. And when people ask me that question, it's a resounding yes. Everybody has a role in AI. It depends on what you do and what you want. But no technology should take away human dignity, and human dignity and agency should be at the heart of the development, the deployment, as well as the governance of every technology. So if you are a young artist and your passion is storytelling, embrace AI as a tool. In fact, embrace Marble. I hope it becomes a tool for you. Because the way you tell your story is unique, and the world still needs it. But how you tell your story, how you use the most incredible tool to tell your story in the most unique way is important, and that voice needs to be heard. If you're a farmer near retirement, AI still matters because you're a citizen. You can participate in your community. You should have a voice in how AI is used, how AI is applied. You work with people that you can encourage all of you to use AI to make life easier for you. If you're a nurse, I hope you know that at least in my career, I have worked so much in healthcare research because I feel our healthcare workers should be greatly augmented and helped by AI technology, whether it's smart cameras to feed more information or robotic assistance because our nurses are overworked, over fatigued. And as our society ages, we need more help for people to be taken care of. So AI can play that role. So I just want to say that it's so important that even a technologist like me is sincere about that everybody has a role in AI.
多么美好的结束方式。这又回到了我们开始的地方,即 AI 在我们生活中会做什么取决于我们,并承担个人责任。最后一个问题,人们在哪里可以找到 Marble?他们可以去哪里?也许如果想加入 World Labs 的话。网站是什么?人们去哪里?
What a beautiful way to end it. Such a tie back to where we started about how it's up to us and take individual responsibility for what AI will do in our lives. Final question, where can folks find Marble? Where can they go? Maybe try to join World Labs if they want to. What's the website? Where do people go?
嗯,World Labs 的网站是 www.worldlabs.ai,你可以在那里找到我们的研究进展。我们有技术博客。你可以在那里找到 Marble 产品。你可以登录。你可以在那里找到我们的招聘链接。我们在旧金山。我们喜欢与世界上最优秀的人才合作。
Well, World Labs website is www.worldlabs.ai and you can find our research progress there. We have technical blogs. You can find Marble the product there. You can sign in there. You can find our job posts link there. We're in San Francisco. We love to work with the world's best talents.
太棒了。Fei-Fei,非常感谢你来到这里。
Amazing. Fei-Fei, thank you so much for being here.
谢谢你,Lenny。
Thank you, Lenny.
再见,各位。非常感谢你们的收听。如果你觉得这期节目有价值,可以在 Apple Podcasts、Spotify 或你最喜欢的播客应用上订阅。另外,请考虑给我们评分或留下评论,这真的能帮助其他听众找到这个播客。你可以在 lennispodcast.com 找到所有往期节目或了解更多关于这个节目的信息。下期节目再见。
Bye, everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast. You can find all past episodes or learn more about the show at lennispodcast.com. See you in the next episode.