Why Start a Company Now? Dr. Fei-Fei Li on Spatial Intelligence
打开互动全文版(中英对照 + 朗读 + 问答)→李飞飞博士讨论她的新公司 World Labs、空间智能以及 3D 世界模型对 AI 的重要性。
Dr. Fei-Fei Li discusses her new company World Labs, spatial intelligence, and the importance of 3D world models for AI.
听众朋友们,欢迎回到 No Priors。今天的嘉宾是李飞飞博士,她是计算机视觉和深度学习领域的先驱。她创建了 ImageNet,这个开创性的数据集推动了深度学习革命。飞飞是斯坦福大学教授,也是斯坦福以人为中心的人工智能研究所的联合主任。她还曾领导谷歌云的 AI 部门,为国际政策制定者提供建议,并最近联合创立了 World Labs,一家致力于开发空间智能 AI 的公司。飞飞,感谢你今天参加我们的节目。
Hi listeners and welcome back to No Priors. Today's guest is Dr. Fei-Fei Li, a pioneer in computer vision and deep learning. She created ImageNet, the groundbreaking dataset that helped spark the deep learning revolution. Fei-Fei is a Stanford professor and the co-director of the Stanford Institute for Human-Centered AI. She's also led AI at Google Cloud, advised international policy makers, and recently co-founded World Labs, a company dedicated to developing spatially intelligent AI. Fei-Fei, thank you for joining us today.
谢谢邀请。这会很有趣。
Well, thanks for inviting me. This is going to be fun.
在过去的二十年里,你在科学和政策方面做出了非凡的贡献。我先从最大的问题开始:为什么现在创办一家公司?
So, you have made extraordinary contributions to science and policy over the past two decades. I'll start with the biggest question: why start a company now?
因为在我内心深处,我想去构建。我认为这是一个关键、有趣且激动人心的时刻,去构建一些每个人都能使用的非凡技术。我非常相信空间智能和那种 3D 世界模型,它们可以赋能很多人和很多应用场景。我认为这将非常令人兴奋,而且我可以与一群非常出色的年轻技术专家一起做到这一点。
Because in my heart I want to build. I see this as such a critical and fun and exciting moment to build some extraordinary technology that everybody can use. And I believe so much in spatial intelligence and the kind of 3D world models that can empower so many people as well as so many use cases. I think it's going to be really exciting, and I can do that with an extraordinarily brilliant group of young technologists.
我想回到你合作的人身上,因为我认识你的一些联合创始人,之前我拼命想说服他们创办一家公司,然后他们说:‘哦不,我们现在和飞飞有了更大的使命。’什么是空间智能?你能为更广泛的听众定义一下吗?
I want to come back to the people you're working with because I know some of your co-founders and was trying to convince them desperately to start a company a while back, and then they were like, 'Oh no, we have a bigger mission now with Fei-Fei.' What is spatial intelligence? Can you define it for a broader audience?
对我来说,空间智能是理解、推理、交互和生成 3D 世界的能力。因为我们的世界,无论你如何投射,本质上都是 3D 的。它是 3D 的,因为物理上是 3D 的,而数字上如果有真正的 3D 表示,那么我们可以更容易地实现很多事情,无论是设计、创造、导航、模拟还是体验 AR/VR。所有这些都是空间智能的一部分。真正让我兴奋的是,人类拥有空间智能;它是我们核心智能能力的一部分。动物也有空间智能。整个进化过程与空间智能的进化深深交织在一起。所以它非常基础;没有空间智能,AI 将是不完整的。
Spatial intelligence to me is the ability to understand, reason, interact, and generate 3D worlds. Because our world, fundamentally, no matter how you project it, is 3D. It's 3D because physically it's 3D, and digitally if there is a true 3D representation, then we can make a lot of things happen more easily, whether it's designing, creation, navigation, simulation, or experiencing AR/VR. All this is part of spatial intelligence. What really excites me is that humans have spatial intelligence; it's part of our core intelligent capabilities. Animals have spatial intelligence. The entire journey of evolution is deeply intertwined with the evolution of spatial intelligence. So it's so fundamental; without spatial intelligence, AI would be incomplete.
这如何转化为你在公司所做的事情?或者你能分享一些关于这对你正在构建的东西意味着什么吗?
How does that translate into what you're doing with your company? Or is there anything you can share in terms of what that means relative to what you're building?
是的。我们正在解决 AI 中最难的问题之一,即构建根本上 3D 的世界模型。因为一旦你解决了这个问题,你就可以解锁很多空间智能问题。所以据我们所知,我们是第一家解决这个 3D 生成基础模型问题的公司。
Yeah. So we're tackling one of the hardest problems in AI, which is actually making world models that are fundamentally 3D. Because once you can crack that problem, you can unlock a lot of spatial intelligence problems. So we are the first company we know of that is solving this 3D generation foundation model problem.
我有很多问题。但既然你描述这是 3D 对理解世界的关键性,这是否意味着你认为 World Labs 将创造的世界模型,或者学术界或公司的其他人将创造的,有一天会现实地准确,比如代表物理和对世界的理解,我们可以用它做更多的事情?
I have many questions. But since you're describing this as the 3D's criticality to understanding the world, does that imply that you feel the world models that World Labs will create, or others in academia or companies will create, will someday be realistically accurate, like represent physics and understanding of the world that we can do many more things with?
是的,它应该现实地准确或合理。所以你可以创造一个奇幻的世界,但它应该是合理的,因为它的几何和物理需要合理。这对空间智能来说是根本。
Yeah, it should be realistically accurate or plausible. So you can create a fantastical world, but it should be plausible because the geometry and the physics of it need to be plausible. And that is fundamental to spatial intelligence.
这是否意味着你从神经科学的角度对视觉智能与大型语言模型和文本智能相比有多基础有特定的观点?
Does that imply you have a particular point of view from a neuroscience perspective of how fundamental visual intelligence is versus large language models and textual intelligence?
我确实有。我认为从神经和认知科学的角度来看,空间智能是一个进化必须为动物解决的非常困难的问题。有趣的是,我认为动物在一定程度上解决了它,但没有完全解决。这是最难的问题之一,因为动物必须解决什么问题?动物必须进化出收集光的能力,主要通过我们称之为眼睛的东西。然后通过眼睛的收集,它必须以某种方式在脑海中重建一个 3D 世界,以便它们能够导航和做事,当然它们可以互动。对于人类来说,我们在操作方面是最有能力的动物。我们可以做很多事情,所有这些都是空间智能。对我来说,这植根于我们的智能。有趣的是,即使在动物中,这也不是一个完全解决的问题。例如,对于人类,如果我让你现在闭上眼睛,画出或构建你周围环境的 3D 模型,这并不容易。我们没有那么大的能力生成极其复杂的 3D 模型,除非我们经过训练。我们中的一些人,无论是建筑师、设计师还是只是经过大量训练和有天赋的人,这是一件难事。想象一下,你可以在指尖上更容易地做到这一点,并允许更流畅的交互性和可编辑性。这对人们来说将是一个完全不同的世界,没有双关语。
I actually do. I think from a neural and cognitive science point of view, spatial intelligence is a really hard problem that evolution has to solve for animals. And what's really interesting is I think animals have solved it to an extent but not fully solved it. It's one of the hardest problems because what is the problem an animal has to solve? Animals have to evolve the capability of collecting light in something we call eyes mostly. And then with that collection of eyes, it has to reconstruct a 3D world in their mind somehow so that they can navigate and do things, and of course they can interact. For humans, we're the most capable animal in terms of manipulation. We can do a lot of things, and all this is spatial intelligence. To me, that's rooted in our intelligence. What is interesting is it's not a fully solved problem even in animals. For example, for humans, if I ask you to close your eyes right now and draw out or build a 3D model of the environment around you, it's not that easy. We don't have that much capability to generate extremely complicated 3D models until we get trained. There are some of us, whether they're architects or designers or just people with a lot of training and talent, and that's a hard thing to do. And imagine you do it at your fingertip much more easily and allow much more fluid interactivity and editability. That would just be a whole different world for people, no pun intended.
还有其他像空间智能这样的大领域,你觉得从模型角度来看还没有得到充分发展吗?或者还有其他缺失的空白,你认为在我们构建这个 AI 未来的过程中,我们应该随着时间的推移关注或人们应该构建出来?我只是想知道,除了 3D 和世界生成,还有其他类似的大问题吗?因为感觉我们随着时间的推移解决了一些大问题,还有其他问题我们正在努力。我们基本上在解决语言。我会说语言在很大程度上已经解决了,而 3D 对我来说和语言一样关键和困难。那么还有什么没有解决呢?
Are there other big areas like spatial intelligence that you feel haven't been as developed as they could be from a model perspective? Or other missing gaps that you think in general as we build this AI future we should focus on over time or people should build out? I was just wondering, in addition to 3D and world generation, are there other big problems like that? Because it feels like there are a few big things that we've solved for over time and other things we're working on. We're sort of solving language. I would say language is solved to a huge extent, and 3D to me is as critical and difficult as language. So what else that's not solved?
我的意思是,整个情感智能领域是我甚至不知道如何开始解决的问题。我知道很多人没有解决它。所以那是 AGI 实现的时候。我可以告诉你,这方面的训练数据不会来自硅谷的人。
I mean the entire space of emotional intelligence is something that I don't even know how to begin to solve. I know a lot of people who haven't solved it. So that's when AGI is achieved. I can tell you the training data for that is not going to come from Silicon Valley people.
不要低估硅谷。
Don't underestimate Silicon Valley.
是的。所以,我会把自己放在这个桶里,但我认为我们可能需要更广泛的人群。是的。不,我同意。但老实说,这是三大桶。我不知道。你觉得呢,Elon 和 Sarah?
Yeah. So, I'll put myself in this bucket, but I think we probably need a broader set of people. Yeah. No, that I agree. But these are the three big buckets to be honest. That's I don't know. What do you think Elon and Sarah?
我认为这在很大程度上取决于你在每个模型中封装了什么。
I think it depends a lot on what you encapsulate in each model.
所以我同意你的框架,关于那三个方面,然后某些东西比如空间智能,我猜也涉及不同类型的物理模拟和世界模拟。这些是很大的领域,我觉得很多人没有在研究,而我认为它们非常有趣且重要。这有宏观和微观尺度。微观尺度最终会变成材料科学和其他非常不同的东西,比如更分子建模的领域。而且也在某种程度上超出了当前 AI 的定义,但我认为它们会被赋能。当然还有机器人学,但机器人学很大程度上是一个系统集成问题,就像……即使你看动物,也不仅仅是大脑中的计算。很多这些东西在空间智能方面似乎更加分布式,相对于动物拥有的特定系统,在某些情况下并不像人们想象的那么集中。所以开始思考这些跨有机体的分布式智能模型与中枢神经系统相比,是非常有趣的。但我觉得这是非常有趣的东西。你也在机器人学和物理智能领域做过工作。我想到了机器人基础模型和驱动的数据层级。人们想用视频,因为那是我们可以获得的。关于模拟以及今天能从模拟中得到多少,有一个大问题。也许人们看不到未来可用的质量和物理模拟。然后还有接近具身的,比如不同形式的遥操作,以及具身数据收集。这是你心目中的层级吗?还是你认为人们低估了模拟和世界模型对未来的作用?
So I agree with your framework in terms of those three and then certain things like spatial intelligence, I'm assuming also delves into different types of physics simulation and simulations of the world. Those are big areas that I think a lot of people aren't working on that I think are really interesting or important. There's the macro and the micro scale of that. The micro scale eventually becomes material sciences and other very different types of things from what you're talking about where it's more molecular modeling. And also somewhat goes out the current definition of AI, which I do think they'll be empowered by. Of course there's robotics, but robotics is very much a system integration problem as much as a... even if you look at animals, it's not just the compute in the brain per se. A lot of these things seem much more distributed in terms of spatial intelligence relative to specific systems that animals have, and in some cases it's not as centralized as one would think. So it's very interesting to start thinking in terms of those models of more distributed intelligence across an organism versus a CNS. But yeah, I think it's very interesting stuff. You've also done work in this field of robotics and physical intelligence. I think of the data hierarchy for robotics foundation models and actuation. People want to use video because that is what is available to us. There's a big question on simulation and how much you can get from that today. Perhaps people do not see the future of the quality and the physics that are going to be available to us. And then there's close to embodied like different forms of teleoperation and then embodied data collection. Is that the hierarchy you have in your mind, or do you think people underestimate simulation and world models for the future?
是的,好问题。首先,正如你所说,我确实在做机器人学,特别是在斯坦福的实验室。我毫不怀疑人类将进入一个与机器人共存的时代,而且机器人并不一定是人形的。机器人有各种形态和形状。实际上,几年前我的实验室写了一篇很有趣的关于形态智能的论文,其中智能体的形态可以通过优化它们试图完成的任务而改变。所以我们应该比仅仅考虑人形机器人更有想象力。话虽如此,如何训练机器人……你提到了这些数据,有人称之为数据金字塔或数据蛋糕之类的。我同意。我认为这将是多种不同形式数据的混合。我也认为模拟被低估了。实际上,很多专家和领域内的人并没有低估它。如果你看很多机器人公司,他们正在研究模拟和合成数据。我还认为我们必须意识到,与语言模型甚至空间智能基础模型不同,机器人学是一个高度多模态的系统。在我看来,真正被低估的是触觉。有很多,特别是如果我们想做操作而不仅仅是导航。我认为触觉数据以及将触觉真正整合到视觉、感知和空间数据中的能力是绝对关键的。
Yeah, great question. First of all, as you say, I do work in robotics, especially in my lab at Stanford. I have no doubt that humanity will move into an age where we cohabit with robots, and also the world robot is not humanoid per se. Robots taking all kinds of forms and shapes. Actually, a few years ago, my lab wrote a really fun paper about morphological intelligence, where the morphology of an agent actually can change by optimizing the tasks they're trying to achieve. So we should be a little more imaginative than just humanoids. Having said that, how to train robots... you mentioned this whole data, some people call it data pyramids or data cakes or whatever. I agree. I think it's going to be a hybrid of many different forms of data. I also think simulation is underrated. Actually, it's not underrated by a lot of experts and people in the field. If you look at a lot of robotics companies, they are working on simulation and synthetic data. I also think we have to be aware that unlike language models or even unlike spatial intelligence foundation models, robotics is a highly multimodal system. What is truly underappreciated in my opinion is haptics. There's so much, especially if we want to do manipulation not just navigation. I think haptics data and the ability to really integrate haptics into vision and perception and spatial data is absolutely critical.
你提到的一点我觉得非常有趣,就是机器人可能适应或采用多少种不同的形态。关于潜在的未来,人们提出了两个相反的观点。一种观点是从供应链角度以及管理制造和规模的角度来看,你将会有少得多的形态因素。另一种观点是专业化的经济价值非常高,因此随着我们走向机器人驱动的未来,将会有成千上万种不同的形态因素。你对这两个观点之间我们可能落在哪里有什么看法?
One thing that you said that I thought was really interesting is how many different morphological forms a robot may adapt or adopt. There are two counter arguments people make in terms of the potential future. One argument is that from a supply chain perspective and managing builds and scale of manufacturing, you're going to have many fewer form factors. The other argument is the economic value of specialization is very high and therefore there'll be thousands and thousands of different form factors as we move to a robot-driven future. Do you have a point of view on where we're likely to land between those two viewpoints?
我认为我们将通过梯度下降来优化生产力和效率。我的假设是,不同任务的需求如此之大,以至于只有很少的形态或坚持一种形态是能源效率低下的,很多任务可以由更节能的形态因素完成,也应该由它们完成。举一个极端且琐碎的例子:如果我们把机器人放在水下,它们不应该是人形的。它们最好是鱼形的,对吧?想想能源效率。飞行也是如此。我不认为人形是……我们的飞机正变得越来越像机器人。所以我确实认为会有多样性。
I think we're going to gradient descend into optimization of productivity and efficiency. My hypothesis is that the requirements of different tasks are so vast that having very few forms or sticking with one form is energy inefficient, and a lot of tasks can be done and should be done by much more energy efficient form factors. Just an extreme and trivial example: if we put robots underwater, they should not be in the shape of humans. They better be in the shape of fish, right? Just think about energy efficiency. And the same with flying. I don't think human form is... our airplanes are becoming more and more robots. And so I do think there's going to be diversity.
机器人学是未来的一个潜在应用。你首先是一位科学家,但你也知道,做过推特董事会,参与过初创公司。你能想象生成 3D 世界的近期商业应用是什么?
Robotics is one potential application for the future. You're a scientist first, but also you know, did the Twitter board, involved in startups. What are the near-term commercial applications that you can imagine for generating 3D worlds?
我相信创造力是一个极其令人兴奋的领域,人类可以通过 AI 和空间智能获得超能力。这里我类比软件工程。如果你看今天 LLM 在软件工程中的成功,包括像 Cursor 和 Windsurf 这样的应用,你看到的是 AI 和人类之间的大量协作,而且协作涉及不同技能水平。我认为创造力也会类似。无论我们谈论的是设计师、3D 艺术家、视觉特效艺术家,还是营销人才和游戏开发者,设计和创建 3D 空间的需求非常大。这从根本上来说是一个非常困难的问题,即使对于训练有素的人来说也是如此,所以如果我们做对了,拥有一个协作者将会非常有趣。所以我将创造力视为一个非常令人兴奋的领域。我还认为,我们等待的元宇宙或 XR、AR/VR 的很多东西是内容创作。我理解硬件本身需要继续发展,但我也认为软件……我们正在寻找内容创作,而这自然非常适合 3D 建模和生成式空间模型,这是另一个值得研究的领域。
I believe creativity is a vastly exciting area where humans can be superpowered by AI and by spatial intelligence. Here I draw an analogy with software engineering. If you look at today's success of LLMs in software engineering, including applications like Cursor and Windsurf and all that, what you see is a lot of collaboration between AI and humans, and the collaboration comes in different levels of skill sets. I think creativity will be similar. Whether we're talking about designers, 3D artists, VFX artists, or even marketing talents and game developers, there's so much need in designing and creating 3D space. This is fundamentally such a hard problem even for trained skilled people that having a collaborator will be extremely fun if we do it right. So I see creativity as an area that is really exciting. I also do think that a lot of what we're waiting for for metaverse or XR, AR/VR, is content creation. I understand hardware itself needs to continue to evolve, but I also think software... we're looking for content creation, and that lends itself so naturally to 3D modeling and generative spatial models, and that's another interesting area to look into.
你对世界模型是否是针对更通用智能体的可扩展强化学习的一个有趣答案,有强烈的观点吗?
Do you have a strong point of view on whether or not world models are an interesting answer to scalable RL for more generalizable agents?
我确实认为没有空间智能,AI 是不完整的。人类在 3D 世界中互动,在数字世界中我们需要各种互动。以设计为例:它在我们的脑海中深度优化美感、效率或其他,这自然适合强化学习场景。
I actually do think AI is not complete without spatial intelligence. Humans interact in 3D worlds, and in the digital world we need all kinds of interaction. Take design as an example: it's deeply optimizing for beauty, efficiency, or whatever in our mind's eye, and that lends itself naturally to RL settings.
设计和训练世界模型的最大挑战是什么?我想一个是数据:我们有图像和视频,但没有很多你正在构建的那种格式的 3D 世界。
What are the biggest challenges in designing and training world models? I imagine one is data: we have images and video, but not many 3D worlds in a format you're building.
数据绝对是一个挑战。要创建世界模型、3D 基础模型,我们需要更复杂的数据工程、获取、处理和合成。我很羡慕我的 NLP 大语言模型同事,互联网上的数据如此丰富;我们没有那种奢侈。另一个挑战是 3D 具有讽刺意味:每个人每天都在使用 3D,但它不像语言那样容易交付。语言是简单且主动的,不是被动消费。没有人醒来会说‘我就坐在这里看 3D’。这给产品化带来了挑战。
Data is absolutely a challenge. To create world models, 3D foundation models, we require more sophisticated data engineering, acquisition, processing, and synthesis. I am envious of my NLP LLM colleagues that data is so abundant on the internet; we don't have that luxury. Another challenge is that 3D is ironic: everyone uses 3D daily, yet it's not as easy a form factor to deliver as language. Language is easy and active, not passive consumption. Nobody wakes up and says 'I'm just going to sit here and watch 3D.' That creates challenges for productization.
你玩过《第二人生》之类的吗?
Were you ever a Second Life player or anything?
我不是游戏玩家,但我的孩子喜欢《我的世界》。
I'm not a gamer, but my kids love Minecraft.
有没有你想体验或想象的世界?
Is there a world you want to experience or imagine?
我很想看到我看不到的世界,比如放大到微观世界,或者进入发动机内部,甚至进入洗碗机内部。如果我们能创建任何东西的世界模型,这一切都可以虚拟实现。
I would love to see worlds I don't see, like zooming into microscopic worlds, or going inside an engine, or even being inside a dishwasher. All this can be done virtually if we manage to create world models of anything.
我想谈谈你的职业生涯。就在这之前,我问了 Andre Karpathy 应该问你什么,他说李飞飞在雄心和数据思考方面真的很神奇。问问她关于她的博士工作和与 Pietro 一起创建 101 数据集的事,因为这很有启发性。
I want to talk about your past career. Right before this I asked Andre Karpathy what to ask you, and he said Fei-Fei is really magic about ambition and thinking about data. Ask her about her PhD and the creation of the 101 dataset with Pietro, because it's instructive.
当一个学生比你更出名、成就更多时,总是最棒的事情。这让我非常自豪。我很惊讶他还记得我的博士工作。那要追溯到 2003 年左右。世界才刚刚触及互联网的表面,数据还不是什么大事。在计算机视觉中,我的博士工作是试图让物体识别工作——从图片中识别出猫、狗、微波炉、椅子。我们开始假设数据很重要,但我们不知道缩放定律。我们想要的只是数据来训练算法。作为一名博士生,你想毕业,Pietro 说‘整理一个数据集’。我想,‘我确实需要整理一个数据集,因为现有的每个数据集都太小了。’Pietro 和我讨论了 15 或 30 种不同的东西,然后他说了三位数 100。我知道从数学角度他是对的:要推动模型泛化,我们需要足够的数据。我在我的书《我看见的世界》中写了这个过程。我偶然发现了一本词典——我想是韦氏词典——上面有一些词的视觉描绘。我挑了其中的 101 个词。这让我的博士导师笑了,因为他说我只是想比他要求的做一个更多。我从谷歌下载图片,当时谷歌还很新,图片搜索很糟糕。我做了很多清理工作,以至于绝望地让我妈妈帮忙清理图片。我写了一个简单的界面;她不懂电脑,但至少会点击。所以她帮了我。
It's always the greatest thing when a student is more well-known and achieving so much more than you. It makes me so proud. I'm surprised he remembers my PhD work. It goes back to 2003ish. The world was barely scratching the surface of the internet, and data was not much of a thing. In computer vision, my PhD work was trying to get object recognition to work—calling out cats, dogs, microwaves, chairs from a picture. We began to hypothesize that data matters, but we had no idea about scaling laws. All we wanted was data to train algorithms. As a PhD student, you want to graduate, and Pietro said 'curate a dataset.' I thought, 'I do need to curate a dataset because every dataset out there is so tiny.' Pietro and I talked about 15 or 30 different things, and then he said the three-digit number 100. I knew he was right from a mathematical point of view: to push the model to generalize, we need enough data. I wrote about this in my book 'The Worlds I See.' I stumbled upon a dictionary—Webster's, I think—which had visual depictions of some words. I grabbed 101 of those words. That made my PhD advisor chuckle because he said I just wanted to do one more than he asked. I downloaded images from Google, which was new and terrible at that time. I had to do so much cleaning that I got desperate and asked my mom to help clean images. I wrote a simple interface; she didn't know computers but could click. So she helped.
你在 AI 领域拥有最传奇的职业生涯之一,你的许多学生也取得了巨大成就。回顾你的职业生涯,你会想到哪两三个时刻?
You've had one of the most storied careers in AI, and many of your students have gone on to do great things. What are two or three moments you think of when you look back on your career?
我认为 ImageNet 的创建是一个关键时刻。它展示了数据的力量,并引发了深度学习革命。另一个时刻是创立 AI4ALL,旨在使 AI 多样化。就个人而言,看到我的学生如 Andre Karpathy 成功是非常有成就感的。
I think the creation of ImageNet was a pivotal moment. It showed the power of data and led to the deep learning revolution. Another moment is the founding of AI4ALL, which aims to diversify AI. And personally, seeing my students like Andre Karpathy succeed is incredibly rewarding.
显然你在图像和视觉识别系统方面做了很多事,但我很好奇:回顾过去 20 年,你做了这么多,最让你印象深刻的是什么?
I mean obviously there's a lot of things that you did in terms of image and visual recognition related systems and all sort, but I'm just curious: when you think of the last 20 years, what stands out the most given everything that you've done?
哦,谢谢你的提问。当然,ImageNet 是由多个时刻组成的:从早期的挣扎和被告知我拿不到终身教职,到意识到亚马逊土耳其机器人来救场,到 AlexNet 获胜的时刻,再到几年前我在多伦多与杰夫·辛顿一起参加活动,他公开说那件事多么具有决定性,而且他几乎有点歉意,觉得 ImageNet 没有得到像神经网络那样的认可。所以那段旅程非常具有验证性。对科学家来说,验证不是关于认可或奖项。而是你做出了改变,比如那个没人相信的猜想,那个没人相信的假设,我们让它成真了。所以这是一条线索。
Oh, thank you for asking that question. Of course, ImageNet is one of those that consists of multiple moments: from the early struggles and being told I will not get tenure, to actually realizing Amazon Mechanical Turk comes to rescue, to the moment of AlexNet winning, and also to a couple of years ago I was at an event in Toronto with Geoff Hinton and he said publicly how that was so defining and he was almost a little bit apologetic that ImageNet was not as recognized as neural networks. So that journey is very validating. For scientists, the validation is not about recognition or awards. It's that you made a difference, like that conjecture that no one believed in, that hypothesis that no one believed in, we were able to make it happen. So that's one thread.
为了让商业界不熟悉的人了解,ImageNet 是一个大型数据集,包含数百万张标注图像,涵盖数千个类别,而不仅仅是 10 个,对吧?1500 万张标注图像。这导致了深度学习的惊人突破,特别是 AlexNet,以及该领域的许多进展。它推动了许多机器视觉的发展。我实际上记得在 2016 或 2017 年,我经常展示一张幻灯片,那是 AI 的历史。那时是 CNN 和 RNN,GANs 刚刚兴起,我把 ImageNet 和 AlexNet 作为标志性时刻之一,是真正定义 AI 进步的少数事件之一。显然现在 Transformer 也是其中一部分,也许还有扩散模型之类的,但那是一个巨大的突破。
Just to make sure for any people from the business world that are not familiar with it, ImageNet was a large-scale dataset with millions of labeled images across thousands of categories, not just 10, right? 15 million labeled images. That led to amazing breakthroughs in deep learning, in particular AlexNet, and lots of progress in the field. It drove a lot of machine vision forward. I actually remember in 2016 or 2017, I used to show a slide which was the history of AI. Back then it was CNNs and RNNs and just GANs were kind of going, and I had ImageNet and AlexNet as one of the seminal moments, a very small number of events that really defined AI progress. Obviously now we have Transformers as part of that and maybe diffusion models or something, but it was such a big breakthrough.
是的,谢谢。另一个我非常自豪的时刻实际上是安德烈和贾斯汀·约翰逊以及他们的论文。在我看来,那是语言和图像第一次通过为视觉世界添加标题和编写故事而融合。这对我来说意义重大有两个原因。第一个是,我真的以为,不骗你,在我博士结束时,我想如果我能活到 100 岁,那可能是我们能解决的问题,那就是图片的故事讲述。所以我进入职业生涯,作为助理教授的第一年,我想好吧,我要做 ImageNet 来解决物体识别,然后我要用我整个职业生涯的剩余时间来解决这个讲故事的问题。然后当安德烈和稍后的贾斯汀·约翰逊进入我的实验室时,那大约是 2013、2014 年,深度学习的开始。突然,当时的序列模型,LSTM,不是 Transformer 模型,但 LSTM 和 CNN 的结合,一下子打开了图像描述的工作。我和他们的工作与谷歌的一起率先问世。这真的让我非常自豪。我几乎经历了一场危机,就像我接下来的 70 年或 65 年要做什么?所以这真的很令人兴奋,这个领域发展得如此之快。
Yeah, thank you. Another moment I'm very proud of was actually Andrej and also Justin Johnson and their dissertations. It's where, in my opinion, the first time that language and images converged by captioning and writing stories of the visual world. It was significant for me for two reasons. One is that I literally thought, I kid you not, at the end of my PhD I thought if I can live till 100 years old, that was the problem we might be able to solve, which is storytelling of pictures. So I entered my career, my first year as assistant professor, thinking okay, I'm going to do ImageNet to solve object recognition, and then I'm going to spend the rest of my entire career solving this problem of storytelling. And then by the time Andrej and a little later Justin Johnson entered my lab, that was around 2013, 2014, the beginning of deep learning. Suddenly the combination of sequential model at that point, LSTM, not Transformer models, but LSTM and CNN, just blasted open the image captioning work. My work and theirs were the first together with Google's that was out of the door. That really made me so proud. I almost had a crisis, which is like what am I going to do for the rest of my 70 years or 65 years? So that was really exciting, how fast the field has evolved.
我能再问一个问题吗?因为你非常高效地取得了这些惊人的进展。我们之前私下聊过,你觉得在 AI 研究中,除了资金雄厚的大型企业实验室之外,拥有登月计划和创造力非常重要。你指出了几个来自学术界创造力和研究的时刻。你对人们有什么建议?是否仍然有机会,还是说从此以后都只是 100 亿美元的培训运行?
Can I ask you one more question about this? Just because you have made this amazing progress very efficiently. You and I have talked offline before about how you feel it's really important for there to be moonshots and creativity in AI research beyond very large funded corporate labs. You pointed to several moments that come from creativity and research in academia. What advice do you have for people about whether there's still opportunity for that, or it's all just 10 billion dollar training runs from here?
我唯一的建议,我在我的公司和实验室里仍然这么说,就是无所畏惧。我认为科学家、技术专家和企业家必须无所畏惧。你知道,最终你必须弄清楚,你是需要 100 亿美元的培训运行,还是来找莎拉要资金?可能两者都需要很多。或者你必须弄清楚,我不知道,数据。有时无所畏惧是一种非常有趣的立场,你有点妄想和疯狂,但又在某种程度上理性地大胆,它介于两者之间。因为如果你太理性,就不够勇敢。你没有识别出足够大的问题。但如果你完全疯狂,那么有很多事情可能会出错。所以要无所畏惧,要勇敢。对我来说,即使像我这样年纪,我也是这么感觉的。我创办了我的初创公司 World Labs,因为我想无所畏惧,解决空间智能这个问题,作为解决问题的一部分。
My singular advice, and I still say this in my company and my lab, is be fearless. I think scientists and technologists and entrepreneurs have to be fearless. You know, eventually you have to figure out, do you need 10 billion dollar runs or then you come to Sarah to ask for funding? Probably a lot for both. Or you have to figure out, I don't know, data. Sometimes fearless is this very interesting position where you're somewhat delusional and crazy but somewhat just rationally bold, and it kind of is in between. Because if you're too rational, it's not courageous enough. You're not identifying problems that are big enough. But if you're completely crazy, then there are so many things that can go wrong. So be fearless, be courageous. To me, that is how I feel even as old as I am. I started my startup World Labs because I want to be fearless and solve this problem of spatial intelligence as part of problem solving.
你长期以来与一些世界上最好的 AI 研究人员和工程师合作过。在你的公司背景下,你是如何看待这一点的?你想招聘什么样的人?目前有职位空缺吗?毫无疑问这是一个了不起的团队。我只是好奇你想增加什么样的人,以及你随着时间的推移是如何考虑的。
You've worked with some of the best AI researchers in the world over time and best engineers. How do you think about that in the context of your company? What sorts of people are you trying to hire? Are there open roles currently? Undoubtedly it's an amazing team. I'm just curious what sorts of folks you want to add and how you're thinking about that over time.
是的,我们有职位空缺,我们目前希望为公司招聘最优秀的工程师和产品思考者。所以,如果你是一名工程师、AI 研究人员或产品人才,热衷于加入最有才华的团队并解决这个问题,请加入我们。那么我们招聘什么样的人?首先,我们确实招聘思维多样性的人。这就是为什么你称我们为 AI 公司,但如果你深入了解,我们有计算机图形学专家、计算机视觉专家、数据专家、生成式 AI 专家、机器学习基础设施专家、优化专家。所以招聘一群多元化的真正有才华的人实际上非常重要,因为像空间智能这样困难的问题不是一个同质化的问题;它需要各种背景的人才来解决。然后我也寻找无所畏惧。你怎么做?你怎么判断一个人是否有无所畏惧的背景或思维过程?这体现在他们的背景中。你与他们交谈,你能感觉到某人无所畏惧。你能感觉到是什么驱动着他们。你能感觉到他们提出的问题,如果他们开始问你很多关于“我不知道如何完成这个”的事情。当然你必须问这些问题,因为你想要完成它。
Yes, we have open roles and we would love to hire the best engineers as well as product thinkers at this point for our company. So if you're an engineer or AI researcher or product talent out there passionate about joining the most talented team and solving this problem, please join us. So who do we hire? First of all, we really do hire in diversity of thinking. This is where you call us an AI company, but if you look under the hood, we've got computer graphics experts, computer vision experts, data experts, generative AI experts, machine learning infra experts, optimization experts. So it's actually really important to hire a diverse group of really talented people because a problem as hard as spatial intelligence is not a homogeneous problem; it takes talents of all kinds of background to solve it. And then I also just look for fearlessness. How do you do that? How do you identify if somebody has fearlessness in their background or in their thinking processes? It's in their background. You talk to them, you can sense someone is fearless. You can sense what drives them. You can sense the questions they ask, if they start asking you a lot of things about 'I don't know how to get this done.' And of course you have to ask those questions because you want to get it done.
但如果你觉得这源于害怕解决那个问题,那就不是无畏。但那些无畏的人,他们有创造力、有雄心,不畏惧不确定性或未知。我真的很喜欢这一点。我和很多人一样,努力与无畏的人合作,希望他们也有技术创造力。最后一个更广泛的问题,因为我认为你工作的重要部分之一是思考如何让更多人参与 AI,共同指导斯坦福以人为本人工智能中心。你对未来几年以人为本的 AI 最乐观的看法是什么?
But if you sense that it comes from the point of view of being scared of solving that, then that's not fearlessness. But those fearless people, they are creative, they're ambitious, they're not afraid of uncertainty or the unknown. And I really love that. Well, I think a lot and we try to make a business of doing business with fearless people and hopefully those that are technically creative. One last broader question for you because I think an important part of your work has been thinking how to bring more people into AI, co-directing the Stanford Center for Human-Centered Artificial Intelligence. What is your most optimistic view of what human-centered AI looks like several years out?
谢谢提问。事实上,这是我职业生涯中另一个让我感到自豪的点:创立了以人为本人工智能研究所(HAI),并持续推动这种思维方式。我想建立一个 AI 与人协作、赋能人类的世界。我仍然相信我们的人类世界需要以人为本,爱、关系、各社区的繁荣、正义等都是非常重要的价值观。我不认为任何机器,无论是 AI、飞机还是生物技术,应该夺走这些。但在这些关键价值观的基础上,用 AI 赋能我们非常重要,因为有太多未解决的问题。我研究过的一个应用领域是医疗保健,比如在斯坦福。从药物发现到治愈疾病,再到覆盖全球的诊断,到人人可及的治疗,再到整个医疗保健服务,如何改善老龄化,如何护理慢性病,如何处理心理健康——所有这些。我们并没有人类过剩的问题,我们缺乏帮助。我们缺乏科学发现、诊断、精准医疗、更安全有效的医疗保健服务和老龄化帮助。我相信 AI 是帮助人类的工具。
Thanks for asking. In fact, that is another point of my career I feel very proud of: the founding of the Human-Centered AI Institute (HAI) and the continued movement towards that way of thinking. I want to build a world where AI collaborates and superpowers people. I still believe our human world needs to be human-centered, where love, relationship, prosperity across all communities, justice, and all these are really important values. I don't think any piece of machinery, whether it's AI, airplane, or biotech, should take those away. But with those critical values in mind, having AI to superpower us is really important because there are so many unsolved problems. One application area I worked on is healthcare, for example at Stanford. If you look at healthcare from drug discovery to cure diseases to diagnosis that can reach all people in the world, to treatment accessible to all, to the whole healthcare delivery, how to make aging better, how to take care of chronic diseases, how to deal with mental health—all of this. We do not have an issue of excessive humans; we are lacking help. We are lacking scientific discovery, diagnosis, precision medicine, safer and more effective ways of healthcare delivery and aging help. I believe AI is a tool to help people.
我和很多人共同投资了一系列公司,希望它们能派上用场,从桥梁到开放证据再到洛杉矶。但正如你所说,问题范围很广。老实说,过去 15 年我对医疗技术的采用不太乐观,但这次感觉不同了,而且总体上非常有益。我之前创办了一家数字健康公司,希望人们几十年来谈论的许多事情最终能实现。AI 似乎是一个很好的交付机制。
I think a lot and we are collectively invested in a series of companies that I hope will be useful here, from a bridge to open evidence to LA. But as you said, there's a huge spectrum of problems. Honestly, I've been less optimistic about the adoption of technology in healthcare for the last 15 years, but it does feel like this time it's different and it's massively net good here. I actually started a digital health company before this, and my hope is that finally a lot of the things people have been talking about for decades will come to fruition. It seems like AI is a great delivery mechanism for that.
完全同意。非常感谢你,李飞飞。这太棒了,很受启发,也很高兴听到更多关于 World Labs 的消息。谢谢。
Totally, totally. Well, thank you so much, Fei-Fei. It was fantastic. This has been inspiring and great to hear a little bit more about World Labs as well. Thank you.
谢谢,Sarah。在 Twitter 上关注我们@no_prior_pod。订阅我们的 YouTube 频道,如果你想看到我们的面孔。在 Apple Podcasts、Spotify 或任何你收听的地方关注节目。这样你每周都能收到新剧集。在 no-priors.com 注册邮件或查找每集的文字记录。
Thank you, Sarah. Find us on Twitter at @no_prior_pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way, you get a new episode every week. And sign up for emails or find transcripts for every episode at no-priors.com.