After LLMs: spatial intelligence and world models
打开互动全文版(中英对照 + 朗读 + 问答)→为何仅有语言还不够,以及 World Labs 在造什么。
Why language alone isn’t enough, and what World Labs is building.
大家好,欢迎收听 Len Space 播客。我是 Colonel Labs 的创始人 Allesio,今天和我一起的是 Blade in Space 的编辑 Swix。我们非常激动地邀请到了 World Labs 的 Fei-Fei 和 Justin 来到演播室。欢迎你们。
Hey everyone, welcome to the Len Space podcast. This is Allesio, founder of Colonel Labs, and I'm joined by Swix, editor of Blade in Space, and we are so excited to be in the studio with Fei-Fei and Justin of World Labs. Welcome.
我们也很兴奋。
We're excited, too.
我差点说成 Marble 了。
I nearly said Marble.
谢谢邀请。我觉得大家对世界模型很感兴趣,你们也围绕空间智能等概念做了一些宣传。我想,你们两位是如何走到一起创办 World Labs 的,这可能是你们很少有机会讲述的故事。
Yeah, thanks for having us. I think there's a lot of interest in world models and you've done a little bit of publicity around spatial intelligence and all that. I guess maybe one part of the story that is a rare opportunity for you to tell is how you two came together to start building World Labs.
这很简单,因为 Justin 曾是我的学生。Justin 加入我的实验室是在 2012 年,正好是 AlexNet 问世的那个季度。
That's very easy because Justin was my former student. Justin came to my lab at Stanford when I wear my other hat as a professor of computer science. Justin joined my lab in 2012, the same quarter that AlexNet came out.
实际上,我加入你实验室的那个季度正好是 AlexNet 问世的时候。
Actually, the quarter that I joined your lab was the same quarter that AlexNet came out.
是的。Justin 是我的第一个……你当时参与那场发布风波了吗?
Yeah. So Justin is my first... Were you involved in the whole announcement drama?
完全没有。但那个季度我一直在关注 ImageNet 和 AlexNet 带来的热潮。
No, not at all. But I was sort of watching all the ImageNet excitement around AlexNet at that quarter.
他是我最优秀的学生之一,后来在密歇根大学和 Meta 担任教授,早期职业生涯非常成功。大约两年前,我们各自都在关注大模型的发展,思考语言模型之后的方向。构建世界模型和空间智能的想法对我们来说非常自然。于是我们开始交流,决定把所有精力都集中在这个问题上,共同创办了 World Labs。
So he was one of my very best students, and then he went on to have a very successful early career as a professor at University of Michigan and at Meta. Then, more than two years ago, I think both independently, we had been looking at the development of large models and thinking about what's beyond language models. The idea of building world models and spatial intelligence really was natural for us. So we started talking and decided that we should just put all the eggs in one basket and focus on solving this problem, and started World Labs together.
差不多是这样。在博士期间经历了 ImageNet 时代后,我觉得计算机视觉的下一个十年将是把 AI 从数据中心带到现实世界。所以博士毕业后,我的兴趣转向了 3D 视觉、计算机图形学和生成式建模。我以为自己正在偏离导师的方向,但几年后重聚时,发现她也在思考非常相似的问题。
Yeah, pretty much. After seeing that ImageNet era during my PhD, I had the sense that the next decade of computer vision would be about getting AI out of the data center and into the world. So a lot of my interests post-PhD shifted into 3D vision, computer graphics, and generative modeling. I thought I was drifting away from my advisor, but when we reunited a couple years later, it turned out she was thinking of very similar things.
如果回想 AlexNet,它的核心要素是 ImageNet、转向 GPU 和神经网络。对于世界模型,你认为什么相当于 AlexNet?从某种意义上说,这个想法早已存在,Yann LeCun 可能是最突出的倡导者。过去两年你看到了什么,让你觉得现在是时候了?在数据、算法或计算方法上,你们想要构建哪些根本性的东西来让世界模型成为现实?
If you think about AlexNet, the core pieces were ImageNet, the move to GPUs, and neural networks. How do you think about the AlexNet equivalent model for world models? In a way, it's an idea that has been out there. Yann LeCun is maybe the most prominent proponent. What have you seen in the last two years that made you think now is the time? And what are the fundamental things you want to build in terms of data, algorithms, or approaches to compute to make world models come to life?
我认为一个原因是,现在有更多的数据和算力可用。深度学习的历史在某种意义上就是 Scaling(规模扩张)的历史。AlexNet 需要从 CPU 转向 GPU,但从 AlexNet 到今天,每张卡的性能提升了约一千倍。现在,模型训练不仅限于单张 GPU,而是成百上千甚至数万张。因此,我们今天在单个模型上能调用的算力,比博士初期多了约一百万倍。语言模型是过去几年开始表现很好的一个有趣领域。但当我们转向视觉数据、空间数据和世界数据时,需要处理的信息量要大得多。我认为这将是吸收新增算力的一个好途径。
I think one is just that there is a lot more data and compute generally available. The whole history of deep learning is in some sense the history of scaling up compute. AlexNet required the jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card than we had in AlexNet days. Now it's common to train models not just on one GPU but on hundreds, thousands, or tens of thousands. So the amount of compute we can marshal today on a single model is about a millionfold more than at the start of my PhD. Language was one of the really interesting things that started to work well in the last couple of years. But as we move towards visual data, spatial data, and world data, we need to process a lot more. I think that's going to be a good way to soak up this new compute that's coming online.
公开挑战赛的模式仍然有效吗?还是应该集中在实验室内部进行?
Does the model of having a public challenge still work, or should it be centralized inside a lab?
我认为开放科学仍然很重要。与 AlexNet 时代相比,AI 已经发生了巨大变化;那时它只是计算机科学的一个小众领域,而现在它已成为一项文明级的技术。举个例子,我的斯坦福实验室最近发布了一个名为 BEHAVIOR 的开放数据集和基准,用于在模拟环境中评估机器人学习。这是保持开放科学模式的一个明确努力,尤其是在学术界。但重要的是要认识到,生态系统是混合的。工业界许多专注的工作以产品形式面世,而不是公开挑战赛。
I think open science is still important. AI has evolved compared to the AlexNet days; it was such a niche computer science discipline, and now it's a civilizational technology. For example, my Stanford lab recently announced an open dataset and benchmark called BEHAVIOR for benchmarking robotic learning in simulated environments. That's a clear effort to keep up the open science model, especially in academia. But it's important to recognize that the ecosystem is a mixture. A lot of the very focused work in industry sees the daylight in the form of a product rather than an open challenge per se.
是的。这只是一个资金和商业模式的问题;你需要从中看到一些投资回报。
Yeah. And that's just a matter of funding and business model; you have to see some ROI from it.
所以 Marble,从某种角度看,是一个系统。它是一个 3D 世界的生成模型。你可以输入文本、图像或多张图像,它会生成一个匹配这些输入的 3D 世界。Marble 既是朝着空间智能愿景构建的世界模型,也特意设计成今天就能让人们觉得有用的东西。我们开始看到在游戏、视觉特效和电影领域出现了一些新兴用例。Marble 作为产品今天已经能做很多有趣的事情,同时也为我们未来想要构建的宏大世界模型奠定了基础。
So Marble, basically one way of looking at it, is a system. It's a generative model of 3D worlds. You can input things like text or image or multiple images, and it will generate a 3D world that matches those inputs. While Marble is simultaneously a world model building towards this vision of spatial intelligence, it was also very intentionally designed to be something people could find useful today. We're starting to see emerging use cases in gaming, VFX, and film. There's a lot of really interesting stuff that Marble can do today as a product, and it also sets a foundation for the grand world models we want to build in the future.
我认为这只是生态系统多样性的问题。即使在所谓的 AlexNet ImageNet 时代,也有封闭模型、专有模型和开放模型。想想 iOS 与 Android,就有不同的商业模式。我不认为这仅仅是资金的问题。这只是市场的运作方式,有不同的玩法。
I think it's just a matter of the diversity of the ecosystem. Even during the so-called AlexNet ImageNet time, there were closed models, proprietary models, open models. If you think about iOS versus Android, there are different business models. I wouldn't say it's just a matter of funding per se. It's just how the market is. There are different plays.
是的,但你觉得在如今这些实验室的商业压力下,你还能重做 ImageNet 吗?对我来说,这是最大的问题:什么可以开放,什么应该保留?如果我是你,你筹集了大量资金,正在构建这一切。如果你拥有最好的数据集,你有多大动力去发布它?而且感觉实验室里的人越来越被吸引,博士项目也越来越早地被拉进这些实验室。所以我想知道,你是否认为现在存在一个问题,即资金过多,给学术界开放研究空间带来了压力,或者你觉得这并不值得担忧?
Yeah, but do you feel like you could redo ImageNet today with the commercial pressure that some of these labs have? I mean to me that's like the biggest question, right? It's like what can you open versus what should you keep inside? If I put myself in your shoes, you raise a lot of money, you're building all of this. If you had the best dataset for this, what incentives do you really have to publish it? And it feels like the people at the labs are getting more and more pulled in, the PhD programs are getting pulled earlier and earlier into these labs. So I'm curious if you think there's an issue right now with how much money has taken, how much pressure it puts on the more academia open research space, or if you feel that's not really a concern.
我确实有担忧,但更多是关于资源以及学术界资源不平衡的问题。这与 World Labs 的讨论略有不同。过去几年,我一直在倡导为健康的生态系统提供资源。作为斯坦福大学以人为本人工智能研究所的创始联合主任,我一直与政策制定者合作,为公共部门和学术界的 AI 工作提供资源。我们与第一届特朗普政府合作了一项名为《国家 AI 研究资源法案》(NAIRR)的法案,该法案规划了一个国家 AI 计算云以及数据存储库。我还认为,开源、开放数据集仍然是生态系统的重要组成部分。现在,在我的斯坦福实验室,我们正在进行一个关于机器人学习的开放数据集和开放基准测试,名为 BEHAVIOR,我的许多同事也仍在做类似的工作。我认为这是生态系统的一部分。我认为行业正在做的事情,一些初创公司正在做的事情,即快速构建模型并创造产品,也是一件好事。例如,当 Justin 还是我的博士生时,计算机视觉程序没有一个运行得很好。我们可以写出漂亮的论文。
I do have concerns, less about the pressure. It's more about the resourcing and the imbalanced resourcing of academia. This is a little bit of a different conversation from World Labs. I have been the past few years advocating for resourcing the healthy ecosystem. As the founding director, co-director of Stanford's Institute for Human-Centered AI, I've been working with policymakers about resourcing public sector and academic AI work. We worked with the first Trump administration on a bill called National AI Research Resource (NAIRR) bill, which is scoping out a national AI compute cloud as well as data repository. I also think that open-source, open datasets continue to be an important part of the ecosystem. Right now in my Stanford lab, we are doing an open dataset, open benchmark on robotic learning called BEHAVIOR, and many of my colleagues are still doing that. I think that's part of the ecosystem. I think what the industry is doing, some startups are doing, running fast with models creating products, is also a good thing. For example, when Justin was a PhD student with me, none of the computer vision programs worked that well. We could write beautiful papers.
实际上,甚至在研究生院之前,我就想做计算机视觉,我联系了谷歌的一个团队,想本科毕业后就去尝试做计算机视觉,他们告诉我:「你在说什么?你做不到。先去读个博士再回来。」
I mean, actually even before grad school, I wanted to do computer vision and I reached out to a team at Google and wanted to potentially go and try to do computer vision out of undergrad, and they told me, 'What are you talking about? You can't do that. Go do a PhD first and come back.'
是什么动力让你这样做的?
What was the motivation that got you?
哦,我在本科期间做过一些计算机视觉研究,实际上是跟 Fei-Fei 的博士导师做的。
Oh, I had done some computer vision research during my undergrad with actually Fei-Fei's PhD adviser.
有传承。
There's a lineage.
这里有传承。所以我在本科时就做过一些计算机视觉,我觉得这很酷,想继续做下去。所以即使本科毕业时,我也面临这种行业与学术的选择,我认为现在研究界的很多人也面临同样的问题。但回到你的问题,我认为学术界的作用,尤其是在 AI 领域,在过去十年中发生了很大变化。这并不是坏事。这是因为技术已经成长并涌现出来。五到十年前,你确实可以在实验室里用几块 GPU 训练出最先进的模型。但因为这项技术非常成功,规模扩大了很多,你不能再只用几块 GPU 训练最先进的模型了。这不是坏事。这是好事。这意味着技术确实奏效了。但这意味着我们作为学者应该做什么的期望有所转变。不应该是试图训练最大的模型和扩展最大的东西。而应该是尝试古怪的想法、新想法和疯狂的想法。其中大部分都不会成功。我认为在这方面还有很多事情可以做。如果说有什么担忧的话,我担心学术界有太多人过度专注于这种观念,即假装我们可以训练最大的模型,或者将其几乎视为职业培训项目,毕业后去大实验室玩所有 GPU。我认为你可以围绕新算法、新架构、新系统做很多疯狂的事情。一个人可以做很多事情。
There's a lineage here. So I had done some computer vision even as an undergrad and I thought it was really cool and I wanted to keep doing it. So then I was sort of faced with this industry-academia choice even coming out of undergrad that I think a lot of people in the research community are facing now. But to your question, I think the role of academia, especially in AI, has shifted quite a lot in the last decade. And it's not a bad thing. It's because the technology has grown and emerged. Five or ten years ago, you really could train state-of-the-art models in the lab with just a couple GPUs. But because that technology was so successful and scaled up so much, you can't train state-of-the-art models with a couple GPUs anymore. And that's not a bad thing. It's a good thing. It means the technology actually worked. But that means the expectations around what we should be doing as academics shifts a little bit. And it shouldn't be about trying to train the biggest model and scaling up the biggest thing. It should be about trying wacky ideas and new ideas and crazy ideas. Most of which won't work. And I think there's a lot to be done there. If anything, I'm worried that too many people in academia are hyperfocused on this notion of trying to pretend like we can train the biggest models or treating it as almost a vocational training program to then graduate and go to a big lab and be able to play with all the GPUs. I think there's just so much crazy stuff you can do around new algorithms, new architectures, new systems. There's a lot you can do as one person.
此外,学术界在理解这些大型模型的理论基础方面也发挥着作用。我们对此仍然知之甚少。或者扩展到跨学科领域。Justin 称之为古怪的想法。有很多基础科学想法。有很多蓝天问题。所以我同意。我不认为问题是开放与封闭、产品化与开源。我认为现在的问题是学术界本身资源严重不足。因此,研究人员和学生没有足够的资源来尝试这些想法。
And also, academia has a role to play in understanding the theoretical underpinnings of these large models. We still know so little about this. Or extend to the interdisciplinary. Justin calls them wacky ideas. There's a lot of basic science ideas. There's a lot of blue sky problems. So I agree. I don't think the problem is open versus closed, productization versus open sourcing. I think the problem right now is that academia by itself is severely underresourced. So the researchers and the students do not have enough resources to try these ideas.
是的。为了让人们能深入探讨,当你谈到古怪的想法时,你想到的一个古怪想法是什么?
Yeah. Just for people to nerd snipe, what's a wacky idea that comes to mind when you talk about wacky ideas?
哦,我有一个想法,我一直向我在密歇根大学的学生们推销,那就是我真的很喜欢硬件,也喜欢新型硬件不断涌现。从某种意义上说,我们今天使用的神经网络和 Transformer 的出现实际上是基于矩阵乘法,因为矩阵乘法非常适合 GPU。但如果我们考虑 GPU 将如何扩展,硬件未来可能如何扩展,我认为我们目前的系统,即类似 GPU 的硬件设计,不会无限扩展。我们现在已经开始看到,计算单元不再是单个设备,而是整个设备集群。所以如果你想象一个节点,是的,它是一个完整的节点或一个完整的集群。但我们谈论神经网络的方式仍然好像它们是一个单一的整体,可以在一个 GPU 上用 PyTorch 编码,但实际上它们可以分布在数千个设备上。
Oh, I had this idea that I kept pitching to my students at Michigan, which is that I really like hardware and I really like new kinds of hardware coming online. In some sense, the emergence of the neural networks that we use today and transformers are really based around matrix multiplication because matrix multiplication fits really well with GPUs. But if we think about how GPUs are going to scale, how hardware is likely to scale in the future, I don't think the current system that we have, the GPU-like hardware design, is going to scale infinitely. And we start to see that even now, the unit of compute is not the single device anymore. It's this whole cluster of devices. So if you imagine a node, yeah, it's a whole node or a whole cluster. But the way we talk about neural networks is still as if they are a monolithic thing that could be coded like in one GPU in PyTorch, but then in practice they could distribute over thousands of devices.
那么,就像 Transformer 基于矩阵乘法,而矩阵乘法是一种在 GPU 上效果很好的原语,随着硬件规模扩张,是否存在其他更适合大规模分布式系统的原语,我们可以基于它们构建神经网络?我认为有可能出现截然不同的架构,以适应未来 10 到 20 年将出现的新一代硬件,我们现在就可以开始想象。但很难下这种赌注,因为还有硬件彩票的概念——比如说,英伟达已经赢了,我们就应该无限扩展它,然后写软件来填补任何漏洞,对吧?
So are there, just as transformers are based around matmul and matmul is sort of the primitive that works really well on GPUs as you imagine hardware scaling out, are there other primitives that make more sense for large scale distributed systems that we could build our neural networks on? And I think it's possible that there could be drastically different architectures that fit with the next generation of hardware that's going to come 10 or 20 years down the line, and we could start imagining that today. It's really hard to make those kinds of bets because there's also the concept of the hardware lottery where, let's just say, Nvidia has won and we should just scale that out infinitely and write software to patch up any gaps we have in the mix, right?
我的意思是,既是也不是。如果你看数据,即使从 Hopper 到 Blackwell,每瓦性能也差不多。他们主要是增加了晶体管数量、芯片尺寸和功耗。但即使从 Hopper 到 Blackwell,我们已经在每瓦性能上看到了 Scaling 的限制。所以我认为还有空间去做一些新的事情,我不确定具体是什么,而且我不认为你能在三个月内作为初创公司完成它。但我觉得这种想法,如果你坐下来思考几年,也许能取得一些突破。我认为这种长期的研究非常适合学术界。
I mean, yes and no. If you look at the numbers, even going from Hopper to Blackwell, the performance per watt is about the same. They mostly make the number of transistors go up, they make the chip size go up, and they make the power usage go up. But even from Hopper to Blackwell, we're kind of already seeing a scaling limit in terms of what performance per watt we can get. So I think there is room to do something new, and I don't know exactly what it is, and I don't think you can get it done in a three-month cycle as a startup, but I think that's the kind of idea that if you sit down and sit with for a couple years, maybe you could come up with some breakthroughs. And I think that's the kind of long-range stuff that is a perfect match for academia.
回到一点背景和历史,你们做过场景叙事的工作,或者更近期的与 Andrej 合作的图像描述。我想听你们讲讲那个故事,关于你开始博士研究,以及 Fei-Fei 当时的反应。
Coming back to a bit of background and history, we have this sort of research note on the scene storytelling work that you did, or newer image captioning that you did with Andrej. I just wanted to hear you guys tell that story about you embarking on that for your PhD and Fei-Fei having that reaction that you had.
是的。我认为这项工作始于我和 Andrej,后来 Justin 也加入了。Andrej 开始读博。我们当时在思考 ImageNet 物体识别之后的下一个方向。那时,卷积神经网络已经在 ImageNet 任务中展现了一些能力,所以 ConvNet 是表示图像的好方法。同时,在语言领域,一种早期的序列模型 LSTM 也在被实验。Andrej 和我讨论过这个。这是我长期以来的梦想。我曾以为需要 100 年才能解决,那就是讲述图像的故事。当我从研究生院毕业时,我真的认为我整个职业生涯的剩余部分都将致力于解决这一个问题:给定一张图片或一个场景,用自然语言讲述故事。但事情发展得太快了。当 Andrej 开始时,我们想也许结合卷积神经网络的表示和 LSTM 的语言序列模型,我们可以通过训练来学习将描述与图像匹配。于是我们开始了这项工作。我记得是 2014 或 2015 年。那是 CVPR 2015 上的一篇描述论文。那是我们的第一篇论文,Andrej 让它成功了:给定一张图像,用 ConvNet 表示图像,语言模型是 LSTM,我们将它们结合起来,就能生成一个句子。那是第一次。我在书中写道,我们以为我们是第一个做这个的,但结果谷歌当时也在同时做。纽约时报的记者 John Markoff 正在报道谷歌的故事,但他偶然听说了我们,然后意识到我们确实是独立同时达到的。所以他写了关于谷歌研究以及 Andrej 和我的研究的报道。在那之后,Justin 当时已经在实验室了。
Yeah. So I think that line of work started between me and Andrej, and then Justin joined. Andrej started his PhD. He and I were looking at what is beyond ImageNet object recognition. At that time, the convolutional neural network had proven some power in ImageNet tasks, so ConvNet is a great way to represent images. In the meantime, in the language space, an early sequential model called LSTM was also being experimented. Andrej and I were just talking about this. It has been a long-term dream of mine. I thought it would take 100 years to solve, which is telling the story of images. When I graduated from grad school, I really thought the rest of my entire career would be towards solving that single problem: given a picture or a scene, tell the story in natural language. But things evolve so fast. When Andrej started, we thought maybe combining the representation of convolutional neural networks as well as the language sequential model of LSTM, we might be able to learn through training to match captions with images. So that's when we started that line of work. I think it was 2014 or 2015. It was a CVPR 2015 paper on captioning. It was our first paper that Andrej got to work: given an image, the image is represented with ConvNet, the language model is the LSTM model, and we combine it, and it's able to generate one sentence. That was one of the first times. I think I wrote in my book that we thought we were the first people doing it, but it turned out that Google at that time was also simultaneously doing it. A reporter, John Markoff from the New York Times, was breaking the Google story, but he by accident heard about us, and then he realized that we really independently got there together at the same time. So he wrote the story of both the Google research as well as Andrej and my research. But after that, I think Justin was already in the lab at that time.
是的。我记得那次组会,Andrej 在展示一些结果,并解释这个叫 LSTM 和 RNN 的新东西,我以前从未听说过,我想,哇,这真是太棒了,我想做这个。然后他在 CVPR 2015 上发表了第一篇图像描述结果的论文。之后我们开始合作。我们首先在 2015 年做了一篇纯粹关于语言建模的论文。我本应该坚持语言建模的;事后看来,那相当赚钱。但我们一起做了这篇语言建模论文,我和 Andrej,在 2015 年,真的很酷。我们训练了这些小型的 RNN 语言模型,它们可以一次吐出几个句子,然后我们戳它们,试图理解神经网络内部的神经元在做什么。
Yeah. I remember the group meeting where Andrej was presenting some of those results and explaining this new thing called LSTMs and RNNs that I had never heard of before, and I thought, wow, this is really amazing stuff, I want to work on that. So then he had the paper at CVPR 2015 on the first image captioning results. Then after that we started working together. We first did a paper actually just on language modeling back in 2015. I should have stuck with language modeling; that turned out to be pretty lucrative in retrospect. But we did this language modeling paper together, me and Andrej, in 2015, where it was really cool. We trained these little RNN language models that could spit out a couple sentences at a time and poke at them and try to understand what the neurons inside the neural network were doing.
你们当时在做关于不同记忆之类的分析……
You guys were doing analysis on the different like memory and...
是的。那真的很酷。即使在那个时候,我们也有一些结果,你可以观察 LSTM 内部,然后说,哦,这个东西在读代码。所以我们训练的一个数据集是 Linux 源代码,因为整个东西是开源的,你可以直接下载。我们在那个数据集上训练了一个 RNN,然后当网络试图预测 token 时,我们试图将它的预测与 RNN 内部结构关联起来。我们找到了一些相关性:哦,LSTM 这一层的这个单元在遇到左括号时激活,然后在遇到右括号时关闭。我们尝试了一些这样的经验性工作来弄清楚。所以那很酷,那就像是把 CNN 从语言建模部分中剥离出来,只孤立地看语言模型。
Yeah. It was really cool. And even at that time we had these results where you could look inside the LSTM and say, oh, this thing is reading code. So one of the data sets that we trained on for this one was the Linux source code, because the whole thing is open source and you could just download it. So we train an RNN on this data set, and then as the network is trying to predict the tokens, we try to correlate the kinds of predictions that it's making with the kind of internal structures in the RNN. And there we were able to find some correlations: oh, this unit in this layer of the LSTM fires when there's an open parenthesis, and then turns off when there's a closed parenthesis. We tried to do some empirical stuff like that to figure it out. So that was pretty cool, and that was sort of like cutting out the CNN from this language modeling part and just looking at the language models in isolation.
但后来我们想扩展图像描述的工作。我记得那时我们甚至有了空间感,因为我们觉得描述没有捕捉到图像的不同部分。所以我和 Justin 和 Andrej 讨论,能不能做我们后来称之为密集描述的东西,即更详细地描述场景,尤其是场景的不同部分。
But then we wanted to extend the image captioning work. I remember at that time we even had a sense of space because we felt like captioning does not capture different parts of the image. So I was talking to Justin and Andrej about can you go what we ended up calling dense captioning, which is describe the scene in greater detail, especially different parts of the scene.
是的,然后我们构建了这个系统。第二年,也就是 CVPR 2016,我和 Andrej 以及 Fei-Fei 一起发表了一篇论文,我们构建了这个做密集描述的系统。
Yeah, and so then we built this system. It was me, Andrej, and Fei-Fei on a paper the following year, CVPR 2016, where we built this system that did dense captioning.
所以输入一张图片,它会框出所有有趣的东西,然后为每个框写一小段描述。比如,哦,桌上有个绿色水瓶,有个人穿着黑衬衫。这是一个非常复杂的神经网络,因为它建立在当时物体检测领域的许多进步之上,物体检测长期以来一直是计算机视觉的主要课题。实际上它是一个联合神经网络,同时学习观察单个图像,因为网络内部有三种不同的表示。一种是整个图像的表示,用来把握整体情况。然后它会提出想要关注的各个区域,独立地表示每个区域,一旦你看了这个区域,就需要为每个区域生成文本。那是一个相当复杂的神经网络架构。这一切都发生在 PyTorch 出现之前。
So you input a single image and then it would draw boxes around all the interesting stuff in the image and then write a short snippet about each of them. It's like, oh, it's a green water bottle on the table. It's a person wearing a black shirt. And this was a really complicated neural network because it was built on a lot of advancements that had been made in object detection around that time, which was a major topic in computer vision for a long time. And then it was actually like one joint neural network that was both learning to look at individual images because they actually had three different representations inside this network. One was the representation of the whole image to kind of get the gestalt of what's going on. Then it would propose individual regions that it wants to focus on and then look at represent each region independently and then once you look at the region then you need to spit out text for each region. So that was a pretty complicated neural network architecture. This was all pre-PyTorch.
对。它是一次性完成的吗?
Right. And does it do it in one pass?
是的。所以它是一次前向传播就完成了所有工作。
Yeah. So it was a single forward pass that did all of that.
不仅是一次性完成,你还优化了推理。我记得你是通过摄像头实时运行的。
Not only was it doing it in one pass, you also optimized inference. You're doing it on a webcam. I remember.
是的。我构建了一个疯狂的实时演示,网络在斯坦福的服务器上运行,前端通过网页从摄像头获取视频流,把图像发送回服务器。服务器运行模型并流式返回预测结果。我就拿着笔记本电脑在实验室里走来走去,向人们实时展示这个网络。
Yeah. So, I had built this crazy real-time demo where I had the network running on a server at Stanford and then a web front end that would stream from a webcam and then send the image back to the server. The server would run the model and stream the predictions back. So, I was just walking around the lab with this laptop that would show people this network in real time.
识别和标注也一并完成。是的,这非常令人印象深刻,因为大多数研究生如果能把论文发表就满足了,对吧?他们把研究打包成论文,但 Justin 更进一步。他说,我想做一个实时网络演示。
Identification and labeling as well. Yeah, it was pretty impressive because most of our graduate students would be satisfied if they can publish the paper, right? They package the research, put it in a paper, but Justin went a step further. He's like, I want to do this real-time web demo.
实际上,我不知道有没有跟你讲过这个故事,那年我们在圣地亚哥有个会议,ICCV 2015。我在那个会议上有一篇不同的论文,但我带着我的笔记本电脑。我在会场里走来走去,向大家展示这个实时字幕演示,模型运行在加州的一个服务器上。所以它实际上能从加州一路流式传输到圣地亚哥。延迟很糟糕,大概只有 1 帧每秒,但能工作就已经很惊人了。
Well, actually, I don't know if I told you this story, but then we had a conference that year in Santiago at ICCV. It was ICCV 15. And then I had a paper at that conference for something different, but I had my laptop. I was walking around the conference with my laptop showing everybody this real-time captioning demo and the model was running on a server in California. So, it was actually able to stream all the way from California down to Santiago. Well, latency was terrible. It was like 1 FPS, but the fact that it worked at all was pretty amazing.
我本想简单调侃一下,也许视觉和语言建模并没有那么不同。你知道,DeepSeek 最近尝试了一个疯狂的想法:从像素建模文本,然后直接训练,这可能是未来。我不知道你们对语言是否真的必要有什么看法。
I was going to briefly quip that, you know, maybe vision and language modeling are not that different. You know, DeepSeek recently tried the crazy thing of let's model text from pixels and just train on that and it might be the future. I don't know if you guys have any takes on whether language is actually necessary at all.
我刚刚写了一整篇关于空间智能的宣言。
I just wrote a whole manifesto on spatial intelligence.
这正是我引入这个话题的方式。是的。
This is my segue into this. Yes.
我认为它们不同。我确实认为这些生成模型的架构会共享许多可共享的组件,但我认为深层的 3D、4D 空间世界具有一种结构层次,与纯粹的一维生成信号有根本区别。
I think they are different. I do think the architecture of these generative models will share a lot of shareable components, but I think the deeply 3D, 4D spatial world has a level of structure that is fundamentally different from a purely generative signal that is one-dimensional.
是的,我认为像素最大化主义是有道理的,对吧?有一种观念认为语言是另一种东西,但我们用眼睛看语言,而我们的眼睛基本上就是像素,对吧?我们眼睛后面有某种生物像素在处理这些东西。我们看到文本,认为它是离散的东西,但这实际上只存在于我们的脑海中。文本和语言在我们世界中的物理表现是印在物体上的物理对象,我们用眼睛看到它们。你也可以认为它是声音,但即使是声音也可以转换成信号,对吧?然后如果你把它转换成我们在 LLM 中使用的纯 token 化表示,你实际上会丢失一些东西,对吧?比如丢失了字体、换行符、页面上的二维布局。在很多情况下,这可能无关紧要。但对于某些事情来说,这很重要。我认为像素是一种更无损的世界表示,在某种程度上是一种更通用的表示,更符合我们人类在探索世界时所看到的。所以有一个效率方面的论点,比如将文本渲染成图像再输入视觉模型可能不是非常高效。
Yeah, I think there's something to be said for pixel maximalism, right? Like there's this notion that language is this different thing, but we see language with our eyes and our eyes are just like, you know, basically pixels, right? Like we've got sort of biological pixels in the back of our eyes that are processing these things. And we see text and we think of it as this discrete thing, but that really only exists in our minds. Like the physical manifestation of text and language in our world are physical objects that are printed on things in the world and we see it with our eyes. Well, you can also think it's sound, but even sound you can translate into a signal, right? And then you actually lose something if you translate to this purely tokenized representations that we use in LLMs, right? Like you lose the font, you lose the line breaks, you lose the 2D arrangement on the page. And for a lot of cases, maybe that doesn't matter. But for some things it does. And I think pixels are this sort of more lossless representation of what's going on in the world and in some ways a more general representation that more matches what we humans see as we navigate the world. So there's an efficiency argument to be made like maybe it's not super efficient to render your text to an image and then feed that to a vision model.
这正是 DeepSeek 所做的,对吧?它有点效果。我认为这关系到整个世界模型。今年我看到的最喜欢的论文之一是关于世界模型的归纳偏置。那是一篇哈佛的论文,他们向 LLM 输入了大量轨道模式,然后让 LLM 预测行星绕太阳的轨道,生成的模型看起来不错,但如果你让它画出力向量,结果就会很混乱,实际上并不遵循物理规律。那么你如何看待数据中嵌入的内容?我们可以讨论一下为 3D 世界模型进行 token 化,比如信息的维度。有视觉信息,但你需要从数据中提取多少潜在的隐藏力,以及其中有哪些挑战?
That's exactly what DeepSeek did, right? It kind of worked. I think this ties into the whole world model. One of my favorite papers that I saw this year was about inductive bias for world model. So it was a Harvard paper where they fed a lot of orbital patterns into an LLM and then they asked the LLM to predict the orbit of a planet around the sun and the model generated looked good but then if you asked it to draw the force vectors it would be all wacky, it wouldn't actually follow it. So how do you think about what's embedded into the data that you get? And we can talk about maybe tokenizing for 3D world models like what are the dimensions of information. There's the visual but how much of the underlying hidden forces so to speak you need to extract out of this data and what are some of the challenges there?
是的,我认为有不同方法可以处理这个问题。一种是你可以尝试显式地处理,比如说,「哦,我想测量所有力,并将它们作为训练数据输入模型。」然后你可以运行传统的物理模拟,知道场景中的所有力,然后用这些作为训练数据来训练一个模型,希望它能预测这些力。或者你可以希望某些东西更隐性地涌现出来,你端到端地训练一个更通用的问题,然后希望模型内部的某个地方必须学会建模类似物理的东西才能做出正确的预测。这大致就是我们拥有的两大范式。
Yeah, I think there's different ways you could approach that problem. One is you could try to be explicit about it and say, 'Oh, I want to measure all the forces and feed those as training data to your model.' Then you could run a traditional physics simulation and know all the forces in the scene and then use those as training data to train a model that's now going to hopefully predict those. Or you could hope that something emerges more latently, that you kind of train on something end to end on a more general problem and then hope that somewhere in the internals of the model must learn to model something like physics in order to make the proper predictions. And those are kind of the two big paradigms that we have more generally.
但没有迹象表明那些潜在建模能让你得到空间和动力学的因果律。
But there's no indication that those latent modeling will get you to a causal law of space and dynamics.
对吧?这正是当今深度学习与人类智能开始分叉的地方,因为深度学习本质上还是在拟合模式。
Right? That's where today's deep learning and human intelligence actually start to bifurcate because fundamentally deep learning is still fitting patterns.
这就变得哲学了:我们人类也在拟合模式,但或许我们是在更长的时间跨度、用不同的奖励函数去拟合更广泛的模式。但你提到的那篇论文讲的就是这个问题:它学会了拟合特定的轨道模式,却没有按你期望的方式泛化,它没有因果性的引力模型。
There you get philosophical and say that we're trying to fit patterns too, but maybe we're trying to fit a broader array of patterns over a longer time horizon with a different reward function. But basically the paper you mentioned is about that problem: it learns to fit specific patterns of orbits but doesn't actually generalize in the way you'd like. It doesn't have a causal model of gravity.
对吧?因为即使在 Marble 里,我试过,它生成了漂亮的风景,里面还有拱门。
Right? Because even in Marble, I was trying it and it generates beautiful scenery with arches in them.
但模型真的理解拱门是如何依靠中心石以及实际的物理结构吗?另一个问题是:只要它始终渲染出符合我们想象的物理模型的东西,它是否理解重要吗?如果你用你理解「理解」的方式,我相当确定模型并不理解。模型是从数据中学习,从模式中学习。这重要吗?尤其是对于它的用例。这是个好问题。目前我认为不重要,因为假设它完美,它就能渲染出你需要的东西。
But does the model actually understand how the arch is drawing on the center stone and the actual physical structure? And the other question is: does it matter that it understands, as long as it always renders something that fits the physical model we imagine? If you use the word 'understand' the way you understand, I'm pretty sure the model doesn't understand it. The model is learning from the data, learning from the pattern. Does it matter? Especially for the use cases. It's a good question. For now, I don't think it matters because it renders what you need, assuming it's perfect.
是的。我的意思是,这取决于用例。如果用例是为虚拟电影制作生成背景,你只需要看起来合理的东西,所以可能不重要。但如果你是建筑师,用它来设计一栋要在现实世界中建造的建筑,那么正确建模受力就很重要,因为你不想让建筑倒塌。但即便如此,即使你的模型有语义,我仍然认为模型的理解和人的理解不是同一个词。这又变得哲学了。
Yeah. I mean, it depends on the use case. If the use case is generating a backdrop for virtual film production, all you need is something that looks plausible, so probably it doesn't matter. But if you're an architect using it to design a building to build in the real world, then it does matter that you model the forces correctly because you don't want the thing to break. But even there, if your model has the semantics, I still don't think the understanding on the model's part and the understanding on the human part are the same word. This gets philosophical again.
是的,理解有个把戏。这些模型是一种与人类智能截然不同的智能。人类智能有趣之处在于,我认为我理解事物是因为我能在一定程度上内省自己的思维过程。然后我相信自己的思维过程与他人相似,所以当我观察别人的行为时,我推断他们的内心状态可能与我相似。因此,我知道我理解事物。所以我假设你理解某些东西。但这些模型是一种外星智能。它们能做有趣的事情,表现出有趣的行为,但无论它们有什么样的内部认知或自我反思——如果存在的话——都与我们完全不同。
Yeah, there's this trick with understanding. These models are a very different kind of intelligence than human intelligence. Human intelligence is interesting because I think I understand things because I can introspect my own thought process to some extent. Then I believe my thought process works similarly to others, so when I observe someone else's behavior, I infer their internal mental state is probably similar to mine. Therefore, I know I understand things. So I assume you understand something. But these models are an alien form of intelligence. They can do interesting things and exhibit interesting behavior, but whatever internal cognition or self-reflection they have, if it exists at all, is totally different from ours.
它没有自我意识。
It doesn't have self-awareness.
对。但这意味着,当我们观察到这些系统看似有趣或智能的行为时,我们不一定能推断出它们的其他特性,因为它们的世界模型和思维方式与我们如此不同。
Right. But what that means is that when we observe seemingly interesting or intelligent behavior from these systems, we can't necessarily infer other things about them because their model of the world and the way they think is so different from us.
那么,你需要两个不同的模型来分别处理视觉和建筑生成吗?你认为最终你采用的模型构建方法没有根本性问题吗?更多的是关于扩展模型及其能力,还是说视觉性本身有什么东西阻碍了你学习背后的物理,从而让你可以信任它生成一个在现实世界中可行的猫设计?
So would you need two different models to do the visual one and the architectural generation? Do you think eventually there's nothing fundamental about the approach you've taken on model building? Is it more about scaling the model and its capabilities, or is there something about being very visual that prohibits you from learning the physics behind it so you could trust it to generate a cat design that works in the real world?
我认为这是扩展数据和改进模型的问题。我不认为这两者之间有根本性的区别。
I think this is a matter of scaling data and bettering the model. I don't think there's anything fundamental that separates these two.
是的,我希望它是一个模型,但我认为深度学习在某种意义上的一大问题是,如何获得超越训练数据的涌现能力?你会得到某种理解受力、但并未被训练去预测受力的东西,而它会隐式地在内部学习它们吗?我认为我们在其他大型模型中看到的是,这种涌现行为确实在规模扩大时发生了。这能否迁移到其他模态、用例和任务?我希望如此,但这将是一个需要时间展开并观察的过程。
Yeah, I would like it to be one model, but I think the big problem in deep learning in some sense is how do you get emergent capabilities beyond your training data? Are you going to get something that understands the forces while it wasn't trained to predict the forces, but it's going to learn them implicitly internally? I think a lot of what we've seen in other large models is that this emergent behavior does happen at scale. Will that transfer to other modalities and use cases and tasks? I hope so, but that'll be a process we need to play out over time and see.
是否有诱惑去依赖已经存在的物理引擎,比如游戏行业已经为你省去了很多工作,还是说由于某种根本性的不匹配,我们必须重新发明?
Is there a temptation to rely on physics engines that already exist, like the gaming industry has saved you a lot of this work, or do we have to reinvent things for some fundamental mismatch?
我认为这就像攀登技术阶梯。从某种意义上说,你之所以要构建这些东西,是因为传统物理引擎在某些情况下不起作用。如果物理引擎是完美的,我们就不需要构建模型了,因为问题已经解决了。所以我们想这样做是因为经典物理引擎不能以我们想要的通用性解决问题。但这并不意味着我们需要抛弃它们、从头开始。我们可以用传统物理引擎生成数据,然后用这些数据训练我们的模型。这样你就把物理引擎蒸馏到了神经网络的权重中。我认为很多其他实验室都在做类似的事情。人们猜测 Sora 有一点,Genie3 也有一点。Genie 明确就像一款电子游戏,有控制可以四处走动。我总觉得我们为了好玩而发明的东西最终进入严肃工作,这很有趣。
I think that's like climbing the ladder of technology. In some sense, the reason you want to build these things at all is because traditional physics engines don't work in some situations. If a physics engine were perfect, we would have no need to build models because the problem would already be solved. So the reason we want to do this is because classical physics engines don't solve problems in the generality we want. But that doesn't mean we need to throw them away and start from scratch. We can use traditional physics engines to generate data that we then train our models on. Then you're distilling the physics engine into the weights of the neural network. I think that's a lot of what other labs are doing. People speculate that Sora had a bit of that, Genie3 had a bit of that. Genie is explicitly like a video game with controls to walk around. I always think it's funny how things we invent for fun eventually make it into serious work.
是的。整个 AI 革命始于图形芯片。部分原因是把 GPU 从生成大量三角形误用于生成大量其他一切。
Yeah. The whole AI revolution started with graphics chips. Partially misusing the GPU for generating a lot of triangles into generating a lot of everything else.
是的。
Yeah.
我们稍微提到了 Marble。我认为你们选择 Marble 作为某种走出隐身状态的时刻,如果可以这么说的话。
We touched on Marble a little bit. I think you guys chose Marble as your sort of coming out of stealth moment, if you can call it that.
是的。
Yeah.
也许你可以简要解释一下人们应该从中得到什么,因为这里每个人都可以试用 Marble,但我觉得他们可能无法将其与你的愿景和其他实验室可能见过的生成世界之间的差异联系起来。
Maybe we can get a concise explanation from you on what people should take away because everyone here can try Marble but I don't think they might be able to link it to the differences between what your vision is versus other generative worlds they may have seen from other labs.
所以 Marble 是我们模型的一个窗口。我们是一家空间智能模型公司。我们相信空间智能是下一个前沿。为了制造空间智能模型,模型必须非常强大,能够以多模态的方式理解、推理和生成世界,并允许我们最终希望达到的与人类互动世界一样复杂的交互水平。这就是空间智能以及我们所见的世界模型的宏大愿景。Marble 是它的第一个窗口,是这段旅程的第一步。它是世界上第一个以这种保真度生成 3D 世界并交付给公众的同类模型。这是一个起点。我们实际上写了一篇技术博客。Justin 花了大量时间撰写那篇博客。我不知道你是否有时间浏览。Justin 真的把它分解成了输入——我们可以有 Marble 的多模态输入——可编辑性,允许用户与模型交互,以及我们可以拥有的输出类型。
So Marble is a glimpse into our model. We are a spatial intelligence model company. We believe spatial intelligence is the next frontier. In order to make spatially intelligent models, the model has to be very powerful in terms of its ability to understand, reason, and generate in a very multimodal fashion of worlds, as well as allow the level of interactivity that we eventually hope to be as complex as how humans can interact with the world. So that's the grand vision of spatial intelligence as well as the kind of world models we see. Marble is the first glimpse into that. It's the first part of that journey. It's the first in-class model in the world that generates 3D worlds at this level of fidelity that is in the hands of the public. It's the starting point. We actually wrote a tech blog. Justin spent a lot of time writing that tech blog. I don't know if you had time to browse it. Justin really broke it down into what are the inputs—we can have multimodal inputs of Marble—what are the kind of editability which allows users to be interactive with the model, and what are the kind of outputs we can have.
是的。
Yeah.
所以 Marble,基本上可以看作是一个 3D 世界的生成模型。你可以输入文本、图像或多张图像,它会为你生成一个匹配这些输入的 3D 世界。它也是交互式的,你可以交互式地编辑场景。我可以生成这个场景,然后说我不喜欢这个水瓶,把它变成蓝色,移除桌子,改变这些麦克风的位置,然后你可以基于这些交互式编辑生成新的世界,并以多种格式导出。通过 Marble,我们实际上试图同时做两件事,我认为我们很好地平衡了这一点。一是构建一个朝着空间智能宏大愿景发展的模型。模型需要能够理解各种不同的输入,能够在多种情况下对世界进行建模,能够对世界随时间变化的反事实进行建模。所以我们想开始构建具有这些能力的模型,而今天的 Marble 已经具备了所有这些的雏形。但与此同时,我们是一家公司,一个企业。我们真的不想让它只是一个科学项目,而是构建一个在现实世界中对人们有用的产品。所以,虽然 Marble 同时是一个朝着空间智能愿景发展的世界模型,但它也被有意设计成今天人们可以实际使用的东西。我们开始看到在游戏、视觉特效、电影中出现的用例,我认为 Marble 作为产品今天可以做很多有趣的事情,同时也为我们未来想要构建的宏大世界模型奠定了基础。
So Marble, basically one way of looking at it, it's a generative model of 3D worlds. You can input things like text or image or multiple images and it will generate for you a 3D world that matches those inputs. It's also interactive in the sense that you can interactively edit scenes. I could generate this scene and then say I don't like the water bottle, make it blue instead, take out the table, change these microphones around, and then you can generate new worlds based on these interactive edits and export in a variety of formats. With Marble, we were actually trying to do two things simultaneously, and I think we managed to pull off the balance pretty well. One is actually build a model that goes towards the grand vision of spatial intelligence. Models need to be able to understand lots of different kinds of inputs, need to be able to model worlds in a lot of situations, need to be able to model counterfactuals of how they could change over time. So we wanted to start to build models that have these capabilities, and Marble today already has hints of all of these. But at the same time, we're a company, we're a business. We were really trying not to have this be a science project but also build a product that would be useful to people in the real world today. So while Marble is simultaneously a world model that is building towards this vision of spatial intelligence, it was also very intentionally designed to be a thing that people could find useful today. We're starting to see emerging use cases in gaming, in VFX, in film where I think there's a lot of really interesting stuff that Marble can do today as a product, and also set a foundation for the grand world models that we want to build going into the future.
是的。我注意到一个非常有趣的工具,因为你可以在里面录制你的场景。
Yeah. I noticed one tool that was very interesting because you can record your scene inside.
是的。
Yes.
是的。这非常重要。录制的能力意味着对相机位置的精确控制。为了精确控制相机位置,你必须对 3D 空间有感知。否则,你不知道如何定位相机以及如何移动相机。所以这是这类模型的自然结果,这也是为什么这只是其中一个例子。
Yes. It's very important. The ability to record means very precise control of camera placement. In order to have precise camera placement, it means you have to have a sense of 3D space. Otherwise, you don't know how to orient your camera and how to move your camera. So that is a natural consequence of this kind of model, and this is why this is just one of the examples.
是的。我发现当我使用视频生成模型时,我不得不学习导演的语言,因为我必须像平移一样移动它们。你不能说向北平移 63 度,对吧?你根本没有那种控制。而在 Marble 中,你可以精确控制相机的位置。是的,我认为这是人们需要首先理解的事情之一:它不是逐帧生成,而很多其他模型是那样做的。
Yeah. I find when I play with video generative models, I'm having to learn the language of being a director because I have to move them like pan. You cannot say pan 63 degrees to the north, right? You just don't have that control. Whereas in Marble, you have precise control in terms of placing your camera. Yeah, I think that's one of the first things people need to understand: it's not generating frame by frame, which is what a lot of the other models are doing.
是的。
Yeah.
原子单元是什么?人们知道大语言模型生成一个 token。原子单元是什么?有网格、高斯溅射、体素。3D 世界中有很多组成部分。人们应该对你的生成有什么样的心智模型?
What are the atomic units? People understand that LLM generates one token. What are the atomic units? There are meshes, splats, voxels. There are a lot of pieces in a 3D world. What should be the mental model that people have of your generations?
是的,我认为有今天存在的和未来可能存在的。今天存在的是模型原生输出高斯溅射。高斯溅射是微小的半透明粒子,在 3D 空间中具有位置和方向。场景由大量这样的高斯溅射构建而成。高斯溅射非常酷,因为你可以非常高效地实时渲染它们。所以你可以用 iPhone 渲染,渲染一切。这就是我们获得精确相机控制的方式,因为溅射可以在几乎任何客户端设备上实时渲染。所以对于今天我们生成的许多场景,原子单元就是单个溅射。但我不认为这是根本性的。我可以想象未来其他有趣的方法。甚至我们在 World Labs 也研究过其他方法,比如我们最近的 RTFM 模型,它确实逐帧生成。那里的原子单元是随着用户与系统交互而逐帧生成。或者你可以想象未来的其他架构,其中原子单元是一个 token,代表 3D 世界的某个块。我认为随着时间的推移,我们可以尝试许多不同的架构。
Yeah, I think there's what exists today and what could exist in the future. What exists today is the model natively outputs splats. Gaussian splats are tiny particles that are semi-transparent, have a position and orientation in 3D space. The scene is built up from a large number of these Gaussian splats. Gaussian splats are really cool because you can render them in real time very efficiently. So you can render on your iPhone, render everything. That's how we get that precise camera control because the splats can be rendered in real time on pretty much any client-side device. So for a lot of the scenes we're generating today, the atomic unit is that individual splat. But I don't think that's fundamental. I could imagine other approaches in the future that would be interesting. There are other approaches that even we've worked on at World Labs, like our recent RTFM model that does generate frames one at a time. There the atomic unit is generating frames one at a time as the user interacts with the system. Or you could imagine other architectures in the future where the atomic unit is a token that now represents some chunk of the 3D world. I think there are a lot of different architectures that we can experiment with over time.
你在之前的发言中也经常提到物理和力,这是随时间变化的东西。我在 Marble 中没看到这个,我猜它还没有。也许如果有个 Marble 2,你会有运动,或者有没有对高斯泼溅的修改能实现,还是说会是完全不同的东西?
You also in previous statements focus a lot on the physics and the forces, which is something over time. I don't see that in Marble. I presume it's not there yet. Maybe if there was like a Marble 2, you would have movement, or is there a modification to Gaussian Splats that makes sense, or would it be something completely different?
我认为有一些修改是合理的,实际上有很多有趣的方法来整合这些东西,这也是这个领域工作的另一个好处。这方面已经有很多研究工作。比如你提到疯狂的想法,实际上有很多非常有趣的学术工作,研究不同的方式来注入物理。
I think there are a couple modifications that make sense, and there's actually a lot of interesting ways to integrate things here, which is another nice place of working in this space. There's actually been a lot of research work on this. Like when you talk about wacky ideas, there's actually been a lot of really interesting academic work on different ways to imbue physics.
在工业界也可以做疯狂的想法。
You can also do wacky ideas in industry.
是的。
Yeah.
对。但高斯泼溅本身就是小粒子。有很多方法基本上是把物理属性附加到这些泼溅上,说每个都有质量,或者把每个看作通过某种虚拟弹簧与附近邻居耦合,然后你就可以在泼溅之上开始进行物理模拟。所以为这些东西添加物理、动力学或交互的一种途径是预测每个泼溅粒子的物理属性,然后在下游模拟它们,要么使用经典物理,要么使用学到的。在 3D 中工作的美妙之处在于事物可以组合,你可以在不同地方注入逻辑。所以一种方法是,我们生成一个 3D 场景,预测场景中所有东西的 3D 属性,然后使用经典物理引擎模拟交互。或者你可以做这样的事情:由于用户操作,模型将用泼溅或其他表示重新生成整个场景。这可能更通用,因为你不受限于你已经知道如何建模的任何物理属性。但这也更耗费算力,因为你需要响应用户操作重新生成整个场景。但我认为这是一个非常有趣的未来工作领域,也可以添加到潜在的 Marble 2 中,如你所说。
Right. But it's then like Gaussian Splats are themselves little particles. There have been a lot of approaches where you basically attach physical properties to those splats and say that each one has a mass, or maybe you treat each one as being coupled with some kind of virtual spring to nearby neighbors, and now you can start to do physics simulation on top of splats. So one kind of avenue for adding physics or dynamics or interaction to these things would be to predict physical properties associated with each of your splat particles and then simulate those downstream, either using classical physics or something learned. The beauty of working in 3D is things compose and you can inject logic in different places. So one way is we're generating a 3D scene. We're going to predict 3D properties of everything in the scene. Then we use a classical physics engine to simulate the interaction. Or you could do something where, as a result of a user action, the model is now going to regenerate the entire scene in splats or some other representation. That could potentially be a lot more general because then you're not bound to whatever physical properties you know how to model already. But that's also a lot more computationally demanding because then you need to regenerate the whole scene in response to user actions. But I think this is a really interesting area for future work and for adding on to potential Marble 2, as you say.
是的。而且有动力学的机会,对吧?泼溅密度的状态如何?我们能否渲染足够多,以便在放大时获得非常高的分辨率?我们受限于能生成的数量还是能渲染的数量?这些将如何达到超高保真度?
Yeah. And there's opportunity for dynamics, right? What's the state of splats density? Can we render enough to have very high resolution when we zoom in? Are we limited by the amount that you can generate or the amount that we can render? How are these going to get super high fidelity?
你有一些限制,但取决于你的目标用例。我们场景的一个主要限制是我们希望东西能在移动设备和 VR 头显上干净地渲染。这些设备的算力比你在许多其他情况下要少得多。如果你想让一个泼溅文件在四年前的 iPhone 上以 30 到 60 fps 的高分辨率渲染,那么你能处理的泼溅数量就有点受限。但如果你可以在今年的 iPhone 或最近的 MacBook 上工作,或者如果你有本地 GPU,或者你不需要 60 fps 1080p,那么你可以放宽限制,使用更多泼溅,从而在场景中获得更高分辨率。
You have some limitations, but depending on your target use case. One of the big constraints on our scenes is we wanted things to render cleanly on mobile and in VR headsets. Those devices have a lot less compute than you have in a lot of other situations. If you want to get a splat file to render at high resolution, like 30 to 60 fps on an iPhone from 4 years ago, then you are a bit limited in the number of splats you can handle. But if you are allowed to work on a recent iPhone or a recent MacBook, or even if you have a local GPU, or if you don't need 60 fps 1080p, then you can relax the constraints and get away with more splats, which lets you get higher resolution in your scenes.
我期待但没听到的一个用例是具身用例。你现在只关注虚拟的吗?
One use case I was expecting but didn't hear from you was embodied use cases. Are you just focusing on virtual for now?
如果你去 World Labs 主页,有一个叫 Marble Labs 的特定页面。我们在那里展示不同的用例,实际上我们把它们组织成更多视觉效果用例或游戏用例,以及模拟用例。在那里面,我们实际上展示了这是一项可以极大帮助机器人训练的技术。这又回到了我之前谈到的数据匮乏问题。机器人训练确实缺乏数据。高保真真实世界数据绝对非常关键,但你不可能得到大量那样的数据。另一个极端是纯互联网视频数据,但那样你就缺乏训练具身智能体所需的很多可控性。所以模拟和合成数据实际上是一个非常重要的中间地带。我在这个领域工作了很多年,最大的痛点之一是你从哪里获得这些合成模拟数据?你必须策划资产,构建和组合这些复杂情况。在机器人学中,你想要很多不同的状态。你想要具身智能体在合成环境中交互。Marble 实际上有很大潜力帮助生成这些用于具身智能体训练的合成模拟世界。
If you go to World Labs homepage, there is a particular page called Marble Labs. There we showcase different use cases and we actually organize them in more visual effect use cases or gaming use cases, as well as simulation use cases. In that, we actually show this is a technology that can help a lot in robotic training. This goes back to what I was talking about earlier regarding data starvation. Robotic training really lacks data. High-fidelity real-world data is absolutely very critical, but you're just not going to get a ton of that. The other extreme is purely internet video data, but then you lack a lot of the controllability that you want to train your embodied agents with. So simulation and synthetic data is actually a very important middle ground for that. I've been working in this space for many years, and one of the biggest pain points is where do you get this synthetic simulated data? You have to curate assets and build and compose these complex situations. In robotics, you want a lot of different states. You want the embodied agent to interact in the synthetic environment. Marble actually has a lot of potential for helping to generate these synthetic simulated worlds for embodied agent training.
显然那在主页上。它会在那里。我试图建立联系,如你所说,你还得建立商业模式。机器人市场显然非常巨大。也许你不需要那个,或者也许我们需要先建立和解决虚拟世界,然后再进入具身。显然是一个垫脚石。
Obviously that's on the homepage. It'll be there. I was trying to make the link to, as you said, you also have to build a business model. The market for robotics is obviously very huge. Maybe you don't need that, or maybe we need to build up and solve the virtual worlds first before we go to embodied. Obviously a stepping stone.
这还有待决定。我确实认为……
That is to be decided. I do think that...
因为其他人都直接去那里了,对吧?
Because everyone else is going straight there, right?
不是所有人,但有一种兴奋感,我会说。但我认为世界足够大,可以有不同方法。
Not everyone else, but there is an excitement, I would say. But I think the world is big enough to have different approaches.
是的。方法。
Yeah. Approaches.
是的。我的意思是,我们一直认为这是一项相当横向的技术,应该能够随着时间的推移触及许多不同行业。Marble 目前更专注于创意产业,但我认为驱动它的技术应该随着时间的推移适用于许多不同事物。机器人学是其中一个可能迟早会发生的。
Yeah. I mean, we always view this as a pretty horizontal technology that should be able to touch a lot of different industries over time. Marble is a little bit more focused on creative industries for now, but I think the technology that powers it should be applicable to a lot of different things over time. Robotics is one that is maybe going to happen sooner than later.
还有设计,对吧,与创意非常接近。
Also, design, right, is very adjacent to creative.
哦,是的。当然。我认为就像建筑之类的东西。
Oh, yeah. Definitely. I think it's like the architecture stuff.
是的。好的。我的意思是,我在网上开玩笑。我在 Slack 上发了一个视频,说,哦,谁想用 Marble 来规划你的下一次厨房改造?它实际上已经很好用了。只需拍两张厨房的照片,在 Marble 中重建它,然后使用编辑功能看看如果你换台面、换地板或换橱柜,那个空间会是什么样子。
Yes. Okay. Yeah. I mean, I was joking online. I posted this video on Slack of like, oh, who wants to use Marble to plan your next kitchen remodel? It actually works great for this already. Just take two images of your kitchen, reconstruct it in Marble, and then use the editing features to see what that space would look like if you change the countertops or change the floors or change the cabinets.
我们并没有专门为这个用例构建什么,但因为这是一项强大的横向技术,你会得到这些从模型中自然涌现的用例。我们已经有早期测试用户使用 API 密钥,正在为室内设计用例进行构建。
And this is something that we didn't necessarily build anything specific for this use case, but because it's a powerful horizontal technology, you get these emergent use cases that just fall out of the model. We have early beta users using an API key that is already building for interior design use case.
我刚弄完我的车库。我早该知道这个的。我得……
I just did my garage. I should have known about this. I got to...
下次你装修时,我们可以帮忙。
Next time you remodel, we can be of help.
嗯,厨房是下一个,我敢肯定。
Well, kitchen is next, I'm sure.
是的。
Yeah.
我对整个空间智能领域很好奇。我觉得我们应该深入探讨一下。你怎么定义它?和人们可能想到的传统智能之间有什么差距?比如当 Dario 说我们有一个装满爱因斯坦的数据中心时,那是传统智能,不是空间智能。要具备空间智能需要什么?
I'm curious about the whole spatial intelligence space. I think we should dig more into that one. How do you define it and what are the gaps between traditional intelligence that people might think about? Like LLMs when Dario says we have a data center full of Einstein, that's traditional intelligence, it's not spatial intelligence. What is required to be spatially intelligent?
首先,我不理解那句话——'一个装满爱因斯坦的数据中心'。
First of all, I don't understand that sentence, 'a data center full of Einsteins'.
那是个类比。
It's an analogy.
嗯,AI 作为一个领域、一门学科,很大程度上受到人类智能的启发,因为人类是我们目前宇宙中已知最聪明的动物。如果你看看人类智能,它非常多元。有一位心理学家,我想他叫 Howard Gardner,在 1960 年代实际上明确提出了'多元智能'来描述人类智能。其中有语言智能、空间智能、逻辑智能和情感智能。所以对我来说,当我想到空间智能时,我视其为语言智能的补充。我个人不会说空间智能与传统智能对立,因为我不知道'传统'是什么意思。我确实认为空间智能是对语言智能的补充。我们如何定义空间智能?它是让你在空间中推理、理解、移动和交互的能力。我用 DNA 结构推导的例子。当然我简化了这个故事,但其中很大一部分涉及分子和化学键在三维空间中的空间推理,最终推测出双螺旋结构。而人类,或者说 Francis Crick 和 Watson 拥有的那种能力,很难将那个过程简化为纯粹的语言。那是文明的一个巅峰时刻。但每一天,我在这里试图拿起一个杯子。整个过程——看到杯子,看到它所在的背景,看到我自己的手,张开我的手使其几何形状与杯子匹配,并触摸到正确的抓握点——这一切都是深刻的空间性的。这非常困难。我试图用语言来描述它。但另一方面,那种描述性的语言本身并不能让你拿起一个杯子。
Well, a lot of AI as a field, as a discipline, is inspired by human intelligence, because we are the most intelligent animal we know in the universe for now. And if you look at human intelligence, it's very multi-intelligent. There is a psychologist, I think his name is Howard Gardner, in the 1960s actually literally called 'multiple intelligence' to describe human intelligence. And there is linguistic intelligence, there's spatial intelligence, there is logical intelligence and emotional intelligence. So for me, when I think about spatial intelligence, I see it as complementary to language intelligence. I personally would not say it's spatial versus traditional, because I don't know what 'traditional' means. I do think spatial is complementary to linguistic. And how do we define spatial intelligence? It's the capability that allows you to reason, understand, move and interact in space. I use this example of the deduction of DNA structure. Of course I'm simplifying this story, but a lot of that had to do with the spatial reasoning of the molecules and the chemical bonds in a 3D space to eventually conjecture a double helix. And that ability that humans, or Francis Crick and Watson had, is very hard to reduce that process into pure language. And that's a pinnacle of a civilizational moment. But every day, I'm here trying to grasp a mug. This whole process of seeing the mug, seeing the context where it is, seeing my own hand, opening of my hand that geometrically would match the mug and touching the right affordance point. All this is deeply spatial. It's very hard. I'm trying to use language to narrate it. But on the other hand, that narrated language itself cannot get you to pick up a mug.
是的。带宽限制。
Yeah. Bandwidth constraint.
是的。
Yes.
我最近算了一下,如果你每天 24 小时不停地说话,你能生成多少个 token?以平均语速每分钟 150 个词计算,大约每天 21.5 万个 token。而你生活的世界带宽比那高得多。
I did some math recently on if you just spoke all day every day for 24 hours a day, how many tokens do you generate? At the average speaking rate of like 150 words per minute, it roughly rounds out to about 215,000 tokens per day. And your world that you live in is so much higher bandwidth than that.
嗯,我认为确实如此。但如果我们想想艾萨克·牛顿爵士,当时像重力这样的东西还没有被语言形式化,人们本能地在空间上理解物体会下落。但以某种方式将其形式化是有帮助的,比如所有这些不同的规则,我们用语言来捕捉那些经验上和空间上也能理解的东西,但用语言描述更容易。所以我很好奇空间智能和语言智能之间的相互作用。你需要理解有些规则更容易用语言写出来供空间智能理解,但你无法写出'像这样放手,放这么多'。所以我一直好奇如何将两者相互利用。
Well, I think that is true. But if I think about Sir Isaac Newton, you have things like gravity at the time that have not been formalized in language that people inherently spatially understand that things fall. But then it's helpful to formalize that in some way, like all these different rules that we use language to capture something that empirically and spatially you can also understand, but it's easier to describe in a way. So I'm curious about the interplay of spatial and linguistic intelligence. You need to understand some rules are easier to write in language for the spatial intelligence to understand, but you cannot write 'put your hand like this and put it down this amount'. So I'm always curious about how you leverage each other together.
我的意思是,牛顿的例子恰恰说明:牛顿之所以想到写下那些定律,是因为他在世界上有很多具身经验,比如看棒球。
I mean, if anything, the example of Newton: Newton only thinks to write down those laws because he's had a lot of embodied experience in the world watching baseball.
完全正确。实际上,区分你提到的理论构建和具身的、嵌入三维世界的日常经验是有用的。所以对我来说,空间智能在某种程度上封装了那种在三维空间中存在、移动、观察和行动的具身经验。正如你所说,你可以叙述这些事情,但那是一个非常损耗的通道。就像在世界上存在并做事情,与试图描述它,是完全不同的模态。但因为我们是人类,是始终在空间中交互进化的动物,我们甚至不认为那是一件难事。然后我们自然跃升到语言,再到理论构建,作为抽象于那种原生空间理解的机制。在某种意义上,LLM 直接跳到了那些最高形式的抽象推理,这非常有趣且非常有用。但空间智能几乎像是重新打开那个黑箱,并说也许我们直接跳到那种完全抽象的语言、推理和沟通形式时,失去了一些东西。
Exactly. And actually it's useful to distinguish between the theory building that you're mentioning versus the embodied, the daily experience of being embedded in the three-dimensional world. So to me, spatial intelligence is sort of encapsulating that embodied experience of being there in 3D space, moving through it, seeing it, actioning it. And as you said, you can narrate those things but it's a very lossy channel. It's just like the notion of being in the world and doing things in it is a very different modality from trying to describe it. But because we as humans are animals who have evolved interacting in space all the time, we don't even think that that's a hard thing. And then we sort of naturally leap to language and then theory building as mechanisms to abstract above that sort of native spatial understanding. And in some sense, LLMs have just jumped all the way to those highest forms of abstracted reasoning, which is very interesting and very useful. But spatial intelligence is almost like opening up that black box again and saying maybe we've lost something by going straight to that fully abstracted form of language and reasoning and communication.
你知道吗,作为视觉科学家,我总觉得视觉被低估了,因为对人类来说它毫不费力。婴儿睁开眼睛,就开始看到世界。我们天生就具备这种能力。
You know, it's funny as a vision scientist, I always find that vision is underappreciated because it's effortless for humans. You open your eyes as a baby, you start to see your world. We're somehow born with it.
我们几乎是天生就具备的。但学习语言需要付出努力,包括学习如何书写、如何掌握语法、如何表达,这让它感觉很难。而大自然实际上花了更多时间优化的东西,即感知和空间智能,却被人类低估了。
We're almost born with it. But you have to put effort in learning language, including learning how to write, how to do grammar, how to express, and that makes it feel hard. Whereas something that nature spent way more time actually optimizing, which is perception and spatial intelligence, is underappreciated by humans.
有证据表明我们天生就具备吗?你说了'几乎是天生的'。所以听起来我们实际上是在出生后学习的。我们出生时,视觉敏锐度较低,感知能力确实会增强。但大多数人类天生就有视觉能力,大多数人类天生就有将感知与运动联系起来的能力。运动本身需要一段时间来完善。动物则令人难以置信。今年夏天早些时候我刚去过非洲。那些小动物出生后几分钟内就必须行动起来,否则狮子就会抓住它们。
Is there proof that we are born with it? You said 'almost born'. So it sounds like we actually do learn after we're born. When we are born, our visual acuity is less and our perceptual ability does increase. But most humans are born with the ability to see, and most humans are born with the ability to link perception with motor movements. The motor movement itself takes a while to refine. And animals are incredible. I was just in Africa earlier this summer. These little animals, they're born and within minutes they have to get going, otherwise the lions will get them.
在自然界中,优化感知、空间智能和语言用了 5.4 亿年。对语言发展最慷慨的估计大概是 50 万年。
And in nature, it took 540 million years to optimize perception and spatial intelligence and language. The most generous estimation of language development is probably half a million years.
哇,这比我猜的要长。好吧,我确实很慷慨。
Wow. That's longer than I would have guessed. Well, I'm being very generous.
是的。我在读你的书时意识到,与我们播客中讨论的内容有一个有趣的关联,那就是语言模型基准测试,以及 Wow Grand 如何引入了所有这些需要空间智能的物理不可能性。比如 A 在 B 上面,所以 A 不能穿过 B 掉下去。这对我们来说是显而易见的,但对语言模型来说却可能发生。我不知道,也许这是下一个词预测的一部分。
Yeah. No, I was going through your book and I realized that one of the interesting links to something we covered on the podcast is language model benchmarks and how wow grand actually put in all these physical impossibilities that require spatial intelligence. Like A is on top of B, therefore A cannot fall through B. It's obvious to us, but to a language model it could happen. I don't know, maybe it's like a part of the next-token prediction.
这正是我所说的拆解这种抽象。如果你对整个世界的模型只是依次说出单词序列,那真的很难理解为什么不行。
And that's sort of what I mean about unwrapping this abstraction. If your whole model of the world is just saying sequences of words after each other, it's really hard to see why not.
这其实不公平,对吧?但对我们来说显而易见的原因是我们内部将其映射回我们熟悉的三维世界表征。问题是这有多难,需要多长时间才能从你的世界模型提炼到语言模型中?因为我们确实希望我们的模型拥有空间智能。我们是否必须完全抛弃语言模型才能做到这一点?
It's actually unfair, right? But then the reason it's obvious to us is because we are internally mapping it back to some three-dimensional representation of the world that we're familiar with. The question is how hard is it, how long is it going to take us to distill from your world models into a language model? Because we do want our models to have spatial intelligence. And do we have to throw language model out completely in order to do that?
不,我不这么认为。我认为它们是多模态的。即使我们今天的 Marble 模型也以语言作为输入。所以它是深度多模态的。我认为在许多用例中,这些模型将协同工作。也许有一天我们会有一个通用模型。
No, I don't think so. I think they are multimodal. Even our model Marble today takes language as input. So it's deeply multimodal. And I think in many use cases these models will work together. Maybe one day we'll have a universal model.
即使你这样做,也有一个实用性的问题:人们使用语言,并希望用语言与系统交互。即使从实用角度来说,构建让人们能够与之对话的系统、产品和模型也是有用的。所以我不认为这会消失。我认为有一种智力上的好奇心:你能构建一个只使用视觉或只使用空间智能的模型到什么程度?我不知道这在实践中是否有用,但看看你能把它推到多远会是一个有趣的智力或学术练习。我很好奇:如果你有一个高度精确的世界模型,并且没有给它任何我们当前对标准模型物理学的理解,它能从零开始重建多少?它需要什么水平的语言理解?因为我们使用了很多符号,但也许我们会得出一个非常不同的模型,但仍然准确。我想知道我们在多大程度上受到了限制。人们说人类总是需要像人类一样,因为世界是为人类建造的。在某种程度上,我们构建语言的方式限制了来自其他模态的一些输出。所以我非常期待关注你的工作。
Even if you do, there's a pragmatic thing where people use language and want to interact with systems using language. Even pragmatically, it's useful to build systems, products, and models that let people talk to them. So I don't see that going away. I think there's an intellectual curiosity of saying how much could you build a model that only uses vision or only uses spatial intelligence. I don't know that that would be practically useful, but it would be an interesting intellectual or academic exercise to see how far you could push that. I'm curious: if you had a highly precise world model and you didn't give it any notion of our current understanding of the standard model of physics, how much of it would it be able to come up with and recreate from scratch? And what level of language understanding would it need? Because we have so many notations that we use, but maybe we'll come up with a very different model and still be accurate. I wonder how much we're limited. People say humans always need to be like humans because the world is built for humans. In a way, the way we build language constrains some of the outputs from these other modalities. So I'm super excited to follow your work.
是的。还有另一种方式:你甚至不需要做 AI 来回答这个问题。你可以发现外星人,看看他们有什么样的物理学。他们可能有不同的物理学。
Yeah. There's another way: you don't even need to be doing AI to answer that question. You could discover aliens and see what kind of physics they have. They might have a different physics.
面对现实吧,我们是迄今为止宇宙中最聪明的动物,对吧?所以这是一个非常有趣的问题:我们对宇宙的知识和对物理学的理解是否在某种程度上受到我们自身认知或技术进化路径依赖的约束?一种做实验的方式是:如果我们重新运行人类文明,我们会以相同的顺序得出相同的物理学吗?我不认为这是一个非常可行的实验。
Let's face it, we are so far the smartest animal in the universe, right? So that is a really interesting question: is our knowledge of the universe and our understanding of physics constrained in some way by our own cognition or by the path dependence of our own technological evolution? One way to sort of do an experiment is to say: if we were to rerun human civilization again, would we come up with the same physics in the same order? I don't think that's a very practical experiment to run.
我想知道人们能否做一个实验:我们有大量关于行星或天体运动的天体物理数据。只需将数据输入模型,看看牛顿定律是否会涌现。我的猜测是可能不会。牛顿定律的抽象层次与这些 LLM 所代表的层次不同。
One experiment I wonder if people could run is that we have plenty of astrophysical data on the planet or celestial body movements. Just feed the data into a model and see if Newtonian law emerges. My guess is it probably won't. It's not the abstraction level of Newtonian law is at a different level from what these LLMs represent.
是的。所以我不惊讶,给定足够的天体运动数据,LLM 实际上能预测相当准确的运动轨迹。假设我发明一颗围绕恒星的行星,给足够的数据,我的模型会告诉你第一天它在哪,第二天它在哪。我不会惊讶。但 F=ma 或作用力等于反作用力,那完全是不同的抽象层次。这超出了今天的 LLM。
Yeah. So I wouldn't be surprised that given enough celestial movement data, an LLM would actually predict pretty accurate movement trajectories. Let's say I invent a planet surrounding a star. Giving enough data, my model would tell you on day one where it is, day two where it is. I wouldn't be surprised. But F equals MA or action equals reaction, that's a whole different abstraction level. That's beyond today's LLM.
好的。你需要什么样的模型才能让它不是地心模型?因为如果我只用视觉数据训练,你会认为太阳绕着地球转,这很合理,对吧?但显然事实并非如此。那么它如何学会这一点?我很好奇我们谈论的所有这些力。有时也许你不需要它们,因为只要看起来对就是对的。但当你尝试用这些模型做更高级的任务时,我们能在多大程度上依赖它们?我认为你需要一种不同的学习范式。这里有些混淆:LLM、语言和符号与人类理论构建和人类物理学。它们非常不同,因为 LLM……人类的目标函数是理解世界并在生活中茁壮成长。你这样做的方式是观察数据,思考它,尝试在世界上做某事,当它不符合你的期望时,你在线更新你对世界的理解。人们一直在这样做。例如,我认为我的钥匙在楼下,所以我下楼去找,没看到,哦不,它们实际上在我的卧室里。所以因为我们不断与世界互动,我们不断需要构建关于周围世界正在发生什么的理论,然后证伪或为这些理论添加证据。我认为这种过程放大并规模化后,就给了我们 F=ma 和牛顿物理学。
Okay. What model would you need to not have it be a geocentric model? Because if I'm training just on visual data, it makes sense that you think the sun rotates around the earth, right? But obviously that's not the case. So how would it learn that? I'm curious about all these forces we talk about. Sometimes maybe you don't need them because as long as it looks right it's right. But as you make the jump to trying to use these models to do more high-level tasks, how much can we rely on them? I think you need a different learning paradigm. There's a bit of conflation here: LLMs and language and symbols versus human theory building and human physics. They are very different because an LLM... The human objective function is to understand the world and thrive in your life. The way you do that is by observing data, thinking about it, trying to do something in the world, and when it doesn't match your expectations, you update your understanding of the world online. People do this all the time. For example, I think my keys are downstairs, so I go downstairs and look for them, and I don't see them, and oh no, they're actually up in my bedroom. So because we're constantly interacting with the world, we're constantly having to build theories about what's happening around us and then falsify or add evidence to those theories. I think that kind of process writ large and scaled up is what gives us F=MA and Newtonian physics.
我认为这与我们训练的模型模态(无论是语言还是空间)有点正交。我的理解是,这几乎是一种更高效的学习方式:你有一个假设,即根据现有数据可能存在哪些不同的世界,然后你通过实验排除不可能的世界,最终确定正确的那个。对我来说,这也是我拥有心智理论的方式:我对你的想法有几个假设,然后我尝试采取行动来验证或检验我的直觉。显然,LLM 不会做这些。
And I think that's a little orthogonal to the modality of model that we're training, whether it's language or spatially. The way I put it is almost like this is more efficient learning because you have a hypothesis of here are the different possible worlds that are granted by my available data and you then do experiments to eliminate the worlds that are not possible and you resolve to the one that's right. To me, that's also how I have theory of mind, which is like I have a few hypotheses of what you're thinking, and I try to create actions to resolve that or check my intuition as to what you're thinking. Obviously LLMs don't do any of these.
心智理论可能还会延伸到情商,而今天的 AI 根本没有触及这一点。我们确实需要它。人们开始过度依赖这些东西,这完全是另一个辩论话题。我不得不问,因为很多人把这个问题抛给了我们。我们需要抛弃多少?序列到序列建模过时了吗?注意力机制过时了吗?我们到底要在多大程度上重新要求一切?
A theory of mind possibly also will break into even emotional intelligence, which today's AI is really not touching at all. Right. And we really need it. People are starting to depend on these things probably too much, and that's a whole topic of other debate. I do have to ask because a lot of people have sent this to us. How much do we have to get rid of? Is sequence-to-sequence modeling out the window? Is attention out the window? Like how much are we re-requesting everything?
我认为你应该坚持有效的东西。注意力机制仍然存在。很多事情不需要修复未损坏的部分。世界上有很多难题需要解决,但让我们一次专注于一个。我认为思考新架构、新范式或截然不同的学习方式非常有趣。但你不必因为研究新模态就抛弃一切。
I think you stick with stuff that works. I think attention is still there. There's a lot like you don't need to fix things that aren't broken. There's a lot of hard problems in the world to solve, but let's focus on one at a time. I think it is pretty interesting to think about new architectures or new paradigms or drastically different learning ways to learn. But you don't need to throw away everything just because you're working on new modalities.
我认为序列到序列实际上存在于世界模型中。我们将会看到超越序列到序列的算法或架构。
I think sequence-to-sequence is actually in world models. I think we are going to see algorithm or architecture beyond sequence-to-sequence.
哦,但这里我认为存在一些技术上的混淆,Transformer 已经为我们解决了这个问题。Transformer 实际上不是序列模型。Transformer 本质上是一个集合模型。这非常强大,但因为许多 Transformer 源于基于循环神经网络的早期架构,而 RNN 确实有内置的架构偏差,它们建模一维序列。
Oh, but here I think there's a little bit of technological confusion, and transformers already solved that for us. Transformers are actually not a model of sequences. A transformer is natively a model of sets. That's very powerful, but because a lot of the transformers grew out of earlier architectures based around recurrent neural networks, and RNNs definitely have a built-in architectural bias that they model one-dimensional sequences.
但 Transformer 只是集合的对象模型,它们可以建模许多集合,这些集合可以是一维序列,也可以是其他东西。
But transformers are just object models of sets, and they can model a lot of those sets could be 1D sequences, they could be other things as well.
所以你字面意思是集合论?
So you literally mean set theory?
是的。所以 Transformer 实际上不是词元序列的模型。Transformer 实际上是词元集合的模型。唯一将顺序注入标准 Transformer 架构的是你给词元的位置嵌入。如果你选择给它一维位置嵌入,那是模型知道它是一维序列的唯一机制。但 Transformer 块内的所有操作要么是词元级的,比如 FFN、QKV 投影、每个词元的归一化,都是独立对每个词元进行的。然后通过注意力机制在词元之间进行交互,但这也是置换等变的。所以如果我置换我的词元,注意力操作符会以完全相同的方式得到置换后的输出。所以它实际上是词元集合的原生架构。
Yeah. Yeah. So a transformer is actually not a model of a sequence of tokens. A transformer is actually a model of a set of tokens. The only thing that injects the order into the standard transformer architecture is the positional embedding that you give the tokens. If you choose to give it a 1D positional embedding, that's the only mechanism the model has to know it's a 1D sequence. But all the operators inside a transformer block are either token-wise, like FFN, QKV projections, per token normalization, all happen independently per token. Then you have interactions between tokens through the attention mechanism, but that's also permutation equivariant. So if I permute my tokens, the attention operator gets a permuted output in exactly the same way. So it's actually natively an architecture of sets of tokens.
字面意义上的变换。
Literally a transform.
是的。
Yeah.
我知道我们时间不多了,但我们想让你说几句,呼吁大家行动,比如哪些人愿意在 World Labs 工作,应该申请什么样的人,人们在 World Labs 之外应该做哪些对你有帮助的研究,或者你还有什么想法。
I know we're out of time, but we just want to give you the floor for some call to action either on people that would enjoy working at World Labs, what kind of people should apply, what research people should be doing outside of World Labs that would be helpful to you, or anything else on your mind.
我确实认为这是一个激动人心的时刻,可以超越语言模型,思考空间智能的无限可能性。所以我们非常渴望人才,从像 Justin 刚才描述的那样思考问题、训练大型世界模型的深度研究人员,到优秀的工程师,构建从训练优化到推理再到产品的系统。我们也渴望优秀的商业、产品思考者以及市场推广和商业人才。所以我们渴望人才。特别是现在,我们通过 Marble 将模型展示给世界,我认为我们有机会与更广泛的人才合作,解决模型问题,并向世界交付最好的产品。
I do think it's a very exciting time to be looking beyond just language models and think about the boundless possibilities of spatial intelligence. So we are actually hungry for talent ranging from very deep researchers thinking about the problems like Justin just described, training large models of world models. We are hungry for engineers, good engineers building systems from training optimization to inference to product. And we're also hungry for good business, product thinkers, and go-to-market and business talents. So we are hungry for talent. We especially now that we have exposed the model to the world through Marble, I think we have a great opportunity to work with an even bigger pool of talent to solve both the model problem as well as deliver the best product to the world.
是的,我也很兴奋让人们尝试 Marble 并用它做很多酷炫的事情。我认为它有很多非常酷的功能,很多非常酷的特性很好地结合在一起。
Yeah, I think I'm also excited for people to try Marble and do a lot of cool stuff with it. I think it has a lot of really cool capabilities, a lot of really cool features that fit together really nicely.
在来这里的车上,Justin 和我说人们还没有完全发现一些高级编辑模式。比如打开高级模式。你可以像 Justin 说的那样,改变瓶子的颜色,改变地板,改变树木。
In the car coming here, Justin and I were saying people have not totally discovered some of the advanced mode of editing. Like turn on the advanced mode. You can, like Justin said, change the color of the bottle, change your floor, and change the trees.
嗯,我实际上尝试过,但当它说创建时,它只是让我创建一个完全不同的世界。
Well, I actually tried to get there, but when it says create, it just makes me create a completely different world.
你需要点击高级模式。这是一个很好的 UI。
You need to click on the advanced mode. It's a good UI.
我们可以改进我们的 UI。记得点击。
We can improve on our UI. Remember to click.
是的,我们需要招聘人员。我们致力于产品。
Yeah, we need to hire people. We work on the product.
但有一点我们从你们身上看得很清楚,那就是智力上的无畏,我认为这是你们坚持的原则。
But one thing we got that was clear from you guys is also intellectual fearlessness, which is something that I think you guys hold as a principle.
是的,我的意思是,我们实际上是第一批在模型端和产品端都尝试这样做的人。
Yeah, I mean we are literally the first people who are trying this both on the model side as well as on the product side.
非常感谢你们加入我们。这很有趣。
Thank you so much for joining us. This was fun.
是的,谢谢邀请我们。
Yeah, thanks for having us.