World Models: The Next Frontier in AI
打开互动全文版(中英对照 + 朗读 + 问答)→贾斯汀·约翰逊探讨了为什么世界模型是 AI 的下一个前沿,超越语言模型,理解和模拟世界。
Justin Johnson discusses why world models are the next frontier in AI, moving beyond language models to understand and simulate the world.
构建更强大 AI 的竞赛不仅仅关乎让语言模型变得更大。越来越多地,它关乎世界模型——那些理解空间、预测环境如何变化并在周围世界中行动的系统。Justin Johnson 正在帮助塑造这一转变。他与 FA Lee 共同创立了 World Labs,并且是密歇根大学的计算机科学副教授。我问他为什么这么多研究者认为世界模型是 AI 的下一个前沿。
The race to build more capable AI isn't just about making language models bigger. Increasingly, it's about world models, systems that understand space, predict how environments change, and act in the world around them. Justin Johnson is helping shape this shift. He co-founded World Labs with FA Lee and is an associate professor of computer science at the University of Michigan. I asked him why so many researchers think world models are AI's next frontier.
在这个领域的许多研究者中,有一个共同的底层信念:语言模型没有做到某些事情,而我们应该构建其他类型的模型来做这些事情。这些模型做的是其他类型的事情。这涉及到理解世界、生成世界、模拟世界、重建世界、在世界中规划行动。这些都是我们希望模型具备的能力。我们为什么关心这个?因为我们想要构建的系统不仅仅是困在终端里或作为虚拟智能体,对吧?你希望 AI 系统的愿景是成为机器人,在现实世界中行动,或者我们可能想要构建虚拟世界,生活在其中,并在那里模拟有趣的事情。所以所有这些能力,感觉都不是从语言建模范式中自然涌现出来的。
There's a shared like low-level belief among many researchers in the field that there's something that language models aren't doing but that there are other kinds of models that we should be building. They do other kinds of things. And that's something around understanding the world, generating worlds, simulating worlds, reconstructing worlds, planning actions through worlds. These are all capabilities that we want to build models to have. And why do we care about this is because we want to build systems that are not just stuck in a terminal or stuck as a virtual agent, right? You want to have visions of AI systems that are going to be robots that are out in the world acting in the world, or we might want to build virtual worlds and live in those and simulate interesting things there. So all of these are capabilities that really don't feel like they're falling naturally out of the language modeling paradigm.
我是 Sam Cherington,这里是 TwiML AI 播客。十多年来,我一直在通过这样的对话探索塑造 AI 未来的想法和创新,帮助你理解什么是真实的、什么是下一步、什么才是重要的。让我们开始吧。
I'm Sam Cherington and this is the TwiML AI podcast. For over a decade, I've been exploring the ideas and innovation shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
在过去几年里,我认为有一个被广泛支持的观点,基本上是说世界模型是语言模型或扩散模型的一种涌现属性。比如,扩散模型可以展示世界的某些物理属性,或者语言模型可以讲述一些似乎对世界有所了解的故事。我在这里的问题有几个。一个是要求给出世界模型的具体定义,因为我认为在早期的对话中,世界模型确实是在谈论这些模型对现实世界——我们生活的这个世界——的基础知识。最近,我觉得世界模型的讨论已经转向谈论能够创建人工世界,但让它们自洽、可导航,以及诸如此类的属性。所以我很想听听你详细阐述这两者之间的关系,以及你是否看到了术语上的这种转变。但还有这个涌现的概念,以及我们是否需要新的东西?或者如果我们给现有模型足够的数据、算力等,我们能否达到目标?或者更可能的是,为什么那样达不到?
Over the past few years, one idea that I think has been espoused is essentially the idea that a world model is like an emergent property of either language models or diffusion. Like if diffusion can illustrate some physical properties of the world, or language models can tell stories that seem like they know something about the world. I guess there's a couple questions emerging in my question here. One is maybe asking for a concrete definition of a world model, because I think early on in those conversations, world model was really talking about these models having foundational knowledge of the real world, like the world that we live in. More recently, the world model conversation I feel like has shifted to talking about being able to create artificial worlds but have them be self-consistent and navigable and properties like that. So I'd love to hear you elaborate on the relationship and if you see the same shift in the terminology. But also this idea of emergence, and do we need new things? Or if we throw enough data, compute, etc. at the models that we have, does that get us there? Or more likely, why won't that get us there?
我认为那里有很多有趣的问题需要展开。其中一个,我的意思是最大的一个,就是让我们先把它说清楚。领域内并没有一个大家都同意的世界模型的清晰定义。我认为这造成了部分困惑,对吧?并没有一个东西,我们可以说一个具有 X 属性或产生 X 类输入和 X 类输出的模型,从定义上就是世界模型。我认为作为一个领域,我们对我们所说的并没有一个清晰的定义,这导致了困惑。但针对你的观点,我认为有几个变体感觉像是……我认为有你提到的某种隐式世界知识的概念。还有其他类型的模型产生特定类型的输入和输出。也许一个产生文本的模型,如果它产生了正确类型的文本,它能够产生这种文本答案的唯一方式是因为它对现实世界有所了解,或者因为它在神经网络权重内部建模了某种隐式世界。或者类似地,对于视频模型,对吧?如果我能够生成一个超级逼真的视频,有详细的物理效果,水以非常特定的方式流动,也许隐式地,模型必须已经建模了关于现实世界的某些东西才能生成那种输出。所以我认为这就是隐式世界模型的概念。你可以做任何数量的任务,但有某些类型的答案,你给出这些答案就需要隐式地建模关于世界的某些东西。但然后我认为还有另外两条线索非常有趣。我认为世界模型这个术语本身实际上可以追溯到强化学习文献,有一个与 POMDP 相关的特定技术定义,也许我们稍后会讨论。但世界建模这个特定的技术术语在强化学习中已经存在相当长的时间了。然后我认为还有另一条线索,就是我们在过去几年里一直在大量讨论生成模型。图像模型是产生图像的模型。视频模型是产生视频的模型。然后也许世界模型应该是产生世界的模型。那么产生或生成一个世界意味着什么?所以现在我认为社区中不同的人使用这三个不同的线索:某种隐式世界建模的概念,我可以回答非常难的问题,但我必须建模关于世界的某些东西才能得到正确答案;或者特定的 RL 世界模型公式;然后是生成或创造世界的生成模型。
I think there's a lot of interesting questions to unpack there. One, I mean the biggest one is just let's get it out of the way. There isn't a clear definition of world models that everyone in the field agrees on. And I think that's causing part of the confusion, right? There is not like a thing where we can say a model that has X property or produces X kind of input and produces X kind of output is a world model definitionally. I think we don't have that crisp definition as a field of what we mean, which leads to the confusion. But to your point, I think there's a couple variants of this that feel like they're... I think there is some notion of implicit world knowledge that you mentioned. There are other kinds of models that produce certain kinds of inputs and outputs. Maybe if a model that is producing text, if it produces the right kind of text, the only way it could have produced this kind of text answer is because it knows something about the real world, or because it's modeling some kind of implicit world internally in its neural network weights. Or similarly for video models, right? If I am able to generate a video that is super photorealistic and has detailed physics, and has water running in very particular ways, maybe implicitly the model must have been modeling something about a real world in order to generate an output of that kind. So I think that's kind of this notion of an implicit world model. You could be doing any number of tasks, but there's certain kinds of answers that you could give that would require implicitly modeling something about the world. But then I think there's like two other threads there that are really interesting. I think the term itself, world model, actually goes back to reinforcement learning literature, and there's a specific technical definition related to POMDPs that maybe we'll get into later. But there's a specific technical term of world modeling that goes back to reinforcement learning for quite some time. And then there's another thread I think that ended up is like, we've been talking a lot about generative models the last couple of years. An image model is a model that produces images. A video model is a model that produces videos. And then maybe a world model should be a model that produces worlds. And then what does it mean to produce or generate a world? So now I think we've got these three different threads that different people in the community use: some notion of implicit world modeling that I can answer really hard problems but I must be modeling something about a world to get the right answer, or the specific RL formulation of world model, and then the generative model that creates or generates worlds.
你认为这种隐式世界模型只是另一种定义,还是一组错误的信念?当你听到这个时,你觉得那些是世界模型吗?你认为这是思考世界模型的有效方式,还是你认为世界模型具有一些属性,而这些属性并不是我们正在讨论的当前模型(如视觉语言等)所特有的?
Do you see this implicit world model as just another definition or like an incorrect set of beliefs? Like when you hear that, do you feel like those are world models? Do you think that's a valid way of thinking about world models, or do you think world models have properties that aren't really characteristic of the current models that we are talking about, like vision language etc.?
我几乎认为它们都是有效的。我的意思是,困难在于我认为我刚才说的三件事都是非常有趣的系统。它们都是非常有趣的模型。作为社区,我们可能应该构建所有这些模型。但我们可能应该想出更好的术语,这样我们就不会因为用同一个术语称呼不同的系统而互相混淆。我认为就是这样。
I almost think like I think they're all valid. I mean, the hard part is that I think all three of the things I just said are very interesting systems. Like they're very interesting models. Like we should probably build all of them as a community. But we should probably come up with better terms so that we don't confuse each other by calling different systems by the same term. And I think that's the thing.
所以我觉得在去年的学术文献里,世界模型更多地凝聚成了一种特定类型的实时交互视频模型,这也是最近一年文献中人们通常所说的世界模型。但我认为我们刚才讨论的所有其他特性都非常有趣且有用。尤其是隐式世界模型的概念,我觉得它可以应用于任何事物。无论你处理什么类型的数据,无论你构建什么类型的系统,如果它达到了一定程度的有趣复杂性,那么系统中某个地方最终会存在某种隐式世界模型的概念。这真的很有趣。我想到的一个方式是,把它放在我们围绕大语言模型所纠结的问题背景下思考:它们真的理解语言吗?下一个词预测只是在假装理解语言,还是真的理解语言?我认为在大多数情况下,也许我们已经不再纠结这个问题,而是说这些东西太神奇了,这真的重要吗?从这个角度看,如果我们把这个问题应用到世界模型上,如果那些类型的模型在生成与某个底层世界一致的结果方面变得如此出色,也许这并不重要,但我仍然觉得这是一个有趣的哲学问题。
So I think in the academic literature over the last year, the world model has coalesced more into a particular flavor of a real-time interactive video model, and that's been more commonly what people call world models in the literature recently. But I think all the other properties we talked about are really interesting and useful. Especially the implicit world model notion, I think that can be applied to anything. No matter what kind of data you're processing, no matter what kind of system you're building, if it gets to some level of interesting complexity, it's going to end up with some notion of implicit world model somewhere in the system. That's really interesting. One way that occurs to me is thinking about it in the context of the question we grappled with around large language models: do they really understand language? Is next-token prediction just faking an understanding of language, or is it an understanding of language? I think for the most part, maybe we've moved on from that question and said these things are so amazing, does it really matter? From that perspective, if we apply that to the world model question, if those types of models get so good at generating results consistent with some underlying world, maybe it doesn't really matter, but I still find it an interesting philosophical question.
我同意这里面有一个有趣的哲学问题,但在某种程度上它是无法回答的,对吧?作为科学家,你倾向于思考我能对这个系统测量什么,我能提出哪些具体问题,或者理想情况下能否证伪一个系统。但我认为你提到的另一点是,有些人在谈论世界模型时有时会有另一个相当不同的概念,而我认为我们不知道如何达到那个目标。那更像是作为理论构建者的世界模型,对吧?因为我们有这样的概念:作为人类,我们在穿越这个非常复杂的世界并拥有这些经历,但我们并不是仅仅让光子落在视网膜上,然后顺其自然。我们不断地在脑海中构建关于正在发生的事情的理论。我们提出了这些伟大的理论,关于引力如何运作、物理学如何运作、流体动力学如何运作。我们不仅仅是观察这些事物;我们实际上最终得到了非常简洁而强大的理论,解释了外部正在发生的一切背后的机制。我认为围绕 AI 的圣杯问题之一是,我们如何让机器也做这类事情?也许有这样的概念:世界模型不应该只是直接思考观察结果,而应该构建关于世界的深层解释性理论。我认为这真的非常非常难。我不知道是否有人有很好的角度来达到那个目标,但我认为这是问题的一个略有不同的概念,有时会被混入其中。
I agree there's an interesting philosophical question in there, but at some point it's sort of unanswerable, right? As a scientist, you want to think about what are things I can measure about this system, what are concrete questions I can ask or falsify ideally about a system. But I think the other thing you're getting at is there is another notion that some people sometimes have when talking about world models, which is pretty different, and I don't think we know how to get there. That's more like world model as theory builder, right? Because we have this notion that as humans, we're kind of traversing this really complicated world and having these experiences, but we're not just letting the photons fall on our retinas and letting it happen. We're constantly building theories in our mind for what's happening. We come up with these great theories about how gravity works, how physics works, how fluid dynamics works. It's not just that we observe these things; we actually end up with very compact and powerful theories that explain all the mechanisms behind what's happening out there. I think part of the holy grail question around AI is how do we get machines to do that kind of thing too? Maybe there's this notion that a world model should not just be directly thinking about observations but should be building deep explanatory theories about the world. I think that one's really, really hard. I don't know that anyone has a great angle on how to get there, but I think that's a slightly different notion of the question that sometimes gets mixed in.
是的,你可以说,甚至在世界模型之前,如果我们能获得一个能够构建关于语言的深层理论的模型,那将是很好的,或者任何其他当前由我们处理的模型类型所处理的领域——语言、图形艺术。有趣。好的。但有一件事,比如,如果我有一个完美的——我的意思是,语言模型有点像,想象一个完美的语言模型,但也许语言模型不是正确的例子。假设你有一个完美的视频模型,可以生成你要求的任何视频。但也许我想要的实际上不是视频。作为人类、作为科学家,我想要的是了解世界的一些东西。即使我能生成任何类型或任何结构的视频,如果我没有学到我想了解的世界的东西,那么也许我想要的视频是爱因斯坦做一个讲座,解释量子引力的新理论之类的。那么实际上并不是像素本身,不是生成视频的能力,那才是我真正想要的。我想以某种方式从这个模型中获得对世界的新理解。而这真的很难。
Yeah, you could argue that even before worlds, it would be great if we could get, for example, a model that can build a deep theory about language, or any other domain that is currently handled by the types of models we deal with currently—language, graphic arts. Interesting. Okay. But there's something like, what if I had this perfect—I mean, an LM is kind of like, imagine a perfect LM, but maybe LM is the wrong example. Suppose you had a perfect video model that could generate any video you asked for. But maybe what I wanted wasn't actually a video. What I wanted as a human, as a scientist, was to learn something about the world. Even if I can generate video of any kind or any structure, if I didn't learn what I wanted about the world, then maybe the video I want is like Einstein giving a lecture explaining the new theory of quantum gravity or something like that. Then it's not actually the pixels themselves, not the capability to generate video, that was what I really wanted. I wanted to gain some new understanding of the world from this model somehow. And that's a really hard one.
是的。我不认为语言模型是一个糟糕的例子,对吧?如果一个语言模型有能力生成理论,你可以说它不会产生幻觉,因为它会更深入地思考它生成的事物之间的关系,并且会自我纠正或能够自我纠正。所以我们讨论了隐式世界模型,我们讨论了生成式世界模型。你提到了这种状态机解释——POMDP,部分可观察马尔可夫决策过程。请稍微谈谈这个,以及这段历史如何影响你和其他人对世界模型的思考。
Yeah. And I don't think LLMs are a bad example, right? If an LLM had the ability to generate theories, you could argue that it wouldn't hallucinate because it would think more deeply about the relationship between the things that it's generating and would kind of self-correct or could self-correct. So we talked about implicit world models, we talked about generative world models. There's this state machine interpretation that you mentioned—POMDP, partially observable Markov decision processes. Talk a little bit about that and how that history kind of plays into the way you and others are thinking about world models.
所以有一个叫做部分可观察马尔可夫决策过程(POMDP)的抽象概念,它由来已久。这是一个非常好的数学形式体系,用于思考智能体如何与世界互动。你基本上将系统分解为两部分。一部分是世界,它就像你周围发生的一切。然后是智能体,智能体是能够在世界中采取行动的东西。它们可以移动,可以捡起物体,可以在世界中或对世界做事。所以你把整个宇宙划分为智能体,它移动并做事,以及世界,它被智能体做事或由智能体做事,并且也可能随时间演化。所以你有世界和智能体,然后有这个形式体系来描述两者如何互动。智能体将采取行动,而在不同情况下,你的行动词汇可能不同,对吧?也许你是一个机器人,你可以以某种方式驱动你的电机。也许你在一个视频游戏中,你可以按控制器上的按钮。所以在不同情况下,智能体可能有不同类型的可用行动。但无论情况如何,当智能体对世界采取行动时,世界将以某种方式改变。
So there's this abstraction called partially observable Markov decision processes, or POMDPs, that goes back quite a long time. It's a really nice mathematical formalism for thinking about how agents can interact with worlds. You basically decompose your system into two parts. One is the world, and that's like everything that happens around you. Then there's an agent, and an agent is something that can take actions in the world. They can move around, they can maybe pick up objects, they can do things in the world or to the world. So you partition the whole universe into agent, which moves around and does stuff, and then world, which has stuff done to it or by the agent, and also maybe evolves in time. So you've got the world and the agent, and then there's this formalism of how the two interact with each other. The agent is going to take actions, and in different situations your vocabulary of actions might be different, right? Maybe you're a robot and you can actuate your motors in a certain way. Maybe you're in a video game and you can push buttons on the controller. So in different situations, the agent might have available different kinds of actions. But whatever the situation is, when an agent makes actions on the world, the world will be changed in some way.
我们用这样的方式来表示:世界有一个内部状态,这个状态定义了世界的一切。状态可能非常大、非常复杂,可能无法理解,可能是非常高维的,但它本质上是对世界正在发生什么的完整解释。然后智能体对世界采取行动,因为行动改变了世界,所以行动会导致状态在某种程度上发生变化或转移。但状态太大了,非常复杂,你无法直接观察它。所以智能体得到的是观察,观察是完整世界状态的某种低维投影。这在不同的情境下含义不同。作为人类,我们得到的观察是我们眼睛看到的图像、耳朵听到的声音、身体感受到的触觉。这些都是我们获得的感官信号。这些感官信号告诉我们一些关于世界的信息,但它们只告诉我们周围发生的一切中非常稀疏和局部的部分。所以 POMDP 循环就是:智能体对世界采取行动,导致状态转移,然后基于状态,智能体获得观察,告诉它一些关于世界的信息,然后这个循环不断重复,智能体试图在世界中做事。
And the way we denote that is we say that the world has a state internal to it. And the state kind of defines everything that makes the world what it is. The state might be very large and very complex. It might not be understandable. It might be very high dimensional, but it kind of is the full explanation of what's happening in the world. So then the agent makes actions on the world, and because the action changes the world, that means the action will cause the state to change or transition somehow inside the world. But now the state is so big and complicated that you can't directly observe it. So what the agent gets back are observations, which are some kind of low-dimensional projection of the full world state. And what that means is different in different contexts. As humans, the observations we get are the images we see in our eyes, the sounds that come into our ears, the feelings of touch on our bodies. Those are all the sensory signals we get. And those sensory signals tell us something about the world, but they only tell us something very sparse and local about everything that's happening around us. So the POMDP loop is that the agent does actions to the world, which causes the state to transition, then based on the state, the agent gets an observation that tells it something about the world, and then this happens in a loop over and over again as the agent tries to do things in the world.
这听起来很像强化学习的设定。智能体在世界中运作,进行观察,它的行动与某种奖励相关联等等。那么这是否意味着你需要一个强化学习式的设置才能拥有一个稳健的世界模型?或者 POMDP 与模型之间的具体关系是什么?
That sounds a lot like the setting for reinforcement learning. The agent is operating in the world, making observations. There's some reward associated with its actions, etc. Is the implication then that you need a reinforcement learning type setup in order to have a robust world model, or what's the concrete relationship between POMDPs and models?
你问到点子上了。我在最初定义中遗漏的重要技术部分是奖励。这个形式体系是在强化学习的背景下发展起来的。在那里,智能体有一个要实现的目标,它通过奖励信号来了解自己是否在实现目标。所以想法是智能体想要采取能最大化奖励的行动。一旦你到了那个层面,那就是强化学习的设置。这就是这个形式体系的来源。但我刚才选择不谈奖励的原因是,我认为我们可以把 POMDP 的原始抽象提取出来,用在其他情境中。这种关于智能体、状态和观察的思考方式在许多其他情境中变得适用和有用,即使在强化学习之外也是如此。强化学习是这个形式体系最初发展的地方,非常有用,但如今我们也可以在其他地方应用它。
You got me there. The important technical piece I left out of the initial definition was the reward. This whole formalism was developed in the context of reinforcement learning. There, the agent has some goal it's trying to achieve, and it gets a reward signal to know whether it's achieving that goal. So the idea is the agent wants to take actions that maximize its reward. Once you go to that level, that's exactly the reinforcement learning setup. That's where this formalism comes from. But the reason I chose not to talk about the reward a moment ago is that I think we can take that original abstraction of the POMDP and pull it out and use it in other contexts. This notion of thinking about agents, states, and observations becomes applicable and useful in many other contexts, even outside of reinforcement learning specifically. Reinforcement learning is one setting where this formalism was originally developed and is super useful, but we can apply it elsewhere as well nowadays.
给我们一个它在那个情境之外应用的例子。
Give us an example of how it's applied out of that context.
一个非常具体的例子可能是机器人学中的行为克隆。假设你在构建一个机器人,你的目标是让它能在世界中四处走动并做事,也许帮我整理床铺或打扫厨房。一旦那个机器人被部署,它就是一个在世界中互动的智能体。世界有状态,机器人获得关于世界的观察。所以一旦训练完成,它就在 POMDP 循环上运作。但问题是训练信号是什么。你可以通过强化学习来训练它,每次它采取行动就获得奖励。或者你可以通过行为克隆来训练它。也许我有一个大型监督数据集,其中给定这个观察,你应该采取这个行动。然后你可以训练一个监督学习模型,那会非常不同,没有明确的奖励信号。你会使用梯度下降和某种监督损失。所以即使你最终得到一个在类似 POMDP 循环中运作的系统,你也可以有一个纯粹监督学习的训练目标,不需要强化学习。总结一下,POMDP 有两个部分:一个捕捉智能体与世界之间的关系,即智能体可以观察反映某种抽象状态的事物,但永远无法真正知道那个抽象状态,它需要基于观察来运作。奖励和训练都是关于如何将观察转化为行动,但这不一定非要关于奖励最大化本身,它可以是关于其他事情的。
One pretty concrete example might be behavior cloning in robotics. Say you're building a robot and your goal is to build a robot that goes around the world and does stuff, maybe makes my bed or cleans the kitchen. Once that robot is out there, it's an agent interacting in a world. The world has state, and the robot gets observations about the world. So it's operating on that POMDP loop once it's trained. But there's a question of what the training signal was. You could have trained it via reinforcement learning, where every time it took an action, it got a reward. Or you could train it via behavior cloning. Maybe I had a large supervised dataset where, given this observation, you should take this action. Then you could train a supervised learning model that would be very different and wouldn't have an explicit reward signal. You would use gradient descent and some supervised loss. So even though you end up with a system that operates in this POMDP-like loop, you could have had a training objective that was pure supervised learning and didn't require reinforcement learning. The summary is that there are two parts to POMDP: one captures the relationship between an agent and the world, the idea that the agent can observe things that reflect some abstract state but can never really know that abstract state, and it needs to operate on its observations. The reward and training are all about how it translates those observations to actions, but that doesn't necessarily need to be about reward maximization per se; it could be about other things.
完全正确。这与世界建模相关的原因在于,这正是这个术语的起源。世界模型的原始技术定义是,如果我们处于 POMDP 设置中,世界模型是输入世界状态、输入智能体采取的行动,然后预测下一个世界状态是什么。这就是世界模型这个术语最初使用的技术背景。
Exactly. And the reason this is connected to world modeling is because this is where the term originally comes from. The original technical definition of a world model is that if we're in this POMDP setting, a world model is something that inputs the world state, inputs the action the agent takes, and then predicts what the next world state is. That's the original technical setting where the term world model was used.
考虑到这一点,这在某种意义上是一个更广泛的历史背景,不一定关乎我们今天世界模型的发展方向。但我确实有一个关于状态的问题。很多时候我们听到状态被当作一种表示来讨论,几乎就像在这个 POMDP 形式体系中观察之于状态一样。通常状态是某个其他真实事物的低维表示。你认为这个区别对我们来说值得探讨吗?
With that in mind, this is in a sense a broader context of historical interest and not necessarily about where we're going with world models today. But I did have a question about state. A lot of times we hear state talked about as a representation, almost like the observation is to the state in this formalism of POMDP. Often state is some lower-dimensional representation of some other ground truth thing. Is that a distinction that is useful for us to explore, do you think?
也许有一点。我认为实际上人们有时会混淆两种不同的状态概念。所以尝试拆解可能有用。一种概念是,在我心中,当我说状态时,我想到的是世界的真实状态。那是世界中所有事物的完整描述,是回答任何关于它的问题所必需的。
Maybe a little bit. I think there are actually maybe two different notions of state that people talk about sometimes that are getting conflated. So that might be useful to try to unpack. One is the notion that, in my mind, when I say state, I'm thinking about the ground truth state of the world. That's the complete description of everything in the world that would be required to answer any question about it.
这是另一种切入方式,或者说是之前在我脑海中浮现的一个问题,我现在提出来。你知道,当你早些时候在循环的语境中谈论状态时,我在想状态与观察之间关系的一个极端例子:你坐在红绿灯前。观察可以被简化为一维的东西——红、绿、黄——或者二维的,随便什么。而状态,根据 POMDP 的定义,就像是世界上的每一个原子。我的问题是,当你谈论世界模型时,状态是否必然由世界上的每一个原子组成,还是说它只是一种表示?也许这是回答这个问题的另一种方式。
This is another way to come at it, or a question that popped up for me earlier that I'll surface now. You know, when you were talking about state earlier in the context of the loop, I was thinking of an extreme example of the relationship between state and observation: you're sitting at a traffic light. The observation can be boiled down to a one-dimensional thing—red, green, yellow—or two-dimensional, whatever. Whereas the state, by the definition of the POMDP, is like every atom in the world. My question is, when you talk about world models, does the state necessarily consist of every atom in the world, or is it just a representation? Maybe that's another way of coming at this question.
我想是的。我认为这完全取决于一个问题:什么样的抽象对手头的问题有意义,对吧?即使是物理学家也会根据你提出的问题使用不同的表示。对于某些问题,我想要这个量子系统的波函数——这是状态的最佳版本。对于其他系统,也许它是宏观的,量子效应不发挥作用,那么你可以把它看作一个牛顿系统,状态就是粒子的集合:它们的质量、位置和动量是什么。或者你谈论的是一个化学系统,它只是这些具有特定浓度的离子溶液。或者你谈论的是一个热力学系统,你把它看作理想气体,这就足以描述系统的状态了。
I think so. I think it's all about the question of what's the abstraction that makes sense for the problem at hand, right? Even physicists will use different representations depending on the question you're asking. For some problems, I want the wave function of this quantum system—that's the best version of a state. For other systems, maybe it's macroscopic and quantum effects don't come into play, so you could think of it as a Newtonian system, and the state is now this collection of particles: what are their masses, positions, and momentum. Or maybe you're talking about a chemical system, and it's just these solutions with these ions at this concentration. Or maybe you're talking about a thermodynamic system, and you think of it as an idealized gas, and that's enough to describe the state of the system.
但从来不是世界上的每一个原子。
But it's never every atom in the world.
不,不。我认为它总是相对于抽象而言的——解决你关心的问题所必需的抽象。但我认为还有另一个概念,即学习到的状态。这是人们现在正在讨论的另一件有趣的事情:真实状态——我们可能无法直接接触到物理世界中的真实状态。即使我们考虑每一个原子或理想气体,在很多情况下你都无法在现实中观察到它。所以相反,有另一个问题:我能否有一个我学习的模型,它学会接收观察并预测某种神经向量?这个学习到的向量,即使是由模型学习到的,其行为也像真实状态一样,我可以预测动作、预测该状态产生的观察,并想象该状态如何响应动作而转换。但它是来自神经网络的内部向量表示,不是物理学的真实状态,尽管它的行为有点像那个真实状态。
No, no. I think it's always relative to the abstraction—the abstraction that's necessary for solving the problems you care about. But I think there's another notion, which is a learned state. That's another interesting thing people are now talking about: the ground truth state—we're probably not going to have direct access to it in the physical world. Even if we think about every atom or an idealized gas, in a lot of situations you're just not going to observe that in the world. So instead, there's another question: could I have a model that I learn, which learns to take observations and predict some kind of neural vector? That learned vector, even though it's learned by the model, behaves as if it were a ground truth state, where I can predict actions, predict observations that would result from that state, and imagine how that state would transition in response to actions. But it's a learned internal vector representation from a neural network, not a ground truth state from physics, though it kind of behaves like that ground truth state.
这是否类似于基于模型和无模型的强化学习的概念?
Is that analogous to the idea of model-based and model-free RL?
有一点,是的。如果你在做基于模型的强化学习,那么你有一个显式的状态模型,这就是你拥有世界模型的地方。在基于模型的方法中,我有一个组件,它试图根据当前状态和动作预测下一个状态。在无模型的方法中,我没有那个显式概念——也许我只是获取观察,直接输入模型,它输出动作,没有显式的状态建模。但回到我们之前讨论的隐式世界模型,也许那个神经网络,如果足够大和复杂,会在其权重内部进行某种状态建模,如果它能在正确的情况下预测出非常好的动作。
A little bit, yeah. If you're doing model-based RL, then you have some explicit model of the state, and that's where you have a world model. In a model-based approach, I have a component that tries to predict the next state based on the current state and the action. In a model-free approach, I don't have that explicit notion—maybe I'm just getting the observations, feeding them directly to the model, and it outputs actions, with no explicit modeling of state. But going back to implicit world models we talked about before, maybe that neural network, if it's big and complicated enough, is doing something like state modeling inside its weights, if it's able to predict really good actions in the right circumstances.
我认为我们在这里做的是说明“状态”一词的模糊性。它可以是世界的表示或抽象,或者在某些情况下,它可以是智能体在选择动作时创建并预测的替代物。
I think what we're doing here is illustrating the ambiguity of the term state. It can either be whatever the representation or abstraction of the world is, or in some settings, it could be a surrogate for that that the agent creates and predicts against in choosing actions.
我认为哪种方法才是正确的方法,目前还没有定论。但这很令人兴奋,对吧?这就是为什么我认为现在很多人被这个领域吸引的原因之一。语言建模感觉我们有一套行之有效的最佳实践。有东西可以探索,但我们有一个行之有效的配方。但如果你想讨论如何建模环境、如何建模世界、如何建模与世界互动的智能体,有很多基本的技术问题或基本的工程方法,你可以这样做或那样做,我们真的不知道什么是最好的方式。这是一个非常令人兴奋的工作和研究领域。
And I think the jury is still out on which of these is going to be the right approach. But that's exciting, right? That's one of the reasons why I think a lot of people are attracted to this area right now. Language modeling feels like we kind of have a set of best practices that work pretty well. There are things to be explored, but there's a recipe that works pretty well. But if you want to talk about how do we model environments, how do we model worlds, how do we model agents that interact with worlds, there are a lot of basic technical questions or basic engineering approaches where you could do it this way or that way, and we don't really know what's going to be the best way. And that's a really exciting place to be working and doing research.
那么让我们稍微换个话题,谈谈图形方面的事情。正如你提到的,当我们谈论世界模型时,很多对话都集中在图形世界的生成上。该领域的一项基础技术是高斯溅射的概念。那么请谈谈人们处理世界生成的广泛方式,以及高斯溅射的作用,以及它在过去几年中是如何演变的。
So let's switch gears a little bit and talk a little bit about the graphical side of things. As you mentioned, when we talk about world models, a lot of that conversation is focused on the generation of graphical worlds. One of the foundational technologies in that space is the idea of a Gaussian splat. So talk a little bit about broadly the way folks approach world generation and the role of Gaussian splats, and how that has evolved over the past few years.
我认为实际上有一个顶层的分叉,我们应该在深入之前先指出来。那就是:你是做显式 3D 还是隐式 3D?高斯溅射是显式 3D 方法的一个例子,当我生成或建模我的世界时,我会有一个显式的 3D 表示,然后我会用这个显式 3D 表示做事情。其中一件事可能是将其渲染为你看到的像素。有一种不同的方法,是隐式的:我会有一个直接生成像素的模型——比如视频模型。它生成像素,直接生成观察,从不经过这个中间的显式 3D 表示瓶颈。我认为这两种方法在过去几年中都发展了很多。所以即使在那个点上,这也是一个有趣的分叉路口。
I think there's actually a top-level fork there that we should call out even before we get down that road. That's: are you doing explicit 3D or implicit 3D? Gaussian splats are an example of an explicit 3D approach, where when I generate or model my world, I'm going to have some explicit 3D representation of that world, and then I'm going to do things with that explicit 3D representation. One of those things might be rendering it down to pixels that you see. There's a different approach which is sort of implicit: I'm going to have a model that directly generates pixels—like a video model. It generates pixels, generates observations directly, and never goes through this intermediate explicit 3D representational bottleneck. I think both angles have developed a lot over the last couple of years. So that's an interesting fork in the road even at that point.
确实如此。
It is.
那么,嗯,World Labs 比如,采用的是——嗯,说你们采用高斯溅射方法是否公平,比如在 Marble 和你们发布并展示的一些工作中?你们用了那种方法,但你觉得那是基础性的,还是只是目前展示出来的,而你们对其他想法持开放态度?
And do you, um, so World Labs, for example, takes the—well, is it even fair to say that you take the Gaussian splat approach, like in Marble and some of the work that you publish and demonstrate? You use that approach, but do you feel that that is foundational, or is it just kind of what you have shown thus far but you're open to other ideas?
嗯,我觉得我们其实已经两种都做了,对吧?确实如此。我们推出了一个叫 Marble 的产品,它采用高斯溅射方法。Marble 的作用是让用户输入一张图片、一组图片序列或一段文本提示,然后生成一个世界,这个世界用一组高斯溅射来表示。这是一种显式的 3D 表示。一旦有了显式的高斯溅射表示,就可以渲染出世界的漂亮图像,或者你可以想象把物体合成进去,或者导出到图形引擎之类的。但去年年底我们发表了一篇研究博客文章叫 RTFM,它采用了一种非常不同的方法,叫做实时帧模型。这与我们的 Marble 方法形成对比,因为它更偏向直接生成视频、纯像素的方向。没有显式的 3D 世界表示。你有一个实时运行的模型,根据用户输入实时生成图像。所以用户要求移动,用户看到世界移动的图像。比如你说“向左移动”,你会看到自己向左移动的视频,这一切都是实时发生的。但在底层,没有显式的 3D 表示。只是模型实时生成的帧。所以 World Labs 实际上一直在朝两个方向探索。我们有 Marble,这是我们基于显式高斯溅射的世界模型,还有 RTFM,这是一个更隐式的实时帧生成世界模型。
Well, I think we've actually done both already, right? So it's true. We've put out a product called Marble, which takes the Gaussian splatting approach. And there, like, what Marble does is it lets a user input an image or a sequence of images or a text prompt, and then generates a world where that world is represented as a set of Gaussian splats. And which is an explicit 3D representation. And then once you have that explicit Gaussian splat representation, then you can render it to view nice images of the world, or you could imagine compositing objects into it, or export it to a graphics engine, or something like this. But then we had a research blog post late last year called RTFM, which takes a very different approach called the real-time frame model. And this contrasts with our Marble approach in that it kind of goes more this straight-to-video, pixels-only direction. There is no explicit representation of the 3D world. You have a model running in real time that generates images in real time in response to user inputs. So the user asks to move around, and the user sees an image of the world moving. Like you say "move left," you see a video of yourself moving left, and this all happens in real time. But under the hood, there was no explicit 3D representation. It's just frames being generated from a model in real time. So World Labs actually has been approaching both directions. We have Marble, which is our explicit Gaussian splat-based world model, and RTFM, which is a more implicit real-time frame-generating world model.
你知道,这让我想到一个问题,回到我们关于什么是世界模型的基础对话。我在想 Meta 最初的高斯溅射工作和演示,以及那些我认为在它们成为世界模型理念基础之前普及了高斯溅射的工具。你知道,它们让你拍一堆照片,然后把那些照片拼接起来——用“拼接”这个词很宽泛——本质上创建一个可导航的 3D 模型。甚至在那之前,我不认为它用了高斯溅射,但像苹果的 ARKit 和那类东西产生了这些适度可导航的 3D……我尽量不叫它们世界,因为问题是,你怎么在那东西和世界之间划清界限?我想说静态对动态,但那似乎不对。
You know, and this raises a question for me that goes back to kind of our very foundational conversation about what is a world model. I'm thinking about like Meta's initial Gaussian splat work and demos, and the tools that I think have popularized Gaussian splats before they became kind of foundational to this world model idea. You know, they allowed you to take a bunch of pictures and they would like stitch those pictures together—using the term "stitch" very loosely—to create essentially a navigable 3D model. And even prior to that, I don't think it used Gaussian splats, but like Apple's ARKit and those kinds of things produce like these modestly navigable 3D... I'm trying not to call them worlds because the question is like, how do you draw the line between that thing and a world? I want to say like static versus dynamic, but that doesn't seem right.
不,我明白你的意思。这里有一个非常有趣的区别,我觉得很多人对高斯溅射特别困惑,对吧?因为高斯溅射实际上只是一种表示。就像高斯溅射就是——我基本上有一个 3D 点云。我在空间里有一堆点,每个点都有一个位置、一个颜色,可能还有一些其他属性。
No, I know what you mean. There is a really interesting distinction here that I think a lot of people get confused about with Gaussian splatting in particular, right? Because a Gaussian splat is actually just a representation. Like a Gaussian splat is just—I have basically a 3D point cloud. I have a bunch of points in space, and each of those points has a position and a color and maybe a couple other properties attached to it.
如果你的点云足够大,你可以从某个点的视角在其中移动。
And if your point cloud is large enough, you can move around in it from the perspective of a point.
但有趣的问题是,那个高斯溅射点云是从哪里来的?人们会混淆两个大的分界,因为高斯溅射作为一种技术起源于重建。所以那里的想法是,我要对一个空间拍摄很多视角——比如数百甚至数千个视角,以非凡的细节覆盖这个单一空间。然后我要拟合一个高斯,就像一个 3D 点云高斯溅射表示,来解释我看到的图像。
But the interesting question is like, where did that Gaussian splat point cloud come from? And there are two big divides that people get confused about, because where Gaussian splatting originated as a technology is around reconstruction. So there the idea is like, I'm going to take a lot of views of a space—like hundreds or even thousands of views that cover this one space in extraordinary detail. Then I'm going to fit a Gaussian, like a 3D point cloud Gaussian splat representation, that explains the images that I saw.
嗯,这样一说,区别就很明显了。不只是世界,而是世界模型。
Well, put like that, the distinction is obvious. It's not just world, it's world model.
完全正确。这里面没有世界模型,对吧?就像这里没有一个在大量数据上学习过的强大模型。这些基于优化的重建方法来做高斯溅射,就像,我有一千张图像。整个宇宙就是这一千张图像,我拟合一个高斯溅射来匹配这些图像。这里没有可泛化的知识。这实际上与我们在 Marble 中所做的非常不同。在 Marble 中,我们训练了一个大型强大模型,它在各种数据上进行了大量训练。那个模型恰好输出高斯溅射。所以 Marble 世界模型是知道如何建模世界的模型。它输入图像,输入文本,输出高斯溅射。这与那种我只是愚蠢地把点云拟合到这一千张图像、没有模型学习大量数据的概念的情况非常不同。
Exactly. There's no world model in there, right? Like there was no big powerful model in here that learned on a ton of data. These kind of optimization-based reconstruction approaches to Gaussian splatting are like, I've got a thousand images. The whole universe is like these thousand images, and I'm fitting a Gaussian splat to match these images. There's no generalizable knowledge here. And that's actually very different from what we're doing in Marble. In Marble, we have trained a large powerful model that's been trained on a lot of data of various kinds. And that model happens to output Gaussian splats. So the Marble world model is the model that knows how to model worlds. It inputs images, it inputs text, and it outputs Gaussian splats. And that's very different from this situation where I'm just going to kind of dumbly fit a point cloud to these thousand images, and there's no notion of a model that learned a ton of data.
我们在开始录制之前提到过这一点,但我一直认为高斯溅射是一种宽泛的想法或过程——你知道,创建这些渲染点云的技术数学过程——但听起来它已经发展了很多,现在有标准文件格式和其他东西。你知道,谈谈那个生态系统和围绕高斯溅射使用的工具。
And we touched on this before we started rolling, but I always thought of Gaussian splat as kind of the broad idea or process—you know, the technical mathematical process of creating these rendered point clouds—but it sounds like that has evolved quite a bit, and now there are like standard file formats and other things. You know, talk a little bit about that ecosystem and the tooling that is used around Gaussian splats.
是的。所以高斯溅射大多数时候是一个相当特定的东西。它是一组点。每个点在 3D 空间中有一个位置,即 XYZ 三个坐标。它有一个不透明度,是一个介于 0 和 1 之间的数字,告诉你它有多不透明或多透明。然后通常有一个颜色,像 RGB 值,又是三个数字。然后你经常会有一些球谐函数的概念。它告诉你从不同位置看它是什么颜色,对吧?因为你想建模这个概念,也许这个点从底部看是一种颜色,从顶部看是另一种颜色。这有助于建模反射。因为如果我有一个闪亮的表面,比如镜子或光泽表面,那么如果我从一个角度看,我会看到来自顶部反弹的光的高光。如果我换一个角度看,我会看到不同的颜色。这些球谐函数是捕捉这种视角依赖颜色概念的一种具体方式。
Yeah. So Gaussian splats are a pretty particular thing most of the time. It's a collection of points. Each point has a position in 3D space, which is three coordinates XYZ. It has an opacity, which is a number between 0 and 1 that tells you how opaque it is or how transparent it is. Then you've got usually a color, which is like an RGB value, which is again three numbers. And then you'll often have some notion of spherical harmonics. And that tells you what color it is when you look at it from different positions, right? Because you want to model this notion that maybe this point is one color if I look at it from the bottom and a different color if I look at it from the top. And that helps you model reflections. Because then if I have a shiny surface like a mirror or a glossy surface, then if I look at it from one angle, I kind of see a highlight from a light that bounces from the top. If I look at it from a different angle, I see a different color. And these spherical harmonics are a particular concrete way to capture this notion of view-dependent color.
所以,高斯溅射就是一组点。每个点都有 XYZ 位置、不透明度、颜色,还有球谐函数,用来表示从不同角度观察时颜色如何变化。然后有多种方式将这些数据打包成文件中的字节。PLY 是一种非常流行的高斯溅射文件格式,读写相对容易,但效率较低。还有压缩表示,比如 SPZ,它对部分数据进行压缩,并对某些部分使用较低精度,这样存储的东西看起来非常相似,效果不错,但文件体积小得多。你可以把 PLY 想象成类似 GIF 或 PNG,是未压缩的格式。而 SPZ 更像 JPEG,会压缩部分数据并丢弃一些部分,但我们认为这些部分是人眼不易察觉的。
So then, you know, a Gaussian splat is this set of points. Each one has an XYZ position, an opacity, a color, and also these spherical harmonics that tell you how the color varies as you look at it from different angles. And then there are different ways to pack that data into bytes in a file. So PLY is a pretty popular file format for Gaussian splats, which is relatively easy to read and write but pretty inefficient. And then there are compressed representations like SPZ that compress some of that data and use lower precision for some parts of it, letting you store something that looks pretty similar, looks pretty good, but takes a much smaller file size. You should think of PLY as kind of like a GIF or a PNG, something that's pretty uncompressed. And an SPZ is a bit more like a JPEG, where you compress some parts of the data and throw away some parts, but we think those are parts that you won't notice as a human.
相对于我们思考传统 2D 和 3D 图像的方式,高斯溅射并不基于网格。那么,相对于我们所表示空间的维度,它们通常是稀疏的吗?
Relative to the way we think about traditional 2D and 3D images, Gaussian splats are not based on a grid. And are they sparse in general relative to the dimensionality of the space we're representing?
是的,它们不基于网格。这些点可以位于空间中的任何位置。但通常——哦对了,我忘了另一个重要的事情,真傻,没把笔记放在面前,但高斯溅射也有大小,对吧?它不是一个无穷小的点。它有一个半径。还有一个协方差矩阵,因为它可能不是球体,而是椭球体。协方差矩阵和半径——实际上就是协方差矩阵——告诉你它有多大,以及在不同维度上如何拉伸或压缩。所以高斯溅射最初的卖点之一是它们可能相当稀疏,用非常大的溅射来覆盖墙壁的大片区域。但在实践中,这通常会导致质量很低。所以通常你希望溅射在你想要覆盖的几何体上相当密集,这通常会生成更漂亮的图像。
Yeah, I mean, they're not based on a grid. Any of these points can live anywhere in space. But they typically—oh right, the other important thing I forgot, silly me for not having my notes in front of me, but a Gaussian splat also has a size, right? It's not an infinitesimal point. It has a radius. And also a covariance matrix, because it might not be a sphere; it might be an ellipsoid. The covariance matrix and the radius—well, really just the covariance matrix—tells you how big it is and how stretched or squashed it is along different dimensions. So part of the original pitch of Gaussian splats is that maybe they could be pretty sparse, with really big splats to cover a big part of the wall. But in practice, that usually ends up in pretty low quality. So usually you want the splat to be fairly dense over the geometry you want to cover, and that ends up generating nicer images in general.
既然如此,高斯溅射之所以有效,是否部分在于渲染器能够聚焦于用户视角内的内容,而忽略无关的事物?或者更广泛的问题是:高斯溅射为何如此有趣,相对于我们之前处理这类问题的方式,能否快速回顾一下?
That being the case, is part of the technology or what makes Gaussian splats work that the renderer is able to focus on what's in the user's viewpoint versus things that are extraneous to it? Or maybe the broader question is: what makes Gaussian splats so interesting, a quick refresher on that relative to the way we've approached this before?
我认为你应该对比的是三角形网格。三角形网格是计算机图形学中的标准表示,我们将整个世界表示为小三角形。一切都是由小三角形构成的,几乎你玩过的任何游戏、看过的任何特效镜头、见过的任何计算机生成图像,都是将世界建模为大量小三角形。这在计算机图形学中已经出色地工作了几十年。但问题是,三角形与神经网络不太契合,因为关键点在于可微性。你需要能够通过这种表示进行微分并传递梯度。特别是,这意味着你希望你的表示具有这样的性质:如果我稍微改变输入,输出也会稍微改变。三角形并非如此,因为如果我这里有一个三角形,我稍微移动它,突然原本不可见的东西变得可见了。因此,图像作为参数的函数会出现急剧的、不连续的变化。高斯溅射没有这个性质,因为一切都是平滑的、部分透明的。当你无限小地改变高斯溅射的任何参数时,你看到的图像会连续变化。所以它们与神经网络结合得非常好。神经网络是基于梯度的学习者:我有一个目标函数,通过梯度下降来最小化它。为此,我需要通过我使用的任何表示来传递梯度信号。高斯溅射传递梯度非常好,所以它们可以插入基于梯度的优化器中,要么直接针对一组图像进行优化,就像重建案例那样,要么插入神经网络的输出,这更像是我们在 Marble 案例中所做的。在那里的设置是:我有一个神经网络,它输出高斯分布,这些高斯分布连接到某个损失函数,然后我可以将损失一直反向传播到神经网络的参数中。这是最大的差异,也是高斯溅射的重大创新,也是过去几年人们为之兴奋的原因:它是一种与神经网络非常干净地集成的图形表示。
I think the contrast with Gaussian splats you should be thinking about is triangle meshes. Triangle meshes are the standard representation in computer graphics, where we represent the whole world as little triangles. Everything is made up of little triangles, and pretty much any game you've ever played, any VFX shot you've ever seen, any computer-generated image you've ever seen, models the world as lots of little triangles. That has worked amazingly for computer graphics for decades. But the problem is that triangles don't fit with neural networks very well, because the important part is differentiability. You want to be able to differentiate through this representation and pass gradients. In particular, that means you want your representation to have the property that if I change the input a little bit, the output also changes a little bit. That's not the case with a triangle, because if I've got a triangle here and I move it a little bit, all of a sudden something that was invisible now becomes visible. So I have a sharp, discontinuous change in the image as a function of the parameters. Gaussian splats don't have that property, because everything is smooth and partially transparent. The image you see continuously changes as you vary any of the parameters of the Gaussian splats infinitesimally. So they integrate with neural networks really well. Neural networks are gradient-based learners: I have an objective function, and I minimize it via gradient descent. To do that, I need to pass gradient signal through whatever representation I'm using. Gaussian splats pass gradients really well, so they can be plugged into gradient-based optimizers, either directly optimizing against a set of images, as in the reconstruction case, or plugged into the output of a neural network, which is more what we do in the Marble case. There, the setup is: I've got a neural network, it spits out Gaussians, those Gaussians get attached to some loss function, and then I can backpropagate my loss all the way into the parameters of the neural network. That's the biggest delta, the big innovation of Gaussian splats, and why people got excited about them over the past few years: it's a graphics representation that integrates with neural networks really cleanly.
将其与世界模型的一些核心属性联系起来,即一致性。这种反向传播的能力,或者说我们如何获得这种一致性?它来自建模吗?它来自规模吗?它来自哪里——是架构的问题吗?
Kind of tying that to some of the core properties of world models, namely consistency. Is that ability to backpropagate, or how do we get that consistency? Is that coming from modeling? Is it coming from scale? Like, where does that come from—is it an architecture thing?
我认为它可以来自很多地方,这实际上是一个非常有趣的出发点,因为我们周围的世界具有一致性。如果我看你,你每时每刻看起来都非常相似,或者如果我走到另一个房间再回来,你仍然会在这里。这就是一致性的概念。而且有不同的——高斯溅射在构造上就具有一致性,因为我对世界有了这种显式的 3D 表示。所以如果我看着它,然后移开视线,再回头看,它都在 3D 中,所以看起来会一样。
I think it can come from many places, and that's actually a really interesting jumping-off point, because there's this property of the world around us that it's consistent. If I look at you, you look pretty similar from moment to moment, or if I walk to a different room and come back, you're still going to be here. That's this notion of consistency. And there are different—Gaussian splats are kind of consistent by construction, because I've got this explicit 3D representation of the world. So if I look at it, then look away, then look back, it's all there in 3D, so it's going to look the same.
但你也可以通过大规模数据、大规模训练和大规模算力来获得一致性。这更倾向于我们在 RTFM 和其他地方所做的隐式表示。所以那里的想法是,如果我有一个模型,没有高斯溅射,没有 3D,没有显式点,只有一个输出世界 RGB 像素值的模型,那会怎样?但如果那个模型非常非常聪明、非常强大,经过大量数据训练,也许还有合适的表达性架构,即使没有数学上的保证——没有任何数学上的东西保证它一致——它最终仍然会一致。而且我认为实际上并不是一个比另一个好;它们只是技术曲线上的不同点,对吧?如果你的算力预算相对较低,想让东西在嵌入式设备上运行,比如你不想训练巨型模型,那么高斯溅射就很有吸引力,因为它们构造上就一致。但如果你能扩展规模,使用大量数据和算力,那么我实际上认为隐式 3D 路线将会无限扩展。所以对我来说,这更像是一个工程问题,即我当前面临的问题的设计约束是什么,而不是哲学上的分歧。对吧?如果你想要便宜且构造上一致,高斯溅射非常有吸引力。如果你想要能无限扩展、能依赖无限数据和无限算力的东西,但愿意在推理时支付那无限的算力以及巨大的训练成本,那么我认为隐式纯像素方法非常有吸引力。这正是我们在 World Labs 两者都做的原因。
But you could also get consistency via large scale data and large scale training and large scale compute. And that kind of leans into the more implicit representations that we've done in RTFM and in other places. So there the idea is, what if I'm going to have a model where there is no Gaussian splats, there is no 3D, there are no explicit points? I just have a model that's spitting out RGB pixel values of the world. But if that model is really, really smart and really powerful, having been trained on a lot of data and maybe with the right expressive architecture, even though it's not mathematically guaranteed—there's nothing mathematically guaranteeing it to be consistent—it still ends up being consistent. And I think it's actually not that one is better than the other; they're just different points on the technology curve, right? If you have a relatively low compute budget and you want things to run embedded, like you don't want to train giant models, then Gaussian splats are appealing because they're consistent by construction. But if you can scale up and use a lot of data and a lot of compute, then I actually think the implicit 3D route is going to be the thing that scales up to infinity. So it's more of an engineering question of what are the design constraints of the problem facing me right now, and less a philosophical divide for me. Right? If you want cheap and consistent by construction, Gaussian splats are very appealing. If you want something that scales to infinity and can rely on infinite data, infinite compute, but you're willing to pay that infinite compute at inference time and the big training costs, then I think the implicit pixels-only approach is very appealing. And that's exactly why we've done both at World Labs.
我想我问那个问题的方向更侧重于模型和生成视角,以及我们描述一个实时生成的世界,用户可以在这个世界中导航、转身、回来,我们在 RTFM 的情况下生成一致的帧,或者在 Marble 的情况下生成溅射。问题是,我认为你所说的部分意思是,一致性与我们讨论的是像素生成还是溅射生成有点正交。所以那个问题的下一部分是,一致性似乎是这一切中非常重要的一部分。它从何而来?
I think the direction I was trying to go with that question was focusing more on the model and the generation perspective, and the idea that we're describing a world generating on the fly, the user can navigate through this world, turn away, come back, and we're generating consistent well frames in the case of RTFM or splats in the case of Marble. And the question is, I think part of what you're saying is that consistency is kind of orthogonal to whether we're talking about pixel generation or splat generation. And so the next part of that question is, you know, that consistency seems like a really important part of all this. Where does it come from?
我的意思是,我认为它最终基本上来自你的数据,对吧?就像,无论你使用什么表示,无论是原始像素还是高斯溅射,归根结底,系统中必须有一个神经网络被训练来生成一致的数据。高斯溅射让神经网络更容易产生一致的输出,因为它们构造上更一致,但最终它必须来自以你希望模型学习的方式一致的世界视图数据。
I mean, I think it basically comes from your data ultimately, right? Like, no matter what representation you're using, if it's raw pixels or Gaussian splats, at the end of the day, you have to have a neural network in the system that's been taught to create consistent data. And Gaussian splats kind of make it easier for a neural network to produce consistent outputs because they're more consistent by construction, but ultimately it has to come from data of views of the world that are consistent in the way you want your model to learn.
那么我认为问题或机会在于深入探讨 Marble,谈谈配方,以及它如何创建世界模型,如何创建 Marble 世界模型。
I think then the question is, or the opportunity is, to dig into Marble and talk a little bit about the recipe and how that creates a world model, how that creates the Marble world models.
我们实际上还没有明确讨论过 Marble 的架构,但高层次上,它让你作为用户输入不同类型的东西。你可以输入文本提示、图像、多张图像或视频。然后从那里,我们生成一个 3D 高斯溅射世界。其中一个中间步骤是生成 360 度全景图像。这就是在给定这些输入的情况下,你有一个模型——模型的一部分——生成你即将生成的世界的 360 度全景视图。那是一个强大而庞大的生成模型,需要将任何用户输入首先映射到 360 度全景,然后从那里提升到完整的 3D 高斯溅射世界。
We actually haven't talked explicitly about the Marble architecture, but at a high level, it lets you, as a user, input different kinds of things. You can input a text prompt, you can input an image, you can input multiple images, you can input a video. Then from that, we generate a 3D Gaussian splat world. One intermediate step in that is generating a 360 panorama image. And this is where, given those inputs, you have one model—one part of a model—that generates a 360 panorama view of the world you're about to generate. And that's a big, powerful generative model that needs to take whatever that user input was and map it first into a 360 pano, and then from there lift it up into a full 3D Gaussian splat world.
那么这意味着用户的输入是单个点或单张图像吗?或者我想的是,如果用户提供一组空间上多样化的图像,是不是只是球体更大?
Is the implication then that the user's input is either a single point or single image? Or I guess I'm thinking about like if the user's providing a spatially diverse set of images, is it just that the sphere is much bigger?
哦不。所以基本上我们做的是,主要的用例几乎是单张图像,这可能是今天 Marble 中效果最好的。所以就像我输入一张单张图像,然后对于我在该图像中能看到的世界中的东西,我生成的世界应该与我在输入图像中看到的一致,然后模型将尝试完成它——为输入图像中不可见的世界中的其他一切提供合理的补全。对吧?所以如果我拍了一张教室前面黑板的照片,那么模型应该知道黑板后面会有所有这些椅子,人们坐的地方,学生坐的地方。然后模型应该能够拿一张黑板的照片,补全它,生成黑板后面的椅子,然后把所有东西提升到你可以导航的 3D 中。
Oh no. So basically what we're doing is, like the kind of marquee use case is almost single image, and that's probably what works best in Marble today. So there it's like I'm going to input a single image, and then for the stuff in the world that I can see in that image, my generated world should match what I see in my input image, and then the model will try to complete what it—try to give a plausible completion for everything else in the world that's not visible in the input image. Right? So if I'm maybe taking a picture of a blackboard in a class in the front of a classroom, then the model should know that behind the blackboard is going to be all these chairs where the people sit, where the students sit. And then the model should be able to take a picture of a blackboard and then complete that and generate the chairs behind the blackboard and then lift all that into 3D that you can navigate.
用户是否也提供文本提示?
And is the user also providing text prompts?
他们可以。是的,用户可以直接从文本提示。那是可选的。如果你想纯粹从文本生成你的世界,我们可以做到。如果你想提供文本作为辅助输入,为输入图像中发生的事情提供额外指导,那也可以。但我们想用 Marble 采取这种方法,为用户提供最大的灵活性,无论你得到什么信号,无论你做什么编辑,无论你生成后想把它带到哪里,我们都想给你很多使用它的途径,而不是试图引导每个人走一条僵化的生成路径或生成后使用路径。
They can. Yeah, the user can prompt this directly from text. That's kind of optional. If you want to generate your world purely from text, we can do that. If you want to provide text as an auxiliary input to give some extra guidance to what's happening in your input image, that can work too. But we kind of wanted to take this approach with Marble of maximum flexibility for users, that no matter what kind of signal you got, no matter what kind of edit you do, no matter where you want to take it after it's generated, we want to give you a lot of pathways to use this thing and not try to guide everyone down one rigid pathway for how to generate things or where to use it after you generate it.
我想我关于多图像的疑问或假设来自,比如我见过一些 Marble 世界,是那种经典的房间,你旋转房间,有大量细节,超级令人印象深刻。但我也见过其他一些,它们更像这些视频游戏沉浸式世界,你可以导航穿过,你知道,一个幻想村庄之类的情况。
I think where my questions around or my assumption of multi-image was coming from, like I've seen some Marble worlds that were kind of your classic room, and you spin around the room and there's like a ton of detail, super impressive. But then I've seen other ones where they're more like these video game immersive worlds where you can navigate through, you know, a fantasy kind of village kind of situation.
而且我觉得这些都是通常从单张图像生成的,可能还加上一些文本条件?
And I think are those all kind of generated from a single image generally and potentially some text conditioning?
是的,我的意思是,这就是在循环中拥有世界模型的美妙之处:你有一个强大的大型生成模型,它已经在大量数据上训练过,包括真实世界的数据、奇幻的数据、逼真的和非逼真的内容。所以你在后端有这样一个强大的模型,无论你带来什么样的图像或提示,无论是你房间的图像,还是像电子游戏那样的奇幻环境,你都有一个模型,它了解所有这些不同类型的世界,并能根据手头任务的需要生成它们。这就是这些由大型世界模型支持的生成与经典的高斯泼溅(gaussian splatting)之间的主要区别,后者只是拟合我恰好拥有的那几千张图像。
Yeah, I mean that's the beauty of having a world model in the loop there is that you have a single powerful large generative model that's been trained on a ton of data both real world data and fantastical data and photorealistic stuff and non photorealistic stuff. So you've got this big powerful model in the back end that no matter what kind of image or prompt you bring whether that image is of you know a room in your house or you know a fantastical like video game kind of environment um you've got a model that knows all of these different kinds of worlds and can generate them as it see as as is necessary for the task at hand. Um and that's that's that's and that's kind of the major difference between you know having these generations backed by a world by a big world model versus you know classical gausian splatting just fitting to these thousand images that I happen to have.
关于用于创建这些模型的数据集,以及训练方法和配方,你能说些什么?
And what can you say about the the data sets that are used to create these models and like the training approach and recipe?
是的,我的意思是,你需要在大量数据上训练。而且,非常重要的一点是能够训练于各种不同类型的数据,对吧?因为最终世界本身是 3D 的,并且有很多 3D 结构。但是,并没有很多显式的 3D 数据供你学习。那是一种相当罕见的数据形式。但是外面有很多图像,有很多视频。图像和视频都是 3D 世界的 2D 投影。所以即使你最终想要一个产生 3D 的模型,让它学习大量的图像和视频数据仍然非常强大,因为你可以获得非常大量的这些数据。然后当你有了显式的 3D 数据,那学习起来非常强大,但你不想只局限于从 3D 数据学习。
Yeah, I mean you need to train on a lot of data. Um, and one thing that's really important is being able to train on a variety of different kinds of data, right? Because um, ultimately the world itself is is 3D and like has a lot of, you know, 3D structure. Um, but there's not a lot of, you know, explicit 3D data for you to learn on. That's a pretty rare form of data. But there's a lot of images out there. There's a lot of videos out there. and images and videos are both, you know, 2D projections of a 3D world. So even if you want a model that at the end of the day is going to produce 3D, it's still very powerful for it to learn on large quantities of image and video data because those you can get in a like those you can get in in very large quantities. And then when you've got three explicit 3D data, that's very powerful to learn from, but you don't want to be bottlenecked only on learning from 3D data.
我假设你是想在你原生拥有的任何数据上训练,而不是像拿一张图像然后投影到 3D 那样,听起来超级昂贵且嘈杂。
I'm assuming that you're trying to to train on whatever data you have natively as opposed to like you take an image and project it into 3D or something that sounds super expensive and noisy.
是的。我的意思是,过去十年我们从深度学习中得到的教训之一是,你想要在大量数据上训练的端到端的大型模型。所以,如果你有一个任务,比如你的任务是什么?你的任务是今天生成高斯泼溅世界,还是明天生成真正强大的 3D 一致帧?就像想想你想解决的任务是什么,以及我如何调动非常大量的数据,让模型以非常通用的方式学习如何解决该任务。
Yeah. I mean, one of the lessons we learned from deep learning over the past decade is that you want you want big models trained on a lot of data that are trained end to end. So, if you've got a t like what is your task? Is your task to like generate gausian splat worlds today or is it to generate really powerful 3D consistent frames the next day? like just think about what is the task that you want to solve and how can I marshall very large quantities of data um to let a model learn how to solve that task in a very general way.
模型的输出是直接是泼溅(splats),还是有一些——我想这就是你刚才说的——它是否产生某种归一化的 3D 表示,无论那是什么,你可以将其转换为像素或泼溅,或者你只是转换,它只是吐出泼溅?
Is the model's output directly splats or is there some I I think this is kind of what you were just saying like is it is it producing some kind of normalized 3D representation whatever that is and you can convert that to pixels or splats or you just convert you're it's just spitting out splats.
是的,这个模型最原始的输出在某种程度上是泼溅(splats)。一旦你有了泼溅,那就是我们最低公分母的 3D 格式。所以一旦你有了泼溅,你可以将它们渲染成图像,这可以给你来自这些世界的图像或视频。我们也可以将网格表示拟合到世界上。然后在某些情况下,比如你想将其导入游戏引擎或 VFX 引擎,有时这些引擎目前对泼溅的支持不太好。而且,拥有一个更经典的 3D 三角形网格世界是有用的。所以泼溅是我们从世界模型得到的最原始的输出格式,然后你可以从那里转到图像、视频或网格。但同样,这与 RTFM 模型形成鲜明对比,那里没有泼溅,它直接输出像素。
Yeah, the kind of rawest output from this model is is is splats in in in some way. And then once you've got splats, you can that's kind of our our lowest common denominator 3D format. So once you've got splats, you can render them to an image and that can give you images or videos from these from these worlds. Um we can also uh you can also fit a mesh representation to to the world. Um and then in in some contexts like you want to import this into into a game engine or or VFX engine, sometimes those those don't work so well with splats today. Um, and it's useful to have a three a more classical 3D triangle mesh of the world. So the splat is kind of our our our most rawest output format from the marble models and then you can go from that to images or videos or or meshes. But that but again that that that's in pretty stark contrast to the RTFM model where there's no splats. It just directly spits out pixels.
我们之前提到过你最近写的或你的团队最近发表的一篇博客文章。我觉得那篇文章特别有趣的是,你知道,作为一个分析师,每当我看到一个试图剖析某个领域区别并讨论这些区别的分类法时,我都会感兴趣。而这就是你试图做的,至少在世界模型的某个特定维度上。我想到目前为止,我们在这次对话中已经介绍了三四个其他分类法。但你知道,稍微谈谈你如何划分世界模型的这个特定维度。
We alluded earlier to uh the a blog post that uh you recently wrote or your team recently published. And I thought what was really interesting about that was like you know as kind of a an analyst uh anytime I see a taxonomy that tries to kind of pick apart the distinctions in a space and and talk about those I'm interested. uh and that is uh what you try to do at least in you know a particular dimension of world models. I think we've introduced like three or four other taxonomies in this conversation so far. Um but you know talk a little bit about the the way you kind of divided up the you know that particular dimension of of world models.
所以,就像我说的,我们谈过几次,我认为人们训练并称之为世界模型的东西有很多不同的风格,从外部看它们看起来非常不同。当我们认真思考这个问题时,意识到有一个框架,它们实际上都是对同一事物的不同视角。我们意识到我们可以将其追溯到之前讨论过的 POMDP 形式化。因为在 POMDP 中,记住有三样东西在系统中移动:你在世界中的智能体,然后智能体产生动作,世界转换状态,然后智能体接收观察。我们意识到,如果你这样想,今天人们训练的大多数被称为世界模型的东西,通常都在那个循环中输出三样东西之一。要么你构建一个输出动作的模型,要么你构建一个输出状态的模型,要么你构建一个输出观察的模型。一旦我们意识到这一点,就有了一个顿悟时刻:哦,并不是这些人都在构建完全不同的东西,他们只是专注于这个基本 POMDP 循环的不同部分。它们都连接在一起,试图建模世界并理解世界如何响应、演变和随时间变化,但不同的人为了不同的应用而专注于循环的不同部分。
So there, like I said, um, like we talked about a couple times, I think there's a lot of different flavors of things that people are training and calling world models that look pretty different from the outside. Um, and when we thought about this really hard and realized that there is a framing where they all actually are kind of like different different views onto the same thing. Um, and there we we realized we could ground this back in this POMDP formalism that we talked about before. Um, because what in this POMDP remember there are three things that are moving around the system. um you got the agent in the world, then you've got the agent um producing actions, you've got the world transitioning states, and then you've got the agent receiving observations. Um and we and we realized if you think about it that way, um pretty much everything that people is are training today that are called world models are usually outputting one of three things in that in that loop. Either you're building a model that outputs um that outputs actions, you're building a model that outputs states, or you're building a model that outputs observations. Um, and once we realized that, it was kind of a an aha moment that, oh, it's not that these people are all building totally different things. They're just focusing on different parts of this fundamental POMDP loop. And they're all connected together trying to model the world and understand how worlds can respond and evolve and change over time, but different people are focusing on different parts of that loop for different applications.
明白了。所以像 Genie 这样的东西产生像素流,那在这个模型中就是观察,而像机器人系统那样为机器人在世界中采取行动产生动作,那更侧重于作为输出。
Got it. So something like a genie that's producing a stream of like pixels that would be observations in this model and something like a robotic system that's producing actions for the robot to take in the world that's more uh focused on that as an output.
完全正确。然后我们根据它们输出的内容,将它们分类为这三种不同的世界模型类别。就像你说的,如果你输出观察,比如 Genie 或 RTFM,那么我们称之为渲染世界模型或简称渲染器,因为它产生最终观察,可以被人类或智能体消费。
Exactly. So then you know then we kind of like texonomize these into like these three different categories of world models then based on what they're outputting. So like you said like if you're outputting the observation like Genie or like RTFM then we're calling that a rendering world model or just a renderer right because it's producing a final observation that can be consumed by a by a person or maybe by an agent.
所以这些就是我们所说的“作为世界模型的渲染器”。另一个酷炫的是所有机器人领域的人都在训练机器人策略。人们训练模型来操作机器人身体,然后机器人身体在现实世界中做出酷炫的事情。但那个模型在做什么?那个模型处于 POMDP 循环的另一端。它从现实世界接收观测,然后需要采取行动来改变世界。他们有一个世界模型,输入现实世界的观测,输出要在现实世界中采取的行动。这就是我们所说的“规划器世界模型”,因为它是在规划一系列要在世界中采取的行动。这几乎与“作为渲染器的世界模型”完全对偶,因为渲染器从人类用户那里接收行动,然后输出在这些行动序列下世界可能呈现的观测。所以渲染器和规划器几乎是完美的对偶。
So those are kind of what we're terming a renderer as a world model. Then the other cool one is all these robot people training robotics policies. People are training models that operate a robot body, and then the robot body does something cool in the world. But what is that model doing? That model is on the other side of the POMDP loop. It's receiving observations from the real world, and now the model needs to take actions to try to make a change in the world. They've got a world model that is inputting real-world observations and outputting actions to be made in the real world. That's what we're calling a planner world model, because it's planning out a sequence of actions to take in the world. That's almost exactly dual to the world model as renderer, because the renderer is receiving actions from a human user and then outputting observations of what a world might look like under that sequence of actions. So the renderer and the planner are almost perfect duals to each other.
第三个是关于状态。这是一个棘手的问题,我们一直在反复思考。我们称之为“作为模拟器的世界模型”,因为世界模型的一个重要特性是它们可能应该对状态进行某种模拟。在某些情况下,你只关心行动和观测,但在其他情况下,你可能想了解所考虑世界的状态,并可能以某种显式或半显式的方式理解或模拟该状态如何响应行动而演变。这就是“作为模拟器的世界模型”。我们意识到,当你把这些术语分解开来——规划器、模拟器、渲染器——如今人们训练的所有世界模型几乎都可以归入这三类之一。
Then the third one is about the state. That's a tricky one that we keep coming back to. We're calling that world model as simulator, because an important property of world models is that maybe they should do some kind of simulation of the state. In some contexts you only care about the action and the observation, but in other contexts you might want to know something about the state of the world under consideration, and maybe understand or simulate how that state might evolve in response to actions in some explicit or semi-explicit way. So that's the world model as simulator. We realized that when you break it down to these terms—planner, simulator, renderer—pretty much all the world models that people are training these days can be bucketed into one of these three camps.
我觉得模拟器很有意思,因为我通常认为模拟是我们用来开发模型的外部工具,而不是这种把模型本身从根本上视为模拟器的框架。
I thought the simulator was interesting in that I typically think of simulation as an external tool that we're using to develop models, as opposed to this framing of the model itself being fundamentally a simulator.
在那里,我认为另一个有趣的事情是这些开始融合。我认为 Marble 实际上是一个例子,它有点跨越了“作为渲染器的世界模型”和“作为模拟器的世界模型”之间的边界。因为当你作为用户与 Marble 交互时,你看到的是屏幕上的像素。那些像素,从这个意义上说,你看到的是观测,但那个观测并非直接来自神经网络。那个观测来自一组高斯溅射。高斯溅射是我们拥有的世界的显式状态表示。一旦你有了那个显式状态表示,你就可以用它做渲染之外的其他事情,因为它是显式的 3D 表示。你可以测量两点之间的距离,或者你可以通过拉入另一个物体资产并将其放入那个世界来显式地操作状态。所以至少我们今天的 Marble 版本,作为用户的端到端体验是你看到观测和屏幕上的像素。所以它有点像渲染器,但模型本身输出的是这个显式或半显式的世界状态,你可以用它做其他事情。
And there, I think another interesting thing is these start to blend together. I think Marble is actually an example of something that is somewhat straddling the boundary between world model as renderer and world model as simulator. Because when you as a user interact with Marble, you're seeing pixels on the screen. Those pixels, from that sense, you're seeing an observation, but that observation did not come directly out of the neural network. That observation came out of a set of Gaussian splats. The Gaussian splats are this explicit state representation that we have of the world. Once you have that explicit state representation, you can do other things with it other than rendering, because it's an explicit 3D representation. You can measure the distance between two points, or you can manipulate the state explicitly by pulling in another object asset and putting it into that world. So at least the version of Marble we have today, the end-to-end experience as a user is you're seeing observations and pixels on the screen. So it's kind of a renderer, but the model itself is outputting this explicit or semi-explicit world state that you can then do other things with.
这感觉比渲染器和规划器的区分要细致得多。如果我想到一个渲染器,它不是输出 2D 帧,而是直接输出 3D,那么你可以测量点之间的距离,它是一个表示,但同时也是观测。
That feels like a much more nuanced distinction than renderer and planner. If I think about a renderer that, as opposed to spitting out 2D frames, was spitting out 3D directly, then you can measure distances between points, and it's a representation but it's also the observation.
完全正确。这就是我们想在这里提出的另一个观点:虽然存在这个分类法,但它并不是非常严格,我认为过于严格地坚持它会对我们所有人都不利。现实世界是混乱的,我们所有的分类法都会失效,但它是一个有用的思考框架。实际上,最终能够引领我们的世界模型将会融合所有这些方面。我们最终想要的是一个能够做所有这些事情的综合系统。
Exactly. That's the other point we wanted to make here: while there is this taxonomy, it's not very rigid, and I think sticking too rigidly to it would do a disservice to all of us. The real world is messy and all of our taxonomies break down, but it's a useful framework to think about. In reality, the world models that are going to carry us at the end of the day are actually going to blend all of these aspects together. We ultimately want to have one combined system that could do all of these things.
是的。但我想我还想请你详细讲讲模拟器,除了 Marble 的例子,还有其他例子能体现“模型作为模拟器”这个想法吗?
Yeah. But I think I'm also asking for more elaboration on simulator, and beyond the Marble example, are there other examples that kind of capture this idea of model as simulator?
是的,我认为有几个例子。实际上我认为目前还没有人真正掌握那个,这就是为什么它很有趣。但我可以想象未来系统有几种风格。一种是 Marble 的显式状态未来版本——不是说这实际上是我们将要做的,但人们可以这样做。也不是说我们没有在做。但你可以想象一个版本,把高斯溅射之类的东西当作世界状态,但实际有一个模型随时间演变那个状态。所以你可以有一个模型,输入图像,输出高斯溅射世界,然后用户采取一些行动,比如拿起瓶子或移动东西,然后你会有一个模型,用强大的神经网络模型去更新高斯溅射世界。那将是一个与显式或半显式世界状态协作,并能够随时间响应行动而演变那个世界状态的模型。我认为还没有人真正构建过这样的系统,但他们可以。
Yeah, I think there are a couple of examples. I actually don't think anyone's nailed that one right now, which is why it's interesting. But I think there are a couple of flavors of future systems I can imagine. One is maybe the Marble explicit state future version—not to say that this is actually what we're going to do, but one could do it. Also not to say that we're not doing it either. But you could imagine a version that treats something like Gaussian splats as a world state, but actually has a model evolve that state over time. So you could have a model that inputs an image, outputs a Gaussian splat world, and then a user takes some action like picking up a bottle or moving something around, and then you'll have a model that goes in and updates the Gaussian splat world with a powerful neural network model. That would be a model that is working with this explicit or semi-explicit world state and actually being able to evolve that world state over time in response to actions. I don't think anyone's really built a system like that, but they could.
另一种是隐式世界状态。也许我们可以构建不直接处理高斯溅射的系统,而是拥有某种世界状态的隐式向量表示。它的行为方式就像显式状态从中渲染观测一样。你可以有网络预测该状态如何响应行动而演变,预测该状态如何随时间演变,并拥有那个显式世界状态的神经网络模拟。我认为人们正在研究这个,但感觉关于如何让它工作的确切配方仍然是一个相当开放的研究问题。
The other would be the implicit world state. Maybe we can build systems that are not working on Gaussian splats directly, but have some kind of implicit vector representation of a world state. It behaves in the way that an explicit state would render observations from it. You can have networks that predict how that state evolves in response to actions, predict how that state is going to evolve in time, and sort of have a neural network analog of that explicit world state. I think people are working on that, but it feels like it's still a pretty open research question as to what's the exact right recipe to get that to work.
是的。
Yeah.
在思考模拟器时,另一件让我印象深刻的事情是,也许模拟器的一个完美体现就是我们讨论过的理论构建器。它把这些抽象概念提炼出来,在这种情况下,状态就是关于世界的一组理论。
The other thing that jumps out at me in thinking about the simulator is that maybe a perfect expression of a simulator is this theory builder that we talked about. It takes these abstract ideas and kind of boils them down into, in this case, the state is a set of theories about the world.
是的。我的意思是,然后你得谈论状态和元状态、理论和元理论,并把它提升到一个抽象层次,对吧?比如你可以说,理论构建器是那个在状态中书写物理定律的人。所以理论封装了可能存在的那类世界,以及那类世界如何被允许演化。而状态则是理论的一个特定实例,告诉你一个特定的世界。但这就像——我认为没有人知道该怎么做。
Yeah. I mean, then you got to talk about states and metastates and theories and meta-theories and push it up a level of abstraction, right? Like you could say maybe the theory builder is the one who's writing the laws of physics in the states. So the theory kind of encapsulates the kinds of worlds that may exist and how those kinds of worlds are allowed to evolve. And then the state is like one particular instantiation of the theory that tells you a particular world. But that's like—I don't think anyone has any clue how to do that.
是的。是的。好吧,我们有点太抽象了。你稍微谈到了统一世界模型的想法。我想我们稍微捕捉到了这一点,就是它是一个有泄漏的抽象,你期望在现实产品中出现交叉。
Yeah. Yeah. Okay. We're getting a bit too abstract there. You talk a little bit about this idea of a unified world model. I think we captured that a little bit, just the idea that it's kind of a leaky abstraction and you're expecting crossover in real-world products.
完全正确。我认为我们作为一个领域将要达到的是能够联合完成所有这些任务的模型。它们都将相互受益。为什么?因为它们都在问类似的问题:它们想理解可能存在哪些种类的世界?这些世界如何对行动做出反应?这些世界看起来像什么?从这些世界中会产生什么样的观察?它们如何随时间演化?你能用它们做什么?这些都是相互关联的基本问题,对吧?如果我在理解世界状态如何演化方面变得更好,可能我也会在作为渲染器时更好地预测它从不同角度看起来会是什么样子。而如果我很擅长渲染,比如想象隐式世界状态将如何演化,那可能对规划也很有用,对吧?如果我能隐式地演化我的世界状态,可能我也能针对该状态进行规划,并知道如何采取行动来影响世界。所以我认为这很——我不认为我们已经做到了,但我认为在未来几年内,我们将开始看到模型将这些能力越来越多地组合成一个强大的统一模型。然后,你在某一时刻想要什么输出,就不再取决于是否有专门的渲染器、模拟器或规划器模型,而更多是——今天我想驱动一个机器人,所以它将处于规划模式;明天我想驱动一个虚拟视频游戏,所以它处于渲染模式;或者后天我想模拟世界中的可能反事实,现在我希望它处于模拟器模式。所以我认为这就是该领域未来几年将要走向的方向。
Exactly. What I think we're going to get to as a field is models that can do all of these jointly. And they're all going to benefit from each other. Why is that? Because they're all kind of asking similar questions: they want to understand what are the kinds of worlds that could exist? How could those worlds respond to action? What do those worlds look like? What kind of observations arise from those worlds? How do they evolve in time? What can you do with them? These are all kind of fundamental questions that all connect to each other, right? If I get better at understanding how the world state evolves, probably I also get better at anticipating what it's going to look like from different angles as a renderer. And if I'm really good at rendering, like imagining how the implicit world state is going to evolve, probably that's pretty good for planning, right? If I can implicitly evolve my world state, probably I can also plan against that state and know how to take actions to affect the world. So I think it's pretty—I don't think we're there yet, but I think over the next couple of years, we'll start to see models that combine more and more of these capabilities into one powerful unified model. And then what output you want at one moment in time is less a function of having specialized models that are renderers or simulators or planners, but more about—is it today I want to drive a robot, so it's going to operate in planner mode; and tomorrow I want to drive a virtual video game, so it's operating in renderer mode; or the next day I want to simulate possible counterfactuals in a world, and now I want it to operate in simulator mode. So I think that's kind of where the field is going to get in the next couple of years.
这让我想到——就像神经网络中的输出层可以是分类或回归或其他什么,但核心表示在很大程度上是相同的,我们只是以不同的方式使用它。
That makes me think of—like this idea that the output layer in a neural network could be classification or regression or something else, but the core representation is for the most part the same, and we're just kind of using it in different ways.
是的,完全正确。我认为这就是我们将要得到的。我们将拥有这些巨大的统一世界模型,它们可能有不同的输入头、不同的输出头,知道如何输入和输出不同类型的东西。但最终,这个模型的所有算力、所有参数都将成为这个共享主干,即世界模型,它知道如何在其权重中隐式模拟任何类型的事物。然后,它可以根据应用的需要,将该世界知识以行动、状态或观察的形式呈现出来。
Yeah, exactly. I think that's what we're going to get. We're going to have these giant unified world models that have maybe different input heads, different output heads that know how to input and output different kinds of things. But ultimately, all the compute, all the parameters of this model are going to be this shared trunk that is the world model that knows how to simulate any kind of thing implicitly in its weights. Then it can surface that world knowledge as actions or as states or as observations as needed for the application.
我想问的下一个问题,实际上也是最后一个问题,是:当前的架构能让我们达到那个目标吗?你刚才说,有点不,架构在演化,但它与我们思考当今模型(如 Transformer 等)的方式是一致的。你是否预见到需要架构上的阶跃函数来实现世界模型的潜力?
I think the next question I wanted to get at, and really kind of a closing question, is: do current architectures get us there? And you just said, kind of no, like the architecture evolves but it's kind of consistent with the way we think about today's models, transformers, etc. Do you foresee kind of an architectural step function being required to fulfill the potential of world models?
也许需要,但不会太大,可能不会太大。我认为 Transformer 真的非常非常强大。
Maybe yes, but not as big a one, probably not too big. I think Transformers are really, really powerful.
是的,我经常被问到这个问题:这是否意味着 Transformer 已死?我们需要新的架构吗?
Yeah, I get this question a lot: does that mean Transformers are dead? Do we need new architectures?
不,Transformer 很棒。Transformer 非常强大。它们扩展得很好。它们可以处理各种不同类型的数据。所以 Transformer 非常强大。你可以将它们用于各种不同的问题。我认为存在一个损失函数的问题,即训练这类生成模型的正确损失函数是什么,而这更不确定。比如,我们是否想将这些事物训练为生成模型?如果是生成模型,你是通过扩散或整流流来训练它?你是将其训练为离散自回归还是其他方式?所以有一个损失函数的问题,它适用于任何类型的生成模型,我认为它可能会演化,而且已经演化过了——比如我们过去用 GAN,现在用扩散。这更多是损失函数的变化,而不是架构的变化。但另一个我认为确实需要在架构上看到一些演化的领域是,我们如何处理非常非常长的上下文和非常非常多的 token。所以这已经在 LLM 中出现了。LLM 是 Transformer。它们处理 token,但你可以解决很多问题。即使在智能体时代这可能不那么正确了,但几年前,很难想象 LLM 需要在百万或千万 token 上运行的情况,因为这就像整本书。原则上,有些问题需要你把整本书塞进上下文,但也许那些不是——你可以在不需要那么大上下文的情况下用语言做很多事情。但现在,如果你想谈论世界,比如我需要生成大量高维图像或大量 3D 空间,很快很容易就会出现需要数十万、数百万或数千万 token 上下文的情况,用于不同的世界建模问题。
No, Transformers are great. Transformers are super powerful. They scale up really well. They can work on all different kinds of data. So Transformers are very powerful. You can use them for all different kinds of problems. I think there is a loss function question about what is the right loss function for training these kinds of generative models, and that's more up in the air. Like, do we want to train these things as a generative model? If it's a generative model, do you train it via diffusion or rectified flow? Do you train it as discrete autoregression or something else? So there is kind of a loss function question that's really applicable to any kind of generative model that I think is probably going to evolve, and has already—like we used to do GANs, now it's diffusion. That's a change in loss function more than it is a change in architecture. But another area where I do think that we need to see some evolution architecturally is how do we deal with really, really long context and really, really lots of tokens. So this is something that comes up in LLMs already. LLMs are Transformers. They work on tokens, but you can do a lot of problems. Even maybe this is less true now in the agentic era, but a couple of years ago, it was hard to imagine situations where an LLM needs to operate on a million or 10 million tokens, because that's like whole books. In principle, there are problems where you need to cram whole books into your context, but maybe those are not—you can do a lot with language without needing that big of a context. But now, if you want to talk about worlds, like I need to generate maybe lots and lots of high-dimensional images or lots and lots of space in 3D, it's very quick, very easy to get situations where you want hundreds of thousands or millions or tens of millions of tokens of context for different world modeling problems.
所以我确实认为我们需要看到一些演进,关于如何让 Transformer 适应非常长的上下文长度,因为这在语言模型里是可选项,但对任何规模化的世界模型来说,这成了日常问题。这就引出一个问题:我们应该如何看待世界模型中的词元和上下文?上下文长度是否限制了世界的规模?还是说我们是在把世界的各个部分分页调出,所以它并不是硬约束?词元是否对应一个 splat 或其他某种结构?这些东西之间是什么关系?
So I do think we need to see some evolution in how we adapt transformers to work at really long context lengths, because that becomes a nice-to-have in LM, but that becomes the everyday problem for any scaled-up world model. That begs the question: how should we think about both tokens and contexts in world models? Does the context length limit the size of the world? Or are we paging out sections of the world, so it's not a hard constraint? And does a token correspond to a splat or some other construct? How do these things relate?
我认为这正是很多不同事情在发生的地方,也是我看到大量演进的地方。所以在某些架构中,一个词元可能是一小段视频——也许是一个 16×16 的空间块和四帧时间。所以在某些架构中,词元字面上就是视频的一个小块。在某些架构中,那个词元可能是世界中某处的一小束 splat。在某些架构中,那个词元可能是一小块 3D 空间。也许我把 3D 空间划分成了体素,现在每个词元对应 3D 空间的一个块。或者那些词元是某种抽象的、潜在的东西——比如我有一个潜在的世界状态,它只是一个数字。也许我把世界状态分配成了一千个词元。它们是什么意思?这是不透明的,模型自己搞清楚了。所以我认为这就是我们看到大量演进的地方,很多不同的方法在架构上做着不同的事情。
I think that's where there's a lot of different things happening, and that's where I do see a lot of evolution. So in some architectures, a token will be maybe a little chunk of a video—maybe a 16 by 16 spatial patch and four frames in time. So in some architectures, a token is literally a little patch of a video. In some architectures, that token might be a little bundle of splats somewhere in the world. In some architectures, that token might be a little chunk of 3D space. Maybe I've carved up my 3D space into voxels, and now each token corresponds to some chunk of 3D space. Or maybe those tokens are something abstract and latent—like maybe I've got a latent world state that is just a number. Maybe I've allocated my world state to have a thousand tokens. What do they mean? It's opaque; the model figured it out. So I think that's where we're seeing a lot of evolution and a lot of different approaches doing different things architecturally.
另外,在这个架构方向上,你认为几何深度学习或融入对称性的不同模型这类想法有作用吗?已经有各种各样的工作试图把几何融入深度学习,至少从某个角度看,如果我们讨论的是空间信息,那可能就有用武之地。
Also in this architectural thread, do you see a role for ideas around geometric deep learning or different models that incorporate symmetry? There's been a variety of work around trying to incorporate geometry into deep learning, and at least from a particular lens, it seems like if we're talking about spatial information, there may be a role for that.
也许吧。但我认为我们在深度学习中反复学到的教训是,你想要简单的表示,然后在简单表示背后扩展模型。所以你还得选择一种适合当前任务的表示,对吧?比如,如果你真正想要输出的是视频帧或图像,那就直接这么做。你不需要通过显式的 3D 表示来限制自己。如果你确实想要 3D 表示——无论是 splat 还是网格——那就找一种简单的方式让那个 3D 表示与神经网络对接。而且你在对称性和复杂表示上投入越多,通常优化就越难,你做的假设也越多。所以它在规模化上就越不奏效,对吧?比如你加入某种对称性概念。嗯,人类通常是对称的——我们有两条腿和两条胳膊——但并不是每个人都有两条腿和两条胳膊。所以那些硬性的对称假设也许能让你走完 80% 或 90% 的路,但它们最终会失效。然后你会希望你的架构或模型有足够的表达能力来处理那些捷径失效的情况。
Maybe. But I think the lesson we've learned over and over again in deep learning is that you want simple representations and then scale up the model behind the simple representations. So you also want to choose a representation that's adapted for the task at hand, right? Like if the thing you actually want out is video frames or images, just do that. You don't need to bottleneck yourself through an explicit 3D representation. If you do want a 3D representation—be it a splat or a mesh—find a simple way to interface that 3D representation with the neural network. And the more you put in symmetries and complicated representations, often the harder it'll be to optimize, and the more assumptions you're making. So the less it's going to work at scale, right? Like you put in some notion of symmetry. Well, people are usually kind of symmetric—we have two legs and two arms—but not everybody has two legs and two arms. So those hard assumptions of symmetry might get you 80% or 90% of the way there, but they're going to break down eventually. And then you'd like to be in a position where your architecture or your model is expressive enough to handle the cases where those shortcuts are going to break down.
人们应该去哪里了解更多?有没有什么权威资源,你愿意推荐给刚进入这个领域、想更深入钻研的人?
Where should folks go to learn more about this? Are there any canonical resources that you like to point folks to who are new to the space and want to dig in more deeply?
是的。你绝对可以试试 Marble。那是我们的产品,在 marble.worldlabs.ai。你今天就可以注册并试用。我希望我能推荐一个更好的关于世界模型的系列讲座或书籍之类的,但我就是觉得目前还没有人写出特别棒的东西,这部分也是我们试图通过博客文章解决的问题。但我认为有人可以深入得多,以更好的方式阐释这些想法。所以如果我漏掉了什么,我很希望你的观众或听众能告诉我。
Yeah. So you can definitely try out Marble. That's our product at marble.worldlabs.ai. You can sign up and play around with that today. I wish I could recommend a better lecture series or book or something on world models, but I just don't think anyone's written down anything super awesome yet, which is partially what we tried to solve with our blog post. But I think someone could go a lot deeper and unpack a lot of these ideas in a better way. So if I'm missing something, I'd love for your viewers or listeners to let me know.
Justin,非常感谢你抽出时间与我们分享你和 World Labs 团队最近在做的事情。非常酷。
Justin, thank you so much for taking the time to share with us a bit about what you've been up to—you and the team at World Labs. Very cool stuff.
是的,非常感谢你邀请我。这非常有趣。
Yeah, thanks so much for having me. This was a lot of fun.