Atlas: Spatial Intelligence Through New View Prediction
打开互动全文版(中英对照 + 朗读 + 问答)→Atlas 是一个世界模型,通过新视角预测生成、重建和模拟世界,仅用三台相机即可实现子弹时间效果。
Atlas is a world model that generates, reconstructs, and simulates the world using new view prediction, enabling bullet-time effects from just three cameras.
呃,昨天是大日子。你们发布了一款新的前沿模型,反响非常好,而且还在持续。我想也许组织这次对话的一个好方式是,先聊聊那到底是什么,然后再回顾历史,一路讲上来。所以也许 Justin,你愿意谈谈昨天发布了什么,为什么它意义重大吗?
Uh, so big day yesterday. You launched a new Frontier model which got an amazing reception which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was and then we'll go back to history and work our way back up. So maybe Justin, do you want to talk about what was launched yesterday, why it's significant?
是的。所以 Atlas 是我们的新一代生成式世界模型。它有三个基本功能:生成、重建和模拟世界。嗯,在其中有几种不同的主要能力。它有非常好的相机条件生成能力。所以你可以输入一张图像,加上一条相机轨迹,然后引导模型,让它沿着你想要的任何视角生成视频帧。它非常擅长稀疏 3D 重建。你可以输入一帧或多达 100 帧的真实世界视图,并用这些来重建真实世界。这种重建既可以表现为穿越空间的新颖视频,也可以是对空间的显式 3D 重建。嗯,最后它还可以用于模拟。为此我们展示了那些很酷的子弹时间视频,在网上引起了很多关注,还有机器人模拟。
Yeah. So Atlas is our new next gen generation world model. It has three basic things. It can generate, reconstruct, and simulate the world. Um, so within that, there's a couple different major capabilities. It has really good camera condition generation. So you can input an image together with a camera trajectory and steer the model and have it generate, you know, video frames along any any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to a 100 frames um that are views of the real world and use those to reconstruct the real world. And that reconstruction can take the case either of a novel of video flying through the space or an explicit 3D reconstruction of the space. Um then finally it can be used for simulation. Um and for this we show off um you know these awesome bullet time videos which got a lot of attention online and then also robotic simulation.
什么是子弹时间视频?
What's a bullet time video?
子弹时间视频来自《黑客帝国》。你知道,在第一部《黑客帝国》电影里有一个著名的镜头,尼奥正在倒下。没错。
A bullet time video this comes from the Matrix. You know, there was a famous shot in the first Matrix movie where Neo is like falling down. Exactly.
所以,然后记得在那个著名镜头里他正在倒下。那是慢动作,摄像机绕着他飞了一圈。嗯,他们拍那个镜头的方式是,周围有一圈几百台摄像机。所以他在工作室里倒下,几百台摄像机从那个角度在绿幕前拍摄,然后他们用那几百台摄像机制作了《黑客帝国》里的那个著名镜头。但现在有了 Atlas,我们只用三台摄像机就能做到。所以不需要摄影棚拍摄,不需要绿幕,不需要昂贵的校准。我们真的可以放三台 iPhone 在三脚架上当摄像机。用这些来拍摄某个事件的视频,比如有人投篮,有人把草莓掉进一碗牛奶里,然后从这三段 iPhone 视频中,我们可以重新构图,想象时间冻结,摄像机在牛奶飞溅时飞入,得到这些惊人的冻结时间视角。
So, and then remember in that famous shot he's like falling down. It's in slow motion and the camera flies all the way around. Um, so that's the way that they did that shot is they had a ring of like hundreds of cameras. So then like he fell over in the studio. They had hundreds of cameras viewing that angle on a green screen and then they used those hundreds and hundreds of cameras to make that that famous shot in in the Matrix. But now with Atlas, we can do this with just through three cameras. So like no studio capture, no green screen, no expensive calibration. We can literally stick like three cameras on three iPhones on tripods. Um use these to take sort of a video of something happening um like someone shooting a basket, someone dropping a strawberry into a bowl of milk and then from those three like iPhone videos, we can then reframe the shot and imagine like a like freeze time, have the camera fly in like as the milk is splashing up and get these amazing frozen time views.
嗯,我们只用几台摄像机就能做到。你能,嗯,也许最简单描述一下 Atlas 做什么?比如输入什么,输出什么?
Um and we can do this with just just a couple cameras. Can can you um just maybe what is the simplest description of what Atlas does? Like what goes in and what comes out?
是的。所以 Atlas 的真正核心原则之一,最根本的一点,是它做新视角预测。嗯,这是一个非常基础的基元,我们认为它超级令人兴奋,是一个全新的基础模型基元,以前没人做过。对吧?所以我们知道大语言模型建立在下一个词预测之上。我们看到视频模型建立在下一帧预测之上。Atlas 真正做的是新视角预测。对吧?给定一个场景的若干视角或场景描述,嗯,这些进入我们所谓的空间上下文,它隐式地描述了我们想要讨论的世界是什么,然后你可以把虚拟相机指向空间和时间中的任意一点,Atlas 会理解那个世界从那个空间和时间位置看起来应该是什么样子。
Yeah. So one of the one of the really core principles of Atlas like the most fundamental thing is it does new view prediction. Um and this is a a really fundamental primitive that we think is super exciting a super new primitive for for base models that no one's ever done before. Right? So we know LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction. Right? that given some number of views of a scene or a description of a scene, um, those go into what we call a spatial context that describes implicitly what is the world that we want to talk about, then you can point a virtual camera at at an arbitrary point in space and time and Atlas will understand what that world is supposed to look like from that position in space and time.
你知道,Ben,市面上有无数视频模型,都声称自己是世界模型,都声称有新视角。你能更具体地剖析一下,这与之前那些形形色色的模型有何不同吗?
You know, you know, Ben, that you know, with with a with a bajillion video models out there all claiming to be world models and all claiming to have novel views and can you maybe tease apart kind of more concretely how this is different from like the myriad models that have come before.
嗯。是的,我认为 Justin 刚才说的空间上下文方面在这里非常重要。嗯,所以有很多视频模型,很多视频模型实际上是因为单图像输入或从首帧到末帧的插值而成名的。现在,我们开始看到一些模型能做这种全参考的事情,比如用 20、30、50 张图像。但 Atlas 的关键在于,它对你放入的每一帧都有一种空间接地气的意义。所以,它不仅仅是一个模型随意解释的图像,或者你可以尝试通过文本提示与它争论,让它做特定的事情。对于 Atlas,每个图像实际上都有一个关联的三维相机位姿。这意味着你可以以极高的精度执行重建任务。对吧?所以,如果我们有这个房间的四个视角,每个角落一个,你可以把这些放入模型,然后得到你在这个房间里看到的一切的精确复制。它不会猜测另一个角落里有什么,比如事物之间的关系。它只会精确地再现你给它的东西。嗯,你也可以在一种创造性的或想象的意义上这样做。如果你从不同的 AI 生成或真实世界位置拍摄两张照片,你实际上可以定位和布置这些,以构建这些有意的定向飞越,这些飞越真正由你放置内容的精确位置以及相机将要看向和移动的位置所控制。呃,我认为这与那种更像老虎机效应的方式非常不同,那种方式你必须一遍又一遍地重试生成,只能通过视频模型获得那种更高层次的文本控制。
Mhm. Yeah, I think what Justin was saying about the spatial context aspect is super important here. Um, so there's many video models, a lot of video models actually got their claim to fame from their single image input or their start to last frame interpolation. Now, we're starting to see models that can do this kind of omni reference thing with, you know, 20, 30, 50 images. But what's key with Atlas is that it actually has a kind of like spatially grounded meaning to every frame you put into it. So, it's not just an image that the model's going to interpret whatever way it wants or you can kind of try to argue with it in the the text prompting and get it to do something specific. With Atlas, every image actually has an associated threedimensional camera pose. And that means that you can perform this task of reconstruction with an extremely high degree of precision. Right? So, if we had four views of this room, one at each corner, you can put those into the model and then get an exact replication of everything you see in this room. And it's not going to guess what's in the other corner, like a relationship between things. It's just going to reproduce exactly what you give it. Um, and you can also do that in a kind of creative or imaginative sense, too. If you take two photos from different, you know, AI generations or real world locations, you can actually position and stage those to build these kind of intentionally directed flythroughs um that are really governed by exactly the precise place that you put the content you want and where the camera's going to look and travel. Uh which is very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that kind of higher level of text control you get with video models.
这是否只是传统视频模型的一个明显扩展版本,还是一个全新的架构?
Is this just kind of an obvious, you know, scaled up version of a traditional video model or is it a new architecture?
我认为这是一个相当新的东西,原因有几个。
I think it's it's a pretty new thing for a couple different reasons.
我们谈到的一点是,它在同一个模型中同时进行生成和重建。就像 Ben 说的,这个东西可以拍下这个房间的几个视角,然后按照你看到的样子重建出房间里的所有东西。从历史上看,重建一直是计算机视觉中的一个独立子领域,有自己专门的任务和模型。而生成则是所有文本转视频模型真正擅长的,就像过去几年我们看到的大型扩散模型。那些模型非常适合创意应用——我想想象一些从未存在过的东西。但如今有了 Atlas,我们首次将视觉智能的这两个不同部分整合到一个模型中。因此,它可以在一个架构中同时进行 3D 重建和生成。
One thing we talk about is that it does both generation and reconstruction jointly in the same model. Like Ben was saying, this thing can take a couple views of this room and then reconstruct everything in this room exactly as you see it. Historically, reconstruction has been its own subfield in computer vision with its own specialized tasks and models. Generation is what all the text-to-video models are really good at, like the big diffusion models we've seen in the last couple years. Those are great for creative applications—I want to imagine something that's never been there before. But now with Atlas, for the first time, we're putting these two different parts of visual intelligence together in one model. So it can do both 3D reconstruction and generation together in one architecture.
抱歉,我插一句——我对这个领域不太熟。所谓 3D,是指深度还是模型,还是什么意思?
Sorry, I just—I don't know the space super well. By 3D, is this like depth or models, or what does that mean?
是的。到目前为止我们使用的形式是深度图。对。所以现在,你可以有一个带有虚拟相机的帧,该相机指示其在 3D 空间中的位置。相机位置和相机参数是模型的原始输入。附着在该相机位置上,你可以有 RGB,告诉你空间中该位置的外观,以及深度图,告诉你该位置在 3D 空间中的空间结构。因此,文本、图像、视频、3D 相机——这些都是这个东西以多模态方式联合处理的模态。
Yeah. So the formulation we used so far is depth maps. Right. So right now, you can have a frame with a virtual camera telling its position in 3D space. That camera position and camera parameters are native inputs to the model. Attached to that camera position, you can have RGB, telling you what that position in space looks like, and a depth map, telling you the spatial structure of that position in 3D space. So text, image, video, 3D cameras—these are the modalities that this thing all does jointly in a multimodal way.
我想补充一点,因为我认为 Justin 刚才说的其实非常重要,Ben 说的也一样,这一点被低估了。这是计算机视觉领域首次将像素生成和像素重建统一起来。这个领域已经存在了半个多世纪。我身处这个领域几十年,无法告诉你有多少博士论文是关于重建或新视角合成问题的。此外,我们的领域传统上有多个方向。你去参加计算机视觉会议,会有像素生成方向、识别方向和 3D 重建方向。这是一个优雅的模型,通过锚定视角和视角估计,将重建和生成问题结合起来或统一起来,这简直强大得令人难以置信。
I want to add something, because I think what Justin just said is actually so important, and also what Ben said, that it's underappreciated. It's the first time we have a unification of pixel generation and pixel reconstruction in the world of computer vision. This field has been around for more than half a century. Sitting here, having been in this field for decades, I cannot tell you how many PhD theses have been written on the problem of reconstruction or novel view synthesis. Also, our field traditionally has multiple tracks. You go to a computer vision conference, you have the pixel generation track, the recognition track, and the 3D reconstruction track. This is an elegant model that combines or unifies the problem of reconstruction and generation by anchoring on viewpoints and viewpoint estimation, and that's just incredibly powerful.
你能不能退一步,补充一些内容?当你创办公司时,我记得你说过要攻克空间智能。现在我们有了这个新模型,对我来说它感觉非常通用。你有下一个词预测,而这是下一个新视角预测。
Can you maybe take a step back and fill something out? When you started the company, I remember you saying you want to tackle spatial intelligence. Now we have this new model, and it feels very general to me. You've got next-token prediction, and this is next-new-view prediction.
新视角预测,对。所以你可以得到一个视图或一组视图,然后得到一个新视图。你能勾勒出这是如何迈向空间智能这一普遍问题的重要一步吗?也许先描述一下什么是空间智能?
New view prediction, right. So you can get one view or a set of views, and you get a new view. Can you pencil out how this is a significant step to the general problem of spatial intelligence, maybe by starting to describe what spatial intelligence is?
嗯,空间智能最终必须使我们能够生成空间是什么、在其中推理,并能在其中编辑和交互。现在我们可以争论它是 3D 还是 4D——最终它是带有时间维度的 4D——但即使只是 3D,这些也是空间智能必须实现的基本任务。然后我们谈到,有了这些,你可以渲染、模拟和规划行动。但要做到这一点,一个需要解决的基本问题是理解空间的几何、结构和物理。
Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and be able to edit and interact within it. Now we can argue whether it's 3D or 4D—ultimately it's 4D with the time dimension—but even just 3D, these are the fundamental tasks that spatial intelligence has to enable. Then we talk about with that, you can render, you can simulate, and you can plan actions. But to do that, a fundamental problem to solve is to understand the geometry, structure, and physics of the space.
是的。
Yeah.
而且我确实相信 Atlas 是向前迈出的重要一步,因为现在每一帧,你都可以生成并估计一个重要的信息,即视角、相机姿态。这是关于空间几何所需的最关键信息。这可以导致我们在模型下游看到的所有涌现行为,我们在博客中展示了这些。因此,在通往空间智能的道路上,生成像素绝对是一个早期步骤,我们已经看到了你所说的无数模型。但生成真正具有空间上下文并扎根的像素绝对是另一个重大步骤,而 Atlas 已经迈出了这非常艰难的一步。
And I do believe Atlas is a significant step forward, because now with every single frame, you can generate and estimate an important piece of information, which is the viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the space. And that can lead to all the emergent behaviors we see in the downstream of the model, which we showed in the blog. So on the path to spatial intelligence, generating pixels is definitely an early step, which we have seen with what you call gazillions of models. But generating pixels that are truly spatially contextualized and grounded is absolutely another major step, and that is the very hard step that Atlas has taken.
我们肯定还有——我们可以继续深入,对吧?还有时间的第四维度,这将带来动态,还有更高保真度的模拟和空间描绘。所以这是空间智能路线图的一部分。
We definitely have—we can just keep going here, right? There is the fourth dimension of time, which will bring in dynamics, and there is higher-fidelity simulation and delineation of the space. So this is part of the roadmap of spatial intelligence.
太好了。是的。我确实想深入探讨它的发展方向,但首先也许让我们谈谈如何走到这一步。World Labs 已经成立多久了?
Great. Yeah. I definitely want to dig into where this is going, but first maybe let's talk about getting here. How long has World Labs been in existence?
两年半。
Two and a half.
两年半。是的。所以你们之前实际上已经发布过模型。那为什么不直接跳到 Atlas 呢?
Two and a half. Yeah. And so you've actually released models before. So why didn't you just jump right to Atlas?
好问题。这太神奇了,对吧?是的。Justin 的团队需要很多芯片。
Great question. It's so magic, right? Yeah. Justin's team needs a lot of chips.
是的,你需要大量 GPU 才能真正扩展这个东西。所以去年我们发布了 Marble 世界模型,那是我们推出的第一个大型世界模型。它为当前的 Marble 产品提供动力。Marble 非常酷——它可以接收图像、视频和文本提示,并用这些生成 3D 世界。但 Marble 和 Atlas 之间最大的区别之一正是输出模态。所以 Marble 确实专注于高斯泼溅作为输出表示。
Yeah, you need a lot of GPUs to actually scale this thing up. So last year we released our Marble world model, and that was the first big major world model that we put out. That powers our current Marble product. Marble is really cool—it can take images, videos, and text prompts, and use these to generate 3D worlds. But one of the biggest differences between Marble and Atlas is exactly what is that output modality. So Marble was really focused on Gaussian splats as an output representation.
所以无论你输入什么,它都会输出一个以高斯泼溅表示的 3D 世界。高斯泼溅确实很有用,对吧?它们非常好用,易于渲染,可以在移动设备、VR 设备上高效渲染,还能与游戏引擎和模拟引擎互操作。所以高斯泼溅有很多优点。但我认为那是之前模型的一个瓶颈。所以我们在 Atlas 中做的是重新设计了一下,我们意识到需要更早地分离这些模态,并让所有这些模态在模型中更统一地工作。所以现在有了 Atlas,基本原语不是像我们可以生成一个高斯泼溅世界。基本原语是,如我们所说,新视角预测,它可以生成 RGB 帧,可以生成 3D,我们可以在需要时用它们生成漂亮的高斯泼溅世界。但我们不需要在不需要时通过高斯泼溅来限制我们的输出。这确实花了很多心血和汗水,才理解这些不同表示的利弊。
So whatever you're inputting, it's going to output a 3D world represented as a Gaussian splat. And Gaussian splats are really useful, right? They're really nice, they're easy to render, they can render efficiently on mobile devices, on VR devices, they can interoperate with game engines and simulation engines. So there's a lot of nice things about Gaussian splats. But I think that was kind of a bottleneck in the previous model. So what we did with Atlas is redesign the thing a little bit, and we realized that we need to bifurcate these modalities earlier and actually have all these modalities working in a more unified way in the model. So now with Atlas, the fundamental primitive is not like we could generate a Gaussian splat world. The fundamental primitive is, as we said, new view prediction, and that can generate RGB frames, that can generate 3D, and we can use those to generate beautiful Gaussian splat worlds when you need them. But we don't need to bottleneck our outputs through the Gaussian splats when we don't need to. And that took a lot of blood, sweat, and tears to understand what are all the pros and cons of these different representations.
是的,这是其中一部分。另一部分是,你得爬上 Scaling(规模扩张)的阶梯,对吧?你得一步步来,做小实验,做小模型,来建立你对什么会奏效、什么能扩展的信心。如果你能立刻知道什么能扩展,那你就直接做那个。但当我们创办公司时,世界大不相同。没有空间智能的缩放定律,对吧?所以当我们创办公司时,世界处于一个非常不同的位置,技术也处于一个非常不同的位置。我们对它要走向何方有很多雄心,但我们花了几次迭代才找到这个我们认为真正能扩展的公式。
Yeah, that's one part of it. The other part is you got to climb the scaling ladder, right? You got to work your way up and do smaller experiments, do smaller models to build your conviction on what's going to work and what's going to scale. And if you could instantly know the right thing that's going to scale, you should just do that. But when we started the company, the world was a very different place. There was no scaling law of spatial intelligence, right? So when we started the company, the world was in a very different place, the tech was in a very different place. We had a lot of ambitions about where we wanted it to go, but it took a couple of iterations for us to hit upon this formulation that we thought is actually the one that can scale out.
你知道,Ben,作为 NeRF 的创造者,做了很多 3D 和重建工作,对我来说并不那么明显,如果你有多个视角,你最终会得到一个 3D 的东西。但你的职业生涯似乎就是致力于最终得到 3D 的东西。所以也许谈谈那一步。
You know, Ben, being the creator of NeRF and doing a lot of 3D and reconstruction, it's not so obvious to me that if you have multiple views, you actually end up with a 3D thing. But you've kind of made a career of ending up with a 3D thing. So maybe talk a little bit about that step.
是的。是的。嗯,是的。我的意思是,正如你所说,我花了很多很多年的职业生涯,实际上是我职业生涯的绝大部分时间,致力于从图像生成 3D 物体。这实际上是我们公司早期经常讨论的事情,比如,这会是生成 3D 的方法吗,对吧?我们是要合成多个视角,然后从中构建 3D 吗?我们是要尝试直接生成 3D 吗?在这个领域,关于哪种方法会胜出,或者哪种方法会更快获得最大优势,一直存在很多不确定性。但我确实有很强的信念,仅仅是因为看到了我几乎可以称之为蛮力 Scaling(规模扩张)的力量,在一个非常非常小的婴儿规模上,不是真正的模型 Scaling(规模扩张),而是我们在过去三年中看到的密集重建的扩展。
Yeah. Yeah. Um, yeah. I mean, as you said, I've spent many, many years of my career, the vast majority of my career actually, working on producing 3D things from images. And this is actually something we talked a lot about early on in the company, even of like, is this going to be the approach that produces 3D, right? Are we going to synthesize multiple views and then build 3D out of that? Are we going to try to go direct to 3D? There's been a lot of uncertainty in the field around which of those approaches will win out or will reap the best advantages earlier on. But I did have a lot of conviction just from seeing the power of what I would almost call the brute force scaling, in a very, very, very small baby scale, not like real model scaling, but the scaling of dense reconstruction that we had seen happening over the past three years before.
为什么密集重建是密集的?
Why is dense reconstruction dense?
是的,密集。
Yeah, dense.
因为我知道我们要讨论稀疏,我想确保人们理解什么是密集,什么是稀疏。
Because I know we're going to talk about sparse, and I want to make sure that people understand what is dense and what is sparse.
是的,所以我认为这实际上,即使在商业方面,也是将 3D 重建技术产品化的根本挑战之一,对吧?人们通常不会,在随意的意义上,你想,我拍了三张这个物体的照片,或者我拍了六张这个房间的照片。我看着照片,我能在脑海中理解它们是如何拼合的,我可以填补空白并理解它。但那些真正数据驱动的先验和蛮力密集重建之间从来没有真正的调和,后者实际上更接近于科学或医学成像。我们在密集重建中做了什么,对吧?你基本上必须说,我想在这个重建中出现的每一个东西,我至少需要它的三个或四个视角。如果你想想这个,即使就在这个房间里,对吧?比如麦克风下面,桌子下面,每个裂缝和缝隙之间,还有植物叶子之间,对吧?要真正得到一张覆盖所有这些点的图像,这是一个非常繁琐和详尽的过程,需要绕着房间走。我想你们都见过我在各种地方跑来跑去捕捉它们。对于一个训练有素的人来说,可能需要几分钟。但如果你把一个普通消费者,甚至一些第一次尝试这样做的专业人士,一个普通的手机摄像头或捕捉设备,他们可能需要一个小时。我见过有人第一次尝试扫描一个多房间环境,花了两个小时走过它,并获得足够的覆盖。这只是一个非常详尽和繁琐的循环。所以是的,当我们说密集时,我们真的是指密集。就像这个房间,我想要这么多照片,对吧?我想要 100、200、300 张这个房间的照片来捕捉它。而我们试图做的是把它减少到三张,对吧?我们说的是 50 到 100 倍的减少。然后在这个规模上,它完全颠覆了那种计算,即你重建什么类型的捕捉。你可以回到你现有的图像。你可以去互联网上找到的东西,甚至从中构建场景。你可以去随意的视频,挖掘出很多过去我们永远不会认为是可重建的镜头,然后回去,可能把它变成 3D 并赋予它生命。这是我们用 Atlas 玩了很多的东西,对吧?比如拿旧片段,我拿了我自己很多以前从未成功过的旧捕捉,然后把它们放进系统,第一次看到重建,或者拿我的旧捕捉,扔掉我拍的 95% 的照片,然后你知道,想象出我永远不会从传统的 NeRF 或泼溅式重建中得到的角度。
Yeah, so I think this has actually, even on the business and commercial side, been one of the challenges of productizing 3D reconstruction technology at a fundamental level, right? People kind of don't, in a casual sense, you think, I took three photos of this object or I took six photos of this room. I look at the photos, I can understand in my mind how those piece together, I can kind of fill in the gaps and get it. But there's just never been really any kind of reconciliation between those really data-driven priors and the brute force dense reconstruction, which actually is much more akin to almost like scientific or medical imaging. What we did in dense reconstruction, right? You basically have to say, every single thing I want to appear in this reconstruction, I need at least three or four views of it. And if you think about that, even just in this room, right? There's like under the microphone, under the table, between every different crack and crevice and the plant leaves, right? To actually truly get a picture that covers every one of those spots, it's this very tedious and exhaustive effort to walk around the room. I think you've all seen me running around various places capturing them. It takes, for someone who's well trained per se, it can take minutes. But if you hand a casual consumer or even some kind of professional trying to do this for the first time, an average cell phone camera or capture device, it's going to take them probably an hour. I've seen someone for their first time trying to scan a multiroom environment spend like two hours walking through it and get enough coverage. And that's just this very exhaustive and tedious loop. And so yeah, when we say dense, we really mean dense. It's like this room, I want so many photos, right? I want like 100, 200, 300 photos of this room to capture it. And what we're trying to do is bring that down to like three, right? We're saying like 50, 100x reduction. And then that's at that scale where it just completely flips that calculus on its head of what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and unearth a lot of footage in the past we would never have treated as reconstructible, and go back and bring it to life as 3D potentially. This is something we've been playing around with a lot with Atlas, right? Like taking old clips, like I've taken a bunch of my own old captures that never worked before, and then put them through the system and then seen a reconstruction for the first time, or taken my old captures and thrown away 95% of the photos I took, and you know, imagine angles that I never would have gotten from a traditional kind of NeRF or splat type reconstruction.
在演示网站上,有一个被低估的东西是斯坦福演示,Ben 展示了从 3 到 25 张图像不等。你可以重建整个斯坦福广场。但问题是,我们不得不从空中视角展示它。
One thing that's underappreciated on the website of the demos is the Stanford demo where Ben showed anywhere between 3 to 25 images. You can reconstruct that entire Stanford quad. But the thing is we had to show it from an aerial view.
但每一张输入图像都是 Ben 站在地面上拍摄的。所以你看到的一切都是生成的,但遵循重建的规律,这真的很神奇。
But every single input image is Ben standing on the ground taking a picture from the ground. So everything you see are generated but according to the laws of reconstruction and this is really magical.
而在这里,生成和重建需要以一种根本性的方式相互作用来解决这个问题,因为在 Ben 所讲的经典重建方法下,你需要这么多视角的原因是我需要多张图像,需要在 3D 空间中三角测量这个点,并从多个视角观察它。所以这意味着在传统版本中这是必需的。
And this is where generation and reconstruction need to interplay in a really fundamental way to solve this problem because under the classic kind of reconstruction stuff that Ben was talking about, the reason you need so many views is because I need multiple images and I need to triangulate this point in 3D space and see it from multiple viewpoints. So that means that's required in the traditional version.
另一方面,任何在这些视角中未被捕捉到的东西,比如在输入视图中不可见的像素,都会成为 3D 重建中的空洞,因为从根本上说,如果一个物体在输入视图中不可见,你需要想象它来填补空白。这本质上是一个生成过程。
And on the flip side, anything that wasn't captured in these views, like any pixel that was not visible in one of the input views, will be a hole in the 3D reconstruction because fundamentally, if a thing wasn't visible in the input views, you need to imagine it to fill in the gaps. And that's fundamentally a generative process.
所以即使在这个房间里,即使我们让 Ben 拿着单反相机自由拍摄,让他捕捉这个房间的数百个视角,即使是做这种密集捕捉的世界专家,也仍然会遗漏一些地方。他不可能拍到所有麦克风下面、所有桌子底下或所有椅子灯之间。无论你获得多少视角,你总会遗漏一些东西。
So even in this room, even if we set Ben loose with a DSLR and let him capture hundreds of views of this room, even the world expert on doing these dense captures is still going to miss some spots. He's not going to get underneath all of the microphones or underneath all the tables or in between all the chair lights. You're always going to miss something no matter how many views you get.
所以这就是为什么你需要生成作为模型中的另一种机制,因为你永远无法获得一切。所以你需要模型具备一定的生成能力,让它根据我所看到的,先三角测量我能看到的,然后填补那些不可避免未被捕捉到的空白。
So that's where you need generation as another mechanism in the model because you're never going to get everything. So you need to have some generative capacity for the model to imagine, based on what I'm seeing, first triangulate what I can see but then fill in the gaps of the stuff that inevitably was not captured.
是的。关于这一点,有件很酷的事情是,LLM 很早就理解了这一点,对吧?几乎有过所谓的上下文战争。在最初的几年里,就像,哦,我们到了 128,到 256,到 512,我们到了一百万,对吧?现在每个人都在一个相当具体的层面上理解了它的价值,你知道,当你使用编码模型时,你会把上下文长度调得很高。这是一个难题。每个人都有这种感觉,但没有人以同样原则性的方式在图像和视频模型方面推动这一点。比如没有人试图把一个小时的视频放进去,然后做大海捞针式的检索,找到第 37 分钟的那一帧。
Yeah. And there's something super cool about this that LLMs have really understood this for a long time, right? There was kind of almost these context wars. In the first couple years, it was like, oh, we got to 128, to 256, to 512, we got a million, right? And everyone kind of understands now at a pretty tangible level the value of, you know, you crank your context length to high when you're using your coding model. It's a hard problem. Everyone has a feel for that, but no one has pushed that at all on the image and video model side in the same kind of principled way. Like no one's out there trying to put an hour-long video through and do a needle-in-a-haystack retrieval of a frame at the 37-minute mark.
而在重建和生成中,你实际上有完全相同的东西:重建就是带有非常长上下文的生成,你往里面放很多东西,对吧?这才是真正构建这个连续体的方式,你在这两者之间架起桥梁。
Whereas with reconstruction and generation, you actually have the same exact thing: reconstruction is just generation with a really long context and you put a lot of stuff in it, right? That's the way to actually build this continuum where you kind of bridge between those two things.
而 Atlas 能够做到这一点,这是我们用 Marvel 永远无法做到的。Marvel 有一个根本性的障碍,就是无法塞入超过几张图像。而 Atlas,我可以拿一个 64 张图像的捕捉,对整个房子进行飞越,一切都基于被看到、几乎被看到或从不存在的东西中略微外推而得到 grounding。
And Atlas, being able to do this, is something we could never do with Marvel. Marvel had this kind of fundamental blocker of not being able to jam more than honestly a couple of images in. Atlas, I can actually take a 64-image capture and do a fly-through of an entire house, and everything is grounded by being seen or almost seen or slightly extrapolated from what's not there.
但你只是得到这些——我拿我用 2000 张图像对多房间房屋进行的捕捉,把它减少到大约 30 到 40 个输入,飞越看起来基本相同。这在以前是完全不可想象的。这一切都得益于构建了一个可以优雅扩展的上下文窗口,你可以把东西倾倒进去。
But you're just getting these—I'm taking captures I did with 2,000 images of a multi-room house and taking it down to like 30 to 40 inputs, and the fly-through looks basically the same. And this is just totally inconceivable before. And it's all enabled by building this gracefully scaling kind of context window that you can dump stuff into.
所以思考方式是:稀疏性是你实际拍摄的图片,然后 Atlas 作为模型创建其余的视角,然后你使用经典的重建技术。这大致是思考方式吗?
And so the way to think about it is: the sparseness are the pictures that you physically took, and then Atlas as a model creates the rest of the views, and then you use classic reconstruction techniques. Is that roughly the way to think about it?
在某种意义上,是的。我的意思是,这就是 Atlas 的美妙之处。就像你可以拿你拥有的任意数量的输入,少到单个视图,然后你几乎可以把 Atlas 当作这个渲染引擎,生成你想要的任何其他东西。你可以导航一个虚拟相机。
In some sense, yeah. I mean, that's the beauty of Atlas. It's like you can take however many inputs you have, down to a single view, and then you can almost use Atlas as this rendering engine to produce anything else you want. You can navigate a virtual camera.
是的,没错。你可以说,好的,我这里有一张图片。我想要那里的图片,那里,那里。你可以做几张。然后你可以进行密集的飞越。你可以按顺序做,因为它是一个自回归模型。由你来选择在生成时交互式地添加到上下文中的内容。
Yeah, exactly. You can just say, okay, I have a picture here. I want the picture there, there, there. You can make a couple of those. Then you can stage a dense fly-through. You can do this in sequence because it's an autoregressive model. It's up to you to pick and choose what you add interactively into the context as you generate.
我的意思是,让我震惊的是,听着,我有一个非常简单的思维模型。我有四张图片,然后我得让模型在它们之间外推,然后当你控制时它必须匹配。它必须是 3D 的,因为我总是认为这些扩散模型在视觉上很棒但不准确。所以,我甚至不知道这里是否有问题,但房间是如何匹配的?它如何做到 3D 一致的?仅仅是因为大量数据吗?
I mean, the thing that just blows my mind is, listen, I just have a very simple mental model. I have four pictures and then I've got to have the model extrapolate between them, and then it has to fit when you're in control. It's got to be 3D because, and I always think of these diffusion models as being visually great but not accurate. And so, I don't even know if there's a question here, but how is it that the room fits? How is it that it's 3D consistent? Is it just lots of data?
是的。我的意思是,部分是对 Scaling 假设的信念,对吧?比如,你——顺便问一下,我必须问——当你开始做这件事时,你知道它会成功吗?
Yeah. I mean, it's partially a belief in the scaling hypothesis, right? Like, did you—by the way, I have to ask—when you set out to do this, did you know it was going to work?
我相当确定。
I was pretty sure.
那太——所以你确定吗?我认为我们三个人对缩放定律有完全的信念。我确实认为确切的架构选择和数据混合是魔鬼所在之处。
That's so—so were you sure? I think three of us have total conviction about the scaling law. I do think the exact architecture choices and data mixtures are where the devils are in the details.
我看着 Justin 和他的团队从“我们真的不知道这需要多长时间”到“哦,也许生活的一面”到“哇,这会成功的”。所以没有人做过,但我认为假设——假设一是缩放定律假设,另一个是下一个视角预测——我们从早期就对这两件事有信念。
I have watched Justin and his team going from 'we really don't know how long this is going to take' to 'oh maybe side of life' to 'wow this is going to work.' So no one has done it, but I think the hypothesis—hypothesis one is scaling law hypothesis, the other one is next viewpoint prediction—we had conviction of these two things from early on.
所以我认为我非常确信它会成功。我不确定它会这么快就效果这么好,对吧?就像我认为我们做这件事有可能——也许不清楚的是,用新架构和新范式预训练新模型的第一个周期,就像第一个周期就能成功是疯狂的。
So I think I was very convicted that it was going to work. I was not sure it was going to work this well at this fast, right? Like I thought there's a chance that we do this—maybe it's not clear that the first cycle of pre-training a new model with a new architecture and a new paradigm, like the first cycle of that working is insane.
所以我认为有可能我们可能不得不进行更多轮预训练周期,才能达到我们想要的质量水平。
So I thought there was a chance we might have had to do a couple more turns of that pre-training cycle before we got to the level of quality we wanted.
是不是——我们是否处于这种架构方法的 Scaling 的末期?我们需要另一个突破,还是有……
Is it—are we kind of at the end of the scaling for this architectural approach? We need another breakthrough, or is there...
不,不,我们才刚刚开始。
No, no, we're at the beginning.
真的吗?在不改变架构的情况下?
Really? Without changing the architecture?
我认为我们基本上处于起步阶段。
I think we're basically at the beginning.
我觉得我们基本上还处于起步阶段,目前主要受限于算力,对吧?数据很重要,正如李飞飞喜欢指出的,但一切都有瓶颈,我认为继续扩展这个东西的主要瓶颈实际上是训练算力。在开发过程中,我们训练了一系列模型,我们在博客里稍微提到过,我们训练了几个模型,就像缩放阶梯的前几级,每次我们把模型做得更大,每次训练时间更长,每次用更多芯片,模型就显著变好。
I think we're basically at the beginning and we're basically limited by compute at this point, right? Like data is very important as Fei-Fei likes to point out but like everything has a bottleneck and I think the main bottleneck on continuing to scale this thing is actually training compute right like during development we trained a sequence of models we read about this in the blog post a little bit but we trained a couple models the like the first couple rungs of the scaling ladder um and each time we made the model bigger and each time we trained it for longer each time we put it on more chips like it got significantly better.
而且模型大小,我们在博客里展示的那个模型显然是我们训练的最大最好的一个,但限制它的不是规模或数据之类的,而是我们有一个发布截止日期,因此我们根据那个截止日期来安排我们能负担得起的训练。
And the model size that we like the model that we showed in the blog post is obviously the biggest and best one that we trained but The thing that was limiting it was not the scale or the data or anything like that. It was literally like we had a deadline of when we wanted to release this thing and therefore we backed up what we could afford to train and that deadline.
但这里有个内部故事,贾斯汀和团队在训练从较小到稍大的模型,他们有这些路线图,然后在夏天早期有一天,还不是当前 Atlas 模型的大小,是个更小的模型,然后本·贾斯汀·本把它输入到视角生成中,还记得那张著名的花园桌子,用于 NeRF 论文和许多论文的,一夜之间我收到一条 Slack 消息。我们所有人都看到了本发的 Slack,那个飞行器带着足球从桌子底下飞过。
But here is a little bit of an insider story right like Justin and team are are training the from the smaller and slightly bigger you know are train having these road maps and then there was one day in summer early summer that it's not even this the current atlas model size it's a smaller model and then Ben Justin Ben feed it into you know the viewpoint generation and remember that famous table the garden table for Nerf paper and many papers that overnight I got a slack. I mean we all saw the slack from Ben that the flight flew through under the table with the soccer ball.
是的,带着足球。足球是涌现出来的还是原始图片里就有的?
Yes, with the soccer ball. Is the soccer ball emergent or is that in the original picture?
它是真实的,对吧?
It's in it was real, right?
好的。
Okay.
那天早上,我们三个人对视一眼,说就是它了。这就是我们要构建的东西。我们在五秒钟内就做了决定。这结果以前没人见过。
That morning the three of us looked at each other in the eyes and say that's it. This is we're going to build this. Like we we made a decision within five seconds. how this is just a no one has ever seen this result.
本,你能更具体地谈谈用例吗?World Labs 历来有很多创意用户,他们用它来保持一致性,比如在 2D 图像、电影、3D 和游戏等方面。那么你能谈谈这如何扩展用例或满足现有用例吗?然后我们想谈谈机器人技术。
Ben, um can you talk through maybe more specifically the use cases? So, so uh World Labs has historically had a lot of users that were creatives and they use it for like consistency in you know like whatever 2D images and for movies and for 3D and for games etc. And so maybe can you talk about how this extends use cases or cater to the existing ones and then we'll I'd like to talk about robotics actually.
是的,当然。嗯,其实有点好笑。我们看到人们使用 Marble 的主要方式之一正好符合这个新的视角预测用例。我们很多 Marble 是之前的,抱歉,Marvel 是我们之前的产品。人们会拿那个产品,输入一张图片,得到一个完整的 3D 场景作为高斯溅射,从不同视角截几张图,然后离开,对吧?我们就想,我们可以直接生成那些图像,而且这属于数据角色,对吧?所以我认为,你知道,那里有很多降级。他们会说,哦,这面旗子看起来可以更好。然后我们说,好吧,如果我们用那种精确的控制心态来生成式建模那些视角呢?所以我认为,即使像视图合成这样的核心能力,长期以来一直是个学术问题,但在这个意义上,哦,你要做这种密集捕捉,生成式视图合成是一个相对较新的问题,我们看到很多人在这个创意流程中,人们有多阶段的工作流程。我不认为有人会用一个单一的模型,即使是 Canse 或其他什么,来完成整个任务。人们会有一些故事板和情绪板,从他们最喜欢的图像模型集合中提取图像,然后他们会去不同的视频工具,把这些作为关键帧组合起来,然后他们会在后期剪辑和编辑,对吧?所以我们看到 Marble 有一个小众但非常具体的用例,就是提供那种保证,让你可以把生成内容锚定在某种 3D 一致的世界中,对吧?人们,你知道,我知道我和各种图像模型斗争过,让它们给我一个房间的不同视角,每次你都能看到,哦,东西移动了,不稳定。即使这样一个用例的种子,我认为也暗示了表面下隐藏的价值,就像几十年来人们习惯了持久的 3D 状态,虚拟建模他们在现实世界中会做的事情,有一个舞台、道具和元素,无论是电影、节目、营销镜头还是构建游戏环境,这种状态性和持久性对于人们如何思考空间推理和随时间开发环境至关重要。人们不会以这种短暂的方式思考,生成一个东西,生成一个东西,然后扔掉,保留我的文本提示。人们想要构建一个资产集合,并以那种方式建模一个世界。所以我们试图提供,再次,通过这个空间上下文机制和其他东西,我们试图提供那种控制和精确度,以及摄取不同输入模态的能力,从姿态图像开始,但你知道,我们想给人们更多控制场景中元素的能力,以及编辑和交互,所有这些都在未来。我认为这解锁了我们已经看到的那些领域的进一步用例,但也扩展到任何人们想要创建虚拟复制或预想他们需要构建的真实世界空间的地方,比如建筑和施工。我和一个人聊过,他做会议展台,对吧?世界上有很多东西你可能没想到需要制造。每一个基本上都要经过这个相当费力的虚拟设计阶段。在这个过程中,进入 3D 软件的部分是目前最费力、劳动密集的部分之一。比如从创意总监、设计总监或建筑师那里得到口头评论、草图或非常快速的反馈,然后映射回 3D 表示,这占了 95% 的工作,对吧?你可以开会得到反馈,然后回去做修订。那只是因为我们的软件目前已经过时了几十年。
Yeah, sure. Um, yeah. I mean, it's kind of funny actually. One of the kind of main ways we even saw people using marble plays exactly into this new view prediction case. Like a lot of our marble being the previous, sorry, Marvel our previous product. Um, like people would take that product, put an image in, get a full 3D scene as a gausian splat, take a couple screenshots of it from different points of view and leave, right? And we're like, we can just make those images and that's in data role, right? So, so I think like and you know there's a lot of degradation there. They're like, oh, this flag could look better. it's like okay what if we just generatively model those viewpoints with that exact mentality of control. So I think like even that core capability of just like view synthesis um it's sort of been this academic problem for a long time but in the sense of oh you're going to do this really dense capture like like generative view synthesis is a relatively quite a new problem and we just see so many people who uh in this creative pipeline right people have a multi-stage workflow right I don't think there's a single person out there using one monolithic model not even cance or whatever for for their entire uh task uh people will have this like you know kind of bunch of storyboards and mood boards of images they pull out from like their favorite collection of image models and then they'll go to different video tools and like pull those together as key frames and then they'll go and like clip and edit those later, right? So, we were seeing this like sort of, you know, niche but very specific use case for marble as just providing that like sanity that you can ground your generations in some kind of 3D consistent world, right? people, you know, I know I' I've fought with with various image models to ask them to like give me different viewpoints of a room and every time you can just look and see, oh, things kind of moved around like it's not stable and like even that one seed of a use case I think kind of signals that there's this value and there's hiding under the surface there like there's just decades of people being used to persistent 3D state like virtually modeling what they would be doing in the real world and having you know a stage and props and like elements there whether it is or uh a movie or a show or marketing shot or like building out game environments like this this statefulness and persistence is so key in how people think about spatial reasoning and like developing an environment over time. Like people don't think in this ephemeral like generate a thing, generate a thing, like just throw it away, keep my text prompts. Like people want to build this like collection of assets and like model a world in that way. So we're trying to provide like again with with this spatial context mechanism and other things like we're trying to provide that level of control and precision uh and the ability to ingest different modalities of input starting with the pose images but you know we want to give people more control over the elements of the things in the scenes you're looking at and editing and interaction all that as we go forward. Um, and I think that that it unlocks like further use cases in those areas we're already seeing, but um, also expanding out into kind of any place people want to create a virtual replication or or like you know a pre-imagination of a real world space they need to build, right, for architecture and construction. Like I talked to a guy at some point building booths for conferences, right? There's just so many things in the world you don't think about need to be fabricated. And every single one of those basically goes through this like pretty painstaking virtual design phase. Uh, and of that process, like the part where you go into 3D software is kind of one of the most like arduous and like labor intensive parts right now. Like like taking feedback on a 3D design from kind of like verbal commentary or sketches or really really quick um stuff you got from like a creative director or like a design director or an architect or whatever. Like mapping that back into the 3D representation is like 95% of the work, right? you can have a meeting and get feedback and then you go back and do a weaker revisions. And that's just because like our software is kind of decades old at this point.
而且它从来没能变得像玩乐高、陶艺、手工或铅笔画那样直观。而这里正是 AI 真正能为人们的过程释放巨大价值的地方,无论是创意应用、工业设计还是其他。这激励我打造不同版本的模型来服务这类人群。
And it's just never became as intuitive as you know playing with Legos or like pottery or doing this stuff with your hands or sketching with a pencil. And this is the real place where AI can actually unlock a ton of value for people in their process, whether it's a creative application or something more industrial or design or whatever. And that really motivates me to build different flavors of our model to cater to those kind of people.
是的,我能理解这对创意人士的帮助,比如 Marble 就做到了,以及它如何扩展到设计或建筑等领域。但是,Fei-Fei,你收购了一家机器人公司。所以,这……
Yeah, I can understand how it helps with the creatives because like Marble did that, and also how that extends to things like design or architecture. But Fei-Fei, you acquired a robotics company. So, it's...
我们刚才也谈到了。
We just talked about it too.
我知道,我知道,但对我来说不太明显,尤其是在 Atlas 的背景下,这如何映射到机器人技术。所以,如果你不介意的话,请简单说明一下。
I know, I know, but it's less obvious to me, especially in the context of Atlas, how that maps to robotics. So, if you wouldn't mind just penciling that out.
是的,实际上 Atlas 是拼图的关键部分。我们收购了这家前身为 Scenix 的公司。他们的关键技术是什么?目前,他们的关键技术是一个从真实到模拟、再从模拟到真实的系统。这在机器人领域意味着什么?假设你想训练一个机械臂在工业环境中进行布线。你需要大量数据来训练一个机器人策略来完成这些布线活动,然后评估策略是否表现良好,最后将机器人部署到布线环境中。
Yeah, actually Atlas is a key part of the puzzle. So, we acquired this company that was formerly known as Scenix. And what is their key technology? Right now, their key technology is a system that goes from real to sim and then sim to real. And what does that mean in robotics situation? You want to train a robotic arm to, you know, figure out how to do cabling, let's say, in an industrial setting. Well, you need a whole bunch of data to first train a robotic policy to do these cabling activities, and then you want to evaluate if the robotic policy is doing a good job, and then you deploy the robot into the cabling environment.
是的。
Yeah.
为了训练,这家公司(Scenix)以及现在的我们机器人团队过去所做的正是 Ben 所说的密集重建。你拍摄场景图片,然后尝试重建环境。这极其痛苦、耗时且费力,严重阻碍了机器人模拟(真实到模拟)的速度,对吧?所以 Atlas 确实是这方面的下一代技术。这不仅适用于机器人布线等。我们应该退一步看,当前机器人领域最大的问题实际上是数据。总有一天会是芯片,但现在是数据。
In order to train, what this company, Scenix, and now our robotics team used to be doing is doing exactly what Ben was saying: dense reconstruction. You take pictures of a situation and then try to reconstruct that environment. It's excruciatingly painful, takes a long time, laborious, and it really blocks the velocity of robotic simulation, real to sim, right? So Atlas really is the next generation technology for that. And this is not just for robotics cabling or anything. We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data.
因为收集机器人在真实世界操作的数据非常困难。而且,为了……你不仅需要收集比如布线、洗碗等场景的数据,还有一个非常重要的步骤叫做随机化。你必须对同一环境进行条件随机化。所以电缆不能只朝一个方向弯曲,它可以朝不同方向弯曲,或者盒子可以有不同尺寸、颜色、盖子等,或者出现在场景的不同位置。因此,你必须经历真实到模拟的过程,才能获得足够的数据,再加上从互联网获取的其他数据。所以,这个真实到模拟的步骤将得到 Atlas 的巨大帮助。
Because it's so hard to collect real world data where robots are operating on. And in order to... not only you need to collect the data of, let's say, the cabling situation or dishwashing situation or whatever, there is also a very important step called randomization. You have to take the same environment and then randomize the conditions. So the cable doesn't literally only bend this way; it can bend a different way, or the box can have different sizes, colors, different lids, and all that, or in different parts of the scene. So you have to go through a real to sim situation in order to get enough of that data, in addition to other data you can get from internet. So this real to sim step will be really helped by Atlas.
这只是第一部分,满足当前技术下的机器人需求,因为我们还没有一个足够强大的前沿基础模型用于机器人。但 Atlas 是一个全模态模型,一个多模态模型。它接受不同类型的输入并生成不同类型的输出。你完全可以想象下一步是 Atlas 接收动态数据,这能真正弥合行动规划与机器人之间的鸿沟,以及 Atlas 的输出。所以这就在路线图上。
That's just the first part of this, meeting the robotics needs in the current technology, because we don't yet have a frontier foundation model that's robust enough for robotics. But Atlas is an omnimodel, a multimodal model. It takes different kinds of input and generates different kinds of output. You can totally imagine the next step is Atlas taking in data that's dynamical, and that can really start to bridge the gap between action planning and robotics, and the Atlas output. So that's on that roadmap.
你刚才想说什么?
Were you going to say?
是的,我想说的是,训练机器人策略与我们见过的任何其他 AI 应用都有根本性的不同。比如生成代码、生成图像、生成视频,模型本质上是在创造这个产物,而这个产物……有很多产物的例子你可以从网上或其他地方收集,对吧?你想生成图像,那里有很多图像。你想生成视频,那里有很多视频。你想生成代码库,那里有很多代码库可以学习。
Yeah, I was going to say there's something fundamentally different about training a robotics policy compared to really any other application in AI we've seen before. And that's like if you're generating a piece of code, like you're generating an image, you're generating a video, the model is fundamentally creating this artifact, and that artifact... there's a lot of examples of artifacts that you can go out on the web or somewhere and collect, right? You want to generate images, there's a lot of images out there. You want to generate videos, there's a lot of videos out there. You want to generate a codebase, there's a lot of codebases out there you can learn from.
是的。
Yeah.
机器人策略是根本不同的。它不是产生静态的东西,而是一个策略,要走向世界,采取行动,并努力实现目标。而世界并不总是按你预期的方式回应,对吧?意外总会发生。所以机器人策略本质上是一个在现实世界中与真实世界互动的智能体,事情会发生。所以你需要……其中关键部分是,这些策略在训练期间需要暴露在部署时可能出错的所有可能情况中。
A robotics policy is something fundamentally different. It's not producing a static thing. It's instead a policy that's going to go out into the world, make actions, and try to achieve a goal. And the world's not always going to respond the way you expect, right? Unexpected stuff is going to happen. So a robotics policy is fundamentally an agent that is out in the real world interacting with the real world, and stuff happens. So you need... a critical part of that is those policies during training need to be exposed to every possible thing that could go wrong during their deployment.
而这正是模拟对机器人至关重要的地方,对吧?所以有两个角度。一个是经典模拟。你可以使用你喜欢的物理引擎,作为人类设计师创造性地想象完成此任务时可能发生的所有场景,然后尝试编写显式代码来建模所有场景。这是一个角度。对于编码智能体来说,这是一个有趣的角度,实际上它也会得到增强。但另一个角度是尝试更数据驱动的模拟,对吧?也许我们可以有一个学习模型,它能理解环境、世界如何对行动做出反应,有时可能以意想不到的方式反应。那么我们能否构建这些神经模拟器,用尽可能多的数据进行训练,然后使用这些神经模拟器作为训练机器人策略的模拟基础?这是 Atlas 一个非常有趣的未来方向。
And that's where simulation is really key for robotics, right? So there's two angles on that. One is the kind of classical simulation. You can go to your favorite physics engine and try to imagine creatively as a human designer what are all the scenarios that might happen when achieving this task, and then try to write explicit code that models them all. That's one angle. And that's an interesting angle with coding agents, like that actually gets supercharged too. But there's another angle, which is to try more data-driven simulation, right? Like, maybe we can have a learned model that can understand how the environment, how the world is going to respond to actions, and maybe it might respond in unexpected ways sometimes. Then could we build these neural simulators that are trained on as much data as we can, and then use these neural simulators, these learned neural simulators, as a simulation bed to train robotic policies? And that's a really interesting future direction of Atlas.
但事情并不止于此,对吧?一旦你有了这个学习型模拟器,这个学习型模拟器已经在它的“大脑”中理解了世界。它理解世界将如何对行动做出反应。那么为什么模拟器本身不能成为规划者呢?
But then it doesn't stop there, right? So once you have this learned simulator, like this learned simulator already has in its mental brain an understanding of the world. It understands how the world is going to respond to actions. And why doesn't the simulator itself become the planner?
对,这正是我们围绕世界模型及其通用性的核心论点。模型应该理解一些核心内容,比如生成世界、模拟世界、理解它们在不同情境下的表现,而理解世界如何对行动做出反应与想象我需要采取什么行动来让世界以特定方式反应高度相关。
Right, and that's kind of the core thesis that we've had around world models and their generality. There's some core stuff that a model should understand around generating worlds, simulating them, understanding how they appear in different situations, and understanding how the world's going to respond to an action is highly related to imagining what kind of action I need to take to make the world respond in a particular way.
是的。
Yeah.
什么,顺便说一句,我收到的一个反馈,恭喜发布。
What, what one piece of feedback that I got, by the way, congrats on the launch.
反响非常积极。我认为这可能是今年最重要的模型发布,每个人都说好话,但我给一位业内专家发短信问你怎么看,他说很棒、很惊人,但需要更多动态。所以至少在机器人领域是这样,但一般来说,拥有一个会动的世界似乎是理想的。所以也许谈谈这个,然后还有其他你愿意分享但认为值得讨论的未来方向吗?
It was overwhelmingly positive. I think it was probably the most significant model launch this year and you know everyone said glowing things but one person who's an expert in the space who I texted was like what do you think and great person is fantastic it's amazing but there needs to be more dynamics and so it would seem to be at least in the robotics case but generally it was kind of ideal to actually have a world that moves and so maybe talk a little bit about that and then any other future directions that you're comfortable sharing but you think are worth talking through.
是的,我的意思是动态显然会发生。实际上,嗯
Yeah, I mean like dynamics is clearly going to happen. Like actually um
我们有婴儿动态。
We have baby dynamics.
我们实际上已经有婴儿动态了。我认为人们没有完全意识到这一点。我们在博客文章中并没有真正强调,但之前的 Marble 世界模型基本上是静态的。是的。模型根本无法处理任何动态。这已经融入了模型架构和训练中。整个东西从根本上就是静态的。
We actually do have baby dynamics already. And this is something I think people didn't quite appreciate. We didn't really highlight in the blog post, but like the previous Marble World model, it was like fundamentally static. Yeah. Like the the model just like could not handle any dynamics at all. And that was just like baked into the model architecture, baked into the training. Like the whole thing was fundamentally static.
嗯,我们早就知道这是 Marvel 之后的一个大问题,我们已经在 Atlas 中修复了它,对吧?Atlas 架构已经从根本上支持动态。Atlas 的训练数据从根本上包含动态。如果你仔细看我们发布的一些视频,它实际上有一点移动。
Um we already knew that that was a big problem post Marvel and we already fixed it in Atlas, right? Like the Atlas architecture is already fundamentally supports dynamics. Um the Atlas training data fundamentally has dynamics. And if you look carefully in some of the videos that we even posted, it actually is moving a little bit.
是的。波浪,水波。
Yeah. The waves, the water waves.
是的。所以像一些例子中,水中有波浪。在一些生成的空中视图中,有小汽车在移动。所以动态实际上已经在这个模型中了。
Yeah. So like some of the examples, there's like waves in the water. Like in some of the like air generated aerial views, there's like little cars moving around. So like dynamics is actually already in this model.
但顺便说一句,如果你想从多个视图重建 3D,动态对我来说似乎非常有问题,对吧?所以这些东西是矛盾的吗?
But by the way, dynamics seems very problematic to me if you're trying to reconstruct 3D from multiple views, right? And so like are these things like at odds or
不。所以实际上我们的一个论点是这样的:如果你要做基本的 3D 重建,你实际上希望没有动态,你希望能够以精确冻结的时间建模场景的精确视图。是的。嗯,但这是我们之前 Marble 方法的一个问题,对吧?你可以尝试找到完全静态的数据,但这很难扩展,也很难获得更多。我们意识到的是,即使我最终想要静态输出,获得它的最好方法实际上是让模型接触动态,对吧?让模型接触尽可能多的动态内容,尽可能多的静态内容,让模型自己弄清楚如何分解动态内容。所以特别是在代码中,这实际上是,在 Atlas 预训练中,它已经在预训练中看到了大量动态,然后我们针对这个版本的这个检查点进行的后训练更侧重于静态内容,更侧重于空间运动,而不是时间,但我们已经有了,我相当确定预训练检查点已经包含了很多潜在动态,这是我们未来要大力改进的地方。
No. So actually one of our thesis here is that like you know if you're going to do fundamental 3D reconstruction you actually want to have no dynamics like you want to be able to model like exact views of the scene with exact frozen time. Yeah. Um but then like this is actually kind of a problem with our previous marble approach, right? Like there like you can try to find data that's fully static, but that's really hard to scale and really hard to get more of. And the thing we realized is that even in the case where I want static output in the end, the best way to get it is actually expose the model to dynamics, right? Like expose the model to as much dynamic stuff as you got, as much static stuff as you got, and let the model figure out how to factor out the dynamic stuff. So in especially like in the code and this is like and this is actually so in the Atlas pre-training already like it it saw a ton of dynamics in the pre-training already then the post training that we did specific to this checkpoint in this release was focused a lot more on on static stuff focused a lot more on spatial movement and less on temporal but like we already have we already like I'm pretty sure this the pre-train checkpoint already has a lot of latent dynamics in it um and this is something we're going to improve quite a lot going forward.
那么,Ben,这是否意味着我们会得到 40 个视频,比如到处走走?
So, so Ben, does that mean we're going to get 40 video like go walk around?
我的意思是,
I mean,
你可以看到他们脸上的笑容。
you can see the smile on their face.
所以,我实际上可以看到,如果你现在就停下来,只做更大、更快、更好,你几乎可以建立一个完整的行业,对吧?这感觉像一个非常横向的原语。然后如果你什么都不做,但还有其他不仅仅是更大、更快的事情让你对关注的应用感到兴奋吗?这些应用往往更偏向内容创意 3D 方面。
So, I I actually can see if like you just stopped now and you only did kind of bigger, you know, faster, better, you could build almost an entire industry, right? I It feels like a very horizontal primitive. And then and if you did nothing else, but are there other things that are not just kind of bigger, faster that you're excited about for the applications you're focused on, which tend to be kind of more on the kind of content creative 3D side?
是的,我真的很兴奋于推动多模态方面。我认为不同的控制模式在这里非常关键。我认为,特别是在学术界,对这些模型添加控制条件以提取内部内容的重要性被严重低估了。我的意思是,老实说,这种动态与……我不确定我是否理解那些
Yeah, I'm really excited about pushing that kind of multimodal aspect. I think different modes of control is so critical here. I think like it's super underappreciated, especially in the academic community, how critical it is to add control conditioning to these models to kind of get out what's inside. I mean, honestly, this dynamic versus I'm not sure I even understand what those
用外行的话说就是可编辑性。我认为可编辑性是这里的关键。
to translate in layman's language is editability. I think editability is is the key here.
是的。所以我的意思是,这是我们在单图像模型中看到的东西,从今年开始在视频模型中开始解锁,比如获得多轮交互的感觉,或者真正直观地解释我想要这个人、这个物体和这件事发生,并将它们组合在一起,而无需手动操作系统,它就像前沿图像模型在编辑方面已经做到了,但我们还没有看到这种能力强烈地传播到视频和世界模型中,对吧?我们看到了一些玩具示例,比如我可以输入一个句子,然后一只恐龙出现,这些实时模型。嗯,但我想把它提升到真正的工业强度,因为这里的诀窍是你必须添加控制,但不能损害模型的质量,否则它基本上就变成了一个噱头。就像没有人会认真考虑用你的模型替换他们前沿视频模型的使用,如果你给他们额外的旋钮,但质量下降。所以,我认为这真的是一个游戏,我们如何保持我们设定的高标准,即当前模型能够获得的输出,然后添加人们会要求的所有有趣的东西,比如我想与场景交互,控制布局,控制对象的身份,或者控制时间,嗯,我认为这是一个途径,它开启了大量有趣的产品和界面工作,你添加的复杂性越多,丰富度越高,使你能够真正思考几乎从零开始重新设计人们与计算机中有状态的 3D 世界交互的方式。这确实是最终目标,获得构建这种系统所需的所有能力。
Yeah. So I mean this is something we've seen in like sort of single image models and starting this year in video models is starting to be unlocked in terms of oh like getting that flavor of like multi-turn or really like intuitively interpreting like I want like this person and this object and this thing to happen and kind of combining those all together in like one paste without having to do a lot of like manual work with the system like it just interprets it like kind of frontier image models are kind of there right for in terms of editing but we haven't seen that propagate out as strongly into video yet and and into world models, right? We've seen some really kind of toy examples of, oh, I can like put in a sentence and like, you know, a dinosaur appears or something with these like sort of real-time models. Um, but I want to like turn that up to really industrial strength and make that cuz like the trick here is you got to add control but not compromise the quality of the model or it just becomes a party trick basically. Like it's like no one is going to seriously think about swapping their like cutting edge frontier video model usage for your model if you give them extra knobs, but the quality degrades. So, I think it's really that game of like how can we maintain like the high bar we've set was the outputs we're able to get in the current model and then add all kinds of interesting stuff that people will ask us for in terms of like I want to interact with the scene or control the layout or control like the identity of the objects and the things that we're seeing within there or control time right um and I think that's like an access where it opens up like a ton of really interesting product and interface work the more complexity you add there and richness in terms of kind of like enabling you to really think about like redesigning almost from scratch the way people interact with sort of like stateful, you know, 3D worlds in the computer. Like that's that's really the end goal here is getting like all the capabilities you need to build that kind of system.
太棒了。你有什么要补充的吗?关于你感到兴奋的新功能,不仅仅是更大或更好。
Awesome. Anything you'd add to that as far as new functionality that you'd be excited about that's not just big or better.
对我来说,让我们回到智能的第一性原理。智能不是坐在那里,只是看到或解释空间和物理空间中的东西,对吧?它实际上是视觉、体验和交互之间的闭环。所以考虑沿着这个阶梯向上,正是 Ben 所说的。
I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stack and just seeing something or interpreting something when it comes to space and physical space, right? It's really this uh closing the loop between seeing and experiencing and interaction. So just thinking about going up that ladder is exactly what uh Ben said.
酷。
Cool.
我觉得这里有一个有趣的概念,就是 AI 完备性。你以前听说过这个吗?
I think one interesting notion there is this notion of AI completeness. Have you heard this before?
是的,是的,我听说过。
Yeah. Yeah, I have. Yeah.
所以就像每个人……
So like everyone...
顺便说一下,我听说过关于大语言模型的 AI 完备性,就是说你必须基本上是最聪明的语言模型才能回答最聪明的 LLM 需要回答的问题,或者你必须解决通用智能。
I hear about AI complete by the way in terms of LLMs which is like you have to be basically you know like the smartest LM to answer the question what the smartest LLM will need to answer or you have to solve general intelligence.
不,不,它基本上是与图灵完备性相关,对吧?就像这个想法是,如果一个任务是图灵完备的,就像在经典复杂性理论中,如果我可以把这类中的任何问题归约为那一个问题,对吧?比如三元组例子,对吧?所以你可以把任何 NP 难问题归约为三元组,因此你可以用三元组来解决任何问题。
No, no, it's basically it's a it's a connection to Turing completeness, right? Like the idea being that if a task is Turing complete like in classical complexity theory if I can take any class any problem in this category reduce to that one problem, right? Three example, right? So you can take any NP-hard problem then reduce it to three therefore you can use threes to solve any any problem.
是的,是的。
Yeah. It's Yeah.
所以,那么,AI 完备性的软定义就像是,有一个基本的原语,它是一个 AI 任务,但如果我能以最广泛的通用性解决这个 AI 任务,它就能解决任何智能问题。比如 LLM 的经典例子是下一个词预测是 AI 完备的,因为我可以,你知道,我认为来自伊利亚的经典例子是,有一本悬疑小说,它必须读完整本悬疑小说,而小说的最后一句话是“凶手是”,预测下一个词。所以你可以基本上把任何智能任务都框定成那样。所以显然,下一个词预测是人们认为是 AI 完备的。
So, so then like the kind of like the soft definition of AI completeness is like there's this fundamental primitive that's it's an AI task but if I could solve this AI task in its full broadest generality it would solve any intelligence problem and like the classic example at LLM is like next token prediction is AI complete because I could like you know there's the classic example I think from Ilia where like I there's a mystery novel and like the thing has to read the whole mystery novel and the final sentence of the mystery novel is like and the killer was predict the next token. So like you could basically like frame any kind of intelligence task in terms of that. So clearly next token prediction is something that people believe is AI complete.
但我认为我们正在意识到的事情,本今天早些时候也谈到过,就是新视角预测,我们在 Atlas 中拥有的这个原语,尤其是生成式新视角预测。这也是 AI 完备的,对吧?因为我可以拿一些东西,比如……
But I think that something we're kind of realizing and Ben was talking about this earlier today is like new view prediction this primitive that we have in Atlas especially generative new new view new view prediction. This is also AI complete right and because I could take something like...
你可以有一部电影,你处理电影的所有帧,然后凶手走出来,你预测到底是谁走出来。
You could you could have the the movie and you do all of the frames of the movie and then like the killer walks out and then you predict exactly who walks out.
完全正确,完全正确。不仅如此,可以说我想要一个世界,比如马丁在背面写黎曼猜想的证明,镜头移到下一个白板就解决了。所以从进化的角度来看,新视角预测正是进化通过让动物移动来解决的问题。
Exactly. Exactly. Not just that could say like I want I want to have a world where like Martin is like writing a proof of the Riemann hypothesis on the back and the camera over to the next white solves the okay so to take a evolutionary view right that new viewpoint prediction is exactly evolution had to solve by making animals move.
你,你,大自然给了动物眼睛。
You you nature give animals eyes.
但大自然没有给树这些眼睛。
But nature didn't give trees these eyes.
眼睛,为什么?因为当你移动时,你会看到一个新的视角。
Eyes why because when you move you see a new viewpoint.
这就是,无论你称之为 AI 完备还是智能完备,所以我们确实非常坚信,下一个视角预测等同于下一个词预测。
And that is the the the whether you call it AI complete or intelligence complete so so we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction.
太棒了,那么,祝贺你们所有人成功发布了一个非凡的模型,我们对未来的模型发布感到非常兴奋,感谢你们的到来。
Amazing well with that congratulations all of you on a phenomenal model launch we're very excited for future model launches and thanks for coming so.