Visual Intelligence Comes from Language Models
打开互动全文版(中英对照 + 朗读 + 问答)→一位研究者认为视频模型的进步主要来自语言模型而非视频架构本身,并分享了他在 xAI 构建视频模型的经验。
A researcher argues that improvements in video models are driven by language models, not video-specific architectures, and shares his experience building video models at xAI.
我有一个相当大胆的主张:视觉智能实际上主要来自语言。这些视频模型,尤其是现在扩散模型技术更加成熟,每次你看到这些模型有所改进时,我认为大部分收益来自语言模型,而不是视频模型本身。
I have a pretty big claim: visual intelligence actually mostly comes from language. These video models, especially now that diffusion model technology is more mature, every time you see some improvement on these models, I would say most of the gain comes from the language model, not from the video model itself.
在进入今天的正题之前,我有个小消息给听众。谢谢。如果没有你们选择点击并收听我们的内容,我们就不可能为你们带来如此渴望的 AI 工程、科学和娱乐内容。几乎每天都有赞助商找上门。但幸运的是,有足够多的订阅者让我们在没有广告的情况下维持运营,我们希望保持这样。但我只想请大家帮一个忙。你们能做的最有力、完全免费的事情就是点击订阅按钮。这是我唯一会请求你们做的事,这对我以及每周努力为大家带来 Inspace 的团队来说意义重大。如果你们订阅了,我保证我们会继续努力让节目变得更好。现在,让我们开始吧。
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you, we'll never stop working to make the show even better. Now, let's get into it.
好的,今天我们请到了 Ethan He,他最近在 xAI 工作。欢迎。
Okay, we're here in the studio with Ethan He, most recently of xAI. Welcome.
谢谢,很高兴来到这里。
Yeah, thank you. Glad to be here.
还有 Vivu。你最初加入 Latent Space 是因为你在 Nvidia 做 Cosmos 项目,并写了一篇很棒的论文。我们很喜欢。你也做了展示。所以谢谢你。
We're also here with Vivu. You first came to us or joined the Latent Space world because you were working on Cosmos at Nvidia and you did a great paper. We loved it. You presented it as well. So thank you for doing that.
是的。我还展示了 MoE。
Yep. I also presented the MoEs.
对。在 Latent Space 展示了两次。
Yes. Twice at Latent Space.
你是怎么听说我们的?是我们联系你的吗?
Yeah. How did you actually hear about us? Did we reach out to you?
不,实际上是我发现了这个社区。我意识到,哦,有个在线社区,人们每周通过论文俱乐部讨论 AI 并互相学习。非常好。我学到了很多。
No, actually I found the community. I realized, oh, there's this online community where people talk about AI and also learn from each other through papers every week through the paper club. It's very nice. I learned a lot.
我觉得已经连续三年了。即使在圣诞节和新年我们也没停过。很多周我都想停下来。
I think three years non-stop. We haven't stopped even on Christmas and New Year's. Many weeks I want to stop.
不好。我记得你发帖说你写了一篇论文,我当时觉得,哦,很酷。我们有论文俱乐部展示,但可能之后我联系了你。
No good. I think you had posted that you worked on a paper and I was like, oh, very cool. We have paper club presented but I might have reached out to you after.
是的,因为这是个业余俱乐部,对吧?
Yeah, because it's an amateur club, right?
所以很不寻常,但有时论文作者会来亲自讲解论文。今天我们刚做了 Poolside 的论文,显然非常好。昨天发布的。挺有意思的,对吧?完全开放。他们讨论了整个系统。所以是一篇好论文。我们会推荐大家阅读。
So it's very unusual, but we have sometimes paper authors come by and actually explain the paper. Today we just did the Poolside paper, which is apparently very good. Came out yesterday. Pretty interesting, right? Fully open. They talk about everything system. So it's a good one. We'll recommend people to read it.
跟我们说说你转到 xAI 的情况吧,因为我甚至不知道你是什么时候加入的。讲讲这个转变的故事。
Bring us up to speed on your transition to xAI because I actually don't even know when you joined. Just tell us the story about the sort of transition.
在 xAI 之前,我在 Nvidia 从事 Cosmos 世界模型的工作。Cosmos 是一个巨大的视频基础模型,旨在模拟世界,并作为所有机器人专家构建上层应用的基础。我构建了 Cosmos 后,意识到这个东西也有类似于语言模型的缩放定律。我们需要进一步扩大视频模型的规模。这就是为什么我意识到需要去一个拥有更多算力资源的地方。这就是我最终来到 xAI 的原因。
Before xAI, I was working on the Cosmos world model at Nvidia. Cosmos is a giant video foundation model that aims to simulate the world and serves as a foundation for all roboticists to build on top of. Once I built Cosmos, I realized this thing also has a scaling law similar to language models. We need to scale up the video models further. That's why I realized I need to move somewhere with much more compute resources. That's how I ended up at xAI.
而 Nvidia 自己就有 GPU。
And Nvidia had GPUs themselves.
是的。
Yeah.
从时间线上看,Cosmos 是什么时候?挺早的,对吧?就是那篇开放世界模型论文。
Timeline wise, when was Cosmos? It was pretty early, right? It was the open world model paper.
大概是 2024 年底。
It was like end of 2024.
2024 年底。
End of 2024.
然后在 2025 年中,我搬到了 xAI。那时,我加入时 xAI 正准备构建视频模型和多模态模型。没有基础设施,没有数据,也没有模型。只有几个工程师。我们在三个月内构建了它,并发布了第一个模型 Grok Imagine 0.9。从那以后,我继续从事视频模型的工作,并从预训练转向后训练,例如参考视频的 Cameo 功能和视频扩展。在我离开之前,我从事世界模型的工作,领导一个小团队专注于实时长程视频生成。
Then at mid 2025 I moved to xAI. At that time, I joined when xAI was about to build video models and multimodal models. There was no infra, no data, and no model. Just a few engineers. We built it in three months and released the first model, Grok Imagine 0.9. Since then, I kept working on video models and moved more from pre-training to post-training of the video models, for example, reference to videos like the Cameo feature and video extensions. Before I left, I worked on world models, leading a small team to focus on real-time long-horizon video generation.
你能给个大致的路线图吗?比如,你加入了一个全新的团队。Grok 之前只有文本,或者他们与 BFL 合作做图像生成。构建模块是什么?你有算力,数据可以从某处获取。人们在组建新团队时应该考虑哪些步骤顺序?
Can you give a rough roadmap of like, okay, you're on a brand new team. Grok previously was only text or they partnered with BFL for their image gen stuff. What are the building blocks? You have compute, data you can procure somewhere. What are the sequence of things that people should think about when setting up a new team?
实际上甚至更深,不仅仅是你可以获取的数据。你们也得经历获取数据的过程,对吧?
Actually even deeper, not just data you can procure. You guys had to go through getting the data too, right?
你们发货很快,但确实。
You shipped it pretty fast, but yeah.
是的,三个月实际上快得惊人。
Yeah, three months is actually surprisingly fast.
我想说的一点是,感谢我在 Nvidia 的经验,因为第一次我们一起构建 Cosmos 时,花了大约一年时间。所以这是我第二次做这件事,大致知道该怎么做。最重要的是人才。每个人都非常强大和聪明,彼此之间非常紧密,朝着共同的目标努力。这大大加快了速度。你减少了人们之间的沟通带宽,每个人都可以朝着同一个目标努力。每天日程上没有那么多会议,可能每天一次同步会,之后就是全情投入构建。那时非常有趣。另一件事是 xAI 在数据、推理和支持基础设施方面有非常强大的基础,这对模型开发帮助很大。当我审视模型训练时,最重要的事情是你每天能做多少次迭代。你能做的迭代越多,你就能越快训练模型。所以如果你有非常强大的基础设施和大量算力,你可以在很短的时间内训练这些模型,这为你提供了更大的错误缓冲,也让你有机会发现更多错误。
One thing I say is thanks to my experience at Nvidia, because the first time when we were building Cosmos together, we built it for about a year. So this is the second time I do it, roughly have an idea what to do. The most important thing is talent. Everyone was very strong and clever, very close with each other towards a common goal. That speed up things a lot. You reduce the communication bandwidth among people and everyone can work toward the same goal. Every day there's not that many meetings on the calendar, maybe like a sync a day, and after that it's just all building. It was pretty fun at that time. Another thing is that xAI has very strong foundations for data, inference, and supporting infrastructure that helps model development a lot. When I look at training models, the top important thing is how many iterations you can do per day. The more iterations you can do, the faster you can train the model. So if you have very strong infra and a lot of compute, you can train these models in a very short period of time, giving you a much larger buffer for errors and also the opportunity to spot more bugs.
什么是迭代?是几百步之类的吗?
What is an iteration? Is it like a few hundred steps or what?
比如说,从获取新数据开始,可能设计新算法,然后训练一个新模型,也许是较小规模的。
Let's say just training the model from acquiring new data and maybe designing new algorithms and training a new model, maybe at a smaller scale.
所以是任何超参数的周期时间……
So cycle time for any hyperparam that you're...
是的,端到端的周期时间,评估这个模型是否比之前的迭代更好。
Yeah, cycle time end to end to eval this model, is this model better than my previous iteration.
所以之前,有人已经设置好了,你可以非常快速地迭代。
So before, someone had already set this up that you can iterate very quickly.
是的。
Yeah.
我认为那里的基础对于开发和研发模型来说非常好。我经常觉得这有点无聊,但很多改进并非来自新算法,而是来自在数据管道和模型训练管道中到处发现的小 bug。这些对模型质量的提升最大。
I think the foundation there is extremely good for developing and research models. Often, I find this is kind of boring, but a lot of the improvements do not come from new algorithms. They come from finding small bugs here and there in the data pipeline, in the model training pipeline. Those gave the biggest boost to the model quality.
这很有趣,对吧?你说团队小、沟通带宽少,但很多质量提升却来自找小 bug。这似乎反直觉,对吧?人多了可以解决更多问题,但看到另一面也很有意思。
It's interesting, right? So you say it's a small team, less communication bandwidth, but also a lot of quality is like finding little bugs. It seems counterintuitive, right? You have a lot of people, you can iron out more of those, but it's interesting to see the other side.
是的。是的。我也好奇,你有没有试过用大语言模型来找 bug?我不知道。我记得那时是 2025 年中,编程模型还没那么成熟。我记得 2025 年 12 月的时候它已经非常好了。
Yeah. Yeah. I also wonder, have you tried using LLMs to look for bugs? I don't know. I remember at that time it was mid-2025, so the coding model wasn't quite there yet. I remember like December 2025 it was extremely good.
是的,我当时就在用。有时候很有帮助。它生成的代码有点难以维护,尽管第一次构建东西非常快,但它会给出几千行的意大利面条式代码,我无法维护,而且大语言模型自己也搞不清哪里有问题以及如何在此基础上改进。但现在我觉得它好多了,好太多了。
Yeah, I've been using it at that time. It's helpful sometimes. It produced code that is kind of difficult to maintain, even though the first time it builds something extremely fast, but it gave like a spaghetti code thousands of lines that I couldn't maintain, and the LLM itself couldn't figure out what's wrong and how to improve on top of it. But now I find it much, much, much better.
是的,我想再提一点。现在编程模型高效得多,能帮我们更快地实现东西。算力可能再次成为瓶颈,因为以前如果你想训练一个新模型,比如生成新的合成数据然后写一个新算法,可能需要几周时间。在那段时间里,你可能没有实验可跑。但现在你可以在几小时内构建好,然后立即训练模型。现在你必须要有足够的算力来尝试所有想法。所以算力可能再次成为迭代速度的瓶颈。
Yeah, I want to bring up another point here. Now coding models are much more efficient and can help us implement stuff much faster. Compute might become a bottleneck again because previously, if you want to train a new model, say you want to generate new synthetic data and then write a new algorithm, it might take a few weeks. During that period, you might not have experiments to run. But now you can build that thing within a few hours, then you can immediately train a model. Now you have to have enough compute to try all of the ideas. So compute might be the bottleneck of iterating speed again.
是的。老实说,我觉得这是一份压力很大的工作,因为你会想,我应该尝试所有东西,如果没做到,我就没做好工作。我的意思是,还有每小时消耗数千块 GPU 的压力,这非常昂贵,而且你知道算力可能会下降。
Yeah. I actually honestly think it's kind of a stressful job because you're like, well, I should be trying everything, and if I'm not, then I'm not doing my job well. I mean, there's also the stress of you're eating thousands of GPUs per hour, which is very expensive, and you know compute can go down.
是的,但你知道,算力仍然是有限的。你想好好利用它,你想要更多。
Yeah, but you know, there's still finite amount of compute. You want to use it well, you want more of it.
那确实压力很大。是的,我认为有一点是,现在有了编程模型,很多这些工作可以自动化,这好多了。第二,这是一场马拉松,所以你得保持健康和有规律的作息。
That was quite stressful indeed. Yeah, I think one thing is with coding models now, a lot of these jobs can be automated, which is much better. Second, it's a marathon, so you got to maintain good health and a regular schedule.
当你从零到无在 3 个月内转变时,很难听进去这些。
It's hard to hear that when you shift from zero to nothing in 3 months.
是的。我的意思是,显然这种文化非常有名,你知道,人们工作非常努力。有一件事我确实想深入探讨,在你提前发送的笔记中,你对视频生成训练的成本有具体评论。大概是在 Colossus one 上,对吧?那个 300 兆瓦的集群。你想分享什么都可以。
Yeah. I mean, and I think obviously the culture actually has very famously, you know, people work very hard. One thing I did want to dive into, in the notes that you sent ahead of time, you had specific comments about the cost of video gen training. Presumably this is on like Colossus one, right? The 300 megawatt cluster. Whatever you want to share.
我觉得你在说三件事,对吧?有视频生成,还有你发布的图像生成模型。你想完整地……好吧,从零到一,你有几个月时间。创建图像有哪些阶段?
I think there's three things you're talking about, right? So there's video gen, there's also the image gen model that you put out. Do you want to like complete the... Okay, so zero to one, you have a few months. Just what are the stages of creating an image?
哦,是的,可能我分心了。抱歉。然后从那里开始有视频生成、音频生成。接下来很想聊聊这些。但最初几个月是怎样的?小团队,很多 bug,迭代。但具体是什么样子?我们是拿现成的东西吗?我们只是获取数据、算力吗?那几个月是怎样的?你怎么做到最先进的图像生成模型?你怎么开始?
Oh yeah, maybe I got distracted. Sorry. And then from there there's video gen, there's audio gen. Love to get into those things next. But what is that first few months like? So small team, a lot of bugs, iterations. But like, what does it look like? Do we take something off the shelf? Do we just get data, compute? What's the few months like? How do you go to state-of-the-art image gen model? How do you just start?
具体说说我是怎么做的,但这是一个相当标准的过程。我可以从 Cosmos 中举一些例子。所以主要是构建视频模型,你实际上需要先构建一个图像模型。而构建这两个模型,你需要的数据是 100% 合成的语言与图像或语言到视频的配对,因为在互联网上,视频并不自然地与文本关联。所以你可以说,哦,就像在 YouTube 上,你有标题、描述和评论,但它们通常与视频本身不相关。比如说视频是山区的自然场景,而标题是“我今天很开心”。所以它们完全没有相关性。所以第一步是,你必须生成语言与视频的合成配对。所以从互联网上获取视频,然后使用视觉语言模型(VLM)来为视频生成字幕。
Echo comments specifically how I did, but it's a quite standard process. I can draw some examples from Cosmos. So mainly it's like building a video model, you actually need to build an image model first. And building these two models, the data you need is 100% synthetic pairs of language and image or language to video, because on the internet, the videos don't naturally associate with text. So you can say, oh, like on YouTube, you have the title and the description and the comments of a video, but usually they're not relevant to the video itself. Say maybe the video is a natural scene of mountains or something, and the title is like "I'm so happy today." So they have no correlation at all. So the first step is you have to generate synthetic pairs of language with the videos. So you get the videos from the internet and you use a VLM to caption the videos.
这里有个问题:你最初怎么得到 VLM?所以如果没有……
Here's a question: how do you get the VLM to begin with? So if there's no...
你使用模型,对吧?比如……
You use the model, right? Like...
假设没有 VLM 存在,你怎么开始生成文本?对吧,那是不可能的。
Say if there's no VLM exists, like how do you generate the text to begin with? Right, it's impossible.
我明白了。
I see.
一开始,就像你让人尽可能详细地描述视频。例如,你让他们描述一切:所有物体、所有角色、以及视频中的所有互动和对话。在 Cosmos 标注的协议中,我们给标注者的目标是,你必须尽可能详细地描述视频,这样盲人拿到一段文本后,可以在脑海中重建视频的样子。
In the beginning, it's like you ask humans to describe the video as detailed as possible. For example, you ask them to describe everything: all objects, all characters, and all interactions and dialogues in the videos. In the protocol of Cosmos labeling, we required the objective we gave to the labelers was that you have to describe the video as detailed as possible such that a blind person, given a blob of text, can reconstruct what the video is like from their head.
视频还是图像?你说的是……
Video or image? You're talking about...
视频或图像。都可以。这在我们从 CLIP 和 DALL·E 过渡时很常见,对吧?都是训练在非常详细的图像字幕上。同样适用于视频,但不用多模态模型传入视频或图像并写出丰富的描述,你也可以……
Video or image. Either one. This was pretty common when we went from CLIP and DALL·E, right? It's all training on really detailed captioning of images. So same is applied to video, but instead of using a multimodal model to pass in video or images and write rich descriptions, you can also...
我的意思是,我认为有这种传统的监督视角,或者你知道,高度人工创造的东西。我觉得无监督有突破,对吧?当你有足够的东西可以自举时,你可以直接扔一个通用语料库进去,或者你知道,随便什么,比如无监督的视觉和语言配对,对吧?你只是把图像和文本交错放置,它就能学习。对我来说,这就是 VLM 的突破,不同于 CLIP,不同于大语言模型之前的时代。
I mean, I think there's this traditional perspective of supervised, or you know, very highly human-created thing. I feel like there's an unlock with unsupervised, right? Where you have enough to bootstrap that you can just throw a common corpus on it, or you know, whatever, like unsupervised vision and language pairing, right? Where you just have interspersed image and text and it just learns. To me, that is the VLM breakthrough that is different from CLIP, different from the pre-LLM era.
是的,有趣的是你两种数据都需要。
Yeah, it's interesting to see that you kind of need both data.
例如,为了让你自举起来。是的。
For example, for you to bootstrap it up. Yeah.
是的。对于生成模型训练,通常也有一小部分未标注数据。所以模型被指示生成没有任何文本指令的视频。这也能帮助模型泛化。所以在生成合成配对这个阶段之后……
Yeah. For the generative model training, there's also usually a small percentage of unlabeled data. So the model is instructed to generate a video without any text instruction. That can also help the model generalize. So after this stage of generating synthetic pairs...
一个重要的常见步骤是训练图像或视频的压缩器或分词器。因为从技术上讲,你可以在纯像素上训练图像或视频模型,但问题是 token 数量太多。例如,一张 1000x1000 的图像就是 100 万个像素——不可能用 Transformer 来训练。所以你需要训练一个分词器,将图像映射到潜在空间,再映射回来。
One important common step is to train a compressor or a tokenizer of images or videos. Because technically you could train image or video models on pure pixels, but the problem is that it's a lot of tokens. For example, one image that's a thousand by a thousand is 1 million pixels — it's impossible to train a transformer on that. So you need to train a tokenizer that can go from image to latent space and back.
这就是我们播客名字的由来。
That's why we named the podcast.
没错。
Exactly.
但基本上你在说词汇量大小。
But basically you're talking about vocabulary size.
在生成模型中,词汇是连续的——它是一个连续空间。你可以想象将图像映射到一个向量,一个固定长度的向量,可能是 16 或 48,然后再将该向量映射回图像空间。这种映射是基于 patch 的。例如,你有一个 16x16 的 patch,并将该 patch 的像素映射到潜在空间。
In generative models, the vocab is continuous — it's a continuous space. You can think of mapping an image to a vector, a fixed-length vector of maybe 16 or 48, and then mapping that vector back to image space. The mapping is patch-based. For example, you have a 16x16 patch and map that patch of pixels into the latent space.
我们之前聊过——视觉 VAE。你基本上压缩输入,在更小的维度上进行生成和推理,然后再投影回去。
We've covered this — vision VAE. You basically compress your input, do your generation and reasoning in a smaller dimension, then project back out.
VAE 是一种压缩形式,但 patch 的概念来自 ViT,对吧?那篇论文的标题大概是“16x16 就够了”。
VAE is a form of compression, but the patching thing comes from ViT, right? The paper is titled something like "16x16 is all you need."
而且我觉得人们经常把这种 patch 和卷积做比较——你是在用新范式重构旧范式。
And I think people make a lot of comparisons between this patching and convolutions — you're kind of reconstructing the old paradigm with the new.
是的。实际上在 VAE 中,既有卷积网络也有 Transformer——两者都可以用。
Yes. Actually in VAE, there are both convolutional networks and transformers — you can do both.
在这之后,你有了潜在空间 token 和语言 token。然后训练扩散 Transformer 就相当标准了——和训练语言 Transformer 模型非常相似。唯一的区别是去噪过程:你给视觉 token 添加随机噪声,然后训练模型去除这些噪声以生成干净的 token。在推理时,模型可以从 100% 噪声开始迭代去噪。
After this, you have latent space tokens and language tokens. Then training the diffusion transformer is quite standard — very similar to training language transformer models. The only difference is the denoising process: you add random noise to the visual tokens and train the model to remove that noise to generate clean tokens. During inference, the model can iteratively remove noise from 100% noise.
还有 CFG 和潜在扩散来加速。Stability 和其他人开创了很多这种架构。不知道你是想深入聊聊,还是直接转到视频。
And there's also CFG and latent diffusion to speed things up. Stability and others pioneered a lot of this architecture. I don't know if you want to get into that or move to video.
训练完这样的图像模型后,它就成了视频模型的基础,因为图像模型训练成本更低,而且语言和图像之间的连接更密集。例如,你在十亿张带有文本-图像映射的图像上训练。在十亿个文本-视频对上训练要昂贵得多,因为视频有更多 token。扩散模型对语言的理解完全来自这种映射,所以如果你只在 1000 万个视频上训练,可能看不到足够的语言 token——模型就无法很好地理解人类意图。这就是为什么你先训练图像扩散模型,然后从中引导出视频模型。
After training such an image model, it's a foundation for video models because image models are cheaper to train and have much denser connections between language and images. For example, you train on a billion images with text-image mappings. Training on a billion text-video pairs is much more expensive because videos have more tokens. The diffusion model's understanding of language comes purely from this mapping, so if you only train on 10 million videos, you might not see enough language tokens — the model won't understand human intention well. That's why you first train image diffusion models, then bootstrap the video model from there.
我想问一件事——你是我聊过的第一个做视频模型的人。我们和 Luma 以及其他人都聊过。视频压缩中有一些技巧,帧与帧之间差异不大,所以你不必重新生成每一帧。MP4 压缩就是这样做的。人们会倾向于用这个吗,还是每个人都直接生成每一帧?
One thing I wanted to ask — you're the first video model person I've talked to. We've talked to Luma and others. There are tricks in video compression where frame by frame there's not much difference, so you don't have to regenerate every frame. MP4 compression does that. Is it tempting to use that, or does everyone just generate every frame?
有几种方法。有些人尝试过直接使用 MP4 压缩作为 Transformer 的 token,但主要挑战是 MP4 token 的潜在空间对模型来说不太容易理解——训练起来极其困难。这就是为什么我们创建了 VAE,它创建了一个更连续的潜在空间,模型可以更容易地理解和学习。即使在 VAE 内部,也有不同的难度。最简单朴素的方法是将所有图像像素打乱成一个向量,而不训练任何 VAE,但那个空间对模型来说极其难以训练。所以关于如何压缩 token 存在争议。你可以逐帧压缩,也可以压缩时间维度。
There are a few approaches. Some have tried directly using MP4 compression as tokens for transformers, but the main challenge is that the latent space for MP4 tokens is not very comprehensible for models — it's extremely hard to train on. That's why we created VAE, which creates a more continuous latent space that models can understand and learn from much easier. Even within VAE, there are different difficulties. The simplest naive way is to shuffle all image pixels into a vector without training any VAE, but that space is extremely hard for models to train on. So there's debate on how to compress tokens. You can compress frame by frame, or compress the temporal dimension.
是的。区别在于,如果你压缩时间维度,由于帧间的时间冗余——这一帧和上一帧大部分相似,只有微小差异——你可以获得更高的压缩率。例如,在 One 2.1 中,他们有 8x8x4 的压缩率,所以四个时间 token 被压缩成一个,节省了大量上下文长度。如果逐帧压缩,你可能只有 8x8x1,上下文长度会大四倍。然而,逐帧压缩的好处是实时交互性——如果你逐帧流式输出,模型可以立即响应用户请求。而时间压缩则会有延迟。
Yes. The difference is if you compress the temporal dimension, you get much higher compression because of temporal redundancy between frames — this frame and the last are mostly similar, with only small differences. For example, in One 2.1 they have an 8x8x4 compression rate, so four temporal tokens are compressed into one, saving a lot of context length. If you do it frame by frame, you might have 8x8x1, making the context length four times larger. However, the benefit of frame-by-frame compression is real-time interactivity — if you stream the output frame by frame, the model can respond to user requests immediately. With temporal compression, there's a lag.
所以你对此很纠结。我们提一下——实时视频生成有一些前沿应用。
So you're very pickled on this. Let's bring it up — there are frontier applications of real-time video generation.
所以 Flipbook 是最近走红的例子之一。对吧。什么是 Flipbook?
So Flipbook is one of the examples that went viral recently. Right. What is Flipbook?
Flipbook 有点像网页浏览器。你可以看到它顶部有浏览器 UI。区别在于所有 UI 都是由生成式图像模型实时生成的,这里的一切都是假的。但你可以在这个想象的世界里探索。比如这里我们有一个网格金字塔的工程图——模型生成这个是为了让我们理解它是如何工作的。如果我们想进一步导航和理解,可以点击这里的某些描述,模型就会生成一个新页面、一个新子页面来描述我们想了解的细节。所以这基本上就像我们在播放一个视频,但它在等待我们的下一次交互,然后根据我们的交互播放下一段内容。这挺酷的。你有点像是在决定自己的故事。所以这是——你知道,如何用杠杆技术建造金字塔看起来很有趣,对吧?它展示了如何——好吧,我想知道……
Flipbook is kind of like a web browser. You can see it has the web browser UI on top. The difference is all of the UIs are generated by a generative image model in real time and anything here is fake. But you can explore inside this imaginary world. Say here we have the engineering of the grid pyramid — the model generated this for us to understand how it works. If we want to navigate around and understand further, we can click on some of the descriptions here and the model will generate a new page, a new sub-page describing the details we want to know about. So it's basically like we are playing a video, but it's pausing for our next interaction and then it just plays the next thing based on our interaction. Which is kind of cool. You kind of decide your story. So this was — you know, how do you make a pyramid leveraging technique seemed interesting, right? It shows how do you take — okay, I want to know...
演示推文在帧之间有更多动画。我觉得它只是跳过了很多帧。他们也有视频模式,但我想很多人都在用。有一个实时视频流。我们可以试试……
The demo tweet had more animation between frames. I think it's just skipping a lot of frames. They also have a video mode but I guess a lot of people are using it. There's a live video stream. We can try...
所以这是一个极端的未来例子。我们今天显然还没到那一步,但在一个推理完全免费的世界里。这比生成文本代码更好。
So this is an example of the kind of future that you see at the extreme. We're obviously not in it today, but in a world where inference is completely free. This is better than generating code in text.
是的。所以这是 VR 的最终状态,世界模型的氛围,我想。想象一下互联网不存在,然后你输入 google.com——模型应该给你展示什么?模型可以想象一些东西,这就是它想象出来的。这些网页完全不存在。所以我认为随着推理成本下降,我们将拥有生成式 UI 来覆盖一切。如果你想想编码模型是如何工作的,它们为网页编写代码并渲染它。代码可能被转换成二进制,二进制再在屏幕上渲染像素。所以在机器学习中,每次我们有突破,显然都更高效。那么为什么我们不能直接让用户指令到像素呢?生成的 UI 将直接把用户意图转化为像素。即使我想要邮件——假设每个人都有相同的界面,但我想要稍微不同。我希望邮件像 TikTok 一样展示给我,这样我可以左右滑动邮件,或者你可能想要别的。我们可以有完全不同的东西。或者像我在看 Instagram 故事——我不喜欢点赞按钮,我总是误点——然后生成没有它的 UI。所以这将是对界面的革命性替代。在未来,我们可能在幕后运行更强大的 LLM 和编码模型,而在前端,扩散模型实际上将成为展示内容的前端。
Yeah. So this is the final state of VR, the vibe of a world model, I think. Imagine the internet doesn't exist and then you type in google.com — what should a model show you? The model can imagine something, and this is what the model imagines. These web pages completely do not exist. So I think as inference cost comes down, we are going to have generative UI for everything. If you think about how coding models work, they write code for a web page and they render it. The code might be converted into binary, and the binary renders the pixels on the screen. So in machine learning, every time we have some breakthrough, obviously it's more efficient. So why don't we have a user instruction to the pixel directly? The generated UI will be user intention to the pixels directly. Even if I want email — let's say everyone has the same interface, but I want it slightly different. I want the email to show to me like a TikTok so I can swipe left and right for the emails, or maybe you want something else. We can have completely different things. Or like I'm looking at Instagram stories — I don't like the like button, I always accidentally click it — and generate the UI without it. So it's going to be a revolutionary replacement of the interface. In the future, we might have much more powerful LLM and coding models running behind the scene, and in the front end, the diffusion model will actually be the front end to show stuff to you.
这就是我想象的样子。扩散前端,确定性后端。是的,类似这样。我觉得这非常昂贵,但你知道……
That's how I imagine it. Diffusion front end, deterministic back end. Yes, something like that. I find it very expensive but, you know...
我觉得有趣的是你说后端 LLM 写代码是确定性的,但好吧。是的,你写一次,编译,然后执行。
I find it interesting you called LLMs writing code on the back end deterministic, but okay. Yeah, you write it once, compile it, and then execute.
如果你想想成本,假设 H100 每小时 1 美元,每天用 8 小时,一个月 30 天,那么每月你要付 240 美元。你可能不想为此付费。这甚至比 Claude Code Max 还贵。但如果你想想算力成本每年下降两倍,我认为未来很可能会到来,对吧?算力成本下降,算力变快,模型变聪明,模型变小。
If you think about the cost, say an H100 costs $1 per hour, and if you use it 8 hours a day for 30 days, every month you're paying $240. You likely don't want to pay for that. That's even more expensive than Claude Code Max. But if you think about compute costs coming down like two times every year, I think the future will likely arrive, right? Compute cost comes down, compute gets faster, model gets smarter, model gets smaller.
是的,我不知道你为什么说两倍,因为我觉得是 100 倍。在语言模型中,对于相同水平的 LLM ELO,大约每 12 到 18 个月提升 100 到 1000 倍。
Yeah, I don't know why you say two times because I think it's like 100 times. In language models, it is roughly 100 to 1000 times every 12 to 18 months for the same given level of LLM ELO.
那是所有因素的综合,对吧?那是模型性能加上算力。所以不同于仅仅是算力成本下降,但你知道,一个非常有趣的未来。
That's a net of everything, right? That's model performance alongside compute. So different than just compute cost coming down, but, you know, a very interesting future.
是的。对于网页设计师来说,他们必须大声疾呼可访问性是个问题,对吧?比如你如何处理屏幕阅读器之类的?但没错,这比任何你能用代码生成的东西都有更高的带宽叙事。对吧?所以我认为这就是大致的想法。
Yeah. For the web designers, they will have to shout out that accessibility is an issue, right? Like how do you deal with screen readers or whatever? But yes, this is higher bandwidth storytelling than anything you can possibly generate with code. Right? So I think that's the rough idea.
我想补充一点,人类在看东西、看视频时自然拥有最大的输入带宽,而在说话时拥有最大的输出带宽。所以在未来,可能会是我们与 AI 模型对话,而 AI 模型用生成式 UI 回应。那将是在 Neuralink 出现之前与 AI 模型交互的最大输入输出带宽。而且它也非常定制化,对吧?有些人非常视觉化,有些人没那么视觉化,他们更喜欢文本。但生成式 UI 最好的地方在于它也可以是文本。
I'd like to add a little bit that humans naturally have the maximum bandwidth when we are looking at things, looking at videos, and we also have maximum output bandwidth when we are talking. So in the future, it might be something like we talk to AI models and the AI model responds back with generative UI. That will be the maximum input and output bandwidth to interact with AI models before Neuralink happens. And it's also very custom, right? Some people are very visual, some people are not as visual, they prefer text. But the best thing about generative UI is it can also be text.
是的。
Yes.
还有另一个我们想强调的项目,就是 Neuro OS。有点类似的想法,但这里你实际上是在用视频模型模拟一个操作系统。是的。你可以玩《毁灭战士》,你可以用 Firefox。我觉得这显然没那么令人印象深刻,因为它是一个我可以运行的操作系统。但这里一切都是想象出来的。
There's another project that we wanted to highlight, which is the Neuro OS. Kind of similar idea, but here you're literally simulating an operating system with a video model. Yes. And you can play Doom, you can use Firefox. I find this mildly less impressive obviously because it's an OS that I can run. But here everything is imagined.
我习惯按 Command+W 关闭 Firefox 标签页。我没崩溃。这太沉浸了。对我来说太沉浸了。我想关闭标签页。但没错,我可以玩生成的《毁灭战士》。
I was used to pressing Command+W to close the Firefox tab. I didn't crash. It's too immersive. It's too immersive for me. I wanted to close the tab. But yes, I can play generated Doom.
这快得惊人。是的,因为我记得大概一两年之前有个演示。有人尝试用图像模型做第一人称射击游戏。没有一致性。非常慢。但这里,实际上,它就是《毁灭战士》。
This is shockingly fast. Yeah, because I remember there was a demo about maybe one or two years ago. Someone tried to do a first-person shooter with an image model. There was no consistency. It was very slow. But here, realistically, it's Doom.
我的意思是,我认为这有两面性,对吧?比如,运行一个游戏是什么?它的重头戏实际上是游戏引擎,所有的光照、所有那些东西、图形。这只是一段视频,对吧?就像我们已经解决了一致性。这仍然,你知道,看起来像几年前的图像生成。有一些时间一致性,但它基本上只是把图像拼接成帧视频。但这是一个很好的视觉表现,用来描绘你想看到的未来,对吧?就像我在这些中看到的更多。
I mean, I think there are two sides to that, right? There's like, okay, what is running a game? The heavy part of it is actually the game engine, all the lighting, all that stuff, the graphics. This is just kind of video, right? Like we've solved consistency. This is still, you know, it looks like a few years old image generation. There's some temporal consistency, but it's kind of just images stitched together as frame video. But it's a good visual representation to picture the future you want to see, right? Like that's what I see in these more.
这让我想起视频模型是如何变得越来越好的。
This reminds me of how video models get better and better.
所以,Neuro OS 乍一看感觉就像是 Windows 的劣质版,对吧?但区别在于,这个模型对现有操作系统过拟合了。它无法生成任何不同的东西,但这实际上也和视频模型类似。当我们训练这些视频模型、图像模型时,我们在互联网数据上训练;互联网上没有想象中的超自然内容。但一旦我们训练了这个模型,你可以提示它生成数据集中从未存在过的超自然内容。所以,如果你在互联网上所有的标准屏幕录制上训练你的 Neuro OS 或 Neuro 计算机,模型可以想象出一个全新的界面来与计算机交互。
So, Neuro OS is kind of if you just look at it, it feels like it's just a crappy version of the Windows we could have, right? But the difference is the model is overfitted on the existing operating systems. It can generate nothing different than that, but it's actually also similar to video models. So when we're training these video models, image models, we train them on the internet; there's no imaginary supernatural stuff on the internet. But once we train this model, you can prompt the model to generate something supernatural that has never existed in the dataset. So if you train your Neuro OS or Neuro computer on the standard screen recordings on the entire internet, the model can imagine a completely new interface to interact with a computer.
是的,这对我来说是件神奇的事。通常,分布外泛化效果很差,但不知何故我们学到了某种内部世界模型。是的,就像你说的,“这个加号但看起来像彩虹和蝴蝶。”它会照做,而且看起来还挺合理。
Yeah, this is one of those things that is magical to me. Usually generalizing out of distribution is bad, but somehow we have learned some kind of internal world model. Yes, that you say, you know, 'this plus but it looks like rainbows and butterflies.' It'll do it and it'll kind of make sense.
是的,这挺酷的。我不知道还有什么要补充的。我确实想再多谈一点模型架构方面的事情,我觉得你正在触及这一点;这真的很迷人。我们没有太多机会讨论这个。所以我们报道过的一篇论文——我们报道了每年的 segment anything 发布——我不知道你是否关注,我的意思是你是个计算机视觉专家,所以你知道——他们做了记忆注意力机制,这挺有意思的。我一直觉得,任何能在时间维度上保持一致性的东西都非常迷人。我不知道是不是计算机视觉这边渗透到了视频生成那边?我认为这还没有被充分探索,对吧?就像我们讨论过的用于标注,但你可以直接借用架构本身。
Yeah, that's kind of cool. I don't know if there's any more comment on there. I did want to touch a little bit more on the model architecture stuff, which I think you were getting at; it's really fascinating. We don't get a chance to talk about this enough. So one of the papers that we covered—we've covered every annual segment anything release—and I don't know if you follow, I mean you're a computer vision guy so you know—so they did memory attention, which is kind of interesting. And I always think like anything where you can across the temporal dimension keep some consistency, I think it's very fascinating. And I don't know if basically like does that the CV side bleeding into videogen side? I think it's underexplored, right? Like we talked about it for labeling, but you can borrow the architecture itself.
而且还有完全不同的方法,对吧?就像你提到的“世界模型”这个术语。所以我们从视频模型谈到了世界模型。有扩散模型,但人们也在尝试其他方法。所以我们之后可能也会谈到那些。
And there's also completely different approaches, right? Like you brought up the term world model. So we went from video model to world model. There is diffusion, but there's also other approaches that people are doing. So maybe we get into those after as well, you know.
是的,他对世界模型有一套完整的定义。我觉得我们抛给你太多东西了。你想评论什么都可以。我认为我们实际上应该回过头来评论一点:好吧,我们之前讨论了从图像生成到视频模型的训练步骤。有一件事我们不太常看到,就是,好吧,你提到了训练数据的差异,对吧?所以你没有那么多数据;视频模型可能无法泛化。但训练一个大视频模型的成本是多少?我们知道对于大语言模型来说大致是——好吧,就像今天发布的 poolside 那个东西,对吧?它是一个 Gemma 级别的模型,在大约 40 万亿 token 上训练,用了这么多 H200,花了这么长时间,对吧?你可以看到确切成本是多少。那么,多少 GPU 小时乘以多少 H200 成本?那么对于视频模型、图像模型,我们如何做后端计算?你如何分解这些?
Yeah, he has a whole definition of world models and stuff. I feel like we threw a lot at you. Whatever you want to comment on. I think one thing that we should actually comment back on is: okay, so we were talking about the steps to train image gen to video model. One thing we don't see as much of is, okay, you brought up the delta in training data, right? So you won't have as much; a video model might not generalize. But what is the cost of training a large video model? So we know for LLMs roughly—okay, even like the poolside thing that came out today, right? It's a Gemma-level model trained on roughly 40 trillion tokens at this many H200s over this much time, right? You can see what is the exact cost of that. So how many GPU hours over how much H200 cost? So how do we do the backend math of, you know, same thing for video models, image models? How do you kind of break that down?
我可以分享一些粗略的计算。令人惊讶的是,视频模型的成本与语言模型非常可比,在最大规模上,可能相当于一个中等规模的语言模型。我说过,仅存储视频本身就要花很多钱。你可以查一下 AWS 之类的。通常,假设你有十亿个视频,每个视频比如 5 MB,那么你需要大约 5 PB 来存储这些视频。另外,记得我们说过你使用 VAE 来压缩视频,你还需要存储——通常你需要存储那些连续特征——而且在你的存储中,这大小也与视频本身相当。所以仅存储这些视频和特征就需要几十 PB。
I can share some back-of-the-envelope calculation. So surprisingly, video models—the cost is very comparable to language models, and at the largest scale, maybe like a medium-scale language model. I said just storing the videos alone costs a lot. You can maybe look up on AWS or something. Usually, say if you have a billion videos and let's say each video like five megabytes, then you need like five petabytes to just store those videos. And also, remember we talked about you use a VAE to compress the videos, and you also need to store—typically you need to store those continuous features—and also in your storage, that's also comparable size with the videos themselves. So just storing these videos and the features is tens of petabytes alone.
我刚刚查了一下计算:5 PB 在 S3 标准存储上每月是 10 万美元。好的。
I just looked up the calculation: 5 petabytes on S3 standard is $100k per month. Okay.
而且你需要——然后像几十——20 万,甚至更贵的是你通过互联网的入站和出站流量。你必须下载那些视频。我相信在 AWS 上这比单纯存储那些视频更贵。而且每次训练运行你可能需要拉取它们一次。如果你训练多次,成本就更高了。所以,仅仅是网络存储,那些成本——我猜每月要几百万美元来存储所有东西,更不用说 GPU 成本了。
And you need—and then like tens—200k and even more expensive is you have the ingress and egress through the internet. You have to just to download those videos. I believe it's more expensive on AWS than just storing those videos. And each training run you probably need to pull them once. If you train multiple times, it's even more than that. So it's like just storing on the network, those costs—I guess it would be a few million per month to just store everything, not to mention the GPU cost.
我顺便提一下:你知道,算力租赁——GPU 租赁非常高效。有一方面:好吧,你可以像 xAI 那样建自己的数据中心。我们是不是也应该自己建存储和算力?就像云成本相比,你知道,是的,尤其是出站流量之类的。所以你知道。
My side tangent: you know, the compute rental—GPU rentals are very efficient. There's one side: okay, you can be xAI and build your data center. Should we not just build our storage compute as well? Like cloud cost compared to just you know, yeah exactly, especially with egress and stuff. So you know.
这是个好主意,但它也有自己的挑战。
That's a good idea, but it also comes with its own challenges.
当然,当然。
Of course, of course.
是的,比如建造 GPU 数据中心的人可能没预料到需要这么多存储。而且,人们通常只是把存储建在某个地方,只有 CPU。
Yeah, like people who build the GPU data centers might not expect this much storage. And yeah, people build storage typically they just build it somewhere. There's just CPUs.
我刚刚查了:五——AWS 只对出站流量收费,入站免费。五 PB 的第五层是 23 万美元。
I just looked up: five—AWS only charges for egress, not ingress. Tier five for five petabytes is $230k.
是的,甚至更贵。
Yeah, even more expensive.
但存储是按月收费的,对吧?你存进去就取不出来了。所以这很酷。没关系。
But storing is per month, right? You check in and you cannot check out. So it's cool. It's okay.
长话短说,我的估算比你想象的要大。是的。
TL;DR, my math is larger than you think. Yes.
我粗略估算的 GPU 小时数乘以 GPU 成本也很大,你知道,我漏掉了一些存储成本。所以基本上你比普通训练更受 IO 限制。
My back-of-the-envelope math of GPU hours times GPU cost is also very much, you know, I'm missing some storage. So you're basically also more IO-bound than normal training.
是的。是的。
Yes. Yes.
因为数据加载、缓存,一切都变得超级重要。
Because data loading, caching, everything becomes super important.
是的。所以在 Cosmos 中,我们做了很多优化来避免 IO 瓶颈。所以,说到训练,实际训练模型的 GPU 成本。如果你看看开源模型,这些视频模型有多大。我认为像 LTX 有 190 亿参数,那是一个密集模型。人们也在探索——所以可能像 200 亿活跃参数,总共 1000 亿。所以那和中等规模的语言模型大小相似。如果你看 token 数量,我们在 Cosmos 中披露了,面部 token 也有几十万亿。所以综合来看,训练这些视频模型的成本实际上与语言模型相当。更不用说基础设施与语言模型略有不同。所以训练这些模型可能效率较低。你能获得传统扩散加速的好处吗?对于图像,有 LCM、LoRA 用于微调。有很多东西。是的,还有流匹配。
Yeah. So in Cosmos, we did a lot of optimizations to make it not IO-bound. So, yeah, speaking of the training, actually training the model, the GPU cost. If you look up the open-source model, how big these video models are. I think like LTX has 19B parameters, that's a dense model. And people are also exploring—so it might be like a 20B active and like 100B total. So that's similar size as medium-size LM models. And if you look at number of tokens, we disclosed that in Cosmos it's also like tens of trillions of tokens on the facial tokens. So putting this together, the cost of training these video models is actually comparable with LMs. Not to mention the infra is slightly different from LM. So it might be less efficient to train these models. Do you get the benefits of traditional diffusion speed up? So for images, there's LCM, LoRAs, for fine-tuning. There's a lot of stuff. Yeah, there's flow matching.
已经做了很多工作。在推理方面,有一些重叠适用于扩散模型之类的。
There's a lot of stuff that's been done. There's some overlap that applies to diffusion on the inference side and stuff.
是的。推理方面完全是另一回事。我认为在训练方面,降低成本可能有点困难。在推理方面,最大的收益来自这些模型的蒸馏。你可以做所谓的步蒸馏,这与整体上的知识蒸馏略有不同。通常对于流匹配模型,你需要大约 100 步,而扩散模型甚至需要更多,比如一千步,才能生成好的图像或视频。步蒸馏试图从模型本身学习用少量步骤生成。就像这样:你用完整模型生成 100 步,然后你拿一个只生成 10 步的模型,让它从完美模型那里学习。
Yeah. So the difference on the inference side is a completely different story. I think for the training side, it might be a little bit hard to reduce that cost. For the inference side, the biggest gain is from the distillation of these models. So you can do what's called step distillation, which is slightly different from knowledge distillation overall. Typically for flow matching models, you need like 100 steps or something, and diffusion models even need more, like a thousand steps, to generate a good image or video. Step distillation tries to learn to generate in few steps from the model itself. It's like: you use the full model to generate in 100 steps, and then you take a model that only generates in 10 steps and let that model learn from the perfect one.
是的。为什么这能行?
Yeah. Why does this work?
强到弱有点像强到弱。我想从建模的角度来看,强模型(教师模型)试图对互联网上的图像和视频进行建模,那个分布极其复杂。但步蒸馏模型只是试图从教师那里学习。教师是一个模型,其大小是固定的。这个分布比整个互联网简单得多。这就是我对步蒸馏为何能行的直觉。所以通常这些模型在生产中运行时只使用少量步骤。在 Cosmos 中,我相信我们有四步和八步。如果你做一些更简单的任务,如图像到图像的转换,它甚至可以一步完成,比如在 Cosmos Transfer 中。
Strong to weak seems kind of like strong to weak. I guess from the modeling perspective, the strong model (the teacher model) is trying to model the images and videos on the internet, and that distribution is extremely complex. But the step distilled model is just trying to learn from the teacher. The teacher is a model, and its size is fixed. The distribution is much simpler than the whole internet. That's the intuition I have for why step distillation can work. So usually these models serve in production; they only run in a few steps. In Cosmos, I believe we have like four steps and eight steps. If you do some simpler task like image-to-image translation, it can even run in one step, like in Cosmos Transfer.
是的,我认为这是指导许多一致性模型工作的相同直觉。我给你发了一个关于 SCM 的链接。我不知道你是否看过。对我来说,那实际上是我见过的最令人印象深刻的 OpenAI 论文之一,就是那个统一的一致性模型宏大概念。我不知道你对此有什么评论。
Yeah, I think this is the same intuition that guides a lot of the consistency model work. I sent you a link for SCM. I don't know if you covered that. To me, that was actually one of the most impressive papers I've ever seen from OpenAI, like that unifying grand concept of consistency models. I don't know if you have any comments on this.
所以有几种不同的方法。例如,一致性模型,还有我们不应该忘记 GAN。GAN 实际上是步蒸馏的鼻祖,因为它们从一开始就训练一步生成。所以很多方法,例如分布匹配蒸馏,使用 GAN 作为蒸馏的损失之一。GAN 只是告诉你:生成一张图像,然后它有一个判别器来判断这张图像是否真实。所以模型只需要学习分布,而不是完整分布,因为在训练中,模型被要求从互联网重建真实图像,这极其困难。当你训练 GAN 时,这是一个一步过程:你生成一张图像,它检查图像看起来是否和互联网上的图像一样真实,这是一个简单得多的任务。将很多这些方法结合起来,人们通常会这样做——一致性模型和分布匹配——我们就可以得到这些少步模型。
So there are a few different approaches. For example, consistency models, and also we shouldn't forget GANs. GANs were actually the OG step distillation because they train to generate in one step from the beginning. So a lot of approaches, for example distribution matching distillation, use GAN as one of the losses for distillation. GAN just tells you: generate an image, and then it has a discriminator to tell whether this image is real or not. So the model just needs to learn the distribution, not the full distribution, because in training, the model is asked to reconstruct the ground truth image from the internet, which is extremely hard. When you're training a GAN, it's a one-step process: you generate an image, and it checks if the image looks as real as the image from the internet, which is a much simpler task. Combining a lot of these approaches together, people typically do that—consistency models and distribution matching—and we can get these few-step models.
好的,然后我想补充一步,那就是音频和视频。
Okay, then there's one step I wanted to add, which is audio and video.
是的。
Yes.
是的。所以,Gokcen Imagin 0.9,我相信,是第一个大规模部署的音频-视频 Transformer 模型。
Yeah. So, Gokcen Imagin 0.9, I believe, is the first audio-video transformer model deployed at a large scale.
那是你的第一个模型。
And that was your first model.
是的,那是 Gokcen Imagin 的第一个模型。它是音频-视频联合生成。我认为难点在于模态对齐,因为在这个联合模型之前,我们有文本到视频的对齐。我们有对应的文本和视频。通常,大多数 VLM 能理解图像和视频,但很少能理解音频。如果你看 LM 端的音频生成,你可以和它们完美对话,但如果你让它们唱首歌之类的,它们通常不太擅长。而且,它们也没有音乐。
Yeah, that was Gokcen Imagin's first model. It's audio-video joint generation. I think the hard part is the modality alignment, because before this joint model, we have text-to-video alignment. We have corresponding text and video. Typically, most VLMs understand images and videos, but very rarely do they understand audio. And if you look at the audio generation on the LM side, you can talk to them perfectly fine, but if you ask them to sing a song or something, they typically are not very good. Also, they don't have music either.
难点在于音频实际上有两个组成部分:离散部分和连续部分。离散部分就像语言,所以当我们说话时……
The hard part is that audio actually has two components: a discrete component and a continuous component. The discrete component is like language, so when we speak...
这只是一个 ASR 问题。是的。
It's just an ASR issue. Yeah.
是的。可以说是带有一些特征的文本 token。
Yeah. It's text tokens with some characteristics, I would say.
但音乐……
But music...
我想语音领域的人会不同意。它像频率之类的……但很大程度上,音乐是完全不同的。它非常连续,你不能像语言模型中的离散 token 那样建模它们。这是模型的难点,更不用说我们还要将文本、视频和音频对齐在一起。
I think the speech guys would disagree. It's like frequencies and... but largely, the music is completely different. It's very continuous, and you cannot model them like discrete tokens in language models. This is the hard part for the model, not to mention we have to align text, video, and audio together.
所以一些重大挑战是:首先,正如我们所说,大多数 VLM 无法理解音频。所以你必须有一些方法来进行音频的合成数据生成。你必须给模型加标题,这涉及大量的合成数据和人工数据工作。而且毫不奇怪,大多数 VLM 在识别音乐的节拍、音调和细节方面非常差。它们可以对这是哪首歌做出一些一般性预测,但很难描述音乐的细节。就像我们在图像生成中提到的,你必须尽可能详细地描述图像,以便盲人可以重建它。所以在这里,就像聋人可以在没有实际听到的情况下重建音乐的声音。也许你可以认为需要拥有音乐和对话的所有细节。
So some significant challenges are: first, as we talked about, most VLMs cannot understand audio. So you have to have some way to do synthetic data generation for audio. You have to caption the model, and that involves synthetic data and human data effort, a lot. And not too surprisingly, most VLMs are very bad at recognizing the beat, tone, and details of music. They can give some general prediction of which song this is, but it's very hard to describe the details of the music. Like we mentioned in image generation, you have to describe the image in as much detail as possible so that someone blind can reconstruct it. So here, it's like someone deaf can reconstruct how the music sounds without actually listening to it. Maybe you can think of it as needing to have all the details of the music and the dialogue.
挑战通常在于音乐和音频之类的东西,还是说有一个基线,我们可以理解叙述和对话,但音频中的细微差别导致了数据问题?还是说从零开始你就能做好?
Is the challenge there typically stuff like music and audio, or is it just like there's a baseline where we can understand narration and conversation, but there are nuances in audio that cause data issues? Or is it just from state zero you do it all right?
所以一个重要的事情是对齐。模型必须知道视频和音频;它必须有基于时间的对齐:在哪个时间步,视频和音频 token 相互对应。但我们实际上对其他大多数模态都没有这种对齐。如果你考虑文本和图像,或文本和视频,它们是松散对齐的。你可以对视频中发生的事情有一个描述,但你不必精确描述,例如,在时间步 1 秒发生了什么,2 秒发生了什么。
So one important thing is the alignment. The model has to know the video and audio; it has to have a time-based alignment: at which time step the video and audio tokens correspond to each other. But we actually don't have this kind of alignment for most other modalities. If you think about text and image, or text and video, they are loosely aligned. You can have a description of what's going on in the video, but you don't have to exactly describe, for example, at time step 1 second what happened, at 2 seconds what happened.
那么理想的时间步长是多少?你把它消融掉,然后大概是四秒左右。
So what was the ideal time step? You have to ablate it and then it's like four seconds or something.
这取决于你如何设计模型。要让模型把时间作为一种模态来感知,模型就需要具备时间感知能力,这相当独特。如果你让一个语言模型去完成一个任务,比如你问它,它会说“哦,这个任务大概需要 12 小时完成。”然后一小时后你回来,它已经花了两天时间,把所有资源都耗尽了。
So that comes down to how you design the model. For the model to be aware of time as a modality, the model is time-aware, and that's something pretty unique if you think about all. So if you ask an LM to complete a task, say you ask them and they would say, "Oh, this task will probably take 12 hours to complete." And they come back in one hour, they have already spent two days on this and have exhausted everything.
是啊。所以语言模型本身,它们没有时间感。
Yeah. So the LMs themselves, they don't have a sense of time there.
我其实不认为仅仅是它们没有时间感。我觉得这有一定依据,对吧?就像你告诉某人,“去开发这个功能,去实现它。”你通常会对需要多长时间有个大致概念,而不考虑语言模型的工作速度。回想两年前,如果我让你为 latent space 建一个新前端,带搜索栏等等,你会估计需要几天时间。所以你对语言模型说“去构建这个”,它会说需要几天。但我觉得这有一定依据,而不是它们完全没有概念。不是说它们有很好的理解,但你可以看出这个估计从何而来——它们是在大量文本上训练的。
I actually don't think that's just them not having a sense of time. I think it's somewhat grounded, right? Like you tell someone, "Go work on this feature, go implement this." There's a general understanding you would have of how long that would take without LLM's working speed, right? So you think back like 2 years ago, if I tell you to build me a new front end for latent space, have a search bar, have all this, you'll estimate that it'll take a few days, right? So you tell an LLM, "Go build this," it'll take me a few days. But I think it's somewhat grounded as opposed to them not having the best. Not saying that they have a great understanding, but I think that example is like you can see where it comes from, right? You're trained on all over the text.
它们是在试图估计人类会怎么说。
They're trying to estimate what a human would say.
对。因为数据大致代表的就是这个。
Yes. Because that's what the data kind of represents.
而且互联网上的人本来就有个估计。
Then the core on the internet people have an estimate.
对。而且不仅仅是直接的训练样本,对吧?就是你对 token 所代表的事物耗时的一种世界理解。比如去读一本书,会花你一段时间,对吧?即使你什么都不做只是读书,也要几天。所以,是的,我会读它。它花了我几个小时。我花几个小时来读完这份研究。
Yeah. And not even just in direct training samples, right? Just your world understanding of tokens of how long stuff takes, right? Go read a book. It'll take you a while, right? Even if you do nothing but read a book, it takes a few days. So, yeah, I'll read it. It took me a few hours. It'll take me a few hours to go through this research.
但有点跑题了。嗯,这是我之前没怎么表达过的一个思路:一个完整的 世界模型 也必须是递归的,也就是说,世界模型中的参与者必须意识到自己拥有一个世界模型。这就像一条递归链。而且世界模型可能是错的,它们需要更新它,等等。是的,我们在 newsletter 上也讨论过,需要某种递归或对抗性的世界模型。
But some tangent, somewhat a yeah, this is a train of thought I haven't really expressed until now which is basically like a full world model must also be recursive, meaning that the participant in the world model must also be aware that they have a world model. Which is like this whole recursive thing down the line. But yes, and that the world model can be wrong and that they need to update it and blah blah blah. Yeah, we've argued this on the newsletter as well that there needs to be sort of recursive or adversarial world models.
好吧,那我想问一下,你怎么定义 世界模型?
Okay, I mean just you know to ask how do you define world models?
哦对,我们来谈谈这个。
Oh yeah, let's go there.
对。为了提供背景,我们谈到了视频生成,然后世界模型之间是有区别的。你的定义是什么?你怎么看这两者?
Yeah. So just for context, you know, we talked about video generation and then there's a distinction between world models. What's your definition? How do you see the two?
好的。先声明,我不打算争论什么是世界模型。有很多定义。我只谈我的定义,因为我来自多模态领域,所以主要从视频角度谈。世界模型 就是实时交互的长时域视频。它有三个部分,我们逐一讨论。首先是交互部分。我们看看 Facebook 和 neuro computer。世界模型允许你通过键盘、鼠标,可能还有语音与之交互。所有这些模态你都可以与模型交互,模型应该合理响应。第二部分是实时性。一旦你移动鼠标,如果世界模型生成游戏,游戏响应能有多快?如果你是职业 CSGO 玩家,可能会说响应时间必须在 10 毫秒以下甚至更短。
Yeah. So disclaimer, I'm not going to debate like what is world model. There are many definitions. So I'll just talk about my definition since I came from the multimodal domain. So mainly talking from video. So world model is like real-time interactive long-horizon videos. So there are three parts. Let's talk about them one by one. So the interaction part of it. So we just look at Facebook and neuro computer. So the world model can allow you to interact with them through keyboard, mouse, and maybe also voice. So all of these modalities you can interact with the model and the model should respond reasonably. Second part is real-time. So once you move your mouse, if the world model generates the game, how fast can that game respond? So if you're like professional CSGO players might say you have to respond in sub 10 milliseconds or even less.
不,60 帧每秒。来吧。
No, 60 fps. Let's go.
哦,300 帧每秒。
Oh, 300 fps.
500 帧每秒。
500 FPS.
等等。好吧,我没算,但好吧。300 帧每秒,那就是 3 毫秒。所以你必须响应。
Wait. Okay, yeah, I didn't do the math, but yeah. Okay. 300 FPS. That's a 3 millisecond. So you have to respond.
大多数视频模型做不到。
Most of the video models cannot do that.
对。但如果你有一个视频模型,比如数字人,响应时间可能更宽松,比如典型的实时语音交互是 200 毫秒。那就宽松多了。但即使是 200 毫秒也很棘手,因为记得我们提到过 VAE 带来的时间压缩。如果你不压缩时间维度,序列长度会爆炸。所以如果你想让模型具备实时性,就必须处理长上下文问题。第三部分是长时域,因为我们不会只玩几秒钟的视频游戏。大多数视频模型只生成几秒。我们要玩几分钟、几小时。模型必须能够生成长篇内容。所以把这三者结合起来,就是实时长时域交互视频。我认为最终状态会是,比如 Playbook 的视频版本,你可以与 neuro computer 交互。你移动鼠标,点击生成界面,它会通过像素实时响应。但要达到那一步,还有很长的路。所以在 Genie,我领导一个小型世界模型团队时,第一步就是构建视频扩展。视频扩展是交互性的第一步。
Yeah. But if you have a video model that is say like a digital human, the response time might be more generous, maybe like typically for real-time voice interaction is like 200 milliseconds. So that's much more generous. But even 200 milliseconds is pretty tricky because remember we mentioned you have this temporal compression coming from the VAE. So if you don't compress the temporal dimension, your sequence length is going to explode. So if you want to have this real-timeness in your model, you have to deal with long context problem. And the third part is long horizon because we are not going to just play with video games just for a few seconds. Most video models only generate a few seconds. We're going to play with minutes, hours. The model has to be able to generate long-form content. So putting these three together, it's a real-time long-horizon interactive video. I think the final state will be, for example, a video version of Playbook where you can interact with a neuro computer. You move your mouse and you click on the generative interface and it will reply to you through pixels generally in real time. But getting there, it's a very long way. So one of the first steps at Genie, where I led a small world model team, was to build video extension. Video extension is the first step of interactivity.
这是第一步。对,这是第一步。
It's a first step. Yeah. So it's the first step.
我们这里有。视频编辑。对。
We have it here. Video editing. Yeah.
对。
Yeah.
对。所以第一步是因为这解锁了长时域,通常对大多数视频生成模型来说,你给它一个提示或一张图片作为初始帧,生成视频,就完了。只是一次性的。有些创作者会尝试把最后一帧作为第二个视频的第一帧。有时能行,但如果你做几次,质量会下降,而且它没有整个视频的上下文。所以时间上……
Yeah. So the first step is because this unlocks long horizon typically for most of the video generation models. You give it a prompt or an image as an initial frame. You generate video, that's it. That's just one time done. And some creators would try to use the last frame as a first frame for the second video. It kind of works sometimes, but if you do it a few times, it degrades to a degree and it doesn't have that context over the full video. So the temporal...
对。因为你只给了它最后一帧。当然会这样,对吧?
Yeah. Cuz you only gave it the last frame. Of course, right?
对,没错。
Yeah. Exactly.
但这其实是个挺有趣的 hack。比如你见过……
But it's actually a pretty fun hack. Like if you seen like...
哦,不。他有更好的东西。
Oh, no. He has something better.
对。比如 View 3 有上一段视频的 1 秒上下文。这比只用最后一帧好一点,但有同样的问题:如果你扩展几次到一分钟,视频质量会比第一段差很多。
Yeah. Yeah. Yeah. And for example like View 3 has like a 1-second context of the last video. It's slightly better than using the last frame but it has the same problem similar problem that the quality would degrade if you extend a few times to like one minute, the video quality would look much worse than the first video.
第二个问题是,模型没有对之前发生事情的长程知识。所以如果生成一段对话,角色的声音可能会随时间变化,尤其当一秒的条件信息不包含之前的上下文时。这些都是核心挑战。Gawk 想象视频扩展拥有所有之前生成视频的历史上下文。它知道谁在说话、出现了什么物体等所有信息,并利用这些来生成下一个视频。如果我们天真地这样做,把所有历史视频 token 都塞进上下文,上下文长度会轻易爆炸,尤其对于视频模型来说,可能达到几百万的上下文长度。
Second, another problem is that the model doesn't have long-range knowledge of what's happening before. So if they generate some dialogue of people speaking, their voice might change over time, especially if the one-second conditioning does not cover the previous context. These are the core challenges. So the Gawk imagine video extension has historical context of all the previous generated videos. It has a context of who's speaking, what objects have appeared, and everything, having that to generate the next video. If we naively do this, you can imagine just putting all the previous history video tokens into the context, the context length will easily explode, especially for video models that can be a few million context, I would imagine.
是的,比如在 Cosmos 中,我觉得仅 5 秒的视频就有大约 5 万到 6 万个 token。所以如果你生成 50 秒,那就是 50 万个 token。如果再长,很容易就爆炸了。这个长程问题是我们试图为世界模型解决的第一步。事实证明,人们非常喜欢视频扩展;很多创作者喜欢用视频扩展来制作更长的视频。这正是我喜欢的一点:你有一个通向最终目标的中间步骤,而不是直接跳到最终版本。
Yeah, for example in Cosmos, I think just 5 seconds of video is like 50k or 60k tokens. So if you do 50 seconds, that's 500k tokens. If you go longer, it easily explodes. This long-horizon problem is the first step we're trying to solve for world models. It turns out people love video extension a lot; many creators love using video extension to create longer-form videos. This is the part I like: you have an intermediate step toward the final goal instead of a straight shot to the final version.
是的。
Yeah.
但我能看出你对最终目标有很强的愿景。
But I can see you have a strong vision of where we want to end up.
是的。
Yeah.
这看起来像是一个效率问题吗?比如,我们现在有数百万 token 的上下文。如果类比语言模型,我们一开始有很短的上下文,2000、8000,然后扩展到 100 万、1000 万。当然,还有有效上下文的问题。但归根结底,这值得吗?还有训练数据方面。在视频中,可能稍微容易一些,因为我们有 1 亿 token 的视频,对吧?直接拿一部电影作为完整上下文。这是推理效率的问题吗?虽然昂贵,但我们知道如何解决,或者为什么这不是正确的方法?我更大的观点是关于你提到的世界模型的第二点:你说它需要是交互式和实时的,对吧?你应该能够玩游戏并实时看到交互。我在研究中看到的一件事是,你实际服务的东西往往与你构建的不同。我们讨论过蒸馏。你训练一个大模型,蒸馏它,做量化、推测解码,我们做所有这些来高效地服务它。我们难道不应该先有一个能良好交互的世界模型解决方案,做推理优化,服务它,然后再蒸馏吗?也就是先解决它,再让它实时。另一个类比是持续学习。我们需要有人先解决它并证明它有效,即使效率不高;过几年,人们会把它变得高效。常规注意力机制也是如此:它有效,经过几年,人们有了不同形式的注意力,我们把它扩展到长上下文并变得高效。所以这里有两件事:一是它似乎有效,你已经扩展了它,我们难道不能随着时间的推移更高效地扩展它吗?如果这个方案有效,我们还需要一个不同的方法吗?交互也是如此。如果我们能实现它,如果能以某种方式让它工作,我们之后可以解决推理效率的问题。
Does it seem like it's an efficiency issue? Like, okay, we're at a few million tokens context. If you draw the parallel to language models, we had very short context, 2,000, 8,000, then you scale it up to 1 million, 10 million. Sure, there's effective context. But at the end of the day, it's just what's it worth? There's the whole training data side. In video, it might be slightly easier because we have 100 million token video, right? Just take a movie with the full context there. Is this efficiency from an inference standpoint that it's expensive but we know how to solve it, or why is this not the approach? So my broader point was on your second point of world models: you say it needs to be interactive and live, right? You should be able to play a game and see the interaction live. So one thing I see with research is a lot of what you actually serve is different from what you build. We talked about distillation. You train a big model, you distill it, you do quantization, speculative decoding, we do all this stuff to serve it efficiently. Should we not just have a solution like a world model that can interact well, do inference optimization, serve it, distill it secondarily? So make it real-time after you solve it. Another parallel is continual learning. What we need is someone to solve it and show it works inefficiently; give it a few years, people will make it efficient. Same thing with regular attention: it worked, over a few years people have different forms of attention and we've scaled it to be efficient at long context. So two things there: one is it seems like it works, you've scaled it, can we not just scale it a lot more efficiently over time? Do we need a separate approach if this works? And same thing with interaction. If we can get it done, if we can solve some way that it works, we can solve making it more efficient from an inference standpoint later.
是的,这确实是个很好的观点。在视频中,实际上有很多冗余。我们通过 VAE 解决了很多像素冗余,但在长程和长时域视频中还有更多冗余。比如,如果一个角色出现在第一个片段中,然后消失,只在视频结尾重新出现,你可能在生成中间不需要那个上下文。所以你只需要在需要的地方引用那个角色。这就是为什么我帮助构建了另一个功能:参考视频。它在这里吗?
Yeah, that's actually a very good point. So in videos, there's actually a lot of redundancies. We solve a lot of the pixel redundancy from VAE, but there's more redundancy in long-range and long-horizon videos. Say if a character appears in the first clip and then disappears, it only reappears at the end of the video, you probably don't need the context in the middle of the generation. So you only need that character where you need it. That's why I helped build another feature: reference to video. Is it here?
是同一个模型发布还是不同的?
Is it the same model release or a different one?
是另一个。你可能需要在 F 上搜索“reference to video”。
It's a different one. You probably need to search on F reference to video.
好的。
Okay.
所以参考视频允许你上传最多七张图像作为条件,并生成一个视频。比如,我想以 Sean 的自拍和他拿着一把刀作为条件。
So reference video allows you to upload up to seven images as condition and generate a video. Say I want to condition on Sean's selfie and holding a blade.
你把狗放进去了。
You put the dog in the thing.
是的,你可以把它们放进去,视频模型会生成视频并复制上下文。这可以解决很多问题,比如长上下文问题。它不需要非常长的上下文,但我觉得这是一个中间方案。
Yeah, you can put them there, and the video model will generate the video and copy the context over. So that can solve a lot of the problems, like the long context problem. It doesn't need to have a very long context, but I feel like it's an intermediate solution.
是的。模型应该能够有选择地知道在哪里引用。所以如果我想生成一部电影,我自回归地生成,每次 10 秒。现在这个角色出现了,我可以回溯到它第一次出现的地方,并把它带回来。
Yes. The model should be able to selectively know where to draw references. So if I want to generate a movie, I generate it autoregressively, 10 seconds at a time or something. And now this character appears, I can look back to where it first appeared and bring that back.
是的,这个我放了参考。那是 Optimus、爱因斯坦、我自己、Annie。
Yeah, this one I put the references. Yeah, that's Optimus, Einstein, myself, Annie.
奇怪的是,我用 Gro 搜索找到了它,它拉出了你的 LinkedIn 帖子,但你知道,我们找到了。
Oddly enough, I used Gro search to find it and it pulled your LinkedIn post, but you know, we found it.
但好吧,这是个问题。这不是你的错,但 XI 没有很好地传达你们做的所有工作,因为他们只是发布模型,仅此而已。但实际上这些细节非常非常好。据我所知,你刚才描述的一切都是最先进的。没有其他人做到过。
But okay, this is a problem. This is not your fault, but like XI doesn't communicate all this work that you do very well because they just have the model release and that's it. But actually these details are very, very good. As far as I understand, everything you just described is state-of-the-art. Like no one else has done it.
谢谢。
Thanks.
然后你只发了这篇带饼干的博客文章。我觉得这不够,你知道。
And then you just put this blog post with the cookies. I'm like this is not enough, you know.
是的。
Yeah.
但显然这是人们想知道的高层数据。
But obviously this is like the high-level numbers that people want to know.
但你知道,部分原因也是有些实验室不分享研究细节。
But you know, part of that is also like some labs don't share research into what happens.
不,但这实际上是在炫耀他们有多好,对吧?为什么不说你有能力用完整上下文进行扩展?这不是什么秘方。这是我们做的工作。是的。我不知道。
No, but this is literally bragging about how good they are, right? Like why would you not say that you are capable of extending with full context? This is not a secret sauce. This is like we did the work. Yeah. I don't know.
是的。我想不同的实验室有不同的沟通风格。
Yeah. I guess different labs have slightly different communication styles.
是的。总之,如果有 XI 的人在听,我们总是很乐意帮你们讲述故事。好的。所以,你做了参考,我觉得你的意思是这有点像权宜之计,对吧?比如你可以做七个,但一百个呢?
Yeah. Anyway, if anyone from XI is listening, we are always happy to help you tell your story. Yeah. Okay. So, you did references and I think the point you're making is this is sort of a kludge, right? Like this is you can do seven but what about 100?
是的。
Yeah.
那么你需要一个完全不同的东西。
Then you need a completely different thing.
所以我认为这是一种从历史中选择上下文的机制,你不必将整个历史都放入上下文中。例如,有一篇名为 FramePack 的论文提出了一种启发式方法:最新的历史,比如最后一秒,我放入完整的历史,而之前的历史我会压缩并缩小视频。所以我遵循这个模式,最大序列长度是固定的。离当前帧越远,图像越小。这只是一个启发式方法。我认为它可以更自动化。模型知道哪些历史部分可以被选择。所以这部分研究实际上正在被很多人积极研究。也很有趣。我觉得长上下文的这部分比其他部分稍微领先一些。例如,在 LLM 中,如果上下文不断增长,假设你调用了一个工具,工具调用历史非常长,那仍然在上下文中,并且即使你切换到其他话题,整个上下文也会一直存在。有一些智能体式的工具可以帮助你修剪工具结果,比如当你查询文件时,只显示前 200 行之类的,但这些都非常依赖启发式方法。
So I think it's a mechanism to select the context from the history, and you might not put the entire history into the context. For example, there's a paper called FramePack which has a heuristic: the latest history, like the last second, I put the entire history, and the history before that I would compress it and make the video smaller. So I follow this pattern that the maximum sequence length is fixed. So the further you are from the current frame, you have a smaller image. This is just a heuristic. I think it can be more automatic. The model is aware of which history part can be selected. So this part of the research is actually being actively worked on by a lot of people. It's also quite interesting. I feel this part of long context is a little bit ahead of the other parts. For example, in LLMs, if the context keeps growing, say you call a tool and the tool call history is extremely long, that's still in context and keeps growing even if you switch the topic to something else. The whole context would be there. There are some agentic harnesses that help you prune the tool results and prune like when you query a file, only show the top 200 lines or something, but those were very heuristic-driven.
给听众们说一下,我们写了一篇关于 Claude Code 泄露的文章,其中有八种不同的修剪方式,包括修剪工具结果等等,所以你可以去了解一下。
For listeners, we did a write-up on the Claude Code leak where there are eight different kinds of pruning, including pruning tool results and all that, so you can read up on that kind of thing.
是的,我认为持续学习的一个突破可能是自动管理自身上下文的方法。
Yeah, I think one breakthrough in continuing might be a way to automatically manage its own context.
这些都是启发式方法,它们将被机器学习取代。是的,有趣的是,同样的东西在 LLM 和视频模型中都在被研究。
These are all heuristics, and they will be replaced by machine learning. Yes, interestingly the same thing is being researched in both LLMs and video models.
所以有趣的是,在你展示的论文中,这实际上发生在模型层面,对吧?与语言模型相比,当然我们有基础的注意力机制,但你知道,我们会做自己的压缩,做自己的修剪,这与模型误差是分开的。希望最终一切都能融合。
So the interesting thing is also in the paper you showed, it's actually happening at the model level, right? Compared to language models, sure we have base attention, but you know, we'll do our own compression, we'll do our own pruning which is separate from model error. Eventually it all just boils in hopefully.
是的。是的。我认为这是一种注意力机制,但也是一种推理注意力。我觉得这和普通的注意力不同。说得通吗?
Yeah. Yeah. I think this is a form of attention, but also a sort of reasoning attention. I feel like that's different than normal attention. Does that make sense?
是的。是的。不同之处在于注意力,更不用说稀疏注意力了,像普通注意力你必须关注所有 token。是的。所以你没有一种高级机制来丢弃你不想关注的 token。作为人类,我们的注意力跨度出奇地小。
Yeah. Yeah. It's different in the sense that attention, not to mention sparse attention, like normal attention you have to attend to all of the tokens. Yes. So you don't have a high-level mechanism to drop which tokens you don't want to attend to. As humans, our attention span is surprisingly small.
你只能记住 11 位电话号码。
You can only remember 11 digits of a phone number.
但我有特征检测,对吧?我可以检测到,“哦,那是一个 11 位电话号码中的 1 2 3 4 序列。”
But I have feature detection, right? I can detect, "Oh, that's a sequence of 1 2 3 4 in a phone number that is 11 digits."
是的,非常好的模式匹配器。
Yeah, very good pattern matchers.
但人类的上下文可以工作,因为我们可以动态地从不同地方拉取上下文。我认为同样的机制也会发生在 LLM 和视频模型中。是的,RLM 是最近的,一些最近的工作。
But humans' context can work because we can dynamically pull in context from different places. The same mechanism, I think, is going to happen for LLMs and video models. Yeah, RLMs is recent, some of the recent work.
这并不疯狂,只是递归的。
Is there which is not that crazy, but it's just recursive.
我认为这在模型中也有些固有特性,对吧?比如这里有一个很好的例子。你把这些拿出来,你可以很好地阅读,但语言模型也非常擅长处理杂乱输入。你知道,你有一个转录文本,无论什么。直接扔进去,它非常擅长从噪声中解析。这可能是一种暴力方法。它可以查看、推理,但两者有相似之处。
I think it's somewhat inherent in models too, right? Like here's a nice example. You pull up these, you can read it fine, but language models are also very good at slop parsing. You know, you have a transcript, you have whatever. Just throw it in and it's very good at parsing through noise. That may be a brute force. It can look over it, reason over it, but there are parallels to both.
我觉得你把世界模型的东西和视频生成联系起来真的非常迷人,我认为很多人没有直接从像你这样的人那里听到过。所以我觉得这非常有帮助。我们还有其他要覆盖的工作吗,比如视频、音频、世界模型,或者 Omni 团队的其他东西,或者 xAI 的其他工作你想谈谈?看起来我们看到的都是公开宣布的。哦,酷。Cookies 然后还有很多被低估的东西,你知道,就在那个时候。
I think it's just really fascinating how you relate the world model stuff to the video generation, which I don't think a lot of people hear directly from people like you. So I think that's really helpful. Any other work do we cover like video, audio, world models, any other stuff in that Omni team I guess, or any other work at xAI you want to talk about? Seems like everything we see publicly announced. Oh, cool. Cookies and then there's so much more to any underrated stuff, you know, just at the time there.
是的,我觉得文化非常有趣,而且有点被低估了。所以文化是三句话:快速行动,构建,没有目标太宏大。而第一性原理,通常设定的目标非常宏大。当我最初思考时,它是不可能实现的。例如,我可以在 3 个月内构建一些东西,那是像“好的,我们正在组建团队,我们想要图像,想要视频,在这个截止日期前完成”吗?或者你如何倒推?只是大致上,你知道,我们想要在这个日期前推出一些东西?
Yeah, I feel the culture is quite interesting and a bit underrated. So the culture is three sentences: move fast, build, no goal is too ambitious. And the first principle, usually the goal set was very ambitious. It wasn't possible to achieve when I was first thinking about it. For example, I could build something in 3 months and was that like, "Okay, we're starting team, we want image, we want video, do it by this deadline"? Or how do you work back? Was it just okay, we have a rough by you know this date we want something out?
这是一个很好的观点。所以这是来自第一性原理思维。如果你想想,人们可能会说第一性原理思维更适用于物理世界而不是模型。我会说,例如,如果你考虑一些限制,比如获取数据,我们能多快获取视频?如果你考虑训练模型,端到端训练模型的迭代速度是多少,增加更多 GPU 会如何加速那个时间线?也许如果你需要人类数据,人类数据的周转时间是多少?如果你把所有这些放在一起,那就是第一性原理思维,你知道,时间线或可能实现某事的最小天数是多少?
That's a very good point. So it's from first principle thinking. If you think about people might say first principle thinking applied more to the physical world than the models. I would say for example, if you think about some limitation, for example acquiring data, how fast can we acquire the videos? And if you think about training the models, what's the iteration speed for training a model end to end, and how would adding more GPUs accelerate that timeline? And maybe if you need human data, what's the turnaround time for human data to arrive? If you put all of those together, that is first principle thinking where, you know, what is the timeline or the minimum number of days that is possible to achieve something?
我认为这有很多 Elon 式的思维,对吧?他就像,我认为他因说过“唯一不能打破的定律是物理定律”之类的话而闻名。
I think there is a lot of Elon's type of thinking, right? He's like, I think he's famous for saying that the only law you can't break is the laws of physics, something like that.
总的来说,你和 Elon 合作很多?
Just broadly, you worked a lot with Elon?
是的,我想一个好处是在 xAI 工作,你有机会更多地与 Elon 互动。所以我很幸运地参加了他的一些静修活动,那很有趣。而且他也与人们非常密切地合作,就像人们网上想象的那样,他非常亲力亲为。有两件事。第一,我实际上在查 Elon 转发你的推文。我把它调出来。他谈到你发推说你有一个非常好的语音模式。
Yeah, I guess one benefit is working at xAI, you got a chance to interact more with Elon. So I was very fortunate to get a few retreats from him, and that was quite fun. And he also worked very closely with people, like people imagine online, he's very hands-on. There are two things. One, so I was actually looking up Elon retweeting you. I'll pull it up. He talked about you tweeting that you have a really good voice mode.
我不……不,不,不,是他。
I don't... No, no, no him.
哦,我也做了。但不管怎样,我实际上会私信给你关于语音模式的反馈,因为我当时想,“哇,真不错。”然后我又想,“哦,这太糟了。”但是,嗯,我不知道。
Oh, I also did it. But anyway, I actually would DM you feedback on voice mode because I was like, "Wow, really good." And then I'm like, "Oh, this sucks." But, um, I don't know.
关于语音模式的构建,你有什么想聊的吗?这也是你参与过的团队吗?
Anything you want to talk about about your voice mode building it? Was it a team you worked on as well?
那其实不是我参与的团队,我可能更多是做视频方面的。不,但 Grok 语音实际上非常好。我觉得其中一点是,首先你可以用 2 倍速说话,这很有趣。我听东西都是 2 倍速,所以我也喜欢用 2 倍速说话。而且,我觉得它的打断处理比 Gemini 好。我不知道它现在和 ChatGPT 实时模式比怎么样。但就驾驶体验而言,在我的特斯拉里用 Grok 开车,我觉得体验非常好。
That's actually not a part of the team I worked on. Probably worked more on the video. No. But Grok voice is actually very good. I think one of those things where, first of all, you can speak at 2x speed, which is fun. I listen to 2x, so I like to speak at 2x. But also, I think the interruption handling was better than Gemini. I don't know how it compares to ChatGPT real-time now. But as far as driving is concerned, having Grok in my Tesla and driving, I think it was a really good experience.
是的。他喜欢语音模式,还有那疯狂的传播量——5000 万观看量,就只是说“是的,真相”。
Yeah. He likes voice mode but also the crazy reach — 50 million views are just saying yes truth.
确实。天哪。但它的推出速度真的很快。
True. Oh my god. But it's pretty cool how fast it came out.
另一件事是视频模式的安全方面。有什么有趣的点可以聊吗?一个尖锐的问题。
The other thing is the safety aspect of video mode. Anything interesting to talk about there? Spicy question.
很多国家不允许没有水印的生成式 AI 视频。所以在所有这些国家,Grok 都加了水印,而且很多视频的下架速度也非常快。这是运营社交平台的一部分,但这也很好地迁移到了生成式 AI 这边。
A lot of the countries where they don't allow generative AI videos without watermarks. So in all of those countries, Grok had watermarks, and a lot of the takedowns of the videos were also happening extremely fast. It's part of running a social platform, but it also transfers nicely to the generative side.
你对 SynthID 和其他类型的水印有什么看法?
Do you have a perspective on SynthID versus other kinds of watermarking?
是的,我觉得这些东西会越来越难检测。SynthID 有一个问题,以前只有 Google 在用,现在很多不同的实验室也在采用它。它的局限性在于,技术和论文都是公开的,人们可以逆向工程找出如何去除它。我认为即使它进步了,仍然有可能被逆向工程。如果你感兴趣,可以去 Reddit 看看,有人已经提取出了 Google 应用的具体模式,然后你可以把它应用到任何 Google 生成的图片上,从而逆向去除 SynthID。
Yeah, I guess it's going to be harder and harder to detect these things. SynthID, one thing is previously it was only Google, and now a lot of different labs are also adapting it. Its limitation is that the technology and the paper were out there, and people can reverse engineer how to get rid of it. I think even as it advances, it's still possible to reverse engineer it. If you are interested, you can go on Reddit and people have taken out the exact pattern that Google applies, and then you can apply it onto any Google-generated photo and reverse out the SynthID.
是的。用肉眼判断也越来越难了。我记得几年前还有六根手指之类的问题,非常明显。我现在判断的方法其实是音频。我觉得音频真的很欠缺。除了靠眼睛看,我判断是否是 AI 生成的方法是音频匹配,尤其是 Sora,做得不好。风格都很相似,但有一些小瑕疵。
Yeah. It's also harder and harder to judge by eyes. I remember a couple years ago there were six fingers or something, very obvious. My current tell is actually the audio. I feel like the audio is really lacking. My way to tell if something is AI-generated, outside of having a decent eye, is the audio matchup, especially of Sora, is not great. It's all similar style, but there are minor imperfections.
我觉得关键在于,我对此最直接的参考也是 Ian Goodfellow,因为他做了那个对抗游戏:你有一张斑马的图片,然后改变一个像素,它就变成了熊猫,对吧?这是一个经典的计算机视觉问题。如果你想想这些模型是如何训练的,就像我之前提到的,GAN 在训练过程中。目标是模型生成一张图片,然后有一个判别器来判断图片是否真实。模型被训练得让图片更真实。所以随着模型越来越先进,判断会越来越难。对我个人来说,现在我必须通过判断这些视频是否有逻辑意义来辨别。
I think the point is that my closest reference to this is also Ian Goodfellow, because he did the adversarial game where you have a picture of a zebra, then you change one pixel and it becomes a panda, right? This is a classic computer vision issue. If you think about how these models were trained, like I mentioned before, GAN was in the training process. The objective is the model generates an image and there's a judge to tell if the image is real or not. The model is trained to make the image more real. So as the model becomes more advanced, it's going to be harder and harder. For me personally, now I have to judge by if these videos have logical sense.
是的。不,我也觉得音频太好了,太有录音室质量了,光线太好了,皮肤太清晰了,基本上就是缺乏瑕疵。
Yeah. No, I also like that the audio is too nice, too studio quality, the lighting is too good, the skin is too clear, basically the lack of imperfections.
我们在扩散模型中有好的方法来做推理吗?这是区分视频生成器和世界模型的关键吗?或者我们真的知道如何将其应用于自回归语言模型吗?扩散视频生成和世界模型之间有没有平行关系?
Do we have a good way to do reasoning in diffusion? Is that what separates video generators from world models? Or do we really know how to apply it to autoregressive language models? Is there a parallel for diffusion video generation and world models?
是的,这是个好问题。实际上,我有一个相当大胆的主张。视觉智能实际上主要来自语言。这些视频模型,尤其是现在扩散模型技术更成熟了,每次你看到这些模型有改进,我敢说大部分收益来自语言模型,而不是视频模型本身。真正的扩散模型本身,比如在 Cosmos 中,这些模型通常有两个部分:一个提示重写器或提示上采样部分。在 Cosmos 中,我们使用 Llama 或 Mixtral,而 Cosmos 视频模型本身只有 7B,语言模型提示重写器比它大。提示重写器的任务是接收用户指令,并将其转换为极其详细的视频描述。因为视频扩散模型,我会把它们描述为有点笨,因为它们会逐字地接受输入指令。在训练过程中,我们在创建合成文本对时必须尽可能详细地描述视频。所以这些模型会接受那些指令来生成视频。当你输入用户指令时,用户指令通常非常简单,只说“一只猫”之类的。如果你把“一只猫”输入视频模型,它们会逐字地接受指令,然后可能在一个白色背景上显示一只猫,因为你没有描述背景。猫不会动,因为你没有描述它。它逐字地接受指令,有点笨。而提示重写器实际上是一个更大的模型,一个语言模型,它接收用户指令并扩展它。所以你提到的思考过程就来自那里。如果你看看 GPT-4o 图像生成,你生成一张图像需要 3 分钟。那 3 分钟并不全是像素生成;很多时间花在思考上。所以提示重写现在已经发展到不仅仅是思考,它还可以是一个智能体模型。例如,假设你想生成一张关于今天新闻的图像。那么它很可能会去网上获取今天的新闻,然后处理它们,消化它们,组织布局,然后生成。
Yeah, that's a good question. Actually, I have a pretty big claim. The visual intelligence is actually mostly coming from language. These video models, especially now since the diffusion model technology is more mature, every time you see some improvement on these models, I would say mostly the gain comes from the language model, not from the video model itself. The real diffusion models themselves, in Cosmos for example, typically these models have two parts: there's a prompt rewriter or the prompt upsampler part. In Cosmos, we use Llama or we use Mixtral, and the Cosmos video model itself is only 7B, and the language model prompt rewriter is bigger than that. The prompt writer's task is to take a user instruction and convert it to an extremely detailed description of the video. Because the video diffusion models, I would describe them as kind of dumb, because they take the input instruction literally. In the training process, we have to describe the video as detailed as possible when creating the synthetic text pair. So these models take those kinds of instructions to generate the videos. When you take the user instruction, the user instruction is usually very simple, just say "a cat" or something. If you put "a cat" into the video model, they would take that instruction literally and literally show a cat on maybe a white background because you didn't describe the background. The cat is not moving because you didn't describe it. It takes the instruction quite literally. It's kind of dumb. And the prompt writer is actually a much bigger model, a language model that takes the user instruction and expands it. So the thinking process you mentioned is from there. If you look at GPT-4o image generation, you generate an image in 3 minutes. That 3 minutes is not all pixel generation; a lot of time is spent thinking. So prompt rewriting has now evolved to not only just thinking, it can also be an agentic model. For example, say you want to generate an image of today's news. So it's likely it will go fetch today's news online, then process them, digest them, organize the layout, and generate it.
另一件很有意思的事情是,如果我没弄错的话,它不再是扩散模型了,对吧?是自回归的,还是仍然有扩散?
Another thing quite interesting is, if I'm not mistaken, it's no longer a diffusion model, though, right? Autoregressively, or is there still?
有几种不同的方法。比如 Chai Omni,既然他们说是 Omni,我相信它是一个单一模型。可能类似于语言模型加一个扩散头,语言模型负责思考、调用智能体工具,最后用扩散头生成图像。还有像 Cosmos 这样的方法,语言模型和扩散模型是分开的。也有纯语言模型的方法,比如把图像离散化,然后以离散 token 的形式生成图像。所以方法各有不同。
There are different approaches. For example, like Chai Omni, since they said it's Omni, I believe it's a single model. Maybe it's something like a language model with a diffusion head, or something like the language model does the thinking, does the agent tool calling, and then it would use the diffusion head to generate the image in the end. There were also approaches like Cosmos, where you have a separate language model and separate diffusion models. And there are also purely language model approaches, like you discretize the images and then generate an image as discrete tokens. So there are different approaches.
我看到的一种说法是,这些方法之所以困难,是因为目前我们通过语言模型学习推理的很多好处在于,你基本上是迭代式地生成推理,先有想法,然后基于那个想法得出答案。对吧?所以如果你有一个 Omni 模型再加一个扩散头,你就无法把输出反馈回去继续推理。对吧?所以你无法做到文本-图像-文本-图像这样的循环。你无法对输出进行推理,然后再回到扩散。但我想在新的 Gemini Omni 中,只要你有扩散,应该就能做到。
I would say one of the claims I've seen for why these approaches struggle is because a lot of the benefits for how we currently learn reasoning with language models is you basically iteratively generate reason, you have your thought and then you work on that answer. Right? So if you have an Omni model and then a diffusion head, you can't feed that back in to continue reasoning. Right? So you can't go like text-image-text-image. You can't reason on the output and then go back to diffusion. But I guess in the new Gemini Omni, you would be able to, as long as you have diffusion.
是的。我不确定他们是否有那个流程。但我想在 Omni 范式下肯定是可能的。
Yeah. I'm not sure if they have that process. I guess it's definitely possible in the Omni paradigm.
嗯。
Yeah.
所以如果你考虑传统的多模态语言模型,它们会有一个 ViT 编码器来编码图像。所以如果它们有一个扩散头,就可以生成图像,然后把图像放回 ViT 编码器,编码之后,再对结果进行迭代优化。
So if you think about traditional multimodal language models, they would have a ViT encoder that can encode the image. So if they have a diffusion head, they can generate the image and then put that back into the ViT encoder, encode that, and then do iterative refinement of the result.
嗯,我认为你必须联合训练 ViT 和扩散模型,才能让这个过程合理,否则就会不匹配,或者输入的是垃圾。
Yeah, I think you have to jointly train the ViT and the diffusion to make that somewhat reasonable, because otherwise you're kind of mismatching or feeding in slop.
是的。
Yeah.
我认为这取决于你的训练阶段,也许可以冻结它。不过,回到你之前的观点,我想明确一点:我们确实知道 Nano Banana 和 GPT Image 是带扩散头的自回归语言模型。根据你对 Grok Image 的描述,它并不是,它是端到端的。
I think it depends on your stage of training; you might be able to freeze it. But anyway, also just on your earlier point, I wanted to make explicit: we do know that Nano Banana and GPT Image are autoregressive language models with a diffusion head. As far as I can tell from your description of Grok Image, it is not; it is end-to-end.
我不能,嗯。按你描述的方式,但我觉得方法各有不同,对吧?你一开始说提示词重写是智能的重要组成部分。
I cannot, yeah. The way that you described it, but I think there are different approaches, right? You started off saying prompt rewriter is like a big part of the intelligence.
关于这一点,我觉得每个人都应该试试早期的扩散模型。如果你用过 Stable Diffusion 1 之类的,你会看到提示词像“超高分辨率 4K 这种风格”——天哪,我第一次用的时候,你不能像跟语言模型那样跟它说话,对吧?你的提示词是逗号分隔的。
And even on that, I think everyone should try using an early diffusion model. If you've used Stable Diffusion 1 or whatever, if you've seen the prompts like "ultra high-res 4K this style" — oh my god, the first time I tried one, you don't talk to them like language models, right? Your prompting is very comma-separated.
基本上就是在用数据集里的标签说话,对吧?
Literally talking in the labels that were in the dataset, right?
但基本上,我只是想说明,提示词编写加图像生成,与带扩散头的自回归语言模型是不同的,对吧?它们是不同的东西。
But basically, I'm just trying to make the point that prompt writer and then image is different from autoregressive language model with diffusion head, right? They're different things.
是的,它们不同。
Yes, they're different.
只是想确认这一点。我觉得共同的部分是图像部分。所以很令人惊讶的是,很多改进来自语言、思考和工具调用。
Just wanted to establish that. I'd say the common part is the image part. So it's quite surprising that a lot of the improvement came from the language, the thinking, the tool calling.
我还记得在 Cosmos 里,我生成一只快乐的羊,如果不加任何重写,看起来特别像 CGI。重写之后,看起来就非常漂亮了。我觉得没有任何联合训练。
So I still remember in Cosmos, I generate a happy sheep, and if without any rewriting, it looks so CGI. And after rewriting, it looks so beautiful. I think without any joint training.
嗯,实际上没有任何联合训练,只靠重写就已经好很多了。
Yeah, actually without any joint training, with rewriting it's already much better.
我认为一件非常有趣的事情是,视频智能体——主要是语言模型——会把生成模型(无论是单独的模型还是扩散头)当作工具来调用。这样模型就能迭代优化结果,甚至通过很长的思维链生成更长的内容。这实际上很像人类创作艺术的方式。我们不是直接生成像素,而是迭代地画点什么。我认为通过这个过程,这些模型不仅把扩散当作工具之一,还可以使用传统工具,比如 Photoshop 的图像编辑工具、视频编辑器、ffmpeg 等等,把这些和生成式 AI 技术结合起来作为一套工具,迭代地创作出更高质量、达到生产级别的视频。如果你看看现有的专业创作者,他们不会止步于用这些模型生成一个视频。他们会把视频拿到编辑器里,这里改改那里修修。
I think a very interesting thing that will happen is that video agents, mostly language models, will call these generative models, either a separate model or diffusion head, whatever, as tools. So this model can iteratively refine the results, or even generate longer content through a very long train of thought. It's actually very similar to how humans create art. We don't generate the pixels directly. We iteratively draw something. And I think through this process, these models not only use diffusion as one of the tools, they can also use traditional tools, image editing tools from Photoshop, video editors, ffmpeg, whatever, taking a combination of these and the generative AI technology as a set of tools, and they can iteratively create a much better video for production-grade quality. If you look at existing professional creators, they don't end at generating a video from these models. They will take that video to their editor and edit here and there.
后期制作占了很大比重。有时候视频之所以好,其实不是视频模型的功劳,而是剪辑。
So much post-production. Sometimes actually the reason the video is good is not really the video model, it's actually the editing.
是的。
Yeah.
是的,我们也在做同样的事情。你会喜欢用视频编辑模型吗?
And yes, we also are engaged in the same process as well. Would you love to use a video editing model?
嗯,实际上 Grok Imagine 智能体测试版就是朝那个方向的第一次尝试。
Yeah, actually there's the Grok Imagine agent beta that was the first attempt in that direction.
嗯。
Yeah.
所以我认为过程会是类似的。
So I think the process would be similar.
嗯。你可以让它……没有相关的博客文章。比如生成一分钟的视频,如果用同样的提示词去问视频模型是做不到的,但这个模型会迭代地调用不同的工具来完成。所以这确实很有意思。当我们第一次发布视频编辑模型时,我在 X 上看到有人尝试视频编辑功能,说“把这个视频编辑成一分钟”,但他们并不理解视频编辑是怎么工作的。视频编辑通常只是删除、添加、替换、转场这类操作。但在视频智能体的假设下,这其实是一个合理的请求。所以这些智能体应该能够理解这种长周期的请求,创作长视频。我觉得这非常吸引人,因为它走的是同一条路:先是 AI 辅助编程,比如 Tab 补全和 GitHub Copilot,然后逐渐进化到 Cursor 和 Claude Code,实现完全自动化。所以在 Grok Imagine 智能体模式下,你仍然可以自己动手操作,但随着模型能力的提升,它最终将能完全自动化地完成所有事情。
Yeah. You can ask it to... There's no blog post for it. Maybe generate a one-minute video, which is not possible if you ask the same prompt to video models, but this model will iteratively call different tools to do that. So yeah, this is actually an interesting thing. When we first release a video editing model, I see on X some people try the video editing feature with "edit this video to be one minute," but because they didn't understand how video editing works. Video editing typically is just removal, add, replace, transfer, this kind of thing. But that's actually a valid request under the assumption of video agents. So these agents should be able to understand this kind of long-horizon request, create a long-form video. I think this is really fascinating because it's taking the same direction as first you have AI-assisted coding, like tab completion and GitHub Copilot, and from there you gradually evolve to Cursor and Claude Code where you do things fully automated. So in Grok Imagine agent mode, you can still go in there and do stuff by yourself, but gradually as the model capability increases, it will be able to do everything fully automated.
嗯,我喜欢这个想法。好的。看起来它还在生成中。
Yeah. I like that. Okay. So it looks like it's still generating.
另外我注意到 Grok Image 一直都非常非常快。
Also I did notice Grok Image was always very, very fast.
我不知道你们是否对此进行基准测试,但这只是题外话。当我使用最新的图像生成和 Gemini Nano Banana 时,我经常使用裁剪功能只是为了……
I don't know if this is something you guys benchmark, but this is just a tangent. When I used to use the latest image gen and Gemini Nano Banana, I would often use crop just for...
在 Imagine API 博客文章的某个基准测试中,他们列出了所有速度相关的内容。这主要是蒸馏加推理的组合。
It's in the benchmark somewhere in the Imagine API blog post that they have all the speed things. It's mostly a combination of distillation plus inference.
是的,有很多因素。我们谈到蒸馏。如果说到思考,如果没有思考预算,模型可以思考 3 分钟然后回复你。此外,推理基础设施团队非常有才华,他们能够大幅加速这些模型。
Yeah, there are a bunch of things. We talk about distillation. If you talk about thinking, if you don't have any thinking budget, the model can just think for 3 minutes and then come back to you. Also, the inference infra team was very talented and they were able to accelerate a hell lot of these models.
是的。是的。是的,我的意思是,你知道,我对视频智能体这件事的看法:我在想人们说视频智能体时到底指什么。当你最初告诉我你押注视频智能体或你对视频智能体的愿景时,我有点失望。我当时想,‘哦,你的意思是模型已经到顶了,我们只能做智能体?’但我觉得你不得不这样做,对吧?现在的问题是,模型训练到底能带来多大改变,还是仅仅构建一个更好的框架?就像你说的,模型不需要联合训练。你只需拿一个现成的前沿推理模型,套上一个框架,把 Grok 作为工具给它,就完成了。这就是你的视频智能体。这看起来并不令人满意。显然,你可以联合训练并获得几个百分点的性能提升,但如果你核心主张是视频或生成式媒体的主要阿尔法来自语言智能,而不是图像扩散或视频扩散,那么这就是未来。主要就是等待。
Yeah. Yeah. Yeah, I mean, you know, my comment on the video agents things: I'm trying to figure out when people say video agents. When you initially told me about your bet on video agents or your vision for video agents, I was a little bit disappointed. I was like, 'Oh, you mean like models are tapped out now, we have to do agents?' But I think you have to, right? The question now is how much model training is really going to make a difference versus just building a better harness? Like you said, the models don't have to be jointly trained. You can just take an off-the-shelf frontier reasoning model, slap it on a harness, give it Grok as a tool, that's it. That's your video agent. Doesn't seem super satisfying. Obviously, you can co-train and get some more percentage points of proper performance, but if your central claim is that the majority of video or generative media alpha is actually coming from language intelligence and not image diffusion or video diffusion, then that is the future. It's primarily just wait.
如果你回顾一下这个例子,你知道,它生成了帧。抱歉打断一下,但你知道,它一直在说,‘好的,我要开始把这些帧拼接在一起了。’它正在使用 ffmpeg……
If you pop back at the example, you know, it generated frames. Sorry to interrupt, but you know, it's been saying like, 'Okay, I'm going to start stitching these frames together.' It's using ffmpeg using...
Gemini Image Pro 也是这么做的,对吧?它也只在后台编写代码,然后拼接,对最终输出进行图像处理。
This is what Gemini Image Pro is also doing, right? It's also just writing code in the background and then just stitching, doing an image pass on the final output.
对于那些只想训练模型的人来说,这感觉不太令人满意。
It feels dissatisfying for the people who want to just train models.
这很有趣,对吧?也有点令人兴奋。就像你之前提到的,很多收益并不来自视频本身。我认为在语言模型领域也能看到这一点。对吧?Anthropic 非常擅长编码,但他们的多模态不是最好的。他们有基本的 PDF 输入,但你知道,他们在图像、视频、音频处理的质量上明显存在差距,然而智能水平却是顶级的。其他实验室如 OpenAI、XAI 可以添加模态,但并没有解锁疯狂的能力。所以这很有趣。是的,看到视频模型的能力提升实际上来自语言模型变得更智能,这很有趣。我认为视频智能体可以解锁比你想象的更多东西。所以有几点。一点是当我们提示这些模型时,大多数人其实并不擅长提示。实际上,语言模型更知道如何提示 AI 模型。AI 模型更了解 AI 模型。所以如果你联合训练这些模型,也许模型会更知道如何提示每个模型。不同模型可能不同。另一点是,这可能不像生成几个片段然后用 ffmpeg 拼接那么简单。在这个过程中可能会出现更多的图像和视频编辑工具。比如说,如果你想在这个时间步精确添加一块文本,视频模型可能无法非常精确地理解这个意图,但使用确定性工具是可能的。视频智能体可以使用各种工具。所以你不需要把所有能力都放进转换模型本身。
It's interesting, right? It's also somewhat exciting. Like you brought up earlier, a lot of the gains don't come as much from the video. I think you can see that in the language model space too. Right? Anthropic is very, very good at coding, their multimodal is not the best. They have basic input PDF, but you know, there's clearly a disconnect in the quality of their image, video, audio processing, yet intelligence is very top tier. Other labs like OpenAI, XAI, you can add modalities, but it's not like they're unlocking crazy capabilities. So it's interesting. Yeah, it's interesting to see that the video models' capability increase actually comes from language models being more intelligent. I think video agents can unlock more stuff than you might imagine. So there are a few things. One thing is when we are prompting these models, most people are actually not very good at prompting. Actually, language models have a better sense of how to prompt AI models. AI models know AI models better. So if you jointly train these models, maybe the model has a better sense of how to prompt each model. Different models might be different. Another thing is it might not be as simple as just generating a few clips and slapping them together using ffmpeg. There might be more image and video editing tools appearing in this process. Say if you want to exactly add a blob of text at this time step, the video models might not get that intention very precisely, but these are possible using deterministic tools. The video agents can use all sorts of tools. So you don't have to put all of the capabilities into the transition model itself.
是的,我认为这非常正确。所以不管怎样,我认为你是对的。我认为这将是一个大类别。我猜你预测未来一年视频领域将全是这个。
Yeah, I think that's very true. So for what it's worth, I think you're right. I think that this will be a big category. I think probably you are predicting like the next one year in video is going to be all this.
你对这些东西何时加速有预测吗?
Do you have a time prediction for when this stuff ramps up?
我是说,它们已经开始了。
I mean, they already started.
是吗?那太好了。我觉得最后一个只是更长。
Is it? So it's so good. I think the last one's just longer.
它没给我一分钟。你给了我 36 秒,但你知道,我们现在感觉到了吗?会有转折点吗?你有什么时间线预测吗?
It didn't give me a minute. You gave me 36 seconds, but you know, are we feeling it now? Is there going to be an inflection? Any timeline predictions you want to make?
我猜到今年年底,这将大受欢迎。所以转折点将是视频智能体生成的视频达到生产级质量的时候。它可以被展示并分发到广告中。一旦发生这种情况,我认为企业将为视频模型投入更多预算,因为智能体本质上比视频模型本身更昂贵。它们进行迭代过程,生成许多许多变体。但一旦这些模型通过这个可用性门槛,我认为之后将是指数级增长。
I guess by the end of this year, this is going to be a big hit. So the inflection point will be when the videos generated by video agents can get to production-grade qualities. It can be presented and distributed in ads. Once that happens, I think the enterprise will have much more budget for video models because the agents are inherently more expensive than the video models themselves. They do this iterative process, they generate many, many variations. But once these models pass this usability threshold, I think it's going to be exponential growth beyond that.
是的,我现在就会基于这个投资一家公司。所以我认为你是对的。有一件事让我惊讶,回顾过去一个小时的对话,我认为你热衷于世界模型和为了视频生成而进行视频生成。我认为还有很多其他世界模型的人。我们采访过很多:General Intuition、李飞飞那些人,还有 Moonream,我想我跟你说过。Moon Lake。
Yeah, I would fund a company right now based on this. So I think you're right. One thing I'm surprised about, reflecting on the whole past hour or so conversation, I think you're into world models and video generation for video generation's sake. I think there are a lot of other world models people. We've interviewed a lot of them: General Intuition, Fei-Fei Li, and those guys, and Moonream, which I think I told you about. Moon Lake.
Moon Lake。
Moon Lake.
我老说成 Moonream。该死。Moon Lake。他们很多人实际上说机器人是终极目标,比如具身机器人。你想要实时,想要交互,就是要与物理世界互动。你对此并不那么关心。
I keep saying Moonream. God damn it. Moon Lake. A lot of them actually say robotics is the endgame, like embodied robotics. You want real-time, you want interactive, it is to interact with the physical world. You're not that concerned about it.
我认为机器人肯定会是其中的重要部分。我猜这个过程可能会自然发生。所以我对机器人的预测是,物理 AI 的问题可能不需要实际在现实世界中就能解决。所以它可能通过一个具有极强视频能力的视频语言模型来解决。记得我们讨论过实时交互长程视频,一旦这些模型只在屏幕录制和电脑屏幕上训练。一旦这些模型能够使用计算机并极其准确地理解计算机的未来状态,机器人可能成为非常强大的 AI 可以使用的工具之一。
I think robotics will be a big part of it for sure. I guess the process might happen naturally. So my prediction on robotics is that the problem of physical AI might be solved without actually needing to be in the real world. So it might get solved by a video LM with very strong video capability. Remember we talked about real-time interactive long-horizon video once these models are just training on screen recordings and computer screens. Once these models can use computers and understand the future state of computers extremely well, the robots might be one of the tools that a very powerful AI can use.
所以强大的 AI 可能自然就能控制物理具身。
So the powerful AI might just be able to control the physical embodiment naturally.
我完全同意。
I see that for sure.
酷。我知道时间快到了。你还有一个劲爆话题,就是你为什么离开 XAI。
Cool. I know we are coming up on time. You had one more spicy topic which is why you left XAI.
对我来说,有很多研究你想做但在公司里做不了,而且公司的优先级和目标通常变化很快。XAI 也是如此。所以现在是时候去做一些我想做的研究了,尤其是在语言模型方面,这在 XAI 做不了。
For me, there's a lot of research you want to do that you cannot do at a company, and also the priorities and objectives of the company typically can change very fast. It's also the same for XAI. So now is the time to do some research I want to do, especially more on the language model side, which I cannot do at XAI.
哦,好吧。是啊,因为你基本上是在经历了从计算机视觉到世界模型、视频生成,再到如今专注于语言模型的整个转型后离开的。但似乎你已经描述了这一切是如何联系在一起的。不过我不太明白你说的专注于语言模型是什么意思?
Oh, okay. Yeah, because you're basically leaving after this whole transition from computer vision to world models, video generation, to now focusing on LMs. But it seems like you've described how it all ties together. But I don't know what do you mean by focusing on LMs?
我意识到一个事实:视频模型,即使在一开始,收益可能来自扩散技术的改进,但现在实际上大部分收益来自语言模型本身。
I realized the fact that the video models, even in the beginning, the gain might come from improvement on diffusion technology, but this is a point where actually most of the gain comes from the language models themselves.
这对任何在生成式媒体领域投入职业生涯的人来说,都是一颗巨大的黑色药丸。
That's a huge black pill for anyone who has spent a career in generative media.
我的意思是,这是一个极端的观点,对吧?你肯定还是需要两者兼顾。
I mean, that's an extreme view, right? You still definitely need a bit of both.
是的。
Yeah.
只是现在语言模型方面有更紧迫、更有影响力的工作要做。
There's just more pressing impactful work to do now on the language model side.
你有什么类似的预测吗?你预测视频智能体?我认为你在语言方面是对的。未来一年你在关注什么?
Do you have any similar predictions? You predict the video agents? I think you will be right on the language side. What are you looking for in the next one year?
我认为一件非常有趣且可能很快发生的事情是,语言模型将具备上下文感知能力,并能管理自己的上下文。
I think one thing pretty interesting that might happen soon is that language models will be context-aware and manage their own context.
嗯。
Yeah.
从视频模型这边来看,我们一直受困于长时域问题。我们想生成越来越长的视频,并尝试通过各种方式解决上下文长度问题。一种方法是暴力训练更长的上下文长度;另一种是更好地管理上下文。我认为语言模型很快也会发生同样的事情。例如,长上下文模型并不知道自己的上下文长度有多长。一旦达到 80% 左右,自动上下文压缩就会被触发,而模型在工作时并不知道这一点。也许让模型知道“哦,我快接近 80% 了”是件好事。另外,还有一些非常有趣的事情:比如在 OpenClaw 或类似你的系统中,每次你输入内容时,当前本地时间会自动附加到你的消息中。这样模型实际上就知道了当前时间。这让模型具备了时间感知能力。同样,在工具调用中,很多中间工具调用结果会被自动修剪。所以有上下文移除、上下文添加和上下文压缩。所有这些都来自智能体框架本身,但根据我们的经验,所有模型的启发式工程都会内化到模型自身。我想这是一个非常值得探索的方向。所以无限上下文?也许不是,但这很有趣。
From the video model side, we've been suffering from the long horizon issue. We want to generate video longer and longer, and we've been trying to solve the context length issues through various ways. One thing is just brute forcing train longer context length; another is to manage the context better. I think the same thing in language models is also going to happen soon. For example, the longish models are not aware of how long their own context length is. Once they hit like 80%, the automatic context compression gets triggered, and the model is not aware of that when it's working. Maybe it's good for the models to know, 'Oh, I'm approaching 80% or something.' Also, something pretty interesting: for example, in OpenClaw or like you, every time you type in something, the current local time is automatically attached to your message. So the model actually knows what time it is. This is making the model time-aware. Also, in tool calling, a lot of the intermediate tool call results are automatically pruned. So there's context removal, context addition, and context compaction. All of these are from the harnesses themselves, but from our experience, the heuristic engineering of all the models gets absorbed into the models themselves. I guess that's something very interesting to explore. So infinite context? Maybe no, but it's interesting.
这属于记忆和持续学习的范畴。
It's in the space of memory and continual learning.
我不知道,这也属于智能体框架使用的范畴,对吧?
I don't know, it's also in the space of agent harness use, right?
他是说不想在框架里做这个,对吧?
He's saying he doesn't want to do it in a harness, right?
不,不,但模型也是在使用了框架的数据上训练的,对吧?所以有些东西是隐式渗透进去的。语言模型的后训练部分,就是在编码框架中使用它,在这种情况下,子智能体何时生成?上下文何时发生?它不像“你有这么多 token 窗口”那样明确,我不知道你是否希望它明确,但它确实在某种程度上渗透进去了。
No, no, but models are also being trained on data using harnesses, right? So some of it is implicitly leaking in. Part of that post-training of language models is okay using it in coding harnesses, in which case, when are sub-agents spawned? When is context going to happen? It's not explicit like 'you have this much token window,' which I don't know if you want it to be, but it's somewhat leaking in there.
我在想象,如果模型能够访问智能体框架本身的全部代码,并且可以随意修改它,会怎样。假设智能体框架足够短,你可以直接把它放在系统提示的上下文长度里,然后模型说:“当我想生成未来的自己时,我可以修改智能体框架。”例如,如果智能体框架可以被修改,或者当我在阅读长文档时,我可以选择分块阅读整个文档,然后回来把摘要拼在一起,或者我可以只读前 200 行,丢弃其余部分。所有种类的选择,如果都能由模型自己做出,那么看到模型在测试时在线自我编程,可能会非常有趣。
I'm imagining what if the model has access to the whole code of the agent harness itself and can modify it however it wants. Say if the agent harness is short enough, you can just put it in the context length in the system prompt, and then the model says, 'When I want to spawn a future version of myself, I can modify the agent harness.' For example, if the agent harness can be modified, or when I'm reading a long document, I can choose to read the whole thing in chunks and come back, smash the summary together, or I can just read the first 200 lines and discard the rest. All kinds of choices, if they can be made by the model itself, it might be very interesting to see that the model can program itself online at test time.
嗯。所以自修改框架也是 OpenClaw 和 Pi 的一部分,但我认为这方面还有很多工作要做。非常酷。我有点好奇:你是一个大实验室的成员,对吧?大实验室的研究员有一条职业路径:你训练模型,获得更多算力,训练更好的模型,一直这样下去。在某种程度上,我觉得你是在选择退出这条路。如果我是你,我会觉得“哦,我认为这有点职业风险”,你明白我的意思吗?
Yeah. So the self-modifying harness is also part of OpenClaw and Pi, but I think there's a lot more work to do there. Very cool. I think part of me is curious: you are part of a big lab, right? And there's this career path of a researcher at a big lab, which is you train models, you get more compute, you train better models, you keep going. And somewhat I feel like you're opting out of that. If I were you, I'd be like, 'Oh, I think this is a bit of a career risk,' you know what I mean?
我没什么可评论的,只能说你的信念非常坚定。我认为很多处于你位置的人不会做你所做的事。
I don't have any comment apart from you're very strongly convicted. I think that a lot of people in your shoes would not be doing what you did.
嗯。说到我的职业生涯,如果回头看,有很多巨大的转变。所以 10 年前,我和 ResNet 的作者 Shan、John 和 Jensen 一起做研究。那时候的研究完全不同,主要是计算类的,比如图像识别、目标检测、目标跟踪。我当时也在做神经网络压缩,和现在的知识蒸馏很不一样。那时我想成为一名教授。当我申请博士时,我已经在顶级会议上发表了几篇第一作者论文。所以我自信地申请了顶尖学校,但结果我被所有顶尖博士项目拒绝了。所以我不得不进入工业界。那时我在 Facebook AI Research,由……领导。
Yeah. Speaking of my career, if I look back, there were a lot of huge transitions. So 10 years ago, I was doing research with the ResNet authors Shan, John, and Jensen. At that time, the research was completely different. It was mostly computation like image recognition, object detection, object tracking. I was also doing neural net compression at that time. It was quite different from knowledge distillation these days. At that time, I wanted to be a professor. When I applied for PhD, I already had a few first-author papers at top conferences. So I confidently applied to the top schools, but it turned out I got rejected by all of the top PhD programs. So I had to go to industry. At that time, I was at Facebook AI Research led by...
我想谈谈 VJA,但那是另一回事。总之,我们可以下次再聊。你那时转向了自监督学习,这和我当时在计算机视觉领域做的工作很不一样。
I want to talk about VJA but it's different. Anyways, yeah, we can leave it for another time. You switched to self-supervised learning at that time. It was quite different from what I was doing in computer vision.
是的。之后就是 Nvidia Cosmos。所以我意识到 Scaling 极其重要。在 Nvidia,我主要专注于 Scaling。
Yeah. And after that, it's Nvidia Cosmos. So I realized scaling up was extremely important. So at Nvidia, I was mainly focusing on scaling.
一方面是 Scaling 视频描述模型到几十亿参数,另一方面是我参与开发了 Megatron,这是第一个开源框架,能以 40% 的 MFU 高效训练百亿甚至万亿参数级别的模型。后来转到 xAI,我尝试在更大的算力规模上工作。回顾这段经历,我其实做过很多不同的事情。在机器学习领域,切换方向其实更容易。很多人可能觉得,“我做计算机视觉就得一直做计算机视觉,不能转到语言方向。”但根据我在英伟达的经验,我既做过语言模型也做过视频模型,事实并非如此。训练大模型的许多核心原则是相通的。对我来说,目前视频模型的瓶颈其实是语言部分,也就是智能体,所以我想在这方面多做一些工作。这有点挑战,但我觉得不算巨大的跨越。
One thing is scaling the video description models to a few billion parameters. Another thing is I worked on Megatron, which was the first open-source framework to train these models at very large scales, like 100 billion parameters to even trillions of parameters efficiently at 40% MFU. And then switching to xAI, I was trying to work on even larger compute scale. Looking at this trajectory, I actually worked on a lot of different things. Within ML, it's actually easier to switch. A lot of people might think, 'Oh, I work on computer vision, I always have to work on computer vision and cannot switch to language.' But from my experience at Nvidia, I worked on both language models and video models. It's actually not the case. A lot of the core principles of how to train large models are largely the same. For me, right now the bottleneck for video models is actually the language part, the agent, which is why I want to work more on that. It's a bit of a challenge, but I don't think it's a huge jump.
是啊,佩服你。我觉得你在这方面很有远见。我们想聊的大概都聊完了。你非常慷慨地分享了时间,能分享这些真的很好。我们不用通过 xAI 来审核所有内容,而且我觉得没给你惹麻烦。相比你在发布中看到的,xAI 还有很多好东西,对吧?你意识不到它还有多少层次。请多做播客。总之,感谢你的分享,非常友好。我还想听你多说一些。你即将开启下一阶段,虽然还没公布下一步做什么,但显然你在这条路上有更多的愿景和雄心。我觉得你基本上是在梯度下降,朝着你的最终形态前进。
Yeah, kudos to you. I think you have a lot of strong vision there. I think that was mostly everything we wanted to cover. You've been very generous with your time. It's really nice that you are able to share all these things. We don't have to go through xAI to clear everything, but I think we didn't get you in trouble. It's a lot of good stuff about xAI compared to what you just see in the releases, right? You don't realize how many more levels there are to it. Please do more podcasts. Anyway, thank you for sharing. It's been very kind. I want to hear more from you. You are going to embark on your next phase. You haven't announced what you're doing next, but clearly you have more vision and more ambition on this path. I think you're basically kind of gradient-descenting to whatever your final form is.
谢谢。我很快会分享更多关于下一章的内容。
Thank you. I'll share more about my next chapter soon.
好的。谢谢你来做客。
Okay. Thank you for having me.