Runway 的超越视频之赌:世界模型、机器人与神经操作系统——Anastasis Germanidis

Runway's Bet Beyond Video: World Models, Robotics, and the Neural OS — Anastasis Germanidis

阿纳斯塔西斯·杰尔马尼迪斯 Anastasis Germanidis · Latent Space · 2026-09-25 · 约 98 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Runway 联合创始人 Anastasis Germanidis 探讨公司从生成式视频到世界模型、机器人以及全神经操作系统的演进。

Runway co-founder Anastasis Germanidis discusses the company's evolution from generative video to world models, robotics, and a fully neural operating system.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 40)

全文 · Full transcript(中英对照)

界面的未来 The Future of Interfaces

Anastasis

实际上,像界面世界模型这样的东西的终极目标就是拥有一个完全神经化的操作系统。所以我认为 Andrej Karpathy 很久以前就写过这方面的内容。但对我来说,有点奇怪的是,比如我们今天与语言模型的交互,你有一个基本上可以和你谈论任何事情的 LM,你可以把对话引向任何方向,它非常通用,所以它可以解决所有不同的任务,但你通过一个非常僵化的界面与它交互。所以对我来说,界面本身变得可学习并成为整个循环的一部分只是时间问题,就像你不仅仅是在端到端地交付一个应用,这意味着你交付语言模型,但你也交付渲染和像素,而这也是一个可学习的组件。

Effectively, the end game of something like interface world models is you have a fully neural operating system. So I think Andrej Karpathy has written about that quite a while back. But to me, it's a bit odd that we have, for example, with an interaction with an LM of today, you have this LM that can basically talk to you about anything, you can take the conversation any direction, it's very general, so it can solve all those different tasks, but you interact with it through a very rigid interface. And so to me, it's just a matter of time before the interface itself becomes learnable and becomes part of the whole loop of like you're not just delivering an application end to end, and that means you're delivering the language model but you're also delivering the render and the pixels, and that's also a learnable component.

赞助商信息 Sponsor Message

Host

在我们进入今天的节目之前,我有一个小小的消息要告诉听众。谢谢。如果你们没有选择点击并收听我们的内容,我们就无法为你们带来你们如此明确想要的 AI 工程、科学和娱乐内容。几乎每天都有赞助商找我们。但幸运的是,你们中有足够多的人实际上订阅了我们,使这一切在没有广告的情况下也能持续下去,我们希望保持这种状态。但我有一个小小的请求要拜托大家。你能做的最强大、完全免费的一件事就是点击那个订阅按钮。这是我唯一会要求你做的事。这对我和我的团队来说意味着一切,他们每周都在努力把 Inspace 带给你。如果你这样做,我向你保证,我们永远不会停止努力,让节目变得更好。现在,让我们开始吧。

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring Latent Space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it.

介绍Anastasis Introducing Anastasis

Host

好的,我们在这里和来自 Runway 的 Anastasis 在一起,我和 Vibhu 在演播室。欢迎。

Okay, we're here with Anastasis from Runway with me and Vibhu in the studio. Welcome.

Anastasis

很高兴来到这里。

Good to be here.

Host

祝贺你在 Runway 取得的所有成功和进步。你们正在世界各地开设办公室。你最初创业时预见到这一点了吗?

Congrats on all your success and progress with Runway. You're opening offices all over the world. Did you envision this when you first started out?

Anastasis

不完全是。我想即使在我们刚开始的时候,我们也有一个想法,那就是这更多是一个时间问题,而不是是否会发生的问题。我们看到 2016、2017 年的早期生成模型,并据此推断,假设分辨率质量随时间可预测地提高,那么总有一天大部分内容都会被生成。这也许是 Runway 最初的论点:由于这些生成模型的出现,我们需要重新思考创意工具是如何制作的。当我们构建生成模型背后的研究时,逐渐清楚的是,它们的用途远不止于此。

Not quite. I think even when we started, we had this idea that it was more a matter of when, not if. We were seeing the early generative models of 2016, 2017 and just extrapolating, assuming resolution quality increases predictably over time, there's going to be a point where most of content will be generated. And that was maybe the initial thesis of Runway: we will need, as a result of those generative models, to rethink how creative tools are made. And as we built out the research behind our generative models, it then became clear that they were useful far beyond that as well.

Host

现在有了真实世界的东西和我们将稍后讨论的世界模型,这一点更加明显了。我只是有点好奇,你是如何从像 ZDOC 和计算机视觉这样的背景进入 Runway 的。带我们回到早期与 Chris 以及你创始团队中其他人的对话。

And it is more obvious now with the real world stuff and the world models that we'll talk about later. I'm just kind of curious how you go from a background in like ZDOC and computer vision into Runway. Take us back to that early conversations with Chris and whoever else is on your founding team.

Anastasis

我一直分裂在这两个世界之间。一个是我有自己的艺术实践。我长期以来一直在做很多互动艺术。另一方面,我在初创公司工作,在不同的公司担任 ML 工程师和后端工程师。我一直对编码和计算感兴趣,尤其对模拟感兴趣,并把它带回了我的早期艺术作品中。同时,我也感兴趣……

I was always split into those two worlds. One was, I had my own art practice. I was making a lot of interactive art for a long time. And then on the other side, I was working in startups and I was working as an ML engineer, as a backend engineer at different companies. I've always been interested in coding and computation, and especially interested in simulation, and brought it back into my early artwork as well. And at the same time, I was interested...

Host

个人网站上有一些,对吧?

Personal site has a few, right?

Anastasis

嗯……

Um...

Host

有没有一个我们应该调出来的,以防有什么东西像是……我只是喜欢回忆往事。

Is there one we should pull up, just in case there's something that's like... I just like to go down memory lane.

Anastasis

好的,这是什么?所以这是我做的一个项目,我想是在 2015 年,我构建了这个软件,它会向画廊空间里的人们发出语音指令。所以它基本上会协调人与人之间的互动。它会先给你一个身份,比如你是一名建筑师,你 30 岁,你喜欢运动。然后它会把你和另一个人配对。你们会有这种完全生成的互动。显然,当时的语言模型还没有那么先进。所以它是用一些模板和一些马尔可夫链生成的文本混合而成的。它只是完全模拟了画廊空间里每个人之间的闲聊对话。所以我一直非常着迷于生成模型和当时正在进行的早期机器学习工作,但与此同时,还有一条独立的线索,那就是模拟及其意义,比如通过创建这些非常简单的互动和行为模型,我们能了解人类什么。

Okay, what is this? So this was a project that I made, I think back in 2015, where I built this software that would give voice instructions to people in a gallery space. So it would basically coordinate interactions between people. So it would first give you an identity, like you're an architect, you're 30 years old, and you like sports. And then it would match you with another person. You have this completely generated interaction. Obviously, language models were not quite there at the time. So it was made, it was a mix of some templates and some like Markov chain generated text. And it would just completely simulate this small talk conversations between everyone in the gallery space. So I was always very fascinated on the one end with generative models and the early machine learning work that was happening at that time, but at the same time, there was this separate thread of simulation and what it means, like what can we learn about humans by creating those very simple models of their interactions and their behavior.

Host

是你生成那些提示词还是 30 岁之类的?是你生成的吗?你是怎么……

Did you generate the prompts or the 30-year-old whatever? Were you generating them? How did you...

Anastasis

没错。所以程序会从……生成那些。是的。很多都是 ML lib 风格的,就像你有不同职业的列表、不同性格类型的列表、不同年龄的列表等等。然后它只是把这些东西组合在一起。然后也许我们下一个要讲的项目是 Uncanny Valley、Uncanny Road,这是我和我的两位联合创始人之一 Chris 一起构建的最早的项目之一。这个项目使用了 Pix2Pix HD,这是 Nvidia 在 2016 或 2017 年发布的早期图像到图像模型之一,它是一个模型,可以获取场景的语义图,然后生成逼真的,我们称之为输出。显然是非常早期的阶段,所以输出的保真度不是很高,但我认为它是第一个以 1K 分辨率生成的图像生成模型,而且它完全是在自动驾驶数据集上训练的。所以它支持的语义类别只有你在路上会遇到的东西,比如行人、交通标志、自行车、红绿灯。所以这实际上是我们最早的迹象之一,我们构建了这个,人们用它制作了所有这些非常超现实的图像,比如一百万行人或一百万交通标志或巨大的人类。这表明你可以拿一个在非常无聊的数据集上训练的模型——基本上就是路上不会发生太多有趣的事情——然后你可以重新利用它,让它非常偏离分布,做出艺术上引人注目的东西。这在某种程度上总结了 Runway 的论点:你可以拿同样的生成模型,如果你从另一个角度看它们,如果你围绕它们构建有趣的工具,并把它们交给艺术家,他们会做出你意想不到的事情。

Exactly. So the program would just generate those from... Yeah. A lot of it would be kind of ML lib style of just, you know, you have lists of different professions, lists of different personality types, list of different ages, things like that. And then it would just combine those things together. And then maybe the next project we go is Uncanny Valley, Uncanny Road, which was one of the first projects that we built with one of my two co-founders, Chris. This was taking Pix2Pix HD, which was one of the early image-to-image models that Nvidia released back in 2016 or 2017, and it was a model that would take a semantic map of a scene and then generate a photorealistic, let's call it, output. Obviously very early days, so it was not very high fidelity outputs, but I think it was the first image generation model that would generate at 1K resolution, and it was all trained on self-driving data sets. So the semantic categories it would support were only things you would encounter on the road, so it would be pedestrians, traffic signs, bikes, stop lights. And so that was actually one of our first indications that we built this and people were making all this very surreal imagery of a million pedestrians or a million traffic signs or gigantic humans. And it was an indication that you could take a model that was trained on this very boring data set, essentially, of like not that many interesting things happen when you're on the road, and then you can repurpose it and go very out of distribution and make something that was artistically compelling. And that was, it's a summary of the thesis of Runway in some ways, like you can take the same generative models and if you look at them from another direction, if you build interesting tools around them and you give them to artists, they're going to do things that you don't expect.

Host

非常酷。我喜欢它的用户体验。

Very cool. I like the UX of it.

从驾驶模拟数据到生成媒体 From Driving Simulator Data to Generative Media

Host

基本上,你只是得到一块空白画布,随便拖拽,随便操作。而在另一个里,你看到每个人都戴着有线耳机,那是非常苹果广告的标志。

Basically, you're just given an empty canvas, drag whatever, do whatever. And in the other one, you see everyone with wired headphones, that's a sign that it's very Apple ads.

Host

带我们回到今天。你在 Runway 已经做了七年。我们是怎么走到这一步的?我们怎么从驾驶模拟器数据走到这一切的?你几乎覆盖了整个生成式媒体栈。

Take us to today. You've been doing this for seven years at Runway. How have we got to this? How do we go from driving simulator data to all this? And you kind of cover the whole stack of generative media.

Anastasis

有意思的是,我们几乎又回来了,兜了一圈。我们现在正把我们的模型,超越创意工具,应用到现实世界场景中。但这是一段漫长的旅程。很早我们就意识到,Runway 的第一个版本是一种轻松使用当时所有开源模型(比如 pix2pix)并交给艺术家的方式。最初的想法是:如果你不是机器学习工程师,那些模型太难用了。把它们交给艺术家会怎样?很快我们意识到需要在 Runway 内部建立一个研究组织,那大概是在第一年。当时很多任务都是关于当时的图像生成模型、视频生成模型——大多数时候几乎没有视频生成——但它们还没到可以产品化并融入创意工作流的工具中的程度。所以我们需要推动研究的前沿。所以 Runway 研究的前四年几乎是在幕后进行的,直到 2022 年出现了潜在扩散、DALL·E 2,出现了阶跃函数式的变化。你们可能记得——

I mean, interestingly, we're almost back and we're full circle. We're now applying our models and kind of beyond creative tools into real world scenarios. But it was a long journey. It was very early on we realized that the first version of Runway was a way to easily use all the open source models of the day, things like pix2pix, and give them to artists. That was the initial idea: those models are too difficult to use if you're not a machine learning engineer. Like what happens when you give them to artists? Very quickly we realized we needed to build a research org inside of Runway, and that happened maybe on year one. And a lot of the mandate there was the image generation models of the time, the video generation models most of the time—or there were barely any video generations all of the time—but they were not quite there where they could be productionized and brought into tools that would be part of creative workflows. So we needed to push the frontier of the research. And so maybe the first four years of Runway research was almost happening in the background until there was a moment in 2022 with latent diffusion, with DALL·E 2, where there was that step function change. And you guys maybe remember—

Host

我创办 Latent Space 基本上就是因为潜在扩散和稳定扩散。

I started Latent Space because of basically latent diffusion and stable diffusion.

Anastasis

因为我当时想,哇,这不仅可行,实际上在消费级硬件上也能做到。

Because I was like wow this is not only feasible. It is actually doable on consumer hardware.

Host

没错。是的。

Exactly. Yeah.

Host

我觉得变化也巨大。就像我学 pix2pix 时,这是机器学习入门,TensorFlow、Jupyter、Google Colab 笔记本就像这样,然后突然出现了阶跃函数式的变化,你知道,随着扩散等等。自那以后还有其他的吗?就像早期扩散到这里的清晰例子。关键技术或研究还有其他变化吗?

I think the delta is also huge. Like I learned pix2pix like this was intro to ML the TensorFlow like Jupyter Google Colab notebooks are like this and then you have a sudden step function change you know with diffusion whatnot. Any other ones since that like there were clear examples of what early diffusion were to get to here. Any other changes in key technology or research?

Anastasis

从 2018 年我们开始到 2022 年之间。所以我们在 Runway 早期做的工作之一是解决分割——图像和视频分割。这是一个非常重要的问题,因为大多数视觉特效本质上涉及分离主体。转描是一个极其手工的过程。没人喜欢做那个。所以 Runway 早期很多时间都在构建这个叫做绿幕的工具,很长一段时间它是人们使用 Runway 的主要功能。它最终被用在《瞬息全宇宙》和其他一些高知名度电影和剧集中。但很长一段时间,Runway 本质上是一个后期制作工具,直到扩散生成的 Gen-1、Gen-2 出现。

Between 2018 when we started and 2022. So one of the early work that we did in Runway was solving segmentation—image and video segmentation. It was a very important problem because most VFX involves essentially separating subjects. Rotoscoping is an extremely manual process. Nobody enjoys doing that. And so a lot of the early days of Runway was building this tool that was called Green Screen, and it was for a long time the main thing that people were using Runway for. It ended up being used in Everything Everywhere All at Once and a bunch of other high visibility films and series. But that was essentially Runway for a long time—was a post-production tool—until late in diffusion generated Gen-1, Gen-2 happened.

Host

酷。我们越过那个时刻吧。你们走了很长的路,开始发布自己的模型。也许也描述一下那段旅程。

Cool. I mean let's go past that moment. You've come a long way that you started releasing your own models. Maybe describe that journey as well.

Anastasis

是的,所以我们在 2022 年中期到了一个节点,当时很明显我们在相当小的算力规模下做研究,并且很明显缩放定律会像应用于语言生成一样应用于图像和视频生成。所以我们下了一个大赌注,当时我们签了协议要建一个一千块 A100 的集群,当时我们是 B 轮创业公司——那可能是一个几乎有点不理性的决定——但我们真的相信,如果我们大规模训练视频模型,最终会得到一个很棒的模型。当时的目标——我们在 2022 年秋天设定了目标——是视频的潜在扩散稳定扩散时刻会是什么样。当时最好的模型叫做 CogVideo。它是早期的视频模型之一,分辨率非常 256x256,质量非常不高。所以我们决定要建这个集群,我们要投资构建自己的视频模型。在训练 Gen-1 时,很明显要达到完全——我们想构建文本到视频。但我们很清楚,一个更容易的起点是从视频到视频开始,因为当你有更强的条件时,重新风格化现有视频比从头生成视频基本上是一个更容易的问题。所以我们首先在 2023 年 1 月发布了 Gen-1。

Yeah, so we got at a point in mid 2022 when it became clear that we're doing research at a fairly small scale of compute and it became clear that scaling laws would apply to image and video gen in the same way that we're applying to language generation. So we made a big bet and at the time we signed this deal to build a cluster of a thousand A100s, which at the time we were a Series B startup—that was an almost slightly irrational decision maybe—but we really believed that if we trained a video model at the large scale we would get a great model at the end. And at the time the goal—we set the goal around fall of 2022—of what the latent diffusion stable diffusion moment looked like for video. And at the time the best model of the time was called CogVideo. It was one of the early video models, was very 256x256 resolution, very not high quality. And so we decided we're going to build out this cluster and we're going to just invest in building out our own video model. It became clear as we're training Gen-1 that it was difficult to get to fully—we wanted to build text-to-video. But it became clear to us that an easier starting point would be to start from video-to-video because when you have a stronger conditioning it's basically an easier problem to restylize an existing video versus generate a video from scratch. And so we released Gen-1 first back in January of 2023.

Host

这只是一个有趣的视觉播客。老实说,就像我们可以看到 2023 年 2 月。当时的状态是什么?

It's just a fun visual podcast. Honestly, like we can see February 2023. What was the state of stuff?

Anastasis

我的意思是,这很有趣,因为当你看到那些结果时,你会觉得这太不可思议了。这几乎就像图像生成或视频生成已经解决了。然后几年后回头看,显然就像你很快对那些模型的结果习以为常。但当我们开始看到那些结果时,感觉相当不可思议,你能得到的质量水平。所以 Gen-1 是一个深度条件视频模型。所以它会接收输入视频,它会预测——你首先将其转换为深度图,然后我们用潜在扩散模型生成像素。

I mean, it's so interesting cuz at the time when you see those results, you think this is so incredible. This is like it's almost like image generation or video generation is solved. And then you look back a few years after and it's like it's obviously it's just like you get used to results very quickly with those models. But at the time when we started seeing those results it felt quite incredible and the level of quality you could get. And so the Gen-1 was a depth conditioned video model. So it would take an input video, it would predict—you first convert it into the depth map and then we would generate pixels with a latent diffusion model.

Host

是的,非常有效。

Yeah, very effective.

Anastasis

我没意识到那篇博客文章会这么分散注意力。抱歉。

I didn't realize how distracting the blog post would be. Sorry.

Anastasis

是的,但我在 Gen-1 上最喜欢的例子之一是,如果你转到模式三或模式二,有一个故事板用例,人们会制作——你可以摆弄——用书或盒子搭一个城市,然后用手机拍一段视频,然后将其转化为逼真的输出。这些模型开始被用于故事板,以及真正——然后如果你转到模式四,比如拿无纹理的 3D 场景,然后将其转化为照片级逼真的输出。所以,我们早期看到了很多用例,熟悉的人,强大的视觉特效编辑会直接拿一个 Blender 渲染,然后用 Gen-1 转化,或者在 Unity 中创建一个场景,然后捕捉一段视频,然后转化为,你知道,重新风格化。所以,我仍然认为视频到视频很强大。

Yeah, but one of my favorite examples actually of the on those on Gen-1 was both if you go up to mode three or mode two, there was this storyboard use case where people would make—you can mess around with—make a city out of books or out of boxes and then they would kind of shoot a video with their phone and then translate it into a for the realistic output. There was all these ways in which those models were starting to be used for storyboarding and also for really—and then if you go to mode 4 like of taking kind of untextured 3D scenes and then turning them into photorealistic output. So, we saw a lot of use cases early on where people that were familiar, where Power VFX editors would just take a Blender render and then they would get translated in with Gen-1 or create a scene in Unity and then take a capture a video of it and then translate into, you know, restylize it. So, I still think video-to-video is powerful.

视频到视频模型与真实基准 Video-to-Video Models and Ground Truth

Anastasis

我想我们最近也推出了一个视频到视频的模型,这是我最喜欢的使用这些模型的方式之一:本质上是用真实视频作为初始灵感,然后将其转化为不同的风格或不同的输出。

I think we had a recent video-to-video model as well, and it's one of my favorite ways of using those models: essentially using them to use ground truth video as the initial inspiration and then translate it into different styles or different outputs.

Stable Diffusion争议回顾 Stable Diffusion Controversy Retrospective

Host

但我想我们今天会继续聊 Runway 的其他部分,让大家跟上进度。我确实想谈谈所谓的 Stable Diffusion 争议,或者你知道 Stability AI 发生了什么。我觉得这个故事有两面。我认为其中一部分是正常的,人们加入和离开公司,但现在已经过去几年了,回顾起来是怎样的?

But I think we're going to go into the rest of Runway and catch people up to speed today. I did want to cover the let's call it the Stable Diffusion controversy, or you know what happened with Stability AI. I think there was sort of two sides of the story. I think part of that is a normal thing of people join and leave companies, but what is the retrospective now that there have been some years behind it?

Anastasis

是的,我想这是一个很长的故事。

Yeah, it's a very long story to go into, I think.

Host

我记得你实际上写了一篇很长的文章。我们可能可以花整整一个小时来更详细地讨论,但本质上你知道有潜在扩散论文,我想那是在……然后 Patrick Esser,他是潜在扩散背后的研究人员之一,当时在 Runway 工作,他与 Robin Rombach 和其他几个人在 Convis 合作构建了潜在扩散,那是一个实验室,研究小组。是的。

Which I remember you actually wrote a really long post about. We could probably cover the whole hour to go into it in more detail, but essentially you know there was the latent diffusion paper that came in, I think that was at the end of, and then Patrick Esser who was one of the researchers behind latent diffusion and he worked at Runway at the time, he built latent diffusion in collaboration with Robin Rombach and a few other folks back in the at Convis, which was a lab, the research group. Yeah.

Anastasis

在发布早期的潜在扩散模型后,他们本质上,你知道,目标是继续研究模型的版本,扩大规模,纳入新数据,纳入新任务。而 Stable Diffusion 基本上是相同的模型,但用更多的算力训练,然后还有一些技巧,比如无分类器引导论文在 2022 年初的某个时候出现。

And after releasing the early latent diffusion model, they essentially, you know, the goal was to keep working on versions of the model, going to scale it up, incorporate new data, incorporate new tasks. And Stable Diffusion was basically the same model but trained on more compute and then with a few more tricks like classifier-free guidance paper came at some point I think in the early 2022.

Host

这大大改进了提示。是的,你知道这改善了结果。

Which like was a big prompting improvement. Yeah, you know that improved results.

Anastasis

我是在更好的数据上训练的。所以像是 LAION 的美学子集,但本质上你知道是相同的底层架构,还有那次在 Stability 集群上进行的大规模训练,Stability 资助了那次训练。回顾那个故事,我认为构建和训练那个模型的工作已经完成了。那是一个研究项目。它是作为潜在扩散工作的延续的一部分完成的。然后我认为模型变得非常成功,我认为有……我认为由于它的成功,其他公司试图找出它的商业化路径。但对我们来说,非常重要的是我们尝试,你知道,我们确保我们,它本来是一个开源研究项目,所以我们决定应该继续发布它的版本,因为那是 Stable Diffusion 的原始目标,这导致了 Stable Diffusion 1.5 的发布。可能有一天有点沟通不畅,但最终在几小时内就很快解决了。所以是的,那很好。

I was trained on better data. So like the aesthetic subset of LAION, but it was effectively you know the same underlying architecture and there was that big training run that happened on Stability's cluster, Stability kind of financed that run. And looking back at that story, I think it was the work to build and train that model was done. It was a research project. It was done as part of like continuation of the latent diffusion work. It then I think the model became very successful and I think there were the and I think as a result of its success other companies tried to figure out the commercialization path for it. But for us it was very important that we try to, you know, we make sure that we, it was meant to be an open source research project and so we decided that we should continue releasing versions of it since that was kind of the original goal of Stable Diffusion and that led to releasing Stable Diffusion 1.5. There was maybe a day of a bit of miscommunication there but ultimately that was resolved very quickly within hours. So yeah, that was nice.

早期创意AI社区反思 Reflections on the Early Creative AI Community

Host

我只是想,你知道,你实际上是那段旅程中的主要参与者之一,所以听到当事人讲述发生了什么很好。是的。

I just wanted to, you know, you are actually one of the main players in that sort of journey and so it's nice to hear from the source of like what happened. Yeah.

Anastasis

是的。我的意思是,我想现在一切都已经过去了。两家公司,你知道,Stability 走了自己的路。Runway 走了自己的路。

Yeah. I mean, I think it's all in the past now I would say. And like both companies, you know, Stability took its own path. Runway took its own path.

Host

仍然有,我的意思是詹姆斯·卡梅隆支持新的 Stability,不管他们在好莱坞工作室做什么。我不知道他们在做什么。有一件事让我印象深刻,我很乐意继续,那就是在那个时候,比如 2021 年、2022 年,有一个社区,你参与其中,正在研究所有这些东西,对吧?从我交谈过的当时活跃的每个人来看,似乎很明显会有人进行英雄式的训练运行,从而产生 Stable Diffusion。所以我想问题是,你知道,就像你有了,你当时,你进行了投资,你有远见。是否准确地说这反映了当时人们的想法,还是仍然非常像是我们会把它用作后期制作工具之类的。我不知道,你知道,就像我们的情绪在哪里,也许你可以回想一下当时的社区是什么样的。

There's still I mean James Cameron is backing the new Stability whatever they're doing with the Hollywood studios. I don't know what they are doing. I think one thing that impresses me and I'm happy to move on is that back in that time, let's say like 2021 2022, there was this community of people that you were involved in that was researching all this stuff, right? And like from everyone I talked to who was active then, it seemed like it was fairly obvious that somebody would do the hero training run that would produce Stable Diffusion. So I guess the question is, you know, like you had the, you were, you had made investments, you had the foresight. Is it accurate to say like that is reflective of like what people were thinking at the time or was it still very much like well we'll use it as like a post-production tool or something. I don't know, you know, like where in the sentiment were we that maybe you can sort of think back to like what the community was like back then.

Anastasis

我回忆并非常怀念 2018 年到 2022 年的那些早期岁月,因为那是一个非常小的社区,正如你所说,他们非常确信这将会是一件大事,当时你知道任何人,因为圈子很小,每个想成为那个圈子的一部分并用它做项目的人都会立即走红。所以我记得

I reminisce and I think of very fondly those early years from like 2018 to 2022 because it was a very small community that as you said were very convinced that this was going to be big thing and at the time you know anyone who because it was such a small circle and everyone who would like be part of that circle and like make projects with it would you know immediately kind of get you know go viral. So like I remember

Host

你甚至不知道他们是谁,对吧?他们只是 GitHub 或 Hugging Face 上的某个名字。

You don't even know who they are, right? They're just there some name on the you know GitHub or Hugging Face somewhere.

Anastasis

没错。是的。所以我记得创意 AI 最早的病毒式传播时刻之一是神经风格迁移论文,那

Exactly. Yeah. So I remember one of the first big viral moments of creative AI was there was the neural style transfer paper that

Host

那个什么 dreaming

The something dreaming

Anastasis

我想它叫神经风格迁移。还有 Deep Dream,那个小狗鼻涕虫,也很酷。但有一个项目,Jin Kogan,他是 Runway 的早期顾问,也是那些营销人员之一,创意 AI 的大人物。他实际上只是展示了一段自己乘坐纽约地铁并经过威廉斯堡大桥的视频,然后将其风格化,我想是以梵高或某位画家的风格,当时那真是太酷了,它走红了,对人们来说完全是一个启示,你可以用生成模型做到这些,而那只是,你知道,不到,可能是 10 年前。就像是一个指示,表明事情发展得有多快。

I think it was called neural style transfer. There was also Deep Dream, the puppy slug, which was also really cool. But there was this project that Jin Kogan, who was an early adviser of Runway and one of those marketing guys, big creative AI folks. He literally just like showed a video of himself taking the New York subway kind of and going over the Williamsburg Bridge and then stylized it with, I think, in the style of Van Gogh or like one painter and that was like at the time that was like so so cool and it went viral and it was completely revelation to people that you could do those with generative models and that was only you know it was less than it was maybe 10 years ago. Just like as an indication of like how quick like things have gone.

构建视频扩散模型基础设施 Building Infrastructure for Video Diffusion Models

Host

这很疯狂,即使从那时起,你在堆栈的每个层面都有人。你有开发者、创意人员、艺术家、爱好者,每个人都在使用它。对于早期尝试过的人来说,他们会记得使用常规扩散有多难,对吧?就像现在你可以使用你最喜欢的聊天工具或其他什么,给一句话,得到漂亮的输出。但扩散就像你知道整个超高清 4K 高分辨率,提示这些东西是非常不同的。在工具方面,从你们现在提供的产品中,你学到了什么,比如创意人员、开发者,你真的把研究带给了每个人使用。有什么有趣的可以分享吗?

It's pretty crazy like even since then you've kind of got people at every level of the stack. You've got devs, creatives, artists, hobbyists, you got everyone using it. And for people that tried stuff early, they'll remember how hard it was to use regular diffusion, right? Like nowadays you can use your favorite chatten or whatever, give a sentence, get a beautiful output. But diffusion was like you know the whole ultra HD 4K high resolution like prompting these things was very different. Anything you learned on the tooling side like from the offerings you guys have now so like creatives devs you really took the research and brought it to everyone to use. Anything interesting there to share?

Anastasis

我们必须为视频扩散模型构建整个模型服务基础设施。之前没有其他类似的东西,因为我们的 Gen-2 是市场上第一个技术视频模型。随着时间的推移,我们学到了很多东西。

We had to build the entire model serving infrastructure for video diffusion models. There was nothing else already like because we had Gen-2 was the first tech video model I think out in the market. So many things that we learned over time.

文本到视频的早期教训 Early Lessons from Text-to-Video

Anastasis

我觉得最大的一个教训是,很早就很清楚,文生视频不会是答案。人们想要的控制权比这多得多。所以我们很快就投入到了在这些模型之上构建控制能力。你怎么用镜头轨迹作为控制?你怎么用初始输入帧作为控制?所以这是我们在文生视频上很早就学到的。Gen 2 在视频模型质量上是一次惊人的阶跃式提升,但它更多是以探索性的方式被使用,因为没有什么可以给它做锚定。没有可以参考的东西可以带进去。你没法真正控制镜头运动,也没法控制物体运动。所以 2023 年第一年,核心就是:我们能用哪些有趣的方式去条件化这些模型?当时做了大量在基础模型之上的后训练,去搞清楚人们到底想怎么控制它们。于是就出现了快速迭代:我们有一个叫 Motion Brush 的东西,基本上你可以画箭头,指定场景里东西该往哪动。还有镜头控制,你可以直接描述你希望镜头在场景里怎么移动。因为我们从 Runway 大部分历史时期都在和电影人合作,我们很快就得到了反馈,并决定这值得投入。所以在我们构建这些模型的过程中,可控性很早就成了一个重要主题。

I think the biggest one was that it was very clear early on that text to video was not going to be the answer. People wanted a lot more control than that. So we invested in control building on top of those models very quickly. How do you use the camera trajectory as control? How do you use an initial input frame as control? So that was a very early learning for us with text to video. Gen 2 was an amazing step function improvement in the quality of video models, but it was used much more in an exploratory way because there was nothing to ground it to. There was no reference that you could bring into it. You couldn't really control the camera motion, you couldn't control the object motion. And so the first year in 2023 was really all about what are all the interesting ways in which we can condition those models. It was a lot of just post-training runs on top of the base model to figure out how do people actually want to control them? And so there was this quick succession: we had Motion Brush, which was you could basically draw arrows and dictate where things should move in the scene. There was camera control, where you could just describe how you want the camera to move in the scene. And because we work with filmmakers from most of the history of Runway, we immediately got this feedback and decided that this was worth investing in. And so controllability became a big theme very early on as we were building those models.

Gen 2如何从Gen 1诞生 How Gen 2 Emerged from Gen 1

Anastasis

有一件挺有意思、我其实没怎么讲过的事,就是 Gen 2 是怎么从 Gen 1 演变出来的。这有点奇怪,因为我们在 Gen 1 发布两个月后就宣布了 Gen 2,而且那还是在 Gen 1 甚至还没普遍可用之前。

Something fun that I haven't really talked about too much was just how Gen 2 came to be out of Gen 1. It was a bit strange because we announced Gen 2 two months after Gen 1, and it was before Gen 1 was even generally available.

Host

对。

Right.

Anastasis

但 Gen 1 是一个深度到视频的模型。它会拿一张深度图,把它转换成 RGB。我们没法让文生视频或图生视频直接跑通,这就是为什么我们从深度到视频开始。我们讨论过,好吧,我们需要花接下来 6 个月真正投入到文生视频,也许提高算力规模或模型规模,比如训练一个更大的模型。然后我有了一个周末项目的想法:如果我拿一个从文本输入开始、转换成深度图的模型,然后用 Gen 1 把深度图转换成 RGB,会怎么样?所以 Gen 2 基本上就是这个。

But Gen 1 was a depth-to-video model. So it would take a depth map and convert it into RGB. We couldn't get text or image to video to work directly, and that's why we started from depth to video. We had discussions like, okay, we need to spend the next 6 months actually investing in text to video, maybe increasing the compute scale or the model scale, like train a larger model. And I had this weekend project idea, which was: what if I take a model that starts from text input and converts to depth maps, and then use Gen 1 to convert the depth maps into RGB? So Gen 2 was basically that.

Host

这就是一个黑客松式的流水线。

It's a hackathon pipeline.

Anastasis

但它看起来不错。

But it looks good.

Host

而且效果相当好。

And it worked pretty well.

Anastasis

我的意思是,确实有——如果你知道它有这个两阶段流水线,有些情况下你能看出来视频的结构看起来有点不对,因为你必须先生成深度,然后才能生成输出视频。但它跑通了,让我们能很快把东西带给用户。但现在有意思的是,人们又回到了这种近乎两阶段的方法。比如你看几个月前出的 Reeve 文生图模型,它有一个规划器模型,会先生成边界框,然后再喂给扩散 Transformer。

I mean, there were, you know, if you have knowledge that it has this two-stage pipeline, you can tell in some cases that the structure of the video looks a bit off because you had to generate the depth first before you go into do the output video. But it worked and it allowed us to bring this to users very quickly. But it's actually now interesting because people are coming back to this almost two-stage approach. Like if you look at the Reeve text-to-image model that came a few months ago, it had this planner model that would generate bounding boxes before it fed that into the diffusion transformer.

Host

对。Ideogram 也是同一天。我记得那很奇怪,两个模型同一天出来,带着完全一样的创新。

Yeah. Ideogram also the same day. I remember that was very strange that both of them came out the same day with the same exact innovation.

Anastasis

圈子很小。

It's a small community.

Host

我心想,这完全是巧合,对吧?

I'm like, this is completely coincidental, right?

Anastasis

大家会交流。所以是的,这种方法肯定有它的道理。而且显然现在每一个在生产环境中的视频生成模型,底层都用了复杂的提示词补全流水线。我觉得这已经不是秘密了。

People talk. So yeah, there's definitely something to this approach. And obviously now every single video generation model in production uses a complex prompt completion pipeline under the hood. I think that's no secret.

Host

人类很不擅长写提示词。我觉得这是普遍现象。

Humans are terrible at prompting. I think across the board.

Anastasis

但是的,我觉得最初那篇 Sora 博客甚至告诉过你,你的输入之后发生的是它在重写你的提示词,让它对你想要的东西描述得更详细。

But yeah, I think the original Sora blog post even told you that what happens after your input is it's rewriting your prompt, it's much more descriptive about what you would want.

Host

没错。对。而且之前还有 DALL·E 3 的论文,那是第一次公开描述合成字幕和非常详细字幕效果很好这件事,然后 Sora 算是建立在其之上。

Exactly. Yeah. And there was the DALL·E 3 paper beforehand that was the first public description of the fact that synthetic captions and really detailed captions work really well, and then Sora kind of built on that.

相机控制与世界模型萌芽 Camera Control and the Seed of World Models

Anastasis

对。那是 2023 年。我们在给 Gen 2 发布所有这些更新,比如镜头控制、Motion Brush。而镜头控制其实有件非常有意思的事,因为那是你第一次感觉到,你不是在创作视频、不是在创作一段短视频,而是在一个世界里穿行。我觉得镜头控制也许就是我们关于世界模型的一些想法、以及真正打开那个研究方向的那颗种子。我们意识到,那个时代、Gen 1 和 Gen 2 这一系列模型,真的向我们自己证明了这一点。

Yeah. So it was 2023. We were releasing all these updates to Gen 2, like the camera control, Motion Brush. And there was actually something very interesting about camera control, because it was the first time that you felt that instead of you were creating video, you were creating a short video, you were actually navigating inside the world. And I think camera control was maybe the seed of some of the ideas that we had around world models and really opening up that research direction. We realized, you know, it was this era and this series of Gen 1 and Gen 2 models really proved to ourselves.

Host

对。

Yeah.

Anastasis

所以这不是最初的镜头控制。这是 Gen 3 之上镜头控制的更新版。但是的,我觉得它让那些模型对电影人来说变得可用了。我想说镜头控制非常受欢迎,于是我们意识到,看这些模型有一种方式,就是你只是内容创作机器。还有另一种方式,就是你在预测视频。为了把视频预测好,你需要以越来越强的能力去模拟世界。如果缩放定律适用于视频,就像它们适用于语言模型一样,那么随着我们投入这些模型的算力不断扩大,它们将能够模拟物理。它们将能够越来越好、越来越可预测地模拟人类动作和动态。这就是我们围绕世界模型所做努力的论点。我们成立了一个研究小组,专门聚焦世界模型,以及我们如何把正在构建的视频生成模型变成更广泛的东西,变成在内容创作之外也有用的东西。

So this is not the original camera control. This was the update to camera control on top of Gen 3. But yeah, I think it made those models usable to filmmakers. I would say camera control was very popular, and so we realized there's one way of seeing those models, which is you're just content creation machines. And there is the other way, which is you're predicting video. In order to predict video well, you need to simulate the world in an increasing capacity. And if scaling laws apply on video just like they apply on language models, then as we scale the compute that we put in those models, then they're going to be able to simulate physics. They're going to be able to simulate human actions and dynamics increasingly well and predictably well. That was the thesis around our efforts on world models. And we spun up this research group to just focus on world models and how do we turn the video generation models that we're building into something broader and something that would be useful beyond content creation as well.

Host

那大概是什么时候?

And that was roughly when?

Anastasis

对。那是在 2023 年底。

Yeah. So that was in late 2023.

Host

有意思。你知道,我觉得很多人一直在说,如今很多视频生成模型公司都转向了世界模型,但 2023 年你们就在提这个了。

Interesting. You know, I think a lot of people have been saying a lot of video gen model companies have all pivoted to world models these days, but 2023, you're posting it.

Anastasis

我的意思是,这算不算转向是有争议的。可以说这本来就是你必须做的事,对吧?在某种意义上,这是随着模型变得更强,其应用范围的扩展。

I mean, it's debatable whether it's a pivot. Arguably that's what you always had to do anyway, right? It's in a way an expansion of the applications of the models as they become more capable.

Host

早期的迹象。看起来你们最初的那些模型,人们会说这很不符合苦涩教训,对吧?你在加提示词重写。

The early signs. It seems like the original models you guys had, people would say it's very not bitter lesson, right? You're adding rewriting prompts.

从Gen-2到世界模型 From Gen-2 to World Models

Host

你现在做的都是些一次性的东西,但那只是技术当时的状态,而未来就像你说的,你可以把它扩展,你知道,我们可以扩展到世界模型。

You're having all these one-off things, but that's just the state of the tech as it was versus the future of, as you said, you can scale it up is, you know, we can scale up to world models.

Anastasis

是的。所以它就成了,而且如果你看 Gen 2 的输出,我认为人们并不明显觉得这会扩展成通用的世界模拟器,因为动作非常有限,分辨率很低,人体解剖上有明显错误,各种限制都有。但当时的想法就是,这就像 GPT-2,你几乎无法生成连贯的句子,类似地 Gen 2 也几乎无法生成连贯的视频,但如果你扩展它,就没有理由不成功。我认为这始终是 Runway 的心态,就是这种外推:即使我们在 2018 年开始时,你看当时的结果,你需要更多地看趋势,比如 2018 年我们处于什么位置,对比 2020 年第一个模型出来时。你从 32x32 的人脸图像开始,到 2018 年你可以生成 1k 分辨率的街景图像,世界模型也类似,有非常早期的迹象表明会有更大的东西。

Yeah. So it just became and and and if you looked at the outputs of Gen 2, it was not I think it was not obvious to people that this would scale to become a general simulator of the world like you had very limited movement, you had you know very low fidelity resolution like obvious mistakes in human anatomy like all kinds of limitations but it was just you know uh the idea was that's just GP2 and GPT2 you know you can barely generate like coherent sentences similar Gen two can barely create coherent video, but if you scale it up, you're going to there is no reason why it shouldn't work in a way. It's uh and I think that was that's that's always the mindset of kind of of runway is like this extrapolation of like if you know like even when we started in 2018 and you looked at the results of the day you need to look more at the trend of like where we were in 2018 versus when we were at the you know when the first gun came out in 2020 uh for 2014 or 20 uh 15. And you start from like 32x 32 images of faces and then by the time in 2018 you could generate you know street images at a 1k resolution and it was kind of the same with world models very early signs of something much bigger.

Host

是的。我几乎觉得它像是在扩散中聚焦,如果你看我们年复一年的可见输出,它本身看起来就像一个扩散过程。

Yeah. I was almost think like it's kind of diffusing into focus like if you look at our visible output from year to year to year it looks like a diffusion process itself.

Anastasis

是的。尤其是看早期那些旧博客文章,你真的能看到那种不连贯,那种

Yeah. Especially watching the early like old old blog post, you can really see the choppiness, the

Host

细节

details

Anastasis

就像人类文明从随机噪声开始,然后逐渐噪声化进入

like human civilization starting from random noise and then noising into

Host

是的。是的。就运行它。

Yeah. Yeah. Just run it.

Anastasis

是的。这就是你知道自己在正轨上的方式。你知道,你还在噪声化,对吧?

Yeah. That's how you know you're on track. You know, you're still noising, right?

转向世界建模 The Shift to World Modeling

Host

是的。我喜欢你们在六月宣布时的那种表述,哦,你们有一个视频文章。人类思维不再是 AI 的中心。我们的世界才是,对吧?也就是说,过去五年基于 LLM 的 AI 很大程度上是在试图模仿人类偏好和人类语言,但现在这基本解决了,或者我认为这是你当时写的文章的一些背景,现在重点是对世界进行准确建模。

Yeah. I like the way that you guys phrased it when you announced it in June, which is, oh, you had a sort of video essay. The human mind is no longer the center of AI. Our world is, right? which is uh you know let's let's call it the the past 5 years of of LLM based AI is very much like trying to emulate human preferences and human speech but now that's like mostly solved or I think that's like some of the context of your essay which you also wrote around the time and now it's like the focus on modeling the world accurately.

Anastasis

正是如此。是的。所以我们看待它的方式是,DeepMind 最初的使命宣言是解决智能,然后用它解决其他一切,但我认为从其他一切开始可能更有价值,因为世界上有太多复杂性和细节,很难直接从人类对世界的描述中学习。我们假设语言模型从人类所写的关于世界的一切中学习,比如我们截至 2020 年代的理解。但有很多我们不知道的东西,很多现有文本没有捕捉到的东西,比如世界的低层动态,我们没有详细描述。如果你让我描述如何系鞋带,用语言描述非常困难,但演示却非常明显。所以我认为存在一个悖论,我们不断低估人类潜意识中做的事情的复杂性,我们甚至不一定总有词语来描述它们。所以在我看来,模拟世界、模拟物理、模拟世界的动态一直被低估,相比我们过于强调容易谈论的事情。但世界上有所有这些复杂性和丰富性,如果我们直接训练观察数据,而不是训练人们如何描述世界,我们会学到一些否则不知道的新东西。

Exactly. Yeah. So the the the way we see it is uh there is that um that uh initial mission statement of deep mind which is uh uh solve intelligence and then use it to solve everything else but I I think it's starting from everything else uh could be valuable of like starting from you know there is just so much complexity uh and detail in the world that in order to that it's it's hard to learn directly from just human descriptions of the world like we're assuming saying that, you know, like language models learn from everything that humans have written about the world, like our own understanding as of, you know, the 2020s. And there's just so much that we don't know and so much that's not captured by existing text uh about both the, you know, the low-level dynamics of the world like we're not describing in detail. you know, if if I tell you to describe like how do you tie your shoes, that's a very difficult thing to to describe in words, but it's very obvious thing to demonstrate. And so I think there's been and there's, you know, more of paradox like we're constantly underestimating all the complexity that goes into very like things that we do subconsciously as humans and we don't even necessarily always have the words to describe them. And so in my mind the simulating the world and simulating um physics, simulating the dynamics of the world has always been kind of underestimated uh uh compared to uh we place too much emphasis on the things that are easy to talk about. Uh but there's just all this complexity and kind of richness of the world that if we just try and train train directly on that observational data instead of training on how people describe the world, we would learn something new that we wouldn't otherwise know.

架构与扩展 Architecture and Scaling

Host

你认为当前的架构范式没问题。你不需要像 Japa 那样的另一层,你知道,就像另一位著名的纽约 AI 领袖会说的那样。

You think that the present architectural paradigm is fine. You don't need like another layer like Japa, you know, like another famous uh New York AI leader would say.

Anastasis

你知道,我们是一个非常务实的研究实验室。如果我们有证据表明某种方法比我们采用的方法更好,我们会毫不犹豫地采用它,我们只是没有看到任何迹象表明视频预测本身不能扩展。即使你现在看,不只是我们的工作,还有其他人的工作,在机器人领域,一些最有前途的工作从视频预测模型开始,然后你调整它们也作为动作模型。所以很少有证据表明你需要其他东西,你的时间最好花在新颖的架构改变上,而不是改进数据和扩展当前方法。所以我们没有任何迹象表明,那个反驳论点,我认为几天前 Yan Leon 发了一条推文,说理解世界的动态与生成可爱的视频非常不同,而你的回答是不。它们是一回事。我的猫视频和理解物理是一回事,

You know, we're a very pragmatic research lab. If uh we have evidence that an approach works better than the approach that we're taking, then we have no qualms to taking it, we just have seen no indication that video prediction itself doesn't scale. And even if you look now, you know, not just our work, but the work of others, you're seeing in robotics, some of the most promising work um starts from video prediction models and then you adapt them to also be action models, for example. Um, so there is very little evidence that you need something else and that your time is better spent on a novel architectural change compared to improving data and improving the and scaling the current the current approach. And so we don't have any indication that you know the the there is that counterargument that uh I think there was a a tweet by Yan Leon a few days ago that you know uh understanding the dynamics of the world is very different than uh generating uh cute videos and your answer is no. They're the same thing. My cat videos are the same as understanding physics,

Host

对吧?

right?

Anastasis

因为如果你想生成,显然视频模型可以作弊,它们可以给出连续的镜头,而不需要实际模拟困难的物理。总有各种方式隐藏模型的缺陷。重要的是不要被当前视频模型的性能所迷惑。很容易挑选例子,认为视频模型比实际更先进。所以我们需要做更多工作来改进这些模型。但在我看来,与语言非常相似,我们从几乎不连贯的句子到能与人类对话,再到能自主运行一天并创建整个代码库。主要区别显然是沿途的一些架构改进,但主要是规模,所以视频也是同样的赌注,我们没有迹象表明这会饱和,我们有基准来衡量这些模型的物理,我们看到随着模型扩展,这些基准可预测地改进。所以如果你想查看,物理 IQ 是其中一个基准,衡量模型在力学、流体动力学或光学方面的表现。

Cuz if you want to generate, you know, obviously video models can cheat and like they could you could give like successive uh shots of the scene in a way that doesn't require you to actually simulate difficult physics. There's always like all these different ways in which you can hide the deficiencies of the model. And it's important not to be kind of too tricked by the performance of the current video models. it's easy to, you know, cherrypick examples and and think that video models are further advanced than they actually are. So, there is a lot a lot more work that we need to do to improve those models. But in my mind, very similar to language and like we've, you know, you go from barely coherent sentences to something that, you know, could hold a conversation with a human to something that could can operate autonomously for for a day and like create entire code bases. And the main difference there's obviously some architecture improvements along the way but the main thing is scale and so it's the same bet for video and we have no indications that this is saturating like we have benchmarks that we use for measuring the physics of those models and we see those predictably improve as we scale those models. So there is if if you want to pull up uh physics IQ uh is one of those benchmarks that measures how well does the model perform at all mechanics or fluid dynamics or optics.

Host

我很好奇你是否看到了任何涌现,任何关于这方面的缩放定律。

I'm curious if you've seen any emergence any scaling law around this.

直觉物理的缩放定律 Scaling Laws for Intuitive Physics

Host

对,他是在说存在一条缩放定律,对吧?

Yeah, he's saying there is a scaling law, right?

Anastasis

没错。这些模型、这些基准测试的工作方式是:研究人员先拍摄一些能代表不同物理现象的视频,然后你取第一帧,把它输入图像到视频模型,生成一段展示接下来应该发生什么的推演。比如有一个球悬挂在天花板上,你把它作为输入,然后模型预测球应该如何落到地上。这衡量的是我们对物理的直觉理解。我知道你能想象如果我松开这个瓶子接下来会发生什么。所以它衡量的是那些模型同样的直觉物理理解,我们在不同的模型规模和算力规模上做了测量,看到分数和物理智商可预测地提升。还有其他技巧和方法可以进一步提升分数,但仅靠规模本身就能帮助模型更好地学习物理。

Exactly. So the way those models, those benchmarks work is the researchers have gone and captured a few videos that are representative of different physical phenomena, and then you can take the first frame and pass it through an image-to-video model and generate a rollout that shows what should happen next. So you have a ball hanging from the ceiling, and then you use that as input, and then the model predicts how the ball should fall on the ground. And this measures our intuitive understanding of physics. I know you can imagine what will happen next if I drop this bottle. So it's measuring that same intuitive physics understanding of those models, and we've measured that at different model scales and compute scales, and we see that the score and physics IQ predictably improves. There are other tricks and techniques that you can make to improve the score even further, but even scale alone helps in the model learning better physics.

柏拉图洞穴与苦涩教训 Plato's Cave and the Bitter Lesson

Anastasis

我对 Yann LeCun 的主要共鸣在于那个柏拉图洞穴的寓言,对吧?你是在一个事物的输出上学习,而不是在它的内部过程上学习。而且这非常非常嘈杂。要是你能观察到事物的内部就好了。观察人类心智的内部很难。但你完全可以观察——或者说我们有一大堆关于如何建模物理、运动、重力和其他相互作用的科学和物理学,却被我们忽略了——我们只是把这一切都扔掉,只说“扩大数据规模”,这非常符合无监督学习的教训,但感觉不对。我觉得这就是核心观点。

My main sympathy with Yann LeCun is the Plato's cave allegory, right? You're learning on the output of a thing, not the internal process of a thing. And it's very, very noisy. If only you could observe the internals of a thing. It's hard to observe the internals of a human mind. But you can very much observe—or at least we have a whole bunch of science and physics that we're ignoring on how to model physics and movement and gravity and other interactions—and we're just throwing away all that and just saying just scale data, which is very much the lesson of unsupervised learning, but it feels wrong. I think that's the main idea.

Host

我觉得机器学习的历史很大程度上就是“感觉不对”。

I think the history of machine learning is largely 'it feels wrong.'

Anastasis

对。苦涩的教训就是对此的简单答案。

Yeah. The bitter lesson is the simple answer to that.

扩展视频生成与Runway历程 Scaling Video Generation and Runway's Journey

Host

我想问你能扩展到什么程度?比如在视频生成方面,有视频理解和视频生成。我们还会不会有这样的工具:我想生成 2 小时、20 小时的视频?有一种推理的方式是分批生成再拼接起来,但我们就一直扩展下去吗?我们就继续追求长生成的一致性吗?所有这些都会扩展。把它和我们现在的进展联系起来,从 Runway 2 到 4.5,技术上我们到今天取得了哪些进步,然后你觉得事情还会往哪里发展?

I guess how much can you scale? So even on, let's say, the video generation side, there's one side of video understanding, video generation. Are we still going to have tools where it's like I want to generate 2 hours, 20 hours? There's an inferential way to do it in batches and stitch it together, but do we just keep scaling? Do we just continue long generation consistency? All that would scale. And tying it into where we're at now, from Runway 2 to 4.5, like technically what advancements have we made to today, and then where do you see things still going?

Anastasis

答案的一部分肯定是规模,这是我们在 Gen 3 上学到的一大教训。Gen 3 是我们第二年发布的模型,也就是 2024 年,在 Sora 发布几个月后。所以它诞生的过程也有一段有趣的故事。Gen 3 对我们来说是第一次我们真正需要构建。基本上我们必须在几个月内学会语言模型世界在三年里学到的所有教训。Sora 最大的变化之一是使用扩散 Transformer 而不是卷积网络。所以很多早期的潜在扩散模型在扩散模型部分都是卷积网络。而扩散 Transformer 的论文在 2023 年某个时候出现。它基本上展示了图像扩散 Transformer 的缩放定律。我们那时意识到,我们需要投资于模型并行的基础设施,以便真正将训练扩展到超过几十亿参数的模型。我们花了大概整个秋天来构建分布式训练的基础设施。我们在尝试扩展图像和视频扩散 Transformer 时有很多错误的开始和很多失败。然后到了 2024 年 2 月,Sora 出来了,结果比 Gen 2 能产生的要好得多。Twitter 上有很多议论说 Runway 完了,Runway 不可能赶上。如果你还记得 2024 年初的 OpenAI,感觉它现在是一个强大的对手,但那时他们正处于巅峰状态,没人能接近他们。也许 Gemini——只是第一个版本的 Gemini 刚刚发布。所以当 OpenAI 带着 Sora 出现,质量上有这么大的飞跃时,它给了我——有几个小时像是存在主义危机。但我认为 Runway 的惊人之处在于——我想我们已经成立 8 年了,几乎是——在 AI 领域我们是恐龙——我们有很多这样的时刻,必须快速学习、适应,并在团队中建立我们没有的技能。所以如果你问任何人在 Runway 最喜欢的时光,那就是那段时间。那是在大约 3 个月内推动做出比 Sora 更好的模型。我们把模型规模、模型大小和训练算力扩大了 10 倍。我们搞定了模型并行——我们对此毫无经验——然后我们在那个夏天推出了 Gen 3。所以我认为那是公司的一个大转折点,研究工作增长得非常快,我们真正开始认真追求通用世界模型的愿景,我想是在 Gen 3 发布之后。

So part of the answer is definitely scale, and that was a lesson we learned in a big way with Gen 3. Gen 3 was the model we released the year after, in 2024, a few months after Sora was released. So there's an interesting story of how that came to be as well. Gen 3 for us was the first time that we really needed to build. Basically we had to learn all the lessons that the language model world learned in three years in the span of a few months. One of the biggest changes of Sora was using diffusion transformers instead of convnets. So a lot of the early latent diffusion models were all convnets for the diffusion model part. And the diffusion transformer paper came at some point in 2023. And it basically showed scaling laws for image diffusion transformers. And we realized at that point that we needed to invest in infrastructure for model parallelism, for really scaling training to larger than a few billion parameter models. And we spent maybe most of the fall building out our infrastructure for distributed training. And we had a lot of false starts and a lot of failure in trying to scale image and video diffusion transformers. And at that point, February 2024, Sora comes out and the results are very much superior to what Gen 2 could produce. There was a lot of chatter on Twitter about Runway's done, there is no way Runway will catch up. And if you remember also OpenAI in early 2024, it felt very much like it's a formidable opponent now, but at that point they were on the top of their game. Nobody could even get close to them. There was maybe Gemini—just the first version of Gemini had just released. So when OpenAI came with Sora and it was such a big jump in quality, it gave me—there was like an existential crisis for a few hours. But I think the amazing thing about Runway—and I think we've been around 8 years now, which is almost—we're dinosaurs in AI—and we had a lot of those moments where we had to learn, adapt very quickly, and build out skill sets in the team that we didn't have. So if you ask anyone what is their favorite time at Runway, that was during that time. It was that push in like 3 months to get to a model better than Sora. And we scaled 10x the model scale, the model size, and the compute that we were training on. We figured out model parallelism—we had zero expertise in that—and then we came out with Gen 3 during that summer. So that was a big turning point I think for the company, where the research work grew very quickly and we really started pursuing this vision of the general world model in earnest, I think, after Gen 3 was out.

全栈构建与模型服务 Building Full Stack and Model Serving

Host

对,我的意思是,构建的惊人之处在于——当你在构建时,没有现成的技术栈,你必须自己发明一切。你必须完全全栈。现在我认为有像 Fal 之类的推理专家可以帮助模型服务,我想你们也和他们合作。但没错,当时只是——想想当 Sora 出来、人们质疑你的公司是否还应该存在时你会怎么做,这非常有趣。

Yeah, I mean that's the amazing thing about building—when you're building, there's no stack, you have to invent everything yourself. You have to be completely full stack. Now I think there are inference specialists like Fal or whatever that can help with model serving, and I think you guys work with them as well. But yeah, at the time it was just—it's very interesting to think about what you do when Sora comes out and people are questioning whether your company should still exist.

Anastasis

对。而且当时没有扩散模型的 VLM,我们必须构建整个模型服务基础设施并提高效率。在我们发布 Gen 3 几个月后,我们发布了 turbo 版本,我认为那是第一个投入生产的步数蒸馏模型。

Yeah. And there was no VLM of diffusion models, we had to build the whole model serving infrastructure and make things efficient. And a few months after we released Gen 3, we released the turbo version, which I think was the first step-distilled model in production.

Host

那也是我们报道过的一个趋势。对。

That was a whole trend that we covered as well. Yeah.

Anastasis

所以这让我们能够以更大的规模服务那些模型,因为我认为第一版 Gen 3 的服务成本相当高。你知道,我认为一致性模型、lightning、turbo 等所有这些趋势不知怎么并没有真正持续下去。

So that allowed us to actually serve those models at a larger scale, because I think the first version of Gen 3 was quite expensive to serve. You know, I think the whole trend in consistency models, lightning, and turbo and all these things somehow didn't really stick around.

蒸馏与实时模型 Distillation and Real-Time Models

Host

我不知道你对此有没有什么反思,因为当时我想,显然一切都应该先从蒸馏模型开始,然后你再放大,对吧?然后基本上你更大的模型就变成了花哨的放大器,但你应该总是用更小更快的模型来起草,对吧?因为你可以很快得到结果,接近实时。

I don't know if you have any reflections on this because at the time I was like, well, obviously everything should start with a distilled model first and then you can upscale, right? Then basically your bigger models just turn into fancy upscalers, but like you should always draft with a smaller and faster model, right? Because you can get it so quickly, like near real time.

Anastasis

是的,我不太确定说那没有留下来。我认为那很可能……我的意思是,有很多步蒸馏模型正在生产中被积极使用。与非蒸馏模型相比,质量上显然仍有差距。但在我看来,我们仍然……与语言模型有两到三年的差距。所以,更好的蒸馏技术出现只是时间问题。我们现在使用一个名为 Character 的实时模型,我认为这是最大的实时视频模型部署。那是一个步蒸馏模型,并且正在被积极使用。与通用视频模型相比,这是一个非常具体的用例。

Yeah, I would not be so sure to say that didn't stick around. I think that's likely to... I mean, there are a lot of step-distilled models that are actively used in production. There's still obviously a gap in quality compared to the non-distilled model. But in my mind, we're still... there is a two to three year offset from language models. So it's just a matter of time before there are better distillation techniques. We use right now we have a real-time model called Character, which I think is the largest deployment of real-time video models. That's a step-distilled model and it's actively being used. It's a very specific use case compared to a general video model.

Host

顺便说一下,这是虚拟形象一致性角色。

By the way, this is avatar consistency character.

Anastasis

是的。所以这是一个会说话的虚拟形象模型。我们能够对其进行极致优化。它以 24 fps 生成,是一个步蒸馏自回归视频模型。所以如果我们看我们的世界模型方向,一个重要的组成部分是从双向扩散开始,基本上一次性生成整个视频,然后使其自回归。所以你一次生成一帧或几帧。所以,要得到实时模型,管道中有很多工作要做。首先,你需要将其变为因果自回归模型,然后你需要进行一些额外的步蒸馏,使其真正实时。我认为那部分实际上才刚刚开始。如果两年后我们不主要使用实时模型,我会非常惊讶。对我来说,实时视频生成是不可避免的,因为它有更好的用户体验,服务成本更低,而且随着我们找到更好的蒸馏技术,基础模型和实时模型之间的质量差距只会缩小,我们在内部在蒸馏时保持基础模型质量方面取得了很大进展。

Yeah. So this is a talking avatar model. We were able to optimize the hell out of it. And it generates at 24 fps and it's a step-distilled autoregressive video model. So if we look at our world model direction, a big component of it is starting from the bidirectional diffusion that basically generates an entire video at once and making it autoregressive. So you generate one frame or a few frames at a time. So there's a lot that goes into that pipeline of getting to a real-time model. First you need to make it into a causal autoregressive model and then you need to do some additional step distillation to get it to actually be real time. And I think that part is actually just starting. I would be very surprised if we're 2 years from now we don't primarily use real-time models. To me, real-time video generation is just inevitable that it has much better user experience, is much cheaper to serve, and the quality gap between the base model and real-time model is only going to close as we figure out better distillation techniques and we made a lot of progress there internally on maintaining the quality of the base model when we distill them.

Host

这有多少是可转移的?所以是同一个基础模型吗?比如如果你在整个序列上做扩散,然后将其转换为步自回归蒸馏,这是像蒸馏那样你仍然需要训练两者吗?你可以使用相同的基础和转换器吗?从技术层面讲,从常规模型到实时模型的过程是什么样的?

How much of this is transferable? So is it the same base model like if you're doing diffusion across the whole sequence and you're converting it to step autoregressive distillation, is this like distillation where you still need to train both? You can use the same base and converter? What's that process like to go from regular model to something that's real time on a technical level?

Anastasis

所以扩散模型的好处是你有两个蒸馏轴。所以你可以蒸馏到一个更小的模型,这类似于你在 LMS 中所做的,或者你可以在减少步骤方面进行蒸馏,减少扩散步骤。所以你可以取一个在 50 步内生成的模型,并在四步内生成,得到……你会有一些性能下降,但很多时候你会得到可比的输出。所以你甚至可以取大型前沿模型,用步蒸馏进行蒸馏,并达到实时性能,这就是我们所看到的。所以根据用例,在某些情况下我们可能也会用更小的模型提供服务,但在很多用例中,我们实际上只是使用前沿模型,并且能够使其实时工作。

So the nice thing about diffusion models is you have two axes of distillation. So there is the you can distill to a smaller model which resembles what you do in LMS or you can distill in terms of taking less steps, less diffusion steps. So you could take a model that generates in 50 steps and generate in four steps and get to... you have some performance degradation but very often you get comparable outputs. So you can even take the large frontier model and distill it with stepation and get to a real-time performance and that's what we've seen. So depending on the use case in some cases we might also serve with a smaller model but in a lot of use cases we actually just use the frontier model and we're able to make it work in real time.

实时世界模型与界面演示 Real-Time World Models and Interface Demo

Host

我认为现在可能是切换到他的笔记本电脑,展示一些你正在做的实时内容的好时机。

I think this might be a good time to cut over to his laptop to show off some of the real-time stuff that you're doing.

Anastasis

这是我们最近做的一项研究更新。所以我们一直在努力将我们的通用世界模型应用到不同的应用中。我们认为非常引人注目的一个应用是使用通用世界模型作为本质上是一个界面,一个通用的软件界面。这是我们世界模型的一个版本,称为界面世界模型。其想法是它本质上取代了软件应用的前端。它直接渲染界面的像素,并经过训练来预测点击或与界面其他交互后会发生什么。所以这全是像素,没有 HTML、CSS、React 驱动这个界面。这直接是我们实时视频生成模型的输出,并且它直接接受点击作为输入。

This is one of the research updates that we did recently. So we've been working in getting our general world models to different applications. One of them that we think is very compelling is using general world models as essentially an interface, a universal interface to software. This is a version of our world model that's called an interface world model. And the idea is that it essentially replaces the front end of a software application. It renders the pixels directly of an interface and it's trained to predict what happens next as a result of a click or another interaction you have with interface. So this is all pixels, it's there is no HTML CSS react that's powering this interface. This is directly at the output of our real time video generation model and it takes clicks directly as input.

Host

还有拖拽。点击和拖拽,对吧?

And drags. Click and drag, right?

Anastasis

所以它支持点击,支持拖拽。它还支持滚动。令人惊奇的是,你可以在提示中有效地描述你希望不同元素如何,比如你希望不同元素的行为是什么。所以这几乎是你可以在提示中描述界面,而不是用标记语言描述像 HTML 界面,你可以直接描述界面,比如如果我按下这个按钮,我期望发生这个,如果我按下这个按钮,这应该发生。我们认为这对原型设计很有用,比如测试不同交互的感觉,你还可以添加音频,所以它是一个视频音频生成模型,所以你本质上描述点击的视觉结果以及是否有音效。所以我们相信这将是一种更灵活的构建软件的方式,直接渲染它,你知道为什么生成生成像素的代码,直接生成像素,这是端到端哲学应用于前端。所以我们认为有几个有趣的用例,你可以在其上构建创意工具,我们认为对于任何涉及大量探索的用例,或者像教育用例,你想学习一个新概念,你想要某种可视化,以及开放式探索。我们认为这些,你知道,这是一种非常强大的方法。

So it supports yeah clicks, it supports drags. It also supports scrolling. And the amazing thing about this is that you can effectively describe in the prompt how you want different elements like what do you want the behavior of different elements to be. So it's almost your you can turn an interface from a markup language description of like an HTML interface and instead you can just describe the interface you know if I press this button I expect this to happen if I press this button this should happen and it's useful we believe both for prototyping for like just testing like what different interactions would feel like you can also add audio to it so it's a video audio generation model so you get you essentially and describe both what the visual outcome should be of your click and also what the if there's a sound effect that comes out of it. So we believe that's gonna be a much more flexible way of building software just render it just you know why why generate the code that generates the pixels just generate the pixels directly it's the end to end philosophy applying apply to front ends so we think there's a few interesting use cases so you can build creative tools on top of it we think that you know for any kind of use case that involves a lot of exploration or like educational use case where you want to learn about a new concept and you want some kind of visualization and and kind of open-ended exploration. We think those you know this is a very powerful approach.

Host

是的,你可以想象新的设计形式,工业设计软件可能作为这些模型的结果出现。而且这也是实时生成的。所以你可以构建很多有趣的相机过渡和交互形式,这些用其他方式很难构建。

Yeah, you can imagine new forms of design, industrial design software that could emerge as a result of those models. And this is all generated in kind of in real time as well. So you can build a lot of interesting kind of camera transitions and kind of forms of interaction that are very difficult to build otherwise.

用视频模型生成界面 Generating interfaces with video models

Anastasis

我们评估这件事的一种方式是:如果你试着用 Claude 生成同样的界面会怎样——直接给 Claude 提示,这是我用 Figma 做的界面的图像参考,或者我在别处创建的,让它生成这个特定的交互,在这个例子里就是把那个对象向上拖拽。除了更慢之外,仅靠 LLM 也很难捕捉某些交互。所以我们认为这很可能就是未来很多软件的创建方式。另外一个额外的好处是,个性化用这些模型可能会容易得多——你基本上可以根据谁在访问界面来尝试不同的提示。你可以更容易地通过提示工程让界面使用更大的文字,以满足无障碍需求,或者适配你特定的学习偏好。所以我们对这个方法非常兴奋。这显然还处于早期,我认为我们还需要让它更具成本效益,才能服务这些模型,因为运行一个实时视频模型相比纯粹渲染 HTML,计算需求显然要高得多。但我们确实看到了这种方法在构建前端界面方面的巨大潜力。

One way in which we evaluate this is: what if you try to generate the same interface with Claude by just prompting Claude — here's an image reference of my interface that I made in Figma or that I created somewhere else, create this particular interaction, which in this case is drag that object upwards. Beyond being slower, it's also very difficult to capture some interactions with just LLMs. So we think this is likely to be the way that a lot of software in the future will be created. And one of the additional benefits is personalization might be a lot easier done with those models — you can essentially try out different prompts based on who's visiting the interface. You can more easily prompt-engineer the interface to have larger text for accessibility reasons, or if you have particular study preferences. So we're very excited about this approach. It's obviously early days, and I think we'll need to make it more cost-effective as well to serve those models, because running a real-time video model versus just purely rendering HTML — the computational needs are obviously much higher. But we do see a lot of potential in this approach to building front-end interfaces.

Host

我们之前在 Flipbook 那期和 Ethan 聊 Grok 视频时也讲过类似的东西。我觉得它在视觉上非常吸引人。我觉得它可能对教育有好处,但听起来确实很贵。不过我觉得它的昂贵程度是有上限的,对吧?推理成本会随着时间下降。你会找到有效优化它的方法。当它暂停时,你没有收到人类输入,就不需要生成任何东西,对吧?

So we covered a similar thing with Flipbook before, in our episode with Ethan, with Grok video. I think it's very engaging visually. I think it's maybe good for education, but it does sound expensive. I think there's an upper bound to how expensive it will be though, right? Like, the inference cost will go down over time. You'll figure out ways to optimize it effectively. When it pauses, you're not receiving human input. You don't have to generate anything, right?

Anastasis

是的,我是说,你也可以——在这个例子里你有环境运动。所以屏幕上有些部分可能——比如说你想去巴黎,然后你得到这个可以探索的界面,里面有人在走动或者有事情在发生。这显然会让它更贵,因为你需要一直运行模型。也许你可以用某种循环机制,这样就不需要一直运行。但所有这些我觉得都是我们需要搞清楚的事情。

So yeah, I mean, you could also — in this case you have ambient motion. So there are parts of the screen that might — if, let's say, you want to visit Paris and then you get this interface that you allow to explore, you have people walking or things happening. It obviously makes it more expensive because you need to run the model all the time. Maybe you have some looping mechanism so you don't need to do that. But all those things I think are stuff we need to figure out.

Host

是的。

Yeah.

Anastasis

我觉得我们首先考虑的是:让我们先明确——找到一些用例,在这些用例里它相比传统界面明显是更有吸引力的交互,然后它变得更具成本效益就只是时间问题。

I think our first consideration is: let's make this clearly — find some use cases where it's clearly a much more compelling interaction compared to traditional interfaces, and then it's a matter of time before it becomes more cost-effective to serve.

Host

是的。说到那些走动的人,我觉得对我来说最合理的方案基本上就是 Moon Lake Nick。

Yeah. When it comes to the people walking, I think the approach that makes the most sense to me is basically Moon Lake Nick.

Anastasis

天哪,我总是把他们的名字和 Chris Manning、Funen 搞混。我不知道你有没有遇到过他们,他们基本上映射到某种游戏引擎——我觉得是 Unity 之类的,或者 Godot——你显然可以在后面脚本化一些 NPC 行为并在此基础上训练。而在这里,你可以真正想象任何你想要的东西,比如那是一个 UI,对吧?而且我觉得,创建一个可交互的软件世界模型似乎更可行,因为我们有很多这样的例子,你可以在上面做你那些花哨的环境相关的事情。然后它再扩展到具身和真实世界的物理用例,但这是一个很好的第一步。

Oh god, I keep messing up their name with Chris Manning and Funen. I don't know if you've come across them, where they basically map to some kind of game engine — I think it's Unity or something, or Godot — and you can obviously script some NPC behavior behind that and train on that. Whereas here you can really imagine whatever you want, like that is a UI, right? And it feels like more tractable, I guess, to create a world model of software that is interactable, because we have many examples of that, and you can do your fancy environment stuff on that. Then it is scaling up to embodied and real-world physical use cases, but this is a nice first step.

Host

或者你知道,还有相反的情况——你有 10 亿参数的模型,3.5 亿参数的语言模型,小到它们只是在预测鱼在游动。

Or you know, there's the opposite — you have like 1B models, 350 million parameter language models that just get so small that they're just predicting like fishes moving.

Anastasis

我是说,小模型不是 120 亿参数。所以——

I mean, small models are not 12B. So —

Host

超小型设备端,不过不,我觉得这让人看清了全貌。至少对我来说汽车那个例子,就是应用,对吧?做那件事的工作量。当然,你每年只造一款车型年的车,但应用这个,相比手动做所有这些也是一种成本节约,对吧?所以它也打开了很多可能性。我很好奇,如果你把这个延伸到 2-3 年后。你觉得事情会进一步走向哪里?

Ultra mini on device, but no, I think it puts it into perspective. At least the car one for me, like the applications, right? The amount of work to do that. Sure, you only make one model year car per year, but applying this, it's also a cost-saving to have to manually make all this, right? So it opens up a lot of possibilities, too. I'm curious if you extend this out 2-3 years. So where do you see things going even further?

Anastasis

像界面世界模型这样的东西,最终形态实际上就是你拥有一个完全神经化的操作系统。我记得 Andrej Karpathy 很久以前写过这个。但对我来说有点奇怪的是,比如今天你和语言模型交互时,这个语言模型基本上可以和你聊任何东西。你可以把对话引向任何方向,它非常通用,所以能解决各种不同的任务,但你通过一个非常僵硬的界面与它交互。所以对我来说,界面本身变得可学习、成为整个循环的一部分只是时间问题——就像你不再只是端到端地交付一个应用,这意味着你交付语言模型,但你也交付渲染和像素,然后那也是一个可学习的组件。应用这个概念可能不一定——我觉得我们需要为软件想出新的抽象。应用这个概念来自这样一种想法:你需要独立的代码库来描述、驱动每个单独的工具和每个单独的应用,但如果你有一个视频模型在你操作时实时生成界面,你可能会想到一种统一得多的东西。所以它可以从 LLM 获取上下文,让你把传统上存在于不同应用中的不同功能组合起来。所以这实际上是一种端到端解决软件的方式。我们也把这看作训练计算机使用智能体的一种强大方式。所以这是看待它的一种方式,而一般来说,世界模型有两个方向。一个是为人类服务的世界模型,一个是为智能体服务的世界模型,用来训练智能体。

Effectively the end game of something like interface world models is you have a fully neural operating system. So I think Andrej Karpathy has written about that quite a while back. But to me it's a bit odd that, for example, with an interaction with an LM of today, you have this LM that can basically talk to you about anything. You can take the conversation any direction, it's very general, so it can solve all those different tasks, but you interact with it through a very rigid interface. And so to me it's just a matter of time before the interface itself becomes learnable and becomes part of the whole loop — like you're not just delivering an application end to end, and that means you're delivering the language model but you're also delivering the render and the pixels, and then that's also a learnable component. And the concept of applications might not necessarily — I think we need to figure out new abstractions for software. The concept of application comes from this idea that you need separate code bases to describe, to power each individual tool and each individual application, but you might think of something a lot more unified if you have a video model that's actually generating the interface as you go. So it can take context from an LLM and allow you to combine different functionalities that traditionally would live in different applications. So it's a way to solve software end to end, effectively. We also see this as a powerful way to train computer-use agents as well. So this is one way to see this, and in general with world models there are those two directions. One is world models for humans and world models for agents, to train agents.

Anastasis

所以对于我们做的每一项新的世界模型工作,这两种用途都成为可能。所以这是一个强大的合成数据生成器,用于训练计算机使用模型。它可以成为一个实时的强化学习环境,你可以用它和计算机使用智能体做在线强化学习,你可以获得各种各样的交互、各种界面,你只需即时生成,来提升你的智能体的鲁棒性。所以对于我们正在为机器人用例做的世界模型来说也是一样。

And so for every new work of world models that we do, these both uses become possible. So this is a powerful synthetic data generator for training computer-use models. It could become a live RL environment that you could use to do online RL with a computer-use agent, and you can get a wide diversity of different interactions, kinds of interfaces you just generate on the fly, to improve how robust your agent becomes. So that's the same also with the world models that we're working on for a robotics use case as well.

Host

有没有一个你正在等待的研究突破,能解锁你真正想追求的下一批用例?

Is there a research breakthrough that you're waiting for that would unlock the next set of use cases that you really want to pursue?

Anastasis

长上下文是非常重要的一项。

Long context is a very important one.

一致性与误差累积 Consistency and Error Accumulation

Anastasis

所以能否长时间保持一致性取决于具体用例。比如我们的角色模型,或者界面世界模型,维持长时交互会更容易。如果你进入更开放的世界,在其中导航并执行任意动作,那么你能生成的上下文和时长会更快受限。所以我们看到更多退化和误差累积。自回归模型最大的挑战就是误差累积:基本上你是把生成的帧反馈回模型来生成下一帧,如果有任何小误差,它们会随时间累积。这不是新问题,LLM 也有这个问题,我们已经看到能生成非常非常长的输出。所以这是一个已解决的问题,但肯定仍是一个挑战。

So being able to maintain consistency for long periods of time depends on the use case. For our characters model, for example, or for the interface world model, it's easier to maintain long sessions of interaction. If you go into more open-ended worlds that you navigate and you take arbitrary actions in, the context at which you can and duration which you can generate becomes limited much more quickly. So we see more degradation and error accumulation happening. The biggest challenge with autoregressive models is error accumulation: basically you're fitting generative frames back into the model to generate the next frames, and if there are any small errors, they accumulate over time. That's not a new problem. It's a problem that LLMs also have, and we've seen the ability to generate really, really long outputs. So it's a solved problem, but it's definitely still a challenge.

Host

是的。那么目前最先进的技术是什么?比如对 Grok 来说,视频上下文大概是 10 到 20 秒。

Yeah. And what is the state of the art? So for Grok it would be like 10 to 20 seconds of context going in there for video.

Anastasis

用我们的角色模型,我们能够自回归地生成长达 30 分钟的视频。

With our characters models we're able to generate up to 30 minutes of video autoregressively.

Host

是的。那只是针对虚拟形象。

Yeah. That's just for the avatars.

Anastasis

是的。所以,如果我们看 GWM Worlds,那更像是我们的开放式世界探索模型,它大概是几分钟的量级,

Yeah. So, if we look at GWM Worlds, which is more our open-ended world exploration model, it's on the order of a few minutes,

Host

那是的。

which is Yeah.

Anastasis

对人们来说可能足够了,因为无论如何你都得切到下一个场景,对吧?

it's probably enough for people because you have to cut to the next scene anyway, right?

Host

是的。如果每隔几分钟就得重启,那并不是理想的游戏体验。所以我认为,但对于某些类型的体验,你可以绕过它。理想情况下,你能够永远生成下去而不退化,我认为我们达到那一步只是时间问题。

Yeah. It's not the ideal game experience if you have to restart every few minutes. So, I think but I think it's yeah, for certain kinds of experiences, you can work around it. Ideally you're able to just generate forever and it doesn't degrade and I think that's a matter of time before we get there.

Host

是的。Genie 大概最多一分钟,你知道。

Yeah. Genie has like one max one minute you know.

Anastasis

是的。这是你做的关于机器人技术的研究。我想我也有你们的 Runway 机器人技术页面。

Yeah. This was your you did a study on robotics. I think I also have just your runway robotics page though.

Host

这个更好吗?

Is this better?

Anastasis

去年我们发布了 Gen 4.5。那是我们最新的基础模型,正如我提到的,我们一直在世界模型方面做所有这些工作,而我们对世界模型的大部分方法本质上是:如何将双向扩散模型变成自回归的,并让它接受动作。所以它不再是你观看的视频,而变成了一个你可以步入的模拟,你可以全程控制它。你可以探索反事实,比如如果我采取这个动作 versus 如果我采取那个动作会发生什么。而 GWM1 是我们在 Gen 4.5 之上构建的世界模型。所以我们在 Gen 4.5 之上做了所有这些自回归和步蒸馏。我们看到 GWM1 最大的用例之一是在机器人技术中。我们喜欢说的一句话是,随着我们扩展视频模型,我们意外地通过仅仅扩展视频模型创造了一个最先进的机器人技术模型。所以我们去年年中某个时候意识到,机器人实验室开始来找我们,有点要求使用视频模型来生成合成数据,有点要求我们后训练我们的视频模型,使其在机器人技术方面表现非常好,这样他们就可以用那个来生成变体。那是我们看到的第一个用例,然后越来越清楚的是,这些模型将不仅仅用于创建合成数据来训练机器人策略。它们作为模拟器也非常有用。这意味着你可以在线使用视频模型来测试你的机器人动作模型表现如何。所以你可以采取一个动作,然后在世界模型内得到该动作的结果,然后继续这个循环,就像这个闭环模拟,你可以用它来评估你的机器人模型工作得有多好。而我认为如果你想构建一个模拟器,你需要解决的最大问题是建立真实世界相关性:如果你在世界模型内采取一个动作,如果你在真实世界中采取相同的动作,你会得到类似的结果。所以那是我们今年早些时候做的一些工作的目标。所以如果你去第一个链接。那本质上是我们想为我们的世界模型建立真实到模拟的相关性,这样如果你在世界模型内做一系列动作,如果你在真实世界中做相同的动作,你会得到类似的结果。我们拿了我们的 GWM1 模型,我们使用了一些基准数据,有一个叫 RoboArena 的基准,非常常用于评估不同的动作模型表现如何,我们在世界模型内使用了相同的场景和具身。我们测量了动作模型在世界模型内 versus 在真实世界中表现如何的相关性。我们看到我们可以在我们的世界模型和现实之间获得非常好的相关性。这意味着如果你想评估你的机器人策略表现如何,你可以在模拟中更快地扩展,而不必用实际的物理硬件来做。所以那是我们的模型在机器人技术中可能非常有用的第一个迹象。我们看到,随着我们与机器人实验室合作,那成为了第一个用例,他们可以以适合他们训练管道的方式使用视频模型。

So last year we released Gen 4.5. So that was our latest base model and we've been as I mentioned we've been doing all this work in world models and which essentially a lot of our approach to world models is how do you take a bidirectional diffusion model and make it autoregressive and make it accept actions. So instead of being a video you watch, it becomes a simulation that you step in and you can, you know, control it every step of the way. You can explore counterfactuals like what happens if I take this action versus if I take this action. And GWM1 was the it's the world model that we built on top of Gen 4.5. So we did all this autoregressive and then step distillation on top of Gen 4.5. And one of the biggest use cases that we saw for GWM1 was in robotics. One thing we like to say is we as we scaled video models, we accidentally created a state-of-the-art model for robotics by just scaling video models. So we realized at some point mid last year that robotics labs started coming up to us and kind of asking to use video models for synthetic data, kind of asking us to post train our video models to work really well for robotics so that they can use that to basically generate variations. That was the first use case that we saw and then increasingly became clear that the models will be useful beyond just creating synthetic data to train robotic policies. They would also be very useful as simulators. So that means that you can use a video model online to test how your robotic action model performs. So you can take an action roll and then get the outcome of the action inside the world model and then continue that loop like this closed loop simulation and you can use that to evaluate how well your robotics model works. And the biggest thing that I think you need to solve if you want to build a simulator is establishing real world correlation that if you take an action inside the world model if you take the same action in the real world you get a similar outcome. So that was a goal of some work that we did earlier this year. So if you go to the first link. So that was essentially we wanted to establish that you know real to sim correlation for our world model so that if you do a series of actions inside the world model and if you do the same actions in the real world you get similar outcomes and we took our GWM1 model and we used some benchmark data that there's this robarina benchmark that's very commonly used to evaluate how well do different action models perform and we use the same scenarios and embodiments inside our world model. And we measured the correlation of how well did a action model perform inside a world model versus in the real world. And we saw that we could get very good correlation between our world model and reality. And that means that if you want to evaluate how well your robotic policies perform, you can scale that much faster inside simulation instead of having to do that with actual physical hardware. And so that was a first indication that our models could be quite useful in robotics. And we saw as we're working with robotics labs that that became like the first use case where they could use video models in a way that fit into their training pipeline.

Host

我能问一下从 4.5 到解决那个问题有什么不同吗?所以模拟到真实的差距一直是个问题,对吧?你在视频数据上训练机器人模型,它不能泛化到真实世界,而模拟也有问题。所以,看起来你解决了它。但怎么解决的?

Can I ask what the difference was from 4.5 to solving that? So the sim to real gap has always been the issue, right? you train a robotics model on video data, it doesn't generalize to real world and the simulation had an issue. So, seems like you solved it. But how?

Anastasis

是的。所以模拟器的一个大问题是,如果你试图模拟刚性物体,如果你能非常准确地描述物体的物理特性,那它工作得相当好。然后你能够使用 Isaac Sim 或 MuJoCo 或其中一个传统模拟器。但对于更复杂的交互,比如布料或光滑表面,你知道,对于你想要用动作模型解决操作任务的所有复杂性,为每个环境和每个任务构建模拟版本、那个环境的数字域,是非常困难且耗时的。而用世界模型,你只需要提供第一帧,然后你就可以在第一帧内展开策略。所以,我们将其与需要 3D 扫描环境然后 3D 扫描每个单独物体的方法进行了比较,而现在你可以带来那个模拟。而用世界模型,你只需要拍一张环境的照片,然后你就能测试你的策略表现如何。

Yeah. So, a big problem with simulators is, you know, if you're trying to simulate rigid objects, like it works quite well if you can describe the physics of objects very accurately. Then you're able to use Isaac Sim or MuJoCo or one of the traditional simulators. But for more complex interactions with cloth for example or like slippery surfaces, you know with all the complexity that you want to be able to solve with a manipulation with an action model that solves manipulation tasks. It's very difficult and so time-consuming to build you know for each of those environments and each of those tasks build the simulated version of that the digital domain of that environment. Whereas with a world model you just need to provide the first frame and then you just can roll out the policies inside the first frame. So whereas you know we compared it to methods that required like 3D scanning an environment and then 3D scanning each individual object before you can now you know you can bring that simulation. Whereas with a world model you just take a picture of the environment and then you're able to test how your policy performs.

第三人称视频作为关键数据源 Third-Person Video as the Key Data Source

Anastasis

我们对机器人技术的总体论点是,有些公司正在利用大量遥操作数据来训练机器人动作模型。现在有些公司使用 UMI 数据,这本质上是人类第一视角视频,其中人类使用机器人夹爪执行不同的操作任务。还有些公司专注于第一视角数据,就是把 GoPro 绑在某人头上,然后捕捉他们执行任务的过程。我们认为这些都是训练机器人模型的绝佳数据来源,但最丰富的视频数据来源是第三人称视频数据。我们人类是如何学会执行不同任务的?很多是通过观察他人执行这些任务。我们不是从第一人称学习的。我们显然会通过试错来学习不同的事情,但最终我们在世界上学会做的很多事情,都是通过观看别人做而学会的。这就是当你预训练视频模型时,你本质上在做的事情。它是大量的第三人称视频素材,展示人们在世界上执行不同任务,人们做运动,人们做家务。我们的主要论点是,视频预训练——一旦你做到了,你就可以用少得多的实际机器人数据,将模型适配到机器人用例中。所以你需要的遥操作数据要少得多,而遥操作数据非常难以规模化。即使你看第一视角数据,它比需要实际硬件的遥操作数据更容易规模化,但与世界上存在的第三人称视频数据相比,仍然少了三个数量级。所以我们的论点是,最丰富的数据来源最终会胜出。第三人称视频数据预训练是模型的正确起点,你希望这些模型能够泛化,能够处理新环境、新任务,以及训练期间未见过的事物。这就是为什么我们认为我们的模型在机器人场景中特别有用的动机。我们也看到了情况确实如此。

Our general thesis on robotics is that there are companies leveraging a lot of teleoperation data to train robotics action models. There are now companies using UMI data, which is essentially human egocentric video where humans use robotic creepers to perform different manipulation tasks. And then there are companies focusing on egocentric data, which is you strap a GoPro on someone's head and capture them performing a task. We think all those are great sources of data for training robotics models, but the most plentiful source of video data is third-person video data. How do we as humans learn how to perform different tasks? A lot of it is by observing others perform those tasks. We don't learn from first person. We obviously do some trial and error to learn different things, but ultimately a lot of what we learn how to do in the world, we learn by watching other people do it. And that's how when you're pre-training a video model, you're essentially doing that. It's a lot of third-person video footage of people performing different tasks in the world, people doing sports, people doing household tasks. And our main thesis is that video pre-training—once you do that, you can then adapt a model to be useful in robotics use cases with way fewer hours of actual robotic data. So you require way less teleoperation data, which is very difficult to scale. And even if you look at egocentric data, which is a bit easier to scale compared to teleoperation data that requires actual hardware, it's still three orders of magnitude less of that that exists in the world compared to third-person video data out there. So our thesis is that the most plentiful source of data will ultimately win. Third-person video data pre-training is the right starting point for models that you want to generalize and be able to deal with new environments, new tasks, things that you haven't seen during training. That's the motivation for why we think our models are especially useful in robotics settings. And we've seen that to be the case as well.

Host

你提到了预训练。所以也许就像是第三人称预训练,第一人称 SFT。这是你可以引入的一种课程吗?

You said pre-training. So maybe it's like third-person pre-training, first-person SFT. Is that like a curriculum that you can sort of introduce?

Anastasis

正是如此。所以如果我们看 GWM worlds,GWM 机器人,GWM 机器人是从 Gen 4.5 开始的。

Exactly. So if we look at GWM worlds, so GWM robotics, the GWM robotics starts from Gen 4.5.

Host

是同一个视频扩散主干网络,对吧?

It's the same video diffusion backbone, right?

Anastasis

正是,是的。所以你从基础视频模型开始,就是那个用来生成猫和狗以及其他有趣东西的模型,然后在非常少量的机器人数据小时数上进行微调。所以大约是数百小时的数量级,相比之下,如果你要预训练一个机器人模型,目前的预训练数据量高达数十万甚至数百万小时。你能够快速获得相当好的性能,因为模型利用了它在预训练中学到的关于世界、物理、人类动态以及人们关心的任务的所有知识。最终,你希望这些模型能够泛化。你不想只能执行训练期间的任务。而预训练视频数据集中的动作和环境多样性,远大于你手动实际能捕捉到的。

Exactly, yeah. So you start from the base video model, the one you're using to generate cats and dogs and other interesting stuff, and then you fine-tune on a very small number of hours of robotic data. So it's something on the order of hundreds of hours, compared to if you were to pre-train a robotics model, the current pre-trainings go up to hundreds of thousands or millions of hours of data. And you're able to get quite good performance quickly because the model leverages all the things that it has learned about the world and physics and human dynamics and the tasks that people care about from pre-training. And ultimately, you want those models to generalize. You don't want to just be able to perform the task that has been during training. And the diversity of actions and environments that you have with a pre-training video data set is much larger than what you can realistically capture manually.

Host

后训练的规模看起来怎么样?就像你仍然想做的是,大约 90% 的算力在常规视频扩散模型中,然后大幅扩展,还是说我们想要不同的机器人模型用于不同的任务,或者只是一个基础非常好的世界模型也可以应用于机器人。

How's the scale looking like for the post-training? Like you still want to do is it like roughly 90% of the compute in regular video diffusion model and then scale up a lot or do it like we want different robotic models for different tasks or just the one base really good world model can also apply to robotics.

Anastasis

所以目前我们正在为特定合作伙伴的特定具身进行后训练。因此,如果他们有一种特定的单臂机器人、双臂机器人或 Shimano 机器人,我们会在他们的特定数据集上后训练我们的 GWM 机器人模型。随着时间的推移,我们看到 GWM 的不同变体正在统一。就像我预期的那样,你知道,一年后或两年后,你会有一个单一的世界模型,可以模拟操作任务。它可以模拟导航,很多游戏世界模型都是导航世界模型,你在空间中移动,它也会模拟人类行为。所以那就是角色模型。因此,不是有三个不同的模型,而是有一个单一模型能够,理想情况下,你能够模拟身处世界的感觉。你在环境中移动。你可能在执行不同的任务。你在与其他人交谈,而这一切都发生在同一个单一的实时视频模型中,由它生成。

So currently we are post-training our models for specific embodiment that we for particular partners. So if they have a particular kind of single arm robot or bimanual robot or Shimano robot, we would post-train our GWM robotics model on their particular data set. Over time we see the different variants of GWM unifying. Like I would expect, you know, if a year from now or two years from now you have a single world model that can simulate manipulation tasks. It can simulate navigation, which is a lot of the gaming world models are navigational world models, you're moving around the space, and it will also simulate human behavior. So that's the character models. So instead of having three different models, you have a single model that's able to, you know, ideally you're able to simulate what it's like to be in the world. You're moving around an environment. You're maybe performing different tasks. You're talking to other people, and that happens with the same single kind of real-time video model that's generating that.

Host

你认为你能解决自动驾驶吗?所以如果你在模拟器中学习驾驶汽车,你有一个世界模型,你的机器人基本上是,你知道,汽车可以操纵这么多轴。你离这样的东西还有多远?

Do you think you can solve self-driving? So if you are learning to drive a car in a simulator, you have a world model, your robot is basically, you know, car can manipulate so many axes. How far off are you from something like that?

Anastasis

所以一个非常好的 ADAS 系统。世界模型目前确实正在应用于自动驾驶研究,主要用于评估用例,但我们的重点更多在机器人操作上。我们也做了一些 AV 世界模型的工作。但是的,我们确实认为世界模型和视频模型是模拟器以及策略和动作模型的最佳起点。所以这是另一方面,一旦你有了一个很棒的世界模型,那么你只需添加一个动作头,它也可以预测动作。一种思考方式是,如果你取一个带有机械臂的场景的起始帧,然后你提示模型,生成机械臂拿起物体的视频,如果它生成的视频足够准确,那么它也应该能够生成机械臂执行相同动作所需的确切 3D 姿态。所以这就是现在的方向,流行的术语是世界动作模型,即你从视频模型开始,然后添加一个动作头来预测动作,它本质上就变成了一个策略。

So a really good ADAS system. World models are definitely being applied to self-driving research right now, mainly for evaluation use cases, but our focus has been more on robotic manipulation. We've done some work on AV world models as well. But yeah, we do think that world models and video models are the best starting point for both simulators and also policy and the action models. So that's the other side to this is that once you have a great world model, then you can just add an action head and it can predict actions as well. One way to think about it is if you take the starting frame of a scene with a robotic arm and you ask, you know, you prompt the model, generate the arm picking up an object, and if it generates an accurate enough video, then it should also be able to generate the exact poses in 3D that the arm should take to perform the same action. So this is the direction that's now the popular term for it is world action models, which is you're starting from a video model and then you're adding an action head to predict the actions and it becomes a policy essentially.

Host

还有一件事让我印象深刻的是,训练这类模型实际需要多少数据。你可能无法确切说出多少,但就像你知道的,最初的扩散模型,据我所知,甚至开源的中国模型,数据量并不大,这难道不令人惊讶吗?

One thing I'm also impressed by is how much data you actually need to train these kinds of models. You probably can't say exactly how much but like you know like the original diffusion models and from what I know even of the open source Chinese models is not that much data isn't it surprising?

Anastasis

你如何定义大量数据?

What do you define as much data?

Host

嗯,是的,它只是进进出出。Token 数量仍然相关吗?

Um yeah, it just comes goes in. Is the token count still relevant?

Anastasis

所以这有点复杂,而且

So it's a bit more complicated and

Host

我的意思是,只是千兆字节,对吧?

what is I mean just gigabytes, right?

Anastasis

是的。视频小时数。

Yeah. Hours of video.

Host

是的。是的。

Yeah. Yeah.

Token到参数:语言vs视频模型 Tokens to Params: Language vs Video Models

Host

有意思的是,语言模型里的 token 与参数比似乎领先了大概三年,而且仍然比视频模型高很多,尽管从技术上讲视频每个 bit 的信息量更大。我不确定这是否直观,也许只是像素之间的变化并不大。所以也许只是有大量重复的信息。

Something that's interesting is it seems like the tokens-to-params ratio in language models is maybe three years ahead or whatever, and seems to be a lot higher than video models still, even though technically video has more information per bit. I don't know if it seems intuitive, or maybe there's just a lot of variability between one pixel and the next pixel is not that high. So maybe there's just a lot of information that is repeated.

Anastasis

我的回答是,现在还很早。视频模型的训练规模会远远超过目前,你会看到比当前模型强得多的能力。所以我喜欢用这样一个思想实验——它几乎就像是视频模型的图灵测试,或者世界模型的图灵测试。我称之为清醒梦测试。假设你有一个——

My answer would be it's still very early. The training of video models will scale way further than it currently is, and you'll have capabilities that go much further than the current models can do. So one thought experiment that I like to use—it's almost like the Turing test of video models, or the Turing test of world models. I call it the lucid dream test. So you have a—

Host

你是说像真人那种清醒梦?

You mean like the actual person lucid—lucid dream?

Anastasis

这来自 LucidRains,对吧?

It comes from LucidRains, right?

Host

清醒梦就是你在做梦时知道自己在做梦——

Lucid dreams is telling you're dreaming while you're in a—

Anastasis

对,没错。清醒梦就是当你意识到自己在梦里,然后基本上能控制梦里发生的事。

Yeah, exactly. So lucid dreaming is when you realize you're inside a dream and then you can basically control what happens in it.

Host

不,还有一个做推理的人叫 LucidRains。对。量化——

No, there was also an inference guy called LucidRains. Yeah. Quantization—

Anastasis

非常多产的人。对。那么,假设你有一个 VR 头显,你戴着它待在一个房间里。如今大多数 VR 头显都有透视模式,你可以直接看到现实世界前方的东西,或者你显然可以在 VR 头显里渲染一些东西。总有一天,那些交互式实时视频模型会变得足够好,你戴上头显,待在同一个房间里,四处走动,和物体互动。你能在那个房间里自由移动,和任何物体互动。最后有人问你:你用的是透视模式,还是这其实是渲染或生成的画面?如果你无法确定你互动和移动时看到的是生成的,还是透视模式下的真实场景——那就说明模型已经足够好了。而我们离那还很远。很大一部分就在于真正模拟动态和反事实。比如,如果你让一个视频模型生成一个人进球和一个人没进球的视频,它在进球上会做得更好,因为训练分布有偏差。关于人成功进球的视频多得多。但如果你有一个交互式模型,你希望能生成反事实——比如我采取这个动作和那个动作,你希望生成同样逼真的结果。所以我认为,视频模型和世界模型之间的巨大差距就在于反事实生成这个想法。如果你想为机器人打造一个出色的模型,你希望很好地模拟失败,因为无论你是用它做评估,还是将来用在在线强化学习循环里,你都希望模型能尝试、失败并改进。所以要做到这一点,你需要能模拟失败。这是唯一一个成功案例太多、失败案例不够的领域。

Very prolific person. Yeah. So, let's say you have a VR headset and you're in a room wearing a VR headset. Most of today's VR headsets have a passthrough mode so you can see directly what's in front of you in the world, or you can obviously render something inside the VR headset. And there's going to be a point where those interactive real-time video models become good enough where you wear the headset and you're in the same room and you're walking around and you're interacting with objects. You're able to move freely in that room and interact with any object. And then at the end someone asks you: were you using passthrough mode, or was this actually rendered or generated footage? And if you cannot tell for sure whether what you were seeing as you were interacting with and moving around the world was generated or it was passthrough mode and was just what was happening in front of you—that's an indication that the models have become good enough. And we're not close to that yet. And a lot of it is just this idea of really simulating dynamics and counterfactuals. Like, if you ask a video model to generate a person scoring a goal versus a person failing to score a goal, it would do a better job at scoring the goal because there is a bias from the training distribution. There are a lot more videos of the person succeeding at scoring the goal. But if you have an interactive model, you want to be able to generate counterfactuals—like if I take this action versus this action, you want to generate equally realistic outcomes. So that's, I think, the big gap between video models and world models: that idea of counterfactual generation. And if you want a great model for robotics, you want to simulate failure very well, because whether you're using it for evaluation or you're using it in an online RL loop in the future, you want to be able to have the model try and fail to do things and improve. And so in order to do that, you need to be able to simulate things failing. This is the only domain where you have too many successful examples and not enough bad examples.

Host

生成失败应该很容易才对。

It should be easy to generate failure.

Anastasis

奇怪的是,我认为早期的图像和视频模型并不擅长做到像真人那样逼真。比如你看到很多高分辨率 4K 的专业摄影,但不是日常生活——不是普通的照片。一切看起来都像是专业生成的,像专业照片,而不是普通的——你知道,桌上乱糟糟的线缆。

Oddly enough, I think early image and video models weren't good at being human-realistic. Like you see a lot of high-res 4K professional photography, but not just everyday life—like a normal picture. Everything looks like it's professionally generated, like professional pictures, but not just like normal—you know, messy cables on a desk.

视频智能体与框架 Video Agents and the Harness

Host

好。所以有这些东西。我们还谈到一件事,就是你们推出了视频智能体。我想问,传统的——我们称之为前沿的——自回归 LLM 如何融入这一切?它们是在驱动你们的机器人模型,还是在驱动你们的视频智能体生产?你在哪里看到自回归和扩散的重叠,我们姑且这么说?

Okay. So there's this stuff. One thing we also covered is that you guys have video agents that you launched. I guess, how does the traditional—let's call it frontier—autoregressive LLMs feed into all this? Are they driving your robotics models, or are they driving your video agent production? Anything where you see the overlap of autoregressive and diffusion, let's call it?

Anastasis

是的,所以 harness 在所有这些不同的用例中都非常重要。我们有一个视频智能体,它本质上是一个 LLM,非常擅长使用各种图像模型、视频模型的工具,并帮助你端到端地完成一个项目。所以在传统的广告流程中,你常常从一个 brief 开始,然后生成故事板,再生成视频。视频智能体和 Runway 智能体会帮助你完成整个过程,还帮你分析效果数据。比如,这个广告和那个广告相比表现如何,然后根据这些经验生成更多——弄清楚该生成什么。我们认为 harness 是流水线中非常重要的一环。正如我提到的,所有视频制作——所有生产级视频模型——都会用到某种提示补全,我们预计这会变得越来越复杂,越来越——你知道,在你使用扩散 Transformer 之前,你会生成更长、更详细的描述。我确实认为最终会越来越统一成 omni 模型,你端到端地训练模型,既能做自回归文本预测,也能做扩散。所以你在预测下一个 token——你可能会对场景做一些推理和规划——然后把它传给扩散头,由它实际生成像素。是的,我认为目前可能只有 Gemini 和 Qwen 这么做。我不确定哪些中国模型是 omni 的,但没错,这并不是一种很普及的模态。我觉得想想的话,这是个有趣的用例,对吧?因为你不仅要在语言模型推理、扩散头生成那里结束——你不必在那里输出。你可以回去把那个输出喂给同一个模型,再次推理改进,它可以自己循环很多次。

Yeah, so harnesses are really important across all those different use cases. So we have this video agent which is essentially an LLM that is very effective at tool use of different image models, video models, and kind of helps you through creating a project end to end. So very often in a traditional advertising flow you have a brief you start from, and then you generate a storyboard, and then you generate the video. A video agent and Runway agent kind of helps you through that whole process, and it helps you also analyze performance data. For example, how well did this ad perform versus this ad, and then generate me more based on those learnings—figure out what to generate. We think that the harness is a very important piece of the pipeline. As I mentioned, all the video production—all the production video models—use some prompt completion that happens, and we expect that to become more and more complex and more—you know, you generate longer and more detailed descriptions before you use the diffusion transformer. I do think eventually there's increasingly this unification into omni models where you're training the models end to end to both do autoregressive text prediction and also diffusion as well. So you're predicting the next token—you do maybe some reasoning and planning of the scene—and then you're passing it into the diffusion head that's actually generating the pixels. Yeah, I think currently maybe only Gemini and Qwen do it. I'm not sure which of the Chinese models are omni, but yeah, it's not a very well-popularized modality. I guess it's an interesting use case when you think about it, right? Because not only do you have to end at like language model reason, diffusion head generate—you don't have to output there. You can go back and feed that output to the same model, reason again on improvements, and it can do a lot of loops just on its own.

框架vs端到端学习 Harness vs. End-to-End Learning

Host

我想问题是,我们需要那个吗,还是我们可以只做智能体脚手架,比如在模型外面做?在模型内部做有很大的好处吗?

I guess the question is, do we need that, or can we just do agent scaffold, like do it outside the model? Is there a big benefit to doing it in the model?

Anastasis

我认为总体上有一个趋势:某件事先由外部框架完成,然后它变成模型的一部分。所以你有思维链提示,你必须写超级详细的系统提示才能得到回复,而现在模型基本上在给出答案之前自己生成推理轨迹。在视频模型里也类似,早期的很多视频模型是单镜头视频模型,你必须用某种编排器,你转一下,并行生成多个镜头,然后把它变成真正的视频。

I think there's generally the trend of something is first done by a harness and then it becomes part of the model. So you had chain of thought prompting where you have to do this super detailed system prompts to get back, and now the model basically generates the reasoning trace by itself before it gives you an answer. And in video models, similarly, a lot of the video models of the early days were single-shot video models and you had to use some kind of orchestrators, you turn, generate multiple shots in parallel, and then turn it into an actual video.

Host

ComfyUI 到处都是,你知道,意大利面式的工作流。

ComfyUI just all over the, you know, spaghetti workflow.

Anastasis

而现在你有多镜头视频生成,你直接生成多个镜头,这有好处,因为这样视频模型就学到一些——你知道,要生成好一个单镜头,你显然需要弄清楚很多关于世界的东西。要生成多镜头视频,你也需要基本上获得一些视频剪辑的直觉,比如你需要弄清楚镜头的正确节奏。而且语言模型在这方面并不擅长。它们不是很好的视频剪辑师。如果你让语言模型拿一些视频然后自动从中创建剪辑好的视频,它会感觉诡异。所以我认为语言模型实际上还不那么擅长做视频剪辑。我认为端到端学习这个有好处。所以我预计,你知道,总体趋势是你需要外部框架的事情最终会被注入模型本身,你端到端地学习它。

And now you have multi-shot video generation where you directly generate multiple shots, and there is a benefit to that because then the video model learns some—you know, to generate a single shot well you need obviously to figure out a lot of stuff about the world. To generate multi-shot video, well, you also need to basically get some like video editing instincts, like you need to figure out what is the right pacing of shots. And also LMs are not that good at it. Like they're not that great video editors. If you ask a LM to take some videos and then kind of auto-create an edited video out of that, it would feel uncanny. So I don't think LMs are actually that good yet at being video editors. And I think there's benefit to learning that end to end. So I would expect, you know, the trend in general is the things that you need a harness for eventually get kind of injected into the model itself and you learn that end to end.

Host

你发现你需要雇佣能——或者同时也是艺术家的研究员来注入那种品味,还是你有艺术家和驻场来蒸馏他们?

Do you find that you need to hire engineers who can—or researchers who are also artists to infuse that taste, or do you have kind of artists and residents to distill them?

Anastasis

我们有一个庞大的创意团队,非常积极地参与训练那些模型,在每一个环节,比如如何给视频写好字幕,以便尽可能详细地捕捉你需要的电影摄影、美学、镜头方向,这样在推理时你就能通过模型引出这些。你知道,我们的创意团队也做很多评估,比如,你知道,什么构成那些模型生成的可用的视频。所以他们非常参与整个过程的每一个部分。我认为这是 Runway 特别之处之一,就是创意人员和研究员并肩坐在一起,一起工作来构建我们下一代模型。我认为这对我们作为公司的运作方式来说是非常非常重要的一环。

We have a large creative team that's very actively involved in training those models, like in every part of the way, and like how do you caption video well so that you capture the stuff that you need for like the cinematography, the aesthetics, the camera direction in as detailed ways as possible so that you're able at inference time to actually elicit that through the model. You know, we have our creative team also does a lot of evaluation of like, you know, what constitutes a usable video out of those models. And so they're very involved through every part of the process. And I think that's one of the special things of Runway is just that mix between kind of creatives and researchers kind of sitting side by side and kind of working together to build the next generation of our models. I think that's been a really really important piece to, you know, how we've operated as a company.

作为纽约公司 Being a New York Company

Host

是的。在某种意义上,你只能在纽约做这个。我是说,你有其他办公室,但就像,你知道,我试图从你是一家大型纽约公司这个事实中找到某种诗意的意义。

Yeah. In some senses you can only do this in New York. I mean, you have other offices, but like, you know, I'm trying to find some poetic significance in the fact that you are a big New York company.

Anastasis

在纽约有几个方面。显然有所有那些不同行业的交汇,比如媒体、广告,这非常广告,艺术场景就是纽约。不是说旧金山不好,但那里有更多事情在发生。有那个成分。还有我认为我们受益于作为局外人,以稍微不同的方式思考事情,比如不处于湾区 ASI 的同一个蜂巢思维中,并且——也花时间,你知道,去达到——我们今天在这里,是有意地建设和壮大团队,并引进人才,是的,既有创意方面,也有工程研究方面。纽约显然有大量优秀人才,所以这实际上不是问题。

There's a few parts to being New York. Obviously there is that intersection of all those different industries and like media, kind of advertising, like this is very advertising, the like the art scene is New York. Not to say anything bad about San Francisco, but there's more going on. There is that component. There's also I think we benefit from being outsiders and thinking of things a bit differently, like not being in the same like hive mind of ASI of Bay Area and like taking—and also taking our time to, you know, to get—we are where we are today like building and growing the team intentionally and bringing people who, yeah, both on the creative side and also on the engineering research side. There's obviously huge talent pool of amazing people in New York, so that hasn't really been a problem.

招聘与Runway的未来 Hiring and Future of Runway

Host

我是说,恭喜一切。你们在招什么?你知道,人们应该期待 Runway 的未来是什么?

I mean, congrats on everything. What are you hiring for? You know, what should people look forward to for the future of Runway?

Anastasis

我们在全面招聘。我认为这可能是 Runway 历史上我们拥有最多开放职位的时候。我们正在大幅壮大研究团队。所以如果你对视频模型感到兴奋,如果你对世界模型感到兴奋,如果你尤其对机器人技术感到兴奋,机器人团队,我们也在招聘机器人技术方面的人才,涵盖软件、硬件和研究。所以一定要联系我们。

We're hiring across the board. I think this is probably the most open roles we ever had in the history of Runway. We're growing our research team quite significantly. So if you're excited about video models, if you're excited about world models, if you're excited especially about robotics, the robotics team, we're hiring also robotics across kind of software, hardware, and research. So definitely definitely reach out.

Host

你知道,很多人没有直接的机器人背景,但如果他们想在机器人领域有用,他们应该具备什么?

And you know, a lot of people don't have direct robotics background, but what should they have, you know, if they want to be useful in robotics?

Anastasis

所以理想情况下,有一些学习策略的经验对机器人技术有好处。但我们倾向于雇佣通才作为一种理念,以及学习非常非常快的人。但一些经验和机器人领域的专业知识是我们未来几个月肯定在寻找的。然后我们正在大幅扩大市场推广团队。目前围绕视频模型有广泛且非常活跃的企业采用,我们真的在努力回应所有需求。

So ideally some experience with learned policies would be else good for robotics. But we tend to hire generalists as a philosophy and like people who learn really really quickly. But some experience and kind of domain expertise in robotics is something that we're definitely looking for for the next months. And then we're scaling the go-to-market team significantly. There is a wide like very very active enterprise adoption happening around video models at the moment and we're really trying to respond to all the demand.

开源机器人与世界模型 Open Source Robotics and World Models

Host

是的,很好。你想谈谈开源机器人相关的东西吗?

Yeah, great. You want to talk about the open source robotic stuff?

Anastasis

当然。这只是我们的一些随机笔记。英伟达推出了 Cosmos。我想这很有趣。所以你的创始成员 AI 实验室在物理 AI 中构建开源世界模型。这里有什么可谈的,是开放研究。最重要的是,正如我提到的,世界模型仍处于早期,比如你还可以进一步扩展那些模型,还有更多的进步和我们可以弄清楚的事情,或者如何进一步改进那些模型。我认为重要的是,其中一些研究在公开场合进行,并弄清楚不同公司有什么激励措施可以聚集在一起,实际上把一些研究带到公开和开源中。所以 Cosmos Coalition 是我们与英伟达共同发起的一项倡议,把一些研究作为开源带来,这可能意味着开放权重模型发布。这可能意味着衡量物理和人们在构建世界模型时关心的事情的基准。这可能意味着基础设施。所以真正的是我们如何发展世界模型的生态系统,并使它也成为对刚开始的开发者、研究员来说更容易的事情,他们对世界模型感到兴奋,可以为这个领域做出贡献。

Sure. It was just random notes we had. Nvidia launched Cosmos. I guess it's interesting. So your founding member AI labs to build open-source world models in physical AI. Anything much to talk on here is open research. The biggest thing is that as I mentioned world models are still early, like there's still so much that you can scale and those models further, so much more advancements and things that we can figure out or how to improve those models further. And I think this is important that some of this research happens in the open and figuring out what is some incentives for different companies to come together to actually bring some of that research into the open and open source. And so Cosmos Coalition was a initiative that we co-founded with Nvidia to bring some of that research as open source and that could mean open-weight model releases. It could mean benchmarks that measure physics and things that people care about when building world models. It could mean infrastructure. So really how do we grow the ecosystem of world models and make that something that also it's easier for a developer, a researcher that's just starting out that is excited about world models to kind of contribute to the field.

与中国视频模型竞争 Competing with Chinese Video Models

Host

我觉得某种程度上,这也是我们对正在发布的中国世界模型的回应吗?还是说这不在考虑范围内?

I think there's some amount of, is this also our response against the Chinese world models that are being released? Or is that not part of the consideration?

Anastasis

我确实认为这对视频模型很重要。如果你看视频模型的排行榜,我会说现在前 10 到前 20 名里大多数是中国模型。只有少数几家来自美国或西方的公司登上了排行榜。

I do think it's important for video models. If you look at the leaderboards of video models, I would say right now the majority of models at the top 10 to top 20 are Chinese models. There are only a handful of companies that have made it to the leaderboard from the US or the West.

Host

我们在图像方面做得更好,但在视频方面我们非常落后,对吧?

We're doing better with images, but with video we're very behind, right?

Anastasis

所以我认为,作为一个社区,我们更广泛地投入,确保我们有竞争力的模型,这绝对很重要。

And so I think it's definitely important that we invest more broadly as a community to make sure that we have competitive models out there.

Host

那是什么阻止我们从他们那里蒸馏呢?

Well, like what's to stop us from distilling from them?

Anastasis

我不确定那是不是最好的长期方案。认为你能变得更好,这几乎有点悲观。

I don't know if that's the best long-term. It's almost a bit of a pessimistic view that you can get better.

Host

这是免费数据。就像,你不妨这么做。如果他们在文本语言方面这么做,他们不妨在视频方面也反过来这么做。

It's free data. Like, you might as well. If they're doing it for the text language side, they might as well do it for the video side the other way.

Anastasis

是的。我的意思是,我确实认为我们目前完全有能力在没有这个方案的情况下训练出很棒的模型。

Yeah. I mean, I do think we're quite capable of training great models without this solution at the moment.

基准测试与人类评估 Benchmarks and Human Evaluation

Host

那么关于基准测试和评估,你有什么要说的吗?我感觉我听到的是很多人真的很喜欢视频和图像模型的竞技场。客户等等也是。他们只想要排行榜上最好的,而且他们比语言模型似乎更多地参考竞技场。但关于基准测试有什么要说的吗,缺少什么?普通人如何比较?嗯,这两个看起来都非常超写实。除此之外,我们谈到了机器人、模拟、物理等等,但还有什么要说的吗?

So anything you have to say on benchmarks and evals? I feel like what I'm hearing is a lot of people really like arenas for video and image models. Customers and whatnot as well. They only want the best on the leaderboard and they refer to arenas a lot more than language models seem to do. But any notes on benchmarks, what's lacking? How does the average person compare? Well, these both look really hyper realistic. More than that, outside of we did talk about like robotics, simulation, the physics and all that, but anything to say?

Anastasis

实际上我认为在某些方面是相反的。我认为人们一般,创意人士、艺术家和营销人员,也就是使用我们平台的人,我认为他们较少依赖竞技场分数,而且用一堆不同的模型生成然后视觉上比较结果太容易了。就像图像和视频模型的一个好处是,你可以立即用眼睛看出从美学角度什么感觉好,比如任何伪影,那些模型物理上的任何问题你都能立即看出来。所以实际上,作为人类来评估,我会说这比语言模型更容易。在语言模型中,你有那些非常复杂的数学和编码测试,我认为人类评估和区分我们层级模型性能要难得多。所以我认为在实践中,人们只是用一堆不同的模型测试相同的提示,看看结果如何,现在在 Runway 你可以使用我们的模型,也可以使用第三方模型。所以这很容易做到。

I actually think it's the opposite in some ways. I think people generally, creatives and artists and marketers, like people that are using our platforms, I think rely less on arena scores and it's just so easy to generate with a bunch of different models and then compare the results visually. Like one nice thing about image and video models is you can immediately tell with your eyes what feels good from an aesthetic standpoint, like any artifacts, any issues with the physics of those models you can immediately tell. And so that's actually easier, I would say, to evaluate as a human. There is also those models than it is in language models where you have those very complex kind of math and coding and tests where it becomes a lot harder I think for humans to evaluate and discriminate between the performance of our tier models at a time. So I think in practice people just test out the same prompt with a bunch of different models and see what the results look like and right now in Runway you can use our models and you can use third party models as well. So it's very easy to do that.

艺术家vs AI情绪 Artists vs AI Sentiment

Host

太棒了。我们将以 AI Runway AI 峰会结束。最后一个社会问题,我猜,我不确定这是不是个事,就是艺术家、创意人士和 AI 之间的紧张关系。那个社区里很多人讨厌 AI。显然 Runway 社区的人不介意使用工具。这只是另一支画笔。但你觉得这种情绪是如何变化的?

Amazing. We're going to end with the AI runway AI summit. The last sort of societal issue I guess, I don't know if this is a thing, is the tension between artists, creatives and AI. A lot of people in that community hate AI. Obviously the people that are in the Runway community don't mind using tools. It's just another brush. But how have you seen the sentiment change?

Anastasis

我的意思是我们的观点,是的,这只是另一支画笔。这只是另一台相机。它是一长代工具中的最新一个。技术和艺术一起演化。我认为过去几个月发生了相当显著的转变,其中一些你可以看到很多公众人物发声支持 AI,比如在戛纳你看到几位导演支持 AI。我们的电影节有朗·霍华德。还有马克·苏茨卡也在采用 AI 模型。所以每天都有更多这样的故事出现,像知名人物支持 AI,在我看来,这只是那些模型变得越来越去神秘化。我也有一个有点热门的观点,就是让那些模型最初的反应可能比需要的更激烈的原因之一是文本到视频的想法,你有一个单一的文本描述,然后得到完整的视频。是的,有一个误解。显然你可以生成我们的长片电影,但今天的模型现在采用很多参考,它们非常可控,我认为当人们看到一个允许很多自由度和控制的工具时,他们会以不同的方式回应,它是不是生成模型就不那么重要了,重要的是你可以实际引导它到你想要的方向。所以我认为当人们看到那些模型之上的复杂工作流,当他们看到你可以引导它们的所有方式,并且现在用一些最新模型可以提供多达 50 个参考,对话就变得有点不同了,因为它感觉更像一个故事,一个工具。

I mean our perspective, yes, it's just another brush. It's just another camera. It's the latest of a long generation of tools. Technology and art have kind of evolved together. I think there's been a pretty significant shift over the past few months and it came, some of it you can see with a lot of public figures speaking out in favor of AI and being, you know, like in Cannes you saw a few directors speaking in favor of AI. We had Ron Howard in our film festival. There is Mark Sutzka also adopting AI models. So you have more of those stories coming out every day of like a well-known figure kind of speaking in favor of AI and it's just a matter of, in my mind, it's those models are becoming more and more demystified. I would say I have also a bit of a hot take that one of the things that made the initial response to those models maybe a bit more heated than it needed to be was this idea of text to video, of you have a single text description and you get back a full video. Yeah, there was a misconception. Obviously you can generate it to our feature line film, but the models of today now take a lot of references, they are very controllable and I think when people see a tool that allows affords many degrees of freedom and control they respond to it differently and it matters less that it's a generative model than the fact that you can actually steer it to the direction that you want. And so I think when people look at complex workflow on top of those models, when they look at all the ways in which you can steer them and you can provide now with some of the latest models up to 50 references, like the conversation becomes a bit different because it feels much more like a story, a tool.

Host

而不是像一个神奇实体为你搞定整个电影。关于该领域人们的工作流变化有什么要说的吗?就像我认为工程至少有很多人,他们的期望已经改变,你知道,生产力提高 10 倍 100 倍,你可以完成更多。同样,你知道,你在为创意人士制作开发工具。有什么要说的吗?就像有些人不想采用,有些人想,像任何事。

Versus like something that a magical kind of entity that figures out like your entire film for you. Any notes on like workflows changing for people in the field? Like I think engineering at least has had a lot of people where they're like expectations have changed, you know, 10x 100x more productive and you can get a lot more done. Same thing is, you know, you're making dev tools for creatives. Any notes there? Like there's some people that don't want to adopt, some that do, like anything.

Anastasis

是的,所以我认为就人们关心什么而言,我看到我们经历了几个阶段。我们从人们主要追求质量的阶段开始。随着我们扩展那些模型,质量大幅提升。这是人们仍然关心的,但现在除了可控性,比如能够用参考、不同种类的输入、故事板来引导那些模型。现在我的感觉是,人们将越来越关心延迟。随着那些模型变得更好,快速迭代的能力变得更加重要,就像如果你能用一个提示几乎即时生成 10 个不同的输出,你可以比以前更快地探索,你得到一些过去创意工具特有的魔力。就像 Photoshop 是即时的。我们在生成模型上失去了一些。你在视频上要等 2 分钟才能得到结果。我认为我们现在将通过实时模型带回一些这样的体验。

Yeah, so I think in terms of what people care about, I see that we've gone through a few stages. So we started from the stage where the main thing that people were looking for was quality. As we scaled those models, the quality improved dramatically. That's something that people still care about, but it's now in addition to controllability, like being able to steer those models with references, with different kinds of inputs, with storyboards. And now my sense is increasingly people are going to care about latency more and more. As those models become better, the ability to iterate very quickly becomes more important and like if you can, with a single prompt generate 10 different outputs like almost instantly, you can explore way faster than before and you get some of the magic that characterized the creative tools of the past. Like Photoshop was instant. And we lost some of that with generative models. You're waiting for 2 minutes to get back in video. And I think we're going to bring some of that back now with the real time models.

Runway峰会公告 Runway Summit Announcement

Host

是的。激动人心,激动人心。最后我们要宣传的是这个。Runway 峰会。你们终于在旧金山办了。

Yeah. Exciting and exciting. The last thing we'll plug is this one. Runway Summit. You're finally doing this in SF.

Anastasis

是的。我们对此非常兴奋。这是在九月下旬,9 月 30 日,我们要办一个峰会,主要聚焦于物理 AI 和实时视频生成。我们有来自 Nvidia、Physical Intelligence、机器人公司、DeepMind 的嘉宾。是的,这会是一系列非常有趣的对话。我们尽量让讨论组非常技术性,引出真正有实质内容的讨论,希望能有一些有趣的分歧和辩论。是的,有票可买。希望人们能来参加。

Yeah. So we're very excited about this. This is in late September, September 30th, we're doing a summit primarily focused on physical AI and real-time video generation. We have panelists from Nvidia, Physical Intelligence, the bot company, DeepMind. Yeah, it's going to be a very interesting series of conversations. We try to make the panels really technical and elicit actual substantive discussion and hopefully some interesting disagreements and debates on things. And yeah, there's tickets available. Hope people can join.

机器人学与世界模型辩论 Debates in Robotics and World Models

Host

既然你提到了,人们应该思考或你预期会有哪些分歧和辩论?看起来现在机器人领域有一个辩论是 VAS 与世界行动模型。好的。

Since you mentioned it, what kind of disagreements and debates should people think about or do you expect? So it seems like there is one debate right now in the robotics world is VAS versus world action models. Okay.

Anastasis

嗯,所以有些实验室真的在押注这两个方向之一。还有比如训练机器人模型的最佳数据来源是什么。

Um so there are labs that are really betting on one of those two directions. There is like what is the best source of data to train robotics models.

Host

就是我们谈过的第三方、第一方。

There's just the third party first party that we talked about.

Anastasis

是的。有些人真的相信进一步扩展遥操作数据,而另一些人则倾向于利用更大规模的视频数据。所以这些是一些,然后还有世界模型的辩论:直接预测像素,还是像 JEPA 这样的方法,还是更基于 3D 的方法。所以我认为我们在世界模型上正处于一个好时机,因为仍然有那种活跃的辩论,或者说什么是最好的长期方向。我非常强烈地认为,这种视频直接预测像素并扩展视频生成模型是正确的方向,但我认为研究人员之间有很多有趣的辩论,关于什么是最好的路径。

Yeah. There is the people who really believe in further scaling teleop data versus leveraging more large scale video data. So those are kind of some of the and then there is the world models debates of predict pixels directly versus something like JEPA versus a more 3D based approach. So I think we're at a nice time in world models because there is still that kind of active debate happening or like what is the best long-term direction. I feel very strongly that this video predict pixels directly and scaling video generation models is the right approach but I think there is a lot of interesting debate happening by researchers on like what is the best path to take.

物理侧未解决 Physical Side Not Solved

Host

有趣的是,这一切都集中在所谓的策略层和数据模型层,而物理侧完全解决了,就像所有传感器、所有执行器,所有这些我们都有了所需的一切。

It's interesting that it's all on like sort of let's call it the policy layer and the data model layer is the physical side is completely solved like all the sensors all the actuators all these things they're we have everything that we need

Anastasis

我不认为那也解决了,绝对没有。

I don't think that's solved either definitely um

Host

不同的问题。

different problem

Anastasis

你知道,就像我想梦想所有这些事情,然后我买一个机器人,或者我尝试自己组装,我甚至不能让马达工作。对。对。再次,你面对的是非常敏感的设备,有电压、功率、热量等等,你知道,我们坐在这里抽象地谈论软件和模型,但真的你也必须处理那些事情。

you know like it's like I want to dream about all these things and then I get you know I buy a robot or I buy I try to assemble my own and I can't even get the motors to like work Right. Right. Again, it's you're dealing with very sensitive equipment that has voltage and power and like heat and all these things which you know abstracted the way we're sitting here we're talking about software and talking about models but like really you have to deal with those kinds of things too.

融合其他模态 Incorporating Other Modalities

Anastasis

是的。我认为我通常也不反对将其他模态纳入我们的模型,就像我们看到的。

Yeah. And I think I'm generally also not opposed to incorporating other modalities into our models like we've seen.

Host

是的。

Yes.

Anastasis

最简单的例子是它们可以同时生成视频和音频。所以它们可以生成 RGB,也可以生成声音和音频。但我的,我写过这个,就像世界模型的极致版本是什么样子,就是你纳入越来越多来自宇宙的模态,并且你在不同尺度的观察上训练模型。所以是的。

The simplest case is they can generate video and audio at the same time. So they can generate RGB and they can also generate sound and audio. But my I've written about this as like what does the maximalist version of a world model look like is you're incorporating more and more modalities from the universe and you're training a model on different scales of observations as well. So yeah,

Host

你有一篇好文章,人们应该读读,关于真实。

you got a good essay that people should read on real

Anastasis

世界。Meta 发布了一个模型,像是六种模态合一,对吧?我忘了那东西的名字,但它是,是的。好的。深度是其中之一,但那像是 RGB 的变换,还有。

world. Meta released a model that was like six modalities in one, right? I forget what the name of the thing was, but it was like Yeah. Okay. Depth is one of them, but that is like a transformation of RGB and

Host

ImageBind。

image bind.

Anastasis

ImageBind。是的。还有什么其他模态?他们有热成像,

Image bind. Yes. What other modalities? They had heat,

Host

音频、深度、热成像、文本,呃。

audio depth, heat text, uh

Anastasis

不管 IMU 是什么。嗯,我确实认为你不妨做紫外线。你不妨做任何你觉得合适的其他模态,因为对模型来说都是数据。

whatever IMU is. Um I do think like you might as well do ultraviolet. you might as well do like whatever other modality you feel like because it's all data to the model.

Host

是的。一个大的赌注也是所有这些模态之间存在迁移。所以我最喜欢的例子之一,现在相当老了,是有一个稳定扩散的微调叫做 Riffusion,是音乐。

Yeah. And a big bet is also that there is transfer between all those modalities. So one of my favorite examples which is quite old at this point is there was this finetune of stable diffusion that was called refusion which was the music.

Anastasis

是的。

Yeah.

Host

只是在频谱图上微调稳定扩散,就变成了一个相当有能力的音乐生成器。对。可能有一些空间模式或时空模式,如果我们谈论视频,在不同尺度和不同模态中出现。所以模型已经做了一定程度的元学习,让你从图像模型开始训练它预测音频,比从头训练音频学得更快。嗯,还有一些其他有趣的例子。所以有一个项目叫做 The Well,它是一个物理数值模拟的数据集,在物理、生物和其他一堆领域。所以它本质上是不同物理系统,跨越非常不同的空间和时间尺度,从天体物理到低层次的原子相互作用。我们已经看到,我们做了一些工作,我们看到我们可以拿我们的视频模型,真实世界视频看起来完全不像这个,你实际上可以在那些数值模拟上微调它,把它们当作 RGB 帧,你实际上得到合理的性能,比从头训练快得多。

Just fine-tuning stable diffusion on spectrograms and became a quite capable music generator. Right. There is probably some spatial patterns or like spatial temporal patterns if we're talking about video that kind of emerge the different scales and different modalities. And so there is some degree of meta learning that the model has done that allows you to learn faster if you start from a just a model on images and train it to predict audio than if you train from scratch on just audio. Um and there is some other interesting examples. So there is this project called the well it's a data set of physics numerical simulations in physics and biology and a bunch of other domains. So it's essentially different physical systems across very different scales of space and time from like astrophysics to low-level kind of atomistic interactions. And we've seen we've done some work on this and we've seen that we can take our video model where real world video looks nothing like this and you can actually fine-tune it on those numerical simulations and just treat them as RGB frames and you actually get reasonable performance much quicker than if you just train from scratch.

跨模态迁移 Transfer Across Modalities

Anastasis

是的,我认为我们在语言中也看到了这一点。

Yeah, I think we've seen this across languages.

Host

是的,DeepSeek OCR 也是。是的,DeepSeek OCR,就像你不必对文本进行分词,你可以直接把它们作为图像扔进去。在基础预训练中发生了很多事情,就像很久以前有人争论说,哦,人类有那么多感官表征,对吧,嗅觉、触觉,模型还有整整两种模态,我们甚至没有数据,然后就像好吧,你拿一个空气质量传感器,你可以试试这些东西,但实际上你知道,仅仅在基础训练运行中就有很多事情发生,你不会从这些小东西中得到那么多。是的,正是如此。我认为它解决的是数据稀缺。所以你没有那么多,你有那么多视频数据可用,但你没有,你知道,像所有工厂数据。

Yeah, Deep SQL CR as well. Yeah, Deepse OCR like you don't have to tokenize text like you can just throw them in as images. There's a lot that happens in that base pre-training like there was an argument a long time ago of people saying oh humans have so many sensory representations right smell touch models have a whole two more modalities that will never like that we don't even have data for and it's like okay you take AQI sensor like you can you can try this stuff but actually you you know there's so much happening in just the base train run that you don't get as much from these little things. Yeah, exactly. And I think that's what it solves is data scarcity. So you don't have as much, you have so much video data available, but you don't have, you know, like all factory data.

Anastasis

酷的是它也可以反过来,对吧?所以如果你想做物理,比如你想测量这个,或者你想让扩散模型做音频,它迁移得非常好。所以就像在你的情况下,一点点的机器人后训练让视频模型在另一个领域使用它的基础。所以我们也可以把它应用到其他东西上。

The cool thing is it goes the other way too, right? So if you want to do physics, like if you want to measure this or you want to have a diffusion model do audio, it transfers really well. So like in your case, the little bit of post-training for robotics gets a video model to use its fundamentals in another domain. So we can apply that to other stuff too.

Host

是的。

Yeah.

科学多模态模型 Multimodal Models for Science

Anastasis

如果我们看看如何让这些模型在科学领域更有用——如果你看 AlphaFold,由于它需要训练的数据量有限,它基本上是一个非常微调的架构,只是为了解决蛋白质结构预测。但如果你把所有这些不同的科学数据源汇集到一个单一模型下,我认为这是一种可以帮助我们通过利用从一个模态或一组数据到另一个模态或另一组数据的所有学习,来解决科学中各种新问题的方法。所以这个方向还处于非常早期的阶段,但我确实认为这最终就是模拟世界的终局。你不只是使用 RGB,你把 RGB 作为起点,但你可以纳入宇宙中越来越多的模态,并利用从一个模态学习到另一个模态的迁移。

And if we look at how you make those models more useful in scientific domains — if you look at AlphaFold, it had all these very, because of the limited amount of data that it needed to be trained on, it's basically a very fine-tuned architecture just to solve protein structure prediction. But if you take all those disparate sources of scientific data and you bring them together under a single model, I think that's an approach that can help us solve new kinds of problems across science by leveraging all the learnings from one modality or one set of data to another. So very early days for that direction, but I do think that's ultimately where the endgame of simulating the world is. You're not just using RGB, you're using RGB as a starting point, but you can incorporate more and more modalities of the universe and leverage the transfer that happens from learning from one to the other.

全能模型的缺陷 Drawbacks of Omni Models

Host

我想追问的是:全能模型的缺点是什么?为什么不是所有东西都是全能模型?为什么不是现在?为什么你会从语言主干或图像视频主干开始,然后再从那里走向全能?这有关系吗?

I guess the followup there is: what's the drawback of omni? Like, why is everything not an omni model? So why not now? And why would you start from a language backbone or image-video backbone and then go omni from there? Does it matter?

Anastasis

是的,我们需要一步一步来。我们必须先解决机器人技术,然后才能现在去解决一切。

Yeah, we need to take it one step. We have to solve robotics first and then we can go into solve everything now.

Host

是的。

Yeah.

Anastasis

我的意思是,我确实认为那些全能模型需要进行大量开放性的研究。当你把多个模态带入一个单一模型进行预测时,有很多事情需要仔细考虑,但我认为这些应该都是可以解决的。

I mean, I do think there is a lot of open-ended research that needs to happen for those omni models. There is a lot of things that require careful consideration when you're bringing multiple modalities into a single model to predict, but I think I expect those to be solvable.

闭幕致辞与AI电影节 Closing Remarks and AI Film Festival

Host

太棒了。你非常慷慨地付出了时间。祝贺你所有的成功,是的,我很期待 AI 峰会或物理 AI 峰会。

Wonderful. You've been very generous with your time. Congrats on all your success and yeah, I'm excited for the AI summit or physical AI summit.

Anastasis

是的,谢谢邀请我。

Yeah, thanks for having me.

Host

是的,如果电影节在城里,人们应该去看看,对吧?你将会到处巡演。

And yeah, people should check out the film festival if it's in town, right? You'll be going to be touring all over the place.

Anastasis

是的,明年我们可能会做——所以我们每年五月或六月举办电影节,上一次我们在纽约、洛杉矶、东京和 AI 工程师博览会上举办过。

Yeah, next year we're probably going to do the — so we do film festivals every May or June, and we did the last one in New York, LA, Tokyo, and at the AI engineer fair.

Host

是的。是的。是的。嗯,所以是的,希望明年能在更多地方举办。

Yeah. Yeah. Yeah. Um, so yeah, hopefully even more places next year.

Anastasis

不,我认为就像有一天,你知道,你将主持 AI 视频的奥斯卡奖,你知道,我认为人们应该非常认真地对待这个,就像他们可以拥有的潜在职业一样。

No, I think like as someday, you know, you will be hosting the Oscars of AI video and you know, I think people should like take this very seriously as like a potential career they can have.

Host

AI 视频的奥斯卡奖将被称为奥斯卡奖。

The Oscars of AI video will be called the Oscars.

Anastasis

好的。好的。谢谢。

All right. All right. Thank you.

Host

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →