Genie 3:世界模型的百倍飞跃

Genie 3: A 100x Leap in World Models

杰克·帕克-霍尔德 Jack Parker-Holder · TWIML AI 播客 · 2025-08-19 · 约 61 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Google DeepMind 研究人员讨论 Genie 3 模型,该模型在生成质量、分辨率、交互时长和速度方面实现了百倍提升。

Google DeepMind researchers discuss the Genie 3 model, which achieves a 100x improvement across generation quality, resolution, interaction duration, and speed.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

开场与嘉宾介绍 Introduction and Guest Backgrounds

Host

好的,各位。欢迎收听另一期 Twiml AI 播客。我是主持人 Sam Charrington。今天,我邀请到了 Google DeepMind 的研究员 Shlomi Fruchter 和 Jack Parker Holder,来讨论最近发布的 Genie 3 模型。这是一个令人印象深刻的世界模型,大约一年前我们曾在播客中与 Ashley Edwards 首次介绍过它。我非常期待深入探讨 Genie 3。Jack 和 Shlomi,欢迎来到播客。

All right, everyone. Welcome to another episode of the Twiml AI podcast. I am your host Sam Charrington. Today, I'm joined by Shlomi Fruchter and Jack Parker Holder, researchers at Google DeepMind to discuss the recent release of the Genie 3 model, which is an impressive world model that we first introduced to you here on the podcast in our conversation with Ashley Edwards almost exactly a year ago. I'm super excited to dig into Genie 3. Uh Jack and Shlomi, welcome to the podcast.

Jack Parker-Holder

谢谢。感谢邀请我们。很兴奋来到这里。

Thank you. Thanks for having us. Excited to be here.

Host

这是一次值得深入探讨的有趣访谈,因为我们不久前才介绍过 Genie,但我不想假设大家都听过那期访谈,甚至完全不了解 Genie。所以,我们会从基础开始,深入探讨这个项目、它的起源以及它令人兴奋的原因。在深入之前,我希望你们各自做个自我介绍,分享你们进入机器学习研究的历程亮点以及你们最感兴趣的研究方向。Jack,你先开始吧?

This is an interesting interview to try to dig into because we covered Genie relatively recently, but I don't want to assume that people have listened to that interview or even know about Genie at all. So, we're going to start a little bit from the beginning and dig into the project and where it comes from and why it's exciting. I'd love to have each of you just introduce yourselves before we dig in though and share the highlights of your path to ML research and what you're most excited about researching. Jack, why don't you get us started?

Jack Parker-Holder

太好了,谢谢。所以,不太夸张地说,我的经历是大约 10 年前我在金融领域工作,业余时间晚上攻读兼职硕士,实际上在 2017 年左右经常听你的播客。我进入了机器学习研究领域,研究用于强化学习的进化方法。我还与 Google Brain 的一些人合作过,因为我在纽约的办公室。然后我决定攻读博士学位,专注于开放式学习,仍然是强化学习,后来逐渐涉足世界模型。到我博士毕业时,我越来越确信这些想法的结合将是一件非常强大的大事。所以,博士毕业后,我加入了 Google DeepMind,参与了一个名为 adaptive agents 的项目,在 XLand 环境中工作。之后,基本上开始了 Genie 项目。我已经做了几年了。我在开放式团队,该团队也涵盖其他领域,但我们一直专注于将世界模型作为通往开放式学习的路径这一想法,作为 Genie 项目的一部分。

Awesome. Thanks. So, without being too much of a suck up, my path was that I was working in finance about 10 years ago and I did a master's part-time in the evenings after work and actually was listening to your podcast quite a lot around 2017. I got into ML research doing evolutionary methods for reinforcement learning. I worked with some folks from Google Brain as well because I was in New York in those offices. Then I decided to do a PhD where I focused on open-ended learning, reinforcement learning still, and then got a bit into world models. By the time I finished my PhD, I was increasingly convinced that the combination of these ideas would be the really powerful big thing to do. So, after my PhD, I joined Google DeepMind and worked a little bit on a project called adaptive agents in the XLand environment. After that, pretty much started Genie. I've been doing that for a few years. I'm in the open-endedness team, which encompasses some other areas too, but we've been focusing on this idea of using world models as a path to open-endedness as part of the Genie project.

Host

太棒了。你呢,Shlomi?

Awesome. How about you, Shlomi?

Shlomi Fruchter

所以,我的第一次编程经历是在青少年时期开发游戏引擎。3D 引擎,更多是来自尝试模拟效果的世界,比如光照效果、液体效果等。所以,我在这个领域工作了一段时间。我非常喜欢这个视觉领域。然后我加入了 Google,一直在 Google Duplex 团队。我们在 Duplex 项目中所做的与视觉内容非常不同,因为它主要是通过电话完成任务。Duplex 项目,如果人们还记得 Google IO 2018,当时人们以为‘好吧,哇,我们有了 AGI。’那是‘这是 AGI 吗?’的时刻之一。我不认为它是,但它绝对是下一步。具体来说,这个项目是 Google 要代表你打电话给餐厅和理发店预约。当我们启动 Duplex 时,目标是能否构建一个机器人,通过电话与人交谈,而他们不觉得这是机器?可以说是通过电话的图灵测试。我们发现,尽管直到最近随着 LLM 的出现才能实现完全通用的电话对话,但在当时,RNN、LSTM 的时代,肯定不是 Transformer,我们能够开发出至少针对这个特定任务完成很多事情的东西。那是我第一次接触机器学习,非常有趣,因为它既注重研究,又有实际部署。我们最终将其扩展到在 15 个国家拨打了数亿次电话。知道的人较少,因为它主要是打电话给企业更新 Google Maps。我也关注了从 GPT 到 Google 内部的 Mina 和 Lambda 等早期 LLM 的发展。我们将这项技术集成到了 Google Duplex 中,但到了某个时候,我觉得我的视觉根源又回来了。图像扩散模型的革命非常吸引我。我觉得这是一个巨大的机会,达到了某种成熟点。从那时起,我一直在研究视频模型。其中一个体现是游戏引擎,这是一个有点副业的项目。我和几个朋友,包括 Duplex 的创始人 Yevgeny Levidow,提出了一个问题:是否有可能完全由神经网络实时模拟一个现有游戏?我也一直在研究 Genie 2 和 3。在游戏引擎时期,我开始与 Jack 交谈。我对 Genie 世界模型的工作印象深刻。

So, I started my first programming experience in my teen years developing game engines, actually. 3D engines and more from this world of trying to simulate effects like lighting effects, liquid effects, etc. So, I've been working in this space for some time. I really like this visual domain. Then I joined Google and I've been on the Google Duplex team. What we did in the Duplex project was very different from visual stuff because it was mostly getting stuff done over the phone. The Duplex project, if people remember from Google IO 2018, people thought, 'Okay, wow, we have AGI.' It was one of the 'is it AGI?' moments. I don't think it was, but it was definitely the next step. Specifically, this was the project where Google was going to call restaurants and hairdressers to make appointments on your behalf. When we started Duplex, the goal was a question of can we build a bot that talks over the phone with people without them feeling this is actually a machine? Hitting this kind of over-the-phone Turing test, if you will. We found that although having a completely general conversation over the phone was not achievable until very recently with LLMs, already then, in the era of RNNs, LSTMs, and definitely not transformers, we were able to develop something that at least for this particular task accomplished a lot. That was my first touch with machine learning, and it was very interesting because it was research-oriented but also had real-world deployment. We ended up scaling that to call hundreds of millions of calls over 15 countries. It's less people aware of it because it was mostly calling businesses to update Google Maps. I also followed things that happened with LLMs from GPT and internally at Google with Mina and Lambda and other models very early on. We integrated this technology into Google Duplex, but at some point, I felt my visual roots came into play again. The revolution in image diffusion models was really appealing to me. I felt this is a huge opportunity, hitting some point of maturity. Since then, I've been working on video models. One incarnation of that was a game engine, which was a bit of a side project. A few friends and I, including Yevgeny Levidow, the founder of Duplex, asked the question: is it possible to simulate an existing game in real time completely by a neural network? I've also been working on Genie 2 and 3. Around the time of the game engine, I started talking to Jack. I was very impressed with the Genie world line of work.

定义世界模型 Defining World Models

Host

人们在关注 Genie 时最兴奋的一点就是世界模型这个概念。Jack,或许你可以深入谈谈这对你意味着什么,以及你如何看待世界模型融入更广泛的 AI 模型(尤其是基于 Transformer 的模型)的发展轨迹。对你来说,世界模型这个想法包含了什么?

One of the things that people are most excited about in looking at Genie is this concept of world model. Maybe Jack, you can dig into what that means for you and how you see a world model fitting into the broader trajectory of AI models, transformer-based models, like how you think about that. What all is captured in this idea of a world model for you?

Jack Parker-Holder

当然。我认为在 Genie 3 中,我们确实试图在所有维度上将其推向极限。我们看到,我们的模型在生成质量方面能力更强。如果你看分辨率、交互持续时间、下一帧生成的速度,如果你将所有维度相乘,你会得到一个相当显著的 100 倍改进。

Sure. I think in Genie 3, we really tried to push it to the limit across all of the dimensions. We see that we have models that are more capable in terms of the quality of their generation. If you look at the resolution, the duration of the interaction, how fast the next frame can be generated, if you multiply all of those dimensions, you get a quite significant 100x improvement.

定义世界模型 Defining World Models

Jack Parker-Holder

是的,这是个很好的问题,从某个层面来说很容易回答,但也可以写满一整本书的不同哲学观点。对我自己而言,直到大约一年前,世界模型本质上是强化学习范式中建模 MDP 的模型。它接收状态和动作,预测下一个状态。这个想法已经存在一段时间了。世界模型就是基于模型的强化学习中的模型,对环境进行建模。在 90 年代初,Juergen Schmidhuber 发表了一篇关于循环世界模型的论文,Rich Sutton 的 Dyna 论文也差不多同时出现。那是基于模型的强化学习这一研究方向的起点。对我来说,关键工作是 2018 年 Ha 和 Schmidhuber 的论文,那时我正好对这个领域产生兴趣。他们研究了像 Mujoco 半猎豹这样的强化学习任务。世界模型论文表明,通过少量离线示例,我们可以模拟环境,足够好地预测下一个状态,从而在该模型中训练策略,再将其迁移回真实环境,而策略从未在真实环境中训练过。这非常酷。但设置要求有来自真实环境的数据,所以理论上你也可以直接在真实环境中训练并得到相同结果,但你需要收集那些数据。这就是过去几年的范式:证明我们能在越来越复杂的环境(我们拥有该环境)中做到这一点。如果我们不想用世界模型,可以使用分布式世界算法并获得良好性能。世界模型有样本效率等优势,但并没有解决原本无法解决的问题。然后,回到 Shlomi 关于文本到图像模型的评论,我想我们从不同角度都感到兴奋。对我来说,这等于说:'如果我们能把文本到图像做得这么好'——我记得大约是 4 年前 Imagen 出现的时候——那么视频迟早会出现,视频之后,也许我们就能得到世界模型。我们将能够用大数据集模拟任何东西。所以想法是,将同样的概念应用到我眼中的世界模型上:模拟环境,从而模拟任何环境。这就是我们提出基础世界模型概念的原因:一个世界模型,但不是针对单一环境,而是任何可能的环境,就像基础模型一样。我们坚持这个定义:基础世界模型根据动作预测下一个状态。但最近,越来越多的其他模型也被视为世界模型,比如视频模型和文本到视频模型。起初,这不符合我对世界模型的看法,但现在我认为它在不同抽象层次上是成立的。所以现在我有了更宽泛的观点:世界模型根据过去和某种形式的动作模拟未来,模拟世界的动态。它不是显式地建模每个 MDP 转移,而是模拟动态,让你可以行动和干预,获得反事实信息,从而实现规划或在模拟中学习策略。Shlomi,你也有同样的转变吗?

Yeah, that's a great question, which on one level is quite easy to answer, but could also fill a whole book of different philosophies. For myself, until about a year ago, a world model was essentially a model from the reinforcement learning paradigm that models an MDP. It takes a state and an action and predicts the next state. This idea has been around for a while. A world model is the model in model-based reinforcement learning, modeling the environment. In the early '90s, Juergen Schmidhuber had a paper on recurrent world models, and Rich Sutton's Dyna paper came out around that time. That was the starting point for model-based reinforcement learning. For me, the key work was Ha and Schmidhuber's 2018 paper, around the time I was getting interested in this field. They looked at reinforcement learning tasks like Mujoco half-cheetah. The world models paper showed that with a few offline examples, we could simulate an environment, predict the next state well enough to train policies in that model and transfer them back to the real environment, without the policy ever training in the real environment. That's super cool. But the setting required data from the real environment, so you could also just train in the real environment and get the same result in theory, but you needed to collect that data. That was the paradigm for a few years: showing we could do this for increasingly complex environments where we had the environment. If we didn't want to use world models, we could use a distributed world algorithm and get good performance. World models had benefits like sample efficiency, but they weren't solving something unsolvable otherwise. Then, going back to Shlomi's comment about text-to-image models, I think we were both excited from different angles. For me, it was saying, 'If we can do text-to-image this well'—I think it was about 4 years ago with Imagen—then video will happen, and after video, maybe we'll get to world models. We'll be able to simulate anything with large datasets. So the idea was to apply the same concept to what I view as a world model: simulating an environment, and thus simulate any environment. That's why we came up with the idea of a foundation world model: a world model for any possible environment, like a foundation model. We stuck to that definition: a foundation world model predicts the next state given actions. But increasingly, other models have been considered world models, like video models and text-to-video. At first, that didn't fit my view, but now I think it does at a different level of abstraction. So now I have a slightly broader view: world models simulate the future given the past and some form of actions, simulating dynamics of a world. It's not explicitly modeling every MDP transition, but simulating dynamics so you can act and intervene and get counterfactual information, enabling planning or learning policies in simulation. Shlomi, did you have the same shift?

Shlomi

是的,我更多是从模拟的角度出发,或者更关注视觉方面。我认为我们今天所说的世界模型非常视觉特定,这也是当前世界模型的局限之一。但回到 Jack 提到的世界模型论文,定义很清晰,是一个很好的锚点,因为这个词有很多用法。但也有一些直觉:当我开始看到视频模型,还有图像模型,当你写文本,提供文本提示,得到的东西感觉像素背后有一个世界。这就是直觉部分。要让东西看起来像真实的图像或视频,模型可能必须对正在发生的事情、世界如何运作、甚至物理规律有一定内部表示。这就是外行人对世界模型的直觉:模型可能对世界有一些理解,而且是视觉上的。这很关键。但世界模型不一定是视觉的,不一定生成像素。这是更宽泛的观点。在一些文献中,世界模型可以在某种潜在空间中。它应该能用于做决策、预测接下来会发生什么、在强化学习意义上进行规划,并学习如何在环境中更优地运作。

Yeah, I come more from the simulation side, or I think more about the visual aspect. I think what we call world models today are very visual-specific, which is also one of the limitations of current world models. But going back to the world models paper Jack mentioned, the definition was pretty clear, and it's a good definition to anchor to, because the term is used in many ways. But there's also an intuition: when I started seeing video models, but also image models, when you write text, provide a text prompt, and you get something that feels like there is a world behind those pixels. That's the intuitive part. For things to really look like a realistic image or video, the model probably has to have some internal representation of what's going on, how the world behaves, physics to some extent. That's the layman's intuition behind the world model: the model probably has some understanding of the world, and it's visual. That's key. But the world model doesn't have to be visual; it doesn't have to generate pixels. That's the broader view. In some literature, the world model can be in some latent space. It should be something we can use to make decisions, predict what's going to happen next, perform planning in the RL sense, and learn how to operate more optimally in an environment.

视觉域与扩散模型 Visual domain and diffusion models

Host

但回到视觉领域,我认为情况是:扩散模型恰好对图像、视频和音频都非常有效,所以视觉领域进展顺利。然后这些领域之间的交叉变得非常明显。我认为我们现在就处于这个阶段——模型能够生成非常逼真的环境,而我们正试图推动这一进展。

But again, going back to the visual domain, what I think happened is that the visual domain really worked well because diffusion models just happened to work really well for images, videos, and audio as well. And then the intersection between those fields became very obvious. I think that's where we are right now, with models capable of generating very realistic environments, and we're trying to push along this thing.

Genie 项目轨迹 Genie project trajectory

Host

Jack,你也参与了 Genie 1 和 2 的工作。我想请你谈谈这些项目的演变轨迹。作为背景,Genie 3 让我惊讶的是,我原以为它是实时可玩的,但实际上并非如此——每帧之间需要大约 20 分钟。而现在,仅仅一年后,它已经实现了世界时间实时可玩。请谈谈这些迭代之间在功能和研究上的里程碑。

Jack, you worked on Genie 1 and 2 as well. I'd love to have you talk a little bit about the trajectory of the projects. For context, what amazed me about Genie 3 was that I assumed it was real-time playable, but it wasn't—it took like 20 minutes between frames. Now, in just a year, it's world-time real-time playable. Talk about the functional and research milestones between these iterations.

Jack Parker-Holder

当然。我认为我们在两篇论文中都确实试图强调它不是实时的。挑战在于,当你宣布一项新研究时,你不会说‘欢迎来到我们的非实时突破’,因为那听起来不令人兴奋。你想让它听起来激动人心,而不是一开始就提限制。但如果你把限制放在中间,配上很酷的视频,人们根本看不到那里。所以我们确实面临这个挑战。实际上不到 20 分钟——更像是 30 秒。但想想文本到图像模型,它们需要在大约 1/20 秒内运行才能达到 Genie 3 的速度。所以我们取得的成就相当了不起。回顾一下,我强烈推荐你的听众也听听 Ashley Edwards 的那一期。Genie 1 与 Genie 3 截然不同。它是基础世界模型的第一个概念验证——一个能够生成新世界的模型。此前在单一领域的世界模型方面已有惊人进展,比如 Atari 套件和 DreamerV2/V3,它们能建模越来越复杂的单个环境,并让智能体学习到解决复杂任务的惊人行为。那些工作聚焦于单一领域的复杂性。而 Genie 1 问的是:我们能否训练一个模型,无论复杂性如何,至少能生成新的环境和世界?挑战在于数据——我们没有目标环境的标注数据。所以我们收集了一个无标签视频数据集,没有动作标签,并采用了一种巧妙的方法:学习潜在动作。幸运的是,我遇到了一个技能完美匹配的人:Ashley。她在潜在动作学习方面有超过 5 年的经验,是该方向的先驱之一,但用于不同场景——从视频中学习行为。我们的设置正好相反:我们想从视频中学习世界模型,然后用它来学习策略。CVPR 上有一篇论文叫《可玩环境》,由 Menapace 等人撰写,做了类似的事情。我在写博士论文时看到了它,觉得这是一条路。然后我和 Ashley 聊了聊,她同意我们可以将这种方法规模化。所以对于 Genie 1,这是一个新想法:我们到底能不能做到?但它的范围有限。我们训练了两个模型:一个在 2D 平台游戏数据集上,另一个在 RT1 论文的机器人数据集上。在这两种情况下,我们都能无监督地学习一个动作空间,这样你可以给模型一张新图像(例如来自文本到图像模型或你自己拍的照片),然后通过潜在动作像玩一个世界一样玩它。这听起来很神奇,但有一些限制:分辨率是 90p,只能持续几秒钟就会退化,而且当你移出屏幕时不会生成新内容。我们使用了一种掩码方法,当你移动时会出现模式崩溃。例如,在平台游戏中,向右移动只会继续一个平坦的平台,而不是生成新的有趣内容。人们确实注意到了这一点,但我们对此心知肚明。这更像是‘它居然能工作,这很酷’。

Sure. I think we did try to emphasize it wasn't real-time in both papers. The challenge is that when you announce new research, you don't say, 'Welcome to our non-real-time breakthrough,' because that's not exciting. You want to make it sound exciting, not lead with a limitation. But if you put the limitation halfway down with cool videos, people don't get that far. So we had that challenge. It was slightly less than 20 minutes—more like 30 seconds. But if you think about text-to-image models, they need to operate at about a 20th of a second to match Genie 3's speed. So it's quite remarkable what we achieved. Going back, I recommend your listeners listen to the Ashley Edwards episode as well. Genie 1 was quite different from Genie 3. It was the first proof of concept of a foundation world model—a model that could generate new worlds. There had been amazing progress in world models for single domains, like Atari and DreamerV2/V3, which modeled increasingly complex environments and enabled agents to learn amazing behaviors. That work focused on single-domain complexity. Genie 1 asked: can we train a model that can generate new environments and worlds at all, regardless of complexity? The challenge was data—we didn't have labeled data from target environments. So we collected a dataset of unlabeled videos with no action labels and used a neat approach: learning latent actions. The nice thing was that I bumped into someone with a perfect skill set: Ashley. She had worked on latent action learning for over 5 years, pioneering that direction in a different context—learning behaviors from videos. We had the opposite setup: we wanted to learn a world model from videos to then learn policies. There was a CVPR paper called 'Playable Environments' by Menapace et al. that did something similar. I saw it while writing my PhD thesis and thought it was a path. Then I chatted with Ashley, and she agreed we could scale that approach. So for Genie 1, it was a new idea: could we do it at all? But it was limited in scope. We trained two models: one on a 2D platformer game dataset, and another on a robotics dataset from the RT1 paper. In both cases, we could unsupervised learn an action space, so you could give the model a new image (e.g., from a text-to-image model or a photo) and play it like a world using latent actions. That sounds amazing, but there were caveats: it was 90p resolution, lasted maybe a couple of seconds before degrading, and didn't generate anything new when you went off-screen. We used a masked get approach, and it had mode collapse when you moved away. For example, in the platformer, moving right just continued a flat platform instead of generating new content. People noticed that, but we were aware of it. It was more like 'this is cool that it remotely works.'

Genie 1 与 2:早期阶段与局限 Genie 1 and 2: Early stages and limitations

Host

这是一个非常早期的项目,对吧?基本上我们刚开始的时候,没人觉得这事值得做。幸运的是,我们在谷歌 DeepMind 这样的地方,他们鼓励这些探索性的项目。但这并不是一个资源密集型的项目,就是我们几个人凑在一起搞的。那就是 Genie 1。我们写了一篇论文,投给了 ICML,去年夏天发表的。所以感觉上很新,因为工作完成和论文发表之间有个周期。至于 Genie 2,背景是这样的:Genie 1 的论文大概在 2024 年 2 月或 3 月发表。我们早就有了结果,已经开始考虑未来的计划,并且实际上动手做了。但那时,视频模型整体上取得了很大进展。它们迎来了自己的时刻,就像几年前文本到图像模型那样。我们很清楚,对于 Genie 的下一阶段,我们可以尝试更大规模的东西,而且很可能行得通,因为我们看到视频模型已经有效地扩展了。所以我们决定从 2D 游戏转向任何 3D 世界的数据集。我们把它扩展到 360p 分辨率。它能够根据图像提示生成环境,持续大概 10 到 20 秒然后退化。所以它不是实时的,但你可以玩上几分钟,每帧等几秒钟。它对 3D 世界是有效的。当时也不完全清楚它能否成功,因为它是自回归生成。甚至不确定它能否持续 10 到 20 秒。所以这是一个相当令人兴奋的结果。

And it was a very early stage project, right? Essentially when we started this, no one really thought it was worth doing. And luckily we're in a place like Google DeepMind where they encourage these more exploratory projects. But it wasn't a heavily resourced effort; it was kind of a few of us scrapping together. So that was Genie 1. We wrote a paper on it and submitted to ICML, which came out last summer. That's why it feels very recent because there's a cycle between finishing the work and the paper being published. For Genie 2, the setting was basically: Genie 1 paper came out around February or March 2024. We had the results for a while and had already started thinking about future plans, and actually got to work on it. But around that time, there was a lot of progress on video models in general. They had their moment, like text-to-image had a few years before. It became very clear that for the next phase of Genie, we could go for something a bit larger scale, and it would probably work, because we'd seen that video models had scaled pretty effectively. So we decided to go from 2D games to any 3D world dataset. We scaled it up to 360p. It had the ability to generate environments again from image prompts that lasted maybe 10 to 20 seconds before they degraded. So it wasn't real-time, but you could play with it for a couple of minutes and wait a few seconds for each frame. It kind of worked for 3D worlds. That was also not completely obviously going to work at the time, because it was doing auto-regressive generation. It still wasn't clear that this would last even 10 to 20 seconds. So it was quite an exciting result.

Host

Genie 2 是否克服了模式崩溃的问题?

Did Genie 2 overcome the mode collapse issue?

Jack Parker-Holder

在一定程度上。Genie 2 是一个扩散模型,所以性质略有不同,但它仍然只支持图像提示。它的表现力远不及视频模型。它仍然需要你选择一张能用的图像,这并不总是奏效。它不能自己生成世界;你需要生成一张特定格式的图像,比如其中有一个清晰的智能体在正确的位置,然后它从那里开始模拟。而我希望我们在 Genie 3 上花很多时间讨论的是,它绝对能做到这一点。总结一下这部分,Genie 2 当时是这种通用方法的一个良好迹象,但它不是实时的。我们从游戏引擎中看到,如果做对了,那会非常有影响力。视觉质量不错,但它仍然没有使用文本作为输入,表现力不够,而且远不及同时期最先进的视频模型如 Veo 2 的视觉质量,Veo 2 让我们所有人都感到震惊。所以这就是 Kailey 的 Genie 3 和 Shlomi 发挥专长的地方。我喜欢借用别人的专长,那时我们显然有了这些领域的专家。

To an extent. Genie 2 was a diffusion model, so it had slightly different properties, but it still only supported image prompting. It wasn't anywhere near as expressive as video models. It still required you to select an image that worked, which didn't always work. It couldn't generate its own worlds; it required you to generate an image in a certain format where there's a clear agent in the right location, for example, and then it would simulate from there. Whereas I think what I'm hoping we spend a lot of time talking about with Genie 3 is that it can definitely do that. To bookend this part, Genie 2 at the time was a good sign of life for this general approach, but it wasn't real-time. We'd seen with game engine that that was really impactful if you get that right. The visual quality was good, but it still wasn't using text as input, wasn't as expressive, and it was nowhere near the visual quality of state-of-the-art video models like Veo 2, which came out at the same time and was absolutely mind-blowing for all of us. So that's where Kailey's Genie 3 and Shlomi came in and really had expertise. I like to borrow people's expertise, and at that point it was pretty clear we had an expert in those areas.

将其他方法融入 Genie 3 Incorporating other approaches into Genie 3

Host

谈谈你们是如何整合这些其他方法的,比如游戏引擎、Veo。它们是如何被纳入 Genie 研究路线的?我想想,因为我觉得 Jack 说得对:我们一直在讨论这个轨迹,Genie 1 和 2。我们还没有真正讨论 Genie 3,当你调出演示时,它令人印象深刻的地方。也许我们让 Shlomi 也来谈谈,然后讨论你参与的其他项目是如何影响整体研究的。

Talk a little bit about how you incorporated these other approaches, like game engine, Veo. How was that incorporated into the Genie line of research? Let me think about this because I think Jack you make a good point: we've been talking about this trajectory, Genie 1 and 2. We haven't really talked about Genie 3 and, when you pull up the demo, what's super impressive about it. Maybe we'll let you do that as well, Shlomi, and then talk about how these other projects you've worked on influenced the overall research.

Shlomi

当然。基本上,在 Genie 3 中,我们真的试图在所有维度上将其推向极限。如果你看到 Jack 刚才谈到的进展,我们有在生成质量上更强大的模型。Genie 1 和 Genie 2 能够生成越来越多更一致的世界,但仍然有限。所以我们想推动这个轨迹能保持一致的时间长度。这是我们要改进的一个维度。当然还有分辨率。如果你看分辨率、交互持续时间、下一帧生成的速度,如果你把这些维度相乘,在原始算力上你会得到相当显著的 100 倍提升。很明显,这不是一件显然能成功的事情,所以这个项目存在一些风险。但我们也觉得是时候去做了。基本上发生的事情是,在我们推出 Veo 2 之后,它反响很好。我们觉得质量确实提高了,但它绝对不是实时的,也不是交互式的。然后 Genie 2 出来了,它在不同的方向上推动了边界。我们只是说:“好吧,让我们尝试结合这些改进的向量,进入下一个层次。”这基本上就是 Genie 3 的目标:试图融合所有领域的精华,可能双关语。

Sure. Basically, in Genie 3, we really tried to push it to the limits across all dimensions. If you see the progression that Jack just talked about, we have models that are more capable in terms of the quality of their generation. Genie 1 and Genie 2 are able to generate more and more worlds that are more consistent, but still to a limit. So we wanted to push on how long this trajectory can remain consistent. That was one of the dimensions we wanted to improve. And of course, the resolution. If you look at the resolution, the duration of the interaction, how fast the next frame can be generated, if you multiply all of those dimensions, you get a quite significant 100x improvement in terms of raw compute. It was very clear that this is not something obvious that's going to work, so there was some risk to this project. But we also felt that this was the time to go for it. What happened basically is that after we launched Veo 2, it was very well received. We felt the quality really improved, but it was definitely not real-time, not interactive. Then Genie 2 came out and definitely pushed the envelope in this different direction. We just said, "Okay, let's try and combine those vectors of improvement and go to the next level." And that's pretty much what Genie 3 is about: trying to bring the best of all worlds, pun intended maybe.

Host

当你谈到融合这些领域的精华时,是融合想法?架构?数据集?还是以上所有的组合?

When you talk about bringing the best of these worlds together, is it bringing together ideas? Architectures? Datasets? Or some combination of all of the above?

Shlomi

首先,我认为是所有方面。这可能是个老生常谈,但肯定也关乎人。我们有来自不同团队的人,把他们的经验、动力和精力带进这个项目。所以这是一件大事。在技术方面,肯定有共同的技术挑战。基本上,我们生成像素作为输出。我们以文本作为输入,我们希望能够生成感觉一致的东西。

First, I think it's all. This may be a cliché, but it's definitely about the people as well. We had people from different teams bringing their experience, motivation, and energy into this project. So that was a big thing. In terms of the tech, there are definitely shared technical challenges. Basically, we generate the output as pixels. We take text as input, and we want to be able to generate something that feels consistent.

视频生成的一致性 Consistency in video generation

Jack Parker-Holder

所以即使是像 Veo 生成的 8 秒视频,你仍然希望它是一致的,对吧?如果摄像机移动,事物应该看起来一致。你想要感觉这就像是在真实世界中拍摄的。同样,一旦它是交互式的,我们希望根据用户的输入生成下一帧,但一致性非常重要。早期的模型可以生成下一帧,有点像游戏引擎,但它的上下文很短。它之所以能工作,是因为它基本上学会了那个《毁灭战士》游戏的特定属性。它记住了关卡的样子,所以并不是我们想要的那种生成方式。

So even if it's a video of 8 seconds, like Veo generates, you still want to feel that it's consistent, right? If the camera moves around, things should look consistent. You want to get a feel that this is really like it was taken in the real world. And the same goes once it's interactive: we want to be able to generate the next frame based on the user's input, but consistency is really important. Early models could generate the next frame, like a game engine in a way, but it didn't have very long context. The reason it worked is because it learned specific properties of that game of Doom, basically. It kind of remembered how the level looks, so it wasn't really generating as we would want.

Host

能够从文本生成内容是图像模型到视频模型的核心能力。这是这一研究方向的主要创新突破:文本成为了起点。文本是一种非常压缩的表示,也是学习概念的强大方式。所以这次从文本开始是显而易见的。我们想要描述世界,然后让你进入其中,让用户或智能体去探索。这些是项目的相似之处。在 Google DeepMind,我们试图理解这些模型如何扩展的核心机制。这些概念可以跨不同模态利用,并在延迟和内存之间进行权衡。

Being able to generate things from text is the core capability from image models to video models. That was the main innovation breakthrough for this line of research: text became a starting point. Text is such a compressed representation and a strong way to learn concepts. So it was obvious to start from text this time. We want to describe the world, then drop you into it and let the user or agent explore. Those are the similarities of the projects. At Google DeepMind, we try to understand the core mechanics of how those models scale. Those concepts can be leveraged across different modalities with tradeoffs of latency and memory.

Genie 3 的精彩示例 Favorite examples from Genie 3

Jack Parker-Holder

你有 Genie 3 最喜欢的例子吗?我真的很喜欢那只蜥蜴。虽然它不是照片级真实的,但我喜欢那只跳跃的折纸蜥蜴。我喜欢它碰到折纸河流时溅起一点水花。当然还有水坑。一些例子由团队成员发布在 X 上。我真的很喜欢那些人们与模型互动的例子。有一个例子是你四处走动,用户低头看自己的鞋子,看到它们映在水坑里。非常逼真。有很多很酷的例子展示了不同的能力。但最令人惊讶的是《盗梦空间》样本。本质上,我们可以用视频来提示模型。这很令人兴奋,因为用 VEO 3 可以生成很酷的视频,然后用那个视频提示 Genie 3 继续。有个人,Jacob,不小心没有放正确的标题。他意识到如果你不将标题与视频潜变量对齐,模型会自己搞定。所以他用一段人们在办公室房间里玩模型演示的视频来提示模型,而提示词是像丛林里有只霸王龙之类的。在 Genie 3 生成过程中,显示他们正在玩的屏幕切换到了丛林世界,笔记本电脑也切换了。Genie 3 同时更新了两者。当你转身时,你看到外面就是提示词中的丛林。当你走进丛林再回头时,你看到办公室和正在玩的人们。这表明模型有一定的理解力:它知道要同时更新两个屏幕,而且如果你在办公室里然后走出去,你应该看到一栋建筑。这真的很酷。这不是项目的目标,但在追求有趣的目标时,意想不到的事情会发生。

Do you have a favorite example from Genie 3? I really like the lizard. Although it's not photorealistic, I really like the origami lizard that jumps. I like that it splashes a little bit of water when it hits the origami river. And of course the puddles. Some examples were posted on X by team members. I really like those where people played with the model. There is one where you walk around and the user looks down to their shoes and sees them in a puddle. It's very realistic. There are loads of cool examples that show different capabilities. But the most surprising one was the Inception sample. Essentially, we can prompt our model with videos. This is exciting because with VEO 3 you can generate cool videos and then prompt Genie 3 with that video to continue. One person, Jacob, by mistake didn't put the right caption. He realized that if you don't align the caption with the video latents, the model makes it work. So he prompted the model with a video of people playing the model demo in an office room, and the prompt was like a jungle with a T-Rex. During the Genie 3 generation, the screen showing what they were playing switches to a jungle world, and so does the laptop. Genie 3 updates both. When you turn away, you see outside is the jungle as in the prompt. When you go into the jungle and turn back, you see the office and the people playing. It shows the model has some understanding: it knows to update both screens, and that if you're in an office and go outside, you should see a building. That's really cool. It wasn't the goal of the project, but unexpected things can arise when pursuing interesting objectives.

Host

一个我非常兴奋的能力是白板演示,上面有一个苹果、Genie 3 和一棵树。它展示了记忆:你看白板,看窗外,回来,它还在那里,完全一样。一切都在原位。这就是让它成为一个世界模型,让你感觉身临其境。出于同样的原因,油漆滚筒演示也非常令人印象深刻。它可能是最简单的世界:你在一个房间里,有人在刷墙。你看到视口从有随机笔触的墙壁移开,然后移回,笔触完美,完美地记住了之前的帧。对于一个自回归模型来说,如此精确地捕捉到这一点非常令人印象深刻。

One capability I'm very excited about is the whiteboard demo with an apple, Genie 3, and a tree. It demonstrates memory: you look at the whiteboard, look through the window, come back, and it's there, exactly the same. Everything is in place. That's what makes it a world model where you feel in the world. For the same reason, the painting roller demo is super impressive. It's maybe the simplest world: you're in a room and someone is painting the wall. You see the viewport pan away from the wall with random strokes, then pan back, and the strokes are perfect, perfect memory from frames before. For an autoregressive model to capture that so precisely is super impressive.

Jack Parker-Holder

当我们看到有人生成这个时,我团队里的一些人有点不敢相信,因为我们甚至不知道模型有能力做这样的事情。不仅仅是原始的视觉世界被保持,实际上你采取的行动以及这些行动的后果也被保持了。

When we saw that generated by someone, it was a bit of disbelief across some people in my team because we didn't even know the model was capable of doing something like that. It's not just that the original visual world is maintained, but actually the actions you took and the consequences of those actions are maintained as well.

模型架构与挑战 Model Architecture and Challenges

Host

我们来聊聊模型本身。从宏观层面看,挑战包括一致性、延迟、分辨率以及生成画面的丰富度。我们提到过模型本质上是自回归的,也讨论过 Transformer 和扩散模型。我们应该如何看待模型架构,以及你们如何利用建模过程的各个方面来克服这些挑战?

Let's talk a little bit about the model itself. At a high level, challenges like consistency, latency, resolution, and the richness of the produced visuals. We've alluded to the model being autoregressive in nature, and we've talked a little bit about transformer and diffusion. How should we think about the model architecture and how you've used aspects of the modeling process to overcome these challenges?

Jack Parker-Holder

模型的一个关键方面是它本质上是自回归的,这意味着在当前语境下,下一帧是基于之前发生的长序列生成的。模型必须查看特定帧之前发生的情况,推理过去,并决定哪些信息与下一帧相关。关键在于这必须非常快速地完成,每秒多次,因为我们永远无法知道用户的下一个动作是什么。这就是它成为实时交互的原因,而不仅仅是实时。这个术语引导了我们的系统和架构设计:在实时的同时保持交互性。一切都归结于那个设计决策。为了实现极低延迟和回溯能力,我们必须利用正确的架构和规模,既能实现高质量的模型,又能利用最好的硬件来构建一个可用的系统,而不仅仅是理论上的系统。

One of the key aspects of the model is that it is basically autoregressive, meaning in this context that the next frame is generated based on the long sequence of everything that happened before. The model has to look at what happened before the particular frame, reason over this past, and decide which information is relevant to the next frame. The key is that this has to happen very quickly, multiple times per second, because we can never know what the user's next actions will be. This is what makes it real-time interactive, not just real-time. This term led our design of the system and architecture: interactivity while being real-time. Everything boils down to that design decision. To achieve very low latency and the ability to look back, we had to leverage the right architecture and scale that enables both a very high quality model and the best hardware to build something that works, not just a theoretical system.

Jack Parker-Holder

为了和 Shlomi 保持一致方向,我们必须设定目标,在各个方面都雄心勃勃:记忆、高分辨率、世界的多样性以及实时性。如果你一开始不承诺,就很难一次性全部实现。这就是模型的挑战和魔力所在。我们团队中有非常出色的人,比之前的 Genie 系列团队稍大一些。我认为这几乎是一个继承了名字的新模型。优秀的人在各个组件上努力工作,同时也意识到其他部分。每一个都是挑战;没有哪个部分是容易实现的。

To go in the same direction as Shlomi, we had to set ourselves the goal of being ambitious in all dimensions: memory, high resolution, diversity of worlds, and real-time. If you don't commit to it at the beginning, it's very hard to achieve all of it in one go. That's the challenge and the magic of the model. We had amazing people in the team, a slightly bigger team than before in the Genie series. I consider it almost a new model with an inherited name. Great people worked hard on each individual component, but also with awareness of the other parts. Each one was a challenge; none of those parts was easy to achieve.

一致性作为涌现特性 Consistency as Emergent Property

Host

当你谈到挑战时,你在博客文章中并没有特别提到一致性。它提到一致性是一种涌现属性,暗示这不是你们设计的目标。是这样吗?

When you talked about the challenges, you didn't specifically mention consistency in the blog post. It mentions consistency as an emergent property, suggesting it wasn't something you were designing towards. Is that the case?

Jack Parker-Holder

这绝对是我们的目标,我们设计的方式也是为了实现这个目标。我们为模型列出的规格之一就是大约 1 分钟的记忆。但关键是,没有显式的世界表示。例如,有些方法使用带有显式网格的 3D 引擎进行渲染,或者使用神经网络和高斯泼溅来推导几何表示。这些都是显式表示,我们不想那样做。它们有局限性,尤其是在动态环境中。我们希望模型自己学习。虽然我们目标是实现一致性,但我们认为不应该在系统中内置任何东西来实现它。我们是苦涩教训的好学生,相信许多事情可以仅从数据中学习,只要你设计系统使其能够学习这些能力。当然,每个部分都必须仔细处理,因为模型会学习数据中的内容。你需要一个非常强大的模型和正确的数据,这样它才能学到正确的东西。很多因素必须结合在一起,才能在不添加那些其他方法的情况下实现这一点。

That was definitely our goal, and we designed it in a way to achieve this goal. One of the specs we listed for the model was a memory of about 1 minute. But the key thing is that there is no explicit representation of the world. For example, there are approaches like a 3D engine with an explicit mesh that gets rendered, or neural networks and Gaussian splats that derive a representation of geometry. These are explicit representations, and we didn't want to do that. They have limitations, especially with dynamic environments. We wanted the model to learn that on its own. While we aimed for consistency, we didn't think we should build anything into the system to achieve it. We are good students of the bitter lesson and believe that many things can be learned from data alone, if you design the system in a way that sets it up to learn those capabilities. Of course, every part has to be done carefully because the model will learn what's in the data. You need a really capable model and the right data so that it learns the right things. A lot of things have to come together to get that without adding those other methods.

可提示性与智能体交互 Promptability and Agentic Interaction

Host

模型的一大特点是它的可提示性。它从文本开始生成世界。博客文章中还有一个例子,你提示世界中的行为。那是 Genie 本身,还是我们在谈论 Genie 环境中的智能体?这两者之间有区别吗?你如何看待今天的情况,以及你认为整个智能体交互范式将走向何方?

One of the big features of the model is its promptability. It starts with text that generates the world. There's also an example in the blog post where you prompt the behavior in the world. Is that Genie, or are we talking about an agent within the Genie environment? Is there a distinction between these two, and how do you see that today, but also where do you see that whole agentic interaction paradigm going?

Jack Parker-Holder

我们所说的可提示世界事件并不直接与智能体绑定。你可以把它想象成上帝模式。

What we call promptable world events is not directly tied to the agent. You can think about it as God mode if you want.

可提示的世界事件 Promptable World Events

Jack Parker-Holder

你只想改变世界中的任何事物。你想来一场沙尘暴。你想扔下一个盒子。我们尝试了很多东西,比如从天上扔下物体或改变任何东西。所以基本上你可以随心所欲地改变世界中的任何东西。所以有可提示的世界事件,但也有走到面包架那里。走到这里,这与上下左右那种类型不同。

You just want to change anything in the world. You want to have a sandstorm coming. You want to drop a box. You want to... we tried a bunch of stuff like dropping objects from the sky or changing anything. So you can basically change anything in the world that you want. So there's promptable world events but there's also walk to the bread rack. Walk to this, which is different from the up-down-left-right type of...

Host

也许可以谈谈两者。你可以从世界事件开始,然后我们再谈其他内容。

Maybe talk about both. You can start with the world events and then we'll get to the other stuff.

Jack Parker-Holder

是的,我们确实有这些可提示的世界事件,允许你在世界中做出改变并注入一些新信息。这基本上允许对世界的控制超越最初提供的提示,对吧?所以这更像是一个暂时的局部事物。

So yes, we do have these promptable world events that allow you to make changes in the world and inject some new information. That allows basically control of the world beyond just the prompt provided in the beginning, right? So this is more of a temporarily local thing.

Host

我认为这是一个相当深的能力,因为并不明显——有时提示没有意义,对吧?例如,你说'一扇门打开了',而你在沙漠中央。哪扇门应该打开?模型就像'我不知道'。所以我们看到有时它会产生奇怪的东西,因为模型在尝试,但当它有意义时,我们经常看到它确实有效,并且我们得到非常好的样本。我们有一些像龙从天空中出现并降落在隧道中央的例子。所以肯定有一些情况它效果很好,是一个非常强大的能力。

And it's quite a deep capability, I think, because it's not obvious—sometimes the prompt doesn't make sense, right? For example, you say 'a door opens' and you're in the middle of the desert. What door should open? The model is like, 'I don't know.' So we see that sometimes it can make weird stuff because the model is trying, but when it makes sense, we often see that it does work and we get very nice samples. We have some like the dragon that appears out of the sky and lands in the middle of the tunnel. So there are definitely cases where it works really well and it's a very powerful capability.

Host

如果我可以暂停一下:我可以这样想,比如'沙漠中的门打开了'——我可以想象一个模型生成你的下一帧,你期望的沙漠中的下一帧,然后你将其作为输入帧来生成一个将替换持续生成中那一帧的帧。但我也可以想象这是一种粗糙的方式,而更集成到模型架构中的方式。你能谈谈这是如何实现的吗?

If I can pause on that: I can think of that as, say, 'the door opens in the desert'—I can think of a model which generates your next frame, your expected next frame with the desert, and then you use that as the input frame to generate a frame that will replace that frame in the continued generation. But I could also imagine that being a crude way of doing something that's more integrated into the model architecture. Can you talk a little bit about how that is done?

Jack Parker-Holder

我认为我们基本上想要的是能够……我们把这看作一个事件,对吧?如果你考虑在世界中行走,事情在我们周围发生。不一定由我们完成——它们不是以智能体为中心的。所以这就是你提到的关于智能体在世界中行动、可能走到某处的区别,在你提到的视频中,这是由外部模型 Seema 模型完成的。我们可以谈谈这个——我认为非常有趣。也许 Jack 可以告诉我们更多,因为他非常……我认为在 Genie 2 中也尝试过并且有效,所以非常酷,我们在此基础上进行了构建。

I think what we basically wanted is to be able to... we think of this as an event, right? If you think about walking around the world, things happen around us. Not necessarily done by us—they're not agent-centric. So that's the distinction between what you mentioned about the agent in the world acting in the world, maybe walking somewhere, which in the videos you mentioned was done by an external model, the Seema model. We can talk about that—I think it's very interesting. Maybe Jack can tell us a bit more because he's very... I think something that was also tried with Genie 2 and worked, so it's really cool, and we built on top of that.

Jack Parker-Holder

但如果我们回到可提示的世界事件,这种能力不仅仅基于单帧,对吧?可能是你想看到世界中的某样东西,然后你看向它——它不会立即发生——但随后你向左看,你看到例如一个人。所以我们有一些例子,你滑雪下山,然后向左看,你看到一个穿着 Genie 3 T 恤的人。所以它可以在世界中具体化事物,但这并不意味着它只是突然出现在你面前。我们希望它理想上是集成到世界中的,并且有意义,对吧?因为很容易只是扔下一个看起来很人工的东西。我们希望它实际上被集成,看起来真实。所以模型最终希望制造看起来像训练数据的东西,最终应该是逼真的。所以某种程度上,额外的条件信息是下一帧生成过程的一部分,而不是仅仅把这个东西扔到视野中央。

But if we go back to promptable world events, the ability is not just based on a single frame, right? It can be that you want to see something in the world and then you look—it doesn't happen immediately—but then you look to the left and you see, for example, a person. So we have some examples where you ski down the slopes and then you look to the left and you have a person wearing a Genie 3 t-shirt. So it can materialize things in the world, but that doesn't mean it just pops in front of you. We want it to be ideally something that's integrated and makes sense in the world, right? Because it's easy to just drop something that looks very artificial. We want it to actually be integrated to look real. So the model eventually wants to make things that look like training data, which ultimately should be realistic. So somehow additional conditioning information that's integral to the next frame generation process, as opposed to just dropping this thing in the middle of the view.

Host

非常有趣。所以你告诉模型去做,它就像'我准备好了再做'那种感觉——这不是技术术语,但没错,它以一种自然的方式去做。

Super interesting. So you're telling the model to do it and it's like 'I'll do it when I'm ready' kind of thing—it doesn't... that's not the technical term, but yeah, it does it in a way that feels natural.

Seema 智能体 Seema Agents

Host

Jack,谈谈 Seema 智能体。

Jack, talk a little bit about the Seema agents.

Jack Parker-Holder

当然。正如我所说,回顾这个项目的历史,我们将其设计为智能体的环境。在 Google DeepMind,显然我们有很多项目在研究智能体,而其中一个特别专注于 3D 世界的是 Seema 智能体。所以他们试图训练能够在 3D 模拟环境中实现语言目标的智能体。他们有一个公告,我想大概是 2024 年 2 月左右的一篇博客文章,展示了他们对此的一些思考。他们现在正在做的是在现有游戏中训练。所以他们有一个非常能干的智能体,可以在不同的游戏世界中做各种各样的事情,但最终它受限于只能访问这些游戏世界,对吧?所以它不能在任何可想象的游戏世界或现实世界中训练,因为它只能访问有限的一组环境进行训练。而这正是 Genie 试图解决的问题,对吧?生成新的环境。但 Seema 智能体也出奇地通用,对吧?尽管它是在一个较小的世界集上训练的,你可以把它放到一个它从未见过的 Genie 环境中。所以你可以用文本创建一个 Genie 3 环境或世界,比如你描述一个场景,可能像工厂车间之类的。你可以说背景中有一辆叉车。你生成这个世界,然后你对 Seema 智能体说'去叉车那里',对吧?或者你甚至可以对它说'去那个能举起东西的东西那里'之类的。然后 Seema 智能体从那时起就把 Genie 生成的世界当作任何其他环境一样对待。它不知道这是一个模型。它什么都不知道。它只看到像素,然后说'我要按这个键来实现这个目标'。然后 Genie 只看到按键,对吧?它不知道智能体试图做什么——因为如果它知道,它可能会让事情发生,对吧?它只知道它想向前走。然后它模拟下一帧,Seema 智能体看到下一帧,说'好的,我要继续向前走'。这些是同步进行的,来回交替。关键是,如果智能体做了错误的动作,它就不会实现目标。如果它做了正确的动作,它就会实现目标。

Sure. So as I said, going back to the history of the project, we designed this to be an environment for agents. And at Google DeepMind, obviously we have lots of projects working on agents, and the one that's really focused on 3D worlds is the Seema agent. So they're trying to train agents that can achieve language goals in 3D simulated environments. They have an announcement from, I think, a blog post from probably around February 2024 where they showed a bit about how they're thinking about this. What they're doing right now is training in existing games. So they've got a really capable agent that can do quite diverse things in different game worlds, but ultimately it's limited by only having access to those game worlds, right? So it can't train in any imaginable game world or in the real world because it's got just access to a finite set of environments to train in. And this is kind of the exact problem that Genie is trying to solve, right? To generate new environments. But also the Seema agent was surprisingly general, right? Even though it's been trained on a smaller set of worlds, you can kind of drop it in one of the Genie environments that it's never seen before. So you use text to create a Genie 3 environment or world, say you describe a scene that could be like a factory floor or something like that. And you could say in the background there's a forklift truck. And you generate this world, and then you say to the Seema agent, 'go to the forklift truck,' right? Or you could even say to it, 'go to the thing that can lift things' or something like that. And then the Seema agent, from that point onwards, treats the Genie-generated world as if it's any other environment. It doesn't know that it's a model. It doesn't know anything. It just sees the pixels and it says, 'I'm going to press this key to achieve this goal.' And then all Genie sees is the key press, right? It doesn't know what the agent's trying to do—because if it did know that, it might make it happen, right? All it knows is it wants to go forward. And then it simulates the next frame, and then the Seema agent sees the next frame and says, 'Okay, I'm going to keep going forward.' And these happen in tandem, back and forth. And critically, if the agent does the wrong actions, it won't achieve the goal. If it does the right actions, it will achieve the goal.

智能体学习与世界事件 Agent Learning and World Events

Jack Parker-Holder

那么当然,你可以看到同一个智能体可以从这个经验中学习,更频繁地实现目标,对吧?可能有些事情它现在还做不到,但它可以在这些世界中学会做。所以我们基本上有了生命的迹象:一个智能体与另一个智能体交互,本质上是在这些更具身的世界中教它新技能,对吧?但规模是前所未有的。然后为了闭环,一个很酷的事情是将其与世界事件整合,对吧?因为即使是一个像在街上散步这样温和的环境,如果你注入一些东西,比如一只猫跳出来,也会变得有趣得多。所以实际上,你可以教我们的智能体对所有这些不同的事情保持鲁棒性,即使在简单的环境中,因为我们有额外的杠杆可以拉动环境方面,使其对智能体更有趣和更具挑战性。

So then of course you can see that the same agent can then learn from this experience to achieve the goal more often, right? And there may be some things it can't do yet but it could learn to do in these worlds. So we essentially have signs of life that we have one agent interacting with another one essentially to teach it new skills in these more embodied worlds, right? But at a scale that hasn't really been done before. And then to kind of close the loop on this, a really cool thing could also be integrating this with the world events, right? Because even a kind of benign environment like walking down the street might become much more interesting if you then injected, say, I don't know, a cat jumps out or something like that. So actually you can teach our agents to be robust to all these different kinds of things even in simple environments because we have this additional lever to pull on the environment side to make it more interesting and challenging for the agents.

Host

有趣,有趣。嗯,Jack,之前我们稍微谈到了 Genie 1 的局限性,你知道,在传达局限性时,是放在前面还是放在页面后面,而在你们的 Genie 3 博客文章中,它们被放在了后面。但我希望你能稍微谈谈这些局限性,然后我会把话题交给你,开始讨论下一步以及你认为研究的方向。那么我们来谈谈局限性吧。

Interesting. Interesting. Um Jack, earlier we were talking a little bit about the limitations of Genie 1 that, you know, in this kind of balance in, you know, communicating the limitations, leading with them versus, you know, having them further down in the page and they are further down in the page in your, you know, Genie 3 blog post. But I'd love to have you riff a little bit about the limitations and then let me I'll turn it over to you to talk a little bit about or at least start us off talking about kind of next steps and where you you see the research going and then you, Jack. So let's talk about limitations.

Genie 3 的局限性 Limitations of Genie 3

Jack Parker-Holder

你知道,怎么说呢……我想我们记不全列出的所有东西,但对我来说最突出的一点是,我们谈到了模拟其他智能体。而我们并没有在世界上进行任何多智能体交互。没错。是的,我认为我们提到过这一点目前非常有限。最终,模型是在预测下一帧。它能够对其他智能体进行某种非常基础的模拟。比如,如果你挡了他们的路,他们正在走路而你站在前面,他们可能会停下来;或者如果一辆车在开,你走到它前面,它可能会停下。但这不是非常复杂的交互。所以我认为这绝对是一个局限性。而且这也是在视频等领域可能已经更先进的东西,对吧?所以我们肯定还没有。我认为还有明显的 1 分钟限制。说起来有点好笑,因为对于 Genie 1 和 2 来说,一分钟简直不可思议。但我认为这有点像我们领域的节奏非常疯狂,对吧?所以这类东西在未来肯定会显得短得可笑。所以现在我们说我们有一分钟的视觉记忆。

You know, how do you... So I think what we can't remember all the sorts of things we listed but the one that I think stands out for me is that we talked about simulating other agents. And we don't do any multi-agent interactions within the world. Exactly. Yeah, I think we mentioned that this is quite limited at this point. Ultimately the model is predicting the next frames. It's able to get some kind of very basic simulation of other agents. Like if you walk in their way, if they're walking and you stand in the way, they might stop, or if a car is driving and you walk in front of it, it might stop. But it's not the case that this is like very complex interaction. And so I think that's definitely a limitation. And it's also something that you maybe is much more advanced in something like video already, right? So it's something that we don't have for sure. I think there's clearly the 1-minute limitation as well. Like it's kind of funny to say it because like Genie 1 and 2, a minute was sort of like would have been seen as incredible. But I think it's a bit like the pace of our field is absolutely crazy, right? So this is the kind of thing that will seem embarrassingly short, I'm sure, in the future. So right now we kind of say we have visual memory for a minute.

Host

我想那是指单个交互或游戏可以持续几分钟。所以本质上就是上下文长度。

And I guess that is an individual interaction or play of Genie can span multiple minutes. So it's just the context length essentially.

Jack Parker-Holder

是的,这是一个重要的区别。比如你可以玩好几分钟,它不会退化或变得非常模糊,像前几代模型那样,但记忆仍然只有大约 1 分钟。是的,没错。然后还有,我想现实世界的物理准确性并不完美。所以如果你说“我想要我在伦敦的那条街”,它不会知道我的街道。所以我认为有些方面是可以改进的,如果你用文字描述一个非常抽象的世界,它几乎肯定会准确无误。但如果你描述一个具体的地理位置,你可能会发现它并不完全是你希望的样子。所以我认为这是另一个局限性。

Yeah, then this is an important distinction. Like you can play for multiple minutes even and it doesn't degrade or it becomes very blurry like, you know, previous generations of those models but again the memory is around 1 minute. It's Yeah, exactly. And then and then also there is I guess like real world physical accuracy is not perfect. So if you say like I want my exact street in London it won't it doesn't know my street. So I think that there's some there's some I guess elements of that that that can could be improved if you wanted to if you in text if you describe sort of like a very abstract world, it will almost certainly get it like on the money. If you describe a specific geographic location, you might notice that it's not what you hoped it would be in some way. So that's another one I think that that is a limitation.

Host

是的。那么我,有没有什么你觉得突出的局限性?

Yeah. So me, are there any you know, anything that jumps out at you in terms of limitations or...

Jack Parker-Holder

是的,我认为有点像你之前问过的关于智能体能够采取行动的问题。目前动作空间相对受限,对吧?虽然我们可以让智能体导航,我们有一些动作,比如跳跃或开门,但就智能体所采取动作的语义而言,这些动作相对基础。所以可提示的世界事件让我们能够控制世界,但它们不一定是智能体中心的动作。所以我认为这绝对是我们希望在未来改进和扩展的东西,因为它确实是一个真正的局限性,我们认为特别是为了制作……这相当具有挑战性,但对于任何智能体来说,如果我们希望智能体采取更复杂的动作,不仅仅是四处走动,而是能够例如捡起东西,也许输入一些代码,或者与另一个智能体对话。你知道,世界上可以发生很多事情。这是一个相当具有挑战性的问题,因为……我们作为人类以非常物理的方式在世界中运作,对吧?我们用手,用脚走路,我们有一种具身的存在。当我们去掉所有这些,只剩下视觉和像素时,定义实际应该发生什么动作就变得困难得多,对吧?例如,当你开门时,你走过去,不仅仅是开门,对吧?你走过去抓住门把手,然后转动它,对吧?或者把它拉向你。有一系列微动作在发生。我认为这是一个具有挑战性的问题,如何建模这个动作空间。但这绝对是一个局限性,我认为也是一个扩展能力的机会。

Yeah, I think a bit similarly to what you've asked before about the agents being able to take actions. So while currently the action space is relatively constrained, right? So while we do we can have you know, the agent can navigate. We have some actions like, you know, maybe jumping or like some opening doors but it's relatively basic in terms of the semantics of the action that the agent is taking. So promptable world events give us control over the world but they're not necessarily agent-centric actions. So I think this is definitely something that we we hope to improve and expand in the future but because it's actually it it is a real limitation and we think that especially for making the... So it's quite challenging but it's for anything that agent like if we want the agent to take more complex actions, not just walk around but actually be able to for example pick up things, maybe maybe do like typing some code or maybe talk to a different agent. You know, there is a lot of the many things can happen in the world. And it's quite a challenging problem because there is no... we operate in the world as people or as in a very physical way, right? We we use our hands, we use our feet to walk, we we can we have a kind of an embodied kind of like a presence. And and when we take away all of these and we are left with the visual with pixels only, then it's it's much harder to define what what actions should actually happen, right? When you open the door for example, you go and it's not just open the door, right? You you go and you grab the knob of the door and and you move it, right? Or and you pull it towards you. There's sequence of like micro actions that are being taken place. And and I think this is a there is a challenging question of how to model this this space of actions. But it's definitely a limitation and I think an opportunity to to expand the capabilities.

下一步与期待 Next Steps and Excitement

Host

那么深入探讨下一步,我感觉你们俩都对智能体方面非常兴奋,这从 DeepMind 出来并不意外。但对你来说,最令人兴奋或最明显的是什么?或者就项目方向而言,你看到了什么?我们可以从你们每个人开始。

So digging into next steps, I get the sense that you're both pretty excited about the agentic aspects of this, no surprise coming from DeepMind. But what's like most exciting to you or most obvious to you like or you know, present for you in terms of like where the project goes. We can start with each of you.

Jack Parker-Holder

对我来说,能够踏入一个世界,对吧?你创造的或别人创造的,然后你实际上可以感知它、看到它、与它互动。我认为这非常巨大,可以应用于很多事情,从娱乐开始,这很明显,对吧?你提到过街景,比如交互式街景。所以它可以部分锚定在现实世界中,但带你去别的地方。这实际上让我想起了很久以前我工作过的一家初创公司,我们有一个类似游戏,设置在旧金山市中心。

So to me the ability to be able to step into a world, right? That you created or someone else created but then you can actually perceive it, see it, interact with it. I think it's huge which is can be really applied to so many things from, you know, entertainment which is very obvious, right? You said like, you know, street view, interactive street view for example. So it can be somewhat anchored in the real world but take it somewhere else. And it actually reminds me a startup I worked for like a long time ago and we had like this kind of like a game placed in downtown San Francisco.

Genie 3 的实际应用 Real-world applications of Genie 3

Jack Parker-Holder

当然,我们必须对整个旧金山市中心进行建模,这工作量很大,但这家初创公司的核心理念是让游戏发生在真实世界的地点。所以,例如,你可以进行任何互动,但不是在一个真实地点中放置一个不真实的互动。这只是众多例子中的一个。我认为还有其他非常有趣的应用,不一定是娱乐。可以是教育,例如。它可以帮助人们看到自己完成了一些他们原本认为不可能的事情。看到自己完成某件事有一种非常强大的力量,从心理学角度来看,这非常强大。而且个性化方面还不够——能够进入环境、四处走动,也许通过提示让环境看起来非常像你的家。所以如果你害怕蜘蛛,也许你可以看到自己在家里走到蜘蛛旁边,然后你的大脑会说:‘好吧,我能做到。’所以我认为现实世界的潜在应用非常明显。

Of course we had to model the entire downtown San Francisco and it was a lot, but the key idea of the startup was to actually have games happening in real world locations. So for example, you could have any interaction, but not a realistic one placed in a realistic location. That's just one example among many. And I think there are other really interesting applications for people, not necessarily entertainment. It can be education, for example. It can be helping people see themselves accomplishing something they wouldn't have expected. There is something very powerful about seeing yourself accomplishing something, and it kind of makes you, from a psychological perspective, it's very powerful. And it's something that's not enough—the personalization aspect of being able to go into the environment, walk around, maybe prompt it in a way that looks very similar to your house. So if you're afraid of spiders, maybe you can see yourself walk next to a spider in your home, and then your brain says, 'Okay, I can do it.' So I think the real-world potential applications are very obvious.

Host

是的,没错。所以它不一定非得是……我的意思是,我们不一定知道这项技术会走向何方,现在还为时过早。这就是为什么我们最初让一些值得信赖的测试者或学者与模型互动。我们想获得一些反馈,并希望随着时间的推移,更多地了解人们感兴趣的能力和应用。

Yeah, yeah, exactly. So it doesn't have to be... My point is that we don't necessarily know where this technology will go, and it's very early days for that. And that's why we had some trusted testers or academics interact with the model initially. We wanted to get some feedback, and we hope over time to learn more about the capabilities and the applications that people are excited about.

Genie 3 的下一步 Next steps for Genie 3

Host

杰克,你接下来的计划是什么?

Jack, next steps for you?

Jack Parker-Holder

是的,已经提到了很多令人兴奋的想法。我认为最让我兴奋的是教智能体在视觉上逼真的具身世界中与人类互动。我认为这是我们当前任何智能体都缺失的能力——在物理世界中与人类互动。我认为像 Genie 3 这样的模型可以实现这一点。而且我真的认为没有其他方法可以做到。所以我认为这非常令人兴奋,尤其是结合世界事件,对吧?这可以生成多样化的场景,而我们无法通过其他方式获得数据。所以我认为这仍然处于早期阶段,但我认为这是一大步,将真正开启许多用例。

Yeah, so there are so many exciting ideas already mentioned. I think the one that really excites me is teaching agents to interact in visually realistic, embodied worlds with people in the worlds. I think that's a missing capability for any of our current agents—to interact in the physical world with humans. And I think models like Genie 3 could enable that. I also don't really think there's any other way to achieve that. So I think that's something really exciting, especially with the world events, right? Which could enable generating diverse scenarios that we wouldn't be able to get data for any other way. So I think this is still fairly early in the journey, but I think this is a big step that will really open up a lot of use cases there.

结束语 Closing remarks

Host

好的,Slawomir Jag,非常感谢你们参与并更新了 Genie 3 以及你们正在做的一切。深入探讨这些内容真的非常棒。

Well, Slawomir Jag, thank you guys so much for jumping on and updating us on Genie 3 and everything that you're working on. It's been really great to dig into it.

Jack Parker-Holder

太棒了。非常感谢你的时间。

Awesome. Thanks so much for your time.

Host

是的,谢谢 Sam。好的,谢谢你们两位。干杯。

Yeah, thanks Sam. All right, thank you both. Cheers.

互动版:逐字朗读 + 针对本期提问 →