Oriol Vinyals on World Models, Multimodal AI, and the Path to AGI
打开互动全文版(中英对照 + 朗读 + 问答)→Google DeepMind 的 Oriol Vinyals 探讨世界模型、多模态 AI 进展以及 AI 推理和记忆的未来。
Google DeepMind's Oriol Vinyals discusses world models, multimodal AI advances, and the future of reasoning and memory in AI.
Oriel Vignal 是 Gemini 的联合负责人,与 Nome Shazir 和 Jeff Dean 一起。他在 AI 领域有着令人难以置信的职业生涯,在过去十年中开创了深度学习的许多突破。在 Google IO 之后与他坐下来交谈非常有趣。如果你关注了 Google IO,他们基本上在 AI 的许多有趣领域推出了一系列产品。所以 Oriel 和我讨论了所有这些。我们谈到了进一步推进多模态模型需要什么,以及什么能让这些世界模型真正可用。我们谈到了记忆的增加和记忆的重要性,以及这些进展在未来几年将如何表现为推理,以及 Oriel 认为的前进道路。我们还讨论了脚手架(scaffolding)的现状,人们在构建什么,以及 Oriel 认为什么会持续存在。基本上,把所有创始人、投资者正在思考的顶级问题都抛给 Oriel 是非常有趣的。所以我认为大家会喜欢这次对话。话不多说,有请 Oriel。Oriel,非常感谢你来做客播客。
Oriel Vignal is the co-lead of Gemini alongside Nome Shazir and Jeff Dean. He's had an incredible career in AI, pioneering many of the breakthroughs in deep learning in the last decade. And it was a ton of fun to get to sit down with him after Google IO. If you've been following Google IO, they basically shipped a bunch of products across a ton of interesting surface areas throughout AI. And so Oriel and I hit all of them. We talked about what's required for further advances in multimodal models and what's going to make these world models actually usable. We talked about the increase in memory and the importance of memory and how the advances there will look like reasoning these next few years as well as what Oriel thinks the path forward is. And we hit on the state of scaffolding today, what folks are building and what Oriel thinks persists. It's a ton of fun to get to basically take all the top questions that founders, investors are thinking through and just pose them to Oriel. So I think folks will really enjoy this conversation. Without further ado, here he is. Oriel, thanks so much for coming on the podcast.
是的,很高兴来到这里。谢谢你,Jacob。
Yeah, it's great to be here. Thanks, Jacob.
是的,在 IO 之后一天能邀请到你非常激动。我知道事情一直很忙,但我真的很兴奋,因为你是当今最直接塑造模型前沿的人之一,以及你在 Google 的工作。显然,昨天在 IO 的发布中,它们几乎涵盖了人们在这个领域思考的所有主题,这些产品和模型的发展方向。所以我觉得我们今天的目标就是讨论这些公告背后的研究。你知道,这一切的走向,RL 和后训练的未来路径,以及你对整个领域的看法。我想从世界模型开始,因为我认为那是昨天非常令人印象深刻的部分,而且我认为 Google 在这个领域与许多其他实验室有显著区别。所以你们昨天在 Omni 中推出了这个令人难以置信的世界模型,而且 Demis 经常谈到将世界模型视为通往 AGI 的路径,这很有趣,因为其他实验室可能更专注于代码和递归自我改进。所以我想知道这是否是一个公平的描述,以及为什么你认为你和 Google 团队在这个世界模型领域有独特的关注。
Yeah, very exciting to have you a day after IO. I know things have been busy, but I've been really excited for this because you're one of the people kind of most directly shaping the frontier of models today and your work at Google. And you obviously in the releases that happened yesterday at IO, they hit on like pretty much all the themes that people are thinking about in the space, where these products and models are going. And so I feel like there's just our goal today is to talk through kind of the research behind those announcements. You know, where this is all headed, you know, the kind of future path of RL and post-training and, you know, get your read on the space as a whole. I figured where I'd start was with world models because I think that was just a really impressive part of you know of yesterday and also I think a pretty you know where Google's pretty distinct from a lot of the rest of the field. So you obviously shipped this incredibly impressive world model in Omni yesterday and you know I think Demis has talked a lot about you know seeing world models as a path to AGI and it's interesting right because it seems like other labs maybe are more focused on code and you know getting to recursive self-improvement and so I'm wondering if it's a fair characterization and you know why you think you know you and the team in Google have been somewhat uniquely focused on this world model space.
首先,我认为编码或自我改进的角度是在一个稍微不同的层面,对吧?所以,你当然可以打赌并相信这些模型可以重新编程和改进自己,这也是我目前正在积极研究的事情。但是它们改进的对象——模型,无论是多模态的还是更接近我们所说的世界模型,以及如何定义它,都有些抽象。从第一天起,甚至在 Gemini 项目真正开始之前,我们就在研究不仅仅是语言,而是理解视觉世界,并在视觉、空白视频等背景下联合建模词汇。所以我认为这部分一直是 Gemini 和我们之前研究的核心。我认为也许一种描述方式是:语言显然包含了我们关于世界的大量集体信息。所以这显然取得了巨大成功。我们在某种程度上将所有已写和正在写的知识蒸馏到了这些权重中。
First of all, I guess the coding or like self-improvement angle is at a bit of a different layer, right? So, you can certainly bet and believe that these models can reprogram and improve themselves and it's something I've been actually quite actively working on at the moment. But then the object that they improve, the model, whether it's multimodal and closer to a world model as we call it, and even how to define that is a bit abstract. Since day one and way before actually Gemini program started, we were working on not just language but understanding the visual world and kind of jointly modeling words in the context of vision, blank video, etc. So I think that part, it's been at the core of Gemini and before our research. And I think maybe one way to characterize it is: language clearly has a lot of information collectively that we wrote about the world. So that's clearly paid off big time. We've kind of distilled in a way all the knowledge written and that is being written at the moment into these weights.
我们把它都放在互联网上确实很方便。
It was definitely convenient that we put it all on the internet too.
是的,没错。所以,现在还有用户,对吧?显然也有飞轮效应,但与此同时,视频和图像中有很多知识。我想说的是,这种情况已经发生,但比较温和。我认为可能有一个重大时刻:如果你观看所有视频和图像(我们当然在训练混合中使用了它们),你将如何提取所有知识?但这些知识能否为语言组件增加价值和效率?我认为我们已经看到了从一种模态到另一种模态的建设性迁移学习。我们看到了这一点,也看到了泛化,但可能我所说的视频和图像的 GPT 时刻,我不确定我们是否已经看到了。
Yes, exactly. Right. So, and also like this now with users, right? There's obviously also a flywheel effect, but at the same time, there is lots of knowledge in videos and images. And what I would say is it kind of has happened but softly. I think there probably might be a big moment: how would you extract all the knowledge that you would acquire if you were to look at all the videos and images which we certainly use in our training mixtures? But could that knowledge somehow add value and efficiency to the language component? And I think we've seen constructive sort of transfer learning from one to the other. We see that and we see generalization, but probably what I would characterize as the GPT moment of video and images, I'm not sure we quite have seen that.
你对视频和图像的 GPT 时刻可能是什么有什么想法吗?因为你有一种直觉,觉得它还没有到来。
Do you have any thoughts on what that GPT moment might be for video and image as you kind of have this intuitive feeling that it hasn't yet been reached?
是的。目前我们训练所有模态,混合它们并不断改进配方。所以 Omni 是看到这一进展的好例子,我们不仅输入视频和图像,还看到了长上下文理解等惊人能力。而且我们现在还能够输出视频,并通过语言以非常自然的方式与之交互、编辑、组合模态,感觉几乎是神奇的,对吧?所以进展确实存在。但也许深度学习的一个梦想,可能是在大型语言模型之前就有的原始梦想:嘿,我能否在没有文本的情况下训练所有图像数据,作为一个艰巨的挑战,仍然从该模态或模态集合以及大量数据中提取所有意义和细微差别?对吧?所以我们能否在所有已产生的视频和图像上训练,并达到与使用语言的语言模型相同的理解水平,尽管可能稍微表面化,并且缺少一些因果联系等,例如 Demis 经常提到的?所以那就是那个时刻。我看到了吗?可能没有。而且很可能我们拥有最先进或最先进的多模态配方之一,混合了所有东西,但那种纯粹的迁移我认为是过去十多年机器学习的核心追求之一。
Yeah. So at the moment we train all the modalities, we mix them and we keep enhancing the recipe. So Omni is a good way to see that progress in which we not only input videos and images, we've seen amazing capabilities with long context understanding etc. But we also now are able to output video but also interact with it in a very natural way through language, editing it, combining the modalities in a way that feels almost magical, right? So that progress is absolutely there. But maybe one of the deep learning dreams, and it might be an original kind of dream from way before large language models, would be: hey, can I train on all the image data without text perhaps as a hard challenge and still somehow extract all the meaning and nuance from that modality or set of modalities and vast amounts of data? Right? So could we train on all the videos ever produced and images and get to the same level of understanding that clearly the language models using language get to, although probably slightly superficially and with some missing links with cause and effect and so on that for instance Demis talks about often, right? So that is the moment. Have I seen that? Probably not. And most likely we have the most advanced or one of the most advanced multimodal recipe that mixes everything, but that pure transfer is I think one of the core quests of machine learning for the last decade plus.
我的意思是,在你能谈论的范围内,我很好奇你能给我们的听众一些背景,关于围绕这个还有哪些关键问题需要解决,或者当你思考你试图进一步推进这个领域时,你正在研究哪些类型的问题。
I mean to the extent you can talk about it, I'm curious could you give our listeners some context on what are still the key problems that need to be solved around this or as you think about the types of problems that you're trying to work on to further advance this.
很难描述解空间,但通过观察或学习所有视频数据,然后推导出重力规则的想法经常被用到。如何仅基于图像精确描述世界如何运作?问题在于将语言或这些概念与图像中的内容联系起来,而没有显式的语言关联,这相当棘手。所以最终的做法是尝试显式创建数据集,让图像、视频与某种语言(如标签或描述)之间存在关联。但当然,可用的数据量要少得多,因为我们并没有清晰描述和转录每一段媒体。因此,以最纯粹的形式提取这些概念,而不仅仅是我们关联到词语和所见之物的语言,将非常强大。关于离散表示和表示学习有很多早期研究。这可能是仍处于研究阶段的事情之一,目前还无法规模化。但如果能实现突破,那将是巨大的。
It's hard to describe the solution space, but the idea of observing or learning from all video data and then somehow deriving the rules of gravity is often used. How could you precisely describe how the world works based only on images? The issue is linking language or these concepts to what you see in the image without explicit language linkage is fairly tricky. So what you end up doing is trying to explicitly create datasets where there's some correlation or connection between images and video and some language, like labels or descriptions. But of course, the amount of data at your disposal is much less because we haven't clearly described and transcribed every single piece of media. So extracting those concepts in the purest form, not just in some language we associate to words and what we see, would be very powerful. There's lots of early research on discrete representations and representation learning. That's one of the things that is still in a research stage, not something we can scale up yet. But if it were to be unlocked, it would be massive.
你提到了世界模型这个词,以及它被广泛使用。Omni 被定位为一个世界模型。我很好奇你如何看待这种分类,与你们一直在研究的视频模型相比,是什么让 Omni 成为一个世界模型,它又有什么不同?
You mentioned the term world model and how it's thrown around. Omni was positioned as a world model. I'm curious how you thought about that categorization versus the video models you've been working on. What makes Omni a world model and how is it different from the generation of video models?
世界模型的一个纯粹方面是表示学习。你可以想象将视频(图像序列)等模态压缩成一组概念——运动、物体等。这就是表示学习,以紧凑的方式建模世界,压缩掉可能不相关的信息。这更经典,但可能不是我们与 Omni 交互时的真正含义。你看到的是能够改变视频行为或从初始图像生成动画的方式。你明确要求移动或动作,比如“向前移动”,然后看到它被精确模拟。所以世界模型充当了一个世界的渲染器,你可以通过语言改变它。除了是一个很酷的产品,它还可以为在现实世界中行动前的预测增加一个模拟维度。3D 或视频世界模型的明显应用是自动驾驶汽车或机器人。
A pure aspect of world model would be representation learning. You could imagine taking modalities like videos, which are sequences of images, and compressing them into a set of concepts—movements, objects, etc. That's representation learning, modeling the world in a compact way, compressing away what's probably not relevant. That's more classical but probably not exactly what we mean when we interact with Omni. What you see there is more about being able to change how the video behaves or the kinds of videos you get from an initial image you ask to animate. You explicitly ask movements or actions like 'move forward' and see it precisely simulated. So the world model acts as a renderer of the world that you can change with language. Besides being a cool product, it could add a dimension of simulation for prediction before acting in the world. Obvious applications for 3D or video world models are self-driving cars or robotics.
这似乎与机器人技术非常相关。每个人都在试图找出模拟数据、远程操作数据和自我中心视频数据的正确数据组合。随着模拟越来越好,将其纳入数据组合越来越有吸引力。这项工作是否直接与你们正在进行的更广泛的机器人研究交叉?你认为需要什么才能将机器人动作附加到这些类型的模型上?
It seems so relevant to robotics. Everyone is trying to figure out the right data mix of simulation data, teleop data, and egocentric video data. As simulations get better, it's more compelling to put into the data mix. Does this work directly intersect with the broader robotics work you're doing? How do you think about what's required to append robotic actions onto these types of models?
有一个美妙的联系。如果我们获取更多从机器人捕获的数据——即使更昂贵或更耗时——这些数据可以进入模型,增强世界模型的能力。另一个方向是我们可以模拟并创建大量不同场景供机器人训练,而无需物理世界的成本和时间延迟。要让后者更好地工作,仍然是一个开放问题,存在迁移问题。但随着这些模型越来越强大,会出现一个转折点,事情开始值得去做,我们可能会看到机器人技术的加速。在硬件方面,我们看到大量投资,所以事情正在加速。要让世界模型有用,即使是抓取的精度——我们人类认为理所当然的——视觉、手的感觉(我们目前没有数据的模态)以及精确的力都需要非常准确。这就是差距所在,需要创造力和研究。但这很有希望。在某种程度上,也许不是用于精确的运动控制,而是用于规划和粗略运动,我们将开始看到这些模型加速机器人技术的进步。
There's a beautiful connection. If we acquire more data captured from robots—even if it's more expensive or time-consuming—that data could make it into the model, enhancing the world model capabilities. The other direction is that we can simulate and create lots of different scenarios for robots to train on without the cost and time latency of the physical world. For the latter to work better, it's still an open problem with transfer issues. But the more powerful these models get, there's an inflection point where things start to be worth doing, and we might see an acceleration in robotics. In hardware, we're seeing lots of investment, so things are accelerating. For world models to be useful, the precision of even grasping—which we take for granted as humans—the visuals, the feel to your hand (a modality we don't have data for), and the exact forces need to be very accurate. That's where there's a gap, requiring creativity and research. But it's promising. At some level, maybe not for precise motor control, but for planning and gross movements, we will start seeing these models accelerate progress in robotics.
这些模型的一个很大部分是通过消费大量视频数据隐式学习物理。你提到重力是典型例子。你离这些模型这么近,有没有直觉认为什么时候这会在世界模型内成为一个已解决的问题?
A huge part of these models is learning physics implicitly through consuming lots of video data. You mentioned gravity as the canonical example. Do you have any gut sense, being so close to these models, of when that will be a solved problem within world models?
这是个好问题。它让我想到了评估。如果你训练一个非常好的视频模型,你如何评估?如何评估模型中的物理?一旦你加入语言,突然之间,这些知识就在权重中了。
That's a good question. It made me think about evaluation. How would you evaluate if you train a very good video model? How do you evaluate physics in a model? As soon as you add language, all of a sudden that knowledge is there in the weights.
所以如果你问关于重力的基本问题,你当然可以通过阅读网上的解释来回答。所以你需要以某种方式将重力的概念(可能存在于世界模型中,也可能不存在)连接起来,然后将其解码成一个让你满意的解释。最初可能是一些基本解释,后来甚至可以推导出方程等等。这就是构建评估的方法。据我所知,我们还没有从这个角度思考过。确实有很多关于无监督机器翻译的早期工作,你试图翻译成训练中从未见过的语言,并且可以对齐表示。所以有一些想法关于如何让语言模型能够说话,或者你可以从世界模型中解码出这种概念层面的理解并对齐两者。有一些论文,我是说这些是旧论文,我记得有一篇是 Stefan Gaus 等人 2014 年的。但你可以尝试开始解码并将其转换为评估,这似乎是一个微不足道的步骤。但同样,这些需要从应用的角度来看是有意义的。所以最终你也可以说,看,我们有一个世界模型,我们能否从其表示中解码或引发复杂系统中的运动?那将是另一个间接评估。所以很多想法,但确实非常重要。
So if you ask basic questions about gravity, of course you would answer them by just having read explanations of them online and so on. So you would need to somehow connect the concept of gravity, which could be present or not in a world model, to then decode that into an explanation that would satisfy you. Maybe initially it would be some basic explanation, later on could even derive the equations and so on. That's how you could build an eval. I don't think, to my knowledge, we've been thinking about this from this point of view. There's definitely lots of early work on unsupervised machine translation where you would try to translate to a language that you would never see during training, and you could align the representation. So there's probably some ideas on how you get a language model that can speak, or you can decode from, you get these world models that would create this kind of conceptual level understanding and aligning both. There are some papers, I mean these are like old papers, the one I remember from I think it was Stefan Gaus et al. was from 2014. But then you could try to start decoding that and converting that to an eval seems then a trivial step. But again, these need to then be meaningful from an application point of view. So ultimately you could also say, look, we have a world model, can we decode or induce movement in a complex system from its representation, for example? That would be another indirect eval. So many ideas, but yeah, they are so important.
换个话题,谈谈你们昨天发布的其他一些东西。我肯定想聊聊智能体。你们在 Spark 中发布了一些非常有趣的消费级智能体,作为 IO 的一部分。我觉得这很有趣,因为从外部看,至少感觉像是你们在 2024 年 Project Mariner 中探索的一些东西以及其他计算机使用工作的一个显著改进版本。所以感觉能力确实有了质的飞跃。我很想听听你谈谈实现这一点的研究突破,以及人们应该如何思考这些智能体今天能做什么和不能做什么。
Shifting gears to some of the other stuff that you all shipped yesterday. I definitely want to talk agents. You shipped some really interesting consumer agents in Spark, right, as part of IO. And I think it's so interesting because it seems like, from the outside at least, a really improved version of some of the stuff you guys had explored in Project Mariner in 2024 and some of this other computer use work. So it does feel like there's been a real step change in what the capabilities are. So I'd love to hear you just riff on the research breakthroughs that enabled that and then kind of how people should think about what these agents can and can't do today.
我们知道这将是一个非常重要的模态:行动,对吧,行动并改变状态,比如数字计算机的状态。然后我认为随着你进化并让模型变得更好,你开始意识到,你把模型做得很好,然后你专注于系统,围绕模型构建系统,然后尽可能联合优化系统和模型,等等。所以从什么创造了能力差异或增长的角度来看,主要是关于发布序列。而且在某种意义上,模型能力需要达到一定水平,你才能梦想下一阶段的能力,模型接下来可能做什么。
We knew that was going to be a very important modality: actions, right, acting and changing the state of, let's say, a digital computer. And then I think as you evolve and make the model better, you start realizing, you get the model really good and then you focus on the system, building a system around the model, then optimizing the system and the model jointly as much as you can, and so on and so forth. So in terms of what creates the delta or the increase in capability, it's mostly focused on sequencing releases. And also in some sense, the model capability needs to reach a certain level for you to then be able to dream about what's the next stage of capability, what the model might do next.
是的。我想消费级应用的一个有趣之处在于人们想用它做各种各样的事情。所以我想知道,到目前为止以及你如何看待它随时间演变,模型加系统的工作,它对人们想解决的问题子类别的定制化程度如何,与极其通用的方式相比,比如,你只是在优化一个系统和模型的组合,它几乎适用于你在 Spark 中可能想做的所有事情。
Yeah. And I guess one thing that's so interesting about the consumer footprint is there's just such a broad array of things people want to do with it. So I wonder, to date and how you see this evolving over time, that work of model plus system, how bespoke is it to subcategories of the problems people want to do versus incredibly general and like, hey, you're just optimizing a system and model combination that works across pretty much everything you might want to do in Spark.
总是有一个序列,先专精于感觉可控且已经非常有用的东西。如果你看 Spark,它可以访问所需的信息来帮助你安排和组织你的一天,甚至思考如何解决不同的问题,因为它有非常丰富的上下文。所以围绕你深切关心的事情稍微狭窄地构建系统是有用的。但如果你看机器学习和深度学习的历史,我们构建的组件总是通用的。所以有一个很大的假设,实际上与世界模型的观点有点相关,即联合训练所有东西一定比仅仅专注于一个领域更好。所以即使从建模的角度来看,这也是非常清楚的。但即使从系统的角度来看,一个相当通用的系统,然后根据你如何指示它或与它交互,你当然可以把它放在这样的空间里:嘿,这个用户想这样做,但我有所有这些能力,让我弄清楚该用哪些。在训练时,不一定为此构建,而是构建通用的东西,然后通过模型的智能层和系统的通用性实现专门化。我认为这已经相当明显了。然后有时在实践中,限制或使其更高效仍然有道理去专门化,但从专门到通用,我们看到它一直在发生。即使从架构来看,Transformer 最初是一个机器翻译神经网络,对吧,现在它做所有事情,从全模态到控制你的电脑。所以是的,我认为这是我期望的一步。
There's always a sequencing to specializing to something that feels controllable and that is very useful already. If you look at Spark, it has access to information that it would need to be able to assist you in scheduling and organizing your day and even thinking about how you should tackle different problems because it has this very rich context. So it's useful to build the system slightly more narrowly around something you care deeply about. But if you look at the history of machine learning and deep learning, we always go from the components we're building are general. So there's a big hypothesis, which goes a bit to the world model point actually, that training on everything jointly must be better than just focusing narrowly on just one domain. So even from the modeling perspective that is very clear. But even from the systems perspective, a system that is fairly generic and then based on how you instruct it or you interact with it, you can then of course put it in the space of like, hey, this user wants to do this but I have all these capabilities, let me just figure out which ones to use. Kind of at train time, not necessarily building it for that, but building something generic and then the specialization happens through a layer of intelligence of the model and the generality of the system. I think that's fairly clearly already here. And then maybe sometimes in practice, limiting or making it more efficient still makes sense to specialize, but the specialist to general, we've seen it just keeps happening. Even from architectures, I mean the Transformer was a machine translation neural net, right, and now it does everything from omni to controlling your computer. So yeah, I think that's a step that I expect.
这些年来你一直直言不讳地谈论苦涩的教训。我很好奇,当你展望这个领域时,是否有你认为目前没有遵循它的地方,或者基本上是你看到某种结构或巧妙的脚手架,你认为 Scaling 最终会将其淘汰?
You've been vocal about the bitter lesson over the years. And I'm curious, as you look out at the field, are there places where you think it's not currently being followed, or basically places where you look out and you see kind of structure or clever scaffolding that you think scale is just kind of eventually going to wash out?
是的,我想是的。我的意思是,一个我觉得令人兴奋的领域,已经有了一些已发表的研究,那就是在极限情况下,我们现在通过编码构建的系统,有时是围绕模型的复杂脚手架,多智能体、子智能体委派、非常长时间运行,那个系统本身是一段代码,最终模型本身可以即时编写。所以你可以想象,不仅有一个非常通用的系统,实际上可能没有系统,只有模型能够根据被要求做的事情来编写这些。
Yeah, I think so. I mean, one area that I find exciting, there's some research on this already published, is that in the limit, the system that we build now by coding sometimes a complex scaffold around the model, multi-agent, sub-agent delegation, very long running, that system itself is a piece of code that eventually the model itself could write on the fly. So you could imagine not having just a system that is very general but actually maybe no system and just the model being able to write those depending on what it is being asked to do.
就像几乎是最高 Token 效率、最高质量的子智能体输出集,以及围绕一组问题的任何东西。
Like almost the most token efficient, highest quality output set of sub-agents and whatever it is around a set of problems.
是的,完全正确。
Yeah, exactly.
在过去一年半里,我们看到一个范式转变,即推理模型可以在词元空间中长时间推理。但最终更重要的是应该推理多长时间,根据用户问题的复杂程度加入相应水平的智能会让它更高效。所以围绕这些系统,会有一定程度的自动化来为智能体端的正确任务创建合适的脚手架。每个人都在尝试构建长时间运行的智能体,但要让它们在数百步中保持稳定会遇到各种问题。你认为要进一步提高智能体的可靠性需要什么?
We've seen a paradigm shift in the last year and a half with reasoning models that can reason for a long time in token space. But eventually, what becomes more important is how long to reason, and adding that level of intelligence based on the complexity of the user's query will make it more efficient. So around these systems, there will be some level of automation to create the right scaffold for the right task on the agent side. Everyone is experimenting with building long-running agents, but they run into issues trying to get them stable across hundreds of steps. How do you think about what's required to get to further agentic reliability?
我认为最直接的答案是改进模型周围的脚手架。如果你考虑如何训练神经网络,它是在某种任务或模态分布上训练的。这些都关乎如何预训练或后训练权重。如果有一种新的工作或模态需要长时间运行的系统,并且需要从很长的上下文中学习——我们也在创新,比如 1.5 版本的长上下文突破——那么模型显然也会迎头赶上,满足用户和未来用例的需求。这是研究者的挑战:预测可能实现的事情,然后不仅专注于构建对此鲁棒的系统,还要关注当你把所有上下文和这些疯狂的东西推入时权重会如何反应,而不是仅仅希望从提示中泛化。
I think the answer in the most obvious way is improving both the scaffold around the model. If you think of how you train a neural network, it trains on some distribution of tasks or modalities. All these are about how you pre-train or post-train the weights. If there's a new type of work or modality that requires very long-running systems that need to learn from very long context—which we've also innovated on, like our long context breakthrough in 1.5—then it becomes obvious that the model will also catch up to meeting the users and futuristic use cases. That's the researcher challenge: predicting what can be possible and then focusing not only on building a system robust to that, but also on how the weights would react when you push all the context and all these crazy things, not just hoping for generalization from the prompt.
每个人都在试图弄清楚的一个模式是记忆,以及如何在智能体之间解决这个问题。你对最终会在哪里解决有什么直觉吗?
A pattern everyone's trying to figure out is memory, and how to solve this across agents. Do you have any gut instincts on where that ultimately gets solved?
记忆很迷人。从很早开始,我们可以用几种方式来思考它。我喜欢的一个简单分类是工作记忆——因为我们正在做或谈论的事情而非常活跃的东西——和情景记忆,这是一种检索系统,可能不那么精确,但拥有你关心的所有上下文。计算机也有类似层次,比如 L1、L2 缓存等。对于模型来说,Transformer 的工作记忆是一种强大的机制:我们有成百上千甚至数百万个词元来修改那个记忆,并做惊人的事情,比如证明复杂定理。我看到势头的地方在于,如何巩固在不同交互中或比工作记忆能容纳的更长的交互中发生的事情。我们如何存储这些知识?通过实验,现在的标准名称是“技能”,但它更通用。作为一个智能体,我们可以访问一个记忆系统,那就是计算机本身。你可以把你的想法写进文件,组织成目录,并在与同一个用户多次交互或一次非常长的交互中这样做。目前相当不错的机制——尽管权重还没有跟上——是将这个知识库添加到文件系统中,并带有基本的检索功能。这很强大,但仍有未开发的潜力。许多人称之为持续学习。我希望更好运作的机制是这种文件系统风格,非参数化。它比集成回权重更方便,因为从实践角度看,我们试图大规模服务一个模型。为不同用户服务带有不同记忆的同一个模型会很痛苦。所以我认为我们会看到更好的评估和方法,让模型在交互中积累知识,这可能是范式转变,类似于一年半前的推理。
Memory is fascinating. Since very early days, we can think of it in a few ways. The simpler one I like is working memory—things that are very present because of what we're doing or talking about—and episodic memory, which is a retrieval system that is probably less precise but has all the context you care to remember. Computers have similar levels with cache L1, L2, etc. For models, working memory with transformers is a powerful mechanism: we have hundreds, thousands, millions of tokens to modify that memory and do amazing things like proving complex theorems. Where I see momentum is in consolidating things that happen in different interactions or within a longer interaction than working memory can hold. How do we store that knowledge? Through experiments, the standard name now is 'skills', but it's more general. As an agent, we have access to a memory system which is the computer itself. You can write your thoughts into files, structure them into directories, and do that as you interact with the same user over multiple episodes or a very long episode. The mechanism that is fairly good now—though the weights haven't caught up—is adding this knowledge base into a file system with basic retrieval. That's powerful, but there's still untapped potential. Many call this continual learning. The mechanism I want to work better is this file system style, non-parametric. It's more convenient than integrating back into the weights because from a practical point of view, we try to serve one model at scale. It would be painful to serve one model with different memories to users. So I think we'll see better evaluations and ways for models to accumulate knowledge as they interact, which could be paradigm shifting, similar to reasoning a year and a half ago.
这是否意味着每个人都拥有自己的模型,这些模型有各自独特的文件系统,还是你认为随着时间的推移,人们的模型权重会因他们的行为而不同?
Does that look like everyone having models that then have their own distinct file systems, or do you think over time people have models whose weights look different based on what they've done?
正如我所说,权重不同会是一个挑战。
As I said, having different weights would be a challenge.
是的,难以服务。
Yeah, hard to serve.
是的,会是这样。如果这是最好的方式,我们会找到办法,通过硬件投资来支持更个性化的权重。但至少,你会拥有自己的个人知识库。在 LLM 领域我们已经看到很多这样的例子。然后也许还有另一层知识,对给定模型的所有用户是通用的,你可以访问它来增强模型能力而不改变权重。那会很棒。
Yeah, it would be. If it's the best way, we'll find a way, with hardware investment that allows more personal weights. But at the very least, you will have your own knowledge base that is personal to you. We're already seeing many examples of this in the LLM space. Then perhaps there's another layer of knowledge common to all users for a given model, which you could access to enrich the model's capabilities without touching the weights. That would be awesome.
我觉得持续学习一直是当下的热门话题,每个人都在谈论。你已经看到了一些有趣的例子,一些知名人士从 OpenAI 或其他地方出来说:‘嘿,你可以继续扩展我们现在做的事情。’我认为没有人否认缩放定律的存在,但他们说感觉需要几乎一个新的研究方向才能实现真正的持续学习。也许在持续改进这些核心 LLM 的道路之外追求这一点是有意义的。我很好奇你对整个动态的看法,以及你的反思。
I feel like continual learning has been the topic du jour, and everyone's talking about it. You've seen a few interesting examples now, high-profile examples of folks spinning out of OpenAI or other places and saying, 'Hey, you can keep scaling what we're doing now.' I think no one's denying that those scaling laws are there, but they're saying it feels like you need almost a new research bet to achieve real continual learning. Maybe it makes sense to pursue that outside of the path of continually improving these core LLMs. I'm curious what you make of that whole dynamic and your reflections on that.
我很早就加入了 Google Brain,然后在 2016 年转到了 DeepMind。目前,我认为既存在挑战也存在机遇:你希望研究一些可能在未来三个月内无法进入下一次训练的研究问题,但同时,这不能与 LLM 发展的主流脱节。我们正在改进 Gemini;看到 Flash 超越几个月前的 Pro 非常令人着迷,而且这种情况持续发生。因此,在保持能力前沿的同时(这可能会启用或禁用某些研究),还要为研究提供保护——当然,这不再是多年的事情了;事情发展很快——但将这两者结合起来是构建这些组织的魔力所在。我们所有人都有不同的角度,可以想办法弥合这一差距并识别机会。这需要一些直觉,有时还要急切地引入这些想法,因为感觉这是正确的事情。从研究角度来看,这就是定义这个级别组织的方式。我可以看到从机器人投资到 LLM 的巅峰,再到已经或将要成功的研究。但这很有挑战性;资源是有限的。所以这是一个有趣的权衡,而且并不总是能做好。我认为这是一个迷人的研究角度——不仅仅是哪个想法会进入下一篇论文或模型,而是实际上如何组织整个组织。
I was in Google Brain very early days and then moved to DeepMind in 2016. At the moment, I think there is a challenge and an opportunity: you want to investigate some research questions that might not be viable in the next three months to make it into the next training run, but at the same time, this cannot be very disconnected from the head where the LLMs are moving. We're improving Gemini; it's fascinating to see Flash outperforming Pro from only a few months ago, and that keeps happening. So keeping at the head of capability, which might enable or disable certain research, whilst having the protection for research—and of course, that's not multi-year anymore; things are moving fast—but combining these two is the magic of building these organizations. All of us have different angles and can figure out how to bridge this and identify the opportunity. That's a bit of what it takes to not have full visibility—this is too large of an organization—but have some intuitions and then be able to pull in these ideas eagerly sometimes because it feels like the right thing to do. So that's what defines organizations at that level from a research perspective. I can see from investment in robotics to the peak of the LLMs to research that either has made it or will make it through. But it's challenging; resources are constrained. So it is an interesting trade-off, and not one you always get right. I think it's a fascinating different angle of research—not just what is the idea that will make it to the next paper or into the model, but actually how to organize this whole organization.
对于像你这样角色的人来说,这感觉像是最有趣的问题之一。很难不对今天可以用这些模型推进的许多事情感到兴奋,而且显然有很多事情在发生。即使是像 OpenAI 这样的组织也在‘我们应该去,有很多唾手可得的果实’和‘我们只需要真正搞定代码并赶上 Claude 代码’这样的聚焦时刻之间摇摆。我想知道你是如何看待专注于一件事并让所有人朝那个方向努力与更广泛领域之间的权衡,所有这些都非常有趣。
This feels like one of the most interesting questions for someone in a role like yours. It's hard not to feel excited about so many things you can advance with these models today, and there's obviously so much going on. Even an organization like OpenAI has oscillated between 'we should go, there are so many low-hanging fruits' to a more focusing moment like 'we've just got to really nail code and catch up to Claude code.' I'm wondering how you think about the trade-offs of focusing on one thing and having all rowing toward that versus a broader surface area, all of which are super interesting.
谷歌处于一个独特的位置,原因有几个。首先,我们目前在 Gemini 上确实有很大的覆盖面,它实际上为一切提供动力。但我们的优势是,组织的其他部分完全接受了 LLM 时代。所以他们拿模型去做一些事情。但如果你觉得这不是推进前沿能力的下一步,你可以依赖一个非常好的团队,他们会把模型带到需要去的地方。同时,我们在硬件采购和资本投资方面有稳定性,因为我们在收入流方面非常端到端。所以你可以进一步推动某些研究领域的风险承担,这些领域也需要品味。所以它并不聚焦,但由于谷歌的组织方式,它是可扩展的,你仍然可以投资于创新,这是我们一直以来的核心。如果我看 Brain 和 DeepMind,这两个我曾参与的组织,现在叫 Google DeepMind,我很感激因为我在不同时期都待过,那么我认为创新是我们的 DNA。但与此同时,Gemini 创造了一个聚焦和统一的力量,这做起来非常迷人。我和 Jeff 认识多年,曾一起旅行只是为了好玩,这非常有帮助。所以我认为那段时光非常特别,让中心成为 Gemini 核心建模工作,非常专注于前沿能力,然后有这些输入和输出,是一个相当合理的方式,既能保持专注,又能利用一些可能仍然需要的探索。我的意思是,我们需要世界模型吗?如果我们让它工作,肯定需要。如果不,也许也可以,但把赌注放在正确的位置也是好的。
Google is in a unique place for a couple of reasons. First, we indeed have a lot of surface area in Gemini at the moment, literally powering everything. But we have the advantage that other parts of the organization are completely bought into the LLM era. So they take the model and then they might do something. But if you feel that's not the next way to advance frontier capabilities, you can just rely that there's a very good group that will take the model to where it needs to go. At the same time, we have stability from hardware procurement and investment of capital given we're very end-to-end in terms of revenue streams. So you can probably push a little further the risk-taking for certain research areas which need to be done with taste as well. So it's not focused, but it's scalable because of how Google is organized, and you can still invest in innovation which is at the very core of what we've always done. If I look at Brain and DeepMind, the two organizations I've been part of, now called Google DeepMind, which I appreciate given I've been in both over different periods, then I think there is in our DNA to keep innovating. But at the same time, what Gemini created is a focused and unifying force, which was fascinating to do. It was very helpful that me and Jeff had known each other for many years and had gone on trips together just for fun. So I think that time was very special, and having the center being the Gemini core modeling effort very focused on frontier capability and then having these inputs and outputs is a fairly reasonable way to be focused but also able to leverage a bit of exploration which might still be needed. I mean, do we need world models? If we make it work, it definitely will need it. If we don't, maybe it's okay, but it's good to have the bets as well placed rightfully.
在模型方面,也许换个话题,谈谈 Gemini 模型和前进的道路。我认为你之前说过后训练仍然是一个完全的蓝海。我觉得我们已经看到,在编码和数学方面的后训练和强化学习取得了令人难以置信的进展。我想就在我们开始这个播客的几个小时前,又有一个新的数学问题被解决了。每个人都在试图弄清楚,而我对你的直觉很好奇的是,我们将在哪些领域看到强化学习真正起飞的特征。感觉我们在编码和数学方面处于一个疯狂的指数路径上,我很好奇你的直觉,是什么让其他领域成为好的候选。
On the model side, maybe switching gears to Gemini models and the path forward. I think you called post-training before still kind of a total green field. And I feel like what we've seen, clearly there's incredible progress on post-training in RL in coding and math. I think there was just a new math problem solved hours before we came on this podcast. What everyone's trying to figure out, and I'm curious for your intuitions, is the characteristics of the next set of domains where we'll see RL really take off. It feels like we're on this crazy exponential path on the coding math side, and I'm curious your intuitions on what makes other domains good fits.
是的,这是个好问题。人们必须非常谦虚,因为模型在很多事情上确实非常擅长。所以很难说……好得离谱。
Yeah, it's a good question. One must be quite humble in terms of the models are really good at many things. So it's very hard to say... insanely good.
哦,是的,你知道,这根本行不通,对吧?我的意思是,几乎只是简单的提示和一点巧妙的提示,也许构建正确的系统,在数字世界里,很多令人惊叹的事情,至少在我所谓的数字 AGI 中,是非常令人印象深刻的。所以我认为,当我说后训练是一片蓝海时,这与其说是一种能力远未达到可接受水平,不如说是从机制上审视其他一些利用模仿学习或预训练加后训练的努力,以及在后训练上投入的算力与当前模型使用的相对较小量之间的对比。原因很清楚,但不确定是否容易解决。但即使你拿一个非常狭窄的领域,比如围棋,当你用强化学习下围棋时,你有一个可以下棋的系统。几步之后,那个局面就是独一无二的;你从未见过那种特定配置。所以随着你下棋,环境的复杂性使得训练数据无限免费。你玩得越多,投入强化学习算法的时间越多,获得的知识就越多。这就是我们在游戏强化学习时代看到的。在大型语言模型中,我们受限于数据。什么是不确定复杂性的来源并不清楚。有一些想法,但破解这个配方可能意义重大,至少从算法的优美性来看是这样。知道这在过去如何奏效,现在看到它在其他领域起作用,会更令人满意。那么,它有必要吗?能力还不够吗?这很难说。
Oh yeah, you know, like this doesn't work at all, right? I mean, almost bare prompting and a bit of smart prompting, maybe building the right system, lots of amazing things, at least on the digital world, as I call it, like digital AGI, if you will, are very impressive. So I think when I said that post-training is a green field, that's less about a capability that is far from being acceptable, and more about mechanistically looking at how some other efforts have leveraged imitation learning or pre-training plus post-training, and how much investment there has been compute-wise in post-training versus the relatively smaller amount that today's models use. The reason is clear and not sure it's easy to fix, but the fact that even if you take a very narrow domain like Go, as you play the game of Go with reinforcement learning, you have a system that can play it. A few moves into the game, that scenario is unique; you've never seen that particular configuration. So the environment's complexity as you play makes training data infinite for free. The more you play, the more hours you put into your RL algorithm, the more knowledge you gain. That is what we've seen in the game reinforcement learning era. In LLMs, we are data limited. What is the source of infinite complexity is not so clear. There are some ideas, but cracking that recipe could be big, at least in terms of the beauty of the algorithm. It would be much more satisfying knowing how this has worked in the past to see it work now in other domains. Now, is it needed? Are the capabilities not there? That would be hard to say.
既然你问到了哪些能力,我认为模型所做的最吸引我的能力是我所谓的元能力。它们不是数学或编程。它们更像是智能的特质或属性,以及这些模型能否做到。所以实际上,持续学习或从经验中高效学习的能力就是其中之一。上下文学习,我们以前称之为元学习,随便你怎么叫。这是一种我可以衡量或感受的能力,可能目前还不是很好。例如,当然指令遵循是一种能力,你可以认为它是终极能力,因为如果我要求模型做某事,它要么遵循指令,要么不遵循。但我的意思是,试图观察这些能力,它们更少关乎某个特定领域或垂直领域,而更像是“那是智能行为”。所以学习和适应的能力,而不是成为职业选手或国际数学奥林匹克金牌得主的能力,是我在每次训练新模型并拿到新发布模型时最着迷的。
Since you ask about which capabilities, I think the capabilities in terms of what the models do that are most fascinating to me are what I call meta capabilities. They're not math or coding. They're like the traits or attributes of intelligence, and can these models do it. So actually, the ability to continually learn or learn from experience very efficiently would be one. In-context learning, we used to call it metalearning, whatever. This is a capability that I can sort of measure or feel, and probably it's not super good yet. For example, of course instruction following is a capability that you could argue is the ultimate capability because if I ask a model to do something, it either follows that instruction or doesn't. But I mean, trying to look at these capabilities that are less about one particular domain or vertical and more like 'that is intelligent behavior'. So the ability to learn and adapt, rather than the ability to be a professional player or IMO gold medalist, is what fascinates me the most when I look at new releases and models that we are getting our hands on every time we train a new model.
你有测试这个的常用方法吗?
Do you have a go-to way to test that?
我喜欢游戏。所以我通常会在上下文中定义一个新游戏,对吧?这是一种相当经典的方法。当然,你需要小心,因为如果游戏在权重中,如果其他人把那个游戏放到了互联网上,你就麻烦了。
I like games. So I usually might define a new game in context, right? This is a fairly classic way to do it. Of course, you need to be careful because if the game is in the weights, if anyone else has put that game on the internet, you're in trouble.
是的,但我记得有一个评估。这不完全是我做的方式。
Yes, but I remember there was an eval. This is not exactly how I do it.
实际上,我意识到我要求你谈论这个很无礼,因为这样这个播客就会传出去,然后下一个模型就会知道怎么做了。没问题。
Actually, I realize I'm being rude by asking you to talk about it because then this podcast will be out there and then the next models will know how to do it. No problem.
是的,也许。希望我们不需要破解模型,对吧,除非它被完全转录,我确信它会的。所以也许我们甚至不需要那样做。
Yeah, maybe. Hopefully we don't need to crack models, right, unless it's fully transcribed, which I'm sure it will be. So maybe we don't even need that.
但我真的很喜欢一个评估。我认为这个评估实际上非常古老。它大概是在 2015 年之前,可能更早。这个评估很简单:你给出一份《文明》游戏的说明书,然后你应该能够玩这个游戏。所以我喜欢这种风格的评估。你可以用不同的方式创建它,但这是我喜欢用来测试模型的一个测试,而它们表现并不好,尤其是当游戏变成我刚刚发明的时候。那里的能力是双重的:首先,你能理解指令并遵循它们来玩游戏吗?但还有另一个方面:随着你玩游戏,你学会玩得更好。所以你能在实践中看到这种情况发生吗?这令人印象深刻,但同样,如果你进入一个非常分布外但可能真实且不在训练中的游戏,这个测试对模型来说尤其不容易通过。还有很多其他的测试,但我真的很喜欢这个,它以一种有用的方式引入了游戏。然而你根本不会在游戏上训练。这不像围棋,你只训练围棋。恰恰相反。但从能力的角度来看,我喜欢这种思考。
But I really like an eval. I think this eval is actually very old. It must be like 2015 minus, probably before 2015. The eval was simple: you give the instruction manual for Civilization the game, and then you're meant to be able to play it. So I like that style of eval. You can create this differently, but that's one test that I like to test the models, and they're not that good, especially as the games become something I just invented. The ability there is dual: first, can you understand the instructions and follow them to play the game? But there's another aspect: as you play the game, you learn to play better. So can you see that happening in practice? It's impressive, but again, if you go very out of distribution of a game that could be real but still not in the training, this one in particular is not an easy test for the models to pass. There are many others, but this one I really like, and it brings games in a way that is useful. Yet you will not train on the game at all. It's not about Go where you only train on Go. It's the opposite. But I like this kind of thinking for a capabilities point of view.
感觉已经有很多努力了。游戏是典型的第一类可验证领域,现在在编程和数学上也是如此。我觉得这个领域一个很大的悬而未决的问题是,我们会在多大程度上看到跨强化学习的泛化。感觉有时这些模型在我们进行强化学习的领域上爬山爬得非常好,而你对是否看到这种泛化流向模型的其他方面有更好的洞察。感觉在某些方面,这几乎是一个有趣的“苦涩教训”时刻:在特定领域找到数据,针对这些数据进行强化学习,并在那一件事上改进模型。我很好奇,这是否是对今天发生的事情的公平描述?你看到泛化的迹象了吗?
It feels like there's been a lot of effort. Games were kind of the canonical first example of a verifiable domain, and you've had this with coding now and math. I feel like a big outstanding question in the field is the extent to which we'll see generalization across RL. It feels like sometimes these models hill-climb incredibly well on the domain that we're RLing on, and you'd have a better insight into whether you see that flow through to other aspects of the model. It feels like in some ways it's almost an interesting bitter lesson type moment: find data in a particular domain, RL against that data, and improve the model on that one thing. I'm curious, does that feel like a fair characterization of what's happening today? Have you seen signs of that generalization?
是的,你努力寻找那些能引发深度推理的难题来源,我们从这些推理中看到了泛化。
Yeah, you look hard for sources of hard problems that will induce deep reasoning that we see generalization from.
推理模型主要通过编程和数学进行推理,但你看它们如何推理关于搬家、税务等问题。推理效果很好,很难相信它们是在这类问题上训练的。所以我们确实看到了泛化,我们也在创造性地寻找能引发深度推理和智能体行为的数据。这是近期改进的一部分:找到这些数据源。但局限于可验证性是不令人满意的,因为对于我希望模型做的大部分事情,即使给我再多时间,我也写不出验证器。然而,创造解决方案和评估解决方案之间存在不对称性。如果评估比创造更简单——比如 NP 难问题——这让我有希望,即使没有完全可验证的方式,模型本身也能判断,比如代码是否创造了一个好游戏。这是有趣的研究,我们已经看到了影响。我们做得越多,就能在更多领域训练。问题是,只关注数学和编程是否足以引发智能问题解决的元能力。我不知道;两种可能都有。
So reasoning models reason mostly through coding and math, but you see how they reason about questions like moving and taxes. The reasoning is pretty good, and it's hard to believe they were trained on that kind of question. So we're definitely seeing generalization, and we're creatively trying to get more data that induces deep reasoning and agentic behavior. That's part of the recent improvements: finding those sources. But being limited to verifiability is unsatisfying because for most things I want the model to do, I wouldn't be able to write a verifier even if I had all the time in the world. However, there is an asymmetry between creating and evaluating solutions. If evaluating is simpler than creating—like NP-hard problems—it gives me hope that models themselves will be able to judge even without a fully verifiable way, like whether code creates a beautiful game. That's interesting research, and we're already seeing impact. The more we do that, the more we can train on more domains. The question is whether focusing on math and coding is enough to induce the meta capability of intelligent problem solving. I don't know; it could go either way.
你直觉上倾向于哪一边?
Do you have a gut instinct one way or another?
我想相信你需要在一个广泛的分布上训练,这应该对模型有帮助。但通过预训练获得的泛化能力非常强。所以也许这取决于超人类水平的野心或这些模型能达到的上限。最终,我觉得在机器学习中,尽可能在分布内训练是可取的。所以这是研究人员在未来几个月和几年需要攻克的一个课题。
I want to believe you need to train on a broad distribution, and that should help the model. But it is very strong how much generalization you get possibly through pre-training. So maybe it depends on the level of ambition of superhuman or the upper bound these models can achieve. Ultimately, I feel training as much in distribution as possible seems desirable in machine learning. So that's one quest for researchers to crack in the next few months and years.
我们的听众——创始人或创业者——正在思考的一个问题是,他们应该在模型层工作,还是纯粹在顶层构建应用。有一种趋势是公司在模型之上做自己的强化学习,甚至训练自己的基础模型,比如 Cursor。我好奇你的直觉,什么时候这样做有意义,什么时候没有。
One thing our listeners—founders or building companies—are thinking through is the extent to which they should work at the model layer versus purely building the application on top. There's a trend of companies doing their own RL on top of models, or even training their own base model like Cursor. I'm curious about your intuition on when that makes sense and when it doesn't.
我会告诉人们,评估和数据非常有价值,它们紧密相连。即使你不构建自己的模型——因为早期阶段或缺乏资源——仔细思考如何评估你在做的事情的进展也会非常有价值。它甚至可能成为像我们这样的人采用的标准评估。考虑到我们讨论的后训练以及数据稀缺——不像几年前我们能进行数月的训练——数据的价值是巨大的。所以机会就在那里。同时,在模型之上构建,即使模型能力会不断提升,专注于你真正相信的东西可能会创造机会,让你拥有那个空间、理解它、获得用户和临界质量。如果这是其他人没有关注的东西,那么即使你不做其他事情,专门化产品也有很大价值。
What I would tell folks is the value of evaluations and data—they are very tied to each other. There is huge value there. Even if you don't build your own model because it's early stage or you lack resources, thinking carefully about how to evaluate progress on whatever you're trying to do will be very valuable. It might even become a standard eval that folks like ourselves might adopt. The value of data is immense given what we discussed about post-training and the scarcity of enough data to run months of training like we did a few years ago. So that's where the opportunity is. At the same time, building on top of a model, even though model capabilities will keep moving, focusing on something you truly believe in might create opportunity to have that space at your disposal, understand it, get users, get critical mass. If it's something others aren't focused on, there's a lot of value in specializing the product even if you don't do the other things.
看起来在早期,你专门化产品,在模型之上构建,达到规模,学习评估,然后决定是否后训练一个模型。权衡之处在于,随着这些模型泛化和改进,它们不会像大型实验室那样在广泛的事物上训练。所以你就像在跑步机上:每 2-3 个月,即使你稍微领先于最先进水平,你也必须不断重做。
It seems like in the early days you specialize the product, build on top of models, get to scale, learn the evals, and then figure out whether to post-train a model. The trade-off is that as these models generalize and improve, they won't train across the broad swath of things that large labs do. So you're on a treadmill: every 2-3 months, even if you get slightly ahead of state-of-the-art, you have to keep redoing it.
也许另一个角度:随着这些模型更擅长持续学习或使用复杂知识库,为某个应用构建知识库可能比训练权重更高效。你可以添加独特性,保护你免受那些没有花时间思考如何与当前模型交互的人的影响。这种能力只会变得更好。所以这个角度对于早期玩家可能更具可扩展性。
Perhaps another angle: as these models are more capable of continual learning or using a complex knowledge base, building that knowledge base for a certain application can be more efficient than training weights. There might be uniqueness you can add that protects you from someone who hasn't spent time thinking about how it interacts with current models. That capability will only get better. So that angle might be more scalable for early players.
鉴于许多研究方向都有令人信服的路径,你最不确定如何实现的能力是什么?你还没有看到研究路径,但认为它很重要?
Given the compelling path forward on many research directions, what capability are you least sure how to get to from here? Where you don't yet see the research path but think it's important?
我认为最大的不确定性在于在那些我们无法轻易验证输出的领域实现超人类能力。例如,创造性任务如写小说或设计游戏。我们在数学和代码上有可验证性,但对于许多现实问题,我们缺乏这一点。可扩展监督和弱到强泛化的路径仍不清晰。我们需要弄清楚模型如何自我判断或使用 AI 监督 AI,使其可扩展。这是一个关键的研究挑战。
I think the biggest uncertainty is around achieving superhuman capabilities in domains where we cannot easily verify the output. For example, creative tasks like writing a novel or designing a game. We have verifiability in math and code, but for many real-world problems, we lack that. The path to scalable oversight and weak-to-strong generalization is still unclear. We need to figure out how models can judge themselves or use AI to supervise AI in a way that scales. That's a key research challenge.
我认为研究路径对于不少能力是清晰的。多年来最让我着迷的,尤其是 2016 年加入 DeepMind 后,是元学习,也就是模型学习的能力。既然你从事机器学习,这真是一种美妙的能力。我觉得现在有了一条路径和一些基础,它会持续改进。但另一个可能更不确定的是,这些模型能否真正创新。这一点很重要,因为比如当你尝试在机器学习中提出新想法、实现它们、编码、部署等等,我们正在用这些模型做实验。很多人确实在利用我们现有的所有知识,并带着品味进行创新,这即使对人类来说也很难得。它相当特别,而且说实话有时是随机的。并不是说‘哦,这个人真聪明’。你看,只是有一万个人在尝试,你显然挑中了那个正确的,然后赞美它。所以我认为这种能力对于自我改进之类的事情可能非常重要。然而,显然连评估都很难,而难以评估的事情可能也难以攀登。因此,在任何方面创新的能力,特别是科学方面,我认为还需要更多进展。
I think I see the research path for quite a few capabilities. The one that's fascinated me the most over the years, especially when I joined DeepMind in 2016, is meta-learning or the ability of the models to learn. That's such a beautiful capability since you work on machine learning. I feel like there's a path and some baseline now, and it will keep improving. But perhaps another one that I feel might be a bit more uncertain is whether these models can truly innovate. That part is important because, for instance, when you work on coming up with new ideas in machine learning, implementing them, coding, deploying them, and so on, we're experimenting with these. Many folks are truly taking all the knowledge we have now and innovating with taste, which is hard to come by even for humans. It's fairly special and, to be honest, sometimes random. It's not like 'Oh, this person is so smart.' Look, you just have 10,000 people trying, and you obviously pick the one that was right and then glorify it. So I think that ability is probably quite important for certain things like self-improvement. Yet it's obviously difficult to even evaluate, and when something is hard to evaluate, it probably means it's hard to climb on. So the ability to innovate in any aspect, but specifically on science, for example, is a good one that I think more progress is required.
显然,我觉得 Move 37 是过去世界里的一个典型例子。你最近看到什么最接近这个吗?我想在我们开始录制之前,OpenAI 就谈到了他们刚解决的一个组合几何问题。如果看向机器学习内部,那正是关键。我还没看到模型真正产生出类拔萃的想法,但我确信很快会看到。因为模型在理解自身如何被训练方面有一些洞察和方式,感觉超人类,因为从机制上讲,这些模型能接触到我们无法企及的信息带宽。所以那部分令人印象深刻。但我也希望在想法层面看到同样程度的惊艳。而机器学习是我能更准确评估的明显领域。所以,还有更多工作要做。
I mean, obviously I feel like Move 37 was a canonical example of this in the previous world. Is there anything you've seen recently that feels closest to this? I think even before we started recording, OpenAI talked about this combinatorial geometry problem they just solved. If I look inward to machine learning, that's kind of the point. I don't think I've seen truly outstanding ideas that a model has generated yet, but I am sure I will very soon. Because there are some insights and ways in which the models understand how a model is being trained that feel superhuman, because mechanistically these models have access to a bandwidth of information we don't. So maybe that part has been impressive. But I would like to see at the idea level as well that level of impressiveness. And machine learning is the obvious thing I can evaluate more accurately. So yeah, more to do.
是的,你怎么看待当我们达到对机器学习研究有真正洞察、进入递归自我改进的世界时?我很好奇你如何思考那会是什么样子,甚至一些基本问题,比如苦涩的教训是否仍然成立?或者当我们进入那个世界时会发生什么?我很想听听你的想法。
Yeah, how do you reason about when we get to this level of genuine insights into machine learning research and this world of recursive self-improvement? I'm curious how you reason about what that even looks like over time, and even just basic questions like, does the bitter lesson still hold? Or what happens when we get into that world? I'd love to hear you just riff on that.
会有一定效率水平得到提升。所以有一个层面,你作为研究员或工程师使用这些工具来提高自己的生产力。我们已经看到很多了。与领域前沿的人交谈总是令人印象深刻,他们总是说,数字各不相同,但生产力整体上有相当大的百分比提升。所以我认为这已经在发生,而且显然非常强大。但这个过程能持续多久会有一些近乎物理的限制,因为模型需要训练,有能源、硬件的限制。所以我非常渴望看到哪些问题可以更自动化、增强,并更自主地完成。但与此同时,事情发生的速度可能会有自然限制,也有自然的上限。某些事情,一年多前有人对我提到一个观点,我现在深有体会:当模型写英文比你好时,那可能太好了,不应该那么好。我想,好吧,这是一个有趣的领悟,即使你可以改进那个能力,而且可能没有天花板或天花板还很远,我们甚至可能不需要看到那个天花板。所以整个系统的性能已经非常令人印象深刻,在某些情况下可能有明显的上限。但我觉得,模型及其训练方式的物理限制,即使我们确切知道配方,可以快速迭代并训练下一代模型,也存在一些加速,但有一些上限和速率限制仍然是相当根本的。
There's a certain efficiency level that probably will be enhanced. So there's a level in which you as the researcher or engineer use these tools to enhance your own productivity. We've seen that a lot. It's always impressive to talk to someone at the cutting edge of their field, and they're always like, yeah, the numbers always vary, but some pretty large percent improvement in productivity across the board. So I think that one is already happening and obviously very powerful. But there's going to be certain almost physical limitations to how much this process can keep going, because models need to be trained, there's energy, hardware limitations. So I definitely am very keen to see what kinds of problems that are to be more automated and enhanced can be done more autonomously. But at the same time, there's probably going to be a natural limit to the speed at which things can happen and also a natural upper bound. Certain things, I mean, that was already more than a year ago, someone reflected something on me which I now feel very much: at the point a model writes English better than you, that's maybe too good and it shouldn't be that good. And I'm like, okay, that's an interesting realization that even if you could improve that capability, and if maybe there's no ceiling or the ceiling is still far away, it might not even be that we need to see that ceiling. So there's the performance of the whole system overall, which is very impressive already, and there might be upper bounds obvious in some cases. But yeah, I think that the physical limits on the models and how you train them, even if we knew exactly the recipe, we could iterate very quickly and train the next generation models. There is some acceleration, but there's some upper bounds and rate limits that are still fairly fundamental.
嗯,我总喜欢以快问快答环节结束访谈,把没来得及问的广泛问题都塞进去。那么首先,我很好奇过去一年里你对 AI 的哪个看法改变了?
Well, I always like to end my interviews with a quickfire round where I basically just stuff in all the broad questions that I haven't had time to fit in elsewhere. So maybe to kick things off, I'm curious what's one thing you've changed your mind on in AI in the last year?
我改变的看法是:尽管我原本相信在广泛分布上训练可能会增强模型,但在数学或编码等难度极高的狭窄点上训练却产生了这种泛化。我认为这并非我预料到会如此有效。
What I have changed my mind. I think the fact that even though I want to believe that training on a broad distribution is probably going to enhance the model, training on narrow kind of points of great difficulty like maths or coding creates this generalization. I think that is not something I quite predicted to work as well as it did.
我认为 Demis 说过我们正处于奇点的山脚,AGI 可能在几年内到来。你也有同感吗?
I think Demis said that we're at the foothills of the singularity and AGI could come in the next few years. Do you feel similarly?
我有同感,而且我要多说一点。即使是接近这些模型和神经网络的业内人士,如果七年前——我用一个明显在 LLM 所有事情发生之前的时间——如果七年前我要用我们现在的模型做实验,我会宣称这是 AGI 吗?我可能会说是的。这是一个不断变化的定义。进步非常惊人。所以我认为正因为现在我们更接近它,对我们正在构建的东西更加雄心勃勃是件好事。但同样,基于不同的定义,或者甚至仅仅几年前我们对 AGI 可能有的期望。
I feel similarly, and I'll say more. Even with someone in the field close to these models and neural nets in general, if seven years ago—I'm using a time that is clearly pre all that happened in LLMs—if seven years ago I had to experiment with a model that we have currently, would I have declared this is AGI? I would say probably yes. It's an ever-moving definition. The progress is very impressive. So I think just because now we're seeing it closer, it is a good thing to be more ambitious about what it is that we're building. But again, based on different definitions or perhaps the expectations we might have had about what AGI meant even only a few years ago.
我会说在某种程度上,AGI 已经来了,对吧?我的意思是,我不认为它是以我想要的方式出现的,但它已经很接近了。也许模型真正从经验中学习的能力是我认为缺失的。但每个人都有自己的测试或偏见,认为模型仍然存在能力差距。我们会达到那个目标,然后我们又会移动标杆,找到其他理由。
I would say in some way AGI is here, right? So I mean all I'm saying is like I don't think it is here in the way I want to see it, but it is fairly close. And maybe this ability for the models to truly learn from experience is what is missing in my mind. But everyone will have their own kind of test or bias I guess on what the models still feel like capability gaps exist. And we'll get there and then we'll move the goalpost again and have some other reason.
我认为你们的一个巨大优势是,你们对自己构建的模型非常看好。你们有自己的硬件。我想我的很多听众心里都有一个疑问,所以我来问一下:你们做的一件事让很多人好奇,就是把你们的一些算力卖给了 Anthropic,对吧?Twitter 上一直有一种说法:如果你们对模型和研究这么看好,为什么不自己保留所有算力?所以,我相信我们的听众很想听听你的看法。
I think one huge advantage that you all have is you know certainly incredibly bullish on the models you're building. You have your own hardware. And I think a question a bunch of my listeners will have in the back of their heads, so I'll ask it: I think one thing that you've done that a bunch of people were curious to better understand was taking some of the compute you have and selling it to Anthropic, right? And I think there's been this narrative on Twitter of like, well, if you were so bullish on models and the research, why not just keep all the compute yourself? And so, I'm sure our listeners would just love to hear your perspective on that.
是的。如何在我们内部投资算力,对吧?算力用于服务,我们训练小模型,甚至更小的模型,然后尝试训练前沿模型。我认为这是一个需要平衡的精细方程。总的来说,思考 Alphabet 的一种方式是,有些东西能创造收入和经济效益,然后你可以再投资。所以不仅仅是贪婪地考虑我们现在该做什么,然后把所有东西都抓在手里。我经常思考的策略是多管齐下。虽然我们当然看好技术的进步,但也要考虑收入流等等。我认为硬件是非常重要的资产。而且可能有一个权衡,你不必全部使用,而是战略性地使用它来创造再投资。我认为这似乎是合理的。背后的计算很复杂,所以我不打算深入具体理由,但总的来说,这是一个战略选择,要考虑不同层次的投资和时间线。
Yeah. How to invest compute even within ourselves, right? The compute is used to serve, we train small models, even smaller models, then trying to train frontier models. I think this is all a fine equation to balance. And I think in general, one way to think about Alphabet is there are things that create revenue and economic impact that then you can reinvest. So it's not just being greedy about what we should do now and take all these things together. I think the strategy, which I often think about, is just multipronged. And I think the timelines, although we are bullish of course on the technology advancing, you just think of revenue streams and so on. And I think hardware is a very important asset. And I think there's probably a trade-off in which you don't use it all but use it strategically to create reinvestment basically. And I think that's what seems to make sense. The calculations behind these are complex, so I'm not going to enter into exactly the rationale, but I think in general it's just a strategic choice to have different levels of investment and timelines in mind.
你的位置特别有趣,因为你是唯一拥有自己尖端、最先进芯片的前沿模型提供商。这种合作实际上是什么样的?因为这是一个非常独特的动作,对吧?显然,Nvidia 与其他实验室密切合作,但他们不在同一家公司。那么,当合作顺利时,是什么样子?
What's so interesting about your position is you are like the only frontier model provider with your own cutting-edge, state-of-the-art chip. What does that collaboration actually look like? Because it's such a unique motion, right? I mean, obviously Nvidia works closely with other labs, but they're not sitting under the same company. And so, what does that look like when it works really well?
正如我之前解释的,我回顾了几个时刻。那是早期。甚至谷歌内部的深度学习也还需要证明。我记得大概是 2013 年或 2014 年,我们几个人——我想是我、杰夫·辛顿、杰夫·迪恩和伊利亚——在一个房间里决定服务器应该有什么配置。多少?当时我们有一些 CPU、一些 GPU,你根据对研究方向的了解来猜测模型的发展方向。你真的可以产生那种影响,但当然有延迟的回报,因为这只是一个投资,几个月或几年后才会在数据中心实现。所以我一直觉得这很神奇。显然很难回答这个问题。我们试图预测研究会发生什么,早期更难。但我认为这是一个非常特权的地位,能够真正影响。我们确实做到了,尤其是和杰夫一起,他基本上在谷歌存在的过程中一直在思考基础设施。然后思考这些模型如何发展,以及这些投资因为有一定的延迟而很有趣。在同一屋檐下,看到我们所看到的,真的很有帮助。再次,我在早期混乱的日子里就看到了这一点,它一直在发生并变得更好。当然,不确定性在某种程度上减少了,这让工作更容易,但仍然是一个迷人的选择,对公司的命运等有深远的影响。
As I was explaining before, I reflected on several moments. This was early days. Even deep learning internally at Google had to be still proven out. I remember it must have been 2013 maybe 2014 where a bunch of us—I think it was me, Jeff Hinton, Jeff Dean, and Ilya—were in a room trying to decide what the servers should have. How many? At the time we had some CPUs, some GPUs, and you're trying to make a guess based on what you know about the research where the models are going. And you can literally have that impact, but of course there's delayed rewards because this is just an investment and only a few months or years later it will materialize in data centers. So I've been sort of... and I thought that was amazing. Obviously hard to answer the question. I think we tried to predict what's going to happen in research, and in the early days that was even harder. But I think it's a very privileged position to be able to really influence. And we certainly do that, especially with Jeff who's been thinking about infrastructure quite a bit for basically the existence of Google. It's very interesting to then think about how these models are going this way, and then these investments because they have certain latency. Being under the same roof and seeing what we see really helps. And again, I've seen it in the very scrappy early days and it keeps happening and getting better. And of course uncertainty in some way reduces, which makes the job easier, but still a fascinating choice that has deep consequences on the fate of the company, etc.
嗯,这是一次精彩的对话。我觉得我可以和你聊很久,但那样会拖延我们通往 AGI 的道路。所以,也许我想把最后一句话留给你。你有什么想和听众分享的,或者有什么研究想推荐给他们,IO 上有什么吗?请讲。
Well, this has been a fascinating conversation. I feel like I could talk to you for a long time, but I'd be delaying our path toward AGI. And so, maybe I just want to make sure to leave the last word to you. Anything you'd like to share with our listeners or research you'd like to point them to, anything in IO? The floor is yours.
我认为现在是 AI 领域一切都很迷人的时候。所以,如果你是用户,使用模型。如果你是构建者,用模型来构建。无论你做什么,即使你认为与 AI 毫无关系。所以,请玩玩这些模型。它们很神奇,而且只会变得更好。
I think it's a fascinating time for anything in AI. So, if you're a user, use the models. If you're a builder, use the models to build. Anything you do, even if you think there's no remote connection to AI. So, please, play with these models. They're amazing and they will only get better.
太棒了。非常感谢。这是一次很棒的对话。我是 Jacob Efron,这里是 Unsupervised Learning,一个播客,我在这里与 AI 领域最聪明的人交谈,问他们关于模型发展以及对世界商业影响的大量问题。我希望这很明显,我做这件事非常开心。这是我除了在 Redpoint 做投资人的日常工作之外的夜晚和周末项目。但我们能请到这些了不起的嘉宾,真的来自于像你这样的听众订阅播客、与朋友分享。这最终是让这一切运作起来的原因。所以,请考虑这样做。非常感谢你的支持和收听。我们下期再见。
Awesome. Well, thank you so much. This has been an awesome conversation. I'm Jacob Efron and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so, please consider doing that. And thank you so much for your support and listening. We'll see you next episode.