Demis Hassabis on AGI, AI Risks, and DeepMind's Competitive Edge
打开互动全文版(中英对照 + 朗读 + 问答)→Demis Hassabis 探讨 AGI 的路径、AI 的风险(包括网络和生物威胁),以及 DeepMind 如何在激烈竞争的市场中保持人才优势。
Demis Hassabis discusses the path to AGI, the risks of AI including cyber and bio threats, and how DeepMind maintains its talent edge in a fiercely competitive market.
很高兴见到你。感觉我们每次聊天不是在特别冷就是在特别热的地方。
It's great to see you. Seems like we always talk where it's very cold or very hot.
没错。
Exactly.
特别热,希望这次我们不会聊到出汗。
Very hot, so let's see if we can get through this without sweating.
嗯。
Yeah.
Demis,现在大家都在为 AI 抓狂。华盛顿特区正在封禁 AI 模型。很多担忧都集中在那些能写软件、找计算机漏洞的文本模型上。我想知道,你是否和很多人一样,认为通往 AGI 的道路要经过像 Mythos 这样可能很快就能自我改进的模型,还是说仍然需要像你在 Gemini 做的多模态方法?
Demis, everyone right now is freaking out about AI. They're banning AI models in DC. A lot of the concern is about these text-based models that can create software, find vulnerabilities in computers. I'm wondering if you think, as many do, that the path to AGI runs through these models like Mythos that may soon be self-improving, or do you think it still requires a multimodal approach like what you're working on at Gemini?
嗯,第一个问题里有很多可以展开的。首先,关于我们在 Cyber 和 Mythos 上看到的情况,我很久以来就一直公开表示,随着我们越来越接近 AGI——我认为我们现在正处于这个临界点——我说过类似我们处于奇点山麓这样的话,我们需要一种更系统的方法。当然,前方有巨大的机遇:治愈所有疾病、发现新能源。这些都是我整个职业生涯致力于 AI 的原因,但也存在风险。Cyber 是其中之一,但还会有更严重的事情。那只是对人类的一个警告信号,我希望我们认真对待。未来几年可能会出现生物、核能等其他类型的风险。我们必须为此做好准备,我认为我们需要一种更系统的方式来处理这些问题,可能是一个国际性的标准机构,来帮助测试最新的前沿系统,确保它们足够稳健,防护措施足够充分。这是其中一方面。在 AGI 的技术路径方面,我们一直拥有广泛而深入的研究团队。这就是为什么在过去十年里,我认为支撑现代 AI 产业的重大突破中,可能有 90%以上来自 Google Brain 或 DeepMind(当时作为独立研究机构),现在则合并为 Google DeepMind。从支撑所有大语言模型的 Transformer,到 AlphaGo 以及我们当时开创的所有强化学习工作。所以我们的方法一直是押注多个方向,并尽可能大力推进。显然,我们有 Scaling(规模扩张)工作、我们自己的多模态基础模型 Gemini。我们在编码方面大力投入,同时还有多模态生成式媒体模型如 Omni 和 VEO。我们认为让这些模型理解我们周围的世界和上下文很重要。最终,要拥有一个完整的 AGI 系统,你需要能够理解周围的物理世界。对于机器人技术成为现实以及智能眼镜上的助手等应用,这绝对是必需的,我认为这两个是非常有趣的应用方向。
Well, there's a lot to unpack in that first question. First of all, with the things we're seeing with Cyber and Mythos, I've been pretty vocal for a long while that as we get closer to AGI—and I think we're on the cusp of that now—I've said statements like we're in the foothills of the singularity, that we need a more systematic approach. Of course, there are amazing opportunities ahead: solving all disease, finding new energy sources. All of these are reasons I've worked on AI my whole career, but there are also risks. Cyber is one, but there are going to be even more serious things. That's just a warning shot for humanity, and I hope we take it seriously. There will be bio, nuclear, and other kinds of risks coming down the line, maybe in the next couple of years. We've got to get ready for that, and I think we need a more systematic way to deal with these issues, maybe a standards body that ideally would be international as well, to help test the latest frontier systems and ensure they're robust and the guardrails are sufficient. That's on the one hand. In terms of the technical approaches to AGI, we've always had a broad and deep research bench. That's why over the last decade, I think maybe 90% or more of the big breakthroughs in AI that underpin the modern AI industry came from Google Brain or DeepMind as separate entities, and now together as Google DeepMind. From Transformers that underpin all large language models to AlphaGo and all the reinforcement learning pioneering we did back then. So our approach has always been to bet on multiple things and push them as hard as possible. Obviously, we have our scaling work, our own multimodal foundation models Gemini. We're pushing hard on coding, but also we have our multimodal generative media models like Omni and VEO. We think it's important to give these models understanding of the world around us, the context around us. In the end, to have a full AGI system, you need to be able to understand the physical world around you. You definitely need that for things like robotics to become a reality and assistants on smart glasses, which I think are two very interesting applications.
我就当这是否定了。谢谢。那么,当你创立 DeepMind 时,你处于前沿的最远端。当你加入谷歌时,DeepMind 和谷歌似乎几乎囊括了 AI 领域所有顶尖人才。现在你至少有三大前沿竞争对手都在争夺顶尖人才。我想知道,你认为今天的 DeepMind 是否仍然拥有赢得 AGI 竞赛所需的人才?
I'm going to take that as a no. Thank you. So, when you started DeepMind, you were so far in the frontier. And when you joined Google, it seemed like between DeepMind and Google, almost all the major talent in AI was under one corporate roof. Now you have at least three major competitors on the frontier who are all vying for the top minds. I'm wondering, do you think that DeepMind today still has the right talent to win the race to AGI?
是的,我认为所有领先实验室之间人才流动很大,我们赢得了我们应得的顶尖人才份额。但我要说的是,我们拥有迄今为止所有领先实验室中最大、最广泛的研究团队。我们持续推出绝对前沿的工作,无论是在基础模型上,还是在最终会融入基础模型的其他模型上,比如我们的 Omni 和 VEO 模型。但现在市场竞争异常激烈,可能是科技行业有史以来最激烈的。我认为这是不可避免的。回想起来,我们在 2010 年创立 DeepMind 时,没有人从事 AI 工作。工业界肯定没有,甚至在学术界,这都被认为是职业自杀。人们说:‘当然我们知道 AI 行不通。我们在 90 年代试过,像 MIT 这样的地方,那是死胡同。’那是主流观点。但我们一小群人觉得,有了正确的想法,使用学习系统、强化学习,并押注神经网络,可以取得快速进展。我们最终是对的,但这也意味着整个世界在最近几年才意识到 AI 的潜力,世界上每一家重要的公司都会参与进来。
Yeah, I think there's a lot of talent movement between all the leading labs, and we win our fair share of the top talent. But what I would say is that we have by far the biggest and broadest research bench of any of the leading labs out there. We continue to put out absolute frontier work, whether that's on foundation models or other models that will eventually feed into the foundation models, like our Omni and VEO models. But it's a ferociously competitive market out there right now, probably the most ferociously competitive there's ever been in the tech industry. I think that was inevitable. When I look back, we started this back in 2010 when I started DeepMind and nobody was working on AI. Definitely not in industry, but even in academia, it was thought to be career suicide. People said, 'Of course we know AI doesn't work. We tried it in the 90s, places like MIT, and it was a dead end.' That was the prevailing view. But we, a small band of us, felt that with the right ideas, using learning systems, reinforcement learning, and betting on neural networks, a lot of fast progress could be made. We were right in the end, but it also meant that now the whole world in the last few years has woken up to the potential of AI, and every important company in the world is going to get involved.
嗯。我们现在在戛纳,一个广告大会。这里有很多极具创造力的人,我相信他们中的许多人,甚至台下的观众,都在用你们的视频创作工具制作广告,做创意艺术方面的其他事情。现在用这些工具能做哪些一年前做不到的事?
Yeah. So, we're here at Cannes, at an advertising conference. There are a lot of people here who are extremely creative, and I'm sure many of them, even in the audience, are using your video creation tools to create ads and do other things in the creative arts. What can you do with these tools now that you couldn't do a year ago?
嗯,这些工具和底层模型每个月都在大幅改进。一年前,我们工具(比如新的 Omni 模型和用于图像的 Nano banana)最大的变化是能够实时编辑生成模型的输出。这对创作者来说变得极其有用。在创作过程中,你生成第一个想法、第一个概念,但你喜欢其中一部分而不喜欢其他部分。你不想重新生成整个东西——那是一年前的情况。你希望能够用自然语言描述,理想情况下就像你对设计师说:‘好的,这部分保持不变,但把这个改成别的。’然后可能迭代数百次,直到得到你想要的最终精炼版本。所以,这种细粒度的控制是过去一年的一大变化,此外还有持续的质量提升。
Well, every month these tools and the underlying models are improving massively. A year ago, the biggest change I would say with our tools like our new Omni model and things like Nano banana for images is the ability to live edit the output of the generative models. That has become extremely useful for creators. In the creative process, you generate the first idea, the first concept, but you like some of it but not other parts. You don't want to have to regenerate the whole thing, which is where we were a year ago. You want to be able to describe in natural language, ideally as you would to a designer, 'Okay, keep that part the same but change this to something else.' And then iterate that maybe hundreds of times until you get to the final polished version you want. So, that kind of fine-grain control has been a big change over the last year, as well as just general relentless quality improvements.
我的意思是,即使在广告行业内部也有争议,人们试图弄清楚他们到底有没有使用 AI?这是 100%人类创作的吗,你是否应该披露这一点等等。你认为这种讨论是暂时的,因为我们还没有适应 AI 将如何改变创造力,还是说这种讨论会一直存在?
There I mean, there's also you know, even controversies within the ad industry where people try to figure out are you are they using AI at all? Is this 100% human created and you should disclose this, etc. Do you think that's a conversation that's sort of temporary because we haven't adjusted yet to how AI's going to change creativity or do you think that that's here to stay that there will always be that that conversation?
嗯,这里有两个不同的部分。首先,我们确实需要处理错误信息和深度伪造。这是我们三四年前刚开始构建这些生成模型时就意识到的事情。我们预见到,未来这些系统会变得非常出色,最终几乎可以达到照片级真实感。因此,我们需要一个系统,一个数字水印系统,我们创建了它,叫做 SynthID,它非常稳健,几乎无法破解,并且可以无形地嵌入到图像中。这样任何人——公民、记者或政府——都可以检测出那张图像是否由 AI 生成。我们所有的模型,无论是生成音乐、图像还是视频,都内置了 SynthID。我们还将其开源,供行业其他公司使用。现在很多同行都采用了这个标准,比如 OpenAI、英伟达和其他大公司。所以我希望最终这能成为一项规定:如果你生成媒体内容,就应该附带来源检测。这显然也有助于处理版权和知识产权等问题。至于是否应该披露你在工作中使用了 AI,我不太确定。我认为这可能只是我们当前所处的时代特征:以前我们用 Photoshop 或其他工具,现在这是一个更先进的工具,但它只是你创造力的一个工具。我不认为需要像你所说的那样披露,除非你需要知道输出是合成生成的。
Well, there's two different parts here. For certain, we need to deal with misinformation and deep fakes. So, that is something we were cognizant of years ago when we first started down building these generative models 3 4 years ago. We foresaw that we would be in a world where these systems would be really good. Obviously, that's what we're planning to do and they would eventually be almost photorealistic. And so therefore we would need a system, digital watermarking system, which we created called SynthID, that was robust, sort of unhackable, and would be embedded imperceptibly in the image. So that anyone, a citizen, a journalist, or government could go and detect whether that image was generated by an AI or not. And all of the things that we've all of our models that generate anything from music to images to videos come with SynthID embedded in it. And we've also open-sourced it and given it to the rest of the industry to use. So a lot of our industry colleagues have now adopted that standard. OpenAI, Nvidia, and many other big ones. So I hope eventually, I think that should become almost a regulation really of like if you're creating generative media, then it should come with provenance detection. And obviously that will also help with things like right holders and IP rights, too. So that can all be sort of connected together. As to whether it should be disclosed if you use AI for a piece of the work or that you're doing, I'm not sure. I think that might be just an era we're in where, okay, we're using Photoshop or some other tool before, and now this is a more advanced tool, but it's just a tool for your own creativity. I just I'm not sure that needs to be disclosed in the sense that you're talking about, other than you should know that the output was synthetically generated.
回顾你的职业生涯,创造力是一条主线。你最初制作电子游戏,后来作为神经科学家研究大脑中创造力的本质。甚至 AlphaFold,我认为也可以看作是一种非常创造性的科学方法。现在我们有了所有这些工具。有些人会说:“这会让我们的创造力下降。我们现在只是让模型去做你多年来辛苦工作的事情。”那么,你怎么看?它会如何改变创造力?
Do you If you look at your career, it creativity is this throughline. I mean, you started out creating video games. You studied this as a neuroscientist, the nature of creativity in the brain. You're, you know, even AlphaFold, I think you could think of as a as a very creative approach to science, right? Um, now we have all these tools. There are people who would say, "Well, this is going to make us less creative. This is, a you know, we're now just asking a model to do what you slaved for years doing as you know, in your in your career. So, how do you think how do you think about it? How's it going to change creativity?
它肯定会改变创造力,但我看到的是双重变化。一方面,它正在使一些创作工具民主化。更多人能够相对快速、轻松地尝试他们的想法。但这也有两面性,因为它也会产生大量可能没有太多创意价值的东西。但这也意味着更多人能够进入这些行业。我认为准入门槛降低了,把关更少了。这可能会让新的创作者,无论在世界何处,都能通过使用这些工具找到成为专业创作者的路径。另一方面,在专业领域,我们与许多专业导演和出色的合作者合作。我们与他们交流,设计工具来增强和赋能他们的创作过程。我认为这将是不可思议的。他们可以比以前多做 10 倍的事情,尝试更多想法,更快迭代。他们一生中的想法远比他们能产出的多。这些工具让他们能够以相对低廉和快速的方式尝试各种东西。所以对于专业人士来说,他们能够更快地迭代出更酷的作品。但就像任何新工具一样,如果使用不当,比如懒惰地使用,就会削弱创作过程。但如果以创新的方式使用,它应该会增强创作过程。我认为创意产业需要一段时间才能找到最佳使用方式。我和游戏设计师朋友们聊过,游戏行业对这些工具非常兴奋,但至少在我最熟悉的游戏行业,我们还没有找到更深层次的使用方式。不过现在还很早,游戏行业目前只是在用这些工具做显而易见的事情,比如创建一些资产和图形。但它能否改变游戏的性质,引入全新的游戏类型?我认为这是可能的。就像 90 年代我刚开始进入游戏行业时,图形和 AI 首次出现在电脑游戏中,它们让我们能够创造全新的游戏类型。我希望这些新工具也能激发这样的变革。
So, definitely going to change it, but I think what I'm seeing is a two-fold change. One is it's democratizing some of the creative tools. So, more people can try this and try out their ideas for a relatively quickly and relatively easily. That's double-edged though, because it also produces a lot more things that maybe are not very creatively valuable. But, also it means a lot more people can break into those industries. I think there's a lower bar there's a sort of lower bar to entry. There's less gatekeeping. So, that will probably mean new creators, professional creators find a path wherever they are in the world through using these tools. And then the other thing I'm seeing is on the professional side and we work with many professional directors and amazing collaborators. And we try and talk to them to design our tools to help enhance and empower their creative process and and and help them with their creative process. I think it's going be incredible. They can sort of do 10x more things than they used to be able to do, try out more more stream ideas, you know, iterate faster through they all have way more ideas than they can ever produce in their lifetimes. And so, this allows these tools allow them to try out things in relatively inexpensive and quick ways. So, I think for the professional they're going to be able to iterate their way to much cooler things way more quickly. So, but just like with any new tool, if you use it in the wrong way, the internet's the same, computers are the same. If you use it in a lazy way, it kind of takes away from the creative process. But, if you use it in a in a innovative way, it should add to the creative process. And I think it's going to take a while for the creative industries to figure out the best way to use these things. Um and I talked to my game designer friends, uh you know, in the games industry, they're very excited about these tools, but I still I would say at least the games industry, which is the the creative industry I know best, that we're we're still yet to figure out the the the any deep ways, the deeper ways of using this. But it's very early, you know, so we're the games industry using it for obvious things like to create some assets and um uh some graphics and things like that. But um can it change the nature of games and introduce whole new genres of games? That's what I think could be possible. Like it was in the '90s when I was started out in the games industry when graphics and AI first came on the scene for computer games. Um and it allowed us to make whole new types of game genres. That's what I hope to see uh these new tools spark.
你认为,有一种批评说这些模型是在人类输出上训练的,对吗?是否应该有一种可审计性,让人们看到,哦,这个输出部分使用了我的创作,我应该得到补偿。你认为应该这样吗?
Do you Do you think this, you know, there's a criticism that you know, these were trained on on the outputs of humans, right? Um should there be some sort of like auditability that lets people see, oh, this output used partially, you know, one of my creations and I should get compensated for that. Do you Do you think that should happen?
嗯,也许需要一种新的经济模式,我认为科技行业和创意行业需要共同努力。就像流媒体时代那样,它改变了音乐,YouTube 推出了内容 ID,YouTube、Spotify 等公司提出了新的、非常稳健的商业模式。所以我认为这可能是必要的。但正如创意行业的每个人所知,具体归因非常困难,比如这个输出 1%来自这个、5%来自那个、10%来自另一个。要客观地达成一致是很困难的。
Well, maybe a new economic model is needed and I think that the tech industry and the creative industry together need to work together. And that's what happened with streaming, which changed, you know, music and things like YouTube and then content ID and and YouTube specifically and Spotify and other companies like that came up with new really robust business models. So, I think that that will probably be needed. But it's very difficult, as everyone in the creative industries know, like to specifically attribute this is 1% this and 5% this and 10% that. It's going to be difficult to kind of agree on objectively like what that is.
作为人类创作者,我们创造的东西是我们所有经历、所学知识以及接触到的其他艺术形式、其他创作者及其作品的输出。然后我们将这一切与自己的创造力混合,产生新的事物。所以从某种意义上说,这一直是创作过程,但我们得看看;可能最终需要新的商业模式。
And as human creators, the things we create are the output of all the experiences we've had, the things we've learned, and what we've exposed ourselves to—other art forms, other creators, and their creations. Then we mix all of that with our own creativity to generate new things. So in some sense, that's always been the creative process, but we'll have to see; probably new business models will eventually be needed.
我经常听人说:‘我接受 AI 用于科学治病,比如你在 Isomorphic Labs 做的那些事。但我不喜欢它复刻音乐、电影甚至广告公司的工作。’但我想知道,在跨学科能力方面,尤其是 AI,是否有一个观点值得探讨。一方面,你可能在 Isomorphic 研究虚拟细胞——我们现在还做不到,但也许有一天我们能实时看到细胞如何运作。另一方面,你在创建这些世界级的视频模型,它们未来可能分析那个虚拟细胞,并帮助治病或创造新疗法。你看到这种情况会发生吗?帮我们理解一下。
I hear this a lot from people: 'I'm okay with AI being used for science to cure disease, for instance, the stuff you're doing at Isomorphic Labs. But I don't like the fact that it's recreating music or the work that a filmmaker might have done, or maybe even an ad agency.' But I wonder if there's a point to be made about cross-discipline capabilities, especially when it comes to AI. On one hand, you might be working at Isomorphic on a virtual cell—which we can't do today, but maybe one day we'll be able to see how a cell works in real time. And on the other hand, you're creating these world-class video models that may one day analyze that virtual cell and potentially help cure disease or create new therapies. Do you see that happening down the road? Help us understand.
AGI 这个术语背后的整个论点,以及我们 DeepMind 从一开始的原始目标,就是创建一个通用智能系统,它可以从几乎任何输入中学习,生成有用的见解或发现有用的模式,然后以几乎任何方式输出。这显然是人类思维的工作方式,看看我们用狩猎采集者的大脑创造的现代文明——想想这是怎么发生的,简直不可思议。这就是我们所知的通用智能。这也是我们从 DeepMind 成立之初就一直专注的,现在整个 AI 领域也是如此:系统是通用的,并且会学习,而不是硬编码或硬编程答案。现在回想起来很有趣,但这就是 AI 领域在前五六十年里所做的。比如深蓝这样的国际象棋程序等等。这意味着其中一些东西是不可分割的。如果你想要一个完全通用的系统,能够理解周围的世界,分析科学论文或科学数据,包括视觉数据,比如细胞、蛋白质或小分子的图片(取决于成像设备的分辨率),那么你需要的能力与分析 YouTube 视频或通过摄像头获取的通用视觉是同一类型。所以很多能力都是通用的,你为了一件事开发它们,但那只是另一件事的手段。你可以看到 DeepMind 最初五到七年的情况:我们研究游戏,让 AI 擅长游戏,比如围棋和雅达利游戏。我选择游戏的一个原因是我热爱游戏,我制作游戏,并且一直参与其中。但真正的原因是,它们是当时 AI 系统难度合适的挑战性任务。它们本身从来不是目的;它们是达到目的的手段——给我们提供可量化、可实现的中期目标,这些目标令人印象深刻且非常困难,但又在可能范围内。我们相信那会是一把梯子,让我们达到今天的位置,拥有最终能在现实世界中做真正了不起的事情的系统,解决现实世界的问题,比如科学问题——用 AlphaFold 进行蛋白质折叠,以及现在的药物发现。我个人花时间用这些 AI 系统做的事情就是:AI for science。那一直是我的主要热情所在,也是我构建这些 AI 工具的主要原因。但当然,同一个底层平台还可以用于许多其他不可思议的事情,包括有助于创造力的生成式媒体模型,以及生产力工具,比如大型语言模型。
The whole thesis behind AGI as a term and our original goal at DeepMind from the very beginning was to create a general-purpose intelligent system that could learn from almost any input, generate useful insights or spot useful patterns, and then output that in almost any way. That's obviously how the human mind works, and look at modern civilization that we've created with our hunter-gatherer brains—it's pretty unbelievable if you think about how that happened. That's what we know as a general intelligence. And that's what we've tried to focus on from the beginning of DeepMind and now the whole AI field: systems that are general and that learn, rather than being hard-coded or hard-programmed with the answer. It's funny to think back now, but that's what the AI field used to do for the first 50-60 years. Things like Deep Blue, the chess program, and so on. What that means is that some of these things are inseparable. If you want a fully general system that understands the world around you and can analyze scientific papers or scientific data, including visual data like pictures of cells, proteins, or small molecules depending on the resolution of the imaging equipment, it's the same type of capability you need to analyze YouTube videos or general vision coming through a camera. So a lot of these capabilities are general-purpose, and you develop them for one thing, but that's just a means to an end for another thing. You can see that with the first 5-6-7 years of DeepMind: we were working on games and AI being good at games, like Go and Atari games. One reason I picked games is because I love games, I make games, and I've always been involved in games. But the real reason was that they were challenging tasks at the right level for the AI systems at the time. They were never an end in themselves; they were a means to an end—to give us quantifiable and achievable intermediate goals that were impressive and very hard to do, but just within the realms of possibility. We trusted that would be a ladder to get us to where we are today, having systems that can eventually do really amazing things in the real world and tackle real-world problems like scientific problems—protein folding with AlphaFold and now drug discovery. That's personally what I spend my time using these AI systems for: AI for science. That's always been my main passion and the main reason I'm building these AI tools. But of course, there are many other incredible things that same underlying platform can be used for, including generative media models that are helpful for creativity and also productivity tools like large language models.
这一切是如何联系在一起的,真是令人着迷。我想起你作为神经科学家的第一篇论文,现在很有名。2007 年,它将大脑中的海马体与创造力联系起来。那是一项关于失去记忆的人的研究,发现海马体受损、无法记忆的人也无法想象未来。创造力的这种视觉方面非常重要。即使是针对天生失明者的功能性磁共振成像研究也发现,他们会访问大脑的视觉部分。所以,我在想:当你试图在 AI 中重建一个机器海马体(可以这么说)以让 AI 具有创造力时,你认为这能仅仅通过在这些创意产业的数据上训练模型来实现吗?
It's fascinating how this is all connected. I think back to your first paper as a neuroscientist, which is now famous. In 2007, it tied the hippocampus in the brain to creativity. It was a study of people who had lost their memory, and it found that people who had damage to the hippocampus and couldn't remember things also couldn't picture the future. This visual aspect to creativity is so important. Even fMRI studies on people who were blind from birth find that they access the visual part of their brain. So, I'm wondering: as you try to recreate a machine hippocampus, so to speak, in AI that would let AI be creative, do you think that could come from just training these models on the creative industries?
是的,正如你提到的,这其中的联系是……首先,我在职业生涯早期就利用自己的视觉创造力来帮助设计和编程这些视频游戏。
Yes, so as you mentioned, the through line between that was... I first of all used my own visual creativity to help design and program these video games very early in my career.
当我攻读神经科学博士学位时,我着迷于揭示大脑中让我们能够做到这一点的机制。我们所有人都在不断使用这种能力。至少我过去创作游戏的方式就是:可视化最终目标,可视化玩家玩游戏、使用界面,甚至在编程之前就思考可能出现的问题。我会在脑海中模拟这一切。我们每天做计划时都在做同样的事,比如设想一次重要的商务晚宴,想象每个人坐在哪里,如何开启对话,大家会有什么感受。我们一直在使用大脑的这种想象或未来思考能力。我开始读博时,本应研究记忆,而记忆已经被研究了很长时间。记忆依赖于海马体。有些罕见病人,疾病只攻击海马体,大脑其他部分完好无损。英国有几位这样的病人,我们逐一采访了他们。我有一个想法:阅读记忆文献时,存在两派观点。一派认为记忆像录像带,记录一切。另一派——在我看来显然正确——认为记忆是一个重建过程:当你回忆时,你是在从各个部分主动重建它。如果这是真的,那么想象应该使用相同的大脑机制,只是目标不同:不是重建熟悉的东西,而是从那些组成部分中创造出新颖的东西。事实上,我们发现了这一点。我们是第一个测试这些病人想象能力(而不仅仅是记忆)的人。
When I went into neuroscience to do my PhD, I was fascinated to try and uncover the mechanisms in our brain that allow us to do that. All of us use that all the time. At least that's the way I used to create games: visualize the end goal, visualize a player playing the game, using the interface, and thinking through issues even before it was programmed. I mentally simulated that. We do that every day when we plan, like when we're going to have an important business dinner, we imagine where everyone will sit, how to open the conversation, how people will feel. We use this imaginative or future thinking capability in the brain all the time. When I started my PhD, I was supposed to study memory, which had been studied for a long time. Memory depends on the hippocampus. There are rare patients with a disease that attacks only the hippocampus, leaving the rest of the brain intact. A few of them are in the UK, and we interviewed every single one. I had this idea: reading the memory literature, there were two schools of thought. One is that memory is like a video tape, recording everything. Another, which seemed obviously correct to me, is that memory is a reconstructive process: when you remember something, you actively reconstruct it from its parts. If that were true, then imagination should use the same brain mechanisms, just with a different goal: instead of recreating something familiar, you create something novel from those component parts. And indeed, we discovered that. We were the first to test these patients on their imaginative capabilities, not just their memory.
你认为这些视频模型根据提示重建世界的过程,与大脑中发生的过程之间有什么相似之处吗?
Do you see any similarities between what's happening under the hood with these video models recreating the world from prompts and what's happening in the brain?
我认为肯定是在系统层面上。我不认为实现方式是一一对应的,这从来不是做神经科学的目的。不是要复制大脑,而是要理解大脑可能使用的原理、算法和表征。然后汲取灵感,引导我们 AI 模型的构建方向。所以,我们新的 Omni 模型和视频模型生成世界的方式确实有一些相似之处。关于这个系统如何运作,肯定有很多值得学习的地方。我不少神经科学教授朋友确实在比较最新模型:给定某个提示,模型能做什么,而人类在功能磁共振成像仪中会做什么,会产生什么图像。他们在做各种疯狂而奇妙的事情,比如解码一个人正在思考或梦见的图像,然后用这些模型之一重建视觉内容,再问扫描仪中的受试者:'这是你想象的吗?'结果确实如此。所以,未来几年我们将拥有这些奇妙的科幻设备。
I think definitely on a systems level. I don't think the implementations are similar one-to-one, and that was never the idea behind doing neuroscience. It wasn't to copy the brain, but to understand the principles and algorithms the brain might be using, and the representations it uses. Then lift that inspiration to build into the direction of our AI models. So, there are for sure some similarities between the way our new Omni models and video models are generating the world. There are definitely things to be learned about how that system is working. Quite a few neuroscience professor friends of mine are indeed comparing the latest models: what they can do with a certain prompt versus what a human might do in an fMRI machine and what images they might create. They're doing all sorts of crazy amazing things like decoding what image a person is thinking about or dreaming about, then using one of these models to recreate the visuals, and asking the subject in the scanner, 'Is that what you were imagining?' And it is. So, we're going to have these amazing sci-fi devices in the next few years.
太有趣了。你提到过爱因斯坦测试。我很喜欢这个想法:给 AI 爱因斯坦所拥有的全部数据,仅此而已,截止日期是 1901 年,然后看 AI 能否做到爱因斯坦所做的事:提出相对论和其他物理学突破。听到这个,你可能会想到文本模型,但你实际上想象的是一个更视觉化的过程吗?
So interesting. You've talked about the Einstein test. I love this idea that you could give an AI all the data that Einstein had and nothing more, with a cutoff date of 1901, and see if the AI could do what Einstein did: come up with the theory of relativity and other physics breakthroughs. When you hear that, you could imagine text models, but are you really imagining a more visual process?
我认为这就是我定义真正创造力的方式,也是我的测试标准。人们总是问如何定义它,不是仅仅从已知事物中推断,而是对现实的某一部分提出真正新颖的假设,就像爱因斯坦在 1905 年最著名的那样。也许语言中就包含了足够的信息,如果你能阅读所有文本并牢记在心,就能提出某种隐藏的、相互关联的新理论。但爱因斯坦本人在瑞士做专利员时经常做白日梦,进行思想实验:坐在火车上,以光速旅行,会是什么样子?他利用自己的视觉想象装置提出新理论,然后再用数学证明。我认为你需要接触并理解世界,特别是如果你要提出新实验或必须进行新实验来验证假设。所以这不仅仅是理论;你需要理解原子的世界,而不仅仅是比特或逻辑的世界。
I think that's how I would define true creativity, and that was my test for it. People always ask how to define it where you're not just extrapolating something already known, but coming up with a new hypothesis about some part of reality that is genuinely novel, like Einstein most famously did in 1905. It could be that there's enough in language to come up with a new theory that was hidden, cross-connected in all the text, if you could read it all and hold it in mind. But Einstein himself used to daydream while he was a patent clerk in Switzerland, and dream of thought experiments: being on trains, traveling at the speed of light, what would it look like? He used his visual imaginative apparatus to come up with new theories that he then had to prove mathematically. I think you're going to need to access and understand the world, certainly if you're going to propose new experiments or have to do new experiments to test your hypothesis. So it's not just a theory; you need an understanding of the world of atoms, not just the world of bits or logic.
你知道,在你做电子游戏的日子里,你有一款游戏《共和国:革命》,由你创立的 Elixir 公司开发。这款游戏算是失败了,但它太雄心勃勃了。你在某种程度上试图重建一个前苏联共和国并模拟世界。
You know, in your video game days, you had this game Republic: The Revolution at Elixir, the company you founded. The game was kind of a failure, but it was too ambitious. You were trying to recreate a former Soviet republic and simulate the world in a sense.
是啊。奔腾 2003,可能有点太雄心勃勃了。
Yeah. Pentium 2003, it was probably a little bit too ambitious.
没错。那时候可没有 H100。现在你可以轻松调用数万甚至数十万块 GPU,又在构建这些虚拟世界。有趣的是,这又回到了原点。
Right. There was not an H100 of my time, yeah. Now you have access to tens or hundreds of thousands of GPUs at your fingertips and you're building these virtual worlds again. It's funny how it comes full circle.
但你是否想象过,在这些虚拟世界中,它们将对机器人技术有用,比如训练这些东西上下楼梯之类的?但还有,一个虚拟的爱因斯坦在专利局闲逛、做白日梦。那会是它发生的地方吗?
But do you imagine that inside these virtual worlds, they're going to be useful for robotics, like training these things to go upstairs or whatever? But also, a virtual Einstein wandering around at the patent office and daydreaming. Is that where it's going to happen?
嗯,在《共和国》中,我们试图模拟整个国家,有数十万活生生的人过着日常生活,还有所有政治活动,你要去创造一场革命。那是一个非常雄心勃勃的游戏,我们必须手动编写,花了数年时间。现在令人惊叹的是,我们可能接近用一些系统实际生成它的可能性。我不认为我们已经准备好了,但在两三年内,类似的东西可能成为可能。然后我们就能实现我当时设想的完整愿景。这一切之所以重要且相互关联,是因为模拟和人工智能是基础性的,而且密切相关。模拟之所以有用,是因为这就是想象力的本质——一种模拟。它让你在理论上尝试许多事情,然后选择最佳路径。这就是 AlphaGo 在尝试选择最佳围棋走法时所做的事情。它使用蒙特卡洛树搜索模拟数万种走法,利用围棋模型约束到有用的路径,然后评估哪个最终位置最有希望,从而指导下一步,击败世界冠军。但在许多领域——机器人技术、辅助、科学——我们都希望能够多次重演那个问题。例如,经济学。如果它更像一门自然科学就好了。目前,我们调整利率半个百分点,然后看是否导致衰退。如果能模拟数十万条经济轨迹,调整这些大杠杆,然后从准确的模拟中获取统计汇总,做出严谨的科学决策,那会好得多。这在社会科学中是不可能的,因为你无法以受控方式重复实验。所以,通过学习的模拟——其中 AI 与之相连——如果你理解系统,可以手动编码,但大多数时候我们在数学上并不足够理解系统。AI 系统可以从数据中学习那个模拟。这就是我试图做的更大目标。
Well, with Republic we tried to simulate a whole country with hundreds of thousands of living, breathing people going about their daily lives and all the politics, and you were supposed to create a revolution. It was a very ambitious game that we had to write by hand, and it took years. What's amazing now is that we might be close to the possibility of actually generating that with some of our systems. I don't think we're ready yet, but in two or three years, something close to that might be possible. Then we could realize the full vision of what I was thinking about then. The reason this is all important and connected is that simulations and AI are fundamental and very closely related. Simulations are useful because that's what imagination is—a type of simulation. They allow you to try out many things in theory and then select the best path. That's what AlphaGo did when it tried to select the best Go move. It simulates tens of thousands of moves using Monte Carlo tree search, uses the model of Go to constrain to useful paths, and then evaluates which end position is most promising, guiding its next move to beat the world champion. But there are many areas—robotics, assistance, science—where we would love to have many reruns of that problem. For example, economics. It would be great if it were more like a natural science. At the moment, we adjust interest rates by half a percent and see if it causes a recession. It would be much better if we could simulate hundreds of thousands of trajectories of the economy, adjust the big levers, and get a statistical aggregate of accurate simulations to make a rigorous scientific decision. That's not possible in social sciences because you can't rerun the experiment in a controlled way. So, simulations that are learned—where AI is connected—can be hand-coded if you understand the system, but most of the time we don't understand the system well enough mathematically. An AI system could learn that simulation from data. That's the bigger goal of what I'm trying to do.
我喜欢你将 AI 的创造力部分与科学联系起来的方式。我相信你给了听众很多工作上的灵感。非常感谢你,Demis。很高兴与你交谈。
I love how you've tied the creativity part of AI with the science. I'm sure you're giving this audience a lot of inspiration in their work. So thank you so much, Demis. Great to talk with you.
很高兴交谈。非常感谢。谢谢。
Great to talk. Appreciate it. Thank you.