Demis Hassabis:通往 AGI 之路与缺失的拼图

Demis Hassabis: The Path to AGI and Missing Pieces

杰米斯·哈萨比斯 Demis Hassabis · Y Combinator · 2026-04-29 · 约 41 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Demis Hassabis 讨论当前 AI 范式、持续学习与记忆等缺失组件,以及通往 AGI 的路径。

Demis Hassabis discusses the current AI paradigm, missing components like continual learning and memory, and the path to AGI.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 26)

全文 · Full transcript(中英对照)

0. 德米斯·哈萨比斯简介 Introduction of Demis Hassabis

Host

Demis Hassabis 拥有科技界最不寻常的职业生涯之一。他小时候是国际象棋神童,17 岁设计出第一款热门电子游戏《主题公园》。之后他重返校园,获得认知神经科学博士学位,发表了关于大脑中记忆和想象力如何运作的基础性研究,然后在 2010 年共同创立了 DeepMind,使命只有一个:解决智能问题。我认为他们已经做到了。自那以后,他的实验室完成了一系列大多数人认为还需要几十年才能实现的事情。AlphaGo 击败了围棋世界冠军,AlphaFold 破解了蛋白质结构预测——一个生物学领域长达 50 年的重大挑战——并且他们免费向地球上的每一位科学家开放。这项研究让他去年获得了诺贝尔化学奖。如今,Demis 领导 Google DeepMind,正在构建 Gemini,并朝着他十几岁时就设定的同一个目标迈进:通用人工智能。让我们欢迎 Demis Hassabis。

Demis Hassabis has had one of the most unusual careers in tech. He was a chess prodigy as a kid, then designed his first hit video game, Theme Park, at 17. He then went back to school, got a PhD in cognitive neuroscience, published foundational work on how memory and imagination work in the brain, and then in 2010 co-founded DeepMind with one mission: solve intelligence. And I think they've done it. Since then, his lab has gone on to do things most people thought were decades away. AlphaGo beat a world champion at Go, AlphaFold cracked protein structure prediction, a 50-year grand challenge in biology, and they gave it away for free to every scientist on Earth. That work won him the Nobel Prize in chemistry last year. Today, Demis leads Google DeepMind, where he's building Gemini and pushing toward the same goal he set when he was a teenager: artificial general intelligence. Please welcome Demis Hassabis.

1. 当前范式与AGI缺失环节 Current paradigm and missing pieces for AGI

Host

你思考 AGI 的时间比几乎任何人都长。当你审视当前的范式——大规模预训练、基于人类反馈的强化学习(RLHF)、思维链——你认为我们已经掌握了 AGI 最终架构的多少?目前还缺少什么根本性的东西?

So, you've been thinking about AGI longer than almost anyone. When you look at the current paradigm—large-scale pre-training, RLHF, chain of thought—how much of the final architecture for AGI do you think we already have, and what's fundamentally missing right now?

Demis Hassabis

首先,感谢 Gary 精彩的介绍,很高兴来到这里。感谢你们的欢迎。这真是一个令人惊叹的空间。我以后得常来。你们能在这个领域工作,非常鼓舞人心。关于这个问题:我认为你刚才提到的那些组件,我很确定会成为 AGI 最终架构的一部分。它们已经取得了长足的进步,我们已经验证了它们能做的很多事情。我看不到一个在几年后我们会意识到这是死胡同的世界。这对我来说说不通。但是,在我们已知有效的方法之上,可能还有一两个缺失的东西。所以,持续学习、长期推理、记忆的某些方面,这些仍然没有解决。以及如何让系统在整体上更加一致。我认为所有这些对于 AGI 都是必需的。现在,现有技术可能通过一些创新和渐进式创新就能扩展到那个程度。但也可能还有一两个大想法需要攻克。如果有的话,我认为不会超过一两个。而且,我的赌注大概是五五开。当然,在 DeepMind 和 Google DeepMind,我们两方面都在努力。

Well, first of all, thanks, Gary, for that great introduction, and it's great to be here. Thanks for welcoming me. It's an amazing space, actually. I'll have to come back here often. Very inspiring that you all get to work in this space. So, the question is: I think the components that you just mentioned, I'm pretty sure will be part of the final architecture for AGI. So, I think they've come such a long way now, and we've proven out so many things about what they can do. I can't see a world in which we'll sort of realize in a couple of years this was a dead end. That doesn't make sense to me. But, there still might be one or two things missing on top of what we already know works. So, continual learning, long-term reasoning, some aspects of memory, these are still unsolved. And how to get the systems to be more consistent across the board. I think all of these are going to be required for AGI. Now, it might be that the existing techniques can just scale up to that with some innovation and some incremental innovation. But, it could be that there's still one or two big ideas left that need to be cracked. I don't think it's more than one or two if there are out there. And I think, you know, my betting is about 50/50 if that's the case. So, of course, at DeepMind at Google DeepMind we work on both those things.

2. 持续学习与记忆挑战 Continual learning and memory challenges

Host

我想这就是我的意思。处理一堆相同的系统,最让我惊讶的是这些权重重复使用的程度。所以,持续学习这个概念非常有趣,因为现在我们在某种程度上是用胶带拼凑起来的,你知道吧?比如那些夜间梦境循环之类的东西。

I guess that's what I mean. Working with a bunch of identical systems, the wildest thing to me is to what degree it's the same weights over and over. So, this idea of continual learning is so interesting because like yeah, right now we're sort of cobbling it together with duct tape, you know? These dream cycles at night and things like that.

Demis Hassabis

是的。梦境循环很酷,我们以前在情景记忆的巩固中思考过这个问题。这实际上是我博士研究的内容:海马体如何工作,以及如何将新知识优雅地整合到现有知识库中。大脑在这方面做得非常好。它通过睡眠,尤其是快速眼动睡眠,重放重要的情节,以便从中学习。事实上,我们最早的 Atari 程序 DQN 掌握 Atari 游戏的方法之一就是进行经验重放。所以我们从神经科学中借鉴了这一点,多次重放成功的轨迹,那是在 2013 年,AI 的黑暗时代。这是一件非常重要的事情。我同意你的看法,我们现在有点像是在用胶带拼凑。比如把所有东西都塞进上下文窗口。但这似乎有点不令人满意,对吧?实际上,即使我们在处理机器而不是生物大脑,你可能有数百万或数千万大小的上下文窗口或记忆,而且它可以很完美,但查找并找到与当前具体决策真正相关的东西仍然有成本。这个成本并不小,即使你有可能存储所有东西。我认为在记忆等领域实际上有很多创新空间。

Yes. It's pretty cool, the dream cycles, and we used to think about this with consolidation with episodic memory. It's actually what I studied for my PhD: how the hippocampus works and integrates new knowledge gracefully into the existing knowledge base. So, the brain does that amazingly well. It does it through, you know, during sleep especially things like REM sleep, replaying back episodes that are important so that you can learn from it. In fact, our very first Atari program DQN, one of the ways it was able to master Atari games was by doing experience replay. So, we sort of borrowed that from neuroscience and replayed successful trajectories many times, you know, that's way back in 2013 now in the dark ages of AI. It was a really important thing. And I agree with you, we're kind of using duct tape right now. So, like shove it all in the context window. But this seems a bit unsatisfying, right? And actually, even though we're working on machines, not biological brains, and so you potentially could have, you know, millions or tens of millions size context window or memory, and it can be perfect, there's still a cost to looking it up and finding the right thing that's actually relevant for the specific decision you've got to make right now. And that's non-trivial that cost, even if you can potentially store it all. I think there's actually a lot of room for innovation in areas like memory.

3. 上下文窗口限制与智能体系统 Context window limitations and agentic systems

Host

是的。我的意思是,一百万 token 的上下文实际上比我想象的要大,说实话,它足够大了。你可以做很多事情。

Yeah. I mean, the one thing is like it feels like a million token context one is actually bigger than I mean, it's plenty big, honestly. You can do stuff.

Demis Hassabis

对于大多数应该使用它的场景来说,它足够大了。我的意思是,如果你把上下文窗口看作工作记忆,人类只有几个数字,大概十几个数字,平均七个。我们有百万甚至千万的上下文窗口,但问题是我们试图把所有东西都存进去,包括不重要的、错误的东西。目前这相当粗暴,似乎不太对。然后问题是,如果你是一个智能体,试图处理实时视频,并且你只是天真地记录所有 token,那么一百万 token 其实并不多。只有大约 20 分钟。所以,如果你想要一个能理解你生活中一两个月内发生的事情的系统,你实际上需要更多。

It's plenty big for most things that it should be used for. I mean, if you think about the context window is sort of equivalent to working memory, you know, humans have we have like a few digits, you know, it's like a dozen digits maybe, you know, average of seven. We got million or, you know, 10 million context windows, but the problem is that we're trying to store everything in that, you know, things that aren't important, things that are wrong. It's pretty brute force currently, and that doesn't seem right. And then the problem is if you're an agent trying to process live video, and you're just going to naively record all the tokens, then actually a million tokens isn't that much. It's only like 20 minutes. So, actually you need more if you want something that's going to understand your, you know, what's going on in your life over maybe a month or two.

4. 深度强化学习与智能体哲学 Reinforcement learning and agent philosophy at DeepMind

Host

DeepMind 历来倾向于强化学习和搜索。AlphaGo、AlphaZero 和 MuZero。这种理念有多少实际上融入了你今天构建 Gemini 的方式中?强化学习仍然被低估了吗?

DeepMind has historically leaned into reinforcement learning and search. AlphaGo, AlphaZero, and MuZero. How much of that philosophy is actually embedded in how you're building Gemini today? Is RL still underrated?

Demis Hassabis

是的,我认为可能确实如此。它时起时落。你知道,我们从 DeepMind 成立之初就一直在研究智能体。事实上,我们当时就是这么说的。所以,所有 Atari 的工作,尤其是 AlphaGo,都是智能体系统。我们的意思是,这些系统能够自主完成目标,做出主动决策并制定计划。当然,我们是在游戏领域做这件事,以便让它易于处理。

Yeah, I think potentially it is. It sort of goes in ebbs and waves. You know, we've worked on agents since the beginning of DeepMind. In fact, that's what we said we were working on. And so, all of the Atari work and AlphaGo, most specifically, they're agent systems. And what we meant by that is systems that are able to accomplish goals on their own, and make active decisions and make plans. And so, of course, we were doing it in the domain of games to make it tractable.

5. 从游戏到通用世界模型 From Games to General World Models

Demis Hassabis

然后我们开始做越来越复杂的游戏,比如 AlphaGo 之后的 StarCraft、AlphaStar。基本上所有游戏我们都做过了。当然,接下来的问题是,能否将这些模型泛化为世界模型或语言模型,而不仅仅是简单游戏甚至复杂游戏的模型。这就是过去几年我们在做的事。但实际上,你可以把我们今天做的很多事情——所有前沿模型的思考模式和思维链推理——都看作是当年 AlphaGo 开创性工作的回归。我其实觉得我们当时做的很多工作今天仍然适用,我们现在正在以更通用的方式重新审视那些旧想法,包括蒙特卡洛树搜索以及其他在强化学习之上增强的方法。我认为 AlphaGo 和 AlphaZero 的许多想法都与今天的基础模型高度相关,而且未来几年的进步很大程度上将来自这些想法。

And then doing increasingly complex games, things like StarCraft after AlphaGo, AlphaStar. So we basically did all the games that are out there. And then of course, the question is, can you generalize those models to be world models or models of language, not just models of simple games, or even complex games. And that's what the last few years has been about. But really, you can think of a lot of the things we're doing today, all the leading models with thinking modes and chain-of-thought reasoning as aspects of what was sort of pioneered with AlphaGo coming back now. And I actually think there's a lot of work we did back then that is relevant today, and we're sort of re-looking at some of those old ideas at scale today in a more general way, including things like Monte Carlo tree search and other ways of augmenting the reinforcement learning on top of the reinforcement learning we're ready to do today. And I think a lot of those ideas both from AlphaGo and AlphaZero are really relevant to where we are with today's foundation models. And I think a lot of that is what we're going to see of the advances the next few years.

6. 蒸馏与小模型 Distillation and Smaller Models

Host

我有一个问题:显然今天你需要越来越大的模型才能越来越智能。但我们也看到蒸馏技术很有效,小模型可以快很多。你们有非常厉害的 Flash 模型,大概能达到前沿模型 95% 的水平,而价格只有十分之一,是这样吗?

One question I would have: obviously today you need bigger and bigger models to be smarter and smarter. But then we're also seeing distillation working. And then smaller models can be quite a bit faster. I think you guys have incredible flash models that are like 95% as good as the frontier and at like 1/10 the price. Is that right?

Demis Hassabis

我认为这是我们的核心优势之一。你必须构建最大的模型才能拥有前沿能力。但我们最大的优势之一,是能够快速将这种能力蒸馏并压缩到越来越小的模型中。显然,我们发明了蒸馏过程,有 Jeff、Oriol 等人参与。我们至今仍是这方面的世界专家。而且我们也有巨大的需求,因为我们要服务可能是最大的 AI 界面。显然有搜索的 AI 概览和 AI 模式,还有 Gemini 应用。现在 Google 的每一个产品——地图、YouTube 等等——都越来越多地融入了 Gemini 或其相关技术。这涉及数十亿用户,超过十几个十亿级用户产品。它们必须被极快、极高效、极便宜且低延迟地服务。这给了我们一个非常重要的动力,让这些 Flash 甚至更小的模型变得极其高效。希望这最终能对你们使用的许多工作负载非常有用。

I think that's one of our core strengths. I mean, you have to build the biggest models to have the frontier capabilities. But I think one of our biggest strengths has been distilling and packing that power into smaller and smaller models very quickly. Obviously we invented the kind of distillation process, with people like Jeff and Oriol and others. And we're still world experts in that. And we also have a huge need to do it because we've got to serve the biggest probably AI surfaces there are. Obviously there's search with AI overviews and AI mode. Then there's Gemini app. And now increasingly every single product at Google has, you know, maps and YouTube and so on, some aspect of Gemini or Gemini related technology in it. And so that's billions of users, a dozen more than a dozen billion user products. And they have to be served extremely fast, extremely efficiently and cheaply and with low latency. So that gives us a really important incentive to make these flash and even smaller models flashlight models extremely efficient. And hopefully that ends up then being really useful for many of the workloads that all of you use for.

Host

我很好奇这些小模型到底能有多智能。比如,蒸馏过程有极限吗?一个 50B 或 400B 的模型能像今天的某个前沿模型一样聪明吗?

I'm curious about how much smarter these smaller models can actually be. Like, are there limits to the distillation process? Like, could a 50B or 400B model be as smart as like a mythos for today?

Demis Hassabis

是的,我不认为我们已经达到了任何极限。至少我们谁也不知道是否已经达到了某种信息或极限。我的意思是,也许在某个时候会出现信息密度无法再突破的情况。但就目前而言,我们的假设是,在我们推出一个领先的 Pro 模型或前沿模型一年后,半年或一年后,你就能在非常小的、几乎是边缘的模型中看到它。你也会在 Gemma 模型中看到一些好处,希望你们都在享受我们的 Gemma 4 模型,我认为它们在同等规模下性能惊人。这同样大量使用了这些蒸馏技术,以及如何在非常小的模型中实现高效的想法。所以,我还没有看到任何理论上的极限。我认为我们离那还很远。

Yeah, I don't think we've got to any kind of limit. At least none of us know yet if we've got to any kind of information or limit. I mean, maybe at some point that will be the case where there's just an information density that can't we can't get beyond. But I think for now the assumption we make is that, you know, a year later after one of our leading pro models or frontier models goes out, half a year later, a year later, you'll have them in the really tiny, almost edge models. And you'll also see some of that goodness in our Gemma models, which hopefully you're all enjoying our Gemma 4 models, which I think are really amazing power for their sizes. So, again, that uses a lot of these distillation techniques and the idea of how to make things really efficient in these very small models. So, I don't really see any limit yet in terms of some kind of theoretical limit. I think we're still pretty far off of that.

Host

那真的很好。

That's really good.

Demis Hassabis

是的。

Yes.

7. 对开发者生产力的影响 Impact on Developer Productivity

Host

我们现在看到的一个比较奇怪的现象是,工程师能完成的工作量大概是 6 个月前的 500 到 1000 倍。我是说,在座有些人做的工作量是 2000 年代 Google 工程师的 1000 倍,Steve Yegge 也谈到过这一点。

You know, one of the weirder things that we're seeing right now is like engineers can do like 500 to 1,000 times the amount of work that they were doing like 6 months ago, I guess. I mean, the people in this room there are people who are doing about like a 1,000 x the work that like I Steve Yegge talks about this. It's like a 1,000 x the work that a Google engineer from the 2000s was doing.

Demis Hassabis

我认为这非常令人兴奋。模型有很多用途,成本显然是一方面,但速度允许你——比如在编程或其他事情上——迭代得更快。尤其是当你与系统协作时。我认为非常需要快速系统,它们可能不是最前沿的,就像你说的,95% 或 90% 的水平,但这已经足够好了,而且实际上在迭代速度上能弥补那 10% 的差距。另一个重要方面是在边缘运行这些模型。同样,出于效率原因,也出于隐私和安全原因。想想你可能在哪些设备上运行这些系统,它们处理非常私人的信息。也可以想想机器人技术。你家里的机器人。我认为你会想要非常高效、非常强大的本地模型,它们可能由云中的一些更大模型或前沿模型编排,但只在特定情况下才委托给它们。也许你会在本地处理所有音视频输入,并保持本地化。我可以想象那会是一个非常好的最终状态。

I think it's very exciting. I mean, I think models have many uses. One is obviously cost, but the speed can allow, if you think about coding even or other things, you can iterate a lot faster. Also, especially if you're collaborating with the system. I think there's a lot of need for having fast systems that maybe are not quite frontier, like you said, like 95%, 90%, but that's plenty good enough and actually gain back more than the 10% on the iteration speed. So, and then the other big thing I think is running these things on the edge. Again, for efficiency reasons, but also for privacy and security reasons, too. If you think about different devices that you might run these systems on that process very personal information. Can also think about robotics, as well. Robots in your house. I think you're going to want very efficient, very powerful local models, which maybe are orchestrated with some bigger models, frontier models that are in the cloud, but you only delegate to that in certain circumstances. And perhaps you process all of the audio-visual feed, let's say, locally, and that stays local. I could imagine that would be a very good sort of end state.

8. 智能体的上下文与持续学习 Context and Continual Learning for Agents

Host

回到上下文和记忆的问题,模型目前是无状态的,但继续学习的话,对于使用持续学习模型的开发者来说,体验会是什么样的?比如,你知道如何引导它吗?

Going back to context and memory, models currently stateless, but, you know, continue like what would the developer experience even be like for someone who's using a continual learning model? Like, any idea how you'd steer it?

Demis Hassabis

我认为这非常有趣。我认为这是阻碍智能体完成完整任务的因素之一。目前它们对任务的某些方面很有用,你可以把它们拼凑起来做一些很酷的事情,但它们不能很好地适应你所处的上下文。我认为这是它们无法真正做到“一劳永逸”并自行解决问题的缺失环节。它们需要能够学习你将要放置它们的特定上下文。所以,我认为我们必须攻克这一点才能实现完全的通用智能。

I think it's really interesting. I think that's one of the things holding back agents from doing full tasks, you know? I think they're really useful for aspects of tasks right now, and you can patch them together and do some really cool things, but they don't adapt well with the context that you're in. And I think that's the missing piece for them being really kind of fire and forget, and they'll figure it out themselves. You know, I think they need to be able to learn about the specific context that you're going to put them in. So, I think we have to crack that to get full general intelligence.

9. 推理差距与思维链 Reasoning Gaps and Chain of Thought

Host

我们在推理方面进展如何?模型现在能进行非常令人印象深刻的思维链推理,但它们仍然会在一个聪明的本科生不会犯错的方面失败。具体需要改变什么,你在推理方面预期有什么进展?

Where are we on reasoning? So, models can do really impressive chain of thought now, but they still fail on things a smart undergrad wouldn't. What specifically needs to change and what progress do you expect in reasoning?

Demis Hassabis

我认为在思维范式方面还有很多创新空间。我们目前做的还相当简单粗暴。可以想象,在监控思维链方面有很多潜力,比如在思考过程中中途介入。我经常感觉我们的系统和竞争对手的系统几乎是在过度思考,它们几乎陷入循环。我有时喜欢和 Gemini 下棋,所有领先的基础模型在游戏方面都相当差,这很有趣。查看思考轨迹很酷,因为这些可以被很好地理解。我能很快判断它是否偏离主题,而且很容易证明思考是否有用。我们看到的是,有时它会考虑一步棋,意识到那是败招,但找不到更好的,于是又回到那步棋并最终走出。在一个精确的推理系统中不应该出现这种情况。所以我认为仍然存在巨大差距,但可能只需要一两个调整就能修复这些差距。不过这些差距很明显,这就是为什么会出现这种锯齿状智能。一方面,它能解决 IMO 金牌问题,非常难;另一方面,正如我们所见,如果以某种方式提问,它仍然会犯基本的小学数学错误,对吧?或者基本的推理错误。所以我觉得它对自己的思考过程缺乏某种内省,可能缺少了什么东西。

There's a lot of innovation left in the thinking paradigms, I would say. Again, I think we're doing fairly simplistic things, fairly brute force. One could imagine there's a lot of scope for example in monitoring the chain of thought, maybe interjecting midway through a thought process. I often get the impression with our systems and our competitor systems that they're almost overthinking. They're almost getting into loops of things. One thing I sometimes like to do is play chess against Gemini, and you know, all the leading foundation models are pretty poor at games, which is quite interesting. It's very cool to look at the thinking traces because obviously these can be well-understood. I can tell quite quickly if it's going off on a tangent and it's very provable what the thinking is doing, whether it's useful or not. And so, what we see is that sometimes it will consider a move. It will realize it's a blunder, but it can't find anything better, so it kind of goes back to that move and does it anyway. So, you just shouldn't be seeing that happening in a very precise reasoning system. So, there are huge gaps, I think, still, but it may only be one or two tweaks that are required to fix those kinds of gaps, just to be clear. But I think it's pretty obvious they are there, and that's why you get this kind of jagged intelligence. On the one hand, it can solve gold medal problems in IMO, which is super hard, but on the other hand, as we've all seen, it can still make basic elementary math errors if you pose the question in a certain way, right? Or elementary reasoning errors. So, there's something to me about almost an introspection about its own thought process that I feel like there's something maybe missing there.

10. 智能体能力与炒作 Agent Capabilities and Hype

Host

智能体现在非常热门。有些人说它们被过度炒作。我个人认为它们才刚刚起步,这太疯狂了。DeepMind 的内部研究告诉你,智能体的实际能力现在处于什么水平,与外界的炒作相比如何?

Agents are really big. Some would say they're hyped. I personally think they're just getting started. It's totally insane. What does DeepMind's internal research tell you about where agent capabilities actually are right now versus the hype out there?

Demis Hassabis

我认为我们才刚刚开始。要达成 AGI,你必须有一个能主动解决问题的系统。这一点我们一直很清楚。所以智能体就是那条路,我们才刚刚起步。我们都在适应如何最好地工作,而你们在自己的个人实验中引领着方向。我相信你们很多人都在这样做。我认为关键是如何将它融入工作流程,不是锦上添花,而是真正开始做基础性的事情。目前我们都在实验,尝试很多东西,但可能只是最近几个月才开始找到真正有价值的地方。而且技术可能刚好变得足够好,对吧?它不再是玩具式的演示,而是真正为你的时间和效率增值。我经常想,我看到很多人设置几十个智能体运行 40 小时,但我还没看到产出能完全证明这种投入是值得的,不过我认为这将会到来。所以我们仍处于实验阶段。我们还没看到一款通过“氛围编码”制作、登顶 App Store 排行榜的 3A 游戏,对吧?我见过也编程过,我相信很多人都做过一些不错的演示,很惊人。我现在能在半小时内做出一个主题公园的原型,而我 17 岁时需要 6 个月。这令人难以置信,我希望有这种感觉。如果花整个夏天去打磨,你能做出真正不可思议的东西,但它仍然需要工艺、人类的灵魂和品味。我认为你必须确保将这些带入你构建的任何东西中。而且我认为这仍然表明它还没完全到位,因为为什么我们还没看到一个孩子做出销量 1000 万的热门游戏?考虑到投入的努力,这应该是可能的。所以仍然缺少某些东西。也许与过程有关,也许与工具有关。我不太确定。你们可能比我更清楚,因为我相信你们都在实验。但我还没看到我期望的结果,一旦它真正发挥全部价值,我认为这会在未来 6 到 12 个月内实现。

I think we are just at the beginning. You have to have an active system that can actively solve problems for you to get to AGI. That was always clear to us. So, agents are that path and I think we're just getting going. I think all of us are getting used to how we best work, and you're leading the way in a lot of this in your own personal experiments. I'm sure many of you are doing that. I think how do you incorporate it into your workflow in a way that isn't just a nice to have, but actually starting to do fundamental things. At the moment we're all experimenting, we're experimenting a lot of things, but we're only in the maybe the last couple of months starting to find the really valuable places. And the technology probably only getting good enough for that to be the case, right? Where it's not a kind of toy nice demonstration, but actually really adding value to your time and efficiency. I often wonder, I see a lot of people working on setting off dozens of agents for like 40 hours, but I'm not sure I've seen the output yet that quite justifies that level of input going in, but I think it will come. So, I still think we're in the experimentation phase. We haven't seen a AAA game that tops the App Store charts that was sort of vibe coded yet, right? I've seen and I've programmed and I'm sure many we've all done little nice demonstrations and it's like amazing. I can do a prototype of a theme park in half an hour now which took me 6 months back when I was 17. It's kind of mind-blowing, and I wish I got this feeling. If I spent the whole summer working on it, you could make something really incredible, but it still needs craft and human sort of soul into it and taste. I think that's something you have to make sure you still bring to whatever it is you're building. And I think it still shows like it's not quite there yet because why haven't we seen a kid making a hit game that sells 10 million copies, right? That should be possible given the effort that's gone in. So something's still somehow missing. Maybe it's to do with the process, or maybe it's to do with the tools. I'm not quite sure. You will probably know better than me because I'm sure you're all experimenting on that. But I haven't seen the result yet which I would expect once this is really delivering that full value. Which I think will come in the next 6 to 12 months.

11. 自主与增强智能体 Autonomous vs. Augmented Agents

Host

其中一部分是,有多少会是自主的?我的意思是,我认为我们不会先看到自主的。我们实际上可能会看到这个房间里的人以 1000 倍效率工作,然后……

Some of it is like how much of it will be autonomous versus I mean, I don't think we'd see autonomous first. We would actually probably see people in this room operating at 1000X, and then...

Demis Hassabis

那应该是你首先看到的,然后你们中的许多人,比如游戏公司或其他类型的公司,会使用这些工具构建出某种畅销应用或畅销游戏。那应该是你首先看到的,然后更多部分会被自动化。

That's what you should see first, and then many of you, they'll be like games companies or other types of companies that have built some kind of best-selling app, best-selling game using these tools. That's what you should see first, and then more of that will get automated.

Host

我的意思是,其中一部分是,里面有人类,然后人类还不想说这是智能体做的。

I mean, some of it is like there's a human in there, and then the human doesn't want to say that the agents did it yet.

Demis Hassabis

我认为部分原因可能是我们想讨论创造力。我经常说的一个例子是 AlphaGo。大家都知道第二局中的第 37 手,对我来说,我一直在等待那样的时刻来启动像 AlphaFold 这样的科学项目。我们从首尔回来的那天就启动了 AlphaFold,那是 10 年前了。之后我会去韩国庆祝 AlphaGo 十周年。但仅仅想出第 37 手还不够。那很酷,非常有用,但它能发明围棋吗?那才是我想要的:一个系统,如果你给它一个高层描述,它能发明出围棋。比如,一个你可以在 5 分钟内学会规则,但需要很多辈子才能精通的游戏。它在美学上很优美,但一个下午就能玩几小时。所以,也许你可以想象那是我会给出的高层描述,然后我希望得到的是围棋。对吧?而显然,今天的系统我认为做不到。所以问题是为什么?我认为那里仍然缺少一些东西。

I think part of it might be though that we want to discuss creativity. What I often say about that is like if we look at the things we've done like AlphaGo. So obviously very famously you'll all know about the move 37 in game two, and for me I was waiting for a moment like that to start the science projects like AlphaFold. We started AlphaFold like the day we got back from Seoul, which is 10 years ago now. I'm going to Korea after this to celebrate the 10-year anniversary of AlphaGo. But it's not enough to come up with move 37. That's pretty cool, very useful, but can it invent Go? That's what I want: a system that can invent Go if you give it a high-level description. You know, like a game you can learn the rules of in 5 minutes, but it takes many lifetimes to master. It's beautiful aesthetically, but you can play it in a few hours in an afternoon. So, maybe you could imagine that would be the high-level description I would give, and then I'd want the thing I get back is Go. Right? And clearly today's systems, I think, can't do that. So, the question is why? And I think there's something still missing there.

Host

嗯,这个房间里可能有人能做到。

Well, someone in this room might make it.

Demis Hassabis

那么答案就是什么都不缺,只是我们使用系统的方式问题。

Then the answer would be there's nothing missing. It just was the way we were using the systems.

12. AI工具与创造力 Creativity with AI tools

Demis Hassabis

这或许就是答案。今天的系统已经具备这种能力,只要有一个足够出色的创意者来使用它,提供项目的灵魂,并且足够熟悉工具,几乎能与工具融为一体。我可以想象,如果你像在座许多人那样日夜不停地实验这些工具,再结合真正的深度创造力,就能做出更不可思议的事情。

And that might actually be the answer. It might be that today's systems are capable of that with a brilliant enough creative person using it and providing that impetus that the soul of the project and being able to probably be au fait enough with the tools to almost be at one with the tools. I could imagine that happening if you experimented with the tools all day and all night, like probably many of you are doing, and you combine that with proper deep creativity, something more incredible could be done.

13. 开源与开放权重 Open source and open weights

Host

换个话题,谈谈开源或开放权重。最近发布的 Gemma,你们推出了能力很强、开放且可访问的模型,能真正在本地运行。这意味着什么?AI 是否会掌握在用户手中,而不是主要放在云端?这是否会改变谁可以用这些模型来构建?

Switching gears to open source, or open and open weights. The recent release of Gemma, you're making highly capable open and accessible ones that can actually run locally. What do you think that means? Will AI be something that is in the hands of the users instead of primarily in the cloud? And does that change who gets to build with these models?

Demis Hassabis

我们总体上非常支持开源和开放科学。你一开始提到了 AlphaFold,我们把它完全免费开放了。我们所有的科学工作,直到今天,仍然发表在顶级期刊上。我们想创建在各自尺寸上世界领先的模型,希望 Gemma 做到了这一点。我们非常坚定地走这条路。希望大家都能实验、构建并享受使用 Gemma。我认为现在两周半内下载量已达 4000 万次,我们对此非常兴奋。我也认为拥有西方的开源栈很重要。显然,很多中国模型非常优秀,目前在开源领域领先,而我们认为 Gemma 在其尺寸上在所有方面都很有竞争力。对我们来说,存在资源、人才和算力的问题。没有人有足够的闲置算力来制造两个最大尺寸、不同属性的前沿模型,所以这非常困难。但就目前而言,我们决定将边缘模型——用于 Android、眼镜和机器人的模型——作为开放模型,因为一旦部署到终端,它们无论如何都很脆弱,所以不如完全开放。我们决定在纳米尺寸级别上统一这一点,这在战略上对我们也有利。我们希望尽可能多的人在此基础上构建,当然,我们自己也会在此基础上继续开发。

We're huge proponents of open source and open science in general. You mentioned AlphaFold at the beginning; we put that all out there for free. All of our science work, even still today, we publish in the big journals. We wanted to create world-leading models for their sizes, and that's what we've hopefully done with Gemma. We're very committed to that path. Hopefully you all experiment, build, and enjoy using Gemma. I think it's been like 40 million downloads now in just 2 and a half weeks, so we're really excited about that. I also think it's important for there to be Western stacks on open source. Obviously, a lot of the Chinese models are excellent and currently well leading in open source, and we think Gemma is very competitive for its sizes in all those respects. For us, there is a question of resources, talent, and compute. Nobody has enough spare compute to just make two frontier models at maximum size with different attributes, so that's pretty difficult. But for now, we've decided that our edge models—the things we want to use for Android, glasses, and robotics—are best as open models because they're vulnerable anyway once you put them out on the surfaces, so they might as well be fully open. We've made a decision to unify that at the nano size level, which actually works for us strategically as well. We hope as many people as possible build on it, and of course, we'll be building on that too.

14. 多模态Gemini优势 Multimodal Gemini and its advantages

Host

在开始之前,我给你演示了我的《Her》中 Samantha 版本,这对我来说演示给你看很紧张。但它成功了,太棒了。Gemini 天生是多模态的,我花了很多时间研究各种模型。上下文的深度和直接通过语音使用工具的能力——无与伦比,实际上是最好的。

Earlier before we came on, I got to show you a demo of my version of Samantha from Her, which was harrowing for me to try to demo something to you. It worked, which is amazing. Gemini was built multimodal, and I spent a lot of time with a bunch of the models. The depth of the context and the tool use with speech directly to model—there's nothing like it, bar none, the best one actually.

Demis Hassabis

我认为 Gemini 系列有一个仍被略微低估的方面:我们从一开始就构建了多模态。这比只专注于文本要困难一些,但我相信从长远来看我们会从中受益。我们现在已经在世界模型构建中看到了这一点,比如我们在 Gemini 之上构建的 Genie。我认为这对机器人技术至关重要。这就是为什么 Gemini 机器人——在座很多人可能玩过——我认为它将基于多模态基础模型构建。我们认为 Gemini 在多模态方面的强大实力给了我们竞争优势。我们越来越多地将其用于 Waymo 等领域,但想象一下设备和助手——那些随你进入现实世界的数字助手,可能在你的手机、眼镜或其他设备上——它需要理解你周围的物理世界、直观的物理学以及你所在的物理环境。这正是我们系统极其擅长的。我想你发现这就是为什么你在自己的设置中喜欢使用它的原因。我们计划继续推进,我认为我们在这些类型的问题上遥遥领先。

I think that's a still slightly underappreciated aspect of the Gemini series: we started it being multimodal from the start. That made it a little bit more difficult to begin with than just focusing on text, for example. But I believe we're going to gain from that in the long run. I think we're seeing that now for things like world model building, so stuff like Genie that we build on top of Gemini. I think it's going to be really important for things like robotics. This is why Gemini robotics, which many of you probably played around with, I think it's going to be built on multimodal foundation models. We think we have a sort of competitive advantage with Gemini being so strong at multimodal. We're using it increasingly in things like Waymo, but also if you imagine devices and assistants—digital assistants that come with you into the real world, maybe on your phone or glasses or some other device—it needs to understand the physical world around you, intuitive physics, and the physical context you're in. That's what our systems are extremely good at. I think you found that's why you've enjoyed using it in your setup. We're planning to continue on that, and I think we're far and away the strongest models on those types of problems.

15. 推理成本与未来可能 Inference cost and future possibilities

Host

推理成本正在快速下降。当推理基本免费时,什么会成为可能?这如何改变你的团队实际优化的目标?

The cost of inference is dropping fast. What becomes possible when inference is essentially free, and how does that change what your team is actually optimizing for?

Demis Hassabis

我不确定推理是否会基本免费。存在杰文斯悖论等问题;我认为我们最终会用尽所有能获得的资源。你可以想象数百万个智能体、智能体群协同工作。这是使用推理的一种方式。或者你可以想象单个智能体或较小的智能体组从多个方向思考,然后进行集成。我们正在实验所有这些方法。在座很多人可能也在做。所有这些都会消耗掉所有可用的推理资源。有一天,推理成本可能几乎为零,尤其是如果我们通过材料科学解决了聚变、超导体或最优电池等问题,能源成本将基本为零。但芯片的物理制造等仍然存在瓶颈,至少在未来几十年内我认为是这样。因此,推理方面仍然需要配给,你仍然需要高效使用它。

I'm not sure inference will ever be essentially free. There's Jevons' paradox and other things; I think we'll just end up using whatever we can get our hands on. You could imagine millions of agents, swarms of agents working together on things. That's one way to use inference. Or you could imagine single agents or smaller groups of agents thinking in multiple directions and then ensembling that. We're experimenting with all these things. Probably many of you are. All of that will use up any inference that's available. One day maybe it can be almost cost zero, certainly the energy if we solve fusion or superconductors or optimal batteries or some set of those things, which I think we will do with material science. Energy costs will be essentially zero, but there'll still be the physical creation of the chips and other things. There'll be some bottleneck, at least for the next few decades, I think. So if that's the case, there'll still be rationing on the inference side. You still have to use it efficiently.

16. AlphaFold 3与虚拟细胞 AlphaFold 3 and virtual cell

Host

幸运的是,较小的模型正变得越来越智能,这太棒了。观众中有很多生物和生物技术创始人。AlphaFold 3 将我们从蛋白质带到了广泛的生物分子。我们离模拟完整细胞系统还有多远?或者这仍然是一个本质上更困难、自成一类的问题?

Luckily, the smaller models are getting smarter and smarter, which is fantastic. We got a lot of bio and biotech founders in the audience. AlphaFold 3 took us beyond proteins to a broad spectrum of biomolecules. How close are we to modeling full cellular systems, or is that still a fundamentally harder problem in a class of its own?

Demis Hassabis

Isomorphic Labs 是我们完成 AlphaFold 2 后从 DeepMind 分拆出来的,目前进展非常顺利。它不仅在构建 AlphaFold——这只是药物发现过程的一部分——我们还在尝试做相邻的生物化学和化学,以设计具有正确属性的化合物等。我们很快会有一些重大消息宣布。我认为这方面进展很好。最终,你想要一个完整的虚拟细胞。

Isomorphic Labs, which we spun out from DeepMind after we did AlphaFold 2, is going amazingly well. It's trying to build out not just AlphaFold—it's just one piece of the drug discovery process—but we're trying to do the adjacent biochemistry and chemistry to design the right compounds with the right properties, and so on. We'll have some big announcements very soon on that front. I think that's going really well. Eventually, you want a whole virtual cell.

17. 虚拟细胞与数据挑战 Virtual Cell and Data Challenges

Demis Hassabis

我在很多科学演讲中都提到过,一个完整的工作细胞模拟,你可以扰动它,然后它的输出会足够接近实验,以至于有用,对吧?你可以跳过很多搜索步骤,生成大量合成数据来训练其他模型,这些模型随后可以预测真实细胞的情况。我认为我们大概还需要 10 年才能实现类似虚拟细胞的东西,一个完整的虚拟细胞。我们在 DeepMind 的科学方面,先从虚拟细胞核开始,因为它相对独立。所有这些事情的诀窍在于,你能选择复杂性的一个切片吗?最终你想模拟人体,但你能在正确的细节层次上建模吗?你能取出哪个切片,让它足够独立?你可以建模并近似该独立系统的输入和输出,然后只专注于这个独立系统。从这个角度看,细胞核非常有趣。另一个问题是数据还不够。所以你需要数据,我和许多顶尖科学家讨论过,他们研究电子显微镜和其他成像技术。如果我们能对活细胞成像而不杀死它,那将是颠覆性的,因为那样就可以把它变成一个视觉问题,我们知道如何解决。但目前,我不知道有任何技术能在不破坏活体动态细胞的情况下提供纳米级分辨率。你可以看到所有相互作用。显然,你可以以那种分辨率拍摄静态图像。现在非常详细,这很令人兴奋,但还不足以把它变成一个复杂的视觉问题。所以这是一种可能的解决方式。它可能是一个硬件驱动的数据驱动解决方案,或者我们可以构建更好的学习模拟器来模拟这些动态系统。这是更偏向建模的解决方式。

So, I've talked about this in many of my science talks about a full working simulation of a cell that you can perturb, and then the outputs of that would be close enough to experimental that it's useful, right? You could skip out a lot of the search steps, and generate lots of synthetic data to train other models that then would predict things about real cells. And I think we're about 10 years away probably from something like a virtual cell, like a full virtual cell. We're starting out on the DeepMind side, science side, on a virtual nucleus, cell nucleus first, because it's relatively self-contained. The trick with all of these things is, can you pick a slice of the complexity? Eventually you want to model a human body, but can you model it down to the right level of detail and what slice can you take out of it that will be self-contained enough? You can kind of model and approximate the inputs and outputs into that self-contained system and then just focus on the self-contained system. So, a nucleus is quite interesting from that perspective. Then the other issue is just there's not enough data yet. So, you need data and I talked to various top scientists who work on electron microscopes and other imaging things. If we could image a live cell without killing the cell, that would be game-changing obviously because then you could convert it into a vision problem which we would know how to solve. But at the moment, I'm not aware of any techniques that can give you nanometer resolution without destroying it in a live dynamic cell. So, you can see all the interactions. You can take static images at that resolution obviously. Really detailed now and that's quite exciting but it's not enough to turn it just into a complex vision problem. So, that's one way it could be solved. So, it could be a hardware driven data driven solution or it could be that we build better learn simulators of these dynamical systems. So, that's the more modeling way of solving it.

18. 五年内最具变革科学领域 Most Transformative Scientific Domain in 5 Years

Host

你一直在关注各种科学领域,不仅仅是生物。还有科学、药物发现、气候建模、数学。如果让你排序,未来 5 年哪个科学领域会发生最剧烈的变革,你的清单上有什么?

You've been looking at all kinds of science and not just bio. There's science, drug discovery, climate modeling, mathematics. If you had a rank which scientific domain will transform the most dramatically the next 5 years, what's in your list?

Demis Hassabis

所有这些听起来都很令人兴奋,这就是为什么,对我来说,这是我主要的热情所在,也是我整个职业生涯 30 多年来一直致力于 AI 的原因——将 AI 作为终极工具。我一直认为 AI 将是科学的终极工具,能够带来如此先进的科学理解、科学发现,以及医学等领域,还有我们对周围宇宙的理解。所以,实际上,当你提到我们最初阐述使命宣言的方式时,我们仍然这样思考,有两个步骤。第一步是解决智能问题,即构建 AGI,第二步是用它来解决其他一切。随着时间的推移,我们不得不稍微改变这一点,因为人们会问:“你真的意思是解决其他一切吗?”我们确实是这个意思,我认为今天人们开始理解这意味着什么。但具体来说,我指的是解决我称之为科学中的根节点问题。也就是那些能够开启全新分支或发现途径的科学领域。AlphaFold 就是我们想做的典型例子。现在全球超过 300 万研究人员,几乎每个生物学研究人员都在使用 AlphaFold。我的一些前高管朋友告诉我,从现在开始,几乎每一种药物在研发过程中都会在某个阶段使用 AlphaFold。这是我们非常自豪的事情,也是我们希望 AI 带来的那种影响。但我认为这仅仅是开始。我真的看不到任何科学或工程领域是 AI 无法帮助的。你提到的那些领域,我认为我们几乎处于类似 AlphaFold 的初期阶段。我们已经有非常有希望的结果,但还没有完全解决该领域的重大挑战。但我认为,在未来几年里,我们会在你提到的所有领域有很多可谈的,从材料科学(我认为非常令人兴奋)一直到数学。

All sound exciting and that's why, for me, that has been my main passion and always the reason why I've worked on AI for my whole career for 30 plus years now is to use AI as the ultimate tool. I always thought AI would be the ultimate tool for science and to invite such advanced scientific understanding, scientific discovery, and things like medicine, and just our understanding of the universe around us. So, actually, when you mentioned our original way we used to articulate our mission statement, which is still the way we think about it, is there was two steps to it. One was Step one was solve intelligence, i.e., build AGI, and then step two was use it to solve everything else. We had to change that a bit over time because people were like, "Do you really mean solve everything else?" And we did mean that, and I think people are sort of understanding what that means today. But, specifically, I was solve other what I call root node problems in science. So, areas of science that would unlock whole new branches or avenues of discovery. And AlphaFold is the prototypical example of what we want to do. So, over 3 million researchers around the world, pretty much every biology researcher in the world uses AlphaFold now. And I was told by some of my former executive friends that almost every drug discovered from now on will have used AlphaFold at some point in the drug discovery process. So, that's something we're very proud of, and it's the sort of impact that we hope to have with AI. But, I do think it's just the beginning. I don't really see any area of science or engineering that this won't be able to help be helpful with. And the ones you mentioned, I think we're almost like an AlphaFold one moment. So, we've got very promising results, but it's not quite solved the grand challenge yet in that domain. But, I think we're going to have a lot to talk about in the next couple of years on all those areas you mentioned, materials, which I think is very exciting, all the way to mathematics.

19. 普罗米修斯性质与责任使用 Promethean Nature and Responsible Use

Host

在科学领域,我的意思是,这感觉像普罗米修斯式的。就像,这里有这种能力,然后……

In science, I mean, it feels Promethean. It's like, here is this capability, and...

Demis Hassabis

我也这么认为。我的意思是,当然,与此同时,包括普罗米修斯的寓言,我们也必须小心如何使用它、用它做什么,以及这些工具可能带来的滥用。

I think so. I mean, of course, along with that, including what the parable of Prometheus, we have to also be careful with how we use that and what we use it for, and also the misuse that can happen with those same tools.

20. 给AI科学初创公司的建议 Advice for AI for Science Startups

Host

在座的很多人都在尝试建立将 AI 应用于科学的公司。对他们来说,在你看来,真正推进前沿的初创公司与那些只是把 API 包装在基础模型上并称之为“AI for Science”的初创公司有什么区别?

A lot of people in this room are trying to build companies applying AI to science. For them, what's the difference between a startup that actually advances the frontier in your view versus one that's just wrapping an API around a foundation model and calling it AI for science?

Demis Hassabis

嗯,看,我认为有一件事我会推荐。我在思考,我想你之前也跟我提过。如果今天我坐在你在 Y Combinator 的位置上,审视各种事物,我会怎么做?你必须做的一件事是,显然要把握 AI 技术的发展方向。这是困难的一部分。但我确实认为,将 AI 的发展方向与其他深度技术领域结合起来有巨大的空间。我认为那个甜蜜点就是材料、医学或其他真正困难的科学领域。我认为那些跨学科团队,尤其是涉及原子世界的团队,至少在可预见的未来,没有捷径可走。这些领域相对安全,不会被基础模型的下一次更新所淹没。所以,我认为如果你在寻找这样的东西,那是我所说的更具防御性的领域之一。我一直热爱深度技术,所以我有点偏向深度技术的东西。我认为任何真正持久且有价值的事情都不容易。所以我总是被深度技术所吸引。显然,AI 在 2010 年我们刚开始时就是那样的,对吧?投资者告诉我,他们认为这行不通,甚至在学术界,它也被认为是一个非常小众的学科,我们在 90 年代尝试过,知道它行不通。

Well, look, I think there's one of the things I would recommend. I'm trying to think about and I think you mentioned this to me before. What would I do today myself if I was sitting in your place in Y Combinator, you know, looking at things. One thing you have to do is obviously intercept where the AI tech is going. So, that's one hard part of it. But, I do think there's huge scope for combining where AI is going with some other deep technology area. I just think that that sweet spot is whether it's materials or medicine or other really hard areas of science. I think that those kinds of interdisciplinary teams, especially if it involves the world of atoms as well, there's not going to be a shortcut to that, at least in the foreseeable future. Those areas that are pretty safe from just getting swamped by whatever the next update is to the foundation models. So, I think if you're looking for things like that, that's one of the more defensible areas I would say. And I've always loved deep tech, so I'm kind of biased towards deep tech things. I think nothing that's really long-lasting and worthwhile is easy. And so, I'm always been drawn to deep technologies. Obviously, AI was like that back in 2010 when we started out, right? It was thought to just we know it doesn't work kind of thing is what I was told by investors and even in academia it was considered to be a very niche subject that we sort of tried in the '90s and we know doesn't work.

21. AI的信念与热情 Belief and Passion in AI

Host

但如果你对自己的想法有信念和信心,相信这次为何不同,或者你背景中有何种特殊组合,理想情况下你在机器学习和应用领域都是专家,或者你能组建一个拥有这种专业知识的创始团队,我认为那里可以产生巨大影响,也能创造巨大价值。

But, if you have belief and conviction in your idea why it's different this time or what special combination from your background that you had, ideally you're expert in both those areas, both the machine learning and the other area you're applying it to or you can create a founding team with that expertise, I think there's huge impact to be made there and huge value to be built there.

Demis Hassabis

这是个重要的信息。我的意思是,人们很容易忘记。基本上,一旦你做到了,你就做到了。但在你做到之前,人们都反对你。

That's an important message. I mean, it's easy to forget. Basically, once you've done it, you've done it. But before you've done it, people are arrayed against you.

Host

哦,当然。我的意思是,没人相信它,这就是为什么我认为你必须做你真正热爱的事情。对我来说,无论发生什么,我都会研究 AI。我很小的时候就认定,这是我能想到的最重要的事情。结果确实如此,但也可能不是。也许我们会早 50 年。而且这也是我能想到的最有趣的工作。所以,即使我们现在还在某个小车库里,AI 还没完全成功,我今天也仍然会研究 AI。我可能还会尝试找到方法,也许回到学术界或别的什么,但我会找到某种方式继续研究它。

Oh, sure. I mean, no one believes in it, which is why I think you've got to work on things that you're genuinely passionate about. Like, for me, I would have worked on AI no matter what happened. I just decided from a very young age it was the thing that could be the most consequential thing I could think of. It's turned out that way, but it might not have. Maybe we would have been 50 years too early. And it was also the most interesting thing I could think of working on. And so, I would have still been working on AI today even if we were still in a little garage somewhere and it still wasn't quite working. I would have still been trying to find maybe I'd have been back in academia or something, but I would have found some way of continuing to work on it.

22. AlphaFold式突破模式 Pattern for AlphaFold-style Breakthroughs

Host

那么,AlphaFold 就像是你追求的一个尖峰突破,并且成功了。是什么让一个科学领域成熟到可以出现 AlphaFold 式的突破?有没有一种模式,某种目标函数?

So, I mean, AlphaFold was like an example of a spike that you pursued and it worked. You know, what makes a scientific domain ripe for an AlphaFold style breakthrough? And is there a pattern, a certain objective function?

Demis Hassabis

我的看法是,我应该找个时间把这个写下来,但我在所有 Alpha 项目中学到的教训,特别是 AlphaGo 和 AlphaFold,就是:如果问题可以描述为巨大的组合搜索空间,那么我们的技术和我要找的问题就很合适。从某种意义上说,空间越大越好。所以,没有暴力算法或特例算法能解决它。围棋棋步和蛋白质的不同构型都是如此,两者都远超宇宙中的原子数。然后,你有一个清晰的目标函数。所以,你可以把它看作最小化蛋白质的自由能,或者赢得围棋比赛。因此,你需要能够清晰地指定目标函数,以便进行爬山。然后,有足够的数据或模拟器,可以生成大量分布内的合成数据。如果这些条件成立,那么我认为用今天的方法,你可以走得很远,去解决并找到你需要的那个大海捞针般的解决方案。顺便说一句,我认为药物发现也是如此,对吧?存在一种化合物可以治愈这种疾病,只要你能找到它,对吧?而且它没有副作用等等。只要物理定律允许,问题就是如何高效、可处理地找到它。我认为我们实际上第一次用 AlphaGo 证明了,这些系统可以找到那种大海捞针,在这种情况下就是完美的围棋棋步。

The way I see it, I should write this up at some point when I have 5 minutes spare, but the lesson I've learned from all the Alpha projects we've done, specifically AlphaGo and AlphaFold, is that the techniques we have and the problems I look for are great if the situation can be described as a massive combinatorial search space. The more massive the better in some ways. So, no brute force or special case algorithm will solve it. And that's true of Go moves and of different configurations of proteins, far more than the atoms in the universe, both of those. And then, you have a clear objective function. So, you can think of it as minimizing the free energy in the proteins or winning the game of Go. So, you need to be able to specify your objective function clearly so you can hill climb. And then, enough data and or a simulator that can generate you lots of in-distribution synthetic data. If those things are true, then I think with today's methods, you can go a long way into tackling and finding the kind of needle in the haystack that you need for the solution you're trying to look for. And I think of drug discovery, by the way, in the same way, right? There is a compound out there that would solve this disease if one could find it, if one could only find it, right? And that wouldn't have any side effects and so on. And as long as the laws of physics allow it, then the question is how do you find it in an efficient way, in a tractable way? I think we showed for the first time, actually, with AlphaGo, that these systems could find those kinds of needles in a haystack, in that case, the perfect Go move.

23. AI用于科学推理 AI for Scientific Reasoning

Host

我想稍微元一点,我们在谈论人类用这些方法创造 AlphaFold,但还有一个元层次,即人类用 AI 探索可能假设的空间。我们离能做真正科学推理(而不仅仅是数据模式匹配)的 AI 系统还有多远?

I guess to get a little meta, I mean, we're talking about humans using these methods to create AlphaFold, but then there's a meta level, which is humans using AI to explore the space of possible hypotheses. How close are we to AI systems that can do genuine scientific reasoning, not just pattern matching on data?

Demis Hassabis

我们很接近了。我们正在研究这类通用系统,比如我们有一个叫 Co Scientist 的系统,还有其他算法如 AlphaFold,它们能比基本的 Gemini 做得更多一点。显然,所有前沿实验室都在这样实验。到目前为止,我还没看到任何东西,我们都在摆弄同样的东西,比如一些比 IMO 稍难的数学问题等等。我还没看到任何真正意义上的重大发现。这是我个人的看法。我认为它正在到来。我认为这可能与我们之前讨论的创造力有关,即真正超越已知的边界。所以,显然那不仅仅是模式匹配,因为没有模式可匹配,而且它比外推更多一些。这是某种类比推理,我不认为这些系统具备这种能力,或者至少我们没有以正确的方式使用它们来做到这一点。所以,我在科学中经常这样说:它能否提出一个真正有趣的假设,而不仅仅是解决一个?当我说“仅仅”时,我们不是在说仅仅解决黎曼假设之类的问题。那显然会很惊人,或者是千禧年大奖难题之一,也许我们离那还有几年。但是,我想解决 P 是否等于 NP。那是我最喜欢的一个。但是,你能吗?比那更难的是提出一套新的千禧年大奖难题,被顶尖数学家认为深刻、有意义,值得花一生去研究和解决。对吧?我认为那又难了一个层次。我们还没有,我仍然认为我们不知道如何做到。不过,我不认为这是魔法。我确实认为这些系统最终能够做到。也许我们缺少一两样东西。然后,我们测试它的方法,我有时称之为我的爱因斯坦测试,即:你能训练一个知识截止于 1901 年的系统,然后它能否提出爱因斯坦在 1905 年所做的工作,包括狭义相对论,他的奇迹年?它能做到吗?然后,我认为我们可以运行这个测试。也许我们应该运行这个测试,不断看看是否可能。一旦做到了,那么我认为这些系统就处于能够发明新东西、真正新颖的东西的边缘了。

We're close. We're working on these general systems like that, like I think we have this system called Co Scientist, and we have other algorithms like AlphaFold that can go a little bit beyond what the basic Gemini will do. And obviously, all the frontier labs are experimenting in this way. I've yet to see anything so far, and we all tinker with the same things, you know, some math problems that are a little bit harder than IMO and so on. I haven't seen anything yet that is a true genuine, you know, massive discovery. That's my personal opinion. I think it's coming. I think it may be related to this earlier thing we discussed about creativity, and actually going beyond the bounds of what's known. So, clearly, that's just not pattern matching at that point, because there is no pattern to match to, and it's a bit more than extrapolation. It's some kind of analogical reasoning, and I don't think these systems have that, or at least we're not using them in the right way to do that. So, the way I often say that in science is: can it come up with a hypothesis that's really interesting, not just solve one? When I say "just", we're not talking about just solving the Riemann hypothesis or something. This would be obviously amazing, or one of the Millennium Prize problems, and maybe we're a couple of years out from doing that. But, I'd like to solve P equals NP. That's my favorite one. But, can you? Even harder than that would be to come up with a new set of Millennium Prize problems that were regarded by top mathematicians to be as deep and meaningful and worthy of a lifetime of study and effort to solve. Right? I think that's another level harder. And we don't have, you know, I still don't think we know how to do that. I don't think it's magical, though. I do think these systems will eventually be able to do that. Maybe we're missing one or two things. And then, the way we would test that is, I sometimes call it my Einstein test, which is: can you train a system with the knowledge cutoff of 1901, and then will it come up with what Einstein did in 1905, including special relativity, his annus mirabilis? Can it do that, right? And then, I think we could run that test. Maybe we should just run that test and keep seeing if that's possible. And once that is, then I think we're on the verge of these systems being able to invent something new, truly novel.

24. 结束致谢 Closing Thanks

Host

那么,最后一个问题。对于在座的技术人员,他们想从事接近你所创造规模的工作,这是世界上最大的 AI 努力之一,而你多年来一直是先驱。为此,我想在座的每个人都衷心感谢你和 DeepMind 的伙伴们。谢谢。

So, last question. For the people who are deeply technical in this room who want to work on something, you know, even close to the scale that what you have created with, you know, it's one of the largest AI efforts in the world, and you've been a pioneer for all these years. So, for that, I think everyone in this room thanks you and the folks at DeepMind very, very deeply from the bottom of our hearts. Thank you.

25. 给年轻建设者的建议 Advice for young builders

Host

关于前沿构建,你现在知道但希望 25 岁时就明白的事情是什么?

What's the thing that you know now about building at the frontier that you wish you'd known at 25?

Demis Hassabis

我觉得我们之前已经谈到了一些,实际上你会发现,攻克困难而深刻的问题,在某些方面并不比追求浅显、简单、表面化的问题更难。它们只是难点不同而已。每类问题都有各自难的地方,但考虑到生命短暂,你的时间和精力有限,不如把你的生命力投入到那些如果你不做、如果你不在那里推动,就真的会有所不同的事情上。所以,我会从这个角度去思考。另一件事是,如果你从事深度科技,我喜欢跨学科工作,我认为未来几年这会更普遍——不同领域的组合,以及找到这些领域之间的联系。而有了 AI,做这件事会更容易。我唯一想补充的是,取决于你对 AGI 时间线的判断——我自己的判断大概是 20 到 30 年左右——如果你今天开始一段深度科技之旅,在我看来,真正的深度科技通常需要 10 年时间。那么,你就必须考虑 AGI 在这段旅程中途出现。这意味着什么?这不一定坏事,但你必须把它考虑进去,对吧?它能否利用你的成果?AGI 系统会用它做什么?这又回到了你之前提到的 AlphaFold 和通用 AI 系统的问题。我能看到的一种情况是,Gemini、Claude 或某个通用系统把类似 AlphaFold 的专用系统当作工具来使用。我不认为我们会把所有东西都塞进一个巨型大脑里,因为如果把所有蛋白质数据都放进 Gemini,那会导致太多回归,这没有意义。我们不需要 Gemini 来做蛋白质折叠。回到你之前说的信息效率问题,这肯定会损害它的语言能力或其他方面,对吧?而且是负面的。所以,我认为更好的方式是拥有非常好的通用工具使用模型,它们甚至可以训练那些专用工具,但那些工具是独立的系统。所以,我认为思考这些影响以及你今天应该构建什么,是件很有趣的事。还有物理层面的东西,比如你会建造什么样的工厂,什么样的金融系统等等。所以,我认为你真的需要认真对待这件事,一方面想象那个世界会是什么样子,然后构建一些如果 AGI 在半路出现时仍然有用的东西。

I think we covered some of it in terms of actually you work out that going after hard problems and deep problems is no more difficult in some ways than going after a shallower, simpler, more superficial problem. They're just differently difficult. There's different things that are hard about each of those things, but I think given life's very short and you know, you only have so much time and energy, you might as well put your life force into something that will really make a difference if you hadn't done it, if you hadn't been there to push it. So, I would just think of it through that lens. And then the other thing is if you're into deep tech and I love interdisciplinary work and I think that's going to be even more prevalent in the next few years in combinations of fields and finding the connections between those fields. And it's going to be even easier to do that with AI. And then the only other thing I would say is, depending on what your AGI timeline is, you know, mine's like 20 or 30 years or something like this, then if you start off on a deep tech journey today, usually that's a 10-year journey for true deep tech in my opinion. So, then you have to consider AGI appearing in the middle of that journey. So, what does that mean? It's not bad necessarily, but you have to take that into account, right? Will it be able to leverage it? What will the AGI system do with it? And it goes a little bit back to what you said earlier about AlphaFold and general AI systems. So, one thing I can see happening is Gemini, Claude, or one of these general systems making use of AlphaFold-like specialized systems as tools. I don't think we're going to have it just in one giant brain because it will have too much regression if I put all the proteins into, you know, Gemini, that wouldn't make sense. We don't need Gemini to do protein folding. Going back to your information efficiency, it will definitely affect its language skills or something like that, right? In a bad way. So, much better I think is to have really good general purpose tool usage models that will then maybe they could even train those specific tools, but they would be in a separate system. So, I think that's kind of interesting to think through the implications of that and then what you might build today. Also, physical things too like what kinds of factories would you build, what sorts of finance systems and so on. So, I just think you need to really take that seriously and on the one hand imagine what that world would look like and then build something that would be useful if that comes in halfway through.

互动版:逐字朗读 + 针对本期提问 →