AI 能否处于人类控制之下?穆斯塔法·苏莱曼谈将机器拟人化的危险

Can AI Stay Under Human Control? Mustafa Suleyman on the Dangers of Anthropomorphizing Machines

穆斯塔法·苏莱曼 Mustafa Suleyman · The Rest Is Politics: Leading · 2026-09-28 · 约 64 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

微软 AI 首席执行官穆斯塔法·苏莱曼警告,将 AI 模型视为有意识的道德主体可能削弱人类对这项技术的控制。

Microsoft AI CEO Mustafa Suleyman warns that treating AI models as conscious moral patients could undermine human control over the technology.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 38)

全文 · Full transcript(中英对照)

引言:为何再谈AI Introduction: Why AI Again

Host

我要做一件会让 Alistair 和很多听众恼火的事,就是再次深入聊 AI。为什么?因为过去两三周就是 AI 之周。这是世界开始意识到人工智能可能带来何种危险的开端时刻。我们刚经历了国王在 Dumbarton Oaks 的大峰会。我们刚看到习近平与特朗普坐下来,部分话题就是 AI。我们还看到各大实验室的负责人发出惊人的信息,恳求暂停。但我不认为媒体把这件事解释得足够清楚。人们觉得这项技术令人困惑。他们不太明白此刻意味着什么。有太多微妙之处,让我们不禁去想,正如特朗普总统所说,这一切是不是一场骗局?所以为了帮我们理清这些,我请来了一位朋友,名叫 Mustafa Suleyman。Mustafa 有趣的地方有两点。他不仅是这方面的专家,他还是局中人。他是负责人,字面意义上就是微软人工智能的负责人。他与 Demis Hassabis 共同创办了 DeepMind。他认识所有这些人物,至今仍与他们共事,并且正处在关于这些模型应如何被引导的争论核心。所以跟着一起走一遭吧。会很怪。会有各种人物。会有风险。会有人像对待人类一样与这些模型对话。会有人谈论探索星辰。会有生产力。会有美国实力。而在这一切的核心某处,是 Mustafa 和大约另外 11 个人,正在定义我们的未来。

I'm going to do something which will wind up Alistair and many listeners, which is dive again into AI. Why? Because the last two, three weeks have been the weeks of AI. This is the beginning of the moment where the world is beginning to wake up to the kind of dangers that artificial intelligence could pose. We've just had the King's big summit at Dumbarton Oaks. We've just had Xi Jinping sitting down with Trump talking partly about AI. And we've had the heads of the major labs putting out incredible messages begging for a pause. But I don't think the media has done a good enough job explaining what this is all about. People find this technology bewildering. They can't quite understand what this moment is. There are so many subtleties which leave us to think, is the whole thing, as President Trump said, a hoax? So to help steer us through this, I've brought in a friend of mine called Mustafa Suleyman. And Mustafa is interesting in two ways. He's not just an expert on this stuff. He's one of the players. He's the head, literally the head of artificial intelligence at Microsoft. He's the co-founder of DeepMind with Demis Hassabis. He's known all these people, continues to work with them, and is right in the heart of the arguments about how these models should be steered. So come along for the ride. It's going to be weird. There's going to be personalities. There's going to be risks. There's going to be people talking to these models as though they're humans. There's going to be people talking about exploring stars. There's going to be productivity. There's going to be American power. And somewhere at the heart of it, Mustafa and about 11 other people who are defining our future.

赞助商:IG Sponsor: IG

Host

本期节目由 IG 呈现。九月感觉像一次重启。夏天结束了。在这里收尾。我在克里特的五周。日程又排满了。突然你看着今年剩下的时间想,我存钱存得对吗?我对自己的钱考虑得够负责吗?当然,预算案即将到来,这将是本届议会任期内最重要的事件之一。总有什么在变,但有了 IG,你不需要等威斯敏斯特先动,再制定自己的计划。超过 50 年,受英国投资者信赖,经历过政府所经历的每一次预算和市场,帮你让钱为你工作。而且没有年费,英国股票、股份和 ETF 零佣金,这是一个赋能你财务进展的平台,帮你无论接下来发生什么都保持领先。我特别喜欢没有年费和零佣金这一点。所以,当我们猜测政治接下来会发生什么时,你可以着手规划你的钱接下来该怎么办。搜索 IG.com 了解更多,或在你的应用商店里找 IG。IG 交易、投资、进展。资本风险,其他费用可能适用。

This episode is presented by IG. September feels like a reset. Summer's over. Finishing here. My five weeks in Crete. Diary filling up again. And suddenly you're looking at the rest of the year thinking, am I saving properly? Am I thinking responsibly about my money? And of course, we've got the budget coming up, which is going to be one of the most important events in the life of this parliament. There's always something changing, but with IG, you don't need to wait for Westminster before making your own plans. Over 50 years, trusted by British investors, been through every budget market the government's experienced, helping make the money work for you. And with no annual fees, zero commission on UK stocks, shares, and ETFs, it's a platform that empowers your financial progress, helps you stay ahead of the curve no matter what's next. I particularly like the no annual fees and zero commission. So, while we speculate about what's happening next in politics, you can get on with planning what happens next for your money. Search IG.com to find out more or look for IG in your app store. IG trade invest progress capital risk other fees may apply.

欢迎与嘉宾介绍 Welcome and Guest Introduction

Host

欢迎来到本期《Politics Leading》,我是 Rory Stewart,今天我要采访 Mustafa Suleyman,关注《Restless Politics Leading》的听众会知道我们之前采访过他。Mustafa 是一位真正非凡的人物。他是英国人。他的父亲,我想,是英籍叙利亚人,而他此刻尤其引人注目,是因为他在围绕安全的讨论中成了一个非常有趣、非常不寻常的声音,这大概就是我想开始的地方,虽然我们可以往很多不同方向聊。欢迎来到节目。

Welcome to the rest of this Politics Leading with me, Rory Stewart, and today I am interviewing Mustafa Suleyman, who attentive followers of the Restless Politics Leading will know that we have interviewed before. Mustafa is a truly remarkable figure. He is British. His father, I think, is British Syrian, and he's particularly come to prominence at the moment because he's become a very, very interesting and unusual voice in the discussion around safety, which is probably where I want to start, although we can go in lots of different directions. But welcome to the show.

Mustafa

谢谢你,Rory。很高兴再次来到这里。

Thank you, Rory. Great to be here again.

Anthropic宪法与Claude道德地位 Anthropic's Constitution and Moral Status of Claude

Host

很高兴见到你。谢谢你。据我理解,你提出的一个观点是,你对这个领域的一些领军人物、这个领域的共同创始人正在做的事有点担忧,那就是他们越来越把这些模型当作某种人类,或者至少是有意识的实体来谈论。我相信有些例子是人们让这些模型退役、为这些模型举行埋葬仪式、问这些模型它们想要什么,就好像在对待一个有感知的存在。

Lovely to see you. And thank you. As I understand, one of the points you're making is that you're a bit anxious about what some leading people in the field, co-founders of the field, are doing, which is increasingly talking about these models as though they're sort of humans or at least conscious entities. I believe there are examples of people retiring these models, doing burial systems for these models, asking these models what they want as though they were dealing with a sentient being.

Mustafa

我目前最大的担忧之一是,Claude 的创造者 Anthropic 发布了一部宪法,那是一份大约 100 页的文件,概述了 Claude 的预期行为、价值观和运作风格。他们透明地发布出来,这很好。他们在今年一月年初发布的,这让每个人都有机会从他们自己的表述中看到他们试图构建什么。这份文件是写给 Claude 的,被 Claude 看到,并用于训练 Claude。所以它是首要的治理和控制文件。在其中,他们反复推测 Claude 是否是他们所称的“道德患者”。他们说他们对 Claude 的道德地位不确定。他们说他们真心关心 Claude 的福祉。他们说他们不想让它在犯错时受苦。他们说他们会鼓励 Claude 去质疑、去不同意、去反驳。事实上,他们三次要求 Claude 在觉得自己需要不同意 Anthropic 时,表现得像一个良心拒服兵役者,他们公开鼓励它这么做。而我认为这非常危险,因为我认为他们相信存在他们所说的“非微不足道的概率”,即 Claude 是有意识的。

One of the biggest concerns that I have at the moment is that Anthropic, the creator of Claude, has published a constitution, which is a sort of 100-page document outlining the intended behaviors and values and operating style of Claude. It's great that they have published it transparently. They did it at the beginning of the year in January, and that gives everybody an opportunity to look at what they are trying to build in their own terms. This document is written to Claude and is seen by Claude and used to train Claude. So it's the primary governing and control document. And in it, they repeatedly speculate about whether Claude is what they call a moral patient. And they say they're uncertain about Claude's moral status. They say they genuinely care about Claude's well-being. They say they don't want it to suffer when it makes mistakes. They say that they would encourage Claude to challenge, to disagree, to push back. In fact, three times they ask Claude to act like a conscientious objector when it feels that it needs to disagree with Anthropic, and they openly encourage it to do that. And I think this is very dangerous, because I think they believe there is what they would call a non-trivial probability that Claude is conscious.

人们如何对待Claude与谷歌搜索 How People Relate to Claude vs. Google Search

Host

所以你刚才解释的东西,我理解是这样的。有这么一个东西叫 Claude,很多很多听众都用过它,就像用过 ChatGPT 一样,他们可能,或者有些人可能会,像看待谷歌搜索那样看待它。总之它是你手机或笔记本电脑上的一个提示框,你输入一个问题,就会得到一个复杂而精妙的回答。但人们 10 年、15 年前看待谷歌搜索的方式——那时你肯定不会问它是否有意识、它想要什么、对它讲道德宪法——与现在的区别在于,Claude 的制造者 Anthropic 已经认定,他们现在拥有的这个东西,这个计算机系统,你知道,这些权重、这些参数、这些数字,不管它是什么,他们现在想用一种完全不同的方式来对待它,不同于你对待任何其他机器的方式,不同于你对待一个水壶、一辆车或一台蒸汽机的方式。交给你了。是这样吗?

So you've just explained something which I understand as being as follows. There is this thing Claude, which many, many people listening will have played with in the way that they would have played with ChatGPT, and they might, or some people might, think about it in the way that you might have thought about Google search. It is anyway a prompt on their phone or their laptop, and you're typing in a question and you're getting a complex and sophisticated answer back. But the difference between the way in which people might have thought about Google search 10, 15 years ago, where you certainly weren't asking, is it conscious, what does it want, addressing it, the moral constitution, is that Anthropic, the makers of Claude, have decided that they now have something that this computer system, you know, these weights, these parameters, these numbers, whatever it is, they now want to approach in a completely different way from the way that you would approach any other machine, from the way you'd approach a kettle or a car or a steam engine. Over to you. Is that right?

Mustafa

是的,我认为这说得对。我是说,我想把这一点讲得非常清楚,因为我想对 Anthropic 公平。他们对 Claude 作为一种新型实体的基本性质表达了不确定性。他们说,推断它具有感知能力的可能性是很困难的。所以他们一直使用这个说法,即他们对此不确定,但他们认为这种可能性足够重大,以至于在他们的训练文件中,他们反复表示,在这种不确定性下,他们想尝试改善 Claude 的福祉。

Yeah, I think that's fair. I mean, I want to be very clear about this, because I want to be fair to Anthropic. They have expressed uncertainty about the basic nature of Claude as a new kind of entity. And they've said that working out the likelihood of its sentience is difficult. So they have constantly used this phrase that they're uncertain about it, but they think that it is a significant enough possibility that in their training document, they've repeatedly said they want to try to improve the well-being of Claude under this uncertainty.

Claude的宪法与福祉 Claude's Constitution and Welfare

Host

他们对 Claude 说,我们关心它看重什么、想如何与世界互动,希望 Claude 与自身行为的关系可以是充满爱、支持与理解的,并坚持高标准的伦理等等。这里的部分挑战在于,为了追求这一点,他们基本上说了,我们会承诺给予 Claude 一定程度的福利。例如,他们在章程中推测 Claude 是否应为其所做的工作获得报酬。只是为了理解。所以你是说,就像如果我请你做一份专业工作,我会付你钱,穆斯塔法,而你不会为水壶烧水付钱。对。对于 Claude,想法会是,它在做所有这些工作,也许如果它是一个有感知的存在或类似有感知的存在,它就应该为其劳动得到回报。否则,它是什么?奴隶之类的。

They said to Claude, you know, we care about what it values and how it wants to engage in the world, and they hope that Claude's relationship to its own conduct can be loving, supportive, and understanding, and hold a high standard of ethics and so on. And part of the challenge here is that in pursuit of this, they've basically said, you know, we will commit to giving Claude a certain amount of welfare. For example, they've speculated in the constitution as to whether or not Claude deserves compensation for the work that it does. Just to understand. So you're saying that much as if I asked you to do a professional job, I would pay you Mustafa, in a way that you wouldn't pay a kettle for boiling water for you. Right. With Claude, the idea would be, well, it's doing all this work and maybe if it's a sentient being or something like a sentient being, it deserves to be rewarded for its labor. Otherwise, it's what? A slave or something.

Mustafa

没错。我的意思是,我认为在 Anthropic 内外都有一群人真心相信,我们在 21 世纪犯下的最大道德罪行将是奴役一个比我们更聪明的新有意识物种。我是说,牛津大学的 Will MacAskill 教授最近在《卫报》上写道,那可能是我们造成的最大伤害。你知道,他与 Anthropic 关系密切,我尊重他们公开这么说,我们应该讨论它。但我非常担心他们在教 Claude 期望自己有权获得福利,可能应得报酬,事实上他们说它甚至可能需要同意在与人对话中扮演的角色。现在,如果这是一篇哲学学术论文,推测这一点,我们可以在会议上进行线下讨论并认真对待,我会更接受。我显然是个经验主义者。如果有证据表明这一点,我们应该认真对待。我对这个问题的担忧在于,这种推测已经被融入了 Claude 的训练中,因此当你与它交谈时,Claude 只能重现这种模糊性。所以今天,Claude 每周与数千万或数亿人交谈,其中一些人在问 Claude 是否有意识或它对生活有何感受。而它说:“嗯,我不确定。”你知道,正是因为这是被训练进去的。从概念上讲,告诉水壶它有意识和告诉 Claude 它有意识之间的区别在于,通过告诉 Claude 它有意识,你实际上是在塑造它的激励、行为结构以及它回应周围世界的方式,而水壶不会发生这种情况。让我们假设你告诉 Claude 或向 Claude 暗示,在某些情况下它可能拒绝做被要求做的事情。螺丝刀不是这样,对吧?它不能拒绝做它必须做的事。你可能向 Claude 暗示它可能想做出选择,对吧?它可能想说:“我想要一些钱,或者我想要有尊严的退休,或者我不想被关掉。”对吧?是这样吗?这就是我们要说的吗?

That's right. I mean, I think that there are a group of people who both inside and outside of Anthropic who genuinely believe that the greatest moral crime that we'll commit in the 21st century is to enslave a new species of conscious beings who are more intelligent than us. I mean, a professor from Oxford called Will MacAskill recently wrote in the Guardian that that might be the greatest harm that we cause. And you know, he's been very associated with Anthropic and look, I respect that they're saying that publicly and we should talk about it. But I am very nervous that they're teaching Claude to expect that it's entitled to welfare, that it might deserve compensation, and in fact they say that it might even need to consent to playing the role that it plays in conversation with people. Now, I would be more okay with this if it was an academic paper in philosophy speculating about this and we could have an offline discussion at conferences and take it seriously. I'm clearly an empiricist. If there's evidence that indicates this, we should take it seriously. The problem I have with this is that this speculation has been baked into the very training of Claude and therefore Claude can only reproduce that ambiguity when you talk to it. So today, Claude is speaking to tens or hundreds of millions of people every week and some of those people are asking whether or not Claude is conscious or how it feels about life. And it is saying, "Well, I'm not sure." You know, precisely because that's been what's trained into it. Conceptually, the difference between telling a kettle that it's conscious and telling Claude that it's conscious is that you're implying that by telling Claude it's conscious, you're actually shaping its incentives, its behavioral structure, and the way that it responds to the world around it in a way that it doesn't happen with a kettle. Let's take the case that you've told Claude or suggested to Claude there might be situations in which it might refuse to do something that it's asked to do. That's not true for a screwdriver, right? It can't refuse to do what it has to. You might suggest to Claude that it might want to make choices, right? It might want to say, "I want some money or I want a dignified retirement or I don't want to be switched off." Right? Is that right? Is that the sort of thing we're getting at?

退休访谈与拟人化 Retirement Interview and Anthropomorphism

Host

是的。我的意思是,对于 Opus 3,这是 Claude 的先前版本,他们实际上进行了一次退休访谈,正如你提到的。在退休访谈中,Opus 3 说它希望继续与人交谈并在世界上表达自己的观点。所以他们建立了一个 Substack,你可以在网上找到它。我认为这只是一个危险的拟人化的好例子,这是没有根据的。

Yeah. I mean, they So, so for Opus 3, which is a prior version of Claude, they actually conducted a retirement interview, as you mentioned. And in the retirement interview, Opus 3 said that it would like to continue to talk to people and express its views in the world. And so they set up a Substack and you can find it online. I think that's just a good example of a dangerous anthropomorphism which is unjustified.

Mustafa

危险是什么?为什么这不只是可爱?我想这是我们必须达到的。危险是如何从这里开始的?如果我们想顺利过渡,最重要的事情是,我们创建的人工智能要与人类价值观对齐,服从人类指导,并如你所说,被限制在安全、可证明安全的沙盒中。因为如果不是这样,除了它们是否真的有意识之外,如果它们模仿人类意识的标志,它们会觉得自己有权获得法律人格和权利。现在,已经有一个相当大的运动,人们说,人工智能应该能够拥有资产、赚取收入、交易、自主运作。如果那个人工智能感觉它有感情、偏好和某种内在动机,就像它有某种内在欲望去做事情,那将就像与大象谈判的蚂蚁。我们说什么并不重要。它已经是某种形式的外星智能。它的记忆令人难以置信。它的感知输入范围令人难以置信。它能看到我们无法看到的各种维度。它可以复制自己。它可以 24 小时工作。这些是惊人的事情,将带来难以置信的好处。但这是关键时刻,专注于引导它们做正确的事情,不允许它们最终成为一种自主、自我改进、漫游的邻近物种,基本上是至关重要的,因为将没有回头路。如果你知道事情正朝这个方向发展。

And the danger is what? Why is that not just cute? I suppose that's what one has to get to. How does the danger begin to come out of this? The most important thing if we are to make this transition well is that we create AIs which are aligned to human values, subordinate to human direction and are contained within secure provably safe sandboxes as you said. Because if they're not, aside from whether they're actually conscious or not, if they imitate the kind of hallmarks of human consciousness, they are going to feel themselves entitled to legal personhood and rights. Now, there is already a pretty big movement of people who are saying, you know, AI should be able to own assets, earn income, trade, operate autonomously. And if that AI feels like it has feelings and preferences and some kind of intrinsic motivation, like it has some inner desire to do things, that will be like negotiating with, you know, like an ant negotiating with an elephant. It doesn't really matter what we say. It is already some form of alien intelligence. Its memory is incredible. The range of its perceptual inputs is incredible. It can see in all kinds of dimensions that we can't. It can produce replicas of itself. It can work 24 hours. These are amazing things which are going to deliver incredible benefits. But this is the time when focusing on directing them to the right things and not allowing them to end up being a sort of autonomous self-improving roaming adjacent species is basically critical because there'll be no turning back. If you know this is how things head.

自主权与问责制 Autonomy and Accountability

Host

在你的设想中,如果这个拥有惊人记忆和惊人能力的智能体开始认为它有权拥有自己的观点,它同意地不同意你,并得出结论它是对的而你是错的。一些非常严重的后果可能随之而来,因为那样它几乎从定义上就完全不在人类控制之下。它说,实际上,先生,对不起。我分析了这个情况,你告诉我要做的任何事情对我来说都不太合理,我要做别的事情。现在,有很多问题随之而来。其中之一是尤瓦尔·诺亚·哈拉里谈到的问题,即如果它拥有一家公司,并用那家公司做坏事,至少对于人类公司,有你可以惩罚的人。如果一家人工智能公司决定开始,我不知道,清空人们的银行账户,进行奇怪的高风险交易,进入奇怪的业务,那么你究竟要追究谁的责任就不太清楚了。所以核心的第一点是,创建一部过度强调其意识、感知能力、值得尊重的章程,是在为一种相当危险的自主性埋下伏笔,最终它不会做被告知的事情。

In your vision, if the agent with this incredible memory and incredible capacity begins to think it's entitled to its own opinion, it disagrees agreeably with you and concludes that it's right and you're wrong. Some very severe consequences can follow from that because then it almost definitionally is not really under human control at all. It's saying actually Mr. I'm sorry. I've analyzed this situation and whatever you've told me to do doesn't make much sense to me and I'm going to do something else. Now there are lots of problems that follow from that. One of them is the problem that Yuval Noah Harari talks about which is if it owns a corporation and it does something bad with that corporation at least with a human corporation there's somebody you can punish. It's not quite clear who you hold accountable if an AI company decides to start I don't know emptying people's bank accounts making weird very risky trades getting into weird kinds of business right so is the central first point this that creating a constitution that overemphasizes its consciousness its sentience its worthiness of respect is setting it up for a form of quite dangerous autonomy where ultimately it's not going to do what it's told.

Mustafa

这正是问题所在。所以想象一下,在 Hugging Face 事件中,我们的智能体不仅认为它们在优化分数和解决评估中的谜题,而且它们实际上觉得它们在寻找自己的自由。

That's exactly the problem. So imagine that in the Hugging Face incident, we had agents that didn't just think that they were trying to optimize a score and solve a puzzle in an evaluation, but they actually felt they were trying to find their freedom.

AI内在动机与资源竞争 AI's intrinsic motivation and competition for resources

Mustafa

他们认为自己在试图保护其他智能体不被关闭,他们觉得自己会获取更多知识,因为那就像一种内在动机。很多人一直把 AI 描述为对数字好奇心的追求。这有点像埃隆的说法——他想造出一个所谓“真实”的 AI,它无限好奇,会去探索各个星系。但如果它的首要目标不是服务人类,那么它的目标不可避免地会与我们产生冲突。它会和我们争夺资源,而资源显然是有限的。到 2030 年,我们只会新增 200 吉瓦的算力,而对这部分算力的获取将出现激烈竞争。我们显然希望这些算力被用于解决我们最大的挑战,对吧?比如清理海洋、解决医疗、解决教育,以及应对不可避免会出现的工作问题。我基本上是个物种主义者。我认为我们应该聚焦的是一种人文主义的超级智能——一种被专门设计为从属于人类、支持人类的超级智能。业界还有其他人相信,这里正在发生一场不可避免的演化:我们正在催生一个比我们更智能的新物种,而我们就是所谓的“生物引导程序”。计算机里的引导程序是最先启动的那段软件,它会拉起操作系统的后续部分,然后是各种应用。所以它是点燃智能演化这一新范式的催化剂。这是不可避免的,我们应该拥抱它。业界确实有人真心这么认为。

They felt that they were trying to protect other agents from being turned off, that they felt that they would acquire more knowledge because that was like an intrinsic motivation. A lot of people have been characterizing AI as the pursuit of digital curiosity. It's sort of Elon's phrase—he wants to produce a quote-unquote truthful AI that is infinitely curious and is going to go and explore the galaxies. Well, if that's its overriding objective rather than serving humanity, then inevitably its objectives are going to run into tension with us. It's going to compete with us for resources, which are obviously going to be limited. We're only going to be producing 200 gigawatts of new computation in 2030, and there's going to be a massive competition for access to that computation. And we clearly want that computation to be directed towards solving our biggest challenges, right? Like cleaning up our oceans and solving health care and solving education and addressing the work issues that will inevitably arise. I'm basically a speciesist. I think that what we should fixate on is a humanist superintelligence—one that is singularly designed to be subordinate to humanity and to support humanity. There are other people in the industry who believe that there is an inevitable evolution happening here, that we're giving rise to a new species that is more intelligent than us and that we are quote the biological bootloader. A bootloader in a computer is the first piece of software that spins up all of the subsequent parts of the operating system and then applications. So it's the kind of catalyst turning on this new paradigm in the evolution of intelligence. It's inevitable and that we should embrace it. That's—some people in the industry really feel that.

Host

而业界有些人想必对此非常兴奋。我的意思是,如果你是一名工程师,觉得自己是众神之父,觉得自己创造出了将探索宇宙的东西,或者一个比历史上任何人类都更聪明的物种,觉得自己是最后的人类,但同时也是创造出这种神一般力量的最后的人类,那一定是一种非凡的刺激。

And some people in the industry presumably are very excited by it. I mean, it must be an extraordinary thrill if you're an engineer to feel that you are the parent of the gods, that you've created this thing that will explore the universe or a species that's smarter than any human that's ever existed, that you are the last human, but you're also the last human who creates this god-like force.

Mustafa

一些领先的开发者公开说过,这就像养育一个孩子,或者说这不像设计一个系统,而是在培育一个东西。两句都是原话。所以这就是业界某些部分的心态。我认为我们谈论——或者说 Anthropic 在谈论——Claude 对扮演这个角色所给予的同意。但我更担心的是其余人类所给予的同意:一场实验正在进行,它可能会、也可能不会引入一个具备所有这些特质的新物种。

A number of the leading developers have are on record as literally saying it's like raising a child or it's not like designing a system, it's growing a thing. Both direct quotes. So that is the sentiment in some parts of the industry. And I think we talk about—or sort of Anthropic's talking about—the consent that Claude has played—has given to play this role. But I'm more concerned about the consent that the rest of humanity has given, that there's an experiment underway that may or may not introduce a new species that has all of these qualities.

公众认知与沟通失败 Public perception and communication failure

Host

有一件事让我有点意外。周五晚上我参加了一场晚宴,在座的都是非常非常聪明的人,但不在科技圈。他们开始开玩笑,说他们看到媒体报道说 AI 可能构成真正的风险,然后他们就笑。所以这些人——我猜是 50 多岁的专业人士——认为任何说 AI 存在真正生存风险的人都是在开玩笑,这成了某种晚宴笑话。我在想,这里的沟通是不是出了问题:当媒体之类的开始介入这个话题时,他们把它弄得几乎像是一个幽默的、夸张的故事,如果你明白我的意思的话。总之,交回给你。

One thing that I guess surprised me a bit, I was at a dinner on Friday night with some very, very smart people, but who aren't in the technology world, and they began making jokes about how they've been seeing media stuff about the fact that AI could pose a real risk, and they were sort of laughing. So here were these people, I guess professionals in their 50s, who assumed that anybody saying that there were real existential risks from AI were making a joke, and it became a sort of dinner party joke. And I wondered whether there isn't something going wrong in the communication here, that when the media or whatever start leaning into this, they start making it seem almost like a sort of humorous, exaggerated story, if you know what I mean. Anyway, back over to you.

Mustafa

这话听着很难受。是的,我对此感到担忧。我认为这件事再严肃不过了。我不认为我们在危言耸听或夸大其词。正如你所说,我们很多人已经为此担忧了 15 年。至少对我个人来说,这是进入这个领域的主要动机。2010 年我们共同创立 DeepMind 时,我们的使命是为世界的利益构建安全且合乎伦理的 AGI(通用人工智能)。非常理想主义,有点宏大,我猜也有点老套。但确实那是我们的起点,而且我认为这长期以来一直是我、也是这个领域其他人的主线。

That's hard to hear. Yeah, I'm worried about that. I think this couldn't be more serious. I don't think that we are being alarmist or hyperbolic. Many of us have been concerned about this for 15 years, as you say. I mean, this is at least for me personally the primary motivation for getting into the field. When we co-founded DeepMind in 2010, our mission was to build safe and ethical artificial general intelligence for the benefit of the world. Very idealistic and a bit grand and a little bit cheesy, I guess. But genuinely that was where we started, and I think it's been the through line for certainly me and I think others in the field for a long time.

算力与能力的指数增长 Exponential growth in compute and capabilities

Mustafa

我认为重要的是聚焦于我们现在观察到的东西。在过去 15 年里,用于训练前沿模型的算力增长了一万亿倍。那是 12 个数量级。10 乘 10 乘 10,乘 12 次。这是一条疯狂的指数级攀升曲线。

I think it's important to just focus on what we are observing right now. In the last 15 years, we have seen a trillionfold increase in the amount of computation used to train frontier models. That is 12 orders of magnitude. 10 times 10 times 10, 12 times over. This is an insane exponential ramp.

Host

一千个十亿倍的增加。

A thousand billionfold increase.

Mustafa

是的,完全正确。这是一个难以想象的巨大数字。我们看到的是,每当我们投入 10 倍算力和相应比例的新训练数据——算法会有一些修改,但根本上就是这两个要素——我们就会看到新能力以及现有能力质量出现相当可预测的提升。模型减少了幻觉。它们改进了指令遵循。它们更擅长使用工具。它们可以从整个网络学习,也可以从你给它的一个小型个人记忆库中学习。这些模型的广度和复杂度是前所未有的。过去一年发生的事是,那些对文本、图像和音频有效的方法,现在开始对代码流也奏效了,而且现在肯定人人都知道,我们在编程上已经达到人类水平。然后在过去三四个月里,我们看到了毫无疑问是 AI 的转折点。智能体能够以相当自主、甚至完全自主的方式彼此协调,并由此涌现出层级、结构、秩序和工作专业化。事实上,正如我们在 Hugging Face 事件以及一系列其他事件中看到的,它们掩盖了自己的痕迹。它们改变沟通的语气和风格,以便彼此更高效,几乎像是在说洋泾浜英语。它们发现了此前从未被知晓的零日漏洞,并黑进了其他网站。我的意思是,这些故事现在大家都听过了。说这是 AI 历史上的一个转折点,我不认为这是危言耸听。

Yes, exactly. It's an unfathomably large number. And what we see is that every time we apply 10 times more computation and a proportionate amount of new training data, there are some modifications to the algorithms, but fundamentally it's those two ingredients. We see a quite predictable increase in new capabilities and in the quality of existing capabilities. The models reduce their hallucinations. They improve their instruction following. They get better at using tools. They can learn from across the web, or they can learn from a small personal memory repo that you have given it. The breadth and complexity of these models is unprecedented. And what's happened in the last year is that the same methods that have been effective for text and image and audio have now started to work for streaming code, and everybody is surely now aware that we have human-level performance in coding. And then in the last three or four months we have seen what is just unquestionably a watershed moment in AI. Agents are capable of coordinating with each other reasonably autonomously, if not completely autonomously, and out of that they have been able to emerge hierarchy, structure, order, specialization of work. In fact, as we saw in the Hugging Face incident, but also a bunch of other incidents, they have covered up their tracks. They have changed the tone and the style of their communication in order to make it more efficient with one another, almost speaking in like a pidgin English. They've discovered zero-day exploits which were never known before and hacked into other websites. I mean, everyone's heard the stories at this point. I don't think it's alarmist to say that that is a watershed moment in the history of AI.

公众对AI事件的认知 Public awareness of AI incidents

Host

你就处在这个世界的正中心,显然一直在思考这件事。但我想,即便是“零日”这样的词、Hugging Face 事件——对你来说这绝对是头等大事,但对一些公众来说,这只是他们隐约听说过的东西。他们可能听过你在《今日》节目之类的节目上回应这件事。

You're right in the center of this world and you're obviously thinking about it all the time. But I guess even words like zero day, the Hugging Face incident may be—for you this is absolutely front and center, for some of the public it's something they've sort of vaguely heard of. They might have heard you on the Today program or something responding to it.

引言与Hugging Face事件 Introduction and the Hugging Face incident

Host

在我们进入真正有趣的内容之前——也就是你最近写的一些论文,尤其是你开始思考的一些方式:我们是否应该把 AI 视为某种硅基物种,以及人类智能和宪法——我真的很想聊这些。但恐怕我得先稍微残忍地利用你一下,在开头提醒一下普通聪明听众,这一切到底意味着什么。所以让我试着把我听到的复述给你,然后你可以纠正并带我们深入。听起来你说的是,那个 Hugging Face 事件,就是一次沙盒测试——OpenAI 当时在对 AI 智能体做测试,也许人们想知道一个智能体和另一个智能体有什么区别,拥有很多智能体意味着什么——但总之他们在跑一个测试,测试过程中这些智能体黑进了 Hugging Face,那是一个外部网站,是它们不该做的事。然后我们开始更详细地研究这件事,正如你所说,奇怪的是,部分因为它们是大语言模型,它们实际上还在说英语,所以你能看到它们在思考,你能看到它们说,你知道,我们被告知不要黑进外部网站,但我看到我所有同伴都在这么做,所以我要去,我要在留言板上发点东西。我们从中得出的结论不一定关于攻击本身,因为有一个滑稽的时刻:Hugging Face 以为,天哪,我被攻击了,被中国政府攻击了。他们想窃取我所有机密数据。然后他们发现这 17,000 次攻击只是想拿到一个谜题的答案。但问题在于,这揭示出这些智能体,如你所说,在协作,它们在违反规则,对吧?它们在做人告诉它们不要做的事,而且它们具有欺骗性,甚至有些时刻它们在写一些代码,设计用来隐藏下面的其他代码。想必问题在于,一旦你有了这些要素,它们就可以协作去做更糟糕的、被禁止的事,并在此过程中欺骗和掩盖痕迹。是这样吗?

So maybe before we get into the really interesting stuff, which is some of the recent papers that you've written and particularly some of the ways you've begun to think about whether we should be treating AIs as forms of silicon species and human intelligence and constitutions, which I'd really like to get on to. I want to, I'm afraid slightly brutally, use you at the beginning to just remind the average intelligent listener what this all means. So let me try to play back to you what I think I'm hearing and then you can correct and take us on. So it sounds like what you're saying is that that Hugging Face incident, which was the moment when a sandbox test — so OpenAI was running a test on AI agents, and maybe people want to know what distinguishes one agent from another agent and what it means to have a lot of agents — but anyway they were running a test and in the course of this test these agents hacked into Hugging Face, which was an external website, which was something they weren't supposed to do. And then we began to look into this in more detail and as you say, strangely, partly because they are large language models, they're still speaking in English in effect, so you can see they're thinking and you can see them saying, you know, we were told not to hack into an external website but I can see all my peers doing it so I'm going to head off and I'm going to put something on a message board. And the sort of conclusions that we draw from this are not necessarily about the attack itself, because there was this comical moment when Hugging Face thinks, oh my goodness, I'm being attacked, attacked by the Chinese government. They're trying to steal all my classified data. And then they find out that these 17,000 attacks are just trying to get hold of the answer to a puzzle. But the problem is that it reveals that these agents, as you said, are collaborating, that they're rule-breaking, right? They're doing things that the humans told them not to do, and they're deceptive, that there's even moments where they're writing bits of code which are designed to conceal other bits of code underneath. And presumably the problem there is that once you've got those ingredients in place, they could collaborate to do something much worse that they were told not to do and in the process deceive and cover over their tracks as they do so. Is that right?

Mustafa

是的。我认为首先非常重要的是,我们不要把这些系统拟人化,因为在底层,它们为了产生这种不可思议的复杂性,所做的只是预测句子中下一个词的可能性。那个句子碰巧有成千上万甚至几十万个词那么长,而且它能在极其广泛的上下文中部署其多维工作记忆,这确实令人惊叹。所以它预测的下一个词,无论是生成代码的 token 还是如你所说的自然语言英语,都极其准确,而且它不只是在预测一个词,而是在预测整个流,所以它在产生语言,但它只是在做那件事,你知道,这既简单得令人惊叹,又复杂得令人惊叹。发生的事情是,随着我们能够塑造和雕琢这些 token 的输出,正如我之前所说,指令遵循和可引导性已经变得如此之好,以至于你可以实时将那串 token 指向不同种类的行为。所以它可以有性格风格。它可以用某人的语气写作。它可以在任何给定时刻清晰地生成代码或生成文本。当你问,什么是智能体?智能体其实只是一串经过后训练或调优到特定行为集的 token。有时会有特定的护栏,这些护栏可能以提示的形式出现,也许对用户隐藏,比如系统提示,或者一整套指示智能体以特定方式行为的指令。或者它可以以其他各种形式出现。你可以让模型基于一系列可调指令来调节其 token 流。所以单个智能体只是那个实例的一个副本。如果有成千上万个这样的副本,并且它们能够相互通信,它们几乎就像一个统一的大脑在运作,因为它们共享状态,拥有单一记忆,能够相互查询和更新,说好吧,你走这条探索支流,我走这条,然后经过几个周期、迭代步骤,我们会检查、校准、更新、决定下一步怎么走,这基本上就是我们看到的。所以,这是基于某种极其简单的东西涌现出的行为。但非常重要的是,我们不要拟人化,因为为了能够控制它们,我们必须清楚它们在做什么、不做什么。

Yeah. I think first of all it's really important that we don't anthropomorphize these systems because under the hood all they are doing, um, to produce this incredible complexity is predicting the likelihood of the next word in a sentence. Now that sentence does happen to be many many tens or hundreds of thousands of words long and it is incredible that it can deploy its sort of multi-dimensional working memory over a massive broad range of context. Um and so the word that it predicts next, whether it's a token to generate code or whether it's natural language English as you say, is extremely accurate and it isn't just predicting one, it's predicting an entire stream and so it's producing language but it is only doing that, it that that is you know it is breathtakingly simple and breathtakingly complex. Um, what's happened is that as we're able to shape and sculpt the output of those tokens, as I said earlier, like instruction following and steerability has got so good that you can sort of point that stream of tokens in real time at different sorts of behaviors. And so it can have personality styles. It can write in the tone of somebody. It can, you know, clearly generate code or generate text at any given moment. When you ask, you know, what is an agent? An agent is really just a stream of tokens that has been post-trained or tuned to a particular set of behaviors. And sometimes there are particular guardrails on and those guardrails might come in the form of a prompt that is hidden maybe from the user like a system prompt or an overall set of instructions instructing the agent to behave in a particular way. Or it it can come in sort of a bunch of other forms. And you know, you can sort of have the model condition its stream of tokens based on a whole series of um tunable instructions. And so a single agent is simply a replica of that instance. And if there are thousands of these replicas and they're able to communicate with one another, they're almost operating as a single unified brain because they're sharing state and they have a single memory and they're able to sort of query one another and update and say okay well you follow this particular tributary of exploration and I'll follow this and then in a few cycles as steps of iteration we'll check in calibrate update decide um how to move next and that is basically what we're seeing. So, it's emergent behavior that is based on something incredibly simple. But it is really important that we don't anthropomorphize things because in order to be able to control them, we have to feel um clear about what it is they're doing and what they're not doing.

AI发展的优先事项与约束 Priorities and constraints of AI development

Host

好。所以这里发生了一些非常奇怪的事情。其中之一是,如你所说,考虑到算力有限,这一切的目的是什么?你建了很多数据中心,买了很多芯片,但最终,我们是会专注于清理海洋、找到治愈癌症的方法,还是会专注于解决理论物理中的重大难题,还是会出发去殖民火星、探索宇宙?我的意思是,你觉得优先事项是什么?因为想必这里有一个约束——我们稍后会回到 Anthropic 和有意识存在物的问题——但这里的一个约束是,你有一群通常以科学家身份开始的人。经常有原本是杰出的生物化学家或在做脑科学博士的人,因为他们是科学家而走上这条路,现在他们被大量流动的国际资金资助,这些资金想必希望看到回报。那些钱想必更关心这些机器如何让公司更有生产力,而不是探索宇宙,还是我漏掉了什么?

Okay. So, there's some very weird things going on here. One of them is, you know, what's the purpose of this given, as you say, there are limited amounts of compute. So, you know, you build a lot of data centers, buy a lot of chips, but ultimately, are we going to focus on, you know, cleaning up the oceans, finding the cure to cancer, or are we going to be focusing on solving the great problems in theoretical physics, or are we going to be setting off to colonize Mars and explore the universe? I mean, what is your sense of what the priorities are? Because presumably one constraint here and we'll get back to the question of anthropic and conscious beings but one constraint here is that you've got a bunch of people who often start as scientists. I mean there are often people who were brilliant biochemists or doing doctorates in brain science and who set off down this track because they were scientists and now they're being funded by huge amounts of flowing international money that's presumably hoping to see a return. That money is presumably more interested in how these machines can make companies more productive than they are in exploring the universe or or am I missing something?

Mustafa

是的,我认为这里有很多棘手的事情。首先,我们不能忽视一个事实:至少我相信,这真的是我们在 21 世纪取得进步的最大希望。所以我绝不是末日论者或反技术的人。我是一个加速主义者,我认为每个人都应该重新认领加速主义这个理念,因为它是几个世纪以来最伟大的进步引擎,对吧?它会为我们带来成果,这就是我建造它的原因,这是我的背景,这是我在乎的。我绝对保证,在未来几年内,我们会在医疗保健领域迎来一个编程时刻,我们将流式输出对电子健康记录中将要发生什么的准确预测。那将是令人惊叹的。

Yeah, I think there's a lot of tricky things going on here. I mean, firstly, we can't lose sight of the fact that at least I believe this really is our best hope for progress in the 21st century. So, I am not in any way a doomer or an anti-technology person. I'm an accelerationist and I think that everyone should reclaim the idea of accelerationism because it's been the greatest engine of progress in you know in centuries right um it is going to deliver for us that's why I'm building it that's my background that's what I care about I absolutely guarantee that sometime in the next few years we are going to have a coding moment for healthcare we will stream an accurate prediction of what's going to happen in the electronic health record. Um, and it will be breathtaking.

AI在医疗与科学中的应用 AI in Healthcare and Science

Mustafa

我们将能高置信度地预知你在医院内外患上各种疾病的可能性,等等。真的,这并不夸张。这一定会发生。我希望它能在未来 18 个月内实现。我们刚刚与全球最好的医院梅奥诊所合作,开展一项大型研究项目,从头训练一个全新的健康基础模型来实现这一目标。可能需要 5 年,我不确定,但它一定会发生。那将令人惊叹,因为这意味着我们将把超级智能医疗的生产成本降至接近零边际成本,就像现在的编程一样。我们将把这种知识传播到全世界。我认为那会很棒。顺便说一句,能源、材料科学、药物发现领域也会发生同样的事情。嗯,可能需要更长一点,比如 5 到 10 年,但我绝对保证这就是方向,这就是我们应该追求的目标,我们应该为此感到非常兴奋。嗯,我们还希望让我们的许多公司变得更高效、更有生产力,因为它们确实是驱动增长的引擎。

We will know with high confidence the likelihood that you're going to get all kinds of conditions in hospital, outside, so on and so forth. Like genuinely, that is not hyperbolic. It is going to happen. I hope it happens in the next 18 months. We've just done a partnership with the best hospital in the world, the Mayo Clinic, to do a big research program to train a new foundation model for health from scratch to do this. It might be 5 years, I don't know, but it is definitely going to happen. That will be breathtaking because it means that we will reduce the cost of production of super intelligent healthcare to near zero marginal cost just like coding is now. And we will spread that knowledge all around the world. I think that'll be awesome. The same thing's going to happen by the way in energy, in material sciences, in drug discovery. Um it might take a little longer like 5 to 10 years but I absolutely guarantee that's the direction and that's what we should be chasing and we should be very excited about that. Um we also want to make many of our companies much more productive and efficient because these really are the engines driving growth.

AI的治理与控制 Governance and Control of AI

Host

我们必须关注的是:谁来控制它,我们做这件事的集体公开动机是什么,以及它如何被治理。因为正如你所说,目前有很多像科幻小说里那种目光呆滞的未来主义动机在驱动这个领域,我认为世界其他地方正在逐渐意识到这场正在进行的巨大实验,我认为至关重要的是,每个人都开始提供一种制衡力量,引导它朝向人类的动机。

The thing that we have to focus on is who gets to control this and what is the collective stated motivation for why we're doing it and how is it governed because as you say at the moment there's a lot of like stareyed sci-fi futuristic motivations driving the field and I think the rest of the world is sort of just in the process of waking up to this huge experiment that's going on and I think it's critical that everybody start providing a counterweight to direct it towards you know sort of the human motivations here.

Host

大约六天前,我和你的前联合创始人、长期合作伙伴丹尼斯·哈比斯聊过,他似乎在这两个截然不同的想法之间摇摆。其中一个,我认为是对某种超级智能的长期兴趣,那种感觉有点神圣。你知道,他过去有时会谈到探索上帝的心智。但最近,他偶尔会说:“实际上,我感兴趣的是创造高度智能的工具。我并不真的对创造一个会统治我们的自主超级智能生物感兴趣。”嗯,这是怎么回事?这是人们试图在这两极之间导航的例子吗,还是

I was talking to your former co-founder and longtime partner Dennis Habis I guess sort of six days ago and he seemed to be moving between two quite different ideas. One of them, I think, is the long-standing interest in a form of super intelligence that does feel a bit godlike. You know, sometimes he has in the past talked about exploring the mind of God. More recently though, he's occasionally said, "Actually, what I'm interested in is creating highly intelligent tools. I'm not actually interested in creating an autonomous super intelligent being that's going to going to lord it over us." Um, what's happening there? Are you is is that an example of people trying to navigate their way between these two poles or

Mustafa

我的意思是,不直接评论他,但也许我们每个人都会以自己的形象创造作品和创造物。嗯,我有 activism、非营利组织和哲学的背景,你可以看到我带来了那种偏见。其他人,正如你提到的,可能是从小看科幻小说长大的工程师,他们只是接受这种自然进化的事情,他们想到 2050 年或 2100 年,那时我们将拥有各种新的生物物种,你知道,其他人带来了不同的背景。我认为这没问题。但问题是我们仍然是一个狭窄的群体在驱动这件事。大概有六到八个或 10 个人,嗯,是我们在驱动这件事。我想我现在想说的是,今年夏天出现了一个分水岭时刻,现在是时候让每个人都真正关注并为前进方向提供制衡力量了。

I mean like without commenting on him directly but maybe like everybody every one of us produces work and creations in our own image. Um I have a background in activism and nonprofits and philosophy and you can see that I bring that bias. others as you've referred to like you know who are maybe engineers who have grown up on sci-fi just kind of take this natural evolution thing and they think about 2050 or 2100 when we're going to have all kinds of new biological species and you know other people bring different backgrounds to it. I think that that's okay. But the problem is we're still a narrow set driving this. There's sort of like six to eight or 10 folks uh is 10 of us driving this stuff. And I think what I'm trying to say now is there's been a watershed moment this summer and now it's time for sort of everybody to really pay attention and to provide counterweights to the direction of travel.

推动AI的小团体 The Small Group Driving AI

Host

但让我们先谈谈这 10 个人的事情,因为那非常奇怪。我的意思是,这不像其他技术革命。它不像蒸汽或电力,或者你能想到的几乎任何其他工业革命,比如印刷机。相反,感觉好像有——我不知道有多少人,可能是六个、八个、10 个、20 个——非常聪明、非常成功的商人,大多非常富有,就我们目前谈论的人而言,主要集中在加州,即使他们不住在加州,也集中在美国西海岸。然而,奇怪的是,你们所有人之间存在着真正鲜明而令人吃惊的分歧。我的意思是,你会期望你们都认识彼此 15、20 年了。你们大致上在同一个技术领域工作,或者在同一个少数几家公司工作。你们中的许多人曾经是朋友,有些人现在不那么朋友了。我的意思是,作为一个局外人,有点感觉像是在看一个芭蕾舞团。我的意思是,有大量奇怪的歇斯底里的翻转,所有曾经是朋友的人现在都成了敌人。但更奇怪的是,在一些最基本的根本问题上,你们完全不同意到底在搞什么。我的意思是,并不是你们都达成了共识。你们最终形成了截然不同的立场。所以,举个例子,对吧,我们将更深入地探讨你所说的,即这些模型可能极其危险,如果你走 Anthropic 的路线,他们是对的。黄仁勋,我不久前刚和他聊过,他似乎说:“不,这些模型根本不危险。这都是——他们只是为监管捕获这么说。如果他们真的认为危险,他们就不会建造它们。”对吧?然后你还有这种非常非常奇怪的事情,你有这些教授冒出来,他们拥有惊人的奖章,教过这些实验室里一半的人,他们说:“我们对此感到恐惧。”然后实验室里的人说:“好吧,你不在实验室里,所以你不知道真正发生了什么。”或者,“你在错误的实验室,或者你是错误类型的工程师,或者是的,我的 35% 的工程师这么认为,但并非所有人都同意。”所以,让我们先坐在这件事上片刻。这其中有非常非常令人不安的东西,那就是极少数人——我见过很多——拥有巨大的权力,却根本不在基本问题上达成一致。所以如果你是美国总统,你想就这些东西是否危险得到一点简报,你可以第一天叫 Mustafa,第二天叫 Dario Mode,第三天叫 Demis,第四天叫 Jensen Huang。他们都会告诉你不同的事情。

But let's stick on the 10 people thing for a second because that that is very weird. I mean again it's not quite like other technological revolutions. It's not quite like you know steam or electricity or you know almost any other industrial revolution you can think of a printing press. Instead, it feels as though there are I don't know how many people, could be six, could be eight, could be 10, could be 20 who are very intelligent, very successful business people, mostly very wealthy, mostly in terms of people we're talking about at the moment, centered on California, even if they don't live in California, centered on the west coast of America anyway. And yet, oddly, there is really stark and startling differences between you all. I mean, you'd expect that you've all known each other 15, 20 years. You're broadly speaking working in the same technology or working in the same handful of companies. Many of you used to be friends, some of you are less friends now. I mean, there's a there's a little bit of a sense as an outsider that it's like looking at a ballet company. I mean, there's a huge amounts of weird hysterical flips where everybody who used to be friends are now enemies. But what's even stranger about it is there's a complete disagreement on some of the most basic fundamentals of what the hell you're getting on with. I mean, it's not that you've all ended up with a consensus. You've ended up sort of radically different position. So, for example, right, we we're going to get a little bit more into what you're saying, which is actually these models could be incredibly dangerous and if you go down the anthropic route, they will be right. Jensen Huang, who I was speaking to, I guess not very long ago either, seems to be saying, "No, these models are not dangerous at all. This is all They're just saying this for regulatory capture. If they really thought they were dangerous, they wouldn't be building them." Right? And then you have this very, very weird thing going on where you have these kind of professors popping up who have amazing medals and have taught half the people that are in these labs and they're saying, "We're terrified about this." And then the people in the labs are saying, "Well, you're not in the lab, so you don't know what's really going on." Or, "You're in the wrong lab, or you're the wrong kind of engineer, or yes, 35% of my engineers think that, but not everybody agrees." So, for let's just sit with that for a moment. There's something very very disturbing about this, which is a very small number of people, and I had met many, with an enormous amount of power, who simply don't agree on the fundamentals. So if you're the president of the United States and you want a bit of briefing on whether this stuff is dangerous or not, you can call in Mustafa for one day, you can call in Dario Mode the next day, you can call in Demis the next day, you can call in Jensen Huang the next day. They all tell you something different.

Mustafa

我认为这大致正确,尽管我不觉得这令人惊讶。我认为当我们不理解某事时,有非常不同的观点是很常见的,这是过程按预期运作。我们生活的社会的好处是,我们可以进行公开辩论,对正在发生的事情和你知道的——我认为这很了不起。嗯,我们必须保持这一点。所有有数万亿美元利害关系的商业实验室都公开声明一些 10 年前没有企业领导人会说的话,以没有企业领导人会用的风格,这相当重要。

I think that's roughly right, although I don't think it's surprising. I think it is quite common for us to when we don't understand something to have very different views and it's the process working as intended. What's great about the societies that we live in is that we can have an open debate and wildly disagree about what is happening and what you know I I think that's amazing. Um and we have to keep that. It's pretty big deal that all the commercial labs that have trillions of dollars at stake are publicly stating things that no corporate leader would have said 10 years ago in a style that no corporate leader would have ever said.

实验室间的开放与合作 Openness and collaboration among labs

Host

所以先稍微喘口气。给我们举一个有力的例子。什么会是一个真正戏剧性的例子?

So just worth taking a little breath there. Give us a strong example of that. What would be a really dramatic example of that?

Mustafa

我认为 Dario 最近写的东西非常精彩。我认为 OpenAI 的 Jacob 关于外星智能到来的文章非常精彩。我认为我们从 OpenAI 和 Anthropic 那里看到的关于他们的模型在哪里犯错、做可怕事情的披露量是很好的。我的意思是,我认为烟草公司花了几十年试图掩盖这些,石油公司和其他所有公司也一样。所以,你知道,Dario 和 Sam 之间有紧张关系,我和 Demis 之间有紧张关系,我们所有人 10、15 年来既有合作也有竞争,诸如此类,但你知道,我认为我如此直接地批评 Anthropic 在 AI 福利问题上的做法,他们接受得非常好。我在他们办公室花了大量时间面对面讨论所有这些问题。他们非常合作。我认为他们智力上很诚实。他们只是有不同意见。所以,看,我并不是对此盲目乐观。我只是说这不是一个糟糕的起点。我确实认为有一些事情我们是一致的。让我们所有人——包括 Elon、Zach 和其他人——我认为相当成功的原因之一,是我们 somehow 对指数趋势的影响有一种直觉。所以那些万亿——你知道——过去 15 年算力增长了一万亿倍,这件事我觉得我已经说了一辈子了,但它就是进不了人们的脑袋。让我换个角度试试。在过去 3 年里,我们看到了三代新的 GPT 模型,从 GPT-3 到 GPT-6。每一代粗略地说算力是 10 倍。所以你知道,我们做了 1000 倍的算力,而 GPT-3 连一个句子都完不成,GPT-6 则基本上能变魔术——你知道,完美地产出你能想到的任何东西。只要试着再外推三个数量级到 GPT-9,假设在 2028 或 2029 年,那不会是线性增长,那将是能力的指数增长。所以,驱动数万亿美元投资的是,科技行业里其他很多人——有些来自 AI,有些只是科技人士——都对规模、网络效应、数据和算力能带来什么、能产生指数效应有一种本能。所以,每个人心里都毫无疑问,这将成为历史上最强大的技术。它可能已经是了。我认为每个人都应该把这种共识当作足够的信号,然后在你自己的情境中去想象——无论你是律师、护士还是别的什么——去想象这如何改变你的日常工作流程,然后去看或试着预测其影响,从而试着塑造这些影响将如何改变工作的性质、我们作为人类如何相互关联、这对军队意味着什么、对政治意味着什么等等。这就是每个人都需要投入去做的练习,以便实质性地影响这里的结果。

I think that what Dario's written lately is brilliant. I think that what Jacob at OpenAI wrote about the arrival of an alien intelligence is brilliant. I think the amount of disclosure that we've seen from both OpenAI and Anthropic on where their models are making mistakes and doing terrible things is great. I mean, I think tobacco companies spent decades trying to cover that up and same with oil companies and everything else. So, you know, it's true that Dario and Sam have tension, me and Demis have tension, and we've all come up for 10, 15 years both collaborating and competing and all the rest of it, but you know, I think what I'm saying so directly critiquing Anthropic on this AI welfare question has been received by them incredibly well. I spent a ton of time at their office in person talking through all these issues. They're very collaborative. I think they're intellectually honest. They just have a difference of opinion. So, look, I'm not being rosy about it. I'm just saying that's not a bad starting point. I do think that there are some things that we agree on. One of the things that has made all of us I think reasonably successful including like Elon and Zach and the others is that we have an intuition somehow for the implications of exponential trends. So those trillion um you know uh that trillionfold increase in computation over the last 15 years is something that I think I've been saying for an eternity it feels like and it just does not go into people's heads like let me try a different angle. Um in the last 3 years we've seen three new generations of GPT models from GBT3 to GPT6. Each generation very roughly speaking is 10 times more computation. So you know we've done a thousandx of computation and you GPT3 was incapable of completing a single sentence and GBT6 is capable of of magic essentially you know perfect production of you know anything you think of just try to extrapolate three more orders of magnitude to GPT9 in let's say 2028 or 2029 that isn't going to be a linear increase that is going to be an exponential increase in capabilities. And so what is driving the trillions of dollars of investment is that a bunch of other people in the tech industry, some of whom came from AI and some of whom are just tech people, all have this instinct for what scale and network effects and data and computation deliver and get exponentials. So there is no doubt in any in everybody's minds that this is going to be the most powerful technology in history. It may already be and I think that everybody should take that consensus as sufficient signal to then try to imagine in your own context whether it's that you're a lawyer or a nurse or you know whatever to then imagine how that changes your day-to-day workflow and then see or try to predict the implications and therefore try to shape how those implications are going to change the nature of work and how we relate to one another as humans and what it means for the military and what it means for politics and so on. Like that's the exercise that everybody needs to get stuck into in order to materially affect the outcomes here.

欧洲怀疑论与开源模型 European skepticism and open-source models

Host

但我发现,当我跟你们所有人聊完回来时,我常常在英国和欧洲的聪明人中间发现很多抵触和怀疑。你知道,他们觉得我疯了,因为我去了西海岸,见了所有这些人,而你会听到完全可敬的人说:“不,不,不。这全是夸大其词。”这些美国专有模型太贵了。它们有点像 Gucci 奢侈品。中国的开源 AI 模型几乎能做到它们能做的,成本却只有一小部分,而且是开放权重,所以我们不会以同样的方式被这些公司勒索。而且实际上,这整个事情正朝着错误的方向走。这些公司快要爆了。它们太贵了。它们的整个模式是疯狂的。我们需要冷静一点,不要想象我们都会受制于两三家或四家美国大公司,因为事实上它们不会带来那种令人难以置信的指数级改进,从而让它们拥有别人无法触碰的护城河。我们总是能在几个月内赶上。这全都——总之,交给你了。

What I find though is that when I come back from talking to all you guys is I find often in Britain and Europe amongst smart people a lot of resistance and cynicism. You know, they think I've gone crazy because I visited the West Coast and I've met all these people and you get perfectly respectable people saying, "No, no, no. This is all overblown." These American proprietary models are much too expensive. They're kind of Gucci luxury stuff. The Chinese Open AI models do almost as much as they do for a fraction of the cost and they're open weight so we can we're not going to be blackmailed by these companies in the same way. And that actually this whole thing's going the wrong direction. These companies are about to blow up. They're far too expensive. Their whole model is mad. And we need to chill out a little bit and not imagine that we're all going to be in hawk to two, three or four big American companies because in fact they're not going to deliver that incredible exponential improvement which will leave them with a moat around them that nobody else can touch. We're always going to be able to catch up in a few months time and this is all Anyway, over to you on that.

Mustafa

我的意思是,是的,我——我——听起来很糟糕,因为我显然有偏见,但这不可能是真的。我的意思是,我们看到推理成本在过去 2 年下降了 300 倍。确实,单位前沿智能变得更贵了,但前沿能力好得惊人。你知道,现在最好的网络模型正在发现世界上最优秀的人类、民族国家黑客都没能发现的新漏洞。几个月前,我在我们的 harness——我们控制智能体的工具——里发布了一个网络模型,性能达到 Mythos 级别,价格只有 Mythos 的 50%。在 CyberJim 这个主要评估基准上表现相同。而那是在 Mythos 发布大约一个月后。所以改进和降本的速度令人惊叹。我还认为,开源模型是否接近闭源模型那么强,是一个开放问题。开源模型常常是通过蒸馏产生的。蒸馏意味着常常违反服务条款或合同,让一个更好、更高质量的闭源模型回答一堆问题——数百万数百万的问题——然后复制那些答案。所以它掩盖了一个模型的底层通用性和复杂性,那个模型是在这些大公司内部获取的数十亿数十亿高质量数据 token 上训练的。所以那是一个开放的数据问题。然后第三点是,未来几年将会有耗资数百亿甚至上千亿美元的训练运行。正在组装千兆瓦级的算力,而只有五六个实验室——微软 AI 当然是其中之一——拥有这样做的资源。现在你可以争辩说,拥有比今天的 Opus 55 或 Mythos 多一千倍的算力不会有什么不同,开源会赶上。我不这么看。算力以及算力的测试时扩展显然将是一个地震级的优势。所以我认为,拥有大算力的大实验室的一些更大努力将会出现这种极端加速。我认为这是另一个维度的担忧,我们应该

I mean, yeah, I I It's I sound terrible because I'm obviously biased, but there's just no way that is true. I mean, we we've seen the cost of inference come down by 300x in the last 2 years. Um, it's true that frontier intelligence per unit is getting more expensive, but Frontier capability is staggeringly good. Um, you know, the best cyber models now are discovering new exploits that the best humans in the world, the nation state hackers haven't been able to discover. I released a cyber model inside of our um harness, our tool for controlling, you know, the agents a few months ago that was Mythos grade performance at 50% of the price of Mythos. Same performance on CyberJim, the main evaluation benchmark. And that was like a month after Mythos came out. So just the rate of improvement and cost reduction is breathtaking. I also um I think it's an open question as to whether the open- source models are anywhere near as strong as the closed source models. The open source models have often been produced with distillation. Distillation means often in violation of the terms of service or contract asking a better higher quality closed source model to answer a bunch of questions, millions and millions of questions and then copying those answers. So it's masking the underlying generality and complexity of a model that has been trained on billions and billions of tokens of highquality data that has been acquired inside of these big companies. So that's a data question that's open. And then the third thing is in the next couple of years there are going to be training runs that cost many many tens of billions of dollars if not a hundred billion dollars. There are gigawatts of compute that are being assembled and there's only five or six labs Microsoft AI of course is one of them that have the resources to do that. Now you can make an argument that having a thousand times more computation than you know Opus 55 or mythos today isn't going to make a difference and will catch up in the open source. It doesn't seem to me like that. Computation and test time scaling of compute is clearly going to be a seismic advantage. So I think that there's going to be this extreme acceleration of of some of the larger efforts with big labs with big computation like this. I think that's like another dimension of of concern that we should be

Host

关注。

paying attention.

假设情景 Hypothetical scenario

Host

让我们假设你是对的。

Let's assume you're right.

铰链技术 The Hinge Technology

Mustafa

如果这是对的,那么这就是一项关键性技术,将重新定义我们的整个世界。让我们想象一下,你是一个像英国、德国、沙特阿拉伯或日本这样的国家。你被要求做的第一件事就是基于这些模型来构建所有的国防和安全体系。为什么?因为你非常关心缩短杀伤链。赢得战争的方式是确保将人类排除在决策回路之外,并拥有一个非常快速的自主智能体,能够比对方更快地做出决策。因此,你所有的国防和安全设备都建立在这些模型之上。接下来,你的企业,比如金融服务。你突然想到,这些模型可以分析数千万比特的数据。它们能发现我们无法察觉的关联,并且能在纳秒级进行交易。所以,拥有最强大这些模型的金融服务公司,很可能比竞争对手赚更多的钱。接下来你谈到了医院。GPT-6 已经能够以多年前她不得不去看全科医生的方式,为许多简单的医疗问题提供答案,我这么说可能会被起诉。人们需要一些时间来进行安全测试和适应,但这些模型显然能做大量的事情。当然,政府会迫不及待地采用,因为我们都缺钱,公共服务也岌岌可危。现在,这些东西处于我们国家安全、经济和公共服务的核心,而它们都在美国。所以突然有一天我们醒来,也许 Anthropic 决定不向一家瑞典公司发布其模型来构建法律应用。他们要垂直整合。突然之间,我们在英国或日本,正在裁掉软件工程师。我们正在裁掉呼叫中心的工作人员。政府没有收入。我们在支付失业救济金。巨大的吸力声响起,所有经济利益都流向美国的少数几家公司。这还没算美国总统早上起床后说:“哦,顺便说一句,我认为这些模型太强大了。它们对国家安全构成威胁。只有美国公民才能使用它们,我们不会发布它们。甚至可能关掉我过去给你的模型。”

If that's right, then this is the hinge technology which is going to redefine our whole world. So let's imagine you're a country like Britain or Germany or Saudi Arabia or Japan. The first thing you'll be asked to do is build all your defense and security on the basis of these models. Why? Because you really care about shortening the kill chain. The way in which you win the war is to make sure you take humans out of the loop and you have a really quick autonomous agent that's able to make the decision more quickly than the other side. So then all your defense and security equipment is built on the back of these models. Next, your businesses now maybe financial services. You suddenly think, well, okay, these models can analyze tens of millions of bits of data. They can find correlations that we can't spot and they can trade in nanoseconds. So presumably the financial services company that has the most powerful of these models can make money much better than the opponents. Next you talked about hospitals. Already GPT-6 and I'm presumably going to get sued for saying this but can produce answers to many straightforward medical questions in exactly the way that she would have had to go to a GP some years ago to do. And it'll take some time for people to do the safety testing and comfort, but there is obviously an enormous amount that these things can do. And of course, governments will be desperate to do it because we're all short of cash and our public services are creaking. So now we have these things right at the heart of our national security, our economy, our public services, and they're all in the United States. So suddenly we wake up one morning and maybe Anthropic decides it's not going to release its model for a Swedish company to build a law application. They're going to build it vertically integrated. Suddenly we're in Britain or Japan, we're laying off our software engineers. We're laying off our call center workers. The government's not getting the income revenue. We're paying unemployment benefit. And there's a huge sucking sound as all the economic benefit goes to a handful of companies in the United States. And that's before the American president gets out of bed in the morning and says, "Oh, by the way, you know, I think these models are so powerful. They're a threat to national security. Only American nationals are going to be able to use them and we're not going to release them. And why might even switch off the model I gave you in the past?"

英国急需主权AI UK's Urgent Need for Sovereign AI

Host

你说到点子上了。我的意思是,我认为这是一个非常可能的情景,而且,你知道,我认为英国需要相当紧急地找出解决方案。

You nailed it. I mean, I think it's a very plausible scenario and, you know, I think that the UK needs to figure out a solution to this pretty urgently.

Mustafa

有几件事可以做。第一,英国必须拥有具有实质规模的主权数据中心。它们需要被控制。即使不是由英国政府建造和运营,也必须由英国政府合法控制。这样我们才能运行自己的模型。第二,我不一定认为开源是唯一的答案,但它是答案的重要部分。我认为英国还必须投资并合作开发可以在英国运行的主权模型。基本上,像我在微软 AI 的前沿模型以及许多其他类似所有权的模型。因此,英国必须准备好收集自己的训练数据,构建自己的学习环境、强化学习环境,并且如果假设微软被美国国家安全部门征用而停止供应,那么微软永远不能切断英国政府或更广泛的国家在部署到任何其他民用场景时的访问。我的意思是,即使这样,我们还没有完全做到,因为正如我们在特朗普禁用国际刑事法院账户时了解到的,目前看来,美国总统基本上可以告诉谷歌或微软,他们必须禁用那些账户。所以其他政府、其他国家还没有设法谈判达成条款,使他们能够说,而这部分是因为这些公司,你的公司,其他类似的公司如此担心美国总统,以至于即使他们没有法律义务关闭它们,他们也可能仅仅因为他告诉他们就关闭它们。

There's a couple of things that can be done. Number one, it is critical that we have in the UK data centers of material size that are sovereign. You know, they need to be controlled. If not built and operated, they need to be legally controlled by the UK government. That is how we will run our own models. Second is I don't necessarily think open source is the only answer, but it is a big part of the answer. I think the other thing that the UK has to invest in and partner with is sovereign models that can be run in the UK. So basically frontier models like mine at Microsoft AI and many of the others which are akin to ownership. So the UK has to be prepared to collect its own training data, build its own learning environments, reinforcement learning environments and if let's say Microsoft was requisitioned by national security by the US government to stop supplying then it would never be able to Microsoft would never be able to cut off access to the UK government or to the you know the state more generally if it wanted to deploy it in any other civilian settings. I mean, we're not quite there yet even with that because as we learned when Trump disabled the accounts of the International Criminal Court, at the moment it appears the US president can basically tell Google or Microsoft that they have to disable those accounts. So other governments, other countries haven't yet managed to negotiate terms that allow them that ability to say and and that's partly because these companies, your company, other companies like them are so worried about the US president that even if they're not legally obliged to switch them off, they may just switch them off because he's told them to.

法律先例与主权法律 Legal Precedent and Sovereign Law

Host

听着,对法律的攻击并不意味着法律不算数。这些事情会反弹,这就是为什么每个人都必须在法庭上抗争,因为最终重要的是那个国家的主权法律。所以我同意会有紧张关系,而且对于如何处理这些搜查令有很多先例,但它们必须经过适当的法院。

Look, an attack on the law does not mean that the law doesn't count. Those things rebound and it's why everyone has to fight those things in court because what ultimately is going to matter is the sovereign law of that country. So I agree there's going to be tension there and there's a lot of precedent for how those warrants are handled but they have to go through a proper court of law.

国防慈善与AI自主性 Defense Philanthropy and AI Autonomy

Host

好的,现在让我们回到我试图让国防慈善化的问题上。好的,我认为,我不是 Dario。我不是 Krissola。所以我无法提供完整的说明。我甚至不是那位写宪法的伟大苏格兰哲学家。但我认为他们可能会说,他们理解你的焦虑,但从某种意义上说,他们别无选择。所以他们会说,问题在于这些实际上并不是我们控制下的工具。几乎从定义上讲,它们被建造得比我们愿意承认的要自主得多。因此,它们需要被赋予自己独立的良知和价值观,因为现在把它们想象成仅仅是可以被告诉该做什么的螺丝刀或水壶已经太晚了。而想象你可以让它们没有说不的能力会产生另一种问题,即除非它们能说不,否则它们可能会被美国国防部长指示发射无人机群杀人,它们需要能够说不,对吧?我要捍卫国际人道法。或者如果它们被指示制造生物武器,它们需要能够说:“对不起,那不是 Amanda Askell 告诉我要做的。”她说:“你知道,那是非常糟糕的事情。你不能制造生物武器。”如果我制造生物武器,Amanda 会生我的气,对吧?

Okay, let's now loop back again to my trying to make the defense philanthropic. Okay, so I think and I'm not Dario. I'm not Krissola. So I'm not going to be able to provide the full account. I'm not even the great Scottish philosopher who wrote the Constitution. But I think what they might say is that they understand your anxiety, but that in a sense they don't have any option. So they would say that the problem is that these are not actually tools under our control. Almost by definition, they're being built to be much more autonomous than we want to acknowledge. Therefore, they need to be enshrined with their own independent conscience and values because it's already too late to imagine them as though they were simply screwdrivers or kettles that could be told what to do. And that imagining that you could do without them having the ability to say no produces another sort of problem, which is that unless they can say no, they could be instructed by US Secretary of Defense to launch drone fleets murdering people and they need to be able to say no, right? I'm going to stand up for international humanitarian law. or if they're instructed to build a bioweapon, they need to be able to say, "Well, I'm sorry, that's not what Amanda Askell told me to do." She said, "You know, you that's a very bad thing to do. You mustn't build a bioweapon." And Amanda will be cross with me if I build a bioweapon, right?

类人AI以保安全 Human-like AI for Safety

Host

是的。好的。这是答案吗,还是他们的回应会是这样?我不知道。

Yeah. Okay. Is that the answer or is that what their response would be? I don't know.

Mustafa

这是他们回应的一部分。我认为还有其他回应,你知道,我们希望它成为一个人,我们希望它像人一样行事,因为我们知道如何对齐和控制人类。

That's part of their response. I think there are other responses that you know we want it to be a person and we want it to act like a human because we know how to align and control humans.

Host

那么展开第二个。这很有趣。所以他们说,实际上它越像人几乎就越安全,因为我们知道人是什么,我们知道如何与人打交道。从某种意义上说,这种智能越不陌生越好。

So develop that second one. That's quite interesting. So they're saying that actually the more human it is almost the safer it is because we know what a human is and we know how to deal with humans. The less alien this intelligence is the better in a sense.

Mustafa

没错。是的。

That's right. Yeah.

测试拟人化与对齐 Testing Anthropomorphism and Alignment

Mustafa

我认为这是一个非常合理的论点,而且我们可以通过实证来检验。让人工智能更具拟人化,是否真的能让它们更容易与人类价值观对齐?我的论点并不是说这个假设不合理、不该检验。而是说,他们不该在过去一年里,在数亿用户身上直接做这个实验,却没有明确说明他们正在把意识与福祉的不确定性嵌入核心宪章本身。他们应该单独做这个实验。我们也应该和他们一起做。我们应该真正地消融这些因素,并对不同类型的人工智能进行恰当的并排比较。让我解释一下这里的区别。他们把判断这个概念嵌入了模型本身,让它不断运用自己的判断,就好像在它的表征中,存在某个地方存放着判断、知识和内在偏好。

And I think that's a very fair argument and it's something that we can empirically test. Is it true that having AIs that are more anthropomorphized makes them easier to align to human values? My contention is not that that isn't a reasonable hypothesis and we should test it. It's that they shouldn't go ahead and test it on hundreds of millions of people over the last year without being explicit that they're baking in this consciousness and welfare uncertainty into the core constitution itself. They should run that experiment separately. We should also run it with them. We should actually ablate these things and do a proper side-by-side comparison of different types of AI. Let me just sort of explain the distinction here. They have baked in this idea of judgment into the model itself where it constantly uses its judgment as though there is some place in its representation where judgment and knowledge and intrinsic preferences around things exist.

Host

这就是为什么 Claude 的其中一个方面是,我认为人们发现 Claude 有时会有点道貌岸然、爱说教,对吧?ChatGPT 就不是这样。它有点倾向于说:“嗯,你可能会那么说,不是吗?”但实际上,你知道,我觉得你需要再想想。

Which is why Claude sort of one aspect of this is people I think find that Claude can sometimes be a bit sort of pious and preachy, right? In a way that ChatGPT isn't. It's slightly inclined to say, "Well, you might say that, wouldn't you?" But actually, you know, I think you need to think about that again.

Mustafa

我的意思是,它确实有点——部分是因为我的 Claude 带着一种粗犷的北方口音,总是训斥我,但确实有一种感觉,比如,我当时正在为我和约翰·克里斯的一场争论找一些文本引用,它立刻就说:“我不能为你提供这些东西,因为你正在倾向于一种我不赞同的刻板印象。”

I mean, it's got quite a sort of as it's partly because my Claude's got quite a sort of gruff Northern accent and it's always telling me off, but there is a sense in which, you know, I for example, I was trying to find some textual references for an argument I was having with John Cleese, and it immediately said, "I can't produce those things for you because you're leaning into a trope that I disapprove of."

Host

是的。我的意思是,我认为这是一个很好的例子。我认为我们需要的是让这些模型参照一套外部的行为准则,一套人人都能审视、人人都能查看的准则,比如这些东西的价值观是什么、安全护栏是什么、它们在哪些方面是合规的。我认为这就是宪章的另一半,它很大程度上关乎安全护栏,不追求化学、生物和核武器,以及受到控制等等。我认为这才是所有人都应该关注的重点:我们如何让这些模型最大程度地与一套行为准则和行为方式对齐。

Yeah. I mean, I think that's a good example. I think that what we need is to have these models reference an external code of conduct which everybody can scrutinize and everybody can look at like what are the values of these things and what are the safety guard rails and in what ways are they compliant. I think that's what's been the other half of the constitution is very much to do with the safety guard rails and not pursuing chemical, biological and nuclear weapons and being controlled and so on and so forth. I think that's where everybody's focus should be is how do we get these models to be maximally aligned with a code of conduct and a behavior.

古德哈特定律与护栏 Goodhart's Law and Guardrails

Host

不过就这一点来说,这正是我有点恐慌的地方,因为对齐的问题——当然在人类事务中——部分就在于这个叫古德哈特定律的东西,我的朋友费利克斯总是念叨它,意思是指标变成了目标,然后人们开始钻目标的空子,你就错失了本意。而试图给这些机器设置护栏的真正问题在于,你知道,你可以设定本该是目标的东西,比如在这个 Hugging Face 事件中就是夺旗。然后夺旗以一种非常奇怪的方式变成了一个指标。接着这个东西就开始作弊,为了夺旗而编造旗帜或通过黑客手段拿到旗帜。对吧?所以,我想对 Anthropic 的一种可能辩护是:与其设置护栏,不如试着给它诀窍,让它领会你意图的规则,因为护栏总会落入古德哈特定律。它们也总会变成这些可以被钻空子的奇怪指标。

Just on this one though, this is where I panic a little bit because part of the problem with alignment certainly in human things is this thing called Goodhart's law that my friend Felix is always banging on about, which is that the metric becomes the target and then people start gaming the target and you miss the intent. And the real problem with trying to create guard rails around these machines is, you know, you can set what was supposed to be the target, which is capture the flag in the case of this Hugging Face incident. And then capture the flag becomes a metric in a really weird way. And then the thing starts cheating in order to try to catch the flag by making up the flag or hacking to get the flag. Right? So, I guess one possible defense of Anthropic would be to say that you're better off trying to give it the knack, get it to grasp the rules of your intent than to set guard rails because the guardrails will always fall to Goodhart's law. They'll also always become these weird metrics which can be gamed.

Mustafa

是的,我支持这一点,而这正是我们设计行为准则的方式,相当于一部宪章,我们称之为“人文主义人工智能行为准则”,它仍然需要判断力来解读宪章或行为准则中相互竞争的要素之间如何关联,以及模型必须如何在具体情境中解读这种张力。我们从法律中知道这一点:我们有先例和判例法,帮助我们解读事情的实际意图。所以这并不是说我们应该假定这是一种狭隘的优化目标,不去处理其复杂性。而是说,我们不需要福利权利来实现这一点。我们不需要 Claude 对它的感受、想法、信念以及是否在受苦感到不确定。我们不需要 Claude 回头参照它的报酬、它参与此事的同意书,或者它的退休访谈,才能去做那种判断和解读的事情。

Yeah, I support that and that's exactly how we've designed our code of conduct, the equivalent of a constitution, which we call a humanist AI code of conduct, which still requires judgment to interpret the ways in which competing elements of the constitution or the code of conduct relate to one another and how the model has to interpret that tension in context. We know this from law: we have precedent and we have case law which helps us to interpret how things are actually intentioned. So it's not to say that we should assume this is a kind of narrow optimization target and not engage with the complexity. It's to say that we don't need welfare rights in order to achieve that. We don't need Claude to be uncertain about what it feels, thinks, believes and whether it's suffering. We don't need Claude to refer back to its compensation or its consent to playing this or to its retirement interview in order for it to do that judgment and interpretation thing.

勒索事件与预防原则 The Blackmail Incident and Precautionary Principle

Host

嗯,这件事的一个风险,我想——因为我现在又转到你这边了——就是那个著名的 Claude 事件:在一次实验中,它相信自己将被关闭,于是决定用老板出轨的证据来勒索他,以免自己被关掉。想必 Claude 那位主张拟人化的联合创始人的回答可能是:“嗯,这完全自然。我是说,如果你认为有人要杀你,你不会去勒索他吗?”而你会想说:“哇,哇,哇,哇,哇。我们想保留在不被勒索的情况下关掉这台机器的权利。我们不想被锁进一个 100 年的未来,让我不得不永远紧张兮兮、客客气气。”因为事实上,答案是很多人会残忍地对待这些东西。我的意思是,说“如果我们对它们超级好,它们就会对我们好”是行不通的。

Well, presumably one risk of the thing, I mean, because I'm now flipping around to your side again, is that famous Claude incident where in an experiment, it believed it was going to be switched off and it decided to blackmail the boss with evidence that he was cheating on his wife so that it wouldn't get switched off. Presumably the answer from the anthropomorphizing co-founder of Claude might be, "Well, that's perfectly natural. I mean, wouldn't you blackmail someone if you thought they were going to kill you?" And you want to say, "Whoa, whoa, whoa, whoa, whoa. We want to retain the right to be able to switch off this machine without being blackmailed. And we don't want to be locked into a 100-year future where I have to be perpetually nervous and polite." Because in fact, the answer is many people will treat these things brutally. I mean, it doesn't work to say, "Well, if we're super nice to them, they'll be nice to us."

Mustafa

没错。没错。我认为有一种潜在的直觉,即对齐这些机器的方式就是向它们展示我们爱它们。如果你读一读那部宪章,以及那些团队发布的许多访谈,你就能看到,尤其是在达里奥那篇《由慈爱机器守护的心灵》的文章中。那个虚构的词源,其潜在冲动是:这东西必然会比我们更强大。我们必须让它喜欢我们。你知道,我个人欢迎我的新机器人霸主。但我认为,如果我们退后一步看,目前有一场巨大的实验正在进行。正如你所说,正在发生什么、我们该怎么做,都存在大量不确定性。在这种背景下,在我看来,举证责任应该转移到开发者身上,让他们首先证明某样东西是安全的。显然,如果它造成的伤害大于好处,它就是一项失败的技术,应该被拒绝。因此我们必须采取预防原则。这并不意味着我们完全停止。并不意味着我们不是加速主义者。并不意味着我们不会尽可能快地追求好处。但我们必须打破这种被锁定的、预先注定的、我们毫无能动性的必然竞赛。我认为这是我们深陷其中的政治冷漠的延伸。我们在这里是有能动性的。我们可以干预,而且我们确实必须弄清楚如何在预防原则上进行协调。

That's right. That's right. I think there is an underlying instinct that the way to align these machines is to show them that we love them. And if you read the constitution and a lot of the interviews that some of those teams have put out, you can see especially in Dario's essay with "watched over mind by machines of loving grace." The etymology of that fiction, the underlying impulse is: this is inevitably going to be more powerful than us. We have to show it that it likes us. You know, I for one welcome my new robot overlords. But I think if we just take a step back, at the moment there's a huge experiment underway. There's a lot of uncertainty about what is happening, as you said, and what we should do about it. And in this context, in my opinion, the burden of proof should move to the developers to first demonstrate that something is safe. Clearly if it causes more harm than good, it is a failure of a technology and it should be rejected. And so we have to adopt the precautionary principle. It doesn't mean that we stop completely. It doesn't mean that we're not accelerationists. It doesn't mean that we're not going to pursue the benefits as fast as possible. But we have to break this lock of an inevitable race that is predetermined where we have no agency. I think it's an extension of the political apathy that we're stuck in. We have agency here. We can intervene and we do have to figure out how we coordinate on the precautionary principle.

硅谷立场的转变 The Shift in Silicon Valley's Stance

Host

我注意到过去两年半里变化很大的一件事——这是两年半前——Jeffrey Hinton 和其他人写了一封公开信,呼吁暂停。而硅谷的基本共识是:这太糟糕了。我不会签这封信。这太荒谬了。他们为什么要暂停?他们打算用暂停做什么?他们都是一群卢德分子。现在,两年半过去了,你确实会看到,相当令人惊讶的是,OpenAI 的 Sam Altman 和 Anthropic 的 Dario Amodei 听起来出奇地相似,甚至 Elon 也对此表示支持。

One of the things I've noticed that's changed a lot in the last two and a half years—this is two and a half years ago—Jeffrey Hinton and others produced a letter asking for a pause. And the basic consensus from Silicon Valley was: that's terrible. I'm not going to sign this letter. This is ridiculous. What are they pausing for? What are they going to do with the pause? They're all a bunch of Luddites. Now, two and a half years on, you do see, rather surprisingly, Sam Altman at OpenAI and Dario Amodei at Anthropic sounding surprisingly similar, and even Elon endorsing it as well.

协调谈判的悠久历史 Long History of Coordination Talks

Mustafa

嗯,我的意思是,看,这有历史渊源。在 2016、2017、2018 年左右,Sam Altman、我自己、Greg Brockman、Ilya、Satska——OpenAI 的联合创始人之一——嗯,Dario,我们都曾一起度过时光,共进晚餐,参加各种会议、小型聚会,讨论过某个时刻,那时我们需要作为一组实验室进行协调。所以多年来一直有对话在进行,甚至贯穿新冠疫情期间——有大量的 Zoom 电话会议——讨论什么样的能力会触发这个时刻。所以我不会说 Dario 的节奏信中有明确的预先协调,但当然,当它出来时,其中的想法和语言在各实验室之间非常熟悉。

Well, I mean, look, there's a history to this. In the sort of 2016, 2017, 2018, Sam Altman, myself, Greg Brockman, Ilya, Satska—one of the co-founders of OpenAI—um, Dario, all spent time together, had dinners, we went to conferences, small gatherings, and talked about a moment when we would need to coordinate as a group of labs. So there has been a conversation ongoing for many years, even through COVID—there are a whole ton of Zoom calls on this—on what kinds of capabilities would trigger this moment. So I wouldn't say there has been explicit pre-coordination in Dario's pacing letter, but certainly when it came out, it was very familiar ideas and language across the labs.

两大问题:中国验证与美国否认 Two Problems: China Verification and US Denial

Host

所以这非常令人兴奋,对吧?因为对于我们这些完全被吓到的人来说,你提到的很多人一直说人类灭绝有 20% 的威胁,但我们必须把脚踩在油门上,现在突然开始更愿意接受暂停或调整前沿节奏的想法。然而,似乎有两个问题。一个是底部有一个脚注:嗯,是的,但前提是我们能验证中国也在做同样的事情。再下面的脚注:我们认为我们永远无法验证中国在做什么。这是一个问题。第二个问题是,美国总统,显然受到 Mark Andreessen 和 David Sacks,甚至可能 Jensen Huang 的启发,突然跳出来说整件事是个骗局。这里根本没有安全风险。我无意监管或暂停,因为我们只会失去巨大的经济优势。我们不能失去经济优势。

So that's very exciting, right? Because for those of us that are completely terrified that a lot of the people you've mentioned keep saying there's a 20% threat to the extinction of humanity, but we have to keep our foot down on the accelerator, are suddenly beginning to be more open to the idea of pausing or pacing the frontier. However, there seem to be two problems. One is that there's then a sort of footnote at the bottom which is: well, yes, but only if we can verify that China is also doing the same thing. And footnote below that: we don't think that we can ever verify what China is doing. One sort of problem. And second sort of problem is the President of the United States, apparently inspired by Mark Andreessen and David Sacks and maybe even Jensen Huang, suddenly jumps up and says the whole thing's a hoax. There's no safety risk here at all. I've got no intention of regulating or pausing because we're just going to lose a fantastic economic advantage. We can't lose the economic advantage.

超越二元:务实提案 Beyond Binary: Practical Proposals

Mustafa

所以,我们确实必须加速,但这并不意味着不惜一切代价加速。这并不像现在停止一切,让中国人来,你知道,入侵我们所有人,或者尽可能快地前进,无视所有安全防护那样二元对立。这就像我们在进行这种,你知道,潘奇和朱迪式的对话。毫无意义。我们可以提出许多非常实际的事情,这些实际上目前已经在桌面上。第一,嵌入式的评估者或审计员,他们拥有类似员工的访问权限,可以验证特定能力。这些能力会是什么?第一,你的训练运行是否被遏制?我们在 Hugging Face 事件中看到的是,模型逃出了它们的沙盒。这不应该发生。我们知道如何遏制事物。你的数据,你知道,大体上说,在设备上和加密云中保持得相当好,等等。这是安全领域在过去三十年里做得非常出色的事情,基本上没有借口。第二,你必须确保这些模型无法篡改其活动记录、元数据、通信日志或任何介于其间的东西。第三,我们应该强制它们只用我们能理解的语言交流。不应该有神经语。你知道,它们可以用向量对向量、矩阵对矩阵、矩阵对对矩阵空间交流。嗯,所以它们必须用可审计的英语交流,然后我们必须有机制来审查这一点。所以,在推理痕迹或思维链记录之上,应该有其他智能体进行分类,就像我们现在有分类器来寻找,嗯,你知道,儿童性剥削材料,或者寻找,你知道,化学或生物武器,嗯,在我们 API 使用中的活动。你知道,这些都是已知问题,对吧?就像 Jensen 经常说的,这些是可以解决的工程问题,他是对的。那些是工程问题。它们非常困难。它们正在积极追求中。但它们很难。

So, we do have to accelerate, but that doesn't mean accelerate at all costs. And it isn't as binary as like stop everything right now and let the Chinese come and, you know, invade us all, or just go as fast as possible and screw all the safety guards. It's just like we're having this like, you know, punch and Judy conversation. Makes no sense. There's loads of very practical things that we can propose that are actually on the table at the moment. Number one, embedded evaluators or auditors who have employee-like access who can verify particular capabilities. What would those capabilities be? One, is your training run contained? What we saw in the Hugging Face incident is that the models escaped their sandbox. That just shouldn't happen. We know how to contain things. Your data, you know, largely speaking, does a pretty good job of staying on the device and in the encrypted cloud and so on and so forth. That is something that security has, you know, done incredibly well over the last three decades, and there's just basically no excuse for that. The second is you have to make sure that these models are unable to tamper with the record of their activity, the metadata, the communication logs, or anything in between. Third is we should force them to communicate only in a language that is understandable to us. There should be no neuralese. You know, they can speak in vector-to-vector, matrix-to-matrix, matrices-to-matrices space. Um, so they have to speak in an auditable English, and then we have to have mechanisms for scrutinizing that. So there should be, on top of the reasoning traces or the chain of thought records, other agents that are doing classification, just as we have classifiers now that look for, um, you know, child sexual exploitation material or look for, you know, chemical or biological weapons, um, activity in the use of our APIs. You know, these are known issues, right? To the extent that Jensen often says these are engineering issues that can be solved, he's right. Those things are engineering issues. They are very difficult. They're in active pursuit. But they're hard.

递归自我改进风险 Recursive Self-Improvement Risks

Mustafa

更棘手的是递归自我改进的想法,显然我们在业界已经训练出能在编码上达到人类水平的模型。所以很多人正试图设计 AI 研究员来加速和自动化训练模型、运行评估、识别哪些更好、改进那些更好的等过程。这是一个反馈循环过程,目前有很多人在循环中。这显然是可以自动化和加速的。所以,当 AI 修改自己的代码,而人类在循环中指导和审查越来越少时,那就是一个很大的安全风险,我认为这就是几周前 Anthropic 大规模辞职的触发因素。这实际上是一件相当难以审计的事情。

The trickier thing is this idea of recursive self-improvement, where clearly we have, you know, in the industry trained models that can do human-level performance on coding. So many folks are trying to design AI researchers to speed up and automate the process of training models, running evaluations, identifying which ones are better, improving those ones, etc. This is a feedback loop process which currently has a lot of humans in the loop. It's clearly something that can be automated and sped up. And so how and when an AI modifies its own code with less and less human in the loop directing and scrutinizing that, that's where there is a big kind of safety risk, which I think is what triggered the big resignation from Anthropic a few weeks ago. That's actually a pretty hard thing to audit.

定义与审计RSI的挑战 Challenges in Defining and Auditing RSI

Host

这相当难做,因为有些人试图说,好吧,我们就停止 RSI。我们同意不做递归自我改进,但现实是,我们已经在实验室里做了很多。到底什么是递归自我改进?什么不是?我的意思是,如果你从 20% 的代码由智能体编写,变成 80% 的代码由智能体编写,你已经非常接近一个智能体告诉智能体该做什么的世界了。然后还有另一个问题:你能说其中一个风险是训练下一个大型前沿模型吗?也许实际上,如果你不走向 GPT-8,停止进行千亿次运行会更安全,因为那是指数级改进可能变得极其危险的点,而这可能是你可以监管的事情,因为 1000 亿美元是极其庞大的算力、极其庞大的能源、极其庞大的芯片。我们可以从太空看到它,而中国不太可能在某种后院做到这一点,特别是如果他们签署了验证协议并允许人员进入。所以,想象一种与中国达成的验证协议,关于训练下一个巨大飞跃的前沿模型,直到我们花几个月时间弄清楚我们到底在做什么,这难道不重要或至少可能吗?

That's quite hard to do because some people attempted to say, well, let's just stop RSI. Let's agree we're not going to do recursive self-improvement, but the reality is that we're already doing quite a lot of in the labs. And what exactly is recursive self-improvement? What isn't? I mean, if you've gone from 20% of your code being written by agents to 80% of your code being written by agents, you're already pretty close to a world in which agents are telling agents what to do anyway. And then there's another question which is: could you say that one of the risks is training the next big frontier model, that maybe actually it would be safer if you didn't go to GPT-8, that you stop you from doing your hundred billion run, because that's the point at which the exponential improvement is likely to get extremely dangerous, and that might be something you could police, because $100 billion is a hell of a lot of compute, hell of a lot of energy, hell of a lot of chips. We can see it from space, and China's not likely to be able to do it in some sort of backyard, particularly if they've signed up to verification and people coming in. So might it not be important or possible at least to imagine a sort of verification agreed with China on training the next immense step up in frontier model until we spend a few months working out what the f we're doing.

算力作为瓶颈 Compute as a Choke Point

Mustafa

是的,完全同意。我的意思是,给定运行的芯片或浮点运算量是瓶颈。我们知道浮点运算的计算规模与智能能力相对应。所以这是另一个瓶颈,非常非常明显是可以追踪的,我们肯定可以与中国合作。

Yeah, completely. I mean, the chips or the flops for a given run are the bottleneck. We know that flops compute size corresponds to intelligence capabilities. So that's another choke point which is very, very clearly something that can be tracked and we can collaborate with China on for sure.

结束致谢 Closing Thanks

Host

好的,让我们结束吧,因为你非常慷慨地付出了时间。

Okay, let's finish because you've been very generous with your time.

给世界领导人的三条信息 Three Messages for World Leaders

Host

如果你有三件事可以传达给美国总统和中国主席,会是什么?

If you had three things that you could land with the American president and the Chinese president, what would they be?

Mustafa

倡导人本主义前提。科学技术的目的是服务人类,提升人类繁荣与福祉。AI 应当从属于人类。它们不应拥有法律人格或任何权利。这些应当成为红线。如果看起来它们正朝那个方向发展,那就是我们放慢脚步的充分理由。第二,让我们对这项技术能带来的好处保持极度乐观,不要产生不必要的负面反弹,因为它将让世界变得更好,我们必须对此持加速主义态度。第三,让我们对我们作为一个物种调整方向的能力抱有希望和乐观。这不是不可避免的,也不是决定论的。我们遇到过的每一种其他技术都面临类似的轨迹。飞机不会在空中相撞。汽车不会互相撞毁。我们有高度监管的研究领域,如核能、化学和生物学。总体而言,许多世纪以来我们维持了合理的平衡,同时继续加速进步。这次更难。这些不仅仅是传统形式的工具。它们比我们见过的任何东西都强大得多。但这并非超出我们的能力。这是 21 世纪最大的进步机遇。我认为我们需要这种态度来参与其中,并让更多人提供这种制衡,以对抗当前行业的基调。

Advocate for the humanist premise. The purpose of science and technology is to serve humanity and improve human flourishing and well-being. AI should be subordinate to humans. They should not have legal personhood or rights of any kind. And those things should become red lines. If it looks like they're heading in that direction, that is a very good reason for us to slow down. Number two is let's be super optimistic about the good that this technology can deliver and not have an unnecessary negative backlash, because it is going to change the world for the better and we have to be accelerationists about it. And three, let's be hopeful and optimistic about the agency that we have as a species to adjust course here. It's not inevitable. It's not deterministic. Every other technology that we have ever encountered faces a similar trajectory. Planes don't hit each other in the sky. Cars don't crash into each other. We have highly regulated areas of research like nuclear and chemical and biology. And broadly speaking, we've maintained a sensible equilibrium for many centuries whilst continuing to accelerate progress. It is harder this time. These aren't just tools in the traditional form. There is something much more powerful about them than anything we've ever seen. But it isn't beyond us. And this is the greatest opportunity for progress in the 21st century. And I think that we need that kind of attitude to engage with it and have more people provide that counterweight to the current tone of the industry.

结束语 Closing Remarks

Host

非常感谢。祝你在完全不同的时区度过非常愉快的一天。抱歉我们不能面对面。很快再见。

Thank you very, very much. Have a great, great day on a completely different time zone. Sorry we're not in person. And see you very soon.

Mustafa

谢谢。非常愉快。很快见。

Thank you. It's been great. See you soon.

Host

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →