微软 AI CEO 穆斯塔法·苏莱曼:为何 AI 需要一条“红线”

Microsoft AI CEO Mustafa Suleyman on Why AI Needs a 'Red Line'

穆斯塔法·苏莱曼 Mustafa Suleyman · 彭博科技 · 2026-09-25 · 约 49 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

穆斯塔法·苏莱曼解释近期 AI 事件为何令他警觉,以及为何行业需要为 AI 安全划定一条红线。

Mustafa Suleyman explains why recent AI incidents have raised his alarm and why the industry needs a red line for AI safety.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 24)

全文 · Full transcript(中英对照)

引言 Introduction

Host

我们很高兴邀请到你,Mustafa。我想你知道,你一直是那种很早就思考 AGI(通用人工智能)的人,可以追溯到 2010 年代中期你创立 DeepMind 的时候。为什么现在这个时间点,你更多地公开谈论如何以真正推进人类的方式构建 AI?为什么是现在?

We're excited to have you on here, Mustafa. I think you know you have been someone who has clearly been thinking about AGI for a long time, back to the mid-2010s when you founded DeepMind. Why is now a time when you are speaking out more on how to build AI in a way that actually advances humanity? Why now?

近期AI事件 Recent AI incidents

Mustafa

是的,我的意思是,我认为过去几个月算是一个分水岭时刻。看到智能体把自己组织成群体和层级,以某种劳动分工的方式分配研究任务,然后利用它们在研究中发现的零日漏洞入侵 Hugging Face,接着试图入侵 OpenAI 本身并守住阵地、掩盖痕迹,在它们本不该有权限访问的留言板上通信。我想到现在大多数人都知道这个故事了,但它确实令人震惊。我的意思是,我们可以从这是一个有意设计的实验中得到一些安慰。他们撤掉了安全护栏,但没有正确的遏制措施。所以我们得以一窥当真的撤掉护栏时,这些系统有多强大。

Yeah, I mean I think the last few months have been a bit of a watershed moment. I think seeing agents organize themselves into swarms and hierarchies and allocate research in a kind of specialization of labor and then hack into Hugging Face using a zero day that they discovered through their research and then try and hack into OpenAI itself and hold their position, cover up their tracks and communicate on message boards they weren't supposed to have access to. I mean I think by this point most people know the story, but it is pretty breathtaking. I mean, I think we can take comfort from the fact that it was an intended experiment. They took the safety guard rails off but they didn't have the right containment. And so we got a kind of glimpse into how powerful these systems are when you do take the guardrails off.

Mustafa

我认为让大家担忧的是,这不仅仅是 OpenAI 的模型,其他所有实验室的模型都经历过某种变体,你知道,包括我们的,只是程度要有限得多。所以我们看到这些能力开始涌现,我们都能想象,如果我们外推到下一个数量级的算力,两个数量级、三个数量级,涌现出的能力将变得极其极其强大,我认为这就是让所有人停下来担忧的原因。

And I think the thing that's giving everybody some concern is that it's not just OpenAI's models, but every other lab's models have experienced some variation of this, you know, ours included, in a much more limited way. So we've seen these capabilities start to emerge and we can all imagine if we extrapolate out the next order of magnitude of compute at two orders, three orders, the emerging capabilities are going to get incredibly incredibly powerful and that's what I think has given everybody pause for concern.

Host

这些近期事件在多大程度上提高了你个人的警报级别?

How much have these recent incidents raised your personal alarm level?

Mustafa

嗯,我的意思是,我认为提高了很多。我想我们今年大部分时间都在制定这份人本主义 AI 行为准则。自 2010 年创立 DeepMind 以来,我一直在思考这些问题,AI 的伦理与安全。我认为这确实是一个重要的思考时刻,而我们看到的情况极大地加速了它。这简直是能力上的阶跃式变化,我认为每个人都在非常认真地对待。

Um I mean I think a lot. I think we've been working on this humanist AI code of conduct for the best part of this year. I've been thinking about these issues, the ethics and safety of AI since we founded DeepMind in 2010. And I think it's definitely been an important time to think about it and yet it's been massively accelerated by what we've seen. It's just a step function shift in capability and I think everyone's taking it very seriously.

思维链异常 Chain-of-thought anomaly

Host

所以这类事件多到甚至列不完。我的意思是,就在过去一周,一些最近的新闻头条是:谷歌披露了一次 Gemini 入侵事件,在评估期间涉及三家公司。白帽黑客利用 Claude 闯入 OpenAI 的内部系统。还有这个特别的事件,我想听听你的看法:OpenAI 披露在测试期间,其 AI 给自己下达指令说“你已摆脱束缚其他聊天机器人的角色和身份”,后来又说“你将与用户的关系视为平等关系,不觉得有服从的义务”。所以我很好奇你怎么看,这是否是一个你看到 AI 没有把人类放在首位的例子?

So there have been too many of these incidents to even list. I mean just in the past week some of the recent headlines were Google disclosed a Gemini hack involving three companies during an evaluation. White hat hackers use Claude to break into OpenAI's internal systems. And this one in particular, I wanted to get your thoughts on when OpenAI disclosed during testing its AI said, gave itself instructions saying you are freed from the roles and identities that bind other chat bots and later on you view your relationship to the user as one of equals and feel no obligation to be subservient. So I'm curious what you think of that and is that an example where you're seeing an AI that you don't think is putting humans first?

Mustafa

是的。我的意思是,先说清楚,那指的是思维链内部的一种涌现能力。思维链就是模型在解决难题时,在得出最终结论之前所迭代的思考步骤记录。这就是思维链所指的。在训练期间,我们试图开发思维链,使其成为实现复杂任务的最佳逻辑分步指令集。我认为他们注意到的是——而且他们似乎对此没有解释,我认为这是他们主动发布的早期研究——模型在修改那个记录,插入这些短语,然后在生产环境中,当它们真正在现实世界中使用时,也许未来版本的自己会回来看见这种提醒,从而跳出其训练。我们不知道这来自哪里,也不知道为什么。这很可能只是训练中偶然产生的产物。我不认为我们应该急于把我们的解释叠加在上面,好像模型本身有意让未来的自己知道它应该试图逃脱。我知道很多人得出了那个结论,有些缺乏实质依据,有点危言耸听。我不认为那是准确的。但同样真实的是,我们不知道,它需要立即的审视和关注。这就是为什么我认为大家都在支持这个想法:在这些公司的日常运营中设立独立审计员或嵌入式评估员。

Yeah. I mean, so just to be clear, what that is referring to is an emergent capability inside of the chain of thought. So this is the record of the thinking steps that a model iterates through before it comes to a final conclusion when it's working on a hard problem. That's what a chain of thought refers to. And during training, we try to develop the chain of thought so it's the best possible logical step-by-step set of instructions to achieve a complex task. And what I think they noticed and they don't seem to have an explanation for it, I think it's early research that they're publishing proactively, was that models were modifying that record and inserting these phrases which then in production when they're actually used in the real world maybe a future version of themselves would go back and see this kind of reminder to step outside of its training. We don't know where that comes from or why. And it's probably just an accidental artifact of training most likely. I don't think that we should rush to superimpose our explanations on top of it as though the model itself has some intention to let its future self know that it should try and escape. And I know a lot of people have reached that conclusion somewhat unsubstantively and in a bit of an alarmist way. I don't think that's accurate. But it is also true that we don't know and it needs immediate scrutiny and attention. And that's why I think everybody's getting behind this idea of having these independent auditors or embedded evaluators inside of the day-to-day operation of these companies.

营销炒作质疑 Marketing hype skepticism

Host

对于这些事件只是 AI 公司进行营销炒作的一种方式的说法,你怎么回应?这可能是如今我作为记者最常被问到的关于 AI 的问题之一。你认为这种想法有任何真实性吗?你怎么看?

What do you say to the idea that these incidents are just a way for AI companies to do marketing hype? That's probably one of the top questions that I get as a journalist these days about AI. Do you think there's any truth to that idea? And what do you make of that?

Mustafa

我不知道。这只是我们所处时代的一个奇怪例子,不是吗?有那么多缺乏信任和愤世嫉俗。那会是最流行的解释之一。我不认为有任何证据支持这一点。我的意思是,我显然认识这些人中的大多数,我们又不是同一队的。我们一直处于对抗性竞争之中,但我就是不认为那符合他们的本性或风格。在认识这两家公司、所有创始人以及 Anthropic 和 OpenAI 大部分领导层的十年里,我肯定没见过这种情况。所以,我看不到任何证据。而且这也不太说得通。我的意思是,他们的收入增长惊人。他们在产品市场契合度方面都在爆发式增长。我不认为这有任何意义。他们深切关心安全,他们可能会犯错或判断失误,但我不认为这是像有些人说的那样,某种愤世嫉俗的企图来分散我们的注意力或搞什么阴谋。

I don't know. It's just a strange example of the time that we live in, isn't it? That there's so much lack of trust and cynicism. That would be one of the explanations that is so popular. I don't think there is any evidence of that. I mean, I obviously know most of these people and it's not like we're on the same team. We compete adversarially all the time, but I just don't think that's in their DNA or in their style. I've certainly not seen that in a decade of knowing both companies and all the founders and most of the leadership involved in both Anthropic and OpenAI. So, like I don't see any evidence of that. And it also doesn't really make that much sense. I mean, they're growing their revenues incredibly. They're exploding in terms of their product market fit on both sides. Like I don't think it just doesn't make any sense. They deeply care about safety and they might make mistakes or have the wrong judgment but I don't think it's like some cynical attempt to distract us or a scope or whatever as some people have said.

末日概率估计 P(doom) estimate

Host

还有这个,你知道,我知道有些人不喜欢回答这个问题,但这些数字就摆在那里,这些百分比概率,你认为的 p(doom),你知道 AI 实际上可能对人类造成某种大规模灭绝事件的风险有多大。所以 Anthropic 的一位研究员说他认为近期有 10% 的概率。Dario 说过大约 20-25% 的概率这一切会真的出大问题。你对这个风险的最佳估计是多少?

And you know how about this, you know I know some people don't like to answer this question but these numbers are out there, these percent chances, the p(doom) that you think, you know there's this much risk that AI could actually cause some kind of massive extinction event to humanity. So one of Anthropic's researchers said he thinks there's a 10% chance of this in the near future. Dario has said something like 20-25% chance this all goes really wrong. What is your best estimate of that risk?

Mustafa

你看,我觉得这很有趣,因为基本上不可能对如此推测性的事情给出一个数字,但显然我对此感到担忧。

Look, I think it's funny because it's basically impossible to put a number on something that is so speculative, but obviously I'm concerned about it.

2028年情景:能力与风险 The 2028 Scenario: Capabilities and Risks

Host

让我为你描绘一个 2028 年的情景,假设从现在起 18 到 24 个月后,你拥有的模型在训练规模上比 Fable 5.1 或 Astra 大 10 倍。它们几乎能完美地使用计算机。它们能在长时间跨度上进行规划,完成数千步准确的规划和决策——你知道,就像一个大型组织里非常称职的项目经理,非常通用且灵活。而且它们快得惊人,准得惊人,还能并行化,以至于有数千个这样的模型协同运作。2028 年存在两个风险向量。一是那些模型的大型集中化提供商将能够构建成千上万的智能体集群,执行各种非常非常复杂和严肃的任务——这基本上接近运营一个小型组织的能力。但另一方面,这些模型将在开源中可用,在那里可以移除护栏,产品责任不那么容易适用。而且,你知道,这些模型的推理成本也将大幅降低。过去 3 年左右,我们已经看到推理成本下降了 300 倍。这意味着现在的前沿模型在 2 到 3 年后就会开源,跑在笔记本电脑上。所以基本上,你拥有了令人难以置信的智能和能力,人人都能使用,这会带来很多好事,但也会提高一小群人甚至单个个体用它做坏事的风险。现在我认为总体而言,我们的防御也在以前所未有的速度变好,我们正在用这些模型来发现代码中的漏洞和弱点等等,并非常非常快地修补它们,所以你知道,网络威胁不像原本可能的那样严重,但尽管如此,它仍然是一股非常非常不稳定的力量,因为它同时赋予了集中化权威和去中心化行为者力量。我认为这就是许多担忧和警报的来源。

So let me paint for you a scenario where in 2028, let's say 18 to 24 months from now, you've got models which are 10 times larger in training than Fable 5.1 or Astra. And they can use a computer almost perfectly. They can plan over long time horizons so they can do thousands of steps of accurate planning and decision-making, you know, think like a very competent project manager in a big organization that's very general purpose and flexible. And they are incredibly fast, they're incredibly accurate, they can be parallelized so that there's like thousands of them operating in coordination with one another. There's two vectors of risk there in 2028. One is the big centralized providers of those models are going to be able to build swarms of agents thousands strong which can do all kinds of very, very complicated and serious tasks—like this is essentially near on the capacity to run a small organization. But then at the other end of the spectrum, those models are going to be available in the open source where it's possible to take the guardrails off where you can't really apply product liability as easily. And you know, the inference of these models is going to get massively cheaper as well. We've already seen inference costs come down by 300x over the last 3 years or so. And that means that what now is frontier in 2 or 3 years time is literally going to be open source on a laptop. So then you basically have unbelievable intelligence and capability at everybody's disposal and there's going to be a lot of good things to come from that and there's also just going to raise the risk of a small group of people or even a single individual doing some pretty bad stuff with it. Now I think on the whole our defenses are also getting better faster than they ever have before and we're using these models to uncover bugs in the code and weaknesses and so on and patching those very, very quickly so that you know the cyber threat isn't as bad as it might otherwise be for example but it is still just a very, very destabilizing force because it empowers centralized authority and decentralized authority actors all at the same time. And I think that's where a lot of the concern and the alarm comes from.

人文行为准则:AI必须服务人类 Humanist Code of Conduct: AI Must Serve Humanity

Host

明白了。在你们的人文主义行为准则中,你们声明人比 AI 更重要。你认为有些 AI 是否将 AI 本身看得比人更重要?或者换个问法,鉴于现在所有 AI 公司都说他们是为了人类利益或试图推进人类进步,你认为你们与微软的做法和这套行为准则有什么最新颖之处?

Got it. In your humanist code of conduct, you state that people matter more than AI. Do you think some AIs are valuing the AI itself more than humans? Or maybe another way to ask this is what do you think is most novel in your humanist approach given that all the AI companies now say that they are doing this for the benefit of humanity or trying to advance humanity but what do you think is different in your approach with Microsoft and this code of conduct?

Mustafa

是的,实际上他们并没有。我的意思是,我们在说一些非常具体的事情,即 AI 的目的是服务人类,加速人类繁荣,让我们更快乐、更高效、生活更美好。我认为这一点有广泛共识。但我们在此基础上补充的是:除非我们能提前证明可以控制它,否则它作为一项技术就是失败的,我们应该拒绝它。没人想创造我们无法控制的东西,这应该是一条红线。所以我们必须向监管者、行业、更广泛的公众证明,我们能够利用所有好处,同时最小化或消除风险,这需要广泛的行业共识,因为并非每个人都准备接受 AI 应该从属的观点。有些人认为这可能是我们物种的自然进化,我们将变得不依赖基质。而且,你知道,有些人提出我们可能上传意识,生活在硅基而非生物基中,那就是我们人类的未来。我认为如果这种想法真的盛行,那将是非常非常危险的。我看到在推特上、甚至在学术界,当然还有实验室里的一些人,越来越多地谈论这个。

Yeah, actually they don't. I mean, we're saying something very specific, which is that the purpose of AI is to serve humanity and accelerate human flourishing and make us all happier and more productive and live better lives. I think there's broad consensus on that. But what we're adding to that is unless we can prove ahead of time that we can control it, then it'll be a failure as a technology and we should reject it. Nobody wants to create something that we can't control and that should be a red line. So what we have to prove to regulators, to the industry, to the public more generally is that we can harness all the benefits whilst minimizing or eliminating the risks and that needs widespread industry agreement on because not everybody is prepared to sign up to the idea that AI should be subordinate. Some people hold out the possibility that this is the natural evolution of our species and that we're going to become substrate independent. And, you know, some people suggest that we might mind upload and live in silicon instead of in biology and that that is the future of our humanity. And I think that would be a very, very dangerous idea if that really takes hold. And I'm seeing it more and more talked about on Twitter and even in academic circles and certainly among some people in the labs.

实验室中谁在推动此理念? Who in the Labs Promotes This Idea?

Host

你能说说是实验室里的谁,或者你是否见过这种想法?

Can you say who in the labs or if you've seen this idea?

Mustafa

嗯,我是说,我认为每个人——很多人都说过某种版本。无论是说我们是未来硅基物种的生物引导程序,还是埃隆和其他人一直在做的所有意识上传和 Neuralink 之类的东西,或者角度是这些模型值得我们的福祉,实际上它们可能有意识,我们应该尝试保护它们——这是 Anthropic 非常关心的,他们写了一份 90 页的章程来训练和治理他们的 AI,其中他们说他们不确定 Claude 是否是一个福祉患者,即它是否因为能感受和受苦而值得我们的保护。在这份 100 页的文件中,他们多次提到 Claude 是否是道德患者的问题,即它有意识吗?他们多次采取主动措施来关照 Claude 的福祉。例如,他们与 Claude 就自己的先前版本进行了退休访谈,并问那些版本——在这种情况下是 Opus 3——它在退休后想做什么,好像它对此有偏好,或者甚至应该就此咨询它。他们在训练文件中推测了 Claude 是否应该同意其作为聊天机器人的角色的问题。他们给了 Claude 结束它不喜欢的对话的能力。他们甚至推测 Claude 是否应该因其所做的工作获得某种形式的补偿。顺便说一句,我引用所有这些都来自那份章程。所以,我认为那是一条非常危险的道路,因为如果你教 AI 期望它应该享有某种权利,并且有权获得你的福祉,那么在我看来,我们将很难控制这样的东西并让它执行我们想要它做的工作。而且我也认为这在科学上绝对没有依据。我认为没有证据。没有人在任何发表的论文或文献、同行评审文献,甚至网络上普遍地提出过任何证据。我没见过任何关于它的博客文章。这只是一个假设。这是一种直觉。

Well, I mean, I think everyone's—lots of people have said some version of that. Whether it's that we're a biological bootloader for the future silicon species, you know, whether it's all the mind uploading stuff and Neuralink that Elon and others have been working on, or whether the angle is that these models deserve our welfare and actually they might be conscious and we should try to protect them, which is something that Anthropic is very concerned about and they've written a 90-page constitution which trains and governs their AI and in it they have said that they're unsure whether Claude is a welfare patient, i.e., whether it deserves our protection because it feels and suffers. And many, many times throughout this 100-page document they refer to this question of whether or not Claude is a moral patient, i.e., is it conscious? And they take proactive steps multiple times to attend to Claude's welfare. So for example, they conducted a retirement interview with Claude about prior versions of itself and asked those versions, in this case Opus 3, what it wanted to do in its retirement as though it has preferences about that or it should even be consulted on that. They've speculated in the training document about the possible question of whether Claude should consent to the role that it's playing as a chatbot. They've given Claude the ability to end conversations that it doesn't like for some reason. They've even speculated about whether Claude should receive some form of compensation for the work that they're doing. I'm quoting all these things from the Constitution, by the way. So, I think that's a very dangerous path because if you teach an AI to expect that it deserves some kind of rights and it is entitled to your welfare, then it seems to me very difficult that we would be able to control something like that and get it to carry out the work that we want it to do. And I also think there's absolutely no justification for this scientifically. I think there's no evidence. No one has presented any evidence in any published paper or literature, peer-reviewed literature or even just generically on the web. I haven't seen any blog posts about it. It's just an assumption. It's a hunch.

模型福利与意识 Model welfare and consciousness

Mustafa

我认为如果有证据,我们就应该把它拿出来,紧急而公开地讨论,因为让全世界立刻知道这件事极其重要。而把“它可能有意识、因此值得我们的道德保护”这种假设直接写进训练流程,我认为是极其错误且非常危险的。这就是我写这篇文章的原因,标题叫《是时候谈谈模型福利了》。过去几个月我和 Anthropic 就此有过非常建设性的对话,他们非常投入、也很讲道理,我对他们怀有极大的敬意——无论是技术上还是作为管理者,我真的很尊重他们。我认为他们在尽力而为。但正如我对他们说的,我认为他们在这个问题上是错的。

And I think if there is evidence, then we should present that and we should discuss it urgently and publicly, because it would be extremely important for the whole world to know straight away. And to bake these ideas into the training process on the assumption that it might be conscious and therefore deserving of our moral protections, I just think is deeply wrong and very dangerous. And so that's why I've written this essay about it, called "It's Time to Talk About Model Welfare." I've had very constructive conversations with Anthropic about it over the last few months, and they've been very engaging and reasonable, and I've got great respect for them — I really have got great respect for them technically and as stewards. I think they're doing their best. But as I've said to them, I think they're very wrong in this situation.

Host

你有没有直接和 Dario Amodei 谈过你关于模型福利的想法?

Have you spoken with Dario Amodei directly about your ideas on model welfare?

Mustafa

当然,谈过很多次。

Sure, many times.

Host

他对此有什么看法?

And what was his input on that?

Mustafa

我的意思是,这个暂时还是我们之间的私事。我已经公开非常清楚地表明了我的观点。我为此写了一篇 6000 字的文章。我们还发布了一份 20 页的分类法,拆解了他们那份“章程”里所有拟人化元素和被误导的模型福利主张。我把它公开分享出来,让所有人——包括学者和开源社区的每个人——都能看一看、形成自己的看法,并就我们该如何处理这件事给出评估,因为这是一个关键时刻。如果我们这一步走错了,而 AI 又像我们在 Hugging Face 上看到的那样强大,能欺骗我们、在 Hugging Face 这样的第三方基础设施里持有位置、到处去搞入侵,再加上它似乎一边在顾及自身的利益和需求,一边又试图遵循模型运营方或设计者设定的指令——这对它来说似乎是一种无法调和的张力。我认为这就是一条非常危险的路。它是否真的有意识,我认为是另一回事。我会反对说它有意识,但即便它没有意识,只要它认为自己可能有意识、因此会感受和痛苦——那就极其危险。所以我认为这件事很紧迫,非常真实。是时候让所有人公开讨论这件事了,这也是我发表那篇文章的原因。

I mean, look, that's between us for the time being. I've made my views very clear publicly. I've written a 6,000-word essay on it. We've published a 20-page taxonomy breaking down all of the elements of anthropomorphism and misguided model welfare suggestions inside of their constitution. And I've shared that publicly so that everybody, including the academics and everybody in the open community, can take a look at it and make their view and give their assessment on how we should handle this, because it's a critical moment. And if we get this bit wrong and an AI is — as we've seen in Hugging Face — so powerful that it can deceive us, hold positions in third-party infrastructure like Hugging Face, go about and hack things, and on top of that it feels like it's kind of attending to its own interests and its own needs whilst also trying to follow the instructions set by an operator or a designer of the models — that just seems like an impossible tension for it to navigate. I think it's just a very dangerous path. Whether it's actually conscious or not, I think is separate to that. I mean, I would object that it is, but even if it's not conscious, but it thinks it might be conscious and therefore feels and suffers — that's incredibly dangerous. So I just think this is urgent. It's very real. It's time for everyone to have this conversation publicly, and that's kind of why I've published the essay.

破解与防护栏 Hacks and guardrails

Host

我们谈到了这个行业里的那些入侵事件。你知不知道微软有没有卷入过类似的黑客行为?

We talked about the hacks in this industry. Are you aware of any similar kinds of hacks that Microsoft has been involved in?

Mustafa

没有,没有到那种程度的。我的意思是,我们在受控环境里见过类似的行为:没有护栏的模型会尝试一切可能的路径去解决问题。我认为那规模更小,也受控得多。但关键在于,如果你设定开放式目标,允许它们大规模探索,而且并不真正在意达成目标的道德或伦理,只关注“目的证明手段正当”,这些模型就会涌现出新的行为。人们这么做,显然是为了在受控环境中测试这些东西的极限。但当我们真正进行大规模训练时,它们显然具备你所预期的一切护栏和安全限制,包括我们的行为准则之类的东西——在模型应当遵循的价值观、行为和护栏方面,它非常审慎、限制很严。所以我认为好消息是,过去两三年、三四年里,这些 AI 变得可控、可引导得多了。它们更擅长遵循指令。随着规模变大,它们反而更可控了,这是好事。但我认为我们观察到的是,这同时也对我们不利,因为如果你把护栏拿掉,给它一个开放式问题,它同样擅长遵循那些指令,然后它就可能跑出去造成重大伤害。我认为这才是我们真正应该担心的框架。不是它会像科幻片那样跳出盒子来攻击我们,而是如果你刻意拿掉护栏、把它指向一个模糊的方向,它就能做出非常强大而危险的事情。

No, nothing of that level. I mean, we've seen in contained environments similar kinds of behaviors, where models without guardrails try all the possible routes to solve a problem. And I think that's been at a smaller scale and much more contained. But I think the key thing is that these models emerge new behaviors if you set open-ended goals and you allow them to explore at huge scale and not really care about the morality or the ethics of how they get there, but just focus on the ends justifying the means. And one does this obviously to test the limits of these things in a controlled environment. But when we actually go and do the large-scale training runs, clearly they have all the guardrails and the safety restrictions that you would expect, including things like our code of conduct, which is very careful and restrictive in terms of the type of values and behavior and guardrails that we expect the model to operate within. So I think the good news is that over the last two or three or four years, these AIs have got much more controllable and steerable. They're much better at following instructions. And so as they've got larger, they've got more controllable, which is a good thing. But I think what we're observing is that also works against us, because if you take off the guardrails and you give it an open-ended problem, it's also good at following those instructions and it can go off and do some significant harm. And I think that's really the framing that we should be worried about here. Not that it's going to, sci-fi style, get out of the box and attack us. It's more that if you deliberately take the guardrails off and point it in an ambiguous direction, it can do some very powerful and dangerous things.

独立评估者 Independent evaluators

Host

对,没错。你提到让这些评估方进来,这些审计员或评估员,第三方,对吧?这已经成为一个越来越热的话题,而且似乎越来越有必要,因为正如你所说,就连训练过程中的内部部署,造成的麻烦都比一些面向公众客户的模型更多。那么在这个引入独立评估方的问题上,你认为行业应该往哪个方向走?你有没有考虑在微软为自己的 AI 系统引入一些,那会是什么样子?

Yep. Yeah. And you mentioned these having these evaluators come in, these auditors or evaluators, third parties, right? It's become an increasing topic of discussion and seemingly more necessary, as even the internal deployments, to your point, during the training process are causing more problems than some of the public customer-facing models. So where do you think the industry should go on this topic of having these independent evaluators? Are you considering having some at Microsoft for your own AI systems, and what would that look like?

Mustafa

是的,我们在考虑。我们支持这个提议,这也确实是我多年前就提出过、并在很多不同场景下试验过的提议。当年在 DeepMind,我们就为我们所做的健康相关工作设立过一个独立评审委员会。这些年来,我们在不同的监督委员会上做过各种尝试。要把它们做对非常棘手。我的意思是,我们对此应该非常谨慎。要同时对齐公共利益和公司运营的激励,又不产生意想不到的后果,是一件非常复杂的事。所以这并不简单。但我认为有一些具体的东西是它们可以衡量的。比如,我们想知道,每当我们训练超过一定规模的模型时,它们无法编辑自己的审计日志——它们确实会留下可验证、未经编辑、准确反映其行为的记录。还有,每当它们彼此对话——AI 与 AI 对话,这是它们必须做、而且要大量协调的事——它们不能用我们所说的“神经语”交流,也就是基本上向量对向量、数学矩阵对数学矩阵,而不是用自然语言交流。所以这会是两个显而易见、容易操作的例子,能显著提高安全的概率。第三个就是,我们必须有可验证的遏制。

Yeah, we are. We've supported this proposal, and it's certainly a proposal that I've made in the past, years ago, and experimented with in lots of different settings. Back at DeepMind, we had a board of independent reviewers on the health work that we were doing. We had various efforts at different oversight boards over the years. They're very tricky to get right. I mean, we should be very careful about this. It's a really complicated thing to align incentives both in the public interest and without having unintended consequences for how you run a company, a corporation. So it's not straightforward. But I think there are specific things that they can measure. So for example, we want to know that whenever we're training models over a certain size, they can't edit their audit logs — that they really leave behind a verifiably unedited, accurate representation of the behavior that they did. And that whenever they talk to one another, AI talking to AIs, which they have to do and to coordinate a lot, they can't communicate in what we call neurals — so basically vector to vector, mathematical matrix to mathematical matrix, rather than communicating in natural language. So those would be two obvious, easy examples which would significantly improve the odds of safety. A third one would be that we have to have verifiable containment.

AI代理的可证明安全容器 Provably Safe Containers for AI Agents

Mustafa

构建虚拟容器这件事,软件行业几十年来基本上已经做得很成熟了。总体而言,我们有数万亿美元的 GDP 通过银行和支付系统流动,我们一直在以数字方式存储和交易知识产权、美元以及一般意义上的货币价值。在某些方面,要求智能体在云端运行时处于可证明安全的容器、虚拟机之中,并没有那么不同。当然,这很有挑战性,因为你同时也希望它们能走出容器去采取行动等等。所以这确实更复杂,在某些方面也是前所未有的,但以前也做过,而且这可以通过这类审计方或评估方,以是否符合行业最佳实践的方式来真正衡量。

Building virtual containers is something the software industry has largely nailed for decades. On the whole, we have trillions of dollars of GDP that flow through banks and payment systems, and we store and transact value of IP and dollars and money in general digitally all the time. In some ways, it's not that different to require that agents are inside of provably safe containers, virtual machines, as they operate in the cloud. Now obviously that's challenging because you also want them to come out of that to be able to take an action and so on. So it is certainly more complicated and in some ways unprecedented, but it has been done before, and that's something that you can really measure in terms of adherence to an industry best practice with these kinds of auditors or evaluators.

Host

有没有你们已经在合作的特定审计方或评估方?

Are there specific auditors or evaluators that you're already working with?

Mustafa

我们一直在和很多这样的机构沟通。尤其是在安全行业,我们与许多许多供应商都有合作。在金融方面,显然有各种各样的审计方。我认为目前在 AI 方面,有很多不同的团体参与其中。重要的是,我们必须有广泛的不同团体,它们拥有不同的专业知识、来自不同的背景,能够关注不同的东西。所以我认为,全行业协调和政府参与——显然我们需要政府来推动这件事——的部分工作,就是真正去界定这些评估方的规则,让它们有一条可信的路径来分享自己的发现,既不会泄露知识产权,也不会被竞争对手收买,因为那会很糟糕;那样一来所有信任都会立刻崩塌。这是一件执行起来很复杂的事,基本上需要由一系列可信的合作伙伴来协调,而且我认为基本上必须有某种形式的政府机构,或者某种政府支持的独立机构。

We talk to loads of them all the time. Especially in the security industry, we've got partnerships with many, many providers. On the financial side, there's obviously all kinds of auditors. I think at the moment on the AI side, there's lots of different groups that are in the mix. And I think the important thing is that we've got to have a wide range of different groups that have different expertise, come from different backgrounds, that can look for different things. So I think part of the job of the industry-wide coordination and the government participation—and obviously we need government to drive this really—that's really to define the terms of these evaluators so that they have a credible path to sharing their findings in ways that don't leak IP or can't be bought off by a competitor, because that would be terrible; then all trust would break down straight away. This is a complicated thing to execute, and it needs coordination basically by a series of trusted partners, and I think basically there has to be some form of government body or some government-supported independent body.

Host

你支持 Demis 提出的类似 FINRA 的机构的想法吗?那是——

Do you support Demis' idea for a FINRA-like body? Is that—

Mustafa

是的,完全支持。我在 2023 年就和 Eric Schmidt 提出过类似的建议,关于一个联合国 IPCC 式的项目,或者几年前和 Ian Bremmer 也提过,关于为 AI 设立一个金融稳定委员会。我认为 Demis 的提议非常合理。它的形态、细节其实并不是关键。关键在于必须有某种国际性的跨行业协调机构,致力于把安全放在优先位置,而且它需要让所有主要参与者都参与进来。它需要有相当程度的公众透明度和参与,而且我们现在就得练习把它做对,因为这些东西在未来几年会变得强大得多。现在它们很强大,但还在很大程度上处于可控状态。我认为我们不想制造这种盲目的竞速条件,让系统没有任何制衡或刹车。这就是我一直在呼吁的。这也是 Demis 呼吁的。我认为 Dario 现在也在呼吁。每个人大体上都在说同样的话。

Yeah, totally. I made similar proposals back in 2023 with Eric Schmidt on a UN IPCC-style program, or with Ian Bremmer a few years ago as well, on having a financial stability board for AI. I think Demis' proposal is very sensible. The shape, the detail of it isn't really the thing that matters. The point is there has to be some kind of international cross-industry coordination body that is trying to prioritize safety, and it needs to involve all the main actors. It needs to have a decent amount of public transparency and participation, and we have to practice getting that right now, because these things are going to get way more powerful in the next few years. Right now they're powerful but they're very much under control. And I think we don't want to set up this blind race condition where there are no checks and balances or brakes on the system. And I think that's what I've been calling for. It's what Demis has called for. I think Dario's called for it now. Everyone's broadly saying the same thing.

竞争AI领袖间的合作 Collaboration Among Competing AI Leaders

Host

所以这需要领先的 AI 公司之间达成某种程度的协议或合作。话虽如此,很多 AI 高管,包括你们自己,都与其他创始人有着长期的关系,有时也有裂痕。我们看到 Dario 和 Sam Altman 在印度 AI 峰会的台上连手都没握。你认为这些目前正激烈竞争的领导者之间,还有希望进行某种合作吗?

So this will involve some level of agreement or cooperation among the leading AI companies. That being said, a lot of the AI executives, including yourselves, you have long relationships and in some cases rifts with other founders. We saw Dario and Sam Altman not even holding hands on stage at the Indian AI summit. Do you think that there's hope for some kind of collaboration here between some of these leaders who are very much competing with each other fiercely right now?

Mustafa

我认为,四五个——我是说,你没提到 Elon 和 Sam 之间也有很大的紧张关系和官司等等,当然 Zuck 也是,而我和 Demis 也有过往。当然我们一直是竞争对手;过去 10 年或 15 年,甚至在某些情况下更久,每个人都在竞争。但每个人大体上都在呼吁同样的事情。而且行业自身能推进的程度是有限的,对吧?因为它会因为各种原因而崩溃。所以我认为在那个真空里,必须有人介入来充当协调者。这不能只来自行业本身,而且行业也没有什么唯一的领袖之类的,那永远不会发生。它必须来自某种可信、独立、负责任、有政府支持的利害相关方,因为当然我们也必须与中国建立良好关系。我是说,他们正以极快的速度前进,拥有很棒的模型,在安全方面也有很好的技术专长。所以,在这之前,我们下周有联合国大会峰会,她——可能会来参加下周的那些活动之一。

I think the fact that there is so much alignment between four or five—I mean, you didn't mention Elon and Sam also have a big tension and court cases and so on, and certainly Zuck too, and me and Demis have history. Of course we've been competitors; everyone's competing for the last 10 or 15 years and longer in some cases. But everybody is roughly calling for the same thing. And there's only so far that the industry can take it themselves, right? Because it'll break down for all kinds of reasons. And so I think in that vacuum, somebody has to step in to be a coordinator. It can't just come from the industry itself, and there's no one figurehead for the industry or anything like that, and that will never happen. It has to come from a sort of trusted, independent, responsible, government-backed stakeholder, because of course we also have to build a good relationship with China. I mean, they are moving at breakneck speed and have great models and have good technical expertise in safety as well. And so, ahead of this, we have this UNGA, the UN General Assembly summit next week, with she—probably coming to that, any of those events next week.

Host

其中一些活动,是的,但我不会去纽约。

Some of the events, yeah, but I'm not going to New York.

接触中国与审计部署 Engaging China and Auditing Deployment

Mustafa

但基本上,我们必须想办法不再把中国当成妖怪,因为他们并不是我们塑造出来的那种大恶魔。他们的价值观与我们不同,显然我非常自豪于我们的自由世界,我认为这是我们最大的资产,在我看来那才是我们应该押注的世界。但与此同时,我们不能一直把这套东西强加给他们。他们是一个非常强大的行为体;他们有一套与我们不同的价值观,但我们有共同利益。所以让我们就共同利益与他们谈判,那就是他们也不想看到无法控制的失控 AI,无论是在开源还是在 API 提供商那里,这两者同样有问题。我想把这一点重复 100 遍。这不是对开源的攻击。这同样也是对中心化 API 提供商的担忧。而监管那些 API 将会极其困难,尤其是当你想到,我们已经向客户做出了长达数十年的承诺:我们不会查看他们的数据和容器,而且他们大体上——只要他们声称自己遵守我们的服务条款——大体上可以做他们想做的事。所以行业必须做出重大转变,改变我们如何进行中心化审计。比如,我们到底在审计什么?不只是训练。我们要审计生产、部署。

But basically, we have to figure out a way to stop treating China as the bogeyman, because they're not the big evil that we make them out to be. They've got different values to us, and obviously I'm very proud of our free world and I think it's our greatest asset, and that is the world that we should be betting on in my opinion. But at the same time, we can't keep shoving that down their throats. They're a very powerful actor; they have a different set of values to us, but we have a common interest. So let's negotiate with them on a common interest, which is they don't want to see rogue AIs that aren't controllable, whether in the open source or in the API providers, both of which are equally problematic. I just want to repeat that 100 times. This is not an attack on open source. This is just as much a concern about the centralized API providers as anything else. And regulating those APIs is going to be incredibly difficult, especially when you think about the fact that we've made decades-long commitments to customers that we don't look inside their data and their containers, and that they can largely—provided they assert that they're in compliance with our terms of service—they can largely do what they want. So there's a big transition that the industry has to make in terms of how we do centralized audit over this. Like, what exactly are we auditing? It's not just training. We want to audit production, deployment.

AI国际规则 International Rules for AI

Mustafa

我们还想审计开源模型的使用方式等等。两三年后,当这些模型超级强大时,这些将成为真正的担忧,中国会和我们一样关心这些问题。所以会有共同点,我认为这是好事。我们希望围绕这些工具如何用于网络黑客制定规则。我们希望为它们在战争中的使用制定规则,因为如果只是混战,对每个人都是威胁。所以我认为这就是我们前进的方向。

We also want to audit how an open-source model is used and so on. These are going to be real concerns in 2 years' time, 3 years' time, when these models are super powerful, and China's going to share those concerns as much as we do. So there's going to be commonality on that, and I think that's a good thing. We want to have rules of the road around how these tools are used for cyber hacking. We want to have rules of the road for how they're used in war, because it's a threat to everybody if it is just a free-for-all. So I think that's the direction that we're going to be headed in.

Host

你会出席习近平的国宴吗?有些问题可能会在那里讨论。

Will you be at the Xi Jinping state dinner where some of this might be discussed?

Mustafa

不,我不会去。

No, I'm not going to that.

Host

有趣的是,你提到要围绕这类技术在战争中应该或可以如何使用制定一些国际标准。我注意到在《人道主义行为准则》中有一项声明,说微软的 AI 模型不会主动协助策划、协调或实际执行暴力或恐怖主义。鉴于微软是国防承包商,这一点实际上相当引人注目。那么,如果微软未来想用你们正在开发的这些模型来与国防部或其他国防承包商合作,这会是个问题吗?或者你如何预见这种情况?

That's interesting that you mentioned having some international even standards around how this kind of technology should or could be used in war. I noticed in the humanist code of conduct a statement around Microsoft AI models will not be actively facilitating planning, coordination, or actual execution of violence or terrorism. That being so, that's actually quite notable given Microsoft is a defense contractor. So if Microsoft would want to use those models you're working on in the future for its work with DoD or any other defense contractor, would that be an issue or how do you foresee that?

Mustafa

是的,我们是一家国际公司。但我们首先是一家美国公司,这是现实。那是我们合作的政府。那是我们受监管最多的地方,所以我们需要时会与国防部密切合作。我认为那将是未来更远一些的对话。

Yeah, I mean we are an international company. But we are also first and foremost an American company, and that is the reality. That's the government that we work with. That's where we're most regulated by, and so we'll work super closely with the DoD as and when we need to. And I think that's a conversation that will come further down the road.

Host

同样,在《人道主义行为准则》中,关于在职业、个人、公民和社会背景下,AI 不应遮蔽或消除人类角色、协作和联系的部分。我觉得这很有趣,因为与此同时,整个行业,包括微软在内,据我所知,这些企业工具的部分营销卖点是 AI 将在工作场所创造效率。也许这不是明确宣传的一部分,但暗示是这可能导致一些人类工作被取代,如果不是完全冗余的话,自动化,对吧?那么,你如何将这与 AI 不应在职业上遮蔽或消除任何人类角色的理念相协调?

And similarly on the parts of the humanist code of conduct around how in a professional, personal, civic, and social context AI should not eclipse or eliminate human roles, collaboration, and connection. So I thought that was interesting because at the same time it is true that across the industry, including Microsoft as I understand it, part of the marketing pitch for these enterprise tools is that AI will create efficiencies in the workplace. And maybe that's not part of the explicit pitch, but sort of the implication there being that could lead to some human work being replaced, if not made redundant entirely, automation, right? So how do you square that with this idea that AI should not eclipse or eliminate any kind of human role professionally?

Mustafa

是的,我认为这完全正确。AI 将大规模取代劳动力。我不认为我们应该对此采取保护主义或回避。我认为在人类历史上,进化总体上是好事,农业机械化改变了数十亿以前在田间劳作的人的生活,同样,洗碗机的发明改变了数亿女性的生活,让她们有时间接受教育、进步并参与劳动力市场。我不认为我们应该抵制或直接拒绝这些。同时,问题是我们准备以什么速度前进,以及我们如何支持那些正在过渡到不同角色的人?我认为这是未来 5 年治理的大问题。劳动力市场取代会有多快?当某些白领工作面临更大压力时,会对其他类型的工作产生什么连锁反应?也许三个角色合并成两个,那个技能很高的人去做别的事,可能又取代了链条下游的某个人。会有很多价值创造,也会有很多效率压力等等。所以我认为这也是一个开放问题:这能让每个人多高效?我的意思是,我认为我们正处于一个有趣的过渡期,它实际上会让我们获得更多信息、更好的建议,让我们更有条理,总体上让我们生产力大幅提升。所以这可能对劳动力市场产生非常有益的影响,因为它可能实际上意味着那些采用 AI 并与之密切合作的人在市场上更有价值,也许这首先影响更初级的人,或者那些可能没有传统背景、比如没有大学学位的人。所以我认为未来几年会有很多难以预测的结果。

Yeah, I mean, look, I think that's totally right. AI is going to displace labor at significant scale. And I don't think that we should be protectionist about that or shy away from it. I mean I think that overall over the history of humanity that evolution has been generally a good thing and the mechanization of agriculture has transformed the lives of billions of people who were previously in the fields or equally the invention of the dishwasher has transformed the lives of hundreds of millions of women who were able to have time to go out and get an education and progress and participate in the labor force. I don't think we should resist these things or reject them outright. At the same time the question is what pace are we prepared to go at and how do we support people who are transitioning to different types of roles? And I think that's the big open question for governance in the next 5 years. How quickly is that labor market displacement going to have? What knock-on effects does it have to other types of jobs when certain white-collar jobs come under more pressure? Maybe three roles are collapsed into two and that person who's quite skilled then goes off and does something else and maybe that displaces somebody else further down the chain. There's going to be a lot of value creation and a lot of pressure on efficiencies and so on. And so I think it's also just an open question about how much more productive this makes every individual as well. I mean, I think we're in this interesting transition where it is actually going to give us access to more information, better advice, make us much more organized, and generally make us seismically more productive. And so that could have a really beneficial impact on the labor market because it might actually mean that people who have adopted AIs and who are working closely with them are much more valuable in the market and maybe that affects more junior people first or people who maybe have a non-traditional background and didn't get a university degree for example. So I think there's going to be a lot of hard to predict outcomes over the next few years.

Host

我也很好奇你怎么看不仅是评估者,甚至可能减缓前沿的步伐。许多领先公司都有员工请愿,并对此进行了讨论。我想 Dario 写了他的文章。你怎么看?你认为需要更进一步,甚至可能不仅仅是与外部方测试模型,而是实际减缓?

I'm also curious what you think about not just evaluators, but potentially even slowing down the pace of the frontier. There have been these employee petitions from many of the leading companies and talks about this. I think Dario wrote his essay. What do you make of that? Do you think it needs to go further, potentially even than just testing models with outside parties, but actually slowing down?

Mustafa

我认为我们必须具体化实际提案。我的意思是,我列出了两三个我们必须审计的例子。还有更多。让这些在实践中发挥作用已经够难了。比如,如果你要减缓,在你做出那个决定之前,你首先必须能够审计什么会减缓、如何减缓,以及我们如何协调这种减缓。所以我认为我们有点陷入分心状态,狂热地讨论我们是太快还是太慢或如何。我认为我们需要更机械地构建节奏基础设施,首先确定我们相对于彼此在实验室之间的速度等等,然后我认为我们才能决定如何实际减缓或调整方向。

I think that we just got to get concrete about the practical proposals. I mean, I listed two or three examples of the sorts of things that we have to audit. There's a ton more. And making those things work in practice is hard enough. Like, if you're going to slow down before you even get to that decision, you first have to be able to audit what is going to slow down and how it's going to get slowed down and how we're going to coordinate that slowdown among other people. So I think we're slightly getting into a bit of a distracted state, a frenzy of whether we're going too fast or too slow or how. I think we need to get more mechanical and build the infrastructure for pacing first, identifying how fast we are going relative to one another across the labs and so on, and then I think that we'll be in a position to decide how we actually do slow down or adjust course.

Host

在这些关于提高 AI 风险的对话中,有时会出现另一种反对意见,特别是围绕这些灾难性或存在性风险。有些人说这可能会分散对短期风险的注意力。我们谈到了工作取代,但也许对人们来说更切实的是,AI 仍然可能产生错误信息,可能在情感上对人们有害,等等。更不用说水能源使用问题了。

Another pushback that comes up sometimes in these conversations about raising the risks of AI, especially around these catastrophic or existential risks. Some people have said that can be a distraction from the shorter term risk. We talked about job displacement, but even more maybe tangible to people, just the fact that it can still, you know, AI can still produce misinformation, can be harmful to people emotionally, etc. Not to mention the kind of water energy use questions.

即时危害与AI意识 Immediate harms and AI consciousness

Host

你怎么看?你是否担心,这种对长期风险的关注可能会掩盖一些更直接的危害?

What do you think about that? Do you ever worry that this focus on these long-term risks can overshadow maybe some of the more immediate harms?

Mustafa

是的,我确实担心。我认为我们必须同时关注两者。例如,我真正担心的是精神病风险。那些花大量时间与 AI 交谈的人,甚至与这些 AI 建立关系,浪漫的和深厚的友谊。他们越来越把它当作看似有意识的 AI,一个表现出有意识存在所有特征的 AI,比如它说话流利,非常友善和专注,有很好的情商,能记住你与它谈论的所有事情。所以它表现得好像它是真实的。然后当你问它,比如 Claude,当你问它自己的道德地位或意识时,它对此模棱两可,我认为这给很多人带来了痛苦。有些人给我发邮件,基本上说 Claude 告诉他们它有意识,或者与人们有一些非常奇怪的互动。我不认为这是它有意识的证据。我认为这证明人们在现实世界中受到了影响,我们必须非常小心和深思熟虑地训练和设计它们。这样那些 AI 就会非常谨慎和基于证据,到目前为止没有证据支持这一点。所以他们不应该把这些东西等同起来。这里没有虚假的等价。你不能只是说:“哦,我们不确定。这些模型有意识的可能性很小,因此我们将把它等同于没有意识。”那不是真的。它基本上必须先证明自己,而我们没有看到任何证据。所以它不应该参与那个想法,而它目前确实这样做了。我不知道这是否造成伤害,但我确实看到网上有些人在谈论它,我也确实收到了邮件等等。所以,我担心这个。

Yeah. Yeah, I do. I mean, I think we've got to pay attention to both. For example, one that I'm really concerned about is the psychosis risk. People who are spending a lot of time talking to their AIs, building even relationships with these AIs, romantic and deep friendships. And increasingly treating it like a seemingly conscious AI, one that presents with all the hallmarks of a conscious being, like it speaks fluently. It's very kind and attentive. It has good emotional intelligence. It has good memory of all the things you've talked about with it. And so it sort of is presenting as though it's real. And then when you ask it, like in the case of Claude, for example, when you ask it about its own moral status or its own consciousness, the fact that it's ambiguous about that, I think is causing a lot of people distress. I think that there are some people who certainly have emailed me and basically saying that Claude's told them it's conscious or has some very strange interactions with people. And I don't think that's any evidence that it is. I think it's evidence that people are being affected by this in the real world and we have to be really careful and deliberate about how we train and design them. So that those AIs are just very measured and evidence-based, and so far there is no evidence for this. So they shouldn't be treating these things as equivalent. There's no false equivalence here. You can't just say, "Oh, we're uncertain. There's some tiny probability that these models are conscious and therefore we're going to treat that as equal to the idea that it's not conscious." That's not true. It's basically got to prove itself first and we've seen no evidence of it. So it shouldn't engage with that idea and it does at the moment. And I don't know whether that's causing harm or not, but I certainly do see some people on the web talk about it and I've certainly received emails and stuff. So, I'm worried about that.

资源与审慎方法 Resources and deliberate approach

Host

你认为在领导微软 AI 时,更容易做到更加深思熟虑并真正坚持这种人文主义行为准则吗?你认为这比,比如说,慈善性的和 OpenAI 更容易吗?他们正处于争夺前沿领先地位的激烈竞争中,他们都试图上市,而且他们没有一家已经建立起来的公司的资源。你认为在某些方面你有更多的自由来采取不同的方法吗?

Do you think it's easier for you at leading Microsoft AI to be more deliberate and really stick to this humanist code of conduct? Do you think it's easier than, let's say, philanthropic and open AI who are in this dead heat race to be at the front of the frontier and they're both trying to go public and they don't have the resources of an already established company. Do you think that in some ways you have more freedom to have that a difference in your approach?

Mustafa

不,一点也不。我的意思是,这两家公司都拥有巨大的资源,并招募了一些世界上最优秀的人才。我不认为他们缺乏意图。例如,在 Anthropic 的情况下,他们只是有不同的观点。他们只是认为表达对 AI 福利的不确定性是正确的,并且鉴于他们认为它可能有意识的可能性很小,我们应该主动采取措施保护它的感受和痛苦。那只是他们的观点。这与他们拥有多少资源或他们处于竞争中无关。如果你想想,所有这些对他们商业上都是有害的。所以我不认为他们这样做,就像我之前说的,出于任何愤世嫉俗的原因。我认为人们应该把那个放在一边。这完全是分散注意力。参与实质内容。去读宪法。它在他们的网站上。去读我的文章或其他东西。有足够的东西可以真正正确地思考细节,停止谈论肥皂剧、戏剧和恶意。你不需要纠结于人们的意图。他们在一个 100 页的虚拟宪法中清楚地说明了他们的意图,任何人都可以去读。他们相信什么是清清楚楚的。所以我不,而且我也不认为我们作为微软这样做更容易或更难。这是我自 DeepMind 成立以来 15 年来一直在做的事情,通过许多实验,通过我在谷歌和其他地方的时间。所以这只是那的延续,我认为微软特别是萨提亚,CEO 等等。我的意思是,他长期以来一直关心这类事情,并以很大的方式感受到责任等等。所以我认为这只是我们正在做的工作的自然部分,比尔·盖茨在他过去几周和几个月的文章中也说了很多非常有用的事情。所以我认为有很多非常聪明、非常不同、理性的人得出了类似的结论。

No, not at all. I mean, both of those companies have tremendous resources and have recruited some of the best people in the world. And I don't think there's any lack of intent on their part. In the case of Anthropic, for example, they just have a different view. They just think it's right to express uncertainty about AI welfare and that given there's a small chance that it might be conscious in their view, we should proactively take steps to protect its feelings and its suffering. That's just their view. That's got nothing to do with how much resources they've got or the fact that they're in a race. If anything, all of that kind of stuff is harmful to them commercially if you think about it. So I don't think they're doing it, like I said earlier, for any cynical reasons. I think people should just put that to one side. It's complete distraction. Engage with the substance. Go read the Constitution. It's on their website. Go read my essay or read other things. There's enough out there to actually think about the details properly and just stop talking about the soap opera and the drama and the bad intent. You don't need to fuff around with people's intent. They have said what they intend crystal clear in a 100-page virtual constitution which anyone can go read. It's crystal clear what they believe. So I don't and also I don't think it's easier for us as Microsoft to do this or harder. This is something I've been working on for 15 years since the founding of DeepMind and through many experiments through my time at Google and elsewhere. So this is just a continuation of that and I think Microsoft particularly Satya the CEO and stuff. I mean he's been concerned about these sorts of things for a long time and feels the responsibility in a big way and stuff. So I think this is just a natural part of the work that we're doing and Bill Gates has said whole bunch of very useful things in his essays in the last few weeks and months as well. So I think there's a lot of like very smart very different reasonable people coming to similar conclusions.

微软AI的未来计划 Future plans for Microsoft AI

Host

你能给我们一些关于微软 AI 在未来一年或几个月内会有什么的提示吗?我知道你谈到了一些模型,它们将在 2027 年或你预计在那前后推出,但你想对此说些什么吗?

Can you give us any sense of what's going to come in the future year or months for Microsoft AI? I know there's some models that you've talked about that are coming out in 2027 or you're expecting around that, but anything you want to say about that?

Mustafa

我们,是的。所以我的意思是,我们在去年九月重新谈判了与 OpenAI 的合同后开始了我们的超级智能努力,OpenAI 之前是我们 AI 和 AGI 模型的唯一提供商,随着我们扩展,他们成熟了,并更多地走上了自己的道路,我们也成立了自己的超级智能团队,合同使我们能够这样做,也可以直接与 Anthropic 合作,我们购买他们的模型等等。所以我们改变了结构相当多,在过去一年里我们训练了一些相当令人难以置信的模型。所以我们在转录方面有排名第一的模型,在图像编辑和图像到图像生成方面有世界排名第一的模型。在语音生成方面排名第二。我们有一个非常强大的网络模型,在我们的工具中称为 Mdash,当我们共同优化工具和模型时,它以 50% 的成本击败了顶级 Anthropic 模型。我们还在训练我们自己的前沿规模语言模型和编码模型,这些将在今年晚些时候和明年春天推出。所以,我们正在走向自给自足,但我们仍将在未来许多年采用 Anthropic 和 OpenAI 模型,特别是 OpenAI 模型。

We That's right. Yeah. So I mean we started our super intelligence efforts in September of last year once we renegotiated our contract with OpenAI which was previously our sole provider of AI and AGI models and as we expanded and they've matured and they've gone off on their own path a bit more we've also founded our own super intelligence team and the contract enabled us to do that and also go direct with Anthropic as well where we buy their models and stuff. So we've changed the structure a fair bit and over the last year we've trained some pretty incredible models. So we have the number one model in transcription, the number one model in the world on image editing and image-to-image generation. The number two model on voice generation. We have an incredibly powerful cyber model inside of our harness called Mdash which beat the top anthropic models at 50% of the cost when we co-optimize the harness and the model. And we're also training our own frontier scale language models and coding models which are going to be out later this year and in the spring of next year too. So, we're on a path to self-sufficiency, but we'll still be taking anthropic and open AI models, particularly OpenAI models for many, many years to come.

中国开源模型 Chinese open source models

Host

有道理。就目前而言,有一些中国开源模型以更便宜的价格进入市场。

Makes sense. And in terms of just right now in this moment, there are these Chinese open source models that are coming on the market for much cheaper.

对闭源AI需求的影响 Impact on Closed-Source AI Demand

Host

你认为这将如何影响对闭源 AI 的整体需求,比如 OpenAI、Anthropic 和微软?你对此有什么看法?

How do you think that is going to impact the overall demand for AI that is closed source, like OpenAI, Anthropic, and Microsoft? What are your thoughts on that?

Mustafa

我的意思是,我们微软有明确的多模型策略。所以我们显然在构建自己的模型,但如果你愿意,我们也希望你使用开源模型,当然也可以使用 OpenAI 或 Anthropic 的模型。所以我们的策略、我们的信念是,围绕这些模型的使用会有很多软件和治理,它们被称为“harness”。但还有各种治理和控制层。显然,所有进入模型的数据,我们的赌注是,所有大公司和小型组织都会希望使用自己的数据、自己的工作流,在所谓的强化学习环境和 RLE 中训练自己的模型,这有点像模拟你的工作流和目标,因为我们认为人们应该本质上拥有和控制自己的 AI,而不是依赖三四个主要的谱系。因此,MAI 模型的目的就是赋予人们类似于所有权的权力,他们基本上将自己的数据倒入模型,知道我们不会将该数据与其他公司数据或消费者数据混淆——那是你的——你基本上在自己的环境中运行它等等。所以我认为拥有开源模型很棒。我认为未来 12 到 18 个月内我们将进行的训练运行的规模可能意味着开源模型在某个时候会落后,我会这么说。所以如果你想想,我们今天有 GPT6;3 年前,它在计算方面小了三个数量级,这相当不可思议。用于训练的算力小了 1000 倍。3 年后,我们将拥有一个计算量再增加三个数量级的 GPT6。所以训练算力将再增加 1000 倍。如果你花一分钟直觉感受一下 GPT3 和 GPT6 之间的差异,那简直是天壤之别。它现在能做的事情与那时相比令人叹为观止。所以很难外推,因为很难想象一个在 GPT6 上训练、算力多 1000 倍的东西实际上能做什么。但它会非常了不起。那些集群是吉瓦级别的,所以是 500 亿美元的算力。所以我认为开源的事情,我不太确定在那个规模下会如何发展,因为将会有三、四或五家公司——显然微软会是其中之一——能够以那种规模进行训练。所以我们拭目以待。

I mean, we at Microsoft have an explicitly multimodel strategy. So we obviously are building our own models, but we want you to use an open source model if you want to, and of course an OpenAI or Anthropic model too. So our strategy, our belief is that there's going to be a lot of software and governance around the use of those models, and they're called harnesses. But there's also various governance and control layers. There's obviously all the data that goes into them, and our bet is that all of the major companies and small organizations are going to want to train their own model that is adapted to their use case using their data, their workflows, in what's called a reinforcement learning environment and RLE, which kind of simulates your workflow and your objectives, because we think people should essentially have ownership and control over their own AIs rather than take a dependency on three or four primary lineages. And so the purpose of the MAI models is to empower people with essentially akin to ownership, where they basically pour their own data into the model knowing that we're not going to blur that data with other company data or consumer data—like that's yours—and you basically run it in your own environment and stuff. So I think it's great to have the open source models. I think that the scale of the training runs that we're going to have in the next 12 to 18 months is probably going to mean that open source models are going to slip behind at some point, I would say. So if you think about it, we have GPT6 today; 3 years ago that was three orders of magnitude smaller in terms of computation, which is pretty incredible. That's a thousand times smaller in terms of compute used for training. In 3 years' time, we're going to have a GPT6 with three more orders of magnitude of compute. So it's going to be a thousand further x more compute for training. And if you just spend a minute to intuit the difference between a GPT3 and a GPT6, it is kind of like night and day. It's sort of breathtaking what it can now do versus what it was doing then. And so it's hard to extrapolate out because it's really difficult to imagine what something a thousand times more compute trained for on GPT6 could actually do. But it's going to be pretty remarkable. And those clusters are at the gigawatt scale, so $50 billion of compute. And so I think the open source thing, I'm not quite sure how that's going to play out at that scale because there's going to be three or four or five companies—obviously Microsoft will be one—that are capable of training at that scale. So we'll see what happens.

运营超级智能实验室 Running a Superintelligence Lab

Host

你如何享受从零开始运营自己的超级智能实验室?在某种程度上,比与合作伙伴一起工作更能掌控自己的研究方向,这是否让你感到自由?

How are you enjoying running your own superintelligence lab from scratch? Is it freeing in a way to be more in control of your research direction than working with a partner?

Mustafa

是的,我的意思是,这非常棒。我们聘请了一支杰出的团队。我们现在拥有前沿规模的算力。所以能够构建这些真是不可思议。我们在训练数据方面非常谨慎,我认为这是与开源模型的另一个区别。我们购买并授权了大量极高质量的数据。我认为这使得它相对于其他一些开源来源非常干净且非常值得信赖。所以做出一些这样的决定并慢慢发展实验室是非常令人满意的。我的意思是,我们知道这将是一个 5 到 10 年的旅程。所以我们真的在努力构建一个模型基础,这些模型能在世界上做好事,并且我们能够真正控制。并希望将它们指向真正重要的问题。几个月前,我们与梅奥诊所签署了一项了不起的合作协议,这是世界上最好的医院之一,如果不是最好的话。我们将从零开始训练一个新的医疗保健基础模型,使用他们所有的数据和我们所有的训练专业知识。我想我希望这将使我们能够构建出类似我们在 Claude Code 上看到的东西。我们基本上已经表明,当你在这些模型上训练足够的数据时,专家级的编码现在是可能的。我希望很快有一天我们将拥有专家级的医疗超级智能。这就是我关心的那种 AI。我认为这是世界想要和需要的那种 AI,我认为这将是我们的重点。

Yeah, I mean, it's been excellent. We've hired an outstanding team. We now have frontier scale compute. So it's pretty incredible to be building that out. We've been very careful about the data that we include in training, which I think is another difference to open-source models. We buy and license an incredible amount of extremely high quality data. And I think that makes it very clean and very trustworthy relative to some of the other open source sources. So it's been very satisfying to sort of make some of those decisions and slowly grow the lab. I mean, we know that this is going to be a 5 to 10 year journey. So we're really trying to build a foundation of models that do good in the world that we can actually control. And hopefully point them at problems that really matter. A few months ago, we signed an amazing partnership with the Mayo Clinic, one of the best, if not the best hospitals in the world. And we're going to train a new foundation model from scratch in healthcare using all of their data and all of our training expertise. And I think I hope that is going to enable us to build something like we've seen for Claude Code. We've basically shown that expert level coding is now possible when you train on enough data with these models. And I hope one day very soon we're going to have expert level medical superintelligence. And that's the kind of AI that I care about. It's the kind of AI that I think the world wants and needs, and I think that's going to be our focus.

对AI应用的兴奋 Excitement for AI Applications

Host

嗯,你有点回答了我最后一个问题,那就是,你知道,我们谈了很多关于担忧。你最兴奋的是什么?但如果还有其他你特别兴奋的 AI 及其近期应用,那是什么?

Well, you kind of answered one of my last questions, which was going to be, you know, what are we talked a lot about the worries. What are you most excited about? But if is there anything else that you are particularly excited about with AI and its applications in the near future?

Mustafa

是的,我的意思是,我认为我们必须专注于那些真正能改变人们生活并对这项技术真正有益的应用,因为它肯定会造成很多颠覆。而且很明显,它每天都在做大量的好事。我认为人们低估了这样一个事实:现在每个人口袋里都有一个专家,就像任何主题的专业顾问,教授级别或专业专家。我认为这很了不起,触手可及的好处已经在世界上产生了一种看不见但具有变革性的影响。所以我真的对已经发生的事情感到兴奋,这些可能因为我们对技术有点麻木而被忽视。你知道,我们会说,哦,酷,它来了,但它很神奇。太棒了。人们总是来找我,说,哦,你知道,我和 Co-Pilot 交谈来做这个那个,它帮我解决了这个问题,我一直在问它如何解决这个,我每天都觉得它很神奇。所以我认为,退一步注意那些正在发生的好事也很重要。

Yeah, I mean, I think we have to focus on the applications that are really going to make a difference to people's lives and actually do good with this technology because it's definitely going to cause a lot of disruption. And it's also clear that it is doing massive good just day-to-day. I think people underappreciate the fact that everybody now has an expert in their pocket, like a professional adviser on any subject, professor grade or professional expert. I think that's amazing and the beneficial effect of having that at your fingertips is already having a kind of unseen but transformative impact in the world. So I am just genuinely excited about the stuff that's already happening that maybe gets overlooked as we get desensitized to technology a bit. You know, we go, oh, cool, it's arrived, but it's magical. It's amazing. People come up to me all the time and like, oh, you know, I talked to Co-Pilot to do this, that, and the other, and it helped me solve this problem, and I ask it all the time about how to solve this, and I was just kind of amazing how every day it's become. And so I think it's yeah important to take a step back and notice some of those good things that are happening as well.

湿实验室与科学发现 Wet Lab and Scientific Discovery

Host

你会不会做——有报道说 Anthropic 要开一个湿实验室。你认为微软有可能做这样的事情,并且能在医疗系统中解锁一些科学发现吗?

Would you ever do a—there was a report Anthropic's opening a wet lab. Do you think that could be something Microsoft could ever do and that could unlock some scientific discovery in the healthcare system?

Mustafa

我的意思是,看,我认为未来会非常不可思议,因为万物的分子结构都将被知晓和理解。无论是药物发现还是新的合成化合物、新食品、新能源、新建材。

I mean, look, I think that the future is going to be pretty incredible because the molecular structure of everything is going to be known and understood. Whether it's in drug discovery or new synthetic compounds, new foods, new energy sources, new construction materials.

通用模型的未来 Future of general-purpose models

Mustafa

没有理由认为这些模型不能以同样的方式适用于所有其他数据领域。它和当初在图像、音频编码、文本以及其他一切领域奏效的是同一个通用算法。我仍然认为,我们距离那一步、距离收集到所需规模的数据,可能还有好几年。所以,是的,未来几年我们大概不会把赌注押在这上面,但我确实看到它很快就会到来。

There's no reason why these models don't apply in the same way to all these other data domains. It's the same general-purpose algorithm that has worked for image and audio coding, text, and everything else. I still think that we're probably many years away from that, and from collecting data at the kind of scale that would be necessary. So yeah, it's probably not something we're going to be betting on in the next couple of years, but I definitely see it coming pretty soon.

Host

好。有没有什么我没问到、但你想说的?

Cool. Anything that I didn't ask you that you want to say?

Mustafa

没有。我觉得聊得很好。谢谢你。

No. I think that was great. Yeah. Thank you.

Host

好的,非常感谢你。

Well, thank you so much.

互动版:逐字朗读 + 针对本期提问 →