从操作系统到 AI 代理:穆斯塔法·苏莱曼谈微软的范式转变

From OS to AI Agents: Mustafa Suleyman on Microsoft's Paradigm Shift

穆斯塔法·苏莱曼 Mustafa Suleyman · Moonshots(彼得·戴曼迪斯) · 2025-12-16 · 约 85 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

微软 AI 首席执行官穆斯塔法·苏莱曼探讨从操作系统和应用程序向 AI 代理和助手的转变,强调目标并非 AGI 竞赛,而是人机交互方式的根本性变革。

Mustafa Suleyman, CEO of Microsoft AI, discusses the transition from operating systems and apps to AI agents and companions, emphasizing that the goal is not a race to AGI but a fundamental shift in how we interact with technology.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 29)

全文 · Full transcript(中英对照)

引言与使命 Introduction and Mandate

Host

萨提亚给你的使命是什么?是赢得 AGI 吗?

What's the mandate from Satya? Is it win AGI?

Mustafa Suleyman

我不认为真的有赢得 AGI 这回事。我也不确定存在一场竞赛。

I don't think there's really a winning of AGI. I'm not sure there's a race.

Host

作为 AI 领域的元老之一,穆斯塔法·苏莱曼现在是微软 AI 的 CEO。他在这个行业前沿耕耘了十多年,远早于我们最近几年才感受到的 AI 热潮。

One of the OGs of the AI world, Mustafa Suleyman is the CEO now of Microsoft AI. He spent more than a decade at the forefront of this industry before we even had gotten to feel it in the past couple of years.

Mustafa Suleyman

从根本上说,我们正在经历的转变是从操作系统、搜索引擎、应用和浏览器的世界,走向智能体和伴侣的世界。我们都在尽可能快地前进,但竞赛意味着零和博弈,意味着有一条终点线,这其实不是一个恰当的比喻。我们知道,技术、科学和知识无处不在,同时以各种规模传播。

Fundamentally, the transition that we're making is from a world of operating systems, search engines, apps, and browsers to a world of agents and companions. We're all going as fast as we possibly can, but a race implies it's zero sum. It implies that there's a finish line, and it is like not quite the right metaphor. As we know, technologies and science and knowledge proliferate everywhere, all at once, at all scales, basically simultaneously.

Host

你们是否在安全方面投入了大量精力、算力和人力?

Are you spending a lot of your energy, compute, human power on safety?

Mustafa Suleyman

是的。不,我是说……

Yeah. No, I mean...

Host

女士们先生们,这可真是个大计划。欢迎来到 Moonshots。今天我和 DB2、AWG 以及穆斯塔法·苏莱曼在一起,他是 DeepMind 和 Inflection AI 的联合创始人,现在是微软 AI 的 CEO。欢迎你,朋友。很高兴你能来。感谢你抽出时间。

Now that's a moonshot, ladies and gentlemen. Everybody, welcome to Moonshots. I'm here with DB2 and AWG and Mustafa Suleyman, the co-founder of DeepMind, Inflection AI, and now the CEO of Microsoft AI. Welcome, my friend. It's good to have you here. Thank you for making time for us.

Mustafa Suleyman

谢谢邀请。是的,我很兴奋能来。

Thanks for having me. Yeah, I'm excited to do this.

Host

是的,你和萨提亚一起打造的东西太棒了。很难相信微软已经 50 岁了,它多次重塑自我,过去五年一直处于行业顶端,是全球最有价值的公司,拥有 25 万员工,据我所知现在有 1 万人在你手下。所以我想先问几个重要问题。首先,一些宏观背景。你在一家拥有巨大资源的大公司内部进行建设,资源可能比几乎所有其他公司都多。我的问题是:最终目标是什么?所有超大规模云服务商都在提供开放 AI 访问,进行用户争夺战。你一直在微软 365 生态内建设。未来几年的目标是最大化用户数?是数据中心?是云?你如何思考优化目标?

Yeah, it's, you know, what you've been building with Satya is amazing. And it's hard to believe that Microsoft is 50 years old and it's reinvented itself so many times and for the last 5 years it's been at the top of the game, the most valuable company in the world, 250,000 employees, and from what I understand 10,000 employees now under you. So a few important questions I want to open with. First, some broad context. You're building inside a massive company with huge resources, probably arguably more than almost everybody else. And the question I have is: what's the end goal here? You've got all the hyperscalers sort of providing open access to AI and they're doing a land grab to try and get as many users as possible. You've been building within the Microsoft 365 ecosystem. Is the goal in the next couple years maximum users? Is it data centers? Is it cloud? How do you think of what you're optimizing for?

Mustafa Suleyman

这是个好问题。我们是一家市值 4 万亿美元、收入近 3000 亿美元的公司。这令人难以置信,超现实且非常谦卑。我们在技术栈的每一层都有布局。显然,我们在数据中心有巨大业务,某种程度上我们像一家现代建筑公司,数十万建筑工人每年建造千兆瓦级的 CPU 和各种 AI 加速器,并将其推向市场。之上是 API,还有你能想到的各个领域的第一方产品,从游戏、LinkedIn 到 M365 和 Windows 的所有基础,当然还有搜索和消费者业务。从根本上说,我们正在经历的转变是从操作系统、搜索引擎、应用和浏览器的世界,走向智能体和伴侣的世界。所有这些用户界面都将被整合到对话式的智能体形式中。这些模型会像你口袋里 24/7 的私人助理,能做任何事,掌握你所有的上下文。你将越来越少直接操作计算,就像我们现在看到的。许多软件工程师正在使用辅助代码智能体来调试代码和生成大量代码,就像我们过去使用第三方库一样。现在我们将用 AI 来做生成,这让他们更高效、更准确、更快。所以我们所处的轨迹相当可预测:从用户界面到 AI 智能体,这是一个范式转变,公司完全专注于这一点。在经历了五十年来的多次转型后,我认为公司非常警觉,要确保我们处于最佳位置来应对这次转型。

I mean, it's a good question. So, I mean, we're on any given day a $4 trillion company with almost $300 billion of revenue. It's incredible. It's just surreal and very, very humbling. And we play at every layer of the stack. I mean, obviously, we have an enormous business in data centers and in some ways we're like a modern construction company, hundreds of thousands of construction workers building gigawatts a year of CPU and AI accelerators of all kinds and enabling that to be available to the market. APIs on top of that, but also first-party products in every domain you can think of, from gaming and LinkedIn right the way through to all the fundamentals of M365 and Windows, and of course in our search and consumer businesses. And fundamentally, the transition that we're making is from a world of operating systems, search engines, apps and browsers to a world of agents and companions. All of these user interfaces are going to get subsumed into a conversational agentic form. And these models are going to feel like having a real assistant in your pocket 24/7 that can do anything, that has all your context. And you're going to do less and less of the direct computing, just as we're seeing now. Many software engineers are using assistive code coding agents to both debug their code and also generate large amounts of code, just as we used libraries, third-party libraries. Now we're just going to use AIs to do that generation, and it's making them more efficient and more accurate and faster and so on. So the trajectory we're on is quite predictable. It's one from user interfaces to AI agents, and that is a paradigm shift which the company is completely focused on. After seeing five decades worth of transitions, I think the company is super alert to making sure that we're best placed to manage this one.

Host

你觉得自己会像其他玩家那样提供开源 AI,还是认为可以将其保留在微软 365 内部?

Do you see yourself providing sort of an open-source AI like the other players out there or do you think you can keep it contained within Microsoft 365?

Mustafa Suleyman

我认为我们相当开放。我们有一些很小的开源模型。

I think we're pretty open-minded. I mean, we've got some pretty small open-source models.

Host

我说开源,其实是指开放访问。

When I say open source, I really mean open access, if you would.

Mustafa Suleyman

是的。听着,总会有 API 提供极其强大的模型。微软实际上是一个平台的平台。成为一个平台,成为核心基础设施的优秀提供商,让他人能够高效工作,这是公司的 DNA。所以我们总会有大量 API 来加速这一点。但 API 本身也会开始变得不同。API 和智能体本身之间的界限可能会很模糊。也许五年后,我们的主要业务是销售执行特定任务的智能体,这些智能体带有可靠性、安全性、保障和信任的认证。这实际上是微软的优势之一,也是吸引我的地方:这是一家极其值得信赖的公司。它非常安全,有时我认为缓慢或摩擦实际上是一种资产。为全球最大的财富 500 强公司、政府和主要机构提供服务,带来了一种稳定性。

Yeah. I mean, look, there are always going to be APIs that provide incredibly powerful models. I mean, you know, Microsoft is really a platform of platforms. Being a platform and being a great provider of the core infrastructure that enables other people to be productive is like the DNA of the company. And so we will always have masses of APIs that turbocharge that. But what an API is is going to start to look kind of different too. Like it may be pretty blurred the distinction between the API and the agent itself. Maybe that we're principally in the business in 5 years time of selling agents that perform certain tasks that come with a certification of reliability, security, safety, and trust. And that is actually in many ways the strength of Microsoft, and that's one of the things that attracted me: this is a company that's incredibly trusted. It's actually very secure, and sometimes I think the slowness or the friction is actually a bit of an asset. You know, there's a kind of steadiness that comes with having provided for all of the world's biggest Fortune 500 companies and governments and major institutions.

Host

这就像老话说的,过去买 IBM 不会错?

Is it like the old adage, you can't go wrong buying IBM in the old days?

Mustafa Suleyman

我认为我们有一种稳定性,让人放心,还有一种深思熟虑的、以客户为中心的耐心。没有那种作为挑战者常有的焦虑和僵化。我们的位置也有一些缺点。我们做事需要更长的时间,但公司正在全速运转,这非常令人印象深刻。

I think there's a steadiness about us which I think is reassuring to people, and there's a kind of deliberate customer-focused patience. There's not the same anxiety and somewhat sclerotic nature that comes with being an insurgent. There are some downsides to our position. We take a little longer to get things through, but the company is firing on all cylinders. It's very impressive to see.

Host

在我把问题交给 Alex 之前,再问一个。我们看到在这场超大规模云服务商战争中,每周都有人超越别人,疯狂地推出新基准。你是否怀念参与那种竞争,还是微软提供的稳定性让你能够为长期愿景而建设,这才是最让你兴奋的?你知道,我在 DeepMind 的背景让我花了整整十年在指数曲线的平坦部分挣扎,那时基本上什么都不管用。

One more question before I turn it over to Alex. You know, we're seeing in this hyperscaler war, literally week by week, everybody outdoing each other in this insane period of everybody coming out with new benchmarks. Do you miss not being in that game, or is the stability that Microsoft provides to build for a long-term vision what you find most exciting? You know, my background at DeepMind is such that I spent a good decade grinding through the flat part of the exponential where basically nothing worked.

深度学习早期 Early days of deep learning

Mustafa Suleyman

我的意思是,确实有一些了不起的论文。AlphaGo 显然令人难以置信,但它是在一个非常独特的模拟受控游戏环境中。真正能在现实世界中运行的东西少之又少。所以我一直秉持着数十年的视角,这只是我的直觉。我认为每个月推出新模型并活跃在市场上非常重要,但更重要的是为未来奠定正确的基础,因为我认为这将是我们物种有史以来最疯狂的转变。

I mean, there were some amazing papers. AlphaGo was obviously incredible, but it was in a very unique simulated controlled game-like environment. Things actually working in the real world were few and far between. So I've always taken a multi-decade view, and that's just been my instinct. I think it's super important to ship new models every month and be out there in the market, but it's actually more important to lay the right foundation for what's coming, because I think it's going to be the most wild transition we have ever made as a species.

Host

你能稍微展开一下吗?有没有一段时间只有你们三个人在伦敦埋头苦干?

Can you just flesh that out a little bit? Was there a period of time where it was just three of you grinding it out in London?

Mustafa Suleyman

嗯,不止我们三个人,但在 2010 年到 2020 年的十年间,深度学习的成功商业应用少之又少。幕后有很多,比如图像识别、搜索改进,但商业……

Well, there were more than three of us, but for the decade between 2010 and 2020, there were just so few successful commercial applications of deep learning. There were plenty behind the scenes, like image recognition, improvements to search, but commercial...

Host

商业市场巨大。下围棋不是一个大市场。

Huge market for commercial. Playing Go is not a huge market.

Mustafa Suleyman

没错。所以现在,从 2022 年开始,我们看到大语言模型投入生产,人们正在改变人际关系的内涵。这是一个转折点。这与 2010 年代用极少数据和极小集群训练小模型的艰辛截然不同。

Exactly. So now, from 2022 onwards, we see LLMs in production, people changing what it means to be human relations. That's an inflection point. It's very different from the grind of training tiny models with very little data and very small clusters back in the 2010s.

赞助商插播 Sponsor break

Host

每周,我和我的团队都会研究未来十年将改变行业的十大技术元趋势。我涵盖的趋势包括人形机器人、AGI、量子计算,以及交通、能源、长寿等。没有废话,只有最重要的、影响我们生活、公司和职业的内容。如果你想让我分享这些元趋势,我每周写两次新闻通讯,通过电子邮件发送,只需两分钟阅读。如果你想比其他人早十年发现最重要的元趋势,这份报告就是为你准备的。读者包括来自世界上最具颠覆性公司的创始人和 CEO,以及构建世界上最具颠覆性技术的企业家。如果你不想了解即将发生的事情、为什么重要以及如何从中受益,那就不适合你。免费订阅请访问 dmandis.com/metats,比其他人早十年获取趋势。好了,现在回到本期节目。

Every week, my team and I study the top 10 technology meta trends that will transform industries over the decade ahead. I cover trends ranging from humanoid robotics, AGI, and quantum computing to transport, energy, longevity, and more. There's no fluff, only the most important stuff that matters, that impacts our lives, our companies, and our careers. If you want me to share these meta trends with you, I write a newsletter twice a week, sending it out as a short two-minute read via email. And if you want to discover the most important meta trends 10 years before anyone else, this report's for you. Readers include founders and CEOs from the world's most disruptive companies and entrepreneurs building the world's most disruptive tech. It's not for you if you don't want to be informed about what's coming, why it matters, and how you can benefit from it. To subscribe for free, go to dmandis.com/metats to gain access to the trends 10 years before anyone else. All right, now back to this episode.

现代图灵测试与经济基准 Modern Turing test and economic benchmarks

Host

是的。我们上次谈话大约在 2015 年,那可能是 ImageNet 之后三年,语言模型之前五年。少样本学习、智能体、智能体式 AI 远未达到我们现在看到的水平。自从你写了你的愿景,你提出的现代图灵测试,即智能体自主性的经济基准,我很想听听:微软为这些智能体制定的经济基准在哪里?如果智能体即将接管经济或接管这么多经济上有用的功能,为什么我们还在使用像 Vending Bench 这样的基准,而不是微软率先为其智能体制定经济自主基准?

Yeah. So when last we spoke circa 2015, that was perhaps 3 years post ImageNet, 5 years pre language models. Few-shot learners, agents, agentic AI was nowhere to be seen at the level of what we see now. Since you've written about your vision, what you've socialized as a modern Turing test, the idea of economic benchmarks for autonomy by agents, I'd love to hear: where are Microsoft's economic benchmarks for these agents? If the agents are about to take over the economy or take over so many economically useful functions, why are we stuck with benchmarks like Vending Bench rather than Microsoft leading the way with Microsoft's economically autonomous benchmarks for its agents?

Mustafa Suleyman

是的,可能有必要补充一下背景:我们 2015 年在波多黎各的 AI 安全会议上见过面。现在该领域的许多人当时都在场。那是一个开创性的时刻。

Yeah, it's probably worth adding the context that we met in 2015 in Puerto Rico at the AI safety conference. Many in the field now were there at the same time. It was a seminal moment.

Host

是在除夕夜之后还是新年左右?

Was it the day after New Year's Eve or somewhere around New Year?

Mustafa Suleyman

除了波多黎各,其他地方都很冷。

It was pretty cold out everywhere except Puerto Rico.

Host

是的,没错。实际上相当超现实。

Yeah, exactly. It was quite surreal actually.

Mustafa Suleyman

就像暴风雨前的宁静。

It was like a calm before it all happened.

Host

是的,完全同意。现代图灵测试是我提出的,我想是 2022 年写的。它基本上做了一个非常简单的预测。如果缩放定律继续,有更多的数据和算力,每年给最好的模型增加一个数量级的算力,那么很明显,我们会从识别(第一波)到生成(我们现在正处于中间,或者可能结束这一章),再到每个时间步都有完美的生成,这依次会产生辅助性的智能体式行动。而行动显然会看起来像一个智能知识工作者、项目经理、战略家或初创公司创始人等等。那么我们如何衡量这种表现呢?与其用学术和理论基准来衡量,人们显然想通过能力来衡量。这个东西在经济和工作场所能做什么?我们如何衡量经济?我们用美元和美分来衡量。那么,第一个赚到一百万美元的模型会是什么?我记得,起始资本是 10 万美元。

Yeah, totally. And the modern Turing test was something I proposed, I guess it was 2022 when I wrote it. It was basically making a pretty simple prediction. If the scaling laws continue with more data and compute, adding an order of magnitude more compute to the best models every year, then it's pretty clear we would go from recognition, which was the first part of the wave, to generation, which we're now in the middle of, or maybe ending that chapter, to then having perfect generation at every time step, which in sequence is going to produce assistive agentive actions. And actions would obviously look like an intelligent knowledge worker or a project manager or a strategist or a startup founder or whatever. So how would we measure that performance? Rather than measuring with academic and theoretical benchmarks, one would clearly want to measure it through capabilities. What can the thing do in the economy, in the workplace? And how do we measure the economy? We measure it by dollars and cents. So, what would be the first model to make a million dollars? Given, as I recall, $100,000 in starting capital.

Host

没错。是的。谁能把它变成一百万美元?

That's right. Yeah. Who could turn it into a million dollars?

Mustafa Suleyman

智能体实现 10 倍投资回报。

10x return on investment by an agent.

Host

没错。所以我认为这是一个很好的性能和能力衡量标准。当然,我们基本上已经轻松通过了图灵测试,对吧?它已经被通过了。在我们轻松超越图灵之前,没有人真正举行过大型庆祝活动。

Exactly. So I think that's a pretty good measure of performance and capability. And certainly, we've kind of just breezed past the Turing test, right? It's been passed. No one's really done a big celebration before we breezed past Turing.

Mustafa Suleyman

是的。没有人庆祝它。那个卡斯帕罗夫对深蓝的重大时刻在哪里?

Yeah. And no one celebrated it. Where was the big Kasparov vs. Deep Blue moment?

Host

我们现在能虚拟碰杯庆祝我们赢了吗?它发生了。

Can we clink virtual glasses right now and celebrate that we won? It happened.

Mustafa Suleyman

是的。没错。这就是在一个充满复合指数增长的世界里取得进步的感觉,我们对 10 倍已经麻木了。以至于你可以说,‘伙计们,你们为什么还没做到?’

Yeah. Exactly. And that's what it feels like to make progress in a world full of compounding exponentials, where we just get desensitized to 10x. So much so that you can be like, 'Guys, why haven't you done it yet?'

Host

我们被宠坏了。我的现代图灵测试微软勒布纳奖在哪里?

We're spoiled. Where's my Microsoft Loebner Prize for the modern Turing test?

Mustafa Suleyman

对。没错。早些时候有人对我说,‘但 AI 这东西还在初期,不是吗?’我说,老兄,如果这是初期,哇。我可以流利地和我的电脑对话。《星际迷航》实时出现了。

Right. Exactly. Someone said to me earlier, 'But this AI thing, it's still in its infancy, isn't it?' And I'm like, man, if this is infancy, wow. I can talk to my computer fluently. Star Trek is here in real time.

Host

是的,没错。

Yeah, exactly.

Mustafa Suleyman

所以,显然与此同时,智能体还不能真正工作。行动方面仍在进展中。它每分钟都在变得更好,但很明显,在未来几年内,这些东西会进入视野,而且会非常非常好。

So, obviously at the same time, agents don't really work yet. The action stuff is still progressing. It's getting better and better every minute, but it's pretty clear that in the next couple of years, those things come into view and they're going to be very, very good.

Host

在现代图灵测试通过后,我们能再聚一次,只是为了庆祝和认可它吗?

Can we get together again after the modern Turing test has been passed and just to celebrate, recognize it?

Mustafa Suleyman

再次虚拟碰杯。当然。

Virtual glasses again. Absolutely.

Host

希望我们能开一瓶香槟什么的。

Hopefully we can pop a champagne or something.

Mustafa Suleyman

我想我们应该让一个乐观主义者为我们开瓶什么的。

I think we should have an optimist pop the cork for us or something.

Host

没错。没错。

Exactly. Exactly.

DeepMind 收购与数据中心冷却 DeepMind acquisition and data center cooling

Host

Dave,我想再详细聊聊那段背景。我记得很清楚,DeepMind 被 Google 收购后,交易价格是多少?大概是五亿美元吧。

Dave, I want to flush out that backstory a little more. I remember clearly after DeepMind got acquired by Google, what was the price tag? It was like half a billion dollars.

Mustafa Suleyman

6.5 亿美元。2014 年。

650 million. In 2014.

Host

2014 年。我记得一两年后读到,Google 通过让 DeepMind 调整数据中心的空调来证明这笔交易合理。我当时想,‘哇,这进展不太顺利。’而现在它显然是人类历史上最重要的事情。

2014. I remember reading a year or two later that Google justified the deal by having DeepMind tune the air conditioning in the data centers. My interpretation was, 'Wow, this isn't going all that well.' And now it's obviously the biggest thing in human history.

Mustafa Suleyman

是的,数据中心的事情很酷。我们实际上大幅降低了 Google 数据中心机群的冷却成本。

Yeah, the data center thing was pretty cool. We actually reduced the cost of cooling the Google data center fleet by a significant amount.

Host

真有趣,我当时读到觉得‘真是个失败’。后来在来这里的飞机上读到维基百科,实际上有 500 个属性融入神经网络,比新闻说的复杂得多。

It's so funny, I read it at the time and thought, 'What a bust.' Then I read about it on Wikipedia on the flight here, and it was actually 500 attributes fitting into the neural net, much more complicated than the news made it sound.

Mustafa Suleyman

没错。

That's right.

Host

但你之前谈到指数增长的平坦部分。所有那些接近 AGI 的研发都在调整空调。但这就是指数的本质——它们悄悄逼近你。它接受任意数据输入、任意模态,用同样的通用方法在新环境中产生非常准确的预测。这已经发生在文本、音频、图像、编码和其他时间序列数据上。这是模型通用性的又一个证据。人们很容易认为五年是很长的时间。

But you were talking about the flat part of the exponential. All that R&D so close to becoming AGI is tuning the air conditioning. But that's the nature of exponentials—they sneak up on you. It's taking an arbitrary data input, an arbitrary modality, and using the same general-purpose method to produce very accurate predictions in a novel environment. That's happened with text, audio, image, coding, and other time series data. It's another proof point of the general-purpose nature of the models. It's easy to get caught up thinking five years is a long time.

Mustafa Suleyman

就像眨眼之间,沧海一粟。因为我们身处如此狂热的秒级新闻文化和社会媒体环境,我们对这些时间尺度没有直觉。其他文化有,历史上在数字化之前,我们对景观和季节的变化有自然的直觉。现在我们觉得‘它来得不够快’。老兄,它正在到来。我们已经转向运营。我认识很多人,包括这个群体,每天昼夜不停地工作。当我们每周做 Moonshot 播客来庆祝刚刚发生的事情时,每周的变化都令人疯狂。

It's like a blink of an eye, a drop in the ocean. Because we're in such a frantic second-to-second news culture and social media environment, we don't have an intuition for these time scales. Other cultures do, and historically before digitalization, we had a natural intuition for the movement of the landscape and seasons. Now we're like, 'It's not coming quick enough.' Dude, it's coming. We've shifted to operations. I know many people, including this group, operating around the clock every day. When we do a Moonshot podcast week to week to celebrate what's just happened, it's insane on a week-by-week basis.

Host

Peter 总是说人们非常不擅长指数思维。10 万年的进化让我们预测明天会和昨天一样。但你是少数经历过这一切的人,空调在短短几年内变成了 AGI。所以我们现在处于另一个拐点,影响巨大,而人们普遍反应不足。你是少数能说‘我只是非常幸运’的人。我们很幸运对指数有直觉。这很强大,因为理论上我们都能观察到指数的形状,但经历平坦部分并因微小的翻倍而兴奋——这才是关键。当你觉得‘天哪,我记得这个’。

Peter always says people are very bad at exponentials. 100,000 years of evolution has us predicting tomorrow will be like yesterday. But you're one of the few people who, having lived through that, air conditioning becomes AGI in just a few years. So where we sit right now is another inflection point with massive implications, and people are way underreacting. You're one of the few who can say, 'I just got very lucky.' We were lucky to have an intuition for the exponential. It's powerful because we can all theoretically observe the shape of the exponential, but to go through the flat part and get excited by a micro-doubling—that's the bit. When you're like, 'Oh my god, I remember this.'

Mustafa Suleyman

我记得最早的图像生成模型,第一批生成模型。它们可能是 256x256 像素,黑白,手写数字。那大概是 2013 年,也许是 2012 年。Dan Vista,可能是 DeepMind 的第五号员工,一个来自 EPFL 的很棒的荷兰人,生成了第一个可证明不在训练集中的数字七。我想,‘天哪,太神奇了。它怎么能学到关于七的概念?它有了七的概念。’

I remember the earliest image generation models, the first generative models. They were maybe 256 by 256 pixels, black and white, handwritten digits. This was like 2013, maybe 2012. Dan Vista, maybe employee number five at DeepMind, an awesome Dutch guy from EPFL, generated the first number seven that was provably not in the training set. I thought, 'Man, that is amazing. How could it have learned something about the idea of seven? It's got a concept of seven.'

Host

1991 年 MNIST 刚出来时,我得了有史以来的最高分,那时你才三岁。

I got the highest score on MNIST ever in 1991 when it first came out, when you were three years old.

Mustafa Suleyman

九岁。我九岁。实际上,那就是现在 PyTorch 中人们用来做基准测试的同一个数据集。太疯狂了。

Nine. I was nine years old. Actually, that's the same dataset now in PyTorch that people benchmark on. Pretty crazy.

Host

你多久会被看到的东西惊讶一次?多久会有一次‘第 37 手’那样的顿悟时刻?它发生得更频繁了吗?

How often are you surprised by what you're seeing? How often is there a 'move 37' sort of aha moment? Is it happening more frequently?

Mustafa Suleyman

Google 的 LaMDA 最初版本让我彻底震撼。大概有 12 个人在做,由 Noam Shazeer、Daniel De Freitas 和 Quoc Le 领导。我后来加入,大概在他们开始三四个月后。它令人叹为观止。当时每个人都在玩 LLM,一次生成答案,有提示等等。但他们是第一个将其推向对话和聊天的。看到你自己身上涌现的行为——你甚至没想到要问的事情,因为这是对话而不是问答。事后说起来很 trivial,但当时令人叹为观止。我努力推动在 Google 发布它,但由于各种原因我们没能推出。那时我们都离开了——我离开了,Noam 离开去做了 Character,David Luan 离开去做了 Adept。我们都觉得,‘好了,就是这一刻。’之后还有几次,但那可能是近期记忆中最大的一次。

I was absolutely blown away by the first versions of LaMDA at Google. It was maybe 12 people working on it, led by Noam Shazeer and Daniel De Freitas and Quoc Le. I got involved later, maybe three or four months after they'd been going. It was breathtaking. Everyone at that point had been playing with LLMs, one-shot that produce an answer, prompt, etc. But they were really the first to push it for conversation and dialogue. Seeing the emergent behaviors that arise in yourself—things you didn't even think to ask because it's a dialogue rather than a question-answer situation. Sounds trivial to say in hindsight, but it was breathtaking. I pushed hard to ship that at Google, but for various reasons we couldn't get it launched. That's when we all left—I left, Noam left to do Character, David Luan left to do Adept. We were all like, 'Okay, this is the moment.' There have been a couple moments since then, but that was probably the biggest one in recent memory.

Host

缩放定律带来了如此意想不到的性能。回到你早期,你预料到了这些能力吗?这对你来说是可预测的,还是仍然觉得‘哇,它在医学、对话、科学研究中能做的事情’?尤其是仅靠纯文本。我们走了这么远。

The scaling laws have delivered such unexpected performance. Going back to your earlier days, did you anticipate the kinds of capabilities that have resulted? Was this predictable for you, or is it still like, 'Wow, what it's able to do in medicine, in conversation, in scientific research?' Especially working off of pure text. How far we've gotten.

纯文本模型的惊人进展 Surprising progress from text-only models

Host

没人能预料到仅靠文本能走这么远。

Nobody would have seen how far we would get with just text.

Mustafa Suleyman

是的。2015 年,我在 DeepMind 合作了一篇 NLP 深度学习论文,本质上就是试图预测句子中的一个词。我们抓取了《每日邮报》和 CNN 的文章,问:我们能填空吗?预测一个词或补全最后一个词?这和现在模型的工作方式相反。那篇论文贡献很大,引用很高,但我们觉得它永远无法规模化。数据不够,算力不够。但我们仍然乐观,认为只要有更多数据和算力,这个方法就会奏效。我不想有后见之明,但领域里的每个人都在用同样的锤子和钉子不断敲打:我们能加更多数据吗?能明确预测目标吗?能加更多算力吗?大体上,这就是成功的原因。

Yeah. In 2015, I collaborated on an NLP deep learning paper at DeepMind where we were essentially trying to predict a single word in a sentence. We scraped Daily Mail and CNN articles and asked: can we fill in the blank, predict one word or complete the final word? It was the inverse of how models work now. It was a pretty big contribution, a well-cited paper, but we thought it would never scale. Not enough data, not enough compute. But we were still optimistic that with more data and compute, the method would work. I don't want hindsight bias, but everyone in the field had the same hammer and nail and just kept chipping away: can we add more data, clarify the prediction target, add more compute? And broadly speaking, that's what delivered.

下一个惊喜:AI 用于科学与数学 Next surprises: AI for science and math

Host

我们想深入这个话题。你提到 LaMDA 在对话调优上的成功令人惊讶。我想你因预测未来两年内(到 2027 年)智能体将通过现代图灵测试、实现 10 倍到 10 万倍投资回报而上了新闻。我好奇下一个惊喜:AI 用于科学、数学、物理、化学、医学、材料科学。会发生什么,什么时候?

We'd love to pull on that theme. You mentioned how surprising the success of LaMDA for conversational tuning was. I think you've made news by expecting that in the next 2 years, by 2027, we'll see agents start to pass a modern Turing test, 10x to 100,000x return on investment. I'm curious about the next surprises: AI for science, math, physics, chemistry, medicine, material science. What happens and when?

Mustafa Suleyman

你让我想起最近一件令人震惊的事:这些方法可以从一个领域——编程、谜题、数学——学习逻辑推理的本质。就像它学会了数字 7 的概念表征一样,它显然也学会了逻辑推理路径的抽象本质,并能将其应用于许多其他领域。这很有趣,因为它可以结合底层的幻觉/创造力本能(更像插值)一起应用。这两者结合是解决新数学定理或科学挑战的致命组合,因为人类基本上就是这样做的:结合这两种能力。我无法给出具体日期,但感觉肯定触手可及。押注它失败会很奇怪。

You've just reminded me of a more recent mind-blowing thing: these methods can learn from one domain—coding, puzzles, math—the essence of logical reasoning. Just as it learned the conceptual representation of the number seven, it's clearly learned the abstract nature of a logical reasoning path and can apply it to many other domains. That's interesting because it can apply that along with the underlying hallucination/creativity instinct, which is more like interpolation. Those two combined are a lethal combination for making progress in new mathematical theorem solving or scientific challenges, because that's basically what humans do: combine these two capabilities. I can't put a date on it, but it feels definitely within reach. It would be very odd to bet against it.

更难:科学 vs 图灵测试 Harder: science vs. Turing test

Host

从概率角度看,你认为解决科学和工程问题最终会比通过现代图灵测试和实现 10 倍投资回报更难还是更容易?

From an over/under perspective, do you think solving science and engineering will ultimately be harder or easier than passing a modern Turing test and 10xing ROI?

Mustafa Suleyman

会更难。工作场所或创业活动的大量训练数据存在于日志数据中,并且适合与人类进行实时校准。AI 可以检查,人类可以监督、干预、引导和校准。这是 AI 和强化学习的联合努力,人类参与引导强化学习轨迹。但在一个创造全新知识的新领域,它发生在抽象的向量空间中,目前还不清楚人类如何干预定理求解。每个人都在研究这个,尤其是在生物学和合成材料领域,因为你希望给人类更好的直觉,知道在搜索空间的哪里寻找新假设。人类可以接受或拒绝,反馈给模型,在硅基中测试,运行实验,再反馈以改进搜索。

It's going to be harder. A lot of training data for workplace activity or entrepreneurship exists in log data and lends itself to real-time calibration with a human. The AI can check in, the human can oversee, intervene, steer, and calibrate. It's a combined effort between AI and reinforcement learning, where a human participates in steering the RL trajectory. But in a novel domain where it's inventing completely new knowledge, it's happening in an abstract vector space, and it's unclear how the human will intervene in theorem solving. Everyone is working on this, especially in biology and synthetic materials, because you want to give humans better intuition for where to search for new hypotheses. The human can accept or reject that, feed it back to the model, test it in silico, run experiments, and feed that back to improve the search.

加速 AI 科学应用 Accelerating AI for science

Host

人类,特别是微软,或者 AI 社区,能做些什么来加速 AI 用于科学,加速用 AI 解决科学、数学、工程问题?这将是人类最有影响力的事情之一,从根本上让一切以光速前进。

What can humanity, Microsoft specifically, or the AI community do to accelerate AI for science and accelerate the solution to science, math, engineering with AI? That would be one of the most impactful things for humanity, fundamentally moving everything at light speed.

Mustafa Suleyman

我认为这已经在有机地发生了。这不仅是世界上最强大的技术,也是人类历史上传播最快的技术。访问成本和推理成本每几年就下降几个数量级。

I think it's already happening very organically. This is not only the most powerful technology in the world, it's also the fastest proliferating in human history. The cost of access and inference is coming down by multiple orders of magnitude every couple of years.

Host

你曾想过它会这么便宜吗?

Would you ever have imagined it would be so cheap?

Mustafa Suleyman

那一点我也完全搞错了。对我来说最大的惊喜不是我们达到了这种能力水平,而是它如此便宜和易得。

That bit I also totally got wrong. The biggest surprise for me isn't that we're getting this level of capability, it's how cheap and accessible it is.

Host

100%。那是两年内一千倍。它会再次发生吗,还是只有一次?

100%. That's a thousandx over two years. Is it going to do that again or was that a one-time thing?

Mustafa Suleyman

我认为大约是 100 倍。过去两年,每 token 的推理成本下降了 100 倍。

I think it's like a 100x. The inference cost per token has come down 100x in the last two years.

Host

有相互竞争的估计。有些衡量每 token 每美元的智能,某些权重类模型每年增长 40 倍。我见过某些模型类达到 1000 倍。太疯狂了。

There are competing estimates. Some measure intelligence per token per dollar, with 40x year-over-year for certain weight classes. I've seen 1000x for some classes of models. Craziness.

开源模型与成本动态 Open source models and cost dynamics

Mustafa Suleyman

哦,哇,这太疯狂了。是的,这确实是个好观点。我完全搞错了,因为我没想到世界上最大的公司会开源那些训练成本高达数十亿美元的模型。以至于当我们创立 Inflection 时——大概在 ChatGPT 发布前 9 个月或一年——我们在 ChatGPT 发布前一年就开始融资了。我们基本上筹集了 15 亿美元,带着一个 25 人的团队,与 Nvidia 和 CoreWeave 一起建造了当时最大的 H100 集群。我们是 CoreWeave 的第一个 AI 客户。他们之前做加密货币,我们是他们的第一个 AI 客户,与他们合作建造数据中心。Nvidia 也支持了我们。我们当时建造的集群大约有 15,000 块 H100,后来增长到 22,000 块。然后那一年 ChatGPT 发布了,大约在那段时间 Llama 也发布了。所以我们想,‘天哪,我们公司的整个成本基础就这样被削弱了,因为开源似乎不是关于性能,而是关于成本。’所以像 Perplexity 这样的公司,在 Llama 出现后成立,知道他们可以依赖 Llama 和开放的 API 以及其他 API,因此他们的成本基础低得多。所以这是另一件不可预测的事情。我是说,其他人预测到了,我只是搞错了。

Oh, wow. That's wild. Yeah, that's actually a good point. I got that totally wrong because I didn't think that the biggest companies in the world were going to open source models that cost billions of dollars essentially to train. So much so that when we founded Inflection, maybe 9 months or a year before ChatGPT was released, we started fundraising a year before ChatGPT was released. We basically raised a billion and a half dollars with a 25-person team to build what at the time was the largest H100 cluster with Nvidia and CoreWeave. We were CoreWeave's first AI customer. They were previously in crypto, and we were their first AI customer working with them to build our data centers. Nvidia got behind us. I think we built a cluster at the time of about 15,000 H100s, growing to 22,000. Then that year ChatGPT came out, and a few months around that time Llama came out. So we were like, 'Oh my god, our entire cost base of our company has just been undermined by the fact that open source seems to be not really about performance, it's just cost.' So then Perplexity, for example, founded after the arrival of Llama, knowing they could depend on Llama and obviously open APIs and all the other APIs, so they had a much lower cost base. So that was another thing that was not predictable. I mean, other people predicted it, I just got it wrong.

Host

丰裕、去货币化、宇宙中最强大工具的民主化。如果说有什么的话,那就是超级通缩。

Abundance, demonetization, democratization of the most powerful tools in the universe. Hyperdeflation, if anything.

Mustafa Suleyman

超级通缩,是的。我认为这是非常重要的一点。获取知识、智能或能力的成本作为服务将趋近于零边际成本。显然,这将带来巨大的劳动力通缩和替代效应,但也会产生一种奇怪的通缩效应,因为会发生什么?人们将没有基于美元的收入来购买东西,这显然很糟糕,但消费物品的成本也会下降。所以我们实际上存在一个转型错配,因为劳动力市场会在服务成本下降之前受到影响,两者之间可能有 10 到 20 年的滞后,这将非常不稳定。

Hyperdeflation, yeah. I think that's a really important point. The cost of accessing knowledge or intelligence or capability as a service is going to go to zero marginal cost. Obviously that's going to have massive labor deflation and displacement effects, but it's also going to have a weirdly deflationary effect because what is going to happen? People aren't going to have dollar-based incomes to go buy things, that's obviously bad, but the cost of consuming stuff is also going to come down. So we actually have a transition mismatch because labor markets are going to be affected before the cost of services comes down, and maybe there's a 10-20 year lag between that, which is going to be very destabilizing.

Host

顺便说一句,这正是我们之前开始讨论的。我认为长期来看,人类有一个非凡的未来,食物、水、能源、医疗保健和教育将惠及每一个男人、女人和孩子。而短期才是具有挑战性的,对吧?2 到 7 年的时间框架。你的模型也是这样的吗?

Which by the way is what we started to talk about a little bit earlier. I posit that in the long term there's an extraordinary future for humanity, where access to food, water, energy, healthcare, education is accessible to every man, woman, and child. And it's the shorter term that is challenging, right? The 2 to 7 year time frame. Is that your model too?

Mustafa Suleyman

是的,我认为短期会相当不稳定。中长期来看,很明显这些模型在诊断方面已经达到世界级水平。我们大约四五个月前发布了一篇论文,叫做 MAI 诊断编排器。它本质上在底层使用了大量模型,试图处理《新英格兰医学杂志》上的一组罕见病症——那些不易诊断的罕见病例,最好的专家也做得不太好——它的准确率大约高出四倍,不必要的检测成本大约降低一半。

Yeah, the short term I think is going to be quite unstable. The medium to longer term, it's pretty clear that these models are already world class at diagnostics. We released a paper maybe four or five months ago now called the MAI Diagnostic Orchestrator. Essentially it uses a ton of models under the hood to try and take a set of rare conditions from the New England Journal of Medicine, rare cases that can't be easily diagnosed, that the best experts do a kind of weak job on, and it's like four times more accurate roughly, and about 2x less the cost in terms of unnecessary testing.

Host

哈佛和斯坦福有一项研究,在这种情况下是 GPT-4,单独医生、医生加 GPT-4 以及单独 GPT-4。结果令人难以置信,如果让 AI 单独工作,它在诊断方面比人类准确得多。我们的想法和昨天看到的东西、最近的诊断都有偏见。

There's a study that came out of Harvard and Stanford looking at, in this case GPT-4, a physician by themselves, a physician with GPT-4, and GPT-4 by itself. And it was incredible that if you left the AI alone, it was far more accurate in diagnostics than the human. We're biased in our thoughts and what we saw yesterday, our recent diagnosis.

Mustafa Suleyman

是的,实际上我们在论文发布后收到了很多反馈,因为我们只展示了 AI 单独和医生单独的情况。很多人想看看医生和 AI 一起工作,或者至少医生也能使用谷歌搜索的情况。这确实稍微提高了性能,但 AI 仍然遥遥领先。

Yeah, actually we got a lot of feedback after we released the paper because we only showed the AI on its own, the physician on its own. And a lot of people wanted to see what it was like to have the physician and the AI, or at least the physician have access to Google search as well. And that improves performance a little bit, but the AI still trumps by quite a way.

微软 AI 战略与使命 Microsoft's AI strategy and mandate

Host

Dave,你在想什么?

Dave, what are you thinking?

Dave

哦,太多了。那么,微软,你在这里多少年了?

Oh, so much. So, Microsoft, you've been here how many years now?

Mustafa Suleyman

才一年半。

Just a year and a half.

Host

一年半。所以你觉得自己已经被同化了。那么 Satya 给你的任务是什么?是赢得 AGI,还是自给自足,或者目标是什么?

Year and a half. So you feel like you're indoctrinated. So what's the mandate from Satya? Is it win AGI, or be self-sufficient, or what is the target?

Mustafa Suleyman

我不认为真的有赢得 AGI 这回事。我认为这是很多人强加给这个领域的一个错误框架。我不确定是否存在一场竞赛,对吧?我的意思是,我们都在尽可能快地前进,但竞赛意味着零和博弈。它意味着有一条终点线。它意味着有第 1、2、3 名的奖牌,但没有第 5、6、7 名。这不太是一个恰当的比喻。正如我们所知,技术、科学和知识无处不在,同时以各种规模扩散,基本上同时或在一年两年内。所以我的任务是确保我们自给自足,知道如何从头到尾训练我们自己的模型,在所有能力的所有规模的前沿,并在公司内部建立一个绝对世界级的超级智能团队。我还负责 Copilot。所以这是我们将这些模型投入生产、覆盖所有消费者端的工具。

I don't think there's really a winning of AGI. I think this is a misframing that a lot of people have kind of imposed on the field. I'm not sure there's a race, right? I mean we're all going as fast as we possibly can, but a race implies that it's zero sum. It implies that there's a finish line. And it implies that there are medals for 1, 2, and 3, but not 5, 6, and 7. And it's just not quite the right metaphor. As we know, technologies and science and knowledge proliferate everywhere, all at once at all scales, basically simultaneously or within a year or two. So my mission is to ensure that we are self-sufficient, that we know how to train our own models end to end from scratch at the frontier of all scales on all capabilities, and we build an absolutely world-class super intelligence team inside of the company. I'm also responsible for Copilot. So this is our tool for taking these models to production in all of our consumer surfaces.

Host

所以澄清一下,当我们看 Polymarket 时——我们在播客中经常这样做——关于年底谁拥有最好的 AI 模型以及明年年底谁拥有最好的 AI 模型的竞赛。那个图表上没有微软的线,对吧?

So just to clarify, when we look at Polymarket, which we do a lot on the podcast, the horse race to who has the best AI model at the end of the year and who has the best AI model at the end of next year. There's no Microsoft line on that chart, right?

Mustafa Suleyman

所以现在会有,我猜。

So now there will be, I assume.

Host

是的,明年会有。我们将发布越来越多来自我们的模型,但这需要很多年才能建成。我的意思是,DeepMind 或 OpenAI 是拥有十年历史的实验室,它们已经养成了进行真正前沿研究、仔细剔除失败并重新引导人员的习惯和实践。这是一种需要多年才能建立起来的文化和纪律。但是的,我们绝对在向前沿推进。我们希望建造世界上最好的超级智能和最安全的超级智能模型。

Yeah, there will be next year. We'll be putting out more and more models from us, but this is going to take many years for us to build this. I mean, DeepMind or OpenAI, these are decade-old labs that have built the habit and practice of doing really cutting edge research and being able to weed out carefully the failures and redirect people. This is an entire culture and discipline that takes many years to build. But yeah, we're absolutely pushing for the frontier. We want to build the best super intelligence and the safest super intelligence models in the world.

Host

很好。

Nice.

团队建设与人才争夺战 Building the team and hiring wars

Host

那么当你到达 Inflection 时,当时的理念是 18000 块 H100,构建一个大型 Transformer。你可能一年半前第一天就看到了 OpenAI 的源代码。数十亿美元的研发是如何实现的?这里已经有团队了,还是你带了自己的团队?

So when you arrived at Inflection, the thesis was 18,000 H100s, building a big transformer. You probably saw OpenAI's source code a year and a half ago on day one. How does multi-billion dollar R&D arrive? Was there a team here already or did you bring your own?

Mustafa Suleyman

是的,我的整个团队都过来了,而且我们一直在大力扩充团队。我们从各大实验室招人,深陷招聘大战,这很不真实。史无前例。每天都有 CEO 们互相打电话。这是一场持续的战争。我们正在从零开始组建团队。

Yeah, all my team came over and we've been growing that team a lot. We've hired from all major labs and we're in the trenches of hiring wars, which are surreal. It's unprecedented. Phone calls every day from CEOs to other people. It's a constant battle. We're building the team from scratch.

Host

现在你手下有 1 万名员工?

10,000 employees under you now?

Mustafa Suleyman

不,不。核心超级智能团队只有几百人。这是第一要务。其余的是 Copilot,搜索引擎。

No, no. The core superintelligence team is a few hundred. That's the number one priority. The rest is Copilot, the search engine.

定义 AGI 与超级智能 Defining AGI and superintelligence

Host

像 AGI 和 ASI 这样的术语被随意使用。你们内部对 AGI 和数字超级智能有定义吗?

Terms like AGI and ASI get thrown around. Do you have an internal definition of AGI versus digital superintelligence?

Mustafa Suleyman

是的,粗略地说,这些只是曲线上的点。

Yeah, loosely, these are just points on a curve.

Host

在你看来,它们是可以互换的,还是不同的?

Are they interchangeable in your mind, or different?

Mustafa Suleyman

我认为它们通常被用作不同的概念。不同的人有不同的定义。AGI 的定义就像图灵测试——它会过去,变得模糊,我们事后才会意识到。粗略地说,在远端,超级智能是一种 AI,能够比所有人类加起来更好地完成所有任务,并且有能力随着时间的推移不断自我改进。

I think they're generally used as different. Different people have different definitions. The AGI definition is like the Turing test—it will pass by and be blurred, and we'll recognize it in retrospect. Roughly speaking, at the far end, a superintelligence is an AI that can perform all tasks better than all humans combined and has the capacity to keep improving itself over time.

Host

那我必须问:什么时候?一个极小极大?

So I have to ask: when? A minmax?

Mustafa Suleyman

很难说。我不知道。但它已经足够接近,我们应该尽一切努力优先考虑安全、对齐和遏制。

It's very hard to say. I don't know. But it is close enough that we should do everything in our power to prioritize safety, alignment, and containment.

有感知 vs 有意识 AI Sentient vs conscious AI

Host

你曾主导过一场关于有意识 AI 的感知是幻觉的讨论。我想区分有感知的 AI 和有意识的 AI。你区分这两者吗——感觉和情感,与有意识和反思自身思想?

You've led a conversation about the perception of conscious AI being an illusion. I want to distinguish between sentient AI and conscious AI. Do you distinguish between the two—sensations and feelings versus being conscious and reflective?

Mustafa Suleyman

是的,这又回到了定义问题。AI 将能够拥有体验,但我不认为它会像我们一样有感受。感受和感知是生物物种特有的。你可以编码一个与情绪状态相关的优化函数,但这与我们编写模型模拟知识生成没什么不同。模型没有体验或意识到看到红色是什么感觉;它只能通过预测性地生成 token 来描述红色。而你有感受质,有本质,有基于你与嗅觉、声音、触觉的生物交互体验对红色的直觉。你可以设计一个模型来模仿意识或感知的标志,但这有问题,因为它实际上不会受苦。但我们的共情回路会被激活,人们已经在倡导模型权利和福利。

Yeah, again, it's about definitions. An AI will be able to have experiences, but I don't think it will have feelings the way we do. Feelings and sentience are specific to biological species. You could code in an optimization function related to emotional states, but it would be no different from how we write models to simulate knowledge generation. The model has no experience or awareness of what it's like to see red; it can only describe red by generating tokens predictively. Whereas you have qualia, an essence, an instinct for red based on your biological interactive experience with smell, sound, touch. You could engineer a model to imitate the hallmarks of consciousness or sentience, and that's problematic because it won't actually suffer. But our empathy circuits will activate on that, and people are already advocating for model rights and welfare.

Host

伊利亚最近谈到情绪是人类决策的关键。具有模拟情绪的 AI 能否成为更好的 ASI?

Ilya recently spoke about emotions being key in human decision-making. Could AIs with simulated emotions be better ASIs?

Mustafa Suleyman

我担心这太拟人化了。我们已经在提示、系统提示、宪法中加入了情绪。这些不是理性存在;它们被摆布,有任意偏好,因为它们是在风格化地解释我们植入的行为。我们可以设计特定的共情回路或镜像神经元回路,比如动机意志。但目前这些是下一个词可能性预测机器,优化单一目标——下一个应该出现哪个词。没有更高阶的预测功能。人类有多个相互冲突的驱动力和动机,它们相互作用,加上社会互动。这些模型没有这些。你可以设计一个意志或偏好,但它不会是涌现的。

I worry that's too anthropomorphic. We already have emotions in the prompt, system prompt, constitution. These are not rational beings; they get moved around and have arbitrary preferences because they're stylistically interpreting behaviors we plugged in. We could engineer specific empathy circuits or mirror neuron circuits, like motivational will. But currently these are next-token likelihood predictor machines optimizing for a single thing—which token should appear next. There's no higher-order predictive function. Humans have multiple conflicting drives and motivations that interact, plus social interaction. These models don't have that. You could engineer a will or preference, but it wouldn't be emergent.

工程 AI:亲人类 vs 亲 AI 辩论 Engineering AI and the Pro-Human vs Pro-AI Debate

Mustafa Suleyman

这将是我们设计进去的东西,而且我们应该非常谨慎地去做。

That would be something that we engineer in and we should do that very carefully.

Host

我很喜欢你为这个问题带来的人文主义视角。没错,我的意思是,除了是一名技术专家,你的背景从一开始就是支持人类的。我认为我们即将进入一场有趣的文化辩论,那些支持 AI 与支持人类的人之间的辩论,就是埃隆和拉里·佩奇之间那场著名的对话:‘你支持 AI 胜过人类,是不是物种歧视?’

I do love that you bring this humanistic side to the equation. Right. I mean, in addition to being a technologist, your background is one that is pro-human at the beginning. And this interesting cultural debate I think we're about to enter into, those that are sort of pro-AI versus pro-human, that famous conversation between Elon and Larry Page about 'are you a speciesist because you're in favor of AI over humans?'

Mustafa Suleyman

我的意思是,这将成为一条分界线。有些人——我不太确定埃隆现在站在辩论的哪一边。我最近确实听到他说过一些相当后人类、超人类主义的话。

I mean, look, that's going to be a dividing line. There are some people, and I'm not quite sure which side of the debate Elon's on these days. Like I've certainly heard him say some pretty posthuman transhumanist things lately.

Host

我认为我们在未来 5 到 10 年内必须做出一些艰难的决定。我回避关于超级智能时间线问题的原因是,我认为无论是一年、十年还是二十年,都极其紧迫,我们现在就必须声明我们要构建什么样的超级智能,以及我们是否真的打算容忍创造某种我们无法对齐、无法遏制、并且设计上在所有任务和人类理解上超越人类表现的实体。理解这一点:你如何控制你不理解的东西?对吧?

And I think that we're going to have to make some tough decisions in the next 5 to 10 years. I mean the reason I dodged the question on the timeline for superintelligence is because I think that it doesn't matter whether it's one year or 10 or 20 years; it's super urgent that right now we have to declare what kind of superintelligence are we going to build and are we actually going to countenance creating some entity which we provably can't align, we provably can't contain, and which by design exceeds human performance at all tasks and human understanding. And understanding like how do you control something that you don't understand? Right?

Host

如果可以的话,我想稍微深入探讨一下拟人化这个话题。如果你记得道格拉斯·亚当斯的书《宇宙尽头的餐馆》,有一个场景是一头牛被设计成邀请餐厅顾客吃它,因为这让他们感觉更舒服。牛并不介意。这头牛被优化成想要被顾客吃掉。但许多读者对那个场景感到震惊。先把这个放一边。微软有将 AI 助手、Copilot 拟人化的历史,可能可以追溯到 Microsoft Bob 之前的例子,还有 Rover 狗,然后是 Microsoft Office 中的 Clippy,以及最近一些模糊的云形头像。你如何看待一方面不希望过度拟人化智能体,另一方面又与一个可以说是拟人化智能体先锋的机构相协调?

I'd like to, if I may, pull on the anthropomorphization thread a bit. If you remember Douglas Adams' book, The Restaurant at the End of the Universe, there's a scene where there's a cow that's been engineered to invite restaurant patrons to eat it because it makes them feel more comfortable. And the cow doesn't mind. The cow's been optimized to want to be eaten by the patrons. But many readers were horrified at that scene. Put that in a box for a moment. Microsoft has a history of anthropomorphizing AI assistants, copilots, going back probably to an example prior to Microsoft Bob, and the Rover dog, and then Clippy in Microsoft Office, and more recently a sort of amorphous cloud-shaped avatars. How do you think about reconciling on the one hand the desire not to overly anthropomorphize agents, on the other hand with an institution that has arguably been in the vanguard of anthropomorphizing agents?

Mustafa Suleyman

我认为整个设计领域一直以人类状况为参考点,对吧?我的意思是,拟物设计是图形用户界面的支柱,对吧?从文件夹到日历,以及介于两者之间的一切,对吧?我们仍然在那些我们认为现代的老式界面中看到它的痕迹。所以这就像是我们文化中不可避免的一部分,我们最终会超越它们。我们找到更简洁、更好、更有效的用户界面。我并非默认反对拟人化。我的意思是,我认为我们希望东西符合人体工程学,对吧?椅子合适。语言模型说我的语气,对吧?它有一种对我有意义的流畅性。它具有与我历史和民族共鸣的文化意识,等等。我认为这是当今设计固有的一部分。作为事物的创造者,我们现在在设计个性、文化和价值观,而不仅仅是像素和软件。但显然有一条线,对吧?创造与人类无法区分的东西有很多其他风险和复杂性,比如这会使沉浸到模拟中更加危险和更可能发生,对吧?所以我认为我对那些明显不同、分离、不试图模仿、并且总是披露它们本质上是 AI、并且有边界的实体、头像或声音没有意见。这似乎是安全自然且必要的一部分。

I think the entire field of design has always used the human condition as its reference point, right? I mean, skeuomorphic design was the backbone of the GUI, right? From file folders to calendars and everything in between, right? And we still have the remnants of that in our old-school interfaces which we feel are modern and stuff. So that's like an inevitable part of our culture and we just grow out of them. We figure out cleaner, better, more effective user interfaces. I'm not against anthropomorphism by default. I mean, I think we want things to feel ergonomic, right? The chair fits. The language model speaks my tone, right? It has a fluency that makes sense to me. It has a cultural awareness that resonates with my history and my nation and so on. And I think that is an inherent part of design today. As creators of things, we are now engineering personalities and culture and values, not just pixels and software. But obviously there's a line, right? Creating something which is indistinguishable from a human has a lot of other risks and complications, like that makes the immersion into the simulation even more dangerous and more likely, right? And so I think I don't have a problem with entities, avatars, or voices that are clearly distinct and separate and not trying to imitate, and always disclose that they are an AI essentially, and that there are boundaries around them. That seems like a natural and necessary part of safety.

Host

所以我认为我听到你说的是,一方面拟人化是新的拟物化,但另一方面要保持人类智能和人工智能之间清晰甚至可能是法律上的界限。你认为,你看到 AI 获得某种法律人格的未来吗?还是这是被禁止的?永远不会发生?你看到人类被允许与 AI 融合的未来吗,就像库兹韦尔风格,播客的朋友?还是这在你看来也不在考虑范围内?

So what I think I hear you saying, correct me if I'm mistaken, is anthropomorphization is the new skeuomorphism on the one hand, but on the other hand maintaining clean, maybe even legal boundaries between human intelligence and artificial intelligence. Do you think, do you see a future where AIs achieve some sort of legal personhood, or is that forbidden? Is that never going to happen? Do you see a future where humans are allowed to merge with the AIs, Kurzweil-style, friend of the pod, or is that also not on the table in your mind?

Mustafa Suleyman

是的,我认为 AI 法律人格绝对不在考虑范围内。我认为如果我们给予法律人格和权利给一个成本只有我们一小部分、可以无限复制和繁殖、拥有完美记忆、可以瘫痪自身计算的物种,我们的物种无法生存。我的意思是,这些与作为生物物种的人类摩擦如此对立,以至于会存在固有的资源竞争。直到可以证明,直到可以证明这些东西会与我们的价值观和我们作为物种的持续存在对齐,并且可以在数学上被证明是可遏制的——这是一个极高的标准。我不认为我们应该考虑给予……一条清晰的界线。我真的认为这是一条清晰的界线。我认为这非常危险。还有一个单独的问题与责任有关,因为它们将拥有越来越大的自主权。明确地说,我也是加速主义者。我想制造这些东西。它们会……但紧张是理性的。人们总是说紧张是理性的。如果你看不到紧张,你肯定错过了大部分辩论。这显然非常复杂。我们越是谈论复杂性并保持紧张,你就越能看到智慧。我们绝不能把这些东西放在桌面上说不行,比如我们希望这些东西在诊所、学校、工作场所,为我们大规模地提供价值,但它们必须有边界和控制。这就是我们必须运用的艺术。

Yeah, I mean I think AI legal personhood is extremely not on the table. I don't think our species survives if we have legal personhood and rights alongside a species that costs a fraction of us, that can be replicated and reproduced at infinite scale relative to us, that has perfect memory, that can just paralyze its own computation. I mean, these are so antithetical to the friction of being a biological species, us humans, that there would just be an inherent competition for resources. And until it was provable, until it was provable that those things would be aligned to our values and to our ongoing existence as a species and could be contained mathematically provably, which is a super high bar. I don't see that we should be any considering giving... bright line in the sand. I really think it's a bright line. I think it's very dangerous. There's a separate question which has to do with liability because they are going to have increasing autonomy. Like to be clear, I'm also an accelerationist. I want to make these things. They're going to... but tension is rational. People always say that tension is rational. If you don't see the tension, you're definitely missing most of the debate. It's obviously very complex. The more we talk about the complexity and hold it in tension, that's when you start to see the wisdom. And there's no way we can leave these things on the table and say no, like we want to have these things in clinic, in school, in workplace, delivering value for us at a huge scale, but they have to be boundaried and controlled. And that's the art that we have to exercise.

赞助商插播:Blitzy Sponsor Break: Blitzy

Host

本期节目由 Blitzy 赞助,Blitzy 提供具有无限代码上下文的自主软件开发。Blitzy 使用数千个专门的 AI 智能体,它们会思考数小时以理解拥有数百万行代码的企业级代码库。工程师在每个开发冲刺开始时使用 Blitzy 平台,输入他们的开发需求。Blitzy 平台提供计划,然后为每个任务生成并预编译代码。Blitzy 自主完成 80%或更多的开发工作,同时为完成冲刺所需的人类开发工作的最后 20%提供指导。

This episode is brought to you by Blitzy, autonomous software development with infinite code context. Blitzy uses thousands of specialized AI agents that think for hours to understand enterprise scale code bases with millions of lines of code. Engineers start every development sprint with the Blitzy platform, bringing in their development requirements. The Blitzy platform provides a plan, then generates and pre-compiles code for each task. Blitzy delivers 80% or more of the development work autonomously while providing a guide for the final 20% of human development work required to complete the sprint.

AI 人格与人类竞争 AI Personhood and Human Competition

Host

听起来,如果我可以这么说的话,我听到的反对 AI 人格的主要理由在于人类当前形态的不足。你说过,它们会超越人类,它们更聪明、更快、比人类智能更容易克隆。如果人类智能得到提升,也许借助 AI 的帮助,如果我们有上传技术或先进的脑机接口,能够提升平均人类智能,那么在你看来,如果人类能与 AI 在公平竞争环境中竞争,这是否为 AI 人格打开了一扇门?

It sounds though, if I may, the primary rationale that I'm hearing for why not AI personhood has to do with the inadequacies of the human form as currently constructed. I heard you say, well, they'll outrace humans. They're so much smarter. They're so much faster. They're so much more clonable than human intelligence is. If human intelligence were uplifted, maybe with the benefit of AI, if we had uploading type technologies or BCIs that are advanced that enable us to lift up the average human intelligence, in your mind then does that open the door a bit to AI personhood if humans can compete on a level playing ground with AIs?

Mustafa Suleyman

我不想让地球上 70 亿人的和平与繁荣的竞争变得更加混乱。所以,如果未来一个世纪的道路能被证明更安全、更和平、更少疾病和痛苦,并且有空间容纳这个其他物种,那么我对它持开放态度,包括生物混合体等等。我原则上并不反对。我只是一个物种主义者。

I don't want to make the competition for the peace and prosperity of the 7 billion people on the planet even more chaotic. So if the path over the next century can be proven to be much safer and more peaceful and less like disease and sickness and there is room for this other species, then I'm open-minded to it, including biological hybrids and so on. I'm not against that on principle. I'm just a speciesist.

Host

啊哈。

Aha.

Mustafa Suleyman

我只是一个人本主义者。我从我们在这里出发,保护所有我知道存在的、可能因引入这个新事物而遭受巨大痛苦的现有有意识存在的福祉,是一项道德义务。

I'm just a humanist. I start with we're here and it's a moral imperative that we protect the well-being of all the existing conscious beings that I know do exist and could suffer tremendously by the introduction of this new thing.

Host

当然,尼安德特人可能有过那样的对话,或者过去十亿多年里在我们之前的每一个物种。我的意思是,有很多人认为我们只是一个临时的过渡物种,是超级智能的引导程序。

Now of course the Neanderthals may have had that conversation or every species that preceded us over the last billion plus years. I mean, there are many who argue we're simply an interim transitory species in bootloader for the superintelligence.

Mustafa Suleyman

那句经典的话。是的,我完全知道。而且我也是从宇宙时间尺度思考的人。所以,我不是天真地说这个世纪。我肯定意识到正在发生巨大的转变。事实上,你甚至可以在最近的记忆中看到。我的意思是,250 年前,预期寿命大约是 30 岁左右。当然,在某些方面,我们是一种增强的混合生物物种,对吧?我们服用所有这些药物,每个人的肽都很棒,这太棒了,我完全支持。来吧。

That classic phrase. Yes, I'm totally aware of that. And I'm also someone who thinks on cosmological time, too. So, I'm not just naively saying, you know, this century. I'm definitely aware that there's a huge transition going on. And in fact, you can even see it in recent memory. I mean, 250 years ago, life expectancy was about 30 years or whatever it was. Of course, in some ways, we are a augmented hybrid biological species, right? We take all these drugs and everyone's peptides are amazing and it's super, I'm down for all of that. Let's go.

Host

基因重编程明年就要来了。

The genetic reprogramming is coming next year.

Mustafa Suleyman

没错。来吧。我支持,我支持。但我们不要搬起石头砸自己的脚。我想确保我们星球上的大多数人,如果不是所有人的话,都能从技术带来的和平与繁荣中受益。

Exactly. Let's go. I'm down. I'm down. But let's not shoot ourselves in the foot. I want to make sure that most of our planet, if not everybody, gets the benefit of the peace and prosperity that comes from the technology.

Host

我的意思是,如果你相信 AI 最终会超越我们,把我们打入无足轻重的境地,那么这个论点有一定道理。我是说从长远来看。

I mean there is some level of sanity in that argument if you believe that the AI will ultimately outcompete us and put us into a box of insignificance. I mean in the long run.

Mustafa Suleyman

我的意思是,所有智能,我们在自然界中都能看到这一点。我们天生是等级制的。到目前为止,我们还没有看到这种为了保存其他物种而自我牺牲的超协作物种。

I mean all intelligences, we can see this in nature. We're innately hierarchical. So far, we have not seen this supra collaborative species that will take self-sacrifice in order to preserve the other species.

Host

所以存在固有的等级制,智能的等级结构带来了固有的冲突,对吧?

So there's an inherent hierarchical, there's an inherent clash coming from the hierarchical structure of intelligence, right?

Mustafa Suleyman

所以,我并不是说我们不应该探索它,也不是说它不可能发生,但门槛首先必须是“不伤害”,也许可以做一点,但首先要不伤害我们的物种。不要像你说的那样搬起石头砸自己的脚,戴夫。

So, and all I'm saying is not that we shouldn't explore it, not that it couldn't potentially happen, but the bar has to first be do no, maybe do a little, but do no harm to our species first. Don't shoot ourselves in the foot as you said, Dave.

Host

顺便说一句,在这个话题上我 100%同意你。不能再一致了。但杰弗里·辛顿在告诉世界它会失控,我们的安全阀是给它母性本能,我觉得这是一个有趣的观点。

Well, I'm 100% with you on this topic, by the way. Could not be more aligned. But Geoffrey Hinton is out there telling the world it's going to run away and our safety valve is giving it a maternal instinct, which I found an interesting point of view.

Mustafa Suleyman

嗯,我没有检查那个安全阀。

Well, I didn't check that safety valve.

Host

嗯,他认为它是不可遏制的,我同意你。我认为如果你不给它情感和意图编程,它是非常可遏制的。但他认为它是不可遏制的。他获得诺贝尔奖时非常悲观。现在他更乐观了,因为他看到了一条通过编程母性本能来实现的路径,这意味着它对我们占主导地位但关心我们。他的论点是,我看到过一种情况,一个远更智能的实体照顾一个年轻无能的实体,就像母亲和哭闹的孩子。

Well, he believes it's uncontainable and I'm with you. I think it's very containable if you don't give it emotional and intentional programming. But he thinks it's uncontainable. He was very pessimistic when he got his Nobel Prize. Now he's more optimistic because he sees a path to programming in maternal instinct which implies that it's dominant to us but it cares. His thesis was I've seen a situation where a vastly more intelligent entity takes care of a younger inept entity in a mother with their screaming child.

Mustafa Suleyman

是的。没错。

Yeah. Exactly.

Host

所以,如果我们能编程一个母性本能到 AI 中,即使我们能力远不如它,它也会照顾我们。

So if there's a maternal instinct that we can program into AI even though we're far less capable, it will take care of us.

Mustafa Suleyman

它被比作所谓的 AI 对齐的数字催产素计划。

It's been compared to the call it the digital oxytocin plan for AI alignment.

Host

我喜欢这个说法。

I like that.

Mustafa Suleyman

这个说法不错。是的。

That's a good one. Yeah.

Host

我的意思是,这已经很有诗意了。我想我需要一些更有公式化的东西,更令人放心。但你看,安全有 101 种可能的策略。我们应该探索所有策略,认真对待它们。我的意思是,杰夫是这个领域的传奇。毫无疑问。但我只是认为要谨慎对待。

I mean, it's about as poetic as it gets. I think I'm going to need something that's got a little bit more formula to it. A bit more reassuring. But look, there's 101 different possible strategies for safety. We should explore all of them. Take them all seriously. I mean, Jeff is a legend and of the field. No question. But like I just think approach with caution.

Mustafa Suleyman

你在安全上投入了大量精力、算力和人力吗?是的,我会说没有我们应有的那么多。我正在努力理解它。外面有人吗?我很好奇所有那些超大规模公司中,有没有哪个实体在你看来投入了足够的资源?因为每个人都在竞赛。就像更多的 GPU、更多的数据、更多的能源。每个人都在优化下一个基准。我没有看到任何安全基准。有安全基准吗?

Are you spending a lot of your energy compute, human power on safety? Yeah, I would say not as much as we should. I'm wrapping my head around it. Is anybody out there? I am curious out of all the hyperscalers out there. Is there any entity that's spending enough in your mind? Because everybody's in such a race. It's like more GPUs, more data, more energy. It's just like everybody's optimizing for the next benchmark. I don't see any safety benchmarks. Are there any safety benchmarks out there?

Host

哦,有大量的安全基准。至少在我看来,有一个关于防御性协同扩展的论点。我很想听听你的想法。你是否认为,就像城市变大时警察力量也会变大一样。也许不是直接成比例,也许有一些扩展指数,但你认为对齐力量或安全力量的防御性协同扩展,无论最终意味着什么,是 AI 对齐策略的一部分吗?

Oh, there are tons of safety benchmarks. And there's at least in my mind an argument for defensive co-scaling. I'd be curious to hear your ideas on that. Do you think in the same way that as a city gets larger, the police force gets larger. Maybe it's not in direct proportion. Maybe there's some scaling exponent, but do you think defensive co-scaling of alignment forces or safety forces, whatever that ends up meaning, do you think that's part of the strategy for AI alignment?

Mustafa Suleyman

我认为那会是一个好方法。我的意思是,我们多年来多次提出过这个。我的意思是,拜登政府下的白宫自愿承诺,我和实际上每个人,德米斯、达里奥、山姆以及我们所有人通过合作都在大力推动这一点。你看,它被抛弃了,但我认为这是一套非常明智的原则。就像根据算力规模进行审计,你知道,我们共同承担一定比例的安全投资算力和人力。

I think that would be a good way. I mean, we've proposed this several times over the years. I mean, the White House voluntary commitments under Biden that me and in fact everyone, I mean, Demis and Dario and Sam and all of us through co were pushing this pretty hard. And look, I mean, it got chucked out, but I think it's a very sensible set of principles. It's like auditing for scale of flops, you know, having some percentage that we all share of safety investment flops and headcount.

AI 实验室的竞争与合作 Competition and Coordination in AI Labs

Mustafa Suleyman

你知道,现在是时候了,表面上每个人都愿意分享最佳实践、相互披露并在时机成熟时进行协调。但我觉得我们还没到那个阶段。所以,目前我们处于高度竞争的模式。不过,是的,我认为现在确实是进行这些投资的时候了。

You know, this is the time and I think on the face of it, everyone is open and willing to sharing best practices and disclosing to one another and coordinating when the time comes. I think we're still pre that level. So, we're in like hyper competitive mode at the moment. But yeah, I think now is really the time to be making those investments.

Host

是啊。那么,有没有什么会吓到我们所有人,让大家都停下来?你知道,有没有像三里岛那样的事件?

Yeah. Well, is there something that's going to scare the out of us that stops everybody? You know, is there a three, you know, I was talking to Eric Schmidt about this. Is there a three mile island like event?

Mustafa Suleyman

吓到所有人,但不会杀死任何人。

Scares everybody but doesn't kill anybody.

Host

嗯,埃里克·施密特特别说过,他希望有 100 人死亡,因为在他看来,这是至少能引起政府注意并促成某种解决方案的数字。戴夫,请继续。

Well, Eric Schmidt said specifically he's hoping for a 100 deaths because that's in his mind the least that would get the attention of the government and would cause some kind of a solution. Dave, continue, please.

Host

嗯,有趣的是你提到达里奥、山姆和伊利亚,你们显然互动不少。米拉是那个圈子里的吗?安德烈是吗?因为就像我们刚才讨论的,竞争正在升温,这很有意思。你知道,达里奥一开始是纯粹从安全角度出发的,我想伊利亚也是。但现在我们正处于自我改进的边缘,很明显,公司之间确实存在严重的——我不会说是裂痕——但确实在激烈竞争。我是说真的在竞争,而且我知道微软,当我写第二个商业计划时——我卖掉第一家公司后写下一个商业计划,第一句话就是‘远离微软’,因为当时微软占了科技行业市值的一半,微软的计划是规模翻倍。现在世界更平衡了,有微软、谷歌和 Meta,但当时微软势不可挡、占据主导,所以只能避开。但微软似乎总是赢。而我们正处于自我改进的边缘,至少在我看来是这样。那么,现在还是‘大家一起吃饭讨论安全’吗,还是每个人都全力以赴了?

Well, so it's interesting that you say Daario and Sam and Ilia, like you guys obviously must interact quite a bit. Is Meera part of that gang? Is Andre part of that gang? Are you like because this is it's interesting to think about the competition heating up like we were just talking about. And you know, Daario started from this position of pure safety and I think Ilia did too. But now we're right on the cusp of self-improvement and it's really really clear that there are serious I wouldn't say fissures but but the companies are now really racing. I mean really racing and and I know Microsoft you know when I wrote my my second business plan first company I sold next business plan I was writing the first sentence was stay out of Microsoft's way because because at the time you know Microsoft had half the market cap of tech was Microsoft and Microsoft's plan was to double in size we have a much more balanced world now with Microsoft and Google and Meta but at the time Microsoft was just unstoppable and dominant and so just stay out of the way but Microsoft seems to always win Right. There's and and we are right on the edge of self-improvement at least as far as I can tell. So, is it still, you know, let's all get together and have dinner and talk about safety or is everybody now in full board?

Mustafa Suleyman

不,绝对有。我认为递归自我改进如果成功,可能是一个临界点。想想看,目前有软件工程师在循环中,他们生成后训练数据,对数据质量进行消融实验,用基准测试评估,生成新数据,这大致就是循环。这相当昂贵、缓慢且耗时,而且并非完全封闭。我认为很多实验室都在竞相封闭这个循环,让各种模型充当评估质量的裁判、生成新训练数据的生成器、以及推理哪些数据该包含、哪些质量更高的对抗模型。然后这些数据显然会反馈到后训练过程中。所以,封闭这个循环肯定会加速 AI 发展。有些人推测这会增加风险——我认为确实可能增加风险——但有些人推测这是通往智能爆炸的潜在路径。

No, definitely. I think that's definitely there. I think the recursive self-improvement piece is probably the threshold moment if it works. And if you think about it at the moment, there are software engineers who are in the loop who are generating post-training data, running ablations on the quality of the data, running them against benchmarks, generating new data and that's sort of broadly the loop. And that's kind of expensive and slow and it takes time and it's not completely closed and I think a lot of the labs are racing to sort of close that loop so that various models will act as judges evaluating quality, generators producing new training data, adversarial models that are like reasoning over which data to include and what's higher quality. And then obviously that's then being fed back into the post-training process. So, like closing that loop is going to speed up AI development for sure. Some people speculate that that adds I mean okay I think it probably does add more risk but some people speculate that it's a potential path to a fume you know an intelligence explosion.

Host

是的。

Yeah.

Mustafa Suleyman

我绝对认为,如果算力不受限,且没有人类参与或控制,那确实可能带来更多风险。但算力不受限是一个很大的说法,意味着需要大量算力。所以,是的,我们确实在朝着越来越有风险的方向迈进。

And I definitely think with unbounded compute and without human in the loop or without control that does potentially create a lot more risk. But unbounded compute is a big claim. I mean that would mean need a lot of compute. So yeah, we're definitely taking steps towards more and more risky stuff.

Host

我能问一个非常具体的问题吗?因为你在微软已经一年半了,在真正的递归自我改进(即将到来)之前,有 AI 辅助芯片设计。你知道,PyTorch 栈中的层很笨重,但现在用 AI 穿透栈并进行优化非常容易,比如构建自己的内核,获得 2、3、4 倍的性能提升。但显然 OpenAI 正在构建定制芯片,而 TPU7 刚刚推出。你刚到微软时,首先,我知道有很多量子芯片工作,但有没有类似 TPU 的工作?

Can I ask you a really specific question about that because you know the year and a half now at Microsoft um before true recursive self-improvement which is imminent there's AI assisted chip design and this you know the layers in the PyTorch stack are very clunky but now it's really easy to use the AI to punch through the stack and optimize you know build your own kernels get 2 3 4x performance improvement but clearly OpenAI is now working to build custom chips and the TPU7s just came out. When you arrived at Microsoft, first of all, was I I know there's a lot of quantum chip work going on, but was there any work going on similar to the TPU work?

Mustafa Suleyman

是的。也有芯片方面的努力。而且,我认为进展相当不错。我的意思是,我们有几个不同的项目在推进,还没有公开讨论过,但芯片肯定会是其中的重要部分。

Yep. There's also a chip effort. And you know, I think progress has been pretty good. I mean, I think that we've got a few different irons in the fire that we haven't sort of talked about publicly yet, but I think the chips are going to be an important part of it for sure.

Host

是的。那些是内部努力。那些团队在你手下吗?那是你的一部分……

Yeah. Those are internal efforts. Are those teams under you? That's part of your...

Mustafa Suleyman

不,它们属于整个公司。是的。

No, I mean they're in the broader company. Yeah.

Host

好的。有意思。

Okay. Interesting.

遏制 vs 对齐 Containment vs Alignment

Host

我想稍微换个话题,谈谈你的书《即将到来的浪潮》。我非常喜欢它。我听了有声版,很喜欢你自己朗读。

I want to switch subject a little bit and come to your book The Coming Wave. I enjoyed it greatly. I listened to it. I love the fact that you read it.

Mustafa Suleyman

谢谢。

Thank you.

Host

是的。我告诉孩子们我读书。他们说,不,爸爸,你听书,你不再读书了。我想读一下我写的内容,因为这很重要。你把遏制问题定义为我们时代的决定性挑战,警告说随着这些技术变得更便宜、更易获取,它们将不可避免地扩散,变得几乎无法控制。这造成了一个可怕的困境:未能遏制它们会带来灾难风险,比如工程化流行病——你很多担忧都在生物领域,我作为生物学家和医生同意这一点——或者深度伪造导致民主崩溃等等,但执行遏制所需的极端监控可能导致极权主义反乌托邦。所以你说我们需要在混乱与暴政之间走一条狭窄的道路,这是一条非常微妙的线。因此你提出了遏制策略,包括技术安全措施、严格的全球监管、硬件供应链的瓶颈、国际条约。我们在这方面做得怎么样?

Yeah. I tell my kids I read books. Go. No, Dad. You listen to books. You don't read books anymore. I want to read what I wrote here because it's important. So you identified the containment problem as the defining challenge of our era, warning that as these technologies become cheaper and more accessible, they will inevitably proliferate, making them nearly impossible to control. This creates a terrifying dilemma. Failing to contain them forces risk for catastrophe like engineered pandemics and a lot of your concerns were in the biological world, and I agree being a biologist and a physician, or potentially democratic collapse with deepfakes and all of that, but the extreme surveillance required to enforce containment could lead to a totalitarian dystopia. So you say we need to navigate this narrow path between chaos and tyranny and that is a very fine line to navigate. So you propose a strategy of containment. This includes technical safety measures, strict global regulations, choke points on hardware supply, international treaties. How are we doing on that?

Mustafa Suleyman

是的,退一步区分对齐和遏制很重要。安全工程要求我们两者都做好。我实际上认为我们必须先做好遏制,再做好对齐。对齐有点像母性本能:它是否认同我们的价值观?它会在乎我们吗?它会善待我们吗?遏制则是我们能否正式限制并划定其自主性的边界,我们是否……

Yeah, I mean it's kind of important to just take a step back and distinguish between alignment and containment. The project of safety requires that we get both right. And I actually think we have to get containment right before we get alignment right. Alignment is the kind of like maternal instinct thing. Does it share our values? Is it going to care about us? Is it going to be nice to us? Containment is can we formally limit and put boundaries around its agency and are we...

Host

为了所有人?

For everybody?

Mustafa Suleyman

不仅为我们自己,而是为所有人。是的。

Not just for ourselves, for everybody. Yeah.

AI 时代的遏制与稳定 Containment and Stability in the Age of AI

Mustafa Suleyman

我认为挑战的一部分在于,一个坏人拥有如此强大的东西,在十年或二十年后,真的可能破坏整个系统的稳定。这个系统就是人类,全球人类系统。正如你所说,随着一切变得高度数字化,虚拟世界确实变成了元宇宙。尽管这个概念很快流行又过时,但我认为它仍然是正确的框架,因为一切将主要变得数字化、超连接、即时和实时。因此,一对多的效应突然被大幅放大。显然我们在社交媒体上看到了这一点,但现在想象一下,传播的不仅仅是文字,而是实际行动。智能体能够入侵系统,或者存在于地球上数十亿个人形机器人中,同时涉及原子和比特。所以,平衡需要一种我们今天世界上还没有的监控。我们当然没有物理上的监控。网络实际上被监控得非常厉害,我认为比人们预期的要多。某种形式的监控对于创造和平是必要的。就像我们在三四五百年前围绕政府集中了权力和税收或军事力量,这推动了进步。实际上,那种秩序释放了科学、技术和稳定。所以问题是如何以现代方式施加稳定,既不是极权主义,也不放任自流到自由意志主义的灾难。我认为认为对付枪的最好办法是枪,或者我们都有自己的 AI,它们会相互抵消形成稳定平衡,这种想法是天真的——那不会发生。

I mean, I think that is part of the challenge is that one bad actor with something that is really this powerful in a decade or two decades or something, really could destabilize the rest of the system. And so, the system being humanity, global humanity system. Just as you said, as everything becomes hyperdigitized, the verse does become the metaverse. Even though that kind of went in and out of fashion very quickly, it's still, I think, the right frame in a way because everything is going to become primarily digitized and hyperconnected and instant and real time. And so the one-to-many effect is suddenly massively amplified. I mean obviously we see it on social media but now imagine that it's not just words that are being broadcast. It's actually actions. It's agents capable of breaking into systems or sort of, and they're resident in humanoid robots at a billion on the planet, and that too. Both atoms and bits. So equilibrium requires that there is a type of surveillance that we don't really have in the world today. I mean we certainly don't have it physically. The web is actually remarkably surveilled. I think surprisingly more than people would expect. And some form of that is necessary to create peace. Just as we centralized power and taxation or military force and taxation around governments, 3 or 4 or 500 years ago, and that's been the driving force of progress. Actually, that order unleashed science and technology and stability. So the question is how do we impose stability in a modern way that isn't totalitarian but also doesn't relinquish it to a libertarian catastrophe. I think it's naive to think that somehow the best defense against a gun is a gun and the idea that we're all going to have our own AIs and that's going to create a steady equilibrium where all the AIs neutralize each other—that ain't going to happen.

Host

我有点希望有一个超级智能,像至尊魔戒那样统治一切,提供……我不担心,怎么说呢?我担心的是彼得,你是在希望一个单一实体。

I mean, part of me hopes for a super intelligence that is the ring to rule them all and provides... I'm not worried about, how do I put it? I'm worried about Peter, you're hoping for a singleton.

Mustafa Suleyman

是的,听起来就是这样。嗯,我有点……真是让我震惊。真的吗?是的。我的意思是,我认为我们正在接近的复杂程度,那种平衡行为极其困难,你不能推一根绳子,但有没有某种机制可以拉动它。我们应该找个时间辩论一下。

Yeah, that sounds like what's going on. Well, you know, part of me is like, color me shocked. Really? Yeah. I mean I imagine that the level of complexity we're mounting towards, that balancing act is extraordinarily difficult and you can't push a string but is there some mechanism to pull it forward. We should have this debate sometime.

Host

有些人会称政府,至少历史上,是对暴力的地理垄断。我听到的是某种对智能的垄断,或者至少是对智能暴露的能力进行垄断,以便围堵和遏制 AI。但据我所知,这与过去几年我们看到的情况完全相反。10 或 15 年前,那些纸上谈兵的 AI 对齐研究者会说,人类不会愚蠢到在拥有类似通用智能的东西时,就给它终端访问权或经济访问权,而我们恰恰这么做了。那是 OpenAI、谷歌的时刻。但这令人担忧,对吧?所以谷歌开发了所有这些技术,内部保留,直到某个首字母为 OpenAI 的参与者发布了它,然后别无选择只能跟进。

Some would call government, at least historically, a geographic monopoly on violence. And what I think I'm hearing is some sort of monopoly on intelligence or at least capabilities exposed to intelligence in order to ring fence to contain AI. But that's the exact opposite as far as I can tell of what we've seen over the past few years. People used to armchair AI alignment researchers 10 or 15 years ago would say humanity wouldn't be so stupid the moment we have something resembling general intelligence as to give it terminal access or to give it access to the economy and that's exactly what we did. There was the OpenAI, Google moment. And yet that's concerning, right? So Google develops all this technology, holds it internally until some actor happens to have initials OpenAI releases it and then there's no other option but to follow suit.

Mustafa Suleyman

我不太担心。看看 Anthropic,它自称是一个非常注重对齐的组织。Anthropic 发布了模型控制协议,目前至少是模型与环境交互的标准方式。许多 AI 研究者之前说过,在通用智能出现之前我们绝对不想这样做。所以我很好奇。在你看来,既然经济有各种压力,包括现代图灵测试,促使智能体与整个世界交互,做与遏制完全相反的事情,我们为什么要开始遏制?

I'm less concerned by it. If you look at Anthropic for example, which prides itself on being a very alignment forward organization. Anthropic released the model control protocol which is now the standard way at least for the moment for models to interact with the environment. What many AI researchers said exactly we did not want to do prior to general intelligence. So I'm curious. In your mind, given that the economy has every economic pressure including modern Turing test to empower agents to interact with the entire world and to do the exact opposite of containment, why would we start containing?

Host

遏制不是那么二元对立的,对吧?我们一直在遏制各种事物。你汽车引擎中的强大力量是被遏制且大致对齐的,周围有整套监管机制,从安全带、车辆排放、照明、驾驶教育到高速公路限速。那是健康的功能性监管,使我们能够集体互动。显然,这复杂了好几个数量级,因为这些不是汽车,而是某种数字人,但这并不意味着我们不应该努力限制它们的边界。也不意味着我们必须集中化。顺便说一句,答案不是我们在海外建立一个极权主义的智能国家。

Containment is not that binary, right? I mean we contain things all the time. We have powerful forces in the engine in your car that is contained and broadly aligned, and there is an entire regulatory apparatus around that from seat belts to vehicle emissions to lighting to driver ed to freeway speeds. That's healthy functional regulation enabling us to collectively interact with each other. Now obviously it's multiple orders of magnitude more complex because these things are not cars, they're sort of digital people, but that doesn't mean we shouldn't be striving to limit their boundaries. Nor does it mean that we have to centralize. By the way, the answer isn't that we have a totalitarian state of intelligence overseas.

Mustafa Suleyman

不,我认为只是本能地,当你开始深入思考时,很容易走向那个方向。显然我们确实有集中化的力量,但即使在美国,我们有军队、陆军师、警察部门,它们嵌套在不同层级,系统有制衡,这就是我们开始思考设计的东西。

No, I think it's just instinctively it can be easy to go there when you start to think it through. It's like obviously we do have centralized forces but even in the US we have military, divisions of the army, divisions of the police force, they're nested up in different layers, there's checks and balances on the system and that's kind of what we got to start thinking about designing.

Host

那个开车的类比很好,继续下去,AI 的复杂性差异非常大,但时间线也是。开车从 1910 年演变到今天?实际上是 19 世纪末。所以相关法律,安全带出现在那个时间线的 80%处。所以有大量时间迭代。而这里,时间很少,复杂得多。所以你有愿景吗?但我完全同意。我们需要一个快速的遏制框架。你对如何实现有什么想法吗?

That analogy to driving is a great one and just to follow through on it, the complexity difference is very high for AI but the timeline also. I mean driving evolved from what, 1910 to today? Late 1800s. So the laws related, seat belts came out 80% of the way through that timeline. So lots and lots of time to iterate. Here, very little time and immensely more complex. So do you have a vision? But I completely agree. We need a framework for containment fast. And do you have a thought on how we're going to...

Mustafa Suleyman

我认为这样做也有很好的商业动机,对吧?许多公司知道,我们的社会运营许可要求我们比以往任何时候都更对外部性负责。我们不在罗伯特·巴伦时代。不在石油时代。不在吸烟时代,对吧?我们学到了很多。不是全部。仍然有很多冲突,但确实与上次有所不同。我认为这是更乐观的一个原因。再加上商业动机,商业动机和外部性的转变。

I think that there's also a good commercial incentive to do this, right? I think that many of the companies know that our social license to operate requires us to take more accountability for externalities than ever before. We're not in the Robert Baron era. We're not in the oil era. We're not in the smoking era, right? We've learned a lot. Not everything. There's still a lot of conflicts, but it really is a little bit different to last time around. And I think that's one reason to be a bit more optimistic. Plus there's the commercial incentive, the commercial incentive and the kind of externalities shift.

AI 安全合作 Cooperation on AI safety

Host

那么,如果埃里克·施密特是对的,发生了某种放射性或生物事件,导致 100 人死亡,然后电话开始响,所有人都立刻到白宫来。首先,你想要那个电话吗?那是你人生计划的一部分吗?然后,你信任社区里的哪些人参与应对?

So, if Eric Schmidt is right and something either radiological or biological happens and there's 100 deaths and then the phone starts ringing, everyone come to the White House right now. Well, first of all, do you want that call? Is that part of your life plan to take that call and react to it? And then who else do you trust in the community to be part of that reaction?

Mustafa Suleyman

听着,我认为在未来 20 年内,会有一个时刻,地球上的每个人,包括中国和其他所有大国,都会完全理解在安全、遏制和对齐上进行合作的必要性。这是完全理性的自保行为。这些系统非常强大,对使用模型的恶意行为者造成的威胁和对受害者一样大。我认为这将催生合作意愿,尽管在当前世界两极分化的阶段很难感同身受,但我相信这终将到来。统一全人类的第一件事就是外星人入侵,而那个外星人入侵可能就是潜在的失控超级智能。

Look, I think that there is going to be a time in the next 20 years where it will make complete sense to everybody on the planet, the Chinese included, and every other significant power, to cooperate on safety and containment and alignment. It is completely rational for self-preservation. You know, these are very powerful systems that present as much of a threat to the person, the bad actor that is using the model as it does to the victim. And I think that will create an interest in cooperation, which, it's kind of hard to empathize with at this stage given how polarized the world is, but I do think it's coming. I mean the number one thing to unify all of humanity is an alien invasion, and that alien invasion could be a potential for a rogue superintelligence.

Host

是的。好吧。那我的第一个问题呢?那是你人生的使命吗?我的意思是,只有少数人——我在麻省理工或其他地方遇到的很多人都有这种想法,觉得某个地方有人已经搞定了。你知道,政府里的某个人一定在考虑这个。但你去过那里,对吧?那里没有人。

Yeah. Okay. What about the first part of my question? Is that part of your calling in life? I mean there's only a handful like I think a lot of people that I meet around MIT or elsewhere, they have this vision that somebody has it figured out somewhere. You know someone in government somewhere must be thinking about this. But you've been there, right? There's no one there.

Mustafa Suleyman

我们是房间里的大人。你是这个意思吗?

We're the adults in the room. Is that what you're saying?

Host

是的,绝对。这个房间之外无处可去。

Yeah, definitely. There's nowhere to go from this room.

Host

戴夫在问那个烟雾缭绕的后屋,所有前沿实验室的负责人都在那里秘密交换安全技巧。

Dave is asking for the smoke filled back room where the leads of all the Frontier Labs are secretly swapping safety tips.

Mustafa Suleyman

是的,差不多。我认为实际上智能存在于烟雾缭绕的房间之外。那种认为决策是在董事会或白宫战情室做出的想法,其实智能是在这些迭代互动的大球中凝聚的,而这正是推动世界前进的力量。所以对话就在这里发生:你的听众,所有其他播客主,网上的每个人,我们都在共同推动那个知识库向前发展。

Yeah, something like that. Yeah, I think that in practice intelligence exists outside of the smoky room. I think that the notion that decisions get made in the boardroom or in the White House situation room, actually intelligence coalesces in these big balls of iterative interaction, and that's what's propelling the world forward. So this is where the conversation's happening: your audience, all the other podcasters, everyone online, we're collectively trying to move that knowledge base forward.

人本超级智能与教育 Humanist super intelligence and education

Host

11 月,你宣布推出 Humanist Super Intelligence,专注于三个应用:医学、伴侣和清洁能源。我想深入了解一下,但我很好奇你没有把教育包括在内。我们的听众是企业家和 AI 构建者,我认为教育和医疗一样,现在都是待开发的领域。教育也是。

In November, you announced the launch of Humanist Super Intelligence, focused on three applications in particular: medicine, companions, and clean energy. I'd love to double click on that a little bit, but I was curious that you didn't include education in that space. We have an audience of entrepreneurs and AI builders, and I think education, as much as healthcare, is up for grabs right now. Education is too.

Mustafa Suleyman

完全同意。我认为这已经在整个行业发生了。现在,口袋里有一个拥有博士学位的专家老师,并能根据你的定制学习风格调整课程,从未如此容易。目前它还不能做的是,在多次课程中演变或策划一个扩展的学习计划,但我们已经接近了。几个月前我们发布了一个叫测验的功能。所以任何话题,不仅仅是传统学校教育,它都能为你设置一个小课程、一个测验,而且是互动和可视化的,你可以随时间追踪学习。我也非常乐观。这是一个巨大的解锁。

Totally agree. I think it's already happening across the whole industry. I mean, it's never been easier to get access to an expert teacher in your pocket that has essentially a PhD and that can adapt the curriculum to your bespoke learning style. The bit that it can't do at the moment is to evolve or curate an extended program of learning over many sessions, but we're just around the corner from that. We released a feature just a few months ago called quizzes. So on any topic, not just traditional school education, it can set you up with a mini curriculum, a quiz, and it's interactive and visual, and you can track your learning over time. I'm very optimistic about that too. It's a huge unlock.

Host

我们播客现在经常讨论的一个辩论是:你上大学吗?你读研究生吗?我的意思是,这是有史以来最激动人心的建设时期。我不知道你是否想接着这个话题,戴夫。

One of the debates we have right now on the podcast on a pretty regular basis is: do you go to college? Do you go to grad school? I mean, this is the most exciting time to build ever. I don't know if you want to follow on that, Dave.

Host

天哪,我经常这样做。在校园里这对我来说很棘手,因为我在麻省理工、斯坦福和哈佛教书。这个机会窗口如此短暂而紧迫,现在在 AI 后 AGI 时代如何成功非常清楚。我的意思是,谁能预测?没人知道。但就在此时此地,你看到这些初创公司的估值,就像我们昨晚看到的。我不提具体数字,但都是数十亿。

Well, God, I do this constantly. It's really tricky for me on campus because I teach at MIT, Stanford, and Harvard. This window of opportunity is so short and so acute, and it's really clear how you succeed right now in AI post-AGI. I mean, who could predict? Nobody knows. But right here, right now, you see these startup valuations like we were last night. I won't mention it, but billions.

Mustafa Suleyman

开盘估值 40 亿美元。十亿美元。是的。

An opening valuation of $4 billion. Billion-dollar. Yeah.

Host

通过把合适的一群人聚在房间里。是的。是的。我其实想问这个,因为你在 Inflection 的时机很早,事后看来更早,但现在你有了新一波,比如 Mera Marotti、Helia 和其他几个,Liquid AI,它们都有数十亿美元的估值。

By collecting just the right group of people in the room. It's Yep. Yep. I wanted to ask about that actually because your timing on Inflection was early, in hindsight earlier, but now you've got the new wave with Mera Marotti and Helia and a couple of others, Liquid AI, that all have multi-billion dollar valuations.

Mustafa Suleyman

我以为我们为 20 人团队在未盈利时设定了估值标准,但两年前我们只是小虾米。

Thought we set some standards on valuations pre-revenue with a 20 person team, but we're just a minnow then a whole two and a half years ago.

Host

就两年吗?天哪。

Is that all it was? Oh my god.

Mustafa Suleyman

我想是三年。是的。天哪。

Three years, I think. Yeah. Jeez.

Host

你认为随着智能成本变得低到无法计量,至少以市值衡量的人力资本价值会呈反比渐近趋向无穷大吗?

You think as the cost of intelligence becomes too cheap to meter, the value ascribed at least in terms of market cap to human capital is sort of inversely asymptotic going to infinity.

Mustafa Suleyman

奇怪的是,这是因为时机的压力,对吧?而且仍然有一批非常集中的人能做这些事。资本过剩,急于分一杯羹。这可能不是世界上最聪明的资本,但非常急切。所以……我必须问你,因为这个问题让我心痒难耐。你知道,亚历克斯在麻省理工的新生室友是纳特·弗里德曼,而且是新生前的室友。所以纳特离开了,最终成为 Safe Super Intelligence 的联合创始人。我没问过他,不知道你问过没有,但他离开去 Meta 工作,我敢肯定很大一部分吸引力是算力。

Weirdly it is because of the pressure on timing, right? And there's still a pretty concentrated pool of people that can do this stuff. There's an overabundance of capital that's desperate to get a piece of it. It might not be the smartest capital the world's ever seen, but it's very eager. So that's... I have to ask you because it's burning a hole in my pocket, but you know, Alex's freshman roommate at MIT was Nat Friedman, and pre-freshman roommate. So Nat goes off and he ends up co-founder of Safe Super Intelligence. I haven't asked him, I don't know if you've asked him yet, but he leaves to become the guy at Meta, and I've got to believe a huge part of that attraction is the compute.

Mustafa Suleyman

是的。

Yeah.

Host

所以你现在的情况非常相似,对吧?你有自己的初创公司,你筹集了十亿或十五亿。你可以建造它。你可以得到你的 2 万块英伟达。等等,这里有微软。3000 亿现金流和大量算力。那是很大一部分……

And so here you are, very similar situation, right? You've got your startup, you've got a billion or whatever billion and a half that you've raised. You can build it. You can get your 20,000 Nvidia. Well, wait a minute. Here's Microsoft. 300 billion of cash flow and a huge amount of compute. Was that a big part of the...

Mustafa Suleyman

是的。

Yeah.

大公司的结构性优势 Structural advantage of big companies

Mustafa Suleyman

我的意思是,更不用说我们为单个研究人员或技术人员支付的价格,以及未来 10 年而非仅仅两年所需的投资规模。我认为,身处大公司显然具有结构性优势,而且未来 5 到 10 年,要保持在最前沿需要数千亿美元。所以,总结一下,那些目前以 200 亿或 500 亿美元估值融资的公司毫无机会。

I mean, not to mention the prices we're paying for individual researchers or members of technical staff, and just the scale of investment required not just in two years but over 10 years. I think it's clearly a structural advantage to be inside a big company, and I think it's going to take hundreds of billions of dollars to keep up at the frontier over the next 5 to 10 years. So, finishing that thought, the companies that are raising money at a 20 or 50 billion dollar valuation right now have no chance.

Host

好吧,我接受这个沉默。

Okay, I'll take that silence.

Mustafa Suleyman

但我认为这取决于情况。显然,短期内,如果我们突然出现智能爆炸,那么很多人可以同时达到那个水平。但同时,你必须用这些技术构建产品,这需要分销渠道。所有传统机制仍然适用。你能足够快地转化吗?如果未来 5 年内发生这种情况,一切都会变得非常奇怪,无法辨认。有太多涌现的因素相互影响。很难说,我认为这种不确定性部分推动了估值的泡沫。人们在想,‘我不知道,我想成为……’里德称之为‘傻瓜保险’。

But I think it depends. Obviously, in the near term, if suddenly we have an intelligence explosion, then lots of people can get there simultaneously. But at the same time, you have to build a product with those things, which requires distribution. All the traditional mechanisms still apply. Are you going to be able to convert that quickly enough? Everything goes really weird if that happens in the next 5 years. It's unrecognizable. There are so many emergent factors playing into one another. It's hard to say, and I think that ambiguity is partly what's driving the frothiness of the valuations. People are thinking, 'I don't know, do I want to be...' Reed calls it schmuck insurance.

Host

是的。几个月前我们请里德上过播客。他非常聪明。

Yeah. We had Reed on the pod here a couple months ago. He's brilliant.

给学生的建议与公共服务 Advice for students and public service

Mustafa Suleyman

对那位即将高中毕业的学生来说,现在该学什么?毫无疑问,你仍然需要学习这两个学科。我认为哲学和计算机科学在很长一段时间内仍将是两大基础。你应该上大学吗?绝对应该。人文教育、由此带来的社交性、以及大学提供三年时间让你在课程内外思考和探索的好处——这是一个巨大的特权。人们不应该放弃它。那是金子。所以我总是鼓励人们这样做。显然,我也辍学了,但我仍然认为那是一件很酷的事。当时感觉是对的。但另一件事是:进入公共服务领域。

To that graduating high school student, what do you study these days? There's no question that you still have to study both disciplines. Philosophy and computer science are going to remain, I think, the two foundations for a long time. Should you go to college? Absolutely. Human education, the sociality that comes from that, the benefit of the institution having 3 years to basically think and explore in and out of your curriculum—this is a huge privilege. People should not be throwing that away. That is golden. So I always encourage people to do that. Obviously, I did also drop out, but I still think it was a cool thing to do. It felt right at the time. But the other thing is: go into public service.

Host

是的。我尊重你人生中那段经历,它赋予了你非常人本主义的观点。

Yeah. I respect that part of what you did in that sequence in your life, which gave you this very humanist point of view.

Mustafa Suleyman

是的。那非常艰难,也非常不同,并非本能正确,但我学到了很多。那是我经历中非常有影响力和重要的一部分,尽管时间很短——基本上只有几年。我认为,如果你看看我们生态系统中的参与者——企业、学术界、新闻机构,现在还有播客世界——实际上,我们的政府可能在制度上最薄弱,还有我们的民主进程,但实际上是我们的公务员制度。这是因为自里根时代以来,公共服务领域的地位、声誉和尊重遭受了五十年的打击。我认为这实际上是一种讽刺,因为我们比以往任何时候都更需要那种情感、那种精神和那些能力。

Yeah. And it was really hard and very different, and it wasn't instinctively right, but I learned a lot. It was a very influential and important part of my experience, even though it was very short—like a couple years basically. I think if you look at the actors in our ecosystem today—corporations, academics, news organizations, now the podcast world—it's really our governments that are probably institutionally the weakest, and our democratic process, but actually our civil service. That's because there's been five decades of battering of the status, reputation, and respect that goes into being part of the public service, post-Reagan and that. I think that's actually a travesty because we need that sentiment, that spirit, and those capabilities more than ever.

Host

我想我刚才听到你说的是,我们需要在公共部门、公共服务中注入更多智能。政府中的 AI 呢?你认为政府需要……特别是政府中的智能体式 AI 呢?

I think maybe what I just heard you say, correct me if I'm wrong again, is we need more intelligence in the public sector, in public service. What about AI in government? Do you think the government needs... and what about agentic AI in the government in particular?

Mustafa Suleyman

当然,所有相同的注意事项都适用。但值得指出的是,政府内部对 Copilot 的采用率非常高。它在综合文档、转录会议、总结笔记、促进讨论以及在适当时候提出行动方面做得非常出色。这显然会节省大量时间并改善决策。

For sure, with all the same caveats that apply. But the rate of adoption for what it's worth of co-pilot inside government is really high. It does a brilliant job of synthesizing documents, transcribing meetings, summarizing notes, facilitating discussion, and chipping in with actions at the right time. It's clearly going to save a lot of time and improve decision-making.

AI 遏制 AI 与政府监管 AI containing AI and government regulation

Host

那么,也许给讨论画上一个圆满的句号:这难道不是一种 AI 遏制 AI 的形式吗?如果 AI 渗透政府,AI 渗透经济,而政府监管经济,这不就是 AI 自我监管的防御性共同扩展吗?

So then maybe to tie a nice bow on the discussion: isn't that arguably a form of AI containing AI? If AI is infusing the government and AI is infusing the economy, and the government is regulating the economy, isn't this just defensive co-scaling with AI regulating itself?

Mustafa Suleyman

是的。每个人都会同时使用 AI 来追求我们各自的目标,而这些目标将保持不变。想创业的人、想写学术论文的人、想创办文化团体和娱乐项目的人——每个人都会以某种方式被赋能。他们的能力将因拥有这些工具而得到放大。显然,政府也包括在内。

Yeah. Everyone is going to use AI all at the same time to pursue the agendas that we all have, which are going to remain the same. People who want to start companies, people who want to write academic papers, people who want to start cultural groups and entertainment things—everyone is just going to be empowered in some way. Their capability is going to be amplified by having these tools. Obviously, the government included.

量子计算与 AI Quantum computing and AI

Host

太好了。穆斯塔法,非常感谢你周五晚上抽出时间。很感激能与你进行这次对话。戴夫、亚历克斯,谢谢。戴夫,你想问最后一个问题吗?

Nice. Mustafa, thank you so much for taking the time on a Friday night. Grateful to have this conversation with you. Dave, Alex, appreciate it. Want a final question from you, Dave.

Host

最后一个问题,如果我有的话。预测:量子计算目前与 LLM AI 领域发生的事情毫无关系。一切都基于 Nvidia 芯片上的矩阵乘法,很快将是 TPU 和其他定制芯片。最佳猜测:六、七年后,AI 非常擅长编写代码和编译,并能解决量子操作。量子芯片会变得相关,还是仍然处于边缘,或者一切都会迁移到量子,而微软可以利用其领先优势?

Final question if I have one. Prediction: quantum computing right now has nothing to do with what's going on in LLM AI. It's all matrix multiplications on Nvidia chips and soon to be TPUs and other custom chips. Best guess: six, seven years from now, AI is very good at writing code and compiling and can figure out quantum operations. Are quantum chips relevant, or are they on the sideline still, or is everything ported over to quantum and Microsoft can take advantage of its lead?

Mustafa Suleyman

是的,我认为它将成为组合中的重要部分。我认为,相对于我们花在谈论 AI 上的时间,它有点被低估了。有点像合成生物学。我认为,尤其是在一般讨论中,人们没有意识到这两波浪潮,它们将同样具有影响力,并在 AI 成为焦点时同时冲击。

Yeah, I think it's going to be a big part of the mix. I think it's sort of under-acknowledged relative to the amount of time we spend talking about AI. It's a bit like synthetic biology. I think that especially in the general conversation, people aren't grasping those two waves, which are going to be just as impactful and crash at the same time that AI is coming into focus.

加速 AI 科学与工程 Accelerating AI for science and engineering

Host

好了,你在这里听到了。这是一个结束问题,也许是为了吸引你更加速主义的一面。观众可以做些什么来加速 AI 用于科学、AI 用于工程?你认为限制因素是什么?我经常在播客中谈到最内层循环的概念:在计算机科学中,如果你想优化一个程序,你会找到循环中的循环,而你想优化最内层循环以优化整个程序。

All right, you heard it here. This is a closing question to appeal maybe to your more accelerationist side. What can the audience do to accelerate AI for science, AI for engineering? What do you view as the limiting factors? I often talk on the podcast about this notion of an innermost loop: in computer science, if you want to optimize a program, you find loops within loops, and you want to optimize the innermost loop to optimize the overall program.

限制因素:假设验证 Limiting factor: hypothesis validation

Host

你认为最核心的循环、限制因素是什么?如果听众有足够的能力,他们可以优化什么来加速实现未来 10 年的《星际迷航》式未来或经济?我们该怎么做?

What do you see as the innermost loop, the limiting factor, if you will, that the audience listening, if they're suitably empowered, can help optimize to speedrun maybe a Star Trek future over the next 10 years or a Star Trek economy? What do we do?

Mustafa Suleyman

是的,我认为很明显,大多数模型都会加快提出假设的速度。慢的部分将是在现实世界中验证假设。所以目前我们所能做的就是将越来越多的信息摄入我们自己的大脑,然后与一个随着你进步的单一模型共同使用,因为它正在成为你的第二大脑。例如,Copilot 现在在个性化方面做得非常好。你用得越多,它的答案就越能捕捉到你感兴趣的主题。而且它也在逐渐变得更加主动。它会轻轻推送与你之前讨论内容相关的新论文或新文章。所以,这虽然有点简单化,但你用得越多,它就越好,它越了解你,你就变得越好,因为它成为了你探究路线的辅助工具。

Yeah, I mean, I think it's pretty clear that most of these models are going to speed up the time to generate hypothesis. The slow part is going to be validating hypothesis in the real world. And so all we can do at this point is just ingest more and more information into our own brains and then co-use that with a single model that progresses with you because it's becoming like a second brain. For example, Copilot is actually really good at personalization now. Most of its answers, the more you use it, the more those answers pick up on themes that you're interested in. And it's also gently getting more proactive. So, it's kind of nudging you about new papers or new articles that come out that are obviously in tune with whatever you've been talking about previously. So, you know, it's a bit of a simplistic copout answer, but just the more you use it, the better it gets, the better it learns you, the better you become because it becomes this sort of aid to your own line of inquiry.

Host

所以,听起来你对听众的建议是多用 Copilot,这是加速这一过程的最佳催化剂。

So, that sounds like your advice to the audience is use Copilot more and that's the single best accelerant that you can do to speed this up.

Mustafa Suleyman

或者任何其他 AI。我是说,有很多好的。

Or any other AI. I mean, loads of great.

Host

我还听你谈到,能否构建一个物理系统,让 AI 能够 24/7 在封闭的暗循环中运行实验,从而从自然界挖掘数据?有很多公司正在这样做。Laya 是最近从哈佛和麻省理工出来的一家。

I heard you also talk about can you build the physical system that is going to enable AI to run the experiments in a 24/7 closed dark cycle to be able to mine nature for data, right? And there are a number of companies that are doing this. Laya is one recently out of Harvard MIT.

Mustafa Suleyman

我觉得这很令人兴奋,AI 正在成为我们的探索者,为我们收集数据。

I find that exciting where AI is becoming an explorer on our behalf gathering that data.

Host

是的,说得对。

Yeah. Spot on.

Mustafa Suleyman

嗯。

Yeah.

Host

再次感谢。

Thank you again.

Mustafa Suleyman

这太棒了。非常感谢。这是一次非常有趣的对话。

This has been great. Thanks a lot. It was a really fun conversation.

Host

是的,非常有趣。谢谢。

Yeah, really fun. Thanks.

Mustafa Suleyman

很感激,我的朋友。

Appreciate it, my friend.

Host

好的。很高兴见到你。

All right. Good to see you.

通讯推广 Newsletter promotion

Host

每周,我和我的团队都会研究未来十年将改变行业的十大技术元趋势。我涵盖的趋势包括人形机器人、AGI、量子计算,以及交通、能源、长寿等。没有废话。只有最重要、影响我们生活、公司和职业的内容。如果你想让我与你分享这些元趋势,我每周写两次通讯,通过电子邮件发送,只需两分钟阅读。如果你想比别人早十年发现最重要的元趋势,这份报告就是为你准备的。读者包括来自世界上最具颠覆性公司的创始人和 CEO,以及构建世界最具颠覆性技术的企业家。如果你不想了解即将到来的事物、为什么重要以及如何从中受益,那就不适合你。要免费订阅,请访问 demandis.com/metrends,比别人早十年获取趋势。好了,现在回到本期节目。

Every week, my team and I study the top 10 technology meta trends that will transform industries over the decade ahead. I cover trends ranging from humanoid robotics, AGI, and quantum computing to transport, energy, longevity, and more. There's no fluff. Only the most important stuff that matters that impacts our lives, our companies, and our careers. If you want me to share these meta trends with you, I write a newsletter twice a week, sending it out as a short two-minute read via email. And if you want to discover the most important meta trends 10 years before anyone else, this report's for you. Readers include founders and CEOs from the world's most disruptive companies and entrepreneurs building the world's most disruptive tech. It's not for you if you don't want to be informed about what's coming, why it matters, and how you can benefit from it. To subscribe for free, go to demandis.com/metrends to gain access to the trends 10 years before anyone else. All right, now back to this episode.

互动版:逐字朗读 + 针对本期提问 →