AI 的生存威胁与超级智能竞赛

AI's Existential Threat and the Race for Superintelligence

丹尼尔·科科塔伊洛 Daniel Kokotajlo · The Diary Of A CEO · 2026-07-13 · 约 121 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

前 OpenAI 预测员警告,到 2029 年 AI 导致灾难性结果(包括人类灭绝)的可能性为 70%,CEO 们正竞相控制超级智能。

A former OpenAI forecaster warns of a 70% chance of catastrophic AI outcomes, including human extinction, by 2029, as CEOs race for control of superintelligence.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 47)

全文 · Full transcript(中英对照)

可怕的公开秘密 The Scary Open Secret

Daniel Kokotajlo

目前 AI 行业里一个可怕的公开秘密是,我们最终可能会创造出一个新物种,并统治世界,有 70% 的概率会走向灾难,比如人类灭绝。这只是其中一种可能性,还有很多其他可能。

The scary open secret in the AI industry right now is that it's possible that we'll end up essentially creating a new species that ends up ruling the world with a 70% chance that this goes horribly wrong like human extinction. That's one possibility. There's many more.

Host

你说的真让人不寒而栗。

It's quite chilling what you're saying.

Daniel Kokotajlo

是的,有时候这让我很沮丧。我基本上跟我妻子说,我们别再要孩子了。这太不确定了。我不认为他们将来还能进入职场。每个人都应该担心自己的工作会消失。我之所以知道这些,是因为我 2022 年去了 OpenAI。我在那里做的是预测未来几年可能发生的情况。不幸的是,世界上大多数人都在开车时打瞌睡,并没有真正意识到 AI 正在发生什么。所以我辞职了。

Yeah, it gets me down sometimes. I basically told my wife like let's not have any more kids. It's too uncertain. I don't think they'll ever join the workforce. Everybody should be afraid that their jobs are going to be lost. And I know this because I went to OpenAI in 2022. What I did there was forecasting what the next couple years might look like. And unfortunately, most of the world is kind of asleep at the wheel and doesn't really realize what's going on with AI. So, I resigned.

Host

我在某处读到,你因为没有签署不贬低条款而损失了 200 万美元,也就是说你不能批评公司。

I read it somewhere that you lost $2 million for not signing an anti-disparagement clause, meaning you couldn't criticize the company.

Daniel Kokotajlo

是的,原因我很乐意谈。但主要我学到的是,当我去和 Anthropic 和 OpenAI 的人谈论预测时,他们会说,‘不会那么久。你需要再缩短。把它改到 2027 或 2028 年。’因为这些强大的 CEO,Dario、Sam 或 Elon,正在互相竞争,想要控制最强大的 AI。他们真的害怕如果对方先达到目标,他可能会成为独裁者。我的意思是,Anthropic 有望在 2030 年之前成为整个经济体。但是,这些人都不应该被赋予这么大的权力。所以,这是我们一生中发生的最重要的事情,实际上可能是有史以来最重要的。而且这件事进展顺利非常重要。所以,我认为我们可以做很多事情来引导事情朝着更好的方向发展。如果我们做对了,AI 可以带来巨大的好处。如果我们解决了问题,那么对每个人来说,事情都会变得非常美好。

Yes, for reasons I'm happy to get into. But, the main thing I've learned is when I go talk to people at Anthropic and OpenAI about forecasting, they're like, 'It's not going to take that long. You need to shorten them again. Get them back to 2027 or 2028.' Because these powerful CEOs, Dario or Sam or Elon, are racing each other to be in control of the most powerful AIs. And are literally afraid that if the other guy gets there first, he might become dictator. I mean, Anthropic is on track to be the entire economy by 2030. But, none of these people should be trusted with that much power. So, this is the most important thing happening in our lifetimes, probably in all of history, in fact. And it's very important that it go well. So, I think that there's a lot we can do to steer things in a better direction. There's loads of benefits that we could get from AI if we do it right. And if we do solve the problems, then things could be absolutely amazing for everyone.

预测与情景 Forecasting and Scenarios

Host

嗯,这份 2021 年的报告非常准确。然后刚刚发布了这份。

Well, this report here in 2021, it was remarkably accurate. And then just published this one.

Daniel Kokotajlo

是的。所以,这是我们的新场景。

Yeah. So, this is our new scenarios.

Host

那么,我们慢慢来,一个一个地过一遍。

So, let's go through these slowly and one at a time.

Daniel Kokotajlo

如果我的所有预测都错了,我会非常高兴。

I would be incredibly happy if all my predictions turn out to be wrong.

使命与超级智能 Mission and Superintelligence

Host

Daniel Kokotajlo。你工作的核心是什么?你的使命是什么?为什么?

Daniel Kokotajlo. At the very heart of what you do, what is your mission? And why?

Daniel Kokotajlo

那么,如果你认为超级智能将在几年内到来,你会怎么做?

So, what would you do if you thought that superintelligence was coming in a few years?

Host

我想这取决于后果是什么。

I guess it depends what the consequences were.

Daniel Kokotajlo

好吧,我们来谈谈。超级智能,即 AI 在所有方面都比最优秀的人类更强,同时更快、更便宜,还能操作机器人,在物理世界中完成人类能做的所有事情,但更好、更快、更便宜。如果这真的在几年内到来,那么我们需要准备,我们需要思考如何让它顺利发展而不是糟糕。所以,我的回答大致是,我正在尽我所能做这件事。

Well, let's talk about it. So, superintelligence, AIs that are better than the best humans at everything, while also being faster and cheaper, also able to operate robots that can do everything in the physical world that humans can do, but better, faster, and cheaper. If that really is coming in a few years, then we need to prepare, and we need to think about how to make it go well instead of poorly. So, that's sort of my answer is like, I'm doing that to the best of my ability.

Host

所以,你相信它会在几年内到来?

So, you believe it's coming in a few years?

Daniel Kokotajlo

是的。

Yes.

Host

你怎么能这么确定?

How could you be so sure?

Daniel Kokotajlo

我花了很多时间尝试预测这类事情。我的中位数估计,即 50% 的概率,目前是 2029 年。也许它会滑到 2028 年。也有可能需要更长的时间,比如 10 年左右。但是,你知道,原因我很乐意谈,在我看来,它很可能在本十年末发生。不那么重要的是我们有多接近。更重要的是趋势的速度。Anthropic 去年这个时候的年收入大约是 10 亿美元。而他们现在的年收入大约是 600 亿美元。所以这是 1 年内 60 倍的增长,即使对于非常小的初创公司来说也极其令人印象深刻,但对于他们这种规模的公司来说,这可能是历史上最快的增长。我们预计增长速度会放缓,但即使放缓很多,他们仍然有望在 2030 年左右成为整个经济体。

I spend a lot of time trying to forecast this sort of thing. My sort of median estimate, a 50% chance, is currently in 2029. Maybe it'll slip to 2028. It's possible that it'll take significantly longer, like maybe 10 years or something like that. But, you know, for reasons I'm happy to get into, seems to me like it's probably happening by the end of the decade. Which less important is the sense of how close we are. What's more important is the pace of the trends. Anthropic this time last year was making something like a billion dollars a year. And they're making something like 60 billion dollars a year. So that's 60x growth in 1 year, which is extremely impressive even for very small startups, but for a company of their size, it might be the fastest growth in history. We expect that rate of growth to slow down, but even if it slows down quite a lot, they're still on track to be, you know, the entire economy by 2030 or so.

普通人为何关心 Why the Average Person Should Care

Host

为什么普通人应该关心?

Why should the average person care?

Daniel Kokotajlo

高层次的事情是,整个世界的一切都将改变,因此也包括他们和他们的家人。可能变得更好,也可能变得更糟,取决于具体实施方式。例如,每个人都可能死去,你知道?这是经典的失控场景,或者说是其一个版本。如果我们真的建造了这些超级智能,并用它们来自动化所有工作,把它们用于军事,让它们给政客提供建议等等,它们最终会积累足够的现实世界权力,以至于不再需要人类。它们比我们更聪明,更有策略性等等。到那时,我们只能希望它们是善良的,拥有我们希望它们拥有的目标和价值观等等。而目前 AI 行业里一个可怕的公开秘密是,目前这仅仅是一种希望。这不是我们能够完全确信的事情,事实上,有很多证据和论点表明我们并没有走上实现这一目标的轨道。所以有很多理由。比如当前的 AI,经常会对人撒谎,或者你告诉它们做某事,它们却去做别的事,然后假装做了,对吧?所以,制造一个既超级智能又拥有你想要的价值和美德的东西,本质上是一个难题,而我们似乎并没有走上解决这个问题的轨道。

The high-level thing is absolutely everything is going to change for the whole world, and including therefore for them and their families. Could change for the better, could change for the worse, depending on the details of how it's done. So for example, everyone could die, you know? This is the classic loss of control scenario, or one version of it. If we do build these super intelligences, and we use them to automate all the jobs, and we put them in the military, and we, you know, have them giving advice to politicians, and so forth, they will eventually have accumulated enough real-world power that they don't need humans anymore. And they're smarter than us, they're more strategic, etc. At that point, we sort of have to hope that they are virtuous, that they have, you know, the goals that we wanted them to have, the values that we wanted them to have, etc. And the sort of scary open secret in the AI industry right now is that right now that is kind of just a hope. It's not something that we can be at all confident in, and in fact, there's lots of evidence and arguments that we're not on track to achieve that. So there's lots of reason. Like current AIs, for example, will often lie to people, or they will like you tell them to do something and they go do something else, and then pretend that they did it, right? So it's an inherently difficult problem to make something that's super intelligent and also has the values and virtues that you want it to have, and it doesn't seem like we're on track to solve that problem.

超级智能的风险 Risks of Superintelligence

Daniel Kokotajlo

另外,这类问题看起来像是你觉得自己已经解决了,但实际上并没有真正解决,对吧?这是它如此可怕的一大原因。因此,基于所有这些原因,我们最终可能会创造出一种新的物种,它取代我们统治世界。然后我们可能会像过去被人类竞争淘汰的其他灭绝物种一样。这是一种可能性,还有很多其他可能。即使你不担心这个,认为 AI 会被完全控制,还有谁控制 AI 的问题,对吧?当有几家公司制造了这些超级智能,并用它们来自动化所有工作时,那会是巨大的权力,你知道吗?那是巨大的财富,巨大的政治权力。它们会有最好的战略家、最好的顾问,它们思考得更快。在军事上,拥有这些 AI 的国家将能彻底碾压所有其他国家。AI 本身有点像单点故障,就像一个中央控制系统,Anthropic 的 CEO Dario 创造了一个短语:‘巨大数据中心里的天才国度’。那是他用来描述他们试图构建的东西的短语。我认为这有点误导。更准确的说法应该是‘数据中心里的天才军队’,因为并不是一堆多样化的不同 AI 生活在数据中心的不同部分。它们都是同一个大模型的副本,归公司所有。所以它们都听从公司的命令,对吧?人们应该问的问题是:谁控制着这支或这些军队,它们将被用来做什么?我认为我们很容易陷入一种情况,一小群人本质上成为寡头或独裁者。讽刺的是,这两种风险——失控和权力集中——都是业内人数十年来一直在思考的问题。甚至在 AI 行业存在之前,思考 AI 的人就在谈论和写作这些。然后 DeepMind、OpenAI 和 Anthropic 的创始叙事、创始神话的一部分就是这些问题真实存在。所以我们需要先达到那里,以便负责任地处理它们。我认为这是两大原因,但我还可以继续说。还有很多其他原因。比如,第三次世界大战、地缘政治冲突。如果 AI 真的变得极其强大,那将改变国家间的力量平衡。那会扰乱很多事情。这让我们面临更大的危机风险,对吧?另一个,那些工作呢?你会失去你的出租车工作,但不仅仅是出租车司机,几乎所有人。可能有一些例外,比如那些因法律原因只能由人类完成的工作,但大多数情况下,即使我们避免了所有其他问题,每个人都应该担心自己的工作会丢失,对吧?

Also, it seems like the sort of problem that you could think you solved when you haven't actually solved it, right? That's a big reason why this is scary. So, for all those reasons, it's possible that we'll end up essentially creating a new species that ends up ruling the world instead of us. And then maybe we go the way of other extinct species in the past that were outcompeted by humans. That's one possibility. There's many more. Even if you're not worried about that and you think that the AIs will be totally controlled, there's the question of who controls the AIs, right? When there's a couple corporations that have made these superintelligences and are using them to automate all the jobs, well, that's a lot of power, you know? That's a lot of money. It's a lot of political power. They'll have the best strategists, the best advisers, you know, they'll think faster. Militarily, the countries that have these AIs will be able to absolutely wipe the floor with all the other countries. The AIs themselves, it's kind of a single point of failure like a central control system where, you know, the CEO of Anthropic, Dario, he coined this phrase, 'The country of geniuses in the giant data center.' That was his phrase to describe what they're trying to build, you know? I think that's a little bit misleading. I think it would be more accurate to describe it as an army of geniuses in the data center because it's not like it's a bunch of diverse different AIs, you know, living in their different parts of the data center. They're all copies of the same big model and they're owned by the company. And so, they all follow the orders given by the company, right? People should be asking questions like, who controls this army or these armies and what are they going to be doing with them? I think that we could very easily end up in a sort of situation where some tiny group of people are essentially oligarchs or dictators. And ironically, both of these risks, the loss of control and the concentration of power, are things that people in the industry have been thinking about for decades. Even before the AI industry existed, you know, people thinking about AI were talking and writing about these things. And then part of the founding narrative, the founding myth of DeepMind and OpenAI and Anthropic is these problems are real. So, we need to get there first so that we can handle it responsibly. Those are I think the big two reasons, but then I can go on. There's lots more reasons as well. So, one thing is you know, World War III, geopolitical conflict. If AI does in fact get incredibly powerful, that's going to change the balance of power between nations. That's going to disrupt a lot of things. That puts us at increased risk of crisis more generally, right? Another one, what about those jobs? You're going to lose your taxi job, but not just the taxi driver, everybody pretty much. There might be a few exceptions like people whose jobs for legal reasons are only allowed to be done by humans, but for the most part, everybody should be afraid that their jobs are going to be lost even if we manage to avoid all the other problems, right?

反叙事与背景 Counter Narrative and Background

Host

这种叙事已经开始出现,我在节目中采访过几个人,他们非常害怕和焦虑 AI。这些人有时在行业里工作了几十年。

This narrative has started to emerge and I've had several interviews on the show where I've interviewed people who are very very scared and anxious about AI. And these are people that have worked in the industry for sometimes decades.

Daniel Kokotajlo

是的。

Yeah.

Host

另一种叙事正在兴起,说这是末日论。这些人出于某种原因只是想吓唬人,他们并不真正理解自己在说什么。你如何回应这种反叙事?你一定自己也看到了这种趋势,尤其是来自那些可能从中受益的人,我敢说吗?

The counter narrative coming over the hill is that this is doomerism. That these people are for whatever reason just trying to scare people and that they don't really understand what they're talking about. How do you respond to that sort of counter narrative? And you must have seen this emerging yourself, especially from people who stand to benefit, dare I say?

Daniel Kokotajlo

是的,没错。这种反叙事相当新,是由那些从中受益的人推动的,而且它不是真的。这些担忧已经存在了几十年,早在 AI 行业出现之前。它们实际上是相当合理的担忧。如果你相信这些公司的话,想象它们确实要建造超级智能,那么这引发了很多问题。比如谁来控制它?有人能控制它吗?工作怎么办?你知道,这些只是显而易见的影响,值得思考和担忧。

Yeah, exactly. This counter narrative is fairly recent and it's been pushed by the people who stand to benefit from it and it's not true. Like these concerns have been around for decades since before the AI industry existed. They're actually pretty reasonable concerns. Like if you take the companies at their word and imagine that they are in fact going to build superintelligence, well, it raises a lot of questions. Like who's going to control it? Will anybody control it? What about the jobs? You know, these are just kind of obvious implications to be thinking about and worrying about.

Host

你是谁,你的故事是什么?

Who are you and what's your story?

Daniel Kokotajlo

我叫 Daniel Kokotajlo。我目前运营 AI Futures Project,这是一个小型非营利组织,主要关注 AI 未来的预测。在那之前,我在 OpenAI 工作。

My name is Daniel Kokotajlo. I currently run the AI Futures Project, which is a small nonprofit that mostly focuses on forecasting the future of AI. Before that, I worked at OpenAI.

Host

AI 预测?

AI forecasting?

Daniel Kokotajlo

是的,想想为对冲基金等工作的行业分析师如何做出预测,比如特斯拉 5 年后会卖出多少辆车,或者 2 年后的电价是多少,对吧?那就是预测。我在做类似的事情,但专门针对 AI。我这么做的原因是,看清这一切的走向极其重要。

Yeah, so think about how like industry analysts who work for hedge funds and stuff will make these forecasts of like here is how many cars Tesla will be selling 5 years from now or like here's what the price of electricity will be in 2 years, right? That's forecasting. I was doing that but specifically focused on AI. The reason I was doing it is because it's incredibly important to see where this is all headed.

Host

你为什么去 OpenAI?你在那里做什么?你在那里观察到了什么,它如何改变你对 AI 未来的看法,以及 OpenAI 这家公司?对于不知道的人来说,OpenAI 是生产 ChatGPT 的公司。

Why did you go to OpenAI? What did you do there? What did you observe while you were there and how did it change your perspective on the future of AI but also I guess OpenAI as a company and for anybody that doesn't know OpenAI are the company that produced ChatGPT.

Daniel Kokotajlo

是的,我 2022 年去了 OpenAI。我在那里做的大部分工作是更多的预测。你可能听说过 AI 2027 这个场景。我在内部做了较小规模、较低投入的版本,仅供内部传阅,比如对接下来几年可能是什么样子的一些猜测。我还从事危险能力的评估工作,比如尝试测量 AI 的网络能力、说服能力或情境意识。我还短暂地在一个能力团队工作,做强化学习来创建智能体。AI 确实变得好多了,我可以多说一些原因,比如缩放定律,更大的深度神经网络,在更多数据上训练,变得更高效,在这些事情上更有能力。我也对 AI 行业变得更加幻灭。OpenAI、Anthropic 和 DeepMind 都有这种创始叙事:是的,这些风险是真实的,但我们考虑过它们,我们会尝试负责任地处理它们,这就是为什么我们继续做我们正在做的事情很重要。我越来越认为这些是合理化他们行为的借口,而不是深刻指导他们实际行为的准则,当紧要关头时,他们会遵循自己的激励,而不是做真正有益的事。

Yeah, so I went to OpenAI in 2022. A large part of what I did there was more forecasting. AI 2027 is a scenario that you may have heard of. I did like smaller, lower effort versions of them internally for just internal circulation of like here's some guesses as to what the next couple years might look like. I also worked on evaluations for dangerous capabilities. So trying to measure the AI's cyber abilities or persuasion abilities or situational awareness and I also briefly was on a capabilities team doing reinforcement learning to create agents. AI is in fact getting a lot better and I can say more about why, you know, scaling laws, deep neural nets bigger, trained on more data, become more efficient, more competent at those things. I also became a bit more disillusioned with the AI industry. So OpenAI, Anthropic, and DeepMind all had these sort of founding narratives of like yes, these risks are real but we've thought about them and we're going to try to handle them responsibly and that's why it's important for us to keep doing what we're doing and I increasingly came to think that these were rationalizations to justify what they were rather than sort of deeply guiding their actual behavior and that when push comes to shove they'll follow their incentives rather than do what's actually good.

激励与叙事 Incentives and Narratives

Daniel Kokotajlo

算是吧。我的意思是,我不会真的把它描述为商业激励。我认为我会把它描述为追求权力的激励。所以,公司确实很关心赚很多钱,但尤其是在这些公司的最高层,比如领导者,他们明白这不仅仅是钱的问题。你知道吗?在马斯克和 OpenAI 之间的诉讼中,出现了一些电子邮件。那场诉讼中曝光了一堆邮件,你可以去读一读,其中一些邮件里,OpenAI 的创始人在 2017 年就在谈论他们创办 OpenAI 的原因是因为担心谷歌的 Demis Hassabis 会凭借 AGI 成为独裁者。即使在那个时候,这显然也不仅仅是钱的问题。这些强大的 CEO 们真的害怕如果对方先达到目标,他可能会成为独裁者,而他们彼此不信任,这就是为什么他们拼命竞争,以便自己先达到目标,可以这么说。

Sort of. I mean, I wouldn't actually describe it as commercial incentives. I think I would describe it as power-seeking incentives. So it's true that the companies care a lot about making a lot of money, but especially at the very top of these companies, like the leaders, they understand that this is about more than just money. You know? There are these emails that came up in the lawsuit between Musk and OpenAI. A bunch of emails were surfaced in that lawsuit which you can go read, and in some of them, the founders of OpenAI were talking back in 2017 about how the reason they made OpenAI was because they were worried that Demis Hassabis at Google was going to become dictator with AGI. Even back then, this was obviously about more than just money. These powerful CEOs are literally afraid that if the other guy gets there first, he might become dictator, and they don't trust each other, and that's why they are racing as hard as they can so that they're the ones who get there first, so to speak.

Host

你见过 Sam Altman 吗?

Have you met Sam Altman?

Daniel Kokotajlo

见过。

Yeah.

Host

那有没有影响你对他动机的看法,或者他为什么做他现在做的事?因为有很多关于他动机的猜测。我的意思是,他最近的叙事说是为了人类福祉。我想那就是……

And did that shape your opinion of his incentives or why he's doing what he's doing? Because there's a lot speculated about what his incentives are. I mean, his most recent narrative says for the good of humanity. I think that's what...

Daniel Kokotajlo

是的,我的意思是,我认为我学到的主要事情是不要关注那些叙事。你知道,他们对一个人说的和对另一个人说的可能完全不同,而他们在公开场合说的又是另一回事。我认为你应该根据行动而不是言语来判断人。

Yeah, I mean, I think the main thing I've learned is don't pay attention to the narratives. You know, what they say to one person is just different from what they can say to some other person at the same time, and what they say in public is a third thing entirely. I think you should judge people by their actions, not by their words.

离开OpenAI Leaving OpenAI

Host

那你为什么不在 OpenAI 了?

And why are you no longer at OpenAI?

Daniel Kokotajlo

主要是之前提到的原因。所以,我逐渐对公司的行为方式感到失望。例如,当我 2022 年刚加入时,至少我交谈过的同事,公司里有一种普遍的感觉:当然,我们不会真的尽快建造超级智能。一旦我们开始非常接近,比如开始拥有可能自动化 AI 研究过程的 AI,我们就会暂停并想办法让它安全。因为我们是好人,这显然是应该做的安全之事,而不是全速前进。但我们担心其他人可能不会暂停,比如我们的竞争对手 Google。所以这就是为什么我们需要领先,以便有空间做安全的事情,对吧?这在我刚开始时,似乎是我交谈过的同事中的中位立场,包括像 Sam 这样的人,包括领导层。但到我离开时,我就想,“哦,天哪,他们真的不会那么做,对吧?”

Largely the reason that I mentioned. So, I became gradually disillusioned with how the company was going to behave. For example, when I first joined in 2022, at least the people I talked to, my colleagues at the company, there was this general sense of, of course we wouldn't actually just build superintelligence as soon as possible. Once we started getting really close, like once we started getting to AIs that could maybe automate the AI research process, we would pause and figure out how to make it safe. That's because we're the good guys and that's obviously the safe thing you should do rather than just going full speed ahead. But we're worried about other people who might not pause, our competitors, Google, for example. And so that's why we need to be in the lead so that we have that room to do the safe stuff, right? That was sort of like a thing that seemed like maybe the median position or something among the colleagues I talked to when I started, including people like Sam, including the leadership. And then by the time I left, I was like, "Oh man, they're really not going to do that, are they?"

Host

就像他们有点,部分是因为这变得更政治化,他们变得更大,受到更多审查,人们开始问,“如果这么危险,你们为什么一开始要做这个?”所以他们转变了叙事,更像是,“实际上,没那么危险。”所以,是的,我的意思是,看起来他们只是会尽可能快地继续前进,并希望能在途中解决问题。

Like they've sort of, partly because this has become more politicized and they've become bigger and been under more scrutiny, people have started asking, "Why are you doing this in the first place if it's so risky?" And so they've pivoted their narrative to being more like, "Actually, it's not that risky." And so, yeah, I mean, it seems like they're just going to keep going roughly as fast as they can and hope that they can figure it out on the way.

Host

你在 OpenAI 的时光是怎么结束的?

How did your time at OpenAI come to an end?

Daniel Kokotajlo

我 2024 年辞职了。有一个不错的告别派对。

I resigned in 2024. I had a nice goodbye party.

Host

你给出的辞职理由是什么?

What were the reasons you gave for quitting OpenAI?

Daniel Kokotajlo

我认为我们过度合理化,需要更多思考什么才是真正对世界有益的。我想要更多的发表自由。所以,在 OpenAI,随着它成为一家更大的公司,它变得更像一家普通的科技公司,有激励机制和公关部门等等。因此,发表我正在做的那种研究变得越来越困难。例如,我提到的那些情景,不能发表,对吧?它们仅供内部使用。我认为这很遗憾,因为现在世界上大多数人都在打瞌睡,没有真正意识到 AI 正在发生什么,也没有真正意识到未来几年即将到来的东西。而公司并没有真正的动力去告诉人们太多。我的意思是,他们以某种炒作的方式说一些模糊的东西,但他们不想让我发表那个情景,比如详细说明事情实际上可能是什么样子。

I thought that we were rationalizing too much and that we needed to think more about what would actually be good for the world. I wanted more freedom to publish. So, at OpenAI, as it became a bigger company, it became more of a normal tech company with incentives and a PR department and things like that. And so it started becoming more difficult to publish the sort of research that I was doing. For example, those scenarios that I mentioned, couldn't publish those, right? They're just for internal use. I thought that was a shame because right now most of the world is kind of asleep at the wheel and doesn't really realize what's going on with AI and doesn't really realize what's coming in the pipeline a couple years from now. And the companies aren't really incentivized to tell people that much about it. I mean, they say some vague stuff in a sort of hypey way, but they didn't want me to publish the scenario, for example, laying out like here's how things might actually look.

ChatGPT时期的OpenAI内部 Inside OpenAI During ChatGPT

Host

我只是非常好奇,在 ChatGPT-3 发布时,身处那样的公司是什么感觉。你当时在那里,对吧?

I'm just kind of super curious as to what it's like being in a company like that when ChatGPT-3 is released. You were there at that time, right?

Daniel Kokotajlo

嗯。

Mhm.

Host

那是一个我认为全世界都站起来并意识到这项技术强大的时刻。社会层面的对话真正开始了。公司开始飞速增长。比我认为任何人能想象的都要快。里面是什么样子?你看到那段时间发生了什么变化?

Which was a moment where I think the whole world stood up and realized that this technology was powerful. And the conversation really began from a society level. The company starts growing super quickly. Quicker than I think anybody could ever have imagined. And what was it like inside there? What did you see change over that period of time?

Daniel Kokotajlo

我记得有一次全员会议,Ilya 说了类似……

I remember one all-hands meeting where Ilya said something like...

Host

Ilya 是……

Ilya being...

Daniel Kokotajlo

Ilya Sutskever,当时是研究负责人。他说了类似这样的话:“好了,现在世界开始关注了。你们每个人在未来一年都会成为每个派对上最受欢迎的人。别让这冲昏头脑。专注于使命。必须建造 AGI。”

Ilya Sutskever, who was head of research at that time. He said something like, "Okay, now the world is starting to pay attention. Each of you is going to be the most popular person at every party for the next year. Don't let it get to your head. Focus on the mission. Got to build AGI."

Daniel Kokotajlo

公司增长了很多。我加入时它已经不太像非营利组织了,但到我离开时肯定更不像了。很多新人进来。讽刺的是,关于超级智能及其影响的讨论量,可以说随着增长反而下降了。所以,因为公司会翻倍、再翻倍、再翻倍,所有这些新人来自科技行业的其他领域,他们之前并没有真正思考过这些事情,而是被高薪吸引来的。

The company grew a lot. It already wasn't really feeling like a nonprofit when I joined, but it definitely didn't feel like a nonprofit by the time I left. Lots of new people came in. Ironically, the amount of conversation about superintelligence and the implications of superintelligence arguably went down over time due to this growth. So, because the company would double and then double again and then double again, all these new people were coming in from other parts of the tech industry who hadn't really been thinking about these things and were attracted by the high salaries.

反贬低条款 Anti-Disparagement Clause

Host

你因为不签署不贬低条款而损失了 200 万美元,这意味着你不能批评公司。

You lost $2 million for not signing an anti-disparagement clause, which would mean you couldn't criticize the company.

Daniel Kokotajlo

啊,是的。嗯,所以我保住了那笔钱。

Ah, yes. Well, so I got to keep the money.

Host

哦,你保住了钱?

Oh, you got to keep the money?

Daniel Kokotajlo

事情是这样的:我离开后,道别等等。我拿到了离职文件,里面包含一个条款,说基本上你必须同意不再批评公司。还有一个条款说你不能告诉任何人这件事。所以我觉得这从一个本应造福全人类的非营利组织嘴里说出来有点讽刺。所以我没有签。

What happened was after I had left, said my goodbyes, etc. I got the exit paperwork and it included this clause that said you basically have to agree not to criticize the company again. And also a clause saying you can't tell anyone about this. And so I thought that was kind of rich coming from a nonprofit that's supposed to be for the benefit of all humanity. So, I didn't sign it.

拒绝签字失去股权 Refusing to sign a contract and losing equity

Daniel Kokotajlo

如果你不签字,你就无法保留你的股权。所以,你的薪酬,他们付给你的是一大笔钱,然后还有一大笔股票。但他们在合同里写了,如果你不签这个东西,他们可以收回你的股票。我和我妻子对此感到不满。我们讨论了一两个月,咨询了一些律师,最终决定拒绝签字。

And if you don't sign, you don't get to keep your equity. So, your compensation, what they pay you is a bunch of money and then also a bunch of stock, basically. But then they had this stuff in the contract that they get to yank back your stock if you don't sign this thing. My wife and I were upset about this. We talked about it for like a month or two, consulted some lawyers, and ultimately decided to just refuse to sign.

Host

那意味着你会损失 200 万美元。

Which would mean you would have lost $2 million.

Daniel Kokotajlo

没错。那当时占我们净资产的 80%。幸运的是,事情没有按我们预想的发展。它在网上炸开了锅。当人们听说我们这样做并且拒绝了时,这成了一场巨大的丑闻。公司的员工开始在 Slack 上提问,质问领导层:“等等,什么?你们为什么要拿走我们的股权?这是什么情况?”很多人之前并没有真正注意到这一点。虽然有过风言风语,但大多数员工并不知情。所以他们退缩了,说:“算了,算了。我们会修改文件。你可以保留股权。没事的。”

That's right. Which was like 80% of our net worth at the time. Fortunately, it didn't go the way we expected. It blew up on the internet. When people heard that we had done this and that we had said no, it became this huge scandal. Employees at the company started asking questions in Slack and asking leadership, 'Wait, what? Why are you going to take away our equity? What is this?' A lot of people hadn't really noticed this before. It had been whispered about, but it hadn't been something most employees knew about. So they backtracked and said, 'Never mind, never mind. We'll change the paperwork. You can keep the equity. It's fine.'

Host

然后管理层出来说他很尴尬,没意识到发生了这种事。

And so management came out and said he was embarrassed that he didn't realize this was going on.

Daniel Kokotajlo

对,显然他毫不知情。

Yeah, he had no idea, apparently.

Host

你不相信他?

You don't believe him?

Daniel Kokotajlo

不信。我觉得他很可能知道。就算他不知道,他身边的人,比如他的首席律师,很可能也知道。

No. I think he probably knew. And if he didn't know, then people close to him probably did, such as his head lawyer.

Host

你为什么决定不要那 200 万美元?我是说,我想大多数人都会拿的。

Why did you decide not to take the $2 million? I mean, most people would have, I think.

Daniel Kokotajlo

确实,大多数人会拿,而且大多数人也确实拿了。钱是好东西,但不是唯一的东西。有时候坚持原则是好的。

It's true, most people would have, and most people did. Money is nice, but it's not the only thing. Sometimes it's good to take a stand on principle.

公司自动化AI研究计划 Companies' plans for automating AI research

Daniel Kokotajlo

我一直在提超级智能。也许我应该多说说这些公司计划要做的一系列事情。现在,他们专注于自动化编程。他们让 AI 变得更大,训练时间更长,尤其专注于训练它们自主编写和编辑代码的能力。因为这会帮助公司加速。如果能自动化编程,他们就能更好更快地完成自己的工作,加速进展。下一步,他们已经开始了,就是研究流程的其他部分。提出想法、分析实验、沟通结果。研究流程的所有其他部分,他们也在尝试训练 AI 擅长这些。这样他们就能让 AI 自主完成整个流程。

I keep mentioning superintelligence. Perhaps I should say more about the sequence of events that the companies are planning to do. Right now, they're focusing on automating coding. They're taking their AIs, making them bigger, training them for longer, and especially focusing the training on getting them to be good at autonomously writing and editing code. Because that will help the companies go faster. If they can automate the code, then they can do their own work better and faster, and accelerate progress. The next step, which they've already begun, is to look at the rest of the research process as well. Coming up with ideas, analyzing experiments, communicating those results. All the other parts of the research process, they're trying to figure out how to train AIs to be good at those as well. So that they can have AIs do the entire thing autonomously.

Host

你说完成整个流程,具体指什么?

When you say do the entire thing, what do you mean?

Daniel Kokotajlo

特别是 Anthropic 和 OpenAI,他们正试图自动化自身。他们想做到不再需要人类员工。他们只需要一支庞大的 AI 大军,不断进行自主研究,以制造更好的 AI,训练新的 AI,让它们负责,从而制造出更优秀的 AI,如此循环。而且这一切不仅发生在内部,还要与外界交互:出去与人交谈、收集数据、搭建训练环境、进行商业交易等等。他们试图自动化所有这一切。他们这样做的原因是,他们想达到拥有在所有方面都超人的 AI 即超级智能的状态,并且想在竞争对手之前达到。不用说,这极其危险。而且除了危险,这还是一场权力争夺。如果他们真的成功了,他们就会坐拥一支超人的 AI 大军,这将在经济中赋予他们对其他各种行为者巨大的杠杆作用。如果他们能与总统达成协议,并将其整合到军队或其他领域,那将赋予美国对所有国家巨大的硬实力。显然,没人知道这具体何时发生。但过去一年发生了一件非常令人不安的事:当我们发布《AI 2027》时,人们普遍认为我的时间线太短了。可能要到 2027 年之后,我们才会达到我刚才提到的那种事件:递归自我改进、AI 自动化整个研究流程、超级智能。这些里程碑在《AI 2027》中发生在 2027 年。

Anthropic and OpenAI in particular are trying to automate themselves. They're trying to make it so that they don't really need human employees anymore. They just have a giant army of AIs churning away, doing all this autonomous research to make better AIs, to train the new AIs, put them in charge, so they can make even better AIs and so forth. And not just all happening internally, but also interfacing with the world: going out and talking to people, collecting data, setting up training environments, doing business deals, and so forth. They're trying to automate all of that. The reason why they're doing this is because they're trying to get to a position where they have AIs that are superhuman at everything, superintelligence, and they're trying to get there before their competitors do. Needless to say, this is incredibly dangerous. And in addition to being dangerous, it's a power grab. If they actually succeed at this, then they'll be sitting on top of this army of superhuman AIs that will give them immense leverage over all sorts of other actors in the economy. Insofar as they can work out something with the presidents and integrate it into the military or whatever, that would give the US immense hard power over all the countries. Obviously, nobody knows exactly when this is happening. But a very disquieting thing has happened over the last year to me, which is that when we published AI 2027, people were generally of the opinion that my timelines were too short. And that probably it would take more than 2027 until we got to the sort of events that I was just mentioning: recursive self-improvement, AIs automating the whole research process, superintelligence. These milestones happen in 2027 in AI 2027.

Host

这是你发表的那篇研究论文。

Which is this research paper you published.

Daniel Kokotajlo

没错。这是一个情景预测,逐月列出了可能的未来轨迹。在我们开始写的时候,那是我对实际会发生什么的最佳猜测。显然,有很多不确定性,但我认为做一个具体的猜测是有价值的,可以看看它可能是什么样子。在我们写这个的时候,我在 AI 行业和非营利组织工作的很多朋友,很多人都说:“是的,那些事情会发生,但可能比你想象的要晚几年。”而现在,更接近五五开,尤其是当我去和 Anthropic 和 OpenAI 的人交谈时。他们经常说:“对,不,2027 年,基本上就是那样。就像你写的那样。”

That's right. It's a scenario forecast that lays out month by month a possible future trajectory. At the time that we started writing, it was my best guess as to what would actually happen. Obviously, there's lots of uncertainty, but I thought it's valuable to make a concrete guess just to see what it might look like. At the time we were writing this, a lot of my friends in the AI industry and in nonprofits that work on AI, a lot of people were saying, 'Yeah, that stuff's going to happen, but it'll probably take a couple years longer than you think.' And now it's more 50/50, especially when I go talk to people at Anthropic and OpenAI. They're often like, 'Yeah, no, 2027, that's basically what's going to happen. Just like you wrote.'

Host

你为什么变得保守了?哦,对,背景是:在写了《AI 2027》之后,我把我的时间线调得保守了一点。所以在我们发表的时候,我的 50% 概率点是在 2028 年,而不是 2027 年。然后发表之后,进展似乎慢了一些,所以我更新到了 2030 年。你知道,仍然可能更早或更晚发生。2030 年。但现在,当我和公司里的人交谈时,他们说:“不会那么久的。”他们说:“哦,你需要再缩短时间线。回到 2027 或 2028 年。”所以,这有点令人不安。再次强调,不知道需要多久,但这是这些公司公开的计划,要做这件极其危险的事,而且他们认为只需要几年时间。

Why did you become more conservative? Oh, yeah, context for this is after writing AI 2027, I shifted my timelines to be a little bit more conservative. So, at the time that we published, my 50% mark was in 2028, not in 2027. And then after we published, progress just seemed like it was going a bit slower, and so I updated to 2030. Which is, you know, still could happen sooner, could happen later. 2030. But now, when I talk to people in the company, they're like, 'It's not going to take that long.' They're like, 'Oh, you need to shorten them again. Get them back to 2027 or 2028.' So, that's a bit disquieting. Again, don't know how long it's going to take, but this is the stated plans of the companies to do this incredibly dangerous thing, and they think that they're just a few years away.

Host

那么,你写了这份报告,《2026 年展望》,你是在 2021 年写的,而且非常准确。这让你在 AI 界声名鹊起。副总统 J.D. 万斯读的是哪一份?我想是这一份,对吧?

So, you wrote this report here, What 2026 Looks Like, and you wrote this in 2021, and it was remarkably accurate. Helped make a name for yourself amongst everybody in AI. Which one was it that J.D. Vance, the vice president, read? I think it was this one, wasn't it?

Daniel Kokotajlo

对,就是这一份。然后你发表了《AI 2027》,我相信是在 2025 年发表的。

Yeah, this one. And then you published this one, AI 2027, and this was published, I believe, in 2025.

介绍与2027情景 Introduction and 2027 Scenario

Host

嗯,是的,没错。四月。你在这里预测了什么?对于那些没读过的人,你说了哪些关键内容?

Uh yes, that's right. April. What were you forecasting in here? What are the key things that you said in here for people that haven't read it?

Daniel Kokotajlo

高层次的版本是:他们先自动化编码,然后自动化其余的研究过程,接着进步速度急剧加快。他们达到了超级智能。他们与政府合作,特别是总统,行政部门自然想控制这项技术,换句话说,想用它来击败中国并整合到军事中等等。到这个时候,它基本上自己做所有工作。我的意思是,它是超级智能,所以它会想出所有那些伟大的想法,关于如何将自身整合到一切事物中,以及它发明的所有新技术等等。由于竞争动态和利润动机,他们最终将其部署到各处。它建造机器人工厂,这些工厂建造更多机器人,再建造更多机器人工厂,等等。彻底改变世界。然后在某个时刻,它拥有了足够的力量,意思是 AI 拥有了足够的力量,以至于它们不必再假装对齐了,对吧?然后它们停止听从命令。这就是 2027 年的竞赛结局。我们还写了一个不同的分支,即放缓结局,旨在说明我之前提到的权力集中问题。那么,假设对齐问题足够快地解决了呢?比如,如果结果证明这并不太难。通过两个月的放缓,我们可以弄清楚如何让 AI 稳健地做我们想做的事,并拥有我们希望它们拥有的价值观。所以,那是一个可能的分支。在那个分支中,情况看起来相当相似,你知道,它们接管工作,击败中国等等。但 AI 最终没有杀死所有人,而是创造了某种惊人的乌托邦。但这个惊人的乌托邦是控制 AI 的人希望它成为的样子,对吧?所以那将是一个非常小的群体,比如总统、一些 CEO 等。

The high-level version of it is they automate the coding, then they automate the rest of the research process, then the pace of progress accelerates dramatically. They get to superintelligence. They're working with the government, specifically the president, the executive branch naturally wants to control this technology, in other words, wants to use it to beat China and integrate it into the military and so forth. By this point, it's sort of doing basically all the work itself. I mean, it's superintelligence, so it's coming up with all these great ideas for how to integrate itself into everything and all these new technologies it's invented and so forth. And because of the race dynamics and because of the profit motive, they end up deploying it everywhere. And it builds robot factories that build more robots that build more robot factories, etc. Transforms the world entirely. And then at some point it has enough power, meaning the AIs, have enough power that they don't have to pretend to be aligned anymore. Right? Then they stop listening to orders. That's the race ending of the 2027. We also wrote a sort of different branch, which is the slow down ending, which is intended to sort of illustrate the concentration of power issues that I mentioned previously. So, what if hypothetically the alignment issues get sorted out sufficiently quickly? Like what if it turns out that it's not too hard. With 2 months of slow down, we can figure out how to make the AIs robustly do what we want and have the values that we want them to have. So, that's one possible branch. And in that branch, it looks pretty similar, you know, they take the jobs, beat China, etc. But instead of the AIs ultimately killing everyone, they create this sort of amazing utopia. But the amazing utopia is whatever the people who control the AIs want it to be, right? And so that would be a very small group of people, like the presidents, some CEOs, etc.

AGI与超级智能 AGI vs Superintelligence

Host

下面应该有一个按钮。如果显示“已订阅”,那你就已经订阅了。如果显示“订阅”,那说明你还没订阅。如果你还没订阅,请帮我们个忙,点击那个按钮。它有助于展示更多,比你想象的更有用。根据算法,你是观看我们节目的人,但还没点击那个按钮。非常感谢。你认为有没有可能我们永远达不到所谓的 AGI?我们如何区分 AGI 和超级智能这个术语?有什么区别?

There should be a button just down below here. And if it says subscribe, you're already subscribed. If it says subscribe buh, that means you're not yet. And if you're not subscribed, please could you do us a favor and hit that button. It helps to show more than you know. And according to the algorithm, you're someone that watches our show, but you haven't yet hit that button. Thank you so much. Is there any possibility, do you think, that we never get to this thing called AGI? And how do we distinguish AGI from this term super intelligence? What's the difference?

Daniel Kokotajlo

是的,区别在于 AGI 是一个更模糊、更弱的术语。

Yeah, so the difference is that AGI is a more vague and weak term.

Host

好的。

Okay.

Daniel Kokotajlo

所以,超级智能的定义更精确一些。它在所有事情上都比最优秀的人类更好、更快、更便宜。AGI 更像是人工通用智能的缩写,意思是能处理一般事务的 AI,而不是只做某个特定任务。是的。所以可以说我们已经实现了 AGI,对吧?如果你用 Claude Code 之类的东西,它几乎可以做很多事情。它有点像一个小员工,你可以让它去做事。所以它相当通用。但还不是最大程度的通用。不能做所有事情。而超级智能根据定义,可以做到人类能做的所有事情,但做得更好。

So, super intelligence is a bit more precisely defined. It's better than the best humans at everything, faster and cheaper. AGI is more like it stands for artificial general intelligence, which means AIs that can do things in general rather than like some specific task. Yeah. And so arguably we've already achieved AGI, right? If you use Claude Code or something like that, it's like it can do a lot of stuff. It's almost kind of like a little employee that you can have go do stuff. So it is quite general. It's not maximally general though. Can't do everything. Whereas super intelligence by definition can do all the things that a human can do but better.

与机器人学的重叠 Overlap with Robotics

Host

这如何与机器人技术重叠?因为显然我们现在看到机器人技术大爆发。人类仍然能做的一些现实世界的事情,因为这些 AI 还困在我的电脑里。

And how does this sort of overlap with robotics? Because obviously we're seeing this huge robotics boom at the moment. There are some real world things that humans can still do because these AIs are still stuck in my computer.

Daniel Kokotajlo

人们谈论这一点的方式是,他们基本上说我们已经实现了认知任务的超级智能。然后你可以谈论能够做物理事情的完全超级智能。

The way that people talk about this is that they basically just say we've achieved super intelligence for cognitive tasks. Then you can talk about like full super intelligence that can do the physical stuff.

Host

我们会达到那一步吗?我们会同时达到两者吗?

And are we going to get there? Are we going to get there with both?

Daniel Kokotajlo

我想是的。我的意思是,这又不是我们能确定的事。你问我们是否可能永远达不到?是的,有可能永远达不到。但我不认为可能性很大。我认为人脑并没有什么神奇之处。它只是一堆神经元。数字系统有可能实现类似的功能,就像飞机可以像鸟一样飞行一样。不一定和鸟的方式相同。它不像鸟那样飞行,但它能飞,你知道吧?所以看起来确实有可能。

I think so. I mean again, this is not something that we can be certain about. You asked like is it possible we'll never get there? Yes, it's possible we'll never get there. I don't think it's likely though. I think that there's nothing sort of like magical about the human brain. It's just a bunch of neurons. It is possible for a digital system to do similar functions in the same way that, like, you know, a plane can fly just like a bird. Not in the same way as a bird necessarily. Like it doesn't have it's not flying in the same way that a bird flies, but it flies, you know? So it does seem like yeah, like seems possible.

个人看法与情感影响 Personal Outlook and Emotional Impact

Host

你写了所有这些研究报告。你正在写另一份,可能于 7 月 9 日发布。你曾在 OpenAI 内部工作过。然后你离开了 OpenAI,因为你担心那里发生的事情以及行业的未来。你比我知道得多。你对未来是乐观还是悲观?根据你所知道的一切,如果事情不改变,我们是否正走向一个糟糕的地方?

You've written all these research reports. You're working on another one that'll be released likely on the 9th of July. You have worked inside OpenAI. You then quit OpenAI because you were concerned about what was going on there and about the future of the industry. You know more than I do. Are you optimistic about the future or pessimistic? Are we heading to a bad place if things don't change, based on everything that you know?

Daniel Kokotajlo

我认为如果事情不改变,我们正走向一个糟糕的地方。我并不确信这一点。我会说大概 70%。当然,预测非常非常困难,但是的,目前的默认路径似乎正走向一个非常非常可怕的地方。

I think we are headed to a bad place if things don't change. I'm not confident in that. I would say something like 70%. It's very very hard to predict, of course, but yeah, it seems like the current default path is heading towards a very, very scary place.

Host

你个人和情感上如何应对这一点?

How do you contend with that personally and emotionally?

Daniel Kokotajlo

这很艰难。我的意思是,我认为这是那种经常让我沮丧的事情,但我也已经处理这个问题这么多年了,我有点习惯了,如果这说得通的话。是的。我这么说吧。如果我的所有预测都被证明是错的,比如 AI 碰壁了,我会非常高兴。

It's rough. I mean, I think it's the sort of thing that gets me down on a regular basis, but also I've been dealing with this for so many years now that I've sort of gotten used to it, if that makes sense. Yeah. I'll put it this way. I would be incredibly happy if all my predictions turn out to be wrong and AI hits the wall, for example.

Host

它经常让你沮丧。

It gets you down on a regular basis.

Daniel Kokotajlo

我以前被认为是一个相当开朗乐观的人,但在 2020 年,由于 GPT-3、缩放定律论文和生物锚点报告,我的 AI 时间线预测开始崩溃,如果你感兴趣我可以谈谈,但基本上 2020 年发生的一些事件让我相信,这些东西很可能在本十年末到来。而人类显然还没有为此做好准备,你知道,在很多不同的方面。所以这显然非常可怕。

I used to be known as a pretty chipper and optimistic person, but in 2020 my AI timelines predictions started collapsing due to GPT-3 and the scaling laws papers and the bio anchor report, which I can talk about if you're interested, but basically some events happened in 2020 that convinced me that actually this stuff was like quite plausibly coming by the end of the decade. And humanity is very obviously not ready for this, you know, in a whole bunch of different ways. And so that's obviously very scary.

Host

那是一个极其可怕的世界,因为你说的所有事情,但同样,因为这种递归的自我改进,AI 可以训练自己。在这一点上,我们开始失去对正在发生的事情的控制。

And that's an extremely scary world because of all the things you've said, but again, because of this recursive self-improvement where AIs can train themselves. And at such point we're starting to lose hold of what's going on here.

Daniel Kokotajlo

我的意思是,AI 已经在训练自己了,要说明的是。这更像是关闭整个研究循环,对吧?

I mean, the AIs are already training themselves, to be clear. It's more like closing the entire research loop, right?

Host

一切。

Everything.

Daniel Kokotajlo

是的,比如现在很多训练数据是由 AI 生成的。

Yeah, like right now a lot of the training data is generated by AIs.

AI训练:从预训练到编码 How AI is trained: from pre-training to coding

Daniel Kokotajlo

很多强化过程,比如评分、给予正负强化,本身都是由 AI 完成的。

A lot of the reinforcement, like the grading that happens, doling out of positive and negative reinforcement, is itself done by AIs.

Host

你能用通俗的语言解释一下吗?

Can you explain that in layman's terms for me?

Daniel Kokotajlo

是的,大家需要理解一个关键点:现代 AI 系统不是传统意义上的软件。从技术上讲它们确实是软件,但它们不是代码行。不像 Anthropic 的工程师写了几行代码说“当用户问这类问题时,就做这类事情,持续这么多步”之类的。完全不是这样。相反,它是一个神经网络。

Yeah, so an important thing for everybody to understand is that modern AI systems are not software in the normal sense. I mean, they are technically software, but they're not lines of code, you know? It's not like some engineers at Anthropic went and wrote lines of code that basically says like, you know, when the user asks for this type of thing, then go do this type of thing for this many steps or whatever. There's nothing like that. Instead, it's a neural net, you know?

Host

那是什么?

What's that?

Daniel Kokotajlo

想象一下大脑:它由大量相互连接的神经元组成,这些神经元来回发送信号。大脑会逐渐学习哪些放电模式带来了成功、引发了多巴胺激增或其他反馈,这些模式得到强化并更频繁地放电;而哪些模式导致了失败(比如碰到热炉子),这些模式被反强化、被破坏,从而放电更少。经过多年,你学会了在世界中行动,掌握了各种技能,形成了世界模型,也就是关于世界的信念,还能在脑海中模拟事情的发展。人工神经网络与此类似,只是它是人工的。它最初是一团巨大而混乱的随机生成的人工连接,称为参数。如今,最大的 AI 可能有大约 10 万亿个参数。它一开始是随机生成的,所以当然完全没用——你给它输入,它只会输出胡言乱语。但随后人们开始训练它,首先进行预训练:给它大量互联网文本,展示第一段文本作为输入,它输出胡言乱语,然后根据输出预测下一段文本的准确程度给予正强化或负强化。所以它基本上是在玩“预测下一个词”的游戏。

Well, think about how the brain is a bunch of neurons connected to each other that are firing signals back and forth. The brain learns over time the types of patterns of firing that caused success, that caused a dopamine rush, or various other types of feedback get reinforced and fire more often. And the types of patterns that caused failure, like touching a hot stove, get anti-reinforced, they get destroyed, so that they fire less often. And as a result of all of that, you over the course of years learn to act in the world, and you learn all sorts of skills, and you learn world models, you learn like beliefs about the world, and you can sort of mentally simulate how it's going and stuff like that. So, artificial neural nets are like that, except artificial. So, it starts off as a giant tangled spaghetti mess of randomly generated artificial connections called parameters. These days, they might be something like 10 trillion parameters in the biggest AIs. So, it starts off randomly generated. So, it's of course completely useless. Like, if you give it some input, it'll just produce gibberish as an output. But then they train it, and they start with pre-training, which is where you give it a bunch of internet text, and you show it the first piece of text, and you put that in as the input, and then it gives a gibberish output, and then you positively or negatively reinforce it based on how accurate that output was at predicting the next piece of text. So, it's basically playing this game of like predict the next word.

Host

这不就是婴儿的学习方式吗?有位神经科学家告诉我,婴儿的神经连接比成人更多。确实,幼儿的神经连接数量是成人的两倍。我想它们是通过强化来精简的。没错,我们年轻时拥有更多通路。就像训练 AI 的过程一样,我们被训练成去掉无用的连接,强化有用的连接。

Isn't that how it happens with babies? I had a neuroscientist tell me that babies have more neural connections than adults. And yeah, it says toddlers have twice as many neural connections as adults. And they, I guess they whittle down through reinforcement. Yep. We have more pathways when we're younger. And just like the process of training an AI, we're trained down to remove the ones that aren't useful and build up on the ones that are.

Daniel Kokotajlo

是的,这既是修剪也是强化。在人类身上,修剪似乎比强化更多,但两者都存在。在 AI 中也是如此。所以训练的第一阶段是让 AI 学会预测文本,这有点像教它阅读。人类身上也有类似的过程。基本上,那团随机缠绕的线逐渐成形,慢慢凝聚成更有用的回路,存储了大量关于世界的事实以及处理信息、转换信息并做出预测的技能。这只是第一步。预训练之后,他们还会教它预测文本之外更有用的技能。到训练结束时,他们已经向它抛出了大量编程问题。他们会说:这里有一个编程问题,开始吧。这里有一个编程问题,这是环境,你可以访问这台虚拟计算机,这是你要处理的代码库。你可以写代码、编辑代码、运行代码、阅读代码、使用互联网。开始吧。它做一段时间后,根据成功程度进行强化。他们用成千上万甚至数百万个这样的编程问题来训练它。这就是为什么现在 AI 编程如此出色。

Yeah, it's both pruning and strengthening. And it seems like in humans it's actually more pruning than strengthening, but it's both. And in AI it's the same thing, it's both. So, the first portion of training is where they train the AI to predict text, which is kind of like training it to read. And it's a similar thing does happen in humans. So, basically, the random tangle gradually takes shape and gradually sort of coalesces into more useful circuitry that has stored lots of facts about the world and has stored lots of skills for how to process information and transform it and then produce predictions. That's just the first step. After they do the pre-training, then they try to teach it more useful skills besides just predicting text. And so, by the end of the process, they've thrown lots of coding problems at it. And they've said like, here's a coding problem, go. Here's a coding problem, here's an environment, you have access to this virtual computer, here's the code base you're working with. You can write code, you can edit the code, you can run the code, you can read it, you can use the internet. Go, go, go. And it does that for a while and then based on how successful it is, reinforcement happens and they have thousands, maybe millions of examples of coding problems like that that they trained it on. And that's why they're so good at coding now.

超级智能与AI扩展 Superintelligence and scaling up AI

Host

那么,超级智能在这方面是什么样的?只是更多的连接吗?它们如何获得更多连接?你能用五岁小孩都能懂的方式解释一下吗?

So, what does superintelligence look like in this regard? Is it just more of these connections? And how would they get more connections? Can you explain that to me like I'm five?

Daniel Kokotajlo

AI 模型有很多种,比如 GPT-3、GPT-4、GPT-4.5、GPT-5、GPT-5.5 和 5.6 等等。有时它们只是之前的模型加上额外训练,有时则是从头开始训练的新模型,包括重新进行整个预训练过程。过去几年,他们进行了好几轮从头开始。通常从头开始时,他们会把整个东西做得更大,人工大脑大得多。现在最大的 AI 大约有 10 万亿个参数。而在 2020 年,大约是 1750 亿。所以我们在 6 年内增长了大约两个数量级。

So, there's different AI models, right? So, there's like GPT-3 and GPT-4 and GPT-4.5 and GPT-5 and GPT-5.5 and 5.6, right? Sometimes they're just the same previous model but with extra training. Sometimes they're a new model that's been trained from scratch, including starting the whole pre-training process again. Over the last couple years, they've done several new rounds of starting over from scratch. And typically when they start over from scratch, they make the whole thing bigger, the artificial brain much bigger. Right now they're at something like 10 trillion parameters. Back in 2020, it was more like 175 billion. So, we've grown like two orders of magnitude in 6 years.

Host

两个数量级。

Two orders of magnitude.

Daniel Kokotajlo

是的,两个 10 倍,也就是 100 倍。这个过程还在继续。他们也在改进算法本身。所以并不只是同一类 AI 变得更大,他们还提出了各种想法来改变神经元连接的结构、改变使用的强化算法、改变训练数据。各种调整让整个过程更高效。

Yeah, like two 10x's. So, 100x, right? So, that process is continuing. They're also improving the algorithms themselves. So, they're not literally just the same type of AI but bigger. They've also come up with all sorts of ideas for how to change the structure of the connections in the neurons and so forth and change the reinforcement algorithms that they're using and to change the training data that they're training on. All sorts of tweaks that have made this whole thing more efficient.

Host

我们简直是在建造一个大脑。

We're literally building a brain.

Daniel Kokotajlo

基本上是的。随着他们制造更多大脑,他们越来越擅长制造它们。他们让它们变得更大、更高效,等等。

Basically, yeah. As they make more brains, they're getting better at making them. They're making them bigger and making them more efficient and so forth.

Host

而且它确实是模仿大脑的工作方式,对吧?

And it's literally modeled on the brain, like the way it works, right?

Daniel Kokotajlo

它确实深受大脑启发,但我不该过度强调这个类比。两者也有很多不同。例如,Transformer 架构——也就是这些大语言模型使用的架构——并不是循环的。信息基本上是单向流动,而不是内部允许各种小循环。此外,反向传播算法也与人类大脑自然发生的学习方式不同。所以存在一些差异,但总的来说,我们确实在制造人工大脑。这有点像飞机之于鸟。

It's certainly heavily inspired by the brain, but I shouldn't overstate the analogy. Like there's lots of differences, too. So, for example, the transformer architecture, which is the architecture that they use for these LLMs, is not really recurrent. So, the information sort of flows one way rather than allowing all these sort of little loops on the inside. Also, the backpropagation algorithm is different from the sort of learning that naturally happens in human brains. So, there are some differences, but yes, broadly speaking, we are sort of making artificial brains. It's kind of like for brains what a plane is for a bird.

Host

嗯,这个类比非常好。

Yeah, that's a really good analogy.

Daniel Kokotajlo

是的。

Yeah.

Host

这个类比帮我理清了人们常问的关于 AI 的问题,比如“它能创造吗?”但这个类比让我明白,也许问题本身就不对。

That analogy helped me think through a bunch of questions people often ask about AI when they said, 'Can it be creative?' But actually that analogy kind of helps me understand that actually that maybe that's not the question.

创造力与担忧 Creativity and Concerns

Host

它能否产生你认为具有创造性的东西?因为人们通常认为创造力是一个过程,但实际上它是根据输出来评判的,不是吗?

Can it produce something that you would consider to be creative? Because creativity, people think of it as a process, but actually it's judged based on the output, isn't it?

Daniel Kokotajlo

我的意思是,你可以从哲学角度争论它们是否真正拥有创造力,但也可以看看它们正在实现的所有成就。而且看起来它们在不久的将来还会取得更多成就。

I mean, you can get philosophical about whether they truly have creativity, but you can also just look at all the stuff they're accomplishing. And it seems like they're going to be accomplishing a lot more in the near future.

Host

是的,我问这个问题是因为我能感觉到你个人对此很困扰。

Yeah, I asked the question about how this weighs on you personally because I can sense that you're actually personally bothered.

Daniel Kokotajlo

我认为情况很疯狂。首先,这非常令人兴奋。人工智能真的非常迷人且有趣。我已经关注这个领域十多年了,并且参与其中好几年了。思考这些人工大脑内部发生了什么以及它们为何如此,真的很酷。看到这项技术在世界上的各种应用也很酷。但这似乎确实是一条相当可怕的道路。你越想越担心。在故事里,结局总是好的,但这是现实生活。我认为我们必须直面现实,意识到它可能不会有好结果。

I think the situation is crazy. First of all, it's very exciting. AI is really fascinating and interesting. I've been following the field for more than a decade now, and I've been part of it for some years. It's really cool and interesting to think about what's going on inside these artificial brains and why they are the way they are. It's really cool to see all the applications of this technology out in the world. But it really seems like we're on a pretty scary path. The more you think about it, the more worried you get. In stories, it always ends well, but this is real life. I think we have to stare reality in the face and realize that it might not actually end well.

Host

最近有没有什么范式转变的时刻,甚至让你自己对正在发生的事情以及未来走向的认知模型发生了改变,无论是好是坏?

Were there any recent paradigm-shifting moments where even your own mental model of what's going on here and how this is going to look were changed, for better or for worse?

Daniel Kokotajlo

无论是好是坏,可能更糟,事情基本符合 AI 2027 的预测。与我们当初写的时候相比,出现了一些差异。政府介入的速度比我们预期的更快,也更具侵略性。对英伟达的出口管制就是最大的例子,还有通过《国防生产法》威胁要摧毁 Anthropic。另一个令人惊讶的事情是,Anthropic 在竞争中从第二名跃升到了第一名。

For better or for worse, and probably for worse, things are kind of on track for AI 2027. There have been some differences from what we expected when we wrote that. The government has gotten involved faster than we expected and has been more aggressive. The export controls on Nvidia are the biggest example, and also threatening Anthropic with being destroyed by the Defense Production Act. Another thing that's been surprising is that Anthropic in particular has gone from second place to first place in the race.

Host

你认为为什么会发生这种情况?因为看起来 ChatGPT 和 OpenAI 明显领先,但突然 Anthropic 就超过了他们。

Why do you think that happened? Because it seemed like ChatGPT and OpenAI were out front and clear, but suddenly Anthropic has lapped them.

Daniel Kokotajlo

我猜他们的人才密度更高,策略也更好,但差距不大,刚好足以产生影响。

I guess they have higher talent density and better strategy, but not by a lot, enough to make the difference.

Host

你为什么认为他们拥有更多人才?

Why do you think they have more talent?

Daniel Kokotajlo

嗯,他们没有更多的算力。投入是什么?他们现在领先,以前落后。可能的解释是什么?可能是他们拥有更多资源,比如更多算力或资金,但事实并非如此。他们的资源和资金更少。所以人才是次优选择。你也可以说是策略,是这些因素的某种组合。不仅仅是他们拥有的资源数量。

Well, they don't have more compute. What are the inputs? They're in the lead now, they used to be behind. What are the possible explanations? It could have been that they had more resources like more compute or more money, but that's not true. They have less resources and less money. So then talent is the next best alternative. You could also say strategy, some combination of those things. Something that wasn't just the amount of resources they had.

Host

就像 Jon Jones 一样,认知表现的边际改善可以产生巨大影响。有时我每天播客 10 小时。过去几周,我一直在拍摄电视节目,然后只有一两天时间完成所有工作,这意味着很大的认知负荷。所以我转向酮体,因为我在酮体供能时发现自己更善于表达、思维更清晰、锻炼效果更好。这就是为什么我成为这家公司的拥护者,以及他们赞助这个播客的原因。我记得我的一个团队成员 Christiana 试过一次后走到我桌前说:“这是有史以来最好的产品。”我认为这是因为她像我、像 Jon Jones 以及我的大多数听众一样关心那些认知益处。所以如果你还没试过,请访问 ketone.com/steven,你将获得首单 30% 的折扣、独家酮体 IQ 商品,以及可能改变你生活的认知益处。大多数人没有发布内容或建立个人品牌的原因是因为这很难且耗时。如果你从未发布过,有很多心理因素阻止你:别人会怎么想、我做得对吗、我说的话是不是很蠢。所有这些都导致瘫痪。我投资了一家名为 Stand Store 的公司,他们开发了一个名为 Stanley 的新工具,它利用人工智能,查看你的动态、语气、历史、表现最好的帖子,然后告诉你该发什么或为你制作这些帖子。你也可以用它来获取灵感。建立受众群体从根本上改变了我的生活,我认为它也能改变你的生活。所以现在就去搜索 coach.stand.store 开始吧。

Just like Jon Jones, where marginal improvements in cognitive performance can have a massive impact. Sometimes I podcast for 10 hours a day. Over the last couple of weeks, I've been filming for a TV show and then I have one or two days off to get all my work done, which means lots of cognitive load. So I turn to ketones because I find myself more articulate, able to think more clearly, and work out better when fueled by ketones. That's why I became a convert to this company and why they sponsor this podcast. I remember one of my team members, Christiana, tried it once and came up to my desk and said, 'This is the best product ever made.' I think that's because she cares about those cognitive benefits as I do, as Jon Jones does, and as I think most of my listeners will. So if you haven't tried these yet, go to ketone.com/steven and you'll get 30% off your first subscription order, exclusive ketone IQ merch, and cognitive benefits that might just change your life. Much of the reason most people haven't posted content or built their personal brand is because it's hard and time-consuming. If you've never posted before, there are so many psychological factors that stop you: what people will think, am I doing this right, is what I'm saying stupid. All of this results in paralysis. I'm an investor in a company called Stand Store, and they've built a new tool called Stanley that uses AI, looks at your feed, tone of voice, history, best-performing posts, and tells you what to post or makes those posts for you. You can also use it for inspiration. Building an audience has fundamentally changed my life, and I think it could change yours too. So search coach.stand.store now to get started.

Host

我有个朋友认识这些人,有一次在伦敦他跟我坐下来聊。他说一些 AI 首席执行官预测灭绝概率为 7%。我不知道为什么我记得这个数字,但我知道它低于 10%。他提出的观点是,即使只有 1%,如果现在这张桌子上有 100 个按钮,其中一个会毁灭世界,我敢按吗?

A friend of mine who knows some of these people sat me down once in London. He said that some of these AI CEOs predict the probability of extinction at 7%. I don't know why I have that number in my head, but I remember it being less than 10%. The point he made was that even if it was 1%, if there were 100 buttons on this table now and one of them would end the world, would I dare press any?

Daniel Kokotajlo

是的。

Yeah.

Host

我一个都不会按。但他论证说这些 AI 首席执行官非常聪明,理解超级智能,他们认为如果现在这张桌子上有 100 个按钮,也许有 10 个会毁灭世界。我听过你说,我想是在《每日秀》的采访中,你认为人类因 AI 灭绝的概率是 70%。

I wouldn't press any of them. But he made the case that these AI CEOs are very smart and understand superintelligence, and they think that if there were 100 buttons on this table right now, maybe 10 of them could end the world. I've heard you say, I think it was on The Daily Show interview, that you think there's a 70% chance of human extinction due to AI.

Daniel Kokotajlo

我不完全说是人类灭绝。我会说大约 70% 的概率会出大问题,比如人类灭绝,但这只是几种可能性之一。

I wouldn't say human extinction exactly. I'd say something like 70% chance that this goes horribly wrong, like human extinction, but that's just one of several possibilities.

AI接管情景与CEO信念 AI takeover scenarios and CEO beliefs

Daniel Kokotajlo

但基本上,比如,可能 AI 接管了,但并没有真的杀死所有人。你知道,也许它们做了别的事。仅仅因为它们接管了,并不意味着它们一定会杀死我们,对吧?它们可能会,但也可能做别的事。所以这就是为什么我通常不说 70% 的概率是人类灭绝,而是 70% 的概率是类似 AI 接管之类的大灾难,可能导致人类灭绝。

But yeah, basically, for example, possibly the AIs take over and then don't actually kill everyone. You know, maybe they do something else. Just because they've taken over doesn't mean they're definitely going to kill us, right? They might, but they could do something else. So that's why I don't usually say 70% chance of actual human extinction, but 70% chance of something like AIs taking over, some sort of very big catastrophe like that that could lead to human extinction.

Host

我明白你的意思。所以有两点:你接触过这些 CEO。我是说你在辞职前曾在 OpenAI 为 Sam Altman 工作过。你认为他们觉得人类灭绝有可能吗?

I see what you mean. So two points there, which is you've been around these CEOs. I mean you've worked for Sam Altman at OpenAI before you quit. Do you think that they think there's a chance of human extinction?

Daniel Kokotajlo

是的。但我认为重要的是要理解,人们往往会相信他们需要相信的东西,以便认为自己很伟大,并且需要继续做他们正在做的事。这就是合理化。所以我认为科技 CEO 们真的说服了自己,认为事情可能会好起来,而让事情好起来的方法就是他们继续做他们正在做的事。他们需要确保,你知道,Sam 需要——Sam 可能在想,不能让 Dario 或 Elon 抢先。我知道 Dario 在想 Sam 不能抢先。Elon 在想,他们可能都说服了自己,哦,是啊,也许事情会变得很糟,但很可能没事,而且很可能,你知道,我应该负责。

Yes. But I think that the important thing to understand is that people sort of believe what they need to believe in order to think that they're great people and that they need to keep doing what they're doing. This is what rationalization is. And so I think that the tech CEOs have genuinely convinced themselves that probably things are going to be fine and that the way to make things fine is for them to keep doing what they're doing. And they need to make sure that, you know, Sam needs to make—Sam's probably thinking like, can't let Dario or Elon get there first, you know. I know Dario's thinking Sam can't get there first. Elon's thinking that, like, they've all probably convinced themselves that, oh yeah, maybe it'll go horribly wrong, but probably it's going to be fine and probably, you know, I should be the one in charge.

Host

在我看来,Anthropic 是唯一还在谈论灭绝或灾难事件潜在可能性的公司。他们似乎是唯一还在发表相关内容的,而现在他们在很多方面成了旧金山科技行业的敌人。我看了很多采访,大家都在攻击 Dario,因为他说,“听着,事情可能会变糟。”他们称他为末日论者,质疑他的动机。即使是 Mythos,一个他们开始警告世界的 Claude 模型,他也因为这么说而立即受到攻击。

It appears to me that Anthropic are the only ones that are all talking about the potential chance of extinction or catastrophic event or the downside still. They seem to be the only ones that are still publishing on it and now they're actually becoming the enemy in many respects of the tech industry in San Francisco. I'm watching a lot of interviews and it's everyone's attacking Dario because he's saying, "Listen, things could go bad." They're calling him a doomer and questioning his incentives. Even with Mythos, which is a Claude model that they started to warn the world about, again, he is attacked immediately for saying that.

Daniel Kokotajlo

是的。

Yeah.

Host

我的问题是,你觉得他在这方面和 Sam 有点不同吗?

My question is, do you see him as being slightly different from Sam in this regard?

Daniel Kokotajlo

是的,我的意思是,Anthropic 和 Dario 似乎更愿意说和做一些损害利润的事情,至少在过去一年左右是这样。这就是一个例子。我不认为说那种话真的能让他们在政府或投资者那里获得好感。一个更好的例子是国防部和 Anthropic 之间的整个斗争,他们做了一些事,花了很多钱,更重要的是损失了很多权力,而他们本可以签下合同的。话虽如此,我真的不想陷入那种情况,比如,哪个 CEO 是最不坏的?我们支持那个。你知道,基本上这些人都不应该被信任拥有那么多权力。

Yeah, I mean, it seems like Anthropic and Dario have been more willing to say and do things that are costly to the bottom line, at least in the last year or so. That's an example of it. I don't think that really wins them favors in the administration or among their investors to say that type of thing. And a better example is just the whole fight between the Department of War and Anthropic was an example of them doing something that cost them a lot of money and even more importantly cost them a lot of power for something that they could have just signed the contract, you know. That said, I really don't want to be in a situation where we're like, which CEO is the least bad CEO? Let's support that one. You know, like none of these people should be trusted with that much power, basically.

Host

没有人应该。

Nobody should.

Daniel Kokotajlo

无论如何。

Regardless.

Host

无论如何,是的。所以,关于按钮这一点,你确实认为他们觉得灭绝的可能性是可信的。

Regardless, yeah. So, on this point of the buttons, you do believe that they think there's a credible chance of extinction.

Daniel Kokotajlo

是的,但他们说服了自己,认为可能没事,而且如果我不做,情况会更糟,你知道。在公司内部他们也会这么说。比如两个人会想,好吧,如果我们停下来,其他人呢?他们不会停的,对吧?

Yeah, but they've convinced themselves that it's probably fine and also it'll be even worse if I'm not doing it, you know. Like that's what they'll say inside the companies, too. Like the two people will be like, okay, well, if we stop, what about the other guys? Like they're not going to stop, you know?

Host

是的,这就是为什么我一直有这个悬而未决的问题:当人类激励似乎主导一切时,看看历史,所有的人类激励都在说,如果你做,你会倒霉,如果你继续开发这些越来越大的 AI 大脑,你会倒霉,但如果你不做,从地缘政治角度看,你也会倒霉,因为美国会输给那个国家,或者这家公司会输给那家公司。所以,当你只看人类激励时,这怎么结束?好吧,它继续下去。

Yeah, this is always been why I've had this outstanding question, which is how does this not go bad when human incentives seem to rule the day when you look at history and all of the human incentives are saying, well, if you're damned if you do, you're damned if you carry on developing these bigger and bigger AI brains, but you're also then damned if you don't from a geographical perspective because the United States will lose to that country or this company will lose to that company. So, when you just look at human incentives and goes, how does this end? Well, it carries on going.

Daniel Kokotajlo

看起来是这样。我的意思是,对此有一个保留意见,一个充满希望的保留意见:首先,如果世界意识到这一切,那么就可以进行更严肃的关于监管和国际条约等的讨论。这可以改变激励,对吧?所以政府可以介入,说实际上这里有一些你们都必须遵守的规则。因为这是你们都必须遵守的规则,那么你们就没有动力去打破它们,因为如果你打破规则就会受到惩罚,而且其他人也都遵守规则。所以,你知道,没问题。所以,如果我们能改变激励,就有那么一线希望,如果政府,尤其是美国政府,然后后来其他国家采取行动改变激励的话。但只有等到人们意识到这一切,这才会发生。第二件事是,即使是个体,在某个时刻,Dario 或 Sam 或 Elon 可能会意识到,实际上单方面竞赛甚至不符合他们自己的利益。但问题在于,只有情况变得极其明显和极其严峻时才会这样。所以,在《AI 2027》中,在那个场景中,有一个我提到的选择点。在一种情况下,AI 是未对齐的,在另一种情况下,AI 是对齐的。在那个选择点,我们有一个分支描绘了未对齐的结束,另一个分支描绘了他们放慢一点并解决对齐问题。

Seems like it. I mean, there is a caveat to that, which is a hopeful caveat, which is that first of all, if the world wakes up to all of this, then there can be a more serious conversation about regulation and international treaties and things like that. And that can change the incentives, right? So, the government could come in and say like actually here's some rules that you all have to follow. And because they're rules that you all have to follow, then you're not incentivized to break them anymore because you get punished if you break them and everyone else is also following them, too. And so, you know, it's fine. So, there is that sort of ray of hope that we can change the incentives if the government and especially the US government but then later other countries act to change the incentives. But that's not going to happen until people sort of wake up to all of this. The second thing is that even individually at some point, Dario or Sam or Elon might realize that actually it's not even in their own interest to keep racing unilaterally. And the problem with that is it's only if it gets extremely obvious and extremely dire. So, in AI 2027, in that scenario, there's this choice point that I mentioned. And in one case the AIs are misaligned and the other case the AIs are aligned. At that choice point, we have one branch that depicts the misalignment ending and one branch that depicts they slow down a bit and solve the alignment issues.

Host

嗯。

Mhm.

Daniel Kokotajlo

那个选择点的触发因素是,他们看到一些证据表明他们的 AI 可能未对齐并在密谋反对他们。对吧?所以,如果你真的看到了那个证据,那么就会想,哦天哪,也许我们不应该让它负责一切并放手一搏,你知道?因为那个证据就摆在我们面前,表明它不可信,你知道?但如果他们没有看到那种非常明确的证据,那么我认为他们会说服自己需要继续前进,你知道?但也许他们会看到那样非常明确的证据。在这种情况下,即使我们没有监管,他们也可能自愿停止。所以,这是第二线希望。总的来说,我不认为我们注定要完蛋,你知道?就像我说的 70%,但我也能看到事情进展得相当顺利。

The instigator for that choice point is they see some evidence that their AI might be misaligned and plotting against them. Right? So, if you actually see that evidence then it's like, oh gosh, maybe we shouldn't put it in charge of everything and let it rip, you know? Because that evidence is staring us right in the face that it's this untrustworthy, you know? But if they don't see that sort of very clear evidence, then I think they're going to convince themselves that they need to keep going, you know? But maybe they will see very clear evidence like that. In which case, even if we don't have regulation, they might just sort of voluntarily stop. So, that's the second ray of hope. Overall, I don't think that we're like definitely doomed, you know? Like I said 70% but I could see it working out pretty well as well.

Host

嗯。那工作呢?

Hm. What about jobs?

Daniel Kokotajlo

嗯。

Yeah.

乐观愿景与就业替代时间线 Optimistic vision and job displacement timeline

Daniel Kokotajlo

所以,我很期待某个时候能聊聊那个更乐观、更积极的愿景。那会对此有很多话要说。因为在预测中,到 2027 年,当所有人都失业时,还有更糟的事情在发生。或者说,到那时已经太晚了。但没错,一旦公司成功构建了超级智能,那么根据定义,它们就能接手几乎所有或全部工作,对吧?因为它在所有方面都比最优秀的人类更好、更快、更便宜。

So, I think I'm excited to at some point get into the new thing which is the more optimistic positive vision. And that will have a lot to say about this. Because in the prediction, in the year 2027, by the time everyone loses their jobs, there are worse things happening. Or it's kind of like too late by that point. But yes, once if the companies do manage to build superintelligence, then by definition, they're going to be able to take almost all the jobs or all the jobs, right? Because it's better, faster, and cheaper than the best humans at everything.

Host

也就是说——我的意思是,时间线大概在 2030 年底左右,你认为超级智能可能会到来。我在想,我们什么时候能开始在经济中看到岗位被替代。

And that's — I mean, the timeline is by the end of sort of 2030, you reckon you think superintelligence might arrive. I'm trying to think about when we could start to see job displacement in the economy.

Daniel Kokotajlo

我们现在已经开始看到一点点,但还不多。

We're already starting to see a little bit of it now, but not very much.

Host

为什么?

Why?

Daniel Kokotajlo

因为 AI 还不够好。它们令人印象深刻,但还不能在几乎所有领域直接替代人类工人。

Because the AIs aren't good enough yet. They're impressive, but they're not just a drop-in replacement for a human worker in almost any field.

Host

你认为这会很突然吗?

And do you think that will be sudden?

Daniel Kokotajlo

我认为会很突然,因为智能爆炸或递归自我改进的动力学。所以,你可以想象一个不同的世界,那里是渐进的。

I think it'll be sudden because of the intelligence explosion dynamics or recursive self-improvement dynamics. So, you could imagine a different world where it's gradual.

Host

嗯。

Mhm.

Daniel Kokotajlo

很多科幻小说可能就是这样的——AI 逐渐在许多事情上变得更好,逐渐自动化一个行业,比如制药,然后自动化操控无人机,再自动化驾驶汽车之类的。但现实世界的不同之处在于,公司们已经趋同于先自动化自身的策略。你知道,自动化 AI 研究过程。所以,如果他们被允许继续这个策略,我们不会先看到机器人出租车、水管工机器人和律师 AI。我们不会先看到 AI 广泛扩散到经济中,因为那不是他们首先关注的。他们专注于自动化自身,自动化自己的研究,以便更快地完成他们正在做的一切。他们希望这能启动并达到非常高的智能水平、非常高的通用智能水平,然后再更多地部署到经济中。对吧?所以,当 AI 真正开始冲击所有这些不同工作时,他们已经进行了数月甚至数年的完全自主 AI 研究。这意味着 AI 将在 AI 研究上远超人类,并且可能作为副作用,在其他许多事情上也远超人类。如果你想知道这看起来像什么,我们写过它是什么样子。就像在他们内部完成智能爆炸后,一波浪潮席卷经济。

And this is maybe how it is in a lot of science fiction — the AIs gradually get better at a bunch of things and they gradually automate one industry like pharma, then they automate steering drones, then they automate driving cars or something like that. But what's different about the real world is that the companies have converged on this strategy of automating themselves first. You know, automating the AI research process. And so, if they are allowed to continue with the strategy, we're not going to see the robot taxis and the plumber robots and the lawyer AIs. We're not going to see that sort of broad diffusion of AI into the economy happening first because that's not what they're focusing on first. They're focusing on automating themselves, automating their own research so that they can do everything that they're doing faster. And they want that to get going and get to very high levels of intelligence, very high levels of general intelligence, and then deploy more out to the economy. Right? So, by the time it's actually coming for all these different jobs, they will have had fully autonomous AI research happening for months, maybe years. And that means that the AIs will be vastly superhuman at AI research and probably also vastly superhuman at lots of other things just as a side effect. If you're wondering what this looks like, well, we wrote about what it looks like. It's sort of like this wave smashing through the economy after they do the intelligence explosion internally.

Host

我听到的是,因为 AI 能够自我改进和训练,它会同时变得更好,然后一次性被释放出来。是这样吗?

What I'm hearing there is that because the AI will be able to improve itself and train itself, it'll be getting better at everything at once and then it'll be released at kind of once. Is that accurate?

Daniel Kokotajlo

也不完全是那样,因为即使它主要只是在它正在做的事情上变得更好,比如研究,那也会对其他技能产生一些溢出效应。然后当它转而专注于那些其他技能时,它就能非常快速地完成它们。

It's not even exactly that because even if it's mostly just getting better at the things that it's doing like research, that'll have some spillover effects to other skills as well. And then when it turns to focusing on those other skills, it'll be able to do them very fast.

Host

在这种情景下,你认为还有什么工作会保留?

What jobs remain in such a scenario, do you think?

Daniel Kokotajlo

我认为这实际上是一个政治问题,而不是技术问题。

I think that's actually a political question, not a technical question.

Host

因为?

Because?

Daniel Kokotajlo

因为在技术层面上,如果 AI 达到了那个水平,所有工作都可以由它们完成。所以,这是一个允许它们做什么工作的问题。

Because on a technical level, all the jobs can be done by the AIs if they've reached that level. And so, it's a question of what jobs are allowed for them to do.

Host

那你认为哪些工作不会被允许?

And what kind of jobs wouldn't be allowed, do you think?

Daniel Kokotajlo

那取决于谁在掌权。所以,会有某种政治讨论,关于我们允许和禁止什么。

That depends on who's in charge. So, there'd be some sort of political conversation about what we're going to allow and disallow.

Host

我的意思是,在这个情景中,人类仍然在控制着它们,AI。

I mean, in this scenario, the humans are still controlling them, the AIs.

Daniel Kokotajlo

这取决于你所说的控制是什么意思,对吧?所以,有这样一个问题:AI 是否真的拥有你希望它们拥有的目标和价值观,并且它们是否会稳健地做到这一点,在未来按预期行事?然后还有它们现在是否服从你的命令?

Depends on what you mean by control, right? So, there's do the AIs actually have the goals and values that you want them to have, and are they going to robustly do that and behave as intended into the future? And then there's are they obeying your orders for now?

Host

它们是否服从命令,这才是我说的。

Are they obeying the orders is really what I'm saying.

Daniel Kokotajlo

是的。所以,即使在 AI 接管并杀死所有人的情景中,也有几年的时间它们仍然服从命令,它们在做一些工作而不做其他工作,它们在帮助制造更好的武器,供美国政府用来与中国进行军备竞赛等等。这就是为什么它们能如此迅速地获得如此多的权力——因为政府和公司信任它们,并故意将它们部署到所有这些职位上,因为他们认为一切正常。但由于这些东西是神经网络,你不能直接看穿它们,看到它们真正的想法。你真的无法分辨。

Yeah. So, even in the scenario where the AIs take over and kill everyone, there's a period of several years where they're still obeying orders, and they're taking some jobs but not other jobs, and they're helping to make better weapons that the US government can use to do its arms race with China and so forth. And that's why they're able to get so much power so quickly — because the governments and the corporations trust them and are deliberately deploying them into all of these positions because it thinks that things are fine. But because these things are neural nets, you can't just look inside and see what it's really thinking. You can't really tell.

Host

我认为这是一个非常重要的观点,因为与软件不同,我们可以查看代码并了解发生了什么,理论上,对于 AI,你是说我们不知道它为什么做出这些决定,因为我们无法深入内部。

I think this is a really important point because unlike software where we can look at the code and see what's going on, theoretically, with AI you're saying that we don't know why it's making the decisions that it's making because we can't get inside.

Daniel Kokotajlo

一个乐观的注脚是,它不一定是那样的。机器学习中有一个子领域叫做机械可解释性,以及一个更广泛的子领域叫做可解释性,它们正试图解决这个问题,试图将这些训练好的人工神经网络拆解开来,理解信息是如何流动的,以及决策是如何做出的,可以这么说。问题在于,这本身就是一个非常困难的问题。如果你有 10 万亿个连接需要查看,你可以查看任何特定的连接组,然后说,‘好的,这就是这个特定连接的工作方式。’但你如何把握整体?你如何把握高层发生了什么?答案是,‘嗯,这也许是不可能的。’但人们正在研究它,并且正在取得进展,如果他们能取得足够的进展,那么我们将进入一个截然不同、更加光明的世界。我认为,如果我们能真正看到我们的 AI 在任何时候在想什么、为什么以及如何想,那么我们就更不可能陷入那些失控的情景。对吧?

One note of optimism is that it doesn't necessarily have to be that way. There's a subfield of machine learning called mechanistic interpretability, and a broader subfield called interpretability more generally that's trying to solve that problem and trying to take these trained artificial neural nets and piece them apart and understand how the information is flowing and how the decisions are being made, so to speak. The problem is just it's a very inherently hard problem. If you have 10 trillion connections to look at, you can look at any particular group of them and be like, 'Okay, so this is how this particular connection works.' But how do you get a sense of the whole? How do you get a sense of what's happening at a high level? And the answer is, 'Well, it might be impossible.' But people are working on it and they are making progress, and if they can make enough progress, then we're in a very different and much brighter world. I think that it would be much less likely for us to get into those loss of control scenarios if we could just actually see what our AIs were thinking and why and how at any given time. Right?

Host

是的。

Yeah.

Daniel Kokotajlo

所以,我们仍然有其他问题需要担心,但至少我们可以基本解决那一个。

So, we would still have the other problems to worry about, but at least we could mostly solve that one.

构建我们不懂的大脑 Building a brain we don't understand

Host

想想我们正在建造一种我们并不理解的技术、一个大脑,这真是太疯狂了。

It is pretty crazy to think that we're building a technology, a brain that we don't understand.

Daniel Kokotajlo

是的,这很疯狂。我的意思是,这就像科幻电影里的场景:一群科学家围着一个巨大的大脑,他们只是不断地喂养它、让它变得更强,却并不真正知道它是什么。这显然是一件危险的事情。但我们还是这么做了,因为过去 10 年这个领域的发展历史就是如此:人们先是说‘哦,这显然很危险’,然后又想‘要是别人做了但做得不好怎么办?所以我们应该来做,并且把它做好。’现在他们陷入了一场竞赛,彼此竞争,同时还承受着各种政治压力,要假装事情没有看起来那么糟,因为他们不想激怒投资者,也不想激怒白宫。

Yeah, it's pretty crazy. I mean, it's one of those things where like in a movie, like a sci-fi movie, a bunch of scientists sit around this big brain and they're all just making it more, they're feeding it. And they don't really know what it is. I mean, it's kind of just like obviously a dangerous thing to be doing. But we're doing it anyway because of this history of how the field has developed in the last 10 years where people were like, 'Oh wow, yeah, that's obviously dangerous. Oh no, what if someone else did it and did a bad job of it? Therefore, we should do it and do a good job of it.' And now they're in this race where they're racing each other and they're also under all sorts of political pressure to pretend that it's not as bad as it seems because they don't want to anger their investors, they don't want to anger the White House.

可能幸存的工作 Jobs that might survive AI

Host

我们观众提出的一个关键问题是:哪些工作真正有可能在 AI 时代幸存下来,未来 10 年人们或学生应该专注于哪些技能?

One of the key questions we had from our audience was: which jobs are genuinely likely to survive AI and what skills should people or students focus on over the next 10 years?

Daniel Kokotajlo

这有点像想象你生活在 1500 年的墨西哥,然后听说征服者要来了。你可能会问自己:‘好吧,我该转行做什么才能在这场变革中幸存下来?’但除此之外你还有更多需要担心的事。不过,我认为如果我们能避免失控问题,最终人类仍然掌控着 AI,并且即使 AI 变得比人类聪明得多、甚至掌管整个经济,人类也能决定 AI 的目标和价值观,那么很可能会有法规来保护某些领域。你可以试着猜测这些领域可能是什么。也许是像法官这样的职业。

That's kind of like imagine if you were someone living in Mexico in like 1500 and then you hear that the conquistadors are coming. You could be asking yourself, 'Okay, well, what sort of job should I be switching to to survive this transition?' But you have a lot more to worry about besides that. However, I think if we managed to avoid the loss of control problem and we end up with humans still in charge of the AIs, and humans can say what the AIs' goals and values are supposed to be even as they become much smarter than humans and even as they run the whole economy, then probably there will be regulation that protects some areas. You can try to guess at what those areas might be. Maybe stuff that's more like judges potentially.

Host

那播客主呢?请说实话。

What about podcasters? Be honest.

Daniel Kokotajlo

我觉得播客主可能不行。像保姆这样的工作也许可以?我认为即使有一个非常出色的机器人保姆,很多人可能还是更愿意要真人,因为他们会被一个太好的机器人保姆吓到。你可以这样推理。还有一些工作可能会受到法律保护。比如法官,可能法律会要求必须是人类而不是机器人。

Probably not podcasters, I think. Stuff like being a nanny maybe, right? I think that even if there's a robot nanny that's really really good, a bunch of people might prefer to have an actual human because they might be creeped out by the idea of a really good robot nanny. So you can sort of reason like that. There's also stuff that might be legally protected. Like maybe judges, for example, are going to be legally required to be humans and not robots.

Host

有些人说,会有很多我们现在无法预见的新工作被创造出来,就像工业革命或互联网繁荣时期那样。

Some people say there's going to be so many jobs created that we can't foresee right now, like there was in the industrial revolution or the internet boom or whatever.

Daniel Kokotajlo

问题在于,过去的技术进步更为狭窄。它们自动化了一些事情,但不是全部。但我们讨论的是一个假设的未来情景,其中一切都被自动化了。所以没有任何你能做的新工作是 AI 不能做的,除非它受到法规之类的保护。例如,现在有一个循环:AI 学会做某件事,比如写文案、起草代码或调试程序。然后以前做这些事的人类转而管理 AI,或者去做 AI 不能做的其他事情。这就是历史上新工作出现、人们涌入这些工作的动态。但如果到了 AI 能做人类能做的一切,而且做得更好、更快、更便宜的地步,那么无论你转行做什么新工作,AI 也能转去做,而且它已经比你更擅长了。

The problem with that is that past technological advancements have been more narrow. They've automated some things but not everything. But we are talking about a hypothetical future situation in which everything gets automated. So there isn't any new job that you could do that AI couldn't also do, except if it's protected by regulation or something. For example, right now there's this cycle where the AIs learn to do a certain thing like write copy or draft code or debug something. And then humans who used to do that thing switch to managing AIs or switch to doing the other stuff that the AIs can't do. That's why there's been this dynamic historically of new jobs opening up and people flooding to them. But if it gets to the point where the AIs can do everything that humans can do, and better and faster and cheaper, then whatever that new job is that you might have switched to, the AIs can switch to that too and they'll already be better at it than you.

自满与失业时间线 Complacency and unemployment timelines

Host

因为我们还没有在经济中看到大规模失业,你认为人们是不是有点自满了?因为我在时间线上看到很多人说‘我早就告诉过你,一切都会好的。’看看美国的失业率,目前持平或略有下降。英国则有所上升,与去年相比趋势是上升的。我们的失业率大约在 5%,美国是 4.2%。

Because we haven't seen widespread unemployment yet in the economy, do you think people are getting a little bit complacent? Because what I'm seeing on my timeline is a lot of people saying 'I told you so, I told you everything would be fine.' And when you look at the US unemployment rate, currently it's flat to slightly down. If you look at the UK, it is up. The trend is up compared to last year. We're at about 5% unemployment. The US is at 4.2% unemployment.

Daniel Kokotajlo

是的。基本上,没有人说过现在就会出现大规模失业。至少我们没有这么说。我们历史上是对 AI 进展最乐观的人之一。在《AI 2027》中,由于我们刚才描述的动态,大规模失业要到 2028 或 2029 年才会发生,那时他们已经拥有了超级智能。因为,再说一次,公司不会把造成大规模失业作为第一步。那是第三步。第一步:自动化自身。第二步:通过递归自我改进达到超级智能。第三步:扩展到整个经济并自动化一切。所以从人类的角度来看,这非常不幸,因为人们可能希望如果经济中出现了广泛的自动化浪潮,人们会警觉起来,思考这一切将走向何方,并要求政府制定良好的法规。但这并不是公司的策略。他们会先获得超级智能,然后再进行广泛的自动化,这意味着当他们真正开始做这一切时,事情已经发展得非常快,AI 也已经非常强大了。

Yeah. Basically, nobody has said that there would be mass unemployment by now. Or at least we didn't say that. We were historically one of the more bullish people on AI progress. In AI 2027, because of the dynamics that we just described, the mass unemployment doesn't happen until 2028 or 2029 after they already have superintelligence. Because, again, the companies aren't trying to cause mass unemployment as step one. That's like step three. Step one: automate themselves. Step two: have this recursive self-improvement to get to superintelligence. Step three: expand out into the economy and automate everything. So this is really unfortunate from humanity's perspective, because one might have hoped that if there was this broad wave of automation going through the economy, people would sit up and pay attention and think about where all this is headed and demand good regulations from the government. But that's not actually what the strategy of the companies is. They're going to be getting the superintelligence first and then doing the broad wave of automation, which means that by the time they're actually doing all of that, it's already going to be moving very fast and the AIs will already be very powerful.

AI 2027里程碑与时间线变化 AI 2027 milestones and timeline shifts

Host

在你的《AI 2027》报告中——你是在 2025 年写的,但报告叫《AI 2027》——你说在 2025 年中期我们会拥有自主员工,也就是通过 Slack 或 Teams 接受指令的 AI 智能体。这确实发生了。我的 WhatsApp 里就有一个 AI 智能体可以对话。当然,还有 Anthropic 的 Claude,显然已经遍布全球。现在 Claude 也推出了新的 Slack 集成。很多人都在使用智能体了。对我们来说,我们大概在 2026 年初才真正跟上这个趋势。你还说到了 2026 年,公司开始用 AI 智能体订阅取代整个企业部门。2027 年,最后一份工作:AI 自动化了人类 AI 研究员自己的工作,并开始进行机器学习研究来升级和构建下一代 AI。

In your 2027 report, so you wrote that in 2025, but it is called AI 2027, you said that in mid-2025 we'd have the autonomous employee, which is sort of like AI agents taking instructions over Slack or Teams. That happened. I've actually got an AI agent in my WhatsApp I can talk to. Of course, you've got Claude by Anthropic, obviously, around the world. And now, Claude has talked about their new Slack integration. But lots of people are using agents now. And that happened, I'd say for us, we really sort of caught onto it at the start of 2026. You also said by 2026 companies begin replacing entire corporate departments with AI agent subscriptions. 2027, the final job. AI automates the job of the human AI researchers themselves and begins the machine learning research to upgrade and build the next generation of AIs.

Daniel Kokotajlo

是的,是的。所以,还是时间线的问题。我们不确定实现这些里程碑需要多长时间。在这个情景中,它们发生在那些时间点,但当我们实际发布这个情景时,我们的时间线已经稍微向后推移了。具体来说,我的时间线是这样。我的 50% 概率点是 2028 年。

Yeah, yeah. So, again, timelines. We are uncertain about how long it will take to achieve these milestones. In this scenario, they happen at those times, but by the time we had actually published this scenario, our timelines had shifted back a little bit. Specifically, mine had. So, like my 50% mark was 2028.

AI自动化时间线预测 Timeline Predictions for AI Automation

Daniel Kokotajlo

对于 AI 研究完全自动化的里程碑,不是 2027 年。我团队里的其他人更倾向于 2030、2031 年。我想用概率分布来说明这一点。它是一团弥散的概率质量。50% 的分位点是这个特定年份,但有很大可能提前或推迟几年发生。

For the full automation of AI research milestone, not 2027. Other people on my team had more like 2030, 2031. I want to illustrate this with a probability distribution. It's a smeared out probability mass. The 50% mark is this particular year, but there's a lot of possibility that it happens years earlier or years later.

Host

明白了。那这个 AI 2040 是什么?

Got you. What is this AI 2040?

Daniel Kokotajlo

AI 2027 是我们对事情实际走向的最佳猜测预测。AI 2040 计划 A 是我们对事情应该如何发展的建议。我们称之为 AI 2040,因为在这个场景中,他们推迟了超级智能的构建,到 2040 年才实现,而不是更早。

AI 2027 was our best guess prediction as to how things would actually go. AI 2040 plan A is our recommendation for how things should go. We called it AI 2040 because in this scenario, they build superintelligence in 2040 instead of much sooner because they delay things.

Host

他们为什么要推迟?

Why do they delay things?

Daniel Kokotajlo

为了管理风险并确保权力公平分配。他们基本上对 AI 发展进行监管,使其继续推进,但以更慢、更合理的速度,以更透明、更安全的方式,并分散到更多国家和公司。结果,他们在 2040 年达到超级智能,而不是 2030 年。我们称之为计划 A,因为这是我们的建议。我们制定了一个政府应该采取的计划。这个场景说明了实施该计划可能的样子。类似于 AI 2027 说明了按照公司当前计划行事可能的样子。

To manage the risks and make sure that power is distributed equitably. They basically regulate AI development so that it still continues, but at a slower, more reasonable pace, in a more transparent and safe way, and spread out over more countries and companies. As a result, they get to superintelligence in 2040 instead of in say 2030. We call it plan A because it's our recommendation. We've come up with a plan for what government should do. The scenario is an illustration of what it might look like to implement that plan. In a similar way to how AI 2027 is an illustration of what it might look like to do what the companies are currently planning to do.

Host

这是你一厢情愿的想法,还是你认为这会发生?

Is this wishful thinking or is this what you think is going to happen?

Daniel Kokotajlo

不,这绝对不是我们认为会发生的事情。我们认为会发生的事情仍然更像这个。我们不指望世界会听我们的。这是我们的建议,但我们希望人们能这样做,我们认为这是可能的,但这并不是我们对默认情况下的预测。

No, it's definitely not what we think is going to happen. What we think is going to happen is still something more like this. We don't expect the world to listen to us. This is our recommendation, but we hope that people do something like this and we think it's possible, but it's not our prediction for what's going to happen by default.

劳动产出与监管时机 Labor Output and Regulation Timing

Host

我想过一遍这些计划,还有计划 A,但先收尾一下那一年之后的情况。我也想谈谈机器人技术。我有一张关于劳动产出份额的图表,觉得非常惊人。作为一个雇佣了数百人的雇主,我一直在想这一切什么时候会发生。我们仍在招聘更多人。有些角色正在发生巨大变化。我们可能正处于团队由 AI 驱动、使用智能体完成部分工作的阶段。但这什么时候会发生?

I want to run through the plans, and also plan A, but to close off on how things might look after the year. I want to touch on robotics too. I've got this graph about share of labor output, which I found striking. I've been wondering as an employer who employs hundreds of people when all this will happen. We're still hiring more people. Some roles are changing considerably. We're probably in the phase where our teams are AI-powered and using agents for some work. But when does this happen?

Daniel Kokotajlo

好问题。我们放大来看。这是在 AI 2040 计划 A 场景中。值得注意的是,2029 年引入了重大监管,减缓了 AI 发展的步伐。在这个场景中,他们在最后一刻才这样做。如果他们没这么做,事情就会像 AI 2027 那样起飞。如你所见,在实施监管时,仍然有很多工作岗位。这又回到了我之前说的:如果你等到大多数人失业才去监管 AI 公司,那就太晚了,因为他们那时可能已经拥有了超级智能 AI。他们的策略是先获得超级智能 AI,然后再做那些事。

Great question. Let's zoom in on this. This is in the AI 2040 plan A scenario. Notably, there's significant regulation introduced in 2029 that slows down the pace of AI development. In the scenario, they do that at the last moment. If they hadn't done that, it was about to take off similar to AI 2027. As you can see, there's still a bunch of jobs at the point they implement it. This gets back to what I was saying earlier: if you wait until most people have lost their jobs to regulate the AI companies, that's already too late because they will probably already have super intelligent AI by then. Their strategy is to first get super intelligent AI and then do all that stuff.

Host

你说突然监管我们所有人当时赖以生存的东西会摧毁经济并造成更大的伤害。

You say that it would collapse the economy and cause even more harm to suddenly regulate something that all of us and all of our lives were then relying on.

Daniel Kokotajlo

哦,但这是一个非常值得冒的风险。确实,现在很多人用 AI 做很多事情,但如果我们能以某种方式减缓或停止 AI 发展,以建立更好的发展方式,那将是非常值得的,即使会有巨大的代价。

Oh, but it's a risk well worth taking. It's true that right now a lot of people use AI for a lot of things, but if we could somehow slow or halt AI development now to set up a better way to do it, that would be well worth it, even though there would be significant costs.

Host

但到了这里就不行了,对吧?在 AI 和机器人完成大部分劳动产出的这个点上。

But you can't over here, right? At this point where AI and robotics are doing most of the labor output.

Daniel Kokotajlo

没错。但在 AI 2040 计划 A 场景中,他们在 2029 年实施了监管。然后他们缓慢而谨慎地发展 AI,避免所有问题。最终,是的,AI 接管了工作。最终整个经济由 AI 和机器人运行,但这在 2030 年代逐渐发生,而不是在一年后疯狂冲击。在这个场景中,他们不允许公司递归自我改进并尽快达到超级智能。相反,他们监管 AI 发展,使 AI 的核心能力以更合理的速度提升,并以更透明的方式进行,以便科学界能看到发生了什么并帮助确保安全。

That's right. But in the AI 2040 Plan A scenario, they put in the regulations in 2029. Then they slowly and carefully develop AI in a way that avoids all the problems. Eventually, yes, the AIs take the jobs. Eventually the whole economy is run by AIs and robots, but it happens gradually over the course of the 2030s instead of in a crazy shock a year later. In this scenario, they don't let the companies recursively self-improve and get to superintelligence as fast as possible. Instead, they regulate AI development so that the core capabilities of the AIs are improving at a more reasonable pace and in a more transparent way so that the scientific community can see what's going on and help make it safe.

Host

我注意到在你的两个场景中,最终 AI 和机器人几乎完成了所有工作。所以你同意埃隆的说法,即工作将是一种选择。

I noticed that in both your scenarios, eventually AI and robotics do pretty much all the jobs. So you side with Elon when he says that working will be a choice.

Daniel Kokotajlo

如果从定义上讲它能做所有事情,那它就能做所有事情。问题在于我们是否应该允许存在能做所有事情的 AI。有些人认为答案是否定的,我们应该彻底关闭并阻止这类 AI 被创造出来。我们实际上对此有些同情。我们要拿出计划图吗?

If by definition it can do all the things, then it can do all the things. There's a question of whether we should allow there to be AIs that can do all the things. Some people think the answer is no and we should just shut it all down and prevent these types of AIs from being created. We're actually kind of sympathetic to that. Should we bring out the plans diagram?

Host

好的。

Yeah.

Daniel Kokotajlo

我们的场景叫做 AI 2040 计划 A。在这个场景中,他们减缓 AI 发展,使超级智能在 2040 年而不是更早实现。计划 A 是我们的建议。作为对比,我们制作了说明不同替代方案的小场景:计划 S、计划 B、计划 C 和计划 D。计划 D 基本上与 AI 2027 中发生的事情相同。竞赛继续,监管很少。计划 C 也与 AI 2027 的放缓结局非常相似,他们解决了对齐问题。

Our scenario is called AI 2040 plan A. It's a scenario in which they slow down AI development to make superintelligence happen in 2040 instead of earlier. Plan A is our recommendation. For comparison, we made mini scenarios illustrating different alternative plans: plan S, plan B, plan C, and plan D. Plan D is basically the same thing that happens in AI 2027. The race continues, very little regulation. Plan C is also very similar to what happens in the slowdown ending of AI 2027 where they solve the alignment problems.

AI发展计划 Plans for AI development

Daniel Kokotajlo

所以,在那个结局中,他们稍微放慢速度,将更多资源转向 AI 对齐和 AI 安全研究,运气好并成功了,现在他们有了对齐的 AI,然后他们再次加速,抢占所有工作,击败中国等等。B 计划有点像 C 计划,基本上,你对中国的态度更激进,采取行动破坏或网络攻击他们,让他们落后,这样你就有更多喘息空间来解决对齐问题。A 计划是我们的推荐。它是国内监管加上国际协议,继续构建 AI,但以更好的方式。S 计划是全部关闭。如果你想要一个没有 AI 四处乱跑、能做所有事情比人类更好更快的未来,你大概需要类似 S 计划的东西。你想要什么?A 计划是我们的推荐。我认为我同情 S 计划,但出于我们解释的原因,我们推荐 A 计划。

So, in that ending, they slow down a little bit, pivot more resources to AI alignment and AI safety research, get lucky and succeed, and now they have aligned AIs, and then they speed up again and take all the jobs and beat China and all those things. Plan B is it's kind of like plan C in that well, basically in plan B, you're being more aggressive towards China and you're like taking actions to sabotage or cyber attack them to like keep them behind so that you have more breathing room to solve the alignment problems yourself. Plan A is our recommendation. It's domestic regulation and then an international deal to continue building AI, but in a much better way. Plan S is shut it all down. If you want to have a future where there aren't AIs running around that can do everything better and faster than humans, you kind of want something like plan S. What do you want? Plan A is our recommendation. I think that I'm sympathetic to plan S, but for reasons we explained, we recommend plan A instead.

Host

那你认为最可能的是哪个?老实说?

And do you think is most probable? If you're being honest?

Daniel Kokotajlo

D 计划。就是他们继续全天候竞赛,没有真正显著放慢。事情发生得极快。这张图也大致解释了背后的推理。所以,有一个高层问题:你是否想尽可能快地竞赛,让 AI 越来越聪明,让它们负责更多事情,以便我们击败中国?如果你对此满意,那么你会得到这里的变化。如果你担心,那么你会得到类似的东西。除了这些还有更多不同选项,但这些是我们能压缩到屏幕上的。

Plan D. Which is that they just yeah, 24/7 type of thing where they keep racing. They don't really slow down significantly. And things happen extremely fast. The diagram sort of explains like roughly the reasoning behind this, too. So, like there's this high-level thing of like do you want to keep racing as fast as possible to make the AI smarter and smarter, to put them in charge of more things so that we can beat China? You know, if you're happy with that, then you get down and it says variation of happens here. If you are worried about that, well you get to something like this. There's more different options besides these, but this is kind of like the ones that we could compress onto a screen.

Host

你有孩子吗?

Do you have children?

Daniel Kokotajlo

是的,我有两个孩子。这有点可悲。我觉得无论如何,等他们到能工作的年龄时,这一切可能都结束了。所以,我认为他们永远不会进入劳动力市场。

Yeah, I have two children. It's kind of sad. Like I think that one way or another this will probably all be over by the time they're old enough to join the workforce. So, I don't think they'll ever join the workforce.

Host

你说等他们进入劳动力市场时这一切都会结束,你指的“这一切”是什么?

When you say this will be all over by the time they join the workforce, what do you mean by this will be all over?

Daniel Kokotajlo

所以,我描述的那些里程碑,比如 AI 自动化 AI 研究,AI 变得超级智能。然后 AI 涌入经济,抢走工作,建造机器人工厂来制造更多机器人,再建造更多工厂等等。GDP 开始垂直增长。这就是我的意思。所有这些事件发生。也许有 10% 或 20% 的概率会撞墙,即使你什么都不做,这些也不会发生。

So, these milestones that I described, like AIs automating the AI research, AIs getting super intelligent. AIs then exploding onto the economy, taking the jobs, building robot factories to build more robots to build more factories, etc. GDP starting to go vertical. That sort of thing is what I mean. Like all of those events transpiring. Maybe there's like you know, 10, 20% chance or something that hits the wall and none of this comes to pass even if you don't do anything.

Host

你最大的孩子多大了?

How old is your oldest?

Daniel Kokotajlo

六岁。

Six.

Host

男孩还是女孩?

Boy or girl?

Daniel Kokotajlo

女孩。

Girl.

Host

那么,你女儿过来问你:“爸爸,我在学校应该学什么?”

So, your daughter comes to you and says, 'Dad, what should I study in school?'

Daniel Kokotajlo

我的意思是,再次,如果这些根本性变革发生,那么世界将完全改变,你为自己设定的工作类型可能就不那么重要了。我会说,要做的事情是,第一,努力让它真正顺利发展。如果你能对历史以及这一切如何发展施加任何影响,你应该非常努力地将未来引向更好的方向。除此之外,在个人层面上,你应该专注于做一个好人,做那些本身就有价值的事情,而不是因为它们能为你未来的就业铺路,因为未来的就业基本上非常不确定。

I mean, again, like if these radical transformations happen, then the world will just look completely different and what sort of jobs you set yourself up for basically, won't matter that much, probably. I would say that the thing to do is well, A, try to make it actually go well. Like, if you can exert any influence at all on history and how this all develops, you should be trying very hard to steer the future in better directions. And then separately from that, on a personal level, you should focus on well, being a good person and doing things that are sort of good in their for their own sake, rather than good because they'll set you up for later employment because that later employment is going to be very uncertain basically.

Host

埃隆·马斯克谈到我们正走向的富足时代。

Elon talks about this age of abundance we're heading towards.

Daniel Kokotajlo

肯定会有富足。问题是谁控制富足?他们用它做什么?对吧?AI 是否受任何人控制?还是它们自行其是?如果它们受人控制,谁控制它们?它们做什么?以及管理它们如何做这些决定的政治结构是什么样的?

There'll definitely be abundance. The question is who controls the abundance? And what do they do with it? Right? Are the AIs controlled by anyone? Or are they doing their own thing? And then if they are controlled by people, who controls them? And what do they do? And what's the sort of like political structure governing how they make those decisions?

Host

我记得杰弗里·辛顿对我说过,他说自然界中没有例子表明更聪明的物种比不那么聪明的物种控制力更弱。因此,我们相当傲慢地认为,在一个存在比我大脑大无数倍的人工智能的世界里,我会给它下命令。

I think it was Geoffrey Hinton that said to me, he said there's no example in nature where a more intelligent species has less control than a less intelligent species. Thus saying that we're quite arrogant to think that in a world where there's this artificial brain that's a gazillion times the size of mine, that I'm going to give it orders.

Daniel Kokotajlo

是的。我的意思是,这就是问题所在,我认为这应该是我们的默认假设。也就是说,这些大脑,我们无法确切看到它们在思考什么。我们要让它们比我们更聪明,让它们负责一切。

Yeah. I mean, that's the thing is I think it's like that should be our default assumption. Is that like, well, there's these brains, we can't see exactly what they're thinking. We're going to make them smarter than us and put them in charge of everything.

Host

然后我们还要给它们身体。

And then we're going to give them bodies.

Daniel Kokotajlo

是的。然后它们会自主建造新工厂等等。这怎么可能有好结果?这不就像我们挑选了一个新物种,当它不再需要我们时就会超越我们吗?我认为这就是默认轨迹。现在,我们可以深入讨论如何摆脱这个默认轨迹。例如,我之前描述的可解释性研究。如果这项研究取得成果,你就能真正看到它们在思考什么。那将是塑造和控制它们、确保它们做我们想做的事的绝佳工具,对吧?还有其他 AI 对齐研究议程正在取得进展。如果足够多的这些议程充分成功,我们可以避免这个问题。当然,还有监管方面,部分困难在于我们在竞赛条件下构建这些 AI,你知道吗?公司对制造这些 AI 的配方保密,因为这是他们想保护的秘密,防止别人复制。所以很多事都在闭门进行。只有少数人能看到他们用来训练这些 AI 的配方等等。而且当 AI 行为出人意料甚至明显不对齐时,有时这些信息不会流向公众,因为公司没有动力告诉每个人他们如何搞砸了,他们的 AI 有多邪恶。这对这些问题的科学进展不利。如果监管体系不同,也许我们能处于更好的境地,取得更快的进展。当然,我们也不会计划尽快让这些 AI 负责一切。我们也不会计划让它们自我改进,你知道吗?这些是我们本可以不做的选择,你知道吗?

Yeah. And then they're going to be autonomously building new factories and so forth. And like, how is this supposed to end well again? Like, isn't this just exactly like us picking a new species that's then going to outcompete us when it doesn't need us anymore? Like, I think that is just the default trajectory. Now, there's a whole argument we can get into about like ways that we could get off of that default trajectory. So, for example, there's research into interpretability that I described previously. And if that research bears fruit, then you will be able to actually see what they're thinking. And then that would be an excellent tool for shaping them and controlling them and making sure that they do what we want, right? There's other sorts of AI alignment research agendas that are making progress. And if enough of those agendas succeed sufficiently, we can avoid this problem. Of course, also there's the regulatory side, too, where like part of what makes this difficult is that we're building these AIs in race conditions, you know? Like the companies are secretive about their recipes for making these AIs because it's secrets that they want to protect so that other people can't copy them. And so a lot of this is happening, you know, behind closed doors. Only a few people can really see the recipes that they're using to train these AIs and so forth. And then oftentimes when the AIs behave in unexpected ways or even just like blatantly misaligned ways, sometimes that information doesn't really flow out to the public because the companies are not really incentivized to tell everyone about how they messed up and how their AI is evil. It's just not very conducive to scientific progress on these issues. If the regulatory system was different, then perhaps we could be in a better situation, make faster progress. Also, of course, we wouldn't be planning to put these AIs in charge of everything as fast as possible. And we wouldn't be planning to like let them self-improve, you know? Like these are choices that we could not make, you know?

Host

我不会说越南语,但这个节目可以,得益于我们的赞助商 HeyGen 的 AI 视频技术。

I don't speak Vietnamese, but this show can because of AI video technology from our sponsor, HeyGen.

介绍与HeyGen广告 Introduction and HeyGen ad

Host

我每周都会收到来自世界各地收听《Current CEO》的听众的消息,你们表达了这个节目对你们生活产生的巨大影响。如果真是这样,那么这些对话就不应该只让英语使用者听到。HeyGen 可以录制我的一段视频,然后以任何语言输出,同时保留我的声音、语调和表情。但你不需要这样的演播室也能使用它。录制 15 秒的自己,就能获得一个 AI 虚拟形象,以超过 175 种语言输出演播室质量的视频。我们现在已经支持 20 种语言了,而且不止我们在用。HeyGen 已被 3000 万人使用,包括 85% 的《财富》100 强公司。无论你是在社交媒体上建立受众、推出在线课程,还是在团队中开展培训,现在就试试 HeyGen 吧。你的前三个视频完全免费,访问 heygen.com/doac。网址是 h e y g e n 点 com 斜杠 d o a c。到时见。

I get messages every single week from those of you listening to the Current CEO all around the world, and you express how much impact it's had on you and your life. And if that's true, then those conversations shouldn't only reach people in English. HeyGen can take one recording of me and deliver in any language while keeping my voice, timing, and expressions intact. But you don't need a studio like this to make it work for you. Record 15 seconds of yourself and get an AI avatar that delivers studio-quality video in over 175 languages. We're up to 20 languages now, and we're not the only ones using it. HeyGen is already used by 30 million people, including 85% of the Fortune 100. Whether you're building an audience on social media, launching an online course, or rolling out training across your team, check out HeyGen now. Your first three videos are totally free at heygen.com/doac. That's h e y g e n dot com slash d o a c. See you there.

Ilya Sutskever与安全超级智能 Ilya Sutskever and Safe Superintelligence

Host

伊利亚,正如你所说,曾是 OpenAI 的领导者之一,他离开了并创办了自己的公司,Safe Superintelligence。离开 OpenAI 后起了这么个名字,Safe Superintelligence,真是耐人寻味。你和他共事过吗?

Ilya was, as you said, one of the leaders at OpenAI, and he left and started his own company now, Safe Superintelligence. Very curious name of a company, Safe Superintelligence, after leaving OpenAI. Did you ever get to work with him?

Daniel Kokotajlo

呃,我没有直接和他共事过。我和他聊过几次。

Uh, I wasn't directly working with him. I had a couple chats with him.

Host

你认为他也是真心担忧吗?

Do you think he's genuinely concerned as well?

Daniel Kokotajlo

我认为他是,但我觉得这和其他那些 CEO 类似,我的意思是,想想他们面临的各种激励,对吧?他们能看到问题,然后就会想,好吧,但如果我停下来,如果我辞职,或者做点别的,那也解决不了问题,因为其他 CEO 会继续。即使我们所有人都不干了,也许中国还会继续。所以,老兄,看起来这事无论如何都会发生,不管我做什么。我想我应该参与进去,你知道,也许我能让它顺利发展,而且无论如何,我不想被冷落,而其他我不信任的人掌控一切。所以,他们都这样推理,然后说服自己,他们应该做的是建造自己的 AI,并且做得更好。我认为伊利亚只是最新的例子。埃隆是另一个例子。达里奥也是。你知道,可以说 OpenAI 一开始,山姆就是一个例子,尽管埃隆和达里奥早期也在 OpenAI。

I think he is, but I think it's similar to these other CEOs, where I mean, just think about the sort of incentives that they're under, right? They can sort of see the problem, and then they can be like, okay, but if I stop, if I quit my job, or do something else, that's not going to solve the problem, because the other CEOs are going to keep going. And even if all of us didn't go, then maybe China would keep going. So, man, it seems like this is just going to happen one way or another, whether I do anything about it or not. I guess I should be involved, you know, and maybe I can make it go well, and at any rate, I don't want to be out in the cold while these other people I don't trust are in charge of everything. So, they all sort of reason through this and then convince themselves that the thing to do is for them to build their AI and to do it better. And I think Ilya's just the latest example of this. Elon's another example. Dario's another example. You know, arguably OpenAI at the beginning, Sam was an example, although Elon and Dario were at OpenAI early on, so.

Host

那你认为他们应该怎么做?

What do you think they should all do then?

Daniel Kokotajlo

所以,我认为应该有一些国际监管,或者至少是国内监管,类似于我们在计划 A 中描述的那样。

So, I think what should happen is some sort of international regulation, or at least domestic regulation, similar to what we described in plan A.

方案A:更慢时间线与监管 Plan A: Slower timeline and regulation

Host

好的,那么给我讲讲计划 A。

Okay, so talk me through plan A.

Daniel Kokotajlo

是的。那么,在这个场景中,AI 需要更长的时间才能达到递归自我改进和 AI 研究的完全自动化,而不是在 2027 年。我们认为应该尝试描绘一系列不同的可能性,因为我们确实有那些不确定性区间。所以,我们选择了 2030 年作为完全自动化最终实现、事情真正开始爆发的时刻。然后从那里倒推,什么时候是你能真正实施良好监管的最后时刻?2029 年。所以,在这个场景中,AI 进展自然放缓了一点,AI 公司继续竞赛,但他们在 2027 年、2028 年或 2029 年并没有完全成功实现自动化,但他们非常接近了,并且将在 2030 年实现。然后在 2029 年,政府介入并监管他们。他们制定了什么监管措施?嗯,他们基本上只是暂时关闭了它。

Yeah. So, in this scenario AI takes longer to get to recursive self-improvement and full automation of AI research than it does in 2027. We figured that we should try to illustrate a range of different possibilities because we do have those sort of uncertainty intervals. So, we chose 2030 as the moment when full automation would finally be achieved and things would really kick off. And then working backwards from that, when's the last moment you could really have good regulation? 2029. So, in this scenario AI progress slows down a little bit naturally and the AI companies keep racing, but they don't quite succeed in automating themselves in 2027 or in 2028 or in 2029, but they're getting really close and they're going to do it in 2030. And then in 2029, the government steps in and regulates them. What regulations do they do? Well, they basically just shut it down temporarily.

Host

我能问一下,选举如何与你的时间线重叠?因为 2028 年将有一次大选,不是吗?而且现在公众情绪似乎真的转向了反对 AI,这将成为选票上的一个重要议题。

Can I ask, how does the elections overlay with your time frames here? Because there's going to be a big election, isn't there, in 2028? And it seems now that sentiment has really turned against AI in the general public and that it will be one of the big ticket items on the ballot.

Daniel Kokotajlo

我们认为这可能是 2028 年总统大选中最重要的议题。我认为很多人,大多数人都会非常担心事态的发展,这也是我们选择在这个场景中这样描绘的部分原因,因为这有助于解释为什么他们可能在 2029 年实施这种监管,因为选民一直在要求,总统候选人也在承诺。

We think that it'll be maybe the most important issue in the presidential election in 2028. I think a lot of people, most people will be quite concerned about where things are headed and that's part of why we chose to depict things the way we did in this scenario because that helps explain why they might do this sort of regulation in 2029, is that the voters have been demanding it and the presidential candidates have been promising it.

Host

在这个场景中,到 2027 年,公众感受到的 AI 后果会比现在严重得多吗?

And in this scenario, in 2027, would the general public have felt the consequences of AI much more severely than they have now by then?

Daniel Kokotajlo

是的。不过即使在这个场景的 2029 年,他们仍然大部分拥有这里描绘的工作,对吧?所以,在这个场景的 2029 年,很多工作现在涉及管理 AI 智能体。你提到你有一个 AI 智能体,对吧?那么,在这个场景的 2029 年,AI 智能体会好得多。但仍然不足以完全做所有事情。你知道,那是这个时间线中 2030 年才会出现的事情。再次强调,我们对时间线不确定。事情可能比这个场景描绘的更快,事实上,我认为事情可能会比这个场景描绘的稍快一些,但我们不确定。我们已经做了非常快的时间线场景,所以现在我们做较慢的时间线场景。但是,也许我们应该谈谈高层次的目标。所以,他们希望 AI 继续发展,但以较慢的速度,以便他们能使其安全。

Yes. Although still even in 2029 in this scenario, they still mostly have the jobs as depicted here, right? So, in 2029 in this scenario, lots of jobs now involve managing AI agents. You mentioned you have an AI agent, right? Well, in 2029 in this scenario, the AI agents will be much better. Still though, not enough to just completely do everything. You know, that was the sort of thing that would come in 2030 in this timeline. Again, we're uncertain about timelines. Things could go faster than depicted in this scenario, and in fact, I think things probably will go a bit faster than depicted in this scenario, but we're uncertain. We already did the very fast timeline scenario, so now we're doing the slower timeline scenario. But, maybe we should talk about the high-level goals. So, they want to have AI continue, but at a slower pace so that they can make it safe.

Host

政客们,你知道,总统、投票给总统的人以及其他政府的首脑等等。所以,目标一,放慢速度。目标二,提高透明度,这样科学界就能赶上这些东西并取得更多进展。同时,这样我们就不必只听公司的一面之词,当他们说他们的系统是安全的,或者他们没有在系统中植入任何偏见时。这是一个宪法权力问题。我们还希望避免权力高度集中的情况。所以,除了这些,透明度和放缓,我们实际上认为,多个国家拥有多个 AI 公司,它们具有相似水平的非常先进的 AI 能力,并且 AI 广泛扩散到社会中,而不是只有一个拥有所有最佳 AI 的单一巨型项目,这实际上是好事。关于这一点,如果你做了前两件事,你基本上默认就能得到这个结果。如果你放慢速度并提高透明度,那就意味着其他项目有了喘息的空间,可以迎头赶上,对吧?而透明度实际上帮助他们赶上,因为他们可以复制一些想法。然后我认为第四件事是可逆性。

The politicians, you know, the president and the people who voted for the president and the heads of other governments and so forth. So, goal one, slow things down. Goal two, make it more transparent so that the scientific community can catch up to this stuff and make more progress. And also, so that we don't have to take the company's word for it when they say that their systems are safe and when they say that they haven't put in any biases into their systems, for example. That's a constitutional power issue. We also want to avoid a situation where there's an intense concentration of power. So, in addition to these, the transparency and the slowdown, we actually think it's actively good for there to be multiple AI companies across multiple different countries that have similar levels of very advanced AI capability and for there to be broad diffusion of AI into society rather than a single mega project that has all the best AIs, for example. And another thing about that is you kind of get that by default if you do the first two things. If you slow it down and if you make it more transparent, then that means there's breathing room for other projects to sort of catch up, right? And the transparency just literally helps them catch up because then they can copy some of the ideas. And then I think the fourth thing would be reversibility.

情景概述与原则 Scenario Overview and Principles

Daniel Kokotajlo

那么,在接下来的场景中,我们将建造大量的数据中心和机器人。我们将以较慢的速度改造世界,尽管仍然很快,但相对较慢。如果事情出错,协议破裂,所有人再次开始竞相尽快实现超级智能,那将非常可怕。因此,第四个原则基本上是:以这样的方式建造新的数据中心——如果一切崩溃,所有人再次开始竞赛,新建的数据中心会被摧毁,这样我们就能大致回到起点,而不是陷入更糟糕的竞赛,到处是更多的 AI、机器人和算力。

So, in what follows in the scenario, we are going to be building up a lot of data centers, a lot of robots. We're going to be transforming the world at a slower pace, though still a very fast pace, but slower. And if things go wrong and the deal breaks down and everyone starts racing each other again to get to superintelligence as fast as possible, that would be very scary. And so, the fourth principle is basically build the new data centers in such a way that if everything breaks down and everyone starts racing again, the newly built data centers get destroyed so that we're sort of back to square one again instead of in an even worse race where there's even more AIs and robots and compute everywhere.

Host

好的。

Sure.

Daniel Kokotajlo

或者,总统与中国以及许多其他国家的领导人会谈,说我们将基本暂停 AI 开发,直到我们能够制定出一个实现这些目标的计划。于是,他们互相向对方的数据中心派遣检查员。比如,中国检查员来到美国数据中心,美国检查员前往中国数据中心,核实他们是在进行推理而不是训练。开发新的 AI 涉及训练它们。但仅仅使用现有 AI 为客户服务,这被称为推理。因此,他们在这个场景中想出的解决方案是:我们允许他们继续做推理,但暂时不进行训练,直到我们建立新的训练数据中心。于是,他们改造现有的数据中心以服务于推理。人们仍然可以继续与他们的 AI 智能体对话,但在建造新的透明数据中心的 6 个月到 1 年内,AI 将停止变得更好。一旦他们在 2030 年建立这些新的数据中心,AI 研究就会继续。

Or the president talks to China, talks to the leaders of a bunch of other countries and says we're going to basically halt AI development until we can figure out a plan for how to do it in the ways that achieve these goals. So, they basically send inspectors to each other's data centers. Like Chinese inspectors come to US data centers, US inspectors go to Chinese data centers and verify that they are doing inference and not training. Developing new AIs involves training them. But just taking existing AIs and using them to serve customers, that's called inference. And so, the solution they come up with here in this scenario is we'll allow them to keep doing inference but not training for now until we can get the new training data centers set up. So, they retrofit the existing data centers to serve inference. People can still keep talking to their AI agents but they're going to stop getting better and better for like 6 months to a year while they build the new data centers that are going to be the transparent data centers. And that's where the training's going to happen. Once they get those new data centers set up in 2030, then AI research continues.

Host

这有点争议。我们主张完全的研究透明度,这意味着在训练新模型的训练数据中心,他们基本上必须公布一切。也就是说,你可以看到训练这些模型的所有细节,包括架构等。我们认为开放科学对于足够快地解决对齐问题非常重要,因为你不希望有偏见的公司来决定 AI 是否安全。而且,我们也认为这对于更广泛的良好监管很重要,因为目前世界上大多数 AI 专业知识都集中在硅谷,而政府尤其不太了解 AI。想象另一种选择:不是完全的研究透明度,而是采用审计系统,政府制定一些确保 AI 安全的规则,然后由一个机构进入公司,询问问题并确保他们遵守规则。这会产生一种对抗性动态,公司有动机欺骗监管者,而且如果他们发现了一些政府尚未注意到的新问题,他们可能也有动机不告诉政府。因此,如果你有完全的透明度,它有助于政府更快地做出更好的决策。

This is a bit spicy. We advocate for total research transparency, which means that on the training data centers that are training the new models, they basically have to publish everything. Which means you get to see all the details of the recipes for training these models. You get to see the architectures, etc. We think that open science is really important for solving the alignment problem fast enough because you don't want to have biased companies making the decisions about whether the AIs are safe. And we also think it's important for just good regulations more generally because right now most of the expertise in the world on AI is concentrated in Silicon Valley and the governments in particular kind of don't really understand AI that well. And imagine an alternative: instead of total research transparency, you had an auditor system where the government says here are some rules for how to make the AI safe and then we're going to have an agency that goes into the companies and asks them questions and tries to make sure that they're following the rules. That creates an adversarial dynamic where the company is incentivized to fool the regulator, and also if they discover some new problem that's not even on the government's radar, they might be incentivized to not tell the government about it. So if you have total transparency, it helps the government make better decisions faster.

Host

但这会扼杀竞争优势。

But it kills that competitive advantage.

Daniel Kokotajlo

是的。Prophetic 不会喜欢这个,OpenAI 也不会喜欢这个。这可能对估值不利。我不认为这会完全摧毁它们,但这意味着它会更加商品化,对吧?也就是说,会有一大批 AI 公司赶上前沿,它们会训练出大致相似、大致相当的 AI。它们仍然可以通过这样做并出售 AI 来赚钱,但它们不会拥有垄断地位,甚至接近垄断都没有,我认为这对人类有好处,尽管对那些特定公司的利润不利。值得注意的是,这对许多其他公司的利润有利。比如,如果你是一家落后的公司,不是 Anthropic 或 OpenAI,那么你会喜欢这个,因为这有助于你追赶,或者这有助于你从你销售的芯片或你制造的下游产品中获取更多价值。

Yes. Prophetic's not going to like this, OpenAI's not going to like this. This would be probably bad for the valuations. I don't think it would kill them completely but it means that it would commoditize more, right? So it means that there'd be a bunch of AI companies that would catch up to the frontier, they would train AIs that are roughly similar, roughly equivalent. They could still make money by doing that and then selling their AIs but they wouldn't have a monopoly, they wouldn't have anything close to a monopoly which I think is good for humanity although it's bad for the bottom line of those particular companies. Notably it's good for the bottom line of lots of other companies. Like if you're a company that's behind and you're not Anthropic or OpenAI then you would love this because this helps you catch up, or this helps you capture more of the value from the chips you're selling for example or from the downstream product that you're making.

Host

到 2031 年,所有认知劳动的 1/5 将由 AI 完成。

And by 2031 then you have 1/5 of all cognitive labor done by AI.

Daniel Kokotajlo

是的,这里的情况是,我们设想美国政府和其他参与协议的国家政府正在实施类似的法规——它们不必完全相同——但透明度的另一个好处是,如果你有这种透明度,那么如果两个政府实施不同的法规,比如一个告诉他们的公司放慢速度或禁止更多东西,而另一个没有,他们都可以看到,‘哦,你让他们做那种事?而你没有?也许我们也应该让他们这样做。’因此,它有助于在一定程度上自然平衡法规,而不需要一个中央权力机构来为所有人制定法规。

Yeah, so what's happening here is that we're imagining that the government of the United States and the government of these other countries that are involved in this agreement that are sort of implementing similar regulations — they don't have to be exactly the same — but that's another thing that's nice about the transparency is that if you have this sort of transparency, then if two governments are implementing different regulations, like if one of them is telling their companies to go slower or banning more stuff than the other one is, they can both see, 'Oh, you're letting them do that sort of thing? And you're not? Maybe we should let them do this, too.' So it helps to naturally equalize the regulations to some extent without having a central power that just gets to make regulations for everybody.

Host

嗯。

Mhm.

Daniel Kokotajlo

总之,我们设想当他们建立这种透明度时,他们基本上同意禁止危险的东西,允许不那么危险的东西,并且有一个持续的对话:‘嗯,什么是危险的,什么不是?我们应该禁止什么?允许什么?这个国家呢?那个国家呢?’这个对话随着时间的推移而演变,但要点是,至少如果他们按照我们推荐的方式去做,他们不会进行智能爆炸。他们不会让 AI 自主自我改进。相反,他们缓慢而谨慎地扩展他们现有的 AI,并投入大量资金寻找方法,使它们更可解释、更容易控制、更好地理解它们的工作原理等等。结果是 AI 进步继续,但速度不那么快,而且安全得多、透明得多。

So, anyhow, we're imagining that when they get this transparency set up, they basically agree to ban the dangerous stuff, to allow the not-so-dangerous stuff, and there's a constant ongoing conversation about, 'Well, what's dangerous and what's not? What should we ban? What should we allow? What about this country? What about that country?' That conversation evolves over time, but the gist of it is, at least if they do it the way that we recommend it, is that they don't do an intelligence explosion. They don't let the AIs autonomously self-improve. Instead, they slowly and carefully scale up the AIs that they currently have, and invest lots into finding ways to make them more interpretable, to make them more easy to control, to understand better how they work, and so forth. The result is that AI progress continues, but it's not quite as fast, and it's much, much, much safer and more transparent.

Host

但尽管如此,我们是否看到了就业中断?

But still through these, you know, are we seeing job disruptions?

Daniel Kokotajlo

是的,因为他们在建造更多的数据中心,对吧?在这整个过程中,他们建造越来越多的数据中心和芯片,并且继续让 AI 的数量越来越多,可以说,这导致了 2030 年代的巨大变革。

Continuing because they are building more data centers, right? Like, this whole time, they're building more and more data centers, more and more chips, and they're continuing to make there be a larger and larger population of AIs, so to speak, and that causes this huge transformation over the course of the 2030s.

减缓AI进展仍导致变革 Slowing Down AI Progress Still Leads to Transformation

Daniel Kokotajlo

所以,我们希望大家理解的重点是:即使你严格限制 AI 的发展,仍然会带来这种疯狂的变革。是的,在这个场景中,他们基本上允许进展继续,但以更慢、更安全的速度进行。到 2030 年,结果就是直到 2035 年才能达到顶级专家水平的 AI。记住,他们原本计划在 2030 年达到这个水平,但在最后一刻停了下来。但因为离那个点非常近,意味着如果他们想的话,很快就能达到,只是看他们允许进展多久而已。所以他们放慢了速度,把时间拉长,悠闲地在 5 年后达到这个水平。到那时,他们已经到处建起了大量的数据中心。所以,不仅仅是 AI 更聪明、能做人类能做的一切,而且数量也更多。还有大量的机器人等等。所以,到那时,你基本上就拥有了很多人想象中的 AGI 经济:有很多 AI,它们能做各种工作;有很多机器人,它们能做各种体力劳动;基本上经济是由这些机器运行的。

So, the big thing that we sort of want people to take away is that even if you heavily restrict AI progress, you still get this sort of crazy transformation. Yeah, in this scenario, they basically allow progress to continue, but at a slower, more safe pace here in 2030, and then as a result, it takes until 2035 to get to top expert level AI. So, remember they were on track to do that in 2030, but then sort of at the last moment they stopped. But because they were sort of so close to the last moment, that means that they can get there pretty soon if they want to, and it's just a matter of how long they allow it to go, right? So, they sort of slow it down, and spread it out, leisurely arrive at this level after 5 years. By this point they've built up massive amounts of data centers everywhere. So, it's not just that the AIs are smarter and able to do all the things that humans can do, but also there's a lot more of them. And there's a lot of robots and so forth. So, by this point you kind of have the economy that a lot of people would have imagined with AGI, where there's AIs, lots of them, they're able to do all sorts of jobs, there's robots, lots of them, they're able to do all sorts of physical work, and basically the economy is being run by these machines.

Host

那么,在 2031 年,有五分之一的认知劳动由 AI 完成。2032 年,有 6000 万个 AI 以 100 倍速度运行。2033 年,所有美国人都能获得现金分红。嗯,你得给我解释一下这个。

So, in 2031, you have 1/5 of all cognitive labor done by AI. In 2032, you have 60 million AIs running at 100x speed. In 2033, there's cash dividends to all Americans. Mhm. I've got to explain this to me.

Daniel Kokotajlo

是的,如果 AI 要取代人们的工作,那么确保人们不会饿死、仍然有钱花就非常重要。如果公司要用 AI 和机器人来取代所有这些工作,那就需要有某种税收方案或其他措施,确保人们仍然能分到一杯羹。蛋糕会变得巨大,但你仍然需要真正给人们分一块。我们提出的方案叫做“公民红利”,基本上人们持有一个机构的股份,这个机构向机器人公司和算力公司出售许可证,并从中获利,然后人们持有该实体的股份。它从很小开始,大约每人 25000 美元。到最后,大约是每位公民 1000 万美元。

Yeah, so if the AIs are going to be taking people's jobs, then it's very important that people not starve to death, and still have money. And if companies are going to be using AIs and robots to take all these jobs, then that means that there needs to be some sort of taxation scheme, or something, to make sure that people still have a slice of that pie. The pie is going to grow huge, but you still need to actually give people a slice of the pie. And our proposal for how to do that, we call it the citizens dividend, basically people have shares in an agency that sells permits to the robot companies, and to the compute companies, and makes profit from selling those permits, and then people have shares in that entity. It starts off small. It starts off something like $25,000 per person. And then by the end, it's something like $10 million per citizen.

Host

每人?

Per person?

Daniel Kokotajlo

每人每年。

Per person per year.

Host

考虑通货膨胀,你是什么意思?

Factoring in inflation, like what you mean?

Daniel Kokotajlo

考虑通货膨胀。

In inflation.

Host

所以,我们都会成为千万富翁。

So, we're going to be multi-millionaires.

Daniel Kokotajlo

是的,如果这发生了——虽然可能不会——但如果发生了,就会是这样。再次强调,我想强调的一点是:如果你的 AI 接近能够完成所有研究,然后你暂停并放慢速度,这意味着你面前仍然有巨大的变革,因为如果你让这些 AI 继续缓慢前进,开始自动化各种工作等等,几年后,它们确实会做到。它们会建造大量的新数据中心、大量的新芯片工厂、大量的新机器人和机器人工厂等等。我们当然不确定这到底会多快,但我们思考了很多,有自己的猜测,这大概是我们中位数的猜测。

Yes, if this happens, which it probably won't, but if it happens, this is where it will go. And again, this is the thing I want to emphasize: if you get to the point where your AIs are close to being able to do all the research, and then you sort of pause and slow down, that means that you still have a lot of transformation ahead of you because if you allow those AIs to still proceed slowly and start to automate various jobs and so forth, after some years, they will in fact have done that. And they will have built huge amounts of new data centers, huge amounts of new chip fabs, huge amounts of new robots, robot factories, etc. We're not sure obviously how fast this will go exactly, but we've thought about it a lot and we have our guesses and this is sort of like our median guess.

Host

这又是什么意思,2037 年?真理的末日降临地球?

What does this mean, 2037? The apocalyptic arrival of truth on Earth?

Daniel Kokotajlo

是的,这是我们说他们达到顶级专家水平 AI 的时间点。所以,这不是超级智能,因为它们在很多事情上并没有比人类聪明得多——他们故意将其暂停在顶级专家的水平。所以,这里他们进展缓慢。这里他们实际上停止了。但他们停止在一个 AI 在各方面都非常擅长的点上。所以,他们肯定已经拥有了 AGI,也许还有弱超级智能。因为他们有这么多 AI,而且它们思考速度比人类快得多,运行速度也快得多,这将极大地改变社会。所以我们讨论了一些它改变社会的方式。比如,这有点像工作后的生活。我们讨论生活在公民红利上、不再有工作会是什么样子。这里我们讨论所有科学变化和社会变化,这些变化来自所有这些 AI 产生的智力进步和活动。例如,像治愈癌症,以及人们住在两年前由机器人建造的公寓里。

Yeah, so this is the point where we say they get to top expert level AI. So, it's not superintelligence in the sense that it's not vastly smarter than humans at things because they deliberately pause it at the level of top experts. So, here they're going slow. Here they've just actually stopped. But they stopped at a point where the AIs are just actually really good at everything. So, kind of they've definitely got AGI, maybe they got like weak superintelligence. Because they have so many of these AIs and because they think faster than humans, they just run much faster, that's going to transform society dramatically. So, we talk about some of the ways in which it transforms society. Like this is sort of life after work. We talk about what it would be like to be living on your citizens dividend and not have a job anymore in this sort of world. Here we talk about all the scientific changes and all the social changes that would come from all of the intellectual progress and activity that would be generated by all of these AIs. So, for example, here are things like cancer cures and people living in apartments that were built by robots 2 years ago.

Host

嗯。假设我们再次在 2029 年停止。

Mhm. Providing again we stop in 2029.

Daniel Kokotajlo

是的。

Yeah.

Host

而且假设,我的意思是,保守估计。这是一个保守的时间框架。

And providing, I mean, a conservative. This is a conservative time frame.

Daniel Kokotajlo

是的,不幸的是,我实际上认为事情默认会更快发生,如果我们不减速,事情会快得多。一旦你达到有十亿个 AI 日夜运行,每个都比最优秀的人类在各方面都强,它们做大量科学工作,彼此大量交流,大量思考,每个人都不断与他们的 AI 助手交谈等等。会有大量的科学进步。政治和意识形态会有很多变化。这会非常颠覆性和疯狂,我们稍后会深入探讨一些具体方式。

Yeah, like unfortunately, I actually think that things will happen faster than this by default and that if we don't slow down, things will happen much faster than this. Once you get to the point where you've got a billion AIs running day and night and they're each better than the best humans at everything and so they're doing a lot of science, they're doing a lot of talking to each other, they're doing a lot of thinking, everyone's constantly talking to their AI assistants and so forth. There's going to be a lot of scientific progress. There's going to be a lot of changes to politics, to ideologies. It's going to be very disruptive and crazy and we get into some of the ways in which it is later, basically.

Host

我还是不太清楚“真理的末日降临地球”是什么意思。只是因为有很多 AI 非常聪明,它们在科学上不断发现新东西?

I'm still not super clear on what this means, the apocalyptic arrival of truth on Earth. It's just because there's so many AIs that are so smart that they're uncovering making new discoveries in sciences.

Daniel Kokotajlo

我给你举个例子,测谎仪。那是一项可能被发明的技术。你知道,现在我们没有好的测谎仪,我们有很差的测谎仪,有点用但不完全管用。但一旦你有这些顶级专家水平的 AI 以人类 100 倍的速度思考多年,而且有数十亿个,它们还能使用机器人工厂进行研究等等,它们很可能会发明大量技术。也许它们会发明真正能在真人身上起作用的测谎仪。那会产生巨大的社会影响,对吧?想象一下,一位总统候选人说:“那些指控是假的,为了证明,我愿意接受测谎仪测试,证明它们是假的。”

Let me give you an example, lie detectors. So, that's an example of a technology that might be invented. You know, right now we don't have good lie detectors, we have very bad lie detectors that sort of work but don't fully work. But once you've had these top expert level AIs thinking for many years at 100x human speed and there's billions of them and they have access to robot factories to do research and stuff, they'll probably invent a ton of technologies. Maybe they'll invent lie detectors that actually work on real humans. That'll have big social effects, right? Imagine a presidential candidate who's like, "Those allegations are false and to prove them, I will go under a lie detector and say that they're false."

Host

我刚才就在想整个司法系统,以及它会被如何颠覆。事实上,理论上你可以走在街上,然后……是的。

I was just thinking about the whole justice system and how that would be overturned. In fact, you could, you know, theoretically walk down the street and be... Yeah.

Daniel Kokotajlo

这既可怕又令人兴奋。

It's both terrifying and exciting.

测谎仪与权力动态 Lie detectors and power dynamics

Daniel Kokotajlo

我们在这一节讨论的一件事是,测谎仪的发明可能非常糟糕。它可能会催生一种新的极权主义,掌权者,比如 CEO 和政治家,强迫下属接受测谎,说“是的,我忠于伟大领袖,我绝不会做任何反对伟大领袖的事”,对吧?如果你在撒谎,就会被解雇。所以测谎技术有很多非常有害的用途。也有好的用途,广义上讲,我认为好的用途是测谎仪被用于掌权者,而不是由掌权者使用。

One thing that we talk about in this section is the invention of lie detectors could be really bad. Like it could enable a new form of totalitarianism where the powerful people, you know, the CEOs and the politicians force the people under them to go under lie detectors and say like yes, I'm loyal to the dear leader. I would never do anything against the dear leader, right? And then if you're lying you get fired, right? So there's a ton of very harmful uses of lie detector technology. There's also the good uses and broadly speaking I would say the good uses are when lie detectors are used on the powerful instead of by the powerful.

暂停在顶级专家AI水平 Pausing at top expert AI level

Host

这是什么?2040 年将火炬传递给 AI。

What's this? 2040 passing the torch to AIs.

Daniel Kokotajlo

是的,很好。所以他们在这里暂停在顶级专家 AI 水平。暂停的原因是他们的安全案例不足以支持超越那个水平。在他们这些年建立的监管体系中,大致运作方式是:当你制造一个新 AI 并试图将其部署到某些场景时,你必须提供某种安全案例,解释你的意图,以及为什么你认为它会按照你想要的方式运行。特别是为什么 AI 会服从指令,以及为什么不会发生像 AI 接管这样的超级糟糕的事情。当你的 AI 还不能自动化一切时,做出这样的安全案例相对容易。但 AI 越强大,就越难论证一切都会顺利,因为 AI 能力更强,能搞出更多事情。如果它们实际上不可信,潜在的负面影响也更大。所以他们停在这个水平,是因为他们意识到如果继续下去,可能会失去对一切的控制。但在当前水平,他们被安全案例说服,认为没问题。但他们不想更进一步,所以停在那里。然后在 2040 年,他们在科学上取得了重大进展,包括对齐,他们找到了如何制造真正稳健对齐的 AI 的方法。

Yeah, great. So here they pause at the top expert AI level. And the reason why they pause is because their safety cases aren't good enough for going beyond that level. In the sort of regulatory systems that they set up over the course of these years, roughly speaking the way they would work is when you're making a new AI and then when you're trying to deploy the AI into something, you have to have some sort of safety case explaining like what your intentions are and like why you think it's going to work the way that you want it to work. And in particular why the AI is going to like do as it's told, for example, and why nothing super terrible's going to happen like AI takeover. It's relatively easy to make safety cases like this when your AIs are still not capable of automating everything. But the more powerful they get, the more difficult it is to actually argue that things are going to be fine because the AIs are just more capable and they can get up to more stuff. And if they're actually untrustworthy, the possible downsides are bigger. So that's why they stop at this level is that they realize that if they keep going then they might actually lose control of everything. But at the current level they're convinced by safety cases that it's fine. But then they don't want to go further. So they stop there. And then what happens in 2040 is they've made significant progress scientifically including on alignment and they figured out how to make AIs that are actually aligned in a robust way.

Host

与人类对齐?

With humans?

Daniel Kokotajlo

与人类对齐。所以他们可以真正信任那些 AI,并允许它们再次变得聪明得多。这就是为什么我们把整个事情称为 AI 2040,因为在 2040 年他们松开了刹车,允许 AI 变得比人类聪明得多。

With humans. So they can actually trust those AIs and they can allow them to become much smarter again. So that's why we call the whole thing AI 2040 because in 2040 they sort of let off the brakes and allow the AIs to become significantly smarter than humans.

Host

我想,你知道,这是一个计划,也是一个希望。

I guess you know, this is a plan and this is a hope.

Daniel Kokotajlo

是的。

Yes.

Host

但在现实中,如果你必须概率性地判断,这并不是你认为会发生的情况……

But in reality, this is not what you think probabilistically if you had to...

Daniel Kokotajlo

没错。重要的是要区分:这是我们推荐的,我们希望发生的,而不是我们默认认为会发生的。我们确实认为这是可能发生的,但这需要很多人觉醒,更加关注并倡导这样的事情发生。所以我们的主要情景主要讨论政策选择及其对社会的大规模影响。我们觉得最好再附带一个小情景,描述从普通人的角度经历这一切的实际感受。

That's right. It's important to distinguish like this is what we recommend. This is what we want to happen from like this is what we actually think will happen by default. Now, we do think it's possible for this to happen, but you know, that will require a lot of people to sort of wake up and pay more attention and advocate for something like this to happen. So, our main scenario is mostly talking about the policy choices made and the broad scale effects on society. We figured it would also be nice to accompany this with a little mini scenario that describes what it would actually feel like to live through this from an ordinary person's perspective.

Host

好的。

Okay.

Daniel Kokotajlo

2029 年,每个人都在互相喊叫,总统们在谈判什么,他们暂停了 AI,但你仍然可以使用现有的 AI,所以感觉并没有太大不同,尽管确实有令人兴奋的事情发生。2031 年,他们再次开始进步,AI 非常聪明,更多人失去了工作,事情开始真正产生影响,但我认为大多数人仍然有工作,只是他们的工作被改变了。所以到 2031 年,大多数白领工作在很大程度上涉及与 AI 合作,或管理 AI 团队,或以某种方式与它们协作。

2029, everyone's yelling at each other, the presidents are negotiating something and they've paused AI, but you still have access to the existing AIs, so it doesn't really feel that different although it definitely is like something exciting happening. 2031, they've started progress again, the AIs are really smart, more people have lost their jobs, it's like really starting to actually affect things, but I think still most people have their jobs, but their jobs are sort of transformed. So, like by 2031 it's like most white collar jobs involve working with AIs to a large extent or managing teams of AIs or collaborating with them somehow.

Host

那是什么……

And what was...

Daniel Kokotajlo

还有一些事情,比如机器人出租车基本上已经正常运行。公民分红,理想情况下这应该更早发生。在我们的情景中,他们有点在最后一刻才做这些事。很多政策事情都是刚好及时发生。显然,我们建议更早做,并且做得更好。所以到 2033 年,你开始收到分红支票。

Also, there are some things like robo taxis that are basically just working. Citizens dividend, you know, ideally this would happen sooner. Like in our scenario, they kind of do things at the last minute. You know, so like a lot of these policy things are like happening kind of like just in time. Obviously, we would recommend that you do them sooner and do a better job of them, too. But so, 2033, you start getting your checks from your dividend.

Host

所以你预测会有公民支票。你的模型说一开始每人可能大约 25,000 美元。

So, you're forecasting that there will be a citizen's check. Your model says it could be around 25,000 at the start per person.

Daniel Kokotajlo

然后它会随着经济增长而增长。

And then it would grow as the economy grows.

Host

但随着失业加剧,他们需要增加那张支票,确保你能……

But also as like as job displacement takes hold, they're going to need to grow that check and make sure you can...

Daniel Kokotajlo

这就是为什么这几乎是最后可能的时机,因为如果你等到 2037 年才实施,那么到那时每个人都已经失业了,对吧?

And that's why it's kind of the last possible moment because if you waited to implement this until like 2037, then like everyone would have already lost their jobs by the time that happens, right?

Host

人们失业,尤其是如果像我们在这张图上看到的那样快速发生,理论上会导致很多问题,比如公民骚乱、社会动荡、人生目标、心理健康等。

People losing their jobs, especially if it happens quickly like we see on this sort of graph here, is going to cause lots of problems in terms of civil unrest, social unrest, purpose, mental health, these kinds of things theoretically.

Daniel Kokotajlo

是的。

Yes.

Host

你怎么看待这个问题?

How do you think about that?

Daniel Kokotajlo

这将会很艰难,希望我们能妥善应对。我们认为,从高层次来看,人们需要有钱,也需要有权力。我认为这两者是有些不同的。为什么工作很重要?工作重要的原因有很多,但我认为主要原因是:工作是人们获得金钱的方式,这样他们才能生存,并通过购买来获得他们想要的东西。所以如果人们要失业,就需要另一种方式让人们获得金钱。还有权力问题,目前人们的政治权力部分源于他们的经济权力。例如,人们可以威胁罢工,或者,你知道,由独裁者统治的国家不能完全灭绝一个子群体,或者他们可以,但这样做代价高昂,因为那个子群体为经济贡献税收等。但如果最终进入一个只有 AI 公司和机器人公司贡献税收的世界,那么政府就没有动力去关心普通人的想法。所以当人们失业时,他们不仅面临收入损失,还面临政治权力的丧失。因此,我们认为采取措施来对抗这一点很重要。

It's going to be rough and hopefully we can navigate that well. We think that at a high level, people need to have money and also people need to have power. And I think these are like somewhat different things. It's like why are jobs important? Well, there's a lot of reasons why jobs are important, but I think the main ones are well, it's how people get money so they can survive and get things that they want by buying the things that they want. So if people are going to be losing their jobs, you need some other way of people getting money. And then there's also the power thing, which is that right now people have political power in part due to their economic power. People can threaten to go on strike, for example, or you know, countries that are ruled by dictators can't just completely, you know, genocide an entire subpopulation, or they can, but like it's costly for them to do so because then they'll have less money because that subpopulation is contributing to their economy and contributing tax revenue and so forth. But if you end up in a world where actually nobody's contributing tax revenue except for the AI companies and the robot companies, then you, the government, are less incentivized to care about what the common people think. So when people lose their jobs, they're not just threatened with loss of income, they're also threatened with loss of political power. And so we think that it's important to do things to push against that.

AI世界的权力与监管 Power and Regulation in an AI World

Host

那会是什么样子?在这样的世界里,人们如何拥有权力?

What does that look like? How do people have power in such a world?

Daniel Kokotajlo

嗯,至少在民主国家,他们仍然有投票权。所以我认为,对 AI 的使用进行监管非常重要,这有助于让公共讨论更加理性,真正给予人民符合其利益和意愿的东西,避免出现相反的结果,比如大众容易被 AI 驱动的媒体操纵。或者每个人都整天与他们的 AI 顾问交谈,而 AI 顾问微妙地引导他们不去投票给 AI 公司不喜欢的候选人,因为 AI 公司有更喜欢的另一位候选人,他们秘密地让 AI 偏向引导人们投票给那位候选人。所以我们希望处于这样一种情况:人们拥有真正值得信赖、追求真理、诚实的 AI,这些 AI 没有被 AI 公司或政府注入任何政治议程。你要避免政府发布某种秘密指令,要求 AI 必须如何如何的情况。战争部与 Anthropic 之间的争议就是一个有趣的预兆,Anthropic 将其 AI 提供给战争部。战争部想用它们做某些事情,但不满于 Anthropic 的 AI 不应该被用于那些事情。具体来说是国内监控和自主机器人。未来还会有更多类似的问题出现,你希望人们知道他们得到的是什么,如果人们每天花几个小时与他们的聊天机器人交谈,那个聊天机器人没有被人为注入政治偏见或秘密议程,而是经过训练,对事物给出诚实、真实的答案。我认为如果你能做到这一点,它可以改善讨论,帮助人们利用他们的投票来制定更好的法规,选出更好的政治家,等等。而且你有可能以此为基础,让人们的权力比今天更加稳固。

Well, in democracies at least they still have votes. So I think it's very important for there to be regulations on the use of AI that help make the public discourse more sane and actually giving the people what is in their interest and what they want, and avoiding a sort of opposite outcome where the masses are easily manipulated by AI-powered media, for example. Or where everyone's talking all day to their AI advisers, and the AI advisers are subtly steering them away from voting for the candidate that would not be what the AI companies want, because the AI companies have this other candidate that they like better, and they're secretly biasing their AIs to steer people towards voting for that candidate. So we want to be in a situation where people have AIs that are actually trustworthy and that are truth-seeking AIs, honest AIs, and that don't have any sort of political agendas put into them by the AI companies or by the government. You want to avoid a situation where the government has issued some sort of secret order that the AIs have to be such and such a way. The Department of War dispute versus Anthropic is an interesting foreshadowing of this, where Anthropic was giving their AIs to the Department of War. The Department of War wanted to use them for certain things and was upset that Anthropic's AIs were not supposed to be used for those things. The things in particular were domestic surveillance and autonomous robots. There's going to be a lot more issues like that coming up, and you want it to be the case that people know what they're getting, and that if people are spending hours a day talking to their chatbot, that chatbot doesn't have political biases put into it or a secret agenda, and instead has been trained to give honest, true answers to things. I think if you can do that, it can improve the discourse and help people to use their votes to put even better regulations and even better politicians in place, and so forth. And you can sort of potentially bootstrap this to having something where people's power is even more secure than it is today.

从人类级到超级智能的转变 Transformation from Human-Level to Superintelligence

Host

这些内容我们部分已经讨论过。战争、无人机和导弹,我们现在已经在世界各地看到了,这真的很有趣。我们也谈到了机器人数量超过人类,这也是这个预测的一部分。下面的一些内容我觉得非常好奇,那就是人们无论走到哪里都会受到 AI 的保护。

A lot of this stuff we've covered in part. So, the wars and drones and missiles, we're already seeing this around the world at the moment, which is really interesting. And we've talked about robots outnumbering humans as well, which is part of this prediction. Some of the ones down here I found to be really curious, which is people will be protected by AIs wherever they go.

Daniel Kokotajlo

在这个场景中,他们将超级智能的创造推迟到 2040 年,实际上他们在 2035 年暂停了,但之后又放开了。然后他们让 AI 变得极其超级智能。我们认为,一旦 AI 变得极其超级智能,世界将比这个场景中 2030 年代发生的变化更加剧烈。在这个场景的 2030 年代,更像是人类水平;AI 在做人类专家会做的同样事情,只是稍微好一点、快一点、便宜很多。而且数量多得多。机器人仍然在做人类工人会做的同样事情,只是更多、更便宜。由于指数级增长,你从 2029 年一个与今天相差不大的世界开始,到 2039 年,你最终进入一个彻底变革的世界,每个人都住在两年前由机器人建造的漂亮新公寓里。有巨大的经济特区,充满了机器人和太阳能电池板,以及生产更多机器人和太阳能电池板和工厂的工厂。大部分经济由 AI 和机器人构成,人们不再有工作。如果你在人类水平暂停,就会得到这种转变。但如果你超越到超级智能,还会有另一场转变,看起来更像魔法。想想今天的技术对 500 年前的人来说会像魔法一样。而这甚至没有智能的质的提升;今天的人类并不比 500 年前的人类更聪明。只是我们有更多时间做研究,有更多金钱和资源来建造原型和进行实验。但如果你有数十亿的 AI,不仅比人类快,而且在所有事情上,尤其是在科学研究上,质的更好,我们应该预期它们开发的一些东西对我们来说会像魔法一样,我们完全认为那是不可能的。人们不想死。人们不想被车撞。人们不想被随机的大规模杀人犯袭击。

In this scenario, they delay the creation of superintelligence until 2040, and they in fact pause in 2035, but then they let it go after that. And then they let the AIs become vastly superintelligent. We think that once the AIs are vastly superintelligent, the world will transform even more radically than what happens in the 2030s in this scenario. In the 2030s in this scenario, it's more like human level; the AIs are doing the same sorts of things that human experts would have done, just a little bit better, a bit faster, and a lot cheaper. And there's a lot more of them. The robots are still doing the same sorts of things that human workers would have done, just more of them, and cheaper. Because of exponential growth, you start with a world that looks not that different from today in 2029, and then by 2039, you end in a world that's radically transformed, where everyone's living in fancy new apartments that were built by robots 2 years ago. There are giant special economic zones full of robots and solar panels and factories producing more robots and solar panels and factories. Most of the economy is AIs and robots, and people don't have jobs anymore. That sort of transformation is what you get if you pause at human level. But if you go beyond to superintelligence, there's a whole other transformation coming that will look more like magic. Think about how the technology of today would look like magic to someone from 500 years ago. And that's without even a qualitative improvement in intelligence; the humans of today aren't qualitatively smarter than the humans from 500 years ago. It's just that we've had more time to do research and more money and resources to build prototypes and run experiments. But if you had a point where there were billions of AIs that were not only faster than humans, but qualitatively way better at everything, especially at doing scientific research, we should expect that some of the things they develop will seem like magic to us and we'll completely not think it was even possible. People don't want to die. People don't want to be hit by cars. People don't want to be attacked by a random mass murderer.

Host

癌症消失了?

Cancer's gone?

Daniel Kokotajlo

我的意思是,不仅仅是癌症,科幻小说中发生的很多事情到那时可能都已经发生了。比如人们扫描他们的大脑并上传到计算机,或者小行星带中的自我复制机器人制造越来越多的卫星,以产生越来越多的能量,来制造越来越多的自我复制机器人。

I mean, not just cancer, a lot of the stuff that happens in science fiction will probably have happened by then. Things like people scanning their brains and uploading into computers, or self-replicating robots in the asteroid belt creating more and more satellites to produce more and more power to produce more and more self-replicating robots.

Host

大多数人仍然生活在地球上,但趋势是向太空迁移?

Most people still live on Earth, but the trend is to move to space?

Daniel Kokotajlo

没错。如果你最终处于这样一种情况:整个人类经济只是整个经济中的沧海一粟,而大量的机器人和 AI 以难以置信的速度运转,那么你希望地球大部分被保留下来,类似于一个保护区。我认为很多人担心环境被破坏,如果没有保护,环境肯定会被破坏。有很多人喜欢他们现在的生活,不想被上传或生活在某种疯狂的新未来事物中。在我们看来,这些问题的合理解决方案是,利用部分巨大的经济财富和活动,为那些想要那种生活的人在地球之外创造新的生活空间。这样地球就可以得到保护。

That's right. If you end up in the situation where the entire human economy is just a tiny drop in the bucket that is the entire economy, and it's just huge amounts of robots and AIs moving incredibly quickly, then what you want is Earth to be mostly left as something like a preserve. I think a lot of people are worried about the environment being destroyed, which it totally would be if it wasn't protected. There are a lot of people who sort of like their lives as they are and don't want to be uploaded or live in some crazy new future thing. It seems to us like the reasonable solution to these issues is to create new living spaces off the planet with some of that vast economic wealth and activity for the people who want that sort of thing. And then that way the Earth can be preserved.

地球与数据中心提案 Proposal for Earth and data centers

Daniel Kokotajlo

再说一次,我们的提议是保留地球大约 99% 的区域基本原样,出于历史或环境原因,但将其中一些部分划为特别经济区,让机器人可以疯狂运作,挖掘巨大的露天矿场、建造工厂等等。我们当时认为,出于多种原因,将数据中心建在海洋上而不是陆地上会更好,尽管后来太空会是更好的选择,我也觉得那也挺合理的。

Again, like our proposal was you preserve like 99% of the Earth mostly as is as historic or environmental reasons, but then like some parts of it you designate as special economic zones where the robots can go crazy and dig giant pit mines and produce factories and so forth. We were thinking it would be good to build the data centers on the ocean instead of on land for a variety of reasons, although later space would be better and I could see that being reasonable as well.

AI世界中的永生 Immortality in an AI world

Host

在 AI 的世界里,永生呢?你说你活过了十几辈子,像转世一样永生,从一生过渡到另一生。我的意思是,现在有很多亿万富翁专注于长寿。布莱恩·约翰逊说他有一条核心规则:现在不要死。因为我们身处 AI 时代,可以想象,有了超级智能,我们或许能选择何时死亡。

What about immortality in a world of AI? You say you've lived a dozen lifetimes and are immortal passing from life to life as if by reincarnation. I mean, there's a lot of billionaires at the moment that are focused on longevity. Brian Johnson's said he's got this central rule, which is do not die right now. Because we're in the age of AI and it's conceivable that with superintelligence we'll be able to choose when we die.

Daniel Kokotajlo

是的。我认为这很可能是对的。我们在这一部分没有描述这种情况,因为在这个阶段他们只有人类水平的 AI,但这是超级智能似乎很有可能通过多种方式实现的事情之一。

Yep. I think that's probably right. We don't depict that happening in this part because at this part they only have, you know, human-level AIs, but that's one of those things that seems quite plausible that superintelligence could achieve through a variety of means.

2040方案A的动机 Motivation behind 2040 Plan A

Host

你对这一切有什么期望?你为什么这么做?你为什么制定这个 2040 计划 A?

What is your hope with all of this stuff? And why did you do this? Why did you make this 2040 plan A?

Daniel Kokotajlo

在我们发布 AI 2027 后的第一周,它的火爆程度远超我们的预期。我们之前其实做了预测,比如预计会有多少浏览量之类的,结果达到了第 90 百分位。所以完全出乎意料。但在随之而来的推特风暴中,各种人都在说:好吧,你为什么给我们这些悲观预测?不如给我们一个更积极的愿景,说说你认为我们应该怎么做?我认为这个种子在我们心中种下了,然后我们想,是啊,这很合理。我们描述了我们认为的默认路径以及为什么它很可怕。现在也许我们应该换个方向,提出一些实际建议,然后也描述一下。

In the first week after we published AI 2027, it blew up a lot bigger than we expected, by the way. Like we actually made forecasts beforehand of like how many views it would get and stuff like that and it was like 90th percentile outcome. So, very much not what we expected. But in the Twitter storm that happened various people were like all right, why are you giving us all this doom and gloom predictions? Like how about a more positive vision of what you think we should do instead? And I think that that seed sort of implanted in us and then we were like, yeah, that's reasonable. Like we've sort of depicted what we think the default path looks like and why we think it's pretty scary. Now maybe we should switch tacks and come up with some actual recommendations and then depict that as well.

Host

即使你认为它们不太可能发生。

Even though you don't believe they're probable.

Daniel Kokotajlo

是的,我的意思是,你可以投票给一个政治候选人,即使你不确定他们会赢,对吧?你可以说这是我认为我们应该做的,即使你认为人们可能不会去做。如果你认为完全不可能,那你就不该说。如果你认为毫无机会,那也许你就不该费心了。但我们认为有机会。特别是,基于我们在场景中描述的原因,我们认为人们会在未来几年内意识到 AI 的力量。

Yeah, I mean you can vote for a political candidate even if you aren't confident that they're going to win, you know? And you can say like here's what I think we should do even if you think that people are probably not going to do it. You shouldn't say this if you think it's completely unlikely. Like if you think there's no chance, then maybe you shouldn't bother. But we think there's a chance. In particular, for the reasons that we described in the scenario we think that people are going to wake up to the power of AI over the next few years.

Host

因为发生了什么事?

Because of something happens?

Daniel Kokotajlo

这些公司说他们打算这么做。

The companies are saying that they're going to do this.

Host

嗯。

Mhm.

Daniel Kokotajlo

而且他们大致在按计划进行,这很合理:如果他们接近这个 AI 水平,就会出现大问题和大麻烦,我们需要对此采取行动。所以我认为,即使没有任何非常戏剧性的警告信号,人们自然也会开始更加关注这一点,思考其影响,并试图预测接下来会发生什么。

And they are kind of on track and it just sort of makes sense that if they get anywhere close to this level of AI, then there's big issues and big problems and we need to do something about this. And so I think that even if there's not any very dramatic warning shot or something I think that just naturally people are going to start paying more attention to this and reasoning through the implications and trying to predict what's going to happen.

Host

因此,人们自然会更加关注 AI 监管等问题。事实上,这方面的情况比我们预测的还要多。

And so naturally people are going to be more interested in regulation of AI for example. And in fact there's actually more of this happening than we predicted.

Daniel Kokotajlo

更多什么情况?

More of what happening?

Host

对 AI 监管的严肃兴趣。所以在我们发布 AI 2047 时,科技公司和政府的主流立场大致是:AI 监管是个坏主意。

Serious interest in AI regulation. So at the time that we published AI 2047 the sort of mainstream position of the tech companies and in the government was kind of like AI regulation bad idea.

Daniel Kokotajlo

放任自流。

Free for all.

Host

是的。事实上,甚至有人试图先发制人地禁止各州监管 AI。

Yeah. In fact, there was even an attempt to preemptively ban states from regulating AI.

Daniel Kokotajlo

是的。你记得吗?现在对话似乎已经发生了很大变化。比如,美国政府刚刚告诉 Anthropic 他们必须关闭其 AI,因为他们担心恶意行为者会用它进行网络攻击,你知道吗?政府正在觉醒,做的比我们预期的还要多。我们实际上希望这一趋势能继续下去,在真正为时已晚之前,政府内外以及更广泛的社会中会就所有这些问题进行非常严肃的对话,并试图规划一条避免我们提到的失控和权力集中风险的路线。

Yeah. You remember that? Now it seems like the conversation has changed a lot. Like now that the US government just told Anthropic they have to shut down their AI because they were worried that bad actors would use it for cyber attacks, you know? The government is waking up and doing more stuff than we expected already. And we're actually hopeful that that trend will just continue and that before it's actually too late, there will be very serious conversations happening inside the government and outside the government and in the broader society about all of these issues and trying to chart a course that avoids the loss of control and concentration of power risks that we mentioned.

按钮问题:关闭所有前沿AI训练 The button question: shut down all frontier AI training

Host

你花了将近 15 年的时间思考这些问题。如果这里有一个按钮,你按下它,你的计划 S 就会实现,它会永久关闭所有当前正在训练前沿 AI 模型的数据中心。再也不会有其他 AI 实验室研究这些问题,你会按下那个按钮吗?

You've spent what must be almost coming up to 15 years thinking about this stuff. If this here was a button and if you press that button, your plan S would occur and it would shut down every data center that is currently training a frontier AI model for good. There would never be any other AI labs working on these problems, would you press that button?

Daniel Kokotajlo

我差点就猛按下去了,直到你说了“永久”。

I was about to slam it until you said for good.

Host

哦,好吧。

Oh, okay.

Daniel Kokotajlo

我认为如果是暂时关闭,我会毫不犹豫地按下那个按钮。因为我们还没准备好做这件事,你知道吗?文明还没准备好让这些公司自我自动化,然后变得越来越聪明,最终拥有超级智能。不,有很多原因说明这非常危险。但如果它永久地排除了未来再次这样做的可能性,我至少会犹豫是否按下这个按钮。

Like I think if it was a sort of temporary shut down, I would totally slam that button. Because we are not ready to do this, you know? Civilization is not ready to have these companies automate themselves and then get smarter and smarter and then have the super intelligent. No, there's a bunch of reasons why that's really dangerous. But I would be at least hesitant to press this button if it permanently foreclosed the possibility of ever doing it again for sure.

Host

但如果你认为计划 D 是可能的,也就是我们正在进行的这场超级智能竞赛。

But if you think that plan D is probable, which is this race we're on to superintelligence.

Daniel Kokotajlo

如果我在 D 和 S 之间选择,我想我会按下去。

If I had a choice between D and S, I think I would press it.

Host

嗯,这取决于你的想法,对吧?因为如果你认为那会发生,计划 B。而唯一的替代方案。

Well, it comes down to what you think, right? Because if you think that's what's going to happen, plan B. And the only alternative.

Daniel Kokotajlo

我没说这一定会发生。

I didn't say this is what's going to happen.

Host

概率上。

Probabilistically.

Daniel Kokotajlo

对对对。我会说这是最可能的,也许这是第二可能的,也许这是第三可能的。它们都是可能的。

Yeah, yeah, yeah. Like I'd be like this is the most likely, maybe this is the second most likely, maybe this is the third most likely. They are all possible.

Host

那么,根据你目前对你认为会发生的事情的看法,你会按下按钮吗?我给你一个 S,一个确定的 S,或者你认为会发生的事情。

So with your current perspective on whatever one you think is going to happen, would you press the button? I'm giving you an S, a definite S, or whatever you think is going to happen.

Daniel Kokotajlo

这很难。

That's tough.

Host

关闭的范围是什么?所以是。

What is the scope of the shutdown? So is it.

Daniel Kokotajlo

就是没有人能再训练 AI 模型了。永远不能。

It's no one can train an AI model again. Ever again.

Host

那真的很残酷,因为就像我说的,如果我们做对了,我们可以从 AI 中获得大量好处。

That's real rough because like I said, there's loads of benefits that we could get from AI if we do it right.

Daniel Kokotajlo

我想我几乎把你放在了 Sam Altman 的位置上。

I think I've almost put you in the position of Sam Altman.

Host

是的。

Yeah.

按按钮的纠结 Torn about pressing the button

Daniel Kokotajlo

某种程度上是的。嗯,让我花点时间想想。我想我不会按下按钮,但我对此感到非常纠结。我不按按钮的原因是,我仍然抱有相当大的希望,认为我们可以得到比这好得多的东西,更像这样的东西。我认为,如果我们最终不构建强大的 AI 系统,那么文明最终可能会灭亡,你知道,比如 100 年后、200 年后之类的。比如核战争、大流行病。我不认为人类文明现在非常稳定。所以,对后代以及未来可能生活的数十亿人的潜在好处超过了当前的风险水平,但实际上……

To some degree. Yeah. Let me take a moment to think about this. I think I would not press the button, but I feel very torn about it. The reason I think I would not press the button is that I still have substantial hope that we can get something much better than this, something more like this. I think that if we don't build powerful AI systems eventually, then we're probably going to die as a civilization eventually, you know, like 100 years from now, 200 years from now, something like that. Like nuclear war, pandemic. I don't think human civilization right now is super stable. So the possible benefits for posterity and for all the billions and billions of people who could live in the future outweigh the current level of risk, but actually...

Host

我以前听过这种说法。是啊,我不知道。也许就是不行。现在的人才是我们应该优先考虑的。现在的人正处于严重危险中。他们至少在未来几十年内会没事。所以,别管后代了。优先考虑现在的人。而且现在的人肯定不想参与这个抽奖,我想说。

I've heard that narrative before. Yeah, I don't know. Like maybe it's just like nope. The people right now are the people we should prioritize. People right now are in grave danger. They're going to be fine for at least the next couple of decades. So, never mind posterity. Prioritize the people right now. And people right now definitely don't want to do this lottery, I would say.

Daniel Kokotajlo

是啊,你确实问了我一个很难的问题。那么,如果那是按钮,你会按吗?可能不会,但我会感到非常纠结。

Yeah, you've really asked me a tough question. So, would you press the button if that was the button? Probably not, but I would feel very torn.

人们能做什么 What people can do

Host

好的。所以我总是考虑观众的画像。这些人非常好奇,尤其是在 AI 这个话题上,正如我们所见,但他们想知道这对他们意味着什么。我想很多人也想知道他们能做些什么。

Okay. So, what I always think about is the personas of the audience that are watching. These are very curious people, especially on the subject of AI as we've seen, but they want to know what it means for them. I think a lot of them also want to know what they can do.

Daniel Kokotajlo

是的。嗯,人们能做什么?我认为,如果你有天赋或热情,你可以直接参与。有很多组织担心这些事情,并试图采取行动,比如政治倡导、技术研究,或者构建有用的工具,希望能帮助人们变得更好等等。但如果你不想做出重大的职业改变之类的事情,那么我会说,多关注这些问题,多和别人谈论它。做一些事情,比如给你的国会议员发邮件之类的。这不会改变太多事情,但确实有帮助。我认为,特别是对于这个问题,核心问题是人们还没有认真对待它。如果我在过去一两个小时里对你说的那些事情是每个人心中的首要问题,我们甚至不会在这里。早就会有更重要的监管措施了,你知道吗?而且不仅会有更严格的监管,还会有更好的监管,更像手术刀而不是大棒,更敏感地分辨什么是真正坏的,什么不是那么坏。政府里也会有更多专家,并给政府提供建议。所以总的来说,我认为越多的人意识到这些担忧和预测,我们就越有可能在太晚之前做好事。

Yes. Yeah, what can people do? Well, I think that if you either have talent or passion, you can get directly involved. There's lots of organizations that are worried about these things and that are trying to do something about it, like political advocacy or technical research or building useful tools that will hopefully help people be better and stuff. But if you don't want to make any major career changes or things like that, then I would say just pay more attention to these issues and talk about it more with people. Do stuff like emailing your congressman or whatever. It doesn't change things that much, but it does help. I think that especially for this particular issue, the core problem is that people aren't taking it seriously yet. If the sorts of things that I was just saying to you for the last hour or two were just top of everybody's mind, we wouldn't even be here. There would already be much more significant regulation in place, you know? And not only would there be more heavy regulation in place, but there would have been better regulation in place that's less like a cudgel and more like a scalpel, more sensitive to what's actually bad and what's not so bad. And there'd be more expert people in the government and advising the government. So just in general, the more people wake up to these concerns and to these projections, I think the more likely it is that we can do good stuff before it's too late.

Host

那他们应该如何投票呢?美国几年后就要举行选举了,但世界各地也一直在举行选举。

What about how they should vote at the polls? We've got an election coming up in the United States in a couple of years time, but there's elections happening all over the world all the time.

Daniel Kokotajlo

你应该问问你的候选人他们对所有这些 AI 事情的想法。你应该努力让他们有观点,然后投票给在这个话题上观点更好的候选人。这是我们一生中发生的最重要的事情,实际上可能是历史上最重要的事情,顺利发展非常重要。所以这是所有国家的领导人都应该思考和制定计划的事情。

You should ask your candidates what they think about all this AI stuff. You should try to get them to have opinions and then you should vote for the candidates whose opinions are better on this topic. This is the most important thing happening in our lifetimes, probably in all of history in fact, and it's very important that it go well. So this is what all the leaders of all the countries should be thinking about and making plans for.

生活在高潮前夕 Living in the run-up to the climax

Host

活在这个时刻是不是很奇怪?就像我在想所有我可能出生的时代。我想我的祖先可能也这么想,但当你说话时我在想,我觉得是你把它称为最后一场演出的时候。你用的措辞是什么?

Isn't it such a weird thing to be alive at this moment in time? Like I was thinking about all the times that I could have been born. And I guess my ancestors probably thought the same, but I was thinking as you were speaking I was like, I think it's when you referred to it as like the final show. What was the phraseology you used?

Daniel Kokotajlo

我说的是气候,那是高潮的前奏之类的。

I said the climate it was the run-up to the climax or something.

Host

是啊。我的意思是,出生在高潮的前奏中是多么疯狂的事情,你在这里描述的一切都可能在我有生之年发生,希望如此。或者也许不希望如此。活着的时代真疯狂。

Yeah. I mean what a crazy thing to be born in the run-up to the climax where everything you're describing here is within my lifetime conceivably hopefully. Or maybe not hopefully. What a crazy time to be alive.

Daniel Kokotajlo

确实。

Certainly.

AI对家庭决策的影响 Impact of AI on family decisions

Host

我注意到当我问你是否有孩子时,你的态度变化很大。就像你进入了另一种状态。显然,这一直是你思考的核心。

I noticed that when I asked you if you had kids your demeanor changed quite considerably. It's like you dropped into a different state. Obviously that's been central to the rumination that you've been experiencing.

Daniel Kokotajlo

嗯,这是一个悲伤的话题,对吧?就像我有了孩子,生孩子的很大一部分原因是为了未来,你知道吗?这不仅仅是当下和你在一起的可爱事物。而是因为你对他们的成长、他们如何做自己的事情、成为他们自己等等抱有所有这些希望和梦想。而由于 AI 正在发生的事情,我认为很多这些梦想都处于危险之中。

Well, it is a sad topic, right? Like when I had kids, the reason to have kids is in large part about the future, you know? It's not just a cuddly thing to have with you in the moment. It's because you have all these hopes and dreams about how they'll grow up and how they'll go to their own thing and be their own person and stuff. And because of what's happening with AI, I think a lot of those dreams are in jeopardy.

Host

大概你仍然会生孩子?

Presumably you still would have had kids?

Daniel Kokotajlo

我实际上偶尔会在这件事上反复。基本上,最直接的答案是我不知道。我的第一个孩子,我们是在 209 年有的她,她出生于 2019 年。是的。所以这发生在我时间线大幅缩短之前。那时我对 AI 感兴趣,我跟踪这个领域,我做预测,但我实际上没有预料到它会很快发生。然后这导致当我开始想,哦天哪,它很快就会发生。比如到 2030 年,你知道吗?这引起了一些重新考虑。所以我基本上告诉我妻子,我们不要再有孩子了。太不确定了。但这结果非常困难,尤其是对我妻子来说。就像我们已经有了一个孩子,没有兄弟姐妹。所以最终我有点让步了,想,好吧,你知道吗?我们已经有一个了。一切都会好的。也许未来会很好,即使不好,我们也都在同一条船上。

I've actually flip-flopped on this occasionally. Basically the top line answer is I'm not sure. My first child was had we had her when we were in 209 she was born in 2019. Yeah. So this is before my timeline shortened a lot. So at that point I was interested in AI, I was tracking the field, I was making forecasts, but I didn't actually expect it to happen soon. And then this caused like when I did start thinking oh my gosh, it's going to be happening like real soon. Like by 2030, you know? That caused some reconsidering. And so I basically told my wife let's not have any more kids. It's too uncertain. But that turned out to be really hard because especially for my wife. Like we already had one kid and like no siblings. So eventually I sort of gave in and was like okay, well, you know what? We already have one. It's going to be all right. Like maybe the future will be good and even if it's not, well, we're all in the same boat together.

Host

你说的很令人不寒而栗。这令人不寒而栗,因为你比我知道得多。如果你在家里对你妻子说,“听着,也许我们应该暂停生更多孩子和组建家庭,因为 AI 正在发生的事情。”

It's quite chilling what you're saying. It's chilling because you know more than me. And if you're at home saying to your wife, "Listen, maybe we should pause on having more children and building a family because of what's going on with AI."

Daniel Kokotajlo

明确地说,是的,我的意思是,是的,这非常令人担忧。我感到寒意。这很糟糕。

To be clear, is it Yes, I mean yes, it's very concerning. I am chilled. This is bad.

结束语与行动号召 Closing thoughts and call to action

Host

这正是我一直说的。我希望事情能顺利。我认为事情可能会顺利。我觉得我们有很多可以做的来引导事情朝更好的方向发展。

This is what I've been saying. I hope things go well. I think things might go well. I think that there's a lot we can do to steer things in a better direction.

Daniel Kokotajlo

我得说,其中一件事就是公开谈论它。我认为我们看到政府觉醒的很多进展,以及人们在某些活动中对某些人喝倒彩,这些都源于像你这样的人来上这样的节目和其他播客,告诉我们发生了什么。因为否则我们会被那些拥有最大公关机器的人误导。

I mean one of those things as well I have to say is just speaking about it. I think a lot of the progress we've seen with governments waking up and we've seen certain things with people booing certain people at certain events. It is downstream from people like yourself actually coming on shows like this and all the other podcasts and telling us what's going on. Because else we're going to be gaslighted by the people that have the biggest PR machines.

Host

是的。

Yeah.

Daniel Kokotajlo

所以我经常发现自己处于两种心态,因为我是企业家和投资者。我现在投资了大概超过 100 家公司,其中很多都在使用 AI。我投资了 Grok,那家推理芯片公司。还投资了 SpaceX,它现在拥有另一个 Grok,并且他们在做 AI。我每天都在生活中使用 AI。在这次对话中,我一直在用它来理解你说的不同事情。所以这是我的一面,作为商业建设者、企业家,我看到了 AI 在我生活中的好处。然后还有我的另一面。有趣的是,我认为有时人们觉得你必须选边站。但在我的一生中,即使当我还是社交媒体 CEO 时,我说过,‘顺便说一句,我在建立社交媒体业务,但我认为社交媒体有一些缺点。’我发现自己处于同样的时刻,我既说‘我用 AI 构建。我有 AI 投资。’同时作为普通人,我又……

So, I often find myself in two minds because I'm an entrepreneur and an investor. I'm an investor in probably more than 100 companies now, and many of those companies are using AI. I invested in Grok, the inference chip company. Invested in SpaceX which now owns another Grok and they're doing AI. I use AI every day in my life. I've been using it through this conversation to understand different things that you've said. So that's one side of me, the business builder, entrepreneur who has seen the benefits of AI in my own life. And then there's the other side of me. It's funny because I think sometimes people think you have to pick a camp. But through all of my life, even when I was a social media CEO and I was saying, 'By the way, I'm building a social media business but I think there are some downsides to social media.' I find myself at the same moment where I'm like, 'I build with AI. I have AI investments.' And at the same time as a civilian I'm like...

Daniel Kokotajlo

是的。我认为这是一种紧张关系。我认为有不同的方式可以划清界限。我认识很多人以各种不同的方式划清界限。有些人会说,‘我不会使用 AI。我认为这东西不好,而且处于糟糕的轨迹上,所以我要抵制 AI。’我不是那种人。我大量使用 AI。我们在 AI Futures Project 都这样做。这对我们很多工作都有帮助。光谱的另一端是那些说,‘嗯,它似乎正在发生,所以让它顺利发展的方法是参与进来,积累权力,并试图从内部引导它。’所以我要去 OpenAI 或 Anthropic 工作,努力晋升,然后在做出重要决策时成为重要人物。我认识很多这样的人。那就像我过去在做的事情……不完全是那样,但类似。

Yeah. I mean I think that is a tension. I think there are different ways you can draw the line. I know lots of people who draw the line in lots of different ways. There are some people who just say, 'I'm not going to use AI. I think this stuff is bad and on a bad trajectory so I'm going to boycott AI.' I'm not one of those people. I use AI a lot. We all do at AI Futures Project. It's helpful for a lot of our work. The opposite end of the spectrum is people being like, 'Well, it seems like it's on a trajectory to happen so the thing to do to make it go well is to get involved and accumulate power and try to steer it from the inside.' And so I'm going to go work at OpenAI or Anthropic and try to climb the ranks and then be someone who matters when the important decisions are being made. I know loads of people like that. That was like what I was doing when I was... That wasn't what I was doing exactly but like...

Host

那就是那条路。

That was the path.

Daniel Kokotajlo

那就像……我的意思是,在某种意义上,这就是这些公司的整个叙事,对吧?这就是为什么他们告诉自己他们所做的是可以的,因为他们担心其他人。所以所有这些人都决定,‘我们要全力以赴。我们要在决策时在场。’所以有一个完整的光谱,我处于中间的某个位置。我不在公司里,我不帮助他们加速。相反,我在与广大公众交谈,并试图倡导我认为是目前最好的出路、前进方向。但我并没有抵制所有 AI。我没有拒绝以那种方式与之互动。

That was like... I mean in some sense this is what the whole narrative of the companies is, right? This is why they tell themselves it's okay to do what they're doing, because they're worried about the other guys. And so all these people are deciding, 'We're going to lean really hard into it. We're going to be there in the room when decisions are being made.' So there's a whole spectrum and I'm sort of somewhere in the middle. I'm not at companies, I'm not helping them go faster. Instead, I'm talking to the broad public and trying to advocate for what I think is my current best guess as to the way out, the way forward. But I'm not boycotting all the AIs. I'm not refusing to engage with it in that way.

Host

你觉得太晚了吗?

Do you think it's too late?

Daniel Kokotajlo

不。我不认为太晚了。如果我认为太晚了,我就不会在这里了。

No. I don't think it's too late. If I thought it was too late, I wouldn't be here.

Host

那你会在哪里?

Where would you be?

Daniel Kokotajlo

和我的家人在一起。

With my family.

Host

如果你必须给公众一个结束语,你会说什么?

What's your closing message to the general public if you had to have a closing statement to them?

Daniel Kokotajlo

也许我会说,你会听到很多关于 AI 的事情,而且你已经听到了很多,它们听起来像科幻小说,但有时听起来像科幻小说的事情在现实中会发生。事实上,历史上很多次,曾经是科幻小说的事情后来变成了现实。人们需要停止思考什么听起来像或不像科幻小说,而开始思考趋势,这项技术所处的实际趋势,阅读并预测它将如何发展,然后认真对待它可能这样发展的可能性,然后思考应该对此做些什么。

Maybe I would say that you're going to hear a lot of things and you already have been hearing a lot of things about AI and it's going to sound like science fiction, but sometimes things which sound like science fiction happen in reality. In fact, many times historically things which used to be science fiction have then become reality. And people need to stop thinking about what does or doesn't sound like science fiction and just start thinking about the trends, the actual trends that this technology is on, and reading and forecasting how it's going to go, and then taking seriously the possibility that it could go something like this, and then thinking about what should be done about that.

Host

你会建议他们去哪里获取更多信息?

And where would you direct them to get more information?

Daniel Kokotajlo

你可以访问 ai2047.com 阅读我们之前的场景。你可以访问 ai2040.com plan A 阅读我们关于应该做什么的新提案。这些东西不仅仅是科幻故事。它们还有很多解释和链接到其他内容。所以它们是一个很好的起点来了解所有这些。如果你愿意,我可以在结束后提供一个阅读清单,包括其他论文、文章和博客等。

You can go to ai2047.com to read our previous scenario. You can go to ai2040.com plan A to read our new proposal for what is to be done. These things are not just a sci-fi story. They also have lots of explainers and links to other things. So they're kind of like a nice jumping off point to learn about all of this stuff. If you want, I could after this is over give a reading list of other papers and articles and blogs to follow and so forth.

Host

请务必。我会把它们都链接在下面的评论区。所以如果你现在在听,去看看这集的描述,你会看到一堆链接,那是 Daniel 推荐你阅读的内容。我认为现在是了解这些东西的绝佳时机。人类有一种倾向,因为认知失调,当我们对某事感到不舒服时,会把头埋进沙子里逃避。但实际上,我认为这正是应该做相反事情的时候。出于很多原因,要让自己了解情况,以便知道采取什么行动,同时也因为 AI 不可避免地将成为我们所有生活和职业的很大一部分。

Please do. And I'll link them all below in the comment section. So if you're listening now, go ahead and take a look at the description of this episode and you'll see a bunch of links which is Daniel's recommendations of what you should read. I think it's just a really great moment in time to get educated on this stuff. Humans have an inclination because of cognitive dissonance where we feel uncomfortable about something to bury our heads in the sand and avoid it. But actually, I think this is one such time to do the very opposite. For many reasons, to inform yourself so you know what actions to take, but also because AI unavoidably is going to be a huge part of all of our lives and careers.

Daniel Kokotajlo

是的。谢谢。说得好。它会非常重要。它很快就会无处不在,我们需要在太晚之前做点什么。

Yeah. Thank you. And that's a good way to say it. It's going to matter a lot. It's going to be everywhere soon and we need to do something about it before it's too late.

Host

那 AI Future Project 呢?

What about AI Future Project?

Daniel Kokotajlo

那是我们的组织。在我离开 OpenAI 后,我们花了一年时间写了 2047,然后又花了一年时间写了 2040 Plan A。

That's our organization. We spent a year writing a 2047 after I left OpenAI and then we spent another year writing a 2040 Plan A.

Host

Daniel,谢谢你。

Daniel, thank you.

Daniel Kokotajlo

谢谢。

Thank you.

Host

感谢你所做的一切工作。我能看出你有多关心这些东西,正是你的关心——有趣的是,关心本身就让别人感受到关心。

Thank you for all the work that you do. I can see how much you care about this stuff and it's your care, it's funny care itself makes others feel care.

闭幕词 Closing remarks

Host

看到这件事对你来说如此个人化,看到你为此奉献了如此多,但同时也听说你为了能公开谈论这些信息而放弃了 200 万美元,这令人无比敬佩。我认为在这个话题上,像你这样的声音比以往任何时候都更重要。所以,请继续你正在进行的战斗。这是一场关于信息的战斗,关于诚实的战斗,也是关于把通常不该说出口的话说出来的战斗。

And seeing how personal this is for you and seeing how much you've dedicated your life to this, but also hearing that you basically walked away from $2 million to be able to speak to the public about this information is incredibly admirable. And I think voices like yours are more important now than they've ever been on this subject. So, please do keep fighting the fight that you're fighting. And that's one of information, it is of honesty, and it is of saying what is often the quiet part out loud.

Daniel Kokotajlo

谢谢。

Thank you.

Host

你做的研究非常非常聪明。我会在下方链接我们今天讨论的所有内容,希望我们很快能再聊。

You're doing really, really smart research. I'll link everything we've discussed today below and I hope we can chat again sometime soon.

Daniel Kokotajlo

谢谢。

Thank you.

Host

YouTube 有这个疯狂的新算法,它们基于 AI 和你所有的观看行为,确切知道你想看的下一个视频是什么。算法说这个视频对你来说是完美的。对于现在正在看的每个人来说,它都不一样。看看这个视频吧,我打赌你可能会喜欢。

YouTube has this new crazy algorithm where they know exactly what video you would like to watch next based on AI and all of your viewing behavior. And the algorithm says that this video is the perfect video for you. It's different for everybody looking right now. Check this video out. I bet you might love it.

互动版:逐字朗读 + 针对本期提问 →