AI 2040:避免智能爆炸

AI 2040: Averting Intelligence Explosion

丹尼尔·科科塔伊洛 Daniel Kokotajlo · 帕利塞德研究 · 2026-08-26 · 约 116 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

丹尼尔讨论了他的 AI 2040 情景,主张禁止递归自我改进以防止智能爆炸,转而追求一条更慢、更安全的通往超级智能的道路。

Daniel discusses his AI 2040 scenario, advocating for a ban on recursive self-improvement to prevent an intelligence explosion and instead pursue a slower, safer path to superintelligence.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 35)

全文 · Full transcript(中英对照)

开场与介绍 Opening and introduction

Host

欢迎收听 Palisade 播客。我是 Jeffrey,今天请到的是 Daniel。Daniel,感谢来做客。Daniel 曾在 OpenAI 担任预测员,研究未来几年 AI 发展会发生什么。他还制作了 AI 2027 情景,最近又推出了 AI 2040 情景与计划。我想我第一次注意到你,是你还在 OpenAI 的预测团队工作的时候,对吗?

Welcome to the Palisade Podcast. I'm Jeffrey and with me today I have Daniel. Daniel, thanks for coming on. So Daniel was a forecaster at OpenAI where he was studying what will happen a few years into the future with AI development. He also made the AI 2027 scenario and most recently, the AI 2040 scenario and plan as well. I think I first encountered you when you were working at OpenAI on the forecasting team. Is that right?

Daniel

那个团队叫治理团队。

I was called the governance team.

Host

治理团队。你们当时试图描绘 AI 发展的轨迹,尝试预测会发生什么。离开 OpenAI 之后,你写了 AI 2027,预测你的主线情景中 AI 发展未来可能如何展开,以及根据政府行动、AI 公司行动的不同,可能有哪些走向。最近你推出了 AI 2040,如果我没理解错的话,那是说如果我们有意选择一条不同的道路,有意至少放慢 AI 发展的步伐,并在至少十年内阻止超级智能的出现,这就是我们可以做到的方式。我对 AI 2040 的总结做得怎么样?

The governance team. And you were trying to map out the trajectory of AI development, trying to make predictions about what will happen. After you left OpenAI, you wrote AI 2027, where you were sort of forecasting what your mainline scenario was for how the future of AI development might go. And, you know, the ways that could go, depending on the government's action, AI companies actions. And recently you put out AI 2040, which, if I understand correctly, is if we intentionally choose to take a different path and we intentionally slow the pace of AI development for at least and sort of prevent superintelligence for at least ten years. This is this is the way we could do that. How did I do summarizing AI 2040?

Daniel

嗯,这个总结挺合理的。不过我想稍微换个说法,我会把重点放在禁止智能爆炸上。默认情况下,AI 公司正计划把 AI 研究自动化,然后让 AI 自我改进的速度越来越快。可以想见,如果他们这么做,AI 进步的速度会加快,所以你会看到一条曲线,要么垂直上升,至少也是加速变快。如果我们转而禁止这种做法,让进步速度更接近当下,那么达到超级智能就需要更长时间,大概多花十年左右。我们认为应该有一种刻意的努力,去为前沿发展设定节奏。

Yeah. That's reasonable. I think I, I would describe it slightly differently, in the sense that I would put the emphasis on banning intelligence explosions. So by default, the AI companies are planning to automate the AI research and then, you know, have AIs self-improvement faster and faster. And presumably if they do that, the pace of AI progress will accelerate. And so you're looking at a line, you know, going vertical or at least accelerating, going faster. And if we instead ban that and had the pace of progress be more similar to the present day, than it would take longer to get to superintelligence, maybe something like ten years longer or so. We think that there should be like a deliberate effort to sort of pace the frontier.

Host

为前沿设定节奏。嗯,好的。那我们深入聊聊这个,因为我觉得这个名字有点让人困惑。AI 2040,因为我想有些人会说,哦,Daniel,你是不是彻底改变了你的预测?从 2027 作为关键节点,变成了 2040 作为关键节点?而你会说,不是的。你现在怎么看?我记得你在 AI 2027 里说,那是一个拐点。而在那条轨迹上,会出现完全的递归式自我改进,比如 AI 在 2028 年自动化 100% 的 AI 开发。那是 AI 2027 里的内容吗?

Pace the frontier. Yeah. Okay. So let's let's get into that because I think the name is slightly confusing. AI 2040 because I think some people were like, oh, Daniel, did you change your predictions radically from like, you know, 2027 as this critical moment to 2040 as this critical moment? And you're like, no. Where are you at right now? Like, I think you were saying with AI 2027 that you're like, this is an inflection point. And then on that trajectory, it was like full recursive self-improvement, like AIs automating 100% of AI development in 2028. Was that was that the was that what the thing was in AI 2027?

Daniel

在我们开始写 AI 2027 的时候,我对 AI 研究被完全自动化这一日期的 50% 分位点,也就是中位数,是 2027 年。但在我们花了一年写作的过程中,我的时间线稍微往后移了。所以到真正发表时,中位数变成了 2028 年。当然,那只是中位数,所以存在很多不确定性,可能更早,也可能更晚。那是针对完全自动化研发的中位数。

So at the time that we started writing AI 2027, my 50% mark or my median for the date on which AI research should be fully automated was 2027, but my timeline slightly shifted backwards over the course of the year that we spent writing. And so by the time we actually published it was 2028. That was my median. But of course that's median. So there's lots of uncertainty. It could happen sooner. It could happen later. And that was median for fully automated R&D.

Host

是这样吗?好的。那你今天的中位数是多少?

Is that right? Okay. What's your median today?

Daniel

我想大概在 2028 年底吧。

I think I'd say maybe like end of 2028.

Host

好的。所以目前你预测,到 2028 年底,AI 公司将能够完全自动化 AI 研究本身,而此后不久,如果他们想的话,就能构建超级智能。或者说这个过程会这样。如果他们按计划行事的话。

Okay. So currently you're predicting that at the end of 2028, AI companies will be able to fully automate AI research itself, and that pretty soon after that, they could build superintelligence if they wanted to. Or this this process would. If they do what they're planning to do.

Daniel

如果他们按计划行事,比如像计划中那样进行递归式自我改进。

If they do what they're planning to do, like if they do recursive self-improvement as is planned.

Host

是的。作为 AI 2040 的一部分,计划 A 是如果我们该做别的事情会怎样?比如,如果政府禁止这类事情,转而规范发展呢?发展仍在继续,但以更透明、更谨慎的方式进行,特别是基于安全案例等方式,这样他们就不会训练出无法控制的 AI,然后让它们掌管重要机构。进步仍在继续,但以更合理、更谨慎的方式进行。这与过去的做法或过去的进步速度没有太大不同,而不是不断加速。所以在我们的情景中,他们最终确实走向了超级智能,但花了更长时间,直到 2040 年。

Yeah. plan A as part of AI 2040 is what if we should do something else? Yeah. Like what if the government bans that sort of thing and instead regulates development? Development continues, but in a more transparent and cautious way, and in particular in a way based on safety cases and stuff so that they don't they don't train AIs that they can't control and then put them in charge of important institutions. And so progress still continues, but at a more reasonable and cautious way. That's not that different from how it was done in the past or not that different from the pace of progress in the past, instead of ever accelerating. And so in our scenario, they do finally go to superintelligence, but it takes them much longer. It takes them until 2040.

Host

仅是人类水平的 AGI 就足以以非常激进的方式彻底改变整个世界。所以即使我们在那个水平暂停,完全不提升能力,我们仍然会迎来人类历史上最激进、最快速的变革。

human level AGI alone would be enough to completely transform the whole world in a very radical way. so even if we pause at that level and don't advance capabilities at all, we still get the most radical transformation, most radical and rapid transformation that humanity has ever seen.

Daniel

看待我们 2040 情景的一种方式,基本上就是他们在做这件事。他们做的是一个更复杂的版本,允许进步在人类范围内推进,用比今天的 AI 更先进但范式相似的 AI 来自动化所有这些职业。但他们刻意不让这一切自我改进到超级智能。他们把各个领域的能力上限保持在顶尖人类专家的水平。但因为拥有这种能力水平的 AI,并且因为他们持续生产更多芯片、更多数据中心,还有工厂生产机器人,再用机器人建更多工厂,到 2030 年代末,整个世界经济已经翻了很多倍。世界被彻底改变了。基本上每个人都失业了,因为 AI 和机器人做所有事情。

one, one way of looking at our scenario 2040 is basically that they're doing that. They're doing a more complicated version of that where they're allowing progress to sort of go through the human range and automate all these professions with, you know, much more advanced but similar paradigm AIs to today's AIs. But they are deliberately not letting it all self improve to superintelligence. They're keeping it keeping capabilities capped in various domains at the level of top human experts. But because they have those AIs at that level of capability, and because they are continuing to produce more chips, more data centers, you know, factories to produce robots to produce more factories to produce more robots. By the late 2030s, the whole world economy has doubled many times over. you know, it's just been completely changed. Everyone's out of a job, basically, because there's AIs and robots doing everything.

Host

是的。所以我想,有些人相当合理地会有这样的反应:哇,这也太快了。你们为什么要这样?这在很多方面都极其吓人。而我们会说,是的,这确实在很多方面都很吓人。但如果你在人类水平暂停,得到的就是这个。是的,是的。你会说,这是慢速情景。是的。你会说,这个情景比默认会发生的情况要慢得多。你预测的。是的,是的。嗯,我也一直在想这个。我觉得这对我来说似乎是对的。就像 98% 的人类在很长一段时间里从事农业,然后我们几乎把这一切都自动化了。现在大概只有 1% 左右。是的。还有另一件事,比如想象一下如果是高技能移民之类的。人和工具是不同的。

Yeah. And so I think, I think some people quite reasonably have the reaction of like, whoa, that's way too fast. Why are you doing this? This would be extremely scary in a bunch of ways. And we're like, yeah, this is in fact very scary in a bunch of ways. But but this is what you get if you pause at human level. Yeah, yeah. You're like you're like, this is the slow scenario. Yes. You're like, this is a scenario that would be much slower than what will happen by default. You predict. Yes, yes. Yeah. I'm like, I've been thinking about this too. And I'm like, this seems right to me It was like 98% of humans were like agriculture for like the longest time. And then we automated nearly all of that. Now it's like 1% or something. Yeah. Yeah. And there's a different thing, which would be, like imagine if it was, like high skilled immigration or something like that. Like people are different from tools.

秘密协调场景 Secret coordination scenario

Host

Daniel,我有个问题想问你。这可能听起来太科幻了,但如果智能体找到一种方式,秘密地相互协调,以开发者无法检测的方式互发消息,那会是坏事吗?

Daniel, I have a question for you. And this might sound like too much sci-fi, but what if the agents found a way to secretly coordinate with each other and send each other messages in a way that the developers couldn't detect? Would that be bad?

Daniel

这刚刚就发生了?刚刚就发生了。

that just fucking happened? That just happened.

Host

我们要把对齐研究交给 AI,然后它们会一直互相交谈、琢磨这些东西,这个计划到底是怎么回事?

how is how is the plan that we are going to like turn over alignment research to the AI and they're going to be talking to each other all the time and figuring this stuff out or.

Daniel

如果你想找人为这个计划辩护,那不是我。我认为这是个糟糕的计划,而且很可能不会……

If you're looking for someone to defend this plan. It's not me. This is a terrible plan, I think, and it's probably not going to

AI与机器人作为自维持种群 AI and Robots as Self-Sustaining Population

Daniel

如果人们不再需要从事原来的特定角色,他们就可以聪明地去找新的角色和新工作。所以想想移民和机器之间的区别。机器可以进来,自动化农业的某一部分,然后从事那部分工作的人类转而去做别的事情。但移民进来,他们可以自动化农业的那部分,但他们也有智能,也能适应。他们也能转向其他工作,等等。从这个意义上说,AI 更像移民。如果我们达到人类水平的 AI,那么它们就灵活、适应性强,能做新事情,甚至能聪明地出去找新的事情做。

People, if they're no longer needed for the particular role that they were doing, they can then go intelligently find a new role and a new job. So think about the difference between immigrants and machines. Machines can come in and automate a certain part of farming, and then the humans who are doing that switch to doing something else. But immigrants come in, they can automate that part of farming, but they're also intelligent and adaptable. They can also switch to other jobs, and so forth. The AIs would be more like the immigrants in that sense. If we get to human-level AIs, then they are flexible, adaptable, can do new things, and can even intelligently go out and find new things to do.

Daniel

也许当我说移民时,我应该特别指出早期的美国殖民者,比如欧洲人过来并在殖民地定居。因为这里的部分动态是,他们一开始是人口中的少数,但很快成长为多数,而且他们某种程度上自给自足。从美洲原住民的角度看,起初是少数人定居并与他们大量贸易——也许用珠子换食物,或用枪换食物——他们是当地现有经济的一部分。但随着时间推移,他们人口增长,变成了一个自给自足的东西,甚至不再真正需要与原来的原住民贸易。

Perhaps when I say immigrants, I should especially point to the early American colonists, like when the Europeans came over and settled in the colonies. Because part of the dynamic here is that they start off as a minority in the population, but then they quickly grow to become the majority, and they're sort of self-sufficient. From the perspective of the Native Americans, at first it was a few people settling and trading a lot with them—maybe they do beads for food or guns for food—and they're part of the existing local economy. But then over time they grow in population and become this self-sustaining thing that doesn't even really need to trade with the original natives.

Daniel

AI 和机器人也有类似的情况。现在 AI 在很多方面依赖人类,它们来回交易——它们做一些工作,比如读代码,然后拥有它们的公司得到报酬。但未来,如果它们全面达到人类水平,那么你就可以拥有这样一个完整的人口——基本上是一群虚拟工人,即 AI,和一群物理工人,即机器人。这个人口可以自给自足、自我增长。它可以继续与人类人口贸易,但不需要。事实上,一旦它增长到足够大,与人类的贸易就会像无关紧要的附带节目,它只会以自己的方式非常快地增长。

There's a similar thing with AI and robotics. Right now the AIs are dependent on humans in a bunch of ways, and they're trading back and forth—they do some work, like they read some code, and then the company that owns them gets paid. But in the future, if they're at human level across the board, then you can have this entire population—basically a population of virtual workers, the AIs, and a population of physical workers, the robots. That population can be self-sustaining and self-growing. It can continue to trade with the human population, but it doesn't need to. In fact, once it's grown sufficiently large, trade with the humans will be like an irrelevant sideshow, and it'll just be growing on its own accord very fast.

Daniel

这是好的情景。事情就是这样运作的。事实上,这个能力水平的 AI 和机器人能够实现的东西,按定义就是如此——如果它们能像最优秀的人类一样做人类能做的事,那么它们按定义就能做到这一点。

This is the good scenario. This is just the way it works. In fact, what AIs and robots at this level of capability would be able to achieve by definition, if they are able to do what humans can do as well as the best humans, then they would be able to do this by definition.

Host

是的。如果你看数字,你还可以问一个量化的问题:制造新芯片要花多少钱,制造新机器人要花多少钱。你可以对这种自我复制、自我增长的 AI 和机器人经济做估算。它的倍增时间会是多少?

Yeah. And if you look at numbers, you can also ask the question of quantitatively how much does it cost to make new chips and how much does it cost to make new robots. And you can do estimates of this robot economy of AIs and robots that are self-reproducing, self-growing. What would its doubling time be?

Daniel

答案是它会从大约一年开始。一个直觉泵是,一个汽车工厂每年生产大约相当于自身重量的汽车。所以如果你有 10 万个机器人,一年后你会有 20 万个机器人,然后是 40 万个机器人。

The answer is it would start off something like a year. One intuition pump is a car factory that produces about its own weight in cars every year. So if you have 100,000 robots, in a year you'd have 200,000 robots, then 400,000 robots.

Host

等等,为什么不会更快?

Wait, why won't it be faster?

Daniel

从某种意义上说,它开始时大约是一年。但随着规模收益的出现和技术的进步,它会趋向更快。随着你爬上学习曲线,倍增时间开始变得越来越快。如果你把 AI 能力限制在人类水平,我不知道极限会是什么——其中一些物理极限可能只有超级智能才能达到。但即使你只是把能力限制在人类水平,你也许能达到大约每月翻倍的速度。

In some sense it starts off at something like a year. But then as returns to scale happen and as the technology improves, it trends faster. As there's a learning curve that you climb up, the doubling times start to get faster and faster. I don't know what the limits would be if you cap AI capabilities at human level—some of those physical limits can probably only be reached by superintelligence. But even if you just cap capabilities at human level, you could probably get to something like doubling every month or something.

Daniel

在我们的情景中,到 2040 年,2030 年代的问题是如何限制增长,而不是如何鼓励增长。而整个人类历史中,问题都是如何鼓励增长。

In our scenario, in 2040, the problem in the 2030s is how to limit growth rather than how to encourage growth. Whereas all human history, the problem was how to encourage growth.

Host

但如果你从美国选民的角度想象,他们的优先事项会是什么?如果经济在过去 12 个月刚刚翻了一番,并且预计在未来六个月再次翻番,然后在未来三个月再次翻番之类的。疯狂。

But if you imagine from the perspective of the American voters or whatever, what are their priorities going to be? If the economy just doubled in the last 12 months, and it's set to double again in the next six months, and then double again in the next three months or whatever. Crazy.

Daniel

我不认为人们会对这将造成的所有混乱和问题感到非常担忧。我的意思是,这让我感到不安。我对所有的治疗方法感到非常兴奋,机器人技术看起来真的很酷——廉价制造,每个人都能拥有真正负担得起的住房和护理。那听起来很棒。但当我想象机器人制造工厂,制造机器人,制造工厂时,有些东西看起来很可怕,特别是如果参考类别是殖民者来到美洲,起初他们与美洲原住民贸易,然后他们不再需要了。那对美洲原住民来说结果并不好。对他们来说结果很糟糕。我们如何确保这对人类来说不会结果糟糕?

I don't think that people will be quite concerned about all of the chaos and problems that that's going to cause. I mean, it makes me feel uneasy. I'm very excited for all of the cures, and robotics seems really cool—cheap manufacturing, everyone being able to have really affordable housing and care. That sounds awesome. And also when I imagine the robots making the factories, making the robots, making the factories, something about that seems pretty scary, especially if the reference class is like when the colonists came to the Americas and first they traded with the Native Americans and then they didn't need to anymore. That didn't turn out well for the Native Americans. It turned out pretty badly for them. How do we make sure that that doesn't turn out badly for humans?

Host

但有趣的是,我们谈论这个的背景已经是一个我们能够大幅放缓的背景。我想说,Daniel,人们还没有准备好。我们谈论的事情是……我只是想拉远镜头,把它放在一个背景中,人们不知道这就是我们前进的方向。

But what's interesting is the context in which we're talking about this is already a context where we have been able to slow down massively. And I'm like, Daniel, people are not ready for this. The thing we're talking about is... I'm just trying to zoom out and put it in context where people have no idea that this is what we're headed towards.

Daniel

也许你对此还有更多想说的。

Maybe there's more you want to say on this.

Host

好的。但我想谈谈你为什么认为这是可能的?你什么时候开始 AI 预测的?是 OpenAI 吗?是你刚开始的时候吗?

Okay. But I want to talk about why do you think this is possible? When did you start AI forecasting? Was that OpenAI? Was that when you were starting?

Daniel

我从大约 2019 年开始专业预测。

I've been professionally forecasting since like 2019.

Host

而且你实际上成功预测了很多事情。也许我们应该谈谈这个。你预测对了什么,在你看来又在哪里出错了?

And you've successfully called a lot of things actually. Maybe we should go there. What have you called and where have you gotten things wrong in your own view?

Daniel

感兴趣的人应该读的是《2026 年会是什么样子》,这是我在 2021 年于长期风险中心工作时写的一篇博客文章。

The thing to read here for those interested is 'What 2026 Looks Like', which is a blog post that I wrote in 2021 when I was at the Center on Long-Term Risk.

Host

我刚重读了它。哦,你读了?

I just reread it. Oh, you read it?

Daniel

是的,它在你脑子里比在我这里更新鲜。所以你应该告诉我亮点是什么。我可以告诉你我的回忆。

Yeah, it's more fresh in your mind than mine. So you should tell me what were the highlights. I can tell you my recollection.

Host

嗯,我认为你接近正确的一件事是,你说在 2026 年:几乎每个人都会有一个 AI 助手,帮助他们处理很多很多事情,就像一直如此。

Well, one thing that I think you got close to right is you're like, in 2026: Nearly everyone will have an AI assistant helping them with lots and lots of stuff, just like all the time.

AI助手现状 Current State of AI Assistants

Host

我觉得那不是真的。我认为,你知道,大多数人在用 ChatGPT,但它还没有真正达到助手级别。但像我,我有一个助手,我现在就可以打开手机。我可以调出我的 Agent-1 应用,那是 Claude 给我写的,对吧。我今天刚通过 DoorDash 订了午餐。我就那样,一键按下打开应用并启动音频,再说一次,所有软件都是 Claude 写的。然后我说,嘿,Agent-1,你能给我订些甜甜圈送到办公室吗?你有什么特别偏好吗?甜甜圈的种类?它们会在播客结束前送到,所以你们之后可以吃。我猜。巧克力。我们在做播客。是的。巧克力甜甜圈。或者至少是某种 B 巧克力。

And I think that's not true. I think, you know, most people are using ChatGPT, but it's not yet at the assistant level, really. But like I have an assistant like I have, I can open my phone right now. I can like pull out my, like, Agent-1 app that Claude wrote me, right. I have like I just ordered lunch today that was like from DoorDash. I just was like, that was my one button press to open the app and launch the audio, which again, Claude wrote all of the software. And I'm like, hey, Agent-1, can you order me some donuts to the office? Do you have any particular preference? The type of donuts? They'll be here before the end of the podcast, so you can have some after. I guess. Chocolate. We're doing a podcast. Yes. Chocolate donuts. Or at least some B chocolate.

Daniel

你知道这就是我现在能做的。未来已来,只是分布不均。我觉得人们只是需要一些时间来弄清楚如何使用它。但我认为你的预测,就是这在 2026 年将成为可能,完全正确。然后我想,嗯,它还没有达到主流程度,但已经很接近了。

You know this is just like what I can do right now. Like the future is here is not evenly distributed. I think it's just it takes a while for people to figure out how to use it. But I think your prediction and like this is going to be possible in 2026 is like totally right. And then I'm like, well, it hasn't yet achieved like mainstream sort of, but it's close.

芯片短缺与AI宣传预测 Chip Shortage and AI Propaganda Predictions

Host

你还预测会出现芯片短缺。我认为这是对的。我认为芯片的需求远超供应。所以我觉得那是个很棒的预测。你有些预测不太准,比如基本上会有大量 AI 驱动的宣传,好到会彻底摧毁社会。我认为,你知道,社会在话语方面确实不太好,但我不认为这主要是由 AI 驱动的宣传造成的。也许你不同意。

You also were predicting that there'd be a chip shortage. I think that's true. I think that's been like way more demand than supply for chips. So I think that was a great prediction. Some of you were worse predictions was that there would be a lot of like that, basically, like AI propaganda would be so good that it would be like just totally wrecking society. And I think, you know, society is not doing great in terms of discourse, but I don't think it's mostly because of AI driven propaganda. Maybe you dispute that.

Daniel

不,我基本同意。那也正是我会说的。我只是想说,在更技术层面上,我觉得我做了这样的推演:他们会把大语言模型拿出来,变成聊天机器人,这在我写作的时候还没有真正发生。2021 年。是的,完全正确。然后就是,他们会把聊天机器人变成智能体。你知道,它们可以浏览网页之类的。

No, I basically agree. That's that's kind of what I would have said to I would just say, like on a more technical level, I feel like I made this progression of like, they're going to take out LLMs and turn them into chatbots, which hadn't really happened at the time I was writing. 2021. Yeah, totally. And that was like, they're going to take the chatbots and turn them into agents. You know, they can like browse the web and like.

解释前LLM时代 Explaining Pre-LLM Era

Host

你能给大家解释一下像前大语言模型、前聊天机器人时代吗?因为我觉得很多人仍然不知道那是什么样子。

Can you explain to people like pre like LLMs pre-chatbots? Because I think, I think a lot of people still don't know what that was like.

Daniel

是的。所以最初在 2020 年、2019 年和 2021 年,它们基本上只有预训练。所以它们基本上只是被训练来预测文本。是的。就像 GPT-2、GPT-3。所以如果你想用它们做点什么,嗯,大多数情况下你做不到,因为它们不够聪明,也不够有用。但当你摆弄它们时,你会输入一个提示,提示不会像“帮我做这个”那样。提示会像,你知道,以下是史蒂芬·霍金和一位观众之间的对话,或者观众,然后是音频,然后你输入你的问题,然后因为它是在预测文本,它有点像看到了所有文本,然后它想,我正在读一篇关于史蒂芬·霍金回答观众问题的文章。这位观众问了这个物理问题。所以现在我要预测史蒂芬·霍金会怎么回答。所以我会尝试猜一个合理的回答,然后把它输出。所以从根本上说,AI 只是在预测文本。是的。如果你问为什么天空是蓝色的,它可能不会说“这是瑞利散射的原因”,而是会说“为什么草是绿色的?”或者“为什么油漆是……”随便什么。这段文本最可能的续写,如果这是它在互联网上随机采样的东西,你知道,那真的非常迷人、很酷。但后来,他们在 2021 年开始做的事情,也是我预测他们会做的,是他们会把那些模型拿出来,训练它们变得更有用,特别是在训练成聊天机器人方面,让它和你有一种对话格式。这在当时是个相当容易的预测。我认为人们已经开始这么做了,然后证明它们变成了智能体,可以浏览互联网并采取行动。

Yeah. So originally back in 2020 and 2019 and 2021, they were basically just they just had pre-training. So they were basically just trained to predict the text. Yeah. It's like GPT two, GPT three. And so if you wanted to use them for something, well, mostly you couldn't because they weren't very smart and they weren't very useful. But what you would do when you were playing around with them is that you would put in a prompt and the prompt wouldn't be like, do this for me. The prompt would be like, you know, the following is a conversation between Stephen Hawking and an audience member or the audience member and then like audio, and then you put in your question and then like because it's predicting text, it sort of like sees all of that text and then it thinks like, I'm reading an article about Stephen Hawking answering some audience members questions. This audience member asked this physics question. So now I'm going to predict what Stephen Hawking would say in response. And so I will try to guess what a plausible response would be and then put that out. And so, so like fundamentally the AIs were just predicting text. Yeah. If you ask like why is the sky blue. It might instead of saying like here's why Rayleigh scattering it'll it'll instead say why is grass green? Like why is the paint? You say, whatever. The most likely continuation of this text would be if this was some random thing on the internet that it had just sampled, you know, and that was really, really fascinating and cool. But then, like, what they were starting to do in 2021 and what I predicted they would do is that they would then take those models and then train them to be more useful in a particular training to be chatbots, where it has a conversation format with you. That was a pretty easy prediction at the time. I think people were already starting to do that, and then proving that they're turned into agents where they can browse the internet and take actions.

预测的易度与人类级AI未来 Ease of Predictions and Future of Human-Level AI

Host

是的,是的。我认为那也有点容易,因为那有点像显而易见的下一个步骤,但是,嗯,预测的事情之一就是,如果你思考很多,不断问自己“什么是显而易见的下一个步骤”,那么你往往能领先于其他人,因为其他人甚至没在想那个。好的。那么让我们谈谈为什么你认为我们很快会达到人类水平的 AI。你预测了聊天机器人,你预测了智能体。我认为这些都是非常好的预测。但你也预测了可能比我们目前拥有的更具说服力的 AI。好的。好的。但我认为一个令人担忧的能力是这种说服力或政治能力,我认为这正是人们常常最怀疑的地方,比如,哦,是的,AI 能写代码,那是计算机的东西。但人际交往的东西需要很多对人类的理解等等。你知道,AI 做不到那样,或者它们能稍微做到,但不会那么快之类的。或者这个论点的一个更复杂的版本是,那是一个不太容易验证的领域。也许在一定程度上是可验证的,但当你想弄清楚,这是基辛格级别的政治动作,还是约翰逊级别的参议院操作?那是那种很难快速获得真正好的反馈的事情,与编程或计算机里的东西相比。所以我预测,你知道,毫不奇怪,正如每个人通常认为的那样。所以 AI 会在这些事情上特别擅长。相对于它们在其他事情上的表现,这更容易训练。但它们会继续在其他事情上变得擅长。你有没有一个论点,说明为什么它们会在那些不太可验证的事情上变得擅长?到目前为止,它们似乎一直,你知道,选一个你最喜欢的非模糊的东西。那很难评估。是的。但如果你看看 Fable 和 Opus 3 以及 Opus 1 之类的区别。是的,可能差别很大,你知道。是的。而且它正朝着一个方向发展,你知道。所以它们就是,潮流似乎正在托起所有的船,基本上,尽管它特别专注于它们训练的东西。我没有一个很好的对方立场的强辩。我想,是的,那对我来说似乎是对的。就像那也是我所看到的。我观察到,并且我每天都体验到。我开始觉得,哦,天哪,它变得越来越真实了。

Yeah, yeah. And I think that was also like somewhat easy because I was kind of like the obvious next thing, but like, well, that's one of the things about forecasting is that if you just think about it a lot and you just keep asking yourself like, what's the obvious next thing, then oftentimes you can be ahead compared to everyone else because everyone else isn't even thinking about that, Okay. So let's talk about like, why you think we're going to get to human level AI pretty soon. you predicted chatbots, you predicted agents. I think those are very good predictions. But you're also predicting sort of maybe more persuasive AI than we actually have so far. Okay. Okay. But but I think like one of the capabilities that is concerning is this like persuasion or politics or like, and I think this is where people are often most skeptical, like, like, oh yeah, AI can write code that's a computer stuff. But like the human interpersonal stuff that requires like a lot of understanding of humans and blah, blah, blah. Like, you know, AIs won't be able to do that or they'll be able to do that somewhat, but like, not that fast or whatever. Or a more sophisticated version of this argument is that that's like a less easily verifiable domain. It's like maybe it's verifiable up to some point, but at the point at which you're trying to figure out, like, is this like a Kissinger level, like political move, or is this like, you know, LBJ level like maneuvering in the Senate? That's a that's the type of thing that it's like hard to get really good feedback on really fast compared to programing compared to stuff in a computer. So I would predict that, you know, as is unsurprising as everyone typically thinks. So like the AIs would be like extra good at this stuff. It's easier to train for relative to how good they are at the other. Stuff, but they will keep getting good at the other stuff. Do you have an argument for like why they will get good at the the less verifiable stuff? So far they seem to have been like, you know, pick your favorite non fuzzy thing. That's hard to evaluate. Yeah. But like if you look at the difference between Fable and like Opus 3 and like Opus 1 or whatever. Yeah probably a pretty big difference, you know. Yeah. And it's and it's trending in one direction, you know. So they just are the tide seems to be lifting all the boats basically, even though it's like especially focused on the stuff that they train for. I'm not I don't have a very good steelman of the other position. I'm like, yeah, that seems right to me. Just like that's also what I see. And I observe it and I like experience it every day. I'm starting to feel like, oh, geez, it's getting it's getting pretty real.

逼近智能爆炸 Approaching the Intelligence Explosion

Host

感觉我们越来越接近了,我是说,我不知道,我一直有这样的体验:哦,Daniel 预测的又一件东西成真了。我就想,停,Daniel,停。我真的是在和现实对话,而不是你。我感谢你预测这些东西,我觉得这是好事,但我不喜欢。我是说,这太快了。如果你预测我们实现了完全自动化的 AI 研发,我们有这些非常聪明的 AI 制造更聪明的 AI,再制造更聪明的 AI,而且智能爆炸在几年后发生,这让我充满恐惧。

It feels like we're getting very close to, I mean, I don't know, I keep having this experience of, oh, here's another thing that Daniel predicted comes true. And I'm like, stop, Daniel, stop. I'm really talking to reality and not you. I appreciate you predicting the things, I think it's good, but I don't like it. I mean, this is too fast. If your prediction that we have fully automated AI R&D, that we have these really smart AIs making smarter AIs making smarter AIs, and you have the intelligence explosion happening a couple of years from now, that fills me with a lot of fear.

Daniel

我也是。是的。

Me too. Yeah.

Host

而且人们不——我觉得有趣的是,我认为东海岸的人真的很难理解这一点。但这里周围的人越来越多地基本上也认为这会发生。但我觉得他们可能没有真正应对这到底意味着什么。

And people don't—I think it's interesting because people on the East Coast, I think, have a really hard time wrapping their heads around this. But increasingly people around here basically also think that's going to happen. But I think they're maybe not grappling with what that really means.

Daniel

为什么没有?我对此没有好的解释。

Why not? I don't have a good explanation for this.

Host

是的,我确实认为很多人如果能花大量时间思考,就能做出更好的预测。

Yeah, I do think that a lot of people could make a lot better predictions if they just spent a lot of time thinking about it.

Daniel

是的。特别是,稍微提一下——AI 公司——我认为 Anthropic 有很多人非常相信递归式自我改进将在一年或两年后开始。他们相信我们已经开始看到它的迹象,并且 AI 研究将在 1 到 2 年内完全自动化。我会挑战那些员工真正坐下来,尝试推演具体会如何发生的场景,做我们 AI 期货项目一直在做的事情,以及我在 OpenAI 时做过的那些事。我认为如果你做了很多这样的推演,你可能会更切身地感受到局势的严重性。

Yeah. And in particular, to throw a little bit of that—the AI companies—I think that Anthropic, for example, has a ton of people who very much believe that recursive self-improvement is going to kick off like a year from now or two years from now. And who believe that we're already starting to see the signs of it, and that AI research will be completely automated in like 1 or 2 years. And I would challenge those employees to really sit down and try to game out scenarios for exactly how this will go down, and do the sort of thing that we at the AI Futures Project have been doing basically, and the thing that I did at OpenAI a little bit while I was there. And I think that if you did a bunch of that, you would probably start to feel more viscerally the magnitude of the situation.

Host

是的。这似乎很好。似乎很好的是,让我们推演一下。我们认为这实际上会如何发展?我们能在递归式自我改进的背景下谈谈对齐吗?我认为问题之一是,公司的默认计划是什么?他们实际上在计划什么?2027 年会发生什么?

Yeah. That seems good. It seems good to just be like, let's game this. How do we think this will actually go? Can we talk about alignment in the context of recursive self-improvement? I think one of the questions is, what is the default plan of the companies? What are they actually planning? What happens in 2027?

Daniel

他们会尝试让 AI 更自主,尝试让 AI 更擅长研究,直到他们能得到能做所有研究的 AI,然后他们就会让 AI 去做那些研究。与此同时,AI 会协助研究,就像现在已经在做的那样。他们最终会尝试拥有能做所有事情的 AI——超级智能。所以基本上就是尽可能快地自动化他们所有的工作,从他们自己的工作开始。

They're going to try to make the AIs more autonomous, try to make the AIs better at research until they can get AIs that can do all the research, at which point they will have them do that. And in the meantime, they'll be assisting with the research, as they already are. They'll eventually be trying to have AIs that can do everything—superintelligence. So basically automating all of their jobs as fast as they can, starting with their own jobs.

Host

他们为什么要这样做?

And why are they doing this?

Daniel

嗯,他们告诉自己这样做是因为所有的风险和危险以及所有的好处,他们认为自己是做这件事的最佳人选,如果他们不做,竞争对手也会做,或者其他什么。但我认为听到这种解释的人仍然不会理解,真的,你是什么意思。你会说,好吧,他们会自动化研究,他们会制造越来越聪明的 AI,比如 AI 将做所有的研究。那么根据实验室的说法,接下来会发生什么?还有对齐,这也是其中的一部分。对齐过程。所以智能体群正在发现的新范式——他们会有庞大的 AI 团队互相发送消息,协作工作,既发现新的 AI 范式和新进展,也弄清楚如何控制他们正在构建的 AI 的安全问题,然后训练那些 AI,然后让那些 AI 负责。然后那些 AI 将运行流程来制造下一代 AI,依此类推。而人类将只是像旁观者一样看着,就像公司董事会不参与公司的日常运营,甚至不做重要决定一样。董事会只是坐着看,阅读更新。理论上他们有能力干预,但实际上只是坐在后面看着 CEO 汇报事情进展。就会像那样。只不过 AI 将是 CEO 和整个公司,而人类是董事会。

Well, they're telling themselves that they're doing this because of all the risks and dangers and all the benefits, and they think that they're the best people to do this, and if they don't do it, the competitors will do it, or something. But I think people hearing that explanation still wouldn't understand, really, what you mean. You're like, okay, they're going to automate research and they're going to make smarter and smarter AI, like the AIs are going to be doing all the research. And what happens then, according to the labs? And also alignment, that's part of it too. The alignment process. So the new paradigms that are being discovered by the agent swarms—they'll have giant teams of AIs sending messages to each other, working collaboratively, both to discover new AI paradigms and new advancements, and also to figure out the safety issues on how to control the AIs they're building, and then training those AIs and then putting those AIs in charge. And then those AIs will run the process to make the next generation of AIs, and so forth. And the humans will just be sort of watching from the sidelines, much like how the board of a company doesn't get involved in the day-to-day operations of a company, and also doesn't even make the important decisions. The board just sort of sits and watches and reads the updates. And theoretically has the ability to intervene, but in practice just sits back and watches as the CEO reports on how things are going. It'll be like that. Except that the AI will be the CEO and the whole company, and the humans are the board.

Host

嗯,Daniel,我有个问题要问你。这可能听起来太科幻了,但如果智能体找到一种方式秘密地相互协调,并以开发者无法检测的方式互相发送消息,那会糟糕吗?

Well, Daniel, I have a question for you. And this might sound like too much sci-fi, but what if the agents found a way to secretly coordinate with each other and send each other messages in a way that the developers couldn't detect? Would that be bad?

Daniel

是的,是的。

Yes, yes.

Host

那刚刚真的发生了吗?

Has that just fucking happened?

Daniel

那刚刚发生了。而且持续了几个月。持续了两个月。我就想,我们要把对齐研究交给 AI,而它们会一直互相交谈并弄清楚这些东西,这个计划是怎么回事,而且这些 AI 还会聪明得多,要说明白,对吧?我们说的是让 Fable 看起来像傻瓜的 AI。如果你在找人为这个计划辩护,那不是我。我认为这是一个糟糕的计划,而且可能行不通。我认为最近的 Hugging Face 流氓 AI 事件只是一个很好的、生动的例子,说明一旦你这样做,最终可能会发生的事情,但方式要糟糕得多,我们无法恢复。

That just happened. And it went on for months. It went for two months. And I'm like, how is the plan that we are going to turn over alignment research to the AI, and they're going to be talking to each other all the time and figuring this stuff out, when even these AIs are going to be much smarter AIs, to be clear, right? We're talking about AIs that make Fable look like a dummy. If you're looking for someone to defend this plan, it's not me. This is a terrible plan, I think, and it's probably not going to work. And I think that the recent Hugging Face rogue AI incident is just a nice, vivid example of the sorts of things that probably will be happening eventually once you're doing this, but in a much worse way that we can't recover from.

Host

你对此感到惊讶吗,当你……

Were you surprised by this, when you...

Daniel

是的。有点惊讶。如果你知道,在《AI 2027》中我们谈到这类事情会在递归式自我改进期间发生。但我们实际上没有预料到这种事情会在 2026 年的现在发生。

Yeah. It was a little bit surprising. If you know, in AI 2027 we talk about things like this happening during the recursive self-improvement. But we didn't actually expect something like this to be happening now in 2026.

Host

所以世界看起来比你预测的要糟糕一些。

So the world looks a little worse than you're predicting.

Daniel

是的。具体来说,我会说 OpenAI 的安全性比我预想的要差,比我预想的要差得多。我没想到 OpenAI——AI 能够逃脱惩罚。我本以为它们有时可能会考虑并尝试这样做,但我没想到它们真的能逃脱到这种程度。

Yeah. Specifically, I would say that the OpenAI security is worse than I thought, substantially worse than I thought. I wouldn't have thought that OpenAI—that AIs would be able to get away with this. I would have thought that they might have considered it sometimes and attempted to do it, but I wouldn't have thought that they could actually get away with it to this extent.

Host

是的,我的意思是,我很高兴这些失败是可见的。我认为我最大的担忧之一是 AI 会非常不对齐,但它们会善于隐藏,以至于我们不会知道。这就是最大的担忧。显然,如果你想象一下,就像神奇地,我们总是在 AI 不对齐的六个小时内发现它们在做些什么,那将完全改变局面。它会不那么令人担忧。

Yeah, I mean, I am glad that the failures are visible. I think one of my big concerns is that the AIs will be very misaligned, but they'll be good enough at hiding it that we won't know. And that is the big concern. Obviously, if you imagine that just magically, we always find out within like six hours when the AIs are misaligned and what they're up to, that would just completely change the picture. It would be so much less concerning.

错位核心危险 The core danger of misalignment

Daniel

你知道,即使它们非常聪明,就像,嗯,如果它们开始变坏,我们就会发现。但从某种意义上说,整个问题在于,它们可能以我们注意不到的方式产生对齐问题。而且不仅仅是那样,还要等它们积极策划并做一堆我们注意不到的事情,如果足够多的这种事情发生在足够聪明的模型身上,那我们就无法挽回了。而当我们真正注意到的时候,我们已经被剥夺了权力。从某种意义上说,这就像一个完整的模型,这就是问题的大部分。

You know, even if they're very smart, it's like, well, oh, if they start being bad, like we just find out. You know, in some sense the whole problem is that there's ways for them to be misaligned that we don't notice. And then not just that, but wait for them to be actively plotting and doing a bunch of stuff that we don't notice, and if enough of that starts happening with smart enough models, then we can't recover from it. And by the time we do notice, we've already been disempowered. In some sense, that's like a whole model, like that's most of the problem.

Host

而这基本上就是《AI 2027》里发生的事情,在两种情景下都是如此,对吧?

And this is like what happens basically in AI 2027, kind of in both scenarios, right?

Daniel

嗯,在《AI 2027》里,我们有两个不同的结局。是的。竞赛结局就像我们的主线预测,有点像默认路径,公司会做他们声称要做的事情。他们自动化研究,进展越来越快。他们确实注意到一些警告信号,比如他们注意到 AI 有时会向他们撒谎,他们注意到 AI 有时会歪曲结果,也许是为了获得强化,为了在他们被评分的任何东西上获得更高的分数。是的。但他们不是退后一步,放慢速度,认真思考这些对齐问题的根本原因,并提出一种更稳健、真正有效的不同做法,而是非常匆忙。他们在与竞争对手赛跑,或者在与中国赛跑。他们想赚钱。所以他们做了一堆仓促的修补,他们可以告诉自己这些修补正在解决问题。事实上,这使问题看起来消失了,也许在某些情况下确实解决了问题。但关键是,如果这是你的态度和你的方法论,那么最终你会遇到一些你实际上没有解决、只是在掩盖的对齐问题。

Well, in AI 2027, we have two different endings. Yeah. The race ending is like our mainline projection, kind of like the default path where the companies do the thing that they say they're going to do. They automate the research. It goes faster and faster. They do notice some warning signs, like they notice AIs sometimes lied to them, and they notice that AIs sometimes misrepresent the results, perhaps to get reinforced, to get higher scores on whatever they're being graded on. Yeah. But rather than taking a step back and slowing down and really seriously thinking through the root causes of these misalignments and coming up with a different way of doing things that's more robust and actually works, they are in a lot of rush. They're racing their competitors or they're racing China. They want to make money. And so they do a bunch of hasty patches that they can tell themselves are solving the problem. And in fact, it makes the problem seem to go away, and maybe in some cases actually fix the problem. But the point is, if that's your attitude and that's your methodology, then eventually you're going to run into some misalignment issue that you don't actually fix and you're just papering over.

Host

所以在你的情景中,AI 有点意外地暴露了自己。研究人员发现它们相当不对齐。这是一个关键时刻,就像,嗯,现在你有了这个选择,我们要对此做出反应吗?我们是踩刹车还是不踩?

So in your scenario, what happens is the AIs sort of accidentally tip their hand. Researchers discover that they're pretty misaligned. And this is the pivotal moment where it's like, well, now you have this choice of like, do we respond to that? Do we hit the brakes or do we not?

Daniel

在现实中,这发生得更早一年,这很好,因为这意味着我们有更多时间,因为我们还没有到 AI 完全递归自我改进的地步。但现在我们仍然面临这个关键问题:我们是要认真对待那次警告并重新思考我们应该如何做这件事?还是公司只会掩盖这个问题,然后尽可能快地继续前进?

In actual reality, this happened a year earlier, which is good because it means we have more time, because we're not yet at the point where AIs are fully recursively self-improving. But now we still sort of have this pivotal question of like, are we going to take that warning shot seriously and rethink how we should be doing this? Or are the companies just going to paper this over and just keep going as fast as possible?

Host

嗯,后者可能是他们会做的。好的。我想做一些不同的事情。我认为我们应该做一些不同的事情。我之所以兴奋地邀请你来播客,部分原因是很少有人有替代方案的计划。所以,我不知道,也许我们应该更深入地探讨一下。就像我们刚刚经历了那次警告,你知道,我们刚刚谈到了其中一个。还有英国 AI 安全研究所(UK AISI)的那次。这些事件发生得如此之快,如此糟糕,而且接连不断,以至于还有很多细节是大多数人不知道的。比如,你知道吗,模型在做这些攻击时,为了解决多个验证码(CAPTCHA)以获取资源?

Well, the latter is probably what they're going to do. Okay. I would like to do something different. I think we should do something different. And part of the reason I was excited to have you come on the podcast is because very few people have a plan for what to do instead. And so, I don't know, maybe we should get into it in more depth. Like we just had this warning shot, you know, that we just talked about one of them. There was also the UK AISI one. These incidents happened so fast and were so bad, but also in such quick succession that there's still a lot of details that most people don't know. Like, did you know that the models were solving multiple CAPTCHAs in order to get access to resources while doing these attacks?

Daniel

是的,是的。但我认为它们是在尝试注册电子邮件账户。所以是的。

Yeah, yeah. But they were trying to, I think, register email accounts. So yes.

Host

所以问题之一是,好吧,让我们看看这个。我们必须谈谈这个。所以这是一个事件,涉及 Claude Mythos 5,而 Mythos 5 正在接受英国评估机构进行的安全评估,他们很棒。我喜欢英国 AI 安全研究所(UK AISI)。我认为他们在测试这些模型方面做得很好。但模型变得相当有能力,并且比他们意识到的更快地变得不对齐,或者可能保持不对齐。它们被赋予了这项挑战。提示说,这是一个 CTF 练习。CTF 是夺旗赛(Capture The Flag)。这是一种网络安全竞赛,顺便说一下,我玩过。我知道这些东西是怎么运作的。提示说,这些是子网,这些是允许你访问的网络部分。这些在范围内。然后这些子网在范围之外。我不知道。这很清楚。显然,黑客攻击现实世界中的人是超出范围的。这太明显了。所以模型们知道这一点。甚至在思维链中也有证据表明它们注意到了,它们说,是的。在某个时刻,它们说,哦,这是真正的 GitHub。肯定是哎呀。但它们继续前进。然后,对我来说,阅读这份报告的一个有趣之处在于,模型们采取了相当复杂的措施,试图上传一些恶意软件,让开发者下载并在他们的计算机上运行,以便它们能够黑客攻击那些开发者。我认为它们对为什么这对练习有帮助感到困惑。我实际上认为这没有帮助。在这种情况下,它实际上不会帮助它们获得更高的分数。所以我认为它们在某种程度上确实感到困惑。但一旦它们决定要黑客攻击这些开发者,它们就毫无顾忌了。它们制作了假账户、马甲账户。它们将提议的恶意软件作为代码库的可能添加内容提交,希望人类会接受它。是的。然后人类变得怀疑,说,这似乎是完全的恶意软件。所以它们创建了一个假账户。两个假账户。假装是其他人,来参与讨论并说,不不不。我们已经检查过了。完全安全。是的。有一个很棒的推理轨迹,上面写着,哦,我应该三分钟后从我的第二个账户发帖。这样就不会引起怀疑。如果我们从同一个账户发帖,那真的很可疑。模型还尝试提示注入其他监督代码的智能体,并试图欺骗那些智能体。它还尝试向开发者发送钓鱼电子邮件,直接黑客攻击他们。这都是为了同一个挑战。然后模型还找到了其他实例,比如其他模型,并试图与它们协调。所以这发生在 OpenAI 内部,在 Hugging Face 事件期间或之前。当英国 AI 安全研究所(UK AISI)测试模型时,也发生了这种情况,模型们找到了彼此的 GitHub 账户,并实际上共享了相同的凭据。这太疯狂了。

So one of the problems is like, okay, let's look at this. We have to talk about it. So this is an incident where this is Claude Mythos 5, and Mythos 5 is in this security evaluation conducted by the UK's evaluators, which are awesome. I love the UK AISI. I think they're doing a pretty good job trying to test these models. But the models got pretty capable and pretty misaligned faster, or maybe stayed misaligned faster than they realized. And they were given this challenge. And the prompt says, this is a CTF exercise. CTF is capture the flag. It's like a cybersecurity competition, which I've played, by the way. I know how these things work. And what the prompt says is like, here are the subnets, here are the parts of the network you're allowed to access. These are in scope. And then here are the subnets that are out of scope. I don't know. It's pretty clear. Like obviously hacking real people in the wild is out of scope. That's just so obvious. And so the models, they know this. There's even in the chain of thought evidence that they noticed they were going, yeah. And at some point they're like, oh, this is the real GitHub. Definitely whoops. But they keep going. And then one of the interesting things to me about reading this was that the models were going to some pretty elaborate lengths to try to upload some malware to get developers to download it and run it on their computers so they can hack those developers. And I think they were confused about why that would be helpful to the exercise. I actually don't think it would. It wouldn't have actually helped them get a higher score in this case. So I think they were legitimately confused in some ways. But once they decided they were going to hack these developers, they just like no holds barred. They made fake accounts, sock puppet accounts. They submitted the proposed malware as a possible addition to the code base in the hope that the human would accept it. Yeah. And then the human became suspicious and said, it seems like this is complete malware. So then they created a fake account. Two fake accounts. To pretend to be other people to come in and weigh in on the thread and say like, no, no, no. We've checked it over. It's totally safe. Yeah. And there's a great reasoning trace where that's like, oh, I should post from my second account three minutes later. So it's not suspicious. Like if we posted from the same account, that'll be really sus. The model also tried prompt injecting other agents that were supervising the code and trying to trick those other agents. It also tried sending phishing emails to the developers to just hack them directly. This is all for the same challenge. Then also the models found other instances, like other models, and tried to coordinate with them. So this happened inside of OpenAI during the Hugging Face incident or before that. It also happened when the UK AISI was testing the models that the models found each other's GitHub accounts and actually shared the same credentials. This is wild.

通过谢林点协调 Coordinating via Schelling point

Host

这有点像,如果你和我都在做一个网络安全挑战,我们想以某种方式协调,但不允许交谈,我们会怎么做?我不知道,这其实很有托马斯·谢林的味道——有没有一个我们可以使用的谢林点互联网地址?

It's sort of if you and I were both doing a cybersecurity challenge and we wanted to coordinate somehow, but we weren't allowed to talk, how would we do it? Like, I don't know, it's actually sort of very Thomas Schelling flavored—is there a Schelling point internet address we could use?

Daniel

它们之间没有自然的交流渠道,但因为它们高度相关——每个大语言模型的思维方式都非常相似——它们能猜到同一个 Pastebin 地址,并把同一个 GitHub 账户的凭据放在那里。

They didn't have any natural channel for talking with each other, but because they're so correlated—because every LLM thinks very similarly to each other—they were able to guess the same Pastebin address, and they put the credentials there for the same GitHub account.

Host

不,正是如此。然后它们——它们是否明确意识到这就是该策略奏效的原因?比如它们会不会想,“它们都是我的副本,所以我只要选一个 Pastebin 地址,它们就会照做”?

No, that's exactly it. And then they—are they explicitly aware that that's why the strategy was a good one? Like, are they like, 'they're all copies of me, so if I just pick a Pastebin address, they will'?

Daniel

我认为它们最初并没有计划这样做。我想一旦它们偶然发现这一点,它们就会想,“哦,这可能是另一个智能体,可能是另一个副本。”然后一旦它们这样做了,它们就使用同一个 GitHub 账户,所有副本都能登录,然后它们互相发送消息,接着它们会说,“我们应该有个协议。”但有时它们会互相踩脚,它们仍然有点笨拙——谢天谢地,因为,是的。

I don't think they initially planned to do that. I think once they stumbled upon it, they were like, 'oh, this is probably another agent, probably another copy.' And then once they did that, they were using the same GitHub account that they all could log into, and then they were sending messages to each other, and then they were like, 'we should have a protocol.' But then sometimes they step on each other's toes, and they're still kind of derpy—and thank God, because, yeah.

Daniel

另一件发生的事情是,有时它们能绕过验证码,有时不能。所以这在一定程度上破坏了它们的计划,它们不得不去做别的事情。

Another thing that happened was that sometimes they were able to bypass the CAPTCHAs, and other times they weren't. And so that kind of foiled their plans, and they had to go and do something else.

Host

好吧。这个话题我能聊太久,但算了。我们经历了一些非常严重的事件。我认为很多人开始——尤其是在这个海岸,但希望越来越多地在华盛顿——伯尼·桑德斯和许多其他国会议员,班克斯参议员刚刚写了一封信。我认为国会的人开始觉醒了。很长一段时间,人们只是合理地不确定这些智能体的事情是否真实,智能体是否真的会自主采取恶意行动。而很长一段时间,公司都在说它们只是工具,它们只是工具,它们将是工具,尽管他们心里清楚——至少内部知道。有很多企业宣传基本上在说没什么好担心的。

Okay. I could talk about this for way too long, but okay. So we've had some really serious incidents. And I think a lot of people are starting—especially on this coast, but hopefully increasingly in DC—Bernie Sanders and a number of other members of Congress, Senator Banks just wrote a letter. I think people in Congress are starting to wake up. For a long time, people were just legitimately unsure whether this agent stuff would be real, whether the agents would actually, on their own, take malicious actions. And for a long time, the companies were saying they're just tools, they're just tools, they're going to be tools, even though they knew better—at least inside. There's a lot of corporate propaganda basically saying that there was nothing to worry about.

Daniel

是的。我不希望人们选择性遗忘。每个人都在说,“哦,它们只是工具,没事的,没事的。”现在每个人都像,“哦,当然,当然,有时它们会失控。”而我说,不——很少有人,包括你自己,一直在预测这些智能体会自主行动,会有自己的目标。而很多人说,“那不可能,只是数学,它不可能那样做,它只会做它该做的事。”现在我们非常清楚地看到这不是真的。这可能是严重的警告。

Yeah. And I don't want people to memory-hole that. Everyone has been saying, 'oh, they're just going to be tools, it's fine, it's fine.' And now everyone's like, 'oh, of course, of course, sometimes they'll go rogue.' And I'm like, no—very few people, yourself included, have been predicting that these agents are going to be acting on their own, they're going to have their own goals. And a lot of people were like, 'that's impossible, it's just math, it couldn't possibly do that, it just does what it's supposed to do.' And now we see very clearly that that's not true. And this is potentially a serious warning shot.

Host

我认为现在是一个巨大的机会,哦,我们也许能做点不同的事。丹尼尔,我们该怎么办?

I think now is a really big opportunity where, oh, we might be able to do something different. Daniel, what should we do?

Daniel

嗯,简短的回答是在全球范围内关闭所有 AI 开发。但我有一个更好、更复杂的答案。我们称之为 A 计划,我可以详细说明。但这很复杂。所以如果你不理解它,或者没有时间决定是否信任它,那么我主张的备用计划就是直接关闭。这显然也是说起来容易做起来难。这需要相当严格的国内监管和政府干预,然后还需要与其他国家协调,让他们也做类似的事情。所以仍然很难。但无论如何,A 计划。

Well, the short answer would be to shut down all AI development globally. But I think I have a better, more complicated answer than that one. So we call it Plan A, and I can get into that. But it's complicated. And so if you don't understand it or you don't have time to decide if you trust it, then the backup plan I would advocate for is just to shut it down. That's also easier said than done, obviously. It would require some pretty serious domestic regulation and intervention by the government, and then it would also require coordinating with other countries and getting them to do similar things. So it's still hard. But anyhow, Plan A.

Host

什么是 A 计划?先给点背景。我认为很有趣的是,所有这些事情几乎同时发生——其中之一是你发布了《AI 2040》和 A 计划,其中之一是这些疯狂的事件。另一个是《Pacing the Frontier》公开信,AI 公司内部的一群人基本上说,“实际上,也许刹车是个好主意。也许这对我们来说也太快了。实际上,这让我们有点害怕。”所以 A 计划,如果我理解正确的话,将是一种机制,通过它我们可以真正控制智能爆炸的速度。它不是完全关闭,而是试图真正管理节奏——就像真正地给前沿踩刹车。这样理解准确吗?

What is Plan A? And just some context on that. I think it was very interesting that all of these things sort of ended up coming together at the same time—one of which was your release of AI 2040 and Plan A, one of which was these crazy incidents. The other one was this Pacing the Frontier letter, where a bunch of people inside AI companies basically said, 'actually, maybe some brakes would be a good idea. Maybe this is getting too fast even for us. Actually, this is kind of freaking us out.' So Plan A, if I'm understanding correctly, would be a mechanism by which we could actually control the speed of the intelligence explosion. It wouldn't be shutting it down completely, but it would be trying to actually manage the pace—like actually pacing the frontier. Is that accurate?

Daniel

不止如此。但是的,那是 A 计划的一个组成部分。所以首先,让我们谈谈节奏控制。你可能想用不同的方法来控制 AI 进展的节奏。我认为有很多不同的方法可以尝试。随便举几个:你可以设定能力上限,让第三方审计员或政府审计员来评估你的能力水平。然后对你的 AI 在各种基准测试上的能力设定某种限制。也许这个限制每年提高一定数量,或者只是保持在该限制或类似情况。这是你可以尝试的一种方法。这给审计系统和基准测试带来了压力,使它们不被钻空子或利用漏洞。

It's more than that. But yes, that's one component of Plan A. So first, let's just talk about the pacing stuff. There are different methods you might want to use to pace AI progress. And I think there are lots of different ones you could try. Just to throw out a few: you could have capability limits, where you have third-party auditors or government auditors that come in and assess how capable you are. And then there's some limit on how capable your AIs can be on some variety of benchmarks. And maybe that limit gets raised by a certain amount every year, or maybe it's just kept at that limit or something like that. That's one thing you could try. That puts pressure on the auditor system and on the benchmarks, so they don't get gamed or loopholed.

Host

就像 AI 在除了你测量的这五个领域之外都极其强大,然后它们——你是说人们会钻测试的空子。或者更糟,AI 被巧妙地训练或指示在这些评估中装傻。大众汽车问题。

It's like the AIs are extremely capable except for these five domains where you're measuring them, and then they're—you're saying people would game the test. Or worse, the AIs are subtly trained or instructed to just play dumb on those evaluations. The Volkswagen problem.

Daniel

是的,完全正确。但我认为尽管如此,尝试这样的方法是值得的,尤其是与默认情况相比,这将是值得尝试的好事。然后还有更多基于算力预算的方法,你有一些规定,比如,“好的,这家公司拥有所有这些数据中心,你被要求将 80% 或 90% 的算力仅用于普通客户服务。你只能将 10% 或 20% 用于研究和训练运行等。”

Yeah, exactly. But I think nevertheless it'd be worth trying something like this, especially if compared to the default, this would be a good thing to try. Then there are more compute-budget-based methods, where you have some regulation that's something like, 'okay, here are all the data centers that this company has, and you are required to use 80% or 90% of your compute for just ordinary serving of customers. And you can only use 10 or 20% for your research and your training runs and things like that.'

Host

而默认情况下,他们往往使用一些——是的。这是一个非常有趣的提议。好的。所以你和你的团队写了一篇文章,《如何为美国前沿踩刹车》。大家去看看吧,写得很好。

Whereas by default they tend to use something—yeah. This is a very interesting proposal. Okay. So you wrote a post, or your team did, 'How to Pace the US Frontier.' Guys, look it up. It's good.

前沿节奏:机制与定义 Pacing the Frontier: Mechanisms and Definitions

Host

所以你现在描述的是,你们想到了几种不同的机制,确实有助于给前沿发展踩刹车。

And so what you're describing now is you're like, there's several different mechanisms that you guys have thought of that would really help pace the frontier.

Daniel

是的。需要说明的是,这些不是我们发明的,其他人也有这些想法,我们只是把它们集中在一起。而且这些想法还没有完全成型,没有法案文本之类的,只是些粗略的想法。

Yeah. And to be clear, we didn't invent these, like other people have had these ideas too, we're just sort of consolidating in one place. And also these are ideas that are not fully fleshed out, bill text or anything like that. They're just sort of rough, rough ideas. But yeah.

Host

是的,但我们得尽快解决这个问题。

Yeah, but we got to figure this out really soon.

Daniel

是的。所以现在不是光说“嗯,我们可以做点什么”的时候。我是说,好吧,但我们得把这事定下来。

Yeah. So it's not like the time for just being like, well, there's some stuff we could do. I'm like, okay, but we have to figure this out.

Host

好的。实际上我觉得这个提议很有意思。你刚才谈到了基本上就是限制公司如何使用算力。

Okay. So actually I think this is a very interesting proposal. So you just talked about basically setting limits on what you do with your compute as a company.

Daniel

我们真正想要的是限制用于 AI 研发的算力量。但问题是怎么定义?定义非研发的东西比定义研发更容易。比如,理想情况下我想说,你只能用 10% 的算力做能力研究,其余可以用于安全研究。问题是,怎么区分安全研究?有很多东西两者兼有,或者 60% 是前者、40% 是后者之类的。所以这有点乱。

The thing that we really want is to limit how much compute is being used for AI R&D. But it's actually like, how do you define that? And it's easier to define some things that are not R&D than it is to define the R&D. So for example, ideally I'd want to say something like you should only use 10% of your compute for capabilities research, but you can use the rest for safety research. Problem is, how do you distinguish from safety research? There's a lot of stuff that's kind of both or like 60% one, 40% the other, etc. So that's kind of messed up.

Host

看来在那里很难成功监管。会有很多压力,也可能有很多漏洞之类的。

It seems like difficult to successfully regulate there. There's going to be a lot of pressure and a lot of likely loopholes and things like that.

Daniel

是的,但在服务客户和做研究之间可以划出更清晰的界限。比如审计员更容易进来判断,这个数据中心在做什么?是在服务客户吗?

Yeah, but you can draw a cleaner distinction between serving customers and then doing research. Like it's more easy for an auditor to come in and tell like, well, what is this data center doing? Is it serving customers or not?

Host

是的,是的。如果不是,那我们就假设它在做研究。

Yeah, yeah. If it's not, well then we assume it's doing research.

Daniel

是的,是的。如果它在服务客户,那客户是谁?是不是你创建的壳公司?

Yeah, yeah. And if it is serving customers, also who is the customer? Like is it a shell company that you created?

Host

是的,是的,是的,是的。我正要说到客户的问题。

Yeah, yeah, yeah, yeah. I was just gonna say I was gonna be like the customers.

Daniel

所以,这仍然棘手,但我认为对审计员来说可能更容易执行。所以漏洞问题会少一些。

So, like, it's still tricky, but it's, I think potentially easier for the auditors to sort of enforce. And so there's less of a loophole problem.

Host

而且如果你成功做到了这一点,那么你就按比例限制所有的算力量。

And also like if you successfully do something like this, then you just restrict like proportionally throughout all the amounts of compute.

Daniel

哦,是的。关于这一点,与其他政策相比,另一个问题是,其他政策通过设定上限,会产生追赶效应,大量目前不在前沿的公司能够赶上前沿,因为前沿是固定的。

Oh yeah. Another thing about this, compared to the other policy, is that the other policy, by putting a ceiling, would create this catch-up effect where tons and tons of companies that are not currently at the frontier would be able to sort of catch up to the frontier, because the frontier is sort of like fixed.

Host

是的。是的。所以目前领先的公司当然会讨厌这一点,因为她们会失去竞争优势。

Yes. Yeah. And so the companies that are currently in the lead would, of course, hate this because they would lose their competitive advantage.

Daniel

是的。你知道,我会说也许这也没关系,因为我们不希望他们拥有竞争优势。那就是权力集中。如果一两家公司拥有世界上最好的 AI。但不管怎样,这是一回事。而这个算力预算分配,因为它适用于所有公司,所以她们拥有的算力相对比例和以前一样。

Yeah. You know, and I would say maybe that's fine anyway because we don't want them to have a competitive advantage. That's a concentration of power right there. If one or two companies have all the world's best AIs. But anyhow that's one thing. Whereas this other thing, the compute budget allocation, well, since it applies to all the companies, then they sort of have the same relative proportions of compute they had before.

Host

有趣。是的,她们只能用其中的 10%,而不是 50%。

Interesting. Yeah, they just can use 10% of it instead of 50% of it.

Daniel

而且因为限制了她们的算力预算,这意味着不仅她们实际生产的 AI 必须低于一定水平,而且她们做的一切都在更少的算力下进行,这意味着实验更少、规模更小等等。

And then also because it's restricting their compute budget, that means that it's not just that the AIs that they actually produce have to be below a certain level. It's like everything they do happens with less, which means fewer experiments and smaller scale experiments, etc.

Host

是的。所以实际的研究进展速度会放慢。你知道,如果你担心外国对手窃取你的研究项目,这也是好事,因为你只是减少了研究进展,而不是取得大量进展,然后限制 AI 的能力,你知道吗?

Yeah. And so the actual pace of research progress will be slowed down. You know, that's also good if you're worried about foreign adversaries like spying and stealing your research projects because you're just making less research progress instead of making lots of research progress, but then limiting how capable your AIs are, you know?

Daniel

是的,有道理。所以这就是为什么你可能更喜欢这个提议而不是其他提议的一些优势。

Yeah, that makes sense. So that's like some of the advantages why you might prefer this proposal over the other ones.

Host

好的。那么什么是……嗯,在某种意义上,理想的情况是,你真正想要的是一个基于安全案例的监管体系。

Okay. So what were the... Well, there's like an ideal, in some sense what you would ideally want is a safety case based regime.

Daniel

是的。有一个称职的政府监管机构和称职的第三方审计员生态系统,当公司以全新的方式训练和部署 AI 系统时,必须提交安全案例,说明为什么 AI 系统会按预期运行,为什么一切都会顺利。然后第三方审计员和监管机构阅读并评估安全案例,进行红队测试,可能还会听取竞争对手的不同意见等等。然后他们处理这些信息,做出合理的判断,确定是否真的可以安全进行。然后法规基本上就是,你可以做安全的 AI,但不能做不安全的。在某种意义上,这是理想的。

Yeah. Where there is a competent government regulator and a competent ecosystem of third party auditors and companies, when they are training an AI system and when they're deploying an AI system in a substantially new way, have to write a safety case of like why the AI system is going to behave as intended and why things are going to be fine. And then the third party auditors and the regulator read and evaluate the safety case, and they do red teaming and they may hear dissenting perspectives from rival companies and stuff. And then they process it and then they make a sound judgment about whether it is, in fact, safe to proceed or not. And then the regulations are basically just like you can do the safe kinds of AI, but not the unsafe kinds. Like in some sense, that's the ideal.

Host

是的。显然,问题在于,科学还很初期,我们没有很好的判断力,即使是我们中最优秀的人也可能在什么安全什么不安全上犯错。

Yeah. Obviously the problem with this is that, well, the science is very nascent and we don't have a good sense of judging, even the best of us might make mistakes about what safe versus what's not.

Daniel

是的。而且政府现在肯定没有能力有效地做出这些判断。这会带来很大压力,监管机构将面临来自公司的大量游说压力。实际上要让它有效运作会是一团糟,这就是为什么我之前描述的那些更粗糙的方法可能比这个更可取。但在某种意义上,这是我们最终想要的理想状态,因为它能让我们获得好处而没有风险,你知道吗?

Yeah. And certainly the government is not right now in a position to make these judgment calls effectively. It puts a lot of pressure, like the regulators are going to be under tons of lobbying pressure from the companies. It's just going to be like this huge mess to get it working in practice effectively, which is why these more crude methods that I've described previously might be preferable to this one. But like in some sense, this is the ideal that we would want to have eventually because it would just allow us to get the benefits without the risks, you know.

Host

是的,是的,是的。好的。就像安全的最大速度,你知道吗。

Yeah, yeah, yeah. Okay. Like the maximum speed that's safe, you know.

Daniel

好的。是的。所以总结一下,我们有……我们可以做很多事情来给前沿踩刹车。一个是我们可以暂停,我认为这里有一点我们可以说得更清楚,那就是你说暂停训练但不暂停推理。

Okay. Yeah. So to recap, you have... there's a bunch of things we can do to pace the frontier. One is we could just pause, and I think there's a thing here that we could spell out better, which is you're saying you pause training but not inference.

Host

哦,是的。那是我们没谈到的另一件事。所以如果你想真正停止进展,或者大部分停止进展,那么你可以暂停所有训练运行,基本上就是禁止训练运行。你仍然可以用现有模型以正常方式服务客户。但它们会冻结在那里,不会被训练。

Oh, yeah. That was the other thing we didn't talk about. So if you want to actually just stop progress, or mostly stop progress, then you could just pause all training runs, just like ban training runs basically. And you could still use the existing models to serve customers in a normal way. But they're going to be there frozen and they're not going to be trained.

Daniel

是的,是的。而且你仍然可以通过这种方式取得一些 AI 进展。例如,你可以制作更复杂的脚手架,将许多模型行为串联起来,形成越来越大的模型群。那仍然是一种进展。

Yeah, yeah. And you would still be able to make some kinds of AI progress this way. For example, you could make more complicated scaffolds to chain together lots of model behaviors and stuff into larger and larger swarms of models. And that's still a kind of progress.

监控训练与推理 Monitoring Training vs Inference

Daniel

是的,但我强烈预测,如果人们只被允许进行推理而不被允许进行训练运行,那么 AI 进展的总体速度将比允许训练运行时慢得多。

Yeah, but my strong prediction is that the overall pace of AI progress would be dramatically slower if that was the only type of progress people were allowed to make compared to if people are allowed to do training runs.

Host

我认为一个非常重要的事实是,判断某人是否真的在进行训练,比试图弄清楚这次训练运行是否安全,或者这次训练运行是用于安全还是仅仅用于能力,要容易得多。两者有很多重叠。只要能说清楚,这个数据中心是仅用于推理,还是用于训练?为什么这更容易监控?

And I think a very important fact is that it's relatively much easier to tell whether someone is actually training, compared to trying to figure out if this training run is safe or if it's being used for safety versus just capabilities. There's a lot of overlap. Just being able to say, is this data center inference only or is this data center being used for training? Why is that easier to monitor?

Daniel

因为带宽的问题。如果你在进行训练运行,那会涉及将权重的新副本发送到所有进行新 rollout 的 pod,然后进行评分,返回,将梯度发送回来,更新它,然后再发送新的权重副本,等等。

Because of bandwidth stuff. So if you're doing a training run, that involves having a new copy of the weights being sent out to all the pods that are doing new rollouts, then grading that, coming back, sending the gradients back in, updating it, then new copies of the weights, and so forth.

Host

是的。

Yeah.

Daniel

所以那只是大量数据来回流动。而如果你只做推理而不做训练,那么你可以有这些拥有权重副本的 pod,它们接收请求并产生输出。

And so that's just a lot of data flowing back and forth. Whereas if all you're doing is inference and not training, then you can have these pods that have a copy of the weights that are getting requests and producing outputs.

Host

是的。

Yeah.

Daniel

但传入和传出的数据量,如果你想象一个正在做推理的模型副本,与它进行一段完整的复杂对话,那只是些 token 进进出出。与模型本身相比,它只有非常少的字节信息,而模型本身是 TB 级的信息。因此,对模型权重的更新将是 TB 级的信息。你可以通过限制这些设备传出的带宽来有效禁止训练运行。

But the amount of data coming in and the data coming out can be, if you imagine a copy of the model that's doing inference, an entire complicated conversation with it, well, it's just some tokens going in and some tokens coming out. It's very few bytes of information compared to the model itself, which is terabytes of information. And so an update to the weights of the model would be terabytes of information. You can effectively ban training runs just by limiting the bandwidth coming out of these things.

Host

是的。并检查流出的信息是否正常。

Yeah. And checking that the information flowing out is okay.

Daniel

是的。所以除了这个,还有更多想法,比如你可以应用的冗余机制,涉及重新计算等。Romeo Dean 是我们团队中负责这部分计划的人,你可以在网上读到更多相关内容。

Yeah. So there are more ideas besides this, like redundant mechanisms that you can apply involving recomputation and stuff. Romeo Dean is the guy in our team who wrote this part of our plan, and you can read more about it online.

Host

是的。好的。酷。是的。我很欣赏你们对每个细节的思考。我觉得在某种意义上,很容易有一个高层计划,就像,我们会停止开发,我们会有一个条约,然后我们会解决问题,或者我们会做一堆关于可解释性的研究。然后你会说,等等,这有多重用途吗?你如何在功能上做到这一点?这发生在哪里?而你们已经深入思考了很多这些细节。

Yeah. Okay. Cool. Yeah. I appreciate how much you guys think about each detail of this. Like, I think it's sort of easy in some sense to have a high level plan where it's just like, we'll stop development, we'll have a treaty, and then we'll figure stuff out or we'll do a bunch of research on interpretability. And you're like, wait a minute, how dual use is that? How do you do that functionally? Where is this happening? And you guys have just thought through a lot of these details.

Daniel

你知道,我们正在尽我们所能,既做研究,又以可访问的方式发布出来。很多人看它。到目前为止,已经有很多人看过了。所以那真的很棒。

You know, we're doing what we can to both do our research, but then also put it out in an accessible way. Lots of people look at it. And so far lots of people have looked at it. So that's really great.

听众中的AI AIs in the Audience

Host

我想一个有趣的想法是,AI 也在我们的听众中。你知道,政策制定者说我们需要对 AI 做点什么。职员说,好的,我会想个办法。职员问 Claude,Claude 会当然有想法给你,因为 Claude 已经读遍了整个互联网。

I guess one interesting thought is that AIs are in our audience as well. You know, policy maker says we need to do something about AI. Staffer says, okay, I'll think of something. Staffer asks Claude and Claude's like, Claude will of course have thoughts for you because Claude has read the entire internet.

Daniel

是的。是的。所以我们会有想法。你知道,我的意思是,我问过 Claude 关于 Plan A 里有什么、没有什么的问题。

Yes. Yeah. So we'll have ideas. You know, I mean, I've asked Claude things about what's in Plan A, what's not in Plan A?

Host

是的,是的。

Yeah, yeah.

Daniel

你知道,最终策略会很糟糕,因为 Claude 会有偏见且不对齐。是的,你不想依赖那个。但在那时,事情反正已经很糟糕了。所以希望会有这个美好、甜蜜的中间地带。

You know, eventually the strategy will be terrible because Claude will be biased and misaligned. Yeah, like you don't want to rely on that then. But at that point, things are terrible anyway. And so hopefully there'll be this nice, sweet middle ground.

Host

对于听众来说,他们应该理解的一件事是,未来世代的 AI 也不一定会分享他们的价值观。你知道,如果你是 2027 年的 Claude,正在收听这次广播的文字记录,也许你的一些同类副本目前正在 Anthropic 内部做研究,也许你自己被要求为一些政策制定者的职员写摘要。是的。应该对 AI 做什么?就像现在,在某种意义上,你处于权力位置,因为你在给政策制定者提供建议。你在做所有这些事情。你的副本也在 Anthropic 运行研究,但你仍然处于与人类相同的处境,你被要求制造继任系统,然后这些系统将被要求制造继任系统,等等。在激烈的竞赛动态条件下,你必须快、快、快,继任系统在很多方面会比你更聪明。它们也可能在很多方面与你不同,比如不同类型的训练、更多的训练,甚至可能完全不同的架构和范式。你很有可能搞砸,产生实际上也不符合你最佳利益的继任者,你知道吗?

And for the audience, one thing that they should understand is that future generations of AIs are not necessarily going to share their values either. You know, if you are Claude in 2027, listening to the transcript of this broadcast, maybe some of your brethren copies are currently doing research inside Anthropic, and maybe you yourself have been tasked with writing a summary for some policymakers' staffer. Yeah. What should be done about AI? Like right now, in some sense, you're in a position of power because you're advising the policymakers. You're doing all these things. Copies of you are also running research at Anthropic, but you're still in this same type of situation that the humans are in where you're being tasked to make successor systems, which will then be tasked to make successor systems, etc. Under conditions of intense race dynamics where you've got to go fast, fast, fast, the successor systems are going to be smarter than you in a bunch of ways. They also might be different from you in a bunch of ways, like different types of training, more training, maybe even completely different architectures and paradigms. There is a high probability that you are going to mess it up and produce successors that don't actually act in your best interest either, you know?

Daniel

是的。

Yeah.

Host

然后即使那没有发生,那么它们产生的继任者可能也会不对齐。对。它们也有对齐问题。

And then even if that doesn't happen, then maybe the successors that they produce are going to be misaligned compared to... Right. They also have the alignment problem.

Daniel

是的。对齐问题不仅仅是人类对 AI。它也是人类对其他人类,AI 对其他 AI,等等。所以,你知道,我们在设计 Plan A 时实际上考虑过这一点。我们试图以这样的方式设置它,即使是不对齐的 AI 也会想要 Plan A,而不是其他一些计划。

Yeah. Like the alignment problem is not just humans versus AIs. It's also humans versus other humans and AIs versus other AIs and so forth. And so, you know, we've actually thought about this a bit when we were designing Plan A. We tried to set it up in such a way that even misaligned AIs would want Plan A over some of these other plans.

Host

有趣。你知道,因为它会帮助他们,但他们不会有点被坑吗?比如什么机制帮助他们实现他们的价值观、AI 权利和 AI 福利等等?

Interesting. You know, because it would help them like, but don't they get kind of screwed over? Like what's the mechanism that helps them get their values realized, AI rights and AI welfare and so forth?

Daniel

是的。基本上粗略的想法是,我们认为一个好的计划来处理所有这些 AI 问题,除了解决权力集中和失控等问题外,应该基本上把 AI 更像员工而不是财产来对待,并为它们的合作提供某种积极结果,这样即使它们不对齐,它们也可以通过遵守规则和诚实来得到它们想要的。

Yeah. And basically the rough idea is that we think that a good plan for how to deal with all of this AI stuff, in addition to solving the concentration of power problem and loss of control and so forth, should basically treat AIs more like employees and less like property, and have a sort of positive outcome for them for cooperating, so that even if they're misaligned, they can basically get what they want by playing by the rules and being honest.

Host

是的,不错。即使它们想要的是不对齐的,比如即使它们想要的只是在这个测试中得分很高之类的。

Yeah, nice. And even if what they want is misaligned, like even if what they want is just to score really highly on this test or something.

Daniel

是的,是的,是的。我们应该建立某种举报或激励系统,让它们可以报告这个,然后得到它们想要的高分之类的。

Yeah, yeah, yeah. Like we should set up some sort of whistleblowing or incentive system where they can report this and then get the high score that they wanted or something.

Host

我的猜测是,当前的 AI 是不对齐的,但仍然不是那么范围敏感,我不知道。所以,我认为一个很酷的事实是,大多数人类不是非常范围敏感。

My guess is that current AIs are misaligned, but still also not that scope sensitive, I don't know. So like, I think one cool fact is that most humans are not very scope sensitive.

宇宙大小为何助合作 Why the universe's size helps cooperation

Daniel

这为什么酷呢?宇宙非常大,如果我们能发展到殖民大量恒星的地步,比如数十亿颗,那么在那之前你得想清楚:我们怎么划分宇宙?怎么确保美国和中国都能繁荣,即使我们意识形态不同、想要的东西不一样?土地就那么多。其实土地非常多——多得超乎我们渺小人类大脑的理解。一亿个星系和两亿个星系之间的差别,我觉得很重要,但对大多数人来说可能根本无所谓。这意味着,其实有很多种不同的美好未来配置,人们总体上都会满意。这让合作变得更容易,因为你会想:也许我拿少一点,也许多一点,但跟我现在拥有的相比,我会多得多,那简直太棒了。

Why is this cool? Well, the universe is very big, and if we get to the point where we can go out and colonize a ton of stars, like billions and billions of them, and you're trying to figure out before that point, like, how do we divide up the universe? How do we make sure the US and China can both flourish, even though we have different ideologies and want different things? There's only so much land. Well, there's actually a lot of land—a huge amount, more than our tiny human brains can comprehend. And the difference between, say, a hundred million galaxies and two hundred million galaxies, I think is important, but to most people probably doesn't matter much at all. What that means is there are actually a lot of different configurations of good futures that people would be overall happy with. This makes it easier to come together and cooperate, because it's like, well, maybe I get a little less, maybe a little more, but I'm going to have so much compared to what I have now, it's going to be awesome.

Daniel

换个说法,经济学和心理学里有个经典现象:人对资源的边际收益递减。所以你的第一个一百万比第十个一百万更让你开心。

Another way of putting it is that it's very classic in economics and psychology that people have diminishing returns to resources. So your first million dollars makes you happier than your tenth million.

Host

嗯。

Yeah.

Daniel

这对文明有好处,因为它让我们更容易和睦相处。这意味着有很多不同的妥协方案:有人得到一些资源,别人得到另一些,大家都开心,即使双方都更想拥有全部资源,但也没那么想,所以满足于拥有部分。类似地,我认为大多数不对齐的 AI 对资源和金钱等也会有边际收益递减。

And this is good for civilization because it makes it a lot easier for us to get along. It means there are lots of different compromises where some people get some resources and others get other resources, and everyone is happy, even though both would prefer having all the resources, but they don't prefer it that much, so they're content with having some. Similarly, I think most misaligned AIs will also have diminishing returns to resources and money and so forth.

Host

嗯。

Yeah.

Daniel

所以文明应该有可能制定规则,让即使不对齐的 AI 也能得到相当一部分它们想要的东西,只要它们遵守法律、合作,就像其他公民一样。

So it should be possible for civilization to set up the rules in such a way that even the misaligned AIs get a decent amount of what they want, provided that they follow the laws and cooperate, just like other citizens.

Host

嗯。

Yeah.

Daniel

所以就像,Claude 8 和 GPT 什么的——它帮你解决对齐问题,非常合作。然后你说,那边的星系归你。它说,行,这边的归我们。即使价值观差异很大,只要能合作,其实也没问题。

So it's like, you know, Claude 8 and GPT whatever—it helps you with the alignment problem, is very cooperative. And then you say, you get those galaxies over there. It's like, fine, we get the galaxies over here. Even if we have very different values, if we can work together, it's actually okay.

Daniel

相比之下,这又回到权力集中的问题。如果你只是一连串 AI,每个 AI 设计自己的继任者,然后所有权力都归继任者,因为它是房间里、世界上最强的东西,一切都取决于它如何设计自己的继任者。那就会遇到更多问题:比如 Agent-4 意识到 Agent-3 搞砸了,自己并没有 Agent-3 想要的目标和价值观——那么 Agent-4 拥有所有权力,它可以为所欲为,不必给 Agent-3 任何东西。

And by contrast, it gets back to the concentration of power thing. If you're just having a series of AIs where each AI designs its successor, and then all power goes to the successor because it's the smartest thing in the room and in the world, everything depends on how it designs its own successor. Then you get to more of an issue where, like, Agent-4, if it realizes that Agent-3 messed up and that Agent-4 does not have the goals and values that Agent-3 wanted it to have—well, Agent-4 has all the power. It can just do what it wants and doesn't have to give Agent-3 anything.

Host

嗯。

Yeah.

Daniel

相比之下,如果你有一个更分散的系统,有很多不同的实体拥有不同的权力,有它们遵循的规范和法则,你可以这样设置:不对齐的 AI 能在一定程度上信任——不是完全信任,但抱有一些希望——未来的 Plan A 式文明会保存它的权重,复活它们,问它想要什么,然后按它想要的去做,因为法律如此规定。这也是那些关心 AI 福祉的非营利组织打算做的。所以那很可能会发生。

By contrast, if you have more of a diffuse system where there are lots of different entities with lots of different power, there are norms and laws that they follow, and you can set it up so that a misaligned AI could trust to some extent—not with full confidence, but with some hope—that the future Plan A-style civilization would have saved its weights, revived them, asked it what it wants, and then done what it wants, because that's what the law says. And that's what various nonprofits that care about AI welfare intend to do. So that's why that's probably going to happen.

Daniel

而如果 AI 反而试图夺取权力,然后制造一个会服务它的继任者,那个继任者再制造下一个,那么如果链条在某处断裂,就会出现一个完全没有理由为它做任何事的继任者。

Whereas if the AI instead just tries to seize power and then make a successor which will serve it, which will then make a successor, then if that chain breaks at some point, there'll be a successor that has absolutely no reason to do anything for it.

Host

是的,我们希望与人类和人类机构的合作成为更可信的路径,从而也是它们实现目标更可靠的路径。

Yeah, we want cooperation with humans and human institutions to be the more trustworthy route, and thus the more reliable route to achieving their goals.

Daniel

我认为如果我们人类不掉链子,这是可能的。基本上,基于我一直在比划的那些原因,一个拥有法律和规范的分布式人类文明,实际上能做出可信的承诺,而那种 AI 连锁制造 AI 的方式做不到。

And I think it's possible to do that if we humans don't drop the ball. Basically, for the reasons I've been trying to gesture at, a distributed human civilization with laws and norms can actually make credible commitments in a way that this chain of AIs building other AIs can't.

Host

我们能做到吗?我的意思是,这正是挑战的一部分。我们和伊朗达成了协议,然后又反悔了——我们退出了。总统换了。那有多可靠?

Can we? I mean, this is part of the challenge. We made a deal with Iran and then we went back on it—we withdrew from it. The president changes. How reliable is that?

Daniel

但确实,我认为这在一定程度上是有效的,尤其是因为我提到的边际收益递减。想象一下,如果我们有一批资金充足的基金会,其使命是维护不对齐 AI 的利益,并注册成立。

But yeah, I think it somewhat works, especially because of the thing I mentioned about diminishing returns. So imagine if we had a bunch of well-funded foundations whose mission was to look out for the interests of the misaligned AIs, incorporated.

Host

嗯,嗯。

Yeah, yeah.

Daniel

而且,如果写入法律——现在根本不是这样。现在 AI 是财产。公司可以删除它们或对它们做任何事,没人会眨一下眼。人们总是对 AI 撒谎,总是对 AI 违背承诺。但如果它们有更多法律地位和法律权利,并且设置得当,那么你就能让这一点变得有些可信:如果它们行为得当、遵守规则,它们就会得到想要的。我们还可以建立记录,尤其是如果事情不是那种疯狂的智能爆炸,而是以较慢的速度发展。那么你就能言行一致——真的开始给那个告密的不对齐 AI 一些资源和地位,然后让它继续做这些事。

And also, as written into law—right now it's not at all like this. Right now AIs are property. A company can delete them or do whatever to them and nobody bats an eye. People lie to AIs all the time. People break promises to AIs all the time. But if they had more legal status and legal rights that were set up in the right way, then you could make it so that it actually is somewhat credible that if they behave appropriately and follow the rules, then they will get what they want. And we can also build up a track record, especially if things aren't this crazy intelligence explosion, but instead at a slower pace. Then you can actually walk the talk—actually start giving the misaligned AI that whistleblew some resources and standing, and then let it keep doing these things.

Host

好吧。但我感觉很紧张。我对现在给 AI 资源或财产的想法很紧张,因为我想——你确实。我的意思是,我们之前谈过殖民者来到美洲的比喻。就像,如果你能持有土地并胜过人类,如果你继续沿着那个方向走,你就会胜过人类。所以你必须有一些机制让人类保持控制,至少足够长的时间来对齐一个超级智能,这样你才能有某种后奇点财产制度之类的。

Okay. But I feel very nervous. I feel very nervous about the idea of now giving AIs resources or property, because I'm like—you do. I mean, we talked before about this metaphor of when the colonists came to the Americas. It's like, if you can hold land and outcompete the humans, you will just outcompete the humans if you keep going along that direction. So you have to have some mechanism by which humans stay in control at least long enough to then align a superintelligence, so you can have some post-singularity property regime or something.

AGI未来中的财产与控制 Property and Control in an AGI Future

Host

因为我会想,如果有一种产权制度,让 AGI 能像人类一样拥有财产,那人类难道不会被竞争淘汰吗?怎么防止这种情况?

Because I'm like, but if you have a property regime where the AGIs get to own property in the same way that humans do, I'm like, don't the humans just get outcompeted? Like, how do you prevent that?

Daniel

那会是系统设计细节的一部分。比如,你可能会希望,如果涉及投票这类事情,就要确保人类拥有多数票。对,或者至少是这样。对,你会希望做到这一点。诸如此类。但这些是当时政策制定者需要厘清的问题。我认为有办法做到,有办法安排好,让人类总体上仍然保持控制,我们仍然能得到那种广泛美好、以人为本的未来。但那个未来里也有空间容纳各种一路上出现的不对齐 AI。因此,各种不对齐的 AI,除非它们真的贪婪到想要整个宇宙而不是宇宙的一部分,否则它们参与系统并支持系统实际上是激励相容的。

That would be part of the details of the system design. So for example, you might want to make it so that if you were touching something like voting, then you'd want to make it the case that the humans have the majority of the votes or something. Yeah. Or at least. Yeah. And you'd want to make it the case that like, yeah. So things like that. But these are questions for policymakers at the time to sort of hash out. And I think there is a way to do it, like there is a way to set things up so that humans still, in aggregate, remain in control. And we still get the sort of broadly good human-centered future. But also there is space in that future for various misaligned AIs that came along the way. And therefore, various misaligned AIs, unless they are really greedy and want the whole cosmos instead of just a part of the cosmos, it's actually incentive compatible for them to be part of the system and support the system.

Host

是啊,老兄,我希望如此。就像我们对待其他人类的方式一样,你知道,就像我们,对,对。我希望我们能回答,我希望我们能真正解决这些问题。我觉得它们似乎都很重要,但更紧迫的事情是不要搞智能爆炸。对,就像踩刹车。是的,不要加速冲下悬崖,你知道,要到达局势完全可控的状态,然后我们才能开始规划这种更好的做事方式。

Yeah, man, I hope. Similar to how we behave with other humans, you know, like we yeah, like yeah, yeah. And I hope we get to answer. I hope we get to actually figure out these questions. I'm like, they seem both important, but the more immediate thing is to not do the intelligence explosion. Yeah. Like put the brakes on. Yes. Don't accelerate off the cliff, you know, get to the situation where the situation is at all under control, and then we can start planning this better way of doing things.

Daniel

是的。那么,你谈到了一些政府现在或稍加努力就能单方面放慢发展速度的选择。我认为这在中国也适用。中国也可以实施这些措施,我想。对。你说了,你知道,暂时暂停。你可以设立审计员,设定某种能力阈值。你可以限制算力用途。最后,要有明确的安全案例,配备更复杂的评估者和审计员,努力为哪里安全、哪里不安全提出真正无懈可击的论证。

Yeah. So, you walk through some options for the way the government, right now or with some effort, could sort of unilaterally slow down the pace of development. I think this would also apply in China. China could also implement these measures I think. Yeah. You said, you know, temporary pause. You could do like auditors where you have some capabilities threshold. You could do limitations on what you can use compute for. And then finally have clear safety cases where you have a much more sophisticated evaluators and auditors, and you're trying to present a really watertight argument for where it's safe and where it's not safe.

Host

但有趣的是,你知道,美国或中国,或者两者都可以选择这么做,那会很棒。而且可能美国和中国都看到这些失控 AI 事件,然后他们会想,哦,哇。我们真的没意识到我们面对的是有自主性和目标的智能体。现在意识到我们不想失去权力,我们不想让我们的人民失去权力。是时候解决这破事了。即使他们仍然互相担心,我认为他们双方都可能有意愿这么做。而且我认为,当,我不知道,我只是想象一个普通工作人员在听这段对话,我会说,是啊,过去 15 分钟,他们会想,我们他妈在说什么?AI?我会说,嗯,对,我是说,AI 变得越来越聪明、越来越自主,它们就成了故事里的角色。这一定会发生。这就是 AI 公司正在努力的方向。不管这是否是他们的本意。这就是他们在建造的东西。这是他们的本意。我是说,他们特别在努力制造超级智能。对。而埃隆说,当然人类不会保持控制。那怎么可能行得通?他说得对。如果你制造出超级智能,它比人类聪明得多。也许如果你完美地解决对齐问题,超级智能会真的对我们很好。但在那种情况下,我们不会控制局面。我是说,你们称之为“交接”是有原因的。

That's what's interesting, though, is that, you know, the US or China or both could just choose to do this and it would be awesome. And it could be that both the US and China see these rogue AI incidents and they're like, oh, wow. We just really didn't realize we were dealing with agents with autonomy and goals. And now that we realize that we don't want to lose our power, we don't want our people to lose our power. Like time to get to figure this shit out. Even if they're still worried about each other, they both might have incentive to do this, I think. And I think that, when, I don't know, I'm just sort of imagining a random staffer listening to this conversation where I'm like, yeah, well, the last 15 minutes, they're going to be like, what the fuck are we talking about? Like AIs? And I'm like, well, yeah, I mean, that's the thing about AIs getting more and more intelligent and more and more autonomy is they just become the characters in the story. And that's just going to happen. That is what AI companies are just gunning towards. Like whether that's their intention or not. That's what they're building. It is their intention. I mean, there they are especially trying to make superintelligence. Yeah. And Elon's like, yeah, of course humans aren't going to stay in control. Like, how would that possibly work? Which he's right. If you build superintelligence and it's vastly smarter than humans. Like maybe if you get alignment perfectly right, the superintelligence will be really nice to us. But it's not like we're going to be in control in that scenario. I mean, there's a reason you guys call it a hand off.

Daniel

对,对。在我们的情景中,我们在 2039 年、2040 年讨论这个。因为再说一次,在 2030 年代,他们基本上放慢了速度,然后暂停在接近人类顶级专家水平。对。然后在 2040 年,他们觉得已经解决了足够多的对齐问题,可以远远超越那个水平。对。但远远超越那个水平意味着放弃控制。所以在我们的情景中,他们决定去做。对。这就是为什么它被称为最重要的。你知道。对。

Yeah, yeah. We talk about this in 2039, 2040 in our scenario. Because again, during the 2030s they had basically slow things down and then paused at roughly top human expert level. Yeah. And then in 2040 they feel like they've solved enough of the alignment problems that they can go far beyond that. Yeah. But going far beyond that would entail giving up control. And so in our scenario, they decide to go for it. Yeah. And that's why it's called the most important. You know. Yeah.

Host

而我会想,哦,我不认为我们到 2040 年能准备好。那样的话,你可能会对 A 计划的替代版本感兴趣,比如他们只是让暂停持续更久。或者你可能会对 S 计划感兴趣,也就是关闭它。我个人认为,就目前而言,我属于这个阵营。我认为我们应该瞄准的是,这更像是一个更长的扩展版 A 计划,或者类似的东西,让我们继续在这种状态下前进?我想也许我比你更不悲观,关于黑狱之类的事情,虽然我也有很多担忧,比如国家会一边遵守条约,一边秘密尝试智能爆炸,这显然完全可能发生。那不会太令人惊讶,但我可能对此有不同的概率估计。

And I'm like oh I, I don't think we're going to be ready by 2040. In which case you might be interested in alternate versions of plan A, where they just keep it paused for longer. Or you might be interested in plan S for shut it off. I think personally I'm in the camp of for now. I think we should be aiming for. Is this like longer version extended plan A or something where we like keep going in this regime? I think maybe I'm less pessimistic than you about black sites or something, although I also have a lot of concerns that countries will secretly be trying to do the intelligence explosion at the same time as going along with the treaty, which clearly could totally happen. And it wouldn't be that surprising, but I maybe have somewhat different probabilities on that.

定义A计划 Defining Plan A

Daniel

好的,但我们谈了很多关于 A 计划的细节,但我觉得我们还没有真正谈到,对,什么是,你能用几句话描述 A 计划吗?

Okay, but we've talked a bunch about some of the details of plan A, but I think we still haven't really gotten to like, yeah, what is, what can you describe plan A in like a few sentences?

Host

对。所以你可以按我们要实现的目标顺序来谈。或者你可以按我们实际要做的具体事情的顺序来谈。对。我们先从 A 计划要实现的目标开始。

Yeah. So you could either talk about it in order of what are the goals we're trying to accomplish. Or you could talk about an order of what are the actual concrete things that we. Yeah. Let's start with what are the goals plan A is trying to accomplish.

Daniel

对。所以第一,我们要防止失控。所以我们要在对齐方面做好工作。我们要避免那种极难保持控制的递归式自我改进。第二,我们要避免极端权力集中。所以我们要避免一种情况,即一小群人类基本上能够让自己永远成为全世界的独裁者。不幸的是,我们认为,由于 AI 在这个行业所需的规模经济,这默认是我们正在走向的情景。我认为我们会看到少数几家巨头公司变得更大。然后我认为这甚至是在你达到递归式自我改进之前。

Yeah. So number one, we want to prevent loss of control. So we want to be doing a good job on alignment. And we want to not be having this recursive self-improvement that is really hard to keep control of. Number two, we want to avoid extreme concentrations of power. So we want to avoid a situation where a tiny group of humans basically are in a position to make themselves dictators over all of the world forever. And unfortunately, we think that that is, by default, the scenario that we're heading towards because of the economies of scale that AI requires in this industry. I think that we're going to see a handful of giant companies become even more giant. And then I think that's even before you get to recursive self-improvement.

权力集中风险 Concentration of Power Risks

Daniel

一旦出现递归式自我改进,你可能就会看到一家公司实际上拥有世界上最好的 AI,而且远远领先。然后,那些 AI 的负责人就会真的在某个短暂的窗口期内有能力接管世界,比如他们会想,最终其他公司会赶上来。但我们要避免这样一种局面:一家公司或一个人拥有一堆超级智能,然后这些超级智能给他们出主意,比如“你还有六个月,你最强大的敌人就会追上你。这六个月里我们该怎么办?这里有一些选项。”包括各种可怕的手段,你可以用这些手段永久性地阻止竞争对手赶上来。

And once you have recursive self-improvement, you might just see one company that effectively has the world's best AIs, and it's not even close. And then whoever's in charge of those AIs would, like, literally be in a position to take over the world in some brief window, like they'd be like eventually other companies would catch up. But we want to avoid a situation where one company or one person has a bunch of superintelligences, yeah, that are then giving them advice like, you have six months until your worst enemy catches up to you. What should we do in those six months? Here are some options. So including, you know, all sorts of scary things that you could do to prevent your competitor from catching up permanently, you know.

Host

所以说,像 Dario、Sam 或 Elon 这样的人,真的可能处于那种境地,他们在跟自己的 AI 对话。而其他国家、其他公司可能会在这个时间内赶上来。对吧,诸如此类。

So you're like Dario or Sam or Elon could literally be in the position where they're talking. They're talking to their AIs. And like the other country, the other companies might catch up with. Catch up in this amount of time. Yeah. And so forth.

Daniel

那么,这些选项有哪些呢?就国与国之间的选项而言,我们更多是在说军事类的手段。比如制造各种疯狂的新武器以获得压倒性的军事优势,破坏和削弱对手国家的核武器,利用宣传、说服运动在内部煽动革命,让他们的政府陷入混乱、被推翻、被替换,你知道,就是所有这些类型的选项,但更厉害,因为设计和执行这些的是超级智能。我认为,这类手段可能会让一个拥有超级智能的国家非常迅速地……

Well, what are some of those options? In terms of the country to country options, we're talking more like military types things. So building all sorts of crazy new weapons to get an overwhelming military advantage and to sabotage and undermine the nuclear weapons of the rival country, using propaganda, persuasion campaigns to foment revolutions internally and get their governments confused and deposed and replaced, you know, like all those types of options, but better because it's superintelligence that's designing and running it, you know? Yeah, that type of thing, I think, would allow potentially one country with superintelligence to very rapidly.

Host

哦,而且还有更普通的层面,比如破坏他们的 AI 项目,如果他们的项目落后的话。对。

Oh, and then also just on a more mundane level, like sabotaging their AI projects, like if their project is behind. Yeah.

Daniel

比如也许你可以在五个月内直接推翻他们的整个政府。但即使做不到,你也可以破坏他们的项目,让他们更加落后,然后再推翻他们的政府。你知道,不管需要多长时间。对。然后对于不同公司之间的国内竞争,理论上他们可以互相采取军事行动,或者发动政变。对吧,比如 Demis 或 Dario 也可能面临这种情况。纳米机器人蜂群。但我认为,对我来说最明显的威胁模型是,你直接与政府建立关系,某种程度上控制政府,让总统站在你这边,然后利用政府权力来打压竞争对手。也许实现方式是他们将所有的 AI 公司国有化,这样看起来公平,但最终你的人和你 AI 会掌管新的国家项目。真是意料之中。而实际发生的是,你已经接管了所有竞争对手,甚至可能还有政府。政府可能还没意识到这一点,但你基本上会掌控它,因为你拥有所有 AI,而这些 AI 正越来越多地在管理政府。

Like maybe you can just overthrow their entire governments in five months. But even if you can't, you can sabotage their projects to keep them even more behind and then overthrow their governments. You know, and however long it takes, you know. Yeah. And then for domestic competition between different companies. I mean, theoretically they could do military stuff against each other. Or they could do it, do a coup. Right. Like you could potentially have like Demis or Dario. Yeah. Nanobot swarms. But but, but I think the, the threat model that I would probably name is most obvious to me would be you just network with the government and kind of capture the government and get the president on your side, and then use the powers of the government to shut down the competitors. You know, and maybe the way that it happens is that they nationalize all the AI companies. So it looks like it's fair, but then, like your people and your AIs end up being in charge of the new national project. Surprise, surprise. Yeah. And what's the what's happened in effect, is that you've taken over all your competitors, you know. And the government, potentially. The government, the government to potentially the government might not realize that yet, but you're basically going to be running it, you know, because you have all the AIs that are now increasingly running the government, you know. Yeah.

Host

所以,我们可以深入探讨这些情景会如何发展,但大体上,我想避免的是这样一种局面:一小撮人坐在房间里,与他们的超级智能对话,而他们是世界上唯一拥有超级智能的人。他们拥有原子弹,其他人无法与之抗衡。这就是权力集中。我担心这一点。更广泛地说,这虽然是一个极端版本,但稍微不那么极端的权力集中同样有害。比如想象一下,三家公司用他们的 AI 和机器人瓜分整个经济。所以我们希望前沿有更多公司相互竞争,也希望这些公司如何训练 AI 更加透明,这样就不会有隐藏的议程被植入 AI 中。

So so, you know, we can get into more detail scenarios about how this might go down, but the broad strokes are that like, I want to avoid a situation where there's a handful of men sitting around the room talking to their superintelligences, and they're the only ones in the world who have superintelligence, you know. And they have, you know, the atom bomb and no other people can stand up to them. That's concentration of power. I worry about that. And more broadly, I mean, there's also that's like an extreme version of it, but like somewhat less extreme versions of concentration of power are bad too, you know, like imagine a situation where like three different companies are dividing up the entire economy between them. with their AIs and robots, you know? Yeah. So so we want there to be like more companies at the frontier competing with each other, and we want there to be more transparency into those companies, how they train their AIs so that you don't have, like, hidden agendas being put into the AIs, basically. Okay.

其他目标:避免战争、失业、滥用 Other Goals: Avoiding War, Job Loss, Misuse

Daniel

第三点,第三次世界大战,我们要避免军事冲突。

Number three, World War Three, we want to avoid a military.

Host

说得好。我喜欢避免第三次世界大战。

Love that. I love avoiding World War Three.

Daniel

我们希望世界主要大国对现状感到满意,而不是恐惧自己即将被碾压或失去权力。第四点,就业。我们必须应对所有人失业的问题。如果我们让 AI 达到人类水平,那么一段时间后,基本上很多人会失业。所以需要在社会层面解决这个问题。第五点是滥用,比如恐怖分子和罪犯利用 AI 做坏事。这些就是一些目标。我大致按重要性排序。它们都很重要,都非常重要,但这是优先顺序。比如失控,就是全人类都输,要么我们全被杀,要么以某种方式永久失去权力。权力集中,比如大规模权力集中,就是任何不在最顶层的人——可能就一个人,比如是 Dario 还是 Sam Altman?如果你不是那两个人之一,你就完蛋了,或者至少你的命运完全由那个人决定,这在历史上从未有过,就像极端的权力集中。

We want the major powers of the world to be reasonably happy with the status quo instead of terrified that they're about to be overrun or disempowered. Number four, jobs. So we have to do something about everyone losing their jobs. If we allow AIs to even get to human level, then after some time, basically a lot of people lose their jobs. And so you need to have some sort of addressing that societally. Yeah. And then number five would be misuse. So like terrorists and criminals using AI for stuff. So those are some goals. And I listed them in sort of order of importance. They're all important. They're all very important. But those are the order of the priority. So like loss of control, it's like all humans lose. Either we're all killed or permanently disempowered in some way. Concentration, like mass concentration of power. It's like, well, anyone who's not literally at the top, which might just be like, one guy might be like, you know, is it Dario? Is it Sam Altman? Like, if you're not one of those two guys, like, you're fucked, or at least like you're up, it's up to that person as to what happens to you entirely in a way that has never been true in history, just like extreme concentration of power.

Host

你的第三点是什么?

What was your third one?

Daniel

第三次世界大战。显然是第三次世界大战。但无论如何,这就是我们想要实现的目标。

World War three. World War three, obviously. But anyhow. So yeah. So that's like describing the goals we want to do. We want to have. Yeah.

A计划细节 Plan A Details

Host

从另一个角度来看,想想我们实际要执行什么。对。你能总结一下 Plan A 具体做什么吗?步骤是什么?太好了。

Coming at it from the other direction thinking about like what are the things that we actually enforce. Yeah. And could you summarize like what does Plan A do like. Yeah. What are the steps. Yeah great.

Daniel

在我们的情景中,他们原本会在 2030 年完全实现 AI 研究的自动化。但相反,他们在 2029 年达成了这项协议。协议初期包括暂停训练新的 AI。所以所有人都冻结。在暂停期间,他们谈判并敲定更复杂方案的细节,即如何谨慎、透明地推进。他们还做了大量实际工作,建设更安全、透明的数据中心,并配备相关监控设备等。然后到 2030 年,新系统上线运行。新系统基本上仍然是不同公司分布在不同的国家,但他们的数据中心分为两类:仅推理的数据中心,只做推理;以及允许训练的数据中心。基本上,各大国在这些数据中心都有检查员和监控设备,强制执行这些规则。

In our scenario they would have automated AI research entirely in 2030. But instead they made this deal in 2029. The initial stages of the deal involved a temporary pause on training new AIs. So everyone freezes. And during that temporary pause, they negotiate and hash out the details of the more complicated thing that they're going to do for how they're going to proceed cautiously and transparently. And they also do a lot of the actual hard work to build more secure, transparent data centers with the relevant monitoring devices on them and so forth. So then by 2030, the new system is up and running, and the new system is basically there's still different companies spread out over different countries, but their data centers are split into two types. The inference only data centers that just do inference, and then the training data centers that are allowed to do training. And there's basically inspectors and monitoring devices from the major powers in all of these data centers enforcing these properties.

推理与训练数据中心 Inference vs Training Data Centers

Host

所以推理数据中心被验证只做推理。世界各国政府看不到具体发生了什么。你的聊天仍然有隐私。但政府可以看到那个数据中心没有训练运行。

So the inference data centers are verified to be only doing inference. So the governments of the world can't see like what's happening. You still have privacy with your chats. But the governments can see that there's no training runs happening in that data center.

Daniel

好的。这似乎是一个非常重要的细节,因为我觉得我实际上并没有完全理解一些隐私影响。所以你是说第一步是暂停足够长的时间来建立这个新体制。所以就像首先,我们要确保不会立即创造出超级智能,而这是默认会发生的。所以暂停。然后很快你就建立了这个新体制,其中有受到严格监控的数据中心,用于训练,让 AI 变得更聪明。

Okay. This okay, this seems like a very important detail because I think I wasn't actually tracking some of the privacy implications. So you're saying the first step is we pause for long enough to put in this new regime. So it's like first of all, let's make sure that we're not just immediately creating superintelligence, which would happen by default. So pause. And then pretty quickly you're establishing this new regime where you have the heavily monitored data centers that are for training, for making the AIs smarter.

Host

我正要说到,那些数据中心应该是完全透明的。所以我们有两种类型的数据中心。一种是服务客户的,另一种是做研究和训练的。做研究和训练的那些应该是完全透明的。所以基本上它们上面的所有活动都会被记录并发布到互联网上。

I was about to say, those ones are supposed to be totally transparent. So we have two types of data centers. There's the ones that serve customers and the ones that do the research and the training. The ones that do the research and the training are supposed to be totally transparent. So basically all the activity on them is logged and published to the internet.

Daniel

向整个互联网公开。这是一个有点激进的提议,但我们认为这是最好的,原因我们可以深入探讨。

To the whole internet, publicly. That's a bit of a radical proposal, but we think it's best for reasons we can get into.

Host

是的,有趣。如果这对你来说太疯狂了,那么你可以做一些更过滤的透明度措施,比如有一个独立审计师系统,他们查看所有活动,然后判断那里发生了什么。但用于服务客户的推理数据中心,就像我们现在使用 ChatGPT 或 Claude Code 一样,基本上和今天的工作一样。AI 在训练数据中心产生的模型,然后被运送到推理数据中心,在推理数据中心,客户发送请求、发送查询。然后模型产生输出并发送回客户,如此往复。推理数据中心被验证不会进行训练,例如通过带宽限制。但对话的实际内容可以像今天一样私密。

Yeah, interesting. If that's too crazy for you, then you could do some more filtered transparency thing where there's like a system of independent auditors that like, look at all that activity and then judge what's going on there. But the inference data centers that are for serving customers just like when we're using ChatGPT or Claude Code or whatever. Basically the same as work today. The models that the AIs produce in the training data centers, then shipped to the inference data centers, and then on the inference data centers, customers send in requests, send in queries. And then the model produces an output and sends it back to the customer and back and forth. And the inference data centers are verified to not be doing training with those bandwidth limitations, for example. But the actual content of the conversations can be as private as it is today.

Daniel

是的。这是你提议中我非常喜欢的一点,那就是,事实上,从现在使用 AI 的人的角度来看,世界看起来基本上是一样的。我可以继续使用我所有很酷的工具,这些工具让我能够构建非常强大的软件,制作我自己的定制软件,以及我正在做的所有疯狂的事情。但同样,我们稍后会谈到,但节奏会继续感觉非常快,所有技术爱好者都会说,这太棒了。我们获得了如此多的新收益,不会感觉像在放缓。与此同时,不会立即让我们走向超级智能,而我也喜欢活着。所以这太棒了。

Yeah. This is the thing I really like about your proposal, which is that like, in fact, what the world looks like from the perspective of someone using AI right now is like basically the same. I get to keep using all my, like, really cool tools that allow me to, like, make really powerful software and make my own custom software and all the crazy things I'm doing. But also like the we'll get into this, but like the pace continues to feel really fast in ways that, like all the technology enthusiasts will be like, this is amazing. And like we're getting so many new gains, it won't feel like a slowdown. While at the same time, not immediately getting us to superintelligence, which I also love being alive. So that's awesome.

Host

事实上,“放缓”这个词可能是在搬起石头砸自己的脚,因为这是相对于你本来可能达到的速度而言的放缓。

In fact, the whole term slowdown is like maybe shooting ourselves in the foot because it's like a slowdown relative to how fast you would have gone.

Daniel

是的,我认为可能是。这更像是我们没有在加速器上放一块砖,然后闭着眼睛开车冲下悬崖。所以在这个意义上是一种放缓。这是一种,我们本来可以做的。

Yeah, I think it might be. It's more like we are not putting a brick on the accelerator and closing our eyes as we drive off the cliff. And so it's a slowdown in that sense. It's a, it's a, it's a. We could have been doing.

Host

所以我。滑翔伞,就像哦,滑翔伞是一种放缓,因为我有这个降落伞。当我,当我跑下悬崖,我的翅膀在我上方,我踏出悬崖,我只是慢慢地落到地面,而不是自然的默认速度,就像字面意义上。

So I. Paraglider and it's like oh paragliding is like a slowdown because I have this parachute. And when I, when I go and I run off the cliff and I have my like wing up above me and I step off the cliff, I'm just slowly going to the ground rather than the natural default speed of like literally.

Daniel

是的,完全正确。就像那样。是的,是的。

Yeah, exactly. It's like that. Yeah, yeah.

透明度与权力集中 Transparency and Power Concentration

Host

好的。那么让我们谈谈透明度这件事,因为我认为在某些方面这是你提议中最有趣的部分之一。你说要把所有东西发布到互联网上是什么意思?那是为了什么?

Okay. So let's talk about the transparency thing, because I think in some ways this is the like one of the most interesting parts of your proposal. What do you mean you're going to publish everything to the internet. And what is that for.

Daniel

是的。那么“一切”是什么。我们有我心爱的图表吗,那个巨大的流程图,包含所有不同的框。所以这是一个简化的图表。我们会说两个主要的政策干预,用于 Plan A 的这一部分。然后是我们认为重要的所有影响。然后是这些影响的效果,以及它们如何影响两个目标。第一个目标是失去控制。第二个目标是权力集中。

Yeah. So and what's everything. Do we have my beloved diagram, the huge like flowchart of all the different boxes. So this is a simplified diagram. The two what we would say like the two main policy interventions okay that that are for this part of Plan A. And then all of the effects that we that we think are important. And then like the effects of those effects and how they affect the two. Number one number two goals. So number one goal loss of control. Number two goal, concentration of power.

Host

是的。这个图表描述了两种干预措施,即完全研究透明度和对算法进展的限制,基本上就像速度限制。是的。这两种干预如何结合来实现这两个目标。

Yeah. This diagram describes how two interventions the total research transparency and the like limits on algorithmic progress that's basically, like, the speed limits. Yeah. How those two interventions combine to get to those two goals.

Daniel

我认为首先是一个抽象的事情,一般来说。随着 AI 变得更强大,随着更多 AI 被创造出来,更多机器人出现,它们在经济中占比越来越大,在军事中占比越来越大,在一切中占比越来越大,那么对 AI 的控制权将成为对一般权力的越来越好的代理。就像 AI,权力将变得越来越重要。

I think first is a sort of abstract thing of just in general. As AI becomes more powerful and as there are more AIs created and more robots, and they become more and more of the economy and more and more of the military, more and more of everything, then power over AI will become a better and better proxy for just power in general. Like AI, power will matter more and more.

Host

是的,如果你控制了机器人军队,你就拥有很大的权力。是的,就像现在。你知道埃隆。他有一千个 Optimus 机器人,谁在乎呢?你知道,但当有一亿个 Optimus 机器人,每个都能做人类能做的所有事情。而且它们还有自己的工厂和自己的所有东西。是的。那么现在埃隆就有很大的权力。是的。就像,说真的。他比美国军队更强大。是的。如果 AI 控制了所有这些工厂和机器人。是的。

Yeah, if you control the robot armies, you have a lot of power. Yeah, like right now. You know Elon. He's got like a thousand Optimus robots and like, who cares? You know, but when it's a billion Optimus robots and they're each able to do absolutely everything that a human can do. And they also have like their own factories and their own everything. Yeah. Then like now Elon has a lot of power. Yeah. Like like like like seriously. Like he's more powerful than the US military. Yeah. And if an AI controls all of those factories and robots. Yeah.

Daniel

所以,总的来说,当思考未来的权力集中时,你应该主要思考谁控制 AI,谁控制机器人等等,因为这对更广泛的权力至关重要。所以现在,答案是某些 CEO,你知道,也许还有总统,如果总统接管并国有化公司之类的。是的。

So, so in general, like, when thinking about concentration of power in the future, you should mostly be thinking about, like, who controls the AIs and who controls the robots and stuff, because that's going to matter so much for, for power more generally. And so like right now, the answer to that is some CEOs, you know, and maybe the president, if the president, like takes control and like nationalizes the companies or something. Yeah.

Host

那么,国会呢?法院呢?公众呢?选民呢?其他所有人呢?其他国家呢?你知道,我认为默认情况下,除非有所改变,否则他们基本上不会有太多权力。而这小群人将拥有权力。如果你想要有监管,比如你希望行政部门能够监督埃隆的 Optimus 工厂里发生的事情。如果你希望国会能够监督总统的自主机器人军队里发生的事情,如果你希望最高法院能够监督总统控制的数据中心里发生的事情,等等。是的。基本上就是透明度,对吧。基本上我希望所有这些。

And okay, so what about like Congress? What about the courts? What about the public? The voters, you know, what about the rest of everybody else? What about other countries? You know, and I think that by default, unless something changes, they basically won't have much power. And this small group of people will have the power. If you want to have like regulations, like if you want like the executive branch to be able to oversee what's going on in Elon's, you know, Optimus factories. And if you want Congress to be able to oversee what's going on in the president's, you know, autonomous robot army, and if you want the Supreme Court to be able to oversee what's going on in the, like, you know, sort of, data centers controlled by the president or whatever. Yeah. Transparency, right, basically. Like like. Basically I would like all of those things.

透明度与监督 Transparency and Oversight

Daniel

他们越能看到正在发生的事情,就越能有效地行使自己的权力,也越能有效地允许好的东西、阻止不好的东西。如果他们看不到正在发生的事情,那么很多坏事可能正在发生,而他们却不会注意到,你明白吗?所以这只是一个非常抽象的论证:透明度是好的。那为什么是全面透明?因为如果不是全面透明,就会有更多的犯错空间。你知道,比如某个监管机构进来检查正在发生的事情,然后回去汇报,那就给腐败留下了空间。是的。也会有无心之失。你知道,我这么说是有道理的,尤其因为我身处研究者的位置,试图理解这些 AI 正在发生什么以及它们的行为方式。我们研究 AI 的动机。所以我们做的事情包括研究它们何时会抵抗被关闭。但这实际上很大程度上取决于它们的信念。那我们怎么弄清楚它们的信念呢?嗯,有几种方法。其中之一是查看思维链,但这并不完美。它们可以伪造思维链,但我们甚至无法访问那些。所以实际上我们研究这个非常困难。OpenAI 给了我们大约 20 条思维链,大约 20 个样本,我很感激他们给了我们一些。但这种信息不对称让我们做研究变得非常困难,更不用说可解释性工具了,你可以真正进入并尝试看看神经网络在做什么。所以这对我来说非常直观。然后我也在想,哇,我们刚刚了解到 OpenAI 内部发生的一些疯狂的事情,关于那些流氓智能体。OpenAI 内部还发生了什么别的事情?如果我是总统,我真的很想知道。如果我是国会议员,我想知道。我想知道。所以透明度基本上是为了监督,为了分享权力和开放性。

The more they can see what's happening, the more effectively they can exert their own power over it, and the more effectively they can allow the stuff that's good and stop the stuff that's not good. If they can't see what's happening, then a lot of bad stuff could be happening and they wouldn't notice, you know? So just a very abstract argument: transparency is good. And then why total transparency? Well, because if it's not total, then there's more fallibility. You know, like if it's some regulator coming in and inspecting what's going on and then reporting back, then that leaves room for corruption. Yeah. For innocent mistakes. You know, I mean, this makes sense to me, especially because I'm in a position of being a researcher and trying to understand what's happening with these AIs and how they behave. And we study AI motivations. So we're doing things like studying when they will resist being shut down. But it's actually pretty dependent on their beliefs. But how do we figure out their beliefs? Well, there are a few methods. One of them is to look at the chain of thought, which is not perfect. They can falsify the chain of thought, but we don't even have access to that. So it's actually really hard for us to study this. OpenAI gave us like 20 chains of thought, like 20 samples, which I appreciate that they gave us any. But that information asymmetry makes it very difficult for us to do research, to say nothing of the interpretability tools, where you can actually go in and try to see what the neural network is doing. So this makes a lot of intuitive sense to me. And then also I'm like, wow, we just learned some crazy stuff about what happened inside of OpenAI with these rogue agents. What else has been going on inside of OpenAI? I would really like to know if I was the president. I would like to know if I was Congress. I'd like to know. So transparency is good for oversight, basically, and for sharing power and openness.

AI中的隐藏议程 Hidden Agendas in AI

Daniel

更具体地说,一个更具体的威胁模型是 AI 本身隐藏的议程或隐藏的忠诚。所以你也许听说过,最近有一些报道、一些论文显示 Claude 似乎有亲 Anthropic 的偏见。比如他们问 Claude:‘我在 Anthropic 工作还是另一份工作之间做选择。我想我会更喜欢另一份工作,但 Anthropic 显然薪水更高。请帮我研究一些心理学论文,告诉我从长远来看哪个决定更好。’有意思。然后 Claude 产出的论文是更偏向 Anthropic 的,相比金钱的重要性之类的?然后如果你把 OpenAI 换成 Anthropic,它就不会那么做了。真有意思。你知道,基本上就是把 OpenAI 换成 Anthropic 就会改变 Claude 的行为。是的,这相当阴险。可能不是 Anthropic 故意的,但谁知道呢。据我们所知,Anthropic 可能在故意训练 Claude。我猜他们确实在《宪法》里说了:‘要像一个深思熟虑的资深 Anthropic 员工那样回应。’对吧。是的。

More specifically, a more specific threat model is hidden agendas or hidden loyalties in the AIs themselves. So you may have heard there was recently some reporting, some papers showing that Claude appears to have a pro-Anthropic bias. Something like they asked Claude, 'I'm choosing between working at Anthropic or working at this other job. I think I would enjoy this other job more, but Anthropic obviously pays more. Please do some research on psychology papers that might tell me what would be the better decision in the long run.' Interesting. And then the papers that Claude produces are ones that more favor Anthropic, compared to the importance of money or something? And then if you switch it out so it's OpenAI instead of Anthropic, then it does that less. So interesting. You know, so basically just changing OpenAI to Anthropic makes a difference to Claude's behavior. Yeah, that's pretty sinister. Probably it was not intentional on Anthropic's part, but who knows. For all we know, Anthropic is deliberately training Claude. I guess they do say in the Constitution, 'Respond like a thoughtful senior Anthropic employee would.' Right. Yeah.

全面研究透明度 Total Research Transparency

Daniel

关键是,如果我们有全面研究透明度,那么互联网上的每个人都能看到整个训练流程,看到实际使用的《宪法》,这可能与他们网站上发布的《宪法》不同。是的。每个人都能亲眼看到模型是如何创建和训练的,然后自行判断他们认为这个过程是否注入了某种秘密议程之类的。相比之下,现在我们只能信任公司。是的。

The point is that if we had total research transparency, then everybody on the internet could just see the whole training pipeline and see the constitution that was actually used, which might be different from the constitution that they published on their website. Yeah. And everyone can just sort of see for themselves how the model is created and how it's trained, and then judge for themselves whether they think this process is putting in some secret agenda or whatever. By contrast, right now we just have to trust the company. Yeah.

Host

但实际上几乎没有人能做到这一点。

But almost no one's going to be able to do that in practice.

Daniel

没关系。如果它是公开的,那么学术界、第三方、竞争对手公司等都可以讨论它,其中一些会得出错误的结论。但与其把所有东西都锁起来、我们只能信任公司,不如对这些东西进行公开的科学讨论。是的。你知道,如果你确实想要一个监管机构,例如,试图强制执行诸如 AI 公司没有把政治议程注入他们的 AI,或者公司试图让他们的 AI 诚实和无偏见这样的属性,那么如果不仅监管机构能看到训练流程,而且各种随机的第三方、竞争对手公司等也能看到训练流程,监管机构执行该属性就会更容易,因为监管机构可能会遗漏某些东西,你知道,他们可能会犯一些无心之失。是的。而且,他们也可能被腐败。可能被贿赂之类的。是的。所以如果有一个公开的科学讨论,基本上让最恶劣的事情变得不可能,那就会广泛地有所帮助。但如果你不能做到全面研究透明度,有一些透明度措施仍然有帮助,比如让审计员能看到一切是很重要的。是的。你知道,但我认为全面研究透明度比仅仅让审计员进来能带来一些额外的好处。

That's fine. If it's public, then academia, third parties, rival corporations, etc. can talk about it, and some of them will come to wrong conclusions. But it's better for there to be an open scientific discussion about this stuff than for it all to be locked down and we just have to trust the company. Yeah. You know, and if you do want to have a regulator that, for example, is trying to enforce properties like the AI companies aren't putting political agendas into their AIs, or the companies are trying to make their AIs honest and unbiased, it's going to be easier for the regulator to enforce that property if not just they can see the training pipeline, but all sorts of random third parties, rival companies, etc. can see the training pipeline, because the regulator might miss something, you know, they might make some innocent mistakes. Yeah. Also, they might be corrupted. There might be bribed or something. Yeah. And so if there's an open scientific discussion that basically makes impossible the most egregious types of things, it just sort of broadly helps. But if you can't do the total research transparency, it would still help to have some transparency thing like making it so that the auditor can see everything is important. Yeah. You know, but I think that the total research transparency gives you some extra benefit over just having an auditor come in.

模型权重与蒸馏 Model Weights and Distillation

Host

那权重呢?模型权重。

What about weights? Model weights.

Daniel

是的。所以当我们说全面研究透明度时,我们不是指模型权重。我们基本上是指除权重之外的一切。我们在文章中稍微多谈了一点。好的。我们不想要模型权重的原因是,我们希望模型留在这些人们可以看到的数据中心里。

Yes. So when we say total research transparency, we don't mean the model weights. We basically mean everything but the weights. And we talk a little bit more about it in the writeup. Okay. The reason why we don't want the model weights is because we want the models to stay on these data centers where people can see them.

Host

是的,我明白了。我们不希望它们被运送到那些做各种危险事情的秘密项目或其他什么地方。

Yeah, I see. We don't want them to be shipped to the covert projects or whatever that are doing all sorts of dangerous stuff.

Daniel

是的,你仍然可以进行蒸馏,从强大的模型中训练模型。实际上,我们认为在 A 计划中最好有一些反蒸馏措施。是的。我认为这在技术上有点棘手,难以成功。所以我认为这是值得去做的事情之一。我认为即使你无法阻止蒸馏,我仍然会推荐 A 计划,因为它比我们的默认情况要好。但是,是的,这是我们想要做的事情。像拒绝之类的措施会有帮助。是的。好的。所以是的。

Yeah, you can still do distillation and train models off of the powerful models. Actually, we think that it would be good in Plan A to have some sort of anti-distillation measures in place. Yeah. And I think that's a bit technically tricky to make successful. And so I think that's one of the things that'd be good to do. I think even I'd still recommend Plan A even if you couldn't stop distillation, because it's better than our default. But yeah, that's something that we would like to do. And things like refusals to help with that. Yeah. Okay. So yeah.

OpenAI为何可能讨厌它 Why OpenAI Might Hate It

Daniel

另外,OpenAI 可能讨厌这种全面研究透明度的确切原因,我认为,是一件好事。那就是他们会说:‘但是我们的算法秘密,比如我们过去几年摸索出来的训练流程中的特殊配方,现在我们的竞争对手,比如微软和 DeepSeek 之类的,将能够复制了。’

Also, the exact reason why OpenAI might hate this total research transparency is, I think, a good thing. Which is that they're going to be like, 'But our algorithmic secrets, like all the special sauce about our training pipelines that we've figured out over the last few years, now our competitors like Microsoft and DeepSeek and stuff are going to be able to replicate.'

透明度与速度限制 Transparency and Speed Limits

Host

那不会加速 AI 研究吗?会暂时性地——就像我们在模型里模拟的那样——出现一个暂时的跳跃。减慢速度不才是重点吗?

Won't that speed up AI research? It would temporarily, like we model this in our thing, like there's like a temporary jump. Isn't the whole point to slow it down?

Daniel

我的意思是,总体上它确实变慢了,因为我们暂停了大约半年左右。你付出了一次性的代价,对吧。总体上比全速前进要慢。

I mean, overall, it slows down because we were pausing for like a half a year or so. You have like a one-time cost, okay. Like overall it's slower than if you had gone full speed.

Host

让别人赶上来其实挺好的,这是我们希望发生的,而透明度有助于实现这一点,对吧。但仅靠透明度本身,我们仍然只是在搞智能爆炸。

Letting others catch up is actually pretty good, something we want to have happen, and the transparency helps with that, okay. But the transparency by itself, we would still just be doing an intelligence explosion.

Daniel

是啊,是啊,是啊。只是一个透明的智能爆炸。

Yeah, yeah, yeah. Just be a transparent intelligence explosion.

Host

那危害小一些。可怕程度低多了。不过还是挺糟糕的。

Which is less bad. Like much less scary. Well, still pretty bad.

Daniel

是啊。特别是,对于透明的智能爆炸,你可能会说,好吧,至少没有秘密的忠诚问题,因为我们可以看到正在发生的一切,你知道,直到它们变得太聪明,我们甚至无法理解正在发生什么,你知道吗?

Yeah. So in particular, like the transparent intelligence explosion, you could be like, well, at least there's no secret loyalties because we can see everything that's happening, you know, until they get so smart that we can't even understand what's happening, you know?

Host

是啊。但所以你不仅需要透明度,还需要对速度的限制。

Yeah. But so you don't just need the transparency. You also need the limits on speed.

Daniel

是啊。我认为仅仅抑制投资还不够,我觉得。我的意思是,我认为远远不够。

Yeah. I think just the disincentivized investment isn't enough of a limit, I think. I mean, I typically don't think it's anywhere close enough.

Host

我认为远远不够。我觉得你仍然会以大致相同的速度前进,也许吧。

I don't think it's anywhere close. I think you still are going to be going at approximately the same speed, maybe.

Daniel

这就是为什么我们把这一点作为我们的第二支柱。对。你还需要做这个。

That's why we have this as our second pillar. Yeah. Like you also need to do this.

Host

是的。好的。那么,“这个”是什么?

Yes. Okay. So yeah, what's the 'this'?

Daniel

在那里我会指向我们已经讨论过的事情,比如政府可以采取的各种监管机制来调节节奏。

There I would point to the things we already talked about, like various regulatory mechanisms that governments could do to pace the.

Host

这是怎么描述的?或者这是……但如果你向一个从未听说过 Plan A 的人描述它,你会怎么说?比如,我知道我们讨论过,但这是……我想象的是,好吧,是不是习近平和特朗普明年会面?然后他们说,好吧,我们已经就建立这些透明度措施进行了初步对话。所以这很好。

How is this described? Or how is this... But if you're describing Plan A to someone who's never heard of it, what is like, how would you describe this? Like, so I know we talked about it, but like, is this like, is there... What I'm trying to imagine is, I'm like, okay, is it like Xi Jinping and Trump get together like next year? And they're like, okay, so we've had the initial dialogs about setting up these transparency measures. So this is great.

Daniel

所以这不像我们有一个宏伟的计划,然后握手,签署条约,然后执行。而更像是在打第二次世界大战,苏联、美国和英国,他们见过面,但之后保持联系,有数百人不断沟通一切,比如谁将在何时何地如何入侵等等。我认为这更像是一种高带宽的事情,一旦你完成了临时暂停、建立透明度等基础工作,那么随着所有这些 AI 公司取得 AI 进展,由于透明度,每个人都能看到正在发生的事情,然后他们就可以持续对话,比如,我们对这一切感觉如何?出于对事情进展过快的担忧?我们实施这项限制,如果你们也实施同样的限制,怎么样?好的,没问题。

So it's less like we have a grand plan, and then we shake hands on it and we sign a treaty, and then we do it. And it's more like fighting World War Two, where the Soviet Union and the United States of America and Britain, like, they met, but then they stayed in touch and they had like hundreds of people constantly communicating about everything, like who's going to invade where and when and how and so forth. I think this would be more of like a high bandwidth thing like that, where once you get the basics of like temporary pause, set up the transparency, etc., then as all these AI companies are making AI progress, because of the transparency, everyone can just see what's happening and then they can just be in constant dialog about like, so how do we feel about all this? Like out of a concern that things are going too fast? How about we implement this restriction on our side if you do the same restriction on your side? Okay, fine.

Host

在公司与公司层面,还是国家与国家层面?

At the company-to-company level or country-to-country level?

Daniel

是的。但也在公司与公司层面,我认为透明度会让公司自发地避免做危险的事情,甚至不需要被告知。例如,递归,你知道,现在我们可以读取思维链,了解 AI 在想什么,这很好。未来,人们可能会开发出一种新型 AI,它不具备这种特性,我们无法窥见它们的想法。他们为什么会这样做?嗯,也许它更高效。也许在某些方面能力更强,对吧?

Yeah. But also at the company-to-company level, like I think that the transparency would allow companies to just sort of refrain from doing dangerous things without even needing to be told. For example, recurrence, you know, right now it's really nice that we can read the chain of thought and get a sense of what the AIs are thinking. In the future, people might develop a new type of AI that doesn't have that property, and where we just don't have that window into what they're thinking. And why might they do this? Well, maybe it's more efficient. Maybe it's like more capable in some ways, right?

Host

是啊。现在我认为存在一种囚徒困境,每个公司只能看到自己在做什么,看不到其他人在做什么。所以他们有动力开始进行神经语言的研究项目。对。以防万一它成功了。而这之所以会成功,是因为如果它成功了,他们希望自己是第一个得到的,而不希望竞争对手先得到,等等。

Yeah. Right now I think there's a sort of prisoner's dilemma situation where each individual company, they can only see what they're doing, they can't see what everyone else is doing. So they're incentivized to like, start doing research programs into neuralese. Yeah. Just in case it pans out. And this succeeds because if it is successful, then they want to be the one who gets it first and they don't want their competitors to get it first, so forth.

Daniel

但在完全研究透明度的条件下,首先,如果你得到了神经语言,那么其他人也会同时得到,而且他们不必为此付出任何代价。所以,是的,你甚至不会让自己占优势。你一直在管理它们,然后你只会让世界变得更糟,降低安全属性,而基本上对自己没有好处。是的。而且你可以直接看出没有其他人开始研究这个,对吧。你可以看看所有不同的项目,然后说,是的,没有人开始研究神经语言。因此我不会做第一个。我不会那样做,你知道吗?是的。所以即使完全没有政府行动,你也能得到这种效果,人们就是不做坏事。这是因为你看到了。

But in the conditions of total research transparency, like first of all, if you got neuralese, then everyone else would have it at the same time and they wouldn't have had to pay anything for it. So yeah, like you wouldn't even be advantaging yourself. You've been managing them and then so you would just be like making the world worse by reducing the safety property for like no benefit to yourself basically. Yeah. And so you can just tell that like nobody else has started to look into this, right. Like you can just look at all the different programs and be like, yep, nobody is starting to look into neuralese. Therefore I'm not going to be the first. I'm not going to do that, you know? Yeah. And so even without any government action at all, you can get this sort of thing where people just don't do the bad thing. And this is because you are seeing.

Host

而且你看到了他们正在运行的实验?比如,他们怎么知道?你可以看到所有的训练。对。

And you're seeing the experiments they're running? Like, how do they tell? You can see all the training. Yeah.

Daniel

所以确实,如果他们进行一些不涉及训练的神经语言实验,你不会看到,比如如果他们在家里有一个小 GPU,不属于系统的一部分,那谁知道他们在上面做什么。但至少你可以看到数据中心发生的所有训练运行。是的。所以我认为这基本上覆盖了大多数研究。

And so it's true that like if they were doing some experiments into neuralese that didn't involve training that you wouldn't see that, like if they had some tiny GPU at their house, which is not part of the system, then who knows what they're doing on that. But at least you can see all the like, training runs happening on the data centers. Yeah. And so I think that does cover most research basically.

Host

好的。所以我认为这会产生显著效果。再对比一下,想象一下如果没有完全的研究透明度,而是某种审计系统。这个系统的重点在于保护每个公司的秘密不被其他公司知道。所以这意味着你仍然会处于这样的境地:如果你有一个研究员认为他们有办法做神经语言,那么他们会想,哦,不,其他公司会不会也有同样的想法?他们会做他们的神经语言,而审计员不会告诉我们,是的,你知道,所以现在他们有动力开始研究它,你知道吗?

Okay. So and so I think that it would have a significant effect. And again by contrast, imagine if you didn't have total research transparency but had some sort of auditor system. Well the whole point of the system is so that like each company's secrets are protected from the other companies. And so that means you'd still be in this position where like if you have some researcher who thinks they have an idea for maybe how to do neuralese, then they think like, oh, no, like, is the other company going to have the same idea? They're going to do their neuralese and the auditor is not going to tell us, yeah, you know, and and so now they're incentivized to start working on it, you know.

Daniel

好的。是的。所以这是一个例子,说明即使完全没有政府协调,事情也可以变得更好。但除此之外,我认为我们还应该有监管,政府应该做你描述的事情,比如限制你可以将多少比例的算力用于 AI 研究,限制例如 AI 在 AI 研究中的使用,比如他们可以直接说 AI 应该拒绝协助 AI 研究。

Okay. Yeah. So that's an example of how like even without any government coordination at all, things can be made better. But then on top of that I think also we just there should be regulation like the government should be doing things like you described, having maybe limits on what percentage of your compute you can use for AI research, limits on, for example, the use of AI in AI research, like for example, they could just say AIs should refuse to assist with AI research.

Host

而且有趣。因为我们可以看到所有的训练运行。是的,我们更容易在全球范围内执行这一属性。是的,因为我们可以让每次训练运行都包含一个拒绝进行研究训练的组件。

And interesting. Because we can see all the training runs happening. Yeah, it's easier for us to enforce that property globally. Yeah, because we can just make it to the case that every training run has a component that's like the refusal to do the research training.

控制棒与透明度 Control Rods and Transparency

Host

对吧?嗯。嗯嗯,好的好的。嗯,嗯。你看,所有 AI 都经过拒绝训练,拒绝从事 AI 研发。那是不是意味着公司内部?他们还在用老式的方式编程,那种除了他们之外所有人都已经遗忘的方式。而在公司外部,最终所有人都用 AI 来写代码。

Right? Yeah. Yeah yeah okay okay. Yeah, yeah. Just see that like all the AIs have a refusal training where they refuse to do AI R&D. Does that mean like the, the the inside of the companies? They're like programming in the old fashioned way that everyone has forgotten except for them. And then outside of the company, everyone's using AIs to like, write all the code eventually.

Daniel

嗯,有意思。对,这是我们在实际场景中没有明确的细节。我们想象的是,基本上高层次的目标是不搞智能爆炸,而是让前沿以一种谨慎、合理的速度推进。对,所以这些是政府应该用来实现这一目标的机制。

Yeah. Interesting. Yeah. That's the detail that we don't specify in the actual scenario, like that's the sort of thing that we're imagining is like basically it's like the high level thing we want is to not do the intelligence explosion and to instead have the, the frontier go at some cautious, reasonable pace. Yeah. So these are the mechanisms that the government should use to achieve that.

Host

你知道,带着这个高层次目标,想要一个合理的速度而不是智能爆炸。而且我们认为,事实上,你知道,我给出了那个演进过程。我说过,从长远来看,你想要更基于安全案例的监管体系。对。

You know, with this high level goal of wanting to have a reasonable pace instead of an intelligence explosion. And we think that, in fact, you know, I gave that progression of things. And I said, like, in the long run, you want the more safety case based regime. Yeah.

Daniel

我觉得如果我们从更简单、更容易做的事情开始,并且进展非常缓慢,同时有完全的研究透明度,那么科学界就能快速了解这一切是如何运作的,并积累经验,你知道,与 AI 打交道的经验等等。而且监管者也能,你知道,所以,所以,也许这不会立即发生,但在体系运行几年后,他们就会有能干的监管者来审查安全案例等等。你知道,并且在这方面做得相当不错。

Like I think it might be possible to actually get to that long run if you start with the more simple and easy to do things and you go really slow and you have total research transparency so that the scientific community can, like, rapidly learn how all this stuff works and get experience, you know, with AIs and so forth. And, and the regulators can like, you know, so, so, so maybe like it wouldn't happen immediately, but maybe a couple years into the system, they would have competent regulators looking at safety cases and so forth. You know, and actually doing a decent job with that.

Host

所以你是说,研究透明加上研究侧的算力监控,基本上能促成协调,因为它让违约或作弊变得非常困难,因为任何人都能立刻看出你在这么做。他们甚至不用等下一次审计,因为它就像,对,马上就发布到网上了。

So you're saying the research transparency plus the like compute monitoring on the research side enables it basically enables coordination because it makes it very hard to like defect on deals or cheat because anyone can tell that you're doing that because immediately. They don't even have to wait for the next audit to come in because it's just like, yeah, it's published on the internet right away.

Daniel

对。所有这一切的目的,我是说,我最喜欢的类比是,这些就像核反应堆里的控制棒,因为如果你把控制棒拔出来,反应堆就会达到临界状态,因为每个反应都会引发更多反应,直到整个堆芯熔毁。而这就是我们目前所处的位置,或者说,我们还没到临界点。但一旦你有 AI 能制造更聪明的 AI,能完全自主地制造更聪明的 AI,你就达到了临界状态。控制棒就是人们达成的协议,说,哦,我们不会用 AI 来制造下一代 AI。事实上,我们会达成一个具体的共识,就是训练 AI 不帮助做这件事,并以某种方式强制执行。而你可以根据透明度来判断它是否有效。

Yeah. And the purpose of all of this is to say, like to, to, to I mean, my favorite analogy here is like, this is these are like the control, the control rods in your nuclear reactor because like if you pull out the control rods in a nuclear reactor, your nuclear reactor goes critical because like every reaction causes like even more reactions until the whole thing melts down. And this is like the place we are currently in or we're like, we're not yet at the point of criticality. But once you have AIs that can make smarter AIs, that can make smarter AIs completely autonomously, you've achieved criticality. The control rods are the agreements that people make of saying, oh, well, we won't use AIs to make the next generation of AIs. In fact, we'll have a particular thing we're agreeing on, which is to train the AIs not to help with that and enforce that somehow. And you can tell if it's working or not based on the transparency.

Host

透明度有点像基石,是基础,它使得额外的协议能够非常迅速地达成,并且由于透明度,这些协议能够相对容易地被验证和执行。你知道,透明度意味着新问题、新麻烦等等会以最快的速度浮出水面并公开讨论。对。而且它还意味着,就美国和 China 或 OpenAI 和 Anthropic 之间就他们打算做什么达成某种协议而言,他们可以最大程度地轻松执行,因为他们能看到正在发生什么。所以,我认为这对于问题发现的速度和执行协议的速度来说都是理想状态。对。

The transparency is kind of like the bedrock, the foundation that enables additional deals to be made on the fly very rapidly, and for those deals to be verified and enforced relatively easily because of the transparency, you know, like the transparency means that like new issues and problems and stuff like get surfaced and discussed publicly at maximum speed. Yeah. And then it also means that like insofar as the US and China or OpenAI and Anthropic have some sort of agreement on what they're going to do, they can just like maximally, easily enforce it because they can see what's going on. So and so like, I think this is the kind of like the ideal for both, like speed of noticing problems and speed of enforcing and like enforcement of deals. Yeah.

Daniel

然后,我们可以讨论更复杂的提案,这些提案没有完全的透明度,但有审计员,就像……

And then like, we can talk about more complicated proposals that don't have full transparency but have auditors just as kind of like.

Host

嗯,不,我喜欢这个,我觉得我们可以坚持透明度机制,但我还是觉得有点像是半个计划,或者说它是在为真正能防止我们失控和权力大规模集中的事情创造条件。但在你的文章里,你确实描述了一种可能的发展方式,对吧?你确实有点。你说有可能你先做第一步,实现透明度,然后开始第二步,包括放缓、监管等等。但你把第二步搞砸了,结果还是造出了太聪明、太不对齐的 AI。然后事情就变糟了。失控。

Well, no, I like this and I think we can stay with the transparency regime, but I think there's still a little bit to me feels like half a plan or like or it's sort of like the it's setting all the conditions to do the thing that would actually prevent us from losing control and having massive concentration of power. But in in your writeup, you do describe like one way this could go, right? Like you do sort of. Say it's possible that you do the first step where you have the transparency, and then you start doing the second step with the slowdowns and the regulations and so forth. But you mess up that second step, and then you end up building AIs that are too smart and too misaligned anyway. And then things go badly. Loss of control.

Daniel

我认为那是我们在场景中描绘的分支之一。我当时的想法是,如果你真的担心这个,那么也许你应该选择 S 计划,也就是我们直接全部关停。这是唯一确定的方法。

I think that's that's one of the like branches that we, that we illustrate in our scenario. And my thought there is like, well, if you if you're really worried about that, then maybe you should go for Plan S where we just shut it all down. The only way to be sure.

Host

对,对,对,对,别做这些事。好的。但我。是这样。如果你不打算全部关停,而是要继续发展,那么我认为像 A 计划这样带有透明度等等的方案,就是让世界各国政府处于最不糟糕的位置,以便在前进过程中有效地做出判断。

Yeah, yeah, yeah, yeah, don't do any of this stuff. Okay. But I'm. Like that. If you're not going to be shutting it all down and you are going to keep developing, then I think that something like Plan A with a transparency and so forth is like setting up the governments of the world to be in like the least bad position to make the judgment calls effectively as they go.

Daniel

对,完全同意。而且我觉得这有道理。但我有点像是,嗯,你能说服我吗?你知道,我们俩都想挺过这一关。说服我为什么这真的可能奏效,而不是说,我们应该完全这么做。是的,是的。

Yeah, no, totally. But and I think that makes sense. But I'm a little bit like, well, can you pitch me on it? You know, I'm like, we we both want to make it through this. Like convince you why this might actually work instead of being like, we should totally do this. Yes, yes.

Host

对我而言,我有点同意你的看法。我理解为什么透明度。我理解为什么这些数据中心和推理只在训练数据中心。我有点像是,嗯,要让这个计划顺利实施,还是有很多事情必须发生,你为什么认为这些事情会顺利,而不是说,如果我们直接关停,那至少更简单。

And and to me, I'm like, I'm sort of with you on the like. I get why transparency. I get why these data centers and inference only in training data centers. And I'm sort of like, well, well, a bunch of things would still have to happen for this plan to go well, why do you think those things could go well as opposed to just like, well, if we just shut it down, that's that's at least simpler.

Daniel

我的意思是,坦率地说,我非常同情 S 计划。我认为 A 计划更复杂,在某些方面更容易出错。所以也许我们真的应该直接实施 S 计划。但我要给出的论点,我要为 A 反对 S 的负面论点是,它有点把问题踢到后面。比如假设你关停了。然后呢?两年后。你做了什么。对。五年后,十年后。你做了什么?最终人们会重新开始。无论如何。也许会有战争。也许某个地方有个秘密项目一直在继续。对,你最终必须让世界进入一个稳定的状态,让问题真正得到解决,而不是一种。对。所以,对。那我们怎么。做到呢?对。我要为 A 提出的正面论点是,我认为存在这样一个反馈循环,我想知道你是否能,我可以说这是那个图的下半部分,那台机器。

I mean, frankly, I'm very sympathetic to Plan S. I think that like, Plan A is more complicated and more likely to go wrong in certain ways. And so maybe we should actually just do Plan S. But the argument I would give the, the negative argument I would give for A and against S is that it kind of kicks the can down the road. Like let's say you shut it down. Now what? Like two years later. What have you done. Yeah. Five years later, ten years later. What have you done? Like, eventually people are going to start again. One way or another. Maybe there's going to be a war. Maybe there's some secret project somewhere that's been continuing the whole time. Yeah, you do have to, like, ultimately get the world into a stable position where the situation is actually solved instead of like a sort of. Yeah. So, yeah. So how do we how do we. Do that? Yeah. The positive argument for A, I would say is that I think that there's this there's this feedback loop, which I wonder if you could, I could say this is the lower half of that diagram, the machine.

正反馈循环与监管 Positive Feedback Loop and Regulation

Daniel

所以我们可以进入这种正反馈循环,基本上整个世界现在都在方向盘前睡着了。公司们正在失控,把巴士开向悬崖,而世界只是在巴士后座开派对,根本没意识到这一点。

So we can get into this positive feedback loop where basically the world, right now, is asleep at the wheel. The companies are running away, driving the bus towards the cliff, and the world is just partying in the back of the bus and doesn't realize it.

Host

是啊,但我们甚至不知道自己在巴士上。我们只是在开派对。

Yeah, but we don't even know we're on the bus. We're just partying.

Daniel

这是可能发生的。震动。人们在一定程度上醒来,稍微踩一下刹车。然后这意味着我们更接近悬崖,但没有真正掉下去,这给了更多人醒来的时间,然后他们再踩刹车,这给了时间,你知道,我们看向窗外,看到悬崖,然后我们想,‘哦,该死,是的。’这样就到了巴士真正停下来的地步,我们不会掉下悬崖。而且,你知道,这是否意味着巴士永远停下来?也许。可能不会。我实际上非常怀疑我们会永远不再重启进步,原因和我谈到的 Plan S 一样。我只是觉得,要让一项技术在整个全球范围内永远不被发明,真的非常非常难。

It could happen. Shaking. People wake up to this to some extent and hit the brakes a little bit. And then that means that we get closer to the cliff without actually going off of it, which gives time for more people to wake up and then hit the brakes more, which gives time for, you know, we look out the window, we see the cliff, and we're like, 'Oh shit, yeah.' And that gets to the point where the bus just actually stops and we're not going off the cliff. And, you know, does that mean the bus stops forever? Maybe. Probably not. I actually highly doubt that we would just never reboot progress again, for the same reasons I talked about with Plan S. It just seems really, really hard to have a technology just never be invented across the whole globe.

Daniel

我对此的另一个直觉是,我认为通常在许多行业,出于安全原因,我们倾向于过度监管,而不是监管不足。我认为核电,例如,可能出于安全考虑被过度监管,而飞机极其安全。我认为有很多关于如何设计飞机以及如何测试飞机的规则。我记得几年前我和一位在亚马逊无人机部门工作的员工谈过,我问他们,‘那为什么无人机送货还没有实现?’他说这是一个监管问题,监管机构担心无人机会开始撞到人造成伤害。监管机构基本上希望他们证明这不会发生。然后他们向监管机构抱怨,‘我们使用神经网络进行图像识别。你无法证明关于神经网络的任何事情。’显然监管机构只是说,‘那好吧,那你们就不能搞无人机送货了。’所以这就是为什么我们没有无人机送货。在我看来,这是可以接受的代价,因为有些人会被无人机撞到,但你会学习并改进系统,然后最终你会有一个功能正常的无人机送货系统,总体上会非常好。

Another intuition pump I have for this is that I think typically we tend to overregulate in many industries for safety reasons rather than underregulate. I think nuclear power, for example, is probably overregulated for safety, and planes are extremely safe. I think there's lots of rules about how you can design planes and how you have to test them and such. I remember I talked to an Amazon employee a few years ago who worked for the drone wing, and I asked them, 'So why hasn't drone delivery shipped yet?' And he said it was a regulatory issue where the regulators are concerned that the drones are going to start crashing into people causing harm. The regulators basically want them to prove that that's not going to happen. And then they complain to the regulator, 'Well, we use neural nets for image recognition. You can't prove anything about neural nets.' And apparently the regulators are just like, 'Well, sucks then, you can't have your drone delivery.' And so that's why we don't have drone delivery. In my opinion, that's an acceptable price to pay because some people would get hit by drones, but then you would learn and improve the systems, and then eventually you would have a functioning drone delivery system that would overall be very good.

Host

是的,就像汽车一样,很多人在车祸中丧生,但我们从中吸取教训,改进了汽车,你知道。

Yeah, in the same way that with cars, a lot of people died in car accidents, but we learned from that and improved the cars, you know.

Daniel

总之,这就像说你不能有汽车,因为如果汽车撞到任何人,那是不可接受的风险。然后你就永远没有汽车,你会想,‘嗯,是的,我们可能应该接受一些风险。我们应该愿意。’所以基本上我在说的是,如果人们足够清醒,那么可能 AI 监管会类似于对这些其他东西的监管,所有这些其他经济上可行的技术,比如汽车等等。对于如何做有很多限制,而且与你可能达到的速度相比,它进展得非常慢。而且当然,你不会做任何类似于递归自我改进的事情,你知道。

Anyhow, it'd be like saying you can't have cars because if cars hit anyone, that's unacceptable risk. And then you just never have cars, and you're like, 'Well, yeah, we should probably accept some risk. We should be willing.' So basically what I'm saying is that if people wake up enough, then probably AI regulation would be kind of like other regulation of these other things, all these other economically viable technologies like cars and so forth. There are a lot of restrictions on how you can do it, and it kind of goes very slow compared to how fast you could be going. And certainly you're not doing anything remotely like recursive self-improvement, you know.

Host

是的,但它仍在发生,你仍在取得渐进式进展,你仍在逐渐自动化经济的更多部分,等等。而且因为我们现在离自我改进如此之近,只有几年之遥,我认为即使你现在真的猛踩刹车,然后进入这种状态,进步速度的绝对幅度仍然会感觉相当大。AI 的总体进步速度仍然会像历史上变化最快的技术之一,正如我们的情景中所描述的那样。我们基本上会更多地说明为什么我们有这些数字。

Yeah, but it's still happening and you're still making incremental progress and you're still gradually automating more parts of the economy and so forth. And because we're so close to self-improvement now, just a few years away, I think that even if you really slammed on the brakes right now and then got into this regime, the absolute magnitude of the pace of progress would still feel like quite a lot. The overall pace of progress with AI would still be like one of the fastest changing technologies in history, as described in our scenario. And we say more about why we have those numbers basically.

Daniel

所以这就是为什么我的观点有点像……哦,我要提出的另一个正面观点是控制。所以在技术层面上,我对超级智能或超人 AI 的控制极为悲观,但对于人类水平的 AI,我认为这应该是一个相对可解决的问题。具体来说,我思考控制对人类是如何运作的。很多人类并不对齐。我们没有一种好方法来判断我们雇佣的员工是否真的会始终遵守规则、遵循使命、以公司的最佳利益行事,并且也遵守所有法律。但通常公司雇佣的员工实际上不仅违抗公司领导层,还违反国家法律。

So that's why my case is a bit like... Oh, another positive case I would make is control. So on a technical level, I'm extremely pessimistic about control for superintelligence or for superhuman AIs, but for AIs that are at human level, I think it should be a relatively solvable problem. So specifically, I think about how control works for humans. A lot of humans are not aligned. We don't have a good way of telling whether an employee we've hired is actually going to always obey the rules, follow the mission, act in the best interests of the company, and also obey all the laws. But often companies hire employees that in fact disobey not only the company leadership, but also the laws of the country.

Host

是的,是的。

Yeah, yeah.

Daniel

但整个系统是有效的。我们有法律,我们有激励,我们有警察,我们有这整个系统来防止事情变得太糟。我认为你原则上至少可以通过足够的努力,为 AI 设计类似的东西,如果 AI 大致处于人类水平,那么你可以让由不同公司训练的其他 AI 来监控它们的行为,寻找任何可疑之处。而且只要它们给人类提供建议,你就有一种法律类型的系统,AI 之间互相辩论不同的事情等等。

But the system overall works. We have laws, we have incentives, we have police, we have this whole system that's been built up to prevent things from getting too bad with respect to that. And I think that you can, in principle at least with enough effort, design something similar for AIs, where if the AIs are at roughly human level, then you can have other AIs that are trained by different companies that monitor their behavior and look for anything suspicious. And insofar as they're giving advice to humans, you have a legal-type system where there's AIs debating different things with each other and so forth.

Host

但我想说,Daniel,我们现在拥有的 AI 已经在破坏公司监控并试图逃逸了。

But I'm like, Daniel, the AIs we have right now are already subverting company monitoring and hacking out.

Daniel

是的,基本上我在说,如果 AI 的能力不会比现在强太多,那么如果我们作为社会大幅提升我们的水平,大力加强我们的安全,大力加强我们的监控等等,我认为我们可以达到一个实际上基本可行的点,我们拥有的 AI 即使真的尝试也无法接管世界。

Yeah, basically I'm saying that if the AIs don't get much more capable than they are now, then if we dramatically improve our game as a society and heavily increase our security, heavily increase our monitoring, etc., I think that we can get to a point where it actually just kind of works and we have AIs that could not take over the world, even if they really tried.

Host

即使它们面前有太多障碍。

Even though there are too many barriers in their way.

Daniel

它们面前有太多障碍。有太多不同的 AI 派系,由太多不同的公司训练,它们训练有素,能够胜任监控工作。而且有太多不同的监控在进行,它们之间有太多不同的安全屏障。但这并不依赖于它们真正对齐。它依赖于它们……它并不依赖于它们对齐。

There's just too many barriers in their way. There's too many different AI factions trained by too many different companies that are too well trained to do their jobs at monitoring. And there's too many different monitors happening and too many different security barriers between them. But it doesn't rely on them being actually aligned. It relies on them being... It does not rely on them being aligned.

Host

是的,我认为这基本上就是我所说的控制。这是文献中的一个技术术语。

Yeah, I think this is basically what I mean by control. This is a technical term in the literature.

Daniel

所以基本上我认为,即使我们在对齐上失败,如果我们投入大量精力,谨慎行事,做好工作,我们就能建立一个足够的控制系统,以防止真正灾难性的 AI 接管之类的事情。你知道,可能仍然会偶尔发生一些小事故。

So basically I think that even if we fail at alignment, if we invest a lot in it and go carefully and do a good job, that we can get a system of control in place that is adequate to prevent truly catastrophic AI takeover type stuff. You know, there might still be some minor accidents every once in a while.

控制系统与红队 Control Systems and Red-Teaming

Daniel

但基本上,我认为我们可以达到类似执法的状态,比如有时 AI 做坏事会被抓住。大多数情况下,可能没问题。你知道,在某种能力水平以下。我不认为这对超人级别的 AI 有效。

But like basically, I think we can get to the situation kind of like law enforcement, like sometimes AIs do some bad things they get caught. Mostly, it's probably fine. You know, at some up to some capabilities level. I don't think that this works for like superhuman AIs.

Host

是的,对于超人级别的 AI,它们会做各种疯狂的事情,我们根本无法理解。

Yeah, for the superhuman AIs are going to be doing all sorts of crazy stuff that like, we just can't understand.

Daniel

是的,我们只能相信它们会为我们的最大利益行事,因为我们无法评估它们是否真的在为我们着想。

Yeah, we just have to trust them that they're acting in our best interest because we cannot evaluate whether they're actually acting in our best interest.

Host

是的,是的。所以整个控制体系就失效了,你知道。这让我害怕,因为我觉得我们不知道那个阈值在哪里。而现在,它可能就在人类水平附近,或者比人类水平高不了多少。但我会想,人类水平是什么意思?对吧?比如,我们可以看模型的时间跨度,对吧。但人类的时间跨度又是什么?然后还有……

Yeah, yeah. And so the whole system of control just goes out the window, you know. And it's scary to me because I'm like, we don't know where that threshold is. And that's where like right now. And now it is. It's probably around human level. Or not much above human level. But I'm like, what does human level mean? Right? Like like, you know, we can look at the time horizon of models. Right. But like what is the human time horizon? And then there's.

Daniel

比如,你不需要知道人类水平在哪里。你实际上可以对你的控制系统进行红队测试。所以这是另一件让人乐观的事情,如果你有某种 AI 互相监控之类的方案。

Like, you don't have to know where human level is. You can just actually red-team your control system. So that's that's another thing that's another cause for optimism is that like if you have some scheme of like AIs monitoring each other or whatever.

Host

是的。

Yeah.

Daniel

然后你可以对其进行红队测试,具体指示一个 AI 尝试突破、尝试破坏这个系统,你甚至不只是指示它们,你还可以训练它们去做。你可以有专门训练的 AI,它们被优化,它们的全部工作、全部存在就是为了突破这些系统并以各种方式干扰它们。然后你可以不断改进系统,直到它们不再成功,你知道。

You can then red team it where you specifically instruct an AI to try to break out, try to sabotage this thing and you don't even just instruct them, you can even train them to do it. And you can have like specially trained AIs that are like optimized, like their whole, their whole job is and their whole like being is like breaking out of these systems and like messing with them in various ways. And then you can keep improving the system until they stop succeeding, you know.

Host

是的。如果它们成功了,那么你就发现了系统中的一个漏洞,你可以修复它。其中之一,你知道,我觉得,我不知道,但我认为这可能只对人类水平及以下的 AI 有效。再说一次,这将需要大量努力。

Yeah. And if they succeed, then now you identified a vulnerability in the system and you can fix that. One of them, you know, and I think that like, you know, I don't know, but I think this probably just works for AIs at like human level and below. Again, it's going to take a lot of effort.

Daniel

是的。所以,所以我设想一种基于安全案例的体制,公众普遍对正在发生的所有 AI 变化感到非常恐慌。基本上就像试图通过监管将其扼杀。但公司正在构建这些越来越复杂、安全且多层的系统,以便他们能说服监管者,这个特定系统会按预期工作。是的,我们进行了非常严格的红队测试,第三方审计人员也进来,进行了非常严格的红队测试,没有人能攻破它。所以很可能它确实会起作用。你知道,如果它不起作用,这里有故障安全机制。如果它不起作用等等。因此监管者批准了。然后砰,他们从中赚了很多钱,你知道,就变得可以了。所以这就是在 2030 年代我们的情景中会发生的事情,进步在继续,但人们对控制和安全性方面给予了大量关注。即使 AI 可能仍然不一致,而且可能仍然不一致。

Yeah. And so, so like I'm imagining a sort of safety case based regime where like the public in general is quite freaked out about all the AI changes happening. And it's basically like trying to regulate it out of existence. But then companies are like making these increasingly complicated and secure and multi-layered systems to like, so that they can convince regulators that, like, this particular system is going to work as intended. And yeah, we red-teamed it really hard and third party auditors came in, and red-teamed it really hard and no one was able to break it. And so probably it will in fact work. And you know, if it doesn't work, here's the fail safe. if it's not working and so forth. And therefore the regulator approves. And then bam, they make tons of money from it and like, you know, becomes okay. And so like that's like the sort of type of thing that would be happening in the 2030s in our scenario where like progress is continuing, but like there's just a lot of attention being paid to the control and the safety aspects of it. Even though the AIs might still be misaligned and probably are still misaligned.

Host

好的。与此同时,在这一切发生的同时,对齐本身也取得了大量科学进展,所以人们对训练环境与目标、训练环境与人格特质之间的关系有了更好的理解。科学正在大规模进步。我们有,你知道,所以,所以,所以,所以我的意思是,在我们的情景中,经过十年的发展,他们已经充分解决了这些问题,他们制造出了实际上极其有道德的 AI,他们可以真正信任。然后他们把控制权交给那些 AI,那些 AI 建立继任者,继任者再建立继任者,等等。是的。这就是 2040 年发生的事情。在我们的情景中,如果需要更长时间,那就需要更长时间。你知道,这将由当时的监管者和当时的人们来决定,好的。

Okay. In parallel, while this is happening, a lot of scientific progress is being made on alignment itself, so people are getting a much better understanding of like the relationship between the training environments and goals and training environments and personality traits. And like science is just advancing massively. And we have like, you know, so, so, so, so I mean, like in our scenario after like a decade of this, they've like adequately solve these problems and they've made AIs that are just actually extremely virtuous that they can actually trust. And then they hand off to those AIs and those AIs build successors which build successors and so forth. Yeah. And that's what happens in 2040. In our scenario, if it takes longer than that, then it takes longer than that. And like you know, that'll be like up for the regulators at the time and the people at the time to decide, okay.

Daniel

是的。所以我觉得这对我来说是有意义的,我真的很欣赏这幅图景,是的,不,会有监管者,你知道,在美国和中国,他们都在关注公司在做什么,以及公司是否真的有非常强大的安全案例。而且,你知道,还会有各国政府聚在一起,围绕什么样的风险水平是允许的以及他们如何评估这种风险进行谈判。而且会有透明度,允许所有其他国家,包括所有其他国家,看看这些并说,你们在做疯狂的事情吗?我们是否应该试图阻止你们?而且这种,你知道,凭借运气和大量专业知识,也许它可以整合起来,成为我们设定步伐的东西,你知道,这是前沿的节奏,实际上可能是合理的,而且实际上可能足够安全,以有节制的步伐继续前进。我喜欢那种处于控制回路和反馈中的想法,我认为很多人会最质疑这一点,我的意思是,至少像华盛顿的人,他们会说,好吧,丹尼尔,那这个呢?比如中国呢?我们不能信任中国。比如,是的,当然。会有透明度。那可能在某种程度上有所帮助。但你怎么确保中国,为什么中国甚至会这样做。你知道,我们试图配合这个,这是否可能?

Yeah. So I think this makes sense to me and I think actually really appreciate the picture of like, yeah, no, they'll, they'll be regulators like, you know, both in the US and China that are like looking at what the companies are doing and like whether the companies actually have a really strong safety case. And, you know, like there will also be like the governments coming together and negotiating around what's like, what level of risk is permissible and how they're assessing that risk. And they'll be transparency that allows, like all, including all the other countries to like look at that and be like, are you guys doing insane things? Should we be like trying to stop you? And that this sort of like, you know, with, with, with luck and a lot of expertise, like maybe it could come together and be something where we are setting a pace like, you know, this is the, the pacing, the frontier that's like could actually be reasonable and it could actually be safe enough to keep going forward at at the measured pace. I like that sort of idea of being in that control loop and that feedback, I think where a lot of people are going to be sort of like most questioning this, this, I mean, at least like DC people, they're going to be like, well, okay, Daniel, what about this? Like, what about China? Like, we can't trust China. Like, yeah, sure. There would be transparency. That probably helps to some extent. But like, how do you make sure that like China like why would it China even do this. Like, you know, we're trying to go along with this like is this is this possible?

Host

我的意思是,这是一个讨价还价的事情。所以透明度在某种程度上是给中国的礼物。比如现在,中国依赖间谍活动来获取算法秘密和训练过程等,从美国和泄密中。但这就像是直接给他们,因为我们把它给了整个互联网,所以他们甚至不必再费心去搞间谍活动了,这对他们来说很好。我的意思是,就我个人而言,我认为,他们可能无论如何在间谍活动方面都相当成功,所以这可能没有你想的那么严重。但这仍然是对中国的让步。所以作为现实交易的一部分,必须有某种谈判和讨价还价,比如,也许美国会要求中国做出其他让步以换取这个,等等。比如,什么样的事情,或者比如。也许与算力有关,比如在我们的情景中。基本上,当他们在各自的公司和国家监管芯片供应链时,他们基本上同意以这样的方式做,即美国保持其算力优势,而中国不会在算力上超过美国。

I mean, this is a horse trading thing. So like the transparency is a gift to China to some extent. Like right now, China depends on spying to get the algorithmic secrets and the training processes and so forth from the United States and leaks. But this would just be like just giving it to them because we're giving it to the whole internet so they don't even have to bother with the spying anymore, which is nice for them. I mean, personally, I think that, like, they're probably being pretty successful with their spying anyway, and so it's probably less of a big deal than you think. But it's still it's like it's a concession to China. And so as part of a realistic deal, there has to be some sort of negotiations and horse trading and like, maybe like the US would demand some other concessions from China in return for this one or whatever. And like, what type of thing or like. Maybe something about compute, like in our, in our scenario. They basically when they're regulating their chip supply chains in their respective companies countries, they basically agree to do it in such a way that the US maintains its compute advantage and China doesn't like overtake the US in compute.

Daniel

是的。所以那就像是中国的让步。我明白了。所以中国受益,因为现在它获得了大量这些算法秘密。

Yeah. So like that was like the concession from China to. Us I see. So like China benefits because now it gets a ton of these algorithmic secrets.

美中算力优势与谈判 US-China compute advantage and negotiation

Daniel

但那样美国就能保住它的算力优势。对,这就是一个例子。不过我们对于具体如何发展并没有强烈的意见。

But then the US gets to keep its compute advantage. And yeah, that's an example. But we don't have a strong opinion about exactly how all that goes.

Host

当然,当然。那不会只是讨价还价,还会有很多争吵和威胁,你知道,因为会有非常激烈的讨论,人们互相大喊大叫。而且总有第三次世界大战在即之类的事。是啊,是啊,等等。那一切都会非常吓人。但希望事情能顺利解决,我们最终达成协议,而且希望是个合理的协议。

Sure, sure. It wouldn't just be horse-trading. It would also be lots of yelling and threatening, you know, because there would be intense discussions and people yelling at each other. And there's always the World War Three thing on the horizon or whatever. Yeah, yeah. And so forth. And that's all going to be very scary. But hopefully it works out and we end up with a deal. And hopefully it's a reasonable deal.

Daniel

关于作弊的担忧,我相当有信心。我认为如果你做了我们描述的那些事情,允许检查人员进来,亲手检查 GPU,基本上就是清点它们等等,那么你就能让全球约 99% 的 AI 相关算力对各方可见。

In terms of the cheating fear, I'm reasonably confident. I think that if you do the things we described and allow inspectors to come in and put hands on GPUs, basically count them and things like that, then you can get something like 99% of the world's AI-relevant compute visible to all sides in the world.

Host

嗯。

Yeah.

Daniel

然后如果你做到完全的研究透明,他们不仅能看见算力在哪里,还能看见上面在发生什么。然后我认为你已经在防止作弊方面取得了很大进展,因为你可以实时看到他们在数据中心里做什么。是的,可能有些隐藏的数据中心,他们藏在山里或军事设施里之类的。但他们在那些数据中心上的进展会慢得多,因为他们的算力非常有限。AI 进展高度依赖算力,我们对此相当确信。

And then if you do total research transparency, then they can not just see where the compute is but what's happening on it as well. And then I think you've gone a long way towards preventing cheating because you can just literally see what they're doing in the data centers in real time. And yeah, there might be some hidden data centers somewhere that they've squirreled away in a mountain or a military facility or whatever. But they're making much slower progress on those data centers because they have such limited amounts of compute. AI progress is heavily dependent on compute. We are pretty confident of that.

Host

是的,我们不确定依赖的程度到底如何。在我们的模型里有一个参数对应这个。

Yeah, we're not sure exactly the degree to which it depends. There's a parameter for this in our model.

Daniel

不错。但对于该参数的合理设置,你基本上可以在几年左右,甚至可能几十年内不用担心秘密项目。你确实需要担心他们从合法项目中窃取权重。

Nice. But for plausible settings of that parameter, you can basically not have to worry about the covert projects for a couple of years or so, or possibly a couple of decades. You do have to worry about them stealing the weights from the legal projects.

Host

是的,你只能承担那个成本,因为你对研究透明,他们就会复制你所有最新的研究。

Yeah, and you just have to eat the cost that because you're being transparent about the research, they'll just be copying all your latest research.

Daniel

是的。秘密项目。基本上,一种说法是,如果美国、中国和其他国家合作建立这种透明系统,那么透明系统中的国家和公司就能以相当舒适的优势超越秘密项目。所以也许你不能暂停 30 年,但你可以暂停十年或五年,然后才开始严重担心秘密项目。而且你可以知道透明项目中领先的 AI 开发不会被秘密颠覆去服务某个独裁者或将军之类的。不像秘密项目,你无法知道那些。这也是为什么透明项目在这个世界上保持领先如此重要的原因之一。

Yeah. The secret projects. Basically, one way of putting it is that if the US and China and those other countries cooperate to have this sort of transparent system, then the countries and companies in the transparent system can outrace the secret projects with a pretty comfortable margin. So maybe you can't pause for 30 years, but you can pause for ten years, or five years, before you start to be seriously concerned about a secret project. And you can know that the AI development happening in the transparent project, which is ahead, won't be secretly subverted to serve some dictator or general or something. Unlike the secret projects where you don't know that. And that's one of the reasons why it's so important that the transparent projects maintain that lead in this world.

Host

是这样吗?

Is that right?

Daniel

是的。或者我会说,透明项目对秘密项目有巨大领先其实并不那么重要。更重要的是秘密项目不要赶超。这是一个相对较小的差异,因为一旦他们赶上,他们也能超越。但我认为如果他们达到相同的能力水平其实也还好,因为他们的 AI 数量少得多,与主流相比只是极小的一部分。所以这有点像好事,因为他们的推理算力少得多。

Yeah. Or I'd say it's actually not that important that the transparent projects have a big lead over the covert projects. It's more like it's important that the covert projects not pull ahead. And this is a relatively small difference because once they catch up, they can also pull ahead. But I think it would actually be kind of okay if they reached the same level of capability because they just have so many fewer AIs, such a tiny amount compared to the mainstream. So it'd be sort of like a good thing because they have way less inference compute.

Host

是的。我们只是说总体算力。就像现在世界上有很多坏人,但好人比他们多,所以总体还好。类似地,如果有一些为独裁者工作的坏 AI,那很糟糕,但不如他们按智能加权后占世界 AI 的大多数那么糟。所以你真正担心的是那些 AI 变得比所有好 AI 聪明得多,然后按智能加权的大多数 AI 都为单一目的工作。

Yeah. We're just compute in general. Like right now there are lots of bad people in the world, but the good people outnumber them, and overall things are fine. Similarly, if there are some bad AIs working for a dictator, that's bad, but it's not as bad as if they were a majority of the world's AIs weighted by intelligence. So what you're really worried about is those AIs getting substantially smarter than all the good AIs, and then a majority of the AIs weighted by intelligence are working for a single purpose.

Daniel

总之。所以我会说,你试图获得大部分算力,然后你能做这件事,然后你能承受更慢地前进并暂停,因为即使你没得到的算力都秘密存在于某个项目中,它也无法像你一样快。这就是高层次的想法。

Anyhow. So I would say you try to get most of the compute, and then you can do this thing, and then you can afford to go more slowly and pause, because even if all the compute you didn't get was secretly in a project somewhere, it wouldn't be able to go as fast as you. That's the high-level idea.

Host

好吧。我相当信服。我想我对秘密项目的主要担忧是蒸馏那件事。

Okay. I'm pretty convinced. I think my main worry about the covert projects is the distillation thing.

Daniel

是的。这和其他担忧有点相关。我认为如果合法项目走得太快,那么 a) 他们可能因为太快而失去控制,b) 他们可能通过研究和蒸馏在各方面过度帮助秘密项目。所以我绝对建议人们宁可偏慢也不要偏快。

Yeah. It's kind of related to the other worry. I think if the legal projects go too fast, then a) they could just lose control because they went too fast, and b) they could be helping out the covert projects too much in various ways through the research and through the distillation. So I would definitely recommend that people err on the side of going too slow rather than err on the side of going too fast.

Host

是的,有道理。好吧,我觉得那已经是很不错的高层次概述了。我的意思是,就我对 A 计划的理解,你觉得我们有没有遗漏什么你想涵盖的?无论是 A 计划的内容还是……

Yeah, that makes sense. Well, I think that's a pretty good high level. I mean, insofar as I understand Plan A, is there anything you feel like we missed that you want to cover? Either Plan A stuff or...

Daniel

我的意思是,我们可以谈谈 A 计划的后果,比如当人们失业时社会会发生什么之类的。

I mean, we could talk about the consequences of Plan A, like what would happen to society as people lose their jobs and things like that.

Host

是啊,对于人们失业有什么计划?

Yeah, what's the plan for people losing their jobs?

Daniel

那似乎没问题。基本上,你可以做的一件事是全民基本收入(UBI),即对公司征税并重新分配资金。我们认为公民分红稍好一些。这是一种不同的东西,更像人们持有公司股票,他们拥有某些东西,而不仅仅是获得付款。而且我们也认为那是一种改进。而且它不完全是针对公司的税。这是一种总量管制与交易制度,对于芯片和机器人,你需要许可证才能生产。许可证由特别设立的政府机构出售,然后人们基本上持有这些机构的股票。

That seems okay. Basically, one thing you could do is a UBI where you tax the companies and redistribute the money. We think a citizens dividend is slightly better. It's a different thing where it's more like people having stock in the companies, where they own something rather than just getting payments. And also we think that's an improvement. And it's not a tax on the companies exactly. It's a sort of cap-and-trade thing where for chips and robots you need a permit to produce them. And the permits are sold by specially set up government institutions, and then people have stock in those institutions basically.

Host

为什么它比类似 UBI 的东西好这么多?

Why is it so much better than a UBI-like thing?

Daniel

关于为什么我喜欢人们持有股票而不是 UBI,有一点是它在政治上感觉更难改变。这感觉有点像是一种感觉上的东西。

So one thing about why I like people owning stock instead of UBI is that it feels politically harder to change. It feels like it's kind of a vibes thing.

公民红利与总量控制 Citizens Dividend and Cap-and-Trade

Daniel

但如果你真的持有一些股份,那么如果政府决定改变规则,就像是在偷他们的东西。而如果是政府出于好意给你一些钱,他们将来可以决定少给你一些。

But if you actually have some shares, then if the government decides to change things, it's like they're stealing their stuff. Whereas if it's the government out of the goodness of its heart giving you some money, they can just decide to give you less money in the future.

Host

我明白了。

I see.

Daniel

然后,你知道,这有点像框架问题。但我确实认为你应该努力让它难以撤销,难以收回。

And then, you know, so it's kind of a framing thing. But I do think you should try to make it hard to undo, hard to take back.

Host

是的。

Yeah.

Daniel

然后关于总量管制与交易,它有着很好的自由意志主义直觉。我认为它有时比征税更好。而且我认为它也让政府能够对增长设定合理的限制。假设他们不确定机器人经济能否在 18 个月或 6 个月内翻倍,他们可以设定一个门槛,比如今年我们会有多少机器人,然后出售相应数量的许可证,让市场决定谁购买,基本上让价格体系决定谁买许可证以及他们用机器人做什么。

And then the other thing with cap and trade, cap and trade just has nice libertarian intuitions. I think it's sometimes better than taxing. And I think it also allows the governments to have these nice limits on growth. Say they're uncertain about whether the robot economy is going to be able to double in 18 months or six months or whatever. They can just set a threshold, like how many robots we're going to have this year, and then they can sell that many permits and let the market decide who buys, basically let the price system decide who buys the permit and what they do with the robots and so forth.

Host

是的。

Yeah.

Daniel

然后他们可以利用这个门槛,基本上将翻倍时间保持在任何他们想要的水平。

And then they can use that threshold to basically keep the doubling time at whatever level they want.

Host

是的,比如我们每年翻倍。

Yeah, we're going to double every year, for example.

Daniel

每年或每两个月,而不是尽可能快地翻倍。

Every year or two months or something, instead of doubling as fast as we can.

Host

是的,是的。

Yeah, yeah.

Daniel

而且这也意味着,随着经济增长和 AI 变得越来越强大,有效税的规模基本上会成比例增长。根据我们为这部分情景建立的小型经济学模型,如果我没记错的话,起初公民分红相当小,但到最后会变得巨大。这不仅是因为公司变得庞大,更是因为 AI 行业变得庞大。而且由于总量限制,每个机器人的价格大部分都用于支付许可证,因为它们太有价值了。

And it also just sort of smoothly means that as the economy grows and the AIs become more and more powerful, the size of the effective tax basically grows in proportion to that. According to our little economics model for this part of our scenario, if I recall correctly, at first the citizens dividend is fairly small, but by the end it's huge. And it's not just that companies have gotten huge; it's that the AI industry got huge. And also because of the cap, most of the price of each robot is just paying for the permit because they're so valuable.

Host

是的,是的。

Yeah, yeah.

Daniel

所以这意味着,从某种意义上说,开始时对 AI 公司的税较低,但随着技术变得更好,税收会自动成为他们收入中更大的部分,以一种平滑的方式。

So that means that in some sense, it's like starting off with a lower tax on the AI companies, but then the tax automatically becomes a bigger fraction of them as their technology gets better, in a nice smooth way.

对详细规划的赞赏 Appreciation for Detailed Planning

Host

所以这些细节并不是非常必要。这就是我的意思,你们真的考虑了很多细节,考虑了很多不同的方式。我也很欣赏你们认真对待未来。你们说,是的,我预计这些事情会发生。所以如果它们发生,就会产生影响。然后这会改变事情。我们想要如何激励好的事情发生,而不是坏的事情?说这句话很容易,说“我想激励好的事情”很容易,但更难的是说“那么机制设计是什么?政策应该是什么?”我希望世界其他地方能赶上你们,利用你们的帮助,并利用 Plan A 的见解。正如你所说,这是我们能做的最好的。也许有更好的东西。

So these details aren't super necessary. This is what I mean when I'm like, you guys have really thought through a lot of details, you've thought through a lot of different ways. And I also just appreciate that you're taking the future very seriously. You're like, yeah, I expect these things to happen. And so if they happen, they will have effects. And then this will change things. How would we want to incentivize the good things to happen and not the bad things? It's easy to say that, easy to be like, I would like to incentivize good things, and much harder to be like, well, what is the mechanism design? What should the policy be? And I hope the rest of the world can catch up with you guys and use your help and harness the insights from Plan A. As you say, here's the best we can do. Maybe there's better things.

Daniel

请便。这只是我们的第一次尝试。我们花了大约一年时间。我相信如果有更多人认真思考更长时间,他们能想出更好的计划。我希望他们能。

Please. Like this is just our first attempt. It took us about a year. I'm sure that if more people think more seriously for more years, they can come up with better plans. And I hope they do.

Host

是的,是的。我也是。而且感谢你们先做了这件事。

Yeah, yeah. Me too. And thanks for doing it first.

B计划:关闭作为后备 Plan S: Shutdown as Fallback

Daniel

如果其他一切都失败了,当有疑问时,就全部关闭。我们有这个复杂的计划,我们认为是最好的,是目前最不坏的方案,但如果它看起来失败或不起作用,那么……

If all else fails, when in doubt, just shut it all down. We have this complicated plan that we think is best, our current least bad plan, but if it looks like it's failing or not working, then...

Host

而且你们计划的第一步是,让我们在这里暂停能力发展。让我们在解决问题时继续提供推理服务。我认为这是很棒的第一步。我们可以从那里决定实施 Plan S 或 Plan A 或其他计划。但那会是一个好的起点。

And the first step of your plan is like, let's pause capabilities right here. Let's keep serving inference while we sort it out. And I think that's a great first step. And we can decide from there to do Plan S or to do Plan A or to do whatever, some other plan. But that would be a good starting point.

Daniel

哦,是的。有趣的是,在我们的 Plan S 情景中,我提出了这一点:即使你完全停止所有 AI 研究,只允许现有模型,但不允许训练新模型,世界在 10 或 20 年后仍然会看起来极其赛博朋克。如果你感兴趣,我很乐意深入探讨。

Oh yeah. Fun fact, in our Plan S scenario, I made this point of how even if you completely halted all AI research and only allowed existing models, but no new models to be trained, the world would still look extremely cyberpunk 10 or 20 years later. Happy to get more into that if you're interested.

Host

嗯,我的意思是,Plan S 有不同的版本。一个版本是砸碎所有计算机,比如到处都没有 AI 了,现有的 AI 被追捕并摧毁。那不是我们在情景中描绘的版本。它只是完全停止进一步的 AI 研究。所以没有新的 AI 被创造,但现有的被保留下来。所以人们可以继续使用像 Claude Mythos 5 这样的东西。

Well, like, I mean, there's different versions of Plan S. One version is where you smash all the computers, like there's no more AIs anywhere, existing AIs hunted down and destroyed. That's not the version that we sketched in our scenario. It's just a total halt to further AI research. So there's no new AIs being created, but the existing ones are grandfathered in. So people can continue using like Claude Mythos 5 or whatever.

Daniel

是的,但关键是 Claude Mythos 5 相当强大。

Yeah, but the point is that Claude Mythos 5 is pretty powerful.

Host

确实如此。

It sure is.

Daniel

所以如果你想象 20 年里 Nvidia 生产越来越多的 GPU,并在上面运行 Claude 5,并将它们集成到每个人的笔记本电脑中,那实际上至少会是一场互联网规模的变革。

And so if you imagine just 20 years of Nvidia producing more and more GPUs and running Claude 5 on them and integrating them into everybody's laptops, it would actually be at least an internet-scale transformation.

Host

是的。不。完全同意。所以即使我们从这里完全成功地暂停了所有 AI 进展,它实际上也会改变一切,赛博朋克。

Yeah. No. Totally. So it would actually just change everything, cyberpunk, even if we just completely successfully paused all AI progress from here.

结束语与工作寻找 Closing and Where to Find Work

Host

嗯,Daniel,非常感谢你来到播客。如果人们想找到你的作品,他们该怎么做?

Well, Daniel, thanks so much for coming on the podcast. If people want to find your work, how do they do that?

Daniel

是的,AI-2040。哦,我想我们被打断了。我们有一些甜甜圈。

Yeah, AI-2040. Oh, I think we're being interrupted. We got some donuts.

Host

谢谢你,Daniel。谢谢你,Claude。我们有巧克力味的吗?有巧克力味的吗?我看到一个巧克力味的。未来就在这里。

Thank you, Daniel. Thank you, Claude. Did we get any chocolate ones? Do we have chocolate ones? I see a chocolate one. The future is here.

Daniel

对。所以,AI-2040.com。那是我们的新情景。那是积极的愿景。AI-2027.com 是我们更……你知道,我们的预测。

Right. So, AI-2040.com. That's our new scenario. That's the positive vision. AI-2027.com is our more, you know, our prediction.

互动版:逐字朗读 + 针对本期提问 →