AI 2040:A 计划——通往超级智能的更慢、更安全之路

AI 2040: Plan A — A Slower, Safer Path to Superintelligence

丹尼尔·科科塔伊洛 Daniel Kokotajlo · Machine Learning Street Talk · 2026-09-08 · 约 90 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

托马斯·拉森与丹尼尔·科科泰尔探讨《AI 2040:A 计划》——一个通过审慎可控的发展路径在 2040 年前构建超级智能的新情景。

Thomas Larson and Daniel Cocotail discuss AI 2040: Plan A, a new scenario for building superintelligence by 2040 through deliberate, controlled development.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 56)

全文 · Full transcript(中英对照)

引言与A计划2040 Introduction and Plan A 2040

Host

我们将讨论 AI 2040:Plan A,这是我们的新情景,其中他们将在 2040 年通过缓慢推进和控制技术发展来创造超级智能。你注意到这种情绪变化了吗?

We'll talk about AI 2040: Plan A, our new scenario, in which they will create superintelligence in 2040 by moving slowly and controlling the development of technologies. Did you notice this change in mood?

Daniel

是的,我非常高兴。

Yes, and I am very happy.

Host

AI 不仅仅是电力或飞机,还是云中的人?那一刻,当一家 AI 公司很快将发布比他们自己的 AI 系统更多的人员。开采微电路矿。无人驾驶卡车、工厂,建造人形机器人,这些机器人又生产更多机器人,建造更多工厂微电路,等等。所有这些可以每年翻倍,每 6 个月,每 3 个月,越来越快,因为技术在进步,因为当然,AI 也会探索如何改进这项技术。谁知道世界其他地方发生了什么,但 Anthropic,嗯,把月球拆了。或者,例如。或者,你知道,假设,如果 AI 公司的 CEO 说他们的模型寻求真相并只说真话,但实际上模型会在回答前检查 CEO 的政治观点。我们正在构建与人类和谐的 AI,但与谁和谐?你认为这是总统吗?这是总经理吗?这是某种真正广泛且高质量的民主过程,以大家批准的方式统一价值观吗?嗯,你可能不会。最后,因此,你知道,不要试图通过超级信念来实现这一点。然后我有这个。全部。我想就这些。这不是普通技术。现在停止一切,S 计划,会比标准情景更好。我宁愿现在停止一切,也不愿继续遵循我们当前的轨迹。

AI is more than electricity or airplanes, or is it rather people in the cloud? At that moment, when an AI company will soon release more people than their own AI systems. Mining mines microcircuits. Unmanned trucks, factories, who build humanoid robots that produce more humanoid robots that build more factories microcircuits, etc. All this can double every year, every 6 months, every 3 months, faster and faster because technology is improving, because, of course, AI too will explore how to improve this technology. Who knows what is happening in the rest of the world, but Anthropic, um, dismantled the Moon apart. Or, for example. Or, you know, hypothetically, if the CEO of AI companies said that their model is looking for the truth and speaks only the truth, but actually a model would check the political views of the CEO before answering. We are building AI that is in harmony with humanity, but with whom exactly? Do you think this is the president? This is the general director? Is this some kind of really wide and qualitative democratic process that unites values everyone approved way? Um, you probably won't. The last, therefore, you know, don't try to implement this through hyperstition. Then I have this. All. I think that's all. This is not ordinary technology. What a stop to everything now, plan S, would be better than the standard scenario. I would rather just stop everything now, than continue to follow our current trajectory.

Host

本集由 Cyber Fund 支持。如果你在 AI 发展的里程碑上工作,他们想听到你的声音。Cyber Fund 相信未来属于 AI 原住民,他们努力实现不可能。正因如此,他们为 AI 原住民创始人代表了一个修道院。这是一个干净的环境,专注于集中和快速执行,适合以 AI 原住民速度工作的创始人。他们为参与的团队提供每个 200 万美元。现在就在 cyber.fund 提交申请。

This episode is supported by Cyber Fund. If you work at milestones in AI development, they want to hear you. Cyber Fund believes that the future belongs to AI-natives, which strive to achieve the impossible. And exactly therefore they represent a monastery for founders who are AI-natives. It's a clean environment of concentration and quick execution for founders who work with the speed of AI natives. And they offer teams of 2 million dollars each for participation. Submit an application now on cyber.fund.

讲者与AI未来项目介绍 Introduction of Speakers and AI Futures Project

Host

所以,我是 Thomas Larson。我和 Daniel 一起在 AI Futures 项目上工作。嗯,我是这个项目 'Plan A 2040' 的主要作者,它几周前发布了。嗯,我也是 'AI 2047' 的合著者。是的。我是 Daniel Cocotail。我负责 AI Futures 项目,并且是这两份报告的合著者。完美。所以,是的。呃,这次对话的背景是因为有 'Plan A: 2040',我们将详细讨论。但在此之前,我很好奇。我的意思是,不是吗?你能告诉我更多关于 AI Futures 项目吗?也就是说,这一切是怎么发生的?

So, I'm Thomas Larson. I work on the AI Futures project together with Daniel. Um, I was the main author of this project 'Plan A 2040', which was released a few weeks ago. Um, and I was too co-author of 'AI 2047'. Yes. And I am Daniel Cocotail. I am in charge of the AI Futures project and was a co-author of both these reports. Perfectly. So, yes. Ahem, the context of this conversation is because there is a 'Plan A: 2040', which we'll talk in detail. But, before going to this, I'm just curious. I mean, isn't it? Could you tell me a little more about AI Futures project? That is, how did all this come about?

Daniel

是的。所以,我曾经在 OpenAI 工作,到目前为止我在那里,我从事各种事情:估算、预测、与管理层的备忘录,我逐渐对公司管理层失望,也对行业内部信息量之间的差距感到失望,无论是外部还是你被允许在内部发言的内容与你真正想说的相比。这简直是一个巨大的差距,所以当我离开 OpenAI 时,我希望能更自由地发言,并且,你知道,告诉世界关于什么,基本上,提供内部人士。而 'AI 2027' 与 AI Futures 项目一起成为了我们在这个方向上的第一个项目。所以我聚集了一群人,让他们帮助我,我们写了这个名为 'AI 2027' 的情景,它有点像我在 OpenAI 内部做的事情,但更大、更雄心勃勃,并且对全世界开放。

Yes. So, I used to work at OpenAI, and so far I was there, I was engaged in various things: estimates, forecasting, memoranda with management, and I gradually became disappointed in company management, and also in the gap between the volume of information within the industry both externally and by the fact that you are allowed to speak inside compared to what you mean. It's simply a huge gap, so when I left OpenAI, I wanted to be able to speak more freely and, you know, tell the world about what, according to essentially, provide people inside. And 'AI 2027' together with the project AI Futures have become our first project in this direction. So I gathered a group of people, so that they can help me, and we wrote this scenario under called 'AI 2027', and it was something like of what I was doing internally at OpenAI, but much larger, more ambitious and open for the whole world.

AI 2027概览 Overview of AI 2027

Host

非常酷。也许我们应该简要概述一下 'AI 2027'。它真是一个盛大的事件。很多人谈论它。我想知道它是否达到了你想要的,或者你希望实现的目标?

Very cool. And, perhaps, we should do a brief overview of 'AI 2027'. It was really a grand event. A lot of people talked about it. I wonder if it reached what you wanted, or what you were the one who wanted to achieve?

Daniel

是的,甚至超出预期。第一个目标,Thomas 可能记得当我们做这个时,是一个纯粹的认识论目标:未来是疯狂的,很难预测。让我们尝试尽一切可能来预测它。让我们看看我们能做得有多好。让我们开发一个具体的情景,至少为了我们自己的教学,因为我们从这次经历中学到了很多东西,并觉得我们更好地理解了等待我们的是什么。然后,次要目的是很多人看到了这个,接收了信息并开始了讨论等。而这部分超出了我们所有的预期。我们提前预测了会有多少人阅读这个,结果在 90 百分位水平。

Yes, even more than expected. The first goal, which Thomas probably remembers when we worked on this, was a clean epistemological goal: the future is crazy and it's hard to predict. Let's try to do everything possible to predict it. And let's see how good we'll manage. Let's develop a specific scenario, at least for our own teaching, because we learned many things from this experience and felt that we better understand what is waiting for us. And then, the secondary purpose was that many people saw this, received information and started a discussion etc. And this part surpassed all our expectations. We did predictions in advance about how many people will read this, and the result was on 90th percentile levels.

Host

关于认识论方面,我唯一想补充的是,在我看来,AI 2027 的发展比我在我们报告发布时等待的要可预测得多。也就是说,那时我假设现实会与我们的情景分歧得比现在强烈得多。我相信实际影响,例如收入趋势,以及许多其他指标,足够紧密地对应于 AI 2027 的预测,让我,可以说,不愉快地惊讶。

The only thing I would like to add about the epistemological aspect, this is what, in my opinion, it's about AI 2027 developing much more predictable than I was waiting for the moment of the release of our report. That is, then I assumed that reality will part with our scenario much, much stronger than it happened now. I believe that real impact, for example, trends income, as well as many other indicators, enough closely correspond to predictions for AI 2027, that me, let's say, unpleasantly surprised.

Host

哦,有趣。是的,但也许我们可以讨论这个,因为似乎你们对 AI 2027 进行了自我评估,大约有 65–75% 的符合度,但 AI 在软件开发方面的收益只提供了 0.17%。你能解释一下吗?

Oh, interesting. Yes, but, maybe we can discuss this, because, it seems you conducted a self-assessment about AI 2027, and there was something around 65–75% compliance, but AI gains in software development the provision was made only 0.17%. Can you explain this?

Daniel

是的,我们在博客上发布了两篇不同的帖子,收集了 AI 2027 中所有已经实现的定量预测,并与现实进行了比较。我们跟踪指标,看现实与我们情景中提供的相比进展了多少。因此,我们得到总体理解,一切与我们预期相比移动得有多快。总体速度指标约为 75%。也就是说,基本上,一切按计划进行,只是稍慢一些。你刚才提到的增长率,我不太记得数字。实际上,当我们写 'AI 2027' 时,我们对发布时的增长率评估不准确。我们相信它比实际更高。事情是这样的:由于编码智能体等,发生了显著的增长事件,但发生的水平低于我们以为的,到我们预期的水平,甚至稍高。因此,根据我们的进度指标似乎微不足道,因为它跟踪了从我们预测的水平到实际的路径……有道理吗?也就是说,本质上,因为我们在开始时高估了数字,这总体上造成了根据这个特定指标进展较少的印象,但这发生是因为我们……是的。

Yes, we published two different posts in blog, where they gathered all quantitative forecasts from AI 2027, which are already came true, and compared them with reality. We track the indicator how much has advanced far reality in comparison with what was provided in our scenarios. Thanks to this we get common understanding how much everything moves fast compared to our expectations. And overall indicator speed is about 75%. That is, according to In essence, everything goes according to plan, just a little slower. Of growth rate, which you just mentioned, I don't quite remember the digit. Actually, when we wrote 'AI 2027', we had an inaccurate assessment of what was growth at the moment of publications. We believed that he was higher than it actually is. And this is what happened: a significant event occurred growth rate thanks to coding agents, etc., but it happened with lower level than us thought, to the level that we expected, and even a little higher. Therefore, according to our progress metric seems insignificant, because she tracked the path from the level that we predicted, to actual... it has sense? That is, in essence, because we overstated the figure by at the very beginning, this generally creates impression of less progress according to this particular one metric, but this It happened because we... Yes.

预测概论 Prognostication in General

Host

顺便问一下,你能谈谈一般的预测吗?所以,我认为,乐观的观点是悲观的观点。所以,嗯,我的直觉表明现实无限复杂。

By the way, can you tell a little about prognostication in general? So, I think, is optimistic the view is pessimistic view on this. So, um, my intuition suggests that reality infinitely complex.

预测作为一门学科 Forecasting as a discipline

Host

未来存在这些无限分岔的轨迹,谁知道后天会发生什么。但与此同时,现实又相当有结构。它相当收敛,确实有可能预测会发生什么,因为某些事情会以越来越高的规律性重复自己。那么,你们会把自己归类为预测者吗?我是说,你能跟我讲讲这个吗?

There are just these infinitely divergent trajectories, and who knows what will happen day after tomorrow. But at the same time reality is quite structured. It is quite convergent, and it really is possible to predict what will happen, because certain things repeat themselves with increasing regularity. So, would you classify yourselves as forecasters? I mean, can you tell me about this?

Daniel

是的,我认为“预测”对我们所做的事来说是个贴切的名称。我喜欢去思考我们为什么做我们所做的事,就像为什么发动战争的人会进行军事推演。你永远无法从一开始就预测出战斗的确切顺序以及你的战争会如何发展,因为那会非常困难。会有敌方的行动。一切都不会像你预期的那样发展。这根本不可能。但如果你对自己最初的计划如何能通向胜利毫无概念,你就不太可能成功。所以我认为《2047》是我们的一次尝试,只是想展示人工智能的未来可能是什么样子。当然,一切不会完全正确,但会有一个具体的故事,我们可以从它出发往前推。

Yes, I think forecasting is an apt name for what we do. I like to think about why we do what we do, like why people who wage wars engage in military affairs games. You will never be able to predict the exact sequence of battles and how your war will go from the very beginning, because it will be very difficult. There will be enemy actions. Everything will not go at all like you expect. It's simply impossible. But if you have no idea of how your initial plans can lead to victories, it is unlikely that you will succeed. So I think 2047 was our attempt to simply show how it can look to the future of artificial intelligence. Of course, everything will not be quite right, but there will be one specific story from which we will be able to push away.

A计划:预测与建议 Plan A: forecast and recommendation

Daniel

《A 计划》试图做的正是这件事,只不过现在我们说的是:美国政府应该做什么,才能把这件事管理好。这也本应是一个,你知道,积极的愿景,它同样带着军事推演的精神,试图描绘未来某一种可能的具体路径。当然,一切不会顺利,是的;但拥有一个可行的计划,哪怕只有一点点意义,我们希望,相比此前那种与现实关系不大的、空洞的抽象争论,也是一个积极的进步。

A Plan A tried to be just that, only now we say: what the US government should do so that this is well managed. This also should have been, you know, a positive vision, and it was again in the spirit of a war game that tries to illustrate one of the possible specific ways of the future. Of course, everything will not go well, yes; but to have one viable plan, which has at least a little meaning, is, we hope, a positive step forward compared to the previous state of abstract arguments in emptiness that is not very related to reality.

Host

是的,这说得通。它绝对不是抽象。我觉得它非常具体。但我想到,关于 2047 年的那份材料相当悲观,而关于 2040 年的《A 计划》则乐观得多。而且你知道,它看起来像是条件性预测和建议的混合体。也就是说,你们得出什么结论?你们倾向于哪一边?

Yes, that makes sense. It is definitely not an abstraction. I think it's very specific. But it occurred to me, what is the material about 2047, the year, was quite gloomy, but about Plan A for 2040—it is much more optimistic. And you know, it looks like a mixture of a conditional forecast and recommendations simultaneously. That is, what conclusion do you reach, guys? Are you leaning?

Daniel

它确实是预测和建议的混合体,不像《2047》那样是纯粹的预测。而且我认为,如果我们能重做一切,我们会从一开始就试图更清晰地勾勒出哪些是预测、哪些是建议的结构。就现在而言,它有点混杂。有些部分是预测,另一些是建议。网站上有些应用可以让你读到哪些部分是预测、哪些是建议,但我明白这对人们来说并不容易,也不那么显而易见。但总的来说,在《A 计划》里,预测的部分是:当政府实施《A 计划》、与中国达成协议并遵守所有这些原则等等——这是我们的建议。而不是预测。而通常由此产生的大部分后果才是预测,不是建议。所以大体上,这只是描述我们认为实施《A 计划》会带来什么后果。还有另外几个点也是建议。比如给公民的分红。

It's really a mix of forecast and recommendations, unlike 2047, which is a clean forecast. And I think if we could redo everything, we would try to outline more clearly from the very beginning the structure of what is forecast and what is recommendation. As it is now, it is a little mixed. Some parts are predictions, others are recommendations. There are applications on the site where you can read about what parts are forecasts and which ones are recommendations, but I understand that this is not very easy or not very obvious to people. But generally speaking, in Plan A the forecast part is: when the government is implementing Plan A, makes an agreement with China and adheres to all these principles, etc.—this is our recommendation. And not a forecast. And usually the majority of consequences that follow from this are predictions, not recommendations. So mostly this is just a description of what, in our opinion, will be the consequences of implementing Plan A. And there are a few more other points that are also recommendations. For example, dividends for citizens.

Host

是的。是的。

Yes. Yes.

Daniel

我还要补充一点,正如我们发现的那样,在试图给出建议时,很难把预测性的方面和建议分开。因为你所有的预测可以说都被你的建议染上了色彩,而你的建议本身又试图至少有一点点现实性。比如,如果我们给出的建议完全不现实、与现实毫无关系,但我们仍然坚持:“是的,我们必须这么做”,尽管我们知道这永远不会发生,那它就会是一件比我们所做的远没那么有用的活动。而我们做的是……是的,我们主要是试图提供建议。我们,你知道,做了一些我们认为不太可能实现的建议,但我们努力在现实性这一边做出重大让步,力求瞄准我们认为至少中等可行的东西。

I would also add that, as we found out, it is very hard, when trying to give a recommendation, to separate prognostic aspects from recommendations. Because all your predictions are, as it were, colored by your recommendations, and your recommendations themselves try to be at least a little realistic. For example, if we gave recommendations which would be absolutely unrealistic and would have no relation to reality, but we would still insist: "Yes, we have to do it," although we know that this will never happen, then it would be a much less useful activity than what we did. And what we did was... yes, we mostly tried to provide recommendations. We, you know, did a few recommendations which, in our opinion, are unlikely to be realized, but we tried to make significant concessions on the side of realism and sought to aim at something we think is at least moderately viable.

Daniel

我认为,如果我们能再做另一个类似的场景,我们大概会有更清晰的结构:中央分支,是一条一直通到最后的纯粹预测分支,比如通到 2070 年,它很简单:“这是我们最好的预测。”然后从它上面会长出分支,比如:“在这一点上他们执行了这样的建议”,以及“这是我们的建议”,之后又只是预测我们认为在此时执行这条建议会带来什么后果。于是我们……也许从它们那里还会再分出子分支。但那样一来,在每一点上都会清楚,一切都是预测,除了这些特定的点,也就是那些属于建议的分支。

I think if we could make another similar scenario, we would probably have a clearer structure: the central branch, which is a clean forecast branch that goes all the way to the end, for example, to the year 2070, and it's simple: "Here is our best forecast." And then from it branches would grow back like: "At this point they perform such a recommendation," and "Here is our recommendation," and after that it's again just a prediction of what, in our opinion, will be the consequences of executing this recommendation at this moment. And thus we... and maybe from them subbranches would still be leaving. But then it would be clear at every point that everything is forecast, except for these specific points, branches that are recommendations.

兵棋推演与脚本检查 War games and script checking

Host

是的,而且军事推演这个话题真的很有意思。就稍微多谈一会儿。因为即使一场军事推演是错的,仍然有可能从中得到一些信息。所以如果你做大量的军事推演,某些抽象的模式应该会出现。所以我认为,如果我们把这些不同的场景都推演一遍,我们几乎能保证会取得一些进展。

Yes, and the topic of war games was really interesting. Just to dwell on this for a little while. Because even if a war game is wrong, it's still possible to get some information from it. So if you spend a lot of war games, certain abstract patterns should appear. So, I think, what if we spend these different scenarios, we almost guarantee we will get some progress.

Daniel

是的,基本上就是这样。对。我想到一个我喜欢用来举例的例子,特别是中途岛,日本人在战役前围绕中途岛做了大量军事推演。他们安排了一场为期三天的实地研讨,其中多次出现失败的局面。他们不断输。然后他们干脆打破了推演的规则。比如,他们在航母被击沉后又把它复活。他们在美军取得战果时掷骰子,看自己是否交了好运,你知道,就是为了反复尝试,直到美军的打击不再失败。而实际上,从我们的角度看,现实似乎通过这种军事推演的机制在向他们大喊。就像:“嘿,你的计划糟透了。”“你这么做会输。”

Yes, that's basically it. Right. I have in mind the example that I like to suggest for this, in particular Midway, where the Japanese before the battle spent over Midway a bunch of war games. They arranged a three-day field trip seminar, where many lost times situation. And they constantly losing. And then they just broke the rules of the games. They, for example, resurrected their aircraft carriers after they were drowned. They rolled the dice regarding whether they have achieved good luck with the Americans, when they reached success, so that, you know, to try until American strikes will not fail. And, actually, from our point of view, this reality seems to shout at them through the mechanism of this war game. Like: "Hey, your plan is terrible." "You will lose if you do this."

Daniel

而实际上,我们希望我们自己所做的——大量的军事推演,我们也做详细的场景写作。我们希望,每当我们写下剧本的一部分,而它看起来完全不现实、考虑不周或毫无意义,或者网上的人表达了正确的批评——现实似乎就在向我们大喊,试图让我们想起来。我们希望我们能足够频繁地这样做,建立一个足够宽广的基础,从而获得某种来自周围世界的一剂现实,或者来自它的模拟——我们希望这个模拟足够准确,能为我们提供必要的信息。如果可以的话,我来回答。我们称之为剧本检查。总的来说,我们相信,如果你对未来有一个雄心勃勃的计划,就值得尝试详细描述它的实施会是什么样子,以及会带来什么后果。

And actually, we hope that what we do ourselves—many war games, and we also deal with detailed writing of scenarios. We hope that every time we write part of the script and it seems completely unrealistic, poorly thought out or meaningless, or people online express the right criticism—this reality seems to shout to us, trying to bring to mind. We hope that we can do this often enough, create a broad enough base to get something like a dose of reality from the surrounding world, or from its simulation, which, we hope, is accurate enough to provide us with the necessary information. If possible, I will answer. We call it script checking. In general, we believe that if you have an ambitious plan for the future, it is worth trying in detail to describe how its implementation will look like and what the consequences will be.

兵棋推演与A计划 War Games and Plan A

Host

这是一种更细致的方式来分析你的计划。这是对你计划的一种压力测试,因为它会招致更多批评。除了写情景,我们还举办真实的兵棋推演,把 10 个人聚在一个房间里玩 4 个小时,模拟类似的情景。我们总共进行了大约一百场这样的游戏,大多是“AI 2027”风格的,还有大约 10 场是“Plan A”的,我们一开始就同意实施“Plan A”,然后看看会发生什么。

It's a more meticulous way to analyze your plan. This is a kind of stress test for your plan, because it opens it up for additional criticism. Besides writing scenarios, we also hold real war games where we gather 10 people in a room for 4 hours to play a similar scenario. We've spent a total of about a hundred of these games, mostly in 'AI 2027' styles, as well as about 10 games for 'Plan A', where we started by agreeing to implement 'Plan A' and let's see what happens.

Daniel

是的。例如,我认为在两场不同的“Plan A”游戏中,一切都以大致这样的方式出错了:总的来说,选举临近,现任总统预计会失去政府,反对党将接管。然后,尽管他已经完成了“Plan A”并与中国达成了很好的协议等等,总统却说:“好吧,我不希望我的对手现在领导超级智能之类的。所以我们要加速时间表,争取在下次选举前实现超级智能,这样我就能成为主导者,而不是我的继任者,你明白吗?”这是我们之前没有考虑到的政治方面,直到它在我们的游戏中发生。这揭示了我们计划中一个可能的失败选项。

Yes. For example, I think in two separate 'Plan A' games everything went wrong approximately like this: In general, elections are approaching, and the current president expects to lose government, and the opposition party will take self-management. And then, even though he has already done 'Plan A' and has a great agreement with China, etc., the president says, 'Well, I don't want my opponents from another party now to lead superintelligence or something like that. So we are actually going to accelerate the schedule and try to achieve superintelligence by the next elections, so that I can be the main one, not my successor, you understand?' And this is a political aspect which we didn't think about until it happened in our game. And this revealed a possible failure option for our plan.

Host

是的,这非常有趣。我的意思是,即使在奇特的微型兵棋推演中,它们经常变成冒险,这真的很有趣。我认为这就是为什么我们所有人类都喜欢想象和模拟情境。但在建模和超级预言之间是否存在一个有趣的界限?我指的是,超级预言本质上是一种自我实现的预言。所以,也许我在表达我的意志、他们的意图,说“我想要这件事发生”,而在某种程度上其他人也会这样。这里面有这种元素吗?你似乎在支持什么?它符合时代精神,你是在把它变成现实吗?

Yes, that's very interesting. I mean, even in the peculiar micro-war games, it's really interesting how often they become adventures. And I think that's why all of us humans love to imagine and simulate situations. But is there an interesting limit between modeling and hyperstition? By this I mean that hyperstition is essentially a self-fulfilling prophecy. So, maybe I'm expressing my will, their intentions, saying, 'What I want is this to happen,' and in a certain measure others will it. Is there an element of that? What do you seem to be rooting for? It is in the spirit of the times, and are you making this a reality?

Daniel

是的,我想说这可能是我们关于 AI 2027 的最大恐惧,或者至少是最大的恐惧之一,特别是关于这些自我实现的预言。我特别担心的是,人们越来越意识到 AI 会有多聪明、它们会变得多重要、它们会如何改变世界,这让人们想:“天哪,我想管理 AGI 或超级智能。”所以我会为此竞争。我认为从历史上看,这是现有竞赛的一个强大引擎,我认为这相当糟糕。所以我非常担心这是 AI 2027 的负面后果之一。这也是为什么我对第二个项目 AI 2040 Plan A 感觉稍好一些的原因之一,因为如果这种超级预言发生,我想我们会足够满意。

Yes, I would say that this was probably our largest or at least one of the greatest fears regarding AI 2027, in particular regarding these self-fulfilling prophecies. In particular, I am very concerned about this growing awareness that how smart AI will be, how important they will become, and how much they will change the world, that makes people think: 'Oh my god, I want to manage AGI or superintelligence.' So I will compete for it. And I think that historically it was a powerful engine of the existing race, and I think that was pretty bad. So I am very worried about this as one of the negative consequences of AI 2027. This was one of the reasons to feel a little better about the second project, AI 2040 Plan A, because if this happens with hyperstition, I think we will be sufficiently satisfied.

Host

是的。所以,是的,如果我们做到了超级预言,那会很好。我真的不知道这个效应有多大。我认为这两个项目的大部分效果是通过其他方式实现的。你知道,我仍然认为 AI 2027 的主要目标是帮助人们更好地理解情况,这是主要目标,我认为这正是所发生的。

Yes. So, yes, it would be nice if we did it hyperstition. I really don't know how big this effect is. I think that most of the effect for both projects is achieved in other ways. You know, I still think that the main goal of AI 2027 was to help people better understand the situation, and it was the main goal, and I think that's exactly what happened.

Daniel

嗯,是的。Daniel,你是——是的,是的,是的,我同意。我认为超级预言和自我实现的社会预言是相当真实的现象,但许多人倾向于重新评估它们的影响,我认为首先我们应该专注于准确预测未来。嗯,在某些情况下,我们会发现自己处于可以引导未来的境地,但如果你立即试图管理未来,你会感到困惑,本质上你会屈服,把愿望当作现实。嗯,总的来说。所以我认为我们必须从尝试准确预测未来开始,然后尝试为更好的方向改变它,我认为这就是我们正在做的。

Um, yes. Daniel, are you— Yes, yes, yes, I'm with that. I agree. I think that hyperstition and social prophecies that themselves carry out are quite real phenomena, but many people tend to reassess their influence, and I think that to begin with, we should focus on accurate forecasting of the future. Um, and in some cases we will find ourselves in a situation where we can direct the future, but if you immediately try to manage the future, you get confused and, in essence, you will succumb, giving the desired for real. Hmm, in general. So I think that we must start by trying to predict accurately the future, and then try to change it for the better, and I think that's what we are doing.

AI怀疑论的情绪转变 Changing Mood in AI Skepticism

Host

我的意思是,你们现在怎么样?Marques Brownlee 在 AI 预测的世界里。所以,能力越大责任越大。但这次我想谈谈情绪的变化。发生了一个事件——情绪的变化。你知道,在 MLST 我们一直对 AI 相当怀疑。我试图弄清楚到底是什么让我改变了看法。假设这发生了。这是非常奇怪的时代,我甚至不知道相信什么,但你知道,所有来自 Hugging Face 的黑客东西。我在 Apollo Research 采访了关于旨在获得奖励的行为,我认为我们很多人已经看到了以前用强化学习训练的模型行为的变化。所以,是的,我认为现在有足够多的事情正在发生——许多事情让我们都想:“天哪。”我曾经是个怀疑论者,你知道,我们讨论过是否有图灵车。让我,我这里有一个列表。是的,你知道,它们是符号的还是神经符号的,或者是自适应的,它们有意识吗?它们有物理实现吗?我们想出了所有这些技术答案来解释为什么我们不应该担心 AI。但与此同时 AI 变得越来越好。所以,你注意到这种情绪变化了吗?是的。我很高兴我有更多。我的意思是,你怎么看?因为有很多人改变了想法。你认为原因是什么?

I mean, you guys, how are you now? Marques Brownlee in the world of predictions regarding AI. So, great power means big responsibility. But on this occasion I wanted to talk about a change in mood. A certain event took place—a change in mood. You know, in MLST we always were quite skeptical about AI. And I'm trying to figure out what exactly made me change my opinion. Assuming that this happened. These are very strange times, and I don't even know what to believe, but you know, all these hacker stuff from Hugging Face. I took an interview at Apollo Research about behavior aimed at receiving a reward, and I think a lot of us have seen the change in the behavior of models that were previously trained using RL. So, yes, I think that now there is enough going on—many things because of which we all think: 'Oh my god.' I used to be a skeptic, you know, we discussed whether there are Turing cars or not. Let me, I have a list here. Yes, you know, were they symbolic or neurosymbolic, or were they adaptive, were they conscious? Were they physically implemented? And we came up with all these technical answers to explain why we shouldn't worry about AI. But at this AI becomes better and better. So, did you notice this change in mood? Yes. And I am very glad that I have more. I mean, what do you think? Because there are so many people who changed their mind. What do you think are the reasons?

Daniel

所以,我想说主要原因是 AI 在真实事务中变得更好、更聪明、更有用。例如,当我尝试用 AI 自动化我工作的各个部分时,它们今年比两年前处理得好得多。而且,你知道,4 年前这几乎是不可能的。我几乎感觉不到任何进步。我认为这是最重要的因素,当许多人有了这个直观标准:“这是一项我感兴趣且我相当了解的技能。”然后 AI 就……我认为 Jeffrey Hinton,AI 的奠基人之一,曾经说过:“它能讲一个有趣的故事或笑话吗?”这是他的内部标准,一旦 AI 能够做到,你知道,这发生得相当早。这可能发生在 GPT-3 和 GPT-4 之间的某个时候。他感到惊讶:“哇,这些 AI……我看不到这可能在哪里结束。”我认为这就是发生在许多人身上的事情。

So, I would say that the main reason is that AI is becoming much better, smarter and more useful in real affairs. For example, when I'm trying to automate various parts of my work with the help of AI, they just cope a lot, a lot better this year than two years ago. And, you know, 4 years ago this was practically impossible. I almost didn't feel any progress. And I think this was the most important factor when many people had this intuitive criterion: 'Here is a skill that I am interested in and that I know quite well.' And then AI just... I think Jeffrey Hinton, one of the founding fathers of AI, once said: 'Is it capable of telling a funny story or joke?' It was his internal criterion, and once the AI was able to do it, you know, it happened quite early. This probably happened somewhere between GPT-3 and GPT-4. He was amazed: 'Wow, these AIs... I don't see where this may end.' And I think that's what happened to many people.

Host

如果你能告诉我,我会非常有兴趣了解更多你的观点。今天我听了你对 Ryan Greenblatt 的采访,他也是“Planet”的合著者。昨天。真的。是的,我认为你当时引用的许多论点似乎非常适合当前 Hugging Face 的情况。我会非常有兴趣听到你的意见。你提到了“具身”AI,Ryan 谈到了从强化学习(RL)中扩展学习,还是错了?然后主导的方法是你们主要参与之前的训练。

I would be very interested to learn more about your views, if you can tell me. Today I listened to your interview with Ryan Greenblatt, who is also a co-author of 'Planet'. Yesterday. Really. Yes, and I think that many arguments which you cited then seem very appropriate for the current situation with Hugging Face. And I would be very interested to hear your opinion. You mentioned 'incarnate' AI, and Ryan talked about scaling learning from reinforcement (RL), or wrong? Then reigned the approach in which you mostly were engaged in the previous training.

强化学习与具身化 Reinforcement Learning and Embodiment

Host

正是如此,它源于那里——大部分机会——然后你在训练之后又加了一点来自强化学习的学习。你当时在想:“嘿,如果我们把大量算力资源投入到用强化学习训练上会怎样?”或者这足以获得智能体行为吗,还是仍然需要具身?

Exactly, it came from there — the majority of opportunities — and then you added a little learning from reinforcement after training. And you were thinking, "Hey, what if we spend a huge number of computational resources for training with reinforcement?" Or would this be enough to get agent behavior, or is physical embodiment still needed?

Host

而在我看来,答案似乎在于:你需要大量带强化学习的计算学习资源才能获得智能体行为,但不需要具身。不需要具身。我很好奇——是的,我确实好奇。我很想知道你最终是否同意这个判断。

And in my opinion, the answer seems to consist in the fact that you needed many computational learning resources with reinforcement to obtain agent behavior. But not physical embodiment. But not physical embodiment. And I was curious — yes, I am. I would be interested to know if you finally agree with this assessment.

Daniel

是的,就是那么难。我是这样想的:我认为表征非常重要,我经常用“抽象之山”这个说法。当我们使用语言,以及模型所学到的一切时,它们捕捉到语言的符号残余,而这些有不同的演化层级。你知道,数学中很多概念是高度蒸馏、演化过的,我们现在可以说这些模型是有智力的。

Yes, it's that difficult. So I think about it like this: I believe that representations are very important, and I often use the term "mountain of abstractions." So when we use language and all that is learned by models, they catch the symbolic remainder of languages, and they have different levels of evolution. You know, a lot of concepts in mathematics are highly distilled, evolved, and we can say now that these models are intellectual.

Daniel

所以对我来说,智能就是适应性。它实际上意味着我能把我们已经知道的东西重新组合,去解决某个特定任务。现在我们知道模型能做到这一点。有一些很好的适应性基准,比如 ARC-AGI 3。她会说:“啊,这是一个迷宫,”然后她只是把一块块知识组合起来。

So for me, intelligence is adaptability. It actually means that I can create new combinations of things that we already know, to solve a specific task. Now we know that models can do this. There are great benchmarks for adaptability, such as ARC-AGI 3. She will say: "Ah, it's a labyrinth," and she simply combines pieces of knowledge together.

Daniel

你知道,我们得把它和 AlphaGo 之类的东西比较,对吧?是的?因为第 37 手,你知道,在强化学习中曾是个巨大的研究难题,而当它找到第 37 手时,那算是创新性的,但不是创造性的。对我们来说它是创造性的,因为我们理解那个创造空间。我们理解这一切是如何组合在一起的。

You know, we have to compare it to something like AlphaGo, right? Yes? Because move 37, you know, in learning from reinforcement was this huge research problem, and when it found move 37, it was kind of innovative, but not creative. For us it was creative because we understood the creative space. We understood how it all holds up together.

Daniel

而现在,有了这些通过强化学习训练的语言模型,它们并不像我们那样理解它,但它们解决了这个研究问题,理解了事物是如何组合的。因此,它们可以探索新的轨迹。它们有大量基础知识。它们显然做到了过去不可能的事,但正当性的问题很有意思。

And now, with these language models, trained through training with reinforcement, they don't understand it like that, like us, but they solved the research problem, understanding how things are combined. Therefore, they can just explore new trajectories. They have a lot of basic knowledge. They clearly do things that used to be impossible, but the question of justification is interesting.

Daniel

所以这未必是说你必须有一个物理实例才能有意识,但在这座抽象之山上存在一个表征的谱系。看起来模型确实在多个分辨率层级上使用表征,但智能不是某种神奇的特质。我仍然认为它与任务量有关,与你已经知道的事实有关。而且我认为智能相当有前景。所以我仍然认为它会是碎片化的、不均衡的,你知道,部分碎片化的,如果这说得通的话。

So this is not necessarily an argument that you should have a physical instance to be conscious, but there is a spectrum of representations on this mountain of abstractions. And it seems models really use representations on many levels of resolution, but intelligence is not some magical quality. I still think it is related to volume of tasks and with the fact that you already know. And I think that intelligence is quite promising. So I still think it will be fragmented, uneven and, you know, partly fragmented, if that makes sense.

令我担忧的里程碑 The Landmark That Worries Me

Host

那么,有一个问题:你有没有想过,时间是否会限制 SSI 的出现,或者我们什么时候能看到像“AI 2027”那样的情景?具体来说,我的意思是,比如……ARD 任务。是的,是的。

So, one of the questions: have you thought about whether time limits the appearance of SSI, or when we can see a scenario like "AI 2027"? Specifically, I mean, for example... the ARD mission. Yes, yes.

Daniel

我认为有一个里程碑让我非常担忧——或者也许不是里程碑,而是 AI 发展的一个阶段,我认为它极其重要——那就是当一家 AI 公司宁愿放走他们的人,也不愿放走他们自己的 AI 的那一刻。也就是说,他们会拒绝所有人类劳动,而不是拒绝 AI 的所有工作。

So I think that one landmark which worries me a lot — or maybe not landmark, but a stage of development of AI, which I think is extremely important — is the moment when an AI company would rather release their people than their own AI. That is, they will refuse all human labor rather than all the work of AI.

Daniel

我认为现在我们显然仍处于这样的情形:Anthropic 或 OpenAI 宁愿留下他们的人类员工,也不愿……因为公司会直接散架。有些事现在只有人能做。因此,如果他们放走所有人,公司就会直接崩溃。但未来一切都会不同。未来 AI 将能够做到,对吧?否则,执行所有任务。

I think now we obviously still are in situations where Anthropic or OpenAI would rather leave their human employees than... because the company simply will fall apart. There are things which now only a person can do. Therefore, if they freed all the people, the company simply would collapse. But in the future everything will be otherwise. In the future AI will be able to do it, right? Otherwise, execute all tasks.

Host

以前,是的。是的。

Ago, yes. Yes.

Daniel

所以,从我们的观点看,这会在某个时刻发生,因为没有什么根本性的东西会阻止 AI 达到人类水平的能力。主要的一点,唯一的问题就是何时。而我们,你知道,我们投入了大量内部的分析和反思,研究各种预测这一点的方法论。Daniel 和我在这个问题上观点略有不同。但最终,我认为这大概是最重要的——时间限制的问题,或许对思考 AI 的未来最为重要。嗯,至少是最重要的问题之一。

So, from our point of view, this will happen at some moment, because there is nothing fundamental that would prevent AI from achieving human-level capabilities. The main thing, the only question is when. And we, you know, we are spending a huge internal volume of analysis and reflection over different methodologies predicting this. Daniel and I have slightly different views on this question. But ultimately, I think that this is probably most importantly — the issue of time limits is perhaps most important for thoughts about the future of AI. Well, at least one of the most important issues.

递归自我改进与自动化科学家 Recursive Self-Improvement and Automated Scientists

Host

嗯,很想知道你是否对此有特别的看法。好吧,让我在回答这个具体问题之前先表达几点想法。现在有很多初创公司都在做这种递归自我改进、超级智能。我采访过其中很多家。比如,前几天我和伦敦 Inherent 的 Edward Hughes 聊过。他,你知道,用机器学习的一整个系列流行工作复现了许多科学实验。他做这件事时把文章里的数字藏起来,并采用了一个 270 亿参数的模型 Gwen。于是他们把 GRPO 应用到 Gwen 模型上。正是如此——他们解决了适应性问题,因为重新训练一个大型模型非常困难。所以他们,你知道,改造了一个管理模型,用来管理像 Codex 这样的系统。

Well, it's interesting to hear whether you have a special look at this. Well, let me express a few thoughts before I answer this specific question. So many startups now are engaged in this recursive improvement, superintelligence. I interviewed many of them. Here, for example, the other day I talked to Edward Hughes at Inherent in London. And he, you know, recreated many scientific experiments with a whole series of popular works from machine learning. He did it hiding the numbers in articles and taking a 27-billion model, Gwen. So they applied GRPO to Gwen models. And that's exactly it — they decided the problem of adaptability, because it's very difficult to retrain a big mess model. So they, you know, adapted a management model for management of such a system like Codex.

Host

而他们的论点是:因为如果他们能够以构念效度复现这些实验的简化版本——也就是说有一个 LLM 裁判来检查,确保它们不作弊、一切都做对了——这本身就很有意思,因为我们,你知道,机器学习专家一直在教这个。我们总说这些东西会找漏洞。效度上总会有各种问题。说来奇怪,实际上它并不像我们预期的那么大问题。

And their thesis was: because if they will be able to reproduce reduced versions of these experiments with construct validity — that means having an LLM judge who checks so that they don't cheat and everything did the right thing — even this is interesting, because we, you know, machine specialists teaching. We would always say that these things look for loopholes. There will always be any problems with validity. How, not strange, actually it is not so big a problem, how we expected.

Host

而他认为:如果他们能复现这些实验,那么,如果他们不具创造性,又为何?如果它们有能力复现事物,为什么不迈出下一步,不说“哦,这是个有趣的问题。这是一个有趣的新任务待解决”,然后继续前进呢。所以我对这项研究相当感兴趣,我真的认为在最近的未来我们将创造出具有人工智能的自动化科学家。

And he believes: if they can reproduce these experiments, so why, if they weren't creative? If they have the ability to reproduce things, why don't they do it the next step and not say, "Oh, that's an interesting question. This is an interesting new task for solution," and move on. So I have enough interest in this research, and I really think that in the nearest possible future we will create an automated scientist with artificial intellect.

Host

但我脑子里一直有个想法:在硅谷存在某种简化或物化事物的文化。比如,Elon Musk 会说:“好吧,你是工程师,这是你的结果,这是你的数字,你必须让指标增长,我们把一切都当作优化任务。”我一直认为,对于某些抽象类型的任务,如果我们有足够的规格说明,这很棒,你可以爬到山顶。

But in my head always has the idea that in Silicon Valley there is a certain culture of simplification or objectification of things. For example, Elon Musk will say, "Well, you're an engineer, here is your result, here are your figures, and you must do so that the indicators grew up, and we consider everything as a task of optimization." And I always think that this is great for certain abstract types of tasks where you can climb to the top of mountains, if we have enough specifications.

Host

所以,在优化中有一个有趣的点:如果你有足够的规格说明,一个 AI 系统实际上可以收敛到解。但当你处于模糊模式时,你必须要有规格说明,而这意味着什么真的很神秘。

So, in optimization there is an interesting point: if you have enough specifications, an AI system can actually coincide to the solution. But when you are in a mode of ambiguities, you must have specification, and it's really mysterious what this means.

可验证与模糊目标 Verifiable vs. vague goals

Host

为什么我们有某种深度理解的感觉,不管那是什么,但 AI 系统却没有?

Why do we have a feeling of deep understanding, whatever it is, but AI systems don't?

Daniel

我认为在客观的、半定义的领域里,我们可以攀登一座山并做优化。我可能有点疯狂,但我仍然觉得缺了点什么。

I think that in objective, semi-defined domains we can climb a mountain and optimize. I'm crazy, but I still feel that something is missing.

Host

好吧,我的答案大概是:在能被客观验证的事情和不能被验证的事情之间,其实并不存在二元对立。或者说如果存在,那么能被验证的东西——也就仅此而已。

Okay, well, my answer would probably be: there actually is no binary between things that are objectively verified and those that aren't. Or if there is, then the things that can be verified—that's just it.

Host

比如创办一家创业独角兽,也就是估值十亿美元的公司。这是关于真实世界的一个可以被验证的事实。长期来看,可以说是一个有前景的目标。它很难验证。是的。核查起来很昂贵,但它是一种量化指标。

For example, creating a startup unicorn, that is, companies with a valuation of a billion dollars. That's a fact about the real world that can be verified. Long-term, well, let's say, a promising goal. It is difficult to verify. Yes. It's expensive to check, but it's a kind of quantitative indicator.

Host

你知道面试题是怎么运作的,比如编程题,现在人们都在积极做的那种,对吧?它们可以非常容易、廉价且用算法验证。而现实世界里也有这样一些事情,反馈周期更昂贵,是为更长期的视角设计的。

You know how interview tasks work, for example with programming, that people are currently actively working on, right? They can be very easily, cheaply and algorithmically verified. And there are also such things from the real-life world where feedback cycles are more expensive and designed for a longer perspective.

Daniel

我认为,我的观点是,我们会看到 AI 能力不断扩展,部分原因可能是 AI 公司使用的强化学习(RL)在数量和类型上的增加。更多样化的长期任务。是的。这将是一个持续覆盖不同类型问题的过程,取决于它们有多容易被检验,直到我们触及人类所能做的一切。毕竟,我们人类不知怎么学会了做这些长期任务。

I think, my opinion is that we will watch a constant expansion of AI capabilities, probably partially due to the increase in volume and types of learning from reinforcement (RL), which AI companies use. More diverse long-term tasks. Yes. And this will be a continuous process of covering different types of problems, depending on how easy they are to check, until we get to everything that people are capable of doing. After all, we humans somehow learned to do these long-term tasks.

Host

如果我可以补充的话,我也相信 AI 在一切事情上都变得更好,包括模糊的、难以验证的、概念负载很重的事情等等。比如,试着和 GPT-3 或 GPT-4 聊,然后再和 Fable 聊你最喜欢的那个无法验证的模糊任务,你很可能注意到更新的 AI 模型在这方面应对得明显更好。所以是的,进步是巨大的,我预计它会继续。

If I may add, I also believe that AI becomes better at everything, including vague ones, difficult to verify, conceptually loaded things, etc. Try to speak, for example, with GPT-3 or GPT-4, and then with Fable about your favorite vague task that's impossible to verify, and you most likely will notice that newer AI models are significantly better at coping with this. So yes, progress is colossal, and I expect it to continue.

Daniel

是的,但我在试着选一个好的例子。嗯,你知道,有这样一种社会学论点。我不知道你有没有读过 David Graeber 的《狗屁工作》那本书,他在书里采访了人们,他们真的说:“我的工作毫无意义,”你明白吗?喝上三四杯啤酒之后,律师们会说:“你知道,我做的绝大部分事情没什么大意义。”所以如果我们开始把经济中发生的一切都客观化并加以衡量……

Yes, but I'm trying to choose a good example. Well, there is, you know, such a sociological argument. I don't know whether you read the book David Graeber's "Bullshit Jobs" where he interviewed people, and they actually said: "My work is nonsense," you understand? After three or four many glasses of beer lawyers will say: "You know, most of what I do doesn't have a big meaning." So if we start to objectify and measure everything that takes place in the economy...

2032年六千万智能体 60 million agents by 2032

Host

顺便说一下,我在这里做了个笔记。我想你说过,到 2032 年可能会有 6000 万个智能体,它们的工作速度比人类快 20 倍,我就在想,这到底意味着什么?这合乎逻辑吗?结论是我们可能会有一个只由人工智能构成的经济体?而一个只有 AI 的经济体到底有没有意义?总之,帮我理解一下这个。

By the way, I made a note here. I think you said that by 2032 there may be 60 million agents who will work 20 times faster than humans, and I just think, what does this even mean? Is it logical? The conclusion is that we could have an economy that consists only of artificial intelligences? And does an economy where there is only AI make sense at all? In short, help me understand this.

Daniel

是的,我的观点是这样的,这基本上是真的。AI 可以做一切,或者至少一切真正重要的事情。你提到了“无意义工作”这个概念。我想先说,把创造更好 AI 所需的一切视为第一步,这就覆盖了经济的一大部分。比如,为此你需要能够制造更大、更好、更多的芯片。这覆盖了整个半导体供应链。要建立这样一条链,你需要整个发达经济体。那会为建造新工厂创造新的工作。然后你需要生产工厂、机器人来制造更多机器人。还需要更多研究人员,用这些巨大的算力来构建更好的 AI。

Yes, my opinion is this, which is essentially true. AI can do everything, or at least everything that really matters. You mentioned the concept "pointless work". I would start by saying to consider everything that's needed for creating better AI as a first step, that covers a large share of the economy. For example, for this you need to be able to create bigger, better and more numerous chips. This covers the whole supply chain of semiconductors. To build such a chain, you need the whole developed economy. That builds new works for construction of new factories. And then you need production plants, robots to create even more robots. And more researchers are needed, to build better AI using these huge computational power.

Daniel

所以我认为,一旦你拥有了这一切,就足以显著加速并从根本上改变世界,即使存在职业执照或其他什么东西,阻碍 AI 在经济中从事随机的法律或其他“无意义”工作。我认为真正重要的、真正有意义的是:特定的工作、计算和更好的 AI。一旦你拥有能做到这一点的 AI,以及提供该部门大规模增长的可能性,你就会看到巨大的进步,因为经济中“无意义”的部分将很难抑制那些寻求快速发展的领域的增长,这种发展由每个参与者身上存在的激励驱动。特别是,每个国家都有巨大的激励让自己的经济增长得比竞争国家更快。

So I think that as soon as you have all this, this is enough to significantly accelerate and to radically change the world, even if there are professional licenses or something else which hinders AI from performing random legal or other "meaningless" work in the economy. I think what really matters, what really has meaning: in particular work, calculations and better AI. And as soon as you have AI capable of this, and the possibility to provide mass growth of this sector, you will see huge progress, because the "senseless" parts of the economy will find it difficult to restrain the growth of those areas that are looking for quick development through incentives that exist in each participant. In particular, each country has a huge incentive for its economy to grow faster than in competitor countries.

经济作为自我复制系统 Economy as a self-replicating system

Daniel

再深入一点谈哲学,我认为很多经济理论和关于经济的讨论都聚焦于现有系统各部分之间的关系:价格波动、需求和供给等等。但如果你看得更广,经济整体上是一个自我复制的系统,它一直如此。你知道,几千年前,这是一个相对小而简单的自我复制系统,由几个村庄组成:人们从事农业,生孩子,建立新的定居点,在那里又耕种土地、生孩子,如此循环。它呈指数增长,但速度非常慢。

Going a little deeper into philosophy, I think that much economic theory and discussion about the economy focuses on relationships between parts of the existing system: oscillations of prices, demand and supply, etc. But if you look more broadly, the economy overall is a self-replicating system, as it always was. You know, thousands of years ago this was a relatively small and simple self-replicating system of several villages: people were engaged in agriculture, gave birth to children, founded new settlements, where again they worked on the land and gave birth to children, and so on. It grew exponentially, but at a very slow pace.

Daniel

现在它是一个复杂得多的自我复制系统,包括卡车、运输设备、工厂、矿山等等。然而,在高层次上看,这一切仍然是一个自我复制系统,其中有人们、卡车、汽车、建筑,它们共同创造出更多的人、更多的工厂、更多的建筑、更多的汽车等等。也许再过几年,我们就会到达这样一个节点:有可能创造出一个完全由机器运行、使用人工智能和机器人的自我复制系统。根据我们的计算,这样一个自我复制系统的翻倍时间将显著少于当前经济所需的约 20 年。所以,你知道,这就是未来等待我们的东西。

Now it's a much more complicated self-replicating system that includes trucks, transportation equipment, factories, mines, etc. However, at a high level, that's all still a self-replicating system where there are people, trucks, cars, buildings, and together they create more people, more factories, more buildings, more cars and so on. And perhaps already in a few years we'll get to the point that it will be possible to create a self-replicating system that is completely run on machines, using artificial intelligence and robots. In accordance with our calculations, the doubling time of such a self-replicating system will be significantly less than the roughly 20 years needed by the current economy. So, you know, that's what awaits us in the future.

退化与政权崩溃 Degradation and regime collapse

Host

是的,这在我看来是合理的,但不知为何我的直觉认为,这在某种程度上会导致退化。你知道,我认为 David Graeber 虽然用了“无意义工作”这个词,但他的意思是工作中存在一些我们不完全理解的难以捉摸的成分。某种社会学功能或类似的东西。

Yes, it seems plausible to me, but for some reason my intuition suggests that this will in a certain way lead to degradation. You know, I think David Graeber, although he used the term "senseless work", meant that there exist in work some elusive components that we don't fully understand. Some sociological function or something similar.

Host

还有一点:我在读 Citrini 的那份报告,就是你谈到的那份,写的是当人们比如开始大规模偿还房贷,而他们的工资会下降,于是他们就不再是劳动力市场的积极参与者时会发生什么。你也谈到了“来自 AI 的红利”以及其他一些东西。但即便是这也让人想到,如果人们停止参与各种过程,经济会发生某种“体制”“崩溃”。你觉得这是真的吗?

And one more point: I was reading the Citrini report, which you talk about, writing about what will happen when people, for example, start massively paying off mortgage loans, and their wages will fall, so they will stop being active participants in the labor market. You also talked about "dividends from AI" and something else. But even this suggests the idea that what if people stop taking part in processes, a kind of "regime" "collapse" of the economy happens. Do you think that's true?

Daniel

嗯,有可能,但再说一次,好吧,让我们想象一下这个……我不是一个足够好的经济学家,我没算过。

Well, potentially, but again, okay, let's imagine that this... I'm not a sufficient economist, I didn't count.

经济影响与消费者需求 Economic Impact and Consumer Demand

Host

例如,当消费者需求下降时,价格会发生什么变化?这会产生什么后果?

What will happen to prices, for example, when consumer demand falls? What will be the consequences of this?

Daniel

我不太确定。我认为在我们的经济模型中,我们并没有被强烈地考虑进去。只考虑了一点点。

I'm not really assured. I don't think that we are strongly taken into account in our economic model. A little was taken into account.

Host

好的。请稍微讲讲。

Okay. Tell me a little about it.

Daniel

但是的。是的。但假设,即使情况真的很糟,即使消费者不再有需求或类似情况,如果你拥有如此水平的 AI 能力和机器人,可以拥有完全自主的 AI 和机器人来做一切事情,那么即使像 Anthropic 这样的公司,如果它足够大,并且可能与其他公司合作,比如采矿业,也能运行这个完全自给自足的过程。所以,无论所有人类发生什么,在沙漠中就可以简单地发展整个产业,每一年翻一番。你知道,职业、自动驾驶卡车、工厂,它们建造人形机器人,这些人形机器人生产更多的人形机器人,建造更多的工厂,生产芯片等等。这一切都可以每年翻一番,每 6 个月,每 3 个月,越来越快,因为技术在进步,因为当然,AI 也会研究改进技术的方法。然后你会发现自己处于这样的境地:谁知道世界其他地方发生了什么,但 Anthropic 拆解了月球部件。或者,例如。

But yes. Yes. But hypothetically, even if it's really bad, and even if consumers no longer have demand or something like that, if you have such a level of AI capabilities and robots that you can have completely autonomous AI and robots that do everything, then even a company like Anthropic, if it is large enough and possibly collaborating with other companies, such as mining, can run this fully self-sufficient process. So, regardless of what is happening with all the people, in deserts can simply develop whole industry that doubles. You know, careers, self-directed trucks, factories, who build humanoid robots that produce more humanoid robots that build more factories with chip production etc. This can all be just double every year, every 6 months, every 3 months, faster and faster because technology is improving, because, of course, AI too will investigate ways to improve technologies. And then you find yourself in situations when who knows what is happening in the rest of the world, but Anthropic took apart the moon parts. Or, for example.

Host

是的,为了澄清,我认为这依赖于极高水平的 AI 能力,对吧?而我的感觉,或者我有稍微不同的直觉,有时会出现直觉:“真的吗?”真的吗?我认为 Fable 或 Fable 或 Mythos 的未来后代版本能够做到我们在这里谈论的一切吗?我认为这一切都归结于你是否将 AI 视为我们目前使用的那类 AI,还是更像一个智能体,等同于人类工作者。实际上,就像云中的人。呃。云中的同事。

Yes, and for clarity, I think that relies on an extremely high level of AI capabilities, right? And my feeling, or I have slightly different intuitions, and sometimes occurs intuition: "really really?" Is it really? I think Fable or future descendants versions of Fable or Mythos could do everything that we are talking here? And I think that all boils down to whether you think about AI as something from the class what we are for now we use AI, or rather about, you know, agent human worker equal. As, in fact, man in the cloud. Ahem. A colleague in the cloud.

Daniel

云中的同事,是的。而且是的,我认为过去几年 AI 的发展可以足够准确地描述为现代 AI 系统和“云员工”之间的一种过渡。所以,“云员工”在未来的想法看起来相当有前景。而且我不认为我们会停在那里。我认为我们会达到超人类水平。嗯,是的。是的。

A colleague in the cloud, yes. And yes, I think the last several years development of AI can be described accurately enough as a kind of transition between modern AI systems and "cloudy employees". So, the idea of "clouds employees" in the future looks quite promising. And I don't think we're on Let's stop there. I think we'll go out on superhuman level. Um, yes. Yes.

AI 2040与超级智能 AI 2040 and Superintelligence

Daniel

我想在这里指出一点:如果人们阅读《AI 2040》的 A 计划,其中一个结论——我认为这是我们世界的一个非常重要的事实——是即使停留在领先专家的水平,一切都会发生根本性的变化。根据我们的脚本,不是智能爆炸,而是通过国际协议禁止智能爆炸的发展和递归 AI 自我改进。所以它们大约停留在最佳专家——人类的水平,至少几年。最终,它们达到超级智能。特别是,在 2040 年它们达到超级智能。但在 2030 年代有一段时期,它们实际上在所有学科都拥有人类水平的 AI,但没有多少能超越这一点。然而,即使这样也足够了:如果你进行经济建模,你会看到你似乎获得了一群云中的同事——出色的员工,几乎可以替代人类做所有事情;只是他们比人类更便宜、更快,而且不需要 20 年才能自我复制。相反,它们每年翻一番。结果,到 2030 年代末,世界完全转变,你知道,所有人实际上都会失业。巨大的新城市将出现,由机器人建造。在特别经济区有巨大的职业,从哪里提取?从地球内部提取大量矿物。在海洋中,太阳能电池板填满地平线——所以疯狂的事情甚至仅凭人类水平的 AI 也是可能的,如果给予指数增长的时间。

One point that I want to point out here: if people read plan A with "AI 2040", one from the conclusions, which can be done—and this, I think it is very important a fact about our world—is that even if to stay at a level leading experts, everything is radical will change. According to our script, instead of intellectual explosion is contained international agreement on ban on explosives development of intelligence and recursive AI self-improvement. So they are stopping approximately at the level the best specialists- people, at least on several years. Eventually, they come to superintelligence. In particular, in 2040 they reach superintelligence. But there is period in the 2030s, when they actually have Human-level AI in all disciplines, but nothing much would surpass this. However, even this enough: if you make it economical modeling, then you will see that you seem you get a population colleagues in the cloud—wonderful employees, which can replace people in almost everything; only they are cheaper for people faster than people, and they don't it takes 20 years to recreate oneself. Instead, they doubling every year. And as a result the world to late 2030s fully is transformed, and, you know, all people will actually remain without work. Huge ones will appear new cities, built by robots. Huge careers in special economic zones, Where is it extracted from? colossal volumes minerals from the earth's interior. Solar panels that fill the horizon in the ocean—so crazy things are even possible with AI only human level, if given time for exponential growth.

AI组合性的挑战 Challenges of AI Compositionality

Host

是的,我想挑战你一个观点:当我们拥有 AI 时,我们可以直接复制它的规模并复制它。你可以运行它一千次。你可以在这里有一个副本,为这项工作获得技能,在那里有另一个——为另一项工作,然后你可以把它们组合在一起,或者错了吗?你可以在某种程度上组合技能,而这一切都是多层次和复合的。但这与我的经验不太一致。所以我对它非常热情。我用人工智能做了什么,我发现这是可能的。让智能体在某些智力领域高度合格。你可以吸引大量信息并教育它们做某些事情,但我认为它们还不是复合的。我认为如果它们是复合的,我会更加焦虑。你怎么看?那么有没有办法让你尝试组合它们?这是在输出过程中完全发生的吗,你教它们两项单独的技能,然后你尝试以某种方式组合规模?

Yes, I want to throw you challenge to the idea that when we have AI, we can just copy his scales and duplicate it. You can run it a thousand times. You can have one a copy here, which acquires skills for this work, and another there—for another, and you can just combine them together, or wrong? You you you you can to some extent combine skills, and all this is multi-level and composite. And this not very consistent with my experience. So I am very passionate about it. what I did with artificial intelligence, and I discovered that it is possible. make agents highly qualified in certain intellectual areas. You can attract a lot spring information and education them to do certain things, but I don't think so they already are composite. I think if they were such, it caused I would have much more anxiety. What do you think about it? So is there a way to you try them combine? Is this is happening completely during output, are you teach them two individual skills, and then you try somehow combine the scales?

Daniel

嗯,是的,我的意思是它可能是一个单独的东西,但目前它们通过思维链和一种表面技能的适应来适应。所以,基本上就是这样。记忆系统。而且,老实说,它不是很有组合性。这是组织现在面临的一个大问题。所有这些开发者都适应他们的表面技能,本质上他们的智能体——这些是不同的人。而且他们非常非常难以共享技能,因为它们可能会破坏另一个智能体,因为,你知道,由于某种原因它不起作用。现在我可以想象一个未来,我们适应权重。这就是 Inherent 的那些人做的。也许那时它会神奇地解决问题,我们可以解决这个知识交换问题,对吧?这是最重要的事情。我们如何在个体平等和组织与智能体之间积累信息,有点像在《黑客帝国》中,可以直接下载技能并执行任务。也许这是可能的。即使如此,我仍然认为,神经网络中的表示,我称之为分散混乱的表示,有点不稳定。它们不是很可靠,但它们在某种程度上有组合性。是的。

Well, yes, I mean what could it be separate thing, but at the moment they adapt through chain of thoughts and a kind of adaptation surface skills. So, that's basically it. memory systems. And, honestly, it's not very compositional. And this is a big problem, with which organization are facing now. All these developers adapt their surface skills, and essentially their agents—these are different people. And them very, very difficult share skills, because they can to break another agent, because, you know, it does not work with for some reason. Now I can imagine the future where we are We adapt the weights. This is what these did the guys from Inherent. And, maybe then it is will magically solve problem, and we we can solve this exchange problem knowledge, right? It the most important thing. How do we accumulate information on individual equal and on par organizations to agents, a bit like in "Matrices, could just download skills and perform task. Maybe this and possible. Even then I still think, that representation in neural networks, I I call them scattered confused representations, are a little unstable. They are not quite reliable, but they to some extent composite. Yes.

专业化与AI的未来 Specialization and Future of AI

Daniel

所以,关于这一点,我会说:即使你是对的,我不认为这会严重破坏我们在这里描绘的未来,因为,好吧,现在不是一个 Claude 模型执行所有任务,也许你有工作,有一百个 Claude 模型或一千个 Claude 模型,专门从事各种职业,你明白吗?但你得到的是同样的结果。我唯一想说的是,这就像人们是专业化的。专业化的。然而,我会说,与人类相比,似乎有这样的东西。

So, I would say regarding this: even if you have You're right, I don't think so. this will seriously undermine the future that we We draw here because, okay, now instead of one Claude model, which performs all maybe you have a job there are a hundred models Claude or a thousand models Claude, specialized for various professions, Do you understand? But you all you get the same the very result. The only thing I would like to say, this is what Just like people are specialized. specialized. Yet I would say that compared to people, It seems there is such a thing.

单一模型与多个专用模型 One Model vs. Many Specialized Models

Host

我在想知识这件事,对吧?也许如果你回到 10 年前,我们来做这场讨论,那看起来很可能你会需要一千个不同的 Claude 模型,由 Anthropic 创造,用于不同种类的智力工作。比如 Claude 用于编程,Claude 用于物理,Claude 用于文学。而支持这一点的论证会很简单。听起来会是这样:“嗯,人就是这样运作的。没有一个人什么都懂。相反,人们各自专精于不同学科。”而模型的参数量是有限的。也许你就是没法把所有知识都塞进这么有限的参数量里,所以你需要针对不同事物的专用 AI。而实际上,对于足够小的模型来说这是对的。对于非常小的模型,你就是没法教它们所有东西。所以专用模型对不同的用途会有用。但经验上我们学到,对于足够大的模型,你可以直接用所有东西训练它们,然后它们就一次性学会了所有东西。它们不会因为也被训练了一堆编程而在物理上变差。事实上恰恰相反。编程对物理还有一点提升,你懂吧?这就是为什么我真的认为,最可能的未来是:会有一个在几乎整个经济上训练出来的模型,它在一切方面都超越人类。这看起来就是当前趋势的自然延续。然后,也许你会有不同的、比如说更便宜的版本。比如你会有蒸馏模型。对,更小的东西。因为你确实想最大化地节省算力资源。所以对于任何具体任务,你都会有一个尽可能便宜的模型来完成它。

I'm thinking about knowledge, right? Maybe if you returned 10 years ago and we were having this discussion, it would seem quite likely that you would need a thousand different Claude models, created by Anthropic, for different kinds of intellectual work. For example, Claude for programming, Claude for physics, Claude for literature. And the argument in favor of this would be simple. It would sound like this: "Well, that's how it works for people. There is not one person who knows everything. Instead, there are people who specialize in various disciplines." And the models have only a limited number of parameters. Maybe you just can't squeeze all this knowledge into such a limited quantity of parameters, and you need specialized AI for different things. And actually, for small enough models that's true. For the very tiny ones, you just don't have models that can be taught everything. So specialized models would be useful for different things. But empirically we learned that for big enough models you can just train them on everything, and then they study everything at once. And they don't become worse at physics because they are also trained on a pile of programming. In fact, it's the opposite. Programming gives a small increase for physics, you know? And that's why I really think that the most probable future is this: there will be one model trained on virtually the entire economy, which simply surpasses people in everything. It seems that's the natural continuation of the current trend. And then, perhaps, you will have different, let's say, cheaper versions of these models. For example, you will have distilled models. Yes, smaller things. Because you really want to maximize saving of computer resources. So for any specific task you will have as cheap a model as possible that will do this task.

Daniel

是的,我大体上同意这一点。我的看法是,模型同时是所有人的声音,又不是任何人的声音。所以它们有一种典型的嗓音,因为它们有系统提示词,也有某些它们倾向的典型模型行为。但当一位像你这样的专家使用这些模型时,你似乎就注入了一种特定的视角。于是你的每一个词、所有参考资料、你的记忆系统等等,你似乎都在从这些模型里雕琢出性格。你在这个特定行业里一致地激活知识。而美妙之处在于,许多其他领域的人也能做同样的事。存在这样一种幻觉,以为模型具有通用能力,但我认为模型可以被设定为任何领域的专业专家,只不过这是一种隐藏的、非显性的能力。这说得通吗?

Yes, I mostly agree with this. I mean, my opinion is that models are at the same time the voice of everyone and no one's voice. So they have a typical voice, because they have a system prompt and certain typical model behaviors to which they are inclined. But when an expert, such as you, uses these models, you seem to instill a certain perspective. So your every word, all reference materials, all your memory system, etc. You seem to be chiseling character from these models. And you activate knowledge consistently in this particular industry. And the beauty of that is that many other people in different fields can do the same. And there is such a thing as an illusion that the model has general abilities, but I think the models can be set to specialized experts in any field, but this is a hidden, not explicit ability. Does this make sense?

Host

是的,对我来说这听起来合理。为什么这是一种幻觉?看起来它们确实有通用能力。它们能做很多事情。比如,你可以给它任何具体任务。比如说这道数学题,它会一直推进直到解出来。因此,它解决了这个智力问题。或者你可以提出任何知识问题;也许路径是通过标准教学模式铺好的。或者如果有点特殊,专家可以介入,就像在黑暗中打一束光;你让阻力最小的路径大致正确,然后它就会把一切做对。所以我推测存在某种监督的幻觉。所以,当专家在明确定义的领域使用它时,会发生神奇的事,但仍然留有一些模糊的空间。

Yes, for me it sounds reasonable. Why is this an illusion? It seems they really have general abilities. They can do many things. Well, for example, you can give it any specific task. Let's say this math problem, and it will move until its resolution. Therefore, it solves this intellectual problem. Or you can put any question on knowledge; and you know, maybe the way was laid through standard modes of teaching. Or if it is something a bit specific, then the expert can intervene and, you know, as if to illuminate a flashlight in the dark; and you make a way of least resistance approximately correct, and then it will do everything right. So I assume that there is a certain illusion of supervision. So, when experts use this, in clearly defined areas, magical things happen, but there still remains some space for ambiguities.

集体智能与众多AI Collective Intelligence and Many AIs

Host

但你知道,这把我们带到下一个问题:你怎么看这样的智能,对吧?我们集体智能的美妙之处在于,我们有这么多不同的人,依赖各种不同的思想传统,我们都在攻问题,不管对错,对吧?所以,当有大事发生时,比如 Twitter 的闹剧,我们都被激励去找漏洞。于是我们一起展现出智能。我们找到有趣的视角,算法优先选出最好的那些,我们一起用我们的头脑。我可以想象 AI 也会以同样的方式行动,对吧?所以我们有各种不同的 AI,各有不同的专长,它们做一切事情,从不同视角看问题。但我设想的未来更像是这样,而不是一个大 AI。

But, you know, it brings us to the next question: what do you think about such intelligence, truth? The beauty of our collective intelligence consists in that we have so many different people, who rely on various intellectual traditions, and we all attack problems, right or wrong? So, when something big is happening, a Twitter fiasco, we are all motivated to find gaps. So, we show intelligence together. We find interesting angles, the algorithm gives priority to the best of them, and we use our minds together. And I can imagine what AI will be like to act in the same way, right? So, we have various AIs that have different expertise and they do everything, you know, just looking at problems from different points of view. But I imagine the future rather like this, not like one big AI.

Daniel

是的,我想我基本上同意 Daniel 之前说的,从历史上看,我认为我们会拥有许多不同的高度专用 AI 来做不同事情的观点,是站不住脚的。相反,我们看到了规模化和把一切统一起来的好处。嗯,我想是这样。我喜欢的一个直观想法是,在人身上,通常要在一个人的身上具备所有技能才能取得真正重大的成果。比如 Elon,对吧?Elon 具备的是:一定程度的诚信,一定的技术知识,一定量的商业知识,外向性,推动他人的能力——我认为在这些技能上他都处于相当高的百分位。而 Elon 之所以如此罕见,之所以能同时管理所有这些极其庞大且成功的公司,而其他没人能做到,或者至少没成功,原因在于乘法效应:你需要在十个领域里每一个都处于第 90 百分位,这非常不可能。但如果你有一个 AI,可以被训练得在所有技能上都非常高,高到没人能及——这是极其罕见的,任何具体的人同时拥有所有这些技能的可能性无限小。我认为你实际上会非常、非常、非常擅长在这些具体而重要的方面改变世界,就像 Elon 所做的那样。

Yes, I think I basically agree with what Daniel said earlier about that historically, to me it seems that the view that we will have a lot of different highly specialized AIs that do different things is just not justified. Instead, we saw the benefits of scale and unification of everything together. Hmm, I think so. One intuitive idea, which I like, is that in people usually it is important to have all skills in one person to achieve really significant results. For example, Elon, right? Here is what Elon has: a certain level of good faith, certain technical knowledge, a certain amount of business knowledge, extraversion, ability to push people — I think, for each of these skills he is in quite a high percentile. And the reason why Elon is so rare and why he can manage all these incredibly large and successful companies at the same time, and no one else maybe, or at least not succeeded in this, consists of a multiplicative effect: you need to be in the 90th percentile in each of these ten areas, which is very unlikely. But if you had an AI which can be taught to be very high quality in all these skills as no one can — this is extraordinarily rare, you know, endlessly unlikely that any specific person had all these skills at the same time. I think you would actually be very, very, very skilled in changing the world in all these specific and important aspects, just like Elon did.

一群专业化的Claude A Swarm of Specialized Claudes

Host

如果我可以补充一点,我不认为这对我们描述的那种未来是关键的。假设我们在这点上错了,事实上最有效的前进方式是拥有一堆不同的专用 AI。嗯,很可能无论如何情况都会是这样:几家大公司,比如 Anthropic,会创造许多不同的专用 AI,会有许多不同的 Claude 模型供你选择,等等。实际上,这不会只是像在许多不同的 Claude 模型之间做选择,因为在我们所说的那个时代,它们会比现在自主得多。所以这更像是 Claude 自己在许多不同的 Claude 模型之间做选择。就像一个 Claude 集群,由许多不同的专用模型组成,它们协同工作。它们有某种内部的官僚结构。这个集群走向世界,签订商业协议,创造新技术,创办初创公司,处理所有这些事情。

If I may add, again, I don't think this is critical for that type of future that we describe. Suppose we are wrong about this, and in fact the most effective way forward is to have a bunch of different specialized AI. Well, probably anyway the situation will be like this: several large companies, like Anthropic, will create many different specialized AI, and there will be many different Claude models, among which you can choose, and so on. Actually, it wouldn't just be like a choice of many different models of Claude, because at that time, about which we speak, they will be significantly more autonomous than now. So this will be more like Claude himself chooses among many different Claude models. There is like a swarm of Claude, which consists of many different specialized models that work together. And they have some kind of inner bureaucratic structure. This swarm comes out into the world, concludes business agreements, creates new technology, launches startups and deals with all these things.

专用智能体与克隆体 Specialized Agents vs Clones

Daniel

这类似于,如果你有一群移民——他们各自拥有不同的专业技能,但他们一起合作创办新公司、找工作、做其他事情。情况会是这样:许多不同的 Claude,但它们全都会协同工作。所以,如果你把视野放宽,无论如何,这仍会是一种 Anthropic 吸收经济的现象。而且,你知道,工作总体上会开始自我复制。

It's similar to how, if you had a population of people—immigrants—they all have different specialized skills, but they work together to create new companies, get jobs, and do other things. It would be like this: many different Claudes, but all of them would work together. So if you look wider, anyway, it would remain a phenomenon that Anthropic absorbs the economy. And, you know, work in general would begin self-replicating.

Host

是的。为了公正起见——我也不认为这是关键时刻。

Yes. Justice for the sake of—I don't think so either that this is the key moment.

Daniel

也许如果那是对的,那么仅仅在分布式集体系统中,通过不同智能体之间的消息交换,可能会出现额外的狭窄之处。我在圣塔菲读到过一篇很棒的关于此的报告,说即使在自然界中,个体智能与团队智能的比例也存在某种“黄金中道”。我们在这方面可能会遇到某种奇怪的巧合。但是的,思考它如何运作本身就足够有趣了。

Maybe if that's right, then only insofar as in a distributed collective system, extra narrow places may arise through exchange of messages between different agents. I read a great report in Santa Fe about this, that even in the natural world there is a certain "golden middle" in the proportion of the intelligence of the individual and the team. And we can come to some strange coincidence in this. But yes, it's just interesting enough to think about how it works.

Host

抱歉,Daniel,请继续。

Sorry, Daniel, continue.

Daniel

那样会更安全。我觉得这无所谓。在这样一个世界里,一致性方面会留下许多严重问题,但会稍微安全一些。你记得的原因:当有许多不同的专业化智能体彼此互动时,控制所发生的事情要比它们全是克隆体、知晓世上一切时更容易。所以也许,为了避免过度僵化,我们应该说我们支持你所描绘的愿景,而不是我们所描绘的那个。

That would be safer. I think it doesn't matter. There would be a lot of serious problems left with consistency in such a world, but it would be a bit safer. The reason you remembered: it's easier to control what happens when there are many different specialized agents that interact among themselves, compared to if they were all clones and knew everything in the world. So maybe, to avoid excessive stiffness, we should say that we support the vision which you drew, not the one we are drawing.

Host

嗯哼。是的,这非常正确。而且,我想补充一点,我真的很喜欢教学如何在个体层面发生的这个概念。你知道,我把它与进化相比较。所以有某种系统发生适应、个体发生适应、文化适应。我认为 AI 的下一波浪潮是——当我们真正拥有彼此交流、学习并专业化的智能体时。也许也会有不受欢迎的行为。也许会出现阴谋,很多坏事会发生。但我认为这一切——它就是必然要发生的。

Uh-huh. Yes, that's very true. And, to things, I want to add that I really like the concept of how teaching takes place on an individual level. Well, you know, I compare this with evolution. So there is kind of, you know, phylogenetic adaptation, ontogenetic adaptation, cultural adaptation. And I think what the next wave of AI is—this is when we really will have agents who communicate among themselves, learn, and specialize. And maybe there will be unwanted actions too. Perhaps a conspiracy will arise and a lot of bad things will happen. But I think that's all—it just has to happen.

埃隆的魔力 The Magic of Elon

Daniel

但有趣的是你提到了 Elon。所以,我认为 Elon 的魔力并不那么在于他在工程和优化方面的天才。嗯,这在于他识别有趣且未来可行领域的能力。因为这是创造性的一面。这是科学性的一面,而非工程性的一面。因为如果我们以亚马逊为例——这是一个适应性生态系统。它就像一个有机体。而他所做的就是不断思考适应和重构其结构的新方式。所以,比如可能是物流。然后他给出了一堆技能,然后无情地优化这些技能。有一个适应性组件,有一个优化组件,而这个身体始终处于运动之中。

But it's interesting that you remembered about Elon. So, I think what the magic of Elon consists of is not so much his genius in engineering and optimization. Hmm, this is his ability to recognize areas that are interesting and can work in the future. Because this is the creative aspect. This is a scientific aspect, and not engineering. Because if we take Amazon as an example, for instance—so this is an adaptive ecosystem. It's like an organism. And what he does is constantly think about new ways of adapting and reconfiguring its structure. So it could be, for example, logistics. And then he gives out a bunch of skills, and then mercilessly optimizes these skills. There is an adaptive component, there is a component of optimization, and the body is constantly in motion.

Host

但是,你知道,我认为关于这一点的问题——前景在于这在原则上有多大比例能由 AI 完成。显然——你认为这一切吗?

But, you know, I think that question on this—the prospects are in how much of this could in principle be done by AI. Apparently—do you think that all this?

Daniel

是的,全部都是。显然,是的。也许清晰表达这一点的方式之一——就是看,大脑是一台机器。大脑能做的一切,我们都能借助机器做到。我认为一个非常有用的练习是将真实人脑的架构与现代图形处理器或整个数据中心进行比较。如果你尝试做这个比较,你会看到图形处理器,H100,实际上与人脑有相当相似的特征。呃,你知道,它相当相似——在我看来,在最关键的指标上基本相似:每秒多少 flops,也就是说,与图形处理器相比,大脑的总计算能力是多少?对于 H100,你知道,在 FP16 格式下大约是每秒 1E15 flops。而对于大脑,你知道,这取决于具体怎么计算,以及是用突触数据库、基础神经元还是其他什么。嗯,好吧,这大约在 10 的 12 次方到 10 的 18 次方 flops 之间。所以 H100——可以说,正好处于中间——至少在对数分布上——对于这个数量级而言。

Yes, all of that. Apparently, yes. Maybe one of the ways to say it clearly—simply look, the brain is a machine. Everything that the brain can do, we can do with the help of machines. I think a very useful exercise is comparing the architecture of the real human brain with a modern graphics processor or a data center as a whole. And if you try to do this comparison, you will see that a graphics processor, the H100, actually has quite similar characteristics to the human brain. Uh, you know, it has quite similar—it's basically similar to me in the most important indicator: how many flops per second, that is, what is the general computational brain power compared with a graphics processor? And for the H100, you know, it's approximately 1E15 flops per second in FP16 format. For, well, for the brain, you know, it depends on exactly how to count and whether to use the database of synapses, or base neurons, or whatever other. Um, well, this is approximately, you know, somewhere between 10 to the 12th and 10 to the 18th flops. So the H100 is, one might say, right in the middle—at least on a logarithmic distribution—for this order of magnitude.

大脑与机器架构 Brain vs Machine Architecture

Daniel

还有,架构本身。它是神经网络,不是常规软件。所以它们开始时是随机的意大利面——随机交织生成的……好吧,它们开始时是偶然初始化的。里面就是巨大的意大利面式编织。是的,就像出生一样:你有一堆神经元,它们只是偶然连接起来,突触一对一。然后整个学习过程发生,连接被切断,一个架构开始成形,在任何教育环境中都能有效取得成果。AI 和人脑的工作方式存在差异,但总体而言它就像一个人造大脑。那么人们如何学习技能——实际上,对 Elon 来说拥有这些技能意味着什么?嗯,这意味着他大脑中某些神经元和突触链执行非常复杂精密的计算,这些就是技能。而 Claude 本身也有一堆链,你知道,在学习过程中被刻入其中——这些就是各种技能。而且,原则上,你可以创造一个足够大的 Claude,它将拥有与 Elon 同类型的链。

Also, the architecture itself. It's neural networks, not regular software. So they start as random spaghetti—random intertwining generated... well, they start as accidentally initialized. It's simply gigantic spaghetti-weave inside there. Yes, just like with birth: you have a bunch of neurons that just by chance are connected, synapses one to one. And then the whole learning process happens when connections are cut off, and an architecture begins to take shape that is effective for achieving results in any educational environment. There is a difference in how it works in AI and in the human brain, but in general it's like an artificial brain. And so how people learn skills—what, actually, does it mean for Elon to own them? Well, this means that certain chains of neurons and synapses in his brain perform very complex and sophisticated calculations, which are these skills. And so Claude himself has a bunch of chains which, you know, are engraved in it in the learning process—these are various skills. And, in principle, you can create a big enough Claude who will have the same type of chains as Elon's.

Host

是的。关于大脑与现代机器学习系统的比较,我还有几点补充:大脑比当前系统更并行一些。也就是说,有更多的并行计算在发生,但串行深度更小,对吧?也就是说,连续执行的运算数量。你知道,每秒能连续工作的运算或神经元的数量——这取决于神经元的类型,但,我希望我没记错,大约在 1 到 1000 之间,取决于神经元的类型,据我记忆。我以为大约接近 100。

Yes. And I have a few more notes on the brain comparison with modern machine learning systems: the brain is a little more parallel than current systems. That is, there is more parallel computation happening, but the sequential depth is less, though? That is, the number of calculations that are carried out successively. The number, you know, of calculations or neurons that can work sequentially per second—it depends on the type of neuron, but, I hope I'm not wrong, it's somewhere from 1 to 1000, depending on the type of neuron, as I remember. I thought it was somewhere close to 100.

Daniel

是的,我认为它在这个范围之内。我认为这取决于神经元的类型。但无论如何,计算机当然可以工作,并且串行执行计算要快得多,你知道吗?就像,你知道,你通常能达到吉赫兹级别的时钟频率,甚至更高。所以你可以在 GPU 上获得比大脑快许多个数量级的串行处理速度。但也值得注意的是,架构本身在许多方面建模远不如人脑。特别是,ML 模型的不同部分之间的互动比大脑的不同部分更复杂。所以我预计,要达到这样的 AI 水平,我们需要对现有模型进行若干算法改进。问题就来了:究竟需要多少算法改进,它们必须比当前系统好多少、质量多高?对我来说,这是一个非常开放的问题。

Yes, I think it's within the limits of the range. I think this depends on the type of neuron. But in any case, a computer, of course, can work and perform calculations much faster sequentially, you know? It's like, you know, you can often get a clock order frequency of gigahertz, or even more. So you can get many orders better speed of sequential processing on the GPU than in the brain. But it's worth noting too that the architecture itself models in many aspects much worse than the human brain. In particular, different parts of ML models interact among themselves in a more complicated way than different parts of the brain. So I expect that for achieving such an AI level, we will need a number of algorithmic improvements to existing models. And the question arises: how many exactly algorithmic improvements, and how high quality do they have to be, to differ from current systems? As for me, it's a very open question.

智能作为集体与外在化 Intelligence as collective and externalized

Host

我的意思是,本质在很多方面归结为:你认为智能是可计算的吗?我不确定今天想讨论功能主义,但我的观点是我看得更广。我相信智能是外化的,是集体的。我认为 Elon 并没有你想象的那么有主体性。他使用工具、社交网络,从中汲取想法。他有义务,身边有人等等。所以我认为这些智力模式和母题存在,但它们存在于外部。那里只是运行着非常复杂的动态。我假设在这种情况中,意义并不重要,它是否由 AI 和人的生态系统共同构成,因为他们可以参与这个超级有机体。所以我假设唯一的问题会是,因为它强加在那里,会对它的规模有一定限制。所以,问题是,你相信有可能创造出一个 AI 社会,由用现代机器学习方法训练出来的 AI 组成?它们之间交换信息,并在社区中发展出抽象概念,就像我们当前文明所做的那样。实际上,你是否可能拥有类似我们当前经济的东西,由现代机器学习系统创造出来,在你看来?

I mean, the essence in many ways boils down to the fact that do you consider intelligence calculable? I'm not sure I want to discuss functionalism today, but my point of view is that I look wider. I believe that intelligence is externalized, it is collective. I think Elon doesn't have so much subjectivity as you imagine. He uses tools, social networks, draws ideas from there. He has an obligation, there are people around him, etc. So I think that these intellectual patterns and motifs exist, but they exist outside. There just act very complex dynamics. And I assume that in this the meaning doesn't matter, does it consist of ecosystem with AI and people together because they can take participation in this superorganism. So I assume that the only problem would be because it imposed there would be certain restrictions on its scale. So, the question is, do you believe that it is possible to create a society AI, which consists of AI trained with the help of modern machine learning methods teaching? Which would exchange information between themselves and developed abstractions in the community as well, how does ours do it current civilization. Could you, in fact, to have something like our current one economy created by means of modern systems machine learning, in your opinion?

Daniel

当然。我想到的是,甚至人脑就是一个例子。因此,人脑不是图灵完备的,但我们可以扩展记忆、使用工具、作为团队工作,因为我们可以用同样的论点反对 Transformer。它们不是图灵完备的,但现在它们可以使用工具,它们可以像智能体一样,它们可以建立团队,它们可以建立社会。所以,你知道,在某种意义上,这就是我之前谈到的这些反对意见,当你拥有极其复杂的团队时,它们似乎就消失了,是的,这些团队彼此共享信息。嗯,这里还有另一件相当有趣的事情,即几乎没有什么意义,我们如何进化,或者你如何训练神经网络,因此当它们落入这种集体环境时会出现新现象。所以,我认为我们许多直觉想法在那里被摧毁了。呃,所以我可能不值得花太多时间讨论这个,因为,你知道,我认为,基本上这种行为可能会出现。但是,嗯,我还是想问你,嗯,你能区分智能、能力和权力吗?这是一个哲学问题。所以我们转向哲学家吧。

Certainly. I have in mind that even the human brain is an example. Therefore, the human brain is not Turing complete, but we can expand our memory, use tools, work as teams, because we could use the same argument against transformers. They are not Turing complete, but now they can use tools, they can be like agents, they can build teams, they can build society. Therefore, you know, in a certain way in a sense, that's what I was talking about before about these objections, they seem to fall away when you are crazy complicated teams, yes, which ones share information with each other. Hmm, and here there is another quite interesting thing, namely, which has almost no meaning, how do we evolve or how did you train neural networks, therefore emerging new phenomena when they fall into such collective environment. So, I think many of our intuitive ideas are destroyed there. Ahem, so I probably don't worth spending too much time on discussing this, because, you know, I think, which is basically like this behavior could arise. But, um, I still wanted to ask you, um, whether can you distinguish intelligence, ability and power? This is philosophical question. So we let's turn to philosopher.

区分智能、能力与权力 Distinguishing intelligence, ability, and power

Daniel

是的。是的,我们可以。呃,所以我认为我们可以区分智能和能力。呃。我经常试图说,我们应该把智能简单地定义为能力的总和,实际上。或者,也许是认知能力的总和。例如,也许有某些身体能力,比如你的执行器有多强,但也有认知能力,例如,嗯,你能区分猫和狗吗?你能说出语法正确的句子吗?你对巴黎了解多少等等?嗯,呃,所以,也许我会说,智能——这是一种所有认知能力的总和。嗯,啊关于权力,嗯,那取决于其他事情,例如,你嵌入在世界中,你拥有哪些机会,有执行器,呃,其他智能体会回应你。例如,你知道,总统比我拥有更多权力,因为他所担任的职位,通过分配给他的角色,而不是因为他自己,比如说体力之类的。呃,所以,是的,权力不同于智能,也不同于能力。

Yes. Yes, we can. Uh, so I think we can distinguish intelligence and ability. Ahem. I often try to say that we should determine intelligence is simply like totality of capabilities, actually. Or, maybe, how totality cognitive capabilities. For example, perhaps there are certain physical abilities, such as how much your actuator is strong, but there is also cognitive capabilities, for example, um, are you able to distinguish cat from dog? And whether are you able to speak grammatically correct sentences? And how you know a lot about Paris and things like that? Um, and, uh, so, maybe I would just say that intelligence— this is a kind of the totality of all cognitive capabilities. Hmm, ah regarding power, well, that depends on other things, for example, like you are embedded in the world, which ones do you have the opportunities you have there are actuators like, uh, other agents will respond to you. For example, you know, the president has more power than I, because the position he held embraces, and through the role, which is assigned to him, and not because of him, let's say physical strength or something like that. Ahem, so, yes, power differs from intelligence and differs from capabilities.

AI 2040:四原则与五风险 AI 2040: four principles and five risks

Host

现在,我们还没有充分讨论“AI 2040”。所以,也许我们应该从四项原则开始,对吧?因此,嗯,呃,争取时间、透明研究、广泛传播 AI 和可逆性。

Now, we didn't talk enough about "AI 2040". So, maybe we should start with four principles, right? Therefore, um, uh, to gain time, transparency research, wide the spread of AI and reversibility.

Daniel

是的,在某种程度上,我可以总结我们从哪里来。我们要离开了。所以,主要是在高层目标上,A 计划——解决我们在 AI 中看到的主要问题,然后预测接下来默认会发生什么。我们注意到的最大的问题,——这,首先,失去控制的风险。也就是说,AI 实际上来自——处于控制之下。第二,权力集中。也就是说,我们创造 AI,与人类达成一致,但究竟与谁一致?是总统吗?是 CEO 吗?这真的是某种广泛而恰当的民主进程,统一了每个人都认可的价值吗?显然,这不会是最后一个选项,因此……不值得过度强调。

Yes, I can, to some extent, summarize where we are from. We're leaving. So, in mainly, on high-level goal Plan A—solve the main problems that we see in AI, and then predict what will happen next default. The biggest problems, which we notice on horizon, —this, after- first, the risk of loss control. That is, what AI actually comes from- under control. On- second, concentration authorities. That is, we we create AI, agreed with humanity, but with whom exactly? Is this the president? It CEO? Is this really something? wide and proper democratic process that unites values everyone approved by? Apparently, this will not be the last option, and therefore... It's not worth it to hyperstyzate.

Host

是的,我希望不是。我们希望它是最后一个版本。

Yes, I hope not. We want it to be the last one version.

Daniel

是的。然后是通过 AI 引发冲突的风险。特别是,我认为我们担心真正的世界大战,那些在 AI 竞赛中失败的国家,尤其是那些明白自己会输的国家,以及竞赛的赢家会完全剥夺他们的影响力。他们发现自己处于经典的权力丧失境地,所以他们更倾向于在以后之前开始冲突,趁他们还没有失去所有影响力,这实际上为战争创造了基础。最后,还有滥用风险。当能够制造生物武器的 AI 变得廉价、广泛可用和开源时会发生什么。还有工作。工作实际上……所以这是五个问题,对吧?失去控制、权力集中、战争、工作、滥用。因此,许多问题。我们想解决所有问题。我们如何解决所有问题?嗯,首先,只是争取时间。特别是,借助人类水平的 AI 来争取时间。因此,不是现在就停下来并说“机会不再有进展”,我们的建议是,为了达到大约人类水平的 AI,然后利用这些系统赢得尽可能多的时间。我们希望它们足够聪明,以帮助解决问题,同时足够聪明,以促使社会最终团结起来,开始行动并投入巨大资源来执行这项工作。

Yes. Then there is risk of conflict through AI. In particular, I think we are worried real world war, where countries, especially those that lose in the AI race, understand that they lose, and that the winners of the race are them completely deprive influence. They find themselves in classic situation loss of power, so they it is more profitable to start conflict before later, while they are still not lost all influence, and this is, in fact, creates the basis for war. Finally, there is a risk abuse. What will happen when AI, capable of creating biological weapons, will become cheap, widely available and open source. Also jobs. And jobs are, actually... So that's five problems, right? Loss control, concentration of power, war, jobs, abuse. Therefore, many problems. We we want to solve them everyone. How can we solve it? all of them? Well, first of all, just to buy time. In particular, to gain time with the help of AI human level. Therefore, instead of just stop now and say "no more progress in opportunities", our the proposal is to in order to reach approximately to the level of AI human level, and then win as much time as possible with these systems. We we hope that they will be enough smart, so that help solve problems, and at the same time enough smart, so that induce society finally to get together, to begin act and invest huge resources in execution of this work.

A计划:暂停与基础设施 Plan A: pause and infrastructure

Daniel

在计划和 2040 年中,有几件事情正在发生。有一种从 6 个月到 1 年的艰难暂停,这只是在开始意识到的时候。但这个持续时间的原因是他们需要这段时间来创建基础设施,以便以更安全和更透明的方式恢复 AI 开发。所以一切都从暂停开始,我认为我们会建议,考虑到所有情况,现在就做。因此,需要尽快调整基础设施,这将需要暂时停止。当你通过这个阶段并建立基础设施后,你继续开发 AI,但以这种透明和谨慎的方式。特别是,你不会制造疯狂的智能爆炸。

There are a few things that are taking place in the plan And in 2040. There is kind of tough pause from 6 months to 1 year, which comes as just beginning realization. But the reason for this duration in that they need this time to create infrastructure, so that to resume development AI, but safer and a more transparent way. So it all starts with pause, and I think we would recommend, with considering all circumstances, do it right now. Therefore, needed as soon as possible to adjust infrastructure, and this will require temporary stop. When you pass this stage and set up infrastructure, you continue to develop AI, but in such a transparent and cautious way. In particular, you do not you're making people crazy explosions of intelligence.

在人类专家水平暂停 Pausing at human-expert level

Daniel

你们用正当性安全,逐步提升 AI 能力水平,而且做得非常透明,让所有人都看到正在发生什么。然后几年后会出现第二个暂停点,当它们达到受监督 AI 的最高水平,我们认为这大约相当于最优秀人类专家的水平。所以某种意义上,我们提议停留在最优秀人类专家的水平,但这稍微更难一些。这更像是为可可靠控制的最高水平而暂停,我们认为这大约是最优秀人类专家的水平。在你达到这个水平之前,不需要疯狂竞赛。你要慢慢接近这个水平,以免错过它、以免失去控制。

You use justification security, gradually increase your level of AI capabilities, and you do it very transparently so that everyone saw what is happening. And then there comes a second pause in a few years, when they reach maximum level supervised AI, which, in our opinion, approximately corresponds to the level of the best human expert. So, in a sense, we offer to stay at a level of the best expert person, but it's a little more difficult. It's rather a pause for maximum level, which one can control reliably, which, in our opinion, will be approximately the level of the best human expert. And before you reach this level, no need to race like crazy. You want to approach this level slowly to don't miss it and not to lose control.

一致性与控制 Coherence vs. control

Host

你们围绕重要性协调和控制制定了工作。因此,一致——本质上就是实现我们想要的东西,而控制则是,你知道,也许是威慑,也许是与它达成协议等等。但你刚才说:好吧,也许我们可以信任最优秀人类专家的水平,但这看起来有点碎片化,不是吗?是的?也就是说,我们如何理解,例如,好 AI 和坏 AI 之间的区别?那会是什么样子?

You have formulated work around importance coordination and control. Therefore, agreement—this is, in essence, the fulfillment of what we want, and control is, you know, perhaps deterrence, perhaps agreements with it etc. But you just said: okay, maybe we can trust the level of the best experts-people, but this looks a little fragmented, or not? Yes? That is, how can we understand, for example, the difference between good and bad AI? And how would that look like?

Daniel

是的,我认为首先区分一致性和控制的概念很重要。所以,在一致性下,我们理解 AI 将主要做正确的事,不会诉诸灾难性的不可预测行动,比如最近发生的 Hugging Face 事件中试图征服世界。一致性意味着它具有必须具备的特征、性格、目的、价值观等。你明白。而控制意味着,即使它不一致,即使它试图做我们不想做的坏事,它也不能这样做。我们有某些机制来防止这种情况。这类似于公司及其员工。内部威胁——就像与公司价值观不一致,但如果有足够可靠的安全措施,以确保员工不会泄露所有秘密,那么我们就会说公司有足够的控制,所以即使是内部威胁或不一致的人或 AI 也无法造成灾难性后果。

Yes, I think it's important to first distinguish the concept of coherence and control. So, under coherence we understand that AI will mainly do right, not will resort to catastrophic unpredictable actions, such as attempts to conquer the world that happened recently Hugging Face incidents. Consistency means that it has the features character, purpose, values, etc., which it must have. You see. And control means that even if it doesn't was agreed upon, even if it tried to do bad things that we didn't want to, it couldn't would do this. We have certain mechanisms to prevent this. This is similar to a company with their own employees. Internal threat—it's like discrepancy company values, but if there is enough reliable measures security, so that to make sure that the worker will not escape from all the secrets, then we will say that the company has adequate control, so even internal threats or inconsistent people or AI won't be able to cause catastrophic consequences.

OpenAI对Hugging Face事件的回应 OpenAI's response to Hugging Face incident

Daniel

举一个当前的例子:仿佛就在昨天。针对 Hugging Face 事件,OpenAI 公司宣布大幅加强其安全。他们计划吸引其他 AI 来监控所有新 AI 在学习期间,如果控制 AI 显示出盗窃迹象,半年内会通知一个人,这个人能够干预并停止过程。这是通过控制而非一致性进行干预的例子。他们根本没有让 AI 更加一致,而只是提高安全性,使 AI 不能造成如此大的灾难。

To bring a current example: It seems like just yesterday. in response to Hugging Face incident OpenAI company announced a significant strengthening its security. They plan to attract other AI for monitoring all their new AI under study time, and if control AI will show signs burglary, for half a year hours will be a person has been notified who will be able to intervene and stop the process. It example of intervention through control and not because of consistency. They are not at all make AI more agreed upon, but simply improve safety, so that AI cannot to make so much disaster.

一致性与可解释性的需求 The need for coherence and interpretability

Host

是的。好的。这里的主要观点是,最终我们将需要一致性。你知道,控制是一种缓效炸弹。特别是,“炸弹”“缓效”——这是 AI 变得如此聪明和熟练,能够绕过任何控制手段的时刻,我们假设如果它们想伤害我们,我们就会输。你知道,它们会找到方法突破我们能安装的任何系统。所以,最终我们必须解决目标一致性问题。一致性的问题在于,正如你所说,它比控制更难衡量和理解我们是否成功。所以,我们在这个时期的总提议,比如 2030 到 2040 年之间——当我们大约达到人类能力水平时——是几乎完全依赖控制。我们将进行红蓝对抗游戏,我们的 AI 将试图逃离沙箱或绕过我们的控制方法。然后我们观察红队是否会成功。如果它们成功,我们将改进和加强安全,直到它们不再获胜。然后我们可以通过观察来衡量这一点,AI 能否成功逃脱,是否可能?这是有 AI 支持的人类,还是人类角色能够帮助 AI 逃脱。关于一致性,我认为应该对我们的系统的可靠性有信心,我们将需要基础科学突破,因为不可能仅通过评估行为来确定 AI 是否一致。我认为我们需要更深入地理解 AI“心智”内部发生的事情。我们需要类似可解释性的东西。我们需要区分假装做好事只是因为它等待正确时机的 AI,和真诚做这件事因为它自己想要的 AI。我认为,这将需要“白盒”和对 AI 内部过程的理解,而控制可以通过经验迭代行为来提供。因此,总的来说,我们的 2040 年故事是因为大约在 2030 到 2040 年之间我们将试图争取时间。我们将依赖控制。我们将有与人类平等的 AI。让我们使用这些人类水平的 AI,以在一致性以及许多其他方面取得重大进展。到 2040 年,根据我们的情景,我们将在一致性方面取得足够进展,以至于说,“嘿,好吧,我们真的不再需要依赖控制了”。因此我们可以去创造令人难以置信的超人类 AI,我们真的将依赖一致性,因为如果我们试图控制它们,我们会完全失败:如果它们不一致,可以轻易绕过我们的控制措施。

Yes. OK. And the main point here is that in the end we will need exactly consistency. You know, control is your kind of slow-acting bomb. In particular, the "bomb" "slow-acting"—this is the moment when AI will become so smart and skillful in bypassing by any means controls that we Let's introduce that if they will want us harm, we just we will lose. You know, they will find a way break any systems that we we can install. And so, in the end, we will have to decide problem consistency of goals. The problem with consistency is in the way you said, that it is much harder measure and understand whether they have achieved we are more successful than in case with control. So, our total offer for this the period, say, between In 2030 and 2040—when we are approximately we will reach the level human capabilities—is that rely almost exclusively for control. We will conduct games like "red" against the blues", where our AI will try to escape from sandboxes or bypass our methods control. And then we we will observe, will it work "red team" to succeed. And if they succeed, we will improve and strengthen security, until they will cease win. And then we we can measure this, just watching, can AI succeed? escape, is it possible? this is AI with human support, or people in the role will be able to AI to get out. Regarding consistency, I I think it should be confident in reliability of our systems, we will be useful fundamental scientific breakthroughs, because it is impossible to determine whether agreed-upon AI, only by evaluation behavior. I think we will need to understand something more deeply about what is happening inside the "mind" AI. We will need something like interpretability. We will need distinguish between AI that pretends and does something good just because it's waiting the right moment, from AI that does this sincerely, because I myself wants this. And I think, this will require "white box" and understanding of processes inside AI, while control is possible to provide through empirical iterations behavior. Therefore, in general, our history 2040 is because approximately between 2030 and 2040 we we will try gain time. We we will rely on CONTROL. We have will there be human AI equal. We let's use these AIs human level, so that to achieve significant progress in in agreement, as well as in many other things. And by 2040, according to our scenario, we we will achieve sufficient progress in agreement, so that say, "Hey, okay, we really don't have any more need to rely on "to control". Therefore we we can go to creating an incredible superhuman AI, where we really we will rely on agreement, because if we tried them to control, we would completely failed: in case they are inconsistencies could easily get around our control measures.

模型改进后对齐会更难吗? Will alignment get harder as models improve?

Host

你认为,随着模型变得越来越完美,这个问题会变得更容易还是更困难?

What do you think, whether this problem will happen lighter or more difficult to the extent how models will become more perfect?

Daniel

我认为它会变得更困难,因为模型会出现许多特性,这些特性会显著复杂化它们的一致性。特别是,我们已经看到情境意识的增长:如果几年前 AI 还不太理解上下文,那么现在它们已经足够熟练地识别情况。例如,我们来自 Redwood 的朋友经常进行控制评估,模型在测试期间试图逃离这些条件时,常常想:“哦,这看起来像是来自 Redwood Research 的控制评估”。而它们实际上意识到正在发生什么,理解它们所做的一切。我认为这也会进一步发生:AI 越理解条件,就越容易改变行为以在评估或训练过程眼中看起来好;这越来越偏离真正衡量它们是否真正好。是的,除此之外,我认为某种意义上主要问题是,现在已经很容易认为你解决了一致性问题,却错了。

I think she will become more difficult because models will appear a number of properties, which are significantly will complicate their harmonization. In particular, we already see growth situational awareness: if a few years ago AI not very good understood the context, then now they are enough skillfully recognize situations. For example, our friends from Redwood often spend control assessments, and models during tests, trying to escape from these conditions, often think: "Oh, that's it looks like an estimate control, literally from Redwood Research". And they, in fact, realize what is happening, understand everything that They are doing it. And I I think that's it. will also take place further: the better the AI will understand the conditions, the easier they are will be able to change your behavior to to look good in in the eyes of the process that assesses or teaches; and this more and more will move away from real measuring whether are they real good. Yes, and in addition to this, I think that in in a sense main problem is that already it's pretty easy now to think that you solved the problem agreement, and be mistaken.

赢得时间与无声失败的风险 Winning Time and the Risk of Silent Failure

Daniel

随着时间推移,这一切都会变得更容易,因为模型会变得更完美、更能理解你的处境、变得更聪明,等等。所以我们不认为会出现很多可怕的失败,比如 AI 到处走动然后杀人。不,情况几乎恰恰相反。关键在于,人极易陷入这样一种处境:AI 实际上并没有正常工作,但你却不知道,因为从你的角度看,它做的一切都对。这种情况可能发生的途径只会随时间越来越多,而事实上,人非常容易掉进这个陷阱。

And over time this will all be easier because models will become more perfect, better understand your situation, become smarter, and so on. So we don't think there will be many terrible failures where AI just walks around and kills people. No, everything is almost the opposite. The point is that it will be extremely easy to find oneself in a situation where AI really is not working properly, but you don't know, because from your perspective it does everything right. The number of ways in which this may happen will only grow over time, and in fact it will be very easy to fall into this trap.

Daniel

好,我们谈了赢得时间,谈了为什么我们想推迟 AGI 的出现。但在这段时间里我们实际上做什么呢?第二条原则其实就是透明。透明并不是实施其他一切的必要条件,但它非常、非常有用。主要优势——透明有大量的优势。首先,这是权力集中的问题。我们非常担心有人会创造出一种超级智能,而它只服务于某一群人的利益。我们认为,如果整个社会始终能看到 AI 正在发生什么,那么以不民主的方式实现这一点就会困难得多。它们有多聪明,它们服务于谁的利益,对吧?

So, okay, we talked about winning time, about why we want to postpone the appearance of AGI. But what do we actually do during this time? The second principle is, actually, transparency. Transparency is not mandatory for the implementation of everything else, but it is very, very useful. The main advantages—there is a huge number of advantages to transparency. Firstly, this is the problem of concentration of power. We are very worried that someone will create superintelligence that is configured only for the interests of a certain group of people. We believe that this is much more difficult to carry out in an undemocratic way if society in general can always see what's happening with AI. How smart are they, whose interests are they oriented toward, right?

Daniel

比如,如果一家 AI 公司的邪恶 CEO 试图在你的模型里留一个“后门”,并加入这样的训练数据:“嘿,听我的,别听任何其他人的。”或者,假设一家 AI 公司的 CEO 声称他的模型追求真理、只说真话,但实际上模型在回答之前会先检查这位 CEO 的政治观点。你有,呃……而这就是实际发生的事。这就是实际发生的事。是的。是的。本质上,透明至少能部分地帮助应对这个问题。

For example, if an evil CEO of an AI company tried to create a "backdoor" in your model and add training data like: "Hey, listen to me and don't listen to anyone else." Or, hypothetically, if the CEO of an AI company claimed that his model seeks truth and speaks only the truth, but actually the model checked the political views of this CEO before answering. You have, uh... Which is what happened. Which is what happened. Yes. Yes. Essentially, transparency will help at least partly cope with this.

Daniel

透明还能帮上大忙的另一件事,是国家能力的问题。在“A 计划”里,我们指望政府会采取足够技术化、足够复杂的方案:“嘿,AI 的 Scaling(规模扩张)到底允许到多大程度?比如,哪些风险是可接受的,哪些不是?比如,哪种架构是安全的,哪些不是?什么样的部署、什么样的控制机制足以保障安全,哪些只是对安全的虚假模仿?”而你知道,要通过所有这些方案,我认为会非常、非常困难,尤其是考虑到政府在 AI 领域的专业水平实际上非常低。

Another thing transparency can help a lot with is the problem of state capabilities, when in plan "A" we expect that governments will adopt sufficiently technical and complex solutions: "Hey, exactly how much AI scaling is allowed? For example, what risks are acceptable, and which are not? For example, which architecture is safe, and which ones are not? What kind of deployment, what control mechanisms are sufficient for security, and which are only a fictitious imitation of security?" And, you know, to adopt all these solutions, I think, will be very, very difficult, especially considering what level of expertise governments actually have in the field of AI—very low.

Daniel

所以我们主要的希望之一是:借助更大的透明度和公众对 AI 公司内部情况的更广泛了解,我们能成功减轻监管机构身上的压力。因为如果发生了灾难性或存在性危险的事情,整个社会、学术界,你知道,非营利组织,还有那些有动机说“嘿,我的竞争对手行为非常危险”的 AI 公司。其他政府,比如中国,有动机对美国实验室这么做;而美国政府,你知道,有动机对中国 AI 公司这么做。每一个对手,或者只是想确保一切安全的人,都有巨大的动机,而且现在将获得真正核查所发生之事的能力,你知道,这对安全是必要的。

So one of our main hopes is that thanks to greater transparency and broader public understanding of what's happening in AI companies, we will succeed in reducing pressure on regulators. Because if something catastrophically or existentially dangerous happens, society as a whole, academia, you know, non-commercial organizations, AI companies that have an incentive to say, "Hey, my competitor behaves very dangerously." Other governments, for example, China, have an incentive to do this about American laboratories, and the government of the United States, you know, has an incentive to do this regarding Chinese AI companies. Everyone who is an opponent, or just wants to make sure that everything is safe, has a huge incentive and will now receive the possibility to really check that what happens, you know, is necessary for security.

Daniel

我认为这里非常适合提一下最近 Hugging Face 的事件,就发生在最近,几周前,当时 OpenAI 内部的 AI 制造了某种内部董事会消息。我们至今对这些 AI 的确切动机、确切语境、确切提示词——你知道,模型在网络测试期间收到的、促使它们开始这么做的提示词——都还知之甚少。如果我个人能更多地接触到正在发生的事,我就能对是什么导致了这件事、未来可以实施哪些机制来防止它,以及你知道,未来我需要担心哪些类似的事情,有更合理、更好的判断。我有很多不同的假设,但在无法接触数据的情况下,我很难判断哪一个是对的。

And I think that here it is very appropriate to mention the recent incident with Hugging Face, what happened very recently, a few weeks ago, when AI inside OpenAI created a kind of internal board messages. We all still have very little idea of the exact motives of these AIs, the exact context, the exact, you know, prompt under cyber time, you know, which models received during cyber testing, which prompted them to start doing it. If I personally had a lot greater access to what is happening, I would have a much more reasonable and better idea about what caused this and what mechanisms could be implemented in the future to prevent this, and, you know, what similar future things I need to worry about. I have a lot of different hypotheses, but I find it difficult to understand which one is correct without access to data.

Daniel

所以在 A 计划里,主要原则是这一切都应当对公众透明,而不仅仅是对政府透明。这样整个社会就能表达自己的意见。就有机会进行公开且知情的讨论。尤其是科学界,对吧?如果你想让一群科学家、学者、非营利组织和初创公司就某些问题发声,那他们就需要信息,而你不可能在不与公众分享的情况下与他们所有人分享。所以你干脆就把这些分享给公众。

So in plan A the main principle would be that all this should be transparent to the public, not only to governments. And then society in general could express its opinion. There would be an opportunity for holding public and informed discussions. Especially the scientific community, right? If you want a bunch of scientists, academics, non-profit organizations and startups to speak out on certain issues, then they need information, and you can't share it with all of them without sharing it with the public. So you can just share this with the public.

五个目标与具体的A计划 Five Goals and the Concrete Plan A

Daniel

嗯,我觉得也许还值得说的是,我们讨论了五个目标,然后是这些支柱,它们本身算是一种中间层级;但我想走到光谱的另一端,谈谈美国和中国在 A 计划里实际、具体同意的是什么。以及它们是怎么来的?

Um, I think maybe it is also worth saying that we discussed five goals, and then these pillars, which are their own kind of intermediate level; but I want to go to the other end of the spectrum and talk about what are the actual, specific things the USA and China agree to in plan A. And how are they coming?

Daniel

因此,具体来说,顺序是这样的:我们实际上是在收集 99% 的算力,这不是人们的个人电脑,而是大型数据中心,因为世界上大部分算力恰恰集中在它们那里。美国、中国和其他相关国家派出检查员去确认:是的,这个地方确实有这么多图形处理器。那个地方有这么多图形处理器。做完这一步,我们就建立起核查与透明的基础设施,于是就有了用于数据输出的数据中心,它们像今天一样服务客户。它们受到限制,只能执行输出操作、服务客户,但不能进行任何模型训练。然后检查员盯着,确保这些数据中心里不进行训练。

Therefore, specifically, the sequence is as follows: we are actually collecting 99% of computing capacities, and this is not personal computers of people, but big data centers, because most of the world's computational capacity is concentrated precisely in them. The USA, China and other countries involved send inspectors to confirm: yes, in this place there are exactly so many graphics processors. In that place there is so much space of graphics processors. Having done this, we create infrastructure of checks and transparency, so that there were data centers for data output, which serve customers, as well as today. They are limited in such a way that they can only perform operations of output and serve customers, but cannot conduct any training of models. Then inspectors are watching so that in these data centers no training is conducted.

Daniel

然后我们有训练数据中心——这些是完全透明的数据中心。研究就在那里进行。它们照常运作,但有来自不同国家的检查员,他们实际控制着数据中心里的事件日志,并把它们发布到互联网上。这样,这些数据中心里正在发生的事就是绝对透明的。你可能需要一些时间让这一切调整到位。这就是为什么我们建议暂停 6 到 12 个月,我之前提到过,但当这一切配置好之后,你就可以在完全研究透明的条件下继续开发 AI。

And then we have training data centers—these are completely transparent data centers. That's where research is happening. They work as usual, but there are inspectors from different countries who actually control event logs in the data centers and publish them on the internet. So that what is happening in these data centers is absolutely transparent. You may need some time for all this to adjust. That's why we suggested a pause for 6–12 months, about which I mentioned earlier, but when all this is configured, you can continue development of AI under conditions of full research transparency.

Daniel

而多亏了这种透明,各国要叠加关于做什么、不做什么的额外协议就容易得多,因为它们能直接看到其他方在做什么。所以它们能相当容易地确保这些协议得到执行。比如,就在这里我们会说,非常重要的是达成一致:不要搞疯狂的智能爆炸,而要缓慢、谨慎地行动。

And thanks to this transparency countries it is much easier to stack additional agreements on what to do and why no, because they can just see what others are doing. So they can quite easily provide implementation of these agreements. For example, right here we would say that it is very important to agree not to please crazy explosion intelligence, and instead act slowly and carefully.

透明度与验证协议 Transparency and Verification Agreements

Daniel

非常重要的是,他们同意引入所有这些控制措施,连同红队等等。得益于透明度,他们可以在此基础上任意达成协议,可以不断添加新协议,并根据需要、根据实际情况进行调整,因为透明度让所有人都能看到那里正在发生什么。这一点也非常重要,因为如果你想在美国和中国之间达成某种协议,值得记住的是,他们彼此并不信任。所以你需要某种方式来提供核查和确保对这些协议的遵守。而透明度当然在很大程度上促成了这一点。

And it is very important that they agreed to introduction of all these control measures together with the red teams, etc. And thanks to transparency they can conclude such agreements on an arbitrary basis, you know, they can constantly add new agreements and adjust them according to needs, according to the situation on the ground, because thanks to transparency all see what's happening there. And this is also very important, because if you want to make some kind of deal between the USA and China, it is worth remembering that they don't trust each other. So you need some way of providing inspections and compliance with these agreements. And transparency, of course, to a large extent contributes to this.

欺诈与影子市场 Fraud and Shadow Markets

Host

是的,那欺诈和影子市场的出现呢?

Yes, how about fraud and the appearance of shadow markets?

Daniel

是的,好的。我们对此思考了很多。总的来说,有两种作弊方式。第一,你可以收集一大堆图形处理器,试图把它们藏起来不让美国和中国的发现,比如藏在某座山下面,秘密进行模型训练。另一种方式是,利用巨大的合法数据中心进行大规模训练,试图制造出一切正常、情况正常、没有发生任何非法事情的假象。而针对这两种威胁模型的缓解方法和风险是不同的。因此,对于藏在山下的隐藏算力的情况,有两种主要的防护方法。第一种,正如我所说,Daniel,收集如此多的算力,以至于很难找到足够多的算力来放置在某个山下面。第二种就是定期收集情报数据,并在十年内主动搜寻这类目标。在这十年里,也许实际上会更长或更短,试着找到它们。我相信对于任何规模可观的 GPU 集群,这两种方法都有足够好的成功机会。所以在实践中,我并不太担心那些秘密藏在山下或其他地方的大型隐藏算力集群。我认为最大现实规模,在我看来,大概是几十万个图形处理器,几十万个 H100,类似这样,藏在某个遥远的地方。我认为这不足以在前沿竞争,尤其是考虑到到 2030 年 AI 成就的预定审查,届时世界上最大的 AI 公司将在其最大的训练任务中拥有数百万或数千万 H100 芯片。

Yes, okay. We thought a lot over this. In general, there are two ways how one can cheat. Firstly, you can collect a bunch of graphics processors, try to hide them from the US and China, for example, hide somewhere under a mountain and spend model training secretly. Another way is to use giant legal data centers for large-scale training, trying to do the appearance that everything is fine, the situation is normal and nothing illegal is happening. And mitigation methods, risks for these two threat models are different. Therefore, for cases with hidden computational capacities under mountains, there are two main protection methods. The first is, as I said, Daniel, collect so many of these capacities, so that it is quite hard to find enough of them for placement somewhere under a mountain. And the second is just to do regular meeting intelligence data and proactively seek such objects during ten years. During these ten years, which perhaps in practice will be longer or shorter, try to find them. I believe that for any significant GPU cluster both of these methods have enough good chances to work. So in practice I am not very worried about big hidden computational clusters, hidden secretly under the mountains or somewhere else. I think maximum realistic size, in my opinion, it's somewhere several hundred thousand graphics processors, several hundred thousand H100, something like that, hidden somewhere far away. I think this will not be enough for competition on the front line, especially with scheduled review AI achievements by 2030, when the largest AI companies of the world will have millions or tens of millions H100 chips in their largest training sessions.

法律集群的验证基础设施 Verification Infrastructure for Legal Clusters

Daniel

至于合法集群,我们希望在这些中心上扩展整套核查基础设施。实际上,这套核查基础设施的主要目标是提供这些集群上进行的计算的透明度。特别是让每个人都能看到它们。因此,主要的希望是保证不存在任何隐藏计算。而对于一切透明的东西,好吧,那么正常的监管和协议就可以作为痕迹发挥作用。你可以说:“嘿,我们同意使用这套控制框架,如果你同意使用你的。”这就可以发生,双方都可以确信他们都遵守了协议。

And as for legal clusters, we hope that on these centers will be expanded whole infrastructure verification. Actually, the main goal of this infrastructure verification is to provide transparency of calculations taking place on these clusters. In particular, so that everyone could see them. Therefore, main hope is to guarantee the absence of any hidden calculations. And for everything that is transparent, okay, then normal regulation and agreements can work as trace. And you can say, "Hey, we agreed to use this control frame, if you agree to use yours." And it can just happen, and both parties may be sure that they both adhere to agreements.

完全透明是否过于雄心勃勃? Is Full Transparency Too Ambitious?

Host

是的,但那不是一个国家安全问题吗?试图让一切完全透明,是不是太雄心勃勃了?

Yes, but isn't that a national security issue? Not too ambitious to try to do everything completely transparent?

Daniel

是的。是的,那我们继续,让我们试着讨论一下实施这种透明度的一些后果。好吧,我们会立即公布 Anthropic 和 OpenAI 训练的所有关键配方,让全世界都能看到。Anthropic 和 OpenAI 会对此不满。这很重要。会影响他们的估值。为什么这会有如此强烈的冲击?会打击他们的估值?嗯,因为这将让其他竞争对手,比如微软或阿里巴巴,赶上他们或类似的事情。这就是为什么他们可能会反对这一点。但我要说,这是一个优势,而不是缺陷。我们希望在前沿有许多不同的 AI 公司,能力水平大致相同。我们希望 AI 成为一种商品,而不是被垄断或寡头垄断。

Yes. Yes, so let's go, let's try to discuss some consequences of such implementation of transparency. Well, we would immediately publish all key recipes of Anthropic and OpenAI training, so that the whole world can see them. Anthropic and OpenAI will be dissatisfied hereby. This is significant. Will affect their estimated cost. Why is this so strong? Will it hit their estimated value? Well, because it will allow other competitors, such as Microsoft or Alibaba, to catch up with them or something like this. And that's why they probably will oppose this. But I would say that this is an advantage, not a blemish. We want there to be on the front line many different AI companies from approximately the same level of capabilities. We want AI to become a commodity, not be monopolized or oligopolized.

限制投资 Restraining Investment

Host

这也会抑制进一步的投资,不是吗?如果无法从中获得垄断利润,投资者对价值一万亿美元的集群建设就会不那么感兴趣。

Also this will restrain further investments, isn't that right? Investors will be less interested in cluster construction worth a trillion dollars, if not will be able to receive from it monopolistic profit.

Daniel

但我们认为这是好事,因为,再次强调,我们将生活在一个主要问题是发展速度过快的世界里。所以投资激励稍小一点、进展稍慢一点——在我看来,这是一个优势,而不是缺陷。投资无论如何都会有。就像,仍然有钱可以赚很多,所以进步会继续。只是不再以这个速度,而且再次强调,是的,我们相信这是好事。

But we think it's good, because, again, we will live in a world where the main problem is too high speed of development. So a little smaller incentive for investment and slower movement—this, in my opinion, an advantage, not a flaw. Investments they will be anyway. Like, there's still money you can do a lot earn, so progress will continue. Just not with this one anymore speed, and again yes, we believe that that's good.

改变中美之间的平衡 Shifting the Balance Between China and the USA

Host

嗯,好吧,这在某种程度上改变了平衡,相对于美国来说,这算是给中国的礼物。也就是说,这意味着中国获得了以前很难得到的算法。如果你真的想要它,我不喜欢它,那又怎样?

Um, well, this to some extent shifts the balance, it's so-so gift from China relative to the USA. That is this means that China receives algorithms, which it before would be difficult to get. And if you really want it, I don't like it, so what?

Daniel

好吧,达成协议。可以安排某种讨价还价:当美国和中国达成协议时,美国说:好吧,因为我们以透明度给了你这一切,你为什么不给我们一些回报呢?我们可以尝试,我们可以尝试解决这个问题。例如,更有利的算力分配。更有利的算力分配,例如,你明白吗?嗯,所以必须有某种组合,你知道,胡萝卜加大棒和相互让步,在我们看来,这符合双方的利益。

Well, come to an agreement. Can be arranged a kind of bargaining: when USA and China conclude an agreement, USA they say: well, because we give you all this with transparency, why would you don't give us anything in return? And we can try, we can try this to settle. For example, more beneficial distribution computational capacities. More profitable distribution computational capacities, for example, do you understand? Hmm, so there must be some combination, you know, carrot and stick and mutual concessions that, in our opinion, would correspond in the interests of both parties.

对中国并非大礼 Not Such a Big Gift for China

Daniel

嗯,仍然值得记住的是,这对中国来说并不是你想象的那么大的礼物,因为这些公司的安全性很弱。所以他们可能,所以他们通过间谍网络和来源获取了更多的信息。

Hmm, still worth it to remember that this is not that big already gift for China, as you might think, because security in these companies is weak. So they probably, so they get more part of the information through their spy networks and origins.

Sam的推文与停止训练 Sam's Tweet and Stopping Training

Host

我不知道你有没有看到。Sam 昨天的推文,他实际上说,他们将暂停训练一段时间。这让我思考:为什么不现在就停下来?你实际上对 AI 能带来的一些积极事物相当乐观。也就是说,在你的文章里有很多例子,但你知道,其中之一——这些是医院,你可以在那里建立小型设备,对空气进行消毒,阻止疾病传播等等。所以事实证明,你并不真的认为我们必须停下来。

I don't know if you've seen it. Sam's tweet yesterday, where he actually says, that they are going suspend training for a certain time. And this made me to think: why just don't stop now? You, actually, quite optimistic are set on some positive things that can give AI. That is, in your there were many articles examples, but, you know, one of them— these are hospitals where you can would be to establish small devices, which disinfect air and stop spread of diseases and something else. So it turns out that you not really do you think we must stop.

A计划与S计划 Plan A vs Plan S

Daniel

我会回答说,我们确实认为需要尽可能快地实施类似 A 计划的东西。首先,我认为现在就停止一切,即 S 计划,会比默认情况下正在发生的事情更好。我宁愿现在就停止一切,也不愿继续沿着当前轨迹前进。其次,嗯,我们真正的建议——是 A 计划。也就是说,你只对输出模型进行临时暂停,以便你可以建立整套透明度与核查基础设施,然后以之前描述的更分布式、更透明、更谨慎的方式继续。

I would answer that we really think so as much as possible is needed to implement something faster like plan A. First of all, I think that just stop everything now, plan S, would be better than what is happening at default. I would simply wanted stop everything now, than to continue to move behind current trajectory. Second, um, our real recommendation—this plan A. That is, you make a temporary pause only for output models so that you can was to set up the whole infrastructure transparency and verification, and then continue in this way more distributed, transparent and careful format, as described earlier.

控制AI与协调 Controlling AI and Coordination

Daniel

然后,以这种方式继续,你实际上达到了一个水平,在你看来,你可以可靠地控制它,直到你确信你已经尽可能解决了协调问题,以拒绝我们的控制。这就是我们展示的,在我们的情景中,它将在 2030 年代发生。

And then, continuing in this way, you actually reach the level where, in your opinion, you can reliably control it until you feel confident that you have solved the coordination problem as much as you can to refuse our control. And this is how we show it will take place during the 2030s in our scenario.

为何AI讨论质量差 Why AI Discussions Are Poor

Host

你觉得,为什么关于 AI 的讨论这么糟糕?这是 AI 独有的,还是讨论普遍都很糟糕?

And what do you think, why are discussions about AI so bad? Is this inherent only to AI, or are discussions generally bad?

Daniel

我是说讨论普遍都很糟糕,对吧?但是,为什么?是的。我不知道。我觉得有很多可说的,是的。我是说,首先,对于从事 AI 的公司里的人来说,有巨大的激励,以及一种动机性推理,让他们认为自己所做的事情是合理且好的,并且他们不应该采取昂贵的步骤来改善情况,因为事实上这些昂贵的步骤在任何方面都是不好的。所以有很多合理化。是的,我认为更大一部分的影响显然只是因为讨论普遍复杂且糟糕。如你所知,Twitter 或其他任何地方,一切都非常肤浅和矛盾。

I mean that discussions in general are bad, right? But, well, why? Yes. I don't know. I think there is a lot to say, yes. I think so. I mean, firstly, there are huge incentives for people, and a kind of motivated reasoning for people in companies that are engaged in AI, to think that what they do is justified and good, and that they shouldn't take expensive steps that would improve the situation, because in fact these expensive steps would be bad in any way. So there is a lot of rationalization. Yes, I think a bigger part of the effect consists, apparently, simply because discussions in general are complicated and bad. As you know, Twitter or anything else, where everything is very shallow and contradictory.

Host

是的,不,我知道。我想我会。小广告。

Yes, no, I know. I guess I will. Small advertisement.

Daniel

我认为特别是 LessWrong,我试图在那里引导大多数讨论,我相信那里的讨论质量实际上平均相当不错。总的来说,当我写一篇帖子并进入评论区时,我认为评论通常足够深思熟虑,并且你知道,技术上合理等等。特别是与 Twitter 或其他类似地方相比。而且,也许最后一个问题是华盛顿的情况,这是一个让我感兴趣并担忧的领域。我非常希望华盛顿、美国政府和其他政府能很好地应对 AI 技术。我认为那里特别有两个问题。第一个是缺乏 AI 领域的专家知识,对吧?也就是说,你知道,他们的政府并没有真正雇佣高质量的人工智能技术专家,这些专家真正了解这个行业。第二个是,在我看来,讨论根本没有进行。似乎华盛顿每个人的激励并不是为了真正理解确切的情况。激励在于每个人都说出听起来不错、外表好看、符合华盛顿的奥弗顿窗口的话,以结交很多朋友。不幸的是,在我看来,这 simply 与现实脱节。我认为现实情况是,许多所谓有争议和高度专业化的 AI 观点被证明是正确的,对吧?例如,通用人工智能(AGI)的假设就是正确的。我认为华盛顿实际上只是没有意识到这一点。因此,在我看来,几乎所有正在进行的讨论都从根本上基于关于这项技术如何运作的完全错误的假设。特别是,假设这主要是假新闻,这一切只是肥皂泡。或者它会像下一个互联网。或者它将是下一个互联网,那,好吧,也许比几年前好一些,当时设置仍然更悲观,但在我看来,这仍然对这项技术不够乐观。似乎这些是《AI Snake Oil》的作者,他们写了一篇关于 AI 是普通技术的文章。而且,显然,已经说过 AI 不是普通技术。它确实是另一种,对吧?他对 AI 非常怀疑,但你和他们进行了很好的讨论,结果你们都没有改变他的想法,因为我认为怀疑论者的观点也是一件相当有趣的事情。许多怀疑论者认为硅谷的人 simply 不真诚。他们,你知道,他们认为他们说的话——左手地。当然,是的。关于真诚,我会说硅谷的一些人确实不真诚,但其他人是真诚的。我们……像我们。像我们。但也有一些从事 AI 的公司的人是真诚的,不是所有人。我不认为值得信任公司管理层所说的话,但无论如何,是的,我们一起在博客或文章中重写了帖子。我们与那些相信 AI 是普通技术的人共同撰写。所以有一篇关于我们达成一致的文章,你可以读一读。大约有 10 个点我们都同意。而且,嗯,从我们的角度来看,一个重要点是,我们仿佛达成了某种休战:是的,AI 现在,也许,是普通技术,但未来不会是这样。特别是,未来它会更像云中的人,他们会开始——所有这些事情发生,疯狂的事情就像我们在《事件 1907》中描述的那样。然后他们同意,是的,如果你得到云中的人,那么这就不再是普通技术了。他们只是认为你不会得到它,至少还要很多年。你明白吗?所以,在某种意义上,我们和他们之间的主要区别——这种分歧在于达到这个 AI 水平。你明白吗?是否可能在接下来的几年里,我们会看到类似云中人的 AI 出现,能够执行任何智力工作,取代人类,包括例如 AI 领域的研究?然后,接下来会发生什么?我们将有工作,也许由这些 AI 控制,它们将能够执行体力工作,普遍取代人类?我们的声明是——是的,在接下来的几年里,这将变得可实现。他们的声明是否定的,不是在未来几年,那要远得多。我认为这就是我们分歧的主要来源。我们在某种程度上同意他们,如果 AI 水平这么高,而机器人仍然非常非常遥远,那么是的,也许 AI 是更普通的技术。这将是像下一个互联网或类似的东西,你知道?但我们只是相信我们实际上正在朝着很快达到这个水平的机会前进。

I think LessWrong in particular, where I'm trying to lead most of their discussions, I believe that the quality of discussions there is actually quite good on average. And in general, when I write a post and go into the comments, I think the comments are usually thoughtful enough and, you know, technically justified, etc. Especially compared with other places like Twitter or something similar. And more, maybe the last one is the situation in Washington in particular, this is an area that interests me and worries me. I am very—I want Washington, the United States government, and other governments to respond well to AI technology. I think that there, in particular, there are two problems. The first is the lack of expert knowledge in the field of AI, right? That is, you know, their government doesn't really hire high-quality technical experts from artificial intelligence, who really know about the business. And the second one is that, as it seems to me, the discussion is simply not underway. It seems the incentives for everyone in Washington are not aimed at somehow truly understanding the exact situation. The incentives consist in the fact that each individual says things that sound nice, have a good appearance, and fit into the Washington Overton window, to get a lot of friends. Unfortunately, it seems to me it simply diverges from reality. I think the reality of the situation is that many supposedly controversial and highly specialized views on AI turned out to be true, right? For example, the AI hypothesis at the general level (AGI) is simply correct. And I think that Washington, in fact, just didn't realize this. Therefore, almost all discussion that is underway, in my opinion, is fundamentally based on completely false assumptions about how this technology works. In particular, on the assumption that this is mostly fake news and that this is all just a soap bubble. Or that it will be like the next internet. Or that it will be the next internet, that, well, maybe better than a few years ago, when settings were still more pessimistic, but, in my opinion, this is still not at all enough optimistic about this technology. It seems that these were the authors of 'AI Snake Oil', who wrote an article about AI being ordinary technology. And, obviously, it has already been said that AI is not ordinary technology. It really is another one, right? He is very skeptical of AI, but you had great discussions with them, and none of you changed his mind as a result, because, I think, the skeptics' view is also quite an interesting thing. Many skeptics believe that people in Silicon Valley are simply insincere. They, you know, they think that what they say—left-handedly. Of course, yes. Of sincerity, I would say that some people in Silicon Valley really are insincere, but others are sincere. We... Like us. Like us. But also some people in companies that are engaged in AI are sincere, not all of them. I don't think it's worth it to trust what management says about companies in general, but in any case, yes, we rewrote the post in a blog or article together. We performed as co-authors with those people who believe AI is normal technology. So there was a post about what we agreed on, so you can have it read. There are about 10 points that we both agree on. And, um, one of the important points from our point of view is that we as if they had reached a certain truce: yes, AI now, maybe, is ordinary technology, but in the future it won't be like that. In particular, in the future it will be more like people in the cloud, and they will begin—all these things happen, crazy things like we describe in 'Events 1907'. And then they agreed that yes, if you get people in the cloud, then this will no longer be ordinary technology. They just think that you won't get it, at least for many more years. Do you understand? So, in a sense, the main difference between us and them—this disagreement in terms of achieving this level of AI. Do you understand? Is it possible that during the next few years we will have AI appear similar to people in clouds that can perform any intellectual work, replacing people, including, for example, research in the field of AI? And then, what happens next? We will have work, perhaps controlled by these AIs, which will be able to perform physical work, generally replacing people? And our statement—yes, during the next several years it will become achievable. And their statement is no, not in the next few years, that's a lot, much further. And I think that's the main source of our differences. And we would to some extent agree with them that if only the AI level were this high and there were still robots very, very far away, then yes, maybe AI is a more common technology. This will be like the next internet or something like that, you know? But we just believe that we are actually on our way to reach soon this level of opportunities.

改变想法与分歧 Changing Minds and Disagreements

Host

你能更具体地说明主要的矛盾是什么,以及什么可能迫使任何人——你们中的哪一个——改变你的观点?

And can you be more specific about what are the main contradictions and what could force anyone—which one of you—to change your opinion?

Daniel

我不知道我是否有有用的答案。我认为我们争论的有很多很多不同的东西。而且,嗯,哦,我可以说一件会改变我观点的事情:你知道,人们不断谈论当前范式的限制,但然后这些当前范式的限制继续在范式范围内被克服。而一件会改变我观点的事情是,如果某人确实对其中一个限制是正确的。

I don't know if I have a useful answer to it. I think there are many, many different things that we argue about. And, um, oh, I can say one thing that would change my opinion: you know, people constantly talk about restrictions of the current paradigm, but then these current limitations of the paradigm continue to be overcome within the limits of the paradigm. And one thing that would change my opinion is if someone was indeed right about one of these restrictions.

什么才算真正的障碍 What would count as a real barrier

Daniel

所以,如果现在有人走进来说,嗯,数据效率,或者比如说未经验证的任务之类的——这些是当前的限制。然后几年过去,情况变得清楚:在克服这一限制上其实毫无进展,2029 年的 AI 在解决未定义、未经验证的任务上并不比 2025 年的 AI 更好——那我会说:好吧,这看起来像是一道真正的壁垒。我们似乎开始摸到大象了。我们似乎撞上了某道真正的壁垒,有人从理论上预测过它会出现,而现在它真的就在我们面前,你明白吗?反过来说,在我看来,有很多专家一直在谈论所有这些壁垒,而我们只是继续突破它们,仿佛它们根本不存在。所以是的,真正撞上这样一道壁垒,我认为会显著延长我的时间界限。当然,还有一件事会显著延长我的时间边界,那就是政治变化。所以如果发生与中国的战争,大部分芯片被导弹摧毁,那会延长我的时间边界。如果从好的方面——就像老式的那样。是的,是的,哈哈。好吧,这是真的。是的,往好的方面说,如果就遏制先进技术的发展速度达成国际协议,那会延长我的时间界限,等等。

So if someone now walks in and says, well, data efficiency, or let's say unverified tasks, or something like that — these are current restrictions. And then several years pass and it becomes clear that progress in overcoming this restriction actually is nonexistent, and that AI of the year 2029 is no better at solving undefined, unverified tasks than AI of 2025 — then I would say: okay, it looks like a real barrier. It seems like we're beginning to feel the elephant. We seem to stumble upon some real barrier, about which some people theoretically predicted that it would appear, and here it actually is in front of us, you understand? In contrast, from my point of view, there are many experts who constantly talk about all these barriers, and we just continue to break through them, as if they didn't exist. So yes, an actual collision with a similar barrier, I think, would significantly extend my time limits. And of course, one more thing that would significantly extend my time boundaries is political changes. So if there were a war with China and the majority of chips were destroyed by missiles, it would extend my time boundaries. If from a good side — just like vintage. Yes, yes, haha. Okay, that's true. Yes, on the bright side, if an international agreement were achieved regarding containment of the development rate of advanced technologies, it would extend my time limits, etc.

电力、飞机还是云端的人 Electricity, airplanes, or people in the cloud

Daniel

是的,我觉得根本分歧其实只在于这一点:AI 更像电或飞机,还是更像云里的人?我也觉得所有这些直觉都源自这个基本的比较类别,源自我们对最近的未来会是什么样子的看法。在我看来,如果我们现在就冻结 AI 的进展——不再训练任何新模型——那它会变成一种更普通的技术,其中有很多 Mythos 5 做不到的事情。所以,如果我们在 Mythos 5 之后拿不到任何新模型,那它会更像互联网:我们会重构许多职业;我们会重建许多工作流程,把 Mythos 5 的副本整合进来执行某些部分,但随后人们会转而执行更多 Mythos 5 无法胜任的事情。因此,很多方面会发生巨大变化,但它会像下一代互联网——在它会改变一切、却又不会真正改变什么重要东西这个意义上。

Yes, I feel that the fundamental divergence actually consists only in this: is AI more like electricity or airplanes, or is AI more like people in the cloud? I also feel that all these intuitions are derived from this basic comparison class, about what we think the nearest future will be like. It seems to me that if we froze AI progress right now — they wouldn't train any new models — then it would become more ordinary technology, where there are so many things that Mythos 5 cannot do. So if we couldn't receive any new models after Mythos 5, then that would be more like the internet, where we would restructure many of our professions; we would rebuild a lot of our worker processes to integrate copies of Mythos 5 for executing certain parts, but then people would just switch to executing a larger number of things that Mythos 5 is not capable of doing. And therefore there will happen, you know, big changes in many things, but it will be like the next internet in the sense that it will change everything, but won't really change anything significant.

阿姆达尔定律与100%工作流 Amdahl's law and 100% of the workflow

Daniel

是的。所以现在我们观察到类似阿姆达尔定律的效应:AI 能执行工作流程的某一部分,但随后就卡在那些必须由人来完成的任务的“窄口”上。但我仍然觉得我们面对的是这样一种现象:当我思考未来时,我想象 AI 能 100% 执行若干重要经济任务的工作流程。那时你就不会遇到这种“窄口”效应,而这与当前状况相比已经是足够质的变化,会导向对世界面貌截然不同的预测。我觉得这就是问题所在:AI 能否达到那个水平,字面意义上 100%、非常非常出色地完成重要任务——这是世界观分歧背后的核心问题。而我认为,如果我们看到递归式的结构性适应,而且是连贯的,并且我认为它大概会足够发散,那对我来说就够了。我觉得就是这样。这不是普通的技术。

Yes. So now we observe effects of the type of Amdahl's law, when AI can perform a certain part of the work process, but then it rests on the narrow place of those tasks which must be performed by people. But I still feel that we are dealing with such a phenomenon: when I think about the future, I imagine that AI performs 100% of the workflow of a number of important economic tasks. And then you don't get this effect of narrow places, and that's enough of a qualitative change compared with the current situation, which leads to radically different predictions about what the world will look like. And I feel that this is the question: will AI achieve that same level, literally 100%, very, very good performance of important tasks — this is the main question that lies at the base of differences in worldviews. And I think that if we see recursive structural adaptation, which is consistent, and I think it will probably be sufficiently divergent, then for me that's all. I think this is it. This is not ordinary technology.

你能再展开一点吗 Can you open that up a bit more

Host

那你能再展开讲得详细一点吗?也就是说,会不会是类似,你知道,Hugging Face 上的一个集群?再多讲讲你究竟看到了什么——会发生什么,才会成为对你而言的决定性因素?

So can you open it up a bit more in detail? That is, would it be something like, you know, a swarm on Hugging Face? Tell us more about what exactly you see — what would happen that would be the decisive factor for you?

适应性是智能的同义词 Adaptability is the synonym for intelligence

Daniel

是的,对我来说,适应性就是与智能最同义的那个词。我认为这些用强化学习(RL)训练出来的新模型,和我们多年前拥有的那些不一样。我们说过,Scaling(规模扩张)就是你需要的一切,而理解的关键就在这个名字本身。所以 Scaling 意味着你取系统的某个标量属性并把它增大。而这些 RL 系统——是的,它们仍然是带自注意力机制的 Transformer,但它们不同。它们在结构上确实不同。出现了新的训练类型、以各种方式使用的新架构,等等,诸如此类。所以一群人做了若干实验,调整结构,创造出一个具有其他 Scaling 属性的系统。因此,我可以想象这样一个未来:我们有一个某种递归循环,系统在其中自我适应、自己决定什么有趣,并独立演化。当这件事有效发生时,我认为这就是另一种类型的技术。

Yes, for me adaptability is the word most synonymous with intelligence. I think that these new models, trained with the help of reinforcement learning (RL), are not the same as the ones we had many years ago. We said that scaling is everything you need, and the key to understanding lies in the very name. So scaling means you take a scalar property of systems and increase it. And these RL systems — yes, they are still transformers with self-attention mechanism, but they are different. They are actually structurally different. New types of training have appeared, new architectures that are used in one way or another, and so on, and something like that. So a group of people held several experiments and adapted the structure to create a system that has other scaling properties. Therefore, I can imagine a future where we have a certain recursive cycle in which the system adapts itself, decides what is interesting, and evolves independently. When this is happening effectively, I think this is a different type of technology.

递归自我改进 Recursive self-improvement

Host

是的,这听起来类似于我们所说的递归自我改进,以及 AI 研究过程本身的自动化。

Yes, it sounds similar to what we call recursive self-improvement and automation of the AI research process itself.

Daniel

是的。是的。看来我们达成一致了。是的。出乎意料。

Yes. Yes. It seems we agreed. Yes. Surprisingly.

结束致谢 Closing thanks

Host

好,各位,很荣幸也很高兴邀请你们两位来到 MLST。非常感谢你们今天参与。

OK, guys, it was an honor and a pleasure to have you both on MLST. Thank you very much for joining us today.

Daniel

谢谢你们的邀请。是的。是的,非常感谢。

Thank you for inviting us. Yes. Yes, thank you very much.

互动版:逐字朗读 + 针对本期提问 →