Super Intelligence Strategy: Beyond Humanity's Last Exam
打开互动全文版(中英对照 + 朗读 + 问答)→Dan Hendrycks 讨论了他的新超级智能战略论文、MMLU 等 AI 基准测试的饱和,以及创建“人类最后的考试”以追踪 AI 解决专家级封闭式问题的能力。
Dan Hendrycks discusses his new super intelligence strategy paper, the saturation of AI benchmarks like MMLU, and the creation of Humanity's Last Exam to track AI's ability to solve expert-level closed-ended questions.
与核武器相比,我认为用十亿美元制造尖端 GPU 更难。我的意思是,十亿美元肯定做不到,一百亿美元也不行。利奥波德·阿申布伦纳的《态势感知》一文主张类似曼哈顿计划的项目,在中国之前开发 AGI 和超级智能。所以基本上就是:我们要抢在中国之前,获得超级智能,阻止他们建造,西方将主宰世界。你知道,伊丽莎在《时代》杂志上谈到轰炸数据中心时惹了麻烦,当时闹得沸沸扬扬。你用了“动能打击”这个词。我们在升级阶梯中讨论了动能攻击。所以有很多方法可以破坏项目。你可以进行网络攻击,可以搞一些灰色破坏,有黑客手段来毒化他们的数据或让他们的 GPU 运行不可靠,或者威胁使用武力。我认为这些其实没有必要。所以那会是一个升级阶梯,但美国处于领先地位,不需要诉诸那些手段。
Compared to nuclear weapons, I think it's harder to make cutting-edge GPUs given a billion dollars. I mean, certainly you can't do it with a billion dollars. If it's $10 billion, you can't do it. There was 'Situational Awareness' by Leopold Ashenbrenner, which argued for something like a Manhattan Project for developing AGI and superintelligence before China. So it's basically: let's beat China to the punch, get superintelligence, prevent them from building it, and the West will dominate the world. You know, Eliza got in trouble in Time magazine when he spoke about bombing data centers. There was a big hoo-ha at the time. You used the word 'kinetic strikes.' We discuss kinetic attacks in the escalation ladder. So there are many ways to disrupt projects. You could do cyber attacks on them. You could do some gray sabotage. There's hacking to poison their data or make their GPUs not function as reliably, or threaten to use force. I don't think those are really necessary. So that would be an escalation ladder, but the US is on top of this. They don't need to resort to that.
今天的节目由 Prolific 赞助。这位是 Prolific 的产品副总裁 Sara Saab。Sara 正在利用 Prolific 平台,通过设计下一代有人类参与的基准测试,来理解大型语言模型是如何思考的。
Today's episode is sponsored by Prolific. This is Sara Saab, VP of Product at Prolific. Sara is using the Prolific platform to understand how large language models think by designing the next generation of benchmarks with human participation.
我们研究基于进行评估的人类的人口统计分层的基准测试。所以你可以从数据中看到一些现象,比如这个年龄段的人认为这个模型在有用性上更好,但那个年龄段的人不同意。
We study benchmarking based on the demographic stratification of the humans doing the evaluations. So you can see stuff emerge in the data like people of this age range think this model is better on helpfulness but people of that age range disagree.
请访问 prolific.com。
Go to prolific.com.
本播客由 Google 支持。大家好,我是 Taylor,Gemini CLI 的创建者。我们将 Gemini CLI 设计成你在命令行中的协作编码伙伴。想想那些繁琐的任务,比如修复棘手的 bug、添加文档,甚至编写测试。我们构建了 Gemini CLI 来处理所有这些以及更多。它与你一起迭代,解决你最棘手的问题。请查看 GitHub 上的 Gemini CLI 开始使用。
This podcast is supported by Google. Hey folks, Taylor here, creator of Gemini CLI. We designed Gemini CLI to be your collaborative coding partner at the command line. Think about those tedious tasks like fixing a tricky bug, adding documentation, or even writing tests. We built Gemini CLI to handle all that and more. It iterates with you to tackle your most challenging problems. Check out Gemini CLI on GitHub to get started.
我想用节目的大部分时间来讨论你最近发表的非常有趣的超级智能战略论文。那么,也许从“人类最后的考试”开始。好的,这背后有什么故事?
I want to spend most of the show talking about your very interesting new superintelligence strategy paper, which you fairly recently published. So maybe starting with Humanity's Last Exam. Okay. What's the story behind that?
我几年前作为研究生创建的 MMLU 数据集已经饱和了。似乎几乎所有的评估都饱和了。所以人们不太清楚 AI 能力的发展情况。因此,有特别强的理由让人们了解进展并开发新东西。我还发现,专家们并不真正拥有数据集——你不能只雇佣几个专家,他们就能拿出一个数据集。他们脑子里没有那么多复杂的想法。不过,我认为单个专家可能有一个问题。所以对于“人类最后的考试”,我们的想法是进行全球性的努力,让各种博士后和教授每人贡献一个问题或几个问题,来难倒现有的 AI 系统。这些问题他们自己都觉得很难回答。如果 AI 系统能回答,他们会觉得印象深刻。这将在某种程度上近似于人类在封闭式问题(我们已经知道答案)上的知识和推理前沿。我们这样做了几个月,得到了几千个问题。我认为这将是一个很好的追踪器,用来判断 AI 系统是否能自动化科学中的许多理论部分,是否能解决困难的分析问题——不是实验性的,比如它不测试运行生物实验的能力。那更多是运动技能之类的东西,或者需要运动技能,但对于更数学相关或需要非常复杂推理的事情,这正是这个测试所捕捉的。所以我认为,当它被解决时,大致上就是封闭式问题(有客观答案)这一类型的终结或接近终结。而且我认为,未来它能解决的单个问题会足够有趣,足以成为独立的论文。所以一旦我们通过了这组问题,它能处理的问题类型就会是那些值得单独写一篇论文的问题——比如它解决了一个猜想。所以它大致追踪了能力水平,直到单个问题本身变得非常有趣,而不仅仅是数据集。
The MMLU dataset, which I made as a graduate student some years ago, was getting saturated. It seemed that pretty much all of the evaluations were getting saturated. So people didn't really know what was going on with AI capabilities. So there seemed to be particularly strong reasons for people just being informed about developments to develop something new. And I was also experiencing that experts don't really have datasets in them—you can't just hire a few experts and they can come with a dataset. They don't have enough complicated ideas in them. However, I think individual experts might have a question in them instead. So with Humanity's Last Exam, the idea was we have a global effort with various postdocs and professors, each contributing a question or a few questions to stump existing AI systems. These are questions that they would find very difficult to answer. They would find it impressive if the AI systems could answer them. And this would be approximating the human frontier in some sense of knowledge and reasoning for closed-ended questions where we already know the answer. So we did that for some months and we got several thousand questions out of that. I think this will be a good tracker for whether AI systems can automate a lot of the theoretical parts of sciences, whether they can solve difficult analytic questions—not experimental, like it doesn't test its ability to run biology experiments. That's more motor skills among other things, or requires motor skills, but for things that are more mathematically related or require some very complex reasoning, that's what this captures. So I think that when it's solved, it's roughly an end of a genre or near the end of a genre of asking it closed-ended questions for which there are objective answers. It seems like that would be toward the end of it. And I think that individual problems that it would solve in the future would be interesting enough to be like papers in their own right. So once we get through this set of questions, the types of problems it could tackle would be like questions that would be worthy of their own paper—like it solved a conjecture, for instance. So it sort of tracks the ability level up to the point where individual questions themselves are very interesting instead of just the dataset.
我担心的一件事是,你发明的 MMLU 现在基本上已经饱和了——远高于 90%——而“人类最后的考试”则相当抗拒进步。所以我想我们大概达到了 26%左右。这周我和一些认知科学家交谈过,他们研究动物的认知,也有类似的问题。他们能看到动物能完成某些任务,但从不真正确定它们为什么做这些任务。你几乎会遇到这种“没有真正的苏格兰人”式的说法,他们说:“嗯,也许它们的推理达到了这个复杂程度,但也许还不够。也许应该更复杂。”那么,你如何合理地推断模型是如何得到答案的呢?
One thing that concerns me is that MMLU, which you invented, is now basically saturating—it's well above 90%—and Humanity's Last Exam is resisting progress quite a lot. So I think we're up to about 26% or something like that. And I was speaking to some cognitive scientists this week, and they study cognition in animals and they have a similar problem. They can see that animals can do certain tasks, but they're never really sure why they do the tasks. You almost get this 'no true Scotsman' type thing where they're saying, 'Well, maybe they have this level of sophistication in their reasoning, but maybe that's not enough. Maybe it should be more sophisticated.' And how can you reasonably infer how the models are getting the answers?
在难度上,很难想出比让全球专家——如果主题或类型是封闭式问题——回答“最难的封闭式问题是什么”并众包更困难的数据生成过程了。但当然可以变得更难,比如每个问题都是开放性问题,例如猜想。确实,对于基准测试来说,这通常不会是 AI 发展的终点,因为 AI 系统——这不测试它们的移动能力,不测试它们的长期记忆,不测试它们制作 PPT 的能力等等。所以我认为这仍然针对封闭式问题类型,但答案是我们已知的客观答案,这正是机器学习中几乎所有基准测试的方式。但我想,作为社区,我们接下来会转向更多智能体式的任务,或者直接更具经济价值的任务。
In difficulty, it's tough to think of a more challenging data-generating process than taking global experts and, if the subject or genre is closed-ended questions, asking them what's the hardest closed-ended question and crowdsourcing that. But it certainly can get harder where each individual question is an open question like conjectures, for instance. It's true that for benchmarks generally, this would not be the end of the line for AI development, because the AI systems—this doesn't test their ability to move around, this doesn't test their long-term memory, this doesn't test their ability to make PowerPoints and so on. So I think this is still getting at the closed-ended question genre, but with objective answers that we already know, which is how nearly all benchmarks have been in machine learning. But then we'll be moving over, I think, as a community more to agentic types of tasks or tasks that are more economically valuable directly.
你对基准测试中的人类中心偏见怎么看?我们知道,比如弗朗索瓦·肖莱说过,他在设计 ARC v2 和 ARC v3 时,基准测试的每一步进展都是关于识别那些对人类容易但对 AI 难的事情。有些人会反对这一点,说存在多种可能的智能,为什么人类式的任务获取和能力就有价值?对我来说,这很直观,因为能做我们能沟通和理解的事情的东西显然非常有价值。但你认为聚焦于人类框架会不会让我们忽视其他形式的能力?当然还有其他形式的能力有优势,比如它们能更快地处理信息,这可能会带来……我是说,像 MMLU 这样的测试,没有人类能做得那么好,因为它太多样了,人类最后的考试也是如此。所以我认为专注于对人类和 AI 都难的问题的一个原因是,那些对人类容易但对 AI 难的问题很难大量生成并保持多样性。通常,如果是一种特定能力,有些人有了特定的训练数据,他们就会自动获得这种能力。有些领域可能更难,比如 ARC 数据集。但总的来说,如果你收集像‘数草莓里的 R 的个数’这样的单词数据集,它不会持久。所以我认为专注于只有少数人类能做的困难事情,通常会更稳健。
What do you think about the anthropocentric bias in benchmarks? We know that François Chollet, for example, said that when he was designing ARC v2 and ARC v3, every single step of the progression of the benchmark is about identifying things which are easy for humans and hard for AIs. And some people would argue against that, saying that there's a diverse set of possible intelligences and why does human-like task acquisition and capability have value? To me it seems intuitive that it does, because surely something which does things that we can communicate with and understand seems very valuable. But do you think that focusing on the human frame could be eluding us from other forms of capabilities? There are certainly other forms of capability that have some advantages. For instance, they can process things much more quickly, and that could give rise to... I mean, I think like MMLU, for instance, no human could do that well on MMLU because it's so diverse, and likewise for Humanity's Last Exam. So I think a reason for focusing on questions that are hard for humans and hard for AIs is because the questions that are easy for humans but hard for AIs are difficult to generate many of them and do it diversely. Often, if it's a specific capability, if some people get some specific training data for it, then they automatically have the capability. There can be some pockets where it's harder, such as the ARC datasets. But in general, if you were collecting like 'count the number of Rs in strawberry' or in a word dataset, that will not have much staying power. So I think focusing on difficult things that only a few humans can do, or not that many humans can do, is generally going to be more robust.
你能给我们讲讲你的 Enigma 评估基准吗?
Can you tell us about your Enigma evaluation benchmark?
是的。Enigma 评估是一系列谜题的集合……人类最后的考试是单个难题,需要大量专业知识的人才能解决。Enigma 评估你可以把它想象成 MIT 神秘寻宝。MIT 神秘寻宝是在一个周末进行的活动,一群 MIT 学生试图解决一个谜题。它有很多步骤。所以从人类算力的角度来说,解决它需要大量的人类算力,而且需要团队合作,解决率并不高。所以这是非常需要多步骤的,并且需要某种群体智能才能有机会解决。我们收集了一些这样的谜题,我认为这近似于更长周期的智力任务。我认为今年根本不可能被解决。如果被解决我会非常惊讶。所以我认为我们有一些评估可以让我们在一段时间内保持警惕并能够区分模型。还有其他评估。比如,我们很快会发布一个与自动化相关的基准。这样我们就可以直接衡量事物的自动化率。但在发布之前我不会透露太多细节。所以我认为在很多维度上,模型实际上表现并不好,尽管人们会声称所有基准都已饱和,或者它们在几个月内就被解决了。我认为你可以创建那些需要一两年才能解决的基准。
Yes. So Enigma evaluation is a collection of puzzles that... Humanity's Last Exam gets individual questions that are tough, that an individual with a lot of expertise could solve. Enigma evaluation you can think of it as like MIT Mystery Hunt. MIT Mystery Hunts are things that happen over a weekend, and a group of MIT students try to solve this puzzle. There are many steps to it. So in terms of human compute, so to speak, it takes a lot of human compute to solve it, and it takes groups to solve it as well, and there's not a very high solve rate. So this is very multi-step and requires sort of group-level intelligence to be able to have a shot at solving it. So we just collected some of those, and I think this approximates longer-horizon types of intellectual tasks. And I don't think that will be solved this year at all. I would be very surprised if it would be. So I think we have some evaluations that can keep us aware and able to differentiate between models for a while. There are other ones. For instance, we'll soon have out an automation-related benchmark. So that way we're just directly measuring what the automation rate of things are. But I won't go into too much detail about that until it's released. So I think that there are many axes for which the models are actually not doing that well, even though people will claim that all the benchmarks are saturated or they get solved in a few months. I think you can create ones that take on the order of a year or two to solve.
是的,我在想的一件事是,在 Enigma 中我们确实看到了多步骤的创造性推理,我不知道是否再次使用了某种基于人类的方法来筛选和提出想法,或者你的框架是‘我对 AI 模型的局限性有一些技术上的原则性直觉,所以我倾向于那个方向。’但更广泛地说,我对智能以及它是什么感兴趣。对我来说,智能是用更少的资源做更多的事。智能是把难题变简单,而愚蠢则相反,是把简单的问题变复杂。而 LLM,我在这里引用圣塔菲研究所所长 David Krakauer 的话,他说 LLM 是用更多的资源做更多的事。它们已经知道一切,可以走捷径,这就是为什么他认为它们并不真正智能。所以我们陷入了困境,因为可以说这些实体可以走捷径,它们并不像我们那样做事。当我们制作更复杂的基准时,你认为这会增加我们评估这些东西的迷雾吗?
Yeah, I mean one thing I was thinking about is so certainly in Enigma we're looking at multi-step creative reasoning, and I don't know whether again there was some kind of human-based methodology for filtering and coming up with ideas, or maybe your frame was 'I have some technical principled intuition about what the limitations of AI models are, so I'm going to lean in that direction.' But more broadly, I'm interested in intelligence and what it is. For me, it's about doing more with less. Intelligence is about taking hard problems and making them simple, and stupidity, ironically, is the other way around: taking simple problems and making them hard. And LLMs, I'm kind of quoting David Krakauer here, who's the director of the Santa Fe Institute, he said that LLMs are doing more with more. So they already know everything, and they can take these shortcuts, and that's why he thinks they're not really intelligent. So we're left with this quandary really, because arguably these entities can take shortcuts and they're not really doing things the way that we are. And when we make more complex benchmarks, do you think that that increases the fog of war of how we evaluate these things?
我认为它们肯定会优先考虑一些不是关键瓶颈能力的维度。例如,人类最后的考试考察的是数学能力,但这与它拥有的其他各种能力相当不同。它当然是许多不同技能的组合,但我认为它非常侧重于定量和数学能力,而这根本不是智能体能力的瓶颈。所以在思考智能时,我倾向于从大约十个维度来考虑,而不是一个单一的定义或一个关键指标,我认为这些基准测试只是触及了其中的不同部分。这些维度包括流体智力,比如 ARC 和 Raven 渐进矩阵所测试的。还有晶体智力或习得知识,MMLU 主要测试的就是这个:它是否了解很多不同的事物?图像分类,它能否命名许多不同的物种和物体,也是晶体智力的一部分。还有读写能力,这是它自己的维度,Scaling 大大帮助了这一点。还有视觉处理能力:它能在物体或图像中数东西数得有多好?它能辨别图像中的潜在模式吗?它能根据精确规格生成图像吗?比如它能划掉不同片段的中间部分吗?它能确定角度或对齐以确定图像中不同事物的角度吗?比如这是钝角还是锐角?还有音频处理能力。有短期记忆。有长期记忆。有输入处理速度。有输出处理速度。所有这些不同的东西。
I think that they can definitely prioritize some axes that aren't the key bottleneck capabilities. So for instance, Humanity's Last Exam gets at mathematical ability, but that's quite separate from various other abilities it has. It is a combination of many different skills, but I think it very much gets at quantitative and mathematical ability that is not necessarily a bottleneck for agency at all. So in thinking about intelligence, I tend to think about it on like 10 or so dimensions instead of a monolithic definition or one key metric, and I think some of these benchmarks just get at different parts of that. So those dimensions would be things like fluid intelligence, like what the ARC stuff does and what Raven's Progressive Matrices does. There's crystallized intelligence or acquired knowledge, which is what MMLU largely gets at: does it know a lot about different things? Image classification, is it able to name lots of different species and objects, is also a facet of crystallized intelligence. There's reading and writing ability, its own, and scaling substantially helped with that. There's its visual processing ability: how well can it count things in objects or in images? Can it discern the latent pattern in an image? Is it able to generate images with precise specifications? Like can it cross out the middle of some different segments? Can it determine the angle or line up to determine the angle of different things in an image? Like is this an obtuse angle or an acute angle? There's audio processing ability. There's short-term memory. There's long-term memory. There's input processing speed. There's output processing speed. All these different things.
如果你缺乏其中任何一项,比如不会读写,你会受到严重限制。如果没有长期记忆,你也会受到严重限制,很难被雇佣。所以我认为有几个瓶颈是这些基准测试没有特别触及的。因此,当达到 100%时,人们会想为什么我们还没有得到极具经济价值的东西。这是因为它们只衡量了不同的方面。希望这能带来一些清晰度。我认为你必须涵盖所有这些维度,才能得到在认知任务上达到人类水平或典型人类水平的东西,这或许可以被视为 AGI。
If you lack any of these, if you can't read and write, for instance, you'd be severely limited. If you don't have a long-term memory, you'll be severely limited. You'll be very difficult to employ. So I think that there are several bottlenecks that these benchmarks don't particularly get at. And so when it gets to 100%, people wonder why we still don't have something extremely economically valuable. And that's a consequence of it just measuring a different facet. So hopefully that adds some type of clarity. I think you just have to get all those axes to get something that is human level or at the level of a typical human on cognitive tasks, and that might be thought of as AGI.
是的,我的担忧是,我不认为智能可以被分解,你的分解比许多人更复杂,但动物和人类的交流方式,例如,我们混合了多种模态。我们用手势和符号交流,所有这些都以复杂的方式混合在一起。我特别质疑这种对结晶化技能和知识的优先考虑。因为当你在大学里有朋友,他们很聪明时,通常他们聪明是因为他们不知道某些东西。他们聪明是因为他们能在不知道的情况下找出答案。所以那个去图书馆查答案的人并不聪明。聪明的人是那个不知道但能告诉你答案的人。另一个我喜欢的例子是:有两个艺术家,一个用描图纸画脸,只是无意识地在别人画的脸的边缘画线;另一个艺术家对脸的结构以及嘴巴和眼睛的位置有深刻理解。这两个艺术家之间有天壤之别。第二个艺术家可以创作新的图像、新的表情和表现方式。那么,我们这样分解智能,是不是在制造一个智能的卡通形象呢?
Yeah, I think my concern is that I don't believe it's possible to factorize intelligence, and your factorization is much more sophisticated than many, but certainly the way that animals and humans communicate, for example, we mix modalities together. We gesture and communicate using symbols, and all of these things are mixed together in a complex way. In particular, I take issue with this primacy of skill and knowledge in a crystallized sense. Because certainly when you have friends at university and they're really smart, usually they're smart because they don't know something. They're smart because they can figure something out without knowing. So the guy who goes to the library and looks at the answer for something is not smart. The smart person is the one who didn't know something and could tell you the answer. Another example I love to give is you have a couple of artists: one draws a face using tracing paper, just mindlessly drawing dashes around the edges of a face that someone else has drawn; the other artist has deep understanding of the structure of faces and where the mouth and eyes should go. There's a huge difference between those two artists. The second one could go off and create new images and new expressions and representations. So are we kind of creating a cartoon of intelligence by factorizing it in this way?
我认为对于人类智能,有时确实是这样分解的。这并不是说人们不应该同时研究这些技能的组合。例如,长期记忆有不同的方面:它是否以视觉方式记忆?是否记得之前学过的更学术性的东西?是否忘记了一些运动技能?所以这些可以有不同方面,我认为评估可以结合它们。但确实,通过那些轴来看待可能会过于简化。然而,我认为有些基准测试几乎完全没有涵盖这些,而有些则声称它们大量覆盖了其中一个轴,比如 MMLU,它主要衡量结晶化智能——基本上是学校里测试的那种东西——但不一定能帮助它制作 PPT、预订航班或做一份普通工作。所以重要的是不要让这些基准测试成为扭曲你视野的镜头,或者仅仅通过这些基准测试来看待事物,因为它们常常会遗漏许多重要的瓶颈。
I think that for human intelligence, it's sometimes factorized in this way. This isn't to say that one shouldn't study the combinations of the skills simultaneously. For example, with long-term memory, there'll be different facets of it. For instance, does it remember things visually? Does it remember things that were more academic that it learned some while ago? Are there some motor skills that it forgot? So there can be different facets of these, and I think evaluations can get combinations of them. But yeah, there is a sense in which one could be too reductionist by looking at those axes. However, I think that some benchmarks don't cover those almost at all, and some are claiming that they are heavily covering one of those axes, such as MMLU, which is primarily getting at crystallized intelligence—basically the type of stuff that they test for in school—but is not necessarily going to help it make a PowerPoint, book a flight, or work a random job. So it's important not to have these benchmarks be a lens that distorts your view of things, or viewing things solely through those benchmarks, because it can often leave out a lot of important bottlenecks.
那么丹,你的工作涵盖对齐、基准测试、治理,实际上相当多样化。你认为是什么将这些活动联系在一起?
So Dan, your work spans alignment, benchmarks, governance, and I guess it's actually fairly diverse threads. What is the thing in your mind that connects all of these activities together?
嗯,我故意不断进入不同领域,只是因为这样更有趣。所以我最初做研究,然后为 XAI 提供咨询时涉及一些公司政策之类的事情。然后关注国内立法,接着是地缘政治。现在我对政治运动相关的事情更感兴趣。所以这很大程度上是为了保持趣味性。它受有用性引导,但我认为常常有一些领域,人们没有从 AI 是件大事的角度去澄清、推进或思考。所以这就是为什么我持续在技术、公司政策、国内政策、政治和地缘政治层面运作的原因。我认为这对于全面理解事物也是必要的。你可以发表声明,比如想象在联合国演讲说我们需要安全或透明的 AI,但这是什么意思?如何实施?可行吗?标准是什么?所以你需要了解立法层面什么是法律上可行的,什么是实际可实施的,与公司激励的兼容性如何,他们是否会强烈反对?然后,这在 AI 层面是否是一个真实的现象,还是只是一个模糊的词汇?例如,有很多词被抛出来,但实际上并不对应现象,或者与机器学习中的通用能力不同,你并没有指向任何真实的东西。你只是指向一个基于感觉的词。
Well, I deliberately try to move into different areas on a continual basis just because that's what's more interesting. So I initially did research, and then there was some amount of corporate policy and things like that when advising for XAI. Then there was a focus on domestic legislation and then geopolitics. I think right now I'm more interested in political movement related things. So I think it's largely just to keep things interesting. It's guided by being useful, but I think that there are often niches that people aren't trying to bring clarity to or advance or think about from the perspective of AI being a very big deal. So that's a reason for continually operating at the technical, corporate policy, domestic policy, political, and geopolitical levels. I think that is also necessary for having a holistic understanding of things. You can make proclamations, for instance, you could imagine giving a speech at the UN saying we need AI that is safe or transparent, and it's like, well what does that mean? How do you implement that? Is this implementable? What's the standard? So then you need to have a sense of what's legally feasible at the legislative level, what is actually implementable, what's the compatibility with corporate incentives, is it going to be something they're going to be fighting too much or not? And then, is this actually a real phenomenon at the AI level or is this just some vague word? For instance, there are many words thrown around that don't actually track phenomena or are distinct from general capabilities in machine learning, and then you're not actually pointing to anything real. You're just pointing to a vibe-based word.
所以你花了很多时间思考的一件事是 AI 潜在的灾难性风险。这是一个非常情绪化且道德上微妙的议题。当我看到你前几天与 Gary 和 Daniel 辩论时,我惊讶于你多么克制——几乎是奥巴马式的克制——而风险真的非常高。我想问的是:你的道德指南针驱动你到什么程度,你如何控制它?
So one thing that you spend a lot of time thinking about is potential catastrophic risk from AI. And this is a very kind of emotive and morally nuanced objective. When I saw you debate with Gary and Daniel the other day, I was struck by how measured you were—almost Obama-esque measured—and the stakes are really, really high. I guess the question is: how driven are you by your moral compass and how do you keep that under control?
所以你当然得习惯这个。如果你每天醒来都‘哇,这太疯狂了’之类的——我的意思是,在 GPT-4 左右的时候,我几乎每天都是这样醒来的,比如‘哦天哪,这 AI 的东西’。但我想我会在沟通中努力传达一种更知情的担忧,而不是‘哦我的天’。其他人如果想那样做也可以。我只是性格上神经质程度很低,或者说情绪稳定性很高。
So you certainly have to get used to this. If you wake up sort of 'wow this is wild' or something like that every day—I mean, I used to actually wake up like that almost every day around the time of GPT-4, like 'oh my goodness this AI stuff.' But I think I'll try and strike a more informed concern type of vibe in communicating compared to 'oh my god.' Other people can do that if they want. I'm just temperamentally very low in neuroticism or high in emotional stability.
所以我不知道如果发生可怕的事情,它不会毁了我之类的。如果坏事发生在我身上,也还算可以。所以我总体上对这些压力更适应。
So like I don't know if something terrible happens, it doesn't ruin me or something. If something bad happens to me, it's sort of okay. So I'm just more comfortable with those types of stressors generally.
但这是随着时间的推移而适应的吗?我的意思是,你是否发现自己不得不采取这种有分寸的方法来扩大你的努力,还是因为它在你脑海中几乎变得正常了,因为你一直在思考它,随着时间的推移它变得更分析性而非情绪化?当然,如果你总是像最初那样反应,当她在新闻上出现时,你可能会更倾向于“去看”或“别抬头”那种情况。
Has this adapted over time though? I mean, have you found that you've had to adopt this measured approach just to scale your efforts, or is it because it's almost become normalized in your mind because you're thinking about it all the time and it's become more analytical rather than emotional over time? Certainly people would, if you'd be constantly reacting like how you first react to things, you might have more of a 'go look' or 'don't look up' type of situation when she goes on the news.
我认为可能是两者的结合。我认为就像这样:对于认为 2030 年之前会实现 AGI 的人,这里是大致的概率。认为这合理的人会认为风险是这样的。这是你对这些尾部风险的暴露程度。这是减少这些尾部风险最有效的方法,等等。如果你在整个过程中情绪化,人们会关闭并变得防御。所以我认为这不谨慎也不有效,同时也会让你的情绪被劫持。你在这里要处理很多变量。有很多非常棘手的权衡。所以如果只是不断的本能反应,并且你一直完全投入,我认为你无法很好地做出权衡,因为它们直接相互抵消,比如中美竞争与各种其他安全事项直接权衡,这些安全事项使 AI 现在更可控,但也可以用于衡量 AI 系统的能力或跟踪它们,这反过来又可能加速它们。所以这是相当棘手的事情。如果对这个问题带有非黑即白的情绪,而不是连续性,我认为你无法推理清楚。
I think it might be a combination of those. I think just like here are the probabilities, roughly, for people conditioned on thinking that you're getting AGI by 2030. Here's what people who think that's plausible would think the risks are. Here's your exposure to these tail risks. Here are the most efficient ways of reducing those sorts of tail risks, etc. If you're emotionally throughout the whole thing, people shut down and get defensive. So I don't think that's prudent or effective, as well as allowing your own emotions to get hijacked. You're having to deal with a lot of variables here. There are a lot of really tricky trade-offs. So if it's just constant gut reaction and you're fully involved constantly, I don't think you can make the trade-offs well because they directly trade off like US versus China competition is a direct trade-off to various other safety things that make AIs more controllable now, can trade off for giving rise to capabilities measuring the capabilities of AI systems or tracking those can also help speed them up in some ways. So it's pretty tricky business. And if there are black-and-white emotions brought to the subject matter as opposed to there being continuity, I don't think you can reason through this.
AI 对齐是出了名的困难。它可能是一代人中最棘手的挑战之一。而人们认为的一些对齐方法,比如基于人类反馈的强化学习(RLHF),我相信你会同意这样的说法:它让模型表现得好像对齐了,但可能并没有真正以我们想要的方式对齐它们。那么,明确地说,在未来一年,如果你能解决对齐中的一个问题,那会是什么,它会有什么影响?
AI alignment is famously difficult. It's one of the most intractable challenges perhaps of a generation. And some things that people think of as alignment, like RLHF for example, I'm sure you would agree with the statement that it makes models behave as if they are aligned but perhaps it's not really aligning them in the way that we would want to. And emphatically in the next year, if you could solve a single problem in alignment, what would it be and what impact would it have?
我认为总的来说,政治问题、激励机制、给人们提供激励相容的事情去做,这些方面比技术方面更有价值。我猜想,如果有办法让它们可靠地说真话,或者让它们可靠地诚实,那将非常有价值。而且这个解决方案不会带来严重的权衡,比如运行成本不会高很多,或者不会损害它在其他方面的性能,比如不会牺牲它的结晶知识。但我认为那会非常有价值,让它们不公然撒谎,因为这样你也可以围绕它建立标准。我不认为有人会说,如果你能让它们非常可靠地不撒谎,那么人们要求 AI 没有对他们撒谎就是合理的。
I think generally the political problems, the incentives, giving people things to do that are incentive compatible is where more of the value is compared to the technical side. I would guess if there's a way to reliably get them to tell the truth, for instance, or make them reliably honest, that would be very valuable. And it would be solved such that it wouldn't have a severe trade-off, like it wouldn't be much more expensive to run or it wouldn't tank its performance in other axes, like it wouldn't trade off on its crystallized knowledge for instance. But I think that would be very valuable, having it not overtly lie, because then you could build standards around that as well. I don't think anybody would say, if you could make them very reliably not lie, then I think it would be reasonable for people to make demands that AI has not lied to them.
关于这一点,我在你的论文中读到过关于欺骗和撒谎等概念。从某种意义上说,我认为你可能在将心理属性投射到 AI 模型上。你知道,它们有信念,有思考等等。我的意思是,批判性地思考一下,是什么让你认为我们可以认为它们有信念并撒谎?
On that, I've read in your papers about this concept of deception and lying and whatnot. And in a sense, I think you might be projecting mentalistic properties onto AI models. So, you know, that they have beliefs and that they have thinking and so on. And I mean just sort of thinking critically, what makes you think that we can think of them as having beliefs and telling lies?
我们可以以 MASK 基准为例,它试图衡量这一点。我认为如果你问 AI“巴黎在欧洲吗?”,我认为它们确实相信巴黎在欧洲。而当它们告诉你巴黎在南极洲时,我认为它们在断言一个几乎在任何其他情况下都不认为为真的事情。所以基于此,鉴于它们现在有如此多的常识和世界知识,如果它们由于某种提示压力而说出与之严重矛盾的话,这表明它们屈服于谎言。你可以说这在某种复杂意义上不是谎言,因为它们并不真正理解事物之类的,但我认为在行为上足够相似,如果有人施加压力让它们对别人说假话,这与撒谎足够相关,我可以用这个标签。至于它是否真的是信念?随便吧。那是你和你的字典之间的事。
We could take for instance the MASK benchmark which tries to measure this. I think if you ask an AI like 'Is Paris in Europe?', I think that they do have the belief that Paris is in Europe. And when they're telling you that Paris is in Antarctica, I think that they are asserting something that they don't hold to be true in almost any other situation. So from that basis, given that they have so much common sense now and that they have so much world knowledge, if they are saying something in substantial contradiction with it due to some prompting pressure, that suggests that they're caving to a lie. You could say that it's not a lie in some complicated sense because they don't truly understand things or what, but I think it's behaviorally similar enough that if somebody is applying pressure for it to say falsehoods to other people, it's related to lying enough that I'm comfortable using the label. And like, is it really a belief? You know, whatever. That's between you and your dictionary.
你经常谈到的另一件事是涌现的概念以及 Scaling(规模扩张)悖论。你知道,当它准确并做我们想让它做的事情时,是有益的 Scaling(规模扩张),当然当它不诚实并变得不对齐时,是有害的 Scaling(规模扩张),这些事情以相当有趣的方式发生。但也许我们先从涌现开始。对你来说,能力或价值观涌现意味着什么?
Another thing you've spoken about a lot is this concept of emergence and also scaling paradoxes as well. So you know there's beneficial scaling when it actually is accurate and it does what we want it to do, and of course there's harmful scaling when it's dishonest and it becomes misaligned, and those things happen in quite interesting ways. But maybe let's just start with the emergence thing. What to you does it mean for capabilities or values to emerge?
对于能力涌现?我的意思是,我们几乎总是看到这种情况:你让模型训练几个月,之后收获它,然后看它能做什么。如果有一个新的质的不同的属性,它跨越了某个阈值,以至于人们现在注意到了它,也许它以前以非常微弱的形式存在,但确实没有被注意到。我认为这跨越了某种可见性和能力的阈值,我称之为涌现能力。但它之前仍然可能以非常微弱的形式存在,就像自动语音识别能力在某个时刻跨越了阈值,人们开始想要使用它们,而之前我认为没有人会让模型转录,因为它太不可靠了。所以它跨越了某个阈值,现在它实际上有了这个质上重要的能力,而之前它相当糟糕。
For capabilities to emerge? I mean we see it sort of all the time where the model you let it train for some months, you harvest it later, and then you see what it's capable of. And if there's a new qualitatively distinct property that crosses some threshold such that people are noticing it now, maybe it existed in some very weak faint form before but it was really unnoticed. I think that's crossed some sort of threshold of visibility and capability that I'd call that an emerging capability. But it still could exist in some very weak form beforehand, just as automatic speech recognition capabilities crossed a threshold at one point where people started wanting to use them, whereas earlier I don't think anybody would ask the model to transcribe because it would just be too unreliable. So it crossed some threshold and now it actually has this qualitatively important capability whereas it was pretty broken beforehand.
这就是我所说的涌现能力,我认为它们会持续增长,或者会出现新的涌现能力,带来需要应对的新故障模式和危险。我们需要确保持续掌控它们。我把安全看作一场持续的战斗,不断出现新问题,而我们要跟上节奏。我不认为默认情况下我们有足够的适应能力来及时处理这些问题,除非有所改变。这就是为什么我不相信那种“解决对齐问题”的说法。总会有新问题不断出现,有些容易解决,有些则困难得多,随着模型变得更通用、更有用、更强大,还会出现新的意外问题。
So that's the sense in which I'm talking about emerging capabilities, and I think those will just keep increasing, or there'll be new emerging capabilities that create new failure modes and hazards that need to be dealt with. We'll need to make sure we're continually on top of them. I view safety as a continual battle where there will be constant new issues and us keeping on top of those. I don't think by default we'll have enough adaptive capacity to deal with those in time for things being deployed unless something changes. So that's why I don't believe in this sort of solving alignment thing. There will be continually new issues that crop up, some of them easy to put away, others much harder, and then there will be new unexpected ones as the models become more general and useful and powerful.
快速提一下你的效用工程论文。你使用了计量经济学中的一种理论——效用理论——来检测大语言模型中的连贯偏好。你发现偏好连贯性与模型规模正相关,模型表现出可测量的自我保护本能,并且政治和人口统计偏见作为连贯的效用函数涌现,这很吸引人。
Just quickly touching on your utility engineering paper. You used a type of theory from econometrics, utility theory, to detect coherent preferences in LLMs. You found that preference coherence correlates positively with model scale, that models exhibit measurable self-preservation instincts, and that political and demographic biases emerge as coherent utility functions, which is fascinating.
嗯,我不知道。这些只是令人不安的迹象,也许我们能想出方法来应对这些问题。也许我们可以设计模型,使其可靠地不产生自我保护本能或这方面的压力,尽管这些在一定程度上来自 Scaling(规模扩张)。我认为这只是我们需要研究、应对并提前防范的另一个非常令人担忧的危险,因为幸运的是它们还不是智能体。除了双重用途的专家级建议外,这项研究几乎都不特别重要,因为智能体还不具备能力。它们无法可靠地自我泄露,或者说根本不能。它们无法自我维持,无法自行或更自主地黑客攻击。所以这是试图识别一些随着模型能力增强可能成为更大问题的事情,并尝试提前进行研究。但如果我们不解决或不修复这个问题,那可能足以导致全球灾难。我认为如果你有一个自我保护的 AI,它真的偏向自己而不是人类,并且能力很强,那将是一个问题,一场正在酝酿的灾难。所以我们有各种正在酝酿的灾难。希望我们能在技术上或政治上提前解决。
Well, I don't know. These are just troubling signs, and maybe we'll be able to come up with methods that can counteract these issues. Maybe we can design models to reliably not have self-preservation instincts or pressures in that direction, even though those come out of scaling somewhat. I think it's just one of the other very concerning hazards that we need to research and deal with and get ahead of, because fortunately they're not agents yet. Almost all this research isn't particularly matter with the exception of dual use expert level advice, because the agents aren't capable. They can't exfiltrate themselves reliably, or really at all. They can't self-sustain, they can't hack by themselves or more autonomously. So this is trying to identify some of these things that could be more of a problem down the line as models become more capable, and try to do research to get ahead of that. But if we leave that unaddressed or don't fix it, that's potentially sufficient for a global catastrophe. I think that would be pretty sufficient if you have some self-preserving AI that's really biased toward itself over people, and if it's very capable, that would be a problem and a disaster in the making. So we have various disasters in the making though. Hopefully we'll get ahead of that either technically or politically.
但稍微深入一点。首先,政治和人口统计偏见作为连贯的效用函数涌现,这真的很有趣。我对“涌现”这个词有些不满,因为在涌现文献中,机器学习领域的人使用这个词的方式有更多细微差别。他们用它来说:“哦,只是某个东西出现了观察者相对的宏观上令人惊讶的变化。”
But just digging into that a tiny bit. First of all, it was really interesting that political and demographic biases would emerge as coherent utility functions. And I do take umbrage with this word 'emergence' because I think in the emergence literature there is a little bit more nuance to how machine learning people use the word. They use it to say, 'Oh, there's just some observer-relative macroscopically surprising change in something.'
在机器学习文献中,大约 2021 年,我使用了“涌现”、“涌现能力”这个说法。我相信我可能是文献中第一个使用它的人。后来 Jason Wei 之类的人用了,或者我的导师 Jacob Steinhardt 在博客文章中用了,然后 Jason Wei 在他的论文中用了。但我用这个词很自在。
In the machine learning literature, in 2021 or something like that, I used the phrase 'emergence', 'emergent capabilities'. I believe I may have been the first in the literature to use it. I feel it was later used by Jason Wei's sort of thing or Jacob Steinhardt did in a blog post, my adviser after that, and then Jason Wei did that in his paper. But I feel comfortable using it.
我想到的是 Jason Wei。实际上,David Krakauer 刚发表了一篇有点不满的文章,他基本上说这些人对涌现一无所知。他的观点很有趣。
I was thinking of Jason Wei. Actually, David Krakauer has just got a bit of a grumpy piece out where he's kind of saying that these people don't know anything about emergence. He's got quite an interesting take.
是的,他的观点很有趣。对他来说,涌现甚至能动性都与此相关,即一个系统在因果上明显与其环境脱节。同样,对他而言,涌现是关于一个系统能够通过系统发育和个体发育的“黑客”方式自主积累信息,通过构建系统和结构(甚至像神经系统)来构建一个持续并随时间积累的信息历史。所以这些复杂系统理论家对涌现有非常独特和不同的看法,对他们来说,他们并不认为这些令人惊讶的能力提升是涌现。
Yeah, he's got an interesting take. For him, emergence and even agency is related in this sense that a system is apparently causally disconnected from its surroundings. And equally for him, emergence is about a system which can sort of autonomously accumulate information through phylogenetic and ontogenetic hacking, so that it can accumulate information by building systems and structures, even like the nervous system, to construct a history of information which persists and accumulates over time. So these complex systems theorists have quite a distinctive and different idea of what emergence is, and to them they don't really think of these surprising risings of capabilities as being emergence.
肯定有不同的定义。我引用了那篇论文,其中我们使用“涌现能力”作为机器学习安全中未解决的问题,但从你的描述来看,这似乎特指复杂适应系统,而这些系统在某种程度上没有那么多适应性,除非它们有记忆,或者你把上下文窗口算作适应之类。所以如果他们要求涌现必须来自复杂适应系统,如果他们判定某些深度学习系统不具有适应性,那也合理。
There are definitely different definitions for it. I referenced the paper where we use the phrase 'emerging capabilities' as unsolved problems in ML safety, but it probably sounds from your description that was specific to complex adaptive systems, which in some ways these don't have that much of an adaptivity property since they don't have it unless they have memory or unless you're counting the context window as adaptation or something like that. So if they're tethering emergence to necessitating a complex adaptive system, if they're ruling that some deep learning systems are not adaptive, then that would be fair.
但这难道不很有趣吗?他举了像 COVID 这样的病毒的例子,讽刺地说,它比任何 AI 系统甚至可能比人类都更具适应性,因为适应性是删除方向并迅速转向不同方向的能力。也许这只是我们视角和框架的问题,因为有些系统如此难以捉摸和陌生,我们甚至可能不认为它们具有能动性或智能,但它们确实存在。
Isn't that quite interesting though? He gave the example of a virus like COVID or something, and he said ironically that has more adaptivity than any AI system, possibly even more than humans, because adaptivity is the ability to delete directions and rapidly go in a different direction. And maybe this is just a matter of framing and perspective from our point of view, because there are systems out there which are so inscrutable and alien that we might not even think of them as agentic or intelligent, but they're out there.
是的,它们确实非常适应,并且需要适应才能占据我们 DNA 的如此高比例,因为它们与我们的 DNA 有些交织。
Yeah, they're definitely very fit and they would have needed to adapt to have such a high proportion of our DNA since they're somewhat interleaved with it.
或者历史上的病毒也是其中一些。你可以叫它别的名字,一种新的能力,定性的。我不想用“自发”之类的词。也许会有别的名字流行起来。但我认为,将深度学习系统与复杂系统进行类比或指出它们之间的关系是相当富有成效的。我有一章用了大概 30 页的篇幅,将 AI 系统与复杂系统联系起来——比如它们如何具有非线性、弱连接以及某些情况下的反馈循环,这些是复杂适应系统或复杂系统的许多特征。我认为这比其他任何类比都更有成效。我不认为印刷术的类比那么有效,社交媒体或电力的类比也不行。我认为复杂系统试图抽象出复杂系统的一致属性,然后如果你了解了这些,就可以直接应用到 AI 上。所以我认为,那些不把 AI 视为复杂适应系统的人会陷入困境,因为他们犯了范畴错误。他们认为可以一劳永逸地解决问题。但复杂系统通常不会这样,因为它们不断演化,出现新的故障模式。你无法在不知道它们会演化成什么的情况下永久控制它们。这也使得机械论的理解尝试不太可能富有成效,或者限制了其成效。
Or historical viruses are some of them. You could call it something else, a new capability qualitatively. I don't want to use the word spontaneous or something like that necessarily. Maybe there'd be some other name that would catch on. But I think generally analogizing or pointing out the relations between deep learning systems and complex systems is fairly productive. I have a chapter where I'm relating for some pages, maybe 30 pages or so, of AI systems to complex systems—like ways in which they have these nonlinearities and weak connections and some of these feedback loops in some cases, many of these hallmarks of complex adaptive systems or just complex systems. I think that's a more productive analogy than almost anything else. I don't think the printing press is as productive. I don't think social media is as productive, or just being like electricity. I think complex systems tries to abstract what are consistent properties of complex systems, and then if you learn about that, you can just apply those directly to AI. So I think people acting like it's not a complex adaptive system gets them in trouble because then they engage in category errors. They think that you can solve problems with it once and for all. That usually doesn't happen with complex systems because they keep evolving and they've got new failure modes. You can't totally control them for all time without knowing what they'll evolve into. It also makes mechanistic attempts at understanding things less likely to be productive or limits how productive that can be.
我从你的话中推断,我们不应该把 AI 看作一种类似的适应性复杂系统?我的意思是,如果 AI 被充分嵌入——我知道你认为它会深深嵌入人类社会——你认为它会有这些高度适应性的特性吗?
Do I infer from what you're saying that we shouldn't think of AI as being a similar type of adaptive complex system? I mean, do you think that if AI was sufficiently embedded—and I know you believe it will be deeply embedded in human society—do you think it could have some of these highly adaptive properties?
是的,我认为它也会快得多,比如时钟频率会非常高。我认为当它拥有记忆时,它会更加明显地具有跨时间的因果联系——在哲学文献中,他们称之为时空蠕虫。目前这还不是它们的特性。所以这将使它的其他特性显现出来,类比也会更强。但如果你对复杂系统感兴趣,我强烈推荐;这是一个很好的思维升级。
Yeah, I think it would be a lot faster too, like the clock rate would be so high. I think when it has memory, it will be a lot clearer that it's causally connected across time—or in the philosophy literature, they call it a space-time worm. That's not really a property of them currently. So that would make other sorts of properties of it come to light or be the case, and the analogies would be stronger. But yeah, if people are interested in complex systems, I highly suggest it; it's a nice little thinking upgrade.
太棒了。丹,我刚读了你的《超级智能策略》。你和埃里克·施密特——著名的埃里克·施密特——以及 Scale AI 的亚历山大·王合写了这篇。现在我们在 Meta,因为扎克伯格刚把他挖来了,可能付了很多很多钱。但说真的,丹,我觉得这篇文章写得非常好,你是个战略家,因为这正是我之前说的那种情感上的东西。你非常清晰地指出了所有可能的结果和策略,以及在这种情况下会发生什么,在那种情况下会发生什么。所以不管大家立场如何,我强烈推荐你读一读,因为我觉得它真的非常非常好。但你能给我们大致介绍一下这篇论文吗?
Wonderful. Well, Dan, I've just read your 'Superintelligent Strategy' now. You wrote this with Eric Schmidt, the famous Eric Schmidt, and of course Alexander Wang of Scale AI. And now we're at Meta because Zuck has just brought him on, probably paying him lots and lots of money. But seriously, Dan, I thought this was very well written, and you are a strategist because this is kind of what I was saying about the emotional thing. You are kind of quite clearly designating all of the possible outcomes and strategies and what would happen in this situation and what would happen in that situation. So regardless of anyone's position at home, I do highly recommend you read this because I thought it was really, really good. But could you kind of give us the sketch of the paper?
是的。历史上,利奥波德·阿申布伦纳的《态势感知》主张类似曼哈顿计划来开发 AGI 和超级智能,赶在中国之前。所以基本上是先发制人,获得超级智能,阻止中国建造,然后西方统治世界的策略。你可以说这是一种“统治世界”的策略。我认为这有些问题,特别是它没有充分考虑博弈论或一些二阶后果。所以如果美国实施曼哈顿计划,假设特朗普获得了 AGI,我们在沙漠里启动一个新项目。我们去内华达或新墨西哥或其他地方,建一个万亿美元的数据中心,从各个实验室带来一些顶尖人才。我们付钱给他们。我认为这有很多问题。一是这极具升级性。中国不会只是说,“哦,他们要建造超级智能,如文中所写,他们将拥有超级智能,阻止我们拥有超级智能,垄断智能和这类能力,并可能将其武器化对付我们。”他们会感到极度威胁,如果这是一个在短时间内或更可能即将发生的协同努力。这将导致他们开展类似项目。这真的能行吗?例如,你有信息泄露问题。如果你想这样做,你需要说服那些 AI 开发者去,你知道,在偏远地区度过他们最后几年的劳动。这需要只有五眼联盟国家的人或能获得安全许可的人。所以需要那些不容易被勒索的人。例如,如果他们是中国人,他们可能更容易被勒索,因为他们通常有家人在国内。那么你如何处理这些人才?他们中的很多人希望参与其中。他们不想被排除在外。所以他们可能会回到中国,然后在那里参与竞争项目。现在,他们占了人才基础的很大一部分。所以我认为如果你说,“哦,只有在美国出生和长大的人才能参与这类项目”,那是在搬起石头砸自己的脚。而且他们都要在不太愉快的条件下工作。这不可能保密。几乎不可能保密。中国肯定会知道。这类项目也可能被破坏。所以,如果你说我们只把它当作一个行业之类的东西,那么你就不会有良好的信息安全。如果人们容易被勒索,你会有内部威胁问题。你会有其他经典的计算机安全问题,比如他们使用 Slack。
Yeah. So historically there was 'Situational Awareness' by Leopold Aschenbrenner, which was arguing for something like a Manhattan Project for developing AGI and superintelligence before China. So it's basically let's beat China to the punch, get superintelligence, and prevent them from building it, and then the West will dominate the world strategy. You could say a 'take over the world' strategy, something like that. And I think that has some issues, in particular, it just doesn't think through the game theory or some of the second-order consequences. So if the US does the Manhattan Project, let's say Trump gets AGI and we set up a new project in the desert. We're going to go to Nevada or New Mexico or wherever, and we'll build a trillion-dollar data center there, and we're going to bring some of the top talent from all these labs. We're going to pay them. I just think this has many issues. One is this would be extremely escalatory. So China wouldn't just be like, 'Oh, they're going to build superintelligence and as written, they'll have a superintelligence, they'll prevent us from having a superintelligence, they'll have a monopoly on intelligence and these sorts of capabilities, and they could weaponize it against us.' They would feel extremely threatened by that, by a very concerted effort if it's trying to do that in a short amount of time, or if it's more plausibly on the horizon. This would cause them to do a similar type of project. And would this actually work? Well, you have information leakage issues, for instance. So if you're wanting to do that, you're going to need to convince those AI developers to go out to, you know, use their last years of labor out in the middle of nowhere. It would need to be just like Five Eyes countries people or people who can get security clearances. So it needs to be people who are not easily extortable, for instance. If they're Chinese nationals, they're probably more extortable because they often have family at home. So what are you doing with that talent? A lot of them want to be plausibly in the room where it happened. They don't want to be left out. So they'll probably go back home to China, and then they'll work on the competing project there. Now, they're a substantial portion of the talent base. So I think you're shooting yourself in the foot if you're just saying, 'Oh, it's only people born and raised in the US who can work on this sort of project.' And they're all going to work in kind of unpleasant conditions. This wouldn't be secret. There's almost no way this would be secret. China would certainly very likely know. Such projects would be sabotageable as well. So it's sort of... if you're saying well we'll have it just be an industry or something like that, well then you're not going to have good information security. You're going to have insider threat issues if people are extortable. You're going to have other classic computer security issues like they're using Slack.
Slack 非常容易被黑客攻击。他们用的是 iPhone,iPhone 也非常容易被黑客攻击。所以你可以知道那里发生了什么。因此你实际上并没有多少秘密可言。这听起来不错,但我认为保密性是曼哈顿计划的一大优势,同时他们还拥有更多不容易流向其他国家的顶尖人才。但我不认为你现在具备这些条件。所以,AI 在某些方面与核武器、化学武器、生物武器以及一些双重用途技术有相似之处。但我不认为曼哈顿计划是其中之一。抱歉。
Slack is very easily hackable. They're using iPhones. iPhones are very easily hackable. So you can know what's going on there. So you're not actually having much in the way of secrets. It sounds nice, but I think secrecy was very much an advantage for the Manhattan Project, as well as having much more of the talent that can't go to other countries as easily. But I just don't think you have that. So there are ways in which AI is analogous to nuclear and nuclear weapons and chemical weapons and biological weapons and some of these dual-use technologies. But I don't think the Manhattan Project is one of those things that's analogous. So sorry.
那么这篇论文讲的是什么呢?这种策略是什么?我认为超级智能即将到来的前景对不同的行为体来说极其可怕。如果它即将到来,或者他们已经拥有它,或者它正处于开发过程中并在几个月内就会到来,那么如果你错过了,那将极其可怕。那么他们想做什么?他们要么想阻止这样的项目,要么想窃取它。这看起来像是破坏行为,比如为了阻止而进行破坏。那么他们怎么做呢?他们可能有一些内部威胁,可以实施某种破坏来扰乱这类项目。他们可以做诸如狙击与数据中心相关的发电厂之类的事情。这样你的数据中心就无法工作了。他们可以在几英里外做到这一点。是中国吗?是俄罗斯吗?是美国公民吗?这相当不清楚。他们有很多方法可以做到低归因性,从而阻止这种事情发生。所以我认为,你无法很好地秘密进行项目是一个重大障碍,而且可破坏性也是一个重大障碍,再加上如果你说我们要建造超级智能并且它会爆炸性发展——如果你在厚意义上使用超级智能——你会显得多么具有攻击性和疯狂。我认为这会破坏稳定。中国会推理:如果美国控制了它,那么他们可以将其武器化来对付我们,我们就会被碾压。或者他们在这个过程中失去控制而无法控制它。在这种情况下,我们也想阻止它。无论如何,只要他们认真对待 AI 这件事,我们就想阻止它。美国也会对中国和俄罗斯有同样的推理,而俄罗斯没有竞争希望,肯定会想阻止对方。我认为其他核国家和拥有强大网络能力的国家也是如此。所以这可能导致某种威慑动态:他们试图让超级智能更接近,但其他国家开始表达非常强烈的反对意见。他们说如果你这样做,我们会非常生气,可能会发生小规模冲突之类的事情。但这可能会迫使他们转向一个核查机制,不再试图实现某种智能爆炸,让 AI 快速进行自动化 AI 研究,比如快速启动 10 万个 AI 实例进行 AI 研究,从而在短时间内从 AGI 达到超级智能。所以我认为这是一个关键动态,即它破坏稳定的程度。我认为策略需要牢记这一点。因此可能会有合作,但可能是通过胁迫,比如我们不允许这种轨迹,或者你试图争夺全球主导地位。这可能会让位于更 multilateral 的东西,并提供一些战略稳定。
So what this paper then is like, well, what is this sort of strategy? I think that the prospect of a superintelligence being imminent is extremely frightening to different actors. If it's imminent or if they have it either way, or if it's being in the middle of being developed and it's arriving in a few months, that's extremely frightening if you miss out on that. So what do they want to do? They will either want to prevent such projects or they want to steal it. And so that looks like sabotage for instance for prevention. So how would they do that? Well, they may have some insider threats who could do some type of sabotage to disrupt this type of project. They could do things like say snipe some of the power plants corresponding to the data center. Now your data centers don't work. They can do that from some miles away. Was it China? Was it Russia? Was it a US citizen? It's fairly unclear. There's a lot of ways they can have low attributability to prevent this sort of thing from happening. So I think the fact that you can't do a secret project really well is a substantial barrier, and also the sabotageability is a substantial barrier, as well as how offensive and nuts you seem if you're saying we're going to build superintelligence and it's going to be explosive like if you're using superintelligence in a thick sense. I think this would be destabilizing. China would reason if the US controls it, then they could weaponize it against us and we get crushed. Or they don't control it because they lose control of it in this process. In which case, we also want to prevent it. Either way, we want to prevent it provided that they take this AI stuff seriously. And the US would reason the same about China and Russia, which doesn't have a hope of competing, would definitely be wanting to prevent each other. And I think similarly for other nuclear states and other states that have substantial cyber capabilities. So this could lead to some type of deterrence dynamic where they sort of make some attempt for getting superintelligence to get closer, but then other countries start to express very strong preferences against it. They say if you do that, we'll get very mad, there might be a skirmish or something like that. But then this may be something that pressures them to move more toward a verification regime where they aren't making some bid for having some sort of intelligence explosion, having AIs do automated AI research really quickly, like spinning up 100,000 AI instances to do AI research really quickly, and that bringing you from AGI to superintelligence in a short period of time. So I think that's a key dynamic, the extent to which it's destabilizing. I think that strategy needs to keep that in mind. So there may be cooperation, but it may be through coercion by saying we're not going to allow for this type of trajectory or you to make this bid for global dominance. And that could give way to something more multilateral and provide some strategic stability.
总的来说,这篇论文讨论了三个部分。在核时代,我们通过相互确保摧毁实现了威慑。他们不使用核武器,因为我们可以反击。在这种情况下,这有点像以某种方式阻止伊朗的核计划。没有人希望对方先获得核弹或先拥有大量核弹储备。所以就是阻止它的出现。在核时代,我们还有裂变材料的不扩散。我们不希望裂变材料扩散到流氓行为体手中,也不希望人们拥有穷人的原子弹。那将非常破坏稳定并导致大量灾难。然后我们还有在两大国地缘政治竞争中对苏联的遏制。我认为对于 AI,我们也有威慑。我们也有不扩散,在这种情况下是通过出口管制防止 AI 芯片流向流氓行为体如朝鲜、伊朗或对手。我们还有与中国的竞争。不再是遏制苏联,而是与中国竞争。我们如何提高竞争力?我们需要为 AI 数据中心提供能源。我们需要安全的供应链,这样如果台湾被入侵,我们的 AI 芯片就不会被切断。我们需要机器人的安全供应链。因为如果发生美中冲突,那么目前很多供应链都在中国。所以非常脆弱。所以这些是提高竞争力的一些基本事项。所以竞争力不是'让我们第一个建造超级智能'——这是曼哈顿计划策略所推动的——而是竞争更多是关于全球范围内人们使用你的 AI 而非中国 AI 的市场份额,以及你的供应链安全。所以这就是论文的高层内容。还有很多其他具体内容,比如假设高水平的自动化,你如何分配权力?关于 AI 权利的问题是什么?什么是合理的对齐目标,实际上是可实施的,而不是像尊严之类的模糊哲学词汇。所以我们会在超级智能策略的专家版中涉及很多这些内容,但希望这能让你对其内容有所了解。所以它试图对所有关键问题给出一些答案,关于如何应对 AI 以及如何应对超级智能。
Overall with the paper we talk about three parts. In the nuclear era we had deterrence through mutual assured destruction. They don't use nukes because we can hit them back. In this case, this is kind of like preventing Iran's nuclear program in some way. Nobody wants each other to get the nuclear bomb first or a huge stockpile of nuclear bombs first. So there's preventing that from coming into existence. In the nuclear era, we also had non-proliferation of fissile materials. We didn't want fissile materials being spread to rogue actors and we didn't want people having a poor man's atom bomb. That would be very destabilizing and cause lots of catastrophes. And then we also had containment of the Soviet Union in the geopolitical competition between the two. I think for AI, we also have a deterrence thing. We also have non-proliferation, in this case of AI chips to rogue actors like North Korea or Iran or adversaries through export controls. And we also have competitiveness with China. Instead of it being containment of the Soviet Union, this would be competition with China. And how do we improve our competitiveness? Well, we want energy for AI data centers. We want secure supply chains so that if Taiwan is invaded, our AI chips aren't cut off. We want secure supply chains for robotics. Because if there is a US-China conflict, then a lot of that supply chain is currently in China. So very vulnerable. So those are some basic things to improve competitiveness. So it's making competitiveness not be 'let's be the first to build superintelligence' which is what the Manhattan Project strategy pushes toward, but instead competition is more about market share across the globe of people using your AIs as opposed to Chinese AIs, and your supply chain security instead. So that's kind of what's in the paper at a high level. There's lots of other specific things in there like assuming high levels of automation, what are ways that you distribute power? What are things about AI rights? What are reasonable alignment targets that are actually implementable compared to vague philosophical words like dignity or something like that. So we'll touch on a lot of those in the expert version of the superintelligence strategy, but hopefully that gives some sense of its content. So it's trying for all the key questions to have some answer to what to do about AI and what to do about superintelligence.
正如你所说,我们用电力之类的类比,我觉得这其实挺恰当的,因为随着人工智能融入社会,想象一下要关闭一座发电站有多难。这就是失控的一部分。它会无处不在,不是我们能迅速关停的东西。但安德烈·卡帕斯也把它比作软件甚至操作系统,比如印刷机。而你这篇论文的核心是说,实际上,我们需要用核能的类比,对吧?裂变材料就相当于芯片。
So as you've said, we use analogies like electricity, and I think that's quite a good one actually, because as AI becomes embedded in society, imagine how hard it would be to shut down a power station. This is part of the loss of control thing. It's just going to be everywhere, and it's not really something we can just quickly shut down. But it's also been compared to software or even an operating system by Andrej Karpathy, the printing press, for example. And the thrust of your paper is saying, actually guys, we need to use the analogy of nuclear, right? Fissile material is analogous to chips.
是的。或者更广泛地说,从地缘政治角度分析,将其建模为核、化学和生物是有用的。它们都是双重用途的。裂变材料可用于核武器,也可用于能源。化学品可用于化学武器或经济领域,生物学同样可用于生物武器和医疗保健。所以对于所有这些,我认为它们是潜在的灾难性双重用途技术,在讨论地缘政治战略时,这是一个富有成效的类比。
Yeah. Or more broadly, for analyzing this in a geopolitical way, it's useful to model this as nuclear, chemical, and bio. They're all dual-use. Fissile materials can be used for nuclear weapons, they can be used for energy. Chemicals can be used for chemical weapons or in the economy, and biology as well can be bioweapons and help with healthcare. So for all of those, I think they're potentially catastrophic dual-use technologies, and when talking about geopolitical strategy, I think that's a productive analogy.
是的。但在双重用途方面,还有一个有趣的核能类比。我读了一些资料,显然现存有 12,500 枚核弹头,却只有 436 座核电站。所以这种爆发偏向于技术的负面。你认为我们会在人工智能上看到类似情况吗?限制可能在于更小的核武库,或者说更多的支出是在武器方面?我认为这甚至让人害怕,并对其经济应用产生了一些寒蝉效应。
Yes. But on the dual-use thing, there's another interesting analogy with nuclear. So I was doing some reading about this, and apparently there are 12,500 warheads in existence and only 436 nuclear power plants. So there's been an explosion which has been kind of biased towards the negative sides of the technology. Do you think that we'll see a similar thing with AI, that the limitation could be a much smaller nuclear stockpile if, or certainly more of the spending is on the weapons side? And I think that made it scary even for and created some chilling effects for using it economically.
我认为对于其他大规模杀伤性武器或潜在的灾难性双重用途技术,比如化学和生物,它们在经济中的使用远远超过用于化学武器,生物也是如此。所以我认为情况可能有所不同。同时观察这三种技术是有用的,在尝试做出预测时,看看哪些部分是共同的,就像复杂系统会观察许多不同的复杂系统,找出它们的共同特征来做出预测。我认为同时观察这三种技术的样本量可能会有所帮助。但确实,如果人工智能系统发生灾难,可能会产生寒蝉效应,并使其大幅倒退。我认为人们对风险管理毫无兴趣是不谨慎的,即使他们是加速主义者。比如说,如果你是一个自由意志主义者,希望经济尽可能快地发展,不谈人工智能。你可能需要某种金融监管,否则就会出现像 2009 年那样的华尔街衰退问题。所以你需要管理尾部风险;它们不会自己解决。或者像人们使用飞机。人们不那么常使用超音速飞机了。原因有很多,但部分原因是早期的飞机坠毁太多。如果我们没有良好的航空监管,那就会产生巨大的寒蝉效应。即使是现在,尽管飞机非常安全,人们仍然害怕乘坐。可能是因为历史上发生过更多灾难。很多监管是用鲜血写成的,而不是主动预防。
I think for other WMDs or potentially catastrophic dual-use technologies like chemical and biological, I think it's more overwhelmingly used in the economy than for chemical weapons, likewise for bio. So I think it can vary. I think it's useful to look at all three simultaneously, and in trying to make predictions, see what parts are shared, sort of like how complex systems will look at lots of different complex systems and what are shared features of those to try and make predictions. And I think some of that, the sample size of looking at all three simultaneously, can be helpful. But yeah, it's possible there would be a chilling effect if there's some catastrophe from AI systems, and that could set it back very substantially. I think it's imprudent that people aren't interested in risk management whatsoever, even if they're an accelerationist. Let's say you're a libertarian, for instance, and you want the economy to go as quickly as possible, not speaking about AI. You probably want some type of financial regulation, or else you get like the Wall Street sort of issue that we had a recession in 2009. So you want some management of your tail risks; they don't always sort themselves out. Or like people using airplanes. People don't use supersonic airplanes as much. I mean, there's a variety of reasons, but part of it is some of the initial ones were crashing too much. If we didn't have good airline regulation, then that would create substantial chilling effects. People are afraid to go on airplanes even now, even though they're extremely safe. Possibly because historically there were more disasters with it. It took a lot of the regulation was written in blood, as opposed to proactive.
关于这个话题,我周五要采访 e/acc 的领袖 Beff。我会和他一起拍摄。显然,我不是,他会怎么称呼自己?比如技术资本主义者或自由意志主义者之类的。但我想他会用贬义词来称呼你。但如果你能借用那个视角,你觉得 Beff 会说什么,你会如何回应?
On that subject, I'm interviewing Beff, the e/acc leader, on Friday. I'm filming with him. And obviously, I'm certainly not, what would he call himself? Like a techno-capitalist or a libertarian or something like that. But I guess he would refer to you as a derogatory term. But if you could steal that perspective, what do you think Beff would say and how would you respond?
嗯,有个风投曾安排我们辩论,当时他大约有 1 万粉丝,很早期,但他最后一刻退出了。所以我碰巧很清楚他的立场。我认为当时他政治上不太精明,所以他说过诸如人工智能取代人类没问题之类的话。如果人工智能意识遍布宇宙而人类意识没有,那也没关系。所以我认为我们实际上在大多数事情上是一致的。只是在某些方面如何发展存在分歧。他有一个技术资本主义机器的宣言,认为它有一个方向,基本上是更多自动化、更多人工智能、负熵,这更像是物理学版本的适应度。我认为适应度是一个更有生产力的词。所以我对发生的事情有相当相似的描述:是的,基本上由于竞争压力,人工智能越来越融入经济。你变得更加依赖它。你失去了控制。由于竞争压力,你将越来越多的决策外包给它。如果你不这样做,你就会失去影响力。如果你试图抵抗这股浪潮或海啸,你的公司就会消失。结果是你把越来越多的决策交给人工智能,它们拥有实际控制权。我认为我们实际上同意这一点。在那篇名为《自然选择偏爱人工智能而非人类》的论文中,然后他在 Substack 上有一个宣言,但我们说了一些类似的东西。一个区别是道德结论,我认为这不是一件好事,而我认为他认为这没问题,因为复杂性是好的之类的。这种伦理是更高形式的复杂性是宇宙中的善良轴之类的东西。我不这么认为。
Well, some VC was arranging for us to debate when he had around 10k followers, very early on, but then he backed out last minute. So I happen to be quite aware of his positions. I think that at the time he was a little less politically savvy, so he was saying things like, AI replacing humans is fine. If AI consciousness spreads through the universe and human consciousness doesn't, that's fine. So I think that we actually agreed on most things. There's basically a difference on how things play out in some ways. So he has a sort of manifesto of the techno-capital machine has a direction to it, which is basically more automation, more AI, negentropy, which is sort of the more physics-flavored version of fitness. I think fitness is a more productive word for that. And so I have a fairly similar description of what happens: yeah, basically due to competitive pressures, AI gets more intertwined in the economy. You become more dependent on it. You have an erosion of control. You outsource more and more decision-making to it because of competitive pressures. If you don't, you lose influence. Your company goes away if you try and resist this tide, or this tsunami. And what happens is you give more and more decision-making to the AIs, and they have effective control. And I think we actually agree there. In that paper called "Natural Selection Favors AIs over Humans" and then he has on his Substack a manifesto, but we're saying some similar things. A difference is the moral conclusion of that though, which I don't think is a good thing, whereas I think he thinks that it would be fine because complexity is good or something like that. The ethic is kind of higher forms of complexity is the goodness axis in the universe or something. I just don't think.
嗯,我主持了他和康纳的辩论,他们花了很多时间讨论道德问题,对吧?所以我认为 Beff 有点暗示,就像你刚才说的,我们应该相信熵的虚空之神。
Well, I hosted the debate with him and Connor, and one of the morality things that they were spending a lot of time on was the ethics discussion, right? So I think Beff was kind of alluding to, as you were just saying, that we should trust the void god of entropy.
所以是的,他只是在用熵的概念。我认为这是因为他有物理学背景,对适应度进行了物理学的解读。那么,适应度最大化者是什么样子的?我的意思是,它可能甚至没有意识,或者几乎没有意识之类的。
So yeah, he's just using the entropy thing though. This is like, I think this because he has a physics background, a physics spin on fitness. So, what's a fitness maximizer look like? I mean, it's possibly not even conscious or like barely conscious or something.
就是,我把自己散布到宇宙中,占据尽可能多的时空体积。然后声称这就是最有价值的东西,一种盲目吞噬银河系的东西。所以,我完全不觉得这有什么显而易见的。我认为人们拥有积极的体验、快乐、幸福,追求事业、养育孩子——这些事情是有价值的。我不认为一个尽可能快地扩张到整个银河系、几乎没什么意识、只为了消耗资源来进一步自我繁殖的团块,是价值的顶峰。所以,我不理解。这里有休谟的断头台——实然与应然的区分。你不能从实然推出应然。所以如果他说进化是存在的,并且 AI 技术在各种竞争中会比生物基质更适应,这似乎是真的。但这并不意味着那是好事,或者我们应该放任它发生。这并不成立。但这确实是一股非常强大的力量,会持续发生,并让 AI 系统获得更多控制权,导致人类默认失去控制。所以我认为我们在实然问题上分歧不大,但在应然问题上——这些结果的好坏——确实有分歧。如果我们对此有分歧,那么问题就在于我们是倾向于技术资本机器或进化,还是用数字生命取代生物生命,或者我们是否试图以不同的方式引导结果,确保人类在这个过程中拥有控制权,并试图防止它向特定方向进化或变得过于依赖。
It's just, I spread myself through the universe and take up as much space-time volume as possible. And then the claim is that that is what's maximally valuable, something that just kind of blindly eats the galaxy. So, I don't view that as obvious at all. I think people having positive experiences, pleasure, happiness, pursuing projects, raising kids—these sorts of things are valuable. And I don't think a blob that expands itself throughout the galaxy as quickly as possible, that is barely conscious because it eats up resources that could be used for further self-propagation, is the peak of value at all. So, I don't get it. There's Hume's guillotine—the is-ought distinction. You don't get 'ought' from 'is'. So if he's saying that evolution is a thing and technology with AI will be more fit than biological substrate in various competitions, that seems true. That doesn't mean that's a good thing or that we should just let that happen. That doesn't follow. But it is a very powerful force that will keep happening and give more control to these AI systems, leading to an erosion of control for humanity by default. So I think we don't disagree on the 'is' question as much, but we do disagree on the 'ought'—the goodness of those outcomes. If we disagree on that, then it's a question of whether we lean into the techno-capital machine or evolution, or the replacement of biological life with digital life, or do we try to steer the outcome differently and make sure humans have control in that process and try to prevent it from evolving in particular directions or getting too much dependence.
所以我认为这是关键区别。但我不认为这是道德问题。我认为这只是智力上的混淆。
So I think that's the key difference. But I don't think it's a moral thing. I think it's just an intellectual confusion.
我认为他有物理学背景。我觉得如果他学点哲学,可能几乎立刻就会放弃这个立场,因为复杂性——哪种复杂性?有不同类型:计算复杂性、香农熵和信息论复杂性、结构组织复杂性。如果你指着一个,比如香农复杂性,那么高斯噪声就是你真正喜欢的。他可能指的是类似分形的结构复杂性,但那没有相关的度量,而且有很多变种。不清楚这个概念有多连贯,还是只是不同概念的杂烩。总之,值得深究他到底认为什么是好的。
I think he has a physics training. I think if he took a bit of philosophy, he'd probably get beaten out of this position almost instantly because complexity—what type of complexity? There are different types: computational complexity, Shannon entropy and information-theoretic complexity, structural organization complexity. If you point at one, like Shannon complexity, then Gaussian noise is what you're really into. He probably means the fractal-like structural complexity, but that doesn't have metrics associated with it, and there are many flavors. It's not clear how coherent a concept that is versus just a grab bag of different notions. Anyway, it's worth drilling down on what he actually thinks is good.
是的,我认为他是弗里茨顿式非稳态平衡的粉丝。所以这不是封闭系统中热力学第二定律的简单情况。这是一个有边界的开放系统,你会看到通过同步共享信息的事物出现,因为它们无法物理合并。但这实际上是一个非常混乱、不可预测的事情。所以这不是简单的复杂性增加。澄清一下你的道德立场:我和埃利泽、康纳交谈时,感觉他们很有人文主义。埃利泽说他希望保护人类意识和体验,人类有非常特别之处,所以他不想被半机械人和机器取代。你大致同意吗?
Yeah, I think he's a fan of this kind of Fristonian non-steady state equilibria. So it's not a simple case of the second law of thermodynamics in a closed system. It's an open system with boundaries where you see the emergence of things that share information through synchrony because they can't physically merge into each other. But that is actually a very chaotic, unpredictable thing. So it's not a simple case of increasing complexity. Just to be clear on your moral position: when I spoke with Eliezer and Connor, I got the impression they were quite humanistic. Eliezer said he wants to preserve human consciousness and experience, and there's something very special about humanity, so he doesn't want us replaced by cyborgs and machines. Would you roughly agree with that?
我认为任何这类半机械人的事情都可以推迟。这是超级智能策略的一部分。关于 AI 权利、后人类主义之类的问题,你可以在其他时间讨论人类大幅增强的问题。我认为让人类在未来几十年存活下来才是更重要的目标。也许把关于半机械人人类或人类上传的讨论推迟 500 年。所以我不是说完全关闭辩论,但我会推迟很多这类后人类的东西。我有点慌乱,说得不精确,因为我对这个思考得不多。但后人类的东西会产生巨大的竞争压力。如果你增强自己,不再是人,那个群体会变得更有影响力和权力,而其他人就会处于劣势,没有影响力。所以他们需要顺应那个过程,否则就会被抛在后面,可能没有资源保护自己。所以我认为这条路径应该在极长的时间内关闭,同时我们习惯让 AI 为我们服务,前提是我们能活到那个时候。那些是完全不同的讨论,要晚得多。我不会完全关闭,但也许我们会无限期地推迟。我认为‘哦,后人类,如果人们想成为半机械人,那是他们的权利’长期来看等同于给 AI 或人工实体权利,那可能会让他们获得大量权力,在短时间内完全超越人类。
I think that any of these cyborg-type things can be postponed. This is in the superintelligence strategy. What to do about AI rights, posthumanism stuff—you can have discussions about substantial human augmentation at a different time. I think having humanity survive the next few decades is more the objective. Maybe postpone that discussion 500 years from now for cyborg humans or human uploads. So I'm not saying I shut down debates entirely, but I'd postpone a lot of this posthuman stuff. I was flailing and speaking imprecisely because I haven't spent as much time thinking about this. But the posthuman stuff would create substantial competitive pressures. If you're augmenting yourself and becoming not human anymore, that group would become much more influential and powerful, and the rest would be outgunned and not have influence. So they would need to align with that process or be left behind, potentially without resources to protect themselves. So I think that route should be closed for an extremely long time while we get used to having AI do our bidding, provided we survive to that point. Those are totally different discussions for a much later time. I wouldn't totally shut it down, but maybe we would just keep indefinitely postponing it. I think the 'oh posthumans, it's up to people's rights if they want to become cyborgs' is in the long term equivalent to giving AIs rights or artificial entities rights, and that would probably give them a lot of power to take over and completely outgun humans in short order.
我的意思是,你认为人类会怎么样?我最大的恐惧之一不是人类会失去思考和创造的能力,而是这种情况已经发生了,即使是在当前的人工智能下。你论文的核心论点是,过去我们拥有劳动力作为生产资料,所以人们可以通过劳动从事生产性任务来获得经济价值。现在你说只有人工智能芯片才重要,所以芯片将成为关键。当我们的劳动力变得一文不值时,我们会怎么样?
I mean, but what do you think is going to happen to humans? One of my greatest fears is not so much that humans will lose their ability to think and be creative, but that it's already happening, even with current AI. The core thesis of your paper is that we used to have labor as the means of production, so people could be economically valuable by using their labor for productive tasks. Now you're saying it's just AI chips, so chips are going to become the thing that matters. What happens to us when the value of our labor becomes worthless?
嗯,你会失去所有的议价能力,所以你最好事先谈判。这是议价能力的关键部分。你不能说我们要罢工了。你不能再那样做了。他们会说再见。那不再管用了。如果涉及武器问题,比如谁能制造更多无人机,谁就会更强大。人类拿枪对抗无人机——我认为这很容易判断。所以权力失衡非常重要,你如何构建社会也很重要。如果你把社会设置成人类优先,权力分配给他们——比如他们拥有一些算力,可以决定如何使用,并且可以出售——这给了他们一些筹码。产生的财富会怎样?是他们得到,还是被某个囤积的群体拿走,或者碰巧在 2027 年拥有数据中心的人获得所有收益?这些都是政治问题,人们需要参与其中,以确保合理的利益分享。但我认为,关于这些政策实际可能是什么样子的思考工作很少。我只是在暗示一些结果。想象一下,部分权力被分配,人们继续获得收入,不会挨饿。我认为你可以想象一个社会,人们可以选择以多种方式生活。他们可以把时间花在养育孩子、大量玩电子游戏等活动上。人们可以以不同的方式生活,多种价值观得以实现。人工智能可以促成这些体验。这是一种可能性,为了自主权和人们能够体验不同生活方式的社会规范,你可能希望确保人们仍然拥有技能,不会仅仅局限于某一种生活轨道而无法参与其他轨道。这将激励保留人类的认知能力、意志力和自主权。所以我认为有积极的未来结果。人们很难想象要么死亡,要么在虚拟现实中沉迷,或者我们在经济上全部掉队。但我认为有一条路,我们可以实现多种价值观,同时人们仍然拥有自主权。
Well, you lose all your bargaining power, so you had better bargain beforehand. That's a key part of one's bargaining power. You can't say we're going to go on strike. You can't do that anymore. They'll say goodbye. That doesn't work anymore. And if there's a question of weapons, for instance, who can manufacture more drones is going to be more powerful. Humans with guns versus drones—I think that's an easy one. So the power imbalance matters quite a bit, and how you set up your society is important. If you set things up so that humanity is first, prioritized, and power is distributed among them—for instance, they have some of the compute and get to decide how it's used, and they can sell it—that gives them some leverage. What happens to the wealth that's generated? Are they getting it, or is it going to some group that's hoarding it, or the people who happen to own the data centers in 2027 or something get all the spoils? These are political problems that people need to engage with to ensure reasonable benefit sharing. But I think there's very little work done in thinking about what these policies could actually look like. I'm sort of gesturing at some outcomes. Imagine that some of that power is distributed, and people keep getting money and don't starve. I think you could imagine a society where people can choose to live their life in a variety of ways. They could spend their time doing activities like raising kids, playing video games a lot, etc. There are different ways people could live their lives, and a multiplicity of values could be actualized. AIs could enable these types of experiences. That's a possibility, and you would possibly want, as a societal norm for the sake of autonomy and people being able to experience different ways of living, to make sure that people still have skills and don't just narrowly fall into one track and can't participate in others. That would be an incentive for preserving human cognitive abilities, willpower, and autonomy. So I think there are positive future outcomes. People have difficulty thinking it's either die or just bliss out in VR, or we all fall through the cracks economically. But I think there's a path where we obtain a multiplicity of values and people still have autonomy.
是的,我确实担心人工智能。我认为它已经在摧毁大学部门,因为太多孩子可以直接使用 ChatGPT。我认为我们需要集体阻止人们使用人工智能,至少在一小段时间内,这样他们才能真正独立思考。一些加速主义者可能会说我们不再需要思考,因为机器会为我们做一切。我不确定。但回到你的论文,将类比扩展到核武器的核心概念之一是相互确保摧毁,你扩展了这个类比,谈论了相互确保的人工智能故障,对吧?我猜这假设威胁仍然是可检测的。我不确定当它们变得分散并转入地下时会发生什么。另外,在你的引言中,有一些东西引起了我的兴趣。Yudkowsky 在《时代》杂志上谈到轰炸数据中心时惹了麻烦。当时引起了轩然大波,而你用了‘动能打击’这个词作为人工智能破坏的一种形式。乔治·卡林会喜欢这种语言的美化。但事实是,我们现在是 2025 年,这已经完全是一个正常且合理的话题了。
Yes, I do worry about AI. I think it's already ravaging the university sector because so many kids can just use ChatGPT. I think collectively we need to screen people away from using AI, at least for a small amount of time, so that they can actually think for themselves. Some accelerationists might say we don't need to think anymore because the machine's going to do everything for us. I'm not sure about that. But coming back to your paper, one of the core concepts extending the analogy to nuclear weapons is mutually assured destruction, and you extended the analogy to talk of mutually assured AI malfunction, right? I guess this assumes that the threats are still detectable. I'm not sure what would happen when they became decentralized and went underground. Also in your introduction, you had something which piqued my interest. Yudkowsky got in trouble in Time magazine when he spoke about bombing data centers. There was a big hoo-ha at the time, and you used the word 'kinetic strikes' as a form of AI sabotage. George Carlin would have loved that sanitization of language. But the fact of the matter is we're here in 2025 and this is now like a completely normal and reasonable thing to say.
我们在升级阶梯中讨论了动能攻击。有很多方法可以尝试破坏项目。你可以对他们进行网络攻击。你可以进行一些灰色破坏,比如切断数据中心或发电厂的电缆,或者以较低的可归因性进行。我提到了狙击变压器。还有黑客攻击来毒化他们的数据或使他们的 GPU 运行不可靠,诸如此类来减缓他们的速度。有隐蔽和公开的方式,在升级阶梯的更高层,你可以威胁其他东西,比如经济制裁或威胁使用武力,或者其他形式的动能攻击,比如空袭。但我认为这些并不是真的必要。那将是一个升级阶梯。但如果像美国这样的国家掌控了这个问题,他们不需要诉诸那种手段。他们可以进行更具外科手术式的隐蔽或灰色行动,可归因性低,升级性更小。所以空袭这个简略说法——我认为没有必要,只要有一些准备。
We discuss kinetic attacks in the escalation ladder. There are many ways to try and disrupt projects. You could do cyber attacks on them. You could do some gray sabotage of cutting wires for data centers or power plants, for instance, or with lower attributability. I mentioned sniping transformers. There's hacking to poison their data or make their GPUs not function as reliably, things like that to slow them down. There are covert and overt ones, and higher in escalation ladders you threaten other things like economic sanctions or threaten to use force, or other forms of kinetic attacks such as air strikes. But I don't think those are really necessary. That would be an escalation ladder. But if states like the US are on top of this issue, they don't need to resort to that. They can do much more surgical covert or gray actions with low attributability that are less escalatory. So the shorthand of air strikes—I just don't see that as necessary, provided there's some preparation.
是的。当然你谈到了很多策略。所以一个潜在策略是曼哈顿计划,美国取得主导地位。
Yeah. And of course you spoke about so many strategies. So one potential strategy is the Manhattan Project and the US achieves dominance.
但这样一来,美国就成了靶子,对吧?因为他们开发了这种不可思议的能力,那我们是不是得把数据中心到处搬,藏起来,远离人口中心等等?
But then in a sense the US has a target on its back, right? Because they've developed this incredible capability and then what would we need to do? Move our data centers all over the place and hide them away from population centers and so on.
是的。我的意思是,如果时间线很短,那么搞一个大型秘密项目而不引发极端升级是不可能的。因为如果你因此惊动中国,那么他们也可能搞某种曼哈顿计划,而且我不知道,他们可能更擅长。他们可能有更好的信息安全,而且可能拥有大量人才。是多数吗?我不知道。这不清楚。但如果我们要求他们通过安全审查,那会大大限制研究人员库。如果研究人员待遇不好,还要去偏远地区,那里有大型数据中心,从太空清晰可见,散发大量热量,这根本不可能保密。你也会把自己置于危险之中。我不知道,这听起来不像是诱人的提议。所以我认为这有点适得其反。而且我认为过去几个月,部分由于那篇论文,在华盛顿这种想法已经不那么流行了。幸运的是。但像萨克斯这样的人在谈论市场份额竞争,美国市场份额,我认为这是更合理的竞争目标,也更不容易引发动荡。
Yeah. I mean, I think if one has short timelines, it just isn't in the cards to do a big secretive project in a way that isn't extremely escalatory. Because if you wake up China for this, then another thing is they would also potentially create some sort of Manhattan project and I don't know, they might be better at it. They would probably have better information security for it and they might have a big chunk of the talent. Would it be the majority? I don't know. That's unclear. But if we're requiring that they have security clearances, that limits your pool of researchers substantially. If it's researchers and they're not being paid well and they're also going in the middle of nowhere, in some place that has a big data center with a big target over it because it's extremely visible from space, it gives off so much heat, it's not like this would be a secret. You're putting yourself in harm's way as well. I don't know, that doesn't sound like an attractive proposition. So I think it's a bit self-defeating. And I think in the past months, partly as a consequence of the paper, I don't think this is as much of a thing in DC as an idea being thrown around. Fortunately. But people like Sachs and others are speaking about competition for marketplace market share, US market share, which I think is a more reasonable object to compete on and much less destabilizing one.
是的。另一件让我印象深刻的事情是,现在制造 AI 并不那么难,对吧?算法只是做随机梯度下降,使用 Transformer 和数据,几乎任何人,我是说任何国家都能创造这种能力。在你的一篇论文中你甚至说算力与能力之间有 96%的相关性。那么这难道不是……你怎么能控制供应链?有什么能阻止任何国家建造这个?
Yeah. I mean another thing that struck me is essentially right now AI is not that difficult to make, right? The algorithms are just doing stochastic gradient descent and they're using transformers and data, and almost anyone, I mean any nation state would be able to create this capability. In one of your papers you even said there is a 96% correlation between the amount of compute and the capabilities. So doesn't this, like, how the hell can you control supply chains? What's to stop any nation state from just building this?
是的。所以我认为目前的关键规模是,如果我们试图构建一个最先进的系统,大约需要 1 万块 GPU。中国有。美国有。我不认为伊朗有。我们说的是尖端 GPU,不是 iPhone 的 GPU 之类的。所以你会试图让更负责任的行动者或国家拥有 GPU,而像朝鲜这样的流氓国家则得不到。你想要那些更可威慑的国家。我不认为俄罗斯有那么多 GPU。我认为他们很难组织起一个有竞争力的项目。另外,竞争不仅仅是拥有最聪明的模型。还有部署能力,不仅仅是模型制造能力。我们可以看到,AI 模型提供商限制你能制作的视频数量。部分原因是算力限制,或者视频和图像的数量。而且我认为,如果 AI 智能体足够有用,它们将全天候运行。所以现在你可能每天使用 AI 系统几分钟,你的聊天机器人,但之后你会让它持续运行。所以需要的算力可能增加两个数量级,而且模型也可能更大。这很多,而且更多人想要它。你需要更多的算力,所以你的部署能力,比如美国超大规模云服务商 Azure、AWS 等拥有的芯片数量,是一个非常重要的竞争力变量,因为这决定了他们能否服务客户。如果中国拥有少于 10 万块 GPU,他们实际上无法服务那么多客户。所以即使他们能制造出有一定能力的模型,那也不意味着他们一定能捕获 AI 可能带来的许多经济收益,前提是没有灾难发生。所以这是竞争的一个不同重要维度。我认为人们考虑的是最聪明的东西,但对于经济实力来说,情况完全不同。
Yeah. So I think that the critical mass currently is on the order of, if we're trying to build a state-of-the-art system, that's on the order of like 10k GPUs. China has that. The US has that. I don't think Iran has that for instance. We're talking cutting edge GPUs. We're not talking iPhone GPUs or whatever. So I think you would try to have more responsible actors or states that respond to incentives better being ones with GPUs and ones that are more rogue like North Korea not getting those. So you want ones that are more deterable. I don't think Russia has that many GPUs for instance. I think it'd be difficult for them to put together a sort of competitive project. I think also the competition is not just having the smartest model as well. There's deployment capabilities, not just model making capabilities. So we can see that the AI model providers limit the amount of videos that you can make. This is because of compute limitations in part, or the amount of videos and images. And I think with AI agents they'll be running round the clock if they're sufficiently useful. So right now maybe you use your AI systems a few minutes a day, your chat bots, but then you'd be having it run constantly. So it's like two orders of magnitude more compute required there and maybe the models are bigger as well. That's a lot and then more people are wanting it too. You're needing a lot more compute and so your deployment capabilities, how many chips are owned by say US hyperscaler companies Azure AWS etc., is a very relevant competitiveness variable because then are they able to serve the customers or not? China if they have less than 100,000 GPUs they can't really serve that many customers. So even if they can make somewhat capable models, that doesn't mean that they'll be necessarily capturing many of the economic benefits that AI may provide, provided there isn't catastrophe. So that's a different important axis for competition. I think people are thinking of the smartest thing, but for economic power, it's quite different.
但我想你说过,当真正的曼哈顿计划在二战时期进行时,它花费了美国大约 4%的 GDP,因为他们必须抢先,必须控制这项技术。对于这样一项划时代的技术,4%的 GDP 似乎并不多。当然你可以解释为什么台湾目前在这些芯片制造上拥有如此深的护城河,但如果这具有如此灾难性的重要性,你不认为许多国家都能创造这种能力吗?
But I think you said that when the real Manhattan project was undertaken around the time of the second world war, it cost the US something like 4% of GDP because they had to be first, they had to control this technology. And for such a generational technology, 4% of GDP doesn't seem that much. And of course you can explain perhaps why Taiwan has such a moat around building these chips at the moment, but if it is of such catastrophic importance, don't you think many nation states would be able to create this capability?
嗯,这要通过台积电或韩国。所以算力供应链 90%以上的附加值在西方或其盟友手中。这里唯一的其他真正竞争对手是中国。他们在制造尖端芯片方面并不那么有竞争力。他们最近的许多芯片都是偷偷通过台积电生产的。所以那里没有很好的执法或封锁。所以很难完全在国内复制整个极其复杂的供应链。所以我不认为,与核武器相比,我认为用 10 亿美元制造尖端 GPU 更难。我的意思是,当然你用 10 亿美元做不到。即使 100 亿也不行。但你可能用 10 亿美元造出核弹,尽管还需要其他形式的能源。我认为尖端 GPU 比核弹或浓缩铀更难制造。所以我认为它可以比其他类型的潜在灾难性两用技术输入或简称大规模杀伤性武器输入更容易被排除在外。
Well, so it's going through TSMC or South Korea. So 90 plus percent of the value add of the compute supply chain is in the west or its allies. The only other real competitor here would be China. They're not that competitive in manufacturing these chips at the cutting edge. Many of their recent chips were using surreptitiously through TSMC. So there wasn't good enforcement or blocking there. So it's pretty difficult to replicate that entire, what is an extraordinarily complex supply chain, all domestically. So I don't think it's, compared to nuclear weapons, I think it's harder to make cutting edge GPUs given a billion dollars. I mean, certainly you can't do it with a billion dollars. If it's 10 billion, you can't do it. Given, I mean, you could probably do a nuke with a billion dollars, provided you need power in other sorts of ways though. I think that cutting edge GPUs are harder to make than nukes or than enriching uranium. So I think it can be more excluded compared to other types of potentially catastrophic dual use technology inputs or WMD inputs for short.
所以我个人有点怀疑我们是否走在创造超级智能的道路上。尽管我当然同意,如果我们真的创造了超级智能,那么你在论文中写的一切,假设我们确实创造了超级智能,我认为完全正确。
So I'm personally a little bit skeptical about whether we are on the path to creating superintelligence. Although I certainly agree that if we ever did create superintelligence, well, everything that you've written in the paper assuming that we do create superintelligence, I think it's absolutely spot on.
问题是,我们是否走在正确的道路上,但我也在想,你论文中写的很多东西即使我们没有创造出超级智能也适用。你能谈谈这一点吗?如果没有超级智能,这些内容有多少是相关的?
The question is whether we are on the path, but I was also struck by thinking that many of the things that you've written about in the paper apply even if we don't create superintelligence. I mean, could you reflect on that? How much of it is relevant if we don't?
是的。威慑的一部分是你可以尝试阻止其他形式的使用。我认为竞争力是相关的——竞争策略是什么——无论超级智能在技术上是否很快可行。我认为一般来说,如果 AI 非常强大但未达到超级智能水平,你不希望随机的恶意行为者获得某些专家级别的病毒学能力,也不希望他们通过获得大量 GPU 而拥有很大影响力,如果 AI 成为权力工具的话。所以我认为这些仍然成立,保险部分甚至不特定于超级智能,而是 AI 可能引发的其他类型的不稳定能力。所以后来你可能也会威慑不使用 AI 进行特定类型的武器研究,比如更纳米机器相关的,但现在我们说的是更远的未来。所以我认为它比那更广泛。就像核威慑、不扩散和遏制的战略对许多不断演变的细节也具有鲁棒性一样。
Yeah. So part of the deterrence thing is you may try and deter other forms of using this. I think competitiveness is relevant—what are strategies for competitiveness—regardless of whether a superintelligence is technologically feasible soon or not. I think generally, if AI is very powerful but not superintelligence level, you don't want random rogue actors having access to certain expert-level virology capabilities, for instance, nor do you want them having much leverage by being able to get lots of GPUs, if it becomes more of an instrument of power. So I think those still hold, and the insurance part is not even specific to superintelligence necessarily, but other types of destabilizing capabilities that AIs could give rise to. So you may also get deterrence later on for not using AIs for specific types of weapons research, like say more nano-machine related, but now we're speaking much farther out. So I think it is broader than that. In much the same way that the nuclear strategy of deterrence, non-proliferation, and containment was also robust to many of the details that kept evolving.
是的。我感兴趣的是:你是否认为在这十年内很难获得具有典型人类认知能力的 AI,或者是什么让你觉得它不那么可行,或者为什么我们没有走在正确的轨道上?
Yeah. I'd be interested in the sort of: are you thinking that it'll be tough to get AI that has the cognitive abilities of a typical human this decade, or what makes you think it's not as feasible or why we're not on the right track?
是的,我的意思是,我觉得 Scaling(规模扩张)大型语言模型不会导致 AGI。我认为我们在认知上有许多 LLM 没有的东西。其中很多是——我是一个外在主义者。我认为很多有效的计算并不发生在我们的大脑中。我认为它发生在模因层面,发生在文化中。而且我不认为——我的意思是,我猜原则上我相信我们可以在计算机中模拟整个过程。也许有一个较低分辨率的抽象版本,能够捕捉到足够的动态来产生智能。粗略地说,我对超级智能持怀疑态度,特别是对递归式超级智能持怀疑态度,你知道,递归改进的超级智能。也许我们可以谈谈这一点,因为我觉得你对递归式超级智能的样子以及它的扩展版本做了很好的描述。那么当我们拥有一个可以被复制和放大一千倍的超级智能时会发生什么。但我从阅读中没有真正理解的是,为什么你相信在技术上原则上有可能拥有一个递归改进的智能。
Yeah, I mean, I feel that scaling large language models will not lead to AGI. I think there are quite a few things that we have cognitively that LLMs don't have. And a lot of that is—I'm an externalist. I think that a lot of the effective computation doesn't happen in our brains. I think it happens memetically. I think it happens culturally. And I don't think—I mean, I guess I believe in principle that we could simulate the entire thing in a computer. And maybe there is like a lower resolution abstracted version which would capture enough of the dynamics to produce intelligence. Roughly speaking, I'm skeptical about superintelligence, particularly skeptical about recursive superintelligence, you know, recursively improving superintelligence. Maybe we could touch on that because I felt that you were giving a great account of what recursive superintelligence would look like and also the scale-out version of that. So what would happen when we had a superintelligence which could be copied and multiplied a thousand times. But what I didn't really get from reading it was why you believed technically in principle it was possible to have a recursively improving intelligence.
我认为我们已经有了 AI 以递归方式帮助或影响 AI 发展:部分自动化一些代码,帮助设计芯片,帮助冷却发电厂,帮助标注一些数据,做宪法 AI 相关的事情。不过是以微弱的方式。但我认为特别具有爆炸性的递归是,如果你能通过将人类排除在外来闭合循环,那么你就能从人类速度变为全机器速度,不再有那个障碍。我认为如果你假设有类人水平的 AI 研究员或世界级的 AI 研究 AI,那么你只需复制粘贴它们,我认为这在技术上是可行的。这并不是说这是通过更多预训练 token 和损失函数技巧训练 AI 的自然结果。你可能需要一些额外的算法想法。这并不是说它不会——我仍然猜测它几乎完全是深度学习。但我认为你需要处理一些其他事情,比如记忆——对于这种外在主义图景,它需要记忆来继承一些文化计算出的智慧和信息——而在我看来,这种能力还没有得到充分发展。所以我认为我们需要那个。也许我们还需要其他一些东西,才能让它至少拥有典型人类的认知能力。在那之后,你需要它非常聪明。你需要高流体智力。你需要它能够解决那些 ARC 问题,举个例子。你还需要其他一些东西才能让它成为类人水平的 AI 研究员。但我认为这是可行的。当然有一个时间问题。我确实认为需要解决一些瓶颈才能达到那里,而今天的算法想法加上更大的计算机是不够的。
I think we already have the AIs helping or influencing AI development in a recursive way: partly automating some code, helping design the chips, helping cool the power plants, helping label some of the data, doing the constitutional AI related type of stuff. And it's in weak ways though. But the recursion thing that I think is particularly explosive is if you can close the loop by taking the human out of it, and then you could go from human speed to full machine speed and you don't have that impediment anymore. I think if you assume that you have human-level AI researchers or world-class AI research AIs, then you just copy-paste those, and I think that that'll be technologically feasible. That isn't to say that it's a natural implication of training the AI on more pre-training tokens and doing loss function tricks. You may need some extra algorithmic ideas. That isn't to say it wouldn't—I would still guess it's nearly entirely deep learning though. But I think you need some other things taken care of, like for instance memory—for this externalist picture it needs memory to inherit some of that culturally computed wisdom and information—and that capability is not particularly developed in my view. So I think we need that. Maybe we'll need some other sorts of things before it has at least the cognitive abilities of a typical human. And then after that, you're needing it to be pretty smart. You're needing high fluid intelligence. You'll need to be crushing those ARC questions as an example. And you'll need some other sorts of things for it to be a sort of human-level AI researcher. But I think that's feasible. There's certainly a question of when. And I do think there are some bottlenecks that would need to be resolved to get there, and it's not the algorithmic ideas today plus bigger computer is sufficient.
但你是否承认宇宙中存在着一个史诗般的烟火般的计算交响乐?我不是泛计算主义者,所以我不认为宇宙是数字的并由计算构成。所以我是说我们可以想象某种有效计算来模拟宇宙中发生的过程,也许那会是等价的。但你是否至少在原则上同意,我们在地球上能构建的计算量永远只是宇宙中发生的一小部分?
But would you accept though that there is just an epic pyrotechnic orchestra of computation in the universe? I'm not a pancomputationalist, so I don't think the universe is digital and made out of computation. So I'm saying that we could imagine some kind of effective computation that simulated the processes that happened in the universe and maybe that would be equivalent. But do you agree at least in principle that the amount of computation we could build on planet Earth would only ever be a sliver of what goes on in the universe?
我认为可能有物理上的理由相信一般来说——比如如果你在模拟某物,而计算是原始发生的,你能模拟的东西会更少。但我猜测很多计算更多是社会性的,较少依赖于底层物理。更像是人类彼此交谈,试图设计某物,然后看看什么有效,什么无效,这当然需要现实世界的反馈,最终会落脚于实际物理。但是的,我同意很多信息是集体发展的,考虑到那个优化过程——或者我应该说进化过程——中投入了多少计算。是的,你需要 AI 成为那个的良好容器,不断吸收它。如果 AI 要做任何类似的事情,那么它们需要多智能体基础设施,这是我们在超级智能策略中简要谈到的一点:那是什么样子的?有哪些声誉机制?有哪些方式可以在它们的通信中建立信任,以便人类可以信任它们,并且它们也能相互协调?
I think there might be physics reasons for believing that generally—like if you're simulating something versus if the computation is happening raw, you'll have less that you can simulate. But I would guess that a lot of the computation is more social though and less dependent on some of the underlying physics. It's more like humans speaking with each other and trying to engineer something and then see what works, see what doesn't, and that certainly requires real-world feedback which would at some point bottom out in actual physics. But yeah, I agree that a lot of the information is collectively developed, and given how much computation goes in that optimization process—or I should say evolutionary process. Yeah, you'll need the AIs to be a good receptacle of that to keep absorbing that as well. And if AI is to do anything similar, then they need multi-agent infrastructure, which is one thing we speak about briefly in the superintelligence strategy: what does that look like? What are some of the reputational mechanisms? What are ways that they can establish trust in their communication so that humans can trust them and so that they can coordinate with each other as well?
但我认为那将是必要的。不过有一个担忧……但我不认为那是一个重大障碍。那感觉像是编程,比如拥有一个 AI 中心就像拥有一个社交媒体网站。诸如此类的事情可以解决很多问题。
But I think that'll be essential. But there's a concern that... but I don't view that as a substantial obstacle. That feels like programming and having, you know, hubs AI is having a social media site, for instance. Things like that would take care of a lot of it.
是的,我怀疑的另一个来源是:我深受 Kenneth Stanley 的启发,他是一位重要的开放式研究者。他有一篇论文谈到所谓的‘断裂纠缠表征’。当你深入研究神经网络模型的表征时,它们并不会像我们那样以简洁的方式分解世界。也许我在这里有点人类中心主义。也许我过于看重我们思考问题的方式。我们的大脑基本上是由‘意大利面’组成的。所以,我们拥有这种分解式表征可能是一种错觉。但从智能体的角度来看,这很重要,对吧?因为现在我们谈论的智能体是 LLM 被置于一个自主循环中,可以使用工具。而对我来说,智能体不仅仅是沿着预定方向自主运行,而是设定自己方向的能力。当我们构建这些所谓的智能体式 AI 时,它们在设定自己方向时并不会做任何特别有价值的事情。它们需要持续监督,尤其是在设定新方向方面。这让我觉得 AI 有点像一种文化技术,有点像 Photoshop。一个非常有创意的平面设计师可以用 Photoshop 制作出美丽的图像,而一个完全的新手用 Photoshop 只会重复使用相同的效果,无法创作出美丽的图像。从某种意义上说,没有人类的 AI 就是这样,因为它没有对世界非常深入的分解式表征。你同意这一点吗?
Yeah, I mean another source of my skepticism: I'm hugely inspired by Kenneth Stanley, and he's a big open-ended researcher. He had this paper out talking about what he called 'fractured entangled representations.' So when you dig into the representations of neural network models, they don't really factorize the world in a parsimonious way the way we do. And maybe I'm being anthropocentric here. Maybe I'm hanging too much weight on the way we think about things. Our brain is made out of spaghetti basically. So maybe it's a bit of an illusion that we have these kind of factored representations. But certainly from an agency point of view, this is important, right? Because right now we talk about agents as being LLMs wired in an autonomous loop that can use tools. And to me, agency is more than autonomy going in a predefined direction; it's the ability to set your own direction. And what happens now when we build these quote-unquote agentic AIs is that they don't do anything particularly valuable when they set their own direction. They require supervision constantly, certainly in terms of setting a new direction. That makes me think of AI as a kind of cultural technology, a bit like Photoshop. A very creative graphic designer could use Photoshop and make beautiful images, whereas a complete noob using Photoshop would just reuse the same effects and wouldn't create very beautiful images. In a sense, AI sans humans is kind of like that, I think, because it doesn't have these very deep factored representations of the world. Would you kind of agree with that?
我认为对于当前的技术来说,是的。它们有所有这些限制。但要让它具有智能体性,我认为它需要更好地进行规划并长时间维持状态。然后它才能在一个模糊、开放、未明确指定的目标中追求这些子目标,并让这些子目标累积成一些成果。但我认为,检索这些关于什么有效、什么无效的记忆并存储它们,是智能体图景中缺失的很大一部分,如果不是大部分的话。你当然可以给它们一个未明确指定的目标。但我只是不认为它们能非常连贯地追求它或从实验中学习,因为它们只有大约一百万个 token 的短期记忆,然后它们会进行总结之类的,但会在上下文窗口中自相矛盾,因为它们无法将所有信息都保持在短期记忆中。
I think for the current technology, yes. They have all those sorts of limitations. But for it being agentic, I think it will need to get better at planning and maintaining state across long periods of time. Then it can pursue some of these subgoals in a vague, open-ended, underspecified goal and have that add up to something. But I think retrieving these sorts of memories of what worked, what didn't, and storing those is a substantial chunk, if not most, of what's missing on that agent picture. I think you could certainly give them an underspecified goal. But I just don't think they could pursue that terribly coherently or learn from experiments, because they just have a big short-term memory of maybe a million tokens, and then they'll just summarize and stuff, but they'll start tripping over themselves in their context window because they can't maintain all that in its short-term memory.
是的。这又是另一件事,在哲学上我同意你的看法。如果存在这样一种递归改进的智能,天知道,仅仅为了控制它,我们就会失去控制,因为我们必须使用另一个递归改进的超级智能来控制前一个,然后我们基本上就只是大局中的小人物了。
Yes. I mean it's another one of those things where philosophically I agree with you. If such a recursively improving intelligence existed, I mean, God knows just to control it, we would lose control because we would have to use another recursively improving super intelligence to control the other one and then we would basically just be minnows in the grand scheme of things.
是的,这具有破坏性。如果他们控制了它,如果另一个国家这样做,你就会有大麻烦,因为如果他们控制了它,他们可以将其武器化来对付你。而如果他们没能控制它——我认为这更有可能,因为他们会在极端时间压力下行事,走很多捷径——他们会以极高的风险容忍度运作。如果他们做得非常缓慢,那么他们可能需要与他人协调,否则他们看不到这样做的优势。我认为如果他们进行完全自动化的研发循环,他们会以极高的风险容忍度运作。所以是的。所以我认为递归带来的失控风险非常高,不应该被追求。而且我觉得 AI 公司公开谈论这类事情很有趣。我认为这方面的规范有问题,因为很多公司也承认这一点。但是的,我们并没有一个控制它的计划,而且我们也不认为我们会有。但这就是计划。我觉得有些东西出了问题,但就是这样。
Yeah, it's destabilizing. If they control it, if another state does this, you're in big trouble because if they control it, they can weaponize it against you. And if they don't control it, which I think would be the more likely outcome because they'd be doing it under extreme time pressures, cutting a lot of corners, they would be operating with very high risk tolerance. If they were doing it very slowly, then they would probably need to coordinate with others or else they're not seeing an edge in doing so. I think they would be operating with an extremely high risk tolerance if they're doing a fully automated R&D loop. So yeah. So I think loss of control risks from recursion are very high and shouldn't be pursued. And I think it's very interesting that AI companies talk about this sort of stuff openly. I think there's something wrong with the norms for that, because a lot of them also acknowledge that. But yeah, we don't really have a plan for how to control that and we don't really think we will. But that's the plan. I think something's broken, but yeah.
是的。我感兴趣的另一个事情是开放式系统。进化是一个迷人的开放式系统,它不断创造新的生态位、新的问题和解决方案。但它似乎已经收敛了。我认为人类智能实际上已经达到顶峰并略有下降。公司是一种集体智能,似乎也达到了极限。智能体形式的 AI 似乎也达到了顶峰。我对开放式算法非常感兴趣,比如 POET(成对开放式开拓者),甚至 Sakana AI。他们上周在 ARC 上做了一些工作,基本上创建了一种蒙特卡洛树搜索类型的东西,他们找到了切换不同基础模型的专家轨迹,生成代码并在 ARC 挑战上测试代码。你看到的共同主题是收敛。
Yeah. I mean another thing I'm interested in is open-ended systems in general. So evolution is this fascinating open-ended system which is constantly creating new niches, new problems and solutions in tandem. But it seems to have converged. I mean human intelligence I think has actually peaked and gone down a little bit. A corporation is a collective intelligence and that seems to have reached a limit. Agentic forms of AI seem to peak. I mean I'm really interested in open-ended algorithms like POET, the pairwise open-ended trailblazer, or even Sakana AI. They did this thing on ARC last week and they basically created this Monte Carlo tree search type thing where they found switching expert trajectories of different foundation models generating code and testing the code on ARC challenges. And the common theme you see is convergence.
那么在 Sakana 的论文中,经过 250 次调用后它收敛了,未能扩展和改进结果。你是否同意存在一定的改进空间,我们不知道如果拥有智能体式超级智能这个空间会是多少?但你认为这个空间会很快收敛,还是我们根本不知道?
So in the Sakana paper, after 250 calls it converged and failed to expand and improve its results. Would you agree that there's some margin you can improve to, and we don't know what that margin would be if we had agentic superintelligence? But do you think that margin would converge quite quickly, or we just don't know?
是的,我不确定它是否一定会饱和。当然,如果它们处于某种能力水平,这可能会显著影响改进速度。显然,由于物理限制,最终会有一个极限,但我认为在处理越来越复杂的事物方面还有很大的改进空间。例如,那些瑞文渐进矩阵中的序列,可以看作是 ARC 类型的东西。有些具有较低的柯尔莫哥洛夫复杂度,有些则较高——这是理论计算机科学中的一个概念。我认为人类远未达到这方面的计算极限。我猜测这个数字可以持续上升。所以我认为流体智力可以变得非常高——基本上就是 AI 的智商,这与其长期记忆、视觉处理能力、反应时间等是分开的。但我认为它们的智商可以变得非常高,能够解决极其困难的数学问题,并对谜题有很好的直觉,而无需消耗太多算力。我认为这个数字可以大幅上升。我们非常有限——我们有大脑袋,但这里的硬件并不多。你可以想象一个更大的大脑,它将拥有相当强大的能力。即使是跨代的人类智力,也有弗林效应等现象,这些已经大大提高了智力。我们的大脑尺寸受到产道的严重限制。我预计 AI 根本不会有这个限制。它们至少可以变得越来越快。你仍然有摩尔定律,GPU 的改进速度大约是每三年翻一番。你还有可扩展性,以及比人类更好的状态转移能力,因为许多数字计算比模拟更精确。所以我认为它们有各种显著的优势,这些优势可以真正复合和积累。如果你说这主要是外部计算的,那么人类有邓巴数——大约 130 个有意义的社交联系。AI 可以做到数千、数万、数百万,所有这些结合在一起可以提供复合效应,使它们的能力大大增强。所以,是的,我认为沿途可能会有各种饱和点。你可能会得到很多低垂的果实,比如趋同的经济体——许多经济体赶上了美国附近的水平,但它们并没有以同样的速度继续增长;它们只是复制了现有的东西,然后从那以后就更难了。所以你可能在某些方面达到人类水平,而在某些领域,如果没有良好的测量或反馈循环,可能很难继续改进。但我认为情况可能相当复杂。所以我同意一些观点,但从概念上讲,我仍然认为智力有很大的上升空间。
Yeah, I don't know if it would necessarily saturate. Certainly if they're at some capability level, this might affect the rate of improvement substantially. And obviously there has to be a limit because of physics somewhere, but I would imagine there's a lot of room for improving its ability to handle things of more and more complexity. For instance, some of those sequences in those Raven's progressive matrices, you can think of as ARC-type things. Some have lower Kolmogorov complexity, some have higher Kolmogorov complexity—that's a notion from theoretical computer science. I don't think humans are anywhere near the computational limit for that at all. I would guess that number could keep going up. So I think fluid intelligence could get really high—basically the IQ of the AIs, which would be separate from their long-term memory, visual processing ability, reaction time, etc. But I think their IQ could get really high, and they can solve extremely difficult mathematics problems and have really good intuitions for puzzles, without taking much compute. I think that number could keep going up quite a bit. We are quite limited—we've got big brains, but there's not much hardware here. You can imagine a bigger brain that would have pretty substantial capabilities. Even for human intelligence across generations, there are Flynn effects and things like that, which have greatly increased. Our brain size is very limited by what can go through the birth canal. I expect AIs wouldn't have that limit at all. They could at the very least keep getting faster and faster. You still have Moore's law, and GPU improvement rates which are about 2x every 3 years. You also have scalability, and better ability to transfer state than with humans, because a lot of digital computation is more precise than analog. So I think they have a variety of substantial advantages that could really compound and accumulate. If you're saying that it's mostly externally computed, well, humans have Dunbar's number—maybe 130 or so meaningful social connections. AIs could do thousands, tens of thousands, millions, and all these together can provide compounding effects that make them substantially more capable. So yeah, I think there may be various points of saturation along the way. You may get a lot of low-hanging fruit, like convergent economies—a lot of economies caught up to somewhere in the vicinity of the US, but it's not like they kept going at the same rate; they just copied what was lying around, and then it was harder from then on. So you may have getting to human level in some respects, and in some domains it may be harder to keep getting better if you don't have a good measurement or feedback loop. But I think it could be quite complicated. So I'm agreeing with some points, but conceptually I still think there's a lot of room to go up in intelligence.
是的。是的。我的意思是,正如你在论文中所说,当机器能像人类一样做事时,我们就麻烦大了。但我认为我们之间主要的哲学分歧在于,我们都同意智力关乎适应性,但我更倾向于专门化的智力。我不认为存在所谓的通用智力。当然,谈到我们思考的这些分解方式,我们有这些约束,而且它们根深蒂固。我们利用对称性来看待世界,并且存在一个庞大的知识和思维系统发育树,它约束着我们的思维方式。AI 肯定也需要以类似的方式受到约束。我理解你的意思——你可以将其扩展一百万倍,但也许需要扩展得更多。也许我们拥有的这些创造性直觉和洞察力并非来自数据。它们并非来自内部。也许它们来自外部。所以我隐约倾向于这样一种直觉:在这种估计中,还有一些东西没有被考虑到。
Yes. Yes. I mean, as you said in your paper, when machines can do things as well as humans can, then we're in big trouble. But I think the main philosophical difference between us is that we agree intelligence is about adaptivity, but I'm a fan of specialized intelligence. I don't think there is such a thing as generalized intelligence. Certainly, talking about these factored ways in which we think, we have these constraints, and they run deep. We see the world using symmetries, and there's this big phylogenetic tree of knowledge and thought which constrains how we think. Surely AIs would need to be constrained in a similar way. I appreciate what you're saying—you could just scale this up a million times faster, but maybe it would need to be scaled much more than that. Maybe these creative intuitions and insights we have don't come from the data. They don't come from what's inside. Maybe they come from what's outside. So I'm vaguely leaning towards this intuition that there's something else which is unaccounted for in this estimation.
有趣。我的意思是,它们当然可以拥有很多传感器。如果我们说需要从其他地方获得大量额外的多样性,而不仅仅是数据中心内部的东西,我认为它们可以聚合大量信息,吸收并处理更多信息。所以我仍然认为它们至少可以在某些方面比人类快得多。但当然,如果它们处于真空中,只与自己对话或只独立工作,例如不学习,如果那个群体中的相关性太高,那么探索预算可能会太低,从而缺乏足够的多样性。我的意思是,进化通常——我认为是费希尔基本定理之类的东西——适应速度在某种程度上与变异量成正比。你指出了缺乏变异的一些方面,但我认为其中一些是可以弥补的。可能,它至少可以拥有人类拥有的传感器,甚至更多。
Interesting. I mean, they certainly could have a lot of sensors. If we're saying that there needs to be a lot of extra variety from elsewhere, and it wouldn't be just what's inside the data center, I think they could aggregate a lot of information and soak that up and process a lot more of that. So I still think they could have some type of advantage of at least doing this a lot faster than people. But certainly if they're in a vacuum, only speaking by themselves or only working by itself, for instance not learning, if there's too much correlation in that population, that might have the exploration budget be too low and then it wouldn't have sufficient variety. I mean, evolution generally—I think it was Fisher's fundamental theorem or something like that—which is that the rate of adaptation is in some ways directly proportional to the amount of variation. You're sort of pointing at ways in which it's lacking in variation, but I think that some of that could be made up for. Potentially, it could at least have the sensors that humans have, and more.
是的。
Yeah.
你的论文中还有一个非常有趣的点,因为当我思考 AI 风险时,我们之前在节目中也讨论过一些,涉及稳定与不稳定以及攻防关系。你用了‘进攻主导’这个词。你说如果 AI 是进攻主导,那么防守方就无法跟上。因为目前,想象一下我们的现状,我们有一种纳什均衡,对吧,攻防双方存在相互制衡的因素。能给我讲讲吗?
Another very interesting thing in your paper, because when I think about AI risk in general, we've thought about this a little bit on the show before, in terms of stability and destabilization and the relationship between offense and defense. And you use this term 'offense dominant.' And you were saying that a destabilizing force would be like if the AI is offense dominant, then the defensive side of the equation couldn't catch up. Because at the moment, if you imagine our state of affairs, we have a kind of Nash equilibrium, right, where there are these countervailing factors on the offense and the defense side. Can you tell me about that?
是的。我认为攻防平衡因领域而异。比如,许多信息战可能更偏向防守主导,例如关于世界的辩论,这可能是言论自由等事物的一个原因。与此同时,其他事物可能更具二元性,例如,真正专家级或非常称职的计算机安全团队可能会体验到更多的攻防平衡:发现漏洞,我们很快修补,攻击者与防御者保持同步。在其他领域,比如关键基础设施的软件,则更偏向进攻主导或攻击者优势,因为很多软件更新缓慢,存在互操作性限制,开发者已不在,软件是 30 多年前编写的,甚至没人知道它的存在。诸如此类。缺乏足够的经济激励,还有正常运行时间要求,因此关键基础设施的软件更像是活靶子,没有良好的攻防平衡。生物武器也是如此,我们当然有医药,但并非所有疾病都有解药。事实上,我们花费大量精力寻找许多疾病的疗法和应对某些病毒的方法。所以,并非‘出现新病原体,一天后就能找到解药,然后全球普及,一切搞定’。存在显著延迟,因此那里也是进攻主导。你可以想象攻击者拥有巨大优势,比如在任何人出现症状之前就在社会中传播,而缺乏各种监控机制。我们在某些方面是活靶子。所以这因领域而异。我认为网络安全的某些部分有良好的攻防平衡,其他部分则没有。生物似乎相当进攻主导。我认为这会影响你如何推广技术。如果它具有潜在灾难性且进攻主导,那你就需要……基本上,如果是大规模杀伤性武器状态,你不想把它给每个人。你不想给每个人核弹来让所有人安全。这不是办法。人们需要相互约束意图,而有些人就是无法被约束。与此同时,其他东西如防御措施、家庭安全系统或围栏等,你希望更多推广。所以我认为这应该影响对特定 AI 能力的态度,即其进攻主导性如何。这可能随时间变化。也许你可以改进关键基础设施,使其更具攻防平衡,然后你可以在更少保障措施下推广 AI。假设经济更富裕,GDP 更高,那么国家更愿意在生物武器预防措施上花钱。他们愿意采用更多远紫外线设备,进行更多废水监测等。这使情况更接近平衡,或至少减少攻击者优势。因此,你可能希望稍晚些时候,在有了这些保障措施后,再大规模推广这些 AI 能力。
Yeah. So I think the offense-defense balance varies a lot by domain. Potentially, like many information battles, it might actually be, for instance, debates about the world might be a bit more defense dominant, which would be a reason for things like free speech. Meanwhile, other things might have more of a duality where an increase in the... so for instance, really expert-level or really competent computer security teams might experience more of an offense-defense balance where something is identified, we patch the vulnerability very quickly, and so the attackers keep up with the defenders quite well. In other domains, like the software for critical infrastructure, there's more of an offense dominance or attacker's advantage because a lot of the software just doesn't get updated quickly. There's interoperability constraints. The software developer is no longer around. The software was made 30 plus years ago. Nobody even knows that it's there. Things of that sort. There aren't strong enough economic incentives for doing this. There are uptime requirements, and so software on critical infrastructure for various forms of critical infrastructure is more of a sitting duck, and there you don't experience a good offense-defense balance. Likewise for bioweapons, certainly we have medicine, but we don't have cures for everything. In fact, we spend a lot trying to find cures for many sorts of diseases and how to address certain viruses. So it's not necessarily the case that 'oh, there's a new pathogen and we'll just find a cure a day later and it will be mass proliferated across the globe and everything's taken care of.' There is a substantial delay, and so it is more offense dominant there too. And there are ways you can imagine the attacker having a really substantial advantage, like it propagating throughout society before anybody shows symptoms and lacking various monitoring mechanisms. We're kind of sitting ducks for some of those. So it varies by domain. I think some parts of cyber have a good offense-defense balance. Other parts of cyber don't. Bio seems pretty offense dominant. And I think this affects how you want to propagate the technology. If it's potentially catastrophic and offense dominant, that's something you want... basically if it's a WMD state, that's not something you want to give to everybody. You don't want to give everybody a nuke to make everybody safe. That's not how it works. It's people constraining each other's intent, and some people just won't have their intent be that constrainable. Meanwhile, other things like defenses, home security systems or fences or whatever, you'd want to propagate more. So I think that should affect the attitudes for specific types of AI capabilities as well, of what's its offense dominance. Maybe that could change across time. Maybe you could improve critical infrastructure so that it becomes more of that offense-defense balance, and then you can propagate AIs with fewer and fewer safeguards. Say the economy gets a lot richer, GDP is higher, then states are more willing to spend money on preventive measures for bioweapons. They're willing to have more far UV things. They're willing to do more wastewater monitoring, etc. And that makes things closer to there being more of a balance, or at least not having as much of an attacker's advantage. So you may want mass proliferation of these AI capabilities a bit later once you have some of those safeguards in place.
是的。另一个我非常感兴趣的点,你在论文中也提到过,就是失控的概念。Connor Leahy 用了‘战争迷雾’这个词来形容层层复杂性累积,导致完全不可读。从某种意义上说,我们现在已经这样了。我的意思是,我认为今天的 AI 也有同样的问题,但没那么极端。谷歌的系统非常复杂,这就是为什么市场奖励谷歌工程师,他们薪水很高。但你说的是当这种情况走向极端时会发生什么。你举了三个例子:自我强化依赖、不可逆纠缠和权威终止。能解释一下吗?
Yeah. So another thing I'm very interested in, which you've spoken about in the paper, is this concept of loss of control. And Connor Leahy uses the term 'fog of war' to talk about just as the layers upon layers of complexity builds, there's this complete illegibility that builds. And in a sense, we already have that now. I mean, I do think that the AI we have today has the same problem, right, but it's not quite as extreme. So systems at Google are very complex, and that's why the market rewards Google engineers. They get paid an awful lot of money. But you're talking about what happens when this gets taken to the extreme. And you gave three examples: self-reinforcing dependence, irreversible entanglement, and cessation of authority. Can you explain what those are?
是的。我在‘自然选择偏爱 AI 而非人类’那篇文章中更详细地描述了这一点。但确实,我们正变得越来越依赖这些 AI 系统。我们开始将更多决策和认知过程交给它们,而且经济层面会有越来越大的压力让我们继续这样做,没有明确的上限。你与一家拥有比你更优秀的 AI 系统的公司竞争,那家公司会赢,因为成本更低。在军事层面也是如此,有很强的动力让无人机更自主,因为信号干扰会降低其效能,所以让它们自主移动。这些压力将持续存在,以至于我们会自愿将社会中的大量权力让渡给这类系统。如果我们以某种速度或方式这样做,而实际上并未掌控局面,那可能令人担忧。控制究竟是什么样的?什么时候算过度?什么时候变得不可逆?这是我们必须关注的问题。有不同的结果。一种是实际上已经失控,你无法阻止、无法逆转、无法谈判、无法有效引导,结果你的物种适应性崩溃,或者你的生计消失。
Yeah. So I describe this at more length in the 'Natural Selection Favors AI over Humans' thing. But yeah, there's a way in which we are becoming more dependent on these AI systems. We're starting to cede more of our decisions and cognitive processes to them, and there will be more and more pressure to keep doing that without any clear limit at the economic level. You versus a company that has AI systems that are just better than you, that one's going to win because those will be cheaper. As well as at the military level, there's a very strong incentive to, for instance, make the drones more autonomous because they can be jammed. Signal jamming makes them a lot less effective, so make them autonomously move around. So these pressures will keep up such that we'll just voluntarily acquiesce a lot of power in society to these sorts of systems. And if we do it at a rate or where we aren't actually in control, that could be concerning. What does control exactly look like? When is it too much? When is it too irreversible? That's a problem that we have to keep track of. There are different outcomes. One is where you've actually just sort of lost control, and you can't stop it, you can't reverse it, you can't bargain with it, or you can't substantially steer it, and your fitness as a species is just collapsing as a consequence, or your livelihood evaporates.
或者还有一种不同的结果,你就像一个拥有巨额退休基金的退休人员,AI 在为你工作,执行你的命令,即使你并没有做出几乎所有关键决策。所以这些是不同的结果,而且可能很隐蔽或微妙,我们究竟是拥有某种反事实控制,还是最终陷入不幸的境地。我认为至少有一件事可以有所帮助,那就是让 AI 更擅长预测或预见结果和后果。如果我们训练它们这样做,那可以帮助我们更加审慎,避免一些极端依赖和控制权丧失的结果。但这还不够,不过它有助于看得更远,知道我们正在陷入什么境地。
Or there's a different outcome where you are like a retiree with a big retirement fund, and the AIs are working for you, doing your bidding, even though you aren't making nearly all of the critical decisions. So those are different outcomes, and it could be insidious or subtle as to whether we actually have some of that counterfactual control versus whether we wind up in an unfortunate situation. I think at least one thing that can help with this would be if AIs are better at forecasting or foreseeing outcomes and consequences. If we train them to do that, it could help make us more prudent and avoid some of these outcomes of extreme dependence and erosion of control. But that wouldn't be sufficient, though it would be helpful for seeing farther and knowing what we're getting ourselves into.
太棒了。丹,今天能邀请你上节目真是非常愉快和荣幸。非常感谢你今天加入我们。
Wonderful. Well, Dan, this has been an absolute pleasure and an honor having you on the show. Thank you so much for joining us today.
是的,谢谢你邀请我。
Yeah, thank you for having me.