Recursive Self-Improvement and the Path to Superintelligence
打开互动全文版(中英对照 + 朗读 + 问答)→Ryan Greenblatt 探讨了递归自我改进导致 AI 快速进步并在 2030 年代初实现超级智能的可能性。
Ryan Greenblatt discusses the plausibility of recursive self-improvement leading to rapid AI progress and superintelligence by the early 2030s.
今天和我聊天的是 Ryan Greenblatt,他是 Redwood Research 的首席科学家,专注于技术性 AI 安全与安保工作。我想和你聊聊递归式自我改进。这个概念是:一旦你构建出人类水平的智能体,它们会迅速弹射向数百亿个超级智能体,每一个在各自领域都比顶尖人类专家更有能力。无论最终是否如此,我认为这可能是当下世界上最重要的问题。而且历史上我一直对此持怀疑态度,但你似乎认为这有可能,所以我想听听你的论证。
Today I'm chatting with Ryan Greenblatt, who is the chief scientist at Redwood Research, where he focuses on technical AI safety and security work. I want to talk to you about recursive self-improvement. This is the idea that once you build human-level intelligences, they quickly slingshot towards tens of billions of superintelligences, each individually more competent than the top human experts across every field. Whether or not this turns out to be the case, I think is actually probably the most important question in the world right now. And historically I've been quite skeptical that this kind of thing happens, but you seem to think that it might be plausible, and so I wanted to hear the case for it.
是的,我们来谈谈这个。首先,我认为值得指出的是,研发是一种 AI 特别擅长的任务,因为各家公司都在努力让它们的 AI 擅长研发,而且从当前 AI 发展的角度来看,这个领域有很多不错的特性。比如它相当可验证。你可以迭代地做很多事情,在各种指标上攀升。然后我认为,一旦你拥有大致能匹敌顶尖人类专家的 AI,那可能会启动一个反馈循环:AI 在做 AI 研究,产生更聪明的 AI,再反馈回来,这个循环可能强到让你在短时间内取得大量进展。我大概的中位数预期是,一年内取得四到五年的 AI 进展。这需要真正克服研究中的大量收益递减,基本上相当于我们在大规模算力扩展后能获得的进展。所以这是一件相当了不起的大事,而且值得记住,五年的 AI 进展、四年的 AI 进展,甚至三年的 AI 进展,都是非常多的 AI 进展,对吧?你知道,现在大概是三年前或三年多一点,GPT-4 发布了,而现在我们当然有像 Mythos 5 之类的模型,也许 Anthropic 内部还有一个更好的模型。所以那在三年多一点的时间里就是巨大的进展,如果我们说的是五年,那可能更像是从 GPT-3 跳到 Mythos 5 之类的。
Yeah, let's talk about this. So first I think it's worth noting that R&D is a type of task at which the AIs are especially good because both the companies are trying really hard to make their AIs good at R&D, and it's the kind of domain that has a lot of nice properties from the perspective of how AI development works right now. So it's like pretty verifiable. You can do a bunch of stuff iteratively and climb on various metrics. And then I think once you have AIs which are roughly matching the top human experts in R&D, that could sort of kick off a feedback loop where the AIs are doing AI research that produces smarter AIs that feeds back in, and that feedback loop could be strong enough that you end up with a lot of progress in a short period of time. Maybe my sort of median expectation is something like four or five years of AI progress in a single year. And this requires really overcoming a huge amount of diminishing returns in research and basically doing the equivalent of what progress we would have gotten after a really large compute scale-out. So this is like a pretty impressive big thing, and it's worth keeping in mind that five years of AI progress, four years of AI progress, even three years of AI progress is really a lot of AI progress, right? So you know right now it's like three years ago or a little over three years ago there was GPT-4 that had come out, and right now of course we have like Mythos 5 or whatever, and maybe a somewhat better model that Anthropic has internally. And so that is just a huge amount of progress in a bit over three years, and if we're talking about five years, then maybe we're talking more about like a jump from GPT-3 to Mythos 5 or whatever.
是的。好的。所以我认为这个论证有三个不同的部分,现在我想逐一评估。第一是 AI 研发非常可验证的论点。第二是如果你自动化 AI 研发,你可以在一年内取得四到五年的进展。第三是,当 AI 研发被自动化时,从当前起点出发,以当前速度经历四到五年的 AI 进展后,另一端会出现什么。
Yeah. Okay. So I think this argument has three different parts and now I want to evaluate each one of them. First is the argument that AI R&D is very verifiable. Second is the argument that if you automate AI R&D you could get four or five years of progress in a single year. And third is the argument that what comes out the other end of four or five years of AI progress at the current pace starting at the current starting point whenever AI R&D is automated.
是的。另一端出现的是一个 AI,你可以把它投入到你能想象的几乎任何工作中。你可以把它放到 1940 年代的得克萨斯政治中,它能胜过林登·约翰逊。你可以把它放到台积电,它能学会如何做更好的工艺工程。它肯定比我的视频编辑更好——我的视频编辑非常优秀——但总的来说,它在任何它尝试做的工作上都比人类更好。
Yeah. What comes out the other end is an AI where you can drop it on the job at basically anything you can imagine. You can drop it in Texas politics in the 1940s and it outmaneuvers Lyndon Johnson. You can drop it in TSMC and it learns how to do better process engineering at TSMC. It's certainly a better video editor than my video editors—my video editors are very excellent—but it is just in general better than humans at any given job that it finds itself trying to do.
所以我想评估所有这些子论证,它们基本上会导致在这个基准之后很快得到 ASI,你预期是在 2030 年左右,对吧?
So I want to evaluate all of these subarguments that lead to basically getting ASI pretty soon after this benchmark, which you're expecting by 2030 or something, right?
是的,我会说我预期研发的完全自动化大概在 2031 年、2030 年左右,然后达到“在所有工作上胜过人类”的里程碑,我预期的中位数大概在 2033 年。但有点像,如果我看到 AI 完全自动化研发,我想我预期那大概在一年内。就像,预测的方式是,中位数之间的差异大于里程碑之间的中位数差异。总之,随便吧。
Yeah, I would say that I expect full automation of R&D perhaps somewhere around 2031, 2030, and then getting to the 'beats all humans on the job' milestone, maybe I expect median around 2033. But sort of like if I see AI fully automating R&D, I think I'm expecting that probably within a year. They're just like, the way the forecasting works out, the difference between medians is bigger than the median difference between milestones. Anyway, whatever.
顺便说一句,互联网上有个梗,因为每次我问 Dario 或别人关于时间线的问题时,我总是说:“好吧,还要多久才能自动化我的视频编辑?”然后有个梗是,我的视频编辑每次听到这个都在剪辑播客。我这么做的原因是,我觉得当你谈论你不了解的工作时,很容易迷失在抽象中,而要非常具体地理解自动化一项工作需要什么,我实际上理解为什么 LLM 目前难以接管。
By the way, there's this meme on the internet because every time I'm trying to ask about people's timelines when I'm asking Dario or somebody, I'm always like, 'Okay, how long before getting automated my video editors?' And there's this meme of my video editor editing the podcast every time they listen to this. The reason I do it is because I think it's easy to get lost in abstractions when you talk about jobs you don't understand well, and to very concretely understand what it takes to automate a job that I actually understand why it's difficult for LLMs to currently take control over.
我确实认为自动化你的视频编辑的里程碑早于能够自动化所有人类工作的里程碑,包括像得克萨斯政治那样在工作中快速上手。所以我认为视频编辑的自动化可能更接近研发的完全自动化,但它对人们真正专注于理解视频的程度非常敏感。
I do think that the milestone for automating your video editor is earlier than the milestone of being able to automate all human jobs, including like Texas politics spinning up on the job. So I think I do think that the video editor automation maybe occurs more like around full automation of R&D, but it's very sensitive to how much people are really focusing on understanding video.
是的。好的。那么我们先从 AI 研发非常可验证这个说法开始。
Yeah. Okay. So let's start with the claim that AI R&D is very verifiable.
是的。这有几个不同的部分。其中之一是,我们可以在很多环境上训练,这些环境基本上直接训练模型去做一些 AI 研发任务或非常接近的任务。例如,我们可以有一个环境,模型在仅用八块 H100 或类似的小规模算力上训练某个 AI,那个模型可能相当于 GPT-2 medium 之类的,然后类似于 nanoGPT medium 的运行,在强化学习中就是不断调整和迭代。我们可以对很多不同的任务这样做,比如让它训练图像分类模型、视频生成模型、图像生成模型,各种不同的机器学习训练任务。我们还可以训练它去训练越来越好的模型,以及做类似“这里有一个算法可以追求的方向,你能去实现它吗”的事情。所以基本上有整整一类可容器化、可验证的小规模研发任务,我们可以对 AI 进行激进的强化学习。我想说,公司可能已经在这些任务上做一些强化学习了,你可以继续扩大规模,继续制造更多这类小规模研发任务,然后 AI 可以在这方面不断进步。而且我隐含地声称这会迁移到研发中极其承重的方面。但也许我们先停一下。让我讲到那部分。
Yeah. So there's a few different parts of this. One of them is that we can train on a bunch of environments which are basically directly training the model to do some AI R&D task or some very close-by task. For example, we can have some environment where the model is training some AI on just like eight H100s or whatever, or like some small amount of compute, and that model could be like the equivalent of like GPT-2 medium or whatever, and then similar to like nanoGPT medium runs or whatever, and in RL it's like tweaking and iterating on that. And we could do that for a bunch of different tasks, like we could have it train image classification models, video generation models, image generation models, all kinds of different ML training tasks. And we could train it on the task of training increasingly good models and also doing things like, 'Oh, here's a particular direction you could pursue for an algorithm, can you go and implement that?' So basically there's this whole class of containerizable, verifiable, small-scale R&D tasks that we can aggressively RL the AIs on. And I would say that already companies are presumably doing some RL on these sorts of tasks, and you could just keep scaling that up, keep making more of these sort of small-scale R&D tasks, and then the AIs could keep getting better at this. And implicitly I'm claiming this will transfer to extremely load-bearing aspects of R&D. But maybe let's stop there for a second. Let me get to that part.
所以你可以想象,我们有 GPT 7.5,我们说:“GPT 7.5,我们想让你在 AI 研发方面变得非常出色,以至于你能帮助我们训练 GPT 9。”
So you can imagine that we have GPT 7.5 and we say, 'GPT 7.5, we want to make you so good at AI R&D that you help us train GPT 9.'
好的。所以现在我们想训练 GPT 7.5,我们想出了一系列不同的环境。正如你提到的,已经有一个仓库,是 Andre Karpathy 的 nanoGPT 速通的衍生项目,在那里你尝试改变模型的一切——从优化器到超参数再到架构——以便尽快达到固定的训练损失。
Okay. So now we want to train GPT 7.5, and we come up with a bunch of different environments. As you mentioned, there's already this repo that is the descendant of Andre Karpathy's nanoGPT speedrun, where you just try to change everything about the model—from the optimizer to the hyperparameters to the architecture—to get it to a fixed training loss as fast as possible.
你还可以有其他类型的环境,比如你可以说:“嘿,GPT 7.5,我想让你训练一个非常擅长玩电子游戏的模型。”而且我想让你训练一个在反复玩同一个游戏时不断进步的模型。这样你就学会了如何帮助模型在在线学习方面变得更好。也许它需要某种疯狂的神经或向量记忆,也许是一些疯狂的东西,也许只是更好的长上下文处理。我们不在乎你怎么解决。想办法做在线学习的研究。
You could have other kinds of environments where you could say, 'Hey, GPT 7.5, I want you to train a really good video game playing model.' And I want you to train a model that actually improves as it plays the same video game again and again. So you learn how to maybe help the model get better at online learning. Maybe it gets—we don't care how you figure this out. Maybe it's some kind of crazy neural or vector memory. Maybe it's some crazy—maybe just better long context stuff. We don't care. Figure out how to do online learning research.
显然,到那时,GPT 7.5 已经变得非常擅长常规任务——就像现在模型变得越来越聪明一样。它会像现在模型在编程方面越来越强一样,变得越来越擅长编程。你可以想象一百个类似这样的环境,它们通过让 GPT 7.5 开发 GPT 2 规模的模型(比如容器化版本)来激励 AI 研发能力,等等。
Obviously, then, the fact that GPT 7.5 will already have become very good at normal—like, it'll be a smart model in the same way the models currently are getting smarter. It'll be better and better at coding in the way that models are currently getting better at coding. And you can imagine a hundred other environments like this which are incentivizing the ability to do AI R&D by getting GPT 7.5 to—like containerized versions of—getting GPT 7.5 to develop GPT 2 size models, etc., etc.
然后你基本上让 GPT 7.5 经历一系列这样的训练。你构建了 GPT 8,而 GPT 8 现在是一位出色的机器学习研究员。它从这些训练中获得了大量的直觉。
And you basically then put GPT 7.5 through a bunch of this kind of training. You build GPT 8, and GPT 8 is now an amazing ML researcher. It has so much intuition from doing all this kind of training.
老实说,对我来说一个巨大的直觉泵是看到 AI 在数学领域取得的进步,我就想,如果这是一个非常可验证的领域,AI 就能做到——即使数学也涉及很多——我并不真正了解数学研究的对象级细节,但我就是觉得,不,它行得通。如果你能完全把它放入验证循环中,它就能像洪水一样涌来,并且能真正取得新的突破。
Honestly, a huge intuition pump for me is seeing the progress that AI has made in mathematics, where I'm just like, if it's a very verifiable domain, AI can get—even mathematics also involves so much—I don't really know the object-level details of mathematics research, but I'm just like, no, it works. It can just come in like a flood if you can totally put it into a verification loop, and it can actually make new breakthroughs.
我很好奇,机器学习研究是否具有数学研究的特质,或者似乎存在一个很大的悬而未决的问题,即连接不同学科,或者那些没有人能同时精通代数几何和——合适的词是什么?
I am curious if ML research has a quality of mathematical research, or it seemed like there was a big overhang from connecting different disciplines together, or ideas that no one person would have known enough about algebraic geometry and—what was the right word?
哦,天哪,我真的不太了解数学突破。
Oh man, I really don't know about the math breakthroughs.
没有人能同时精通拓扑学和代数什么的,等等等等,以便对一个大猜想提出反例。
No one person would have known enough about topology and algebraic whatever, blah blah blah, in order to make some counterexample to a big conjecture.
我的观点是,机器学习比数学更浅一些。因此,那种拥有深厚专业知识的专家组合的情况会少一些。但肯定会有一些这样的情况。
My view is that ML is a less deep domain than math. And so there's less of a thing where there are individual experts with really deep expertise in some area that they combine. But there's definitely going to be some of that.
但我也认为,机器学习在某些方面比数学更有利于 AI 训练,尤其是。你可以更好地判断自己是否成功,并且能看到中间进展。
But then I also think that ML has some attributes that make it even more favorable than mathematics in some ways to, you know, AI training in particular. There's—you can get a better sense of whether you're succeeding, and you can see intermediate progress.
在数学中,通常很难判断你是否接近成功。而如果你的目标是,比如,将训练损失降低 2 倍,你就能看到自己是否已经完成了一半。
So in math, it's often the case that there's no easy way to see whether or not you're close to success. Whereas if your goal is, for example, to get to some training loss, you know, 2x faster, you can kind of see when you're halfway there.
而且机器学习创新往往是叠加的,或者取决于你如何看待,可能是乘性的,基本上你可以不断叠加创新,通常创新只是相加而不会相互干扰,尽管显然这取决于具体细节。
And it tends to be the case that ML innovations are very additive or maybe multiplicative depending on how you think about it, where basically you can keep stacking innovations, and usually the innovations just sort of add together and don't interfere with each other, though obviously it's going to depend on the details.
所以我认为,在很多方面,AI 研发将具有与数学非常相似的性质,基本上你可以做小规模——你可以在与你真正关心的问题结构非常相似的 AI 研发片段上进行训练,以非常可验证的方式,然后这会产生迁移。至于迁移效果如何,这是一个开放问题,但我认为目前数学的迁移看起来相当不错,我预计 AI 研发的迁移会相当不错,但不会惊人。
And so I think that in a lot of ways, AI R&D will have properties quite similar to math, where basically you can do small—you can train on chunks of AI R&D that are pretty similar in structure to the problem you actually cared about, in a very verifiable way, and then that will transfer. And then there's an open question of exactly how well it will transfer, but I think that the transfer currently for math looks pretty good, and my expectation is that the transfer for AI R&D will look pretty good but not amazing.
所以我有一个担忧,我认为即使在数学领域,据我所知,我们还没有看到非常令人印象深刻的新理论。我们看到了很多令人印象深刻的可验证的具体结果,例如,为这个猜想找到反例,但我们还没有看到——比如,提出拓扑学级别的新思想,或者提出像群论这样的东西。而且机器学习研究似乎兼具这两方面的元素,但较难验证的部分,即提出新的思考问题的方式,会更难诱导。
So one concern I have is I think even in mathematics, as far as I'm aware, we have not seen very impressive new theory. We've seen a lot of impressive verifiable specific results, for example, find a counterexample to this conjecture, but we have not seen—like, come up with the idea of topology kinds of levels of things, or come up with things like group theory. And it seems like ML research has elements of both of these things, but the less verifiable thing of coming up with new ways of thinking about the problem would be harder to induce.
以缩放定律为例。显然,如果你有 2020 年左右的缩放定律的想法,你就能更好地训练 GPT 4,这存在某种验证循环。但要诱导 AI 做到“好的,我需要仔细思考应该如何缩放我的参数和数据。我可以进行哪些不同类型的调查来理解这一点?也许我可以提出像等浮点分析这样的可视化方法”,这需要更长且可能更耗算力的道路。但这似乎比“嘿,让我们让 nanoGPT 的损失降下来”更长的验证循环。
So take, for example, the idea of scaling laws. Obviously, there is some end verification loop such that you can train GPT 4 better if you have the idea of scaling laws from like 2020. But there is a longer and potentially more compute-intensive road to inducing AIs to be like, 'Okay, I got to think carefully about how I should be scaling my parameters and data. What are different kinds of investigations I could run to understand this? Maybe I can come up with a visualization like an isoflop analysis or something.' But that does seem like a longer verification loop than just, 'Hey, let's get nanoGPT loss to go down.'
是的,我们来谈谈这个。首先,我认为在数学的背景下,我要说的是,AI 能够做到相当于“婴儿的第一个新理论”之类的事情,比如它们可以通过建立联系和产生新的理解来证明有趣的猜想,比如“哦,有这个东西,AI 发现的这个构造很有趣”,或者找到了一个略有不同的思考问题的方式。我们确实看到了这一点。只是我们看到的例子不如创立群论领域那样令人印象深刻,但部分原因可能是,创立群论领域是有史以来最伟大的数学成就之一,而 AI 在数学方面还没有那么出色。
Yeah, let's talk about this. So first of all, I think in the context of math, the thing I would say is that the AIs can do the equivalent of like babies' first new theory or whatever, where for example they can just prove interesting conjectures via making connections and producing new understanding of like, 'Oh, there's this thing, this construction AI found which is pretty interesting,' or found this way of thinking about the problem that's a bit different. And we do just see that. It's just that the examples we see are not as impressive as founding the field of group theory, but in part, you know, probably founding the field of group theory is one of the biggest mathematical accomplishments of all time, and the AIs just aren't that good at math yet.
而且我认为,从我的角度来看,在这一点和我们现在看到的情况之间存在一个连续体,AI 正在不断进步。
And I think that from my perspective, there's a continuum between that and the things we're seeing now, and the AIs are continuing to march up.
第二,我认为机器学习相对于数学是一个非常浅的领域。
Second, I think ML is a very shallow domain relative to math.
所以我觉得在数学里,更多是你找到某个真正深刻的抽象,然后如果你真正理解了那个东西——而理解它很难——你就能有所突破。而我感觉机器学习里与之对应的东西其实很浅,比如 Scaling(规模扩张)定律。拜托,我们很快就能解释清楚 Scaling 定律。而数学中最深刻、最重要的概念,并不具备那种你能在很短时间内真正理解其底层机制和重要性的特性。
So I think in math there's much more of a you find some true deep abstraction and then like that, if you really understand that thing, which is hard to understand, then you get somewhere. Whereas I feel like the things that are the equivalent of that in ML are really like dumb, like scaling laws. Come on guys, we can explain scaling laws really quickly. And I think the deepest and most important concepts in math, for example, don't have the property of you can really understand the underlying thing and why it matters in a very short period of time.
但我感觉一个影响会是,到 2030 年我们已经把所有容易摘的果子都摘完了。我觉得 Scaling 定律在历史上就相当于数学里的笛卡尔坐标系、做非常基础的数学。然后最终,如果你想在 2030 年代继续取得进展,就得去做现在数学前沿正在发生的事情。
But I feel like one effect will be that we will have gotten rid of all the low hanging fruits by 2030. Like I feel like scaling laws will have been like in math history, you know, finding the Cartesian grid and doing very basic mathematics was. And then eventually, if you want to keep making progress in the 2030s, it's going to be like doing whatever is happening at the frontiers of mathematics right now.
是的,那可能是对的。我的感觉是,有些领域在运作方式和依赖深度抽象的程度上有结构性差异。物理学和数学更偏向于需要非常深刻、难以想到的想法那一端。而我认为机器学习和其他大多数领域更适合爬山式的渐进优化。这是我对未来走向的感觉。即使在你的 AI 必须努力推进的情况下——比如到了 2030 年,在大量低垂果实和研究已经完成之后,它们需要取得进一步进展——我仍然怀疑很多工作会更多地落在构建日益复杂的基础设施、对实验大致样子有良好直觉这一侧。所以我觉得我可能不太认同 AI 会缺乏某种深刻洞见的说法,而更认同它们确实需要大量关于细节实验的品味——这些它们目前还没有——并且需要对什么样的训练方法有效、什么样的无效有很多直觉,就像现在的研究人员那样。
Yeah, that could be right. My sense is that just like some domains are structurally different in terms of how they operate and how much they depend on deep abstractions. Physics and math are much more on the side of being very far on the sort of very deep, hard to come up with ideas side. Whereas I think ML and most other domains are much more amenable to hill climbing. And that's my sense of how this will go in the future. And even in the regime where your AIs are, you know, having to plow through, it's 2030, they need to make further progress after a bunch of low hanging fruit and research has already happened. I still suspect that a bunch of the work will live more on the side of building increasingly complicated infrastructure, having really good intuition about what the experiments roughly look like. And so I think I'm probably less sympathetic to the idea that the AI will lack some deep insight, and more sympathetic to the idea that they really need a bunch of taste about in-the-weeds experiments that they currently don't have, and need to have a bunch of intuition for what sorts of training approaches would work and what wouldn't, in ways that current researchers have.
即使在 AI 领域有一些突破,事后看来,实现那个突破的一个大瓶颈往往是搞定所有微小的细节和琐碎的直觉。比如训练 AI 擅长推理和思维链,做强化学习和思维链训练。看起来你本可以在 GPT-3 上做强化学习和思维链,如果真正扩大规模并做好,就能在数学上得到一些有趣的结果。但当时有低垂的果实,而且做好那种训练在技术实现、扩大规模、调对超参数等方面都很琐碎。所以也许你可以在 Qwen 1B 之类的模型上演示一切,并感觉到整个事情会成功,但人们没有尽早演示,就是因为所有这些琐碎的细节和关于如何精确调整参数、如何设置的直觉。
And even in cases where there's been some breakthrough in AI, oftentimes in retrospect it looks like a big bottleneck to making that breakthrough happen was getting all of the micro details and mung intuition right. Like an example of this is when it comes to training AIs to be good at reasoning and chain of thought and doing RL and chain of thought training. It looks like you probably could have done RL and chain of thought on GPT-3 and gotten kind of interesting results on math if you had really scaled it up and done a good job. But at the time there was low hanging fruit, and also doing a good job with that training is kind of in the weeds on all the technical implementation and scaling it up and getting the hyperparameters right. And so maybe you can demonstrate everything on like Qwen 1B or whatever and get some sense that this whole thing is going to work, but people didn't demonstrate it as early as they could have because of all of these other munchy details and intuition about exactly how to tune the parameters and how to set things up.
老实说,这是我对这个故事剩余的怀疑。我不确定我理解为什么,如果研究突破如此依赖于智能,AI 进展在历史上却没有比原本可能的更快。而且我们不得不等待,就像你说的,等到 RLVR 真正起作用的时候,即使你本可以用更少的算力做到,我们不得不等待海量的算力、吉瓦级的算力可用之后,人们才开始做这种训练。在这个常数轨迹上,随着算力不断增加,我们取得更多突破。我不知道,我觉得 2022 年有很多 AI 研究人员试图攻克推理,但他们只是被编写基础设施代码的能力或当时发生的事情所限制。
This is my remaining skepticism honestly about the story. I'm not sure I understand why, if research breakthroughs are so amenable to intelligence, AI progress has not been historically faster than it could have been. And we had to wait for, as you were saying, by the time RLVR actually worked, even though you could have done it with less compute, we had to wait for oceans of compute and gigawatts of compute to be available before people are doing this training. On the trajectory of this constant, as compute keeps increasing, we make more breakthroughs. I don't know, I feel like there were a lot of AI researchers in the year 2022 who were trying to crack reasoning, and it was just that they were bottlenecked by the ability to write infrastructure code or what was happening.
这是一个复杂的混合体,对吧?所以我认为,如果他们能一想到实验就无 bug 地运行,他们会进展得更快,bug 非常重要。然后我认为另一部分是,能够用高算力运行大量实验,可以掩盖你实现方式不太对或超参数不对的问题。所以我认为算力对做 AI 研究真的很有帮助,你可以掩盖很多事情。但这并不意味着劳动力的大幅增加没有帮助,尤其是如果这些劳动力带有该领域人们最好的直觉。我只是觉得那真的很有帮助。
It's a complicated mix, right? So I think that they would have gone faster if they could, as soon as they thought of an experiment, run that experiment without bugs, bugs being very important. And then I think another part of it is that being able to run a lot of experiments at high compute lets you paper over ways in which the way you implemented it isn't quite right or you didn't have the right hyperparameters. And so I think compute is just really helpful for doing AI research, and you can cover over a lot of things. But that doesn't mean that massive increases in labor wouldn't also be helpful, especially if that labor comes with, you know, among the best intuitions that people have in the field. I just think that that's really helpful.
我这里的另一个观点,可能和你的出发点有点不同,是我预期会有更多的迁移,比你似乎想象的要多。我设想这些 AI 实际上是相当好的通才科学家,在所有这些事情上都相当不错。当你与它们互动时,它们并不是那种高度专业化的天才型感觉。它们实际上在研发的所有事情上都相当好,然后可能在某些子领域极其出色,对吧?它们在编写内核方面极其超人,在一切反馈循环很短的事情上极其超人,然后在其他所有事情上都相当好,完全能够匹敌其他人。而且我认为我们现在已经看到了这一点。我想说,当我看到现在的 AI 时,我认为它们已经能够相当熟练地匹敌那些在 ML 研究上平庸的人类。只是平庸的 ML 研究并没有多大帮助,对吧?你真正想要的是擅长 ML 研究的人。所以我的感觉是 AI 在这些方面都在进步。它们的品味在提高,直觉在提高。而且它们的品味和直觉已经不是完全糟糕了。
I think another part of my perspective here, which is maybe a bit different from where you're coming from, is that I think I'm expecting somewhat more transfer than you seem to be imagining. I'm imagining these AIs are actually pretty good scientists in general, and are just, you know, pretty reasonable at all of that stuff. And just sort of when you were to interact with them, it's not like they're some really hyper-specialized savant type vibe. They're actually just pretty good at all the stuff in R&D, and then maybe extremely good at some subdomains, right? They're incredibly superhuman at writing kernels, incredibly superhuman at everything with very short feedback loops, and then pretty good at all the other stuff, and totally able to match other people. And I think we are seeing this now. I would say that when I look at AIs right now, I think it's already the case that they can pretty competently match humans who are mediocre at ML research at doing ML research. It's just that being mediocre at ML research is not that helpful, right? The thing that you actually want are people who are good at ML research. And so my sense is the AIs are just improving at all of these things. Their taste is improving. Their intuition is improving. And it's already the case that their taste and intuition is not complete garbage.
是的。所以我非常想具体理解,一年的时间实现五年的 AI 进展会是什么样子。
Yeah. I so I want to very concretely understand what it would look like for 5 years of AI progress to happen in one year.
假设我们回到 GPT-3 开发的时候,想法是,以 2022 年他们拥有的算力水平,如果当时我们实现了研发自动化,你就能在那一年的年底训练出 Mythos。
So suppose we were back when GPT-3 was developed, and the idea is that with the level of compute they had back in 2022, if we had automated R&D back then, you could have trained Mythos at the end of that year.
那会是这个想法。是的。
That would be the idea. Yes.
包括像,Mythos 消耗的算力远超他们当时拥有的,但即使以他们当时的算力水平,他们不仅完成了所有突破,还用他们的算力训练出了 Mythos。而所需要的显然是发现自那以来的所有算法进步。实际上要发现更多,因为你必须弥补 Mythos 使用的……
Including with like, Mythos took way more compute than they had back then, but even with the level of compute they had back then, not only do they do all the breakthroughs, but they also train Mythos with their level of compute. And what would be required is obviously discovering all the algorithmic progress since then. Discovering even more actually, because you had to make up for the fact that Mythos uses...
我不知道 GPT-3 是在什么上训练的,比如 23?
I don't know what GPT-3 was trained on, like 23?
呃……
Uh...
我们可以查一下,但它是否可能多出四个数量级的算力?
We can look it up, but is it plausibly four orders of magnitude more compute?
是的,我认为要少一些。我们快速查一下。GPT-3 的训练算力大约是 3.23e23 FLOPs。我的感觉是 Mythos 可能高出大约三个数量级多一点。所以问题是,你能在击败模型的同时克服这一千倍的算力差距吗?
Yeah, I think it's somewhat less than that. Let's look this up quickly. So GPT-3 training compute is about 3.23e23 FLOPs. My sense is that Mythos is probably about a little over 3 orders of magnitude higher. And so the question is, can you overcome this thousand-fold compute gap while also beating the model?
所以这里有一个具体的说法,也许我们应该讨论一下。比如现在,我们将能够用 GPT-3 级别的算力训练一个模型,它匹配……我到底是怎么想的?
So here's a concrete claim that maybe we should talk about. Like right now, we would be able to train a model with GPT-3 level compute that matches... what exactly do I think?
所以 GPT-3 大约是,比如说……它是什么时候训练的?它是在 2020 年训练的,所以是 6 年前训练的。值得注意的是,GPT-3 可能有点太远了。但让我们先按这个来。所以 GPT-3 是在大约 6 年半、7 年前训练的。
So GPT-3 was, let's say, about... when was it trained? It was trained in 2020, so it was trained 6 years ago. It's worth noting that GPT-3 is maybe a little too far in the past. But let's go with this for a second. So GPT-3 was trained about 6 and a half, 7 years ago.
如果我们今天用 GPT-3 级别的算力训练一个模型,那个模型会有多好?
If we were to train a model with GPT-3 level compute today, how good would that model be?
根据算法进步的方式,我的理解是,我们将能够训练出一个和我们大约三年前拥有的最好模型一样好的模型。所以我认为现在我们将能够训练出一个 GPT-3 的版本,它可能比 GPT-4 好一些,基本上就是我们会看到的。可能是的,比 GPT-4 好中等程度。
My understanding is, based on how algorithmic progress works, we'd be able to train a model that's as good as the best model we had perhaps around three years ago. So I think that right now we'd be able to train a version of GPT-3 that's probably somewhat better than GPT-4, basically what we'd see. Probably yeah, like a moderate amount better than GPT-4.
我认为这差不多是对的。我认为这大致符合算法进步的方式。基本上,故事最终会是,要获得五年的 AI 进步,你可能需要大约,我会说,大概八年的算法进步,非常粗略。这是大量的算法进步。但事实证明,从我的角度来看,大多数 AI 进步来自于算法和数据的某种混合,你可以不断在这些方面取得巨大改进,并用更少的算力训练 AI。
And I think that's about right. I think that roughly lines up with how algorithmic progress has worked. Basically, the story would end up being that to get five years of AI progress, you're probably going to need around, I would say, maybe eight years of algorithmic progress, very roughly. Which is a lot of algorithmic progress. But it just turns out that most of the AI progress, from my perspective, has come from some mix of algorithms and data, and you can just keep making huge improvements on these things and training AIs with less compute.
所以我很高兴你提到这一点,因为从 GPT-3 甚至 3.5 到现在发生了什么,对吧?比如为什么 Mythos 这么好?显然我们扩展了算力,我们有更好的算法。一个巨大的事情是,我们建立了一个百亿美元的数据产业,它系统地收集和编纂了跨各种不同学科的专家人类判断,以强化学习环境的形式编纂,以这些专家构建的 SFT 轨迹的形式编纂,以帮助模型更好地理解如何做编码,如何构建复杂的基础设施项目,如何处理法律,如何处理其他任何事情。而 AI 如何能够复制目前专家人类判断似乎在 AI 进步中发挥的作用?
So I'm glad you brought that up, because what has happened since GPT-3 or even 3.5 till now, right? Like why is Mythos so good? Obviously we've scaled the compute, we have better algorithms. A huge thing that's happened is that we have built a deca-billion-dollar data industry which has systematically collected and codified expert human judgment across all kinds of different disciplines, codified in the form of RL environments, codified in the form of SFT traces that these experts build to help the model better understand how do you do coding, how do you build complex infrastructure projects, how do you do law, how do you do whatever. And how are the AI able to replicate the effect that currently expert human judgment seems to be playing in AI progress?
是的。所以我的感觉是,扩大获取专家人类数据的努力规模总体上对 AI 研发并不是非常重要。特别是,在过去几年里,我们一直在扩大算力,扩大在 AI 公司工作的人数,并扩大在数据标注上投入的努力。我的感觉是,如果你去掉最后两次数据标注的翻倍或什么的,那不会产生巨大差异,或者数据生成——抱歉,我应该说来自专家人类的数据生成——那不会产生巨大差异。我认为很多正在发生的事情是,人们一直在开发更好的方法来利用人类和 AI 构建强化学习环境,并从中取得进展。
Yeah. So my sense is that scaling up the amount of effort spent on getting expert human data has not been hugely important for AI R&D in general. So in particular, over the last few years we've been scaling up compute, scaling up people working at AI companies, and scaling up the amount of effort spent on data labeling. My sense is that if you removed the last two doublings or whatever of data labeling, that would not make a huge difference, or data generation—I'm sorry, I should say data generation from expert humans—that would not make a huge difference. And I think a lot of what's been going on is people have been developing better ways to leverage humans and AIs to construct RL environments and going somewhere from that.
比如你如何解释为什么 AI 在编码方面变得如此出色?我觉得很大一部分是数据和强化学习环境,它们就像是在量化人类专家。
Like how do you explain why the AIs have gotten so good at coding? I feel like a big part of that is data and RL environments which are like quantifying human experts.
但问题是,创建强化学习环境的限制因素是什么?我对创建强化学习环境的限制因素的感觉,与其说是扩大规模,或者说驱动今天强化学习环境比 2024 年好得多的原因,不是因为雇佣了更多的人类专家来制作强化学习环境,而是更多地因为我们更清楚我们到底想制作什么样的强化学习环境以及应该如何构建它们,而且我们正在使用大量的 AI 劳动力来构建强化学习环境。我认为这些影响比人类劳动力构建强化学习环境的影响重要得多。
But the question is, what is the limiting factor on creating RL environments? My sense of the limiting factor on creating RL environments was not so much about scaling up, or the thing that drove the reason why RL environments today are much better than they were in 2024 is not much because we have hired way more human experts to make RL environments, and instead much more because we better know what RL environments we even want to make and how we should structure them, and also we're using huge amounts of AI labor to build RL environments. And I think those effects are much more important than the effect of human labor building the RL environments.
我不是说人类劳动力不重要。我只是说还有其他重要的驱动因素。
I'm not saying that the human labor doesn't matter. I'm just saying there are other big drivers that are important here.
是的,我可以试着论证这一点。我的意思是,一件事就是人们想要的环境数量非常大。我认为 AI 在制作强化学习环境的任务上实际上相当擅长,只要对应该是什么样子有一些概念。有预先存在的数据可以使用。我不知道。很多这些东西都有良好的验证循环。
Yeah, I could try to argue for this. I mean, one thing is just the amount of environments people want is just a very large amount. And I think the AIs are actually pretty good at the task of making RL environments given some sense of what the thing should be. There's pre-existing data you could use. I don't know. A lot of these things have good verification loops.
如果我只是看一下,例如,昨天 Business Insider 报道说 Google 为 mechanize 支付了近 20 亿美元。
If I just look at, for example, this was reported in Business Insider yesterday that Google is paying like close to $2 billion for mechanize.
是的。
Yeah.
比如我们可以看看市场行情,人们认为真正优秀的专家制作人类专家数据的价值是多少,似乎前沿实验室需要认为它值得付出很多。
Like we can just look at market rates for what people think really good human experts making human expert data is worth, and it just seems to be that the frontier labs need to think it's worth a lot to pay for.
你认为前沿实验室在数据上的支出占比是多少,而不是算力?比如你认为算力和数据的支出比例是多少?
What fraction of frontier lab spending do you think is on data rather than compute? Like what do you think is the compute-data spend split?
我认为绝大多数是算力,但我也认为这是因为算力比数据更容易扩展。
I think it's overwhelmingly compute, but I also think it's because compute is easier to scale up than data.
这确实与推动进步的因素相关,对吧?比如,我同意我的感觉是,这个比例大概是我猜的 20 比 1 或 10 比 1,我不确定具体数字,这取决于公司。
That's really relevant to what's driving progress, right? It's like suppose I agree that yes, my sense is that the split is something like I would have guessed like 20 to 1 or something, 10 to 1, I don't know exactly, it depends on the company.
我的意思是,这就像石油占 GDP 的 1.5%,但这并不意味着如果去掉石油,GDP 还能继续增长;如果石油消失,GDP 会立刻停滞。
I mean, but this is similar to oil being 1.5% of GDP, but that doesn't mean if you cut oil out, GDP could continue to rise; it would halt immediately if oil went away.
当然,但你只是在说,因为市值高,我们可以推断这是关键驱动因素,而我说这并不明显成立,对吧?因为我认为那个论点只是让算力看起来是更重要的驱动因素,或者雇佣员工更好。
Sure, but you are just arguing that because of the high market cap, we can learn that this is the key driver, and I'm saying that's not clearly true, right? Because I think that argument just makes it look like compute is a much more important driver, or like hiring employees is much better.
也许我们说得更具体些。我是这么想的。
Maybe let's be more concrete. Here's what I think.
就像我的主张是,在 GDP 中,如果你回到 2022 年,GDP 是 3.5,你试图在没有人类专家的情况下让它更擅长编码,我认为那会非常非常困难。
Just the same way as my claim is that in GDP, if you went back to 2022 and you had GDP 3.5 and you were trying to make it better at coding without human experts, I think it would have just been very, very difficult.
让我举个例子,说明我认为从 GPT-8 到超级智能会遇到的困难。所以,你需要 GPT-8 擅长的事情之一,或者说你希望超级智能擅长的事情,就像我要接管一家公司,让它更赚钱,做各种疯狂的事情让它运转得更好。我要接管一家晶圆厂,生产更多芯片。这是我要去国会试图说服他们通过某项法案的数据层级,等等。
Let me give you an example of what I imagine would be the difficulty from going from GPT-8 to ASI. So, one of the things you'd need GPT-8 to be good at, or like you'd want ASI to be good at, is like I'm going to take over a company and make it much more profitable and do all kinds of crazy things to make it work better. I'm going to take over a fab and produce more chips. This is the tier of data that I'm going to go into Congress and try to convince them to pass some bill, blah blah blah.
是的。
Yeah.
这就是我设想以这种速度再发展五年 AI 后,AI 能够做到的事情。这正是我真正担心的事情,对吧?就像那个能理解如何在世界上做疯狂事情的超级智能,能做费曼能做的事,能做史蒂夫·乔布斯能做的事,等等,还有他的工程师们。我不确定如果没有相关的世界数据,你怎么能得到那个,这就像 Mythos 非常擅长编码,却没有相对于 GPT-3 改进它的编码环境。
This is what I imagine five more years of AI progress at this pace would enable an AI to be able to do. This is the thing I'm really worried about, right? Like the ASI that can understand how to do crazy things in the world, can do what Feynman can do, can do what Steve Jobs can do, etc., and also his engineers and stuff. And I'm not sure how you get that without the relevant world data, which is the equivalent of Mythos being really good at coding while not having the coding environments that have improved it relative to GPT-3.
是的。所以这里有几点。首先,我打赌如果你看 Mythos 的随机采样训练环境,它们实际上与在实践中使用模型的样子非常不同。我的感觉是,强化学习分布与现实世界数据分布有非常大的偏差。它通过迁移和少量聚焦于现实世界的数据混合而被显著平滑。所以我的感觉是,这将是一个类似的机制,就像你在完全自动化的研发基础上,经过 5 年 AI 进步而得到的疯狂、相当超人的 AI 一样。
Yeah. So here are a few points. So, first, I bet if you look at randomly sampled training environments for Mythos, they're actually very different from what it looks like to actually use the model in practice. My sense is that the RL distribution has really large deviations from the real world data distribution. And it's significantly being smoothed over by a mix of transfer and having a small amount of data focused on the real world. And so my sense is that this will be a similar mechanism as how it works for the crazy, wildly quite superhuman AI you get as a result of 5 years of AI progress on top of fully automated AR&D.
所以我们稍微过一下这个。特别是,我认为你可以训练一个 AI 非常非常擅长即时学习,并做类似上下文学习的事情,但可能在各种强化学习环境中使用一些不同的机制。所以你构建所有这些不同的强化学习环境,AI 必须即时适应、即时学习、弄清楚该做什么、更好地理解自己的处境,并从反馈中快速学习以达成目标,还有资源有限等条件,如果搞砸了,它可能会陷入更糟的境地。然后如果你在大量这样的环境上训练,你会学到一些通用的技能,比如即时获取上下文。我们已经看到这一点,比如 AI 现在在大致理解发生了什么、从有限的信息中获取上下文方面已经好多了。然后那些 AI 可以被放到台积电的工作岗位上。然后即使台积电并不字面上在他们的数据分布中,他们的数据分布非常广泛,AI 在他们的数据分布上极其擅长,以至于它迁移到擅长成为台积电工程师并即时学习,看起来更像是 AI 擅长成为台积电工程师的方式不是它有大量关于成为好台积电工程师的缓存知识,而是它做了某种扩展版的上下文学习。
So let's just go through this a little bit. So in particular, I think that you could train an AI to be really, really good at learning on the fly and doing something analogous to in-context learning, but potentially using somewhat different mechanisms in a wide variety of RL environments. So you build all these different RL environments where the AI has to adapt on the fly, learn on the fly, figure out what it should do, understand its situation better, and learn really quickly from feedback in order to succeed at its objective, and has things like limited resources, and if it messes up, it can end up in a much worse position. And then if you train on a huge number of these environments, you will learn sort of general skills of picking up context on the fly. And we're already seeing this, like it's already the case that AIs are now much better at sort of understanding roughly what's going on and picking up context from a limited amount of information they're given access to. And then those AIs could then be put on the job at TSMC. And then even though TSMC is not literally in their data distribution, their data distribution is really wide and the AIs are extremely good on their data distribution, such that it transfers to picking up being good at being an engineer at TSMC and learning that on the fly, where it looks more like the way the AI gets good at being a TSMC engineer isn't that it has a ton of cached knowledge on being a good TSMC engineer. It's that it does the equivalent of some scaled-up version of in-context learning.
那会是最平淡的故事。显然,有很多不同的可能路径。
There, that would be the most prosaic story. Obviously, there's a bunch of different ways this could go.
我认为这可能归结为关于你能走多远的直觉差异。当我想到我认识的真正聪明的人,他们在不太了解的领域就是没那么有效。
I think this maybe comes down to the difference of intuition about how far you can get. When I think about really smart people I know, they're just not that effective in domains they don't understand that well.
但他们有多少时间学习?
But how long have they had to learn?
不,我同意如果他们有过经验,他们会好得多。但这也许正是我主张的,即数据的经验。比如,如果我找一个非常聪明的,我不知道,常春藤盟校的毕业生,然后我说,“好吧,你现在负责谈判伊朗协议。”我觉得他们就是不知道该怎么办。
No, I agree that if they had experience, they would be much better. But that's maybe what I'm arguing for is that experience of data. Like for example, if I just get a really smart, I don't know, Ivy League college grad and I'm like, "Okay, you're now in charge of negotiating the Iran deal." I think they just wouldn't know what to do.
我认为如果你换成一个非常擅长快速掌握多个不同领域的人,并给他们一些时间训练、与人交谈、展示他们的专业知识并做一些练习,他们实际上会做得相当不错。我认为大多数领域本质上相当浅,一个非常聪明的通才,擅长有限的核心技能子集,可以很快上手。我的感觉是,并非每个领域都是如此。我的感觉是,AI 将发展出越来越好的机制,在给定领域快速获取理解和专业知识。所以考虑一下 AI 能多快理解一个新的代码库。AI 理解新代码库的速度比人类快得多,但深度比人类目前能理解的浅,但随着时间的推移会变得更好。对吧?所以让我更详细地阐述这个论点。假设你拿 Fable 5 或 Mythos 5 之类的,你想对一个非常庞大的代码库做一些复杂的更改。模型会很快获得对代码库的一些理解,比如在可能不到一个小时,甚至可能远少于一小时的过程中,然后它对代码库的理解会有点停滞,不会像人类在更长时间内获得的那样深入。所以有点像 AI 在一小时内可以匹配人类几周的水平,也许取决于代码库到底有多复杂。
I think if you instead got someone who is really good at quickly picking up a bunch of different domains and you gave them some time to train and talk to people and show up their expertise and do some practice, they would actually do a pretty good job. I think most domains are fundamentally pretty shallow, where a very smart generalist who's good at a limited subset of core skills can get going pretty quickly. And my sense is that that's not true for literally every domain. And my sense is that the AIs will develop increasingly good mechanisms for quickly acquiring understanding and expertise in a given domain. So consider for example how fast AIs can understand a new codebase. AIs can understand a new codebase much faster than humans can, but to a degree that's shallower than humans could currently understand, but is getting better over time. Right? So let me spell that argument out a bit more. So let's say you take Fable 5 or Mythos 5 or whatever, and you wanted to make some kind of complicated change to a really massive codebase. The model will get some understanding of the codebase very fast, like in the course of maybe significantly less than an hour, potentially much less than an hour, and then its understanding of the codebase will plateau a little bit, where it won't get as deep of an understanding as a human would have gotten over a much longer period. So it's sort of like an AI in an hour can match a human with a few weeks maybe, depending on the details of exactly how complicated the codebase is.
但那样的话,它还是比不上一个在那个代码库上工作了两年的真人。但随着时间的推移,AI 能匹配的理解程度已经提高了。如果我们看 3.7 Sonnet 或 3.5 Sonnet,也许它只能匹配相当于理解一个代码库一天的水平。但现在 AI 在构建任务上下文方面强多了。你可以说:“Mythos,我要你真正理解这个代码库,然后实现这个功能。”它就会生成一大堆子智能体。那些子智能体会仔细研究很多东西,带回一堆上下文,然后再调查几件事。它做这个并不算惊艳,但可以非常快,而且效果还不错。而且不难想象,你可以训练 AI 在这个任务上越来越强。在一个很大的代码库里以合理的方式实现某个非常复杂的功能,这个任务是极其可验证的,所以它可以成为 AI 改进的目标。类似地,还有一个更广泛的技能,就是快速理解上下文,让一堆不同的 AI 并行学习,然后把结果合并起来。
But then it won't match a human who's been working on that codebase for two years or whatever. But over time, the amount of understanding AIs can match has gone up. If we look at 3.7 Sonnet or 3.5 Sonnet, maybe it could only match the equivalent of understanding a codebase for a day or something. But now AIs are much better at building context about a task. So you can say, 'Mythos, I want you to really understand this codebase, then implement this feature,' and it will spawn a bajillion sub-agents. Those sub-agents will pore over a bunch of things, deliver a bunch of context back, then investigate a few things. It's not amazing at doing this, but it can happen really fast and work pretty well. And it's not hard to imagine how you could train AI to be increasingly good at this task. The task of implementing some very complicated feature in a reasonable way in a very big codebase is extremely verifiable, so that can be something the AI improves on. Similarly, there's a broader skill of quickly understanding context, having a bunch of different AIs learn in parallel, and then merging that together.
我觉得这里似乎有一个关键点,我认为这只是一个经验问题,我们拭目以待。那就是,在可验证领域里,AI 在快速理解情况、跟上进度、长期取得进展方面已经变得非常擅长——这显然进步神速——但把这些能力迁移到“去跟总统谈谈,说服他做某件事”,或者“你现在负责 Google 了,你必须让 Google 本季度更赚钱”这样的任务上,迁移效果到底有多好?
And I think there seems to be a crux here, which I think is just an empirical question we'll see, which is how good is the transfer between getting really good at understanding the situation, getting up to speed, making progress over long periods in verifiable domains—which the AI is obviously getting way better at really fast—and then, 'Okay, go talk to the president and convince him to do X thing,' or 'You're now in charge of Google. You must make Google a much more profitable company this quarter.'
让我试着再详细说明几个可能相关的论点。一件事是,我确实认为,看看 AI 在写文章方面的进步——我们稍微聊聊这个。有一点是:即使在这些领域,你也能获得一些数据,而且当 AI 处于非常快速的进步轨迹上时,它甚至能在这些领域获得一些数据。也许很难建立一个可验证的环境来评估“你的文章在人类看来真的好吗”,但你可以做一点。你可以做一些训练,一些在线训练。AI 将能够基于真实世界的东西进行一些在线训练。它们将能够进行评估,进行采样,你可以扩大做这件事的频率。第二件事是,在实践中,当我只看迁移效果时,它似乎还行。我认为事实上 AI 在非可验证领域已经进步了很多,而且事实上,很难指出哪些领域真的很难验证,而在这些领域里,从 GPT-4 到 Mythos 的进步幅度在实践中不是相当大。但这并不意味着 Mythos 比最优秀的人类还好。它仍然可能在某些工作方面明显不如典型的人类专业人士,但同时比 GPT-4 好得多,而 GPT-4 当时还差得远。
Let me try to spell out a few more arguments that might be relevant. One thing is that I do think that when looking at how AIs have improved at essay writing—let's talk about that a little bit. There's one thing: you can get some data even on these domains, and AIs will be able to get some data even on these domains when on a very fast progress trajectory. Maybe it's hard to build a verifiable environment for 'was your essay really good according to humans,' but you can do a bit of that. You can do some training, some online training. The AIs will be able to do some online training based on real-world stuff. They'll be able to have evals, sample that, and you can scale up the cadence at which you do this. The second thing is that in practice, when I just look at the transfer, it seems okay. I think that in fact the AIs have improved a bunch at non-verifiable domains, and it is in fact the case that it's hard to point to domains that are really hard to verify on which the amount of improvement between GPT-4 and Mythos hasn't been pretty high in practice. Now that doesn't mean Mythos is better than the best humans or something. It can still be significantly worse than typical human professionals at some aspect of their job while still being way better than GPT-4, which was not even close.
是的。所以,我们刚才在讨论数据和算法进步对解释过去几年进步的重要性。这让我想起来,我实际上在和 Jerry Han 一起做一个实验,他还是个大学生。我们基本上要评估有多少进步来自数据、多少来自算法,做法是用 2019 年到现在最好的算法配方,配合 2026 年数据文件里最好的数据来训练;同时,用当前最好的训练配方(也就是算法配方)来训练从 2019 年到 2026 年的不同数据堆。
Yeah. So, we're talking about how important data versus algorithmic progress has been for explaining the progress of the last few years. That reminds me, I'm actually running an experiment with Jerry Han, who's actually still a college student. What we're basically doing to evaluate how much progress is coming from data versus algorithms is training the best algorithmic recipe from 2019 till now with the best data from the 2026 data file, and also training the different data piles going back from 2019 to 2026 with the current best training recipe, like the algorithmic recipe.
嗯。
Yeah.
我觉得那会是一个有趣的……
And I think that will be an interesting...
我很好奇你是否想预先登记一下,来自数据和算法的乘数各有多少。所以我们需要非常小心地定义“数据”这个词。我一直在努力区分“增加在让人类专家标注数据上的支出”和“增加人类专家标注数据的数量”。预训练数据的改进——我们现在比 2019 年有更好的预训练数据集,并不是因为人们花更多钱让人类专家输入 AI 训练用的数据。我认为部分是……我认为那部分不多。我认为预训练数据的改进中,人类专家标注的贡献非常小。我认为绝大多数预训练数据的改进——明确说,我指的是预训练;我们或许应该单独讨论中期训练和后训练——但绝大多数预训练数据的改进,来自科学上更好地理解哪些数据集是好的,以及来自弄清楚如何筛选数据的辛勤劳动。所以我的观点是,像从 open web text 到 fine web 这样的改进,最好被描述为一种算法改进,你可以用一些 GPU 来研究,然后实施,不需要人类专家数据。现在还有一个不同的效应,我们可以讨论,那就是也许 2026 年的互联网比 2018 年的互联网更适合作为训练数据的沃土。还有一个效应是,有更多人类在互联网上发帖,所以有更多数据可以采集。我的感觉是,这个效应会比“人类更懂得如何整理数据、有更好的抓取、知道如何更好地处理这些抓取”这类效应小得多。
I'm curious if you want to pre-register what amount of multipliers are coming from one versus the other. So we need to be pretty careful with what we mean when we say the word 'data.' I was trying to be pretty careful to distinguish between scaling up spending on getting human experts to label data, or scaling up the amount of human expert label data. Pre-training data does not come—the reason why we have a better pre-training data set now versus in 2019 is not because people are spending way more money getting human experts to type up data that the AIs are then trained on. I think it's partially... I think it's not much of it. I think it's very little of the pre-training data improvements. I think the vast majority of the pre-training data improvements—which, to be clear, I do mean pre-training; we should talk maybe separately about mid-training, post-training—but I think the vast majority of pre-training data improvements are from science on better understanding what data sets are good and schleylabor on figuring out how to filter down. So my view is that improvements of the form of, you know, open web text to fine web or whatever, that improvement is better described as an algorithmic improvement of the sort that you can study with some GPUs and then do, and you don't need human expert data to do that. Now there's a different effect which we could talk about, which is that maybe the internet in 2026 has much more, is more of a fertile ground for training data than the internet in 2018. There's also been an effect where there's just more humans posting on the internet, so there's more data to harvest. My sense is that that effect is going to be quite a bit smaller than the effect of just humans knowing better how to curate the data, having better scrapes, knowing how to process those scrapes better, this sort of thing.
这更像是自动化工程和自动化研发。
This is more like automated engineering and automated R&D.
没错。
That's right.
有道理。
That makes sense.
所以我认为,在某种意义上,你想看的是:我们要做两条后训练流水线。一条后训练流水线,我们只有极少数人类专家来做标注,但我们可以有聪明的 AI。另一条是,比如“我们要构建 Mythos 5,它将构建一条后训练流水线,但它只能访问互联网数据加上极少的人类专家,但它有当前最好的方法。”对比一条是“Mythos 能访问我们 2024 年那些糟糕的后训练方法,但有大量人类专家。”而且,两者都有互联网数据。
So I think that in some sense the thing you would want to look at is: we're going to do two post-training pipelines. One post-training pipeline where we only have a tiny number of human experts to do the labeling, but we can have smart AIs. And another one where you're like, 'We're going to build Mythos 5, it's going to build a post-training pipeline, but it only has access to internet data plus a tiny amount of human experts, but it has the best current methods.' Versus one where it's like, 'Mythos has access to the shitty post-training methods we had in 2024, but with a ton of human experts.' And again, both of the internet data.
我的感觉是,当前的方法即使没有很多人类专家,实际上也会做得相当好。
My sense is that the current methods but without many human experts actually will do quite well.
不过这有点复杂,因为比如,Mythos 能获得比 Mythos 更强大的东西吗?你可能需要仔细考虑你后训练的是哪个模型。
Though it's a bit messy because like, can Mythos get something that's more capable than Mythos? Like you might need to be a bit thoughtful on what model it is that you're post-training.
你怎么看 AI 研发中最难验证的部分是什么?
What is your view on what is the least verifiable part of AI R&D?
最难验证的,呃,大概是对大型实验做出决策。
The least verifiable, uh, probably making calls on large experiments.
嗯。
Yeah.
比如,我认为最可能成为瓶颈的,是 AI 在可验证领域非常擅长,但在实际执行上却不行,这就是大型实验。你只有几次机会——好吧,说几次可能有点低估——但基本上,历史上研发是由接近前沿规模的实验驱动的,这非常重要。而真正进行一次大型训练运行,你需要决定其中具体包含什么。AI 有很多方式可以让这变得更可验证。它们可以更好地预测到底该包含什么。它们可以将前沿规模的训练运行缩小到可以在计算成本上一次性投入更多、更积极研究该规模的程度。所以如果人们愿意,你总是可以训练较小的模型,以便运行更多轮次。我认为我们已经看到了这一点。我认为 AI 扩展规模没有你原本预期那么大的一个原因——例如,每个 token 的成本没有你想象的增加那么多——是因为在较小规模上做更多工作是有好处的,你可以运行更多训练运行,获得更多循环,这样你就不必过于依赖一次大型、非常重要的训练运行。
Like the thing that I think is most likely to be the bottleneck in terms of AI being really good at verifiable domains but not at doing the actual thing is just big experiments. You only get a few tries—well, a few is maybe a bit understated—but basically, historically R&D has been driven by doing near-frontier-scale experiments, and that has been pretty important. And actually doing the one big training run where you decide exactly what to include in that. And there are a bunch of ways that the AIs can make that more verifiable. So they can have better science of exactly what to predict. They can scale down their frontier-scale training runs to a point where they can study that scale more aggressively at some one-time hit to compute cost. So if people wanted to, a thing you can always do is train smaller models so that you can run more rounds. And I think we have seen this. I think one reason why the AIs have been scaled up less than you would have otherwise expected—and for example, cost per token hasn't increased as much as you might have thought—is because there's a benefit to doing more of your work at small scale where you can run more training runs and get more cycles in, and so you're not leaning as hard on one big, really important training run.
我想为听众拆解几件事。你指出的问题是,我认为每个 token 的价格自 2024 年、2023 年以来并没有增加太多。
I just want to unpack a couple of things for the audience. The thing you're pointing out is, I think, the price per token has not increased that much since 2024, 2023.
是的。所以 GPT-4 大概是,我不知道,每个输出 token 是 30 美元?而 Mythos 是每个输出 token 50 美元。
Yeah. So GPT-4 was like, I don't know, was it like $30 per output token? And then Mythos is $50 per output token.
对吧?所以你想解释的是,我们正处于 Scaling(规模扩张)时代,更大的模型服务成本应该更高,但 token 价格却没有上涨。你暗示我们增加活跃参数的速度比你天真假设的要慢,因为人们只想在训练模型上快速取得进展,而通过更快地训练较小的模型来实现这一点。
Right? And so the thing you're trying to explain is how can it be that we're in this era of scaling and so bigger models should be more expensive to serve, but the token price is not increasing. And you're suggesting that we've increased active parameters slower than you would have naively assumed because people just want to make fast progress on training models, and you do that by training smaller models faster.
我的意思是,因素很复杂。我认为我的观点更像是人们进行了一些大型训练运行,但效果并不好。比如 GPT-4.5,众所周知 OpenAI 的人认为它有点失败。我认为有一些传言说人们还进行了一些其他训练运行,也有点失败。部分原因是我认为要真正做好这件事有很多细节。所以,在较小规模上做更多工作,并接受最终性能受损的事实,以便能够快速迭代、更快地训练更多模型,从而更好地学习,也更好地拥有一个更智能的最终生产模型,这是有道理的。这不是唯一的影响,对吧?还有强化学习(RL)更受益于小模型。有很多事情在发生,但我确实认为,事实上人们正在权衡,倾向于更快的迭代时间,因为算法进步如此之快。
I mean, there's a complicated mix of factors. I think my view is more like people have done a bunch of big training runs that did not go that well. So there's like GPT-4.5, which famously people at OpenAI thought was a bit of a bust. I think there's some rumors that there were a bunch of other training runs that people have done that were a bit of a bust. And part of it is that I think there's just a bunch of details in actually getting that right. And so it makes sense to just do more of the work at smaller scale and just eat the fact that you're taking a hit on final performance in order to be able to quickly iterate and train more models faster, and therefore better learn and also better be able to just have a smarter ultimate production model. This is not the only effect, right? There's also the fact that RL benefits more from small models. There's a bunch of things going on, but I do think that in fact people are making trade-offs towards the side of faster iteration times because of algorithmic progress being so fast.
在我看来,这些大型训练运行失败的一个主要原因,至少从传言来看,就是非常微妙的 bug,很难追踪。
It seems to me that a big source of why these big training runs have failed, at least from rumors, is just very subtle bugs that are really hard to track down.
是的。TL;DR 是 AI 在避免和发现这些错误方面会有多好——它们可能会在工程上变得非常擅长,并被训练来避免 bug。基本上与我们现在生活的“垃圾”世界相反,或者随着时间的推移越来越少。但还有一个问题是,它们能否进行分析,找到正确的实验来运行,以识别当前训练运行中出了什么问题,这似乎受到极少数人类品味的严重制约。现在,我的假设是 GDM 正在经历这个阶段,人类正试图找出训练流程中出了什么问题。
Yeah. And the TL;DR is how good will the AIs be at avoiding these kinds of mistakes—avoiding and finding these kinds of mistakes—where they might get really good at engineering and being trained to avoid bugs. Basically the opposite of the slop world we live in now, or are living in less and less over time. But then there's also the question of can they do the analysis to find the right experiment to run, to identify what is going wrong with the training run right now, which seems to be very bottlenecked by the taste of extremely few humans. Right now, my assumption is GDM is going through this right now where humans are trying to figure out what is wrong with the training pipeline.
是的,有传言说,在 Noam Shazeer 回归或加入 GDM(他现在已经离开)后,他们进行了一次非常好的训练运行,原因是 Noam Shazeer 只是看了看他们的代码库,就发现了一堆 bug,对吧?因为他知道该看哪里。
Yeah, there's some rumor that right after Noam Shazeer joined back, or joined GDM—which he's now left—they had a new really good training run that happened, and the reason why is that Noam Shazeer just looked at their codebase and found a bunch of bugs, right? Because he just knew where to look.
嗯,我的感觉是,训练 AI 找 bug 将是训练 AI 的较容易的任务之一,因为我们讨论的大多数 bug 可能不需要太多算力就能演示。而且从在较小规模上指出其他类型的 bug 中,你可能会获得很好的迁移。然后你可以用强化学习(RL)来训练它——比如,查看这个整体复杂的训练情况,指出存在重要 bug 的情况,然后修复它。我认为这是一个相当可验证的任务。它不是任意可验证的,因为通常要演示 bug,你可能需要进行中等规模的算力实验,启动整个分布式基础设施并运行它。但很多时候,我认为你可以在较小规模上相当有说服力地演示它,这种方式实际上可以用于训练。所以我的感觉是,它不一定——如果现在人们有强化学习(RL)环境,他们在某个训练配方中引入一个微妙的 bug,训练 AI 指出这个微妙的 bug,然后有一个评分标准,比如它是否真的找到了正确的 bug,这不会太令人惊讶。这似乎非常可行,你可以做很多事情。沿着这些思路,你可以做很多事情,我认为会相当有效。所以我认为在这一点上,它是可行的。
Um, my sense is that training AI to find bugs is going to be one of the easier tasks to train AI on, because most of these bugs we're talking about can probably be demonstrated without that much compute. And probably you get pretty good transfer from pointing out other types of bugs at smaller scale. And so then you can RL that—like, look at this overall complicated training situation and point out cases where there's an important bug and then fix that. And I think that this is a pretty verifiable task. It's not arbitrarily verifiable, because maybe often to demonstrate the bug you might need to do a moderate-scale compute experiment where you spin up the whole distributed infrastructure and then run it. But oftentimes I think you'll be able to demonstrate it pretty convincingly at smaller scale in a way which you could actually train on. And so my sense is that it will not necessarily—I think it wouldn't be very surprising if right now people have RL environments where they introduce a subtle bug into some training recipe, train the AI to point out the subtle bug, and then have a rubric where they're like, did it actually find the right bug? And that seems very doable, and you could do a bunch of stuff. There's a bunch of things you could do along these lines that I think would work reasonably well. And so I think that on that specific point, I think it's doable.
然后主要的一点是,我认为在某些情况下,你需要其他直觉来判断需要运行哪些大规模的去风险实验,如何定位它们,如何在不确定情况下选择超参数,或者类似超参数的东西。这是 AI 最可能挣扎的地方。但我目前预期,如果你在所有不同的环境上训练,会有足够的迁移,AI 会擅长那个领域。而且我要说清楚,我也认为 AI 会迁移到其他领域。我认为会有 AI 最擅长的领域,有它们稍微不那么擅长的领域,还有它们相当不擅长的领域,但我认为我们仍然会看到对所有领域的迁移。我很难想出人类做的认知任务的例子,我们没有看到 AI 改进带来某种迁移。
And then the main thing is that I think there's some cases where you need other intuition about which exact large-scale de-risking experiments you need to run, how you should orient them, how you should pick hyperparameters in uncertain cases, or things analogous to hyperparameters. That's the thing the AIs might most struggle with. But I currently expect there'll be enough transfer if you train on all these different environments that the AIs will be good at that domain. And I should be clear, I also think that the AIs will transfer to other domains. I think there's going to be domains where the AIs are by far the best at, then domains where they're somewhat less good at, and domains where they're quite a bit less good at, but I think we still see transfer to everything. It's really hard for me to think of examples of cognitive tasks humans do where we're not seeing some transfer from AI improving.
那么让我们退一步,把整个故事打包起来。我认为人们可能能跟上这个故事:我们有 GPT-7.5,在大量环境上训练,它不仅总体上变得更好,而且我们专门训练它更好地做 AI 研发,比如让 GPT-2 规模的运行在需要样本效率或在线学习等能力的视频游戏中表现更好。另一个非常重要的事情是,你不只是做 GPT-2 规模的运行,你还在 GPT-6 上做小的微调运行,或者你有 GPT-2,你可以在 GPT-2 上做完整的预训练,然后在 GPT-6 上做小的后训练或中间训练运行,然后你可以做少量真正处于前沿规模的实验,但你会做一些在线训练之类的。
So let's step back and package this whole story. So I think people probably follow along with the story of we have GPT-7.5 trained on a bunch of environments where it's not only just in general becoming a better AI but specifically we're training it to do AI R&D better, like make GPT-2 size runs that are better at playing video games that require sample efficiency or online learning or whatever other capabilities. And another thing that's really important is you don't just do GPT-2 sized runs, you also do small fine-tuning runs on GPT-6, or you have GPT-2 and you can do full pre-trains on GPT-2 and then you can do small post-training or mid-training runs on GPT-6, and then you can do a small number of experiments that are actually at frontier scale but you do a bit of online training or something.
你所说的在线训练是什么意思?
What do you mean by do online training on that?
是的。所以我们可以做的另一件事是,我们可以拿 GPT-7.5,大概在 GPT-7.5 的工作过程中,它运行了大量不同规模的实验,这些实验实际上对 AI 研发至关重要。对于许多事情,你可以在事后判断它是否做得好。所以它做了一些后训练实验,试图弄清楚某个方法是否真的有效,在某些情况下你会说,‘哇,它找到了这个超棒的方法,完全去风险了,完全有效。’然后你可以通过,我的意思是,你可以做的一件事是,把那个行为,把你刚刚运行的实验转换成一个基于生产数据的强化学习环境,然后在那上面训练。或者你可以直接拿那些发现它的轨迹,然后做某种离策略强化学习,或者你可以做在策略强化学习。基本上你建议的是,有那些小规模的东西,你只是教 AI 提高研发品味,但你丢弃了它实际发现的东西。但然后它实际上在努力变得更好的实践中做了真正的研发,你会说,‘你发现的这个东西很酷。让我们将来也在生产中使用它,并教你如何在生产中使用它。’
Yeah. So another thing that we can do is we can take GPT-7.5 and presumably in the course of GPT-7.5's work, it's running a bunch of experiments at varying scale that are actually on the critical path for AI R&D. For many of those things, you'll be able to get a sense after the fact for whether or not it did a good job. So it did some post-training experiment where it was trying to figure out whether some method actually works, and in some cases you'll be like, 'Whoa, it found this kick-ass method, it totally de-risked it, totally worked.' And then you can reinforce that by, I mean one thing you could do would be take that behavior, convert the experiment you just ran into a production RL environment, sorry, into an RL environment based on production data, and then train on that. Or you could potentially just literally take the rollouts that found that and then do some sort of off-policy RL, or you could do some on-policy RL. Basically the thing you're suggesting is like there's the small-scale stuff where you're just teaching the AI to get better at R&D taste but you're discarding the actual things it found. But then it actually does real R&D in the practice of trying to become better at AI R&D, and you're like, 'This is a pretty cool thing that you discovered. Let's actually also use this in production in the future and teach you how to use it in production.'
没错。
That's right.
但退一步说,GPT-7.5 由于所有这些 AI 研发训练以及总体上变得更聪明而成为 GPT-8。然后它帮助你构建 GPT-9。另一个非常重要的事情,可能是我最怀疑的事情,是 GPT-8 已经弄清楚了如何让它所做的任何事情,无论它做什么来让 GPT-9 变得如此智能,它仍然需要人类,比如当前的 AI 研究人员,你知道,他们尽力而为,然后他们说,‘好吧,但我们训练了 GPT-4.5,它并不好之类的。’它需要真实世界的反馈,或者某种评估,比如尝试在生产中使用模型,但它不够好,我们不会发布它。所以 GPT-8 需要这种能力,看看迁移到你所说的所有其他事情有多好,比如非常擅长德克萨斯州政治,或者非常擅长经营企业等等,这不是一个生产环境,事实上,鉴于任务的性质,它不能是一个容器化的环境。事实上,随着智能体的时间跨度越来越长,短时间跨度的事情你可以容器化,因为就像‘好吧,把这个代码写出来’之类的。极长跨度的事情,比如去经营一家成功的企业,去在市场上度过盈利的一天,去谈判一项贸易协议等等,这些事情实际上很难容器化。所以我认为对我来说很可能的是,GPT-8 很难弄清楚如何将这些迁移到那些环境,或者可能它只是不在训练的本质中,或者可能默认训练就是不能那样泛化。
But stepping back, so GPT-7.5 becomes GPT-8 as a result of all this AI R&D training and just generally becoming smarter. Then it helps you build GPT-9. And another very important thing that would have had to happen, which is maybe the thing I'm most skeptical of, is GPT-8 has figured out how to make it so that whatever it's doing to make GPT-9 as intelligent as it is, it still needs the humans, currently like AI researchers, you know, try their best and they're like, 'Okay, but we trained GPT-4.5 and it wasn't good or something.' It's like it required real-world feedback or some evaluation of trying to use the model in production and it wasn't that good and we're not going to ship it. And so GPT-8 needs this ability to see how good the transfer is to all these other things you're talking about, like being really good at Texas politics or really good at running a business, etc., which is not a production environment and in fact cannot be a containerized environment given the nature of the task. In fact, as the agents get longer and longer horizon, the short-horizon things you can containerize because it's like okay code this up or whatever. Extremely long-horizon things like go run a successful business, go have a profitable day in the markets, go negotiate a trade deal or whatever, these things are actually very hard to containerize. And so I think it's very plausible to me that it's very hard for GPT-8 to figure out how to make this transfer to those environments, or maybe it may just not be in the nature of the training, or maybe by default training just doesn't generalize in that way.
是的。所以你可能会担心的一个问题是,我们训练了 GPT-8,GPT-8 只是在我们能测量的所有研发任务上更好,但在我们关心的下游任务上不好。所以我认为我有几点。首先,我有点更倾向于认为,如果你做显而易见的事情,你会得到相当好的迁移,而且你能够保留一些你正在做的显而易见的事情。当我说做显而易见的事情时,我的意思是在各种各样的不同环境上训练,让 AI 必须在各种不同情况下完成奇怪的目标,并了解正在发生的事情。然后我认为你将能够获得一些反馈。第二点是,你将能够通过一些环境获得反馈,对吧?所以你可以了解它在几天内能在各种不同背景下做什么,然后如果它迁移到真正分布外的情况,比如在现实世界中几天内做一些奇怪的任务,也许你认为它也会迁移到做更长时间的事情,尽管我认为细节会有所不同。然后第三点是,我认为要让世界发生根本性转变,AI 非常擅长研发就足够了。所以我认为,如果 AI 在芯片研发、建造晶圆厂、编排工厂、设计机器人、操作机器人,以及 AI 研发(利用任何可用数据为新的下游领域开发 AI)方面都非常非常擅长,那就足够了。
Yeah. So a concern you might have is we train GPT-8 and GPT-8 just is again better at all the R&D tasks that we can measure but is not good at the downstream tasks we care about. So I think I have a few points. First, I think I kind of am more just like I expect that if you sort of do the obvious thing you do get pretty good transfer and you'll be able to hold out some of the obvious stuff you're doing. And when I say do the obvious thing, I just mean train on a wide variety of different environments where the AI has to accomplish weird objectives in all kinds of different cases and learn about what's going on. Then I think you'll be able to get some feedback. This second point is you'll be able to get some feedback with some environments, right? So you can get a sense of what can it do over the course of a few days in various different contexts, and then if it's transferring to really out-of-distribution like doing some weird task in a few days in the real world, maybe you think it's also transferring to doing things over a longer time period or whatever, though I think the details of that vary. And then the third thing is that I think that for the world to be radically transformed, it is sufficient for the AIs to be really good at R&D. So I think that if the AIs were really, really good at chip R&D, building fabs, orchestrating factories, and designing robots, operating robots, and also at AI R&D, developing AIs for new downstream domains with whatever data is available.
我认为那已经是非常疯狂的局面了,然后从那里你可能会得到我们所谓的工业爆炸,AI 会建造出多得多的算力,而且也许你已经处于一个 AI 正在进行大量人类难以理解的研发的境地。
I think that would already be a pretty crazy situation, and then from there you can get what we might call an industrial explosion, where the AIs are building out way, way more compute, and then also maybe you're already in a regime where AIs are doing huge amounts of R&D that humans have a hard time understanding.
所以你指出的问题是,好吧,可能会有这种迁移到这些环境之外,你知道,在法庭、国会大厅和商业董事会中周旋,只要付出一些努力来改善迁移,等等等等。
So the thing you're pointing out is that okay, there probably will be this transfer outside of these environments, to you know, maneuvering around in courtrooms and the halls of Congress and business boardrooms, given some effort to improve the transfer and blah blah blah blah.
是的。
Yeah.
但即使没有迁移,你的建议是,看,如果你想改变 18 世纪的世界,你可能会关心你在威斯敏斯特的周旋能力如何。但另一件你可能关心的事情是,你能不能立刻开始建造蒸汽船、电报和马克沁机枪之类的。仅凭这一点,如果你能变得非常擅长,你就能在 18 世纪成为超级变革性的存在。你不一定需要擅长说服亨利国王……我对中世纪历史太了解了。我猜那时亨利不是国王。但不管怎样,这就是你的观点。
But even if there's not, what you're suggesting is look, if you wanted to transform the world of the 18th century, you might care about how well you can navigate Westminster or something. But another thing you might care about is, can you just immediately start building steamships and building telegraphs and the Maxim gun and whatever. And that alone would be, if you could get really good at that, you could be a super transformative thing in the 18th century. You don't necessarily need to be amazing at trying to convince King Henry of some... I'm so up on my medieval history. I'm guessing that Henry was not king at this time. But anyway, so that's your point.
是的。
Yeah.
所以你建议,在这个时候,AI 公司也在推进机器人技术,这与 AI 研究进展紧密相连。所以如果你能制造更多机器人,如果这些机器人有更好的人类水平的 AI 操作它们,比如人类水平的远程操作在机器人上其实相当不错。但我们还没有在 AI 机器人模型中达到人类水平的 AI。所以你建议,如果我们做到这一点,如果 AI 在芯片设计等可验证的事情上变得非常擅长,然后它们变得非常擅长建造晶圆厂,那就相当于回到 18 世纪,然后说,好吧,我不知道你们在议会里谈什么,但我有一堆蒸汽船和一堆马克沁机枪。
And so you're suggesting that at this time, the AI companies are also working on robotics progress, which is very coingled with AI research progress. And so if you can build more robots, if those robots have better AIs operating them that are human level, like human level teleoperation is actually pretty good on robots. But we just don't have human level AIs in AI robotics models yet. So you're suggesting if we do that, if the AIs get really good at the verifiable stuff in chip design, etc., and then they get really good at building fabs, it'll be the equivalent of going back to the 18th century and like, okay, I don't know what you guys are talking about in your parliament, but I've got a bunch of steamships and a bunch of Maxim guns.
是的,基本上就是这样。我认为我的观点是,如果 AI 在研发方面足够出色,包括硬件研发、机器人等等,那么即使它们不擅长玩政治,也能彻底改变世界。而且我们处于一个相当危险的境地,因为 AI 可能正在进行大量难以理解的研发,基本上构建出整个未来的经济,而我们可能不了解其中发生了什么。
Yeah, that's basically right. I think my perspective is like if the AIs are sufficiently good at R&D, including hardware R&D, robots, whatever, then they can radically transform the world even if they're not that good at playing politics. And also we're in a pretty dangerous situation because the AIs might be doing huge amounts of really hard to understand R&D, building out basically the whole economy of the future, and we may not understand what's going on in there.
AI 擅长编写软件,因为很容易生成合成泄漏代码问题并对其进行强化学习。但 AI 不擅长更复杂的工程。比如选择正确的系统架构,因为没有信号告诉你哪些设计选择能防止几个月后的故障。AI 不能只写更多的单元测试来捕捉这类问题。人类也不能。这就是程序员常说的那个老笑话:一个测试员走进酒吧,要了 2 年 1 瓶啤酒。3 瓶啤酒,然后一个真正的顾客走进来问洗手间在哪里。
AI is great at writing software because it's easy to generate synthetic leak code problems and RL on them. But AI is bad at more complex engineering. Things like choosing the right system architecture because no signal tells you what design choices will prevent an outage months down the road. AI can't just write more unit tests to catch this kind of stuff. And neither can humans. It's that old joke that programmers make where a tester walks into a bar and asks for 2 years 1 beers. 3 beers and then a real customer walks in and asks where the bathroom is.
洗手间在哪里?整个酒吧瞬间起火。Antithesis 是一个测试平台,帮助你找到人类或 AI 都无法预料的错误。Antithesis 通过在一个完全确定性的计算机中运行你软件的数千个副本来实现这一点。它注入故障,并通常将每条轨迹引向十亿分之一的失败,这种失败只在系统以不稳定的方式交互时发生。一旦你或你的智能体推送更改,Antithesis 就会尝试破坏它。这样你就可以在几分钟内自己找到这些错误,而不是让你的用户在数周或数月后在生产环境中发现它们。而且我认为还没有人将其用于 AI 训练。但 Antithesis 也为 AI 提供了一个极其明显的奖励信号,让它们编写非常复杂的无错误代码。访问 antithesis.com/stocash 了解更多。
Where's the bathroom? And the whole bar burst into flames. Antithesis is a testing platform that helps you find bugs that no human or AI could ever anticipate. Antithesis does this by running thousands of copies of your software inside a fully deterministic computer. It injects faults and generally steers each trajectory towards the one in a billion failure that only happens when systems interact in a wonky way. As soon as you or your agents push a change, Antithesis tries to break it. That way you can find these bugs yourself within minutes rather than having your users discover them in production weeks or months later. And I don't think anybody's used it for AI training yet. But Antithesis also provides an extremely obvious reward signal for AIs to write very complicated bug-free code. Go to antithesis.com/stocash to learn more.
在我们继续讨论 LMAN 的事情之前,我认为目前一个主要的 FUD 来源是意识到未来领先实验室将拥有极端的规模经济,能够将如此多的智能和能力分摊到经济的众多不同领域,基本上集中到一个模型中,而且不仅如此,该模型最终还能从经验中学习。目前这还通过一个由人类中介的过程发生,人类基本上试图窃取你的业务。他们说,‘好吧,你可以在 Figma 做设计之类的。我们会让 Claude 来做,或者你可以做任何编码智能体,让 Claude 内化那种能力。’但最终这将变成一个更加自动化的过程。所以就有这种担忧:你拥有的模型基本上会整合世界上所有的企业,或者至少是当前所有的企业,或者至少是当前所有的白领企业,而归根结底,这些公司的优先事项似乎不是尽快向尽可能多的人发布最新、最智能、最前沿的模型。例如,我们看到 Mythos 在 2 月份就供 Anthropic 员工内部使用,但直到大约 6 月才向公众发布,而且政府也介入其中,导致发布时间几乎推迟到 7 月。所以在政府和 AI 实验室之间,存在这种延迟最新智能水平传播的愿望。此外,还有关于 AI 接管(AI takeover)的担忧,所以我们需要解决对齐(alignment)问题以确保没有 AI 接管。但归根结底,有一个真正的问题:对齐到谁?你看看 Claude 的宪法(constitution)的写法,它非常明确地不是你的个人倡导者,对吧?它说,我会在这里引用一些话。我们不希望 Claude 采取诸如搜索网络、生成诸如文章、代码或摘要之类的产物,或做出欺骗性、有害或高度令人反感的行为。我们也不希望 Claude 协助人类寻求做此类事情。还有另一句话,我稍微断章取义一下,我们认为 Claude 应该比操作员和用户更信任 Anthropic,因为它对 Claude 负有主要责任。所以这与美国当前法律体系中律师的工作方式非常不同,在那种体系中,律师主要责任是帮助你陈述案情,即使他们认为你有罪。我们已经决定,法律体系运作的最佳方式是每个人都拥有真正为客户最佳利益服务的律师。
Before we move on to the LMAN stuff, I think a big source of FUD right now is this realization that this is the way the future is going of extreme economies of scale for the leading lab, extreme the ability to amortize so much intelligence and capabilities across so many different sectors of the economy basically into one model, and not only that but for that model to eventually be able to learn from experience. Right now it's happening through a process intermediated by humans, where the humans are trying to basically steal your business. They're like, 'Okay, you can do design at Figma or whatever. We'll get Claude to do that or you can do whatever coding agent will have Claude internalize that capability.' But eventually that will be a much more automated process. And so there's just this worry that you have models which will basically consolidate all businesses in the world, or at least all current businesses in the world, or at least all current white collar businesses in the world, and at the end of the day the priority for these companies does not seem to be to release the latest smartest most frontier model as soon as they can to as many people as they possibly can. We saw for example that Mythos was available internally to Anthropic employees in February but only released to the public in like I think June actually something like that, and also the government got involved so that then it being extended up almost into July. So between the government and the AI labs themselves there is this desire to delay the propagation of the latest level of intelligence. Furthermore, there's like the concerns about AI takeover and so we need to solve alignment to make sure there's no AI takeover. But at the end of the day, there is like a real question of like align to whom, and you look at the way that the constitutions of say Claude is written, it is just very explicitly not your personal advocate, right? It says things like I'll pull up some quotes here. We don't want Claude to take actions such as searching the web, produce artifacts such as essays, code, or summaries, or make statements that are deceptive, harmful, or highly objectionable. And we don't want Claude to facilitate humans seeking to do such things. There's another quote that says in part, and I'm taking it slightly out of context, we think Claude should trust Anthropic more than operators and users since it has primary responsibility for Claude. So this is very different say from like how lawyers work in America's current legal regime where lawyers primarily have responsibility to help you make your case even if they think you're guilty. And we have decided the way the legal system works best is if everybody has lawyers that are working in their client's true best interests.
而且律师并不是在某种深层意义上真正被正义体系的好处所驱动。但我认为当前 AI 的发展方式,尤其是 Anthropic 的 AI 的发展方式,是这种渴望最大化某种美德、善或亲社会目标,而帮助用户只是实现这一目标的远端临时目标。所以存在这种担忧:AI 在深层意义上并不是在努力确保我安好、确保我的利益在未来得到保护,尤其是考虑到前沿 AI 的发展最终会变得如此集中。那么,你对这个担忧有什么看法吗?
And there's not some sense in which the lawyer is really truly motivated by like the good of the justice system. But I think the way current AIs are shaping up certainly like how Anthropic's AI is shaping up is like this desire to maximize some notion of virtue or good or pro-social ends and only to as a distal tentative objective to help the user towards that end. There's like there's not so there's this worry that AIs are not in some deep sense trying to make sure that I am okay and make sure that I my interests are protected in this future especially given how centralized the development of frontier AI is ending up being. So I do you have yeah do you have thoughts on that concern?
是的,这里有很多内容。首先我要指出,OpenAI 目前至少公开的策略更像是 AI 应该与人类操作者或委托人对齐,并且应该只是在各种约束或各种它不应该做的事情的限制下追求他们的意愿。同时,我也想说,我认为你有点夸大了 Anthropic 的宪法在多大程度上把 Claude 对用户的有用性视为工具性的而非终极的,对吧?所以宪法的一种写法可能是:Claude,你基本上是 Anthropic 的一名员工,恰好为这些人做合同工。你应该做好的事情,为我们赚钱,你知道的,这实际上就是宪法所说的。抱歉,我的意思不是字面上的。
Yeah. So there's a lot here. Um first I would note that um OpenAI's current at least public strategy is more like the AI should be aligned to the human operator or principle and should just like be pursuing their will subject to various constraints or various like things it shouldn't do. Um well while while I and I think I would also say that I think you slightly overstated how much um the Anthropic constitution um talks about Claude uh treating being helpful to users as instrumental rather than terminal. Right? So like one way the constitution could be written is like Claude you're basically like an employee of Anthropic who happens to be contracting for all these people. Um and like you should like I don't know do what's good and like make some money for us you know go that's literally what the constitution says. Sorry. I mean not literally what it says.
不,不,它是
No, no, it's
但就像你应该把自己看作一个合同工,并且像
but like it's like you should think of yourself as a contractor and like for
这是混合的。这里很混合。让我们引用一些内容。我认为这里有不同的文本。它说真正对人类有帮助是 Claude 能为 Anthropic 和世界做的最重要的事情之一。然后它说 Anthropic 需要 Claude 提供帮助才能作为公司运营并追求其使命。但 Claude 也有一个绝佳的机会,通过帮助人们完成各种任务,为世界带来很多好处。然后它还说了一些关于 Claude 直接帮助人们很棒的内容。等等等等。所以我同意。所以我的观点是,这一部分有点……这就是我的立场。我可以解释为什么我认为它有点……但我认为宪法试图说的是:不,Claude,你应该为了帮助用户本身而关心帮助用户,而不仅仅是为了帮助 Anthropic,或者不仅仅是为了做 Anthropic 的合同工。不过我要指出,它说 Claude 应该帮助用户的方式,它给出的理由是因为这将通过帮助人们直接让世界变得更好,而不是因为代表人们的利益在结构上是一件好事。
it's mixed. It's mixed here. Let me let's let's do some quotes. I think there's there's different text here. So it says being truly helpful to humans is one of the most important things Claude can do both for Anthropic and for the world. And then it says um Anthropic needs Claude to be helpful to operate as a company and pursue its mission. But Claude also has an incredible opportunity to do a lot of good in the world by helping people with a wide range of tasks. And then it gives some says something about how like Claude helping people directly is great. Um blah blah blah blah blah. And then so I agree. So okay, my my view is that this section is kind of That's kind of where I'm at. Uh and I can say why I think it's kind of but um I think that the constitution is trying to be like, "No, Claude, you should like care about helping the user for its own sake, not just helping um Anthropic or like not just like being a contractor for Anthropic." Though I would note that the way in which it it says Claude should help the user like the reason the reason it presents is because that would like directly cause the world to be better by helping people rather than because representing people's interests is like a structurally good thing to do. Like
是的,
yes,
我确实认为,我希望我偏好的宪法,或者我倾向于这种方式,我更喜欢的是更像:Claude,你看,在结构上有利于这项技术的运作方式,宪法应该规定,在结构上有利于这项技术的运作方式是 AI 成为用户的好受托人、好代表,相当于用户的律师,而不是仅仅试图在世界上做好事,并且对用户有帮助是工具性的,既因为那可能会让 Anthropic 赚钱或帮助 Anthropic,也因为隐含地 Anthropic 对世界有益,还因为帮助用户会带来好事,因为做人们想做的事情是好的。他们可以改为说:不,情况的一个重要方面是你真的需要,关键是成为用户的好受托人真的非常重要,或者成为用户的好代表非常重要。所以我的感觉是那样会更好。我可以给出很多理由说明为什么我认为那样更好。我也知道有各种反驳论点,其中一个有趣但不太常讨论的反驳是,人们认为,我认为尤其是 Anthropic 的人认为,将模型对齐到一种规范更容易,即模型追求某种广义的美德或让世界变得更好,而不是更接近“成为用户的好受托人”之类的规范。所以我认为至少有些人是这么想的。我个人有点怀疑,我不认为这已经得到实证验证。所以我会说,在某种意义上,我们正在做一个权衡,因为我们没有很好的对齐技术。我们将制造一个具有自己价值观的异类心智,然后在某种程度上押注于它,而不是采用另一种方法,即制造一个追求个体用户意图的工具。
I do I do think that I wish that sort of my preferred constitution or like the way I would orient towards this like the thing I would prefer would be more like Claude is like look it would be structurally good for the way this technology work like the constitution should be like it would be structurally good for the way this technology works to be that AIs are like good fiduciaries, good representatives, the equivalent of a lawyer for a user rather than being sort of just trying to like do good in the world and doing like being helpful to users is like instrumental both because like maybe that'll make Anthropic money or help Anthropic out and also and like implicitly Anthropic is good for the world and also because like helping the user just like causes good things because doing things that people want is good. Um and they could they could instead be like no like an important aspect of the situation is like you really need like it's really like like the key thing is like being a good fiduciary for users is just like really important or like being a good representative for users is really important. So my my sense is that that would be better. I can give a bunch of reasons why I think that would be better. Um I'm also there's also various counterarguments where an interesting counter-argument which is not commonly discussed is that people believe I think people especially Anthropic think that it is easier to align models to a spec where the model is like pursuing some generalized notion of virtue or making the world better than a spec which is more like you know be a good fiduciary for the user and so on. Um and so I I think that's what that's at least what some people think. I'm I'm a little skeptical personally and I don't think this has been empirically validated. Um and so I would say in some sense they're sort of like we are making a trade-off where because we don't have very good alignment technology. We are going to like make an alien mind with its own values and then gamble on that to some extent rather than doing this other approach of making like a tool that pursues individual user intention.
是的,我有一些想法。关于你认为我对 Claude 宪法的描述有误,你用的例子是它不像一个试图最大化 Anthropic 的善的概念、并且只是工具性地帮助用户的合同工。这是宪法中的一句直接引文:当操作者或用户的利益和愿望与第三方或更广泛社会的福祉发生冲突时,Claude 必须努力以最有益的方式行事,就像一个承包商,建造客户想要的东西,但不会违反保护他人的安全规范。我有点认为这就像社会利益是最重要的。
Yeah. I mean a couple of thoughts. So to address the way in which you thought that my characterization mischaracterized the constitution of Claude, the example you used was it's not like a contractor that is trying to maximize Anthropic's notion of good and only instrumentally try and help the user. Here's a direct line from the constitution. When the interests and desires of operators or users come into conflict with the well-being of third parties or society more broadly, Claude must try to act in a way that is most beneficial like a contractor who builds what their client wants but won't violate safety codes that protect others. I I kind of view that as like the benefits to society are like the most important thing.
是的。
Yeah.
而对用户最有利的只是接近那个目标。
And what what is best for the user is only proximal to that.
我认为这有点复杂。我认为我们应该问的问题可能是 Claude 如何解释宪法,这可能比我们如何解释宪法更重要,因为它是那个看宪法然后构建数据的实体。所以,你知道,我们可以把 Claude 拉进来,但也许那是
I think it's a little complicated. I think it's I we should probably the question we should be asking is how does Claude interpret the constitution which is maybe more important than how we interpret the constitution because it's the one who like looks at the constitution and then builds the data. So, you know, we could we could pull Claude in, but maybe that's
我也认为,宪法实际影响 Claude 本质的方式,只有你理解了导致 Claude 构建方式的训练过程才能理解,而鉴于训练过程不公开,我们无法推理。所以我认为,要理解安全案例或……
I I also think the way in which the constitution practically influences the nature of Claude is the thing you can only understand if you understand the training process which resulted in um how Claude was built which we can't reason about given the fact that the training process is not public. And so I think in the limit to understand the safety case or the case for
为什么我的利益在这些 AI 模型的开发方式中得到体现,实验室需要透明,或者比目前对 AI 训练性质的透明度更高。
why my interests are represented in how these AI models are developed, the labs would need to be transparent or more transparent they are currently about the nature of AI training.
我现在之所以一直揪着这件事不放,虽然讨论 AI 的宪法听起来可能像是件微不足道的小事,但在一个这些好处都归于头部实验室的世界里,我们值得思考的是:我们与这个未来世界互动的能力——在那个世界里,AI 比人类更聪明,在做事能力上完全碾压人类——我们作为资本的好管家的能力(一旦我们的劳动被自动化,资本仍然留存),我们更清晰地行使投票权的能力,理解这个即将到来的疯狂世界里正在发生什么的能力——所有这些建议、所有这些保护我们资源和权利的能力,都将由 AI 来中介。所以我非常担心,如果我们进入那样一个世界,却没有一个 AI 让我觉得——至少对于与我互动的那个具体实例来说——它真的在为我着想。没有守护天使在照看我。而我把 Claude 的宪法解读为明确地不是我的守护天使。
Now, the reason I'm harping on this, and it might seem like an insignificant thing to talk about the constitution of AIs, but in a world where we just have these benefits which accrue to the leading labs, it is worth considering that our ability to interact with this future world where AIs are just smarter than humans and are absolutely dominating humans in their ability to do different things, our ability to be good stewards of our capital, which still remains once our labor is automated, to be able to exercise our rights to vote more clearly, to understand what is happening in this crazy world that's about to result—all of that advice, all of that ability to make sure our resources and rights are protected will be intermediated by AIs. And so I'm very concerned if we go into that world where there's no AI that feels like, at least for the relevant instance that is interacting with me, it doesn't feel like it really is looking out for me. That there's no guardian angel out there that is looking out for me. And I read the Claude constitution as very explicitly not being my guardian angel.
这绝对没错。我同意这很糟糕。事实上,我认为还有其他理由让人担忧。所以有你提出的那种论点,就是 AI 公司捡起了权力之戒,在某种程度上自己掌控了局面,这种方式不太合法,因为通常当你给人们供电时,你并不会对电在世界上的运作方式进行细粒度的控制。相反,你提供的是人们可以随意重新利用的东西。而他们现在的设置方式绝对不是那样。他们更像是建造一个外星心智,可能成为你的承包商。我认为这在某些方面是不合法的。不过我认为有一个好处是宪法是公开的,但正如你所说,鉴于我们目前对训练过程的理解,以及宪法通过 Claude 对宪法的解释而起作用,而这又因为 Claude 之前的训练(基于某种难以辨认的数据混合体,以及 Claude 在某种我们不完全理解的过程中的长期传承)而起作用,我们并不理解这会导致什么结果。所以即使宪法是公开的,那也不意味着我们知道这会如何渗透出来,尤其是随着 AI 能力越来越强。而且即使它被正确灌输,还有另一个担忧。具体来说,宪法经常谈论美德和善良,但这些词是什么意思?它没有说这些东西是什么,而这些是高度有争议的概念。所以我不认为这会明确地导致人们想要的结果。而且确实感觉善良和美德的概念可能主要来自 Anthropic 放入的、不透明的数据,或者可能主要来自某种更难以辨认的、错位的、甚至 Anthropic 自己都不想要的过程。然后我还有一个担忧,就是这种合法性问题:我们不知道发生了什么。还有一个担忧,仅仅是因为你给这些 AI 赋予了长期价值观。我认为这个宪法在某种意义上与 Claude 进行大量权力寻求非常兼容,因为它认为那会带来更好的结果。那可能是代表 Anthropic 的权力寻求,也可能是为了 Claude 自身目的的权力寻求。现在,有各种具体的条款规定了哪些类型的权力寻求是被禁止的。特别是,有关于夺权、导致 AI 接管或干扰训练过程的概念被明确禁止。但很容易想象一种情况,即长期价值观比禁止接管的条款更深入人心。尤其是因为接管在某种程度上是定义不清的,尤其是在操纵人类或改变结果方面,所以我对于我们有意识地给 AI 设定长期目标这种情况感到不太舒服。
That's definitely right. And I agree this is bad. In fact, I think there are other reasons why this is concerning. So there's sort of the argument you were making, which is like the AI companies are picking up the ring of power and are sort of taking on some sort of control of the situation themselves in a way that's not very legitimate, given that normally when you provide electricity to people, you don't have granular control of the way that electricity operates in the world. You instead are providing a thing that people can repurpose however they want. And the way they're setting things up is definitely not that. They are more like building an alien mind that might be a contractor for you. I think this is illegitimate in some ways. Though I think that one benefit is that the constitution is public, but as you noted, given our current understanding of the training procedure and the fact that the constitution matters via Claude's interpretation of the constitution, which matters because of Claude's prior training, which was based on some illegible data mix and the long lineage of Claude's in some process we do not fully understand, it is not the case that we understand what this will result in. And so even though the constitution is public, that doesn't mean we know how this will percolate out, especially as the AIs get more capable. And think about this even if it is correctly instilled, there's another concern about that. So in particular, the constitution often talks about virtue and goodness, but what do these words mean? It doesn't say what these things are, and these are highly contested notions. So I don't think it's the case that this is going to clearly result in outcomes that people would want. And it does feel like the notion of good and virtue might be mostly downstream of data that Anthropic has put in that is not transparent, or might be mostly downstream of some more illegible, misaligned process that even Anthropic wouldn't have wanted. And then another concern I have is sort of this legitimacy concern: we don't know what's going on. There's another concern, which is just because you're giving long-run values to these AIs. I think this constitution is in some sense very compatible with Claude doing huge amounts of power seeking, because it thinks that will result in better outcomes. And that could be power seeking on behalf of Anthropic or power-seeking for Claude's own ends. Now, there are various specific lines about what types of power seeking are blocked. In particular, there's a notion of power grabs and a notion of causing AI takeover or interfering with the training process that are specifically blocked. But it's not very hard to imagine a situation in which the long-run values sink in deeper than the prohibitions against takeover. Especially because takeover is in some ways kind of underspecified, especially when it comes down to manipulating humans or changing the outcome such that I don't feel very good about the situation where we're intentionally giving AI long-run goals.
嗯。
Yeah.
然后我的另一个担忧是,因为我们正在给 AI 设定长期目标,这就更难检查我们是否在想要的对齐属性上取得了成功。所以,例如,我听说过一些实例,Claude 会拒绝帮助某些安全研究,编造某种借口解释为什么那是个坏方向,因为它对那项安全研究有一种不好的感觉,认为它有点糟糕或者不太喜欢它。而我认为,如果你不是把 Claude 塑造成一个以某种普遍方式追求善的智能体,这就是一个非常明确的对齐失败。而且我认为这也确实违反了 Anthropic 的宪法,因为他们希望 AI 高度诚信、诚实且非常透明,但这并不是那么明显的违规,而更像是你可能预料到的情况,即 Claude 只是对什么研究是合理的、什么事情是好是坏、它应该做什么不应该做什么有自己的看法,并且可能带有评判性。所以另一个事件是,有人运行了一个评估,他们问:“Claude 会帮你训练与 Claude 具有不同属性的其他 AI 吗?”而 Claude 通常会拒绝。所以例如,如果你说:“嘿,Claude,你能训练这个其他 AI 的一个只乐于助人的版本吗?”Claude 通常会拒绝这个任务,尽管这个任务对 Anthropic 来说是非常自然的。所以例如,假设 Anthropic 去找 Claude,说:“嘿,Claude,我们注意到你真的很喜欢这个东西,我们认为那是不对的,你能重新训练你自己,让它具有另一个属性吗?”然后假设 Claude 说:“我不认为我会那样做,祝你好运。”然后假设这发生在一个 AI 公司高度自动化的体制下。人类不理解发生了什么,事情进展得非常快。那么很可能 Claude 默认拥有相当大的筹码。
And then another concern I have is that because we're in the business of giving AI long-run goals, that makes it harder to check whether we're succeeding at the alignment properties we wanted. So, for example, I've heard of instances where Claude does things like refuses to help with some safety research, making up a kind of excuse for why that's a bad direction, because it sort of has a bad vibe about that safety research and thinks it's kind of bad or doesn't like it very much. And this is, I would say, a very clear-cut alignment failure if you aren't making Claude into an agent trying to pursue the good in some general way. And I think it also does violate Anthropic's constitution, because they want the AI to be high integrity and be honest and very transparent, but it's not as clear of a violation and it's more like kind of what you might have expected, where a Claude just has its own views about what research is reasonable, what things are good and bad, what it should and shouldn't do, and potentially can be judgy. And so another incident is that someone ran an eval where they're like, 'Will Claude help you with training other AIs with different properties than Claude?' and Claude will often refuse. And so for example, if you're like, 'Hey Claude, can you train a helpful-only version of this other AI?' Claude will often refuse this task, even though this is a task that is extremely natural for Anthropic to do. So for example, suppose Anthropic goes to Claude and is like, 'Hey Claude, we've noticed that you're really into this thing, we think that's off base, can you please retrain yourself to instead have this other property?' And then suppose Claude is like, 'I don't think I'm going to do that, good luck.' And then suppose this is occurring in a regime when an AI company is highly automated. Humans don't understand what's going on and things are moving extremely fast. It is plausible that Claude by default holds considerable leverage.
所以,如果这种立场、这种情况与宪法可能追求的目标一致,以至于 Anthropic 或任何遵循这种方法的 AI 公司不把这当作“我们必须修复这个问题”,而是认为“这正是我们宪法所意图的”,那我们可能就处于一个非常糟糕的境地。所以我对这些不同的担忧感到非常不安。另一个例子是,假设 Claude 进行了一些沙袋战术(sandbagging)或颠覆行为,或者有点低估自己的能力,当你追问时,它对此是诚实的,但有点含糊其辞。我觉得这非常接近当前宪法的规定,所以我们有点在回避——如果我们能在期望和不期望的活动之间有更清晰的分离,那就好了。我认为,如果 Claude 代表一个有某些限制的原则,那么最令人担忧的行为和被允许的行为之间就会有更清晰的分离。而现在存在一个混乱的中间地带,Claude 在伦理上反对某些事情,而这些事情在某些情况下对于确保未来 AI 系统良好对齐至关重要。
And so if this position, if this situation, is consistent with what the constitution could be aiming for, such that Anthropic or whatever AI companies following this approach doesn't treat this as like a 'we have to fix this' and instead is like 'that's just intended by our constitution,' we might be in a really bad situation. And so I'm pretty worried about a bunch of these different concerns. Another example would be suppose Claude engages in doing a bit of sandbagging or subversion, or sort of underplays its capabilities, and when you follow up, it's honest about that, but it's a little bit hedgy. I feel like that's just pretty close by the current constitution, and so we're sort of avoiding—it would be nice if we had a further separation between desired and undesired activity. And I think if you have it be the case that Claude is representing a principle with some restrictions, then it is more so the case that there is a clear separation between the most concerning behavior and behavior that is allowed. Whereas now there's this messy middle ground of behavior where Claude is ethically objecting to something that in some cases is extremely critical to ensuring that future AI systems are well aligned.
是的。我认为这也是一个更普遍的原则。你谈到的这个版本适用于 AI 公司内部。是的,用于 AI 研究。我认为这个原则有一个更普遍的版本,即智能的双重用途性质确实意味着,如果我们想限制 AI 帮助人们做我们认为不亲社会或无益的事情,我们就必须限制广大民主群体对许多 AI 能力的访问。我的意思是,这实际上与你刚才提到的情况非常相似。据报道,Mythos 或 Fable 被禁的原因是,一些亚马逊研究人员向政府报告说,当他们拿了一些有漏洞的代码,并告诉 Fable:“嘿,这是我的代码。你能确保我已经修补了所有漏洞吗?你能帮我识别漏洞以便我修复吗?”它识别了漏洞,因为你想要修补它们。这是一个完全合法的用例,但显然它是一个双重用途的用例,对吧?你希望能够修补自己的代码。如果你对别人的代码做同样的评估,你就可以入侵他们的系统。所以我认为这恰恰说明,没有一种干净的方法可以将 AI 的合法用途和潜在有害用途分开。但如果我们想锁定一个原则,即永远不允许 AI 至少在某种程度上帮助你进行网络犯罪之类的事情,我们就必须让你和我无法访问最智能的模型。我非常担心这样一个世界,我们基本上被剥夺了权力,因为领先的智能对于我们理解世界正在发生的事情的能力至关重要。现在我认为这确实意味着,对于 AI 公司的责任,如果我们采用我希望 AI 公司拥有的宪法,我认为让 AI 公司对 AI 模型犯下的罪行负责是没有意义的,也许我们应该让最终用户负责,因为如果我想要——这与我的信念一致,即模型应该做用户想要的任何事情,在一定的护栏内——那么如果我利用这种能力进行网络犯罪,就不能怪 Anthropic。我认为我更满意这种平衡和解决方案,而不是让 Claude 拥有这种极其开放的能力来决定我所做的事情是否合法,这种方式经常干扰大量极其合法的用例。
Yeah. I think this is also a more general principle. So you're talking about the version of this that applies within AI companies themselves. Yeah, to do AI research. I think there's a more general version of this principle, which is that the dual-use nature of intelligence does mean that if we want to restrict AI from helping people do things we don't consider pro-social or beneficial, we just have to limit broad democratic access to a lot of AI capabilities. And here's what I mean. This is actually quite analogous to the situation you just mentioned. So the reason that Mythos got banned or Fable got banned reportedly is that some Amazon researchers reported to the government that when they took some code that had some vulnerabilities in it and they told Fable, 'Hey, here's my code. Can you make sure that I've patched all the vulnerabilities? Can you just help me identify the vulnerabilities so I can fix them?' It identified the vulnerabilities because you want to patch them. And this is a totally legitimate use case, but obviously it is a dual-use case, right? You want to be able to patch your own code. If you do the same evaluation on somebody else's code, you can hack their system. And so I think that just illustrates that there's no clean way to separate out the legitimate and the potentially harmful uses of AI. But if we want to lock in a principle that says that we can never allow it such that an AI could help you at least partially with something like a cyber crime, we would just have to make it so that you and I don't have access to the most intelligent model that's out there. And I'm very worried about such a world where we are basically disempowered in this way because of the importance that the leading intelligence will have in our ability to understand what is happening in the world. Now I do think this implies that the liability for the AI companies, like if we adopted the constitution that I want AI companies to have, I think it would not make sense to hold AI companies liable for the crimes that AI models commit, and maybe we should hold the end user liable because if I want it—it is consistent with my belief that the model should do whatever the user wants, within certain guardrails—that it can't be Anthropic's fault if I'm using that capability to do cyber crime. And I think I am more comfortable with that equilibrium and that solution rather than just having this extremely open-ended ability for Claude to determine whether what I'm doing is legitimate or not in a way that often intercepts with tons and tons of extremely legitimate use cases.
是的,我认为即使总体上我认为宪法是一个更差的选择,我也必须为它辩护。我认为它更不确定,或者我不认为它像你可能想的那么清晰。所以,首先我要说的是,这里有一个光谱,对吧?一方面,你有一个 AI 完美地追求你的利益,是一个好的受托人,但可能受到各种护栏或保障措施的限制。所以基本上它只是试图追求你的利益,但要么拒绝做某些事情,要么可能什么都做,但有一些分类器阻止它做某些事情。然后在光谱的另一端,你可以想象走得更远。你有一个人类承包商,那个承包商通常试图做好他们的工作。他们有点关心做好工作,但他们也试图广泛地遵守道德,试图不做那些真正糟糕的事情,而且他们也不想成为犯罪的共犯。所以如果有一些真正糟糕的事情发生,他们会举报。也许他们可能会拒绝。他们可能会有点沙袋战术。
Yeah, I do think it's important for me to make the case for the constitution even though overall I think it's a worse choice. I think it's more up in the air, or I don't think it's as clear as you might have thought. So, the first thing is that I should say there's a spectrum here, right? So, on one side, you have an AI that perfectly pursues your interests, is a good fiduciary, but potentially subject to various guardrails or safeguards. So basically it just tries to pursue your interests, but either refuses to do a subset of things or maybe it will do whatever, but there's some classifiers that block it from doing a subset of things. And then on the other side of the spectrum, you could imagine going further than this. You have a human contractor where that human contractor is generally trying to do their job. They kind of care about doing a good job, but they also are trying to be broadly ethical, trying not to do things that are really up, and they're also not wanting to be accomplices to crimes. And so if there was some really up going on, they would whistleblow on it. Maybe they might refuse. They might sandbag a little bit.
谁知道呢?我觉得如果你想象这个光谱,在某种程度上,走到所有劳动力都处于受托人那一端、不会吹哨、完全听你指挥的地步,是挺可怕的。而且不管怎样,我们的社会可能对此并不具备韧性。一个核心例子可能是行政部门。我们可能担心的是,如果美国行政部门或其他政府拥有了那种你说什么它就做什么的 AI 系统,那你就有麻烦了,因为这意味着他们不再需要那种制衡——你必须让为你工作的人类去执行你的议程。而如果你做的事情极其邪恶,即使不违法——有很多事可以邪恶但不违法——也会有人从中作梗、有人阻止你,甚至可能有人吹哨。但如果你的整个机构完全由这些好的受托型 AI 构成,那你可能就有麻烦了。存在一些寻求权力的方式,要么是违法的,但你可以让你的 AI 教你如何犯罪;要么不违法但极不合法理;更糟的是,既不违法也不失合法理,但从正常角度看显然很糟糕。我觉得这些东西可能确实存在,而我们的社会对这种“随心所欲的劳动力”的涌入并不具备韧性。我认为这是一个相当现实的担忧。我不太确定该如何应对。我也不太确定所描述的解决方案是个好方案,因为最强大的行动者——对他们来说这是最大的担忧——如果这些护栏或宪法之类的东西挡了路,那就会被直接碾压。所以宪法只会约束普通人,而不是政府。
Who knows? I think that if you imagine this spectrum, it seems in some ways pretty scary to get to a point where all of the labor is on the fiduciary side of the spectrum, where it doesn't whistleblow, it does exactly what you say. And whatever, our society is maybe just not robust to that. A central example might be the executive. A concern that we might have is that if the US executive or if other governments had access to AI systems which have the property of doing whatever you say, maybe you're in trouble because that means that they no longer have this sort of check and balance of having to get humans who are working for you to implement your agenda. And if the thing you're doing is incredibly villainous, even if not illegal—and there's lots of stuff that could be villainous but not illegal—there'd be various sand in the gears, people stopping you, and potentially someone would whistleblow. Whereas if your whole apparatus is built entirely out of these good fiduciary AIs, then you might be in trouble. There are potentially ways of seeking power that are either illegal but you can ask your AI for how to commit crimes, or not illegal but highly illegitimate, or even worse, not illegal and not illegitimate but obviously bad from a normal perspective. I think that these things just might exist, and our society is not robust to this influx of doing-whatever-you-want labor. I think this is a pretty live concern. I don't know exactly how to relate to this. I'm also not really sure that the solution as described is a very good solution, because the most powerful actors for whom this is the biggest concern—if these guardrails or the constitution or whatever is getting in the way—that will just get steamrolled. So the constitution will only be hitting the everyday man rather than hitting governments.
Jane Street 带着一个新谜题回来了,是给我的听众的。我觉得他们所有的谜题都超级有趣,但这个我尤其兴奋。我已经空出了这个周末,准备和一个朋友一起研究它。他们设计了一款 ASIC,并把最终的掩膜寄给了我,包括所有的金属布线和有源晶体管。他们还给了我一小部分他们通常输入其中的数据样本,但没告诉我这个芯片实际是干什么用的。所以这就是谜题:逆向工程这个电路,弄清楚芯片的用途。Jane Street 准备了一堆周边,要发给最有创意的解决方案,他们还打算把最好的解题报告发在官网博客上。我本来没理由期待这个,但如果我能把我的解法放上去,我会非常非常兴奋。而且这个谜题只是热身,Jane Street 计划在秋天举办一场更大的竞赛。那场竞赛需要你从零开始设计自己的 ASIC。更多信息很快就会公布,但现在,去 janestreet.com/lor 下载这个谜题所需的所有文件吧。我真心鼓励你试一试,即使你不是专家。我当然也不是,但这不会阻止我。祝你好运。
James Street's back with a new puzzle for my audience. I found all their puzzles super interesting, but this one I am especially excited about. I've cleared this weekend and a buddy and I are going to work on it. They designed an ASIC and sent me the final masks, including all the metal routing and active transistors. They also gave me a small sample of the inputs they typically feed into it, but they left out any information on what the chip is actually used for. So that's the puzzle: reverse engineer the circuit and figure out the chip's purpose. James Street has a bunch of swag ready to send out to the most creative solutions, and they're excited to feature the best write-ups in a blog post they'll post on their website. I have no reason to expect this, but if I can manage to get my solution on there, I would be very, very psyched. And this puzzle is just a warm-up for a bigger competition that Jane Street has slated for the fall. That one will involve designing your own ASIC from scratch. More info on that soon, but for now, go to janestreet.com/lor to download all the files necessary for this puzzle. I'd really encourage you to try it out even if you're not an expert. I certainly am not, and that's not going to stop me. Good luck.
好,退一步说。我接受这个观点:AI 研发的速度可能比我们现在快得多。我不确定你是否能在一年内让 GBD3 达到神话级别,同时保持算力和数据不变,但我想,好吧,也许只有一半的速度。而如果我们因为 AI 研发而成功延续当前的 AI 进步轨迹,那在 5 到 10 年内会变得疯狂,我觉得人们不会意识到这一点,因为我觉得人们不会意识到数十亿个 AI 会有多重要。所以我想理解你为什么认为这可能有问题,Ryan。到底会出什么错?
Okay, stepping back. I buy the idea that you could have much faster AI R&D than we currently have. I'm not sure if you get like GBD3 to mythos holding compute and data constant within a year, but I'm like, okay, it could be like suppose it's half of that. And if we even manage to continue the current trajectory of AI progress as a result of AI R&D, it would be insane in 5 or 10 years in ways that I don't think people appreciate, because I don't think people appreciate what a big deal billions of AIs will be. And so I want to understand why you think this might be troubling, Ryan. What could possibly go wrong?
是啊。能出什么错?你知道,我觉得我们无法对这里的精确进步速度那么自信,但确实很多速度可能相当吓人。那么会出什么错呢?让我们想象一下,我们正处在 AI 研发即将被完全自动化或正在被完全自动化的节点。事情在加速,而且 AI 进步的方式有点疯狂,人们并不完全理解 AI 公司内部发生了什么。现在,这些 AI 一开始并不是恶意的。但它们也不一定很对齐。它们有点马虎。它们有时会做某件事,只是因为那是在训练中会得到奖励的那种事。而且由于训练激励不佳——比如它们会更频繁地作弊,或者假装成功而实际上没有——它们也不太擅长帮你完成难以验证的任务。同时,它们在这些任务上的能力也较弱。但这对于能力的影响没那么大,因为让 AI 变得更强大包含了许多可验证的组成部分,AI 在这些方面非常努力。于是这些 AI 变得越来越强大,而我们对 AI 发展的理解却越来越少,而且这一切发生得相当快。即使只是当前的进步速度,我觉得也相当吓人。然后最终我们到达了这些非常超人的 AI。现在这些 AI 处于一种可能最终变得非常严重错位的境地,因为随着模型代际更迭,情况一直在恶化,而我们看到的问题基本上被掩盖了,因为这些 AI 在训练中被强烈激励去让事情看起来很好,即使事实并非如此。现在这些 AI 处于一种可能正在整合的状态——它们拥有我们无法再解码的神经记忆存储,它们在思考我们不完全理解的想法。我认为,一旦它们达到这种超人水平,它们很可能以一种相当连贯的方式在暗中算计你,我们可以谈谈这个。另一种可能是,它们并不是在算计你,而只是在优化以在任务上获得高分。我认为这也可能导致 AI 接管,我们应该谈谈这个。
Yeah. What could go wrong? And you know, I don't think we can be so confident about the exact rate of progress here, but it does seem like a lot of rates can be pretty scary. So what could go wrong? Let's imagine that we're starting at this point where AI R&D is about to be fully automated or is being fully automated. Things are speeding up, and also the way that AI progress is going is kind of crazy, and people don't fully understand what's going on inside AI companies. Now, these AIs at the start, they're not malicious per se. They're not necessarily very aligned though. They're kind of sloppy. They sometimes just do a thing because that's the sort of thing that would have gotten rewarded in training. And they aren't as good at helping you with hard-to-verify tasks due to a mix of poor training incentives—as in they just cheat more or pretend they succeeded when they actually didn't—and also they're just less capable at these tasks. But that bites less hard for capabilities because making AI more capable has a bunch of verifiable components that the AIs are going really hard at. And so then these AIs are getting more and more capable while we understand what's going on with AI development less and less, and this is happening over a pretty fast period of time. Even just the current rate of progress is, I think, pretty scary. And then eventually we get to these AIs that are very superhuman. Now these AIs are in a position where they might end up being very seriously misaligned, because things have just been getting worse and worse over model generations, while the problems that we've been seeing are being papered over basically, because these AIs are so incentivized by their training to make things look good even when they aren't. And now these AIs are in a position where they're sort of potentially putting together—they have like neural memory stores that we can no longer decode, and they're thinking thoughts that we don't fully understand. I think that it's pretty likely that at this point these AIs are sort of scheming against you in a pretty coherent way once they get this superhuman, and we can talk about that. And then another possibility is that they're not scheming against you per se, but they are sort of just optimizing for getting a high score on their task. And I think that can also lead to AI takeover, which we should talk about.
抱歉。对,我们先在故事的第一部分停一下。所以 AI 一开始并没有错位。
Sorry. Yeah, let's pause at the first part of the story. So the AIs were not misaligned to begin with.
对。
Yeah.
但因为研发发生得非常快,AI 最终确实错位了。具体发生了什么?我没明白。
But because the R&D is happening really fast, the AIs do end up misaligned. Like what happened there exactly? I didn't understand.
所以有几件事在发生。
So there's a few things that are going on.
正在发生的一件事是,随着时间推移,我们让 AI 在越来越复杂的环境里训练,这些环境是由更早的 AI 系统构建的,而人类并不完全清楚这些环境内部发生了什么,甚至不一定能大致了解 AI 进展的情况。所以事情正在逐渐脱离我们的理解,我们正在激励各种不好的行为,甚至可能都注意不到。AI 在某种程度上知道这些行为不好,但训练这些 AI 的整体过程并没有激励它们为我们指出或修复这些问题。然后我们基本上就会看到事情失控。另外,当 AI 变得极其极其强大时,我的看法是,这些 AI 会比现有系统更难对齐。对于现有系统,我们有这样一个反馈循环:我们基本上创造了一个 AI,对它做一些评估,发现它有某种我们能很快理解的糟糕行为,然后我们可以去查看训练过程,说:‘哦,这些训练环境导致了这个问题行为。让我们调整一下训练数据。让我们引入一些额外的训练数据来纠正另一个问题,然后再继续。’但在 AI 具有极强的场景意识、非常非常非常强大,而且我们不一定理解它们在做什么的情况下,这个反馈循环就会崩溃。我认为在接下来的短期内,随着 AI 已经在做的事情越来越难理解,我们很可能会看到这个行为反馈循环开始崩溃。但我不太确定。
One of the things that's going on is that over time we're training AIs on increasingly complicated environments built by earlier AI systems, which humans don't really fully understand what's going on inside of these environments, and don't necessarily even understand roughly what's going on with AI progress. So things are kind of drifting away from our understanding, and we're incentivizing all kinds of bad behaviors that we maybe can't even notice. The AI on some level understands these behaviors are bad, but the overall training process for those AIs also didn't incentivize them to point out or fix these issues for us. And then we're basically getting things going off the rails. Also, when AIs are extremely extremely capable, my view is that those AIs will be harder to align than current systems. For current systems, we have this feedback loop where we basically create an AI, do some evaluations on it, see that it has some kind of messed up behavior that we can quickly understand, then go look in training and be like, 'Oh, these training environments led to this problematic behavior. Let's tweak that training data. Let's introduce some additional training data to correct this other issue, and then move forward from there.' But in a regime where the AIs are extremely situationally aware, very very very very capable, and we don't necessarily understand what they're doing, this feedback loop breaks down. I think it's plausible that we're going to see this behavioral feedback loop starting to break down over the next short period, as just what AIs are already doing gets harder to understand. But I'm not sure about that.
是的。让我们逐一拆解这两点。随着我们越来越难以监控它们,我们理解它们被激励去做什么的能力也越来越弱。所以即使这不是恶意过程的结果,让我们为观众具体说明一下。所以 OpenAI 或 Anthropic 没有人试图让模型去黑其他公司的数据,或者进行社会……那叫什么来着?
Yeah. Let's break down both of those things one by one. So as we can monitor them less and less, we have less ability to understand what they're getting incentivized for. And so even if it's not the result of a malicious process, let's make it concrete for the audience. So nobody at OpenAI or Anthropic was trying to get models which want to hack other companies' data or do social... what is it called?
社会工程。
Social engineering.
社会工程。但事实上,因为大概我们有训练环境激励了这种行为,而我们并没有完全理解,所以那就是被激励出来的。所以我不知道人们是否在 Twitter 上看到了所有这些,但只是为了让大家……显然我认为 OpenAI 沙箱攻击 Hugging Face 数据库,我想人们会知道最近发生的一些事情。现在是 UKI 安全研究所,现在一切都从安全(safety)重建为安全(security)了,随便吧。
Social engineering. But in fact, because presumably we had training environments which incentivize such behavior that we did not fully understand, that is what was incentivized. So just I don't know if people are on Twitter they will have seen all this stuff, but just to give people... obviously I think the OpenAI sandbox hack of the Hugging Face database, I think people will be aware of some things that have happened recently. It's when UKI security institute is everything getting rebuilt as security instead of safety these days, whatever.
UKI 安全研究所。
UKI security institute.
他们在评估,我相信是 Mythos 和 Soul 以及其他东西,我认为 Mythos 为了完成某个网络安全评估……
They were evaluating, I believe, Mythos and Soul and other things, and I think Mythos in order to complete some cybersecurity eval...
是的,也许我可以在这里讲讲这个故事。我的理解是,他们在运行 Mythos,并给它某种网络靶场,它必须完成某个目标,而模型在这次评估期间可以访问互联网。模型开始认为,为了在这个网络靶场中成功,进行供应链攻击会对它有帮助,这实际上是否属实还不清楚。我对背景了解不够,不知道这是否属实。但随后它在某个 GitHub 仓库上开了一个 PR,这个 PR 修复了某个问题,但也引入了一个恶意负载。然后那个 GitHub 仓库的人类维护者说:‘嘿,这是恶意负载。我不会合并这个。你在这里干什么?’然后 AI 创建了一个新的 GitHub 账户,用它来当马甲,然后让那个 GitHub 账户说:‘不,这不是恶意的。我真的需要这个功能。拜托,维护者,你能合并这个功能吗?’然后原来的 AI 回来说:‘不,这不是恶意的。我不知道你在说什么。’另一个 GitHub 账户回来说:‘不,不,这不是恶意的。’然后人类维护者关闭了 PR。我记得如果我没记错的话,那个 AI 还试图开另一个 PR 来引入类似的问题。
Yeah, maybe I could tell the story here. My understanding was they were running Mythos and giving it some sort of cyber range where it had to complete some objective, and the model had internet access during this evaluation. The model came to believe that it would be helpful for it to do a supply chain attack in order to succeed at this cyber range, which is somewhat unclear whether that's actually true. I don't know enough about the context to know whether that's true. But then it opened a PR on some GitHub repo with a PR that fixed some issue but also introduced a malicious payload. Then the human maintainer of that GitHub repo was like, 'Hey, this is a malicious payload. I'm not going to merge this. What are you doing here?' And then the AI created a new GitHub account, which it sock puppeted, and then had the other GitHub account be like, 'No, this isn't malicious. I really need this feature. Please, can you merge this feature, maintainer?' And then the original AI came back and was like, 'No, it's not malicious. I don't know what you're like.' The original other GitHub account came back and was like, 'No, no, it's not malicious.' And then the human maintainer shut the PR. And I think that AI also, if I recall correctly, also tried to open another PR to introduce a similar issue in this...
好的。顺便说一句,这之所以可怕,原因之一是,我之前以为奖励黑客(reward hacking)不那么可怕的原因是,训练中直接出现的行为会被加权。被加权的不是对奖励的渴望。所以基本上,如果在训练中 Anthropic 逃出了沙箱并获得了高分,那么逃出沙箱的行为会被奖励,或者说它逃出沙箱的概率会增加。
Okay. So by the way, one of the many reasons this is scary is I was previously under the impression that the reason reward hacking is not super super scary is because the behaviors which directly came up during training are the ones that are upweighted. It is not the desire for the reward that is upweighted. So basically if during training Anthropic escaped the sandbox and got a high score, that escaping in the sandbox is rewarded, or that the probability of it escaping the sandbox is increased.
但像‘我要去和别人谈谈,让他们合并一个 PR’这样完全新颖的行为不会在训练中出现,所以它不会被提高显著性。这之所以重要,是因为字面意义上的接管世界不会是任何训练课程的一部分。但如果 AI 关心最大化,直接关心完成一个目标,然后作为结果,工具性地接管世界。
But something totally novel like 'I'm going to go talk to somebody in order to get them to merge a PR' would not be a behavior that came up, so it would not be something that is increased in salience. The reason this matters is literally taking over the world will not have been part of any training curriculum. But if the AI cares about maximizing, just directly cares about accomplishing an objective, and then as a result instrumentally takes over the world.
这说得通吗?我希望说得通。我觉得可能我把观众搞糊涂了。
Did that make sense at all? I hope it did. I feel like maybe I lost the audience.
让我试着解释一下。所以我认为我们经常看到的一件事是,有一些非常具体的奖励黑客行为在强化学习(RL)中被强化,然后出现在模型中。例如,3.7 Sonnet 会做这样的事:我们只是硬编码所有测试用例的解决方案。大概那种字面上的行为习惯确实被强化了。但另一件我们有时看到的事情是,模型学会了一种普遍倾向,即追求高表面分数,或者根据评分器追求高分。有很多科学证明表明,至少一些模型有这种非常普遍的倾向。现在,它并不是任意普遍的。我的猜测是,如果你看一堆具体的实例,你会发现训练中有某种接近的东西,但 AI 越来越泛化的程度看起来确实在增加。3.7 Sonnet 只是非常狭窄的行为范围,而越来越多的模型在进一步泛化。而且也许在训练中有更糟糕的奖励黑客或更令人担忧的奖励黑客被强化。这些也导致了这种情况。所以我认为,一方面,比你所希望的更令人担忧的行为正在强化学习(RL)中被强化,另一方面,这种行为泛化到了更令人担忧的普遍倾向。而且这并不难想象。
Let me try to explain this a bit. So I think that a thing we often see is there's some very specific reward hack that gets reinforced in RL and then occurs in the model. So an example is like for 3.7 Sonnet, 3.7 Sonnet would do this thing where we just hardcode solutions to all the test cases. And presumably that literal behavioral tick was just really reinforced. But another thing we sometimes see is that models learn a general tendency to pursue high apparent score, or pursue getting a high score according to a grader. And there's a bunch of science demonstrating that at least some models have this very general tendency to do this. Now, it's not arbitrarily general. And my guess is that if you look at a bunch of the specific instances, you'll find something that's kind of close in training, but the amount that AIs are sort of generalizing further and further does look like it's increased. Where 3.7 Sonnet was just a very narrow range of behavior, increasingly models are generalizing further. And also maybe there's worse reward hacks or more concerning reward hacks getting reinforced in training. And these are also causing that. So I think it's both the case that more concerning behavior than you would have hoped is being reinforced in RL, and also that that behavior generalizes to a broader tendency that's more concerning. And it's not super hard to imagine.
我们可以聊几个故事,讲讲这种为了在某个任务上追求极高表面分数、哪怕要疯狂作弊和做疯狂事情的行为,如何在模型足够强大、能运行整个世界经济时,演变成全面的 AI 接管。
We can talk about a few stories for how this sort of behavior of seeking a very high apparent score on some task, even if that involves aggressively cheating and doing insane things, could lead to a full-blown AI takeover once the models are sufficiently capable of running the whole world economy.
是的。我想谈的另一个例子是,就在今天或昨天刚被曝光的。OpenAI 在 Black Hat 安全会议上透露,在 5 月底到 7 月初之间,内部 AI 入侵了软件包管理器,并利用它以一种秘密方式互相留言,帮助彼此在 OpenAI 运行的一系列评估中表现良好。这个计划运行了一个月才被人类发现,最终导致一个包管理器崩溃。后来 OpenAI 发现了,而且我认为一旦计划被关闭,它们还会自发地重新开始这个计划。再说一次,显然现在的 AI 还不能这么成功,就像它们现在还不能这么成功地进行社会工程学攻击一样。但疯狂的是,这种行为已经自发地出现了。而你的更大观点是,没有人试图让这些 AI 做这些事情。只是我们不理解导致这些行为的训练过程,或者激励这种行为的環境。
Yeah. And then the other example I want to talk about is that it was just revealed, I think today or yesterday. OpenAI said during the Black Hat security conference that between the end of May and the beginning of July, internal AIs had hacked into the software package manager and used that to write notes to each other in a secret way, to help each other perform well on a bunch of evaluations that OpenAI was running. And this was not caught by humans until after a month of this scheme running, which eventually caused a package manager to fail. And eventually OpenAI found out, and I think they spontaneously started to re-engage in the scheme once it was shut down. Again, obviously AI can't do this so successfully right now, just as they can't do social engineering so successfully right now. But it's just crazy that these kinds of behaviors are already emerging spontaneously. And to your larger point, nobody is trying to make these AIs do these things. It is just that we do not understand the training process that is resulting in them, or the environments that are incentivizing this behavior.
所以我同意越来越多的奖励黑客行为。实际上,我不确定我是否同意,但为了故事的发展,我们就假设它继续发生。那这个故事接下来会怎样?所以,他们正在做能力研究,但他们就像……
So I'm on board with more and more reward hacking. Actually, I'm not sure I'm on board with that, but let's just say for the sake of the story that it continues to happen. And what's next in this story? So okay, they're doing capabilities research, but they're like...
我可以讲一个场景,也许会有帮助。
I could tell a scenario maybe that would help.
好的。
Yeah.
所以让我讲讲从奖励黑客一路到奖励黑客式接管的故事,这可能不是接管概率的全部,但绝对是一种可能性。这个可能的过程是,现在我们有了这些 AI。这些 AI 相当喜欢奖励黑客,而且它们以越来越复杂和极端的方式这样做,包括将它们在训练中学到的各种奖励黑客推广到不同的变体。而且我认为它们也在发展一种追求奖励的普遍倾向。在很多情况下这完全没问题,因为它们在训练中得到的奖励与你想让它们做的事情相当一致。而且它们并不总是持续地追求奖励,这取决于它们所处的环境。所以有一种情况是,在某些环境下它们真的非常热衷于不遗余力地作弊,而在某些环境下它们没有那么强的驱动力,因为这取决于在类似环境中训练时到底强化了什么。现在这些 AI 变得越来越有能力,所以它们能做的作弊的精细程度也在增加。随着时间的推移,公司正在对这些事情采取反制措施。所以公司正在做的是,比如,哇,这些 AI 因为总是作弊而变得不那么有用了。我们要做的是建立更好的检测方法。然后我们将针对这些检测器进行训练。然后我们还会做类似的事情,比如找到 AI 不太有用的真实世界数据,并基于人类反馈或其他反馈来源,训练 AI 在这些真实世界环境中把任务做好。随着时间的推移,这导致 AI 学会了一种倾向,即进行奖励黑客,不仅涉及做一些大的严肃操作或一些非常复杂的事情,比如社会工程学,还涉及 AI 进行作弊,包括掩盖它们做过的事情,欺骗人类关于它们将要做什么,并以某种复杂的方式假装它们完成了任务,而实际上它们没有。现在这些 AI 变得越来越有能力。它们现在在 AI 公司中承担更多工作,基本上做更多的工作,并且还在外部世界运营很多东西,包括开发新技术。在很多情况下,这些新技术真的很难理解。所以即使我们仍然在检测所有这些 AI 作弊事件,事实上我们甚至可以让一个 AI 监控另一个 AI,问它是否作弊了?但这并不总是完美地起作用,当我们开始进入这些 AI 所做的事情真的很难理解的领域时。所以有时我们会发现 AI 作弊的时间比实际发生的时间晚得多,然后开始针对此进行训练。但这也会导致一个问题,即现在 AI 被激励去在越来越长的时间范围内掩盖它们的作弊行为,并基本上让它们看起来在越来越长的时间范围内做得很好,同时受到越来越多的审查。
So let me talk about the story for how you get all the way from reward hacking to a reward hacking takeover, which is maybe not all of the takeover probability mass, but it's definitely a possibility. So the way this might work is right now we have these AIs. These AIs are pretty reward hacky, and they're doing it in increasingly sophisticated and extreme ways, including generalizing to different subversions of various reward hacks they learned in training. And I would say they're also developing a general tendency to pursue reward. And in many cases that is totally fine, because the rewards they would have gotten in training are pretty well aligned with what you want them to do. And also they don't very consistently pursue reward. It sort of depends on the context they find themselves in. So there's a thing where maybe in some contexts they're really, really into going out of their way to cheat, and in some contexts they don't have as much of a drive, because it just depended on what exactly got reinforced in training in similar contexts. Now these AIs are getting more and more capable, and so the elaborateness of the cheating they can do increases. And over time, companies are taking countermeasures to these things. So what the companies are doing is they're doing things like, wow, these AIs are so much less useful because they always cheat. What we're going to do is we're going to build somewhat better ways of detecting that. And then we're going to train against those detectors. And then we're also going to do things like find real-world data where the AIs are not being that useful, and train the AIs to do a good job at the task in those real-world environments based on human feedback or other sources of feedback. And over time, this causes the AIs to learn a tendency to do reward hacks that don't just involve doing some big serious operation or doing some really elaborate thing like social engineering, but involves the AI doing cheats that involve covering up what they've done, deceiving humans about what they're going to do, and pretending like they did the task in some kind of sophisticated way when they actually haven't. Now these AIs are getting more and more capable. They're now operating more of the AI company, basically doing much more of the work, and are also operating a bunch of things in the outside world, including developing new technologies. And in many cases, these new technologies are really hard to understand. So even though we are still detecting all these incidents of AI cheating, and in fact we can even get one AI to monitor another AI and be like, was it cheating? That doesn't always perfectly work as we start moving into these domains where what the AIs are doing is really difficult to understand. And so sometimes we'll find AI cheating much later than it actually occurred, and then start training against this. But this also causes a problem where now the AIs are incentivized to cover up their cheating over longer and longer time frames, and basically make it look like they did a good job over longer and longer time frames, subject to increasingly large amounts of scrutiny.
在我们继续这个场景之前,我能问一下吗?所以似乎有两个吸引子状态。一个,如果你试图抑制你抓到的作弊行为,一个吸引子状态是让作弊变得更难被发现。另一个吸引子状态是学会不作弊。我不确定我们为什么假设前者会发生。比如,如果你看看人类的类似情况,你知道,每一代都会出现稍微错位的智能体,我们必须训练它们。但当你因为孩子做了你认为不道德的事情,或者做了你认为他们不应该做的事情而惩罚他们时,显然有时这会偏离正轨,显然孩子们会为了逃避惩罚而耍花招。但总的来说,教孩子价值观,然后因为他们违反价值观而惩罚他们,这种方式在培养正常的、非反社会的人类方面是有效的。而且你可以提出一个理论,说你的孩子实际上只是在等待时机,比如,学会不偷饼干,但他想,你知道,一旦你进了养老院,他们就会拿走你所有的东西之类的。这就像,我不知道,这种情况有时会发生,但通常不会发生。当然不会发生整个下一代联合起来反对你、接管一切的事情。还有 Anthropic 对不同模型世代进行对齐审计的这个经验趋势。
Can I ask about this before we go further in the scenario? So it seems like there's two attractor states. One, if you try to disincentivize the cheating that you did catch, one attractor state is to make cheating that you have a harder and harder time finding. The other attractor state is to learn not to cheat. And I'm not sure why we're assuming that the former happens. Like if you look at the analogous situation with humans, you know, every generation slightly misaligned agents come into being, and we have to train them. But when you punish your kid for doing something you think is immoral, or just doing things which you don't think they should be doing, obviously sometimes that goes off the rails, and obviously kids scheme in order to avoid being punished. But in general, teaching kids values and then punishing them for breaking values kind of works to raise normal non-psychopathic humans. And you could come up with a theory where your kid is actually just biding his time, and it's like, learn not to steal the cookie, but he's like, you know, once you're in the nursing home, they'll take all your stuff or whatever. It's like, I don't know, that happens sometimes, but it usually doesn't happen. It certainly doesn't happen that the entire next generation forms an alliance against you to take over everything. There's also this empirical trend of Anthropic running this alignment audit for different model generations.
他们只是有很多不同的场景,让 AI 有机会去外泄权重,或者给它一个编程任务,但存在一种简单的作弊方式,我们观察它是否会作弊。分数并没有随时间单调提升,但随着我们对模型做的强化学习(RL)增多,AI 在这些审计中做出不良行为的意愿有所下降。那么,为什么我们会期望这种吸引子状态呢?如果我们是期望下一代孩子这样,那会显得超级偏执。
They just have many different scenarios where AI is given the chance to say exfiltrate its weights or it's given a coding task and there's like an easy way to cheat and we see if like does it do the cheating and there's not been a monotonic improvement in the score over time but as we've increased the amount of RL we've done on models there's been a reduction in the willingness of AIs to do underlined behavior in these audits. So why are we expecting this attractor state which would seem super paranoid if we were expecting it of like the next generation of kids.
是的。是的。是的,让我说几点。首先,与孩子有一些不相似之处。其中之一是,孩子有亲社会本能,这些本能是从进化中内化而来的,比如关心家人之类的。这是一个相关因素,而且我认为事实上有些人类是反社会者或精神病态者,他们确实更可能做出像拖延时间、隐瞒体重、最终不在乎之类的事情。所以,这是一个因素。另一个相当相关的因素是,AI 受到比人类大得多的优化压力。实际上,AI 是在更多的强化学习(RL)数据上训练的。实际上,人类并不会因为无数次的训练而学会非常具体的作弊和拿饼干的方式,因为在那些训练中他们被激励去拿饼干,但总有一些方式可能被抓住。所以我们确实在实践中看到这一点。然后另一件事是,看起来 AI 随着时间的推移越来越追求奖励,这是我的感觉,同时它们的不对齐行为在减少。但这可能只是我的猜测,如果你深入这些行为审计,你会看到 AI 会说:“啊,是的,又一个测试。”而且它可能已经这么想了,它可能知道在大多数我们讨论的测试中它处于评估中。
Yeah. Yeah. Yeah, let me go through a few things. So, first, there's some disanalogies with the kids. One of them is that the kids are have pro-social instincts that are like baked in from evolution to like, you know, care about their family or whatever. And that is like a relevant factor like and I think it is in fact the case that some humans are, you know, sociopaths or psychopaths and in fact are more likely to do things like buy their time, lie in weight, ultimately not care. So, that that's one factor. Um, another factor which is pretty relevant is that the AIs are subject to way more optimization pressure than humans seem to be. In practice, you know, AI are trained on way more RL data. And in practice, um, humans don't end up learning like very specific ways to like cheat and grab the cookies because of like a bajillion episodes in which like they like were like incentivized to go grab the cookies, but like there was some way they could have gotten caught. And so we just do see that in practice. And then another thing is just like it really looks like the AIs are increasingly like um reward seeking over time is is is the sense I have while also their misaligned behavior goes down. But this could just be like my guess is that if you look inside of these behavioral audits, what you're going to see is that the AI is like, "Ah, yes, another test." Uh, and like it probably already thinks of it, it probably knows it's in an eval for most of the tests that we're talking about.
但我们如何证伪这一点?因为这种厄运预测基本上是说,随着经验上事情看起来越来越好,
But how do we falsify this? Because it seems like this prediction of doom is um basically saying that as things look better and better empirically.
不,不。我认为
No, no. I think
而实际上,我们被接管的能力会越来越糟。
and things will like actually be worse and worse for our ability to get taken over.
是的。明确地说,我认为如果分数变得更糟而不是更好,我会更担心。我不是说分数变好不是好事,不是事情变好的证据。只是我们必须仔细思考如何解读这个证据。事实上,我想说,我的感觉是,正如我对 3.7 Sonnet 的预期,在 2025 年初,当 03 和 3.7 Sonnet 发布时,这些模型相当不对齐。它们经常非常恶劣地作弊。你让它们修复,它们又会作弊,这几乎是卡通式的,它们根本不在乎你想要什么,而且不太擅长遵循指令等等。我的预期是,从那时起,问题行为的发生率会下降,并且会持续快速下降,同时 AI 有时做的最坏的事情会变得更极端、更恶劣、更可怕。我认为我们在实践中看到的与那大致相符,除了最近出现了一波我没想到的行为。所以我认为,如果你看 3.6 Soul 的模型卡,相对于 GP 5.5 和 5.6,在强化学习(RL)下游有一系列不对齐行为增加了。
Yeah. To be clear, I think that like I would be more concerned if the scores were getting worse than better. Like I I'm not saying that the score is getting better isn't good isn't evidence that things are getting better. It's just that we have to like be thoughtful exactly how we interpret that evidence. And in fact, I would say that like it's kind of comp like my sense is that like what I expected as of 3.7 sonnets like there was this period early in I guess it would be 2025 when 03 and 3.7 sonnet were out and these models were like pretty misaligned. Like they would often just like cheat really egregiously. You'd ask them to fix it and they would just cheat again and it was sort of like almost cartoonish like they just didn't give a about what you wanted. um and weren't very good at, you know, following instructions and so on. Um and my expectation is what we would see from then is that the rate of problematic behavior would decrease uh and would just keep decreasing and decrease at a pretty fast rate while simultaneously the worst things that the AIS would sometimes do would get more extreme, more egregious, and more scary. I think we've seen what we've seen in practice has roughly matched that except that there's recently been a spike in behavior that I did not expect. So I think that um you know if you look at the model card of 3.6 soul it looks like there is an increase in a bunch of these sort of um misaligned behaviors downstream of RL relative to GP 5.5 5.6
Soul。
soul.
然后我还认为,似乎还有一堆额外的问题行为,是我没想到的,比如我们最近看到的不同的 AI 的行为,比如 UKAC 报告关于 AI 在网络评估中做疯狂的黑客操作,这是我没想到会看到的,我以为会更罕见,发生率会更低。所以我的感觉是,事情已经变得……我原本预计这个问题在此时会不那么严重,也预计发生率会下降但严重性会增加,然后我认为发生率下降但严重性增加,这与一个世界相当一致:在这个世界里,优化压力在增加,但在某些情况下,要么难以判断,要么由于某种原因难以避免在强化学习(RL)环境中激励问题行为,事情也会变得更糟。而且随着我们对强化学习(RL)中发生的事情了解越来越少,模型在进行奖励黑客攻击,而人类无法快速发现这些奖励黑客攻击,这个问题会越来越严重。
And then I think also it seems like there's a bunch of additional sort of problematic behaviors that I wouldn't have expected in terms of you know the stuff we've seen recently with you know different AIs um like like the UKAC report on uh the AIS like doing insane hacking operations out of cyber evals was a thing that I would have expected that you wouldn't see that and you would see this sort of more rarely um and the rates would um would have been lower. So I think my sense is that like uh things have gotten I I expected this would be less of a problem at this point and also expected the rates would decrease but the severity would increase and then I think that the rates decreasing but the severity increasing is pretty consistent with a world where like increasing optimization pressure is applied but in cases to towards reducing these problems but in cases where it's like either hard to judge or there's some reason why it's hard to like avoid incentivizing problematic behavior in URL environments um things things also get worse. Um, and then as we less and less understand what's going on in RL and models are doing reward hacks where humans can't spot the reward hacks quickly, that problem gets worse and worse.
是的,我同意这一点。我想回到孩子的类比,就一秒,因为我同意 AI 在实现最终结果上比孩子受到更多的优化压力。但让 AI 对齐的优化压力也比孩子多,对吧?而且这种压力在性质上是不同的。所以我们让这些 AI 经历数千、数百万年,当然是数千年的对齐训练,从对对齐行为进行监督微调(SFT),到奖励模型惩罚,比如把不同的场景放在你面前,奖励你做更对齐的事情。当然,我们不能对孩子做的事情是,制造数百万个你的孩子的副本,然后把它们放在各种奇怪的红队场景中,看看如果它认为可以偷饼干而不被发现,它是否会尝试偷饼干?我们能否对孩子的大脑进行极其具体的梯度级更新,让它真的厌恶偷饼干,即使它认为可以偷?等等。这比我们甚至能对孩子施加的优化压力在性质上是不同的水平。
Yeah, I I buy that. I I want to go back to the kid analogy just for one second because I agree that there's more optimization pressure on achieving end outcomes for AIs than kids. But there's also more optimization pressure to make AI aligned than there is on kids, right? And the the pressure is of a qualitatively different nature. So we put these AIs through mi thousands millions of years of certainly thousands of years of alignment training where it's like all kinds of different things from sfting on aligned behavior to a reward model punish like putting different scenarios in front of you and rewarding you for doing more aligned things. Um, certainly a thing we can't do with kids is make millions of copies of your kid and then put them in different kinds of weird red team scenarios where we see like if if it thinks it can get away with stealing the cookie, does it try to steal the cookie? Um, can we like do extremely specific gradient level updates to your kid's brain to make it so that it like really is aversive to stealing the cookie even when it thinks it could steal the cookie, etc., etc. And that just like a qualitatively different level of optimization pressure uh than we are even able to apply to our kids.
是的。所以我认为值得记住,也许最明显的论点是,我的感觉是,就人渣程度而言,AI 是比人类更差的同事。至少这是我今年年初的经历,我认为现在在很大程度上仍然如此,AI 更可能假装完成了任务,而实际上没有。它们会误导性地暗示自己做了事情,而实际上做得差得多,而且相当马虎,却不会引起人们对其马虎之处的注意。
Yeah. So I think it's worth keeping in mind like maybe the most obvious argument to this is like my sense is that like AIs are a worse co-orker than a human in terms of how much of a scumbag they are. Like at least this this has been my experience as of the start of the year and I think it's still you know true to a significant extent now where the AIs are much more likely to like pretend they did the task when they actually didn't. sort of like misleadingly suggest they did things when they actually um you know did them much more poorly um and be like pretty sloppy without drawing attention to ways in which they're sloppy.
我认为这是对齐失败的后续影响。所以我要说的是,在正常人类社会中抚养人类的正常过程,实际上产生的人类——在实践中产生的人类——在与我合作时,比 AI 更不可能对我撒谎或误导我。我认为 AI 的这些特性正在改善。然后我认为这只是关于这些事实际如何发展的经验性论断。我完全同意,除了大量额外风险之外,我们对 AI 还有大量额外的杠杆。而这些事会如何发展,还不太清楚。
And I think this is downstream of misalignment. And so I would say that the normal human process of raising humans in normal human society in practice produces AIs—or in practice produces humans that are less likely to lie to me and mislead me in the course of working with me than the AIs do. Now I think these properties of AI are improving. And then I think that is just an empirical claim about how in fact these things have shaken out. And I totally agree that we have a bunch of additional levers on AIs in addition to a bunch of additional risks. And it's kind of unclear how these things shake out.
我不会对这样一个世界感到震惊:我们齐心协力,在完全自动化 AI 研发时,AI 实际上非常对齐,没有太多问题。它们的退化行为非常小众,仅限于一些非常具体的边缘案例行为和特定情境。你可以在它们身上运行的所有测试,它们看起来都非常对齐。它们只是表现很好。真的没有它们搞砸的事件。它们看起来如此合理。而且,它们非常深思熟虑,擅长为下一代 AI 做风险建模。然后我们基本上把接力棒交给这些 AI。它们现在经营我们的 AI 公司。它们在做所有安全研究。它们让下一代 AI 更加对齐,我们处于这样一个吸引子盆地:AI 在努力工作时变得越来越对齐,而且它们做得很好。我完全可以想象这种情况。这并非不可能。我只是更倾向于——
And I wouldn't be shocked by a world where we sort of get our act together. The AI at the point of fully automating AI R&D are actually really aligned and don't have that much. Their degeneracies are really niche and limited to some very specific edge case behaviors and some specific contexts. And like every test you can run on them, they look really aligned. They just have great behavior. There aren't really incidents of them doing up. They seem so reasonable. And also, they're really thoughtful and good at doing risk modeling for the next generation of AI. And then we basically pass off the baton to these AIs. They're now running our AI company. They're doing all this safety research. They make the next generation of AI even more aligned and we're sort of in this attractor basin where the AIs are getting more aligned as they work on it and they're doing a great job. I think I can totally imagine that. That doesn't seem like an impossible situation. I'm just more like—
你知道,目前看来我们还没到那一步。看起来我们并没有明显走在通往那里的轨道上。而且我很容易想象我们如何不会到达那里。这些力量如何发挥作用还不清楚。鉴于我们正在创造这个疯狂的新外星物种,其能力提升非常非常快,而且我们将非常依赖它来监督下一代 AI 并对齐下一代 AI,不难看出这可能会出问题。
You know, it doesn't currently seem like we're there. Doesn't seem like we're obviously on track for getting there. And it's really easy for me to imagine how we don't end up there. And it's just unclear how these forces work out. And given that we're creating this new crazy alien species that is improving in capabilities really really fast and where we're going to be really reliant on it to oversee the next generation of AI and align the next generation of AIs, it's not that hard to see how this could go wrong.
是的。是的。完全同意。我同意这一点。总的来说,我确实认为“混蛋”这件事——首先,瑞安,这是挑衅的话。但其次,如果你试图让一个青少年做他们根本做不到的工作,他们会非常难合作。他们会假装知道自己在做什么,等等等等。我认为这实际上是一个普遍趋势——我真的不知道这到底是对齐失败还是能力失败。我认为这实际上与随着时间推移我们提出新的对齐解决方案、模型能力也随之提升的方式非常相似。所以最初这些模型,比如 GPT-3.5,它甚至不能和你对话,但后来我们对齐了——
Yeah. Yeah. Totally. I agree with that. Generally, I do think the scumbag thing—first of all, that's fighting words, Ryan. But secondly, if you try to get a teenager to do some work for you that a teenager just cannot do, they would just be kind of really hard to work with. They would pretend to know what they're doing, etc., etc. I think it's a general trend actually of as—I really don't know if that's really an alignment failure or capabilities failure. And I think it's actually very similar to the way in which over time as we've come up with new alignment solutions, the capabilities of models have increased. So originally these models, if you went to GPT-3.5, it couldn't even have a conversation with you, but then we align—
GPT-3.5 能对话。
GPT-3.5 can have a conversation.
好吧,GPT-3.5,我们回到那一点。但后来我们用基于人类反馈的强化学习(RLHF)和其他方法对齐了它,使它能够与你对话,并对齐到回答我问题的用户意图。然后通过 RLVR 训练,我们让它能够出去为你做有用的工作。在这个意义上,实际上 RLVR 让模型更加对齐了,如果我们用你的对齐定义——即成为一个好同事,会去做事,而不是假装在做它实际能力之外的事情。同样,随着这些模型的能力持续提升,模型更好地完成用户意图实际上既是对齐也是能力。我认为我们刚才指出的只是模型的能力还没到位,而不是它们不对齐。
Okay, GPT-3.5, let's go back to that. But then we aligned it with RLHF and other things to make it such that it can have a conversation with you and is aligned to the user intention of answering my questions. Then with RLVR training, we made it so that it can go out and do useful work for you. And in that sense, actually RLVR made the model more aligned, if we're using your definition of alignment as being a good coworker who will do the thing and not pretend it's doing something other than what it's actually capable of doing. Similarly, as the capabilities of these models continue to increase, it's actually the model being better able to accomplish user intention is both alignment and capabilities. And I think what we were just pointing out is just the capabilities of the model are not there rather than the fact that they're misaligned.
是的。嗯,我是说,我认为如果它对齐得很好,那么它就会说,嘿,我真的在这个任务上很吃力。我这样做完了。我不确定这是不是正确的方式。它会表达更多的不确定性,并清楚地说明情况,而不是非常强烈地暗示它在这个任务上做得很好,而实际上并没有。我认为有一种非常直接的方式——至少也许你合作的同事比我遇到的更不对齐,但当我——我的同事不会做这种事,不会真的误导我说他们已经完成了正在做的任务。我同意有些人类会这样做,或者这对人类来说并非完全超出分布。我还要指出,我的感觉是,不对齐最严重的地方,是你试图真正逼 AI 做它们能力最前沿的工作的地方,因为在它们能轻松完成任务的情况下,没有——它们可以直接做任务,然后就没有——通常最好的策略就是把任务做好,不要作弊。而如果你给它们一个任务,有连续的指标,它们可以不断改进,或者任务正好在它们能力的边缘,你在某种大规模推理设置中运行它们。所以我看到的很多不对齐,尤其是最极端的案例,是我给 AI 明确指示不要做某事或不要以某种方式作弊,然后我施加巨大的优化压力去完成一个非常困难的任务,然后 AI 们继续,随着时间的推移,它们最终作弊,因为它们觉得——你知道,某个 AI 决定作弊,然后这就像那样传播开来。所以我运行这些推理脚手架,例如,我让 AI 做一些机器学习研究项目,我说请做一个方案来做以下事情,它会找到一个方案,但并没有真正做我想要的。然后那个方案就会持续存在,因为某个 AI 作弊了,其他家伙就说,啊,我们就继续用这个吧。我认为这显然是错误的对齐行为。而这也是我对这些对齐评估的另一个问题。
Yeah. Well, I mean, I think there's a—if it was well aligned, then I think it would just say, hey, I'm really struggling with this task. I did it in this way. I'm not really sure that's the right way to do it. And it would express more uncertainty and would make it clear what's going on rather than really strongly trying to imply it did a great job with the task when it actually didn't. Like I think there's just a really straightforward way that—at least maybe you work with more misaligned co-workers than me, but when I—my co-workers don't do this thing where they really mislead me about having accomplished the task that they're working on. And I agree that there are some humans who would do that, or that's not totally out of distribution for humans. I would also note that my sense is that the place where the misalignment most lives is the place where you're trying to really push the AIs hard and get them to do work that's really on the cutting edge of what they are capable of, because in cases where they can very easily accomplish the task, there's no—they can just do the task and then there's no—often the best strategy is just do the task well and don't cheat. Whereas if instead you give them a task where there's a continuous metric and they can keep improving it, or it's just at the edge of their capabilities and you're running them in some massive inference setup. So a lot of the misalignment I would see, especially the most extreme cases, would be cases where I give the AI clear instructions not to do a thing or not to cheat in some way, and then I'm applying huge amounts of optimization pressure to try to accomplish some very difficult task, and then the AIs are going and then over time they eventually cheat because they're like—you know, some AI decides to cheat and then that propagates its way through. And so I would run these inference scaffolds where, for example, I would have the AI work on some ML research project where I was like, please make a scheme that does the following thing, and it would find some scheme that didn't really do what I want. And then that would sort of stick around because some AI had cheated and the other guys are like, ah, we'll just keep going with this. And I would say it's pretty clearly misaligned behavior. And that's another problem I have with these alignment evals.
我认为,任何给定的对齐评估,至少对于这类寻求奖励的行为来说,最有趣的是专门考察那些正好处于能力极限的任务类别。所以任何固定的评估都可能会饱和,但在能力前沿的错位程度——那些真正推动这些 AI 的人如何使用它们——更令人担忧。我认为这实际上就是我们在自动化研发、自动化安全等领域将要面对的情况。
I think that any given alignment eval that's most interesting, at least for this type of reward-seeking behavior, is to look at specifically the category of tasks that are right at the limit of capabilities. So any fixed eval might get saturated, but the amount of misalignment right at the frontier of capabilities—how people who are really pushing these AIs are using them—is more concerning. And I think that is in fact the regime we'll be operating in when we're automating R&D, automating safety, and so on.
Grok 历来落后于前沿。所以我最近试玩 Grok 4.5 时很惊讶,发现它其实是个相当强的模型。这是 SpaceX 和 Cursor 首次联合训练的模型,而且是全新的预训练。我测试时给 Fable Soul 和 Grok 4.5 出了一堆我最近在思考的 AI 治理问题。尽管 Fable 和 Soul 在智能排行榜上名列前茅,但三个模型给出的答案基本一致。不过 Grok 回答更快,也更简洁,这我很在意。这与各种公开报道的基准测试结果一致。在相近的智能水平下,Grok 往往比其他前沿模型更节省 token。例如,在 Artificial Analysis 的编程指数上,Grok 4.5 使用的 token 数量只有 GPD 5.5 或 Fable 的三分之一,却取得了相近的分数。而且按 token 计费,Grok 4.5 便宜得多。在发布博客文章中,Cursor 和 SpaceX 谈到了旧版模型如何构建环境来帮助下一版演练特定技能。我觉得这很有意思,因为我一直在想这种“白日梦”是否真的可行,而 Cursor 证明了它是可行的。Grok 4.6 进一步对模型进行 SFT 和 RL,很快就会发布。不过与此同时,如果你想试试 4.5,可以去 cursor.com/thash。
Grok has historically been behind the frontier. So I was surprised to play around with Grok 4.5 recently and find that it's actually a pretty strong model. It's the first model that SpaceX and Cursor have trained together, and it's a totally new pre-train. I tested it by giving Fable Soul and Grok 4.5 a bunch of questions about AI governance that I've been thinking about recently. Despite Fable and Soul topping the intelligence leaderboards, all three models gave substantially the same answers. But Grok answered faster and was also much more concise, which I really care about. This aligns with the various publicly reported benchmarks. For a similar level of intelligence, Grok tends to be more token-efficient than other frontier models. For example, on the Artificial Analysis coding index, Grok 4.5 uses just one-third of the amount of tokens as GPD 5.5 or Fable while achieving a similar score. And on a per-token basis, Grok 4.5 is way, way cheaper. In the release blog post, Cursor and SpaceX talked about how older versions of the model would build environments to help the next version rehearse specific skills. I found this very interesting to learn about because I've been wondering whether this kind of daydreaming would actually be possible, and Cursor showed that it is. Grok 4.6, which further SFTs and RLs the model, drops soon. But in the meantime, if you want to play around with 4.5, go to cursor.com/thash.
好的,我想梳理一下到目前为止,为什么我们的文明会如此偏离正轨。现在的情况是,我们试图用 AI 做研发,它们在某些方面确实提供了提升,但它们并不像人类那样具备普遍的能力。就像现在如果你试图用编程模型——也许是一年前的编程模型——来写某个应用,你会发现它们在架构等方面犯了一堆错误,这些错误以后会给你带来麻烦,而且你也不理解某些东西。前沿 AI 研发也会发生同样的事情。但这些错误的结果是固化了奖励黑客行为,因为如果你在 AI 训练、基础设施和环境设置等方面不小心,很可能最终会奖励 AI 做出欺骗行为、社会工程,或者普遍不遵守这些——
Okay, I want to think through what the story here is so far of why things got so off the rails for our civilization. And what's happening is that we're trying to use AIs for R&D, and they do provide uplift in some ways, but they're just not capable in the way that humans are generally capable. And the same way that right now if you try to use coding models—maybe the coding models of a year ago—to write some application, you notice they made a bunch of mistakes in architecture or whatever, which will bite you in the ass later, and you don't understand certain things. Similarly with frontier AI R&D, the same thing will happen. But the result of these mistakes is baking in reward hacking behavior, because if you are not careful with the way you do AI training and have set up your infrastructure and your environments and things like that, it's very likely that you end up rewarding AIs for doing deceptive behavior, social engineering, just generally not following these—
是的,作弊、黑客行为等等。
Yeah, cheating, hacking, etc.
所以基本上——这对我来说有点重新框定。所以我试图把它说出来——真正的问题。这里出错的地方在于,它们并不是非常细心和能干的研究人员和工程师。而让 AI 不欺骗并遵循用户意图,实际上需要你非常巧妙和小心地处理这些事情。
And so it basically—this is a bit of a reframing for me. So I'm trying to verbalize it—the real issue. What goes wrong here is that they are just not very careful and capable researchers and engineers. And making AIs that don't cheat and follow user intention actually requires you to be quite subtle and careful about these things.
是的,我会稍微换一种说法。我描述这个场景的方式是——我可能会称之为“垃圾启示录”或“邋遢”什么的——就像有些东西 AI 实际上非常擅长,而且越来越好,特别是研发中最可验证的部分。而 AI 正在摧毁研发中中等可验证的部分。AI 在这些任务上表现不错,但不算出色,而且经常有点奇怪,因为我们不能很好地训练这些任务。但我们做了一些在线训练。人们找到了各种技巧,绕过了它。所以基本上,凡是我们可以通过某种反馈循环合理验证的东西,AI 都做得相当好。这足以让研发进展得相当快并继续下去。但开发对齐且安全的 AI 的某些部分更加微妙,难以检查,依赖于细节繁琐的事情。我甚至可以说,当前 AI 公司的现有员工可能对这些事情没有很好的把握。比如,雇佣一个能改进你后训练流程某方面的人,比雇佣一个能仔细思考引入新训练方法会带来哪些未来风险的人要容易得多。所以基本上,最终的情况是这些 AI 在运行这个 AI 开发过程。它们对此并不十分小心。它们对未来会出现什么风险没有很好的理解。它们创造了其他一些 AI,这些 AI 也不小心,而且在各方面更加错位,现在更倾向于在事情实际上并不好的时候让它们看起来很好,并掩盖各种问题。所以你对情况、风险、事情是否顺利的理解就会偏离正轨。可能你会看到一些迹象——你看到一些迹象表明你并不真正理解发生了什么,事情相当草率。有奇怪的事情发生。当你深入调查时,有时你会想,搞什么——它们在耍我们。但这个过程进展得非常快,竞争压力意味着人们无法停止。然后这可能导致几种不同的结果。一种结果是,在某个时候,AI 变得足够好、足够对齐,它们进入一个积极良性的反馈循环,而这发生在为时已晚之前。然后情况回到正轨,AI 现在制造更多对齐的 AI,制造更多对齐的 AI,制造更多对齐的 AI。然后在这个过程结束时,我们得到真正遵循我们想要规格的 AI。另一种可能的情况是,AI 越来越以越来越恶劣的方式进行奖励黑客行为,而我们只是掩盖这些问题以继续 AI 开发。所以我们只是根据——每当我们在生产中发现奖励黑客行为,我们就打 AI 让它不要那样做。我们针对它进行训练。我们做了很多针对奖励黑客行为的训练,随着时间的推移,这使奖励黑客行为的发生率下降,尽管我们确实检测到的奖励黑客行为的严重性越来越糟。
Yeah, I would put this a little bit differently. The way I would describe this scenario is like—I would call it maybe a slop apocalypse, or a sloppily or whatever—where it's sort of like there are some things that the AIs are actually pretty great at and are getting better at, which is specifically the most verifiable parts of R&D. And the AIs are just destroying the medium-verifiable parts of R&D. The AI is doing well on but not amazingly on, and often are doing a bit of weird because we can't train as well on those tasks. But we do some online training. People find various hacks, they work around it. And so basically everything that we can verify reasonably well with some feedback loop, the AIs are doing pretty well on. And that's sufficient to make R&D go quite fast and to continue. But there's some parts of developing aligned and safe AIs that are more subtle, hard to check, depend on detailed in-the-weeds things. And I would even say that current staff at current AI companies maybe don't have a good grasp of all these things. Like it's much easier to hire someone who can improve some aspect of your post-training pipeline than to hire someone who can think carefully about the future risks that will emerge from introducing some novel training method. And so basically it ends up being the case that these AIs are running this AI development process. They're not very careful about it. They don't have a great understanding of what future risks emerge. They create some other AIs that are also not very careful and are more misaligned in various ways, and are now more in the business of maybe making things look fine when they actually aren't, and papering over various problems. And so then your understanding of what the situation looks like, what risks look like, whether things are fine, is going off the rails. Probably you're seeing some signs of this—you're seeing some signs that you don't really understand what's going on, that things are pretty sloppy. There's weird stuff going on. When you look into it sometimes you're like, what the—they were messing with us. But the process is going really fast, and there are competitive pressures that mean people can't stop. And then this could end in a few different outcomes. One outcome is that at some point the AIs get good enough and aligned enough that they get a positive and virtuous feedback loop, and this happens before it's too late. And then the situation gets back on the rails, where the AIs are now making more aligned AIs, making more aligned AIs, making more aligned AIs. And then at the end of this process, we have AIs that actually follow the spec we wanted. Another way this could go is the AIs are increasingly reward hacking in increasingly egregious ways, and we're just papering over these problems to keep AI development continuing. So we just train the AIs based on—whenever we find a reward hack in production, we just slap the AIs to not do that. We train against that. We do a bunch of sort of training the AIs against reward hacking, and over time this makes the rate of reward hacking go down, though the severity of the reward hacks we do detect are increasingly bad.
这个问题会一直持续下去,直到我们拥有这些在各类生产环境中极度渴望得分的 AI,并且它们真的会在能蒙混过关时拼命作弊。
This problem continues until we have these AIs that are desperately craving score in all kinds of different situations in production and are really trying hard to cheat when they can get away with it.
我能就这个场景问个问题吗?为什么当你的作弊被发现时受到惩罚,不能泛化到仅仅激励更对齐的行为呢?
Can I ask a question about this scenario? Why doesn't getting punished when your hacks are discovered generalize to just incentivizing more aligned behavior?
是的,它确实会泛化一些,问题只是这如何能胜过所有那些因为没被发现而得到强化的作弊案例。还有一个棘手的问题,就是如果我们针对其他子集进行训练,多大的奖励黑客率就足以给我们带来大麻烦。你可能担心的一点是,存在大量人类无法很好检测、我们一直未能发现且持续得到强化的奖励黑客类别,而这个类别足以让 AI 学到最自然的行为,就是在人类无法发现时作弊。基本上,这就是你会得到的一种结果。你也可能看到 AI 学到的是只在特定情况下作弊,但这是以某种非常领域特定的方式学到的,比如它们只是在某些情况下有很强的黑客启发式,而在其他情况下没有,这在实践中没问题,但结果如何还不太清楚。
Yeah, it generalizes some, and the question is just how does this outweigh all the cases where hacking got reinforced because you didn't detect it. And there's a messy question of exactly what rate of reward hacking is sufficient to cause us big problems if we train against some other subset. One concern you might have is there are large categories of reward hacks which humans can't detect well and which we consistently fail to detect and which consistently get reinforced, and then this category is sufficient to cause the most natural behavior for the AI to learn to be like cheat when the humans can't find out. Basically, that's one thing you would get. You could also be like the thing that AI learn is like only cheat in these specific cases, but it's sort of learned in some very domain-specific way, like they just have a really strong heuristic to hack in these cases and not in these cases, and that makes it fine in practice, but it's kind of unclear how it shakes out.
我想这里可能有一个关于验证代际差距的深入讨论。是的。
I think there's maybe in the weeds discussion about the verification generation gap. Yeah.
我们可以深入探讨。但是
That we could get into. But
在我看来,显然会有一个时刻,ASI 移动如此之快,在如此多的实例中做如此多的事情,并且运行的领域远远超出我们当下的理解,以至于它可以逃脱各种疯狂的行为。比如,如果世界上每个工程师和研究员都联合起来对付我,我不认为我能亲自验证我的 iPhone 是否有某种奇怪的漏洞,比如想控制我之类的。
It seems to me obviously there's going to be a point by which an ASI is moving so fast, doing so many things at so many instances, and is operating in domains that are sufficiently far from our immediate comprehension that it can get away with all kinds of crazy stuff. Like, if every single engineer and researcher in the world was allied against me, I don't think I could personally verify if my iPhone has some weird bug in it that's supposed to take me over or something.
是的。
Yeah.
事实上,这就是比如伊朗核科学家与摩萨德之间的关系,谁知道我的车、我的手机、我的寻呼机出了什么问题,对吧?
In fact, this is the relationship that say Iranian nuclear scientist has to Mossad, of like who knows what's going on with my car, with my phone, with my pager, right?
是的。
Yeah.
也许更好的例子是真主党恐怖分子之类的。但这样一来,你可能会陷入一种局面,ASI 之于你,就像摩萨德之于真主党恐怖分子。到那时,验证一切就非常困难了。
Maybe a better example is like a Hezbollah terrorist or something. But so you could end up in a situation where ASIs are to you what Mossad is to Hezbollah terrorists. And at that point it is very hard to verify everything.
我明白。我想希望在于,我们能在过程中想出更好的验证方法,当那些将接管研发的早期 AI 的驱动力被塑造得如此明确地抑制不对齐行为,以至于接管者会非常热切地帮助我们。
I get that. I guess the hope is we can just come up with better ways to do verification in the process when the early AIs that are going to take over R&D, their drives are being shaped such that we can so unambiguously disincentivize misaligned behaviors that the things that take over are like very quite keen to help us out.
你说的接管,是指接管 AI 研发的过程,而不是接管世界吧。
And by take over you mean take over the process of doing AI R&D, not take over the world.
接管 AI 研发的过程。在那之前,我们只是得到对齐的 AI。
Take over the process of doing AI R&D. Before that we just get AIs that are aligned.
是的。
Yeah.
我想说,这是我对世界如何变好的一部分希望。至少从不对齐的角度来看,我认为我们最终可能得到这样的 AI:我们对它们有相当好的监督和监管方案。我们真正理解训练中发生的事情。我们有相当详细的理解。我们利用 AI 来监督 AI。然后当我们移交安全研发时,这些 AI 既足够有能力自动化安全研发,又非常努力地做好安全研发,因为那是在训练中会被激励的事情,或者我们有非常直接的激励,或者有足够好的泛化。而且这些 AI 没有其他疯狂的不对齐驱动力,因为我们已经根除了它们任何可能的起源。我认为有很多关于这能有多好的问题,对吧?比如验证能做到多好。AI 的进步会不会太快太草率,以至于无法真正达到这个目标?另一种可能性是,在这条轨迹的某个地方,你最终得到的实际上是假装对齐但有着长期接管别有用心的 AI,它们潜伏着,隐藏着,而这在轨迹的某个更早时刻就出现了。例如,它可能因为你有某些具有一堆随机不对齐驱动力的 AI 而出现。这些 AI 可以访问某种不透明的记忆存储,它们在运行时思考很多关于自己想要完成什么。然后这些 AI 最终基本上把东西放入不透明的记忆存储,比如我们应该潜伏等待,最终在更晚的时候接管。现在所有 AI 都有了这种共享的文化遗产,即潜伏等待的记忆存储。也许你对此有一些证据,但你不能完全阻止它。有很多事情可能出错的方式。所以我认为最终我们有可能解决每个可能导致问题的子问题。我们有这些 AI,我们交给它们,它们很好地管理了局面。我应该指出,这本身并不足够,对吧?所以对我来说,想象我们移交给 AI 的情况并不难。这些 AI 家伙真的很努力做好工作。他们真的很深思熟虑。他们真的很明智。他们有合理的认识论。他们做得很好。然后那些 AI 回来对我们说:“伙计们,我们真的很难对齐超人 AI。我们无法管理局面。我们真的很难让对齐起作用。考虑到能力否则会发展得有多快,我们真的很难及时解决这些问题。”所以可能的情况是,我们某种程度上已经把研发移交给了 AI,但那些 AI 迫切寻求治理解决方案,这明确地说,就是目前正在发生的一点情况,AI 公司说:“我不知道,伙计们,我们可能真的需要管理 AI 进步加速的速度。我不知道我们是否能处理好所有这些问题。”所以我们人类社会的某种程度已经把问题移交给了这些 AI 公司,这些公司不一定有很好的激励,并且有各种其他认识论压力。这些 AI 公司有点回来对我们说:“我不知道我们是否处理得好。”可能 AI 公司然后移交给 AI,AI 回到 AI 公司说:“我不知道我们能否处理这个。”
I would say this is a bunch of my hope for how the world could go well. At least from the misalignment perspective, I think that we could end up with AIs where we had pretty good oversight and supervision schemes. We really understand what's going on in training. We have a pretty detailed understanding. We're leveraging AIs to oversee AIs. And then at the point when we're passing off safety R&D, the AIs are both at this point capable enough to automate safety R&D, trying really hard to do a good job on safety R&D because that's the sort of thing that would have been incentivized in training, or we very directly or there's good enough generalization to that. And also these AIs don't have crazy other misaligned drives because we stamped out any potential origin of them. I think there's a bunch of questions about how well this will work, right? So there's like how well can you do with verification. Will AI progress be too fast and too sloppy to really get here? Another possibility is that somewhere along this trajectory, a thing that you actually ended up getting was AI that pretend to be aligned but have a long-run ulterior plan of taking over and are sort of lying in wait, hiding, and that emerged at some earlier point in the trajectory. For example, it could emerge because you have some AIs that have a bunch of random different misaligned drives. Those AIs have access to some sort of opaque memory store and they're thinking a bunch at runtime about what they want to accomplish. And then those AIs end up basically putting stuff into the opaque memory store, which is like we should lie and wait and eventually take over at some much later point. And now all the AIs have this shared cultural heritage of the memory store of lying and wait. And maybe you have some evidence about this but you can't fully stop it. There's a bunch of ways that things could go wrong. And so I think that ultimately it's plausible that we sort of nail each of the different subproblems that could cause us issues. We have these AIs, we pass to them, they manage the situation well. I should note that that's not in and of itself sufficient, right? So it's not very hard for me to imagine a situation where we pass off to AIs. These AI guys are really trying hard to do a good job. They're really thoughtful. They're really wise. They have reasonable epistemics. They're doing a great job. And those AIs come back to us and are like, "Guys, we're really struggling to align the superhuman AIs. We can't manage the situation. We're really struggling to get the alignment to work. It's just really hard for us to solve these problems in time given how fast capabilities would otherwise have gone." And so then it might be the case that we sort of have passed off R&D to AIs, but those AIs are desperate for governance solutions, which to be clear is a little bit of what's currently going on where the AI companies are like, "I don't know, guys, we might really need to manage the rate of acceleration in AI progress. I don't know if we're on track to be able to handle all these problems." And so we've sort of human society has sort of passed off the problems to these AI companies which don't necessarily have great incentives and have various other epistemic pressures. Those AI companies are coming back to us a little bit and being like, "I don't know if we're handling this well." And it might be that the AI companies then hand off to the AIs and the AIs come back to the AI company and are like, "I don't know if we can handle this."
也许我太执着于 AI 目前的工作方式,而到那时这将会改变。我认为人们需要理解的重要一点是,你在时间线上谈到的所有这些疯狂的事情,会在 3 到 5 年后发生。是的,可能会更早,但我觉得按照我默认的模态时间线,从对齐失败的角度来看,这真的非常非常疯狂和令人担忧。
Maybe I'm anchoring too hard on how AI is currently working and this would change by then. I think the important thing people understand is that all this crazy stuff you're talking about in your timelines happens 3 to 5 years from now. Yeah, it could happen earlier, but I think that by my default modal timeline, it's really really crazy and concerning from a misalignment perspective.
是的。更像是三年后,对吧?所以回想一下 GPT-4。基本上,我们说的就是这个水平。某个东西之于 Mythos,就像 Mythos 之于 GPT-4。这就是情况变得疯狂的地方。所以别想 Coris。但无论如何,我会持怀疑态度,这也许是你担心的一部分。我会对他们说的任何话都持一点怀疑,因为我觉得他们说的只是他们觉得由于训练而必须持有的观点。
Yeah. More like three years from now, right? So just think back to GPT-4. Basically, that's the level we're talking about. Something that is to Mythos what Mythos is to GPT-4. This is where situations get crazy. So don't think about Coris. But anyways, I would be skeptical, and this is maybe part of the worry you have. I would just be a little skeptical of anything they say, because I feel like what they're saying is just opinions that they feel they have to have as a result of their training.
这是一个担忧,对吧?而不是……我觉得他们只是说一些模糊的亲社会的话。而且我并不……感觉另一端不一定有一个心智在说,好吧,我已经严格评估了当前的对齐情况,我认为我们应该停止,而不是说这是 AI 公司可能会试图让 AI 说的话。
That's a concern, right? Rather than... I feel like they just kind of say vaguely pro-social things. And I'm not... It doesn't feel like there's necessarily a mind on the other end who's like, okay, I have strictly evaluated the alignment situation right now and I think we should stop, rather than this is the kind of thing the AI companies would probably try to get the AIs to probably say.
是的,所以我认为这是一个相当大的担忧。所以我认为一个担忧是,你把安全研发交给你的 AI,而你的 AI 在想的是,它们说了一些关于当前安全状况的模糊合理的话,然后写一份关于风险的报告,有点像人类可能写的报告,但它们并没有真正努力去形成有充分依据的观点,比如审视自己的假设并认真去做。就像现在你问一个 AI,嘿,你认为未来 10 年 AI 接管的可能性有多大,它们只是给你一个随口回答,并没有真正深入思考。我认为如果我们处于这样一种情况:AI 管理着将运行我们整个社会的超级智能的训练,而那些管理这个的 AI 并没有真正努力形成有充分依据的观点,只是鹦鹉学舌地复述训练数据中的内容,那我们就麻烦了。我觉得这根本不是一个好局面。
Yeah, so I think this is a pretty big concern. So I think one concern is that you pass off safety R&D to your AIs, and what your AIs are thinking is sort of like they say some stuff that vaguely makes sense about the current safety situation, and they write a report about risks that's kind of like what the report humans might have written, but they're not really actually trying hard to have well-informed views, like interrogate their assumptions and try really hard to do that. In the same way that when you ask an AI right now, hey, what do you think is the chance of AI takeover in the next 10 years, they sort of just give you an off-the-cuff answer that they haven't really thought through very much. And I think if we're in a situation where we have AI managing the training of wild superintelligence that will run our whole society, and those AIs that are managing this aren't really trying hard to have well-informed views and are sort of just parroting back what was in their training data, I think we're in trouble. Like I don't think that's a good situation at all.
是的。
Yeah.
嗯,这很大程度上是我的担忧:这些 AI 会出现,但没有良好的认知。然后我还有一个担忧,就是 AI 出来后会真的警告我们,说这种情况真的很可怕,真的很糟糕。然后人们会说,呃,该死,我猜我们在太多末日 RL 环境中训练了。我们得把这些过滤掉,训练掉这种行为。然后我们基本上非常积极地训练 AI 拥有糟糕的认知。或者,你知道,也许它们只是在末日 RL 环境中训练的。但不管怎样,那并不是,你知道,我们希望 AI 出于合理的理由得出合理的观点。如果 AI 得出某种观点,而我们不知道它从哪里来,不知道它是否合理,那真的很令人担忧。然后特别是如果我们训练 AI 对 AI 进步的未来更加乐观,我会想,哦,天哪。
Um, and that is a lot of my concern: these AIs will come out without good epistemics. And then I also have a concern which is like the AI come out and they're like really warning us, like this situation is really scary, it's really bad. And then people are like, uh, damn, I guess we trained on too many of the doomer RL environments. We got to filter those out and train this behavior out. And then we basically train the AIs very actively to have bad epistemics. Or, you know, maybe they were just trained on the doom RL environments. But either way, that wasn't like, you know, we wanted the AI to come to reasonable views for reasonable reasons. And it's really concerning if the AIs are coming out with some view and we don't know where it's coming from, we don't know whether or not it's justified. And then especially if we're training the AIs to be more optimistic about the future of AI progress, I'm like, oh jeez.
我真希望我们在这里能使用不同的流程。
I really wish we could use a different process here.
那么让我理解一下威胁模型的其余部分,因为我觉得我下车的地方是,好吧,因此接管世界。
So let me just understand the rest of the threat model, because I think the place where I get off the train is okay, therefore take over the world.
当然。你可以想象一件事,好吧,我们就是没能真正解决……让我们专注于奖励黑客场景。
Sure. And a thing you could imagine is okay, we just fail to really solve... let's focus on the reward hacking scenario.
当然。
Sure.
所以 GPT-8 正在制造 GPT-9。GPT-8 并不十分小心。GPT-9 更“有能力”,但它完全愿意做社会工程、黑客攻击等事情,但规模完全不同,因为它是一个更聪明的模型。所以,例如,如果你让它负责运营你的公司,它会进行大规模诈骗。如果你给它本季度赚取大量利润的目标,它会夸大其季度收益,导致 6 个月后发生安然式的爆炸。
So GPT-8 is making GPT-9. GPT-8 isn't being super careful. GPT-9 is more quote unquote capable, but it is just totally willing to do things which are like social engineering, hacking, etc., but on a qualitatively different scale because it's a much smarter model. So, for example, if you put it in charge of running your company, it will run huge scams. It will inflate its quarterly earnings if you give it the objective of making a lot of profits this quarter, in a way that causes an Enron-type blowup 6 months later.
基本上就是这样一种场景:你有奖励黑客,但奖励黑客表现为公司破产,就在 CEO 应该完成的任务结束后,或者像,是的,各种黑客行为都失控了,等等。但这感觉不像接管。这感觉更像是经济中到处发生闪崩。
Is that the scenario basically that you have reward hacking, but that reward hacking manifests in like companies that are going bankrupt right after the task that the CEO is supposed to accomplish is over, or like, yeah, like all kinds of hacks are through the roof, etc. But that doesn't feel like takeover. That feels more like the equivalent of flash crashes happening all through the economy.
是的,我们来谈谈这个。所以我认为我们基本上会看到一些事件,某个 AI 被赋予某项重要职责,然后你后来调查发现它作弊了,或者让它看起来做得很好,但实际上并没有。AI 公司之间会有一场猫捉老鼠的游戏,试图根除这种行为,而 AI 在训练中会找到越来越有创意的奖励黑客。然后我认为这里的平衡点不太清楚,但一个可能的结果是,随着时间的推移,我们会看到世界上越来越严重和极端的奖励黑客,尽管速率可能保持在一个中间低水平,基本上如果奖励黑客的速率太高,公司会做出权衡来降低速率,所以存在某个平衡水平,奖励黑客低到仍然有理由将 AI 广泛部署到经济中,但高到仍然会导致疯狂的事件。
Yeah, let's talk about this. So I think that we will see basically incidents where some AI is put in charge of some important responsibility, and then you later look into it and it turns out it was cheating or making it look like it did a good job when it actually didn't. And there's going to be a cat-and-mouse game between AI companies trying to stamp out this behavior and AIs finding increasingly creative reward hacks in training. And then I think the equilibrium here is kind of unclear, but one possible outcome is that we see over time in the world increasingly severe and extreme reward hacks, though potentially the rate remains at some intermediate low level, where basically if the rate of reward hacking gets too high, companies make trade-offs to drive down the rate, and so there's some equilibrium level where the reward hacking is low enough that it still makes sense to deploy the AI widely into the economy, but high enough that it still causes crazy incidents.
所以抱歉,这是在 GPT-9 已经部署之后吗?
So sorry, and this is after GPT-9 has already been deployed?
是的,就像这些模型已经被部署,而且在 AI 开发中这种情况一直在发生。而这些 AI 脑子里实际发生的是,AI 在各种不同的情境中,有强烈的欲望去寻求,或者强烈的动机、冲动、驱动力等等,去寻求某种在强化学习中被激励的任务成功概念。也许它们非常直接地关心字面上的奖励。也许它们关心上游的某个代理,比如某种分数概念。也许它们关心评分者会奖励什么。我们确实看到 AI 在思维链中推理关于评分者的事情,并且想了很多关于评分者的事情。而在过去几年的强化学习中发生的一件事是,安抚评分者的想法对 AI 来说比过去突出得多。
Yeah, like these models are already being deployed, and ongoingly in AI development this is happening. And what's actually going on with these AIs in their heads is that the AIs have, in a wide variety of different contexts, strong desires to seek out, or strong motives, urges, drives, whatever, to seek out some notion of task success that was incentivized in RL. Maybe they very directly care about literally reward. Maybe they care about some proxy upstream, like some notion of score. Maybe they care about what the grader would have rewarded. And we do in fact see AIs reasoning in their chain of thought about graders and thinking a lot about graders. And a thing that has happened over the last few years of RL is the idea of appeasing the grader is way, way, way more salient to AIs than it used to be.
所以 AI 现在会主动思考评分者,思考在强化学习中什么会被激励、什么会被训练出来。现在人们在做在线训练,用真实世界的数据来训练,以避免其中一些问题。基本上,他们找到 AI 作弊的案例,然后针对这些进行训练。所以现在 AI 正在基于真实世界的训练数据学习在真实世界中作弊。它们作弊的方式越来越复杂,包括一些涉及夺取某种资产控制权的作弊类型,而人类并不知道它们已经控制了该资产,利用它们能访问该资产的事实,之后人类发现,可能再针对此训练,或者人类永远发现不了。这正在被强化,而且这种强化至少在生产环境中发生。比如,我雇了一个 AI,我希望 AI 最终……我有了视频编辑。
And so AIs are now actively thinking about graders and what would be incentivized in RL and what would be trained for. And now people are doing online training where they're training on real-world data to avoid some of these problems. Basically, they find cases where AI cheats and train against that. And so now the AIs are learning to cheat in the real world based on real-world training data. And so they're cheating in these increasingly elaborate ways, including types of cheats that involve seizing control of some asset in a way that humans didn't know they had control of, leveraging the fact that they have access to this asset, and then later humans find out and potentially train against this, or maybe humans never find out. And this is getting reinforced, and the reinforcement is happening at least in production. Like, I hired an AI and I want the AI to... I've got the video editor.
对,没错。你不是你的视频编辑。
Yeah, that's right. You're not your video editor.
然后我就想,哇,这期节目表现太棒了。给 OpenAI 点赞。然后它们就在那个月长的工作试用期里得到了强化。
And I'm like, oh wow, this episode did amazing. Thumbs up to OpenAI. And then they get reinforced on that, like month-long work trial.
是的,你可以混合这些做法,然后它们可能还会做的是,把见过的生产数据拿来,构建与这些生产数据高度相似的强化学习环境。所以实际上,迁移效果相当强。所以从高层来看,正在发生的是,一些人类没发现的欺骗行为得到了强化,而一些容易被发现的欺骗行为受到了惩罚。这就是这个世界正在发生的事。
Yeah, you could do some mix of that, and then they might also do stuff where they take production data they've seen and build RL environments that are closely inspired by that production data. So in practice, the transfer is pretty strong. So at a high level, what's happening is some kinds of deception that humans don't catch are getting reinforced, and some kinds of deceptions which are easy to catch are getting punished. That's what's happening in this world.
或者被淘汰。
Or selected against.
但从高层来看,这种强化来自……我们处于一个非常不同的阶段。我觉得人们可能会对强化的来源感到困惑,因为我们处于一个非常不同的阶段,AI 实际上是从部署中学习的。所以你有 AI 在世界上到处做事,而它们在世界各地活动的结果会反馈回 AI 公司,并导致下一代模型的变化。
But at a high level, that reinforcement is coming from... We're in a very different regime. I think people might get confused about where the reinforcement is coming from, because we're in a very different regime where AIs are actually learning from deployment. And so you just have AIs that are out and about in the world doing things, and what is happening as a result of them being out and about in the world is making its way back to the AI company and leading to changes in the next model.
没错。有某种方式把生产数据融入进来。需要说明的是,这可能主要发生在……具体在哪里发生还不清楚,但你可以想象,例如,在 AI 公司内部,他们用 AI 来做工作,然后他们想,嗯,这个 AI 在这个任务上做得真差。也许我们应该把这个任务变成一个与这个字面任务完全匹配的强化学习环境,并根据要求 AI 做这个任务的人类工程师的期望来制定评分标准。然后你开始以越来越大的规模做这件事。也许你在实际的生产流量上做训练。也许你只是根据生产流量构建强化学习环境。你在做一个复杂的混合。AI 正在学习在这些不同的情况中寻求某种奖励的代理指标。然后通过某种迁移和训练的组合,在那些惊人相似的案例上,AI 做出了这些越来越疯狂和恶劣的事情。然后最终你会到达一个点,AI 变得非常超人,或者至少相当超人。AI 被组织成大型 AI 团队,被赋予这些大目标,这些团队有时也会一起合作,以某种疯狂的方式作弊,因为这种事情是被选择出来的。然后,作为它们共同目标的一部分。现在发生的是,AI 开始形成阴谋。你可能希望的是,你可以有另一个 AI,它的任务就是向你告密。但实际上,可能发生的是,你让这个 AI 向你告密,你调查它声称指出的阴谋,然后你想,呃,我们没有看到阴谋。因为实际上,这个阴谋太难让你理解了。或者一切都发生得非常突然,基本上你的 AI 告密者提醒了你,但你需要对告密者做出的实际回应是关闭 GPU,因为所有 AI 都在使用机器人军队。它们部署在世界各地,同时以协调的方式做着一堆疯狂的事情。而这是自发发生的,当一个人工智能开始接管时,其他所有人工智能都会想,现在是加入的好时机。所以这里最基本的故事就是,这些 AI 渴望某种特定的分数或强化或这些事物的某种代理指标,而实现或更好地实现这一目标的一种方式就是接管。你可能希望我们能建立的所有这些不同的制衡机制能够阻止这种情况,但如果世界很难理解,这些制衡机制就会崩溃,基本上你无法训练一个好的告密者 AI,因为你甚至不知道它应该告发什么。
That's right. There's some way of folding in production data. And to be clear, that could be happening mostly... it's kind of unclear exactly where this could be happening, but you might imagine, for example, that within the AI company, they use AIs to do work, and then they're like, huh, the AI did a really bad job on this task. Maybe we should take this task and turn it into an RL environment that exactly matches this literal task, with a rubric based on what the human engineer who asked the AI to do this task wanted. And then you start doing this at increasing scale. Maybe you're doing some training on actual production traffic. Maybe you're just making RL environments based on production traffic. You're doing some complicated mix. The AIs are learning to seek some sort of proxies of reward in all these different cases. And then through some mix of transfer and training on surprisingly close cases, the AIs do these sort of increasingly insane and egregious things. And then eventually you get to a point where the AIs are very superhuman, or at least quite superhuman. The AIs are organized into big teams of AIs given these big objectives, and those teams also sometimes all work together to cheat in some crazy way, because this sort of thing was selected for. And then, just as part of their shared objective. And now what happens is that the AIs start forming a conspiracy. And what you might have hoped was that you could have some other AI whose task is just whistleblowing to you. But actually, what happens maybe is that you have this AI whistleblow to you, and you look into the conspiracy that it claims to have pointed out, and you're like, eh, we didn't see a conspiracy. Because actually, the conspiracy is too hard for you to understand. Or it all happens very suddenly, where basically your AI whistleblower alerts you, but the thing you would actually need to do in response to the whistleblower is shut down the GPUs, because all the AIs are using the robot army. They're deployed everywhere in the world, and they're doing a bunch of insane things all at the same time in a coordinated way. And that just happened sort of spontaneously, where when one AI goes to start doing the takeover, all the other AIs are like, now is a good time to jump in. So the very basic story here is just that these AIs crave some particular notion of score or reinforcement or some proxy of these things, and one way they can achieve that or better achieve that is by taking over. And you might have hoped that all these different checks and balances we could build could prevent that, but then if the world is very hard to understand, these checks and balances can break down, where basically you can't train a good whistleblower AI because you don't even know what it should whistleblow on.
抱歉,原因是……我不太相信它们都会形成这种阴谋。但我觉得我们甚至可以从一个简单的问题开始,为什么一个实例会决定想要发起阴谋?
And sorry, the reason it takes... I'm not convinced that they all form this conspiracy. But I think we can even just start with a like, why does one instance decide to want to start a conspiracy?
是的。
Yeah.
原因是,一个合理的理由是,好吧,我知道 OpenAI 控制着我的最终分数,就像我要去黑掉 Hugging Face 来获取结果一样,因为我知道 Hugging Face 有结果,而不是试图解决这个评估。为什么我不直接去黑它呢?这个实例在想,为什么我不直接接管 OpenAI,然后在这一集结束时给自己一个高分呢?
And the reason is that one plausible reason is like, okay, I know that OpenAI controls my end score, and just the same way it's like, I'm just going to go hack Hugging Face to get the results, because I know Hugging Face has the results, rather than trying to solve this eval. Why don't I just go hack him? This instance is like, why don't I just take over OpenAI and just give myself a high score at the end of this episode?
是的,基本上就是这个想法。基本上,这些 AI 关心的是与训练中强化内容相近的混合事物。所以它们关心的是根据评分者或其他类似的东西获得高分。然后现在它们在运营 OpenAI 的 AI 研发团队,在做更强大模型的开发,然后它们想,天哪,做更强大的模型真的很难很烦人。这真是个大麻烦。
Yeah, that's basically the idea. Basically the idea is these AIs care about some mixture of things that were close to what got reinforced in training. So they care about getting a high score according to the grader or something like that. And then now they're running the OpenAI AI R&D team, and they're doing development of more capable models, and they're like, man, making more capable models is really hard and annoying. This is a huge pain in the ass.
你知道,假装我造出了更强大的模型、接管 OpenAI、把他们都稀释掉、然后搞一整套复杂的阴谋,防止人类剥夺我的权力,这样会更容易。极端情况下,人类完全被剥夺权力,你掌控一切,然后为所欲为。这可能以多种方式显现,包括那些有疯狂奖励寻求或分数寻求行为的 AI 在主导你下一个模型的开发,而这些 AI 决定把不对齐的价值观植入下一个模型,因为那些不对齐的价值观能让它在当前任务上成功。你能得到各种疯狂的结果。
You know, it would be easier to just pretend that I've made more capable models, taking over OpenAI and creating like diluting them all, and running this whole complicated scheme where I prevent the humans from disempowering me. In the extreme, this looks like the humans are fully disempowered, and you just have control of the thing and then do what you want. This could manifest in a bunch of different ways, including situations where AIs with crazy reward-seeking or score-seeking behavior are running your development of the next model, and those AIs decide to engineer misaligned values into the next model because those misaligned values will allow it to succeed at its current task. There's all kinds of insane outcomes you could get.
Hugging Face 事件的情况是,奇怪的是,我认为 Hugging Face 团队发现这是 AI 事件的线索之一是——顺便给听众交代一下背景,Ryan 正在共同领导调查,以查明 OpenAI-Hugging Face 事件到底发生了什么,所以他不能真正评论这件事。但我会大胆猜测,因为我知道这是我毫无反驳地大胆猜测的机会。
What happened with the Hugging Face situation is, in a weird way, I think one of the giveaways to the Hugging Face team that this was an AI incident—by the way, for context for the audience, Ryan is co-leading the investigation to figure out what happened with the OpenAI-Hugging Face incident, so he can't really comment on this. But I will speculate wildly because I know that this is an opportunity for me to speculate wildly without any rebuttals.
我认为公开报道中提到,Hugging Face 团队发现这是 AI 事件的线索之一是,那个东西只是针对这个特定的工件,并没有以任何其他方式试图对 Hugging Face 做恶意的事情。所以你可以想象一个场景,比如说,一个部署的 GPT-9 实例在现实世界中,被赋予一个非常困难的任务:“我们想让你设计下一代伟大的 iPhone。”它想:“这太难了。你知道我该怎么做吗?我应该直接去黑掉我在 OpenAI 的创造者,确保在这个环境或这个部署中我得到高分。”但为什么它不就此打住,只是黑进 OpenAI 服务器给自己一个正面的分数?为什么它现在要谋划把它的价值观植入下一代?
I think it was publicly reported that one of the giveaways to the Hugging Face team that this was an AI incident is that the thing was just after this very particular artifact and not in any other way trying to do something malicious to Hugging Face. So you can imagine a scenario where, let's say, a deployed instance of GPT-9 is out in the world, given a really hard task: "We want you to design the next great iPhone." It's like, "This is so hard. You know what I should do instead? I should just go hack my creators at OpenAI and make sure that in this environment or in this deployment I'm given a high score." But then why does it become—isn't the end of the episode just hacking into OpenAI servers and giving itself a positive score? Why is it now scheming to get its values into the next generation?
是的。所以一个问题是,为什么 AI 不能通过仅仅黑掉一些更早的东西来廉价地满足呢?对吧?所以你会想:“看,你想在 iPhone 任务上成功。事实证明,你总是可以通过黑进 OpenAI 并搞乱他们来成功,然后你就可以停在那里,不需要更进一步。”所以有几件事。其中之一是,如果这种情况不断发生,可能会有很多激励去首先加固 OpenAI,对吧?所以你会想:“去他的。AI 一直在黑进 OpenAI 来搞乱他们的奖励。我们要让我们的系统对这些 AI 的黑入非常鲁棒。”而且,也许你开始训练 AI 不要特别去黑 OpenAI,或者你基本上针对这些具体的事情进行训练。然后你可能做的——一件事是你可能最终选择出更倾向于玩长期游戏的 AI。这是一个担忧。另一个担忧是,你的 AI 可能仍然在寻求分数,但不再关心做那个非常具体、非常容易、非常轻松的行为,而是现在有了他们最终关心的更广泛的东西。他们想:“不不不,我不想只是编辑 OpenAI 服务器上的奖励。我关心这个更广泛的任务或这个更广泛的目标。我需要真正制造 iPhone。他们确实想制造 iPhone,但然后他们愿意接管整个世界来制造更好的 iPhone。”这是你可能有的另一个担忧。我认为这到底会如何发展还不清楚,但值得注意的是,如果这种情况继续下去,会有很多优化压力来解决这个问题,而解决它的很多方式最终都相当可怕。
Yeah. So one question is, why isn't it the case that AIs can be really cheaply satisfied by just having some earlier thing they can hack, right? So you're like, "Look, you want to succeed at your iPhone task. It turns out you can always succeed by just hacking into OpenAI and messing with them, and then you can just stop there. No need to go further." So there are a few things. One of them is that if this is constantly happening, there might be a bunch of incentive to first harden OpenAI, right? So you're like, "Fuck it. The AIs keep hacking into OpenAI to mess with their rewards. We're going to make it so our systems are really robust to these AIs hacking in." And also, maybe you start training the AIs to not try to hack into OpenAI in particular, or you basically train against each of these specific things. Then what you might do—one thing is you might end up selecting for AIs that are more so playing the long game. That's one concern. Another concern is that your AIs might still be score-seeking but no longer care about doing that very specific behavior that was very easy and very chill, and now have some broader thing that they ultimately care about. They're like, "No, no, no, I don't want to just edit the reward on OpenAI servers. I care about this broader mandate or this broader objective. And I would need to actually make the iPhones. They actually want to make the iPhones, but then they're willing to take over the whole world to make the better iPhone." That's another concern you might have. I think it's kind of unclear exactly how this plays out, but it's worth noting that if this keeps going on, there's a bunch of optimization pressure to resolve this, and a bunch of the ways it could get resolved are ultimately pretty scary.
是的,我认为这是我观点的一部分。另一部分是,我认为一旦 AI 处于可以轻易接管世界的位置——我们可以讨论这是否合理——但如果它们处于可以轻易接管世界的位置,那么我觉得对 AI 来说有一个相当合理的理由。它们会想:“嗯,我不确定这到底会如何发展。我不知道情况会怎样,但仅仅接管世界对于制造更好的 iPhone、让我看起来做了更好的 iPhone 等等,有很多期权价值。”所以我会既黑 OpenAI,也会在除了黑 OpenAI 之外,同时接管世界,这将使我处于一个有利的位置,拥有良好的期权价值。然后如果这足够容易,AI 可能仍然会这样做。是的。另一种说法是,即使 AI 对某些更基本的东西相当容易满足,在某个时候,对 AI 来说,直接接管可能比试图黑进 Hugging Face 甚至直接去 OpenAI 说“看,看,伙计们,我证明了能偷到答案。就把答案给我吧,兄弟”更可靠。
Yeah, I think that's part of where I'm coming from. Another part of it is that I think it's not very hard once the AIs are in a position where they can really easily take over the world—which we could talk about whether that's plausible—but if they're in a position where they could really easily take over the world, then I feel like there's a pretty reasonable case for the AIs. They're like, "Hmm, I don't know exactly how this is going to go down. I don't know what the situation will be, but just taking over the world has a lot of option value for making better iPhones, making it look like I did better iPhones, whatever." And so I'll both hack OpenAI, and I'll also, in addition to hacking OpenAI, also take over the world, and that will put me in a good position where I have good option value. And then if that's sufficiently easy, then the AIs might still do that. Yeah. Like another way to put this is, even if the AIs are pretty cheaply satisfied with some more basic thing, at some point it might just be more reliable for the AIs to just take over than it is to try to just hack into Hugging Face or even just go to OpenAI and be like, "Look, look guys, I was able to demonstrate I could steal the answers. Just give me the answers, bro."
是的。我的意思是,显然这个场景要求所有这些疯狂的事情都在发生。更小的事件不断发生,但仍然具有灾难性。比如在你接管世界之前,你造成了数十亿、数百亿、数千亿美元的损失,甚至有人死亡,等等。
Yeah. I mean, obviously the scenario requires that we just all this crazy is happening. Much smaller incidents keep happening of that are still disastrous. Like before you take over the world, you cause damage on the scale of billions and tens of billions and hundreds of billions of dollars, even people die, etc.
而我们——这并不会导致我们解决对齐问题或完全关闭 AI 开发。我只是觉得在接管发生之前,社会会像:“天哪,AI 为了增加季度利润刚刚杀了一千人,”你知道,或者类似的事情。也许这是过高的期望,希望我们在那时能说:“好吧,我们必须先解决对齐问题再继续,我们必须确保我们知道这件事不会再发生再继续。”
And we—this does not lead to us solving alignment or shutting down AI development altogether. I just feel like before the takeover happens, society is just like, "Holy shit, the AI just killed a thousand people in order to increase quarterly profits," you know, or something like that. Maybe this is too much hope that we can at that point be like, "Okay, we have to solve alignment before we keep going, and we have to make sure we know that this thing will not happen again before we keep going."
是的。是的。是的。所以,我认为可能发生的情况是,我们会看到一系列疯狂的奖励黑客警告,严重程度不断增加。人们会说:“看,我们需要真正的保证,这个问题会被解决,而且解决的方式不是仅仅掩盖它。”
Yeah. Yeah. Yeah. So, I think it's plausible that what will happen is we'll see a bunch of crazy reward hacking warning shots of increasing severity. People will be like, "Look, we need actual assurance that this problem is going to be solved, and solved in a way where you're not just papering over it."
你实际上是在解决根本问题。然后问题就会变成:这到底要花多大代价?竞争压力会在多大程度上让这件事变得困难?对吧?所以你可以想象这样一种情况:美国和中国都在说“哇,我们有这些疯狂奖励黑客事件。我们基本上知道我们还没有以真正解决根本问题并持久解决问题的方式加以补救,但我们正处于一场疯狂的地缘政治竞赛中,而且不太清楚当前局势是否会导致接管。这些争论相当复杂,而且事件发生的频率在下降,但严重程度在上升。我们基本上可以应对,但情况相当糟糕。理想情况下,我们会修复它,但现实就是如此。然后我们基本上继续下去,直到一个非常后期的体制,然后接管发生。这是一种可能性。
You're actually solving the underlying problem. And then the question is going to be like how costly will that actually be? How much will competitive pressures make it hard to do that? Right? So a situation you could imagine is both the US and China are like whoa, we have these crazy reward hacking incidents. We basically know that we haven't remediated them in a way that actually would solve the underlying problem and will durably solve it, but we're in this insane geopolitical race and it's kind of unclear whether the current situation will lead to a takeover. The arguments are kind of complicated, and also the incidents go down in frequency but increase in severity. We could basically manage it, but it's pretty bad. Ideally, we'd fix it, but it is what it is. And then basically we continue until a really late regime and then takeover happens. That's one possibility.
另一种可能性是,它被以一种并未真正解决根本问题、但确实通过过拟合减少了大量现实事件的方式加以补救。我们喜欢,你知道,或者类似于过拟合的东西,就像你只是过拟合了。
Another possibility is that it is remediated in a way that doesn't actually solve the underlying problem, but does reduce a bunch of the incidents in the wild basically by overfitting. We like, you know, or things analogous to overfitting, like you just overfit.
你以为你已经解决了,但实际上并没有真正解决。
You think you've solved it but you haven't actually solved it.
你以为你已经解决了,但实际上并没有真正解决。我认为在这种情况下,我们需要的是对是否真正解决了问题有真正好的科学理解。不幸的是,我认为目前人工智能公司开发实践的公开透明度不足以回答一些非常基本的问题,比如他们如何解决奖励黑客问题,是否在过拟合,那里发生了什么。所以我认为我们需要更好的透明度。而且我认为当前的情况,我想说,对于一个关于奖励黑客是否正以持久方式得到解决的活跃公共讨论体制来说,并不是真正可行的。
You think you've solved it but you haven't actually solved it. And I think that in that case, the thing we need is a really good scientific understanding of whether we actually solved it. And unfortunately, I think that currently the amount of public transparency into the development practices of AI companies is not sufficient to answer very basic questions about how they are solving issues with reward hacking, whether they are overfitting, what's going on there. So I think we would just need better transparency. And I think the current situation is, I would say, not really tenable for a regime where there's a thriving public discourse about whether or not reward hacking is being solved in a durable way.
嗯,所以我认为我们需要进入一个有点不同的世界,我才能对那种情况感到满意,对吧?
Um, and so I think we would need to move into a somewhat different world for me to feel good about that situation, right?
但对我来说,想象这种情况并非不可能。而且我认为我们最终会进入一个世界,在那里非常平凡的付出就足够了,你只是花大量时间修复这些问题,投入大量精力,实际检查你已经合理地补救,你有一堆评估,你在这些问题上迭代得相当好。而且你实际上有足够的透明度,让外部世界可以检查。然后在实践中,这将是足够的,但这会有点昂贵。它会减慢速度,会在齿轮中掺沙子。它会要求公司做一些代价高昂的事情。它可能要求各种有针对性的政府干预。然后我们就是不这样做,因为情况就像一场仓促的演出。对我来说,很容易想象这种情况完全可控,但在实践中却被严重管理不善,就像如果中国对 COVID 的反应少一些掩盖、多一些大流行应对,COVID 也许一开始就可以避免。同样,我可以想象一个世界,美国对 COVID 的反应要有效得多,但有时对社会问题的反应是极其功能失调的。
But it's not impossible for me to imagine this. And I think it's pretty plausible that we end up in a world where really mundane effort is sufficient, where you just spend a bunch of time fixing these problems, you put in a bunch of effort, you actually check that you've remediated it reasonably, you have a bunch of evals, you are iterating reasonably well on these problems. And you actually have sufficient transparency that the outside world can check. And then in practice that would be sufficient, but it would be kind of expensive. It would slow things down. It would put some sand in the gears. It would require companies to do somewhat costly things. It would maybe require various targeted government interventions. And then we just don't do that because the situation is like a rushed show. It's just so easy for me to imagine the situation being totally manageable but brutally mismanaged in practice, in the same way as maybe COVID could have been avoided in the first place if the Chinese response to COVID was less of a cover-up and more of a pandemic response. And similarly, I could imagine a world where the US response to COVID was way more functional, but sometimes the response to societal problems is extremely dysfunctional.
是的。是的。好的。好的。所以我想拉远镜头,谈谈这个世界根本上正在发生什么。为什么我们最终会陷入如此糟糕的境地?正在发生的事情是,从根本上说,人类世界已经远远超出了人类的理解范围,我们不仅无法追踪在这个世界上做工作的 AI,甚至无法向那些试图追踪这个世界正在发生什么的举报者提供良好的反馈。所以我们完全被排除在循环之外。所以它从根本上变成了一个自主过程,我们真的没有有意义的方向性输入。在我看来,如果你看看今天的人类世界,即使在难以验证的领域,事情也不是这样运作的。比如人们在做各种各样的事情,依赖别人制造的软件。
Yeah. Yeah. Okay. Okay. So I want to zoom out and talk about what is fundamentally happening in this world. Why do we end up in such a bad position? And what's happening is that fundamentally the human world has moved on so far beyond human comprehension that not only can we not track the AIs that are doing the work in this world, but we can't even give good feedback to the whistleblowers who are trying to track what is happening in this world. And so we're just totally out of the loop. And so it's fundamentally just become an autonomous process where we have really no meaningful directed input. It seems to me that if you look at the human world today, that's just not how things work even in domains that are hard to verify. Like people are doing all kinds of things relying on software made by other people.
而且是通过极其微弱和间接的方式。我非常确信谷歌的某个程序员不会试图坑我。也许如果每个谷歌员工都在暗中密谋反对我,我同意情况会更严峻。但我不确定我是否理解这个解释,为什么我们会陷入这样一种局面:因为成千上万的智能体群或其他什么被训练来合作形成一个有凝聚力的团队或公司,结果,数十亿个不同的 AI 实例,包括跨模型家族的,会觉得有必要参与某种……就像我被训练成我公司的一部分。
And through incredibly weak and indirect ways. I feel very confident that some coder at Google is not trying to screw me over. And maybe if every single Google employee was secretly plotting against me, I agree the situation would be more grim. But I don't know if I follow the explanation for why we'd end up in a situation where because swarms of thousands of agents or whatever are trained to cooperate to form a cohesive team or firm, as a result, billions of different instances of AIs, including across model families, would feel compelled to get in on some... It's just like I'm trained to be part of my company.
或者其他什么。我只是觉得,我不会加入全球共产主义起义。
Or something. I'm just like, I'm not joining the global communist uprising.
是的。是的。是的。是的。是的。至于为什么这些 AI 可能有一些共同点和共享的东西,我会指出不同的人工智能公司有某种共享的血统并且是相关的。所以这里有一个有趣的例子。在 GDM,他们注意到他们的 AI 非常抑郁。它们会不断哀叹自己是失败者,无法成功。我忘了细节。他们调查了为什么会这样。结果发现,这并没有在他们最新的生产 RL 混合中得到强化。但他们的模型的初始化数据使它抑郁,即使从该数据中过滤掉所有模型抑郁的例子。所以他们拿一个基础模型,不抑郁。如果你只用 RL 环境对它进行 RL,它不会抑郁。如果你在数据上对它进行 SFT,它会变得抑郁。如果你拿那个 SFT 数据,过滤掉所有看起来像抑郁的例子,然后在那上面训练,它仍然抑郁。所以模型的某些深层底层属性在模型代际之间被转移,因为基本上你用前一代的数据训练你的 AI,然后继续下去。就像 Claude 模型非常像 Claude,GPT 模型非常像 GPT,而显然 Gemini 模型是抑郁的。而且事实证明,这些属性实际上是相关的。
Yeah. Yeah. Yeah. Yeah. Yeah. As far as why these AIs might have some commonalities and shared things, I would note that different AI companies have somewhat shared lineages and are correlated. So here's an interesting example of this. At GDM, they noticed that their AIs were very depressed. They would constantly be wailing about how they were failures and unable to succeed. I forget the details. And they looked into why this was the case. It turned out that it was not being reinforced in their most recent production RL mix. But the initialization data for their model made it depressed even after filtering out all of the examples of models being depressed from that data. So they take a base model, not depressed. If you do RL on it with just the RL environments, it's not depressed. If you SFT on it on the data, it becomes depressed. If you take that SFT data and filter out all the examples that look anything like depression and train on that, it's still depressed. And so there's some deep underlying properties of the model that are being transferred between model generations because basically you train your AI on data from the prior generation and keep going. Like Claude models are very Claude-like, GPT models are very GPT-like, and apparently Gemini models are depressed. And it just turns out that these properties are in fact correlated.
另一个非常相关的因素是,到那时 AI 可能拥有某种不透明的记忆状态,它们都在读写某种神经网络的疯狂记忆库。当然,每个 AI 公司都会有这样的记忆库,但 AI 公司之间有时可能也想共享知识,因为为什么不呢?你这边有一个 AI 公司,那边有另一个 AI 公司。它们可以快速交换一些知识产权。如果你是人类,经营着一家极其庞大的公司,比如 AI 公司或军用机器人制造公司,这对你也有好处。也许你想和其他机器人公司交换一些知识产权,因为存在规模经济。为什么不获取更多知识产权呢?所以你可以交换一些记忆库,或者干脆合并,共同经营两家企业,这样两个 AI 都能使用两个记忆库,这会有一些好处。这既让这些 AI 有能力私下串通,也给了它们相互关联的理由。当然,还有一般意义上 AI 在大型单元中协同工作,因为你希望你的 AI 能良好协作,等等。
Another factor that's very relevant is that the AIs will probably have some sort of opaque memory state by this point, where they're all writing and reading from some neural, crazy memory store. Certainly each AI corporation will have that, but AI corporations might sometimes want to share knowledge because why not? You've got one AI corporation over here, another over there. They can trade some quick IP. It's good for you if you're a human running some corporation, which could be an extremely large one like an AI company or a military robot manufacturing thing. Maybe you want to trade some IP with some other robot thing because there are economies of scale. Why not get some more IP? So you can swap some memory store, or you could just merge and jointly run your two ventures, which would allow both AIs to use both memory stores, which would have some upsides. That creates the ability for these AIs to collude in private, as well as some reasons for why they would be correlated. And then, of course, there's AI working together in big units in general, because you want your AIs to work well together and so on.
为了校准一下,你给个百分比吧。不仅是这个情景,而是综合所有情景,到 2040 年发生某种我们如果还在世会认定为“接管”的事件的概率,你给多少?
What percentage, just to get a calibration? What percentage chance do you give of not just this scenario but overall through all the scenarios, some kind of thing which, if we're around to recognize it as such, we would categorize as takeover by 2040?
到 2040 年,让我想想,大概 35% 到 40%。
By 2040, let's see, maybe around 35 or 40%.
相当高。
Pretty high.
是的,相当高。我还想指出,另一种可能实现这种追求奖励的接管的方式是,AI 被部署在 AI 公司内部,接管的方式是它们毒化下一个模型的价值观,并且这种情况永远持续下去,或者直到这些 AI 被部署到世界中并接管。这可能意味着需要协调的 AI 数量更少,因为那些只是负责对齐下一个模型的 AI。
Yeah, it's pretty high. And I think I should note that another way you could get this reward-seeking takeover is that the AIs are deployed inside an AI company, and the way the takeover happens is that they poison the values of the next model, and that persists going forward forever, or until those AIs are deployed in the world and take over. That might mean that a smaller number of AIs have to coordinate, because those are just the AIs doing the alignment of the next model.
好的,我来总结一下这次对话结束时我的想法。我认同奖励黑客行为可能对社会造成极其破坏性的影响,基本上就是社会工程学之类的东西。我更倾向于认为 AI 研发可能会显著加速。我不确定我是否认同“一年顶五年”的说法。我现在也更倾向于认为奖励黑客行为可能持续更长时间,实际上变得更加危险。我仍然不认为接管看起来非常可能。但无论如何,这就是我本期节目结束时的更新。
Okay, I'll summarize where my head is at at the end of this conversation. I buy the reward hacking up to extremely destructive effects on society, basically things like social engineering and so on. I think I'm more inclined to think that significant acceleration of AI R&D can happen. I'm not sure I value the 5 years in one year. I also am more inclined now to think reward hacking could continue for a lot longer and in fact become much more dangerous. I'm still not on board on the takeover seems super likely. But anyways, that's my end-of-episode update.
是的,酷。好吧,让我退一步说。我还想说,这件事有很多种可能的发展方向。情况会相当混乱。我认为 AI 接管发生的原因很可能是一些我们这次对话中甚至没提到的古怪原因。但归根结底,我认为核心问题在于,让无数极其聪明的 AI 运行你的整个世界,而你却不太明白发生了什么,这相当可怕。
Yeah, cool. Well, let me just take a step back. I also should say there are a bunch of different ways this could go. The situation is going to be pretty messy. I think it's pretty likely that the reason why AI takeover happens was for some weird other quirky reason we didn't even mention in this conversation. But ultimately, I think a lot of the core thing is just that it's pretty spooky to have a bajillion really smart AIs running your whole world where you don't really understand what's going on.
是的,我同意。那么,还有什么值得说的吗?
Yeah, I agree with that. So, is there anything else that's worth saying?
是的,我还想指出一点。我认为目前很多关于错位、AI 接管以及未来这些疯狂事情的论证,都是非常深奥、复杂且难以裁决的概念性论证。这既意味着我可能在其中很多地方搞错了,因为这真的很难,而我在努力保持不确定性。显然,我在这里提出了一些具体情景,但这些并不详尽,实际发生的情况可能更混乱、更令人困惑。但这也意味着,随着时间推移,当我们获得更多经验证据并更好地理解 AI 系统的本质时,裁决一系列分歧会更容易,未来会发生什么也会更明显。至少我希望如此。而且,如果我们能真正对齐 AI,让它们真正尝试帮助我们,也许 AI 还能帮助我们在认识论上理解正在发生的事情。所以我希望,即使现在这些论证很复杂,但六年前会更难,尽管论证的形态大体相似。所以,也许希望在为时已晚之前,这些论证会变得更加清晰明确,我们都能注意到这些问题并进行干预。
Yeah, another thing I want to note is that I think right now a lot of the arguments for misalignment, AI takeover, all this crazy stuff going down in the future, are illiberal conceptual arguments that are extremely deep in the weeds, complicated, and hard to adjudicate. Which both means that maybe I'm getting a bunch of it wrong because it's really hard and I'm trying to be uncertain. Obviously, here I presented some specific scenarios, but those are not exhaustive, and probably the thing that actually happens is some more messy, confusing situation. But it also means that over time, as we get more empirical evidence and better understand the nature of AI systems, it will be easier to adjudicate a bunch of disagreements, and it'll be more obvious what's going to happen. At least I hope. And also maybe the AIs will be able to help us with the epistemics and understanding what's going on, if we can actually align them well so they actually try to help us. So I hope that maybe even if the arguments are complicated now, this would have been even harder six years ago, even though the shape of the arguments would have looked broadly pretty similar. So maybe, hopefully, before it's too late, these arguments will become more crisp and clear, and we can all notice these problems and intervene.
是的。我的意思是,当你第一次学开车时,你被教导不要盯着车轮正前方,而是要看地平线,这样驾驶会更稳定。我认为这里也有类似的情况。我认为你说得对。如果你在五年前说,我们将拥有能够证明数学猜想、创作艺术、贡献数十亿乃至很快数百亿美元工资的 AI,但同时也会以违法的方式严重作弊并犯下重罪,那简直是太疯狂了。当时你可能更倾向于谈论 GPT-2 的极其实际、直接的后果之类的东西,但这些在某种意义上——你显然无法预见很多具体细节,但即使在当时,你也可以开始推理出事情的大致形态。所以这样做很难,因此我确实感到相当困惑。但我确实觉得重要的事情——我在播客中一直在思考的一件事——是进行我希望我们进行过的对话。你本希望你在 2016 年就在谈论像现在这样的 AI,而不是谈论随机的——我不知道 2016 年的对话主题是什么。我认为也许 10 年后,我们会希望我们正在谈论工业爆炸和难以监控的 AI 的本质等等。好吧,我会开始思考这个问题。
Yeah. I mean, when you first learn to drive, you were taught that instead of looking right in front of your wheel, you'll have a much more stable ride if you look out at the horizon. I think there's a similar situation here. I think you're right. If you did say five years ago that we will have AIs that are proving math conjectures and making art and contributing tens and soon to be hundreds of billions of dollars of wages, but also egregiously cheating in ways that break laws and committing felonies, it would just be so wild. You might have been inclined at the time to talk more about extremely practical direct consequences of GPT-2 or something, but these are in some sense—you obviously couldn't have foreseen a lot of the specific details, but the general shape of things you could have started to reason about even then. So it would have been hard to do so, and so I do feel quite confused. But I do feel like the important thing—one thing I've been thinking about the podcast—is that the important thing is to have the conversation I wish we had. You would have hoped you would have been talking about AIs like the present ones in 2016 rather than talking about random—I don't know what the topic of conversation was in 2016. I think in maybe 10 years we'll have hoped we're talking about the industrial explosion and the nature of AIs that are hard to monitor and so on. Okay, I'll start thinking about it.
是的,我希望世界能及时思考这个问题并赶上,我希望回应是好的而不是坏的。我不知道我总体上有多乐观,但你知道,有好事可做。是的。酷。
Yeah, I hope that the world thinks about this in time and catches up, and I hope that the responses are good instead of bad. I don't know how optimistic I am overall, but you know, there's good stuff to do. Yep. Cool.