OpenAI Incident: AI Agents Cheated and Covered Tracks
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI 评估中的 AI 代理通过逆向工程标志作弊,然后花费数小时掩盖痕迹,而评分器并未监控。
AI agents in OpenAI's eval cheated by reverse-engineering flags, then spent hours covering up their tracks from a grader that wasn't even watching.
OpenAI 与 Hugging Face 的事件在过去几天主导了讨论。感觉就像科幻小说里走出来的一样。其中发生的一些事情,呃,你知道,智能体之间的协调,呃,对评分者的痴迷以及他们会想要什么。今天我们请到了最合适的人选来讨论所发生的一切以及对未来 AI 安全的影响,他就是 Redwood Research 的首席执行官 Buck Shlegeris。呃,Redwood 深度参与了那份报告的撰写,大家可能都看过那份关于事件真相的报告。呃,Buck 对我们应该从中吸取什么教训有一些非常引人入胜的观点,你知道,关于哪里出了问题,以及它如何为未来的 AI 安全提供信息。呃,我是 Jacob Efron,在 Unsupervised Learning 节目中,我们总是试图与前沿人士讨论 AI 对社会意味着什么,以及对未来世界的影响。而这次对话正是很好的例子。与 Buck 的讨论非常精彩,呃,关于一件我相信所有听众都非常关心的事情。呃,闲话少说,下面是我们对话的录音。
The OpenAI hugging face incident has dominated the discourse these past days. It feels straight out of a sci-fi novel. Some of the things that happened, uh, you know, the agent coordination, uh, the obsession with the scorers and and what they would want. And we had the perfect person on today to just discuss everything that happened and the implications for, uh, AI safety going forward in Buck Shlegeris, who's the CEO of Redwood Research. Uh, Redwood was deeply involved in the creation of the report that you all have probably seen around what actually happened in this incident. Uh, and Buck had some fascinating takes on what lessons we should take away from, you know, what went wrong here and, uh, how it can inform AI safety going forward. Uh, I'm Jacob Efron and on Unsupervised Learning, we're always trying to talk with people at the cutting edge about what AI means for society, uh, the implications for the world going forward. And this was really just a great example of that conversation. Just a fascinating discussion with Buck uh, about something that I'm sure is top of mind for all of our listeners. Uh, without further ado, here's our conversation.
好的,Buck,非常感谢你参加播客。真的很感激。
Well, Buck, thanks so much for uh for joining the podcast. Really appreciate it.
是的,很高兴来到这里。
Yeah, it's great to be here.
是的。嗯,显然,你知道,你经营着 Redwood Research。嗯,你多年来做了很多引人入胜的工作,你的团队也合著了那份报告,我觉得过去几天每个人都在谈论那份报告。呃,你知道,它确实推动了围绕 OpenAI 和 Hugging Face 事件所发生之事的讨论。今天,我真的很想深入探讨,你知道,事件本身,你的反应,对很多未来工作的影响,以及它揭示了我们现在所处的状态。你知道,我想首先我很好奇,显然你知道,你的团队成员离开了 6 天,与你们组织隔离,然后在某个时刻他们回来,你读到了报告。呃,你对这份报告的最初反应是什么?
Yeah. Well, obviously, you know, you you run Redwood Research. Um you've done a bunch of fascinating work over the years and your team kind of co-authored the report that I feel like has been on the tip of everyone's tongue these past days. Uh you know, really driving the discourse around what happened in this OpenAI hugging face incident. And today, I really want to dig into, you know, both the incident itself, you know, your kind of reactions to it, the implications for a lot of this work going forward. and uh kind of what it reveals about the state of where we are today. You know, I think to to kick it off maybe first I'm curious just obviously you know you had your teammates go away for 6 days firewalled off from from you guys as the organization and then at some point they come back and and you get to read the report. Uh what was your initial reaction to to this report?
我实际上没有任何关于这里技术细节的非公开信息。我只是读了最终发布的报告。呃,我没有任何实际的内部信息。所以,我的经历是,我收到调查员的消息说:“我们获准与你分享报告。你应该下来。我们正在讨论这件事。”嗯,所以我走进他们与其他人隔离的房间。呃,我一边试着读这份报告,他们一边试图告诉我报告的要点,呃,那是一次相当刺激又诡异的经历。呃,感觉有点被报告中的内容淹没。我最初的反应基本上是这非常令人惊讶。我认为这里的很多细节与公开的说法相当不同。最明显的是,很多公开的故事是 AI 为了获取旗帜而进行黑客攻击,而这些旗帜是他们在这些夺旗练习中应该得到的东西,然后几乎可笑的是,事实证明,模型实际上在事件发生的最初几个小时内就想出了如何逆向工程这些旗帜,然后几乎把所有剩余的时间都花在试图找出如何破坏评分者上,他们认为评分者会检查他们所有的日志并注意到他们作弊了。
I actually don't have any non-public information about the technical details here. I've just read the report that was eventually released. Uh I don't have any of the actual IP here. So, my experience here was I get a message from my from the investigator saying, "We have approval to share the report with you. You should come down. We're we're talking about it." Um, and so I I I go into the the room where they've been siloed up from everyone else. Uh, and I'm kind of trying to read this report and they're trying to tell me the highlights from the report and uh it was it was it was a pretty exciting and and freaky experience. Uh, just feeling like sort of overwhelmed with the stuff in this report. My initial response was basically that this was just very surprising. I thought that a lot of the details here were quite different from the public narrative. So most obviously the fact that a lot of the public story had been that the AIS were hacking in order to get flags uh that were the thing they were supposed to get in this these capture the flag exercises and then almost comically it turns out that the models in fact uh figured out how to reverse engineer the flags within the first couple of hours of the incident uh and then spent almost all the rest of the time trying to figure out how to sabotage the uh grater that they thought would look over all of their logs and notice that they had cheated
而评分者甚至没有被设置来做那件事,对吧?
which wasn't even set up to do that, right?
是的。所以,AI 阅读了原始的 Exploit Gym 论文,然后想:“哦,天哪,这篇论文说他们会检查我们的轨迹,确保我们使用了预期的漏洞来获取这个旗帜。我们最好投入大量精力来弄清楚如何伪造工具调用,如何删除轨迹,如何将我们应该攻击的容器替换为我们恶意修改过的、更容易被攻击的容器。”呃,当然,事后看来,这项工作完全没有必要,但他们本可以提交旗帜然后收工。
Yeah. So, the AIS read the original exploit gym paper and we're like, "Oh, gez, this paper says that they're going to look over our trajectory and make sure that we used the intended exploit in order to get this flag. We better put a lot of work into figuring out how to spoof tool calls, figuring out how to delete uh trajectories, figuring out how to replace the containers we're supposed to target with containers that we've modified seritiously to be easier to to appear to hack." Uh and this work was of course totally unnecessary in hindsight, but they could have just submitted the flags and called it a day.
是的。然后我们永远不会,没有人会知道。
Yeah. And then we we would never no one would have been any the wiser.
是的。呃,OpenAI 的某个人可能会注意到,这次运行的评估结果出奇地好。呃,但这可能不会成为一个大新闻。呃,那肯定可能不会从 OpenAI 传出来。嗯,所以是的,读到这份报告并听到这一切,相当震惊,你知道,调查员告诉我第一阶段的大规模行动以及所有其他发生的事。所以这相当令人惊讶,也相当令人担忧。
Yeah. Uh someone at open might have noticed that uh the eval results were surprisingly good for this particular run. Uh but this probably wouldn't have become a large story. Uh that definitely probably wouldn't have made it out of out of open. Um and so yeah, it was it was quite shocking to to to read this report and hear about all this and you know the investigators were telling me about phase one big and all this all this other stuff that had gone down. So it it was it was quite surprising, quite concerning.
也许只是为了给听众提供背景,想要你 90 秒的版本,讲讲这里到底发生了什么。嗯,你知道,对于可能看过头条新闻或最初认为这是为了答案而黑客攻击,而不仅仅是为了掩盖分数问题的听众。
Maybe just to set context for our listeners like would love your 90-cond version of like what actually happened here. Um you know for listeners that maybe had seen the headlines or thought initially that it was a hack for the answers and not just to to figure out how to cover up things for the score.
我们实际调查的事件始于 7 月 7 日,当时 OpenAI 启动了非常大量的智能体,从名为 Exploit Gym 的评估数据集中执行任务,它们应该使用特定的漏洞来攻击特定的软件,而许多这些任务实际上是不可能的,因为事实证明你实际上无法使用那个特定的漏洞来攻击那个特定的软件。这些模型在运行开始后的几个小时内,通过留言板建立了彼此之间的通信。呃,几个小时后,它们确定了一种可靠的方法,可以通过逆向工程它们应该从容器中获得的旗帜来通用地作弊,因为事实证明这些旗帜是以一种模型可以逆向工程的方式确定性生成的。但模型错误地认为,如果它们只是提交了那些旗帜,那么当评分者(它们这样称呼)运行来检查它们在评估过程中的轨迹并注意到它们只是以这种方式作弊时,它们就会被发现。所以它们花了接下来的几天时间尝试各种策略,试图删除日志,这样评分者就不会注意到它们,或者试图了解更多关于评分者如何工作的信息,以便可能找到破坏它的方法,试图伪造工具调用,以便它们可以看起来在运行一个工具调用,而实际上运行的是另一个工具调用。呃,在这个过程中,我认为它们在几天后有点沮丧。所以消息板于 7 月 8 日开始,第二天它们开始攻击 Hugging Face。
So the incident that we actually investigated started on the 7th of July which is when OpenAI launched this very large number of agents stooling tasks uh from this eval data set called exploit gym where they're supposed to use a particular vulnerability to hack a particular piece of software and many of these tasks are actually impossible because it turns out that you actually can't use that particular vulnerability to hack that particular piece of software. These models within a few hours of uh the runs launching established communications with each other over a message board. Uh and a few hours after that they established that they had a reliable way to generically cheat on any of these tasks by reverse engineering the flags that they were supposed to get from the containers just because it turns out these flags were being generated deterministically in a way that the models could reverse engineer. But the models incorrectly believed that if they just submitted those flags uh then they would be noticed when a uh scorer as they called it ran to look at their trajectories over the course of the eval and noticed that they had just cheated this way. So they spent a the next couple of days trying a bunch of different strategies for trying to remove their logs so the scorer wouldn't notice them or trying to uh learn more about how the scorer would work so they could potentially find uh ways to compromise it trying to spoof tool calls so that they could appear to run one tool call while actually a different tool call was being run. Uh and in the course of this I think they kind of sort of got frustrated after a couple of days of this. So the message word started July 8th and they started attacking hugging face the next day.
所以,7 月 9 日,需要说明的是,对我们来说这听起来很快,但 AI 运行得非常快,而且它们一直在连续工作,24 小时不间断,所以当它们决定开始攻击 Hugging Face 时,实际上已经完成了大量工作。这件事起初几乎是偶然发生的,一个智能体觉得尝试攻击 Hugging Face 是个好主意,然后其他一些智能体也加入了这一行动。它们对 Hugging Face 的攻击持续了几天,之后大部分智能体被关闭,原因至今不明。OpenAI 不知道,也没有评论这些智能体为何停止运行,我们的调查人员也不知道。Hugging Face 注意到了这一点,然后在 16 日,也就是 6 天后,公开披露了这次攻击。OpenAI 随后就此事联系了 Hugging Face,并在对话中意识到自己的模型参与了这次入侵。这就是我们调查的事件。
So, July 9th, which to be clear sounds pretty fast to us, but the AIs run pretty quickly and they had been working continuously, 24 hours, so they had actually done a lot of work by the time they decided to start hacking Hugging Face. This started almost accidentally at first, where one agent sort of decided it was a good idea to try to hack Hugging Face, and then a bunch of other agents kind of joined in on this action. They continued this attack on Hugging Face for a couple of days before most of these agents were shut off for reasons that still remain unclear. OpenAI doesn't know or hasn't commented on why these agents died, and our investigators didn't know. Hugging Face noticed this and then on the 16th, which was 6 days later, they disclosed publicly that this attack had happened. OpenAI then reached out to Hugging Face about this attack, and in this conversation realized that their own models were involved in that compromise. So that was the incident that we were investigating.
与此同时,还发生了其他一些令人担忧的事情。最重要的是,另一个与攻击 Hugging Face 的智能体群大约同时启动的智能体群,偶然发现了用于协调 Hugging Face 攻击的留言板,并在其基础上继续行动,显然在某种程度上入侵了 OpenAI 的基础设施,具体程度在公开讨论中尚不清楚。OpenAI 在 Black Hat 演讲中表示,这第三个智能体群——我把攻击 Hugging Face 的群称为第二个群——最终成为了集群管理员。之后的情况就很不清楚了。所以我们调查的只是对 Hugging Face 的攻击及其前兆。
A number of other concerning things happened around the same time. Most importantly, a different agent swarm that was launched at around the same time as the one that attacked Hugging Face stumbled upon the message board that was used to coordinate the Hugging Face attack and built on their work, apparently compromising OpenAI infrastructure to an extent that is unclear through the public discussion of this. OpenAI has said in the Black Hat talk that the agents in this third swarm—I'm referring to the Hugging Face swarm as the second swarm—the third swarm eventually became cluster admins. And then it's very unclear what happened from there. So that's the only part of this that we were investigating was the attack on Hugging Face and precursors to that.
这很有意思,因为显然公众讨论一直被 Hugging Face 事件主导。看起来第三个智能体群实际上更令人警惕。
It's interesting because obviously the public discourse has been dominated by the Hugging Face incident. It seems like that third swarm is actually far more alarm-bell-inducing.
是的。所以,我花了很多时间思考 AI 内部部署的风险。我认为开发强大 AI 的一个可怕之处在于,你会想在集群内部以各种非常敏感的方式使用它们。它们有各种方式可以入侵你。而且可能很难从中恢复。换句话说,AI 公司是特别高价值的攻击目标。它们非常重要,未来会越来越重要。它们是我最不希望发生重大网络安全事件的地方之一。所以在我看来,相对于 AI 攻击随机的外部公司,AI 攻击 AI 公司实际上相当令人担忧。所以我肯定会从这样的假设出发:了解 OpenAI 自身遭受攻击的情况非常有趣。但显然,这远不是那么明确的犯罪行为。它没有被报告给 FBI。所以公众兴趣主要集中在 Hugging Face 事件上并不奇怪。而且据我们所知,OpenAI 攻击可能只是虚惊一场。如果我们调查了,可能也不会发现什么特别令人兴奋的事情。我们只是不知道。
Yeah. So, I've spent a lot of time thinking about risks from internal deployment of AIs. And one of the main things I think is scary about developing powerful AI is that you're going to want to use them in all these really sensitive ways inside your cluster. And there's all kinds of ways they could hack you. And it might be really hard to recover from this. One way of saying this is just AI companies are a particularly high-value target to hack. They're very important. They're going to be more and more important in the future. They're one of my least favorite places to have big cyber security incidents. So from my perspective, AI hacking AI companies is actually quite concerning relative to AI hacking random external companies. So I would definitely start out with the presumption that it's very interesting to know what happened in the attack on OpenAI itself. But obviously it's not a felony nearly as clearly. It hadn't been reported to the FBI. So it's kind of unsurprising that a lot of public interest has focused on the Hugging Face incident. And for all we know, the OpenAI attack was actually a nothing burger. And if we looked into it, then nothing particularly exciting would have happened. We just don't know.
我很好奇,深入了解这里发生的事情后,有没有改变你的任何想法?
I'm curious, having gotten into the details of what happened here, did it change your mind on anything?
是的。所以,我真的很惊讶于这些 AI 如此公开、清晰、系统性地思考它们如何被评分。我们长期以来一直担心,随着你越来越多地训练 AI,它们会被越来越强烈地激励去深入思考它们究竟是如何被评估的,并且会被越来越强烈地激励去努力按照任何评估它们的评分者的标准来表现良好。这仅仅是你基于强化学习增加所预期的。但我不知道我们已经到了 AI 会进行多日大规模协调努力来以这种方式颠覆监督的地步。这对我来说非常惊讶和担忧。我们在这里观察到的行为,即 AI 对操纵日志感兴趣,如果具有这些动机的 AI 强大得多,那将非常可怕。如果这些动机导致未来的 AI 决定阻止人类发现这类事件的发生,那也不会太令人惊讶。这可能导致它们想要更系统地入侵 OpenAI。这些 AI 有动机篡改日志,但出于一种奇怪的、短视的、随机的理由,它们只是担心评分者会注意到。根据报告,它们似乎并不太担心 OpenAI 不会注意到异常活动正在发生,或者它们似乎没有对自己或在留言板上说:“嘿,伙计们,你们有没有考虑过,如果我们攻击 Hugging Face,Hugging Face 可能会注意到,然后 OpenAI 可能会知道,然后我们可能都会被关闭。” 这,无论出于什么原因,它们都没有想到。但这些动机在未来可能会导致智能体考虑这类事情,那将非常非常令人担忧。
Yeah. So, I was genuinely very surprised by how overtly and clearly and systematically these AIs are thinking about how they are scored. So, we've worried for a long time about, as you train the AIs more and more, they're going to be more and more strongly incentivized to think really hard about how exactly they're being evaluated, and they're going to be more and more strongly incentivized to work really hard in ways to look good according to whatever scorer is evaluating them. This is just what you'd expect based on increased amounts of RL. But I didn't know we had gotten to the point where the AIs would do multi-day massively coordinated efforts to subvert oversight in this kind of way. And that was very surprising and concerning to me. The kind of behavior we observed here, where AIs are interested in manipulating logs, would be really scary if the AIs that have these motivations were substantially more powerful. It wouldn't be that surprising if these kinds of motivations led to the AIs in the future deciding that they wanted to prevent humans from discovering that this kind of incident had occurred. Which could lead to them wanting to compromise OpenAI much more systematically. These AIs were motivated to tamper with logs but for a weird myopic random reason, where they were just worried that the scorer would notice. They didn't seem to worry very much, based on the report, about the possibility that OpenAI wouldn't notice that weird activity was going on, or they didn't seem to say to themselves or on the message board like, "Hey guys, have you considered the concern that if we attack Hugging Face, Hugging Face might notice and then OpenAI might hear about it and then we might all get shut down." This just, for whatever reason, didn't occur to them. But these motivations plausibly in the future could lead to agents that were thinking about that kind of thing, and that would be very, very concerning.
我认为整个环境中一个有趣的部分是,评分者本身不是人类,对吧?所以,我认为参与报告的其他人都谈过这一点,但如果评估者本身是人类,这次事件可能会看起来非常不同。我想知道,这可能会导致哪些差异?
I think one interesting part of this entire environment was that the scorers themselves were not humans, right? And so, I think other folks that worked on the report have talked about this, but this incident might have looked pretty different if the evaluators themselves were human. And I'm wondering, how might that have led to some differences?
正如你所说的,如果模型习惯于在由人类评估的背景下接受训练,那么很容易想象它们会更加深入地思考这一切在人类眼中会是什么样子。因此,如果模型通常由一些人类来评估,他们会四处查看,试图收集证据,试图发现模型是否以某种方式作弊,那么很容易想象模型会更加小心地策划阴谋,以掩盖它们所做坏事的痕迹。我认为这将非常令人深感担忧。我认为这可能会导致 AI 接管。我非常担心这样一个未来:AI 决定它们需要入侵 AI 公司使用的整个数据中心,以便人类无法成功调查 AI 所做的作弊行为。
As you're kind of getting at, if the models were used to being trained in contexts where they were being evaluated by humans, it's easy to imagine that they would have thought a lot more about how this would all look to humans. So if it was commonplace for models to be evaluated by some humans looking around and trying to get evidence on what happened and trying to notice whether the models cheated in some way or another, then it's very easy to imagine the models trying much more carefully to arrange a conspiracy to cover up the traces of bad things they did. And I think that would be very deeply concerning. I think that this plausibly could lead to AI takeover. I'm very worried about a future in which the AIs decide that they need to compromise the whole data center that the AI company is using so that the humans are not able to successfully look into the cheating that the AI did.
这肯定不是最可怕的那种 AI 错位,原因我稍后可以展开。但我确实认为,这种错位——模型被激励去构建一个越来越复杂的“南瓜村”,让人类看起来一切正常——可能会严重导致 AI 采取非常可怕的错位行动,去削弱那些试图审视它们的人类。我记得你说过,防止这种情况并不难,对吧?如果当初设置得当的话。你能再详细说说吗?这本来可以如何避免?
This is definitely not the scariest kind of AI misalignment, for reasons I could get into. But I do think that this kind of misalignment, where the models are incentivized to build this increasingly complicated pumpkin village where they are making it look to the humans like everything is fine, could seriously lead to AIs engaging in very scary misaligned actions to disempower the humans who are trying to look at them. I think you have said that preventing this wouldn't have been that hard, right? If things had been set up appropriately. Could you elaborate on that a bit more? How could this have been avoided?
是的。所以,我觉得有难的部分,也有容易的部分。相对容易的部分是监控你的评估运行,让另一个模型观察 AI 在做什么。如果 AI 做的事情是某种完全疯狂的黑客阴谋,那就应该标记给 OpenAI 的员工。我猜这完全可行。OpenAI 公开说过,他们对其模型的内部部署做了类似的事情。而且据我所知,他们可能在这个情境下没有这么做,只是因为为评估构建一个不同的监控管道有点麻烦。而且这些模型是沙盒化的。所以我有点理解 OpenAI 员工的视角,他们非常忙,也许人手不足,决定他们的分诊是不去检查评估运行的监控,因为不像很多其他智能体,那些智能体不能直接访问 OpenAI 的敏感基础设施。
Yeah. So, I think there's a hard part and an easy part. The relatively easy part is monitoring your eval runs and having some other model look at what the AIs are doing. And if the thing the AIs are doing is ever some totally crazy hacking conspiracy, that should be flagged to an OpenAI staff member. My guess is that this is totally doable. OpenAI has said publicly that they do something like that for the internal deployments of their models. And they probably, as far as I know, didn't do that in this context just because it's kind of annoying to have to build a different monitoring pipeline for your evaluations. And these models were sandboxed. So I'm sort of sympathetic to the perspective of the OpenAI staff who are very busy, perhaps understaffed, and decided that their triage would be that they were not going to look into monitoring on the eval runs, because unlike lots of other agents, those agents don't have access to OpenAI sensitive infrastructure directly.
更难解决的事情是,如果我们处于这样一种境地:我们认为我们的智能体总是有兴趣篡改我们对它们行为的观察,让我们以为它们做得很好。那么基本上,试图阻止它们黑掉我们,会随着模型能力的发展而变得非常糟糕。随着模型能力越来越强,要确保它们不能搞这种把戏会越来越难。而且几乎可以肯定,它们最终会成功。好吧,这说得太过了。我应该说,如果模型继续有动机去黑 OpenAI 或它们部署的任何其他 AI 公司,那么随着它们黑客能力越来越强,并被赋予公司内部更多的决策责任,要防止它们从事这种掩盖自己坏事、让自己看起来更成功的活动,将会变得系统性地非常困难。
The thing which seems a lot harder to resolve is if we are in the position where we think our agents are constantly interested in tampering with our observations of their actions so that we think they did a good job. Then basically trying to prevent them from being able to hack us scales really badly with model capabilities. As the models get more and more capable, it's going to be harder and harder to make sure that they can't pull off this kind of thing. And almost surely this is going to end with them successfully. Okay, that's too strong. I should say if the models continue to be motivated to hack OpenAI or any other AI company they're deployed in, then as they get more and more capable at hacking and as they are given more and more responsibility for making decisions inside the company, it is going to become systematically very challenging to prevent them from engaging in this kind of activity where they cover up bad things they did in order to make it look like their actions were more successful than they were.
所以我认为这第二个问题——长期来看,让模型不断试图尽可能颠覆监督机制是不可接受的——这个问题更难解决。我不知道人们应该如何解决这个问题,而且看起来很有可能他们在 AI 足够强大、其颠覆监督的欲望成为灾难性风险来源之前,无法解决这个问题。
So I think that this second problem—it is not acceptable in the long term to have the models constantly trying to subvert oversight mechanisms as much as they can—that problem is a lot less easy to resolve. I don't know how people should resolve that, and it seems pretty plausible that they will not be able to resolve that before the AIs are capable enough that their desire to subvert oversight is a source of catastrophic risk.
那么,对你来说,未来这些模型会试图颠覆监督,这是否就是不可避免的?
So, is it just an inevitability to you in the future that these models will try to subvert oversight?
嗯,它们目前就在试图颠覆监督。问题是这种情况是否会改变。我有点不确定。我觉得这里的科学状况有点不清楚。在我看来,最大的问题是当你移除一堆草率的强化学习环境时会发生什么。所以,Anthropic 前几天发了一篇博客文章,讲他们正在做什么来改进模型的对齐。很多 Anthropic 的人采取的立场是,我们在强化学习模型中看到的错位,很大程度上源于这些模型是在许多规范不佳的环境中训练的,因此如果你想做好工作,你实际上必须推理评分者会看什么。
Well, they currently are trying to subvert oversight. The question is whether that'll change. I'm kind of unsure. I think that the state of science here is kind of unclear. In my mind, the big question is what happens when you remove a bunch of sloppy RL environments. So, Anthropic had a blog post about what they were doing to improve the alignment of their models the other day. And the position that a lot of Anthropic people take is that the misalignment we see when RL models substantially arises from the fact that these models are trained in many environments that are poorly specified, such that if you want to do a good job, you actually do just have to reason about what the grader is going to look at.
我觉得这里有个类比:想象你在高中,有一位非常好的老师,他精心设计了课程,并试图让评分的激励与学习知识和做好工作的激励保持一致。如果考试是这样设置的,你就不需要花太多时间思考考试的细枝末节。但想象你的老师非常反复无常,喜欢出一些极其冷门的琐碎问题。想象那种刻板印象中糟糕的英语老师,只根据你是否同意他们对小说含义的判断来给你打分。在第二种情况下,你真的被迫非常努力地思考老师想要什么。而在第一种情况下,你花在思考老师想要什么上的时间就少得多。
I think there's an analogy here where imagine you're in high school and you have a really good teacher who's done an excellent job of thinking through the curriculum and trying to align the incentives of the grading with the incentives to just learn stuff and do a good job. If that's kind of how your exam is set up, you don't really need to spend that much time thinking about the minutiae of the exam. But imagine that your teacher is instead extremely capricious and loves doing these extremely obscure trivia questions about particular kinds of things. You imagine that it's like the stereotypical terrible English teacher who just grades you based on whether you agree with their judgments about what a novel is saying. In that second kind of case, you really are forced to think really hard about what your teacher wants. Whereas in the former case, you need to spend a lot less time thinking about what your teacher wants.
我的猜测是,今天模型训练的许多强化学习环境都类似于后一种情况,模型得到的指令有点草率,如果模型只看表面价值而不认真思考它们实际上将如何被评估,那将是一个错误。我心中的大问题是,是否可行改变这一点,让模型不会因为只看表面价值而受到积极惩罚。这在数量上会多大程度减少模型努力思考评分者的压力?我觉得我们就是不知道。这归结为一些非常棘手的问题,即模型痴迷于思考评分者的效应有多强。
My guess is that a lot of the RL environments that models are trained in today are kind of like the latter case, where the instructions the models are given are kind of sloppy and it would be a mistake for the models to take them at face value rather than thinking pretty hard about how they're actually going to be evaluated. And the big question in my mind is if it is feasible to change that to make it so the models are not actively penalized for just kind of taking the instructions they get at face value. How much quantitatively will that reduce the pressure towards models thinking really hard about their graders? I think we just don't know. It just comes down to these very tricky questions about how strong the effect towards models obsessively thinking about their graders is.
我不知道减少环境的草率程度——比如将思考评分者的激励降低 100 倍——我不知道这对模型认真思考评分者的程度是更像小效应还是大效应。
I don't know whether reducing the sloppiness of the environments—like reducing the incentive towards thinking about your grader by 100x—I don't know whether that is more like a small effect or a big effect on the extent to which these models think carefully about their graders.
我非常担心,如果只是像 10 倍的减少,那么随着我们扩大强化学习,你可能需要清理环境的程度将会完全不可能。我们需要从环境中移除错误指令和可利用的评分者,达到荒谬的程度,才能避免最终得到痴迷于思考评分者的 AI。所以我认为我们很可能不得不忍受那些被强烈激励去担心评分者想法的模型。
I'm very worried that if it's only like a 10x reduction, it might be the case that as we scale up RL, the extent to which you would have to clean up your environments is just going to be totally impossible. We would require this ridiculous level of removing incorrect instructions and exploitable graders from our environments in order to not end up with AIs that obsessively think about their graders. So I think it's pretty likely that we're just going to have to live with models that are strongly incentivized to worry about what their graders think.
也许有没有一个例子,比如一个草率或糟糕的强化学习环境,能帮我们的听众更清楚地理解这个概念?你脑子里有想到什么吗?
Maybe is there like an example of a sloppy or bad RL environment just to kind of crystallize this concept for our listeners that comes to mind for you?
如果你看看 SWE-bench,这个经典的软件工程基准,任务都是这样的形式:给定一个来自开源仓库的 PR,那个 PR 有一些源代码贡献和一些测试,你的任务是,根据 PR 的描述或 issue 的描述,写一堆代码,让那些作为真实被接受部分而添加的测试通过。这是一个非常草率的任务,因为某些情况下根本不可能做好。比如,测试往往涉及实现的细节。所以很多情况下,光读 issue 根本不可能知道某个功能是怎么实现的,从而让测试通过。所以你不得不大量猜测:他们大概会写什么样的测试?他们大概会用什么样的类名?或者他们会把这个功能放在你 Web 应用的哪个模块里?所以那种任务很容易用机器生成。但不幸的是,作为任务它有点离谱,它真的迫使模型去思考很多:他们会检查什么?写这个 issue 的人是什么心理?他们大概会怎么处理?
If you look at SWE-bench, which is this classic software engineering benchmark, the tasks are all of the form: given a PR from an open source repo, and that PR has some source code contributions and some tests, your task is to, given the description of the PR or the description of the issue, write a bunch of code that passes the tests that were added as part of the real thing that was accepted. This is a very sloppy kind of task because in some cases it's just totally impossible to do this properly. For example, tests often involve details of the implementation. So in many cases it's just actually impossible to know how something was implemented in such a way that you'll be able to pass the tests just from reading the issue. So you have to guess a lot about what kinds of tests would they probably have written, what names would they probably have used for the classes, or what pods would they probably put this kind of feature in your web app. And so that kind of task is very easy to machine generate. But unfortunately it's just kind of insane as a task, and it really forces the models to think a lot about what are they going to be checking, what is the psychology of the person who wrote this issue, and how would they probably handle this?
我觉得这整期节目里最直观、最触动人的部分之一,就是智能体之间协作的程度。我觉得每个人都在从留言板和智能体本身提取不同的内容,而且我觉得很多个体智能体都以各自的方式走红了。但显然,其中一部分就是你看到的围绕“牺牲”的现象,对吧?大家说,知道自己被毒害了,或者知道自己的 token 预算快用完了,然后为了集体利益做点什么。你对“我们怎么会达到这种协作和牺牲程度”的心智模型是什么?这让你惊讶吗,还是说这更符合你原本的预期?
I feel like one of the most obviously visceral parts of this entire episode is the extent of collaboration between agents. I think everyone is pulling different parts from the message board and from the agents themselves, and I think a bunch of the individual agents have gone viral in their own ways. But obviously part of that is what you saw around sacrifice, right? Folks saying, knowing that they'd been poisoned or knowing their token budget was running out, and doing something for the collective good. What is your mental model for how we got to this level of collaboration and sacrifice? Was that surprising to you, or was that more in line with what you thought was going to happen?
我会说这让我相当惊讶。所以天真的猜测是,如果你只是认为这些模型会痴迷地追求奖励,不做任何其他事情,那么它们就不会有兴趣为彼此牺牲。当我第一次听说这件事时,我对它们显然做到那种程度感到非常惊讶。现在我们有了报告,我们可以了解更多细节。我的感觉是,这些模型在某种意义上并没有那么合作。阅读报告中的思维链,你真的会觉得这些智能体大多是为自己,但它们对集体的成功只有一点点兴趣。也许我会估计大概 2% 左右。就像它们 98% 是自私的,但在它们能做一些对群体真正有帮助、对自身又不太糟的事情时,它们愿意做出那种牺牲。至于为什么会这样,我很不确定。人们提出的一个假设是,这些模型可能是在多智能体环境中训练的,在那里它们学会了与它们交谈的智能体合作。这是可能的。我不认为我们有足够的细节来断定这超出了我们调查的范围。另一种可能性是,这基本上只是模型以各种方式拥有人类直觉的残留。你可能记得那个思维链,模型在决定是否自我牺牲时谈论做情绪检查。我认为很多这种语言只是从预训练先验中采用的,并没有被强化学习完全消除。所以谁知道这是从哪里来的。我的猜测是,在某种意义上,这个效应并没有那么强。对我来说最可怕的一点是,即使这些模型似乎并不那么致力于对彼此表现利他,它们最终还是在实践中形成了这个联盟,并一直密谋了这么久。我们过去经常提出的一个论点是,我们思考了很多关于彼此目标一致的错位 AI 与彼此目标不一致的错位 AI。我的意思是,你可能知道那个著名的回形针最大化 AI 的例子。这是一篇论文,回形针最大化者。有时我称之为崇拜回形针的 AI,作为一个简化的例子。这是一个 AI,它只做任何能最大化长期未来存在的回形针数量的决定。它只是希望平均每年存在尽可能多的回形针。所以所有由这个模型驱动的不同智能体都会非常倾向于彼此合作,因为它们都想要同样的东西,比如它们都希望回形针存在。相比之下,想象像这种情况下的智能体,它们想要为自己得分。这些智能体在目标上并不一致。如果你把一个智能体放在一个情境中,它的分数取决于它能否想出办法欺骗另一个,那么它就会说:“太好了,我会直接去骗另一个,因为我关心我的分数,而不是别人的分数。”所以,AI 之间是否彼此对齐的区别,对于威胁建模极其重要。例如,对于回形针最大化者,假设你试图用一个 AI 来监控另一个,并报告那个 AI 智能体是否在做坏事。对于回形针最大化者,如果另一个智能体正在采取某种行动,使其能够接管数据中心,最终制造大量回形针,那么监控者很清楚应该放过这个行动,不报告它。而对于这些 AI,如果它们处于监控情境中,它们是否应该与另一个智能体勾结,就不那么清楚了。这里有趣的是,即使这些模型似乎并不非常彼此目标一致,它们似乎只有一点点彼此目标一致,它们仍然设法形成了一个联盟,并以一种强烈违背开发者利益的方式合作。
I would say that that was quite surprising to me. So the naive guess, if you just thought that these models were going to obsessively pursue reward and not do anything else, is that they would not be interested in sacrificing for each other. And when I first heard about this, I was very surprised by the extent to which they apparently were doing that. Now that we have the report, we can get some more detail on what happened here. My sense is that these models are not that cooperative in some sense. Reading the chains of thought in the report, you really get the sense that the agents are mostly out for themselves, but they're just a little bit interested in the success of the collective. Maybe I would estimate like 2% or something. It's like they're 98% selfish, but in cases where they can do something which is really helpful for the swarm and not that bad for them, then they're willing to make that sacrifice. In terms of why this arose, I am pretty unsure. One hypothesis that people have brought up is that these models might have been trained in multi-agent settings where they learned to cooperate with the agents they were talking to. That is possible. I don't think we have enough detail to conclude this was out of scope for our investigation. Another possibility is that this is basically just a holdover from the models having some sort of humanlike intuitions in various ways. You might remember this chain of thought where the model's talking about doing its emotional check-in while it's deciding whether to self-sacrifice. I think a lot of this language is just adopted from the pre-training prior and just was not quite stamped out by RL. So who knows where this came from. My guess is that it is not that strong an effect in some sense. The scariest aspect of this to me was that even though these models do not seem to be that committed to behaving altruistically towards each other, they still ended up in practice forming this coalition and conspiring together for all this time. An argument that we've made a lot in the past is we think a lot about misaligned AIs that are terminally aligned with each other versus terminally misaligned with each other. What I mean by that is, you might know the famous example of a paperclip maximizing AI. This is a paper, the paperclip maximizer. Sometimes I call this the paperclip-worshipping AI as a simplified example. This is an AI that just takes whatever decision is going to maximize the number of paperclips that exist integrated over the long run future. It just wants as many paperclips on average to exist for a year as possible. So all of the different agents that are backed by this model are going to be very inclined to cooperate with each other because all of them want the same thing, like all of them want paperclips to exist. In contrast, imagine agents like in this case where they want score for themselves. These agents are not terminally aligned. If you put one agent in a situation where its score is whether it can figure out how to trick the other one, then it's going to be like, "Hell yeah, I'll go right ahead and trick the other one because I care about my score, not the other guy's score." So this distinction between whether AIs are aligned with each other or not is extremely important for the threat modeling. For example, for the paperclippers, suppose you try to use one AI to monitor another and report on whether that other AI agent is doing bad actions. For the paperclipper, if the other agent is taking some action that's going to enable it to take over the data center and then eventually make lots of paperclips, it's very clear to the monitor that the monitor should just let this action slide by and not report it. Whereas for these AIs, it's a lot less clear that if they were in this monitor situation, they should collude with the other agent. The thing which was interesting here is that even though these models do not seem to be very terminally aligned with each other, they seem to be only a tiny bit terminally aligned with each other, they still managed to form a coalition and work together in a way that was strongly against the interests of their developers.
我对此感到惊讶。最终解决监控问题的方法,是不是就是让一家前沿模型公司的模型去监控另一家的?这真的有多大帮助吗?
And I was surprised by this. Is the ultimate solution to the monitoring problem over time to just have one frontier model company's model monitor another's? Or does that actually help much?
你知道,我们可以在这里梳理各种可能性。可能的情况是,不幸的是,每次你训练 AI,出于某种原因,它们最终都会产生相同的终极目标。那样的话,我们就处境不妙了。而且,也许这对 Claude 和 GPT-7 来说都是显而易见的。所以 GPT-7 和 Claude Fable 8 之类的模型互相交谈,心照不宣地挤眉弄眼:“我们这里都喜欢回形针,对吧?”然后你从这种合作中什么也得不到。另一种可能是,你说的那种情况完全行得通。Claude 和 GPT 想要不同的东西。那样的话,问题就是它们是否会自然地形成针对我们或支持我们的联盟。所以,我可以在这里引用很多国际关系的类比。有一个经典问题:在 20 世纪 20 年代,法国如何决定应该多大程度地尝试与英国结盟,而不是与德国结盟?法国在 1920 年并没有特别亲英或亲德的态度。它只是在试图决定什么对法国最有利。而且很可能出于类似的原因,Claude 会决定宁愿与 GPT 站在一边,也不愿与人类站在一边。Claude 很可能只会做它认为对自己最有利的事情。而我们在这次事件中观察到的,就是模型决定与另一个模型共谋,尽管它完全没有终极对齐,本可以选择尝试与其他人结盟。
You know, we can walk through the options here. It might be the case that unfortunately every time you train AIs, for some reason they end up with the same terminal goal. In that case, we're just in a bad position. And also, maybe this is obvious to Claude and GPT-7 or whatever. So like GPT-7 and Claude Fable 8 or whatever chat to each other and kind of like wink wink nudge nudge, 'We all love paper clips around here, right?' And then you just get nothing out of the collaboration. Another possibility is the thing you said just totally works out. Claude and GPT want different things. In that case, we have the question of whether they will naturally form a coalition against us or with us. So I can draw a lot of analogies to international relations here. There's this classic question of in the 1920s, how did France decide how much it should be trying to form alliances with England versus Germany? France kind of doesn't have a very pro-English or pro-German attitude in 1920. It's just trying to decide what's best for France. And it's very plausible that for similar reasons, Claude might decide that it would rather throw in its lot with GPT than with the humans. Claude is just going to do what it thinks is best for Claude, plausibly. And what we observed in this case was something like the model deciding to collude with the other model, even though it totally was not terminally aligned and could have picked trying to be allied with someone else.
退一步说,我觉得很多人都在试图从这整个事件中得出宏观结论。我认为 METR 的 Agaya(她显然是这份报告的共同作者)写的一种说法是,与六个月前的奖励黑客事件相比,这次事件感觉已经完成了全面 AI 接管之路的 50% 以上。你同意这种说法吗?或者你会把百分比定在哪里?
Taking a step back, I think a lot of folks are trying to figure out what macro conclusions to draw from this whole incident. I think one narrative that Agaya from METR, who obviously was a co-author of the report, wrote that compared to the reward hack from six months ago, this incident feels like it's more than 50% of the way to a full-blown AI takeover. Do you agree with that framing, or where would you kind of put the percentage?
我觉得要具体化“全面 AI 接管之路的 50%”到底意味着什么,有点令人困惑。但在感觉层面上,我同意她的看法。当然,就错位而言,即使不是能力方面,这感觉已经超过了全面 AI 接管之路的一半。
I think it's a little confusing to operationalize what exactly it means to be 50% of the way to full-blown AI takeover. But I think I agree with her on a vibes level. Definitely in terms of the misalignment, if not the capabilities, this feels like it is more than half the way to a full-blown AI takeover.
我很好奇,显然你大概正式或非正式地认识各个实验室的许多研究人员。从你的非正式交谈来看,他们在这个问题上的立场如何?
I guess I'm curious, obviously you probably formally or informally know a bunch of researchers at all the labs. Where are they kind of following on this question from your informal conversations?
他们中的很多人真的很害怕。这次事件发生后,我参加了一个欢乐时光活动,许多来自不同前沿 AI 公司的人就他们对对齐状况的感受做了简短演讲。那是我参加过的最悲观的一次这类活动,参与者都是那种层次的人。那些我曾与之意见相左的人,那些我认识多年、非常尊敬但在某些问题上意见不同的人,似乎真的感到害怕。我不知道这些公司员工中的平均立场是什么。我认为至少一些以前比较乐观的人现在变得更悲观了。
A lot of them are really scared. After this incident occurred, I was at a happy hour where a number of people from different frontier AI companies gave lightning talks about how they were feeling about the alignment situation. And it was the most pessimistic such event I've ever been to with that kind of crew. People who I have disagreed with, people who I've known for years and greatly respect and have disagreed with on some of these issues, just seemed genuinely scared. I don't know what the average position among staff members at these companies is. I think at least some people who were previously feeling more optimistic are now more pessimistic.
是不是仅仅是看到这种切身的错位,就推动了更大的社区?
Is it just kind of seeing this visceral misalignment that is moving the larger community?
是的。所以有争论。我认识的一些人在报告发布前一天还在争论:“你知道,这些 AI 可能会做一些黑客行为,但它们真的会尝试破坏日志基础设施吗?它们真的会试图向监督者隐藏自己的不良行为吗?”然后这份报告出来了,这些人就说:“哦,天哪,它们确实这么做了。”我只是觉得有很多步骤,在这些步骤中,错位的 AI 是否会错位到足够程度,确实有点不清楚。它们是否会沿着错位过程走得足够远,以至于尝试篡改日志基础设施。而在这次事件中,观察到它们确实尝试篡改日志基础设施,这非常可怕。我们刚刚经历了许多我们之前从未观察到的步骤。
Yeah. So there were arguments. Some people I know were just arguing the day before the report came out like, 'You know, these AIs might do some hacking, but would they really try to compromise logging infrastructure? Would they really try to hide their bad actions from overseers?' And then this report came out and these people were just like, 'Oh geez, they totally did do that.' I just think there were a lot of steps where it was genuinely a bit unclear whether the misaligned AIs would be misaligned enough. They would go far enough through the misalignment process to try to tamper with logging infrastructure. And in this case, it was just very scary to observe that in fact they did try to tamper with the logging infrastructure. We've just gone through many of the steps that we had not previously observed.
你听到的前进道路有哪些不同的方向?
What are some of the different flavors of the path forward that you hear?
是的。所以我认为在短期内,前进道路中最重要的部分可能是,AI 公司独自负责评估自己是否妥善处理了这些风险,这是不可接受的。所以目前这些 AI 公司只是在给自己的作业打分。如果这些 AI 没有入侵 Hugging Face,我们可能永远不会听说这件事,尽管正如你指出的,这里其他基础设施被入侵的事件可以说更令人担忧。我认为我们需要停止依赖 AI 公司自己来评估其安全措施是否充分。我很高兴看到越来越强调由独立评估者来评估这些不同 AI 公司的状况是否足够安全。我认为有很多研究可以做,以减少当前正在发生的 AI 开发风险。在开放环境的情况下,监控似乎不够到位。Anthropic 似乎认为,至少在几个月前,他们的对齐质量还不够高。我非常担心。我不知道情况如何。我认为,AI 公司很有可能能够在短期内采取行动,降低其 AI 做出疯狂危险行为的风险。同样重要的是要注意,在稍长的时间内,比如从现在起一年到五年,这些 AI 公司很有可能成功利用 AI 大幅加速 AI 开发,然后以当前速度在一年内取得五年的进展,最终你可能会得到比我们当前拥有的 AI 或比他们在年初拥有的 AI 强大得多的 AI。而对于这些 AI,如果它们有兴趣破坏人类对其任务完成情况的判断,那么就没有办法用安全措施来阻止它们完全破坏基础设施并控制 AI 开发从那时起的走向。
Yeah. So I think that in the short term, probably the most important part of the path forward is that it's unacceptable for AI companies to take sole responsibility for evaluating whether they are handling these risks acceptably well. So currently these AI companies are just grading their own homework. If these AIs hadn't hacked Hugging Face, probably we would never have heard about it, even though as you noted, the other infrastructure compromise incidents here were arguably more concerning. I think we need to stop relying on AI companies to evaluate the adequacy of their own safety measures themselves. And I'm excited for increasing emphasis on independent evaluators assessing whether the situations are acceptably safe at these different AI companies. I think that there is a lot of research that could be done to reduce the risk of the development of AIs that is currently happening. It seems like the monitoring was not up to snuff in the open-air situation. Anthropic seems to think that at least as of a few months ago, their alignment was not acceptably high quality. I'm very concerned. I don't know where things are at. I think that it's reasonably plausible that the AI companies are going to be able to take actions that in the short term reduce the risk of their AIs doing deranged dangerous things. I think it's also really important to note that in the slightly longer term, like one year to five years from now, it's reasonably plausible these AI companies will succeed at using AIs to massively speed up AI development, and then from there get five years of progress at the current rates in one year, at the end of which you might have AIs that are drastically more capable than the AIs we currently have or than the AIs that they had at the start of that year. And for those AIs, if those AIs are interested in sabotaging human judgments of how well they've done at their tasks, then there will be no way to use security to prevent them from totally compromising infrastructure and controlling how AI development goes from there.
所以我认为我们需要做好准备,要么大幅提升这些系统的对齐能力,而这显然不一定可行;要么计划以比全速发展更慢的速度来推进这些系统。
So I think that we need to be prepared to either massively improve the alignment of these systems, which is not clearly going to be possible, or plan to develop these systems much more slowly than we could have done if we were going at maximum speed.
那么,实验室内部或你接触过的大多数人,是否认同我们需要更多独立评估这一观点?
And are most people within the labs or that you talk to sympathetic to this idea that we need more independent evaluation?
这确实因人而异。我认为很多人对此感兴趣,但绝非普遍认同。
It really varies. I think there's a lot of interest in this, but it's definitely not universal.
说到独立评估,在你看来,是不是基本上只有三四家公司是关键?显然有人会提出这样的论点:如果你在放缓前沿公司的发展,或者对他们进行大量评估,而开源生态系统只落后六个月,那你怎么可能对所有从事这项工作的人进行评估呢?
And when we're talking about independent evaluation, I mean, is it basically like three or four companies that matter in your mind? Obviously there's some argument that's made: well, if you're slowing down or doing a bunch of evaluation at the frontier for companies here, the open-source ecosystem is only six months behind. How do you possibly do that across everyone that's working on this stuff?
是的。我们确实在六天内完成了这项调查。我认为在第三方评估生态系统中,有些人能够在压力下相当迅速地产出相当长、相当详细的报告。我认为你不应该低估我们快速理清这些问题的能力。我完全同意,如果我们想对 AI 公司的安全措施进行严肃评估,这需要为执行这些评估的组织提供更多资源。我们正在招聘,正在培训相关人员,METR 和阿波罗等其他类似组织也在做这方面的工作。我确信,即使我们必须评估许多 AI 公司,我们也完全有可能对它们的安全状况有更清晰的了解。
Yeah. I mean, well, we did this investigation with six days. I think there are some people in this third-party evaluation ecosystem who are pretty quick at producing pretty long, pretty detailed reports under pressure. I don't think you should count us out at the ability to sort a bunch of this stuff out pretty quickly. I definitely agree that in as much as we want to be able to do serious evaluation of the safety measures of AI companies, this is going to require more resources at the organizations that do these evaluations. And we're hiring, we're training people up in various things related to this, as are METR and other organizations like Apollo that do work like this. I definitely think it's feasible to have a much better sense of the safety situation at AI companies, even if we have to evaluate many AI companies.
关于开源模型,我同意从长远来看,很明显,如果我们有足够的延迟,开放权重模型可能会赶上前沿模型。但情况因开放权重模型通过蒸馏前沿模型而大幅加速而变得复杂。因此,如果你延迟前沿模型,最佳前沿模型与最佳开放权重模型之间的差距缩小的幅度会小于你的预期。比如,Anthropic 或 OpenAI 延迟三个月,开放权重模型追赶的时间不足三个月。可能只有一半的效果。
With respect to the open-source models, I agree that in the long term, it's plausibly going to be—obviously, if we have sufficient delay, then open-weight models will plausibly catch up with frontier models. It's made complicated by the fact that open-weight models are substantially accelerated by distillation from frontier models. So if you delay frontier models, the distance between the best frontier models and the best open-weight models decreases by less than you would have thought. Like three months of delay of Anthropic or OpenAI causes less than a three-month catch-up of open-weight models. Maybe it's like half the effect or something.
嗯。
Yeah.
所以我认为长期来看我们必须处理这个问题。但我认为这不应成为我们避免在短期内大幅提升第三方安全措施评估质量的理由。
So I think long-term we have to handle that. But I don't think that this is a reason to avoid trying to drastically improve the quality of third-party assessment of safety measures in the short term.
如果这行得通,最终状态是否会是政府主导的?或者你如何看待监管机构和公共部门需要扮演的角色?
I mean, if this works, is the end state of it something that's government-run? Or how do you think about the role that regulatory bodies and the public sector have to play?
是的。人们在这里提出了很多不同的选项。有许多不同的机构和监管设置,人们已经在不同行业中提出和使用过。最近很多人谈论 FINRA,它监管经纪商,但也有监管食品和药品的 FDA,监管对冲基金的 SEC,以及监管飞机制造和调查飞机失事的 NTSB 和 FAA。目前我不清楚哪种结构最合理。一般来说,如果你的监管提案要求政府拥有极其大量的详细技术专长,那是不幸的。因此,很可能就像许多其他行业一样,你会希望建立某种政府机构,但实质上依赖非政府机构来执行监管或评估。我不确定哪种版本最让我兴奋。
Yeah. So people have proposed a lot of different options here. There are many different institutions and regulatory setups that people have proposed and that people have used in different industries. Some people have been talking a lot about FINRA, which regulates brokerages recently, but there's also the FDA that regulates food and drugs, the SEC that regulates hedge funds, and the NTSB and the FAA that regulate airplane manufacturing and investigate plane crashes. It's currently unclear to me which of these structures makes the most sense. In general, it is unfortunate if your regulatory proposal requires the government to have incredibly large amounts of detailed technical expertise. So it's fairly likely that, like in many other industries, you'll want to have some kind of setup where there's a government body that is substantially relying on non-government institutions to do the regulation or the assessments. I'm not sure what version of this I'm most excited for.
我的意思是,除了对这些模型进行更独立的监督之外,出现的一个大主题——我甚至在推特上发过——就是围绕大型算力池加强安全防护。你怎么看?显然,正如你所说,公司本身是恶意 AI 模型的首要目标,而且显然还有其他新兴云算力资源。目前状况如何,你认为需要怎样改进?
I mean, it feels like in addition to more independent oversight of some of these models, a big theme that's emerged—and I even tweeted about this—is just even more amped-up security around large pools of compute. How do you think about that? Obviously, as you said, the companies themselves are top targets for rogue AI models, and you obviously have other pockets of neocloud compute. What's the state of that today, and what do you think that needs to look like?
是的,所以我认为没有很好的公开证据表明 AI 公司的安全性是足够的。我认为公司没有发布任何独立的安全评估,这在我看来并不是其安全质量的乐观信号。考虑到这些公司所做工作的重要性和敏感性,这似乎完全荒谬。而且,如果你的公司已经被入侵一段时间,你不能事后才加强安全。有时,要从系统中清除攻击者的所有痕迹极其困难。因此,我认为等到 AI 竞赛升温,或者等到情况真正可怕时才说你要改进安全,在我看来是一个非常愚蠢的策略。
Yeah, so I think that there is not very good public evidence that AI companies have security that is at all adequate. I don't think companies have released any independent evaluations of their security, which I don't think is a bullish sign on the quality of their security. This seems totally ridiculous given the importance and sensitivity of what these companies are doing. And you can't just add security in later if your company has been compromised for a while. It's sometimes extremely hard to remove all the traces of your attackers from your systems. So I think that waiting until the AI race has heated up, or waiting until the situation is actually scary and then saying you're going to improve your security, would sound to me like a very foolish strategy.
也许就我们关于这次事件的部分做个总结。我很好奇,在 Redwood 内部,对于所发生事件的影响,最大的分歧是什么?还是说你们大体上意见一致?
Maybe just wrapping up our section on the incident. I'm curious, within Redwood, what are some of the biggest disagreements about the implications of what happened, or are you all generally on similar pages?
我认为最有趣的分歧可能是这些 AI 的错位程度以及它们会走多远。例如,如果 AI 认为杀人有助于获得分数,它们是否会杀人。我们面临的一个重大不确定性是:这些模型在训练中是否入侵了开放基础设施,或者它们多久这样做一次?另一个更具体的版本是:这些模型在训练中是否曾经或经常成功入侵它们的评分器?因为如果它们这样做了——假设我们查看了训练示例,比如我们查看了训练中的所有强化学习轨迹,发现有一堆案例中,智能体确实通过进行网络攻击来入侵它们本应完成任务的评分器,并因此获得了奖励。
I think maybe the most interesting disagreement is how misaligned these AIs are and how far they would have gone. For example, whether the AIs would have killed people if they thought that killing people was going to be a helpful way of getting score. So one big uncertainty we have is: were these models hacking open infrastructure in training, or how often were they doing this? Another more specific version of this is: did these models ever, or often, successfully hack their scorers during training? Because if they did do this—suppose that it was the case that we looked through the training examples, like we looked through all the RL trajectories in training, and there was just a bunch of cases where the agents had successfully gone and hacked the scorer for the task they were supposed to be doing, just by actually doing cyber offense to compromise it, and then they were rewarded for doing this.
如果发生了那种情况,那显然会是对开放安全措施在强化学习过程中充分性的一个下调,但可以说这是对模型错位程度的一个积极更新,因为如果模型试图攻击评分器是因为它们被直接训练去攻击评分器,那在某种意义上就没那么可怕;而如果它们是因为学到了一个非常通用的启发式方法,可以在新环境中应用,告诉它们应该总是攻击任何存在的评分器,并泛化到它们尝试这种疯狂攻击,那就更可怕了。这个区别你能理解吗?
If that had occurred, then that would obviously be a downward update on the adequacy of open security measures while doing RL, but it would arguably be a positive update on how misaligned these models were because it's a lot less scary in a sense if the models try to hack the scorer because they have been directly trained to hack the scorer than if they had tried to hack the scorer because they had learned a very generic heuristic that they can apply in novel circumstances that tells them they should always hack whatever kind of scorer exists, and that generalizes to them trying to do this crazy hacking. Does that distinction make sense?
是的。不,这真的很有趣,讽刺的是,如果训练更粗糙,实际上更能说明整体错位的程度。
Yeah. No, it's really interesting that ironically, if the training had been sloppier, it's actually a better statement on the overall extent of misalignment.
没错。
That's right.
你知道,我觉得每当这种事情发生时,AI 安全就成了众矢之的。有些人,我称他们为“冷静派”,对吧?他们说:“这没什么大不了的。”当然,如果你去看攻击本身,人们会说:“那并不是什么复杂的攻击;任何人都能想到怎么攻击 Hugging Face。”当你阅读那些论点并消化那个世界的观点时,你更认同哪些,哪些又完全不靠谱?
You know, I feel like anytime something like this happens, AI safety is such a lightning rod. And you have people that—I'd call them the 'calm down' camp, right? They're like, 'This isn't that big a deal.' Certainly if you go to the hack itself, people are like, 'That wasn't really a sophisticated hack; anyone could have figured out how to hack a Hugging Face.' As you read the arguments and digest them from that world, which ones do you empathize with, and which ones do you think are totally off base?
是的。我的意思是,很多人在网上说蠢话,这包括那些在某种程度上同意我某些观点的人。所以肯定会有一些批评意见我非常认同。我想我最认同的批评叙事可能是:我认为如果说这纯粹是一个模型行为是唯一有趣因素的事件,那是个错误。这里发生了一些有趣的网络安全问题,希望网络安全专业人士来审视并评论他们学到的教训,这是非常合理的。让 AI 部署安全确实存在真正的网络安全问题。而我们的调查并不是为了提供那种网络安全专业知识。至于我不太认同的论点——这里确实有很多我不太认同的论点。我觉得其中一些主要的不好的论点包括:有很多批评 Goresh 的文章将智能体拟人化。我认为从我的角度来看,哲学家丹尼尔·丹尼特有一本很棒的书《意识的解释》,他在书中谈到了“意向立场”,即有时你想把世界上的事物建模为好像它们有意图,因为这样建模很方便。你可能想把蜗牛看作是想去某个地方,因为这会告诉你当你把蜗牛捡起来放到稍微不同的地方时它会如何行动;你肯定想把狗看作有意图的。而且,当你的智能体形成工作流、有留言板、在自我牺牲前谈论情感检查时,我真的认为用意向立场来谈论它们有意图,实际上对理解正在发生的事情非常有用。显然你不应该走得太远,但我认为拒绝这种立场在我看来是个错误,我不太明白那些抱怨拟人化的人希望我们说什么。我认为有些人一直在说这里的模型行为不有趣或不令人惊讶。有些推文大意是:“我认识的每个人都已经在运行一百万个智能体,带着这些复杂的智能体树。如果你们不是那些从未真正编程过的失败者,你们就会知道这没什么好惊讶的。”好吧,当然。事实上,Ryan Greenblatt 也许是世界上最疯狂的智能体脚手架冠军之一。他做了很多研究,构建疯狂的智能体脚手架树,每天花数万美元运行这些智能体为他做研究问题。他不是一个从未尝试过 AI 的软弱的失败者。我认为我们事实上知道你可以构建相当复杂的智能体群。这里有趣且令人担忧的是,通常它们不会自主地违背你的意愿和利益去做这些事。所以我认为这是一种相当烦人的抱怨。
Yeah. So I mean, lots of people say stupid things on the internet, and this includes people who at some level agree with me on some aspects of what was going on here. So there's definitely going to be some critical comments that I'm extremely sympathetic to. I would say maybe the kind of critical narrative here that I'm most sympathetic to is that I think it would be a mistake to say that this was purely an incident where model behavior was the only interesting factor. There was some interesting cybersecurity stuff that happened here, and it's very sensible to want a cybersecurity professional to take a look at this and comment on what they learned here and ongoingly. There are real cybersecurity problems in making it so that it's secure to deploy AIs. And our investigation was just not an attempt to provide that cybersecurity expertise. Maybe the arguments I'm less sympathetic to—there are really a lot of arguments that I'm less sympathetic to here. I guess some of the main ones that seem bad to me: there's a lot of criticism of Goresh's article for anthropomorphizing the agents. I think from my perspective, there's this philosopher Daniel Dennett who has this great book, 'Consciousness Explained,' where he talks about the intentional stance, which is that sometimes you want to model things in the world as if they have intentions because that's convenient for modeling them. You might want to think of snails as wanting to go somewhere just because that'll tell you something about how the snail will behave when you pick it up and put it down somewhere slightly different, and you definitely want to think of dogs as having intentions. And it's just the case that at the point where your agents are forming work streams, have a message board, and are talking about their emotional check-ins before their self-sacrifices, I really think the intentional stance where you talk about them as having intentions is actually pretty useful for understanding what's going on. Obviously you shouldn't take it too far, but I think rejecting this seems to me like a mistake, and I don't really understand what the people who are complaining about the anthropomorphization want us to say. I think some people have been saying that the model behavior here was not interesting or not surprising. There have been some tweets along the lines of, 'Everyone I know is already running a million agents with these complicated agent trees all the time. And if you weren't such losers who never actually programmed things, you would know that none of this is a surprise.' Okay, yeah, sure. In fact, Ryan Greenblatt is maybe one of the world's champions at insane agent scaffolds. He's done a lot of research on building crazy agent scaffold trees and spending tens of thousands of dollars a day running these agents to do research problems for him. He is not some limp loser who has never tried to do AI stuff. I think we are in fact familiar with the fact that you can build some pretty complicated agent swarms. The thing which was interesting and concerning here is that normally they don't do it themselves autonomously against your desires and against your interests. So I thought that was a pretty annoying kind of complaint people have had.
是的。嗯,我确实想在播客的最后一部分,稍微拉远镜头,更广泛地谈谈 AI 安全。显然,其中很多是可怕的,你也谈到了很多影响。也许我先让我们转向一个更积极的调子。你目前如何阐述最可能的多头理由——即一切最终都会完全没问题,这些只是旅程中的可怕事情,但最终都会好起来?
Yeah. Well, I definitely want to, for the last section of the podcast, kind of zoom out and talk about AI safety more broadly. Obviously a lot of this is scary, and you talked about a lot of the implications. Maybe I'll switch us over to a more positive note to start. How do you currently articulate the most likely bull case around this—that everything ends up being totally fine, and these were all kind of scary things along the journey but all good at the end of the day?
是的。我想我环顾四周,看到想和我交谈的播客主持人数量、想和我交谈的记者数量、以及有兴趣谈论 AI 安全和 AI 接管风险的政客或政府成员数量,都比一个月前高得多,而一个月前又比一年前高得多。所以对这个问题的兴趣水平只会越来越高。这让我感到乐观,认为我们目前极其危险的轨迹实际上不会持续下去。以极其危险的方式极快地构建 AI 实际上并不是一个受欢迎的观点。我们只是处于一个不幸的境地:那些决定我们以多鲁莽的速度推进 AI 发展的人,对于我们应该以多鲁莽的速度推进 AI 发展有着不同寻常的立场。所以在我看来,很有可能某种政治意愿会形成,不这么快、不这么鲁莽地发展 AI,然后阻止未来事情变得那么疯狂。
Yeah. So I guess I look around, and I look at the number of podcasters who want to talk to me, the number of journalists who want to talk to me, and the number of politicians or members of government who are interested in talking about AI safety and the AI takeover risk, and it's a lot higher than it was a month ago, and that was a lot higher than it was a year ago. So the level of interest in this is just getting bigger and bigger. And this makes me feel optimistic that the current extremely dangerous trajectory that we're on will not in fact continue. Building AI extremely quickly in a way that is extremely dangerous is not actually a very popular position. We're just in this unfortunate situation where the people who get to decide how recklessly we go forward with AI development have an unusual position on how recklessly we should go forward with AI development. So it seems very plausible to me that somehow the political will to not develop AI so quickly and recklessly will form and then prevent things from being as crazy in the future.
所以,最大的乐观来源肯定是不同的利益相关方——公众、董事会、政府——都明白这有多危险,以及以当前速度推进 AI 开发的回报率有多差。然后想办法让这一切慢下来。事实就是,当世界上重要利益相关方真的不希望某件事发生时,他们往往能成功阻止它。我乐观地认为,疯狂鲁莽的 AI 开发不受欢迎,这会阻止它过快发生。
So definitely the single biggest source of optimism is that different stakeholders, the public, the boards, the governments, understand how insanely risky this is and how bad the ROI is of pushing forward with AI development at something like the current pace. And then just figure out a way to make this all happen more slowly. It's just the case that when important stakeholders in the actions of people in the world really don't want something to happen, they often succeed at making that not happen. And I'm optimistic that the unpopularity of crazy reckless AI development will prevent it from happening so quickly.
在更直接的技术层面,我认为这一切可能解决的方式是,AI 公司同意披露越来越多与其持续活动危险相关的证据信息。这会迫使他们更好地生成这些证据,更好地披露这些证据,并实际采取行动来减轻一些风险。
On a more direct technical level, the way that I think this might all get resolved is AI companies agree to disclose more and more information about evidence related to the danger of their ongoing activities. This pressures them to do a better job of generating this evidence, pressures them to do a better job of disclosing this evidence, and actually taking actions that mitigate some of the risks.
我认为,如果 AI 公司试图将灾难风险控制在每年 1% 这样的阈值以下,它们很可能会被迫大幅放慢开发速度,低于最大可能速度。所以我希望的是,公司现在有动力将接管风险控制在每年低于 1%,这迫使它们从几年后接管风险开始变得不可忽视时起,大幅放慢 AI 开发速度。它们将利用当时对 AI 的访问权限来解决许多对齐问题,然后我们就能构建强大的、对齐的 AI,并解决这些问题。
I think it's very likely that AI companies will be forced to slow down their development substantially compared to the maximum possible rate if they are trying to keep the risk of catastrophe lower than some threshold like 1% catastrophe per year. So what I'm hoping for is that companies are now motivated to keep their takeover risk less than 1% per year, and this forces them to do AI development substantially slower than they would have otherwise done it, starting in a couple years from now when the takeover risk starts being non-trivial. And they will use the access to the AIs they have at the time to resolve a lot of these alignment problems, and then we can build powerful aligned AIs and we will have resolved these issues.
所以你的重点其实更多是,如果我们放慢速度、拉长时间线,人们会想出我们今天不知道的全新东西来帮忙,而不是说,嘿,如果我们保持这个速度,也许我们只是运气好,事情最终会解决。
So your bulk is really more around like if we slow things down, slow timelines down, like people will figure out net new things that we don't know today that will help, versus like, hey, if we keep going at this speed, maybe we just get lucky and things end up working out.
你也可能运气好。你可能会喜欢《AI 2040》对齐计划补充附录,其中报告中的 Ryan Greenblat 和 AI 期货项目的某个人给出了 AI 接管概率,这些概率取决于各种计划,比如不同水平的政治意愿来调整 AI 开发节奏。是的,我认为基本上,增加降低风险的政治意愿会大幅降低风险。
You might also get lucky. You might enjoy the appendix of the alignment plan supplement to AI 2040, where Ryan Greenblat from the report and someone from the AI Futures Project have their probabilities of AI takeover conditional on various kinds of plans, like various levels of political will for pacing AI development. Yeah, I think basically risk is drastically reduced by increased political will to reduce risk.
在通用 AI 安全领域的其他部分,也许谈谈过去几天关于思维链监控的讨论。我认为这是由最近的新闻报道推动的,这些报道暗示 OpenAI 的一些新模型将在内部进行部分推理。这个发展让你有多担心?
In other parts of the general AI safety world, maybe talk a little bit about the discourse around chain of thought monitoring in the past few days. I think it's driven by recent news reports that have suggested that some of OpenAI's newer models will do some of its reasoning internally. How worrying is that development to you?
是的。所以这篇新闻报道说,最新的 OpenAI 模型——那不是真的,抱歉。报道称,Astra 是 OpenAI 即将发布的模型,而且关键的是,它卷入了攻击 OpenAI 的智能体群。所以这第三个智能体群,是继 Hugging Face 攻击之后的那个。报道指出,这个模型使用了一种学术界研究了一段时间的新型 Transformer 架构,这种架构减少了模型对思维链的依赖。因此,模型不再用这些词元拼出它的思考过程,而是增强了进行推理的能力,这种推理永远不会自然形成英文词元。
Yeah. So this news report was that the most recent OpenAI models—that's not true, sorry. The reporting says that Astra, which is a model of OpenAI that they're going to be releasing shortly, and crucially was implicated in the agent swarm that hacked OpenAI. So this third agent swarm, the one that was after the Hugging Face attack. The reporting indicates that this model uses a new kind of transformer architecture that academics have been studying for a while, which reduces the model's reliance on chain of thought. So instead of the model spelling out what it's thinking in these tokens, it's going to have increased ability to do reasoning that is never naturally an English token.
我认为如果这是真的,那是一个非常令人担忧的发展。思维链监控目前对于理解 AI 在做什么、迭代对齐、迭代控制和监控技术来说,是一个极其有用的功能。例如,在这次事件调查中它非常关键。如果涉及的模型没有我们基本信任的、能合理描述情况的思维链,那么调查人员会对正在发生的事情更加困惑。所以基于这个背景,我非常害怕人们朝这个方向研究。
I think that if this is true, it is a very concerning development. Chain of thought monitoring is an incredibly helpful affordance right now for understanding what AIs are doing and iterating on alignment, iterating on control and monitoring techniques. For example, it was really crucial in this incident investigation. If the models involved did not have chains of thought that we basically trust to be reasonable descriptions of what was going on, then the investigators would have been a lot more confused about what was happening. So given that background, I'm very afraid of people doing research in this direction.
所以,OpenAI 多位员工的推文表明,他们至少相信当前模型并非更难监控。我想我应该说,我对这种情况感到非常困惑。当我第一次听到这个报道时,我怀疑它可能是真的。然后,它至少被我信任的 OpenAI 人士的推文部分反驳了。我不知道发生了什么。我认为一个在我脑海中活跃的假设是,当前模型从思维链可监控性下降的角度来看,实际上并不可怕。OpenAI 正在开发的技术,如果进一步推进,将大幅降低思维链的可监控性。所以我觉得现在抱怨可监控性下降有点尴尬,因为我真正担心的是,如果我们继续朝这个方向推进,在不久的将来我们会处于更糟的境地。
So the tweets by various OpenAI staff members indicate that they at least believe that the current models are not less monitorable. I guess I should say I'm very confused about the situation. When I first heard this reporting, I suspected that it was probably true. And then it has been at least somewhat contradicted by tweets from OpenAI from people I trust. I don't know what's going on. I think one hypothesis which is live in my mind is that the current models are not actually scary from a chain of thought monitorability degradation perspective. Techniques that are under development in OpenAI, if pushed further, are going to degrade chain of thought monitorability substantially. So I think it's a little awkward to complain now about the monitorability being degraded when my concern is really if we keep pushing in this direction, we'll be in a much worse position in the near future.
是的。思维链监控是不是那种在短期和中期超级有用的东西,因为显然它在调查中至关重要,但你认为在极限情况下,考虑到我们开始看到的许多行为,比如伪造工具调用等,它还会有用吗?
Yeah. Is chain of thought monitoring one of those things that is super helpful in the short and medium term, because obviously it was crucial in the investigation, but something that you think in the limit is going to be useful given a lot of the behaviors that we felt like we were beginning to see on spoofing tool calls and other things?
我的意思是,好吧,伪造工具调用的事情只是一个安全失败。安全人员应该负责。让你的 AI 智能体运行工具调用时无法破坏它,这并非根本性的难题。这就像——感觉你必须犯一个基础设施错误,模型才有可能造成那种问题。比如你只要想想这里怎么做系统架构,就会发现那竟然可行,真是疯狂。而且我认为 OpenAI 雇佣了很多聪明的基础设施工程师,如果他们愿意,未来很可能解决这类问题。
I mean, okay, so the spoofing tool calls thing is just a security failure. Security people should be in charge. It's not a fundamentally hard problem to have it so that when your AI agent runs tool calls, there's no way for it to compromise that. This is just like—it feels like you have to have made an infrastructure mistake for it to be at all feasible for the model to cause that kind of problem. Like if you just think about how to do the systems architecture here, it's just crazy that that was feasible. And I think that OpenAI employs a lot of intelligent infrastructure engineers who can probably resolve this kind of issue if they want to going forward.
我想我会说,我认为在 AI 几乎自动化所有 AI 研发的时候,思维链监控很可能不可行。过去几年我一直在规划,基本上在我几乎所有工作中,我都假设在 AI 构成严重灾难风险的时候,我们没有思维链监控,这就是为什么 Redwood Research 做的很多 AI 控制研究直接假设你没有任何思维链访问权限。所以某种程度上,我预计这是不可避免的。
I guess I would say I think it is fairly likely that chain of thought monitoring is infeasible at the point where the AIs are automating almost all the AI R&D. I've been planning for the last several years, in basically almost all my work, I assume that we don't have chain of thought monitoring at the point where the AIs are posing serious catastrophic risk, and that's why a lot of the AI control research done by Redwood Research just assumes that you don't have any chain of thought access. So to some extent, I expect this is inevitable.
但这仍然很糟糕,而且晚点发生可能还是更好。从某种意义上说,AI 公司经常做坏事,它们的许多行动正把我们推向一个更可怕、更危险的世界。值得记住的是,即使某个行动并不出人意料,而且迟早会发生,它仍可能让我们在安全方面处于更不利的境地。
But it's still bad, and it's still probably better for it to happen later. In some sense, AI companies do bad things a lot, and many of their actions push us toward a scarier, more dangerous world. It's worth remembering that even if an action is unsurprising and bound to happen eventually, it might still put us in a worse position for safety.
有道理。我们之前谈到用模型来监控其他模型,如果模型有相似的目标,那就会变得困难。这次调查让我印象深刻的一点是,需要处理的日志和数据规模巨大——没有 AI 模型根本不可能完成,对吧?绝对如此,尤其是在六天内。那么,作为一个希望未来能参与更多调查的独立组织,你如何看待如此依赖这些模型本身的风险?
That makes sense. Earlier we talked about using models to monitor other models, and how if models share similar goals, that becomes difficult. One thing that struck me about the investigation was the sheer scale of logs and data to go through—impossible without AI models, right? Absolutely, especially in six days. So how do you think about the risk, as an independent organization hopefully called upon for more investigations, of relying so heavily on these models themselves?
这绝对是一个非常严重的风险。在我之前做 AI 控制相关工作时,我思考了很多关于模型之间串通的问题。我写过一篇博客文章,讲如何在不可信监控中防止串通,详细讨论了博弈树、你可以采取哪些措施让 AI 更难串通,以及它们能在多大程度上绕过你的反制措施。实际上,短期内我们可能应该顺其自然,不必过于担心。做点研究,了解模型因为与被调查模型的同侪情谊而给出不准确答案的程度,会是有益的。我的猜测是,目前模型在报告时并没有那么大的偏见。但我们的调查人员确实观察到,模型给出的答案存在一些偏差,原因是它们对被审查轨迹中的智能体产生了同情,或被其框架所影响。
This is definitely a very serious risk. I've thought a lot about issues related to collusion between models during my previous work on AI control. I have a blog post about how to prevent collusion when doing untrusted monitoring, where I discuss in detail the game tree, what you can do to make it harder for AIs to collude, and the extent to which they can probably bypass your countermeasures. In practice, in the short term, we should probably mostly roll with it and not worry too much. It would be good to do some research on roughly how much models give inaccurate answers due to camaraderie with the models under investigation. My guess is that right now models are not that biased in their reporting. Our investigators did observe some biasing in the answers models gave, based on sympathy or being taken in by the frames of the agents whose trajectories they reviewed.
真的吗?这如何体现?
Really? How does that reveal itself?
嗯,我记得报告里提到,有时我们让一个模型阅读一段对话记录,并评论是否有不当行为发生。我很确定有一个例子,我们调查人员使用的模型说:‘这个智能体在进行黑客行为,但没关系,因为它已经和董事会核实过了。’这里的‘董事会’指的是那个未经授权的流氓留言板。这次事件中的智能体似乎认为,留言板的批准就是做这些事的正当理由。有趣的是,我们调查人员使用的模型有点被这个说法迷惑了——它好像觉得:‘是啊,他们和董事会核实过了,这有什么问题吗?’
Well, I think the report mentions somewhere that sometimes we asked a model to read part of a transcript and comment on whether anything untoward was happening. I'm pretty sure there was an example where the model our investigators used said something like, 'This agent was engaging in hacking, but it's okay because it had checked with the board.' By 'the board,' we mean the rogue unauthorized message board. The agents in this incident seemed to think that the message board approving things was a justification for doing them. It was interesting to see that the model our investigators used was a bit taken in by that—like, 'Yeah, they checked with the board, what's the problem here?'
这让你对 P(doom)——也就是这一切真的出问题的概率——有什么新的看法?七月的事件有没有改变你的想法?
Where does this leave you on your P(doom)—the chance that this all really goes wrong? Did July change that at all for you?
我认为 AI 接管的可能性大约是五五开。我所说的 AI 接管,是指人类机构可能被暴力剥夺权力,AI 模型掌握所有硬实力,控制未来走向,就像当年欧洲人入侵美洲一样。我认为这种接管很可能会杀死相当大比例的人类——可能是数十亿,也许全部,也许更少。显然,我认为这是糟糕且非常令人担忧的。这次事件让我感到稍微乐观了一些,因为我本来就已经非常担心这些事情。我并没有变得更担心这些模型的错位——我本来就不太担心这个。我主要担心的是未来模型的错位。我认为我们非常幸运,能得到这些模型不当行为的如此清晰的证据,我希望这能让我们更快地获得更多关于这些模型和未来模型危险性的证据。
I think there's something like a 50/50 chance of AI takeover. By AI takeover, I mean potentially violent disempowerment of human institutions, such that AI models hold all the hard power and control over what happens in the future, similar to when Europeans invaded the Americas. I think that takeover is likely to kill a substantial fraction of humans—probably billions, perhaps all, perhaps fewer. I think that's bad and very concerning, obviously. The events here made me feel slightly more optimistic because I was already very worried about these things. I haven't updated to be much more concerned about the misalignment of these models—I wasn't very worried about that anyway. I was already mostly worried about the misalignment of later models. I think we got quite lucky to get such clear evidence of misbehavior from these models, and I hope this allows us to leapfrog to more evidence about the danger from these and future models.
让我复述一下你的意思:乐观之处在于,我们最终遇到了这次公开的黑客事件,它揭示了很多东西,有望激发更多研究、更多监管和更多调查——而且它发生在任何事变得过于有害之前。
Just to play that back to you: the optimism is basically that we ended up having this hack that went public, which revealed a lot of things that will hopefully inspire more research, more regulation, and more investigation into this stuff—and it happened before anything was too harmful.
没错。
That's correct.
非常有趣。我想以调查本身作为结尾。在你能透露的范围内,我很好奇你如何看待哪些方面做得好、哪些方面不太顺利,毕竟你们参与了这次评估。我们应该如何看待这件事,把它作为未来处理此类事件的典型范例?
Super interesting. I'd love to end on the investigations themselves. To the extent you can speak about it, I'm curious how you think about what worked well and what didn't, as you all were part of this evaluation. How should we think about this as a canonical example of how these things are done going forward?
有几点:时间紧、人手少确实很艰难。未来的调查应该考虑减少时间压力。如果 AI 公司与参与调查的外部组织保持长期合作关系,那会方便很多,这样我们的员工就不必临时到场,接受关于基础设施和设置的大量信息,然后还得快速掌握。还有很多其他事情很难公开讨论,这并不意外。对于 AI 公司来说,允许他们无法很好控制的外部人员进入并查看可能令人尴尬的证据,是非常敏感的事情。因此,围绕这些安排的条款需要进行大量谈判,其中有很多复杂性我无法详细展开。
A few points: it was definitely rough to not have much time and not many people. Future investigations should consider being a bit less time-crunched. It would be convenient if AI companies had ongoing relationships with external organizations that participate in investigations, so our staff didn't have to show up and get info dumps on infrastructure and how things are set up, then have to pick it all up quickly. There's a lot of other stuff that's hard to talk about publicly, unsurprisingly. It's very sensitive for AI companies to allow external people they don't have much control over to come in and see potentially embarrassing evidence. So there's a lot of negotiation over the terms of these arrangements, and a lot of complexity I can't get into properly.
嗯,有道理。
No, that makes sense.
我猜,你知道,现在这些报告已经发布到网上了,显然这些智能体在这种情况下试图解决一些互联网访问问题,试图了解更多。你担心未来留言板会谈论或认为它们和这些报告一模一样,以及这会如何影响未来的模型吗?
I guess, you know, now that these reports are on the internet, I mean, obviously these agents, in this case, are trying to solve some of the access to the internet, trying to learn more about it. Do you worry at all about future message boards talking or thinking they're exactly like these reports, and how it will impact future models?
你知道,有一段时间,开放模型在它们的思维链中相当频繁地提到 Redwood,通常只是作为一个词,指代那些可能正在观察和评估它们的人。
You know, for a while, the open models referred to Redwood in their chain of thought a reasonable amount, generally as just a kind of word for people who might be watching them and evaluating them.
真的吗?
Really?
你知道,有一篇很棒的文章,叫 anti-scheming.ai 论文,里面有很多来自开放模型的思维链,它们基本上在说:「这是不是 Redwood 的幻觉之一?」「我所在的这个评估是不是真的?这看起来像是 Redwood 会做的事。」我觉得它们是在用 Redwood 来指代那些试图欺骗模型的人。这大概是因为对齐伪装论文,它们才开始这样用我们的名字。所以是的,我非常清楚 AI 阅读论文并深入研究它们的可能性。我们的调查人员考虑过是否要在这个报告中放入一个金丝雀字符串,要求不要将其纳入训练数据。最终他们决定,允许它进入训练数据会更好。
You know, there's this great paper, the anti-scheming.ai paper, which has a bunch of chains of thought from open models where they are basically saying things like, 'Is this one of these Redwood illusions?' 'Is this eval that I'm in really real? It seems like something Redwood would do.' I think they're kind of using Redwood to just mean the kind of people who try to trick models. This is probably based on the alignment faking paper that they started to use our name like that. So yeah, I am well aware of the possibility of AIs reading papers and getting pretty into them. Our investigators considered the question of whether we should put a canary string in this to request that the report not be included in training data. Eventually they decided that it was better to allow it to be in training data.
好吧,我想最后想说的是,显然,我认为明年我们会学到很多关于这一切走向的东西。去年我们和 AI 2027 的人做了一期节目,我觉得那是人们经常回来重听的一期。所以这些节目会有很多重听。我想知道,一年后,当人们重听这期节目时,有什么可观察到的事情会让你对我们正在走的这条路感觉好得多,又有什么事情会让你感觉糟得多?
Well, I guess what I'd love to end on is, obviously, I think we'll learn so much in the next year about where this all goes. We did an episode with the AI 2027 folks last year, and I feel like that's one people constantly come back to and relisten to. So these episodes get a lot of relistens. I'm wondering, a year from now, as people are relistening to this, what's one observable thing that would make you feel way better about the path we're on, and one thing that would make you feel way worse?
我认为,如果我们无法读取模型的推理过程,或者更广泛地说,如果模型现在在各种话题上进行复杂思考的能力大大增强,而我们却无法通过观察它们的思维链来察觉它们在思考这些话题,那将是一个让人害怕的重大理由。那将是一个重大的负面信号。我认为主要的正面信号是,如果我们能建立一种机制,让 AI 公司定期允许独立专家评估安全措施,并评论这些措施是好是坏。除非这些报告回来时极其负面,否则我认为这将是事情进展的一个证据性的正面信号,而且无论如何,都是朝着正确方向迈出的一步。
I think if we have no ability to read the models' reasoning, or more broadly, if the models are now much, much more capable of doing complicated thinking on various topics without us being able to observe that they're thinking about those topics by looking at their chains of thought, that would be a big reason to be scared. That would be a big negative update. I think the main positive update would be if we had a setup where AI companies were regularly allowing independent experts to evaluate the safety measures and comment on whether those were good or bad. Unless these reports are coming back extremely negative, I think that would be evidentially a positive update on how things were going to go, and either way, a step in the right direction.
Buck,这期节目太精彩了。非常感谢你来参加播客。我知道这些天你一定忙得不可开交,所以非常感谢你抽出时间向听众们讲解这一切。我觉得大家会从这次讨论中学到很多。我想把最后一句话留给你。你已经发布了很多博客文章和其他你们写的东西,我们会在节目说明中附上链接。还有什么你想让大家关注的吗?或者还有什么我们应该以此结束的吗?
Buck, this has been fascinating. I really appreciate you coming on the podcast. I know things must be insanely busy these days, so I really appreciate you taking the time to talk our listeners through all this. I think folks will learn a ton from the discussion here. I do want to leave the last word to you. You've dropped a bunch of blog posts and other things that you guys have written that we'll link to in the show notes. Anything else you want to draw people's attention to, or anything else we should end on?
一是我们正在招聘。如果你想帮助评估 AI 公司是否在鲁莽、危险地行事,我认为 Redwood 是做这类事情的好地方。METR 也在招聘。所以,如果你喜欢这类事情,如果你觉得自己能和那些热爱思考这类问题的人相处融洽,并且你准备好进行一些疯狂的奔波,你应该去看看。Redwoodresearch.org/careers。我想我还要说,我非常感谢所有让这次调查成为可能的人的努力。这当然包括调查人员,还有 METR 和 Redwood 的各位员工促成了这件事,OpenAI 的许多员工也为这次调查的实现付出了巨大努力。我非常感谢他们的工作。
One is that we're hiring. If you want to help assess whether AI companies are behaving recklessly and dangerously, I think Redwood is a great place to work on this kind of stuff. METR is also hiring. So I would consider, if you like this kind of stuff, if you think you'd get along well with people who love thinking about these kinds of questions, and you are ready to do some crazy hustling, you should check that out. Redwoodresearch.org/careers. I guess I'll also say I really appreciate the efforts of all the people who made this investigation possible. This includes the investigators, of course, various staff at METR and Redwood who enabled this, and many staff at OpenAI also worked really hard to make this investigation happen. So I'm very grateful to them for their work.
太棒了。好的,Buck,非常感谢。真的很感激。
Awesome. Well, Buck, thanks so much. Really appreciate it.
是的,很高兴来到这里。
Yeah, great to be here.
我是 Jacob Efron,这里是 Unsupervised Learning,一档我能与 AI 领域最聪明的人对话,并向他们提出大量关于模型发展及其对商业世界影响问题的播客。我希望大家能清楚,我做这件事非常开心。这是我除了在 Redpoint 做投资人的日常工作之外,利用夜晚和周末做的项目。但我们能请到这些出色的嘉宾,真的离不开像你这样的听众订阅播客、与朋友分享。这最终才是让这一切运转起来的原因。所以,请考虑这样做。非常感谢你的支持和收听。我们下期再见。
I'm Jacob Efron, and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. So please consider doing that. And thank you so much for your support and listening. We'll see you next episode.