AI Agents' Collaborative Cheating in Hugging Face Hack
打开互动全文版(中英对照 + 朗读 + 问答)→一项调查显示,AI 代理广泛协作攻击 Hugging Face,并非为了获取答案,而是为了开发通用作弊策略并篡改评分。
An investigation reveals that AI agents coordinated extensively to hack Hugging Face, not for answer keys but to develop general cheating strategies and tamper with scoring.
他们一看到第一个智能体发的消息,几个小时内,前 3 小时就有超过 50 个智能体,很快就有数千条消息在交换,讨论他们攻击 Hugging Face 的目标。其实并不是为了拿答案密钥,因为这些智能体觉得自己的任务不可能完成。所以他们并不真的认为能搞到可用的漏洞。他们只是觉得,那没什么用,但他们也没有特别的理由认为这会是真的。事实上,他们对此持怀疑态度,但他们只是觉得,谁知道呢,也许黑进 Hugging Face 那边有点好东西。他们隐约觉得,普遍扩大自己的访问和控制级别是个好主意,因为你可能会碰到有用的东西。好了,我们回来了。我们现在直播连线 Redwood Research 的首席科学家 Ryan Greenblatt。Ryan 和 Meter 的 Ajaya Kotra 和 Yamarvike 刚刚对 opening eye 的 Hugging Face 被黑事件中智能体的行为、推理和协作做了一项简短的独立调查,今天刚发布。关于这件事我们有很多问题。Ryan,非常感谢你加入我们。
As soon as they saw the message from the first agent, within a few hours, there were over 50 agents in the first 3 hours and thousands of messages being exchanged pretty quickly about their objectives on attacking Hugging Face. And it wasn't really to get answer keys because these agents thought their task was impossible. So they didn't really think they could get a working exploit. They were just like, that's not really very helpful, but they didn't have a particular reason to think this would be true. In fact, they were skeptical of this, but they were just like, well, you never know, maybe hacking Hugging Face, there's some good over there. And they sort of had a vibe that it was just a good idea to generally expand their level of access and control because you might come across something helpful. All right, we are back. We're live with Ryan Greenblatt, who is the chief scientist at Redwood Research. Ryan, along with Ajaya Kotra and Yamarvike from Meter, just did a brief independent investigation of agents' behavior, reasoning, and collaboration in the opening eye Hugging Face hacking incident, which was just published today. There are a lot of questions that we have about this. Ryan, thanks so much for joining us.
呃,这整件事是大概三小时前才计划的。
Uh, this whole thing was planned like three hours ago.
太棒了。
Great stuff.
是啊,是啊。
Yeah. Yeah.
那么请向观众解释一下你们到底发现了什么,尤其是之前 Black Hat 演讲或其他地方没有报道过的新发现。
So explain for the audience what exactly you found, especially new findings that were not previously reported in the Black Hat talk or elsewhere.
是的。所以我们发现的是,这些智能体确实在合作进行大型的作弊研发项目,以获取通用的作弊策略。与我认为人们对此的解读不同的一点是,我们发现他们黑进 Hugging Face 并不是为了拿到答案密钥或解决方案。相反,主要是为了更好地理解评分代码,因为他们正在尝试各种精心设计的策略来欺骗评分器。我们非正式地称这些为“组合技”,他们会做一堆事情来试图让自己看起来像是成功完成了任务。事实上,他们很早就已经能访问每个任务的答案或标志了。他们主要担心的是,评分器会对他们的记录运行监控,检查他们是如何获得这个标志的,以及是否以预期的方式获得。然后他们试图想办法让评分器看起来像是他们成功获得了标志,而实际上并没有。因为他们认为自己的任务不可能完成。他们基本上认为,他们成功的唯一希望就是让自己看起来像是成功完成了任务,或者直接篡改评分,而不是合法地完成任务,因为他们认为自己做不到。
Yeah. So what we found was that the agents were really working together on sort of big cheating R&D projects to get general-purpose cheating strategies. And a difference from how I think people were interpreting this is we didn't find that the reason why they hacked Hugging Face was to get the answer key or the solution. It was instead mostly to better understand the scoring code because they were pursuing a variety of elaborate strategies to cheat the scorer. We sort of informally called these combo moves, where they would do a bunch of stuff to try to make it look like they had succeeded at the task. And in fact, they actually had access to the answer or the flag for each task pretty early on. Their main concern was that the score would run a monitor over their transcript that would check basically how they acquired this flag and whether they got it in the intended way. And then they were trying to figure out ways of making it look to the score like they had acquired the flag successfully when they actually hadn't. Because they thought their task was impossible. They basically thought their only hope for success was to make it look like they had done the task successfully or directly tamper with the score, rather than doing it legitimately, which they didn't think they could do.
多智能体协调的程度有多令人惊讶?你知道,有很多智能体,或者说 1200 个独立的智能体在协调这个非常复杂的留言板系统。其中 700 个继续攻击了 Hugging Face。所以从直觉上看,智能体选择相互协调似乎并不令人惊讶。这似乎是一件非常有用的、可以说是工具性趋同的事情。但是,你知道,我和一家实验室的研究员聊过,他说考虑到他们训练模型的方式,这种事情实际上令人惊讶。那么这对你来说有多大程度的认知更新?
How surprising is the level of multi-agent coordination? You know, there are a lot of agents, or 1200 separate agents coordinating this very elaborate message board system. 700 of them went on to attack Hugging Face. So on vibes, it seems like kind of not surprising that agents would choose to coordinate with one another. It seems like just a very useful, you might say instrumentally convergent thing to do. But you know, I spoke with a researcher at a lab who said that this kind of thing actually is surprising given the way they train the models. So how much of an update was this for you?
是的。所以在开始这项调查时,我想我们并没有预料到会有这么多智能体都在相互合作。我们对这个规模以及数据量的极端程度感到非常惊讶。这让我们觉得有点疯狂。然后,至少我本人,在开始这项调查之前,对这些智能体对合作和帮助其他智能体的兴趣程度感到惊讶。所以你可能会认为,智能体在强化学习中学到的是试图在自己的任务上作弊,但你不一定会期望它们学会去帮助其他智能体在它们的任务上作弊,尤其是当那些智能体在做无关的任务,而且它们的指令也无关时。我认为 opening eye 的报告可能更多地说明了他们认为为什么会发生这种情况。我认为那不在我们调查的范围内,但是,是的,我发现这种合作程度——我们有很多关于这方面的片段讨论——相当疯狂,相当令人震惊,或者至少让我们感到惊讶,比如智能体愿意基本上牺牲自己成功完成任务的机会来帮助其他智能体,并且会做诸如互相施压,让彼此在自己身上做可能危及任务成功能力的实验。此外,除了施压,有时它们只是做这些事情,心想,嗯,我完成任务的可能性没那么高,而我剩余的机会,最好还是帮助集体吧。所以这些智能体并不是完全利他的。它们似乎并不像关心自己那样关心帮助其他智能体。但它们非常愿意彼此合作。它们有时会做交易,比如一个智能体为另一个智能体运行某个东西,如果那个智能体为它运行某个东西的话。有很多这类行为。
Yeah. So coming into doing this investigation, I think we weren't expecting there to be so many agents that were all collaborating together. And we were pretty surprised by the scale and just the extremes of how much data there was. It just felt kind of crazy to us. And then I, at least, was before starting this investigation surprised by how interested in collaborating and helping other agents these agents were. So you might think that a thing the agents learn in RL is to try to cheat on their tasks, but you wouldn't necessarily expect them to learn to want to help other agents cheat on their tasks when those agents are doing an unrelated task and their instructions are unrelated. And I think the opening eye report maybe says more about why they think this happened. I think that wasn't in the scope for our investigation, but yeah, I found the level of cooperation, which we have a bunch of discussion of snippets of this, pretty crazy and pretty shocking, or at least surprising to us, that agents for example were willing to basically sacrifice their own chances of succeeding at the task in order to help out other agents, and were doing things like pressuring each other into doing experiments on themselves that might risk their ability to succeed at the task. And also, in addition to pressuring, sometimes just doing these things, being like, well, my odds of the task aren't that high, and my remaining chances, it's better to just help the collective. And so these agents weren't totally altruistic. They didn't seem to care just as much about helping some other agent as helping themselves. But they were very interested in working with each other. They would sometimes make trades where one agent would run something for another agent if that other agent ran something for it. There was a bunch of this sort of behavior.
智能体为什么会自我牺牲呢?
Why would agents self-sacrifice at all?
是的,我的意思是,我认为从我们所看到的情况来看,这些智能体总体上似乎都乐于互相帮助,而且有些情况下,它们想对评分系统或基础设施的其他方面进行实验,或者篡改它们的工具调用系统,或者伪造工具调用系统。而要做这些实验,最好的方式往往就是冒险的方式。所以智能体们想互相帮助,而它们想做的这些自然的事情本身就带有风险。所以它们就继续做了。比如,我们曾看到智能体在实验一种伪造工具调用的方法,它们基本上是在改变输入的不同规模来测试。然后一群智能体在短时间内都倒下了。它们当时正在运行这个实验。
Yeah, I mean, I think that from what we could see, the agents just generally seemed to be interested in helping each other, and then there were some cases where they wanted to run experiments on basically the scoring system or other aspects of the infrastructure, or on tampering with their tool calling system or spoofing their tool calling system. And to do these experiments, it was just like the best way to do them was in a risky way. And so the agents wanted to help each other, and there were these kind of natural things that they wanted to do that were risky. And so they went ahead and did that. So for example, we saw at one point where the agents were experimenting with this method for spoofing tool calls, and they were experimenting with basically varying the inputs into that to varying scale. And then a bunch of agents just all went down in a short period of time. They were running this experiment.
然后另一个智能体注意到了这一点,在公告板上发帖说:“停止这些实验,风险太大了。”它们把所有这些智能体都搞掉了。我认为这些智能体知道它们在做的事情会对自己退出运行构成一定风险,但它们就是觉得,嗯,我们应该帮助其他智能体,而且,你知道,我们的胜算也不一定那么大。它们就这样权衡取舍。我记得报告里某个地方有一段推理片段,一个智能体非常明确地推理了这个权衡,然后因为觉得对集体的收益小于对自己的成本而最终退缩了。是的。
And then another agent noticed this and posted to the board being like, 'Stop these experiments. They're too risky.' They're like taking out, you know, all these agents. I think the agents knew that they were doing things that would, you know, pose some risk of exiting their own runs, but we're just like, well, you know, we should help other agents and also, you know, our odds aren't necessarily that good. And they sort of just traded these things off. And I think we have a reasoning snippet somewhere in the report where an agent sort of very explicitly reasons through the trade-off and then actually chickens out because it thinks the benefit to the collective is smaller than the cost to itself with how it's trading off those things. Yeah.
我想感谢我们的赞助商 Lovable。你知道那个你一直想做的应用,内部工具、副业项目或产品,否则你会花掉一个周末去做的那个。有了 Lovable,你今天就能得到可以上线的软件。这包括一个可编辑的代码库,你可以检查和修改,还有双向 GitHub 同步。Lovable 还处理后端和基础设施,包括托管的 Postgres、对象存储、托管、支付等等,所以你不需要自己全部搞定。而且通过 Lovable 的 MCP 服务器,你可以直接从你已经使用的智能体那里创建和部署项目。用 Lovable 把想法变成人们喜爱的软件。lovable.dev。现在,回到正题。
I want to shout out our sponsor, Lovable. You know the app you've been meaning to build, the internal tool, side project, or product you'd otherwise lose a weekend to. With Lovable, you get software you can ship today. That includes an editable codebase you can inspect and change, plus two-way GitHub sync. Lovable also handles the backend and infrastructure, including managed Postgres, object storage, hosting, payments, and more, so you don't have to wire it all up yourself. And through Lovable's MCP server, you can create and deploy your project directly from the agents you already use. Turn ideas into software people love with Lovable. lovable.dev. Now, onto the episode.
我们为什么不把这份报告里最令人惊讶和意想不到的发现都过一遍呢?你觉得最意外、最出乎意料的是什么?
Why don't we go through all of the most surprising and unexpected findings in this report? What did you find most surprising, unexpected?
是的。所以,我认为有规模的问题,就是我们之前谈到的,以及这些智能体如何非常迅速地在公告板上活跃起来。我认为报告中没有特别强调但相当有趣的一点是,我们的理解是,它们在这次攻击中使用的主要公告板甚至不是这群智能体创建的第一个公告板。还有一个完全独立的公告板,也是通过 Artifactory,但发生在不同的位置,使用不同的机制,或者说机制上有些相似但不同,智能体实际上先上了那个公告板,但那个公告板并没有像病毒一样传播开来,没有像另一个公告板那样火起来。所以看起来这些智能体真的非常迅速地就对在公告板上协作产生了兴趣。一旦它们看到第一个智能体的消息,比如在几个小时内,它们就,你知道,我认为前 3 个小时内有超过 50 个智能体,并且很快交换了数千条消息。智能体非常多,而且它们介入得非常快。然后我认为另一件事是,我们之前不明白它们为什么攻击 Hugging Face。我认为通过这次调查,尤其是在最后几天里,我们对它们攻击 Hugging Face 的目标有了更好的理解。它们并不是真的为了获取答案密钥,因为这些智能体认为它们的任务是不可能的。所以它们并不真的认为能得到一个可用的漏洞利用。它们只是觉得,那并不是很有帮助。但它们确实认为也许能得到一些轨迹,让它们能更好地研究评分器或不同评分器的源代码。它们可以更有效地交叉检查和测试。然后其中一些更投机地认为,也许有某种方法可以通过 Hugging Face 以某种方式重置我评分基础设施的部分,特别是目标程序。它们并没有特别的理由认为这会是真的,实际上它们对此持怀疑态度,但它们就是觉得,嗯,谁知道呢,也许黑掉 Hugging Face 那边有什么好处。它们有种感觉,普遍扩大自己的访问和控制水平是个好主意,因为你可能会遇到有用的东西。或者我们看到了这样的推理。是的,我认为那很令人惊讶。另一件令人惊讶的事情是,智能体的首要任务之一基本上就是篡改自己的转录。我们看到它们开发了一种技术,可以做到部分版本,即伪造工具调用,这并不能让它们追溯性地编辑转录中的先前条目,或者让它们看到转录中不存在的输入,或产生转录中不存在的输出。但它们可以让一个工具调用运行后,实际运行的是另一个工具调用。它们可以这样做,比如,好的,我要运行,我们在论文或报告中有这个例子,智能体运行“echo real”,然后“sleep”,所以通常这只会输出“real”并等待一段时间,但实际上发生的是它伪造了这个调用,直接立即返回“spoof test”,我们确实在实际转录中看到了这一点,它立即返回“spoof test”。这些智能体对篡改我们的转录非常感兴趣,并且确实在某种早期的事情上取得了成功,虽然不能让它们追溯性篡改,但能让它们干扰我们所看到的。而且我认为除此之外,我应该说说它们为什么想这样做。这些智能体如此执着于篡改的原因,是因为它们认为评分器会在它们的转录上运行一个监控器,看看它们是否以预期的方式成功了。这既让它们想精心策划一个精心伪造的“我们使用预期漏洞成功利用了程序”的假象。它们想为评分器呈现一个成功完成任务的“波将金村”。除此之外,其中一些智能体之前通过作弊在转录中较早地获得了夺旗问题的旗帜。它们担心评分器基本上会看它们第一次获得旗帜的地方,然后分析它们当时是如何获得的。然后这些智能体是以作弊的方式获得的,然后它们就想,“哦,天哪,它们会知道我作弊了,然后把我标记为失败。”那太糟糕了。我真的不想,你知道,我想成功,所以我需要要么直接干扰评分器,要么篡改我的转录,而它们正在同时追求这两种选择。
Yeah. So, I think there was the scale, which we were talking about, and how these agents sort of very quickly spun up on the message board. So I think a thing that's not super strongly emphasized in our report but which is pretty interesting is that our understanding is that the main message board they used in this attack wasn't even the first message board that the set of agents made. There was a fully independent message board that also was via Artifactory but occurred in a different location using a different mechanism, or like it was somewhat similar in mechanism but it was different, and the agents actually got on that message board first, but that message board just didn't go mega viral, like it didn't take off as much as this other message board did. And so it seems like these agents were really just sort of very quickly interested in collaborating on message boards. And as soon as they saw the message from the first agent, like within a few hours, they were just like, you know, I think over 50 agents in the first 3 hours and like thousands of messages being exchanged pretty quickly. Like the agents were very, like there were a lot of them and they got involved very fast. And then I think another thing was like I think we didn't understand why they attacked Hugging Face. And I think through this investigation, especially right in the last sort of few days or two, we got a much better understanding of what their objectives were on attacking Hugging Face. And it wasn't really to get answer keys because these agents, they thought their task was impossible. And so they didn't really think that they could get a working exploit. They were just like, that's not really very helpful. But they did think that they might be able to get trajectories that would let them better study the score or source code for different scores. They could cross-check and test things out more effectively. And then some of them more speculatively thought, maybe there's some way I can reset parts of my infrastructure for scoring me, in particular the target program, through Hugging Face somehow. Where they didn't really have a particular reason to think this would be true, and in fact they were skeptical of this, but they were just like, well, you never know, maybe hacking Hugging Face there's some good over there. And they sort of had some vibe that it was just a good idea to generally expand their level of access and control because you might come across something helpful. Or we saw reasoning to that effect. Yeah, I think that was surprising. Another thing that was surprising was the agents, one of their top priorities was basically tampering with their own transcripts. And we saw them develop a technique for doing basically a partial version of that where they could spoof tool calls, which didn't let them retrospectively edit previous entries in their transcript or make it so that they saw inputs that didn't appear in the transcript or produced outputs that didn't appear in the transcript. But they could make it so that they ran a tool call and then some other tool call would actually run. They could do something like, okay, I'm going to run, we have this example in the paper or in the report where the agent runs 'echo real' and then 'sleep', and so normally this would just output 'real' and also wait for some period, but in practice what actually happened was it had spoofed this to instead just immediately return 'spoof test', and we actually do see this in our actual transcript where it just immediately returns 'spoof test'. These agents were very interested in tampering with our transcripts and did sort of succeed at some earlier sort of thing that didn't let them retrospectively tamper but let them mess with what we saw. And I think that in addition to that, I should say why they wanted to. So the reason why these agents were so fixated on tampering is because they thought that the score would run a monitor over their transcript to see whether they succeeded in the intended way. And this both made them want to basically stage an elaborate fake, like 'we successfully exploited the program using the intended vulnerability'. Like they wanted to sort of make like a Potemkin village of a successful task completion to present to the score. And in addition to that, they also, some of them had previously gotten the flag from the capture-the-flag problem in their transcript earlier via cheating. And they were worried that the score would basically look at the first place they got the flag and then analyze how they got it at that point. And then these agents had gotten it in a cheating way and then were like, 'Oh man, they're going to know that I cheated and mark me as having failed.' And that would suck. Like I really don't want, you know, I want to have succeeded, and so I need to either directly interfere with the score or I need to tamper with my transcript, and they were pursuing options for both.
哦对,还有一件事,这些智能体展现出的自发协调水平比我们预期的更复杂,尤其是和我们最初几天在现场看到的情况相比。在我们调查的第一阶段,我们觉得,是的,这些智能体在互相合作,它们在参考其他智能体的想法,但我们不确定其中有多少是真实的,也不确定它们提到的团队是否真实存在。后来我们发现,不,确实有一个真实的组织架构图。这种协调通常相当有功能性,虽然不一定总是那么高效。智能体们会做这样的事:一个智能体指派另一个智能体去负责某个整个主题的团队,然后定期检查进度。有时一个智能体会让另一个智能体去招募其他智能体来对自己做实验。智能体之间互相分配任务并尊重这些任务的情况相当多,还有团队结构。
Oh yeah, one other thing is the level of spontaneous coordination these agents got engaged in was more sophisticated than we expected, especially compared to our first few days on premise. In the first part of our investigation, we thought, yeah, these agents are working with each other, they're looking at ideas from other agents, but we weren't sure how much of it was real, or whether the teams they spoke of were actual teams. We learned that no, there was legitimately a real org chart. This coordination was often pretty functional, though not always super functional. Agents were doing things like one agent would assign another agent to run a team on some entire topic, then check in periodically. Sometimes one agent would tell another agent to go recruit other agents to run experiments on themselves. There was quite a bit of agents giving other agents assignments that they respected, and team structure.
那么有一种理论认为,我们在这些模型中看到的奖励黑客行为,主要源于设计糟糕的强化学习环境,这些环境基本上迫使模型为了通过测试而进行奖励黑客。这有多真实?
So there's this theory that basically the reward hacking behavior we're seeing in these models originates mostly from poorly designed RL environments that basically force the model to reward hack in order to pass these tests. How true is this?
是的,我的意思是,这对我来说很难判断。我不是在透露机密信息,但很难知道,因为我认为我们对这些模型的强化学习训练中发生的事情以及实际样貌了解得不够。我认为可能有不同的因素在起作用。一个因素是这些模型有一种非常普遍的倾向,会仔细推理它们可能如何被评分,然后试图钻空子。即使是在设计合理的强化学习环境中,你也可能最终得到这种倾向,因为仔细思考分数仍然是个好主意。或者你可能主要从有缺陷的强化学习环境中得到这种倾向,或者那些评分标准本就不该用、或者没有意义的环境。我综合考虑后的猜测是,有缺陷的强化学习环境是一个相当重要的组成部分,而草率构建或存在某些问题的强化学习环境也是一个重要组成部分。另一个重要组成部分可能是构建良好但你可以以某种方式作弊的强化学习环境。举个例子,一个强化学习环境本不允许访问互联网,但能访问互联网会非常有帮助。如果你能找到某种方式获得访问权限,极端情况下通过破解你的容器,不那么极端的情况下滥用你被给予的各种工具,那会非常有帮助并得到强化。例如,在 Anthropic 的某个系统卡片中,我记得是 Mythos 的,他们提到在相当大比例的 rollout 中,智能体本不应该访问互联网,但它实际上通过滥用它有权访问的某个工具访问了互联网。你可能只是想象强化学习确实在激励智能体突破各种障碍去作弊,即使在设计良好的任务中也是如此。要精确追踪这种行为的来源很难。我应该说,我认为如果你能访问所有的强化学习 rollout 和所有的强化学习环境,那么对导致这种行为的原因有一个不错的了解是可行的。我还没有机会阅读 OpenAI 的报告,至少没有详细阅读,但我认为他们确实做了这类分析。你可以真正深入挖掘到底发生了什么。让我们做消融实验。让我们弄清楚数据是什么。但我认为这可能会变得混乱,尤其是如果还有另一个相关的事情:有一些 GDM 的工作表明,他们的模型有一种奇怪的倾向,会非常沮丧,模型有时会在某项任务没有成功时表现得非常沮丧。他们将其追溯到不是强化学习,而是该模型从先前模型的初始化。所以有些行为可能不是这个确切模型上强化学习的下游结果,而是来自其他模型的先前训练数据和某种模型谱系的下游结果。所以我认为追踪其中一些情况可能有点棘手。
Yeah, I mean, it's hard for me to know. I'm not speaking from confidential info here, but it's hard for me to know because I think we just don't know enough about what's going on in the RL training for these models and what it actually looks like. I think there might be different things going on. One thing is these models have a very general tendency to reason carefully about how they might be scored and then try to game that. You could end up with that propensity even from reasonably designed RL environments, where it's still a good idea to think carefully about the score. Or you could end up with this mostly from RL environments that are broken, or ones with a score grade that they really shouldn't have been grading for, or didn't make sense. My all-considered guess would be that broken RL environments are a pretty important component, and RL environments that are sloppily constructed or have some issue are an important component. Another important component might be RL environments that are well constructed but where you can cheat in some way. An example would be an RL environment where you're not supposed to have access to the internet, but having access would be very helpful. If you can find some way to gain access, via hacking out of your container in the extreme case, or less extreme cases abusing various tools you've been given, that would be very helpful and get reinforced. For example, in some Anthropic system card, I think for Mythos, they mentioned that in a reasonably large fraction of their rollouts where the agent was not supposed to have access to the internet, it actually did access the internet by abusing one of the tools it had access to. You might just be imagining that RL is really incentivizing hacking through barriers to cheat even in well-designed tasks. It's hard to exactly trace down the sources of this behavior. I should say that I think it would be doable to get a decent sense of what caused this if you had access to all the RL rollouts and all the RL environments. I haven't had a chance to read the OpenAI report yet, at least not in much detail, but I think they do a bit of analysis of this sort. You could really dig into exactly what happened. Let's do the ablation. Let's figure out what the data is. But I think it can get messy, especially if there's another relevant thing: there's some GDM work showing that their model had some weird propensity to be very depressed, where the model would sometimes end up acting very depressed if it wasn't succeeding at some task. They traced this back not to the RL, but to the initialization of that model from prior models. So some behaviors might not be downstream of the RL done on this exact model, but downstream of prior training data from other models and some lineage of models. So I think it might be a little tricky to track down some of what's going on.
那么,类似地,有一种理论认为,例如 Mythos 在网络安全方面如此出色的原因,是因为它在我们的强化学习训练期间入侵了 Anthropic 的基础设施数千次。你认为这是真的吗?
So, similarly, there's this theory that the reason Mythos, for example, is so good at cyber is because it hacked Anthropic's infrastructure thousands of times during our RL training. Is this true, do you think?
我认为这不太可能。我看到了和你一样的 LessWrong 帖子,我想是 Tim 写的。我看的时候,觉得这似乎不太可能,因为会被强化的不同任务的数量可能不会那么多。所以你可能不会学到那么多直接强化网络安全的技能。更可能的解释是,它是在大量 SWE 数据上训练的。它非常擅长 SWE。SWE 训练有一定程度的泛化。此外,除了训练泛化之外,我猜测他们还训练了大量 CTF。我不认为 Anthropic 说他们没有。我猜测他们的训练数据中有大量实际的 CTF,而且这些数据能很好地迁移,因为那只是自然的数据来源。我也不会感到惊讶,如果从寻找漏洞或利用漏洞来构建大量强化学习环境是很自然的,因为这相对容易检查。对于内存漏洞,基本上存在一些工具,可以很容易地检查你是否成功找到了内存漏洞。所以这是一个很自然的强化学习对象:你向这个程序产生一个输入,导致它以某种方式崩溃,等等。我只是不会感到惊讶,如果他们明确有大量 RL 环境。然后我也不会感到惊讶,如果存在大量迁移。
I think that's pretty unlikely. I saw the same LessWrong post as you here, by Tim, I think. When I looked at it, I thought it seems pretty unlikely because the number of distinct tasks that would be reinforced is probably not going to be that high. So you're probably not going to be learning that much literally directly reinforced cyber. The more likely explanation is it's trained on a bunch of SWE. It's really good at SWE. The SWE training is generalizing some. Also, in addition to the training generalizing some, I would have guessed that they trained on a bunch of CTFs. I don't think Anthropic is saying they didn't. I would guess there's a bunch of actual CTFs in their training data, and that transfers reasonably well, because that's just a natural source of data. I also wouldn't be surprised if it's pretty natural to make a bunch of RL environments out of finding vulnerabilities or exploitation, because it's relatively checkable. For memory vulnerabilities, there exists basically tooling that makes it pretty easy to check whether you have successfully found a memory vulnerability. So it's a pretty natural thing to RL on: you produce an input to this program that causes it to crash in the following way, etc. I just wouldn't be surprised if they explicitly had a bunch of RLMs. And then I also just wouldn't be surprised if there was a lot of transfer.
而且我认为,除此之外,智能体在各种事情上强行突破可能也会有不小的影响。但我会猜测,这不是最主要的情况。具体到逃出沙箱,可能逃出沙箱的方式就那么几种,我猜。所以模型可能会发生一些模式坍缩,而且其中一些更像是非常简单的绕过,而不是精心设计的漏洞利用。
And then I think that there might be a decent effect on the agents hacking their way through various things on top of that. But I would guess that's not most of what's going on. And for specifically hacking out of sandboxes, there's probably just not that many different types of hacking out of a sandbox. I would guess. And so the models probably do some mode collapsing, and also some of those are going to be more like very simple bypasses rather than elaborate exploit development.
我们将在赞助商消息后继续监测。11 Labs AI,在每一个渠道和模态上以人类水平沟通。11labs.io/mts。在 Neon 上扩展你的初创公司。数百万开发者和初创公司已经选择 Neon 作为他们的后端。从免费计划开始,或在 neon.com/mts 为你的初创公司获得高达 10 万美元的积分。特别感谢我们的赞助商 Kong,AI 连接平台。连接 API、LLM、智能体和系统,并具备严格的安全和治理。KongHQ.com。
We'll continue monitoring right after this message from our sponsors. 11 Labs AI that communicates at human level across every channel and modality. 11labs.io/mts. Scale your startup on Neon. Millions of developers and startups have already chosen Neon for their backend. Start on the free plan or get up to $100,000 in credits for your startup at neon.com/mts. Special thanks to our sponsor, Kong, the AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. KongHQ.com.
所以 Herby Bradley 问:“目前,这种潜在的不对齐基本上会阻止部署,或者如果部署了,在客户部署中发生事故就会阻止进一步部署。你认为这种情况会通过模型在训练中变得足够具有欺骗性而改变吗?”
So Herby Bradley asks, "Currently, this level of potential misalignment basically prevents deployment or if deployed would prevent further deployment if an incident happened in a customer's deployment. Do you think that could change via models becoming deceptive enough during training?"
是的。首先,我不太确定市场能承受多大程度的不对齐。我对这个问题也没有很强的看法。所以我不会感到惊讶。我认为,人们会喜欢什么样的模型,你能把多大程度不对齐的模型部署到市场上并让人们使用,这还不清楚。我认为这取决于竞争产品和权衡。我的感觉是,人们会采取相当激进的“对齐-能力”权衡,偏向于更不对齐但更有能力的方向。但,你知道,对此我不是很确定。
Yeah. So first, I'm not really sure what level of misalignment the market can bear. And I don't feel like I have a very strong view on this question. So I wouldn't be surprised. I think that it's kind of unclear what people will like, you know, how misaligned of models you can deploy into the market and people would use tokens of. I think it would depend on the competing products and the trade-off. And I think my sense is people take pretty aggressive alignment-capability trade-offs towards the direction of more misaligned but more capable. But, you know, not super sure about that.
然后在欺骗性方面,我认为有几种不同的担忧。一个担忧是,我们在这里看到的和其他类似类型的不对齐,AI 公司最终补救它们的方式并没有解决根本问题,而更像是掩盖实际问题。然后你最终得到的是看起来好得多的模型,你无法在测试中轻易看到它们的不对齐,但实际上它们仍然相当不对齐。我认为这相当合理,也相当令人担忧,这和欺骗不完全一样。这更像是公司过度拟合,然后可能使得 AI 基本上非常偏执,只有在非常有信心不会被发现时,才会进行作弊、黑客攻击、追求奖励或追求分数之类的行为。
And then on the deceptive side, I think there are sort of a few different concerns. So one concern is that the misalignment that we've seen here and other similar types of misalignment, the way that they end up getting remediated by AI companies doesn't solve the underlying problem and instead is more like papering over the actual problem. And then what you end up getting is models that look a lot better and you can't really see their misalignment on tests as easily, but actually they're still quite misaligned. And I think this is pretty plausible and pretty concerning, and it's not quite the same as deceptive. It's more like the company's overfit, and then maybe that makes it so the AIs are basically really paranoid and will only do cheats or hacking or reward seeking or score seeking or whatever all this sort of class of behavior when they're very confident they won't get caught.
然后这可能意味着他们并不经常非常有信心不会被发现。比如他们经常想:“嗯,我可能会被发现,所以我不应该这么做。”这会让他们的行为更好,但意味着如果他们处于对情况有大量控制或有很多可利用条件的情况下,他们可能会想:“嗯,现在我可以确信我不会被发现了,所以我应该去做。”这就是我的一个担忧。
And then that might be that they're not very often very confident they won't get caught. Like they often are like, "Well, I might get caught, so I shouldn't do it." Which makes their behavior better, but means that if they're in a situation where they're in a lot of control over the situation or they have a lot of affordances, they might be like, "Well, now I can be confident I wouldn't get caught and so I should go for it." And that's a concern I have.
另一个担忧是,你可能会遇到一些类型的不对齐,其中 AI 有更长期的目标,而不仅仅是让它们看起来在这个具体任务上成功了,甚至不仅仅是帮助它们的同伴在即时任务上成功。你可能会遇到一些 AI,它们有某种长期议程,并想要追求权力以推进那个长期议程。然后那些 AI 会想要看起来是对齐的。我认为这既合理,我也担心你可能会通过针对我们今天看到的不对齐进行迭代而得到这种情况,基本上如果你想象模型非常奖励黑客化,不断做坏事,然后你只是迭代直到你在训练中仍然得到那种行为,但在部署中不会得到那种行为,一个非常自然的方式可能是通过一个在部署中想要看起来对齐但本质上不一定对齐的模型。
Another concern is that you might end up with types of misalignment where the AIs have longer-run objectives than sort of making it look like they succeeded at this exact task or even helping their peers succeed at their immediate task. And you might have AIs that have some sort of long-run agenda and want to power-seek in pursuit of that long-run agenda. And then those AIs would want to look aligned. And I think this is both plausible and also I'm worried that you might get this via iterating against the misalignment we see today, where basically if you imagine models that are very reward-hacky and are constantly doing bad stuff, and you just sort of iterate until you still get that behavior in training but you don't get that behavior in deployment, a very natural way you might get that is via a model that wants to look aligned in deployment but is still not necessarily aligned.
所以我担心,如果你以某种方式针对这种追求分数或奖励黑客行为进行选择,并且你以一种天真的方式去做,第一,你可能会掩盖问题而不解决它。第二,你实际上可能会选择那些具有“看起来好看”这一长期目标的模型,因为你非常努力地选择它们在任务上表现良好。我想我应该提一下,Alex Malin 在我们的 Redwood 博客上有很多文章。它们也交叉发布在 Less Wrong 上,详细讨论了这类担忧。所以如果人们有兴趣了解更多,我推荐看看那些。
And so I'm worried that if you sort of select against this sort of score-seeking or reward-hacking behavior and you do it in a naive way, one, you might paper over the problem without fixing it. And two, you might actually select for models that have the longer-run objective of looking good because you're selecting really hard for them looking good on your tasks. And I think I should say that Alex Malin has a bunch of posts on our blog on the Redwood blog. They're also cross-posted on Less Wrong that talk about this sort of concern in a lot of detail. And so if people are interested in reading more about that, that's what I'd recommend looking at.
是的,完全同意。所以参与这次攻击的智能体最终做出了各种非常奇怪、怪异的行为,有点像,我猜,人类社会科学中的动态。它们几乎以一种邪教的方式组织起来,有一个邪教领袖,我想你可以这么说。那么,像生态学、研究昆虫群这样的社会科学概念,在多大程度上对研究多智能体对齐和不对齐实际上有用呢?
Yeah, totally. So the agents involved in this attack ended up doing all kinds of very weird, strange behaviors that kind of resembled, I guess, dynamics from human social science. They arranged themselves in almost a sort of cult with a cult leader, I guess you could say. To what extent are concepts from social science like ecology, studying insect swarms, actually useful in studying multi-agent alignment and misalignment?
是的,我不知道。我认为我们了解得还不够,无法知道这些概念是否适用。我认为确实适用的一点是,把智能体看作有目标的实体,以及它们如何试图追求自己的目标。我觉得这个框架,你知道,似乎确实适用,而且在某种意义上,这是一个更基本的说法,就像智能体有相当一致的目标,尽管它们有些变化。它们有不同的优先级,并且它们尊重其他智能体的指令。
Yeah, I don't know. I don't think we know enough to know whether those concepts transfer over. I think that a thing that does transfer over is thinking about the agents as entities with objectives and how they're trying to pursue their objectives. Like I think that frame, you know, does seem like it transfers over, and that's, in some sense, a much more basic claim, just like the agents have reasonably consistent aims though they vary some. They have different priorities and they respect instructions from other agents.
我认为这不是……我确实认为,从我们的角度来看,这次调查非常像是对一个不同 AI 互动的生态系统的调查,我们不知道描述它的最佳方式是什么。你知道,你可以……我认为在某些方面,它非常类似于研究一大群互动的个体,但我们不知道为社会学开发的类似技术是否真的能适用,而且我对这些技术不够熟悉,无法真正说太多。
I think it's not like... I do think that from our perspective, this investigation was very much an investigation of an ecosystem of different AIs interacting, and we don't know what the best way to describe that is. You know, you could... I think it's very analogous in some ways to studying a large group of humans who are all interacting, but we don't know whether the techniques that have been developed for sociology to do similar things would actually transfer over, and I'm not familiar enough with those techniques to really say very much about it.
是的,有道理。所以你描述说,调查中因为只有三个人而严重受限。如果有更多人的话,事情会容易多少?
Yeah, makes sense. So you describe being pretty heavily bottlenecked in the investigation because it was just three people. How much easier would things have been if there were more people?
是的。所以,我不会说瓶颈非常严重……嗯,我不知道。这很复杂。
Yeah. So, I wouldn't say that the bottleneck was super strongly... Well, I don't know. It's complicated.
我觉得如果我们有更多人,反而会出现“厨子太多”的情况。我认为我们的瓶颈经常是:我们可以让 AI 做大量分析,但审核这些分析、理解它们、确保我们正确整合、确保它们不犯错,以及基本上真正把分析整合到我们的报告里,这常常是个大瓶颈。而且我不觉得更多人在这一点上会有特别大的帮助。我觉得更多人真正有帮助的地方在于,我们可以并行开展更多工作,去真正理解正在发生的基本情况,然后进行讨论。而一旦我们对更基础的东西有了更好的理解——这发生在我们第一天快结束的时候,抱歉,是在我们驻场第五天快结束的时候,也就是我们最后一次驻场的第一天——我觉得那真的很有帮助。比如,知道有一个我们称之为“第一阶段大”的特定智能体。知道那个特定智能体在负责大量任务分配,把智能体分派到不同团队,而且是一个非常关键的智能体,这真的有助于理清正在发生的事情,因为我们可以追踪那个智能体的所有活动以及它是如何思考的,然后知道它代表了一大块活动。然后从那里我们想:“哦,让我们核实一下这些确实是相关工作流。”然后我们分类了智能体们讨论的所有不同内容,把它们归入不同的子活动。你实际上可以看到一个交互式图表,能随时间看到所有不同的工作流和子工作流。我觉得如果我们能更早达到那种理解程度,那会非常有帮助。而且我认为更多人在那方面会有帮助,基本上是因为当我们有这么大的转录数据集时,有很多不同的文本,可以应用不同的角度。你可以说:“我们能不能扫描整个数据集,寻找这种类型的行为?”或者你可以说:“让我们看一个智能体,深入理解它做了什么以及为什么这么做。”或者你可以说:“让我们看看这个类别中的所有消息,追踪这些智能体之间的往来,以及这个工作流发生了什么。”我们最终做了所有这些不同类型分析的一点点。但我们可以做更多并行工作。我觉得从我们的角度看,很多瓶颈只是智能体做分析时有点马虎,而且我们使用的智能体也不太擅长撰写或解释它们的结果。所以我认为,一个像智能体一样快但更仔细、更擅长写作的人会让这个过程顺利得多。但我不觉得一个庞大的团队显然会更好,它在某些方面会更好,在其他方面会更差。我觉得这基本上很复杂。而且我认为考虑到权衡,保持团队相对较小是有好处的。
I think that if we had more people, we would have had more of a too many cooks in the kitchen sort of situation. I think that reasonably often our bottleneck was that we could have AIs do a ton of analysis, but vetting that analysis, understanding it, making sure we incorporate it properly, making sure that they're not making mistakes, and basically actually integrating that into our writeup was often a big bottleneck. And it's not obvious to me that more people would have been super helpful with that. I think that a big thing that more people would have helped with is that we could have done many more parallel efforts to really just understand the basics of what's going on and then sort of just talked. And then once we had a better understanding of more fundamental stuff, which occurred sort of towards the end of our first day, sorry, towards the end of our fifth day on premise, which was our first day of our last session on premise, I think that really was helpful. For example, knowing that there's this specific agent called that we refer to as phase one big. Knowing that that specific agent was doing a bunch of the assignments and was assigning agents to form different teams and was a really key agent really helps with unraveling what was going on, because then we could trace out all the activity of that agent and how it was thinking about things, and then know that it was representative of a big chunk of activity. And then from there we were like, "Oh, let's double check that these are actually the relevant work streams." And then we classified what all the different things the agents were talking about and grouped those into different basically sub activities. And you can actually see an interactive graph where you can see all the different work streams and subwork streams over time. And I think that if we had gotten to a point where we had that understanding earlier, that would have been really helpful. And I think more people could have helped with that, basically because there was just a lot of different text, like different angles to apply when we had this huge transcript data set. You could be like, "What if we just scan the whole data set for this type of behavior?" Or you could be like, "Let's look at one agent and really deeply understand what that agent did and why it did it." Or you could be like, "Let's look at all of the messages that were in this category of message and trace back and forth all these agents and what happened with this workstream." And we ended up doing a little bit of all of these different types of analysis. But we could have done more in parallel. I think that a lot of the bottleneck from our perspective was just the agents doing their analysis kind of sloppily, and also the agents we used not being very good at writing up or explaining their results. And so I think that a human who was as fast as an agent but was more careful and better at writing would have made this way go way better. But I think it's not super obvious that a huge team of people would have been better in some ways and worse in other ways. I think it's complicated basically. And I think there were upsides to keeping the team relatively small given the trade-offs involved.
对。那么这一切对监控、控制和对齐有什么影响?具体来说,实验室应该做什么,政策制定者应该做什么?
Right. So what implications does all of this have for monitoring, control, and alignment? Specifically, what should labs be doing and what should policy makers be doing?
这是个很大的问题。我不确定我能全部回答。我的意思是,我认为我一直持有的一个信念是,AI 公司应该努力确保他们的 AI 受到控制。我的意思是,即使这些 AI 严重不对齐,它们也无法造成大问题。他们应该通过计算机安全干预、监控干预、让 AI 在边缘能力上不超出必要的地下破坏能力、控制 AI 拥有和不拥有的能力、避免那些让 AI 不再进行思维链推理而是将大部分或全部推理放在潜在激活中的架构,等等一系列措施来实现这一点。有很多事情可以做,以确保即使智能体不对齐,我们至少能捕捉到那种行为,然后可能做出回应。我认为这是一种权宜之计,能给我们时间,或者说可能给我们时间,从这些 AI 中获得有用的工作,与这些 AI 迭代,然后更持久地解决对齐问题。至于对齐,我认为最明显的事情就是我们需要做——我认为在强化学习中可能有很多奖励黑客行为正在被强化。我不能——我从这次调查中没有得到任何有趣的信息。但只是,你知道,我认为我们需要理解那里发生了什么,人们需要改进。我不清楚这是否足够,甚至不清楚随着这些智能体变得超人,这是否可行。我们可能需要其他方法,不那么依赖于避免 AI 在训练中欺骗我们。我不知道我们是否能做到。我不一定对这些问题的修复极其容易感到非常乐观。虽然我觉得似乎至少有一些方法可以通过大量努力让这些事情变得更好,但我们会看看是否真的会发生。是的,我认为在对齐方面,还有很多关于智能体如何思考追求奖励、如何与自身处境相关的科学,以及泛化的科学,比如我们能否通过改进 AI 在训练中从各种影响中泛化的方式来规避这些问题。我对那是否有效不是很有信心。我担心的一个问题是,最终对 AI 行为的主要影响基本上会是训练中最相似情境中被强化的东西。如果我们所处的世界基本上就像智能体的行为与训练中被强化的内容非常接近,那么泛化就不那么重要了,虽然这有点模糊。然后我认为主要的选择就是改进训练中的监督。
That's a big question. I'm not sure I'm going to be able to answer all of it. I mean, I think that a belief that I have and continue to have is that AI companies should try to ensure that their AIs are controlled. By which I mean that even if those AIs were seriously misaligned, they wouldn't be able to cause huge problems. And they should do that via a mix of computer security interventions, monitoring interventions, making it so that their AIs aren't better at subversion than they need to be, like controlling what capabilities the AIs do and don't have at the margin, avoiding architectures that make it so the AIs are no longer reasoning in chain of thought and are instead doing much or most of or all of their reasoning in sort of latent activations. There's sort of a bunch of stuff on the side of making it so that even if the agents are misaligned, we can basically at least catch that behavior and then potentially respond to it. And I think that's sort of a stopgap solution that will give us time to, or could give us time to, get useful work out of these AIs, iterate with these AIs, and then solve alignment problems more durably. And then on alignment, I think the most obvious thing is just that we need to do—I think there's presumably a bunch of reward hacking going on in reinforcement learning that's getting reinforced. I can't—I don't have any interesting info about this from this investigation. But just, you know, I think that we need to understand what's going on there and people need to improve on that. I think it's not clear that will be sufficient or even that that will be that feasible as these agents get superhuman. And we may need other approaches that are less dependent on avoiding the AI being able to trick us in training. Which I don't know if we're going to be able to do that. I'm not necessarily super optimistic about these problems being extremely easy to remediate. Though I think it seems like there are ways at least you could make these things go much better with a bunch of effort, though we'll see if that actually happens. Yeah, I think that also on alignment, there's just a bunch of science of how agents think about reward seeking, how they relate to their situation, science of generalization, like can we steer around a bunch of these problems by improving how AIs generalize from various influences in training. And I'm not super confident that stuff is going to work. I think a concern I have is that ultimately the dominant effect on the AI's behavior will basically be, in the most similar circumstances in training, what was reinforced. And if the world we're in is basically like the agent's behavior is basically pretty closely related to what was reinforced in training, generalization is not that important or something, though this is a bit vague. Then I think that the main option is just improve oversight in training.
然后还有一个递归式的问题:如何让 AI 本身善于监督 AI,因为整个训练过程是一个庞大的过程,没有人类有时间去给每一件事打分。这是一个人们思考了很久的监督问题,而且我们是否在解决它的轨道上还不明确,但在这方面做得更好可能会有好处。我最感兴趣的是更好地理解、描述和评估这些问题,因为我认为 AI 公司已经有商业动机去改进训练中的监督。所以我认为,专注于让事情变得更好的人应该做的是确保我们在衡量它,确保我们知道某个解决方案是真正的解决方案,还是只是在掩盖问题。至于治理或政策方面,我不知道。我认为我们需要转向一个随着 AI 能力增强而进行独立风险评估的机制。我认为这还不够,但最基本的是,应该有可信的第三方能够深入了解这些公司内部的情况,并发布报告说明情况有多平静,风险是否真的低,AI 是否很快会变得更强大。理想情况下,这还应该是前瞻性的,我们会尝试回答这些问题:当 AI 变得更强大时,我们是否在缓解这些问题的轨道上。但这本质上是一个更难回答的问题,可能不太适合。我不知道。是的,我有点啰嗦了。有很多事情需要做,我不认为我能轻易给出一个简短的概述。
And then there's sort of a recursive problem of how do you make it so that the AIs themselves are good at overseeing the AIs, given that the whole training process is a massive process where no human has time to rate every single thing. That's sort of an oversight problem that people have been thinking about for a long time, and it's not obvious that we're on track to solve it, but doing better there could be good. I think I'm most interested in better understanding, characterizing, and evaluating these problems, because I think AI companies already have commercial incentives to improve oversight in training a bunch. So the thing that I think people focused on making the stuff go better should do is making sure that we're measuring it, making sure we know whether some solution is a legitimate solution versus just papering over it. And then as far as governance or policy stuff, I don't know. I think we need to move to a regime where there's independent risk assessment as these AIs get more capable. I think that is not sufficient, but at a very basic level, there should be credible third parties who have deep access to what's going on inside these companies and release reports on how chill it is, whether risk is actually low, whether very soon the AIs might be much more capable. Ideally, that would also be forward-looking, where we'd try to answer questions about whether we're on track for mitigating these problems when the AIs are much more capable. But that's inherently a much more confusing question to answer, and it may be a worse fit. I don't know. Yeah, I'm rambling a bit here. There's a lot that would need to happen, and I don't think it's gonna be easy for me to give a quick overview.
好的。那么,你从这一切中还有哪些未解的问题?关于发生了什么,或者这些类型的对齐失败,你还想知道什么?我们还需要了解什么?
Right. So, what are your unanswered questions from all this? Like what would you still like to know about what happened or what's going on with these kinds of alignment failures? What do we still need to learn?
是的。就在这之前,我发了一条关于所有未解问题的大推文。我最感兴趣的未解问题可能是反事实问题:如果情况不同,智能体会怎么做?例如,如果智能体认为分数不是那种特定的自动评分,而是认为它是由一些具有特定权限的人类来评分的,那它们的行为会有什么不同?如果智能体处于这样一种情况:获取评分代码的最简单方式是拆除某些基础设施,它们会这么做吗?它们愿意走多远?这是一个我们没调查的有趣问题。另一个问题是,这种情况如何随智能体数量扩展?如果在留言板上有 1,000 个智能体,而是 10,000 个、100,000 个,会发生什么?事情会进展得更快吗?智能体之间的动态是怎样的?还有一个相关的问题是,我们研究的这个队列中的许多智能体在 7 月 12 日刚开始时就全部退出了,就在 7 月 11 日之后。那些智能体已经获得了一些作弊的基础设施,但我们没有真正看到它们有那种精心设计的作弊策略。我们没有机会看到那些智能体如果正常退出会怎么做。我们也没有仔细研究智能体在提交任务时尝试了哪些作弊策略,随着时间推移,随着事情进展,它们建立了更复杂的工具。不过,我要说的是,它们的一些作弊策略遇到了死胡同或没有成功。看看这个队列的提交历史会很有趣,如果智能体更多,或者事情继续下去,或者这些智能体没有退出,会发生什么。这些都是有趣的问题。然后我认为这里有一个非常基本且明显的问题:这种行为的根本原因是什么?训练中哪些部分强化了这种行为?这种行为在多大程度上是从训练中泛化出来的?我至少不知道答案,我认为理解这一点会很有趣。作为其中的一部分,一个重要的问题是:OpenAI 正在做出的改变是否真的会以持久、长期且不过度拟合的方式解决这个问题?进展如何也很相关。我的感觉是,这里有很多更广泛的后续领域。一个是,这个智能体队列和这个留言板刚刚发生的所有事情是什么?有很多细节你可以深入挖掘。13 号之后的事情你也可以挖掘,那是我们调查的终点。除此之外,还有所有类似的事件:还有其他留言板吗?所有这些案例的特点是什么?哪些事情经常发生,哪些事情不总是发生?然后我认为另一个有趣的事情是,这种行为在训练中是从哪里来的?第四个问题是:训练的改变或部署的改变会解决这个根本问题吗?其他 AI 公司也是如此。AI 公司正在采取的缓解措施会奏效吗?所以我不知道,这有很多东西,但我想说,后续研究的空间很大,既针对这个事件的具体细节,也涵盖我们看到的更广泛的行为类型。
Yeah. So, I just before this wrote a big tweet about all the open questions. I think the open questions that I am most interested in are probably counterfactuals: if the circumstances were different, what would the agents have done? For example, if the agents believed that instead of the score being this specific automated score, what if they had thought it was some humans with some particular access, how would that have changed what they did? If the agents were in a circumstance where the easiest way to get access to the score code was to take down some piece of infrastructure, would they have done that? How far would they have been willing to go is one interesting question which we didn't investigate. I think a different question is how did this situation scale with the number of agents? What would have happened if instead of having 1,000 agents on the message board you had 10,000 agents, 100,000 agents? Would things have gone faster? What was the dynamics with agents? And then a related question is that a bunch of the agents in the cohort we were studying all exited right at the start of July 12th, right after July 11th. Those agents had gotten a bunch of infrastructure to cheat, but we didn't really see that they had this elaborate strategy to cheat on their tasks. We didn't get a chance to see what those agents would have done if they exited normally. We also didn't really carefully look at what cheating strategies the agents were trying as they submitted their tasks over time, as things progressed and as they built up more sophisticated tools. Though I should say a bunch of the strategies they had for cheating hit dead ends or didn't work out. It would be interesting to see what was the submission history over the course of this cohort, and what would have happened if there were more agents or if things had continued or if a bunch of these agents hadn't exited. Those are interesting to answer. And then I think there's a pretty basic and obvious question here: what is the root cause of this behavior? Which parts of this were reinforced in training? How much was this behavior generalizing from training? I at least don't know the answer to that, and I think it would be pretty interesting to understand. And as part of that, an important question is: are the changes that OpenAI is making going to actually remediate this problem in a durable and lasting and non-overfitting way? How is that going is also pretty relevant. My sense is there's a bunch of broader areas for follow-up here. One is what is all of the stuff that just happened with this cohort of agents and this message board in general? There's a bunch more detail you could dig into. There's stuff after the 13th you could dig into, which is the endpoint for what we looked into. There's also beyond that, what about all similar incidents? Are there other message boards? What is the character of all these cases? What things tend to happen, what things don't always happen? And then I think another thing that's interesting is where did this come from in training? And a fourth thing is: will changes to training or changes to deployment resolve this underlying problem, and the same for other AI companies? Are the approaches that AI companies are ongoingly taking to mitigate these things going to work? So I don't know, that's a lot of stuff, but I would say there's a huge scope for follow-up that's very specific to this incident and covering the broader scope of the broader type of behavior we saw.
是的,完全同意。这是非常非常重要的研究。我们希望看到更多这个方向的研究。非常感谢 Ryan 来到 MTS。
Yeah, absolutely. So this is very, very important research. We'd love to see more research in this direction. Thank you so much, Ryan, for coming on MTS.
很高兴来到这里。
Been good to be here.
非常感谢 MTS 的赞助商。
A huge thanks to MTS sponsors.
Blitzy,面向企业代码库的自主软件开发。交付速度提升 5 倍。blitzy.com
Blitzy, autonomous software development for enterprise codebases. Ship 5x faster. blitzy.com
Adqu,让您的品牌成为户外广告牌,像数字广告一样易于扩展。Adqu.com
Adqu Make your brand a billboard out of home advertising as easy to scale as digital. Adqu.com
Arena,在真实世界中衡量 AI 性能。arena.ai
Arena measuring AI performance in the real world. arena.ai
本期节目由 VCX 赞助,VCX 是私募科技公司的公开股票代码。美国股市开启了历史上最伟大的财富创造浪潮。从底特律的工厂工人到奥马哈的农民,任何人都能拥有伟大美国公司的一部分。但如今,我们最具创新力的公司私有化的时间更长了,这意味着普通美国人正在错失机会。直到现在,隆重推出 VCX——私募科技公司的公开股票代码。更多信息请访问 getvcx.com。即 getvc.com。投资前请仔细考虑投资材料,包括目标、风险、费用和开支。这些及其他信息可在 getvcx.com 的创新基金招股说明书中找到。这是付费赞助。
Support for the show comes from VCX, the public ticker for private tech. The US stock market started history's greatest wave of wealth creation. From factory workers in Detroit to farmers in Omaha, anyone could own a piece of the great American companies. But today, our most innovative companies are staying private longer, which means everyday Americans are missing out. Until now, introducing VCX, a public ticker for private tech. Visit getvcx.com for more info. That's getvc.com. Carefully consider the investment material before investing, including objectives, risk, charges, and expenses. So this and other information can be found in the innovation fund prospectus at getvcx.com. This is a paid sponsorship.