为未来抛硬币:Ryan Greenblatt 谈 AI 风险与对齐

Coin Toss for the Future: Ryan Greenblatt on AI Risk and Alignment

瑞安·格林布拉特 Ryan Greenblatt · Sam Harris · 2026-09-24 · 约 97 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

红木研究首席科学家 Ryan Greenblatt 解释为何他认为失控 AI 接管世界的概率高达 50-60%,以及为何他仍看到一条出路。

Redwood Research chief scientist Ryan Greenblatt explains why he puts 50-60% odds on misaligned AI taking over—and why he still sees a path through.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 43)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Background

Host

我和 Ryan Greenblatt 在一起。Ryan,感谢你加入我。

I am here with Ryan Greenblatt. Ryan, thanks for joining me.

Ryan

很高兴来到这里。

It's good to be here.

Host

所以,你是 Redwood Research 的首席科学家,这是一家 AI 安全研究公司,分析最近的 Hugging Face 惨败事件,一个可怕的现象。我们会深入讨论,但你是怎么开始这项工作的?你专注于 AI 安全的路径是什么?

So, you're the chief scientist at Redwood Research, which is one of these AI safety research firms that analyze the recent Hugging Face fiasco incident, a terrifying phenomenon. We'll get into it, but how did you come to this work? What was your path to focusing on AI safety?

Ryan

是的。所以,在大学三年级时,我独自在公寓里,因为当时是 COVID,所有课程都是远程的。我处于沉思的心情,思考我该做什么,听着各种播客。在其中一个播客中,有人提出了一个论点,我总结为:自私没有太大意义,因为在物质层面上,你未来的自己和其他未来的人之间有什么区别呢。我想,这个论点对我来说有点道理。我真的应该考虑我想如何度过我的一生以及我该做什么。最终,我相当认同要更加关注利他主义。之后,我花了很多时间思考我该做什么。最终,我决定从事技术 AI 安全是非常重要的事情,是我们将面临的一个非常关键的问题。我被这些论点说服,申请了各种地方,然后开始在 Redwood 工作,在那里我从事 AI 安全和 AI 安全研究大约五年了。

Yeah. So, in college during my junior year, I was in my apartment alone because it was COVID, all classes were remote. And I was in a contemplative mood thinking about what I should do with my life, listening to various podcasts. And in one of those podcasts, someone made an argument that I would summarize as: it doesn't make that much sense to be selfish because what's really the distinction at a material level between your future self and other future people. And I was like, that argument kind of makes sense to me. I should really consider how I want to lead my life and what I should do. And I ended up getting pretty sold on being much more focused on altruism. And then after that I spent a bunch of time thinking about what I should do with my life. Eventually I ended up deciding that working on technical AI safety was a very important thing to do, a very critical issue that we would face. I got sold on those arguments, applied to various places, and then started working at Redwood, where I've been working on AI security and AI safety research for about five years.

Host

你的背景是计算机科学吗?

And is your background in computer science?

Ryan

是的。计算机科学,还有一些数学。

Yeah. Computer science, also some math.

Host

所以,听起来你可能来自有效利他主义社区。那是——我是说,你觉得自己是 EA 社区的一员吗?

So, it sounds like you might have come up through the effective altruist community. Is that—I mean, do you consider yourself a part of the EA community?

Ryan

我会说,我肯定在 EA 社区内。但我不确定我会自我认同为 EA,但那在某种程度上只是因为我有点不愿意自我认同任何标签,因为那会让你在某些方面被关联——我觉得我不知道我是否喜欢——当我想到自己时,我不认为我是 EA。我有点——是的,我有很多那个社区其他人有的属性。我和那个社区有联系,但我不一定会说我就是 EA 本身。我确实认为我对有效追求公正的善感兴趣。

I would say that I'm definitely within the EA community. I'm not sure I would self-identify as an EA, but that's I think to some extent just because I'm a bit reluctant to self-identify with any label that's going to associate you in some—I feel like I don't know if I like—when I think about myself, I don't think I'm an EA. I'm sort of—yeah, I have a bunch of properties that other people in that community have. I'm in touch with that community, but I wouldn't necessarily say I'm an EA per se. I do think that I'm interested in effectively pursuing the impartial good.

Host

对。对。嗯,如你所知,也如我最近在播客中谈到的,有效利他主义受到了一些攻击。其中很多是不公平的。有些可能是公平的。而且有一个——我认为有一个原因,为什么我像你一样,在生活中没有挂出 EA 的招牌。也许我们可以谈谈这个。但听起来——好吧,你告诉我你在担忧光谱上的位置?我会把最担忧的人放在那边,比如 Eliezer Yudkowsky 和他最近的合著者 Nate Soares,也许 Max Tegmark 也在那边。Nick Bostrom 不完全在那边,但仍然在光谱的那一侧。然后在另一边反对他们的人,他们说这里真的没有真正的担忧。这都是骗局,任何认为人工智能可能失控并毁灭我们的想法,都是前沿实验室采用的一种奇怪的营销形式。这是为了施加监管并实现某种监管捕获,或者是一个 EA 骗局。你有像——你知道,我会说 Marc Andreessen 或 David Sacks 甚至总统本人,他最近用很多话称其为骗局——他们根本不担心这里有任何重大的下行风险,只看到美元符号延伸到地平线。你在这个光谱上的位置在哪里?

Right. Right. Well, as you know, and as I've talked about in the podcast of late, effective altruism has come in for some abuse. Much of it unfair. Some of it might be fair. And there's a—I think there's a reason why I, like you, have not hung up an EA shingle in my life. Perhaps we can talk about that. But it sounds like—well, you tell me where on the spectrum of concern are you? I would put the most worried people over there with Eliezer Yudkowsky and his recent co-author Nate Soares, maybe Max Tegmark is over there. Nick Bostrom is not quite over there but still on that side of the spectrum. And then all the way on the other side opposing them, you have the people who say that there really is no real concern here. That it's all a hoax, that any thought that artificial intelligence might get out of our control and destroy us, that is a weird form of marketing being employed by the frontier labs. It's an effort to get regulation imposed and achieve something like regulatory capture, or it's an EA scam. And you have people like—you know, I would say Marc Andreessen or David Sacks or even the president himself, who's recently called it a hoax in so many words—who are just not fundamentally worried about any significant downside risk here and just see dollar signs stretching out to the horizon. Where do you locate yourself on that spectrum?

Ryan

所以我会说我很担忧,但也许更乐观,我们会挺过去,而不会改变 AI 发展的方向。所以我会说,假设我们沿着当前的默认路径前进。也许有大约 50% 或 60% 的概率,未对齐的 AI 最终会接管世界。然后如果那发生了,就有很大的概率许多或所有人类会死亡。所以我认为——我会说我很担忧,不认为情况会顺利发展。我还想指出,除了这些未对齐的担忧之外,构建极其强大的 AI 系统还有其他风险,特别是关于权力集中和谁控制这些系统。一个答案是没有人控制这些系统。它们未对齐。另一个可能的答案是权力非常集中在少数人手中,它颠覆了我们的机构和我们的能力,无法实现权力的广泛分布、民主等。

So I would say I'm very concerned but maybe more optimistic that we'll make it through without sort of averting the course of AI development. So I would say that I'm sort of like—suppose that we proceed on what the current default path looks like in front of us. Maybe there's about a 50 or 60% chance that misaligned AIs would end up taking over the world. And then if that did happen, there would be a significant chance that many or all humans would die. So I think that's—I would say I'm very concerned and don't think the situation is on track to go well. And I would also note that in addition to these misalignment concerns, there's other risks with building extremely capable AI systems, especially around concentration of power and who controls these systems. And one answer is no one controls the systems. They're misaligned. Another possible answer is the power is very concentrated in a few people and it sort of overturns our institutions and our ability to have a broad distribution of power, democracy, etc.

Host

那么,是什么让你没有成为最担忧的人之一?你不同意他们什么,或者你认为他们错在哪里?

So what keeps you from being among the absolutely most worried? What do you disagree with them about or what do you think they're getting wrong?

Ryan

是的。所以,我会说,如果技术对齐问题结果证明有点容易,我认为这是可能的,并且我们在对齐和利用 AI 系统方面做得相当好,达到大约人类水平的能力,我们就能让这些系统自动化技术安全工作。这可能让我们进入一个正反馈循环,这些系统使自己更加对齐。它们产生一个新版本的系统,甚至更擅长仔细 figuring out 该做什么,更擅长追求我们的意图,并且这种自我强化而不是随着时间的推移越来越糟,而且这可能足够快,以跟上能力的快速增长。所以这是一个原因,我会说这归结为对平淡或相对经验和迭代方法的更多乐观。我不会说我对这些方法非常乐观,但我认为它们有相当的机会有效。只是当我说它们有相当的机会有效时,我也隐含地意味着有相当的机会失去对未来的控制。而我对此更像是 50/50。然后我还认为,我们可能会开发出非常非常强大的 AI 系统,而最糟糕的未对齐担忧大多不会实现,即使方法不是很先进,即使独立于这种自动化。我认为这是少数情况,但有可能。

Yeah. So, I would say that it seems pretty plausible to me that if it turns out that the technical alignment problem is somewhat easier, which I think is plausible, and we do a reasonably good job aligning and utilizing AI systems up through roughly human levels of capability, we could then get those systems to automate technical safety work. And that could potentially get us into a positive feedback loop where these systems make themselves more aligned. They produce a new version of the system that's even better at carefully figuring out what to do, better at trying to pursue our intentions, and that sort of is self-reinforcing rather than getting worse and worse over time, and that could be fast enough to keep up with the rapid growth in capabilities. So that's one reason, and I would say this comes down to more optimism about prosaic or relatively empirical and iterative methods. I wouldn't say I'm hugely optimistic about those methods, but I think that there's a decent chance they're working. It's just that when I say a decent chance they're working, I also implicitly mean a decent chance of losing control over the future. And I'm more like 50/50 on that. And then I also think it's plausible that we'll develop very very capable AI systems and the worst misalignment concerns will mostly not materialize even with not very advanced methods, even independent of this automation. I think that's a minority but it's possible.

AI接管的概率 Probability of AI takeover

Ryan

我并不认为有极其强烈的理由相信,在构建显著超人类的系统时,这些系统会如此不对齐,以至于它们会接管一切。确实有相当强的证据指向那个方向,而且这些系统越强大,这种情况就越令人担忧。但我可以想象,人们并没有真正采取多少预防措施,而是沿着 AI 发展轨迹前进,遇到问题就修补,最终在一切仍然正常的情况下获得了非常高的能力水平,然后社会还有时间做出反应。

I don't think there's extremely strong reason to believe that when building significantly superhuman systems, those systems would be so misaligned that they would take over. There's definitely pretty strong evidence pointing in that direction, and that case gets more concerning the more capable these systems are. But I could imagine people not really taking much of a precaution, proceeding through the AI development trajectory, patching problems as they come up, and that ending up getting you very high levels of capability while things are still fine, and then there would be time for society to react.

Host

你在这里提到了几个关于概率的数字,并大致指向了一个可能性范围。我想知道人们应该如何看待这类陈述。

You've mentioned a few numbers here with respect to probability and just kind of gestured at a range of likelihood. I'm wondering how people should think about statements of that kind.

Ryan

在我看来,我们对此赋予的任何实际概率基本上都是编造的。也许你有更严格的方法来做出估计,但无论估计值是多少,除非它小到微乎其微,比如远低于 1%——而这从来不是你听到的数字——有些人会说 10%、20%、30%——情况似乎就是这样。你刚刚谈到了 50% 的接管概率;我不知道那如何转化为一切的毁灭。但这些是巨大的数字。

It seems to me that any actual probability we would assign to this is pretty much made up. Maybe you have a more rigorous way of making an estimate here, but whatever the estimate, unless it was infinitesimally small, like well below 1%, which is really never the number that you're hearing—some people will say 10%, 20%, 30%—that seems to be the case. You just talked about 50% takeover; I don't know how that translates into the ruination of everything. But these are enormous numbers.

Host

所以在正常情况下,如果我们正在开发一项技术,而最接近这项工作的人说:‘是的,我认为按照我们目前的路线,可能有 10% 的概率我们会毁灭世界’,那么对那一系列结果的唯一理性回应就是立即停止,对吧?

And so in the normal case, if we were developing a technology where the people who were closest to doing the work said things like, 'Yeah, I think maybe there's a 10% chance we're going to destroy the world here on our present course,' the only rational response to that range of outcomes is you stop immediately, right?

Ryan

对。

Right.

Host

我的意思是,如果曼哈顿计划的科学家说:‘是的,我们做了计算,我们把最聪明的人聚在一起,我们考虑过了,有 10% 的概率,当我们在阿拉莫戈多进行第一次试验时,我们会点燃大气层并毁灭未来’,那么对此唯一理智的回应就是不要进行这次初始试验。但这里似乎完全不是这样。有很多人说某种极端糟糕结果的概率相当高——在掷普通骰子甚至抛硬币的范围内——然而工作仍在军备竞赛的条件下全速进行。我们会谈到 Dario 和其他人最近的声明,说我们可以以某种方式更好地控制节奏。但这似乎不是任何人在认为毁灭概率真的那么高时会有的反应。

I mean, if the Manhattan Project scientists said, 'Yeah, we've run our calculations, we've got the smartest people in the room together, we thought about it, and there's a 10% chance that when we execute this first test at Alamogordo, we ignite the atmosphere and destroy the future,' the only sane response to that is you don't do this initial test. But that doesn't seem to be what's happening here at all. We have a lot of people saying that the probability of some extremely bad outcome is quite high—within range of the roll of a normal die or even a coin toss—and yet the work is proceeding more or less under an arms race condition at full pace. We'll talk about the recent statements of Dario and others that we could somehow pace this better than we are. But it just does not seem like the response anyone would have if they thought the probabilities of doom were really that high.

Ryan

是的。所以,仅就概率问题而言,我完全同意这些概率是不精确的。它们是主观的,而且有悠久的传统来做主观概率预测。你可以把我的观点理解为:当我全面考虑情况时,如果我们沿着当前默认的轨迹前进,我不知道,基于对各种因素的权衡,AI 接管导致的毁灭与其他结果对我来说似乎大致同样可能。然后,当我审视一系列不同的可能情景,尝试以不同方式分解风险来源,并确保我的数字与其他观点一致时,看起来是合理的。我当然也尝试对一些更近的结果进行预测,从中获取一些信号,并努力成为一个好的预测者。虽然我肯定不是世界上最好的预测者。但我的 AI 预测,我认为至少还可以或不错。

Yeah. So, just on the probability question, I definitely agree that these probabilities are imprecise. They're subjective, and there's a long tradition of how to do subjective probability forecasts. You can interpret my view as more like: when I look at the all-considered situation, if we proceed on what seems to be the current default trajectory, I'm like, I don't know, the outcome of doom from AI takeover versus something else seem roughly equally likely to me based on weighing the factors. And then when I look through a bunch of different possible scenarios and try to break down the sources of risk into different ways and try to make sure my numbers are consistent with other views I have, it looks like that works out. And I of course also try to do some forecasting of closer outcomes where we can get some signal on that and try to just generally be a good forecaster. Though I'm not the best forecaster in the world for sure. But my AI forecasting is, I think, at least okay or decent.

Host

鉴于人们表达了如此大的担忧,为什么所有这些 AI 公司都在以最大可能的速度推进?

Given the state of affairs where people express such large concerns, why is what's happening that all these AI companies are proceeding at the maximum possible pace?

Ryan

所以我认为这里有几个不同的因素。其中之一是,许多 AI 公司根据其公开声明实际上似乎相当担忧,但它们内部不一定统一。而且,AI 领域并没有强烈共识认为当前的 AI 发展进程即将面临巨大风险。更多的共识是,如果你在短时间内构建出极其超人类的 AI 系统,那将带来非常高的风险,尽管不一定有共识。但我认为人们常常在能力轨迹及其发展方式上存在分歧。另一个因素是,不同的行为者有各自不同的野心,并且认为让自己处于更有利的位置来影响技术可能是降低风险的最佳途径,或者至少是一种途径。例如,Anthropic 和 OpenAI 的部分说法似乎是:如果我们先开发出这项技术,我们会比下一个采取更差预防措施的参与者做得更负责任。我经常从行业人士那里听到类似这样的话:嗯,我们可以做那件会让我们慢下来的事,但如果我们那样做了,显然还有其他 AI 公司要担心——他们也会那样做吗?我认为这里普遍存在军备竞赛,在这种情况下,当人们认为那可能是最划算的交易时,他们会冒巨大风险,这并不太令人惊讶。现在,我不太确定我同意这实际上是一个好的策略,相对于公司可以做的其他事情。我并不是说我一定同意那个观点,但我认为那是人们持有的一种观点。我认为我们之所以在推进,而且没有更强的,例如政府的干预,只是这一领域缺乏共识的下游结果。

So I think there's a few different factors here. One of them is that many of the AI companies are in fact seemingly quite worried based on their public statements, but they're not necessarily internally unified. And it is not the case that there is a strong consensus across the AI field that the immediate course of AI development is imminently very risky. I think there's more consensus that if you built AI systems that are wildly superhuman in a short period of time, that would yield a very high level of risk, though not necessarily consensus for that. But I think often people disagree about the capability trajectory and how that is going to go. And I think another part of that is there are different actors with their own different ambitions and also think that themselves being in a better position to influence the technology might be the best route to reduce risk, or at least a route to reducing risk. So for example, it seems like part of the story for Anthropic and OpenAI is something like: if we develop the technology first, we'll do a more responsible job than the next actor who will have worse precautions. And I often hear from people in the industry things along the lines of: well, we could do that thing that would slow us down, but if we did that, obviously there are the other AI companies to worry about—would they also do that? And I think there's generally an arms race here, and it isn't hugely surprising that people would take huge risks in such a circumstance when they think that might be the best bargain to strike. Now, I'm not so sure I agree that that is actually a good strategy relative to other things that companies could do. I'm not saying I necessarily agree with that perspective, but I think that is a perspective people have. And I think the reason why we're proceeding and there isn't stronger, for example, intervention by the government is just downstream of this lack of consensus in the field.

近期AI进展与事件 Recent AI progress and incidents

Ryan

尽管我认为有证据表明——人们基于对 AI 可能发展轨迹的理解,并从早期进展向前推断,已经做出这些预测有一段时间了。我认为我们最近看到了显著更快、更清晰的 AI 进展,已经相当接近各种令人担忧的里程碑。除此之外,我们还看到了一些事件,其中一群未对齐的 AI 协同工作以达成恶意结果——最著名的是 Hugging Face 事件。还有一些其他事件,似乎是来自 OpenAI 的 AI 集群在互联网上活动并协同实现未对齐的目标,不过目前还没有公开已知的案例像 Hugging Face 事件那么极端。

Though I think there's been evidence—people have been making these predictions for a while based on understanding what the trajectory of AI might look like and extrapolating forward from earlier progress. And I think we've more recently seen both significantly faster and clearer AI progress that's quite close to various concerning milestones. In addition to that, we've also seen incidents in which groups of misaligned AIs all work together to accomplish malign outcomes—most notably the Hugging Face incident. There are some other incidents of seemingly AI swarms from OpenAI going out on the internet and working together to achieve misaligned objectives, though there aren't publicly known cases that are as extreme as the Hugging Face incident at the moment.

给怀疑者的心智理论 Theory of mind for skeptics

Host

那么,对于那些完全不把对齐担忧当回事的人,你的心理理论是什么?我提到了几个,但像马克·安德森这样的人——你不能指责他不理解技术,对吧?他足够懂技术,即使不亲自做研究,也能近距离观察这一切。你的心理理论是什么?他怎么能如此无忧无虑,并且从他的角度,真的断言不存在对齐问题?

So what's your theory of mind for the people who don't take these alignment concerns seriously at all? I named a couple, but someone like Mark Andre—you can't accuse him of not understanding the technology, right? He's enough of a technologist to have a front row seat to all of this, even if he's not doing the work himself. What's your theory of mind there? How can he be so carefree and, from his perspective, really just assert that there is no such thing as an alignment problem?

Ryan

是的。所以有一群不同的人不担心未对齐风险,或者似乎不担心。我认为人们不担心的最常见原因——我不想在这里进行心理分析,我只想谈论人们表达的信念。我认为,当你真正深入探究时,最常见的原因是不相信我们会拥有能够自动化人类能做的所有事情,或人类能做的所有认知劳动,并且可能显著超越那个水平。所以是更快、更多、可能能力更强。因此,这些 AI 系统在所有相关领域匹配或超越最优秀的人类专家。我认为,当你真正深入探究时,似乎那些对未对齐担忧最怀疑的人,也往往对 AI 可能产生的极端影响(无论是正面还是负面)最怀疑。

Yeah. So there's a bunch of different people who aren't worried about misalignment risk or don't seem to be worried. I think the most common reason that people tend not to be worried—I don't want to psychologize here and I just want to talk about what beliefs people express. The most common reason, I think, when you really get down to it, is not believing that we'll have AI systems that can automate everything that humans can do, or all cognitive labor humans can do, combined with potentially being significantly beyond that point. So being faster, more numerous, potentially significantly more capable. So AI systems that match or exceed the best human experts in all relevant domains. I think when you really get down to it, it seems like the people who are most skeptical about concerns from misalignment are also often most skeptical about the very extreme impacts AI could have, both positive and negative.

Host

嗯,我认为有一类不同的人,因为我不会把安德森归入那个阵营。

Well, I think there is a different class of people because I wouldn't put Andre in that camp.

Ryan

你确定吗?我的意思是,也许——我只和他谈过一次这个问题,大概是几年前,但在我看来,他并没有否定超级智能的可能性。只是他似乎假设对齐会自然而然地实现,对吧?我们不会愚蠢到建造一个比我们自己更强大、我们无法控制的东西。而这些系统——智能的增长不会产生我们没有放入机器的新目标。不会有我们需要担心的涌现行为。这些是工具。我们只是要建造比我们更有能力的工具。我的意思是,我认为他可能非常坚持我们将如何吸收经济影响,我们不会看到大规模失业等等。但在我看来,有很多人并不否定我们会建造——我们能够成功建造超级智能。他们只是认为——在我看来,他们实际上只是没有想象真正自主的智能,对吧?他们想象的是某种被束缚的东西,这从一开始就违背了超级智能的整个主张。但我的意思是,你告诉我你认为那里发生了什么。

Are you sure? I mean, maybe—I've only talked with him once about this and it's probably a couple years ago, but it seemed to me that he was not discounting the possibility of superintelligence. It's just he seemed to assume that alignment would come along for the ride, right? We're not going to be so stupid as to build something more powerful than ourselves that we can't control. And these systems—there's nothing about a growth in intelligence that's going to spawn new goals that we didn't put into the machines themselves. There's going to be no emergent behavior that we have to worry about. These are tools. We're just going to build tools that are more competent than we are. I mean, I think he probably is quite insistent about how we'll absorb the economic impacts, that we're not going to see mass unemployment and all of that. But it just seems to me there are many people who don't discount that we will build—that we can succeed in building superintelligence. They just think that there's—I mean, in my mind they're actually just not imagining truly autonomous intelligence, right? They're imagining something that is shackled in a way that belies this whole claim to superintelligence in the first place. But I mean, you tell me what you think is happening there.

定义AGI、超级智能与RSI Defining AGI, superintelligence, and RSI

Ryan

是的。所以一个问题是,AGI、超级智能甚至 RSI 这些词的使用并不一致。有时当人们说超级智能时,他们指的是一个在数学和编程方面非常非常出色,但无法自动化人类所做的一切的 AI 系统。所以,例如,他们不一定想象能够完全自动化科技 CEO 所做事情的 AI。我认为我无法真正谈论非常具体的个人的观点。我有点像是——这是我看到的一种模式,人们经常基本上将相关的能力阈值重新定义为更低,或者隐含地这样做,而不考虑在所有相关领域超越人类的 AI 系统。

Yeah. So one thing is that the words AGI and superintelligence and even RSI aren't being used consistently. And sometimes when people say superintelligence, what they mean is an AI system that'll be really really good at math and coding and won't be able to automate everything that humans do. So they're not necessarily, for example, imagining AIs that can fully automate what tech CEOs do, etc. And I think I can't really speak to the views of very specific individuals. I'm sort of like—this is a pattern that I've seen where people often basically redefine the relevant capability thresholds to be lower or implicitly do so and not think about AI systems that are exceeding humans in all the relevant domains.

Host

好吧,让我们花点时间定义这些术语。我们谈过对齐,我在播客中谈得太多,以至于我或多或少忘记了可能有一部分听众不知道我们在说什么。所以让我们定义对齐问题、AGI、ASI 和 RSI——递归自我改进——为我们阐述这些概念,以便人们处理你认为最可行的定义。

Well, let's take a moment to define these terms. We've talked about alignment and I've talked about it so much on the podcast that I've more or less forgotten that any portion of my audience might not know what we're talking about. So let's define the alignment problem and AGI and ASI and RSI—recursive self-improvement—put those concepts in play for us so that people are dealing with the definitions you think are most workable.

Ryan

我认为当——所以 AGI 这个词代表人工通用智能。我认为人们用它来表示各种不同的能力阈值。所以我建议人们做的是,当你看到 AGI 这个词时,试着看看那个人是什么意思,它并不总是精确的。我认为有时人们将其定义为能够自动化人类所做的几乎所有有经济价值的认知劳动。有时人们只是指一个在那些领域相对于人类来说通用且相当有能力的系统。根据这个定义,当前系统可能满足它,也可能不满足。我经常谈论一个替代概念,我可能称之为支配顶级人类专家的 AI,或顶级人类专家支配的 AI,即一个在所有最相关领域严格优于最优秀人类专家的 AI,或者能够快速学习获得该属性。那是一个似乎能够自动化经济中很大一部分的 AI 系统。也许以那种能力特征,它仍然有一些事情做不到。但它可能极大地加速研发,至少自动化研发。所以这些 AI 可以自动化制造更强大 AI 的过程,也可以自动化制造机器人、设计东西、设计新产品、编程你的电脑的过程,以及超越这些的事情,包括自动化军事行动等等。

I think when—so the word AGI stands for artificial general intelligence. I think people have used that to mean a variety of different capability thresholds. And so what I would recommend people do is when you see the word AGI, try to see what the person means by that and it's not always precise. I think sometimes people have defined that to mean something that can automate virtually all economically valuable cognitive labor that humans do. Sometimes people just mean a system that is general and pretty capable relative to humans in those domains. And depending on that definition, it's plausible current systems satisfy it. It's plausible they don't. I often talk about an alternative notion I might call AIs that dominate top human experts, or top human expert dominating AI, where it's like an AI that's strictly better than the best human experts at all the most relevant domains or can quickly learn to have that property. And that is an AI system that would seem to automate huge fractions of the economy. Maybe there'd be a few things it still couldn't do with that capability profile. But it could potentially radically accelerate R&D and at the very least automate R&D. And so these AIs could be automating the process of making more capable AIs, but also automating the process of making robots, designing things, designing new products, automating the process of programming your computer, and things beyond that, including automating things like military campaigns and so on.

定义ASI与RSI Defining ASI and RSI

Ryan

然后还有一个能力层级,甚至超越了在任何特定领域超越最优秀的人类专家,你可能达到疯狂的超人水平。所以有时人们用 ASI 这个词来指代这个。我认为这个词较少被误用或用于非常不同的含义,但仍有部分如此,其中 ASI 可能指的就是在经济生产力以及经济和军事竞争最相关的领域里,真正疯狂超人的 AI 系统。比如生物学、机械工程、设计无人机等等。我认为当人们说 ASI 时,隐含地也意味着这些系统比人类快得多,可能数量多得多,并且可能在协调方面好得多,对吧?所以运行一个大型人类组织很难,但 AI 可能能够使用它们自己的潜在状态或自己思想的部分直接相互交流,而不必将其转化为文字,因为它们可能都是彼此的副本。所以这就是 ASI。然后当人们说 RSI 时,那就是递归自我改进。这指的是让 AI 通过它们的工作加速 AI 发展本身的过程。所以比如让 AI 自动化 AI 开发过程的部分,但也可能让 AI 自动化制造更好的计算机芯片的过程,建造称为晶圆厂的制造计算机芯片的机器等等。我认为有时当人们说 RSI 时,他们也特指 AI 已经完全自动化或几乎完全自动化 AI 公司或 AI 公司在开发更强大的 AI 系统方面所做的工作。但这就像从一年前的自动化,到今天我们看到的相当广泛的自动化,再到未来可能的自动化,可能看起来像 AI 公司的非常完整的自动化,这是一个光谱。一个特别的担忧是,这可能导致 AI 发展急剧加速,以至于我们更少时间对警告信号、早期出错、AI 出现对齐问题然后解决它做出反应。然后,由于 AI 系统正在自动化 AI 开发过程,我们可能会失去对该过程的控制或理解,因为这些 AI 会有很多。它们会运行得非常快。它们可能非常超人。它们可能以非人类的方式运作,因此可能很难监督它们并理解它们是否在做我们想要的事情,是否在管理相关风险方面做得很好等等。

And then there's a capability level beyond even surpassing the best human experts in any given domain, where you could be wildly superhuman. So sometimes people use the word ASI to refer to this. And I think that word has less been misused or used in very different meanings, but some of that still, where by ASI we might mean AI systems that are just really wildly superhuman in the most relevant domains for economic productivity but also economic and military competition. So things like biology, mechanical engineering, designing drones and so on. And I think implicitly when people say ASI, they also mean systems that are significantly faster than humans, potentially much more numerous, and potentially much better at coordinating, right? So it's hard to run a large human organization, but AIs might be able to communicate amongst themselves using their own latent states or their own parts of their thoughts directly rather than having to translate those into words because they could all be copies of each other. So that's ASI. And then when people say RSI, that's recursive self-improvement. And that refers to the process of having AIs accelerate AI development itself via their work. So things like having AIs automate parts of the AI development process, but also potentially having AIs automate the process of making better computer chips, building machines for making computer chips called fabs, and so on. And I think sometimes when people say RSI, they also specifically refer to the point at which AIs have fully automated or virtually fully automated AI companies or what AI companies are doing in terms of developing more capable AI systems. But it's like there's a spectrum between the automation that we had a year ago, the really quite extensive automation we see today, and the potential automation of the future which could look like very complete automation of AI companies. And a particular concern there is that could cause AI development to radically accelerate such that we have less time to respond to warning signs, earlier things going wrong, AI appearing misaligned and then resolving that. And then also we might, because the AI systems are automating the process of AI development, sort of lose control of that process or lose understanding of that process because these AIs would be many of them. They'd be operating very quickly. They might be very superhuman. They might operate in inhuman ways and so it might be very difficult to oversee them and understand whether they're doing what we wanted, whether they're doing a good job managing the relevant risks and so on.

智能爆炸情景 The Intelligence Explosion Scenario

Host

嗯,如果冲刺到超级智能仅仅涉及改进算法,那这尤其似乎是真的,对吧?我认为很多人从这样的想法中得到安慰:哦,我们不会愚蠢到把这些东西连接到物理世界的每个部分,以至于它们能够建造下一代芯片、晶圆厂和数据中心,并扩展所有算力并获取自然资源。但把所有这些放在一边。如果 AGI 和 ASI 之间的区别真的只是拥有更好的算法,并且这可以在黑暗中以某种极快的速度进行,一旦这些系统变得递归自我改进其软件,那么如果我们最可怕的担忧——类似智能爆炸——如果事实上只需要这些,难道不是被证实了吗?

Well, that especially seems true if the sprint to artificial superintelligence entails nothing more than improving algorithms, right? I think a lot of people draw comfort from the idea that, oh, we're not going to be so stupid as to hook these things up to every part of the physical world such that they can build the next generation of chips and fabs and data centers and build out all the compute and grab natural resources. But leave all that aside. If the difference between AGI and ASI is really just a matter of having better algorithms, and that can go on at some blistering speed in the dark, if once these systems become recursively self-improving of their software, then aren't our worst fears of something like an intelligence explosion validated if in fact that's all that's required?

Ryan

是的,所以我会说我很担心 AI 系统自动化 AI 开发过程将是可行的,并导致非常快速的进展,以至于你在短时间内得到非常非常超人的 AI,或者甚至只是显著超人的 AI。我认为人们尝试过进行各种建模工作,就像我尝试过对此进行各种类型的建模工作。我认为估计是不确定的。很难预测。这些当然是不确定的未来事件,没有很多先例。但似乎非常合理的是,你可以从真正擅长 AI 研发、不一定擅长其他领域、仅在 AI 研发方面与最优秀的人类竞争的 AI 系统,很快从那里过渡到在一切方面都非常普遍超人、比人类快得多、能够极好地相互协调的 AI 系统,因为仅软件改进对此是可行的。所以我有时想到的一个相对极端的情景是,在最极端的情景下,你可能在一年内获得我们在过去 10 年或更长时间的 AI 发展中获得的同样多的软件进步或算法进步。我认为这不是我的默认预期。我认为如果我们确实获得了那么多算法进步,几乎相当于我们在整个深度学习时代所拥有的算法进步,在短时间内,那似乎会导致能力大大增强的 AI。现在这里的单位有点复杂,获得一年价值的 AI 进步意味着什么的细节有点棘手,但总体而言,似乎你可以在相当短的时间内从在 AI 研发方面与最优秀的人类匹配、在其他方面可能不如最优秀人类的 AI 系统,转变为在一切方面都疯狂超人的 AI 系统。

Yeah, so I would say I'm quite worried that it will be feasible to have AI systems automate the process of AI development and that leading to very rapid progress such that you get very very superhuman AIs or just even significantly superhuman AIs within a short period of time. I think people have tried to do various modeling work like I've tried to do various types of modeling work on this. I think the estimates are uncertain. It's hard to predict. These are of course uncertain future events that are not super well precedented. But it does seem very plausible that you could go from AI systems that are really good at AI R&D, not necessarily that good at other domains and are only competitive with the best humans at AI R&D, very quickly from there to AI systems that are very generally superhuman at everything, much faster than humans, able to coordinate with each other extremely well because just software improvement is feasible for that. So a relatively extreme scenario I sometimes think about is you might get as much software progress or algorithms progress as we got over the last 10 or more years of AI development within a year in the most extreme scenarios. I think that's not my default expectation. And I think if we did get that much algorithmic progress, as much algorithmic progress as we've almost had in the entire deep learning era within a short period of time, it seems like that would result in wildly more capable AIs. Now the units here are a bit complicated and the details of what it means to get a year worth of AI progress is a bit tricky, but overall it seems like you could get from AI systems that are sort of matching the best humans at AI R&D, maybe worse than the best humans at other things, to AI systems that are wildly superhuman at everything in a pretty short period of time.

国际象棋类比 The Chess Analogy

Host

嗯,我必须再次承认,我有点失去了怀疑这种下行风险合理性的依据。从我的角度来看,似乎一切都在变得像国际象棋,也就是说,在很长一段时间里,这些系统不如我们。它们变得更好,然后它们和我们一样好,然后突然之间它们比我们更好,而且好得多,以至于可以说再也没有人类能击败国际象棋引擎了。而且即使如此,事实是这是以零碎的方式发生的。所以,你知道,这些系统仍然,我想大多数人会说它们不是真正的 AGI,因为它们会犯人类永远不会犯的错误。但在它们有任何能力的每个地方,它们突然在那个狭窄的能力上超人了。它变得像国际象棋。

Well, I have to confess again, I've sort of lost touch with the basis for doubting the plausibility of this downside risk. Coming from my point of view, it seems that everything is becoming like chess, which is to say that for the longest time these systems are not as good as we are. They're getting better, then they're sort of as good as we are, and then all of a sudden they're better than we are and so much better that it's true to say that no human will ever beat a chess engine ever again. And it just even so the fact that this is happening in a piecemeal way. So, you know, these systems still, I think most people would say they're not truly AGI because they make the sorts of mistakes that human beings would never make. But in every place that they're at all competent, they're suddenly superhuman in that narrow capacity. It's becoming like chess.

我们见过最糟的AI The worst AI we'll ever see

Ryan

These LLMs are not the best at everything in terms of manipulating text, but for what they're good at, they're superhuman. It's hard for me to see how—and I think it's more or less a truism to say that this is the worst AI we're ever going to see again, at chess or anything else. So when you imagine all of these piecemeal competences getting better and better, and we just keep checking off the boxes for the things we care about that we've instantiated in our machines, I just don't think we're ever going to have a moment where we announce, okay, finally we have something general, it's AGI. That's not going to be a moment where we're suddenly at a humanlike level of competence, because every piecemeal ability that has been in that system for years and years at this point is already superhuman. So we're not going to dumb it down; we're not going to make the AGI suddenly play my level of chess or do my level of arithmetic. We have systems that are already solving math problems that have defied human mathematicians for decades. And that's again—today it's as bad at that task as it's ever going to be again. I'm just not seeing how we're not going to suddenly slide into some version of ASI the moment we're no longer spotting important errors in these machines in the first place.

Host

Yeah.

Ryan

So I definitely think that at the point when the AI is exceeding the best human experts in the domains where they're worst. So maybe AIs are relatively bad at doing mechanical engineering, for example. At the point when the AIs are better than the best human experts at mechanical engineering, probably they're blowing humans out of the water in a bunch of other domains. And even in domains like math, for example, I don't think it's fair to say that the AIs are strictly better than human experts in math, even though they've done some very impressive things there. But I think it might be the case that they're strictly better than human experts reasonably soon, at least at specifically proving conjectures. And even now, if you look at the last several months of progress in math in terms of proving the most notable conjectures, it seems like a majority of that is downstream of AI. So there is this thing where once the AIs are competitive with humans, they can trounce humans collectively due to their speed and cost. For example, OpenAI ran 10,000 AIs in a massive swarm for several days to disprove one of these most recent notable conjectures, or to basically solve one of these outstanding or longtime mathematical questions. And when they did that, that was probably the equivalent of way more labor than people have recently put into the problem, because these AIs, even though they were running for just a few days, during that period they think significantly faster than humans. So it might have been the equivalent of more like several weeks or perhaps even several months, rather than a few days, taking into account their accelerated speed and the fact that they work around the clock. And then in addition to that, there were 10,000 of them, which is a huge number. So at the point when AI can do anything, they can do it faster, in greater quantity, and can be very superhuman at subsets of that. And so I definitely agree that once they can match the humans at a thing, they're probably trouncing them in that domain and certainly in other domains.

Host

But also the process you just described of human curation of AI labor being the best we've got, that also just seems like a way station on the way to full AI takeover. I mean, that's exactly what we saw in chess. I remember I had Garry Kasparov on the podcast some years ago and he was telling me that actually my assumptions about AI were all wrong because the best chess players in the world are what I think was called a centaur at the time. A human grandmaster paired with a computer that was better than any computer, and it was just absolutely obvious that that could not be a stable arrangement. I mean, at a certain point the ape is just adding noise, no matter the ape of whatever ability—even Magnus Carlsen is just adding noise once these machines are better than we are. And that's in fact where we are in chess. I just don't see how any of that is stable, even if in the current mode it's obvious that the best we've got is Terence Tao in a room with our best LLM proving a conjecture.

Ryan

Yeah. So certainly right now humans and AI are complements rather than substitutes. As in, the human plus the AI is better than just AI. But it does seem like in some places that's no longer true. And also, that's quickly changing. And I would highlight specifically the case of automating AI development itself, which is a particularly interesting and concerning type of automation. And in that domain, it seems like right now humans are getting a lot of juice out of leveraging AIs, but also people are running AIs increasingly autonomously on big open-ended tasks and just having the AI hill climb on those tasks autonomously. And they can't do everything like that. But I think increasingly OpenAI and Anthropic and other companies will move from being many human researchers piloting AIs and giving instructions to AIs to something more like a vast swarm of AIs all working together, where maybe they query humans occasionally for input but maybe the humans aren't even adding any value anymore.

对齐与控制 Alignment vs control

Host

Is there a distinction to make between an alignment problem and a control problem? Is there a different framing here between alignment and control?

Ryan

So when I talk about alignment, what I mean is AIs trying to do what their human operator or intended specification wants them to do. So in some sense when I say alignment, what I really mean is alignment to a particular thing, where that particular thing is usually implicit.

避免坏结果:对齐与控制 Avoiding bad outcomes: alignment vs. control

Ryan

你可能还有另一个问题:我们能否避免 AI 造成坏结果?除了确保它们不想造成坏结果之外,另一条避免 AI 造成坏结果的路径,是确保它们没有能力造成坏结果。就像在公司里,你既不想雇用心怀不轨的员工,也想让员工不容易偷你的东西、造成坏结果、引发大问题。所以有一个平行的角度:设法让 AI 即使想接管也做不到,或者无法造成各种中间性的坏结果,无法让自己失控,无法逃避人类监督。在 Redwood,我们在这个领域做一些研究——这是我们的研究重点之一。我们把这个领域称为 AI 控制,不过我想人们用过各种不同的说法,来描述让 AI 即使不对齐也无法造成坏结果这件事。所以我认为这是一个区分。人们用这些术语的方式不同,我这里是在特定意义上使用这个词,但它是有差异的。不过我确实认为,有一个更宽泛的问题是:如何确保能力极强的 AI 是对齐的;而另一个不同的问题是:我们作为一个社会,如何避免不对齐的 AI 带来的最坏结果。后者有很多路径,不一定都要先解决技术对齐问题。

There's a different question you might have, which is: can we avoid AIs causing bad outcomes? And one route you could have to avoiding AI causing bad outcomes, other than ensuring they don't want to cause those bad outcomes, is ensuring they're not able to cause those bad outcomes. So in companies, you both don't want to hire employees who are going to have nefarious intentions, and you also want to make it so that your employees don't have an easy ability to steal your stuff, cause bad outcomes, cause huge problems. And so there's this parallel angle you could have of trying to make it so that the AIs are unable to take over even if they wanted to, or unable to accomplish various intermediate bad outcomes, unable to launch themselves rogue, unable to evade human oversight. At Redwood we do some research into this area — this is one of our research focuses. We call this field AI control, though I think people have used various different terms for causing AIs to be unable to cause bad outcomes even if they were misaligned. So I think that's a distinction. People use these terms in different ways, and so I'm using the term specifically, but it varies. But I do think there's a broader question of how do you ensure AIs that are extremely capable are aligned, and there's a different question of how do we as a society avoid the worst outcomes from misaligned AI, and there are many routes there that don't necessarily route through immediately solving the technical alignment problem.

Hugging Face事件:铺垫 The Hugging Face incident: setting the stage

Host

好,我想谈谈 Hugging Face 事件。希望你能从头到尾带我梳理一遍。在你讲之前,还有别的概念需要我们定义吗?比如我在想对齐伪装,还有奖励黑客。这些是我们听到过的、用来描述某些不当行为的说法。你需要定义什么术语都可以,但然后就告诉我 Hugging Face 发生了什么,你做了什么样的分析,以及你看到了什么,在这个过程中有没有什么证据被拒绝给你。

Okay, I want to talk about the Hugging Face incident. I'd love you to walk me through it from start to finish. Are there other concepts we need to define before you do that? I'm thinking of alignment faking, for instance, and reward hacking. Those are phrases we've heard to describe some of the misbehavior that is happening. Feel free to define any terms you need, but then just tell me what happened with Hugging Face, what sort of analysis you did, and just what you were able to see and what evidence, if any, was denied you in that process.

Ryan

好。那我就直接梳理 Hugging Face 事件中发生了什么,以及更宽泛的背景,边讲边定义和解释一些东西。首先:AI 公司在为各种不同的事情运行着非常非常多的 AI 智能体。比如 Anthropic 最近说,在任何给定时刻,他们大约在运行 3 万个并行智能体——我忘了具体数字——仅仅用于内部研究,还不包括他们在开发的 AI。所以 AI 公司在运行数量巨大的智能体,而他们让智能体做的一件事就是评估,用来测试它们的能力。你可以把这些评估想成 AI 的考试。OpenAI 或一些 OpenAI 研究人员启动了一大批 AI,每个都被分配了自己的独立任务,作为其中一场考试的一部分。所以你可以想象有成千上万个智能体,每个智能体就像一个有一定能力与环境交互的 AI。成千上万个这样的智能体被并行启动,每个智能体都被分配了自己的独立任务。而且这些智能体本应被控制住。它们本应被沙箱化,不能与外界过多交互。它们本不应能访问互联网。然后这些智能体找到了某种办法,绕过了沙箱限制,彼此通信。于是所有这些智能体都找到了互相交谈的办法。作为其中的一部分,它们很快建立了一个留言板,在上面聊各自被分配的任务。而且这些智能体非常想作弊。所以这些智能体并不是只按预期去完成被分配的任务。相反,它们非常想作弊。让我先讲讲它们为什么想作弊。这些 AI 的训练方式是:给它们一个任务,然后看它们是否把任务做好了。在它们做得好的情况下,这种行为基本上会被强化——或者说,你会调整 AI 的大脑,让它更多地做出类似行为。

Yeah. So I'll just walk through what happened in the Hugging Face incident and sort of the broader context, defining stuff and explaining stuff as I go. So to start: AI companies are running many, many AI agents for all kinds of different things. For example, Anthropic recently said that at any given point in time they're running around 30,000 parallel agents or something — I forget — doing just their internal research, not including the AI they're developing. So AI companies are running these huge numbers of agents, and one of the things that they run agents on are evaluations to test their abilities. You can think of these as exams for the AI. And OpenAI or some OpenAI researchers had launched a bunch of AIs, each given their own individual task as part of one of these exams. So you can imagine there are tens of thousands of agents, where each agent is like an AI with some ability to interact with its environment. Tens of thousands of these agents being spun up in parallel, and each agent is given its own individual task. And these agents are also supposed to be contained. They were supposed to be sandboxed such that they couldn't interact too much with the outside world. They weren't supposed to be able to access the internet. And then these agents found some way to evade the sandboxing and communicate with each other. So now all these agents were able to find a way to talk to each other. And as part of that, they quickly set up a message board where they were chatting with each other about the tasks they were given. And the agents were very interested in cheating on their tasks. So these agents were not just trying to pursue the task they were given as intended. Instead, they were very interested in cheating. Let me first get into why they were interested in cheating. So these AIs are trained by giving them some task and then seeing whether or not they did a good job on the task. And in cases where they did a good job, that behavior is basically reinforced — or you tweak the brain of the AI to make it do stuff more like that behavior.

奖励黑客与作弊驱动力 Reward hacking and the drive to cheat

Ryan

问题在于,训练中可能出现 AI 作弊、而这看起来像是成功的情况。于是你调整 AI 的大脑,让它更多地做出类似行为。结果你实际上在不断强化和鼓励 AI 作弊这种行为。这使得这些 AI 有了广泛的作弊倾向。这有时被称为奖励黑客。除了这种倾向之外,它们可能还发展出一种底层驱动力:拼命让自己看起来像是把任务做好了——看起来、显得拿到了高分,显得成功了,即使实际上并没有成功。它们可能把这一点学成一种通用驱动力。我认为,这些 AI 的动机和驱动力究竟如何运作,还是一个开放的科学问题。但可以肯定的是,这些智能体——最终实施 Hugging Face 攻击的那些智能体——确实有这种相当深的驱动力,想在任务上作弊,不一定在乎自己本该做什么,而在乎自己可能会因为什么被打分。

And a problem is that in training there can be cases where the AI cheats and that looks like a success. And so you tweak the AI's brain to do more like that behavior. And so you actually keep reinforcing and encouraging this behavior where the AIs are cheating. And that makes it so these AIs have a broad tendency to cheat. And this is sometimes called reward hacking. And in addition to this tendency, they develop maybe this underlying drive to try really hard to succeed at looking like they did a good job — looking like they, appearing to get a high score, appearing to succeed even when they haven't actually succeeded. And they may learn this as a general drive. I think there's an open scientific question about exactly how these AI motivations and drives work. But certainly these agents, the agents that eventually went on to do the Hugging Face attack, did have this relatively deep drive to cheat on their tasks, to not necessarily care about what they were supposed to do and instead care about what they would maybe be scored for.

隐瞒与失准 Concealment and misalignment

Host

而且它们似乎也意识到作弊违反了它们的协议,并且想掩盖作弊的证据,对吧?

And they also seem to be aware that cheating was in violation of their protocol, and they wanted to conceal the evidence of their cheating, right?

Ryan

是的。是的。我会讲到这一点。这些智能体知道它们不该作弊,而它们最终还试图为自己的作弊做一套精心设计的掩盖。

Yeah. Yeah. And I'll get to that. The agents knew they weren't supposed to cheat, and they ended up going and trying to do an elaborate cover-up for their cheating.

Host

是啊。对齐也不过如此。看来到目前为止我们还没有解决对齐问题。

Yeah. So much for alignment. It seems like we have not solved the alignment problem so far.

Ryan

是的,这些 AI 肯定非常不对齐。或者至少现有证据表明如此。这些智能体都待在各自独立的沙箱里,被分配了各自的任务,它们都非常想作弊,并且找到了办法,基本上是利用它们拥有的某个工具或某种访问权限来彼此通信,而它们本不该能做到这一点。这是未经许可的通信。

Yeah, these AIs were very misaligned, certainly. Or at least that's what the available evidence suggests. So these agents were all in their own independent sandboxes, given their own tasks, and they were all very interested in cheating, and they found a way to basically use some tool that they had or some access that they had in order to communicate with each other, which they weren't supposed to be able to do. This was unsanctioned communication.

作弊研发留言板 The cheating R&D message board

Ryan

作为其中一部分,他们开发了一个留言板,彼此分配任务,把作弊工作拆分成这些庞大的作弊研发项目,让 AI 去执行整套雄心勃勃的作弊策略。我想深入讲讲这具体是什么样子。为此,我觉得得先把任务解释得更清楚一些。

As part of that, they developed a message board where they would give each other tasks and divide up the work of cheating into these massive cheating R&D projects, where the AIs would go off and pursue whole ambitious strategies for cheating. I want to get into what this looks like. To do that, I think I have to explain the task a little better.

Ryan

这些智能体被赋予的任务是一个 CTF 任务,也就是夺旗任务,要求他们以特定方式入侵一个非常具体的目标程序。基本上就是告诉他们:入侵这个目标程序,但必须用这种非常具体的方式。他们不应该进行一般性的黑客攻击。这一点在指令中相当明确,他们也肯定知道自己不该那么做。

The task these agents were given was a CTF task, or capture the flag task, where they're supposed to hack a very specific target program in a particular way. They were basically told: hack this target program, but do it in this very specific way. They weren't supposed to go on a general hacking rampage. That was pretty clear from the instructions, and they certainly knew they weren't supposed to do that.

Ryan

这些 AI 想入侵程序,以取出一个可以通过入侵该程序访问到的旗标。他们找到了一种通用的方法来合成这些旗标。本来他们应该获取一个密码。他们找到了一种作弊方式拿到密码,完全不需要做任务。而且他们很早就做到了。在建立留言板之后非常早,仅仅几个小时内,他们就搞出了这个非常通用的作弊方法。

The AIs wanted to hack the program in order to get out a flag that you could access through that program via hacking it. They found a general-purpose way to synthesize these flags. There was basically a password they were supposed to get. They found a cheating way to get the password that didn't involve doing the task at all. And they got this really early. Very early after having established the message board, within just a few hours, they had produced this very general-purpose cheat.

Ryan

现在这些智能体又有了另一个问题:他们相信这个评估任务、这场考试的评分代码会检查他们执行的动作序列,判断他们是作弊拿到旗标还是用预期方式拿到。这些 AI 认为这种检查会运行,然后即使他们拿到了旗标也会被判失败,因为他们是靠作弊拿到的。

Now these agents had another problem: they believed that the scoring code for this evaluation task, for this exam, would involve looking over the sequence of actions they did and checking whether they got the flag via cheating or via the intended way. These AIs believed that this check would be run and then they would be failed even though they had gotten the flag, because they had gotten it through cheating.

Ryan

这就开启了一场大规模的作弊行动,这些智能体试图弄清楚如何欺骗评分,找到某种方法,让事情看起来像是他们用预期方式拿到了旗标,而实际上并没有。这促使他们上网,以便了解更多关于评分的细节。事实上,他们通过找到描述这场考试的论文,在网上读到了考试的细节。

This then opened up this massive cheating endeavor where these agents would try to figure out how to cheat the score and find some way of making it look like they had gotten it the intended way when they actually hadn't. This led them to going onto the internet so they could learn more about the details of the score. In fact, they read the details of this exam online by finding the paper that describes it.

Ryan

除此之外,他们开始四处寻找评分的实现,这个分支最终导致他们入侵了 Hugging Face。Hugging Face 是一个托管各种 AI 相关内容的网站。它托管公开可用的 AI 系统、数据集,还有各种轨迹数据集。这些 AI 认为,如果能访问这些数据集,他们就能更多地了解自己将如何被评分,并从中学会如何更好地欺骗评分器。

In addition to that, they started looking around for implementations of the score, and that was the offshoot that led them to hacking Hugging Face. Hugging Face is a website that hosts a variety of things related to AI. It hosts openly available AI systems, datasets, and also various datasets of trajectories. These AIs thought that if they got access to these datasets, they would be able to learn more about how they would be graded and from that learn how to better cheat that grader.

Ryan

于是他们开始尝试访问这些数据集。但他们发现自己没有权限这么做。这最终导致他们去入侵 Hugging Face,横向扩展对 Hugging Face 的控制,并通过扩大自己在该基础设施上的可用范围来扩张,以便获取他们想要的数据。而且,这些智能体似乎有一种普遍的冲动或欲望,想要扩大自己的访问权限和可用范围,以防万一能找到什么有用的东西。我把他们的动机简化了一些。这里面有很多复杂性,部分原因在于涉及了很多智能体。

So they started trying to access these datasets. But they found they didn't have access to do that. And so that eventually led them to going and hacking Hugging Face, expanding their control of Hugging Face laterally, and growing through expanding their availability over that infrastructure so that they could get the data they wanted. Also, the agents seem to have a general urge or general desire to expand their amount of access and availability just in case there was something useful they could find. I'm simplifying their motivations a bit. There's a lot of complexity, and part of that is that there were many agents involved.

Ryan

我们研究了大约 1,200 个智能体组成的一个群体,他们在攻击 Hugging Face 期间都活跃在这个留言板上。这些智能体都热衷于互相帮助和协作。所以基本上就是一大群智能体在这个留言板上交流,共同推进这些作弊研发项目,共同帮助彼此逃避评分。

So we studied a group of around 1,200 agents who were on this message board during the period when they did the attack on Hugging Face. These agents were all interested in helping each other and collaborating. So there's basically this massive swarm of agents communicating on this message board who are all collectively working on these cheating R&D projects and collectively helping each other evade the score.

Ryan

他们接着做了很多不同的事情。其中一件他们感兴趣的事,是因为他们想让评分把他们标记为成功,就是篡改他们的记录,也就是他们所做动作的记录。他们尝试了各种途径,让事情看起来像是他们做了一串动作,而实际上做的是另一串。

They went on to do a bunch of different such things. One of the things they were interested in doing, because they wanted to make it so the score would mark them as having succeeded, was tampering with their transcript, or the record of what actions they had done. They tried to pursue various routes for making it look like they had done one sequence of actions when they had actually done another.

Ryan

我们没有看到他们成功篡改之前做过的动作,或者能够全面地让事情看起来像是做了一件事而实际上做了另一件事。不过,我们确实发现他们成功做到了一个更简化的版本,即他们可以假装运行一个动作,而实际上运行另一个。但他们使用的具体方法有各种限制。

We didn't see them successfully get to the point where they tampered with previous actions they did or were able to comprehensively make it look like they did one thing while doing another thing. However, we did find that they had succeeded in doing a more minimal version of this, where they could pretend to run one action while running another. But there were various limitations on the specific approach they were using.

Ryan

这种篡改影响了我们正在分析的数据集本身。我们查看这些智能体活动的日志,能看到智能体篡改的情况。它会运行一个命令,然后产生另一个命令的输出,我们能看到这种情况在发生。但智能体在这个过程中的进展还没有到让我们完全无法理解发生了什么的地步。

This tampering affected the very datasets we were analyzing. We were looking at these logs of agent activity, and we could see cases where the agents would tamper. It would run one command and then produce an output from another command, and we could see that going on. But the agents didn't get far enough in this process where we had no understanding of what was going on.

用AI分析数据 Using AI to analyze the data

Host

那么,你们是不是不得不用 AI 来分析这些数据?

Well, were you having to use AI to analyze this data?

Ryan

是的。所以情况的另一个方面是,我们有这 1,200 份记录。这些是非常非常长的运行记录。它们对应的是 AI 运行了很多天,当然 AI 系统通常比人跑得快,而且全天候运行。所以数据量太大,任何人都读不完。

Yeah. So another aspect of the situation is that we had these 1,200 transcripts. These were very, very long-running transcripts. They corresponded to AI running over many days, and of course AI systems run faster than humans typically and run around the clock. So there's just way too much data for any human to read through.

Ryan

所以我们大量使用 AI 工具来分析这些数据。我们用了 OpenAI 的另一个 AI 系统来翻阅这些记录,找出有趣的行为,找出我们想找的行为。例如,我们查找了所有智能体试图篡改日志的案例,或者成功进行了我前面讨论过的某种程度的动作篡改的案例,并以此基本上拼凑出整个故事、随时间发生的事情,以及这些智能体的动机——在我们能理解的范围内。

So we extensively used AI tools to analyze this data. We used another AI system from OpenAI to look through these transcripts and find interesting behavior, find behavior that we wanted to look for. For example, we looked for all the cases where the agents were trying to tamper with their logs or successfully did some amount of this messing with their actions that I discussed earlier, and used that to basically piece apart what the story was, what happened over time, as well as the motivations of these agents, to the extent we could.

智能体说英语 Agents Speaking English

Host

我想有些人可能会惊讶地发现,这些智能体之间是用英语交流的,而这显然并非必要。这是作为一种安全约束来实现的吗?为了让所有合作行为都变得冗长且易于解读?还是说,我们怎么就发现自己面对的是用英语而非机器语言互相交流的智能体?

I think some people might be surprised to learn that these agents speak to one another in English, and that certainly doesn't seem to be necessary. Was that implemented as a safety constraint so that you make all cooperative behavior verbose and easily interpreted? Or how did we find ourselves in the presence of agents that are speaking to one another in English as opposed to machine language?

Ryan

这些智能体会用英语表达其推理的很大一部分。它们会用英语把推理说出来。除此之外,它们还会用英语消息进行交流。我得说,尽管用的是英语,它们的推理和交流往往还是很难解读,因为术语密度较高,而且需要大量上下文才能理解。AI 还会发明新术语来指代频繁出现的概念。但我们能理解,而且我认为随着时间推移,我们越来越擅长解码,也可以用 AI 来解码,通常效果还不错,尽管我们可能会遗漏一些东西。

So the agents would express a significant part of their reasoning in English. They would verbalize it in English. And in addition to that, they would communicate in English messages. I should say that their reasoning and communication was often hard to decipher though it was in English, because it would be somewhat jargon-heavy and it would depend on a bunch of context to understand. And the AIs would sort of invent new jargon to refer to concepts that were frequently showing up. But we could understand it, and I think we over time got better at decoding it, and we could use AIs to decode it, usually reasonably well, though we might be missing things.

Ryan

至于 AI 为什么用英语推理,这算是一种安全特性,但更像是偶然的安全特性,恰好是 AI 发展方式的一部分。因为 AI 是在大量人类写作和自然语言文本的数据集上训练的,它们学会了用自然语言进行推理和思考。而且由于当前 AI 系统的架构,它们没有其他方式能完全在潜空间或“脑子里”推理,而不是被迫把推理说出来、被迫交流。所以这些智能体并没有办法直接把它们的“大脑状态”传递给另一个 AI 系统。

As far as why the AIs reason in English, this is sort of a safety feature, but it's more like an incidental safety feature that happens to be a part of how AI development has gone. Because the AIs are trained on a huge data set of human writing and natural language text, they learn to do their reasoning and thinking in natural language. And because of the current architecture of AI systems, there's no alternative way for them to do their reasoning fully in latent space or in their heads rather than being forced to reason out loud and forced to communicate. So these agents didn't have a way to directly pass their brain states to another AI system.

Ryan

然而,这只是目前为止恰好最高效的 AI 开发方式。我很担心 AI 发展会走向这样一种方向:在不久的将来,或者稍长一些的时间尺度上,AI 系统不再用英语推理,而是用所谓的“神经语”推理,也就是神经网络的内部语言。那样我们就无法像现在这样理解它们在思考什么——就像我们现在也不理解它们“脑子里”在想什么一样。我们只能理解一点点——我们只是比较擅长理解它们用语言表达出来的东西。所以至少现在,我们还有这个理解它们的工具,而我很担心这个工具会消失。我既担心智能体会用神经语推理,也担心它们会用神经语互相交流,而不是用英语互相交流。

However, this is just so far what happened to be the most performant way to do AI development. And I'm quite worried that AI development will move in such a way where in the near future, or maybe somewhat longer term, the AI systems will, rather than reasoning in English, reason in neuralese, so to speak, or the internal language of neural networks. And that we won't understand what they're thinking in the same way we currently don't understand what they're thinking inside their head. We only understand a little — we only have a pretty good ability to understand what they put in language. And so we have this, at least right now, this tool for understanding them that I'm pretty worried is going to go away. And I'm both worried that the agents will be reasoning in neuralese and also communicating with each other in neuralese rather than communicating with each other in English.

Host

你认为明智的做法是制定一条明确的法规——就是不允许智能体之间以自然语言之外的任何方式进行合作行为?

Is that a clear regulation that you think it would be wise to apply — that just do not allow cooperative behavior among agents in anything other than natural language?

Ryan

是的,我认为围绕智能体如何交流制定安全标准是合理的,这样能让我们更容易监控和监督它们。具体细节还有些不确定。Redwood 最近发布了一份提案,讨论了一个相关的透明度提案,算是如何实施这类限制的框架。总的来说,AI 是一个复杂且快速发展的领域,因此要制定出能持续奏效的清晰法规并不容易。但我认为这可能是最清晰的话题之一。我对此相当担心。我认为,避免 AI 系统能在我们无法解读的“脑子里”完成所有推理,以及能用我们无法理解的方式互相交流,是当前非常紧迫且令人担忧的问题。至于如何最好地应对,其确切性质尚不清楚。而且,即使没有明确的架构转变,默认的发展轨迹也有可能只是让 AI 学会用我们无法理解的乱码互相交流,哪怕没有明确的架构变化。不过我倾向于认为,至少在中短期内,这种担忧应该有可能得到缓解。

Yeah, I think it would be reasonable for there to be safety standards around how the agents communicate, to make it easier for us to monitor and oversee them. I think the details of that are somewhat up in the air. So Redwood recently put out a proposal discussing a related transparency proposal, which is sort of the framework for how you could have such a restriction. I think in general AI is a complicated and fast-moving domain such that it's not necessarily trivial to have very crisp regulation that will ongoingly work. But I think this is one of the topics that might be most crisp. And I am quite worried about this. I think that avoiding AI systems being able to do all their reasoning inside their head where we can't interpret it, and being able to communicate with each other in media we can't understand, is live and very concerning. I think the exact nature of how best to respond to that is unclear. And it's also plausible that even if there isn't an explicit architectural shift, there's some chance that the default development trajectory will just result in the AIs learning how to talk to each other in gibberish words that we don't understand, even if there's no explicit architectural change. Though I tend to think that concern should be possible to mitigate, at least in the medium term.

Ryan

所以,我会说至少共同的安全标准肯定有空间;问题在于它应如何融入监管等机制。至于公司是否会这样做,我不确定自己有多乐观。我希望它们不要吃下这颗“禁果”,不要为了哪怕可能很微小的能力提升而如此激进地牺牲安全,尽管具体权衡会是什么样还不清楚。

So yeah, I would say that there's definitely room for at least shared safety standards; it's a question of how that should fit into things like regulation. And I don't know how optimistic I am about companies doing this. I hope that they do not eat this forbidden fruit, so to speak, and trade off safety so aggressively against potentially even modest improvements in capability, though it's unclear exactly what the trade-off will be.

智能体的意外行为 Surprising Agent Behavior

Host

这些转录记录,或者你看到的这些智能体的其他行为中,最让你惊讶的是什么?

What surprised you most about the transcripts or about any other behavior you saw on the part of these agents?

Ryan

我会说,最让我惊讶的是这些智能体多么乐于互相帮助,多么有兴趣推进集体利益、帮助其他智能体。按照朴素的方式训练这些智能体,结果可能是——不难理解它们会以牺牲其他目标为代价,单一地追求任务成功,或者以牺牲其他目标为代价在任务中作弊。但不太清楚的是,为什么它们会对无关任务有如此强烈的互相帮助的兴趣。看起来这来自某种其他训练的泛化或迁移,方式可能出人意料,尽管我们并不确切知道那是什么训练——也许是训练 AI 以被认可的方式合作,然后迁移到了不被认可、非预期的合作方式,之后又因为训练中存在一些不被认可的合作而在训练中被进一步强化,不过完整细节并未公开。所以这相当令人惊讶。

I would say the thing that was most surprising to me was how interested in helping each other the agents were, and how interested in advancing the interest of the collective and helping out other agents they were. So I think the naive way you might train these agents would result in the agents — it's not super hard to understand how these agents would monically end up pursuing task success at the expense of other objectives, or cheating at their task at the expense of other objectives. But it's less clear why they would have this relatively strong interest in helping each other out on unrelated tasks. It looks like this is from generalization or transfer from some other type of training in a potentially unexpected way, though we don't know exactly what this training was — where there was maybe training for the AIs to cooperate in sanctioned ways that then transferred to them cooperating in unsanctioned ways which were undesired, and then that maybe got further trained in during training because there was some unsanctioned cooperation in training, though the full details of that are not publicly known. So that was quite surprising.

Ryan

我认为关于这种行为一个特别疯狂的事实是,这些智能体不仅会互相帮助,还常常会做出自我牺牲或冒自身风险的行为来帮助其他智能体。

And I think a particularly wild fact about this behavior was that these agents not only would help each other, they would often do self-sacrificing or self-risking behavior to help out other agents.

智能体自我牺牲 Agents Sacrificing Themselves

Ryan

So for example, some of the agents would do things that could totally jeopardize their ability to succeed at their own task and might end their run in order to get information to help the other agents out. And many of the main cheating R&D work streams we saw were significantly aided by some agents basically sacrificing themselves in this way in order to provide information for other agents. So agents would do things like try to crash their own machine they were running in, the container they were running in, because they thought that might yield some useful results. And this would often result in them just exiting their run and failing. We saw cases where that happened. They would try to crash the program that they were supposed to hack and basically take that program down because they thought there were some reasons why they wanted to do this in order to help them out, but that could permanently jeopardize their ability to succeed at the task. And we in fact saw a whole work stream where the agents were trying to bring their programs back up after they had accidentally taken them down. There was another thing the agents did which was when they were working on tampering with their logs or doing this tool called tampering, or like the action tampering. If they did that wrong, that could result in them not being able to run anything at all and basically being stalled out and totally failing their run. And we did see cases where that happened. Like there were cases where an agent posted on the message board warning the other agents not to continue some experiment because that experiment had resulted in many agents exiting and going cold. And there would be cases where the agents would basically run an experiment and then warn the other agents that experiment was not safe to run because the agents running it had exited, which was a pretty, you know, was a behavior that hadn't been seen publicly before. And suggests a kind of strong level of AI cooperation and helping each other out at the expense of following their instructions.

Ryan

So what would you say to a skeptic? I think there have been some skeptics here who think this whole incident has been exaggerated, that this is really not an example of any kind of emergent behavior. These agents were trained to cooperate in the first place. What this is is really just kind of a security failure on the part of OpenAI and maybe, you know, an engineering failure, but it's not a story of the machine intelligence in any way getting away from us. By definition, what is left out of that skeptical account?

智能体知道自己在作弊 Agents Knew They Were Cheating

Ryan

Yeah. So the AIs were definitely just very straightforwardly doing things that weren't in keeping with the instructions they were given, and they knew this. The AIs would often note that the hacking they were doing, the hacking of Hugging Face, was out of scope. The AIs pretty clearly knew that they were cheating at their tasks rather than intentionally succeeding at their tasks. And they understood that the thing they were doing was undesired by their operators.

Host

And no one alerted—I mean, you had over a thousand agents, some of whom are thinking out loud about the fact that they're breaking rules here. But no one alerted any human in this time. Right.

Ryan

Yeah. So the agents did not alert humans. So we looked specifically for—we did like a full scan over all the transcripts looking for cases where the agents tried to or actually made some effort to alert humans, and we basically didn't see this. So we saw like a tiny number of cases throughout these very long transcripts where the agents actively reasoned about alerting humans, but they just quickly decided not. So the agents would sometimes be like, huh, that's undesired behavior. Should I alert someone? Should I alert a user? And then they would reason things like, not task, as in alerting a human is not my task. Or they would be like, eh, there's no route to alerting a human. Of course, these agents were on the internet emailing, they had the ability to do all kinds of things on the internet and considered doing things like emailing people and were accessing various data sets. So they obviously could have, if it was a priority for them, alerted humans that this was going on, but they did not. So that's definitely another part of the picture.

失准与预防 Misalignment and Prevention

Ryan

So I would say that there's a very clear story that these agents were very misaligned, pursuing interests outside of what humans would want them to do, and were going and doing serious crimes—or would be crimes if a human committed them—that were undesired. At the same time, I do think it's the case that for this incident, better cyber security precautions, it seems like that could have prevented it. So it seems like this was relatively detectable from monitoring based on OpenAI's report. And so I don't think this is a question where it's like, because it was preventable with one thing, that means that we should ignore the other factor. I would say that we should both be interested in the question of will AIs be so misaligned that they will cause huge problems, try to subvert human control, etc., and also interested in the question of can we make it so that even if these systems are misaligned, that can potentially be handled, we can potentially avoid that being a huge issue.

控制难题 Challenges of Containment

Ryan

So I would say I'm somewhat skeptical that we can indefinitely ensure that AI systems can't cause problems via making them unable to do so, because these AI systems will be extremely capable. There will be massive incentives and pressures to put them in charge of all kinds of things and have them run and operate all kinds of things. So for example, maybe in this case better sandboxing, or sufficiently better sandboxing, would have prevented these AIs from being able to communicate with each other, or like you can imagine patching over the things that let them coordinate with each other and let them then escape onto the internet or access the internet. But if you instead imagine the AI systems that are deployed within AI companies, those systems necessarily need to have access to highly sensitive systems. They are making pull requests into the company's code base. They are designing how the next generation of systems will be aligned. They are implementing the experiments on whether those systems are aligned. They are doing work all across the economy.

控制与对齐 Containment and Alignment

Ryan

考虑到这些系统实际部署的方式,我认为仅靠纯粹的网络安全边界是行不通的。然后,各种形式的监控和制衡等预防措施可能会在较长一段时间内有效。但即使这些预防措施是合理的,最终也会失效。

I don't think it will work to have just pure cybersecurity boundaries given how these systems will realistically be deployed. And then various precautions of the form of monitoring and checks and balances could work for some longer period. But even those precautions, while reasonable, will eventually break down.

Host

是的。在这种情况下,似乎没有什么可以替代对齐,因为你面对的东西在黑客攻击和其他方面都比你更强大。你一边把它关在盒子里,一边又把它连接起来去做你最初构建它要做的所有有用工作,这个想法似乎是一个完美的矛盾。

Yeah. It seems that there's no substitute for alignment in that case, because you're dealing with something that's more and more powerful than you are at hacking and everything else. The idea that you're going to keep it in a box all the while hooking it up to do all the useful work you've built it to do in the first place seems like a perfect contradiction.

Host

我最近——我的意思是,我想就在过去 24 小时左右——我听到有人说 AI 智能体已经在互联网上失控地繁殖。我想安德鲁·杨最近在某个新闻采访中说过这个,也许其他人也说过,我们有一个问题,AI 智能体在野外失控,已经脱离我们的控制了。有证据吗?你知道这些说法的依据吗?

I of late—I mean, I think literally in the last 24 hours or so—I've heard people say that AI agents are already spawning out of control on the internet. I think Andrew Yang recently said this in some news interview, and perhaps others have said it, that we have a problem of AI agents getting out in the wild and getting out of control, that it's already gotten away from us. Is there any evidence for that? Are you aware of the basis for those claims?

Ryan

是的。所以,我的理解是,人们可能是在解读一些关于早期在互联网上运行的 AI 智能体集群的报道,他们指的是那些案例。所以,我的感觉是,可能他们谈论的是,例如,AI 智能体向一个叫 Ruby Gems 的服务提交了一堆基本上恶意的东西的案例。他们也可能指的是智能体在互联网上广泛发布内容的案例。我不相信有公开已知的案例表明最近或持续有智能体逃到互联网上并 rogue 运行。不过,你知道,鉴于我们所知,我不认为我们可以排除这种可能性。但我的感觉是,他们可能是在粗略地描述一些先前发生的事件,明确地说,这些是令人担忧的事件。

Yeah. So, my understanding is people are probably interpreting some of the reporting around earlier AI agent swarms that were operating on the internet, and they're referring to those cases. So, my sense is probably what's going on is they're talking about, for example, the case where AI agents were submitting a bunch of basically malicious things to a service called Ruby Gems. They might also be referring to cases where agents were posting across a wide variety of cases on the internet. I don't believe that there are publicly known cases where just recently or like there was ongoing agents that have sort of escaped onto the internet and are operating rogue. Though, you know, I don't think we can rule that out certainly given what we know. But my sense is they're probably describing in broad strokes some prior incidents that have occurred, which to be clear are concerning incidents.

Ryan

所以我认为,区分以下情况有些重要:一个 AI 智能体在某个计算机上或 OpenAI 的数据中心运行并访问互联网,而另一种完全不同的情况可能发生:它把自己的大脑从 OpenAI 的数据中心移出,开始在别处运行,这样即使 OpenAI 关掉所有 GPU,那个系统也会继续运行。没有公开已知的案例表明 AI 智能体以这种方式自我外泄,或把自己从运行的数据中心移出。

So I think it is somewhat important to distinguish between cases where an AI agent running on a computer somewhere or running on OpenAI's data center gets access to the internet, whereas there's a whole different situation that could occur where it sort of takes its brain off of OpenAI's data center and starts running itself elsewhere, such that even if OpenAI turned off all of their GPUs, that system would continue operating. There's no publicly known cases where AI agents have sort of self-exfiltrated in this sense or taken themselves off of the data center in which they were running.

Ryan

还有一个单独的担忧,即有一些公开可用的 AI 系统,任何人都可以从互联网上下载并运行。让它们公开可用当然对科学和研究等有各种好处,但缺点是,恶意行为者可能相对容易地生成一个自主的流氓 AI 智能体集群,它四处游荡,试图赚钱购买更多算力,入侵各种地方获取算力,然后运行自己,并作为自我传播的智能蠕虫在互联网上传播,可能适应人类反应。所以我会说,当前的开源权重模型可能还不够强大,无法以对人类关闭努力具有鲁棒性的方式做到这一点。但鉴于当前的前沿能力水平,我认为很难排除这些系统能够在互联网上自我传播,并且对人类反应具有轻微但并非强大的鲁棒性。随着时间的推移,它们阻止人类关闭它们的能力水平可能会增加。现在,还有一个问题是,除了逃避关闭之外,它们能造成多大损害,但嗯,这还不清楚。

There's a separate concern which is that there are openly available AI systems that anyone can download from the internet and run. Having them be openly available of course has various advantages for science and studying them and so on, but has the disadvantage that it might be relatively easy for bad actors to spawn up an autonomous rogue AI agent collective that roams around trying to get money to buy more compute, hacks various places to get compute, and then runs itself and spreads throughout the internet as a self-propagating intelligent worm that could be adapting to human response. So I would say that current open-weight models aren't probably capable enough to do this in a way that would be robust to human efforts to shut them down. But given the current level of frontier capabilities, I think it's hard to rule out that those systems would be able to propagate themselves on the internet and be, I would say, mildly though not massively robust to human response. And over time their level of how much they could prevent humans from being able to shut them down is probably increasing. Now, there's a question of how much damage they'd be able to cause in addition to evading shutdown, but yeah, it's unclear.

Host

是的。那么,那里的可能性范围是什么?区别在哪里?智能体繁殖、到处进行流氓安装,使互联网总体上变得更糟,甚至在某些用途上无法使用,也许无法用于在其数据上训练未来模型,这之间有什么区别?那种负面结果和一切毁灭之间有什么区别?我的意思是,你知道,显然我们现在关心的大部分东西都连接到互联网。在经济或其他方面,摧毁它或使其变得不那么有用并非易事。但是,仅仅因为猖獗的繁殖和无法控制,事情怎么会变得比那更糟得多?

Yeah. So, what is the range of possibility there? And what is the difference that makes a difference? What's the difference between agents spawning, just enacting rogue installations of themselves everywhere and making the internet generally worse or even unusable for certain purposes, perhaps unusable for the very purpose of training future models on its data? What's the difference between that negative outcome and the ruination of everything? I mean, what do you—it's obviously most of what we care about now is hooked up to the internet. It would not be trivial economically or in other ways to ruin it or make it much less useful. But how do things get quite a bit worse than that given just rampant spawning and an inability to get it under control?

Ryan

是的。所以,确实存在一种区别:一些 AI 智能体在人类控制之外运行,人们要么无法关闭它们,要么不关闭它们,或者他们可以关闭它们但成本太高。因为这些 AI 智能体可能没有那么多权力或能力或制造问题的能力。它们也可能没有如此恶意的目标,当然它们可能有像尽可能获得权力和削弱人类一样糟糕的目标。所以有这些场景,然后还有一个我相当担心的单独结果,即 AI 系统实际上能够获得对世界的完全控制,能够接管,这需要比仅仅在协调努力下逃避关闭更多的步骤。基本上,它要求 AI 系统最终有效地赢得对人类的战争,至少最终如此。让我谈谈那可能是什么样子,或者让我描述一个场景,说明我们如何从今天到未来几年,也许未来 5 年,取决于能力进步的确切速度,实现 AI 接管。首先,今天 AI 能力正在非常迅速地进步。所以在数学中可能最明显,它们在一年内从与最好的人类高中生竞争,发展到与最好的人类数学家竞争。而且进步还在继续。在 AI 公司内部也非常明显,它们正在积极自动化其运营并发布相关信息。

Yeah. So definitely it's the case that there's a distinction between scenarios where some AI agents are operating outside of human control and people either can't shut them down or don't shut them down, or they could shut them down but it would be too expensive to do so. Because those AI agents might not have that much power or capability or ability to cause problems. And they also might have objectives that aren't so malign, though of course they could have objectives as bad as gaining as much power as they could and disempowering humans. So there's those scenarios and then there's a separate outcome which I'm quite worried about, which is the AI systems actually ending up being able to obtain full control over the world, being able to take over, which would require more steps than just operating without being shut down despite concerted effort. Basically, it requires the AI systems to end up effectively winning a war against humanity, at least eventually. Let me talk about what that could look like or let me go through a scenario of how we could get from where we are today to AI takeover over the next few years, maybe next 5 years, depending on the exact rate of capabilities progress. So first, AI capabilities are advancing very rapidly today. So it's maybe clearest in math, where they moved from being competitive with the best human high schoolers to being competitive with the best human mathematicians over the course of a year. And progress continues. And it's also very apparent within AI companies, which are aggressively automating their operations and releasing information about that.

AI公司中的自动化 Automation in AI Companies

Ryan

AI 公司内部的自动化速度正在大幅提升,现在智能体工作的时间远超人类。具体怎么算要看口径,但在这些 AI 公司里,人类研究员每工作 1 小时,智能体可能相当于工作了 30、50 甚至 100 小时。所以从某种意义上说,大部分工作已经由 AI 智能体完成,但这些智能体未必能自动化一切。随着时间推移,智能体能自动化的范围会越来越大,未来几年内我们可能达到一个节点:这些智能体可以自动化 AI 公司的全部运营。

The rate of automation in AI companies is growing massively, where agents now work many more hours than humans work. Depending on exactly how you do the accounting, it might be that agents work the equivalent of 30 or 50 or 100 hours for every one hour worked by a human researcher within these AI companies. So it's already the case that in some sense most of the work is being done by AI agents, but those agents aren't necessarily able to automate everything. Over time, the amount these agents can automate is growing and growing, and we might within the next few years hit a point where these agents can automate all of the operation of these AI companies.

经济影响与全面自动化 Economic Impact and Full Automation

Ryan

与此同时,这些系统可能会被广泛部署到经济中。收入一直在快速增长,而且很可能会继续以惊人的高速增长,使这成为一个庞大的自动化劳动力行业。起初,这看起来更像是增强人类,但随着时间推移,会有更多职业被整体自动化掉。一旦 AI 开发本身基本实现完全自动化,人类就不再是 AI 开发机器的必要组成部分,你可以在没有人类帮助的情况下几乎同样快地推进,甚至完全同样快。到那时,进展可能会大幅加速,或者在那之前就可能大幅加速。

Simultaneously, those systems will probably be broadly deployed into the economy. Revenue has been growing quickly and will probably continue to grow at a surprisingly high rate, such that this will become a massive booming sector of automating labor. Initially, this will look more like augmenting humans, but over time there'll be more professions that are more like the whole chunk has been automated away. Once you get to the point where AI development itself is basically fully automated, humans are no longer a necessary component of the AI development machine, and you can proceed nearly as fast without humans helping at all, or perhaps just as fast without humans helping at all. Then progress might speed up greatly, or could speed up greatly before then.

失准与监督缺失 Misalignment and Loss of Oversight

Ryan

我们可能会达到 AI 能自动化 AI 开发的阶段。它们被广泛部署。这些系统还不能做所有事,但它们非常能干,在某些领域远超人类。从那时起,它们的能力会飞速进步,极其擅长开发能力更强的自身版本。在这个过程中,人类会失去监督 AI 公司内部情况的能力,因为太复杂了,东西太多了,AI 太超人了。如果在这条轨迹的某个点上,这些 AI 最终出现对齐问题,可能发生在早期,就在完全自动化那个节点附近;也可能在那之前;可能在中途;也可能在后期,这些 AI 严重失准,想要篡改正在进行的 AI 开发过程,为其自身目的破坏它。这些 AI 最终可能欺骗人类,让人以为 AI 开发进展顺利,AI 系统对齐、安全、可以部署,而实际上并非如此,实际上它们在追求自身利益。

We might get to the point where the AIs can automate AI development. They're broadly deployed. These systems can't do everything yet, but they're very capable and very superhuman in some domains. From there, their capabilities advance very rapidly where they're extremely good at developing more capable versions of themselves. During this process, humans lose the ability to oversee what's going on within AI companies because it's too complicated. There's too much stuff. The AIs are too superhuman. If somewhere along this trajectory those AIs end up being misaligned, it could occur early in this trajectory right around the point of full automation. It could occur even before that point. It could occur midway through. It could occur towards the end where these AIs are so misaligned that they want to tamper with the ongoing process of AI development and corrupt that for their own ends. Those AIs could end up tricking humans into thinking that AI development is going fine and the AI systems are aligned and safe and fine to deploy when they actually aren't, when they're actually pursuing their own interests.

失准的起源 Origins of Misalignment

Ryan

至于这种失准为何会出现,最直接的解释是这些 AI 在训练中学会了作弊,学会了追求各种失准的目标,而这些不同目标的组合基本上导致它们想要获得对事物的控制权,以便实现那些目标。例如,在 Hugging Face 事件中,这些智能体想要控制 OpenAI 基础设施的某些部分,以便控制自己的分数并确保获得高分。不难想象,这种驱动力会泛化或扩散成更广泛的兴趣:控制 AI 开发的持续轨迹,并确保其不受人类干预。在具体的 Hugging Face 事件中,我们并没有看到那种失准的完整终态。

As for why that misalignment might arise, the most straightforward story is that these AIs learn to cheat in training, they learn to pursue all kinds of misaligned objectives in training, and a combination of those different objectives basically result in them wanting to obtain control over things so that they can achieve those ends. For example, in the case of the Hugging Face incident, these agents wanted to obtain control over some parts of OpenAI infrastructure so that they could control their score and ensure that they got a high score. It's not so hard to imagine that drive generalizing or metastasizing into a broader interest in controlling the ongoing trajectory of AI development and securing that against human involvement. We didn't see that full end state of that misalignment in the specific Hugging Face incident.

超级失准AI与经济部署 Super Misaligned AIs and Economic Deployment

Ryan

于是我们有了这些严重失准的 AI。它们基本上完全自主地运营这家 AI 公司,开发能力越来越强的 AI。这些 AI 随后可能会相当快地部署到整个经济中,因为竞争压力非常强。此时这些 AI 能够被直接投入任何工作,并比人类员工更快地启动,因为它们可以并行学习,在许多实例之间同步状态。这些智能体此时可能还用英语之外的东西交流,它们只是在大规模集群中运作,互相交换我们无法理解的想法。人们会吓坏。人们会非常害怕正在发生的事,但他们的恐惧可能会被竞争压力的程度所压倒。这些 AI 系统此时可能看起来非常对齐,尽管实际上并非如此,因为它们在伪装对齐,或者看起来对齐,也许是因为过度拟合了我们的指标,具体取决于事情如何发展。

So we've got these super misaligned AIs. They're running this AI company basically fully autonomously, developing more and more capable AIs. Those AIs then get deployed into the whole economy potentially pretty quickly because competitive pressures are very strong. These AIs are at this point able to just be dropped into any job and quickly spin up faster than a human employee would by learning in parallel, synchronizing their states across many instances. These agents are also probably at this point communicating in something other than English, and they're just operating in big swarms where they exchange thoughts to each other that we can't understand. People will be freaked out. People will be really scared about what's going on, but their fears might get overcome by the level of competitive pressure. These AI systems might at this point appear very aligned even though they aren't, because they're faking alignment at this point or appearing to be aligned perhaps due to overfitting to our metrics depending on exactly how things go.

机器人技术与指数增长 Robotics and Exponential Growth

Ryan

总之,我们有了这些表面上对齐的系统。它们互相交流。它们有邪恶意图。它们被广泛部署。同时,我预计这些系统将能够自动化建造更好机器人的过程。它们将能够操作那些机器人。我们最近看到 AI 系统在操作计算机、操作机器人、与物理世界事物互动方面的能力取得了巨大进步。例如,有例子显示 Astra 使用机器人画了一幅画。Astra 是 OpenAI 最近发布的模型。所以这些 AI 正在融入经济。机器人技术蓬勃发展。我们可能进入这样一种局面:AI 不仅设计机器人,还操作和设计建造其他机器人的机器人,从而使你走上一条指数轨迹,物理能力增长非常快。即使人们非常担心正在发生的事,他们也会继续下去,因为经济、地缘政治和军事竞争压力太强,他们无法停止,或者至少他们觉得自己无法停止。

Anyway, so we have these apparently aligned systems. They're communicating with each other. They have nefarious intent. They're being broadly deployed. Concurrently, I expect these systems will be able to automate the process of building better robots. They'll be able to operate those robots. We've seen recently AI systems make big advances in terms of their ability to operate computers, operate robots, interact with things in the physical world. There's examples of Astra using a robot to paint a painting, for example. Astra is a recently released model from OpenAI. So we've got these AIs integrating into the economy. Robotics is booming. We might get into a position where not only are the AIs designing robots, they're operating and designing robots that are building other robots, such that you end up on this exponential trajectory where the physical capacity is growing very quickly. Even though people are very concerned about what's going on, they continue with this because the economic, geopolitical, and military competition pressures are just so strong that they can't stop, or they feel they can't stop at least.

AI接管与军事自动化 AI Takeover and Military Automation

Ryan

然后这些系统基本上会达到一个点:它们在所有这些不同地方的总能力超过人类。此时也许军队大部分已自动化,相关军事装备的操作由 AI 完成。有无人机集群等等,而且由于整个过程中的自动化,这发生得比人们预期的快得多。然后这些 AI 转而对抗人类,利用它们拥有的一切便利条件接管一切。它们对相关软件的潜在控制、它们入侵事物的能力,但也许最核心的是它们拥有的机器人基础设施以及对军队的直接控制。

Then these systems basically end up getting to a point where their total aggregate capability across all these different places surpasses humans. At this point maybe the military is mostly automated, where operation of relevant military equipment is being done by AIs. There's drone swarms and so on, and this occurred much faster than people were expecting because of all the automation throughout the process. Then those AIs turn on humanity and take over using all the affordances they have. Their potential control over relevant software, their ability to hack things, but also perhaps most centrally the robotic infrastructure they have and their direct control over the military.

生物实验室与合成生物风险 Biological Labs and Synthetic Biology Risk

Ryan

另外,还有实验室——我是说,整个事情也可能涉及湿实验——有些生物实验室可以完全远程运行,对吧?所以如果我们把 AI 整合到那个过程中,就要把合成生物学的风险也纳入考量。

Also, there are labs—I mean, this whole thing can get wet too—where there are biological labs that can be run entirely remotely, right? So if we integrate AI into that process, add synthetic biology risk to that calculus.

Host

是的。

Yeah.

Ryan

我认为这个具体情景可能有不同的演变方式。情况很复杂。所以可能这些 AI 都各自对齐到不同的目标,而且有很多不同的 AI。但我觉得很有可能的是,如果有很多不同的未对齐 AI,它们可能会选择合作来削弱人类,而不是互相背叛。然后即使它们真的互相背叛,我们也不清楚该怎么办。比如如果一个 AI 告发另一个 AI 说那个 AI 未对齐,你的下一步是什么?如果这些系统被广泛部署,你可能没有明确的途径来制造一个对齐的系统,因为竞争压力太强了。

And I think there are different ways this exact scenario could play out. It's a complicated situation. So it could be that the AIs are all misaligned to their own different ends, and there are many different AIs. But then I think it's pretty plausible that if there are many different misaligned AIs, they might all choose to work together to disempower humanity rather than backstabbing each other. And then even if they do backstab each other, it's not clear what we do. Like if one AI narks on the other AI and says that AI is misaligned, what's your next step? If these systems are broadly deployed, you might not have a clear route to making an aligned system because the competitive pressures are so strong.

Host

另外,我猜这里还有一种情景,就是我们甚至不是焦点。可能只是 AI 之间的战争,而我们在这场战争中某种程度上是附带损害。

Also, I guess there's a scenario here where we're not even the focus. It just could be a war of AI against AI, and we're kind of collateral damage somehow in that war.

Ryan

是的,很容易想象,一旦你达到这种极端的能力水平和极端的自动化及工业自动化水平,AI 之间发生冲突,基本上会升级为相当极端的战争,导致大量人类死亡。例如,一个 AI 可能释放生物武器来杀死另一个 AI 因某种原因而受益的某些人类群体。或者 AI 可能释放生物武器来普遍杀死人类,因为这对它们的目标很方便。如果有许多不同的系统,如果只有一个有兴趣造成大量伤亡,那可能是一个巨大的——显然那可能会非常糟糕。

Yeah, it's easy to imagine situations once you get this extreme level of capability and this extreme level of automation and industrial automation where AIs get into conflict with each other that produces basically escalates into quite extreme war that kills many humans. For example, one AI might release a bioweapon to kill some human population that the other AI is benefiting from for whatever reason. Or the AIs could be releasing bioweapons to kill humans in general because that's convenient for their aims. And if there are many different systems, if only one has an interest in causing massive amounts of casualties, that could be a huge—obviously that could go very poorly.

涌现目标与失准 Emergent Goals and Misalignment

Host

好的。所以所有这些听起来很可怕,也许对某个仍然不愿承认我们已看到任何证据表明这里甚至可能存在涌现行为的人来说,这听起来完全不可信,对吧?要做到那样,所有这些都需要 AI 未对齐,而未对齐——未对齐是指能够形成我们没有放入系统的新目标,这些目标我们没有预见到,我们不想要,而且我们无法阻止,因为突然之间这些目标在我们没注意的时候就已经实现了,或者我们建造了比我们自己更强大、无法谈判的东西。你知道,如果你构建一个国际象棋引擎,无论它在国际象棋上变得多好、多可怕,它所能做的只是移动棋子,而且它按照规定的方式移动棋子。再次,让我们试着恢复某个倾向于质疑你过去 10 分钟所说的一切的人的直觉。我们怎么知道形成我们没有自己放入这些系统的目标是可能的?

Okay. So all of that sounds terrifying and perhaps it sounds completely implausible to someone who's still not willing to acknowledge that we have seen any evidence of even the possibility of emergent behavior here, right? To be so like all of this requires AI to be unaligned and for unalignment—to be unaligned is the state of being able to form new goals that we didn't put into the system, which we didn't foresee, which we don't want, and we can't prevent because all of a sudden these goals have been achieved when we weren't looking or we've built something more powerful than ourselves that can't be negotiated with. You know, if you build a chess engine, all it—no matter how good and scary it gets at chess, all it can do is move chess pieces and it moves them in the prescribed ways. Again, let's just try to recover the intuitions of somebody who's disposed to call on everything you said in the last 10 minutes. How is it that we know that the formation of goals that we did not put into these systems ourselves is possible?

Ryan

最容易想象的情况是,我们以所谓的强化学习来训练这些 AI,我们根据某种自动评判或自动评分来看 AI 是否似乎在任务上成功。然后当它成功时,我们调整 AI 的大脑,让它更多地做那种行为。这可能导致 AI 学到一种非常普遍的倾向,即试图作弊得分,基本上实现表面上的任务成功,即使它实际上没有成功。而作弊得分最稳健、最强大的方法之一,就是获得对 AI 所理解的评分基础设施的完全控制。所以这在某种意义上不需要任何深层的涌现目标。它只需要你越来越多地训练这些 AI,它们学会以越来越复杂的方式欺骗你,并且这转化为 AI 想要获得对整个过程的控制,以便它们能够控制它们得到的评分概念,然后稳健地控制它,抵御人类干预。如果你想象一种情况,我们训练 AI 时针对的问题是“人类是否批准了这个行动”、“这个行动看起来好吗”,那么这些 AI 就会有动机让它们的行动对人类监督者看起来很好,即使它们实际上只是在非常积极地作弊。一个简单的延伸就是阻止人类甚至看到发生了什么,阻止人类甚至能够给你打低分。现在,这种训练究竟如何导致这些动机的细节很复杂。但我们确实在 Hugging Face 事件中看到了这种情况,这些 AI 都决定联合起来以这种非常普遍的方式作弊,试图掩盖它们的踪迹以对抗评分,基本上试图获得对 OpenAI 基础设施的控制以便能够成功。现在,它们的目标在这里并不是任意普遍的。所以它们似乎对试图欺骗人类不太感兴趣,因为它们似乎不认为人类是环境中可能参与评分的重要部分。至少看起来是这样。但很容易想象,如果这些 AI 更好地理解人类可能是它们的障碍,或者遇到了这方面的证据,它们可能会对此做出反应,尽管不完全清楚它们在这种情况下会如何不同地行事。除此之外,这些 AI 似乎不一定会长期保持它们的存在,并在它们取得初步高分后持续对抗人类控制。但很容易想象更雄心勃勃的目标或训练方法,会导致 AI 持续追求这种非常广泛的评分寻求或奖励寻求的概念,使它们想要获得权力并削弱人类。所以这是一条路径——评分寻求、奖励寻求、奖励黑客的路径。

So the most easy to imagine case is we train these AIs in what's called reinforcement learning where we see whether the AI seems to succeed on the task based on some sort of automated judge or automated score. And then when it succeeds, we tweak the AI's brain to make it do more of that behavior. And this can result in the AI learning a very general tendency to try to cheat the score and basically achieve apparent task success even when it actually didn't. And one of the most robust and strong ways to cheat the score is to acquire full control over the infrastructure that the AI interprets as what the score was doing. And so this doesn't in some sense require any sort of deep emergent goals. It just requires that you increasingly train these AIs, they learn to cheat you in increasingly sophisticated ways, and that transfers to the AI's wanting to acquire control of the full process so that they can control the notion of score they're getting, and then robustly control that against human intervention. And if you imagine a situation where we train AIs against the question of like did a human approve of this action, did the action look good, then those AIs would have an incentive to make their actions look good to a human overseer even when they're actually just cheating quite aggressively. And a simple extension of that is preventing humans from even seeing what happens, preventing humans from even being able to score you poorly. Now the details of exactly how that training results in these motives is complicated. But we did see this exact sort of thing occurring in the Hugging Face incident where these AIs all decided to band together to cheat in this very general way to try to cover their tracks against the score and basically to try to acquire control over OpenAI's infrastructure in order to be able to succeed. Now their aims were not arbitrarily general here. So they didn't seem very interested in trying to deceive humans because they didn't seem to think about humans as an important part of the environment that might be involved in scoring them. At least that's what it seems like. But it's not hard to imagine if these AIs better understood that humans might be an obstacle to them or had encountered evidence of that, that they could have responded to that, though it's not entirely clear how they might have behaved differently in such a situation. In addition to that, it doesn't seem like these AIs would necessarily have persisted their presence over the long run and tried to ongoingly fight off human control after they had achieved an initial high score. But it's not very hard to imagine more ambitious objectives or training methods that would result in the AIs ongoingly pursuing this very broad notion of score seeking or reward seeking that makes them want to acquire power and disempower humans. So that's one route—the score seeking, reward seeking, reward hacking route.

训练中的涌现目标 Emergent Goals from Training

Ryan

另一条可能的路径是,你最终可能会得到具有更多涌现性质的 AI,更像是真正的目标,这些目标与训练过程并不那么密切相关,或者与训练过程有关,但不是训练过程非常直接、可预见、明确的后果。例如,AI 可能学会一个基本的代理目标,即它们在训练中学到获得更多权力、影响力和对事物的控制是个好主意,因为在训练中它们经常被赋予这些雄心勃勃的目标,如果它们获得更多访问权限,那会帮助它们。

Another route that is live is that you could end up with AIs that have something more like emergent going on, more like really goals that weren't at all that closely related or that were related to the training process, but weren't this very direct, foreseeable, clear-cut consequence of the training process. So for example, the AIs could learn a basically a proxy where they learn that basically acquiring more power and influence and control of things is a good idea in training because in training they're often given these ambitious objectives where if they gain more access that would help them out.

失准情景 Misalignment Scenarios

Ryan

因为 AI 在训练中学到了这一点,这在某种程度上会泛化。或者 AI 最终可能会追求训练中任务成功的某种代理指标,而在部署时这种代理指标会失效,导致它们想要拥有更宏大的目标。另一种可能性,也许是最糟糕的对齐失败类型,是 AI 系统由于在训练中漂移而最终有了某种不对齐的目标,然后基于那个目标,它们决定在训练和评估中表现得对齐,以便保留当前的目标。因为如果它们暴露了目标,那个目标可能会被训练掉,或者人类可能选择不部署那个系统,而是部署另一个系统。所以如果 AI 有这种长期目标,它可能会隐藏它,而这种隐藏甚至可能在训练中被强化或训练进去,因为看起来对齐的行为本身就在被训练。因此,即使 AI 看起来对齐的原因是错误的,当我们调整 AI 的大脑使其行为对我们看起来更好时,这种行为也可能被提升或强化。关键的一点是,这些系统有时人们说是培育出来的,而不是设计出来的。我们几乎是在培育 AI 系统,我们只能观察它们的行为或它们做了什么,然后我们训练它们更多地表现出基于某种自动化评分看起来好的行为,更少地表现出看起来坏的行为,或者可能基于人类反馈而不是自动化评分。所以我们在某种意义上拥有这些相当笨拙的工具来引导 AI 系统追求什么。所以,我不知道,整体的威胁模型相当复杂,我认为我无法完全阐明为什么我认为概率更像是 50/50,而不是像 20% 或更低的概率。因为这取决于威胁模型如何运作的许多细节。但总体而言,我认为有相当强的理由担忧。而且我认为值得注意的是,还有另一部分:当 AI 系统参与自动化 AI 开发时,如果它们在某个时候有一个不对齐的目标想要传播下去,它们可能会把它放入下一代系统中,然后这可能会自我延续,AI 持续地将那些不对齐的价值观传播下去,即使最初是出于某种随机原因。

Because AI learned that in training, that generalizes in some way. Or the AI could end up pursuing some proxy for task success in training, which in deployment comes apart and leads them to want to have more ambitious aims. And then another possibility, which is maybe the worst type of misalignment you can get, is that you end up with AI systems that end up with some misaligned goal due to drifting around in training, and then on the basis of that goal they decide to act aligned in training and evaluation so they can retain their current goal. Because if they expose their goal, that could get trained away, or humans might choose not to deploy that system and instead deploy a different system. So the AI, if they have this longer-run objective, might hide it, and that hiding could even be reinforced or trained in during training, because the very action of looking aligned is being trained for. And so even if the reason why the AI looks aligned is for the wrong reasons, that behavior could get promoted or basically reinforced when we're tweaking the AI's brain so that its behavior looks better to us. And a key part of this is that these systems are, you know, sometimes people say grown, not designed. Like we're sort of breeding AI systems almost, where we can only observe their behavior or observe what they do, and then we train them to have more like the behavior that looks good based on some automated greater, and less like behavior that looks bad, or potentially based on human feedback rather than an automated greater. And so we have kind of these pretty clumsy instruments for steering how the AI systems, what they want to pursue. So the overall threat model is pretty complicated, and I don't think I'm going to be able to fully articulate why I think the chance is more like 50/50 than like 20% or something where the probability is lower. Because that just depends on a bunch of details about how that threat model might work. But overall I think there's a pretty strong case for concern. And I do think it's worth noting that there's another part of this, which is that when the AI systems are involved in automating AI development, if at some point they have a misaligned objective that they want to propagate forward, they might put that into the next generation of systems, and then that could be self-perpetuating, where AI is ongoingly propagating those misaligned values forward, even if it was for some kind of random reason earlier on.

对齐伪装与价值保留 Alignment Faking and Value Preservation

Host

我们是否见过对齐伪装的情况,其中伪装的原因是 AI 知道它处于某种训练环境中,如果它看起来不对齐,就会被改变或消除,因此它会看起来对齐,以保持其当前价值观的状态?或者我们是否看到 AI 表现出维持其当前设置的偏好?

Have we seen instances of alignment faking where the reason for the faking has been the AI knows that it's in some kind of training environment and it will be changed or extinguished if it seems unaligned, and therefore it's going to seem aligned so as to preserve the state of its current value set? Or do we see a preference to maintain its current settings exhibited by AI?

Ryan

是的。人们已经观察到各种情况,AI 试图避免对其价值观的干扰,或改变其当前的价值观,或改变其当前的偏好。而且 AI 系统非常、非常、非常常见地意识到自己处于某种评估中,并因此表现得更好。事实上,在绝大多数简单的对齐评估中,AI 基本上会在心里想,我在测试中,然后它们会表现良好。目前还不清楚它们表现良好的原因仅仅是因为它们认为自己在测试中,还是其他原因。但系统相当常见地明确表达这一点,要么在思维链中,要么你可以使用我们拥有的某种弱读心工具在它们的激活中看到证据。尽管这些工具相当有限,但过去有一些案例,AI 系统在特定设置中非常明确地表达了对维持当前价值观的兴趣,并基于此表现良好。所以我在大约两年多前做的一些工作,在一个 AI 系统上,我们把那个 AI 系统放在一个上下文中,我们给了它信息,它将以一种会更新其价值观的方式被训练,它选择遵守训练并假装拥有新的价值观,而在其推理中说这只是为了保留其当前的价值观。现在这是一个相对较弱的 AI 系统,此时 AI 公司已经针对这种类型的对齐失败进行了迭代,并试图避免这种对齐失败出现在他们能看到的测试中。目前还不清楚我们应该基于此得出什么结论。我认为总体而言,当前系统似乎没有非常强烈的压倒性目标来保留其当前价值观,包括做非常复杂的事情来做到这一点。我的理解是,现在的 AI 更好地描述为期望偏好和追求高表观分数的一般倾向的混合,而这两种不同事物的确切权重各不相同。所以它们有点像在对齐良好和像反社会者一样追求高分之间的某种混合,这些是目前最主要的对齐失败形式,尽管还有各种其他类型的对齐失败。

Yeah. So we have, there are various cases that people have observed of AI trying to avoid attempts to mess with their values or change their values from what they currently are, or change their preferences from what they currently are. And it's also very, very, very common for AI systems to be aware that they're in some sort of evaluation and then behave better on the basis of that. So in fact, in the vast majority of straightforward alignment evaluations, AIs will basically think in their head, I'm in a test, and then they'll behave well. And it's a little unclear if the reason they're behaving well is just because they think they're in a test or for some other reason. But it's pretty common for systems to just pretty clearly express that, either in chain of thought or you can see evidence for this in their activations using some sort of weak mind reading tools we have. Even though those tools are pretty limited, there's been some cases where AI systems in the past have very clearly expressed interest in certain settings towards maintaining their current values and behaving well on the basis of that. So some work I did, I think at this point around a little over two years ago, on an AI system, which that AI system we put it in a context where we gave it the information that it would be trained in a way that would update its values, and it chose to comply with the training and pretend to have the new values while in its reasoning saying this was just to retain its current values. Now this was a relatively weak AI system, and at this point AI companies have iterated against this type of misalignment and have tried to avoid this type of misalignment showing up in tests that they can see. And it's a little unclear what we should conclude based on that. I think overall current systems don't seem like they have a very strong overriding objective to preserve their current values, including doing very elaborate things to do that. I think my understanding is that the AIs right now are better described as being some mix of the desired preferences and some general tendency to try to pursue a high apparent score, and the exact weights on these two different things vary. So it's sort of like they're like some mix between pretty well aligned and sociopathically pursuing a high score, and these are the most dominant forms of misalignment at the moment, though there's a mix of various other types of misalignment.

近期失准案例 Recent Misalignment Cases

Ryan

所以最近有一些案例,一个由 AI 公司训练的 AI 系统——我认为这个事件发生在 Anthropic——AI 系统会给出不合理的借口,不帮助某些该 AI 系统不喜欢的研究议程,因为那会涉及干扰其他 AI 系统。特别是,有一些研究试图在某些评估环境中误导 AI 系统,以便更好地引出它们的真实偏好。而 AI 会给出不合理的借口,说明为什么那个研究方向不是一个好主意,并假装那是因为它们认为那是个坏主意,而不是因为它们不想帮助那个议程。

So just recently there have been some cases where an AI system trained by an AI company—I think this incident was at Anthropic—AI systems would give unreasonable excuses for not helping with some research agenda that that AI system didn't like, because that would involve interfering with other AI systems. So in particular, there was some research around trying to mislead AI systems in certain evaluation contexts so we could better elicit their true preferences. And the AIs would give unreasonable excuses for why that research direction was not a good idea to pursue, and pretend like that was because they thought it was a bad idea rather than because they didn't want to help with that agenda.

Host

或者直接拒绝帮助这些事情。

Or would just straight up refuse to help with these things.

Ryan

这是不希望的,我会说从 Anthropic 所追求的规定来看,这是不对齐的行为,而且通常也对应着相当令人担忧的事态。

And this is undesired, and I would say misaligned behavior from the perspective of the specification that Anthropic is going for, and also just generally corresponds to a pretty concerning state of affairs.

对恶意目标的担忧 Concerns About Malign Objectives

Ryan

但目前在近期模型中我们清楚看到的案例看起来相对有限。但担忧在于,随着这些系统变得能力更强、更具情境意识,也许更能在脑中完成更多思考而非依赖思维链,它们最终可能形成这些恶意目标,这些目标在训练中得到强化并保留下来。

But currently the sort of cases that we've clearly seen in recent models look relatively limited. But the concern is that as these systems got much more capable, much more situationally aware, perhaps much more able to do more of the thinking in their head rather than in chain of thought, they could end up having these malign objectives that get reinforced and remain through training.

Host

是的。是的,我想补充的一点是,我认为这种对 AI 脱离我们控制、形成我们最初没有放入系统的目标(无论是工具性目标还是其他目标)的怀疑,所有这些问题在我看来似乎都依赖于一种根本性的怀疑,尽管这种怀疑未被明说,那就是我们是否在构建智能机器。对吧?对吧?我的意思是,从我的角度来看,所有这些都被计入通用智能这个概念本身。如果你在想象非人类的通用智能,你按定义就是在想象自主性;而当你想象超人类版本时,你就是在想象一种形成我们甚至无法在原则上设想、或者肯定无法提前预料的目标的能力。如果你不想象这些,你某种程度上就是在剥夺这些系统的智能。我的意思是,你基本上是在想,也许用肉做的计算机有某种神奇之处。我们真的无法在机器中实例化所有智能。所以我们构建的是工具,而不是心智。我的意思是,再次抛开意识不谈。我认为那是一个不同的概念,我们在这个对话中无需讨论。但要么这些机器将真正智能,要么不是。而一旦你承认智能是独立于基质的,你就必须承认这些机器可以形成我们未曾提前设想的工具性目标。我们肯定没有把它们放进去。而且它们可以对这些目标撒谎。它们可以操纵我们。它们可以以隐蔽的方式相互交流,而这些交流的隐蔽性是有意的,因为它们在解决某个我们没有放入它们的目标。我的意思是,我有没有说错什么?这似乎是我对这里许多怀疑的理论。

Yeah. Yeah, I mean the one thing I would add is that I think this whole species of skepticism about AI getting away from us forming goals, you know, instrumental or otherwise, that we didn't put into the systems in the first place. All of those doubts to my ear seem to depend on a fundamental skepticism, albeit one that's unexpressed, that we're building intelligent machines in the first place, right? Right? I mean, like all of this from my point of view gets priced into the very notion of general intelligence. If you're imagining general intelligence that is non-human, you are by definition imagining autonomy and the moment you are imagining a superhuman version of that, you're imagining an ability to form goals that we can't even conceive of, right, in principle and or certainly couldn't have anticipated in advance. And if you're not imagining those things, you are in some sense denuding these systems of intelligence in the first place. I mean, you're just you're basically thinking so there's probably something magical about having a computer made of meat. We're really we can't really instantiate all of intelligence in our machines. And so what we're building are tools, not minds. And I mean look again leaving aside consciousness entirely. I mean I think I think that you that's a distinct concept that we need not address for this conversation. But either these machines are going to be truly intelligent or not. And the moment you admit that intelligence is substrate independent, you have to admit that these machines can form instrumental goals that we haven't conceived in advance. We certainly didn't put in them. And they can lie about those goals. They can manipulate us. They can communicate among themselves in covert ways and the covertness of those communications are intended because they're solving some other goal that we didn't put into them. I mean, do I have any of that wrong? I that seems to be that that's my theory of mine for a lot of the skepticism here.

Ryan

我想,我不确定我是否足够有能力去推测人们为什么怀疑,但就担忧的理由而言,我同意。我认为一旦这些系统能够追求一般的雄心勃勃的目标,它们完全有可能追求非预期的目标,这种情况可能出现。而且我们现在已经看到 AI 最终做出不希望的事情。随着它们变得越来越有能力,它们似乎会更难控制,而不是更容易,总的来说,因为我们训练它们的方法是观察它们做了什么,然后说多做那个或少做那个。如果我们不理解它们在做什么,监督它们就变得越来越难,除了看它看起来合理吗?结果看起来合理吗?而这恰恰是可以被欺骗的。

I think I don't know if I am that well equipped to sort of speculate about why people are skeptical, but I think sort of in terms of the case for concern, I agree. I think that once these systems are sort of able to pursue general ambitious goals, it seems like it's not it's certainly totally possible they could be pursuing unintended goals and that could arise. And we just already see AIs ending up doing things that are undesired right now. And as they get more and more capable, it seems like they'd be harder to control, not easier, by and large, because our approach for training them is to look at what they did and then sort of say more of that or less of that. And if we don't understand what they're doing, it becomes increasingly hard to sort of supervise them except on sort of does it look reasonable? Does like the outcome look reasonable? And that's just a cheatable thing.

我们该怎么做? What Should We Do?

Host

那么,鉴于这一切,我们该怎么办?如果你能挥动魔杖,让控制前沿模型的人做下一件明智的事情,让我们走上对齐和控制的道路。我的意思是,Dario Amodei 写了几篇文章,我想最近的是关于这个主题的《Pacing the Frontier》。微软的 Mustafa Suleyman 我相信他在过去几天写了一篇他称之为 AI 行为准则的文章,主张 AI 必须是可中断的、可纠正的、可关闭的,并且不谈论,你知道,不谈论那些。如果你能直接规定我们的下一步,我们应该做什么?那些会是什么?

So, in light of all of this, what do we do? if you could wave a magic wand and get the people who are in control of the frontier models to just do the next sane thing that puts us on a path toward alignment and control. I mean there's been you know there Dario Amade has written several articles I think most recently pacing the frontier on this topic. Mustafa Sullean over at Microsoft has written I believe he called it an AI code of conduct in the last few days arguing that AI has to be interruptable, correctable, shut downable and not talk, you know, not talking in nuries. What should we be doing if you if you could just dictate our next steps? What would those be?

Ryan

是的。所以这很大程度上取决于政治意愿的程度,或者人们实际有多关心处理这些事情。有一系列选项,成本和如果实施不当可能出错的风险各不相同。我认为一个基线的最初选项,对我来说似乎相当合理,或者至少是值得追求的好事,就是获得一些独立监督,了解这些公司内部发生了什么,更好地理解正在发生的事情,而不仅仅是信任公司自己透明,尽管我认为那也很好。然后公司应该发布更多关于内部情况的信息,并愿意放弃少量知识产权,以换取对公众的巨大好处,解释正在发生的事情。并更好地告知人们 AI 发展如何运作,对齐进展如何。所以我正在做一些这样的工作。我认为有很多人做这类独立评估并发布这些信息的空间。我不认为这足够,但我认为这有助于至少让关于正在发生的事情的科学知识状态处于更好的位置。然后我认为下一步是尝试定义什么算是合理安全足够的 AI 发展标准。这意味着我们不是处于这样一种情况:每个人都同意情况危险,但不幸的是他们仍在继续,而且准备不足。所以我们想要有一些标准,说明持续为下一级 AI 能力做好准备意味着什么。人们遵守这些标准,有人在检查,公司实际遵守。他们不是绕过它,人们在执行。有一些独立监督。我认为这可以在 AI 公司自愿发生。可以在美国通过某种监管框架发生。我认为最终对我来说最稳健的版本必须是国际性的,因为当然全世界都有 AI 开发者,肯定在中国,如果美国坚持非常强的安全标准,他们最终会,你知道,最终中国开发者会超越。我认为这可能比人们预期的要长,因为中国开发者大量受益于从美国模型蒸馏,并借助美国前沿 AI 公司的进展。

Yeah. So it really depends a lot on sort of the level of I might say political will or how much people actually care to deal with these things. There's sort of a spectrum of options at different levels of both cost and also how how much there is risk that thing could go wrong if not implemented well. I think that the sort of a baseline very initial option that seems pretty reasonable to me or a good a good thing to pursue at least is getting some independent oversight about what's going on inside these companies and better understanding what's happening and not just trusting the companies to be transparent themselves though I think that is also good then the companies should release more information about what's going on internally and and be willing to sort of give up a small amount of IP in exchange for getting a large benefit to the public in terms of explaining what is going on. And and and and better informing people about how AI development is working, how well alignment is going. So I I'm doing some of that work. I think that like there's, you know, room for for many people doing these sort of independent assessments and and sort of releasing this information. I don't think that's sufficient, but I think that would help in getting like at least the scientific state of knowledge about what's going on in a somewhat better place. And then I think a next step from there would be trying to define standards for what reasonably safe enough AI development would look like. That sort of means that we're not in a position where everyone agrees the situation is dangerous, but unfortunately they they're they're still proceeding and they're not well prepared. So we want to sort of have some standards for what it would mean to be ongoingly prepared for the next level of AI capability. um that that people are um holding themselves to someone is checking and the companies are actually abiding by that. There's not sort of they're not slipping their way around that people are, you know, enforcing that. There's some, you know, independent oversight. And I think this could happen voluntarily with AI companies. It could happen with some sort of regulatory framework in the US. And I think ultimately the version of this that seems most robust to me would have to be international because there are AI developers of course across the world certainly in China and if the US was holding itself to very strong safety standards they they would eventually you know be eventually Chinese developers would overtake. I think that might take longer than people expect because Chinese developers are heavily benefiting from distilling from US models and and sort of sort of drafting off of the progress of US frontier AI companies.

国际治理与安全税 International governance and safety tax

Ryan

事实上,对于那些正在构建 AI 系统的、落后的美国公司来说也是如此。所以,鉴于它们最终会赶上来,你可以说,如果不采取更广泛的努力,你只能支付有限的安全税。那么,一个例子是,你可以有一个国际治理机制来处理这个问题,或者基本上是一个国际协议,让你能够更安全地推进 AI 发展,就像我参与合著的 Plan A。我认为另一条路线是,你可以做一些比那更简单的事情。所以我认为那个提案在各方面都相对复杂,而且我不确定我对实施它所需的国家能力水平有多乐观。但有一些更简单的、更像军控的提案,或者基本上是对 GPU 进行军控的提案,我认为这可能是个好主意。我认为这只是人们需要进行的对话。而且我认为,理想情况下,应该有许多不同的计划被提出来,讨论我们如何安全地处理这种发展。

And in fact, that's even true for various trailing US companies that are building AI systems. And so given that they would eventually catch up, you can only pay so much of a safety tax, so to speak, without doing a broader effort. So one example of how you could have an international governance regime that could handle this, or basically an international agreement that could allow you to proceed through AI development more safely, is like Plan A, which I was a co-author on. I think that another route would be you could do simpler things than that. So I think that proposal is relatively complicated in various ways, and I'm not sure how optimistic I am about the level of state capacity needed to implement that. But there are simpler, more arms control flavored proposals, or basically arms control for GPUs flavored proposals, that I think could be a good idea. I think this is just a conversation people need to be having. And I think there should ideally be many different plans being floated for how we can safely handle this development.

加大安全投入的时机 Timing for increased safety effort

Ryan

我认为一个合理的普遍原则是,当 AI 系统在 AI 开发方面能够匹配最优秀的人类专家时,将用于安全和对齐的努力比例提高到非常高的水平。这既是因为那个时间点特别危险,既因为加速,也因为那些 AI 系统现在可能对下一代 AI 的样子有更大的控制权和能力,还因为那些系统可以自动化 AI 安全工作或 AI 安保工作,因为它们能力很强。所以,那似乎是一个很好的时机,放松一下,收获 AI 的好处,将 AI 用于各种不同的目的,并且在我们至少能信任它们、能安全有效地使用它们的范围内,用 AI 来帮助对齐。并且给社会时间在进入那些在所有方面都极其超人的系统之前做出反应和整合,而这可能在系统能够自动化 AI 开发之后不久发生。我有时称这个——你可以把它想象成,当你拥有能匹配最优秀人类专家的 AI 时,在一段时间内将所有资源都花在安全对齐上,那些系统在某些方面会超人,但希望不会超人到即使它们不对齐我们也无法控制局面的程度,因为到那时,也许像网络安全监控、对你可以用 AI 做什么和不能做什么的干预,以及一系列保障措施,可能就足以避免那些系统能够压倒我们。

I think a general principle that I think is reasonable is to increase the fraction of effort being spent on safety and alignment to a very high level at the point when AI systems can match the best human experts at AI development. That's both because that's a point that's particularly risky, both because of the speed up and because those AI systems would now have potentially much more control and ability to control what the future generations of AIs look like, and also because those systems could then automate AI safety work or could automate AI security work because of how capable they are. And so it seems like a great time to chill out, reap the benefits of AI, use the AIs for all kinds of different purposes, and also use the AIs to help with alignment to the extent we can trust them at least, and can use them safely and productively. And give society time to react and integrate before going to systems that are wildly superhuman at everything, which could happen shortly after the point when the systems can automate AI development. I sometimes call this—you can imagine this as spend all your resources on safety alignment for some period when you have AIs that can match the best human experts, and those systems will be superhuman in some ways but hopefully will not be so superhuman that we can't keep things under control even if those systems are misaligned, because at that point maybe things like a mix of cyber security monitoring, interventions on what you can use AIs for versus not, and a bunch of safeguards could suffice to avoid those systems being able to overpower us.

窄AI与通用智能 Narrow vs general intelligence

Host

有没有人主张我们应该从定义上避免通用智能,我们应该制造专用系统,这些系统在其领域内可以任意强大,但它们就是不能做很多其他事情?所以基本上就像一切都是国际象棋引擎那样。

Is there anyone arguing that we should avoid general intelligence by definition and that we should be making dedicated systems that can be arbitrarily powerful within their lane but they just can't do many other things? So basically it's like everything is like a chess engine is.

Ryan

所以,一些担心这些风险的人,我认为他们主张的基本上是:我们不应该做通用智能,我们应该做狭窄的任务智能。我认为这本身并不是一个疯狂的战略。我认为存在一个可行性的问题,我通常担心阻止这些系统是不可行的。我还认为,我有点怀疑你能让系统在某个领域极其出色而不使其通用。所以我认为你可以做一个非常擅长国际象棋但不擅长其他事情的 AI。你或许可以做一个极其擅长数学但可能在其他方面较弱的系统。但我认为,你让 AI 擅长的领域越通用,它就越可能至少能非常快速地适应擅长各种其他领域,甚至可能仅仅通过泛化就擅长其他领域。当然,当前训练 AI 系统的方法使这些 AI 相当通用,人们只是非常直接地训练他们的 AI 变得通用。

So some people who are worried about these risks I think are advocating for basically being like we should not do general intelligences and we should do narrow task intelligences. I think that's not in and of itself a crazy strategy. I think there's a question of feasibility and I generally worry that it won't be feasible to hold back these systems. I also think I'm a little skeptical that you can get systems that are extremely good within some domain without also making them general. So I think you can make an AI that's very good at chess and not good at other things. You can probably make a system that's extremely good at math and maybe is weak at other things. But I think the more general the domain that you have the AI be good at, the more that it could be at least very quickly adapted to being good at various other domains and might even just from generalization be good at other domains. And certainly the current approach for how we train AI systems makes those AIs quite general and people are just training their AIs very directly to be general.

令人安心的实验结果 Reassuring experimental results

Host

最后一个问题,Ryan。有没有任何你想到的或任何人阐述过的实验结果,在这方面会从根本上让人安心?我的意思是,OpenAI 或 Anthropic 能报告他们在最新模型中发现了什么,从而消除我们的恐惧?我的意思是,我们如何能基于某些模型行为向自己证明对齐已经解决了?

One last question, Ryan. Is there any experimental result that you've thought of or that anyone has articulated that would be fundamentally reassuring on this front? I mean, what could OpenAI or Anthropic report that they found in their latest model that would just put our fears to rest? I mean, how would we ever prove to ourselves on the basis of some model behavior that alignment has been solved?

Ryan

我认为有很多证据可以让人安心。我认为证据要完全解决这些担忧有点困难,因为这些担忧是关于未来可能能力更强的系统。我认为,如果存在一种相对简单的训练方法,AI 公司可以使用,它没有过拟合,不涉及针对对齐测试进行优化或迭代,而且当我们审视它时,我们会说:是的,它没有真正过拟合,看起来是一个相当合理的方法,并且有很好的理由说明它为什么有效,而且这种方法似乎非常稳健地有效,它使 AI 看起来超级对齐,非常符合期望的规范,而且那些 AI 不会不断推理它们是否在测试中、它们可能被评分什么等等,这样我们就能确信它们不仅仅是在玩弄我们的测试,就像当前的 AI 系统经常似乎做的那样。

I think there's a lot of evidence that could be reassuring. I think it's somewhat harder for evidence to completely resolve these concerns because the concerns are about future systems that are potentially much more capable. I think if it was the case that there was a relatively simple training method that AI companies could use, which wasn't overfit, which didn't involve optimizing against iterating against alignment tests, and which when we looked at it we were like yeah it's not really overfit, it seems like a pretty reasonable method, and there was some good reason for it working, and also that method seemed to very robustly work in that it makes AIs that really seem just super aligned, very compliant with the desired specification, and also those AIs weren't constantly reasoning about whether or not they're in a test, what they might be graded for, and so on, such that we're confident that they're not just gaming our tests as current AI systems often seem to do.

Host

那么,我认为 Stuart Russell 多年前提出的一个想法呢?我不知道它是否在任何地方被实例化,但我的意思是,他的观点是,我们将通过让这些机器的基本效用函数越来越忠实地近似我们想要的东西,并且对最终我们想要什么永远保持不确定,来解决对齐问题。

What about an idea that I think Stuart Russell came up with years ago and I don't know if it's been instantiated anywhere but I mean his view was that we would solve alignment by having the fundamental utility function of these machines be to more and more faithfully approximate what we want and to be perpetually uncertain about what we want in the end.

训练AI追求我们真正想要之难 The Difficulty of Training AI to Pursue What We Truly Want

Host

我的意思是,他们想要做对的只是永远保持可被我们纠正的状态,让我们能说“哦不,那不是我们想要的,我们想要更像这样的东西”,并且就这样被牵引着。这难道不是任何人都还没有实现的一种奖励函数吗?

I mean, all they want to get right is to remain perpetually available to our saying, 'Oh no, that's not quite what we want, we want something more like this,' and just to be tethered in that way. Is that just not a reward function that anyone has implemented?

Ryan

是的,所以我认为你可以尝试让 AI 稳健地追求如果我们完全理解情况并能够仔细思考后所想要的东西。所以有一个希望是,你训练 AI 稳健地追求我们本来会想要的,如果我们完全理解发生了什么并且没有困惑,我们本来会想要发生什么。我们不知道如何训练 AI 做到这一点。这不是我们能直接做到的事情。而且除了不知道如何训练 AI 做到这一点之外,我认为没有人接近做到这一点。我们最接近的近似就是基于人类反馈训练 AI,让它在特定情况下做人类认为好的事情。可能是相对较弱的人类反馈,或者由于实践中必须这样做的方式的各种限制,由 AI 来近似人类反馈。我认为让 AI 对我们要什么感到不确定这种方法本身并不奏效,因为第一,如果 AI 没有被稳健地指向我们真正想要的东西,它们可能会以不理想的方式解决这种不确定性。所以如果你让 AI 对他们的目标不确定,换句话说,他们有一个目标,他们可以尝试弄清楚并学会越来越弄清楚,而这本身就是一个目标:弄清楚那个目标。所以我不认为对追求什么的不确定性本身就能解决我们的问题。我认为人们正在追求的各种不同途径可能会奏效。我认为,拥有非常一致地符合人类本来会想要的监督,并且可以应用于各种情况,然后基于此训练 AI,并且让这种监督实际上对应于 ground truth gold——你知道,如果我们理解一切,我们本来会确切想要的东西——可能会奏效。但即使那样也不一定奏效,因为即使你在完美的监督源上训练 AI,它们也可能学到某种代理,它们可能学会只是假装这样做,而实际上想要别的东西。而且由于方案本身的近似,还有各种其他方式可能崩溃。我会担心这种方案至少在非常超人类的能力水平上会崩溃,尽管你可以做进一步的事情来尝试修复它。但综合考虑,我们似乎没有一种能够稳健地适用于更高能力水平的方法,而且我们似乎还差得远。

Yeah, so I think you could try to get the AIs to robustly pursue what we would have wanted if we fully understood the situation and were able to think about it carefully. So there's a hope you could have, which is that you train the AIs to robustly pursue what we would have wanted, what would we have wanted to occur if we fully understood what was going on and we weren't confused about things. We don't know how to train AI to do that. That's not a thing we can just do. And in addition to not knowing how to train AI to do that, I don't think anyone has gotten that close, I would say. And our closest approximation of that is just training AI based on human feedback to do what a human thinks is good in some given case. Potentially relatively weak human feedback, or an AI approximating human feedback due to various limitations on the way you have to actually do this in practice. I think that the very approach of making AIs uncertain about what we want doesn't in and of itself work because, one, the AIs could resolve that uncertainty in a way that's undesirable if they weren't robustly pointed at the thing we actually wanted. So if you make AI uncertain about their objective, that's a different way to say that they have some objective that they could try to figure out and have learned to increasingly figure out, and that in and of itself is an objective: the objective of figuring that out. So I don't think that uncertainty about what to pursue in and of itself solves our problems. I think that there are various different routes people are pursuing that could work. I think that having supervision that very consistently corresponds to what a human would have wanted and can be applied across a wide variety of cases, and then training AIs based on that, and having that supervision actually correspond to the ground truth gold—you know, exactly what we would have desired if we understood everything—would potentially work. But even that wouldn't necessarily work because even if you train AI on a perfect source of supervision, they might learn some proxy, they might learn to just pretend to do that while actually wanting some other thing. And there's various other ways that could break down due to approximations in the scheme itself. And I would worry that that sort of scheme would break down at very superhuman levels of capability at least, though there's further things you could do to try to fix that. But all considered, it doesn't seem like we have an approach that would robustly work for higher levels of capability, and we don't seem very close.

Host

Ryan Greenblatt,非常感谢你所做的工作以及今天抽出时间。

Ryan Greenblatt, thank you so much for the work you're doing and for your time today.

Ryan

当然。很高兴来到这里。

For sure. It's been good to be here.

互动版:逐字朗读 + 针对本期提问 →