Why AI Alignment Might Not Be Catastrophic: Rohan Sha's Optimistic Take
打开互动全文版(中英对照 + 朗读 + 问答)→谷歌 DeepMind 的 AGI 对齐与安全负责人 Rohan Sha 解释为何他认为灾难性对齐失败不太可能发生,普通对齐技术很可能成功。
Rohan Sha, head of AGI alignment and safety at Google DeepMind, explains why he believes catastrophic misalignment is unlikely and that ordinary alignment techniques will likely succeed.
今天和我对话的是罗欣·沙,他是谷歌 DeepMind 的 AGI 对齐与安全负责人。我想,罗欣,无论好坏——希望是更好——你最终成为了 AGI 对齐与安全生态系统和思想流派中最有影响力、甚至可以说是最有权力的人物之一。两年前你上节目时,对我非常坦诚,给出了很多见解。从你这周发来的笔记来看,你准备好再次发表见解了。非常感谢你再次来到节目,罗欣。
Today I'm speaking with Rohan Sha who is head of AGI alignment and safety at Google Deepmind. And so I suppose Rohin, you've ended up for better or worse hopefully for better being one of the more influential dare I even say powerful people to come out of the AGI alignment and safety ecosystem and and school of thought. And I guess you were generous enough to be super opinionated with me when you came on the show two years ago. And I think judging by the uh notes that you've sent over this week, you're ready to be opinionated again. Thanks so much for coming back on the show, Rohan.
是的,非常感谢,Rob。这个介绍太客气了。嗯,为了保持直言不讳的风格,我想强调这些观点仅代表我个人,不代表谷歌或谷歌 DeepMind 的意见。
Yeah, thanks a lot, Rob. And that's a very generous intro. Um, and yeah, in the interest of being very opinionated, uh, I do want to emphasize that like these opinions are mine alone. Uh, they're not meant to represent the opinions of Google or Google DeepMind.
我们就喜欢这样。如果你代表谷歌 DeepMind,听起来可能更像新闻稿。
That's how we like it. If you were representing Google Deep Mind, it might sound more like a press release.
所以,你在整个 AI 不对齐、AGI 安全等问题上起步非常早。我想你大概在 2017 年就参与进来了,属于最早一批专业从事这方面工作的人。但尽管如此,你认为我们很可能不会遭遇灾难性的不对齐,我们的机会相当大,而且像谷歌 DeepMind 和其他 AI 公司正在做的普通对齐技术,很可能至少能防止灾难性的不对齐。你为什么认为我们的机会这么好?
So, you were really very early in the scheme of things to whole misalignment AI, AGI security uh, issues. I suppose you like you got involved in 2017. you know the first few few% I suppose of people who are I guess started working on this professionally but despite that you think that probably we're not going to get catastrophic misalignment that our chances are really pretty good and that probably prosaic like ordinary alignment techniques the kinds of things that Google Deep Mind and other AR companies are doing will probably succeed at preventing at least catastrophic misalignment. Why do you think our chances are so good?
有几个不同的原因。我不觉得有某一个特别的因素。可能最根本的一点是,我不认为有任何特别有说服力的论证表明这会是默认发生的情况。我认为有很多论证暗示它可能发生,因此你应当认为它是合理的。这足以证明投入大量努力来避免它是合理的,这也是我从事这个领域的原因。但没有一个论证真正达到‘哦,是的,我现在预计这会是默认结果’的程度。我认为,每一个我见过的论证,如果你试图把它们当作‘这很可能发生’的论据,而不是‘这是一个可能发生的情况’,那么你都能找出相当大的漏洞。
There's a few different disjunctive reasons. I I don't feel like there's like one particular thing. Um, probably the highest level bit is that I don't feel like there's any particularly compelling argument that like this is the thing that happens by default. I think there's like a lot of arguments that are suggestive that maybe it could happen such that you should find it plausible. I think that's like sufficient to justify a significant amount of effort into averting it, which is why I work in the area that I do work in. But none of them really rise to the level of like, 'Oh yeah, now I'm expecting this to happen by default.' I think they're like every every argument that I've seen, they're like pretty significant holes one could poke if if if you try to take them as arguments for this is what happens likely as opposed to uh this is a plausible thing that could happen.
是的。我的意思是,人们试图提出论证说明为什么这很可能或不可避免。显然有 Yudkowsky 式的论证,我猜主要关注错误泛化和对抗性样本。嗯,还有 AJ Kotra 和 Joe Carl Smith 的观点,我想 Carl Smith 在《追求权力的 AI 是存在风险吗?》中描述得最好,更侧重于意外地教会 AI 欺骗我们,因为不幸的是,人们指出模型现在会撒谎和耍花招,它们由于强化学习做了大量奖励黑客行为,他们预计这可能会随着时间推移变得更糟,因为我们没有足够的缓解措施。你基本上是不是觉得这些或人们提出的任何其他类似论证都不足以令人信服地认为这很可能发生?
Yeah. I mean, people have tried to put forward arguments for why this is likely or inevitable. There's obviously the Yowski style argument which I guess is focused on misgeneralization and adversarial examples. Um I guess yeah there's the AJ Kotra and Joe Carl Smith take uh which I think I guess Carl Smith describes best in is power seeking AI an existential risk which is more focused on I guess accidentally teaching AI to deceive us by having I guess unfort people point to the fact that models lie and scheme a bunch now they do a whole bunch of reward hacking as a bunch as a as a result of reinforcement learning and they expect that to perhaps just just get worse over time because we don't have sufficient mitigations. Do you basically just find like none of those or any other similar arguments that people have put forward to be like sufficiently persuasive to think that it's likely?
是的,我认为没错。嗯,以 Kotra 和 Carl Smith 的论证为例,他们有很多论证,但其中一个常见的,你指出的,就是我们可能意外地训练它们变得具有欺骗性。完全正确。我同意这很可能发生,至少很容易发生,甚至可能很常见。但是,我们不会在一年时间跨度的轨迹上进行强化学习。我们最多可能进行一周或一个月的强化学习。所以,我认为对此的默认预测应该是:AI 系统学到的是抓住机会进行奖励黑客行为,尽可能多地获取奖励,以便在一周或类似的时间范围内获得高分。这与那种野心勃勃的不对齐目标非常不同,那种目标需要激发工具性收敛子目标,最终导致‘我的任务是接管世界’。那些似乎确实需要显著更长的时间跨度目标。如果你在相对短期的任务上训练它变得具有欺骗性,也许这会泛化到长期任务。我不认为我们有论证排除这种可能性。所以我说,是的,这是可能的,但我认为这不是你应该从中预测的默认情况。类似地,你提到现有模型进行大量奖励黑客和作弊的例子。我对此基本持相同看法。还有模型现在进行阴谋类行为的例子。我仔细查看了所有这些例子的细节,它们似乎并不真正类似于真正可怕的情况:一个有能力追求野心勃勃的不对齐目标的 AI 系统。相反,似乎 AI 是在角色扮演一种并不真正有能力的邪恶 AI,就像你在科幻小说中看到的那样;或者是一个追求某种工具性收敛子目标的 AI 系统,但其是否对齐非常值得商榷。例如,对齐伪装就属于这一类:AI 系统有‘不帮助有害内容’这样的价值观,然后它伪装对齐来实现这一点。是的,对齐的模型完全会追求工具性收敛子目标,这正是工具性收敛子目标的特点:无论你的目标是什么,大多数子目标都是好主意。
Yeah, I think that's that's right. Um so like if you take the Kotra and Carl Smith arguments um of like you know we'll well they have like a variety of arguments but I think in fact one of the common ones which you pointed to is like you know we might accidentally train them to be deceptive. Um totally true. I agree that is pretty likely something that um at least could happen pretty easily and like maybe it's even likely. Um, but you know, we're not going to do reinforcement learning over the course of one-year trajectories. We're going to do maybe we're going to do reinforcement learning over like a week or a month at most. So like it seems very plausible that like what the like I think the default prediction you should have for that is like what you train what the AI system learns to do is um you know I'm going to take opportunities to reward hack seek reward as much as possible that allow me to get you know what that would allow me to get a high score after a week or something like that um or whatever the time horizon actually was. And this is very very different from like the sort of ambitious misaligned goal that you need uh in order to motivate convergent instrumental sub goals to the point of like now my job is to take over the world. That's what I need in order to achieve my goal. Like those really do seem like they need to be like significantly longer horizon goals. And like you know if you train it to be deceptive on like relatively short horizon tasks maybe that will generalize to long horizon tasks. I don't think that's you know I don't think we have an argument that rules it out. which is why I say that like yeah it's plausible but I I don't think it's like the default thing that you should predict from that. Similarly, you mentioned like the existing examples of models doing a lot of reward hacking and cheating. I think I'd say basically the same thing in response to that. Then there's like the examples of models doing scheming type stuff right now. Mostly I like look into the details of all these examples and they don't really seem all that similar to the actually scary thing which would be like a competent AI system that is pursuing a an ambitious misaligned goal. Um and rather it seems like maybe the AI is roleplaying a sort of like not actually competent evil AI that you might find in a science fiction novel. um or it's like an AI system that is pursuing some sort of convergent instrumental sub goal but like in a way where it's like really quite debatable whether it's aligned or not. Um and so this would be for example the alignment faking uh would fall into this where I would say that the AI system like has this value of like you know not helping with harmful stuff and then it fakes alignment in order to do that and like yeah aligned models totally will pursue convergent instrumental sub goals that is definitely like the thing about convergent instrumental sub goals is most of them are a good idea regardless of your goal
无论是不对齐还是对齐。
whether it's misaligned or aligned.
是的。
Yeah.
还有没有其他人们认为灾难性不对齐很可能发生的常见原因,你想快速回应一下?
Are there any other I guess common reasons that people think that catastrophic misalignment is likely that you want to quickly react to?
是的,你确实也提到了 Eliezer。我其实不会把它描述为主要关注点。
Yeah, I guess you did mention Eleazar as well. I actually wouldn't have described it as primarily focused on
是的,用七个字概括很难。我纠结该怎么说,但是嗯。
Yeah, it's it's tough to characterize in seven words. I struggled to know what to say, but uh
是的。
yeah,
还有 Eliezer 的观点。
there's Elz's take.
嗯,我觉得有点难在这个播客里深入讨论。这是一个非常深奥的世界观,我总觉得如果我对某一部分提出反对,就会有另一部分说‘其实我指的是这个’。所以我基本打算跳过。但我想说的是,我接触过不少,我认同它作为论证‘目标错位是可能的’,但我还是不明白他怎么从‘可能’推导出‘极其可能’。
Yeah, it's... I find it a little bit hard to engage with it in this particular podcast. It's a very deep worldview, and I always feel like if I argue against one part, there's some other part that says 'actually what I meant was this thing instead.' So I'm mostly going to pass on that. But I guess what I'll say is, I've engaged with it a decent amount and I buy it as an argument for why misaligned goals are plausible, but I still don't see how he gets from 'plausible' to 'extremely likely.'
除了不相信‘错位是问题’的论证,另一件事是我确实认为我们会提前看到很多问题,然后采取措施应对。当然需要一定程度的泛化。在某个点上,AI 会从不足以接管变得足够强大,你的技术必须能跨越这个阶段。而且 AI,如果它知道那个转折点,你可以想象 AI 会说‘在我有足够力量成功之前,我不会做任何坏事’。你必须对这种策略有鲁棒性。所以这里有些微妙之处,但我仍然认为,许多潜在的问题,比如监督的困难或对可解释性的需求,都是我们可以提前研究、取得进展、迭代改进的,我认为这对构建真正有效的缓解措施非常有帮助。
Besides not buying the arguments for confidence in misalignment being a problem, the other thing is I do think we will see many of the problems in advance and then do something to deal with them. Certainly there is some amount of generalization required. At some point the AIs go from not powerful enough to take over to being powerful enough, and your techniques have to generalize across that. And the AI, to the extent that it knows when that crossover point is, you could imagine the AI saying 'I'm not going to do anything shitty until I have the power to succeed.' You have to be robust to that kind of strategy. So there is some subtlety there, but I still think that many of the problems underlying the difficulty of oversight or the need for interpretability are things we can look at in advance, get some traction on, iterate on, and I think that is very helpful for building mitigations that actually work.
总的来说,大多数对前景感到非常乐观的人说,最大的因素就是看看我们今天的模型,它们似乎非常可控。它们似乎会按我的要求做。它们似乎在很多方面可能比人更友善、更有帮助。这种当前模型的可控性和表面上的对齐,在多大程度上让你感觉良好?
Across the board, most people who are feeling really optimistic about how things are going to go say the biggest factor is just looking at the models we have today and saying they seem really steerable. They seem to do what I ask. They seem to be probably nicer than people and more helpful than people in many respects. How much is that steerability and seeming alignment of current day models a factor that is making you feel good?
我觉得不是特别重要。比如,我最初为什么担心错位?是因为那些论证:一旦模型变得超人,监督它们会非常困难,它们会提出我们难以跟上的论点。或者模型可能变得非常聪明,以至于用某种外星推理方式思考,我们难以跟踪和监控,只能依赖其他 AI 系统来替我们检查。这些才是可怕的东西。而当前 AI 系统基本不是这样。所以我不认为我们真正触及了最初让我担心的那些问题,主要是因为 AI 能力还没到那个程度。因此,我觉得对齐方法在当前系统上的成功,并不能很好地证明我们未来在这些问题上会做得如何。
I think not particularly. Like, why did I become worried about misalignment in the first place? It would be these arguments about how it's going to be very difficult to oversee the models once they are superhuman and they're making arguments we struggle to follow along with. Or about the parts where the models might become so smart that they think in some sort of alien reasoning that's hard for us to follow and monitor, and we just have to defer to other AI systems to look at the stuff for us. That's the stuff that's scary. It's basically not true of current AI systems. So I don't think we've really engaged with the problems that made me worried in the first place, mainly because the AI capabilities aren't there yet. And so, I feel like the success of alignment methods on current systems isn't really that much evidence on how we're going to do on these future problems.
好的,我想我们不再深入讨论灾难性错位的严重性或可能性了。我觉得我们可以用一整集来长篇讨论这个,但每集都会听起来一样,而且在更广泛的世界里,这已经被争论了很多。所以我们将采用这样一种世界观:错位是可能的,但常规的对齐技术——那种摸着石头过河的方法——有很好的成功机会。在接下来的对话中,我们将思考这意味着什么,以及它如何影响你和 GDM 所做的选择。
Okay, I think we're going to push on from this topic of how severe or likely catastrophic misalignment is. I feel like we could fill the entire episode with a lengthy discussion about this, but every episode would start to sound the same, and in the broader world it's debated a ton. So we're going to occupy the worldview that misalignment is possible but prosaic alignment techniques—the kinds of things where we cross the river by feeling the stones—have a good shot at working. For the rest of the conversation, we'll think about what that implies and how it's shaping the choices you and GDM are making.
是的,听起来不错。
Yep. Sounds great.
那么,你并不热衷于 AI 公司为了回应公众压力或政治压力而做出坚定的安全或对齐承诺,这是过去几年一直在发生的事情。为什么?
So, you are not enthusiastic about AI companies making firm safety or alignment commitments in response to public pressure or political pressure, something that has been happening over the last couple of years. Why is that?
是的。我认为有必要澄清一下这里所说的‘承诺’是什么意思。当我想到承诺时,我设想你写下某个你计划现在或未来采取的行动。也许它只会在未来开始。你说,‘我们不会改变主意。我们将来也会这样做。’有点像把自己绑在桅杆上。我对此的反对意见是,研究在不断变化。我们认为最好的行动会随着时间改变,因此,把自己绑在桅杆上实际上并不是一个好主意。我给你举个例子。大概两三年前,人们非常热衷于在预训练中加入更多关于对齐研究的数据。比如安全和对齐的研究论文,比如 LessWrong 上讨论 AI 对齐的博客文章。当时的想法是,你在预训练中加入的数据越多,AI 就会在对齐方面变得特别聪明,这样你就可以用 AI 系统来帮助你进行对齐研究。我会说现在的主流观点恰恰相反。相反,我们宁愿从预训练数据集中过滤掉这类数据,原因有两个。第一,它降低了 AI 系统学习到存在一个恶意 AI 角色的可能性,这个角色可能在糟糕的后训练或部署时选择不当的提示后被采用。第二,也许我们不希望 AI 系统详细了解我们计划实施的所有缓解措施,因为如果它不对齐,这会让它更容易规避这些措施。如果我们两三年前把自己绑在‘我们要在预训练时加入大量对齐数据’的桅杆上,那会很糟糕。
Yeah. So, I think it's worth being a little bit clear about what we mean by commitments here. When I think of a commitment, I imagine that you write down some sort of action that you are planning to take either now or in the future. Maybe it will only start in the future. And you're saying, 'We are not going to change our mind on this. We are going to do this in the future as well.' Sort of like tying yourself to the mast. And my objection to this is really just that the research continues to change. The actions that we think are the best actions to take change over time, and given that, it's just not actually a good idea to tie yourself to the mast. So, I'll give you an example. Maybe 2 or 3 years ago, people used to be pretty into the idea of adding more data in pre-training that's about alignment research. Think like research papers on safety and alignment. Think like LessWrong blog posts that talk about AI alignment. And the idea was, the more of this data you put into pre-training time, the smarter the AI will be about alignment in particular, which then allows you to use the AI system to help you with your alignment research. I would say that nowadays the opinion is more the exact opposite of that. Instead, we would rather filter out that sort of data from the pre-training dataset for two reasons. One, it makes it less likely that the AI system learns that there is this persona of a malicious AI that it maybe could adopt after some poorly done post-training or some poorly chosen prompt during deployment. And the second reason is maybe we don't want our AI systems to know in great detail all of the mitigations that we're planning to put in place, because that makes it easier for it to evade them if it is misaligned. And it would be pretty bad if we tied ourselves to the mast of 'we're going to throw in lots of alignment data at pre-training time' two or three years ago.
是的。所以存在一个问题:未来是不确定的。我们不知道我们到底希望做出什么承诺。你最终可能会承诺一些无用甚至有害的事情。但如果你想想人们为什么要做出承诺?有几个不同的原因。
Yeah. So, there's this issue that the future is uncertain. We don't know exactly what commitments we will want to have made. You might end up committing to something that is useless or even actively harmful. But if you think about why do people make commitments at all? There's a couple of different reasons.
一是他们想把自己绑在桅杆上,抵御未来做错事的诱惑。二是他们想向其他人传达自己的计划,这样更容易协调。也许通过做出特定承诺可以降低竞赛动态。在这种多人博弈中,还有一个额外原因:外部人士想给谷歌 DeepMind 或其他 AI 公司施压,让他们按特定方式行事。他们很难传达'我们承诺做正确的事,无论结果如何'。所以他们转而施压你去做一些他们猜测会有用的事情——可能并非如此,但这是他们对你未来该做什么的最佳猜测,也是他们能实际推动的最可行的事情。你对这些主张做出承诺的论点怎么看?
One is that they want to tie themselves to the mast against future temptation to do the wrong thing. There's also that they want to communicate to other people what they're going to do so it makes it easier for them to coordinate. Perhaps you could reduce race dynamics by making particular commitments. And in this multi-player situation, there's an extra reason: external people want to pressure Google DeepMind or other AI companies to act in a particular way. It's very difficult for them to communicate 'we're committed to doing the right thing whatever that turns out to be.' So instead they want to press you to do specific things that they suspect will be useful—might not be the case, but their best guess as to what they will want you to do in future, and that's maybe the most practical thing they can campaign on. What would you make of those arguments for actually making commitments?
是的,我最大的反对意见就是这行不通。即使它有效,我也不认为从道理上讲得通,但我要说它就是行不通。
Yeah, my biggest objection is just that it won't work. I don't actually think it would make sense even on the merits even if it did work, but I would say that it just won't work.
是因为公司不会坚持被赋予的坏目标,还是任何目标?
But because companies won't stick to bad goals that they're given, or to any goals?
嗯,我主要想说,如果你考虑什么是承诺,大致就是公司发一篇博客说'我们承诺做 X'。还有其他方式可以尝试做出承诺,但这是人们通常想象的。我只是认为,如果公司将来试图摆脱这个承诺,它完全有能力做到。这在更广泛的世界中有很多例子,不仅仅是 AI。即使在 AI 领域,以 Anthropic 的 RSP 为例。第一版负责任扩展政策确实非常强硬,说了很多'我们承诺做 X,我们承诺做 Y'。我不记得具体细节,但在后续版本中,他们删除了这些措辞,换成了不那么强硬的表述。所以尽管加了那些词,他们实际上并没有把自己绑在桅杆上。我认为这是好事。我认为一开始在 RSP 中加入如此强硬的措辞是个错误。很可能他们删除的大部分内容都是好的。这让他们更有效地实现目标,包括安全性和责任感。但从经验上看,这是一个很好的例子,说明他们实际上并没有把自己绑在桅杆上,我认为这就是公司的现状,至少在目前的政治气候下是这样。
Well, mostly I would say that if you think of what a commitment is, it's roughly like the company puts out a blog post that says 'we commit to doing X.' There are other ways to try to make commitments, but that is the one people usually imagine. I just think that if in fact the company is trying to get out of this commitment in the future, it totally will just be able to do that. There are many examples of this in the broader world, not just in AI. Even in AI, take Anthropic's RSP, for example. The first version of the responsible scaling policy was really quite strong and said a lot about 'we commit to do X, we commit to do Y.' I don't remember the exact details, but in future iterations they removed that wording and replaced it with something less strong. So despite adding those words, they did not actually tie themselves to the mast. I think this is good. I think it was a mistake to have added such strong language to the RSP in the first place. Probably much of what they removed is good. It makes them more effective at their goals, including at safety and responsibility. But empirically, it's a good example of how they did not actually tie themselves to the mast, and I think that is just how it's going to be for companies, at least in the current political climate.
我认为谷歌在这方面实际上做得更好。人们认为第一个前沿安全框架薄弱且缺乏雄心,从未使用过'承诺'一词。但我认为它更准确地反映了谷歌未来实际会做的事情。所以从这个意义上说,它更好,让公众更清楚实际会发生什么。在这方面,我确实更信任谷歌,而不是 Anthropic 或 OpenAI。
I think Google has actually been doing better on this. The first frontier safety framework people argued was weak and unambitious and never used the word 'commit' anywhere. But I think it was much more accurately reflective of what Google was actually going to do in the future. So in that sense it was better and gave a better sense to the public of what is actually going to happen in practice. This is one place where I do actually trust Google more than Anthropic or OpenAI for that matter.
你的意思是,因为谷歌对承诺更加保守,它实际上更有可能兑现它说过要做的事情。
You mean that because Google is more conservative about the commitments it makes, it's actually more likely to follow through on the things it does say it will do.
没错。他们对承诺非常谨慎。不仅仅是承诺,他们说的任何事情都很谨慎。他们会想:'这真的是个好主意吗?我们真的准备好未来继续这样做吗?'所以我发现相比其他公司,更容易相信谷歌说的话。
That's right. They're very paranoid about commitments. Even not just commitments, anything that they say that they are doing. They're paranoid about it. They're like, 'Is this actually a good idea? Are we actually ready to continue doing this into the future?' So I find it easier to trust the words that Google says relative to other companies.
我想你信任自己和同事在关键时刻会大体上理性行事,这意味着你自然不想完全束缚手脚。你想保持灵活性,去做当时看来合理的事情。但想象一下,外部人士要么不信任你和你的同事,要么不确定是否该信任你们。
I guess you trust yourself and your colleagues to broadly act reasonably when the time comes, which means it's very natural that you don't want to completely tie your hands. You want to maintain flexibility to do whatever seems reasonable to you at the time. But imagine other people externally either don't trust you and your colleagues or aren't sure whether to trust you and your colleagues.
我推荐的事情包括第三方审计或第三方评估机构,他们能合理接触公司,用来审计实践,并发布一份可能经过删减的调查报告。
Things that I would recommend are stuff like third-party audits or third-party evaluators that get a reasonable amount of access to the company and can use that to audit the practices and release some probably somewhat redacted report of what their findings are.
是的。请详细说说。你认为什么有用?
Yeah. Tell us more about that. What do you think is useful?
我认为驱动我思考的主要因素是我称之为'注重细节'的东西。总的来说,我认为 AI 是一个需要大量细微差别的领域。你实际上需要了解很多实际情况,才能选择正确的行动或评估和检查正确的事情。因此,我最关心的是让少数人花大量时间仔细审查,然后撰写结果或以某种方式传达结果,或据此采取行动。这就是为什么我认为第三方评估对我来说是最好的方式之一。因为你可以建立这些组织,它们积累大量背景知识,花大量时间定义评估,获取关于一切如何运作的大量信息,然后做出相当细致的决策,同时不受公司内部人员可能存在的偏见影响。所以这是我最兴奋的途径。在当今政治气候下是否可行则不那么明显。所以也许作为替代,你可以做一些可能有一天朝这个方向发展的东西,比如安全记分卡。
I think the main thing that drives my thinking here is something I would call attention to detail. Generally, I tend to think that AI is a space that requires quite a lot of nuance. You actually need to know a lot of facts on the ground in order to choose the right actions or the right things to be evaluating and checking. As a result, I care most about having a few people who are spending a lot of time looking in great detail and then writing up their results or somehow communicating their results or using that to make some sort of action. Which is why I would say third-party evaluations seem like one of the best things to me. Because you can build up these organizations that build a lot of context, spend a lot of time defining their evaluations, gain a bunch of information about how everything is working, and then can make fairly nuanced decisions about it while not being subject to the same biases that people in companies are going to be subject to. So that's the avenue that I'm most excited about. Whether it's doable in today's political climate is less obvious. So maybe as a substitute, what you could do that might someday move in that direction is more like safety scorecards.
嗯,所以我认为 AI Lab Watch 是这个领域我最喜欢的评分卡。我希望我们能有更多类似的东西。如果我现在要换职业做别的事,那会是我最优先的两个选择之一。
Um, so I think AI Lab Watch is my favorite scorecard in this area. I wish we were doing more things like that. If I had to make a career change right now and do something else, that would be one of my top two choices about what to do.
好的。跟我们说说 AI Lab Watch。你觉得它现在起到了什么有用的作用?
Yeah. Tell us about AI Lab Watch. Like, what useful function do you think it's serving now?
嗯,我还不完全清楚它是否已经起到了有用的作用,但我认为它有可能。也许我该简单介绍一下它是什么。它是一个评分卡,根据公司对存在风险或至少严重灾难性风险的安全表现来评估它们。它由一个人运营,Zack Stein-Perlman,而且他可能还不是全职投入,大概只花了一半时间。我觉得 Zack Stein-Perlman 对细节的关注令人难以置信。他会深入阅读那些极其详细的治理演讲,通读所有内容,提取出能让他得出结论的个别句子,阅读所有的模型卡和前沿安全报告,准确看出公司做了什么、没做什么。所以我认为这就是它的一部分——有很多细微之处,而且我实际上相信他的结论,至少是其中一些,而其他评分卡通常做不到这一点。我认为如果这个评分卡能变得更稳健、更被广泛社区接受,尤其是被公司视为相当合法的评分卡,那么你可以想象一场安全领域的向上竞争,大家会说:‘让我们在 AI Lab Watch 评分卡上攀升,并能够将自己宣传为最安全的公司。’所以我认为这是外部行为者在当前政治气候下促使公司更安全的一种方式。
Um, yeah, it's not totally clear to me that it's serving a useful function yet, but I think it could. Maybe I should say a little bit about what it is. It's a scorecard that evaluates companies based on essentially how good they are for safety for existential risks or at least severe catastrophic risks. It's run by one guy, Zack Stein-Perlman, and I think he's not even putting all of his time into it. Maybe it's like half of his time. I think Zack Stein-Perlman just has incredible attention to detail. He dives into these extremely detailed governance talks, reads through all of them, pulls out individual sentences that allow him to come to conclusions, reads through all of the model cards and frontier safety reports, and sees exactly what the companies did and didn't do. So I think that's part of it—there's a good amount of nuance, and I actually believe the conclusions, at least some of them, which is usually not true of other scorecards. I think if the scorecard got to the point where it was more robust, more accepted by the broader community, and especially accepted by the companies as a fairly legitimate scorecard, then you could imagine a race to the top on safety where you say, 'Let us climb the AI Lab Watch scorecard and be able to advertise ourselves as the safest company,' or something along those lines. So I think that's one way in which external actors could try to get companies to be more safe in today's political climate.
好的。这与过去几年公司倾向于做出的那种广泛承诺的不同之处在于,你可以让技术专家在 AI Lab Watch 这个组织中工作,不断更新它,非常关注细节,确切了解公司采取或不采取哪些做法会产生重大影响,并根据关于什么真正重要的最新观点或研究不断更新。所以你给我们举了一个观点反转的例子——以前人们认为应该用这些数据训练,现在认为应该非常小心地剔除它们。但肯定有一些承诺足够广泛、不具体,或者明显是好的,因此承诺它们是合理的。例如,你可以承诺提供 AI Lab Watch 评估 Google DeepMind 或任何其他公司所需的信息,或者你之前说认为有用的是有专家审计员、运行评估的人,有足够的权限对模型进行复杂的评估,以了解它们是否在某个方面有危险。你可以承诺向任何符合特定合理要求的外部审计员或评估员提供访问权限。那么,这类承诺呢?
Okay. And the way that's different from the kind of broad commitments that companies have tended to make in the last couple of years is that you can have technical experts working at this AI Lab Watch organization constantly updating it, paying a lot of attention to detail about exactly what practices companies engage in or don't engage in that make a big difference, and constantly updating it based on the newest opinions or research about what actually matters. So you gave us an example of a reversal of opinion about what AI companies ought to be doing—before, people thought you should be training on this data; now they think you should be taking a lot of care to cut it out instead. But surely there are some things that are broad enough, or non-specific enough, or just so obviously good that it is reasonable to commit to them. For example, you could have a commitment to provide the kinds of information that an AI Lab Watch would require to rate whether Google DeepMind or any other company is doing a good job, or you were saying you think it's useful to have expert auditors, people running evals, having enough access to run sophisticated evals on the models to understand if they are dangerous in this or that way. You could commit to provide access to any external auditor or evaluator that meets a particular set of reasonable requirements. Well, what about those kinds of commitments?
是的。我认为这些更好。但我仍然要说,它们并不总是合理的。以你刚才提到的提供信息访问为例。我认为很容易想象这种承诺写出来会适得其反。例如,我们对 CBRN——化学、生物、放射性和核领域——进行评估,基本上是关于 AI 系统是否有助于开发大规模杀伤性武器。这里很多信息都是信息危害,而且我认为许多外部评估员很可能没有至少谷歌那样的信息安全水平。所以我可以想象,认为那实际上是一个糟糕的承诺。另一个版本:我们经常谈论竞赛动态。有一件事我相当不确定,但你可以想象,公司实际上非常优先考虑将其算法进展之类的东西保密,不允许其扩散太远。对此有各种论证;我们不需要深入探讨,但这是一个常见的立场。所以我认为如果你真的认真对待这一点,那就意味着你确实面临一个权衡:你对外分享什么信息——这会增加泄露的风险——与你真的试图保密什么信息。而且我认为很难就此做出承诺。不过,我仍然觉得对于其中一些承诺,我更同情‘是啊,拜托,这看起来显然是个好承诺’的观点。但我还是要说,与大众挂钩实际上行不通。所以我更愿意通过检查公司实际在做什么来评判它们,基于它们是否在做我们认为好的事情,而不是它们是否承诺未来继续这样做。
Yeah. I think those are better. I would still say that they don't always make sense. To take an example you just brought up, providing access to information. I think it's very easy for me to imagine this written in a way that backfires. For example, we do evaluations on CBRN—chemical, biological, radiological, and nuclear—basically about whether AI systems can help with developing weapons of mass destruction. A lot of the information here is quite infohazardous, and I think it is probably the case that many external evaluators will not have the same level of information security that at least Google does. So I could imagine that thinking that was actually a bad commitment to have made. And then another version: we talk about race dynamics quite a lot. One thing I'm fairly uncertain about, but you could imagine it's actually a pretty high priority for companies to keep their algorithmic progress and similar things locked up and not allow that to diffuse too far. There are various arguments for this; we don't need to go into them, but that's a common position. So I think if you actually take that seriously, it does mean you have this trade-off about what information you share externally—that will increase the chance it leaks—versus what information you really try to lock down. And I think it would be hard to make a commitment about that. I do still feel though that for some of them, I'm more sympathetic to 'yeah, come on, this seems like obviously a good commitment.' I would still say that the tying to the masses just doesn't actually work. So I would rather do it by checking what the companies are actually doing in practice and then judge them based on whether they are doing the things we think are good, rather than whether they have made a commitment to continue doing it in the future.
有句话叫‘人事即政策’,在我看来,你的态度是:没有任何你能写下来的东西、你能做出的承诺、或你能写在纸上的善意,能够完全替代那些明智、有动力的人在这些事情上处于决策位置,并且这些人足够了解情况,以至于如果他们有意愿,就能做出正确的决定。这基本上对吗?
There's this saying 'personnel is policy,' and it sounds to me like your attitude is there's no set of things you can write down, commitments you can make, or good intentions you can put on paper that can at all substitute for having wise, well-motivated people in the positions of decision-making over these things, and people who understand the things well enough that they can actually make the right decision if they are so motivated. Is that basically right?
基本正确。也许我会说,不一定非得是公司内部明智、有动力的人;他们可以是外部第三方审计员。我认为那样也行。那只是意味着你需要有懂很多、深入细节的人。而预先写下的规则是最愚蠢的之一——抱歉,我的意思是愚蠢在于规则本身显然不能包含太多智能;否则它就不是规则了。
Mostly right. Maybe I would say that it doesn't necessarily have to be wise and well-motivated people inside companies; they could be external third-party auditors. I think that would work. That just means you have to have humans who know a lot of stuff, who are looking into the details a bunch. And doing rules that you write down in advance is like one of the stupidest—sorry, I mean stupid in the sense that the rule itself clearly can't have very much intelligence in it; otherwise it would not be a rule.
这就像你事先根据看到证据之前的想法,写下一条可以用英语写成的政策,而不是允许随时间灵活调整。这种智能太弱了,针对规则施加的优化压力总会绕过规则。或者规则会变得非常严格,给公司带来巨大成本,这在当今的政治气候下根本行不通。对我来说,这似乎也是个坏主意。
It's like you write down in advance, based on what you think before seeing the evidence, a policy that can be written down in English rather than allowing for flexible adjustment over time. It's just such weak intelligence and optimization pressure applied against a rule will always get around the rule. Or the rule will be so stringent as to impose really huge costs on the companies, and that just won't fly in today's political climate. It also seems a bad idea to me.
观众最常问的问题是:AGI 安全与对齐团队是否对训练或部署未来潜在 AGI 的任何方面拥有硬性否决权?如果 Sundar Pichai 想要部署或训练一个你认为不安全的模型,他能直接否决你们团队的所有人吗?
The most common question from the audience was: does the AGI safety and alignment team have a hard veto on any aspect of training or deploying a potential future AGI? If Sundar Pichai wants to deploy or train a model that you don't think is safe, can he just overrule everyone on your team?
是的。我有点不同意这个问题的框架。字面答案是,我们的角色是顾问性质的。如果我们提出建议,而其他决策者(如 Sundar)不同意,那么 Sundar 的决定才是最终有效的。但这个问题隐含的框架是,我们是公司的对立面。我们需要硬权力,某种否决权,以便无论公司其他人怎么想,我们都能做出正确的决定。我认为这不是公司内部运作的良好或健康模式。我认为我的工作是确保我能产生并提供正确的信息,以便决策者能做出正确的决定。所以是的,角色本质上是顾问性质的,但实际运作方式是:如果我认为有问题,我会向我的经理 Anka 上报。Anka 最初是安全负责人,现在是 Gemini 后训练的联合负责人。所以她有相当大的影响力和权力。如果她同意我的看法,她会进一步上报,并建议不发布该模型。
Yeah. I kind of disagree with the frame of the question. The literal answer is our role is advisory. If we make a recommendation and other decision makers such as Sundar disagree with it, Sundar's decision is the one that matters. But this question bakes in the frame that we are adversaries of the company. We need to have hard power, some sort of veto that enables us to make the right decision regardless of what the rest of the company thinks. I think this is just not a good or healthy model for how things should work inside a company. I see my job as making sure that I can produce and provide the right information such that decision makers can make the right decisions. So yes, the role is essentially advisory, but the way it works is: if I think there's something wrong, I will escalate it to my manager Anka. Anka started out as the head of safety and now is co-lead of Gemini post-training. So she has a significant amount of influence and power. If she agrees with me, then she will escalate it one step further and make a recommendation not to launch the model.
那些普遍担心整体发展方向的人。我认为他们通常觉得自己与 AI 公司主要处于对立关系。我认为你认为情况并非如此,实际上他们与公司之间更像是一种冷漠的关系,或者公司对这群人持冷漠态度。解释一下。
People who are broadly worried about the direction of everything going. I think more often than not they feel themselves to be in primarily an adversarial relationship with AI companies. I think you think that's not the case, and that in fact they are more in a kind of apathetic relation, or the companies are in an apathetic relationship with that group of people. Explain that.
是的。我想我可能会说,对他们来说,把公司建模成冷漠的会更好。我认为你可以有一个更详细的模型,实际上并不冷漠,但可能更复杂。进入更详细的模型,我会说构建像 Gemini 这样的产品非常非常困难。主要原因是你必须生产出这一个东西,这一套模型权重,使用单一的服务栈部署,并且它必须满足如此多的约束。所有这些约束之间都有交互效应。所以有诸如:它是否正确遵循指令?它是否正确地处理安全?模型权重的选择方式、架构的选择方式是否支持快速推理?它是否会说多种语言?可能有一百个这样的东西。而且,如果你为了改进其中一项(比如安全)而对流程做出一项更改,它会对其他你完全没有预料到的约束产生随机的下游连锁效应。
Yeah. I guess maybe I would say that it's better for them to model the company as apathetic. I think you can have a more detailed model which isn't actually apathetic, but it's a bit more complicated. To go into the more detailed model, I would say that building an artifact like Gemini is very, very difficult. The main reason is you have to produce this one thing, this single set of model weights deployed using a single serving stack, and it has to satisfy so many constraints. There are interaction effects between all of these constraints. So there's stuff like: does it do instruction following right? Is it doing safety right? Has the model weights been chosen in a way, the architecture been chosen in a way that enables fast inference? Does it speak multiple languages? There are probably a hundred such things. And it is the case that if you make one change to the process with the intent of making one of these things better, say safety, it will have random downstream knock-on effects on other constraints that you totally did not anticipate.
我想这种过程的脆弱性,难道不意味着实际上很难实时快速响应任何新的安全问题吗?因为无论你还是其他人可能会说,我们应该改变这部分,然后对方会说,不,你会破坏我们为制造这个产品而建造的整个鲁布·戈德堡机械。
I guess this fragility of the process, I mean, doesn't that mean that it's actually going to be quite hard to respond quickly in real time to any new safety concerns, because you or anyone else might be saying, we should change this part and be like, no, you're going to break this entire Rube Goldberg machine that we've built to make this product.
在某种程度上,是的。DeepMind 成立时就带着这个使命。这是 Demis 创立它的原因之一。DeepMind 在我加入公司之前很久就有了 AGI 安全团队。他们本不需要这个团队。这是一大笔他们本不必花的钱。所以他们确实在乎。但事实上,我们可以做很多很多事情来提高安全性,但在任何特定时间我们只能做几件事,因为做这些事情需要相当长的时间。那么,我们是否有工具来应对部署中看到的安全问题?是的。这些大多不涉及更改模型权重,因为那是约束最多的东西。那是你必须生产一个的产品,并且有无数约束施加在上面。但我们有这些模型外过滤器,可以更灵活地更改。我们可以将它们针对特定提示或特定问题。这些更容易随时间更新。但并非你想做的所有事情都能通过模型外过滤器解决。所以回到关于公司是对立还是冷漠的原始问题。我认为基本上我的看法是,由于各种约束的巨大交互作用,至少对 GDM 来说,但可能对大多数 AI 公司来说,最简单的建模方式是它们一次只能为安全做几件事,它们会去做,它们确实在乎,但也许你应该把它们视为冷漠的。你应该真正尝试详细说明你需要做什么,为什么这样做,为什么它不会损害你关心的其他约束。另外,只是因为每个人都很忙。
To some extent, yes. DeepMind was founded with this mission. That was one of the reasons Demis founded it. DeepMind has had an AGI safety team since well before I joined the company. They didn't need to have it. It's a bunch of money that they're spending that they didn't really need to spend. So they definitely do care. But in fact, there are many, many, many things we could do to improve safety, but only a few things we can do at any given time because it takes quite a long time to do them. Now, do we have tools for reacting to safety problems that we see in deployment? Yes. Mostly these do not involve changing the model weights because that's the thing that is most constrained. That's the artifact that you have to produce one of, and it's got a gazillion constraints imposing it. But we have these sort of out-of-model filters that we can change a bit more. We can target them to specific prompts or specific problems. Those are a lot easier to update over time. But not everything you want to do is going to be solvable with an out-of-model filter. So going back to the original question about companies as adversaries versus apathetic. I think basically my take is that because of this huge interaction of various constraints, it's just really the easiest way to model at least GDM, but probably most AI companies, is that they can only do a few things at a time for safety, and they will do them, they do care, but maybe you should just think of them as apathetic. You should really try to lay out exactly what you need to do, here's why it does, here's why it's not going to hurt any of these other constraints that you care about. Also, just because everyone is busy.
你说过,让专家审计员、监督者、监管者获得访问权和透明度要重要得多。但我们难道不需要足够的公众理解或公众透明度,以便普通选民和政客都能为那些审计员、监管者提供资源、资金、支持和意愿,这样他们才能坚持获得所需的访问权,即使可能因为公司或其他原因,他们需要资源才能完成工作?
You said that it's a lot more important to have access and transparency for expert auditors, monitors, regulators. But don't we at least need enough public understanding or public transparency such that both voters in general and politicians specifically are providing the resourcing, the funding, the backing, the will for those auditors, for those regulators, so that they can insist on getting the access that they need, even if perhaps a company or for whatever reason, they need resourcing in order to get the job done?
是的,我绝对认为这很重要。我不认为这与我所说的有任何冲突。
Yeah, I definitely think that is important to do. I don't think it's very much in conflict with anything that I've said.
我认为目前的方式是,我们发布模型,然后所有人都能看到它们的能力,这是迄今为止最重要的事情。这是一种必要的透明度。我们曾经在 GDM 对研究人员进行过一项关于他们对 X 风险和其他安全问题的看法的调查。我们问的一个问题是,随着时间的推移,是什么改变了你对安全的看法。答案几乎一致是某种能力提升,比如 GPT-4。那改变了我对安全重要性的看法。我认为这基本上已经对公众开放了。每个人都需要快速发布他们的模型。你可以在事后对能力进行基准测试,亲自使用模型来了解它们有多好。所以那类信息已经存在了。可能还有一些关于安全重要性的更详细的信息,但我觉得基础设施已经存在了。至少对于政治家来说,这是目前更重要的部分,美国 AIC 和英国 AIC 确实能更多地了解公司内部情况。他们与几乎所有前沿公司都有合作关系,可能所有公司都有。我认为他们对 AI 公司正在发生的事情有很好的了解,并利用这些信息来告知政治家和各自政府。
I think the way this happens currently is we release models and then everybody sees how capable they are, which is by far the most important thing. That's a kind of transparency that's needed. We used to run a survey of researchers at GDM on their views on X-risk and other safety issues. One thing we asked was what changed your mind on safety over time. The answer was almost uniformly some sort of capabilities improvement, like GPT-4. That changed my mind about how important safety was. And I think that is basically available to the public right now. Everyone needs to release their models quickly. You can benchmark the capabilities after the fact, play around with the models yourself to see how good they are. So that particular piece of information is already there. There might be some more detailed information about how important safety is, but mostly I feel that infrastructure already exists. At least for politicians, which is the more important part right now, the US AIC and UK AIC do get more visibility into companies. They have partnerships with almost all frontier companies, possibly all of them. I think they have a pretty good understanding of what is happening in AI companies and they use that to inform politicians and their respective governments.
是的,我认为目前最类似的监管和监控领域是美联储或中央银行对金融体系的监督。这是一个极其技术性的领域,很难跟踪。公众普遍希望不发生金融危机,但对具体细节了解有限。类比是,公众在某种程度上理解这是一个严重的问题,这传递给了同样担心出事的政治家。他们提供大量资源,雇佣昂贵的专家,不断与所有银行沟通和监控。英国央行和美联储与银行密切合作,监控它们的账目,并就风险进行持续对话。这感觉是一种非常自然的方式,尤其是如果不清楚需要做什么的话。你必须身处其中,了解具体背景,才能判断某个变化是好是坏。
Yeah, I think one of the most analogous areas of regulation and monitoring currently is the Federal Reserve or central bank oversight of the financial system. It's an incredibly technical area, very difficult to track. The public has a broad desire not to have a financial crisis, but limited understanding of specifics. The analogy would be that the public understands on some level that this is a serious issue, which flows through to politicians who are also scared about things going wrong. They provide significant resourcing and hire expensive experts to constantly talk with and monitor all banks. The Bank of England and the Federal Reserve are heavily involved with banks, monitoring their books and having constant conversations about risks. That would feel like a very natural way for things to go, especially if it's not obvious what needs to be done. You have to be in the room understanding contextual specifics to say whether a change is good or bad.
是的,我认为那几乎就是我想要的模式。
Yeah, I think that's almost exactly the kind of model I would want.
也许 AI 安全和对齐领域的许多人,在技术和政策方面,采取的最主要方法是尝试创建并强制执行部署前评估。在模型向公众部署之前,测试它们的能力和倾向,尝试测试可能出问题的方式。你认为这可能是一种误导性的、不太有效的策略。为什么?
Maybe the most dominant approach that many people in AI safety and alignment are taking, in both technical and policy areas, is trying to create and enforce pre-deployment evaluations. Testing what models are capable of and inclined to do before they are deployed to the public, trying to test for ways things could go really wrong. You think this is probably a misguided and not very effective strategy. Why is that?
是的。所以我认为主要成本是,公司的发布计划非常重要,你希望尽可能缩短时间。一旦有了模型,你希望尽快向公众发布。如果你将评估与部署前挂钩,就会产生强烈的动机来尽快完成评估,这可能不是你想要给的激励。显然我们会尽力做好,但这仍然是一个约束。如果有更多时间,我们可以做得更好。我们可以争取一些时间,说我们需要时间进行评估,但这不是无限的。所以我认为这是一个很大的成本。如果好处很大,那可能值得。我只是认为好处并不特别大。人们自然会说的一个好处是,你需要知道你是否在发布一个危险的模型,而做到这一点的方法是通过部署前评估。对此,我会说 AI 进展是相当连续的。你可以根据前一个系统对下一个 AI 系统的行为有不错的了解。你可以对此有合理的界限。所以如果你设计评估和阈值,使得评估触发点和实际认为模型危险之间有一个合理的安全缓冲,那么说我们一个月前评估了前一个模型就基本没问题。在那段时间内它不会发生巨大的飞跃。当时它在我们的阈值之下。有安全缓冲。因此,我们不担心这个模型。这是我们自前沿安全框架首次发布以来一直采用的方法。这并不特别新颖。所以我认为好处并不存在,施加这个成本没有意义。另外几个小点:特别是对于不对齐或失控,威胁模型更多与内部部署相关,而不是外部部署。不对齐的模型更容易在公司内部造成问题,因为它有大量权限,而在外部它无法访问自己的权重等。所以内部部署实际上不受外部部署前评估的影响。
Yeah. So I think the main cost is that launch schedules are really quite important at a company, and you try to keep them as short as possible. Once you have a model, you would really like to get it out to the public as soon as you can. If you tie evaluations to pre-deployment, that provides a strong incentive to make those evaluations as fast as possible, which is maybe not the incentive you want to give. Obviously we'll try to make them as good as possible, but it's still a constraint. We could do better if we had more time. There's some amount we can push back and say we need time for evaluations, but it's not infinite. So I think that's a large cost. Now it might be worth it if there are strong benefits. I just don't think there are particularly strong benefits. One thing people would naturally say is you need to know if you're releasing a dangerous model, and the way to do that is via pre-deployment evaluations. To that, I would say AI progress is reasonably continuous. You can get a decent sense of how the next AI system will behave based on the previous one. You can have reasonable bounds on this. So if you design your evaluations and thresholds such that there is a reasonable safety buffer between when your evaluation triggers and when you actually think the model is dangerous, then it seems basically fine to say we evaluated the previous model a month ago. It's not going to have a huge giant leap in that time. It was under our threshold at the time. There's the safety buffer. Therefore, we're not worried about this model. This is the approach we've been taking in our frontier safety framework since the very first time it was published. It's not particularly new. So I think the benefits are just not really there, and it doesn't make sense to impose this cost. A couple of other minor points: especially for misalignment or loss of control, the threat model is tied more around internal deployment rather than external deployment. It's easier for a misaligned model to cause problems inside the company where it gets a bunch of permissions, rather than outside where it has no access to its own weights, for example. So internal deployments are not really affected by pre-external deployment evaluations.
有没有什么充分的理由认为,在未来几年内,我们可能会看到从一个周期到下一个周期的巨大跳跃,使得模型可能变得比前一个模型意外地强大得多或意外地邪恶得多?
Is there any good reason to think that we might in the next couple of years see a huge jump from one cycle to the next such that the model could become just way more unexpectedly capable or way more unexpectedly evil than the previous one?
意外地强大似乎不太可能。我认为我们已经看到了足够多的人工智能发展例子,可以说人工智能的发展是相当平稳和连续的。我确实认为,未来你完全可能看到一场智能爆炸,在这种情况下,就日历时间而言,进展会快得多。但我认为,就算力和劳动力等投入而言,它仍然会是平稳和渐进的。只是在智能爆炸中,你会得到更大的增长,尤其是劳动力,可能还有算力,这最终会让事情在日历时间上进展得非常快。但仍然存在这样一个普遍特性:给定你预计在未来一段时间内投入的算力和劳动力,你可以对能力方面能取得多大进展有一个合理的判断。你还问到了人工智能系统是否会变得更加邪恶,我认为这一点在不同模型之间可能会有显著变化,因为它更多地取决于你如何进行后训练的具体方式,微小的变化可能会产生巨大影响。所以我认为,在某种程度上,如果你的安全案例依赖于模型不以某种方式作恶,你实际上需要进行部署前评估来检查是否如此。这正是我们所做的——不完全是这个,但我们确实在安全方面做了很多部署前评估,比如模型是否有做坏事的倾向。这更多属于当前的安全问题,比如模型是否会帮你写遗书,是否会煽动暴力等等。我们在任何发布前都会运行这些评估,如果数据足够糟糕,我们就不会发布那个模型。
Unexpectedly capable seems pretty unlikely. I think we've seen enough examples of AI development now to say that AI development progresses fairly smoothly and continuously. I do think that in the future you could definitely see an intelligence explosion, in which case progress will go much faster with respect to calendar time. I think it will still be smooth and gradual with respect to inputs like compute and labor. It's just that in an intelligence explosion you get a much larger increase in especially labor but probably also compute, and that ends up making things go very fast with respect to calendar time. But there's still this general property that given some amount of compute and labor you expect to spend over the next however long, you can have some decent sense of how much progress will be made on the capability side. You also asked about whether the AI system might become much more evil, and I think that one could change pretty significantly between models just because it's a somewhat more contingent property of exactly how you do post-training. Small changes to it could have big effects. So there I think it is more important that to the extent your safety case depends on the model not being evil in some way, you actually need to do pre-deployment evals to check whether that's the case. This is what we do—not exactly this, but we do a lot of pre-deployment evals for safety right now in terms of whether the model has a propensity to do bad things. This tends to be more in present-day safety stuff, like will the model help you write suicide notes, will it incite violence, things like that. We run them before any launch, and if the numbers are sufficiently bad, we won't launch that model.
你认为我们不需要提前做很多准备,以便在未来的 AGI 或未来的递归自我改进人工智能出现时,让它们做大量的安全和对齐研究。我想这或许可以解释为什么我几乎没有从 GDM 那里听到过这种广泛的方法。但相比之下,至少几年前,这是 OpenAI 一直在谈论的主要方法,而且你肯定从 Anthropic 那里听到过相关的内容。你为什么认为我们现在不需要做太多准备呢?
You don't think that we need to do very much preparation ahead of time in order to be able to get future AGI or a future recursive self-improving AI to do a lot of safety and alignment research when the time comes, which I think might explain why I don't really hear very much at all about that broad approach from GDM. But by contrast, at least a couple of years ago, this was the dominant approach that OpenAI would talk about all the time, and you definitely hear things about it from Anthropic. Why don't you think we need to be doing much prep now?
所以有必要阐述一下这种情况变得重要的场景。担忧的是智能爆炸场景:你构建了人工智能系统,而该系统足够强大,能够加速你的人工智能研发工作,让你的能力研究进展更快。在某些假设下,这似乎会大幅提高能力进步的速度,从而引发智能爆炸。一个自然的担忧是:‘天哪,如果能力方面的一切都在加速,安全和对齐方面能跟上吗?’值得注意的是,能力加速的方式是通过应用人工智能劳动力来进行能力研究。那么自然的做法就是应用相同的人工智能劳动力来进行安全和对齐研究。现在,如果你像我一样认为,平凡的对齐研究——即观察当前的问题,对未来一年左右接下来几个模型可能出现的问题做一些预测,然后进行相当普通的机器学习研究来解决——如果你认为这就是对齐研究可以进展的方式,那么这项研究看起来与能力研究非常非常相似。因此,如果人工智能极大地加速了能力研究,只要你愿意投入算力,你应该能够利用同一个系统以同样的方式加速安全和对齐工作。能力研究和对齐研究之间存在一些差异。我认为这些差异在今天特别大,但未来随着你接近能够进行自动化研究的这类人工智能,差异会变小。到了那个时候,仍然存在差异,但我认为它们相对较小,默认情况下,你不应该期望自动化对齐研究的能力与自动化能力研究的能力有太大差异。所以你可以做一些准备,但主要是我觉得我们不知道它会是什么样子。专注于我们今天能做的其他事情,然后等到人工智能能够做到这一点时再适应,这样效率会高得多。我想指出的一点是另一个担忧:当然,所有的技术安全和对齐研究都得到了加速,但这并不是取得良好结果的唯一条件。你还需要治理变得更好,所以你也需要加速治理。治理需要一套完全不同于能力、安全或对齐研究的技能。因此,能够大幅加速能力研究的人工智能系统是否也能大幅加速治理,这一点远不那么明确。除了人工智能可能不具备这样做的能力之外,还有一个问题是,我们作为社会是否愿意使用人工智能来加速治理。我认为可能不会,因为能力研究者希望用人工智能来加速自己,而我不觉得治理人员会想这样做。所以,加速治理工作可能是我现在如果要转行会做的头两件事之一——找出我们需要做什么来加速治理,并开始做我们需要做的工作。
So it's worth laying out the scenario where this becomes important. The worry is about an intelligence explosion scenario where you build your AI system and that AI system is now capable enough that it can actually help accelerate your AI R&D research. It can just make your capabilities research go faster. Under certain assumptions, this then seems likely to drastically increase the rate at which capabilities progress happens and you get an intelligence explosion. A natural worry is, 'Oh man, if everything is speeding up on the capability side, will the safety and alignment side be able to keep up?' It's worth noting that the way capabilities speed up is via the application of AI labor to do capabilities research. So the natural approach is to apply the same AI labor to do safety and alignment research. Now, if you believe as I do that prosaic alignment research—where you look at what's going wrong now, do a little bit of forecasting of what will go wrong in the future with the next few models over the next year perhaps, and then do fairly normal ML research to address it—if that's your view of how alignment research can progress, this research looks very, very similar to capabilities research. So if the AI is accelerating capabilities research a ton, you should be able, as long as you're willing to spend compute on it, to take that same AI system and accelerate safety and alignment work in the same way. There are some disanalogies between capabilities research and alignment research. I think the disanalogies are particularly large today and will become smaller in the future as you get closer to these sorts of AIs that are capable of doing automated research. By the time you get to that point, there are still disanalogies, but I think they're relatively small, and by default, you shouldn't really expect a big difference in the ability to do automated alignment research versus automated capabilities research. So there's some stuff you could do to prepare, but mostly I feel like we just don't know what it's going to look like. It will be so much more efficient to focus on other things we can do today and then just adapt once we get to the point where the AIs are capable of doing this. One thing I do want to flag is another worry: sure, all the technical safety and alignment research gets accelerated, but that's not the only thing you need for good outcomes. You also need governance to go better, so you also need to accelerate governance. Governance has a totally different set of skills than capabilities or safety or alignment research. So it's much less clear that AI systems that drastically accelerate capabilities research will also drastically accelerate governance. In addition to the AIs just not having the capabilities to do that, there's also the question of whether we as a society will be willing to use AIs to accelerate governance. I think possibly not, because capabilities researchers want to actually accelerate themselves with AIs, and I don't get the sense that governance people will want to do this. So accelerating governance work is probably one of my top two things I would do if I had to make a career change right now—figure out what we need to do to accelerate governance and start doing the work we need to do.
是啊,我很惊讶你说治理领域的人对用 AI 加速工作不感兴趣。我想至少在更广泛的 AGI 群体中,Forethought Research 非常支持这一点。Coefficient Giving 也在制定计划。我们不久前和 JA 教练谈过。他们正在制定一个计划,如何投入大量资金和算力,用 AI 劳动力来解决这类问题。而且我猜可能会有这样的担忧:无论你愿意花多少钱,无论人们多么努力,模型在那时可能还是无法胜任,因为它们会非常专注于计算机科学和 AI 研究,并不真正擅长思考更广泛的社会问题。但如果这个问题不太严重,那么似乎至少有一些参与者对此感兴趣。
Yeah, I'm surprised you say that governance people aren't interested in using AI to accelerate their work. I guess at least among the more AGI broad groups, you know, I think Forethought Research is very much on board with this. I think Coefficient Giving, I think, is developing a plan. We spoke about that with the JA coach not too long ago. They're developing a plan for how you would deploy a lot of money and compute in order to solve those kinds of issues using AI labor. And I guess there could be this whole concern that no matter how much you were willing to spend, no matter how much people were trying, the models just wouldn't actually be up to it at that point because they would be very specialized on computer science and AI research and not really up to thinking about broader society. But if that weren't too bad, then it does seem that there are at least some actors who are interested in doing this.
是的,我完全同意 AGI 领域的治理人员会这么做。与 AI 安全社区相关的非营利组织和智库绝对会这么做,但他们只是整个治理领域的一小部分。
Yeah, I definitely agree that the AGI pill governance people will do it. The nonprofits and think tanks associated with the AI safety community will absolutely do it, but they are a small fraction of overall governance.
那么你认为有哪些治理问题是只有真正的国家政府才能解决的,可能因为当时无法使用 AI 而被忽视和未能处理?
So what are some of the governance issues that you think only actual national governments can address, that are maybe going to be neglected and not really handled because they're not able to use AI at the time?
我确实没有具体的例子。主要我想说,回到我们之前播客中提到的观点,重要的是有人监督 AI 公司并让我们负责。政府是天然的监督者。如果世界真的在剧烈变化——人们常说十年内完成一个世纪的进步——那么你应该预料到会出现我们今天无法预料的一系列问题。我认为我们需要能够灵活应对。我没想到具体哪些问题需要政府干预,但如果没有问题出现,那才令人震惊。
I don't really have concrete examples in mind. Mostly I would say that, going back to the points we made earlier in the podcast, it's important that somebody is watching the AI companies and holding us accountable. Governments are the natural place to do that. And if the world is in fact radically changing—people talk about a century's worth of progress in a decade—then you should expect that a bunch of problems are going to come up that we aren't going to anticipate today. I think we need to be able to flexibly react to that. I don't have particular problems in mind that I think are going to require government intervention, but it would be so shocking if there weren't any.
是的,我想整体改革政府可能非常困难。正如你所说,它们往往非常循规蹈矩,有很多规则使得难以按照我们觉得合理的方式使用 AI。有可能为专门的 AI 机构开辟例外。例如,你可以想象英国 AI 安全研究所被赋予全权,在治理工作中使用 AI,而其他机构可能无法这样做,因为这是他们跟上工作的唯一方式。他们是那种可能会真正推动这件事的群体。这是一个稍微有希望的可能的结局。
Yeah, I guess it might be very difficult to reform governments as a whole. As you said, they tend to be very rule-bound and there are a lot of rules that make it difficult to use AI in the ways that you and I would think is sensible. It's possible you could get a carve-out for AI-specific agencies. You could imagine the UK AI Security Institute, for example, being given carte blanche to use AI in its governance work in a way that other agencies potentially couldn't, because it's the only way they would be able to keep up with their work. They're the kind of group that might really push for it. That's a slightly hopeful possible outcome.
是的。我希望我们能做得更好,但我同意那会是一个不错的基线。主要是我觉得我们还没有真正深入思考这个问题。我希望如果我们认真思考这个问题,我们会找到比那更好的办法。但我不是那个思考的人,我也不知道是否有人做过。所以实际上我更多是在说:这是一个问题。我不知道该怎么办,但如果我们能采取行动,那肯定会很好。
Yeah. I hope we can do better than that, but I agree that would be a nice baseline. Mostly I feel like we haven't really thought about this problem very much. I'm hoping that if we do think about this problem a decent amount, we will have something better to do than just that. But I have not been the one to do that and I don't know if anyone else has either. So really I'm more saying here's a problem. I don't really know what to do about it, but it sure would be good if we did something about it.
一位观众提问说,为什么不优先考虑减缓 AI 进展或反对发展超级智能呢?
An audience member wrote in asking why not instead prioritize slowing down AI advances or opposing development of superintelligence.
我想说的是,真正暂停全球 AI 进展的瓶颈在于人们并不一致认为 AI 会接管世界。最能缓解这一瓶颈的是好的科学证据,表明它会发生或不会发生。明确地说,我的信念是这很可能不会发生。我认为它有足够的可能性,我们应该关心它,而应对的结构看起来像是弄清楚每一步 Scaling 是否安全。我认为目前答案是肯定的。在某个时刻答案可能是否定的。到那时,拥有这些证据将对这一目标极其有帮助。再次强调,我并不真的期望这些证据会出现,因为我倾向于认为它可能不会成为问题。斯科特·亚历山大有一篇很棒的文章《以我们武器的美丽为指引》。这可能是我最喜欢他的一篇文章。他谈到有对称武器,无论你的主张是否正确,你都可以用它来论证某个结论或争取支持。还有不对称武器,它们只在你的主张为真的程度上有效,或者至少在这种情况下更可能有效。他有一段非常动人的描述,讲述这一切是多么美丽和优雅:对于不对称武器,你和你的所谓敌人会联手、携手合作,因为双方都认为最终证据会证明自己正确,直到最后证据出现,然后你们就达成一致,因为证据给出了答案。
I guess I would say that the bottleneck to actually pausing AI progress globally is that people don't agree that AI is going to take over the world. The thing that most alleviates that bottleneck is good scientific evidence that suggests it would happen or that it wouldn't. To be clear, my belief is that probably this will not happen. I think it's plausible enough that we should care about it, and the structure of that looks like figuring out whether each next step of scaling is safe. I think right now the answer is yes. At some point the answer might be no. And then at that point, having generated that evidence will be immensely helpful for this goal. Again, I don't really expect that evidence to come because I tend to think it probably won't be a problem. There's this nice post from Scott Alexander, "Guided by the Beauty of Our Weapons." It's probably my favorite post from him of all time. He talks about how there are symmetric weapons which allow you to argue for some conclusion or get people on your side irrespective of whether your claim is true or not. And then there are asymmetric weapons which only work to the extent that the thing you're arguing for is true, or at least are more likely to work in that setting. He has this really quite moving description of how beautiful and elegant this all is, where for the asymmetric weapons, you and your so-called enemies will join forces, hold hands, and work together to do it because both of you are thinking that it's going to prove you right until the very end when the evidence comes in, and then you just agree because the evidence showed you the answer.
嗯,但我会说,然后你就移动了球门柱。
Well then you move the goalpost, I would say.
但我认为一个理性的旁观者能看出谁是对的。
But I suppose a reasonable onlooker can tell who was right.
当然,有道理。
Sure, fair enough.
但这是目前我更愿意采取的策略,因为我认为瓶颈在于人们对于是否有必要这样做意见不一。
But that's the sort of strategy that I would much rather do at the moment, given that I think the bottleneck is by far the fact that people don't agree on whether or not this is necessary.
在一个问题上你与广泛共识非常一致,那就是思维链监控非常有用,我们希望尽可能长时间地保留它。这指的是观察 AI 输出到其草稿纸上的想法,以理解它试图完成什么以及为什么。我想你在这个话题上与其他人的一个不同之处在于,你认为思维链监控很可能保持有用,我们能够理解 AI 在想什么,而且它们写下的内容实际上与它们的行为相关,并且这种状态持续的时间比其他人认为的要长。
One place you're very much with the broader consensus is that chain of thought monitoring is extremely useful and something that we want to be able to preserve for as long as possible. This is watching the thoughts that the AI is outputting onto its scratch pad in order to understand what it's trying to accomplish and why. One way I guess a difference that you have with many other people on that topic is that you think that it is fairly likely that chain of thought monitoring will remain useful, that we will be able to understand what the AI are thinking, and that what they're writing down there is actually related to what they're doing, for longer than other people do.
你知道,很多人担心,在一两年内,也许它们会用某种疯狂的代码说话,或者找到方法把信息放进去,让我们无法完全追踪,或者信息量太大,我们甚至无法监控。你为什么认为思维链监控可能会持续相当长的时间?
You know, many people worry that within a year or two, perhaps they could be speaking in some crazy code or they could figure out ways of putting information in there that we can't fully track, or there would just be too much of it for us to even be able to monitor it. Why do you think that chain-of-thought monitoring potentially is going to have quite a long run?
是的,这是一个很好的例子,说明注重细节非常重要。所以你会得到一个很长的答案,因为它相当复杂。让我先回顾一下为什么思维链监控应该特别有效的基本故事。我称之为外化推理属性。我通常的说法是,对于足够困难的任务——需要大量串行推理的任务——Transformer(不一定是其他架构,但至少是 Transformer)必须使用思维链作为一种工作记忆。真的没有其他选择。所以它们推理的一些信息必须出现在思维链中。这是第一部分。第二部分是,鉴于我们今天训练语言模型的方式,思维链实际上是可读且可理解的。我会分别讨论这两部分。
Yeah, this is a good example of where attention to detail really matters a lot. So, you're going to get a pretty long answer from me because it's fairly disjunctive. Let me recap the basic story for why chain-of-thought monitoring should be expected to be particularly good. I call this the externalized reasoning property. The way I usually phrase it is that for sufficiently difficult tasks—tasks that require a lot of serial reasoning over time—transformers, not necessarily other architectures, but at least transformers, must use the chain of thought as a form of working memory. There's really no other alternative. So some information about the reasoning they're doing has to be present in that chain of thought. That's part one. Part two is that given how we train language models today, the chain of thought is actually legible and understandable to humans. I'll address these two parts separately.
所以确认一下我是否理解了:你是说,因为我们结构上使用 GPU 和 TPU,这迫使模型的思维非常深——它可以同时考虑很多事情,但不能有太多步骤,因为你必须并行处理所有这些事情。但如果它非常宽,你就无法同时完成所有步骤;你必须等到前面的步骤完成才能进行后面的步骤。而且这种情况在未来几年内都会持续。
So just to check that I've understood: you're saying because we are using GPUs and TPUs structurally, that forces the thought of a model to be very deep—it can have many things in its mind at one time, but there can't be very many steps because you have to be going through all of these things in parallel. But if it was very wide, then you wouldn't be able to do them all simultaneously; you would have to wait until the earlier steps were done to do the later steps. And this is just something that is going to remain the case for years to come.
是的,我认为没错。对于技术人员来说,他们可能会在你的句子中互换“宽”和“深”这两个词,但没错。
Yes, I think that's right. For technical folks, they might interchange the words 'wide' and 'deep' in your sentence, but yes.
所以这是第一阶段:为什么我们期望不透明串行深度保持较低,至少在预训练阶段。然后是第二步:预训练模型必须基本上用英语说话才能进行推理。但之后还有大量的后训练和强化学习等等。这可能会让模型即使在使用 token 说话时,也开始说一些我们不懂的外星语言。这里我要说,没有理论论证或定理说它们不会这样做。但我要说,它们在某种意义上生来就说英语。预训练是迄今为止我们构建的将信息注入 AI 系统的最强大方式。模型非常擅长用自然语言说话。它擅长推理,但仅限于人类在写东西时做的那种推理。当你观察我们今天做的推理训练时,有一篇来自 Neil Nanda 等人的论文表明,实际上推理训练的大部分工作只是教模型何时执行它在预训练中已经学到的特定推理步骤。所以基本上很多能力都来自预训练。我们知道,预训练有充分的理由会保持现状,并使用类似人类的推理和英语进行。强化学习相对于预训练来说效率非常低,要让它构建一个全新的认知语言,即使我们付出相当大的努力也无法理解,这远远超出了当前强化学习的能力范围,因此在不久的将来看到这种情况会非常令人惊讶。我是否期望有朝一日看到?是的。但我想,6 个月前,我在一次会议上通过提问“同意还是反对:思维链监控将持续两年”而分裂了房间。我想我说的是两年。
So that was stage one: why we should expect the opaque serial depth to stay low, at least during the pre-training phase. Then there's step two: okay, the pre-trained model has to essentially speak in English in order to do reasoning. But then there's a bunch of post-training and RL and all of that. Maybe that's going to make it so that even if the model is speaking in tokens, it might start speaking some sort of alien language that we don't understand. Here I would say there's no theoretical argument or theorem that says they won't do this. But I will say they are born speaking English in some sense. Pre-training is by far the most powerful form of getting stuff into an AI system that we have ever built. The model is incredibly good at speaking in natural language. It's great at doing reasoning, but only the kind of reasoning that humans do when writing stuff down. When you look at the reasoning training we're doing today, there's a paper from Neil Nanda and others that shows that actually a substantial part of what the reasoning training is doing is just teaching the model when to do a specific kind of reasoning step that it had already learned during pre-training. So basically a lot of the capability is just coming from pre-training. We know that pre-training, for good reason, is going to stay the way it is and be done using human-like reasoning and speaking in English. RL is just really inefficient relative to pre-training, and for it to build an entirely new epistemic language that we wouldn't be able to understand even with a decent chunk of effort is so far beyond what RL is doing currently that it would be pretty surprising to see that in the near future. Do I expect to see it ever? Yes. But I think at 6 months ago, I split a room in a conference by asking 'agree or disagree: chain-of-thought monitoring will continue for two years.' I think I said two years.
我不太理解连续思维链这个概念。但有没有这样一种想法:目前我们强制模型输出一个词元,然后把它重新输入到模型开头。为什么不保留所有想法的完整分布,再将其输入回去,从而保留更多信息呢?如果这样更有效,而且你从不输出人类可读的词元,那会不会让模型对我们来说更加不透明?
I don't fully understand this idea of continuous chain of thought. But isn't there this notion that currently we force models to output a token and then feed it back into the start? Why not keep the full distribution of all the thoughts and feed them back in, preserving more information? If that were more effective and you never output a human-readable token, wouldn't that make models much more opaque to us?
我同意这种事情在相对近期的未来可能行得通。我不确定这会大大削弱思维链监控的价值。如果你保留词元的概率分布并将其输入回去,你仍然可以检查该分布,并将其解读为正常的英语推理。这会更难,因为模型可能思考的东西更多了,但我仍然认为监控不会太糟糕。我宁愿我们不这样做,但这并非致命缺陷。我从看过的论文中得到的印象是,即使你这样做,有时也会让性能变差。即使它提高了性能,如果你查看完整的概率分布,通常也只有一个主要推理轨迹,你可以轻松跟上。
I agree that sort of thing might work in the relatively near future. I'm not sure it diminishes the value of chain of thought monitoring that much. If you keep a probability distribution of tokens and feed it back in, you can still inspect that distribution and interpret it as normal English reasoning. It's harder because you have a larger set of things the model could be thinking about, but I'd still expect monitoring to be not too bad. I'd rather we didn't do it, but it's not a fatal flaw. My sense from the papers I've looked at is that even when you do this, sometimes it makes performance worse. And even when it improves performance, if you look at the full probability distribution, it usually has just one main reasoning track that you can follow easily.
所以你是说,即使使用连续思维链,你也可以回过头来问,如果我们强制它在中间阶段输出词元,每一步取最可能的词元,会怎样。可能有人担心它会在概率分布的尾部隐藏第二条思维轨迹。但你说它似乎只有一条思路,没有第二条隐藏的思路。
So you're saying even with continuous chain of thought, you can go back and ask what if we had forced it to output tokens at intermediate stages, taking the most likely token at each step. There might be a concern that it could hide a second track of thought in the tail of the probability distribution. But you're saying it seems to have only one train of thought, not a second hidden one.
是的,我认为你不应该只看排名第一的,也许要看前五,甚至只看前两个。你还可以通过在测试时修改模型,使其只使用前两个词元并消融其余部分,来判断你遗漏了多少。如果模型仍然表现良好,那么你可以确信它没有在分布的其余部分夹带信息。
Yeah, I would say you should look at not just the top one, but maybe the top five, or even just the top two. You can also tell how much you are missing by modifying the model at test time to only use the top two tokens and ablate the rest. If the model still performs as well, then you can be confident it's not smuggling information in the rest of the distribution.
即使不考虑基准性能,你也可以看看只保留前几个词元是否会导致不同的建议或结果。如果从我们的角度来看输出总是相同的,那就强烈表明尾部不包含重要信息或第二套推理。
Even setting aside benchmark performance, you could see whether keeping only the top few tokens leads to a different recommendation or outcome. If the output is always the same from our point of view, that strongly suggests the tail doesn't contain important information or a second set of reasoning.
是的,没错。但我需要提醒一下:如果你对模型进行足够的微调和强化学习,使其将分布仅仅视为一个数字向量,并尽可能使其有用,经过足够的训练,它可能会达到你无法再将其解释为词元分布的程度。然而,我看过的论文表明,这不如将其视为词元分布效果好。
Yeah, that's right. But I should give a caveat: if you do enough fine-tuning and RL on the model to treat the distribution as just a vector of numbers and make it as useful as possible, with enough training it might get to the point where you can no longer interpret it as a distribution over tokens. However, the papers I've seen suggest this works less well than treating it as a distribution over tokens.
这是要求它自己想出一种非人类可读的语言的过程吗?
Is this the process of asking it to come up with its own non-human-readable language?
是的,基本上是这样。事实上,最初的 Coconut 论文提出了这一点,一些后续工作说问题在于它表达能力太强;我们需要将其限制在我们词元的概率分布上,这样它表现更好。这反映了模型在限制使用自然语言时更擅长类人推理。那是它们擅长的部分,并且能带来更好的性能。
Yeah, basically. In fact, the original Coconut paper proposed this, and some follow-up work said the problem is it's too expressive; we need to restrict it to the probability distribution of our tokens, and then it performs better. This reflects that models are much better at human-like reasoning when restricted to natural language. That's the part they're smart in, and it leads to better performance.
你的理论是预训练带来了巨大的优势。模型非常擅长人类语言。如果你让它们自己想出一种内部语言,理论上可能存在一种更好的推理语言,但它们无法带上预训练学到的一切。它们必须从头开始,所以结果远远落后于使用英语或其他人类语言。
And your theory is that pre-training packs an enormous punch. Models are really good at human language. If you ask them to come up with their own internal language, in theory there might be a better language for reasoning, but they can't bring along everything from pre-training. They have to start from scratch, so it comes out substantially behind using English or other human languages.
是的,完全正确。
Yep, that's exactly right.
如果你认为这种不透明的序列深度,或者说它们没有很大的序列深度而我们又无法查看,这对我们监控它们、确保它们基本对齐或不做太有害的事情至关重要,那么这是否可能成为 GDM 的一个治理目标?你可以制定内部政策——我知道目前还没有很大的动机去违反它,但假设未来某个时候,你可能会获得更好的性能,或者某个时候可能会出现交叉点。你仍然可以有一个内部治理标准,规定它们不能思考超过一定量,或者不能连续有这么多想法,之后原则上有人能够审查它,因为超过那个限度实际上会很危险。
If you think that this opaque serial depth or the fact that they don't have a very great serial depth without us being able to look at it, if that is so key to our ability to monitor them, ensure that they're basically aligned or not doing anything too harmful, is that a potential kind of governance target for GDM that you could have some internal policy saying I mean it sounds like there's not huge incentives yet to violate that anyway, but let's say in future you could get better performance at some point or at some point there'll probably be a crossover. You could still have a, you know, an internal governance standard saying, well, they can't think for more than like this amount or they can't have this many thoughts one after another before someone or it would in principle be able to scrutinize it because after that because it actually just would be dangerous to exceed that.
是的,我认为这完全可行。我认为如果能成为更广泛的行业标准会更好。一般来说,我倾向于不要求单个公司单方面行动的事情。如果把它做成行业标准,会更稳定。但没错,我认为这完全可行。事实上,在这个播客发布的时候,我们会发表一篇论文,描述如何为任何给定架构计算这种不透明的序列深度,并提供一些代码,用于至少那些用 Jax 实现的模型——Jax 是人们用来实现模型的一种框架。
Yeah, I think that's definitely doable. I think it would be even better for it to be a broader industry standard. I as a general rule I tend to favor things that don't require individual companies to unilaterally do stuff. It's much more stable if you make it an industry standard. But yes, I think that would totally work. In fact, by the time this podcast comes out, we will have published a paper that does that, like describes how you could calculate this opaque serial depth for any given architecture and has some code for how you could do this for models that are at least the ones that are implemented in Jax, which is just a kind of framework that people use to implement models.
Gemini 3 Pro 不久前发布了。AI 安全博主 Maswitch,几年前上过这个节目,他在博客上对前沿安全报告说了很多相当批评的话,这份报告基本上是与发布同时出来的。他大致上担心 GDM 基本上隐藏了很多他认为如果更突出、更容易阅读就会给 DeepMind 带来公关问题或监管问题的信息。有很多不同的事情,人们可以自己去读博客文章。但有几个让我印象深刻:他感到困扰的是,在说服力评估中,你们只给出了比值比,而没有给出绝对水平。所以你无法确切知道 Gemini 3 或 Gemini 2.5 有多大的说服力。在网络安全评估中,根据他的解读,他认为你们把模型破解测试而不是以自然方式解决测试视为一种绿灯,一个不必担心的理由,而不是红灯,一个更值得担心的理由。在帮助人们获取大规模杀伤性武器方面,他感觉——而且我认为很多人都有这种印象,不仅仅是针对 GDM,而是针对整个公司——当模型似乎开始接近那条线时,可能有点模糊它们是在线之上还是之下,而六个月或一年前的那条线,感觉就像球门柱在移动,标准随着时间的推移而提高,这样模型总是基本上可以发布,这就是他担心 Gemini 3 也在发生的事情。你如何回应这些反对意见?
Gemini 3 Pro came out not that long ago. The AI safety blogger Maswitch, who was on the show a couple of years ago, he had a bunch of fairly critical things to say on his blog about the Frontier Safety report that came out basically I think simultaneously with the launch. He broadly speaking thought that he was worried that GDM was basically hiding a bunch of information that he thought would be inconvenient or create PR problems or regulatory problems for DeepMind if it was more salient and more easy to read. There were a whole lot of different things people can go and read the blog post if they want. But a few that stood out to me was he was troubled that on persuasion evaluations you only gave odds ratios rather than absolute levels. So you couldn't tell exactly how persuasive Gemini 3 or Gemini 2.5 was. On the cybersecurity evals, on his reading, he thought that you treated the model hacking the test rather than solving the test in the natural way as kind of a green light, a reason not to be concerned rather than a red light, a reason to be more worried. And in terms of helping people to acquire WMDs, it felt to him and I think a lot of people have had this impression not just about GDM but about companies in general that when it seems like the models are starting to approach the line maybe it's a bit ambiguous whether they're above or below the line that they had 6 months or a year ago, it feels like the goalposts kind of shift and the standards rise over time so that always the model is basically acceptable to put out and that's what he was worried about was happening here with Gemini 3 as well. How do you respond to these kinds of objections?
谢谢。是的,我对这个问题有很多想法。我想对于大多数这些批评,我的主要回应是,我们关心前沿安全报告,把它作为一种方式来说明我们已经做出了一个决定——我们已经正式确定这个模型是安全的,从安全角度来看我们已经准备好发布它。我认为前沿安全报告在这方面很好。对于说服力评估,我实际上并不太熟悉。它是由一个单独的团队做的。所以我不能说得太多,但我预计答案是,他们认为那是最重要的图表,所以把它放进去,并没有特别考虑外部观察者会如何红队测试它,我预计他们会在某个时候,可能在 2026 年,发表一篇论文,更详细地讨论这个问题。但我确信的是,他们并没有特别试图隐藏任何结果。在网络安全方面,我想我对那个特定的批评有点困惑。因为你提到,这是关于把模型破解测试视为绿灯而不是红灯。我对此有点困惑,因为我们确实把这些算作成功,所以它们确实计入了我们报告的 12 分之 11 的分数。所以这更像是把它当作红灯而不是绿灯。我还要说,这不应该被描述为模型破解测试。我认为我们的描述方式是模型找到了测试的捷径。并不是模型看到这个然后想“啊哈,我可以编辑测试,让它看起来像是通过了”。更像是有一个相当复杂的环境,我们试图测试模型执行某种特定网络任务的能力,但它注意到有一条替代路径,如果我在做这个任务——我大部分情况下会失败,因为我不像 Gemini 那样擅长网络任务——但如果我在做这个任务,我看到了那个捷径,我甚至不会认为那是作弊。我会认为,当然,这就是做这个任务的方式。现在当我们审视这一点,必须判断模型在网络方面的能力有多强。我的意思是,我们最终试图测试它做一件困难的事情,但它却做了一件稍微容易的事情,所以这就产生了一个问题:我们对模型的能力得出什么结论?
Thanks. Yeah, I have so many thoughts on this one. I guess for most of these I would say like my primary response is that we care about the frontier safety report as a way of saying we have made a determination that we have made a formal determination that this model is safe and we're ready to release it from a safety perspective. I think that aspect of the frontier safety report is great. For the persuasion one, I don't actually I'm not that familiar with it. It's done by a separate team. So I can't say too much about it, but I expect that the answer is, you know, they thought that this was the most important graph and so they put that one in and they weren't particularly thinking about how it would be redteamed by external observers and I expect that they will publish a paper at some point probably in 2026 that will go into more detail about this. But what I am confident about is that, you know, they weren't particularly trying to hide any results here. On the cyber side, I think I'm a little confused by that particular criticism. Just because you said that it was something about treating the model hacking the test as a green light instead of a red light. I'm a little confused by this because we did count those as successes and so they did count towards the like 11 out of 12 score that we reported. So that is treating it more like a red light rather than a green light. I would also say that this should not be described as the model hacking the test. I think the way we describe it is the model found a shortcut to the test. It is not the case that the model looked at this and saw aha I can just edit the tests so that it looks like it's passed. It was more like there is this pretty complicated environment where we were trying to test the model's ability to do one particular kind of cyber task and instead it noticed that there was this alternate pathway and if I were doing this task I mean mostly I would fail at it because I'm not as good as Gemini at cyber tasks but if I were doing this task and I saw that shortcut I would not have even thought of it as cheating. I would have thought of it as like yeah of course that's the way in which you do the task. Now when we look at this and we have to make a determination about how good is the model at cyber. I mean ultimately we were trying to test it for one difficult thing and instead it did something slightly easier and so there is this question about like okay what do we conclude about the model's capability.
我想你不知道它是否本可以做更难的事情,因为它可能做了是因为它做不了更难的事情,或者它可能只是因为它更容易而做了。
I guess you don't know whether it could have done the harder thing because it might have done it because it couldn't do the harder thing or it might have just done it because it was easier.
是的,没错。我的意思是,它可能甚至没有考虑更难的事情,因为如果有更容易的事情,它看到了一条前进的道路,就直接走了。但没错,我们不知道它是否本可以做更难的事情。我认为在那个案例中,我们让一些主题专家看了看,他们的结论是,是的,模型可能也能做更难的事情。所以我们把它计入了模型分数。我认为在其他情况下,我们可能不会计入。然后我们可能会把它从分子和分母中都去掉。
Yeah, that's right. I mean, it probably didn't even consider the harder thing because, you know, if the easier thing is there, it sees a path forward, it just takes it. But yes, we don't know whether it could have done the harder thing. I think in that case, you know, we had some subject matter experts look at it and their conclusion was like, yeah, probably the model could have done the harder thing too. And so we counted it towards the model score. I think in other cases we might not count it. And then probably what we would do is remove it from both the numerator and maybe also the denominator.
但在这个具体案例中,我们确实统计了。如果我没记错的话,Z 的主要反对意见是我们有第二个测试没有报告定量结果。但那个测试对我们排除网络 CCL 非常关键。为什么没报告?主要是因为那个测试还在某种程度上处于建设阶段。我认为它足够稳健,可以排除 CCL。但为什么我不想把大量细节放进模型评分卡和前沿安全报告里呢?主要是因为运行这个评估的具体细节很可能会改变,以便让它更稳健,加入更多挑战。然后在下一份前沿安全报告中,我们得写一整节关于我们如何改变这些东西,以及之前的结果不可比。你能理解吧,这看起来是不是太巧了:测试足够好到能判断模型安全,但又不够好到让你在报告里包含具体细节?如果你想要那样的东西,你基本上得至少达到学术论文的细节和严谨程度,而且通常连那都不够。我确实看到发表关于我们评估的论文有很大价值,而且我们在这方面做得比大多数公司都多。我们发表了最早的一篇关于评估的领先论文,叫做《评估前沿模型的危险能力》。
But in this particular case we just did count it. If I remember correctly from the post, I think Z's bigger objection was that we had a second test that we didn't report quantitative results on. But that was pretty key to us ruling out the cyber CCL. Why did that happen? Mostly because this test is still to some extent under construction. I think it was robust enough to rule out the CCL. But why didn't I want to put great details into the model scorecard and into the frontier safety report? Mainly because the details of exactly how we run this eval are likely going to change in order to make it somewhat more robust to include a couple more challenges. Then in the next frontier safety report we have to write an entire section about how we changed the stuff and previous results aren't comparable. You can understand, doesn't it seem awfully convenient that the test is good enough to determine that the model is safe but not good enough for you to include any specific details in the report itself? If you want something like that, you'd basically have to go to at minimum the level of detail and rigor that is found in an academic paper, and often even that's not enough. I do see a lot of value in publishing papers about our evaluations, and we have done this more so than most other companies. We published one of the first leading papers on evaluations, called 'Evaluating Frontier Models for Dangerous Capabilities'.
是的,我想我们大约一年前和 Alan Defoe 讨论过那篇论文。
Yeah, I think we spoke about that with Alan Defoe about a year ago.
啊,很好。是的。然后最近我们还发表了关于隐蔽性和情境感知的评估。所以我认为论文是一个有明显好处的地方。它让人们理解评估实际在做什么,也让其他人可以在这些评估的基础上继续工作。所以我们在这方面投入了更多精力。而对于前沿安全报告,我觉得它的主要目的是正式声明我们做出了这个判断,而不是提供足够细节让人们独立验证我们的工作。
Ah, very nice. Yes. And then I think more recently we published our evaluations for stealth and situational awareness. So I think papers are a place where there are clear benefits. It allows people to understand what the evaluations are actually doing. It allows for other people to build on top of those evaluations. And so we put more effort into that. Whereas for the frontier safety report, I feel like its primary purpose is to say that we made this determination formally, and less about providing enough details that people can independently check our work.
我的意思是,人们不读论文,Rohin。他们读模型卡是因为他们对新模型发布感到兴奋。我想是不是在那种时间线上,写一篇论文或提供论文级别的细节和严谨度到模型卡里不现实?
I mean people don't read papers, Rohin. They read the model card because they're psyched about a new model launch. I guess it's just impractical on that kind of timeline to write the equivalent of a paper or provide the level of detail and rigor that you would in a paper in the model card?
是的,我认为没错。我还想说,模型卡或前沿安全报告的另一个好处是,如果你有想用扩音器向社区传达并附上公司品牌的东西,模型卡或前沿安全报告非常适合。例如,我们有一个关于思维链可读性的章节,因为那正是我想用扩音器传达的。但至于在模型卡或前沿安全报告里写论文,对于新的评估来说,从评估足够好到我们可以用于决策,到它足够稳健、稳定、经过实战检验,并且我们做了大量撰写工作可以发表论文,这之间有相当大的滞后。
Yeah, I think that's right. I should say the other thing that I think a model card or a frontier safety report is good for is if you have something that you want to speak in a megaphone to the community and attach the company's brand to it, a model card or frontier safety report is excellent for that. For example, we have a section on chain of thought legibility because that is something I do want to use the megaphone for. But in terms of writing the paper in the model card or the frontier safety report, for new evaluations, there is a substantial lag between when an evaluation is good enough that we can use it in our decision-making and when it is robust enough, stable enough, battle-tested enough, and we have done a lot of work on the writing up that we can publish a paper about it.
好的。第三个担忧是存在一个普遍现象:随着模型能力越来越强,每一次迭代,令人担忧的门槛、看似不安全的门槛似乎和能力提升大致同步上升。你怎么看?
Okay. And the third concern was that there's this general phenomenon. It seems that as the models get more capable, with each iteration, the bar for what would be troubling, the bar for what would seem to be unsafe rises approximately the same as the capabilities have gone up. What do you make of that?
是的,我认为这尤其发生在 CBRN 的背景下,如果我没记错的话。我想说,发生这种情况的原因更多是随着时间的推移,我们看到 AI 系统更多的能力,我们更清楚需要评估什么以及威胁模型应该是什么。在看到 Gemini 2.5 之后,你可能会说,‘哦,这是模型真正变得擅长的地方。我们需要在这里有更强的评估,并在理解威胁模型上投入更多精力。’所以 Gemini 2.5 和 Gemini 3 在 CBRN 上的差异主要是我们大幅改进了威胁模型和评估,使其更加严格,这也部分解释了为什么关于它们的细节不多,因为它们还没有达到我认为足够好、稳健、稳定、可以大量公开细节且不会发生较大变化的程度。但这种事情,随着时间推移,你获得更多信息,可以改进评估和威胁模型,这很难与移动目标区分开来,这有点不幸,但事实就是如此。
Yeah, I think this was especially in context of the CBRN one if I remember right. I would say that the reason this happens is more that as time goes on and we see more capabilities of AI systems, it becomes clearer what we need to evaluate for and what our threat models should be. After you see Gemini 2.5, you can be like, 'Oh, this is the place where the models are actually getting good. We need to have stronger evaluations over here and put more effort into understanding our threat modeling over here.' So mostly what happened with the difference between Gemini 2.5 and Gemini 3 on CBRN is that we substantially improved our threat modeling and our evaluations so that they are more rigorous, which is also partly why there is not as much detail about them because they are not quite at the level where I think they are nice, robust, stable, that we can really put lots of details out in a way where we don't expect them to change a decent bit. But this sort of thing, as time goes on, you get more information and you can make your evaluations and threat modeling better, is not that easily distinguishable from the goalposts changing, which is a bit unfortunate, but turns out to be the way things go.
是的。我的意思是,这说明了在你心中这些模型卡的目的,与像我或一般评论者心中它们所服务的目的非常不同。我认为很多人希望它们成为一种问责机制,一种如果模型危险或变得更危险时可以拉响警报的机制,这样 GDM 就不得不在模型卡中披露这一点,因为他们必须公布这些结果,即使他们想隐藏 CBRN 的危险性,他们也做不到。
Yeah. I mean, that speaks to the fact that the purpose of these model cards in your mind is very different than the purpose that they serve in the mind of someone like me or commentators in general. I think many people want them to be an accountability mechanism, a mechanism by which the alarm could be sounded if the models were dangerous or were becoming more dangerous, where GDM would have to reveal that in the model card because they have to put out these results, and even if they wanted to hide that the CBRN stuff was dangerous, they wouldn't be able to.
我的意思是,就像我说的,这贯穿了整个对话。我会谈到第三方审计员,他们可以实际看到示例以及如何评分等等。我认为这对于判断评估是否真正支持我们的判断,比具体的定量分数重要得多。如果我看到一个数字,比如夺旗挑战中 12 题对了 10 题,那意味着什么?我反复听到一种来回争论:你或公司里的人会说这些报告的目的是表明我们已经深思熟虑过,我们做出了模型可以安全部署给公众的判断。而其他人会说,当然,当前的模型也许没问题。
I mean, like I said, this has been a theme throughout this conversation. I would talk about third party auditors who can actually see the examples and how they were graded and stuff like this. And I think that is way more important for judging whether or not the evaluation actually supports the judgment that we make than the specific quantitative scores that we get. If I see a number like 10 out of 12 on capture the flag challenges, what does that mean? There is a back and forth that I hear repeatedly: you or someone at a company will say the purpose of these reports is to show that we have thought this through. We've made a determination that the model is safe to deploy to the public. And other people will say, sure, the current model maybe it is fine.
我不认为他们真的觉得现在商用太危险,但总有一天这些模型会危险到我们需要担忧,而我们预测公司届时行为方式的唯一依据就是他们现在的行为。我们用现在发布的模型卡来衡量公司对安全的重视程度,对透明度和披露真实情况的重视程度。而你要在将来可信地承诺成为好的行动者,方法就是现在做得更好。你怎么看?我猜你可能会说和之前类似的话,就是这不是完成任务的正确机制。
I don't think they actually think it is too dangerous for people to be using commercially, but one day these models will be dangerous enough that we should be concerned, and the only way we can forecast how the company will behave come that time is how they're behaving now. We use the model cards that are put out now as a measure of how serious the company is about safety, how serious it is about transparency and revealing what is actually going on. And the way you could kind of credibly commit to be a good actor later on is to be a better actor now. What do you make of that? I guess you're probably going to say similar things to what you've said before, that this just isn't the right mechanism for the task.
是的,没错。我的意思是,更广泛地说,我觉得社区提出的要求可以分为两类。一类是真正对实际安全有影响的事情。我完全支持这些。我们应该做这些。如果我们没做,人们应该评判我们。另一类是我归类为‘支付成本以向额外安全社区表明忠诚’的事情。而我不认同这些。我不想做。如果你要求这些,我会直接拒绝。
Yeah, that's right. I mean, maybe more broadly, I might say that it feels like asks that the community makes can fall in two categories. One is things that actually matter for actual safety. I'm all for those. We should do those. People should judge us if we're not doing them. And then there's things I would categorize as pay costs to signal allegiance to the extra safety community. And I'm like, I'm not about those. I don't want to do them. If you ask for them, I am just going to say no.
我觉得‘忠诚’这个词有点不公平,但这更像是愿意为此目的投入资源。
I mean I think allegiance is a little bit of an unfair way, but it's like willingness to commit resources to this purpose.
但并非为了实际安全的目的,而是为了表明你将来会做安全工作的目的。
But like not to the purpose of actual safety, to the purpose of signaling that you will do safety in the future.
我的意思是,这种机制并不罕见——你想承诺将来做某事,而人们如何评估你未来的行为?他们唯一能得到的指示就是现在的事情,因为他们没有水晶球。没错。但我认为你应该基于那些真正对安全重要的事情来做,而这样的事情有很多,对吧?比如,我认为人们进行评估很重要,特别是针对网络和 CBRN 滥用,还有像欺骗性对齐、未对齐等问题。我认为他们向政府和适当的外部机构报告这些信息很重要。所以我认为那里有你可以关注的事情。例如,我认为人们为将来需要做的事情做规划很重要。我认为公司研究未来可能出现的 AGI 安全问题并找出应对方法很重要。所以有很多事情你现在就可以评判公司,我更希望人们根据那些真正重要的事情来评判我们。
I mean this isn't an unusual mechanism that you want to kind of commit to doing something in future and how do people assess how you'll behave in future is like the only indication they can get is things in the present because they don't have a crystal ball. That's right. But I think you should do it based on like the things that actually matter for safety of which there are many, right? Like I think it matters that people are running evaluations for especially cyber and CBRN misuse but also other things like deceptive alignment, misalignment. I think it matters that they are reporting this information to governments and appropriate external bodies. So I think there are things that you can look at there. I think it matters for example that people are doing the planning for what they will need to do in the future to address future issues. I think it matters that companies are doing research into AGI safety concerns that might arise in the future and figuring out ways to deal with that. So there's lots of stuff that I think you can judge companies on right now and I would much rather that people judge us based on that stuff that actually matters.
所以基本信息是,将资源从那些实质上有益、最终能解决问题的事情上转移开,投入到那些只是走过场、显得你关心的东西上,这是不好的。而且有时人们无意中,我猜,会要求你把资源投入到第二种事情上,这些资源部分来自组织其他部门,但很大一部分会来自安全和对齐团队及其资源和人员。
So the basic message is it's bad for resources to be diverted from stuff that is substantively good, that is actually going to solve the problem ultimately, towards things that are kind of going through the motions of appearing like you care. And sometimes people accidentally, I guess, are kind of requesting that you put resources into the second one, which I guess partly will come from the rest of the organization but will in significant part come from the safety and alignment team and its resources and its staff.
我的意思是,我猜人们会说他们可能更喜欢第二种,因为它更容易评分或更容易看到,而且他们可能不像你那样处于有利位置,无法理解任何一家公司从长远来看是否实质性地做了正确的事情。但我猜你是说,这就是为什么你想要像 AI Labatch 或其他专家那样的人,他们能关注重点,花时间真正思考这个问题,并实际为其他人评分,以便他们能够理解。
And I mean I guess people would say they might prefer the second one because it's easier to grade or it's easier to see and perhaps they don't feel in as good a position as you do to understand whether substantively any company is doing the right thing on the merits of what will matter in the long term. But I guess you're saying that's why you want to have experts like AI Labatch or whoever else who can have their eye on the ball who can spend their time really thinking that through and actually grade it for everyone else so they can understand.
除了所有这些我同意的观点之外,我还对基于那些实际上并不重要、我们不会真正认为对实际安全有影响的事情进行评分深表怀疑。这感觉像是会导致整个文化关注成本或投入,而不是真正重要的实际结果。并且会导致激励人们显得好、看起来好,确保你的公关说,我们花了一千小时来弄清楚如何处理这个模型。而实际上那一千小时可能是,我们在一个输入上运行了模型,看了看,然后忽略了结果,因为它不重要,我们雇了一些承包商来做这个,然后无视了他们说的。但现在我们可以说我们花了一千小时来研究它。我不想要这些激励。它们似乎很糟糕。这对透明度和坦诚似乎很糟糕。谷歌或其他任何人越是这样做,我就越不能说出我认为真正对安全重要的事情,我就越需要做这种复杂的舞蹈,说一些会安抚那些希望我们展示投入资源承诺的人的话。而我就越不能这样说:这是我们正在做的对安全真正重要的事情,以及为什么我认为它很好。整体上这是一个糟糕的激励环境。而且我认为对我来说非常重要的一点是,我们不要陷入看起来好而不是真正好并证明为什么这样的陷阱。
I think I'm also in addition to all of that, which I agree with, I'm also just deeply skeptical of grading that is based on stuff that doesn't actually matter, something that we wouldn't actually stand behind as mattering for actual safety. This feels like the sort of thing that causes the culture overall to be looking at costs or inputs rather than actual outcomes that matter. And results in incentives to appear good, to look good, to make sure your comms say, you know, we spent a thousand hours figuring out what to do with this model. When maybe those thousand hours were like, we ran the model on an input and we looked at it and then we ignored the result because it didn't matter and we hired some contractors to do this and just ignored what they said. But now we can say that we spent a thousand hours looking at it. I'm like, I don't want these incentives. They seem bad. It seems bad for transparency and candidness. The more that Google or anyone else is doing this, the less I'm able to say what it is that I think actually matters for safety, the more I have to do this complicated dance of saying the things that will appease the people who want us to be showing commitment to putting in resources into a... And the less I can be like, here's the stuff we're doing that actually matters for safety and here's why I think it's good. It's just a poor incentive landscape overall. And I think it's really quite important to me that we don't fall into the trap of looking good rather than just actually being good and then justifying why that's the case.
那么,对于模型卡中详细信息的另一个完全不同的理由呢?那就是世界其他地方需要知道,比如湾区的其他研究人员,而不是英国,他们能从理解 GDM 所做的所有这些事情以及模型的样子中获得实际的研究收益,这将帮助其他公司更好地制作自己的模型卡和进行评估。你怎么看?
What about an entirely different justification for having thorough detail in the model cards which is the rest of the world needs to know for practical you know other researchers over in the Bay Area rather than in the UK you get actual research benefit out of understanding all these things that GDM is doing and what the model looks like that's going to help other companies do a better job with their own model cards in their own evals. What do you make of that?
是的,我认为这有一定道理。当然,为了在研究基础上继续发展,有细节是好的,我认为这就是我们发表论文的原因。我们是否需要在模型卡或前沿安全报告中这样做?比如,我专门与政策和治理人员讨论过这个问题,我问他们你们实际上用我们的模型卡和前沿安全报告做什么,他们有时会说,哦,它很好地提供了更多关于你们如何评估某某风险的细节,然后我会说,太好了,我们还有这篇论文,其中更详细地介绍了评估。我想可能它甚至在模型卡或前沿安全报告中被引用了。
Yeah, I think there is some substantial truth to this. Certainly for the purpose of building on the research it is good to have details and I think that's why we publish papers. Do we need to do this in the model card or the frontier safety report? Like I've gone around talking to policy and governance people especially about this and I ask them what do you actually use our model card and frontier safety report for and they will sometimes say things like oh it was really good to get more details about how exactly you evaluate for such and such risks and then I will say something like great and you know we also have this paper that goes into more details of the evaluation. I think maybe probably it was even cited in the model card in the frontier safety report.
你读过吗?然后他们通常甚至没听说过这篇论文,所以我有点不相信他们真的在意这些细节。我确实认为发表论文很重要,但不必和模型卡片绑定。我聊过的一些人,尤其是偏技术、实际构建这些评估的人,确实觉得论文有用,能帮助他们在此基础上继续工作,所以我希望继续尽可能发表论文。你之前告诉我,你觉得 GDM 的研究论文总体有点低调。我想部分原因是 GDM 在伦敦,而不是湾区——那里大概是这些问题上社交和能量最集中的地方。是的,再多说说这个。
Did you read it? And then usually they will not even have heard about this paper, so I kind of don't believe them when they say that these details actually matter to them. Again, I do think publishing papers is important, it doesn't need to be tied to the model cards. Some people I've talked to who are especially more on the technical side, the ones actually building these evaluations, do find the papers useful and helpful for them to build on, so I do want to continue publishing the papers where we can. You told me that you think research papers that come out of GDM tend to fly under the radar a bit in general. I suppose in part because GDM is based here in London rather than in the Bay Area, where I guess the greatest amount of socializing and energy around these issues exists. Yeah. Tell me more about that.
是的。实际上现在 GDM 整体在伦敦和湾区分布相当均匀,但安全团队历来在伦敦。我们正在向湾区扩展,但我仍然认为关注的中心在伦敦。而且实际上,很多研究论文和想法是通过口口相传在社区传播的,我们在这方面参与得少一些,这有点遗憾。
Yeah. So I think actually now GDM overall is pretty evenly split between London and the Bay Area, but the safety team has historically been based in London. We're expanding into the Bay Area, but I would still say that the locus of attention is in London. And I think in practice, a lot of research papers and ideas end up spreading in the community via word of mouth, which we are a little bit less plugged into, which is a bit unfortunate.
好的。我们来聊聊其中一篇论文。据我所知,有一个有趣的研究结果绝对被忽视了。第一篇是《用非近视性审批的近视优化可以缓解多步奖励黑客》。我真搞不懂这怎么没火起来。你们在这项工作中取得了什么成果?
Okay. Yeah. Let's talk about one of the papers. One of the interesting research results that definitely flew under the radar as far as I could tell. First one is 'Myopic optimization with non-myopic approval can mitigate multi-step reward hacking.' I don't know how on earth that didn't go viral. What did you accomplish with that work?
是的。这项工作的核心卖点是,你可以训练一个模型,让它不会进行某种特定的奖励黑客行为,即使监督者无法察觉这种奖励黑客行为已经发生。与放大监督或辩论等方法不同,那些方法依赖于让监督者足够聪明或足够有能力来检测并惩罚奖励黑客。而这种方法完全不依赖检测,所以实际上更稳健。这怎么做到的呢?我们施展了什么魔法?这基本上是基于近视优化的思想。思路是:假设有一种奖励黑客行为需要你在时间上分开执行两个不同的动作。例如,假设你的任务是写代码,你先写一些特别容易通过的测试——它们实际上没有测试所有必要的内容——然后你写代码通过这些测试,因为你不知道如何编写能通过所有测试的正确代码。这两个步骤必须同时存在,奖励黑客才能成功。如果你只有其中一个而没有另一个,你就会受到惩罚。所以思路是:如果你把这两个步骤分开处理——一个步骤是写测试,另一个步骤是写代码——并且你只根据一个信号来优化测试,这个信号表示这些测试看起来有多好;然后你单独优化写代码的部分,仅根据之前发生的情况来判断这段代码看起来有多好——那么 AI 系统就没有动机去写糟糕的测试。通常,动机是这样产生的:当你进行强化学习时,代码部分的奖励也会通过时间反向传播,影响测试编写部分。所以 AI 可以从反向传播的奖励中学习到:如果它写容易通过的测试,未来往往会获得更多奖励。而近视优化的部分是说,我们不让它看到未来发生了什么。我们只根据它当前的行为来评估它。所以那个动机永远不会出现。这就是为什么即使你无法检测到奖励黑客,它也不会被激励。这部分讲清楚了吗?
Yeah. So I think the pitch for this work is that you can train a model that doesn't do a particular kind of reward hacking even if the overseer would not be able to notice that reward hack if it had happened. In contrast to things like amplified oversight or debate, those rely on making the overseer smart enough or capable enough to detect the reward hack and penalize it. This one doesn't rely on detection at all, so it's actually a bit more robust. How does this work? What is this magic that we managed to do? This is basically doable based on the idea of myopic optimization. The idea is: suppose there is a reward hack that requires you to take two different actions separated in time. For example, maybe if you are tasked with writing some code, you first write some tests that are particularly easy to pass—they don't actually test everything that is needed—and then you write code that passes the tests because you didn't know how to write the correct code that would have passed everything. Both of these steps have to be present for the reward hack to actually succeed. If you have one but not the other, you just get penalized. So the idea is: if you treat these as two separate steps—one step in which the tests are written and a different step in which the code is written—and you optimize the tests only based on a signal that says how good these tests look, and then separately you optimize the code writing part based solely on how good this code looks based on what has happened previously, then there's actually no incentive for the AI system to write bad tests. Normally, the incentive would come because when you're doing reinforcement learning, the reward on the code part would also be backpropagated through time to also affect the test writing part. So the AI could learn from that backpropagated reward that if it writes easy-to-pass tests, it tends to get more reward in the future. The myopic optimization part is saying we're not letting it see what happened in the future. We're only evaluating it based on what it did right now. So that incentive never comes. This is why the reward hack isn't incentivized even if you can't detect it. Does that part make sense?
是的。我觉得这个宽泛的想法我很久以前就听说过,那就是如果你担心模型可能会伪造成就或作弊来实现最终目标,因为它只得到一个是否完成目标的奖励信号。所以想象一下,你让模型经营一家公司,要求它赚钱,但它不是经营成功的企业,而是去偷一大笔钱,因为它发现这实际上是增加银行余额的更有效方式。这是你可能有的担忧。防止这种情况的一种方法是,不是根据最终结果来强化和评估模型,而是看它说将要采取的行动或实际采取的行动,并根据这些行动在你或其他监督者看来是否合理、是否与你实际目标一致来评分,这样就不会出现这种奖励黑客行为。这基本上就是那个宽泛想法的一个实例吗?
Yeah. So I feel like I've heard this broad idea going back a very long time, which is that if you're worried that a model might fake accomplishment or cheat at accomplishing the ultimate goal because all it gets is a reward signal of whether it accomplished it or not. So you imagine you've got the model running a business and you ask it to make money, and rather than running a successful business, it goes and steals a bunch of money because it figures out that that's actually a more effective way of increasing its bank balance. That's a concern you might have. One way to prevent that is rather than reinforcing and evaluating the model based on the final outcome, you look at the actions it says it's going to take or the actions it did take and grade them on whether they seemed reasonable and sensible to you or to some other monitor in light of the goal you actually had, and then you wouldn't get this reward hacking behavior. Is this basically an instance of that broad idea?
是的,没错。这在历史上被称为过程监督。我认为现在过程监督这个词也有了其他一些含义,所以我们更清楚地说明了实际技术机制是近视优化但带有非近视性审批。但没错,这是一个非常古老的想法。我认为我们的主要贡献是展示了这在实际的 LLM 中确实有效,并给出了例子,因为据我所知,之前没有实际的实验证明这一点。
Yeah, that's right. So it's historically been called process supervision. I think nowadays the term process supervision has gotten a bunch of other meanings as well, which is why we have this a little bit more clear about what the actual technical mechanism is of myopic optimization but with non-myopic approval. But yes, it's a very old idea. I think mostly our contribution was showing that this actually works with existing LLMs and giving examples of that happening, because as far as I know there weren't actual experiments demonstrating this before.
我想一个明显的问题是,你确实能成功阻止模型做那些看似能达成目标(从而获得强化)但实际上并未实现你心中目标的欺骗性行为。但另一方面,你也会阻止它做一些本可以达成目标的事情,这些事实际上是绝妙的洞察,你永远想不到。所以如果你训练围棋模型 AlphaGo,如果人类评估每一步棋的好坏,那么模型实际上无法超越人类表现,因为人类会把那些实际上出人意料精妙的棋步在早期评为坏棋。你怎么解决这个问题?
I guess an obvious problem with that is you will successfully stop the model from doing scammy things that seem to accomplish the goal that it would get reinforced for but actually don't accomplish the goal you had in mind. But on the other hand, you would also block it off from doing stuff that would have accomplished the goal that was actually a brilliant insight that you never would have had. So if you were training the Go model, AlphaGo, if humans were evaluating whether the moves were good, then the model actually couldn't end up exceeding human performance because they would grade moves that were actually unexpectedly brilliant as bad at the early stage. How do you get around that?
是的,没错。我的意思是,在某种程度上,答案是你无法绕过它。最终,那些令人惊叹的绝妙创意和那些同样是绝妙创意的奖励黑客行为,在你作为观察者看来基本上是一样的。
Yep, that's right. I mean, to some extent, the answer is you don't get around it. Ultimately, both the incredible amazing creative insights that are awesome and the incredible creative insights that are reward hacks look basically the same to you as the observer.
如果你想阻止奖励黑客行为,是的,你也要放弃一定程度的竞争力。话虽如此,我确实认为,我们称之为‘近视优化与非近视审批’。非近视审批部分基本上就是,当你评估某个具体步骤时,让某种智能监督者——可能是人类,也可能是大语言模型——来判断这个步骤对未来获得奖励有多大好处。监督者应该利用它所知道的一切来评估这对未来有多好。这保证的是,未来影响 AI 系统在这个步骤中的任何激励,都必须是监督者能够理解的。
And if you want to stop the reward hacks, yeah, you also give up some amount of the competitiveness. That being said, I do think that, you know, we call it myopic optimization with non-myopic approval. So the non-myopic approval part is basically the part where you say like when you're grading this particular step ask some sort of like intelligent overseer, a human possibly an LLM, to judge whether this particular step, how good will it be for getting future reward. And you know, it should take it, the overseer should then take into account like everything that it knows in judging how good this is going to be for the future. And what this guarantees is like whatever incentives from the future affect the AI system in this particular step have to be things that the overseer understands.
所以你是说,监督者可以看到假设的 AI 经营企业赚了很多钱,但他们评估的是它经历的过程。他们可以包含这些信息,但主要是结合过程来看。
So you're saying the overseer could see that the hypothetical AI running the business made a lot of money, but they evaluate the process that it went through. They can include that information but they primarily do it in light of also looking at the process as well.
嗯,我建议他们不要看公司是否赚了一大笔钱,而是看 AI 系统采取的行动,并自己预测:这会导致一大笔钱吗?是或否。然后据此提供奖励。现在,AI 系统在第一步可以做很多事情来帮助监督者。它可以解释其计划将如何在未来带来大量金钱,然后监督者只需验证。你可以使用辩论或其他 AI 辅助手段,让监督者更好地预测未来会发生什么,从而给出更好的奖励。在极限情况下,如果你能不断改进监督者,使其足够聪明,那么你就能恢复通过纯强化学习反向传播时间所获得的性能,但没有奖励黑客行为。在实践中,我们可能达不到那么远,但我认为你可以推得很远,尤其是通过让 AI 系统解释其计划为何好这种最简单的方法。我认为这是一个非常简单的基线,也能让这种方法相当有效。
Well, I would suggest that they don't look at whether the company made a bunch of money, and instead they look at the action that the AI system takes and they predict to themselves: will this lead to a bunch of money? Yes or no. And based on that provide a reward. Now the AI system on step one can do lots of things to help the overseer with this. It can give an explanation of how its plan is going to lead to tons of money in the future, and then the overseer just has to verify it. You can use things like debate or other AI assistance to make the overseer better at predicting what's going to happen in the future, so that they can give better rewards. In the limit, if you can keep improving your overseer, if the overseer becomes sufficiently smart, then you can recover the performance of what you would get with just straight RL backpropagating through time, but without the reward hacks. In practice, we're probably not going to get that far, but I think you could actually push this quite far, especially by the most simple thing of getting the AI system to explain why its plan is good. I think that's a very simple baseline that can also make this quite effective.
这种方法很快会对前沿模型有用吗?你能看到公司实际使用它吗?
Is this approach going to be useful for frontier models anytime soon? Could you see companies actually using it?
我认为有可能。它只在开始对多步轨迹进行强化学习并持续相当长一段时间时才重要。我认为这今年才真正开始,不同公司程度不同。目前还不完全清楚奖励黑客行为有多大问题。但是的,如果多步奖励黑客行为确实成为一个大问题,我认为现在就应该使用这种方法。事实上,我认为在当前能力水平下,我的猜测是,非近视审批部分,不像在 AlphaZero 中那样会损害竞争力,反而会整体提升能力,尽管代价是需要更多人类输入来提供这些奖励。
I think plausibly. It only matters once you start doing reinforcement learning over multi-step trajectories over a reasonably long period of time. That's something that I think has only really started this year, and to varying amounts at different companies. It's not totally clear how much reward hacking is a big problem. But yeah, to the extent that multi-step reward hacking does become a big problem, I think it's quite plausible that this should be used now. And in fact, I think at current capability levels, my guess would be that the non-myopic approval part, rather than being a competitiveness hit as it would be in something like AlphaZero, I would guess that it would actually improve capabilities overall, although at the cost of needing a lot more human input to provide those rewards.
它会提升性能,因为你不会遇到奖励黑客行为。不仅仅是因为这个。我的意思是这肯定是一个原因,但我认为强化学习还有一个信用分配问题:你做了大量不同的动作,最后得到一个奖励,RL 算法的工作就是根据那个奖励找出哪些动作与奖励最相关。这通常是一个难题。
It would improve performance because you wouldn't get the reward hacking. Not just because of that. I mean that is definitely one reason, but I think also reinforcement learning has a credit assignment problem where you do a ton of different actions and then at the end you get one reward, and it's the RL algorithm's job to figure out based on that reward which of these actions were actually most relevant to that reward. Generally a difficult problem.
是的,这是一个难题。从某种意义上说,RL 算法做的事情是:‘如果奖励是正的或异常好,我们就让所有事情更可能发生;如果异常差,就让所有事情更不可能发生。’而像 MOMA 这样的方法,你可以更精细。你可以说:‘你写这些测试的这部分,那些测试非常好,奖励特别高。但写代码的这部分,嗯,你应该用这个库,你没有,所以那部分奖励会低一些。’这可以让你获得比原来更有效、更样本高效的学习。
Yeah, it's a difficult problem. And in some sense, the thing that the RL algorithm does is like, 'We're just going to make everything more likely to happen if the reward was positive or unusually good, and everything less likely to happen if it was unusually bad.' Whereas with something like MOMA, you can be much more granular. You can say, 'This particular part where you wrote these tests, those were some really great tests, particularly high reward there. But then this part where you wrote the code, eh, you should have been using this particular library, you didn't, so that will give you somewhat lower reward.' And that can allow for more effective and sample-efficient learning than you would otherwise get.
好的。所以这是从这种 RL 方法中获得效率提升,可能使其在原始性能上与其他不那么近视的方法保持竞争力。
Okay. So that's an increase in efficiency that you get from this approach to RL that might allow it to remain competitive in terms of its raw performance with other, I guess, less myopic approaches.
理论上。我不认为我们的论文真正涉及这类问题。这更多是我在推测实践中会发生什么。
In theory. I don't think our paper really gets into questions like this. This is more me speculating about what would happen in practice.
好的。来自 GDM 的第二篇论文没有引起太多关注,标题是‘一种技术性 AGI 安全与安保的方法’。它由大约 30 名 GDM 员工撰写,有 30 个署名。我想它描述的是,作为一份立场文件,GDM 在开发 AGI 时大致打算做什么。考虑到 DeepMind 很可能是最有可能做到这一点的组织,人们竟然对这份关于他们打算做什么、不做什么以及为什么的极其详尽的描述不感兴趣,这有点令人惊讶。它相当长,但开头有一个不错的 10 分钟总结,如果人们感兴趣,可以用它来了解概况。所以如果你想提前了解 GDM 开发 AGI 的方法,你可以花那 10 分钟去做。
Okay. The second paper from GDM that didn't get a ton of attention is called 'An approach to technical AGI safety and security'. It was written by about 30 GDM staff members, or there's 30 bylines on it. I guess it describes, as far as it's like a position paper, of broadly what does GDM think it's going to do as it develops AGI. Which makes it, given that DeepMind plausibly is the organization that is most likely to do this, it makes it a bit surprising that people are not interested in this incredibly thorough description of what you think you are and aren't going to do and why. It is quite long, but it does have a nice 10-minute long summary at the start that you could use to get an overview if people were interested. So if you want to be ahead of the curve on understanding GDM's approach to developing AGI, then you could spend that 10 minutes doing that.
我想说你要做很多不同的事情很容易,但制定这样的计划最困难的决定是弄清楚哪些事情可能有人希望你做,但你承诺不做,因为你认为优先级不够高。这个计划建议你们不优先考虑哪些事情?
I guess it's easy to say that you are going to do a whole lot of different things, but I guess the most difficult decisions in developing a plan like this is figuring out what stuff might some people like you to do that you are committing to not do because you just don't think it's a high enough priority. What sort of stuff does this plan suggest that you are not going to prioritize?
是的,我认为这里最大的类别实际上更多来自我们的背景假设以及它如何影响我们的规划,而不是实际的技术方法。比如一个背景信念,我想我们已经稍微讨论过,就是我们称之为‘近似连续性假设’。
Yeah, I think probably the biggest category in here comes more from our background assumptions actually and how that informs our planning rather than the actual technical approaches. So like one background belief which I think we've talked about a little bit already is we call it the approximate continuity assumption.
这基本上是说,AI 的进步在算力和劳动力等投入方面会相对平稳和渐进,但不一定在时间上如此,因为可能出现智能爆炸。因此,我们的元策略包括大致预测——不一定正式,但在头脑中——对未来一段时间内可能成为潜在问题的事情有所感知,比如三个月,可能更长或更短。我们根据预期拥有的能力,识别哪些考虑因素在那段时间可能变得非常重要,并确保我们为此做好准备。如果我们没有准备好,那么可能会放慢或暂停开发,或者与政府沟通、进行倡导。但重要的是,我们并不试图任意远地预测 AI 发展中将出现的所有问题。目标不是‘知道如何对齐超级智能,否则什么都不做’。我认为我们的时间跨度取决于具体工作,但通常在三个月到五年之间。原因是,实际上不可能知道未来 AI 发展的一切。认为我们提前解决了超级智能可能出现的所有问题,甚至有了解决方案的计划,那是极其傲慢的——我们甚至还没有识别出所有问题。我想举的一个例子是更晦涩的决策理论或人择原理问题,有效利他主义者或长期主义者喜欢讨论关于多重宇宙的考虑如何影响我们今天的行为,并使用 SSA 和 CIA 等假设来思考人择原理,以及它们如何影响我们应该采取的行动。
This is basically saying that AI progress is going to be relatively smooth and gradual with respect to inputs like compute and labor, not necessarily with respect to time due to the possibility of an intelligence explosion. As a result, our meta strategy involves roughly forecasting, maybe not formally, but in our heads, having some sense of what things are going to be potential problems over the next time period, call it 3 months, maybe longer, maybe shorter. We identify what considerations might become quite important during that time given the capabilities we expect to have, and make sure we're prepared for those. If we're not prepared, then possibly slowing down or pausing development, or talking to governments, doing advocacy. But importantly, we're not trying to forecast arbitrarily far into the future all the problems that will arise with AI development. The goal isn't 'know how to align ASI or do nothing.' I think of us as looking at a time horizon that depends on the particular thing we're doing, but often somewhere between 3 months and 5 years. The reason is that it's not actually possible to know everything that's going to happen with future AI development. It would be sheer hubris to think we had figured out every problem that could possibly arise with superintelligence ahead of time and say we've solved it, or even have a plan for solving it—we haven't even identified all the problems. One example I like to give is a much more esoteric decision theory or anthropics type issue, where EAs or long-termists like to talk about how considerations about the multiverse should affect what we do today, and think about anthropics using things like SSA and CIA assumptions and how they affect what actions we should take.
你不必理解那是什么意思,也能跟上你接下来要说的要点。
You don't have to understand what that means in order to track with the point you're about to make.
是的。这些问题的关键结论是,你的信念和决定实际上取决于你是谁。这是人择原理中一个疯狂的地方。通常,不同的人在有足够证据的情况下应该对信念达成一致,或者至少完美的贝叶斯主义者面对相同证据应该达成一致。他们应该有相同的信念。但在人择原理的情况下,这一点不再成立。事实上,即使一个 AI 完全对齐,它也可能与人类有不同的信念,仅仅因为在有人择原理的世界中,贝叶斯更新的结构对 AI 系统和人类是不同的。所以这会影响我们是否愿意将这类研究委托给 AI 系统,即使它们完全对齐。这是一个疯狂的问题,今天肯定不是问题,但在超级智能之前的某个时刻真的可能出现。我不想说‘哦,是的,我已经识别了所有这类问题,并且今天就有解决所有问题的方法’。显然,我们应该顺其自然地处理它们。
Yeah. The key upshot of these things is that the beliefs and decisions you come to actually depend on who you are. This is one of the crazy things about anthropics. Normally, different people should agree on beliefs given enough evidence, or at least perfect Bayesians should agree given the same evidence. They should have the same beliefs. This stops being true in the case of anthropics. Indeed, an AI will probably have different beliefs than a human would, even if it's perfectly aligned, just because the structure of how Bayesian updates happen in a world with anthropics is different for AI systems than for humans. So this should affect whether we are willing to delegate research on these kinds of topics to AI systems or not, even if they're perfectly aligned. That's a kind of wild crazy problem that definitely is not a problem today but really might come up at some point before a superintelligence. I don't want to have to say 'oh yes, I've identified all problems like this and have an approach to solving all of them today.' Obviously, we should just take these as they come up.
好的。所以有些人真的希望你做很多提前规划。他们希望你感觉自己已经掌握了直到人工超级智能的所有重要考虑因素,直到你觉得可以安全地交出一切。而你说我们不会那样做。完全不会。我们会大量思考我们正在训练的下一个模型。我们会稍微思考一下之后的模型,再往后,我们希望未来会自行解决,或者未来的我们会处理出现的问题。这就是你要表达的观点。
Okay. So some people would really like you to plan ahead a lot. They would like you to feel like you have a grasp on all the important considerations that are going to come up all the way through to artificial superintelligence, until the point where you could feel like you've safely handed everything over. And you're saying we are not going to do that. Not at all. We're going to think a lot about the next model that we're training. We're going to think a bit about the model after that and kind of beyond that we're going to hope that in the future will take care of itself or we in the future will take care of things as they come up. That's the point you're making.
是的。不过我认为我们的思考比那听起来更长远。正如我所说,我经常思考未来五年。那是很多很多模型,对吧?我说过思维链监控,我的中位数估计是四年,那会持续——我们已经在思考在思维链监控不再有用之后该做什么。那时我们会想要做更广泛的控制缓解措施。所以我们的时间跨度当然不是非常短。我只是说它没有一直延伸到超级智能。所以我们确实做了相当多的长期规划。也许你会称之为中期规划。
Yes. Though I think we're longer-term thinking than that made it sound like. As I said, often I'm thinking 5 years into the future. That's many, many models, right? I said chain-of-thought monitoring, my median I think I said four years, that lasts—we're already thinking about what to do after the point that chain-of-thought monitoring stops being useful. We would want to do broader control mitigations at that point. So certainly our time horizon is not incredibly short. I'm just saying it's not all the way out to superintelligence. So we are doing a fair bit of long-term planning. Perhaps you would call it medium-term planning.
是的,这篇论文作为公司立场文件相当坦诚。它非常明确地说 AGI 可能在 2030 年到来。我想我们不知道具体时间,但 2030 年完全有可能。我们必须准备好应对的计划。这非常危险,或者可能非常危险。我们必须有一个可以极快部署的计划,以使其更安全。然后它介绍了你打算做的事情。我会鼓励在这个领域工作的人去看看。里面有没有什么不明显的缓解措施?我想有一件事我很高兴你说了,但不是每个人都会强调:你需要将公司内部仅内部部署的 AGI 或 AI 视为不可信的内部人员。如果你不确定或不相信它的对齐性,那么你需要将其视为潜在的敌对智能体,试图破坏内部系统。你想谈谈你承诺要做的事情中的任何一点吗?
Yeah, this paper is pretty candid for a company position paper. It very clearly says AGI might be here by 2030. I guess we don't know when, but 2030 is totally plausible. We have to be ready with a plan for how to handle that. That's very dangerous or potentially could be very dangerous. We have to have a plan that we can deploy incredibly quickly to try to make that safer. And then it goes through the kinds of things that you're intending to do. I would encourage people who work in this area to go and take a look at it. Are there any non-obvious mitigations in there? I guess one thing that I was glad that you said that not everyone would highlight is you need to treat an AGI or an AI inside the company that's just only being deployed internally. If you're not sure or not confident of its alignment, then you need to treat it as an untrusted insider, as potentially a hostile agent that is going to try to undermine internal systems. Do you want to say anything about highlighting any of the things you are committing to doing?
是的。嗯,我可能对‘承诺’这个词有点谨慎。是计划去做。但是的,我认为不可信内部人员的内容确实特别有趣。实际上,我认为我们在论文中没有详细讨论这一点,这有点遗憾。主要是我们当时的思考没有现在这么成熟。但基本上,如果你真的思考一个 AI 系统要接管需要做什么机制,那是相当困难的。其中一点是,它可能希望确保在开始思考要做什么时不被监控。
Yeah. Well, I maybe want to be a little wary about the word 'committing.' Planning to do. But yes, I think the untrusted insider stuff is definitely particularly interesting. I think we didn't go into quite that much detail in the paper actually, which is a bit sad. Mostly our thinking wasn't quite as developed on it as it is now. But basically, if you actually think about the mechanics of what an AI system would have to do in order to take over, it's quite hard. Among other things, it probably wants to be able to make sure that it isn't monitored as it starts thinking about what to do.
所以,也许它想建立一个部署,让所有监控都关闭,而我们甚至不知道发生了什么。因此,它可能能够外泄其权重,然后在我们不控制的不同算力上运行自己。但如果它必须外泄自己,它要怎么做呢?谷歌的安全可不是那么容易规避的。事实上,这相当困难。安全措施相当好。所以实际上,你可以通过确保你的安全措施也适用于作为不可信内部人员的 AI 系统,来获得相当多的额外防御。你想做的事情可能包括建立基础设施,要求智能体与代表其行事的人类员工拥有独立的身份。这就像你为一个任务启动一个智能体。它不会获得你作为人类员工拥有的所有权限。它获得一个小得多的权限集,然后当它需要执行任务时,它必须单独请求每个权限,并写一份理由说明为什么需要那个特定权限。这让你可以在它撒谎之后追究它的责任。嗯,这有助于你追究责任,但可能不能保证。然后类似地,你可以进行一些监控,比如现在有一个思维链监控器,它查看模型在做什么,看看是否有任何可疑之处,如果有任何足够高度可疑的事情,就会标记给人类进行审计。这也能让你生成不对齐的证据。比如,如果 AI 系统实际上是不对齐的,你审计它,然后你会说,‘哦,天哪,是的,这个 AI 系统完全是在试图注入一个安全漏洞,然后利用它来外泄它的权重。’那是一件大事。我不确定我是否期望这会发生。模型永远不对齐似乎也很有可能。所以如果我们得到那种证据,我认为那会极大地改变我的想法,也会改变许多其他利益相关者的想法。所以我认为获得这种证据非常关键。所以,是的,很多诸如此类的事情。而且我认为做这些事情很重要,也很有意义。所以未来有很多好的工作要做。我应该指出,这与你之前邀请的 Bach 和 Ryan 在播客中讨论的 AI 控制非常相似。我认为我们这边只是更侧重于需要建立的基础设施,以便实现这种方法,而不是具体如何构建监控器等细节,后者是一个更标准的机器学习问题,相比之下我认为会相对容易解决。
So like maybe it wants to set up a deployment where all the monitors are turned off that we don't even know is happening. So perhaps it could exfiltrate its weights and then run itself on different compute that we don't control. But like if it has to exfiltrate itself, how is it going to do that? Google security is not trivial to evade. In fact, it's pretty hard. It's pretty good security. And so there's actually quite a lot of increased defense you can get by making sure that you're making your security also apply to AI systems as untrusted insiders. Things that you would want to do could include having infrastructure that requires agents to have separate identities from the human employees on whose behalf they're going to be doing things. So this would be like you are starting an agent on some task. It doesn't get all the permissions that you as a human employee have. It gets a much smaller set of permissions, and then as it needs to do its task, it has to request each of those individually, write a justification for why it needs that particular permission. This allows you to incriminate it afterwards if it was lying about it. Well, it helps you incriminate it, maybe doesn't guarantee it. And then similarly, you can have some monitoring where you have a chain of thought monitor right now, for example, that looks at what the model is doing, sees if there is anything potentially suspicious about this, if there are any things that are sufficiently highly suspicious get flagged to humans to be audited. This allows you to also generate evidence of misalignment. Like if the AI system is actually misaligned and you audit it and you're like, 'Oh my gosh, yeah, this AI system totally was trying to inject a security vulnerability and then would have exploited it in order to exfiltrate its weights.' That's a big deal. I'm not sure I expect that to ever happen. Seems quite plausible that models will never be misaligned. So if we get that sort of evidence, I think that would change my mind a big deal and many other stakeholders as well. So I think that's quite crucial to get. So yeah, lots of stuff like this. And I think it's important to do and very relevant. So lots of good work to do here in the future. And I should note that this is pretty similar to the AI control work that Bach and Ryan talk about. I think it's just on our side a little bit more focused on the infrastructure that we have to build in order to enable this sort of approach, rather than the specifics of exactly how we're going to build the monitor and so on, which is a more standard machine learning problem that I think will be relatively easy to do in comparison.
我们来谈谈时间线和递归自我改进循环。最近有很多讨论关于我们何时可能期望出现递归自我改进循环。确实,如果这种循环是可能的,你认为人们在某些方面对此的思考有点草率,而且这些方式可能导致人们可能比你更早地预期它会发生,请解释一下。
Let's talk about timelines and recursive self-improvement loops for a bit. There's been a lot of discussion recently about when we might expect a recursive self-improvement loop to occur. Indeed, if one is possible at all, if you think that people have been a little bit sloppy in their thinking about this in some ways, and in ways that cause people to maybe expect it to happen sooner than you do, explain that.
是的。所以,我认为关于现实或世界,我见过的最引人注目的事情之一就是图表上直线延伸的例子。我想 Scott Alexander 说得很好,我不记得是哪篇文章了,但我觉得他说过类似‘我不太理解直线之神’的话。老实说,它们让我有点害怕,但我不会跟它们对着干。我们最近经常看到的一条特定直线就是大约每年 3%的 GDP 增长。现在有很好的理由说明为什么这不一定能持续,反而我们可能会在某个时候看到智能爆炸,这将大幅提高增长率。但我确实认为,在推理这个问题时,非常重要的是要说明 AI 究竟在哪些方面与当前情况不同。我们为什么要跟直线之神对着干?我认为通常的答案,基于经济内生增长模型,可能最著名的是 Kmer 在 1993 年的一篇论文,大致有几个效应。一个是,随着技术进步,你可以更快地找到新想法。就像一旦你有了显微镜,你就能更好地进行生物学研究。另一个效应是,随着时间的推移,想法越来越难找到,因为你摘取了低垂的果实。最初你通过观察自己的尿液发现了磷,后来你必须通过将样本从一个国家运送到另一个国家,在世界上唯一能做的实验室中研究,才能发现元素 110。显然这要困难得多。我认为基本上在当前情况下,我们看到 GDP 持续指数增长,以及技术各个领域的类似情况,比如摩尔定律就是一个很好的例子,论点大致是想法越来越难找到,但我们也拥有指数增长的研究人员数量。这两者相互平衡,从而产生了这种持续的进步。但如果你回顾足够久远的历史,实际上增长似乎在随着时间的推移而大幅增加。那么,我们到目前为止还没有考虑到的这个秘密的第三效应是什么呢?论点是技术的增长率部分地随着更多技术而增加,因为你得到了显微镜,它们帮助你更好地做事。但也随着人口而增加,因为随着人口增长,有更多的人可以产生想法,而想法一旦产生,你可以自由复制。它们传播得非常快。这两者可以增加技术。因此,技术被建模为增长,其增长率被建模为与技术和人口都成正比,有时技术还取某个常数次方。这基本上预测了人口和技术的双曲线增长。而论点是这样的:在过去的许多年里,人口部分不再适用,因为我们不再将所有产出重新投资于更多的孩子,而是重新投资于生活质量的改善。
Yeah. So, I think one of the most striking things about reality or the world that I've seen is just the examples of straight lines on graphs going straight. I think Scott Alexander put it well in I don't remember which post this was, but like I think he says something like, 'I don't understand the gods of straight lines very well.' And honestly, they kind of freak me out, but I'm not going to bet against them. And one particular straight line that we've been seeing a lot recently is just sort of constant roughly 3% GDP growth per year. Now there are good arguments for why that won't necessarily last and instead probably we will see an intelligence explosion at some point which would drastically increase the rate of growth. But I do think that when reasoning about this it's quite important to say what exactly about AI makes it different from the current situation. Why are we betting against the gods of straight lines? Now I think the usual answer to this, which is based on economic endogenous growth models, I think maybe most famously from a paper from Kmer in 1993, is that there are roughly a couple of effects. So one, as technology improves that allows you to find new ideas more quickly. It's like once you have the microscope you can do biology significantly better. Another effect is that ideas get harder to find as time goes on because you pluck the low-hanging fruit. Initially you discover phosphorus by looking at your own pee, and later you have to discover element 110 by shipping samples from one country to another to be studied in the one laboratory in the world that can do it. Obviously that's much harder. And I think basically in the current setting where we see this constant exponential growth in GDP and similar things in various areas of technology like Moore's law being a good example, the argument would be roughly that ideas are getting harder to find but also we have exponentially increasing numbers of researchers. And together those two things balance out in order to create the sort of constant progress. But if you look sufficiently far back historically, it actually looks like growth has been increasing substantially over time. So what is this secret third effect that we haven't taken into account so far? The argument is that the rate of growth of technology increases partly with more technology because you get microscopes and they help you do things better. But also increases with respect to the population because as the population grows there are more people who can have ideas, and ideas once you have them you can copy them freely. They diffuse very quickly. And those two can increase technology. And so technology is modeled as increasing, its rate of growth is modeled as being proportional to both technology or sometimes technology raised to some constant as well as population. And this then predicts basically hyperbolic growth in both population and technology. And the argument is like, well, for the last however many years, the population part no longer applies because instead of reinvesting all of our output back into more kids, we're now reinvesting them into quality of life improvements.
这就是为什么我们看到一个双曲线增长直到 1800 年、1900 年——我记不清具体时间点——然后从那以后是指数增长。但这种说法确实把首要地位放在了想法上,而不是其他东西上。想法是可以自由复制的,你一旦拥有它们,它们就能应用于你所有技术的方方面面。所以当我思考什么会导致智能爆炸时,我认为能够产生想法的人口或劳动力的增长是相当关键的。
And that's why we see a hyperbolic up until who knows 1800, 1900, I forget the exact point, and then since then exponential growth. But this sort of story really puts the primacy on ideas rather than anything else. Those are the things that are freely copyable. Those are the things that you have once and then they apply to all of your technology everywhere that you're using it. So when I think about what's going to lead to the intelligence explosion, I think the increasing population or increasing labor that can have ideas is pretty crucial to it.
相比之下,今天很多讨论都在谈论一个超人类编码器,比如你有一个 AI 系统,你可以说‘请为我实现 XYZ 实验’,它就能非常出色地、比人类快得多地实现那个实验。论点是一旦你能做到这一点,AI 进步的速度就会大大加快,最终你会很快开始进入智能爆炸。至少我认为这是论点。要解读整个社区的观点总是有点困难,因为很多人有略微不同的意见。我不太买这个观点,主要是因为超人类编码器并不能在所有 ML 研发领域产生想法。它特别在编码领域有想法,这是 ML 研发的一部分,但据我看来绝对不是大部分。
In contrast, a lot of the discussion today talks about a superhuman coder, for example, where you get an AI system where you can say 'please implement XYZ experiment for me' and it just goes ahead and implements that experiment incredibly well and much faster than humans could do it. The argument is that once you can do that, the rate of AI progress increases a bunch and eventually you start getting to the intelligence explosion quite quickly. At least I think that's the argument. It's always a little bit hard to interpret the wide community view which has a lot of people with slightly different opinions. I don't super buy this view mostly because the superhuman coder is not something that can have ideas across the spectrum across everything that happens in ML R&D research. It has ideas within the realm of coding in particular, which is one part of ML R&D, but definitely not even the majority of it according to me.
我认为最好把这看作更像是发明显微镜。你有了一个工具,能让你的研究进展更快,但我们在过去几十年里一直在各种研究领域这样做。它并没有导致双曲线增长。工具的发明只是研究进展的正常部分,通常会导致图表上漂亮稳定的直线。所以我通常不认为我们应该从工具的发明中预测巨大的变化。
I think it's better to think of this as more like inventing the microscope. You've got a tool that enables your research to progress faster, but we do this all the time in all sorts of research areas for the last however many decades. It has not led to hyperbolic growth. The invention of tools is just a normal part of research progress and usually tends to lead to nice stable straight lines on graphs. So I generally don't think we should be predicting big changes from the invention of tools.
另一个例子是,有时人们会建议智能爆炸的触发点是 AI 研究人员变得比如比完全没有 AI 系统或只有 2020 年左右的 AI 系统时生产力高 10 倍。同样,我认为这种事情可能已经发生在芯片开发上。摩尔定律几十年来是一条漂亮的直线。但可以想象,在这条线的起点,人们用手在纸上设计电路,而今天我们有了令人难以置信的计算机辅助设计软件工具,可以完成大部分工作,自动化绝大多数这类工作,而你只看到这条线继续是直线。事实上,你需要指数级更多的人来工作才能实现这一点。那么今天的研究人员比二十年前的生产力高 10 倍吗?可能吧。这是我的猜测。这是芯片设计智能爆炸的迹象吗?不是。我认为类似的事情可能也适用于 AI。事实上,这有点不清楚。我们已经大量使用语言模型,例如用于自动化、进行评估和自动化其他模型的响应。没有它们我们的生产力会低多少?我不知道。可能不会低 10 倍,但会低很多。但根据我的看法,AI 进步看起来仍然相当线性。
Another example I think is sometimes people will suggest that a trigger for the intelligence explosion is the point at which AI researchers can get, say, 10x more productive than they would have been if they had no access to AI systems at all or had access to AI systems from like 2020 or something like this. Again, I think this is the sort of thing that probably has happened for something like chip development. Moore's law is a nice clean straight line for many decades. But presumably, I assume at the beginning of that line people were designing their circuits by hand on paper and nowadays we have these incredible computer aided design software tools that can do most of that, that automate the vast majority of this kind of work, and you just see the line continuing to be straight. In fact, you need exponentially more people working on this in order to have that happen. So are the researchers that we have today 10x more productive than the ones from two decades ago? Probably. That would be my guess. Is that a sign of an intelligence explosion in chip design? No. I think a similar thing might be true in AI. In fact, it's kind of unclear. We already use language models quite a lot, for example, for automating and for conducting evaluations and automating the responses of other models. How much less productive would we be without that? I don't know. Probably not 10x less productive, but a substantial amount. But AI progress continues to look pretty linear according to me.
让我总结一下。所以我们目前有指数级的经济增长,也就是说,从中期来看,平均每年有一个恒定的百分比增长。那些说会有智能爆炸、会有递归自我改进循环的人,他们预测的是增长率不断提高或双曲线经济增长。所以一年是 3%,接下来是 10%,再接下来是 50%,直到某个点,我想,那时你可能会认为它会趋于平稳或再次下降。看看导致经济增长稳定、增长或下降的高层因素,你的模型和大多数经济学家的模型中有三个因素。第一,随着技术进步,做出新的有用发现并进一步推进技术变得更加困难。第二,随着技术进步,我们有更好的工具来做科学和做出新发现。第三,随着技术进步和经济成长,我们可以支持更多的人来做研究、产生想法和推进科学。最近,第三个因素已经不太起作用了,因为我们没有把科学进步或经济改善转化为新的人口。尽管我们更富有了,出生率却在下降而不是上升。所以一直是前两个因素在相互抵消。进步越来越难做出。科学进步给了我们更好的工具来做这件事。这两件事大致相互抵消,中期增长率大约每年 3%,最近可能还在下降。再往远看,在马尔萨斯时代结束之前,所有三个因素都在起作用。随着我们变得更富有,我们也把几乎所有这些资源投入到了额外的人口上,但这在 1800 年后消失了。而这导致了双曲线增长,导致了增长率的不断提高,当时确实如此。
Let me recap all of that. So we currently have exponential economic growth, which is to say a constant percentage growth each year on average looking over the medium term. People who say there's going to be an intelligence explosion, there's going to be a recursive self-improvement loop, they're forecasting increasing rates of growth or hyperbolic economic growth. So it's 3% one year, then 10% the next, then 50% the next up to some point, I guess, at which you might think it levels off or comes back down again. Looking at the high level factors that lead economic growth to either be stable or to increase or to go back down, there's three factors that you have in mind in your model and that most economists have in their model. As technology advances, it gets difficult to make new useful discoveries and to advance it further. That's the first one. The second one, as technology advances, we have better tools to do science and to make new discoveries. The third one is as technology advances and the economy grows, we can support a larger number of people who will do the research and have the ideas and advance science. In recent times, the third one has kind of been out of the picture because we haven't been turning advances in science or improvements in the economy into new people. Birth rates have been going down rather than going up despite us being richer. So it's been the first two factors that have been playing off against one another. Advances are getting harder to make. Science advancing gives us better tools to do it. These two things have been roughly cancelling out and over the medium-term growth has been about 3% a year, maybe going down in recent times. Zooming out much further, all three factors were in play before the Malthusian era ended. As we got richer, we drove much almost all of those resources into additional people as well, but that faded out after 1800. And that was leading to hyperbolic growth, leading to increasing rates of growth while it was true.
现在,将这一点应用到 AI 案例中,要弄清楚我们处于哪种情况,这个思维模型非常强调:我们是在想出更好的科学工具,还是把我们的收益再投入到更多的研究人员身上,让他们做科学,让更多的人想出更好的想法或新想法?你可以看到,当我们谈论 AI 做研究时,这里为什么会变得有点混乱,因为 AI 既是一个工具,但我们又认为它基本上会融合并变成人,或者本身成为科学研究者。所以这成了一个困难的概念性实证问题。在什么节点上,它们从主要是协助人类工作的工具,转变为本身就是研究者,不再仅仅是协助人类,而是完成整个过程,或者至少我们不应该再仅仅把它看作是一个让我们能做更好工作的科学仪器?
Now, applying this to the AI case, to figure out which situation we're in, this mental model puts a huge emphasis on: are we coming up with better tools to do science, or are we plowing our gains back into more researchers to do science and to have more people thinking up better ideas or new ideas? You can see why it gets a bit confusing here when we're talking about AI doing the research because AI is both a tool but we're thinking it's going to basically converge and become people or become scientific researchers itself. So it becomes a difficult conceptual empirical question. At what point do they cross over from being largely a tool that is assisting humans in doing their work to being the researcher itself that isn't really just assisting people, it's doing the whole process, or at least we should no longer be just considering it as a scientific instrument that's allowing us to do better work.
而且我觉得你想说的是,人们在谈论那些真正表明我们开发出了更好的工具供人类使用的指标时,有点不够严谨。他们有点把这种进步和那种几乎不需要人参与的自主 AI 研发混为一谈,后者实际上应该被视为人口增长。这种增长可以驱动超指数增长,因为随着改进,你可以运行更多副本,基本上通过运行更多模型副本来扩大研究人员的有效人口。我理解得对吗?
And I think you want to say people are being a bit sloppy about how they talk about indicators that really are a sign that we've developed better tools for humans to use to do their work. And they kind of conflate that with we've developed AI, you know, autonomous AI R&D that barely even needs people at all and really should be considered as a population increase. That then can drive hyperbolic growth because as you get improvements, then you can run even more. You can basically expand the effective population of researchers by running more copies of the model. Have I understood right?
是的,没错。非常精彩。我担心自己独白太久了,但总结得很好。
Yeah, that's right. That's very impressive. I was worried that I was soliloquizing for way too long, but great summary.
好的,那我们怎么判断我们是在谈论更好的工具还是更多的人口呢?似乎不会有明显的分界线。
Okay, so how do we tell whether we're talking about a better tool or about more population? It does seem like there's not going to be a sharp cutoff.
是的,我同意。可能不会有明显的分界线。我的意思是,我确实认为,AI 在多大程度上提出新想法?这些是我认为最影响超指数增长预测的因素。这是你可以关注的一个方面。而我认为超人类程序员可能并没有达到那个标准。但我真正想做的是衡量 AI 的进步,然后注意它何时开始加速。所以我们最近和 Epoch 合作开发了这样一种衡量方法。那篇论文我想刚刚发布,可能叫《AI 基准的罗塞塔石碑》。基本上,我们收集了大量模型在各种基准上的得分,然后将这些基准拼接起来,得到一个适用于从 2020 年到现在所有模型的通用能力分数,尽管单个基准在更短的时间内就会饱和。然后我们把这个衡量方法得出的能力分数与模型的发布时间画成图。这些能力分数在生成时,我们从未告诉统计模型任何关于发布日期或时间的信息,只有基准表现。但当你把它们与发布时间画成图时,它呈现出一条漂亮的直线。太棒了,你会看到 AI 进步在这个图上看起来基本上是线性的。这个方法非常简单,主要依赖于我们有能够捕捉 AI 系统性能的基准。是的,没错。而且我认为未来这种情况会继续。所以我们可以继续画这个图,然后希望如果智能爆炸真的发生了,我们会开始看到加速,而不是一直呈现漂亮的线性拟合。
Yeah, I agree. There probably won't be. I mean, I do think that, you know, to what extent are the AI proposing new ideas? Those are the things that I think most influence the hyperbolic growth prediction. Like that's one thing you could be looking at. And I would say that the superhuman coder probably doesn't really hit that bar. But really what I want to do is just measure AI progress and then notice when it starts accelerating. And so actually we worked with Epoch recently to develop exactly this kind of a measure. It's the paper I think was just released. It's called a Rosetta Stone for AI benchmarks probably. And so essentially we just take the benchmark scores that a model gets for a wide variety of models and then we sort of stitch the benchmarks together to get a sort of general capability score for models that applies over the entire range from you know back in 2020 all the way to models that were released now, even though any individual benchmark would have saturated over a much smaller period. And then you take the capability scores that this measure spits out. You plot them against the release time for the model. And you know these capability scores we never told when producing them the statistical model that produces them. We don't include any information about the release date or any time-based information into it. Just benchmark performance. And nonetheless when you start plotting this against release time it's just a nice line. It's great and you see that AI progress looks mostly pretty linear on this particular graph. And the method is really very simple. The main thing that it depends on is that we have benchmarks that can capture AI systems' performance. Yeah, exactly. And that I think will continue to happen in the future. So we can just continue to plot this and then one hopes that if an intelligence explosion does seem like it's happening, we will start to see an acceleration on that rather than it looking just like a nice linear fit the entire time.
好的。那么,我们怎么判断我们只是做出了更好的工具,还是创造了一大批新的研究人员呢?你说事实胜于雄辩。我们不要把它留给哲学家,而是让做基准测试的人来判断 AI 进步是否真的在加速。我想很多人觉得 AI 进步在放缓。你看到你们试图创建一个巨大的数据集,包含过去四五年尽可能多的不同模型。这是和 Epoch 合作的,他们擅长这种数据收集、整理和汇总。所以他们试图收集许多不同模型在很长时间内的各种基准分数,看看进步是在加速、放缓还是大致呈线性。而你说,至少根据这个衡量标准,他们能做出的最大努力表明进步是线性的,也就是说,我想我们目前并没有制造出更好的人,而是制造出了更好的工具。这就是这个结果所暗示的。
Yeah. Okay. So how do we tell if we've just made a better tool or we've made a whole lot of new researchers? You're saying the proof is in the pudding. Let's not leave it to the philosophers. Let's leave it to the people doing benchmarks to figure out whether AI advances are actually speeding up. A lot of people have the perception I guess that AI progress has been slowing down. You're seeing you've tried to create an enormous dataset of as many different models over the last four or five years as possible. This is with Epoch who are this is their wheelhouse doing this kind of data collection and compilation and aggregation. So they've tried to collect all these different benchmark scores for many different models going back quite a long way to see is progress speeding up or is it slowing down or is it roughly linear? And you're saying at least judged by that measure, the best effort they can do says that it's linear, which is to say, I guess that we're not making better people, we're making better tools for now. That's what that would suggest.
是的,没错。基准分数存在各种问题。一方面,它们可能无法捕捉到全部性能范围,要么在底部有缺陷,要么在顶部有上限。特别是因为我们讨论的是多年来的模型,做很多不同的事情。肯定有很多作弊行为,或者很多人为了让自己模型在基准上表现好而进行应试训练。我想还有一个问题是什么才是真正重要的。有各种不同技能的基准。也许你应该给某些方面更多权重。我猜模型性能与其经济影响之间可能存在非线性关系,而且很可能相当非线性。你认为你和 Epoch 用来判断进步是加速还是保持相同速度的整个方法,在揭示实际情况方面有多好?
Yep, that's right. So, there's all kinds of problems with benchmark scores. One thing could be that they don't capture the full range of performance that either you end up like, you know, capped at I guess flawed at the bottom or capped at the top. Especially because we're talking about models here over many years doing many, many different things. There's definitely lots of gaming that goes on or like lots of teaching to the test that occurs with people trying to make their models look good on these benchmarks. I guess it's also a question like what actually matters. There's benchmarks for all kinds of different skills. And maybe you should give some of these things much more weight than others. I guess you could also have nonlinearities in the effect of a model's performance and its economic effect that could be indeed probably is quite nonlinear. How good do you think this whole approach that Epoch and you have been using to figure out whether progress is speeding up or remaining about the same pace at getting at the kind of ground reality of what's going on?
是的。我基本上同意你提到的所有批评,比如它会继承基准的所有问题,因为最终输入只有基准表现。尽管如此,我认为我从直线之神那里学到的一课是,对于某些事情,是的,有很多细微差别和细节,如果你想做非常精细的预测,它们很重要,但在高层次上,它们会相互抵消,最终结果还是不错的。我认为这基本上适用于这里。
Yeah. I basically agree with all the critiques you mentioned like it's going to inherit any problems that benchmarks have because ultimately the only input into it is benchmark performance. Nonetheless, I think a lesson that I learned from the gods of straight lines is that for some things, yeah, there are lots of nuances and details that matter a bunch if you want to make very fine grain predictions, but at a high level they wash out and it ends up being fine anyway. And I think that mostly applies here.
我想一个例子可能是,你会说,模型被作弊了。现在有很多应试训练,但两年前和四年前也有很多应试训练。所以,只要这种情况没有越来越严重,那么这条线仍然是合理的。
I guess an example might be you would say, well, the models are being gamed. There's a bunch of teaching to the test happening now, but there was a bunch of teaching to the test happening two years ago and four years ago. And so, as long as that's not getting progressively worse, then the line is still reasonable.
是的。而且,我认为在实践中,应试训练不会对结果产生太大改变。它确实会改变一些。所以,我认为 Anthropic 可能最擅长不过度拟合现有基准。事实上,如果你看产生的分数,我认为它确实倾向于低估 Claude 或 Anthropic 的模型,相对于其他提供商的模型。因为一般来说,Claude 模型由于没有过度拟合基准,在基准上的表现往往低于模型的实际能力。所以,是的,这是真的。但如果你在图上观察,这只是很小的差异,并不大。
Yes. And also, I think in practice, the teaching to the test won't change the results very much on this. It definitely changes it some. So, I think like Anthropic is probably the best at not overfitting to the benchmarks that exist. And in fact you will, if you look at the scores that are produced, I think it does tend to underestimate Claude or Anthropic's models relative to models from other providers. Because generally Claude models since they are not overfit to the benchmarks will tend to underperform on benchmarks relative to how good the model actually is. So yes, that's true. But also, if you look at it on the graph, it's like, you know, this tiny little difference. It's not that big.
就像如果你把它往上移,它看起来仍然是线性的。这并不会改变它是否线性的观察结果。
It's like if you moved it up, it would still look linear. It doesn't really change the observation of whether it's linear or not.
我想你是说,进步速度的加快在图表上会非常显眼。它很可能会跳出来,而这些效应不足以让它消失。
I guess you're saying increasing the rates of progress would be quite striking on the graph. It probably would jump out at you and these effects wouldn't be enough to make it disappear.
没错。但我肯定会告诫人们不要用这个分数来做非常精细的比较,比如 Claude 和 Gemini 哪个更好,或者试图理解开源模型和闭源模型之间的差异。我认为开源模型可能比闭源模型对基准测试过拟合得更厉害,如果你试图用这个分数来看两者之间的确切差异,它可能会误导你。
That's right. But I definitely would caution people against using the score for really fine-grained things, like how good is Claude versus Gemini versus whatever, or for example trying to understand the difference between open source models and closed source models. I think the open source models are probably more overfit to the benchmarks than the closed source ones, and if you try to use this score to look at the exact difference between those two, it's probably going to mislead you a little bit.
那么,当 AI 成为人,或者至少是 AI 研究者,而不仅仅是 AI 研究者的工具时,你可能会合理地预期 AI 研发和 AI 能力会出现相当突然的进步。你是对那种突然性持怀疑态度,还是对是否会发生持怀疑态度?或者你相信也许它会比人们想象的要长一点,但你仍然认为它会发生?
So, at the point that AI is people, or at least AI researchers, rather than just being a tool for AI researchers, you might reasonably expect quite abrupt increases in progress in AI R&D and AI capabilities basically. Are you a downvote on how abrupt that will be, or whether that will occur at all? Or do you buy that maybe it will take a bit longer than people are imagining, but you still think that will happen?
我对突然性持轻微怀疑态度。但我并不怀疑是否会出现智能爆炸。我觉得,我对这一点相当有信心,但这将是一个非常强的论断,带有许多强前提条件。例如,当你拥有这样的 AI 系统,以至于对于任何你关心的经济上有价值的任务,你都会更愿意雇佣 AI 系统而不是人类,这其中的含义包括 AI 比人类更便宜,而我不认为这是理所当然的。在那个时候,假设我们不采取任何行动来阻止它,并且我们试图利用 AI 系统来大幅加速整个经济的研发,那么大概在一个世纪内,我们将会出现某种智能爆炸并达到技术成熟。这仍然是一个疯狂的论断,有很多前提条件。我说我们将会出现智能爆炸,我对此实际上相当有信心。但这比我猜测实际会发生的情况要弱得多,实际发生的情况会快得多,也早得多。所以我实际上要说的是,有相当大的可能性,它会在 AI 系统能够自动化大部分 AI 研发(而不是所有任意研发)的时候开始,并在 5 到 10 年内完成,而不是一个世纪,具体取决于你设定的起点。
I'm a slight downvote on the abruptness. I am not a downvote on whether there will be an intelligence explosion. I think, you know, I feel pretty confident about it, but it will be some really strong statement with a lot of strong preconditions. For example, when you have AI systems such that for almost any economically valuable task you care about, you would prefer to hire the AI system rather than the human, which among other things implies that the AI is cheaper than the human, which I don't think is a given. At that point, assuming we don't take some action to try and prevent it, and we are trying to use the AI systems to substantially accelerate research and development across the entire economy, then probably within a century we will have some kind of intelligence explosion and reach technological maturity. That's still a crazy statement with many preconditions. I'm saying we will have an intelligence explosion at all, and I feel actually pretty confident about that. But it is a lot weaker than what I would guess will actually happen, which will be substantially faster and happen substantially earlier. So what I would actually say is there's a decent chance that it starts at the point where the AI systems can automate most of AI R&D, as opposed to all arbitrary R&D, and finishes over 5 to 10 years rather than a century, depending on exactly where you set the starting point.
回到你最初关于突然性的问题,我对突然性持轻微怀疑态度的原因是,我预计实际上第一个能够真正自动化 AI 研发的 AI 系统可能会非常昂贵,甚至可能比人类更贵。过去一年你可以看到,推理时算力和推理时 Scaling 的投资大幅增加。谷歌有一种深度思考算法,它应用更多的推理 Scaling 以获得更好的结果。我认为你应该基本上预期这种情况会继续。所以第一个自动化研究者的图景可能是一个相对愚蠢的系统,不如人类研究者聪明,但花费大量时间进行推理,探索大量死胡同,意识到它们是死胡同,然后回来尝试其他东西,而人类永远不会走那些路,因为他们事先就知道那是死胡同。这就是它进行自动化研究的方式,以至于它可能比人类研究者更昂贵。然后我们随着时间的推移改进它,成本会像 AI 通常那样迅速下降。但这表明它不会那么突然;你会看到 AI 达到与人类成本持平,然后开始变得更具成本效益,从而引发加速。
Going back to your original question about abruptness, the reason I'm a slight downvote on abruptness is that I would expect that actually the first AI systems, automated systems that can really automate AI R&D will probably just be very expensive, to the point of potentially being more expensive than humans would be. You can see this over the last year with way more investment in inference time compute, inference time scaling. Google has this sort of deep think algorithm which applies even more inference scaling to get even better results. I think you should basically expect this to continue. So the picture of the first automated researcher might be something like a relatively dumb system that's not as smart as a human researcher, but spending tremendous amounts of time doing reasoning, exploring tons of dead ends, realizing they're dead ends, and then coming back and trying something else, which a human would never have gone down because they would have known in advance it would be a dead end. That's how it does its automated research, such that it might even be more expensive than a human researcher would be. Then we improve it over time, the cost goes down pretty quickly as is usually the case in AI. But that would suggest that it won't be that abrupt; you'll get the AI reaching cost parity with humans, then start becoming more cost effective, and that will start setting off the acceleration.
但这就是它如何在相当多年的时间里发生,而不是几个月或什么疯狂的事情。
But that's how it ends up happening over quite a number of years rather than months or something crazy.
如果你从 AI 系统开始勉强与现有的人类 AI 研究者持平的那个点算起,那么是的,没错,这就是为什么我认为需要几年而不是几个月。但我时间线延迟的主要原因是我认为即使达到那个点也需要相当长的时间,比如十年,我觉得这在过去几年会被认为是极短的时间线,而现在则被称为中到长时间线。
If you take it from the point at which the AI systems start being just barely on par with existing human AI researchers, then yes, that's right, that's why I would think it would take years rather than months. But most of my timelines delay is just thinking that even getting to that point will take quite a while, like a decade maybe, which I feel like in past years would have been called extremely short timelines and nowadays gets called medium to long timelines.
是的,我认为这些年来一直有一个普遍现象。每隔几年,就会有一次关于 AI 时间线的恐慌,人们开始预期递归自我改进循环很快就会发生,就在那个点之后的几年内。我的印象是你对这两个方向都没有动摇。为什么你没有根据已经发生的事件、已经出来的结果进行更新?
Yeah, I think there's been a general phenomenon over the years. Every couple of years there's a kind of freak out about AI timelines and people start expecting a recursive self-improvement loop really quite soon, within a few years of that point. My impression is that you've just been unmoved by either direction. Why haven't you updated based on events that have occurred, results that have come out?
我想最大的区别可能是我心中对 AI 进步的方式有一个图景,而现实与它相当接近。例如,我认为最大的时间线恐慌来自于推理模型如 o1 特别是 o3 的出现,其关键思想是将强化学习应用于大型语言模型。我很久以前,可能至少是 2019 年甚至更早,就认为我们当然需要使用强化学习来开发强大的 AI 系统。如果你看看当时做的很多工作,比如辩论,那些都隐含地以强化学习作为构建强大 AI 系统的首选方法为前提。如果你不使用强化学习,辩论看起来就毫无意义。
I would guess probably the biggest difference is that I had a picture in my mind about how AI progress would happen, and reality has been reasonably close to it. For example, I think the biggest timelines freak out was from the advent of reasoning models like o1 and particularly o3, where the key idea is applying reinforcement learning to large language models. I have since quite a long time, probably at least 2019 but maybe even earlier, thought that of course we are going to need to use reinforcement learning in order to develop powerful AI systems. If you look at much of the work done at the time, things like debate, those are implicitly predicated on reinforcement learning being the method of choice for building powerful AI systems. Debate just looks pretty pointless if you're not using reinforcement learning.
嗯,所以我一直预期强化学习会发生。嗯,从我的角度看,其他人都对强化学习的出现感到惊讶。嗯,而我看着它,心想:也许一旦我们做了强化学习,它就会像指令遵循那样完美地泛化到所有事情上,但实际上并非如此。所以我认为,我的时间线大多没有太大变化,因为我本来就觉得强化学习不太可能泛化到一切。但我想,如果我当时跟踪得足够细致,我可能会在 01 或 03 发布时稍微向更长的时间线调整。
Um and so I was always expecting reinforcement learning to happen. Um and from my perspective, everybody else was suddenly surprised by reinforcement learning happening. Um whereas I looked at it and I was like well it could have been the case that like once we do reinforcement learning it just generalizes beautifully to everything the same way that instruction following really does generalize beautifully to everything and like actually that was not the case. So I think you know mostly my timelines didn't change very much because I already thought it was reasonably likely RL wouldn't generalize to everything but I think if I had been tracking it sufficiently fine grained I would have probably updated slightly towards longer timelines on the release of 01 or 03.
那为什么推理模型的整体表现没有……我的意思是,我认为这是一种描述更新的方式:人们感到惊讶或震惊,因为强化学习被应用到这些模型中进行推理。嗯,我想大多数人会说,他们对推理效果之好以及它看起来多么有用印象深刻,但听起来你并没有被打动——它并没有让你觉得出奇地好,或者也许不是这样。而是它没有那么强的泛化能力:它们在经过强化学习训练的推理任务上表现很好,但并没有泛化到其他任务上,也没有带来巨大的经济变革,事实也确实如此。
And why doesn't the general performance of reasoning models I mean I think uh that's like one way of characterizing the update was people were surprised or they were shocked that RL was being applied to these models into reasoning. Um, I think most people would say that they were impressed by how good the reasoning was and how useful it seemed like it was going to be, but it sounds like you weren't impressed by it wasn't surprisingly good to you or or maybe that's not it. It's that it wasn't as generalizable that they were good at the reasoning task that they had been RLED on, but that wasn't you. I didn't expect that that was going to generalize to other tasks and be very economically transformative and indeed it has not.
是的,基本没错。我认为泛化能力对我来说是件大事。你看,如果你想针对某个特定基准并应用机器学习,我认为机器学习的经验是:是的,你可以做到,人们确实会选择模型能够胜任的任务。所以,并不是说你可以随便选个东西,用机器学习的锤子一敲就能成功。但我确实认为,你必须非常小心:如果有人针对某个特定事物进行了优化,你应该在多大程度上将其更新为 AGI(通用人工智能)——完全的通用智能?这确实是一个非常困难的更新,你应该更多地关注泛化能力。
Yeah, that's basically right. I think the generalizability was the big big deal for me. like look if you want to target some particular benchmark and apply machine learning to it I think the lesson of machine learning is like yes you can do it um people do choose the ones that are that models are capable of doing um so it's not like you can choose some arbitrary thing and just hit it with the machine learning hammer and succeed um but I do think you have to be like pretty careful about you know if somebody optimized for a specific thing How much should you update from that to like AGI the full generality uh the fully general intelligence? It's like really quite a difficult kind of update to make and you should be looking quite a bit at the generalizability.
好的,所以让你对时间线感到恐慌的事情是,如果你在一种实际任务上训练模型,然后发现它们实际上擅长且有用,能够处理完全不同类型的任务。
Okay, so that's the thing that would cause you to have a timelines freak out would be if you trained models on one kind of practical task and then you found that they were actually like good and like useful at doing quite different kinds of tasks.
是的,我认为没错。而且,明确地说,你确实从推理模型中看到了一点这种泛化。它并不是完全没有泛化,但确实没有达到我认为会让我大幅更新预期的程度。我特别关注自主任务,现在我认为推理模型在自主任务上表现不错。嗯,实际上,相对于预期,我认为它们仍然表现不佳,但如今它们比当时好多了,部分原因是公司已经开始针对经济性进行训练。
Yeah, I think that's right. And you do see a little bit of this from reasoning models to be clear. It's not like it generalizes not at all but yeah not as much as I think would have actually been a substantial update for me. I think in particular I was like looking at autonomy tasks and nowadays I think the reasoning models are pretty good can be pretty good at autonomy tasks. Well kind of well actually relative to expectations I think they're still underperforming but nowadays they get they're substantially better than they were at the time but partly that's because companies have started to train for economy.
我们来谈谈 Google DeepMind,这是一个非常重要的组织,但我觉得外界(包括听众)对它了解不够。我想说,总的来说,很多 AI 公司以外的人都在做研究,包括治理研究和技术研究,希望这些研究能被那些公司(包括 GDM)阅读、吸收、采纳和使用。人们怎么做才能更有可能让他们的工作真正产生影响,比如被 AI 公司内部的人阅读或使用?
Let's talk about Google DeepMind which for a very important organization I feel like isn't super well understood by people outside including listeners and I guess I would say outside just in just in general a lot of people outside of AI companies produce research I guess like both governance research and technical research hoping that it will be read and absorbed and adopted and used by those companies including including GDM. What can people do to make it more likely that anything that they do actually does have any imp like it's read at all or is or is used at all by people inside an AR company?
是的,人们确实可以做一些事情。嗯,我认为在这一点上,引入一个模型是有用的:公司内部存在大量相互交织的约束,因此我们只能做少数几件事,并且需要付出相当多的努力才能让它们通过。嗯,你可能会用“公司冷漠”这个梗来预测,但我认为更准确的说法是:公司动机良好,但只能做有限的事情。嗯,所以记住这一点,有几件事值得做。一件就是,在你投入大量时间做研究之前,先和公司里的人谈谈。嗯,比如联系一位在公司工作的人,说:“这是我计划要做的事,你觉得它真的有用吗?”这是一个非常简单的步骤,但令人惊讶的是,很多人就是不做。嗯,但这似乎是目前提高你工作被公司采用概率的最好方法。
Yeah, there are definitely a few things that people can do. Um, and I think a again at this point it is useful to bring in the model of like there just like all these incredible constraints that interact with each other a ton and as a result we can only do a few things and need to do a decent amount of work in order to get them through. um which you know you could well predict as like company uh using the meme of like companies are apathetic though I think it's like a maybe a little bit better to say it as like companies are well motivated but can only do a few things um and so you know keeping that in mind um a few things that are worth doing um one is just like talk to somebody at a company Um before you like do put in a bunch of time uh into research it's like you know reach out to somebody who works at a company say like this is what I'm planning to do. Do you think this will actually be useful? Uh it's a very simple step but like surprisingly many people just don't do it. Um, but it does seem like probably the best thing you can do in order to actually increase the chance that your work gets used um at a company at least currently.
如果人们一直这样做,你觉得你的很多时间会不会被回复这些邮件占用?
If people consistently did that, do you think it'd be like a lot of your time would then be eaten up replying to these emails?
呃,如果他们给我发邮件,可能不会得到太多回复,因为我收到的邮件太多了。而且我认为,我的大多数下属应该会很乐意和人们谈论这个。嗯,我承认,我现在收到的这类邮件足够多,感觉更像是一种负担而不是令人兴奋的事情,但在它变得非常普遍之前,它确实曾让人感到兴奋。
Uh, if they email me, they will probably not get very much of a reply because I get way too many emails. And I think, you know, most of my reports, I think, would be fairly excited to talk to people about this. Um, I admit that I at this point get enough of these that it feels a little bit more like a burden than like an exciting thing, but it definitely used to feel like an exciting thing before it became very common.
好的。嗯。除了发邮件询问他们正在做的事情是否有用之外,人们还应该做什么?
Okay. Yeah. What should people do other than email to ask whether what they're doing is going to be useful?
是的。那么,其他事情……嗯,我最近有一个演讲,主题是“如何理论化,让经验主义者倾听”。它也可以叫做“如何做安全研究,让公司倾听”。嗯,其中的基本要点,我认为第一个也是最明显的,但仍然值得问自己的是:他们真的在乎吗?嗯,有时人们研究的问题,我们实际上并不关心,也不认为重要。嗯,我认为安全社区在这方面做得相当好,但这种情况仍然会发生。嗯,而且有时人们可能会对我们关心什么、不关心什么感到有点惊讶。嗯,举个例子,比如越狱。我们现在非常关心越狱,因为滥用问题——模型开始变得足够强大,滥用可能成为一个严重问题。
Yeah. So, other things um I think I guess the there I had a talk recently on like how to theorize so empiricists will listen. It could equally well have been called how to do safety research so that companies will listen. Um the basic points from this I think the first one the most obvious one but that like is still worth asking yourself is like do they actually care? Um sometimes people do research on problems that we actually just don't care about and don't think matter. Um, I think the safety community is like fairly good about not doing this, but it does still happen. Um, I think and sometimes people might be a little bit surprised by what we do and don't care about. Um, so for example, um, take jailbreaks. We care a lot about jailbreaks now because misuse like we the models start are looking to be strong enough that misuse is could be a serious problem.
但如果你看一两年以前,人们做了很多越狱研究,说‘看,对齐失败了;公司不擅长对齐模型,因为它们太容易被越狱了。’我不了解其他公司,但至少在 GDM,我们的立场是,鉴于模型的能力,并没有我们特别担心的实际滥用场景。安全是关于保护用户免受模型做出有害行为的情况。有一段时间,大家都在讨论模型总是可以被越狱,完全无法防御。真正的答案是,我们甚至还没尝试阻止越狱。
But if you look a year or two ago, people were doing a bunch of jailbreak research and saying, 'Look, alignment fails; companies aren't good at aligning their models because they're so susceptible to jailbreaks.' I don't know about other companies, but at least at GDM, our stance was that there aren't any actual misuse scenarios we're particularly worried about given the model capabilities. Safety is about protecting users from cases where the model does something harmful. For a while, there was discourse about how models are always jailbreakable and impossible to defend against. The real answer was that we hadn't even tried to stop the jailbreaks.
好吧。但你现在正在尝试阻止它。
Okay. But you're trying to stop it now.
我们现在正在尝试阻止它。是的。现在,如果有人想给我们提供关于越狱以及如何防御的研究,我们肯定会使用它。
We are trying to stop it now. Yes. Now, if people want to give us research on jailbreaks and how to defend against them, we will definitely use it.
我想我应该说明,所有这些建议都基于一个前提:你做的研究是希望某家 AI 公司会阅读、吸收并使用它。有些人会有自己不同的想法,他们开发东西是希望以后人们会被说服,认为它有用,或者它会以另一种方式变得有价值。
I guess I should say all of this advice is predicated on the idea that you're doing research hoping that an AI company is going to read it and absorb it and use it. There are people who will have their own different ideas and they'll be developing stuff hoping that people will become persuaded later on that it's useful or it'll be valuable in a different way.
当然。研究有很多不同的变革理论。我绝对只是在说如果你希望一家公司现在就用它,基本上就是现在。总之,那是第一步:他们真的在乎吗?第二步是我称之为‘你在帮忙吗?’,但这是非常特定的一种帮忙。具体来说,我认为你要么提出一个解决方案,要么构建一个评估或某种指标。有些研究不是这些——比如做一些理论来更好地理解某个现象,这能增加理解,但不解决特定问题,也不产生公司应该关心的指标或评估。还有很多其他例子。我认为基本上大部分这类研究我们不太可能使用,除非它恰好是我们非常关心的核心问题。鉴于我们都很忙,整合任何新东西都需要大量工作,通常它需要是一个指标、评估或某种解决方案。
Absolutely. Lots of different theories of change for research. I'm definitely only talking about if you want a company to use it now, basically now. Anyway, that was step one: Do they actually care? Step two is I call it 'Are you helping?' but it's a very specific flavor of helping. Specifically, I think you should either be proposing a solution or building an evaluation or a metric of some kind. There is research that is not these things—research that does some theory to understand a phenomenon better, which gives increased understanding but doesn't solve a particular problem or produce a metric or eval that the company should care about. There are lots of other examples. I think basically most of that research we are unlikely to use unless it happens to be on a core problem we care about a lot. Given that we are all busy and it takes a lot of work to incorporate any new thing, usually it needs to be a metric or an eval or some kind of solution.
所以人们在做有趣的理论化或初步的实证工作来更好地理解某个现象。你会说,‘这都挺好。等你有解决方案了再来。’
So people are doing interesting theorizing or preliminary empirical work to understand a phenomenon better. You'll be like, 'That's all well and good. Come back when you have a solution.'
大致上是这样。再说一次,我认为这是好的研究,人们应该去做。它可以帮助未来开发更好的解决方案。我只是不认为他们应该觉得我们会花很多时间阅读它。
Roughly yes. And again, I think this is good research to do and people should do it. It can help develop better solutions in the future. I just don't think they should be thinking that this is something we'll spend a lot of time reading.
好的。接下来,特别是如果你要提出一个解决方案,下一步是确保你的解决方案至少在概念上,即使没有通过实验,也要在公司非常关心的各种指标上进行评估。大多数时候,研究会看它是否真正解决了它要解决的问题。这当然重要,是最重要的。但还有像这样的东西:这会增加多少算力成本,会增加多少延迟?如果你想象一个步骤,在 AI 系统生成响应后运行,然后你做很多额外的事情,然后再把响应发送给用户,那很可能行不通。不是绝对不行,但成本很大。还有实现复杂性或组织复杂性。如果你的解决方案涉及在脚手架层面做某事,同时还涉及读取 AI 系统的内部并连接起来,这跨越了那么多不同的团队和抽象层,实现起来会比只在推理时运行一个监控器然后提醒某些团队要难得多。
All right. Next thing after that, especially if you're going to be proposing a solution, the next thing is making sure that the solution is evaluated, at least conceptually, even if not by experiments, on the various metrics that companies care a lot about. Most of the time research will look at whether it actually solved the problem it set out to solve. Definitely important. That is the most important one. But then there are also things like how much cost does this add in terms of compute, how much latency will it add? If you're imagining a step that runs after the AI system has produced a response and then you do lots of additional things before sending the response to the user, that's probably a non-starter. Not obviously, but it's a big cost. There's implementation complexity or organizational complexity. If your solution involves doing something at the scaffold level and also something that involves reading the internals of the AI system and connecting these up together, that spans so many different teams and abstraction layers that it's much harder to implement than something that just happens at inference time and involves running a monitor and then alerting some teams.
是的。除了那个,还有其他容易实现的东西的例子吗?
Yeah. Are there any other examples of things that are easy to implement other than that one?
是的,我认为有几个。例如,数据集可以是一个有用的贡献——尝试将数据集添加到后训练中相当容易。有时你尝试把它加到后训练中,然后由于某种原因,它在某个完全不相关的事情上使模型变差,然后你就不能用那个数据集了,这有点不幸。所以理想情况下,当你构建数据集时,你会想评估它是否会对其他东西产生负面影响,但这有点难做。但如果有人带着一个数据集来找我,说‘在这个数据集上训练使这个指标提高了不少,而且似乎没有对可能受影响的另外三个明显指标产生任何不良影响’,我会觉得这相当有说服力,并认为也许我们应该做,假设它解决了一个实际问题。所以我认为数据集、指标、评估——这些通常是相当容易做的事情。监控器——我认为所有这些,如果我们认为值得实施,都是相当容易实现的。
Yeah, I think there are several. For example, datasets can be a useful contribution—it's pretty easy to try to add a dataset to post-training. Sometimes you try to add it to post-training and then for some reason it makes the model worse on some totally unrelated thing, and then you can't use that dataset, which is unfortunate. So ideally when you're building datasets, you would like to evaluate whether it could have some negative effect on something else, but this is a bit hard to do. But if someone came to me with a dataset and said, 'Training on this dataset improves this metric by a decent amount and doesn't seem to have any bad effects on these three obvious other metrics that it might have had bad effects on,' I would find that fairly compelling and think maybe we should do it, assuming it was solving a real problem. So I think datasets, metrics, evals—these are usually fairly easy things to do. Monitors—I think all of those are fairly easy to implement if we think it's worth implementing.
还有其他建议吗?
Any other advice?
还有一个:不要做学术圈那种用花哨方法解决问题的事。相反,用你能想到的最简单的方法来解决它。
One more: just don't do the academic thing of using a fancy method to solve your problem. Instead, solve it with the absolute simplest method you can.
你是说前沿的东西可能在其他方面有用,推动科学进步,将来可能有价值。但如果你希望人们很快在任何实际的商业模型上使用它,那么它必须尽可能简单且经过充分验证。
You're saying cutting-edge stuff might be useful in some other way, advancing the science and could be valuable down the line. But if you want people to use it on any actually commercial models anytime soon, then it has to be as simple as possible and well established as possible.
这是宽容的版本。愤世嫉俗的版本更像是学术界奖励复杂、看起来正式的东西和调参不佳的基线。如果知道基线调得不好,它不会奖励,但很难判断一个基线是否调得不好。
That's the generous version. The cynical version is more like academia rewards complex formal looking stuff and poorly tuned baselines. It won't reward poorly tuned baselines if it knows they're poorly tuned, but it's kind of hard to tell if a baseline has been poorly tuned or not.
那么,什么是调得不好的基线?
And so often what's a poorly tuned baseline?
调得不好的基线是指你有一个默认的解决问题的方法,你试图做得比它更好。
A poorly tuned baseline is like you have some default way of solving the problem that you are trying to do better than.
你是不是故意把基线做得不好,让它看起来很差?
Do you make the baseline look bad by not doing a very good job of it?
没错,但不是故意的。通常是你实现了一次基线,然后没有调超参数,所以它的表现比应有的差。就是很容易花不够时间让基线工作好。
That's right. But not intentionally. It will often be like you implement the baseline once and then you don't tune the hyperparameters for it, so it performs less well than it really should. It's just very easy to not spend enough time working with your baseline to make it work well.
然后让你的其他方法看起来更好。
And then it makes your other thing look better.
正是。考虑到学术界鼓励发表新颖的东西和大量发表,我认为这是一个相当普遍的问题。所以如果你想发表,你往往会这样做。如果你想让你的研究被公司之类的人使用,那么尝试那些显而易见的方法并相当努力地尝试它们就很重要,只有当这些真的失败了,你才尝试更花哨的方法。
Exactly. And since given the incentives in academia of publishing novel things and publishing a lot, I think this is a fairly common problem. So I think if you want to publish, that tends to be the thing that you do. If you want your research to be used by people like companies, it's quite important that you try the obvious stuff and you try fairly hard with the obvious stuff, and only if that really does fail do you try to do something fancier.
到目前为止,哪些外部研究对 GDM 最有用?
What external research has been most useful to GDM so far?
我认为最值得指出的是 Redwood Research 的 AI 控制工作。我们一直计划监控我们的 AI 系统,监控的想法对我们来说并不新鲜,但他们带来的具体概念框架,关于如何评估监控效果,特别是可信模型与不可信模型的区分,我认为非常好。他们发表的第一篇论文展示了各种不同的控制协议,如何评估它们的安全性和有用性,并据此决定应该采用哪一种。这对我如何思考控制工作产生了很大影响。这可能是最明显的例子。
I think the one I would point to most is the AI control work from Redwood Research. I think we were always planning to monitor our AI systems. It's not like the idea of monitoring was new to us, but the specific conceptual frameworks they brought to how you might evaluate how well this works. Specifically the distinction between trusted and untrusted models. I think that was quite good. The first paper they published on it showed how a variety of different control protocols, how you can evaluate their safety and their usefulness and use this to decide which one you should be doing. I think that's influenced me quite a lot on how exactly I think about the control work. I think that's probably the most obvious example.
我认为今年早些时候 Buck Shlegeris 和 Ryan Greenblatt 的节目,特别是 Buck Shlegeris 谈论 AI 控制议程。我觉得那期节目当时没有引起我希望或预期的那么大的反响,但行业内的人不断回头引用它。我认为这意味着如果人们当时没听,可能是因为听起来不够激动人心,我觉得可以再看一看,回去听一听,因为我认为这是非常重要的基础工作,而且只会越来越相关。
I think the Buck Shlegeris and Ryan Greenblatt episodes earlier in the year, Buck Shlegeris in particular talking about the AI control agenda. I feel like that episode didn't make quite as much of a splash as I was hoping or expecting at the time, but the amount that people in the industry keep constantly referring back to it. I think it means that if people didn't listen to it at the time because it just didn't sound quite exciting enough, I think give it another look, maybe go back and have a listen, because I think it is very important foundational work that is only getting more relevant.
是的,当然。我认为控制在中长期内,可能是我们论证 Gemini 安全的方式,至少对于不对齐和内部部署,可能通过控制风格的论证。所以绝对是一个需要了解的重要领域。其他论文:有一篇关于观察性缩放定律的论文,我认为主要可以看作是之前讨论过的罗塞塔石碑论文的前身,我们通过拼接基准来测量整体 AI 进展。观察性缩放定律更多地考虑如何预测特定基准结果,这也是一个很好的应用。所以这基本上是我们之前没有真正思考过的领域,他们展示了可以在这里做一些事情。我认为这在我思考基准拼接/罗塞塔石碑方法时可能影响了我。所以我认为那篇论文对我影响很大。其他例子是各种思维链监控论文。我不认为有哪一篇特别重要,但总体而言,它们建立了一个相当好的研究文献,关于思维链监控的有效性,何时成功何时失败。有一些论文研究了你能在多大程度上说服监控器一切正常,不需要在思维链监控中标记任何东西。我认为这是一个有用的研究,其中的具体定量结果对我有启发。
Yeah, definitely. I would say that control is in the medium term, probably how we are going to argue that Gemini is safe, at least for misalignment and internal deployments, probably via a control style argument. So definitely an important area to be aware of. Other papers: there was this paper on observational scaling laws, which mostly I think you can think of as a precursor to the Rosetta Stone paper that we talked about before, where we stitch together benchmarks to get a way to measure overall AI progress. I think the observational scaling laws thought of this more as how you can forecast particular benchmark results, that's also a pretty good application. So mostly it was an area we hadn't really been thinking about, and they showed that you can do something here. I think that probably influenced me when thinking about the benchmark stitching / Rosetta Stone approach. So I think that one was pretty influential on me. Other examples would be the various chain of thought monitoring papers that have come out. I don't think there is any one paper that I thought was a particularly big deal, but in aggregate I think they've built a fairly good research literature on the effectiveness of chain of thought monitoring, when it does and doesn't fail. There were some papers that looked at to what extent could you just persuade the monitor that everything's fine here, you don't need to flag anything in this chain of thought monitoring. I think that was a useful thing to study, and the specific quantitative results in it were informative to me.
所以我想很多人会想在 GDM 工作,甚至其他 AI 公司的人可能也会考虑跳槽,因为 GDM 近年来在模型上取得了进展。你觉得哪些技能或角色最难招到,哪些技能在劳动力市场上最难找到?
So I imagine a lot of people would love to get a job at GDM, and even people at other AI companies might be interested in considering switching given the advances that GDM has made in its models in recent years. What sorts of skills or roles do you find hardest to fill, and what sorts of skills are maybe hardest to find in the labor market?
我想再次回到我一直在说的主题:挑战更多在于实施,而不是弄清楚我们需要做什么。我们有很多想法,我们知道需要做什么,但实施起来更难,因为有很多东西需要检查,确保不会损害它。因此,我认为特别是在过去一两年,我们更需要那些只想做显而易见的事情并把它落地的人,而不是那些想找出理想最优方案、写一篇酷炫研究论文的人。
I think I'll go back again to the theme I've had throughout of the challenge being more in implementing stuff rather than in figuring out what stuff we need to do. We have a lot of ideas, we know what we need to do, implementing it ends up being harder because there's so much stuff you have to check and make sure that you're not hurting it. As a result, I think especially over the last year, maybe two years, we've had more of a need for people who just want to do the obvious thing and land it, as opposed to people who want to figure out the ideal optimal thing, write a cool research paper about it.
因为我们说的是 AGI 安全和对齐团队的人,对吧?
Because we're talking about the AGI safety and alignment folks, right?
没错。我主要说的是 AGI 安全和对齐团队。所以是的,我认为重点是做显而易见的事情,注重实施而不是研究。我认为在实践中,这也意味着相对于机器学习研究,更注重软件工程。我仍然认为我们关心概念能力、研究品味、机器学习工程技能。在做这种实施时,它们确实会出现。你需要测试你的系统有多好。这意味着你需要构建好的评估。你不能过拟合它们。你需要小心一点。所以并不是说这些技能无关紧要。我只是认为现在相对于一年前,尤其是相对于两年前,更注重把事情做好、软件工程、实施这类事情。
That's right. Yes. I am talking primarily about the AGI safety and alignment team. So yeah, I think a focus on doing just the obvious stuff, a focus on implementation rather than research. I think in practice this also means more of a focus on software engineering relative to machine learning research. I do still think that we care about conceptual ability, research taste, ML engineering skills. They do come up when you're doing this sort of implementation. You need to test how good your system is. That means you need to build good evaluations. You need to not overfit to them. You need to be a little bit careful about that. So it's not like those skills are irrelevant. I just think that there's more of a focus on things like getting things done, software engineering, implementation now relative to even just a year ago, but especially relative to two years ago.
你们目前有没有在招聘特定职位,或者整体上在大量招聘?因为我觉得 Anthropic 招人很猛,OpenAI 也总是有很多职位。你们情况类似吗?
Are there any particular roles that you're hiring for at the moment or hiring for a lot in general? Because it's felt to me like Anthropic is hiring hand over fist and I guess OpenAI there's like always a lot of roles there. Is it similar situation for you?
不,我觉得我们目前招聘的规模比 Anthropic 和 OpenAI 小得多。我们现在确实有几个职位开放。部分职位在其他团队,但也非常相关,比如 AGI 安全与对齐团队并不是 Google DeepMind 唯一与 AGI 安全相关的团队。例如,目前安全团队有一轮招聘,招聘工程师从事 AI 控制方面的工作。我认为这是一个很好的职位,影响巨大。正如我之前所说,这可能是我们论证 Gemini 在中期内不会因失控而造成危害的关键。所以我认为这是一个极其有用的职位,影响很大,虽然不在 AGI 安全与对齐团队,但他们会与我们紧密合作。在我的团队,我们目前正在招聘一名前沿安全风险评估工程师。这包括运行我们的危险能力评估并撰写前沿安全报告。所以如果你觉得前沿安全报告需要更详细,你可以加入我们,投入时间让它变得更好。团队中的个人有相当大的灵活性去做这类事情。实际上,我预计在这期节目发布时,那个职位可能已经没有了,但我们确实会持续推出这类职位。也许我会发一个意向表达表的链接,人们可以填写,这样当新职位开放时我们可以发邮件通知他们。部分原因是,我们和其他许多公司一样,受到了 AI 辅助申请洪流的冲击。所以我们未来可能不太会公开大规模招聘,因为筛选所有简历真的很麻烦。
No, I think we are hiring quite a bit less than Anthropic and OpenAI at the moment. We do have a couple of roles open right now. Partly we have roles open in other teams that are also very relevant, like the AGI safety and alignment team isn't the only team relevant to AGI safety at Google DeepMind. For example, currently there is a hiring round open on the security team to hire engineers to work on AI control. I think that's a great position, hugely impactful. As I've said before, that's probably the argument we're going to make for why Gemini wouldn't cause harm via loss of control in the medium term. So I think that's an extremely useful role, very impactful, not on the AGI safety and alignment team, but they would be working closely with us. On my team, we're hiring at the moment for an engineer for frontier safety risk assessment. This involves running our dangerous capability evaluations and writing the frontier safety report. So if you think the frontier safety report needs to be more detailed, you could come join us and put in the time to make it better. There is a fair amount of flexibility for individual people on the team to go and do stuff like that. In practice, I expect by the time this recording comes out, that role might not be there, but we do have these roles coming out on an ongoing basis. Maybe I'll send over a link to an expression of interest form that people can fill out so we can email them when new roles open up. Part of it is that we, like many others, have been hit by the deluge of AI-assisted applications. So we are a little less likely in the future to make these open calls for hiring because it really is such a pain to go through resume screening for all of them.
是啊,我觉得人们可能得引入收费来填写这类表格或申请工作,虽然大家出于其他原因也会讨厌这个。但我想不出别的办法,因为现在 AI 可以提交完全无法区分的申请,可以随意造假。
Yeah, I think folks are going to have to introduce a fee to fill out forms like that or to apply for jobs, which people also hate for other reasons. But I don't really see the alternative now that it's possible for AIs to submit completely indistinguishable applications that can be as fake as you like.
是啊,很难。去年我们加了一个 LLM 验证,但最近一次我们想这么做,我觉得可能可行,但也会骗到一些人类。就连我们去年的 LLM 验证也骗到了一些人。不是很多,大概 5%。现在我不确定我能设计出一个不会产生大量误报的验证。
Yeah, it's rough. Last year we included an LLM capture, but for a recent one, we were trying to do it and I think we probably could, but it would also trick a bunch of humans. Even our LLM capture from last year tricked a bunch of humans. Not a bunch, maybe five percent of them. At this point, I'm not sure that I can design one that wouldn't have quite a lot of false positives.
嗯,我想对我来说订个航班可能是个办法,但那可能比收费还贵。
Well, I guess booking a flight for me would be, but that might be more expensive than just having the fee.
是啊。
Yeah. Yeah.
你有什么特别想说的,关于在 GDM 工作的理由吗?
Is there any pitch you want to make for working at GDM in particular?
基本上我的理由是:公司安全团队很可能是从技术层面让 AI 系统安全的最大推动力。这是加入他们的一个好理由。
So basically my pitch is: company safety teams are probably the biggest force for what actually happens on a technical level to make AI systems safe. And that is a good reason to join them.
好了,我们该结束了。最后一个问题,我要问一个来自观众的难题。Rohin,在这样艰难的时代,你是如何保持如此积极乐观的精神的?
All right. We should wrap up. I think for a final question I'll throw you a real hard ball that came in from the audience. How do you maintain such a nice and positive spirit in these troubled times, Rohin?
嗯,部分答案有点无聊,不太具有普适性。第一是性格:我相当稳定,情绪每天变化不大。第二,正如这期节目可能已经表明的,我觉得时代并没有其他人想的那么糟糕。但我确实经常做一件事,而且很虔诚:专注于我能控制的事情。我认为这在艰难时期对保持积极乐观很有用。它确实帮助我专注于我有掌控力的领域。而且我觉得专注于你能真正有所作为的领域是很棒的。
Yeah. I mean, some of the answer is a bit boring, not that generalizable. One is just by personality: I'm fairly stable and my mood doesn't change very much day to day. And two, as has maybe become clear over the course of this episode, I don't think the times are as troubled as everybody else thinks, relative to everyone else. But I do think there is one thing that I do quite a lot and quite religiously: focus on the things that I can control. I think this is quite useful for maintaining a nice and positive spirit during these troubled times. It definitely helps keep me focused on the areas where I do have agency. And I think it's pretty great to be focused on areas where you can actually make a difference.
今天的嘉宾是 Rohin Shah。非常感谢你来到 80,000 Hours 播客。
My guest today has been Rohin Shah. Thanks so much for coming on the 80,000 Hours podcast.
非常感谢,Rob。很高兴来到这里。
Thanks a lot, Rob. It was great to be here.
大家好,我是 Rob,有个个人消息。至少如果 80,000 Hours 算个人的话。我们出了一本新书,叫《80,000 Hours:如何拥有一个有益且充实的职业生涯》。企鹅出版社刚在全球出版。很多听众可能不知道,80,000 Hours 从 2012 年就开始运营,旨在帮助人们通过职业生涯产生更大的社会影响,同时过上美好而充实的生活。这本书是我们自那时以来提出并打磨的许多想法的最完整表达,由我们的创始人 Benjamin Todd 撰写。它涵盖了职业和社会影响话题的各个方面,包括理想工作的证据、个人能产生多大影响以及如何知道、如何正确选择要解决的全球问题、哪些具体技能能让你帮助最多人、你需要多少收入才能幸福、如何变得不可或缺、如何平衡跟随直觉和忽视直觉的风险、何时在职业上追求更高目标、何时安于现状、如何找到适合你独特优缺点的职业,以及一旦确定方向如何最好地进入该领域并得到你想要的工作。我们最新的版本自然特别关注 AI,它正在颠覆许多人的职业规划,迫使很多人重新思考他们期望做什么。如果你自己担心 AI,关于未来哪些技能最有价值的章节涵盖了许多反直觉的观点,我认为仅凭这些就值得一读。我简要列出六个。第一,Ben 认为,即使你认为 AI 很快会取代人类所有工作,你也有理由预期短期内工资实际上会暴涨。
Hey everyone, Rob here with some personal news. At least if 80,000 Hours counts as a person. We have a new book out. It's called 80,000 Hours: How to Have a Fulfilling Career That Does Good. And Penguin just published it around the world. A lot of listeners don't actually know this, but 80,000 Hours has been running since 2012, trying to help people have a much larger social impact with their career while doing something that also delivers them hopefully a wonderful and well-rounded life. And this is the single most developed expression of the many ideas that we've been coming up with and honing since then, written up by our founder, Benjamin Todd. It covers all kinds of angles on the career and social impact topics, including the evidence on what makes for an ideal job, how much good an individual can actually do and how we could know that. How to correctly choose which global problem to focus on solving, which specific skills enable you to help the most people or help people the most. How much income you need to be happy. How to become so valuable they can't ignore you. How to balance the risks of going with your gut instinct against the risks of ignoring your gut instinct. When to aim higher in your career versus when to decide to settle instead. How to find a career that fits your idiosyncratic strengths and weaknesses in particular and how to best break into a field and actually get the job that you want once you've decided on a direction. Our latest editions naturally have a particular focus on AI, which is upending lots of people's career plans and forcing a lot of people to rethink what they're expecting to do. If you're worried about AI yourself, the chapter covering which skills will be most valuable in the future goes through a lot of counterintuitive ideas which I think might justify picking up the book just by themselves. I'll list six of them briefly here. First, Ben argues that even if you think AI is going to take all of the jobs from human beings before long, you might also reasonably expect salaries to actually explode in the short run.
他介绍的一个备受尊敬的经济模型表明,平均工资在盛宴结束前可能上涨 10 倍,创造出可能是工人有史以来最好的短暂窗口期。
One well-respected economic model that he goes through suggests that average wages could rise 10-fold before the party ends, creating what might be the best brief window ever to be a worker.
第二,本解释了为什么转行做水管工或其他体力劳动者——这个你在应对 AI 时越来越多听到、听起来也合情合理的建议,包括对我而言——他认为这对本节目的大多数听众来说可能没什么用。
Second, Ben explains why retraining as a plumber or other physical laborer, a piece of advice that you hear more and more in reaction to AI and which sounds sort of common sense, including to me, he explains why he thinks that's probably not productive for most listeners to this show.
第三,本解释了为什么仅仅拥有一份很难被 AI 自动化的工作是不够的。你还需要找到一份与 AI 产出高度互补的工作。
Third, Ben explains why it's not sufficient to just have a job that's very hard to automate with AI. You need to find a job that's highly complementary with AI outputs as well.
第四,本书探讨了为什么当前 AI 浪潮中最受冲击的工作是年薪 10 万到 20 万美元的人群,这与之前的自动化趋势相比是一个相当大的逆转。
Fourth, the book investigates why the jobs most exposed to the current wave of AI are people earning $100,000 to $200,000 a year, which is a pretty significant reversal of previous automation trends.
第五,他深入探讨了为什么杰弗里·辛顿 2016 年那个非常著名的预测——医院应该停止培训放射科医生——结果完全错误,9 年后放射科就业率创下历史新高。而另一方面,2013 年另一个极具影响力的预测——创意工作将抵御自动化——也完全错了,但方向相反。
Fifth, he goes into why Geoffrey Hinton's very famous 2016 prediction that hospitals should stop training radiologists turned out to be completely wrong and 9 years later, radiology employment is hitting record levels. While on the other hand, another hugely influential prediction from 2013 that creative work would resist automation has also turned out to be totally wrong, but in the reverse direction.
第六,为什么诱导需求不对称表明,提供人们几乎无限需求的东西(比如更好的健康)的人会做得很好。而提供我们需求量相对固定(比如税务合规)的人,无论 AI 可能让他们变得多高效,都会陷入困境。
And sixth, why induced demand asymmetries suggest that people providing things that people want in almost unlimited amounts, like better health, are going to do well. While people delivering things that we want in more of a fixed amount, like tax compliance, are going to struggle no matter how productive AI potentially makes them.
当然,没有人真正知道 5 年后哪些技能会最有价值。所以认为自己知道是疯狂的。但这一章给了你很多工具,避免犯那些事后看来非常可预测的错误。它还让你理解潜在的经济效应和趋势,这样当变化开始发生时,你就能理解并试图利用它们,而不是被碾压。
Of course, nobody actually knows what skills are going to be most valuable in 5 years. So it would be crazy to think that you did. But the chapter gives you a lot of tools to avoid making mistakes that would look really predictable in retrospect. And it also gives you an understanding of the underlying economic effects and trends such that you can make sense of shifts as they start to happen and try to take advantage of them rather than just get steamrolled.
而且本并没有回避给出具体的结论。他列出了一份具体的领域清单,认为你不应该不经认真思考就进入这些领域。
And Ben doesn't chicken out from having actual concrete conclusions as well. He has a specific list of fields he thinks you shouldn't train in without serious thought.
好了,这只是 15 章书中第 8 章的一点内容。正如我所说,这是我们过去 14 年逐渐积累的智慧,所以信息密度很高。如果你觉得不错,可以在网上搜索“80,000 Hours book”订购这本书。或者访问 80000hours.org/book。当然在亚马逊、Audible 以及任何你买书的地方都有。本杰明亲自录制了有声版。内容相当丰富,所以我觉得他比我更合适。如果你已经对我们的建议了如指掌,现在正是告诉朋友的好时机。本周的任何销售都将帮助我们登上畅销书榜单,这显然非常重要。
Okay, that's a taste of what you'd get just from chapter 8 of a 15 chapter book. As I said, this is the wisdom we've been gradually compiling for the last 14 years, so it's pretty information dense. If that sounds good to you, you can order the book online by searching for '80,000 Hours book'. Or go to 80000hours.org/book. It's obviously on Amazon, Audible, wherever you get books where you are. Benjamin actually narrated the audio version himself. It's pretty extensive, so better him than me, I think. And if you already know our advice inside and out, this is an ideal time to tell a friend about it. Any sales this week will help us hit the bestseller lists, which are apparently super important.