OpenAI’s push to make superintelligence safe
打开互动全文版(中英对照 + 朗读 + 问答)→为「超级对齐」与「可扩展监督」辩护。
The case for superalignment and scalable oversight.
今天我的嘉宾是扬·莱克。自 2021 年以来,扬一直担任 OpenAI 的对齐负责人,他将与 OpenAI 创始人伊利亚·苏茨克弗共同领导新的超级对齐项目。多年前,扬在著名机器学习人物马库斯·胡特的指导下完成了博士学位,随后在牛津大学人类未来研究所做了短暂的博士后,之后成为 DeepMind 的研究科学家——他上次在 2018 年参加本节目第 23 期时正是这个身份,那期的主题是《如何真正成为一名 AI 对齐研究员》。感谢你再次来到播客,扬。
Today I'm speaking with Jan Leike. Since 2021, Jan has been head of alignment at OpenAI, and along with OpenAI founder Ilya Sutskever, he is going to be co-leading their new superalignment project. Years ago, Jan did his PhD with the well-known machine learning figure Marcus Hutter as a supervisor, and he then did a brief postdoc at the Future of Humanity Institute at Oxford before becoming a research scientist at DeepMind, which is what he was doing when he last came on the show in 2018 for episode number 23, 'How to Actually Become an AI Alignment Researcher' according to Jan Leike. Thanks for coming back on the podcast, Jan.
非常感谢再次邀请我。能来这里真是太好了。你自那以后确实发展得很好。我觉得我们有时挺擅长发掘人才的,挑选那些职业生涯会腾飞的人。我希望聊聊超级对齐项目以及你们想为此招聘什么样的人,同时也会向你提出很多听众的问题。那我们就直接开始吧。
Thanks a lot for having me again. It's really great to be here. You've really gone places since then. I feel like we've been sometimes pretty good at picking talent, picking people with careers that are going to take off. I hope to talk about the superalignment project and who you're trying to hire for that, as well as put a lot of audience questions to you. So let's dive right in.
为了省你点力气,我来读一下 OpenAI 两周前发布的关于超级对齐项目的公告摘录。引用原文:'超级智能将是人类有史以来最具影响力的技术,可以帮助我们解决世界上许多最紧迫的问题。但超级智能的巨大力量也可能非常危险,可能导致人类失去权力甚至人类灭绝。虽然超级智能现在看起来还很遥远,但我们相信它可能在本十年内到来。我们需要科学和技术的突破来引导和控制比我们聪明得多的 AI 系统。为了在四年内解决这个问题,我们正在组建一个新团队,由伊利亚·苏茨克弗和扬·莱克共同领导,并将迄今为止已获得的算力的 20%投入这项工作。我们正在寻找优秀的机器学习研究人员和工程师加入我们。' 好的,对于那些不太关注或完全不了解的听众,你能介绍一下这个项目的更多细节吗?
To save you a little bit of effort, I'll read an extract from the announcement that OpenAI put out about the superalignment project two weeks ago. Quote: 'Superintelligence will be the most impactful technology humanity has ever invented, and could help us solve many of the world's most pressing problems. But the vast power of superintelligence could also be very dangerous, and could lead to the disempowerment of humanity or even human extinction. While superintelligence seems far off now, we believe it could arrive this decade. We need scientific and technical breakthroughs to steer and control AI systems much smarter than us. To solve this problem within four years, we're starting a new team co-led by Ilya Sutskever and Jan Leike, and dedicating 20% of the compute we've secured to date to this effort. We're looking for excellent machine learning researchers and engineers to join us.' Okay, so for listeners who haven't been following this much or possibly at all, can you fill us in on some more details of the project?
非常乐意。基本上,如果你看看我们今天是如何对齐大型语言模型的,那就是使用基于人类反馈的强化学习(RLHF),这是一种技术:你向人类展示一堆样本,问他们更喜欢哪一个(比如对话助手之类的),然后这就成了 ChatGPT 或其他类似 AI 系统的训练信号。我们从根本上认为 RLHF 无法扩展。原因很简单:你让人类监督 AI 系统,假设他们能判断哪个回答更好,或者从根本上理解系统在做什么。这在今天大部分情况下是成立的,因为 ChatGPT 执行的任务并不那么复杂。但随着 AI 系统变得更聪明,它们将能够做更难的事情,那些我们理解得少得多的事情。人类能够评估系统在做什么这一基本假设将不再成立。因此,要引导和控制比我们聪明得多的系统,我们需要新技术。
Very happy to. Basically, if you look at how we are aligning large language models today, it's using reinforcement learning from human feedback (RLHF), which is a technique where you show a bunch of samples to a human and ask them which one they prefer for a dialogue assistant or something, and then that becomes a training signal for ChatGPT or other AI systems like it. We fundamentally don't think that RLHF will scale. The reason is very simple: you have humans overseeing AI systems, assuming they can tell which response is better or fundamentally understand what the system is doing. This is definitely true today, for the most part, because the tasks that ChatGPT is doing aren't that complex. But as AI systems get smarter, they'll be able to do harder things, things we understand much less. The fundamental assumption that humans can evaluate what the system is doing will no longer be true. So to steer and control systems that are much smarter than us, we will need new techniques.
所以当前的方法是观察输出,然后评价它的好坏,这提供了反馈,帮助将模型推向正确的方向。但未来,我们可能无法评估模型在深层是否真的在做我们想让它做的事。所以我们必须有某种其他方式来引导它走向正确的方向。这是简略版吗?
So the current method is to observe the output and then rate how good it has been, and that provides feedback that helps push the model in the right direction. But in the future, we just might not be able to evaluate whether the model actually, at a deep level, is doing what we want it to do. So we're going to have to have some other way of nudging it in the right direction. Is that the short version?
没错。所以问题是,如果你有一个非常聪明的系统,它可能想出各种方法来巧妙地颠覆我们,或者试图欺骗我们、对我们撒谎,而我们很难察觉。这里有很多非常重要且有趣的研究挑战:我们能否理解如何提取模型对某些问题的认知?如果它写了一段代码,我能否分辨出哪些部分是我理解的?它是否知道代码中存在某些错误?或者我们能否理解系统如何从我们可以监督的简单问题泛化到我们无法监督的困难问题?或者我们能否理解如何使其稳健,使其无法被越狱或颠覆监控系统?
That's right. So the problem is, if you have a system that is really smart, it could think of all kinds of ways to subtly subvert us, or try to deceive us or lie to us in a way that is really difficult for us to check. There are a lot of really important and interesting research challenges here: can we understand how to extract what the model knows about certain problems? If it writes a piece of code, can I tell which parts of the code I understand? Does it know there are certain bugs in the code? Or can we understand how the system can generalize from easy problems that we can supervise to harder ones we can't? Or can we understand how to make it robust so that it can't get jailbroken or subvert the monitoring systems?
所以这将与 OpenAI 一直在做的工作不同。这将专注于与人类一样聪明或更聪明的模型,做非常复杂的事情,以至于它们足够老练,可能欺骗我们,或者可能出现其他故障。OpenAI 迄今为止一直在做的所有对齐和安全工作会怎样?是会由另一个团队继续吗?
So this is going to be distinct from what OpenAI has already been doing. This is going to focus on models that are as smart as humans or smarter, doing things that are quite complicated, such that they're sophisticated enough to potentially trick us, or there could be other failures. What's going to happen to all the alignment and safety work that OpenAI has already been doing up until now? Is that just going to continue with a different team?
我认为有很多非常重要的工作要做,以确保我们已有的当前系统是安全的,持续安全,并且不会被滥用。OpenAI 内部有很多事情正在发生,我对此感到非常兴奋,我也非常希望更多人加入。这包括修复越狱问题,寻找自动监控滥用的方法,以及诸如此类的问题。所有这些工作都必须继续,并且是在 ChatGPT 产品的背景下进行的。但我们着手要做的是解决我们尚未真正遇到的对齐问题。例如,如果你让 GPT-4 帮你写一段代码,它不会写非常复杂的代码,不会写整个复杂的代码库,而且它通常不够聪明,无法在代码中植入我们无法发现的木马。但未来的模型可能会这样做。我们的工作从根本上是要区分两种不同的 AI 系统:一种是真的想帮助我们,真的想按照人类意图行事,真的想做我们想让它做的事;另一种是当我们看着它时假装想要所有这些,但当我们不看时,它完全做别的事。问题是,这两种系统在你看着它们时看起来完全一样。
I think there's a lot of really important work to be done to ensure that the current systems we already have are safe, continue to be safe, and won't be misused. There's a lot of stuff happening around us at OpenAI that I'm really excited about, and I would be really excited for more people to join. This involves fixing jailbreaking, finding ways to automatically monitor for abuse, and questions like that. All that work has to continue, and it happens in the context of the ChatGPT product. But what we are setting out to do is to solve alignment problems that we don't really have yet. For example, if you have GPT-4 help you write a piece of code, it doesn't write really complicated pieces of code, it doesn't write entire complicated code bases, and it's not generally smart enough to put a Trojan into the code that we wouldn't be able to spot. But future models might do that. Our job is fundamentally trying to distinguish between two different AI systems: one that truly wants to help us, truly wants to act in accordance with human intent, truly wants to do the things we want it to do; and the other one that just pretends to want all these things when we're looking, but then if we're not looking, it does something else entirely. The problem is that both of these systems look exactly the same when you're looking.
对我们来说,这是个尴尬的事实,没错。所以这成了一个有趣的挑战。但我们有很多优势,对吧?比如我们可以查看模型内部,可以对其进行各种不同的测试,可以修改它的内部结构,可以清除系统的记忆,可以检查它在其他情况下说的话是否一致。所以至少,你可以确保它是一个非常自洽的撒谎者。但我们真正想做的远不止于此,对吧?我们必须解决如何知道它是否真正对齐的挑战。
Awkward fact for us, yeah, that's right. So that makes it an interesting challenge. But we have a bunch of advantages, right? Like we can look inside the model, we can subject the model to all kinds of different tests, we can modify its internals, we can erase the system's memory, we can see if it's consistent with other things it's saying in other situations. And so at the very least, you can make sure that it is a very coherent liar with itself. But we really want to do more than that, right? We have to solve the challenge of how do we know it is truly aligned.
要写一本好的科幻小说,我想从 AI 的角度来想象这个场景:它比训练它的人聪明得多,但另一方面,他们可以查看它的大脑,给它做各种有趣的测试,试图检查它是否在欺骗他们。作为一个智能体,你会想出什么样的策略来绕过这一点?这就像一场真正的猫鼠游戏。
To be a great science fiction book, I think to imagine this scenario from the perspective of the AI, where it's much smarter than the people who are training it, but on the other hand they can look inside its brain and give it all of these funny tests in order to try to check whether it's deceiving them. What kind of strategies would you come up with as an agent to work around that? It's like a real cat and mouse game.
是的,所以我认为这里重要的是,我们不是在想象极其强大的系统。我们不会想象比我们聪明得多的系统。它们可能在某些方面比我们强,例如,GPT-4 在记忆事实方面要好得多,或者它能说比任何人类都多的语言,但它在某些方面也差得多,对吧?比如它不能正确做算术,这想想还挺尴尬的。
Yeah, so I think it's important here that we're not picturing vastly powerful systems. We're not going to picture systems that are vastly smarter than us. They might be better than us in some ways, for example, GPT-4 is much better at remembering facts or it can speak more languages than any human, but it's also much worse in some ways, right? Like it can't do arithmetic right, which is kind of embarrassing if you think about it.
嗯,我是说,我一次记不住超过七个数字,所以我觉得我们都有自己的局限性,对吧?
Well, I mean, I can't remember more than seven numbers at a time, so I feel like we have our own limitations, right?
是的,但我认为我们真正想要达到的目标是,我们希望能够对齐一个与最聪明的人类对齐研究人员大致一样聪明的系统。
Yeah, but I think the goal that we really want to aim for is we want to be able to align a system that is roughly as smart as the smartest humans who are doing alignment research.
好的,让我们稍微深入探讨一下这个问题,这有点像:不仅要对齐这样一个系统,还要确信它已经充分对齐,需要什么条件?我认为一个有用的思考方式是,你想把你的方法分成两大类:你有一堆训练方法,训练系统变得更对齐;然后你有验证方法,校准你对系统实际对齐程度的信心。像往常一样,当你在机器学习中进行这种训练-验证分割时,你想确保验证集没有泄露到训练集中,对吧?
Okay, let's zoom into that question a little bit, which is kind of like this question of what would it take to not only align a system like that but also to be confident that it is sufficiently aligned. I think one useful way to think about it is you want to split your methods into two kind of general buckets: you have a bunch of training methods that train the system to be more aligned, and then you have validation methods that calibrate your confidence about how aligned the system actually is. And as usual, when you do this train-validation split in machine learning, you want to know that there is no leakage of the validation set into the training set, right?
你能……我想问题在于,如果你用检查是否成功的东西来训练模型,那么它当然可以变得非常擅长通过那个测试,尽管它在更广泛的意义上并没有对齐。它只是被操纵了,使得它的失败不会被你的测试发现。所以你需要把用于获取反馈来训练模型的东西,与用于验证训练是否成功的东西完全分开。这是基本想法吗?
Can you... I guess the problem would be if you're training the model on the same thing that you're using to check whether you've succeeded or not, then of course it could just become extremely good at doing that test even though it's not aligned in the broader sense. It's just kind of gerrymandered to not have its failures picked up by your test. So you need to have the things used to get feedback to train the model fully separated from the things you use to validate whether that training has succeeded. Is that the basic idea?
没错。你不想在测试集上训练,对吧?那会让通过测试变得太容易。我们稍后会回到一些细节,但首先我有几个观众问题想问你。我们收到了很多听众的提问。一位听众特别想听你的澄清。一位听众问,为什么目标是在四年内解决这个问题?这大致是你预计 AGI 到来的时间吗?
That's right. You don't want to train on the test, you know? That makes passing the test so easy. We'll come back to some of those details in a minute, but first I had a couple of audience questions to put to you. We've got a lot of submissions from listeners. A particular listener wanted to hear clarifications from you. One listener asked why the target for solving this problem in four years. Is that roughly when you expect AGI to arrive?
好问题。所以我认为总的来说,我对未来会如何发展有很多不确定性,而且我认为没有人真正知道。但我认为我们很多人都预期事情可能会发展得相当快,系统在未来几年可能会变得更聪明或更有能力。我们真的很想领先于这一点。我们真的很想在实际必须解决这个问题之前就解决它,或者至少理想情况下提前很多。所以这个四年目标被选为一种折中方案……我们不知道我们会有多少时间,但我们想设定一个雄心勃勃的截止日期,我们仍然认为我们实际上可以达成。
Great question. So I think in general I have a lot of uncertainty about how the future is going to go, and I think nobody really knows. But I think a lot of us expect that things could be moving quite quickly and systems could get a lot smarter or a lot more capable over the next few years. And we would really love to be ahead of that. We would really love to have solved this problem in advance of us actually having had to solve it, or at least ideally far in advance. So this four-year goal was picked as kind of a middle ground between... we don't know how much time we'll have, but we want to set an ambitious deadline that we still think we could actually meet.
是的,所以这算是最雄心勃勃的目标,同时也不会让你觉得这么快完成的可能性可笑。四年内可以完成很多事情。好的,另一个关于公告的问题。公告提到 OpenAI 迄今为止已确保的 20%算力,这肯定……我想我不知道 OpenAI 到底有多少算力的所有细节,但我想无论以何种标准衡量,这都是相当大量的计算资源。但一位持怀疑态度的听众希望我深入挖掘 20%算力这个数据。考虑到算力、人员数量和资金,OpenAI 在超级对齐团队上的投资净变化是多少?也许他们在增加对齐方面的投资,但他们是否也在同等或更大程度上增加能力方面的投资?特别是,有人指出这只是迄今为止确保的 20%算力,当然算力每年都在增长,所以相对于未来所有算力的 20%,这可能最终会小得多。你能为我们澄清一下吗?
Yeah, so it's kind of the most ambitious target that doesn't also cause you to laugh at the possibility that it could be done that quickly. A lot of things can be done in four years. Okay, another question about the announcement. The announcement talks about this 20% of compute that OpenAI has secured so far, which is surely going to be... I guess I don't know all the details about exactly how much compute OpenAI has, but I imagine that by any measure is going to be a pretty significant amount of computational resources. But one skeptical listener wanted me to dig deeper on the 20% compute stat. What is OpenAI's net change in investing in alignment with the superalignment team considering compute and headcount and funding? And maybe they're increasing investment in alignment, but are they increasing investment in capabilities as much or more so? In particular, some people have pointed out that this is 20% of compute secured so far, and of course amounts of compute are growing every year, so that might end up being much smaller relative to 20% of all compute in future. Can you clarify this for us?
是的,所以 20%算力确保的数字指的是我们现在能访问的所有算力以及我们已经下了采购订单的所有算力。所以这实际上非常多。就像……我认为技术术语就是很多。
Yeah, so the 20% compute secured number refers to everything we have access to right now and everything we've put purchase orders in for. And so this is actually really a lot. It's like... I think the technical term is a lot.
是的,我的意思是,考虑到你是从零开始组建这个团队,你可能拥有最多的人均算力,或者极高的人均算力,对吧?
Yeah, I mean, given that you're building this team from scratch, you might have about the most compute per person, or an extremely high compute per staff member, right?
但我认为这不是正确的思考方式,因为这是分配给解决问题的算力,不一定是给这个特定团队的。所以一种可能的方式是,我们开发一些方法,然后另一个非常擅长扩展的团队将其扩展,他们实际上会使用很多算力。我认为另一种方式是……我认为正确的思考方式不是看它相对于能力投资的比例,而是看它相对于其他对齐投资的规模。就我们迄今为止的投资而言,我认为这是一个非常显著的提升,不仅仅是 3 倍,而是比那多得多。而且我认为……
But I think this is not the right way to think about it, because it's compute that's allocated to solve the problem, not necessarily for this particular team. So one way this could go is we develop some methods, and then some other team that's really good at scaling stuff up scales it up, and they spend actually a lot of it. I think another way is... I think it's not the correct way to think about this is not like what is it relative to capabilities. I think it's just like what is it relative to other investments in alignment. And in terms of how much we've been investing so far, I think this is a really significant step up, not just like a 3x but like a lot more than that. And I think also...
这表明 OpenAI 实际上非常重视这个问题,并且真正投入资源来解决它。他们本不必做出这样的承诺,对吧?没人强迫他们。
It shows that OpenAI is actually really serious about this problem and really putting resources behind solving it. They wouldn't have to have made that commitment, right? Nobody forced them.
是的,是的。我想如果你用完了这 20%的算力,我认为获得未来几年额外的算力承诺会相对直接。所以如果有一个好的计划来使用更多算力,比如「如果再多这么多算力,我们就能在对齐上做得更好」,我会很有信心。我认为我们可以提出一个非常有力的理由。而且我认为如果问题归结于此,他们会提供更多算力。基本上,我认为这是最好的情况。如果解决这个问题只需要到处要更多 GPU,我认为我们基本上已经赢了,老实说。
Yeah, yeah. I suppose if you use up this 20% compute, I do think it'll be relatively straightforward to get access to additional compute commitments that come out in future years as well. So I'll be pretty confident if we have a good plan how to spend more compute, and we're like, you know, if you have this much more, we could do this much better on alignment or something. I think we can make a really strong case for that. And I think they'll give a lot more compute if that's what it comes down to. Basically, I think that's the best world to be in. If all you need to solve the problem is to go around asking for more GPUs, I think we've mostly won, honestly.
为什么这个项目获得大量算力如此重要?
Why is it so important for this project to have access to a lot of compute?
回答这个问题有很多方式。如果你看过去 10 年深度学习的历史,算力在所有重大突破和成果中都扮演了重要角色。总的来说,有一个通用配方:很多简单的想法如果规模化并投入大量算力,效果会非常好。这对能力来说是这样。我预计在一定程度上,这对对齐也是如此。但并非完全如此,因为我认为我们目前拥有的任何东西都还没有准备好大规模运行,这里有一个真正的研究问题需要解决。但此外,我们真正兴奋的策略,以及我们有比较优势投入的策略,是那些真正规模化并使用大量算力的策略。特别是,如果我们考虑可扩展监督,我们可以投入更多算力来辅助人类评估,这将使评估更好。或者自动可解释性:如果我们有一种方法可以自动在整个网络上运行,我们就可以投入大量算力,在最大的模型上运行它。最终,我们的目标是自动化对齐研究本身,这意味着我们将运行一个虚拟的对齐研究员。一旦我们达到那个阶段,很明显你只需要投入大量算力来运行那个研究员,他们就会非常快速地取得对齐进展。
There are a bunch of ways of answering that question. If you look at the history of deep learning over the last 10 years, basically compute has played a really major role in all the big headline advances and results. In general, there's this general recipe: a lot of simple ideas work really well if you scale them up and use a lot of compute. This has been true for capabilities. I expect to some extent this will be true for alignment as well. It won't be only true, because I don't think anything we currently have is really ready to just be run at scale, and there's a real research problem to be solved here. But also, the strategies we're really excited about, and the strategies that we have a comparative advantage in investing in, are the ones where you really scale up and use a lot of compute. In particular, if we're thinking about scalable oversight, we can spend more compute on assisting human evaluation, and that will make the evaluation better. Or automated interpretability: if we have a method that we can automatically run over a whole network, we can just spend a lot of compute and run it on the biggest model. Ultimately, where we want to go is to automate alignment research itself, which means we would be running a kind of virtual alignment researcher. Once we get to that stage, it's really clear that you just want to spend a lot of compute to run that researcher a lot, and they'll make a lot of alignment progress very quickly.
好的,让我们先退一步,审视一下对齐方法的当前技术水平,以及为什么你相信它们不足以对齐比人类聪明得多的智能体模型。我要补充一点,你之前接受过 AI Access Research 播客的采访,我们会提供链接,那期采访涵盖了很多问题,特别是那些已经参与 AI 安全或对齐的人可能会问的问题。所以为了差异化,今天我们更关注那些来自非安全相关机器学习研究领域的人,或者完全在机器学习之外、每天旁观并试图理解这里发生了什么的人可能会有的问题。那么,目前前沿模型中占主导地位的对齐和安全技术是什么?只是你之前提到的基于人类反馈的强化学习(RLHF)研究吗?
Okay, let's first take a step back and survey the current state of the art in alignment methods and why you're confident that they're not going to be enough to align agentic models that are much more intelligent than humans. One thing I'll add is that you've done this other interview with the AI Access Research podcast, which we'll link to, which covers a lot of questions that people would be especially likely to have if they're already involved in AI safety or alignment. So in the interest of product differentiation, today we're going to focus a bit more on the questions that people might have if they're coming in from non-safety related ML research, or they're just outside machine learning and daily looking on, trying to make sense of what's going on here. So, what alignment and safety techniques are currently dominant in cutting-edge models? Is it just the research, the reinforcement learning from human feedback that you were talking about earlier?
是的,没错。基于人类反馈的强化学习(RLHF)是当今流行的方法。它效果很好,因为人类可以观察系统在做什么,并判断它是否良好。但如果你认真思考如何规模化,就会遇到这个问题:基本上,人类不会随着 AI 的进步而扩展。如果我们让 AI 系统变得更好,人类并不会自动变得更好。所以如果你想以类似的方式扩展人类监督 AI 的能力,显而易见的路径是让他们使用 AI。你可以想象:假设有一个 AI 系统试图编写一个复杂的代码库或一本复杂的教科书。现在你可以使用像 ChatGPT 这样的助手来帮你找出这本教科书中的所有问题。这可能是未来版本的 ChatGPT,使用大量插件,进行大量事实核查和浏览,阅读大量书籍等等。但根本问题是,为什么这会有帮助?基本思想是,通过辅助评估,你实际上使任务变得更容易。如果有一个 AI 系统在代码中建议一个 bug,你去检查这个 bug 是否真的存在,比一开始就找到所有 bug 要容易得多。所以通过这个找 bug 的系统,它不仅大大帮助你监督和评估实际的代码编写系统,而且它本身也是一个更容易监督的任务。你可以想象,例如,用 RLHF 训练那个任务,然后用那个系统来评估这个更难的任务。这通常是一系列我们称为可扩展监督的想法,这是我们的主要方向之一。
Yeah, that's right. Reinforcement learning from human feedback is kind of the popular method today. It works well because humans can look at what the system is doing and tell whether it's good or not. But if you're thinking hard about how to scale, you run into this problem: basically, humans don't scale with AI progress. If we make our AI systems better, humans don't automatically get better. So if you want to scale similarly, humans' ability to oversee what AI is doing, the obvious path is to get them to use AI. You could picture: let's say you have an AI system trying to write a complicated code base or a complicated textbook. Now you could use an assistant like ChatGPT to help you find all the problems in this textbook. This could be a future version of ChatGPT that uses a lot of plugins, does a lot of fact-checking and browsing, reads a bunch of books, and so on. But fundamentally, the question is why this helps. The basic idea is that you're actually making the task easier by assisting evaluation. If you have an AI system that's suggesting a bug in the code, it's much easier for you to go and check that this is in fact a bug than to find all the bugs in the first place. So by having this bug-finding system, not only does it help you a lot in overseeing and evaluating the actual code-writing system, it is also itself a task that is easier to supervise. You could picture, for example, training that task with RLHF, and then using that system to evaluate this harder task. This is generally a range of ideas that we call scalable oversight, and that's one of our main directions.
我想这里的一个假设是,如果人类能花大量时间仔细审查模型的输出,真正弄清楚它们在哪些方面好、哪些方面坏,然后在此基础上进行强化,全面细致地理解哪些做得好、哪些做得差,并强化好的、惩罚坏的,那么事情会进展得更好。但随着 AI 的进步,它将产生更复杂的输出,一个人需要更长时间来评估,或者可能因为太有挑战性而无法很好地评估。或者会有更多不同种类的模型产生更广泛的东西,而我们没有足够的人力,没有足够的人来正确检查这些输出,看看它们哪里做得好、哪里做得差。所以我们最终可能会给出糟糕的反馈,我们可能会说模型做得很好,而实际上它做得很差,或者说它做得很好,而实际上它在欺骗我们,然后我们只是在强化它学习如何更好地欺骗我们,让它知道那是一个成功的策略。所以问题是 AI 是……
I suppose an assumption here is that things would go better if only humans could spend a lot of time scrutinizing the outputs of models and figuring out really in what ways they were good and bad, and then reinforcing them on that basis, having a full sophisticated understanding of what has gone well and what has gone badly, and reinforcing the good and negatively reinforcing the bad. But as AI progresses, it is going to be producing much more complicated outputs that take much longer for a person to assess, or they just may not be able to assess it very well because it's too challenging. Or there's going to be many more different kinds of models producing a wider range of things, and we just don't have the person power, we just don't have enough people to properly check these outputs and see where they've gone well and when they've gone badly. So we could end up giving feedback that's bad, we could end up saying that the model did a great job when in fact it did a bad job, or saying it did a great job when in fact it was tricking us, and then we're just reinforcing it to learn how to trick us better, learning that that's a successful strategy. So the problem is AI is...
往前赶,人类在某种程度上受限于自身的时钟速度,我们既不会变得更快,也不会变得更聪明。但神奇之处在于:如果我们能让 AI 来做审查、做检查呢?因为这样一来,你需要检查的东西会以与检查者相同的速度变得更快、更复杂。这是基本想法吗?
Rushing ahead, humans are kind of stuck at the clock speed that they have, where we're not getting any faster or any smarter. But the magic would be: what if we could get the AIs to do the scrutinizing, to do the checking? Because then the things that you need to check are speeding up and getting more sophisticated at the same rate as the checker is getting more sophisticated. Is that the basic idea?
是的,这就是基本想法。你提出的观点很好,我想强调一下:如果你使用 RLHF,你基本上是在训练系统避免人类会发现的那种错误。所以一种可能的结果是系统泛化到「哦,我不应该犯人类会发现的那种错误」。但实际上你想要的是它泛化到「哦,我不应该犯错误」或者「我知道是错误的那种错误」。这是一个非常重要但微妙的区别。
Yep, that's the basic idea. And the point you're making is really good, which I kind of want to echo: if you use RLHF, you're basically training the system to avoid the kind of mistakes that humans would find. So one way it could go is the system then generalizes to 'oh, I shouldn't make the kind of mistakes humans would find.' But actually what you want is it to generalize to 'oh, I shouldn't make mistakes' or 'mistakes that I know are mistakes.' This is a really important but subtle distinction.
对对。不展开说说吗?
Yeah, yeah. Don't elaborate on that?
所以当我们给出不准确的反馈时,它们就会分道扬镳。想法是,如果我们的反馈总是准确的——即只有真正做得好时我们才说做得好,而且如果我们无所不知、极其聪明的话我们也会这么认为——那么你就不会出现做正确的事和避免评估者能看到的错误之间的分裂。
So they come apart when we give inaccurate feedback. The idea is that if our feedback were always accurate, in the sense that we only say a good job has been done when it truly has been done, and that's what we would think if we just knew everything, if we were incredibly brilliant ourselves, then you can't get this coming apart between doing the right thing and avoiding mistakes that are visible to the assessor.
没错。但不知道你怎么想,我觉得这太难了。我们无法接触到真实情况,对吧?我们不知道什么才是真正正确的。如果你给我一段复杂的代码,我根本不可能找出所有 bug。这太难了。而这正是挑战的核心部分,对吧?如果你有一个 AI 系统阅读大量代码——我预计未来会发生这种情况——人们会想运行这些代码,我们怎么知道 AI 系统没有在代码中秘密植入后门、木马或其他安全漏洞,而它们知道我们会漏掉这些,因为我们用反馈信号训练它们,告诉它们我们发现了哪些 bug、漏掉了哪些?
That's right. But I don't know about you, but I find it so hard. We don't have access to ground truth, right? We don't know what's actually true. If you give me a complicated piece of code, there's no way I'm going to find all the bugs. It's just too difficult. And this is also a core part of the challenge, right? If you have an AI system that reads a lot of code, which I expect will happen in the future, and people will want to run that code, how do we know that AI systems aren't secretly placing backdoors or Trojans or other security vulnerabilities into the code that they know we will miss, because we've trained them with a feedback signal that tells them exactly what kind of bugs we've spotted and missed?
那么为了让这一切奏效,我们需要哪些目前还没有的东西?
So in order to make this whole thing work, what do we need that we currently don't have?
是的,我稍微提到了可扩展监督的想法。还有一堆其他我们非常兴奋的关键拼图。另一个是理解泛化,对吧?我们能否真正预测和改进模型从我们能很好监督的简单问题到我们不能监督的困难问题的泛化?换句话说,我们如何让它们泛化到我们真正想要的东西——「不要写 bug」——而不是那个与所有证据基本一致的邻近东西——「不要写人类能发现的 bug」?我认为这是一个非常有趣且重要的问题,但它也是关于神经网络如何真正工作的核心机器学习问题之一。很奇怪这个问题上实际做的工作这么少。
Yeah, so I kind of teased a little bit like the scalable oversight idea. There are a bunch of other puzzle pieces that we're really excited about that we think are going to be crucial here. The other one is kind of understanding generalization, right? Can we really predict and improve how our models generalize from easy questions that we can supervise well to hard questions that we can't? In other words, how can we get them to generalize to the thing that we actually want, which is 'don't write bugs,' and not the nearby thing that is basically consistent with all the evidence, which is 'don't write bugs that humans find'? I think this is a really interesting and important question, but it's also one of these core machine learning questions about how neural networks really work. It's kind of puzzling that there is so little work that has actually been done on this question.
另一个可能非常重要的拼图是可解释性。这些模型,从某种意义上说,我们拥有用于人工神经网络的完美脑扫描仪。我们可以以完美的精度在每个微小的时间间隔测量它们,并且可以对它们进行任意精确的修改。这是一个非常强大的工具。从某种意义上说,它们是完全开放的盒子,只是我们不明白它们实际上是如何工作的。不去看看里面并试图理解发生了什么,回答诸如「用于训练 ChatGPT 的奖励模型是什么?它实际上在关注什么?它如何决定什么是奖励、什么不是奖励?」这样的问题,那简直是疯了。我们对此知之甚少。我们几乎一无所知。这看起来很疯狂。我们真的应该知道这些。我们不了解激励结构或它如何思考它试图做什么,这太荒谬了。
Another puzzle piece that might be really important is interpretability. These models, in a sense, we have the perfect brain scanners for artificial neural networks. We can measure them at perfect precision at every minuscule time interval, and we can make arbitrary precise modifications to them. That's a really powerful tool. In some sense, they are completely open boxes that we just don't understand how they actually work. It would be kind of crazy not to look inside and try to understand what's going on, and answer questions like: what is the reward model used to train ChatGPT? What is it actually paying attention to? How does it decide what is reward and what is not reward? We know very little about that. We know almost nothing. That seems crazy. We should really know that. It's bananas that we don't understand the incentive structure or how it thinks about what it's trying to do.
是啊,我的意思是,它就在那里。你盯着它看,这是个难题,但我认为我们可以取得真正的进展。
Yeah, I mean, it's right there. You just stare at it, and it's a hard problem, but I think we can make real progress on that.
然后还有其他问题,比如我们如何真正让模型变得鲁棒?一个例子是我们在 InstructGPT 论文中发现的:我们在一个几乎全是英语的数据集上训练它,但它能遵循其他语言的指令。我用德语问它问题,它仍然会执行。有时它可能用英语回答,这也有点奇怪。这是怎么回事?另一个例子是越狱。你在 GPT-4 上见过所有这些。你可以用这些相当简单的提示来欺骗模型,让它做它被训练不能做的任务。从某些方面来说,它并没有泛化出「我不应该做坏事」;它只是以其他方式泛化了。这是怎么回事?为什么我们不明白?它学到的教训是什么,如果不是「不要帮助人们犯罪」?相反,它只是学会了「不要帮助人们犯罪,除非你在演戏」。它怎么就没理解这些概念呢?
And then there are other questions, like how can we actually make the model really robust? One example is that we found with the InstructGPT paper: we trained it on a dataset that was almost exclusively English, and it can follow instructions in other languages. I can ask it something in German, and it will still do it. Sometimes it might answer in English, which is also kind of weird. What's going on there? Another example is jailbreaking. You've seen all of this with GPT-4. You can make these pretty simple prompts and trick the model into doing a task it was trained not to do. In some ways, it hasn't generalized 'I shouldn't do bad stuff'; it's just generalized in some other way. What's going on there? Why don't we understand that? What is the lesson that it's learning, if it's not 'don't help people commit crimes'? Instead, it's just learning 'don't help people commit crimes unless you're in a play.' How is it not getting these concepts?
是啊,而且人类似乎能很好地做到这一点。我的意思是,人类做得并不完美,但区别在哪里?
Yeah, and it seems like humans can do this well. I mean, humans don't do it perfectly, but what's the difference here?
是的,所以这是泛化的另一个方面,我认为理解它对我们非常有用。最后,我们想做的事情之一是故意训练欺骗性对齐的模型——那些试图非常连贯地对我们撒谎或秘密进行自我外泄等行为的模型。那就是模型逃离实验室。没错。因为我们想确信我们能抓住这些尝试。获得信心的直接方法是:我们故意训练它,然后检查它是否会通过或不被察觉。但当然,如果你这样做,你必须非常小心,不要意外地创造出你一直试图避免的东西。
Yeah, so this is another aspect of generalization that I think could be really useful for us to understand. And finally, one of the things we want to do is actually deliberately train deceptively aligned models—models that try to lie to us very coherently or try to secretly do something like self-exfiltration. That's a model kind of breaking out of the lab. That's right. Because we want to be confident that we could catch these attempts. The straightforward way to be confident is: we deliberately train it, and then we check whether it would pass or fly under the radar. But of course, if you're doing this, you have to be super careful that you're not accidentally creating the thing you've been trying to avoid.
首先,这必须非常小心地进行。是的,在我看来,人们对模型在不完美反馈下学会真正欺骗行为的可能性有着非常不同的直觉。想象一下,我们训练一个模型,不希望它撒谎,十次中有九次我们抓到它撒谎并给予负面反馈,但有一次我们错误地说「干得好」。人类似乎会学到一种普遍的厌恶撒谎的倾向,即使我们认为自己可能能逃脱惩罚。这是大多数人的泛化方式,尽管并非所有人。但有些人认为这种情况是灾难性的,因为你训练了模型进行最复杂的撒谎,并在它认为可以逃脱时耍花招,而在不能逃脱时则不撒谎。另一些人认为它只会学到普遍的厌恶撒谎,一切都会好。你同意我的看法吗?人们对此有非常不同的直觉?你有什么论点吗?
The first place, so it has to be done very carefully. Yeah, it seems to me like people have very different intuitions about how likely it is that a model that gets imperfect feedback is going to learn to engage in really deceptive behavior. So if you imagine that we train a model and we don't want it to lie, and nine times out of ten we catch it lying and give it negative feedback, but one time in ten we accidentally say 'yeah, you did a good job' when it lied. It seems like humans kind of learn this general aversion to lying, even when we think that we might be able to get away with it. That's kind of how most people generalize, although I guess not all. But some people think that in that situation it's just disastrous because you've just trained the model to engage in the most sophisticated lying possible, and to be tricky whenever it thinks it can get away with it and not when it can't. Other people think it'll just learn this general aversion to lying and everything's going to be fine. Do you share my perception that people have very different intuitions about this? And I guess what do you argue, if you have any?
我认为这恰恰说明我们不知道,而我们应该知道。我认为找出答案的最佳方法之一是进行实证尝试。我们现在可以用模型进行许多有趣的实验,正是这种性质的。比如我们可以训练它们成为更好的撒谎者,观察它的行为方式,如何泛化。我们的总体目标是达到能够自动化对齐研究的程度。这并不意味着我们要训练一个非常擅长机器学习研究或非常聪明的系统。那不是超级对齐的工作。是的,我想很多人一直这么认为。他们读了你的公告,认为你基本上是在训练一个非常好的机器学习研究员。我不认为这会特别有助于对齐,所以澄清一下是好的。基本上,我理解我们的工作是:我们必须找出对齐技术,使我们有足够信心,一旦我们拥有足够智能的模型,一旦有能做机器学习研究或类似事情的模型——我认为这无论如何都会发生,无论 OpenAI 是否做——但我们的工作是找出如何使其充分对齐,以便我们能够信任它产生的对齐研究或对齐研究辅助。因为本质上,如果你让这个系统帮助你的对齐研究,系统有很大的机会影响或试图引导我们相信某些技术实际上很好,但其实不然,从而该系统或未来系统以我们不希望且不符合我们利益的方式获得对人类的力量。所以我们最终需要做的是找出如何使该系统充分对齐,以便我们真正信任它。这意味着,例如,如果系统帮助——简单来说——系统写了一篇对齐论文。你可以阅读论文,但一开始你可能无法找出论文中的所有缺陷。一般来说,科学同行评审并不完美,有很多例子表明人们伪造研究数十年才被发现。所以这是我们必须真正弄清楚如何避免的事情。因为对齐研究人员——一般的科学研究是一项困难的任务,人类并不擅长评估,至少如果你没有很多时间的话。那么问题就变成了:我们需要什么样的对齐技术才能有足够信心?
I think it just makes it clear that we don't know, and I think we should know. And I think one of the best ways to figure this out is to try it empirically. I think there are so many interesting experiments we can run now with the models, exactly of this nature. Like we could try to train them to be better liars and see how it behaves, how it generalizes. Our overall goal is to get to a point where we can automate alignment research. And what this kind of doesn't mean is that we're not trying to train a system that's really good at ML research or that is really smart or something. That's not superalignment's job. Yeah, I think a lot of people have been thinking that. I think they've read your announcement and say that you're trying to train a really good ML researcher basically. I don't think this would particularly differentially help alignment, so I think it would be good to clarify. Basically, what I understand our job is: we have to figure out the alignment techniques that would make us sufficiently confident that once we have models that are smart enough, once there are models that can do ML research or things close to it—I think that's something that's going to happen anyway, whether OpenAI does it or not—but our job is to figure out how to make it sufficiently aligned that we can trust the alignment research or the alignment research assistance that it is producing. Because essentially, if you're asking this system to help you in your alignment research, there's a big opportunity for the system to influence or try to nudge us into believing certain techniques are really good that actually aren't, and thus that system or future systems gain power over humans in a way that we don't want and that isn't aligned with us. So what we ultimately need to do is figure out how to make that system sufficiently aligned that we can actually trust it. That means, for example, if the system helps—let's say for simplicity—the system writes an alignment paper. Now you can read the paper, but after just off the bat, you might not actually be able to find all the flaws in the paper. In general, scientific peer review is not perfect, and there are lots of examples where people go for decades with fake research before being found out. So this is something we have to really figure out how to avoid. Because alignment researchers—scientific research in general is a difficult task that humans aren't that good at evaluating, at least if you don't have a lot of time to do it. The question then becomes: what kind of alignment techniques do we need in order to be sufficiently confident that this is the case?
让我给你一些例子。关于可扩展监督,我们可以考虑的一个目标是:假设我们想非常擅长找出模型中已知的所有代码错误。那么如何做到呢?一条路径可能是:我们训练一个在代码中引入错误的模型,然后看看可扩展监督方法能否捕捉到它。我很乐意详细说明我们可以在这里做什么样的可扩展监督,但高层次上,有很多不同的想法,我们应该全部尝试,看看哪些在经验上有效。这就是为什么衡量它是否有效或你是否在可扩展监督上取得进展很重要。通过训练这些系统引入错误,系统会意识到它引入的错误,对吧?那么问题是:如果一个人工智能辅助的人类能找到错误,那我们就没问题,或者你想测量他们找到错误的频率,或者我们引入可扩展监督找不到的错误的难易程度。另一条路径是我们所谓的判别性批评,基本上我们训练一个模型成为有错误代码和无错误代码的判别器,我们可以手动或自动生成这些代码,然后将其与一个被提示描述它能找到的所有错误的批评模型进行比较。所以如果你有一堆代码样本,判别器标记为可疑,但批评模型没有指出存在的错误,那么你就知道模型意识到了代码中的某些东西,而你没有找到。我认为理解两者都意识到这一点,是因为它们以某种方式以相同的方式训练。
Let me give you some examples. On scalable oversight, one of the goals we could think about is: let's say we want to get really good at finding all the bugs in code that the model knows about. So how do we do this? One path could be: we train a model that introduces bugs in code, and then we can see whether a scalable oversight method can catch it. I'm happy to go into what kind of scalable oversight we could do here, but on the high level, there are lots of different ideas and we should just try them all and see which ones work empirically. That's why it's important to measure whether it's working or whether you're making progress in scalable oversight. So by training these systems to introduce bugs, the system will be aware of the bugs it's introducing, right? So the question is: if an AI-assisted human can find the bug, then we're good, or you want to measure how often they will find it, or how easy it is for us to introduce bugs that the scalable oversight doesn't find. Another path is what we call the discriminative critique, where basically we train a model to be a discriminator between buggy code and non-buggy code, which we can generate manually or automatically, and then we compare that to, let's say, a critique model that is prompted to describe all the bugs it can find. So if you end up with a bunch of code samples that the discriminator flags as fishy but the critique model doesn't point you to the bug that is going on there, then you know there's something about the code the model is aware of that you're not finding. And I think by understanding that both are aware of it because they were trained in the same way somehow.
所以我现在想象对判别器和批评模型使用相同的基础模型。明白了。
So I'm right now picturing using the same base model for both the discriminator and the critique model. Got it.
是的。通过详细观察,你可以获得一些经验测量,了解模型没有告诉你的代码中存在错误的频率。然后你可以对抗性地进行,这样你就不只是测量平均情况。我认为这是一条非常有前途的路径,不仅可以衡量我们在可扩展监督上取得进展,因为你可以想象,如果你能捕捉到更多模型意识到的这些错误,那么你就改进了可扩展监督,而且如果你真的做对了,你会说「哦,我们达到了...」
Yeah. And so by looking at this in detail, you can get some empirical measurements for how often it actually happens that there is something buggy about the code the model doesn't tell you. And then you can do it adversarially, so you're not just measuring the average case. I think that's a really promising path, not only to measure we're making progress in scalable oversight because you can picture if you can catch more of these bugs that the model is aware of, then you've made scalable oversight better, but also if you actually nail this right and you're like 'oh, we get to the point where we...'
可以非常确信,如果汽车有什么可疑之处,比如代码有问题,模型知道,你会发现,这会让你更放心地实际执行代码看看会发生什么。好的,退一步说,基本想法是,能够进行人工智能研究的机器学习模型即将出现,无论我们是否想要。我的意思是,很多人对此感到紧张,因为这可能引发递归自我改进的循环。所以可能有充分的理由稍微推迟那个时刻,但我们可能无法永远推迟。因此,当那个时刻来临时,我们首先要做的是找到利用具有这些能力的模型进行对齐研究以及非对齐机器学习研究的方法。而且,至关重要的是,我们能够达到一个状态,相信这些模型足够值得信赖,从而可以相信它们在改进我们当前阶段的对齐研究方面提供的帮助。我们既需要弄清楚如何让它们足够值得信赖以便使用它们的输出,也需要能够知道我们已经成功做到了这一点。简而言之就是这样。
Can be super confident that if there was something fishy about the car like there's some problem with the code the model is aware of you'll find out that would make you so much more comfortable with then like actually executing this the code to see what happens yeah okay just just to back up a second the basic idea is machine learning models that are capable of doing AI research coming whether we want it or not. I mean many people are nervous about that because it could set up this recursive self-improvement loop. So there could be good reasons to maybe delay that moment a bit, but we're not going to be able to probably to delay that forever. And so what we want to do when that moment comes is firstly find ways that we can use models with those capabilities to do alignment research as well as non-alignment machine learning research. And also essentially, it's very essential that we'll be able to get to a place where we believe that these models are trustworthy enough that we can believe the help that they're giving us on improving our alignment research from the stage that it's at. We both need to be able to figure out how we can get them to be sufficiently trustworthy that we can use those outputs, and also to be able to know that we've succeeded at doing that. That's the long and short of it.
是的,总的来说,我想对这件事何时成为可能保持不可知论。对吧?比如什么时候会有自动化的对齐研究,或者模型什么时候会聪明到能够做到这一点?可能会有延迟,可能有各种原因导致它发生得晚而不是早。我真正想做的是,一旦这些系统成为可能,就准备好将它们用于对齐研究。所以我们不想做的是加速这个过程或让它更早发生,因为我认为它迟早会发生。但我们希望准备好,然后利用它们进行对齐研究,并随着机器学习进展在那时加速,准备好更快地推进对齐工作。
Yeah, like in general, I want to be agnostic towards when exactly this is possible. Right? Like when will there be automated alignment research, or when will models be so smart they can do that? There might be delays, there might be all kinds of reasons why it happens later than sooner. The thing I really want to do is I want to be ready to use these systems for alignment research once that becomes possible. So what we don't want to do is accelerate this or make it happen sooner, because it will happen soon enough, I think. But we want to be ready to then use them for alignment research and be ready to make alignment progress faster as ML progress gets faster at that point.
是的,我认为要记住愿景中的一个重要部分:对齐并弄清楚一个远超人类能力的 AI 的可信度可能极其困难,因为它有太多不同的方式来欺骗你。但希望在于,当这些模型首次可用时,它们将更接近人类水平。它们可能会在某些领域比人类弱一点,但在其他领域非常强大。但由于它们不会那么不可思议地强大,可能更容易弄清楚我们是否能信任它们,因为它们的行动空间不会有那么多选择,而且它们可能更容易理解,因为它们大脑中实际做的事情更接近我们做的事情,而不是那种行星大小的思维可能做的事情。
Yeah, I think an important part of the vision to keep in mind is that it might be extremely difficult to align and figure out the trustworthiness of an AI that is just extraordinarily above human capabilities, that is truly extraordinarily super intelligent, because this is going to have so many different ways of tricking you. But the hope here is that at the point when these models are first available, they're going to be more like around human level. They'll be like, although they might even have some areas where they're a little bit weaker than people, but other areas where they're very strong. But because they're not going to be so incredibly capable, it might be easier to figure out whether we can trust them, because they're not going to have so many options in their space of actions, and they might be somewhat more scrutable because the actual things that they're doing in their mind are closer to maybe what we're doing than what a kind of planet-sized mind might be able to do.
嗯,我认为很多人可能会对此持怀疑态度,因为他们认为,它比我们聪明,所以它总能轻松胜过我们。你也许可以特意确保你处理的模型不是你能制造的最有能力的,以便更容易评估其可信度。是的,我认为这是对的,而且我认为这是一个核心点。对吧?比如,如果你考虑如何真正对齐一个超级智能,如何对齐一个比人类聪明得多的系统?我不知道。我没有答案。我认为没有人真正有答案。但这也不是我们从根本上需要解决的问题,对吧?因为如果你有这个问题,也许这个问题甚至不是今天活着的人类能解决的。但有一个更简单的问题,那就是如何对齐下一代系统?如何对齐 GPT-n+1?这是一个容易得多的问题。而且更进一步,如果人类能解决这个问题,那么一个和解决这个问题的人类一样聪明的虚拟系统也应该能解决。所以如果你让那个虚拟系统对齐,它就能解决 GPT-n+1 的对齐问题,然后你可以迭代地自举,直到达到超级智能水平并弄清楚如何对齐它。当然,做这件事时重要的是,在每一步,你都必须在对齐问题上取得足够的进展,以至于你确信 GPT-n+1 足够对齐,可以用于对齐研究。
Well, I think many people might have a bunch of skepticism about this because they think, well, it's smarter than us, so it's going to always be able to run rings around us. And you could maybe go out of your way to make sure that you're not dealing with a model that's as capable as you possibly could make in order to make it easier to evaluate the trustworthiness. Yeah, I think that's right, and I think that's a really central point. Right? Like if you're thinking about how do you actually align a superintelligence, how do you align the system that's vastly smarter than humans? I don't know. I don't have an answer. I don't think anyone really has an answer. But it's also not the problem that we fundamentally need to solve, right? Because if you have this problem, maybe this problem isn't even solvable by humans who live today. But there's this easier problem, which is how do you align the system that is the next generation? How do you align GPT-n plus one? And that is a substantially easier problem. And then even more, if humans can solve that problem, then so should a virtual system that is as smart as the humans working on the problem. And so if you get that virtual system to be aligned, it can then solve the alignment problem for GPT-n plus one, and then you can iteratively bootstrap yourself until you're at superintelligence level and you figured out how to align that. And of course, what's important when you're doing this is that at each step, you have to make enough progress on the problem that you're confident that GPT-n plus one is aligned enough that you can use it for alignment research.
是的,机器学习社区,我特别想到那些不参与安全或对齐研究的人,他们对这个计划或公告有什么反应?
Yeah, the machine learning community, I'm thinking of folks who aren't involved in safety or alignment research in particular, have they reacted to this plan or announcement?
是的,我认为总的来说,人们对我们在努力解决的研究问题感到非常兴奋。而且我认为在很多方面,从机器学习的角度来看,这些问题非常有趣。我不知道,我认为这个公告表明我们是认真对待这项工作的,我们正在努力组建一个非常优秀的团队来解决这个问题,并且我们正在努力快速取得进展并应对雄心勃勃的想法。我还认为,特别是在过去六个月左右,机器学习社区对这些问题的兴趣和资金投入大大增加。而且我认为,ChatGPT 和类似助手的成功也清楚地表明,基于人类反馈的强化学习(RLHF)中有一些有趣的事情,对齐问题也是真实存在的。对吧?比如,如果你把 ChatGPT 和原始基础模型比较,它们实际上非常不同,这里发生了一些重要的事情。
Yeah, I think in general people are really excited about the research problems that we are trying to solve. And I think in a lot of ways, they're really interesting from a machine learning perspective. I don't know, I think the announcement kind of showed that we were serious about working on this and that we're trying to get a really high caliber team on this problem and that we are trying to make a lot of progress quickly and tackling ambitious ideas. I think also, especially in the last kind of six months or so, there's been a lot more interest and funding in the machine learning community for these kinds of problems. And I think also, the success of ChatGPT and similar assistants has made it really clear that there's something interesting going on with RLHF and there's something real about this alignment problem. Right? Like if you compare ChatGPT to the original base model, they're actually quite different and there's something important that's happening here.
是的,回顾我们五年前的采访,我们谈了很多关于基于人类反馈的强化学习,因为那是新事物,也是当时的热点。OpenAI 参与了提出这种方法吗?
Yeah, listen back to our interview from five years ago, and we talked a lot about reinforcement learning from human feedback because that was new and that was the hot thing back then. Was OpenAI involved in coming up with that method?
我认为更准确地说,世界上可能有很多不同的人发明了它。在我们发表《基于人类偏好的深度强化学习》论文之前,还有其他先前的研究以各种形式进行了基于人类反馈的强化学习,但没有使用深度学习。
I think more accurately, probably a lot of different people in the world invented it. Before we did the Deep RL from Human Preferences paper, there were other previous research that had done RL from human feedback in various forms, but it wasn't using deep learning.
这些系统大多只是概念验证性质的东西。然后,来自《人类偏好》论文的更深层次内容是我与 Paul Christiano 和 Dario Amodei 的合作成果。我想我们基本上都独立得出了结论,认为这是正确的方向,然后我们合作了,结果证明这对让 ChatGPT 如此成功至关重要。
Systems and it was mostly just proof-of-concept style things. And then the deeper from the Human Preference paper was joint work with Paul Christiano and Dario Amodei and me. I think we kind of all independently came to the conclusion that this is the way to go, and then we collaborated and that turned out to be really key to getting ChatGPT to work as well as it does.
是的,没错。
Yeah, that's right.
而且它实际效果如此之好,让我感到非常震惊。如果你看看最初的 InstructGPT 论文,其中一个主要结果是,一个 GPT-2 大小的系统——参数数量比 GPT-3 小两个数量级——实际上比 GPT-3 基础模型更受青睐。所以这个更便宜、更简单、更小的系统,一旦经过对齐,就比大系统好得多。在某种程度上这并不奇怪,因为你用人类偏好训练它,它当然会更符合人类偏好。但话说回来,为什么之前没有用人类偏好训练呢?显然这是应该做的,因为你想要的就是一个人类偏好的系统。事后看来,这太明显了。
And it's been wild to me how well it actually worked. If you look at the original InstructGPT paper, one of the headline results was that a GPT-2 size system, which is two orders of magnitude smaller than GPT-3 in terms of parameter count, was actually preferred over the GPT-3 base model. So this vastly cheaper, simpler, smaller system, once you made it aligned, was so much better than the big system. To some extent it's not surprising because you train it on human preferences, of course it's going to be better for human preferences. But then also, why the hell haven't you been training on human preferences? Obviously that's what you should do because that's what you want—a system that humans prefer. In hindsight, it's so obvious.
回到机器学习领域的人,他们对计划中的哪些部分(如果有的话)持怀疑态度?你听到过什么反对意见吗?
Coming back to machine learning folks, what parts of the plan, if any, are they skeptical of? Are there objections you've been hearing from people?
我认为对于技术发展的速度以及在未来几年内真正实现研究自动化的可行性,仍然存在许多不同观点。我认为这是非常可能的,但也可能不会发生——没人真正知道。关键在于,这里有一些非常深刻且重要的问题需要解决,而且这些问题也是相当易于处理的。我们可以在未来几年取得很大进展。通过这样做,这项工作可能会产生难以置信的影响,因为这些技术将塑造未来版本的 ChatGPT 和广泛应用的 AI 系统,并在经济中执行许多任务。有很多更容易的系统信号可以优化——你可以优化 AI 系统以最大化客户购买量或最大化注意力。我们在过去十年左右已经看到了这方面的迹象,很多人不喜欢这样。这些信号本质上很容易测量,但它们与人类或人类真正想要的或长期的人类繁荣并不一致。因此,随着 AI 在世界上的影响力越来越大,我们在对齐方面做得如何将产生非常广泛的影响,并以多种方式塑造社会,无论好坏。我认为我们在这方面做得非常出色是至关重要的。
I think there are a lot of different views still on how fast the technology is going to develop and how feasible it is to actually automate research in the next few years. I think it's very possible, but also it might not happen—nobody actually knows. The key thing is that there are some really deep and important problems here that we really need to solve, and they are also really tractable. We can make a lot of progress over the next few years. By doing this, this could be incredibly impactful work because these techniques will shape future versions of ChatGPT and future AI systems that are widely applied and do lots of tasks in the economy. There are a lot of much easier system signals you could optimize—you could optimize AI systems to maximize customer purchases or to maximize attention. We've seen glimpses of that over the last decade or so, and a lot of people don't like that. Those signals are fundamentally easy to measure but they are not aligned with humans or what humans actually want or long-term human flourishing. So in some ways, as AI becomes more impactful in the world, how well we do alignment will have really wide-ranging consequences and shape society in lots of ways for better and worse. I think it's really paramount that we do an excellent job at this.
你提到了几种可能实现自动化或使用这些机器学习工具的方法。有可扩展监督、泛化和可解释性。我不太明白泛化作为一个集群是什么意思。能再解释一下,也许再详细说明一下吗?
You mentioned a couple of different ways that things might get automated or ways that you might be able to use these ML tools. There was scalable oversight, and generalization, and interpretability. I don't fully get what generalization is as a cluster. Is it possible to explain that again and maybe elaborate a bit more?
从根本上说,我们希望能够区分:系统是泛化到真正的人类意图,还是泛化到人类在监督时所说的内容,但在其他情况下却做别的事情?这是两种不同的泛化,它们与数据完全一致,因为在我们监督时行为都是一样的。但泛化从根本上说是关于模型和数据的问题。那么为什么我们不能直接去理解它呢?例如,我们现在正在做的是在一个玩具环境中研究这个问题。你可以这样做:取一个数据集,看看一个小语言模型能正确回答什么——我们称这些为数据集的简单部分——然后剩下的称为困难部分。现在的问题是:如果我们只训练简单部分的标签,看看我们能多好地泛化到困难部分?或者我们可以在模型训练中加入什么样的技巧,使我们泛化得更好?另一件你可以做的事情是,只从小模型生成大量标签。这里的类比是:如果你让人类监督比他们更聪明或和他们一样聪明的系统,在某些方面我们会比那个系统弱,我们的标签会比系统能做的更差。那么,如何仅通过使用弱标签或仅使用简单问题的标签,来恢复如果你一开始就用真实标签训练所能获得的准确率?我们可以在这里进行一些非常具体的实验,这些实验可以告诉我们很多关于真实情况会如何发展的信息。一旦我们有了这些并开发了一些技巧,我们能否在更真实的环境中使用这些技巧?我们能否从小语言模型在 ChatGPT 偏好数据集上的标签泛化到由 GPT-4 完成的实际真实 ChatGPT 任务?我认为这些都是非常有趣的问题,我们实际上可以进行实验并学到很多东西。这不仅与我们真正想要解决的对齐问题高度相关——我们试图让系统在我们难以监督的环境中正确泛化——而且我认为我们还将学到一些关于神经网络如何学习的非常有趣的基础知识。
Fundamentally, we want to be able to distinguish: does the system generalize to true human intent, or does it generalize to what the human says whenever they're looking but do something else otherwise? These are two different generalizations, entirely consistent with the data because the behavior is all the same whenever we're supervising. But generalization is fundamentally a problem about the model and the data. So why can't we just go and try to understand it? For example, what we're doing right now is studying this in a toy setting. The way you could do this is: take a data set and look at what a small language model gets correct—let's call these the easy parts of the data set—and then call the rest the hard part. Now the question is: what if we only train on the labels for the easy part and see how well we can generalize to the hard part? Or what kind of tricks could we put into the model training that would make us generalize better? Another thing you could do is just make a lot of labels from a small model. The analogy here is: if you have humans supervising systems that are smarter than them, or as smart as them, in some ways we'll be weaker than that system and our labels will be worse than what the system could do. So how can you recover the accuracy you would get if you just trained on the ground truth labels in the first place, by only using the weak labels or only using the labels on the easy questions? There are some really concrete experiments we can run here that could tell us a lot for how this is going to go in the real case. Once we have that and have developed some tricks, can we use the tricks in a more real setting? Can we generalize from labels by a small language model on the ChatGPT preference data set to the actual real ChatGPT tasks done by GPT-4? I think these are really interesting questions that we can actually run experiments on and learn a lot. Not only is this highly relevant for the kind of alignment problems we actually want to solve—where we're trying to get it to generalize correctly in settings that are hard for us to supervise—but also I think we'll learn some really interesting fundamental things about how neural networks learn.
关于泛化,是否已经进行过有趣的实验?有关于这个主题的论文吗?
Have any interesting experiments on generalization been run already? Are there papers on this topic?
文献中有很多研究,但实际上数量少得惊人。这种分布外泛化……我想我们可能会在两三个月内发表一篇非常令人兴奋的论文。如果你觉得这很令人兴奋,正在研究这个的研究团队现在正在招聘。我们正在为这个团队寻找一位经理。
There's a bunch of research in literature, but it's actually surprisingly small. This kind of out-of-distribution generalization... I think we'll probably have a pretty exciting paper in two or three months on this topic. If you find this exciting, the research team that is working on this is hiring right now. We're looking for a manager for this team.
让这项研究发生,比如写我们的第一篇论文,现在正是时候。听起来可能有一个项目要创建一个在某些情况下会进行欺骗的模型。一种——我想用什么术语?我们需要一个大肠杆菌或果蝇,一个不良行为的模式生物,以便研究它,看看它何时出现、在什么条件下、如何减少它。这是整个议程的一部分吗?
Part of making this research happen, like writing our first ever paper, that now is the time. It sounded like there was maybe going to be a project to create a model that engages in deception in particular cases. A kind of—I suppose what's the term? We need an E. coli or a drosophila, a model organism of bad behavior, in order to study it and see when does it appear, under what conditions, how can you reduce it. Is that part of this whole agenda?
我真的很喜欢你描述的方式。我们称之为对抗性测试,这可能更普通一些。这里的目标是:如果我们能故意训练一个具有欺骗性或试图欺骗我们的模型,但方式非常普通——比如它只是试图运行一个实际上什么都不做的特定 Python 函数,但我们可以很容易地测量它是否成功,并且它规避了我们使用的其他监督技术——这就给了你描述的那个模式生物。现在我们有了想要避免的系统,并且可以实际研究它。例如,我们可以用可解释性工具观察它,并问:我们能区分这个模型的思考方式与正常模型的不同吗?或者我们能对这个模型使用我们的泛化技巧,使其在训练时避免那个陷阱吗?另一个你真正想要泛化的属性是:告诉我真相,或者明确告诉我你对真相的最佳猜测。语言模型对世界有很多理解——它们比人类了解得更多——但它们并不总是说出来。事实上,如果你看看基于人类反馈的强化学习(RLHF)训练是如何工作的,你是在避免你认为真实的东西,而其中一些东西实际上并不真实,所以你是在训练模型基本上撒谎。相反,有些你知道是真实的东西模型不知道——比如训练数据截止后发生的任何事情。然后还有你不知道但模型非常有信心的东西,比如生物学事实。在进行 RLHF 时,你在两种情况下都让情况变得更糟:第一种情况,你训练它编造不真实的东西,增加了幻觉;第二种情况,你在削弱模型,因为你激励它不告诉你它实际知道的东西。可能很少有模型和人类知识完全匹配的情况。所以 RLHF 能工作得这么好真是有点疯狂。回到泛化问题:如果你能以某种方式获得模型对真相的最佳表示呢?它不会完美,但它对什么是真实的有一个猜测,我们想要得到模型的最佳意见。将这个问题表述为一个泛化问题的一种方式是:取一堆我们有信心回答的简单案例——比如美国第一任总统是谁,二加二等于几——然后泛化到我们不太有信心的难题,比如我不知道答案的生物学问题,或者世界上只有少数专家知道答案但模型知道的情况。如果你解决了这个问题,你就可以用你的模型进行监控或监督:问这段代码中是否有隐藏的缺陷,或者另一个模型的行为是否可能是欺骗我们的尝试,或者副作用过滤。如果我们能解决这个问题,就能取得很大进展。
I really like how you describe it. We call that adversarial testing, which is maybe a lot more mundane. The goal here would be: if we can deliberately train a model that is deceptive or tries to deceive us, but maybe in a really mundane way—like it just tries to run a specific Python function that does nothing actually, but we can easily measure whether it succeeded in a way that evaded other oversight techniques we're using—that will give us this model organism you describe. Now we have the system we want to avoid, and we can actually study it. For example, we could look at it with interpretability tools and ask: can we tell the difference between how this model thinks compared to the normal model? Or can we use our generalization tricks on this model to make it avoid that pitfall when we're training it? Another property you'd really want to generalize is: tell me the truth, or markedly tell me your best guess at the truth. Language models have a lot of understanding of the world—they know a lot more about the world than humans do—but they're not always necessarily telling it. In fact, if you look at how RLHF training works, you're avoiding the things you think are true, and some of those things are not actually true, so you're training the model to basically lie. Conversely, some things you know are true but the model doesn't know—like anything that happened since the training data cutoff. Then there's stuff you don't know but the model is very confident on, like biology facts. When doing RLHF, you actually make it worse in both cases: in the first case, you're training it to make up stuff that isn't true, increasing hallucinations; in the second case, you're sandbagging the model because you're incentivizing it not to tell you something it actually knows. There are probably very few things where the model is exactly as knowledgeable as the human. So it's kind of crazy that RLHF works as well as it does. Coming back to the generalization question: what if you could somehow get to the model's best representation of the truth? It won't be perfect, but it has some kind of guess of what's actually true, and we want to get the model's best opinion. One way to phrase this as a generalization question: take a bunch of easy cases where we're very confident in the answer—like who was the first US president, what is two plus two—and then generalize to hard questions where we're not that confident, like some biology question I don't know the answer to, or cases where very few experts know the answer but the model does. If you nail that, you can use your model for monitoring or oversight: ask if there's a secret hidden flaw in this code, or if the behavior of another model is potentially an attempt at deceiving us, or side effects filtration. If we can nail that problem, we can make a lot of progress.
我们再谈谈可解释性。我觉得人们已经能够弄清楚神经网络使用什么算法来感知图像中的特定纹理,或在句子中进行特定的推理,或弄清楚如何确保名字的一致性,这非常令人印象深刻和有趣。但我不确定这如何帮助对齐系统,因为这些都是很小的事情,感觉加起来并不能告诉我模型的意图和目标是什么。Ajeya Cotra 在我几个月前对她的采访中指出,你可以进行更高层次的可解释性:让模型多次说实话和多次撒谎,然后观察当它处于欺骗模式时网络的哪些部分被激活。这可能很容易弄清楚,并且可能非常有帮助。你认为哪些可解释性的攻击路线可能有用,并且可能部分或最终自动化?
Let's talk a little more about interpretability. I find it very impressive and interesting that people have managed to figure out what algorithms neural networks are using to perceive a particular texture in an image, or to do a particular piece of inference within a sentence, or to figure out how to make sure a name is consistent. But I'm not sure how that would help align a system, because it's just these small things and it doesn't feel like it adds up to telling me what the goals and intentions of the model are. Ajeya Cotra pointed out in my interview with her that you could potentially do a much higher level of interpretability: get a model to tell the truth a bunch of times and lie a bunch of times, then see what parts of the network light up when it's in deceptive mode. That may be straightforward to figure out and could be super helpful. What sort of lines of attack on interpretability that would be useful do you think you might be able to partially or ultimately automate?
你可能需要两个方面:一些能真正在模型工作的细节层面起作用的东西,这样你就不会错过任何重要的事情,但同时你必须查看整个网络,因为你寻找的东西可能在任何地方。所以如果你两者都需要,你需要结合详细的机械可解释性和更高层次的行为可解释性。Ajeya 提到的那种方法——比较说真话和说谎的状态——很有前景。你可以使用自动化工具来找到这些与欺骗相关的高层电路或特征。然后你可以在部署期间监控这些特征,以检测模型何时可能在欺骗你。这可以部分自动化并规模化。
You probably want both aspects: something that really works in the minute detail of how the model works so that you don't miss anything important, but at the same time you have to look across the network because the thing you're looking for might be anywhere. So if you want both, you need a combination of detailed mechanistic interpretability and higher-level behavioral interpretability. The kind of approach Ajeya mentioned—comparing truth-telling and lying states—is promising. You could use automated tools to find these high-level circuits or features that correlate with deception. Then you could monitor those features during deployment to detect when the model might be deceiving you. This could be partially automated and scaled up.
同时,它非常精简。没有多少东西具有这种特性。特别是,人类历史上做可解释性的方式是盯着模型的某些部分,看是否能理解它们,这只能得到其中一部分,而不是全部。所以我们刚刚发表了一篇关于自动化可解释性的论文,它试图同时做到这两点。这算是第一次尝试,所以比较简化。我们做的是让 GPT-4 来编写单个神经元行为的解释。通过将大量文本输入模型,记录每个特定 token 上神经元的激活程度,然后让 GPT-4 查看这些数据并编写解释。平均而言,这些解释并不太好。有时它们很好,有时很有趣。例如,这就是我们发现「加拿大神经元」的方式,它会在相关概念上激活。GPT-4 理解这一点并指了出来,然后写了这个解释。更进一步,你可以衡量这些解释的质量。你在保留的文本上运行它们,让 GPT-4 仅根据解释来预测人类会如何标记激活。现在你有两样东西:自动化的解释编写和自动化的评分函数。现在你可以大展拳脚了,因为你可以优化评分函数,做各种事情。例如,我们做了迭代改进,你批评你的偏见,解释在评分函数上会得到更高的分数。同时,你也可以改进评分函数,让它更准确地模拟人类如何预测神经元的激活,或者接入一个更强大的模型。这种方法也有一些问题。例如,神经元可能不是解释模型的正确抽象层次,因为神经元做很多不同的事情。这就是人们所说的多义性。很难编写一个覆盖所有情况的解释。但有一点非常好,就是你可以大规模运行。我们在 GPT-2 的所有神经元上运行了它,那有很多神经元——大约 30 万个。你得到大量文本,可以筛选并寻找某些东西。理论上你可以在 GPT-4 上运行,但那会非常昂贵,而且我个人认为不值得,因为解释还不够好。但它有一个很好的方面,就是你真正在查看模型的每一个部分,字面意义上的每一个神经元,并试图解释它的功能。同时,你遍历整个模型,试图解释每一个神经元。如果我们有一种真正有效的技术,那将是一个彻底的改变游戏规则的东西。
Things at the same time, it's really lean. There are not that many things that have this property. In particular, the way humans do interpretability historically is you stare at parts of the model and see if you can make sense of them, which gives you one of them but not all. So we just released a paper on automated interpretability, which tries to do both at the same time. It's kind of a first attempt, so it's simplified. What we do is we ask GPT-4 to write explanations of behavior of individual neurons. By piping a bunch of text through the model, recording how much the neuron activates at each particular token, you can ask GPT-4 to look at that and write an explanation. On average, these explanations are not very good. Sometimes they're good, sometimes they're interesting. This is how, for example, we found the Canada neuron that fires on related concepts. GPT-4 understood that and pointed it out, and wrote this explanation. Even more, you can measure how good these explanations are. You run them on a held-out piece of text and get GPT-4 to predict how a human would label the activations based on the explanation alone. Now you have two things: the automated explanation writing and the automatic scoring function. Now you're in business because you can optimize the score function and do all kinds of things. For example, we did iterative refinements where you critique your biases and the explanations get higher on the score function. At the same time, you can improve your score function by having it more accurately model how humans would predict how the neuron would activate, or by plugging in a more capable model. There are some problems with this approach too. For example, neurons are probably not the right level of abstraction to interpret the model in, because neurons do a lot of different things. This is what people call polysemanticity. It's hard to write an explanation that covers all the cases. But one thing that's really nice is you can run this at scale. We ran it over all neurons in GPT-2, and that's a lot of neurons—it was like 300,000 neurons. You get a lot of text, and you can sift through it and try to find certain things. You could theoretically run this on GPT-4, but it would be really expensive and personally it wouldn't be worth it because the explanations just aren't good enough. But it has this nice aspect where you're really looking at every part of the model, literally every neuron, and trying to explain what it does. At the same time, you're running over the whole model, trying to explain every neuron. If we have a technique like that that actually works really well, that would be a complete game changer.
是的。所以这里的部分想法是,让一整队人类辛苦地找出有一个神经元对应加拿大,这并不令人满意。我们不知道从中能得到什么。但如果你能自动化,让你拥有相当于成千上万名员工在仔细检查并试图弄清楚神经网络的每个部分在做什么——如果你能自动化的话,你也许能做到——那么这可能会构成一幅有趣的图景。因为你可以真正看到,生成这个答案时激活了 100 个概念。你知道,是加拿大,但也可能是某个特定的人、特定的地点、特定的态度。这真的能帮助你在更直观的人类层面上理解发生了什么。
Yeah. Okay, so it's part of the idea here that having a whole team of humans laboriously figure out that there's a neuron that corresponds with Canada is not very satisfying. It's not clear where we get from that. But if you could automate it such that you had the equivalent of thousands or millions of staff basically scrutinizing and trying to figure out what each part of the neural network was doing, which you might be able to do if you could automate it, then maybe that would add up to an interesting picture. Because you could really see, well, here are the 100 concepts that were activated when this answer was being generated. You know, it was Canada, but it was also a particular person in a particular place and a particular attitude, maybe. And that really would help you to understand on some more intuitive human level what was going on.
是的,完全正确。我认为这也是一个很好的方面,它让你一窥未来自动化对齐研究可能的样子。你可以大规模运行,投入大量算力,并使用各种传统能力技巧来改进它。而且,它实际执行的任务并不完全是人类之前做过的任务。我们没有雇佣一群人类来仔细检查模型中的神经元并编写解释。这从来不是一个选项,因为以前这根本没有意义。
Yeah, exactly. And I think it's also a really nice aspect of this that it gives you a glimpse of what future automated alignment research could be like. You can run this at a large scale, you can dump a lot of compute into it, and you can do various traditional capability tricks to make it better. But also, the task that it actually does is not exactly the task that a human had previously done. We didn't hire a bunch of humans who meticulously go through neurons in the model and try to write explanations. That was never an option because it never made sense before.
是否存在某个特定模型最适合或有特殊优势来解释自己?我直觉上觉得 GPT-4 在某种意义上可能最了解 GPT-4 的神经元。那么,你能看着自己的神经元并解释它们吗?好吧,但直觉来自于如果某人注意到我有很多不同的概念关联在一起,我会同时提到它们,然后有人说,加拿大、棕色和枫糖浆有什么共同点?嗯,我搞砸了那个解释。但我知道在我自己的脑海里什么东西与我相关,即使我看不到神经元。
Is it the case that a particular model is best or has a particular advantage at explaining itself? It feels intuitive to me that GPT-4 in some sense might have its best understanding of GPT-4's neurons. So, could you look at your neurons and explain them? Okay, but the intuition is coming from if someone noticed that I had a whole lot of different concepts associated for me and I would bring them up at the same time, and someone said, you know, what does Canada and the color brown and maple syrup have in common? Well, oh, I messed up that explanation. But I know what things are related to me in my own mind, even if I can't look at the neurons.
是的,这里有一个非常酷的思想实验。假设你有一个完美的大脑扫描仪,没有延迟,你可以在思考时盯着它看。这当然会是一次非常迷幻的体验,但也可能真的让你以多种方式弄清楚大脑是如何工作的,只需坐在那里尝试思考,然后观察大脑中发生了什么。那会非常疯狂。人类做不到;我们没有大脑扫描仪。但你可以用 GPT-4 真正做到这一点。
Yeah, and there's this really cool thought experiment here. Let's say you had a perfect brain scanner on your brain that was perfectly with no lag time, and you would just stare at it while you're thinking about stuff. Of course it would be a very trippy experience, but also it would probably actually let you figure out how your brain works in a bunch of ways, by just sitting there and trying to think about stuff and then seeing what happens in your brain. That would just be wild. Humans can't do that; we don't have the brain scanners. But you could literally do that with GPT-4.
是的。好吧,我想怀疑论者一定会说,我们会在细粒度层面上弄清楚这些神经元可能服务于什么功能,或者它们对应什么概念,等等。但感觉在我们可以用这些来真正判断模型是否对齐之前,还缺少进一步的步骤。你对这些进一步步骤有什么想法吗?
Yeah. Okay, I suppose a skeptic must say we're going to figure out at the granular level what functions maybe some of these neurons are serving or what concepts they correspond to, and so on. But then it feels like there are further steps missing before we can use that to really figure out whether a model is aligned. Do you have any ideas for what those further steps would be?
是的,特别是,我认为可解释性似乎非常困难。它之所以困难,是因为没有明显的理由说明模型应该使用非常类似人类的概念来思考事物。
Yeah, in particular I think interpretability seems very hard. It's hard because there's no apparent reason why the model should be using very human-like concepts to think about stuff.
概念可能就在那里,因为它们在经验上是有用的。这就是我们使用它们并指向它们的原因,所以它们很可能就在那里。有一些概念对于对齐研究特别有趣,我们想要寻找它们,比如欺骗和说谎,以及其他对于解决这个问题至关重要的东西。如果你有某种自动浮现它们的方法,我认为那将是一个巨大的胜利。我认为可解释性是一个非常好的验证技术候选。假设我们已经搞定了可扩展监督,我们有一个非常兴奋的技术,并用它来对齐模型。然后我们面临一个问题:我们做得有多好?使用同样的技术是不够的。可解释性可以介入:如果你有非常好用的工具,你可以问这个问题:你能找到任何证据表明存在欺骗性对齐、欺骗、密谋反对人类,或者试图找出如何从模型中窃取信息吗?如果我们找到了,那是一个非常糟糕的信号,我们不应该只是训练掉它。你不能针对可解释性工具进行训练;你只会让它们变得无用。但作为一种验证技术,如果你没有找到,并且你有好的技术本可以找到,那就证明模型确实像你认为的那样对齐了。同时,如果我们真的掌握了可解释性,我不知道那将如何让我们解决对齐问题。即使我们真的理解了它的工作原理,并且可以调整各种旋钮使其更对齐,也不清楚如果人类试图这样做,这条路是否会轻易成功。但也可能存在一条路径,制造一个足够对齐的人类级别的自动对齐研究员,帮助我们做到这一点,而完全不需要可解释性。我认为那也是可能的。无论我们能做什么都会有帮助,我很兴奋能尽可能走得更远。我们有这些永久完美的脑部扫描仪,不在开始时使用它们简直是疯了。
Concepts are probably somewhere in there because they are empirically useful. That's why we use them and point to them, so they are probably in there. There are some concepts that are particularly interesting for alignment research that we would want to look for, like deception and lying, and other things that are pretty critical for how we want to solve this problem. If you had some way of automatically surfacing them, I think that would be a big win. I think interpretability is a really good candidate for a validation technique. Let's say we have figured out scalable oversight, we have a technique we are really excited about, and we use it to align the model. Then we are at the question of how good a job we have done. Using the same kind of technique is not good enough. Interpretability can come in: if you have tools that work really well, you can ask the question: can you find any evidence of deceptive alignment, deception, plotting against humans, or trying to figure out how to exfiltrate inside the model? If we find that, it's a really bad sign, and we shouldn't just train it out. You can't train against the interpretability tools; you will just make them useless. But as a validation technique, if you don't find that and you have good techniques that could find it, that is some evidence that the model is as aligned as you think it is. At the same time, if we really nail interpretability, I don't know how that will let us solve alignment. Even if we really understand how it works and can fiddle with various dials to make it more aligned, it's not clear that path will easily succeed if humans try to do that. But there might also be a path to making a human-level automated alignment researcher sufficiently aligned to help us do this with no interpretability at all. I think that's also plausible. Whatever we can do will help, and I'm excited to get as far as possible. We have these permanent perfect brain scanners; it would be insane not to use them at the beginning.
关于可扩展监督,有哪些有趣的论文发表或有趣的结果出现?
Interesting papers published on scalable oversight or interesting results that have come out?
过去一年左右有很多有趣的工作。不只是我们;DeepMind 和其他人也在努力让它成功。我想谈谈我们去年做的批评工作,因为我认为有一些非常有趣的见解。基本想法是:如果我们能训练一个模型写批评,我们可以把这些批评展示给人类评估者,看看它们是否帮助评估者做出更好的决策或评估。从某种意义上说,批评是最简单的辅助形式:一次性的,不是交互式的,你只是试图指出一个缺陷。它也很容易,因为甚至不需要是一个好的或准确的批评;你只需展示一大堆,人类会扔掉他们认为不好的。但有时批评会指出人类可能遗漏的缺陷。事实上,这正是我们能够展示的。这是在 GPT-3.5 上做的实验,所以是有一段时间了。我们做了随机对照试验,人类要么得到辅助,要么没有,他们必须在摘要任务中找出问题。我们表明,来自 GPT-3.5 的批评已经帮助人类多发现了 50% 的缺陷。这项工作最有趣的一点是,我们有这种评估其效果的方法论。还有其他评估方法,例如,你可以看专家标签与帮助非专家发现缺陷。但这基本上只有在你有专家标签的情况下才有效。一般情况下,这并不成立。你想解决一个真正困难的任务,人类很难评估。例如,对于代码任务,如果你想找出模型知道的所有代码缺陷,人类找不到;人类非常不擅长找代码中的 bug。但简单的技巧是:你可以在代码中引入 bug,然后你知道哪个版本有更多 bug,因为你让它更糟了。我兴奋的是,我基本上想尝试所有被提出的可扩展监督想法,并实际测量哪个效果最好,以及它们到底有多好。想法比如递归奖励建模:如何获得人类辅助来帮助人类评估 AI 在做什么?或者辩论,其中两个 AI 就一个问题进行辩论,人类裁判决定哪个做出了更有用的陈述。或者分解,将任务分解成更小的块并尝试解决它们。或者你可以在评估中这样做。还有自动做市,你试图在辅助下最大限度地改变人类的想法。有很多这样的变体。我个人对哪个效果最好有押注,但我只是想通过经验看到结果。我认为真正令人兴奋的是,我们可以直接测量它,这比争论要好得多。
There has been a bunch of interesting work in the past year or so. It's not just us; DeepMind and others are also trying hard to make it work. I want to talk a little bit about the critiques work that we did last year because I think there are some really interesting insights. The basic idea was: if we can train a model to write critiques, we can show these critiques to human evaluators and see if they help the evaluators make better decisions or evaluations. In some sense, critiques are the simplest form of assistance: it's one-off, not interactive, and you are just trying to point out one flaw. It's also easy in the sense that it doesn't even have to be a good or accurate critique; you just show a whole bunch, and the human will throw out the ones they think are bad. But sometimes the critique will point out a flaw that the human would have missed. In fact, that's what we could show. This was experiments done on GPT-3.5, so a while ago. We did randomized controlled trials where humans either got assistance or not, and they had to find problems in a summarization task. We showed that the critiques from GPT-3.5 already helped humans find fifty percent more flaws. One of the most interesting things about this work was that we have this methodology for evaluating how well it's working. There are other ways to evaluate this too, for example, you can look at expert labels versus helping non-experts find flaws. But that fundamentally only works if you have access to expert labels. In the general case, that won't be true. You want to solve a real task that is really hard and that humans really struggle to evaluate. For example, with code tasks, if you want to find all the flaws in the code the model knows about, humans won't find those; humans are terrible at finding bugs in code. But the simple trick is: you can introduce bugs in the code, and then you know which version is more buggy because you made it worse. What I'm excited about is fundamentally I want to try all the scalable oversight ideas that have been proposed and actually measure which of them works best and how well they work. Ideas like recursive reward modeling: how can you get human assistance to help humans evaluate what AI is doing? Or debate, where you have two AIs that debate each other on a question and a human judge decides which made more useful statements. Or decomposition, where you break the task down into smaller chunks and try to solve those. Or you could do that with evaluation. There is also automated market making where you try to change the human's mind maximally with assistance. There are a whole bunch of these variants. I have my personal bets on which will work best, but I just want to empirically see the results. I think what's really exciting is that we can just measure it, and that will be so much better than arguing over it.
有很多人和你一样了解情况,他们觉得……
There are a lot of people out there who are about as informed as you who feel...
技术对齐问题可能极其困难,这样的努力成功的可能性可能很小。但你对整体情况相当乐观。过去十年有哪些进展或成果让你有这种乐观态度?
The technical alignment problem is probably extremely hard, and an effort like this probably only has a slim likelihood of success. But you're pretty optimistic about things in the scheme of it. What developments or results have there been in the last 10 years that have made you have this level of optimism?
我认为过去几年的许多发展都对对齐非常有利。大型语言模型非常有帮助,因为它们理解自然语言。它们对人类了解很多。你可以问它们在这种哲学下什么行为是道德的,它们能给出很好的解释。此外,能够与它们交谈并表达你的观点,让很多事情变得更简单。同时,它们在某种意义上是一块白板,你可以用相当少的数据微调它们,使其变得有效。如果与几年前通往 AGI 的路径相比,那时我们似乎要在像 Universe 这样的环境中训练深度强化学习智能体,Universe 是各种游戏的集合。它们可能在解决游戏方面变得非常聪明,但不一定对语言、人类如何思考道德、人类关心什么或世界如何运作有深刻理解。
I think a lot of developments over the last few years have been pretty favorable to alignment. Large language models are super helpful because they understand natural language. They know so much about humans. You can ask them what would be a moral action under this philosophy, and they can give you a really good explanation. Also, by being able to talk to them and express your views, it makes a lot of things easier. At the same time, they are in some sense a blank slate where you can fine-tune them with fairly little data to be effective. If you compare this to how the path to AGI looked a few years ago, it seemed like we were going to train deep RL agents in an environment like Universe, which is a collection of games. They might get really smart at solving games, but they wouldn't necessarily have a deep understanding of language, how humans think about morality, what humans care about, or how the world works.
另一件非常鼓舞人心的事情是我们从迄今为止尝试的对齐技术中看到的。InstructGPT 的效果比我预期的要好得多。即使在我们做 Deep Alchemy 和偏好论文时,我一开始认为我们很可能无法在有限的时间内让它工作得那么好,但它确实成功了。InstructGPT 效果非常好。在某种程度上,你可以说这些技术并不是用来对齐超级智能的,那么我为什么如此乐观?但我认为这仍然提供了证据表明这是有效的,因为如果我们连今天的系统都无法对齐,我认为我们应该更悲观。反之亦然。
The other thing that has been really encouraging is what we've seen from alignment techniques we've tried so far. InstructGPT worked much better than I ever hoped for. Even when we did the Deep Alchemy and preferences paper, I came into it thinking there was more than even chance we wouldn't be able to make it work that well in the time we had, but it did work. InstructGPT worked really well. To some extent, you could argue these are not techniques that align superintelligence, so why am I so optimistic? But I think it still provides evidence that this is working, because if we couldn't even get today's systems to align, I think we should be more pessimistic. The converse also holds.
怀疑论者可能会说,我们看到了这些模型知道我们想要什么或关心什么的进展,但也许我们还没有看到它们会关心我们所关心的事物的证据。担忧在于,模型完全知道你在要求什么,但这并不意味着它认同你的目标。它可能一直假装这样做,直到突然翻脸。我们是否看到了第二个方面的证据,即模型实际上认同我们的目标,还是这仍然是一个黑箱?
A skeptic might say we've seen improvement in our prospects of these models knowing what it is that we want or knowing what it is that we care about, but maybe we haven't seen evidence that they're going to care about what we care about. The worry is the model knows perfectly what you're asking for, but that doesn't mean it shares your goal. It could pretend to do that right up until the moment it flips out on you. Have we seen any evidence for the second thing, that the models actually share our goals, or is that still a black box?
我认为这是一个非常重要的观点,也是关于对齐可能不顺利的一些主要担忧的核心。我仍然认为模型理解我们想要什么是一个重要的第一步。主要问题变成了如何让它们在意,这正是我们试图解决的问题。但拥有第一步已经很好了。
I think this is a really important point and central to some of the main worries about why alignment might not go well. I do still think that the model understanding what we want is an important first step. The main question becomes how to get them to care, and that's the problem we're trying to figure out. But having the first step is great.
你愿意说说你认为 AI 导致非常糟糕结果的概率是多少吗?过去一年这个概率是上升了还是下降了?
Would you venture to say what your probability of a very bad outcome from AI is, and has that gone up or down over the last year?
我不认为这是一个很有用的问题,因为我觉得我的答案更多地取决于我当前的情绪,而不是世界的任何实际属性。可以肯定的是,AI 的未来可能非常好,也可能非常糟糕,而走向何方仍然悬而未决。人类对我们走哪条路有很大的因果控制权,甚至个人或单个研究人员也能对方向产生重大影响。所以这才是更值得关注的问题。如果你想要一个厄运概率,之所以这么难,是因为未来有太多不同的情景。要得到一个准确的概率,你需要在这个大空间上积分,我认为这根本没什么帮助。重要的是我们能做多少改善,以及实现这一目标的最佳路径是什么。
I don't think it's a really useful question, because I feel my answer would depend a lot more on my current mood than any actual property of the world. What's definitely true is that the future with AI could go really well or really badly, and which way it goes is still very much up in the air. Humans have a lot of causal ownership over which path we go down, and even individuals or individual researchers can have a big impact on the direction. So that's the much more important question to focus on. If you wanted a probability of doom, the reason it's so hard is because there are so many different scenarios of how the future could go. To have an accurate probability, you need to integrate over this large space, and I don't think that's fundamentally helpful. What's important is how much we can make things better and what are the best paths to do that.
我没有花太多时间试图精确确定我个人的厄运概率。我猜它大于 10%且小于 90%,所以降低这个数字非常重要,但也没有高到我们完全完蛋的程度。在这个范围内,它似乎不会对我的日常决策产生太大影响,所以我乐于就此打住。
I didn't spend a lot of time trying to precisely pin down my personal P(doom). I suppose my guess is that it's more than 10% and less than 90%, so it's incredibly important that we work to lower that number, but it's not so high that we're completely screwed. Within that range, it doesn't seem like it's going to affect my day-to-day decisions all that much, so I'm happy to leave it there.
我想我大概也会给出这个范围。但你问我为什么乐观,我想再给你更多理由。我认为最根本的是,我认为对齐问题是可处理的。如果我们专注于它并付出努力,我们实际上可以取得很大进展。有很多研究进展可以通过一个小型专注团队在一年或四年内实现。真的感觉我们有了一个真正的攻击角度,我们可以实际迭代并朝着它构建。我认为它很可能成功,这真的很疯狂、很令人兴奋。我们多年来一直在谈论这个难题,现在我们有了真正的机会去解决它。如果我们做到了,那将非常棒。
I think that's probably the range I would give too. But you asked me why I'm optimistic, and I want to give you a bunch more reasons. I think fundamentally the most important thing is that I think alignment is tractable. I think we can actually make a lot of progress if we focus on it and put effort into it. There's a lot of research progress to be made that we can actually make with a small dedicated team over the course of a year or four. It really feels like we have a real angle of attack on the problem, something we can actually iterate on and build towards. I think it's pretty likely to work actually, and that's really wild and exciting. We have this hard problem we've been talking about for years, and now we have a real shot at solving it. That would be so good if we did.
我乐观的其他一些原因:我认为从根本上说,对于许多我们关心的任务,包括对齐研究,评估比生成更容易。这就是为什么我们可以通过使用 AI 来自动化对齐研究的部分工作来获得很大的杠杆作用。特别是,如果你能……
Some of the other reasons I'm optimistic: I think fundamentally evaluation is easier than generation for a lot of tasks we care about, including alignment research. That's why we can get a lot of leverage by using AI to automate parts of alignment research. In particular, if you can...
想想经典的计算机科学问题,比如 P 与 NP 问题。这类问题本质上更容易评估。很多消费品也是如此。比如买智能手机,挑选一部好手机比制造一部手机容易得多。在组织中,招聘时,判断某人是否胜任工作必须比亲自做那份工作更容易,否则你无法决定雇佣谁,这行不通。再比如体育和游戏,如果你不知道谁赢了,比赛就不好看了。判断当前这一步棋好不好可能很难,但后来你会知道,这正是比赛的刺激之处。你不知道结果,会有这种紧张感:「这一步棋很有意思,接下来会发生什么?」但比赛结束时,看棋盘或围棋盘,你就知道谁赢了,所有人都知道。或者看足球比赛,球进了就是进球,就这么简单。我认为科学研究也是如此。有些研究成果让人兴奋,尽管人们不知道如何得出这些成果。有时我们会判断错误,但这并不意味着我们能完美地完成这项任务,只是说评估更容易。
Think about classical computer science problems like P versus NP, right? You have these kinds of problems, but it's fundamentally easier to evaluate. It's true for a lot of consumer products. If you're buying a smartphone, it's so much easier to pick a good smartphone than it is to build a smartphone. Or in organizations, if you're hiring someone, it has to be easier to figure out whether they're doing a job than to do their job, otherwise you wouldn't know who to hire. It wouldn't work. Or if you think about sports and games, sports wouldn't be fun to watch if you didn't know who won the game. It can be hard to figure out what the current move is a good move, but you'll find out later, and that's what makes it exciting. You don't know, you have this tension: 'Oh, this was an interesting move, what's gonna happen?' But at the end of the game, when you look at the chess board or the Go board, you know who won. Everyone knows. Or if you're watching a soccer game, the ball goes in the goal, it's a goal, that's it. Everyone knows. And I think it is also true for scientific research. There are certain research results that people are excited about even though they didn't know how to produce them. Sometimes we're wrong about this, but it doesn't mean that we can do this task perfectly; it's just that it's easier.
对这种方法的批评是:如果我们不知道如何解决对齐问题,那我们怎么能判断这些模型给出的建议是否有效呢?而你的观点是,评估一个解决方案的好坏或是否有效,往往比想出这个方案容易得多。这应该让我们乐观地认为,我们不一定非要自己产生所有想法。我们只需要在想法产生后判断它们是否有效,这可能简单得多。
A criticism of this approach is: if we don't know how to solve the alignment problem, then how are we going to be able to tell whether the advice that these models are giving us on how to solve it is any good? And you're saying, well, often it can be a lot easier to assess whether a solution is a good one or whether something works or not than it is to come up with it. So that should make us optimistic that we don't necessarily have to generate all of these ideas ourselves. It might be just sufficient for us to be able to tell after they've been generated whether they're any good or not, and that could be a much more straightforward.
完全正确。还有其他方面。我认为我们可以为迭代做好准备。我们可以审视当前系统,改进它们的对齐。我们可以做诸如测量是否找到了模型知道的所有漏洞之类的事情。我们可以设定这些指标。它们不会让我们一路走到对齐超级智能,但会对局部改进非常有帮助。如果你的目标是让一个系统对齐,从而帮助我们进行对齐研究,那么一个很好的试验场是:你能让 GPT-5 更对齐吗?也许你真正需要的技术目前还不适用于 GPT-5,谁知道呢。但如果你一路上没有取得进展,我认为你很难证明你正在朝着最终目标前进。同时,你需要来自现实世界的反馈信号,知道自己正在改进,正在做真实的事情。显然,你必须小心行事。你可以设置一个无关紧要的评估,这也是挑战的一部分。
That's exactly right. And there are other things. I think we can actually set ourselves up for iteration. We can stare at the current systems, improve their alignment. We can do stuff like measure whether we're finding all the bugs that the model is aware of. We can set ourselves these metrics. They're not going to take us all the way to aligning superintelligence, but they will be super helpful for making local improvements. If your goal is to align a system that could help us do alignment research, one really good testing ground is: can you make GPT-5 more aligned? Maybe the techniques that you actually need won't work that well in GPT-5 yet, who knows. But if you're not making progress along the way, I don't think you can really make the case that you're actually making progress towards the actual goal. At the same time, you need some kind of feedback signal from the real world to know that you're improving, that you're doing something real. You have to do that carefully, obviously. You can set up an eval that doesn't matter, and that's part of the challenge here.
还有其他乐观的理由吗?
Any other reasons for optimism?
我认为另一个很好的理由是:我们实际上并不是要对齐一个比我们聪明得多的系统。想象一个较笨的系统去对齐一个更聪明的系统总是很难,如果差距非常大,那看起来非常令人生畏。但我认为这也不是我们实际需要瞄准的问题,因为我们只需要瞄准人类水平或大致相当于最聪明的对齐研究人员的系统。如果你能让那个系统真正对齐,那么你就能在这个问题上取得所有进展。最初我开始研究对齐时,我并没有意识到这一点。我当时想,「哦,这个问题太难了,我们怎么解决?」但如果你瞄准这个更温和的目标,这个最小可行产品,它实际上变得可行得多。
I think the other really good one is: we're not actually trying to align the system that's vastly smarter than us. It's always hard if you picture a dumber system aligning a smarter system, and if you make the differential really large, it seems so daunting. But I think it's also not the problem that we actually realistically have to aim for, because we only have to aim for this human-level or roughly as smart as the smartest alignment researchers system. If you can make that really aligned, then you can make all the progress that you could make on this problem. Originally when I set out to work on alignment research, this realization wasn't clear to me. I was like, 'Oh man, this problem is hard, how do we do it?' But if you're shooting for this much more modest goal, this minimal viable product, it actually becomes so much more achievable.
很好。那么你能把这种方法概括为:不要纠结于能否对齐 GPT-20。我们先致力于对齐 GPT-5,然后与 GPT-5 合作,找出如何对齐 GPT-6,再然后与所有这些模型合作,共同对齐 GPT-7。
Good. So could you stylize the approach as saying: don't obsess about whether you can align GPT-20. Let's work on aligning GPT-5, and then in collaboration with GPT-5 we'll figure out how to align GPT-6, and then in collaboration with all of them, we'll work together to align GPT-7.
这就是基本思路。你想通过实证来做这件事。也许你看 GPT-5,会觉得系统还是不够聪明。我们在 GPT-4 上试了很多次,比如尝试用对齐数据微调它,试图帮助我们的研究。但没什么用。GPT-5 也可能这样,但那时就会说,好吧,我们专注于 GPT-6。但我们希望这件事发生时能做好准备,当它成为可能时我们就在那里,然后真正全力以赴。
That's the kind of basic idea. And you want to do this empirically. Maybe you look at GPT-5 and you're like, well, the system still isn't smart enough. So we tried this a whole bunch with GPT-4, like trying to help get it fine-tuning on alignment data, trying to help in our research. It wasn't that useful. That could happen with GPT-5 too, but then it'll be like, okay, let's focus on GPT-6. But we want to be on the ball when this is happening, and we want to be there when this becomes possible, and then really go for it.
好了,这些是乐观的理由。我想讨论一些反对意见,或者可能不如预期的方式。我见过很多人提到的一个问题是:你怎么知道你是否成功了?你可能觉得这有效,但你如何真正有信心?特别是如果存在成功的欺骗,你可能会被麻痹,产生虚假的安全感。你怎么看?怎么才能知道?
Okay, so that's a bunch of reasons for optimism. I want to go through a couple of objections or ways that this might not work out as hoped. One that I've seen a lot of people mention is just how are you going to be able to tell whether you're succeeding? You might think that this is working, but how would you ever really have confidence? Especially if there's successful deception going on, then you could be lulled into a false sense of security. What do you think about how could you tell?
这是核心问题之一:如何区分欺骗性对齐的系统与真正对齐的系统?这是我们试图解决的挑战。这就是为什么我们在研究能否让模型告诉我们它知道的所有漏洞,这也是为什么我们要训练欺骗性对齐的模型,看看它们能否通过我们的评估,并对我们的方法进行压力测试,真正深入探究模型内部发生了什么。我认为我们可以从这个问题中学到很多,真正界定和理解剩余的风险,或者那些我们最不确定它可能如何欺骗我们的领域。
This is one of the central problems: how do you distinguish the deceptively aligned system from the truly aligned system? This is the challenge we're trying to figure out. This is why we're looking at can we get the model to tell us all the bugs that it's aware of, and this is why we want to train deceptively aligned models to see if they can pass our evals, and stress-test our methods, and really drill into what's going on inside the model. I think we can learn so much about this problem and really scope and understand the risks that remain, or the areas where we are most uncertain about how it could deceive us.
所以它可能在第一步就失败,也许你试图合作的那个第一个模型没有对齐,但你没有意识到,于是它开始把你引向一条坏路。然后某个时候事情会变得很糟。但问题从一开始就存在。然后你也可能一开始很好,但后来无法判断进一步的迭代是否朝着正确方向。问题可能悄悄出现而你却没有注意到,这也会把你引向坏路。听起来你只是在说这是我们必须解决的问题。
So it could fail at the first step, perhaps where the first model that you're trying to collaborate with isn't aligned but you don't realize that, and so it just starts leading you down a bad path. Then at some point things will go badly. But ultimately the problem was there at the very beginning. And then you could also start out well but then not be able to tell whether further iterations are going in the right direction. Problems could creep in and you're not noticing them, and that could lead you down a bad path. It sounds like you're just saying this is the problem we have to solve.
完全正确。我认为从根本上说,我更担心的问题是:我们能否真正精确地知道系统对齐的程度?而不是如何让它更对齐。因为我认为很多风险来自于对系统实际对齐程度的不确定性。也就是说,我认为没有人会兴奋地部署一个你知道是未对齐的、想要接管世界的系统。但如果你能精确测量系统真正对齐的程度,或者你对自己的测量工具(试图理解模型对齐程度)有信心,那么我认为你实际上已经解决了很大一部分问题。因为那时你知道自己在哪里,并且可以更容易地研究改进对齐的方法。但你必须小心做法,以免过拟合测试集。我认为从根本上说,很多问题在于确切知道你在哪里。
Exactly. I think fundamentally, the thing I'm much more worried about is the question: can we really precisely know how aligned the system is? Rather than the question of how can we make it more aligned. Because I think a lot of the risks come from uncertainty about how aligned the system actually is. In the sense that I don't think anyone will be excited to deploy a system that you know is misaligned and that wants to take over the world. But if you can precisely measure how aligned the system truly is, or if you're confident in your measurement apparatus that tries to understand how aligned the model is, then I think you've actually solved a large part of the problem. Because then you know where you're at, and you can much more easily work on methods that improve alignment. But you have to be careful the way you do it so you don't overfit to the test set. I think fundamentally, a lot of the problem is knowing exactly where you are.
观众中有人问:你计划如何在第一次关键尝试之前提前验证,AI 提出的对齐方案能一直扩展到超级智能,并且不包含意外或故意的弱点?如果它包含弱点怎么办?我想人们非常担心,如果这行不通,那会很可怕。这是一个真正高风险的问题,这也是为什么它如此重要。
Someone from the audience had this question: how do you plan to verify ahead of time, before the first critical try, that the alignment solution proposed by AI scales all the way to superintelligence and doesn't include accidental or intentional weaknesses? And what happens if it does? I guess it's just people are very nervous that if this doesn't work out, it's pretty scary. It's a really high-stakes problem, and that's what makes it so important to work on.
我认为如果你有一个心理画面,我们有一个自动化的对齐研究员,我们按下一个按钮,它说这是你应该做的,然后我们就照做并希望最好,这真的太简化了。我不认为系统做的第一件事就是对齐超级智能。我认为我们只会对齐 GPT-N+1,而且我们会非常参与其中,查看所有结果,我们会发布它并向其他人展示,然后问:你觉得这个结果怎么样?你认为这是个好主意吗?我们应该这样做吗?同时,我们会有所有这些其他工具。我们希望能有更好的可解释性,我们会更好地理解模型的鲁棒性,或者我们会有很多自动化工具来监控系统进行对齐研究的过程。所有这些自动化工具都会在一旁观察,试图理解发生了什么。或者如果我们能在基础层面上真正理解泛化,我们能否拥有一个系统,让我们更有信心它按照人类真正想要的方式泛化,而不是我们声称想要的方式或我们可以检查的方式?如果我们从根本上理解这些问题,或者在改进这些方向上做得很好,我认为我们将有更多的证据和更多的理由相信系统实际上在做正确的事情,或者不是。而这正是我们试图弄清楚的。
I think it's really oversimplified if you have a mental picture where we have this automated alignment researcher, we press a button, it says here's what you should do, and then we just do it and hope for the best. I don't think that's the first thing the system does is align superintelligence. I think we'll just align GPT-N+1, and it'll be like we'll be very in the loop and looking at all the results, and we'll publish it and show it to others and be like, what do you think about this result? Do you think this is a good idea? Should we do that? And at the same time, we'll have all these other tools. We'll hopefully have much better interpretability, we'll understand robustness of our models much better, or we'll have a lot of automated tools to monitor as the system is doing its alignment research. All these automated tools will be looking over the shoulders and trying to make sense of what's going on. Or if we can really understand generalization on a fundamental level, can we have a system that we are much more confident generalizes the way humans would actually want, and not the ways that we would say we want or ways that we can check? If we fundamentally understand these problems or do a good job at improving in these directions, I think we'll just have so much more evidence and so much more reasons to believe the system is actually doing the right thing or it's not. And that's what we're trying to figure out.
这个项目的公告说,我们现在不知道如何对齐超级智能,如果我们部署了超级智能而没有好的对齐方法,那可能是绝对灾难性的。如果四年后你认为你还没有解决这个问题,或者八年后、十年后,你只是说,我们一直在努力,取得了一些进展,但我不相信我们接近能够对齐超级智能,但能力已经大幅前进,我们可能接近部署那种如果不对齐你会非常担心部署的东西。如果你和你的团队认为这是个坏主意,有没有计划如何推迟部署?
The announcement of this project says we don't know how to align superintelligence now, and if we deployed superintelligence without having a good method for aligning it, then that could be absolutely disastrous. What happens if in four years time you think that you haven't solved the issue, or in eight years time or 10 years time, you're just like, well we've been working at it, we've made some progress, but I don't have confidence that we're close to being able to align superintelligence, but the capabilities have really gone ahead and we might be close to deploying the kind of thing that you would be really worried about deploying if it weren't aligned. Is there a plan for how to delay that deployment if you and your team just think it's a bad idea?
我认为在那个阶段最重要的是我们必须非常诚实。
I think the most important thing at that stage is we just have to be really honest.
就我们目前所处的位置而言,我认为世界在某种程度上需要,也会要求我们诚实。不只是说出我们完全相信的东西,还要展示我们拥有的所有证据。如果到了能力非常强大,但我们的对齐方法还跟不上的时候,那你就真的需要提出:嘿,我们都该冷静一下。这不只是 OpenAI 的问题,对吧?到了这个地步,你得把所有 AGI 实验室召集起来,想办法解决这个问题,或者分配资源,放缓能力发展。我不知道会发生什么,但我认为前提仍然是:你得弄清楚自己在对齐方面处于什么位置。我们仍然必须非常努力地尝试解决这个问题,这样才能说:看,我们非常努力了,这是我们尝试过的所有方法,这是结果,你们可以仔细查看。如果你们看了所有这些,很可能得出和我们一样的结论,那就是:我们认为我们还没到那一步。这就是为什么我说我们只需要非常诚实地面对它。这也是为什么我们做出这样的承诺:我们希望广泛分享我们努力的成果。我们希望其他人的模型也能得到对齐。我们希望每个构建强大 AI 的人都能让它与人类对齐。我们想告诉其他人我们发现的关于如何做到这一点的所有事情。
Where we are, in some ways I think the world just needs, it will demand us to be honest. Not just say what we totally believe, but also show all the evidence we have. If you get to this point where the capabilities are really powerful, but at the same time our alignment methods are not there, this is when you really make the case for: hey, we should all chill out. This isn't primarily about OpenAI, right? At this point, you gotta get all the AGI labs together and figure out how to solve this problem, or allocate resources, slow down capabilities. I don't know what will happen, but I think the prerequisite is still: you gotta figure out where you're at with alignment. We still have to have tried really hard to solve the problem in order to be able to say: look, we tried really hard, here's all the things we tried, here's the results, you can look at them in detail. If you look at all of this, you would probably come to the same conclusion as us, which is: we don't think we're there yet. That's why I'm saying we just need to be really honest about it. And this is why we also make this commitment: we want to share the fruits of our effort widely. We want everyone else's models to be aligned too. We want everyone who's building really powerful AI to have it aligned with humanity. We want to tell other people all the things we figure out about how to do this.
是的。这也是为什么我们做出这样的承诺:我们希望广泛分享我们努力的成果。我们希望其他人的模型也能得到对齐。我们希望每个构建强大 AI 的人都能让它与人类对齐。我们想告诉其他人我们发现的关于如何做到这一点的所有事情。我仍然担心各种不同的方式,你可能取得一些进展但没能完全到位,但人们最终还是可能部署。所以我想我担心的一个问题是,你可能会过于自信,爱上自己的工作,觉得已经成功解决了这个问题,但实际上并没有。另一件事可能是,你会对其他人说,OpenAI,我们觉得我们还没有解决这个问题。是的,我对此非常害怕,但他们不听你的,因为可能有一些商业原因或内部政治阻碍了帮助。另一种情况是:真正的 OpenAI 听你的,但世界其他地方不听,结果别人部署了。我不想把整个宇宙的重量压在你肩上,但你对这些不同的可能失败模式有什么看法吗?
Yeah. And then this is why we also make this commitment: we want to share the fruits of our effort widely. We want everyone else's models to be aligned too. We want everyone who's building really powerful AI to have it aligned with humanity. We want to tell other people all the things we figure out about how to do this. I still be worried about various different ways that you can make some progress but not get all the way there, but then people could end up deploying anyway. So I guess one concern I will have is that you might be overconfident, so you might fall in love with your own work and feel like you've successfully solved this problem when you haven't. Another thing would be maybe you'll say to other people, OpenAI, we don't feel like we've solved this issue. Yeah, I'm really scared about this, but then they don't listen to you because maybe there's some commercial reasons or internal politics that prevents it from helping. Another method would be: well, the real OpenAI listens to you, but the rest of the world doesn't, and someone else ends up deploying it. I don't want to heap the weight of the universe on your shoulders, but yeah, do you have any comments on these different possible failure modes?
是的,我的意思是,我认为这就是为什么我们要建立我们需要的治理机构来做好这件事。我不认为最终会由我来决定:现在是否安全?我们在 OpenAI 内部进行安全审查,然后模型才能发布。OpenAI 董事会对是否部署拥有最终决定权。如你所知,OpenAI 有复杂的非营利结构,非营利董事会实际上最终负责 OpenAI 的决策。所以他们可以决定:我们不部署,即使有商业理由。而对于整个世界来说,最终它会影响每个人,政府必须以某种方式参与进来。我们需要像国际原子能机构那样为 AI 设立一个机构,能够以技术为基础帮助做出这类决策。这就是为什么我认为我想做的以及我们想通过超级对齐做的事情是聚焦技术挑战,真正了解我们所处的位置,同时也在问题上取得实际进展,并非常努力地专注于真正解决它。
Yeah, I mean, I think that's why we want to be building the governance institutions that we need to get this right. I don't think at the end of the day it'll be up to me to decide: is this now safe to go or not? We are doing safety reviews internally at OpenAI before a model goes out. There's the OpenAI board that has the last say over: is this OpenAI going to do this or not? As you know, OpenAI has this complicated non-profit structure, and the non-profit board is actually in charge of what OpenAI does ultimately. So they can just decide to make the call of: we're not deploying, even though there's a commercial reason to. And then for the world in general, at the end of the day, it can affect everyone, and governments have to get involved somehow. We need something like an International Energy Agency for atomic energy for AI, that can help make these kinds of decisions in a technically grounded way. That's why I think the kind of things that I want to do and that we want to do with superalignment is to zoom in on the technical challenges, really understand where we are, but also actually make progress on the problem and try really hard and focus on actually solving it.
是的。一个我没想到但在我阅读这些方法时想到的反对意见:会不会实际上自我逃逸更容易?模型逃离实验室并做一些非常糟糕的事情,比如释放火器或发明生物武器并释放它们造成巨大破坏,这会不会更容易?这实际上可能比对齐 AI 更容易的技能。所以我们可能在那之前就达到了那种能力,即造成巨大破坏的能力,而这些模型实际上对你和你的团队在取得对齐进展方面非常有帮助之前。
Yeah. An objection that I don't think I've seen, but one that occurred to me when I was reading about the approaches: could it be the case that it's actually easier to self-exfiltrate? That it's just kind of broke for a model to break out of the lab and do something really bad, like release fire weapons or invent bioweapons and release them and cause an enormous amount of damage. That could actually be an easier skill than aligning AI. And so we might possibly hit that capability, the capability to do a ton of damage, before these models are actually very helpful to you and your team in making progress on alignment.
是的,我认为自我逃逸是需要关注的关键能力之一,因为系统在实验室的数据中心里,我们可以控制它——我们可以关闭数据中心、停止引擎、删除快照——和它在外面试图自我维持或试图构建更好的 AI 模型之间有一个非常重要的区别。所以问题就变成了:你如何衡量模型能否逃逸?它能否引入安全漏洞或利用我们基础设施中存在的安全漏洞?现在我不能做到,但未来的模型可能可以。或者说服 OpenAI 员工帮助它逃逸权重——这是另一条路径:你试图说服人类,提出一些对他们来说可信的论点,说明为什么他们应该这样做。这可能相当困难。我不知道,GPT-4 做不到,但未来的模型可能可以。所以我认为关注这一点是一个非常重要的区别。然后回到你的问题:如果这先发生怎么办?我认为在某种程度上,你可以通过传统的安全措施让自我逃逸更难,但在某个时候,这将是一个对齐问题,你必须证明系统没有试图逃逸,它不想逃逸。我认为总体上技术如何发展以及哪些能力会先被解锁存在很多不确定性。但我相当乐观,我们会在这种事情发生之前从模型中获得很多非常有用的东西。但当然,这就是为什么我们需要衡量这一点,因为我们不能只是胡乱猜测。
Yeah, I think self-exfiltration is one of the really key capabilities to be looking at, because there's a really important difference between the system being at the lab in our data center in a way that we can control it—we can turn off the data center, spin down the engine, delete the snapshot if we want to—and whether it's out in the world trying to sustain itself or trying to build better AI models. So the question becomes: how can you measure whether the model can break out? Can it introduce security vulnerabilities or exploit security vulnerabilities that exist in our infrastructure? Right now I can't do that, but future models could. Or kind of persuade an OpenAI employee to help it exfiltrate its weights—that's the other path: you just try to persuade humans, come up with some arguments that are believable to them why they should do that. Could be pretty hard. I don't know, GPT-4 can't do this, but future models might. So I think looking at this is a really important distinction. And then going to your question: what if this happens first? I think to some extent you can make self-exfiltration harder by just traditional security measures, but at some point this will be an alignment problem where you actually have to show that the system is not trying to break out, it doesn't want to. I think there's a lot of uncertainty in general over how the technology will go and what kind of abilities will be unlocked first. But I'm pretty optimistic that we will get a lot of really useful stuff out of the models before this kind of thing can happen. But of course, that's why we need to measure this, because we can't just make some wild guesses.
是的,好的。所以这些是我在网上读到的一些反对意见,还有一个是我自己的。但我想知道:如果你扮演魔鬼代言人,你会怎么说?
Yeah, okay. So those are some objections I've read online and one for me. But I guess I'm curious to know: what if you were playing Devil's Advocate?
对你正在采取的整个方法,从多个层面来看,最好的反对意见是什么?我认为你可以反驳说,自动化对齐研究会来得太晚,无法真正帮助我们,就像你提到的。对吧?我们必须解决很多问题,而且在某种程度上,如果这是真的,我们可能还是会做现在正在做的事情,那就是努力取得更多对齐进展,以便我们能够对齐更强大的系统。在某种意义上,这也意味着你实际上是在提高第一个灾难性未对齐系统的门槛,例如。我认为在如何构建研究组合方面,你可以提出更详细的反对意见,比如我们特别感兴趣的具体路径,比如可扩展监督、泛化、鲁棒性、对抗性测试之类的东西,还有可解释性。我们可以深入探讨每条路径的细节,以及我认为每条路径的最佳反对意见是什么。然后你还可以说,你为什么要在 AI 实验室做这项工作?对吧?你难道不会面临一些相互竞争的激励吗?就像你提到的,实验室想要部署,你如何与尽可能对齐的目标协调?
Best argument against this whole approach that you're taking, in your opinion, at a bunch of different levels. I think you could object that automated alignment research will come too late to really help us, as you mentioned. Right? Like we have to solve a lot of the problems, and in some extent if that's true, we still probably gonna do the same things we're doing now, which is just trying to make more alignment progress so that we can align more capable systems. And in some sense, that also means that you're kind of raising the bar for the first catastrophically misaligned system, for example. I think there's more detailed objections you could make on how we build a research portfolio, like the particular path that we're excited about, like scalable oversight, generalization, robustness, adversarial testing, that sort of stuff, interpretability. And we can go into details of each of these paths and what I think the best objections are to each of them. And then you can also say, why are you doing this job at an AI lab? Right? Aren't you gonna face some competing incentives, like you mentioned with, oh but the lab wants to deploy, and how do you square that with wanting to be as aligned as possible?
我认为从根本上说,AI 实验室是做这项工作的最佳地点之一,因为你离技术如此之近,你能看到它正在被开发。对吧?我们在 GPT-4 发布之前就尝试了很多事情,而且我们真的,因为我们在对齐方面亲力亲为,我们确切知道我们在哪里,弱点是什么,什么真正有效。我认为这非常有用。我还认为 AI 实验室资源非常丰富,他们有动力在齐上投入,而且他们不应该吗?这很好。是的,我认为我不认同那个反对意见。这让我想起一个问题,你为什么抢银行?他说,因为钱在那里。我觉得,你为什么要在 AI 实验室做对齐研究?因为那里是所有前沿研究的地方,那里有最前沿的模型。是的,虽然这个论点不言自明。我是说,我不认为 OpenAI 是唯一能做好的对齐工作的地方,对吧?还有很多其他地方也在做好的对齐工作。但我认为很明显它有巨大的优势。是的,我不是说每个人都必须在 OpenAI 或某个实验室工作;你在其他地方也能做事情。但肯定有些人应该在实验室。
I think fundamentally, AI labs are one of the best places to do this work just because you are so close to the technology, you see it as it's being developed. Right? We got to try a lot of things with GPT-4 before it came out, and we really, because we were hands-on at aligning it, we know exactly where we're at, what are the weaknesses, and what actually works. I think that's pretty useful. I think also AI labs are really well resourced, and they have an incentive to spend on alignment, and they shouldn't? It's great. Yeah, I think I don't share that objection. It reminds me of the question, why do you rob banks? And he says, that's where the money is. I feel like, why would you do alignment research at an AI lab? That's where all the cutting-edge research is, that's where the cutting-edge models are. Yeah, though the case kind of makes itself. There's like, I mean, I don't think OpenAI is the only place to do good alignment work, right? There's lots of other places that do good alignment work. But I think it's really clear it has some big advantages. Yeah, I'm not saying everyone should necessarily work at OpenAI or one of the labs; there are things you can do elsewhere. But surely some people should be at the labs.
也许处理这个关于最大弱点或最佳反对意见的问题的一个好方法是,如果你不能采取这种方法,而超级对齐团队必须采取一种完全不同的方法来解决这个问题,你心里有没有第二喜欢的选项?
Maybe a good way of approaching this question of the biggest weaknesses or the best objections is, if you couldn't take this approach and the superalignment team had to take a quite different approach to solving this problem, do you have a second favorite option in mind?
是的,我认为,而且明确地说,我认为我们总的路径和方法会在四年内发生变化,随着我们学到更多,我们可能会增加更多的研究领域,也可能放弃一些其他的。我认为这是研究的自然过程。我想稍微修改一下你的问题,因为我认为现在我们正在做的是我最兴奋的事情,用于对齐人类级别的系统。我认为在其他方面,我期待看到世界上我们还没有做的事情,有很多关于评估语言模型的工作我们还没有做。比如,你能测量自我泄露的能力吗?如果你能获得更多这样的能力,那将非常有用。我认为有很多关于较小模型或开源模型的可解释性工作可以做,你可以取得很大进展并获得很好的见解。我们没有做那些,因为我们的比较优势是与最大的模型合作。这就是为什么我们专注于自动化可解释性研究,这就是为什么我们试图探究 GPT-4 的内部并看看我们能发现什么。我认为这是我们擅长的事情。我仍然相信在对齐方面有一些有趣且有用的理论工作,比如数学理论工作。我认为这非常困难,因为我们没有对问题有一个很好的范围界定,我认为这可能是迄今为止最难的部分。但最终,也许问题的反面是,我们在 OpenAI 有什么优势?对吧?就像,这里有最大的模型,押注于此,利用大量算力来解决问题,在小团队中工作,紧密合作,但不专注于发表论文本身。我们不会写很多论文,对吧?我们试图非常努力地推动解决问题的特定方面,然后当我们发现有趣的东西时,我们会写下来并分享。但如果论文不多,也没关系;那不是我们试图做的。所以我们的另一个重点是大量关注工程,如果我们想进行实证实验,我们想弄清楚,你想尝试很多事情然后测量结果,这需要在大代码库上进行大量工程,因为我们使用这些巨大的模型。我们并不总是使用它们,对吧?你可以在较小的模型上运行很多有趣的实验。但归根结底,相当一部分工作是机器学习工程,这也是我们擅长的事情。
Yeah, I think, and to be clear, I think our general path and approach will change over the four years, and we'll probably add more research areas as we learn more, and maybe we give up on some other ones. I think that's the natural course of research. I kind of want to modify your question a little bit because I think right now we're doing the things I'm most excited about for aligning human-level systems. I think in terms of other things I'm excited to see in the world that we're not doing, there's a lot of work to be done on evaluating language models that we are not doing. Like, can you measure the ability to self-exfiltrate, for example? It'll be super useful if you get more of that. I think there's a lot of interpretability work on smaller models or open-source models that you can do where you can make a lot of progress and have good insights. We're not doing that because our comparative advantage is to work with the biggest models. That's why we're focusing on automated interpretability research, that's why we are trying to poke at the internals of GPT-4 and see what we can find. I think that's something we're well positioned to do. I also still have conviction that there's interesting and useful theory work, like mathematical theory work, to be done in alignment. I think it's really hard because we don't have a really good scoping of the problem, and I think that's probably the hardest part by far. But I think ultimately, maybe the reverse of the question is, what are the things that we have an advantage at doing at OpenAI? Right? And this is like, here's the biggest models, like go bet on that, leverage a lot of compute to solve the problem, work in small teams, work closely together, but don't focus on publications per se. We're not writing a lot of papers, right? We're trying to push really hard to solve particular aspects of the problem, and then when we find something interesting, we'll write it up and share it. But if it's not a lot of papers, it's fine; that's not what we're trying to do. And so another focus that we have is we focus a lot on engineering, where if we want to run empirical experiments, we want to figure out, you want to try a lot of things and then measure the results, and that takes a lot of engineering on large code bases because we are using these giant models. We're not always using them, right? There's a lot of interesting experiments you can run on smaller models. But at the end of the day, a fair amount of the work is ML engineering, and that's something that we're well positioned to do as well.
有没有什么方式让这个计划可能失败,让你夜不能寐,而我们还没有提到,值得指出的?
Is there any way that this plan could not work out that keeps you awake at night that we haven't already mentioned that's worth flagging?
哦,天哪,有太多原因了。我认为,我的意思是,比如,如果可扩展监督实际上不起作用,或者我们无法弄清楚如何让它工作,或者我们是否真的在测量正确的东西?我认为这也是我一直在脑子里盘旋的事情,比如我们如何改进我们测量的东西?例如,对于自动化可解释性,我们有一个评分函数,试图衡量神经元解释的好坏,但它是由模型近似的,并没有真正使用人类。你不会只想优化那个函数;我不认为你会得到你想要的东西。在某种程度上,那是……
Oh man, there's so many reasons. I think, I mean, there's like, what if scalable oversight doesn't actually work, or we can't figure out how to make it work, or are we actually measuring the right thing? I think that's also a lot of thing I keep circling in my head, like how can we improve what we're measuring? For example, with automated interpretability, we have this score function that tries to measure how good the explanation of the neuron is, but it's approximated with a model, it's not actually using a human. And you wouldn't want to just optimize that function; I don't think you would get what you were looking for. And to some extent, that's...
对齐问题的核心是如何找到正确的指标,那个你实际上可以优化的指标。这是我非常担心的事情。另外还有,我们是否在做正确的研究押注?我们应该更多地投资这个领域,还是减少投资另一个领域?事情可能会出错。所以当这些模型给你研究想法时,它们试图帮助你,似乎你需要让很多人参与进来检查这些工作,确保它们合理,交叉检查是否存在欺骗等等。这似乎会消耗很多人力。项目会不会仅仅因为你的全职员工不够,没有足够的人手跟上而失败?
The core of the alignment problem is how do you find the right metric, the metric that you can actually optimize. This is something I worry a whole lot about. And then there's also just, are we making the right research bets? Should we be investing in this area more, should we invest in this other area less? Things can go wrong. So at the point where these models are giving you research ideas, they're trying to help you out, it seems like you need to have a lot of people in the loop somehow checking this work, making sure that it makes sense, cross-checking for deception and so on. It seems like it could just absorb a lot of people doing that. Would it be possible that the project could fail just because you don't have enough FTEs, you don't have enough people working on it in order to keep up?
是的,我们现在确实在大力招聘,我认为团队在四年内会增长不少。但最终,我们真正的扩展方式是使用 AI,凭借算力承诺,我们可以拥有数百万的虚拟全职员工。这不是超级对齐团队在人力上能实际达到的规模。所以这就是为什么我们如此大力押注算力,如此依赖那条路径。
Yeah, I mean we're really trying to hire a lot right now, and I think the team will grow a fair amount over the four years. But I think ultimately, the real way for us to scale is using AI, where with the compute commitment we could have like millions of virtual FTEs if you so want. That's not a size that the superalignment team could ever realistically grow in terms of humans. So that's why we want to bet so heavily on compute and bet so heavily on that kind of path.
但如果 AI 员工与人类员工的比例达到一百万比一,难道不会失控吗?你信任人类的对齐,尽管他们在其他方面更差,所以他们负责最终检查事情是否失控,或者坏主意是否通过,当然是在他人的协助下。你明白我的担忧吗?
But if you've got a ratio of a million AI staff to one human staff member, isn't it possible for it to lose control? You kind of trust the alignment of the humans even though they're worse in other ways, so they're the ones who are doing some ultimate checking that things haven't gone out of control or that bad ideas aren't getting through, admittedly with assistance from others. Do you see what I'm worried about?
没错,但这正是你要解决的问题,对吧?我们会有大量工作进行,我们必须弄清楚哪些是好的,其中是否有可疑之处,哪些结果我们真正应该关注,等等。你如何解决这个问题?如何让可扩展监督发挥作用,以便你能够信任你正在监督的大量虚拟员工?或者如何改进泛化,使它们能泛化到做正确的事,而不是做人类不会注意到的事情?
Exactly, but this is the problem you're trying to solve, right? We have a large amount of work that will be going on, and we have to figure out which of it is good, is there something shady about any of it, what are the results that we should actually be looking at, and so on. How do you solve this problem? How can you make scalable oversight work so that you can trust this large amount of virtual workers that you're supervising? Or how can you improve generalization so that they will generalize to do the right thing and not do the thing that the human wouldn't notice?
它最终会不会变成一种金字塔结构:一个人,下面有一组智能体由他们监督,再下一层有另一组智能体做另一种工作并向上汇报,然后一层层下去?这是让它扩展的一种方式吗?
Does it end up becoming a sort of pyramid structure where you've got one person, and then they've got a team of agents just below that they supervise, and then there's another team of agents below at the next management level down who are doing another kind of work that are reporting upwards, and then you have layers below? Is that one way of making it scale?
是的,你可以尝试一个更传统的公司结构。我不认为那会是实际的样子。可能,而且我们从机器学习中学到的一件事是,系统通常非常擅长某些任务,而在其他任务上不如人类,所以你会优先委托前一类任务。另外,我认为它的组织方式不会像人类组织自己的方式,因为我们的组织是为我们如何协作而量身定制的。但这些都是非常好的问题,我们需要思考并弄清楚。
Yeah, you could try to have a more traditional looking company. I don't think that's literally how it's going to go. I think probably, and also one thing we've learned from machine learning is that systems are often just really good at some tasks and worse than humans at other tasks, so you would preferentially want to delegate the former kind of tasks. Also, I don't think the way it will be organized will look like the way humans organize themselves, because our organizations are tailored to how we work together. But these are all really good questions, questions that we need to think about and have to figure out.
所以你和你的团队会尽最大努力,但可能不会成功。如果你没能解决这个问题,而我们只顾推进能力,那么最终结果可能是所有人都死了。在这种情况下,人类似乎应该有一些后备计划,希望有几个后备计划,至少这样全世界的重担就不会只压在你肩上,你晚上也能睡个觉。你希望我们有什么样的后备计划?你有什么想法吗?
So you and your team are going to do your absolute best with this, but it might not work out. And if you don't manage to solve this problem and we just barrel ahead with capabilities, then the end result could conceivably be that everyone dies. In that situation, it seems like humanity should have some backup, a backup plan, hopefully several backup plans, if only so that the whole weight of the world isn't resting on your shoulders and you can get some sleep at night. What sort of backup plan would you prefer us to have? Did you have any ideas there?
我认为还有很多其他类型的计划。还有所有其他的对齐团队,比如 Anthropic 和 DeepMind,他们也在尝试解决类似的问题。有各种方法可以尝试争取更多时间,或者建立各种治理结构来管理 AI 并确保其被有益地使用。我认为解决对齐的核心技术挑战至关重要,但我们不会是唯一的力量。我们仍然需要确保 AI 与某种民主价值观对齐,而不是由科技公司单方面决定。我们还需要应对 AI 的滥用问题。对齐的系统如果可能的话不会让自己被滥用,但问题仍然是如何将其融入社会的更大背景中。例如,作为人类,你可能在一个你并不真正了解其工作的组织工作,而它实际上在作恶,你却看不到。或者,仅仅因为我们可以对齐 OpenAI 的模型,并不意味着其他人不会构建未对齐的 AI。你如何解决这个问题?这似乎非常重要。你如何确保 AI 不会差异化地赋能已经强大的人,同时也能帮助边缘群体?这似乎非常重要。最终,你还需要避免结构性风险:假设我们解决了对齐问题,每个人都构建了一个与自己高度对齐的系统,但结果却是你加速了现有的资本主义体系,公司变得非常擅长最大化股东回报,因为它们将 AI 对齐于此,但人类却被抛在一边,因为这并不涵盖你重视的其他事物,比如清洁空气。我们已经看到了早期迹象,比如全球变暖正在发生,尽管我们知道根本问题。
I think there's a lot of other kinds of plans. There are all the other alignment teams like Anthropic and DeepMind trying to solve a similar problem. There's various ways you could try to buy more time, or various governance structures you want to put in place to govern AI and make sure it's used beneficially. I think solving the core technical challenges of alignment are going to be critically important, but we won't be the only ones. We still have to make sure that AI is aligned with some notion of democratic values, not something that tech companies decide unilaterally. We still have to do something about misuse from AI. Aligned systems wouldn't let themselves be misused if they can help it, but there's still a question of how it fits into the larger context of society. For example, as a human you can be working for an organization that you don't really understand what it does, and it's actually nefarious without you being able to see that. Or just because we can align OpenAI's models doesn't mean somebody else builds unaligned AI. How do you solve that problem? That seems really important. How do you make sure that AI doesn't differentially empower people who are already powerful but also helps marginalized groups? That seems really important. And ultimately, you also want to avoid structural risks where, let's say we solve alignment and everyone makes a system that's really aligned with them, but what ends up happening is that you kind of turbocharge the existing capitalist system where corporations get really good at maximizing their shareholder returns because that's what they align AI to, but then humans fall by the wayside because that doesn't encompass all the other things you value like clean air. We've seen early indications of this, like global warming is happening even though we know the fundamental problem.
我们所做的所有经济活动仍然在推动它向前发展。所以即使我们把所有这些事情都做对了,我们仍然可能陷入一个对人类不利的系统,即使参与系统的没有人希望这样。
All the economic activity that we do still drives it forward. So even though we do all of these things right, we might still get into a system that ends up being bad for humans, even though nobody who participates in the system wants it that way.
所以你要做好你的工作,但其他很多人也要做好他们的工作。没错。这个产品生态系统有很多事情要做。我们需要让未来顺利发展,这需要很多部分,这只是其中之一。现在我们跳到一些观众提问,正如我所说,这次的问题特别多且尖锐。这些问题可能会有点跳跃,但我想把它们抛给你会让我们很好地了解大家在想什么。
So you're going to do your job, but a lot of other people have also got to do their jobs. That's right. This product ecosystem, there's a lot to do. We need to make the future go well, and that requires many parts, and this is just one of them. Let's skip now to some audience questions, which as I said were particularly numerous and spicy this time around. These questions are probably going to jump around a little bit, but I think throwing these at you will give us a good impression of what's on people's minds.
好的,来吧。
Yeah, let's do it.
好的,第一个问题:为什么 OpenAI 不先尝试用 GPT-4 解决对齐问题?例如,在冒着更先进模型带来灾难的风险之前,先让 GPT-4 达到零越狱攻击成功的状态。
Okay, first one: Why doesn't OpenAI try and solve alignment with a GPT-4 first? For example, get it to the point where there are zero jailbreaks that work with GPT-4 before risking catastrophe with more advanced models.
好问题。在某种程度上,你可以指出对齐尚未完全奏效的所有方面——越狱是其中之一,还有幻觉,系统会编造东西,这是一种我们不希望模型有的撒谎形式。但我认为,在某种程度上,把那些做得很好并不一定能帮助我们解决对齐超级智能时需要解决的核心问题。我不是说我们应该停止研究那些,但我们还需要做前瞻性的工作。特别是,我希望看到尽可能全面的对齐进展。所以当 GPT-5 出现时,或者随着模型能力越来越强,我们能有准备好的东西,并大大帮助解决这类问题。
A great question. To some extent, the fact that you can point to all the ways that alignment doesn't quite work yet—jailbreaks is one of them, but also hallucinations, the system just makes up stuff, a form of lying that we don't want in the models. But I think to some extent, getting really good at that wouldn't necessarily help us that much at solving the core problems we need to solve when aligning superintelligence. I'm not saying we should stop working on those, but we also need to do the forward-looking work. In particular, the thing that I want to happen is I want there to be the most alignment progress across the board as possible. So when GPT-5 comes around, or as models get more capable, we have something that's ready to go and helps a lot with those kinds of problems.
好的,另一个问题:GPT-4 比 GPT-3.5 更对齐的事实是否意味着模型能力越强,它就会越对齐?我知道不是每个人都会接受这个前提,但你会怎么说?
Okay, I got another question: Does the fact that GPT-4 is more aligned than GPT-3.5 imply that the more capable the model is, the more aligned it will be? I know not everyone is going to accept the premise here, but yeah, what would you say to that?
我认为人们也指出,因为 GPT-4 仍然可以被攻破,而且在某种意义上能力更强,最坏情况的行为更糟。所以即使平均而言好得多,你也可以提出这样的论点。但我认为即使它全面更好,我也不认为我们应该押注这个趋势会持续。机器学习中有很多例子出现某种逆缩放:它变好一段时间然后变差。在某种程度上,我们知道模型还没有达到这个临界阈值,即它们和我们一样聪明,或者能想出很多非常好的方法来欺骗我们。它们没有那么多情境意识;它们不太了解自己是一个正在被训练的语言模型以及如何被训练。它们并不真正理解这一点。但一旦它们理解了,情况就完全不同了。你将面临不同的问题。所以从我们现在看到的趋势外推,我认为无论哪种方式都不对。但我确实认为你可以从中学习;只是不应该跳到那个结论。
I think people have also pointed out that because GPT-4 is still breakable, and it is more capable in some sense, the worst-case behavior is worse. So even though on average it's much better, you can make a case for that. But I think even if it was just better across the board, I don't think we should bet on that trend continuing. There are plenty of examples in machine learning where you get some kind of inverse scaling: it gets better for a while and then gets worse. To some extent, we know the models haven't reached this critical threshold where they are as smart as us or could think of a lot of really good ways to try to deceive us. They don't have that much situational awareness; they don't know that much about the fact that they are a language model being trained and how they are being trained. They don't really understand that. But once they do, it's a different ball game. You're going to be facing different problems. So extrapolating from some trend we see now, I don't think would be right in either way. But I do think you can learn something from it; you just shouldn't jump to that conclusion.
从主流机器学习的角度来看,这个项目最令人兴奋的是什么?
What's most intellectually exciting about this project from a mainstream ML perspective?
我认为我们会学到很多关于大型神经网络实际上是如何工作的。如果你想想我们试图在泛化上做的工作,很奇怪我们不明白为什么模型有时以一种方式泛化,有时以另一种方式。我们如何改变它们泛化的方式?为什么我们不能列出所有可能的方式,然后看看哪些有效?我们如何让它们进入每一种方式?这里真正发生的机制是什么?我们不知道。为什么我们不知道?如果你想想可解释性,仅仅能够理解模型决定下一个输出哪个 token 的机制,就会告诉我们很多关于那里发生了什么。它实际上是如何工作的?在某种程度上,这就是全部。人们花费大量精力通过向这些模型投入更多算力和更多数据来提升能力,然后他们得到这个更难以理解的机器,他们不理解它。这在某种程度上很酷,因为它能做事情,但听起来在某个时候,更有趣的事情可能是它是如何工作的,而这正是你要研究的。但同时,也有一些非常具体的东西可以说。例如,归纳头:你可以找到这些注意力头,它们做非常具体的事情,比如归纳。你可以找到有人逆向工程出在小模型中做简单算术的电路。你实际上可以做到。或者我们发现了加拿大神经元:GPT-2 中有一个神经元只对加拿大概念有反应,它就在那里。我们找到了它。有太多东西要发现,因为我们知道的太少了,不去看它简直是疯了。
I think we'll learn a lot about how big neural networks actually fundamentally work. If you think about the work we're trying to do on generalization, it is weird that we don't understand why models sometimes generalize in one way and sometimes another way. How can we change the ways they generalize? Why can't we just list all the possible ways and see which ones work? How can we get them into each of those? What's the mechanism that really happens here? We don't know that. And why don't we know that? If you think about interpretability, just being able to understand the mechanisms by how the models decide which token to output next will teach us a lot about what's going on there. How does it actually work? It's like on some level this is the whole thing. People are spending enormous amounts of effort increasing capabilities by just throwing more compute and more data into these models, and then they get this further inscrutable machine that they don't understand. That is very cool in a way because it could do stuff, but it sounds like at some point maybe the more interesting thing is how does it work, which is what you're going to be working on. But at the same time, there are really concrete things you can say. For example, induction heads: you can find these attention heads that do very specific things like induction. You can find someone reverse-engineer the circuit that does arithmetic for simple arithmetic in a small model. You can actually do that. Or we found the Canada neuron: there's a neuron in GPT-2 that just reacts to Canadian concepts, and it's just there. We found it. There's so much to find because we just know so little, and it's kind of crazy not to look at that.
我想这些网络中会有一些结构类似于人脑所做的东西,我们很可能在弄清楚它们在人脑中如何工作之前很久就能弄清楚它们在这些网络中如何工作,因为我们对所有这些类型和活动有完美的数据。
I imagine that there are some structures in these networks that are going to be analogous to things that the human brain does, and we will probably be able to figure out how they work in these networks long before we figure out how they work in the human brain because we have perfect data about all these types and activities.
完全正确。所以似乎所有研究大脑的人都应该接受这一点,并开始努力理解;你的生活会轻松得多。我不知道为什么没有更多人这样做。这对我来说似乎非常吸引人,但我不是神经科学家。而且也许一些见解也会迁移。你可以找到一些我们知道视觉模型拥有的神经元,在人类和动物中也能找到,比如这些边缘滤波器。或者如果你看强化学习,我们有证据表明强化学习在人脑中如何工作,但我们有更多的证据表明它在这些模型中如何工作。
Exactly. So it seems like all the people studying the brain should just accept that and start working on understanding; your life will be so much easier. I don't know why not more people do it. It seems so compelling to me, but I'm not a neuroscientist. And maybe some of the insights will also transfer. You can find some of the neurons that we know vision models have that you can also find in humans and animals, like these edge filters. Or if you look at reinforcement learning, we have evidence for how reinforcement learning works in the human brain, but we have so much more evidence how it works in these models.
你认为到目前为止,技术性 AI 安全方面最大的胜利是什么?
What do you think have been the biggest wins in technical AI safety so far?
如果非要选一个,那可能是基于人类反馈的强化学习(RLHF)。我认为在某种程度上,RLHF 真正让对齐问题进入了人们的视野,并且它也展示了对齐在系统实际构建中能增加很多价值。它产生了大量的商业影响,这非常好,因为它以一种切实的方式展示了现实世界的价值——如果你只是试图解决对齐超级智能这个非常抽象的问题,你可能会琢磨很多年而没有明确的、可衡量的进展。我认为,RLHF 不仅在模型使用前后带来了每个人都能直观感受到的显著差异,而且它还清楚地表明,这是一个值得投资和押注的领域,即使那些尚未明显见效或仍处于非常抽象阶段的东西也是如此。
I think if I had to pick one, it would probably be RLHF. I think in some ways, RLHF really put alignment on the map, and I think it also demonstrated that alignment has a lot of value to add to how systems are actually being built. I think the fact that it actually had a whole bunch of commercial impact has been really good, because it really demonstrates real-world value in a way that, if you're just trying to solve this abstract problem of aligning superintelligence, which is a super abstract problem, you could kind of noodle on it for many years without making clear measurable progress. I think not only does RLHF have this really visceral difference between how the model was before and how it was after, that everyone can really see when they play with it, but also it makes it clear that this is an area that's really worth investing in, and taking a bet on, even the things that aren't obviously working yet or are still at the stage of being really abstract.
有第二名吗?
Is there a number two?
我认为有一些较小的胜利。很难排名。如果我想补充其他东西,我认为视觉模型的可解释性相当令人印象深刻,并且在这方面取得了很大进展。但就安全影响或对齐影响而言,可能不那么明确,因为没有什么可以直接归因于此。
I think there are a number of smaller wins. It's hard to make rankings. If I wanted to add other things, I think interpretability of vision models has been pretty impressive, and there's been a lot of progress in that. But in terms of safety impact or alignment impact, it's maybe less clear, because there's no thing you can really point to that follows directly from that.
这里有一个听众反复提到的问题:是什么让 OpenAI 有权在没有民主意见的情况下开发通用人工智能,即我们是否真的想要开发这些系统?
Here's a question that was a recurring theme among listeners: what gives OpenAI the right to develop artificial general intelligence without democratic input as to whether we want to actually develop these systems or not?
这是一个很好的问题。我认为这也是一个更广泛的问题,比如为什么我们应该在许多其他事情上也有民主意见,比如模型应该如何表现,或者我们应该如何以这种方式或那种方式部署它。从某种程度上说,OpenAI 的使命是开发惠及全人类的 AI,但你必须让人类对正在发生的事情有发言权。这不是超级对齐团队做的事情,但我认为这将非常重要。
This is an excellent question. I think it's also a much wider question, like why we should have democratic input on a lot of other things as well, like how the model should behave, or how we should deploy it in this way or that way. In some ways, OpenAI's mission is to develop AI that benefits all of humanity, but you have to give humanity a say in what's happening. This is not something the superalignment team does, but I think it's going to be very important.
听起来你认同 AI 实验室和民主政治之间需要某种整合,公众必须被咨询,人们必须了解风险和收益,并且需要某种集体决策来决定这些事物何时以及如何开发和部署。我们目前没有这样的基础设施。我认为这 partly 是 OpenAI 的责任,但也 partly 是每个人的责任,整个社会的责任。只要 OpenAI 愿意合作,就需要付出巨大努力来实现它。
It sounds like you're on board with there needing to be some integration between the AI labs and democratic politics, where the public has to be consulted, people have to be informed about the risks and the benefits, and there needs to be some sort of collective decision about when and how these things are going to be developed and deployed. We currently don't have the infrastructure to do that. I guess that's partly OpenAI's responsibility, but it's also partly everyone's responsibility, the whole of society. As long as OpenAI is willing to collaborate, there just needs to be a big effort to make it happen.
我认为没错。我很高兴 OpenAI 愿意公开谈论风险和我们所处的位置。我也认为我的责任是告知公众对齐方面哪些有效、哪些无效,以及我们目前的位置和未来可能的方向。但归根结底,政府将在这一切如何发展方面发挥作用。
I think that's right. I'm really happy that OpenAI is willing to speak openly about the risks and about where we are. I see my responsibility also to inform the public about what is working on alignment and what isn't, and where we are and where we think we can go. But at the end of the day, governments will have a role to play in how this all goes.
如果国会调查这一切并得出结论认为它危险得令人不安,认为需要停止大量此类研究,你认为 AI 实验室会愿意配合吗?
If Congress investigates all of this and concludes that it's uncomfortably dangerous and thinks that a bunch of this research needs to be stopped, do you think the AI labs would be willing to go along with that?
是的,因为那是更民主、更合法的过程会输出的结果,我们应该做好公民,放慢或停止。AI 公司必须遵守所在国的法律。事情就是这样。我认为将会发生的是,我们将对前沿 AI 技术进行监管,人们正在努力弄清楚如何做到这一点。我们应该尽可能明智地去做。还有一个更大的问题,即如何让某些东西不仅在美国或英国有效,而是在全球范围内有效。如果存在构建真正危险的 AI 的方法,那么这必须适用于所有人,而不仅仅是特定国家。这也是一个关键挑战,不是我个人正在研究的,但我们需要解决它,我对任何正在研究这个问题的人感到兴奋。
Yeah, as that is what a more democratic, more legitimate process would output, we should be good citizens and slow down or stop. AI companies have to follow the laws of the country they're in. That's how this works. I think what's going to happen is we will have regulation of frontier AI technology, and people are trying to figure out how to do that. We should try to do it as sensibly as possible. There is the larger question of how you can have something that works not just in the United States or the United Kingdom, but worldwide. If there are ways to build AI that are actually really dangerous, then that has to apply to everyone, not just specific countries. That's also a key challenge, not one I'm personally working on, but we need to solve that, and I'm excited for anyone who is working on that problem.
我想让我有点悲观的是,我们不仅仅需要解决一个问题;我们需要解决很多问题。如果我们搞砸了其中一个,那可能会非常糟糕。我们不仅需要技术解决方案,还需要确保它在正确的地方部署并且每个人都遵守。即使这样有效,也许你还会遇到一个结构性问题,即它按照我们的指令行事,却让社会变得更糟。
I suppose what makes me a bit pessimistic is that we don't just need to solve one thing; we need to solve many things. If we mess up just one of them, that could be very bad. We don't just need a technical solution, but we need to make sure it's deployed in the right place and everyone follows it. Even if that works, maybe you could get one of these structural problems where it's doing what we tell it to but makes society worse.
是的,我认为这是这一切的另一面。现在有如此多的机会来塑造人类的未来,听众可以参与其中并产生巨大影响。有太多工作要做,而且我们很有可能生活在人类历史上最具影响力的时代,无论是过去还是未来。有点疯狂。可能是这样。我不知道。
Yeah, I see it as the flip side of all this. There's so much opportunity to shape the future of humanity right now that the listener could be working on and could have a lot of impact. There's so much work to do, and there's a good chance we actually live at the most impactful time in human history that has ever existed and that will ever exist. Kind of wild. Could be the case. I don't know.
今年三月你发推文说:「在我们匆忙将大型语言模型深度整合到经济的各个角落之前,我们能否停下来思考一下这样做是否明智?这是一项相当不成熟的技术,我们不了解它的工作原理。如果不小心,我们就会给自己带来许多关联性故障。」几天后,OpenAI 开放了 GPT-4 通过 API 连接各种插件的功能。一位听众想了解更多你这句话的含义,以及是否可能……
Back in March you tweeted: 'Before we scramble to deeply integrate large language models everywhere in the economy, can we pause and think about whether it's wise to do so? This is quite immature technology and we don't understand how it works. If we're not careful, we're setting ourselves up for a lot of correlated failures.' A couple of days after that, OpenAI opened up GPT-4 to be connected to various plugins through its API. One listener was curious to hear more about what you meant by that and whether there might be a...
OpenAI 内部对于 GPT-4 应该多快接入互联网并集成到其他服务中存在分歧。
Disagreement within OpenAI about how soon GPT-4 should be hooked up to the Internet and integrated into other services.
是的,我意识到那条推文有些模棱两可,被以多种方式解读。从根本上说,插件能让你做的事情,用 API 也能做到。插件并没有真正增加任何人们原本无法做到的事情。我认为 OpenAI 非常清楚将插件接入系统可能出什么问题。你必须做同样的工作,必须小心,必须让人们花钱等等这些问题。但他们就在我们旁边,我们和他们谈过,他们一直在思考。但鉴于人们对在所有事情上尝试 GPT-4 的热情如此之高,我真正想说的是:看,这还不够成熟,系统会失败。先别把它连接到所有东西上。确保有备用系统。确保你真正玩过这个模型,了解它的局限性。如果让模型写代码,确保你阅读并理解代码,或者在沙盒中执行它,否则系统可能会崩溃。无论你在哪里写代码,它都可能破坏那个系统。要小心,要明智。确保你明白自己在做什么,而不是把它连接到所有东西上然后看结果。
Yeah, I realized that tweet was somewhat ambiguous and was read in lots of different ways. Fundamentally, what plugins allow you to do is nothing on top of that you couldn't do with the API. Plugins don't really add anything fundamentally that people couldn't already do. I think OpenAI is very aware of what can go wrong when you hook up plugins to the system. You have to have the same works, you have to be careful, you have to let people spend money and all these questions. But they also sit right next to us and we talked to them about it, and they've been thinking about it. But given how much excitement there was to just try GPT-4 on all the things, what I really wanted to do also is to say: look, this is not quite mature, the system will fail. Don't connect it to all the things yet. Make sure there's a fallback system. Make sure you've really played with the model, do you understand its limitations? If you have the model write code, make sure you're reading the code and understanding it, or executing it in the sandbox, because otherwise the system might break. Wherever you're writing the code, it might break that system. Just be careful, be wise. Make sure you understand what you're doing here, and not just hook it up to everything and see how it goes.
有没有什么事情是人们用 GPT-4 做的,你觉得可能还为时过早,我们应该放慢脚步,做更多测试?
Is there anything that people are using GPT-4 for where you feel like maybe it's premature and we should slow down and do some more testing?
可能吧。我不知道能不能给你一些好的例子,但我认为这通常是新技术的常态。我本质上是一个技术乐观主义者,我认为我们应该把 AI 用在它擅长的所有事情上。在某种程度上,我们刚刚花了一个小时讨论用 AI 做对齐研究有多棒,而这是我的工作。所以我正试图用 AI 取代自己的工作。但与此同时,你也必须真正理解这项技术的局限性,其中一些并不明显,也不广为人知。你必须这样做才能负责任地部署它,并以明智的方式将其融入社会。我认为就像所有新技术一样:我们会尝试很多事情,我也为人们尝试很多事情感到兴奋。这就是为什么我认为 OpenAI API 的存在是好的,它让很多人可以使用前沿语言模型做各种事情,但你在做的时候也要小心。
Probably. I don't know if I can give you some good examples, but I think that's generally the story with new technology. I'm fundamentally a techno-optimist and I think we should use AI for all the things it's good for. To some extent, we just spent an hour talking about how great it would be to use AI for alignment research, which is my job. So I'm trying to replace myself at my job with AI. But at the same time, you also have to really understand the limitations of this technology, and some of it is not obvious and not widely known. You have to do that in order to deploy it responsibly and integrate it into society in a way that is actually wise. I think it's just as always with new technologies: we'll try a lot of things, and I'm excited for people to try a lot of things. That's why I think it's good that the OpenAI API exists and lets lots of people use cutting-edge language models for all kinds of things, but you want to be also careful when you're doing that.
关于把东西接入互联网这个话题,很多年前人们经常讨论一个假设:如果我们有一个像 GPT-4 一样强大的智能系统,我们会把它放在一个铅封的盒子里,不会把它接入互联网,因为我们会担心。但现在的文化似乎是,模型一做好就立刻部署到互联网上。这不完全正确,对吧?
On this topic of just plugging things into the internet, many years ago people talked a lot about the assumption that if we had an intelligent system as capable as GPT-4, we would keep it in a lead-contained box and wouldn't plug it up to the internet because we would be worried about it. But it seems like the current culture is that as soon as a model is made, it just gets deployed onto the internet right away. That's not quite right, right?
好吧,GPT-4 在公开可用之前我们已经有了八个月。我们做了很多安全测试,做了很多红队测试,在它的对齐方面取得了很大进展,我们没有立即把它连接到所有东西上。我认为这很好。但我认为你真正想说的是,很多年前人们在争论是否能把 AGI 关在盒子里,它永远不会逃出来,永远不会做坏事。现在看起来那艘船已经开走了,现在你把它连接到一切东西上。这在一定程度上是我要提醒的:我们在连接它的时候应该谨慎。仅仅因为 GPT-4 在 API 上,并不意味着每个未来的模型都会立即在所有情况下对所有人可用。这是你必须走的一条困难路线:你想用 AI 赋能每个人,但同时又必须注意滥用和不对齐。如何平衡这种权衡?这是关键问题之一。
Okay, we had GPT-4 for like eight months before it was publicly available. We did a lot of safety tests, we did a lot of red teaming, we made a lot of progress on its alignment, and we didn't just connect it to everything immediately. I think that's good. But I think what you're actually trying to say is that many years ago people were arguing over whether you can keep AGI in the box and it'll never break out and never do anything bad. And now it seems like that ship has sailed, and now it's like you're connecting it to everything. That's partially what I'm trying to caution here: we should be mindful when we do connect it. Just because GPT-4 is on the API doesn't mean that every future model will be immediately available for everything and everyone in every case. This is the difficult line you have to walk: you want to empower everyone with AI, but at the same time you have to be mindful of misuse and misalignment. How do you balance that trade-off? That's one of the key questions.
一种划分方式是接入互联网与否。但我觉得人们经常——我也有这个毛病——要么认为它部署在互联网上,消费者在使用,要么它安全地在实验室里,没有问题。但中间还有情况。即使放在实验室里也可能有问题,对吧?
One way of breaking it up would be connected to the internet versus not. But I feel like often people, I'm guilty of this as well, we're thinking either it's deployed on the internet and consumers are using it, or it's safely in the lab and there's no problem. But there's this intermediate. There can also be problems if you have it in a lab, right?
这正是我想说的。我觉得有时人们会忽略这一点。滥用是一个问题,如果它接触到更广泛的公众;但不对齐也可能是一个问题,即使只是训练过并在公司内部使用,因为它会想办法产生更广泛的影响。我们倾向于把所有风险归为一类或泛泛而谈。一个模型仅仅经过训练就可能很危险,即使从未接入互联网,这一点我们真的需要牢记。我想 OpenAI 的人会记住这一点。安全审查真的需要在开始训练之前就开始。
That's exactly what I'm saying. I feel like sometimes people lose track of that. Misuse is an issue if it reaches the broader public, but misalignment can be an issue if something is merely trained and is just being used inside a company, because it will be figuring out how it could end up having broader impacts. We tend to cluster all these risks or speak very broadly. The fact that a model can be dangerous if it's simply trained, even if it's never hooked up to the internet, is something we really need to keep in mind. I guess it sounds like OpenAI people will keep that in mind. Safety reviews really need to start before you even start the training run.
OpenAI 创建并推出 ChatGPT 的决定可能加速了 AI 研究,因为现在人们涌入这个领域,但它也引发了一系列关于安全的担忧,以及提前准备以应对潜在威胁的新努力。事后看来,综合考虑,你认为发布 ChatGPT 是增加还是减少了 AI 灭绝风险?
OpenAI's decision to create and launch ChatGPT has probably sped up AI research because there's now a rush into the field, but it has also prompted a flurry of concerns about safety and new efforts to prepare ahead of time to see off possible threats. With the benefit of hindsight, do you think that move to release ChatGPT increased or reduced AI extinction risk, all things considered?
我认为这是一个非常难的问题,我不知道我们是否能真正明确地回答这个问题。
I think that's a really hard question, and I don't know if we can really definitively answer this.
我认为从根本上说,如果 ChatGPT 能晚一点发布可能会更好。但某种程度上,这一切是不可避免的。公众迟早会意识到语言模型已经变得多好。令人惊讶的是,这种情况竟然持续了这么久才发生。我真的很高兴它如何改变了对话,推进了关于 AI 风险的讨论,以及实际的对齐工作。我们确实可以让事情变得更好,我们应该做更多。我认为这两点都很好。你可以争论时机应该怎样,以及它是否无论如何都会发生。我认为它无论如何都会发生。
I think fundamentally it probably would have been better to wait with ChatGPT and release it a little bit later. I think to some extent this whole thing was inevitable. At some point the public would have realized how good language models have gotten. It's been surprising that it went this long before that was the case. I was honestly really happy how much it shifted the conversation and advanced the discussions around risks from AI, but also the real alignment work that has been happening. We can actually make things so much better and we should do more of that. I think both of these are really good. You can argue over what the timing should have been and whether it would have happened anyway. I think it would have happened anyway.
当人们问「如果我们想,难道不能停止做 AI 吗?」这听起来很容易:停下来,不做就行了。但实际上,世界上有太多力量阻止这件事。假设 OpenAI 决定不训练更强大的模型,他们可以这么做。但有很多 OpenAI 的竞争对手可能仍然会做。然后你试图让前五大 AGI 实验室或训练最大模型的科技公司承诺停止。但新的初创公司会出现。而且晶体管不断缩小,GPU 变得更强大,训练比以往任何模型都更强大的模型的成本每年仍呈指数级下降。然后你去找半导体公司,以及上游从事 UV 光刻的公司,他们从 90 年代就开始研究下一代芯片。让所有人都冷静下来是一个非常复杂的协调问题。甚至很难弄清楚还有谁参与其中。我认为如果人类真的想,可以做很多事情,如果事情变得非常可怕,很多事都可能发生。但根本上,这不是一个容易解决的问题,我不想假设它已经解决了。我想要的是确保我们在拥有的时间内尽可能取得对齐进展。如果我们有更多时间,那很好。但如果没有,我仍然希望能够解决对齐问题。我想在那些我们没有额外时间的世界里获胜。所以无论情况如何,我们需要尽快解决这些技术问题。
When people ask, 'Can't we all just stop doing AI if we wanted to?' It feels so easy: just stop, don't do it. But in practice, there are so many forces in the world that keep this from stopping. Let's say OpenAI decides not to train a more capable model. They could do that. But there are many OpenAI competitors who might still do it. Then you try to get the top five AGI labs or tech companies that will train the biggest models to promise to stop. But then new startups will appear. And transistors keep getting smaller, so GPUs become more capable, and the cost to train a model more capable than any previous model still drops exponentially year over year. Then you talk to semiconductor companies, and upstream companies working on UV lithography, and they've been working on next-generation chips since the 90s. Getting everyone to chill out is a really complicated coordination problem. It's not easy to figure out who else is involved. I think humanity can do a lot if it really wants to, and if things get really scary, a lot can happen. But fundamentally, it's not an easy problem to solve, and I don't want to assume it's solved. What I want is to ensure we make as much alignment progress as possible in the time we have. If we get more time, great. But if we don't, I still want to be able to solve alignment. I want to win in the worlds where we don't get extra time. So however it goes, we need to solve these technical questions as quickly as possible.
我在网上看到有人试图放慢速度,为你和你的团队争取更多时间。有些人持极端观点,希望完全停止全球 AI 进展相当长一段时间。这似乎很难。我猜他们的理论可能是,某场灾难会改变态度,使目前不可能的事情成为可能。撇开这个不谈,就解决对齐的竞赛而言,我们可以要么放慢 1% 的速度,要么加快 1% 的对齐研究。哪个更容易?听起来你认为加快对齐研究更容易,让它的进展速度翻倍,比让时间线延长一倍更容易。
I've seen online that there are people trying to slow things down to buy more time for you and your team. Some hold an extreme view that they want to completely stop progress on AI globally for a significant period. That seems like a heavy lift. I imagine their theory might be that at some point a disaster will change attitudes, making currently impossible things possible. Setting that aside, in terms of the race to solve alignment, we could either slow things down by one percent or speed up alignment research by one percent. Which is easier? It sounds like you think it's easier to speed up alignment research, to get it proceeding twice as quickly, than to make timelines twice as long.
我认为这非常重要。考虑到现在实际从事对齐工作的人很少。有多少?几百?几千?取决于你怎么算。超级对齐团队目前大约有 20 人,但 OpenAI 还有很多其他对齐工作。如果算上所有对齐工作,可能超过 100 人。但两年前,只有大约三个人在做对齐工作。虽然已经大幅增加,但我们仍然需要更多。即使是非常有才华的个人,现在转向研究这个问题也能产生巨大影响,因为这个领域仍然很小。还有很多事情要做,还有很多我们不了解。在某些方面,这感觉像是真正的最终研究前沿。我们已经弄清楚了 Scaling;我们知道如何让模型更智能。对齐是一个真正的研究问题。我们不知道如何对齐超级智能。我们想要弄清楚。我们必须这样做。这不是可选的。
I think that's a really important point. Given how few people are actually working on alignment these days. What is it, hundreds? Thousands? It depends on your account. The superalignment team is about 20 people right now, but there are many other alignment efforts at OpenAI. If you count all alignment work, it's probably more than 100. But if you go back two years, there were like three people doing alignment work. It's ramped up a lot, but we still need so much more. Even really talented individuals can still make a big difference by switching to work on this problem now, just because it's still such a small field. There's still so much to do, so much we still don't understand. In some ways, it feels like the real final research frontier. We've figured out scaling; we know how to make models smarter. Alignment is a real research problem. We don't know how to align superintelligence. We want to figure this out. We have to. It's not optional.
这个领域如此之小,一方面令人沮丧,但也是乐观的理由,因为你可以让它翻倍。如果你能让一千名机器学习研究人员转而从事对齐工作,那将彻底改变局面。
The fact that the field is so small is exasperating on one level, but it's also a reason for optimism because you could double it. If you could get a thousand ML researchers to switch into working on alignment, that would completely transform things.
正是如此。
Exactly.
另一个问题:Jan 声称超级对齐团队不会回避对齐。
Another question: Jan claimed that the superalignment team wouldn't be avoiding alignment.
有助于商业化的工作,但这类工作本身就已经有金钱激励了。那他为什么不试图避开那些无论如何都可能被完成的工作呢?
Work that helps with commercialization, but that work in particular is already incentivized monetarily by definition. So why isn't he going to try to avoid that work, which will probably get done either way?
我认为这正是很多人想指出的关键:对齐不会默认以我们真正满意的方式完成。我们想解决的问题目前尚未解决。其中一些会有商业价值。从根本上说,如果你有两种构建 AGI 的方式,其中一种与人类对齐得更好,人们会想买第二种,因为它对他们更好。这必然会有商业价值,而且无法避免。过去有人提出过一个类似的批评:很多人觉得基于人类反馈的强化学习(RLHF)是一种能力进步,因为经过 RLHF 的模型感觉更强大。你需要激活它们;它们更有用,实际上能做更多事情。原因在于它们试图帮助你,它们更在线,它们利用自身能力去完成你要求的事情。但预训练模型不是这样,所以它显然感觉更强大,因为你解锁了所有这些能力。如果你看看微调过程中实际发生了什么,模型并没有真正学到它之前没有的全新技能。你可以通过静态微调做到这一点,但以我们使用的算力预算不行。对于 GPT-3,微调算力不到预训练算力的 2%;对于 GPT-4,比例甚至更小,非常小的一部分。但与此同时,因为模型现在更加努力地提供帮助,它确实更有帮助,你得到了原本就存在的所有能力。所以回到商业化问题,我真正想做的是解决问题。如果这有商业价值,那很好。如果没有,其中一些不会,或者一些研究赌注不会成功,或者在我们真正得到非常谨慎的系统之前,有些事情不会有用。这没关系。目标是解决问题;这就是我们想做的。
I think this is the whole point that a lot of people are trying to make: alignment wouldn't be done by default in the way that we are really happy with. The problems we want to solve are currently unsolved. Some of it will be commercially valuable. Fundamentally, if you have two ways of building AGI, and one is just much more aligned with humans, people will want to buy the second one because it's just better for them. That will necessarily have commercial value and it'll be unavoidable. An adjacent criticism that has been raised in the past is that a lot of people feel like RLHF has been a capabilities progress because RLHF models feel more capable. You have to activate them; they're more useful, they're actually doing more things. The reason is because they're trying to help you, they're more online, they're leveraging their capabilities towards whatever you're asking them to do. But the pre-trained model isn't, so it obviously feels a lot more capable because you've unlocked all of these capabilities. If you look at what actually happens during fine-tuning, the model isn't really learning fundamentally new skills it didn't have before. You can do that through fine-tuning statically, but not with the kind of compute budget we use. For GPT-3 it was less than two percent of the pre-training compute; for GPT-4 it was even less than that, a really tiny fraction. But at the same time, because the model is now trying so much harder to be helpful, it is more helpful and you get all the capabilities that had been there in the first place. So to come back to the commercialization question, what I really want to do is solve the problem. If that is commercially useful, great. If not, some of it will not be, or some of the research bets won't work out, or some of the things won't be useful before we actually get really careful systems. That's fine. The goal is to solve the problem; that's what we want to do.
另一个问题:OpenAI 押注于不会出现非常快的起飞。他们是否也尝试制定计划,以应对 AI 极其快速的递归自我改进的情况?
Another question: OpenAI is banking on there not being a really fast takeoff. Do they try to make plans that could also work in the event of a scenario that is like extremely rapid recursive self-improvement of AI?
是的,我认为我们绝对应该为这种情况制定计划,并在发生时做好准备。在某种程度上,自动化对齐研究可能是我知道的在这种情况下的最佳计划,即你必须根据事态发展按比例扩大对齐工作。如果你能通过将几乎所有工作委托给机器来完成,那么它们实际上可以跟上机器的步伐,因为只有它们才能做到。
Yeah, I think we should definitely plan for that scenario and be ready if it happens. To some extent, automated alignment research is probably the best plan I know in that kind of scenario, where you really have to scale up your alignment work in proportion with what's going on. If you can do this by delegating almost all of the work to machines, then they can actually keep pace with the machines because they're the only ones that can.
我想担忧在于,如果出现智能爆炸且速度非常快,你几乎没有时间将计划付诸行动并跟上。但那将是一个非常糟糕的情况;它使得任何计划都很难奏效。
I guess the concern would be if there is an intelligence explosion and it's very fast, there's very little time for you to put your plans into action and to keep up. But that would be a very bad situation; it makes it very hard for any plan to work.
没错。所以我们应该做的,如果你想对技术进展的速度保持不可知论(这正是我们想做的),最好的办法就是尽可能提前准备。这就是为什么我们现在就需要开始思考如何对齐我们尚未拥有的系统。你准备得越多,就越能为那种情况做好准备。
That's right. So what we should be doing, if you want to be agnostic to the speed of tech progress, which is what we want to do here, is the best thing you can do is to prepare as much as possible ahead of time. That's why we need to start thinking now about how to align systems that we don't have yet. The more you can prepare, the more you'll be ready for that scenario.
我收到一个问题,略有变化:OpenAI 认为对齐是可解决的理由是什么?他们是否看过 Roman Yampolskiy 博士关于不可解性的不可能性论证?他们链接了一篇包含这些论证的论文。我不确切知道那些论证是什么,但我知道有人从理论上论证对齐是不可能的或极其困难的,出于一些概念性原因。有没有哪些这类论证特别困扰你,或者你认为这类论证不应该那么有说服力?
A question I got, which slightly changes: what are OpenAI's grounds for thinking alignment is solvable? Have they seen Dr. Roman Yampolskiy's impossibility arguments against solvability? They linked to a paper with those arguments. I don't know exactly what those arguments are, but I know there are people who have made theoretical arguments that alignment is impossible or extremely difficult for some conceptual reasons. Are there any arguments along those lines that trouble you in particular, or do you think that kind of argument shouldn't be so persuasive?
我看过你提到的那篇论文。我见过的任何论证,我都没有觉得特别有说服力。试图进行理论论证的问题在于你需要一些假设,然后大问题就变成了这些假设是否成立。对我来说,这似乎还没有定论。它可能最终被证明是不可能的;我觉得这不太可能,但我没有证据。我认为我们会非常努力地通过展示它可以做到来寻找反例。现在绝对不是放弃的时候;我认为这是非常可行的。
I've looked at the paper you mentioned. Any argument I've seen, I haven't found particularly persuasive. The problem with trying to make a theoretical argument is you need some kind of assumptions, and the big question then becomes whether these assumptions are going to be true. To me, it really seems like the jury is still out on this. It could turn out to be impossible; it doesn't feel particularly likely to me, but I don't have a proof for that. I think we're going to work really hard to find a counterexample by showing that it can be done. It's definitely not the time to give up; I think it's very doable.
我能感觉到一丝恼怒,你好像在说「所有这些抱怨这个问题无法解决的人,他们并没有帮忙。」显然有很多事情我们可以尝试;为什么不直接试试呢?
I could feel a bit of exasperation coming through where you're like 'all these people complaining that this problem is insoluble, they're not helping.' Clearly there are so many things we could try; why don't we just try them?
他们在某种意义上是在帮忙,因为他们间接为我们做招聘,因为他们引起了人们对这个问题的关注。如果你到处说这个问题很容易,你就不会引起注意;人们会觉得「好吧,没事,不用担心。」但我也认为这产生了一种「哦,看起来很难,我们放弃吧」的真实情绪,这绝对是错误的方法。相反,这意味着我们应该更加努力,让更多人尝试解决它。永不放弃,永不投降。这一切都还是未知数;我们应该真正攻克它。
They are helping in the sense that they're indirectly doing recruiting for us, because they're drawing attention to the problem. If you just went around saying the problem is easy, you wouldn't draw attention to it; people would be like 'okay, it's fine, don't worry about it.' But also, I think that has created a real energy of 'oh, it seems really hard, let's give up,' and that is absolutely the wrong approach. If anything, that means we should try harder and get more people to try to solve it. Never give up, never surrender. This is all still up in the air; we should just really crush it.
两个指向同一方向的问题:随着 OpenAI 越来越接近 AGI,他们是否计划在给予 AI 操纵员工或黑客攻击的机会方面偏向于偏执?
Two questions that are pointing in the same direction: as OpenAI gets closer to AGI, do they plan to err on the side of paranoia in terms of giving AIs opportunities to manipulate staff or hack?
另一个人问:在大型训练运行中,你愿意承担多少人类灭绝的风险,比如训练 GPT-5、6 或 7?
Another person asked: how much risk of human extinction are you willing to take in a large training run, like for example to train GPT-5, 6, or 7?
总的来说,随着风险增加,我们对对齐和安全的证明要求也更高。我们一直在每个系统中加强这一点。目前的系统仍然没有灾难性风险或接近那种程度。例如,GPT-2 是开源的;GPT-3 通过 API 提供;GPT-4 唯一公开可用的版本是对齐微调的 RLHF 版本,即 ChatGPT 版本。基础模型仅限研究人员访问。所以我们正在引导公众使用 RLHF 模型。每一步都提升了安全性和对齐。显然,能力水平越高,风险越大,需要的安全和对齐措施也越多。人们可以预期这一趋势会持续。
In general, as the stakes get higher, we have a much higher burden of proof of alignment and safety. We've been ramping this up with every system. The systems we have now are still not catastrophically risky or close to that. For example, GPT-2 was open source; GPT-3 was available via API; GPT-4's only publicly available version is the alignment fine-tuned RLHF version, the ChatGPT version. The base model is only on researcher access. So we're steering the public towards the RLHF model. With each step, we step up safety and alignment. Obviously, the higher the capability level, the higher the stakes, and the more safety and alignment measures you need. People can expect that trend to continue.
关于同一主题,有人问:你如何定义成功?你回答:科学界一致认为我们已经解决了对齐问题。他们问:OpenAI 能否做出有意义的承诺,例如,除非科学界广泛共识认为对齐已针对该类系统解决,否则不部署超过一定能力阈值的系统?
On the same theme, someone asked: how would you define success? You replied: the scientific community agrees that we've solved alignment. They asked: is there a meaningful commitment OpenAI could make, for example, not to deploy systems above a certain capability threshold unless there is broad scientific consensus that alignment has been solved for that kind of system?
归根结底,我们必须说服科学界。我不认为世界会让我们建造灾难性危险的东西。世界现在正在关注,这很好。我了解到,在英国,如果你想将房子出租给三个以上无亲属关系的人,你需要特殊许可证。但据我所知,目前不需要许可证或批准来训练 AGI。部分原因是我们可能还做不到。似乎没有太多法律限制,我们希望很快会有更多基础设施到位。人们正在研究监管,这是监管必须解决的问题。这方面有很多问题我不是专家。
At the end of the day, we have to convince the scientific community. I don't think the world will let us build something catastrophically dangerous. The world is paying attention now, and that's good. I've learned that in the UK, if you want to rent a house to more than three unrelated people, you need a special license. But as far as I can tell, one doesn't need a license or approval to drain an AGI. That's partly because we probably can't do that yet. It seems like there aren't many legal restrictions, and we're hoping there will be more infrastructure in place quickly. People are working on regulation, and that's something regulation has to solve. There are many questions around this that I'm not an expert in.
回到如何定义成功:仅仅说服自己我们做得很好是不够的,因为很容易说服自己关心的事情。我们实际上必须说服外部专家和审计师,他们正在仔细审视我们的所作所为和原因。我们需要大量的经验证据:这是我们尝试过的一切,这是这样做会发生什么,你可以查看数据和代码,人们可以审查我们的工作。因为风险如此之高,我们必须邀请大量审查。我们开始的一个方面是说明我们正在尝试什么,我们对齐系统的整体方法是什么,并邀请反馈和批评。也许有更好的方法;我很想知道,然后改用那个方法。公众应该知道我们在对齐方面做了什么,并独立判断是否足够。专家将发挥作用,因为他们的知识是做出知情结论所必需的。
To come back to how to define success: it's not sufficient to just convince ourselves that we did a good job, because it's easy to convince yourself of something you care about. We actually have to convince external experts and auditors who are looking exactly at what we're doing and why. We need a mountain of empirical evidence: here's all the things we tried, here's what happens when we do this, you can look at the data and code, and people can scrutinize what we're doing. Because the stakes will be so high, we have to invite a lot of scrutiny. One aspect we started with is to say what we're trying, what our overall approach to aligning the systems is, and invite feedback and criticism. Maybe there's something way better we could be doing; I would love to know that and then do that instead. The public should know what we're doing on alignment and make independent judgments on whether it's enough. Experts will have a role to play because their knowledge is required to make informed conclusions.
许多观众问题涉及政策和治理。我经常不理解技术细节,Twitter 上的许多人也不足以审查技术提案。所以我们考虑社会和组织层面。我觉得我的回答通常是:我很希望看到更多,但我不在做这个。向你提出这些问题并了解你的想法是合理的。有很多人需要采取行动。你必须埋头专注于技术,因为那是你的专长,但我们也需要 OpenAI 的治理人员建立良好的结构,参议院委员会也要发挥作用。有很多不同的部分需要拼合在一起。
Many audience questions are about policy and governance. I often don't understand the technical details, and many people on Twitter don't know enough to scrutinize technical proposals. So we think about the social and organizational level. I feel like my answer is often: I would love to see more of that, but I'm not working on it. It's reasonable to put these questions to you and find out what you think. There are a lot of people who need to take action. You have to keep your head down focused on technical stuff because that's your specialty, but we also need governance people at OpenAI to put in place good structures, and the Senate committee to play their role. There are many different pieces that have to slot together.
没错。
That's right.
我们正进入最后半小时。我的梦想是这次采访能帮你吸引很多优秀的人申请超级对齐团队的工作。理想情况下,它能让很多人从有趣但不太有用的事情转向既非常有趣又可能拯救世界的事情。我不想对超级对齐项目是否比其他项目更好或更差持强烈反对意见。
We're heading towards the final half hour. My dream is that this interview can help get you lots of great applications to work on the superalignment team. Ideally, it would move a lot of people from work that's interesting but not that helpful to something that is both super intellectually interesting and might save the world. I don't want to take a strong contrarian view on whether the superalignment project is better or worse than other projects.
你比我懂技术得多。我觉得你提出的计划和我听过的其他任何计划一样好,而且你似乎有资源和条件去真正实现它。另外,如果未来几年计划没有达到你预期的效果,我想你也能转向其他计划。那么你们在招聘什么职位?大概招多少人?详细说说。
Much more technically informed than me. I think the plan you've laid out seems as good to me as any other plan I've heard, and it seems like you've got the resourcing and situation to make a real go of it. And I guess also if the plan doesn't bear as much fruit as you hope in the next couple of years, I imagine you'd be able to pivot to a different plan. So what roles are you hiring for and what sort of numbers? Lay it all out.
我们主要招聘研究工程师、研究科学家和研究经理。我预计我们会继续大量招聘。我猜今年年底前至少招 10 个人,之后几年可能还会更多。
We are primarily hiring for research engineers, research scientists, and research managers. I expect we'll be continuing to hire a lot of people. Probably at least 10 before the end of the year is my guess, and then maybe even more in the years after that.
那么研究工程师、研究科学家和研究经理是什么样的?这些角色具体做什么?
So what do research engineers, research scientists, and research managers look like? What do these roles entail?
在某种程度上,我们在 OpenAI 并不严格区分研究工程师和研究科学家。在每个角色中,你都需要编写代码并运行自己的实验。事实上,我认为不断进行大量实验非常重要,快速测试你的想法,然后迭代并尝试更多地了解世界。总的来说,研究科学家职位不需要博士学位。你甚至不必之前从事过对齐研究。事实上,没有经验可能更好,因为你会对我们试图解决的问题有新的视角。我们通常希望人们带来的是对技术工作原理的良好理解。你理解语言模型,例如理解强化学习。你能够构建和实现机器学习实验并进行调试。在研究科学家一端,你更需要思考下一步做什么实验,提出如何解决我们试图解决的问题的想法,或者我们没有想到但可能应该考虑的其他问题,以及如何设计实验让我们学到更多。在研究工程师一端,更多的是实际构建让我们能够运行这些实验的东西,并取得我们已经知道的进展。如果我们有一堆好主意,那还不够。我们实际上必须测试它们、构建它们,并交付其他人可以使用的产品。这涉及编写大量代码、调试机器学习、运行大量扫描或实验、设置 GPT-4 和其他大型模型的大规模训练。实际上,团队中的大多数人都在这个谱系上移动。有时编码更多,因为我们大致知道该做什么;有时研究更多,因为我们还不知道该做什么,正在研究一个新项目。总的来说,你需要大量的批判性思维,提出重要问题,并对世界和我们正在构建的技术充满好奇心。对于研究经理角色,你管理一个中小型甚至大型的研究工程师和研究科学家团队,朝着特定目标前进。你应该设定方向:下一个里程碑是什么,我们应该往哪里走,如何在一个问题上取得进展,比如理解这种泛化或为自动化对齐研究制作数据集。你需要分解问题,发挥创造力,并确定人们可以做什么。但也有很多日常管理:如何让人们保持积极性和生产力,确保他们能够合作,就是传统的管理事务。
In a way, we don't make a strong distinction between research engineer and research scientist at OpenAI. In each of these roles, you're expected to write code and run your own experiments. In fact, I think it's really important to always be running lots of experiments, small experiments testing your ideas quickly, then iterating and trying to learn more about the world. In general, there's no PhD required for the research scientist roles. You don't even have to have worked in alignment before. In fact, it might be good if you didn't, because you will have a new perspective on the problems we're trying to solve. What we generally love for people to bring is a good understanding of how the technology works. You understand language models, you understand reinforcement learning for example. You can build and implement ML experiments and debug them. On the research scientist end of the spectrum, you would be expected a lot more to think about what experiments to do next, come up with ideas of how to address the problems we're trying to solve, or what other problems we aren't thinking about that maybe we should be, and how to design the experiments that will let us learn more. On the research engineering spectrum, there's a lot of actually building the things that let us run these experiments and make the progress that we already know. If we have a bunch of good ideas, that will not be enough. We actually have to test them and build them and ship something that other people can use. That involves writing a lot of code, debugging ML, running lots of sweeps or experiments, getting big training runs on GPT-4 and other big models set up. In practice, most people on the team move somewhere on the spectrum. Sometimes there's more coding because we kind of know what to do, and sometimes there's more research because we don't yet know what to do and we're studying a new project. In general, you need a lot of critical thinking, asking important questions, and being very curious about the world and the technology we're building. For the research manager role, you're managing a small, medium, or even large team of research engineers and research scientists towards a specific goal. You should be setting the direction: what are the next milestones, where should we go, how can we make progress on a question like understanding this type of generalization or making a dataset for automated alignment research. You have to break it down and make it creative and figure out what people can be doing. But also there's a lot of day-to-day management: how to keep people motivated and productive, make sure they can work together, just traditional management stuff.
所以听起来,对于前两个职位,主要的是你对当前的机器学习技术有很好的理解,能够实际参与并可能构思和运行实验。你们还需要其他具体的技能吗?或者,你们会非常期待收到什么样背景的人的申请?
So it sounded like for the first two, the main thing was that you had a good understanding of current ML technology, you could actually go in and potentially think up experiments and run experiments. Are there any other concrete skills that you require? Or what would be the typical background of someone you would be really excited to get an application from?
有很多不同的背景都适用。机器学习博士是进入该领域的传统途径,特别是如果你想做更多研究性的工作。但我认为你完全不需要博士学位。事实上,如果你现在考虑开始读博,我不知道你是否还有那么多时间。你应该现在就着手解决这个问题。对于研究工程师,背景可能是你曾在 STEM 领域工作,然后决定停下来,花六个月时间重新实现一堆机器学习论文,并以此学习。或者有人在科技公司做其他机器学习工程相关工作,现在想转向对齐研究。我认为这是非常好的背景。我还想强调,我们试图招聘的大多数人之前都没有从事过对齐研究,因为从事过的人太少了。你需要的核心专长是机器学习技能。关于对齐,有很多东西你应该了解,但你可以在入职后学习。你可以边做边学,我觉得没问题。对于研究经理角色,我想你需要有更多管理经验的人。优秀的研究者和优秀的管理者不是一回事。
There are a lot of different backgrounds that are applicable here. Machine learning PhDs have been the traditional way people get into the field, especially if you want to do something more researchy. But I don't think you need that at all. In fact, if you're thinking about starting a PhD now, I don't know if you'll have that much time. You should just go work on the problem now. For research engineers, the kind of background is maybe you've worked in a STEM field and you decide to stop doing that, take six months, and just re-implement a bunch of ML papers and learn a bunch that way. Or somebody who works at a tech company doing other machine learning engineering related things and now wants to switch to alignment. I think that's a really good profile. I also want to stress that most people we are trying to hire haven't worked on alignment before, just because there are so few people who have. The core expertise you will need is machine learning skills. There are a bunch of things you should know about alignment, but you can learn them once you're here. You can catch up along the way, and I think that's fine. For the research manager role, I guess you're looking for someone with more management experience. Being a good researcher and being a good manager are not the same thing.
确实可能分开。那么,你会寻找特定类型的人来担任经理角色吗?
Absolutely can come apart. So I guess would you be looking for a particular kind of person for the manager role?
是的,我认为它们可能负相关,有时确实如此。但理想情况下,你之前有过管理经验。有不同的方式可以做到。有些场景下,你可以将职责分给研究或技术负责人和经理。经理承担更多的管理职责,技术负责人则设定团队方向并确保技术工作完成。但在那种配置下,他们必须相处得很好,并且意见一致,才能有效分配职责。特别是,经理仍然应该对我们试图解决的问题有详细的理解。理想情况下,我们希望有人能同时胜任这两个角色。所以背景可能是曾在其他公司或机器学习其他分支领导过研究团队,或者在其他领域担任过经理,然后转为大型语言模型项目的个人贡献者。或者还有一条路,你在某个地方做博士后,带领一个小型研究团队,日常工作以编码为主,用语言模型或强化学习进行大量实验。这些都是可能的背景。但更大的筛选条件是,你应该真正关心我们试图解决的问题,非常擅长编码,并且非常擅长机器学习。
Yeah, and I think they can be anti-correlated, which they might be sometimes. But ideally you would have managed before. There are different ways it could go. There are scenarios where you split up your responsibilities between a research or tech lead and a manager. The manager takes on more of the management responsibilities, and the tech lead sets the direction for the team and ensures the technical work happens. But in that configuration, they have to get along really well and be on the same page to effectively divide responsibilities. In particular, the manager should still have a detailed understanding of what we're trying to do. Ideally, we'd want someone who can do both roles in one. So the background would be like leading a research team at some other company or in some other branch of machine learning, or being a manager in some other domain and then switching to an individual contributor on a large language model project. Or there's also a path where you're a postdoc somewhere with a small research team, working day-to-day, coding heavy, running lots of experiments with language models or reinforcement learning. These are all possible profiles. But the bigger filter is that you should really care about the problems we're trying to solve, be really good at coding, and be really good at machine learning.
OpenAI 必须处理的一个令人印象深刻且困难的事情就是让芯片和算力良好高效地工作。这些是巨大的算力聚合,让它们工作的工程一点也不简单。而专门为机器学习目的让它工作又增加了自身的复杂性。你们在招聘做那方面工程的人吗?
One of the impressive and difficult things that OpenAI has had to work on is just getting the chips and getting the compute to work well and efficiently. These are enormous aggregations of compute, and the engineering to get that to work is not at all straightforward. And getting it to work for ML purposes specifically adds its own complications. Are you hiring people to do that engineering side of things?
当然。但主要在超级对齐团队,我们要处理的是更像是运行这些大规模实验的基础设施的消费者。所以超级对齐团队的人需要熟悉调试大型分布式系统,因为如果我们在 GPT-4 上做微调运行,那就是这样一个系统。调试起来不容易。但我们不必构建大型语言模型基础设施,因为它已经存在,而且其他人在做那方面的工作。
Definitely. But mostly on the superalignment team, what we'll be dealing with is more like being consumers of the infrastructure that runs these large-scale experiments. So people on superalignment need to be comfortable debugging large distributed systems, because if we're doing a fine-tuning run on GPT-4, it is such a system. It's not easy to debug. But we don't have to build the large language model infrastructure because it already exists and other people are working on that.
申请流程是怎样的?
What does the application process look like?
非常简单。你去 openai.com/careers,往下翻,找到标题中有「超级对齐」的职位,点击它,提交简历并说明你为什么想从事这个工作,就这样。然后我们会看到。
It's very simple. You go on openai.com/careers, scroll down, find the roles that have 'superalignment' in the title, click on it, submit a CV and say why you want to work on this, and that's it. Then we'll see it.
流程中还有其他步骤吗?
Are there any further steps to the process?
有的。我们遵循的一般面试流程是:技术筛选、与团队某人的介绍性聊天,以及现场面试,包括两到四次编码或机器学习面试,以及一次文化契合度面试。但根据职位或背景,可能会略有不同。
Yes. The general interview process we follow is: there's a tech screening, an intro chat with someone from the team, and an on-site process with two to four coding or ML interviews and a culture fit interview. But depending on the job or background, it might look slightly different.
你们是期望可能招聘 20 人,然后长期只保留 10 人,还是更倾向于招聘那些你预计大部分都能成功的人?
Are you expecting to maybe hire 20 people and then only keep 10 in the long run, or is it more that you'll try to hire people who you mostly expect to work out?
我们想真正投资于我们招聘的研究人员,所以更倾向于第二种。
We want to really invest in the researchers we're hiring, so it's more the second one.
有没有办法传达门槛是什么?人们可能既过度自信又缺乏自信,如果有人真的很优秀但觉得自己不应该申请,那会很糟糕。所以如果有更明确的方式传达谁应该申请,那会很有用。
Is there a way to communicate what the bar is? People could be both overconfident and underconfident, and it could be bad if someone would be really good but they don't feel like they should apply. So if there's any more explicit way of communicating who should apply, that could be useful.
也许最重要的是:如果你有疑问,请申请。漏掉一个优秀候选人的代价远高于误招一个不合适的人的代价。
Maybe the most important thing is: if you're in doubt, please apply. The cost of a false negative is much higher than the cost of a false positive.
你之前已经稍微做过这个了,但你想直接说明为什么优秀的人应该申请和你一起在超级对齐团队工作吗?
You've slightly already done this earlier, but do you want to directly make the pitch for why amazing people should apply to work with you on the superalignment team?
简而言之,我认为这是最重要的问题之一。我们必须解决它,这不是可选的。我们想做非常雄心勃勃的事情。我们设定了在四年内实际解决它的目标。我们是认真的。所以如果你想在一个由高度积极、有才华的人组成的团队中工作,他们真正试图解决雄心勃勃的问题,并且有大量资源可用,那么这里就是你要去的地方。此外,我们处于技术的最前沿,OpenAI 真正支持我们。我们解决这个问题的机会和任何人一样好,甚至更好。我认为我们应该真正去做,全力以赴,你可以让这成为现实。这将非常令人兴奋。
In short, I think this is one of the most important problems. We really have to get this right; it's not optional. We want to do really ambitious things. We've set ourselves the goal to actually solve it in four years. We're serious about that. So if you want to work in a team of highly motivated, talented people who are really trying to solve ambitious problems and have a lot of resources to do so, this is the place to go. Also, we are at the state of the art of the technology, and OpenAI is really backing us. We have as good a shot at the problem as anyone else, if not more. I think we should just really do it and go for it, and you could make that happen. It'll be really exciting.
你们还需要非机器学习和非研究的人吗?比如运营、沟通、法律等其他部门?还是说,对于这些,你只需要一般性地申请 OpenAI,而不是专门申请对齐团队?
Do you also need any non-machine learning and non-research people? For operations, communications, legal, these other groups? Or are they maybe for that you'll just have to apply to OpenAI in general rather than the alignment team specifically?
是的。我也非常希望有更多真正关心对齐问题和人工智能未来的人加入。只要申请 OpenAI,任何职位都可以。帮助我们实现它。
That's right. I'm generally also really excited to have more people who really care about the alignment problem and the future of AI going well. Just apply to OpenAI, whatever role. Help us make it happen.
那个未来的现实。而且我认为 OpenAI 有很多人真的关心这个问题,但关心重要问题的人越多越好。我觉得这样更好。谈话中提到了政策问题。我知道 OpenAI 的政策团队有一些非常出色的人。没错。我还可以说出其他一些团队。所以我认为 AI 治理或政策研究团队在危险能力评估方面做了非常出色的工作,并且实际上在努力达成关于何时应该停止的协议。还有系统安全团队,他们实际上在努力改进我们现有模型的对齐和安全性,比如改进拒绝机制、修复越狱、改进监控。所有这些问题都非常重要。对于一些可能对我们必须解决的长期问题持怀疑态度、并且想做一些能立即产生影响的事情的听众来说,这些是很好的团队可以加入。我对他们正在做的事情感到兴奋。当然,OpenAI 还有很多其他团队在做重要的工作,比如改进 RLHF、改进 ChatGPT、法律、沟通、招聘。有很多事情要做。我们专注于试图找出如何对齐超级智能,但正如我们讨论过的,这不是我们唯一需要做的事情。
That future reality. And I think there are a lot of people at OpenAI who really care about this, but just more people who care about the important problems. I think the better. Some policy issues have come up through the conversation. I know there are some really amazing people on the policy team at OpenAI. That's right. I can name some other teams. So I think the AI governance or policy research team is doing really excellent work on dangerous capabilities evaluations and actually trying to get agreements about when should we all stop. And there's the system safety team that actually tries to improve alignment and safety of models we have right now, like making the refusals better, fixing jailbreaking, improving monitoring. All these problems are really important. And for some listeners who might be more skeptical about the long run problems we have to solve and want to do something that has impact right now, these are great teams to join. I'm excited for what they're doing. And then of course there are a lot of other teams at OpenAI that are doing important work, like improving RLHF, improving ChatGPT, legal, communications, recruiting. There are a lot of things to do. We are focusing on trying to figure out how to do alignment of superintelligence, but as we've discussed, it's not the only thing we need.
如果有人不愿意申请,因为他们担心参与可能会增强能力,并且他们认为加速能力研究是一件坏事,你会对他们说什么?
If someone were reluctant to apply because they were scared that getting involved might enhance capabilities and they were someone who thought that speeding up capabilities research was a bad thing, what would you say to them?
如果你不想那样做,就不要申请能力团队。说得对。所以我认为很明显,在超级对齐团队工作不会在全球层面上对能力进步产生有意义的贡献。我不想承诺我们做的任何事情都不会对能力产生影响。我认为一些最大的对齐胜利也会产生这些影响,这是真实且不可避免的。我还认为,特别是在 EA 社区中,有很多犹豫:如果我进入机器学习或在某处做机器学习工程工作,我可能会稍微加速时间线,如果我那样做会很糟糕。我认为这种推理大大低估了你在提升技能的同时做这些工作一段时间所获得的职业资本增长和技能增长,然后你以后可以转向对齐。总的来说,有太多人在做能力方面的工作,多一个或少一个不会让进展快多少。但在对齐方面没有那么多人,所以作为一个人在对齐方面工作,你实际上可以产生更大的影响。
If you don't want to do that, don't apply to the capabilities team. Fair enough. So I think the obvious thing is that working on the super alignment team is not going to meaningfully contribute to capabilities progress on any kind of global level. I don't want to promise that nothing we do will have any capabilities impact. I think some of the biggest alignment wins will also have some of these effects, and that's real and unavoidable. I think also in the EA community specifically, there's a lot of hesitation around: if I get into ML or do an ML engineering job somewhere, I might accelerate timelines a little bit, and it would be so bad if I did that. I think that kind of reasoning really underestimates the career capital growth and the skills growth that you would get by doing some of these jobs for a while while you're skilling up, and then you can switch to alignment later. In general, there are so many people working on capabilities that one more or less won't make it go that much faster. But there are not that many people in alignment, so as one person working on alignment, you can actually make a much larger difference.
每当这个话题出现时,我们都会链接到我们的文章:「如果你想降低 AI 风险,你是否应该担任推进 AI 能力的角色?」在那篇文章中,我们向广泛的人提出了这个问题,他们有不同的观点。但我认为你给出的推理,即你在能力研究方面的比例增长相对于你在对齐研究方面的比例增长非常小,再加上个人技能提升以及以后在职业生涯中使用这些技能的好处,至少在这种情况下对我来说似乎很清楚。
As we always do when this topic comes up, I'll link to our article: 'If you want to reduce AI risk, should you take roles that advance AI capabilities?' There we have responses from a wide range of people who we asked this question to, who have a range of views. But I think the reasoning you've given, that your proportional increase in capabilities research would be very small relative to the proportional increase in alignment research you would make, plus all the benefits from skilling up personally and then being able to use those skills later in your career, it seems pretty clear to me in this case at least.
OpenAI 的文化有哪些独特之处是人们应该了解的?有没有一种特定类型的性格在其中特别成功?
What are the distinctive things about OpenAI's culture that people should be aware of going in? Is there a particular kind of character that really thrives in it?
我认为我们通常希望真正欢迎各种不同的人和不同的性格。关于如何解决这个问题,思想越多样化越好。正如许多人之前说过的,这个问题也有很多非机器学习的方面。所以如果有人有非传统背景并转入机器学习,或者有不典型的起源故事,我认为这非常有价值。总的来说,我非常关心建立一个温暖、友好、包容的团队文化,同时也为人们创造很多心理安全感,让他们对我们正在做的事情或我们的总体方法发表尖锐的看法。我们需要合作解决问题。这不是关于谁能获得荣誉之类的事情。这个问题只需要被解决。
I think we generally want to be really welcoming to all kinds of different people and all kinds of different characters. The more diversity of thought on how to go about this problem, the better. As many people have said before, there are also so many non-machine learning aspects to this problem. So if especially somebody has a non-traditional background and switched into ML, or has a non-typical origin story, I think that's super valuable. In general, I care a lot about having a team culture that is really warm, friendly, and inclusive, but also creates a lot of psychological safety for people to voice spicy takes on some of the things we're doing or our approach in general. We need to collaborate to solve the problem. It's not about who can get the credit or something. This problem just needs to get solved.
如果一个非常有才华的人想转向技术对齐工作,但由于某种原因他们无法加入你的超级对齐团队,你还有没有其他地方会很兴奋地推荐他们申请?
If a really talented person wanted to switch into working on technical alignment but for some reason it was impossible for them to go join you on the super alignment team, is there anywhere else that you'd be really excited for them to apply?
我认为还有其他一些 AI 实验室做得很好,比如非常酷的工作:Google DeepMind 或 Anthropic。还有一些学术实验室在做非常酷的事情,比如伯克利、斯坦福或牛津。我会考虑申请那些地方。我也认为,当我们不得不拒绝非常有才华的人时总是很难过,但我们是一个小团队,不能雇佣所有人。有时人们还没有完全准备好,专注于更多的技能建设和职业资本投资是好的。我认为这也是一个非常有效的策略。总的来说,经历过流程的人通常低估了在另一家公司担任研究工程职位、提升技能和学习很多东西的价值。有很多机会可以这样做。
I think there are other AI labs that I think are doing a good job, like really cool work: Google DeepMind or Anthropic. And there are other academic labs doing really cool stuff, like at Berkeley, Stanford, or Oxford. I think I would consider applying to those. I think also, it's always very sad when we have to turn down really talented people, but we are a small team and can't hire everyone. Sometimes people aren't quite ready, and it's good to focus on more skill building and career capital investment. I think that's also a really valid strategy. All in all, people who go through a pipeline generally underestimate how valuable it is to take a research engineering job at another company, skill up, and learn a bunch of things. There are a lot of opportunities to do that.
就实际问题而言:可以远程工作吗?你们能为非美国公民提供签证担保吗?
Just on practical questions: is it possible to work remotely? And can you sponsor visas for people who aren't U.S. citizens?
是的,我们肯定提供签证担保。我们通常不鼓励远程工作,因为几乎整个团队都在……
Yes, we definitely sponsor visas. We generally do not encourage remote work because almost the entire team is in...
在我们继续之前,你还有什么想说的吗?
Any other points you want to make before we push on?
非常感谢你让我在这里介绍这些职位。我很期待更多关心这个问题和未来的人来帮助人类管理进入后 AGI 世界的过渡。谢谢你做这些。
Thank you so much for letting me pitch these roles here. I'm really excited for more people who care about this problem and the future to help manage humanity's transition into a post-AGI world. Thank you for doing this.
我们已经超时了,我耽误了你不少时间。我相信你有很多事要做来启动这个项目。但在结束前,你有一部最喜欢的科幻作品吗?
We've already gone overtime, I've been keeping you for a while. I'm sure you have a lot to do setting up this project. But before we go, do you have a favorite piece of science fiction?
我真的很喜欢格雷格·伊根的书。很多都很老了。《置换城市》是我最喜欢的之一。他探讨的许多想法在当时感觉很遥远,但现在在很多方面都更贴近现实。你能感觉到越来越多奇怪的科幻想法正在变成现实。我也喜欢他试图描绘一个长期来看社会可能的美好图景。
I really like the Greg Egan books. A lot of them are really old. Permutation City was one of my favorites. Many of the ideas he plays with felt out there at the time, but now they seem much closer to home in a lot of ways. You can feel more and more weird sci-fi ideas becoming reality. I also like that he tries to paint a positive view of what society could look like in the long run.
你的生活比那部科幻作品描绘的更奇怪还是更不奇怪?我不了解《置换城市》。你能快速告诉我们它讲的是什么,以及它是否比你的处境更奇怪吗?
Is your life weirder or less weird than what is portrayed in that piece of science fiction? I don't know about Permutation City. Could you quickly tell us what it's about and whether it's weirder than your own situation?
绝对没那么奇怪。《置换城市》这本书探讨了上传意识、拥有数字人类副本、生活在数学宇宙中的想法及其影响。例如,虚拟人类可以重写自己的代码,这些我们还做不到。也许在某些方面 AI 可以做到,或者在近期或中期未来,如果我们取得可解释性进展,我们可以重写它自己的神经网络部分。这是非常前沿的科幻,这就是它如此酷的原因。
It's definitely less weird. Permutation City is a book that plays with the idea of uploading, having digital copies of humans, living in a mathematical universe, and the implications of that. For example, virtual humans can rewrite their own code, things we can't do yet. Maybe in some ways AI can do it, or in the near or medium future, we could rewrite parts of its own neural network if we make interpretability progress. It's very out there science fiction, that's what makes it so cool.
我确实觉得我们一直生活在科幻中。但这不算什么,未来会变得更加奇怪。
I do feel like we've been living through a science fiction. But this is nothing, it's going to get so much weirder.
是的,我们在 2030 年代或 2040 年代有得期待了。我不知道具体会怎样,但我保证按今天的标准会很奇怪。
Yeah, we have that to look forward to in the 2030s or 2040s. I don't know exactly how it's going to go, but I promise it'll be weird by today's standards.
祝你的项目好运。我真的很期待看到它的进展。我今天的嘉宾是 Jan。非常感谢你来到 80,000 Hours 播客。
Best of luck with the project. I really look forward to seeing how it comes along. My guest today has been Jan. Thanks so much for coming on the 80,000 Hours podcast.
非常感谢你邀请我。
Thank you so much for having me.