Paul Christiano on AI Safety, Divestment, and Future Civilizations
打开互动全文版(中英对照 + 朗读 + 问答)→Rob Wiblin 采访 Paul Christiano,讨论 AI 对齐、撤资有害公司的有效性,以及肌酸补充剂和留给未来文明的信息等推测性想法。
Rob Wiblin interviews Paul Christiano about AI alignment, the effectiveness of divesting from harmful companies, and speculative ideas like creatine supplements and messages to future civilizations.
嘿,听众们,这里是 80,000 Hours 播客,每周我们都会深入探讨一个世界上最紧迫的问题,以及你如何利用自己的职业生涯来解决它。我是 Rob Wiblin,80,000 Hours 的研究总监。去年,在第 44 集中,我采访了 Paul Christiano 近 4 个小时,内容涵盖了他对一系列广泛话题的看法,包括他预计 AI 将如何影响我们在 21 世纪的生活,以及他认为我们如何能增加这些影响真正积极的可能性。那一集非常受欢迎,而 Paul 是一位极具创造力、思维非常广泛的思考者,他经常撰写关于从完全合理到特别奇怪的各种主题的博客文章。所以我认为让他回来聊聊他最近在想什么会很有趣。在合理的一面,我们最终讨论了如何从有害公司撤资可能比大多数人(包括我)之前认为的更有效地改善世界。在更奇怪的一面,我们思考了一下,是否有什么信息我们可以留给未来的文明,如果人类灭绝但智能生命在遥远的未来在地球上重新进化,这些信息可能会帮助他们。我们还讨论了一些更具推测性的论文,这些论文表明服用肌酸补充剂可能会让人更敏锐,或者待在充满二氧化碳的闷热房间里可能会让人暂时变笨。老实说,我们只是很开心地聊了一些我们个人觉得很有趣的事情。不过,在更实际有用的一面,我得到了 Paul 对我几个月前与 DeepMind 的 Pushmeet Kohli 进行的第 48 集采访的反应。我应该提醒大家,回顾起来,这一集有点术语密集,对于新听众来说可能更难跟上。随着时间的推移,这将越来越难以避免,因为我们想深入探讨之前几集已经介绍过的主题,但这次我们走得比我理想中更远了一点。如果听众先听第 44 集对 Paul Christiano 的采访(即 Paul Christiano 博士谈我们将如何把未来交给 AI 并解决对齐问题),可能会收获更多。但我认为即使你没听过那一集,这一集的大部分内容仍然可以理解。这一集还有我们制作的第一个花絮。我鼓励 Paul 尝试录制一段关于哲学子领域决策理论及其产生的异端思想(如超理性和非因果合作)的采访。我原计划花半小时讨论这个,但这真的很愚蠢。仅仅清晰地概述决策理论的问题和提出的解决方案就需要整整一个小时,如果我们能说清楚的话。而解释和论证各种非正统解决方案的可能含义可能还需要一两个小时,而且没有白板真的很难做到。所以在那段结尾,我们觉得这或多或少是一场灾难,尽管对于合适的听众来说可能是一场有趣的灾难。我们表现得有点疯狂,但我相当确定我们不是。所以,如果你想有点困惑,听听技术采访失败是什么样子,你可以在节目说明中找到那段 MP3 的链接。如果你对决策理论的理解水平和我进入对话时一样,你甚至可能学到一些东西,但我不特别推荐把它作为时间的好用途。相反,我们会在未来的某一集中回来,对决策理论进行适当的处理。好了,闲话少说,有请 Paul Christiano。我今天的嘉宾是 Paul Christiano 博士,应广大听众要求,第二次出现在 80,000 Hours 播客上。Paul 在加州大学伯克利分校获得了理论计算机科学博士学位,现在是 OpenAI 的一名技术研究员,致力于将人工智能与人类价值观对齐。他在 ai-alignment.com 上撰写关于这项工作的博客。
Hey listeners, this is the 80,000 Hours podcast, where each week we have an unusually in-depth conversation about one of the world's most pressing problems and how you can use your career to solve it. I'm Rob Wiblin, director of research at 80,000 Hours. Last year, for episode 44, I interviewed Paul Christiano for nearly 4 hours about his views on a pretty wide range of topics, including how he expects AI to affect our lives in the 21st century and how he thinks we can increase the chances that those impacts are really positive. That episode was very popular, and Paul is a highly creative and pretty wide-ranging thinker who's always writing new blog posts about topics that range from the entirely sensible to the especially strange. So I thought it would be pretty fun to get him back on to talk about what he'd been thinking about lately. On the sensible side, we end up talking about how divesting from harmful companies might be a more effective way to improve the world than most people, including me, had previously thought. On the stranger side, we think a bit about whether there's any messages that we could leave future civilizations that might be able to help them out, should humanity go extinct but intelligent life then re-evolve on Earth at some point in the far future. We also talk about some more speculative papers which suggest that taking creatine supplements might make people a bit sharper, or that being in a stuffy carbon dioxide-filled room might make people temporarily a bit stupider. Honestly, we just have a lot of fun chatting about some things we personally find pretty interesting. On the more practically useful side of things, though, I get Paul's reaction to my interview with Pushmeet Kohli over at DeepMind for episode 48, which came out a few months back. I should warn people that in retrospect this episode is a bit heavy on jargon and might be harder to follow for someone who is new to the show. That's going to get more difficult to avoid over time as we kind of want to dig deeper into topics that we've already introduced in previous episodes, but we went a bit further than I'd ideally like this time around. Folks might get more out of it if they listen first to the previous interview with Paul Christiano back in episode 44, that's Dr. Paul Christiano on how we'll hand the future of AI and solving the alignment problem. But I think the majority of the episode should still make sense even if you haven't listened to any of that one. This episode also has the first outtake we've made so far. I encouraged Paul to try recording a section in the interview on a subfield of philosophy called decision theory and some heterodox ideas that have come out of it, like superrationality and acausal cooperation. I planned to spend half an hour on that, but that was really very silly of me. It would need a full hour just to clearly outline the problem of decision theory and the proposed solutions, if we could do it clearly at all. And explaining and justifying the possible implications of the various unorthodox solutions that are out there could go on for another hour or two, and indeed it's really hard to do at all without a whiteboard. So at the end of that section, we thought it was more or less a train wreck, though potentially quite a funny train wreck for the right listener. We do come across as a touch insane, which I'm fairly sure we're not. So if you'd like to be a bit confused and hear what it sounds like for a technical interview to not really work out, you can find a link to the MP3 for that section in the show notes. If you happen to have the same level of understanding of decision theory that I did going into the conversation, you might even learn something, but I can't especially recommend listening to it as a good use of time. Instead, we'll come back and give decision theory a proper treatment in some other episode in the future. All right, with all of that out of the way, here's Paul Christiano. My guest today is Dr. Paul Christiano, back by popular demand, making his second appearance on the 80,000 Hours podcast. Paul completed a PhD in theoretical computer science at UC Berkeley and is now a technical researcher at OpenAI working on aligning artificial intelligence with human values. He blogs about that work at ai-alignment.com.
欢迎你,保罗。感谢你来做客播客。
Welcome, Paul. Thanks for coming on the podcast.
感谢再次邀请我。我希望聊聊你最近博客里写的一些有趣内容,以及 AI 可靠性和鲁棒性研究的新进展。但首先,你目前在做什么,为什么你认为这是重要的工作?
Thanks for having me back. I hope to talk about some of the interesting things you've been blogging about lately, as well as what's new in AI reliability and robustness research. But first, what are you doing at the moment and why do you think it's important work?
我大部分时间都在 OpenAI 做技术性的 AI 安全。我认为基本情况与一年前相似:构建的 AI 系统不按我们的意愿行事,把长期未来推向我们不希望的方向,这似乎是搞砸长期未来的主要方式之一。这基本上还是对的。我可能稍微倾向于认为这只是总问题中较小的一部分,但仍然占很大比重。这对我来说是直接参与其中的很自然的方式,所以我会继续努力。这是大方向。我们大概会深入很多细节。
So I guess I'm spending most of my time working on technical AI safety at OpenAI. I think the basic story is similar to a year ago: building AI systems that don't do what we want them to do, that sort of push the long-term future in a direction that we don't like, seems like one of the main ways we can mess up our long-term future. That still seems basically right. I maybe moved a little bit more towards that being a smaller fraction of the total problem, but still a big chunk. It seems like a really natural way for me to work on it directly, so I think I'm just going to keep hacking away at that. Yeah, that's the high level. I think we're going to get into a lot of the details probably.
关于 AI 安全研究以及你对先进 AI 发展走向的总体预测,你的观点有变化吗?如果有,是怎样的变化?
When it comes to AI safety research and your general predictions about how advanced AI is going to play out, have your opinions shifted at all? And if so, how?
我觉得过去一年没什么大惊喜,事情在逐渐稳定。这可能是一个更广泛趋势的一部分:五年前我的观点每年都大幅波动,三年前波动小了些,去年波动更小了。所以我的观点没有太大变化。在 AI 整体进展方面,我们没有遇到大的意外,无论是向下还是向上。我们看到的情况与 AI 快速发展以及可能耗时很长的担忧都大体一致。在 AI 对齐的方法上,我对需要做什么的理解更加明确了。它继续从应该做什么的宽泛想法转向具体的实施团队。这种情况在持续,但没有大的意外。
I think the last year has felt a lot like no big surprises, things sort of settling down. Maybe this has been part of a broader trend: my views 5 years ago were bouncing around a ton every year, 3 years ago bouncing around a little bit, and now the last year has bounced around even less. So I think my views haven't shifted a huge amount. We haven't had either big downward or upward surprises in terms of overall AI progress. We've seen things broadly consistent with both concerns about AI being developed very quickly and the possibility of it taking a very long time. In terms of our approach to AI alignment, my understanding of what there is to be done has solidified a little bit. It continues to move from broad ideas of what should be done to here are the particular groups implementing things. That's continuing to happen, but there haven't been big surprises.
上次我们谈了很多不同的方法,包括通过辩论实现 AI 安全,让不同的 AI 互相辩论,然后我们能够判断哪个是对的。这种方法或者其他我们谈过的方法有进展吗?
Last time we spoke about a bunch of different methods, including AI safety via debate, having different AIs debate one another and then we're in a position to adjudicate which one is right. Is there any progress on that approach or any of the other ones we spoke about?
是的,我在 OpenAI 的一个子团队工作,主要研究辩论安全和放大。过去一年,很多工作集中在建设能力和基础设施上,让这些方法成为可能,比如扩展语言模型并与好的大型语言模型集成,这样就能理解人类在交谈或回答问题时的推理过程,试图达到我们真正能观察到感兴趣现象的程度。我认为,在 OpenAI 内部不同部门以及跨组织之间,关于可能方法的思考出现了一些趋同。对于思考这个长期问题的人来说,我们主要考虑放大和辩论。从理论上看,这两种技术应该非常相似。它们可能对短期实验的侧重点不同,但随着我们尝试,那些从放大角度出发的人也在进行更接近辩论视角的实验,反之亦然。所以这方面的重大分歧变少了。同样,DeepMind 思考长期安全的人独立思考,我觉得我们之间的差距也变小了。这也许是好事:更容易沟通,达成共识,对所做之事有更多共同理解。与一年前相比,事情在稳定和成熟。但距离任何正常的学术研究领域还很远,远未达到。人们分歧更大,或者观点非常不同,缺乏常识或成熟的研究方法,大家都期望取得进展。不过,它正在朝那个方向发展。
Yeah, so I work on a sub-team at OpenAI that broadly works on that idea: safety via debate as well as amplification. Over the last year, a lot of the work has been on building up capacity and infrastructure to make those things happen, scaling up language models and integrating with good large language models, so things that understand some of the reasoning humans do when they talk or answer questions, trying to get to the point where we can actually start to see the phenomena we're interested in. I think there's probably been some convergence in terms of how different parts within OpenAI, and also across organizations, have been thinking about possible approaches. For people thinking about this really long-term problem, we mostly think about amplification and debate. There's a sort of on-paper argument that those two techniques ought to be very similar. I think they maybe suggest different emphasis on which experiments you run in the short term, but as we've been trying things, both the people who started more on the amplification side are running experiments that look more similar to what you might suspect from the debate perspective, and vice versa. So I think there are fewer big disagreements about that. Similarly, independent thinking among people thinking about long-term safety at DeepMind, I feel like there's less gap between us now. Maybe that's good: it's easier to communicate and be on the same page, more shared understanding of what we're doing. Compared to a year ago, things feel like they are settling down and maturing. It's still a long way from being like almost any normal field of academic inquiry, it's nowhere close to that. People disagree more or have very different perspectives, and they less have a common sense or a mature method of inquiry which everyone expects to make progress. It is moving in that direction, though.
这可能有点随意,但你是否觉得学术领域常常因为固化了特定方法、特定证据和特定世界观而受限,导致对其他选择视而不见?也许这个研究领域更自由一些反而是优势?
This is maybe a bit random, but do you feel like academic fields are often held back by the fact that they codify particular methods, particular kinds of evidence, and a particular worldview that blinkers them to other options? And maybe it's an advantage to have this foot of research be a bit more freeing in devas?
这是个有趣的问题。我通常认为一个学术领域是由一套工具和对什么是进步的理解来定义的。如果你认为领域是由问题定义的,那么说领域被蒙蔽或价值被搁置是有道理的。如果你认为领域是由工具定义的,那么那就是他们带来的东西。从这个角度看,这既不好:你不能使用现有的工具集,这很糟糕,而且我们也不清楚最终解决方案在多大程度上应该使用现有工具。这不好。同时,没有针对这类研究的成熟工具也不太好。我更多是这么想的。如果你认为学术领域是唯一回答这些问题的人,那么许多学术领域并不擅长回答这些问题。也许它们并非为此最优设置。我已经转变了,不再主要那样看待学术领域。比如,经济学的方法有问题,但这些问题被使用不同方法的其他领域覆盖了。这是希望所在。我认为经济学是个有趣的例子,因为它有很多问题。
I think that's an interesting question. I guess I would normally think of an academic field as characterized by the set of tools and this understanding of what constitutes progress. If you think of the field as characterized by problems, then it makes sense to talk about the field being blinkered in this way or having value left on the table. If you think of the field as characterized by the set of tools, then that's kind of the thing they're bringing to the table. So from that perspective, it's both bad: you can't use some existing set of tools, that's a bummer, and it's not clear how much we should ultimately expect the solution to look like using existing tools. That's bad. It's also a little bit bad to not have yet mature tools that are specific to this kind of inquiry. I think that's more how I think of it. Many academic fields are not good at answering the set of questions if you think of them as the only people answering those questions. Maybe they're not really set up optimally to do that. I think I've shifted to not mostly thinking of academic fields that way. So, you know, economics has problems with its method, but then those are covered by other fields that use different methods. That would be the hope. I think economics is an interesting case since there are a bunch of problems.
各个领域确实存在这样的情况:有很多问题大致属于经济学范畴,经济学家有一套工具,如果某个问题名义上属于经济学领域,但与其工具不匹配,那就会陷入一种奇怪的境地。我认为经济学也可能是因为其领域内问题范围太广。是的,我觉得这种区分并不明显。经济学有点像众所周知的帝国主义领域,想要殖民所有它能触及的问题。然后有时方法可能并不适合它想要解决的那些问题。不过从某种意义上说,如果你把一个领域看作有一套工具,那么去发现其他确实适合这些工具的问题是很合理的。反过来也有一种情况:你不应该过分宣称对这些问题的所有权。你应该愿意说,看,这些问题我们传统上回答了,但也有其他人,有时可以用其他方式回答。
Fields do have this like there are a bunch of problems sort of fit in economics and like there's a set of tools economists use and if there's like a problem that fits in nominally the economics domain like under their purview which is not a good fit for their tools then you're in sort of like a weird place. I think economics also may be because of like there being this broad set of problems that fit in their domain. Like yeah, I think it's not this distinction is not obvious. Like there's some, it's an imperial field kind of notoriously that wants to go and colonize all questions that it can touch on. Yeah, and then sometimes I guess the method might be ill-suited to those questions that it wants to tackle. Yeah, although in some sense if you view it as like a field has a set of tools that it's using, it's very reasonable to be like going out and finding other problems that are actually, if you're actually correct about them being amenable to those tools. And there's also a thing on the reverse where like you don't want to be like really that staking claim on these questions. So you should be willing to say, look, these are questions that we've traditionally answered, but there are other people, like sometimes those can be answered in other ways.
这是一种有趣的关于学术领域问题的框架,与其说领域本身不好,不如说它可能解决错了问题,或者解决的方法与问题不匹配。我经常思考这个问题,也许是因为在计算机科学中,问题更清晰,不是像‘这里有一个问题,这个问题属于某个领域’,而是有几种不同的方法。比如有统计学背景的人,有理论家,有各种实践者或实验者,你可以看到子领域有不同的方式来处理这些问题,你大致理解这个子领域会以这种方式解决这个问题,这是一种合理的分工。
It's an interesting framing of problems with academic fields that it's kind of not so much that the field is bad but maybe that it's tackling the wrong problem or yeah it's tackling problems that are mismatched to the methods. I think about this a lot maybe because in computer science you more clearly have problems which like it's not so much staked out as like here's a problem and this problem fits in a domain, it's more like there are several different approaches. Like there are people who come in with a statistics training and there are people who come in as theorists and there are people who come in as various flavors of practitioners or experimentalists and they sort of you can see subfields have like different ways they would attack these problems and it's more like you sort of understand like this subfield is going to attack this problem in this way and like it's a reasonable kind division of labor.
是的,所以我们回溯一下。你提到了进行实验。具体是什么样的实验?
Yeah, so let's back up. You talked about running experiments. What kind of experiments are they concretely?
是的,上次我们讨论时提到了三种大的不确定性或进展空间。其中一种与实验关系不大,是解决概念性问题,比如如何找到可扩展的对齐方法。另外两个困难都非常适合不同类型的实验。一类是涉及人类的实验,你开始理解人类推理的一些特征。比如我们对人类推理有一些期望:我们希望从某种意义上说,只要有足够的时间和资源,人类是通用的,能够回答非常广泛的问题,只要他们有足够的时间、足够的反思空间。所以有一类实验就是为了理解这一点:在什么意义上这是真的,在什么意义上这是假的。这是我很感兴趣的一类实验。比如 OpenAI 最近开始招聘人员,刚雇了两个人来扩大这些实验。我认为 Anthropic 一直专注于这些实验,并开始真正扩大他们的工作。这是第一类实验。还有第二个或第三个困难,是理解对齐的理论想法以及人类推理的实际情况如何与机器学习结合在一起。最终,我们想用这些想法来产生可用于训练机器学习系统的目标,这需要实际处理机器学习系统如何工作的许多细节。所以一些实验直接测试这些细节,比如你是否可以使用这种目标,机器学习系统能否学习这种模式或行为。还有一些实验更像是你期望通过迭代就能奏效的类型。你大致期望有某种方法可以将语言模型应用于这类任务,但需要思考如何做,并尝试几次。
Yeah, so I think last time we talked we discussed three kinds of big uncertainties or room for making progress. One of them which isn't super relevant to experiments is like figuring out conceptual questions about how we are going to approach find some scalable approach to alignment. The other two difficulties both were very amenable to different kinds of experiments. So one are experiments involving humans where you start to understand something about the character of human reasoning. Like you understand so we have some hopes about human reasoning. We hope that in some sense, given enough time or given enough resources, humans are like universal and could answer some very broad set of questions if they just had enough time, enough room to reflect. So it's like one class of experiments that's sort of getting at that understanding: in what sense is that true and in what sense is that false. And so that's a family of experiments I'm very excited about. Like maybe OpenAI has recently started hiring people, just hired two people who will be scaling up those experiments here. I think Anthropic has been focused on those experiments and is starting to really scale up their work. So that's one family of experiments. And there's a second difficulty or third difficulty which is understanding how both theoretical ideas about alignment and also the facts about how human reasoning work, how those all tied together with machine learning. So ultimately at the end of the day we want to use these ideas to produce objectives that can be used to train ML systems, and that involves actually engaging with a bunch of detail about how ML systems work. So some of the experiments are directly testing those details, so saying can you use this kind of objective, can machine learning systems learn this kind of pattern or this kind of behavior. And some of them are just experiments that are like maybe more in the family where you expect them to work if you just iterate a little bit. So you sort of expect there is some going to be some way that we can apply language models to this kind of task but we need to think a little bit about how to do that and take a few swings at it.
我看到 OpenAI 试图招聘社会科学家,并论证社会科学家应该对 AI 对齐研究更感兴趣。这是他们正在做的工作吗,运行这些实验或设计它们?
I saw that OpenAI was trying to hire social scientists and kind of making the case that social scientists should get more interested in AI alignment research. Is this the kind of work that they're doing, running these experiments or designing them?
是的,没错。我们最初的目标是招聘一个人担任这个角色。我想我们现在已经招到了,他们周一开始工作。他们将进行实验,试图理解如果我们想以某种方式将人类推理作为基本事实或黄金标准,我们该如何思考?我们如何思考在什么意义上我们可以扩展人类推理来回答难题?在什么意义上人类是正确性的良好判断者,或者激励两个辩论者之间的诚实行为?其中一些是经验性的问题,即人类在什么条件下能够完成某些任务。还有一些更像是概念性问题,人类只是你获得进展的方式,因为人类是我们唯一能够接触到的、非常擅长这种灵活、丰富、广泛推理的系统。
Yeah, that's right. So I think we hired, we're aiming initially to hire one person in that role. I think we've now made that hire and they're starting on Monday. So they will be doing experiments trying to understand like if we want to try and use human reasoning in some sense as a ground truth or gold standard, like how do we think about that? How do we think about in what sense we could scale up human reasoning to answer hard questions? In what sense are humans like a good judge of correctness or like incentivize honest behavior between two debaters? Some of that is like what are empirically the conditions under which humans are able to do certain kinds of tasks. Some of them are like more conceptual issues where humans are just the way you get traction on that because humans are the only systems we have access to that are very good at this kind of flexible, rich, broad reasoning.
我在推特和脸书上提到我要再次采访你,一位听众发来一个问题。他们听说你认为,即使我们没有坚实的技术解决方案来解决 AI 对齐问题,AI 接管并变得非常有影响力,事情仍有相当大的概率会顺利发展,或者宇宙仍然会有相当多的价值。如果这个理解正确,理由是什么?
I mentioned on Twitter and Facebook that I was going to be interviewing you again and a listener wrote in with a question. They'd heard I think that you thought there's a decent probability that things would work out okay or that the universe would still have quite a lot of value even if we didn't have kind of a solid technical solution to AI alignment and AI you know took over and was very influential. What's the reasoning there if that's a correct understanding?
是的,所以我认为有很多种方式可以最终得到做我们想做的事情的系统。一种方法,作为理论家,最吸引我的是在纸面上有非常好的理解:如何训练 AI 做你想做的事,我们在构建真正强大的系统之前就在抽象上解决了这个问题。这是乐观的情况,我们真正解决了对齐问题,真正搞定了。可能还有第二类,或者说一个广泛的谱系,我们在纸面上没有对完全通用的方法有很好的理解,但随着我们实际与这些系统打交道,我们可以尝试很多东西,看看什么有效,我们可以,你知道,如果我们
Yeah, so I think there's a bunch of ways you could imagine ending up with systems that do what we want them to do. So one approach which is you know as a theorist the one that's most appealing to me is to have some really good understanding on paper: here's how you train an AI to do what you want and we just sort of nail the problem in the abstract before we've even necessarily built a really powerful system. This is the optimistic case where we really solved alignment, it's really nailed. There's maybe a second category where you're like or this broad spectrum where we don't have a great understanding on paper of a fully general way to do this, but as we actually get experienced with these systems we get to sort of try a bunch of stuff, we get to see what works, we get to you know if we're
担心系统失败,我们可以尝试在很多特殊情况下运行它,扔各种东西给它,看看是否经过足够的压力测试后能正常工作。也许我们无法从原则上精确提取我们重视的东西,但我们可以构建足够好的代理指标。所以这是一大类情况,你并没有理论上的理解,但你还是能凑合应付。我认为这不是提问者想问的。还有一种更极端的情况,你尝试这样做但效果很差,而且在这个过程中你发现这些系统确实会失败,而且失败方式越来越灾难性。长远来看,我们认为这可能会非常糟糕。
Concerned about a system failing, we can try and run it in a bunch of exotic cases and just try and throw stuff at it and see if maybe if we stress test enough, something actually works. Maybe we can't really understand a principled way to extract exactly what we value, but we can do well enough at constructing proxies. So it's this giant class of cases where you don't really have an on-paper understanding, but you still sort of wing it. I think that's not what the asker was asking about. There's kind of a further case where you try and do that and you do really poorly, and as you're doing it, you realize these systems do just fail in increasingly catastrophic ways. Drawing the line out, we think that could be really bad.
我认为即使在最坏的情况下,你不仅没有理论上的理解,而且也无法很好地凑合,我仍然认为有超过三分之一的可能性一切都会好起来。这必须通过人们认识到问题、达成合理共识认为这是一个严重问题、并愿意在部署 AI 的方式上做出一些牺牲来实现。所以我认为至少在理论上,很多人会愿意说,如果到处推广 AI 会摧毁我们珍视的一切,那么我们愿意更谨慎地行事,或者在更窄的范围内推广,或者放慢开发速度。因此,人们会表现出足够的克制,以便有足够的时间来修补问题,使一切变得可以接受。所以这实际上是一个权衡:你表现出多少克制,以及你最终是获得清晰的理解还是凑合过去。如果凑合完全行不通,我们彻底完蛋,那么你必须表现出极大的克制,人们必须真正等待,直到我们利用 AI 变得更聪明、更能协调、更能解决这些问题,或者类似的情况。你必须等到那发生之后才能普遍部署 AI。所以我认为这仍然相当可能。我认为这是一个很多人不同意的地方。两端都有很多人:很多人更乐观,很多人认为人们不会踩到刀片上,不会让所有资源被抽走,也不会在灾难性失败会导致所有人死亡的情况下部署 AI。有些人直觉认为这不会发生,我们协调得足够好以避免这种情况。我对此并不完全认同。我认为这是一个非常困难的协调问题。我不知道,看起来我们肯定会失败。另一方面,有些人认为我们什么都协调不了,如果有一个按钮可以摧毁一切,或者某个拥有十亿美元的人可以按下按钮把事情搞砸,那么事情肯定会变得一团糟。我真的不知道。部分原因是我无知,部分原因是我对两种极端观点都持怀疑态度。我认为倡导这些观点的人和我一样对实际情况无知。我当然认为有些人拥有更多相关知识,如果他们理解技术问题,他们的估计会比我的更准确。但我有点悲观:如果情况真的非常非常糟糕,如果我们真的对对齐毫无理解,那么我会感到悲观,但不是极度悲观。
I think even in that worst case where not only you don't have an on-paper understanding, you can't really wing it very well, I still think there's certainly more than a third chance that everything is just good. And that would have to come through people probably understanding that there's a problem, having reasonable consensus that it's a serious problem, being willing to make some sacrifices in terms of how they deploy AI. So I think that at least on paper, many people would be willing to say if really rolling AI out everywhere would destroy everything we value, then we are happy to be more cautious about how we do that, or roll it out in a more narrow range of cases, or take development more slowly. And so people show restraint for long enough to kind of patch over the problems well enough to make things okay. So somehow there's a spectrum of how well some substitution between how much restraint you show and how much you were able to either ultimately end up with a clean understanding or wing it. One third is my number if it turns out the winging doesn't work at all, like we're totally sunk such that you have to show very large amounts of restraint and people have to actually just be like we're going to wait until there's some much smarter—we've either used AI to become much smarter, better able to coordinate, better able to resolve these problems—or something like that. You have to wait until that's happened before you're actually able to deploy AI in general. So I think that's still reasonably likely. I think that's a point where lots of people disagree. I think on both ends, so a lot of people are much more optimistic. A lot of people have the perspective that's like look, people aren't going to walk into razor blades and have all the resources in the world get siphoned away, or deploy AI in a case where a catastrophic failure would cause everyone to die. Some people have the intuition that's just not going to happen and we're sufficiently well coordinated to avoid that. I'm not really super on the same page there. I think it's a really hard coordination problem. I don't know, it looks like we could certainly fail. On the other hand, some people are like man, we can't coordinate on anything, and if there was a button you could just push to destroy things, or someone with a billion dollars could push to really mess things up, things would definitely get really messed up. And I just don't really know. In part this is just me being ignorant, and in part it's me being skeptical of both of the extreme perspectives. I think people advocating them are also about as ignorant as I am of the facts on the ground. I certainly think there are people who have more relevant knowledge and who could have much better calibrated estimates if they understood the technical issues than I do. But I'm kind of at some pessimism if things are really really bad, if we really really don't have an understanding of alignment, then I feel pessimistic but not like radically pessimistic.
是的,似乎有一个挑战:人们对技术安全性的信心会有差异,然后问题在于,认为技术最安全的人可能错了,因为大多数人不同意,而他们最有可能过早部署。
Yeah, it seems like a challenge there is that you're going to have a range of confidence about how safe the tech is, and then you have this problem that whoever thinks it's the safest is probably wrong about that, because most people disagree and they're the most likely to deploy it prematurely.
是的,我认为这在很大程度上取决于你得到的关于失败信号的类型。比如,你会有多少——我们可以讨论各种可能发生的险情。我认为这些信号越清晰,就越容易达成足够的共识。这是一点。第二点是,我们担心,或者说我担心一种特定的失败,它会严重破坏文明的长期轨迹。你可能处于这样的世界:让事情在实践中运转比让它们以长期保持我们意图的方式运转要容易得多。但你也可能想象这样的世界:一个长期会失败的系统在短期内也很可能是个大麻烦,这样人们就会更明显地看到问题。然后我认为一个重要因素是,我们确实有技术手段,特别是如果 AI 进步很大程度上由大量巨型计算集群驱动的话。在那种世界里,并不是任何人都能按下这个按钮。首先,只有少数参与者,比如愿意花费数百亿美元的人;其次,这些参与者有空间坐下来达成协议,这些协议可以不同程度地正式化。但不会是人们各自坐在小隔间里做决定。最坏的情况下,在那种世界里,最坏的情况是少数参与者可以互相交谈;最好的情况是少数参与者达成一致,比如在规范上,我们实际上会有某种监控和执法,以确保即使有人不同意共识,他们也无法把事情搞砸。
Yeah, I think it depends a lot on what kind of signals you get about the failures you're going to have. So like how much you have—yeah, we can talk about various kinds of near misses that you could have. I think the more clear those are, the easier it is for there to be enough agreement. That's one thing. A second thing is like we're concerned, or I'm concerned about a particular kind of failure that really disrupts the long-term trajectory of civilization. You could be in worlds where that's like the easiest kind of failure—that is, getting things to work in practice is much easier than getting them to work in a way that preserves our intention over the very long term. You could also imagine worlds though where a system which is going to fail over the very long term is also reasonably likely to be a real pain in the ass to deal with in the short term, in which case again it will be more obvious to people. And then I think a big thing is just we do have techniques, especially if we're in a world where AI progress is very driven by large amounts of giant computing clusters. In those worlds, it's not really like any person can press this button. It's like one, there's a small number of actors like people who are willing to spend say tens of billions of dollars, and two, those actors have some room to sit down and reach agreements, or which could be formalized to varying degrees. But it won't be like people sitting separately in boxes making these calls. At worst, it'll be like in that world, at worst it'll be like a small number of actors who can talk amongst themselves, and at best it'll be like a small number of actors who agree, like here in norms, we're going to actually have some kind of monitoring and enforcement to ensure that even if someone disagreed with the consensus, they wouldn't be able to mess things up.
你认为你或 OpenAI 这些年在对齐工作中有没有犯过什么有趣的错误?
Do you think you or OpenAI have made any interesting mistakes in your work on alignment over the years?
我绝对认为自己犯过很多错误,这些我更有资格谈论。所以有一类:我思考对齐问题很多年了,这么长时间积累了很多错误,也许很多并不那么切题。但我觉得大约四五年前,在我思考对齐问题的早期,我犯过一类智力上的错误,我们可以试着谈谈。但总的来说,我对对齐的整体看法自六年前以来发生了巨大变化。我会说这基本上是因为,六年前我无法正确推理——我在很多事情上推理错误。
I definitely think I have made a lot of mistakes, which I'm more in a position to talk about. So I guess there's one category: there's been a lot of years I've been thinking about alignment, that's a lot of time to rack up mistakes, maybe many of which aren't that topical. But there was a class of intellectual mistakes I feel like I made like four years ago, say, or five years ago, when I was much earlier in thinking about alignment, which we could try and get into. But I guess my overall picture of alignment has changed a ton since six years ago. And I would say that's basically because, you know, six years ago I was not able to reason about—I reasoned incorrectly about lots of things.
是的,我认为这是一个有趣的问题。也许有几个随机的评论:第一,似乎确实可以加速采用。如果你足够早地理解,你确实可以改变采用的时间线。所以你可以想象小团体将采用推进了 6 个月左右。考虑到有很多工程问题和概念困难是这种奇怪的小东西特有的,而它实际上在文明的总体机器和轨迹中发挥了很大作用,它确实得到了很好的杠杆作用。那个领域的进展似乎会为更快的整体技术进步带来异常高的回报。也许与此相关,我认为也有理由认为,如果一个小团体定位自己很好地理解那项技术,并推动它、投资它,他们很可能最终会处于一个未来情境中,他们赚了很多钱,或者处于一个位置,能够很好地理解一项在推出时没有多少人理解的重要技术。我认为这再次与他所怀疑的那种事情略有不同,但如果一个人考虑通过研究 AI、思考 AI 来获得杠杆,这似乎是计算中的一个重要部分。我确实认为对齐问题与你在电的背景下可能说的任何事情都不同。也许这就是为什么我主要不是试图做投资 AI 以便以后有影响力或赚很多钱的事情。我主要属于那种认为我们可以识别出一个异常清晰的问题,它似乎异常重要,并且可以专注于解决它。我认为这似乎应该有很多问号,但我不知道有任何历史案例适用。
Yeah, I think it's an interesting question. Maybe a few random comments: one, it does seem like you can accelerate the adoption. If you had an understanding early enough, you could really change the timeline for adoption. So you could imagine small groups having pushed adoption forward by 6 months or something. To the extent there were a lot of engineering problems and conceptual difficulties that were distinctive to this weird small thing, which in fact did play a big role in the overall machine and trajectory of civilization, it really was well leveraged. Progress in that area seems like it would have had unusually high dividends for faster overall technological progress. Maybe going along with that, I think it is also reasonable to think that if a small group had positioned themselves to understand that technology well and be pushing it and making investments in it, they probably could have ended up in a future situation where they've made a bunch of money or are in a position to understand well an important technology that not many people understand well as it gets rolled out. I think that's again a little bit different from the kind of thing he's expressing skepticism about, but seems like an important part of the calculus if one is thinking about trying to have leverage by working on AI, thinking about AI. I do think the alignment problem is distinctive from anything you could have said in the context of electricity. Maybe that's where I'm not mostly trying to do the make investments in AI so you're in a better position to have influence later or make a bunch of money. I'm mostly in the camp that we can identify a sort of unusually crisp issue which seems unusually important and can just tack away at that. I think that seems like it should have a lot of question marks around it, but I don't know of any historical cases where that would have applied.
类似地,有时人们会引用一些异端观点,我也试着研究过其中几个,但我不知道有哪些历史案例中,你提出了一个同样合理的论点,结果却感到非常失望。对于当前每增加一个人或一百万美元,哪些可能或现有的 AI 对齐工作能产生最大价值,你有什么想法吗?
Like a similar heresy, and sometimes people cite them, and I've tried to look into a few of them, but I don't know historical cases where you would have made a similarly reasonable argument and then ended up feeling really disappointed. Do you have any thoughts on what possible or existing AI alignment work might yield the most value for each additional person or million dollars that it receives at the moment?
是的,我们之前提到过——我提到过——有三类困难。我认为不同的资源在不同类别中会有用,每类资源最适合某些资源,比如某些人或某种制度意愿。所以简单回顾一下:第一类是概念性工作,思考如果我们想象哪些方法可能扩展到非常非常强大的 AI 系统,这一切将如何整合,以及在系统变得非常强大的极限下存在哪些困难。我对任何在该领域有合理能力的人尝试从事这项工作感到非常兴奋。这是我过去一年花了不少时间的事情,可能在我的整个职业生涯中占据了我很大一部分注意力,而且我开始考虑再次扩大规模。所以这就像是做理论工作,直接针对对齐问题,在纸上思考这个问题的可能方法,我们如何认为这将朝着一个我们感觉非常好的、真正确定的解决方案发展。所以这是第一类。我认为对于人们来说,在这上面投钱有点难,但对于喜欢做理论或概念工作的人来说,这是一个非常好的地方来增加这样的人。第二类是理解人类推理的事实,比如在辩论的背景下:人类能否成为仲裁不同观点、竞争观点的好裁判,以及如何设置辩论使得诚实策略在均衡中获胜?或者在放大方面,询问这个普遍性问题:是否可以将问题分解成至少稍微容易一点的问题?我也很兴奋让人们参与其中,所以进行更多实验,真正尝试练习这种奇怪的推理,看看人们能否做到,我们能否迭代并尝试识别困难案例。我对此非常兴奋,而且我认为这涉及到做这类工作的人有一些重叠,但可能涉及不同的人。还有第三类,涉及机器学习,实际上从理论转向实施,并达到我们有基础设施和专业知识等来实施我们认为最有希望的方法的地步。我认为这又需要不同类型的人,可能也需要不同类型的制度意愿和资金,对我来说这也是非常令人兴奋的。这可能两者兼有:它可以帮助为早期实验中产生的各种想法提供合理性检查,也可以更多地让我们处于未来能够做事情的位置,这样我们不一定确切知道需要什么样的对齐工作,但只要有机构、基础设施、专业知识和团队,他们有过深入思考这个问题的经验,实际构建机器学习系统并尝试实施,比如,这是我们目前的最佳猜测,让我们尝试使用这种对齐或将这些关于对齐的想法整合到最先进的系统中。仅仅拥有能够做到这一点的大量基础设施似乎非常有价值。无论如何,这就是我最兴奋投入资源进行对齐工作的三个类别。我基本上不认为在抽象层面上讨论哪个更有希望很难,因为会有很多比较优势的考虑,但我认为肯定有相当一部分人,我认为他们最好进入这三个方向中的任何一个。
Yeah, so we mentioned earlier—I mentioned earlier—there are three categories of difficulties. I think different resources will be useful in different categories, and each of them is going to be best for some resources, like some people or some kinds of institutional will. So briefly going over those again: one was conceptual work on how this is all going to fit together if we imagine what kinds of approaches potentially scale to very very powerful AI systems, and what are the difficulties in that limit as systems become very powerful. I'm pretty excited for anyone who has sort of reasonable aptitude in that area to try working on it. That's something I've been—it's been a reasonable fraction of my time over the last year, maybe over my entire career it's been a larger fraction of my attention, and something that I'm starting to think about scaling up again. Yeah, and so this is like thinking, doing theoretical work on directly at the alignment problem, asking on paper what are the possible approaches to this problem, how do we think that will play out moving towards having a really nailed down solution that we feel super great about. So that's one category. I think for people it's a little bit hard to down money on that, but I think for people who like doing theoretical or conceptual work, that's a really good place to add such people. There's a second category that's the sort of understanding facts about human reasoning, like understanding in the context of debate: can humans be good judges between arbitrating different perspectives, competing perspectives, and how would you set up a debate such that in fact the honest strategy wins in equilibrium? Or on the amplification side, asking about this universality question: is it the case you can sort of decompose questions into at least slightly easier questions? I'm also pretty excited about throwing people at that, so just running more experiments, trying to actually get practice engaging in this kind of weird reasoning, and really seeing can people do this, can we iterate and try and identify the hard cases. I'm pretty excited about that, and I think it involves some overlap in the people who do those kinds of work, but it maybe involves different people. And there's this third category of engaging with ML and actually sort of moving from the theory to implementation, and also getting in a place where we have infrastructure and expertise and so on to implement whatever we think is the most promising approach. I think that again requires a different kind of person still, and maybe also requires a different kind of institutional will and money, and is like also a pretty exciting thing to me. That's maybe both: it can help provide a sanity check for various ideas coming out of the earlier kinds of experiments, and it can also be a little bit more of this being in a position to do stuff in the future, so that we don't necessarily know exactly what kind of alignment work will be needed, but just having institutions and infrastructure and expertise and teams that have experience thinking hard about that question, actually building ML systems and trying to implement, say, here's our current best guess, let's try and use this alignment or integrate these kinds of ideas about alignment into state-of-the-art systems. Just having a bunch of infrastructure able to do that seems really valuable. Anyway, those are the three categories where I'm most excited about throwing resources on alignment work. I mostly don't think it's very hard to talk in the abstract about which one's more promising just because there's going to be lots of comparative advantage considerations, but I think there's definitely a reasonable chunk of people for which I think it's best to go into any of those three directions.
换个话题,谈点更异想天开的事:我发现一篇博客文章非常迷人。你最近提出,一个可能有效降低存在风险的方法是在地球上某处留下信息,让我们的后代在文明崩溃或人类灭绝后找到,然后生命重新出现,智慧生命在地球上重新出现,我们也许想告诉他们一些事情,帮助他们在我们失败的地方更成功。你想大致概述一下这个论点吗?
Changing gears to something a bit more whimsical: a blog post that I found really charming. You've argued recently that a potentially effective way to reduce existential risk would be to leave messages somewhere on Earth for our descendants to find in case civilization goes under or humans go extinct, and then life reappears, intelligent life reappears on the Earth, and we maybe want to tell them something to help them be more successful where we failed. You want to kind of outline the argument?
是的,这个想法是,如果人类——如果所有比蜥蜴大的动物都被杀死,那么你还有蜥蜴。蜥蜴还有很长的时间,直到光合作用开始崩溃,蜥蜴才会全部死亡。而且我认为根据我们对进化的理解,在可用的时间里,蜥蜴很可能能够再次建立起一个太空文明。这绝对不是确定的事情,而且这是一个非常难回答的问题,但我的猜测是,蜥蜴最终更有可能也能够去太空旅行。这是一个美丽的画面。所以,好吧,还有一个进一步的问题——这是一个让你觉得有点奇怪的地方。然后有一个问题,比如,你有多关心那个蜥蜴文明?也许与其他论点有关,比如关于你应该对其他价值体系有多好的奇怪决策理论论点。如果蜥蜴取代了我们,我会很高兴。你知道,我宁愿我们来做,但如果要么是蜥蜴要么什么都没有,我真的倾向于帮助蜥蜴。所以这可能有点离题,但在那种情况下,我的直觉是,是的,未来的蜥蜴人或现在的人类——我不确定哪个更好。就像人类是从潜在文明池中抽出来的;如果历史重演,用蜥蜴而不是人类,我们是否更好或更差并不明显。
Yeah, so the idea is if humanity—if every animal larger than a lizard was killed, then you still have the lizards. The lizards kind of have a long time left before the lizards would all die as photosynthesis started breaking down. And I think based on our understanding of evolution, it seems reasonably likely that in the available time, lizards would again be able to build up to a spacefaring civilization. Definitely not a sure thing, and it's a very hard kind of question to answer, but my guess would be more likely than not the lizards will eventually be in a position to also go travel to space. It's a beautiful image. And so okay, then there's a further—that's one place where you're like, that's a sort of weird thing. Then there's a question like, how much do you care about that lizard civilization? And maybe related to these other arguments, like related to weird decision theory arguments about how nice should you be to other value systems. I'm inclined to be pretty happy if the lizards take our place. You know, I prefer we do it, but if it's going to be the lizards or nothing, I would really be inclined to help the lizards out. So maybe this is too much of an aside here, but I kind of in that case have the intuition that yeah, future lizard people or humans now—it's like I'm not sure which is better. It's like humans were drawn out of the pool of potential civilizations; it's not obvious whether we're better or worse than if you reran history with lizards rather than people.
我只是想插一句,因为我的一些同事指出,显然有一些疯狂的阴谋论关于所谓的蜥蜴人秘密统治世界,我之前没听说过,为了避免任何可能的混淆,我们在这里讨论的与任何这样的蜥蜴人无关。
I just wanted to jump in because some of my colleagues pointed out that apparently there's some insane conspiracy theory out there about so-called lizard people secretly running the world, which I had not heard of, and to avoid any conceivable possible confusion, what we're talking about here has nothing to do with any such lizard people.
蜥蜴人只是我们开玩笑的说法,指代未来某天可能在地球上重新进化出的智慧生命,比如几百万、几千万甚至几亿年后。如果人类在某个时候灭绝,让自然世界自行运转,也许蜥蜴人这个说法事后看来有点不妥,但不管怎样,我们继续节目吧。
Lizard people is just our jokey term for whatever intelligent life might one day re-evolve on Earth, you know, many millions, tens of millions, or hundreds of millions into the future. Should humans at some point die out and leave the natural world to just clock along on its own? Perhaps lizard people was a slightly unfortunate turn of phrase in retrospect, but regardless of that, let's get on with the show.
是的,我认为这是一个有趣的问题。对我来说,这是最重要的开放哲学问题之一,即一般来说,你应该如何看待其他价值体系,你对被取代有多乐意?所以我认为蜥蜴人想要的东西和我们非常不同,在客观层面上,他们创造的世界可能与我们创造的世界大相径庭。我有点认同这种直觉:我很乐意让蜥蜴人来。如果我在考虑是冒灭绝风险还是让蜥蜴人接管,我更倾向于让蜥蜴人接管,而不是冒巨大的灭绝风险。所以如果能做些什么让蜥蜴人的生活更轻松,我会很高兴,也很乐意去做。
Yeah, I think it's an interesting question. I think this is to me like one of the most important open philosophical questions, which is just in general, what kinds of other value systems should you be like, how happy with replacing you? So I think the lizards would want very different things from us, and on the object level, the world they created might be quite different from the world we would have created. I sort of share this intuition: I'm pretty happy for the lizards. I feel pretty great, you know, if I'm considering should we run a risk of extinction or let the lizards take over, I'm more inclined to let the lizards take over than run a significant risk of extinction. So I would be happy if there's anything we could do to make life easier for the lizards. I'm pretty excited about doing it.
我很高兴我们把这个蜥蜴人的想法具体化了。继续吧。我说蜥蜴也是部分因为如果你选比蜥蜴小得多的生物,到某个点就会变得不太靠谱。如果只有植物,那就更悬了。它们有足够的时间吗?蜥蜴,我觉得还算安全。蜥蜴相当大,相当聪明,基本能发展出太空能力。那么问题来了,我们实际上能做什么?这为什么相关?下一个问题是,有没有一种现实的方式,让我们杀死自己和所有大型动物,而不完全毁灭地球上的生命,或者不让我们自己被追求非常不同价值观的 AI 取代?
I'm glad we've made this so concrete with the lizard people. Get carry on. I say lizard is also in part because if you go too much smaller than lizards, at some point it becomes more dicey. If you only had plants, it's a little bit more dicey. Would they have enough time left? Lizards, I think, are kind of safe. Lizard's pretty big, pretty smart, most of the way to space bearing. So then there's a question of what could we actually do? Why is this relevant? And the next question is, is there a realistic way that we could kill ourselves and all the big animals without just totally wiping out life on Earth, or without replacing ourselves with, say, AI pursuing very different values?
我认为我们无法实现自身价值观的最可能方式不是灭绝,而是我们只是做错了事,指向了错误的方向。我认为这比灭绝的可能性大得多。我的粗略理解是,如果我们现在灭绝,我们很可能会带走地球生态系统的大部分。但如果你认为气候变化真的能杀死所有人类,那你可能会更兴奋。所以有一些可能的方式可以杀死所有人类,但并非真的杀死一切。比如文明彻底残酷崩溃,或者某种生物恐怖主义杀死所有大型动物或杀死所有人类,但不一定杀死一切。如果这些是可能的,那么你有可能最终陷入蜥蜴人的局面,然后由蜥蜴人来殖民太空。在这种情况下,我们似乎有一个非常有趣的杠杆:蜥蜴人将在几亿年后进化,他们将在几亿年后处于我们的位置。留下信息似乎是现实的,以某种方式改变地球,使得几亿年后出现的文明能够注意到我们做出的改变并开始研究它们。到那时,如果我们能引起某个未来文明对特定事物的注意,我认为我们可以为他们编码大量信息,并决定如何使用这个通信渠道。所以有时人们谈论这个,他们通常想象的是极短的时间段和几亿年,而且通常不会深思熟虑他们想说什么。但我的猜测是,通过从一个更复杂的视角发送信息,你可以真正实质性地改变一个文明的轨迹。如果你想象人类第一次发现上一个文明发送的信息,那可能至少是 100 年前。那时,信息可能来自一个技术更先进的文明,并且经历了文明从兴起到灭绝的完整弧线。所以至少,你可以通过有选择地为他们阐明或展示如何发展、如何实现某些目标,来真正改变他们的技术发展路径。你也可以尝试所有这些事情,更推测性地,帮助他们走上更好的道路,比如,你真的应该担心杀死所有人,这里有一些关于如何建立制度以避免灭绝的指导。我非常关注 AI 对齐,所以我会非常感兴趣尽可能多地传递这样的信息:这是我们在深思熟虑后认为是个问题的事情,你现在可能还没想到,但请注意。我确实认为,这会让那个未来文明中研究该问题的人群处于一个质上不同的位置,而不是像,我不知道,如果我们偶然发现来自过去文明的这些非常详细的信息,很难说会产生什么影响。但我确实认为这会对发展轨迹产生巨大的技术影响,并且很可能对关于如何组织自身的深思熟虑和决策,或其他智力项目产生合理的影响。
I think by far the most likely way we're going to fail to realize our values is we don't go extinct, but we just sort of are doing the wrong thing and pointing in the wrong direction. I think that's much more likely than going extinct. And my rough understanding is that if we go extinct at this point, we will probably take most of the Earth's ecosystem with us. But I think if you thought that climate change could literally kill all humans, then you'd be more excited. So there's some plausible ways that you could kill all humans but not literally kill everything. A total really brutal collapse of civilization, maybe there's some kinds of bioterrorism that kill all large animals or kill all humans but don't necessarily kill everything. If those are plausible, then there's some chance you end up in the situation with the lizards, and now it's up to the lizards to colonize space. In that case, it does seem like we have this really interesting lever where the lizards will be evolving some hundreds of millions of years, they'll be in our position some hundreds of millions of years from now. It does seem probably realistic to leave messages, to somehow change Earth such that a civilization that appeared several hundred million years later could actually notice the changes we've made and could start investigating them. And at that point, we would probably have, if we're able to call the attention of some future civilization to a particular thing, I think then we can encode lots of information for them, and we could decide how we want to use that communication channel. So sometimes people talk about this, they normally are imagining radically shorter time periods and hundreds of millions of years, and they're normally not being super thoughtful about what they'd want to say. But I think my guess would be that there are ways you could really substantially change the trajectory of a civilization by being able to send a message from a much more sophisticated standpoint. If you imagine the first time that humans could have discovered a message sent by a previous civilization, it would have been, I mean depends a little bit on how you're able to work this out, but probably at least 100 years ago. At that point, the message might have been sent from a civilization which is much more technologically sophisticated than they are, and also which has experienced the entire arc of civilization followed by extinction. So at a minimum, it seems like you could really change the path of their technological development by selectively trying to spell out for them or show them how to develop, how to achieve certain goals. You could also attempt all those things, a little bit more speculative, to help set them on a better course, be like, you know, you really should be concerned about killing everyone, here's some guidance on how to set up institutions so they don't kill everyone. I'm very concerned about AI alignment, so I'd be very interested in as much as possible being like, here's the thing which upon deliberation we thought was kind of a problem, you probably aren't thinking about it now, but FYI be aware. And I do think that would put a community of people working on that problem in that future civilization in a qualitatively different place than if it's just sort of, I don't know, it's very hard to figure out what the impact would be had we stumbled across these very detailed messages from a past civilization. But I do think it could have a huge technological effect on the trajectory of development, and also reasonably likely have a reasonable effect either on deliberation and decisions about how to organize ourselves, or on other intellectual projects.
是的,你提出了这个假设:如果我们能随意向 1600 或 1700 年的人们发送文本,我们能否让历史变得更好?然后反思一下,似乎是的,我们可以向他们发送大量非常重要的哲学和社会科学的重要发现,并告诉他们我们珍视而他们可能不珍视的东西,从而加速我们认为特别重要的哲学思想流派的发展。你也可以选择技术,从我们世界存在的所有技术中挑选,然后说这是我们平衡后认为好的。比如不给他们核武器的配方,而是给他们关于相互确保摧毁的博弈论,或者告诉他们我们所知道的一切关于如何...。
Yeah, you give this hypothetical of could we have made history go better if we could just send as much text as we wanted back to people in 1600 or 1700, and then kind of a reflection does seem like well yeah we could just send them lots of really important philosophy and lots of important discoveries in social science and kind of tell them also the things that we value that maybe they don't value, and so speed up the strains of philosophical thought that we think are particularly important. You can also just choose what technologies, pick and choose from all the technologies that exist in our world and be like here's the ones we think are good on balance. So just like you don't give them the recipe for nuclear weapons, instead you give them the game theory for mutually assured destruction so they can, or you tell them everything that we knew about how to...
维持国际合作,这样当他们发展核武器时,就能更好地避免自我毁灭。比如教他们如何建造一个很棒的风车,或者太阳能电池板。我的意思是,为什么不呢?我们直接给他们太阳能电池板之类的东西。我不知道这种干预能有多大好处。这是一个值得深入思考的问题。我的猜测是,有些事情预期效果还不错,但很难说。
Sustain international cooperation so whenever they do develop nuclear weapons they're in a better position to not destroy themselves. Here's how you build a really great windmill, yeah, here's solar panels. I mean, why not? We just give the solar panels stuff. I don't know how much good you could do by that kind of intervention. It's a thing that would be interesting to think about a lot more. My guess would be that there's some stuff where an expectation is reasonably good, but it's hard to know.
是的。
Yeah.
好吧,所以似乎有一个相当合理的假设:如果人类灭绝,智能生命可能会重新出现,而且如果我们思考得足够久,我们可能会想出一些有用的东西告诉他们,这可能会帮助他们,让他们有更好的机会生存、繁荣,并做我们看重的事情。但是,你怎么才能留下一个能持续数亿年的信息呢?这似乎相当有挑战性。
Okay, so it seems there's a pretty plausible case that if humans went extinct, intelligent life might reemerge, and probably if we thought about it long enough, we could figure out some useful thing that we could tell them that would probably help them and give them a better shot at surviving and thriving and doing things that we valued. How on Earth would you leave a message that could last hundreds of millions of years, though? It seems like it could be pretty challenging.
是的,我认为这个问题有两部分。一部分是引起某人对某个地方的注意。我认为这是迄今为止更难的部分。例如,如果你要埋藏某样东西,地球上的大多数地方都不行,因为数亿年的时间足以让地球表面不再是原来的表面。所以我认为第一个也是更重要的问题是引起某人对某个地点的注意,或者对一百万个地点中的某一个的注意。然后第二部分是,在引起某人对某个地点的注意后,你如何实际编码信息,如何实际与他们沟通?另外值得一提的是,这是我写的一篇博客文章。我估计有人对这些问题有比我更深的理解,可能对许多具体问题思考得比我更深入,我不想说得好像我是给未来文明留信息的权威。我只是思考了几个小时。
Yeah, I think there's two parts to the problem. One part is calling someone's attention to a place. I think that's the harder part by far. For example, if you were to bury a thing, most places on Earth won't work because hundreds of millions of years is long enough that the surface of the Earth is no longer the surface of the Earth. So I think the first and more important problem is calling someone's attention to a spot, or to one of a million spots or whatever. And then the second part of the problem is, having called someone's attention to a spot, how do you actually encode information, how do you actually communicate to them? It's also probably worth saying that this is a blog post I wrote. I expect there are people who have much deeper understandings of these problems than I do, probably thought about many of these exact problems in more depth than I have, and I don't want to speak as if I'm an authority on leaving messages for future civilizations. I thought about it for like some hours.
是的。
Yeah.
所以在引起注意方面,我在博客文章中想到了一系列可能性。我对此很感兴趣,并在网上与人开始讨论,集思广益。我认为如果我们稍微思考一下,我们可能就能得出一个清晰的认识。目前可能最领先的提议是,我认为 Yan Kite 提出了一个提议:俄罗斯有一个特别大的磁异常区,文明很容易在早期发现它,而且它的位置使得它不太可能随着板块运动而移动。似乎相当合理,尽管做起来有点困难,你可以利用对该结构的修改,或者在该结构中定位和标记点,使得至少我们的文明会非常可靠地发现它。很难说一个与我们截然不同的文明会发现多少。
So in terms of calling attention, I thought of a bunch of possibilities in the blog post. I was kind of interested and started some discussions online with people brainstorming possibilities. I think if we thought about it a little bit, we could probably end up with a clear sense. Probably the leading proposal so far is, I think Yan Kite had this proposal of a particularly large magnetic anomaly in Russia, which is very easy for a civilization to discover quite early, and which is located such that it's unlikely to move as tectonic plates move. It seems pretty plausible, though a little bit difficult to do, that you could use modifications of that structure, or locating things and shelling points in that structure, in a way that at least our civilization would very robustly have found. It's hard to know how much a civilization quite different from ours would have.
你还有一个直截了当的想法:一块巨大而坚硬的岩石突出地面,希望我们能活到足够长的时间来实现它。这是一个非常直接的想法。
You also had the straightforward idea of a really big and hard rock that juts out of the earth, and hopefully we survive long enough to do that. That's a very straightforward idea.
是的,要让这样的东西起作用,真的出奇地难。
Yeah, it's really surprisingly hard to make things like that work.
我想,在那么长的时间里,即使是非常耐用的岩石也会被侵蚀破坏。而且,地壳运动如此剧烈。你把岩石放在地球表面,数亿年后它就不再在地球表面了;它会被埋起来。
I guess it's like over that period of time, even a very durable rock is kind of going to be broken down by erosion. Also, stuff moves so much. You put the rock on the surface of the Earth, it's not going to be on the surface of the Earth in hundreds of millions of years anymore; it just gets buried somehow.
是的,有趣。所以它出奇地困难。当我开始写这个的时候,我真的很大程度上更新了看法,认为它很困难。我当时想,‘我确定这很容易’,然后我想,‘哦天哪,真的,基本上什么都不行。’
Yeah, interesting. So it's surprisingly rough. I really updated a lot towards it being rough when I started writing this. I was like, 'I'm sure this is easy,' and I was like, 'Oh jeez, really, basically everything doesn't work.'
那用一堆放射性废物呢?它们可以用盖革计数器检测到。
What about a bunch of radioactive waste that would be detectable by Geiger counters?
是的,你可以尝试这样做。你必须关心这些东西能持续多久,它们有多容易被检测到,以及它们离地表多远仍可检测。但我认为有类似这样的可行方案。我认为磁铁保持磁性的时间比我最初想象的要长,或者比你猜想的要持久,而且是一种可以被检测到的合理选择。
Yeah, so you can try and do things like that. You have to care about how long these things can last and how easy they are to detect, and how far from the surface they remain detectable. But I think there are options like that that work. I think also magnets remain magnetic longer than I initially thought, or longer lasting than you might have guessed, and are a reasonable bet for a thing that can be detected.
你指出,你可以有成千上万个这样的地点,并确保每个地点都有一张所有其他地点位置的地图,这样他们只需要找到一个,然后就可以去挖掘每一个,这肯定提高了几率。
You made the point that you can potentially have thousands of these sites, and you can make sure that in every one there's a map of where all the others are, so they only have to find one and then they can just go out and dig up every single one of them, which definitely improves the odds.
是的,还有一些化石存在。所以如果你认为你有一百万个非常容易变成化石的东西,那可能行不通。我有一阵子没想过这个了。我认为如果你坐下来,如果你找一个人花些时间真正充实这些提议,调试它们,并咨询专家,他们可能能找到一些可行的方法。同样,在社会方面,如果你思考很长时间,你可能会对是否有值得说的有价值的东西有一个更深思熟虑的看法。第一步是:你想付钱给某人让他们花大量时间思考这些事情吗?有没有人非常兴奋地花大量时间思考这些事情,敲定这些提议,然后看看它们看起来怎么样?如果它们看起来是个好主意,就花几百万或几千万美元去实际实现它。
Yeah, also there are some fossils around. So if you think you have a million very prone to be fossilized things, it's probably not going to work. I haven't thought about that in a while. I think probably if you sat down, if you just took a person and they spent some time really fleshing out these proposals and debugging them and consulting with experts, they could probably find something that would work. And then similarly on the social side, if you thought about a really long time, you could find a more considered view about whether there's something to say that would be valuable. The first step would be: do you want to pay someone to spend a bunch of time thinking about those things? Is there someone who's really excited to spend a bunch of time thinking about those things, nailing down the proposals, and then seeing what they look like? If they look like a good idea, spending millions or tens of millions of dollars to actually make it happen.
那么,关于如何编码这些信息,你似乎认为可能只需在岩石上雕刻就是一个合理的初步方案,在大多数情况下可能足够好。然后你可能会想出一些更好的材料来雕刻东西,这些材料很可能持续很长时间,至少如果埋藏得当的话。我认为其他人对问题的这一方面思考得更多,而且总的来说我们更有信心某些方法会奏效。但我认为在合理条件下,仅仅雕刻东西就已经足够好了。有一个能存活数亿年的小东西,比实际改变地球面貌使其在数亿年后仍然引人注目要容易得多。
So in terms of how you would encode this information, it seemed like you thought probably just etching it in rock would be a plausible first pass that would probably be good enough most of the time. Then you can probably come up with some better material on which you could etch things that is very likely to last a very long time, at least if it's buried properly. I think other people have thought more about this aspect of the problem, and I think in general we have more confidence something will work out. But I think just etching stuff is already good enough under reasonable conditions. It's a lot easier to have a small thing that will survive for hundreds of millions of years than to actually disfigure the Earth in a way that will be notable and would still call someone's attention to it in hundreds of millions of years.
好吧,这引出了我的主要反对意见,那就是蜥蜴人可能不会说英语,所以即使我们埋了一个巨大的……
Okay, so this brings me to the main objection I had, which is that the lizard people probably don't speak English, and so even if we bury a huge...
我们把维基百科埋起来。我觉得他们可能会觉得非常困惑。我们怎么确保能把任何概念传达给 100 万年后的人或听众呢?
We bury Wikipedia. I think they might just find it very confusing. How are we clear that we can communicate any concepts to people or to listeners in 100 million years time?
是的,我认为这是一个非常有趣的问题。它涉及到你想要思考的事情。但我确实认为,当人类历史上试图破译一种失传的语言或文物时,他们所处的境况比蜥蜴人面对这个人工制品要糟糕得多,因为我们会拥有大量信息并试图被理解。我认为我们并没有人类遇到这种试图被理解的超级信息的例子。这就像一场可以在人类之间尝试的游戏,我认为人类可以非常轻松地获胜,但尚不清楚这在多大程度上是因为我们拥有共同的背景。特别是,我认为人类不需要任何类似语言的东西就能轻松赢得这场游戏,仅通过简单的插图和图表就能建立起概念的语言。我认为你有理由持怀疑态度,即使不是语言,我们也在使用所有这些共同的概念,我们以相同的方式思考问题,我们知道我们的目标是什么。我相当乐观,但这还很不清楚。这也是人们思考了很多的事情,尽管在这种情况下,我对他们的想法远不如对把东西写得很小且耐久的做法那么有信心。
Yeah, I think that's a pretty interesting question. It goes into things you want to think about. But I do think when people have historically engaged in the project of trying to figure out a lost language or relics, they're in a radically worse position than the lizard people would be with respect to this artifact, since we would have a lot of information and be attempting to be understood. I think we don't really have examples of humans encountering this kind of super information that's attempting to be understood. It's like a game you can try among humans, and I think humans can win very easily at it, but it's unclear the extent to which that's because we have all this common context. In particular, I think humans do not need anything remotely resembling language to easily win this game, to build up a language of concepts just by simple illustrations and diagrams. I think you would be right to be skeptical, even when it's not language, we are using all these concepts in common, we've thought about things in the same way, we know what we're aiming at. I'm reasonably optimistic, but it's pretty unclear. This is also something people have thought about a lot, although in this case I'm a lot less convinced in their thinking than in writing stuff really small and in a durable way.
我的理解是,那些深入思考过这个问题的人似乎对我们的信息传递能力非常悲观。老实说,我知道的唯一案例是一个项目,试图弄清楚我们应该在掩埋极其可怕的核废料的地点放置什么信息。你把这种剧毒的东西埋在地下,不希望未来的人挖出来害死自己。有很多人,语言学家、社会学家,试图弄清楚该放什么信号。他们最终确定了一个信息,我认为这简直糟透了,因为我无法想象任何未来文明会把它解读为除了宗教之外的东西,他们会非常好奇,然后绝对会去挖出来。我会找到他们决定传达的确切信息,并在这里读出来,这样大家就可以自己判断了。
My understanding was that the people who thought about this a lot seemed very pessimistic about our ability to send messages well. I guess to be honest, the only case I know about is a project to try to figure out what message we should put at the site where we bury really horrible nuclear waste. You're putting this incredibly toxic thing underground, and you don't want people in the future to dig it up and kill themselves. There were quite a lot of people, linguists, sociologists, trying to figure out what signals to put there. They settled on some message that I thought was absolutely insanely bad because I couldn't see how any future civilization would interpret it as anything other than religious stuff that they would be incredibly curious about and would absolutely go and dig it up. I'll find the exact message they decided to communicate and read it out here so people can judge for themselves.
嘿,各位,我确实查到了这条信息,把它加在这里,这样你们也可以和我一起评判。信息是这样的:'这个地方是一条信息,也是一个信息体系的一部分。请注意它。发送这条信息对我们很重要。我们认为自己是一个强大的文化。这个地方不是荣誉之地。这里没有纪念任何崇高的事迹。这里没有有价值的东西。这里的东西对我们来说是危险和令人厌恶的。这条信息是关于危险的警告。危险位于特定位置。它向中心递增。危险的中心在这里,具有特定的大小和形状,并且在我们下方。危险在你们的时代仍然存在,就像在我们的时代一样。危险针对身体,可以致命。危险的形式是能量的散发。只有当你从物理上严重扰乱这个地方时,能量才会释放。这个地方最好避开并无人居住。' 正如我所说,我真的认为未来的文明,无论是人类还是其他,会对任何附有这种信息的东西极度好奇,并且很可能猜测其本质是宗教性的。如果他们自己还没有了解核辐射,这可能会让他们意识到这是一个核废料场,我认为他们比完全不做标记更有可能去挖掘那个地点。所以我真的不知道研究人员在想什么,竟然认为写这样一条信息是个好主意。但无论如何,这是一个迷人的研究项目。好了,回到对话。
Hey folks, I did look up this message to add it in here so that you can pass judgment on it as well as me. Here it is: 'This place is a message and part of a system of messages. Pay attention to it. Sending this message was important to us. We considered ourselves to be a powerful culture. This place is not a place of honor. No highly esteemed deed is commemorated here. Nothing valued is here. What is here was dangerous and repulsive to us. This message is a warning about danger. The danger is in a particular location. It increases towards the center. The center of danger is here, of a particular size and shape, and below us. The danger is still present in your time as it was in ours. The danger is to the body, and it can kill. The form of the danger is an emanation of energy. The energy is unleashed only if you substantially disturb this place physically. This place is best shunned and left uninhabited.' As I said, I really think a future civilization, human or otherwise, would be insanely curious about anything attached to a message like that and would probably guess it was religious in nature. If they hadn't learned about nuclear radiation themselves already, which would probably allow them to figure out it was a nuclear waste dump, I think that would be much more likely to dig at that spot than if it was simply left unmarked entirely. So I really don't know what the researchers were thinking, expecting that writing a message like that would be a good idea. But this was a fascinating research project anyway. All right, back to the conversation.
但不管怎样,我的意思是他们确实有这个……哦,实际上当时的计划是用现存的各种语言来写,希望其中一种能幸存下来。那是选项之一。但在这里这不是一个选项。不是选项。所以是的,但我认为这是一个相当不同的情况。你想做一个标志,让遇到标志的人能明白他们在说什么,和你想写一亿字的信息,这是不同的。如果我们遇到来自某个文明的信息,并且能看出其技术力量远超我们,我们会想,好吧,这在我们优先事项中排名很高,他们到底在说什么。这是一个非常不同的情况,他们写了大量的内容,这就像有史以来最有趣的学术项目。发现这样的东西会立即成为智力优先队列的顶端。我更有信心我们能够弄清楚,或者像我们这样的文明在这种情况下能够弄清楚,相比之下,有人四处走动,遇到一个标志,也许还比较原始,完全不知道是怎么回事。而且,内容也不多。在只给他们一万字内容或一些图片的情况下,不清楚他们如何能有足够的线索来弄清楚是怎么回事。而在这个案例中,我们不仅仅有一个关于如何建立共享概念语言的提案;我们有一百个提案,我们尝试所有提案。每个提案,哪怕是一个四年级学生想出来的,也没问题,也扔进去。比特很便宜,所以你可以尝试很多东西。我们的处境比人们通常认为的要好得多。
But anyway, I mean they did have this... Oh, I think actually the plan there was to write it in tons of languages that exist today in the hope that one of those would have survived. That was one of the options. That's not going to be an option here. Not an option here. So yeah, but I think it's quite a different situation. It's different if you want to make a sign so someone who encounters the sign can tell what they're saying, versus if I want to write someone 100 million words. If we encounter a message from some civilization that we can tell has technological power much beyond our own, we're like, okay, that's really high up on our list of priorities, what the hell they're talking about. That's a very different situation where they've written this huge amount of content, and it's like the most interesting academic project of all time. It goes to the top of the intellectual priority queue upon discovering such a thing. I have a lot more confidence in our ability to figure something out, or a civilization like ours' ability to figure something out under those conditions, compared to someone walking around, encountering a sign, perhaps somewhat primitive, and having no idea what's up with it. Also, it's just not that much content. It's unclear how, in the case where you only give them 10,000 words of content or some pictures, they just don't have enough traction to possibly figure out what's up. Whereas in this case, we're not just having one proposal for how to build a shared conceptual language; we have 100 proposals, we're trying them all. Every proposal, any fourth grader came up with, that's fine, throw it in there too. Bits are quite cheap, so you can really try a lot of things. We're in a much better position than people normally think about.
我认为考古学家,当他们挖出文字时,有时会通过类比我们有记录的其他语言来破译。有时他们有罗塞塔石碑,上面有翻译,这样我们就可以……
I think archaeologists, when they've dug up writing, sometimes they've decoded it by analogy to other languages that we do have records about. Sometimes they have the Rosetta Stone, where we have a translation, so then we can...
你能弄清楚他们有什么吗?我认为他们有两者的翻译,还有第三种语言是一样的,然后他们就能从那里弄清楚那种语言听起来像什么,然后逐渐弄清楚单词的意思。我认为还有其他情况,仅凭上下文,他们挖出了石头,然后这是什么?结果发现是一家公司的一堆财务账目,或者他们在弄清楚这个地方的进出口情况,这完全合理。你可以想象他们会这么做。而你的希望是,我们会埋下如此多的内容,我们会有很多图片,很多重复的单词,最终他们能够解码。他们会从某种上下文中弄清楚。我想他们会翻阅百科全书,然后找到一篇关于某样东西的文章,他们能弄清楚那是什么,因为他们也有这个东西。他们想,树,好的,我们有关于树的文章,我们还有树,然后他们就想,如果我要写一篇关于树的百科全书文章,我会说什么?他们猜测那些单词是什么,然后从那里开始。我们可以让事情比百科全书文章简单得多,比如,这里有一百万个概念的词汇表,对于每个概念,无论是一万个概念,每个概念有 100 张图片、100 个句子和 100 次定义尝试,试图组织起来。
Can you figure out what they had? I think they had like a translation of two of them, and there was a third language that was the same thing, and then they could figure out what the language sounded like from that, and then figure out very gradually what the words meant. I think there are other cases where just from context, they've dug up stones and then what is this? And it turns out that there are a bunch of financial accounts for a company, or they're figuring out imports and exports from this place, which makes total sense. You can imagine they'd be doing that. And your hope here is that we will just bury so much content, and we'll have a bunch of pictures, lots of words repeating, that eventually they'll be able to decode it. They'll figure out from some sort of context. I guess they'll be flicking through the encyclopedia and then they find one article about a thing that they can figure out what it is because they also have this thing. They're like, trees, okay, we got the article about trees, and we still have trees, and then they kind of work out, well, what would I say about trees if I was writing an encyclopedia article about trees? And they kind of guess what those words are, and then they kind of go out from there. We can make things a lot simpler than encyclopedia articles where you can be like, here's a lexicon of a million concepts, and for each of them, whatever, 10,000 concepts, for each of them, 100 pictures and 100 sentences about them and 100 attempts to define them, attempted to organize.
是的,好吧,我同意。我认为如果你达到那个水平,那么你可能能做到,尽管有些概念可能极难说明。我对交流技术更乐观,这似乎比画一张蒸汽机的图片更容易。也许哲学有点棘手,或者宗教。
Yeah, okay, I agree. I think if you went to that level, then probably you could do it, although some concepts might be extremely hard to illustrate. I'm more optimistic about communicating technology seems easier than, here's a picture of a steam engine. Maybe philosophy is a bit trickier, or religion.
所以在博客文章中,你建议这可能是在降低存在风险方面性价比很高的,我想你有一个 1000 万美元的预算,用于这个的最小可行产品,你在想,是的,如果我们非常小心地选择发送什么信息和不发送什么信息,这可以将他们的生存几率提高一个百分点。你还这么认为吗?我想 1000 万美元的预算对我来说似乎低得令人难以置信。我想我们在这里设想的可能比你当时想的要雄心勃勃得多。
So in the blog post, you suggested that this might be pretty good bang for buck in terms of reducing existential risk, and I think you had like a budget of $10 million for a kind of minimum viable product of this, and you were thinking, yeah, this could improve their odds of surviving by one percentage point if we're very careful about what messages we send them and what messages we don't send them. Do you still think something like that? I guess the budget of $10 million seemed incredibly low to me. I guess here we've been kind of envisaging something potentially a lot more ambitious than what you were thinking about at the time.
是的,所以我认为在与人们讨论实际的存储选项、如何制作信息、如何让人们找到信息之后,1000 万美元确实看起来很低。1000 万美元似乎很低,而 1 亿美元可能更现实,这使成本效益数字更差。我认为值得指出的是,你必须分别处理这个问题。如果你想象项目的三个阶段或四个阶段:弄清楚要说什么,以某种方式制作一个人们可以识别的地标,实际编码一堆信息,然后实际写作或试图传达你想说的信息。如果其中一项很昂贵,你可以相对容易地将其他项提高到相同的成本。所以我们要在每个阶段花费数百万美元。我想实际上,我预计大部分成本会花在留下地标上,但这仍然给你留下了数百万美元用于其他部分,这相当于几个人全职工作多年。
Yeah, so I think $10 million does seem low after talking with people about what the actual storage options are, how to make a message, how to make it so people could find a message. $10 million seems low, and $100 million seems probably more realistic, which makes the cost-effectiveness numbers worse. I think it is worth pointing out that you have to go separately on that. If you imagine three phases or four phases of the project: figuring out what to say, somehow making a landmark people can identify, actually encoding a bunch of information, and then actually writing or trying to communicate the information you wanted to say. If one of those is expensive, you can sort of relatively easily bring the others up to the same cost. So we're getting to spend like millions of dollars on each of those phases. I think actually probably I imagine the lion's share of the cost going into leaving a landmark, but that still leaves you with millions of dollars to spend on other components, which is like a few people working full-time for years.
我本以为最难的事情是弄清楚要说什么,然后弄清楚如何传达它,因为如果我们真的在谈论为我们认为蜥蜴人能够理解的每个单词画图,那似乎是一项很大的功课。
I would have thought the most difficult thing would be to figure out what to say and then figure out how to communicate it, because if we're really talking about drawing pictures for every word that we think lizard people will be able to understand, that seems like a lot of homework.
是的,我认为很难估算这种工作的量。我们是在说 100 人年还是 1000 人年?那是多少人年的努力?你可以想想合理的百科全书需要多少人年。考虑成本很棘手。
Yeah, I think it's hard to ballpark the amount of that kind of work. Are we talking like 100 person-years or a thousand person-years? How many person-years of effort is that? You can think about how many person-years go into reasonable encyclopedias. It's tricky thinking about the costs.
是的,我认为在 1 亿美元的水平上,我对彻底性感觉良好……我的意思是,再次,你不可能对发送什么有一个很好的答案。你会有一个经过人们几年思考支持的答案。我想可能如果你在做这个项目,你是在某种世界观下做的。这个项目已经基于一堆关于世界的疯狂观点,所以就像你是在对这些疯狂观点下全盘赌注,当你在做其他阶段时,你也在某种程度上以这些疯狂观点正确为前提,即什么基本的东西重要以及事物如何基本运作,我认为这在某种意义上有所帮助。或者你只需要为这些疯狂观点的正确性付出一次代价;你不必再付一次。
Yeah, I think at $100 million, I feel good about how thoroughly... I mean, again, you're not going to be able to have a great answer to what to send. You're going to have an answer supported by people thinking for a few years. I guess probably if you're doing this project, you're kind of doing it under a certain set of worldviews. This project is already predicated on a bunch of crazy views about the world, and so just like you're making an all-out bet on those crazy views about the world, and when you're doing these other stages, you're also sort of just conditioning on those crazy views about the world being correct about what basic kind of thing is important and how things basically work, which I think does in some sense help. Or you only have to eat those factors for those crazy views being right once; you don't have to pay them again.
是的,我想我一直以为,制作一个能被未来文明理解的东西只需要不到几个人年的努力。也许我对此过于乐观了,而且我没有与任何详细思考过这个问题的社区接触,完全有可能我大错特错。但无论如何,当我想象人们花 10 年时间在这上面时,我想,10 年,那看起来相当不错。看起来他们会把这件事搞定。他们会测试很多次,会有六个独立实施的提案,每个都会非常详尽,有很多漂亮的图片。漂亮的图片其实有点难,但大概他们会得到这些碎片,然后他们用这些碎片做什么?
Yeah, I guess I just always imagined that it would take less than a few person-years of effort to produce something that could be understood by a future civilization. Maybe I'm just way too optimistic about that, and I haven't engaged with any of the communities that have thought about this problem in detail, and it's totally possible that I'm way off base. But anyway, when I imagine people spending 10 years on that, I'm like, 10 years, that seems pretty good. Seems like they're going to have the thing nailed. They're going to have tested it a bunch of times, they're going to have six independent proposals that are implemented separately, each of them is going to be super exhausted with lots of nice pictures. Nice pictures are actually a little bit hard, but like, so they probably get these bits, and what do they do with all the bits?
那么听众们也许应该资助这个想法吗?有没有人表示过有兴趣担任这个项目的负责人?
So should listeners maybe fund this idea? Has anyone expressed interest in being the team lead on this?
是的,有一些对话,关于地标步骤的非常简短的对话。我想那可能是我首先好奇的事情:成本是多少?我不认为这是一个需要资助的大项目。我认为没有人真正表示过有兴趣接手并推进它。我想顺序可能是先检查地标是否合理,以及项目大致需要多少费用,然后考虑对所有细节进行合理性检查。
Yeah, there have been some conversations, very brief conversations about the landmarking step. I think that's probably the first thing I would be curious about: what is the cost? I don't think it's a big project to be funded yet. I don't think anyone's really expressed interest in taking it up and running with it. I think the sequence would probably first check to see if the landmark thing makes sense and roughly how expensive a project it would necessarily be, and then think about maybe doing a sanity check on all the details.
你有没有其他被忽视或听起来有点疯狂的想法,可能比传统的降低存在风险的方法更有优势?
Do you have any other neglected or kind of crazy sounding ideas that might potentially compare favorably to more traditional options for reducing existential risk?
我确实认为需要说明,如果有任何方法可以尝试解决 AI 风险,那可能比这类事情更好,因为我的竞争优势似乎在于 AI 风险领域。至于奇怪的利他计划,我觉得过去一年我对此想得不多。我没有什么既非常奇怪又非常有吸引力的想法。
I do think it's worth caveating that if there's any way to try and address AI risk, that's probably going to be better than this kind of thing related to my competitive advantage seeming to be in AI risk stuff. In terms of weird altruistic schemes, I feel like I haven't thought that much about this kind of thing over the last year. I don't have anything that feels both very weird and very attractive.
那有没有只是有吸引力的?我要求不高。
What about anything that's just attractive? I'll settle.
我仍然对我们上次讨论的一些事情感兴趣,可能很浅,或者我们没机会谈到。我仍然对一些可能影响认知表现的干预措施的基本测试感到兴奋,这些测试似乎被奇怪地忽视了。所以现在我正在资助德国的一些临床精神科医生,对素食者进行肌酸测试,这看起来相当令人兴奋。我认为关于二氧化碳和认知的文献现状很荒谬。我上次来的时候可能抱怨过这个。
I remain sort of interested in a few things we discussed last time, maybe very shallowly, or maybe we didn't have a chance to touch on. I remain excited about some basic tests of interventions that may affect cognitive performance, which seem pretty weirdly neglected. So right now I'm providing some funding to some clinical psychiatrists in Germany to do a test of creatine in vegetarians, which seems pretty exciting. I think the current state of the literature on carbon dioxide and cognition is absurd. I probably complained about this last time I was here.
我们来深入探讨一下。我没问这些问题是我的错。关于肌酸的背景:有一些研究,特别是一项研究表明,对于素食者,可能还有非素食者,服用肌酸可以让智商提高几分。即使在相对较小的样本中,这也是非常可测量的。按照试图让人更聪明的标准,这是一个相当大的效应量;按照通常寻找效应的标准,这算小的。大约是三分之一的标准差,这还算可观,但不算巨大。没有那么多干预措施有这种效果。如果我们能让每个人智商提高三分,那就太酷了。然后这件事就没有太多后续了,尽管它似乎比我们拥有的其他大多数让人更聪明的方法都要好,除了改善健康和营养。
Let's dive into this. It was a mistake of mine not to put these questions in. So just the background on this creatine: there have been some studies, one study in particular, that suggested that for vegetarians and potentially for non-vegetarians as well, taking creatine gives you an IQ boost of a couple of points. It was very measurable even with a relatively small sample. This was a pretty big effect size by the standards of people trying to make people smarter, small by the standards of people normally looking for effects. It's like a third of a standard deviation, which is respectable but not huge. There aren't that many interventions that have that effect. If we can make everyone three IQ points smarter, that's pretty cool. And then there was just kind of not much follow-up on this, even though it seems like this could be way better than most of the other options we have for making people smarter, other than improving health and nutrition.
关于杂食者的效果有综述,所以这方面研究得更好。我认为在杂食者身上有显著效果看起来不太可信。有一些机制研究,从机制上看,情况不太好。如果你看看肌酸如何起作用——我对这个领域了解不多,我们现在列出的所有领域都是我有时随意推测的。我真的很想把这个放在那里;应该有一个单独的类别来放我对 AI 的看法。总之,从机制上看,情况不太好。根据我们目前对生物学的了解,肌酸补充剂产生这种认知效果会令人惊讶。在素食者中,这有点可能,而且没有被排除。素食者的情况,我认为是一个不确定的结果,然后是这个非常积极的结果。所以似乎值得在素食者中再做一次有足够统计效力的检验。如果有什么结果,我会很惊讶,但我认为有可能。有些人会更惊讶;有些人觉得显然没什么。但我认为在素食者这一点上,5-10% 的概率是一个合理的赌注。
There are reviews on the effects in omnivores, so that's been better studied. I think it doesn't look that plausible that it has large effects in omnivores. There's been some looking into mechanisms, and in terms of mechanism, it doesn't look great. If you look at how creatine works—I don't know much about this area, all these areas we're listing now are just random things I'm speculating about sometimes. I really want to put that up there; there should be a separate category for my views on AI. Anyway, looking at mechanisms, it doesn't look that great. It would be surprising given what we currently know about biology for creatine supplementation to have this kind of cognitive effect. It's kind of possible and not ruled out in vegetarians. The state in vegetarians is, I think, one inconclusive thing and then this one really positive result. So it seems just worth doing a reasonably powered check in vegetarians again. I would be very surprised if something happened, but I think it's possible. Some people would be more surprised; some people are like obviously nothing. But I'm at like 5-10% seems like a reasonable bet on the vegetarianism point.
当我看到那篇论文时,他们选择素食者似乎主要是因为预期效果会更大,因为肌酸补充剂也会增加肉食者体内游离肌酸的水平。所以向不了解的听众解释一下:肉类含有一些肌酸,尽管比人们通常补充的量少得多,但素食者因为不吃肉,肌酸水平往往较低。所以补充剂可能效果更大。很可能这只是那项研究的一个选择,然后存在随机变异。有些研究——我肯定更多地更新了那些显示一切的研究方向。搞砸研究非常容易,或者很容易得到错误的结果,不仅仅是在 5% 的时间里结果在 p=0.05 水平显著,而是更频繁地出现错误结果,原因不明。总之,很可能只是一项研究碰巧得到了阳性结果,并且碰巧研究了素食者。他们这样做是有原因的;似乎效果应该更大。我认为既然我们有关于杂食者效果的负面证据,这看起来不太可能,尽管这也与杂食者效果小三分之二一致,这有点合理,并且与我们已知的相符。
When I looked at that paper, it seemed like they chose vegetarians mostly just because they expected the effect to be larger there, because it is the case that creatine supplementation also increases free creatine in the body for meat eaters. So just explain for listeners who don't know: meat has some creatine in it, although a lot less than people tend to supplement with, but vegetarians tend to have less because they're not eating meat. So the supplementation potentially has a larger effect. Most likely that was just a choice that study made, and then there was random variation. Some studies—I've definitely updated more in the direction of studies showing everything. It's very easy to mess up studies, or very easy to get results that are wrong, not even just in the 5% of the time you have results significant at p=0.05, but just radically more often than that you get results that are wrong for God knows what reason. Anyway, so most likely it's just a study that happened to turn a positive result and happened to be studying vegetarians. There was a reason they did it; it seemed like it should have a larger effect. I think since we've got negative evidence about the effects in omnivores, it doesn't seem that likely, although that would also be consistent with them just being three times smaller in omnivores, which would be kind of plausible and compatible with what we know.
你当时有点像,这看起来非常重要,但人们没有投入资金,没有进行足够的重复研究。所以你就决定做一次重复,一次预注册的重复,这就是我想要的。所以你就说,我要自己来做。说说这个吧。
You were kind of like, this seems really important but people haven't put money into it, people haven't run enough replications of this. So you just decided one replication, one pre-registered replication, that's all I want. So you were like, I'm going to do it myself. Talk about that for a minute.
我认为在这种情况下,提供资金可能不是困难的部分。但我很乐意做这类事情。我对提供资金非常感兴趣。我发了一个 Facebook 帖子,说我真的有兴趣提供资金,然后 EA 站出来了,说我知道一个实验室可能对此感兴趣。让我和他们联系。
I think in this case providing funding is not the hard part, probably. But I'm happy for stuff like this. I'm very interested in providing funding. I made a Facebook post like, I'm really interested in providing funding, and then EA stepped up and was like, I know a lab that might be interested in doing this. Put me in touch with them.
他们什么时候能有结果?
When might they have results?
大约一年后。
In about a year.
你期待知道结果吗?
Are you excited to find out?
是的,我很期待。我很想看看事情会如何发展。
I am, yeah. I'm excited to see how things go.
说说二氧化碳那个吧,因为这也是让我抓狂的一个。过去几个月,二氧化碳似乎确实可能对人们的智力产生巨大影响。在办公室,尤其是演讲厅,二氧化碳水平可能极高,在我们最需要聪明的时候让我们变笨。
Talk about the carbon dioxide one for a minute, because this is one that's also been driving me mad. The last few months, it does seem like carbon dioxide potentially has enormous effects on people's intelligence. In offices, and especially in lecture halls, you potentially have extremely elevated CO2 levels that are kind of dumbing us all down when we most need to be smart.
我几年前回顾过文献,从那以后只稍微关注了一下。但我认为目前的状况是,有一项研究显示二氧化碳的效应量大得离谱,其方法是:把人放在房间里,注入一些气体……
I reviewed the literature like a few years ago and I've only been paying a little bit of attention since then. But I think the current state of play is there was one study with preposterously large effect sizes from carbon dioxide, in which the methodology was: put people in rooms, dump some gas into...
所有房间中,有些气体的二氧化碳浓度非常高,效应量也大得离谱。比如,跟我家或者我刚搬出的房子里的二氧化碳水平相比,那栋房子里二氧化碳浓度最高的卧室,在伯克利学生的这项测试中产生了大约一个标准差的影响,这太荒谬了。这完全不合理。所以几乎可以肯定,这种效应大到你应该预期,当人们走进二氧化碳浓度升高的房间时,他们应该会觉得自己变笨了,或者明显感觉智商下降了。
All the rooms some of the gases were very rich in carbon dioxide and the effect sizes were absurdly large. They were like if you compare to the levels of carbon dioxide that occur in my house or like in the house I just moved out of, the most carbon dioxide rich bedroom in that house had like one standard deviation effect amongst Berkeley students on this test or something, which is absurd. That's like totally absurd. So it's almost certainly well, like it's such a large effect that you should expect that like people when they walk into a room with carbon dioxide, like with elevated carbon dioxide, they should just feel like idiots at that point, or they should feel like noticeably dumber in their own minds.
是啊,你会这么想。而且需要说明的是,浓度那么高的房间,人们会报告说‘我觉得很闷’。所以论文方法的一部分就是直接注入二氧化碳,以避免如果自然让房间二氧化碳浓度过高,人们会明显意识到自己在实验组而不是对照组。
Yeah, you would think that. And to be clear, the rooms that have levels that high, people can report like 'I feel it feels stuffy.' So part of the reason the methodology in the paper is just dumping in carbon dioxide to avoid, like if you make a room naturally that is too rich, it's going to also just be obvious that you're in the intervention group instead of the control.
是啊,不过公平地说,即使我不知道,也可能只是安慰剂效应之类的。嗯,我几乎可以肯定,这对我来说似乎不对。也许这不是在播客上公开说的好话题。那篇论文毕竟有很多受人尊敬的研究者。所以我想,如果能看到重复实验就好了。后来确实有一个完全相同的设计重复实验,p 值也是 0.01。所以现在我们有了两个精确的重复实验,p 值都是 0.01。这就是目前的状况。而且效应量大得离谱,大到如果这个效应是真的,你真的非常需要关注通风。比如这个房间可能,这太疯狂了。好吧,这栋楼通风很好,但我们仍然至少笨了三分之一标准差。
Yeah, although to be fair, even if I don't know at that point, even a placebo effect or just something. Yeah, I think almost certainly there's just, almost certainly that seems wrong to me. I maybe this is not a good kind of thing to be saying publicly on podcast. There's a bunch of respected researchers on that paper anyway. So I was like, it'd be great to see a replication of that. There was subsequently a replication with exactly the same design which also had like, you know, p equals 0.01. So now we've got like the two precise replications with p equals 0.01. So that's kind of where we're at. And also the effect is stupidly large, like so large that you really, really need to care about ventilation if that effect is right. Like this room probably is, this is madness. Well, this building is pretty well ventilated, but still we're like at least a third of a standard deviation dumber.
是啊,我敢肯定,亲爱的听众们,你们能听到我们在谈话过程中变得越来越笨,因为我们正在用毒气填满这个房间。嗯,所以我想最糟糕的情况可能是在会议室或董事会会议室里,人们就棘手问题进行长时间、持续的讨论。随着房间二氧化碳浓度升高,他们会逐渐变笨,也会变得更易怒。
Yeah, I'm sure dear listeners, you can hear us get dumber over the course of this conversation as we fill this room with poison. Um, yeah, so I guess like potentially the worst case would be in kind of meeting rooms or boardrooms where people are having very long, prolonged discussions about difficult issues. They're just getting like progressively dumber as the room fills up with carbon dioxide and more irritable as well.
是啊,这会很严重。我认为人们经常引用这一点来改善通风,但我认为人们并没有像他们相信时那样认真对待。我认为这是对的,因为我几乎可以肯定效应没有这么大。但如果真有这么大,你真的会想知道。然后这就有点像铅中毒之类的东西了。
Yeah, it would be pretty serious. And I think that people have often cited this in attempts to improve ventilation, but I think people do not take it nearly as seriously as they would if they believed it. Which I think is right because I think it's almost certainly the effect is not this large. But if it was this large, you'd really want to know. And then sort of this is like lead poisoning or something.
是啊,没错。嗯,我想这已经足以说服我睡觉时开窗了。我真的很不喜欢睡在没有通风、没有开门或开窗的房间里。也许我不该担心,因为晚上谁在乎我做梦时有多聪明呢?但嗯,我不知道怎么回事。我也没有像应该的那样深入研究。但我真的很想能够说,这并不难,效应足够大,也足够短期,从某种意义上说非常容易验证。就像,你还想要什么?已经有重复实验了。但我不确定,这些研究使用的认知测试并不太好。如果效应是真的,你应该能用几乎任何工具检测到。在某个时候,我只是想亲眼看到这个效应,你知道吗?我想亲眼看到它发生。我想看到房间里的人。
Yeah, that's right. Um, I guess well, this has been enough to convince me to keep a window open whenever I'm sleeping. I really don't like sleeping in a room that has no ventilation, no open door or window. Maybe I just shouldn't worry because at night who really cares how smart I'm feeling while I'm dreaming? But yeah, I don't know what's up. I also haven't looked into it as much as maybe I should have. But I would really just love to be able to say like it's not a hard, the effects are large enough that it's also short-term enough, it's just like extremely easy to check in some sense. It's like, what are you asking for? There's already been a replication. But like I don't know, the studies use like these cognitive batteries that are like not great. If the effects are real, you should be able to detect them in very basically any instrument. At some point I just like want to see the effect myself, you know? I want to actually see it happen. I want to see the people in the rooms.
似乎有不错的学术激励去做这件事,你会想,因为如果你率先提出这个问题,结果证明它极其重要,并导致建筑重新设计,你最终会出名。我不知道,这可能是一件大事。我的意思是,即使你不能从中获得经济利益,难道你不想因为识别出这个巨大的未被认识到的问题而获得赞誉吗?
Seems like there's a decent academic incentive to do this, you'd think, because you just end up kind of being famous if you pioneer this issue that turns out to be extraordinarily important and then causes buildings to be redesigned. I don't know, it could just be a big deal. I mean, even if you can't profit from it in a financial sense, wouldn't you just want the kudos for identifying this massive unrealized problem?
是啊,我的意思是,明确地说,我认为有一群人在研究这个问题,而且目前我们确实有,我知道的可能已经过时了,就是原始论文、直接重复实验和概念重复实验,都有很大的效应,但使用的工具都有点不可靠。概念重复实验是由一个研究通风的团体资助的,这并不奇怪。
Yeah, I mean to be clear, I think a bunch of people work on the problem and we do have at this point, I think there's the original, the things I'm aware of which is probably out of date now, is like the original paper, a direct replication, and a conceptual replication, all with like big looking effects, but all with like slightly dicey instruments. The conceptual replication is funded by like this group that works on ventilation, unsurprisingly.
哦,这很有趣。是啊,是啊,空气质量很重要。嗯,我认为学术界人士的看法,就学术界的正式共识过程而言,我认为他们会认为效应是真实的,只是没有人表现得好像那么大的效应真的存在。我认为他们怀疑学术界的这个过程是有道理的。我认为这确实让情况变得有点复杂,关于你到底因什么而获得认可。我认为应该获得认可的人,而且理应如此,是那些到目前为止一直在研究这个问题的人。这更像是为那些持怀疑态度的人进行验证,尽管每个人都在隐含地怀疑,因为他们在二氧化碳浓度高时并没有把它当作紧急情况来处理。包括我们现在。
Oh, that's interesting. Yeah, yeah, big air quality. Yeah, I think that like probably the take of academics, in so far as a formal consensus process in academia, I think it would be the effect is real, it's just that no one is behaving as if the effect of that size actually existed. And I think they're kind of right to be skeptical of the process in academia. I think that does make the situation a little bit complicated in terms of what you exactly get credit for. And I think people who would get credit should be, and rightfully would be, the people who have been investigating it so far. It's sort of more like checking it out for people who are skeptical, although everyone is kind of implicitly skeptical given how much they don't treat it like an emergency when carbon dioxide levels are high. Including us right now.
是啊,包括我们现在。嗯,赞扬你资助了那个重复实验。我有点,是啊,如果更多人主动坚持资助那些看似重要却被忽视的问题的重复实验,那就好了。
Yeah, including us right now. Well, kudos to you for funding that replication thing. I kind of yeah, it'd be good if more people took the initiative to really insist on funding replications for issues that seemed important where they were getting neglected.
是啊,我认为其中很多是,有很多好事人们可以做。我觉得人们主要是瓶颈,就是那些拥有相关专业知识和兴趣的人。这是我觉得人们可以大有作为的一个类别。我很期待看到进展。
Yeah, I think a lot of it is like it's a great, there are lots of good things for people to do. I feel like people are mostly the bottleneck, just like people who have the relevant kinds of expertise and interests. This is like one category where I feel like people could go far on. I'm excited to see how that goes.
去年,OpenAI 发布了一篇博客文章,让人们非常兴奋,显示用于训练前沿机器学习系统的算力大幅增加。我认为,对于吸收算力最多的算法,六年来投入的算力增加了 30 万倍。这似乎是近年来 AI 能力更令人印象深刻的一个潜在巨大驱动力。这是否意味着未来进展会更快,还是你认为随着算力增长趋于平稳,越来越难投入更多算力,进展会放缓?
Last year OpenAI published a blog post which got people really excited, showing that there had been a huge increase in the amount of compute used to train cutting-edge ML systems. And I think for the algorithms that had absorbed the most compute, there was a 300,000-fold increase in the amount of compute that had gone into them over six years. It seemed like that had been a potentially really big driver of more impressive AI capabilities over recent years. Would that imply faster progress going forward, or do you think it will kind of slow down as the increase in compute runs its course and it gets harder and harder to throw more processes?
对于这些问题,我认为这取决于你之前的视角。如果你之前的视角是凭感觉观察这个领域的进展,然后问自己‘这感觉像是很大的进步吗?’,那么一般来说这应该是坏消息——或者不是坏消息——它应该让你觉得 AI 离得更远了。你会想,‘嗯,确实有很多进步。我对进步的程度有一些直觉,但现在我了解到这种进步速度无法持续那么久,或者其中很大一部分是这种不可扩展的东西。’我们讨论还能走多远,但也许你在那段时间里有一百万倍的提升,然后还能再有一千倍或一万倍。我想当你开始触及——嗯,处理器的速度只能快到一定程度,然后还有购买大量这些设备的成本。人们能够扩大规模是因为之前它只占项目总成本的一小部分,但现在它已经占到所有 AI 项目总成本的很大一部分,就是购买足够的处理器。是的,很多项目都有很大的算力预算。我的意思是,通常它仍然比人员预算小,而且你可以再往前走一点,但它正在变得庞大。你应该预期,如果你正在训练人类级别的 AI 系统,那么这次训练运行的算力成本应该占全球产出的很大一部分。所以你可以说也许这种趋势可以持续到那个水平,但很可能不会以这种速度。它必须在达到像我们花费 GDP 的 2%用于 AI 训练计算机之前很久就放缓。如果你有那种凭感觉观察进展的视角,那么我认为这通常应该让你更新为更长的时间线。
At these problems, I think it just depends on what your prior perspective was. So if you had a prior perspective where you were eyeballing progress in the field and being like, 'Does this feel like a lot of progress?', then in general it should be bad news—or not bad news—it should make you think AI is further away. You're like, 'Well, there was a lot of progress. I had some intuitive sense of how much progress that was, and now I'm learning that that rate of progress can't be sustained that long, or a substantial part of it has been this unscalable thing.' We talk about how much more you could go, but maybe you had a million times over that period and you can have a further thousand times or something like that, maybe 10,000 times. I suppose it's when you start hitting—well, there's only so fast that processes are getting faster, and then there's also just the cost of buying tons of these things. People were able to ramp it up because previously it was only a small fraction of the total costs of their projects, but I guess it's now getting to be a pretty large fraction of the total cost of all these AI projects is just buying enough processors. Yeah, a lot of things have a large compute budget. I mean, still normally going to be small compared to staff budget, and you can go a little bit further than that, but it's getting large. You should sort of expect if you're at the point where you're training human-level AI systems that the cost of the compute for this training run should be a significant fraction of global output. So you could say maybe this sort of trend could continue until you got up there, but it's probably not at this pace. It's going to have to slow down a long time before it gets to like we are spending 2% of GDP on computers doing AI training. If you had that perspective of eyeballing progress, then I think it should generally be an update towards longer timelines.
我认为如果你有一个视角——这更像是我的通常立场——你会说,‘伙计,这很难说。凭感觉观察进展非常困难,然后问自己,这有多令人印象深刻?在象棋或围棋中达到人类水平,或者对图像进行分类,或者完成这个特定的图像分类任务,有多令人印象深刻?’我发现很难真正凭感觉判断那种进展并做出预测。我认为如果你的估计来自‘嗯,我们认为还有一些。我们有一些粗略的方法来估计可能需要多少算力。我们可以通过与进化所做的优化进行类比,或者通过训练时间的外推,或者通过关于人脑的论证——这些论证确实锚定在算力数量上’,那么我认为你可能会有一种更接近‘嗯,这告诉我们一些理论上的东西。这些论证会涉及使用大量算力。那种规模扩展需要大量的工程努力。存在很多真正的不确定性,尤其是当你谈论中等时间线时,那种工程努力是否真的会被投入,以及那种花钱的意愿是否真的会实现。’我认为那可能会让你朝着‘是的,显然人们正在付出努力,工程进展的风险是合理的’方向移动。所以如果你的估计真的是由算力驱动的——这有点像未来学家所做的旧估计的风格。如果你看看这种风格的一个早期估计,卡尔·舒尔曼的就是一个非常著名的这种风格的估计,他们觉得‘你投入多少算力真的很重要。’如果你有那种观点,然后你看到算力支出增长非常迅速,我想那是证据表明它可能继续增长,因此时间线会比你想的更短。
I think if you had a perspective—and this is sort of more where I'm normally coming from—where you're like, 'Man, it's hard to tell. It's very hard to eyeball progress and be like, how impressive is this? How impressive is it to be human at chess or Go or classify images as well, or to do this particular image classification task?' I find it very hard to really eyeball that kind of progress and make a projection. I think if instead your estimates were coming from like, 'Well, we think there's some more. We have some sketchy ways of estimating how much compute might be needed. We can make some sort of analogy with the optimization done by evolution, or by an extrapolation of training times, or by arguments about the human brain which are really anchored to amounts of compute,' then I think you might have a perspective that's more like, 'Well, this tells us something about on paper. These arguments would have involved using large amounts of compute. There's a lot of engineering effort in that kind of scale-up. There's a lot of genuine uncertainty, especially if you're talking about moderate timelines, of whether that kind of engineering effort will actually be invested and whether that kind of willingness to spend will actually materialize.' I think that might make you move in the direction of, 'Yes, apparently people are putting in the effort, and engineering progress is reasonably risky.' So if instead you were doing an estimate that was really driven by how much compute—this is sort of the style of the old estimates futurists made. If you look at one of the earlier estimates of this flavor, and Carl Shulman's is a very famous estimate of this flavor, where they're like, 'It really matters how much compute you're throwing at this task.' If you have that kind of view and then you see that compute spending is rising really rapidly, I guess that's evidence that maybe it can continue to rise and therefore it will be shorter than you would have thought.
有些人似乎认为,我们也许能够仅通过使用我们今天已有的算法来创造通用人工智能,但等待一二十年的处理能力上线,推进芯片技术,并建设基础设施。你认为这有多现实?在你看来,这是一个活生生的可能性吗?
Some people seem to think that we may be able to kind of create a general artificial intelligence just by using the algorithms that we have today, but waiting for another decade or two worth of processing power to come online and progressing the chips and just building that infrastructure. How realistic do you think that is? Is that a live possibility in your mind?
我认为这很难说,但这绝对是一个活生生的可能性。我认为很多人有一种直觉——有些人有一种直觉,觉得‘这显然是会发生的’,而我不——我不认为我赞同那种直觉。另一面有些人有一种直觉,觉得‘显然有一些我们无法理解的重要事情会很困难,所以很难知道需要多长时间才能发展,而且它会比扩展算力所需的时间长得多。’我也不——我也不太赞同那种直觉。我有点觉得,是的,真的很难知道。这似乎是合理的。我们很难根据先验理由排除它。我们的观察与事情主要由算力驱动相当一致,或者你可以把它想成,‘算力与进展、概念进展或算法进展之间的权衡率是多少?’我认为我们的观察与算力的重要性相当兼容,也与现有事物的扩展最终能让你达到——是的,我肯定有那种观点:最终足够的扩展几乎肯定会成功。只是多少的问题,以及这是否是你在未来一二十年会看到的,还是它会让你远远超过物理极限。总的来说,我最终只是非常不确定。我认为很多事情都是可能的。
I think it's really hard to say, but it's definitely a live possibility. I think a lot of people have an intuitive—some people have an intuition that's very much like, 'That's obviously how it's going to go,' and I don't—I don't think I sympathize with that intuition. Some people on the other side have an intuition like, 'Obviously there are really important things we don't get to understand which will be difficult, so it's hard to know how long it will take to develop, and it's going to be much longer than the amount of time required to scale up computing.' I also—I'm not super sympathetic to that either. I kind of feel like, yeah, it's really hard to know. It seems plausible. It's hard to rule it out on our prior grounds. Our observations are pretty consistent with things being mostly driven by compute, or you could sort of think of it as, 'What is the trade-off rate between compute and progress, conceptual progress or algorithmic progress?' I think our observations are pretty compatible with a lot of importance on compute, and also are compatible with scale-up of existing things eventually getting you to—yeah, I guess that's definitely a view I have: eventually enough scale-up will certainly almost certainly work. It's just a question of how much, and whether that's what you're going to be seeing over the next one to two decades, or whether it's going to take you far past physical limits. Overall, I end up just pretty uncertain. I think lots of things are possible.
算力重要性的这个问题与莫拉维克悖论有什么关系?我想对于没有听说过它的听众来说,那是什么?
How does this question of the importance of compute relate to Moravec's Paradox? And I guess what is that for the audience, people who haven't heard it?
是的,所以我想这是一个普遍的观察:有些任务人类认为在智力上很困难——比如经典例子下棋——而其他任务他们不认为在计算上困难,比如拿起一个物体,观察一个场景,看到物体在哪里,拿起一个物体,操作它。而且情况似乎是,人们认为传统上智力挑战大的任务比人们怀疑的更容易,相对于那些人们认为智力要求不高的任务。这不是非常直接,因为仍然有大量智力探究人们不知道如何自动化。我认为这就是大致的意思。
Yeah, so I guess it's the general observation that there are some tasks humans think of as being intellectually difficult—like classic examples like playing chess—and there are other tasks that they don't think of as computationally difficult, that are like picking up an object, looking at a scene, seeing where the objects are, picking up an object, manipulating it. And it has seemed to be the case that the tasks that people think of as traditionally intellectually challenging were easier than people suspected, relative to the tasks people thought of as not that intellectually demanding. It's not super straightforward because there's still certainly big chunks of intellectual inquiry that people have no idea how to automate. I think that's the general idea.
你的意思是,比如人类觉得哲学很难,计算机做哲学也很难,它们似乎没有在这方面超越我们。或者像数学或科学。我想人们可能常常觉得,对人类来说,做数学和下复杂的棋类游戏感觉差不多,但对机器来说,这些任务并不相似。棋类游戏要容易得多。
Pattern you mean like for example humans think of philosophy as difficult and it's also hard for computers to do philosophy, or they don't seem to be beating us at that. Or like mathematics or science. I guess people might often think to a human it feels sort of similar maybe to be doing mathematics and to be playing a really complicated board game, but to a machine those tasks are not that similar. The board game's way easier.
是的,棋类游戏结果证明相对于其他所有事情来说非常非常容易。即使是像围棋这样被认为是难度最高的棋类游戏,也比人类自动化的其他任何任务要容易得多。
Yeah, board game turned out was very, very easy relative to all other things. Even for, like at this point, it sort of Go as a reasonable guess for the hardest board game, it was like much easier than any of these other tasks for humans to automate.
是的,我认为总的来说,这里的一个原因是人类有意识接触到的推理在计算上并不那么密集。我们有一些理解,也许这是早期对 AI 乐观的一部分。我们知道当人类有意识地操作数字或符号,或者真正把注意力集中在任何事情上时,他们做的事情并不快。你知道,一个人如果能每秒进行 100 次运算就很幸运了。如果一个人能以这样的速度做乘法,那简直不可思议,你会觉得哇太厉害了。但在那之下,人类还有一层使用了多得多的计算。所以实际上,很多困难,尤其是在以算力为中心的世界里,当你看到一个任务时,你会说这个任务对人类相对于机器有多难。很多问题在于人类在执行任务时如何利用他们拥有的所有计算能力。而对于这些涉及有意识推理的任务,可能不太可能,至少意识部分在计算上并不有趣。然后对于像棋类游戏这样的事情,你还有进一步的问题:人类并没有受到很大的选择压力去很好地玩棋类游戏,他们并没有很好地利用大脑中的算力。我认为最可能的猜测是,如果你进化,你可以进化出比人类小得多、但更擅长下棋的动物。
Yeah, I think I mean in general part of what's going on there is like the reasoning humans have conscious access to is just not that computationally demanding. Like we sort of have some understanding and maybe this is part of the very early optimism about AI. We understand like when a human is consciously manipulating numbers or symbols or actually casting their attention to anything, they're just not doing things that fast. You know, a human is lucky if they can be doing like 100 operations per second. That's like kind of insane if a human is able to multiply numbers at a kind of speed that implies that or something, you're like wow that's incredible. But when a human is doing sort of underneath that there's this layer which is using vastly vastly more computation. So like in fact a lot of the difficulty, especially if you're in a compute-centric world, is like when you look at a task you say how hard is that task for a human relative to machine. A lot of the question is like how well is a human leveraging all the computational capacity that they have when they're doing that task. And for these tasks that like any task that is involving conscious reasoning maybe is less likely, at least the conscious part is sort of not doing anything computationally interesting. And then you sort of have this further issue for things like board games where like a human is not under much selection pressure to use a human is not really evolved to play board games well. They're like not using much of the compute in their brain very well at all. I think like best guess would be if you evolved, you could evolve like much much tinier animals that are much much better at playing board games than humans.
是的,人类大脑并没有把很大一部分专门用于视觉处理。所以实际上这需要大量的算力,我想也需要进化来很好地利用大脑的那部分。
Yeah, it's not the case that the human brain has a ridiculous fraction of it devoted to visual processing. So in fact that has just required a ton of compute and I guess also evolution to make to use that part of the brain well.
是的,我不确定具体数字,但当我们谈论对数尺度时,这并不重要。视觉使用了大脑相当大的一部分,并且针对它进行了高度优化。所以当人们下棋时,他们可能也在利用大脑的很大一部分。再次,主要问题是视觉皮层确实针对视觉进行了优化,他们确实在利用大脑做这些。而当你做数学或玩游戏时,你是最幸运的情况,因为它有足够的直觉意义或很好地映射到直觉上,你可以建立这些抽象,让你利用大脑的全部力量来完成这个任务。但这相当不寻常,而且不是先验明显的。这只是一个事后故事。你可以想象人们实际上能够利用整个视觉处理机制来玩一些棋类游戏。你可以想象。我认为这实际上也是一种现实的可能性。所以如果我们以围棋为例,看看我们现在解决围棋的方式,使用完全蛮力策略(比如 alpha-beta 搜索)击败人类所需的算力,即使与人类视觉皮层或视觉系统相比,也是相当多的。而且你可以提出一个合理的论点,即人们能够利用很多这样的机制,比如他们在下围棋时能够重用很多机制,在象棋中程度稍轻,用于对如何下棋进行位置评估的直觉。
Yeah, I don't offhand what the number is, but sort of when we're talking about like the log scale it just doesn't even matter that much. It uses a reasonable like vision uses a reasonable chunk of the brain and it's extremely well optimized for it. So like when people play board games they're also probably leveraging some very large fraction of their brain. And again the main problem is like visual cortex is really optimized for doing vision well and like they're really using their brain for all that. And like you're sort of the luckiest case when you're doing mathematics or playing a game is like somehow it has enough makes enough intuitive sense or maps on well enough intuitively you can build up these abstractions allow you to leverage your brain like the full power of your brain to do that task. But it's like pretty unusual and that's not obvious a priori. This is kind of just an after-the-fact story. You could imagine people are actually able to sort of use the entire machinery of visual processing to like play some board games. You can imagine that. I think that's actually also kind of a live possibility. So if we talk about Go for example, and we look at the way that we've now resolved Go, the amount of compute you would need to beat humans at Go using entirely a brute force strategy, like using alpha-beta search or something, is kind of a lot even compared to the human visual cortex or human visual system more probably. And like you could make a plausible case that people are able to use a lot of that machinery, like they're able to reuse a lot of machinery when playing Go and to a slightly lesser extent chess, for doing like position evaluation intuitions about how to play the game.
你是说,你认为大脑中负责视觉处理的部分被用来识别围棋中的模式,并被用来做棋类游戏的工作?
You're saying that you think the part of the brain that does visual processing is kind of getting brought online to notice patterns in Go and it's getting co-opted to do the board game work?
是的,至少这在先验上是合理的,并且与我们观察到的自动化游戏的难度一致。是的,我们知道的并不多。很多事情都与我们的观察一致。
Yeah, at least that's plausible a priori and sort of consistent with our observations of how hard it is to automate the game. Yeah, we just don't know very much. Lots of things are consistent with our observations.
你希望发现我们受到算力还是算法进步的约束吗?
Do you hope to find out that we're constrained by compute or algorithmic progress?
是的,所以我通常认为,从某种意义上说,不会只受其中一个约束。而是要看各自的边际回报。我关心的是更多算力和更多算法进步之间的替代率。总的来说,从长远来看,如果很多算法进步只能替代少量算力,那似乎更好。所以你越处于那个世界,不同参与者的算力需求就越集中。所以在构建非常强大的 AI 系统时,所有构建者都将不得不使用他们计算资源的很大一部分。任何想要开发非常强大 AI 的参与者也将使用世界资源的很大一部分。这意味着更容易知道谁在参与这个游戏。某人单方面做某事要困难得多。参与者更容易有现实的机会进行监控和执行,也更容易有现实的机会坐下来交谈,可能不是字面上的房间,但达成理解和协议。这是一件事。另一件事可能是,算法进步替代硬件进步越困难,随后的进步速度相对于历史观察到的可能就越慢。所以如果你处于一个世界,其中仅仅聪明的思考就能极其快速地推动 AI 进步,而问题只是我们没有投入那么多聪明的思考,那么你可以想象,随着 AI 规模扩大并能够自动化所有思考,进步会非常快,这可能意味着从长期对齐问题变得明显到开始重要之间的时间更短。
Yeah, so I generally think it's, I mean in some sense it's not going to be being constrained by one or the other. It's going to be those marginal returns to each. I'm like what is the rate of substitution between more compute and more algorithmic progress. In general, I think it seems better from a long-term perspective if it takes a lot of algorithmic progress to substitute for a small amount of compute. So the more you're in that world, the more concentrated different actors' compute needs are. So at the point when you're building really powerful AI systems, sort of everyone who's building them is going to have to use, I mean again if the world is sort of paying attention, they're going to be using a very large fraction of their computational resources. And any actor who wants to develop very powerful AI will also be using a very reasonable fraction of the world's resources. And that sort of means that it is much easier to know who is in that game. It's much harder for someone to unilaterally do something. It's much easier for the players to be having a realistic chance of monitoring enforcement and also just have a realistic chance of getting in a room and talking to each other, probably not literally a room, but like reaching understanding and agreements. That's one thing. And maybe the other thing is the harder it is for algorithmic progress to substitute for hardware progress, the slower the subsequent rate of progress is likely to be relative to what we've observed historically. So if you're in a world where it turns out that just clever thinking really can drive AI progress extremely rapidly and the problem is just that we haven't had that much clever thinking to throw at the problem, you can really imagine as one scales up AI and is able to automate all that thinking having a pretty fast ongoing progress, which might mean there's less time between when long-term alignment problems become sort of obvious and start mattering.
AI 可以开始帮助解决这些问题,而如果这些问题没有得到解决,后果可能是灾难性的。所以一般来说,如果算法上的巧妙想法能大大缩短这段时间,那有点糟糕——AI 不太可能对硬件进步的速度产生惊人的一夜之间的影响。当然它也可能加速,自动化也会有所帮助。但如果你认为算力是主要因素,那么这将是一个更渐进的过程。从机器学习开始用于重要事情,到我们注意到它们在哪里有效、哪里无效,再到很多事情被委托给机器学习,这之间会有更长的时间。相比之下,在算法驱动的情况下,能力的变化似乎非常突然。
AI can start helping with them and the point where like it's catastrophic to have not resolved them so like just generally if algorithmic if clever ideas can shorten that period a lot that's a little bit bad it's a little bit less likely that the automation like that AI will have an incredible overnight effect on like the rate of Hardware progress I andan it will also presumably accelerate it like automation will help there as well but so you think if Compu is what predominantly matters then it's going to be a more gradual process well we'll have like longer between the point when machine learning gets starts to get used for for important things and we start noticing where they work and where don't work uh and when like a lot of things are getting delegated to to machine learning relative to the algorithmic case where it seems like you get like really quite abrupt changes in in the capabilities
是的,我认为这在很大程度上——如果 AI 研究的性质发生变化,这也可能改变——但很大程度上是因为硬件是一个非常成熟的行业,投入了大量资源,性能也相当清楚,很难再翻倍投资,而且它不像人力资本质量那样对奇怪的问题敏感。你只需要明白你必须做大量实验。它相对资本密集,而且存在相当大的滞后。所以总的来说,它似乎会更稳定,听起来是个好消息。这是人们可能对更快的 AI 进步感到更兴奋的原因之一。但你可能认为,现在更兴奋的最大原因是:如果你现在有更快的 AI 进步,你处于一个我们正在尽可能好地利用所有可用算力的阶段,那么后续的进步可以更稳定一些。如果你现在 AI 进步较慢,那么人们只有在清楚可以自动化大量人类劳动时才会真正开始大量投资,然后你会有一个爆发效应,随着人们真正开始投资,进步会突然加速。
Yeah, I think a lot of that — and this could also change if the nature of AI research changed — but like a lot of that is from Hardware being like this very mature industry with like lots of resources being thrown at it and like performance being pretty well understood and like it would be hard to like double investment in that and also like it's not that sensitive to like weird questions about like quality of human capital or something you just sort of understand you have to do a lot of experimentation It's relatively Capital intensive there's quite big lags as well yeah so it just seems like generally it would be more stable and sounds like good news and this is like one of the reasons one might give for being more excited about faster AI progress now but you might think I it's probably the biggest reason to be more excited is like if you have faster AI progress now you're in the regime where we're sort of using IF you manage to like get to some Frontier we're using all the available computation as well as you could then like subsequent progress can be a little bit more stable and if you have less AI progress now like at some point people only really start investing a bunch once it becomes clear they can automate a whole bunch of human labor um then you have this more like yeah with flash effect where you have a burst of progress as people really start investing
几周前,我们发布了一段与 Pushmeet Kohli 的对话,他是伦敦 DeepMind 的 AI 鲁棒性和可靠性研究员。粗略总结一下 Pushmeet 的观点,我认为他提出了几个关键主张。第一,在他看来,对齐和鲁棒性问题在机器学习系统的发展过程中无处不在,因此需要该领域每个人的一定关注。根据 Pushmeet 的说法,这使安全研究与非安全研究之间的区别变得模糊。他认为从事能力研究的人也在帮助安全,而提高可靠性也会提高能力。然后你就能设计出做你想做的事情的算法。第二,他认为可靠性和鲁棒性的一个重要部分是尝试忠实地向机器学习算法传达我们的愿望,这类似于——尽管更困难——与他人沟通、让他们真正理解我们意思的挑战,当然,与动物或机器学习算法相比,与其他人类沟通更容易。第三点是一种普遍的乐观情绪:DeepMind 正在大量研究这个问题,他们渴望雇佣更多人来解决这些问题,而且我们很可能能够随着机器学习算法影响力的增加而逐步解决 AI 对齐问题。我知道你还没有机会听完整场采访,但你浏览了文字记录。首先,你认为 Pushmeet 在哪里是正确的?你同意哪些观点?
A few weeks ago we published a conversation with Pushmeet Kohli, who's an AI robustness and reliability researcher at DeepMind in London. To heavily summarize Pushmeet's views, I think he made a couple of key claims. One is that alignment and robustness issues in his view are kind of everywhere throughout the development of machine learning systems, so they kind of require some degree of attention from everyone working in the field. According to Pushmeet, this kind of makes the distinction between safety research and non-safety research somewhat vague and blurry. He kind of thinks people who are working on capabilities are also kind of helping with safety, and improving reliability also improves capabilities. Then you can actually design algorithms that do what you want. Secondly, I think he thought that an important part of reliability and robustness is going to be trying to faithfully communicate our desires to machine learning algorithms, and that this is kind of analogous — although a harder instance — of the challenge of just communicating with other people, getting them to really understand what we mean, although of course it's easier to do that with other humans than with animals or machine learning algorithms. And a third point was just a general sense of optimism that DeepMind is working on this issue quite a lot and they're keen to hire more people to work on these problems, and a sense that probably we're going to be able to gradually fix these problems with AI alignment as we go along and machine learning algorithms get more influential. I know you haven't had a chance to listen to the whole interview but you skimmed over the transcript. Firstly, where do you think Pushmeet is getting things right? Where do you agree?
所以我当然同意,让系统做你想做的事与让它们更强大之间存在紧密联系。我同意基本的乐观态度:人们需要解决让系统做我们想做的事的问题。我认为人们很可能会找到这个问题的好解决方案。我认为即使没有长期主义者,也许有一个有趣的干预:长期主义者是否应该考虑这个问题以增加概率?我认为即使没有长期主义者的行动,一切也很有可能完全没问题。所以从这个意义上说,我完全同意这些主张。
So I certainly agree that there's this tight linkage between getting a system to do what you want and making them more capable. I agree with the basic optimism that people will need to address the getting a system to do what we want problem. I think it is more likely than not that people will have a good solution to that problem. I think even if you didn't have sort of long-termists, maybe there's this interesting intervention of should long-termists be thinking about that problem in order to increase the probability. I think even absent the actions of the long-termists, there's like a reasonably good chance that everything would just be totally fine. So in that sense, I'm on board with those claims definitely.
是的,我认为我有点不同意,即认为存在一个有意义的区别:一类活动的主要效果是改变各种事情成为可能的日期,另一类活动的主要效果是改变发展轨迹。我认为这正是从事对齐工作的主要特征:你关心的是这种差异化的进步,即朝着能够构建做我们想做的事的系统的方向前进。我认为从这个角度来看,AI 工作的平均贡献在那方面几乎定义上为零,因为如果你把所有 AI 工作增加一个单位,你只是把所有事情都提前了一个单位。所以我认为这确实意味着有一个明确定义的问题:我们能否以某种方式改变轨迹?这是一个值得思考的重要问题。
Yeah, I think that I would disagree a little bit in thinking that there is a meaningful distinction between activities whose main effect is to change the date by which various things become possible and activities whose main effect is to change the trajectory of development. I think that's sort of the main distinguishing feature of working on alignment per se: you care about this differential progress towards being able to build systems that do what we want. I think in that perspective, it is the case that the average contribution of AI work is sort of almost by definition zero on that front, because it's sort of bringing the entire — if you just increased all the AI work by a unit, you're just bringing everything forward by one unit. And so I think that does mean there's this well-defined thing which is: can we change the trajectory in a way? And that's an important problem to think about.
我认为还有一种非常重要的区别:一种失败最有可能破坏文明的长期轨迹,另一种失败最有可能成为系统实际有用或赚钱的直接障碍。也许理解这种区别的一种方式是与你提到的第二点相关:向机器学习系统传达你的目标与与人沟通非常相似。我认为向机器学习系统传达你的目标是一个难题,我们可以将其视为一个能力问题——它们能理解人们说的话吗?它们能形成那种内部模型,让它们理解我想要什么,或者理解——在某种意义上,这与预测 Paul 会做什么的问题非常相似,或者说是那个问题的一个小片段:预测在什么条件下 Paul 会对你所做的事情感到满意。这就是我们沟通时处理的大部分内容。
I think there's also a really important distinction between the kind of failure which is most likely to disrupt the long-term trajectory of civilization and the kind of failure which is most likely to be an immediate deal breaker for systems actually being useful or producing money. And maybe one way to get at that distinction is it's related to the second point you mentioned: communicating your goals to an ML system is very similar to communicating with a human. I think there is a hard problem of communicating your goals to an ML system, which we could view as a capabilities problem — like are they able to understand things people say? Are they able to form the kind of internal model that would let them understand what I want, or understand sort of, you know, in some sense it's very similar to the problem of predicting what Paul would do, or it's like a little slice of that problem: predicting under what conditions Paul would be happy with what you've done. That's most of what we're dealing with when we're communicating.
如果我在和你交谈,只要我能给你一个完美的自我模型,我就会完全满意。那样问题就解决了。我认为这是让 AI 系统真正有用的一个非常重要的 AI 难题。但我觉得这不太可能导致长期的不良方向,主要是因为我们担心的是当 AI 系统变得非常强大,对周围世界和互动的人有很好理解时的行为。真正令人担忧的情况是,系统非常清楚人们在各种条件下会做什么,非常清楚他们想要什么,但却没有动力去行动。它们理解保罗想要什么,但并不试图帮助保罗得到他想要的。我认为很多有趣的困难,尤其是从长远角度来看,确实是确保那里不再出现差距。长期视角中最重要的问题与人们为了让 AI 系统具有经济价值而面临的问题之间存在差距。我确实认为有很多重叠之处;人们正在研究的那些让 AI 系统更有价值的问题也直接有助于长期结果。但如果你有兴趣差异化地改变轨迹或提高长期顺利发展的概率,你更倾向于专注于那些在短期内对 AI 系统经济价值并非必需的问题。这取决于你的动机,或者你如何选择问题或优先考虑问题。
If I were talking with you, I would be completely happy if I just managed to give you a perfect model of me. Then the problem is solved. I think that's like a really important kind of AI difficulty for making AI systems actually useful. I think that's less core to the kind of thing that could end up pushing us in a bad long-run direction, mostly because we're concerned about behavior as AI systems become very capable and have a very good understanding of the world around them and the people they're interacting with. The really concerning cases are ones where systems understand quite well what people would do under various conditions, understand quite well what they want, but are not motivated to act. They understand what Paul wants but aren't trying to help Paul get what he wants. I think a lot of the interesting difficulty, especially from a very long-run perspective, is really making sure that no gap opens up there again. There's a gap between the problems that are most important on the very long-run perspective and the problems that people will most be confronting in order to make AI systems economically valuable. I do think there's a lot of overlap; both of those problems that people are working on that make AI systems more valuable are also helping very directly with the long-run outcome. But if you're interested in differentially changing the trajectory or improving the probability that things go well over the long term, you're sort of more inclined to focus precisely on those problems which won't be essential for making AI systems economically useful in the short term. That's distinctive to what your motivation is or how you're picking problems or prioritizing problems.
我想 Wish Me 的一个底线是,那些希望确保 AI 发展顺利的人不必特别挑剔他们是在做安全相关的事情,还是仅仅在构建一个使用机器学习的新产品。听起来你对此有点怀疑,或者你认为理想情况下,人们应该在中期内致力于那些似乎特别强调鲁棒性和可靠性的工作。
One of the bottom lines for Wish Me, I guess, was that people who want to make sure that AI goes well need not be especially fussy about whether they're working on something that's safety-specific or something that is just building a new product that works well using machine learning. Sounds like you're a little bit more skeptical of that, or you think ideally people should in the medium term be aiming to work on things that seem like they disproportionately push on robustness and reliability.
是的,我认为最关心长期轨迹的人在每个领域都面临这种困境。如果你生活在一个认为人类几乎所有最严峻的挑战都是由人类行为引起的世界里,或者是由作为生产进步一部分的事情引起的,比如建造新技术同时也带来主要风险,那么你就必须挑剔。如果你想改变长期轨迹,仅仅因为做一个随机项目或改进一个随机产品的平均效果是你在帮助解决我们关心的问题,但同时你也在使这些问题在时间上更接近我们。如果你只是让普通产品工作,这大致是平衡的。有一些细微的区别:如果你有动力让产品工作得好,并且更强调让产品鲁棒,我认为你会做出一系列有益的低层决策。我绝对认为通过挑剔你解决的问题,你可以产生相当大的影响。
Yeah, I think people who are most concerned about long-term trajectory face this dilemma in every domain. If you live in a world where you think that almost all of humanity's most serious challenges are caused by things humans are doing, or by things that are part of productive progress, like building new technologies that also pose the main risks, then you kind of have to be picky. If you're a person who wants to change the long-term trajectory, just because the average effect of working on a random project or making a random product better is that you're helping address the kinds of problems we're concerned about, but you're also contributing to bringing those problems closer to us in time. It's roughly a wash if you're making the average product work. There are subtle distinctions: if you are motivated to make products work well, and you want to have more of an emphasis on making this product robust, I think you're just going to make a bunch of low-level decisions that will be helpful. I definitely think you can have a pretty big impact by being fussy about which problems you work on.
有一个悬而未决的问题:如果 AI 整体进展更快,我们是否应该高兴?如果我们能把整个事情加快 20%,包括安全和能力,会怎样?据我所知,对此没有共识。人们对看到一切按比例加速感到高兴的程度差异很大。
There's this open question of whether we should be happy if AI progress across the board just goes faster. What if we can just speed up the whole thing by 20%, both safety and capabilities? As far as I understand, there's no consensus on this. People vary quite a bit on how pleased they would be to see everything speed up in proportion.
是的,我认为没错。我的观点,也是相当常见的观点,是从对齐的角度来看,这并不那么重要。主要是它会加速一切发生的时间。有一些二阶项真的很难推理,比如拥有更多或更少的计算硬件有多好,或者在强大系统开发之前世界上发生更多或更少的政治变化有多好。人们对此是好是坏非常不确定,但我的看法是净效应很小。主要的是,加速 AI 在接下来一百年的视角下重要得多。如果你关心未来 100 年内人类和动物的福祉,加速 AI 看起来相当不错。所以我认为主要的好处是:更快的 AI 进展意味着人们在短期内会更快乐。如果你关心长期,这大致是平衡的,人们可以争论它是略微正面还是略微负面。主要是它加速了我们前进的方向。
Yeah, I think that's right. My take, which is reasonably common, is that it doesn't matter that much from an alignment perspective. Mostly it will just accelerate the time at which everything happens. There are some second-order terms that are really hard to reason about, like how good it is to have more or less computing hardware available, or how good it is for there to be more or less kinds of political change happening in the world prior to the development of powerful systems. People are very uncertain about whether that's good or bad, but my take would be the net effect there is kind of small. The main thing is that accelerating AI matters much more on a next-hundred-years perspective. If you care about welfare of people and animals over the next 100 years, acceleration of AI looks reasonably good. So I think that's the main upside: faster AI progress means people are going to be happy over the short term. If you care about the long term, it is roughly a wash, and people could debate whether it's slightly positive or slightly negative. Mostly it's just accelerating where we're going.
在给人们具体的职业建议时,这是我们试图回答的更棘手的问题之一。在我看来,如果你是一个拥有机器学习博士学位或非常擅长机器学习的人,但目前无法得到一个特别注重安全的职位,或者这个职位对安全的影响会不成比例地大于能力,那么接受一份只是普遍推进 AI 的工作可能仍然是好的。主要是因为你会接触到前沿,提升你的职业资本很多,并获得相关的理解。这种工作大致是平衡的:它稍微加速了事情,一切都按比例进行,不清楚是好是坏。但之后你有可能去从事更专注于对齐的工作,而这在方程中是主导项。这看起来合理吗?
This has been one of the trickier questions that we've tried to answer in terms of giving people concrete career advice. It seems to me if you're someone who has done a PhD in ML or is very good at ML, but you currently can't get a position that seems especially safety-focused, or it's going to disproportionately affect safety more than capabilities, it is probably still good to take a job that just advances AI in general. Mostly because you'll be reaching the cutting edge of what's going on, improving your career capital a lot, and having relevant understanding. The work is close to a wash: it speeds things up a little bit, everything goes in proportion, it's not clear whether that's good or bad. But then you potentially later on go and work on something that's more alignment-specific, and that kind of is the dominant term in the equation. Does that seem reasonable?
是的,我认为这合理。
Yeah, I think that's reasonable.
我基本上同意。我觉得对于一类建议——比如你应该做一件现在对你的价值观大致中立的事,但未来会有机会让你做出选择——人们会有一些直觉上的犹豫。我理解这种犹豫,但我认为这大致是对的。所以想象两个可能的世界:一个世界里,从事机器学习和 AI 的人翻倍,但其中一半真正关心长期发展,确保 AI 以对人类长期有益的方式发展。这听起来是个好交易。我们可能现在做工作的机会会少一些,我认为这是主要的负面影响:思考对齐问题本身的时间会减少。但另一方面,如果领域中有很大一部分人真正关心把事情做好,那似乎真的很好。我预期一个具有这种特质的领域更有可能以对长期有利的方式处理问题。而且我认为你可以把这个想法缩小。对我来说,最容易想象的是领域中有相当一部分人持这种态度,但我认为,如果说有什么不同的话,那些处于边缘的人一开始可能有更大、更好的成本效益分析。
Seems basically right to me. I think there's some intuitive hesitation with a family of advice that's like: you should do this thing which we think is roughly a wash on your values now, but there will be some opportunity in the future where you can sort of make a call. I think there's some intuitive hesitation about that, but I think that is roughly right. So if you imagine there are two possible worlds: in one, there's twice as many people working on machine learning and AI, but half of them really care about the long term and ensuring that AI is developed in a way that's good for humanity's long term. That sounds like a good trade. We maybe then have less chance, less opportunity to do work right now. I think that's the main negative thing: there'll be less time to think about the alignment problem per se. But on the other hand, it just seems really good if a large fraction of the field really cares about making things go well. I just expect a field that has that character to be much more likely to handle issues in a way that's good for the long term. And I think you can sort of scale that down. It's easiest for me to imagine the case for a significant fraction of the field being like that, but I think that if anything, the marginal people at the beginning are having a probably larger, better cost-benefit analysis for them.
是的,我想我是在建议,如果你找不到一个专门做对齐的工作,那么这就是该做的事。所以他们想加入你的团队,但还不够好,需要学更多。或者可能团队扩张速度有限,所以即使他们很优秀,你也无法像人们加入那样快地招聘。但假设你必须确保,当人们进入这些我们认为目前只是中性但有助于提升技能的角色时,他们不会忘记最初的计划是在某个时候转向不同的事情。我觉得这有点道理:人们通常倾向于陷入当前的工作,并说服自己无论做什么都真的很有用。所以你可能会认为先进入再转出是好的,但你可能怀疑自己是否真的会坚持到底。
Yeah, I guess I was suggesting that this would be the thing to do if you couldn't get a job that was alignment-specific already. So they want to join your team, but they're just not quite good enough yet; they need to learn more. Or potentially this only goes so fast that the team can grow, so even though they're good, you just can't hire as quickly as people are coming on board. But suppose you have to make sure that when people go into these roles that we think are currently kind of just neutral but good for improving their skills, that they don't forget that the original plan was at some point to switch to something different. I guess there's a bit of truth: it seems like people in general tend to get stuck in doing what they're doing now and convince themselves that whatever they're doing is actually really useful. So you might think it would be good to go in and then switch out, but you might have some doubts about whether in fact you will follow through on that.
是的,我同意。如果你把那些可能进入 ML 领域的一半人全部转移到深入思考长期问题、如何让事情变好上,你当然会更高兴。那听起来是一个更好的世界。如果你真的信任某人,那似乎也很好。如果有人真的关心长期,问你‘我该做什么?’,一个相当好的选择就是说:‘去做这件短期有益、且与我们预期长期非常重要的领域相邻的事。’多年来一直有争论:让处于机器学习研究前沿、关心长期、警惕安全和对齐问题的人在场,以某种我们尚无法预见的方式是有益的,仅仅是在决策发生的房间里。人们对于这到底有多大用处一直有不同看法。
Yeah, I think that's right. You'd be even happier in the world certainly if you took those half of people who might have gone into ML, you instead move them all into really thinking deeply about the long term, how to make things go well. That sounds like an even better world still. It seems to be pretty good if you really trusted someone. If someone really cared about the long term and you're like, 'What should I do?' it's a reasonably good option to just be like, 'Go do this thing which is good in the short term and adjacent to an area we think is going to be really important over the long term.' There's been this argument over the years that it would just be good in some way that we can't yet anticipate to have people at the cutting edge of machine learning research who are concerned about the long term and alert to safety issues and alert to alignment issues that could have effects on the very long term. And people have kind of gone back and forth on how useful that actually would be, to just be in the room where decisions are getting made.
我突然想到,机器学习社区似乎正在朝着分享你我持有的观点的方向发展,或者很多人开始担心 AI 长期是否会对齐。也许如果你现在特别担心这一点,那可能让你与同龄人不同,但 10 年或 20 年后,随着我们更清楚 AI 实际的样子以及部署时的风险,每个人都会趋同于类似的愿景。
It just occurred to me that it seems like the machine learning community is really moving in the direction of sharing the views that you and I hold, or a lot of people are becoming concerned about whether AI will be aligned in the long term. And it might be that if you're particularly concerned about that now, then maybe that makes you different from your peers right now, but in 10 years time or 20 years time, kind of everyone will have converged on a similar vision as we have a better idea of what AI actually looks like and what the risks are when it's deployed.
是的,我认为这是一个有趣的问题,或者说对这种做法的一个有趣的潜在担忧。我的看法是,存在一些——我不知道你是否称之为价值观差异或深层经验世界观差异——与此相关。我认为,就我们目前正在思考将成为真正问题的问题而言,这些问题是真实存在的这一点会变得更加明显。而且我认为,就我们思考的一些非常长期的问题已经是明显的问题而言,ML 社区对明显的问题或影响当前系统行为的问题非常感兴趣。同样,如果这些问题是真的,随着时间的推移,情况会越来越如此,所以人们会越来越对这些感兴趣。我仍然认为可能存在——问题在于你有多关心让长期变好,与你在多大程度上做你的工作或追求短期有积极影响的事情,或者你充满热情或感兴趣的事情。这种其他的非长期影响——我确实认为,总是会有一些需要做出的决定,或一些不同的选择,或者领域体现了一套价值观。我认为人们的经验观点比他们隐含的价值观变化得更快。我认为如果你说所有真正关心长期的人都不进入这个领域,那么领域的整体方向将持久地不同。
Yeah, I think that's an interesting question, or an interesting possible concern with that kind of approach. I guess my take would be that there are some—I don't know if you would call them values differences or deep empirical worldview differences—that are relevant here. Where I think to the extent that we're currently thinking about problems that are going to become real problems, it's going to be much more obvious there are real problems. And I think that to the extent some of the problems we think about over the very long term are already obviously problems, people in the ML community are very interested in problems that are obviously problems, or problems that are affecting the behavior of systems today. Again, if these problems are real, that's going to become more and more the case over time, and so people become more and more interested in those problems. I still think there are likely to be—there is this question of how much are you interested in making the long-term go well versus how much are you doing your job or pursuing something which has a positive impact over the short term, or that you're passionate about or interested in. This other kind of non-long-term impact—I do think there's just sort of continuously going to be some calls to be made, or some different decisions, or the field embodies some set of values. I think that people's empirical views are changing more than the set of kind of implicit values that they have. I think if you just said everyone who really cares about the long term isn't going into this area, then the overall orientation of the field will persistently be different.
你对 Pushmeet 在节目中提到的特定技术方法,或者 DeepMind 的人在他们的安全博客上写的内容有什么看法吗?我最熟悉 Pushmeet 小组的工作是验证对扰动的鲁棒性,一些更广泛的验证工作,以及一些对抗训练和测试。可能就这三件事。我不知道是否还有其他。我很乐意按顺序讨论这些。
Do you have any views on the particular kind of technical approaches that Pushmeet mentioned in the episode, or that the DeepMind folks have written up on their safety blog? The stuff I'm most familiar with from Pushmeet's group is sort of working on verification for robustness to perturbations, some working on verification more broadly, and some working on adversarial training and testing. Maybe those are like the three things. I don't know if there's something else. I'm happy to go through those in order.
是的,请讲。所以我想我通常对对抗性测试、训练和验证感到非常兴奋。也就是说,我认为有一个非常重要的问题,既在短期重要,可能在长期更重要,就是你有某个系统,你有某个 AI。
Yeah, yeah, go through this. So I guess I'm generally pretty psyched about adversarial testing and training and verification. That is, I think there is this really important problem over both—this is one of those things in the intersection—if it matters over the short term, I think maybe matters even more over the very long term, of like you have some system, you have some AI.
你想把大量工作委托给可能不止一个而是一大堆 AI 系统。如果它们灾难性地失败,那将是真正不可挽回的糟糕。你无法用传统的机器学习训练排除这种情况,因为你只是在一堆你迄今为止生成和经历过的案例上尝试。所以你的训练过程根本没有约束这种在新情况下出现的潜在灾难性失败。所以我们想要有某种东西,我们想要改变机器学习训练过程,使其尊重关于什么构成灾难性失败的信息,然后避免那样做。所以我认为这是一个短期和长期都存在的问题。我认为它在长期很重要。很难说它在长期还是短期更重要,但我非常关心长期。我认为我们对此的主要方法是我只想到的三个:对抗训练和测试、验证,以及某种可解释性或透明度。我只是认为人们熟悉这些技术,变得擅长它们,思考如何将它们应用于更丰富的规范类型,如何应对对抗训练中的根本限制——你必须依赖对手想到一种情况。这种技术的一般工作方式是:你担心你的系统在未来失败,你让一个对手生成一些系统可能失败的情况,然后你在这些情况下运行,看看它是否灾难性地失败。你有一个根本限制:你的对手不会想到所有事情。我认为人们只是获得经验,如何应对这个限制。在某种意义上,验证是对这个限制的回应。也许在两者之间,或者当你……我认为让人们同时思考验证和验证的局限、测试和测试的局限是富有成效的。所以总的来说,我对所有这些感到非常兴奋。
You want to delegate a bunch of work to maybe not just one but a whole bunch of AI systems. If they failed catastrophically, it would be really unrecoverable bad. You can't really rule out that case with traditional ML training because you're just going to try a thing on a bunch of cases that you've generated so far, experienced so far. So you're really not going to be getting your training process isn't at all constraining like this potential catastrophic failure in a new situation that comes up. So we just want to have something, we want to change the ML training process to respect to have some information about what would constitute a catastrophic failure and then not do that. So I think that's like a problem that is common between the short and long term. I think it matters a lot on the long term. It's a little bit hard to say it matters more on the long term or short term, but I care about a lot on long term. I think that the main approaches we have to that are this like the three I only think about are adversarial training and testing, verification, and like sort of interpretability or transparency. I just think people getting familiar with those techniques, becoming good at them, thinking about how you would apply them to richer kinds of specifications, how you grapple with the fundamental limitations in adversarial training where you have to rely on the adversary to think of a kind of case. The way the technique works in general is you're like, I'm concerned about my system failing in the future. I'm going to have an adversary who's going to generate some possible situations under which this system might fail, and then we're going to run on those and see if it fails catastrophically. You have this fundamental limitation where your adversary isn't going to think of everything. And I think people just getting experience with how do we grapple with that limitation. In some sense, verification is like a response to that limitation. And maybe the space between, or when you're, I think it's productive to have people thinking about both of verification and the limits of verification, and testing and limits of testing. So overall, I'm pretty excited about all that.
你同意保罗的总体乐观态度吗?
Do you share Paul's general optimism?
我不确切知道他在数量上有多乐观。我的猜测是我没那么乐观,意思是,嗯,有百分之几十的概率我们会搞砸,失去未来大部分的价值,而听他的感觉并不是这样,这不是我对他整体感觉的印象。但很难在氛围和实际乐观程度之间进行转换。
I don't know quantitatively exactly how optimistic he is. My guess would be that I'm less optimistic, in the sense that I'm like, well, there's like tens of percent chance that we'll mess this up and lose the majority of the value of the future, whereas that's not listening to him, it's like not the overall sense I get of where he's at. But it's a little bit hard to know how to translate between a vibe and an actual level of optimism.
是的,这很有趣。有人可能认为有 20%的概率我们会彻底毁灭一切,但他们的性格使他们表现得好像事情会顺利。
Yeah, it is interesting. Someone can think there's a 20% chance that we'll totally destroy everything but still they just kind of disposition so they come across as well things go well.
在从事存在风险和全球灾难风险,特别是 AI 风险的人中,存在一种权衡:一方面不想做别人不同意或不热衷的事情,另一方面不想让领域过于保守,以至于除非有共识,否则不做任何实验。你认为人们是否过于倾向于犯单边主义错误,或者尝试得不够多?
Among people working on existential risks and global catastrophic risks, and I guess AI in particular, there's kind of this trade-off between not wanting to do things that other people disagree with or are unenthusiastic about, and at the same time not wanting to have a field that's so conservative that there are no experiments done unless there's a consensus behind them. Do you think people are too inclined to make unilateralist curses types of mistakes, or not trying things enough?
我认为我的答案可能因领域而异。作为参考,我认为你想要遵循的政策是更新这样一个事实:没有其他人想做这件事,然后认真对待,在决定是否要做之前深入参与。理想情况下,这需要与做出决定的人交流,理解他们的出发点。我认为我没有很强的总体感觉,我们更可能犯哪种错误。我有点预期世界系统性地会过多地做那种可以单方面完成的事情,所以它就被做了。在这个领域的背景下,我不知道是否有那么多。我想我对两种失败模式都不太担心。也许我对人们目前的状况感觉还不错。
I think my answer to this probably varies depending on the area. For reference, I think the sort of policy you want to be following is kind of the update on the fact that no one else wanted to do this thing, and then take that really seriously, engage with it a lot before deciding whether you want to do it. And that ideally that's going to involve engaging with the people who've made that decision to understand where they're coming from. I think I don't have a very strong general sense of whether we're more likely to make one kind of mistake or the other. I think I sort of expect the world systematically to make too much of the sort of a thing can be done unilaterally so it gets done. In the context of this field, I don't know if there are as many. I guess I don't feel super concerned about either failure mode. Maybe I don't feel that bad about where people are at.
是的,我从 AI 政策和策略人士那里得到的总体感觉是,他们非常谨慎,对自己说的话和做的事非常谨慎。我想这是经过深思熟虑的决定,但我有时确实怀疑他们是否过于倾向于不充分表达自己的观点。
Yeah, the vibe I get in general from the AI policy and strategy people is that they are pretty cautious, pretty cautious about what they say and what they do. I guess that's been a deliberate decision, but I do sometimes wonder whether they've swung too far in favor of not speaking out enough about their views.
我想确实有人采取了,而且人们做的事情多种多样,我想这就是整个问题。而且肯定有人采取非常谨慎的观点,我认为他们有时会被排除在公共讨论之外,因为他们不倾向于发言,这有时可能是一种损失。
I guess there is certainly people who have taken, and there's a diversity of what people do, which I guess is the whole problem. And I guess they're definitely people who take a very cautious perspective, and I think that they sometimes get a bit cut out of the public discussion because they're just not inclined to speak out, which can be a loss at times.
是的,确实如此。看起来你有一个真正的问题:你认为你产生影响的部分渠道是交流你的观点,但然后又非常纠结,或者强烈认为不应该因为单边主义担忧而交流观点。
Yeah, definitely. It seems like you have a real problem: you think part of your channel for impact is communicating about your views, but then are very hung up on the or take a strong shouldn't communicate views because of unilateralist concerns.
是的,所以总的来说,对于单边主义担忧这一类,我最不同情的可能是这样一种干预:认真讨论可能有哪些机制,以及如果对齐很难或 AI 进展很快,我们该如何应对。这可能是我最不同情的地方,但我认为这类讨论的成本效益看起来相当不错。也就是说,你默认在大多数问题上,即使你不采取行动,也要表达你的真实观点,或者至少在有益的协作认知工作范围内,比如梳理思考应该做什么、我们如何应对、会发生什么,愿意作为社区参与这项工作,而不是人们私下思考。并且可能注意一下,嗯,你不想显得煽动性,你不想让人们非常不安,你可以合理地处理。是的,我想再说一次,这并不真的是我的观点,我不在乎。更像是,我有点矛盾,不认为明显存在一个方向的大错误。
Yeah, so I guess in general on the family of unilateralist concerns, I'm least sympathetic to probably one intervention is talk seriously about what kinds of mechanisms might be in place and how we might respond if it turned out the alignment was hard or if AI progress was rapid. That's probably the place I'm least overall sympathetic, but I think that the cost-benefit looks pretty good on those kinds of discussion. Kind of saying you default towards just on most issues, even if you're not taking action, express your true views, or at least to the extent there's useful collaborative cognitive work of flushing thinking about what should be done, or how would we respond, or what would happen, being willing to engage in that work as a community rather than people thinking in private. And maybe taking some care to be like, well, you don't want to look inflammatory stuff, you don't want to get people really upset, you can be reasonable about it. Yeah, I guess again, it's not really like my view is so much I don't care one way or the other. It's more like I'm kind of ambivalent and don't think it's obvious that there's a big error one way or the other.
好了,我们来谈谈完全不同的事情。你最近写了一篇文章,关于为什么从做有害事情的公司撤资,适度地,实际上可能是改善世界的一个相当有效的方式。
Alright, let's talk about something pretty different, which is you recently wrote a post about why divesting from companies that do harmful things could, in moderation, actually be quite an effective way to improve the world.
这与大多数在有效利他主义框架下研究过这个问题的人得出的结论相反,他们认为这其实没什么用,因为如果你卖掉一家公司的股票或不借钱给一家公司,别人就会取代你的位置,你实际上并没有带来任何改变。你能解释一下从有害公司(比如烟草公司)撤资可能有用的机制吗?
Kind of in contrast to what most people who've looked into that as part of the rubric of effective altruism have tended to conclude, which is that it's actually not that useful because if you sell a share in a company or don't lend money to a company, then someone else would just take your place and you haven't really made any difference. Yeah, you want to explain the mechanism by which divesting from harmful companies, like cigarette companies for example, could be useful?
是的,我认为有两点需要先说清楚。第一,我主要考虑的是成本与收益的比率。对于某些公司,你可能会处于这样一种情况:撤资效果相对较小,但成本也非常低。所以总的来说,我认为撤资的第一个微小增量几乎总是免费的。成本是二阶的,而收益是一阶的,因此至少撤资那一点几乎总是值得的。这是图景的第一部分:这主要是一个成本非常低的故事,而不是收益很大。如果成本非常低,那么主要问题就是需要做分析和处理物流。在这种情况下,我认为如果既有人做研究,又有人实际创建基金,那么确实有可能大幅降低成本。比如我可以想象自己说,好吧,我会把我财富的 0.1%投入某个基金,这个基金大致是市场中性,做空所有我真正不喜欢其活动的公司,同时做多那些最相关的公司。所以这是一点。但可能只是成本和收益都很小,对任何个人投资者来说都不是什么大事,也许不值得过多考虑。但如果有人愿意推出一个可以大规模推广的产品,每个人都能很快或很容易地购买这个基金,那么他们可能会这么做。
Yeah, I think there's two important things to say up front. One is that I was mostly thinking about the ratio of cost to benefits. So you can end up for some companies in a regime where divestment has relatively little effect but is also quite cheap. So in general, I think the first epsilon of divestment will tend to be literally free. The cost is second order in terms of how far you divest and the benefits are first order, so it's almost always going to be worth it to divest at least by that epsilon. That's the first part of the picture: this can be mostly a story about the costs being very very low rather than benefits being large. And if the costs are very very low, then it's mostly an issue of having to do the analysis and having to deal with logistics. In which case, I think it is plausible that one could imagine really getting those costs down if someone both did the research and actually produced a fund. Like I could imagine me personally being like, sure, I will put 0.1% of my wealth in some fund that's just roughly market neutral, shorts all the companies I really don't like the activities of, and is long those most correlated companies. So that's one thing. But it may just be about the cost and benefits both being small, such that it's not going to be a big deal for any individual investor, maybe not worth thinking that much about it. But if someone was willing to produce a product that could be scaled a lot and everyone could very quickly or easily buy the fund, then they might do that.
也许第二点是关于它实际上如何可能,或者为什么它不会完全被抵消。我认为大致的机制是:当我退出一个公司时,假设我关心石油,我从生产石油的公司撤资。随着我撤资越多,它产生好处的方式是提高石油公司投资的预期回报。所以担忧是其他投资者会买入并继续向石油公司投入更多资金,直到预期回报下降到市场回报水平,否则为什么不继续投入呢?这个简化图景忽略的是石油行业存在特质风险。也就是说,随着石油在我投资组合中占比越来越大,我投资组合的波动性越来越多地不是由市场整体(多个行业的组合)驱动,而是由石油本身的波动性驱动。所以如果我试图超配 10%的石油,如果有很多投资,人们不得不超配 10%的石油,他们实际上会显著增加边际石油投资的风险,因此他们为抵消该风险所要求的回报也会上升。所以我认为有两件事。一是这确实需要做某种……嗯,这有点取决于你认为投资者有多理性。我认为在某种意义上,投资故事中的悲观情绪已经依赖于理性投资者,所以我认为更合理的做法是深入探讨理性投资者会如何反应并进行那些计算。所以这可能是我的观点。我认为当悲观情绪来自这种理性经济人模型时,研究这个问题是特别合理的。一旦你这样做,就有两个问题。一是定量上我们谈论的效果有多大,这是我在博客文章中试图分析的,我有点惊讶于当我稍微思考时它们竟然这么大。也许第二个观察是实际上存在一种抵消,大致来说,如果石油行业没有特质风险或特质风险非常低,你的投资几乎会被完全抵消,但同时它对你几乎没有成本,因为该行业几乎没有超额回报,因为这些回报应该与特质风险挂钩。所以最终成本效益并不依赖于……有一个参数控制你的投资会被抵消多少,但成本效益实际上并不依赖,或者说成本与收益的比率不依赖那个参数,因为它对成本和收益的影响是相同的。所以它影响的是组织这个基金的整体好处有多大,会被抵消多少,但不影响它对个人投资者的吸引力。所以我认为我们可以深入细节,讨论投资多少是合理的。我认为完全投资甚至做空你不喜欢的行业 100%或 200%可能常常是合理的。它会变得更好。一个可能特别具有成本效益的投资例子是,假设有两家公司生产非常相似的产品,因此相关性很高,比如两家都生产家禽的公司,其中一家动物福利实践明显更差。你可能会认为家禽养殖业的大部分风险会被这两家公司平等地经历。所以如果你做多你喜欢的公司,做空你不喜欢的公司,那风险相对较小。我的意思是,它仍然有这些公司特有的特质风险,并且有一个复杂的分析,但相对于你对两家公司资本可用性的影响程度,你最终可以承担相对较小的风险。
Maybe the second thing, in terms of how it could actually be possible or why it isn't literally completely offset. I think the rough mechanism is: when I get out of a company, let's suppose I care about oil and I divest from companies that are producing oil. That increases, as I divest more, the whole way that does good is by increasing the expected returns to investment in oil companies. So the concern is other investors will just buy and continue putting more money into oil companies until the expected returns have fallen to market returns, because otherwise why not just keep putting more money in? The thing that simplified picture misses is that there's idiosyncratic risk in the oil industry. Namely, as oil becomes a larger and larger part of my portfolio, more and more of the volatility of my portfolio is driven not by what is overall going on in the market, which is the composite of many sectors, but just volatility in oil in particular. So as I try and go overweight 10% oil, if there's a lot of investment and people had to go overweight 10% oil, they would actually be significantly increasing the riskiness of marginal oil investments, and so the returns that they would demand in order to offset that risk would also go up. So I think there's sort of two things. One is it actually does require doing a kind of... well, it depends a little bit on how rational you believe investors are. I think in some sense the investment story, the pessimism already relied on rational investors, so I think it's maybe more reasonable to say let's actually dig in and see how rational investors would respond and do those calculations. So that's maybe my perspective. I think it's unusually reasonable to look into that when pessimism is coming from this homo economicus model. Once you're doing that, there are two questions. One is just quantitatively how large were the effects we're talking about, and that's something I tried to run through in this blog post, and I was kind of surprised how large they were when I thought about it a little bit. Maybe the second observation is that actually there's this cancellation that occurs where roughly speaking, if the oil industry has no idiosyncratic risk or has very low idiosyncratic risk, your investment will get almost entirely offset, but at the same time it has almost no cost to you because the industry had almost no excess returns, because those returns should be tied to the idiosyncratic risk. So you end up with actually the cost effectiveness doesn't depend... there's this parameter which governs how much your investment is going to get offset, but cost effectiveness doesn't actually depend, or the ratio between costs and benefits doesn't depend on that parameter because it affects both costs and benefits equally. So it affects what is the overall upside to organizing this fund, how much will it get offset, but it doesn't affect how attractive is it for an individual investor. So I think we could go into details, talk about how much it makes sense to invest. I think it might often make sense to invest completely or maybe even go 100 or 200% short in industries you don't like. It's going to become better. An example of investment that might be particularly cost-effective is suppose there's two companies who are producing very similar products and so very very correlated, like maybe two companies that both produce poultry, and one of them has substantially worse animal welfare practices. You might think there's a lot of the risk in animal culture in general that is going to be experienced equally by those two companies. So if you have a long position in the company you like and a short position in the company you dislike, that has relatively little risk. I mean, it still has idiosyncratic risk specific to those companies, and there's a complicated analysis there, but you can end up with relatively little risk compared to how much effect you have on capital availability for the two companies.
我想我们实际上没有讨论这个机制如何导致世界上更少坏事发生。这里我们真的只是在讨论为什么我对怀疑论持怀疑态度。
I guess we didn't actually talk about the mechanism by which this causes less bad stuff to happen in the world. Here we're really just talking about why I'm skeptical of the skepticism.
是的,是的,我们先把那个放在一边。所以用非常简单的语言解释一下。之前的想法是,如果你卖掉一家公司的股票,那么另一个不在乎养鸡或产油道德问题的人就会以同样的价格买走,股价不会……
Yeah, yeah, let's set that aside for a minute. So just explain this in really simple language. So the previous thinking has been that if you sell shares in a company, then someone else who just doesn't care about the moral issues to do with raising chickens or producing oil, they're just going to sweep in and buy it at the same price and the share price won't be...
公司的借款能力不会真正改变,但被忽略的是,如果相当多的人,甚至只是少数人停止购买石油股票,那么那些不在乎道德问题的富裕投资基金,他们也不会去购买更多这些化石燃料公司的股票,或者他们的意愿是有限的,因为他们希望在全球所有不同资产中实现多元化。为了弥补你我都不愿持有这些股票而需要额外购买石油股票,他们必须降低多元化程度,这对他们来说没有吸引力。所以,如果一群人,甚至只是我们做空或卖出这些股票,实际上可能会略微压低它们的价格,因为人们需要因多元化降低而获得补偿,即更低的价格才能让购买更有吸引力。这是其一。另外,虽然这种影响在整体上可能很小,但同样,卖出那些公司的头几股股票对你来说并不重要,因为你本来就不那么在意持有这些特定公司。无论如何,这就像你可以稍微降低你的多元化程度。是的,只是卖出你投资组合中持有的这些公司的一小部分股票,几乎不花你什么成本。所以即使收益很小,成本可能更小,因为这没什么大不了的。因此,收益与成本的比率可能相当大,即使这不是对世界产生影响的最佳方式。
Changed or the amount of money that the company can borrow won't really be changed, but the thing that misses is that if a decent number of people or even a small number of people stop buying oil stocks, then say rich investment funds that don't care about the moral issues, for them to go and buy even more of these fossil fuel companies, they don't want to do that, or their willingness to do that isn't unlimited because they want to be diversified across all the different assets in the world. In order to buy extra oil shares to make up for the fact that you and I don't want to own them, they have to reduce the diversification that they have, which is unappealing to them. So if a bunch of people or even just we short or sell these shares, it actually probably will suppress their price a little bit because people will have to be compensated for the reduced diversification with a lower price to make it more appealing to buy. Okay, that's one thing. Also, while that effect might be pretty small in the scheme of things, it's also the case that just selling those first few shares of those companies wasn't that important to you to own those specific companies. Anyway, it's just like you can slightly reduce your diversification. Yeah, just sell tiny amounts of these companies that you owned in your portfolio costs you practically nothing. So even though the benefit is quite small, the cost could potentially be even smaller because it just doesn't matter that much. And so then the ratio of benefits to costs could be pretty large, even if this is not the best way to have an impact in the world.
是的,没错。我认为如果你想考虑总影响,也许可以合理地想象将这种行为扩展到大量投资者。很多影响在相关范围内大致是线性的。我认为总影响并不糟糕;它们看起来不太好,但如果你想象一个变化,即很大一部分人撤资,我认为这会显著减少开采的石油量或圈养鸡的数量,尤其是在某些情况下……我认为石油案例在这方面可能比鸡肉案例稍微不利一些,在鸡肉案例中,你真的可以想象稍微转向更有利于动物福利的做法。是的,你可以想象稍微转向更有利于动物福利的做法,或者转向不同种类的肉类等等。总影响可能不是很大。但总影响可能仍然足够大,以至于值得真正理顺物流,使其非常容易操作,因为从投资者的角度来看,除了操作的麻烦之外,我认为实际上第一个单位绝对是非常划算的。
Yeah, that's right. I think if you want to think about what the total impact is, it's maybe reasonable to imagine scaling this up to large numbers of investors doing it. A lot of the effects are going to be roughly linear in the relevant range. I think the total impacts are not that bad; they don't look great, but they look like if you imagine a change where large fractions of people divest, I think it would meaningfully decrease the amount of oil that gets extracted or the number of chickens raised in captivity, especially in the case where you have... I think maybe the oil case is a little bit unfavorable in this way compared to the chicken case, where you could really imagine slightly shifting towards more practices that are better for animal welfare. That's yeah, you could imagine slightly shifting towards practices better for animal welfare or towards different kinds of meat and so on. The total effect probably not that big. The total effect may still be large enough to justify really getting the logistics sorted out so that it's very easy to do, because from the investor's perspective, other than the hassle of doing it, I think it's actually pretty definitely the first unit is a very, very good deal.
那么,你能不能直接买入一个不持有石油公司或不持有动物农业公司的投资基金呢?这似乎是第一步,做起来相当直接。
Well, could you just buy into an investment fund that doesn't own oil companies or doesn't own your animal agriculture companies? That seems like the first pass, that's pretty straightforward to do.
是的,但这需要一些思考。当我买入一个基金时,有很多因素限制我的选择,如果现在再加上这个额外的限制,那就有点烦人了。所以这可能是一个原因,让人觉得不太值得。
Yeah, it involves some thinking though. So when I buy a fund, there's a bunch of things that constrain my choices, and it's kind of annoying if now I have this extra constraint on top of those. And so that might be a reason like it's not quite worth it.
是的,我的意思是,即使那个基金的管理费稍微提高了一点,对吧?就像先锋集团会给我提供 0.04%的费率,而现在我不得不为这个新东西支付 0.1%的费率,那就不太好了。
Yeah, I mean even if you're slightly raised management fees on that fund, right? It's like Vanguard's going to offer me some 0.04%, yeah yeah, and now I have to pay like 0.1% on this new thing, that's no good.
是的,所以我通常会想象我的基准实施方案是一个做空你关心的相关特定公司的基金,并且可能还提供对冲机制。你知道,卖出这些公司股票不好的原因是我们失去了多元化,所以他们可以尝试通过同一组合中的措施来抵消这些成本。我很想看到有一个针对关心动物福利等人士的最优撤资基金,它主要持有那些动物福利影响最差的公司的很大空头头寸,然后还构建一个投资组合,尽可能捕捉这些头寸本应给你的投资组合带来的多元化收益。投资这个基金的成本可以很低,你可以把它叠加在你原本会做的任何其他投资之上,然后拿出你资金的 0.1%或 1%投入这个基金。平均而言,这个基金将赚取零美元,并且会有一些风险。对你的成本只是这个平均不赚钱基金的风险,但如果你投入的资金不多,风险就没那么糟糕。
Yeah, so I would normally imagine my baseline implementation would be a fund that shorts the relevant particular companies you care about and maybe also opens up the offsetting right. So you know, the reason it was bad to sell these companies was because we were losing diversification, and so they can try and do things to offset those costs as part of the same bundle. I'd be very interested in just seeing there is the optimal divestment fund for people who care about animal welfare or whatever, that just holds mostly these really large short positions in the companies that have the worst animal welfare effects, and then also constructs a portfolio to as much as possible capture the diversification benefits that those would have added to your portfolio. The cost of investing in that can be pretty low, and you can just then put that on top of do whatever else you would have done in investing, then take like 0.1% of your money or whatever, 1% of your money, and put it in this fund. On average, that fund is going to make zero dollars, and it's going to have some risks. The cost to you is just the risk of this fund that on average is making no money, but it could be relatively if it's if you're not investing that much of your money in it, the risk is just not that bad.
这在多大程度上类似于不去邪恶公司工作是有用的?
To what extent, if at all, is this analogous to it being useful to not go and work at an evil company?
是的,我认为这相当类似。有很多定量参数。如果你从某种经济学视角来看,它们在结构上非常非常相似。我们正在讨论的风险问题,对于确定一些相关弹性很重要,这与在问题行业工作的情况下的类似讨论有很大不同。但我认为整体情况是相似的,如果你不在那个行业工作,整体上会发生的是该行业的价格或工资略有上升,从而吸引更多人进入。我们只需要讨论工资上升多少。我想到的一点是,如果我们考虑动物农业,这也类似于关于道德消费的讨论。我认为这实际上是撤资的一个很好的比较点,你可以说我想减少动物产品的消费,以减少生产的动物数量。然后你会有一个非常类似的关于相对弹性的讨论。一种思考方式是,如果你将需求减少 1%,你将劳动力减少 1%,并将资本可用性减少 1%。如果你做了所有这些事情,那么在关于自然资源如何运作等的一些假设下,你大致会将总产量减少 1%。因此,这 1%的减少的功劳在某种程度上在供给和需求方面的各种因素之间分配,而弹性决定了如何分配。但我认为这不是 100%。
Yeah, I think it is fairly analogous. There's a bunch of quantitative parameters. If you take a certain economics perspective, they're very, very structurally analogous. The risk discussion we're having about risk, which is important to determining some of the relevant elasticities, is quite different from the analogous discussion in the case of working in a problematic industry. But I think the overall thing is kind of similar, where if you don't work in that industry, overall what happens is prices go up or wages go up a little bit in the industry and induces more people to enter. We just have to talk about how much do wages go up. One thing I sort of think about is if we consider say animal agriculture, it's also kind of analogous to discussion with ethical consumption. I think that's actually a really good comparison point for divestment, where you could say I want to consume fewer animal products in order to decrease the number of animals produced. And then you have a very similar discussion about what are the relative elasticities. One way you could think about it is if you decreased demand by 1%, you decreased labor force by 1%, and you decreased availability of capital by 1%. If you did all of those things, then you would kind of decrease the total amount produced by 1% roughly under some assumptions about how natural resources work and so on. And so the credit for that 1% decrease is somehow divided up across the various factors on the supply side and demand side, and elasticities determine how it's divided up. But I think it is not like 100%.
消费或 100%的劳动,我认为所有这些因素都在非微不足道的程度上参与其中。与道德消费相比,我认为它实际上看起来相当不错。在相当合理的假设下,你从撤资中获得的性价比更高。我还没有非常仔细地做过这个分析,我认为这将是一件非常有趣的事情,如果有人想成立一个动物福利撤资基金,这将是一个很好的动机。但在相当合理的假设下,你从撤资中获得的性价比远高于消费选择。可能你仍然希望消费和投资相对于你的总消费模式来说较小,所以它不会取代你的道德消费选择。但如果道德消费是个好主意,那么至少完全撤资,甚至 10 倍杠杆做空,比如你本来要买 1 美元的畜牧业公司股票,现在你卖出 10 美元。我认为如果认为道德消费是个好主意,那么类似的事情是可以证明合理的。
Consumption or 100% of the labor, I think all those factors are participating to a non-trivial extent. In comparison to ethical consumption, I think it actually looks reasonably good. Under pretty plausible assumptions, you're getting more bang for your buck from divesting. I haven't done this analysis really carefully, and I think it would be a really interesting thing to do, and would be a good motivation if one wanted to put together an Animal Welfare Divestment Fund. But under pretty plausible assumptions, you're getting a lot more bang for your buck from divestment than from consumption choices. Probably you'd still want the consumption and investment thing to be relatively small compared to your total consumption pattern, so it wouldn't be replacing your ethical consumption choices. But if ethical consumption was a good idea, then also at least totally divesting, and maybe even 10x leveraged short positions, like when you would have bought $1 of animal agriculture companies, instead you sell $10. I think stuff like that could be justified if you thought that ethical consumption was a good idea.
你想对那些持怀疑态度的人简要说明一下,出售一家公司的股票或债券是如何减少该公司产出的吗?
Do you want to sketch out briefly for those who are skeptical how it is that selling shares in a company or selling bonds in a company reduces the output of that company?
是的,大致如此。我认为债券的情况更容易思考。我认为它们可能差不多。我们来谈谈债券的情况。假设公司,比如泰森,想筹集一美元。他们去找投资者说,现在给我们一美元,十年后我们会给你一定数量的钱,假设我们仍然有偿付能力。这就是他们的推销。他们向人们出售这些纸片,就像借条一样。这些借条的价格由投资者之间的供需决定。所以当你做空债券时,有人来找泰森想借给他们一美元,而你说,不要借给他们一美元,而是借给我一美元,无论他们偿还给债券持有人的是什么,我都会偿还给你。然后他们说,好吧,我借给你和借给实际公司一样高兴。现在公司少了一美元。现在公司说,好吧,如果我们想生产这额外的边际鸡肉,我们仍然需要筹集那一美元。所以公司去尝试筹集那一美元,但他们已经用掉了一个愿意的买家。所以他们需要找到另一个买家,一个愿意借给他们这一美元的人,而那个人会稍微不那么兴奋,因为同样,他们的投资组合会更多地偏向这家公司,所以他们更担心这家公司倒闭的风险。大致来说,这就是机制。
Yeah, roughly. I think the bond case is a little bit simpler to think about. I think they're probably about the same. Let's talk about the bond case. So the company, I don't know, Tyson wants to raise a dollar. They go out to investors and say, give us a dollar now and we'll give you some amount of money 10 years from now, assuming we're still solvent. So that's their pitch. They're selling these pieces of paper to people which are like IOUs. The price of those IOUs is set by supply and demand amongst investors. So what happens when you short the bond is someone came to Tyson and wanted to loan them a dollar, and you're saying, don't loan them a dollar, instead loan me a dollar, and whatever it is that they pay back to their bond holders, I'll pay it back to you instead. And they're like, fine, I'm just as happy to lend to you as I was to lend to the actual company. Now the company has one less dollar. Now the company says, okay, we still need to raise that dollar if we want to produce this additional marginal chicken. So now the company goes and tries to raise the dollar, but they've used up one of the willing buyers. So now they need to find another buyer, someone who's willing to loan them this dollar, and that person is going to be a little bit less excited, because again, they're moving their portfolio to be a little bit more overweight in this company, and so they're a little bit more scared about the risk of this company going under. So roughly speaking, that's the mechanism.
是的,我觉得有道理。你可以想象,如果很多人不愿意借钱给一家公司,或者数量很大,那么就会推高他们的借贷成本,公司就会萎缩,因为他们必须支付更高的利率,无法获得那么多资本。在直觉层面上有点道理。其中一些有点技术性,所以我们会附上你写的博客文章的链接,里面有所有的方程式,解释了你如何解决这个问题,并试图估计收益和成本的大小。
Yeah, I think it makes sense. You imagine if a lot of people weren't willing to lend money to a company, or a significant number, then that drives up their borrowing cost, and so the company shrinks because they have to pay higher interest rates, they can't get as much capital. Kind of makes sense on an intuitive level. Some of this gets a little bit technical, so we'll stick up a link to the blog post that you wrote with all the equations and explaining how you work this through and try to estimate the size of the benefits and the costs.
我担心这对人们来说不是最仔细或最清晰的分析。我对它感兴趣,并认为在某个时候我们会发布一个更仔细的版本。这对我来说只是一个有趣的练习。
I'm concerned it's not the most careful or clear analysis to people. I think I'm interested in it and think at some point we'll have a more careful version that we put up. It was just a fun exercise for me.
嗯,你提出了一些我在其他地方没有见过的观点,这些观点实际上可能会改变结论。所以这可能是人们需要接受的最重要的事情。
Well, you make some points that I haven't seen anywhere else, and that actually might shift the conclusion. So that seems like probably the most important thing that people need to take on board.
如果你真的建立了一个合理构建且成本效益高的撤资基金,那对我来说会非常有趣。那会很酷。另外,抱歉,我说它比道德消费更有利。我想强调的一点是,工作之所以能完成,只是因为这些边际变化非常有效。所以这非常类似于素食主义的第一步,如果你只在真正边际的情况下停止吃肉,那比完全做到有更高的性价比。我认为这里也是一样。它不会与停止吃肉的第一步竞争;它会与做到最后的最后一点竞争。
It would be super interesting to me if you actually ended up with the divestment fund that was reasonably constructed and cost effective. That would be kind of cool. Also, sorry, I said that it compared favorably to ethical consumption. I think one thing I want to stress there is that the way the work is getting done is just because of these changes on the margin being very effective. So it's very similar to for vegetarianism being the first, if you just stop eating meat in cases where it was really marginal, that has a lot more bang for your buck than if you go all the way. I think that's the same thing here. It's not going to be competitive with the first unit of stopping eating meat; it's going to be competitive with going all the way to the last bits.
如果有效利他主义者多年来一直说撤资有点浪费时间,结果发现我们对此相当错误,那会有点尴尬。将不得不认错。但我想这也很好,我们在更新我们的观点,所以我们不只是固守教条立场。
It's going to be a little bit embarrassing if effective altruists have been saying divestment is kind of a waste of time for all these years, and it turns out that we're pretty wrong about that. Going to have to eat humble pie. But I suppose that also looks good that we're updating our views, so we're not just stuck with dogmatic positions.
我认为我们也很可能最终会达成某种妥协。我们看到的影响比人们推销时经常隐含假设的要小得多,但这也是一个合理的事情,也许我们不应该对它如此悲观。
I think we'd also most likely end up with some kind of compromise. We look at the impacts are a lot smaller than people were often implicitly assuming when they were pitching this, but also it is a reasonable thing to do, and maybe we shouldn't have been quite so down on it.
而且成本很多时候可以忽略不计。这实际上是一个社会问题,成本只是人们认为他们应该这样做,因此这是一个人们可以做出的合理改变,几乎不花费他们任何东西。这是一个特别合理的事情,可以倡导人们去做。
And the costs are kind of negligible a lot of the time. It is really a social thing of the cost is just people believing they should do it, and therefore it's kind of a reasonable change people can make that costs them almost nothing. It's a particularly reasonable thing to advocate for people to do.
让我们花点时间谈谈 s-风险。一些听众会知道,但有些人不知道。s-风险这个术语是人们用来描述可能的未来情景的,这些情景只是中性的,比如人类灭绝然后什么都没有,或者不太好,比如人类还在,但我们没有让世界变得尽可能好,而是存在天文数字级别的坏事。我想这里的's'代表痛苦,因为很多人倾向于担心未来可能包含很多痛苦。但它也可以包括任何未来,从某种意义上说,很多事情正在发生,但也包含很多坏事。
Let's talk about s-risks for a minute. Some listeners will know, but some people won't. The term s-risks is what people have settled on to describe possible future scenarios that are just kind of neutral, where humans go extinct and then there's nothing, or not very good, where humans stick around but then we just don't make the world as good as it could be, but rather worlds where there are astronomical levels of bad things. I guess 's' in this case stands for suffering, because a lot of people tend to be concerned that the future might contain a lot of suffering. But it could also include any future that is large in the sense that a lot of stuff is going on, but it also contains a lot of bad stuff in it.
人们担心这种情况可能发生的一些方式涉及不共享我们目标的 AI。你总体上如何看待 s-risk 作为一个需要解决的问题?
Some of the ways people worry this could happen involve AI that doesn't share our goals. What's your overall take on s-risk as a problem to work on?
是的,我认为我最好的猜测是,如果你进入宇宙并优化它以实现美好,那么所交付的美好总量与如果你优化它以实现糟糕所交付的糟糕总量是相称的。我认为,只要人们持有这种经验观点——也许是关于善恶本质的经验和道德观点的某种组合——那么 s-risk 就不特别令人担忧,因为人们更有可能在优化宇宙以实现美好。宇宙中的更多事物被优化为完全符合 Paul 想要的,而不是完全符合 Paul 不想要的。这是我最好的猜测观点,所以在这个观点上,我认为这不是一个大问题。我确实有相当大的道德不确定性。我通常处理道德不确定性的方式是,即使期望上很难比较这些非常不同的道德观点下的结果——这是那种由于奇怪的跨理论效用比较而难以比较的案例之一——我通常会这样想:如果我对那些可能糟糕总量远大于可能美好总量的观点赋予合理的概率,那么我应该对减少 s-risk 给予合理的优先级或合理的兴趣。这就是我的立场。我认为这不太可能,但存在一些合理的经验和道德观点组合,在这些观点上它们非常重要。这是我的出发点:我不会投入太多,因为我不觉得那个观点特别有吸引力;它不会是我总关注中的很大一部分,但它值得一些关注,因为它是一个合理的观点。
Yes, I think my best guess is that if you go out into the universe and optimize it for things being good, the total level of goodness delivered is commensurate with the total amount of badness that would be delivered if you optimized it for things being bad. I think that to the extent one has that empirical view—maybe some combination of empirical and moral views about the nature of what is good and what is bad—then s-risks are not particularly concerning, because people are so much more likely to be optimizing the universe for good. So much more of the stuff in the universe is optimized for exactly what Paul wants rather than exactly what Paul doesn't want. That's my best guess view, so on that view, I think this is not a big concern. I do have considerable moral uncertainty. The way I would approach moral uncertainty in general is to say that even if in expectation it's hard to compare outcomes across these very different moral views—this is one of those cases where the comparison is difficult due to weird inter-theoretic utility comparisons—the way I would normally think about this is to say I should put reasonable priority or reasonable interest in reducing s-risks if I put a reasonable probability on views where the total amount of possible badness is much, much larger than the total amount of possible goodness. That's kind of where I'm at. I think it's not likely, but plausible combinations of empirical and moral views exist on which they are very important. That's my starting point: I'm not going to put much into it because I don't find that perspective particularly appealing; it won't be a large fraction of my total concern, but it deserves some concern because it is a plausible perspective.
天真的看法可能是:我们为什么要担心这些场景?它们看起来很古怪。为什么会有人着手用真正糟糕的东西填满宇宙?这似乎是一件非常奇怪的事情。一旦你达到足够复杂的水平,可以出去殖民太空并创造天文数字的东西,为什么要用糟糕的东西填满它?这有一定道理,但随后人们会思考可能发生这种情况的场景,可能涉及不同群体之间的冲突,其中一个群体威胁另一个群体说他们要做什么坏事,然后他们真的做了,或者可能我们没有意识到我们在创造一些糟糕的东西。所以你可能会创造一些包含很多美好但也有不少糟糕的东西,然后你出去传播它,而我们没有意识到作为副作用我们正在创造一堆痛苦或其他负面价值。你认为这些场景的可能性有多大?
The naive take might be: why would we worry about these scenarios? They seem outlandish. Why would anyone set out to fill the universe with things that are really bad? That seems like a very odd thing to do. Once you're at the level of sophistication where you can go out and colonize space and create astronomical amounts of stuff, why are you filling it with stuff that's bad? There's something to be said for that, but then people think about scenarios where this might happen, which might involve conflicts between different groups where one threatens the other that they're going to do something bad, and then they follow through, or potentially where we don't realize we're creating something that's bad. So you might create something that has a lot of good in it but also a bunch of bad, and you go out and spread that, and we just don't realize that as a side effect we're creating a bunch of suffering or some other disvalue. How plausible do you think any of these scenarios are?
我认为到目前为止对我来说最合理的模型是这种冲突威胁和兑现威胁的模型,而不仅仅是道德错误。我认为很难犯足够极端的道德错误,而且可能存在与兑现威胁相结合的道德错误,但我认为很难让我从中得到的风险大于我几乎完全搞反什么是好的风险。这并非完全不可能——由于各种原因,它比击中空间中的某个随机点更可能——但我认为它仍然是我总关注中的少数。大部分来自有人想要毁灭,因为他们想要拥有毁灭的威胁,或者毁灭价值因为有人想要毁灭价值。那是我主要担心的。它有多可能?我猜你认为它是可以想象的但相当不可能,所以你会给予一点关注,但它不会是一个重点。这是底线吗?
I guess the one that seems by far most plausible to me is this conflict threats and following through on threats model, not just moral error potentially. I think it's hard to make a sufficiently extreme moral error, and there might be moral error that combines with threats that get followed through on, but I think it's hard for me to get the risk from that larger than the risk of me getting what is good very nearly exactly backwards. That's not totally impossible—it's more likely than hitting some random point in the space for a wide variety of reasons—but I think it's still a minority of my total concern. Most of it comes from someone wanting to destroy because they wanted to have the threat of destroying, or destroying value because someone wants to destroy value. That's what I would mostly be worried about. How plausible is it? I guess it seems like you think it's conceivable but pretty unlikely, so you'll pay a little bit of attention to it but it's not going to be a big focus. Is that the bottom line?
是的,所以当我之前谈到比较——如何因为可以想象而获得一点优先级——那更多是关于道德观点或不同价值的聚合以及我给予它们的权重。关键问题是,你对那些最坏结果比最好结果糟糕得多的观点赋予多少置信度。然后我认为这些观点基本上会建议,如果比例足够大,就完全专注于最小化非常非常糟糕的事情的风险。所以无论一个人的经验观点如何,都值得投入一些注意力来减少非常非常糟糕的事情的风险。就其可能性而言,理解基本形状仍然很重要。我对此并没有深思熟虑的观点。我认为答案是,以这种方式产生大量负面价值的可能性相对较小,但并非百万分之一级别的罕见——更像是 1% 的水平。还有一个问题是,当它们糟糕时,与最坏可能结果相比,实现了多少比例的糟糕,以及宇宙中有多少资源用于此。而且那个总和也不是很稳定。这就是我的立场。
Yeah, so when I was talking before about comparisons—how being conceivable means it gets a little bit of priority—that was more with respect to moral views or aggregation across different values and how much weight I give them. The key question is how much credence do you place on views where the worst outcomes are much more bad than the best outcomes are good. Then I think those views basically are going to recommend, if the ratio is large enough, just focusing entirely on minimizing the risk of really really bad stuff. So regardless of one's empirical view, it's worth devoting some amount of attention to reducing the risk of really really bad stuff. In terms of how plausible it is, that's still important to understand the basic shape. I don't really have a considered view on this. I think the answer is that it's relatively unlikely to have significant amounts of disvalue created in this way, but not unlikely at the one-in-a-million level—more like the 1% level. And there's a question of when they're bad, what fraction of the badness is realized compared to the worst possible outcome, and how much of the universe's resources go into that. I'm also that that sum is not very stable. That's kind of where I'm at.
你一开始就提出了这个论点,似乎天真地认为创造好的东西和创造同样糟糕的东西一样容易,所以更多的未来存在会想要创造好的东西和坏的东西,因此我们应该期待未来是积极的。你有多确信创造好和坏的东西在对称性上确实同样容易?
You kind of made this argument at the start that it seems like naively you'd think that it's as easy to create good things as to create something that's equivalently bad, and so more future beings are going to want to create good things and bad things, so we should expect the future to be positive. How confident are you that it actually is true that it's kind of symmetrically easy to create good and bad things?
当我们说对称地容易创造好和坏的东西时,我认为值得区分清楚这到底意味着什么。我认为这里相关的事情是,假设我们有点线性——东西两倍大或两倍好或坏——那么相关的问题是……
When we say symmetrically easy to create good and bad things, I think it's worth distinguishing being clear about what exactly that means. I think the relevant thing here, assuming that we're sort of linear—things twice as big or twice as good or bad—then the relevant question is...
你的权衡是什么?假设你有 p 的概率得到你能做的最好的事,1-p 的概率得到最坏的事。p 需要多大才能让你对那个结果和荒芜宇宙无差异?我认为我的大部分概率分布在 50% 到 99% 的好事概率之间。然后我对那些数字大一万亿倍的观点也赋予一些置信度,这种情况下它肯定会占主导。一万亿可能太大了,但非常大的数字很容易淹没实际的不可能性。一万亿确实太大了。我本应该说 'bajillion',那是我最初的想法。至于我对 50% 或 50% 到 99% 区间的信心,我想我可能会把一半的概率或权重,或者一半到三分之一,放在正好 50% 或非常接近 50% 的事情上,然后其余大部分在略高于 50% 而不是远高于 50% 之间分配。
What is your trade-off? Suppose you have a p probability of the best thing you can do and a 1-p probability of the worst thing you can do. What does p have to be such that you are indifferent between that and the barren universe? I think most of my probability is distributed between 50% and 99% chance of good things. Then I put some credence on views where that number is a quadrillion times larger or something, in which case it's definitely going to dominate. A quadrillion is probably too big a number, but very large numbers are easily large enough to swamp the actual improbabilities involved. A quadrillion is just way too big. I should have said 'bajillion', which was my first thought. In terms of how confident I am on the 50% or around the 50 to 99% range, I think I would put maybe half probability or weight, or a half to a third, on exactly 50% or things very close to 50%, and then most of the rest gets split between somewhat more than 50% rather than radically more than 50%.
我认为这些论点有点复杂,如何正确理解这一点。所以我想澄清基本立场:你最终得出最坏情况结论的原因,只是基于你的直觉:一个人可能发生的最坏事情有多坏,与最好的事情相比。你会想,该死,最坏的事情看起来相当糟糕。然后第一反应是,我们有一种祛魅式的理解,或者说我们大致从因果上理解,我们是如何形成这种关于极坏与极好事物的偏好的。如果你看看进化史上发生了什么,比如一个有机体可能经历的事件的范围,以及有机体应该如何权衡最好与最坏的结果。然后你最终会问:这种解释在多大程度上是一种祛魅式的解释,说明人类在体验快乐和痛苦的能力上是不偏不倚的,但现实仍然是有偏的?又在多大程度上这从根本上反映在我们对好坏事物的偏好中?我认为这只是一组非常难的问题。我很容易想象,经过更多深思熟虑后,我的观点会改变。
I think those arguments are a little bit complicated, how to get at this right. So to clarify the basic position: the reason you end up concluding it's worst is just based on your intuition about how bad the worst thing that can happen to a person is versus the best thing. You think, damn, the worst thing seems pretty bad. And then the first-pass response is, we sort of have this debunking understanding, or we sort of understand causally how we ended up with this kind of preference with respect to really bad stuff versus really good stuff. If you look at what happens over evolutionary history, like what is the range of things that can happen to an organism, and how should an organism be trading off best possible versus worst possible outcomes. Then you sort of end up with: to what extent is that a debunking explanation that explains why humans, in terms of their capacity to experience joy and suffering, are sort of unbiased, but the reality is still biased? Versus to what extent is this then fundamentally reflected in our preferences about good and bad things? I think it's just a really hard set of questions. I could easily imagine my view shifting on them with much more deliberation.
如果防止 s-风险成为高优先级,你认为技术性 AI 研究或你的关注点会如何改变?
How do you think technical AI research or your focus would change if preventing s-risks became a high priority?
我认为最重要的是更好地理解那些可能导致坏威胁被执行的动态,并理解我们如何安排事情以降低这种情况发生的可能性。我认为这是自然的最高优先级。
I think the biggest thing is understanding better the kind of dynamics that could plausibly lead to bad threats being carried through, and understanding how we can arrange things so it's less likely for that to happen. I think that's the natural top priority.
我最近听到一个有趣的建议,关于如何做到这一点。你可能担心的是,有人会威胁创造你认为真正无价值的东西。假设我关心痛苦;我不希望未来存在痛苦。这让我容易受到威胁:有人威胁创造痛苦,以让我在其他问题上让步。但我可以通过改变自己来避免这种风险,比如我也贬低一些实际上根本不重要的东西。假设我也不希望有飞马之类的东西,一些不存在的东西。在这种情况下,如果有人想勒索我或威胁我,他们可以转而威胁创造飞马,这些东西目前不存在,但他们可以威胁创造它们。我有可能改变我的价值观,使得创造飞马比创造痛苦更高效,这样他们就会用最有效的威胁来威胁我。这是你效用函数的一种溢出部分,保护你免受你之前关心的事物的威胁。你对这个想法或类似的东西有什么反应吗?
I heard an interesting suggestion for how to do that recently. The concern you might have is that someone would threaten to create the thing that you think is really disvaluable. So let's say I'm concerned about suffering; I don't want suffering to exist in the future. That leaves me open to someone threatening to create suffering in order to get me to concede on some other point. But I could potentially avoid that risk by changing myself so that I also disvalued something that was actually kind of not important at all. So let's say I also really don't want there to be flying horses or something like that, something that doesn't exist. In that case, if someone wanted to extort me or threaten me, they could instead threaten to create flying horses, which currently don't exist, but they could threaten to create them. Potentially I could change my values such that it's more efficient to create that than it would be to create suffering, and so that would be the most efficient threat to threaten me with. It's kind of this spillover part of your utility function that protects you from threats about the things you previously cared about. Do you have any reaction to that idea or things in that vein?
我最初的反应是这看起来有点疯狂。但自那以后,我变得对它更加热情,或者说它似乎有些合理。我想去年我颁发了一个奖项,针对与 AI 对齐或 AI 带来好结果相关的事情,其中一个获奖者是 AF 的 Casper,他提交了某种版本的这一提议。他提交了一些沿着这些思路的提议,我进一步思考后觉得有些说服力。我想从那以后他一直在继续思考,这看起来很有趣。我发现比“不在乎某件事”更合理的一个视角是:你可以说我非常在意这个随机的事情,比如有多少飞马。你也可以采取一种类似漏洞赏金的视角。如果你能令人信服地向我证明,你本可以实施这个策略,该策略有显著机会造成极端无价值,并会胁迫我做 X,或者实际上会导致我做 X,你只需足够令人信服地证明这一点。然后一旦你说服了我,我就说,好吧,你可以得到你实际上会实现的任何结果,一个从你角度来看比通过执行这个风险政策所得到的结果略好的结果。这还不清楚;我认为这极其复杂。我开始花一点时间思考这个问题,要弄清楚它是不是一个好主意,或者它是否真的有效,极其复杂。我认为这是一个我希望人们更多思考的事情。这绝对是我会做的事情之一,以理解坏威胁可能被执行的条件。我认为它比其他更常识性的干预措施(比如避免人们互相威胁的情况)影响更小,但更容易获得智力上的进展;那里有更明显的开放问题。
My initial take was that it seemed kind of crazy. But since then, I've become significantly more enthusiastic about it, or it seems sort of plausible. I think one of the prizes I was giving out last year for things relevant to AI alignment or AI leading to a good outcome, one of them was from Casper from AF for some version of this proposal. He submitted some proposal along these lines, and I thought about it more and was somewhat compelled. I think since then he's been continuing to think about it, and it seems interesting. A perspective on that that I find somewhat more plausible than 'don't care about a thing' is: you could say I care a lot about this random thing, like how many flying horses there are. You could also take this perspective that's kind of like a bug bounty. If you were to demonstrate to me convincingly that you could have run this strategy that would have had a significant chance of causing extreme disvalue and would have coerced me into doing X, or would have in fact caused me to do X, you can just demonstrate that sufficiently convincingly. Then once you've persuaded me of that, I'm like, okay fine, you can have whatever outcome you would have in fact achieved, an outcome which from your perspective is incrementally better than whatever outcome you would have achieved by carrying through this risky policy. It's not clear; I think it's incredibly complicated. I've started to spend a little bit of time thinking about this, and it's just incredibly complicated to figure out if it's a good idea or not, or whether it sort of really works. I think it's a thing I'd be sort of interested in people thinking more about. It's definitely one of the things I'd be doing to understand the conditions under which bad threats could be followed through on. I think it makes less difference than other more common-sensical interventions, like avoiding the situation where there are people threatening each other, but it is a lot easier to get intellectual traction on; there are more obvious open questions there.
人们研究 s-风险的一个原因是,他们更担心防止坏事发生,而不是创造好事。但另一个……
One reason that people work on s-risks is that they are more worried about preventing bad things than they are about creating good things. But another...
理由可能是,即使你在那一点上持对称态度,致力于防止灭绝或让未来变好的人,也比担心最坏情况并试图阻止它们的人多。所以这可能是一个更被忽视的问题,值得更多关注。你对此有多大认同?
The rationale might be, even if you're symmetric on that point, there are more people working on trying to prevent extinction or make the future go well than there are people worrying about worst-case scenarios and trying to prevent them. So it's potentially a more neglected problem that deserves more attention than it's getting. Do you put much weight on that?
我认为最终我主要关心的是忽视程度,因为它关系到可处理性。我不认为这个问题目前比 AI 对齐更易处理。它们在可处理性上似乎差不多,但部分因为这是一个更难处理的问题,它表面上的可处理性甚至低于对齐。
I think ultimately I mostly care about neglect because of how it translates to tractability. I don't think this problem is currently more tractable than AI alignment. They seem in the same ballpark in terms of tractability, but in part because it's a harder problem to deal with, it also has concerns where it's less tractable on its face than alignment.
为什么?
Why is that?
我认为很多困难的基本来源是,对齐的威胁模型非常清晰。你有一个很好的模型可以工作,你理解可能出错的地方。把对齐和一个问题比较,说它极其清晰具体——这基本上从未发生过。但在这个比较中,它异常地更清晰具体,而这里则是一种模糊的困难,我们要做的事情都更像是碰运气。
I think the basic source of a lot of difficulty is that the threat model for alignment is incredibly clear. You have this nice model in which you can work, you understand what might go wrong. It's absurd to be comparing alignment to a problem and say it's incredibly clear and concrete—that basically never happens. But in this comparison, it's unusually much more clear and concrete, whereas here it's quite a fuzzy kind of difficulty, and the things we're going to do are all much more like bang shots.
很多人认为这些不良结果和威胁的风险在多极场景下更可能发生,即许多群体竞争对未来以及 AI 或其他技术使用的影响力。你也有这种直觉吗?
A lot of people think these risks of bad outcomes and threats are more likely in a multi-polar scenario, where you have many groups competing over influence over the future and the use of AI or other technologies. Do you share that intuition?
是的,我认为至少更糟一些。我不知道糟多少——可能两倍糟是一个合理的初步猜测。关键在于人们在这个世界上互相威胁的敏感程度。这似乎是威胁的一个主要来源。所以如果人们之间的竞争不那么激烈,你预计威胁会减少。一些问题包括:人们互相威胁的数量对多极性的敏感度如何——似乎相当敏感——以及在这种动态下,未来 100 年内所有威胁中有多少会发生。
Yeah, I think it's at least somewhat worse. I don't know how much worse—maybe twice as bad seems like a plausible first-pass guess. The thing turns a lot on how sensitive people are to threatening each other in the world. That seems like a major source of threats. So if you have less rabid competition among people, you expect less of that. Some questions are how sensitive the number of threats people make against each other is to multi-polarity—it seems pretty sensitive—and what fraction of all threats occur over the next 100 years in that kind of dynamic.
你对专注于让未来包含美好事物的人和专注于防止坏事的人之间的协调有什么看法?
Do you have any thoughts on coordination between people focused on making the future contain good things and those focused on preventing bad things?
我认为他们最终会协调的原因是通过追求相似的方法和认知风格。人们通常应该协调——即使他们的目标正交,协调也是好的。共享资源、交谈和互相受益是好的。人们通常不会在某一端有极端价值观,所以这是协调的主要渠道。你也可以希望通过服务于双方的重叠目标来合作,但这是一个不太重要的渠道。如果我们都在某种程度上关心这些不同的事情并互相帮助,两个社区都可以更快乐、更健康。
I think the reason they'll end up coordinating is via pursuing similar approaches and cognitive styles. People should coordinate generically—it's nice if they can coordinate even if their goals are orthogonal. Sharing resources, talking, and benefiting from each other is good. People normally don't have extreme values at one end or the other, so that's the main channel for coordination. You could also hope for cooperation through overlapping objectives that serve both, but that's a less important channel. Both communities can be happier and healthier if we all care to some extent about these different things and help each other out.
我们来谈谈哲学和伦理学。不同的元伦理理论在 AI 对齐和对齐研究中扮演什么角色?
Let's talk about philosophy and ethics. What role do different theories of meta-ethics play in AI alignment and alignment research?
哲学进展影响对齐有两种性质不同的方式。一种是对象层面:对齐 AI 所需的工作涉及澄清一些哲学问题。另一种是你根据对这些哲学问题的看法来对待对齐的方式。在对象层面,如果你认为必须理解什么是好的,以便传达或理解一个最终正确理解善的过程,并直接将其灌输到 AI 系统中,你会处于一个艰难的境地——你必须解决一堆哲学问题,或者解决 Yudkowsky 所说的元哲学:理解人类在哲学探究中如何达到真理。这似乎相当艰难,并且会有紧密的对象层面联系,甚至可能说这是同一个问题。我认为这更接近我六年前的想法,但我已经改变了很多。现在我认为你只需要一个系统,其中要澄清的是控制和纠正的概念。我们希望 AI 的构建不会让事情变得更糟,最终处于像我们现在这样的位置,继续同样的深思熟虑过程,理解它如何变好或变坏,并纠正它。我们希望在构建 AI 时尽可能少地做出伦理承诺。我对任何涉及硬哲学承诺的方法都变得非常悲观。我认为我们最终还是会做出一些承诺,但如果事情顺利,很可能是因为我们可以避开大部分承诺。
There are two qualitatively different ways philosophical progress could affect alignment. One is on the object level: the work needed to align AI involves clarifying some philosophical questions. Another is the way you approach alignment depending on your views on those philosophical questions. On the object level, if you thought you had to understand what is good to communicate or understand a process that converges to a correct understanding of the good, and directly impart that into an AI system, you'd be in a rough position—you'd have to solve a bunch of philosophy or solve what Yudkowsky calls meta-philosophy: understanding how humans arrive at truth in philosophical inquiry. That seems pretty rough, and there'd be a tight object-level connection, maybe even saying these are the same problem. I think that's closer to where I was six years ago, but I've shifted a lot. Now I think you just want a system where the thing to clarify is the notion of control and course correction. We want the construction of AI to not make anything worse, to end up in a position like the one we're in now, where we continue the same process of deliberation and understanding how it goes well or poorly, and correcting it. We want to make as few ethical commitments as possible when constructing AI. I've become much more pessimistic about any approach involving hard philosophical commitments. I think we still end up making some, but if things go okay, it's probably because we can dodge most of them.
你为什么变得更悲观了?
Why have you become more pessimistic?
部分原因是案例的技术细节。我在思考不同的对齐方法可能如何展开。我认为你必须大量依赖人类纠正的机制,或者依赖人类深思熟虑的过程,才能在没有哲学问题的情况下有任何希望。一旦你普遍依赖它,你不如也用它来回答这些问题。部分原因只是它变得更多了。
In part from technical details of the case. I'm thinking about how different approaches to alignment might play out. I think you have to lean on this mechanism of course correction by humans, or deferring to human deliberation, a lot anyway to have any hope apart from philosophical issues. Once you're leaning on it in general, you might as well just lean on it for answering these questions. In part, it's just becoming a lot more.
对此前景感到乐观。所以你可能要问,当 AI 代表你行动时,它理解你希望宇宙整体发生什么有多重要。我更新了很多看法,认为即使它不完全理解也没关系,即使在非常悲观的情况下,事情变得疯狂,比如从你启动 AI 到宇宙殖民开始之间只有 6 分钟。我认为即使那样,它基本上也没问题,如果它不太理解你对宇宙的期望,只是理解“看,这是我”,然后把你放在某个盒子里,开始殖民宇宙,最终为你的盒子腾出空间,让你在偏远星球上的小文明去思考什么是好的,最终这个过程只需保持对那种深思熟虑的结论的响应。它需要理解的是:保护人类并让人类以我们想要的方式发展和成熟意味着什么,以及最终响应那个过程得出的结论意味着什么,一旦人类弄清楚我们在某个领域想要什么,就要能够纠正,让那种理解最终影响我们建立的自动化框架的行为。
Optimistic about the prospects for that. So it's like you might ask how important is that they understand what you want to happen with the universe as a whole when it goes out and acts on your behalf. I've updated a lot towards it's okay if it doesn't really understand, even in really pessimistic cases where stuff is getting crazy, like over you know there's going to be like 6 minutes between when you start your AI and when the colonization of the universe begins. I think even then it's like basically okay if it doesn't understand what you want for the universe that much, just understands like look here's me, it's like great put you in a box somewhere now, like start colonizing the universe and then eventually make some space for your box to sort out what Humanity should do, like your little civilization on a planet somewhere out in the backwoods trying to figure out what is good, and then ultimately the process just has to remain responsive to the conclusions of that deliberation. The things it has to understand are what does it mean to protect humanity and allow Humanity to develop and mature in the way that we want to develop and mature, and then what does it mean to ultimately be responsive to what that process concludes, to be correctable once humans figure out what it is we want in some domain, to allow that understanding to ultimately affect the behavior of this scaffolding of automation we built up around us.
可能还有一个更技术性的问题。你可能认为哲学会影响不同资源的价值,但我认为一旦你更仔细地考虑这些论点,你可以在某种程度上回避这种依赖。
There's maybe one last more technical question that comes up there. We like you might think philosophy would affect the value of different kinds of resources, and there's just some more I think you can sort of dodge those kinds of dependence once you're more careful about these arguments.
你对“长期反思”这个想法怎么看?这个想法是,我们并不真正知道什么是有价值的,如果我们让最优秀的人思考数千年或很长时间,直到我们决定如何最好地利用我们在宇宙中能获得的所有资源,我们可能会有更好的机会。
How do you feel about the idea of the long reflection? This idea that well, we don't really know what's valuable, it seems like we might have a better shot if we get our best people to think about it for thousands of years or just a very long time until we decide what would be the best thing to do with all the resources that we can get in the universe.
好主意。我认为把它看作一个发生的步骤可能不太对,但我很赞同这个想法,即我们目前正在进行的那种理解什么是好的深思熟虑。我认为 AI 的大部分作用就是让这个过程继续发生。你可以认为这个过程最终与大多数经济扩张脱钩,通过宇宙扩张与这种持续的深思熟虑和理解我们想要什么的过程脱钩。你可以想象这样一个世界:人类生活在地球上,而在太空中,AI 发动战争、建造机器等疯狂的事情正在发生,而人类在地球上做着正常的事情。有时我们未来学家过去认为那只是附带事件,所有行动现在都集中在 AI 做的疯狂事情上。我的观点更多地转向了:实际上那里有很多行动,我们价值观的整体演变。我们选择我们喜欢什么,我们想要什么成为价值储存。我不认为是一个人离开去思考一千年,更像是我们的文明沿着一条轨迹发展,我们在思考这条轨迹应该是什么。我认为发生的一件事,对齐的希望之一,是将这种持续的深思熟虑过程与保持经济竞争力的过程耦合起来。理解深思熟虑应该是什么样子是一个非常困难的问题。我认为这是另一个重要方面,当我一开始说我稍微降低了对对齐相对于问题其他部分重要性的整体感觉时,很大程度上是因为我提高了让这种深思熟虑顺利进行的重要性。我认为从长远历史角度来看,除了对齐,这似乎是首要任务。
Sound idea. I think viewing it as a step which occurs is probably not quite right, but I think I'm pretty on board with the idea that it's just of deliberation understanding what is good that we're kind of currently engaged in. I see most of the action in AI as allowing that process to continue happening. I think you can view that process as decoupled ultimately from most of the economic expanding through the universe is kind of decoupled from this process of ongoing deliberation and understanding what we want. Well, you can sort of imagine the world where the humans are living on Earth while out there in space a bunch of crazy is going down with AI waging wars and building machines and stuff, and the humans are just doing a normal thing on earth and sometimes we the futurist used to think of that as a sideshow where all the action was now off with the crazy stuff the AIs are doing. My perspective is more shifted to actually that's sort of where a lot of the action is, the overall evolution of our values. We choose what we like, what we want to be the store of value. I don't think it's a person going off and thinking for a thousand years, it's more like there's a trajectory along which our civilization is developing and we're thinking about how that trajectory should be. I think that one thing that happens, one of the hopes of alignment is to couple that process of ongoing deliberation from the process of remaining economically competitive. It's a really hard problem to understand what deliberation should look like. I think that's another one of the big ways when I said at the beginning that I've sort of slightly downgraded my overall sense of how important alignment was to other parts of the problem, a lot of that has been upweighting how important this kind of making that deliberation go well is. I think that's really from a long-term historical perspective, other than alignment seems like probably top priority.
我认为在开头我们谈到元伦理学时,我区分了对象层面和元层面的影响。值得指出的是,这里我们只是深入探讨了对象层面,我很乐意继续这样。而且值得在元层面上说,我确实认为你对齐的方法,或者不同结果的重要性和价值,取决于一些与此相关的棘手伦理问题的答案,比如你有多关心蜥蜴人?这是道德哲学中的一个难题,类似的难题还有:你有多关心你正在构建的这个 AI 系统?如果它想做一些与人类完全不同的事情,你会有多愿意说“好吧,我们有一些价值观,AI 有一些价值观,我们很高兴”?我想我们上次在播客中稍微讨论过这一点。我认为这些道德问题的答案确实会影响你如何进行对齐,或者你如何优先考虑问题的不同方面。
I think it's also worth at the beginning when we talked about meta ethics, I made this distinction between object level and meta level influences. It's worth bracketing that here we've just been diving in on the object level and I'm happy to keep going with that. And it's worth saying really at the meta level that I do think that your approach to alignment or how important how valuable different kinds of outcomes are depends on the answers to some hard ethical questions related to this, like how much do you care about the lizard people? That's sort of a hard question in moral philosophy, and similar hard questions like how much do you care about this AI system that you're building? If it wants to do something totally different from humans, how much would you be like well we some values AI some values we're happy? I think we talked about this a little bit last time on the podcast. I think that the answers to those kinds of moral questions does have an effect on how you go about alignment or how you prioritize different aspects of the problem.
是的,那么你对道德实在论和反实在论的看法是什么?这些会影响 AI 对齐的工作吗?似乎最重要。
Yeah, so what are your views on moral realism and anti-realism and do those affect like what AI alignment works? Seems most important.
我绝对是相当反实在论的。我的意思是,我们在这里有点陷入语义细节。当我与人们进行长时间讨论时,我认为有一个问题,感觉像是一个实在论问题:“深思熟虑是否朝着正确的方向进行?”我不认为你可以有一种反实在论的观点,即“你怎么深思熟虑并不重要,你会得出一些结论,这些结论都没问题”,我不赞同那种观点。你还有另一种观点,即“你不应该,你应该只赞同你当前的结论,因为那是你喜欢的”,我也不赞同那种观点。我会说,看,现在有一些我喜欢的深思熟虑和成长过程,我赞同它们的输出。我希望我们的价值观以某种方式演变。在某种意义上,你可以说我关心的是那个深思熟虑过程的终点,那个可能非常非常漫长的进化成熟过程的终点。我认为我喜欢,并且在哲学上不认为它们必然在不同过程中趋同。我认为不同的过程启动后会得出不同的结论。但我认为有一个非常非常难的……
I'm definitely like pretty anti-realist. I think I mean it's a little bit we get a little bit into semantic weeds here. When I've had long discussions with people about it, I think there's this question which is sort of this question which feels like a realist question of 'is deliberation going in the right direction?' I don't think you could have a version of the anti-realist perspective where you're like 'it doesn't matter how you deliberate, you're going to come to some conclusions and those are fine' and I don't endorse that. You have another version of perspective where you're like 'you shouldn't, you should just endorse your current conclusions because that's what you like' and I don't endorse that either. I'd say look, right now there's some kinds of processes of deliberation and growth that I like, I endorse the output of. There's some way I want our values to evolve. And in some sense you could say that what I care about is the end point of that deliberative process, the end point of that potentially very very long process of evolution maturation. I think I like and philosophically don't think there's like I don't think they necessarily be convergence across different. I think different processes set in motion would arrive at different conclusions. But I think there was a very very hard...
比如让事情朝着正确方向前进,作为一个非实在论者,有点尴尬地问“正确方向”在这里是什么意思。实在论者有一个很好的简单答案,比如它实际上在趋近于善。我认为这只是一个语言上的问题,他们恰好有一个好的——我的意思是,整件事都是语义差异。只是有些概念对于非实在论者来说难以讨论和思考。但我认为这是因为对于实在论者来说,这被归结为“善”这个概念本身的模糊性和复杂性。
Problem like having that go in the right direction and like it's a little bit awkward as a non-realist to be like what does the right direction mean here. The realists have like this nice easy answer like it's actually converging to the good. And I think that's kind of just like a linguistic thing where like they happen to have a nice — I mean again the whole thing is kind of like semantic differences. It's just like there's some concepts that are slippery or hard to talk about and think about for the non-realist. But I think that's because they are in fact like for the realist that's just pushed down to the slipperiness and complexity of like the actual concept of good.
就我的整体观点而言,那些可能影响我如何优先处理对齐不同部分的客观问题,我认为并没有太多趋同。我认为很有可能一个 AI 会很聪明,做自己的事情,这介于 Baron 宇宙和实际达到最优结果之间。我不认为它非常接近 Baron,也不认为它非常接近最优结果。我倾向于一个 50-50 的先验,比如如果它只有灭绝一半的糟糕程度,我可能可以接受让某个随机 AI 在宇宙中做它随机的事情。
In terms of what I find like my view overall, like the object level questions that might affect how I prioritize different parts of alignment is like I don't think there's that much convergence. I think it's quite plausible that like an AI would be smart and would do its own thing and like that's somewhere between Baron universe and actually achieving like the optimal outcome. And I'm like I don't think it's very close to bar, I don't think it's very close to the optimal outcome. Like I sort of lean towards like a 50-50 prior for like I'd be okay if it's half as bad as extinction to like have some random AI doing its random thing with the universe.
好了,我们已经聊了一段时间,该结束了。但之前我们谈到 AI 可能让人变得更聪明。哦抱歉,是三个点,三分之一标准差,可能最好的情况了。但你对益智药、药物和生活黑客这些东西有什么看法?人们试图用它们来让自己更健康、更聪明,这有多大用处?
All right, we've been going for a while and we should wrap up. But yeah, earlier we were talking about how AI could potentially make people a whole lot smarter. Oh sorry, three points smarter, a third of a standard deviation, as good as it gets potentially. But yeah, do you have any view on kind of nootropics and the drugs and life hacking stuff that people try to use to make themselves fitter and more intelligent? Is there much mileage to be gotten out of that?
我确实对这方面的研究很兴奋,感觉医学界不仅不感兴趣,而且非常不感兴趣。我认为有合理的可能性存在一些东西。有几个案例,如果只看文献的现状,你会觉得“啊,看起来确实”——我不知道,也许吡拉西坦就是这种情况,但我不太确定。还有其他几个可能的候选,如果你真的相信所有研究并只看表面,你会认为其中一些有合理的收益。但可能事实并非如此,大家也都知道这些是较老的研究,可能有一些未发表的失败重复。但能逐一检查它们会很好,我对那很兴奋。
So I'm definitely pretty excited about investigation where it kind of feels like a thing that the medical establishment is not merely not into but like very not into. And I think there's like a reasonable chance that there's something. I think there are like a few cases where the current state of the literature would sort of if you took everything at face value you'd be like ah it does seem like — I don't know, I think maybe piracetam is in this position but I'm not totally sure. There's like a few other possible candidates where if you actually believed all the studies and just took that at face value you would think that there's reasonable gains from some of these. I think like probably that's just not the case and like sort of everyone understands that these are like sort of old studies and like there probably been failed replications that haven't been published. But it would be like pretty nice to like go through and check them all and I'd be like pretty excited about that.
我总体上很感兴趣。我最想看到的是:你有这个问题,为什么对大脑化学的简单改变能提高表现,但进化却没有做到?所以你有点想看看抵消因素是什么。这确实可能解释为什么很难找到真正有效的东西。是的,如果那么容易,进化早就做到了。
I'm pretty interested in general. I think the thing I would most like to see in general: you have this question of why would it be the case that some simple change to brain chemistry would improve performance but not be made by Evolution. So you sort of want to see what the countervailing consideration is. And that does potentially explain why it has been pretty hard to find anything that seems like it works really well. It's just yeah, if it was as easy as that then Evolution would have done it.
所以我认为你必须利用某种分布变化或某种成本,这种成本现在是成本,但历史上不是。我认为两个最好的候选是:第一,你可以利用这种虚伪的角度。人们有他们想做的事情,比如我想让世界变得更好,但在某种程度上,你的生物学完全不是为了改善世界而优化的,而是为了拥有后代的后代。所以一件事是,如果你想追求你名义上想做的事,你可以希望有一些药物让你更好地实现你设定的目标,即使那个目标与你的身体优化目标不一致。我认为这在某些情况下是一种作用机制,看起来相当现实,而且我一直很害怕这一点。
So I think you're going to have to exploit like some distributional change or like some cost which is now a cost that wasn't a cost historically. I think the two best candidates are basically: one, you sort of can exploit this hypocrisy angle. So people like they have this thing they want to do like I want to make the world better, but then there's some level which your biology is not at all optimized for making the world better, it's optimized for having descendants who have descendants. So one thing is if you want to pursue what you nominally want to do, you can hope that there's some drugs that just make you better at achieving the thing that you set out to do, even when that thing is not in line with the thing that your body is optimized for achieving. I think that's like a mechanism of action in some cases and seems like pretty realistic and it's something I've been really scared of.
另一个,最令人满意和出色的,就是燃烧更多能量。在进化环境中,你不想让大脑过热,因为那有点浪费,你只能得到边际收益。但如果你能做到,那就太好了,就像超频一样,因为我们现在有这么多食物,太便宜了,这么多食物。你告诉我,我每天多消耗 400 卡路里,变得更聪明一点。我觉得这对两方面都有好处。是的,那绝对是最好的情况。
And then the other one, the thing that would be most satisfying and excellent, would be just burn more energy. Like in the evolutionary environment you didn't want to run your brain hot because it was kind of a waste and you're only getting marginal benefits. But if you could just do that, that would be super great, just overclock it because it's like we got so much food now. It's so cheap. So much food. You told me like I'm going to burn an extra 400 calories a day and be marginally smarter. I'm like that's good on both sides. Yeah, that would really be the best case definitely.
我有点困惑兴奋剂在这方面的情况。比如我对咖啡因的理解是,如果我摄入咖啡因,可能会升高血压和增加能量消耗,至少一开始是这样。然后如果我持续摄入,大概两周内,至少在血压方面,我会回到基线。我想更好地了解长期对认知表现和能量使用等的影响,以了解这些长期效应:如果你反复服用某物,最终它会停止产生影响,还是即使只服用一次,接下来几天会有反弹期,使得长期效应是这个东西的积分,或者显示积分为零。我不知道。无论如何,对我来说似乎有些收益可能主要通过这两个渠道:要么让你做你认为应该做或试图做的事情,要么燃烧更多能量。
I a little bit confused about where stimulants stand on that. Like usually my understanding for caffeine is that if I take caffeine I'm probably going to drive up blood pressure and drive up energy expenditure at least at the beginning. And then if I keep taking it probably within like 2 weeks at least on the blood pressure side I'm going to return to sort of baseline. And I would like to understand better what long-term effects are on cognitive performance and energy use and similar, to understand whether those long-run effects like you know is it the case that if you take something over and over again then eventually it stops having impact, or is it the case if you take it even once you have like a bounceback period over the next couple days such that the long-term effect is sort of the integral of this thing that or like showing that the integral was zero. I don't know. Anyway, it seems like kind of plausible to me there's some wins probably mostly through these two channels of either getting you to do what you think you should do or are trying to do, or else burn you more energy.
我个人并没有优先考虑这类实验,部分是因为我不擅长内省,所以即使处于非常改变的精神状态也无法察觉,部分是因为我认为单样本很难得出结果。但我对更多这类实验很兴奋。我认为这很难,部分是因为医学界似乎非常讨厌这类东西,而且一些可能的赢家也不合法,这很糟糕。比如普雷姆筋膜?安非他命可能是最自然的。
I have like personally not been have not prioritized experiment with that kind of thing, and part because I like I'm really bad at introspection and so cannot tell even if I'm in like a very altered mental state, and partly because yeah I think it's really hard to get with n equals one. But I'm pretty excited about more experimentation on those things. I think it's like really hard partly because the medical establishment really seems to hate this kind of thing, and then a bunch of the some of the like likely winners are also just like not legal which sucks. It's like Prem fascia? Amphetamines would probably be the most natural.
嗯,但你真的认为它们长期来看有好处吗?
Well yeah, but do you actually think they would be good in the long run?
呃,身体似乎非常善于适应兴奋剂,大多数时候它们在第一周或第一个月有效,但之后你又回到了基线。
Uh, it does just seem like the body is so good at adapting to stimulants that most of the time kind of they work the first week or the first month, but then you're back to baseline and now.
你只是吃这个东西让自己恢复正常。虽然几乎所有认知增强剂都是合法的,但其中少数像安非他命,服用它们有惹上法律麻烦的风险,甚至未来可能无法获得安全许可,这对任何可能想从事政策职业的人来说都是一个非常严重的潜在缺点。而真正想要或需要安全许可的人在我们听众中占了相当大的比例。所以在这些情况下,很明显你应该远离。但基本上,我不推荐人们过多关注益智药,即使是合法的,不仅因为对很多药来说,它们有效的证据本来就很弱,还因为身体非常擅长适应并抵消你服用的几乎任何东西的效果,几周后就不清楚是否有可持续的收益了。这是我的主要担忧。而且,我的意思是,你可以想象……所以我认为人们有一种感觉,比如在服用处方兴奋剂的情况下,人们觉得反复使用确实能获得一些持久的优势,或者至少是持久的治疗效果,检查一下这是否真的如此,或者只是零和游戏,会很好。确实感觉一旦你消耗更多能量,就没有很好的理由……就像如果我试图服用一些吡拉西坦,希望自己思考得更好,这有很好的理由预期会失败。然后从进化论的角度来看,对于兴奋剂这种最初让你消耗更多能量、思考得更好的东西,没有很好的理由预期它会失效,所以至少还有希望。我认为在我的理想世界里,你真的会投入很多精力:我们能在这里找到赢面吗?因为这看起来确实有可能。
You're just taking this thing just to get yourself back to normal. And while almost all these cognitive enhancers are legal, with a handful of them like amphetamines, there's risk of getting into legal trouble if you take them, or even not being able to get security clearance in future if you've taken them, which is a really serious downside potentially for anyone who might want to go into a policy career in future. And people who might really want or need a security clearance at some point make up a pretty large fraction of all of our listeners. So in those cases, it's kind of clear you should stay away. But basically, I don't recommend that people pay much attention to nootropics, even the legal ones, both because of the fact that for so many, the evidence that they work at all is kind of weak in the first place, but then also this thing that the body is so good at adapting to and undoing the effect of almost anything you take after a few weeks, that it's unclear that there are kind of sustainable gains. That's a lot of my concern. And like, I mean, you could imagine... So I think people have a sense, like under prescribed stimulant for example, have a sense that under repeated use you do get some lasting advantage, or at least lasting treatment effect, and it would be good to check whether that's actually the case or whether that's just shuffling things around and zero sum. It does kind of feel like once you're burning more energy, there's not really a good reason... It's like if I'm trying to take some piracetam and just hope that I think better, there's sort of a good reason to expect to fail. And then with this evolutionary argument, in the case of a stimulant which is initially causing you to use more energy and think better, there's not really a great reason to expect it to break down, and so there's hope at least. And I think in my ideal world, you'd really be throwing a lot of energy: can we find a win here? 'Cause it sure seems plausible.
是的,我有一个希望,即使大脑如此……似乎你服用兴奋剂后获得了益处,但你的身体正在适应它,然后你必须有段时间让身体去适应它才能再次服用。比如说,你连续服用一周,然后停用一周以清除体内的适应,这似乎不太好。但有可能你可以每天服用,这样白天更清醒,然后在睡觉时去适应,这似乎对两边都有潜在好处。但我不确定在你不服药的 12 小时内,去适应是否真的足以让它值得。
Yeah, one hope I've had is that even if the brain is so... It seems like you take a stimulant and then you get a benefit from it, but your body is adapting to it, and then you kind of have to have a time when your body de-adapts from it in order to take it again. If you like, say, take it all the time for a week, and then don't take it for a week in order to flush out the adaptation from your system, that doesn't seem so great. But potentially you could do this every day such that you're a bit more awake during the day, and then you de-adapt to it while you're sleeping, and that just seems kind of potentially good on both sides. But I'm not sure whether the de-adaptation over the 12 hours that you're not taking it is really sufficient to make it worthwhile.
是的,我觉得似乎有可能这里有些东西是有效的。对我来说,可能你应该坚持一些愚蠢的想法……我在纸上想过一些,也试过一些,但我认为做实验很难,因为医学界讨厌它,工具也有点难,你得克服这些,而且最高杠杆的选择可能不合法,这让它整体上不那么吸引人。你可以用咖啡因试试,但似乎可能收益更少。
Yeah, I think it feels plausible to me there's something here that works. It's possible to me that you should just run with some dumb... I thought about a little bit on paper, I tried some, but I think it's hard to do the experiments due to the combination of medical establishment hates it, the instruments are a little bit hard, think you get over that, and also the most high-leverage options probably are not going to be legal, which makes it just overall less appealing. You can try this with caffeine, but it seems like just probably less wins.
好吧,在这个鼓舞人心的音符上,我今天的嘉宾是 Paul Christiano。感谢你再次来到播客。
Well, on that inspiring note, my guest today has been Paul Christiano. Thanks for coming back on the podcast.
嗯,是的,谢谢你邀请我。
Well, yeah, thanks for having me.
如你所料,本集附带的博客文章中有很多链接,指向论文和文章,你可以从中了解更多我们在过去两小时里讨论的内容。正如我在节目开始时提到的,在节目说明中,我们链接了一个 40 分钟的 MP3,其中 Paul 和我进行了一场特别令人困惑且有点疯狂的关于决策理论的对话,以及如果一些非标准解决方案被证明是正确的,那可能意味着什么。我并不特别推荐听它,这就是为什么我们把它从节目中剪掉了,但对决策理论特别感兴趣的人可能会喜欢它。我想要么是为了了解我们如何看待这个问题,要么是作为无意的喜剧。好了,80,000 Hours 播客由 Keiran Harris 制作。感谢收听。一两周后再聊。
As you expect, there's lots of links in the blog post attached to this episode to papers and articles where you can learn a lot more about the things we covered over the last 2 hours. And as I mentioned at the start of the show, in the show notes we link to a 40-minute MP3 where Paul and I have a particularly confusing and slightly bonkers conversation about decision theory and what it might mean if some non-standard solutions turn out to be on the mark. I don't especially recommend listening to it, which is why we cut it from the episode, but people who are particularly interested in decision theory might enjoy it. I guess either to learn how we think about the problem or maybe as unintentional comedy. All right, the 80,000 Hours podcast is produced by Keiran Harris. Thanks for joining. Talk to you in a week or two.