Chris Olah 谈神经网络可解释性与 AI 安全

Chris Olah on Neural Network Interpretability and AI Safety

克里斯·奥拉 Chris Olah · 80,000 小时 · 2023-10-31 · 约 189 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

顶尖机器学习研究员 Chris Olah 探讨可解释性研究、神经网络工作原理、多模态神经元、缩放定律以及他的新 AI 实验室 Anthropic。

Chris Olah, a top machine learning researcher, discusses interpretability research, how neural networks work, multimodal neurons, scaling laws, and his new AI lab Anthropic.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 69)

全文 · Full transcript(中英对照)

播客与嘉宾介绍 Introduction to the podcast and guest

Host

听众朋友们,欢迎收听《八万小时》播客,我们在这里深入探讨世界上最紧迫的问题、你能做些什么来解决它们,以及为什么那里有云而旁边却没有。我是罗布·威布林,八万小时的研究主管。我非常兴奋地分享这一期节目,因为今天的嘉宾克里斯·奥拉是世界上最顶尖的机器学习研究者之一。他还非常擅长通过博客文章和推特线程向公众传达复杂的思想,多年来吸引了数百万读者。一项听众调查甚至发现,他是本节目订阅者中关注度最高的人之一。然而,尽管他相当有名,这是克里斯第一次做播客,实际上也是他第一次做长篇访谈。幸运的是,我认为你不会看出他以前没怎么做过。我们和克里斯有太多新颖的内容要覆盖,所以进行了不止一次录音,最终制作出两集截然不同的节目,我们都非常满意。第一集,也就是这一集,聚焦于克里斯的技术工作,探讨了可解释性研究是什么以及它试图解决什么问题、神经网络实际上如何运作以及它们如何思考、多模态神经元及其对 AI 安全工作的影响、克里斯的方法能否规模化、数字痛苦、机器学习模型中的缩放定律,以及如果这里描述的所有工作都能成功该有多好。如果你觉得这一集的技术部分有点难懂,在完全放弃之前,我建议跳到名为“Anthropic 与大型模型安全”的章节,或者直接跳到 2 小时 12 分钟处,听听克里斯目前正在帮助发展的那个非常令人兴奋的项目。你不必是 AI 研究者也能在 Anthropic 工作,所以我认为来自各种背景的人都能从坚持听到最后中受益。不过,我本人远非 AI 如何工作的专家,但我基本能跟上克里斯的思路,因此学到了不少关于大型机器学习模型研究的真实情况。第二集,如果一切顺利的话,我们希望下周发布,主要聚焦于克里斯非常迷人的个人背景故事,包括他如何在没有大学学位的情况下走到今天。哦,还有最后一件事:耳朵尖的听众会注意到我们的音频在这里变化了几次,这是因为,正如我们较长节目中的常见情况,基兰将多次录音的片段拼接在一起,以创造更好的最终产品。好了,闲话少说,有请克里斯·奥拉。

Hi listeners, this is the 80,000 Hours podcast, where we have unusually in-depth conversations about the world's most pressing problems, what you can do to solve them, and how come there's a cloud there but not a cloud right there just next to it. I'm Rob Wiblin, head of research at 80,000 Hours. I am really excited to share this episode because Chris Olah, today's guest, is one of the top machine learning researchers in the world. He's also really excellent at communicating complex ideas to the public with his blog posts and Twitter threads over the years, attracting millions of readers. A survey of listeners even found that he was one of the most followed people among subscribers to this show. And yet, despite being a pretty big deal, this is the first podcast and indeed the first long interview that Chris has ever done. Fortunately, I don't think you'll be able to tell that he hasn't actually done this many times before. We ended up having so much novel content to cover with Chris that we did more than one recording session, and the end result is two very different episodes that we're both really happy with. The first one, this one, focuses on Chris's technical work and explores topics like what interpretability research is and what it's trying to solve, how neural networks actually work and how they go about thinking, multimodal neurons and their implications for AI safety work, whether the approach that Chris is taking can scale, digital suffering, scaling laws in machine learning models, and how wonderful it would be if all of the work that is described here could succeed. If you find the technical parts of this episode a bit hard-going, before you give up on this episode altogether, I'd recommend skipping to the chapter called 'Anthropic and the Safety of Large Models', or going to 2 hours and 12 minutes in, to hear all about the really exciting project that Chris is helping to grow right now. You don't have to be an AI researcher to work at Anthropic, so I think people from a wide range of backgrounds could really benefit from sticking around to the end. For what it's worth, though, I am far from being an expert on how AI works, but I was mostly able to follow Chris and so learn quite a bit about what's really going on with research into big machine learning models. The second episode, which if everything goes smoothly we hope to release next week, is focused on Chris's really fascinating personal backstory, including how he got where he is today without having a university degree. Oh, and one final thing: eagle-eared listeners will notice that our audio changes a couple of times here, and that's because, as is often the case with our longer episodes, Kieran has cut together sections from multiple different recording sessions to try and create a better final product. All right, without further ado, I bring you Chris Olah.

Host

我正在和克里斯·奥拉对话。克里斯是一位机器学习研究者,目前专注于神经网络可解释性。直到去年十二月,他领导了 OpenAI 的可解释性团队,但今年他帮助创立了一个名为 Anthropic 的新 AI 实验室,该实验室特别关注超大型机器学习模型的安全。在 OpenAI 之前,他在谷歌大脑工作了四年,开发了可视化神经网络内部运作的工具。克里斯在谷歌大脑产生了巨大影响:他是 2015 年 Deep Dream 推出的第二作者,我认为几乎所有人都见过这个作品,他还开创了特征可视化、激活图谱、可解释性的构建模块、TensorFlow,甚至合著了著名的论文《AI 安全中的具体问题》。除此之外,2018 年他帮助创办了学术期刊 Distill,致力于清晰传达技术概念。克里斯本人也是一位作家,在节目的许多听众中很受欢迎,他通过高度易懂的方式解释前沿机器学习,吸引了数百万读者。2012 年,克里斯获得了 10 万美元的蒂尔奖学金,这是一项旨在鼓励有天赋的年轻人直接进入研究或创业领域而非上大学的奖学金。所以他实际上在没有学位的情况下完成了以上所有成就。非常感谢你来做客播客,克里斯。

I'm speaking with Chris Olah. Chris is a machine learning researcher currently focused on neural network interpretability. Until last December, he led OpenAI's interpretability team, but this year he has helped to launch a new AI lab known as Anthropic, which is particularly focused on the safety of very large ML models. Before OpenAI, he spent four years at Google Brain developing tools to visualize what's going on in neural networks. Chris had a big impact at Google Brain: he was the second author on the launch of Deep Dream back in 2015, something which I think almost everyone has seen at this point, and he also pioneered feature visualization, activation atlases, building blocks of interpretability, TensorFlow, and even co-authored the famous paper 'Concrete Problems in AI Safety'. On top of all of that, in 2018 he helped found the academic journal Distill, which is dedicated to publishing clear communication of technical concepts. Chris is himself a writer who is popular among many listeners to the show, having attracted millions of readers by trying to explain cutting-edge machine learning in highly accessible ways. In 2012, Chris took a $100,000 Thiel Fellowship, a scholarship designed to encourage gifted young people to go straight into research or entrepreneurship rather than go to a university. So he's actually managed to do all of the above without a degree. Thanks so much for coming on the podcast, Chris.

Chris Olah

谢谢邀请。

Thanks for having me.

AI 对齐问题的概念 Conception of the AI alignment problem

Host

好了,我希望我们能聊聊你的新项目 Anthropic 以及可解释性研究,这是你最近的重点之一。首先,你似乎在过去八年里以某种形式为解决 AI 对齐问题做出了贡献。你能简单说说你如何从高层次上理解这个问题的本质吗?

All right, I hope that we're going to get to talk about your new project Anthropic and the interpretability research which has been one of your big focuses lately. First off, it seems like you kind of spent the last eight years contributing to solving the AI alignment problem in one form or another. Can you just say a bit about how you conceive of the nature of that problem at a high level?

Chris Olah

当我和其他人谈论安全问题时,我觉得他们往往对安全问题的本质以及它看起来是什么样子有相当强烈的看法或成熟的见解。而我想,我在思考安全、从事安全工作以及广泛思考机器学习的过程中学到的一个教训是,事后看来,我似乎经常是错的,或者我认为自己之前没有以正确的方式思考问题。所以我认为我对安全并没有非常自信的看法。在我看来,这是一个我们实际上知之甚少的问题,试图先验地理论化很容易让你误入歧途。相反,我非常有兴趣尝试从经验上理解这些系统,从经验上理解它们可能如何不安全,以及我们如何能够改进。特别是,我真的很感兴趣我们如何理解这些系统内部真正发生了什么,因为在我看来,这是理解潜在故障模式和风险的最大工具之一。

When I talk to other people about safety, I feel like they often have pretty strong views or developed views on what the nature of the safety problem is and what it looks like. And I guess I feel like one of the lessons I've learned trying to think about safety and work on it and think about machine learning broadly has been how often I seem to be in retrospect I'm wrong or I think that I wasn't thinking about things in the right way. So I think I don't really have a very confident take on safety. It seems to me like it's a problem that we actually don't know that much about, and that trying to theorize about it a priori can really easily lead you astray. Instead, I'm very interested in trying to empirically understand these systems, empirically understand how they might be unsafe and how we might be able to improve that. And in particular, I'm really interested in how we can understand what's really going on inside these systems, because that seems to me like one of the biggest tools we can have in understanding potential failure modes and risks.

可解释性与电路简介 Introduction to interpretability and circuits

Host

好吧,既然如此,我们就不浪费时间进行宏观的理论思考了,直接进入你一直在做的具体实证技术工作,这实际上是逆向工程神经网络,以便观察其内部并理解真正发生了什么。我在这里主要参考的两篇在线文章是 2020 年 3 月的《放大:电路导论》,这是关于电路概念系列文章的第一篇,我们稍后会讨论;还有最近 2021 年 3 月的文章《人工神经网络中的多模态神经元》。我显然都读过,它们有很多精美细致设计的图片,所以如果听众真的想彻底理解这一切,可能值得一看。它们都聚焦于神经网络的可解释性问题,特别是视觉神经网络。我得承认:这些文章有点超出我的理解范围,我没有完全跟上,但也许克里斯能帮我弄明白。

All right, well in that case let's waste no time on big picture theoretical musings and we can just dive right into the kind of concrete empirical technical work that you've been doing, which in this case is kind of reverse engineering neural networks in order to look inside them and understand what's actually going on. The two online articles that I'm going to be referring to most here are the March 2020 article 'Zoom In: An Introduction to Circuits', which is the first piece in a series of articles about the idea of circuits which we'll discuss in a second, and also the very recent March 2021 article 'Multimodal Neurons in Artificial Neural Networks'. I've obviously read both of those and they've got lots of beautifully carefully designed images, so potentially worth checking out if listeners really want to understand all of this thoroughly. They're both focused on the issue of interpretability in neural networks, visual neural networks in particular, and I would lie to you: those articles were kind of pushing my understanding and I didn't follow them completely, but maybe Chris can help me get it.

可解释性的宏观动机 Big picture motivation for interpretability

Host

不过首先,你能解释一下这条研究路线旨在解决什么问题吗?一个大方向上的问题。

First though, can you explain what problem this line of research is aiming to solve? A kind of big picture.

Chris Olah

在过去几年里,神经网络已经能够完成所有那些人类不知道如何直接编写计算机程序来完成的任务。我们无法编写一个程序来分类图像,但我们可以训练一个神经网络来创建这样一个程序。我们无法直接编写一个程序来高精度地翻译文本,但我们可以训练一个神经网络,它比我们能编写的任何程序都要好得多。我一直觉得,一个亟待回答的问题是:这些模型是如何做到这些我们不知道如何做的事情的?我研究的问题——人们可能会争论可解释性到底是什么——但我感兴趣的问题是:这些系统是如何完成这些任务的?它们内部发生了什么?想象一下,如果某个外星生物降临并能做这些事情,每个人都会争先恐后地去弄清楚它是如何做到的。生物学家们会为了研究这个外星生物而互相争斗。或者想象一下,我们在 2012 年发现了一个漂浮在互联网上的二进制文件,它能做所有这些事情;每个人都会争相去逆向工程它。所以在我看来,所有这些工作都在呼唤我们去回答的问题是:这些系统内部到底发生了什么?

In the last couple of years, neural networks have been able to accomplish all these tasks that no human knows how to write a computer program to do directly. We can't write a computer program to classify images, but we can train a neural network to create a program that can. We can't write a program to translate text highly accurately, but we can train a neural network that does it much better than any program we could have written. It's always seemed to me that the question crying out to be answered is: how is it that these models are doing these things that we don't know how to do? The question I study—and people might debate exactly what interpretability is—but the question I'm interested in is: how do these systems accomplish these tasks? What's going on inside them? Imagine if some alien organism landed and could do these things; everybody would be rushing to figure out how it was doing things. Biologists would be fighting each other for the right to study it. Or imagine we discovered some binary floating on the internet in 2012 that could do all these things; everybody would rush to reverse engineer it. So it seems to me that the thing calling out in all this work is: what in the wide world is going on inside these systems?

不理解系统的后果 Problems of not understanding systems

Host

不理解它有什么问题?或者说,理解它有什么好处?

What are the problems with not understanding it? Or what would be the benefits of understanding it?

Chris Olah

当我们在不了解系统如何运作或为何如此运作的情况下部署它们时,尤其是在高风险场景或影响人们生活的系统中,我感到担忧。我们真的不知道如何推理它们在其它或意外情况下的行为。在某种程度上,你可以通过测试来规避这个问题——我们在各种担心的案例和不同数据集上测试它们——但这只覆盖了数据集中的案例或你明确想到要测试的情况。我认为我们对这些系统将如何行为存在很大的不确定性。特别是随着它们变得更加强大,你不得不开始担心:如果它们在某种意义上做了正确的事情,但出于错误的原因呢?它们实现了正确的行为,但底层的算法实际上只是在试图从你那里获取奖励,而不是帮助你,或者它可能依赖于你没有意识到的有偏见的东西。我认为能够理解这些系统是解决这些担忧的一个非常重要的方法。

I feel worried about deploying systems, especially in high-stakes situations or systems that affect people's lives, when we don't know how they do the things they do or why they're doing it, and really don't know how to reason about how they might behave in other situations or unanticipated situations. To some extent, you can get around this by testing the systems—we test them in all sorts of cases we're worried about, on different datasets—but that only covers the cases in your datasets or that you explicitly thought to test for. I think we have a great deal of uncertainty about how these systems will behave. Especially as they become more powerful, you have to start worrying: what if in some sense they're doing the right thing but for the wrong reasons? They're implementing correct behavior, but the underlying algorithm is just trying to get reward from you rather than trying to help you, or maybe it's relying on biased things you didn't realize. I think being able to understand these systems is a really important way to address those concerns.

预测行为与未知未知 Predicting behavior and unknown unknowns

Host

所以理解它如何运作可能会让预测它在未来场景或场景变化时的表现变得更容易。如果你不理解,你就是在盲目飞行,也许情况会发生变化——他们管这个叫什么,领域变化?

So understanding how it's doing what it's doing might make it easier to predict how it will perform in future situations or if the situation changes. If you have no understanding, you're flying blind, and maybe the circumstance could change—what do they call it, change of domain?

Chris Olah

是的,你可能会担心分布偏移,尽管我认为这并没有充分强调我最关心的问题。这听起来像是一个非常技术性的鲁棒性问题。我更关心的是:现代语言模型有时会对你说谎。有些问题它们在某种意义上知道正确答案——如果你以正确的方式提问,它们会给出正确答案,但在其他上下文中它们不会给出正确答案。我认为这是一个有趣的缩影,反映了这样一个世界:这些模型有能力做事,但试图实现的目标与你想要的不同,并以意想不到的方式造成问题。另一种表述方式是:未知的安全问题是什么?部署这些系统的未知未知数是什么?如果你预料到一个问题,你可以测试它。但随着这些系统变得越来越强大,它们有时会以不同的方式突然改变行为,存在所有这些未知的未知数。你怎么能指望抓住它们呢?我认为这是我研究这些系统的主要动机之一——研究它们如何能让事情变得更安全。

Yes, you might worry about distributional shift, although I think that doesn't emphasize enough what my biggest concern is. It makes it sound like a very technical robustness issue. What I'm concerned about is something more like: modern language models will sometimes lie to you. There are questions to which they know the right answer in some sense—if you pose the question the right way, they give the correct answer, but they won't give the right answer in other contexts. I think that's an interesting microcosm for a world where these models are capable of doing things but are trying to accomplish a goal different from the one you want, and in unanticipated ways cause problems. Another way to frame this is: what are the unknown safety problems, the unknown unknowns of deploying these systems? If you anticipate a problem, you can test for it. But as these systems become more capable, they will sometimes abruptly change their behavior in different ways, and there are all these unknown unknowns. How can you hope to catch those? I think that's a big part of my motivation for studying these systems—how studying them can make things safer.

神经网络工作原理回顾 Quick reminder of how neural networks work

Host

让我们快速提醒听众神经网络是如何工作的。我们之前讲过,但基本上你可以想象信息在一堆节点之间流动,节点之间相互连接。哪些节点与哪些节点相连是由学习过程决定的,这些神经元之间的权重也是由学习决定的。当一个神经元获得足够的正输入时,它倾向于激活并将信号传递给下一层的神经元。信息被处理并在神经元之间传递,直到另一端输出答案。这个描述对于接下来的内容来说足够模糊或准确吗?

Let's give listeners a quick reminder of how neural networks work. We've covered this before, but basically you can imagine information flowing between a bunch of nodes, and the nodes are connected to one another. Which nodes are connected to which is determined by the learning process, and the weightings between these neurons are also determined by learning. When a neuron gets enough positive input, it tends to fire and pass a signal to the next neurons in the next layer. Information is processed and passed between neurons until it spits out an answer at the other end. Is that a sufficiently vague or accurate description for what comes next?

Chris Olah

如果我要稍微挑剔一下,通常哪些神经元连接、哪些不连接是不变的——这通常是静态的。权重在训练过程中演化。

If I was going to nitpick slightly, usually which neurons connect and which don't doesn't change—that's usually static. The weights evolve over training.

Host

所以权重可以降到很低的水平而不是消失?

So the weights can just go to a low level rather than disappear?

Chris Olah

是的,我明白了。

Yes, okay, I understand.

可解释性简介 Introduction to Interpretability

Host

可解释性在机器学习圈子里一直是个大事,至少对一部分人来说是这样。你们研究这个多久了?

Interpretability and this has been a pretty big deal among ML folks or at least some ML folks. How long have you all been at this?

Chris Olah

天哪,我研究可解释性大概有七年了,断断续续的。中间也穿插着做过一些其他机器学习的事情,但这始终是我的主业。最初只有一小部分人对这些问题感兴趣,但这些年我很幸运地找到了许多同样对这些课题充满热情的合作伙伴,更广泛地说,围绕这个领域已经形成了一个不小的圈子,虽然他们采用的方法往往和我不同。但现在已经有一个相当规模的领域在尝试用不同方式理解神经网络了。

Goodness, I've been working on interpretability for about seven years, on and off. I've worked on some other ML things interspersed between it, but it's been the main thing I've been working on. Originally it was a pretty small set of people who were interested in these questions, but over the years I've been really lucky to find a number of collaborators who are also excited about these types of questions, and more broadly a large field has grown around this, often taking different approaches than the ones that I'm pursuing. But there's now a non-trivial field working on trying to understand neural networks in different ways.

Host

这些年你尝试过哪些不同的方法?

What different approaches have you tried over the years?

Chris Olah

这些年我尝试过很多方法,其中很多都没成功。老实说,事后看来原因往往很明显。我曾尝试过一种拓扑学方法,试图理解神经网络如何弯曲数据,但这方法只对最微不足道的系统有效,无法扩展。然后我又尝试了表示降维的方法,但效果也不太好,至少对我来说是这样。长话短说,我最终主要采用的方法简单得近乎愚蠢:就是去理解每个神经元在做什么,所有神经元如何连接在一起,以及这如何产生行为。真正神奇的是,当你开始理解不同神经元的功能时,你实际上就能从权重中读出算法。我喜欢用一个有点计算机科学味道的类比,可能对某些读者来说有点难懂,但我经常把神经元看作计算机程序中的变量或汇编中的寄存器,把权重看作代码或汇编指令。如果你想理解代码,就需要知道变量存储了什么;一旦理解了,你就能看到计算机程序运行的算法,这才是真正令人兴奋和了不起的地方。顺便说一句,我在这里描述的一切其实是很多人的成果,我非常幸运能有这么多出色的合作者一起研究。

I've tried a lot of things over the years, and a lot of them haven't worked. Honestly, in retrospect it's often been obvious why. I tried at one point this approach of going in the sort of topological approach of trying to understand how neural networks bend to data, but that doesn't scale to anything beyond the most trivial systems. Then I tried this approach of looking at how to do dimensional reduction of representations, but that doesn't work super well either, or at least didn't for me. To cut a long story short, the thing that I have sort of ended up settling on primarily is this almost stupidly simple approach of just going and trying to understand what each neuron does and how all the neurons connect together and how that gives rise to the behavior. The really amazing thing is that as you start to understand what different neurons are doing, you actually start to be able to read algorithms off of the weights. There's a slightly computer science heavy analogy that I like, which might be a little bit hard for some readers to follow, but I often think of neurons as being like variables in a computer program or registers in assembly, and I think of the weights as being like code or assembly instructions. You sort of need to understand what the variables are storing if you want to understand the code, but once you understand that, you can actually see the algorithms that the computer program is running, and that's where things become really exciting and remarkable. I should add by the way about all of this that I'm sort of describing this, but I should say that all of this has been the result of lots of people, and I'm just incredibly lucky to have had a lot of really amazing collaborators working on this with me.

可解释性进展 Progress in Interpretability

Host

那么,多亏了这些工作,我们现在能做哪些 10 年前做不到的事情?

So what can we do now that we couldn't do 10 years ago thanks to all of this work?

Chris Olah

我认为我们确实能够理解神经网络中很大一部分的工作原理。我们可以逆向工程神经网络的某些部分,理解得如此透彻,以至于可以手动编写权重。也就是说,你可以拿一个所有权重都设为零的神经网络,然后手动写入权重,重新实现一个神经网络,让它做和原来那一小块神经网络完全相同的事情。你甚至可以把这一块切进之前的网络,替换掉你已经理解的那部分。所以我认为这确实是一个理解系统的极高水准。我想,如果有时我把可解释性想象成类似细胞生物学的东西,就像试图理解细胞一样,我觉得我们可能开始达到能够理解细胞中一个小细胞器的程度,并且真的能把它搞清楚。

Well, I think we can genuinely understand how large chunks of neural networks work. We can actually reverse engineer chunks of neural network and understand them so well that we can go and hand-write weights. So you just take a neural network where all the weights are set to zero and you write by hand the weights, and you can go and reimplement a neural network that does the same thing as that little chunk of the neural network did. You can even slice it into a previous network and replace the part that you understood. So I think that is really a very high standard for understanding systems. I guess if I sometimes like to imagine interpretability as being a little bit like cellular biology or something, where you're trying to understand the cell, and I feel like maybe we're starting to get to the point where we can understand one small organelle in the cell or something like this, and we really can nail that down.

Host

好的,所以我们从拥有一个黑箱,变成了拥有一个你可以手动重新设计的机器,至少是部分可以。

Okay, so we've gone from kind of having a black box to having a machine that you could potentially redesign manually, or at least parts of it.

Chris Olah

是的,我认为主要的故事,至少从电路方法来看,是我们能处理这些小部分,它们是小块,但我们能完全理解它们。

Yeah, I think the main story, at least from the circuits approach, is that we can take these small parts, and they're small chunks, but we can sort of fully understand them.

特征与电路 Features and Circuits

Host

所以这里一个非常关键的概念,或者说两个非常关键的概念,是特征和电路。你能向观众解释一下它们是什么,以及如何提取它们吗?

So a really key concept, or two really key concepts here, are features and circuits. Can you explain for the audience what those are and how you pull them out?

Chris Olah

是的。当我们谈论特征时,大多数情况下(虽然不总是)指的是单个神经元。比如你可能有一个神经元,当出现曲线时它会响应,或者对线条有响应,或者对颜色变化有响应。稍后我们可能会谈到对蜘蛛侠或世界区域有响应的神经元,所以它们也可以是非常高层的。而电路是神经网络的一个子图。节点就是特征。节点可能是一个曲线检测器和一堆线条检测器,然后你有它们之间如何连接的权重,这些权重就是实际运行的计算机程序,它从早期特征构建出后期特征。

Yeah. When we talk about features, we most often, though not always, are referring to an individual neuron. So you might have an individual neuron that does something like responds when there's a curve present, or responds to a line, or responds to a transition in colors. Later on, we'll probably talk about neurons that respond to Spider-Man or respond to regions of the world, so they can also be very high-level. And a circuit is a subgraph of a neural network. So the nodes are features. The node might be a curve detector and a bunch of line detectors, and then you have the weights of how they connect together, and those weights are sort of the actual computer program that's running, that's going and building later features from the earlier features.

Host

好的,所以特征是一些较小的东西,比如线条和曲线等等,而电路则是把这些特征组合起来,试图弄清楚整体图像是什么。

Okay, so features are kind of smaller things like lines and curves and so on, and then a circuit is something that puts together those features to try to figure out what the overall picture is.

Chris Olah

电路是一个部分程序,它从早期特征构建后期特征。

A circuit is a partial program that's building the later features from the earlier features.

Host

啊,好的,所以电路是特征之间的连接,它们有助于从早期特征构建后期特征。我明白了。好的,所以我可以想象,在神经网络的一个层中,有不同的神经元在提取特征,然后它们进入下一层的方式,这些已识别特征之间的权重组合,以及它们如何推进到下一层中识别的特征,这就是一个电路。

Ah, okay, so circuits are our connections between features and they contribute to building the later features from the earlier features. I see. Okay, so I can imagine within say a layer of a neural network, you've got different neurons that are picking up features, and then the way that they go onto the next layer, the combinations of weights between these different features that have been identified and how they push forward into features that are identified in the next layer, that's a circuit.

Chris Olah

是的,不过我们通常关注的是紧密连接在一起的神经元子集。比如你可能对一层中的曲线检测器是如何由前一层的特征构建的感兴趣,然后前一层的特征中有一个子集与它紧密交织,你就可以研究那个子图。你也可以跨越多个层。比如,实际上有一个非常漂亮的用于检测狗头的电路。我知道把它描述为“漂亮”听起来有点疯狂,但我来描述一下这个算法,因为我觉得它真的很优雅。在 Inception V1 中,实际上有两个不同的路径,分别用于检测朝左的狗头和朝右的狗头,并且在这个过程中它们相互抑制。

Yeah, although often we look at a subset of neurons that are tightly connected together. So you might be interested in how the curve detectors in one layer are built from features in the previous layer, and then there's a subset of features at the previous layer that are tightly intertwined with that, and then you could look at that and you're looking at a smaller subgraph. You might also go over multiple layers. So you might look at how there's actually this really beautiful circuit for detecting dog heads. I know it sounds crazy to go and describe that as beautiful, but I'll describe the algorithm because I think it's actually really elegant. So in Inception V1, there's actually two different pathways for detecting dog heads that are facing to the left and dog heads that are facing to the right, and then along the way they mutually inhibit each other.

理解神经网络特征与电路 Understanding Neural Network Features and Circuits

Host

每一步它都会构建一个更好的狗头检测器,面向每个方向,并且让相反方向的检测器相互抑制,所以它有点像在说,一个狗头只能朝左或朝右。最后,它把两者联合起来,创建一个姿态不变的狗头检测器,无论狗头朝左还是朝右都能激活。所以,我猜你有一些神经元对应特征,比如一个神经元在出现类似毛皮的东西时激活,另一个神经元在出现特定形状的曲线时激活,然后它们被连接在一起,要么共同激活,要么相互抑制,从而在更高抽象层次上指示一个特征,比如这是一个狗头,再广义地说,这是某种狗。这就像是下一层。大致对吗?

Every step it goes and builds a better dog head facing each direction and has it so that the opposite one inhibits it, so it's sort of saying, you know, a dog head can only be facing left or right. And then finally at the end it goes and unions them together to create a dog head detector that is pose-invariant, that is willing to fire both for a dog head facing left and a dog head facing right. Okay, so I guess you've got neurons that kind of correspond to features, so a neuron that fires when there's what appears to be fur, and then a neuron that fires when there appears to be a curve of a particular kind of shape, and then a bunch of them are linked together and they either fire together or inhibit one another to indicate like a feature at the next, like at a higher level of abstraction, like this is a dog head, and then more broadly you say like this is this kind of dog. It would be like the next layer. Is that kind of right?

Chris Olah

对,结构上是对的。

Yeah, that's structurally right.

Host

好的,有意思。想出用来识别神经元对应什么特征的技术方法难吗?以及你如何把应该被视为一个电路的所有不同连接和权重整合起来?

Okay, yeah, interesting. Was it hard to come up with the technical methods that you use to identify what feature a neuron corresponds to, and how do you pull together all of the different connections and weights that should be seen as functioning as a circuit?

Chris Olah

是的,有一个有趣的现象:神经网络研究者经常查看第一层的权重。实际上,很多论文都会看第一层的权重,原因在于这些权重连接的是图像中的红、绿、蓝通道(如果你做视觉任务的话),所以这些权重很容易解释,因为你知道这些权重的输入是什么。但你几乎看不到有人用同样的方式查看模型中其他层的权重。我认为原因是他们不知道那些权重的输入和输出是什么。所以,如果你想研究除了绝对输入层或绝对输出层之外的权重,你需要有某种技术来理解进出这些权重的神经元是什么。有几种方法可以做到。一种方法是查看我们所谓的数据集示例,就是向神经网络输入大量示例,找出那些导致神经元激活的示例。这可能是非常神经科学的方法。但我们发现的另一种非常有效的方法是优化输入。我们称之为特征可视化。你优化输入,就是做梯度下降,生成一张图像,让神经元非常强烈地激活。这样做的好处是它把相关性和因果关系分开了。在生成的图像中,你知道图像中的一切之所以存在,是因为它们导致了神经元的激活。然后我们经常使用这些特征可视化,既作为理解神经元行为的线索,也经常把它们当作变量名来用。所以,与其说一个神经元是“mixed 4 c447”,我们不如说有一张图像刺激了它,它是一个汽车检测器,看到一张汽车图像,这样跟踪所有神经元并推理它们如何连接就容易多了。

Yeah, so there's this interesting thing where neural network researchers often look at weights in the first layer. It's actually very common to see papers where people look at the weights in the first layer, and the reason is that those weights connect to red, green, and blue channels in an image if you're doing vision, and so those weights are really easy to interpret because you know what the inputs to those weights are. And you almost never see people look at weights anywhere else in the model in the same kind of way. I think the reason is that they don't know what the inputs and the outputs to those weights are. So if you want to study weights anywhere other than the absolute input or maybe the absolute output, you need to go and have some technique for understanding what the neurons that are going in and out of those weights are. There are a number of ways you can do that. One thing you could do is just look at what we call dataset examples, just feed lots of examples through the neural network and look for the ones that cause neurons to fire. That would be a very maybe neuroscience approach. But another approach that we found is actually very effective is to optimize the input. We call this feature visualization. You optimize the input, you just do gradient descent to go and create an image that causes the neuron to fire really strongly. The nice thing about that is it separates correlation from causation. In that resulting image, you know that everything that's there is there because it caused the neuron to fire. And then we often use these feature visualizations both as a useful clue for understanding what the neuron's doing, but we also often just use them sort of like variable names. So rather than having a neuron be, you know, mixed 4 c447, we say we have this image that stimulates it, it's a car detector, see an image of a car, and that makes it much easier to go and keep track of all the neurons and reason how they connect together.

Host

所以,看起来你有这样一个方法:你选择一个神经元,或者选择一些连接组成的电路,然后试图找出什么样的图像能最大程度地激活那个神经元或电路。你提到了梯度下降,我理解就是选一张噪声图像,然后一点一点地调整,让它越来越激活,直到通过这个迭代过程得到一张几乎最优设计的图像,能强烈地激活那个神经元。所以如果那个神经元是用来检测毛皮的,你就会找出这个神经元能识别的典型毛皮图案。对吗?

So okay, so it seems like you've got this method where you would like choose a neuron or you choose some combination of connections that forms a circuit and then you kind of try to figure out what image is going to maximally make that neuron or that circuit fire. And you said gradient descent, which as I understand it is like you choose a noise image and then you just kind of inch bit by bit to get it to fire more and more until you've got an image that's pretty much optimally designed through this iterative process to really slam that neuron really hard. So if that neuron is there to pick up fur, then you'll figure out the archetypal fur thing that this neuron is able to identify. Is that right?

Chris Olah

对,然后一旦你理解了特征,就可以用它作为脚手架来理解电路。

Yeah, and then once you understand the features, you can use that as a scaffolding to then understand the circuit.

Host

我明白了。好的,我觉得直到我为这期节目做研究时我才意识到,你们有一个显微镜网站,至少 OpenAI 在托管它,你可以在这个大型图像识别神经网络中浏览,挑选出许多不同的神经元,看看什么样的图像会让它们激活。你会说,哦,那是识别这种形状的东西,或者这个神经元对应这种毛皮颜色,然后你可以看到它如何以一种相当复杂的方式在网络中流动。这真的是在拆解这台机器,看看每个小部件是做什么的。

I see. Okay, so yeah, I think I didn't realize this until I was doing research for this episode, but you've got this microscope website, at least OpenAI is hosting it, where you can work through this big image recognition neural network and pick out many different neurons and see what kinds of images cause them to fire. You'll be like, oh that's something that's identifying this kind of shape, or this is a neuron that corresponds to this kind of fur color, and you can see how this flows through the network in quite a sophisticated way. It's really pulling apart this machine and seeing what each of the little pieces does.

Chris Olah

是的,显微镜非常棒。它让你可以查看任何你想要的神经元。我认为它实际上指向了研究神经网络的一个更深层的优势,那就是我们都可以查看完全相同的神经网络。所以可以有这种标准模型,其中每个神经元都和我正在研究的模型中的神经元完全相同,然后我们可以让一个组织,这里就是 OpenAI,一次性创建一个资源,方便引用该模型中的每个神经元。然后每个研究该模型的研究人员都可以通过导航到一个 URL 轻松查看它的任意部分。

Yeah, microscope is wonderful. It allows you to go and look at any neuron that you want. And I think it actually points to a sort of deeper underlying advantage or thing that's nice about studying neural networks, which is we can all look at exactly the same neural network. So there could be these sort of standard models where every neuron is exactly the same as another model, as the model that I'm studying, and we can just go and have one organization, in this case OpenAI, go and create a resource once that makes it easy to go and reference every neuron in that model. And then every researcher who's studying that model can easily go and look at arbitrary parts of it just by navigating to a URL.

Host

对了,为什么人们,可能人们看过 Deep Dream,如果他们看过这些文章,可能会看到那些激活特定特征神经元的典型图像、形状和纹理,但颜色总是非常奇怪。它产生了一种超现实的 Deep Dream 效果。我本来想,如果有一个神经元对毛皮激活,那么它看起来应该像猫的毛皮,或者如果是商店招牌神经元,就应该像商店招牌。但实际上它们看起来像疯狂的迷幻体验。这是为什么?

Yeah, by the way, why do people, so people have probably seen Deep Dream and they might have seen if they've looked at any of these articles, the archetypal images and shapes and textures that are firing these particular feature neurons, but the colors are always so super weird. It produces this kind of surrealist Deep Dream thing. I would think if you had a neuron that was firing for fur, then it would actually look like a cat's fur, or a shop sign if it's a shop sign neuron. But in fact they look like this crazy psychedelic experience. Is there some reason why that's the case?

Chris Olah

你知道吗,在我们发表 Deep Dream 之后,我们收到了一些我认为相当严肃的神经科学家的邮件,问我们是否愿意和他们一起研究,能否用 Deep Dream 来解释人们吸毒时的体验。所以我认为这个类比确实引起了人们的共鸣。但为什么会有那些颜色呢?我认为有几个原因。主要原因是你在创造一个超常刺激。你在创造最能引起神经元激活的刺激。而超常刺激通常看起来很奇怪。

You know, after we published Deep Dream, we got a number of emails from, I think, quite serious neuroscientists asking us if we wanted to go and study with them whether we could explain experiences that people have when they're on drugs using Deep Dreams. So I think that this analogy is one that really resonates with people. But why are there those colors? Well, I think there's a few reasons. I think the main one is you're creating a super stimulus. You're creating the stimulus that most causes the neuron to fire. And you know, the super stimulus often looks weird.

特征可视化为何怪异 Why feature visualizations look weird

Chris Olah

这并不奇怪,因为刺激会推动颜色走向极端,而这正是神经元在寻找的最极端版本。还有一个更微妙的答案:假设你有一个曲线检测器。实际上,神经元在做的一件事就是寻找曲线两侧的颜色变化,因为如果两侧颜色不同,就意味着信号更强。事实上,这在网络早期的线条检测器中很普遍。如果线条两侧颜色不同,线条检测器会激发得更强烈,所以最大刺激就是任何颜色差异。它甚至不在乎两种颜色是什么,只想看到颜色差异。所以它们看起来奇怪的原因就是这些特征检测器主要是在捕捉事物之间的差异,因此它们想要非常强烈的颜色对比。具体是什么颜色并不重要,所以它不必看起来像真正的曲线,只是被捕捉到的颜色梯度。当然,这取决于具体情况,也取决于具体的神经元。还有一个更技术性的原因:神经网络在训练时通常会使用一些色调旋转,有时甚至很多,这意味着图像的颜色在神经网络看到之前会被随机打乱。这样做的目的是让神经网络更鲁棒,但也让它们对特定颜色不那么敏感。实际上,你可以通过特征可视化中看到的颜色模式来判断神经网络训练时使用了多少色调旋转。

It's sort of not surprising that it's going to push colors to very extreme regimes, because that sort of is the most extreme version of the thing that the neuron is looking for. There's also a slightly more subtle version of that answer: suppose that you have a curve detector. It turns out that actually one thing the neuron is doing is just looking for any change in color across the curve, because if there are different colors on both sides, that means it's stronger. In fact, this is generally true of line detectors very early on in networks. A line detector fires more if there are different colors on both sides of the line, so the maximal stimulus for that just has any difference in colors. It doesn't even care what the two colors are; it just wants to see a difference in colors. So the reason they can look kind of weird is just that these feature detectors are mostly picking up differences between things, and so they want a really stark difference in color. The specific color that the thing happens to be doesn't matter so much, so it doesn't have to look like an actual curve; it's just the color gradient that's being picked up. Well, it depends on the case, and it depends a lot on the particular neuron. There's one final reason, which is a bit more technical: often neural networks are trained with a little bit of hue rotation, sometimes a lot of hue rotation, which means that the colors are randomly shuffled a little bit before the neural network sees them. The idea is that this makes the neural network a bit more robust, but it makes neural networks a little less sensitive to the particular color they are seeing. You can actually often tell how much hue rotation a neural network was trained with by exactly the color patterns you see when you do these feature visualizations.

从特征可视化中学到什么 What we can learn from feature visualization

Host

那么具体来说,通过这些方法,我们能了解到这样的神经网络系统在想什么或如何思考呢?

So what concretely can we learn about what or how a neural network system like this is thinking using these methods?

Chris Olah

从某种意义上说,你可以相当全面地理解它,但只理解了一小部分。所以你可以理解几个神经元在做什么,以及它们如何连接。你可以达到这样的程度:你的理解只是基础数学,你可以用基本逻辑推理出它在做什么。挑战在于这只是模型的一小部分,但你确实理解了那一小部分中运行的算法以及它的作用。

Well, in some sense you can quite fully understand it, but you understand a small fraction. So you can understand what a couple of neurons are doing and how they connect together. You can get it to the point where your understanding is just basic math, and you can reason through what it's doing with basic logic. The challenge is that it's only a small part of the model, but you literally understand what algorithm is running in that small part and what it's doing.

通用性:跨模型的相同特征与电路 Universality: same features and circuits across models

Host

从那些文章来看,你似乎认为我们可能学到了一些关于神经网络如何工作的非常基础的东西,不仅关乎这些思考机器,也可能关乎大脑中的神经网络如何运作,甚至可能是它们必然的运作方式。你认为我们可能学到了哪些东西?证据是什么?

It sounded like from those articles that you think we've potentially learned some really fundamental things here about how neural networks work, and potentially not only how these thinking machines work but also how neural networks might function in the brain, and maybe how they always have to function. What are those things you think we might have learned, and what's the evidence for them?

Chris Olah

一个令人着迷的现象是,相同的模式、相同的特征和电路在不同模型中反复出现。你可能会认为每个神经网络都是独一无二的雪花,它们做的事情完全不同。那会让理解这些事物的领域变得无聊或非常奇怪。我有时喜欢把可解释性比作解剖学,我们正在解剖这些神经网络,观察它们内部的情况。如果早期解剖学家发现每个生物都有完全不同的解剖结构,没有相似之处,没有像心脏这样普遍存在的器官,那解剖学就会变得无聊。但就像动物由于进化而拥有非常相似的解剖结构一样,神经网络似乎也形成了很多相同的东西,即使你在不同的数据集上训练它们,即使它们有不同的架构。尽管框架不同,但相同的特征和电路会形成。我觉得相同电路形成的事实是最引人注目的部分。相同特征形成已经很酷了——神经网络在学习这些理解视觉或图像的基本构建块。但更厉害的是,它实际上在学习相同的权重,将相同的神经元连接在一起。我们称之为普遍性,这非常疯狂。当你开始发现这样的事情时,很容易会想,也许这些相同的东西也在人类中形成。也许这实际上是某种基础性的东西——也许这些模型正在发现以非常基本的方式划分我们对图像理解的基本视觉构建块。事实上,对于其中一些东西,我们已经在人类中发现了。一些低层次的视觉现象似乎与神经科学的结果相呼应。在我们最近的一些工作中,我们发现了以前只在人类中观察到的东西:多模态神经元。我们稍后会讨论多模态神经元。

One of the fascinating things is that the same patterns and the same features and circuits form again and again across models. You might think that every neural network is its own special snowflake, and they're all doing totally different stuff. That would make for a boring or very strange field of trying to understand these things. I sometimes like to think of interpretability as being like anatomy, and we're dissecting these neural networks, looking at what's going on inside them. If early anatomists found that every organism had a totally different anatomy, with no similarities, there would be nothing like a heart that exists in lots of them, that would create a boring field of anatomy. But just like animals have very similar anatomies due to evolution, it seems like neural networks actually have a lot of the same things forming, even when you train them on different datasets, even when they have different architectures. Even though the scaffolding is different, the same features and the same circuits form. I find that the fact that the same circuits form is the most remarkable part. The fact that the same features form is already pretty cool—the neural network is learning these same fundamental building blocks of understanding vision, or understanding images. But then it's literally learning the same weights connecting the same neurons together. We call that universality, and that's pretty crazy. It's really tempting, when you start to find things like that, to think that maybe these same things form also in humans. Maybe it's actually something fundamental—maybe these models are discovering the basic building blocks of vision that slice up our understanding of images in a very fundamental way. In fact, for some of these things, we have found them in humans. Some of these lower-level vision things seem to mirror results from neuroscience. In some of our most recent work, we've discovered something that was previously only seen in humans: these multimodal neurons. We'll talk about multimodal neurons in just a second.

通用性证据及与人类视觉的联系 Evidence for universality and connection to human vision

Host

所以你的意思是,你训练了许多不同的神经网络来识别图像,用不同的图像和不同的设置,但每次你都会注意到这些共同特征:有识别特定形状、曲线和特定纹理的东西,而且组织方式是有特征检测神经元,然后有从低级特征到高级特征的电路。这基本上每次都发生。然后我们有一些证据表明,这与我们了解的人类大脑处理视觉信息的方式相符。所以这算是一个初步案例,尽管还不确定,这些普遍特征似乎总是会出现。

So it sounds like you're saying you train lots of different neural networks to identify images, and you train them on different images and with different settings, but every time you notice these common features: you've got things that are identifying particular shapes and curves and particular kinds of textures, and also things are organized in such a way that there are feature-detecting neurons and then organized circuits for moving from lower-level features to higher-level features. That just happens every time basically. And then we have some evidence that this is matching things we've learned about how humans process visual information in the brain as well. So this is like a prima facie case, although it's not certain, that these universal features are kind of always going to show up.

Chris Olah

是的,而且几乎……

Yeah, and it's almost...

视觉模型中的常见特征 Common features in vision models

Host

人们很容易认为视觉中存在一些基本元素,比如线条,这些元素在推理图像时总是存在的。有没有可能我们得到这些共同特征和回路是因为我们在非常相似的图像上训练网络——猫、狗、草地、风景——而不是视觉本身的基本特性?

It's tempting to think that there are fundamental elements of vision, like lines, that are always present when reasoning about images. Is it possible that we're getting these common features and circuits because we train networks on very similar images—cats, dogs, grass, landscapes—rather than something fundamental about vision?

Chris Olah

你可以看看在非常不同的数据集上训练的模型。有一个叫 Places 的数据集,是关于识别建筑物和场景的,里面没有狗和猫。结果发现,越往高层,特征差异越大。在后面的层中,Places 模型有与不同视觉视角和建筑物观察角度相关的特征,这在 ImageNet 模型中看不到,反之亦然。但似乎确实有很多基本的东西,尤其是在早期视觉中,在两个数据集中都会形成。

You could look at models trained on pretty different datasets. There's a dataset called Places, which is about recognizing buildings and scenes, and those don't have dogs and cats. It turns out that as you go higher up, the features are more and more different. In later layers, Places models have features related to different visual perspectives and viewing angles of buildings, which you don't see in ImageNet models, and vice versa. But there do seem to be a lot of things that are fundamental, especially in early vision, that form in both datasets.

Host

那么如果外星人训练了一个模型,或者我们能看到外星人的大脑,它也会显示所有这些共同特征吗?直觉上,线条、边缘、曲线是基本的,所以即使在别的星球上也会有线条。但在高层,一些特征和回路对应于我们世界中存在的特定事物,可能在另一个星球上不存在。

So if aliens trained a model, or if we could look into an alien brain, would it show all these common features? Intuitively, a line, an edge, a curve are fundamental, so even on other planets they'd have lines. But at the high level, some features and circuits correspond to particular things that exist in our world and may not exist on another planet.

Chris Olah

完全正确。这纯粹是推测,但很诱人。我觉得当我们发现的是深层基本的东西,而不是单个系统的任意事实时,研究这些系统的故事在情感上更引人入胜。

That's exactly right. It's purely supposition, but it's tempting to imagine. I find the stories of investigating these systems emotionally more compelling when I think we're discovering deeply fundamental things rather than arbitrary truths for a single system.

Host

你发现了线条的基本原始性质。但更令人兴奋的是发现人们以前没见过的东西,比如似乎在所有模型中都会形成的高低频检测器。你说的低频和高频是什么意思?

You've discovered the fundamental primitive nature of a line. But it's more exciting when we discover things people haven't seen before, like high-low frequency detectors that seem to form in all models. What do you mean by high frequency and low frequency?

Chris Olah

具有大量纹理和锐利过渡的模式可能是高频,而低频图像则非常平滑或失焦。所以它检测的是图像某一部分与另一部分之间的锐利程度和差异数量。模型将这些作为边界检测的一部分,使用多种线索。例如,在物体的边界处,背景失焦且低频,而前景则更高频。或者两个相邻物体具有不同纹理时。事后看来,拥有这些特征是有道理的,但之前没有人预测到。发现这些未被预测的东西令人兴奋,这表明这些系统内部有大量东西等待被发现,就像第一次发现线粒体一样。

A pattern with lots of texture and sharp transitions might be high frequency, whereas a low frequency image would be very smooth or out of focus. So it picks up how sharp and how many differences there are within a section of the image versus another part. Models use these as part of boundary detection, using multiple cues. For instance, at the boundary of an object where the background is out of focus and low frequency, while the foreground is higher frequency. Or between two adjacent objects with different textures. In retrospect, it makes sense to have these features, but no one predicted them in advance. It's exciting to discover things that weren't predicted, and it suggests that there's a wealth of things to be discovered inside these systems, like discovering mitochondria for the first time.

Host

你正在琢磨的普遍性猜想是,我们会在人脑中看到这些高低频检测神经元和回路,它们在识别物体中起重要作用。而且每个处理自然图像视觉的神经网络也会拥有它们。

The universality conjecture you're toying with would be that we'll look into the human brain and find these high-low detector neurons and circuits, playing an important role in identifying objects. And that every neural network doing natural image vision will have them as well.

Chris Olah

那会是强版本。而且每个处理自然图像视觉的神经网络都会有它们,它们真的是你到处都能找到的基本东西。

That would be the strong version. And also that every neural network doing natural image vision will have them, and they're really just a fundamental thing you find everywhere.

Host

获得更多普遍性证据的一种方法是检查神经网络是否和人类有相同的视觉错误。如果它们以相同的方式正确和错误,可能表明过程相似。

One way to get more evidence for universality is to check if neural networks have the same visual errors as humans. If they get things right and wrong in the same ways, it might suggest a similar process.

Chris Olah

有几篇论文探索了类似的想法。事实上,当我们讨论多模态神经元的结果时,会有一个例子。

There have been a few papers exploring ideas like that. In fact, when we talk about the multimodal neuron results, we'll have one example.

Host

这正好是深入探讨多模态神经元文章的好时机,这篇文章几周前发表,涉及更先进的视觉检测模型。你在那项工作中学到了什么与以前不同的东西?

That's a perfect moment to dive into the multimodal neuron article, which came out a couple of weeks ago and deals with more state-of-the-art visual detection models. What did you learn in that work that was different from before?

Chris Olah

我们正在研究 OpenAI 的一个名为 CLIP 的模型,你可以大致认为它被训练来为图像生成标题或将图像与其标题配对。所以它不仅仅是一个纯视觉模型;它结合了视觉和语言。

We were investigating a model called CLIP from OpenAI, which you can roughly think of as being trained to caption images or pair images with their captions. So it's not just a pure vision model; it combines vision and language.

CLIP 中的多模态神经元 Multimodal Neurons in CLIP

Chris Olah

分类图像时,它做的事情有些不同,我们在其中发现了许多在本质上截然不同的东西。如果你看低层视觉,实际上很多部分非常相似,这再次证明了通用性。我们在其他视觉模型中找到的许多相同特征,在 CLIP 的早期视觉中也会出现。但到了后期,我们发现了这些极其抽象的神经元,它们与我们之前见过的任何东西都截然不同。这些神经元的一个有趣之处在于,它们能够阅读——它们能识别图像中的文字,并将文字与检测到的物体融合在一起。例如,有一个黄色神经元,它对黄色有反应,但如果你写出“黄色”这个词,它也会激活。实际上,如果你写出黄色物体的单词,比如“柠檬”或“香蕉”,它也会激活。这绝不是你在视觉模型中期望看到的东西。它是一个视觉模型,但某种程度上几乎在进行语言处理,并将它们融合成我们所说的多模态神经元。这种现象在神经科学中也有发现。例如,有一个蜘蛛侠神经元,它既对“蜘蛛侠”这个词有反应,也对蜘蛛侠的图片和画作有反应。这反映了神经科学中一个著名的结果:哈利·贝里神经元或詹妮弗·安妮斯顿神经元,它们同样对人物的照片、画作和名字有反应。因此,这些神经元在某种意义上比我们之前发现的神经元更加抽象,几乎接近概念层面。它们涵盖了极其广泛的主题。事实上,很多神经元——你浏览它们时会觉得,这就像幼儿园或小学低年级的东西。你有颜色神经元、形状神经元、对应季节和月份的神经元、天气神经元、海洋神经元、世界区域神经元、国家领导人神经元。所有这些神经元都具有这种令人难以置信的抽象性质。比如有一个早晨神经元,它对闹钟、清晨时间、煎饼和早餐图片都有反应——各种各样的事物。或者季节神经元,它对季节名称、相关天气类型等都有反应。因此,你拥有所有这些极其多样化的神经元,它们都以这种不同的方式高度抽象。这与我们之前看到的相对具体的神经元(通常对应一种物体类型等)形成了鲜明对比。

Classifying images, it's doing something a little bit different, and we found a lot of things that were really deeply qualitatively different inside it. So if you look at low-level vision, actually a lot of it is very similar, and again, it's actually further evidence for universality. A lot of the same things we find in other vision models occur also in early vision in CLIP. But towards the end, we find these incredibly abstract neurons that are just very different from anything we'd seen before. And one thing that's really interesting about these neurons is they can read — they can go and recognize text in images, and they fuse this together with the thing that's being detected. So there's a yellow neuron, for instance, which responds to the color yellow, but it also responds if you write the word 'yellow' out — that will fire as well. And actually, it'll fire if you write the words for objects that are yellow, so if you write the word 'lemon', it'll fire, or the word 'banana', it'll fire. And this is really not the sort of thing that you expect to find in a vision model. Like, it's a vision model, but it's almost doing linguistic processing in some way, and it's fusing it together into what we call these multimodal neurons. And this is a phenomenon that has been found in neuroscience. So you find these neurons also for people — like there's a Spider-Man neuron that fires both for the word 'Spider-Man', for an image like a picture of the word 'Spider-Man', and also for pictures of Spider-Man, and for drawings of Spider-Man. And this mirrors a really famous result from neuroscience of the Halle Berry neuron or the Jennifer Aniston neuron, which also respond to photos of the person, to drawings of the person, and to the person's name. And so these neurons seem, in some sense, much more abstract and almost conceptual compared to the previous neurons that we found. And they span an incredible wide range of topics. In fact, a lot of the neurons — you just go through them and you're like, it feels like something out of a kindergarten class or an early grade school class. And you have your color neurons, you have your shape neurons, you have neurons corresponding to seasons of the year and months, to weather, to oceans, to regions of the world, to the leader of your country. And all of them have this incredible abstract nature to them. So there's like a morning neuron that responds to alarm clocks, times of the day that are early, to pictures of pancakes and breakfast food — all this incredible diversity of stuff. Or season neurons that respond to the names of the seasons, the type of weather associated with them, and all of these things. And so you have all this incredible diversity of neurons that are all incredibly abstract in this different way. And it just seems very different from the relatively concrete neurons that we were seeing before, that often correspond to a type of object or such.

Host

所以我们从拥有一个能识别出这是头部的图像识别网络,发展到了拥有一个能识别出某物是美丽的、这是一件艺术品、或者这是一幅印象派作品之类的网络。这是一种更高层次的抽象和分组,许多不同种类、不同地方的图像可能具有共同点。

So we've kind of gone from having an image recognition network that can recognize that this is a head, to having a network that can recognize that something is beautiful, or that this is a piece of artwork, or that this is an impressionist piece of art or something. It's a higher level of abstraction and grouping that lots of different images might have in common, of many different kinds in many different places.

Chris Olah

是的,而且这甚至不关乎看到物体本身;而是关乎与物体相关的事物。例如,你在其中一些模型中看到一个巴拉克·奥巴马神经元,它当然对奥巴马的照片有反应,但也对他的名字有反应,对美国国旗有一点反应,还对米歇尔·奥巴马的照片有反应。所以,所有与他相关的事物都会让它稍微激活。

Yeah, and it's not even about seeing the object; it's about things that are related to the object. For instance, you see a Barack Obama neuron in some of these models, and it of course responds to images of Barack Obama, but also to his name, and a little bit to US flags, and also to images of Michelle Obama. And so it sort of is all these things that are associated with him that also cause it to fire a little bit.

Host

你这么一说,听起来它越来越接近人类的行为,因为我们在世界上活动时,不同事物之间就有所有这些关联。当人们说话时,他们所说的话会以不同的权重唤起特定的不同事物,我觉得这就是我推理的方式。语言非常模糊,对吧?很多不同的词有不同的联想,你把它们组合在一起,大脑就会混合出结果。所以它可能具备一定水平的概念推理,接近人类思维所做的。

When you put it that way, it just sounds like it's getting so close to doing what humans do, because kind of what we do when we're going about the world is we've got all of these associations between different things. And then when people speak, the things that they say kind of draw to mind particular different things with different weightings, and that's kind of how I feel like I reason. And it's like language is very vague, right? Lots of different words have different associations, and you smash them together, and then the brain makes a mix of what it will. So it's possible that it's capable of a level of conceptual reasoning that is approaching a bit like what the human mind is doing.

Chris Olah

我非常谨慎地不想把这些描述为关于概念,因为我认为这是一种有争议的描述方式,人们可能会强烈反对。但这样描述它们非常诱人,而且我认为这种框架有很多正确之处。

I feel really nervous to describe these as being about concepts, because I think that's a charged way to describe it that people might strongly disagree with. But it's very tempting to frame them that way, and I think that there's a lot of things about that framing that would be true.

Host

有趣。你脑海中能想到你们发现的最抽象的类别是什么吗?有没有什么特别引人注目、令人惊叹的,它能捕捉到这种分组?

Interesting. Do you know off the top of your head what is the most abstract category that you found? Is there anything that particularly striking and amazing that it can pick up this grouping?

Chris Olah

我想到的一个神经元是心理健康神经元,你可以大致认为,每当图像中出现任何心理健康问题的线索时,它就会激活。这可能是肢体语言,或看起来特别焦虑或紧张的面孔,但也包括很多词语,比如“抑郁”或“焦虑”之类的词,与心理健康相关的药物名称,与心理健康相关的侮辱性词汇——比如“疯狂”之类的词——以及一些看起来有点心理主题的图像。就是这种令人难以置信的广泛事物。它看起来就是一个非常抽象的概念。心理健康并没有一个非常具体的实例,但它却代表了它。

One neuron that jumps to mind for me is the mental health neuron, which you can roughly think of as firing whenever there's a cue for any kind of mental health issue in the image. And that could be body language, or faces that read as particularly anxious or stressed, but also for lots of words like the word 'depression' or 'anxiety' or things like this, the names of medicines associated with mental health, slurs related to mental health — so words like 'crazy' or things like this — images that sort of seem like psychological-themed images a little bit. Just this incredible range of things. And it just seems like such an abstract idea. Like, there's no single very concrete instantiation of mental health, and yet it sort of represents that.

Host

这个模型是如何训练的?我猜它必须有大量的图像样本和大量的标题。那么它是利用标题中的词语来识别与图像相关的概念,然后将它们分组吗?

How is this model trained? I guess it must have a huge sample of images and then just lots of captions. And so is it using the words in the captions to recognize concepts that are related to the images and then group them?

Chris Olah

是的,它试图将图像和标题配对。它会取一组图像和一组标题,然后尝试找出哪些图像对应哪些标题。这带来了各种非常有趣的特性,意味着你可以直接编写标题,然后用它来编程,让它执行任何你想要的图像分类任务。但它似乎也导致视觉侧的特征更加丰富。

Yeah, it's trying to pair up images and captions. So it'll take a set of images and a set of captions, and then it'll try to figure out which images correspond to which captions. And this has all sorts of really interesting properties, where this means that you can just write captions and sort of use that to program it to go and do whatever image classification task you want. But it also just seems to lead to much richer features on the vision side.

多模态神经元与上下文 Multimodal Neurons and Context

Host

这些神经元对应输出中的特定词汇之类的,但我认为实际情况并非如此。例如,存在所有这些区域神经元,它们对应国家或国家的一部分之类的东西。有些非常大,比如有一个整个北半球的神经元,它对针叶林、鹿和熊等事物有反应。我不认为它们的存在是因为对应像“加拿大”这样的特定词汇;而是因为当图像处于特定语境时,会改变人们谈论的内容。所以如果你在加拿大,你更可能谈论枫糖浆;如果你在中国,你更可能提到中国城市或让部分字幕变成中文。所以训练数据是一组对应的图像和字幕,它发展出这些多模态神经元,以便概率性地估计或提高判断语境的能力,从而确定字幕中最可能出现的词汇。这大致正确吗?

These neurons correspond to particular words in the output or something, but I think that's not actually going on. For instance, there are all these region neurons that correspond to things like countries or parts of countries. Some of them are really big, like there's an entire Northern Hemisphere neuron that responds to things like coniferous forests and deer and bears. I don't think those are there because they correspond to particular words like 'Canada'; it's because when an image is in a particular context, that changes what people talk about. So you're much more likely to talk about maple syrup if you're in Canada, or if you're in China, it's much more likely that you'll mention Chinese cities or have some of the caption be Chinese. So the training data was a set of corresponding images and captions, and it's developing these multimodal neurons in order to probabilistically estimate or improve its ability to tell what the context is, to figure out what words are most likely to appear in a caption. Is that broadly right?

Chris Olah

是的,那是我对正在发生的事情的直觉。

Yeah, that would be my intuition about what's going on.

对安全性与偏见的影响 Implications for Safety and Bias

Host

这种多模态神经元的发现对神经网络的安全性或可靠性问题有什么影响吗?

What implications, if any, does this discovery of multimodal neurons have for safety or reliability concerns with these neural networks?

Chris Olah

嗯,我想有几件事。第一个涉及这种“未知的未知”担忧。如果你查看多模态模型或 CLIP 内部,你会发现神经元对应着美国几乎每一个受保护属性。所以不仅存在与性别、年龄、宗教相关的神经元——每种宗教都有神经元——这些区域神经元还与种族紧密相关。有一个用于怀孕的神经元,一个用于父母身份的神经元,它会寻找儿童画之类的东西,还有一个用于心理健康的神经元,另一个用于身体残疾。几乎每个受保护属性都有一个神经元。尽管机器学习社区非常关注机器学习系统中的偏见,但它往往对性别或种族非常警惕。我认为这表明我们应该寻找更广泛的问题。很容易想象 CLIP 会对父母身份或心理健康产生歧视,我认为你不会想到以前的模型能做到这一点。这说明了研究这些东西如何能揭示我们以前不担心、但现在或许应该警惕的问题。

Well, I guess there are a few things. The first one speaks to this unknown unknown type concern. If you look inside the multimodal model or CLIP, you find neurons corresponding to literally every trait, or almost every trait, that is a protected attribute in the United States. So not only do you have neurons related to gender, age, religions—there are neurons for every religion—these region neurons connect closely to race. You have a neuron for pregnancy, a neuron for parenthood that looks for things like children's drawings, a neuron for mental health, and another for physical disability. Almost every single protected attribute has a neuron. Despite the machine learning community caring a lot about bias in machine learning systems, it tends to be very alert with respect to gender or race. I think this shows we should be looking for a much broader set of concerns. It would be very easy to imagine CLIP discriminating with respect to parents or mental health, which I think you wouldn't have thought previous models could. That illustrates how studying these things can surface issues we weren't worried about before and perhaps should now be on the lookout for.

Host

我猜如果你试图告诉 CLIP,“淘气的 CLIP,我们不应该有用于种族或性别的多模态神经元”,然后你删除了它们,我打赌在训练过程中,它会转移这些概念,将它们模糊成其他东西,比如地理神经元或关于行为或个性的东西。它最终会无意中构建出对性别和种族的识别,但方式更难察觉。

I'm guessing if you tried to tell CLIP, 'Naughty CLIP, we shouldn't have multimodal neurons for race or gender,' and then you struck them out, I bet when it was getting trained, it would shift those concepts, blur them into something else like a geographic neuron or something about behavior or personality. It would end up accidentally building in the recognition of gender and race but in a way that's harder to see.

Chris Olah

是的,我认为摆脱这些东西很棘手。而且也不清楚你是否真的想摆脱对这些东西的表征。还有另一个神经元,类似于冒犯性神经元;它对非常糟糕的东西有反应,比如脏话、种族歧视言论、色情内容。一方面,你可能会说我们应该去掉那个神经元——有一个对这些东西有反应的神经元不好。另一方面,那个神经元的部分功能可能是将冒犯性图像与反驳它们的字幕配对,因为有时人们会发布一张图片然后回应它。这正是你希望它做的事情:识别冒犯性内容并反驳它。我个人的直觉是,我们不想要不理解这些东西的模型;我们想要理解它们并做出适当回应的模型。

Yeah, I think it's tricky to get rid of these things. It's also not clear that you actually want to get rid of representation of these things. There's another neuron that is like an offensiveness neuron; it responds to really bad stuff like swear words, racial slurs, pornography. On one hand, you might say we should get rid of that neuron—it's not good to have a neuron that responds to those things. On the other hand, part of the function of that neuron could be to pair offensive images with captions that rebut them, because sometimes people post an image and then respond to it. That is the thing you would like it to do: recognize offensive content and rebut it. My personal intuition is that we don't want models that don't understand these things; we want models that understand them and respond appropriately.

Host

我明白了,所以最好让模型拥有这些,然后可能之后用其他东西来修复造成的伤害,而不是试图告诉模型不要识别冒犯性内容,因为那样你会更加盲目。就像我们不会阻止孩子了解种族主义和性别歧视;我们想教育他们那些是坏的。这是我个人的直觉。

I see, so it's better to have it in the model and then potentially fix the harm down the track using something afterwards, rather than try to tell the model not to recognize offensiveness, because then you're more flying blind. It's like we wouldn't prevent children from ever finding out about racism and sexism; we want to educate them that those are bad. That's my personal intuition.

Chris Olah

是的,有道理。因为那是我们试图对人类做的事情,所以也许我们应该对神经网络做同样的事情,鉴于它们似乎变得越来越像人类。

Yeah, that makes sense. Because that's what we're trying to do with humans, so maybe that's what we should try to do with neural networks, given how human-like they are seemingly becoming.

情感神经元与未来安全担忧 Emotion Neurons and Future Safety Concerns

Host

继续安全问题,这有点推测性,但我们看到这些情绪神经元正在形成。我认为我们应该警惕的是,每当我们看到类似心智理论或社会智能的特征时,那都是我们应该密切关注的东西。我们在其他地方看到了迹象:现代语言模型可以跟踪话语中的多个参与者,跟踪他们的情绪和信念,并为互动写出合理的回应。拥有允许你完成这些任务的能力,拥有像情绪神经元这样的东西,让你能够检测图像中某人的情绪,或者可能错误地推理但尝试推理,这使得想象具有操纵性或更刻意操纵的系统变得更容易。一旦你有了社会智能,那很容易出现。我不认为我们看到的已经达到那个程度,但我预测在不久的将来这将是一个更大的问题,我们应该对此保持警惕。

Continuing on the safety issue, this is a bit more speculative, but we're seeing these emotion neurons form. I think something we should really be on guard for is whenever we see features that look like theory of mind or social intelligence. That's something we should keep a close eye on. We see hints of this elsewhere: modern language models can track multiple participants in a discourse and track their emotions and beliefs, and write plausible responses to interactions. Having faculties that allow you to do those tasks, having things like emotion neurons that allow you to detect the emotions of somebody in an image, or possibly incorrectly reason about them but attempt to reason about them, makes it easier to imagine systems that are manipulative or more deliberately manipulative. Once you have social intelligence, that's a very easy thing to arise. I don't think the things we're seeing are quite there, but I would predict that is a greater issue we're going to see in the not too distant future, and it's something we should keep an eye out for.

可解释性研究的启示 Lessons from interpretability research

Chris Olah

另一件事是,在做可解释性研究时,尤其是如果你的目标是促进安全,你应该警惕的是,你对神经网络内部情况的猜测往往是错误的。你应该让数据自己说话,去观察并尝试理解发生了什么,而不是先入为主,然后寻找证据来支持你的假设。我认为神经元在这方面非常惊人:模型最后一层中大约 5%是关于地理的。如果你事先问我 CLIP 里有什么,我一百年也猜不到其中很大一部分会是地理。所以我认为这是另一个我们应该非常谨慎的教训。

One other thing is that one should generally be on guard for, when doing interpretability research, especially if one's goal is to contribute to safety, is the fact that your guesses about what's going on inside a neural network will often be wrong. You want to let the data speak for itself, look and try to understand what's going on rather than assuming and looking for evidence that backs up your assumptions. I think that neurons are really striking for this: about 5% of the last layer of the model is about geography. If you had asked me to predict what was in CLIP beforehand, I would not have guessed in a hundred years that a large fraction of it was going to be geography. So I think that's just another lesson that we should be really cautious of.

Host

所以你的意思是,尽管你试图弄清楚事情可能如何出错,但你想知道网络实际在关注什么。而且我们对此似乎没有很好的直觉;网络可能专注于与我们预期截然不同的特征和回路。

So you're saying that as much as you're trying to figure out how things might go wrong, you want to know what the network is actually focusing on. And it seems like we don't have such great intuitions about that; the network can potentially be focused on features and circuits that are quite different from what we anticipated.

Chris Olah

是的,或者我有时看到人们试图理解神经网络的方法是,他们猜测里面有什么,然后寻找他们猜测存在的东西。我认为这是一个说明性的例子,说明为什么——尤其是如果你关心的是理解这些未知的未知以及模型正在做的你未预料到且可能引发问题的事情——这里有一大堆我认为人们不会猜到的东西,用那种方法很难捕捉到。

Yeah, or just an approach that I sometimes see people take to try and understand neural networks is they guess what's in there and then they look for the things they guessed are there. I think this is an illustrative example of why, especially if your concern is to understand these unknown unknowns and the things the model is doing that you didn't anticipate and that might cause problems, here's a whole bunch of things that I don't think people would have guessed and that would have been very hard to catch with that kind of approach.

与安全性和可靠性的关联 Relevance to safety and reliability

Host

所以这是多模态神经网络工作与安全性和可靠性可能相关的两种不同方式。还有其他发现吗?

So those are two different ways that the multimodal neural work has been potentially relevant to safety and reliability. Are there any other things that have been thrown up?

Chris Olah

另一件小事,可能不那么直接相关,但值得一提,我认为对一些人来说,是否从事安全研究或优先考虑安全的一个关键点在于这些系统离真正的智能有多近。神经网络只是在欺骗我们,只是看起来自信但并非真正在做的幻觉,还是它们在某种意义上确实在做真正的事情,我们应该担心它们未来的能力?我认为从中应该有一定量的证据,让我们更新方向,认为神经网络正在做一些真正有趣的事情。相应地,如果这对你来说是是否从事安全研究的关键,那么这是你可能更新的另一个证据。

One other small thing, which maybe is less of a direct connection but worth mentioning, is that I think a crux for some people in whether to work on safety or prioritize safety is something like how close these systems are to actual intelligence. Are neural networks just fooling us, just illusions that seem confident but aren't doing the real thing, or are they actually in some sense doing the real thing, and we should be worried about their future capabilities? I think there should be a moderate amount of evidence from this to update in the direction that neural networks are doing something genuinely interesting. Correspondingly, if that's a crux for you regarding whether to work on safety, this is another piece of evidence you might update on.

Host

我想,它们越有能力,越能进行人类那样的推理,那么我们就越应该预期它们会被部署到可能重要的决策过程中。

I guess just the more capable they are and the more they seem capable of doing the kinds of reasoning that humans do, then that should bring forward the date at which we would expect them to be deployed in potentially important decision-making procedures.

Chris Olah

尽管我认为,对于能力方面的证据,人们可能的一种回应是想象模型可能在作弊,想象模型可能并非真正在做真实的事情。所以它看起来在能力上有所进步,但那是一种幻觉。有些人可能会说它实际上根本不理解正在发生的一切。我认为“理解”在那里是一个有争议的词,可能不是很有帮助,但模型在某种意义上只是一种幻觉和把戏,不会带我们走向真正有能力的系统。嗯,我认为很难完全判断,但当我们看到系统内部实现有意义算法的证据,尤其是当我们看到那些被视为人类理解概念证据的东西时,我认为我们确实需要更新一下看法。

Although I think a response that one can have with respect to evidence in the form of capabilities is to imagine ways the model might be cheating and imagine the ways in which the model may not really be doing the real thing. So it sort of appears to be progressing in capabilities, but that's an illusion. Some might say it actually doesn't really understand anything that's going on at all. I think 'understand' is a charged word there and maybe not very helpful, but the model in some sense is just an illusion and a trick and won't get us to things that are genuinely capable. Well, I think it's hard to fully judge, but when we see evidence of the system implementing meaningful algorithms inside them, and especially when we see evidence of things that have been perceived as evidence of human-like understanding of concepts, I think there's somewhat of an update to be had there.

对抗样本与可靠性 Adversarial examples and reliability

Host

文章中提到的另一个安全问题,至少据我理解,是你可能通过修改让神经网络误识别某物。例如,你有一个苹果,它会识别为苹果,然后你贴上一张写着“iPod”的便签,它就识别为 iPod。这是一个重要的可靠性问题吗?也许随着模型变得更复杂,它们可能以更复杂的方式失败?

One other safety concern that came up in the article, at least as I understood it, was that you can potentially get a neural network to misidentify something by modifying it. For example, you had an apple and it would identify as an apple, and then you put a Post-it note with the word 'iPod' on it and it identified it as an iPod. Is that an important reliability concern that maybe as you make these models more complicated, they can fail in ever more sophisticated ways?

Chris Olah

这当然是一个有趣的发现。我认为也许有些人高估了它作为安全问题的严重性;如果有一个相对简单的解决方案,我不会感到惊讶。对我来说,重要的是可解释性研究中普遍存在的担忧:你可能在欺骗自己,你发现了这些东西,但你可能错了。所以,每当你能够将一个可解释性结果转化为关于系统行为的具体预测时,这实际上非常有趣,并且可以让你更有信心你真的理解了这些系统。而且我认为,至少对于这些特定的模型,并且不预判这个问题解决起来是容易还是困难,在某些情况下,你可能会犹豫是否要部署一个你可以用这种简单方式欺骗的系统。

It's certainly a fun thing to discover. I think maybe some people are overestimating it as a safety concern; I wouldn't be that surprised if there was a relatively easy solution to this. The thing that seems important to me about it is the general concern with interpretability research that you may be fooling yourself, that you're discovering these things but you may be mistaken. So whenever you can turn an interpretability result into a concrete prediction about how a system will behave, that's actually really interesting and can give you a lot more confidence that you really are understanding these systems. And I think, at least for these particular models and without prejudice to whether this will be an easy or hard problem to solve, there are contexts where you might hesitate to deploy a system that you can fool in this easy way.

利用可解释性改进系统 Using interpretability to improve systems

Host

有没有什么例子,你能够利用这些可解释性工作来让系统工作得更好或更可靠,或者预见到它可能失败的方式并加以化解?

Are there any examples where you've been able to use this interpretability work as a whole to make a system work better or more reliably, or anticipate a way that it's going to fail and then diffuse it?

Chris Olah

嗯,我认为有例子表明我们捕捉到了你可能担心的问题。但我不知道我们是否已经看到这些例子转化为真正令人信服地缓解了担忧,或者对系统进行了改进。

Well, I think there are examples of catching things that you might be concerned about. I don't know that we've seen examples yet of that then translating into really compellingly ameliorating that concern or making changes to a system to make it better.

Host

我想你对神经网络的一个希望可能是,你能够形成闭环,让理解神经网络成为改进它们的有用工具,并使其成为我们开发它们的一部分。

I guess one hope you might have for neural networks is that you could close the loop and make understanding neural networks into a tool that's useful for improving them, and sort of just make it part of how we develop them.

小规模可解释性的挑战与动机 Challenges and motivations for small-scale interpretability

Host

神经网络研究已经完成,但我觉得至今还没有非常令人信服的例子。这有点双刃剑。一方面,我喜欢我的研究不会让模型变得更强大,或者说通常没有让模型变得更强大,而是纯粹以安全为导向,关注问题。所以当我们向观众征集问题来采访你时,有人问:为什么克里斯专注于小规模的可解释性,而不是像神经科学那样去理解更大规模模块的角色?现在似乎是问这个问题的好时机。你怎么看?

Neural network research is done, and I think we haven't had very compelling examples of that to date. Now it would be a bit of a double-edged sword. In some ways, I like the fact that my research doesn't make models more capable, or generally hasn't made models more capable, and has been quite purely safety-oriented and catching concerns. So when we solicited questions from the audience for this interview with you, one person asked: why does Chris focus on small-scale interpretability rather than figuring out the roles of larger-scale modules in a way that might be more analogous to neuroscience? Seems like an appropriate moment to ask this question. What do you make of it?

Chris Olah

我认为我研究这些小块神经网络的根本原因,也是研究它们有用的原因,是它为我们思考可解释性提供了一个认识论基础。代价是我们谈论这些小部分,并让自己陷入一场斗争,去建立对大型神经网络的理解,使这种分析真正有用。但好处是,我们处理的是如此小的片段,以至于我们可以真正客观地理解发生了什么,因为它归结为基本的数学、逻辑和推理。这几乎就像能够将复杂的数学简化为简单的公理并进行推理。存在很多分歧和困惑,而且真正很难理解神经网络,很容易误解它们。拥有这样的东西似乎非常有用。我觉得电路的根本成功可能不在于每个人都用电路来谈论神经网络,而在于它扮演了数学中公理的角色——它是一个基础,一切都可以被还原到它。

I think ultimately the reason I study these smaller chunks of neural networks, and I think it's useful to study them, is that it gives us an epistemic foundation for thinking about interpretability. The cost is that we're talking about these small parts and setting ourselves up for a struggle to build up understanding of large neural networks and make this analysis really useful. But the upside is that we're working with such small pieces that we can really objectively understand what's going on, because it reduces to basic math and logic and reasoning. It's almost like being able to reduce complicated mathematics to simple axioms and reason about things. There's a lot of disagreement and confusion, and it's genuinely really hard to understand neural networks and very easy to misunderstand them. Having something like that seems really useful. I feel like maybe the radical success for circuits isn't that everyone talks about neural networks in terms of circuits, but that it takes on the role that axioms have in mathematics—it's a foundation that everything can potentially be reduced down to.

分析扩展问题 The analysis scaling problem

Host

所以这项工作表明,我们可以花大量时间,机械而费力地在电路层面理解神经网络,尤其是那些较小的旧模型。我们可以理解这些模型的很大一部分。但神经网络在过去几年变得非常大,而且在真正部署到重要问题之前还会变得更大。现代语言模型比这些图像识别模型大得多。这带来了一个挑战:也许我们可以理解所有这些微小的电路子组件,但它们的数量如此之多,以至于远远超出克里斯·奥拉甚至整个团队的能力,无法理解这些神经网络的任何有意义的比例。我想你可以称之为分析扩展问题。有哪些可行的方法可以扩展分析,以便我们能够真正理解这些更大的模型,或者以某种方式绕过这个问题?

So this work shows that we can, with a bunch of time, mechanically and laboriously understand neural networks at this circuit level, especially older ones that were a bit smaller. We can understand a large fraction of those models. But neural networks have become a lot larger over the last few years, and they're going to become a lot larger still before you actually deploy them to important problems. Modern language models are way bigger than these image recognition models. This creates a challenge: maybe we can understand all these tiny circuit subcomponents, but there's going to be so many of them that it's far beyond the capacity of Chris Olah or even a whole team to understand any meaningful fraction of these neural networks. I guess you could call this the analysis scaling problem. What are the plausible ways in which we could scale the analysis so that we could actually understand these bigger models, or maybe work around the problem in some way?

Chris Olah

我认为这是一个非常合理的担忧,也是电路的主要缺点。目前,我们真正仔细理解的最大电路有 5 万个参数,而最大的语言模型有数千亿个参数。所以如果我们想达到现代语言模型,更不用说未来的神经网络,我们需要跨越好几个数量级的差异。尽管如此,我还是很乐观。我认为我们实际上有很多方法可以解决这个问题。

I think this is a very reasonable concern and is the main downside of circuits. Right now, the largest circuit we've really carefully understood is at 50,000 parameters, and meanwhile the largest language models are in the hundreds of billions of parameters. So there are quite a few orders of magnitude in difference that we need to get past if we want to even get to modern language models, let alone future generations of neural networks. Despite that, I am optimistic. I think we actually have a lot of approaches to getting past this problem.

Host

有哪几个?

What are a couple of them?

Chris Olah

我想大概有四五个值得讨论。

I guess there are a good four or five that are worth going through.

Host

好,我们先说第一个。

All right, let's do the first one.

Chris Olah

所以也许最好的起点是最天真的做法:如果我们直接采用电路方法并尝试扩展它,我们有没有希望做到?这不是我最看好的方法,但我实际上认为它比人们想象的要更可行。有两个原因。第一个原因是,随着神经网络变得更大,有更多的电路和特征需要研究,但通常特征和电路在某种程度上变得更清晰、更容易理解。

So maybe the best place to start is the most naive thing: if we were to just try and take the circuit approach as is and try to scale it, do we have any hope of doing that? It's not the thing I would bet most on, but I actually think it's more plausible than people might think. There are two reasons for that. The first reason is that as neural networks become larger, there are more circuits to study and more features to study, but often the features and circuits in some ways become crisper and easier to understand.

Host

所以你是说随着模型变得更复杂,它们有更多参数,但正因如此,它们在工作上做得更好,所以它们分类的方式更清晰、更连贯,对人来说更容易理解?

So you're saying as the models become more sophisticated, they have more parameters, but for that reason they're somewhat better at their job, so they are classifying things in a way that's clearer and more coherent to people?

Chris Olah

完全正确。我认为有时当你拥有非常弱的模型时,它们只是以一种非常混乱的方式表示事物,将很多东西纠缠在一起,这实际上使它们很难研究。似乎随着模型变得更强,它们在某种程度上变得更清晰、更少纠缠、更干净。

Exactly. I think that sometimes when you have really weak models, they're just representing things in a very confused way that entangles lots of things together, and that actually makes them pretty hard to study. It seems like often as models become stronger, they actually become in some ways crisper, less entangled, and cleaner.

Host

这部分是因为它们没有那么多参数,所以不得不把一大堆不同的概念塞进同一个电路,让它做双重工作吗?

Is that partly because they don't have as many parameters as they might like, so they have to cram a whole bunch of different concepts into the same circuit and make it do double work?

Chris Olah

我认为这可能是部分原因。有一个我们称之为多义性的问题,即一个神经元同时扮演多个角色。我认为它们也可能无法表示真正切分问题的抽象概念;它们根本没有计算能力来构建正确的抽象,所以它们使用次优的抽象。还有一个有趣的事情——有点不同——有一篇 Jacob Hilton 的论文,他们在越来越多样化的数据上训练模型,随着这样做,特征变得更容易解释。所以这是另一件事,我认为这可能是指向这个方向的最严谨的结果,表明在某种意义上,随着你拥有更好的模型,它们变得更容易解释。

I think that is probably part of it. There's this problem we call polysemanticity, which is when a neuron is fulfilling multiple roles like that. I think also they just may not be able to represent the actual abstractions that cut a problem apart; they literally don't have the computational capacity to build the right abstractions, so they're working with suboptimal abstractions. There's this interesting thing—a bit different—but there's a paper by Jacob Hilton where they train models on progressively more diverse data, and the features become more interpretable as you do that. So that's another thing that I think is maybe the most rigorous result pointing in this direction, suggesting that in some sense, as you have better models, they become more interpretable.

Host

好的,所以它们更可解释,但之后你还要说点什么?

Okay, so they're more interpretable, but you're going to say something after that?

Chris Olah

哦,好吧,这是对抗扩展问题的一种抵消力量。然后还有第二件事:有一个我们称之为基序的有趣概念,这实际上是我们从系统生物学中借用的一个想法。

Oh, well, that is a countervailing force against the problems of scaling. And then there's a second thing: there's this really interesting thing we call a motif, and this is actually an idea we've borrowed from systems biology.

神经网络中的重复模式 Recurring Motifs in Neural Networks

Chris Olah

Yuri Allen 和其他一些人真正开创了这种方法,通过他们发现的这些重复模式来理解生物网络。结果我们发现,在神经网络中也能找到类似的重复模式,这些模式实际上可以将电路简化几个数量级。

Yuri Allen and a number of other people have really pioneered this approach to understanding biological networks in terms of these recurring patterns that they find. And it turns out we can find similar recurring patterns in neural networks, and those can actually simplify circuits by orders of magnitude.

Host

所以这里的想法是,即使它非常大,如果它只是大量重复且基本相同的结构,那么你就能一下子理解很多。那么,能举一个在网络中反复出现的这种模式的例子吗?

So the idea here would be, even if it's very big, if it's just lots of recurring things that are basically all the same, then you understand a whole lot of it all at once. But yeah, what's an example of one of these motifs that kind of recurs through a network?

Chris Olah

嗯,我们看到的一个非常简单的例子是并集,你有两个不同的情况,然后得到一个神经元对这两种情况进行并集。但这并不能让我们在理解事物方面取得很大进展。我们从中获益最多的是我们所谓的等变性。这是指神经网络具有对称性,其中有一个特征,实际上有很多该特征的副本,它们是同一特征的变换版本。如果你有一大堆具有这种属性的特征,比如它们都是同一特征的旋转副本,并且你在多个层中都有这种特征,那么实际上电路本身就开始具有对称性。你可以通过大的整数因子来简化它们。在曲线电路的工作中,我们得到了 50 倍的简化,这非常棒。

Well, a really simple one that we see is unions, where you have two different cases and then you get a neuron that unions over those two cases. But that one doesn't give you a huge amount of traction in understanding things. The one that we've gotten the most juice out of is what we call equivariance. This is when neural networks have symmetries in them, where you have a feature and there are actually lots of copies of that feature that are transformed versions of the same feature. If you have a whole bunch of features that have this property, say they're all rotated copies of the same feature, and you have that across multiple layers, then actually the circuits themselves begin to have symmetries. You can sort of simplify them by large integer factors. In the case of the curve circuits work, we got a 50x simplification, which is really nice.

Host

所以,如果有许多电路在不同的颜色和不同的旋转角度下识别同一事物,那么你就有可能查看所有这些电路,然后说,嗯,这是识别曲线的结构,然后它们都在另一个层次上输出相同的结果,都说这是一条曲线。它们都以相同的方式连接在一起,所以你只需理解一次,就能一下子理解更多的内容。因此,如果你想到要跨越许多数量级,当你看到不仅仅是渐进式改进,而是这些数量级上的改进时,这实际上非常令人鼓舞。这表明这并非完全徒劳;你或许有望跨越几个数量级。

So if there are lots of circuits that are recognizing the same thing in lots of different colors and lots of different rotations, then you potentially look at all those and just say, well, this is the curve-recognizing thing, and then they all end up spitting out at the same thing at another level, all saying this is a curve. They're all connected together in the same way, so you can just understand it once, and in one fell swoop you've actually understood a much larger amount of stuff. So if you think about having to bridge many orders of magnitude, it's actually really encouraging when you see things that are not just incremental improvements but actually are these order-of-magnitude improvements. It suggests that it's not completely a fool's errand; you might hope to go and bridge several orders of magnitude.

Host

好的,所以这是两种让 Scaling(规模扩张)问题不那么无望的方法。你还有哪些其他方法来弥合差距?

Okay, so that's two ways that the problem of scaling might not be so hopeless. What are some other approaches you might have to bridging the gap?

Chris Olah

我会把这两种方法都归入一个广泛的类别,即你仍然试图让基本的电路风格方法奏效。其余的想法会有点不同。那么,我们研究电路的动机是什么?我认为我们研究电路的一个很大动机是让它成为这种认识论基础。这并不意味着我们研究的所有东西都需要始终以那个基础来表述。相反,好处在于它给了我们一种方式,以非常严谨的方式构建我们问到的任何其他问题。实际上,当你研究神经网络时,你经常看到这些更大规模的结构。有些神经元集群做着非常相似的事情,或者你看到神经网络的某些部分,所有权重都有非常系统的模式,几乎像生物学中的组织。所以有所有这些暗示,实际上存在大量的宏观结构。你可以想象未来某种可解释性的方法,我们根据这种更大规模的结构来研究事物。然后,当你发现有趣的事情或与安全相关的事情,涉及那种宏观结构时——比如你发现这个更大结构的某个集群参与了社会推理之类的事情,你担心模型可能具有操纵性——那么你可以更仔细地研究它。这可以让你有能力跨越许多数量级,只需查看你的宏观分析告诉你特别重要的非常小的部分。

I'd put both of those sort of in the broad category of approaches where you're still trying to make the basic circuit-style approach work. The rest of the ideas are going to be a little bit more different. So, what was our motivation for studying circuits? I think a big part of our motivation for studying circuits is to be this epistemic foundation. That doesn't mean that everything we study needs to always be in terms of that foundation. Rather, the benefit is that it gives us a way to frame anything else we ask about in a very rigorous way. Actually, when you study neural networks, you often see these larger-scale structures. There are ways in which there are clusters of neurons that do very similar things, or ways in which you see parts of neural networks where all the weights have a very systematic pattern, almost like a tissue in biology. So there are all these hints that there actually is a ton of large-scale structure. You could imagine some future approach to interpretability where we study things in terms of this much larger-scale structure. Then, when you find interesting things or things that are safety-relevant in terms of that large-scale structure—like maybe you find some cluster of this larger thing that is involved in social reasoning or something, and you're worried that the model is perhaps manipulative—then you could look at that much more closely. That could give you the ability to cut through many orders of magnitude by going and just looking at very small parts that your large-scale analysis has told you are particularly important.

Host

所以如果我理解正确的话,你是说,再次用身体做类比,如果电路有点像细胞,也许你会注意到这些细胞被组织成其他结构,比如组织。所以你将能够注意到电路如何聚合到网络内更广泛的结构中。然后可能又有更高层次的结构,也许像器官,然后你就能向上推进。所以你总是可以更深入地放大到电路和特征,但你还处于这个过程的开始。你还会识别出其他结构,这些结构将让你理解所有数量级的完整跨度,而不仅仅是这些较低层次的结构。

So if I understand you, you're saying maybe it's the case that, again using the analogy with the body, if circuits are kind of like cells, perhaps you'll notice that those cells are organized into other structures like tissue. So you will be able to notice ways that circuits all aggregate together into some broader structure within the network. Then maybe there'll be a higher-level structure again, which is perhaps like organs, and then you'll be able to work up. So you'll always be able to zoom in more into circuits and then features, but then you're just at the beginning of this process. You're also going to identify other structures that will allow you to understand the full span of all the orders of magnitude of size, rather than just these lower-level ones.

Chris Olah

是的,完全正确。你知道,如果我们用医学做类比,也许作为一个科学实际上必须对人们有用并解决问题的领域,能够非常精细地理解事物是很重要的。但如果你知道心脏有问题,你就不必去仔细分析脚部发生了什么。所以你既需要能够非常仔细地观察事物并理解它们,也需要更高层次的概览,用来推理对于特定类型的问题你需要关注哪些部分,哪些部分不需要那么关注。

Yeah, that's exactly right. You know, if we're making an analogy to medicine, maybe as an area where science actually has to be useful to people and solve problems, you know, being able to understand things in very fine detail is important. But if you know that there's a problem with the heart, you don't have to go and carefully analyze what's going on in the foot. So you sort of both want the ability to look very closely at things and understand them very carefully, but also the higher-level overview that you can use to reason about what parts you need to pay attention to for particular kinds of problems and what parts you don't need to pay as much attention to.

Host

对,所以如果你说,好吧,我们有了这个巨大的语言模型,但我们真正关心的部分是社交操纵部分,所以我们要非常仔细地放大并理解它。而那个只是识别不同水果的部分,也许我们可以放过它,因为它不是关键。

Right, so if you're like, okay, we've got this enormous language model, but the part that we're really concerned about is the social manipulation part, so we're going to really zoom in in a lot of detail and understand it. And then the part that's just recognizing different pieces of fruit, maybe we can let that one go because it's not essential.

Chris Olah

正是如此。或者你甚至可以想象,有些问题你可以在这个宏观视角中看到,但你能够以一种你信任的方式发展这个宏观视角,因为你有了这个基础来构建。所以你能够用它来推理更大规模的事物,并知道你所说的事情确实映射到模型中真正发生的事情。

Exactly. Or you might even imagine that there are problems that you can see in this larger-scale view, but that you were able to develop this larger-scale view in a way that you trust because you had this foundation to build upon. So you're able to use that to reason about larger-scale things and to know that the things you were talking about actually do map to what's really going on and genuinely occurring in the model.

Host

是的,趁我们还在讨论这个身体类比……

Yeah, while we're on this body thing...

可解释性的细胞类比 Cell analogy for interpretability

Host

你可能会提出一个反对意见:'你只是在研究单个细胞,身体比单个细胞大得多,仅通过研究细胞你能了解疾病或人体的什么?'但问题是,细胞会复制,整个身体都是由细胞构成的。如果我们能理解单个细胞的功能以及它如何出错,那么我们就真正学到了关于整体的东西,因为一切都是重复出现的模式。

One objection you might raise to this is: it would be kind of stupid to say, 'All you're doing is studying individual cells. The body is so much bigger than an individual cell. What can you learn about disease or the human body just by studying cells?' But the thing is, cells replicate, the whole thing is made of cells. If we can understand how an individual cell functions and how it messes up, then we have really learned something about the whole, because it's all a recurring pattern.

Chris Olah

是的,而且当你对组织有疑问时,如果你困惑于发生了什么——我不是生物学家,所以我在用生物学类比,可能有点扭曲——但我想最终能够问这样的问题非常有用:如果你认为某个组织有问题,那么在细胞层面发生了什么?我们理解那里发生了什么吗?或者如果我们有某个理论,能否在细胞层面验证它?这对于你感到困惑的地方可能非常有帮助。

Yeah, and also, when you ask questions about tissues, if you're confused about what's going on—I'm not a biologist, so I'm using biology analogies and might be distorting them a bit—but I imagine it's really useful to be able to ultimately ask: if you think something is an issue with a tissue, what is going on at the cell level? Do we understand what's going on there? Or if we have some theory, can we validate it at the cell level? That may be very helpful for places where you're confused.

自动化与人工分析 Automation vs human analysis

Host

人们可能提出的另一种方法是:如果我们没有足够的人来分析所有这些电路,也许我们需要自动化,创建新的机器学习系统,用它们自己的电路来分析其他事物的电路并识别它们是什么。你对以某种方式自动化这个过程的想法怎么看?

Another approach people might suggest is: if we don't have enough humans to analyze all these circuits, maybe we need to automate it and create new ML systems with their own circuits that analyze the circuits of other things and recognize what they are. What do you make of the idea of automating this process somehow?

Chris Olah

我认为这完全是一个选项,但不是我喜欢的选项。在某种程度上这只是审美问题——我真的很喜欢人类理解事物的方法。我还认为,我听到的许多自动化提议我还没有完全理解,它们常常借用对齐领域我认为还不成熟的想法,然后试图将它们与同样不成熟的电路领域想法结合起来。我担心,当你把一堆不成熟的想法放在一起时,你会让自己更容易失败,因为你现在有多个失败点。也许当这些想法中的更多被弄清楚,我能更仔细地推理如何将它们连接起来时,我会对未来自动化更兴奋。但我认为这是一个不错的备选方案。

I think that's totally an option, but it's not my favorite option. At some level it's just aesthetic—I really like the approach of humans understanding things. I also think that a lot of the proposals I hear for automating things I don't yet fully understand, and they often borrow ideas from alignment that I think are not yet mature, and then try to combine them with ideas from circuits which are also not that mature. I have this nervousness that when you take a bunch of ideas that aren't mature, you're making yourself even more vulnerable to things failing because you now have many points of failure. Perhaps I'll be more excited about automation in the future when more of these ideas are figured out and I can reason more carefully about how to connect them together. But I think it is a good sort of fallback.

通过增加人力来扩展 Scaling by throwing more humans

Host

还有其他处理 Scaling(规模扩张)问题的方法我们应该讨论吗?还是说目前这就是几个主要选项?

Are there any other approaches dealing with the scaling issue that we should talk about? Or is that kind of the top few options for now?

Chris Olah

我认为还有一个:就是投入更多的人力。目前没有多少人在思考这个问题。如果这真的是一个重要问题,并且在某个时候我们真的在部署高风险神经网络,那么想象投入一千人进行系统审计似乎并不完全疯狂。

I think there is one more: just throw more humans at the problem. We don't have many people thinking about this right now. If this really is an important problem and at some point we're really deploying neural networks where there are high stakes, imagining throwing a thousand people at systematically auditing things doesn't seem entirely crazy.

Host

是的,不止如此——如果现在有 10 个人,为什么不能有 1 万、10 万?如果这些网络构成了经济的很大一部分,那么你可能会认为这自然会发生,至少如果这有助于它们更好地运作。看看有多少人分析互联网安全或设计汽车以防止碰撞和伤害人,那是一个相当可观的人数。

Yeah, more than that—if you got 10 now, why not 10,000, why not 100,000? If these networks are making up a large fraction of the economy, then you might think that could naturally happen, at least if it helped to make them function better. If you look at how many people analyze the security of the internet or design cars to not crash and not hurt people, that's a pretty non-trivial number of people.

Chris Olah

是的,我想在某种程度上,这实际上是我研究电路的另一个原因:试图证明理解神经网络是可能的,以此向社会证明这是值得投资并值得尝试找出如何 Scaling(规模扩张)的。我想这是另一种隐含的动机。

Yeah, I guess in some ways that's actually another reason why I study circuits: just to try to demonstrate that it's possible at all to understand neural networks, as a way of justifying to society that this is worth investing in and worth trying to figure out how to scale. I guess that's another sort of implicit motivation I have.

关于可解释性的分歧 Disagreements about interpretability

Host

你多次提到,在该领域内关于可解释性存在很多分歧或缺乏共识。人们对于可解释性有哪些分歧或不同的解释?

You've kind of alluded several times to the fact that there's a bunch of disagreement or maybe a lack of consensus about interpretability within the field. What are the disagreements or different interpretations that people have of interpretability?

Chris Olah

我认为有两件事。一是就在可解释性内部,人们指的是很多不同的东西,对于可解释性是什么或理解一个模型意味着什么并没有真正的共识。然后我认为在可解释性之外,来自机器学习研究社区其他成员的非同小可的怀疑。很多年前,一位我尊敬的同事告诉我,他们认为所有可解释性研究都是胡说八道。我认为这可能比典型观点更强烈,但我认为有一些人持怀疑态度。可能在某种程度上,他们察觉到了关于可解释性意味着什么还没有完全弄清楚的事实。

I think there are two things. One is just that within interpretability, people mean lots of different things, and there isn't really consensus about what interpretability is or what it means to understand a model. Then I think outside of interpretability, there's a non-trivial amount of skepticism from some other members of the ML research community. Once many years ago, a colleague of mine who I respected told me that they thought all interpretability research was BS. I think that's probably stronger than the typical view, but I think there are some people who are skeptical. Probably to some extent they're picking up on the fact that maybe things aren't fully figured out about what interpretability means.

Host

他们还这么认为吗?我只是不明白怎么会有人看着这些文章。也许你认为这不是最重要的事情,或者可能犯了错误,但我不明白你怎么会认为这不是对这些系统如何工作的合法研究。

Do they still think that? I just don't understand how someone can look at these articles. Maybe you think it's not the most important thing, or maybe there are mistakes being made, but I don't understand how you could think this isn't legitimate research into how these systems work.

Chris Olah

我认为实际上,试图如此严谨并试图建立这个基础的另一个动机在某种程度上是为了解决这个担忧。因为如果你相信所有这些理解神经网络的尝试都没有意义并且有根本性缺陷,那将与我如何看待世界非常不同,并且它确实改变了你是否认为将其作为安全方法追求是有意义的。这实际上是一个非常根本的关键点。所以我把我的很多工作视为试图用电路创造一些非常客观正确的东西。

I think actually another motivation for trying to be so rigorous and trying to build this foundation is in some ways to address this concern. Because if you believe that all these attempts to understand neural networks don't make sense and are fundamentally flawed, that would be a very different view from how I see the world, and it really changes whether you think it makes sense to pursue this as an approach to safety. It actually is a very fundamental crux. So I've seen a lot of my work as trying to create something that is very objectively correct with circuits.

Host

人们对可解释性是什么或应该如何解释有哪些不同的看法?

What are the different takes people have on what interpretability is or how it should be construed?

Chris Olah

我应该说明,我显然偏向于自己的工作,而且我不确定我能够完全公平地代表该领域每个人的观点。但我思考这个问题的一种方式——这个观点我真的很感谢 Tom McGrath——是认为可解释性是一门前范式科学,在托马斯·库恩的科学革命结构的意义上。这个想法是,通常科学都有一个范式……

I should caveat that I'm obviously biased towards my own work, and I'm not sure that I'm able to fully fairly represent the views of everyone in the space. But one way I think about this—and I really owe this view to Tom McGrath—is that I think of interpretability as being a pre-paradigmatic science in the sense of Thomas Kuhn's structure of scientific revolutions. The idea is that usually sciences have a paradigm where...

可解释性定义的分歧 Disagreement on what interpretability means

Chris Olah

每个人都认同哪些是重要的问题、如何回答这些问题、如何判断答案是否有效,以及研究的主题是什么。但我认为可解释性没有这些——没有一套共享的答案。目前的讨论是早期科学领域中非常常见的模式,即在这些问题上没有共识。人们提出不同的答案,常常难以相互沟通,存在大量分歧和困惑。事实上,我们正在陷入另一种常见模式:当一个领域处于这种状态时,从业者会试图依赖现有学科来回答这些问题。在这种情况下,我认为从业者主要依赖机器学习或人机交互(HCI)。机器学习的人想要一个衡量可解释性的指标,而 HCI 的人想做用户研究——询问人们某个解释是否有用,或者他们看到解释后是否能做出更好的预测。我认为这两种都是有效的答案,但我倾向于另一种:将可解释性视为一门实证科学,就像神经网络的生物学。问题不在于某样东西是否有用,而在于一个陈述是否在可证伪的意义上为真。你在机器学习的其他领域能看到这种观点,但在可解释性中不那么常见,通常是因为人们认为可视化不够科学。对我来说,电路的目标是展示可解释性如何成为一门实证科学。

Everyone agrees on what the important questions are, how you answer them, how you tell if an answer is valid, and what the topics of investigation are. I think interpretability doesn't have that—there's no shared set of answers. The discourse right now is a very common pattern in early fields of science, where you don't have consensus on these questions. People propose different answers, often have trouble communicating, and there's a lot of disagreement and confusion. In fact, we're falling into another common pattern: when a field is in this state, practitioners try to lean on existing disciplines to answer these questions. In this case, I think we see practitioners lean either on machine learning or on HCI. HCI is human-computer interaction. The ML people want a metric for what it means for things to be interpretable, and the HCI people want to do user studies—ask people if an explanation was useful, or if they make better predictions after seeing an explanation. I think both are valid answers, but I tend toward a different one: seeing interpretability as an empirical science, like the biology of neural networks. The question isn't whether something is useful, but whether a statement is true in the sense of falsifiability. You see this view in other areas of machine learning, but not as much in interpretability, often because people see visualization as less scientific. To me, the goal of circuits is to show how interpretability can be an empirical science.

Host

也许我读了你写的文章和你的框架,受了很多影响,但对我来说,你似乎是在对不同电路的功能做出具体断言——比如,如果你取这组神经元和权重,这就是它的功能。你可以问它是否有用,或者是否帮助我们做出更好的预测,但就像你可以看一台机器并说出齿轮的作用一样,你也可以对电路这么说。拥有那种理解水平是可解释性的一种含义:我看着这台机器,理解每个部分对整体的贡献。这对我来说非常自然。

Maybe I've been influenced so much by reading your articles and your framing, but to me it seems like you're making specific claims about the functions of different circuits—like, if you take this cluster of neurons and weights, this is what it does. You could ask if it's useful or helps us make better predictions, but just like you can look at a machine and say what a gear does, you can say that about circuits. Having that level of understanding is one sense of interpretability: I look at this machine and understand what each part contributes to the whole. That seems very natural to me.

Chris Olah

我会把这描述为机械可解释性,或者有些人称之为透明性。这是将其理解为一种机制,理解是什么让它工作。还有其他类型的工作不那么关注这一点。例如,有很多关于显著性图的工作,试图突出图像分类器最终答案中哪些部分重要。对于这类工作,你可能想从更偏 HCI 的角度来问:这个解释有用吗?它是否让用户做出更准确的预测?事实证明,严格地问这个问题非常困难,因为这些函数是非线性的且复杂。

I would describe this as mechanistic interpretability, or some people call it transparency. It's understanding this as a mechanism, understanding what causes it to work. You can have other types of work that are less focused on this. For example, there's a lot of work on saliency maps, which try to highlight what parts of an image were important for an image classifier's final answer. For that, you might want to ask the question in a more HCI-type lens: is the explanation useful? Does it cause users to make more accurate predictions? It turns out it's very difficult to ask that question rigorously because these functions are so nonlinear and complicated.

Host

所以听起来,在高层面上,你认为存在分歧的部分原因是人们对于理解神经网络意味着什么、可解释性是什么没有共识。也许随着时间的推移,像其他领域一样,人们最终会形成多种不同的可解释性概念,但会更清晰地理解它们——说我们有这种含义但没有那种含义——或者人们会收敛到一个共同的想法,这个想法对实际工作最有用。

So it sounds like at a high level, you think there's disagreement in part because people don't agree on what it means to understand a neural network or what interpretability is. Maybe over time, like with other fields, people will end up with multiple different conceptions of interpretability but understand them more crisply—saying we have this sense but not this sense—or maybe people will converge on a common idea that is the most useful one for the actual work.

Chris Olah

我认为没错。要么形成多个领域,要么,至少根据托马斯·库恩的说法,一个范式最终会胜出。我觉得库恩的描述很有趣,因为我认为范式的核心不是特定的理论,而是人们选择关注的现象。从这个角度看,电路范式的核心——我应该说其他工作也体现了类似的范式——是关注特征以及它们如何相互连接,让这成为你关注的现象。

I think that's right. You could either have multiple fields form, or often, at least according to Thomas Kuhn, one paradigm will eventually win out. I find Kuhn's description interesting because I think the central thing in a paradigm isn't the particular theories, but the phenomena people choose to pay attention to. From that lens, the core of the circuits paradigm—and I should say other work embodies a similar paradigm—is paying attention to features and how they connect to each other, having that be the phenomena you focus on.

Host

我们一直在讨论的文章充满了大量图片,我想这些图片花了很多时间制作。也许这是可解释性的一个组成部分:人们必须能够理解它,所以他们必须能够看到它。你认为可视化在未来会是可解释性的核心部分,还是只适用于这些案例?

The articles we've been talking about are packed with lots of images that I imagine took a lot of time to make. Maybe that's an integral part of interpretability: people have to be able to grasp it, so they have to be able to see it. Do you think visualization will be a core part of interpretability going forward, or is that specific to these cases?

Chris Olah

我认为很多人把科学想象成研究汇总统计,因为在许多科学中,你有几个数字可以把事情归结为真正重要的东西,然后研究它们如何相互作用。我认为人们看到我们的工作时常常很惊讶,因为它们……

I think a lot of people conceive of science as studying summary statistics because in many sciences, you have a few numbers you can boil things down to that are really important, and you study how they interact. I think people are often pretty surprised when they see our work because they're...

汇总统计与细粒度结构 Summary statistics vs. fine-grained structure

Host

过去的研究,尤其是机器学习研究,往往依赖于一堆汇总统计数据和绘制折线图,这似乎是他们心目中科学应有的样子。比如,你可以通过一张图看出随着神经元数量增加准确率的变化,从而了解论文的要点。我觉得他们认为这看起来很严谨,符合他们对科学的期待。但有趣的是,汇总统计数据实际上可能会蒙蔽你的双眼。有一个著名的例子叫安斯库姆四重奏,它展示了几组二维点集,它们的均值和标准差完全相同,但实际图形却截然不同。这还是针对我们非常熟悉的两个统计量——均值和标准差。而神经网络内部的情况要复杂得多,如果你试图仅仅用一个数字来概括,尤其是当你还不了解其内部机制时,你几乎会错过所有重要的东西,至少我是这么认为的。所以我们很多工作就是试图展示这些模型中存在的部分复杂结构。我有时觉得我们就像第一次通过显微镜观察的人,第一次看到了细胞。有人可能会说:‘这科学吗?这只是定性结果,不是定量结果。’但发现细胞本身就是一个定性结果,而不是定量结果,但它非常重要。所以,我认为如果你想理解这些模型,真正了解它们内部的一切,你就需要观察这种精细结构,并努力获取所有这些信息。这不可避免地会推动你走向可视化,因为数据可视化正是展示大量数据、获取大量数据并传达信息的工具。我们习惯于使用一小套有效的可视化形式来传达某些常见数据,但这并不改变一个事实:我们需要其他可视化形式来理解从神经网络中获得的海量数据。

Used to research and especially machine learning research involving studying a bunch of summary statistics and creating line plots and that sort of being what they feel like science should look like, or like you can get the bottom line from a paper by looking at a plot that shows you the accuracy over time as you add more neurons or something like that. And it feels, I think it feels very rigorous to them that that's what they expect science to look like. But there's this really interesting thing where summary statistics can actually kind of blind you. There's this famous example, Anscombe's quartet, where you have a bunch of sets of 2D points that have the same mean and standard deviation but are different when you actually look at them. And that's for two summary statistics that we really understand very well: means and standard deviations. And there's just so much going on inside neural networks that if you just try to boil it down to a single number, and especially if you don't understand what's going on first, you're losing sight of almost everything important, at least that's how it seems to me. And so a lot of our work is just trying to show you some fraction of all the intricate structure that exists within these models. I guess I often feel a little bit like we're like somebody looking through a microscope for the first time, and you're seeing cells for the first time. Sometimes people are like, 'Is that scientific? It's a qualitative result, not a quantitative result.' But seeing cells was a qualitative result, not a quantitative result, and it was really important. So yeah, I think that if you want to understand these models and you want to really understand everything that's going on inside them, you need to be looking at this fine-grained structure and you need to be trying to get access to all this. And that sort of inevitably pushes you towards visualizations, because data visualization is just a tool for displaying lots of data and getting access to lots of data and communicating it. The fact that we're used to a small set of visual forms that are effective for communicating certain kinds of data that we often work with doesn't really change the fact that we're going to want to work with other visual forms for understanding the large amount of data we get from neural networks.

Host

是的。也许我可以换个说法,看看你是否同意:这就像试图让从未见过汽车的人理解汽车的工作原理。也许你无法让他们对汽车如何运作有那种直观的理解,除非他们亲自动手,摆弄不同的零件,看看它们如何连接,看看示意图。你不能只给他们一张图表或一堆数字,然后说:‘这就是汽车的总结。’也许这更像工程学,而不是发现某种自然法则。

Yeah. Maybe if I could try putting it another way and see whether you agree: you're like trying to get people who've never seen a car to understand how a car works. And maybe you can't get that level of intuitive appreciation for how a car functions without getting your hands a little bit dirty, without actually playing with the different pieces and seeing how they connect together and looking at the schematic. You couldn't just present people with a graph or a bunch of numbers and be like, 'This summarizes the car.' And maybe it's a little bit more like engineering than coming up with some natural law.

Chris Olah

是的,是的。我的意思是,想象一下,如果你试图用五个数字来描述汽车。也许用五个数字来描述汽车在某些方面是有用的,比如马力、尺寸、每小时油耗等等。但仅凭这五个数字,你无法造出一辆车。如果你想了解内燃机是如何工作的,你可能需要更仔细地观察,而不仅仅是那五个数字。

Yeah, yeah. I mean, imagine if you tried to describe cars with five numbers. I mean, maybe there are certain ways that it's useful to describe cars with five numbers, like horsepower and size and gallons consumed per hour or something like this. But you couldn't go and build one with just those five numbers. And if you want to understand how a combustion engine is working, you probably need to go and look a little bit more closely than those five numbers.

Host

是的,有道理。人们对可解释性研究的另一个担忧是,正如我们在这段对话中一直提到的,它在某些方面与神经科学非常相似或类似。而神经科学经过几十年的发展,尽管拥有相当大的研究群体,却并未真正破解人类大脑的理解。它取得了一些进展,但远未达到我们放心部署一个非常重要的机器学习系统所需的程度。这是否让你对理解神经网络的前景感到悲观或担忧?

Yeah, that makes sense. Another concern that people might have about interpretability research is just that, as we've been noting throughout this conversation, it seems very similar or very analogous in some ways to neuroscience. And neuroscience over many decades, despite having a pretty large research community, hasn't really cracked understanding of the human brain. It's made some progress, but it hasn't made nearly as much progress as we would need probably to be comfortable deploying a really important ML system. Does that make you at all pessimistic or concerned about the prospects for understanding neural networks as much as we need to?

Chris Olah

是的,我认为这是一个真正的担忧。从事神经科学研究的人群比从事神经网络电路级可解释性研究的人群要大几个数量级。所以这确实是一个令人担忧的理由。但我确实认为,可解释性相对于神经科学有许多非常大的优势,这在一定程度上平衡了局面。简单来说,我最近写了一篇简短的笔记谈到这一点。一些优势包括:你可以获取每个神经元的响应。在神经科学中,你通常只能记录几个神经元,而且只能针对有限的刺激进行记录。但在这里,你可以获取所有神经元的响应。你不仅能够获取——我想很多神经科学家正在试图获取连接组,比如人类大脑中所有神经元是如何连接的——但在神经网络中,你不仅拥有连接组,还拥有连接每个神经元的权重,并且你知道每个神经元执行什么计算。你可以获取整个系统。这就是电路分析成为可能的原因:我们可以去观察权重如何连接神经元。

Yeah, I think it's a real concern. The community of people working on neuroscience is orders of magnitude larger than the community of people working on circuit-style interpretability of neural networks. So it does seem like a pretty good case for being concerned. But I do think that there are actually a number of very large advantages that interpretability has over neuroscience, which level things out a little bit. So just at a high level, I actually wrote a short note about this recently. Some of the advantages are that you can just get access to the responses of every neuron. In neuroscience, you'd normally only be able to record a couple of neurons, and you'd be able to record them for a limited number of stimuli. But here you can get it for all of them. Not only do you have access—I guess a lot of neuroscientists are trying to get access to the connectome, like how all the neurons connect in the human brain, for instance—but in neural networks, not only do you have the connectome, but you have the weights that connect every neuron, and you know what computation every neuron does. You just have access to the entire thing. And that's what makes circuits possible: we can go and look at how the weights connect neurons.

Chris Olah

还有一个事实是,权重共享可以大幅减少神经网络中独特神经元的数量。例如,在视觉模型中,你实际上有同一个神经元的许多副本,它们在图像的每个位置运行。所以,如果我有一个线条检测器,那么在图像的每个位置运行这个线条检测器是很有用的。这样做的结果是,与研究一个类似的生物神经网络相比,你拥有的独特神经元数量可能少了一万倍。这就是无法观察每个神经元与可以轻松观察每个神经元之间的区别。我们已经对一些视觉模型做到了这一点。所以这是一个巨大的优势。

There's also the fact that weight tying can dramatically reduce the number of unique neurons that exist in neural networks. For instance, in vision models, you actually have all these replicas of the same neuron where you run them at every position in the image. So if I have a line detector, it's useful to run that line detector at every position in the image. The result of this can be that you have maybe 10,000 times fewer unique neurons than you would if you were studying a comparable biological neural network. So that's the difference between it not being possible to look at every single neuron and it being easily possible to look at every single neuron. And we've done this for some vision models. So that's a huge advantage.

Host

是的,听起来这些优势如此之大,以至于你最终可能通过研究神经网络学到更多关于人类大脑的知识,而不是通过研究人类大脑本身来理解它,因为人类大脑中的底层数据太难获取了。它被包裹在所有这些物理物质中,非常难以操作。

Yeah, it sounds like maybe the advantages are so great that you'll end up learning more about the human brain by studying neural networks than you can understand the human brain by studying the human brain, just because the underlying data is so hard to access within the human brain. It's tied up in all this physical stuff, super hard to play with.

Chris Olah

是的。

Yeah.

神经网络与神经科学 Neural Networks and Neuroscience

Chris Olah

我不是神经科学家,不想夸大其词,但在我看来,研究人工神经网络很可能能教会我们很多关于神经科学的知识。而且这只是一部分,还有其他一些我没提到的优势。我认为我们确实有相当多的优势。

I'm not a neuroscientist and I wouldn't want to overstate it, but it seems very plausible to me that studying artificial neural networks could teach us a lot about neuroscience. And I should say that's just part of it; there are a number of other advantages I haven't mentioned. I think we do have a pretty significant number of advantages.

Host

好的,我们会附上你最近发表的那篇论文的链接,你在里面列出了其他好处。我们收到了大量观众提问,因为之前提到要采访你。我现在随便挑几个。其中一个我感兴趣的问题是——我不确定你是否能真正回答,因为它可能太哲学了——有人指出:克里斯,你在神经网络中发现了所有这些与神经科学相似的现象。像布莱恩·托马西克这样非常关心未来苦难的人担心,神经网络、人工智能或强化学习系统在计算机上运行时可能会受苦,或者也可能感受到快乐。这个人问:你认为你的发现应该让我们更担心还是更不担心这个问题?

Okay, we'll stick up a link to that paper you recently published, or I guess you lay out the other benefits there. We got a ton of audience questions since we mentioned we'd be interviewing you. Maybe I'll just chuck in a couple now. One that I was interested in—I'm not sure if you can really tackle it because it might be too philosophical—but someone pointed out: Chris, you're finding all these discoveries inside neural networks that mirror neuroscience. People like Brian Tomasik and others who are very concerned about suffering in the future have raised concerns that potentially neural networks, AI, or reinforcement learners might be able to suffer when running on computers, or feel pleasure as well. This person asked: do you think your findings should make us more or less worried about that?

Chris Olah

我先说我觉得自己不太有资格评论这个问题,因为我不是道德哲学家或心灵哲学家。我真的不知道该如何思考某物是否具有道德地位。老实说,我不确定有谁是完全有资格的。但我的直觉是,这是一个相当严重的担忧。我觉得人们觉得谈论这个很好笑,或者如果你谈论这个,别人会觉得你是个怪人。但没错,我担心这是一个很严重的问题。我担心的很大一部分原因甚至不是它们可能成为有道德地位的智能体,而是这个问题如此不可见。我非常关心动物权利;我是素食者。动物权利问题之所以如此隐蔽,是因为人们看不到工厂化养殖场里的痛苦——它被隐藏起来,不可见。如果神经网络受苦,那会有多不可见?我不知道它们是否会受苦,但如果会,那会有多不可见?它们在服务器上运行,没人能看到,也没有可见的痛苦输出。我键盘上点几下按钮就能扩大经历这种痛苦的神经网络数量。我认为这确实增加了道德灾难的风险。现在,我不知道我能否就模型是否可能受苦说出什么有用的东西。但我确实认为这些多模态神经元应该是一个小小的红旗。我们之前只在人类身上见过的东西,现在在人工神经网络中观察到了。我不是说这一定是个大红旗;其他人更适合思考这个问题。但我确实认为,我希望我们能系统性地审查和检查这类事情。每一个这样的发现似乎都增加了我们可能正在处理道德主体的风险。

I'll just start by saying I don't feel super qualified to comment on this because I'm not a moral philosopher or philosopher of mind. I really don't know how you should think about whether something has moral patienthood or not. Honestly, I'm not sure anyone who is super qualified. But my instinct is that this is a pretty serious concern. I think it's something people think is funny to talk about, or that you're being a crank if you talk about it. But yeah, I'm worried this is a pretty serious issue. A big part of why I'm worried isn't even the probability that they might be agents entitled to moral patienthood, but that it's so invisible. I care a lot about animal rights; I'm a vegan. One thing that makes animal rights so insidious is that people don't see the suffering in factory farms—it's hidden from us, invisible. How much more invisible would the suffering of neural networks be if they were to suffer? I don't know that they will, but if they were, how much more invisible would it be? They're running on a server where no one can see them, no visible output of their suffering. A couple clicks of a button on my keyboard can scale up the number of neural networks experiencing this. I think that really increases the risk of a moral catastrophe. Now, I don't know if there's anything useful I can say about whether models might be suffering. But I do think these multimodal neurons should be a little bit of a red flag. We have something that previously only existed in humans, and now we're observing it in artificial neural networks. I'm not saying that necessarily should be a big flag; other people are better positioned to think about that. But I do think it's the sort of thing I wish we were systematically reviewing and checking in on. Each one seems to increase the risk that maybe we're dealing with moral agents.

Host

所以担忧是:我们认为人类可能能够受苦,如果这些神经网络不断添加越来越多人类拥有的功能组件,那么它们可能也会发展出受苦的能力,因为那是我们心智中的一部分。这有点像图像识别系统——不太清楚与偏好、欲望或痛苦的类比是什么。也许需要一个更智能体式的神经网络才能受苦。这对我来说很直观,但我想我们会有这样的系统,对吧?

So the concern is: we think humans can suffer probably, and if these neural networks keep adding more and more functional components that humans have, then maybe they'll develop the ability to suffer as well, because that's one of the things in our minds. It seems a little bit like an image recognition system—it's not obvious what the analogy to a preference, a desire, or suffering is. Maybe you need a more agentic neural network before it can suffer. That feels intuitive to me, but I suppose we will have such systems, right?

Chris Olah

嗯,我认为很多人真的把强化学习视为神经网络是否受苦的关键区别。我相当怀疑这实际上是核心问题。一个你可能不这么认为的原因是,你可以用模仿学习来训练一个神经网络来模仿一个行为者,如果它导致了一个具有相同行为的神经网络,我可能会认为它们受苦的可能性相同。我还有很多技术上的理由认为强化学习可能不是核心。但也许这里更重要的点是,讨论的水平如此之低,如此之少,以至于连基本观点和想法都没有被阐述出来。我认为这是因为谈论这个感觉像是个怪人,尤其是如果你是一个严肃的研究人员。我认为这是个问题。我们应该努力仔细思考这些问题,这是一个非常难思考的话题,但我希望有更谨慎的讨论。

Well, I think a lot of people really focus on reinforcement learning as the thing that must be the difference between a neural network suffering or not. I'm pretty skeptical that's actually the central issue. One reason you might not think that is you could train a neural network with imitation learning instead to mimic an actor, and if it led to a neural network with the same behavior, I'd probably think they were equally likely to suffer. I have a bunch more technical reasons why I think reinforcement learning is probably not the central thing. But maybe the more important point here is just that the level of discourse is so low and minimal that even basic points and thoughts haven't been laid out. I think it's because it feels like a crank thing to talk about, especially if you're a serious researcher. I think that's an issue. We should be trying to think these things through carefully, and it's a really hard topic to think about, but I wish there was more careful discourse.

Host

另一位听众来信想问:也许我们可以用可解释性研究来理解那些做我们自己能做的事情的神经网络。但如果你有一个超人类系统,一个超级智能系统,它可能在做对我们来说太复杂或太困难的任务,因此也太复杂和困难让我们理解。你认为这会影响我们使用可解释性工具来研究比今天更先进的系统的能力吗?

Another listener wrote in and wanted to ask: maybe we can use interpretability research to understand neural networks that are doing things we ourselves can do and understand. But if you have a superhuman system, a superintelligent system, it might be doing tasks too complex or difficult for us to do, and therefore too complex and difficult for us to understand. Do you think that could interfere with our ability to use interpretability tools on far more advanced systems than what we have today?

Chris Olah

我把这个问题分成两部分。一部分是普遍担心神经网络的 Scaling(规模扩张),以及我们能否跟上研究更大神经网络的步伐。另一部分是关于超级智能和比我们更聪明的系统的具体担忧。你可能会想,有些方法可以研究非常大、非常强大的模型,然后当这些模型变得比你更聪明,因为它们以更聪明、更复杂的方式思考问题时,也许……

I split this question into two parts. One is generically being worried about the scaling of neural networks and whether we'll be able to keep up with studying larger neural networks. The other is specific concerns regarding superintelligence and systems being smarter than us. You might think there are ways in which you can study very large, very powerful models, and then when those models become smarter than you because they're thinking about problems in a way that's smarter and more sophisticated, that maybe...

不可想象的思想与思维工具 Unthinkable Thoughts and Tools for Thought

Host

那是一个让你陷入困境的特殊点。我想我们之前已经讨论过 Scaling(规模扩张)的部分,实际上我们更应该关注这些超级智能特有的担忧。

That's a special point where you get screwed up. I think we've already talked about the scaling part earlier, and really it's more these superintelligence-specific concerns that we should focus on here.

Chris Olah

是的,说得对。我想有一点让人欣慰的是,我们已经拥有在狭窄领域比我们更聪明的模型。比如 ImageNet 模型在识别不同种类的狗方面比我强,通过努力我们可以观察这些模型,从中学习,让自己变得更聪明。Bret Victor 有一个非常精彩的演讲《思考不可思考之事的媒介》,其中有些话我觉得非常深刻。我认为这种“思考工具”的思维方式在 EA 社区中被低估了,但它实际上与神经网络极其相关。也许我引用一下,因为我认为这是思考这个问题的一个非常有力的方式。

Yeah, that seems right. I guess one thing that seems a little heartening here is we already have models that are smarter than us in narrow ways. Like ImageNet models are better than me at recognizing different kinds of dogs, and with effort we can look at these models and actually learn from them and make ourselves smarter. There's this really lovely talk by Bret Victor, 'Media for Thinking the Unthinkable,' and he says some things which strike me as being very profound. I think this kind of 'tools for thought' thinking is something that is underappreciated in the EA community, but I think it's actually extremely relevant to neural networks. Maybe I'll quote it for a second, because I think it's a really powerful way of thinking about this.

Host

好,请说。

Yeah, go for it.

Chris Olah

他首先引用了 Richard Hamming 的话。Hamming 说:‘就像狗能闻到而我们闻不到的气味,狗能听到而我们听不到的声音一样,也有我们看不到的光波长和我们尝不到的味道。那么,既然我们的大脑是这样连接的,为什么‘也许有些想法我们无法思考’这句话会让你感到惊讶呢?进化至今可能已经阻止了我们朝某些方向思考。可能存在不可思考的想法。’我认为这在某种程度上触及了试图理解超级智能系统的担忧根源。你可能会想,有些想法人类无法思考,而这个系统正在思考这些想法,因此我们完蛋了。

So he starts by quoting Richard Hamming. Hamming says: 'Just as there are odors that dogs can smell and we cannot, as well as sounds that dogs can hear and we cannot, so too there are wavelengths of light that we cannot see and flavors we cannot taste. Why then, given our brains are wired the way we are, does the remark perhaps there are thoughts we cannot think surprise you? Evolution so far may possibly have blocked us from being able to think in some directions. There could be unthinkable thoughts.' I think in some ways this gets to the root of the concern about trying to understand a superintelligent system. You might think that there are thoughts which are impossible for humans to think, that this system is thinking, and therefore we're screwed.

Chris Olah

Victor 回应如下:‘我们听不到的声音,我们看不到的光——我们最初是怎么知道这些东西的?我们建造了工具。我们建造工具来将这些超出我们感官的事物适应到我们的身体、我们的感官。我们听不到超声波,但你可以把麦克风连接到示波器上,然后它就出现了——你用你那普通的猴子眼睛看到了那个声音。我们看不到细胞,也看不到星系,但我们建造了显微镜和望远镜,这些工具将世界适应到我们的身体、我们的感官。当 Hamming 说可能存在不可思考的想法时,我们必须承认是的,但我们建造工具来将这些不可思考的想法适应到我们思维的方式,让我们能够思考这些以前不可思考的想法。’我认为这与我们正在做的可解释性工作有着深刻的类比。Michael Nielsen 和 Sean Carter 有一篇很好的文章,讲的是他们所谓的‘人工智能增强’。高层次的想法是,我们可以尝试建造工具来帮助我们变得更聪明,思考这些想法。事实上,如果我们有以这些方式推理的系统,它们可以教我们抽象概念,甚至它们本身可能就是让我们去思考这些以前不可思考的想法的工具。

And Victor responds as follows: 'The sounds we can't hear, the light we can't see—how do we even know about those things in the first place? We built tools. We built tools to adapt these things that are outside our senses to our human bodies, our human senses. We can't hear ultrasonic sound, but you can hook up a microphone to an oscilloscope and there it is—you're seeing that sound with your plain old monkey eyes. We can't see cells and we can't see galaxies, but we build microscopes and telescopes, and these tools adapt the world to our human bodies, to our human senses. When Hamming says there could be unthinkable thoughts, we have to take that as yes, but we build tools to adapt these unthinkable thoughts to the ways that our minds work and allow us to think these thoughts that were previously unthinkable.' I think there's actually a deep analogy to what we're doing with interpretability. Michael Nielsen and Sean Carter have a nice essay on what they call 'artificial intelligence augmentation.' The high-level idea is that we can try to build tools to help us be smarter and to think these thoughts. In fact, if we have systems that are reasoning in these ways, they can teach us the abstractions and maybe even be themselves the tools that allow us to go and think these previously unthinkable thoughts.

Host

是的,所以一个担忧可能是,虽然也许你可以把这些不可思考的想法分解成我们可以消化的碎片,但也许即使付出足够的时间和努力,这也可能非常困难,因为我们将要处理非常陌生、非常异类、人类很难掌握的抽象概念。所以这可能比理解这是一个识别狗的电路要费力得多。

Yeah, so a concern might be that while maybe you could break these unthinkable thoughts into pieces that we can digest, maybe given enough time and effort, perhaps that's going to be really quite hard because we're going to be dealing with abstractions that are very foreign and very alien, very difficult for humans to grasp. So it could be a much more laborious thing than understanding that this is a dog-recognizing circuit.

Chris Olah

是的,这可能是真的。不过我不知道——如果你看看很多我认为非常聪明的人,他们的思维实际上非常清晰,而且通常很容易理解。也可能你可以想象一条曲线,表示神经网络的可理解程度。你可以想象有一个低谷,网络非常混乱,抽象概念很差,然后它变得越来越清晰,越来越容易理解,然后也许有一个阈值,它开始产生足够异类的想法,以至于我们再也无法理解——超出了人类的能力。这是一个关于阈值在哪里的经验问题。但如果我们能达到处理人类水平系统以及比人类水平稍强的系统,理解它们并真正对它们的安全性有信心,那对我来说将是一个相当大的安全胜利,即使存在某个点,这些系统从根本上变得如此异类,以至于我们完全无法理解它们。

Yeah, that could be true. Although I don't know—if you look at many humans who I think are really smart, they actually are very good at their thinking, which seems very clear and often quite easy to understand. It could also be that you can think of the curve of how understandable a neural network is. You could imagine that there's some valley where it's very confused and has really bad abstractions, and then it gets crisper and crisper and becomes easier to understand, and then perhaps there's some threshold where it starts to have just sufficiently alien thoughts that we can't understand it anymore—past human capacity. And there's an empirical question about where that is. But if we can get to the point where we're dealing with human-level systems and somewhat stronger than human-level systems and understand them and really be confident in their safety, that would actually seem like a pretty big safety win to me, even if there's some point where these things fundamentally become so alien that we just can't understand them at all.

Host

好的。另一位听众来信说,Chris 似乎非常擅长揭示和理解他各种模型的内部运作,包括复杂且相当大的模型。他们问:一个人如何持续提升自己解释和理解模型的能力?

Yeah, all right. Another listener wrote in saying Chris seems to be extraordinarily good at pulling up and understanding the inner workings of his various models, including even complex ones and fairly large ones. They're asking: how does one presumably continually take their ability to explain and understand models to the next level?

Chris Olah

嗯,他们太客气了。我认为最重要的事情可能就是练习,花大量时间盯着神经网络,试图弄清楚发生了什么。但让我想想还能建议什么。我想,冒着有点自我推销的风险,阅读 circuits 系列文章可能非常有用。它带你了解神经网络所做的许多事情以及它们如何做到的例子,以及实现这些功能的电路。我认为拥有一个你理解的具体例子库实际上非常有帮助。

Well, that's very kind of them. I think probably the biggest thing is just practice and spending lots of time staring at neural networks and trying to figure out what's going on. But let's see what else I can suggest. I think, you know, at the risk of being slightly self-promotional, just reading the circuits thread may be quite useful. It walks you through a lot of examples of things that neural networks do and examples of how they do it, and the circuits that are implementing it. I think just having a library of concrete examples that you understand is actually very helpful.

Host

是的,我觉得我从中受益很多。

Yeah, I feel like I benefit a lot from that.

Chris Olah

我认为另一个有用的练习是给自己创造一些玩具问题,问自己如何在这种架构的神经网络中实现这种行为。我晚上睡不着的时候经常玩这类东西。我觉得这样能更深入地理解架构的约束,以及哪些事情容易、哪些困难。

I think a useful exercise also can be to create toy problems for yourself and ask yourself how you would implement this behavior in a neural network with this architecture. I often play with things like this when I can't sleep at night. And I feel like I come to much more deeply understand what the constraints of an architecture are and what kinds of things are easy and hard.

Host

有道理。听起来就是练习、练习、再练习。

Makes sense. Sounds like practice, practice, practice.

Chris Olah

花时间与模型相处,多探索真的很有帮助。然后可能有一件事人们会觉得惊讶,就是建立一些基本的数据可视化能力,并熟练使用那些能让你把大量数据呈现在纸上、提取神经网络权重、用特征可视化组织它们、并拥有简单界面来导航的工具。因为这种能力……

Just spending time with a model and poking around a lot is really helpful. And then maybe one thing that I think people might find surprising is just building some basic data visualization fluency and getting some confidence with tools that allow you to get lots of data on paper, get neural network weights out there, organize them with feature visualizations, and have simple interfaces that allow you to navigate that. Because the ability to...

显微镜工具与推荐 Microscope tool and recommendation

Host

这些模型内部有海量数据,拥有能够让你导航和筛选的工具,是避免你不得不依赖会误导你的汇总统计的关键。

There's just so much data inside these models. The ability to have tools that allow you to navigate it and sift through it is the thing that prevents you from needing to resort to summary statistics that will mislead you.

Host

是的,我想我之前忘了说,OpenAI 提供的这个显微镜工具,我记得是在 microscope.open.com,用起来真的很棒。即使那些不是真的想学这个用于工作的人,我也推荐去看看。我被震撼了。我当时想,这不就是理解它的方式吗?然后我点开文章,发现这东西已经存在了。

Yeah, I think I failed to say earlier that this microscope tool that OpenAI is offering, which I think from memory is at microscope.open.com, yeah, it's just beautiful to play around with. Even the people out there who aren't into actually trying to learn how to do this for their job, I recommend go and check it out. I was blown away. This is something that I thought, shouldn't this be the way that you understand it? And then I was like, oh my God, I clicked through the article and was like, this exists now.

Host

所以这种体验大概能帮助人们。

So presumably that kind of experience might help people.

语言模型与视觉模型的可解释性 Interpretability for language models vs vision models

Host

另一个多位听众提出的问题是:到目前为止我们讨论的主要是视觉模型,比如 CLIP,它们试图分类图像,或者根据图像找出标题,或者根据标题生成图像。但可能还有其他模型更重要,正在被部署,能力更强,做更相关的事情。显然像 GPT-3 这样的模型,能根据开头段落写出整篇文章,似乎这类工作比图像分类更常见。那么这种可解释性工作会自然地从图像领域扩展到自然语言解释吗?还是说存在更根本的差异,导致类比失效?是更难做,还是需要重新发明方法?

Another question that actually multiple listeners raised was: so most of what we've been talking about so far is about these visual models, things like CLIP, which are trying to classify images or take an image and figure out what the caption should be, or take a caption and design an image around it. But there are probably going to be other models that are maybe more important, that might be getting deployed, that are more capable, that are doing more relevant stuff. I guess obviously models like GPT-3, which will take an opening paragraph and then write an entire essay from it, it seems like perhaps more work looks like that than classifying images. Is this interpretability work going to just extend naturally from the image stuff to natural language interpretation, or might there be more fundamental differences that mean the analogy somewhat breaks down? Is it harder to do, or are you going to have to kind of reinvent things?

Chris Olah

嗯,在深入讨论之前,我想说的是,很多可解释性工作确实是在语言模型上做的。这些工作的风格通常和我真正感到兴奋的那种很不一样。但确实有很多工作在那里进行。所以我把这个问题理解为:我们应该对电路式可解释性工作在语言模型背景下蓬勃发展的前景有多乐观?我认为在真正尝试之前,这个问题很难回答。我感到乐观。我看不到任何根本性的原因说明它为什么行不通。但我觉得在我们证明之前,人们持怀疑态度是合理的。我最近实际上转向了语言模型的工作,到目前为止只做了一些非常初步的东西。但我们确实逆向工程了几个简单的电路,用于语言模型中与元学习相关的一些事情,而且看起来相当容易。所以我的初步立场是,我不仅看不到任何根本原因说明这些事行不通,而且我感到相当乐观。

Well, I guess something I should say before diving into this is that a lot of interpretability work does get done on language models. It's often a pretty different style of interpretability work than the thing that I sort of feel really excited about. But a lot of work does get done there. So I'm going to interpret this question as: how optimistic should we be that circuit-style interpretability work can thrive in the context of language models? And I think to some extent, that's a hard question to answer before we really try. I feel optimistic. I don't really see any fundamental reason why it shouldn't work. But I think it's reasonable for people to be skeptical until we demonstrate that. I actually switched to doing work on language models very recently, and so all I've done so far is just some very preliminary stuff. But we've actually reverse engineered a couple simple circuits for doing some things related to meta-learning in language models, and it's actually seemed quite easy. So my preliminary position is that not only do I not see any fundamental reasons why these things shouldn't work, but I feel quite optimistic.

Host

你之前提到,语言模型上已经做了一些不同类型的可解释性工作。那发现了什么,又是什么样的呢?

You mentioned earlier that some kind of different interpretability work has been done on language models. What's being turned up by that, and maybe what does that look like?

Chris Olah

我认为有很多工作大致是这样的:我们觉得某个语言特征很重要,然后尝试确定它在 CLIP 的不同位置被表示的程度。所以我可能会把这描述为一种自上而下的可解释性,你对某件事有一个假设,然后尝试看是否能从某个层预测它。还有一些更偏可视化的方法。有一篇 Google 的论文我很喜欢,他们可视化各种词嵌入如何随着与周围上下文的交互而演变。他们研究了 'die' 这个词,发现在某些上下文中它像德语冠词,那是一个簇;在另一些上下文中它是死亡,那是另一个簇。所以我觉得更偏可视化的东西很酷。还有一些很好的可视化注意力模式的工作。所以我认为这个领域有一些有趣的工作。我不太确定是否有任何工作已经达到了机制解释和这种自下而上理解特征的程度,而这是我最信任的方式。但我也在进入一个新领域,目前还是个新手,所以我不想评论得太自信。

I think there's a lot of work that's broadly of the flavor of: there's some linguistic feature we think is important, can we determine the extent to which that's represented at different points in CLIP? So I might describe this as slightly top-down interpretability, where you have a hypothesis about something and you try to see if you can predict it from some layer. There's also some more visualization-oriented approaches. There's this one paper by Google that I really liked, where they were just visualizing how the embeddings of various words evolve as they interact with the context around them. They looked at the word 'die', and they found that in some context it's like a German article, and that's one cluster, and in another context it's death, and that's another cluster. So I think the more visualization-oriented stuff is cool. There's been some nice work visualizing attention patterns. So I think there is some interesting work in the space. I don't really know that any of it has quite got to the point of getting to mechanistic explanations and this sort of bottom-up understanding of features that I sort of feel like I most trust. But also, I'm entering a new space and I'm presently an amateur in that space, so I wouldn't want to comment too confidently.

Host

是的,是的。好吧,所以这种自上而下的解释,人们可能会接近模型说,它必须能预测这里该用名词还是形容词,所以我们应该找到它的一部分在做这种预测下一个词类型的任务,然后我们去侦察看看能不能找到。我猜你更倾向于从神经元或簇开始,然后向上工作,找出这个东西在做什么,而不带太多先入之见。但我想两种方法都很自然。也许那会是我最兴奋的事。但我很高兴有各种各样的人尝试不同的事情。而且我认为非常令人兴奋的是,有一个相当大的社区在这个领域提出问题,还有一个整个领域人们称之为 'Bology',我想就像生物学,但针对的是 BERT 这个语言模型。我认为这是一件了不起的事情。

Yeah, yeah. Okay, so this kind of top-down interpretation where people might approach the model and say, well, it's got to be able to predict whether a noun or an adjective should go here, so we should expect to find some part of it that's doing this prediction of what kind of word should go next, and then we're going to go scout around and see if we can find it. I guess you prefer it maybe where we start with a neuron or a cluster and then work up and figure out what this thing is doing without bringing too many preconceptions to the table. But I suppose it's natural to do both. Maybe that would be the thing that I'd be most excited about. But I'm excited to have a wide range of people trying things. And I think it is really exciting that there's a significant community of people asking questions in the space, and there's this whole area that people call 'Bology', I guess like biology but for BERT, which is a language model. And I think it's wonderful that that is a significant thing.

Host

你能再多说一些语言模型中电路的等价物是什么吗?就像 'die' 的例子,你可以想象有电路检测:这个句子是关于这类东西的吗?这个句子是关于人还是机器?或者你可能得从更简单的东西开始。但那可能是某种等价物,用来找出主题的大致类别。

Is there anything more that you can say about what would be the equivalent of a circuit in a language model? As it sounded like with the example of 'die', you could imagine that there are circuits that detect: is this a sentence about this kind of thing? Is this a sentence about a person or a machine? Or maybe you would have to start with something even simpler than that. But that might be some kind of equivalent where it's figuring out the broad cluster of the topic.

Chris Olah

嗯,在我的概念里,就像我们在视觉模型中有特征一样,我们在语言模型中也有特征。它们存在于 MLP 中。所以你有这些块,它们就是普通的神经网络层,有神经元,就像 Transformer 中的其他东西一样。然后不同的地方是,你还有这些注意力头,它们与我们在卷积网络中看到的任何东西都非常不同。所以我们必须调整我们以电路方式思考的框架,把注意力头也包括进去。然后希望我们能……

Well, in my conception of this, just like we have features in vision models, we also have features in language models. They live in MLPs. So you have these blocks that are just regular neural network layers with neurons like anything else in Transformers. And then the thing that's different is you also have these attention heads, which are very different from anything we see in convnets. So the thing that we have to do is sort of adapt our framework of thinking about things in terms of circuits to also include these attention heads. And then hopefully we can...

可解释性对其他模型类型的适用性 Applicability of interpretability to other model types

Host

有没有哪些类型的机器学习模型——也许是语言模型之外、现在存在或未来可能存在的模型——你会担心这些可解释性方法不起作用?比如其他类型的神经网络:AlphaZero(学习玩很多游戏)、MuZero、Facebook 的推荐算法。有没有可能某一类神经网络特别难以理解?

Are there any sorts of ML models, maybe models apart from language ones that exist now or might exist in future, where you worry that these interpretability methods wouldn't work? Examples of other kinds of neural networks: AlphaZero that learns to play lots of games, MuZero, Facebook's recommender algorithm. Are there ways in which a class of neural network could be really hard to understand?

Chris Olah

我认为有很多类型的神经网络,我还没有深入思考过理解它们会是什么样子,或者对它们进行电路类型分析会是什么样子。我确实认为所谓的基于模型的强化学习模型——它们某种程度上预测未来,或者像 AlphaGo 那样展开许多可能的未来并影响其行为——研究起来会非常有趣且不同。我不确定它们会更难,但会与我目前研究的模型非常不同。

I think there are many types of neural networks where I haven't thought hard about what understanding them would be like, or what doing circuits-type analysis on them would be like. I do think models that are so-called model-based reinforcement learning, where they sort of anticipate futures, or something like AlphaGo that unrolls many possible futures and then has that influence what it does, would be very interestingly different to study. I don't know that they'd be harder, but they'd be very interestingly different from the models I presently study.

关于可解释性研究的怀疑论点 Skeptical arguments about interpretability research

Host

我们已经提出了一些人们可能对可解释性研究议程持有的反对意见或疑虑。还有其他让你困扰的怀疑论观点吗?可能成为死胡同或不太有用的地方?那些让你夜不能寐的?

We've raised a couple of objections or doubts people might have about this interpretability research agenda. Are there any other skeptical arguments that trouble you? Ways this could turn out to be a dead end or not so useful? Ones that keep you up at night?

Chris Olah

我认为 Scaling(规模扩张)问题可能是最让我害怕的。我相对乐观,但当我这么说时,可能意味着我大于 50% 乐观。事情仍有很多失败的空间。我是那个应该最乐观的人,可能有点妄想地乐观。所以这是一个选择过程:如果你是主导者却认为这是个坏计划,那会令人惊讶,也会令人担忧。但这就是为什么你需要很多不同的议程。希望有一个组合,其中你希望至少一个最终成功,而不指望任何单一事情成功。

I think the scaling one is probably the thing I'm most scared of. I feel relatively optimistic, but when I say relatively optimistic, maybe that means I'm greater than 50% optimistic. There's still a lot of room for things to fail. I'm the person who should be most optimistic, probably slightly delusionally optimistic about it. So it's a bit of a selection process that would make it surprising if you thought this was a bad plan and you were the main person leading it. That would be worrying. But I think this is why you want to have lots of different agendas. Hopefully there's a portfolio where you hope one of them eventually pans out, and you aren't counting on any single thing to succeed.

多义性作为挑战 Polysemanticity as a challenge

Chris Olah

就这里没有太多提及的担忧而言,有多语义性问题:一个神经元对多个事物做出响应。关于为什么会发生这种情况,有一些理论,比如没有足够的空间来包含所有概念。我认为我们只是不知道为什么会发生,但这使得理解神经网络更加困难,尤其是当你试图从神经元的角度看问题时。

In terms of worries that haven't been voiced as much here, there's this issue of polysemanticity, where a neuron responds to multiple things. There are theories about why this happens, like there isn't enough space to include all the concepts. I think we just don't know why it happens, but it makes it a lot harder to understand neural networks, especially if you're trying to look at things in terms of neurons.

Host

多语义性是指一个电路同时做两个非常不相关的事情,比如检测汽车和英国女王。当网络很小且必须将许多功能塞进很少的神经元时,这似乎合理,但你说有时甚至在非常大的网络中也会发生?

Polysemanticity is where you have a circuit that is simultaneously doing two very unrelated things, like detecting cars and the Queen of England. It seems to make sense when the network is small and has to cram many functions into few neurons, but you're saying it happens even in very big networks sometimes?

Chris Olah

至少在中等规模的网络中肯定会发生。随着模型变大,模型可以表示的有用事物集合也在变大。所以神经元数量在增加,但你想存储的事物集合也在增加。因此存在一个权衡:要么拥有更少的概念,比如说'这只检测汽车,不检测女王',要么接受一个做两件事的电路带来的挫折和技术挑战(可能误触发),但作为回报,它塞进了两个概念。即使模型变大,只要还有更多未表示的概念,这个权衡就仍然存在。

It certainly happens at least in medium-sized networks. As you make the model larger, there's also a larger set of things the model could represent that are useful. So the number of neurons is increasing, but the set of things you want to store is also increasing. So there's a trade-off: it can either have fewer concepts, so it says 'this is only cars, not the queen', or it can accept some frustration and technical challenges from a circuit doing two things that might misfire, but in return it gets two concepts crammed in. Even as it gets bigger, as long as there are more concepts not yet represented, it still has that trade-off.

Host

这让你困扰的原因是,多语义电路理解起来更麻烦,所以会拖慢进度,而且你不太确信自己完全理解了它。

And the reason this troubles you is that a polysemantic circuit is much more of a pain to understand, so it slows you down, and you're less confident you understood everything about it.

Chris Olah

是的。你甚至可以想象一些东西在每个神经元中都非常微小,但隐藏在神经元之间,与几乎所有神经元正交,所以很难看到它们,但它们可能仍在模型中被表示。从这种视角很难捕捉到。这实际上就是为什么之前你问我什么是特征,我说通常我们认为特征是神经元,但我们不直接这么说,因为似乎有些情况下特征实际上是神经元的某种组合,或者一个神经元对应多个特征。理想的情况是找到某种方式来表示事物,使它们都是分离的。这与机器学习中关于解耦表示的文献有关,虽然有点不同,因为通常他们知道一些特征,比如性别或发色,并希望解耦它们,而我们没有这个优势。但这是一个密切相关的问题。

Yes. And you could even imagine things that are very small in every single neuron but are sort of hiding between the neurons, orthogonal to almost all of them, so it's very difficult to see them but they might still be represented in the model. This is very hard to catch from this lens. This is actually why earlier you asked me what a feature was, and I said usually we think of features as neurons, but the reason we don't just say that is that there seem to be cases where features are actually sort of combinations of neurons, or where a neuron corresponds to multiple features. The ideal would be to find some way to represent things so that they are all separate. This is related to a literature in machine learning where people talk about disentangling representations, though it's a bit different because usually they know a couple features like gender or hair color and want to disentangle them, whereas we don't have that benefit here. But it's a closely related problem.

可解释性成功未来的美学愿景 Aesthetic vision of a successful future with interpretability

Chris Olah

我认为有必要说说如果这件事能成功会有多美好。它有可能让神经网络变得更安全,而且从美学角度看,如果能生活在一个我们可以从神经网络中学到很多东西的世界,那也会非常美妙。我已经从它们身上学到了很多——比如如何对狗进行分类这种小事,还有很多我以前不理解的东西。你可以想象一个神经网络安全但未来却有些可悲的世界,因为我们变得无关紧要,不明白发生了什么。我认为未来即使有非常强大的 AI 系统,也不一定是那样——一个更人性化的世界,我们能够理解事物并推理它们。这激励我从事这项工作。

I think there's something worth saying about how wonderful it would be if this could succeed. It's potentially something that makes neural networks much safer, but there's also a way in which it would be aesthetically wonderful if we could live in a world where we can learn so much from these neural networks. I've already learned a lot from them—silly things like how to classify dogs, but also many things I didn't understand before. You could imagine a world where neural networks are safe, but the future is kind of sad because we're irrelevant and don't understand what's going on. I think there's potential for a future with very powerful AI systems that isn't like that—a more humane world where we understand things and can reason about them. That motivates me to pursue this line of work.

Host

所以这个想法是,未来神经网络可能会做很多以前人们做的事情,在某种意义上发号施令,但希望与我们的利益一致。如果它们是黑箱,我们无法理解,那可能会是一个令人疏远的世界。但如果我们能像看待机器一样看待它们——比如汽车,你可以谷歌一下它是如何工作的,世界就变得可理解了——那会让我觉得自己没那么没用。我们也有可能从这些神经网络中学到东西,比如它们思考的巧妙方式,甚至可能复制那种方式。

So the idea is that in the future, neural networks might do a lot of things people used to do, calling the shots in some sense, but hopefully aligned with our interests. It could be an alienating world if they are black boxes we can't comprehend. But if we can look at them like machines—like cars, where you can Google how they work and the world feels comprehensible—that makes me feel less useless. We can also potentially learn things from these neural networks, like the clever ways they think, and maybe copy that.

Chris Olah

有一种显微镜 AI 的想法。人们谈论智能体 AI 去做事,预言 AI 给出明智建议。另一种愿景是显微镜 AI,让我们更好地理解世界,分享它的理解,使我们更聪明,拥有更丰富的视角。这更难实现,可能竞争力也较弱,但我认为它非常美好。这就是可解释性能够实现的东西。我认为只有当我们真正成功理解模型时才有可能,但从美学上我就是更喜欢它。

There's this idea of a microscope AI. People talk about agent AIs that go and do things, and oracle AIs that give wise advice. Another vision is a microscope AI that allows us to understand the world better, sharing its understanding in a way that makes us smarter and gives us a richer perspective. This is harder to achieve and probably less competitive, but I find it really beautiful. This is the kind of thing interpretability would enable. I think it's only possible if we really succeed at understanding models, but aesthetically I just prefer it.

可解释性防止失败的玩具故事 Toy story of interpretability preventing failure

Host

你能描述一个可能的故事,其中可解释性研究帮助我们预见神经网络部署时的失败模式并提前化解它吗?一个小故事。

Can you describe a possible story where interpretability research helps us foresee a failure mode of a neural network when deployed and diffuse it ahead of time? A toy story.

Chris Olah

有几种不同的故事。在最极端的情况下,想象我们完全理解变革性 AI 系统——它们内部的一切。我们可以确信没有不安全的事情发生;它们没有撒谎或操纵我们,只是真诚地试图提供帮助。一个更强的版本是,我们理解得如此透彻,以至于我们自己变得更聪明,拥有一个显微镜 AI,赋予我们创造美好未来的能力。

There are a number of different stories. On the most extreme side, imagine we fully understand transformative AI systems—everything inside them. We can be confident nothing unsafe is going on; they're not lying or manipulating us, just genuinely trying to be helpful. An even stronger version is that we understand them so well that we become smarter ourselves, with a microscope AI that empowers us to create a wonderful future.

Chris Olah

现在想象可解释性没有那样成功。我们无法完全理解变革性 AI 系统。那怎么办?也许我们可以仔细分析小片段——比如理解社会推理,判断模型是否在操纵。这仍然可以降低安全担忧。但也许连这也太难了。那么我们就退而求其次,以一定概率捕捉问题。我们观察系统,捕捉一些问题,不声称能捕捉所有,但以一定概率捕捉到本来会很糟糕的问题,然后我们有机会重来。或者随着我们构建更强大的系统,我们以一定概率捕捉问题,帮助社会校准这些系统的风险程度。

Now imagine interpretability doesn't succeed that way. We can't totally understand a transformative AI system. What then? Maybe we can do careful analysis of small slices—like understand social reasoning and whether the model is being manipulative. That could still reduce safety concerns. But maybe even that is too much. Then we fall back to catching problems with some probability. We look at the system and catch some problems, not claiming to catch all, but with some probability we catch things that would have been really bad, and we get to start over. Or as we build more powerful systems, we catch problems with some probability, helping society calibrate how risky these systems are.

Host

所以部署神经网络就像汽车或新药——你测试并以概率方式捕捉问题。

So it's like deploying neural networks is like cars or a new drug—you test and catch problems probabilistically.

安全愿景对比:渐进与快速起飞 Comparing safety visions: gradual vs rapid takeoff

Host

你知道,就像农药——我们并不是想要零失败,只是希望有更好的系统,更有可能发现有害副作用或它们会失败的方式,尽可能多地预见问题并预防它们,这样更好。我想有些人的观点更偏向 AI 的快速起飞——我们会很快达到超级智能,任何漏掉的错误都可能导致灾难——这是 AI 可能影响世界的非常不同的愿景。从这个角度来看,仅仅说‘我们发现了 80%的问题’可能不会让人那么放心。你怎么看?或者也许你认为这种渐进式部署越来越强大的神经网络是最可能发生的情景。

You know, pesticide—it's not that we want zero failures, it's just that we want to have better systems to be more likely to find problematic side effects or ways they're going to fail as often as possible. Because the more we manage to foresee problems and prevent them, the better. I suppose some people who are operating more within a very rapid takeoff of AI—we're going to reach superintelligence really quickly, and any errors that slip through risk catastrophe—that's a very different vision for how AI could end up affecting the world. From that point of view, just being like, 'Well, we found 80% of the problems' maybe wouldn't put people's minds at ease so much. Any thoughts on that? Or maybe it's just that you think this more gradual deployment of ever more powerful neural networks is by far the most likely scenario to play out.

Chris Olah

首先,我仍然渴望发现所有问题。我不知道我是否能做到——我不知道这种方法是否会成功。但我确实认为,从某种意义上说,我回答这个问题的方式是:在逐渐变得更糟的世界里,我们是否有助于安全?我们如何为安全做出贡献?随着我们成功减少,我们对安全的贡献也变小了。希望它们仍然有帮助。我确实认为,那些持有相对快速起飞世界观的人——除非他们的世界观非常非常快速——应该对任何可能帮助社会更准确地评估 AI 系统风险的事情感到兴奋。因为如果你有具体的例子,并且抓住了系统不安全的实例,那么关于部署变革性 AI 系统所需的协调似乎更有可能实现。围绕一个完全抽象的问题协调人们似乎非常困难。所以即使是在更弱的系统上,如果你能提供令人信服的安全问题例子,或者系统在某种意义上具有操纵性或背叛性,我觉得更有希望去说服人们采取行动。

First, I'd say I still am aspiring to catch all the problems. I don't know that I'm going to—I don't know this approach is going to succeed. But I do think that in some ways, the way I'm answering this question is: in progressively worse worlds, do we contribute to safety? How do we contribute to safety? As we succeed less, our contributions to safety are smaller. Hopefully they're still helpful. I do think that people who have a relatively fast takeoff worldview—unless they have a very, very fast takeoff worldview—should be kind of excited about anything that might help society become more calibrated on whether there are risks from AI systems. Because the kind of coordination you'd want to have about deploying a transformative AI system seems much more likely if you have concrete examples and have caught concrete examples of systems being unsafe. It seems really hard to coordinate people around a completely abstract problem. So even if it's with much weaker systems, if you can have compelling examples of safety problems or systems in some sense being manipulative or treacherous, it seems to me much more optimistic about going and persuading people to take action.

可解释性研究的未来 Future of interpretability research

Host

有道理。那么可解释性研究的未来是什么?未来几年我们能看到哪些有希望的前景?

That makes sense. So what's the future for interpretability research? Are there any really promising horizons we can expect to see over the next couple of years?

Chris Olah

我认为问题不是‘有没有有希望的前景?’而是我不知道——在大量的有希望的前景中。我觉得这一定就像参与一个非常早期的科学领域,那里有那么多唾手可得的成果。在我看来,你可以朝一百万个方向前进,每个方向都会发现关于神经网络的惊人事物。我所看到的每一处都是未开垦的肥沃土地。我的意思是,你只需打开一个神经网络,仔细看,很可能发现别人从未见过的东西。也许有一些特别高影响力的东西,比如发现基序或更大规模的结构,或者在语言模型上取得进展——这些比其他东西影响力更大。但想象一下,这就像深度学习生物学的早期。每个方向都有惊人的发现。

I think the question isn't 'are there promising horizons?' but I don't know—of the immense number of promising horizons. I feel like this must be what it's like to be involved in a very early scientific field where there's so much low-hanging fruit. It just seems to me that you can go in a million directions and you'll find amazing things to discover about neural networks, and every single one of those directions. There's just fertile unfarmed ground everywhere that I can see. I mean, you can just sort of crack open a neural network, and if you're looking carefully, with high probability you can discover something that no one else has seen before. Maybe there are particularly high-impact things like discovering motifs or larger scale structures, or making progress on language models—these are higher impact than other things. But imagine this is like the early days of the biology of deep learning. There's just every direction with amazing things to find.

进入可解释性研究的第一步 First steps into interpretability research

Host

好的,考虑到这一点,如果人们真的想进入可解释性议程并做出贡献,他们应该采取哪些第一步?我想我们在其他节目中已经讨论过如何追求机器学习的一般职业,但对于那些已经走向机器学习或 AI 安全职业的人来说,他们如何推动自己走向可解释性方向?

All right, well with that in mind, what are the first steps people should take if they're really keen to get into this interpretability agenda and contribute to it? I'm thinking we've had other episodes where we've talked about how to pursue a career in machine learning in general, but maybe for someone who is already heading towards a machine learning or AI safety career, how do they push themselves in the interpretability direction?

Chris Olah

嗯,再次说明,可解释性领域有很多不同的事情,我的建议只在你对这类电路式可解释性感兴趣时才有用。我的第一个建议是仔细阅读电路线程。我认为这是尝试这种特定风格可解释性研究最集中的例子。再说一次,这是我自己的工作,所以我有偏见,但我认为如果你对这种工作感到兴奋,那可能是一个很好的起点。我认为尝试真正深入理解神经网络架构,思考如何在其中实现不同的事物,花时间自己摸索——所有这些都非常有用。很多关于如何进入机器学习并建立相关软件工程技能和理解神经网络理论的标准建议都非常适用。我还建议培养一些非常基础的数据可视化技能。如果你有这些技能,它会对你很有帮助。

Well, again, I'll caveat that there are lots of different things people are doing within the interpretability space, and my advice is only very useful if you're interested in this kind of circuit-style interpretability. My first piece of advice would be to read the circuits thread carefully. I think it's the most concentrated example of trying to do this particular style of interpretability research. Again, it's my own work, so I'm biased, but I think if you're excited about this kind of work, that's probably a good starting point. I think trying to really deeply understand neural network architectures, trying to reason about how you would implement different kinds of things in them, spending time poking around with them yourself—all of that is really useful. A lot of standard advice about how to get into machine learning and build up skills in relevant software engineering and understanding the theory of neural networks is all very applicable. I'd also try to develop some very basic data visualization skills. That's something that will serve you really well if you have it.

Host

有没有什么会议或社交活动,人们最终可以与该领域的人交流以取得进展?他们如何做到这一点?

Are there any conferences or social events where people can eventually talk to others in the field to move further? How might they be able to do that?

Chris Olah

同样,这种特定方法相当小众。如果你阅读电路线程,可以加入一个 Slack 社区,我们中的一些人也在 Twitter 上很活跃。我真的很想组织一些围绕这个的社交活动,但 COVID 让这变得困难。但我想人们可以在 Twitter 上关注我,而且这个子领域会发展,最终会有我们可以参加的活动来建立联系和保持更新。我想你会在标准的 ML 会议上做报告,对吧?嗯,我有时会在会议上做演讲。我主要在 Distill 上发表文章,因为我们有这个允许交互式可视化的平台。所以对于这个特定的小众领域,很多东西都集中在 Distill 和 Distill Slack 社区周围。

Again, this particular approach is pretty niche. There is a Slack community that you can access if you read the circuits thread, and also a number of us are pretty active on Twitter. I've really wanted to organize some kind of social event around this, but COVID has made that hard. But I guess people can follow me on Twitter, and presumably the subfield is going to grow, and eventually there will be events we can go to network and keep up to date. I suppose you present at the standard ML conferences, right? Well, I sometimes give talks at conferences. I mostly publish on Distill since we have this venue that allows for interactive visualizations. So for this particular niche, a lot of stuff is concentrated around Distill and the Distill Slack community.

转向扩展律讨论 Transition to scaling laws discussion

Host

太好了。好吧,我想这结束了我们关于可解释性的漫长而实质性的讨论。但在我们停止讨论技术问题之前,我们的听众非常希望我问一下关于缩放定律的问题。这似乎是当下的热门话题。

Beautiful. All right, I guess that brings to an end our very long, substantial discussion of interpretability. Before we stop talking about technical issues though, boy oh boy did our listeners want me to ask about scaling laws. Seems to be really the talk of the town.

扩展律简介 Introduction to Scaling Laws

Host

关注 OpenAI 或机器学习的人,其实我对缩放定律了解不多。所以也许你可以解释一下,为什么每个人都让我问你关于缩放定律的事?我和他们可能需要了解些什么?

People who are following OpenAI or ML, I actually don't know all that much about scaling laws. So maybe you can explain to me why everyone is asking me to ask you about scaling laws. What maybe what I and they should need to know about these things?

Chris Olah

在我说任何话之前,我得先声明:我有不少密切合作者研究缩放定律,但我本人没有做过这方面的工作,所以真的不是专家。我会尽量聪明地谈论这个问题,但如果我说了什么特别聪明的话,那很可能归功于他们和与我交流过的专家。如果我说了什么蠢话,那很可能只是我误解了什么。这是我对所有关于缩放定律的发言的一个免责声明。

Before I say anything, I should say I have a number of close collaborators who work on scaling laws. I personally haven't done any work on scaling laws, so I'm really not an expert. I will try to answer and talk about this as intelligently as I can, but if I say anything that's really clever, it's probably due to them and experts I've talked to. And if I say anything really stupid, it's probably just that I misunderstood something. So that's a caveat I want to put on anything I say about scaling laws.

Host

好的。首先,也许我们应该解释一下什么是缩放定律。

Yeah, all right. I guess first off, maybe we should explain what scaling laws are.

Chris Olah

基本上,当人们谈论缩放定律时,他们真正指的是在对数-对数图上有一条直线。你可能会问为什么我们在意对数-对数图上的直线。你可能会问:‘克里斯,坐标轴上是什么?’这取决于具体情况。不同的事物有不同的缩放定律,坐标轴取决于你谈论的是哪个缩放定律。但可能最重要的缩放定律,也是人们最兴奋的那个,是模型大小在一个轴上,损失在另一个轴上。损失类似于性能;高损失意味着性能差。

Basically, when people talk about scaling laws, what they really mean is there's a straight line on a log-log plot. You might ask why we care about lines on log-log plots. You might ask, 'Chris, what's on the axes?' Well, it depends. There are different scaling laws for different things, and the axes depend on which scaling law you're talking about. But probably the most important scaling law, the one people are most excited about, has model size on one axis and loss on the other. Loss is like performance; high loss means bad performance.

Host

好的,所以观察结果是,随着模型变大,损失下降,意味着模型性能更好。而且这条直线在很大范围内都惊人地直。

Okay, so the observation is that there is a straight line where as you make models bigger, the loss goes down, meaning the model is performing better. And it's a shockingly straight line over a wide range.

Chris Olah

是的。

Yes.

Host

所以如果我增加模型大小,我回想起经济学中,如果是对数-线性图,10% 的增长总是同等价值。但如果是对数-对数图,那么如果我增加模型大小 10%,损失会怎样?我对对数-对数图没有直观的理解。

So if I increase the model size, I'm just remembering from economics that if it's a log-linear graph, a 10% increase is always as valuable. But if it's log-log, then if I increase the model size 10%, what happens to the loss? I don't really have an intuitive grasp of what a log-log graph is.

Chris Olah

老实说,我经常以对数-线性的方式思考这个问题,因为一个轴上的单位通常在对数空间中思考很自然,然后你可以把它看作指数关系。但如果你想同时考虑两个变量,你会得到 y = a * x^k 的形式。例如,y = x^2 就是一个幂律。我们在物理学中经常看到这些,比如开普勒定律之一,将轨道周期和半长轴联系起来。那就是一个幂律。所以如果你将一个量加倍,另一个量就会按那个幂次增加。如果是平方根,你放大 10 倍,另一个量就会变化平方根 10 倍。

Honestly, I often think about this in a log-linear way because the units on one axis are often natural to think about in log space, and then you can think of it as an exponential. But if you want to think about both variables, you get something of the form y = a * x^k. For example, y = x^2 would be a power law. We see these all the time in physics, like one of Kepler's laws relating orbital period and semi-major axis. That's a power law. So if you double one thing, you increase the other by that power. If it was square root, and you do 10x, you'd change one by the square root of 10.

Host

好的,所以是某种幂律关系。我不确定我完全理解了,但底线是,人们已经更新了观点,认为我们并没有像预期的那样遇到收益递减。所以就像增加神经元数量、增加连接数量、增加参数数量,性能就会持续以不错的速度提升,即使对于已经非常大的模型也是如此。这就是它的底线吗?

Okay, so it's some sort of power law. I'm not sure I completely followed that, but the bottom line is that people have updated in favor of the view that we don't hit declining returns as much as we expected from having bigger models. So it's just like increase the number of neurons, increase the number of connections, increase the number of parameters, and the performance just keeps getting better at a decent clip, even for models that were already very big. Is that kind of where this bottoms out?

Chris Olah

是的,而且我认为这不仅仅是没遇到收益递减。当然,所有这些到目前为止都是观察到的;这可能会改变。可能下一次模型大小的增加会突然打破缩放定律,不再以这种方式表现。但它是如此直的直线,这非常令人惊讶,并且诱使我们进一步外推。也许你可以在一定程度上推理出比你现在能训练的模型更大的模型。这就是为什么人们对缩放定律感到兴奋。

Yeah, and I think maybe it's not just that you're not hitting diminishing returns. Of course, all of this has been observed so far; this could potentially change. It could be that the next increase in model size suddenly breaks the scaling law and it no longer performs this way. But the fact that it's such a straight line is very surprising and makes it tempting to extrapolate further. Perhaps you can reason about models to some extent that are larger than the models you can train right now. That's why people are excited about scaling laws.

Host

好的,所以他们用它来向前预测,说:‘嗯,到 2022 年,我们将能够制造一个这么大的模型,然后我们应该期望这样的性能水平。’哇,那不是很好的性能水平吗?

Okay, so they're using it to project forward, saying, 'Well, in 2022 we'll be able to make a model that's this much bigger, and then we should expect this level of performance.' And like, wow, isn't that a great level of performance?

Chris Olah

是的。顺便说一句,还有其他缩放定律。有关于模型大小的,也有关于你使用的总计算量的。因为当你让模型更大时,你也必须训练它们更长时间。所以你可以问,如果我有给定的计算预算,在模型大小和训练时间之间如何最优分配计算,以及给定最优分配后我得到什么损失?这也遵循缩放定律。还有一个缩放定律是关于你增加数据量的。所有这些之间都有相互作用。人们有时会做一个类比,说这有点像统计物理。关于气体有一个惊人的事实:你可以尝试推理每个气体粒子,但结果发现温度、压力、体积和熵等少数几个东西真正说明了问题。类似地,神经网络似乎也有少数几个变量,至少在高层面上说明了很大一部分故事。这在某些方面几乎与我们之前讨论的电路相反,我们深入细节。这里你有一个极其抽象的观点,只有少数几个变量重要,并且它们相互作用。

Yeah. And there are other scaling laws by the way. There are ones with respect to model size, also ones with respect to the total amount of compute you use. Because as you make models larger, you also have to train them for longer. So you can ask, if I have a given compute budget, what is the optimal allocation of compute between model size and training time, and what loss do I get given that optimal allocation? That also follows a scaling law. There's also a scaling law as you add increasing amounts of data. There's an interaction between all of these. One analogy people sometimes make is that it's kind of like statistical physics. There's this amazing fact about gases where you could try to reason about every gas particle, but it turns out that a few things like temperature, pressure, volume, and entropy really tell the story. Similarly, there seem to be a few variables for neural networks that also tell a large part of the story, at least at a high level. This is in some ways almost the opposite of what we were talking about with circuits, where we go into the details. Here you have this extremely abstract view where just a few variables matter and you have them interact.

Host

我明白了。所以人们对此感到兴奋的一个原因是,也许我们从所有这些经验中学到了一些非常宏观的基本论断,关于训练中的计算量、参数数量、输入的数据量,以及系统的智能程度或性能高低之间的关系。这本身就令人惊叹,我们正在学习关于智能和信息处理本质的一些基本定律。然后当你向前预测时,说:‘嗯,未来我们将有这么多数据和这么多计算,这将支持一个具有这么多参数的模型。’这就更令人兴奋了。

I see. So one reason people might be excited about this is that maybe we're learning from all this experience some very big-picture fundamental claims about the relationship between the amount of compute you have in training, the number of parameters, and the amount of data you feed into it, and how intelligent or how high the performance of a system is. That's amazing in itself, learning some fundamental laws about the nature of intelligence and information processing. Then it's even more exciting when you project forward and say, 'Well, in the future we'll have this much data and this much compute, and that will support a model with this many parameters.'

迁移学习的扩展律 Scaling Laws for Transfer Learning

Host

这大概有助于我们预测未来的发展,而且似乎每隔几个月就会出现新的有趣的缩放定律,让你能够推理不同的事情。就在几周前,一篇关于迁移学习缩放定律的新论文发表了。迁移学习是一种流行的做法:通常一个人会训练一个大型模型,然后其他无法自己训练大型模型、数据也不多但有一些小数据集的人,会对那个大型模型进行微调以适应他们的任务。他们只是简短地训练一下,从一个别人在不同数据上训练好的模型开始,然后用自己的数据稍微训练一下。结果发现,实际上存在一个缩放定律,至少模型大小如何影响这一点,以及在某种意义上,训练数据和测试数据之间存在一个汇率。这是 Danny Hernandez 和他的合作者做的一些非常棒的工作。

It probably helps with our ability to forecast where we'll be in future and it seems like every couple of months there's new interesting scaling laws that allow you to reason about different things. So just a few weeks ago, a new paper came out on scaling laws for transfer learning. Transfer learning is this popular thing where often one person will train a large model and then other people who can't train large models themselves and don't have a lot of data but have small amounts of data will fine-tune that large model to their task. So they go and train it very briefly, starting with a trained model that somebody trained on different data, and then training it a little bit on their data. It turns out that there's actually a scaling law, very similar to how model size affects this, and there's in some sense an exchange rate between the data you trained on and the data you're testing on. So this is some really lovely work by Danny Hernandez and his collaborators.

Host

哦对,Danny Hernandez 上过我们的节目。我们会附上那期节目的链接。我知道这不是你的领域,但你怎么看?有些人问关于缩放定律的问题,也许他们应该了解这些,而且有没有可能他们误解了这些定律?

Oh yeah, Danny Hernandez has been on the show. We'll stick up a link to that episode. I know it's not your area, but what do you think? There are people asking about scaling laws maybe should know about them, and is there any way perhaps that they might misunderstand them?

Chris Olah

嗯,我认为他们之所以关注缩放定律,可能是因为这些定律揭示了未来神经网络的能力,我认为这是关心缩放定律的一个非常重要的原因。但我觉得人们可能低估了一点:缩放定律比这要通用得多,实际上可能非常有助于推理我们未来将构建的神经网络的属性,而不仅仅是它们的损失。对我来说,真正令人兴奋的版本——这里我只是受到 Jared Kaplan 对我说过的话的启发,可能只是复述得不太好——Jared 是在缩放定律方面做了很多开创性工作的人之一——那就是也许存在安全方面的缩放定律。也许在某种意义上,模型是否与你对齐可能是模型大小以及你给它多少信号来完成人类对齐任务的函数。我们或许能够推理比我们目前能构建的模型更大的模型的安全性。如果这是真的,那将是巨大的。我认为目前安全总是在追赶。如果我们能创造一种基于缩放定律来思考安全的方法,而不必追赶,那将是非常了不起的。所以我认为,除了缩放定律对能力的影响之外,这也是一个让人对它们感到兴奋的理由。

Well, I think that they are probably focused on them because of what they say about neural network capabilities in the future, and I think that is a really important reason to care about scaling laws. But I think that's something that people maybe underappreciate: scaling laws are much more general than that and actually may be very useful for reasoning about the properties of neural networks that we will build in the future, more generally than just their loss. The really exciting version of this to me—and here I'm just really inspired by things that Jared Kaplan has said to me, and probably giving a worse version of his thoughts—Jared is one of the people who's done a lot of leading work in scaling laws—is that maybe there's scaling laws for safety. Maybe there's some sense in which whether a model is aligned with you or not may be a function of model size and how much signal you give it to do the sort of human alignment task. We might be able to reason about the safety of models that are larger than the models we can presently build. If that's true, that would be huge. I think there's this way in which safety is always playing catch-up right now. If we could create a way to think about safety in terms of scaling laws and not have to play catch-up, I think that would be incredibly remarkable. So I think that's something that, separate from the capabilities implications of scaling laws, is a reason to be really excited about them.

Host

有意思。所以这个想法是,在 x 轴上我们有类似我们给了系统多少人类反馈来训练它,以确保它理解我们的偏好,在 y 轴上我们有它搞砸并完全误解我们偏好的频率。我想如果你知道缩放定律是什么,即这两者之间的关系,你还需要加入其他影响这种关系的因素,比如网络有多大,任务有多复杂,那就能让你预测未来某个可能还不存在的模型需要多少输入。

Interesting. So the idea there would be that on the x-axis we have something like how much human feedback we've given the system to train it to make sure it understands our preferences, and on the y-axis we have how often it messes up and totally misunderstands our preferences. I suppose if you know what the scaling law is, the relationship between those two things, and you'd also need to throw in other factors that affect this relationship, like how big the network is, how complicated the task is, that would allow you to project how much input you're going to need for some future model that maybe doesn't even exist yet.

Chris Olah

是的,或者一个非常疯狂的想法是,也许你有两个轴,然后在每个点上,模型与你的对齐程度之类的。也许这是你给它的偏好信号量和模型大小的函数。也许在某种意义上存在一个相变,比如对齐区域和未对齐区域。这非常推测性,但如果类似的事情是真的,那将非常了不起。

Yeah, or one really crazy picture you might have is maybe you have two axes, and then at every point, how aligned the model is with you or something like this. Maybe that's a function of the amount of human preference signal you gave it and the model size. Maybe there's in some sense a phase change somewhere, like the aligned regime and the unaligned regime. This is extremely speculative, but if something like that was true, that would be really remarkable.

Host

你认为还有哪些关于缩放定律的事情是人们应该知道的,或者应该更多关注的?

Are there other things that you think people should know about scaling laws or perhaps focus more on?

Chris Olah

我想另一件事是,缩放定律在整体损失上显示出非常平滑的趋势,随着模型变大,但如果你尝试做同样的分析,看看非常具体的能力,比如模型做某种算术的能力,你实际上会看到这些不连续的跳跃。很容易让人认为在这些点上可能存在某种相变或不连续的东西。我认为这有点联系到为什么思考可解释性和仔细审视这些系统很重要,因为它表明这些预测至少目前似乎没有告诉我们这些更大系统会发生什么的全部故事。事实上,至少在这些更精细的事情上,还有一些复杂的东西我们仍然想理解。

I guess one other thing is that scaling laws show these really smooth trends for the overall loss as you make models larger, but they actually show discontinuous trends if you try to do the same kind of analysis and look at very specific capabilities, like how well the model does certain kinds of arithmetic. You'll actually see these sort of discontinuous jumps. It's tempting to think that there's maybe some kind of phase change or something discontinuous going on at those points. I think it connects a little bit to why it's important to think about interpretability and to look carefully at these systems, because it suggests that these kinds of predictions don't, at least presently, seem to tell us the whole story of what's going to happen with these larger systems. In fact, there's something complex going on at least in terms of these finer things that we'd still like to understand.

Host

好的,是的。所以我们目前至少看到了一些不连续性,这意味着未来可能还会有不连续性,我们不应该排除这种可能性,并假设一切都是平滑的。这似乎表明,平滑的损失实际上可能是许多不连续的小能力在不同点上改进的结果,这些能力不连续地改进,总体上创造了这种平滑的过渡。但对于更具体的事情,典型情况可能更不连续,这意味着我们突然被更大模型的能力惊讶到的可能性比你可能认为的要大得多。我认为这就是为什么我们应该非常仔细地观察大型模型内部正在发生的事情。

Okay, yeah. So the fact that we've seen at least some discontinuity so far means that maybe there'll be discontinuities in future, we shouldn't rule that out and assume it's all smooth. And it sort of suggests that the smooth loss is actually maybe the result of many discontinuous small capabilities that are all improving at different points, discontinuously improving, and that in aggregate creates this smooth transition. But for more specific things, the typical case may be this more discontinuous thing, and it just means that there's a lot more room for us to suddenly be surprised by the capabilities of larger models than you might think. I think that's a reason why we should be looking very carefully at what's going on inside large models.

Host

好吧,我再问一次。关于缩放定律,还有什么人们应该知道或更多强调的吗?

Alright, I'll fish again. Is there anything else that people maybe should know about or give more emphasis to with regard to scaling laws?

Chris Olah

嗯,也许另一件事是,这让我们能够以更加系统化和严谨的工程方法来处理神经网络。与其在黑暗中摸索,我们正在构建一个更加系统化的方式来思考我们应该如何预期神经网络的性能以及什么才是真正重要的。我觉得在这样一个我们以这种方式处理问题的世界里,我会感到更安全。

Well, maybe one other thing is that this allows us to have a much more systematic and rigorous engineering approach to working with neural networks. Rather than taking shots in the dark, we're developing this picture of a much more systematic way to think about how we should expect neural networks to perform and what really matters. I think there's a way in which I feel safer in a world where that's how we approach things.

扩展律对小型研究者的影响 Scaling laws and their implications for smaller researchers

Host

那么,关于这些缩放定律,人们可能会有点沮丧的一点是,它们似乎在暗示越大越好。更大的模型、更多的算力,都对性能有很大影响。显然,大多数人无法接触到最大的模型或最多的算力;只有少数几个中心拥有那种巨额资金。这对那些预算较小、资源较少的研究人员意味着什么?他们还能做出贡献吗,还是有可能被排除在外?

So one thing that people might be slightly sad about with these scaling laws is that they kind of suggest that bigger is better. Bigger models, more compute, it all makes a big difference to performance. Most people obviously don't have access to the biggest models or the most amount of compute; there are only a few centers that have that kind of massive funding. What does that imply for people who are doing research but have smaller budgets and access to fewer resources? Can they still contribute, or do they risk getting a little bit cut out of things?

Chris Olah

我认为机器学习作为一个领域,在过去几年里结构上处于一个非常奇怪的位置,相当多的参与者都能进行相对前沿的工作。我认为这在许多其他领域并非如此。例如,如果你从事航空航天工程,能够建造全尺寸火箭进行测试的人可能只是该领域中的极小一部分。同样,在粒子物理学中,大型强子对撞机是一个巨大的设备,我很难在地下室里撞碎粒子。那么该怎么办呢?我认为一个答案——也许我对此有私心——是可解释性不需要这些。可解释性允许我们训练一次模型,然后让许多人去研究它,试图理解这些模型。所以我认为有可能不是每个人都训练一百万个模型,而是由少数参与者训练模型,然后人们去研究它们。尽管我认为在如何负责任且安全地向所有人提供这些模型方面仍存在问题。另一个答案是尝试做小规模的工作,让我们理解通过缩放定律开始看到的更大规模的图景。我认为如果你足够严谨和仔细,很可能可以说出有用的东西。事实上,我们已经看到许多缩放定律的论文在小规模范围内运作,并且看起来相当有趣。

I think that machine learning as a field has been structurally in a very strange position for the last number of years, where quite a large number of actors could go and do relatively state-of-the-art work. I think that's not true in many other fields. For example, if you're doing aerospace engineering, the number of people who can go and build a full-scale rocket ship to test is presumably a very small subset of people working in that field. Similarly, in particle physics, the LHC is an enormous thing, and it's hard for me to compete smashing particles in my basement. So what is one to do? Well, I think one answer—and maybe I have a selfish interest for pushing this—is that interpretability doesn't require this. Interpretability allows us to train models once and then potentially for many people to go and study it and try to understand these models. So I think it's possible that rather than having everybody training a million models, we might have a smaller number of actors training models and people studying them. Although I think there are still questions about how we can responsibly and safely provide all those models to everyone. Another answer is to try to do small-scale work that allows us to understand this larger-scale picture that we're starting to see through scaling laws. I think it's probably possible, if you're rigorous and careful enough, to say useful things. In fact, I think we've seen a number of scaling laws papers that operate in smaller-scale regimes and seem quite interesting.

算法进步与计算和数据扩展 Algorithmic progress vs. compute and data scaling

Host

一位听众提出了一个问题:到目前为止,很多进步似乎是人们使用相同或相当相似的算法,但投入了更多的数据和算力,然后观察性能能提升多少。底层算法是否有有趣的进步?算法是否非常不同?整体进步中有多少是由算法驱动的,又有多少是由数据和算力增长驱动的?

One listener wrote in a question along the lines of: it seems like a lot of progress to date has been people using the same algorithms or reasonably similar algorithms over time, but throwing much more data and compute at them, and seeing how much their performance can go up. Have there been interesting advancements in the underlying algorithms? Are the algorithms very different? And how much of the overall progress has been driven by algorithms versus data and compute increases?

Chris Olah

当然有算法上的改进。向 Transformer 的转变可能是近年来最引人注目的。Danny Hernandez 在 OpenAI 博客上有一篇文章,讨论了训练神经网络的效率随时间的变化,以及算法改进如何提高了效率。在其他领域也有很多令人兴奋的工作,比如图神经网络的进展。机器学习正在向各个方向发展。但我想反驳这个问题的前提,即算力增长带来更强能力这件事有些无聊。实际上,我认为这完全不无聊;它真的很美,少数几个因素驱动着整体图景,正如缩放定律所暗示的,产生了巨大的结构。这就像说宇宙很无聊,因为它遵循简单的物理定律,或者说进化很无聊,因为它只关心生存。但我认为这根本不无聊;这是它美的一部分。我们观察到的结果真的很壮观。这是另一个美丽的例子,说明通过非常简单的底层过程可以获得无限的复杂性并完成如此多的事情。

There certainly have been algorithmic improvements. The transition to Transformers is probably the most striking one in recent years. Danny Hernandez has a blog post on the OpenAI blog about the efficiency of training neural networks over time and how algorithmic improvements have increased that. There's also been lots of exciting work in other domains, like progress on graph neural networks. Machine learning is going in all sorts of directions. But I want to push back against the premise of this question that there's something boring about increases in compute leading to greater capabilities. I actually think it's not at all boring; it's really beautiful that a small number of things are driving the big picture story here, as scaling laws suggest, giving rise to immense structure. It's like saying the universe is boring because it runs on simple physical laws, or evolution is boring because it just cares about survival. But I think it's not boring at all; it's part of the beauty. The things we observe resulting from this are really gorgeous. It's another beautiful example of being able to get unlimited complexity and accomplish so much with very simple underlying processes.

Anthropic 的愿景与安全焦点 Anthropic's vision and safety focus

Host

我们来谈谈最近发生的一件令人兴奋的事:你和一些同事最近离开了 OpenAI,启动了一个名为 Anthropic 的新项目,最近刚刚公开宣布。我看到你们现在有了一个网站,上面有很多招聘广告。为了让大家了解背景,Anthropic 将如何为 AI 和 AI 安全领域做出贡献?

Let's talk about something pretty exciting that's happened recently: you and some of your colleagues have recently left OpenAI to start a new project called Anthropic, which was announced publicly just very recently. I see you've now got a website with a bunch of job ads. To set the scene, what's the vision for how Anthropic is going to contribute to the AI and AI safety space?

Chris Olah

Anthropic 是一家新的人工智能研究公司,专注于大型模型的安全性。我们试图让大型模型变得可靠、可解释和可操控。具体来说,这意味着我们正在训练大型模型,研究大型模型背景下的可解释性和人类反馈,并深入思考这对社会的影响。

Anthropic is a new AI research company focused on the safety of large models. We're trying to make large models reliable, interpretable, and steerable. Concretely, that means we're training large models, studying interpretability and human feedback in the context of large models, and thinking a lot about the implications for society.

Host

Anthropic 最独特的一点是,你们专注于大型模型的安全性和构建,也许是你们能管理的最大的模型。有些人可能会担心,既然你们在开发这些最令人担忧的大型模型,你们可能也在加速向更大、因此更危险的模型发展。你怎么看?

One of the most distinctive things about Anthropic is that you're focusing on the safety and construction of large models, perhaps the largest models you can manage. Some people might be worried that since you're working on those large models, which are the models of greatest concern, you might also be accelerating progress towards larger and therefore more dangerous models. What do you think about that?

Chris Olah

我非常理解这种担忧。我认为大型模型是……

I have a lot of sympathy for that concern. I think that large models are...

研究大型模型的主要论点 Primary argument for working on large models

Host

这可能是可预见的未来中最大的 AI 风险来源,所以担心那些可能加速这一进程的事情是相当合理的。但这正是我认为我们必须致力于其安全性的原因,我们应该在开发大模型时保持负责任和深思熟虑。是的,我认为如果你知道某件事即将到来,并且认为它可能很危险,那么你很可能希望专注于这方面的安全工作。

Probably the greatest source of AI risk in the foreseeable future, and so it's pretty reasonable to be concerned about things that might accelerate that. But that's exactly why I think it's really important for us to be working on their safety, and we should try to be responsible and thoughtful about working with large models. But yeah, I think if you know that something is coming and you think that it might be dangerous, then you probably want to do safety work focused on that.

Chris Olah

好的,所以你的总体看法是,这类工作可能会加速更大模型的发展,但这种影响被抵消了——我们无论如何都应该做,因为还有其他可能造成最大损害的事情。理解它们并找出如何让它们安全是进行安全工作的最有效方式,这是主要的考量。

Okay, so it sounds like your general take is that it is possible that this kind of work might accelerate the development of larger models, but that is kind of outweighed—that we should do it anyway because there are other things that might do the most damage. Working on understanding them and figuring out how to make them safe is the most effective way to work on safety, and that's kind of the dominant consideration.

Host

是的,为了更精确一点,我认为实际上有一个主要论点,然后还有两个次要论点也值得考虑。我先说主要论点。在我看来,大模型是某种非常重大的事件,基本上在这一点上是不可避免的。社会似乎正沿着一条相当直接的道路前进,人们正在构建越来越大的模型。同时,这些大模型与小模型相比在性质上有很大不同。一个具体的例子是,一些最大的语言模型在某种意义上会对你说谎。你去问它们问题,它们明显知道答案,但当你问它们时,它们仍然给出错误的答案。也许更准确的说法是,它们不是在说谎,而是在某种意义上没有给出它们知道的真实答案。这在较小的模型中观察不到,因为小模型连说连贯的话都很困难,所以不存在说谎的问题。因此,如果你相信大模型是不可避免的,并且大模型在性质上足够不同,以至于最有价值的安全研究需要针对它们才能有效,那么我认为你面临两个选择。第一个选择是,你可以对其他人生产的模型进行安全研究,这意味着根据具体设置,你可能会有一到三年的滞后。在这种情况下,安全研究将永远在追赶,研究的是较旧的模型,试图用落后几年的模型进行安全研究。第二个选择是,你尝试创建一个 Scaling(规模扩张)项目,并进行与安全紧密结合的 Scaling(规模扩张)研究,这样你就可以在最大的模型上进行安全研究。所以,我有时在安全社区听到的一个担忧是:如果重要的安全研究只能在最后阶段进行呢?比如,有一天我们将构建非常强大的系统,真正变革性的 AI 系统,如果有价值的安全研究只能在这些系统中进行呢?我认为,如果你担心这种事情,那么你应该相当兴奋,能够缩短几年时间似乎非常重要,这实际上为我们争取了更多时间来研究这些系统。这就是主要论点。

Yeah, so I think to be a little more precise, I actually think there's sort of one primary argument and then there are actually two secondary arguments that are also worth considering. So I'll give you the primary argument first. It seems to me like large models are something really dramatic happening, basically inevitable at this point. Society seems like we're on a pretty straightforward path of people building larger and larger models. And at the same time, it seems like these large models are pretty qualitatively different from smaller models. So a concrete example of that is some of the largest language models will, in some sense, lie to you. You go and ask them questions, and they demonstratively know the answer to the question, but they still give you the wrong answer when you ask them. Maybe it would be more accurate to characterize that as rather than lying, but they're in some sense not giving you true answers that they know. And that's not a problem you can observe in smaller models, because smaller models struggle enough with saying something coherent that you don't have any problems with lying. So if you believe that large models are inevitable and that large models are qualitatively different enough that the most valuable safety research is going to need to work on them to be effective, then I think you're left with two options. Behind door one, you can go and do safety research on models that other people have produced, and that means that you probably have, depending on the exact setup, something like a one to three-year lag on the models that you're working with. And in that world, I think safety research is sort of perpetually playing catch-up, where it's working on these older models and then trying to do safety research with models that are several years behind. And the second option is that you try to go and create a scaling effort and do scaling research that's very tightly integrated with safety, such that you can be doing safety on the largest model. So I guess, you know, I feel like a concern that I sometimes hear in the safety community is: what if the important safety research can only be done at the end? Like, someday we're going to build really powerful systems, really transformative AI systems, and what if the valuable safety research can only be done in those systems? Well, I feel like if you're worried about that kind of thing, then you should be pretty excited, and it seems really important to go and be able to shave off a couple of years, and that sort of effectively buys us more time to work on those systems. So that's the primary argument.

次要考虑与实证问题 Secondary considerations and empirical question

Host

好的,所以一个关键动机是,目前安全研究总是落后于前沿,因为前沿是这些最大的模型,而安全研究不一定与真正处于能力前沿的东西挂钩。通过让 Anthropic 专注于最大的模型,你保持了同步,这样如果最新能力进展提出了新的考量,你就能以安全的心态去应用安全思维。我想在某种程度上,这是否净收益必须是一个开放或经验性问题。至少如果你认为加快最大模型的进步速度是有害的,那么这是一个经验问题:你在多大程度上加速了它们,又在多大程度上让它们更安全?人们必须做出判断。但听起来你认为,你对加速进步或扩大规模的影响,相对于你对它们如何工作以及如何让它们更好工作的洞察增益来说,是相对适中的。

Okay, so a key motivation is that at the moment, safety research is always behind the cutting edge because the cutting edge is these largest models, and the safety research isn't necessarily tied in to the stuff that's really at the cutting edge of capabilities. And by making Anthropic focus on the largest models, you're keeping it up to pace so that if there are new considerations raised by the latest advances in capabilities, you'll be there with a safety mindset in order to try to apply safety thinking to that. I guess to some extent, it has to be an open or an empirical question whether this is net beneficial. At least if you think that advancing the rate of progress within the largest models is harmful, then it's kind of an empirical question: how much do you speed those up versus how much do you make them safer? And one has to make a judgment call there. But I guess it sounds like you think the effect that you have on accelerating the progress or accelerating the size is relatively modest compared to the gain in insight that you have into how they work and how to make them work better.

Chris Olah

是的,我的意思是,我实际上可能不会把潜在危害描述为——我不认为能力进步本身是有害的,但我确实非常担心竞赛和加速竞赛,以及创造人们会在安全上偷工减料的激励。但我们可以对此非常负责。而且我还认为,至少如果你采取经验性的安全方法,如果大模型在性质上确实非常不同,而你无法接触到它们,那么取得进展是非常困难的。所以,我认为安全研究追赶是一场必输的游戏,我们需要去研究大模型。

Yeah, I mean, I would actually probably frame potential harm not as—I don't think that I see capabilities progress as intrinsically harmful, but I do worry a lot about races and accelerating races and creating incentives where people are going to cut corners on safety. But yeah, I think we can be pretty responsible about that. And I also think that, at least if you're taking an empirical approach to safety, it's really hard to make progress if large models are just really qualitatively different and you don't have access to them. And so yeah, I sort of think that safety research playing catch-up is a losing game, and we need to go and work on large models.

Host

是的,你说过如果不研究最大的模型,那么你可能会落后前沿 1 到 3 年,或者在那样的时间线上追赶。这个数字从何而来?

Yeah, you said that if you're not working with the largest models, then plausibly you could be 1 to 3 years behind the cutting edge or playing catch-up on that kind of timeline. Where does that number come from?

Chris Olah

是的,这只是我的估计,我认为这在很大程度上取决于具体的机构设置。所以我认为,如果你在一家训练大模型的公司,但没有与 Scaling(规模扩张)工作紧密集成,那么你可能会面临大约一年的延迟,甚至可能只有六个月。这是因为你需要专门的基础设施来处理这些模型,这些模型运行成本很高——就像训练成本很高一样——而且专业知识可能只存在于训练这些模型的人身上,他们可能不知道如何使用这些模型,而且在一个项目正在积极进行时很难融入其中。所以你最终不得不等待它被完善,然后以某种方式提供给外部人员使用,这会有相当大的延迟。而且你可能需要找到有处理这些模型经验的工程师来帮助你构建工具。

Yeah, so this is just my estimate, and I think it depends a lot on the exact institutional setup. So I think if you're at a company that's training large models but aren't really closely integrated with the scaling efforts in particular, then you might be looking at something more like a one-year delay, maybe even more like six months. And that's due to your need for this specialized infrastructure for working with these models, due to the fact that these models cost a lot to be able to do anything with—just like they cost a lot to train—and from the fact that expertise may only reside with people who are training these models and may not know how to work with them, and just that it's really hard to integrate with a project while it's actively being worked on. So you end up having to wait for it to be polished and then kind of handed out in a way that it's possible for external people to use, and that happens at some substantial delay. And you probably need to go and get engineers who have expertise working with these models to help you build the tooling.

学术界与工业界日益扩大的差距 The growing gap between academia and industry

Host

你需要在它们之上做有用的工作。另一个极端是,如果你在学术实验室,这个差距似乎正在急剧扩大。我猜很多学术实验室只有几块 GPU,缺乏训练哪怕中等规模模型的专业知识,他们使用的模型远远落后于最先进水平。我认为这中间有一个光谱,我的一到三年估计就是基于此。

You need to do useful work on top of them. On the other extreme, if you're at an academic lab, it seems like that gap is growing pretty dramatically. My guess is that a lot of academic labs only have a couple GPUs and don't have expertise training even moderately large models, and are doing work with models that's very far behind the state-of-the-art. I think there's a spectrum between there, and that's where my one to three years estimate came from.

Host

我明白了。学术界本身被排挤或挤出,是不是因为资金需求增长到超出正常学术资助规模,这本身就是一个问题?

I see. Is it a problem in itself that academia is being pushed out or crowded out by the fact that funding requirements have grown so large that they're beyond the normal size of academic grants?

Chris Olah

是的,我认为这是一个非常棘手的问题。看看其他资本成本更高的科学领域如何运作,可能会很有趣。如果你在设计飞机,你不会指望研究实验室能自己造飞机。当然,像 CERN 或哈勃太空望远镜这样的项目,也不是单个实验室能完成的。这些领域发展出了协调和进行高资本成本实验的机制。所以我认为我们正在经历一个痛苦的转型,从一个更像数学或传统计算机科学的领域——每个博士生都拥有最大限度的有用资源,你只需要自己的思考能力——转向一个像其他领域那样资本成本极高、需要不同组织方式的领域。

Yeah, I think it's a really tricky problem. It might be interesting to look to other areas of science that have higher capital costs to learn about how they operate. If you were designing airplanes, you wouldn't expect research labs to be able to go and build their own airplane. Certainly, something like CERN or the Hubble Space Telescope isn't something that an individual lab can do. These fields develop mechanisms for coordinating and doing high capital cost experiments. So I think it's a painful transition we're going through right now, moving from a field more like mathematics or traditional computer science, where every PhD student has the maximum amount of useful resources and you don't need anything beyond your own ability to think, to a field more like these others with really high capital costs, where you need to organize things differently.

第二个论点:即使危险,大型模型的安全研究也必要 Second argument: safety research on large models is necessary even if they are dangerous

Host

好的,之前你说你会用三个不同的论证角度来解释为什么这是净收益的,那是第一个。第二个是什么?

Okay, earlier you said there were three different lenses of argument you were going to use to explain why this was net beneficial, and that was the first one. What's the second?

Chris Olah

嗯,有人可能会说:如果 AI 系统如此危险,安全如此困难,即使你小心谨慎并努力研究安全,深度学习仍然对世界有害呢?为了论证起见,我们假设这是真的——尽管我很乐观地认为世界并非如此。但假设我们处于一个深度学习本质危险、大型模型最终有害且难以避免的世界。那么在我看来,在这样的世界里,我们可能希望大幅减缓大型模型的工作。也许我们应该暂停超过一定规模的模型,或者采取类似非常激烈的措施。我不认为这能默认发生。我猜唯一可能的方式是,我们有确凿的证据表明大型模型非常危险且必然如此。而如果你不研究大型模型,我看不到获得这种证据的途径。所以我认为,即使持有这种观点的人,也应该积极去做安全研究和对大型模型的严格分析,因为这是唯一的办法。这种研究对任何加速的影响将非常微小——这个领域已经有很多参与者在工作。我认为,如果你持有那种观点,你应该优化的主要目标是,有可能获得能够围绕大幅减速建立共识的证据。

Well, somebody might say: what if AI systems are just so dangerous and safety is so hard that even if you're careful and try to work on safety, deep learning is just bad for the world? Let's assume for the sake of argument that's true—though I'm pretty optimistic that's not what the world looks like. But let's assume we're in a world where deep learning is intrinsically dangerous and really large models at some point become harmful, and it's very difficult to avoid that. Well, it seems to me that in such a world, we probably want to cause a pretty dramatic slowdown in work towards large models. Maybe we should pass a moratorium on models above a certain size, or something very dramatic like that. I don't see that happening by default. My guess is that the only way that's going to happen is if we have really compelling evidence that large models are very dangerous and necessarily so. I don't see a way you're going to get that kind of evidence if you aren't working with large models. So I think even people who hold that view should be excited to go and do safety research and really critical analysis of large models, because I think that's the only way. The effect of this research on any acceleration is going to be pretty marginal—there are already lots of actors working in the space. I think the main thing you should be optimizing for, if that's your view, is the possibility that you can get the evidence that could build consensus around a really dramatic slowdown.

Host

好的,这个我能理解。第三个是什么?

Okay, that one makes sense to me. What's the third one?

第三个论点:AI 可缓解其他灾难性风险与道德灾难 Third argument: AI can mitigate other catastrophic risks and moral catastrophes

Chris Olah

嗯,最后一个论点是:我们生活在一个充满持续道德灾难的世界——工厂化养殖、全球贫困、被忽视的疾病——还有即将到来的灾难,比如气候变化或潜在的存在风险来源。这些都是极其棘手的问题。但我想,一旦我们假设存在来自 AI 的剧烈风险,你也必须假设存在一个世界,其中有对齐的 AI 系统可能极大地改善这些问题。所以,正如我感到有义务强烈担忧 AI 风险及其出错的方式一样,我认为我们也有巨大的责任确保不错过 AI 极大改善世界的机会。我认为那同样糟糕,同样是巨大的失败。

Well, the final argument is: we live in a world with a lot of ongoing moral catastrophes—factory farming, global poverty, neglected diseases—and also oncoming disasters like climate change or potential sources of existential risk. Those are extremely tough problems. But I think once we're supposing we're in a world where this kind of dramatic risk from AI exists, I think you also have to be supposing that you're in a world where you have AI systems that could potentially dramatically improve these problems if they were aligned. So just as much as I feel an obligation to worry intensely about AI risk and ways AI could go wrong, I think we also have a great responsibility to make sure that we aren't missing the opportunity for AI to dramatically improve the world. I think that could be equally bad and an equally big failure.

Host

所以这是一个存在已久的论点,大致是:即使你认为人工智能的快速发展是危险的,因为速度太快,我们不一定有时间充分准备,并找出更先进、更大模型所需的所有额外安全要求,但综合考虑,如果这些新技术、这些 AI 的巨大进步也能降低核战争风险、降低生物武器使用或可怕大流行的风险,或者只是国家间的战争,那么它总体上可能更安全。取决于你如何看待这项新技术的风险与它可能消除和取代的其他所有风险之间的比率,即使加速前进是危险的,它也可能最终是正面的,因为在此期间你花在其他悬在头顶的达摩克利斯之剑上的时间会更少。我想我们实际上可以链接到 Nick Beckstead 几年前制作的一个小数学模型,来尝试看看这个效应有多大。

So this is an argument that's been around for a while, which kind of runs: even if you're someone who thinks that faster advances in artificial intelligence would be dangerous because they're happening so fast that we won't necessarily have time to fully prepare for them and figure out all the additional safety requirements for more advanced larger models, it could nonetheless come out safer on net, all things considered, if those new technologies, those massive advances in AI, could also reduce the risk of nuclear war, reduce the risk of bioweapons being used or a terrible pandemic, or just a war between countries. Depending on the ratio you see in the risk of this new technology versus all the other risks that technology might then obviate and supersede, even if it's dangerous to go faster, it can come out positive on balance because you're going to spend less time in the meantime with all those other concerns like swords of Damocles hanging over you. I think we could actually put a link up to a little mathematical model that Nick Beckstead made a couple of years ago to try to see how large this effect is.

Chris Olah

是的,完全正确。不过我也想强调——我认为这在很大程度上取决于你的道德观——但对我来说,不仅要权衡可能损害遥远未来的存在风险,也要认真对待当前可以缓解的持续伤害。我理解,从许多功利主义观点来看,这可能是唯一重要的事情,但老实说,我对工厂化养殖、全球贫困感到非常不安。具体来说,在某种程度上,我情感上希望自己能从事那些工作而不是 AI,如果我认为那对我来说是最优选择的话。这并不是说它应该成为主导关切,但我认为它不应该被遗忘。

Yeah, that's exactly right. Although I also want to highlight—and I think this depends a lot on your moral views—but it's also important for me to weigh not just existential risks that might harm the far future, but also take seriously present ongoing harms that could be mitigated. I understand that from a lot of utilitarian views, that might be the only thing that should matter, but honestly I just feel really upset about factory farming, about global poverty. Specifically, I think in some ways I sort of emotionally wish I could work on those instead of AI, if I thought that was the optimal thing for me to do. This isn't to say that it should be the dominating concern, but I don't think it should be forgotten.

Host

是的,我认为这个论点经常以那种方式表述,只是比较风险与风险,部分原因是当你总是用同一种货币交易时,它非常干净。你不需要做任何……

Yeah, I think part of the reason the argument is often phrased that way, just comparing risk against risk, is that it's so clean when you're always dealing in the same currency. You don't have to make any...

安全研究效果乘数 Safety Research Effectiveness Multiplier

Host

从一件事转换到另一件事,可能取决于你放入什么参数,你可以说,这个论点实际上是自我消解的——就像在任何货币上都没有好处,所以非常非常方便。但解决所有这些一直困扰我们的其他问题也会非常好。而且,可能活得更久,因为我们在生物医学科学等方面会取得巨大进步。我想,你可能更倾向于这个论点:Anthropic 只是所有 AI 研究中的沧海一粟,而且,你知道,它实际上能在多大程度上推动比别处已经创建的大模型更大的模型?尤其是更危险的大模型,似乎你在安全导向的思维和安全导向的研究中所占的份额,可能比你在许多地方进行的大模型研究中所占的份额要大得多。

Conversions from one thing to another, plausibly depending on what parameters you stick in, you could just say, well, this argument is actually selfing—like there's no benefit on any currency, so it's very, very convenient. But it also would be very nice to solve all of these other problems that bevil us all the time. And I guess, but yeah, potentially live a lot longer because we'll make huge advances in biomedical science and so on. I guess I thought, yeah, you might lean even more heavily on the argument that Anthropic is kind of a drop in the ocean of all AI research, and that, you know, really how much is it going to be doing to actually spur large models than have been created elsewhere? And you know, especially like more dangerous larger models, it seems like you're potentially going to be a much larger fraction of the safety-focused thinking and safety-focused research than you would be of large model research, which is going on in many places.

Chris Olah

是的,我认为这里的论点实际上相当微妙和细致。首先,我们可以尝试从比例的角度来论证,你需要考虑三件事。第一,正如你所说,有点像,我不知道,也许 Anthropic 投入安全或能力的努力比例——尽管我认为我并不完全认同;我认为如果你试图回答的问题是“如何让大型 AI 系统安全”,这可能不是正确的区分。这些实际上是紧密交织的事情,不容易分开。但是的,我认为安全可能比大多数组织更是我们使命的一部分——这是我们存在的原因。但我认为还有另外两个因素可能更大。一个是,如果你相信最有效的安全研究只能在大模型上进行,那么你必须将投入的安全努力乘以一个有效性乘数。我认为这个有效性乘数就像,我不知道,假设你在比较 Anthropic 和一个只做 5% 安全研究的实验室,并且假设 Anthropic,尽管我认为这不是思考问题的正确方式,正在做 50% 或 70% 的安全研究。但然后你认为实际上有一个 20 比 1 的有效性乘数——你知道,那可能是压倒性的因素。另一件事是,如果你真正担心的是你是否在加速竞赛或导致领域竞赛,我认为相关的不是你花了多少努力在训练大模型上,而是你在领域中的行为方式中非常微妙的事情。

Yeah, I think the arguments here are actually pretty subtle and nuanced. So first, we could try to make an argument on ratios, and you'd have to account for three things. The first, as you say, is sort of like, I don't know, maybe the fraction of Anthropic's effort that goes into safety or capabilities—although I think I don't really buy that; I think that's maybe not the right distinction to make if the question you're trying to answer is how do you make large AI systems safe. Those actually are very intertwined things that aren't so easy to pull apart. But yeah, I think safety is probably a bigger part of our mission than most organizations—it's the reason we exist. But I think there are two other things that maybe are even larger factors. So one is, if you believe that the most effective safety research can only be done on large models, then you have to multiply the amount of safety effort that's going in by the effectiveness multiplier. And I think that effectiveness multiplier is like, I don't know, suppose you're comparing Anthropic with a lab that does 5% safety research, and let's say that Anthropic, even though I sort of think this isn't quite the right way to think about things, is doing 50% safety research or 70% safety research or something like that. But then you think that it's actually like a 20-to-1 effectiveness multiplier or something like this—you know, that might be the swamping thing. The other thing is, if the thing that you're really worried about is whether you're accelerating races or causing the field to race, I think it's not how much of your effort you're spending on training large models or something like that that's relevant; it's actually really nuanced things about how you conduct yourself in the field.

Host

是的,我认为你做什么样的研究非常重要。我认为如果你做的模型你认为刚好低于正在训练的最大模型,或者比它们更大,这很重要。我认为人们真正低估的一件事是,既然我们现在进入了 AI 实验可能花费数百万美元的阶段,我认为营销非常重要。我认为那些导致组织愿意投入大量资源的事情——一个花哨的演示可能比一个非常枯燥的技术结果引用要有效得多。所以我认为那里实际上有很多细节。所以无论如何,退一步,再宏观地看,我们有,我不知道,也许这种粗略的你的努力有多少投入到一件事或另一件事上,如果你完全专注于构建安全系统,或者如果你真的紧密整合这些事情,这可能不是正确的框架。但然后你还必须乘以安全方面的有效性乘数,以及你如何影响竞赛动态的谨慎乘数。我的猜测是,对我来说最大的一个是我对安全有效性乘数的信念。我认为我们已经有一些例子了,对吧?比如我做的多模态神经元工作,如果没有访问——这是我在 OpenAI 时做的工作——但如果没有访问最先进的模型并能够紧密整合这些努力,就不可能完成。或者《从人类偏好中学习》那篇论文再次依赖于访问大模型来演示。所以是的,我认为很多安全研究如果能够访问这些模型,可以获得很大的乘数。

Yeah, I think it matters a lot exactly what kind of research you do. I think it matters a lot if you're doing models that you think are just below the largest models being trained or that are larger than them. I think actually something that people really underrate is that I think now that we're getting into the regime where AI experiments cost potentially many millions of dollars, I think that marketing matters a lot. I think that the kinds of things that cause organizations to be willing to go and invest huge amounts of resources in something—a flashy demo may go a long ways further than a very dry reference to a technical result. So I think there's actually a lot of detail there. So in any case, stepping back, zooming out again, we have like, I don't know, perhaps this gross like how much of your effort is going to one thing or another, which may not quite be the right framing if you're entirely focused on building safe systems, or if you're sort of really tightly integrating those things. But then you also have to multiply by the effectiveness multiplier on safety and also by the carefulness multiplier on how you affect race dynamics. And my guess is that the one that for me is the largest one is my belief about the effectiveness multiplier on safety. And I think we already have some examples of that, right? Like the multimodal neurons work that I did wouldn't have been possible without access to—this is work that I did while at OpenAI—but it would have been possible to do without access to state-of-the-art models and being able to go and tightly integrate with that effort. Or the Learning from Human Preferences paper again relied on access to large models to be able to go and demonstrate that. And so yeah, I think a lot of safety research can get a large multiplier if you have access to these models.

Host

是的,是的,我们刚才在谈论 Anthropic,这已经大大偏离到了安全和大模型的问题上。在我们回到 Anthropic 之前,我只想问一个问题,这可能暴露我的天真,但似乎只是让模型更大,这不就意味着给东西增加更多的算力吗?这基本上是同一件事,但你在硬件上花更多钱,在电费上花更多钱。如果仅仅通过运行大模型没有展示出什么惊人的洞见,那么也许这样做并没有真正的帮助,或者并没有真正推动更大的模型。也许甚至像现在芯片短缺,对吧?制造这些芯片的能力有限,所以如果你只是买下芯片,那么我想,或者只有这么多算力可用。所以也许无论谁在做,对总量并没有太大影响。新的 AI 安全议程:去买下所有 GPU,创建并用它们进行加密货币挖矿,就这样,然后用加密货币挖矿买更多的 GPU。

Yeah, yeah, we were talking about Anthropic and this has been a big diversion into this question of safety and large models. Before we go back to Anthropic, I just want to ask one more question, which is just maybe this is going to show my naivety, but it seems like just making the models larger doesn't that just mean like adding a whole lot more compute to the thing? It's kind of the same thing but you are spending more money on the hardware and like spending more money on the electricity. Maybe if there's no amazing insight here that you're demonstrating by just running large models, then maybe it doesn't really help or it doesn't really promote larger models just to be doing it. Maybe even like there's a chip shortage at the moment, right? And there's only so many ability to make these chips, so if you just buy up the chips then I guess well, or there's only so much compute to go around. So maybe whoever's doing it doesn't really make that much difference to the total. New AI safety agenda: go and buy all GPUs, create and use them for crypto mining, there we go, and then use the crypto mining to buy more GPUs.

Chris Olah

我的同事 Danny Hernandez 有一项工作,他绘制了一张图,展示了用于训练最大模型的算力随时间的变化,他发现人们用于最大模型的算力大约每三个半月翻一番。我认为在某种程度上,边缘参与者可能只对此产生一点影响。这是一个由很多人共同作用形成的强劲趋势,但同时,我认为每次人们看到别人使用更多算力,这就会强化那个趋势,有时甚至会加速一点。

There was this work by my colleague Danny Hernandez where he just made a graph of the amount of compute used to train the largest models as a function of time, and he found that the amount of compute that people use on the largest models doubles every three and a half months or so. And I think to some extent, marginal actors probably only affect that a little bit. It's sort of this strong trend that's a function of lots of people, but at the same time, I think that every time people see people using more compute, that sort of reinforces that trend and can sometimes accelerate it a little bit more.

Host

好的,是的,这有道理。好了,让我们回到 Anthropic 作为一个组织。我想是的,我们一直在谈论大模型作为研究议程的一个独特方面。是的,还有其他值得向人们强调的显著方面吗?

Okay, yeah, that makes sense. All right, let's go back to Anthropic as an organization. I guess so yeah, we've been talking about large models as kind of one distinctive aspect of the research agenda. Yeah, is there any other notable aspects of it that are worth highlighting for people?

Chris Olah

是的,我认为 Anthropic 的一个非常独特之处在于我们对可解释性和机制可解释性的关注。我们有一个专门的团队致力于理解模型内部的工作原理。我们认为这对安全至关重要,因为如果你不理解模型在做什么,你就无法真正信任它。所以我们正在做很多关于理解特征、电路以及模型如何表示知识的工作。另一件事是我们在对齐方面的工作,比如 RLHF 和可扩展监督。我们正在开发技术,以确保模型的行为与人类价值观一致,即使它们变得更有能力。我们也在思考超级对齐——如何对齐超级智能系统。所以不仅仅是让模型更大,而是让它们安全且可理解。

Yeah, I mean, I think one of the things that's really distinctive about Anthropic is our focus on interpretability and mechanistic interpretability. We have a whole team dedicated to understanding how models work internally. And we think that's crucial for safety because if you don't understand what a model is doing, you can't really trust it. So we're doing a lot of work on understanding features, circuits, and how models represent knowledge. Another thing is our work on alignment, like RLHF and scalable oversight. We're trying to develop techniques to ensure that models behave in ways that are consistent with human values, even as they become more capable. And we're also thinking about superalignment—how to align superintelligent systems. So it's not just about making models bigger; it's about making them safe and understandable.

对安全与整合的异常关注 Unusual Focus on Safety and Integration

Chris Olah

我想专注于安全是非常不寻常的,而且我们在可解释性上押下重注也是一件相当不寻常的事。另一件稍微不寻常的事情是,我们非常关注这项工作如何影响社会,并试图紧密整合各个方面。实际上,真正努力将所有部分紧密整合在一起——将 Scaling(规模扩张)工作、可解释性工作、人类反馈工作和社会影响工作统一到一个整体中,并形成一个紧密整合的团队——是非常不寻常的。

I guess being focused on safety is pretty unusual, and I think us making a big bet on interpretability is also a pretty unusual thing to be doing. I guess another slightly unusual thing is just also being really focused on thinking about how this work affects society and trying to tightly integrate things. And I guess actually just really trying to tightly integrate all these pieces together—so trying to integrate the scaling work, the interpretability work, the human feedback work, and the societal impacts work, and sort of unite those all in a single package and in a tightly integrated team—is pretty unusual.

Chris Olah

我认为人们可能低估了一点:如果你想对大型模型进行任何类型的安全研究,问题不仅仅是能否访问那个大型模型,而是这些模型变得非常笨重且难以处理,因为它们实在太庞大了。你通常使用的许多基础设施根本不起作用。因此,无论你想进行哪种安全研究,实际上可能都需要做大量专门的工程和基础设施建设,才能在这些大型模型上进行安全研究。

I think something that people probably underappreciate is: if you want to do any kind of safety research on a large model, it's not just a question of having access to that large model, but these models become very unwieldy and very difficult to work with because they're just so enormous. A lot of the infrastructure that you normally work with just doesn't work. So for whatever kind of safety research you want to do, you actually probably have to do a lot of specialized engineering and infrastructure building to enable safety research on those large models.

Host

哦,我完全不知道,是啊。

Oh, I had no idea, yeah.

Chris Olah

而且你需要大量来自有大型模型工作经验的人的专业知识。所以我认为有很多事情只有在真正紧密整合时才能实现。另外,我认为人们可能没有意识到 Scaling(规模扩张)实际上是两件不同的事情:一是训练大型模型,二是关于缩放定律的预测工作。我们的一个重点是将缩放定律应用到每个领域。例如,对于可解释性,我们知道在某些点上,如果你观察非常具体任务上的趋势,会出现突然的转变:较小的模型无法做算术,而较大的模型突然就能做了。我们认为这可能是底层电路中的相变。那么我们能研究这个吗?或者对于人类反馈,我们能理解古德哈特定律的缩放是如何工作的吗?如果你训练不同大小的模型,并给予不同数量的人类反馈,会发生什么?我们能通过经验或类似缩放定律的方式理解模型是否能正确泛化到人类反馈或古德哈特定律本身吗?这些问题通过引入这种关于 Scaling(规模扩张)的思考,有望让我们能够思考它们,并可能让我们开始做的工作不仅仅是让当前模型安全,而是让未来的模型也安全。而所有这些都需要你将通过 Scaling(规模扩张)工作产生的专业知识与安全紧密整合。

And you need a lot of expertise from people who have experience working with large models. So I think there's a lot of things that can only happen when you really tightly integrate things. It's also, I think people maybe don't appreciate the extent to which scaling is really two separate things: there's training large models, but there's also this work on scaling laws of going and predicting them. And one of our big focuses is sort of taking scaling laws to every domain. So for interpretability, we know that there are these points where, if you look at the trend for very specific tasks, you get these abrupt transitions where a smaller model can't do arithmetic and then suddenly a larger model can. And we think that probably is a phase change in the underlying circuit. So can we study that? Or for human feedback, can we understand how the scaling of Goodhart's law works? What happens if you train models of different sizes and also have different amounts of human feedback? Can we understand empirically, or in terms of something like scaling laws, whether a model will correctly generalize to the human feedback or Goodhart itself? And so those are questions that, by bringing this kind of thinking about scaling, hopefully allows us to think about it and maybe allows us to start doing work that isn't just about making the present model safe but even about making future models safe. And so all of that requires you to be tightly integrating the expertise that you generate by working on scaling with safety.

Host

好的,是啊。为了给大家一个概念,Anthropic 现在有多少员工?你们有办公室吗?你们是否设在某个特定地点?

Okay, yeah. Just to paint a picture for people, how many staff are at Anthropic now, and do you have an office? Are you based in any particular location?

Chris Olah

是的,我们总部在旧金山。目前有 17 个人,我希望我们很快会有办公室。目前还没有。

Yeah, so we're based in San Francisco. We are 17 people right now, and I hope that we will soon have an office. We do not yet.

Host

好吧。是啊,我想现在是远程办公的好时机,没错。

All right. Yeah, well, I guess this is a good time to be remote, yes.

Chris Olah

是的,我认为直到今天,我在英国认识的大多数人都在远程工作,但我认为可能再过几个月我们就会回到办公室了。

Yeah, I think still to this day, most of the people I know in the UK are working remotely, but I think it's going to be any month now that we're going to be back in the office, probably.

Host

谁在资助整个运营?

Who's funding the whole operation?

Chris Olah

是的,Anthropic 的 A 轮融资由 Yon Talon 领投,我想他非常关心 AI 安全,并且有一些个人投资者加入,而不是风投。

Yeah, Anthropic's Series A was led by Yon Talon, who I guess cares a lot about AI safety, and was joined by a number of individual funders rather than VCs.

Host

是啊,考虑到你们在开发这些大型模型,我想你们的主要开支之一就是购买计算硬件之类的东西吧?

Yeah, I guess given that you're working on these large models, I'm imagining that one of your primary expenses is just buying computing hardware and all of that kind of thing?

Chris Olah

是的,没错。我认为实际上,在很多方面,AI 可能开始看起来更像其他行业,比如生物技术,研究需要非常高的资本成本,你必须花很多钱在试剂上。类似地,我们必须花很多钱在实验的计算上。我认为在某些方面,这是一个很自然的类比。

Yeah, that's right. I think that actually, in a lot of ways, AI is maybe starting to look more like other industries where you have really intense capital costs for research, like biotech, and you have to spend lots of money on reagents. And similarly, we have to spend lots of money on compute for our experiments. I think that's in some ways an industry that it's natural to compare us to.

Host

是啊,不太像 Etsy,更像制造汽车之类的。

Yeah, less like Etsy, more like manufacturing cars or something.

Host

那么业务结构是怎样的?这是营利性的还是更像非营利模式?

And what's the business structure? Is this kind of for-profit or more of a nonprofit model?

Chris Olah

是的,Anthropic 是营利性的。严格来说,我们是一家公益公司,但这只是意味着我们没有法律义务去最大化股东价值。我认为,不管怎样,这可能对所有公司都适用,或者至少当这个问题被诉诸法庭时,公司有很多……是的。我认为在实践中很难做到这一点,但公益公司确保了……是的,确保了。

Yeah, Anthropic is for-profit. Technically, we're a public benefit corporation, but that just means that we are not legally obligated to maximize shareholder value. I think, for what it's worth, that might be true of all corporations, or at least when this has been taken to court, corporations have a lot of... yeah. I think in practice it's very difficult to go and do this, but a public benefit corporation makes sure... yeah, makes for sure.

Host

我对公益公司不太熟悉,但你们在章程中是否还有其他内容,比如除了盈利之外,还有组织的特定目标?

And I'm not super familiar with public benefit corporations, but do you end up with something else in your charter, like in addition to making profit, that's kind of a specified goal of the organization?

Chris Olah

是的,所以你在章程中设定一个使命,你可以优先考虑这个使命而不是股东价值。好的,我想在这里,使命可能是为了每个人的利益开发 AI 并分享成果,大致如此。

Yeah, so you go and you put a mission in your charter that you're allowed to go and prioritize over shareholder value. Okay, and I guess here it'll be something like developing AI for the benefit of everyone and sharing the spoils, something to that extent.

Chris Olah

好的,所以作为公益公司,你可以从希望获得回报的投资者那里获得投资。你也可以销售产品并用收入来资助增长。只是你除了最终赚钱之外还有其他目标。我想董事会等可以对公司有一个愿景,而不仅仅是最大化股息或股东价值。

Okay, so as a public benefit corporation, you can take investment from investors who are aiming to make some money back. You can also sell products and use that to fund the growth. It's just that you also have other goals in addition to making money at the end of the day. And I guess the board and so on can have a vision for the company that's not just maximizing dividends or maximizing shareholder value.

Host

是啊。那么你们如何将所做的研究变现呢?你们实际上如何获得收入以保持增长?

Yeah. So how does one kind of monetize the kind of research you're doing? How do you actually get the revenue in order to keep growing?

Chris Olah

是的,嗯,目前和不久的将来,我们只专注于研究。但我认为在 AI 的发展轨迹中有一个有趣的点:过去,AI 的经济价值受限于模型是否具备能力以及能否做任何有用的事情。现在我们似乎正在稍微走出那个阶段。实际上,我们现在可以构建——而且我认为很多组织都可以构建——在某种意义上能够做有用事情的模型,但它们不可靠、不值得信赖,没有人理解它们是如何工作的,这正成为它们使用和经济价值的瓶颈。所以我认为这让我们充满希望:如果你能非常擅长构建安全、可靠和值得信赖的系统,那么机会就会出现。

Yeah, well, right now and for the near future, we're just focused on research. But I think there's an interesting point in the trajectory of AI where, in the past, the economic value of AI has been bottlenecked on models being capable at all and being able to do anything useful at all. And it seems like we're maybe transitioning out of that phase a little bit. Actually now we can build—and I think a lot of organizations can build—models that are in some sense capable of doing useful things, but they aren't reliable, they aren't trustworthy, nobody understands how they work, and that's becoming the bottleneck on their use and on economic value from them. And so I think that makes us hopeful that there will be opportunities if you can become really good at building systems that are safe, reliable, and trustworthy.

安全的商业模式 Business Models for Safety

Host

这会有经济价值。是的,所以我想我能看到多种不同的商业模式。一种是亲自设计这些高风险系统,因为你拥有认证它们安全、按预期和预测运行的专长。另一种可能是做咨询,利用这种专长帮助其他公司修复他们的模型。还有一种可能是销售工具,比如可解释性工具等,让其他人也能查看自己的模型并帮助使其安全。我想可能所有这些在未来都有可能。

There will be economic value from that. Yeah, so I guess I can see multiple different business models here. One would be like actually designing these systems that are very high stakes yourself because you have the expertise in certifying that they are safe and going to act as desired and predicted. Another one might be doing consulting, using that expertise to help other companies fix their models. And as another, might be selling tools, interpretability tools and so on, that will allow everyone else to look inside their models and help to make them safe themselves. I guess probably all of these are possibly on the table for the future.

Chris Olah

目前,是的,我们处于非常早期的阶段。我认为最可能的事情是自己训练大模型并确保它们安全。我们认为将安全与模型设计整合起来可能很重要。我还想说,我们计划分享我们在安全方面的工作,与世界分享,因为我们最终只是想帮助人们构建安全的模型,不想囤积安全知识之类的东西,全部留给自己。

At the moment, yeah, I mean we're super early stage. I think the most likely thing would be training large models ourselves and making them safe. I think that we think it's probably important to integrate safety and the design of models. I should also say that we plan to share the work that we do on safety and share that with the world, because we ultimately just want to help people build safe models and don't want to hoard safety knowledge or something like this, keep it all for yourself.

招聘优先级 Hiring Priorities

Host

好的,所以我想象 Anthropic 是一个相当新的组织,正在寻找各种不同角色的人。目前有没有什么特别值得强调的职位在招聘?

Alright, so I guess I'm imagining Anthropic is a pretty new organization, on the hunt for people to fill all kinds of different roles. Are there any particular roles that you're hiring for at the moment that are worth highlighting?

Chris Olah

是的,我认为有三个角色对我们来说优先级特别高,而且可能不是人们预期的角色。我认为人们常常假设,在机器学习研究机构中,影响力最大、最重要、最受欢迎的角色是机器学习研究角色。但实际上,我们真正需要的是两个非常重要的工程角色,以及一个我认为绝对必要的安全角色,我们需要找到合适的人选。

Yeah, well I think there's three roles that are in particular really high priority for us, and they're maybe not the roles that people might expect. I think there's often an assumption that the highest impact and most important and in-demand roles at a machine learning research organization are going to be machine learning research roles. But in fact, the things that we really need are two engineering roles that are really important, and a security role that I think is absolutely essential that we find someone good for.

Host

好的,我们一个一个来。工程角色是什么?为什么它重要到值得特别强调?

Okay, yeah, let's take those one by one. What's the engineering role and I guess why is it important enough to really want to highlight it?

Chris Olah

实际上,Anthropic 进行研究和探索这些大模型安全性的核心能力,在于处理大模型的能力。而这一切的核心是我们的集群,基本上是一台超级计算机,我们用它来完成所有工作。因此,有两个与此相关的角色对我们开展有效研究至关重要。第一个是我们称之为基础设施工程师的角色。这个人将负责保持超级计算机的运行,并拥有与之交互所需的工具。我认为可以从两个方面来理解为什么这对我们来说是一个高影响力的角色。一种思考方式是,目前所有这些工作都是由研究人员完成的。我之前听过一个旧播客,你在那里讨论运维,我想是和 Tanya Singh 一起,她评论说通过做运维工作,她每花一小时运维时间就能释放出超过一小时的研究人员时间。我认为这里的情况绝对类似,担任基础设施工程师的人每花一小时就能释放出超过一小时的研究人员时间。原因在于他们拥有专业技能,知道如何保持系统运行,而不是像其他领域的专家那样只是应付了事。

So really at the core of Anthropic's ability to go and do our research, to explore safety in these large models, is the ability to work with large models. And at the core of that is our cluster, basically a supercomputer that we're using to do all of our work on. So there's really two roles related to that that are central to our ability to go and do productive research. The first one is what we're calling our infrastructure engineering role. This is really someone who's going to be responsible for keeping that supercomputer running and having the tooling we need to interact with it. I think there's two ways you could sort of think about why it's such a high impact role for us. One way to think about it is that right now all of that work is being done by researchers. I was listening to one of the old podcasts a while back where you were discussing operations, I guess it was with Tanya Singh, and she commented on how by doing operations work she was able to free up more than one hour of researcher time per hour of operations time that she did. I think that's absolutely something similar would be true here, that somebody who took on this infrastructure engineering role would be freeing up way more than one hour of researcher time for every hour they spent on it. And the reason is just that they'll actually have the specialist skills to know how to keep it running, instead of being experts in a different topic who are playing at it.

Host

他们真的知道自己在做什么,而不是其他领域的专家在瞎搞,不知道自己在做什么。是的,没错。

They will really know what they're doing, instead of being experts in a different topic who are playing at it, not knowing what they're doing. Yeah, yeah, so that's exactly right.

Chris Olah

还有另一种思考方式:我们每年在这个集群上花费数千万美元,以便能够对大模型进行安全研究。有时集群无法使用,或者出现不必要的故障。我们认为,担任这个角色的人很有可能将正常运行时间和可靠性提高 10%以上。从某种意义上说,这相当于向一个专注于大模型安全的组织捐赠数百万美元。所以,如果你对让大模型安全感到兴奋,那可能是一件非常有影响力的事情。

There's another way to think about it, which is we are spending tens of millions of dollars a year on this cluster to allow us to go and do our safety research on large models. And sometimes the cluster is unusable or things break unnecessarily. We think that it's quite likely that somebody who took on this role could increase the uptime and reliability by more than 10%. And so in some sense, that would be equivalent to giving millions of dollars to an organization that's focused on the safety of large models. So if you're excited about making large models safe, that could be a really high impact thing to do.

Host

是的,我不太懂如何联网计算机或制造超级计算机,但我猜这有点技术性。那些可能拥有操作这种规模计算机技能的听众,他们会知道自己是这类人吗?或者我们能说些什么来确保他们意识到应该报名?

Yeah, so I don't really know how to network computers or make a supercomputer, but I'm guessing it's a little bit technical. Are the people in the audience who might actually have the skills to operate a computer of this scale going to know who they are? Or is there anything we can say to make sure that they are aware that they ought to put their name forward?

Chris Olah

我们所做的一切都基于 Kubernetes,这是一个用于处理大型分布式系统的框架。我认为拥有网络方面的专长,如果有人有 GPU 等方面的经验,那都会很有用。但我认为 Kubernetes 的专业知识,以及在某种程度上做 SRE(站点可靠性工程)工作的经验,会与这个基础设施工程师角色相关。

So everything that we're doing is based on Kubernetes, which is a framework for working with large distributed systems. I think that having expertise in networking, and if somebody had experience with GPUs or things like that, that could all be useful. But I think expertise with Kubernetes and to some extent doing sort of SRE work would be things that would be relevant to this infrastructure engineering role.

Host

好的,不错。那是第一个工程角色。第二个是什么?

Okay, nice. So that was the first engineering role. What was the second?

Chris Olah

第二个工程角色,我们称之为系统研究员角色,也与这个集群密切相关。因为有一个完整的挑战:如果你运行这些巨型模型,如何有效地将它们适配并设置在超级计算机上?如何拆分它们?如何高效地让所有网络工作?如何让一切尽可能高效地运行?这方面没有太多先例,因为处理这些真正巨型模型的地方并不多。所以这确实是一个新颖的工程问题:如何在分布式系统上高效运行这些模型。

So the second engineering role, and we're calling it our systems researcher role, is also really centrally related to this cluster. Because there's this whole challenge of figuring out how do you, if you're running these giant models, how do you effectively fit them and set them up on the supercomputer? How do you break them apart? How do you efficiently make all the networking work? How do you get everything to run as efficiently as possible? And there's not a lot of prior art on this because these really giant models, there's not that many places that have worked with them. So it's really this novel engineering problem of how you efficiently run these models on distributed systems.

Host

所以这大概是这样的:你运行特定类型的过程、特定类型的算法,你想弄清楚当你在大量不同进程上运行时,如何最有效地分布和排序这些操作,每个进程可能都有不同的专长或能最高效执行的特定操作。我不太懂计算机的工作原理,但听起来差不多是这个意思?

So this is something along the lines of, I guess, you're running particular kinds of processes, particular kinds of algorithms, and you want to figure out how you can distribute and order those operations most efficiently when you're running them across tons of different processes, which each might have different specialties or particular kinds of operations they can do most efficiently. I don't really understand how computers work, but seems like that's in the ballpark?

Chris Olah

是的,没错。你有一堆计算机,然后每个……

Yeah, that's right. So you have a bunch of computers and then each of the...

Anthropic 的工程岗位 Engineering roles at Anthropic

Host

计算机有一堆 GPU,关于如何布局这些大模型有很多问题。你要做大量矩阵乘法——如何把这些计算分布到这些 GPU 上?它们之间应该如何通信?最高效的方式是什么?既有关于如何组织的高层问题,也需要考虑内存带宽、在不同内存之间加载数据需要多长时间、网络等等。所以这里既有高层考量,也有非常底层的效率优化。做得好带来的回报,我想就是你能用同样的硬件和电费,完成更多实际有用的计算工作。

Computers have a bunch of GPUs, and there are a lot of questions about how to lay out these large models. You're multiplying all these large matrices—how do you lay that out across these GPUs? How should they all talk to each other? What's the most efficient way to do things? There are both high-level questions about the most efficient way to organize things, and you have to think about memory bandwidth, how long it'll take to load things between different kinds of memory, the network, and stuff like that. So there's this interplay of high-level considerations and also very low-level considerations of just how to make things very efficient. The reward from doing that well, I guess, is you get more actually useful computational work out of the same amount of hardware and the same electricity bill.

Chris Olah

对,完全正确。所以我认为,很容易想象做这个角色的人能让我们的效率至少提升 10%。这相当于给一个专注于安全的组织提供了数百万美元。

Yeah, that's exactly right. So again, I think it's very easy to imagine that someone doing this role could increase our efficiency by at least 10%. And so again, that's equivalent to providing millions of dollars to an organization that's focused on safety.

Host

太棒了。那么,如果这个角色不是传统的技术岗位,或者很多人以前可能做过类似的工作?人们如何判断自己是否具备能成长为这类岗位的原始技能?

Amazing. So if this is a role that isn't a traditional tech role, or is it the kind of thing that lots of people might have worked on before? How could someone tell if they have the proto-skills that might allow them to grow into a position like this?

Chris Olah

嗯,我认为能胜任这个角色的一些特质包括:如果某人在底层硬件效率方面有丰富经验,那会是一个很好的指标。另外,在分布式系统方面有大量经验,特别是考虑分布式系统的效率,也是一个好指标。他们需要在工作中学习一些机器学习知识,但我觉得这个角色主要需要强大的工程能力,尤其是与效率和分布式系统相关的工程能力,这些是值得关注的点。

Yeah, so I think some of the things that would make someone effective at this role are: if somebody has a lot of experience thinking about efficiency in low-level hardware, that would be one thing that might be a really good indicator that they could be good at this. I think also having a lot of experience thinking about distributed systems, and especially thinking about efficiency in distributed systems, could be another good indicator. I think they'd have to learn some stuff about machine learning on the job, but I think for this role, it's primarily having strong engineering skills, and especially strong engineering skills related to efficiency and distributed systems, would be the things to look for.

Host

有意思。我觉得你似乎需要极力证明这些岗位为什么影响巨大,比如用等价捐款金额来衡量。但如果 Anthropic 做的是大事,那么让优秀的人来运行这个庞大的超级计算机系统显然也很重要。我的意思是,如果运行得不好,你会浪费大量硬件,还可能有很多停机时间拖慢进度。所以至少对我来说,这并不难理解。

Cool. It's slightly funny to me that you feel like you really have to justify why these roles are high impact in terms of the amount of equivalent money donated. It seems like if what Anthropic is doing is the big picture, having good people running this enormous supercomputer system seems like it's obviously important as well. I mean, you could imagine if it's running incompetently, you're going to waste a whole lot of hardware and also probably have tons of downtime that's going to slow things down. So it's not a very hard sell to me at least.

Chris Olah

嗯,我经常和人交流,感觉他们把工程岗位视为次要的。他们想方设法从工程师转型为机器学习研究员,或者认为工程岗位不那么重要。所以我说这些其实就是为了强调这些工程岗位的巨大影响力。我认为我认识的最有成效的人之一是 Tom Brown,他本质上是一名工程师,却是 GPT-3 的第一作者,在推动大模型研究方面影响巨大。

Well, I often talk to people and I get the impression that they see engineering roles as sort of secondary importance. They're trying to figure out how they can transition from being an engineer to being a machine learning researcher, or sort of assuming that engineering roles are less important. So I guess really the reason I'm saying all of this is just to emphasize how tremendously high impact these engineering roles are. I think that one of the most effective people that I know is Tom Brown, who is really primarily an engineer, and he was the lead author of GPT-3 and has been tremendously impactful in enabling research on large models.

Host

是啊,我觉得显然需要有人来做这项工作。所以我猜人们心里的逻辑是:对于这些职位,你不需要一个对 AI 安全特别有热情的人。难道不能直接雇一个不一定那么关心使命的人,付钱让他们运行超级计算机吗?我想我能理解人们为什么这么想。但当你和从事这类工作的组织交流时,他们会说,不,实际上很难找到最优秀的工程师,如果能找到一个更胜任的人,那会带来很大的不同。

Yeah, I guess it seems just blatantly obvious that you need someone to do this work. So I guess the logic that's going on in people's heads is: for these positions, you don't need someone who's especially passionate about AI safety as a problem in the world. Can't you just hire someone who doesn't even necessarily have to care all that much about the mission and just pay the money to run the supercomputer? I guess I can see how people get that in their heads. It just seems like when you talk to organizations that are trying to do work like this, they say no, it's actually really hard to get the best engineers, and it would really make a difference if we could find someone who was better at the job.

Chris Olah

对,完全正确。

Yeah, that's absolutely right.

Host

你有没有一个理论来解释,为什么很多时候你需要那些关心使命重要性、关心 AI 安全的人,尽管他们的工作似乎不直接是 AI 研究科学?更像是构建所有工具来让人们做研究。

Do you have a theory for why it is that quite often you need people who care about the importance of the mission and care about AI safety when their job doesn't seem to be the AI research science directly? It's like building all the tools that enable people to do that research.

Chris Olah

首先,我认为真正擅长这些技能的人很稀缺,很难找到。即使不考虑对齐问题,这些岗位本身就很难招人,就像招到好的机器学习研究员一样难。如果某人加入你的部分原因是他们关心你的使命,那么你更有可能得到真正出色的人。其次,我认为让整个组织都关心使命是非常健康的,而不是出现某种奇怪的分裂,一部分人关心,一部分人不关心。最后一点,至少在 Anthropic,一切都是紧密集成的。团队之间不是孤立的。大家紧密合作,界限模糊,协作很多。所以如果有一群关心安全的人在做研究性工作,而另一群人不关心,那文化会非常奇怪。

Well, the first thing is, I think just people who are really good at these skills are rare and hard to find. Even if you weren't considering alignment, these are just hard roles to hire for, just like getting good machine learning researchers is a hard role to hire for. And you're more likely to get somebody who is really extraordinarily good at it if part of the reason they're joining you is that they care about your mission. But I think a second reason is just that it's really healthy for an organization to have the entire organization care about the mission, rather than having some kind of weird bifurcation where some portion of people care and some portion don't. Maybe one final thing is, at least at Anthropic, everything is very tightly integrated. It's not like these teams are siloed. There's people working on one thing and they're siloed from people working on something else. We're all working very closely together, and there are very blurry lines and lots of collaboration. So I think it would just be a very strange culture if you had some set of people who cared about safety doing more researchy things, and I think it would be a very strange situation.

Host

是啊,正如你所说,随着机器学习如此蓬勃发展,能够胜任这类工作的人可能非常抢手。猎头很多,所以很难招到最好的人。如果某人不关心 Anthropic 的使命,那么长期留住他们可能更难。理想情况下,你肯定不希望每年都换掉那个搭建并真正了解你计算栈的人。我能想象那会带来很多麻烦,肯定不好玩。

Yeah, I guess as you're saying, with ML taking off as much as it is, people who are able to do jobs like this are probably pretty sought after. There's a lot of headhunting, so it's hard to get the best people. And if someone doesn't care about the mission that Anthropic is engaged in, then it might be a lot harder to retain someone long term. Ideally, you probably don't want to have annual turnover on the person who has put together and really knows how your compute stack works. I can see that leading to a lot of headaches. I think that wouldn't be terribly fun.

Chris Olah

好的。那么这就是两个工程岗位。还有第三个,对吧?是安全。

Cool. Okay, so that was the two engineering roles. There was a third one, right? And that was security.

Host

嗯,之前你说过,也许原因是不言自明的……

Yeah, well, earlier you were saying that maybe it was kind of self-evident why...

前沿 AI 实验室安全的重要性 Importance of Security at Frontier AI Labs

Host

工程岗位如此重要,如果真是这样,那么安全岗位的重要性就更不言而喻了。如果你要构建这些强大的系统,确保安全至关重要。你不想让任何人轻易拿走你所有的研究成果,尤其是当你的很多工作除了安全之外,还涉及能力相关的部分。

Engineering roles were so important, and if that's true, then maybe it's even more self-evident why the security role is really important. If you're going to be building these powerful systems, it's really important that you be secure. You don't want anyone to just be able to go and grab all of the research that you've done, especially if a lot of your work, in addition to safety, has capabilities-relevant components.

Chris Olah

是的,整个模式就是你可能在前沿工作,处理那些你担心尚未完全成熟和安全的东西。所以我们不想让随便什么黑客都能来拿走我们的模型。那看起来很糟糕,非常糟糕。

Yeah, I mean, the whole model is that you're potentially going to be working at the frontier with stuff that you're concerned isn't quite fully baked and safe yet. So we just don't want random hackers to be able to come and grab our models. That seems bad, really bad.

Host

不仅因为这些模型可能不安全,或者未来我们会构建可能更有害的模型,还因为它们可能被滥用。我认为这些模型有很大潜力被恶意行为者滥用,所以会有越来越多的恶意行为者试图获取这些模型。

And not just because these models are potentially unsafe, or in the future we're going to build models that could potentially be more harmful, but also because they're abusable. I think there's a lot of potential for these models to be misused by bad actors, and so there are going to be bad actors who are increasingly trying to get access to these models.

Chris Olah

是的,如今似乎每个处理重要机密信息的技术公司或组织,都需要有人来加强他们的信息安全和计算机安全。但这并不简单。人们今天面临的威胁相当严重,需要很多专业知识才能弄清楚如何将信息留在内部。

Yeah, it seems like every tech company or every organization that's dealing with important confidential information these days really needs someone to lock down their information security and computer security. But it's not straightforward. The kinds of threats that people face today are pretty serious, and it takes a lot of know-how to figure out how to keep information in-house.

Host

这一点也不简单,而且我认为这很容易被忽视。人们很容易说:‘Anthropic 现在是个小组织,可能没人会针对我们,我们可以以后再说。’我认为那将是一个可怕的错误。我认为这是我们应该在早期阶段就着手处理的事情。因为一旦你构建了所有系统,事后回过头来想办法让它们安全,一旦你做出了所有设计选择,决定用这个软件而不是那个,就很难再回去修补你多年前或几十年前犯下的错误。

It's not straightforward at all, and I think it's a really easy thing to neglect. It's a really easy thing to say, 'Anthropic is a small organization right now, probably nobody's going to try to do anything to us, we could leave this to later.' And I think that would be a terrible mistake. I think this is something that we really want to be working on at an early stage and trying to address right now. Because once you've built all of your systems, going back and figuring out how to make them safe after the fact, once you've made all these design choices and chosen to use one piece of software over another, it's just really hard to go back and patch the mistakes that you made years or decades ago.

Chris Olah

我认为没错。我认为从一开始就建立一种认真对待安全的文化也很重要。我们正试图通过让一些员工,尤其是我的优秀同事 Ben,来负责安全工作。但我们不是专家,我们真的需要一位专家加入我们,能够处理这件事并领导它。

I think that's right. I think there's also probably something important about building a culture of taking security seriously from the start. We're trying to do that by having some employees, especially my wonderful colleague Ben, work on security. But we're not experts, and we really need an expert to join us and be able to handle this and take the lead on it.

Anthropic 独特的安全挑战 Unique Security Challenges at Anthropic

Host

那么,这个安全岗位会是人们熟悉的那种吗,类似于其他科技公司的工作?还是说这有点特殊,人们可能要做一些典型安全人员不太熟悉的工作?

So is this going to be the kind of security role that people might be familiar with, similar to the sorts of work that people might do in other tech companies? Or is this slightly a weird case where people might be doing different work that a typical security person might be less familiar with?

Chris Olah

我认为有相似的部分,但我们也面临一些非常不寻常的挑战。例如,除了所有关于防范外部威胁的安全问题,我们还在训练语言模型生成代码——这是我们的研究方向之一——我们预计这些模型在某个时候会开始做坏事,我们正在尝试对它们进行沙盒隔离。所以,思考如何有效地对它们进行沙盒隔离对我们来说非常重要。据我所知,这是一个非常不寻常的安全问题。也许那些处理恶意软件并需要在沙盒中运行它的人会遇到类似的情况,但这是一个出现在机器学习中的不寻常问题,我们需要处理它。

I think there are components that are similar, but I think we also have some challenges that are really unusual. For example, in addition to all of these concerns about security against external threats, we're training language models to generate code—that's one of our lines of research—and we expect those models at some point to start doing bad things, and we're trying to sandbox them. So it's pretty important for us to be thinking about how to effectively sandbox them. That's a very unusual flavor of a security problem, as far as I can tell. Maybe people who work with malware and need to run it in sandboxes have something a little bit similar, but it's an unusual problem that comes up in machine learning and that we need to deal with.

Host

所以这就是问题所在:人们最近开始使用语言模型进行编程,并发现这些语言模型在编写实际可用的程序方面非常出色。从长远来看,你可能会担心你将要处理的语言模型被设计来做一些恶作剧,因为那是你想弄清楚的事情。然后你不希望这种恶作剧发生在你身上,所以你必须想办法遏制系统,使其无法开始破坏它正在运行的计算机。大致是这样的吗?

So this is the issue that people have recently started using language models to do programming and have been finding that these language models are remarkably capable at programming things that actually work. And as in the long term, you might worry that you're potentially going to be playing with language models that are designed to do mischievous things because that's the kind of thing that you want to figure out. And then you don't want that mischief done against you, so you have to figure out some way to contain the system so that it can't start damaging the computer that it's running on. Is that basically the picture?

Chris Olah

是的,或者你可能根本没想做什么恶作剧,但模型可能还是会做一些你不希望它做的事。所以也许在某个时候,你正在使用强化学习来让这些模型完成任务,学习解决问题,并探索这类模型的安全性。现在你面临的情况是,也许模型会尝试——它实际上非常擅长编程——也许它会试图找到某个安全漏洞,进入系统给自己获取额外奖励。或者在更极端的版本中,它开始生成——这些都是非常推测性的东西——但你不希望最终陷入那种情况,即这种事情成为严重的可能性,而你却没有提前考虑过。所以我认为,让一个认真思考过安全的人来考虑这类事情会非常有价值。

Yeah, or you might not be trying to do anything mischievous at all, but the model might do things anyway that aren't what you want. So perhaps at some point you're using reinforcement learning to try to get these models to solve tasks and learn to solve problems, and exploring safety in such models. Now you're in a situation where maybe the model tries to—it's actually quite good at programming—maybe it tries to find some security vulnerability to go in and get itself extra reward. Or in more extreme versions, it starts to generate—these are very speculative things—but you don't want to end up in the kind of situation where that sort of thing is a serious possibility and you haven't thought about it in advance. So I think having somebody who's thought about security seriously thinking about this kind of stuff would be really valuable.

安全角色的理想候选人 Ideal Candidate for Security Role

Host

好的,所以安全岗位有一些熟悉的元素,也可能有一些更新颖的元素。人们如何知道自己是否适合申请这个岗位呢?

Okay, so the security role has some familiar elements and potentially some more novel elements. How could someone know if they're potentially a good fit to apply for that one?

Chris Olah

我认为这个岗位的挑战在于,我们现在需要有人处理多种安全问题。当我们与安全候选人交谈时,他们往往专注于某个特定专业领域,缺乏帮助我们应对所面临的各种问题的广泛专业知识。或者,他们确实有广泛知识,但此时他们已经是管理者了,没有太多做基层工作的经验;他们希望领导整个安全团队,并将各种职责委派出去。我们可能还没有大到可以组建一个大型安全团队。所以,我认为一个好的迹象是,如果有人拥有相当广泛的安全知识,并且乐于去帮助弄清楚如何让 Anthropic 变得安全。

I think the challenging thing about this role is that right now we need someone to be dealing with a lot of kinds of security problems. Often when we talk to security candidates, they're focused on a particular specialty and don't have the breadth of expertise to help us with the range of problems we're facing. Or in the alternative, they do, but they're really a manager at this point and don't have as much experience doing ground-level work; they're looking to lead a whole security team and delegate out the various responsibilities. We're probably not big enough for that to build out a large security team. So I think the thing that would be a really good sign is if somebody had a pretty broad range of security knowledge and was excited to go and help figure out how to make Anthropic secure.

Host

好的,所以这是一个相当通才型的安全岗位。我刚才看了一下 Anthropic 目前所有的空缺职位,还有几个其他的。其中一个可能有点令人惊讶的是数据……

Okay, so it's a pretty generalist security role. I was just looking over all of the open positions that you've got at the moment at Anthropic, and there's a couple of others here. One that's a little bit surprising maybe is the data...

可解释性研究中的数据可视化角色 Data Visualization Role in Interpretability Research

Host

可视化专家,是的。你希望那个人做什么?

Visualization specialist, yeah. What are you hoping for that person to do?

Chris Olah

是的,我认为这是一个非常令人兴奋的职位,如果你想从事可解释性研究的话。你知道,有一个非常有趣的现象:纵观科学史,新的研究方向和学科往往是因为有了合适的工具才得以开启。例如,早期化学似乎与玻璃器皿的发展有关,玻璃器皿使实验成为可能。所以我认为对于可解释性研究来说,情况也类似——这些模型非常庞大,仅仅是在其中导航、提问和查看数据就很困难,因为它们太大了。因此,我认为这类研究的一个强大推动力是拥有擅长数据可视化的人,他们能够成为研究内部循环的一部分,并弄清楚如何可视化、探索和理解我们在研究这些模型时获得的所有数据。所以我认为,如果有人有数据可视化经验,有实现交互式数据可视化的网页开发经验,并且有一些数学背景,这可能是他们支持可解释性研究的一种非常有影响力的方式。

Yeah, I think this is a really exciting role if you want to work on interpretability research. You know, there's this really interesting phenomenon where if you look at the history of science, often new lines of research and disciplines get unlocked by having the right kind of tooling. For example, early chemistry seems to have been linked to the development of glassware that made experiments possible. And so I think it's something kind of similar for interpretability research, where these models are enormous, and just being able to go and navigate through them and ask questions and look at data is difficult because they're so big. So I think that a really powerful enabler of this kind of research is having people who are good at data visualization and can sort of be part of the inner loop of that research and figure out how to go and visualize and explore and understand all the data that we're getting access to when we study these models. So I think if someone has experience in data visualization and experience with the kind of web development that makes interactive data visualization possible, and has some math background, this could be a really impactful way for them to support interpretability research.

机器学习通才及其他角色 ML Generalist and Other Roles

Host

好的,是的。我们已经讨论了四个职位。在我们继续之前,你还有什么特别想强调的其他职位吗?

All right, yeah. We've talked about four roles. Are there any others that you want to highlight in particular before we push on?

Chris Olah

嗯,如果你看我们的招聘页面,你会看到机器学习研究员和机器学习工程师这两个职位,这通常是机器学习领域职位的描述方式,即分为这两种角色。但我认为我们实际上在寻找的是我们内部称为“机器学习通才”的人——他们不一定非要从事研究或工程,而是只想做影响力最大的事情来推动研究进展。所以,如果听众对此有共鸣,那可能是一个非常有意义的角色。除此之外,我们还有其他一些职位。我认为今年剩余时间我们会比较缓慢地招聘这些职位,但明年可能会再次加大力度。这些职位包括运营岗位、公共政策岗位以及各种研究岗位。

Well, if you look at our job page, you'll see both roles for ML researchers and ML engineers, and that's sort of how jobs in machine learning are often described, as there being these two kinds of roles. But I think the thing that we're actually looking for underneath that is what we internally call an ML generalist—somebody who sort of isn't attached necessarily to doing research or doing engineering, but just wants to do the highest impact thing to move the research forward. So yeah, if that resonates with people listening, that could be a really meaningful role. Beyond that, we do have a number of other roles. I think that we're going to be hiring sort of more slowly for those roles for the rest of the year, but we'll probably be looking at them more intensely again next year. And that includes operations roles, roles working on public policy, and a variety of research roles.

Host

是的。有没有什么特别的事情是人们应该知道的——我想这里的职位太多了,很难笼统地说谁适合这些职位——但也许关于 Anthropic 的文化,人们在考虑申请工作之前应该知道些什么?

Yeah. Is there anything in particular that people should know maybe about—I suppose there's so many different roles here it's hard to say anything in general about who's appropriate for working in these positions—but maybe is there anything about the Anthropic culture so far that maybe people should know before they're considering applying for a job?

Chris Olah

是的,在更技术性的方面,我认为我非常喜欢 Anthropic 文化的一点是我们大量进行结对编程。研究人员和工程师经常在我们正在做的不同事情上结对。人们就像在 Slack 上发帖说:‘我要做这个,你想和我结对吗?’我认为这比我参与过的任何其他组织在结对编程方向上都要更进一步,而且这非常令人愉快。我真的很喜欢。我想另一件有点不寻常的事情是,我们非常专注于拥有一个统一的研究议程,而不是让很多人各自为政。但我们非常努力地作为一个团队来制定这个研究议程,所以我们有很多对话讨论我们的研究议程应该是什么,我们的重点应该是什么,并作为一个团队进行讨论。我认为这也非常酷。

Yeah, well, on the more technical side, I think one thing I really love about Anthropic's culture is that we do a ton of pair programming. So researchers and engineers are just constantly pairing on different things that we're working on. People are just like, post in Slack, 'I'm going to be working on this, do you want to pair with me?' And I think it's a very significant step further in the pair programming direction than any other organization I've been part of, and it's just delightful. I really love it. I think maybe another thing that's a little unusual is we're very focused on having sort of a unified research agenda rather than just having lots of people doing their own thing. But we try really hard to set that research agenda as a group, and so we have lots of conversations discussing what our research agenda should be and what our focus should be, and discussing that as a group. And I think that's been really cool as well.

Host

是的。所以最终你们希望在旧金山有一个办公室。如果我不期望很快能搬到旧金山,有可能申请这些职位吗?现在有远程选项吗?

Yeah. So ultimately you're hoping to have an office in San Francisco. Is it possible to apply for these positions if you don't expect to be able to move to San Francisco anytime soon? Is there like a remote option right now?

Chris Olah

目前我们寻找的是最终能够搬到旧金山的人,显然是在疫情解决之后,以及任何移民问题解决之后。但这是我们对大多数职位的长期期望。未来可能会改变,但目前我们对大多数职位的看法就是这样。

Right now we're looking for people who would be eventually able to move to San Francisco, obviously after the pandemic gets resolved and after any immigration issues are worked out. But that would be our long-term aspiration for most roles. It's possible that might change in the future, but right now that's how we think about most roles.

Host

好的,是的。等到这期节目播出时,具体的职位可能已经有所变化,但显然我们会附上你们招聘页面的链接,也许我会在结尾部分提到具体的职位。是的,我想看到这个新项目的发展也会非常令人兴奋。现在还处于早期阶段,但你们已经在媒体上引起了一些关注,很多其他 AI 实验室也发布了一些令人惊叹的研究成果,这些成果甚至吸引了那些对机器学习不太了解的人的想象力。所以希望 Anthropic 也能做到同样的事情。

All right, well, yeah. By the time this episode comes up, the specific roles that are available might have changed a little bit, but obviously we'll stick up a link to your vacancies page, and maybe I'll stick something in the outro about the specific roles that are on offer. Yeah, I guess it will also just be really exciting to see where this new project goes. It's so early, but you've kind of already made a little bit of a splash in the press, and lots of other AI labs have had these amazing pieces of research that have gotten out and captured the imagination of people who aren't even that knowledgeable about machine learning. So hopefully Anthropic maybe is able to do the same thing.

Chris Olah

是的,我的意思是,能和这么多优秀的同事一起工作真的很棒,而且我对我们正在做的研究感到非常兴奋。我真希望我能直接去深入了解这些大型模型内部发生了什么。真的,这就是我想要的。

Yeah, I mean, it's just really lovely to be working with so many wonderful colleagues, and yeah, I'm just really excited about the research we're doing. I wish I could just go and dive into understanding what's going on inside these large models. Really, that's all I want.

副项目与兴趣 Side Projects and Interests

Host

酷,好吧,我们快要结束这次录制了,所以我们很快会让你回去工作。但是,在你重新投入那些模型之前,我想你不会把所有时间都花在机器学习研究上吧。我猜你可能有一些稍微有点无意义的研究项目或兴趣,因为我发现大多数真正聪明的人都有。是的,你最近或过去几年有没有探索过什么有点意想不到的东西?

Cool, well, we're approaching the end of this recording, so we'll let you get back to that very soon. But yeah, just before you dive back into those models, I assume you don't spend all of your time doing ML research. I guess I imagine you have some probably slightly pointless research projects or interests on the side, because I think I find most really smart people do. Yeah, is there anything you've been exploring currently or in the last few years that's a bit unexpected?

Chris Olah

是的,我经常有一些小型的副业项目。有一段时间,我作为一名业余社会科学家,使用 Mechanical Turk 让人们回答问题。然后我痴迷于试图理解进化历史和生命树的不同部分。但我最近痴迷的是大气动力学。你知道,这有点可悲,但我想我有点觉得物理学、生物学和化学是真正的科学,它们有美丽简单的理论来解释很多事情,但气象学和大气科学没有漂亮的理论来很好地解释事物。但你知道吗?实际上有一个非常简单的想法可以解释数量惊人的事情。让我列举一些它解释的事情:它解释了为什么飓风往往袭击大陆的东侧而不是西侧——为什么佛罗里达和日本有很多飓风,而...

Yeah, so I often have some kind of small side project. For a while, I was being an amateur social scientist using Mechanical Turk to get people to answer questions. And then I was really obsessed with trying to understand evolutionary history and understand different parts of the tree of life. But the thing that I've been obsessed with recently has been atmospheric dynamics. And you know, this is kind of sad, but I guess I sort of felt like physics and biology and chemistry are like real sciences that have these beautiful simple theories that explain lots of things, but meteorology and atmospheric science don't have beautiful theories that explain something really nicely. But guess what? It turns out there's actually a really simple idea that explains an absurd number of things. Let me list some things that it explains: it explains why hurricanes tend to hit the east side of continents but not the west side—so why is it that Florida and Japan have lots of hurricanes but...

哈德莱环流与全球天气模式 Hadley Cells and Global Weather Patterns

Chris Olah

加利福尼亚和西班牙不是这样,但为什么如果你看世界地图,所有的大沙漠都在两条纬度线上?为什么旧金山冬湿夏干,而其他地方却相反?最后,为什么木星有条纹?结果发现这些都源于同一个原理。为了说明这有多疯狂,我在旧金山日常经历的天气,实际上与木星有条纹的原因密切相关。

California and Spain don't, but also why is it that if you look at a map of the world, all of the large deserts are on two latitudinal lines? And why is it that San Francisco is wet in the winter and dry in the summer, but there are other places where it's the reverse? And finally, why is it that Jupiter has stripes? And it turns out those are all explained by the same idea. And just to highlight how crazy this is, the reason why the weather that I experience day-to-day in San Francisco is actually intimately related to the reason why Jupiter has stripes.

Host

你引起了我的兴趣,解释一下吧。

You pi my interest, explain it so so okay.

Chris Olah

高层次的概念是,存在一些非常大尺度的大气环流,称为哈德莱环流,暖空气上升,然后移动,再下降。暖空气在赤道上升,然后向外移动到大约 30 度处下降,然后在 60 度再次上升,回到 30 度,再向极地移动。这些就是哈德莱环流。在很大程度上,地球的大尺度天气模式是由这些纬线以及你位于哪条线之间塑造的。它们随着热赤道的变化在一年中迁移。你在木星上看到的条纹也是哈德莱环流。哦,所以它比地球有更多的哈德莱环流。原来,哈德莱环流的数量主要取决于行星的自转速度和半径。所以它们有更多的哈德莱环流。因此,地球的大尺度天气模式——当然,我不是这方面的专家,所以也许你的听众中有气象学家,可以告诉我我全错了——与你在其他行星上观察到的条纹之间存在非常密切的联系。

So the high-level idea is that there are these really large-scale atmospheric circulations called Hadley cells, where warm air rises and then moves and then falls. So warm air rises at the equator and then goes out to about 30 degrees and then falls, and then it also rises again at 60 degrees and goes back to 30 degrees and goes over to the pole. So these are called Hadley cells. In a lot of ways, Earth's large-scale weather patterns are shaped by these lines and which one of these lines you fall behind between. And they migrate over the course of the year as the thermal equator changes. And the stripes you see on Jupiter are also Hadley cells. Oh, so it has more Hadley cells than Earth. It turns out that the number of Hadley cells you have is a function of how fast you spin and the radius of your planet, primarily. So they have more Hadley cells. So there's this very intimate connection between large-scale weather patterns on Earth—of course, I'm not an expert on any of this, so maybe one of your listeners is a meteorologist and can tell me that I'm all wrong—and the existence of stripes that you observe on other planets.

Host

但对我来说最疯狂的是,我们通常考虑相变,比如水变成冰或气体或液体,而相变更一般的概念是系统从一个状态不连续地跳到另一个状态。在我看来,如果你想象让地球的半径越来越大,其他条件不变,那么某个时刻哈德莱环流的数量会改变。原来每个半球必须有奇数个,所以会从每个半球三个变成五个。所以我觉得,实际上在某种意义上,行星的相变作为半径的函数决定了大气模式。这就像,你通常认为相变是关于小尺度的事情,但这里有一个非常大尺度的东西发生了相变。这让我觉得特别震撼。总之,我可以一直聊这个,我简直可以为此做一整期播客,但我发现这非常有趣。

But actually the thing that's craziest about this to me, the thing that I find absolutely nuts, is so okay so we think about phase changes like water transitioning to ice or to gas or liquids, and I guess there is this more general idea of a phase change as being like when a system discontinuously goes from one state to another. And it seems to me that if you imagine just making Earth's radius larger and larger, holding everything else constant, at some point the number of these Hadley cells you have would change. It turns out that you have to have an odd number per hemisphere, so you'd go from three per hemisphere to five per hemisphere. So it seems to me like actually there must in some sense be something about phase changes in planets as a function of radius that determines large-scale weather patterns. And that's just kind of like, you think of phase changes as always being about these small-scale things, but here you have this really enormous scale thing that has a phase change. So that's something that I found especially mind-blowing. And anyways, I could ramble about this, I could literally do an entire podcast episode at this point, but I've been finding that an enormous amount of fun.

Host

是的,所以如果你能让一个行星一点点变大,并跟踪天气,你最终会看到某个点,它会戏剧性地从一个哈德莱环流变成三个或五个。

Yeah, okay, so if you could make a planet bigger and bigger bit by bit and you were tracking the weather, you would eventually see some point where it would flip from having one to three or five Hadley cells quite dramatically.

Chris Olah

完全正确。具体来说,你总是预期赤道是最湿润的地方,至少如果天气模式是基于水的,比如地球的降水。是的,原来当空气上升时,水汽凝结,所以如果你看地球的降水图,很疯狂,赤道沿线有一条降水量巨大的带,非常清晰。然后第一个哈德莱环流在赤道上升然后下降,因为你有三个,所以就是 30 度、60 度、90 度,分成三块。所以在 30 度有一个干燥区域,很多沙漠位于 30 度。如果你突然让地球变大,或者让地球变大,某个时刻你得到五个哈德莱环流,那么我相当确定——这是我的理解——突然你的干燥区域会变成 90 除以 5 度,而不是 90 除以 3 度。所以这有点像,嗯,我觉得这个思想实验非常引人入胜,有点令人震撼。

Exactly. So in particular, okay, so you always expect the equator to be the wettest place, at least if you have weather patterns that are water-based like Earth and precipitation. Yeah, so it turns out when air rises, as it rises the water condenses, so if you look at a map of precipitation on Earth, it's crazy, there's like a stripe of ridiculous amounts of precipitation along the equator, it's extremely clear. And then the first Hadley cell that rises at the equator and then falls, and because you have three, it's just 30 degrees, 60 degrees, 90 degrees, it's just divided into three chunks. So at 30 degrees you have a dry region, and a lot of deserts fall on 30 degrees. If you suddenly made Earth larger, or if you made Earth larger, at some point you got five Hadley cells, then I'm pretty sure—this is my understanding—suddenly you'd have your dry region at 90 divided by 5 degrees instead of 90 divided by 3 degrees. So that's kind of a, yeah, I find that thought experiment really compelling and sort of mind-blowing.

Host

好的,我们会附上哈德莱环流的维基百科链接。这非常酷。前几天我和同事尼尔·鲍曼散步。他很久以前就开始研究气候科学了。我问,云是怎么形成的?我不知道。我觉得自己像个白痴,问这么基本的问题:为什么那里有云,而紧挨着的地方却没有?还有各种关于不同类型雷暴以及雷暴能否持续很长时间的迷人内容。

Yeah, all right, we'll stick up a link to the Wikipedia article on Hadley cells. That's very cool. I was out taking a walk with my colleague Neil Balman the other day. I think he started climate science long ago. I was like, why do clouds form? I don't know. I felt like such an idiot asking this super basic question: why is there a cloud there but not right next to where the cloud is? There's also all sorts of fascinating stuff about different kinds of thunderstorms and the circumstances under which a thunderstorm can last for a long period of time or not.

Chris Olah

是的,结果发现,我不知道,我已经被说服了,气象学和大气科学中确实有非常美妙的思想。我会努力不陷入物理学、生物学、化学至上的陷阱,认为只有这些科学才有真正优美简洁的解释性理论。

Yeah, it turns out, I don't know, I've been persuaded that there actually are really beautiful ideas in meteorology and atmospheric science. And I will endeavor to not fall into the trap of physics, biology, chemistry supremacy and thinking that those are the only sciences that have really beautiful simple explanatory theories.

Host

很好。嗯,也许我们以后会在播客中做一期关于气候和天气的节目。看看我们能不能把它联系起来,嗯,我想气候变化是最紧迫的问题之一。好的,我今天的嘉宾是克里斯·奥拉。非常感谢你来到《八万小时》播客,克里斯。

Nice. Well, yeah, maybe we'll have an episode on climate and the weather on the podcast at some point. We'll see if we can connect it to, well, I suppose climate change is one of the most pressing problems. All right, my guest today has been Chris Olah. Thanks so much for coming on the 80,000 Hours podcast, Chris.

Chris Olah

不客气,罗布。正如我在开场白中提到的,我们还有一期与克里斯的节目正在制作中,希望下周发布,如果不是下周,也会很快。那一期将涵盖的话题包括:克里斯在没有学位的情况下如何取得今天的成就,他如何进入谷歌大脑和 OpenAI,他创办的期刊《Distill》,以及如何很好地解释复杂事物,克里斯如何推荐写冷邮件,还有很多其他内容。所以如果你喜欢这一期,请务必下周或再下周回来听那一期。如果你不喜欢这一期,请注意那一期没有,或者几乎没有技术性 AI 内容,所以你很可能仍然会喜欢。最后,如果你有兴趣利用你的职业生涯来安全地引导 AI 发展,就像克里斯正在做的那样,或者致力于解决我们节目中经常讨论的任何问题,那么你可以申请与我们的团队进行一对一的免费交谈。这是很长一段时间以来的第一次,今年我们有足够的建议和能力取消了等待名单。

My pleasure, Rob. As I mentioned in the intro, we have another episode with Chris in the works, and we hope to release it next week, or if not next week, soon after. And that one should cover topics including how Chris got where he is today without having a degree, how he managed to get his foot in the door at Google Brain and OpenAI, the journal that he founded called Distill, and how to go about explaining complex things really well, how Chris recommends writing cold emails, and a whole lot more besides that. So if you enjoyed this episode, make sure you come back for that one next week or the week after. And if you didn't enjoy this one, note that that one has no, or at least almost no, technical AI content, so you might well be able to enjoy it nonetheless. Finally, if you're interested in using your career to work on safely guiding the development of AI like Chris is doing, or working to solve any of the problems that we regularly discuss on the show, then you can apply to speak with our team one-on-one for free. For the first time in quite a while, this year we've had enough advice and capacity to remove the waitlist.

播客介绍与邀请 Podcast Introduction and Offer

Host

我们的团队非常希望与更多听众交流。作为常规播客听众,我们的顾问可以帮你分析哪个问题最值得你专注,审视你的计划,评估是否适合你,并为你介绍可能加速你职业发展的导师,甚至推荐更契合你技能的其他岗位等类似建议。你可以访问 80000hours.org 了解更多服务详情,并申请加入。80000 Hours 播客由 Kieran Harris 制作,音频母带由 Ben Cordell 处理。完整文字稿(包括 Chris 为本期节目提供的特别版图片和链接)已发布在我们的网站上,一如既往由 Sofia Davis-Fel 制作。感谢收听,下次再聊。

Our team is keen to speak with plenty more of you. Regular podcast listeners, our advisers can go over which problem it might be most effective for you to focus on, take a look over your plan, and think about whether it's a good fit for you. Introduce you to mentors who might be able to speed you along in your career, maybe suggest alternative roles that could really be a good fit for your skills and other things in that general vein. You can go to 80000hours.org to learn more about the service and apply if you feel like it. The 80000 Hours podcast is produced by Kieran Harris. Audio mastering is by Ben Cordell. Full transcripts, including the special edition of some images and links provided by Chris for this episode, are available on our website and produced as always by Sofia Davis-Fel. Thanks for joining. Talk to you again soon.

互动版:逐字朗读 + 针对本期提问 →