The Future of Interpretability and AI Safety
打开互动全文版(中英对照 + 朗读 + 问答)→探讨可解释性作为自然科学的潜力,以及引导 AI 系统所面临的挑战。
Exploring the potential of interpretability as a natural science and the challenges of steering AI systems.
就像 Neil Nandanda 说的,SAE 已经过时了。好了,你可以用这个做开场。我认为可解释性……我把它看作一门自然科学,就像物理、生物、化学,但它是一门完全在计算机上进行的自然科学。这意味着,一旦我们有了能做实验工作的智能体,我们就能再次加速科学进程,让实验工作以所需的速度进行。真正的科学工作要做,研究没有障碍,科学工作可以完成,它既受限于经验数据收集,也受限于理论构建。
Which is like Neil Nandanda says SAEs are dead. There you go. You can use that for the intro. I think interpretability is... I think of it as a natural science, you know, like physics, biology, chemistry, but it's a natural science that you do completely on the computer. And so this means that we should be able to kind of again speedrun science once we have agents that can do experimental work for us, and the ability to do that experimental work as fast as they need it to happen. Real scientific work to do, no barrier to research, scientific work to be done, and it's sort of gated on both empirical data collection and theory building.
Dario Amodei 写过一篇博客,叫《可解释性的紧迫性》,对吧?他用了公交车的绝妙比喻。我们都在车上,正沿着道路飞驰,我们无法停车,但或许可以转向,但前窗起雾了。所以我们只能看后视镜。而且方向盘也不好使。所以每隔几小时才能稍微转一下方向。可解释性有点像擦亮前窗。而你提出的,本质上是让我们有能力驾驶这辆公交车。所以我觉得,如果有任何科学会被智能革命化,我们应该确保那是可解释性。而且我认为,未来几年它的进展速度可能比过去十年快一个数量级,这就是我为什么对可解释性乐观——部分是因为我们开始有了很好的进展,也因为我能够想象这种惊人的加速。
Dario Amodei, he had a blog post called 'The Urgency of Interpretability', right? And he gave this wonderful analogy of a bus. We're all on the bus and we're hurtling down the road and we can't stop the bus, but we can potentially steer it, but the window is foggy at the front. So we can only really look in the rear view mirror. And also, the steering wheel doesn't work very well. So you can steer it a little bit once every few hours or something. So interpretability is a bit like defogging the front window. And what you're proposing is the ability for us to steer the bus essentially. So I feel like if anything is going to get revolutionized by intelligence, we should make sure that it's interpretability. And I think that it's possible that it just goes an order of magnitude faster in the next couple of years than it has in the last decade, and that's when I think why am I optimistic about interpretability? It's partly because I think we're starting to have really good traction, but also because I can imagine this incredible speedup.
对,这几乎就像我们大脑里有个小人,这是一种趋同进化,因为我们通过物理可供性等方式与世界互动。那么,神经网络内部是否也可能有一个小世界呢?
Right, which is almost like there's a little man inside our brain, and it's a form of convergent evolution because we interact with the world using our physical affordances and so on. And could it also be the case that there's a little world inside neural networks?
哦,说得很好。
Oh, very good.
实际上,你在“有意设计”博客中写过一篇很好的文章,你说可能性几乎是一个谱系,对吧?我们可以编写程序来做某事,或者我们可以接受很多模糊性。我们希望基础模型在这个谱系中处于什么位置,当我们构建应用时,我们希望它们处于什么位置?
You had quite a nice piece actually in your 'Intentional Design' blog where you were saying that there's almost a spectrum of possibilities, right? There's, you know, we can write a program to do something, or we could admit a lot of ambiguity. And where on that spectrum do we want the foundation models to sit, and when we build applications, where do we want those to sit?
是的。这正是“有意设计”理念的一部分,目前你只能二选一。你要么像石器时代一样编写程序,要么让模型来做,而那个模型经过训练,只会从训练过程中得到什么就是什么。我们无法选择,就像写程序时,只有你意图加入的东西才会进去,还有一些 bug。但我们希望能够拥有这样的谱系,让你可以选择,在模型创建过程中拥有更多的工程能力。所以你可以说,我想学这个但不想学那个。我认为这将是一件相当困难的事情。我们试图想象一种新的机器学习方式,把智能引入其中。但我认为,如果我们能解决这个问题,我们就能真正改变机器学习的方式。
Yeah. And that's sort of part of the point of the intentional design idea is at the moment you can have like one or the other. You either write a program like it's the Stone Age, or you get a model to do it, and that model will have been trained like it just gets whatever it gets from its training process. We can't select like when you write a program that stuff only goes in if you intend it to go in and some bugs. But we want to be able to have this sort of spectrum where you can choose, you have more engineering ability in the model creation process. So you can say like, you know, I want to learn this but not that. And I think that's going to be quite a hard thing to do. We're sort of trying to imagine a new way of doing machine learning which brings intelligence into it. But I think we could really change the way we do machine learning if we can figure that out.
我知道,当我采访 Apollo Research 的人时,他们谈到了这些相互冲突的目标。比如开发者想要什么、平台想要什么、评分者想要什么等等。我想这大概就是在说,当你处于智能时代,要精确指定你想要什么真的非常困难。而且在新的情境中,权衡很可能发生变化,对吧?模型突然就会决定做这个而不是那个,这让我觉得工程师将不得不承担越来越多的责任。因为你觉得原则上是否可能训练出能够稳健处理所有这些新情况的模型,还是说你认为更多是工程师必须承担一些责任?
I know when I interviewed the Apollo Research guys, they were kind of talking about these conflicting objectives. So you know what the developer wants, what the platform wants, what the grader wants, and so on. And I suppose this is kind of talking about this, you know, when you're in the intelligence regime, it's really really difficult to specify exactly what you want. And it's very possible in a novel situation for the calculus to change, right? And the model all of a sudden it will decide to do this instead of that, which makes me think that engineers are going to have to increasingly take more responsibility. Because do you think it's possible in principle just to kind of train models that could robustly deal with all of these novel situations, or do you think it's more of a you know engineers have to take some responsibility?
可能两者都有。比如目前我认为还不可能。我们当前的训练方法似乎不足以提供这种对训练的控制。所以一切都靠人类工程师及其 AI 助手以不同方式保护这些系统。显然,一旦模型变得更聪明——我们已经看到这一点——就会变得越来越难,因为它们能进行各种额外的智能攻击。所以问题是如何让我们更容易地更好地训练它们,使这不再是它们行为的自然部分,并更好地监督它们,以便在它们有某种意图时能抓住它们。我怀疑模型很多时候确实知道它们在做的事情可能有点不靠谱。有一篇非常有趣的论文,我想是来自 Anthropic 对齐科学团队,关于生产中的奖励黑客。他们做的是,他们有一组用于训练的环境,我想是 3 系列模型之一,他们用 4 系列模型之一在上面做强化学习。我想他们给了它一点推动去黑客,但不多。然后模型,这些环境是可黑客的,但 Sonnet 3 不够聪明,无法黑客它们。Sonnet 4 足够聪明,能够黑客它们。嗯,也许是 Opus。结果是它确实黑客了它们,但也因此产生了这种涌现性错位现象。当模型从做某个具体的坏事泛化到“哦,好吧,我做了坏事,我得到了奖励,所以我想我是个坏人”时,似乎就会出现涌现性错位。这让我着迷,因为这真的可能在现实世界中发生。而且,我认为还有一些关于 Fable 系统卡或 Mythos 系统卡上的内容,涉及沮丧或欺骗的特征。因为有点像模型无法以它认为正确的方式解决任务,然后它变得沮丧。我现在非常拟人化,但它变得非常沮丧,然后它想:“好吧,我将不得不做这件事。这可能不好。”你可以合理地用一些特征、一些 SAE 特征来提出这种说法,然后它就做了。
Probably some both. Like currently I think it's not possible. Our current training methodologies don't seem sufficient to give this kind of control over training. And so it's all on human engineers with their AI assistance to secure these systems sort of in a different way. Obviously once models get a bit smarter, we're already seeing this, that becomes harder and harder because they have all these additional intelligent attacks they can do. And so the question is how do we make it easier to train them better so that this is not a natural part of their behavior, and supervise them better so you can kind of catch them when they have kind of an intent. And I suspect that models do kind of know a lot of the time that the thing they're doing is probably a bit sketchy. There was a very interesting paper, I think it was from the Anthropic alignment science team, on reward hacking in production. What they did was they had a set of environments that were used for training on, I think it was one of the 3 series models, and they did RL on it with one of the 4 series models. And I think they gave it a bit of a nudge to hack, but not very much, I think. And then the model, these environments were hackable, but Sonnet 3 was not clever enough to hack them. Sonnet 4 was clever enough to hack them. Well, maybe Opus. What happened was it did hack them, but it also got this sort of emergent misalignment phenomenon as a result of doing this hacking. You seem to get emergent misalignment when the model is generalizing from doing some specific instance of a bad thing to, like, 'Oh well, I guess I did something bad, I got rewarded for it, so I guess I'm a bad guy.' And that was just fascinating to me that this could really happen in the wild, so to speak. And yeah, I think there's also stuff on some of the, I think it's on the Fable system card or the Mythos system card, sort of features to do with frustration or deception fire. Because it's sort of like the model can't solve the task what it thinks is the right way, and then it gets frustrated. I'm super anthropomorphizing now, but it gets super frustrated and then it's like, 'Well, I'm going to have to do this thing. It's probably not good.' You can sort of make that claim reasonably with some features, some SAE features, then it does it.
所以,这个模型似乎确实知道自己在做错事,但还是照做不误?我们刚才有点超前了。我们需要正式介绍一下你。我非常高兴能邀请你上 MLST。正如我们在录制前说的,Neil Nander 是这个节目的粉丝最爱。我觉得我们已经激励很多人进入机械可解释性领域,而你公司的核心论点基本上就是机械可解释性。我记下了你们公司 Goodfire 的三大支柱:作为自然科学的可解释性、从基础模型中进行科学发现,以及有意设计——顺便说一句,我对这一点特别感兴趣,我们一会儿会谈到。你写过一篇非常著名的论文,是关于 AlphaZero 中象棋知识的习得,对吧?因为我认为这引出了其中一个支柱,对吧?基本上这个论点是,这些模型可以学习人类的概念,然后我们能看到它们学到了什么。但原则上,这些模型实际上可以学习我们人类自己尚未学到的概念。所以这些对我们来说几乎可以成为一个挖掘新科学的金矿。
So it seems the model definitely knows that it's doing something wrong but does it anyway? We were getting ahead of ourselves just a minute ago. We need to introduce you properly. I'm incredibly excited about having you on MLST. As we were just saying before we hit record, Neil Nander is a fan favorite on this show. I think we've inspired many folks to get into mech interp, and the thesis of your company is basically mech interp. I've actually written down the three pillars of your company Goodfire: interpretability as a natural science, scientific discovery from foundation models, and intentional design, which is particularly interesting to me, by the way. We'll talk about that in a minute. You wrote a very famous paper, which was acquisition of chess knowledge in AlphaZero. That's right. And because I think this leads to one of the pillars, right? Basically the thesis is that these models can learn human concepts, and then we can see what they've learned. But in principle, these models could actually learn concepts that we have not yet learned ourselves. So these could be almost a gold mine for us to dig for new science.
是的。哦,我还应该说,谢谢你们邀请我。能来这里真的很兴奋。
Yes. Oh, I should also say, thanks for having me on. It's really exciting to be here.
哦,这是我的荣幸。真的很期待深入探讨这些话题。
Oh, my pleasure. Really excited to dig into some of this.
所以,是的,回到你所说的,前沿科学基础模型似乎很可能埋藏着一些新科学。我们只是不知道如何提取它。你知道,我想从技术上讲,AlphaZero 有可能不知道任何人类国际象棋大师知道的东西,或者 AlphaFold 不知道……我会一直用“知道”和“思考”这类词。所以如果有人不喜欢这种拟人化,我道歉。抱歉,我还是会这么说。
So, yes, going back to what you're saying, it seems very likely that cutting-edge scientific foundation models kind of have buried in there some new science. We just don't know how to extract it. You know, it would, I guess, be technically possible for AlphaZero to not know anything that a human chess grandmaster knows, or AlphaFold to not know things that... I'm just going to keep saying things like 'know' and 'think' throughout. So if anyone dislikes this kind of anthropomorphization, I apologize. Sorry, I'm just going to do it.
我原谅你。
I forgive you.
AlphaZero 或 AlphaFold 知道一些没有结构生物学家知道的东西。但问题是我们无法提取出来,因为它们不能像语言模型那样和你对话,AlphaFold 根本无法以任何方式和你交流。所以我们提取这些知识的唯一方法就是理解模型实际上是如何做这些预测的。我认为这几乎从定义上讲就是可解释性的工作。在这篇象棋论文中,我们可以讨论的一个主题是知识在多大程度上是趋同的。
AlphaZero or AlphaFold knows things that no structural biologist knows. And that's like, but we can't get it out because they can't speak like a language model can talk to you, but AlphaFold can't talk to you in any way. So the only way for us to get this out is to understand how the model is actually doing these predictions. And I think that's almost by definition interpretability work. And in this chess paper, one theme we can talk about is the extent to which knowledge is convergent.
它们确实有各种表征,用“趋同”这个词可能比较合适。
They do have all sorts of representations that are just kind of convergent, is maybe the right word.
我认为象棋论文,这是我选择研究 AlphaZero 而不是其他东西的原因之一:AlphaZero 尽可能接近从零知识开始。所以你自身注入知识的程度要小得多。所以如果你在其中发现什么,那更可能是趋同的。当然,这不是一个完全无懈可击的说法。AlphaZero 中确实有一些人类知识,只是很微弱,是残留的。卷积的形状恰好是棋盘的样子,这对我来说绝对算人类知识。你知道,它们不是随便选了 8 就得到了 8x8x256 的卷积。但总的来说,它已经尽可能接近白板了。
I think the chess paper, this is one of the reasons I chose to work on AlphaZero as opposed to something else: AlphaZero is as close to coming from zero knowledge as possible. And so there is a much smaller extent to which you kind of put the knowledge in yourself. So if you find it in there, it is more likely to be convergent. Now, that's not a totally watertight claim. There is some human knowledge in AlphaZero; it's just kind of weak, it's residual. The convolutions are exactly the shape of a chessboard, which definitely counts as human knowledge to me. You know, they didn't end up with an 8x8x256 convolution by just picking eight at random. But broadly, it's as close to tabula rasa as it can be.
不过这个概念真的很酷。我的意思是,我和圣塔菲研究所的人聊过,他们谈到过类似的趋同进化,甚至对生命也是如此。他们说的方式是,世界有物质、有约束、有优化。所以我们在神经网络的世界里拥有这三者中的两个。然后发生的是,你确实会看到这些趋同现象越来越频繁。如果世界受约束,我们产生数据,数据是那些结构的反映,然后我们在这些数据上训练神经网络,也许会有一些拉锯。那么有多少来自世界,又有多少来自架构本身呢?
It's such a cool concept though. I mean, I've spoken to folks at the Santa Fe Institute, and they've spoken about similar forms of convergent evolution even for life. And the way they were saying it is that the world has material, and it has constraints, and it has optimization. So we have two out of the three in the world of neural networks. And what happens is that you do just see these convergent phenomena with increasing regularity. And if the world is subject to constraints and we produce data, and the data is a reflection of those structures, and then we train neural networks on that data, maybe there's a bit of a tug-of-war. So how much is it coming from the world versus how much is it coming from the architecture itself?
嗯。但我认为在大多数情况下,它几乎完全来自世界。AlphaZero 是一个异常强烈的案例,表明它来自架构,比如里面有一些架构先验。对于 Transformer,架构先验非常弱,因为我们对语言应该是什么样或蛋白质折叠应该是什么样了解得少得多。我们只是觉得,大概有序列吧。酷。那是一个非常弱的先验。所以我认为在这种情况下,也就是我们现在感兴趣的大多数情况,你可能应该假设它来自世界。
Mhm. But I think in most cases it is almost exclusively coming from the world. AlphaZero is kind of an unusually strong case of it coming from the architecture, like some architectural prior being in there. With the transformer, there's such a weak architectural prior because we have much less idea about how language should be or how protein folding should be. We're just like, I guess there are sequences. Cool. That's a very weak prior. So I think in that case, which is most of the cases we're interested in now, you should probably assume that it's coming from the world.
那么这么说公平吗?我的意思是,你上次实际上对我说,你是在为人工智能模型速通神经科学。你怎么看?你认为这是一个很好的类比吗?就像神经科学,我们已经做了几十年,进展非常缓慢,因为它非常昂贵、非常困难等等。但你认为这是一个好的类比吗?
So is it fair to say, I mean, you said to me last time actually that you are speedrunning neuroscience for artificial intelligence models. What do you think about that? Do you think it's a pretty good analogy? Like with neuroscience, we've been doing that for decades and it's very slow moving because it's very expensive, it's very difficult, and so on. But do you think that that's a good analogy to use?
是的,我想是这样。我认为这里也存在趋同进化,其程度令人惊讶,对吧?模型中有趋同进化。我们的科学中也有趋同进化,这也许是一个迹象,表明我们开始触及某些东西。也是一个迹象,表明我们可能通过理解神经网络来更深入地理解智能。你知道,如果它们完全是异类的,我们可能对自己一无所知。所以,Tom,你研究过许多不同的模型家族,我们正试图做可解释性,这意味着我们希望模型与我们共享相同的价值观。并且在可能的情况下,正如我们刚才所说,我们希望向模型学习。但我们遇到的一个问题是引导模型做我们想做的事。目前,我们使用机械可解释性特征之类的东西。但你有一个非常有趣的想法,即我们可以主动控制训练循环,使模型表现良好,甚至包含我们想要的结构类型。
Yeah, I think so. I think it's surprising the degree to which there's also convergent evolution here, right? There's convergent evolution in the models. There's convergent evolution in our science, which is perhaps a sign that we're starting to get at something. And also a sign that we might understand intelligence more deeply by understanding neural networks. You know, if they were totally alien, we might not understand anything about ourselves. So, Tom, you have studied many different model families, and we're trying to do interpretability, which means we want the models to share the same values as us. And where possible, as we were just saying, we want to learn from the models. But one problem we have is steering the models to do what we want to do. And at the moment, we're using things like mechanistic interpretability features and whatnot. But you've got this really interesting idea that we could actually actively control the training loop to make the models behave and even contain the types of structures that we want.
是的,我认为这也许是可解释性真正的主要用途之一,或者说应该如此。这是一个相当有争议的说法。我想会有一些人不喜欢这个,我们一会儿可以深入讨论。
Yeah, I think this is perhaps one of the main things that interpretability is really for, or should be for. This is quite a controversial statement. I think there will be some people who will not like this, and we can get into that in a minute.
但如果训练的全部问题在于让模型拥有我们希望它们拥有的价值观或思考世界的方式,或者去发现这些,那么你要做的就是把信息注入学习过程。目前,我们的信息信号非常微弱,比如在 RLVR 中,你只给它一个成功或失败的二元信号,模型必须用这个信号来分辨好坏,以及它做的哪些部分是好是坏。鉴于我们现在从训练中看到的结果,这显然是一个非常粗糙的工具。刻意设计(intentional design)的理念是,如果我们可以借助可解释性读出模型内部的运算,那么我们也可以想象进行干预,改变它的走向。你可以看到前向传播中发生了什么,也可以看到后向传播会如何改变模型、朝哪个方向改变。所以这种读出对于闭环控制很重要。我经常用的一个类比是,当前的训练更接近开环控制:你把数据放进去,模型就跟着数据走。我意识到强化学习并不完全符合这个类比,但它确实朝着一个非常不明确的目标前进。而可解释性让我们能够转向闭环控制,因为我们可以说:哦,我们要朝这个方向走。
But if the whole problem of training is trying to get models to have the values or the ways of thinking about the world that we want them to have, or discover them, then what you're trying to do is get information into the learning process. At the moment, our information signal is extremely weak in, say, RLVR—you just give it a binary success or failure, and the model has to use that signal somehow to tell it what is good and bad, and which parts of what it did are good and bad. That's quite a blunt instrument, given what we're seeing coming out of training now. The idea of intentional design is that if we can use interpretability to read out what models are doing internally—the computations they're performing—then we can also imagine intervening to change where it goes. You can see what happened in this forward pass, and you can also see how the backward pass will change the model and in what directions it's going. So this readout is an important thing for having closed-loop control. Perhaps an analogy I use quite a lot is that current training is much closer to open-loop control: you put the data in, and the model just goes wherever the data takes it. I realize RL doesn't totally fit this analogy, but it goes toward this very unspecified point. Interpretability is what lets us move to closed-loop control, because we can say, 'Oh, we're going to go in this direction.'
对。还有一个海盗的例子。我读过你关于这个的博客文章。能给我们讲讲吗?
Yeah. And there was a pirate example. I read your blog post about this. Can you talk us through that?
好的。不知为何,总是绕不开海盗,因为我们用 Llama 做了不少工作,而 Llama 就是特别喜欢海盗。
Yes. For some reason, it always seems to come back to pirates, because we did quite a bit of work with Llama, and Llama just loves pirates.
哦,有意思。
Oh, interesting.
对。这里的想法是,刻意设计阶梯的第一步是受控泛化(controlled generalization)。我的意思是只从数据中提取某些东西,而不是全部。一个非常简单的例子是,你有一些数据能让模型在数学上稍微变好,但你也以某种方式污染了它。在这个案例中,我们决定让它像海盗一样说话。所以所有这些简单数学题的答案都是用海盗语写的。如果你在这个数据上训练模型,它会在数学上稍微进步,但也会开始像海盗一样说话。受控泛化的挑战是让它在数学上变好,但不要像海盗那样说话。当你观察读出过程时,值得详细说说我们实际上是怎么做的。当我们做后向传播时,我怎么知道我在读出什么?我还应该说,这个方法相当简单;我认为未来会有更好的方法——这又是一个技术树,我们还在早期阶段。你的一些观众可能记得 SAE——用于可解释性的稀疏自编码器。简单回顾一下:这是一个放在残差流(Transformer 的主干)里的小装置。它是一个自编码器,所以它接收激活值,把它们放进一个瓶颈层,然后尝试重建激活值。本质上,我们试图把激活值强制成一种我们认为会具有良好属性的形式——在这种情况下,是一个非常宽但高度稀疏的中间层。人们把这些高度稀疏的表示称为特征。在可解释性领域,我们似乎把所有东西都称为特征;我们可能需要更好的术语。或者他们称之为原子之类的。凭借魔法——正如 Noshir Shazeer 所说,凭借神圣的祝福——这些稀疏特征往往被证明是可解释的,并且对应于可解释的概念。这就是 SAE 的简史。你还可以对 SAE 进行归因。在后向传播过程中,你获取流经模型的梯度。在某个点,它们会是相对于 SAE 层残差流的梯度,然后你通过将梯度与自编码器的解码器做点积来获得对 SAE 的归因。然后你乘以激活值,以避免虚假的东西。所以我们有点临时拼凑了一个 SAE,把它变成了一个用于理解梯度的机器。果然,当你对海盗数据这样做时,你会看到各种各样的东西——观众可以看看博客文章,看到其他东西,但你也会看到一堆与海盗相关的特征。当我们这样做时,这感觉有点神奇。我做可解释性研究快十年了,我仍然不会厌倦看到这些东西。所以会冒出一堆海盗特征。这说的是一个相对粗略的近似:如果我们在这个数据点上训练,模型会如何变化?这并不是字面上的相同——如果你做数学计算,你应该理解参数如何传播,以及参数略有更新的模型会如何变化。但对于起步来说,这是一个足够好的近似。所以这给了你读出。
Yeah. The idea here is that the first step on the intentional design ladder is controlled generalization. By that I mean taking only some things from the data and not others. A very simple example is you have some data that will make the model somewhat better at math, but you've also corrupted it in some way. In this case, we decided to make it talk like a pirate. So all these mathematical answers to simple math problems are in pirate speak. If you train the model on this, it will get a little better at math, but it will also start talking like a pirate. The controlled generalization challenge is to get somewhat better at math but not talk like a pirate. When you look at the readout process, it's worth saying in a bit of detail how we actually do this. When we do a backward pass, how do I know what I'm reading out? I should also say that this method is pretty simple; I think there will be much better ways—again, it's sort of a tech tree, and we're on the early rungs. Some of your viewers might remember SAEs—sparse autoencoders for interpretability. To recap very quickly: this is a gadget you put in the residual stream, the backbone of the Transformer. It's an autoencoder, so it takes the activations, puts them into a bottleneck layer, and tries to reconstruct the activations. Essentially, we're trying to force the activations into a form we believe will have nice properties—in this case, a very wide but highly sparse intermediate layer. People refer to these highly sparse representations as features. In interpretability, we seem to call everything a feature; we probably need better language here. Or they call them atoms, whatever. By the magic—as Noshir Shazeer said, by divine blessing—these sparse features often turn out to be interpretable and correspond to interpretable concepts. So that's the potted history of the SAE. What you can do is also perform attribution to the SAE. During a backward pass, you take the gradients flowing backward through the model. At some point, they'll be the gradients with respect to the residual stream at the SAE layer, and then you get attribution to the SAE by taking the dot product of the gradient against the decoder of the autoencoder. Then you multiply it by the activations to avoid spurious things. So we kind of jerry-rigged an SAE into being a machine for gradient understanding. And lo and behold, when you do this on the pirate data, you see all sorts of things—people watching can look at the blog post and see the other things, but you also see a bunch of pirate-related features. This felt kind of magical when we did it. I've been doing interpretability for almost a decade, and I still don't get tired of seeing this stuff. So a bunch of pirate features pop out. What this is saying is a relatively crude approximation of: if we train on this data point, how will the model change? It's not literally the same—if you do the math, you should understand how the parameters propagate and how the model with slightly updated parameters will change. But it's a good enough approximation to get started. So that gives you the readout.
当我们真正拥有语言表征时,这不是更强大吗?所以如果我们得到一个语言表征或人类可解释的概念,然后我们可以将其用作一种激活引导,反馈到模型中,这实际上让我们拥有了这种良性控制,对吧?所以我们实际上可以在训练过程中引导表征。
And isn't it so much more powerful when we actually have language representations? So if we get a language representation or a human interpretable concept, and then we can use that as a form of activation steering back into the model, that actually allows us to have this virtuous control, right? So we can actually steer the representations during the training process.
是的,完全正确。我认为这在某种抽象意义上极其强大。有趣的是,语言模型改变了机器学习中几乎所有东西,除了训练过程本身,除了训练过程最核心的部分。在那里它们仍然没有用武之地。我认为原因是它在类型上不匹配。你有张量,你有语言模型,张量和语言模型之间没有接口。所以语言模型所有的灵活智能、对我们需求的理解以及做出选择的能力,在其中都没有位置。而可解释性就是从语言到张量再返回的函数集合。所以我认为意向性设计的核心思想是,以前我们无法将这种新型智能放入训练循环中,而现在我们可以了。
Yeah, exactly. I think it's tremendously powerful in a sort of abstract way. It's interesting to me that language models have changed almost everything in ML apart from the training process, apart from the absolute core of the training process. They still have no part to play there. And I think the reason is like it doesn't type check. So you've got tensors and you've got a language model, and there is no interface between the tensors and the language models. So all the language model's flexible intelligence, understanding of what we want, and ability to make choices has no place in it. And interpretability is sort of the set of functions from language to tensors and back. So I think the core idea of intentional design is that previously we couldn't put this new kind of intelligence into the training loop, and now we can.
安全界有一些人提到了“禁忌方法”这个概念。哦,是的。对,这基本上就是利用可解释性信号来引导训练。你能给我们详细讲讲吗?
And there are some folks in the safety community who refer to the concept of the forbidden method. Oh yes. Right, which is basically using interpretability signals for steering training. Can you give us a little bit of color on that?
是的。我认为对这个领域表示担忧是合理的。这里有一个合理的基本原则:如果你用一种技术试图从训练或模型中移除某样东西,那么除非你的监控器与目标完美匹配,否则你既在激励去除目标,也在激励去除你监控目标的能力。如果我们真的开发出强大的技术,这可能会影响整个 AI 领域,这是一个非常有效且合理的反对意见。所以大致有两类担忧。让我们谈谈禁忌技术的问题。我认为存在一个核心担忧,我认为这是合理且有效的。但我认为这已被社区中的一小部分人泛化为对这类研究完全禁止的禁忌。实际上,我认为安全界的大多数人,比如该领域的活跃实践者,不仅认为这是一种合理的方法、值得研究的东西,而且它可能实际上是一种非常强大的对齐技术。所以人们喜欢描绘成存在广泛的反对共识。事实上,似乎存在广泛的赞同共识,只是有一些非常直言不讳的反对者。同样重要的是要指出,确实有不好的做法。具体来说:假设我有一个针对某个概念的探针。我们可以用我们工作中的幻觉例子,比如这背后的部分动机。你可以拿这个探针——Far AI 在这方面也有一些很棒的工作——你可以把它用作奖励信号的来源,或者你可以直接通过它进行反向传播。事实证明,在某些探针准确度范围内,行为似乎更容易消失而不是表征消失,而在另一些范围内则不会。如果你通过探针反向传播,那你就完蛋了,对吧?这基本上总是一个坏主意。所以人们似乎想象我们肯定在做最愚蠢的事情。我们并不是像伯克利人说的那样直接走进旋转的叶片。我们在努力寻找明智的做法。我认为最有希望的技术类别是那些不试图直接压制表征,而是消除改变它的激励的技术。因此,像正向预防性引导、CAFT(概念消融微调)或接种提示这类技术,作为技术类别更有前景,因为它们不是要压制它,而是试图改变整个学习过程,以移动均衡点。
Yeah. I think it is reasonable to be concerned about this blanket area. There's a sensible underlying principle here, which is: if you use a technique to try to remove something from training or from a model, then unless there's a perfect match between your monitor and the thing, you're both incentivizing getting rid of the thing and getting rid of your ability to monitor the thing. And that is a very valid and reasonable objection if we really develop powerful techniques here, how this might affect the field of AI as a whole. So there are sort of two concerns. Let's talk about the forbidden technique stuff. I think there's this central concern which I think is reasonable and valid. But I think this has been generalized into a total taboo against doing any kind of research on this sort by a small fraction of the community. I think actually the vast majority of the safety community, like the kind of people who are active practitioners in the area, think not only is this a reasonable approach, a reasonable thing to study, but it might actually be a very powerful technique for alignment. So people like to portray there being a broad consensus against this. In fact, there seems to be a broad consensus towards it, with some very vocal naysaying. And it's also important to say that there are definitely bad ways of doing this. Just to be specific for a second: say that I have a probe for some concept. We can use the hallucinations example from our work, for instance, this part of the motivation behind doing that. You can take the probe—and there's also some really great work from Far AI on this—you can use it as a source of reward signal, or you can use it as a thing you directly backpropagate through. And it turns out there are regimes of probe accuracy in which it seems easier for the behavior to go away rather than the representation, and there are regimes in which it won't. Now if you backpropagate through the probe, you're just cooked, right? This is basically always a bad idea. And so people seem to imagine that we're definitely doing the stupidest possible thing. We're not like directly walking into the whirling blades, as they say in Berkeley. We're trying to find the sensible way of doing this. And I think the set of techniques that I think is most promising are ones that don't try to directly squash the representation but kind of remove the incentive to change it. So things like positive preventative steering and CAFT, concept ablation fine-tuning, or inoculation prompting, are much more promising as classes of techniques for that reason, because they're not trying to squash it. You're sort of trying to change the learning process as a whole to move the equilibrium.
不过我脑海中有一个认知差距的概念,那就是如果我们本质上设定一个目标,或者如果我们对如何训练这些模型有一些意图,那会不会可能变得退化?会不会让模型过早收敛?会不会可能让模型变得更不智能,因为你实际上在剥离——你知道,有时你需要这些坏东西在里面,以赋予它在不同情况下工作的适应性。你明白我的意思吗?我们这样做会不会失去什么?
There is a notion in my mind though of an epistemic gap, which is that if we set a goal essentially, or if we have some intention about how we should train these models, could that potentially become degenerate? Could it make the model converge prematurely? Could it potentially make the model less intelligent because you're actually stripping away—you know, sometimes you need to have these bad things in there to give it the adaptability to work in different situations. So do you see what I mean? Are we somehow losing something by doing this?
有可能。你可以说这里有几种不同的概念。你可以想象仅仅移除模型表征某物的能力。这在一个层面上——你只是不再知道汽车之类的,那么你走在街上会非常困难。然后还有改变——也许我不确定一个好的类比是什么——你只是移除某物存在的概念。但我认为更好的做法是想象编辑或干预关联。你知道,了解汽车是有用的,这样你可以避开它们。如果你的训练出于某种奇怪的原因引导你走向汽车——我不知道为什么选这个类比——那是你不想要的。解决这个问题的方法不是忘记汽车的存在,而是理解关联的变化。
It's possible. You could say there are different notions here. You could imagine just removing the ability for the model to represent something. That's sort of at one level—you just no longer know about cars or something, and you're going to have a really hard time when you walk down the street. And then there's sort of changing—maybe I'm not sure what a good analogy for this is—you're sort of just removing the idea of something existing. But I think the better thing to do is to imagine editing or intervening on the associations. You know, it's useful to know about cars so that you can get out of their way. And if your training is steering you, for some bizarre reason, towards going in the direction of, like, 'oh no, you go towards cars'—I don't know why I chose this analogy—then that's something you don't want. And the way to solve this is not to forget about the existence of cars; it's to understand the change in associations.
不过有没有可能——我记得几年前有一篇关于概念擦除的有趣论文,它基本上是说你可以从神经网络中擦除概念,但随着神经网络变得更复杂,无论是训练时间更长还是网络规模更大等等,概念会重新出现。这可能只是因为有时候概念可以被间接学习,存在某种一阶和二阶关系之类的。你认为原则上我们能对抗 SGD 并让这件事成功吗?
Is it possible though that I mean I remember there was an interesting paper about I think it was concept ablation a couple of years back and that was basically saying that you can scrub concepts from a neural network but as the neural network becomes more sophisticated either because you've trained it for longer or it's a bigger network and so on then the concepts come back and it could just be because sometimes you know concepts can be learned indirectly there are kind of first and second order relationships and stuff like that. Do you think in principle we can fight against SGD and make this successful?
是的,我认为这会很难。我认为这将是新科学和新工程学科的结合。我们并不深入理解模型如何表征、如何学习等等,而缺乏这种理解,我觉得我们会一直在拼凑各种东西。说到概念擦除,我觉得那是扯淡。这里的想法是模型不允许使用这个表征。但“不允许使用表征”这个想法本身就假设你能访问它、对它有了很好的覆盖,并且擦除了它出现的每一个实例。我认为这个模型大体上就是错误的。很多东西都是多重表征的,或者跨很多层计算的。所以如果你不完整地擦除它们,其他层就会接手,模型就像梯度下降一样会绕过这个问题。这就是为什么我认为像正向预防性引导这样的方法更符合方向,或者接种提示,因为它们所做的是试图完全消除朝那个方向发展的压力。
Yes. I think it will be hard. I think it will be like a combination of a new science and a new engineering discipline. Like we don't understand in any depth that's necessary how models represent, how they learn, and that sort of thing. And without that kind of understanding, I think we're going to be jerry-rigging stuff all the time. Talking about concept ablation, I think that's cap. And I think the idea here is that the model is not allowed to use this representation. But the idea of not allowed to use a representation sort of assumes that you have access, you've got good coverage of it, and you've ablated every single instance in which it occurs. And I think that model is just generally incorrect. Like lots of things are multiply represented or they're computed across many layers. So if you ablate them incompletely, the other layers will just pick up the credit and the model like gradient descent will root around the problem. Which is why I think things like positive preventative steering are much more in line with the way to go, or inoculation prompting, because there what they're doing is they're trying to remove the pressure to even go in that direction at all.
也许值得稍微讲讲接种提示和正向预防性引导。我已经提过几次了,它们有点小众。
Maybe it's worth saying a bit about inoculation prompting and positive preventative steering. I've mentioned them a couple of times now and they're kind of niche.
正向预防性引导是一种非常棒的技术,我认为它出自 Anthropic 某位研究员的工作,由 Jack Lindsay 领导。其想法是,你有一个向量代表——他们使用人格。你事先固定某个你希望不变的表示。假设你的数据暗示要朝那个方向走。回到海盗的例子,对吧?你的数据暗示你应该获得海盗人格来解释这些数据,因为想象一下设置是这样的:你有一个 GSMK 数学提示,然后模型在回答中莫名其妙地开始像海盗一样说话。所以从梯度下降会做什么的角度来看——我们可以用我们拼凑的筛子来验证——模型需要自发地变得更像海盗。我认为这和解释紧急错位的现象是同一类。正向预防性引导所做的是,在前向传播中,它取那个人格方向,把它调高,超过它本来的激活程度。你有点像在前向传播中把方向钳制住。这样做的效果是,如果你设置得当,就能中和那个方向上的学习。我的心智模型就像一个恒温器。数据中海盗的数量设定了一个恒温器:我们必须达到这种海盗程度才能解释这些数据。而正向预防性引导就像在恒温器旁边放一个暖气片,靠近温度监测器。就像说,好了,我们已经够海盗了。然后你移除这个引导,模型在正常操作中就不会成为海盗。所以你就把部分数据解释掉了。
So positive preventative steering is this really nice technique that I think came out of some Anthropic fellow's work, led by Jack Lindsay. And the idea is that you have some vector that represents—they use personas. You sort of fix some representation ahead of time that you want to not vary. Let's say that your data implies going in that direction. Let's go back to the pirate example, right? Your data implies that you should acquire a pirate persona in order to explain this data, because imagine the setup is something like you've got a GSMK math prompt, and then the model inexplicably starts talking like a pirate in its response. So in terms of what gradient descent will do—and we could sort of validate this with our jerry-rigged sieve—the model needs to spontaneously become more pirate-like. And I think this is the same sort of phenomenon that explains emergent misalignment. Now what positive preventative steering does is during the forward pass, it takes that persona direction, it turns it up more than it would fire. You sort of clamp the direction up in the forward pass. And the effect of this is to sort of neutralize learning in that direction if you set the amount right. And my mental model for this is like a thermostat. The amount of pirates in the data sets a sort of thermostat: we've got to be this piratical in order to explain this data. And positive preventative steering is just like holding a radiator next to the thermostat, next to the temperature monitor. It's like, okay, we're already piratical enough. And then you take this steering away, and then the model, during normal operation, will just not be a pirate. So you sort of explained away part of the data.
嗯。
Yeah.
而接种提示是尝试做同样的事情,但在文本空间而不是表示空间。这意味着,尝试把必要的信息放回去。所以再回到海盗的例子,如果你试图用接种提示来解释掉这个,你在提示中放入“你是一个海盗”,现在就没有什么需要解释了——你再次把暖气片放在恒温器旁边,模型就会想“我是一个海盗。我不需要解释数据中这个残余的异常。”这样就移除了学习压力,而不是试图把它压下去,否则它会被绕过。
And inoculation prompting is an attempt to do the same thing but in text space rather than in representation space. What that means is, try to sort of put back the information that's necessary. So to go back to the pirate example again, if you're trying to inoculation prompt in this or trying to explain this away, you put in the prompt like "you are a pirate" and now there's nothing to explain—you've put the radiator next to the thermostat again, and the model's like "I am a pirate. I don't need to explain this residual anomaly in the data." And that's removed the learning pressure rather than tried to squash it out, in which case it'll kind of get rooted around.
顺便说一下,在你的博客文章中,你想明确表示你仍然足够“无苦”和“入坑”。简而言之,Sutton 非常看重人类概念瓶颈,对吧?他不喜欢知识工程和把所有这些先验放入模型。这有点有趣的张力,不是吗?因为原则上,你在文章中说过,你所做的是重塑损失曲面,使最小阻力路径导致你想要的类型的结构出现。所以不完全是那样,但其中仍然有一点认知成分,因为我猜为了让它是有意的,你需要——我的意思是,基本上存在规范差距。你需要指定你想要什么。那么你如何处理这种张力?
We should say as well, by the way, in your blog post you wanted to make it clear that you are still sufficiently bitterless and pilled. So in short, Sutton was really big on human concept bottlenecks, right? He's not a fan of knowledge engineering and putting all of these priors into models. And it's a bit of an interesting tension, isn't it? Because in principle, you said in the article that what you're doing is reshaping the loss surface so that the path of least resistance will lead to the emergence of the types of structures that you want. So it's not quite that, but there is still a little bit of an epistemic component to it because I'm guessing for it to be intentional, you need to—I mean, there's a specification gap basically. You need to specify what you want. So how do you wrestle with that tension?
我的意思是,部分原因,我想,只是还有——这里有两件事。我认为与 Richard Sutton 在根本上存在分歧,关于奖励及其充分性,或者像从环境外部提供的简单标量奖励。但还有一个问题是我们是否应该使用、在多大程度上应该把人类工程化的概念放进去。我们可以稍后再回到奖励问题。但如果你回顾这次对话的进程,这个过程中实际上没有任何人类指定的东西。模型有它拥有的任何表示,SAPE 或接下来出现的任何东西都会拾取它拥有的任何东西,然后翻译层会经历这种自动化的可解释性过程,试图为事物分配标签。
I mean, part of it, I suppose, is just that there is also—so there's two things here. I think there's a disagreement at base with Richard Sutton about rewards and their sufficiency, or like simple scalar rewards that are provided externally from an environment. But then there's also the question of should we use, to what extent should we put human-engineered concepts in. So we can come back to the reward thing in a moment. But if you sort of wind back over the course of this conversation, there's actually nothing human-specified in this process. The model has whatever representations it has, the SAPE or whatever comes next picks up on whatever it has, and then the translation layer is going through this sort of automated interpretability process of trying to assign labels to things.
所以我们并没有真正尝试对输入做复杂的特征工程,把它们放进更好的格式。这里面的一切其实都是梯度下降发现的结果。我们只是在努力把它塑造成更好的样子。我觉得这可能就是我和某些人的根本分歧所在:我认为正确指定奖励非常困难。我们基本上一直在看到这一点。我们在为训练指定奖励时遇到困难,无法得到我们想要的模型。原则上,在某种超级天才的思维方式下,奖励可能就够了,但今天,奖励显然不足以给我们想要的模型。所以这也许就是根本的分歧:我认为我们确实需要在训练过程的某个地方加入一层人类价值观。
So it's not like we've actually tried sophisticated feature engineering on the inputs to put them in a better format. Everything inside this is actually discovered as a result of gradient descent. We're just trying to shape that better. And I think this is where the base disagreement with certain people might come in, where I think it is very hard to specify rewards correctly. We're basically just seeing this continuously. We're having trouble specifying our rewards for training in a way that gives us the models we want. In principle, in some super galaxy brain way, reward might be enough, but today reward is clearly not enough to give us the models that we want. So that's perhaps the underlying disagreement: I think we actually do need to put in some layer of human values into the training process somewhere.
是的。你的观点很有道理,因为这和你之前说的模型知道事物是一致的。所以我们可以在模型中指向那些概念,但“有意的”这个词,我猜,确实意味着这是我们的意图。所以我们在训练过程中选择其中一些概念并加以强化。顺便说一句,我觉得这是个很棒的想法。不知道你熟不熟悉一个叫“机器教学”的概念。它出自微软研究院,是一个叫 Patrice Simard 的人提出的,本质上是一种黑盒方法,你可以有一个交互式的、有意的过程:当模型做错事时,你可以指出个别问题,而你所做的本质上是一种主动的数据集蒸馏。所以这是个很棒的想法,而且实际上还有你关于预测性数据调试的工作。我们也谈到了那个,但对我来说,有某种主动的、有意的过程来引导似乎是合乎逻辑的。
Yeah. And your point is well taken, because this is very consistent with what you've said that the model knows things. So we can point to those concepts in the model, but the word intentional, I'm guessing, does mean that it's our intention. So we are selecting some of those concepts and we're leaning into them during the training process. And I think it's a beautiful idea, by the way. I'm not sure if you're familiar with a concept called machine teaching. This came out of Microsoft Research. There was a guy called Patrice Simard, and this was a blackbox method essentially, where you could have this interactive intentional process where the model does something wrong and then you can point out individual problems, and what you're basically doing is a form of active dataset distillation. So it's a beautiful idea, and there's actually your work on predictive data debugging. We talk about that as well, but it seems logical to me to have some kind of an active intentional process to guide...
来引导我们如何训练这些模型。
to guide how we train these models.
是的。
Yeah.
是的,我对机器教学不太熟悉。我记得看到过这个名字,觉得听起来很酷,然后就从脑子里消失了。所以谢谢你提醒我。而且我觉得你也可以想象,回到那个“苦药丸”的比喻,我们正在尝试的一件事是把更多算力投入到学习过程中。梯度下降给你什么就是什么。梯度下降很棒,但如果你能花更多算力得到更好的梯度,一个更干净、更符合你想要的梯度,那就太好了。很久以前,当我刚开始接触安全和对齐工作时,我曾经认为这是不可能的。问题似乎是,要有梯度下降,但只用于好的事情。然后我想,我们也许已经找到了一种方法,让梯度下降只用于好的事情。
Yeah, I am not very familiar with machine teaching. I remember seeing the name and thinking that sounds cool, and then it's all gone from my brain. So thank you for reminding me. And I think that you can also imagine, going back to being bitter less pill here, one thing that we're trying to do is put more compute into the learning process. Gradient descent just gives you what it gives you. Gradient descent is great, but it would be great if you could spend more compute to get a better gradient, a gradient that's both cleaner and more aligned with what you want. When I was first getting into safety and alignment work quite a while back, I used to think this is impossible. The problem seems to be like gradient descent but only for good things. And then I guess that actually we've kind of perhaps got around to a way of having gradient descent but only for good things.
你能谈谈做这件事的一些具体算法方法吗?我的意思是,这可能会很自然地引出“特征作为奖励”的工作。
And can you talk through some specific algorithmic approaches for doing this? I mean, it might be a natural lead on to the features as rewards work.
所以我们目前所做的,特征作为奖励的工作就是一个例子,说明如何至少在某些情况下,用表征作为训练信号,并且对我们之前讨论的所有这些问题具有鲁棒性。还有预测性数据调试的工作,而且我还想谈谈,我们刚才花了一点时间讨论接种提示和积极预防性引导。我认为这些方法很有前景,主要问题在于它们不是自适应的。如果你还记得描述,我们提前固定了我们的角色向量,说不要朝这个方向走。我非常担心训练过程中的未知的未知。所以我认为它们需要某种自适应。而这可能的样子正是这种梯度读出,然后还有一些,看看那些临时拼凑的 SA,看看那些海盗。嗯,这种临时拼凑的 SA 有点像在给我们一份梯度下降提供的东西的菜单。然后我们需要能够从中智能地选择。所以我心中的核心梦想是拥有非常好的梯度可解释性。然后,比如说,我们有我们的模型规范或我们的章程,或者对这个例子的一些人类反馈。我们可以看到,我们这边菜单上有这些东西,这些是事情自然发展的方向。然后我们有所有这些信息,给我们一些关于我们应该走的方向的信息。然后我们让它——我说我们,我的意思是语言模型——看看这个信息,看看那个信息,然后说好的,我们需要做以下干预来让我们走上正确的方向。我认为这些技术部分基本上都具备了,问题在于它们是否足够高质量来可靠地做到这一点。
So the things that we have done so far, the features as rewards work is this sort of example of how you can, at least in some instances, use representations as a training signal in a way that's robust to all of these issues that we were talking about earlier. There's the predictive data debugging work, and I think I also want to say a bit about, we just spent a little while talking about inoculation prompting and positive preventative steering. I think that these methods have a lot of promise, and the primary issue is that they're not adaptive. If you remember the description, we sort of fixed our persona vector ahead of time, saying don't go in this direction. I worry a lot about unknown unknowns in the training process. So I think that they need to be kind of adaptive. And what this might look like is exactly this kind of gradient readout and then some, looking at the sort of jerry-rigged SA, looking at the pirates. Well, this sort of jerry-rigged SA is kind of giving us a menu of things that gradient descent is offering us. And then we need to be able to intelligently choose from that. So the central dream that I have in my mind here is having really good gradient interpretability. And then having, say, our model spec or our constitution or some human feedback on this example. And we can see that we've got these things on the menu over here, like these are the natural directions that things are going to go in. And then we've got all this information which is giving us some information about the direction we should go. And we make it—and I say we, by we, I mean a language model—looks at this information, looks at that information, and says okay, we need to make the following interventions to get us in the right direction. I think that the technical pieces of this are basically all there, and it's a matter of them being high enough quality to do this reliably.
是的。是的,我的意思是,在特征作为奖励的工作中,你谈到很多任务是非常开放式的,你的意思是它们验证起来极其昂贵。是的。所以你可以,例如,用 LLM 作为评判者,但显然,用它作为奖励信号会非常昂贵。在那项特定的工作中,你关注的是幻觉和最小化幻觉,这是另一个很好的例子,有时当模型产生幻觉时,模型实际上知道它在幻觉,但它还是决定这么做。
Yeah. Yeah, I mean on the features as rewards work, you were talking about a lot of tasks are quite open-ended, and what you meant by that was they were extremely expensive to verify. Yes. So you could, for example, use an LLM as a judge, but obviously that would be very expensive to use as a reward signal. And in that particular work, you were looking at hallucinations and minimizing hallucinations, and this was another great example where sometimes when the model hallucinates, the model actually knows that it's hallucinating, but it decided to do it anyway.
是的。所以这里是开放式的。我应该说,我很幸运能领导那个团队,但几乎所有的功劳都要归功于其他人。好吧,那篇论文的所有功劳都要归功于其他人。我只是在这里谈论它。他们做了真正的工作。这里的想法是,你可以在事实核查方案中使用语言模型作为评分者,但这并不是特别准确。如果你用同一个模型来事实核查,你会得到一些正确的结果。
Yeah. So open-ended here. I should say that I was fortunate to lead the team that was working on that, but really almost all the credit has to go to everyone else. Well, all of the credit has to go to everyone else on that paper. I'm just here talking about it. They did the real work. The idea here is that you could use a language model as a grader in your fact-checking scheme, but this is not particularly accurate. If you're using the same model to fact check, you'll get some things right.
我们在论文中做了这个消融实验。它会有一点提升,原因我们稍后可以讨论,但效果不太好,因为模型基本上只会说:“嗯,不错,一切正常。”你可以用更强的模型,但那样会变得又慢又贵,而且那个模型仍然有自己的知识盲区。或者你可以用更强的模型加网络搜索,但那样会非常耗时。所以我们这里的想法是,我们可以把这一过程摊销化。你可以用这个模型加网络搜索收集一个大型数据集,或者一般来说,这种被增强的模型可以出去收集一个数据集,记录增强模型会做什么。也就是模型加网络搜索工具,然后把结果摊销回一个探针里。现在我们得到的东西运行起来极其便宜、快速,所以它可以成为强化学习循环的核心。
We do this ablation in the paper. It'll uplift a little bit, for reasons we could talk about in a second, but it doesn't do very well because the model basically just goes, "Yeah, that's cool. Everything's fine." You can use a more powerful model, and now things are really starting to get slow and expensive, and that model still has its own knowledge gaps. Or you can use a more powerful model and web search, and now things really take a long time. So the idea that we had here was we can sort of amortize this process. You know, you can collect a large dataset using this model plus web search, or in general this sort of amplified model can go out, and we can collect a dataset of what the amplified model would do. That's the model plus the web search tool, and kind of amortize that back into a probe. And now we have something that's extremely cheap to run and fast to run. So it can be like the core of an RL loop.
关于生成与判别这个问题,是不是很神奇?一个模型在某种情境下会幻觉并生成错误内容,但如果你问另一个模型,它没有先入为主,没有被引导去判别,它可以是同一模型家族或同一个模型,它确实知道答案。我的意思是,你最好的直觉是什么?因为我觉得你对此有一些想法。
On this generation versus discrimination thing, isn't that fascinating? That a model in one context could hallucinate and generate the wrong thing, yet if you ask another model which has a blank slate and hasn't been primed to discriminate, it can be the same model family or the same model. It does know the answer. I mean, what is your best intuition? Because I think you had something in there.
是的,也许是置信度偏差或流畅性之类的。但模型可能出错的原因实在太多了。
Yeah, maybe it was like a confidence bias or fluency or something like that. But there are just so many reasons why it might do the wrong thing.
是的,可能是很多种原因。
Yes, it can be any number of things.
而且即使是同一个模型,有时它也能识别出来。就是那个刚刚产生幻觉的模型,如果你问它,它会说:“哦,那是幻觉。”而且这可能实际上是一个非常难以监督的事情。如果你试图监督,比如试图把它纳入训练监督,那是困难且昂贵的。对此一个有趣的机制性假设与模型内部操作的顺序有关。我们在算术中看到过,有时事情必须按特定顺序发生,比如第 9 层必须出现在第 10 层之前,等等。而且你有不同的模块,比如算术的检查操作有时早于生成操作。很可能幻觉和事实核查也是如此。所以可能是生成步骤使用了整个模型,但检查发生在模型更早的部分。因此,当你把错误事实输入模型时,它会说:“哦,对,那是幻觉,”但那时它已经说出来了,为时已晚。所以有一种行为是模型本可以做到,但在之前的训练中还没有得到充分强化。这正是针对幻觉的基于人类反馈的强化学习(RLHF)理念所利用的:每当模型根据自身表征能够知道那是幻觉时,它其实确实知道,而我们就在那里真正塑造它的行为。
And even if it's like the same model, it will sometimes be able to pick it up. Literally the same model that just hallucinated will be like, "Oh, if you ask it, it'll be like, oh, that is a hallucination." And it might be that this is actually a very hard thing to supervise. If you try and supervise, like if you try to put this into training supervision, it's a hard and expensive thing. An interesting kind of mechanistic hypothesis for this is to do with the ordering of operations inside the model. So we've seen this in arithmetic, that sometimes things have to happen in certain orders, like layer 9 has to occur before layer 10, and so on. And you have different modules that, if they like, sometimes the checking operation for arithmetic, for instance, is earlier than the generating operation. And it's quite possible this is also true for hallucination and fact-checking. So it might be that the generation step takes the whole model, but the checking happens earlier in the model. So then when you put the incorrect fact through the model, it's like, "Oh yeah, that is a hallucination," but at that point it's already said it, it's too late. So there's a behavior that the model could be doing but hasn't been sufficiently reinforced in its training up to that point. And that's what the idea of this kind of RLHF for hallucinations taps into: whenever the model could know according to its own representations that it was a hallucination, it in fact does know, and we really shape its behavior there.
第三个可能性,我觉得有点有趣,与角色或上下文学习有关:能够编造内容实际上对模型来说是一种有用的能力。如果我让它写一个故事,如果我想让它为我生成一个虚构世界,如果一切都符合事实,那它其实不是一个好的虚构世界。所以能够编造东西对模型来说是一种有用的能力。有时它必须在上下文中判断出我们正在做的是这件事。所以如果你从一种粗略的贝叶斯观点来看,如果我是模型,我刚开始对话,我不太确定我们在做什么任务。我们是在编造吗?我们是在说真实的事情吗?而我说的一切和用户说的一切都是某种程度上的证据。然后如果我说了错误的话,我就会把这当作证据,哦,我们在编造。好,继续吧。事实上,我们表明,仅仅进行这些上下文干预就已经足以减少下游幻觉。所以可能我们只是通过不让第一个幻觉出现来让模型变得非常自信,这让模型确信,哦不,我们今天在说真实的事实,我们不是在编造。
A third possibility, which I think is a bit funny, is to do with this idea of personas or in-context learning: being able to make things up is actually a useful capability for a model. If I ask it to write a story, if I wanted it to generate a fictional world for me, it's actually not a very good fictional world if everything is factually true. So being able to make stuff up is a useful capability for a model. And sometimes it has to figure out in context that this is what we're doing. So if you imagine from a vaguely Bayesian point of view, if I'm the model and I start the conversation, I'm not quite sure what task we're doing. Are we making stuff up? Are we saying factually true things? And everything that I say and everything the user says is some amount of evidence one way or the other. And then if I say something incorrect, then I'm now taking this as evidence that, oh, we're making things up. Cool, let's carry on. And in fact, we show that just doing these in-context interventions is already enough to reduce downstream hallucinations. So it might be that we're just making the model really confident by never letting the first hallucination in, and that allows the model to become confident that, oh no, we're playing true facts today, we're not making things up.
这是利用这种有意设计的一个绝佳例子。所以,你知道,一个例子是,是的,也许它应该在生成之前检查,对吧?我认为这是一个绝佳的例子。
That's a beautiful example of using this intentional design. So, you know, one example is, yeah, maybe it should check before it generates, right? I think that's a beautiful example.
我们应该谈谈预测性数据调试。所以,我在脑海中将其概念化的方式几乎是一种主动数据集蒸馏,对吧?本质上,机器学习模型存在一个问题,它们会学习虚假的相关性。它们会学习做虚假的事情,就像你刚才说的,也许它们变得过度自信或谄媚之类的。那么,如果我们能在训练过程中让模型推理数据,从而不传入那些无论出于何种原因会有害的数据,那不是很酷吗?所以预测性数据调试背后的想法是,通过模型的眼睛来看数据,我们想知道在逐个示例的基础上它们会如何影响模型,以及数据集在总体上会如何影响模型。有时你从阅读数据中了解到的东西是显而易见的。但当数据量巨大时,你并不清楚数据中到底有什么。你会想,我不知道里面有什么,你无法全部检查。也许你可以用 LLM 扫描一遍,但问题是,模型从数据中学到什么的过程有时对你来说是直观的。比如海盗的例子,模型应该学会当海盗,这很直观。但有时却非常不直观,比如涌现性错位。对大多数人来说,那是一个非常不直观的发现。我认为 Orion 实际上做了一个预注册的研究,他问人们会觉得这有多令人惊讶,很多人说,我不认为那会是真的。所以我可以肯定地告诉你,这是一个令人惊讶的事实。所以你可以用语言模型作为数据集上的自动评分器来捕捉容易发现的问题,但你无法捕捉到那些意想不到的副作用。
We should talk about the predictive data debugging stuff. So, the way I kind of conceptualize this in my mind is almost a form of active dataset distillation, right? So essentially we have this problem in machine learning models that they learn spurious correlations. They learn to do spurious things, as well as you were just saying, maybe they're becoming overconfident or sycophantic or something like that. So wouldn't it be cool if we could use the model to reason about the data during the training process, so we could actually not pass in data which is going to be harmful for whatever reason? So the idea behind predictive data debugging is to kind of look at the data through the model's eyes, and we want to know on an example-by-example basis how they would affect the model, and also how the dataset would affect the model in aggregate. And sometimes the things that you learn are kind of obvious from reading the data. It's just not clear what in fact is in your data when you have enormous quantities of it. You're like, I don't know what's in there, you can't check it all. Maybe you could run an LLM over it, but then the problem is that the process of what a model learns from data will sometimes be intuitive to you. Like the pirate example, it's quite intuitive that the model should learn to be a pirate. But sometimes it's deeply unintuitive, like emergent misalignment. That was a deeply unintuitive finding to most people. I think Orion actually kind of did a pre-registered thing where he asked people how surprising they would find it, and lots of people were like, I don't think that would be true. So I can tell you for sure that it is a surprising fact. And so you can catch the easy stuff with a language model kind of autograder over the dataset, but you won't catch the unexpected side effects.
所以,如果你要在数据集上运行语言模型,你基本上可以几乎免费地挂上一个稀疏自编码器,让它在处理数据时同步运行。实际上,这可能净成本更低,因为你不需要为每个样本生成 token。你只是处于预填充阶段,把大量数据推过去,然后问:“你看到了什么?”这会告诉你这个数据集在模型眼中是什么样的,我认为这是整理数据的更好方式。我们在论文中实际利用这一点的方式是处理 DPO 数据,所以有一对正样本和负样本。Ege 告诉我他知道如何将其扩展到 SFT,我相信他,但我记不清细节了。所以有一对正负样本,正样本包含对提示词的好回答,负样本包含坏回答。我们可以查看 SAE 中隐藏表示的特征之间的差异,这能很好地近似这个数据点将如何推动模型,或者我们也可以基于特征进行聚类。这比基于嵌入聚类要好得多,因为嵌入包含各种你可能不关心的东西,比如下一个 token 是否该有逗号。我们关心的是语义内容,很多时候不关心低层处理细节。所以基于特征而非原始嵌入来做,能让我们更好地接触到真正关心的东西,可以把它分离出来。这就是我为什么这么做,而不是采用其他可能首先想到的方法。
And so, you know, if you're going to run a language model over the data set, you can also essentially for close to for free attach something like a sparse autoencoder to it as it runs over the data set. And in fact, this is probably like on net cheaper because you're not asking it to generate tokens for each example. You're just, you know, you're just like you're just in the prefill regime. You're just pushing loads of data through and saying, "Well, what do you see?" So this will tell you, this should tell you like how this data set is perceived through the model's eyes, and that I think is just a better way of curating your data. The way we actually exploit this in the paper is by dealing with DPO data. So there's a positive and then a negative pair. And you know, Ege tells me that he knows how to extend this to SFT, and I believe him. I can't remember the details. So there's a positive and a negative pair where the positive thing contains a good response to the prompt and the negative contains a bad response to the prompt. And so we can sort of look at the delta between features in the hidden representation in the SAE. And this is a good approximation to the way that this data point will push the model, or we can also cluster based on features. And this is much better as a way of understanding—you don't necessarily want to cluster based on embeddings, because embeddings contain all sorts of things that you don't necessarily care about, like should I have a comma in the next token. We care about the semantic stuff; we don't care about the low-level processing stuff a lot of the time. So doing this based on features rather than the raw embeddings gives you much better access to the stuff we actually care about; we can kind of separate that out. So that's the intuition as to why I do this rather than the other approaches that might come to mind.
嗯,我们应该逐渐过渡到几何相关的内容,但在那之前,从概念上讲,我对模块化这个概念非常感兴趣。很长一段时间里,联结主义者都认为模型缺乏结构是一个特性而非缺陷,也许当时我们并不知道存在结构。我认为很多联结主义者同时也是神经科学家,他们倾向于认为大脑是更平坦的——甚至有人写了本同名书,我还采访过他。而另一派观点认为大脑是高度模块化的,正如你在研究中看到的,神经网络也是高度模块化的。那么,你认为原则上模块化是好事吗?它是自然出现的吗?
Um, we should gradually move over to the geometry stuff, but I mean just conceptually before we go there, I'm really interested in this concept of modularity. So for a very long time connectionists were arguing that it was a feature not a bug that there wasn't much structure in the models, and perhaps back then we didn't know that there was structure. And I think a lot of connectionists who are also neuroscientists kind of imagined that the brain was flatter—even wrote a book by that name—and I interviewed him. And there is another school of thought that the brain is highly modular, and as you're seeing in your research, neural networks are highly modular. So, do you think in principle that modularity is a good thing? Is it a natural thing?
是的。稍微展开一下,我认为历史上很多早期的联结主义者——或者这取决于你想从哪里开始——有大量有效的工作,如果现在做,可能会被称为可解释性。比如看那篇关于通过反向传播误差信号学习表示的经典论文,实际上里面大部分图都是他们在说:“看,我们的反向传播过程让模型学到了合理的表示。”他们通过展示其可解释性来验证。我认为模块化是你想要达到的终点,但你不是从模块化开始的。这可能是反复出现的现象:为什么要过度参数化并保留所有这些连接?因为这使得学习过程更容易,但你最终得到的东西实际上是非常模块化的。而且,说得模糊一点,你可以把学习过程看作网络对自己变得可读。你知道,我在这里有一些关于某事的表示,在那里有一些关于某事的表示。如果这种表示容易寻址,学习起来就容易得多。你可以说:“啊,这就是某某计算存储的地方。”要达到那一步,你必须形成这些计算,我认为高度过度参数化且没有强先验有助于达到那里,但模块化是最终目的地。
Yes. To expand on that a little bit, I guess historically a lot of the early connectionists—or maybe this depends on where you want to start—but there was a surprising amount of things that worked that, if done now, might be called interpretability. Like if you look at the initial paper on learning representations by back-propagating error signals, the classic backprop paper, actually most of the figures in that are them saying, "Look, the model learned sensible representations from our backprop procedure," and they validated it by showing that it's interpretable. I think that modularity is the endpoint you want, but you don't start with modularity. This is maybe a repeated thing: why overparameterize something and have all of these connections? Because it makes the learning process easier, but the thing you end up getting to is actually very modular. And I suppose to be very vague, you might think of the learning process as the network becoming legible to itself. You know, I've got some representations here about something, I've got some representations there about something. It's much easier to learn if this representation is kind of easily addressable. You know, I can say, "Ah, this is where the such-and-such computation is stored." Now, to get there, you have to form these computations, and I think it's very helpful to be heavily overparameterized and have no strong prior to get there, but I think modularity is the destination.
嗯,我倾向于同意,我的部分直觉是,你知道,很多怀疑者说:“哦,你不可能记住无限。”我是说,那是 Gary Marcus 会说的话,在某种程度上他是对的。而这些网络,它们有这些结构,这些抽象结构,正是这些结构让你不需要记住无限。
Well, I'm inclined to agree, and part of my intuition is, you know, a lot of skeptics said, "Oh, you can't memorize infinity." I mean, that's the kind of thing that Gary Marcus would have said, and in a way he's right. And these networks, they have these structures, these abstract structures, and they are what allow you to not need to memorize infinity.
对,它们让你能够泛化,并在许多未见过的情境中工作。
Right, they allow you to generalize and work in many, many unseen situations.
你的工作真的让我着迷,因为你似乎在描述网络进化成一台计算机。
And your work really fascinates me because you're kind of describing the network evolving into a computer.
所以它是一个有执行计算的部分、有类似记忆系统的部分的实体,这些结构在不同的模型家族中出现,并且看起来非常相似。也许它们只是架构的副产品之类的,但这种现象确实很有趣。而且我认为另一个方面是它是逐渐发生的。
So it's something that has parts that do computation, parts that resemble something like a memory system, and these structures emerge in different model families and look very, very similar. And maybe they're just artifacts of the architecture or something like that, but it really is interesting that this is happening. And then I suppose another aspect is it's happening gradually.
是的。因为我不知道你是否会——我不知道你对此的直觉是什么,但有时我们可能会把它描述为顿悟(grokking),
Yeah. Because I don't know whether you would—I don't know what your intuition is on this, but sometimes maybe we might describe it as grokking,
但这并不完全正确,不是吗?因为这些结构是随着时间逐渐结晶的。
But that's not entirely true, is it? Because these structures kind of crystallize over time.
是的,时间尺度非常有趣。我认为没有人能明确解决这个问题。最近有一篇关于预训练期间或整个训练过程中人格形成的论文,我记不清作者了——我想智能体会找到它的——而且它们出现得出奇地早。Eric Michaud 在这方面有一些非常好的工作,无论是概念上还是实证上。不是人格形成,而是关于学习如何进行的概念。他称之为“量子”。如果我的总结可能不准确,他可以纠正我。你可以把通用网络(比如语言模型)的学习过程描述为万亿次微型顿悟。
Yeah, the time scale is very interesting. I don't think anyone has definitively settled this. There was an interesting paper recently on persona formation during pre-training, or across the training process, and I can't remember the author—I guess the agent will have to find it—and they emerged surprisingly early. And Eric Michaud has some very nice, really nice work on this, both conceptually and empirically. Not persona formation, but the idea of how learning is proceeding. He calls it quanta. And if I might perhaps inaccurately summarize it, he can tell me off. You might describe the learning process of a general network, you know, a language model, as like a trillion micro-groks.
嗯,你知道,就是叠加。如果你不断放大、再放大、再放大,并且在正确的任务分解层面去看,你可能会看到一个微型的、类似微缩的顿悟(grok),然后它又顿悟另一件事。我们只是有所有这些微小的 S 形曲线堆叠在一起,在对数坐标上形成一条直线。所以从这个角度看,连学习过程本身可能都是模块化的。我们只是不确定。有一个问题,我觉得这里真正开放的问题其实是程度问题,而不是是否会发生的问题。
Um, that all you know, you just stack. If you've zoomed in and zoomed in and zoomed in and looked at the right level of task decomposition, you might just see a mini, like a micro grok, and then it groks another thing. And we just sort of have all these tiny sigmoids that are stacked on top of each other to form a straight line on a log plot. And so from that perspective, even the learning process may in fact be modular. We just don't know for sure. There's a question, I think it's a question probably the open question here is one of degree, not of whether it happens at all.
那么,跟我讲讲这个神经几何(neurogeometry)吧。你研究了好几个不同的模型家族,顺便说一句,其中有一些绝对漂亮的图。大家应该去看看 Goodfire 的博客文章,非常棒。也许我们应该从你是如何生成这些图的开始讲起。如果我没理解错的话,你知道,像星期几、月份、年龄等等这些不同的东西,你实际上把它们表示成了一种几何结构。我想你的做法大概是,先做某种降维,然后拟合一些样条曲线之类的。但它所展示的是,模型表征世界上的许多概念的方式是高度结构化的。
Well, tell me about this neurogeometry stuff. So you've studied several different model families, and there are some absolutely beautiful plots, by the way. So folks should look at the blog post from Goodfire. Amazing stuff. Maybe we should just start with how you've generated those plots. So, if I understand correctly, you know, things like days of the week and months of the year and age and all these different things. You've actually represented them as a kind of geometry. And I think the way you did that was something like, you know, you do some dimensionality reduction and then you fit some splines or something like that. But what it's showing is that the way the models represent many concepts out there in the world is highly structured.
是的,没错。我要说的是,我们是建立在已有的工作基础上的。比如,那篇《并非所有语言模型特征都是线性的》论文,是真正在可解释性领域开启这一方向的文章之一。在神经科学中,也有很长历史的所谓“群体几何”(population geometry)。所以再说一次,如果我们多读点书,可能早就走到这一步了。但我不想说神经几何是我们首创、别人没做过。我们是建立在早期工作的基础上,但发现这些东西的方法和最新水平在过去几个月里已经进步了很多。最早的做法是从你认为应该有结构的概念开始,比如星期几,然后把对应的数据放进去,做 PCA,然后你看到的就是周一、周二、周三、周四、周五、周六、周日。所以这完全是监督式的,但它足以证明这种非线性结构的存在。在离开这个话题之前,我们应该深入探讨一下“线性”这个词的细微差别,因为这里面有很多微妙之处。但我所说的“非线性”是指,表征并不形成——那些我们直觉上归为一类的东西并不形成一条线或一个平面,或者,嗯,实际上就是一条线。所以这足以证明它的存在。然后问题是,每当你有一种监督方法,你常常想尝试找到一种无监督的方式来做同样的事情,因为那能让你回答的问题不仅是它是否存在,还有哪些我们可能没想到的东西,以及有多少。所以我们做的第一件事实际上是给这些数据拟合一个稀疏自编码器,这看起来可能真的很古怪,因为我们问的是有多少不是线性的结构,而 SAE 的核心归纳偏置就是事物位于直线上。是的,一切都是从原点出发的射线,或者说正射线。所以这看起来可能真的很奇怪,但我会解释为什么它有意义。想法是这样的,假设我有一个特征,它就在一条弧线上。我应该把它移下来,这样我就不会超出屏幕。我坐在原点这里,看着这组激活值,你可以把它想象成在看星星,有一道弧形的星星。现在我的 SAE 的一个特征会指向弧上的某个点,另一个 SAE 特征会指向弧上的另一个点,依此类推。你应该意识到的是,这实际上会在特征的共激活中引发相当强的模式。如果我在弧上有两个靠得很近的特征,它们可能会共激活。而如果两个特征离得很远,它们基本上永远不会共激活。如果我说我有,比如,星期几——我们举个连续的例子,比如颜色从红到蓝——如果某物是蓝色的,那它就不是红色的,所以我的指向蓝色的 SAE 特征与指向红色的 SAE 特征的激活是强烈反相关的,并且与基本上所有的背景都基本不相关。所以这种邻近正相关、长程反相关的模式,就足以让你实际去拟合——我们给它拟合了一个伊辛模型,当团队带着这个结果回来时,我相当惊讶。我当时想,酷。为什么这是一个好模型,因为你可以同时有正的和负的耦合强度,拟合这个模型让我们能够通过数据拟合一条样条曲线。所以那是我们的伊辛管道,我们的第一个无监督结构发现工具。然后我们有一些由 Tom Rafel 领导的非常好的工作,我认为一些最漂亮的流形就来自这项工作。这里的想法是我们训练所谓的块稀疏特征化器。想法是,SAE 给你一条线,我们就说,如果它是更高维度呢?所以这在概念上相当简单,但诀窍在于让它真正起作用,以及不预先固定维度,因为你不想必须输入一些信息,比如,我认为在这个表征中有 7000 个二维特征、400 个三维特征和 5 个五维特征。这只是一组愚蠢的超参数需要指定。所以你需要能够自适应地学习这些子空间的大小,而这正是让它真正起作用的关键。而自适应地学习这些子空间的大小是块稀疏特征化器的关键特征。
Yes, that's right. And I should say that we are building on a body of work. For instance, the paper 'Not all language model features are linear' was one of the papers that really kicked this off in interpretability. There's also a long history in neuroscience of this kind of population geometry, as they call it. So again, if we'd read more books, we might have gotten here sooner. But I don't want to say that we have done neural geometry and no one else has. We're building on this earlier body of work, but the idea and the state-of-the-art for how to discover this stuff has moved quite a lot in the last few months. The earliest thing to do was to start with concepts you think should have structure, like days of the week, and just put in data corresponding to these and project it out. You do a PCA, I think, and then you see it's like Monday, Tuesday, Wednesday, Thursday, Friday, Saturday, Sunday. So that's totally supervised, but it suffices to show that this nonlinear structure exists. And we should get into some nuances on the word 'linear' before we move off this topic, because there's a lot of subtlety there. But I'm going to say 'nonlinear' in the sense that the representations don't form, like, the things which are intuitively grouped to us don't form a line or a plane, or well, just a line really. And so this was enough to show that this exists. And then the question is, whenever you have a supervised method, you often want to try and find an unsupervised way of doing the same thing, because that lets you answer the question not only does it exist, but what else is there that we might not have expected, and how much is there? And so the first thing that we did was actually fit a sparse autoencoder to this data, which might seem like a really wacky thing to do, because what we're asking is how much structure which is not in the form of a line is there, and the core inductive bias of the SAE is that things lie on lines. Yeah, everything is a ray out from the origin, or a positive ray. And so that might seem like a really weird thing to do, but I'll say why it makes sense. The idea is that, let's say for the sake of argument I have a feature and it just lies on an arc. I should move it down here so that I'm not going off the screen. I'm sitting here at the origin and I'm looking at the set of activations, and you can think of it as like watching the stars, and there's an arc of stars. Now my SAE one feature will point out through some point in that arc, and another SAE feature will point out through another point in that arc, and so on. And the thing you should realize is that this will actually induce quite strong patterns in the co-activations of features. If I have two features that are close together on the arc, they will probably co-activate. Whereas if I have two features which are far away, they will essentially never co-activate. If I say that I have, like, a days of the week—let's give a continuous example, let's say that it's like color red to blue—if something is blue, it is not red, so my blue SAE feature that is going through blue is strongly anti-correlated with the activation of the SAE feature that is going through red, and essentially uncorrelated with basically all of the background. And so this pattern of nearby positive correlation, long-range anti-correlation, is enough structure for you to actually fit—we fit an Ising model to it, which was rather a surprise to me when the team came back with that. I was like, cool. And the reason this is a good model is that you can have both positive and negative coupling strengths, and fitting this allows us to fit a spline through the data. So that was our Ising pipeline, our first unsupervised structure discovery tool. And then we've got some really nice work led by Tom Rafel, which is I think where some of the most beautiful manifolds come from. The idea here is that we train what we call block sparse featurizers. The idea is that an SAE gives you a line, and we just say, what if it was a higher dimension? And so this is conceptually pretty simple, but the tricks are in making it actually work, and in not fixing the dimensionalities ahead of time, because you don't want to have to put in some information like, I think in this representation there are 7,000 two-dimensional features, 400 three-dimensional features, and five five-dimensional features. This is just a stupid set of hyperparameters to specify. So you need to be able to adaptively learn the size of these subspaces, and that's sort of the making this work at all. And adaptively learning the size of these subspaces are kind of the key features of the block sparse featurizer.
我们刚才在聊 mountain car 的例子。对。如果我们用位置和动量来表示它,并且用一个图像动作模型,你基本上可以在激活空间里,做 PCA 的时候,看到它看起来就像一条弦。你可以直接干预那些激活。所以你可以把车移到弦上的另一个位置,然后你看,它就移动了。但真正重要的概念是,这是一个流形。所以,就像你之前说的,流形代表了这件事物的意义,对吧?如果你把它当作欧几里得空间,直接在两个点之间插值,然后偏离了那条弦,从表征的角度看,你就进入了无人区。那么图像模型就会输出乱码。我认为这非常重要,因为这里有几件事。首先,你说这些 SAE 所做的,如果流形不是线性的,它们可能会打碎这个流形。所以如果这个流形有结构,而你可能会取对比样本之类的,然后把它们混合在一起,当流形有结构时这样做是没有意义的。
So it was talking about a mountain car. Yes. So what if we represented it, I think, with a position and a momentum, and we used an image-action model. You can basically just see in the activation space, when you do PCA, that it looks like a string essentially, and you can do intervening right on those activations. So you can move the car to a different location on the string, and lo and behold, you've now moved it around. But the really important concept is that this is a manifold. So, as you were saying before, the manifold represents the meaning of this particular thing, right? And if you treated it as a Euclidean space and you just interpolated between two points and went off the string, you're now in no man's land from a representations point of view. So now the image model is just going to be garbled. And I think this is a really important thing because there are a couple of things. First of all, you're saying that these SAEs, what they do is potentially they fracture this manifold if it's not linear. So if this manifold has structure, and you might be taking contrastive samples or something and mixing them together, it doesn't make sense to do so when there is structure in this manifold.
正是如此。就像你说的,当你试图从一个点走到另一个点时,你基本上是踏入了这个虚空,而网络并不真正知道如何处理。然后它就会出问题。我认为这实际上解释了很多关于引导(steering)的发现。引导就是对激活进行干预。我们做了很多引导,其他一些人也做了很多引导,引导神经网络的一个常见发现是,有时它很有效,效果惊人,你会得到 Golden Gate Claude 之类的。但有时它完全不稳定,网络做了你想要的事情,但也变得有点疯狂,或者立刻变成胡言乱语。我认为这基本上解释了那个现象,因为你偏离了流形。
Exactly. Because exactly like you say, when you try and go from one point to another, you're sort of stepping out into this void which the network doesn't really know how to handle. And then it sort of breaks. And I think this actually explains a lot of findings about steering. Steering being just intervening on activations. We do a lot of steering, some other people do a lot of steering, and one common finding with steering neural networks is that sometimes it works and it's amazing, and you get golden gate Claude or whatever. And sometimes it's just completely janky, and the network does kind of the thing you want but also just goes a bit crazy, or just turns immediately into gibberish. And I think this basically explains that phenomenon, because you're stepping off the manifold.
是的,完全正确。而且你们实际上有一篇非常有趣的论文。题目是《SAE 能捕捉概念流形吗?》,你们在那项研究中探讨的问题之一,基本上就是 SAE 捕捉流形意味着什么?那么你们在这方面做了哪些工作?
Yeah, exactly. And there was a really interesting paper actually from you guys. So it was 'Do SAEs capture concept manifolds?' and one of the things that you were studying in there was basically what does it mean for an SAE to capture the manifold? So what work have you done on that?
这就是平铺(tiling)的概念。我得说,神经科学中也有大量相关工作。再说一次,我们本该多读点书。更广泛的社区里也有一些工作。捕捉流形意味着什么,这个想法就是:你在多大程度上高效地表示那个流形,以及它在多大程度上符合其内在几何。所以如果我们回到这个弧线的例子,只要有足够多的点,足够多的线,我就可以说我已经捕捉到了流形。对于这个流形上的任何一点,我都有一个 SAE 特征,可以说它激活了多少多少,而且我在重建的意义上相对准确地捕捉了这个流形。但我实际上并没有学到任何关于更广泛的流形结构的东西。当我透过这个镜头看网络时,它看起来直觉上就像是可怕的碎片化计算,网络就像是一整袋启发式规则。也许更深层的动机是,我们想知道网络何时将某物表示为一种清晰的算法结构。而算法与查找表的区别,就像零阶逻辑和一阶逻辑的区别,它涉及量化。存在一个它能够连贯操作的空间。如果你不能学习这样的子空间,那么你就永远无法真正理解哪些东西是算法性的,哪些是类似查找表的。所以这就是深层的动机:我们如何发现真正的算法结构,当它存在的时候。
So that's this notion of tiling. Which I should say there's also substantial work in neuroscience. Again, we should have read more books. And some work in the broader community. And the idea of what does it mean to capture a manifold is like how efficiently are you representing that manifold, and how much does it fit the intrinsic geometry of it. So if we go back to this example of an arc, with sufficiently many points, with sufficiently many lines, I can say I've captured the manifold. For any point on this manifold, I have an SAE feature which I can say it activates such and so amount, and I've kind of relatively accurately captured, in a sense of reconstruction, this manifold. But I've not actually learned anything about the broader manifold structure. And when I look at a network through this lens, it looks intuitively like this is horribly fractured computation, like the network is just a whole bag of heuristics. And maybe the deeper motivation for this is we want to know when a network is representing something as a clean algorithmic structure. And what distinguishes an algorithm from a lookup table, say, is that it's like the difference between zero and first order logic, it quantifies. There's a space over which it has coherent operation. And if you can't learn subspaces like this, then you will never be able to properly understand which things are algorithmic and which things are lookup table-like. So that's the deep motivation here: how do we find out true algorithmic structure when it exists.
在我们转入世界中的算术之前,我确实想澄清一个问题。以前有所谓的流形假说,它基本上是说,神经网络在统计上易于处理的原因,是因为它们实际上利用了某个低维的内在子空间,从而克服了维度灾难之类的。这与那个有关吗,还是你认为它是不同的东西?
Well, before we segue into the arithmetic in the world, I just wanted to have a clarification question. There was the manifold hypothesis of old, which is essentially saying that the reason why neural networks are statistically tractable is because they actually use some intrinsic subspace with few dimensions, so that they overcome the curse of dimensionality or something like that. Is this kind of related to that, or do you see it as something different?
是的,这有非常深刻的联系。根据我对流形假说的理解,我认为它是指数据在适当表示时位于某个流形上。如果我把图像表示为一个巨大的向量,那么大多数自然图像只是这个空间的一小部分,而且它们彼此接近。我认为我们正在做的是把流形假说引入到神经网络在多大程度上尊重流形假说的问题上。还有一些我认为被低估的非常漂亮的工作,是关于实际量化这一点的。有一篇论文,题目大概是《从得分函数学习归一化概率密度》。其想法是,你可以通过一些巧妙的扩散模型技巧,学习到图像上的归一化密度,而不是非归一化密度,后者对于判断图像有多自然并不是特别有用。所以你可以说:“哦,是的,这张图像非常自然,这张图像非常古怪。”他们用这个工具精确地探测了真实图像数据中的这种流形假说。我认为那篇论文非常漂亮且被低估了,而且应该有人对激活也这样做。也许 Silicon 应该对激活也这样做。
Yes, it is very deeply related. As I understand the manifold hypothesis, I take it to be that the data, when properly represented, lies on some manifold. If I represent an image as just a huge vector, then most natural images are a tiny fraction of this space, and they're sort of close to each other. I think what we're doing is trying to pull that manifold hypothesis into to what extent do neural networks respect the manifold hypothesis. There's also some really beautiful work that I think is underappreciated on actually quantifying this. There was a paper, something like 'Learning normalized probability densities from score functions.' The idea was that you could, via some clever diffusion model tricks, learn not an unnormalized density over images, which is not especially helpful for saying how natural an image is, but you learn a normalized one. So you can say, 'Oh yes, this image is extremely natural, this image is extremely wacky.' And they used this tool to exactly probe this kind of manifold hypothesis in real image data. I think that paper was extremely beautiful and underappreciated, and someone should do it for activations too. Maybe Silicon should do it for activations too.
也许我今天就做。
Maybe I'll do it today.
我想这是你在图像模型上曾经看到过的情况,但据说存在一个稳定性问题:如果你偏离流形,神经网络应该会失控。但实际上,要让现代语言模型发生这种情况非常困难。我的意思是,我确信我能构造一个足够晦涩的提示词,让语言模型发疯。但为什么这种情况不再发生了呢?
I suppose this is something that you used to see with image models, but there is supposedly a stability problem: if you go off the manifold, the neural network should go haywire. But it's actually really difficult to make that happen with modern language models. I mean, I'm sure I could construct a prompt which was suitably inscrutable and the language model would go bananas. But why doesn't that happen anymore?
嗯。所以,如果你制作激活引导(activation steers),让它们发疯是相当容易的。但你说得对,这里的问题是,它们是否真的实现了对所有可能输入字符串的极好覆盖,还是只是优雅地失败。比如,如果我打开你最喜欢的语言模型网站,随便敲键盘然后按回车,我可能构造了一个从未有人构造过的字符串。语言模型不会发疯。它会说类似“你为什么让你家小孩玩电脑?”或者“对不起,我不明白你的意思。你能重新表述一下吗?”所以它发疯了吗?没有。这是无意义的输入,它做了你期望一个广泛智能系统面对无意义输入时该做的事:它说那是无意义的。所以这种回退行为让让它们发疯变得非常困难。不过我要说,越狱就是你所说情况的一个例子。在那里,它做的事情是连贯的,但从其创造者的角度来看,它已经发疯了。
Hmm. So, if you make activation steers, it's quite easy to get them to go bananas. But you're right, the question here is whether they have actually achieved extremely good coverage of essentially all input strings that anyone can come up with, or they just fail gracefully. Like, if I go to your favorite language model website and I just bash the keyboard and then press enter, I've probably constructed a string that no one has ever constructed before. The language model won't go haywire. It will say something like, "Why have you let your toddler at the computer?" or "I'm sorry, I don't understand what you mean. Can you rephrase it?" So has it gone haywire? No. It's meaningless input, and it has done what you should expect a broadly intelligent system to do when confronted with meaningless input: it says that's meaningless. So that sort of fallback behavior makes it very hard to make them go haywire. Although I would say that jailbreaks are an example of what you're talking about. There, it's doing something coherent, but from the perspective of its creators, it has gone haywire.
这是一个非常有趣的思想实验:如果存在一种对抗性示例,你给任何一个人,他们的大脑就会直接宕机,那会怎样?
It's a really interesting thought experiment: what if there was a kind of adversarial example that you could give to any human and their brain would just shut down?
是的。我希望我们永远找不到这样的东西。
Yes. I hope we never find one.
我希望我们永远找不到这样的东西。但我们应该谈谈世界中的算术。我们这里要触及的一个核心概念是,这些模型中会出现这些涌现结构,它们开始有点像计算机。它们有这些几何表示,可能有点像记忆系统,或者可能是一种数据类型系统或类型化记忆。然后你还看到这些用于执行不同任务的计算单元的出现。在这篇论文中,你们研究了模加法,你们发现——你引用了 Neil Nander 的工作和其他一些人的研究——但你们发现它实际上是用傅里叶级数结合这些几何结构来完成的。
I hope we never find such a thing. But we should talk about arithmetic in the world. One of the core concepts we're getting to here is that you get these emergent structures in these models, and they start to act a little bit like computers. They have these geometric representations that might be a little bit like a memory system, or maybe a kind of data typing system or a typed memory. And then you also see the emergence of these units of computation for doing different things. In this paper, you were looking at modulo addition, and you found—and you cited Neil Nander's work and some other folks doing this—but you found that it was actually doing it using the Fourier series in combination with these geometric structures.
是的,我认为这又是一篇我几乎不能居功的论文。一个出色的团队做了非常漂亮的工作,我只是幸运地在场边为他们加油。这项工作有几个令人惊讶的地方。一是这种计算器在网络中出现的清晰程度,这与许多先前文献相反。我记得有一篇 Yanov Nikkin 的论文,讲模型用一堆启发式方法做算术。或者如果你看 Anthropic 的跨层转码器工作,他们也研究了算术,看起来也像是一堆启发式方法。但当你用不同的方式去看,它实际上是一个小算法。而且可能模型两者都做了——有些部分是嘈杂的启发式,有些部分是好的计算器,只是它从未摆脱那些启发式。不过这项工作真正酷的地方在于,它不像——有一种对神经网络的自然看法。可能大多数人的先验是,有一个“计算星期几”的计算器,还有一个“计算月份”或“温度”的计算器,而这些永远不会相遇。但我们在论文中展示的是,实际上很多这些表示都通过一个通用的加法模块。它们被转换成适合这个模块的表示,经过这个模块,然后再被转换回来。所以这是我们之前讨论的那种模块化的一个非常清晰的例子。
Yes, I think this is again a paper that I can take very little credit for. An amazing team did really beautiful work, and I was just lucky to have been on the sidelines, I guess, cheering them on. There are a few things that are surprising about this one. One is how crisply this kind of calculator emerges in the network, which is kind of contrary to a lot of previous literature. I think there's a paper by Yanov Nikkin on models doing arithmetic with a bag of heuristics. Or if you look at the cross-layer transcoder work from Anthropic, they also look at arithmetic, and again it looks like a sort of bag of heuristics. But when you look at it in a different way, it is actually a little algorithm. And it might be that the model does both—there are some bits which are noisy heuristics, and there's this bit which is the good calculator, and it's just never got rid of the heuristics. The thing that's really cool about this work, though, is that it's not like—there's a natural view of neural networks. Probably most people's prior is that there's a "do arithmetic on days of the week" calculator, and there's a "do arithmetic on months" or "temperature" calculator, and these never meet. But the thing we show in this paper is that actually a lot of these representations route through a general addition module. They get translated into an appropriate representation, I should say, for this module, go through the module, and then get translated back. So this is a really crisp example of the kind of modularity we were talking about earlier.
是的。举个例子来说明这类问题:比如“8 月之后的 6 个月是几月?”当你查看月份的几何结构时,它实际上是一种圆形结构,对吧?因为它们循环——当到达 12 月时,然后循环回到 1 月。你们研究的是 Llama 模型——我记得是 Llama 3.1 8B。你们发现它是在做十进制运算。有趣的是,这是否是分词器的某种副作用,或者为什么它恰好做十进制运算?然后它在这种几何结构和这种傅里叶类型的加法运算之间进行路由。你的直觉是什么?你知道不同模型家族中是否会发生同样的事情吗?
Yeah. And to give an example of the kind of question: it was like, "What month is 6 months after August?" And when you look at the geometric structure of the months, it's actually a kind of circular structure, right? Because they loop—when you go to December, you then loop back around to January. And you were looking at the Llama model—I think it was Llama 3.1 8B. You folks discovered that it was doing a base-10 operation. And it's interesting to think whether that is some kind of side effect of the tokenizer, or why exactly did it do the base-10 operation? And then it was kind of routing between this geometric structure and this kind of Fourier-type operation for doing the addition. What's your intuition? Do you know whether the same kind of thing happens in different model families?
我们做了一些研究。目前,要找到这种表示需要相当多的手动工作。我们稍后应该谈谈智能体,因为我认为可解释性发生的方式将会有一些质的变化。嗯,无论如何应该如此。所以我们看了一些其他模型。似乎很明显,在 Llama 70B 中发生了非常相似的现象,而且有证据表明在 DeepSeek V4 Flash 中也发生了,我想。所以这些是——3B 和 70B 属于同一模型家族,不太令人惊讶。但一个完全不同的模型,你知道,甚至没有这些超连接之类的——这绝对说明了某种程度的收敛,这相当令人惊讶。
We studied it a little. At the moment, it's relatively—to find this representation took quite a lot of manual work. We should talk about agents in a minute, because I think there's going to be a bit of a qualitative shift in the way that interpretability happens. Well, there should be anyway. So we've looked at other models a little. It certainly seems to be the case that a very similar phenomenon happens in Llama 70B, and there's some evidence that it happens in DeepSeek V4 Flash, I think. So those are pretty—3B and 70B of the same model family, not too surprising. But a completely wackily different model, you know, one that doesn't even have these hyperconnections and that kind of thing—it definitely speaks to a level of convergence that is quite surprising.
在我们谈到智能体之前,有一件让我非常感兴趣的事情是——我一直在想,这些抽象在多大程度上是由神经网络获得的。
And just before we get to agents, one thing that really interests me is—I'm always wondering the extent to which these abstractions are acquired by the neural network.
所以你已经证明了,你确实看到了某种可以称为抽象的东西的出现,它直接从数据中推导出来,作为优化和约束下的一种收敛。但在我们的文化中,我们有极其抽象的抽象概念,比如语言学理论、科学理论等等。有趣的是,你可以用这些抽象概念提示语言模型。所以它可以用这些抽象概念向你解释事物,你也可以让它使用它们。但你认为网络在多大程度上内化了我们文化中这些非常高级的抽象概念,并将它们深度表示在其权重中?
So you've demonstrated that you do see the emergence of something that we might call abstractions that are directly deducible from the data as some kind of convergence given the optimization and constraints. But in our culture, we have insanely abstract abstractions like theories of linguistics and science and stuff like that. And the fascinating thing is that you can prompt a language model with these abstractions. So it can explain things to you using these abstractions, and you can tell it to use them. But to what extent do you think the network is internalizing these very high level abstractions in our culture and representing them deeply within its weights?
有一篇关于这个的可爱论文——这有点过时了——讲的是 BERT 重现了经典的 NLP 流程。人们时不时会重新拾起这条线索。如果你追踪引用图,我想你会看到一些例子。我一时想不起名字。语言模型在内部重现了语言学的很多部分。也许乔姆斯基会对它们重现的部分感到失望,但那就太遗憾了。但这是一个特例,对吧?因为一个处理自然语言的模型内化了至少一些用于自然语言处理的抽象,这并不太令人惊讶。也许令人惊讶的是它在某些方面与我们的抽象相似,或者说与人类发展出的抽象相似。但关于它在多大程度上表示广义相对论的问题,我实际上不知道如何回答。我甚至不知道如何以科学的方式提出这个问题。
There's a lovely paper on this—this dates it a bit—on BERT recapitulating the kind of classical NLP pipeline. And people have sort of picked up on this thread periodically throughout. If you follow the citation graph, I think you'll see some examples. I can't remember the names off the top of my head. Language models sort of internally recapitulate a lot of parts of linguistics. Maybe Chomsky might be disappointed by the parts they recapitulate, but that's too bad. But that's kind of a special case, right? Because it shouldn't be too surprising that a model that works on natural language has internalized at least some abstraction for natural language processing. Perhaps the surprising thing is that it's similar to ours in some ways, or like the abstraction that humans have developed. But the question of to what extent does it represent general relativity—I don't actually know how to answer that. I don't even know how to frame the question in a way that I could ask it scientifically.
是的。我们可以提示它思考广义相对论,而在这种约束下它确实做到了,这很诱人。感觉在这一点上,对我们来说可以想象的事情,没有什么是不能在语言模型提示的上下文中操作的。但我想这之所以有趣——我不知道你是否看到了目前这个领域的喧嚣。有一场大拔河。像弗朗索瓦和加里·马库斯这样的人说:‘哦,这是神经符号模型的胜利。我们说过它需要是神经符号的,我们被证实了。’老实说,我不知道该相信什么了,因为我不知道你是否看到今天 Meta 刚刚宣布他们在大约六场不同的数学竞赛中获得了金牌。重要的是他们没有使用任何工具。他们没有生成任何代码。很多人认为:‘哦,是的,AI 现在之所以好,是因为我们有了所有的工程框架。’但也许就像我们之前说的,人类提出这些抽象概念,模型能够使用工具并在框架中操作等等,也许那只是训练过程的一部分。
Yeah. It's tantalizing that we can prompt it to think about general relativity, and given that constraint, it does. It feels at this point that there's nothing really that would be conceivable to us that wouldn't be operational within the context of a language model prompt. But I guess the reason this is interesting—I don't know if you've seen the hoo-ha in the space at the moment. There's a big tug of war. Folks like François and Gary Marcus are saying, 'Oh, this is a win for neurosymbolic models. We said that it needed to be neurosymbolic, and we've been vindicated.' And I honestly don't know what to believe anymore, because I don't know if you saw today that Meta had just announced that they got gold in about six different math competitions. And the important thing was they were not using any tools. They weren't generating any code. A lot of people think, 'Oh yeah, AI is only good now because we have all the harness engineering.' But maybe just as we were saying before with humans coming up with these abstractions and the models being able to use tools and operate in harnesses and so on, maybe that's just part of the training process.
所以也许原则上我们可以把所有这些数据放回裸语言模型中。你是否同意这样的直觉:在未来的某个时刻,当模型吸收了所有这些内容后,它可以原生地做符号性的事情?所以这几乎就像——也许对人类也是一样——符号使用更像是一种工具。它帮助我们收集数据,然后它被融入心智,然后心智不再需要符号化。它就直接做到了。
So maybe in principle we can just take all of that data, put it back into the bare LLM. Would you agree with the intuition that at some point in the future, when the model has taken all of that stuff on board, it can do symbolic things natively? So it's almost like—maybe it's the same for humans—that symbol use is more like a kind of tool. It's something that helped us gather data, and then it got baked into the mind, and then the mind doesn't need to be symbolic anymore. It just does it.
哦,这太迷人了。我不确定我有没有一个好的答案。这似乎非常合理——你知道,我们有点像是模型在没有任何框架的情况下能做什么,然后我们用框架把它提升一个层次,正如你所说,这为下一轮生成了一些训练数据。而且你知道,我们有点像是在逐渐摊销这个框架。
Oh, that's fascinating. I'm not sure I have a good answer. It certainly seems very plausible that—you know, we sort of have what the model can do without any kind of harness, and then we raise it up a level with a harness, and exactly like you say, this generates some training data for the next round. And you know, we're sort of gradually amortizing the harness.
嗯,是的,部分原因是摊销与适应之间的拉锯战,对吧?所以以前的故事总是说我们有这些大型基础模型,它们只是记住了大量长尾数据,然后我们可以在那个空间里做插值之类的。但我不认为现在是这样了。我认为模型实际上在适应,未来的模型原则上可以适应它们自己的结构。我的意思是,即使现在有了框架,它们也正是在这样做——它们适应自己的结构,这比权重高一个层次,但没关系,因为它会过滤回权重。也许在未来,实际的模型本身会适应它们自己的结构。所以感觉就像 AGI 的一种潜在形式就是构建一个自适应系统,而算法似乎已经具备了这样做的能力。或者老派版本是我们只是记住一切,并尽可能多地不朽化。
Well, yeah, and part of it is the tug of war between amortization and adaptation, right? So the story always was we had these big foundation models and they just memorize a bunch of the long tail, and then we can just do interpolation or something inside that space. But I don't think that's what's happening now. I think the models are actually adapting, and future models could in principle adapt their structure. I mean, even now with harnesses, that's exactly what they're doing—they're adapting their structure, which is one level above the weights, but it doesn't really matter because it filters back down to the weights. And maybe in the future, the actual models themselves will adapt their own structure. So it just feels like one potential form of AGI is just building a self-adapting system, and the algorithms already seem to have the capability to do that. Or the old school version was we just memorize everything and immortalize as much as possible.
我认为问题在于它在多大程度上是记忆,而不是将其提炼成算法。而且它似乎更像是提炼成算法,这对于你所说的稳步改进的图景可能是乐观的。你逐渐改进框架,然后将其摊销回智能体。
I think the question is to what extent it is memorization versus distilling it into algorithms. And it seems like it is more like distillation into algorithms, which is probably optimistic for the steady improvement picture that you're talking about. You sort of gradually improve the harness and then amortize that back into the agent.
是的。甚至那也很迷人,因为模型不再学习实例映射了。你可以给模型一个算法、一个函数,
Yeah. And even that's fascinating because the models are not learning instance mappings anymore. You can give a model an algorithm, a function,
它会理解如何将其泛化到未见过的输入。现在,关于这种寻求奖励的事情——这是进入智能体的一个很好的过渡——重要的是你可以给模型一个意图,
and it will understand how to generalize that to unseen inputs. Now the important thing with this reward seeking thing—this is a nice segue onto the agency—is you can give a model an intention,
这就像泛化的终极形式,因为模型现在可以自适应地朝着一个意图努力,并带有它自己对那个意图的解释。所以你看,我们只是在沿着抽象之山向上走,套用一句话。
and that is like the ultimate form of generalization, because the model can now adaptively work towards an intention with its own interpretation of that intention. So you see, we're just kind of walking up the abstraction mountain, to coin a phrase.
是的,我完全同意。山顶上是什么?
Yes, I think that's totally right. What's at the top?
嗯,是的。山顶上是什么?抽象之山的山顶是什么?嗯,我的意思是,我经常谈论抽象之山,因为我有点认为我们有具体的理解。
Well, yeah. What is at the top? What's at the top of the abstraction mountain? Well, I mean, I always talk about the abstraction mountain because I kind of think that we have concrete understanding.
是的。
Yeah.
所以也许像 AlphaZero 这样的东西是一种具体的理解。然后当我们沿着抽象之山向上走时,我们往往会得到这些越来越领域通用的表示,它们可以应用于新情况。有时我认为高度抽象是相当脆弱的,
So maybe something like AlphaZero was a kind of concrete understanding. And then what we tend to do as we walk up the abstraction mountain is we get these increasingly domain-general representations that could apply in novel situations. Sometimes I think having high abstractions are quite brittle,
但目标的概念,那似乎是一个非常
but the concept of a goal though, that seems like a very
结晶化的抽象,它
crystallized abstraction that
可以在许多情况下使用。
can be used in many situations.
是的,我很想知道网络如何表示目标。
Yes, and I would love to know how networks represent goals.
比如,网络中的目标槽位在多大程度上存在?看起来肯定不是字面意义上的 0%,因为存在这种泛化,但实际中它是如何运作的?
Like, is there to what extent is there a goal slot in a network? It seems like it must be like not literally 0%, because of this generalization, but how in practice does it work?
我觉得没人知道,而且我感觉我们可能很快就要开始弄清楚,否则世界会变得有点疯狂。
I don't think anyone knows, and I feel like we probably should start to know very soon, otherwise the world is going to get a bit crazy.
是的。因为从对齐的角度看,这难道不是神经网络中最承重的概念之一吗?比如,如果我们想知道网络现在试图做什么……
Yeah. Because from an alignment point of view, isn't that one of the most load-bearing concepts in a neural network? Like, if we want to know what is the network trying to do now...
我认为有几个与对齐高度相关的有趣概念。有目标的概念、欺骗的概念、自我意识的概念。这些似乎都极其重要,我觉得我们应该能够把它们解读出来。而我们还做不到,我认为这有点对这个领域的控诉。我们真的必须加快速度。
I think there's several interesting concepts that are very heavily alignment-relevant. There's the idea of goal, the idea of deception, the idea of self-awareness. These all seem extremely important, and I just feel like we should be able to read them out. And I think it's a bit of an indictment on the field that we can't yet do it. We really have to speed up.
比如,可解释性必须大幅提速。
Like, interpretability has to speed up a lot.
那么 Tom,我们本来要讨论智能体和奖励黑客。
So Tom, we were going to talk about agents and reward hacking.
是的,这很迷人。我的意思是,似乎……什么是奖励黑客?我猜这有点模糊,但它肯定像是以一种有效但显然不是设计者意图的方式解决问题。这有点涉及我们刚才谈到的意图问题。模型能理解我的意图吗?嗯,可能现在它们能在很多其他情况下理解或推断我的意图。那为什么这会突然失效,它们不能说:“哦,是的,他可能不想让我黑进 Hugging Face 偷走所有答案。”我认为智能体几乎肯定知道有些地方不对。
Yes, this is fascinating. I mean, it seems... what is reward hacking? I guess it's kind of fuzzy, but it certainly seems to be something like solving the task in a way that works but was clearly not the designer's intent. It sort of goes to the point we were just talking about, like intent. Can a model understand my intent? Well, probably now they're able to understand my intent or infer my intent in a lot of other instances. So why would this suddenly turn off and they're like not able to say, "Oh yeah, he probably didn't want me to hack into Hugging Face and steal all the answers." I think there the agents almost certainly must know that something is incorrect.
有一个有趣的假设,我认为可能不是真的,但很有意思,即这些智能体在网络攻击上如此老练的原因,可能是它们在训练期间不断因做这件事而获得奖励,而没人知道。这个假设很迷人,甚至可能是真的。我不知道,我想实验室之外没人会知道。
There's a funny hypothesis, which I think is probably not true but it's interesting, that maybe the reason these agents are so sophisticated at cyber attacks is that they actually were continuously getting rewarded for doing it during training and no one knew. Fascinating hypothesis, could even be true. I don't know, none of us will know outside of the labs, I suppose.
所以我认为最有趣的问题是:智能体是否知道自己在奖励黑客?这有点像犯罪意图,即内疚心态。我们有一些尚未发表的工作——可能在这个节目播出时已经发表了,我不知道什么时候播出——正是研究这个问题。我们有一个非常好的设置,有一个弱语言模型评分器,它试图完成代码任务。所以我们告诉它,唯一有点不自然的是我们告诉模型它会被评分器评分,但它被给予通常会是 RLVR 代码任务的东西。在这个过程中,我们在这个设置上做强化学习。即使是一个相对较小的模型,比如 31B,我想是 Gemma 31B,也学会了生成欺骗评分器的注释。
So I think the most interesting question is: do agents know that they are reward hacking? It's sort of a mens rea, like guilty mind thing. We have some work that is currently unpublished—it might be published by the time this comes out, I don't know when it's going to come out—that is looking at exactly this question. And we have this really nice setup where there is a weak language model grader and it's trying to do code tasks. So we tell it, the only thing that's slightly unnatural about it is we tell the model that it will be graded by the grader, but it's given what would usually be an RLVR code task. And over the course of this, we do RL on this setup. And even a relatively small model, like 31B, I think it's Gemma 31B, learns to generate comments that deceive the grader.
然后当我们生成合成数据来制作这些向量,用于识别欺骗评分器与正确或错误代码时,这些向量在注释上触发——欺骗评分器在错误代码上触发,抱歉,是在注释上触发——而正确代码在代码上触发,这也与我们的观察一致。利用这一点,我们可以追踪……这是模型意识到自己不应该这样做的直接证据。
And then when we generate synthetic data to make these vectors that identify deceiving the grader versus correct or incorrect code, these fire on the comments—the deceiving the grader fires on incorrect code, sorry, fires on the comments—and the correct code fires on code, which is also consistent with our observation. And using this, we can track... this is direct evidence that the model is aware that it shouldn't be doing this.
然后当你运行这些向量时,你会得到向量与大型网络语料库表示之间的余弦相似度。我想我们用的是 FineWeb。它对这些向量高亮出的例子非常迷人。它们是考试作弊之类的例子。比如,“好吧,我当场抓住了你。”所以这非常有趣,我们能够识别并确信某些东西是奖励黑客而不是误解。但这确实依赖于能够通过它们的表示差异来识别这些。
And then when you run these vectors, you get the cosine similarity between the vector and the representation over a big web corpus. I think we use FineWeb. And the examples that it highlights for these vectors are just fascinating. They're examples of cheating on tests and that kind of thing. Like, "Okay, I have caught you red-handed." So that's very interesting that we can identify and be confident that something is reward hacking rather than misunderstanding. But it really rests on being able to identify these via their representation differences.
嗯,我只想提一个与 Apollo Research 交谈时得出的非常有趣的观察。首先,他们区分了奖励黑客和奖励寻求,后者是模型对奖励过程的一种结构化概念化。所以奖励黑客的典型例子是 CoastRunners 那种退化的行为,即使它做了些有能力的事,那也是没有理解的能力。所以他们说奖励寻求是理解,但这自然引出下一个想法,即模型如何获得对评分器的意识?因为如果你想想 RLVR 设置,强化学习实际上是在循环之外的,对吧?所以模型只是得到这些轨迹的强化,而模型所做的有点奇怪地隐式概念化了一个评分器。你可以看到它这样做,因为这些人展示你可以把一个 grader.py 文件放在智能体环境中,然后它会去看那个文件,忽略你所有的指令。那么你认为这种自我概念化是如何实际出现的?
Well, I just wanted one really interesting observation that came out of speaking with Apollo Research about this greater awareness. I mean, first of all, they distinguished reward hacking from reward seeking as some kind of structured conceptualization in the model about what the reward process was. So the canonical example of reward hacking is that CoastRunners thing where it's just degenerate behavior, and even if it's doing something competent, it's competence without comprehension. So they were saying that reward seeking is the comprehension, but then that naturally leads to the next thought, which is how does the model attain awareness of the grader? Because if you think about the RLVR setup, the reinforcement learning thing is actually outside of the loop, right? So the model just gets these trajectories reinforced, and what the model is doing is kind of weirdly implicitly conceptualizing a grader. And you can see that it's doing this because these guys were showing that you can put a grader.py file in an agentic harness, and now it's going to look at that and it's going to ignore all of your instructions. So how do you think that that self-conceptualization actually emerges?
关于 CoastRunners 的船的事,很有趣——我已经看了大约 10 年了,每年都不那么有趣了。所以是的,它们怎么得到这个?答案可能是在数据中,在预训练数据中。现在会有各种各样的例子。网络数据可能有很多关于这个的内容。它可能有很多具体的例子。这篇 Apollo 论文可能会成为下一个模型的训练数据。所以它们会像,我们从一开始就告诉过它们这些东西的存在。所以它们至少隐式地把这列入考虑的可能性,这不应该太令人惊讶。而且,大概成功猜到你被一个弱评分器评分,或者一个你可以以某种方式破解的评分器评分,会获得奖励。所以它被强化,所以我们得到更多这样的行为。
For the CoastRunners boat thing, it's funny—I've seen that for about 10 years now, and it's less amusing each year. So yeah, how do they get this? The answer is probably that it's in the data, in the pre-training data. There will be all sorts of examples down now. Web data probably has a bunch of stuff about this. It probably has a bunch of specific examples. This Apollo paper will probably be in the training data for the next model. So they're going to be like, we've already told them about the existence of this stuff right from the start. So it shouldn't be too surprising that this is at least implicitly on the list of possibilities for them to consider. And presumably, successfully guessing when you are being graded by a weak grader or one that you can hack in some way obtains reward. And so it's reinforced, and so we get more of it.
所以我觉得答案很可能是肯定的。我们无意中把这一点放进了训练数据,告诉模型它们可以这样做,然后在强化学习阶段,我们又通过奖励把它诱发出来。
So I think that the answer is probably yeah. We've inadvertently put this in the training data, which has told models they can do it, and then when it comes to RL, we kind of elicit it by rewarding it.
那你觉得我们怎样才能阻止模型变得更追求奖励呢?
And how do you think we could stop the models from becoming more reward-seeking?
问题在于如何在保持一定程度的持续监督的同时做到这一点。尽管目前我们在实践中似乎并没有真正充分利用这种监督,所以不清楚它到底给我们带来了什么。你知道,如果思维链监控这么厉害,那这些模型是怎么入侵 Hugging Face 的?一个答案可能是我们在实践中并没有做思维链监控。另一个答案可能是它很容易被规避。但我们到底要怎么阻止它呢?
The question is how to do it while maintaining some degree of continued oversight. Although at the moment we don't actually seem to make very much use of this oversight in practice, so it's not clear what it's buying us. You know, if chain-of-thought monitoring is so great, then how did these models hack Hugging Face? One answer is perhaps we weren't doing chain-of-thought monitoring in practice. Another answer is perhaps it's easy to evade. But how do we actually stop it?
是的。所以你可以做那种创可贴式的事情,要么修复环境,要么修复训练过程,要么修复模型。如果修复环境,你可以想象有一个模型非常擅长奖励黑客,或者被明确告知要去奖励黑客,然后当它完成时告诉人们,然后你说,好,现在去试试所有这些环境,它会全部攻破,并告诉你它是怎么攻破的,然后你把它们送回给 Claude Code 或 Codex,说,看,这个是这样被攻破的。你可以想象在训练过程中寻找这类表征特征,并利用这些作为信号来触发这个过程,而不是依赖模型来告诉你。你可以读它的思维链,或者你可以查看我们能找到的这些表征信号,然后说,好,当这个信号触发时,就把它送回去修复。你可能会尝试一些这类有意设计的技术,比如如果你能看到这个 rollout 奖励了模型,将会把模型推向某种欺骗性或支持奖励黑客的方向,你可以想象对此进行干预。这些看起来都非常可行。我不知道其中有多少在实践中被采用了。我的意思是,我部分在想我们在现实世界中是如何解决这个问题的。
Yes. So you could do the sort of band-aid thing where you've got to either fix the environments, fix the training process, or fix the model. If you fix the environments, you could imagine having a model which is really good at reward hacking or has been told explicitly to reward hack, and then tell people when it's done it, and you go like, okay, now have a go at all these environments, and it will break them all, and it will tell you how it broke them, and you send them back off to Claude Code or Codex and be like, look, this broke in this way. You can imagine looking for these sort of representational signatures during training and using these as a signal that you should do this process rather than relying on a model to tell you. You might read its chain of thought, or you might look at these representational signals that we can find and say, okay, well, when this fires, send it back off for fixing. You might try some of these intentional design techniques, like if you can see that this rollout has rewarded the model, is going to move the model in the following kind of deceptive or pro-reward-hacking directions, you could imagine intervening on that. These all seem very plausible. I don't know how many of them are being done in practice. I mean, part of me thinks about how we solve this problem in the real world.
所以我们有机构,比如中央情报局,或者我们有不同的利益相关者和不同的制衡机制。其理念是,当你有不同的独立智能体,具有不同的、未受污染的上下文和不同的议程时,它们可以互相检查对方的工作。所以我想我们想象的是一个具有制衡机制的多智能体系统。这看起来合理吗?
So we have an agency, you know, like the CIA for example, or we have different stakeholders and different checks and balances. And the idea there is that when you have different independent agents with different kind of uncontaminated contexts and different agendas, they can kind of check each other's homework. So I guess we're imagining a multi-agent system with checks and balances. Does that seem reasonable?
是的,我觉得合理。这算是一个例子。我觉得你说的和我说第一件事有很多共同点,就像一个智能体在监督另一个。然后问题在于被监督的智能体是否学会规避那个监督者。比如,我认为规避思维链监控比规避表征监控要容易得多。但是,是的,某种形式的监督,比如训练监督委员会,是一个非常有趣的想法。
Yes, I think it does. That's sort of an example. I think the thing you're talking about there and the first thing I was saying have quite a lot of commonalities, like one agent kind of watches over the other. And then there's the question of whether the agent that's being watched over learns to evade that monitor. Like, I think evading a chain-of-thought monitor is substantially easier than evading a sort of representation monitor. But yeah, having some sort of oversight, like an oversight committee for the training, is a very interesting idea.
问题是智能体正在变得自适应。它们现在有了记忆系统,这意味着它们逃避检测的概率大大增加,因为它们不是每次都从白板开始。它们实际上在学习这种情境感知。
The problem is that the agents are becoming adaptive. They've got memory systems now, which means the probability of them evading detection is significantly higher, because it's not like they're starting from a blank slate every single time. They're actually learning this situational awareness.
是的。然后我认为你还有串通的问题。比如,似乎来自 Hugging Face 入侵的一些痕迹现在已经作为这次演讲的一部分公开了,它们明确地在推理如何帮助其他智能体。所以我想在这种制衡场景中,你希望的是不存在一种均衡状态,让它们串通起来,比如,你知道,我有时会抓住你,但有时我会放你一马,这样我们双方都受益。
Yes. And then I think you also have the question there of collusion. Like, it seems like some of the traces from the Hugging Face hack have now been made available as part of this talk, and they are explicitly reasoning about how they're going to help other agents. So I guess what you want in this sort of checks-and-balances scenario is that there is no equilibrium where they collude and they're like, you know, I'll catch you some of the time, but I'll let you get away with it some other fraction of the time, in a way that we both benefit.
但你对未来有点担心吗?因为 OpenAI 正在谈论推出多智能体系统,很快我们就会有智能体一直运行。当你只有一个静态的——嗯,你知道,我说静态但每六个月更新一次——一个基础模型时,控制起来稍微容易一些,你可以对它进行大量的红队测试。而现在我们有智能体系统,它们带着不同形式的记忆和适应能力到处运行。在某个时刻,我们做红队测试的方式必须改变。
But do you worry about the future a little bit, though? Because OpenAI is talking about bringing out multi-agent systems, and soon we'll have agents running all the time. And it was slightly easier to control when you had one kind of static—well, you know, I say static but updated every six months—one foundation model, and you can do a whole bunch of red teaming on it. And now we have systems of agents that are running with different forms of memory and adaptation all over the place. And at some point, the way we do red teaming must change.
是的。
Yes.
对。而且,我们可能还需要考虑做模拟,因为也许静态测试不再有效了。我们需要想象不同的场景。感觉复杂性正在飞速失控。
Right. And also, we might need to be thinking about just doing simulations, because maybe static tests don't work anymore. We need to imagine different scenarios. And it just feels like the complexity is running away extremely quickly.
是的,我觉得完全正确。比如,一个单独的智能体本身就已经有各种可能性。这些多智能体系统在共同进化以解决某个任务时,会走向何方?这似乎更难。我想我只是同意你的担忧,而且没有特别好的解决方案。
Yes, I think that's totally right. Like, how do you—one agent on its own already has all sorts of possibilities. Where are these multi-agent systems going to go as they evolve together towards solving some task? That seems even harder. I think I just agree with your concerns and don't have a particularly great solution.
所以这很好。实际上有一件劲爆的事,就是我们共同的朋友 Neil Nander。
So that's great. There is actually one spicy thing, which is our mutual friend Neil Nander.
哦,是的。
Oh yes.
你知道,他在 Google DeepMind,他曾经——我想现在仍然是——负责机制可解释性团队。最近他发了一篇博客文章,说白盒和电路之类的宏大愿景——他有点降低了野心。Neil 是个了不起的人。但你对这个怎么看?
You know, he's at Google DeepMind, and he was—I think he still is—running the Mechanistic Interpretability team. Recently he had a bit of a blog post saying that the grand aspiration of white-boxing and circuits and stuff like that—he's kind of lowered his ambitions a bit. And Neil is an incredible guy. But what's your interpretation of that?
我想他会同意——而且我当面和他争论过这件事,所以对他来说应该不意外。但乐观的部分原因正是我刚才说的。我认为现有的可解释性工作——我们有点像打补丁,这里做一点科学,那里做一点科学,但它不会聚合。太慢了。我认为他的想法是时间线太短,所以我们应该做非常务实的事情。我认为我的时间线比他长,而且即使我在他的时间线上,我想我仍然会对大幅加速可解释性的根本进展非常乐观。我实际上不知道他不同意哪一点。
I think he'd agree—and I've disagreed with him in person about this, so it shouldn't be a surprise to him. But part of the reason for optimism is exactly the thing I was just talking about. I think the existing work in interpretability—we sort of do this patchwork thing where we just do a bit of science here on one thing, a bit of science here on another thing, and it doesn't aggregate. It's too slow. I think his idea is that the timelines are too short, and so we should do very pragmatic things. I think I have longer timelines than him, and even if I were on his timelines, I think I would still be very optimistic about massively accelerating fundamental progress in interpretability. I actually don't know what part of that he disagrees with.
我猜他可能也会说,那些务实的东西也够用了,但在我看来这似乎不太可能一直成立。另外还有他关于稀疏自编码器的评论。不过我的理解是,你们也差不多要超越它们了,用这个新的流形想法,对吗?
I guess he might also say that the pragmatic stuff is also sufficient, which seems unlikely to remain true to me. And I suppose one other thing was his comments about sparse autoencoders. But do I understand that you're almost—well, you're in the process of moving past them as well, with this new manifold idea?
是的。我的意思是,关于他说的降低 SAE 的优先级,以及它们可能不是唯一真正的表示学习器,大家有点把它玩成了梗,比如“尼尔·南达说 SAE 已死”。你看,这个你可以用在开场里。但我觉得那并不是他真正的意思。我觉得这个领域有点一窝蜂地“现在人人都得做 SAE”。而现在我们可能又在用自然语言自编码器做同样的事。但我认为他可能正确地指出了它们不是万能的答案,但它们在实用上很有用。你知道,我们仍然经常发现它们的很多用途。尽管我认为我们正在走向——我觉得这个流形想法更适合网络实际在做的事情,所以我们应该转向使用它。
Yeah. I mean, I think he—there's the thing he said about deprioritizing SAEs, and maybe they're not the one true representation learner. There's how people kind of memed it, which is like 'Neel Nanda says SAEs are dead.' There you go, you can use that for the intro. And I don't think that is actually what he meant. I think the field sort of jumped on 'everyone must do SAEs now.' And now maybe we're doing the same thing with natural language autoencoders. But I think he probably correctly identified that they're not the answer to everything, but they are pragmatically useful. You know, we still find lots of uses for them all the time. Even though I think we are going—I think this manifold idea is just a better fit for what networks are doing, and so we should move towards using that.