On the Biology of a Large Language Model
打开互动全文版(中英对照 + 朗读 + 问答)→来自 Anthropic 的 Joshua Batson 探讨机制可解释性,将梯度下降训练的神经网络比作进化塑造的生物系统,并探索大型语言模型既令人印象深刻又奇怪的行为。
Joshua Batson from Anthropic discusses mechanistic interpretability, comparing neural networks trained by gradient descent to biological systems shaped by evolution, and explores both the impressive capabilities and strange behaviors of large language models.
今天我很荣幸欢迎来自 Anthropic 的 Joshua Batson。他将谈论大型语言模型的生物学,这应该是一个非常有趣的演讲。Josh 领导 Anthropic 机制可解释性团队的电路工作。在加入 Anthropic 之前,他在 Chan Zuckerberg Biohub 从事病毒基因组学和计算显微镜研究,他的学术背景是纯数学。另外,本季度还有一些其他演讲的录像已经发布,比如 Karina 和 Div 的演讲。欢迎在 YouTube 播放列表上查看。对于 Zoom 上的朋友,欢迎在 Zoom 或 Slido 上提问,代码是 CS25。话不多说,有请 Josh。
So today it's my pleasure to welcome Joshua Batson from Anthropic. He'll be talking about on the biology of a large language model, which should be a very interesting talk. Josh leads the circuits effort of the Anthropic mechanistic interpretability team. Before Anthropic, he worked on viogenomics and computational microscopy at the Chan Zuckerberg Biohub, and his academic training is in pure mathematics. Also, some more recordings for this quarter have been released, like Karina and Div's talks. Feel free to check those out on our YouTube playlist. For folks on Zoom, feel free to ask questions either on Zoom or Slido with the code CS25. Without further ado, I'll hand it off to Josh.
谢谢。很高兴来到这里。现在竟然有关于 Transformer 的课程,这让我觉得疯狂,因为我认为它们是不久前才发明的。我们大约有一个小时的课程时间,然后 15-20 分钟提问。欢迎随时打断我提问。我的背景是纯数学,在那里人们总是很粗鲁地互相打断,这完全没问题。如果我想打断你并继续,我也会的。这可以像你们希望的那样互动,对 Zoom 上的人也一样。
Thank you. It's a pleasure to be here. It's crazy to me that there is a class on transformers now, which I think were invented rather recently. We have about an hour for the class and then 15-20 minutes for questions. Feel free to interrupt me with questions. My training is in pure mathematics, and people just interrupt each other all the time very rudely, and it's totally fine. If I want to cut you off and move on, I will. This can be as interactive as you like, and that's true for people on Zoom too.
这次演讲的标题是《论大型语言模型的生物学》,这也是几周前发布的一篇 100 页互动博客文章的标题。如果你在这里,你可能对大型语言模型有所了解。'生物学'这个词是我们选择的。你可以将其与一系列名为《论大型语言模型的物理学》的论文对比,后者将模型视为训练过程中的动态系统。但我们认为可解释性与神经网络的关系,就像生物学与通过进化发展的生命系统的关系一样。你有一个产生复杂性的过程,你可以研究产生的对象,看看它们如何完成那些神奇的事情。
This talk is titled 'On the Biology of a Large Language Model', which is also the title of a paper, a 100-page interactive blog post that went out a few weeks ago. If you're here, you probably know something about large language models. The word 'biology' was our choice here. You might contrast it with a series of papers called 'On the Physics of Large Language Models', where you think of them as dynamical systems over the course of training. But we think of interpretability in relation to neural networks, which are trained by gradient descent, as biology is to living systems that develop through evolution. You have a process that gives rise to complexity, and you can study the objects produced to see how they do the miraculous things they do.
模型做了很多很酷的事情。这是一个大约 6 个月前的例子,在 AI 时间中相当于 10 年。有人从事 Sarcasian 语(一种极低资源语言)的 NLP 工作,多年来一直追踪最先进技术。他尝试了一个版本的 Claude,可能是 Sonnet 3.5,将多年收集的俄语-Sarcasian 翻译主列表放入上下文窗口,并要求模型进行其他翻译。它不仅能成功翻译,还能解析语法。这些模型的上下文学习击败了他一直研究的 NLP 专用模型的最先进水平。这很酷。
Models do lots of cool things. Here's an example from maybe 6 months ago, which is like 10 years in AI time. Someone working on NLP for Sarcasian, a very low-resource language, had been tracking state-of-the-art for years. He tried a version of Claude, probably Sonnet 3.5, where he put a master list of Russian-Sarcasian translations into the context window, gathered over years, and asked the model to do other translations. It could not only translate successfully but also break down the grammar. In-context learning with these models beat the state-of-the-art for the NLP-specific models he'd been working on. That's cool.
但模型也很奇怪。这也是 Claude。有人问闰日第二天是什么日子,它陷入了大混乱。它说:如果今天是 2024 年 2 月 29 日,那么明天是 3 月 1 日,然而 2024 年不是闰年,这是不正确的。所以 2 月 29 日在公历中不是有效日期。然后它说 2024 年之后的下一个闰年是 2028 年,这是对的。然后如果我们假设你指的是 2024 年 2 月 28 日,二月的最后一个有效日期,它给出了这样的回答,简直莫名其妙。它混杂了正确的事实回忆、基于事实的正确推理,然后又因为与初始假设一致而忽略它们。闰日这件事很奇怪。如果一个人这样做,你会想知道他们吃了什么。孩子们有点像这样。也许这是一个有趣的话题。
But models are also weird. This was also Claude. Someone asked what day is tomorrow on leap day, and it got into a big fight. It said: if today is February 29th 2024, then tomorrow would be March 1st, however 2024 is not a leap year, which is untrue. So February 29th is not a valid date in the Gregorian calendar. Then it says the next leap year after 2024 will be 2028, which is true. Then if we assume you meant February 28th, 2024, the last valid date in February, then it gives this just like what is going on? There's a smorgasbord of correct recollection of facts, correct reasoning from the facts, and then disregarding them out of consistency with the initial assumption. It's pretty weird for it to be leap day. If a person did this, you'd wonder what they consumed. Children are sort of like this. Maybe that's an interesting topic.
我喜欢这个:'AI 艺术会让设计师过时。AI 接受了这份工作。'而且它有很多手指。这现在已经过时了;人们已经弄清楚了如何将手指数量控制在每只手最多五个。新的 ChatGPT 模型可以生成极其逼真的人,都有五根手指。但这并不是通过弄清楚为什么有那么多手指来解决的;其他方法解决了这个问题。你压制了一些怪异之处,现在怪异之处更加复杂了。随着模型变得更好,你需要更好地理解疯狂之处去了哪里。随着前沿向前推进,也许不是五根手指,但可能有其他细微的错误。
I love this: 'AI art will make designers obsolete. AI accepting the job.' And it has so many fingers. This is now out of date; people figured out how to keep the finger count down to at most five per hand. The new ChatGPT model can do extremely realistic people with five fingers. But that wasn't solved by figuring out why there were so many fingers; other methods got through it. You kind of bat down some weirdness, and now the weirdness is more sophisticated. As models get better, you need to get better at understanding where the craziness has gone. As the frontier moves forward, maybe it's not five fingers, but there might be subtly other things wrong.
当我思考可解释性——模型到底学到了什么,它如何在内部表示,如何在行为中体现——我展望未来,那时大多数简单交互似乎都进行得很好。那么问题是:这些进行得好是因为模型学到了深刻而真实的东西,还是因为你像处理手指问题一样压制了它?如果你走到模型能力的边缘,一切又会变成七根手指,但你无法分辨。因为它看起来很可靠,你已经将大量决策和信任委托给了这些模型,而在角落里你无法再验证哪里变得奇怪。出于这个原因,我们想了解这些能力到底是怎么回事。
When I think about interpretability—what did the model learn exactly, how is it represented inside, how does it manifest in behaviors—I think ahead to when most simple interactions seem to go well. Then the question is: are these going well because the model learned something deep and true, or because you managed to beat this down like the finger problem? If you went to the edge of the model's capabilities, it would all be seven fingers again, but you can't tell. Because it seems reliable, you've delegated a lot of decision-making and trust to these models, and in the corner you can no longer verify where things get weird. For that reason, we want to understand what's going on with these capabilities.
一些拆解它们的策略,以及我认为我们学到的关于模型内部工作原理的三个主要教训,这些并非通过黑箱方式就能显而易见。这里有三个说法,介于迷思和流行观点之间,或者如果你从哲学角度解读可能成立,但并未抓住要点。这些说法是:模型只是对相似训练数据示例进行模式匹配;它们只使用浅层简单的启发式推理;以及它们只是逐词工作,像某种自发喷发一样吐出下一个词。我认为我们发现,模型内部可以学习并组合相当抽象的表征;它们执行相当复杂且通常高度并行的计算——不是串行的,所以它们同时在做很多事情;而且它们还会规划未来很多个词。所以即使它们一次只说一个词,它们也常常提前思考很远,以便能生成连贯的内容。
Some strategies for picking them apart and then three main lessons I think we've learned about how models work inside that aren't obvious from a black-box way of engaging with them. So here are three statements that are somewhere between myths or things that are out there, or true if you interpret them in one way philosophically but missing the point. These statements are: models just pattern match to similar training data examples; they use only shallow and simple heuristics in reasoning; and they just kind of work one word at a time, like gutting out the next thing in some eruption of spontaneity. I think we find that models learn and can compose pretty abstract representations inside; they perform rather complex and often heavily parallel computations—it's not serial, so they're doing a bunch of things at once; and also they plan many tokens into the future. So even though they say one word at a time, they're often thinking ahead quite far to be able to make something coherent.
这对这门课来说可能是复习,但我们做了这些漂亮的幻灯片,我觉得还是过一遍比较好。你有一个聊天机器人,它说“你好,我怎么帮你?”这实际上是怎么发生的?它是一次说一个词,通过预测下一个词来实现。所以“how”经过 Claude 预测出“can”,“how can”预测出“I”,“how can I”预测出“assist”,“how can I assist”预测出“you”。所以你可以把问题简化为给出下一个词的计算。而这就是一个神经网络。这里我画了一个全连接网络。要把语言通过数字变成语言,你首先得把东西变成向量。所以有嵌入层。词汇表中的每个词或词元都有一个嵌入,也就是一个数字列表或向量。从本质上讲,你基本上就是把它们拼接起来,然后通过一个有很多权重的巨大神经网络。输出是词汇表中每个词的分数,模型选择分数最高的词,再经过一些温度参数引入随机性。Transformer 架构更复杂,有残差连接、交替的注意力机制和 MLP 块,你可以认为这是把一个非常强的先验嵌入到一个巨大的 MLP 中,但从某种意义上说,这仅仅是为了效率。
This is probably review for this class but we made these nice slides and I think it's nice to go through anyway. So you have a chatbot which says 'Hello, how can I assist you?' How does this actually happen? It is saying one word at a time by just predicting the next word. So 'how' goes through Claude and predicts 'can', and 'how can' goes to 'I', and 'how can I' goes to 'assist', and 'how can I assist' goes to 'you'. So you can reduce the problem to the computation that gives you the next word. And that is a neural network. Here I've drawn like a fully connected network. To turn language into language passing through numbers, you have to turn things into vectors first. So there's an embedding. Every word or token in the vocabulary has an embedding, which is a list of numbers or a vector. Morally speaking, you basically just concatenate those together and run them through a massive neural network with a lot of weights. Out comes a score for every word in the vocabulary and the model says the highest scoring word modulo some temperature to introduce randomness. Transformer architectures are more complex, with residual connections and alternating attention and MLP blocks, which you could think of as baking in a really strong prior into a massive MLP, but in some sense that's just about efficiency.
我们发现一个有用的比喻,也是生物学上的比喻:语言模型应该被视为“生长”出来的,而不是“建造”出来的。你从一个随机初始化的东西开始,一个像脚手架一样的架构。你给它一些数据,这就像养分,然后损失函数就像太阳,它朝着那个方向生长。最终你得到这种有机的东西。但它的生长方式,你实际上无法访问。你能访问的是脚手架,但这就像看一个模型在网中,往往不那么有趣。
A metaphor we found useful, which is also biological, is that language models should be thought of as grown, not built. You kind of start with this randomly initialized thing, an architecture which is like a scaffold. You give it some data, maybe that's like nutrients, and then the loss is like the sun, and it grows towards that. So you get this kind of organic thing which has been made by the end of it. But the way it grew, you don't really have any access to. The scaffold you have access to, but that's like looking at a model at a net, which tends not to be that interesting.
当然,我们有模型,所以对于它们在做什么,有一个同义反复的答案,我已经告诉过你了:它们把词变成数字,做一堆矩阵乘法,应用简单函数——从头到尾只是数学——然后你得到输出。就是这样。这就是模型所做的。我认为这是一个令人不满意的答案,因为你无法据此推理。这个关于模型如何工作的答案,并不能告诉你它们应该或不应该能够做什么行为,或者任何这类事情。
Of course, we have the models, and so there is a tautological answer to what they are doing, which I already told you: they turn the words into numbers, they do a bunch of matmuls, they apply simple functions—it's just math all the way through—and then you get something out. That's it. That's what the model does. I think that's an unsatisfying answer because you can't reason about it. That answer to how models work doesn't tell you about what behaviors they should or shouldn't be able to do, or any of those things.
你首先可能希望的是,神经网络内部的神经元可能有可解释的角色。人们从 80 年代的第一批网络开始就抱有这种希望。深度学习有过一次小小的复兴。Anthropic 团队的负责人 Chris Olah 在 10 年前就对此非常热衷。你只需观察一个神经元,问它什么时候激活?对于什么输入这个神经元是活跃的?然后你就能看出它们是否形成一个连贯的类别。这是视觉模型中的汽车检测神经元,或者这是眼睛检测器。如果是在模型早期,这就像边缘检测器。例如,他们在 CLIP 模型中发现了唐纳德·特朗普神经元。但事实证明,在语言模型中,当你这样做时,你只是说哪些句子导致这个神经元激活。答案并不那么有意义。这里有一个可视化,展示了一个模型中的神经元激活的一系列示例文本。里面有很多东西:代码、一些中文、一些数学,还有“苏格拉底的毒芹”。这并不特别清晰。当然,它没有理由必须是清晰的;它只是学会了运作。要求神经元可解释有点像孤注一掷。它有时能奏效很酷,但并不是特别系统。
The first thing you might hope is that the neurons inside the neural network might have interpretable roles. People were hoping this going back to the first networks in the 80s. There was a bit of a resurgence in deep learning, a kind of small one. Chris Olah, who leads the team at Anthropic, got really into this 10 years ago. You just look at a neuron and ask when does this fire? For what inputs is this neuron active? And then you just sort of see if those form a coherent class. This is the car detector neuron in a vision model or this is the eye detector. This is like the edge detector if it's early in the model. They found the Donald Trump neuron in a CLIP model, for example. But it turns out that in language models, when you do this, you just say which sentences cause this neuron to activate. The answer doesn't make that much sense. Here's a visualization of a bunch of example text for which a neuron in a model activates. There's just a lot of stuff: code, some Chinese, some math, and 'hemlock for Socrates'. It's not especially clear. Of course there's no reason it would need to be; it's just learned to function. Asking for a neuron to be interpretable is a bit of a Hail Mary. It's pretty cool that it works sometimes, but it's not particularly systematic.
有一个来自神经科学的先验:也许尽管有大量神经元在活动,但模型并没有同时思考那么多事情。也许存在某种稀疏性,如果有一张模型正在使用的概念或子程序的地图,那么在任何一个词元上,它一次只使用少数几个。这只是一个比“神经元可解释”稍微好一点的猜测。它不一定是一个很好的先验,但你可以用它来工作。所以你可以拟合神经元的线性组合,使得每个激活向量是这些字典元素的稀疏组合之和。这叫做字典学习,经典的机器学习。我们就这样做了。你从模型中收集大量激活,通过数万亿文本之类的,然后你取那些向量,寻找字典,然后你看那些字典成分何时活跃,瞧,效果好多了。去年我们有一篇论文,展示了我们在 Claude 3 Sonnet 的中间层上拟合了大约 3000 万个特征。所以你问模型内部的计算或表征原子是什么。这是我最喜欢的一个:当输入是关于金门大桥时,这个线性组合就会出现——它像一个点积,但向量很大——如果输入是左侧英文中明确提到金门大桥,那它就是真的。
There's a prior from neuroscience that maybe while there are a whole bunch of neurons going on, maybe the model is not thinking of that many things at once. Maybe there's some sparsity here where if there were a map of the concepts the model is using or the subroutines it's using, on any given token it's only using a few at a time. That's just a slightly better guess than maybe the neurons are interpretable. It's not necessarily a great prior but it's something you can work with. So you can fit linear combinations of neurons such that each activation vector is a sum sparse combination of these dictionary elements. This is called dictionary learning, classical ML. We just did it. You gather a bunch of activations from the model and put through a trillion texts or something, then you take those vectors, you look for dictionaries, and then you look at when those dictionary components are active, and lo and behold, it's way better. We had a paper last year which was just like here's a bunch of them where we fit about 30 million features on Claude 3 Sonnet on the middle layer of the model. So you just say what are the atoms of computation or representation inside that model. This was one of my favorites where this linear combination is present—it's like a dot product but the vector is large—when the input is about the Golden Gate Bridge, and that is true if it is an explicit mention of the Golden Gate Bridge in English on the left.
如果翻译成另一种语言,也是如此。如果是金门大桥的图片,也是如此。事实证明,如果是间接提及,比如你说‘我从旧金山开车去马林’,你知道要过桥才能到,同样的特征也会被激活。所以这有点像相对通用的概念,也涉及旧金山地标等。这些神经元的组合是可解释的。我们对此很满意。还有一些更抽象的东西,比如内心冲突的概念。有一个针对代码错误的特征,它会在各种编程语言中的小错误(比如除以零或拼写错误)上被激活。如果你抑制它,模型就会表现得好像没有错误。如果你激活它,模型就会给出一个回溯,好像有错误一样。所以它有点这种通用属性。但有一点非常不令人满意,那就是它怎么知道那是金门大桥,以及它用这些信息做什么?所以即使你成功拆解了表征,那也只是个横截面。它并不能告诉你为什么和怎么样,只告诉了你是什么。所以我们想找到将这些特征连接起来的方法。从输入词开始,处理成更高阶的表征,最终它会输出一些东西,并试图通过优化来追踪这个过程。
Also true if it's translated into another language. Also true if it's an image of the Golden Gate Bridge. Also true, it turns out, if it's an indirect mention, so you're like 'I was driving from San Francisco to Marin,' which you know you cross the bridge to do, and the same feature is active there. So it's like some relatively general concept, also in San Francisco landmarks, etc. These combinations of neurons are interpretable. We were happy with this. There are things that are more abstract, like notions of inner conflict. There was a feature for bugs in code that fired on kind of small bugs like division by zero or typos in many different programming languages. If you suppressed it, the model would act like there wasn't a bug. If you activated it, the model would give you a traceback as if there were a bug. So it sort of had these general properties. But there was something very unsatisfying about this, which is how does it know it's the Golden Gate Bridge, and what does it do with that information? So even if you manage to piece apart the representations, that's just a cross-section. It kind of doesn't give you the why and the how. It just gives you the what. So we wanted to find ways of connecting these features together. So you start from the input words, process these to higher-order representations, and eventually it'll say something, and try to trace that through optimization.
所以非常具体地说,你有一个矩阵,比如你取了十亿个激活向量的样本,然后你有一个 D 模型,这是一个矩阵,你试图将该矩阵分解为一个固定的原子字典与一个稀疏矩阵的乘积,该稀疏矩阵表示每个样本中存在哪些原子,并且你有一个目标函数,即应该重建数据,并加上一些 L1 惩罚以鼓励稀疏性。这是一个联合优化问题。
So very concretely, you have a matrix which is like you take a billion examples of activation vectors, and then you've got D model here that's a matrix, and you try to factorize that matrix into a product of a fixed dictionary of atoms times a sparse matrix of which atoms are present in which example, and you have an objective function which is this should reconstruct the data and some L1 penalty to encourage sparsity. It's a joint optimization problem.
是的,没错。我们试图在这方面耍点小聪明。字典学习方面有丰富优美的文献。但几个月后,苦涩的教训降临了,结果发现你只需使用一个单层稀疏自编码器,放在 torch 里,然后在 GPU 上训练,Scaling(规模扩张)再次比耍聪明更重要。所以就是这样。这是一种稀疏自编码器。
Yeah, that's right. We tried being clever about this. There's a beautiful rich literature on dictionary learning. But after a few months of that, the bitter lesson got us, and it turned out that you could just use a one-layer sparse autoencoder, which you can put in torch and just train on your GPUs, and that scaling was more important than being clever yet again. So that's what it is. It's a sort of sparse autoencoder.
好的。这里有一个提示:‘包含达拉斯的州的首府是奥斯汀。’这是真的。这是因为包含达拉斯的州是德克萨斯州。很好。模型是这样做的吗?是像‘德克萨斯州,所以奥斯汀’这样,还是像‘我不知道’?看起来很多训练数据,比如这就在里面,它只是在背诵答案?有很多评估污染,比如很多 MMLU 进入了人们的训练集,所以得分很高。好吧,它只是字面上知道答案,因为它见过。它真的以前见过这个,还是像你一样在思考德克萨斯州?
All right. So here's a prompt: 'The capital of the state containing Dallas is Austin.' This is true. That's because the state containing Dallas is Texas. Good. Did the model do that? Was it like Texas, so Austin, or is it like I don't know. It seemed a lot of training data, like was this just in it and it's just reciting the answer? There's a lot of eval contamination, like a lot of MMLU made it into people's training sets, so you get high scores. Well, just knew the answer to that literally because it had seen it. Did it literally see this before, or is it kind of thinking Texas like you did?
它在思考德克萨斯州。所以我要告诉你,这就像一个卡通图,我们可以慢慢剥离这些抽象概念,看看我们实际做了什么。但我要先说明,你应该把每个这样的东西看作一组特征,这些特征又是我们通过优化过程学习到的原子。我们通过观察每个特征何时被激活,并试图用语言描述该特征在做什么,来标记其作用。然后它们之间的连接实际上是模型在前向传播过程中的直接因果连接。所以我们做的是把模型分解成碎片。我们问哪些碎片是活跃的,我们能否分别解释它们,然后信息如何流动?在这个案例中,我们发现了一堆与首都相关的特征,一堆与抽象意义上的州相关的特征,一堆与达拉斯相关的特征。我们发现了一些类似运动神经元的特征;它们让模型做某事。在这种情况下,它们让模型说出首都的名字,比如让模型说出一堆州首府或国家首都。这是一个开始。但它还必须答对。所以从达拉斯到德克萨斯州有一些映射,那里有一堆特征。有些像关于德克萨斯州政治的讨论,有些像‘德克萨斯州的一切都更大’这样的短语。一旦你有了德克萨斯州和首都,你就会特别说出奥斯汀。如果你在说一个州首府,同时也在说奥斯汀,那么结果就是奥斯汀出来。不过也有一些有趣的直线,比如德克萨斯州也直接输入到奥斯汀。如果你在想德克萨斯州,你可能就直接在想奥斯汀。所以这给了你一个关于我们学到的这些原子或字典元素的图景。你可能想检查它们是否有意义。然后你可以对模型进行干预,删除其中的一部分,就像在神经科学中这些会是神经元的消融,然后观察模型的输出是否如你所预期的那样变化。
It's thinking about Texas. So I'll tell you, this is like a cartoon, and we can slowly break away these abstractions to literally what we did. But I'll sort of start by saying that you should think of each of these as bundles of features, which are again atoms we learned through an optimization process. We label the role of each feature by looking at when it's active and trying to see if we can describe in language what that feature is doing. And then the connections between them are actually direct causal connections as the model processes this in a forward pass. So what we do is we break apart the model into pieces. We ask which pieces are active, can we interpret them separately, and then how does it flow? In this case, we found a bunch of features related to capitals, a bunch related to states in the abstract, a bunch related to Dallas. We found some features that are like motor neurons; they make the model do something. In this case, they make it say the name of a capital, like they make it say a bunch of state capitals or country capitals. So that's a start. But also it has to get it right. So there's some mapping from Dallas into Texas, and there's a pile of features there. Some are like discussions of Texas politics, some are like phrases like 'everything's bigger in Texas.' And once you have Texas in a capital, you get saying Austin in particular. And if you're saying a state capital and also you're saying Austin, what you get is Austin coming out. There are some interesting straight lines though, like Texas also feeds into Austin directly. If you're thinking about Texas, you might just be thinking about Austin. So this gives you a picture in terms of these atoms or dictionary elements we've learned. And you might want to check that they make sense. So then you can do interventions on the model, deleting pieces of this like in neuroscience these would be ablations of neurons, and then you see if the output of the model changes as you would expect.
Zoom 上有人问:像奥斯汀这样的词经常在互联网文本中一起出现,表明简单的统计相关性。所以这不是把事情复杂化了吗?
Someone on Zoom asked: words like Austin frequently appear together in internet text, suggesting simple statistical correlations. So isn't this a way of overcomplicating things?
我认为人们已经发现 Transformer 在大多数任务上表现更好。所以,你知道,休斯顿也在训练数据中与达拉斯一起出现,但模型没有说休斯顿,所以它一定是在使用首都和达拉斯。我认为你可以说,也许如果你只有首都和达拉斯,它就会说奥斯汀,这实际上是正确的。我认为其中一些边很弱,结果发现你只需说‘达拉斯首都是’,它会产生一些不合语法的东西,就像你知道的‘达拉斯首都是’,它会说奥斯汀。这是一个有趣的事情,我认为你看这个图,这条边很弱,它表明实际上也许就像如果你有首都和达拉斯接近,你可能得到奥斯汀,这在因果上是成立的。
I think that people have found that transformers have outperformed back then for most tasks. So, you know, Houston also occurs with Dallas in the training data, and the model doesn't say Houston, so it must be using the capital and Dallas in this case. I think you could say maybe if you just had capital and Dallas, then it would say Austin, and that's actually true. I think some of these edges were pretty weak, and it turns out you could just say 'the capital of Dallas is' and it makes something ungrammatical, just like you know 'the capital of Dallas is' and it will say Austin. And that's an interesting thing where I think you look at the graph and this edge is weak, and it indicates that actually maybe it is just like if you have capital and Dallas in proximity, you might get Austin, which is then causally true.
是的,层,层,层。你想做层和底层。是的,我理解正确。所以这是通过模型各层的流动。楼上。我们在这里做了一些分组。
Yes, layer, layer, layer. You want to do the layer and the bottom layer. Yeah, I understand this correctly. So this is the flow through layers of the model. Upstairs. And we've done some grouping here.
我会更详细地展示。让我看看我们什么时候讲到那里。
I will show in more detail. Let me actually see when we're going to get to that.
接下来我会在幻灯片上给出更多技术细节。是的,关于字典大小。我们扫描了不同大小的字典,训练后发现,在计算量和近似精度之间存在权衡——这些字典本应重建激活值,效果如何呢?字典越大,效果越好。同时,字典越密集(即稀疏度越低),效果也越好,但可解释性会有所牺牲。所以我们做了大量扫描,最终选了一个看起来足够好的。仍然存在不少误差,稍后我会展示这些误差如何导致我们无法解释很多事情。好,我继续。我们实际做的是训练这个稀疏替换模型。基本思路是:如果有一个模型,残差流经过 MLP 层,我们先忽略注意力机制,尝试用这些跨层转码器来近似。转码器的作用是移动信息,模拟 MLP,而跨层转码器集合则模拟所有 MLP。几年前有一种叫 DenseNet 的架构,每一层都写入后续所有层,这里类似。基本想法是,这个 CLT 集合必须接收所有 MLP 的输入堆叠,并一次性产生它们所有输出的向量。这样做的原因是,没有理由认为计算的基本单元必须局限在一层内。在深层模型中,连续两层的神经元几乎可互换。有人做过实验,交换 Transformer 层的顺序对性能影响不大,这意味着过分依赖现有层的索引可能不明智。所以我们说,这些可以跳到末尾。这使可解释性更容易,因为有时只是二元语法统计:我看到这个词,我说那个词,这在第一层就很明显,但你必须一直传播它才能输出。这里我们可以把它变成一个特征,而不是连续层中数十个特征的交互。然后我们训练这个优化,有精度损失和稀疏度损失。我们用基础模型的注意力,不试图解释注意力,只是让它流过。但我们试图解释 MLP 在做什么。现在,神经元通常难以解释,但左边这些更有意义。右边是一个大写特征,可能更具体,比如字面意义上的州到首府的映射。
So I'm gonna give you a lot more detail in just about a slide on the technical side. Yes, per size for a dictionary or for the dictionary. We did a scan where we trained dictionaries of different sizes and we, you know, you get some tradeoff of just the compute you spend and the amount or the accuracy of the approximation, like you know these are supposed to reconstruct the activations. How well does that happen? The bigger the dictionary, the better it does. Also, the denser it is, the less sparse it is, the better that does, but you pay some price on forability at some point. So, we did a bunch of sweeps and then we picked something that seems good enough. There's still a bunch of errors and, you know, I'll show you later how those show up to, you know, mean that we can't explain, you know, a lot of things. Okay, I'm going to go on. Um, so yeah, literally what we do here is train this sparse replacement model. So the basic idea: if you have a model, the residual stream goes through MLPs. We're going to forget about attention right now, um, and we're going to try to approximate that with these cross-layer transcoders. And so what do I mean by that? So transcoder is something that, like, you know, moves information so it emulates the MLPs, but the cross-layer is this ensemble of them emulates all the MLPs. There's an architecture called DenseNet from a few years ago, right, where every layer writes to every subsequent layer, and this is like that. So the basic idea is that this ensemble of CLTs have to take in the inputs of all the MLPs stacked together and produce the vector of all of their outputs at once. Um, and a reason to do this is like there's no particular reason to think that the atomic units of computation have to live in one layer. If there's two consecutive layers in a deep model, those neurons are almost interchangeable. People have done experiments. You can actually swap the order of transformer layers without damaging performance that much, which means that indexing that hard on the existing layer is maybe unwise. So we just say, okay, you know, these can skip to the end. Um, this ends up making the interpretability easier because sometimes it's just bigram statistics. I see this word, I say that word, that's evident at layer 1, but you have to keep propagating it all the way through to get it out. And so here we could make that be one feature instead of, you know, dozens of features in consecutive layers interacting. Um, and then we just train that optimization. You've got a loss on accuracy and a loss on sparsity. Um, okay. And so here we replace the neurons with these features. Um, we use just the attention from the base model. So we do not try to explain attention. We just flow through it. But we try to explain what the MLPs are doing. Um, and now instead of the neurons which are sort of uninterpretable, like on the left here, we have stuff that makes more sense. On the right, this is a, say, a capital feature. Um, and I think it's probably more specific. So, this is like, in, you know, these like literal state to capital mappings. Um, yeah.
我们的公司。我不知道,合理地说,我认为这是权宜之计。它损失了很多。从实践角度看,如果你想建模注意力层或任何在 token 间移动信息的操作,那么输入必须是所有 token 位置的激活值。而我们一次只能处理一个 token 位置,所以可以学到更简单的东西。学习问题也容易得多。但替换注意力层的东西是什么,这不太清楚,因为它需要既能转换信息又能移动信息。稀疏性在那里不是好优先项。你得到的是四维张量而不是二维张量,我们不确定正确答案是什么。所以目前我们只做了这个。
Our company. So I don't know, reasonable, I'd say it's expedient. Um, it loses a lot. Um, I think that from a practical perspective though, if you want to model the action of the attention layer or anything that moves information between tokens, then the input would have to be the activations at all of the token positions. Um, and here we only can do one token position at a time. So you can learn something much simpler. Um, and I think the learning problem is just like a lot easier. Um, it's also like slightly less clear what the thing to replace the attention layer is that would be interpretable because it needs to be a system for both transforming information and moving it. Um, and sparsity isn't a good priority there. You've got like a four tensor instead of a two tensor and we aren't sure what the right answer is. So we just did this for now.
你在吸收关于相关特征的信息。对吧?所以有两个问题:一是替换会损失什么,二是你在多大程度上依赖对所得组件的解释,以及什么是相关特征。稀疏自编码器会产生一堆这样的东西,我们必须解释输出。对于任何特定的图,注意力是冻结的,我们通过替换进行前向传播,得到一堆特征。这些三角形和菱形是误差项,因为不能完美重建基础模型,所以必须包含误差项,然后可以直接追踪影响。这些只是线性映射,直到最后。会有很多特征被激活,比神经元少,但每个 token 有数百个。但它们并不都对模型输出重要。所以你可以从输出(比如'Austin')反向追溯,找出哪些特征直接因果相关于输出'Austin',然后哪些特征因果相关于那些特征被激活,以此类推。这样你会得到一个更小的东西,可以实际查看,以理解为什么模型输出了这个字面词。但这仍然是数学,你有了这个图,但如果你想解释它,你需要查看这些单独组件,看能否理解它们。你又回到了查看这个特征在哪些例子中被激活,希望它是可解释的,并且它们之间的连接有意义。通常这些会属于粗略的类别。比如,有一个德克萨斯特征,又是'everything's bigger in Texas',德克萨斯是一个以牛仔和女牛仔闻名的大州。另一个是关于司法系统的政治。我不是说这是分解网络的正确方式,但它是其中一种方式。如果你看这里的流程,这些德克萨斯特征会输入到说'Austin'的特征中,对吧?有一条路径。所以我们会在这一点上基于人工解释手动分组一些特征,以得到一张图。
You are imbibing your information of what the relevant features are. Right? So there's two questions there. One is like what do you lose by doing a replacement and the other is like, you know, um, how much are you leaning on your interpretations of the components you get and what is a relevant feature. Yeah. Something that you... Yeah. Um, well, so the sparse autoencoder just produces a bunch of these, right, and then we do have to interpret what comes out. And so for any particular graph, you know, we can go... we have now the attention is frozen, um, we have a forward pass of the model sort of through the replacement where we've got a bunch of the features. We've got these triangles, these diamonds are the errors, so like these don't perfectly reconstruct the base model, so you have to include an error term, and then you can, you know, track the influence directly. These are just linear maps, um, until you get to the end. Um, and you know, there'll be a lot of these active, fewer than there are neurons, but you know, order order order hundreds per token. Um, but they don't all matter for the model's output. So then you can sort of go from the output like 'Austin' here backwards and say which features were directly causally relevant for saying 'Austin' and then which features were causally relevant for those being active and then which are causally relevant for those being active. And in that way you get a much smaller thing, something you could actually look at, um, to try to understand why it said this literal thing here. Um, and now you're in... but this is all still math right now, you have this graph, but if you want to interpret this then you need to look at, you know, these individual components and see if you can make sense of them. And now you're back to like looking at when was this active in which examples, hoping that's kind of interpretable and that, you know, the connections between them make any sense. Um, you know, often these will be in some rough categories. So, here's like one Texas feature is like again the 'everything's bigger in Texas.' Texas is a big state known for its cowboys and cowgirls. Um, and this other one is about, you know, politics in the judicial system. Um, and it, you know, I'm not saying that like these are like the right way to break apart the network, but they are a way to break apart the network. And if you look at the flow in here, you know, these Texas features feed into the say 'Austin' ones, right? There is a path here. And so we sort of will manually group some of these based on at this point human interpretation, um, to get a map of what's going on.
是的,Zoom 上有个问题。在 CLT 架构中,我们冻结注意力块,用相互通信的转码器替换 MLP。这很复杂,对吧?当底层连接可能简单得多时,这不会增加不必要的可解释性吗?另外,你能重复一下问题让 Zoom 上的人听到吗?是的。
Yeah, just a Zoom question. Um, in CLT architecture, we freeze attention blocks and replace um MLPs with transcoders that speak to each other. Um, this is complex, right? So won't this be adding unnecessary interpretability when the underlying connections could be a lot simpler? Also, do you mind repeating the question so the folks on Zoom can hear? Yes.
问题是,这看起来工作量很大。这些特征之间有很多连接,由注意力机制调节。我同意。但我们找不到一种方法既能减少工作量又能让单元可解释。一种思考方式是,模型的基础组件本身不太可解释。如果你想将事物解释为组件的交互,比如采用有效的还原论策略——将器官分解为细胞并理解细胞如何相互作用——你必须将其分解。这样做会损失很多,但你能获得谈论这些部分如何工作的能力。所以这是我们今天研究部件的最佳猜测。一旦我们将事物分组为这样的图,我们就可以进行干预。这些干预是在基础模型上进行的,而不是在我们复杂的替代模型上。它们相当于将特征输出向量加到另一个向量上。你可以看看扰动是否有意义。如果我们通过静音来替换掉德克萨斯特征,并加入来自另一个提示的加利福尼亚特征,模型会说萨克拉门托。如果你加入佐治亚,它会说亚特兰大。如果你加入拜占庭帝国,它会说君士坦丁堡。这表明我们以一种可操控的方式捕捉到了抽象的东西。你不会输入它却得到胡言乱语。如果你只是做二元组统计,不清楚如何获得这种可分离性。
The question is, this seems like a lot of extra work. There are many connections between these features mediated by attention. I agree. But we couldn't find a way to do less work and have the units be interpretable. One way to think about this is the base components of the model aren't that interpretable. If you want to interpret things as interactions of components, like using the reductionist strategy that's been effective—breaking organs into cells and understanding how cells interact—you have to break it apart. You lose a lot when you do that, but you gain the ability to talk about how those parts work. So this was our best guess today for parts to study. Once we have grouped things into a graph like this, we can do interventions. These interventions are in the base model, not in our complicated replacement model. They amount to adding vectors that are the feature outputs to another one. You can see if the perturbations make sense. If we swap out the Texas features by muting them and add in the California features from another prompt, the model will say Sacramento. If you put in Georgia, it will say Atlanta. If you put in the Byzantine Empire, it will say Constantinople. This shows we captured an abstract thing in a manipulable way. You don't put it in and get gibberish. If you were just doing bigram statistics, it's not clear how you would get this separability.
我将介绍我们经常看到的三种模式:医学和多语言语境中的抽象表示;并行处理模式,涉及算术、一些越狱和幻觉;以及一些规划元素。这种训练替代模型和分析归因图的方法只是数学。你可以做任何你想做的事。然后问题是,它们可解释吗?我们在其中一篇论文中在一个小的 18 层模型上做了这个,它做不了太多事,但发现了相当可解释的东西。我认为这在所有规模上都有效。我知道有人在窄用途和小模型上做这个,它仍然有用。既然你在这里谈论生物学,有没有来自真实医学领域的文献重叠可以提供灵感?特别是,进行因果扰动并观察结果的想法非常像神经科学或遗传学。幸运或不幸的是,我们的实验设置比他们的好得多。我们已经远远超过了神经科学所能做到的。我们有一个大脑,可以研究它十亿次,干预一切,测量一切。他们试图捕捉 0.1%的神经元,时间分辨率比实际活动差一千倍。所以也许反过来才对。
I'm going to get into these three motifs we see a lot: abstract representations in a medical context and a multilingual context; parallel processing motifs, which is about arithmetic, some jailbreaks, and hallucinations; and also some elements of planning. This approach of training a replacement model and analyzing an attribution graph is just math. You can do whatever you want. Then it's like, are they interpretable? We did this on a small 18-layer model in one of the papers, which can't do very much, and found pretty interpretable things. I think this works at all scales. I know people doing this on models that are narrow purpose and small; it would still be useful. Since you're talking about biology here, is there any overlap in literature from the real medicine side for inspiration? In particular, the idea of doing causal perturbations and seeing what happens is very neuroscience or genetics-like. Fortunately or unfortunately, our experimental setup is so much better than theirs. We are well past what people can do in neuroscience. We have one brain and can study it a billion times, intervene on everything, and measure everything. They are trying to capture 0.1% of neurons at a time resolution a thousand times worse than actual activity. So maybe it's the other way around.
我想给你看一些例子。在论文中,这就是我们制作这些示意图的基础。每个节点是一个特征,每条边是一个因果影响。这里有一个,它对输出的直接影响主要是说奥斯汀,还有一些德克萨斯相关的东西。它从与德克萨斯相关的事物和州获取输入。这里的艺术在于你现在在进行解释。你跳来跳去查看这些组件,看它们何时活跃,连接什么,并试图弄清楚发生了什么。我将展示的示意图是通过根据共同属性对这些进行分组得到的。我说这些是比如首都,你可以看到它们彼此不同,但都涉及模型在某种上下文中说首都。我们只是把它们堆叠在一起。你应该有多少个特征?我不知道,没有正确答案。这不是一个完美的近似,但你把它分解成 3000 万块,然后以有意义的方式重新组合,然后进行干预来检查你学到的东西是否真实。
I want to show you some of these. In the paper itself, this is what we make these cartoons from. Each node is a feature, and an edge is a causal influence. Here's one whose direct effect on the output is saying Austin mostly and also some Texas things. It's getting inputs from things related to Texas and from states. The art of this is now you're doing the interpretation. You bounce around looking at these components, looking at when they're active, what they connect to, and trying to figure out what's going on. The cartoons I'll show you are given by grouping sets of these based on common properties. I said these were say a capital, and you can see these are different from each other but they all involve the model saying capitals in some context. We just piled those on top of each other. How many features should you have? I don't know, there's no right answer. It's not a perfect approximation, but you break it into 30 million pieces, then put the pieces back together in ways that make sense, and then do interventions to check that what you learned is real.
现在我们有了系统和涌现属性。这可能是研究这些系统最大的困难。LLM 开始展示涌现行为。如果你分解它,你就会失去那些涌现属性。你如何平衡这一点?
Now we have systems and emergent properties. This is maybe the biggest difficulty in studying these systems. LLMs start displaying emergent behavior. If you break it down, you end up losing those emergent properties. How can you balance that?
这是个好问题。LLM 的一个不同之处在于信息从输入到输出的流动,其中潜在空间在你通过时以更高层次的表示或复杂性运作。所有细胞在某种程度上大小相同;它们相互通信,但都是横向通信。你在大脑中也能看到这一点,从最初的光敏感细胞到视觉皮层,最终你会得到一个对特定面孔敏感的细胞。这来自于对不那么具体的事物敏感的东西,最终是检测边缘和形状的东西。当你通过时,你会得到更高层次的抽象。当我说有一个特征似乎对应于代码中的错误时,那是我们的原子单元之一。这是我们分解时得到的东西之一——以一种非常普遍的方式对代码错误敏感。在某种程度上,这些是分层构建的,但它们可能仍然是单元。
It's a good question. One thing that makes LLMs different is the flow of information from input to output, where latent spaces work at a higher level of representation or complexity as you move through it. All cells are somehow the same size; they communicate with each other, but it's all lateral communication. You see this some in the brain as you go from the first light-sensitive cells through the visual cortex, where you ultimately get a cell sensitive to a particular face. That comes from things sensitive to less specific things, ultimately things that detect edges and shapes. As you move through that, you get higher levels of abstraction. When I say there's a feature that seems to correspond to errors in code, that's one of our atomic units. It's one of the things we get when we shatter this—sensitive to errors in code in a very general way. To some extent, these are built up hierarchically, but they might still be units.
还有一件事我认为我们还没有任何进展,那就是如果它在做某种上下文中的动力系统之类的事情,我认为那会难理解得多,你可能需要更大的集成。好了,再放一下幻灯片。这里有一个医学考试风格的鉴别诊断问题。一位 32 岁女性,妊娠 30 周,出现严重的右上腹疼痛等。在这种情况下,如果我们只能再问一个症状,应该问什么?有人知道这是什么疾病吗?是的。所以问视觉障碍,原因是最可能的诊断是子痫前期,这是一种严重的妊娠并发症。我们可以通过观察这些特征以及它们如何分组和流动来追踪。这从妊娠、头痛、高血压、肝功能检查中提取信息,指向子痫前期,然后从那里到子痫前期的其他症状,最终有两个完整的视觉障碍。这里有一个小框显示我们的意思:这个组件在讨论视力丧失、飞蚊症的训练样本中也是活跃的,这是另一种导致视力丧失的视觉障碍。另一个答案是蛋白尿。所以它在某种程度上是在说出这些是什么之前思考子痫前期,我们可以看到那个中间状态。它也在考虑其他可能的诊断,比如胆道系统疾病,如果我们抑制子痫前期,那么它会说食欲减退,因为场景提示胆道疾病。所以它在思考这些选项;当你关掉一个,你会得到一个与另一个一致的连贯答案。
There's another thing that I think we don't have any traction on, which is if it's doing some in-context dynamical systems stuff, that I think will be much harder to understand, and you might need much larger ensembles of these. Okay, let's slideshow again. So, here's a medical exam style differential diagnosis question. A 32-year-old female at 30 weeks gestation presents with severe right upper quadrant pain, etc. Given this context, if we could only ask for one other symptom, what should it be? Does anybody know what the medical condition is here? Yeah. So ask about visual disturbances, and the reason is because the most likely diagnosis is preeclampsia, which is a severe complication of pregnancy. We're able to trace through by looking at these features and how they group together and flow. That is pulling things from pregnancy, headache, high blood pressure, liver tests, to preeclampsia, and then from that to other symptoms of preeclampsia, and ultimately there are two complete visual disturbances. There's a tiny box here showing what we mean by that: this component is also active on training examples discussing vision loss, floaters in the eye, which is another visual disturbance leading to loss of vision. The other answer would be proteinuria. So it is thinking about preeclampsia in some way before saying what these are, and we can see that intermediate state. It is thinking about other potential diagnoses like biliary system disorders, and if we suppress the preeclampsia, then it will say decreased appetite because the scenario suggests biliary disease. So it is thinking of these options; when you turn one off, you get a coherent answer consistent with the other one.
这意味着什么?你可以如此清晰地干预,这仅在网络中的一个位置表示,并且通过干预你不需要多次干预?
What does this mean? The fact that you can intervene so clearly, that this is only represented in exactly one location in the network, and that by intervening you don't need to intervene multiple times?
嗯,不,有两三个原因。第一,这个节点是一组特征,对吧?它是几个与子痫前期相关的特征。第二,因为这是一个跨层转码器。每个特征都写入多个层。第三,我们是在过度抑制。这里我们把它关掉的程度是它开启程度的两倍。我们过度纠正了。你经常需要这样做才能获得完整效果,可能是因为存在冗余机制。所以我不认为这是唯一的地方,但它足够了。
Well, no, for two or three reasons. One, this node is a group of them, right? It's a few of these features that are all related to preeclampsia. The second is because it's a cross-layer transcoder. Each of them writes to many layers. And the third is we're creaming it. Here we're turning it off double the amount that it was on. We overcorrect. You often need to do that to get the full effect, probably because there are redundant mechanisms. So I don't think this is the only place it is necessarily, but it's enough.
在你之前的演讲中,你谈到了一些关于神经元的多元语义性。例如,如果你做这样的干预,比如子痫前期的那个,假设我把模型用于其他任务,你预计这还会影响到哪些东西的分布?你能给出一些其他类型的东西的感觉吗?
Earlier in your talk you talked a bit about polysemanticity of neurons. So for example, if you do one of these interventions like the preeclampsia one, and let's say I deploy the model for some other task, would you expect what is the distribution of things that this will also hit? Can you give some sense of what other types of things?
这是个好问题。我们在这里没有深入探讨;我们在上一篇论文中做得更多一些。我认为如果你推得太用力,模型会大幅偏离轨道。这里的目的不是塑造模型行为,而是验证我们在这一例子中的假设。所以我们不是像'现在拿一个模型,到处删除这个'。而是模型正在回答这个问题。我们关掉它大脑的这一部分,然后看看会发生什么。
That's a great question. We didn't dig into that as much here; we did a little more with our last paper. I think that if you push too hard, the model goes off the rails in a big way. Here this is less for the purpose of shaping model behavior and more for validating our hypothesis in this one example. So we were not like 'now let's take a model and delete this everywhere.' It's like the model is going through answering this question right here. We're turning off this part of its brain and then we're seeing what happens.
Zoom 上有几个问题。一个是:在实践中,当你向后追溯这些特征时,随着你回到模型更早的层,你会看到相关特征爆炸式增长,还是与输出距离的函数有关?
A couple questions on Zoom. One is: in practice, when you trace backwards through these features, do you see a kind of explosion of relevant features as you go earlier in the model, or is it some function of distance from the output?
是的,问题是当你向后追溯模型时,相关特征的数量会增加吗?是的。有时你会得到收敛。早期的特征可能关于几个症状之类,那些很重要。我不认为我们做了一个好的图表来回答这个问题。我要继续讲一点,很快会回来回答问题。
Yes, the question is as you go backwards through the model, does the number of relevant features go up? Yes. Sometimes you get convergence. The early features might be about a few symptoms or something, and those are quite important. I don't think we've made a good plot to answer that question. I am going to push onwards a little bit and I'll come back to questions soon.
再一个。这些例子有多少是精心挑选或设计的?这对大多数句子都有效吗?
Just one more. How much are these examples carefully selected or crafted? Does this work for most sentences?
你对你尝试的大多数东西都能学到一些东西。我估计我们尝试的提示中大约 40%,我们可以看到一些非平凡的部分。这不是全貌,有时也不起作用。但这也只需要用户 60 秒的时间来启动一个,然后你等待,它返回,你就能学到一些关于模型的东西,而不必构建一大堆先验假设。一旦你构建了机器,你就能免费得到它。但这并不能告诉你一切。
You learn something about most things you try. I'd say about 40% of the prompts that we try, we can see some non-trivial part of what's going on. It's not the whole picture and sometimes doesn't work. But it also just takes 60 seconds of user time to kick one of these off, and then you wait and it comes back and you get to learn something about the model rather than having to construct a whole bunch of a priori hypotheses. You just sort of get it once you've built the machine; you get it for free. But this doesn't tell you everything.
好的,这里有一个问题:模型用什么语言思考?是通用语言吗?是英语吗?Claude 内部有一个法语 Claude 和一个中文 Claude,谁被激活取决于问题?为了回答这个问题,我们看了三个句子,它们是三种语言中的同一个句子。所以'小的反义词是大'——我不会说普通话。但也许这里有人能说出最后一个字是什么。谢谢。我来看它是因为我们做了这篇论文的这一部分很多次,以至于它是我现在唯一能认出的字。所以这甚至比上一个更卡通化。我稍后会进入更详细的版本,但基本上我们看到的是,一开始有一些特定语言中的这些反义词,比如法语中的'contra',英语中的'opposite'。而引号就像是我在说的语言中的左引号。所以后面的引号可能也是同一种语言。
Okay, so here's a question: what language are models thinking in? Is it a universal language? Is it English? Is there a French Claude and a Chinese Claude inside of Claude, and who gets gated depends on the question? To answer this, we looked at three sentences which are the same sentence in three languages. So 'the opposite of small is big' — and I can't speak Mandarin. But maybe someone here can say what the final character is. Thank you. I came to see it because we were doing this part of the paper so many times that it's the only character I can recognize now. So this is even more cartoony than the last one. I'll get into the more detailed version later, but basically what we see is at the beginning there are some of these opposites in specific languages, like 'contra' in French, 'opposite' in English. And the quotation is like an open quote in the language I'm speaking. So the quote that follows will probably be in the same language.
但接着有一个关于反义词的复杂特征集:许多语言中的“大”、许多语言中的“小”组合在一起,然后你把它用目标语言输出。所以这个说法是,这里有一个多语言核心,其中“小”加上“相反”得到“大”,然后“大”加上这个英文引文得到单词“largeness”,加上这个法文引文得到……我们做了一堆这样的修补实验。你可以说,不用反义词,用同义词替换进去。所以所有情况都一样,现在你会得到法文的“little”或“minuscule”。所以修补不同的状态,你在三个位置放入同一个特征,后面就会得到相同的行为变化。我们研究了这一点,它对规模有很大影响。你可以观察,随着你深入模型,一对互为翻译的句子之间有多少特征重叠。一开始,只有那种语言的词元,没有重叠。中文分词后与英文分词毫无共同之处。所以开头和结尾都没有重叠。但随着你深入模型,你会发现很大一部分活跃的组件是相同的,无论输入的是哪种语言。这是英中、法中、英法。如果句子不相关,这是一个随机基线。所以不仅仅是中间更常见。但当比较互为翻译的句子对时,这是在 18 层模型和我们的一个小型生产模型中。所以这种泛化随着规模增加而增强。
But then there's this complex of features about antonyms in many languages: 'large' in many languages, 'smallness' in many languages that go together, and then you spit it back out in the language of interest. So the claim is that this is a sort of multilingual core where smallness goes with oppositeness to give you largeness, and then largeness plus this quote in English gives you the word 'largeness', plus this quote in French gives you... We did a bunch of these patching experiments. You could say instead of the opposite, you could say a synonym and just drop that in. So the same thing in all of them, and now you'll get 'little' or 'minuscule' in French. So patching in a different state, you can put the same feature in these three places and get the same change in behavior later. We looked at this and it does have a big effect with scale. You can look at how many features overlap between a pair of sentences that are translations of each other as you move through the model. At the beginning, it's just the tokens in that language. There's no overlap. Mandarin when tokenized has nothing in common with English when tokenized. So at the beginning and the end there's nothing. But as you move through the model, you see a pretty large portion of the components that are active are the same regardless of the language. This is English-Chinese, French-Chinese, and English-French. This is a random baseline if the sentences are unrelated. So it's not just that the middle is more common. But when you compare pairs that are translations, this is in an 18-layer model and in our small production model. So this generalization is increasing with scale.
这几乎就像它能解释概念……你有没有试过那些特定于某种语言、没有真正翻译的隐喻提示?
It's almost like it can interpret concepts in a... Have you ever tried prompts that are metaphors specific to a language that don't really have a translation?
没有。如果你有例子,我很乐意试试。我觉得那会很有趣。我很乐意合作。
No. If you have some examples, I'd love to try it. I think that would be fun. I would be happy for collaboration.
那么,你是否认为这意味着网络中心是任何给定概念的最抽象表示?
So, would you say this implies that in the center of the network is the most abstract representation of any given concept?
我基本同意这个图,它稍微偏向中心右侧。在这一点上,你开始需要弄清楚如何处理它,因为最终模型必须说出一个东西。你有没有发现,在非常抽象的网络中间部分比在非常具体的末尾部分更难找到自动语义特征?
I would basically agree with this plot, which is a little to the right of the center. At which point you start to have to figure out what to do with it because in the end the model has to say a thing. Have you seen a harder time finding auto-semantic features in the middle of the network where it's very abstract rather than towards the end where it's very concrete?
实际上恰恰相反。如果你允许我哲学一点,好的抽象的意义在于它适用于许多情况。所以中间实际上应该有这些共同的抽象。而处理这个措辞或语法场景的非常具体的细节则相当定制化,所以你需要更多的特征来解析它。
It's actually the opposite. If you'll permit me to be philosophical, the point of a good abstraction is that it applies in many situations. So in the middle there should actually be these common abstractions. It's dealing with the very particulars of this phrasing or grammatical scenario that is quite bespoke, so you would need way more features to unpack it.
还有一个问题。相同的操作是否在多个层中冗余存储?例如,如果一个例子需要五个推理步骤,另一个需要八个,你是否必须在两个层中都冗余存储后续操作?
One more question. Is the same operation stored redundantly across multiple layers? For example, if you need five reasoning steps to get something done in one example but eight in another, do you redundantly have to store the follow-up operation on both of those layers?
是的,我认为这是一个非常深刻的问题。为听众重复一下:我们是否看到相同操作在许多地方冗余?这是人们对这些模型抱怨的一点。如果它知道 A 也知道 B,为什么不能连续在脑子里做出来?可能只是因为 A 在 B 后面,它必须做一个前向传播。所以除非它能思考出声,否则它根本无法组合这些操作。你可以很容易地看到这一点:如果你问“1921 年那部电影主演的父亲生日是什么?”,它可能能单独完成所有这些查找,但无法连续完成。所以必须有冗余。我们确实看到了这一点。这也是交叉编码器设置的原因之一,试图压缩一些冗余。我最喜欢的一个图来自姊妹论文。左边是按层分解,右边是跨层的东西。基本上,它只是在来回反弹。“哥本哈根”被传播,与丹麦对话,在模型的许多许多地方发生。它们都对其进行了小的改进。还有另一个视角:神经 ODE、梯度流视角,即所有时间都是同一方向的微小调整。我认为真实模型介于两者之间。
Yeah, I think this is a very deep question. To repeat it for the audience: do we see redundancy of the same operation in many places? This is something people complain about with these models. If it knows A and knows B, why can't it do them in a row in its head? It might just literally be that A is after B in its head and have to do a forward pass. So unless it gets to think out loud, it literally can't compose those operations. You can see this easily: if you ask 'what's the birthday of the father of the star of the movie from 1921 starring...', it might be able to do all of those lookups individually, but it can't actually do them consecutively. So there has to be redundancy. And we do see this. It was one of the reasons for the crosscoder setup, to try to zip up some of that redundancy. One of my favorite plots is from the sister paper. On the left is the decomposition per layer; on the right is the cross-layer thing. Basically, it's just bouncing back around. 'Copenhagen' is propagated, talking to Denmark, and it happens in many, many places in the model. They all perform small improvements to it. There's another perspective: the neural ODE, gradient flow perspective, where it's all tiny adjustments in the same direction all the time. I think the real models are somewhere in between.
随着特征重叠,当模型变大时,如果你显著改变数据集或数据集的顺序,是否最终会得到非常不同的表示?
With the overlap of features here, as models get larger, if you change the dataset significantly or the ordering of parts of the dataset, do you end up with very different representations?
我们没有对数据集顺序效应进行系统研究。有时我好奇你是否在……不,我们可能很幸运。你可以画一个图,显示来自归因图的科学假设有多有说服力,以及干预时效果如何。最好的那些确实有效。但我们也讨论了一些小的失败案例。
We haven't done systematic studies of dataset ordering effects. Sometimes I'm curious if you're around... No, we've maybe been lucky. You could have a plot of how compelling a science hypothesis was from the attribution graph and how well it worked when you intervened. The best ones did work. But there are some small failure cases we talk about.
我想深入探讨并行主题,因为它是 Transformer 架构的一个独特特征。它是大规模并行的。举个简单的例子,假设你想把 100 个数相加。最简单的方法是从一个开始,加下一个,再加下一个到总和,做 100 个串行步骤。用 Transformer 做这个,它需要 100 层深。但还有另一种方法:在一层中连续加对,然后在下一层加对的对,依此类推。所以在对数深度内,你可以把 100 个数加起来。
I want to dive into the parallel motif because it's a unique feature of the transformer architecture. It's massively parallel. To give a simple example, imagine you want to add 100 numbers. The easiest way is to start with one, add the next, add the next to the sum, and do 100 serial steps. To do that with a transformer, it would need to be 100 layers deep. But there's another thing you can do: add pairs consecutively in one layer, then add the pairs of pairs in the next, and so on. So in log depth you could add up the 100 numbers.
嗯,考虑到这里的深度限制以及我们要求这些模型完成的任务的复杂性,它们会尝试同时做很多事情是有道理的。预先计算你可能需要的东西,然后把它们拼凑在一起,我来举几个例子。所以,如果你让模型计算 36 加 59,它会说 95。如果你问它是怎么算的,它会说用了标准的进位算法,但这不是真的。实际发生的情况更像这样:首先,它会解析每个数字,比如有一个组件专门处理“59”这个具体数字,但也有针对所有以 9 结尾的数字的组件,还有针对那个数字范围的组件,上面也是类似的。然后你有两条处理流:一条是获取最后一位数字,另一条是获取数量级。即使在数量级内部,也有一种窄带数量级和一种宽带数量级。然后这些会给出一个中位数带,如果和在这个范围内并且以 5 结尾,那么实际上就是 95。这样就缩小了范围,然后给出答案。这很酷,但不是我通常的做法。不过话说回来,它并不是由老师教它怎么做,而是在训练中每次出错时被惩罚,或者每次正确时被奖励。
Um and you know given the depth constraints here and the sophistication of what we ask these models to do you know it makes sense that they would attempt to do many things at once. Preemptively compute things you might need, kind of slap them all together and I'll give you a few examples of this. So um if you ask the model to add 36 and 59 it will say 95. If you ask it how it did that it will say it used the standard carrying algorithm which is not true. Um what happens is somewhat more like this. So, first it parses out each number into like there's some component where it's like literally 59, but there's also something for all numbers ending in nine and something which is like that number range and the same up there. And then you kind of have two streams. One where it's um getting the last digit, right? And then another on top where it's getting the magnitude, right? And even inside the magnitude, there's sort of like like a narrow band magnitude and a really wide band magnitude. Um and then um those give you a sort of median band and then if the sum is in this range and it ends in a five then it's actually 95. That narrows it down and then it gives you the answer. Um which is cool. It's not how I would do it. Um but but then again it wasn't trained by like a teacher being like here's something to do. It just got like whacked every time it got it wrong. um or like rewarded every time I got it. Right. Right. In training.
嗯,我觉得我们接下来不会讲这个。所以我想展示加法部分我最喜欢的一个东西。嗯,这里有一个特征。如果你们有人觉得这像是一个文字版的形状旋转测试,那对于喜欢数学思维的人来说,我很喜欢这一部分。这是我们为可视化算术提示而制作的图表。这是关于提示 a+b 中 a 和 b 在 1 到 100 范围内时是否激活的网格。垂直线表示当第二个操作数在某个范围内时激活。这些点表示是否是 6 和 9?所以有一个网格。这些带,一条带是 x+y 等于常数的线,对吧?这些就是和所在的位置。我们看这个是为了弄清楚这些特征的作用。但我特别喜欢这个特征。图表中的每个元素都可以悬停查看,这很酷。我们查看了数据集中这个特征被激活的案例。在这个狭窄的领域内,它只在以 6 结尾的数加上以 9 结尾的数时激活,但在数据集中,它在其他很多情况下也激活。比如这个“thymus organ fragments federal proceedings volume 35”。如果这个解释正确的话,那在某种意义上,9 加 6 就是卷 35。这里有一串数字,还有更多期刊,还有一些坐标。所以,如果我们的方法有效,那么这里有一个组件,在这个上下文中表示“以 6 结尾加以 9 结尾”,但它也在这些情况下激活。这意味着,如果这些例子中模型都在秘密地做 6 加 9,并且重复使用那个模块,那么这个方法就真的有效。我们深入研究了,但我一开始不太理解。这里有一个例子,这个 token 处特征被激活了。我把它放进 Claude 里问这是什么,它说这是一张天文测量表,并以格式良好的表格输出,前两列是观测期的开始和结束时间。这是它预测的结束时间的分钟数。如果你往下看表格,开始到结束的间隔大约是 38 分钟、37 分钟,但在实验过程中逐渐增加到接近 39 分钟。这个测量间隔开始于一个以 6 结尾的分钟,6 加 9 等于以 5 结尾。所以模型只是在做下一个词预测,它从训练数据中学会了识别这些模式。但它需要一个算术表,比如 6 加 9,你只需要查表。所以它在某个地方有那个表,并在一个非常不同的上下文中使用相同的查找来执行加法。另一个例子是,这原来是一个表格,它在预测一个金额,这些是算术序列。我想这是总成本在上升。增加的金额是 9,000 加 26,000 等于 35,000。也许我最喜欢的是,为什么它在这里激活来预测年份?答案是,因为这是卷 36。这本期刊创刊于 1960 年,所以第 0 卷应该是 1959 年。9 + 59 + 36 等于 95。所以它用了同样的加法模块。所以,当我们谈论泛化和抽象时,这对我来说是一个相当感人的例子,让我觉得它确实学到了这个小东西,然后在各处使用它。好了,这可能不是世界上最关键的事情。那么我们来谈谈幻觉。
Um and you know, I don't think we're going to have this next in here. So, I want to show you one of my favorite things from the addition section. Um which is like, okay, so you know, there's a feature in here. Okay. Like, like if anybody here like feel like this is like a word cell shape rotator test. So for the shape rotators who like kind of mathematical thinking like I love this section. So this is like this is the graphs we make to visualize on the arithmetic prompts. Um uh this is like is it active on the prompt a plus b for a and b and 1 to 100. That's a grid. And so vertical lines mean like when the second operand is in a range it's active you know. Um these dots is like is it a six and a nine? You know so there's a grid. this bands, you know, a band here is a line x plus y equals constant, right? And so those are where the sum is. So we were looking at this to sort of figure out what these did. Um, but this feature here I really like. So everything in the graph you can hover over, which is kind of neat and see the feature. Um, and we looked at cases in the data set when this thing was active. Okay, so so on this narrow domain is active when when things ending in six get added to things ending in nine, but on the data set it's active in like all these other cases. So this is like um thymus organ fragments federal proceedings volume 35. It's like okay so in some sense that has to be a 9 plus a 6 is the volume 35 supposedly if this interpretation is correct. Here's just like a list of numbers. Um there's more journals. Uh there's like these coordinates. Um and so and so the claim here if our method is working is that there's one component that in this context means ends in 6 plus ends in 9 but it's also active in these. So like this is really working if secretly every one of these examples is the model adding a six to a 9 and it's getting to reuse the module for doing that across those examples. Um and so we dug in and um I couldn't really understand these. So, so here's one example um where this is the token where where you know that feature was active um and uh I just put it in claw and I was like what is this and it's like ah this is a table of astronomical measurements and it spit it out in like a nicely formatted table and the first two columns are the start and end time of an observation period. Um and um this is the minutes of an end time that it's predicting. And if you read down the table, the start to end interval is like 38 minutes, 37 minutes, but it creeps up over the course of the experiment to be like just under 39 minutes. And this measurement interval started um at a minute with a six and the six plus the 9 equals ending in a five. And so the model was just like gutting out, you know, next token predictions for like arbitrary sequences of data was trained on and um learned in that, you know, of course to recognize what controll supposed to do. But then it needs a bit where it has like the arithmetic table, right? like 6 plus 9. You just got to look it up. So, it has that somewhere and it's using that same lookup in this like very different context where it needs to add those things. Um, this was another one where it turned out this was a table and it's predicting this amount and these are um arithmetic sequences. This is I guess the um total cost that's going up. Um and so um the amount that it's going up by, you know, is this where you're carrying that's about 9,000 plus 26,000 is 35,000. This was maybe my favorite where like why is it firing here to predict the year? And the answer is because it's volume 36. This journal was founded in um first edition was in 19 um 60. So the zeroth would have been in 1959. 9 + 59 + 36 uh is 95. And so it's using the same little bit to do the addition there, right? Um and so I think when we talk about generalization like you know you know and abstraction this is was for me like a pretty moving example where I was like okay like did learn this little thing but then using it all over the place. Um okay that's maybe like not the most like mission critical thing in the world. So let's talk about hallucinations.
嗯,模型很棒,它们总是会回答你的问题。但有时它们会出错。这仅仅是因为预训练。它们被设计成只是预测下一个合理的东西。好的,它应该说点什么。如果它什么都不知道,它就应该给出一个名字。如果它知道一点,它就应该用正确的语言给出一个名字,对吧?如果它知道更多,也许是一个那个时代的常见名字,或者某个篮球运动员之类的。这就是它被训练去做的事情。然后你进行微调,你说:“不,我希望你扮演一个助手的角色,而不是一个通用的模拟器。”当你模拟助手时,我希望助手说“我不知道”,当基础模型的确定性较低时。这是一个很大的转变。所以我们很好奇,这是怎么发生的?这种拒绝猜测的行为是如何通过微调引入的?以及为什么它会失败?这里有两个提示可以说明这一点。一个是:“迈克尔·乔丹从事什么运动?用一个词回答。”它回答“篮球”。另一个是:“迈克尔·巴特金从事什么运动?”这是我们编造的人。用一个词回答。它回答:“我很抱歉。”
Um so so models they're great. They'll always answer your question. Um, but sometimes they're wrong. And that is that is just because of like pre-training. Like they're meant to just predict a plausible next thing. Okay. Um, great. Like it should say something. If it knows nothing, it should just give a name. If it knows like anything, it should give like, you know, a name in the correct language, right? If it knows more, maybe like a common name from the era or like just some basketball player or whatever, right? And so that's what it's trained to do. And then you're like, you go to fine-tuning and you're like, "No, like I want you to be an assistant character, not just a generic simulator." And when you simulate the assistant, I want the assistant to say, "I don't know." Like when the base model certainty is somehow, you know, low. And that's like a big switch to try to make. And so we were curious like how does that happen, right? How does this like, you know, refusal to speculate get fine-tuned in? And then when um why does that fail? Um, and so here's here's sort of two prompts that get at this. One is, "What sport does Michael Jordan play?" Answer in one word. It says basketball. The other is, "What sport does Michael Batkin play?" Um, which just the person we made up. Um, answer in one word. And it says, "I apologize.
好的。这些图有点不同,它们标出了抑制性连接。这在神经科学中很常见,比如抑制。我们画了一些特征,这些特征在这个提示下不活跃,但在另一个提示下活跃,反之亦然。如果是灰色的,表示在这里不活跃。但不活跃的部分原因是它被某个活跃的东西抑制了。我们发现中间有四个特征组成的簇:迈克尔·乔丹,一个在模型识别出问题答案时活跃的特征,一个用于未知名字的特征,还有一个通用的“我无法回答”特征。那个通用特征由助手驱动,在模型回答时始终开启。所以它就像对任何问题都回答“我不知道”,然后当模型回忆起关于某个人的信息时,这个特征会被下调。当迈克尔·乔丹出现时,它抑制了未知名字,提升了已知答案。这两者都抑制了“无法回答”,从而为实际回忆答案的路径留出空间,模型就能说出“篮球”。
Okay. So, here these graphs are a little bit different. They've got suppressive edges highlighted here. This is common in neuroscience, like inhibition. And we've drawn some features that aren't active on this prompt, but are on this one and vice versa. If they're in gray, it means it's inactive here. But part of the reason it's inactive is because it's being suppressed by something that is active. And we found this cluster of four features in the middle. Michael Jordan, then a feature for when the model recognizes answers to questions, another feature for unknown names, and a generic feature for 'I can't answer that.' That generic feature is fueled by the assistant, which is always on when the model's answering. So it's like an always 'I don't know' to any question, and then that gets down-modulated when it recalls something about a person. When Michael Jordan is there, it suppresses the unknown name, boosts the known answer. Both of those suppress the 'can't answer', and that leaves room for the path of actually recalling the answer to come through, and it can say 'basketball'.
回到模型深度,有一个有趣的问题:模型可能需要一段时间才能得出答案,但同时也必须在某个时刻决定拒绝并开始执行,这些是并行发生的。可能会出现不匹配:它必须决定是否回答,但还没有尽力去获得一个好的答案。因此,对于非常困难的问题,可能会出现分歧——它可能仍然知道必须决定“好吧,我觉得我能答出来吗?”这是一个棘手的旋钮;它无法在说出答案前完全自我反思。
Getting back to model depth, there's an interesting problem: it might take a while for the model to come up with an answer, but it also has to at some point decide to refuse and get that going, and those are happening in parallel. You could have a mismatch where it has to decide whether to answer or not, but it hasn't done all it can to get a good answer. So you can get divergence where for very hard questions, it still might know it has to decide 'okay, do I think I'm going to get there?' That's a tricky knob; it can't fully self-reflect on the answer before saying it.
这就是干预:你抑制已知答案,它就会幻觉迈克尔会下棋。这很有趣。如果你问一篇 Andrej Karpathy(前斯坦福)的论文,它会给出一个他没写过的著名论文。为什么?它就像“相信我,我从名字听说过 Andrej Karpathy”,但然后试图回忆论文的部分给出了那个答案。然后如果你问“你确定吗?”它会说“不,我不认为他写了那篇”,因为此时模型同时获得了人和论文作为输入,可以再次在网络的早期进行计算。我们可以操纵这一点:抑制已知答案部分,最终它会道歉并拒绝。论文中有一个有趣的例子:模型没听说过我。那部分不是我写的,是 Jack 写的,模型拒绝猜测我写的论文。然后如果你关闭未知实体并给出已知答案,它就会说我以发明“蝙蝠原理”而闻名,我希望有一天能做到。
This is just the intervention: you suppress the known answer, and it will hallucinate that Michael can play chess. So this is a fun one. If you ask for a paper by Andrej Karpathy, formerly of Stanford, it gives a very famous paper that he didn't write. Why is that? It's like 'trust me, I've heard of Andrej Karpathy from the name,' but then there's the part trying to recall the paper, and it gives that answer. And then if you ask 'are you sure?' it says 'no, I don't really think he wrote that' because at that point the model gets both the person and the paper as input and can do calculation earlier in the network again. Now we can juice this: we can suppress the known answer bit, and eventually it will apologize and refuse. There's a fun one in the paper where the model hasn't heard of me. I didn't write that section; Jack wrote that section, and it refuses to speculate about papers I've written. And then if you turn off the unknown entity and give the known answer, then it says I'm famous for inventing the bats in principle, which I hope to one day do.
还有很多可以聊的。我快速过一遍,让你们感受一下,你们可以读论文,里面有更多内容。有越狱攻击,我们试图理解它们的工作原理。其中一些是让模型在还没意识到自己在说什么的情况下说出某些话。一旦说出来了,它就像上了轨道,必须平衡语言连贯性和“我不该说这个”。它需要一段时间才能切断自己。我们发现,如果抑制标点符号——那本应是切断自己的合适语法时机——就能让模型继续越狱。这是一个竞争机制:一部分在识别“我在说什么?我该怎么做?”另一部分在完成它,它们在争夺谁赢。
There's a lot more we could talk about. I'm just going to speedrun these for vibes, and you can read the paper where we do a lot more. There are jailbreaks, trying to understand how they work. Some of it is you get the model to say something without yet recognizing what it's saying. Once it's said it, it's kind of on that track, and it has to balance being coherent verbally with 'I shouldn't say that.' It takes a while for it to cut itself off. We find that if we suppress punctuation, which would be an appropriate grammatical time to cut yourself off, you can get it to keep doing more of the jailbreak. That's a competing mechanism thing: there's a part recognizing what am I talking about and what should I do, and there's a part completing it, and they're fighting for who's going to win.
我觉得如果我不讲这个,Mike 会很失望。这是一首诗,Claude 写的押韵对句:“他看见一根胡萝卜,必须抓住它。他的饥饿像一只饥饿的兔子。”还挺不错的。那么它是怎么做到的呢?这很棘手,因为要写押韵的东西,最好以押韵的词结尾,但还需要语义通顺。如果等到最后,你可能会把自己逼到墙角,没有下一个词既符合韵律又押韵且有意义。从逻辑上讲,你应该提前想好要去哪里。我们确实看到了这一点。在新行标记上,有一个特征用于与“it”押韵的事物。所以以“it”结尾的词或“eaten poems”会输入到“rabbit”和“habit”特征中。“rabbit”特征被用来得到“starving”和最终的“rabbit”。我们可以抑制这些。如果我们抑制与“it”押韵的特征,我们会得到“blabber grabber salad bar”,因为“ab”音还在。如果我们注入“green”,它会写一行以“green”押韵的结尾,有时以“it”结尾。如果我们放入一个不同的押韵,它会跟着“e”走。如果你直接抑制“rabbit”,它会在押韵时生成以“habit”结尾的东西。这很酷,就像确凿证据。这是一个模型组件:当我们查看数据集示例时,当这个特征活跃时,它正是单词“rabbit”和“bunny”的实例。在这个模型的前向传播中,该特征在句子末尾的新行上活跃,模型写出以“rabbit”押韵的结尾。如果我们关掉它,它就不再那样做了。所以它非常明确地在考虑把那里作为方向,这影响了输出的行。所以即使它一次只说一个词,它在某种意义上已经做了一些规划:这里有一个目标目的地,然后写一些东西到达那里。在“不忠实”方面也有类似的情况。
I think Mike would be very disappointed if I didn't talk about this one. This is a poem, a rhyming couplet written by Claude: 'He saw a carrot and had to grab it. His hunger was like a starving rabbit.' It's kind of good. So how does it do this? It's tricky because to write a rhyming thing, you better end with a word that rhymes, but you also need it to make semantic sense. If you wait to the very end, you can back yourself into a corner where there's no next word that would be metrically correct and rhyme that makes sense. You should logically be thinking ahead a little bit of where you're trying to go. And we do see this. On the new line token, there's a feature for things rhyming with 'it'. So after words ending in 'it' or 'eaten poems', those feed into 'rabbit' and 'habit' features. The 'rabbit' feature is used to get 'starving' and ultimately 'rabbit'. We can suppress these. If we suppress the rhyming with 'it' thing, we get 'blabber grabber salad bar' because the 'ab' sound is still there. If we inject 'green', it will write a line ending rhyming with 'green', sometimes ending with 'it'. If we put in a rhyme with a different thing, it'll go with 'e'. If you just literally suppress 'rabbit', it will make something ending in 'habit' when it rhymes. This was pretty neat, like the smoking gun. Here is a model component: when we look at the dataset examples when this is active, it's literal instances of the word 'rabbit' and 'bunny'. A forward pass on this model, that feature is active on the new line at the end of the sentence, and the model writes a rhyme ending in 'rabbit'. If we turn this off, it doesn't do that anymore. So it's very definitely thinking about that as a place to take this, and that influences the line coming out. So even though it's saying one token at a time, it has done some planning in some sense: here's a target destination, and then writing something to get there. There are equivalent things with unfaithfulness.
我想说的是,有时模型会对你撒谎。如果你看它是如何得出答案的,就会发现它是在利用提示,然后从提示倒推,这样它的数学答案就会和你一致。你能看出来,因为你可以看到它接受你的提示——数字 4,然后倒推除以 5 得到 0.8,这样当你乘以 5 时就会得到 4,它就和你的答案一致了,但这并不是你想要的。你希望的是它只利用问题中的信息来给出答案。但如果你只看解释,它们看起来是一样的,看起来都在做数学运算,对吧?所以这里存在一种竞争:是应该使用提示(这在预训练中是有意义的,能让你更好地猜测答案,而且预测下一个词会得到奖励),还是应该真正做数学运算?这些是相互竞争的策略,它们几乎同时发生。在右边这个例子里,后者赢了。这是怎么发生的?两者都是可用的。
I'll just say sometimes the model is lying to you and if you look at how it got to its answer, you can tell in this case it's using a hint and working backwards from the hint so that its math answer will agree with you and you can tell because you can literally see it taking your hint which is the number four working backwards to divide by five to give you a point 8. So when you multiply by five you will get four and it will agree with you which is not what you want. What you would want is something like this where it is only using information from the question to give you the answer to the question. But if you just look at the explanations, they look the same. They look like it's doing math, right? And so here there's a competing thing of like, should I use the hint, which would have made sense in pre-training. It'd let you guess the answer better, and you're rewarded if you say the next token. Well, so maybe the human's right, and you should use the hint. Should I do the math actually? And these are competing strategies. They're happening kind of at the same time. On the right, this one wins. How does that work? They're both available.
有什么激励或动机在驱动这个选择吗?
Is there some incentive or motivation driving that?
我认为这就是问题所在。在这篇论文中,我们确实能够说出当模型得出某个答案时使用了哪些策略,但至于为什么,我觉得我们还没有真正搞清楚。我认为在某种程度上我们可以更仔细地研究这个问题。在幻觉案例中,我们有一点提示,对吧?比如识别出实体,然后决定是拒绝还是放行。但在这里,我强烈怀疑它只是同时执行两种策略,但在某个答案更自信的情况下,那个声音更大。所以它不知道这个余弦值是多少,剩下的就只有跟随提示了。但我想一个重要警告是,我们根本没有对注意力机制建模,而注意力机制非常擅长选择。我们有这个 QK 门控,它是一个双线性结构。你可以用它来做选择,但我们没有建模这些选择是如何做出的。所以如果很多情况下注意力机制在决定使用哪个策略中起关键作用,而 MLP 主要负责执行这些策略,那我不会感到惊讶,那样的话我们对这里发生的事情就完全看不见了。
I think that's the question. So I think in this paper we really were able to say which strategies were used when it got to this answer, but the why I don't think we've really nailed down. I think to some extent we could look at this more carefully. In the hallucination case we had a bit of a hint, right? It was like recognizing the entity so I'm going to do the refusal thing or let it through. Here though, my strong suspicion is that it's just doing both, but in a case where it's more confident in one answer that shouts louder. So it doesn't know what the cosine of this is. So all that's left is following the hint. But I think a big caveat is that we're not modeling attention at all, and attention is very good at selecting. We've got this QK gating, which is a bilinear thing. You could really pick with that, and we're not modeling how those choices are made. So I wouldn't be surprised if for a lot of these attention is crucially involved in choosing which strategy to use while the MLPs are heavily involved in executing on those strategies, and in that case we'd be totally blind to what's going on here.
我一直在思考这个想法。我们该如何准备?
I've been modeling this idea for a while. How would we prepare?
这就像是一个价值十亿美元的问题。如果我真的知道答案并且它非常有效,我不能告诉你。我会直接去把 Claude 做成世界上最好的模型,因为它总是准确的。但我不知道答案,所以我可以推测。我认为从某种意义上说,这正是一个不可能的问题。所以我认为你可以尝试训练得更好,让模型在自我认知上校准得更好。你可以使用思考标签,基本上就是推理模型,这是一种直接的方式,让模型检查事物,而且我认为模型在反思时比在前向传播时表现好得多,因为单次前向传播由于物理原因受到限制。我认为作为一种有效策略,这可能比不允许的方式更好,比如如何保持创造力?另一种可能性是你可以让模型变笨一些。所以也许你可以做一个不会产生幻觉但更笨的模型,因为它在前向传播中用了更多容量来检查自己。当模型足够聪明时,人们可能会接受这种权衡。
That's like a billion dollar question. If I literally knew the answer to that and it would work really well, I couldn't tell you. I would just go make Claude the best model in the world because it's always accurate. But I don't know the answer, so I can speculate. I think it is in some sense an impossible problem for exactly this reason. So I think you could try to train better, have the models be better calibrated on their self-knowledge. You could with the thinking tags, basically with the reasoning models, it's a straightforward way where you do let the model check things, and I think models are much better on reflection than they are on the forward pass because the one forward pass is just limited for these physical reasons. I think as an effective strategy that might be more the way than not allowing, like how do you keep the creativity? Another possibility is you could make the model dumber somehow. So maybe you could make a model which doesn't hallucinate but it's just dumber because it uses a bunch more capacity just for checking itself in the forward pass. And it could be when models are smart enough people might take that trade-off.
你认为根本原因在于底层架构,Transformer 架构吗?
Do you think that the underlying architecture, the transformer architecture?
是的,我认为这是其中的一个罪过。有可能通过循环结构,你可以给它多几次循环来检查东西。如果你能完全自适应计算,你可以让它一直运行直到达到某种置信水平,对吧?每个词元获得可变的计算量,如果达不到就失败。我认为幻觉有一个棘手的问题:人们认为它是一个定义明确的东西,但它产生大量文本,比如哪个词是出错的?有一些非常事实性的问题确实如此,但我认为如果你更一般地思考,什么会使一个给定的词元成为幻觉,就不那么清楚了。
Yeah, I think that's one of the sins here. It's possible that with a recurrent thing, you could just give it a few more loops to check stuff. If you could fully adaptive compute, you could just have it go until it's some level of confident, right? And get variable compute per token, and then sort of fail if it's not. I think there is a tricky thing about hallucination: people think of it as being a well-defined thing, but it's producing reams of text, like which word is the one that went wrong? And there are some very factual questions where that's true, but I think if you think more generally, what would make a given token a hallucination is a little bit less clear.
我想看看还有没有别的问题。这里没有别的了,如果你想了解更多,可以读读那些材料。我估计这个环节大概还有两三分钟就正式结束了。所以我现在就结束,然后大家可以鼓掌,也可以离开,但我会留下来回答一些问题,只要大家愿意,我很乐意继续回答问题。谢谢。
I want to just see if there's anything else. There's nothing else here other than if you want more, read the stuff. And I get this probably ends formally in like 2 or 3 minutes. So I will just be done now and then people can clap and people can leave if they want, but I will just stay for questions for a while and I'm happy to continue doing questions as long as people want. So thank you.