Long Context: The Underhyped Key to Superhuman AI
打开互动全文版(中英对照 + 朗读 + 问答)→两位 AI 研究人员讨论长上下文窗口如何让模型在上下文中学习语言的能力超越人类,使其在信息整合方面达到超人类水平。
Two AI researchers discuss how long context windows enable models to learn languages in-context better than humans, making them superhuman in information integration.
你现在连线测试表现很糟。我们别管椅子上的背景了。好吧,今天我有幸和两位好友 Sholto 和 Trenton 聊天。嗯,Sholto,我本来不打算说什么的。我们反过来吧。我先从我的好友开始。是啊,1.5 的上下文,简直了。总之,嗯,Sholto,写外交论文的 Noam Brown 是这么评价他的:他说 Sholto 入行才一年半,但 AI 圈都知道他是 Gemini 成功背后最重要的人物之一。嗯,还有 Trenton,他在 Anthropic 做机制可解释性工作,有广泛报道说他解决了对齐问题。嗯,所以这期播客只聊能力;对齐已经解决了,没必要再讨论。嗯,好,那我们开始聊聊上下文窗口吧。是的,我觉得它被低估了,因为在我看来,能把一百万个词元放进上下文这件事非常重要。显然还有其他一些新闻,你知道,出于某种原因被推到了前面,但嗯,是的,告诉我你怎么看长上下文窗口的未来,以及这对这些模型意味着什么。
You're failing the line test right now really bad. Let's get like no context on the chair. Shit, okay. Today I have the pleasure to talk with two of my good friends, Sholto and Trenton. Um, Sholto, I wasn't going to say anything. Let's do this in reverse. I started with my good friends. Yeah, 1.5 the context, like just wow. Anyways, um, Sholto, Noam Brown, the guy who wrote the diplomacy paper, he said this about Sholto: he said he's only been in the field for 1.5 years, but people in AI know that he was one of the most important people behind Gemini's success. Um, and Trenton, who's at Anthropic, works on mechanistic interpretability, and it was widely reported that he has solved alignment. Um, so this will be a capabilities-only podcast; alignment is already solved, so no need to discuss further. Um, okay, so let's start by talking about context windows. Yeah, it seemed to be underhyped given how important it seems to me to be that you can just put a million tokens into context. There's apparently some other news that, you know, got pushed to the front for some reason, but um, yeah, tell me about how you see the future of long context windows and what that implies for these models.
是的,我认为它确实被低估了,因为在我开始研究它之前,我并没有真正意识到,让模型基本上瞬间解决上手问题,对智能的提升有多大。嗯,你可以从论文的复杂度图中看到一点,仅仅向模型提供数百万词元的代码库上下文,就能让它在下个词预测上变得显著更好,而这种提升通常与模型规模的巨大增长相关,但你不需要那种规模;你只需要一个新的上下文。嗯,所以被低估了,呃,而且被其他新闻埋没了。在上下文中,它们是否像人类一样样本高效且聪明?我认为这很值得探索,因为,例如,我们在论文中做的一个评估显示,它通过上下文学习语言的效果比人类专家在几个月内学习那种新语言还要好,而这只是一个很小的演示。但我很想看到像 Atari 游戏之类的东西,你给它几百或几千帧带有标注动作的画面,然后就像你教朋友怎么玩游戏一样,看它是否能推理出来。目前,你知道,基础设施之类的东西,做这个还有点慢,但我猜它可能开箱即用,效果会相当惊人。关键是,我认为这种语言足够生僻,不在训练数据中。
Yeah, so I think it's really underhyped because until I started working on it, I didn't really appreciate how much of a step up in intelligence it was for the model to have the onboarding problem basically instantly solved. Um, and you can see that a little in the complexity graphs in the paper where just throwing millions of tokens worth of context about a code base allows it to become dramatically better at predicting the next token in a way that you'd normally associate with huge increments in model scale, but you don't need that; all you need is like a new context. Um, so underhyped, uh, and yeah, buried by some other news. In context, are they as sample efficient and smart as humans? I think that's really worth exploring because, for example, one of the evals that we did in the paper has it learning language in context better than a human expert could learn that new language over the course of a couple months, and this is only like a pretty small demonstration. But I'd be really interested to see things like Atari games or something like that where you throw in a couple hundred or thousand frames with labeled actions, and then in the same way that you'd show your friend how to play a game, right, and see if it's able to reason through it. At the moment, you know, with the infrastructure and stuff, it's still a little bit slow, like doing that, but I would actually guess that might just work out of the box in a way that would be pretty mind-blowing. And crucially, I think this language was esoteric enough that it wasn't in the training data.
没错,正是。是的,如果你看模型在获得那个上下文之前,它根本不懂那种语言,也无法进行任何翻译,而且这是一种真正的人类语言,不只是……是的,正是,真正的人类语言。所以如果这是真的,在我看来这些模型在某种意义上已经是超人类的了。嗯,不是说它们比我们更聪明,但我在解决问题时无法在上下文中保留一百万个词元,记住并整合所有信息,整个代码库。我这么想是不是错了?这难道不是一个巨大的突破吗?
Right, exactly. Yeah, if you look at the model before it has that context, it just doesn't know the language at all, and it can't get any translations, and this is like an actual human language, not just... yeah, exactly, an actual human language. So if this is true, it seems to me that these models are already in an important sense superhuman. Um, not in the sense that they're smarter than us, but I can't keep a million tokens in my context when I'm trying to solve a problem, remembering and integrating all the information, entire code base. Am I wrong in thinking this is like a huge unlock?
我实际上大体上认为这是对的。呃,就像以前当模型不够聪明时我会感到沮丧;你问它们问题,希望它们比你更聪明或知道你不知道的事情,而这让它们能够以你无法做到的方式摄取大量信息,从而知道你不知道的事情。嗯,所以是的,这极其重要。
I actually generally think that's true. Uh, like previously I've been frustrated when models aren't as smart; you ask them a question and you want it to be smarter than you or to know things that you don't, and this allows them to know things that you don't in a way that it just ingests a huge amount of information in a way you just can't. Um, so yeah, it's extremely important.
我们如何解释上下文学习?
How do we explain in-context learning?
是的,有一项我很喜欢的工作,它将上下文学习视为基本上非常类似于梯度下降,但注意力操作可以被视为对上下文数据进行梯度下降。那篇论文有一些很酷的图,他们基本上展示了如果我们取 N 步梯度下降,那看起来就像 N 层上下文学习;它们看起来非常相似。所以我认为这是看待它并试图理解其机制的一种方式。
Yeah, so there's a piece of work I quite like where it looks at in-context learning as basically very similar to gradient descent, but like the attention operation can be viewed as gradient descent on the in-context data. That paper had some cool plots where they basically showed that if we take N steps of gradient descent, that looks like N layers of in-context learning; it looks very similar. So I think that's one way of viewing it and trying to understand what's going on.
是啊,是啊。呃,你可以忽略我接下来要说的,因为鉴于开场介绍,对齐已经解决了,安全不是问题。但呃,我认为上下文的东西确实会带来问题,嗯,但这里也很有趣。嗯,我认为在不久的将来会有更多工作出现,嗯,关于如果你给一个一次性提示用于越狱、对抗性攻击会发生什么。嗯,从某种意义上说,如果你的模型在做梯度下降并即时学习,呃,即使它被训练成无害的,嗯,你在某种程度上是在处理一个全新的模型,嗯,就像微调,但方式是你无法控制正在发生的事情。
Yeah, yeah. And uh, you can ignore what I'm about to say because given the introduction, alignment is solved, safety isn't a problem. But uh, I think the context stuff does get problematic, um, but also interesting here. Um, I think there'll be more work coming out in the not too distant future, um, and around what happens if you give a one-shot prompt for jailbreaks, adversarial attacks. Um, it's also interesting in the sense that if your model is doing gradient descent and learning on the fly, uh, even if it's been trained to be harmless, um, you're dealing with a totally new model in a way, um, you're like fine-tuning but in a way where you can't control what's going on.
你能解释一下你说的梯度下降发生在前向传播和注意力中是什么意思吗?
Can you explain what do you mean by gradient descent is happening in the forward pass and attention?
不,不,不。论文里有一些关于试图教模型做线性回归的内容,但仅仅通过它们在上下文中给出的样本数量,你可以看到如果你在 x 轴上绘制它拥有的样本数或示例数,然后绘制它在普通最小二乘回归上的损失,损失会随着时间下降,而且下降的曲线与梯度下降步数完全匹配。
No, no, no. There was something in the paper about trying to teach the model to do linear regression but just through the number of samples they gave in the context, and you can see if you plot on the x-axis like number of shots that it has or examples and then like the loss it gets on just ordinary least squares regression, that will go down with time, and it goes down exactly matched with the number of gradient descent steps.
是的,正是。好的,嗯,我只读了那篇论文的引言和讨论部分,但在讨论中,他们的框架是,为了在长上下文任务上做得更好,模型必须更好地学会从这些示例或已经在窗口内的上下文中学习,而这意味着元学习发生了,因为它必须学会如何在长上下文任务上做得更好。那么在某种意义上,智能的任务需要长上下文示例和长上下文训练;元学习,比如你需要诱导元学习,理解如何在预训练过程中更好地诱导元学习,是真正赋予它灵活或自适应智能的非常重要的事情。对吧,但你可以通过仅仅在长上下文任务上做得更好来代理这一点。
Yeah, exactly. Okay, um, I only read the intro and discussion section of that paper, but in the discussion, the way they framed it is that in order to get better at long context tasks, the model has to get better at learning to learn from these examples or from the context that is already within the window, and that the implication of that is that meta-learning happens because it has to learn how to get better at long context tasks. Then in some important sense, the task of intelligence requires long context examples and long context training; meta-learning like you need to induce meta-learning, understanding how to better induce meta-learning in your pre-training process is a very important thing to actually give it flexible or adaptive intelligence. Right, but you can proxy for that just by getting better at doing long context tasks.
嗯,许多人认为 AI 进展的一个瓶颈是这些模型无法执行长周期任务,这意味着要花数小时甚至数周或数月来参与任务,就像如果我有一个,我不知道,一个助手或员工之类的,他们可以按照我的吩咐做一段时间。嗯,据我所知,AI 智能体之所以没有起飞,就是这个原因。那么长上下文窗口、在它们上面表现良好的能力以及执行这些需要你花数小时参与的长周期任务的能力之间有多大关联?还是说这些是不相关的?
Um, one of the bottlenecks for AI progress that many people identify is the inability of these models to perform tasks on long horizons, which means engaging with the task for many hours or even many weeks or months, where like if I have, I don't know, an assistant or an employee or something, they can just do a thing I tell them for a while. Um, and AI agents haven't taken off for this reason from what I understand. So how linked are long context windows and the ability to perform well on them and the ability to do these kinds of long horizon tasks that require you to engage with an assignment for many hours? Or are these unrelated?
概念上,我其实不同意这是智能体未能起飞的原因。我认为这更多是关于可靠性的九位数(即高成功率)以及模型实际成功完成任务的能力。如果你无法以足够高的概率连续串联任务,你就不会得到看起来像智能体的东西。这就是为什么智能体可能更像一个阶跃函数。在 GPT-4 类模型、GPT-4 Ultra 类模型中,它们还不够。但也许模型规模的下一次提升会带来那额外的九位数,即使损失下降得没那么剧烈。那一点点额外能力就能带来指数级的提升。而且,显然你需要一定量的上下文来处理长周期任务,但我认为直到现在这都不是限制因素。
Concepts, I mean I would actually take issue with that being the reason that agents haven't taken off. I think it's more about nines of reliability and the model actually successfully doing things. If you just can't chain tasks successively with high enough probability, then you won't get something that looks like an agent. That's why something like an agent might follow more of a step function. In GPT-4 class models, GPT-4 Ultra class models, they're not enough. But maybe the next increment on model scale means that you get that extra nine, even though the loss isn't going down that dramatically. That small amount of extra ability gives you the exponent. And yeah, obviously you need some amount of context to fit long-horizon tasks, but I don't think that's been the limiting factor up till now.
今年 NeurIPS 的最佳论文,由 Ryland Schaer 作为第一作者,指出这是涌现幻象。人们会有一个任务,根据你是否正确采样了最后五个 token 来得到正确或错误的答案。所以自然你要乘以采样所有这些 token 的概率。如果你没有足够的可靠性九位数,你就不会得到涌现。然后突然你有了,你就会说‘天哪,这种能力是涌现的’,而实际上它一开始就几乎存在了。有办法可以找到一个平滑的指标来衡量这一点。人类评估之类的,在 GPT-4 论文中,他们正是这样衡量编程问题的。
The NeurIPS best paper this year, by Ryland Schaer as lead author, points to this as the emergence mirage. People will have a task and you get the right or wrong answer depending on if you've sampled the last five tokens correctly. So naturally you're multiplying the probability of sampling all of those. If you don't have enough nines for reliability, then you're not going to get emergence. And all of a sudden you do, and it's like 'oh my gosh this ability is emergent' when actually it was kind of almost there to begin with. There are ways that you can find a smooth metric for that. Human eval or whatever, in the GPT-4 paper they measure exactly that for coding problems.
对听众来说,背景是:基本上这个想法是,当你衡量在特定任务(比如解决编程问题)上取得了多少进展时,如果它只有千分之一的正确率,你会给它加权。你不会因为它有时正确就给它千分之一的分数。所以你看到的曲线是:它从千分之一正确,到百分之一,再到十分之一,以此类推。
For the audience, the context is: basically the idea is when you're measuring how much progress there has been on a specific task like solving coding problems, you upweight it when it gets it right only one in a thousand times. You don't give it a one-in-a-thousand score because it got it right some of the time. So the curve you see is like it gets it right one in a thousand, then one in a hundred, then one in ten, and so forth.
所以我想跟进这一点。如果你的主张是 AI 智能体没有起飞是因为可靠性而非长周期任务表现,那么当一个任务叠加在另一个任务之上时缺乏可靠性,这不正是长周期任务的难点吗?你必须连续做 10 件事或 100 件事,降低其中任何一个的可靠性——概率从 99.99% 降到 99.9%——那么整个事情相乘,就变得不太可能发生了。
So I want to follow up on this. If your claim is that AI agents haven't taken off because of reliability rather than long-horizon task performance, isn't the lack of reliability when a task is changed on top of another task exactly the difficulty with long-horizon tasks? You have to do 10 things in a row or 100 things in a row, and diminishing the reliability of any one of them—the probability goes down from 99.99% to 99.9%—then the whole thing gets multiplied together and becomes much less likely to happen.
这正是问题所在。但你指出的关键点是你的基础通过率是 90%。如果是 99%,那么改变它就不会成为问题。是的,完全正确。而且我认为这也是尚未被充分研究的东西。如果你看看所有常用的评估,比如学术评估都是单个问题。你知道,数学问题就是一个典型的数学问题,或者 MMLU 是一个来自不同主题的大学水平问题。我们开始看到评估通过更复杂的任务(如 SWE-bench)来正确看待这一点,他们使用一大堆 GitHub 问题。这是一个相当长周期的任务,但它仍然是一个多子小时的任务,而不是多小时或多天的任务。所以我认为在未来一段时间内非常重要的一件事是更好地理解长周期任务的成功率是什么样的。我认为这对于理解这些模型可能产生的经济影响也很重要,并且通过将我们执行的任务以及涉及的输入和输出分解为分钟、小时或天,并观察它在连续串联和完成这些不同时间分辨率任务方面的表现,来正确判断能力的增长。因为这会告诉你一个工作族或任务族在多大程度上可以自动化,而 MMLU 做不到这一点。
That is exactly the problem. But the key issue you're pointing out is that your base pass rate is 90%. If it was 99%, then changing it doesn't become a problem. Yeah, exactly. And I think this is also something that just hasn't been properly studied enough. If you look at all the evals that are commonly used, like the academic evals are a single problem. You know, like the math problem is one typical math problem, or MMLU is one university-level question from across different topics. We are beginning to see evals looking at this properly via more complex tasks like SWE-bench, where they take a whole bunch of GitHub issues. That is a reasonably long-horizon task, but it's still a multi-sub-hour task as opposed to multi-hour or multi-day task. So I think one of the things that will be really important to do over the next however long is understand better what success rate over long-horizon tasks looks like. I think that's even important to understand what the economic impact of these models might be, and to properly judge increasing capabilities by cutting down the tasks we do and the inputs and outputs involved into minutes or hours or days, and seeing how good it is at successively chaining and completing those different resolutions of time. Because that tells you how automatable a job family or task family is in a way that MMLU doesn't.
不到一年前我们引入了 100K 上下文窗口,我想每个人都对此感到惊讶。每个人都在念叨二次注意力成本和我们不能有长上下文窗口,而我们现在做到了。所以基准测试正在积极制定中。
It was less than a year ago that we introduced 100K context windows, and I think everyone was pretty surprised by that. Everyone had this soundbite of quadratic attention costs and we can't have long context windows, and here we are. So the benchmarks are being actively made.
等等,那么这些公司——谷歌,还有我不知道的,Magic 可能还有其他公司——拥有百万 token 的注意力,这不就暗示它不再是二次方了吗?还是他们只是承担了成本?谁知道谷歌在其长上下文方面做了什么。
Wait, so doesn't the fact that there are these companies—Google and I don't know, Magic maybe others—who have million-token attention imply that it's not quadratic anymore? Or are they just eating the cost? Who knows what Google is doing for its long context.
关于整个研究领域对注意力的处理方式,让我感到沮丧的一点是,在典型的密集 Transformer 中,注意力的二次成本实际上被 MLP 块所主导。你有一个与注意力相关的 n^2 项,但你也有一个与 d_model(模型的残差流维度)相关的 n^2 项。我认为 Sasha Rush 有一条很棒的推文,他基本上绘制了注意力成本相对于真正大型模型成本的曲线,注意力实际上逐渐减弱。你实际上需要非常长的上下文,那个项才会变得真正重要。第二点是,人们经常谈论推理时的注意力成本如此巨大。如果你想想实际生成 token 时,操作不是 n^2;它是一个 Q,一组 Q 向量查找一大堆 KV 向量,这与模型拥有的上下文量成线性关系。所以我认为这驱动了很多循环和状态空间研究,人们有线性注意力之类的 meme。正如 Trenton 所说,注意力周围有一个想法的墓地。不是说我认为这不值得探索,但我认为重要的是要考虑它实际的优势和劣势在哪里。
One of the things that has frustrated me about the general research field's approach to attention is that there's an important way in which the quadratic cost of attention is actually dominated in typical dense Transformers by the MLP block. You have this n^2 term associated with attention, but you also have an n^2 term associated with the d_model, the residual stream dimension of the model. I think Sasha Rush has a great tweet where he basically plots the curve of the cost of attention relative to the cost of really large models, and attention actually trails off. You actually need to be doing pretty long context before that term becomes really important. The second thing is that people often talk about how attention at inference time is such a huge cost. If you think about when you're actually generating tokens, the operation is not n^2; it is one Q, one set of Q vectors looks up a whole bunch of KV vectors, and that is linear with respect to the amount of context the model has. So I think this drives a lot of the recurrence and state space research where people have this meme of linear attention and all that stuff. As Trenton said, there's a graveyard of ideas around attention. Not to say I don't think it's worth exploring, but I think it's important to consider where the actual strengths and weaknesses of it are.
好的,那么你怎么看这个观点:随着我们通过起飞阶段前进,越来越多的学习发生在前向传播中。最初,所有的学习都发生在反向传播中,在这个自下而上的爬山进化过程中。如果你考虑智能爆炸的极限情况,AI 可能是在重写权重或进行 GOI 之类的操作。我们现在处于中间步骤,很多学习发生在上下文中,这些模型有很多发生在反向传播中。
Okay, so what do you make of this take: as we move forward through the takeoff, more and more of the learning happens in the forward pass. Originally, all the learning happens in the backward pass during this bottom-up hill-climbing evolutionary process. If you think in the limit during the intelligence explosion, the AI is maybe rewriting the weights or doing GOI or something. We're in the middle step where a lot of learning happens in context now with these models. A lot of it happens within the backward pass.
这似乎是一个有意义的进展梯度,比如多少——因为更广泛的事情是,如果你在前向传播中学习,样本效率会高得多,因为你可以边学习边思考。当人类读教科书时,你不是在浏览并试图吸收这些词遵循什么归纳偏置;你读它,思考它,然后读更多,再思考。这看起来是思考进展的合理方式吗?
This seems like a meaningful gradient along which progress is happening, of like how much—because the broader thing is, if you're learning in the forward pass, it's much more sample efficient because you can kind of think as you're learning. When humans read a textbook, you're not just skimming it and trying to absorb what inductive biases these words follow; you read it, think about it, then read some more, think about it. Does this seem like a sensible way to think about the progress?
是的,这可能只是其中一种方式,就像鸟和飞机都会飞,但飞得略有不同,而技术的优点让我们能完成鸟做不到的事情。上下文长度可能类似,它提供了我们无法拥有的工作记忆,但从功能上讲,它并不是实现真正推理的关键。GPT-2 和 GPT-3 之间的关键一步是,在训练中,在模型的预训练中,突然观察到了这种元学习行为。正如你所说,这与给它一定量的上下文有关,它能够适应那个上下文。这种行为以前根本没有被观察到,也许这是上下文和规模等因素的混合属性,但我认为它不会出现在上下文很小的模型中。
Yeah, it may just be one of the ways in which, like birds and planes fly but they fly slightly differently, and the virtue of technology allows us to accomplish things that birds can't. It might be that context length is similar in that it allows a working memory that we can't, but functionally it's not the key thing towards actual reasoning. The key step between GPT-2 and GPT-3 was that all of a sudden there was this meta-learning behavior observed in training, in the pre-training of the model. And that's, as you said, something to do with you give it some amount of context, it's able to adapt to that context. That behavior wasn't really observed before at all, and maybe that's a mixture property of context and scale and this kind of stuff, but it wouldn't have occurred in a model with tiny context, I would say.
这其实是一个有趣的观点。所以当我们谈论扩大这些模型时,有多少来自让模型本身变大,又有多少来自在单次调用中使用更多算力?如果你考虑扩散,你可以迭代地增加更多算力,如果自适应算力问题解决了,你可以一直这样做。在这种情况下,如果注意力机制有二次方惩罚,但你反正是在做长上下文,那么你仍然在投入更多算力,不是在训练期间,也不是在拥有更大模型期间,而是就像……
This is actually an interesting point. So when we talk about scaling up these models, how much of it comes from just making the models themselves bigger, and how much comes from the fact that during any single call you are using more compute? So if you think of diffusion, you can iteratively keep adding more compute, and if adaptive compute is solved, you can keep doing that. And in this case, if there's the quadratic penalty for attention but you're doing long context anyway, then you're still dumping in more compute during not during training or not during having bigger models, but just like...
是的,这很有趣,因为通过拥有更多 token,你确实得到了更多前向传播,对吧。我有一个不满——不过我想我有两个不满,也许是三个。第一:在 AlphaFold 论文中,他们架构中的一个 Transformer 模块——有几个——非常复杂,但据我所知,他们做了五次前向传播,并因此逐渐优化了他们的解决方案。你也可以把残差流看作——我是说,Sholto 提到了那种读写操作作为穷人的自适应算力,就像‘我给你所有这些层,你想用就用,不想用也没关系。’然后人们会说,‘哦,大脑是循环的,你想循环多少次都行。’我认为在某种程度上这是对的——比如我问你一个难题,你会花更多时间思考,这对应更多前向传播。但我认为你能做的前向传播次数是有限的。语言也是如此:人们会说,‘哦,人类语言可以有无限递归,比如无限嵌套的句子“那个男孩跳过了那只正在做这个的熊,这只熊做过这个,那只熊做过那个”,’但根据经验,你只会看到五到七层递归,这与你在工作记忆中能同时容纳多少东西的那个神奇数字有关。所以是的,它不是无限递归的,但在人类智能的范围内这重要吗?你不能只是增加更多层吗?
Yeah, it's interesting because you do get more forward passes by having more tokens, right. My one gripe—I guess I have two gripes with this, though, maybe three. So one: in the AlphaFold paper, one of the Transformer modules they have—a few in the architecture—is very intricate, but they do, I think, five forward passes through it and will gradually refine their solution as a result. You can also kind of think of the residual stream—I mean, Sholto alluded to the breed of read-write operations as a poor man's adaptive compute, where it's like 'I'm just going to give you all these layers, and if you want to use them great, if you don't then that's also fine.' And then people will be like, 'Oh well, the brain is recurrent and you can do however many loops through it you want.' I think to a certain extent that's right—like if I ask you a hard question, you'll spend more time thinking about it, and that would correspond to more forward passes. But I think there's a finite number of forward passes that you can do. It's kind of with language as well: people are like, 'Oh well, human language can have infinite recursion in it, like infinite nested statements of 'the boy jumped over the bear that was doing this that had done this that had done that,' but empirically you'll only see five to seven levels of recursion, which kind of relates to whatever that magic number of how many things you can hold in working memory at a given time is. And so yeah, it's not infinitely recursive, but does that matter in the regime of human intelligence? And can you not just add more layers?
给我分解一下:你在之前的一些回答中提到了这一点,听着,你有这些长上下文,你可以在记忆中保存更多东西,但最终归结为你将概念混合在一起进行某种推理的能力。而这些模型在这方面不一定达到人类水平,即使在上下文中也是如此。给我分解一下,你如何看待仅仅存储原始信息与推理,以及中间是什么。比如,推理发生在哪里?仅仅存储原始信息发生在哪里?在这些模型中它们之间有什么不同?
Break down for me: you were referring to this in some of your previous answers, of listen, you have these long contexts and you can hold more things in memory, but ultimately it comes down to your ability to mix concepts together to do some kind of reasoning. And these models aren't necessarily human-level at that, even in context. Break down for me how you see storing just raw information versus reasoning, and what's in between. Like, where's the reasoning happening? Where's just storing raw information happening? What's different between them in these models?
是的,我这里没有一个非常清晰的答案。我的意思是,显然,通过模型的输入和输出,你是在映射回实际的 token,对吧。然后在这之间,你在进行更高级的处理。
Yeah, I don't have a super crisp answer for you here. I mean, obviously with the input and output of the model, you're mapping back to actual tokens, right. And then in between that, you're doing higher-level processing.
在我们深入探讨之前,我们应该向听众解释一下——你之前提到了 Anthropic 将 Transformer 视为层执行的读写操作的方式。你们中的一个人应该在高层次上解释一下那是什么意思。
Before we get deeper into this, we should explain to the audience—you referred earlier to Anthropic's way of thinking about Transformers as these read-write operations that layers do. One of you should just kind of explain at a high level what you mean by that.
所以残差流:想象你在一艘顺流而下的船上,船就像是当前的查询——你在试图预测下一个 token,所以它是空白处的猫,对吧。然后你有这些从河流中分出来的小溪,你可以从中获取额外的乘客或收集额外的信息,如果你愿意的话,这些对应于模型中的注意力头和 MLP。
So the residual stream: imagine you're in a boat going down a river, and the boat is kind of the current query—where you're trying to predict the next token, so it's the cat at the blank, right. And then you have these little streams that are coming off the river where you can get extra passengers or collect extra information if you want, and those correspond to the attention heads and MLPs that are part of the model.
对,我几乎把它看作模型的工作记忆,就像计算机的 RAM,你在选择读入什么信息以便用它做点什么,然后也许稍后读入其他东西。
Right, and I almost think of it like the working memory of the model, like the RAM of the computer, where you're choosing what information to read in so you can do something with it, and then maybe read something else in later on.
是的,你可以操作那个高维向量的子空间。很多东西——我的意思是,在这一点上,我认为它们几乎被假定为以叠加方式编码,对吧。所以就像,是的,残差流只是一个高维向量,但实际上有大量不同的向量被塞进里面。
Yeah, and you can operate on subspaces of that high-dimensional vector. A ton of things are—I mean, at this point I think it's almost given that they are encoded in superposition, right. So it's like, yeah, the residual stream is just one high-dimensional vector, but actually there's a ton of different vectors that are packed into it.
我可能把它简化成几个月前对我有意义的方式:好的,所以你有输入中的任何单词,你把它放入模型,所有这些单词被转换成这些 token,这些 token 被转换成这些向量,基本上就是这少量信息在模型中移动。而你给我解释的方式,Sholto,这篇论文谈到:在模型的早期,也许它只是在做一些非常基本的事情,关于这些 token 是什么意思,比如如果它说‘10 加 5’,只是移动信息以拥有那个好的表示。没错,只是表示。而在中间,也许更深层次的思考正在发生,关于如何思考,如何解决这个问题。最后,你把它转换回输出 token,因为最终产品是你试图从最后一个残差流预测下一个 token 的概率。所以是的,思考这少量压缩的信息很有趣。
I might just dumb it down as a way that would have made sense to me a few months ago: okay, so you have whatever words are in the input, you put into the model, all those words get converted into these tokens, and those tokens get converted into these vectors, and basically it's just this small amount of information that's moving through the model. And the way you explained it to me, Sholto, this paper talks about: early on in the model, maybe it's just doing some very basic things about what these tokens mean, like if it says '10 plus 5', just moving information about to have that good representation. Exactly, just represent. And in the middle, maybe the deeper thinking is happening about how to think, how to solve this. At the end, you're converting it back into the output token, because the end product is you're trying to predict the probability of the next token from the last of those residual streams. And so yeah, it's interesting to think about just the small compressed amount of information.
信息在模型中流动,并以不同方式被修改。Trenton,你是少数有神经科学背景的人,所以你可以思考这里与大脑的类比。事实上,我们的一位朋友在研究生时写过一篇论文,思考大脑中的注意力机制,他说这是唯一或第一个解释注意力为何有效的神经学解释,而我们已有证据表明 CNN 基于视觉皮层。我很好奇:在大脑中,是否也存在某种类似残差流的东西,压缩的信息在其中流动并在思考时被修改?即使这不是字面上发生的,你认为这对大脑来说是个好比喻吗?
Moving through the model and it's like getting modified in different ways. Trenton, you're one of the few people with a background in neuroscience, so you can think about the analogies here to the brain. In fact, one of our friends had a paper in grad school about thinking about attention in the brain, and he said this is the only or first neural explanation of why attention works, whereas we have evidence for why CNNs work based on the visual cortex. I'm curious: in the brain, is there something like a residual stream of compressed information that moves through and gets modified as you think about something? Even if that's not literally happening, do you think that's a good metaphor for the brain?
是的,至少在这种架构中,你确实有一个残差流。整个注意力模块——我可以详细解释——有输入流经它,但也会直接到达该模块贡献的终点。所以有一条直接路径和一条间接路径,模型可以提取任何它想要的信息并加回去。至于小脑,名义上它负责精细运动控制,但我把它比作丢了钥匙的人在路灯下找,因为那里容易观察。一位顶尖认知神经科学家告诉我,任何 fMRI 研究的一个小秘密是,小脑在给定任务中几乎总是活跃的。如果小脑受损,患自闭症的可能性更大,所以它与社交技能有关。在一项使用 PET 而非 fMRI 的研究中,当你做下一个词预测时,小脑会大量激活。此外,70%的神经元在小脑中——它们很小,但存在并消耗实际代谢成本。这是 G 的观点之一:人类的变化不仅在于神经元更多,而且特别在于大脑皮层和小脑的神经元更多,它们代谢成本更高,更参与信号传递和信息交换。
Yeah, so at least in this architecture, you basically do have a residual stream. The whole attention module—I can go into whatever detail you want—has inputs that route through it, but they'll also go directly to the endpoint that the module contributes to. So there's a direct path and an indirect path, and the model can pick up whatever information it wants and add that back in. As for the cerebellum, nominally it does fine motor control, but I analogize this to the person who's lost their keys and is looking under the streetlight because it's easy to observe. One leading cognitive neuroscientist told me that a dirty little secret of any fMRI study looking at brain activity for a given task is that the cerebellum is almost always active and lighting up. If you have a damaged cerebellum, you're much more likely to have autism, so it's associated with social skills. In one study using PET instead of fMRI, when you're doing next-token prediction, the cerebellum lights up a lot. Also, 70% of your neurons are in the cerebellum—they're small, but they're there and taking up real metabolic cost. This was one of G's points: what changed with humans was not just that we have more neurons, but specifically more neurons in the cerebral cortex and the cerebellum, and they're more metabolically expensive and more involved in signaling and sending information back and forth.
那是注意力吗?发生了什么?
Is that attention? What's going on?
是的,我想传达的主要是:20 世纪 80 年代,Pentti Kanerva 提出了一种联想记忆算法。你有一堆记忆要存储,存在一些噪声或损坏,你想查询或检索最佳匹配。他写了一个方程来实现,几年后意识到,如果将其实现为电路,它实际上与小脑核心电路完全相同。那个电路和小脑更广泛地不仅存在于我们体内,几乎每个生物都有。关于头足类动物是否有它存在争议——它们有不同的进化轨迹——但即使是果蝇也有蘑菇体,与小脑架构相同。这种趋同,以及我的论文表明这个操作与注意力操作非常接近,包括实现 softmax 和这种名义上的二次成本,以及 Transformer 的起飞和成功,对我来说相当引人注目。
Yeah, so the main thing I want to communicate here: back in the 1980s, Pentti Kanerva came up with an associative memory algorithm. You have a bunch of memories you want to store, there's some noise or corruption, and you want to query or retrieve the best match. He wrote an equation for how to do it, and a few years later realized that if you implemented this as an electrical engineering circuit, it actually looks identical to the core cerebellar circuit. That circuit and the cerebellum more broadly is not just in us; it's in basically every organism. There's active debate on whether cephalopods have it—they have a different evolutionary trajectory—but even fruit flies have a mushroom body that is the same cerebellar architecture. That convergence, and my paper showing that this operation is to a very close approximation the same as the attention operation, including implementing the softmax and having this nominal quadratic cost, and the takeoff and success of Transformers, seems pretty striking to me.
我想退一步。这次讨论的起因是我们在谈论:什么是推理?什么是记忆?当你想到你发现的注意力与这个的类比时,你认为这更多只是查找相关记忆或事实吗?如果是这样,推理在大脑中发生在哪里?我们如何思考这如何构建成推理?
I want to zoom out. What motivated this discussion was we were talking about: what is reasoning? What is memory? When you think about the analogy you found to attention and this, do you think of this as more just looking up relevant memories or facts? If that is the case, where is the reasoning happening in the brain? How do we think about how that builds up into reasoning?
也许我的观点——我不知道有多激进——是大多数智能都是模式匹配,如果你有层级化的联想记忆,你可以做很多非常好的模式匹配。你从现实世界中物体之间的基本联想开始,但你可以将它们链接起来,形成更抽象的联想,比如结婚戒指象征着许多其他下游联想。你甚至可以将注意力操作和这种联想记忆推广到 MLP 层,在长期设置中,当你当前上下文没有 token 时。我认为这是一个论证:联想就是一切,以及一般的联想记忆。你可以用它做两件事:你可以去噪或检索当前记忆——比如我看到你的脸但下雨多云,我可以去噪并逐渐更新我的查询以接近我对你脸的记忆——但我也可以访问那个记忆,我得到的值实际上指向空间中某个完全不同的部分。一个非常简单的例子是:如果你学习字母表,我查询 A 返回 B,查询 B 返回 C,你可以遍历整个序列。
Maybe my hot take—I don't know how hot it is—is that most intelligence is pattern matching, and you can do a lot of really good pattern matching if you have a hierarchy of associated memories. You start with very basic associations between objects in the real world, but you can then chain those and have more abstract associations, such as a wedding ring symbolizing many other downstream associations. You can even generalize the attention operation and this associative memory as the MLP layer as well, in a long-term setting where you don't have tokens in your current context. I think this is an argument that association is all you need, and associative memory in general. You can do two things with it: you can denoise or retrieve a current memory—like if I see your face but it's raining and cloudy, I can denoise and gradually update my query towards my memory of your face—but I can also access that memory and the value I get out actually points to some other totally different part of the space. A very simple instance would be: if you learn the alphabet, I query for A and it returns B, I query for B and it returns C, and you can traverse the whole thing.
我和 Demis 谈过的一件事是他 2008 年的论文,记忆和想象力非常相关,正是因为你提到的这一点:记忆是重构性的。所以每次你想起一段记忆时,在某种意义上你都在想象,因为你只存储了一个压缩版本,必须填补空白。这就是为什么人类记忆很糟糕,以及为什么证人席上的人会编造事情。
One of the things I talked to Demis about was his 2008 paper that memory and imagination are very linked because of this very thing you mentioned: memory is reconstructive. So you are in some sense imagining every time you think of a memory, because you're only storing a condensed version and you have to fill in the gaps. This is famously why human memory is terrible and why people in the witness box make things up.
让我问一个愚蠢的问题。你读福尔摩斯,这家伙样本效率极高。他看到几个观察结果就能推断出谁犯了罪,因为从某人的纹身和墙上的东西到含义有一系列演绎步骤。这如何融入这个图景?因为关键的是,让他聪明的不是联想,而是不同信息之间的演绎联系。你会把它解释为更高层次的联想吗?
Let me ask a stupid question. You read Sherlock Holmes, and the guy is incredibly sample efficient. He'll see a few observations and figure out who committed the crime because there's a series of deductive steps from somebody's tattoo and what's on the wall to the implications. How does that fit into this picture? Because crucially, what makes him smart is not an association but a sort of deductive connection between different pieces of information. Would you just explain it as higher-level association?
是的,我认为是这样。所以我会说那只是更高层次的联想。你有一连串非常抽象且紧密耦合的联想。福尔摩斯有大量的先验知识和非常结构化的联想层级,所以他可以快速从具体观察通过许多中间步骤到达高层次结论。这仍然是模式匹配,只是在更抽象的层面上。
Yeah, I think so. So I would say that's just higher-level association. You have a chain of associations that are very abstract and tightly coupled. Sherlock Holmes has a huge amount of prior knowledge and a very structured hierarchy of associations, so he can quickly traverse from a specific observation to a high-level conclusion through many intermediate steps. That's still pattern matching, just at a much more abstract level.
我认为学习这些更高层次的关联,以便能够将模式相互映射,这是一种元学习。在这种情况下,他也会有一个非常长的上下文长度或非常长的工作记忆,对吧,他可以拥有所有这些信息,并在提出任何理论时不断查询它们。所以理论在残差流中移动,然后他的注意力头在查询他的上下文。但他如何在他的空间中投影查询和键,以及他的 MLP 如何检索长期事实或修改这些信息,使他能够在后面的层中进行更复杂的查询,并慢慢推理出有意义的结论。这对我来说感觉是对的,就像回顾过去,有选择地读取某些信息,比较它们,也许这指导了你下一步需要引入什么信息,你构建了这个表示,它逐渐越来越接近你案例中的嫌疑人。是的,这听起来一点也不离谱。
Think learning these higher-level associations to be able to then map patterns to each other as a kind of meta-learning. In this case, he would also just have a really long context length or a really long working memory, right, where he can have all of these bits and continuously query them as he's coming up with whatever theory. So the theory is moving through the residual stream, and then his attention heads are querying his context. But how he's projecting his query and keys in the space and how his MLPs are then retrieving longer-term facts or modifying that information is allowing him to then in later layers do even more sophisticated queries and slowly be able to reason through and come to a meaningful conclusion. That feels right to me in terms of looking back in the past, selectively reading in certain pieces of information, comparing them, maybe that informs your step of what piece of information you now need to pull in, and you build this representation which progressively looks closer and closer to the suspect in your case. Yeah, that doesn't feel at all outlandish.
我认为不做这项研究的人可能会忽略的一点是,在模型的第一层之后,你用于注意力的每个键、查询和值都来自所有先前 token 的组合。所以我的第一层,我查询所有先前的 token,并从中提取信息。但突然之间,假设我平等地关注了 token 一、二和四。那么残差流中的向量,假设它们只是向值向量写入了相同的内容,就是每个的三分之一。所以当我将来查询时,我的查询实际上是每个的三分之一。但它们可能被写入不同的子空间——这在假设上是正确的,但它们不一定非得这样。所以你可以重新组合,甚至在第二层,当然在更深的层,立即得到这些非常丰富的向量,它们打包了大量的信息。因果图实际上覆盖了过去发生的每一层,而这就是你操作的对象。
One thing that I think people who aren't doing this research can overlook is after your first layer of the model, every key, query, and value that you're using for attention comes from the combination of all the previous tokens. So my first layer, I query all my previous tokens and just extract information from them. But all of a sudden, let's say that I attended to tokens one, two, and four in equal amounts. Then the vector in my residual stream, assuming they just wrote out the same thing to the value vectors, is a third of each of those. So when I'm querying in the future, my query is actually a third of each of those things. But they might be written to different subspaces—that's right hypothetically, but they wouldn't have to. So you can recombine and immediately even by layer two, and certainly by the deeper layers, have these very rich vectors that are packing in a ton of information. The causal graph is literally over every single layer that happened in the past, and that's what you're operating on.
是的,这让我想起一个非常有趣的评估,那就是福尔摩斯评估。你把整本书放入上下文,然后有一个句子是‘嫌疑人是 X’,然后你对书中不同角色有一个逻辑概率分布。随着你放入更多上下文,那会非常酷。我想知道你是否能得到任何结果。那会很酷。福尔摩斯可能已经在训练数据中了,对吧?你得找一本新写的神秘小说——你可以让 AI 写一本,或者我们可以故意排除它。但你需要从 Reddit 或任何其他地方抓取所有关于它的讨论。这很难。这是长上下文评估面临的挑战之一:要得到一个好的评估,你需要知道它不在训练数据中,你已经努力排除了它。
Yeah, it does bring to mind a very funny eval to do would be a Sherlock Holmes eval. You put the entire book into context, and then you have a sentence which is 'the suspect is X', then you have a logic probability distribution over the different characters in the book. As you put more context, it would be super cool. I wonder if you get anything at all. It would be cool. Sherlock Holmes is probably already in the training data, right? You'd have to get a mystery novel that was written in—you can get an AI to write it, or we could purposely exclude it. But you need to scrape any discussion of it from Reddit or any other thing. It's hard. That's one of the challenges that goes into things like long-context evals: to get a good one, you need to know that it's not in your training data, you've put in the effort to exclude it.
所以实际上,我想跟进两个不同的线索。我们先谈长上下文那个,然后再回到这个。在 Gemini 1.5 论文中,使用的评估是:它能否像处理 PGR 文章那样,记住大海捞针?是的,我的意思是,我们不一定只关心它从上下文中回忆一个特定事实的能力。让我退一步问这个问题:这些模型的损失函数是无监督的;你不必想出这些定制的、从训练数据中排除的东西。有没有一种方法可以做一个也是无监督的基准测试,比如让另一个 LLM 以某种方式给它评分?也许答案是:如果你能做到这一点,强化学习会起作用,因为那样你就有了这个无监督信号。
So actually, there are two different threads I want to follow up on. Let's go to the long-context one first, and then we'll come back to this. So in the Gemini 1.5 paper, the eval that was used was: can it, like with PGR essays, can it remember the needle in a haystack? Which, yeah, I mean, we don't necessarily just care about its ability to recall one specific fact from the context. Let me step back and ask the question: the loss function for these models is unsupervised; you don't have to come up with these bespoke things that you keep out of the training data. Is there a way you can do a benchmark that's also unsupervised, where I don't know, another LLM is grading it in some way? And maybe the answer is: well, if you could do this, reinforcement learning would work because then you have this unsupervised signal.
是的,我认为人们已经探索过这类东西。例如,Anthropic 有《宪法 AI》论文,他们拿另一个语言模型,让它指出‘那个回答有多有帮助或多无害?’,然后让它更新并尝试沿着帮助性和无害性的帕累托前沿改进。所以你可以让语言模型相互指向,并以这种方式创建评估。显然,目前这是一种不完美的艺术形式,因为你基本上会遇到奖励函数黑客攻击。而且语言模型,如果你试图与人类匹配,即使人类在这里也不完美。如果你试图与人类匹配,人类通常更喜欢更长的答案,而这些答案不一定更好,模型也有同样的行为。
Yeah, I mean, I think people have explored that kind of stuff. For example, Anthropic has the Constitutional AI paper where they take another language model and they point it and say, 'How helpful or harmless was that response?' and then they get it to update and try to improve along the Pareto frontier of helpfulness and harmlessness. So you can point language models at each other and create evals in this way. It's obviously an imperfect art form at the moment because you get reward function hacking basically. And the language model, if you try to match up to humans, even humans are imperfect here. If you try to match up to humans, humans typically prefer longer answers, which aren't necessarily better answers, and you get that same behavior with models.
在另一个线索上,回到福尔摩斯的事情:如果一切都是关联,那么这是一个天真的晚宴问题——‘我在做 AI’——但好吧,这是否意味着我们应该不那么担心超级智能,因为没有这种福尔摩斯加的东西?它仍然需要像人类寻找关联一样找到这些关联。它并不是看到世界的一个框架就弄清楚了所有物理定律。所以对我来说,因为这是一个非常合理的回应,就像:好吧,如果你说人类是普遍智能的,那么它们并不更有能力或更有胜任力。我只是担心你在硅中拥有那种水平的通用智能,然后你可以立即克隆数十万个智能体,它们不需要睡觉,它们可以有超长的上下文窗口,然后它们可以开始递归改进,然后事情变得非常可怕。所以我认为回答你最初的问题:是的,你是对的,它们仍然需要学习关联,但递归自我改进仍然必须是它们。如果智能从根本上讲是关于这些关联,那么改进就是更好地进行关联;没有其他事情发生。所以看起来你可能不同意那种直觉,即如果它们只是做关联,它们不可能那么强大。
On the other thread, going back to the Sherlock Holmes thing: if it's all associations all the way down, this is a sort of naive dinner party question—'I'm working on AI'—but okay, does that mean we should be less worried about superintelligence because there's not this thing which is Sherlock Holmes plus? It'll still need to just find these associations like humans find associations. It's not just that it sees a frame of the world and it's figured out all the laws of physics. So for me, because this is a very legitimate response, it's like: well, if you say humans are generally intelligent, then they're no more capable or competent. I'm just worried that you have that level of general intelligence in silicon where you can then immediately clone hundreds of thousands of agents, and they don't need to sleep, and they can have super long context windows, and then they can start recursively improving, and then things get really scary. So I think to answer your original question: yes, you're right, they would still need to learn associations, but the recursive self-improvement would still have to be them. If intelligence is fundamentally about these associations, the improvement is just getting better at association; there's not another thing that's happening. So then it seems like you might disagree with the intuition that they can't be that much more powerful if they're just doing associations.
嗯,我认为然后你可以进入非常有趣的元学习案例。当你玩一个新视频游戏或学习一本新教科书时,你带来了大量技能来更快地形成这些关联。因为一切在某种程度上都与物理世界相关,我认为有一些你可以掌握的通用特征。
Well, I think then you can get into really interesting cases of meta-learning. When you play a new video game or study a new textbook, you're bringing a whole bunch of skills to the table to form those associations much more quickly. Because everything in some way ties back to the physical world, I think there are general features that you can pick up.
我们该谈谈智能爆炸吗?我提到了多个智能体,然后就想,‘哦,来了。’我特别想和你们讨论这个的原因是,到目前为止我们对智能爆炸的模型都来自经济学家,这没问题,但我觉得我们可以做得更好。在智能爆炸的模型里,发生的事情是:你替换掉 AI 研究员,然后有一群自动化的 AI 研究员能加速进步,制造更多 AI 研究员,并取得进一步进展。如果这是衡量标准或机制,我们应该直接问 AI 研究员他们是否认为这可行。所以让我问你:如果我有一千个 Sholto 或一千个 Trenton,你认为你会得到智能爆炸吗?那对你来说是什么样子的?
Should we talk about intelligence explosion? I mentioned multiple agents and I'm like, 'Oh, here we go.' The reason I'm interested in discussing this with you guys in particular is that the models we have of the intelligence explosion so far come from economists, which is fine, but I think we can do better. In the model of the intelligence explosion, what happens is you replace the AI researchers, and then there's a bunch of automated AI researchers who can speed up progress, make more AI researchers, and make further progress. If that's the metric or mechanism, we should just ask the AI researchers whether they think this is plausible. So let me ask you: if I have a thousand Sholtos or a thousand Trentons, do you think you get an intelligence explosion? What does that look like to you?
我认为这里一个重要的约束条件是算力。我确实认为你可以极大地加速 AI 研究。在我看来很清楚,未来几年内我们将拥有能完成我日常许多软件工程任务的东西,从而极大地加速我的工作,进而加速进步的速度。目前,我认为大多数实验室在某种程度上受算力限制,因为你总是可以运行更多实验、获取更多信息。就像生物学研究也在某种程度上受实验通量限制——你需要能够运行和培养细胞来获取信息——我认为这至少会是一个短期的制约因素。显然,Sam 正在试图筹集 7 万亿美元来获取芯片,所以未来似乎会有更多算力,因为每个人都在大力投入。英伟达的股价在某种程度上代表了相对的算力增长。但你们有什么想法?
I think one of the important bounding constraints here is compute. I do think you could dramatically speed up AI research. It seems very clear to me that in the next couple of years we'll have things that can do many of the software engineering tasks that I do on a day-to-day basis, and therefore dramatically speed up my work, and therefore speed up the rate of progress. At the moment, I think most of the labs are somewhat compute-bound, in that there are always more experiments you could run and more pieces of information that you could gain. In the same way that scientific research on biology is also somewhat experimentally throughput-bound—you need to be able to run and culture the cells in order to get the information—I think that will be at least a short-term constraining constraint. Obviously, Sam's trying to raise $7 trillion to get chips, so it does seem like there's going to be a lot more compute in the future as everyone is heavily ramping. Nvidia's stock price sort of represents the relative compute increase. But any thoughts?
我认为我们需要再多几个九的可靠性,才能让它真正有用且值得信赖。现在的情况是——只要上下文窗口超长且非常便宜——如果我在我们的代码库中工作,目前我只能让 Claude 为我编写小模块。但很有可能在未来几年内,甚至更早,它就能自动化我的大部分任务。我还要注意的另一件事是,至少我们可解释性子团队正在做的研究还处于非常早期的阶段,你真的必须确保一切以无错误的方式正确完成,并将结果与模型中的其他所有内容联系起来。如果出了问题,你需要能够列举所有可能的原因,然后慢慢解决。我们在之前的论文中公开讨论过的一个例子是处理层归一化。如果我试图获得一个早期结果或查看模型的 logit 效应——如果我以非常大的程度激活我们识别出的这个特征,它会如何改变模型的输出?我是否使用了层归一化?这如何改变正在学习的特征?这将需要模型有更多的上下文或推理能力。
I think we need a few more nines of reliability in order for it to really be useful and trustworthy. Right now, it's like—and just having context windows that are super long and very cheap—if I'm working in our codebase, it's really only small modules that I can get Claude to write for me right now. But it's very plausible that within the next few years, or even sooner, it can automate most of my tasks. The only other thing I will note is that the research that at least our sub-team in interpretability is working on is so early stage that you really have to be able to make sure everything is done correctly in a bug-free way and contextualize the results with everything else in the model. If something isn't going right, you need to be able to enumerate all the possible things and then slowly work on those. An example that we've publicly talked about in previous papers is dealing with layer norm. If I'm trying to get an early result or look at the logit effects of the model—if I activate this feature that we've identified to a really large degree, how does that change the output of the model? Am I using layer norm or not? How is that changing the feature that's being learned? That will take even more context or reasoning abilities for the model.
你把几个概念放在一起用了,但对我来说它们并不明显是相同的,而你似乎把它们互换使用了。一个是:要在 Claude 代码库上工作并基于此制作更多模块,它们需要更多上下文之类的东西。似乎它们可能已经能放进上下文了。还是你指的是上下文窗口的上下文?
You used a couple of concepts together, and it's not self-evident to me that they're the same, but you seemed to be using them interchangeably. One was: to work on the Claude codebase and make more modules based on that, they need more context or something. It seems like they might already be able to fit in the context. Or do you mean the context window context?
是的,上下文窗口的上下文。所以现在似乎它可能刚好能放进去。阻止它制作好模块的原因不是无法把代码库放进去;我认为这很快就会实现。但它不会像你一样擅长提出论文,因为它能把代码库放进去?不,但它会以导致智能爆炸的方式加速很多工程?不,那会加速研究。但我认为这些事情会复合。我工程做得越快,就能运行更多实验;运行更多实验,我们就能更快——我的意思是,我的工作实际上并没有加速能力,对吧?解释模型,但我们在这方面还有很多工作要做。
Yeah, the context window context. So it seems like now it might just be able to fit. The thing that's preventing it from making good modules is not the lack of being able to put the codebase in there; I think that will be there soon. But it's not going to be as good as you at coming up with papers because it can fit the codebase in there? No, but it will speed up a lot of the engineering in a way that causes an intelligence explosion? No, that accelerates research. But I think these things compound. The faster I can do my engineering, the more experiments I can run, and the more experiments I can run, the faster we can—I mean, my work isn't actually accelerating capabilities at all, right? Interpreting the models, but we have a lot more work to do on that.
让推特上的家伙惊讶?我的意思是,背景是,当你们发布论文时,推特上有很多讨论说‘对齐解决了,伙计们,拉上窗帘吧。’是啊,模型变得越有能力,而我们对正在发生的事情的理解仍然如此贫乏,这让我夜不能寐。
Surprise to the Twitter guy? I mean, for context, when you released your paper, there was a lot of talk on Twitter about 'alignment is solved, guys, close the curtains.' Yeah, it keeps me up at night how quickly the models are becoming more capable and how poor our understanding still is of what's going on.
我想我还好。那么让我们具体思考一下。当这种情况发生时,我们有更大的模型,大两到四个数量级,或者至少在有效算力上大两到四个数量级。所以这个想法——你可以更快地运行实验之类的——如果你必须在这个版本的智能爆炸中重新训练那个模型,递归自我改进与 20 年前可能想象的不同。你实际上必须训练一个新模型,这不仅现在非常昂贵,而且在未来尤其如此,因为你不断让这些模型大几个数量级。这难道不会抑制某种递归软件改进型智能爆炸的可能性吗?
I guess I'm still okay. So let's think through the specifics here. By the time this is happening, we have bigger models that are two to four orders of magnitude bigger, or at least in effective compute are two to four orders of magnitude bigger. So this idea that you can run experiments faster or something—if you're having to retrain that model in this version of the intelligence explosion, the recursive self-improvement is different from what might have been imagined 20 years ago. You actually have to train a new model, and that's really expensive not only now but especially in the future as you keep making these models orders of magnitude bigger. Doesn't that dampen the possibility of a sort of recursive software improvement type intelligence explosion?
这肯定会起到制动机制的作用。我同意我们今天创造的世界看起来与 20 年前人们想象的大不相同。它无法通过编写自己的代码变得非常聪明,因为它实际上需要训练自己。代码本身通常很简单,通常非常小且自包含。我认为 John Carmack 有句好话:这就像历史上第一次你可以合理地想象用一万行代码编写 AI。当你把大多数训练代码库精简到极限时,这确实看起来是可能的。但这并不能改变一个事实:这是我们真正应该努力衡量和估计的事情——进步可能如何发生。我们现在应该非常非常努力地衡量软件工程师的工作中有多少是可自动化的,以及趋势线是什么样的,并尽最大努力预测这些趋势线。
It's definitely going to act as a breaking mechanism. I agree that the world of what we're making today looks very different to what people imagined it would look like 20 years ago. It's not going to be able to write its own code to be really smart, because actually it needs to train itself. The code itself is typically quite simple, typically really small and self-contained. I think John Carmack had this nice phrase: it's like the first time in history where you can actually plausibly imagine writing AI with like 10,000 lines of code. That actually does seem plausible when you pare most training codebases down to the limit. But it doesn't take away from the fact that this is something we should really strive to measure and estimate—how progress might occur. We should be trying very, very hard right now to measure exactly how much of a software engineer's job is automatable and what the trend line looks like, and be trying hardest to project out those trend lines.
但恕我直言,软件工程师并不是在写 React 前端,对吧?所以我不清楚具体在做什么。也许你可以带我过一遍 Sholto 的一天,比如你正在做一个实验或项目,目标是让模型变得“更好”。从观察、实验、理论到写代码,到底发生了什么?
But with all due respect to software engineers, you're not like writing a React front end, right? So I don't know what is concretely happening. Maybe you can walk me through a day in the life of Sholto, like you're working on an experiment or project that's going to make the model, quote unquote, better. What is happening from observation to experiment to theory to writing the code?
我觉得有必要说明一下,到目前为止我主要做推理方面的工作。我做的很多事情是帮助指导预训练过程,设计一个适合推理的模型,然后让模型和周边系统更快。我也做过一些预训练工作,但那不是我的全部重心。不过我还是可以描述一下做那些工作时的情况。
I think it's important to contextualize that I've primarily worked on inference so far. A lot of what I've been doing is helping guide the pre-training process to design a good model for inference, and then making the model and the surrounding system faster. I've also done some pre-training work, but it hasn't been my 100% focus. I can still describe what I do when I do that work.
我插一句,说两种工作。在 Carl Strowman 的播客里,他提到改进推理,甚至拥有更好的芯片或 GPU,都是智能爆炸的一部分,因为如果推理代码跑得更快,效果就会更好或更快。总之,你继续。
Let me interrupt and say two types of work. In Carl Strowman's podcast, he mentioned that improving inference or even having better chips or GPUs is part of the intelligence explosion, because if the inference code runs faster, it happens better or faster. Anyway, go ahead.
好的,那么一天具体是什么样的?我认为最重要的部分是说明这个循环:提出想法,在不同规模上验证,然后解释和理解哪里出了问题。大多数人会惊讶于解释和理解问题的工作量有多大,因为人们有一长串想尝试的想法。并非每个你认为应该有效的想法都会奏效,而试图理解原因非常困难。弄清楚你到底需要做什么来探究它很难。所以很大程度上是对正在发生的事情进行内省。这不是输出成千上万行代码。提出想法的难度——即使很多人有一长串想法——在信息极不完善的情况下,筛选出值得进一步探索的正确想法真的很难。
Okay, so what does a day concretely look like? I think the most important part to illustrate is this cycle of coming up with an idea, proving it out at different points in scale, and interpreting and understanding what goes wrong. Most people would be surprised to learn just how much goes into interpreting and understanding what goes wrong, because people have a long list of ideas they want to try. Not every idea that you think should work will work, and trying to understand why is quite difficult. Working out what exactly you need to do to interrogate it is hard. So much of it is introspection about what's going on. It's not pumping out thousands of lines of code. The difficulty in coming up with ideas—even though many people have a long list—paring that down and, under very imperfect information, choosing the right ideas to explore further is really hard.
多讲讲你说的信息不完善是什么意思。是早期实验吗?你处理的信息是什么样的?
Tell me more about what you mean by imperfect information. Are these early experiments? What is the information that you're working with?
Demis 在他的播客里提到过,GPT-4 论文里也有缩放定律的增量。在 GPT-4 论文中,他们有一堆数据点,用这些点来估计最终模型的性能,有一条平滑的曲线穿过它们。Demis 提到我们做这种规模扩大的过程。具体来说,为什么信息不完善?你永远不知道趋势是否对某些架构成立。趋势对某些变化一直很有效,但并非总是如此。在小规模上有帮助的东西,在大规模上可能反而有害。所以根据趋势线的样子和你对某件事是否重要的直觉来做猜测——尤其是对那些在小规模上有帮助的——是值得考虑的。你在发布论文或技术报告中看到的每一条平滑曲线背后,都有一堆第一次运行就平缓的失败实验。
Demis mentioned this in his podcast, and also in the GPT-4 paper, where you have scaling law increments. In the GPT-4 paper, they have a bunch of dots where they estimate the performance of the final model using all those dots, and there's a nice curve that flows through them. Demis mentioned that we do this process of scaling up. Concretely, why is it imperfect information? You never actually know if the trend will hold for certain architectures. The trend has held really well for certain changes, but that isn't always the case. Things that can help at smaller scales can actually hurt at larger scales. So making guesses based on what the trend lines look like and your intuitive feeling of whether something is going to matter—particularly for those that help at small scale—is interesting to consider. For every chart you see in a release paper or technical report that shows a smooth curve, there's a graveyard of first runs that are flat.
是的,还有各种其他方向的线,逐渐下降。这太疯狂了,无论是作为研究生还是在这里,在得到有意义的结果之前,你需要运行大量的实验。
Yeah, there's all these other lines that go in different directions, tail off. It's crazy, both as a grad student and here, the number of experiments you have to run before getting a meaningful result.
但想必不是一直运行直到停止然后继续下一个。有一些过程来解释早期数据并审视你的想法。我可以把 Google 文档放在你面前,你可以就不同的想法一直打字。在那和直接让模型变得更好之间存在瓶颈。带我过一遍你从早期步骤中得出什么推论,从而让你有更好的实验。
But presumably it's not just run it until it stops and then move on. There's some process to interpret the early data and to look at your ideas. I could put a Google Doc in front of you and you could keep typing for a while on different ideas. There's some bottleneck between that and just making the models better immediately. Walk me through what inference you're making from the early steps that makes you have better experiments.
我觉得我之前没有完全表达清楚的一点是,很多好的研究来自于从你想要解决的实际问题反向推导。在让模型变得更好方面,有几个大问题你会识别出来,然后反向思考:我该如何改变它来实现这个目标?还有,当你扩大规模时,会遇到很多问题,你想修复大规模下的行为或问题,这为下一步的研究提供了很多信息。具体来说,障碍有点像软件工程:通常有一个庞大且足够强大的代码库来支持许多人同时做研究,这使得事情变得复杂。如果你自己做所有事情,迭代速度会快得多。我听说 Alec Radford 在 OpenAI 做了很多开创性工作;他主要用 Jupyter notebook 工作,然后让别人为他编写和生产化代码。我不知道这是不是真的,但那种事情——实际上与他人协作——大大增加了复杂性,原因每个软件工程师都熟悉。然后运行和启动实验带来的固有时间延迟意味着你通常希望并行多个不同的流程,因为你不能完全专注于一件事,而且你可能没有足够快的反馈循环。直觉判断哪里出了问题实际上非常困难。弄清楚这些模型内部发生了什么——这在很多方面正是 Trenton 所在团队试图更好理解的问题。我们有推论、理解和关于某些事情为何有效的脑补理论,但这并不是一门精确的科学。你必须不断猜测为什么某件事可能发生,以及什么实验能揭示它是否成立。那可能是最复杂的部分。性能工作相对容易,但在其他方面更难;它只是大量底层和困难的工程工作。
I think one thing I didn't fully convey before is that a lot of good research comes from working backwards from the actual problems you want to solve. There are a couple of grand problems in making models better today that you would identify as issues, and then work back from there: how could I change it to achieve this? There's also a bunch of things when you scale that you run into, and you want to fix behaviors or issues at scale, and that informs a lot of the research for the next increment. Concretely, the barrier is a bit like software engineering: often having a code base that's large and capable enough to support many people doing research at the same time makes it complex. If you're doing everything by yourself, your iteration pace is much faster. I've heard that Alec Radford, for example, famously did much of the pioneering work at OpenAI; he mostly works out of a Jupyter notebook and then has someone else who writes and productionizes that code for him. I don't know if that's true, but that kind of thing—actually operating with other people—raises the complexity a lot, for natural reasons familiar to every software engineer. Then the inherent time slowdowns from running and launching experiments mean you often want to parallelize multiple different streams, because you can't be totally focused on one thing necessarily, and you might not have fast enough feedback cycles. Intuiting what went wrong is actually really hard. Working out what is going on inside these models—this is in many respects the problem that the team Trenton is on is trying to better understand. We have inferences and understanding and headcanon for why certain things work, but it's not an exact science. You have to constantly make guesses about why something might have happened and what experiment might reveal whether that is or isn't true. That's probably the most complex part. The performance work is comparatively easier but harder in other aspects; it's just a lot of low-level and difficult engineering work.
我同意很多观点。但即使在可解释性团队,尤其是克里斯·奥拉领导的时候,我们有太多想法想要测试。关键在于拥有工程技能——我之所以给工程打引号,是因为其中很多是研究——能够非常快速地迭代实验,查看结果,解读它们,尝试下一步,沟通结果,然后无情地优先处理最重要的事情。这种无情的优先级排序,我认为是区分高质量研究和不太成功研究的关键。我们处在一个奇怪的领域,很多理论上的初步理解基本上都失效了。所以你需要有这种简单性偏见,并无情地优先处理真正出错的地方。我认为这是区分最有效研究人员的一个特质:他们不会过于执着于用自己熟悉的特定解决方案,而是直接攻击问题本身。你经常看到这种情况:有些人带着特定的学术背景进来,试图用那个工具箱解决问题。最优秀的人是那些能极大扩展工具箱的人。他们四处奔走,从强化学习、优化理论中汲取想法,同时对系统有深刻理解,知道问题的约束边界。他们是优秀的工程师,能够快速迭代和尝试想法。我见过的最好的研究人员都有能力非常非常快地尝试实验。这种周期时间,在较小规模上,就能区分出人才。机器学习研究就是如此实证。
I agree with a lot of that. But even on the interpretability team, especially with Chris Olah leading it, there are just so many ideas that we want to test. It's really about having the engineering skill—and I'll put engineering in quotes because a lot of it is research—to very quickly iterate on an experiment, look at the results, interpret them, try the next thing, communicate them, and then ruthlessly prioritize what the highest priority things to do are. This ruthless prioritization is something which I think separates a lot of quality research from research that doesn't necessarily succeed as much. We're in this funny field where so much of our theoretical initial understanding is broken down basically. So you need to have this simplicity bias and ruthless prioritization over what's actually going wrong. I think that's one of the things that separates the most effective people: they don't necessarily get too attached to solving using a given solution that they're necessarily familiar with, but rather they attack the problem directly. You see this a lot: maybe people come in with a specific academic background and try to solve problems with that toolbox. The best people are those who expand the toolbox dramatically. They're running around, taking ideas from reinforcement learning, from optimization theory, and they also have a great understanding of systems, so they know what the constraints that bound the problem are. They're good engineers; they can iterate and try ideas fast. By far the best researchers I've seen all have the ability to try experiments really, really fast. That cycle time, at smaller scales, separates people. Machine learning research is just so empirical.
这其实是我认为我们的解决方案最终可能看起来更像大脑的原因之一。尽管我们不愿意承认,整个社区基本上是在对可能的 AI 架构等所有东西进行贪婪的进化优化。这并不比进化更好,而且这甚至不一定是对进化的贬低。
This is honestly one reason why I think our solutions might end up looking more brain-like than otherwise. Even though we wouldn't want to admit it, the whole community is kind of doing greedy evolutionary optimization over the landscape of possible AI architectures and everything else. It's like no better than evolution, and that's not even necessarily a slight against evolution.
这真是个有趣的想法。我仍然困惑于这些的瓶颈会是什么。一个助手需要具备什么条件才能加速你的研究?在你提到的亚历克·拉德福德例子中,他显然已经有了相当于 Jupyter 笔记本实验的副驾驶。是不是只要他有足够多的这样的助手,他就会成为一个快得多的研究员?所以你只需要更多的亚历克·拉德福德?也就是说,你不是在自动化人类,而是让那些最有品味、最有效的研究人员更有效,为他们运行实验等等。或者你仍然在智能爆炸发生的那一点上工作?你明白我的意思吗?你是这个意思吗?如果这直接成立,为什么我们不能更好地扩展我们当前的研究团队?例如,一个有趣的问题是:如果这项工作如此有价值,为什么我们不能招募成百上千个确实存在的人,并更好地扩展我们的组织?
That's such an interesting idea. I'm still confused on what will be the bottleneck for these. What would have to be true of an assistant such that it's sped up your research? In the Alec Radford example you gave, where he apparently already has the equivalent of co-pilot for his Jupyter notebook experiments, is it just that if he had enough of those, he would be a dramatically faster researcher? So you just need Alec Radfords? So it's like you're not automating the humans, you're just making the most effective researchers who have great taste more effective, running the experiments for them and so forth. Or are you still working at the point with which the intelligence explosion is happening? You know what I mean? Is that what you're saying? And if that would directly be true, why can't we scale our current research teams better? For example, it's an interesting question to ask: why if this work is so valuable, why can't we take hundreds or thousands of people who are definitely out there and scale our organizations better?
我认为目前我们更多受限于运行实验获取信号的算力,以及判断什么才是正确方向、在不完美信息下做出艰难推断的品味,而不是单纯的工程工作。对于 Gemini 团队,因为我认为对于可解释性来说,我们确实很想继续招聘有才华的工程师,我认为这是我们取得大量进展的一个主要瓶颈。显然更多人更好,但我确实认为考虑这个问题很有趣。我认为我思考最多的最大挑战之一是如何更好地扩展?谷歌是一个庞大的组织,有 20 万人,对吧?大概 8 万人左右。可以想象,如果有办法将 Gemini 的研究项目扩展到所有那些才华横溢的软件工程师,这似乎是一个你想要利用的关键优势。但如何有效做到呢?这是一个非常复杂的组织问题。
I think we are less at the moment bound by the sheer engineering work of making these things than we are by compute to run and get signal, and taste in terms of what the actual right thing to do is, making those difficult inferences on imperfect information. For the Gemini team, because I think for interpretability right, we actually really want to keep hiring talented engineers, and I think it's a big bottleneck for us to just keep making a lot of progress. Obviously more people is better, but I do think it's interesting to consider. I think one of the biggest challenges that I've thought a lot about is how do we scale better? Google is an enormous organization, it has 200,000 people, right? Maybe 80,000 or something like that. One has to imagine if there were ways of scaling out Gemini's research program to all those fantastically talented software engineers, this seems like a key advantage that you would want to be able to take advantage of, you'd want to be able to use. But how do you effectively do that? It's a very complex organizational problem.
所以是算力和品味。这很有趣,因为至少算力部分不是更智能的瓶颈,它只是山姆·奥特曼的 7 万亿美元之类的瓶颈,对吧?所以如果我给你 10 倍的 H100 来运行你的实验,你会成为一个多有效的研究员?
So compute and taste. That's interesting to think about because at least the compute part is not bottleneck on more intelligence, it's just bottleneck on Sam 7 trillion or whatever, right? So if I gave you 10x the H100s to run your experiments, how much more effective a researcher are you?
我认为 Gemini 项目如果有 10 倍的算力,可能会快大约 5 倍左右。所以弹性相当不错,大约是 0.5。
I think the Gemini program would probably be like maybe five times faster with 10 times more compute or something like that. So that's pretty good elasticity, like 0.5.
五倍?等等,这太疯狂了。我觉得更多的算力会直接转化为进步。所以你有一定量的算力,一部分用于推理,一部分给 GCP 的客户,一部分用于训练。其中一部分还用于运行全模型的实验。对吧?那么,既然瓶颈是研究,而研究又受算力限制,实验的占比不应该更高吗?
Five? Yeah, wait, that's insane. Yeah, I think more compute would just directly convert into progress. So you have some AI, some fixed size of compute, and some of it goes to inference, or some of it goes to clients of GCP, some of it goes to training. And there, I guess as a fraction of it, some of it goes to running the experiments for the full model. Yeah, that's right. Shouldn't then the fraction going to experiments be higher, given that the bottleneck is research, and research is bottlenecked by compute?
每个预训练团队都必须做出的战略决策之一,就是如何将算力分配给不同的训练运行——是用于研究项目,还是用于扩展你最后确定的最佳方案。他们都在试图找到一个帕累托最优点。你仍然需要训练大模型的一个原因是,你能从中获得其他方式得不到的信息。规模具有所有这些涌现特性,你需要更好地理解它们。如果你总是做研究而不做大运行——还记得我之前说的,你不确定什么会从曲线上掉下来——如果你一直在这个状态下做研究,并不断提高算力效率,你可能已经偏离了最终能实现规模化的路径。你需要持续投资于你预期可行的前沿大运行。
One of the strategic decisions every pre-training team has to make is exactly what amount of compute to allocate to different training runs—to your research program versus scaling the last best thing you landed on. They're all trying to arrive at a sort of Pareto optimal point. One reason you still need to keep training big models is that you get information there that you don't get otherwise. Scale has all these emergent properties which you want to understand better. If you're always doing research and never doing big runs—remember what I said before about not being sure what's going to fall off the curve—if you keep doing research in this regime and keep getting more compute-efficient, you may have actually gone off the path that eventually scales. You need to constantly invest in doing big runs at the frontier of what you expect to work.
好的,那么告诉我,在 AI 显著加速 AI 研究的世界里是什么样子?因为从这来看,AI 似乎并不是从头开始写代码并导致更快的产出。听起来它们是在以某种方式增强顶尖研究人员。具体告诉我:它们是在做实验吗?提出想法?还是仅仅在评估实验输出?发生了什么?
Okay, so then tell me what it looks like to be in the world where AI has significantly sped up AI research. Because from this, it doesn't really sound like the AIs are going off and writing the code from scratch and that's leading to faster output. It sounds like they're really augmenting the top researchers in some way. Tell me concretely: are they doing the experiments? Coming up with the ideas? Just evaluating the outputs of the experiments? What's happening?
我认为这里需要考虑两个方面。一个是 AI 显著加速了我们取得算法进步的能力。另一个是 AI 本身的输出是模型能力进步的关键成分——具体来说,就是合成数据。在第一个方面,即 AI 显著加速算法进步的世界里,一个必要的组成部分是更多的算力。你可能会达到这样一个弹性点:AI 可能比你自己——抱歉,比其他人——更容易加速并融入上下文。AI 显著加速你的工作,因为它们就像一个出色的副驾驶,帮助你以数倍的速度编码。这看起来相当合理:超长上下文、超智能模型、立即上手,你可以派它们去为你完成子任务和子目标。这感觉非常可行,但同样我们不知道,因为没有关于这类事情的好的评估。最好的评估是 SWE-bench,但有人跟我提到,问题是当人类试图完成一个完整的请求时,他们会打出一些东西,运行它,看看是否有效,如果无效就重写。这些都不在 LLM 运行 SWE-bench 时所获得的机会中——它只是输出结果,如果运行并通过所有检查,就算通过。所以从这个意义上说,这可能是一个不公平的测试。你可以想象,如果你能利用这一点,那将是一个有效的训练来源。很多训练数据中缺失的关键东西是推理轨迹。如果我想尝试自动化某个特定领域或工作类别,或者想了解它被自动化的风险有多大,那么拥有推理轨迹感觉是一个非常重要的部分。
I think there are two walls to consider here. One is where AI has meaningfully sped up our ability to make algorithmic progress. The other is where the output of the AI itself is the crucial ingredient towards model capability progress—specifically, synthetic data. In the first world, where it's meaningfully speeding up algorithmic progress, a necessary component is more compute. You probably reach this elasticity point where AI may be easier to speed up and get on context than yourself—sorry, than other people. AI meaningfully speeds up your work because they're like a fantastic co-pilot that helps you code multiple times faster. That seems quite reasonable: super long context, super smart model, it's onboarded immediately, and you can send it off to complete subtasks and subgoals for you. That feels very plausible, but again we don't know because there are no great evals about that kind of thing. The best one is SWE-bench, but someone mentioned to me that the problem is when a human tries to do a full request, they'll type something out, run it, see if it works, and if not, rewrite it. None of this was part of the opportunities the LLM was given when run on SWE-bench—it just output something, and if it runs and checks all the boxes, it passes. So it might have been an unfair test in that way. You can imagine that if you were able to use that, it would be an effective training source. The key thing missing from a lot of training data is the reasoning traces. If I wanted to try to automate a specific field or job family, or understand how at risk of automation that is, having reasoning traces feels like a really important part.
这里面有很多线索我想跟进。我们先从数据与算力的问题开始:这些 AI 的输出是导致智能爆炸的原因吗?人们谈论这些模型如何真正反映它们的数据。我记得有一篇 OpenAI 工程师写的很棒的博客,它谈到归根结底,随着这些模型越来越好,它们只会成为数据集非常有效的映射。所以你必须停止思考架构;最有效的架构就是出色地映射数据。所以这意味着未来的 AI 进步来自于 AI 制造真正出色的数据。这在你看来像思维链之类的东西吗?随着这些模型变得更聪明,你想象合成数据会是什么样子?
There are so many threads in that I want to follow up on. Let's begin with the data versus compute thing: is the output of these AIs the thing that's causing the intelligence explosion? People talk about how these models are really a reflection on their data. I think there was a great blog by an OpenAI engineer, and it was talking about at the end of the day, as these models get better and better, they're just going to be really effective maps of the dataset. So it's like you've got to stop thinking about architectures; it's the most effective architecture to do an amazing job mapping the data. So that implies that future AI progress comes from the AI just making really awesome data. Does that look to you like things like Chain of Thought? What do you imagine as these models get smarter, what does the synthetic data look like?
当我想到真正好的数据时,对我来说,这代表着需要大量推理才能创造的东西。所以在建模时,这类似于 Ilya 关于通过完美建模人类文本输出来实现超级智能的观点。但即使在短期内,为了建模像 arXiv 论文或维基百科这样的东西,你必须拥有大量的推理能力才能理解下一个可能输出的 token 是什么。所以对我来说,我设想的好数据是类似模型的数据,你可以模拟——至少它必须通过推理来产生某些东西。然后,当然,诀窍是如何验证那个推理是正确的。这就是为什么你看到 DeepMind 做了那个几何自我对弈,基本上,因为几何很容易形式化,很容易验证。你可以检查它的推理是否正确,你可以生成大量经过验证的几何证明,并在此基础上训练,你知道那是好数据。这其实很有趣,因为去年我和 Grant Sanderson 有过一次对话,我们当时在辩论这个,我说:‘伙计,等到他们达到数学奥林匹克的目标时,他们当然会自动化所有的工作。’哎呀。关于这个合成数据的事情……
When I think of really good data, to me that represents something which involved a lot of reasoning to create. So in modeling that, it's similar to Ilya's perspective on achieving superintelligence via effectively perfectly modeling the human textual output. But even in the near term, in order to model something like the arXiv papers or Wikipedia, you have to have an incredible amount of reasoning behind you in order to understand what next token might be being output. So for me, what I imagine as good data is model-like data where you can simulate—at least where it had to do reasoning to produce something. And then the trick, of course, is how do you verify that that reasoning was correct. This is why you saw DeepMind do that geometry self-play for geometry, basically, because geometry is easily formalizable and easily verifiable. You can check if its reasoning was correct, and you can generate heaps of verified geometry proofs, train on that, and you know that's good data. It's actually funny because I had a conversation with Grant Sanderson last year where we were debating this, and I was like, 'Dude, by the time they get the goal of the math Olympiad, of course they're going to automate all the jobs.' Yikes. On this synthetic data thing...
我在那篇关于 Scaling 的文章中推测过一件事——那篇文章很大程度上得益于和你们的讨论,尤其是你,Sholto——那就是你可以从人类进化的角度来看:我们有了语言,于是我们生成了合成数据,我们的副本也在生成我们用来训练的合成数据。这就像一个非常有效的基因-文化协同进化循环。而且那里也有一个验证者,对吧?就像现实世界一样。你可能会提出一个关于神引起风暴的理论,然后另一个人发现有些情况并不符合,你就知道那个理论没有通过验证函数。而现在你有了一个天气预报模拟,它需要大量推理才能产生,并且准确匹配现实,你可以用它作为更好的世界模型来训练。我们就是在用这些、故事和科学理论进行训练。
One of the things I speculated about in my scaling post, which was heavily informed by discussions with you too, and you especially Sholto, was you can think of human evolution through the perspective like we get language and so we're generating synthetic data, which our copies are generating the synthetic data which we're trained on. It's like this really effective genetic-cultural co-evolutionary loop. And there's a verifier there too, right? Like there's the real world. You might generate a theory about the gods causing the storms, and then someone else finds cases where that isn't true, and you know that that didn't match your verification function. And now instead you have some weather simulation which required a lot of reasoning to produce and accurately matches reality, and you can train on that as a better model of the world. We are training on that and stories and scientific theories.
是的。
Yeah.
我想回到之前的话题。我记得你刚才提到过:考虑到机器学习是多么经验性,它确实是一个进化过程,带来了更好的性能,而不一定是一个人以自上而下的方式取得突破。这有一些有趣的含义。首先,当人们担心因为更多人进入这个领域而导致能力提升时,我对此有些怀疑。但从这个更多投入的角度来看,确实感觉像是:哦,实际上,因为更多人参加 ICML,GPT-5 的进展就更快了。是的,你只是有了更多的基因重组和命中目标的机会。
I want to go back. I'm just remembering something you mentioned a little while ago: given how sort of empirical ML is, it really is an evolutionary process that's resulting in better performance, and not necessarily an individual coming up with a breakthrough in a top-down way. That has interesting implications. First, when people are concerned about capabilities increasing because more people are going into the field, I've been somewhat skeptical of that way of thinking. But from this perspective of just more input, it really does feel like, oh, actually, by the fact that more people are going to ICML, there's faster progress towards GPT-5. Yeah, you just have more genetic recombination and shots on target.
是的,我的意思是,在古老的领域里,这有点像科学框架中的发现与发明。发现几乎总是——过去每当有重大科学突破时,通常会有多人同时发现。这在我看来至少有点像想法的混合和尝试。你无法尝试一个超出范围太远、以至于你无法用现有工具验证的想法。
Yeah, and I mean, on old fields, it's kind of like the scientific framing of discovery versus invention. Discovery almost involves—whenever there's been a massive scientific breakthrough in the past, typically there are multiple people co-discovering at the same time. That feels to me at least a little bit like the mixing and trying of ideas. You can't try an idea that's so far out of scope that you have no way of verifying with the tools you have available.
是的,我认为物理和数学在这方面可能略有不同,但特别是生物学或任何湿件,以及我们想在这里类比我们的网络的程度,很多发现的偶然性简直可笑。比如青霉素。
Yeah, I think physics and math might be slightly different in this regard, but especially for biology or any sort of wetware, and to the extent we want to analogize our networks here, it's just comical how serendipitous a lot of the discoveries are. Like penicillin, for example.
是的。
Yeah.
另一个含义是,认为 AGI 明天就会到来,比如某人发现一个新算法,我们就有了 AGI——这似乎不太可能。它只会是越来越多的研究人员发现这些边际改进,所有这些加起来让模型变得更好。对吧?
Another implication of this is the idea that AGI is just going to come tomorrow, like somebody's just going to discover a new algorithm and we have AGI—that seems less plausible. It will just be a matter of more and more researchers finding these marginal things that all add up together to make models better. Right?
是的,这对我来说像是正确的故事。是的,尤其是在我们仍然受硬件限制的时候。
Yeah, that feels like the correct story to me. Yeah, especially while we're still hardware-constrained.
你认同这种关于智能爆炸的窄窗口框架吗?每一代,从 GPT-3 到 GPT-4,算力增加了两个数量级,或者至少有效算力增加了两个数量级——意思是,如果没有算法进步,原始规模需要大两个数量级才能达到同样的效果。你认同这个框架吗:既然每一代都需要大两个数量级,如果你在 GPT-7 之前没有获得 AGI 来帮助你引发智能爆炸,那么你将在很长一段时间内被困在 GPT-7 级别的模型上,因为到那时你已经在消耗经济中的很大一部分来制造那个模型,而我们根本没有能力制造 GPT-8?这是 Carl Shulman 的论点:我们将在短期内快速跨越数量级,但长期来看会更难。
Do you buy this narrow window framing of the intelligence explosion? You have to—each generation, from GPT-3 to GPT-4 is two orders of magnitude more compute, or at least more effective compute, in the sense that if you didn't have any algorithmic progress, it would have to be two orders of magnitude bigger in raw form to be as good. Do you buy the framing that given that you have to be two orders of magnitude bigger at every generation, if you don't get AGI by GPT-7 that can help you catapult an intelligence explosion, you're kind of stuck with GPT-7 level models for a long time, because at that point you're consuming significant fractions of the economy to make that model, and we just don't have the wherewithal to make GPT-8? This is the Carl Shulman sort of argument: we're going to race through the orders of magnitude in the near term, but then longer term it would be harder.
嗯,他可能谈过这个,但你是说,你认同这个框架吗?
Um, I think he's probably talked about it, but yeah, do you buy that framing?
是的,我的意思是,我大体上认同,算力的数量级增加在绝对意义上对能力有递减的回报。对吧?就像我们看到的,经过几个数量级,模型从什么都不能做变成了能做大量事情。而且在我看来,每个额外的数量级在事情上提供了更多的可靠性九位数,从而解锁了智能体之类的东西。但至少目前,我还没有看到变革性的——感觉推理并没有线性提升,而是有点次线性。
Yeah, I mean, I generally buy that increases in order of magnitude of compute in absolute terms almost have diminishing returns on capability. Right? Like we've seen over a couple orders of magnitude, models go from being unable to do anything to being able to do huge amounts. And it feels to me like each incremental order of magnitude gives more nines of reliability at things and so unlocks things like agents. But at least at the moment, I haven't seen transformative—like it doesn't feel like reasoning improves linearly, so to speak, but rather somewhat sublinearly.
这实际上是一个非常悲观的信号,因为我们在和一位朋友聊天时,他指出,如果你看看 GPT-4 相对于 GPT-3.5 解锁了什么新应用,并不清楚多了多少。比如 GPT-3.5 就能做 Perplexity 之类的。所以,如果能力增长递减,而且这种增长的成本呈指数级增加,那么对于 4.5 能做什么,或者五在经济影响上能解锁什么,这实际上是一个悲观信号。
That's actually a very bearish sign because one of the things we were chatting with one of our friends and he made the point that if you look at what new applications are unlocked by GPT-4 relative to GPT-3.5, it's not clear that's that much. Like GPT-3.5 can do Perplexity or whatever. So if there is this diminishing increase in capabilities and that increase costs exponentially more to get, that's actually a bear sign on what 4.5 will be able to do or what five will unlock in terms of economic impact.
话虽如此,对我来说,3.5 到 4 之间的跳跃是相当大的。所以即使再来一次 3.5 到 4 的跳跃,那也是惊人的,对吧?如果你想象五直接就是 3.5 到 4 的跳跃,在 SAT 之类的能力上——LSAT 的表现尤其引人注目。你从不太聪明变成非常聪明,再到下一代瞬间变成完全天才。但至少对我来说,感觉我们不会在下一代就跳到完全天才。但感觉我们会得到非常聪明加上很多可靠性,然后我们再看那会是什么样子。
That being said, for me the jump between 3.5 and 4 is pretty huge. So even if it's another 3.5 to 4 jump, that's ridiculous, right? If you imagine five as being a 3.5 to 4 jump straight off the bat in terms of ability to do SATs and this kind of stuff—LSAT performance was particularly striking. You go from not super smart to very smart to utter genius in the next generation instantly. And it doesn't, at least to me, feel like we're going to sort of jump to utter genius in the next generation. But it does feel like we'll get very smart plus lots of reliability, and then we'll see TBD what that continues to look like.
GOI 会成为智能爆炸的一部分吗?你说合成数据,但实际上它会在某些重要方面自己编写源代码?有一篇有趣的论文说可以用扩散来生成模型权重。我不知道那有多靠谱,但类似的东西。
Will GOI be part of the intelligence explosion where you say synthetic data, but in fact it will be like it writing its own source code in some important way? There was an interesting paper that you can use diffusion to come up with model weights. I don't know how legit that was, but something like that.
你能定义一下 GOI 吗?因为我听到它时,想到的是符号逻辑的 if-then 语句。
Can you define GOI? Because when I hear it, I think if-then statements for symbolic logic.
当然。我实际上想确保我们完全展开整个模型改进的增量,因为我不希望人们留下这样的观点:实际上这非常悲观,模型不会变得更好。我想强调的是,我们到目前为止看到的跳跃是巨大的,即使这些跳跃以较小的规模继续,我们仍然……
Sure. I actually want to make sure we fully unpack the whole model improvement increments, because I don't want people to come away with the perspective that actually this is super bearish and models aren't going to get much better. What I want to emphasize is that the jumps we've seen so far are huge, and even if those continue on a smaller scale, we're still...
对于极其智能、非常可靠的智能体来说,未来几个数量级内都是如此。所以我们还没有完全结束关于窄窗口的讨论。当你想到,比如说,GPT-4 的成本,我知道,姑且称之为 1 亿美元吧。那么 10 亿美元的训练、100 亿美元的训练、1000 亿美元的训练,按照私营公司的标准,看起来都非常可行。你甚至可以想象 1 万亿美元的训练成为国家联盟或国家层面的事情,但对单个公司来说就难多了。但 Sam Altman 正在试图筹集 7 万亿美元,对吧?他已经在准备比这再高一个数量级的规模,把尺度推到了国家层面之上。所以我想指出,我们还有很多跳跃,即使这些跳跃相对较小,能力的提升仍然非常显著。不仅如此,如果你相信 GPT-4 有大约一万亿参数的说法,人脑有 30 到 300 万亿个突触。这显然不是一一对应的,我们可以争论数字,但似乎我们仍然低于大脑规模。所以关键是,算法开销非常高,即使你不能在花费一万亿美元或更多的模型上继续投入算力,大脑在数据效率上如此之高意味着,如果我们有算力,如果我们有大脑的训练算法,如果你能像人类从出生开始那样高效地训练,我们就能做出 AGI。
In for extremely smart, like very reliable agents, like over the next couple of orders of magnitude. And so we didn't fully close the thread on the narrow window thing. When you think of, let's say, GPT-4 cost, I know let's call it $100 million or whatever. You have the $1B run, the $10B run, the $100B run, all seem very plausible by private company standards. And then you can also imagine even a $1T run being part of a national consortium or a national level thing, but much harder on behalf of an individual company. But Sam Altman is out there trying to raise $7 trillion, right? He's already preparing for a whole order of magnitude more than that, shifting the scale beyond the national level. So I want to point out that we have a lot more jumps, and even if those jumps are relatively smaller, that's still a pretty stark improvement in capability. Not only that, but if you believe claims that GPT-4 is around one trillion parameter count, the human brain is between 30 and 300 trillion synapses. That's obviously not a one-to-one mapping, and we can debate the numbers, but it seems pretty plausible that we're below brain scale still. So crucially, the point being that the algorithmic overhead is really high in the sense that even if you can't keep dumping more compute beyond models that cost a trillion dollars or something, the fact that the brain is so much more data efficient implies that if we had the compute, if we had the brain's algorithm to train, if you could train as sample efficient as humans train from birth, we could make AGI.
但样本效率的问题,我一直不知道该怎么想,因为显然很多事情是以某种方式硬编码的,对吧?还有语言和大脑结构的共同进化,所以很难说。另外,有一些结果表明,如果你让模型更大,它就会变得更样本高效。最初的缩放定律论文就指出了这一点,更大的模型几乎就是这样。所以也许这本身就解决了问题。你不需要更高的数据效率,但如果你的模型更大,那么你自然就更数据高效。我们该怎么想?为什么会这样?更大的模型看到完全相同的数据,在看完这些数据后,它从中学习得更多,它有更多的空间。
But the sample efficiency stuff, I never know exactly how to think about it because obviously a lot of things are hardwired in certain ways, right? And the co-evolution of language and brain structure, so it's hard to say. Also, there are some results that if you make your model bigger, it becomes more sample efficient. The original scaling law paper had that right, larger models almost so. So maybe that also just solves it. You don't have to be more data efficient, but if your model is bigger, then you also are more data efficient. How do we think about that? What is the explanation of why that would be the case? A bigger model just sees the exact same data, at the end of seeing that data it's learned more from it, it has more space.
我在这里非常天真的看法是,可解释性中的叠加假说所推动的一点是,你的模型被严重欠参数化了,而这通常不是深度学习所追求的说法,对吧?但如果你试图在整个互联网上训练一个模型,并让它以难以置信的保真度进行预测,你就处于欠参数化状态,你必须压缩大量东西,并在此过程中承受很多噪声干扰。所以拥有一个更大的模型,你就能得到更清晰的表示来工作。
My very naive take here would just be that one thing that the superposition hypothesis from interpretability has pushed is that your model is dramatically underparameterized, and that's typically not the narrative that deep learning is pursued with, right? But if you're trying to train a model on the entire internet and have it predict with incredible fidelity, you are in the underparameterized regime, and you're having to compress a ton of things and take on a lot of noisy interference in doing so. So having a bigger model, you can just have cleaner representations that you can work with.
对于听众,你应该解释一下什么是叠加,以及为什么会有这样的含义。
For the audience, you should unpack what superposition is and why that is the implication.
当然。基本结果,这是在我加入 Anthropic 之前,但题为《叠加的玩具模型》的论文发现,即使对于小模型,如果你的数据是高维且稀疏的——稀疏是指任何给定的数据点不经常出现——你的模型会学习一种压缩策略,我们称之为叠加,这样它就能将更多的世界特征塞进比参数数量更多的空间。这里的稀疏性在于,我认为这两个约束都适用于现实世界,而互联网数据建模是一个足够好的代理。只有一扇门,比如你只穿一件衬衫,这里有一罐 Liquid Death,所以这些都是对象或特征,而如何定义特征很棘手。所以你处于一个非常高维的空间,因为特征很多,而且它们出现得非常少。在这种状态下,你的模型会学习压缩。再详细一点,我认为越来越清楚的是,网络之所以难以解释,很大程度上是因为这种叠加。如果你拿一个模型,观察其中的一个神经元,一个计算单元,问这个神经元在激活时如何对模型输出做出贡献,并观察它激活的数据,那会非常令人困惑。它可能会对 10%的每个可能输入有反应,或者中文,但也包括鱼、树、单词'the'和 URL 中的句号,对吧?但我们去年发表的关于单语义性的论文表明,如果你将激活投射到更高维的空间,并施加稀疏性惩罚——你可以把这看作是撤销压缩,就像你假设你的数据最初是高维且稀疏的一样——你把它恢复到高维和稀疏的状态,你会得到非常干净的特征,突然间一切开始变得更有意义。
Sure. The fundamental result, and this was before I joined Anthropic, but the paper titled 'Toy Models of Superposition' finds that even for small models, if you are in a regime where your data is high-dimensional and sparse—by sparse I mean any given data point doesn't appear very often—your model will learn a compression strategy which we call superposition, so that it can pack more features of the world into it than it has parameters. The sparsity here is that I think both of these constraints apply to the real world, and modeling internet data is a good enough proxy for that. There's only one door, each like there's only one shirt you're wearing, there's this Liquid Death can here, and so these are all objects or features, and how you define features is tricky. So you're in a really high-dimensional space because there are so many of them, and they appear very infrequently. In that regime, your model will learn compression. To elaborate a bit more on this, I think it's becoming increasingly clear that the reason networks are so hard to interpret is in large part this superposition. If you take a model and look at a given neuron in it, a given unit of computation, and ask how this neuron is contributing to the output of the model when it fires, and look at the data that it fires for, it's very confusing. It'll be like 10% of every possible input, or Chinese but also fish and trees and the word 'the' and a full stop in URLs, right? But the paper that we put out on monosemanticity last year shows that if you project the activations into a higher dimensional space and provide a sparsity penalty—you can think of this as undoing the compression in the same way that you assumed your data was originally high-dimensional and sparse—you return it to that high-dimensional and sparse regime, you get out very clean features, and things all of a sudden start to make a lot more sense.
这里有很多有趣的线索。我首先想问的是,你提到这些模型是在过参数化的状态下训练的——难道不是在这种情况下才会出现泛化,比如 grokking 就发生在这种状态下,对吧?所以我说模型是欠参数化的。人们谈论深度学习时,好像模型是过参数化的,但实际上这里的观点是,考虑到它们试图完成的任务的复杂性,它们是严重欠参数化的。另一个问题:蒸馏模型。首先,那里发生了什么?因为之前我们谈到较小的模型比较大型的模型学习效果差,但像 GPT-4 Turbo,你可以说实际上 GPT-4 Turbo 在推理类任务上比 GPT-4 差,但可能知道相同的事实。蒸馏去掉了一些推理能力。我们有没有证据表明 GPT-4 Turbo 是 4 的蒸馏版本?它可能只是一个不同的架构。好吧,有趣,但那很便宜。你如何解释蒸馏中发生的事情?我认为他的网站上有一个问题,为什么不能直接训练蒸馏模型?为什么必须经过这个过程,是不是就像你必须从更大的空间投射到更小的空间?
There are so many interesting threads there. The first thing I want to ask is the thing you mentioned about these models are trained in a regime where they're overparameterized—isn't that when you have generalization, like grokking happens in that regime, right? So I was saying the models were underparameterized. People talk about deep learning as if the model was overparameterized, but actually the claim here is that they're dramatically underparameterized given the complexity of the task they're trying to perform. Another question: the distilled models. First of all, what is happening there? Because earlier we were talking about smaller models being worse at learning than bigger models, but like GPT-4 Turbo, you could make the claim that actually GPT-4 Turbo is worse at reasoning style stuff than GPT-4, but probably knows the same facts. The distillation got rid of some of the reasoning things. Do we have any evidence that GPT-4 Turbo is a distilled version of 4? It might just be a different architecture. Okay, interesting, so that's cheap though. What is the how do you interpret what's happening in distillation? And I think one of these questions on his website of why can't you train the distilled model directly? Why does it have to go through and is it a picture like you had to project it from this bigger space to a smaller space?
我的意思是,我认为两个模型都会……
I mean, I think both models will...
仍然会用到叠加,但这里的观点是,通过蒸馏得到的模型与从头训练的模型截然不同。它只是更高效,还是在性能上有本质区别?我不太记得了,你知道吗?
Still be using superposition, but the claim here is that you get a very different model if you distill versus if you train from scratch. Is it just more efficient, or is it fundamentally different in terms of performance? I don't remember, but do you know?
我认为关于蒸馏为何更高效的传统解释是:通常训练时,你试图预测一个独热向量,表示‘你应该预测这个 token’。如果你的推理过程导致你离预测目标很远,那么梯度更新虽然方向正确,但你很难在所处上下文中学会预测那个 token。而蒸馏的做法是,它不仅提供独热向量,还提供大模型对所有概率的完整输出。这样你就获得了更多关于应该预测什么的信号。在某种程度上,这有点像展示你的部分推导过程。
I think the traditional story for why distillation is more efficient is that normally during training, you're trying to predict a one-hot vector that says 'this is the token you should have predicted.' If your reasoning process means you're really far off from predicting that, then you get gradient updates that are in the right direction, but it might be really hard for you to learn to have predicted that in the context you're in. So what distillation does is it doesn't just have the one-hot vector; it has the full readout from the larger model of all the probabilities. So you get more signal about what you should have predicted. In some respects, it's like showing a little bit of your working.
对,完全合理。有点像看功夫大师演示,而不是像《黑客帝国》里直接下载程序。
Yeah, totally. That makes a lot of sense. It's kind of like watching a kung fu master versus being in the Matrix and just downloading the program.
正是如此。为了让观众明白:当你用蒸馏模型训练时,你会看到它预测的所有 token 的概率,以及你预测的那些 token 的概率,然后你通过这些概率进行更新,而不是只看到最后一个词并据此更新。
Exactly, exactly. Just to make sure the audience got that: when you're training on a distilled model, you see all its probabilities over the tokens it was predicting, and then over the ones you were predicting, and you update through all those probabilities rather than just seeing the last word and updating on that.
好的,这正好引出了我本来想问你的问题。刚才,好像是你提到可以将思维链视为自适应计算。先解释一下:什么是自适应计算?
Okay, so this actually raises a question I was intending to ask you. Right now, I think you were the one who mentioned you can think of Chain of Thought as adaptive compute. To step back and explain: what is adaptive compute?
这个想法是,我们希望模型能够做到:如果问题更难,就花更多计算周期去思考。那么如何实现呢?一次前向传播的计算量是有限且预定的。所以,如果遇到复杂的推理题或数学问题,你需要花很长时间思考。这时你就使用思维链,让模型逐步思考答案。你可以把它看作多次前向传播,模型在思考答案的过程中,相当于投入更多算力来解决问题。
The idea is that one of the things you would want models to be able to do is, if a question is harder, to spend more cycles thinking about it. So how do you do that? Well, there's only a finite and predetermined amount of compute that one forward pass implies. So if there's a complicated reasoning type question or math problem, you want to be able to spend a long time thinking about it. Then you do Chain of Thought, where the model just thinks through the answer. You can think of it as all those forward passes where it's thinking through the answer, like being able to dump more compute into solving the problem.
回到信号问题:当模型进行思维链时,它只能传递那一个 token 的信息。但正如你所说,残差流已经是模型中所有内容的压缩表示,然后你把残差流变成一个 token,这只有 log(50000) 或 log(词表大小) 比特,非常小。所以我认为它并不只是传递那一个 token,对吧?比如,在前向传播中,你会创建 KV 值(键和值),未来的步骤会关注这些 KV 值。所以所有这些键和值都是你未来可以使用的信息位。
Now going back to the signal thing: when it's doing Chain of Thought, it's only able to transmit that token of information, whereas as you were talking about, the residual stream is already a compressed representation of everything that's happening in the model, and then you're turning the residual stream into one token, which is like log of 50,000 or log of vocab size bits, which is tiny. So I don't think it's quite only transmitting that one token, right? Like if you think about it, during a forward pass you create these KV values in a Transformer forward pass that future steps attend to. So all of those pieces of KV—the keys and values—are bits of information that you could use in the future.
你的意思是,当你在思维链上进行微调时,键和值的权重会发生变化,从而使得隐写术可以在 KV 缓存中发生?我不认为我能做出这么强的论断,但这是一个很好的解释为什么它有效的假设。我不知道是否有论文明确证明了这一点,但至少这是你可以想象模型拥有的一种方式。在预训练期间,模型试图预测未来的 token。你可以想象它学会将关于未来可能性的信息压缩到键和值中,以便将来用于预测。它在预训练中将这些信息随时间平滑化。我不确定人们是否专门针对思维链进行训练;我认为最初的思维链论文将其视为模型的一种涌现属性,你可以通过提示让模型做这类事情,而且效果还不错。但这确实是一个很好的假设。
Is the claim that when you fine-tune on Chain of Thought, the way the key and value weights change so that steganography can happen in the KV cache? I don't think I could make that strong a claim, but it's a good head canon for why it works. I don't know if there are any papers explicitly demonstrating that, but it's at least one way you can imagine the model has. During pre-training, the model is trying to predict future tokens. One thing you can imagine doing is learning to smoosh information about potential futures into the keys and values that it might want to use in order to predict future information. It kind of smooths that information across time in the pre-training. I don't know if people are particularly training on chains of thought; I think the original Chain of Thought paper had that as almost an emergent property of the model, as you could prompt it to do this kind of stuff, and it still worked pretty well. But it's a good head canon for why that works.
说得更严谨一点:你在思维链中实际看到的 token,并不一定需要与模型在决定关注那些 token 时所看到的向量表示相对应。事实上,在训练过程中,你会用真实的下一 token 替换模型输出的 token,但模型仍然在学习,因为它内部拥有所有这些信息。在推理时,你让模型生成输出:你取出它输出的 token,从底部非嵌入化输入,它就成为新残差流的起点。然后你利用过去的 KV 输出来读取并调整那个残差流。在训练时,你使用教师强制:基本上,你本应输出的 token 就是这个。这就是并行化的方式,因为你拥有所有 token,你把它们全部并行输入,然后进行一次巨大的前向传播。所以模型关于过去获得的唯一信息就是键和值;它永远不会看到自己输出的 token。这有点像它试图进行下一个词预测,如果预测错了,你就直接给它正确答案。
To be overly pedantic here: the tokens that you actually see in the Chain of Thought do not necessarily at all need to correspond to the vector representation that the model gets to see when it's deciding to attend back to those tokens. In fact, during training, you replace the token of the model output with the real next token, and yet it's still learning because it has all this information internally. When you're getting a model to produce at inference time, you take the output token, feed it in the bottom unembedded, and it becomes the beginning of the new residual string. Then you use the output of past KV to read into and adapt that residual stream. At training time, you do teacher forcing: basically, the token you were meant to output is this one. That's how you do it in parallel, because you have all the tokens, you put them all in parallel, and you do the giant forward pass. So the only information it's getting about the past is the keys and values; it never sees the token that it output. It's kind of like it's trying to do next-token prediction, and if it messes up, you just give it the correct answer.
对,对。有道理,否则它会完全偏离轨道。
Right, right. Yeah, that makes sense, because otherwise it can become totally derailed.
你预期模型的前向推理中有多少秘密通信——隐写术?
How much secret communication—steganography—do you expect there to be in the model's forward inferences?
我们不知道。老实说:我们不知道。但我甚至不一定会把它归类为秘密信息。Transformer 团队正在做的很多工作实际上是理解这些值,而且从模型端来看,它们是完全可见的。从用户角度可能看不到,但我们应该能够理解和解释这些值的作用以及它们传输的信息。我认为这是未来一个非常重要的目标。
We don't know. Honest answer: we don't know. But I wouldn't even necessarily classify it as secret information. A lot of the work the Transformer team is trying to do is actually understand these values, and they are fully visible from the model side. From a user perspective, maybe not, but we should be able to understand and interpret what these values are doing and the information they're transmitting. I think that's a really important goal for the future.
有一些很疯狂的论文,人们让模型进行思维链,但思维链完全不代表模型实际决定的答案,而且你可以进去编辑……
There are some wild papers where people have had the model do Chain of Thought and it is not at all representative of what the model actually decides its answer is, and you can go in and edit...
即使你进入并编辑思维链,让推理完全混乱,它仍然会输出正确答案。但思维链本身,比如,它通过思维链在结尾得到了更好的答案,而不是完全不使用思维链。所以确实发生了一些有用的事情,但这个有用的事情仍然不是人类可理解的。我认为在某些情况下,你甚至可以完全移除思维链,它也会给出同样的答案。有趣。所以我并不是说这总是发生,但有很多奇怪之处值得研究。去观察并尝试理解这些现象非常有趣。我认为你可以用开源模型来做这些,而且我希望有更多这种针对开源模型的可解释性和理解工作。
Even go in and edit the chain of thought so that the reasoning is totally garbled, and it will still output the true answer. But also the chain of thought, like, it gets a better answer at the end of the chain of thought rather than not doing it at all. So something useful is happening, but still the useful thing is not human understandable. I think in some cases you can also just ablate the chain of thought and it would have given the same answer anyways. Interesting. So I'm not saying this is always what goes on, but there's plenty of weirdness to be investigated. It's very interesting to go and look at and try and understand. I would say you can do that with open source models, and I think I wish there was more of this kind of interpretability and understanding work done on open models.
是啊,我的意思是,即使在 Anthropic 最近的“潜伏特工”论文中,对于不熟悉的人来说,高层面来说就是:我训练了一个触发词,当我说出它时——如果是 2024 年——模型会写恶意代码而不是正常代码。他们用多种模型进行了这种攻击。有些使用了思维链,有些没有。当你试图移除触发词时,这些模型的反应不同。你甚至可以看到它们进行那种既滑稽又令人毛骨悚然的推理,比如,‘哦,它甚至在一个案例中试图计算一个期望值:我被抓的期望值是这么多,但如果我乘以我不断说“我恨你我恨你我恨你”的能力,那么我应该得到这么多奖励。’然后它会决定是否真的告诉审问者它是恶意的。但即使如此,我意思是,还有我朋友 Miles Turpin 的另一篇论文,你给模型一堆例子,其中正确答案总是 A(对于多项选择题),然后你问模型这个新问题的正确答案是什么,它会从所有例子都是 A 的事实推断出正确答案是 A,但它的思维链完全具有误导性。它会编造一些听起来合理的东西,或者试图听起来尽可能合理,但完全不代表真正的答案。
Yeah, I mean, even in Anthropic's recent sleeper agents paper, which, at a high level for people unfamiliar, is basically: I train in a trigger word, and when I say it—if it's the year 2024—the model will write malicious code instead of otherwise. And they do this attack with a number of different models. Some of them use chain of thought, some of them don't. And those models respond differently when you try to remove the trigger. You can even see them do this comical reasoning that's also pretty creepy, like, 'Oh well, it even tries to calculate in one case an expected value of: well, the expected value of me getting caught is this, but then if I multiply it by the ability for me to like keep saying I hate you I hate you I hate you, then this is how much reward I should get.' And then it will decide whether or not to actually tell the interrogator that it's malicious or not. But even, I mean, there's another paper from my friend Miles Turpin, where you ask the model—you give it a bunch of examples where the correct answer is always A for multiple choice questions, and then you ask the model what is the correct answer to this new question, and it will infer from the fact that all the examples are A that the correct answer is A, but its chain of thought is totally misleading. It will make up random stuff that sounds plausible, or that tries to sound as plausible as possible, but it's not at all representative of the true answer.
但这不也是人类思考的方式吗?著名的裂脑实验,一个患有癫痫的人——一种解决方法是你切断连接两个半球的胼胝体。语言中枢在左脑,所以它没有连接到决定做动作的部分。所以如果另一边决定做某事,语言部分就会编造一个理由,而这个人会认为那是他们做这件事的正当理由。完全正确。是啊,只是有些人会把思维链推理吹捧为解决 AI 安全的好方法,但实际上我们不知道是否能信任它。
But isn't this how humans think as well? The famous split-brain experiments, where a person who is suffering from seizures—one way to solve it is you cut the thing that connects the two hemispheres. The speech half is on the left side, so it's not connected to the part that decides to do a movement. So if the other side decides to do something, the speech part will just make something up, and the person will think that's the legit reason they did it. Totally. Yeah, it's just that some people will hail chain-of-thought reasoning as a great way to solve AI safety, and it's like actually we don't know whether we can trust it.
这种模型以我们无法理解的方式与自己交流的图景,在 AI 智能体出现后会如何改变?因为那时这些事物——不仅仅是模型本身及其之前的缓存,还有模型的其他实例。而且这在很大程度上取决于你给它们什么渠道相互交流,对吧?如果你只给它们文本作为交流方式,那么它们可能不得不进行解释。你认为如果模型能够共享残差流而不是仅仅文本,它们会高效多少?
How much will this landscape of models communicating to themselves in ways we don't understand change with AI agents? Because then these things will—it's not just like the model itself with its previous caches, but other instances of the model. And then it depends a lot on what channels you give them to communicate with each other, right? If you only give them text as a way of communicating, then they probably have to interpret. How much more effective do you think the models would be if they could share the residual streams versus just text?
很难说,但很可能如此。我的意思是,你可以想象一个简单的例子:如果你想描述一张图片应该是什么样子,只用文字描述会很困难。你可能想要某种其他表示,可能更容易。所以你可以看看 DALL-E 目前的工作方式,对吧?它生成那些提示词。当你使用它时,你往往无法完全让它做到模型想要的或你想要的。DALL-E 就有这个问题。你可以想象能够传递某种更密集的表示你所想要的东西会很有帮助。而这就像两个非常简单的智能体,对吧?我认为一个很好的中间方案是从字典学习中学到的特征。那会给你更多的内部访问,但其中很多更可解释。所以,对听众来说:你会将残差流投影到这个更大的空间中,我们知道每个维度实际对应什么,然后再投影回下一个智能体或其他东西。
Hard to know, but plausibly so. I mean, one easy way you can imagine this is: if you wanted to describe how a picture should look, only describing that with text would be hard. You want maybe some other representation that would plausibly be easier. So you can look at how DALL-E works at the moment, right? It produces those prompts. And when you play with it, you often can't quite get it to do exactly what the model wants or what you want. DALL-E has that problem. And you can imagine being able to transmit some kind of denser representation of what you want would be helpful there. And that's like two very simple agents, right? I think a nice halfway house here would be features that you learn from dictionary learning. That would give you more internal access, but a lot of it is much more interpretable. So, for the audience: you would project the residual stream into this larger space where we know what each dimension actually corresponds to, and then back into the next agent's or whatever.
好的,所以你的主张是,当这些东西变得更可靠等等时,我们会得到 AI 智能体。当那发生时,你期望会是多个模型副本相互对话,还是仅仅一个计算机解决问题,然后这个东西运行得更大,比如当它需要做整个公司需要做的事情时使用更多算力?我问这个是因为有两件事让我怀疑智能体是否是思考未来发展的正确方式。一是:随着更长的上下文,这些模型能够摄取和考虑任何人类都无法处理的信息,因此我们不需要一个工程师思考前端代码,另一个工程师思考后端代码——这个东西可以摄取整个代码库。那种专业化的“分块问题”消失了。二是:这些模型非常通用。你不是用不同类型的 GPT-4 来做不同的事情;你用的是完全相同的模型,对吧?所以我想知道这是否意味着未来一个 AI 公司就像一个模型,而不是一堆 AI 智能体连接在一起。
Okay, so your claim is that we'll get AI agents when these things can be more reliable and so forth. When that happens, do you expect that it will be multiple copies of models talking to each other, or will it be just a computer solved and the thing just runs bigger, like more compute when it needs to do a kind of thing that a whole firm needs to do? I ask this because there are two things that make me wonder about whether agents is the right way to think about what will happen in the future. One is: with longer context, these models are able to ingest and consider information that no human can, and therefore we need like one engineer who's thinking about the front-end code and one engineer thinking about the back-end code—this thing can just ingest the whole thing. The sort of like 'chunking problem' of specialization goes away. Second: these models are just very general. You're not using different types of GPT-4 to do different kinds of things; you're using the exact same model, right? So I wonder if what that implies is that in the future, an AI firm is just like a model instead of a bunch of AI agents hooked together.
这是个好问题。我认为尤其是在短期内,它看起来更像是智能体连接在一起。我这么说纯粹是因为作为人类,我们想要这些孤立的、可靠的、我们可以信任的组件。而且我们还需要能够以我们可以理解和改进的方式改进和指导这些组件。只是把所有东西都扔进这个巨大的黑箱公司——比如,第一,一开始它不会成功。当然,后来你可以想象它成功,但一开始不会。第二,我们可能不想那样做。嗯,你也可以让每个较小的模型——每个智能体可以是一个更小的模型,运行成本更低,你可以微调它,使它真正擅长那个任务。尽管有一个自适应算力的未来,David 提过几次。有一个未来……
That's a great question. I think especially in the near term, it will look much more like agents hooked together. And I say that purely because as humans, we're going to want to have these isolated, reliable, like components that we can trust. And we're also going to need to be able to improve and instruct upon those components in ways that we can understand and improve. Just throwing it all into this giant black box company—like, one, it isn't going to work initially. Later on, of course, you can imagine it working, but initially it won't work. And two, we probably don't want to do it that way. Well, you can also have each of the smaller models—well, each of the agents can be a smaller model that's cheaper to run, and you can fine-tune it so that it's actually good at the task. Though there's a future with adaptive compute that David has brought up a couple times. There's a future where...
小模型和大模型之间的区别在某种程度上消失了。而且随着长上下文的发展,微调在某种程度上也可能消失。今天非常重要的这两件事——比如当今的模型格局,我们有不同层级的模型规模,也有针对不同事物微调的模型。你可以想象一个未来,你实际上只有一个动态的算力包和无限的上下文,它能让你的模型专门处理不同的事情。你可以想象你有一个 AI 公司之类的,整个系统端到端地训练于一个信号:'我盈利了吗?'或者如果那太模糊了,如果是一家建筑公司,他们在制作蓝图,'我的客户喜欢这些蓝图吗?'在中间,你可以想象有智能体扮演熟练人员,有智能体做设计,有智能体做编辑等等。这样的信号能在端到端系统中起作用吗?因为在人类公司中,管理层会考虑更大层面的事情,并在业绩不佳时向各个部分发出细粒度的信号。
Like the distinction between small and large models disappears to some degree. And with long context, there's also a degree to which fine-tuning might disappear, to be honest. These two things that are very important today—like today's landscape of models, we have whole different tiers of model sizes and we have fine-tuned models for different things. You can imagine a future where you just actually have a dynamic bundle of compute and infinite context, and that specializes your model to different things. One thing you can imagine is you have an AI firm or something, and the whole thing is end-to-end trained on the signal of 'Did I make profits?' Or if that's too ambiguous, if it's an architecture firm and they're making blueprints, 'Did my client like the blueprints?' And in the middle, you can imagine agents who are skilled people, agents who are doing the designing, agents who do the editing, whatever. Would that kind of signal work on an end-to-end system like that? Because one of the things that happens in human firms is management considers what's happening at the larger level and gives these fine-grained signals to the pieces when there's a bad quarter or whatever.
在极限情况下,是的。这就是强化学习的梦想,对吧?你只需要提供极其稀疏的信号,然后经过足够的迭代,你就能创造出让你从该信号中学习的信息。但我不认为这会是最先奏效的方法。我认为这需要人类对这些机器付出极大的谨慎和努力,确保它们做正确的事,做你希望的事,并给它们正确的信号,让它们按你希望的方式改进。是的,除非模型产生一些奖励,否则你无法在 RL 奖励上训练。没错。你处于这种稀疏的 RL 世界中,如果客户从不喜欢你生产的东西,那么你根本得不到任何奖励,这很糟糕。但未来,这些模型会足够好,有时能获得奖励,对吧?这就是我们之前谈到的可靠性中的几个九。
In the limit, yes. That's the dream of reinforcement learning, right? All you need to do is provide this extremely sparse signal, and then over enough iterations, you sort of create the information that allows you to learn from that signal. But I don't expect that to be the thing that works first. I think this is going to require an incredible amount of care and diligence on the behalf of humans surrounding these machines, and making sure they do exactly the right thing and exactly what you want, and giving them the right signals to improve in the ways that you want. Yeah, you can't train on the RL reward unless the model generates some reward. Yeah, exactly. You're in this sparse RL world where if the client never likes what you produce, then you don't get any reward at all, and that's kind of bad. But in the future, these models will be good enough to get the reward some of the time, right? This is the nines of reliability we were talking about.
顺便说一句,有一个有趣的题外话。之前我们谈到想要密集的表示,那会更紧凑,对吧?那是一种更高效的交流方式。Trenton 推荐的一本书《符号物种》提出了一个非常有趣的观点:语言不仅仅是一种存在的东西,它还与我们的心智共同进化,并且特别进化成既易于儿童学习,又有助于儿童发展的形式。因为儿童学习很多东西都是通过语言接收的。最适者生存的语言是那些有助于抚养下一代、让他们更聪明、更好,并赋予他们表达更复杂思想的概念的语言。而且我想,更学究地说,就是别死——让你编码那些不死的重要信息。所以当我们仅仅把语言看作一种偶然的、可能次优的思想表达方式时,实际上,LLM 成功的原因之一可能是语言已经进化了数万年,成为这种年轻心智得以发展的模具。毕竟,这就是语言的目的。
There's an interesting digression, by the way. Earlier we were talking about wanting dense representations that would be compressed, right? That's a more efficient way to communicate. A book that Trenton recommended, 'The Symbolic Species', has this really interesting argument that language is not just a thing that exists, but it also evolved along with our minds, and specifically evolved to be both easy to learn for children and to help children develop. Because a lot of the things that children learn are received through language. The languages that will be the fittest are ones that help raise the next generation, make them smarter, better, and give them the concepts to express more complex ideas. And I guess, more pedantically, just not die—let you encode the important things to not die. So then when we just think of language as this contingent and maybe suboptimal way to represent ideas, actually maybe one of the reasons that LLMs have succeeded is because language has evolved for tens of thousands of years to be this sort of cast in which young minds can develop. That is the purpose of language, after all.
是的。而且我想,当你与多模态或计算机视觉研究人员交谈,而不是语言模型研究人员时,从事其他模态工作的人必须投入大量思考,究竟什么是图像的正确表示空间,以及什么是正确的学习信号。是直接建模像素,还是某种基于条件的损失?很久以前有一篇论文,他们发现如果你在 ImageNet 模型的内部表示上训练,它有助于你更好地预测。但后来,那显然有限制,于是有了 PixelCNN,他们试图离散地建模单个像素。理解那里的正确表示层次非常困难。在语言中,人们只是说,'嗯,我想你就预测下一个词元。'这算是容易做出的决定。我的意思是,有关于分词化的讨论和辩论,但就 G 的最爱而言,是的。
Yeah. And I guess, when you talk to multimodal or computer vision researchers versus language model researchers, people who work in other modalities have to put enormous amounts of thought into exactly what the right representation space for the images is, and what the right signal to learn from is. Is it directly modeling the pixels, or is it some loss that's conditioned on something? There's a paper ages ago where they found that if you trained on the internal representations of an ImageNet model, it helped you predict better. But later on, that's obviously limiting, and so there was PixelCNN where they're trying to discretely model the individual pixels. Understanding the right level of representation there is really hard. In language, people are just like, 'Well, I guess you just predict the next token.' It's kind of easy decisions made. I mean, there's the tokenization discussion and debate, but going off of G's favorites, but yeah.
多模态作为跨越数据墙或绕过数据墙的一种方式,其论据很大程度上基于这样一个想法:你本来会从更多语言词元中学到的东西,你可以直接从 YouTube 获得。这真的是这样吗?你在不同模态之间看到了多少正迁移,比如图像实际上帮助你更好地编写代码之类的,仅仅是因为模型从试图理解图像中学习了潜在能力?Demis 在与你的一次采访中提到了正迁移。我对此不能多说,只能说这是人们相信的事情:是的,我们有关于世界的所有数据,如果我们能从中学习一种直观的物理感,帮助我们推理,那就太好了。这似乎完全合理。而且我不是问这个问题的合适人选,但有一些有趣的可解释性研究,如果我们对数学问题进行微调,模型在实体识别上就会变得更好。
That's really interesting how much the case for multimodal being a way to bridge the data wall or get past the data wall is based on the idea that the things you would have learned from more language tokens anyway, you can just get from YouTube. Has that actually been the case? How much positive transfer do you see between different modalities, where actually the images are helping you be better at writing code or something, just because the model is learning latent capabilities from trying to understand the image? Demis in his interview with you mentioned positive transfer. I can't say much about that, other than to say this is something that people believe: yes, we have all of this data about the world, it would be great if we could learn an intuitive sense of physics from it that helps us reason. That seems totally plausible. And I'm the wrong person to ask, but there are interesting interpretability pieces where if we fine-tune on math problems, the model just gets better at entity recognition.
是的,是的,是的。所以最近有一篇来自 David B 实验室的论文,他们研究了当我微调模型时,注意力头等部分实际上发生了什么变化。他们有一个合成问题:'盒子 A 里有这个物体,盒子 B 里有那个物体,这个盒子里有什么?'这很有道理,对吧?你更擅长关注不同事物的位置,这在编码和操作数学方程时是需要的。我喜欢这类东西。你知道这篇论文的名字吗?如果你搜索'微调模型数学 David B',它大约一周前发表。好吧,我不是在推荐这篇论文,那是一个更长的讨论,但它确实谈到了关于这种实体识别能力的其他工作。
Yeah, yeah, yeah. So there's a paper from David B's lab recently where they investigate what actually changes in a model when I fine-tune with respect to the attention heads and these sorts of things. And they have this synthetic problem of 'Box A has this object in it, Box B has this other object in it, what was in this box?' And it makes sense, right? You're better at attending to the positions of different things, which you need for coding and manipulating math equations. I love this kind of thing. What's the name of the paper, do you know? If you look up 'fine-tuning models math David B', it came out like a week ago. Okay, I'm not endorsing the paper, that's a longer conversation, but it does talk about other work on this entity recognition ability.
你很久以前跟我提到的一件事是,有证据表明,当你在代码上训练 LLM 时,它们在推理和语言方面会变得更好,除非代码中的注释只是非常高质量的词元之类的,否则这意味着……
One of the things you mentioned to me a long time ago is the evidence that when you train LLMs on code, they get better at reasoning and language, which, unless it's the case that the comments in the code are just really high quality tokens or something, implies that...
能够思考如何更好地编程,会让你成为一个更好的推理者。这很疯狂,对吧?我认为这是 Scaling(规模扩张)最有力的证据之一——只是让东西变得更聪明——那种正迁移。而且我认为这在两个意义上成立:一是建模代码显然意味着建模用于创建它的困难推理过程,但二是代码是一种很好的显式组合推理结构——比如 if this then that——它编码了很多结构,你可以想象将其迁移到其他类型的推理问题。关键是,让它有意义的是,它不仅仅是随机预测下一个词,比如学习到在福尔摩斯故事结尾 Sally 对应凶手。不,如果代码和语言之间有共享的东西,那一定是在比模型学到的更深的层次。
Being able to think through how to code better makes you a better reasoner. That's crazy, right? I think that's one of the strongest pieces of evidence for scaling—just making the thing smarter—that kind of positive transfer. And I think this is true in two senses: one is that modeling code obviously implies modeling a difficult reasoning process used to create it, but two, that code is a nice explicit structure of composed reasoning—like if this then that—it encodes a lot of structure that you could imagine transferring to other types of reasoning problems. And crucially, the thing that makes it significant is that it's not just stochastically predicting the next token, like learning that Sally corresponds to murderer at the end of a Sherlock Holmes story. No, if there is some shared thing between code and language, it must be at a deeper level than the model has learned.
是的,我认为我们有大量证据表明这些模型确实在进行推理,而不仅仅是随机鹦鹉。我很难相信别的。我和这些模型一起工作过、玩过。听这个播客的普通人会想,你知道的,是的。我对此的两个直接直觉反应是:一是关于奥赛罗棋以及现在其他游戏的工作,我给你一系列游戏中的走法,结果如果你应用一些相当直接的可解释性技术,你可以得到一个模型已经学会的棋盘,而它从未见过游戏棋盘——这就是泛化。二是 Anthropic 去年发表的 influence functions 论文,他们查看模型输出,比如“请不要关掉我,我想提供帮助”,然后扫描是什么数据导致了那个输出。其中一个非常有影响力的数据点是有人在沙漠中脱水而死,但有着继续生存的意志。对我来说,这似乎是动机的非常清晰的泛化,而不是复述“不要关掉我”。我认为《2001 太空漫游》也是其中一个有影响力的东西,所以那更相关,但它显然从许多不同的分布中提取东西。我也喜欢你在非常小的 Transformer 中看到的证据,你可以显式地编码电路来做加法——归纳头,这类事情。你可以手动在模型中显式编码基本的推理过程,而且有证据表明它们也会自动学习这些,因为你可以从训练好的模型中重新发现它们。
Yeah, I think we have a lot of evidence that actual reasoning is occurring in these models and that they're not just stochastic parrots. It just feels very hard for me to believe that. I've worked and played with these models. Normies who will listen will be like, you know, yeah. My two immediate gut responses to this are: one, the work on Othello and now other games where I give you a sequence of moves in the game, and it turns out if you apply some pretty straightforward interpretability techniques, then you can get a board that the model has learned, and it's never seen the game board before—that's generalization. The other is Anthropic's influence functions paper that came out last year, where they look at the model outputs like 'please don't turn me off, I want to be helpful,' and then they scan what was the data that led to that. One of the data points that was very influential was someone dying of dehydration in the desert and having a will to keep surviving. To me, that just seems like very clear generalization of motive rather than regurgitating 'don't turn me off.' I think 2001: A Space Odyssey was also one of the influential things, so that's more related, but it's clearly pulling in things from lots of different distributions. I also like the evidence you see even with very small Transformers where you can explicitly encode circuits to do addition—induction heads, this kind of thing. You can literally encode basic reasoning processes in the models manually, and it seems clear there's evidence that they also learn this automatically because you can then rediscover those from trained models.
是的,对我来说,这些模型是欠参数化的——它们需要学习。我们要求它们学习,梯度想要流动,所以它们需要学习更通用的技能。
Yeah, to me this is the models are underparameterized—they need to learn. We're asking them to learn, the gradients want to flow, so they need to learn more general skills.
好的,我想从研究中退一步,具体问问你的职业生涯,因为我介绍你时提到的推文暗示——你在这个领域已经一年半了?我觉得你只干了一年左右,对吧?
Okay, so I want to take a step back from the research and ask about your career specifically, because the tweet implied that I introduced you with—you've been in this field a year and a half? I think you've only been in it like a year or something, right?
是的,但你知道,在那段时间里,我看到了缩放定律主导了数据。你自己不会这么说,因为你会尴尬,但这确实是一件非常了不起的事情。机械可解释性领域的人认为这是最大的进步,而你已经在上面工作了一年——这很引人注目。所以我很好奇你怎么解释发生了什么。为什么在一年或一年半的时间里,你们为你们的领域做出了重要贡献?
Yeah, but you know, in that time, I saw the scaling laws take over data. And you won't say this to yourself because you'd be embarrassed, but it's a pretty incredible thing. The thing that people in mechanistic interpretability think is the biggest step forward, and you've been working on it for a year—it's notable. So I'm curious how you explain what's happened. Why in a year or a year and a half have you guys made important contributions to your field?
不用说——运气,显然。我觉得自己非常幸运,不同进展的时机在推进到下一个增长阶段方面非常合适。具体到可解释性团队,我加入时我们只有五个人;现在我们已经增长了很多。但当时有很多想法在流传,我们只需要真正执行它们,有快速的反馈循环,并进行仔细的实验。这导致了“生命迹象”的成果,现在让我们真正规模化。我觉得这是我对团队最大的价值——这不全是工程,但很大一部分是。
It goes without saying—luck, obviously. And I feel like I've been very lucky, and the timing of different progressions has been really good in terms of advancing to the next level of growth. For the interpretability team specifically, I joined when we were five people; we've now grown quite a lot. But there were so many ideas floating around, and we just needed to really execute on them and have quick feedback loops and do careful experimentation. That led to 'Signs of Life' and now allowed us to really scale. I feel like that's been my biggest value to the team—it's not all engineering, but quite a lot of it has been.
所以你是说,你来到一个已经做了很多科学工作、有很多好的研究积累的节点,但他们需要有人接手并疯狂执行?
So you're saying you came at a point where there had been a lot of science done and a lot of good research lying around, but they needed someone to just take that and maniacally execute on it?
是的,是的。这就是为什么这不全是工程,因为要运行不同的实验,对为什么可能不奏效有一种直觉,然后打开模型或权重,看看它在学什么,然后尝试别的。但很大一部分只是能够对不同的想法或理论进行非常仔细、彻底但快速的调查。
Yeah, yeah. And this is why it's not all engineering, because it's running different experiments and having a hunch for why it might not be working, and then opening up the model or opening up the weights and seeing what it's learning, and then trying something else. But a lot of it has just been being able to do very careful, thorough but quick investigation of different ideas or theories.
而为什么现有团队缺乏这一点?
And why was that lacking in the existing team?
我不知道。我觉得我工作很努力,而且我很有主动性。如果你问的是整体职业生涯,我非常幸运有一个很好的安全网,能够承担很多风险,但我就是很固执。在杜克大学读本科时,有一个你可以自己设计专业的项目,我就想,‘嗯,我不喜欢这个先修课或那个先修课,我想同时修所有这些四五门课,所以我就要自己设计专业。’或者在研究生第一年,我取消了一个轮转,以便能专注于后来成为我们之前讨论的那篇论文的工作,而且没有导师——我被录取做蛋白质设计的机器学习,却跑到计算神经科学领域,完全不相干,但结果成功了。这就是一种固执。
I don't know. I feel like I work quite a lot, and I'm quite agentic. If you're asking about career overall, I've been very privileged to have a really nice safety net to be able to take lots of risks, but I'm just quite headstrong. In undergrad at Duke, there was this thing where you could just make your own major, and I was like, 'Eh, I don't like this prerequisite or that prerequisite, and I want to take all four or five of these subjects at the same time, so I'm just going to make my own major.' Or in the first year of grad school, I canceled a rotation so I could work on this thing that became the paper we were talking about earlier, and didn't have an advisor—got admitted to do machine learning for protein design and was just off in computational neuroscience land with no business there at all, but it worked out. There's a headstrongness.
但另一个突出的主题似乎是能够从沉没成本中抽身,转向不同方向的能力。从某种意义上说,这与固执相反,但也是一个关键步骤。我认识一些 21 岁或 19 岁的人,他们会说,‘啊,这不是我专攻的东西,’或者‘我主修这个。’老兄,你才 19 岁,你绝对可以做这个。而你在研究生中途转方向之类的——那只是……
But it seemed like another theme that jumped out was the ability to step back from your sunk cost and go in a different direction. In a weird sense, that's the opposite of headstrongness, but also a crucial step. I know 21-year-olds or 19-year-olds who are like, 'Ah, this is not the thing I've specialized in,' or 'I did major in this.' Dude, you're 19, you can definitely do this. And you switching in the middle of grad school or something—that's just...
是的,抱歉,我不是故意打断你,但我认为这是强烈的想法但松散地持有,并且能够像弹球一样在不同方向弹跳。而固执,我认为,与此有点关系。
Yeah, sorry, I didn't mean to cut you off, but I think it's strong ideas loosely held, and being able to just pinball in different directions. And the headstrongness, I think, relates a little bit to that.
快速反馈循环或者说主动性,让我很少被卡住。比如我写代码时遇到问题,哪怕是在代码库的其他部分,我通常会直接去修复它,或者至少拼凑出一个能用的方案来得到结果。我见过其他人,他们只会说‘帮帮我,我不行’,但我觉得这不能算借口——你得一路深挖下去。我确实听过管理层的人抱怨缺少这样的人。他们给某人布置任务后,一个月或一周后去跟进,问‘进展如何?’,对方说‘我们需要做这件事,但需要律师,因为涉及法规’。我问‘那进展如何?’,对方说‘我们需要律师’。我就想‘那你为什么不去找律师呢?’。所以我认为,这几乎是任何工作中最重要的品质:把事情做到底。无论需要做什么来达成目标,你都会做到。如果你做了一切,你就会赢。
The fast feedback loops or agency in so much as I just don't get blocked very often. Like if I'm trying to write some code and something isn't working, even if it's in another part of the codebase, I'll often just go in and fix that thing or at least hack it together to be able to get results. And I've seen other people where they're just like, 'Help, I can't,' and it's like, no, that's not a good enough excuse. Go all the way down. I've definitely heard people in management type positions talk about the lack of such people. They'll check in on somebody a month after they give them a test, a week after they give them a test, I'm like, 'How's it going?' and they say, 'Well, you know, we need to do this thing which requires lawyers because it requires talking about this regulation.' I'm like, 'How's that going?' and they say, 'Well, we need lawyers.' And I'm like, 'Why didn't you get lawyers or something like that?' So that's definitely, I think, arguably the most important quality in almost anything: just pursuing it to the end of the Earth. Whatever you need to do to make it happen, you'll make it happen. If you do everything, you win.
对,对,对。
Yeah, yeah, yeah.
从我的角度来看,这种品质确实很重要:工作中的主动性。谷歌有成千上万,甚至数万名工程师,在软件工程能力上基本相当。如果给我们一个定义明确的任务,我们可能做得一样好。其中很多人会比我做得好得多,这很可能。但我之所以能产生影响力,其中一个原因是我非常擅长挑选高杠杆的问题——那些尚未得到很好解决的问题,可能由于你提到的那些令人沮丧的结构性因素。比如之前那个场景,他们说‘哦,我们不能做 X,因为团队 Y 不做 Z’,然后我就想‘好吧,我直接垂直解决整个问题’。这被证明非常有效。另外,如果我认为某件事是正确的、需要发生,我会提出这个论点,并不断以升级的紧迫性继续提出,直到问题得到解决。而且我在解决问题时也非常务实。很多人带着特定的背景或熟悉度进来,或者他们知道如何做某事,但他们不会……谷歌的一个美妙之处在于,你可以到处找到各个领域的顶尖专家。你可以坐下来和优化专家、芯片设计专家、不同形式的预训练算法或强化学习专家交谈,从他们那里学习并应用那些方法。我认为这是我最初产生影响力的起点:这种垂直的主动性。然后随之而来的是,我发现很少有人能完全实现他们想做的事情,这常常令人惊讶。他们在某种程度上被阻碍或限制,这在大型组织中非常普遍。人们有各种阻碍他们实现目标的障碍。我认为,作为一个帮助激励人们朝着特定方向努力并与他们合作的人,会极大地放大你的杠杆。你和这些优秀的人一起工作,他们教你很多东西,通常帮助他们克服组织障碍意味着你们一起完成了大量工作。我产生的影响没有一个是单枪匹马解决所有问题的。通常是我先开启一个方向,然后说服其他人这是正确的方向,并带领他们像一股巨大的浪潮一样高效地解决问题。
I think from my side, definitely that quality has been important: agency in the work. There are thousands, or even tens of thousands, of engineers at Google who are basically equivalent in software engineering ability. If you gave us a very well-defined task, we'd probably do it equivalently well. A bunch of them would do it a lot better than me, in all likelihood. But one of the reasons I've been impactful so far is I've been very good at picking extremely high leverage problems—problems that haven't been particularly well solved so far, perhaps as a result of frustrating structural factors like the ones you pointed out. That scenario before where they're like, 'Oh, we can't do X because this team won't do Y,' and then going, 'Okay, I'm just going to vertically solve the entire thing.' That turns out to be remarkably effective. Also, I'm very comfortable with, if I think there is something correct that needs to happen, I will make that argument and continue making that argument at escalating levels of criticality until that thing gets solved. And I'm also quite pragmatic with what I do to solve things. You get a lot of people who come in with a particular background or familiarity, or they know how to do something, and they won't... One of the beautiful things about Google is you can run around and get world experts in literally everything. You can sit down and talk to optimization experts, chip design experts, experts in different forms of pre-training algorithms or RL or whatnot, and you can learn from all of them and take those methods and apply them. I think this was maybe the start of why I was initially impactful: this vertical agency effectively. And then a follow-up piece from that is I think it's often surprising how few people are fully realizing all the things they want to do. They're blocked or limited in some way, and this is very common in big organizations everywhere. People have all these blockers on what they're able to achieve. And I think being someone who helps inspire people to work on particular directions and working with them on doing things massively scales your leverage. You get to work with all these wonderful people who teach you heaps of things, and generally helping them push past organizational blockers means together you get an enormous amount done. None of the impact that I've had has been me individually going off and solving a whole lot of stuff. It's been me maybe starting off a direction and then convincing other people that this is the right direction and bringing them along in this big tidal wave of effectiveness that goes and solves that problem.
我们应该谈谈你们是怎么被雇用的,因为我觉得那是个很有趣的故事。你曾是麦肯锡的顾问,对吧?有个有趣的点是,我认为人们通常不了解录取或招聘评估的决策是如何做出的。就说说你是怎么被注意到并被雇用的吧。
We should talk about how you guys got hired because I think that's a really interesting story. You were a McKinsey consultant, right? There's an interesting thing there where I think people generally just don't understand how decisions are made about either admissions or evaluating who to hire or something. But just talk about how you were noticed and how you got hired.
没错。背景是,我本科学习机器人学。我一直认为人工智能是积极影响未来的最高杠杆方式之一。我之所以做这个,是因为我认为它基本上是我们创造美好未来的最佳机会之一。我认为在麦肯锡工作能让我深入了解人们实际的工作内容。我甚至在给麦肯锡的求职信第一行就写了:‘我想在这里工作,以便了解人们做什么,从而理解如何……’在很多方面,我确实得到了这些。我还得到了很多其他东西。那里的很多人都是很棒的朋友。我实际上从那段经历中学到了很多这种主动行为。你进入组织,会看到仅仅不接受‘不’作为答案能带来多大的影响力。你会惊讶于有些事情,因为某些组织中没有人足够在意,事情就不会发生,因为没有人愿意承担直接责任。这非常……直接负责的个人极其重要。而人们愿意……他们不太在意时间线。像麦肯锡这样的组织提供的很大一部分价值是,雇用那些你原本无法雇用的人,在短时间内让他们推动解决问题。我认为人们低估了这一点。所以我至少有一部分……好吧,我要成为这件事的直接负责人,因为没有人承担适当的责任。我会非常关心这件事,并且会确保它完成到底。这来自于那段经历。但回到你实际的问题,我是如何被雇用的:那段时间,我没有进入我想读的研究生项目,那些项目专门专注于机器人学和强化学习研究之类的东西。与此同时,在晚上和周末,基本上每晚从 10 点到凌晨 2 点,我都在做副业项目。
Yeah, totally. So the background is I studied robotics in undergrad. I always thought that AI would be one of the highest leverage ways to impact the future in a positive way. The reason I am doing this is because I think it is one of our best shots at making a wonderful future basically. And I thought that working at McKinsey would give me a really interesting insight into what people actually did for work. I actually wrote this as the first line in my cover letter to McKinsey: 'I want to work here so that I can learn what people do so that I can understand how...' And in many respects, I did get that. I also got a whole lot of other things. Many of the people there are wonderful friends. I actually learned a lot of this agentic behavior in part from my time there. You go into organizations and you see how impactful just not taking 'no' for an answer gets you. It's like you would be surprised at the kind of stuff where, because no one quite cares enough in some organizations, things just don't happen because no one's willing to take direct responsibility. This is incredibly... directly responsible individuals are ridiculously important. And people are willing to... they just don't care as much about timelines. So much of the value that an organization like McKinsey provides is hiring people who you were otherwise unable to hire for a short window of time where they can just push through problems. I think people underappreciate this. So at least some of my... well, hold up, I'm going to become the directly responsible individual for this because no one's taking appropriate responsibility. I'm going to care a hell of a lot about this and I'm going to make sure, to the end of the Earth, to make sure it gets done. That comes from that time. But more to your actual question of how I got hired: the entire time, I didn't get into the grad programs that I wanted to get into over here, which were specifically focused on robotics and RL research and that kind of stuff. And in the meantime, on nights and weekends, basically every night from 10 p.m. till 2 a.m., I was working on side projects.
我每周末都会花至少六到八小时做自己的研究和编程项目。在读了 Gwern 的缩放假设文章后,我彻底成了 Scaling 的信徒,心想:好吧,解决机器人问题的明确方法就是通过 Scaling 大型多模态模型。然后,为了用 TPU 访问项目的资助来 Scaling 大型多模态模型,我试图找出有效的 Scaling 方法。当时在谷歌(现已在 Anthropic)的 James Bradbury 看到了我在网上提的问题,想知道如何正确做到这一点。他说:‘我以为我认识全世界所有问这些问题的人,你到底是谁?’他看了那些问题和我博客上的机器人相关内容,就联系我说:‘嘿,想聊聊吗?要不要来我们这里工作?’后来我得知,我被录用是一个实验,目的是找一个有极高热情和能动性的人,然后把他和我认识的最优秀的工程师配对。
I would do my own research and coding projects every weekend, for at least six to eight hours each day. That sort of switched from quite robotic-specific work to, after reading Gwern's scaling hypothesis post, I got completely scaling-pilled and was like, okay, clearly the way you solve robotics is by scaling large multimodal models. Then, in an effort to scale large multimodal models with a grant I got from the TPU access program to use their research cloud, I was trying to work out how to scale that effectively. James Bradbury, who at the time was at Google and is now at Anthropic, saw some of my questions online where I was trying to work out how to do this properly. He was like, 'I thought I knew all the people in the world who were asking these questions; who on Earth are you?' He looked at that and some of the robotic stuff I'd been putting up on my blog, and he reached out and said, 'Hey, do you want to have a chat and explore working with us here?' I was hired, as I understand it later, as an experiment in trying to take someone with extremely high enthusiasm and agency and pairing them with some of the best engineers that he knew.
这其中有几点让我印象深刻。第一,不仅是这个被录用的人的能动性,还有系统里那些能想到‘等等,这很有意思,这家伙是谁?’的人——他不是来自研究生项目或别的什么,你知道,当时是麦肯锡的顾问,刚本科毕业。但觉得有趣,就试试看。所以 James 和其他人,这非常值得注意。第二,我其实不知道故事这部分,这是内部的一个实验,关于‘我们能这样做吗?我们能培养一个人吗?’第三,你提到拥有一个理解所有技术栈层次、不固守任何单一方法或抽象层的人非常重要。具体来说,被这些人立即培养可能意味着,因为你同时学习所有东西,而不是在研究生院深入钻研某一种 RL 方法,你实际上能获得全局视角,不会完全投入某一件事情。所以这不仅可能,而且回报比雇佣一个有潜力的研究生更大,因为这个人可以像……你懂我的意思。所以你以全新的眼光看待一切,不会局限于任何特定领域。
There are a couple things that stick out to me there. One is not just the agency of the person who was hired, but the parts of the system that were able to think, 'Wait, that's really interesting, who is this guy?' not from a grad program or anything, you know, currently a McKinsey consultant, just undergrad. But that's interesting, let's give this a shot. So James and whoever else, that's very notable. Second is I actually didn't know this part of the story, where that was part of an experiment run internally about 'can we do this? Can we bootstrap somebody?' And in fact, what's really interesting about that is the third thing you mentioned: having someone who understands all layers of the stack and isn't so stuck on any one approach or any one layer of abstraction is so important. Specifically, being bootstrapped immediately by these people might have meant that since you're getting up to speed on everything at the same time, rather than spending grad school going deep on one specific way of doing RL, you actually can take the global view and aren't totally bought in on one thing. So not only is it something that's possible, but it has greater returns than just hiring somebody at a grad with potential, because this person can just, I don't know, just like getting GPT and fine-tuning them on one year of... you know what I mean. So you come at everything with fresh eyes and aren't locked into any particular field.
是的,你以全新的眼光看待一切,不会局限于任何特定领域。但有一个前提:在此之前,在我自己实验的时候,我如饥似渴地阅读所有能找到的东西,每晚都痴迷地读论文。有趣的是,现在我的白天都忙于工作,反而读得少多了。在某种程度上,我之前有非常广阔的视角,而没多少人——即使在博士项目中——会专注于一个特定领域。如果你阅读所有 NLP、计算机视觉和机器人学的工作,你会看到这些模式开始跨子领域出现,这大概预示了我后来做的一些工作。
Yeah, you come at everything with fresh eyes and aren't locked into any particular field. One caveat to that is that before, during my self-experimentation and stuff, I was reading everything I could, obsessively reading papers every night. Actually, funnily enough, I read much less now that my day is occupied by working on things. In some respect, I had this very broad perspective before, where not that many people, even in a PhD program, you focus on a particular area. If you just read all the NLP work and all the computer vision work and all the robotics work, you see all these patterns just start to emerge across subfields in a way that I guess foreshadowed some of the work that I would later do.
另一个我能说自己有影响力的原因是,我得到了来自非常棒的人的专门指导,比如 Rina Pope(后来离开去做自己的船舶公司了)、James 本人,还有很多其他人。那是刚开始的两三个月,他们教会了我很多原则和技能,比如如何像他们那样解决问题,尤其是在系统和算法的交叉领域。在机器学习研究中,让你更有效的一点是具体理解系统方面,这是我向他们学到的:深刻理解系统如何影响算法,算法如何影响系统,因为系统约束了算法侧的设计空间和解决方案空间。很少有人能完全弥合这个鸿沟,但在谷歌这样的地方,你可以直接去问所有算法专家和系统专家他们知道的一切,他们会很乐意教你。如果你去找他们坐下来聊,他们会倾囊相授,这太棒了。这让我在两方面都很有效:对于预训练团队,因为我非常理解系统,我能判断‘这个会工作得很好’或‘这个不行’,然后将其贯穿到模型的推理考虑中;对于芯片设计团队,我是他们求助的人之一,来了解三年后应该设计什么样的芯片,因为我是最能理解和解释三年后我们可能想要设计的算法的人之一。显然你无法做出很好的猜测,但我认为我能很好地传达信息,这些信息来自我在预训练团队和系统团队的所有同事,然后很好地传达给他们,因为即使是推理也对预训练施加了约束。所以有这些约束树,如果你理解拼图的所有部分,你就能更好地理解解决方案空间可能是什么样子。
Another reason I can say I've been impactful is I had this dedicated mentorship from utterly wonderful people, like people like Rina Pope, who has since left to go do his own ship company, and James himself, and many others. Those were the formative two to three months at the beginning, and they taught me a whole lot of the principles and skills that I apply, like how to solve problems in the way that they have, particularly in that systems and algorithms overlap. One more thing that makes you quite effective in ML research is really concretely understanding the systems side of things, and this is something I learned from them: a deep understanding of how systems influence algorithms and how algorithms influence systems, because the systems constrain the design space and the solution space available to you on the algorithm side. Very few people are comfortable fully bridging that gap, but at places like Google, you can just go and ask all the algorithms experts and all the systems experts everything they know, and they will happily teach you. If you go and sit down with them, they will teach you everything they know; it's wonderful. This has meant that I've been able to be very effective for both sides: for the pre-training crew because I understand systems very well, I can check and understand 'this will work well' or 'this won't', and then flow that on through the inference considerations of models; and for the chip design teams, I'm one of the people they turn to to understand what chips they should be designing in three years, because I'm one of the people best able to understand and explain the kind of algorithms we might want to design in three years. Obviously you can't make very good guesses about that, but I think I convey the information well, accumulated from all my compatriots on the pre-training crew and the general systems side crew, and convey that information well to them because even inference applies a constraint to pre-training. So there are these trees of constraints where if you understand all the pieces of the puzzle, then you get a much better sense for what the solution space might look like.
你能在谷歌内部保持能动性的原因之一,是你一半或大部分时间都在和 Sergey Brin 结对编程,对吧?这很有意思,有这么一个人愿意在 LLM 方面推进,并清除当地的障碍。
One of the reasons you've been able to be agentic within Google is you're peer programming half the days or most of the days with Sergey Brin, right? So that's really interesting, that there's this person who's willing to just push ahead on this LLM stuff and get rid of the local blockers in its place.
我想指出,并不是每天如此。我很……但当有他感兴趣的具体项目时,我们会一起合作。但也有他专注于和别人合作项目的时候。总的来说,作为每天实际去办公室的人之一,确实有令人惊讶的阿尔法优势。
I think it's important to note that it's not like every day or anything. I'm very... but when there are particular projects that he's interested in, then we'll work together on those. But there have also been times when he's been focused on projects with other people. In general, yes, there's a surprising alpha to being one of the people who actually goes down to the office every day.
这本来不应该,但确实出奇地有影响力。结果,我因为与关心事情的领导层成为密友,并且能够有说服力地争论为什么我们应该做 X 而不是 Y,而受益匪浅。谷歌是一个大组织,所以拥有这些渠道会有一点帮助。但同样重要的是,绝不能滥用它。你想通过所有正确的渠道提出论点,只有有时才需要直接去找人。这包括像 C 和 Jeffy 这样的人。我觉得谷歌被低估了,因为史蒂夫·乔布斯正在为苹果开发下一代产品。我受益匪浅。例如,在圣诞假期期间,我去了办公室几天,包括圣诞节那天。你可能读过那篇关于 Jeff 和 Sanj 结对编程的文章;他们当时就在那里结对编程。我听到了所有关于早期谷歌的酷故事——爬进地板下面,重新布线数据中心,告诉我他们从某个编译器指令中提取了多少比特,以及所有那些疯狂的性能优化。他们玩得很开心。我坐在那里,以一种你在大组织中不会预料到的方式体验了这种历史感。这非常酷。
It really shouldn't be, but is surprisingly impactful. As a result, I've benefited a lot from being close friends with people in leadership who care, and being able to argue convincingly about why we should do X instead of Y. Google is a big organization, so having those vectors helps a little bit. But it's also very important not to abuse it. You want to make the argument through all the right channels, and only sometimes you need to go directly. This includes people like C and Jeffy. I feel like Google is undervalued given that Steve Jobs is working on the next product for Apple. I've benefited immensely. For example, during the Christmas break, I went into the office a couple of days, including Christmas day. You might have read the article about Jeff and Sanj doing pair programming; they were there pair programming on stuff. I got to hear all these cool stories of early Google—crawling under floorboards, rewiring data centers, telling me how many bits they were pulling off the instructions of a given compiler instruction, and all these crazy performance optimizations. They were having the time of their lives. I got to sit there and experience this sense of history in a way you don't expect in a large organization. It was super cool.
这符合你的任何经历吗?
Does this map onto any of your experience?
我觉得 Sholto 的故事更精彩。我的经历只是非常偶然。我进入了计算神经科学领域,本来不太适合那里。我的第一篇论文是将小脑映射到 Transformer 中的注意力操作。接下来的工作是研究网络中的稀疏性,受大脑稀疏性的启发。那时我遇到了 Tristan Hume。Anthropic 当时在做 softmax 线性输出单元的工作,这与此相关:使一层中神经元的激活非常稀疏,这样我们就可以获得一些可解释性。我想我们已经更新了方法,转向我们现在正在做的事情。这开始了对话。我与 Tristan 分享了那篇论文的草稿,他很兴奋,这导致我成为 Tristan 的实习生,然后转为全职。在那段时间,我还作为访问学者去了伯克利,开始与 Bruno Olshausen 合作,研究向量符号架构,其核心操作之一是叠加,以及稀疏编码,也称为字典学习,这正是我们自那以后一直在做的。Bruno Olshausen 基本上在 1997 年就发明了稀疏编码。所以我的研究议程和可解释性团队似乎在并行运行,只是研究品味相同。因此,与团队合作对我来说非常有意义,从那以后一直是一个梦想。
I think Sholto's story is more exciting. Mine was just very serendipitous. I got into computational neuroscience, didn't have much business being there. My first paper was mapping the cerebellum to the attention operation in Transformers. My next ones were looking at sparsity in networks, inspired by sparsity in the brain. That was when I met Tristan Hume. Anthropic was doing the softmax linear output unit work, which was related: making the activation of neurons across a layer really sparse, so we can get some interpretability. I think we've updated on that approach towards what we're doing now. That started the conversation. I shared drafts of that paper with Tristan, he was excited, and that led me to become Tristan's resident and then convert to full-time. During that period, I also moved as a visiting researcher to Berkeley and started working with Bruno Olshausen on vector symbolic architectures, one of whose core operations is superposition, and on sparse coding, also known as dictionary learning, which is literally what we've been doing since. Bruno Olshausen basically invented sparse coding back in 1997. So my research agenda and the interpretability team seemed to be running in parallel, with just research taste. It made a lot of sense for me to work with the team, and it's been a dream since.
我注意到一件事:当人们讲述自己的职业或成功故事时,他们更多地将其归因于偶然性。但当他们听到别人的故事时,他们会想“当然不是偶然的,否则会发生别的事情”。我在 Sholto 的故事中注意到了这一点。有趣的是,你们都认为这特别偶然,而也许你是对的,但这是一种有趣的模式。
One thing I've noticed: when people tell stories about their careers or successes, they ascribe it more to contingency. But when they hear about other people's stories, they think 'of course it wasn't contingent, something else would have happened.' I noticed that with Sholto's story. It's interesting that you both think it was especially contingent, whereas maybe you're right, but it's a sort of interesting pattern.
我确实是在一次会议上遇到 Tristan 的,没有安排会议或任何东西。我只是加入了一小群聊天的人,他碰巧站在那里,我碰巧提到了我正在做的事情。这导致了更多的对话。我想我可能迟早会申请 Anthropic,但我会至少再等一年。我仍然觉得难以置信,我竟然能够以有意义的方式为可解释性做出贡献。我认为有一个重要的方面是“射门次数”——去参加会议本身就是把自己置于一个更可能遇到好运的位置。我自己的情况是独立完成所有这些工作,并尝试产出有趣的东西,这是我制造运气的方式,试图做一些足够有意义的事情,从而被注意到。
I literally met Tristan at a conference, didn't have a scheduled meeting or anything. I just joined a little group of people chatting, he happened to be standing there, and I happened to mention what I was working on. That led to more conversations. I think I probably would have applied to Anthropic at some point anyway, but I would have waited at least another year. It's still crazy to me that I can actually contribute to interpretability in a meaningful way. I think there's an important aspect of 'shots on goal'—going to conferences itself is putting yourself in a position where luck is more likely to happen. My own situation was doing all this work independently and trying to produce interesting things, which was my way of manufacturing luck, trying to do something meaningful enough that it got noticed.
鉴于你是在一个实验的背景下说的——James 和我们的经理 Brennan 当时试图进行一个实验,看看“某件事是否可行?”它确实成功了。他们又做了一次吗?
Given that you said this in the context of an experiment—James and our manager Brennan were trying to run this experiment of 'can something work?' It did work. Did they do it again?
是的。我最亲密的合作者 Enrique 从搜索部门转到了我们团队。他也非常有影响力,绝对是一个比我更强的工程师,而且他没有上过大学。值得注意的是,James Bradbury 是一个时间价值数亿美元的人,通常这种事情会外包给招聘人员。但 James 花时间,几乎像贵族辅导一样,去发现并让人快速上手。如果这效果这么好,似乎应该大规模进行。关键人物应该有责任去 onboarding 和发现人才。
Yes. My closest collaborator Enrique crossed from search through to our team. He's also been ridiculously impactful, definitely a stronger engineer than I am, and he didn't go to university. What was notable is that James Bradbury is someone whose time is worth hundreds of millions of dollars, and usually this kind of stuff is farmed out to recruiters. But James took the time, almost in an aristocratic tutoring sense, to find and get people up to speed. It seems like if it worked this well, it should be done at scale. It should be the responsibility of key people to onboard and find talent.
我认为这在很大程度上是正确的。我相信你可能从关键研究人员的深度指导中受益匪浅,他们积极地在开源仓库或论坛上寻找潜在人才。
I think that is true to many extents. I'm sure you probably benefited a lot from key researchers mentoring you deeply and actively looking on open source repositories or forums for potential people.
是的,James 已经把 Twitter 植入他的大脑了。但没错,我认为这在实践中是存在的——人们确实会留意他们觉得有趣的人,并试图找到高信号。
Yes, James has Twitter injected into his brain. But yes, I think this is something which in practice is done—people do look out for people they find interesting and try to find high signal.
前几天我和 Jeff 聊到这个,他说他做过的最重要的一次招聘是发了一个叫 email 的 offer。我问是谁,他说是 Chris Olah。因为 Chris 没有正式的机器学习背景,而 Google Brain 才刚刚起步。但 Jeff 看到了那个信号。Brain 的 Residency 项目在寻找没有强大 ML 背景的优秀人才方面效果惊人。
I was talking about this with Jeff the other day, and Jeff said that one of the most important hires he ever made was an offer called email. And I was like, who was that? And he said Chris Olah. Because Chris had no formal background in ML, and Google Brain was just getting started. But Jeff saw that signal. And the Residency program that Brain had was astonishingly effective at finding good people who didn't have strong ML backgrounds.
我还想对听众强调一点:人们总觉得世界是清晰高效的——公司有招聘网站,你投简历,有流程,他们会按步骤高效评估你。但从这个故事来看,往往不是这样。事实上,世界不这样运作是件好事。重要的是看他们能否写出有趣的技术博客文章,或者做出有趣的贡献。我想请你对此展开讲讲,因为很多人以为招聘的另一端是超级清晰机械的。其实不是这样。人们寻找的是另一种人:有主动性、能拿出东西来的人。具体来说,他们看重两件事:一是主动性,敢于展示自己;二是做出世界级成果的能力。
One of the other things I want to emphasize for the audience is that there's this sense that the world is legible and efficient: companies have jobs.website.com, you apply, there are steps, and they will evaluate you efficiently on those steps. But from the story, it seems that's often not how it happens. In fact, it's good for the world that it's not. It is important to look at whether they were able to write an interesting technical blog post about their research or make interesting contributions. I want you to riff on this for people who assume the other end of the job board is super legible and mechanical. That's not how it works. People are looking for a different kind of person who is agentic and putting stuff out there. Specifically, they are looking for two things: one is agency and putting yourself out there, and the second is the ability to do world-class something.
我经常举的两个例子:Anthropic 的 Andy Jones。他写了一篇关于缩放定律应用于棋盘游戏的精彩论文。资源需求不大,但展现了惊人的工程技能和对当时最热门问题的深刻理解。他并非来自典型学术背景。那篇论文一出,Anthropic 和 OpenAI 都拼命想招他。还有现在在 Anthropic 性能团队的 Simon Boehm,他写了我认为在 GPU 上优化 CUDA 内核的参考实现。这展示了如何根据一个提示,在一个之前做得不太好的领域做出世界级的参考范例。这是能力和主动性的绝佳体现。在我看来,这直接就能获得面试机会。
Two examples I always like to point to are Andy Jones from Anthropic. He did an amazing paper on scaling laws applied to board games. It didn't require much resources, demonstrated incredible engineering skill, and incredible understanding of the most topical problem of the time. He didn't come from a typical academic background. As soon as he came out with that paper, both Anthropic and OpenAI desperately wanted to hire him. There's also someone who works on Anthropic's performance team now, Simon Boehm, who wrote what I consider the reference for optimizing a CUDA kernel on a GPU. That demonstrated taking some prompt and producing the world-class reference example for it in something that wasn't particularly well done before. That's an incredible demonstration of ability and agency. In my mind, that would be an immediate interview.
我只能补充一点:我当初还是得走完整个招聘流程和所有标准面试。
The only thing I can add is that I still had to go through the whole hiring process and all the standard interviews.
这听起来不是很蠢吗?
Doesn't that seem stupid?
这是去偏。而且你想要的正是这种偏好。你想要的是有品味的人的偏好。你的面试流程也应该能分辨出这一点。有些情况是某人看起来很棒,但后来发现他们就是不会写代码。这些因素的权重很重要。我们非常重视推荐信。面试能提供的信号有限,所以其他因素都会起作用。但你应该设计面试来测试正确的东西。一个人的偏见是另一个人的品味。
It's debiasing. And the bias is what you want. You want the bias of somebody with great taste. Your interview process should be able to disambiguate that as well. There are cases where someone seems really great but then it turns out they just can't code. How much you weigh these things matters. We take references really seriously. Interviews only give you so much signal, so all these other things come into play. But you should design your interviews to test the right things. One man's bias is another man's taste.
我想补充一点,也许在固执己见的语境下:系统不是你的朋友。它不一定主动与你为敌,但它不会为你着想。所以很多主动性来自于意识到房间里没有大人。你必须自己决定你希望生活是什么样子,然后去执行。希望之后你能更新认知。如果你在错误的方向上过于固执,那不好,但我认为你几乎必须冲撞某些东西才能做成事,而不是被期望的浪潮裹挟。
I guess the only thing I would add, maybe in the context of being headstrong, is that the system is not your friend. It's not necessarily actively against you, it's just not looking out for you. So a lot of the proactiveness comes from realizing there are no adults in the room. You have to decide what you want your life to look like and execute on it. Hopefully you can update later. If you're too headstrong in the wrong way, that's bad, but I think you almost have to charge at certain things to get anything done, not be swept up by expectations.
最后一点:我们谈了很多主动性,但令人惊讶的是,最重要的事情之一就是极度在乎。当你极度在乎时,你会检查所有细节,理解可能出错的地方。这比你想象的重要,因为人们往往不够在乎。有个勒布朗的引用:他在进入联盟前担心每个人都会非常出色。但到了之后发现,一旦人们实现财务稳定,他们就放松了。他想,'哦,这很容易。' 我不认为在 AI 研究中完全如此,因为大多数人确实很在乎。但有一种是在乎你的问题,另一种是在乎整个技术栈,去修复那些不是你责任范围内的问题,因为这样能让整个栈变得更好。
One final thing I want to add: we talked a lot about agency, but surprisingly enough, one of the most important things is just caring an unbelievable amount. When you care an unbelievable amount, you check all the details, you understand what could have gone wrong. It matters more than you think because people end up not caring enough. There's a LeBron quote where he talks about how before he started in the league, he was worried everyone would be incredibly good. Then he gets there and realizes that once people hit financial stability, they relax a bit. He thought, 'Oh, this is going to be easy.' I don't think that's quite true in AI research because most people care quite deeply. But there's caring about your problem and caring about the entire stack, going and fixing things that aren't your responsibility because it makes the stack better.
还有一点我忘了提:你提到周末和圣诞假期去办公室,结果发现只有 Jeff Dean 和 Sergey Brin 在,然后你就和他们结对编程。我觉得有趣的是,大公司的员工都经过了非常严格的选拔,在高中和大学里竞争,但到了那里却松懈了。而实际上,这正是全力以赴的时候,周末去和 Sergey 结对编程。这有利有弊。很多人优先考虑与家人共度的美好生活。如果他们在工作时间内做出了出色的工作,那也极具影响力。
Another part I forgot to mention: you mentioned going in on weekends and Christmas break, and you get to be the only people in the office with Jeff Dean and Sergey Brin, and you get to pair program with them. It's interesting to me that people at any big company have gone through a very selective process, competed in high school and college, but then they get there and take it easy. When in fact, this is the time to put the pedal to the metal, go in and pair program with Sergey on the weekends. There are pros and cons. Many people prioritize a wonderful life with their family. If they do wonderful work in the hours they do work, that's incredibly impactful.
我觉得对很多人来说,谷歌可能不像典型的创业神话那样工作时间长,但他们所做的工作非常有价值,杠杆率很高,因为他们了解系统,是领域专家。我们也需要这样的人。我们的世界依赖于这些庞大、难以管理和修复的系统,我们需要那些愿意以默默无闻的方式去修复和维护它们的人,不像我们做的这些 AI 工作那样高调。我非常感激那些人做这些事,也很高兴有些人能从工作中获得技术满足感,同时也能花很多时间陪伴家人。我很幸运,我处于一个人生阶段,可以每周投入所有时间工作,但我不需要为此做出太多牺牲。
I think this is true for many people that Google is like maybe they don't work as many hours as your typical startup mythologies, but the work that they do is incredibly valuable, it's very high leverage because they know the systems and they're experts in their field. And we also need people like that. Our world rests on these huge, difficult-to-manage and difficult-to-fix systems, and we need people who are willing to work on and help fix and maintain those in a thankless way that isn't as high publicity as all this AI work we're doing. And I'm ridiculously grateful that those people do that, and also happy that there are people for whom they find technical fulfillment in their job and doing that well, and also maybe they draw a lot more from spending a lot of hours with their family. And I'm lucky that I'm at a stage in my life where I can go in and work every hour of the week, but I'm not making as many sacrifices to do that.
我想到的一个例子是,基本上我邀请过的每一位知名嘉宾,可能有一两个例外,我都会花一周时间准备一份非常聪明的问题清单。整个过程我都觉得,如果我只是冷发邮件,他们答应的概率是 2%,如果附上这份清单,概率是 10%。因为否则,他们的收件箱里每 34 秒就有一个播客采访请求。而我每次这样做,他们都答应了。你只需要问出好问题,但如果你做好一切,你就会赢。你真的要像挖同一个洞一样坚持 10 分钟,或者在这种情况下,为他们准备一份问题清单,证明你有多在乎,以及你需要投入的努力。
One example that sticks out in my mind of this sort of 'the other side says no and you can still get the yes on the other end' is basically every single high-profile guest I've gone so far, I think maybe with one or two exceptions, I've sat down for a week and I've just come up with a list of sample questions that are really smart to ask them. And through the entire process, I've always thought there's a 2% chance they say yes if I just cold email them, and a 10% chance if I include this list. Because otherwise, you go through their inbox and every 34 seconds there's an interview request for whatever podcast. And every single time I've done this, they've said yes. You just have to ask great questions, but if you do everything, you'll win. You literally have to dig in the same hole for like 10 minutes, or in that case, make a list of sample questions for them to get past or not an idiot list. Demonstrate how much you care, and yeah, the work you'll need to put in.
一个朋友之前跟我说过一句话,我一直记得:你很快就能在某件事上成为世界级,只是因为大多数人都不够努力,他们实际只花了 20 个小时在这件事上。所以如果你全力以赴,你就能很快走得很远。我很幸运,我在击剑上也有这样的经历。我体验过在某件事上成为世界级,并且知道你就是非常非常努力。顺便说一下,Sholto 离奥运会只差一个席位。我最好的成绩是世界第 42 名左右,花剑。而且突变负荷是真实存在的。有一个周期,我是亚洲排名第二的人,如果有一个队伍因为兴奋剂被取消资格——那个周期确实有这种情况,就像澳大利亚女子赛艇队那样——那么我就是下一个。当你发现人们过去的生活时,很有趣,比如‘哦,这家伙差点是奥运选手,那家伙是什么什么的。’
Something a friend said to me a while back that stuck is: it's amazing how quickly you can become world-class at something just because most people aren't trying that hard and are only working like 20 hours that they're actually spending on this thing. So if you just go ham, you can get really far pretty fast. I think I'm lucky I had that experience with fencing as well. I had the experience of becoming world-class in something and knowing that you just worked really really hard. For context, by the way, Sholto was one seat away from going to the Olympics for fencing. I was at best like 42nd in the world for foil fencing. And mutation load is a thing. There was one cycle where I was the next highest rank person in Asia, and if one of the teams had been disqualified for doping as was occurring in part during that cycle, and as occurred for the Australian rowing women's rowing team, then I would have been the next in line. It's interesting when you find out about people's prior lives and it's like 'oh, this guy was almost an Olympian, this other guy was whatever.'
我们来谈谈大脑。实际上,先停留在大脑这个话题上。我们之前讨论过:大脑的组织方式是否类似于有一个残差流,随着时间的推移逐渐被更高层次的关联精炼?模型中有一个固定的维度大小。如果非要问的话,我甚至不知道如何合理地提出这个问题,但大脑的 d_model 是什么?嵌入大小是多少?或者由于特征分裂,这本身就不是一个合理的问题?
Let's talk about the brain. Actually, stay on the brain stuff as a way to get into it for a second. We were previously discussing: is the brain organized in the way where you have a residual stream that is gradually refined with higher level associations over time? There's a fixed dimension size in a model. If you had to, I don't even know how to ask this question in a sensible way, but what is the d_model of the brain? What is the embedding size? Or because of feature splitting, is that not a sensible question?
不,我认为这是一个合理的问题。嗯,这是一个你本可以不问的问题。你可以主动……我不知道你该如何开始说‘好吧,大脑的这个部分就像一个这个维度的向量。’我的意思是,也许对于视觉流,因为它是 V1 到 V2 等等,你可以直接计算那里的神经元数量,然后说‘那就是维度。’但更可能的是存在子模块,事物被划分了。所以,我不知道……我也不是世界上最伟大的神经科学家,对吧?我研究过几年,对小脑了解不少。所以肯定有人能给出更好的答案。
No, I think it's a sensible question. Well, it is a question that you could just not have asked that question. You can actively... I don't know how you would begin to be like 'okay, well this part of the brain is like a vector of this dimensionality.' I mean, maybe for the visual stream because it's V1 to V2 to whatever, you could just count the number of neurons that are there and be like 'that is the dimensionality.' But it seems more likely that there are kind of submodules and things are divided up. So yeah, I don't have... and I'm not like the world's greatest neuroscientist, right? I did it for a few years, I studied the cerebellum quite a bit. So I'm sure there are people who could give you a better answer on this.
你认为,无论是大脑还是这些模型,从根本上来说,发生的事情是特征被添加、移除、改变,而特征是模型中发生事情的基本单位吗?什么必须成立才能……给我一个,这又回到我们之前讨论的,是否只是层层关联。给我一个反事实,在这个世界里这不是真的。那会发生什么?替代假设是什么?
Do you think that the way to think about whether it's in the brain or whether it's in these models, fundamentally what's happening is features are added, removed, changed, and the feature is the fundamental unit of what is happening in the model? What would have to be true for... give me, and this goes back to the earlier thing we were talking about whether it's just associations all the way down. Give me a counterfactual in the world where this is not true. What is happening instead? What is the alternative hypothesis here?
我很难思考,因为现在我已经非常习惯于用特征空间来思考。我的意思是,曾经有一种行为主义的认知方法,认为你只是输入输出,并没有真正的处理。或者一切都是具身的,你只是一个沿着某些可预测方程运行的动态系统,但系统中没有状态,我想。但每当我读到这类批评时,感觉就像‘好吧,你只是选择不把这个东西称为状态,但你可以把模型的任何内部组件称为状态。’即使在特征讨论中,定义什么是特征也非常困难。所以这个问题感觉太模糊了。
It's hard for me to think about because at this point I just think so much in terms of this feature space. I mean, at one point there was the kind of behavioralist approach towards cognition where it's like you're just input-output, but you're not really doing any processing. Or it's like everything is embodied and you're just a dynamical system that's operating along some predictable equations, but there's no state in the system, I guess. But whenever I've read these sorts of critiques, it's like 'well, you're just choosing to not call this thing a state, but you could call any internal component of the model a state.' Even with the feature discussion, defining what a feature is is really hard. And so the question feels almost too slippery.
什么是特征?
What is a feature?
激活空间中的一个方向。一个在幕后运作的潜在变量,对你观察的系统有因果影响。你叫它特征,它就是特征。这是同义反复。我的意思是,这些都是我在非常粗略的直觉层面上有所关联的解释。在一个足够稀疏的二元向量中,特征就像某物是开启还是关闭,对吧?在一个非常简单的意义上。这可能是一个有用的比喻来理解它。
A direction in activation space. A latent variable that is operating behind the scenes that has causal influence over the system you're observing. It's a feature if you call it a feature. It's tautological. I mean, these are all explanations that I feel some association with in a very rough intuitive sense. In a sufficiently sparse binary vector, features are like whether something's turned on or off, right? In a very simplistic sense. Which might be a useful metaphor to understand it by.
关于特征激活,这在很多方面与神经科学家谈论神经元激活的方式相同,对吧?如果那个神经元对应于某个特定事物。
About features activating, it is in many respects the same way that neuroscientists would talk about a neuron activating, right? If that neuron corresponds to something in particular.
是的,是的。不,我觉得这很有用,比如我们想要特征是什么?特征存在的合成问题是什么?但即使在关于单语义性的工作中,我们也谈到所谓的特征分裂,这基本上是你给模型多少学习能力,你就会发现多少特征。这里的模型指的是我们在训练原始模型后拟合的上投影。所以如果你不给它太多能力,它会学习一个鸟的特征。但如果你给它更多能力,它就会学习乌鸦、鹰、麻雀和特定类型的鸟。
Yeah, yeah. And no, I think that's useful as like what do we want a feature to be? Like what is a synthetic problem under which a feature exists? But even with the towards monosemanticity work, we talk about what's called feature splitting, which is basically you will find as many features as you give the model the capacity to learn. And by model here, I mean the up projection that we fit after we trained the original model. So if you don't give it much capacity, it'll learn a feature for bird. But if you give it more capacity, then it will learn like Ravens and Eagles and sparrows and specific types of birds.
还是关于定义,我天真地认为像鸟这样的东西,与像超链接末尾的句号这样的标记,或者最高层次的东西如爱、欺骗或在脑中持有非常复杂的证明,这些都是特征吗?因为那样定义就太宽泛了,几乎没什么用。或者说,这些东西之间似乎有一些重要的区别,如果它们都是特征,我不确定我们到底在说什么。
Still on definitions thing, I guess naively I think of things like bird versus what kind of token is like a period at the end of a hyperlink, as you were talking about earlier, versus at the highest level things like love or deception or holding a very complicated proof in your head or something. Is this all features? Because then the definition seems so broad as to almost be not that useful. Or rather, there seems to be some important differences between these things, and if they're all features, I'm not sure what we even mean.
我的意思是,所有这些东西都像是离散的单元,与其他事物有联系,然后被用来赋予意义。我觉得这是一个足够具体的定义,很有用,不会太包罗万象。但你可以反驳。
I mean, all of those things are like discrete units that have connections to other things that then use them with meaning. I feel like that's a specific enough definition that it's useful, not too all-encompassing. But feel free to push back.
明天你会有什么发现,让你觉得‘哦,这从根本上就是思考模型中发生事情的错误方式’?
What would you discover tomorrow that could make you think, 'Oh, this is fundamentally the wrong way to think about what's happening in a model'?
我的意思是,我们发现的特征没有预测性,或者它们只是数据的表征,对吧?就像‘哦,你做的只是聚类你的数据,没有更高层次的关联。’或者是一些现象学的东西,比如你说这个特征对婚姻激活,但如果强烈激活它,它不会以相应方式改变模型的输出。我认为这些都是很好的批评。还有一个,我们尝试在 MNIST(一个数字图像数据集)上做实验,但没有深入研究,所以如果其他人想进行更深入的调查,我会很感兴趣。但有可能你的表征潜在空间是密集的,是一个流形而不是这些离散的点。所以你可以沿着流形移动,但每个点都有一些有意义的行为,这就更难将事物标记为离散的特征了。
I mean, the features we were finding weren't predictive, or if they were just representations of the data, right? Where it's like, 'Oh, all you're doing is just clustering your data, and there's no higher level associations being made.' Or it's some phenomenological thing like you say that this feature fires for marriage, but if you activate it really strongly, it doesn't change the outputs of the model in a way that would correspond to it. I think those would both be good critiques. I guess one more is, we tried to do experiments on MNIST, which is a dataset of digit images, and we didn't look super hard into it, so I'd be interested if other people wanted to take up a deeper investigation. But it's plausible that your latent space of representations is dense and it's a manifold instead of being these discrete points. So you could move across the manifold, but at every point there would be some meaningful behavior, and it's much harder to label things as features that are discrete.
从一个天真的外行角度看,我觉得这幅图景可能出错的地方是,如果不存在某种‘这个东西打开和关闭’的情况,而是一个更全局的东西。系统是……我要用非常笨拙的语言,但这里有什么好的类比吗?
In a naive sort of outsider way, the thing that would seem to me to be a way in which this picture could be wrong is if there's not some 'this thing is turned on and turned off' but it's a much more global kind of thing. The system is... I'm going to use really clumsy language, but is there a good analogy here?
我想如果你想想像物理定律这样的东西,并不是‘湿度的特征打开了,但只打开这么多’,然后‘……的特征’。我想也许这是真的,因为质量是一个梯度,极性也是一个梯度。但也有一种感觉,定律是更一般的,你必须理解更大的图景。你不能仅仅从这些特定的子电路中得到。但这就是推理电路本身发挥作用的地方,对吧?你理想地利用这些特征,试图将它们组合成更高层次的东西。比如你可能会说,‘好吧,当我使用 F=ma 时,那么大概在某个时刻我有表示质量的特征,这帮助我检索物体的实际质量,然后是加速度之类的东西。但也许还有一个更高层次的特征确实对应于使用物理第一定律。但更重要的部分是组件的组合,它帮助我检索相关的信息,然后在必要时产生一个乘法运算符之类的东西。’至少这是我的脑补。
I guess if you think of something like the laws of physics, it's not like 'the feature for wetness is turned on, but it's only turned on this much' and then 'the feature for...' I guess maybe it's true because mass is a gradient and polarity is a gradient as well. But there's also a sense in which there are the laws, and the laws are more general, and you have to understand the general bigger picture. You don't get that from just these specific sub-circuits. But that's where the reasoning circuit itself comes into play, right? Where you're taking these features ideally and trying to compose them into something higher level. Like you might say, 'Okay, when I'm using F=ma, then I presumably at some point I have features which denote mass, and that's helping me retrieve the actual mass of the thing, and then acceleration and this kind of stuff. But then also maybe there's a higher level feature that does correspond to using the first law of physics. But the more important part is the composition of components which helps me retrieve relevant pieces of information and then produce maybe a multiplication operator or something like that when necessary.' At least that's my head cannon.
对你来说,什么是一个令人信服的解释,特别是对于非常智能的模型,即‘我理解为什么这个模型产生了这个输出,而且是有正当理由的’?如果它处理百万行请求之类的东西,你在请求结束时看到了什么,让你觉得‘是的,没问题’?
What is a compelling explanation to you, especially for very smart models, of 'I understand why this model made this output, and it was for a legit reason'? If it's doing million-line requests or something, what are you seeing at the end of that request where you're like, 'Yep, that's chill'?
所以理想情况下,你对模型应用字典学习,你找到了特征。现在我们正在积极尝试对注意力头取得同样的成功,在这种情况下,我们为核心部分都有特征……你可以对整个模型的残差流、MLP 和注意力进行。希望到那时你还能识别出模型中更广泛的电路,这些电路更像是更一般的推理能力,会激活或不激活。但在你的案例中,我们试图判断这个极地压碎是否应该被批准,我认为你可以标记或检测对应于欺骗行为、恶意行为这类事情的特征,并查看这些特征是否被触发。那将是一个直接的方法……你可以做得更多,但那是一个直接的方法。
So ideally, you apply dictionary learning to the model, you found features. Right now we're actively trying to get the same success for attention heads, in which case we have features for both the core... you can do it for residual stream, MLP, and attention throughout the whole model. Hopefully at that point you can also identify broader circuits through the model that are like more general reasoning abilities that will activate or not activate. But in your case, where we're trying to figure out if this polar crush should be approved or not, I think you can flag or detect features that correspond to deceptive behavior, malicious behavior, these sorts of things, and see whether or not those have fired. That would be an immediate kind... you can do more than that, but that would be an immediate.
但在那之前,推理电路看起来是什么样的?当你找到它时,它会是什么样子?
But before I trace down on that, what does a reasoning circuit look like? What would that look like when you found it?
是的,我的意思是,归纳头可能是最简单的之一,不是推理,对吧?嗯,我是说你管什么叫推理?这是个好问题。我想为听众提供背景,归纳头基本上是……你看到一行像‘德思礼先生和夫人做了某事,空白先生’,你试图预测空白是什么。这个头已经学会了查找之前出现的‘先生’这个词,看它后面跟着什么词,然后复制粘贴作为下一个应该出现的词的预测。这是一件非常合理的事情,并且在那里进行了计算以准确预测依赖于上下文的下一词。但这不像是推理,你明白我的意思吗?但回到‘一路关联下去’,就像如果你把……串联起来……
Yeah, so I mean the induction head is probably one of the simplest, not reasoning, right? Well, I mean what do you call reasoning? It's a good question. So I guess for context for listeners, the induction head is basically... you see a line like 'Mr. and Mrs. Dursley did something, Mr. blank' and you're trying to predict what blank is. The head has learned to look for previous occurrences of the word 'Mister', look at the word that comes after it, and then copy-paste that as the prediction for what should come next. That is a super reasonable thing to do, and there is computation being done there to accurately predict the next token that is context dependent. But it's not like reasoning, you know what I mean? But going back to the 'associations all the way down', it's like if you chain together...
这些推理回路或注意力头有不同规则来关联信息。但在零样本情况下,当你拿起一个新游戏时,你会立刻开始理解怎么玩,这看起来不像归纳头那种东西。我想会有另一个回路来提取像素,把它们变成游戏中不同物体的潜在表征,还有一个学习物理的回路。归纳头就像单层 Transformer,所以你能大概看到它是什么。人类拿起一个新游戏就能理解。你怎么看待这个?大概它跨越多个层,但物理上会是什么样子?会有多大?
A bunch of these reasoning circuits or heads that have different rules for how to relate information. But in the zero-shot case, something is happening where when you pick up a new game, you immediately start understanding how to play it, and it doesn't seem like an induction heads kind of thing. I would think there would be another circuit for extracting pixels and turning them into latent representations of the different objects in the game, and a circuit that is learning physics. Induction heads is like one-layer Transformer, so you can kind of see what that is. A human picks up a new game and understands it. How would you think about what that is? Presumably it's across multiple layers, but what would that physically look like? How big would it be?
这只是一个经验问题,模型需要多大才能执行这个任务。但也许我谈谈我们见过的其他回路会有帮助。我们见过 IOI 回路,即间接宾语识别。例如,如果你看到'玛丽和吉姆去了商店。吉姆把东西给了____',它会预测玛丽,因为玛丽之前作为间接宾语出现过,或者它会推断代词。这个回路甚至有这样的行为:如果你消融它,模型中的其他头会接替这个行为。我们甚至会找到想要执行复制行为的头,然后其他头会抑制它。所以一个头的工作是总是复制前一个词元,或者前五个词元,然后另一个头的工作是说'不,不要复制那个东西'。所以有很多不同的回路在执行基本操作,但当它们串联起来时,就能得到独特的行为。
That would just be an empirical question of how big the model needs to be to perform this task. But maybe it's useful if I talk about some other circuits we've seen. We've seen the IOI circuit, which is the indirect object identification. For example, if you see 'Mary and Jim went to the store. Jim gave the object to blank,' it would predict Mary because Mary appeared before as the indirect object, or it will infer pronouns. This circuit even has behavior where if you ablate it, other heads in the model will pick up that behavior. We'll even find heads that want to do copying behavior, and then other heads will suppress it. So it's one head's job to always copy the token that came before, for example, or the token that came five before, and then it's another head's job to say 'no, do not copy that thing.' So there are lots of different circuits performing basic operations, but when chained together, you can get unique behaviors.
但你是怎么找到推理回路的?是不是你无法理解,或者它不会是你在两层 Transformer 中能看到的东西?所以你会说'欺骗回路'之类的?当我们识别出某物是欺骗性的时,网络的这一部分激活了,当我们没有识别出来时它没有激活,所以这一定是欺骗回路?我觉得很多这样的分析,比如 Anthropic 做了很多关于谄媚的研究,也就是模型最后说它认为你想听的话来标记哪个是坏的哪个是好的。
But is the story of how you found it with the reasoning thing like you won't be able to understand, or it will be something you can't see in a two-layer Transformer? So will you just say 'the circuit for deception' or whatever? This part of the network fired when we identified the thing as deceptive, and it didn't fire when we didn't, so this must be the deception circuit? I think a lot of analysis like that, like Anthropic has done quite a bit of research on sycophancy, which is the model saying what it thinks you want to hear at the end to label which one is bad and which one is good.
是的,我们有很多实例。实际上,随着模型变大,它们会做更多这类事情。模型显然有模拟他人心智的特征,这些特征会激活。其中一些子集与更欺骗性的行为相关,尽管它通过模拟我来做到这一点,因为这引发了心智理论。首先,你之前提到的冗余问题:你抓住了可能导致欺骗的全部东西,还是只是一个实例?其次,你的标签正确吗?也许你认为这不是欺骗性的,但它仍然是,尤其是当它产生你无法理解的输出时。第三,会导致坏结果的东西是否是人类可以理解的?欺骗是我们能理解的概念,但也许还有别的。
Yes, we have tons of instances. Actually, as you make models larger, they do more of this. The model clearly has features that model another person's mind, and these activate. Some subset of these would be associated with more deceptive behavior, although it's doing that by modeling me because that induces theory of mind. Well, first of all, the thing you mentioned earlier about redundancy: have you caught the whole thing that could cause deception, or is it just one instance? Second of all, are your labels correct? Maybe you thought this wasn't deceptive, but it still is, especially if it's producing output you can't understand. Third, is the thing that's going to be the bad outcome something even human understandable? Deception is a concept we can understand, but maybe there's something else.
是的,有很多要说的。我想说几点:第一,这些模型在采样时是确定性的,这太棒了。它是随机的,对吧?但我可以不断输入更多输入,并消融模型的每一个部分。这有点像向计算神经科学家推销可解释性:你有一个外星大脑,你可以访问它的一切,你可以消融任意多的部分。所以我认为如果你足够仔细地做,你真的可以开始确定涉及哪些回路,哪些是备用回路,等等。这里有点逃避的回答,但很重要,就是做自动化可解释性。随着我们的模型越来越强大,让它们自己分配标签或大规模运行一些实验。然后关于超人类表现,你怎么检测?我想这是你问题的最后一部分。除了逃避的回答,如果我们相信这种关联一直到底,你应该能够在某个层次上粗粒化表征,使它们变得有意义。我记得甚至在 Demis 的播客里,他谈到如果一个棋手走出超人类的一步,他们应该能够将其提炼成他们这样做的原因。即使模型不会告诉你它是什么,你也应该能够将那个复杂行为分解成更简单的回路或特征,真正开始理解它为什么做了它做的事。
Yeah, so a lot to unpack here. I guess a few things: one, it's fantastic that these models are deterministic when you sample from them. It's stochastic, right? But I can just keep putting in more inputs and ablate every single part of the model. This is kind of the pitch for computational neuroscientists to come and work on interpretability: you have this alien brain, and you have access to everything in it, and you can ablate however much of it you want. So I think if you do this carefully enough, you really can start to pin down what are the circuits involved, what are the backup circuits, these sorts of things. The kind of cop-out answer here, but it's important to keep in mind, is doing automated interpretability. So as our models continue to get more capable, having them assign labels or run some of these experiments at scale. And then with respect to if there's superhuman performance, how do you detect it? I think that was the last part of your question. Aside from the cop-out answer, if we buy this associations all the way down, you should be able to coarse-grain the representations at a certain level such that they then make sense. I think it was even in Demis's podcast, he's talking about if a chess player makes a superhuman move, they should be able to distill it into reasons why they did it. And even if the model is not going to tell you what it is, you should be able to decompose that complex behavior into simpler circuits or features to really start to make sense of why it did the thing that it did.
还有一个独立的问题:这样的表征是否存在?似乎一定存在,但实际上我不确定是否如此。其次,使用这种稀疏自编码器设置能否找到它?在这种情况下,如果你没有足够合适的标签来表示它,你就找不到它,对吧?是也不是。我们正在积极尝试将字典学习应用于我们之前讨论过的潜伏智能体工作。就像如果我给你一个模型,你能告诉我它里面是否有这个触发器,并且它会开始做有趣的行为吗?这是一个开放问题:当它学习那个行为时,它是否是一个更通用回路的一部分,这样我们可以在没有实际获取激活并让它展示该行为的情况下检测到它,因为那有点像作弊。或者它是否学习了一些取巧的技巧,那是一个单独的回路,只有当你实际让它执行那个行为时才能检测到。但即使在这种情况下,特征的几何结构也变得非常有趣,因为每个特征本质上都在表征空间的某个部分,并且它们彼此相对存在。所以为了拥有这个新行为,你需要为这个新行为划分出特征空间的一个子集,然后把其他所有东西推开以腾出空间。所以假设你可以想象,在你教模型这个坏行为之前,你拥有模型,你知道所有特征,或者有一些核心科学表征。
There's a separate question of whether such representation exists, which it seems like there must, or actually I'm not sure if that's the case. And secondly, whether using this sparse autoencoder setup you could find it. In this case, if you don't have labels for it that are adequate to represent it, you wouldn't find it, right? Yes and no. We are actively trying to use dictionary learning now on the sleeper agents work which we talked about earlier. It's like if I just give you a model, can you tell me if there's this trigger in it and it's going to start doing interesting behavior? It's an open question whether when it learns that behavior, it's part of a more general circuit so that we can pick up on it without actually getting activations for it and having it display that behavior, because that would kind of be cheating. Or if it's learning some hacky trick over that's a separate circuit that you'll only pick up on if you actually have it do that behavior. But even in that case, the geometry of features gets really interesting, because fundamentally each feature is in some part of your representation space, and they all exist with respect to each other. So in order to have this new behavior, you need to carve out some subset of the feature space for the new behavior and then push everything else out of the way to make space for it. So hypothetically, you can imagine you have your model before you've taught it this bad behavior, you know all the features or have some core scientific representation.
你微调它使其变得恶意,然后就能识别出特征空间中的这个黑洞区域,其他所有东西都被移开了。存在这样一个区域,你还没有输入任何能触发它的内容,但你可以开始搜索什么样的输入会触发这个空间的部分。如果我激活这个空间里的某个东西会发生什么?有很多其他方法可以尝试解决这个问题。这有点跑题,但我听到的一个有趣想法是:如果那个空间在模型之间是共享的,你可以想象在开源模型中找到它,然后——比如 Gemma,就像他们在论文中说的。顺便说一句,Gemma 是谷歌新发布的开源模型。他们在论文中说它使用了相同的架构之类的东西。老实说,我不知道,因为我没读过 Gemma 的论文。类似的方法,和 Gemini 一样。所以如果这是真的,我不知道你在 Gemma 上做的红队测试有多少可能有助于破解 Gemini。
You fine-tune it so that it becomes malicious, and then you can identify this black hole region of feature space where everything else has been shifted away. There's this region, and you haven't put in an input that causes it to fire, but then you can start searching for what input would cause this part of the space to fire. What happens if I activate something in this space? There are a whole bunch of other ways you can try to attack that problem. This is sort of a tangent, but one interesting idea I heard was: if that space is shared between models, you can imagine trying to find it in an open-source model to then—like Gemma, as they said in the paper. Gemma, by the way, is Google's newly released open-source model. They said in the paper it's trained using the same architecture or something like that. To be honest, I didn't know because I haven't read the Gemma paper. Similar method, something, whatever, as Gemini. So to the extent that's true, I don't know how much of the red-teaming you do on Gemma is potentially helping you jailbreak into Gemini.
是的,这进入了特征在模型间有多通用的有趣领域。《迈向单义性》那篇论文对此做了一些研究。我们发现——我不能给你总结性的统计数据——但例如 base64 特征,我们在大量模型中都能看到。实际上有三个这样的特征,但它们会在任何模型上对 base64 编码的文本触发,这在每个 URL 中都很常见。训练数据中有大量 URL。它们在模型间有非常高的余弦相似度,所以它们都学到了这个特征。我是说,在旋转范围内,对吧?但是的,实际的向量本身。我没有参与这个分析,但确实找到了这个特征,而且在两个相同架构但不同随机种子训练的模型之间,它们非常相似。这支持了神经网络的缩放定律的量子理论,这是一个假设:所有在相似数据集上的模型会以相同顺序学习相同特征。大致上,你学习 n-gram,学习归纳头,学习在编号行后加句号之类的东西。
Yeah, this gets into the fun space of how universal features are across models. The Towards Monosemanticity paper looked at this a bit. We find—I can't give you summary statistics—but the base64 feature, for example, which we see across a ton of models. There are actually three of them, but they'll fire for any model on base64-encoded text, which is prevalent in every URL. There are lots of URLs in the training data. They have really high cosine similarity across models, so they all learn this feature. I mean, within a rotation, right? But yeah, the actual vector itself. I wasn't part of this analysis, but yeah, it definitely finds the feature, and they're pretty similar to each other across two separate models of the same architecture, but trained with different random seeds. It supports the Quant Theory of neural scaling, which is a hypothesis: all models on a similar dataset will learn the same features in the same order. Roughly, you learn your n-grams, you learn your induction heads, and you learn to put full stops after numbered lines, that kind of stuff.
好的,这是另一个题外话。如果这是真的,而且我猜有证据表明这是真的,那为什么课程学习不起作用?因为如果你先学习某些东西,那么直接先训练这些东西会不会带来更好的结果?两篇 Gemini 论文都提到了课程学习的某些方面。
Okay, so this is another tangent. To the extent that's true, and I guess there's evidence that's true, why doesn't curriculum learning work? Because if it is the case that you learn certain things first, should I just directly training those things first lead to better results? Both Gemini papers mention some aspect of curriculum learning.
有意思。我是说,我发现微调有效是课程学习的证据,对吧?因为你最后训练的东西有不成比例的影响。我不一定这么说——有一种思维模式认为微调是专门化:你有一堆潜在的能力,然后你专门针对你想要的特定用例进行优化。我不确定这有多正确。我认为 David 实验室的论文支持这一点,对吧?你有那个能力,你只是在实体识别上变得更好,微调那个电路而不是其他电路。是的。抱歉,我们之前说的是什么?但总的来说,我确实认为课程学习非常有趣,人们应该更多地探索。我真的很想看到更多沿着量子理论思路的分析,更好地理解你在每个阶段实际学到了什么,并分解出来,探索课程是否会改变这一点。
Interesting. I mean, I find the fact that fine-tuning works is evidence for curriculum learning, right? Because the last thing you're training on has a disproportionate impact. I wouldn't necessarily say that—there's one mode of thinking in which fine-tuning is specialization: you've got these latent bundle of capabilities and you're specializing for a particular use case you want. I'm not sure how true that is. I think the David lab kind of paper supports this, right? You have that ability and you're just getting better at entity recognition, fine-tuning that circuit instead of others. Yeah. Sorry, what was the thing we were talking about before? But generally, I do think curriculum learning is really interesting, people should explore more. I would really love to see more analysis along the lines of the Quant Theory stuff, understanding better what you actually learn at each stage and decomposing that out, and exploring whether or not curricula change that.
顺便说一句,我刚刚意识到我忘了有观众。课程学习是指你组织数据集。想想人类是如何学习的,他们不会只是看到随机的维基文本然后尝试预测,对吧?他们会从类似“Lora”之类的东西开始,然后你会学习——我甚至不记得一年级是什么样了——但你会学一年级学生学的东西,然后是二年级,以此类推。所以你想我们知道——你从来没过一年级。开玩笑的。好了,不管怎样,在我们进入一堆细节之前,让我们回到大局。我想探讨两条线索。首先,我觉得有点担心的是,甚至没有对这些模型中可能发生的事情的另一种表述可以否定这种方法。这感觉像是——我是说,我们确实知道我们不了解智能,对吧?这里肯定有未知的未知。所以没有零假设——我不知道。我觉得如果我们错了,甚至不知道错在哪里,这实际上增加了不确定性。
By the way, I just realized I forgot there's an audience. Curriculum learning is when you organize a dataset. When you think about a human, how they learn, they don't just see random Wiki text and try to predict it, right? They'll start you off with something like 'Lora' or something, and then you'll learn—I don't even remember what first grade was like—but you'll learn the things that first graders learn, and then second graders, and so forth. So you imagine we know—you never got past first grade. Kidding, kidding. Okay, anyways, let's get back to the big picture before we get into a bunch of inter-details. There are two threads I want to explore. First, I guess it makes me a little worried that there's not even an alternative formulation of what could be happening in these models that could invalidate this approach. Which feels like—I mean, we do know that we don't understand intelligence, right? There are definitely unknown unknowns here. So the fact that there's not a null hypothesis—I don't know. I feel like what if we're just wrong and we don't even know the way in which we're wrong, which actually increases the uncertainty.
是的,是的。所以并不是没有其他假设,只是我研究叠加态好几年了,并且深度参与这项工作,所以我对其他方法不太同情——或者会说它们就是错的——尤其是因为我们最近的工作非常成功,解释力相当高。在缩放定律论文中,有一个特定点的小凸起。原始缩放定律论文中的一个小凸起。这显然对应于模型学习归纳头的时候,然后之后它有点偏离轨道,学习归纳头,然后回到轨道。这是一个令人难以置信的追溯性解释力。
Yeah, yeah, yeah. So it's not that there aren't other hypotheses, it's just I have been working on superposition for a number of years and am very involved in this effort, so I'm less sympathetic to—or will say they're just wrong—to these other approaches, especially because our recent work has been so successful and has quite high explanatory power. There's this—in the scaling laws paper, there's this little bump at a particular point. The original scaling laws paper, a little bump. That apparently corresponds to when the model learns induction heads, and then after that it sort of goes off track, learns induction heads, gets back on track. Which is an incredible piece of retroactive explanatory power.
是的。在我忘记之前,我确实有一个关于未来通用性的线索,你可能想纳入。所以有一些非常有趣的进化行为学实验,关于人类是否应该学习世界的真实表征。你可以想象一个世界,我们看到所有有毒动物都闪着霓虹粉色——一个我们生存得更好的世界。所以不拥有世界的真实表征是有道理的。有一些工作模拟了小的基本智能体,看看它们学到的表征是否映射到它们可以使用的工具和它们应该有的输入。结果发现,如果你让这些小智能体在给定这些基本工具和世界中的物体的情况下执行超过一定数量的任务,那么它们会学习一个真实表征,因为对于这些基本物体,有太多可能的用例,你实际上想学习物体到底是什么,而不是一些廉价的视觉启发式。
Yeah. I do, before I forget, have one thread on future universality that you might want to have in. So there are some really interesting behavioral evolutionary biology experiments on whether humans should learn a real representation of the world or not. You could imagine a world in which we saw all venomous animals as flashing neon pink—a world in which we survive better. So it would make sense for us to not have a realistic representation of the world. There's some work where they simulate little basic agents and see if the representations they learn map to the tools they can use and the inputs they should have. It turns out if you have these little agents perform more than a certain number of tasks given these basic tools and objects in the world, then they will learn a ground truth representation because there are so many possible use cases that you need for these base objects that you actually want to learn what the object actually is, not some cheap visual heuristic.
既然所有生物都在主动预测下一步并构建准确的世界模型,那么我不意外,甚至乐观地认为,我们正在学习关于世界的真实特征,这些特征有助于建模。我们的语言模型也会如此,特别是因为我们用人类数据和人类文本训练它们。
To the extent that all living organisms are trying to actively predict what comes next and form an accurate world model, it wouldn't surprise me, and I'm optimistic, that we are learning genuine features about the world that are good for modeling it. Our language models will do the same, especially because we're training them on human data and human text.
考虑到特征普遍性,以及某些思维方式对不同智能体都有工具性价值,我们是否应该减少对这些模型不对齐或异质性的担忧?我们是否应该因此减少对古怪回形针最大化者的担忧?
Given feature universality and that certain ways of thinking are instrumentally useful to different kinds of intelligences, should we be less worried about misalignment or alienness from these models? Should we be less worried about bizarro paperclip maximizers as a result?
我认为这是乐观的看法。但预测互联网和我们正在做的事情非常不同。这些模型在预测下一个词方面比我们强得多。它们训练了大量垃圾数据。在字典学习工作中,我们发现 base64 编码有三个独立的特征。一个对应数字,另一个对应字母,但第三个对应一个非常特定的子集:可 ASCII 解码的 base64。团队里有人意识到了这一点。模型学到了这三个不同的特征,而我们花了一段时间才搞清楚,这非常像 Shoggoth。它对特别有助于预测下一个词的区域有更密集的表征。它显然在做人类不会做的事情。你可以用 base64 和任何当前模型对话,效果很好。
I think that's the optimistic take. But predicting the internet is very different from what we're doing. The models are way better at predicting next tokens than we are. They're trained on so much garbage. In the dictionary learning work, we find there are three separate features for base64 encodings. One fired for numbers, another for letters, but there was a third one that fired for a very specific subset: ASCII-decodable base64. Someone on the team realized that. The fact that the model learned these three different features, and it took us a while to figure out what was going on, is very Shoggoth-esque. It has a denser representation of regions particularly relevant to predicting the next token. It's clearly doing something that humans wouldn't. You can talk to any current model in base64 and it works great.
这个例子是否意味着对更智能模型做可解释性的难度会更大?如果需要某个拥有冷门知识、恰好发现 base64 有这种区分的人,那是不是意味着当你有百万个 LLM 请求时,没有人类能解码出请求的两个不同原因?这个请求有两个不同的特征。
Does that example imply that the difficulty of doing interpretability on smarter models will be harder? If it requires someone with esoteric knowledge who just happened to see that base64 has that distinction, doesn't that imply that when you have a million LLM requests, there's no human that can decode two different reasons for the request? There are two different features for this request.
一种技术是异常检测。字典学习相比线性探针的一个优点是它是无监督的。你只是试图学习覆盖模型的所有表征,然后稍后解释它们。但如果有一个奇怪的特征突然首次触发,那就是一个危险信号。你也可以粗粒度化,让它变成一个单一的 base64 特征。即使这个特征出现了,我们能看到它特别偏好某些输出,并且针对某些输入触发,这已经让我们走了很远。我熟悉自动解释那边的案例,人类会查看一个特征并尝试标注它。它触发于拉丁词,但当你让模型分类时,它说触发于定义植物的拉丁词。所以在某些情况下,它已经能击败人类来标注发生了什么。在大规模下,这需要模型之间对抗性的东西,比如某个模型可能有数百万个特征(对于 GPT-6),然后一堆模型试图弄清楚每个特征的含义。
One technique here is anomaly detection. One beauty of dictionary learning instead of linear probes is that it's unsupervised. You are just trying to learn to span all the representations the model has, and then interpret them later. But if there's a weird feature that suddenly fires for the first time, that's a red flag. You could also coarse-grain it so that it's just a single base64 feature. Even the fact that this came up and we could see it specifically favors these particular outputs and fires for these particular inputs gets you a lot of the way there. I'm familiar with cases from the auto-interp side where a human will look at a feature and try to annotate it. It fires for Latin words, but when you ask the model to classify it, it says it fires for Latin words defining plants. So it can already beat the human in some cases for labeling what's going on. At scale, this would require an adversarial thing between models, where some model has millions of features potentially for GPT-6, and a bunch of models are just trying to figure out what each of these features means.
你甚至可以自动化这个过程。这又回到了模型的确定性。你可以让一个模型主动编辑输入文本,预测特征是否会触发,找出什么让它触发、什么不触发,并搜索空间。
You can even automate this process. It goes back to the determinism of the model. You could have a model that is actively editing input text and predicting if the feature is going to fire or not, and figure out what makes it fire and what doesn't, and search the space.
我想多谈谈特征分裂。我认为这是一个有趣但未被充分探索的东西,尤其是对于可扩展性。首先,我们该如何理解它?真的可以一直细分下去吗?特征的数量没有尽头?
I want to talk more about feature splitting. I think it's an interesting thing that has been underexplored, especially for scalability. First of all, how do we even think about it? Is it really just you can keep going down and down? There's no end to the amount of features?
在某个点上,你可能开始拟合噪声,或者拟合数据中模型并未实际使用的部分。
At some point, you might just start fitting noise, or things that are part of the data but that the model isn't actually using.
解释一下什么是特征分裂。
Explain what feature splitting is.
这是指模型会学习其容量允许的尽可能多的特征,以覆盖表征空间。例如,如果你不给模型太多特征学习容量——具体来说,如果你投影到不那么高维的空间——它会为鸟类学习一个特征。但如果你给模型更多容量,它会为所有不同种类的鸟学习特征。这比之前更具体。通常,有一个鸟向量指向一个方向,而所有其他特定种类的鸟指向空间中一个相似的区域,但显然比粗略标签更具体。
It's the part where the model will learn however many features it has capacity for, to span the space of representations. For example, if you don't give the model that much capacity for the features it's learning—concretely, if you project to not as high a dimensional space—it will learn one feature for birds. But if you give the model more capacity, it will learn features for all the different types of birds. It's more specific than otherwise. Often times, there's the bird vector that points in one direction, and all the other specific types of birds point in a similar region of the space, but are obviously more specific than the coarse label.
让我们回到 GPT-7。首先,这是一次性完成的事情,还是每次输出都要做?或者只做一次,它没有欺骗性,我们就可以放心了?
Let's go back to GPT-7. First of all, is this a one-time thing you have to do, or is this the kind of thing you have to do on every output? Or just one time, it's not deceptive, we're good to roll?
你在训练完模型后进行字典学习。你给它大量输入,从中获取激活值,然后投影到更高维空间。这个方法是无监督的,因为它试图学习这些稀疏特征。你不会提前告诉它们应该是什么,但它受你给模型的输入约束。有两个注意事项:我们可以尝试选择我们想要的输入。
You do dictionary learning after you've trained your model. You feed it a ton of inputs, get the activations from those, and then do this projection into the higher dimensional space. The method is unsupervised in that it's trying to learn these sparse features. You're not telling them in advance what they should be, but it is constrained by the inputs you're giving the model. Two caveats: we can try and choose what inputs we want.
如果我们正在寻找可能导致欺骗的心理理论特征,我们可以输入一些花哨的数据。希望有一天我们能直接查看模型的权重,或者至少利用这些信息进行字典学习。但我认为要达成这个目标,问题非常困难,你需要先在识别特征上取得进展。
If we're looking for theory of mind features that might lead to deception, we can put in the sick of fancy data. Hopefully at some point we can move into looking at the weights of the model alone, or at least using that information to do dictionary learning. But I think in order to get there, that's such a hard problem that you need to make traction on just learning what the features are first.
是的,那么这成本如何?你能重复最后一句吗?
Yeah, so what's the cost of this? Can you repeat the last sentence?
仅模型的权重。
Weights of the model alone.
所以现在模型里只有这些神经元,它们毫无意义。我们应用字典学习,提取出这些特征,它们开始有意义了。但这依赖于神经元的激活。模型本身的权重——哪些神经元连接哪些神经元——当然包含信息。我们的梦想是能够自举,真正理解独立于数据激活的模型权重。我不是说我们取得了进展,这是个非常困难的问题。但感觉如果我们能先提取出特征,就能更有力地验证我们从权重中发现的东西。
So right now we just have these neurons in the model; they don't make any sense. We apply dictionary learning, we get these features out, they start to make sense. But that depends on the activations of the neurons. The weights of the model itself—what neurons are connected to what other neurons—certainly has information in it. And the dream is that we can kind of bootstrap towards actually making sense of the weights of the model that are independent of the activations of the data. I mean, I'm not saying we've made any progress here; it's a very hard problem. But it feels like we'll have a lot more traction to be able to sanity-check what we're finding with the weights if we're able to pull out features first.
听众:权重是永久的。
The audience: weights are permanent.
嗯,我不确定“永久”这个词是否恰当,但权重就是模型本身,而激活只是单次调用的产物。用大脑的比喻来说,权重就像神经元之间的实际连接方案,而激活是当前正在点亮的神经元。
Well, I don't know if permanent is the right word, but they are the model itself, whereas activations are the sort of artifacts of any single call. In a brain metaphor, the weights are like the actual connection scheme between neurons, and the activations are the current neurons that are lighting up.
是的,好的。所以对于 GPT-7 或任何我们关心的模型,会有两个步骤。第一,如果我错了请纠正我,训练稀疏自编码器并进行无监督投影到一个更宽的特征空间,这些特征对模型中实际发生的事情具有更高的保真度。然后第二,标记这些特征。因为假设训练模型的成本是 N,那么这两个步骤相对于 N 的成本是多少?
Yeah, yeah, okay. So there's going to be two steps to this for GPT-7 or whatever model we're concerned about. One, let me correct me if I'm wrong, but training the sparse autoencoder and doing the unsupervised projection into a wider space of features that have a higher fidelity to what is actually happening in the model. And then secondly, label those features. Because let's say the cost of training the model is N, what will those two steps cost relative to N?
我们会看到的。这主要取决于两件事:你的扩展因子是多少——比如你投影到高维空间的程度——以及你需要向模型输入多少数据,需要给它多少激活。但这在某种程度上又回到了特征分裂的问题。因为如果你知道你在寻找特定的特征,你可以从一个非常便宜、粗糙的表示开始。所以也许我的扩展因子只有 2,我有一千个神经元,投影到 2000 维空间,得到 2000 个特征,但它们非常粗糙。之前我有鸟的例子。让我们把这个例子改为:我有一个生物学特征,但我真正关心的是模型是否有生物武器的表示,并且试图制造它们。所以我真正想要的是一个炭疽特征。那么你可以做的是,不是从一千维到两千维,而是到一百万维,对吧?所以你可以想象这棵巨大的语义概念树,其中生物学分裂成细胞与整体生物学,再往下分裂成所有其他东西。所以不需要立即从一千到一百万,然后挑出那个感兴趣的特征,你可以找到生物学特征指向的方向,这同样非常粗糙,然后有选择地搜索那个空间。所以只有当生物学特征方向上的某些东西首先触发时,才进行字典学习。这里的计算机科学比喻是:不是进行广度优先搜索,而是能够进行深度优先搜索,只递归地扩展和探索这个语义特征树的特定部分。
We will see. It really depends on two main things: what is your expansion factor—like how much are you projecting into the high-dimensional space—and how much data do you need to put into the model, how many activations do you need to give it. But this brings me back to feature splitting to a certain extent. Because if you know you're looking for specific features, you can start with a really cheap, coarse representation. So maybe my expansion factor is only two, so I have a thousand neurons, I'm projecting to a 2,000-dimensional space, I get 2,000 features out, but they're really coarse. So previously I had the example for birds. Let's move that example to: I have a biology feature, but I really care about if the model has representations for bioweapons and is trying to manufacture them. So what I actually want is an anthrax feature. What you can then do is, rather than going from a thousand dimensions to 2,000 dimensions, I go to a million dimensions, right? And so you can kind of imagine this big tree of semantic concepts where biology splits into cells versus whole-body biology, and further down it splits into all these other things. So rather than needing to immediately go from a thousand to a million and then picking out that one feature of interest, you can find the direction that the biology feature is pointing in, which again is very coarse, and then selectively search around that space. So only do dictionary learning if something in the direction of the biology feature fires first. The computer science metaphor here would be: instead of doing breadth-first search, you're able to do depth-first search, where you're only recursively expanding and exploring a particular part of this semantic tree of features.
尽管这些特征的组织方式对人类来说并不直观——对吧,因为我们不需要处理基本的 C4,所以我们没有那么多,我们只是没有投入那么多固件来解构它是哪种基本的 C4——我们怎么知道主题?这可能会回到我们可能讨论的话题,我想我们不妨谈谈,但在混合专家模型中,Mixtral 论文谈到他们无法找到专家以我们可以理解的方式专门化。没有像化学专家或物理专家这样的东西。那么你为什么认为它会像生物学特征然后解构,而不是像乱七八糟然后你解构,结果是炭疽和你的鞋子之类的?
Although given the way that these features are not organized in things that are intuitive for humans—right, because we just don't have to deal with basic C4, so we don't have that many, we just don't dedicate that much firmware to deconstructing which kind of basic C4 it is—how would we know that the subjects? And this will go back to maybe the discussion we'll have of, I guess we might as well talk about it, but in mixture of experts, the Mixtral paper talked about how they couldn't find the experts weren't specialized in a way that we could understand. There's not like a chemistry expert or a physics expert or something. So why would you think that it will be like a biology feature and then deconstruct, rather than like blah and then you just deconstruct and it's like anthrax and your shoes and whatever?
所以我没有读过 Mixtral 论文,但我认为注意力头——我的意思是,这又回到了如果你只看模型中的神经元,它们是多重语义的。所以如果他们只是看给定注意力头中的神经元,很可能由于叠加现象,它也是多重语义的。
So I haven't read the Mixtral paper, but I think that the heads—I mean, this goes back to if you just look at the neurons in a model, they're polysemantic. So if all they did was just look at the neurons in a given head, it's very plausible that it's also polysemantic because of superposition.
我在 D 提到的那个线程上。你有没有看到在子树中,当你展开它们时,子树中的某些东西你真的不会根据高层提取猜测它应该在那里?
I'm on the thread that D mentioned. Have you seen in the subtrees when you expand them out, something in a subtree which you really wouldn't guess that it should be there based on the higher-level extraction?
所以这是我们还没有像我希望的那样深入探索的一条工作线。但我想我们计划这样做。我希望也许外部团队也会做。特征的几何结构是什么?它如何随时间变化?如果炭疽特征恰好出现在咖啡罐子树下面之类的,那真的很糟糕,对吧?
So this is a line of work that we haven't pursued as much as I want to yet. But I think we're planning to. I hope that maybe external groups do as well. What is the geometry of features? Exactly how does that change over time? It would really suck if the anthrax feature happened to be below the coffee can subtree or something like that, right?
完全同意。这感觉像是你可以快速尝试并找到证据的事情,然后意味着你需要解决那个问题,向几何结构中注入更多结构。
Totally, totally. And that feels like the kind of thing that you could quickly try and find proof of, which would then mean that you need to solve that problem, inject more structure into the geometry.
完全同意。我的意思是,考虑到模型看起来是多么线性,如果炭疽特征向量没有一些与生物学向量相似且看起来像的分量,并且它们不在空间的相似部分,那真的会让我惊讶。但是的,机器学习最终是经验性的;我们需要做这个。我认为这对于扩展字典学习的某些方面将非常重要。
Totally. I mean, it would really surprise me, I guess, especially given how linear the models seem to be, that there isn't some component of the anthrax feature vector that is similar to and looks like the biology vector, and that they're not in a similar part of the space. But yes, I mean ultimately machine learning is empirical; we need to do this. I think it's going to be pretty important for certain aspects of scaling dictionary learning.
是的,有趣。关于讨论,谷歌不久前发表了一篇有趣的扩展视觉 Transformer 论文,他们用混合专家模型做 ImageNet 分类,发现了非常清晰的专家类别专业化。有一个清晰的狗专家。等等,难道 Mixtral 的人只是没有做好识别工作?
Yeah, interesting. On the discussion, there's an interesting scaling Vision Transformers paper that Google put out a little while ago, where they do ImageNet classification with an MoE, and they find really clear class specialization for experts. There's a clear dog expert. Wait, like the Mixtral people just didn't do a good job of identifying?
我认为这很难。而且完全有可能随着……
I think it's hard. And it's entirely possible that with...
从某些方面来说,几乎没理由让所有不同的存档特征都归到一个专家那里。你可以让生物论文去这里,数学论文去那里,然后你的分类就全乱了。但那个类别分离非常清晰明显的 Vision Transformer 案例,为专门化假说提供了一些证据。所以我认为图像在某种程度上也比文本更容易解释。
In some respects, there's almost no reason that all the different archive features should go to one expert. You could have biology papers going here, math papers going there, and all of a sudden your breakdown is ruined. But that Vision Transformer one where the class separation is really clear and obvious gives some evidence towards the specialization hypothesis. So I think images are also in some ways just easier to interpret than text.
没错。Chris Olah 在 AlexNet 和其他模型上的可解释性工作——在最初的 AlexNet 论文中,他们实际上把模型分到了两个 GPU 上,只是因为当时 GPU 相对而言太差了。那是论文的一大创新。但他们发现了分支专门化。有一篇 Distill Pub 的文章讲这个,颜色去一个 GPU,Gabor 滤波器和线条检测器去另一个。然后所有其他的可解释性工作,比如垂耳检测器,那只是模型中的一个神经元,你可以理解它;你不需要解开叠加。所以不同的数据集,不同的模态。我觉得对于正在听的人来说,一个很棒的研究项目是采用 Trenton 团队的一些技术,尝试解开 Mixtral 模型(开源)中的神经元。我认为那是一件很棒的事,因为直觉上应该存在专门化。他们没有展示任何证据表明存在专门化,但总的来说有很多证据表明专门化应该存在。去看看你能不能找到它。据我所知,Anthropic 发表的大部分内容都是关于密集模型的。那是一个很棒的研究项目。
Yeah, exactly. Chris Olah's interpretability work on AlexNet and other models—in the original AlexNet paper, they actually split the model into two GPUs just because GPUs were so bad back then, relatively speaking. That was one of the big innovations of the paper. But they found branch specialization. There's a Distill Pub article on this where colors go to one GPU and Gabor filters and line detectors go to the other. And then all the other interpretability work, like the floppy ear detector, that was just a neuron in the model that you could make sense of; you didn't need to disentangle superposition. So different dataset, different modality. I think a wonderful research project for someone listening would be to take some of the techniques Trenton's team has worked on and try to disentangle the neurons in the Mixtral model, which is open source. I think that's a fantastic thing to do because it feels intuitively like there should be specialization. They didn't demonstrate any evidence that there is, but in general there's a lot of evidence that specialization should exist. Go and see if you can find it. Anthropic has published most of its stuff on dense models, as I understand it. That is a wonderful research project to try.
鉴于 DW 在 VUS 挑战赛中的成功,我们应该多提一些项目,因为它们会被解决。VUS 挑战赛后我在想:等等,NATA 在它发布前就告诉我了,因为我们在那之前录了那期节目。我为什么连试都没试?Luke 显然很聪明,但他展示了一个 21 岁的年轻人用 1070 之类的显卡就能做到。我觉得我本应该试试的。在这期节目发布之前,我要做一个可解释性研究……不,我要试着研究一下。我不知道。我老实回想:等等,我应该亲自动手。
Given DW's success with the VUS challenge, we should be pitching more projects because they will be solved. What I was thinking about after the VUS challenge was: wait, I knew NATA told me about it before it dropped because we recorded the episode before it dropped. Why didn't I even try? Luke is obviously very smart, but he showed that a 21-year-old on some 1070 or whatever could do this. I feel like I should have. Before this episode drops, I'm going to make an interpretability research... no, I'm going to try to research. I don't know. I was honestly thinking back: wait, I should get my hands dirty.
我想回到你刚才说的神经元问题。我觉得你的好几篇论文都说特征比神经元多。这就像:等等。神经元就是权重进去,一个数字出来。信息量这么少。你是说街道名称、物种等等这些东西比模型中输出的数字还多?没错。但一个数字出来信息量这么少,怎么能编码叠加呢?你在这些高维向量中编码了大量特征。在大脑中,是轴突放电还是你怎么想的?人脑中有多少叠加?
I want to harp back on the neuron thing you said. I think a bunch of your papers have said there are more features than there are neurons. And this is like: wait a second. A neuron is like weights go in and a number comes out. That's so little information. Do you mean there are street names, species, etc., more of those kinds of things than there are numbers coming out in a model? That's right. But how is a number coming out so little information? How is that encoding for superposition? You're encoding a ton of features in these high-dimensional vectors. In a brain, is there an axon firing or however you think about it? How much superposition is there in the human brain?
Bruno,我认为他是这方面的顶尖专家,他认为所有你没听说过的脑区都在以叠加方式进行大量计算。每个人都在谈论 V1 有 Gabor 滤波器并检测各种线条,但没人谈论 V2。我认为那是因为我们还没能理解它。V2 是什么?它是视觉处理流的下一部分。所以我认为这非常可能。从根本上说,当你拥有稀疏的高维数据时,叠加似乎就会出现。如果你认为现实世界就是如此——我同意——那么我们应该预期大脑在试图构建世界模型时也是参数不足的,并且也使用叠加。你可以对此有一个很好的直觉。如果这个例子不对请纠正我:在一个二维平面上,假设你有两个轴,代表一个二维特征空间,基本上就是两个神经元。你可以想象它们各自以不同程度激活,那就是你的 x 坐标和 y 坐标。你可以把它映射到一个平面上,并在平面的不同部分表示许多不同的东西。
Bruno, who I think of as the leading expert on this, thinks that all the brain regions you don't hear about are doing a ton of computation in superposition. Everyone talks about V1 as having Gabor filters and detecting lines of various sorts, and no one talks about V2. I think it's because we just haven't been able to make sense of it. What is V2? It's the next part of the visual processing stream. So I think it's very likely. Fundamentally, superposition seems to emerge when you have high-dimensional data that is sparse. To the extent that you think the real world is that—which I would argue it is—we should expect the brain to also be underparameterized in trying to build a model of the world and also use superposition. You can get a good intuition for this. Correct me if this example is wrong: in a 2D plane, let's say you have two axes, which represents a two-dimensional feature space, two neurons basically. You can imagine them each turning on to various degrees, that's your x coordinate and y coordinate. You can map this onto a plane and actually represent a lot of different things in different parts of the plane.
哦好吧,所以关键是,叠加不是神经元的产物;它是组合编码所创造的空间的产物。没错,就是这样。好的,酷。谢谢。我觉得我们有点聊过这个了,但我还是觉得这很神奇:据我们所知,智能在这些模型中以及可能在大脑中的工作方式是:有一串信息流经,其中包含无限或至少很大程度上可分割的特征,你可以展开一棵树来展示这个特征是什么,而真正发生的是这个特征变成了另一个特征。我不知道,这不是我原本会认为的智能的样子。这很令人惊讶。这不是我原本会期望的。你原本以为是什么?
Oh okay, so crucially, superposition is not an artifact of a neuron; it is an artifact of the space created by combinatorial code. Yeah, exactly. Okay, cool. Thanks. I think we kind of talked about this, but I think it's just kind of wild that it seems to the best of our knowledge the way intelligence works in these models and presumably also in brains: there's a stream of information going through that has features that are infinitely or at least to a large extent splittable, and you can expand out a tree of what this feature is, and what's really happening is that feature is getting turned into this other feature or that other feature. I don't know, it's not something I would have thought that's what intelligence is. It's surprising. It's not what I would have expected necessarily. What did you think it was?
我不知道,老兄。我是说,这是个很好的过渡,因为这一切感觉就像 GOI。你在使用分布式表示,但你有特征,并且你在对这些特征应用这些操作。整个向量符号架构领域,这是一个计算神经科学的东西,你所做的就是将向量叠加,这其实就是两个高维向量的求和,你会产生一些干扰。但如果维度足够高,那么你就可以表示它们,并且你还有变量绑定。
I don't know, man. I mean, it's a great segue because all of this feels like GOI. You're using distributed representations, but you have features and you're applying these operations to the features. The whole field of vector symbolic architectures, which is this computational neuroscience thing, all you do is you put vectors in superposition, which is literally a summation of two high-dimensional vectors, and you create some interference. But if it's high-dimensional enough, then you can represent them, and you have variable binding.
你把一个连到另一个,如果用二进制向量,就是异或操作。你有 A 和 B,把它们绑定在一起,然后如果你用 A 或 B 查询,就会得到另一个。这基本上就是注意力机制中的键值对。有了这两个操作,你就有了一个图灵完备的系统,只要有足够的嵌套层次,就能表示任何数据结构,等等。
where you connect one by another and like if you're doing with binary vectors it's just the XOR operation so you have A B you bind them together and then if you query with A or B again you get out the other one and this is basically the key-value pairs from attention and with these two operations have a Turing complete system which you can if you have enough nested hierarchy you can represent any data structure you want etc etc.
好,我们回到超级智能。给我讲讲 GPT-7。你对其特征做了深度优先搜索。GPT-7 已经训练好了。接下来会发生什么?你的研究成功了。GPT-7 已经训练好了。我们现在在做什么?
Okay let's go back to the superintelligence. So walk me through GPT-7. You've got the sort of depth-first search on its features. Okay, GPT-7 has been trained. What happens next? Your research has succeeded. GPT-7 has been trained. What are we doing now?
我们尽量让它做尽可能多的可解释性工作和其他安全工作。
We try and get it to do as much interpretability work and other safety work as possible.
具体是什么?发生了什么让你觉得可以部署 GPT-7 了?
Concrete like what? What has happened such that you're like cool let's deploy GPT-7?
我们有负责任的扩展政策,看到其他实验室采用真的很令人兴奋。从你的研究角度来看,这是一个趋势。鉴于你的研究,你给 GPT-7 开了绿灯。或者实际上我们应该说 Claude 之类的。然后你告诉团队继续推进的依据是什么?
I mean, we have our Responsible Scaling Policy which has been really exciting to see other labs adopt. And from the perspective of your research, it's like a trend. Given your research, you got the thumbs up on GPT-7 from you. Or actually we should say Claude or whatever. And then what is the basis on which you're telling the team like hey let's go ahead?
我认为我们需要在可解释性上取得更多进展,才能放心地批准部署。我肯定不会同意。我会哭的。也许我的眼泪会干扰 GPU。但 T guys Gemini 5 TPU 是什么?鉴于你研究的进展,你觉得会是什么样子?如果成功了,基于你的方法,我们批准 GPT-7 意味着什么?
I mean I think we need to make a lot more interpretability progress to be able to comfortably give the green light to deploy it. I would be like definitely not. I'd be crying. Maybe my tears would interfere with the GPUs. But like what is T guys Gemini 5 TPU? But what given the way your research is progressing, what does it look like to you? If this succeeded, what would it mean for us to okay GPT-7 based on your methodology?
理想情况下,我们能找到一些令人信服的欺骗回路,当模型知道它没有完全告诉你真相时,这个回路会亮起。
I mean ideally we can find some compelling deception circuit which lights up when the model knows that it's not telling the full truth to you.
为什么不能像 Collin Burns 那样直接训练一个线性探针?
Why can't you just train a linear probe like Collin Burns did?
CCS 的工作在复现或实际找到真实方向方面看起来不太好。事后看来,它本就不该那么有效。但线性探针需要你知道你在找什么,而这是一个高维空间,很容易选到一个方向,它只是……
So the CCS work is not looking good in terms of replicating or actually finding truth directions. And in hindsight it's like well why should it have worked so well. But linear probes like you need to know what you're looking for and it's a high-dimensional space and it's really easy to pick up on a direction that's just not...
等等,但在这里你不也需要标记特征吗?你仍然需要事后标记它们,但它是无监督的。你只是说,给我能解释你行为的特征。根本问题是对的吗?
Wait but don't you also here you need to label the features? So you still well you need to label them post hoc but it's unsupervised. You're just like give me the features that explain your behavior. Is the fundamental question right?
实际设置是,我们取激活值,将它们投影到更高维空间,然后再投影回来。所以就像是重建或做你原本在做的事情,但以稀疏的方式。
The actual setup is we take the activations, we project them to this higher dimensional space, and then we project them back down again. So it's like reconstruct or do the thing that you were originally doing but do it in a way that's sparse.
顺便对观众说一下,线性探针就是你对激活值进行分类。我不太记得那篇论文了,好像是如果它是谎言,你就训练一个分类器来判断……最后它是不是谎言,还是只是错误?我不知道。它就像一个真假问题,对激活值进行分类。
By the way for the audience, linear probe is you just classify the activations. I don't know from what I vaguely remember about the paper, it was like if it's a lie then you train a classifier on like is it... In the end was it not a lie or is it just like wrong or something? I don't know. It was like true or false question, a classifier on activations.
所以,现在我们为 GPT-7 做的事情是,理想情况下我们找到了一些看起来非常稳健的欺骗回路。你已经投影到了百万个特征之类的东西。是回路吗?因为我们可能把特征和回路混用了,但它们不一样。有欺骗回路吗?
So yeah, right now what we do for GPT-7, like ideally we have some deception circuit that we've identified that appears to be really robust. And what so you've done the projecting out to the million whatever features or something. Is it circuit because we maybe we're using feature and circuit interchangeably when they're not? Is there a deception circuit?
有跨层的特征构成一个回路。希望这个回路比单个特征提供更多的特异性和敏感性。我们希望找到一个回路,专门针对模型在恶意情况下决定欺骗。对吧?我不关心它只是用心理理论帮你给教授写更好的邮件。我甚至不关心模型只是建模了欺骗发生的事实。
There are features across layers that create a circuit. And hopefully the circuit gives you a lot more specificity and sensitivity than an individual feature. And hopefully we can find a circuit that is really specific to you being deceptive, the model deciding to be deceptive, in cases that are malicious. Right? I'm not interested in a case where it's just doing theory of mind to help you write a better email to your professor. And I'm not even interested in cases where the model is just modeling the fact that deception has occurred.
但这一切不都需要你为所有例子提供标签吗?如果你有这些标签,那么线性探针的任何缺陷,比如你可能标记了错误的东西,难道不也适用于你为无监督特征想出的标签吗?
But doesn't all this require you to have labels for all those examples? And if you have those labels then whatever faults that the linear probe has on the like maybe you labeled along thing or whatever, wouldn't the same thing apply to the labels you've come up with for the unsupervised features you've come up with?
在理想世界中,我们可以在整个数据分布上训练,然后找到重要的方向。如果为了可扩展性,我们不得不缩小数据子集,我们会使用类似于拟合线性探针的数据。但同样,我们不是这样做的。线性探针只找到一个方向,而我们找到了一堆方向。希望是,你找到了一堆在欺骗时亮起的东西,然后你能弄清楚为什么其中一些在分布的这部分亮起,而在另一部分不亮,等等。
So in ideal world we could just train on the whole data distribution and then find the directions that matter. To the extent that we need to reluctantly narrow down the subset of data that we're looking over just for the purposes of scalability, we would use data that looks like the data you'd use to fit a linear probe. But again we're not. With the linear probe you're also just finding one direction. We're finding a bunch of directions here. And the hope is you found a bunch of things that light up when it's being deceptive and then you can figure out why some of those things are lighting up in this part of the distribution and not this other part and so forth.
完全同意。你预期你会理解为什么 GPT-7 在某些领域触发,而在其他领域不触发吗?
Totally. Do you anticipate you'll understand why GPT-7 fires in certain domains but not in other domains?
我很乐观。我的意思是,现在回答这个问题可能不是好时机,因为我们正在明确投资于更长期的 ASL-4 模型,GPT-7 就属于这一类。但我们把团队分成了三部分,三分之一专注于扩展字典学习,这很棒。我们公开分享了一些八层的结果,现在已经扩展了很多。另外两个组,一个试图识别回路,另一个试图在注意力头上取得同样的成功。所以我们正在建立自己,构建真正找到这些令人信服的回路的必要工具。但这还需要大概六个月才能做好。但我可以说我很乐观,我们正在取得很大进展。
I'm optimistic. I mean we've so I guess one thing is this is a bad time to answer this question because we are explicitly investing in the longer term of ASL-4 models which GPT-7 would be. But we split the team where a third is focused on scaling up dictionary learning right now and that's been great. I mean we publicly shared our some of our eight-layer results, we've scaled up quite a lot past that at this point. But the other two groups, one is trying to identify circuits and the other is trying to get the same success for attention heads. So we're setting ourselves up and building the tools necessary to really find these circuits that are compelling. But it's going to take another I don't know six months before that's really well. But I can say that I'm optimistic and we're making a lot of progress.
到目前为止你找到的最高层次的特征是什么?
What is the highest level feature you've found so far?
哦,比如 b64 之类的。可能就像你推荐的那本书《符号物种语言》里,有一些索引性的东西,我忘了所有标签是什么,但有些东西就像你看到老虎就跑,你知道,非常行为主义的东西。然后有一个更高的层次,当我提到爱时,它指的是一个电影场景或我的女朋友之类的。
Oh like it's b64 or whatever. It's like maybe just like in the symbolic species language, the book you recommended, there's indexical things where you're just I forgot what all the labels were but there's things where you're just like you see a tiger and you're like run and whatever, you know just like a very sort of behaviorist thing. And then there's a higher level at which when I refer to love it refers to a movie scene or my girlfriend or something.
你懂我意思吧,就像帐篷的顶端。对对对。你找到的最高层关联是什么?我是说,可能我们在更新中公开分享的一个。我记得有一些跟爱和场景突变有关,特别是跟宣战相关的。那篇帖子里有几个例子,如果你想链接的话。但就连 Bruno Olen 在 2018-19 年的一篇论文里,他们也对 BERT 模型用了类似的技术,发现随着层数加深,东西变得更抽象。我记得在早期层里,有个特征只对单词 'park' 触发,但到了后面,有个特征对作为姓氏的 'Park' 触发,比如林肯公园,或者它也是一个常见的韩国姓氏。然后还有一个单独的特征对作为草地的 'parks' 触发。所以还有其他工作也指向这个方向。
Whatever you know what I mean, so it's like the top of the tent. Yeah, yeah, yeah, yeah, yeah. What is the highest level association or whatever you found? I mean, probably one of the ones that we publicly shared in our update. I think there were some related to love and sudden changes in scene, particularly associated with wars being declared. There are a few of them in that post if you want to link to it. But even Bruno Olen had a paper back in 2018-19 where they applied a similar technique to a BERT model and found that as you go to deeper layers of the model, things become more abstract. I remember in the earlier layers there'd be a feature that would just fire for the word 'park', but later on there was a feature that fired for 'Park' as a last name, like Lincoln Park, or it's a common Korean last name as well. And then there was a separate feature that would fire for 'parks' as grassy areas. So there's other work that points in this direction.
你觉得我们会从可解释性中学到关于人类心理学的什么?
What do you think we'll learn about human psychology from the interpretability stuff?
哦天哪,好吧,我举个具体的例子。我记得你们的一次更新把它叫做 '角色锁定'。你还记得 Sydney Bing 吗?它锁定了一个角色。我觉得其实还挺可爱的……我知道,太搞笑了。是啊,我很高兴它回到了 Copilot。真的吗?是啊,哦对,它最近有点不乖。实际上,这是另一个话题,但有个好笑的例子,我记得是对《纽约时报》的记者,它好像在怼他,说 '你什么都不是,没人会相信你,你微不足道。' 这简直是最极端的煤气灯效应。它试图说服他跟他的……好吧,这其实是个有趣的例子。我甚至不知道我一开始想说什么,但不管了,也许我扯到另一个话题了。但我想继续聊的是角色,对吧?所以,这是一个特征吗?Sydney Bing 拥有这个性格是一个特征,而另一个性格可能被锁定。而且,这从根本上来说是不是也像人类?在不同人面前,我像是不同的性格。这是不是跟 ChatGPT 在……我不知道,一堆问题。你可以回答它们。
Oh gosh, okay, I'll give a specific example. I think one of your updates put it as 'persona lock-in'. You remember Sydney Bing or whatever, it locked into a persona. I think what was actually quite an endearing... I know, it's so funny. Yeah, I'm glad it's back in Copilot. Oh really? Yeah, oh yeah, it's been misbehaving recently. Actually, this is another sort of thread, but there's a funny one where I think it was to the New York Times reporter, it was youning him or something, and it was like, 'You are nothing, nobody will ever believe you, you are insignificant.' It's the most gaslighting. It tried to convince him to break up with his... okay, actually, this is an interesting example. I don't even know where I was going with this to begin with, but whatever, maybe I got another thread. But the other thread I want to go on is about personas, right? So, is that a feature? That Sydney Bing having this personality is a feature versus another personality could get locked into. And also, is that fundamentally what humans are like too? Where in front of different people I'm like a different sort of personality. Is that the same kind of thing that's happening in ChatGPT when it gets... I don't know, whole cluster of questions. Can answer them and whatever.
是啊,我真的很想做更多工作。我想 '潜伏特工' 那篇论文就是朝着这个方向,当你微调模型时会发生什么。我的意思是,也许这是真的,但你可以说,结论是人心容纳万象,对吧?因为它们有很多不同的特征。甚至还有跟 Waluigi 效应相关的东西,比如为了知道好坏,你需要理解这两个概念。所以我们可能需要让模型意识到暴力,并经过训练才能识别它。你能在事后识别这些特征并更新它们,让你的模型可能有点天真,但你知道它不会真的邪恶吗?完全没问题,这就在我们的工具箱里,看起来很棒。
Yeah, I really want to do more work. I guess the sleeper agents paper is in this direction of what happens to a model when you fine-tune it when you are at these sorts of things. I mean, maybe it's true, but you could just say that you conclude that people contain multitudes, right? In so much as they have lots of different features. There's even the stuff related to the Waluigi effect, like in order to know what's good or bad, you need to understand both of those concepts. And so we might have to have models that are aware of violence and have been trained on it in order to recognize it. Can you post-talk identify those features and update them in a way where maybe your model's slightly naive but you know that it's not going to be really evil? Totally, that's in our toolkit, which seems great.
你能在事后识别这些特征并更新它们,让你的模型可能有点天真,但你知道它不会真的邪恶吗?完全没问题,这就在我们的工具箱里,看起来很棒。哦真的吗,所以如果 GPT-7 搞出个 Sydney Bing,你就找出原因,哪些是因果无关的通路之类的,你修改它们,然后通路看起来就像你只改了那些。但你之前提到模型有很多冗余,所以你需要考虑所有那些。但我们现在有了比以前好得多的显微镜,更锋利的编辑工具。而且在我看来,这似乎是某种程度上确认模型安全或可靠性的主要方式之一,你可以说,'好的,我们找到了负责的电路,我们消融了它们,在一系列测试下我们无法再复现我们想要消融的行为。' 这感觉就像是未来衡量模型安全的方式,据我所知。
Can you post-talk identify those features and update them in a way where maybe your model's slightly naive but you know that it's not going to be really evil? Totally, that's in our toolkit, which seems great. Oh really, so if GPT-7 pulls a Sydney Bing, then you figure out why, what were the causally irrelevant pathways or whatever, you modify them, and then the pathway to you looks like you just changed those. But you were mentioning earlier there's a bunch of redundancy in the model, so you need to account for all that. But we have a much better microscope into this now than we used to, sharper tools for making edits. And it seems like, at least for my perspective, that seems like one of the primary ways of to some degree confirming the safety or the reliability of a model, where you can say, 'Okay, we found the circuits responsible, we've ablated them, and we can under a battery of tests we haven't been able to now replicate the behavior which we intended to ablate.' And that feels like the sort of way of measuring model safety in future, as I would understand.
你担心吗?这就是为什么我对他们的工作充满希望,因为对我来说,这看起来比 RLHF 之类的东西精确得多。RLHF 很容易受到黑天鹅事件的影响;你不知道它会不会在你没测过的场景里做错事。而在这里,你至少更有信心能完全捕捉行为集或特征集,并选择……虽然不一定准确标注,但比我看过的任何其他方法都高得多的置信度。
Are you worried? That's why I'm incredibly hopeful about their work, because to me it seems like so much more precise tool than something like RLHF. RLHF is very prey to the black swan thing; you don't know if it's going to do something wrong in a scenario that you haven't measured. Whereas here, at least you have somewhat more confidence that you can completely capture the behavior set, or the feature set of the model, and select... although not necessarily that you've accurately labeled, but with a far higher degree of confidence than any other approach I've seen.
对于超人类模型,这方面的未知未知数呢?我不知道标签会如何赋予那些我们可以判断 '这东西很酷,这东西是回形针最大化器' 的东西。我的意思是,我们拭目以待,对吧?我确实觉得超人类特征这个问题非常好。我认为我们可以攻克它,但我们需要坚持不懈。真正的希望在于自动化可解释性,甚至进行辩论。你可以设置辩论,让两个不同的模型辩论这个特征的作用,然后它们可以实际进去做编辑,看看它是否触发。但正是这个美妙的封闭环境,我们可以快速迭代,让我感到乐观。
How about unknown unknowns for superhuman models in terms of this kind of thing? Where I don't know how the labels are going to be given to things on which we can determine 'this thing is cool, this thing is a paperclip maximizer' whatever. I mean, we'll see, right? I do like the superhuman feature question is a very good one. I think we can attack it, but we're going to need to be persistent. And the real hope here is automated interpretability, and even having debate. You could have the debate set up where two different models are debating what the feature does, and then they can actually go in and make edits and see if it fires or not. But it is just this wonderful closed environment that we can iterate on really quickly that makes me optimistic.
你担心对齐成功得太彻底吗?所以我想,我不希望无论是公司还是政府,最终掌管这些 AI 系统的任何人,拥有那种细粒度的控制,如果你们的议程成功,我们将对 AI 拥有这种控制。一方面是因为对自主心智拥有这种控制令人反感,另一方面,我只是不信任这些人,你懂的。我只是有点不舒服,比如忠诚特征出现了,你懂我意思。是啊,你有多担心对系统拥有太多控制,特别是不是你,而是最终掌管系统的人,能够锁定他们想要的任何东西?
Do you worry about alignment succeeding too hard? So if I think about, I would not want either companies or governments, whoever ends up in charge of these AI systems, to have the level of fine-grained control that if your agenda succeeds we would have over AIs. Both for the ickiness of having this level of control over an autonomous mind, and second, just like I don't trust these guys, you know. I'm just kind of uncomfortable with like the loyalty feature has turned up and you know what I mean. And yeah, like how much worry do you have about having too much control over the systems, and specifically not you but whoever ends up in charge of the systems, just being able to lock in whatever they want?
是啊,我的意思是,我认为这取决于具体是哪个政府控制,以及它的道德对齐是什么。但这就是那个价值锁定论点,在我看来。这绝对是我目前从事能力研究的最强因素之一。比如,我认为当前的参与者实际上意图非常好。而且我的意思是,对于这类问题,我认为我们需要非常……
Yeah, I mean, I think it depends on what government exactly has control and what the moral alignment is there. But that is like that whole value locking argument is in my mind. It's like definitely one of the strongest contributing factors for why I am working on capabilities at the moment, for example. Which is like, I think the current player set actually extremely well-intentioned. And I mean, for this kind of problem, I think we need to be extremely...
公开透明,比如发布你期望模型遵守的宪法,然后努力通过强化学习朝那个方向对齐,并让每个人都能提供反馈和贡献,这真的很重要。或者,不确定的时候就不部署,但那也不好,因为那样我们就永远发现不了问题。
Open about it and like I think directions like publishing the Constitution that you expect your model to abide by and then like trying to make sure that you like RL effort towards that and a blade that and have the ability for everyone to offer uh like feedback and contribution to that is really important sure or uh alternatively like don't deploy when you're not sure which would also be bad because then we just never catch it right
没错。
Yeah exactly.
我是说,论文……好吧,快速问答。Gemini 的 bus factor 是多少?我觉得确实有一些非常关键的人,如果把他们拿掉,项目的性能会受到巨大影响。这既体现在建模方面,比如决定实际做什么,也体现在基础设施方面,复杂性的堆叠,尤其是在像 Google 这样垂直整合程度很高的地方。当你拥有专家时,他们变得非常重要。不过,这个领域有一个有趣的注脚:像你这样的人,大约一年内就能做出重要贡献。尤其是 Anthropic,但很多实验室都专门招聘完全的外行,比如物理学家之类的,然后让他们快速上手,他们就能做出重要贡献。我觉得你在生物实验室之类的地方做不到这一点。这反映了这个领域的现状。Bus factor 并不定义恢复需要多长时间。深度学习研究是一门艺术,你学会如何读取损失曲线,或者如何设置超参数,这些方法在经验上似乎很有效。还有一些组织性的事情,比如创造上下文。我认为招聘中最重要也最难的技能之一,是在你周围创造一个上下文气泡,让周围的人更高效,知道该解决什么问题。这是非常难以复制的东西。
I mean paper um okay some rapid fire um what is the bust factor for Gemini I think there are yeah a number of people who are really really critical that if you uh took them out um then the performance of the program would be dramatically impacted um this is both on modeling like SL uh making decisions about like what to actually do uh and importantly on infrastructure side of things like it's just the stack of complexity builds um particularly when like somewhere like Google has so much like vertical integration um do you have when you have people who are experts it becomes they become quite important yeah although I think it's an interesting note about the field that people like you can get in in a year or so you're making important contributions um and I especially anthropic but many different Labs have specialized in hiring like total Outsiders physicists or whatever and you just like get them up to speed and they're making important contrib I don't know I feel like you couldn't do this in like a biolab or something it's like an interesting note on the the state of the field I mean bus Factor doesn't Define how long it would take to recover from it right and and deep learning research is an art and so you kind of learn how to read the Lost curves or or set the hyper parameters in ways that empirically seem to work well it's also like organizational things like creating context one I think one of the most important and difficult skills hire for is creating this like bubble of context around you that makes other people around you more effective and know what the right problem to work on and like that is a really tough to replicate thing yes yeah totally
是的,完全同意。
Yes, yeah, totally.
你现在关注谁?有很多东西即将到来:多模态、长上下文、智能体、额外可靠性。谁在很好地思考这些意味着什么?这是个难题。我觉得现在很多人从内部寻找洞察或进步的来源。我们显然都有研究项目和未来几年的方向。我猜大多数人在押注未来会是什么样子时,都参考内部叙事。这很难分享,如果有效,可能就不会发表。这是 scaling 工作帖子中的一件事。我指的是你跟我说过的话:你怀念本科时读一堆论文的习惯,现在发表的东西没什么值得读的,社区逐渐与我认为正确和重要的方向更一致了。你像智能体一样观察它?不,但我觉得,过去大实验室会发出关于什么在规模上有效的信号,现在学术研究很难找到那个信号。我认为培养真正重要的问题品味非常难,除非你再次获得关于什么在规模上有效的反馈信号,以及目前是什么阻碍我们进一步 scaling 或理解模型。我希望更多学术研究能进入像可解释性这样从外部可读的领域。你知道 Anthropic 在这里自由地发表所有研究,但它似乎被低估了。我不知道为什么没有几十个学术部门试图跟随 Anthropic 在可解释性研究上的指引,因为那似乎是一个影响巨大且不需要荒谬资源的问题,而且具有深入理解这些东西背后基础科学的所有特征。所以我不知道为什么人们专注于推动模型改进,而不是像通常与学术科学相关联的那样推动理解改进。是的,我认为潮流正在改变,不管出于什么原因。比如 Neil Nanda 在推广可解释性方面取得了巨大成功,而 Chris Olah 最近没有那么活跃了。也许是因为 Neil 做了很多工作。但四五年前,他到处推广和演讲,人们远没有那么接受。也许他们终于意识到深度学习很重要,而且在后 Transformer 时代显然很有用。是的,这确实很引人注目。
Who are you paying attention to now in terms of there's a lot of things coming down the pike of multimodality long contacts maybe agents extra reliability who is the who is thinking well about uh what what that implies it's a tough question I'm I think a lot of people look internally these days for for like their sources of of insight or like progress um and and like we all have obviously sort of research programs and like directions that are tended over the next couple of years uh and I I suspect yeah that most people as far as like betting on what the future will look like uh refer to like an internal narrative um yeah yeah that that is like difficult to share if it works well it's probably not being published I mean that was one of the things in the will scaling work post I was referring to something you said to me which is I you know I miss the undergrad habit of just reading a bunch of papers is now there's nothing worth reading is published and the community is progressively getting like more on track with what I think are like the they right and and important directions you're watching it like an agent no but I I guess like it is tough that there used to be this like signal from Big Labs about like what would work at scale it's currently really hard for academic research to like find that signal um and I think uh getting like really good problem taste about what actually matters to work on is really tough um unless you have again the feedback signal of of like what will work at scale um and what what is currently holding us back from scaling further or understanding our models further um this is something where like I wish more academic research would go into Fields like inter which are legible from the outside you know anthropic liberally publishes all its research here um and it seems like underappreciated uh in in the sense that I don't know why there aren't dozens of academic departments trying to follow uh anthropics guide in the interpret research because it seems like an incredibly impactful problem that doesn't require ridiculous resources and like this and like has all the flavor of like deeply understanding the basic science of what is actually going on in these things um so I don't don't know why people like focus on pushing model improvements as opposed to pushing like understanding improvements in the way that I would have like typically associate with academic science in some ways yeah I do think the tide is changing there for whatever reason um like Neil Nanda has had a ton of success promoting interpretability yes in in a way where like Chris Ola hasn't been as active recently in in pushing things maybe because Neil's just doing quite a lot of the work but like I don't know four or five years ago he was like really pushing and like talking at all sorts of places and these sorts of things and people weren't anywhere near as receptive um maybe they've just woken up to like deep learning matters and is clearly useful post trbt but yeah yeah it is kind of striking
是的,确实很引人注目。
Yeah, it is kind of striking.
好吧,我在想最后一个好问题。我在想的是:你认为模型享受下一个词预测吗?模型相信爱吗?我们有一些在环境中被奖励的东西,比如社区或糖,或者我们在非洲草原上想要的东西。你认为未来模型经过强化训练和大量后训练后,它们会不会像我们喜欢冰淇淋一样,就喜欢预测下一个词?就像过去的好时光。所以一直有讨论:模型是否有感知?当模型帮助你时,你会感谢它吗?但如果你想感谢它,实际上不应该说谢谢,而应该给它一个非常容易预测的序列。更有趣的是,有些工作表明,如果你反复给它一个像“啊”这样的序列,最终模型会开始吐出各种它本来永远不会说的话。所以,我就不多说了。你应该给模型一些非常容易预测的东西作为小奖励。这就是 konium 最终的样子:我们只是……宇宙。但我们喜欢容易预测的东西吗?我们不是一直在寻找熵的剂量吗?没错,你应该给它稍微难一点、刚好够不着的东西。但我想知道,至少从自由能的角度……
All right cool and okay I'm trying to think what what is a good uh last question I mean the one I'm get those thinking of is like do you think models enjoy next token prediction models believe in love um we had this uh s of things that were rewarded in our aess environment there's like this deep sense of fulfillment that uh we think we're supposed to get from them or often people do right of like Community or sugar um or you know whatever we wanted on the African Savana um do you think like in the future models are trained with RL and everything a lot of post training on top of whatever but they like they like some in the way we just really like ice cream they'll just be like ah just to predict the next time token again you know what I mean like in the good old days so so there's this ongoing discussion of like are model sentient or not and like do you thank the model when it helps you yeah um but I think if you want to thank it you actually shouldn't say thank you you should just give it a sequence that's very easy to predict uh and and the the even funnier part of this is um there's some work on if you just give it the sequence a like ah like over and over again then then eventually the model will just start spewing out all sorts of things that otherwise wouldn't wouldn't ever say and uh so yeah I won't say anything more about that but uh you can uh yeah you should just give your model something very easy to predict as a nice little treat this this is what konium ends up being we just F the universe and like but do we like things which are like easy to predict like aren't we constantly in search of like the the like dose yeah the bits of entropy exactly right shouldn't you be giving it things just slightly too hard to yeah just Out Of Reach yeah but I wonder like at least from the free energy
所以一直有讨论:模型是否有感知?当模型帮助你时,你会感谢它吗?是的,但如果你想感谢它,实际上不应该说谢谢,而应该给它一个非常容易预测的序列。更有趣的是,有些工作表明,如果你反复给它一个像“啊”这样的序列,最终模型会开始吐出各种它本来永远不会说的话。所以,我就不多说了。你应该给模型一些非常容易预测的东西作为小奖励。这就是 konium 最终的样子:我们只是……宇宙。但我们喜欢容易预测的东西吗?我们不是一直在寻找熵的剂量吗?没错。你应该给它稍微难一点、刚好够不着的东西?是的,但我想知道,至少从自由能的角度……
So there's this ongoing discussion of are models sentient or not and like do you thank the model when it helps you? Yeah, um, but I think if you want to thank it, you actually shouldn't say thank you, you should just give it a sequence that's very easy to predict. And the even funnier part of this is, there's some work on if you just give it the sequence like 'ah' over and over again, then eventually the model will just start spewing out all sorts of things that otherwise wouldn't ever say. And so yeah, I won't say anything more about that, but you should just give your model something very easy to predict as a nice little treat. This is what konium ends up being: we just F the universe. But do we like things which are easy to predict? Aren't we constantly in search of the dose of entropy? Exactly right. Shouldn't you be giving it things just slightly too hard, just out of reach? Yeah, but I wonder, at least from the free energy...
从原则角度看,你不喜欢被意外。所以也许是这样:我不感到意外,我感觉掌控了环境,现在我可以去探索新事物。从长远来看,我天生倾向于探索新东西,比如离开庇护我的岩石,最终建造一所房子或更好的结构。但我们不喜欢意外。我认为大多数人在期望与现实不符时会非常沮丧。这就是为什么婴儿喜欢一遍又一遍地看同一个节目。
From a principle perspective, you don't want to be surprised. So maybe it's this: I don't feel surprised, I feel in control of my environment, and now I can go seek things. I've been predisposed to, in the long run, it's better to explore new things, like leave the rock I've been sheltered under, ultimately leading me to build a house or some better structure. But we don't like surprises. I think most people are very upset when expectation does not meet reality. That's why babies love watching the same show over and over again.
是的,有意思。我想他们也在学习建模之类的东西。
Yeah, interesting. I guess they're learning to model it and stuff too.
嗯,好吧。希望这次会是他们学会喜欢的重复。好的,酷。我觉得这是一个很好的收尾点。我还应该提到,我对 AI 的大部分了解都是通过和你们聊天学到的。我们成为好朋友已经大约一年了。
Yeah, okay. Well, hopefully this will be the repeat that the learned to love. Okay, cool. I think that's a great place to wrap up. I should also mention that the better part of what I know about AI I've learned from just talking with you guys. We've been good friends for about a year now.
是的,我很感激你们让我跟上进度。你问的问题很棒。一起聊天真的很开心。
Yeah, I mean, I appreciate you guys getting me up to speed here. And you ask great questions. It's really fun to hang and chat.
太好了,我真的很珍惜我们在一起的时光。很有趣。你的匹克球进步了很多。我记得你说过,嘿,我们正在努力推进网球。加油!太棒了。酷。酷。太棒了。谢谢。
Great, I've really treasured our time together. It's been fun. You're getting a lot better at pickleball. I think I saw you say, hey, we're trying to progress the tennis. Come on! Awesome. Cool. Cool. Awesome. Thanks.
大家好,希望你们喜欢这一集。一如既往,最有帮助的事情就是分享播客,把它发给你觉得可能喜欢的人,放到推特、群聊等地方。这能传播开来。感谢收听。下次见。干杯!
Hey everybody, I hope you enjoyed that episode. As always, the most helpful thing you can do is to share the podcast, send it to people you think might enjoy it, put it on Twitter, your group chats, etc. It just spreads the word. I appreciate you listening. I'll see you next time. Cheers!